LLM Inference Benchmark Explorer

Qwen3.8-27B on 1× DGX Spark — FP8, vLLM, TP1 inference benchmark

Open this configuration in the LLM Inference Benchmark Explorer

ModelQwen3.8-27B
Parameters27B
Intelligence Index33.7
Agentic Index45.8
Device1× DGX Spark
QuantizationFP8
Max C0
Chat Capacity0
Agentic Capacity0
TP—
DP—
PP—
EnginevLLM
Speculative Decoding—
CTTFT (ms)TPS (tok/s)Status
11727.90FAIL

In OpenZeka's measurement, Qwen3.8-27B (27B parameters), served in FP8 format with vLLM on 1× DGX Spark, reached a generation speed of 7.9 tok/s per request and a time to first token (TTFT) of 172 ms with a single request.

At the default speed targets (at least 20 tok/s per request, TTFT at most 1,000 ms) there is no capacity estimate for this configuration: no measured load level met them.

Notes: Single C=1 test (config matrix)

Intelligence Index and Agentic Index values are published by Artificial Analysis and are reproduced here with attribution.