LLM Inference Benchmark Explorer

Qwen3.8-27B on 1× DGX Spark — FP8, vLLM, TP1, speculative decoding inference benchmark

Open this configuration in the LLM Inference Benchmark Explorer

ModelQwen3.8-27B
Parameters27B
Intelligence Index33.7
Agentic Index45.8
Device1× DGX Spark
QuantizationFP8
Max C0
Chat Capacity0
Agentic Capacity0
TP—
DP—
PP—
EnginevLLM
Speculative DecodingYes
CTTFT (ms)TPS (tok/s)Status
133018.45FAIL

In OpenZeka's measurement, Qwen3.8-27B (27B parameters), served in FP8 format with vLLM and speculative decoding on 1× DGX Spark, reached a generation speed of 18.4 tok/s per request and a time to first token (TTFT) of 330 ms with a single request.

At the default speed targets (at least 20 tok/s per request, TTFT at most 1,000 ms) there is no capacity estimate for this configuration: no measured load level met them.

Notes: eugr nightly. Single C=1 test (config matrix)

Intelligence Index and Agentic Index values are published by Artificial Analysis and are reproduced here with attribution.