LLM Inference Benchmark Explorer

Qwen3.6-27B on 1× DGX Spark — FP8, vLLM, TP1, speculative decoding inference benchmark

Open this configuration in the LLM Inference Benchmark Explorer

ModelQwen3.6-27B
Parameters27B
Intelligence Index21.4
Agentic Index18.5
Device1× DGX Spark
QuantizationFP8
Max C0
Chat Capacity0
Agentic Capacity0
TP—
DP—
PP—
EnginevLLM
Speculative DecodingYes
CTTFT (ms)TPS (tok/s)Status
132319.50FAIL
251619.83FAIL
453118.56FAIL
862116.62FAIL
1681313.84FAIL

In OpenZeka's measurement, Qwen3.6-27B (27B parameters), served in FP8 format with vLLM and speculative decoding on 1× DGX Spark, reached a generation speed of 19.5 tok/s per request and a time to first token (TTFT) of 323 ms with a single request.

At the default speed targets (at least 20 tok/s per request, TTFT at most 1,000 ms) there is no capacity estimate for this configuration: no measured load level met them.

Notes: MTP gives 2.4x TPS improvement

Intelligence Index and Agentic Index values are published by Artificial Analysis and are reproduced here with attribution.