LLM Inference Benchmark Explorer

Qwen3.8-Flash-Next on 1× DGX Spark — NVFP4, SGLang, TP1, speculative decoding k=3 inference benchmark

Open this configuration in the LLM Inference Benchmark Explorer

ModelQwen3.8-Flash-Next
Parameters176B
Intelligence Index39.8
Agentic Index53.6
Device1× DGX Spark
QuantizationNVFP4
Max C2
Chat Capacity8
Agentic Capacity3
TP—
DP—
PP—
EngineSGLang
Speculative DecodingYes
CTTFT (ms)TPS (tok/s)Status
130228.50PASS
239323.61PASS
456617.13FAIL
876311.80FAIL

In OpenZeka's measurement, Qwen3.8-Flash-Next (176B parameters), served in NVFP4 format with SGLang and speculative decoding on 1× DGX Spark, reached a generation speed of 28.5 tok/s per request and a time to first token (TTFT) of 302 ms with a single request.

Considering the speed targets alone (the KV cache limit was not calculated for this configuration), the estimated capacity is 8 users for chat use and 3 for agentic use. The realistic capacity will likely fall between these two values. In scenarios dominated by coding, tool use, long workflows and multi-agent use, capacity approaches the agentic estimate; where shorter interactions, standard conversations and lighter tasks dominate, it approaches the chat estimate.

Notes: NEXTN MTP k=3. PLE N-gram table (47.7 GiB) file-backed on NVMe. mem-frac 0.85, max 8 running.

Intelligence Index and Agentic Index values are published by Artificial Analysis and are reproduced here with attribution.