LLM Inference Benchmark Explorer

Qwen3.5-397B-A17B on 3× DGX Spark — INT4, vLLM, TP1, PP3 inference benchmark

Open this configuration in the LLM Inference Benchmark Explorer

ModelQwen3.5-397B-A17B
Parameters397B
Intelligence Index21.4
Agentic Index—
Device3× DGX Spark
QuantizationINT4
Max C0
Chat Capacity0
Agentic Capacity0
TP—
DP—
PP3
EnginevLLM
Speculative Decoding—
CTTFT (ms)TPS (tok/s)Status
150017.05FAIL
260614.47FAIL
478510.44FAIL
811016.99FAIL

In OpenZeka's measurement, Qwen3.5-397B-A17B (397B parameters), served in INT4 format with vLLM on 3× DGX Spark (PP=3), reached a generation speed of 17 tok/s per request and a time to first token (TTFT) of 500 ms with a single request.

At the default speed targets (at least 20 tok/s per request, TTFT at most 1,000 ms) there is no capacity estimate for this configuration: no measured load level met them.

Notes: Ring topology, PP=3

Intelligence Index and Agentic Index values are published by Artificial Analysis and are reproduced here with attribution.