LLM Inference Benchmark Explorer

Qwen3.6-35B-A3B on DGX B300 — BF16, vLLM, TP1 inference benchmark

Open this configuration in the LLM Inference Benchmark Explorer

ModelQwen3.6-35B-A3B
Parameters35B
Intelligence Index18.2
Agentic Index13.1
DeviceDGX B300
QuantizationBF16
Max C64+
Chat Capacity256+
Agentic Capacity96+
TP—
DP—
PP—
EnginevLLM
Speculative Decoding—
CTTFT (ms)TPS (tok/s)Status
154241.37PASS
2110191.91PASS
4152168.21PASS
81042128.78FAIL
16201111.84PASS
32222103.20PASS
6427280.42PASS

In OpenZeka's measurement, Qwen3.6-35B-A3B (35B parameters), served in BF16 format with vLLM on DGX B300, reached a generation speed of 241.4 tok/s per request and a time to first token (TTFT) of 54 ms with a single request.

Considering the speed targets and the available KV cache capacity, the estimated capacity is at least 256 users for chat use and at least 96 for agentic use. The realistic capacity will likely fall between these two values. In scenarios dominated by coding, tool use, long workflows and multi-agent use, capacity approaches the agentic estimate; where shorter interactions, standard conversations and lighter tasks dominate, it approaches the chat estimate.

“At least” means the configuration met the speed targets even at the highest load tested, so the real capacity may be higher.

Intelligence Index and Agentic Index values are published by Artificial Analysis and are reproduced here with attribution.