LLM Inference Benchmark Explorer

Qwen3.8-27B on RTX PRO 6000 — NVFP4, vLLM, TP1, speculative decoding k=3 inference benchmark

Open this configuration in the LLM Inference Benchmark Explorer

ModelQwen3.8-27B
Parameters27B
Intelligence Index33.7
Agentic Index45.8
DeviceRTX PRO 6000
QuantizationNVFP4
Max C64
Chat Capacity71
Agentic Capacity24
TP—
DP—
PP—
EnginevLLM
Speculative DecodingYes
CTTFT (ms)TPS (tok/s)Status
172128.33PASS
2102128.21PASS
4114115.32PASS
817594.92PASS
1619678.35PASS
3230652.64PASS
6457029.34PASS
128459614.92FAIL

In OpenZeka's measurement, Qwen3.8-27B (27B parameters), served in NVFP4 format with vLLM and speculative decoding on RTX PRO 6000, reached a generation speed of 128.3 tok/s per request and a time to first token (TTFT) of 72 ms with a single request.

Considering the speed targets and the available KV cache capacity, the estimated capacity is 71 users for chat use and 24 for agentic use. The realistic capacity will likely fall between these two values. In scenarios dominated by coding, tool use, long workflows and multi-agent use, capacity approaches the agentic estimate; where shorter interactions, standard conversations and lighter tasks dominate, it approaches the chat estimate.

Notes: Speculative MTP k=3.

Intelligence Index and Agentic Index values are published by Artificial Analysis and are reproduced here with attribution.