LLM Inference Benchmark Explorer

DeepSeek-V4-Pro on DGX B300 — FP4, vLLM, TP4 inference benchmark

Open this configuration in the LLM Inference Benchmark Explorer

ModelDeepSeek-V4-Pro
Parameters1.60T
Intelligence Index30.4
Agentic Index26.3
DeviceDGX B300
QuantizationFP4
Max C32
Chat Capacity128
Agentic Capacity48
TP4
DP—
PP—
EnginevLLM
Speculative Decoding—
CTTFT (ms)TPS (tok/s)Status
130386.84PASS
232576.84PASS
442151.41PASS
842940.94PASS
1685925.56PASS
3265520.22PASS
64105614.14FAIL

In OpenZeka's measurement, DeepSeek-V4-Pro (1.60T parameters), served in FP4 format with vLLM on DGX B300 (TP=4), reached a generation speed of 86.8 tok/s per request and a time to first token (TTFT) of 303 ms with a single request.

Considering the speed targets and the available KV cache capacity, the estimated capacity is 128 users for chat use and 48 for agentic use. The realistic capacity will likely fall between these two values. In scenarios dominated by coding, tool use, long workflows and multi-agent use, capacity approaches the agentic estimate; where shorter interactions, standard conversations and lighter tasks dominate, it approaches the chat estimate.

Intelligence Index and Agentic Index values are published by Artificial Analysis and are reproduced here with attribution.