LLM Inference Benchmark Explorer

Tencent-Hy3 on DGX B300 — BF16, vLLM, TP4 inference benchmark

Open this configuration in the LLM Inference Benchmark Explorer

ModelTencent-Hy3
Parameters299B
Intelligence Index25.3
Agentic Index24.1
DeviceDGX B300
QuantizationBF16
Max C64+
Chat Capacity144
Agentic Capacity36
TP4
DP—
PP—
EnginevLLM
Speculative Decoding—
CTTFT (ms)TPS (tok/s)Status
137149.61PASS
24394.29PASS
451103.12PASS
86482.08PASS
1610363.09PASS
3210048.42PASS
6410640.43PASS

In OpenZeka's measurement, Tencent-Hy3 (299B parameters), served in BF16 format with vLLM on DGX B300 (TP=4), reached a generation speed of 149.6 tok/s per request and a time to first token (TTFT) of 37 ms with a single request.

Considering the speed targets and the available KV cache capacity, the estimated capacity is 144 users for chat use and 36 for agentic use. The realistic capacity will likely fall between these two values. In scenarios dominated by coding, tool use, long workflows and multi-agent use, capacity approaches the agentic estimate; where shorter interactions, standard conversations and lighter tasks dominate, it approaches the chat estimate.

Intelligence Index and Agentic Index values are published by Artificial Analysis and are reproduced here with attribution.