LLM Inference Benchmark Explorer

GPT-OSS 120B on RTX PRO 6000 — MXFP4, vLLM, TP1 inference benchmark

Open this configuration in the LLM Inference Benchmark Explorer

ModelGPT-OSS 120B
Parameters120B
Intelligence Index11.6
Agentic Index3.7
DeviceRTX PRO 6000
QuantizationMXFP4
Max C64+
Chat Capacity55
Agentic Capacity13
TP—
DP—
PP—
EnginevLLM
Speculative Decoding—
CTTFT (ms)TPS (tok/s)Status
152169.64PASS
259125.81PASS
46498.77PASS
88076.57PASS
1612255.27PASS
3216041.88PASS
6420631.76PASS

In OpenZeka's measurement, GPT-OSS 120B (120B parameters), served in MXFP4 format with vLLM on RTX PRO 6000, reached a generation speed of 169.6 tok/s per request and a time to first token (TTFT) of 52 ms with a single request.

Considering the speed targets and the available KV cache capacity, the estimated capacity is 55 users for chat use and 13 for agentic use. The realistic capacity will likely fall between these two values. In scenarios dominated by coding, tool use, long workflows and multi-agent use, capacity approaches the agentic estimate; where shorter interactions, standard conversations and lighter tasks dominate, it approaches the chat estimate.

Notes: 4K context

Intelligence Index and Agentic Index values are published by Artificial Analysis and are reproduced here with attribution.