LLM Inference Benchmark Explorer

Muse-Glimmer-30B on RTX PRO 6000 — BF16, vLLM, TP1, speculative decoding k=15 inference benchmark

Open this configuration in the LLM Inference Benchmark Explorer

ModelMuse-Glimmer-30B
Parameters30B
Intelligence Index17.5
Agentic Index8.5
DeviceRTX PRO 6000
QuantizationBF16
Max C32
Chat Capacity128
Agentic Capacity47
TP—
DP—
PP—
EnginevLLM
Speculative DecodingYes
CTTFT (ms)TPS (tok/s)Status
1118165.38PASS
2163155.93PASS
4178138.39PASS
8215121.33PASS
1632883.91PASS
3258046.20PASS
64330123.13FAIL

In OpenZeka's measurement, Muse-Glimmer-30B (30B parameters), served in BF16 format with vLLM and speculative decoding on RTX PRO 6000, reached a generation speed of 165.4 tok/s per request and a time to first token (TTFT) of 118 ms with a single request.

Considering the speed targets and the available KV cache capacity, the estimated capacity is 128 users for chat use and 47 for agentic use. The realistic capacity will likely fall between these two values. In scenarios dominated by coding, tool use, long workflows and multi-agent use, capacity approaches the agentic estimate; where shorter interactions, standard conversations and lighter tasks dominate, it approaches the chat estimate.

Notes: DFlash, native MTP k=15

Intelligence Index and Agentic Index values are published by Artificial Analysis and are reproduced here with attribution.