LLM Inference Benchmark Explorer

MiMo-V2.6-Pro-RL on 8× DGX Spark — FP4, SGLang, TP8, speculative decoding k=7 inference benchmark

Open this configuration in the LLM Inference Benchmark Explorer

ModelMiMo-V2.6-Pro-RL
Parameters1.02T
Intelligence Index46.3
Agentic Index—
Device8× DGX Spark
QuantizationFP4
Max C2
Chat Capacity8
Agentic Capacity3
TP8
DP—
PP—
EngineSGLang
Speculative DecodingYes
CTTFT (ms)TPS (tok/s)Status
137238.83PASS
246126.35PASS
466017.52FAIL
852427.61FAIL

In OpenZeka's measurement, MiMo-V2.6-Pro-RL (1.02T parameters), served in FP4 format with SGLang and speculative decoding on 8× DGX Spark (TP=8), reached a generation speed of 38.8 tok/s per request and a time to first token (TTFT) of 372 ms with a single request.

Considering the speed targets and the available KV cache capacity, the estimated capacity is 8 users for chat use and 3 for agentic use. The realistic capacity will likely fall between these two values. In scenarios dominated by coding, tool use, long workflows and multi-agent use, capacity approaches the agentic estimate; where shorter interactions, standard conversations and lighter tasks dominate, it approaches the chat estimate.

Notes: DFlash speculative decoding k=7, 256K context (max_model_len 262144), SGLang. Model: XiaomiMiMo/MiMo-V2.6-Pro-RL (served locally as mimo-v2.6-pro)

Intelligence Index and Agentic Index values are published by Artificial Analysis and are reproduced here with attribution.