LLM Inference Benchmark Explorer

MiniMax-M2.7 on 3× DGX Spark — NVFP4, vLLM, TP1, PP3 inference benchmark

Open this configuration in the LLM Inference Benchmark Explorer

ModelMiniMax-M2.7
Parameters229B
Intelligence Index22.8
Agentic Index15.3
Device3× DGX Spark
QuantizationNVFP4
Max C0
Chat Capacity0
Agentic Capacity0
TP—
DP—
PP3
EnginevLLM
Speculative Decoding—
CTTFT (ms)TPS (tok/s)Status
121118.04FAIL
226215.32FAIL
447810.82FAIL
84327.77FAIL

In OpenZeka's measurement, MiniMax-M2.7 (229B parameters), served in NVFP4 format with vLLM on 3× DGX Spark (PP=3), reached a generation speed of 18 tok/s per request and a time to first token (TTFT) of 211 ms with a single request.

At the default speed targets (at least 20 tok/s per request, TTFT at most 1,000 ms) there is no capacity estimate for this configuration: no measured load level met them.

Notes: Ring topology, PP=3

Intelligence Index and Agentic Index values are published by Artificial Analysis and are reproduced here with attribution.