LLM Inference Benchmark Explorer

MiniMax-M2.7 on 2× DGX Spark — NVFP4, vLLM, TP1, PP2 inference benchmark

Open this configuration in the LLM Inference Benchmark Explorer

ModelMiniMax-M2.7
Parameters229B
Intelligence Index22.8
Agentic Index15.3
Device2× DGX Spark
QuantizationNVFP4
Max C0
Chat Capacity0
Agentic Capacity0
TP—
DP—
PP2
EnginevLLM
Speculative Decoding—
CTTFT (ms)TPS (tok/s)Status
139617.20FAIL
237114.65FAIL
450110.73FAIL
86707.52FAIL

In OpenZeka's measurement, MiniMax-M2.7 (229B parameters), served in NVFP4 format with vLLM on 2× DGX Spark (PP=2), reached a generation speed of 17.2 tok/s per request and a time to first token (TTFT) of 396 ms with a single request.

At the default speed targets (at least 20 tok/s per request, TTFT at most 1,000 ms) there is no capacity estimate for this configuration: no measured load level met them.

Notes: PP=2 (pipeline parallel)

Intelligence Index and Agentic Index values are published by Artificial Analysis and are reproduced here with attribution.