LLM Inference Benchmark Explorer

Gemma-4-31B-it on 2× DGX Spark — NVFP4, vLLM, TP2 inference benchmark

Open this configuration in the LLM Inference Benchmark Explorer

ModelGemma-4-31B-it
Parameters31B
Intelligence Index19
Agentic Index4.2
Device2× DGX Spark
QuantizationNVFP4
Max C0
Chat Capacity0
Agentic Capacity0
TP2
DP—
PP—
EnginevLLM
Speculative Decoding—
CTTFT (ms)TPS (tok/s)Status
113619.42FAIL
222719.59FAIL
424619.17FAIL
828018.20FAIL
1629315.81FAIL
3233713.21FAIL
644209.16FAIL

In OpenZeka's measurement, Gemma-4-31B-it (31B parameters), served in NVFP4 format with vLLM on 2× DGX Spark (TP=2), reached a generation speed of 19.4 tok/s per request and a time to first token (TTFT) of 136 ms with a single request.

At the default speed targets (at least 20 tok/s per request, TTFT at most 1,000 ms) there is no capacity estimate for this configuration: no measured load level met them.

Intelligence Index and Agentic Index values are published by Artificial Analysis and are reproduced here with attribution.