LLM Inference Benchmark Explorer

Gemma-4-31B-it on 1× DGX Spark — NVFP4, vLLM, TP1 inference benchmark

Open this configuration in the LLM Inference Benchmark Explorer

ModelGemma-4-31B-it
Parameters31B
Intelligence Index19
Agentic Index4.2
Device1× DGX Spark
QuantizationNVFP4
Max C0
Chat Capacity0
Agentic Capacity0
TP—
DP—
PP—
EnginevLLM
Speculative Decoding—
CTTFT (ms)TPS (tok/s)Status
120210.81FAIL
228311.06FAIL
429410.89FAIL
831610.51FAIL
1634710.04FAIL
323998.87FAIL
645327.36FAIL

In OpenZeka's measurement, Gemma-4-31B-it (31B parameters), served in NVFP4 format with vLLM on 1× DGX Spark, reached a generation speed of 10.8 tok/s per request and a time to first token (TTFT) of 202 ms with a single request.

At the default speed targets (at least 20 tok/s per request, TTFT at most 1,000 ms) there is no capacity estimate for this configuration: no measured load level met them.

Notes: Model: RedHatAI/gemma-4-31B-it-NVFP4 (31B dense)

Intelligence Index and Agentic Index values are published by Artificial Analysis and are reproduced here with attribution.