LLM Inference Benchmark Explorer

Gemma-4-31B-it on Thor — NVFP4, vLLM, TP1, speculative decoding k=3 inference benchmark

Open this configuration in the LLM Inference Benchmark Explorer

ModelGemma-4-31B-it
Parameters31B
Intelligence Index14.7
Agentic Index4.2
DeviceThor
QuantizationNVFP4
Max C0
Chat Capacity0
Agentic Capacity0
TP—
DP—
PP—
EnginevLLM
Speculative DecodingYes
CTTFT (ms)TPS (tok/s)Status
133616.16FAIL
247815.38FAIL
449114.59FAIL
850714.72FAIL

In OpenZeka's measurement, Gemma-4-31B-it (31B parameters), served in NVFP4 format with vLLM and speculative decoding on Thor, reached a generation speed of 16.2 tok/s per request and a time to first token (TTFT) of 336 ms with a single request.

At the default speed targets (at least 20 tok/s per request, TTFT at most 1,000 ms) there is no capacity estimate for this configuration: no measured load level met them.

Notes: Speculative MTP k=3. Served as gemma-4-31b-it. Model: nvidia/gemma-4-31B-it-NVFP4. 256K context

Intelligence Index and Agentic Index values are published by Artificial Analysis and are reproduced here with attribution.