LLM Inference Benchmark Explorer

Laguna-S-2.1 on Thor — NVFP4, vLLM, TP1 inference benchmark

Open this configuration in the LLM Inference Benchmark Explorer

ModelLaguna-S-2.1
Parameters118B
Intelligence Index—
Agentic Index—
DeviceThor
QuantizationNVFP4
Max C0
Chat Capacity0
Agentic Capacity0
TP—
DP—
PP—
EnginevLLM
Speculative Decoding—
CTTFT (ms)TPS (tok/s)Status
157812.92FAIL
267411.10FAIL
48478.69FAIL
810326.48FAIL

In OpenZeka's measurement, Laguna-S-2.1 (118B parameters), served in NVFP4 format with vLLM on Thor, reached a generation speed of 12.9 tok/s per request and a time to first token (TTFT) of 578 ms with a single request.

At the default speed targets (at least 20 tok/s per request, TTFT at most 1,000 ms) there is no capacity estimate for this configuration: no measured load level met them.

Notes: Jetson Thor. max_model_len 262144

Intelligence Index and Agentic Index values are published by Artificial Analysis and are reproduced here with attribution.