VLM Inference Benchmark Explorer
How many cameras can this device watch with this vision-language model? Pick the image size your cameras send and set how long a camera may wait for its answer; the table updates for every configuration we have measured. Expand any row to see its measurements for every number of cameras, the results explained and a chart.
How to use the benchmark explorer
What a row is
A row is one deployment configuration: a vision-language model, on specific hardware, in a specific number format, served by a specific engine, measured end to end. The same model therefore appears in several rows, and two rows are directly comparable only once you know which of those settings differ between them.
Filters and the target do different jobs
Filters decide which rows you see. The device, model, parameter-count, quantization and engine filters, and the response-time and camera sliders, only show or hide rows.
The target and the assumption decide what the numbers say. They sit under Performance Target and Assumptions. Changing one leaves the rows in place but recalculates the response time, Max Cameras and the green and red colouring of every row.
1. Describe your cameras
Two selectors describe the workload, and every number in the table is read at them:
- Image Size — the size of each image a camera sends: 480p, 720p, 1080p or 2K.
- Cameras — how many cameras send requests at the same moment. This is the concurrency the system was measured at.
Rows not measured at the selected combination are hidden, and dimmed buttons mark combinations that have no measurement. Several cameras were measured at 480p, 720p and 1080p; 2K with one camera only.
2. Narrow and sort
Filters combine, and every column sorts. The most instructive comparisons change a single setting: one model at two quantizations on one device, the same model under llama.cpp and vLLM, or one model on two devices.
3. Set your target
The target is the longest a camera may wait for its answer to start — default 3 seconds. A response time exactly on the target passes.
Beside it, Images per Camera sets how many images each request carries: one snapshot by default, or several frames of the same camera sent together, as video is often sent to a vision-language model. Several images per request were measured with one camera only, so with three or five images Max Cameras cannot be higher than 1.
4. Read Max Cameras
Max Cameras is the highest measured number of cameras at which every camera’s answer starts within your target. A plus — 16+ — means the configuration met the target even at the highest number it was tested with, so its real maximum was not reached. 0 means not even a single camera is answered in time.
5. Open a row
Click a row to see:
- on the left, the measurements at your image size and image count for every number of cameras tested, marked PASS or FAIL against your target;
- on the right, the results explained in plain language, at the default settings;
- below, a chart of response time as cameras are added, one line per image size, with your target as a dashed line. It appears where several cameras were measured, so with one image per camera.
The chart downloads as a PNG for reports and presentations.
Where to start
“How many cameras can one Jetson AGX Orin watch?” Select the device, set the image size your cameras send, and read Max Cameras. Set the minimum cameras slider to the number you need to see only the configurations that keep up.
“Is 1080p worth it?” Switch the image size and watch Max Cameras, or open a row: its chart has one line per image size.
“What do quantization and the engine change?” Hold the model and the device fixed and compare the rows that differ in that one setting: Qwen3-VL-8B-Instruct at Q8_0 and Q4_K_M on Jetson AGX Orin, or Qwen3-VL-4B-Instruct under llama.cpp and vLLM on RTX PRO 6000. Changing the number of cameras shows how the gap develops under load.
“We already own this hardware; what can we run on it?” Start with the device filter, sort the remaining models by size or by Max Cameras, and narrow with the response time or the number of cameras you need.
What the numbers mean and how they are calculated
How the numbers were measured
Each request carries one or more photos and the prompt Describe the scene., and asks for up to 128 tokens. The photos are sixteen fixed pictures, cycled through the requests and scaled to 854×480 (480p), 1280×720 (720p), 1920×1080 (1080p) and 2560×1440 (2K), sent as base64 images through the OpenAI-compatible chat API. Hugging Face checkpoints were served with vLLM and GGUF files with llama.cpp.
Every combination of image size, images per request and number of cameras is measured on its own, after a warm-up batch that is discarded. The figures are means of the requests that completed; a request that failed during the benchmark is left out of the mean and does not count against a configuration.
- Several cameras were measured with one image per request at 480p, 720p and 1080p: 1, 2 and 4 cameras on Jetson Orin NX (up to 8 for one configuration), up to 8 on Jetson AGX Orin and up to 16 on RTX PRO 6000, with 8 requests per level on a Jetson and 24 on RTX PRO 6000, never more cameras’ requests in flight than the level.
- One camera was measured at all four sizes with one, three and five images per request, 5 requests each (3 for one configuration). 2K and several images per request were measured this way only.
A token is the unit a model reads and writes, about three quarters of an English word. An image is read as tokens too, and a larger image becomes more of them.
Response time
Response time is how long a camera waits until its answer starts — the time to first token (TTFT), in seconds. For a vision-language model this is mostly reading the images, so it grows with their size and number, and with the number of cameras sharing the device. How much it grows with size depends on the model: some turn every image into a fixed number of tokens, others into more tokens the more pixels there are. Lower is better.
TPS (tokens per second) is how fast the answer is then written, per camera. It barely depends on the image, and it is shown for reference: Max Cameras is decided by the response time. A long answer adds its writing time on top — 128 tokens at 25 tokens per second take about five more seconds.
Max Cameras: the number of cameras is the concurrency
A camera sends its next request as soon as the previous one is answered, so it always has exactly one request in flight. The number of concurrent requests is therefore the number of cameras.
Max Cameras is the highest measured number of cameras at which the response time meets your target. Only measured numbers count, nothing is interpolated, and if none passes it is 0. It depends on your target and moves with it: at 720p with one image per camera, Cosmos3-Edge on Jetson AGX Orin keeps up with 2 cameras at 1 second and 8 — the highest number measured, so 8+ — at 3 seconds.
The columns that describe the setup
Parameters — the model’s total number of weights, vision encoder included, as published. For a mixture-of-experts model this is the total, not the part active for each token, because all of it is held in memory.
Quantization — the number format the weights are stored in. BF16 is full precision, FP8 uses 8 bits and NVFP4 4 bits. Q8_0, Q4_K_M and Q4_0 are GGUF formats for llama.cpp with about 8 and 4 bits per weight. Fewer bits means less memory and usually more speed, at some risk to quality.
Engine — the server software that loads the model and schedules requests. It affects speed as much as the hardware does. Qwen3-VL-4B-Instruct on RTX PRO 6000 at 720p starts answering a single camera after 0.25 s under llama.cpp (Q8_0) and 0.16 s under vLLM (BF16); at 16 cameras the gap is 1.09 s against 0.46 s.
Device — Jetson Orin NX (16 GB) and Jetson AGX Orin (32 GB) are embedded modules whose CPU and GPU share one memory; RTX PRO 6000 Blackwell is a workstation GPU with 96 GB, measured in its 600 W Workstation and 300 W Max-Q editions.
What the numbers do not tell you
The table uses one fixed workload — sample photos, a short prompt, answers of up to 128 tokens — and mean values, so that configurations can be compared without a site survey. It is not a substitute for a test with your own cameras and prompts: other image content, longer prompts or answers, other engine settings and the slowest requests rather than the mean all change real capacity. The response time is the wait until the answer starts; an application that needs the whole answer waits for its writing time too, and a row measured with one camera only says nothing about how it behaves with several.
Cosmos-Reason1-7B (8.3B parameters), served in Q4_K_M format with llama.cpp on Jetson Orin NX, starts answering a single camera that sends one 720p image after 7.68 s, then writes 12 tok/s.
With 5 images per camera at 720p, the answer starts after 37.76 s. With one 1080p image, after 23.3 s.
At the default target of an answer starting within 3 s (720p, one image per camera), not even one camera is answered in time.
Cosmos-Reason2-2B (2.4B parameters), served in Q8_0 format with llama.cpp on Jetson Orin NX, starts answering a single camera that sends one 720p image after 2.02 s, then writes 30.4 tok/s.
With 5 images per camera at 720p, the answer starts after 9.94 s. With one 1080p image, after 5.45 s.
At the default target of an answer starting within 3 s (720p, one image per camera), it keeps up with 2 cameras at once.
Cosmos-Reason2-8B (8.8B parameters), served in Q4_K_M format with llama.cpp on Jetson Orin NX, starts answering a single camera that sends one 720p image after 6.87 s, then writes 11.6 tok/s.
With 5 images per camera at 720p, the answer starts after 34.19 s. With one 1080p image, after 21.75 s.
At the default target of an answer starting within 3 s (720p, one image per camera), not even one camera is answered in time.
Gemma-4-E2B-it (5.1B parameters), served in Q4_0 format with llama.cpp on Jetson Orin NX, starts answering a single camera that sends one 720p image after 1.38 s, then writes 29.8 tok/s.
With 5 images per camera at 720p, the answer starts after 6.31 s. With one 1080p image, after 3.43 s.
At the default target of an answer starting within 3 s (720p, one image per camera), it keeps up with 2 cameras at once.
Gemma-4-E4B-it (8.0B parameters), served in Q4_0 format with llama.cpp on Jetson Orin NX, starts answering a single camera that sends one 720p image after 1.86 s, then writes 17.2 tok/s.
With 5 images per camera at 720p, the answer starts after 8.32 s. With one 1080p image, after 4.45 s.
At the default target of an answer starting within 3 s (720p, one image per camera), it keeps up with one camera at a time.
Qwen3-VL-4B-Instruct (4.4B parameters), served in FP8 format with vLLM on Jetson Orin NX, starts answering a single camera that sends one 720p image after 1.67 s, then writes 14.2 tok/s.
With 3 images per camera at 720p, the answer starts after 4 s. With one 1080p image, after 4.34 s.
At the default target of an answer starting within 3 s (720p, one image per camera), it keeps up with 2 cameras at once.
Qwen3-VL-4B-Instruct (4.4B parameters), served in Q8_0 format with llama.cpp on Jetson Orin NX, starts answering a single camera that sends one 720p image after 3.04 s, then writes 14.3 tok/s.
With 5 images per camera at 720p, the answer starts after 14.62 s. With one 2K image, after 18.39 s.
At the default target of an answer starting within 3 s (720p, one image per camera), not even one camera is answered in time.
Qwen3-VL-8B-Instruct (8.8B parameters), served in Q4_K_M format with llama.cpp on Jetson Orin NX, starts answering a single camera that sends one 720p image after 6.92 s, then writes 11.4 tok/s.
With 5 images per camera at 720p, the answer starts after 34.43 s. With one 1080p image, after 21.78 s.
At the default target of an answer starting within 3 s (720p, one image per camera), not even one camera is answered in time.
Qwen3.5-4B (4.7B parameters), served in Q4_K_M format with llama.cpp on Jetson Orin NX, starts answering a single camera that sends one 720p image after 3.72 s, then writes 15.6 tok/s.
With 5 images per camera at 720p, the answer starts after 17.25 s. With one 1080p image, after 8.48 s.
At the default target of an answer starting within 3 s (720p, one image per camera), not even one camera is answered in time.
Cosmos-Reason1-7B (8.3B parameters), served in Q4_K_M format with llama.cpp on Jetson AGX Orin, starts answering a single camera that sends one 720p image after 3.21 s, then writes 26.7 tok/s.
With 5 images per camera at 720p, the answer starts after 15.84 s. With one 2K image, after 18.62 s.
At the default target of an answer starting within 3 s (720p, one image per camera), not even one camera is answered in time.
Cosmos-Reason2-2B (2.4B parameters), served in Q8_0 format with llama.cpp on Jetson AGX Orin, starts answering a single camera that sends one 720p image after 0.95 s, then writes 57.8 tok/s.
With 5 images per camera at 720p, the answer starts after 4.45 s. With one 2K image, after 7.07 s.
At the default target of an answer starting within 3 s (720p, one image per camera), it keeps up with 4 cameras at once.
Cosmos-Reason2-8B (8.8B parameters), served in Q4_K_M format with llama.cpp on Jetson AGX Orin, starts answering a single camera that sends one 720p image after 2.71 s, then writes 25.4 tok/s.
With 5 images per camera at 720p, the answer starts after 13.21 s. With one 2K image, after 20.51 s.
At the default target of an answer starting within 3 s (720p, one image per camera), it keeps up with one camera at a time.
Cosmos3-Edge (3.9B parameters), served in BF16 format with vLLM on Jetson AGX Orin, starts answering a single camera that sends one 720p image after 0.51 s, then writes 45.6 tok/s.
With 5 images per camera at 720p, the answer starts after 2.11 s. With one 2K image, after 2.71 s.
At the default target of an answer starting within 3 s (720p, one image per camera), it keeps up with at least 8 cameras at once: it still met the target at the highest number measured, so the real limit is higher.
Gemma-4-26B-A4B-it (26B parameters), served in Q4_0 format with llama.cpp on Jetson AGX Orin, starts answering a single camera that sends one 720p image after 2.34 s, then writes 32.7 tok/s.
With 5 images per camera at 720p, the answer starts after 10.78 s. With one 2K image, after 9.94 s.
At the default target of an answer starting within 3 s (720p, one image per camera), it keeps up with 2 cameras at once.
Gemma-4-E4B-it (8.0B parameters), served in Q4_0 format with llama.cpp on Jetson AGX Orin, starts answering a single camera that sends one 720p image after 0.86 s, then writes 34.2 tok/s.
With 5 images per camera at 720p, the answer starts after 3.58 s. With one 2K image, after 2.55 s.
At the default target of an answer starting within 3 s (720p, one image per camera), it keeps up with 4 cameras at once.
Qwen3-VL-4B-Instruct (4.4B parameters), served in Q8_0 format with llama.cpp on Jetson AGX Orin, starts answering a single camera that sends one 720p image after 1.31 s, then writes 28.4 tok/s.
With 5 images per camera at 720p, the answer starts after 6.4 s. With one 2K image, after 8.32 s.
At the default target of an answer starting within 3 s (720p, one image per camera), it keeps up with 2 cameras at once.
Qwen3-VL-8B-Instruct (8.8B parameters), served in Q4_K_M format with llama.cpp on Jetson AGX Orin, starts answering a single camera that sends one 720p image after 2.75 s, then writes 25.2 tok/s.
With 5 images per camera at 720p, the answer starts after 13.24 s. With one 2K image, after 20.61 s.
At the default target of an answer starting within 3 s (720p, one image per camera), it keeps up with one camera at a time.
Qwen3-VL-8B-Instruct (8.8B parameters), served in Q8_0 format with llama.cpp on Jetson AGX Orin, starts answering a single camera that sends one 720p image after 2.64 s, then writes 19.1 tok/s.
With 5 images per camera at 720p, the answer starts after 13.19 s. With one 2K image, after 20.48 s.
At the default target of an answer starting within 3 s (720p, one image per camera), it keeps up with one camera at a time.
Qwen3.5-4B (4.7B parameters), served in Q4_K_M format with llama.cpp on Jetson AGX Orin, starts answering a single camera that sends one 720p image after 1.57 s, then writes 33.6 tok/s.
With 5 images per camera at 720p, the answer starts after 7.14 s. With one 2K image, after 9.04 s.
At the default target of an answer starting within 3 s (720p, one image per camera), it keeps up with 2 cameras at once.
Cosmos3-Edge (3.9B parameters), served in BF16 format with vLLM on RTX PRO 6000 Max-Q, starts answering a single camera that sends one 720p image after 0.09 s, then writes 310.3 tok/s.
With 5 images per camera at 720p, the answer starts after 0.34 s. With one 2K image, after 0.43 s.
At the default target of an answer starting within 3 s (720p, one image per camera), it keeps up with at least 16 cameras at once: it still met the target at the highest number measured, so the real limit is higher.
Cosmos3-Nano (16B parameters), served in BF16 format with vLLM on RTX PRO 6000 Max-Q, starts answering a single camera that sends one 720p image after 0.14 s, then writes 88 tok/s.
With 5 images per camera at 720p, the answer starts after 0.74 s. With one 2K image, after 0.75 s.
At the default target of an answer starting within 3 s (720p, one image per camera), it keeps up with at least 16 cameras at once: it still met the target at the highest number measured, so the real limit is higher.
Cosmos-Reason1-7B (8.3B parameters), served in BF16 format with vLLM on RTX PRO 6000, starts answering a single camera that sends one 720p image after 0.25 s, then writes 98.2 tok/s.
With 5 images per camera at 720p, the answer starts after 0.71 s. With one 2K image, after 0.61 s.
At the default target of an answer starting within 3 s (720p, one image per camera), it keeps up with at least 16 cameras at once: it still met the target at the highest number measured, so the real limit is higher.
Cosmos-Reason2-2B (2.4B parameters), served in BF16 format with vLLM on RTX PRO 6000, starts answering a single camera that sends one 720p image after 0.11 s, then writes 318.3 tok/s.
With 5 images per camera at 720p, the answer starts after 0.56 s. With one 2K image, after 0.5 s.
At the default target of an answer starting within 3 s (720p, one image per camera), it keeps up with at least 16 cameras at once: it still met the target at the highest number measured, so the real limit is higher.
Cosmos-Reason2-8B (8.8B parameters), served in BF16 format with vLLM on RTX PRO 6000, starts answering a single camera that sends one 720p image after 0.15 s, then writes 92.5 tok/s.
With 5 images per camera at 720p, the answer starts after 0.75 s. With one 2K image, after 0.68 s.
At the default target of an answer starting within 3 s (720p, one image per camera), it keeps up with at least 16 cameras at once: it still met the target at the highest number measured, so the real limit is higher.
Eagle2.5-8B (8.1B parameters), served in BF16 format with vLLM on RTX PRO 6000, starts answering a single camera that sends one 720p image after 0.26 s, then writes 105 tok/s.
With 5 images per camera at 720p, the answer starts after 1.35 s. With one 2K image, after 0.42 s.
At the default target of an answer starting within 3 s (720p, one image per camera), it keeps up with at least 16 cameras at once: it still met the target at the highest number measured, so the real limit is higher.
Gemma-4-31B-it (31B parameters), served in BF16 format with vLLM on RTX PRO 6000, starts answering a single camera that sends one 720p image after 0.24 s, then writes 24.7 tok/s.
With 5 images per camera at 720p, the answer starts after 0.7 s. With one 2K image, after 0.32 s.
At the default target of an answer starting within 3 s (720p, one image per camera), it keeps up with at least 16 cameras at once: it still met the target at the highest number measured, so the real limit is higher.
Gemma-4-31B-it (31B parameters), served in NVFP4 format with vLLM on RTX PRO 6000, starts answering a single camera that sends one 720p image after 0.27 s, then writes 45.1 tok/s.
With 5 images per camera at 720p, the answer starts after 0.58 s. With one 2K image, after 0.3 s.
At the default target of an answer starting within 3 s (720p, one image per camera), it keeps up with at least 16 cameras at once: it still met the target at the highest number measured, so the real limit is higher.
Llama-3.1-Nemotron-Nano-VL-8B (8.7B parameters), served in BF16 format with vLLM on RTX PRO 6000, starts answering a single camera that sends one 720p image after 0.29 s, then writes 94.9 tok/s.
With 5 images per camera at 720p, the answer starts after 1.79 s. With one 2K image, after 0.49 s.
At the default target of an answer starting within 3 s (720p, one image per camera), it keeps up with at least 16 cameras at once: it still met the target at the highest number measured, so the real limit is higher.
MedGemma-1.5-4B-it (4.3B parameters), served in BF16 format with vLLM on RTX PRO 6000, starts answering a single camera that sends one 720p image after 0.2 s, then writes 171.5 tok/s.
With 5 images per camera at 720p, the answer starts after 0.45 s. With one 2K image, after 0.25 s.
At the default target of an answer starting within 3 s (720p, one image per camera), it keeps up with at least 16 cameras at once: it still met the target at the highest number measured, so the real limit is higher.
Nemotron-Nano-12B-v2-VL (13B parameters), served in BF16 format with vLLM on RTX PRO 6000, starts answering a single camera that sends one 720p image after 0.2 s, then writes 68.5 tok/s.
With 5 images per camera at 720p, the answer starts after 0.78 s. With one 2K image, after 0.61 s.
At the default target of an answer starting within 3 s (720p, one image per camera), it keeps up with at least 16 cameras at once: it still met the target at the highest number measured, so the real limit is higher.
Qwen3-VL-30B-A3B-Instruct (30B parameters), served in FP8 format with vLLM on RTX PRO 6000, starts answering a single camera that sends one 720p image after 0.16 s, then writes 194.4 tok/s.
With 5 images per camera at 720p, the answer starts after 1.04 s. With one 2K image, after 0.58 s.
At the default target of an answer starting within 3 s (720p, one image per camera), it keeps up with at least 16 cameras at once: it still met the target at the highest number measured, so the real limit is higher.
Qwen3-VL-4B-Instruct (4.4B parameters), served in BF16 format with vLLM on RTX PRO 6000, starts answering a single camera that sends one 720p image after 0.16 s, then writes 156.6 tok/s.
With 5 images per camera at 720p, the answer starts after 0.64 s. With one 2K image, after 0.56 s.
At the default target of an answer starting within 3 s (720p, one image per camera), it keeps up with at least 16 cameras at once: it still met the target at the highest number measured, so the real limit is higher.
Qwen3-VL-4B-Instruct (4.4B parameters), served in Q8_0 format with llama.cpp on RTX PRO 6000, starts answering a single camera that sends one 720p image after 0.25 s, then writes 231.4 tok/s.
With 5 images per camera at 720p, the answer starts after 0.76 s. With one 2K image, after 0.94 s.
At the default target of an answer starting within 3 s (720p, one image per camera), it keeps up with at least 16 cameras at once: it still met the target at the highest number measured, so the real limit is higher.
Qwen3-VL-8B-Instruct (8.8B parameters), served in BF16 format with vLLM on RTX PRO 6000, starts answering a single camera that sends one 720p image after 0.23 s, then writes 92.4 tok/s.
With 5 images per camera at 720p, the answer starts after 0.76 s. With one 2K image, after 0.68 s.
At the default target of an answer starting within 3 s (720p, one image per camera), it keeps up with at least 16 cameras at once: it still met the target at the highest number measured, so the real limit is higher.
Qwen3-VL-8B-Instruct (8.8B parameters), served in Q8_0 format with llama.cpp on RTX PRO 6000, starts answering a single camera that sends one 720p image after 0.33 s, then writes 153.8 tok/s.
With 5 images per camera at 720p, the answer starts after 1.04 s. With one 2K image, after 1.79 s.
At the default target of an answer starting within 3 s (720p, one image per camera), it keeps up with at least 16 cameras at once: it still met the target at the highest number measured, so the real limit is higher.
Qwen3.6-35B-A3B (35B parameters), served in FP8 format with vLLM on RTX PRO 6000, starts answering a single camera that sends one 720p image after 0.21 s, then writes 220.8 tok/s.
With 5 images per camera at 720p, the answer starts after 0.84 s. With one 2K image, after 0.59 s.
At the default target of an answer starting within 3 s (720p, one image per camera), it keeps up with at least 16 cameras at once: it still met the target at the highest number measured, so the real limit is higher.
Qwen3.6-35B-A3B (35B parameters), served in NVFP4 format with vLLM on RTX PRO 6000, starts answering a single camera that sends one 720p image after 0.22 s, then writes 282.9 tok/s.
With 5 images per camera at 720p, the answer starts after 0.84 s. With one 2K image, after 0.8 s.
At the default target of an answer starting within 3 s (720p, one image per camera), it keeps up with at least 16 cameras at once: it still met the target at the highest number measured, so the real limit is higher.
Qwen3.8-27B (27B parameters), served in FP8 format with vLLM on RTX PRO 6000, starts answering a single camera that sends one 720p image after 0.29 s, then writes 50 tok/s.
With 5 images per camera at 720p, the answer starts after 1.01 s. With one 2K image, after 0.91 s.
At the default target of an answer starting within 3 s (720p, one image per camera), it keeps up with at least 16 cameras at once: it still met the target at the highest number measured, so the real limit is higher.