Batch and multimodal economics

GPU batch job cost calculator

Turn your measured cluster throughput into time, total cloud cost and cost per task. Compare the scenario with exact Qwen3-VL MLPerf Offline and Server results without pretending that those results transfer to your workload.

  • No estimated TFLOPS conversion
  • Explicit retry and checkpoint overhead
  • Published samples/s and queries/s kept distinct

Your measured scenario

Estimate one complete batch

A task can be an image, request, render frame or document only if your throughput measurement uses that same unit.

Time to complete
Cluster rate
Total job cost
Cost per task

Complete every field with measured or planned inputs.

Exact published system

Published Qwen3-VL batch reference economics

Every result below uses the same Nebius eight-GPU B300 system and is re-priced with the current official on demand $7.85 per-GPU-hour catalog rate, equal to $62.80 per hour for the exact eight-GPU reference system. This is a transparent reference calculation, not a promise for another workload.

batchOffline

qwen3-vl-235b-a22b

$0.2229 / 1K samples

Published result
78.28 Samples/s
Reference time
12.78 s / 1K samples
System
8× NVIDIA B300
Reference cluster rate
$62.80 / h
Reference yield @ $100 compute
448.72K samples
Purchase model
on demand
Price observed
2026-08-13 02:40 UTC
Capacity
8 GPUs available together: not verified
Evidence
1 verified result
Source revision
4d3916ac9cf4

F1_HIERARCHICAL: 0.7880721729948121

Inspect source evidence ↗
batchServer

qwen3-vl-235b-a22b

$0.3864 / 1K queries

Published result
45.15 Queries/s
Reference time
22.15 s / 1K queries
System
8× NVIDIA B300
Reference cluster rate
$62.80 / h
Reference yield @ $100 compute
258.83K queries
Purchase model
on demand
Price observed
2026-08-13 02:40 UTC
Capacity
8 GPUs available together: not verified
Evidence
1 verified result
Source revision
4d3916ac9cf4

F1_HIERARCHICAL: 0.7865237675409595

Inspect source evidence ↗

Calculation

Formula you can audit

Seconds = tasks ÷ measured cluster tasks/second × (1 + overhead). Total cost = seconds ÷ 3,600 × GPUs × price per GPU-hour. Unit cost = total cost ÷ tasks.

OfflinePublished throughput is reported in samples per second.
ServerPublished throughput is reported in queries per second and is never merged with samples.
BoundaryYour task definition and measured throughput must remain internally consistent.