Pinned benchmark × current cloud price
qwen3-vl-235b-a22b Offline cost
Reprice the verified qwen3-vl-235b-a22b Offline result with current Nebius AI Cloud B300 rates, transparent formulas and pinned evidence. The benchmark result and cloud price keep independent provenance, so a price refresh never rewrites the measured performance.
The joined cloud price is older than 18 hours. The pinned benchmark remains valid, but repricing is paused until Guardian refreshes the rate.
Reproducible workload economics
Scale the measured result
The published throughput is converted into elapsed cluster-hours per one million measured units, then multiplied by the matched cluster rate. Change only the number of equivalent runs or output batches; the evidence-backed performance stays fixed. The primary estimate uses the lowest compatible observed rate; the range shows sensitivity across all compatible published rates and does not imply capacity.
Measured boundary
Exactly what was benchmarked
Nebius B300 n1 (8x B300-SXM-270GB, TensorRT) reported 78.2775 Samples/s for qwen3-vl-235b-a22b in the Offline scenario. The quality target was F1_HIERARCHICAL: 0.7880721729948121. Software was TensorRT 10.14, CUDA 13.1, cuDNN 9.17, TensorRT-LLM feat/1.2-mlpinf, NVIDIA Dynamo mlperf-v6.0-dynamo-v0.8.0, vLLM CentML:mlperf-inf-mm-q3vl-v6.0. This page does not extrapolate to a different model, implementation, accelerator count or quality target.
Interpretation guardrails
Use this estimate inside its evidence boundary
Comparable workload
The cost applies to this named benchmark, scenario, software stack, quality target and published accelerator system. A production workload can produce different throughput or time.
Serial time only
The displayed time scales benchmark-equivalent runs or output batches serially on the measured system. It excludes queueing, provisioning and any parallel execution strategy.
Compute line only
Storage, networking, data preparation, queueing, engineering time, failed jobs, support and taxes are excluded. Add them in a complete deployment budget.
Capacity separate
The price source proves a published rate, not immediate capacity. Availability remains a separate signal and must be confirmed before scheduling work.
- The result applies only to the published model, software, scenario and 8-GPU system.
- The current Nebius normalized per-GPU rate is joined separately and multiplied by the benchmark GPU count; identical topology and simultaneous capacity are not asserted.
- Storage, data transfer, orchestration, queueing, support, taxes and discounts are excluded.
- A benchmark result is not a guarantee of performance for a different workload or configuration.
Related evidence
Change one decision variable
Move to another scenario or inspect the exact provider and accelerator pricing behind this result.
workload cost
gpt-oss 20b pretraining time to train to quality target gpu cost
2 verified rates from $7.850 per GPU-hour.
Open decision pageworkload cost
llama 3.1 8b pretraining time to train to quality target gpu cost
2 verified rates from $7.850 per GPU-hour.
Open decision pageworkload cost
deepseek-r1 offline gpu cost
2 verified rates from $7.850 per GPU-hour.
Open decision pageworkload cost
deepseek-r1 server gpu cost
2 verified rates from $7.850 per GPU-hour.
Open decision pageworkload cost
gpt-oss-120b offline gpu cost
2 verified rates from $7.850 per GPU-hour.
Open decision pageworkload cost
gpt-oss-120b server gpu cost
2 verified rates from $7.850 per GPU-hour.
Open decision page