Pinned benchmark × current cloud price
Llama 3.1 8B pretraining Time to train to quality target cost
Reprice the verified Llama 3.1 8B pretraining Time to train to quality target result with current Nebius AI Cloud B300 rates, transparent formulas and pinned evidence. The benchmark result and cloud price keep independent provenance, so a price refresh never rewrites the measured performance.
The joined cloud price is older than 18 hours. The pinned benchmark remains valid, but repricing is paused until Guardian refreshes the rate.
Reproducible workload economics
Scale the measured result
The published time-to-quality is converted into cluster-hours, then multiplied by the current matched cluster rate. Change only the number of equivalent runs or output batches; the evidence-backed performance stays fixed. The primary estimate uses the lowest compatible observed rate; the range shows sensitivity across all compatible published rates and does not imply capacity.
Measured boundary
Exactly what was benchmarked
Nebius B300 n1 (8x B300-SXM-270GB) reported 4,320.7418 seconds for Llama 3.1 8B pretraining in the Time to train to quality target scenario. The quality target was 3.3 log perplexity. Software was NVIDIA NeMo Framework Release 25.09. This page does not extrapolate to a different model, implementation, accelerator count or quality target.
Interpretation guardrails
Use this estimate inside its evidence boundary
Comparable workload
The cost applies to this named benchmark, scenario, software stack, quality target and published accelerator system. A production workload can produce different throughput or time.
Serial time only
The displayed time scales benchmark-equivalent runs or output batches serially on the measured system. It excludes queueing, provisioning and any parallel execution strategy.
Compute line only
Storage, networking, data preparation, queueing, engineering time, failed jobs, support and taxes are excluded. Add them in a complete deployment budget.
Capacity separate
The price source proves a published rate, not immediate capacity. Availability remains a separate signal and must be confirmed before scheduling work.
- The result applies only to the published model, software, scenario and 8-GPU system.
- The current Nebius normalized per-GPU rate is joined separately and multiplied by the benchmark GPU count; identical topology and simultaneous capacity are not asserted.
- Storage, data transfer, orchestration, queueing, support, taxes and discounts are excluded.
- A benchmark result is not a guarantee of performance for a different workload or configuration.
Related evidence
Change one decision variable
Move to another scenario or inspect the exact provider and accelerator pricing behind this result.
workload cost
gpt-oss 20b pretraining time to train to quality target gpu cost
2 verified rates from $7.850 per GPU-hour.
Open decision pageworkload cost
deepseek-r1 offline gpu cost
2 verified rates from $7.850 per GPU-hour.
Open decision pageworkload cost
deepseek-r1 server gpu cost
2 verified rates from $7.850 per GPU-hour.
Open decision pageworkload cost
gpt-oss-120b offline gpu cost
2 verified rates from $7.850 per GPU-hour.
Open decision pageworkload cost
gpt-oss-120b server gpu cost
2 verified rates from $7.850 per GPU-hour.
Open decision pageworkload cost
qwen3-vl-235b-a22b offline gpu cost
2 verified rates from $7.850 per GPU-hour.
Open decision page