Pinned benchmark × current cloud price

GPT-OSS 20B pretraining Time to train to quality target cost

Reprice the verified GPT-OSS 20B pretraining Time to train to quality target result with current Nebius AI Cloud B300 rates, transparent formulas and pinned evidence. The benchmark result and cloud price keep independent provenance, so a price refresh never rewrites the measured performance.

Reproducible workload economics

Scale the measured result

The published time-to-quality is converted into cluster-hours, then multiplied by the current matched cluster rate. Change only the number of equivalent runs or output batches; the evidence-backed performance stays fixed. The primary estimate uses the lowest compatible observed rate; the range shows sensitivity across all compatible published rates and does not imply capacity.

Lowest compatible compute cost$87.3772
Compatible published-rate range$87.3772
Serial time on benchmark-equivalent system1.3914 h
Cost basisper benchmark-equivalent run

Inspectable formula

cost/run = measured seconds ÷ 3,600 × accelerator count × USD/GPU-hour

1 × $87.3772 = $87.3772

1 × 1.3914 h = 1.3914 h serial benchmark-equivalent system time

Measured boundary

Exactly what was benchmarked

Nebius B300 n1 (8x B300-SXM-270GB) reported 5,008.8865 seconds for GPT-OSS 20B pretraining in the Time to train to quality target scenario. The quality target was 3.34 log perplexity. Software was NVIDIA NeMo Framework Release 25.09. This page does not extrapolate to a different model, implementation, accelerator count or quality target.

Benchmark evidenceMLPerf commit eabf23a07b2a0c60a289ff871dc3a46fff0d0421. Aggregation: trimmed mean discard lowest and highest; 10 published runs.Open pinned result ↗
Join ruleProvider, accelerator family, GPU count context, market type, USD currency and per-GPU-hour unit must remain compatible. Missing evidence fails closed.Read the full methodology

Interpretation guardrails

Use this estimate inside its evidence boundary

Comparable workload

The cost applies to this named benchmark, scenario, software stack, quality target and published accelerator system. A production workload can produce different throughput or time.

Serial time only

The displayed time scales benchmark-equivalent runs or output batches serially on the measured system. It excludes queueing, provisioning and any parallel execution strategy.

Compute line only

Storage, networking, data preparation, queueing, engineering time, failed jobs, support and taxes are excluded. Add them in a complete deployment budget.

Capacity separate

The price source proves a published rate, not immediate capacity. Availability remains a separate signal and must be confirmed before scheduling work.

Related evidence

Change one decision variable

Move to another scenario or inspect the exact provider and accelerator pricing behind this result.

workload cost

llama 3.1 8b pretraining time to train to quality target gpu cost

2 verified rates from $7.850 per GPU-hour.

Open decision page

workload cost

deepseek-r1 offline gpu cost

2 verified rates from $7.850 per GPU-hour.

Open decision page

workload cost

deepseek-r1 server gpu cost

2 verified rates from $7.850 per GPU-hour.

Open decision page

workload cost

gpt-oss-120b offline gpu cost

2 verified rates from $7.850 per GPU-hour.

Open decision page

workload cost

gpt-oss-120b server gpu cost

2 verified rates from $7.850 per GPU-hour.

Open decision page

workload cost

qwen3-vl-235b-a22b offline gpu cost

2 verified rates from $7.850 per GPU-hour.

Open decision page