Training and fine-tuning economics
GPU training cost calculator
Calculate time and cloud cost from the tokens you will process and throughput measured on your exact cluster. Then inspect 6 published MLPerf time-to-quality references across 3 exact systems—kept separate from your estimate.
- No default training throughput
- Transparent epochs and overhead
- 60 hashed source runs across three independent submitters
Training reference price snapshot:Checking the declared provider-SLA boundary.Valid through ; historical after that instant.
Your measured scenario
Estimate one complete training job
Cluster throughput must come from your own model, precision, sequence length, software and hardware test. Overhead adds checkpointing, evaluation and restart allowance.
Complete every field with measured or planned inputs.
Before you commit
Test this estimate against the costs the calculator cannot know
A buyer-funded review challenges your assumptions against redacted quotes, capacity evidence, contract terms and excluded costs without changing the technical ranking.
Exact published system
6 reproducible training reference profiles
These 6 results cover 3 published systems from 3 matched cloud providers. Each result was re-priced at price-catalog generation only with the same provider, exact normalized GPU model and on demand market, then multiplied by that benchmark's published GPU count. This is a transparent reference calculation, not proof of identical topology, capacity or performance for another workload. The contributing price snapshot is valid through and remains visible as historical evidence afterward.
GPT-OSS 20B pretraining
$247.5209 / reference run
- Published result
- 1,618.96 seconds
- Reference time
- 26.98 min to published quality target
- System
- 64× NVIDIA Blackwell GPU (B200-SXM-180GB)
- Cloud price match
- CoreWeave · B200
- Reference cluster rate
- $550.40 / h
- Reference yield @ $100 compute
- 0.4 reference runs
- Purchase model
- on demand
- Price observed
- 2026-09-09 22:35 UTC
- Capacity
- 64 GPUs available together: not verified
- Evidence
- 10 verified runs
- Source revision
eabf23a07b2a
C4 · 3.34 log perplexity
Inspect source evidence ↗Llama 3.1 8B pretraining
$151.6928 / reference run
- Published result
- 992.18 seconds
- Reference time
- 16.54 min to published quality target
- System
- 64× NVIDIA Blackwell GPU (B200-SXM-180GB)
- Cloud price match
- CoreWeave · B200
- Reference cluster rate
- $550.40 / h
- Reference yield @ $100 compute
- 0.66 reference runs
- Purchase model
- on demand
- Price observed
- 2026-09-09 22:35 UTC
- Capacity
- 64 GPUs available together: not verified
- Evidence
- 10 verified runs
- Source revision
eabf23a07b2a
C4 · 3.3 log perplexity
Inspect source evidence ↗GPT-OSS 20B pretraining
$87.3772 / reference run
- Published result
- 5,008.89 seconds
- Reference time
- 83.48 min to published quality target
- System
- 8× NVIDIA Blackwell Ultra GPU (B300-SXM-270GB)
- Cloud price match
- Nebius AI Cloud · B300
- Reference cluster rate
- $62.80 / h
- Reference yield @ $100 compute
- 1.14 reference runs
- Purchase model
- on demand
- Price observed
- 2026-09-09 22:35 UTC
- Capacity
- 8 GPUs available together: not verified
- Evidence
- 10 verified runs
- Source revision
eabf23a07b2a
C4 · 3.34 log perplexity
Inspect source evidence ↗GPT-OSS 20B pretraining
$86.0397 / reference run
- Published result
- 5,787.43 seconds
- Reference time
- 96.46 min to published quality target
- System
- 8× NVIDIA Blackwell GPU (B200-SXM-180GB)
- Cloud price match
- Lambda · B200
- Reference cluster rate
- $53.52 / h
- Reference yield @ $100 compute
- 1.16 reference runs
- Purchase model
- on demand
- Price observed
- 2026-09-09 22:35 UTC
- Capacity
- 8 GPUs available together: not verified
- Evidence
- 10 verified runs
- Source revision
eabf23a07b2a
C4 · 3.34 log perplexity
Inspect source evidence ↗Llama 3.1 8B pretraining
$76.0547 / reference run
- Published result
- 5,115.79 seconds
- Reference time
- 85.26 min to published quality target
- System
- 8× NVIDIA Blackwell GPU (B200-SXM-180GB)
- Cloud price match
- Lambda · B200
- Reference cluster rate
- $53.52 / h
- Reference yield @ $100 compute
- 1.31 reference runs
- Purchase model
- on demand
- Price observed
- 2026-09-09 22:35 UTC
- Capacity
- 8 GPUs available together: not verified
- Evidence
- 10 verified runs
- Source revision
eabf23a07b2a
C4 · 3.3 log perplexity
Inspect source evidence ↗Llama 3.1 8B pretraining
$75.3729 / reference run
- Published result
- 4,320.74 seconds
- Reference time
- 72.01 min to published quality target
- System
- 8× NVIDIA Blackwell Ultra GPU (B300-SXM-270GB)
- Cloud price match
- Nebius AI Cloud · B300
- Reference cluster rate
- $62.80 / h
- Reference yield @ $100 compute
- 1.33 reference runs
- Purchase model
- on demand
- Price observed
- 2026-09-09 22:35 UTC
- Capacity
- 8 GPUs available together: not verified
- Evidence
- 10 verified runs
- Source revision
eabf23a07b2a
C4 · 3.3 log perplexity
Inspect source evidence ↗Calculation
Formula you can audit
Processed tokens = dataset tokens × epochs. Seconds = processed tokens ÷ measured cluster tokens/second × (1 + overhead). Cost = seconds ÷ 3,600 × GPUs × price per GPU-hour.