Training and fine-tuning economics

GPU training cost calculator

Calculate time and cloud cost from the tokens you will process and throughput measured on your exact cluster. Then inspect 6 published MLPerf time-to-quality references across 3 exact systems—kept separate from your estimate.

  • No default training throughput
  • Transparent epochs and overhead
  • 60 hashed source runs across three independent submitters

Training reference price snapshot:Checking the declared provider-SLA boundary.Valid through ; historical after that instant.

Your measured scenario

Estimate one complete training job

Cluster throughput must come from your own model, precision, sequence length, software and hardware test. Overhead adds checkpointing, evaluation and restart allowance.

Tokens processed—
Time to complete—
Cluster rate—
Estimated job cost—

Complete every field with measured or planned inputs.

Before you commit

Test this estimate against the costs the calculator cannot know

A buyer-funded review challenges your assumptions against redacted quotes, capacity evidence, contract terms and excluded costs without changing the technical ranking.

Exact published system

6 reproducible training reference profiles

These 6 results cover 3 published systems from 3 matched cloud providers. Each result was re-priced at price-catalog generation only with the same provider, exact normalized GPU model and on demand market, then multiplied by that benchmark's published GPU count. This is a transparent reference calculation, not proof of identical topology, capacity or performance for another workload. The contributing price snapshot is valid through and remains visible as historical evidence afterward.

trainingTime to train to quality target

GPT-OSS 20B pretraining

$247.5209 / reference run

Published result
1,618.96 seconds
Reference time
26.98 min to published quality target
System
64× NVIDIA Blackwell GPU (B200-SXM-180GB)
Cloud price match
CoreWeave · B200
Reference cluster rate
$550.40 / h
Reference yield @ $100 compute
0.4 reference runs
Purchase model
on demand
Price observed
2026-09-09 22:35 UTC
Capacity
64 GPUs available together: not verified
Evidence
10 verified runs
Source revision
eabf23a07b2a

C4 · 3.34 log perplexity

Inspect source evidence ↗
trainingTime to train to quality target

Llama 3.1 8B pretraining

$151.6928 / reference run

Published result
992.18 seconds
Reference time
16.54 min to published quality target
System
64× NVIDIA Blackwell GPU (B200-SXM-180GB)
Cloud price match
CoreWeave · B200
Reference cluster rate
$550.40 / h
Reference yield @ $100 compute
0.66 reference runs
Purchase model
on demand
Price observed
2026-09-09 22:35 UTC
Capacity
64 GPUs available together: not verified
Evidence
10 verified runs
Source revision
eabf23a07b2a

C4 · 3.3 log perplexity

Inspect source evidence ↗
trainingTime to train to quality target

GPT-OSS 20B pretraining

$87.3772 / reference run

Published result
5,008.89 seconds
Reference time
83.48 min to published quality target
System
8× NVIDIA Blackwell Ultra GPU (B300-SXM-270GB)
Cloud price match
Nebius AI Cloud · B300
Reference cluster rate
$62.80 / h
Reference yield @ $100 compute
1.14 reference runs
Purchase model
on demand
Price observed
2026-09-09 22:35 UTC
Capacity
8 GPUs available together: not verified
Evidence
10 verified runs
Source revision
eabf23a07b2a

C4 · 3.34 log perplexity

Inspect source evidence ↗
trainingTime to train to quality target

GPT-OSS 20B pretraining

$86.0397 / reference run

Published result
5,787.43 seconds
Reference time
96.46 min to published quality target
System
8× NVIDIA Blackwell GPU (B200-SXM-180GB)
Cloud price match
Lambda · B200
Reference cluster rate
$53.52 / h
Reference yield @ $100 compute
1.16 reference runs
Purchase model
on demand
Price observed
2026-09-09 22:35 UTC
Capacity
8 GPUs available together: not verified
Evidence
10 verified runs
Source revision
eabf23a07b2a

C4 · 3.34 log perplexity

Inspect source evidence ↗
trainingTime to train to quality target

Llama 3.1 8B pretraining

$76.0547 / reference run

Published result
5,115.79 seconds
Reference time
85.26 min to published quality target
System
8× NVIDIA Blackwell GPU (B200-SXM-180GB)
Cloud price match
Lambda · B200
Reference cluster rate
$53.52 / h
Reference yield @ $100 compute
1.31 reference runs
Purchase model
on demand
Price observed
2026-09-09 22:35 UTC
Capacity
8 GPUs available together: not verified
Evidence
10 verified runs
Source revision
eabf23a07b2a

C4 · 3.3 log perplexity

Inspect source evidence ↗
trainingTime to train to quality target

Llama 3.1 8B pretraining

$75.3729 / reference run

Published result
4,320.74 seconds
Reference time
72.01 min to published quality target
System
8× NVIDIA Blackwell Ultra GPU (B300-SXM-270GB)
Cloud price match
Nebius AI Cloud · B300
Reference cluster rate
$62.80 / h
Reference yield @ $100 compute
1.33 reference runs
Purchase model
on demand
Price observed
2026-09-09 22:35 UTC
Capacity
8 GPUs available together: not verified
Evidence
10 verified runs
Source revision
eabf23a07b2a

C4 · 3.3 log perplexity

Inspect source evidence ↗

Calculation

Formula you can audit

Processed tokens = dataset tokens × epochs. Seconds = processed tokens ÷ measured cluster tokens/second × (1 + overhead). Cost = seconds ÷ 3,600 × GPUs × price per GPU-hour.

Your estimateUses only your measured throughput and explicit inputs.
MLPerf referenceUses ten successful runs per workload, discarding the fastest and slowest before averaging.
Not includedStorage, transfer, CPU/RAM add-ons, queueing, tax, discounts and capacity guarantees.