Measured workload economics

Verified GPU cost benchmarks

Convert exact MLPerf 6.0 results into current cloud economics. See cost per million reported tokens, cost per batch unit and cost to reach a published training quality target—without transferring the result to a different model or hiding the assumptions.

  • Official MLCommons source paths pinned to full commits
  • 20 raw training runs retained by SHA-256 hash
  • Exact Nebius 8× B300 system and current official price match

Transparent repricer

Change the rate, never the published result

The benchmark throughput or time stays locked. Change only the per-GPU price and optional target volume to audit the arithmetic.

Reference unit cost
Cluster rate
Target timeAdd a target
Target costAdd a target

Scenario summary is ready to copy.

Exact published system

Published reference economics

Every result below uses the same Nebius eight-GPU B300 system and is re-priced with the current official on demand $7.85 per-GPU-hour catalog rate, equal to $62.80 per hour for the exact eight-GPU reference system. This is a transparent reference calculation, not a promise for another workload.

batchOffline

qwen3-vl-235b-a22b

$0.2229 / 1K samples

Published result
78.28 Samples/s
Reference time
12.78 s / 1K samples
System
8× NVIDIA B300
Reference cluster rate
$62.80 / h
Reference yield @ $100 compute
448.72K samples
Purchase model
on demand
Price observed
2026-08-13 02:40 UTC
Capacity
8 GPUs available together: not verified
Evidence
1 verified result
Source revision
4d3916ac9cf4

F1_HIERARCHICAL: 0.7880721729948121

Open workload cost sheet

Inspect source evidence ↗
batchServer

qwen3-vl-235b-a22b

$0.3864 / 1K queries

Published result
45.15 Queries/s
Reference time
22.15 s / 1K queries
System
8× NVIDIA B300
Reference cluster rate
$62.80 / h
Reference yield @ $100 compute
258.83K queries
Purchase model
on demand
Price observed
2026-08-13 02:40 UTC
Capacity
8 GPUs available together: not verified
Evidence
1 verified result
Source revision
4d3916ac9cf4

F1_HIERARCHICAL: 0.7865237675409595

Open workload cost sheet

Inspect source evidence ↗
inferenceOffline

deepseek-r1

$0.2517 / 1M reported tokens

Published result
69,318.9 Tokens/s
Reference time
14.43 s / 1M reported tokens
System
8× NVIDIA B300
Reference cluster rate
$62.80 / h
Reference yield @ $100 compute
397.37M reported tokens
Purchase model
on demand
Price observed
2026-08-13 02:40 UTC
Capacity
8 GPUs available together: not verified
Evidence
1 verified result
Source revision
4d3916ac9cf4

exact_match: 81.58614402917047 TOKENS_PER_SAMPLE: 3722.1116681859617

Open workload cost sheet

Inspect source evidence ↗
inferenceServer

deepseek-r1

$0.2888 / 1M reported tokens

Published result
60,413.4 Tokens/s
Reference time
16.55 s / 1M reported tokens
System
8× NVIDIA B300
Reference cluster rate
$62.80 / h
Reference yield @ $100 compute
346.32M reported tokens
Purchase model
on demand
Price observed
2026-08-13 02:40 UTC
Capacity
8 GPUs available together: not verified
Evidence
1 verified result
Source revision
4d3916ac9cf4

exact_match: 81.58614402917047 TOKENS_PER_SAMPLE: 3721.7894257064722

Open workload cost sheet

Inspect source evidence ↗
inferenceOffline

gpt-oss-120b

$0.1632 / 1M reported tokens

Published result
106,885 Tokens/s
Reference time
9.36 s / 1M reported tokens
System
8× NVIDIA B300
Reference cluster rate
$62.80 / h
Reference yield @ $100 compute
612.72M reported tokens
Purchase model
on demand
Price observed
2026-08-13 02:40 UTC
Capacity
8 GPUs available together: not verified
Evidence
1 verified result
Source revision
4d3916ac9cf4

exact_match: 82.959

Open workload cost sheet

Inspect source evidence ↗
inferenceServer

gpt-oss-120b

$0.1737 / 1M reported tokens

Published result
100,437 Tokens/s
Reference time
9.96 s / 1M reported tokens
System
8× NVIDIA B300
Reference cluster rate
$62.80 / h
Reference yield @ $100 compute
575.75M reported tokens
Purchase model
on demand
Price observed
2026-08-13 02:40 UTC
Capacity
8 GPUs available together: not verified
Evidence
1 verified result
Source revision
4d3916ac9cf4

exact_match: 83.337

Open workload cost sheet

Inspect source evidence ↗
trainingTime to train to quality target

GPT-OSS 20B pretraining

$87.3772 / reference run

Published result
5,008.89 seconds
Reference time
83.48 min to published quality target
System
8× NVIDIA B300
Reference cluster rate
$62.80 / h
Reference yield @ $100 compute
1.14 reference runs
Purchase model
on demand
Price observed
2026-08-13 02:40 UTC
Capacity
8 GPUs available together: not verified
Evidence
10 verified runs
Source revision
eabf23a07b2a

C4 · 3.34 log perplexity

Open workload cost sheet

Inspect source evidence ↗
trainingTime to train to quality target

Llama 3.1 8B pretraining

$75.3729 / reference run

Published result
4,320.74 seconds
Reference time
72.01 min to published quality target
System
8× NVIDIA B300
Reference cluster rate
$62.80 / h
Reference yield @ $100 compute
1.33 reference runs
Purchase model
on demand
Price observed
2026-08-13 02:40 UTC
Capacity
8 GPUs available together: not verified
Evidence
10 verified runs
Source revision
eabf23a07b2a

C4 · 3.3 log perplexity

Open workload cost sheet

Inspect source evidence ↗

Evidence boundary

What this proves—and what it does not

These are results for a published Nebius system with eight NVIDIA B300 accelerators, a named model, scenario, software stack and quality result. The economics layer multiplies the current official per-GPU price by eight and applies only the published throughput or time.

ReproducibleFull repository commit, source paths, raw-run hashes and aggregation method.
Current price joinSame provider, GPU family and on-demand market; pricing evidence stays separate.
Honest uncertaintyNo claim about immediate eight-GPU capacity, storage, transfer, queueing or another workload.

Machine-readable evidence

Audit or reuse the full dataset