Workload price snapshot:Checking the declared provider-SLA boundary.Valid through ; historical after that instant.

Pinned benchmark × bounded price snapshot

deepseek-r1 Server cost

Reprice this verified deepseek-r1 Server result with an exact Nebius AI Cloud B300 price deadline, formula and pinned evidence. The benchmark result and cloud price keep independent provenance, so a price refresh never rewrites the measured performance.

Reproducible workload economics

Scale the measured result

The published throughput is converted into elapsed cluster-hours per one million measured units, then multiplied by the matched price-snapshot cluster rate. Change only the number of equivalent runs or output batches; the evidence-backed performance stays fixed. The primary estimate uses the lowest compatible observed rate; the range shows sensitivity across all compatible published rates and does not imply capacity.

Lowest compatible compute cost$0.2888
Compatible published-rate range$0.2888
Serial time on benchmark-equivalent system0.0046 h
Cost basisper 1M tokens

Inspectable formula

cost/1M units = 1,000,000 ÷ measured units/second ÷ 3,600 × accelerator count × USD/GPU-hour

1 × $0.2888 = $0.2888

1 × 0.0046 h = 0.0046 h serial benchmark-equivalent system time

Validate the workload boundary

Check this estimate against your workload and quotes

A human review tests benchmark transferability, workload assumptions, capacity evidence and non-compute costs while keeping unknowns explicit.

Measured boundary

Exactly what was benchmarked

Nebius B300 n1 (8x B300-SXM-270GB, TensorRT) reported 60,413.4 Tokens/s for deepseek-r1 in the Server scenario. The quality target was exact_match: 81.58614402917047 TOKENS_PER_SAMPLE: 3721.7894257064722. Software was TensorRT 10.14, CUDA 13.1, cuDNN 9.17, TensorRT-LLM feat/1.2-mlpinf, NVIDIA Dynamo mlperf-v6.0-dynamo-v0.8.0, vLLM CentML:mlperf-inf-mm-q3vl-v6.0. This page does not extrapolate to a different model, implementation, accelerator count or quality target.

Benchmark evidenceMLPerf commit 4d3916ac9cf474b679cdfcf492d43a0559418ad1. Aggregation: official reported result; 1 published run.Open pinned result ↗
Price evidence3 exact provider/model/market price matches; lowest published normalized rate $7.8500 per GPU-hour, valid through 2026-09-10T22:35:29Z.Open official price source ↗
Join ruleProvider, accelerator family, GPU count context, market type, USD currency and per-GPU-hour unit must remain compatible. Missing evidence fails closed.Read the full methodology

Interpretation guardrails

Use this estimate inside its evidence boundary

Comparable workload

The cost applies to this named benchmark, scenario, software stack, quality target and published accelerator system. A production workload can produce different throughput or time.

Serial time only

The displayed time scales benchmark-equivalent runs or output batches serially on the measured system. It excludes queueing, provisioning and any parallel execution strategy.

Compute line only

Storage, networking, data preparation, queueing, engineering time, failed jobs, support and taxes are excluded. Add them in a complete deployment budget.

Capacity separate

The price source proves a published rate, not immediate capacity. Availability remains a separate signal and must be confirmed before scheduling work.

  • The result applies only to the published model, software, scenario and 8-GPU system.
  • The current Nebius normalized per-GPU rate is joined separately and multiplied by the benchmark GPU count; identical topology and simultaneous capacity are not asserted.
  • Storage, data transfer, orchestration, queueing, support, taxes and discounts are excluded.
  • A benchmark result is not a guarantee of performance for a different workload or configuration.

Related evidence

Change one decision variable

Move to another scenario or inspect the exact provider and accelerator pricing behind this result.

workload cost

gpt-oss 20b pretraining time to train to quality target gpu cost

3 published official-source rows. Open the page to inspect its own price range and validity deadline.

Open decision page

workload cost

llama 3.1 8b pretraining time to train to quality target gpu cost

3 published official-source rows. Open the page to inspect its own price range and validity deadline.

Open decision page

workload cost

deepseek-r1 offline gpu cost

3 published official-source rows. Open the page to inspect its own price range and validity deadline.

Open decision page

workload cost

gpt-oss-120b offline gpu cost

3 published official-source rows. Open the page to inspect its own price range and validity deadline.

Open decision page

workload cost

gpt-oss-120b server gpu cost

3 published official-source rows. Open the page to inspect its own price range and validity deadline.

Open decision page

workload cost

qwen3-vl-235b-a22b offline gpu cost

3 published official-source rows. Open the page to inspect its own price range and validity deadline.

Open decision page