workload cost2 rates
deepseek-r1 offline gpu cost
$7.850–$7.850 per GPU-hour · 2 regions.
Pinned benchmark result: 69,318.9 Tokens/s · Offline.
Completed-workload cost range: $0.2517 per 1M tokens on the pinned 8× NVIDIA B300-SXM-270GB system across the current compatible observed rates.
Serial benchmark-equivalent time: 14.43 s per 1M tokens serially on the pinned 8× NVIDIA B300-SXM-270GB system.
Deployability evidence: 2/2 rows (100%). This is not a live-stock claim.
workload cost2 rates
deepseek-r1 server gpu cost
$7.850–$7.850 per GPU-hour · 2 regions.
Pinned benchmark result: 60,413.4 Tokens/s · Server.
Completed-workload cost range: $0.2888 per 1M tokens on the pinned 8× NVIDIA B300-SXM-270GB system across the current compatible observed rates.
Serial benchmark-equivalent time: 16.55 s per 1M tokens serially on the pinned 8× NVIDIA B300-SXM-270GB system.
Deployability evidence: 2/2 rows (100%). This is not a live-stock claim.
workload cost2 rates
gpt-oss-120b offline gpu cost
$7.850–$7.850 per GPU-hour · 2 regions.
Pinned benchmark result: 106,885 Tokens/s · Offline.
Completed-workload cost range: $0.1632 per 1M tokens on the pinned 8× NVIDIA B300-SXM-270GB system across the current compatible observed rates.
Serial benchmark-equivalent time: 9.36 s per 1M tokens serially on the pinned 8× NVIDIA B300-SXM-270GB system.
Deployability evidence: 2/2 rows (100%). This is not a live-stock claim.
workload cost2 rates
gpt-oss-120b server gpu cost
$7.850–$7.850 per GPU-hour · 2 regions.
Pinned benchmark result: 100,437 Tokens/s · Server.
Completed-workload cost range: $0.1737 per 1M tokens on the pinned 8× NVIDIA B300-SXM-270GB system across the current compatible observed rates.
Serial benchmark-equivalent time: 9.96 s per 1M tokens serially on the pinned 8× NVIDIA B300-SXM-270GB system.
Deployability evidence: 2/2 rows (100%). This is not a live-stock claim.