You are about to spend a quarter negotiating 15% off a GPU hour, and lose 2,000% on the hour you never used.
Here is the verdict before the analysis. Buy on-demand from a neocloud until you can prove a duty cycle above roughly 50%, then commit — and never sign a commitment whose break-even utilisation is higher than the utilisation you have actually measured. For most enterprise AI teams that means Nebius or Together AI on the meter for the first two quarters, a 3-to-6-month committed rate after that, and spot reserved strictly for checkpointed training. Google Cloud's A4 is the loser on this table and it is not close: the highest list rate of the nine, and the only one you cannot buy on demand at all.
What 64 B200s Cost for a Year, by How You Buy Them
Normalised to one workload for every row: 64 NVIDIA B200 GPUs — eight nodes of eight — held for twelve months. That is 560,640 GPU-hours at a 100% duty cycle. Same silicon, same node topology. The only variable is the purchasing instrument. All rates were read off each vendor's live pricing page, or a public price tracker where the vendor publishes none, on 8 September 2026.
| Purchase mode | Product | $/GPU-hr | 12 months, 64 GPUs | Break-even duty | Can you exit? |
|---|---|---|---|---|---|
| Committed term | CoreWeave Reserved (up to 60% off) | ~$3.44 | $1.93M | 40% | Contract term |
| Preemptible | Together AI preemptible B200 | $4.09 | $2.29M | Tolerates 50% waste | Instantly — against you |
| Reserved, 91–180d | Together AI Reserved Capacity | $6.79 | $3.81M | 83% | At term end |
| On-demand | Nebius HGX B200 | $7.15 | $4.01M | n/a | Hourly |
| On-demand | CoreWeave HGX B200 | $8.60 | $4.82M | n/a | Hourly |
| Reserved cluster | Lambda 1-Click Cluster, 64 GPU | $9.36 | $5.25M | n/a | 2-week minimum |
| Reserved block | AWS EC2 Capacity Blocks, p6-b200 | $12.36 | $6.93M | 87% | No — non-cancellable |
| On-demand | AWS p6-b200.48xlarge | $14.24 | $7.98M | n/a | Hourly |
| Reservation-only | Google Cloud A4 (a4-highgpu-8g) | $16.11 | $9.03M | n/a | No on-demand option |
Sources, in order. CoreWeave lists HGX B200 at $68.80 per 8-GPU node — $8.60 per GPU — and advertises Reserved Compute Capacity at "up to 60% discounts over On-Demand prices for committed usage"; $3.44 is that ceiling applied, and it is the only estimated figure in the table. Together AI publishes the rare public ladder: $8.19 on-demand, $7.99 at 7–30 days, $7.79 at 31–90 days, $6.79 at 91–180 days, and $4.09 preemptible. Nebius lists HGX B200 at $7.15 on-demand and $3.95 preemptible. Lambda charges $9.86 per GPU at 16 GPUs, $9.36 at 64 and $8.87 at 256+. AWS prices Capacity Blocks for p6-b200 at $12.355 per accelerator-hour, against an on-demand p6-b200.48xlarge rate of $113.933 per instance-hour, which is $14.242 per GPU. Google's a4-highgpu-8g carries eight B200s at roughly $16.11 per GPU-hour — an on-demand-equivalent rate, since Google publishes no on-demand price for the machine type.
The spread is 4.7x on identical hardware. None of it is a performance difference. Almost all of it is the difference between a rate card and a contract.
The Number That Decides This Is Not on Any Price List
Every row above prices a GPU-hour you hold. Your finance model cares about a GPU-hour you use, and those are not the same number — they are not even the same order of magnitude.
Cast AI's 2026 State of Kubernetes Optimization Report measured roughly 23,000 production clusters across AWS, Azure and Google Cloud and found GPU utilisation averaging 5%, with CPU at 8% and memory at 20%, measured before any optimisation or autoscaling. Cast AI's own framing of the same finding is blunter: "95% of GPU capacity is doing nothing".
Read the footnotes before you read that number into your own business case, because three things about it are load-bearing. Cast AI sells the automation that fixes the problem it measured — the finding is credible at that scale but it is not disinterested. The metric is defined as the "percentage of total provisioned GPU compute cycles that produce useful output across a 24-hour period," which is cycle efficiency, not hours-held-versus-hours-busy. And the report explicitly excludes the population most like the cluster priced above: "clusters from AI Labs are not included in the statistics as AI Labs usually have a very intense use of GPUs for training and inference." Dedicated training fleets run far better — published model-FLOPs-utilisation figures for production LLM training sit around 38–43% for Llama 3.1 and 55% for ByteDance's MegaScale.
So treat 5% as the enterprise-Kubernetes floor it is, not as your number — and note that the two failure modes it blends have different cures. Idle hours are a procurement problem, and buying by the hour fixes them. Idle cycles inside hours you genuinely need the GPU are an engineering problem, and no purchasing instrument touches them.
Take the floor literally against the table anyway, because it sets the scale of the risk. At a 5% duty cycle, those 64 GPUs do 28,032 hours of real work in a year. Buy them on the committed rate and you pay $1,928,602 for that work — $68.80 per useful GPU-hour, against a $3.44 line item on the invoice. Buy the identical 28,032 hours on Nebius's on-demand meter and you pay $200,429.
A duty cycle is the only variable in GPU procurement that moves cost by 20x. The discount ladder moves it by 15–60%. Teams spend a quarter on the second one.
Here is the same 64 GPUs across duty cycles, comparing the deepest commitment available against paying the neocloud meter for exactly the hours you run:
| Actual duty cycle | Useful GPU-hours | Pay-per-use at $7.15 | 12-month commit at $3.44 | Cheaper |
|---|---|---|---|---|
| 5% (the enterprise-Kubernetes floor) | 28,032 | $200,429 | $1,928,602 | Pay-per-use, by 9.6x |
| 25% | 140,160 | $1,002,144 | $1,928,602 | Pay-per-use |
| 48% | 269,107 | $1,924,116 | $1,928,602 | Dead heat |
| 70% | 392,448 | $2,806,003 | $1,928,602 | Commit |
| 95% | 532,608 | $3,808,147 | $1,928,602 | Commit, by 2.0x |
Your break-even is 48%. That is the whole decision, and it is a measurement, not a negotiation. Nobody at the vendor can tell you what it is. Before you sign anything, instrument the fleet and get thirty days of allocated-versus-busy GPU-seconds on paper — the same discipline that makes a seat count meaningful only when you reconcile it against weekly actives.
Why the Shallowest Discount Is the Most Dangerous
Break-even utilisation is just the discounted rate divided by the on-demand rate. Run it across the instruments and an uncomfortable pattern falls out.
| Instrument | Discount vs on-demand | Duty cycle needed to break even |
|---|---|---|
| AWS Capacity Blocks (p6-b200) | 13% | 86.8% |
| Together AI, 91–180 days | 17% | 82.9% |
| Nebius reserved cluster ("up to 35% less") | 35% | 65.0% |
| CoreWeave Reserved ("up to 60%") | 60% | 40.0% |
The deeper the discount, the more forgiving the commitment. A 13% saving that bills you for every hour in the window demands near-perfect utilisation to pay for itself; a 60% saving pays off at four days a week. This inverts the instinct that a small commitment is a small risk.
AWS Capacity Blocks are the sharp end of it. The documentation is unambiguous: "Capacity Block cancellations aren't allowed", a block runs one to 14 days or a multiple of seven up to 182 days, you can reserve a start time only up to eight weeks in the future, each block holds up to 64 instances with 256 across all blocks, and the block ends at 11:30 UTC on its final day with instance termination beginning at 11:00. Covering a year means at least two consecutive purchases, neither of which you can cancel, each requiring 87% duty to beat simply paying by the hour.
That is not a criticism of the product. Capacity Blocks are a capacity guarantee wearing a discount's clothes, and in a market where SemiAnalysis reports that "On-Demand GPU rental capacity is sold out across all GPU types", a guarantee is worth paying for. Just do not book one as a savings measure.
Who should not buy Capacity Blocks: anyone whose duty cycle is under 85%, anyone who needs more than eight weeks of forward visibility, and anyone whose project might be cancelled — because the block will not be.
The Charges That Are Not in the Headline Rate
Three line items reliably arrive after the rate is agreed, and one of them is larger than the discount you fought for.
The cluster premium. Lambda's B200 headline is $6.69 per GPU-hour for an 8-GPU node. The moment you need those nodes wired together with InfiniBand, the price is $9.36 at 64 GPUs — a 40% premium, larger than most reserved discounts on the market. If your workload spans nodes, the standalone rate was never your rate. Ask every vendor for the clustered figure at your actual node count before you compare anything.
Storage. Amazon FSx for Lustre persistent SSD lists at $0.145 per GB-month in US East. A hundred terabytes of datasets and checkpoints is $178,176 a year — $0.318 on every one of your 560,640 GPU-hours, before a token moves. Nebius's shared filesystem at $0.08/GB-month works out to $0.175 per GPU-hour, and CoreWeave's distributed file storage at $0.070/GB-month to $0.153. FSx Intelligent-Tiering drops frequent-access storage to $0.023/GB-month but bills throughput separately at $0.52 per MBps-month, which is a different bill, not a smaller one.
Egress. Nebius publishes free network egress and $0.015/GB from object storage; CoreWeave states data transfer between CoreWeave and the internet is free; Lambda advertises no egress fees. The hyperscalers meter it. If your architecture pulls training data in from one cloud and serves inference from another, that is a recurring transfer bill nobody quoted you — and it is the same trap that shows up when you cost a production RAG pipeline end to end instead of pricing the model call.
Spot Is a Market Price, Not a Rate Card
A spot discount is the market clearing price for capacity nobody else wanted this hour. It is not a term you have been granted, and in a shortage it goes away.
The published discounts are real and large: Together AI's preemptible B200 is $4.09 against $8.19 on-demand (50%), Nebius's is $3.95 against $7.15 (45%), CoreWeave's is $34.11 per node against $68.80 (50%), and AWS spot on p6-b200.48xlarge was $42.063 per instance-hour against $113.933 on-demand (63%). The useful way to read a spot discount is as a waste budget: at 50% off you can lose half of every job to preemption and still be level with on-demand.
Then read the other side. SemiAnalysis reports customers "fighting to pay $14/hr/GPU for p6-b200 spot instances in AWS" — essentially the on-demand rate — and by March it had become "increasingly impossible to find any H100s, H200s or B200 rental capacity for any term." Azure's newest silicon tells the same story from the tracker side: ND96isr H200 v5 shows a spot price identical to its $110.240 on-demand rate, with no 1-year or 3-year reserved rate published at all.
The operational limit is harder than the price one. AWS issues a Spot interruption notice two minutes before it stops or terminates the instance, delivered on a best-effort basis, and its Spot Instance Advisor buckets interruption frequency into ranges up to ">20%".
- Spot is right for checkpointed training, hyperparameter sweeps, batch evaluation, and anything where a lost hour costs an hour. Set your checkpoint interval so that the expected rework is smaller than the discount.
- Spot is wrong for anything behind a latency SLA. Two minutes is not enough to drain connections and reschedule under load, and preemptions correlate — the capacity reclamation that takes one node tends to take the family.
- The honest middle is Google's flex-start, which queues a job for capacity and then runs it up to seven days without preemption at up to 53% off. It is a batch queue, not a server, and priced accordingly. Calendar mode books up to 90 days and a maximum of 80 VMs.
Google's A4 Is the Loser, and the Reason Is Structural
Two things are true about Google Cloud's A4 at once, and together they make it the row to avoid.
It carries the highest list rate in the table at roughly $16.11 per GPU-hour — 2.25x Nebius for the same B200. And it cannot be launched on demand at all. Google's own documentation states that "when provisioning A4 machine types, you must reserve capacity to create instances or clusters, use Spot VMs, use Flex-start VMs, or create a resize request in a MIG," and for A4X (GB200) it is stricter still: "you must reserve capacity to create instances and clusters".
A buyer therefore gets hyperscaler pricing and neocloud-style availability planning: you hold a reservation and pay for it, you take preemptible Spot, or you queue. The escape valve that justifies a hyperscaler premium — burst now, sort the commitment out later — is the one thing on offer.
Who should still pick it: teams whose data genuinely cannot leave Google Cloud, or who are deep enough into BigQuery and Vertex that the egress and rework cost of moving exceeds the premium. That is a real constraint and it is worth real money. It is just not a price decision, and it should not be argued as one.
Who should not pick each of the others:
- Nebius, Lambda, CoreWeave and Together AI — not if a regulator requires your workload inside a named hyperscaler region, and not if you need a support organisation that answers a 512-GPU job failing at 3am. The GPU cloud comparison covers where these tiers actually differ on reliability.
- AWS Capacity Blocks — not under 85% duty, not beyond eight weeks of visibility, not if you might cancel.
- Azure Reserved VM Instances — not if your cancellable exposure needs to exceed $50,000 a year. See below.
- Any 12-month commitment at all — not before product-market fit. H100 on-demand rates fell from a mid-2024 hyperscaler peak of about $9.34 per GPU-hour to roughly $6.26, while marketplace providers dropped to $1.92–$2.00 — a third off at the hyperscalers and about 80% off in the open market. A rate you fix today is a floor you keep paying when the market moves under you.
How to Write a GPU Contract You Can Actually Leave
Exit terms are where these deals are won, and they are almost never in the quote.
Azure caps how much you can cancel. Microsoft's reservation policy is explicit: the total cancelled commitment "can't exceed 50,000 USD in a 12-month rolling window for a billing profile or single enrollment". Microsoft's own worked example is the one to read to your CFO: a three-year reservation at $3,000 a month is a $108,000 commitment, and you cannot cancel it at all until you have already spent $58,000 of it. There is no early termination fee today, but the same document reserves the right to a 12% one.
Azure's exchange escape hatch is closing. From 1 February 2027, reservations purchased after that date are not eligible for exchange where the service is covered by savings plans — which includes Azure Virtual Machines. Reservations bought before that date retain the right to one final exchange. If you are buying Azure GPU reservations in the next few months, that deadline is a term of your deal whether or not anyone mentions it.
A spend commitment is not a capacity commitment. AWS Compute Savings Plans discount EC2 usage "regardless of instance family, instance size, OS, tenancy, or AWS Region" for one or three years — accelerated instances included. What they do not buy is a GPU. In a market where on-demand capacity is reported sold out, a discount you cannot exercise is worth exactly zero, and you will still owe the hourly commitment. Buy the discount and the capacity as two separate, separately-priced things, and know which one you actually have.
Ask for these five clauses by name:
- A published rate ladder, or a written one. Together AI publishes its 7-30 / 31-90 / 91-180 day tiers. Everyone else quotes. Put a rival's public ladder on the table as your anchor.
- A downward re-open on list price. If the vendor's public on-demand rate for your SKU falls more than 15%, your committed rate re-prices. Vendors resist this and it is the single most valuable clause in the document.
- A substitution right to next-generation silicon at the same dollars-per-hour. You are buying throughput, not a part number, and the part number changes every eighteen months.
- Termination for convenience with a stated fee. A number you can model — 12%, 20% — beats a cap you cannot cross and beats "not permitted" outright.
- Storage, egress and cluster networking quoted in the same document, per GB-month and per GB, for the term. Otherwise the 40% cluster premium arrives after signature.
Treat the commitment term itself as the negotiable item, not just the rate — the same discipline that applies when a vendor sets a date for capacity but not a price.
What to Do in the Next 90 Days
This week: Pull thirty days of GPU allocation-versus-busy time from your existing fleet. One number: busy GPU-seconds divided by allocated GPU-seconds. If you cannot produce it, you cannot evaluate a commitment, and that is the finding.
This month: Rebuild your capacity model on useful GPU-hours rather than held ones. Divide every quoted rate by your measured duty cycle and re-run the comparison — the ranking will change. Then add the clustered rate at your real node count, storage at $0.07–$0.145 per GB-month, and egress.
Before your next renewal: Compute the break-even duty cycle for every commitment on the table — discounted rate divided by on-demand rate — and refuse any instrument whose break-even is above your measured number. Get the five clauses above into the redline. And check whether your Azure GPU reservations were bought before 1 February 2027, because their exchange right is an asset with an expiry date.
Structurally: utilisation is an engineering problem before it is a procurement one. Multi-tenancy, right-sized inference runtimes, honest memory sizing for mixture-of-experts models, and batch sizing that reflects real context lengths move the denominator far harder than any discount moves the numerator. Moving from 25% to 50% duty is worth more than every commitment discount in this article, combined.
The Bottom Line
The GPU market has spent two years teaching buyers to think about scarcity, and scarcity thinking says: lock it in, lock it in long, take the discount. SemiAnalysis's index says that instinct is being priced — H100 one-year rental contracts rose from $1.70 to $2.35 per GPU-hour between October 2025 and March 2026, up almost 40% in five months. That is a term rate climbing into a shortage, not away from one: the same report has on-demand capacity sold out "despite recent price hikes," and spot bid up toward the on-demand rate. Signing a long commitment at the top of that curve locks in a peak, and committed capacity is the expensive end of the market for anyone who does not fill it.
This is the same shape as every previous infrastructure cycle. Enterprises bought three-year server refreshes at 40% off list and ran them at 12% CPU. They bought reserved instances in 2015 and wrote them off in 2017. The instrument changes; the error does not. It is always the same error: treating capacity you hold as capacity you use.
The rate on the invoice is the only number in this business that is easy to find, easy to compare, and almost never the one that decides anything.
Your duty cycle is the price. Everything else is a rounding error with a salesperson attached.
Continue Reading
- GPU Clouds Compared: Nebius on Price, CoreWeave at 3AM
- Microsoft's GPUs Sit Unplugged. Buy Delivery, Not Capex.
- Equinix Set a Date, Not a Price. Cap Your Commit Term.
- NVIDIA Alternatives for Inference: Only Trainium Pays Off
- Agentic AI Pricing: Don't Buy Consumption Without a Cap
- What RAG Actually Costs: $1,308 a Month at 10M Tokens/Day
- North Carolina Taxed the Power. Inference Pays.
- GPT-5.6 Sol Is $20 Until Nov 21. Budget Both Rates.
