A one-year GPU reservation only saves money if you keep the GPUs busy for 65-80% of every hour you hold them, and most teams being asked to sign one have never measured that number.
The verdict: commit only to the GPUs your fleet has kept busy in every hour of the last 90 days, on a term of six months or less, and put checkpointed training on spot or preemptible capacity. Sign twelve months only when the discount is at least 40% and the contract lets you swap to next-generation silicon at the same dollars per hour. For most teams buying 8 to 64 H100s, Together AI's published 91-180 day tier is the default commitment, because it is the only ladder you can price without a sales call. Azure's ND H100 v5 is the loser on this table: an on-demand rate nearly twice the next-highest, and a $50,000-a-year cap on cancelling reservations.
All prices below are per GPU-hour for NVIDIA H100 and were checked on each vendor's live pricing page, or a public tracker where the vendor publishes per-instance rates, on October 2, 2026.
| Provider | On-demand | Best published commitment | Break-even utilisation | Spot / preemptible | Skip it if |
|---|---|---|---|---|---|
| Together AI | $3.99 | $3.19 (91-180 days) | 80% | $1.99 | You need a 12-month rate on the page |
| Nebius | $4.50 | up to 35% off (~$2.93), quote only | 65% | from $0.79 | You can't use a multi-month cluster |
| CoreWeave | $6.16 | up to 60% off (~$2.46), quote only | 40% | $2.46 | You buy fewer than a few hundred GPUs |
| Lambda | $3.99-$4.29 | clusters $5.85-$6.16, 1-year quote only | n/a | none published | You need spot at all |
| AWS p5.48xlarge | $6.88 | Capacity Blocks $5.19 ($5.97 from Oct 7) | 75% (87% from Oct 7) | $2.99 | Your duty cycle is under 75% |
| Azure ND96isr H100 v5 | $12.29 | none published | n/a | $2.27 | You are not burning down an Azure commit |
Break-even utilisation is the committed rate divided by the on-demand rate: the share of held hours you must actually use before the commitment beats paying the meter only for the hours you run.
What Each Provider Charges to Commit
The published discounts for committing range from 20% to 60%, and only Together AI and AWS let you see the rate before you talk to sales.
Together AI prints the whole ladder for HGX H100: $3.99 on-demand, $3.69 for 7-30 days, $3.45 for 31-90 days, $3.19 for 91-180 days, "contact" for 181 days and up, and $1.99 preemptible. That makes it the anchor for every other negotiation. Its weakness is the same as its strength: the deepest published tier is a 20% discount, which is the shallowest commitment discount on the table, and a shallow discount demands near-perfect utilisation.
Nebius lists HGX H100 on-demand at $4.50 per GPU-hour under a column headed "GPU-hour (Effective October 1, 2026)", up from $3.85. That is a 17% increase on H100 that took effect yesterday, and B200 moved from $7.15 to $8.50. Nebius advertises that you can "pay up to 35% less than on-demand rates by reserving large-scale clusters for multiple months" and sends you to sales for the number. Its preemptible H100 is listed "from $0.79", and "from" is doing real work in that sentence; treat it as a floor you may not see. Nebius has a tool page here.
CoreWeave prices by the 8-GPU node: HGX H100 at $49.24 per hour on-demand ($6.16 per GPU) and $19.71 spot ($2.46 per GPU). It states it "offers up to 60% discounts over our On-Demand prices for committed usage." The ~$2.46 in the table is that ceiling applied to list, and it is an estimate: nobody buying 32 GPUs should expect the ceiling. Note the coincidence anyway. CoreWeave's best possible committed rate and its spot rate are the same number, which tells you what the vendor thinks a twelve-month promise is worth relative to capacity it can reclaim.
Lambda lists H100 SXM on-demand at $3.99-$4.29 and 1-Click Clusters, sold for two weeks to one year, at $6.16 per GPU for 16 GPUs and $5.85 for 64. Its one-year reserved column is a dash and a "contact us for reserved capacity at our lowest prices." Lambda publishes no spot product. Lambda has a tool page here.
AWS lists p5.48xlarge (8x H100) at $55.04 per instance-hour on-demand and $23.93 spot in us-east-1, which is $6.88 and $2.99 per GPU. The Capacity Blocks pricing page shows p5 at $5.191 per accelerator-hour in both US East regions, under a notice that rates are "dynamic and are updated periodically based on supply and demand." The same notice sets P5 at $5.970 from October 7, 2026, a 15% rise that moves the break-even against on-demand from 75% to 87%. A block is charged at the rate on the day you buy it. AWS's documentation is blunt that "Capacity Block cancellations aren't allowed".
Azure's ND96isr H100 v5 tracks at $98.32 per hour pay-as-you-go and $18.17 spot in East US: $12.29 and $2.27 per GPU, with no 1-year or 3-year reserved rate shown on that tracker.
Your Break-Even Is a Measurement, and the Penalty Is Lopsided
A reservation beats on-demand only when your measured utilisation is higher than the committed rate divided by the on-demand rate, and missing that line costs more than clearing it saves.
Normalise everything to one workload: 32 H100s, held for 12 months, which is 280,320 GPU-hours. At Together AI's rates, the 91-180 day tier bought for two consecutive terms costs $894,221 for the year. Paying $3.99 on-demand only for the hours you use costs:
| Measured utilisation | GPU-hours used | On-demand, pay per use | Reserved at $3.19 | Result |
|---|---|---|---|---|
| 50% | 140,160 | $559,238 | $894,221 | Reservation loses $334,983 |
| 60% | 168,192 | $671,086 | $894,221 | Reservation loses $223,135 |
| 80% | 224,256 | $894,781 | $894,221 | Dead heat |
| 90% | 252,288 | $1,006,629 | $894,221 | Reservation saves $112,408 |
| 100% | 280,320 | $1,118,477 | $894,221 | Reservation saves $224,256 |
Read the rows at 60% and 100%. Twenty points under break-even costs as much as twenty points over it saves, and there is no 120% row. The upside is capped at full use; the downside runs all the way to an idle cluster. A shallow discount on a long term is the riskiest instrument a buyer can sign, because it combines the highest break-even with the most hours exposed to it.
The case for committing anyway is real and you should weigh it. SemiAnalysis reported this year that "On-Demand GPU rental capacity is sold out across all GPU types", and its H100 contract index showed one-year H100 contracts at $2.10-$2.70 per GPU-hour as of April 2026, after one-year rentals rose from $1.70 in October 2025 to about $2.35 in March 2026. Nebius's 17% on-demand increase this week points the same way. If on-demand capacity is not there when you need it, the pay-per-use column above is a fiction, and a reservation is buying a guarantee of capacity. That is a legitimate purchase. Price it as insurance, size it to the load you cannot afford to lose, and stop there.
What Happens When the Next Generation Ships Mid-Term
A twelve-month H100 or B200 commitment signed today runs through the Vera Rubin ramp, and the last time H100 supply grew, its hyperscaler price fell by about a third.
NVIDIA says Rubin is in full production, with Rubin-based products available from partners in the second half of 2026, and names AWS, Google Cloud, Microsoft, OCI, CoreWeave, Lambda, Nebius and Nscale among the first clouds to deploy it. A reservation signed in October 2026 expires in October 2027, inside that window.
The H100 shows what a supply wave does to a part you have locked. Silicon Data's history puts the hyperscaler median at $9.34 per GPU-hour in the second half of 2024, falling to $6.26 in the second half of 2025, while marketplace rates fell to $1.92-$2.00. Silicon Data attributes the fall to more NVIDIA supply, the rise of marketplaces and hyperscaler response, and one vendor move was part of it: AWS cut P5 prices by up to 45%, effective June 1, 2025 for on-demand and for Savings Plans bought after June 4. A P5 Savings Plan bought on June 3 would have locked in the pre-cut rate for its whole term.
Run that forward. Say you commit at Nebius's ceiling discount, about $2.93 against today's $4.50, a 65% break-even. If H100 on-demand falls 30% six months in, to $3.15, your break-even for the remaining half of the term jumps to 93%. If instead prices keep rising the way they did this week, the same commitment looks smart. You do not know which happens, and nor does the vendor, which is why the term length is worth more of your negotiating time than the rate.
Two contract terms change this risk. A substitution right lets you move the same dollars per hour onto newer silicon when it arrives; without it, you hold H100 hours while your competitors rent Rubin. A downward re-open reprices your commitment if the vendor's own list rate for your SKU falls past a stated threshold. Vendors resist both. If you cannot get either, the term should not exceed six months, which is exactly where Together AI's published ladder stops.
What Spot Interruptions Cost a Checkpointed Job
Spot loses you a share of every hour equal to your interruption rate times the work since the last checkpoint plus restart time, and for most training jobs that share is smaller than the spot discount.
The notice periods are short. AWS sends a Spot interruption notice two minutes before it stops or terminates the instance, and its Spot Instance Advisor buckets interruption frequency at "< 5%, 5-10%, 10-15%, 15-20% and >20%," measured over the trailing month. DoiT's tracker currently shows p5.48xlarge in us-east-1 in the "<5%" bucket. Azure gives you 30 seconds' notice, no SLA, and quotes eviction rates per hour, so a 10% rate means a 10% chance of eviction within the next hour based on the last seven days.
Two facts make the arithmetic kinder than it looks. First, you already pay for checkpointing on reserved hardware. Meta's Llama 3 paper recorded 419 unexpected interruptions over 54 days of pre-training on up to 16,384 H100s, roughly one every three hours, about 78% of them hardware. Any serious training job on any purchase model needs checkpoints, so spot adds to a cost you already carry. Second, checkpoints got cheap. PyTorch's distributed asynchronous checkpointing cut a 7B model's checkpoint downtime from an average of 148.8 seconds to 6.3 seconds, because the write to storage happens on CPU threads while the GPUs keep training.
The working formula for the share of each hour you lose:
lost share ≈ (nodes × hourly eviction rate) × (checkpoint interval ÷ 2 + restart time) + (checkpoint stall ÷ checkpoint interval)
Multiply by node count because a multi-node job dies when any one node goes. For a 4-node, 32-GPU job checkpointing every 30 minutes with a 20-minute restart:
| Per-node eviction rate (per hour) | Lost share, async checkpoint (6.3s) | Lost share, blocking checkpoint (148.8s) | Effective Together preemptible rate (async) |
|---|---|---|---|
| 2% | 5.0% | 12.9% | $2.09 |
| 5% | 12.0% | 19.9% | $2.26 |
| 10% | 23.7% | 31.6% | $2.61 |
Even at a 10% hourly eviction rate per node, Together AI's $1.99 preemptible H100 works out cheaper than its $3.19 reserved rate at full utilisation. What usually breaks a spot plan is availability: in a shortage there may be no spot capacity for your SKU at all, which is why the plan below survives that case.
SkyPilot is the practical tool for this. Its managed jobs detect the preemption, provision a replacement, and restart the job so your code can resume from the last checkpoint, across clouds, so a preemption on one provider can restart on another's capacity.
Spot is wrong for anything behind a latency SLA. Thirty seconds or two minutes is not enough to drain an inference replica under load.
A Hedged Split That Survives Both Price Cases
Split the fleet into three layers: commit the floor, run training on preemptible, and pay on-demand for the burst. It costs less than a full reservation whether prices rise, fall, or spot dries up.
Take the same 32-GPU fleet with a realistic workload mix: 12 GPUs serving inference around the clock, 12 GPUs training for 70% of hours, and 8 GPUs of burst used 25% of the time. That is 70% average utilisation, ten points under the Together AI break-even. All at Together AI's published rates:
| Scenario | Committed floor (12 GPUs, $3.19) | Training (73,584 GPU-hrs) | Burst (17,520 GPU-hrs) | Year total |
|---|---|---|---|---|
| Reserve all 32 for the year | $894,221 | |||
| Split, prices flat, 12% preemption loss | $335,333 | $166,402 (preemptible) | $69,905 (on-demand) | $571,640 |
| Split, uncommitted prices up 17% | $335,333 | $194,682 | $81,783 | $611,798 |
| Split, no spot capacity, training on-demand | $335,333 | $293,600 | $69,905 | $698,838 |
| Split, no spot and prices up 17% | $335,333 | $343,490 | $81,783 | $760,606 |
The 17% is Nebius's actual H100 move this week, applied to everything you did not lock. Even the worst row, no spot capacity and a price rise together, comes in $133,615 under the full reservation, and if prices fall when Rubin arrives, the split collects the drop on 20 of its 32 GPUs.
The committed floor is the minimum number of GPUs that were busy in every single hour of the last 90 days. An average overstates it. Pull it from your scheduler or your GPU chargeback data. If you run Kubernetes, fix quota stranding before you size the floor, because static quotas idle GPUs that a scheduler could share and a commitment locks the waste in.
Who Should Not Pick Each Option
Every option on the table is the right buy for someone and an expensive mistake for someone else.
- Together AI reserved. Skip it if you need a published 12-month rate or anything past its published 180-day tier without negotiating, or if your floor is above a few hundred GPUs, where a negotiated neocloud rate should beat 20% off.
- Nebius. Skip it if you need a price you can model before signing. The commitment rate is quote-only, the preemptible rate is a "from" figure, and on-demand moved 17% this week.
- CoreWeave. Skip it at small scale. The 60% headline discount is a ceiling for committed usage at volume, and its own on-demand H100 rate is the third-highest on the table. Our GPU cloud comparison covers where it earns that premium on operations.
- Lambda. Skip it if your plan depends on spot, which it does not publish. Its cluster rates also run roughly 35-55% above its single-instance rates, so get the clustered rate at your node count.
- AWS Capacity Blocks. Skip them under 75% utilisation (87% for blocks bought from October 7) and for any project that might be cancelled, because the block cannot be. Buy one for a guaranteed training window, and book it as a capacity cost.
- Azure ND H100 v5. Skip it unless you are burning down an existing Azure commitment or your data cannot leave Azure. At $12.29 on-demand it costs 3.1 times Together AI for the same GPU. Its exit terms are tight too: Microsoft caps total cancelled reservation commitment at $50,000 in a 12-month rolling window, reserves the right to a future 12% early termination fee, and from February 1, 2027, VM reservations bought after that date lose exchange rights.
Azure is the loser here because it is the row where you pay the most per GPU-hour, with no published reserved rate to bring that down. Its spot rate of $2.27 is competitive, which makes it a reasonable place for evictable batch work and the wrong place to sign a reservation.
The Decision: Five Questions That Predict Regret
Teams regret GPU commitments for five reasons, and you can check each one before signing.
- Is the committed count at or below your 90-day floor? If you are committing to average or peak, you are committing to idle hours.
- Is your measured utilisation above committed rate ÷ on-demand rate? At Together AI that is 80%, at AWS Capacity Blocks 75% (87% from October 7), at a 35% neocloud discount 65%.
- Does the term end before the next generation is broadly rentable? With Rubin shipping to clouds in the second half of 2026, a 12-month term needs a substitution clause.
- Can you leave? AWS Capacity Blocks: no. Azure reservations: up to $50,000 a year. A neocloud contract: whatever you negotiated, so negotiate it.
- Does spot capacity exist for your SKU and region? Check the AWS Spot Instance Advisor bucket and Azure's seven-day eviction history before you plan on it, and assume it can vanish.
What changes the answer: a vendor offering 40% or more for twelve months with a substitution right and a re-open clause makes the full-year commitment reasonable for your floor. A workload with no tolerance for interruption and no on-demand fallback in your region makes a capacity reservation worth its premium. Neither one justifies committing above the floor.
What to Do Before You Sign
This Week: Pull 90 days of hourly GPU allocation from your scheduler. Compute two numbers: the floor (minimum GPUs busy in every hour) and the average utilisation. Divide each quoted committed rate by its on-demand rate and compare it to the second number.
This Month: Move one training job to preemptible capacity with asynchronous checkpointing and SkyPilot managed jobs. Measure your actual lost share for two weeks instead of trusting the formula above.
Before the Contract Goes to Legal: Cap the commitment at the floor, the term at six months, and redline in a substitution right to next-generation silicon and a downward re-open on list price. Put Together AI's published ladder on the table as the anchor, and ask every other vendor to beat it in writing.
The Bottom Line
GPU reservations run against a short product cycle: H100 gave way to Blackwell, and Blackwell is now giving way to Rubin, inside roughly one three-year term. This week's evidence points both ways at once, with Nebius raising prices and NVIDIA shipping the generation that historically pushes them down. Commit the floor you measured, for no more than six months, and rent the rest.
Continue Reading
- GPU Hour Pricing: A $3.44 Rate Cost $68.80 to Use
- GPU Clouds Compared: Nebius on Price, CoreWeave at 3AM
- Run:ai vs Kubernetes vs Slurm: Static Quotas Strand Your GPUs
- Best FinOps Tools for AI Spend: No Dashboard Fixes an Untagged Key
- Equinix Set a Date, Not a Price. Cap Your Commit Term.
- AWS Gets Paid to Buy. Don't Lock Your Instance Family.
- Nebius Buys Inferize, Whose 20-Second Restore Has No Ship Date
