GPU Clouds Compared: Nebius on Price, CoreWeave at 3AM

Sixty-four B200s for thirty days costs $329,472 on Nebius and $742,349 on Google Cloud — identical silicon. Rent from Nebius or Together; CoreWeave's 20% premium is a question of cluster scale, not nerve. And never mistake Lambda's single-node price for its cluster price.

By Rajesh Beri·August 29, 2026·17 min read
Share:
A single 8-GPU server node pulled halfway out of a datacenter rack in a cold aisle, one thick orange InfiniBand cable unplugged and hanging loose beside an amber fault light, the rest of the row lit blue.

Illustration generated using AI

The same eight nodes of B200s cost $329,472 or $742,349 for thirty days, depending only on whose logo is on the invoice.

Same silicon. Same NVLink domain inside the node. Same InfiniBand between them. The 2.25x spread is not a performance difference — it is a pricing difference, and most of it is not buying you anything you can measure. But some of it is, and telling those apart is the entire job.

Here is the recommendation before the analysis: start on Nebius or Together AI for anything under a year, move to CoreWeave when your cluster is big enough that node failures stall it for days a month rather than hours, and use a hyperscaler only when your data cannot leave it. Google Cloud is the loser on this table and it is not close. Lambda is the trap — its famous price is for a machine that cannot do the job most people buy it for.


What 64 B200s Cost for 30 Days

Normalised to one workload: 64 NVIDIA B200 GPUs — eight nodes of eight, InfiniBand-connected, running continuously for 30 days. That is 46,080 GPU-hours. No spot, no preemption, no partial-node fractions. The neocloud prices below were read off each vendor's live pricing page on 30 August 2026; the AWS, Azure and Google Cloud rates come from outside trackers, because those vendors do not publish a simple per-GPU-hour rate. Treat the Google row with the most caution: it is the one number here sourced to a rival GPU cloud's blog rather than a neutral tracker, and it is worth knowing that before it is used to call Google expensive.

The tier column is ClusterMAX 2.1, SemiAnalysis's independent rating of more than 80 GPU clouds. A ClusterMAX tier is an outside assessment of networking, storage, reliability, security, observability and support quality — not a vendor's own claim, which is why it is the only quality signal on this table.

Provider Tier B200 $/GPU-hr 46,080 GPU-hrs Network egress Shortest term
Nebius Gold $7.15 $329,472 Free Hourly
Together AI (31–90d reserved) Silver $7.79 $358,963 Not published 31 days
Together AI (on-demand) Silver $8.19 $377,395 Not published Hourly
CoreWeave Platinum $8.60 $396,288 Free Rarely short-term
Lambda 1-Click Cluster (64 GPU) Silver $9.36 $431,309 Free 2 weeks
AWS p6-b200.48xlarge Silver $14.24 $656,254 $0.09/GB tiered Non-cancellable block
Google Cloud a4-highgpu-8g Silver $16.11 $742,349 Paid Reservation, Spot or Flex-start

Sources, in order: Nebius lists HGX B200 at $7.15/GPU/hr on-demand and $3.95 preemptible. Together AI publishes a rare public reserved ladder — $8.19 on-demand, $7.99 at 7–30 days, $7.79 at 31–90 days, $6.79 at 91–180 days. CoreWeave lists HGX B200 at $68.80 per 8-GPU node, which is $8.60 per GPU. Lambda charges $9.86 at 16 GPUs, $9.36 at 64, $8.87 at 256+. AWS's p6-b200.48xlarge carries 8 B200s at $113.933/hour. Google's a4-highgpu-8g is $128.88/hour, or $16.11/GPU — a figure that source itself labels "on-demand-equivalent," because Google's own documentation states that for A4 machine types you must reserve capacity, use Spot VMs, use Flex-start VMs, or create a resize request in a MIG. There is no plain on-demand path to that rate. Trackers disagree wildly on this row — some publish an a4-highgpu-8g rate near $35–40/hour — but those figures imply a B200 node costing less than half Google's own prior-generation a3-highgpu-8g, which the same source puts at $87.83/hour for eight H100s. A Blackwell node at half the price of a Hopper node is not credible list pricing; $128.88 is the figure consistent with the generational step.

Azure is absent because it publishes no on-demand B200 rate. Its closest listed part, ND96isr H200 v5, is $110.24/hour — $13.78 per H200-hour against Nebius's $4.50 for the same chip. That is a 3.1x list premium on last-generation silicon.


Lambda's Headline Price Is Not Its Cluster Price

Lambda's $6.69 B200 is real, and it is for a machine that cannot train across nodes. This is the single most expensive misreading in GPU procurement, and Lambda's own two pricing pages make it plainly, if quietly.

GPU Single node, on-demand Cluster, 16 GPUs 64 GPUs 256 GPUs
H100 SXM $3.99 $6.16 $5.85 $5.54
B200 SXM6 $6.69 $9.86 $9.36 $8.87

The on-demand rates buy you one isolated 8-GPU box. The 1-Click Cluster rates buy you Quantum-2 InfiniBand between boxes, a 16-GPU minimum and a two-week minimum reservation. The gap runs 33% to 54%.

That matters because the InfiniBand is the product. Fine-tuning a 70B model on eight GPUs works fine on a single node. Anything that needs 64 GPUs needs the fabric, and once you are paying for the fabric, Lambda at $9.36 costs 9% more than CoreWeave's Platinum-rated $8.60 and 31% more than Nebius. You are paying a premium for a Silver-tier service.

ClusterMAX's Lambda review is blunt about what that Silver covers: a product surface split across "new-mslurm, old-mslurm, new-mk8s, old-mk8s, private cloud, 1-Click Cluster, and on-demand Instances," a Kubernetes offering that "feels like an early-stage offering, marked by technical debt and a challenging user experience," and "a general degree of disorganization." It also credits Lambda with the largest on-demand fleet in the market and full model FLOPs utilisation once a cluster is up — both true, and both reasons Lambda is genuinely the right answer for single-node work.

Lambda's real product is the $3.99 H100 you can have in ninety seconds. Buy that. Do not buy its cluster tier without pricing CoreWeave next to it.


CoreWeave Charges a Platinum Premium and Mostly Earns It

CoreWeave is the only provider ClusterMAX has ever rated Platinum, and the premium it charges is smaller than the reputation suggests — 20% over Nebius, 8% under Lambda's cluster price.

The CoreWeave review calls it "the only Neocloud with a hyperscaler mentality," singles out a custom Slurm fork that fixes upstream memory leaks at scale, and describes a "sophisticated correlation engine" built to identify "transient physical-layer problems like signal integrity degradation between compute trays and switches." A neocloud is a GPU-only cloud with no general-purpose compute business — the category CoreWeave, Lambda, Nebius, Crusoe and Together all sit in, and the reason they can undercut hyperscaler list prices at all.

That reliability work is what you are actually buying, and there is a number for why it matters. Meta's Llama 3 paper reports 419 unexpected interruptions across a 54-day pre-training run on 16,000 H100s. Roughly 78% traced to confirmed hardware issues; GPU faults alone were 58.7% of all unexpected interruptions. Meta still hit above 90% effective training time — because it had the tooling to detect, drain and replace a bad node automatically. Scale that same failure rate down to 64 GPUs for 30 days, though, and it predicts roughly one interruption in the month. The shape is identical; the frequency is not, and that difference is what decides whether a reliability premium pays for itself.

The catch is structural, and it is the reason CoreWeave's published rate card may be irrelevant to you. ClusterMAX notes CoreWeave "do not offer on-demand instances or autoscaling, and rarely accept short-term rentals." That $68.80/hour is a reference price, not a self-serve button. CoreWeave's Q2 2026 results show $2.575 billion of quarterly revenue against an approximately $104 billion revenue backlog as of 30 June 2026 — a company whose capacity is committed years out to a small number of very large contracts. If you are asking for 64 GPUs for a month, you are not the customer that backlog was built for.

The same review flags the other side of Platinum: security posture "way to limiting for a lot of power users," with systemd and profiling tools unavailable by default. If your team debugs with perf and nsys, ask about that before you sign.


Together AI Sells Flexibility With a Non-Refundable Clause

Together is the only provider here that publishes a full reserved-price ladder, and it is the right default when you do not yet know how long you need the machines.

Together Instant Clusters provision "in minutes" from a single 8-GPU node up to hundreds of interconnected GPUs, over non-blocking Quantum-2 InfiniBand, with terms of hourly, 1–6 days, or 1 week to 3 months. The published ladder for H200 runs $5.99 on-demand, $4.99 at 7–30 days, $4.15 at 31–90 days, $3.99 at 91–180 days. That last number is 37% below CoreWeave's H200 on-demand rate of $6.31 per GPU and 50% below AWS's $7.91 for p5en.48xlarge.

Read the billing terms before you celebrate. Together's docs state it flatly: "reservations are non-refundable. The full reservation period is charged upfront and cannot be cancelled or partially refunded." Storage also "continues to accrue charges even when no cluster is using it." A 90-day, 128-GPU H200 reservation at $4.15 is a $1.15 million cheque written upfront that you cannot unwind if the model you were training gets cancelled in week three.

That is not a scandal — it is how capacity businesses work. It is a thing to know before your CFO discovers it.


Nebius Is the Value Pick on This Table

Nebius wins on price at every generation and it is not a discount-brand tradeoff: ClusterMAX rates it Gold, above Lambda, Together, AWS and Google Cloud.

Its price list is the cheapest here for H100 ($3.85), H200 ($4.50), B200 ($7.15) and B300 ($7.85), with preemptible rates roughly 45% lower and shared filesystem storage at $0.08/GiB-month. Network egress is free; note that its object storage is not — that carries a separate $0.015/GiB egress charge, which is the kind of line item that only shows up on the invoice. The Nebius review credits a Kubernetes-native platform delivering "bare-metal-equivalent performance," a managed Slurm service "completely set up and ready for use out of the box," and notes "There is no other Neocloud as open and transparent with their development as Nebius."

Two honest caveats. The same reviewers criticise Nebius for cherry-picking reliability data — "We are generally frustrated by providers who cherry-pick reliability data in this manner" — and note it "struggles in securing colocation deals and establishing a credit rating." The second is a supply risk on a multi-year commitment: a provider that cannot secure datacenter capacity cannot deliver capacity you reserved. The review also records that some customers remain fixated on the company's Russian corporate origins despite all staff being based outside Russia. If your procurement or security review will raise that, raise it yourself, early, in writing, rather than at contract signature.


The Hyperscalers Cost Twice as Much on List

AWS at $14.24 and Google Cloud at $16.11 per B200-hour are 2.0x and 2.3x Nebius for identical silicon, and both hold a Silver ClusterMAX tier — below the cheaper Nebius. On raw compute economics there is no argument to make for them.

There are three arguments that are not about compute economics.

Your data already lives there. If the training corpus is 400 TB in S3, moving it out costs real money and real weeks. AWS charges $0.09/GB for the first 10 TB, $0.085 for the next 40 TB, $0.07 for the next 100 TB and $0.05 above 150 TB, after 100 GB free per month aggregated across services and regions. A 100 TB one-way exit is roughly $8,000. That is under two percent of the gap between the cheapest and most expensive rows on this table — not a lock-in moat, but the transfer time and the re-plumbing of every downstream pipeline usually is.

Compliance is already argued. If your FedRAMP boundary, your BAA, your data residency commitment and your security review all terminate at an existing hyperscaler contract, re-running that process for a neocloud is months of work by people you cannot spare.

Capacity Blocks are a genuinely good product with genuinely hard edges. EC2 Capacity Blocks let you reserve up to 64 instances (256 across all blocks) with a start date up to eight weeks out, placed inside an UltraCluster, and you can find offerings starting in as little as 30 minutes. Then read the constraints: "Capacity Block cancellations aren't allowed," blocks always end at 11:30 UTC with termination beginning at 11:00 UTC, GB200 UltraServer blocks require you to terminate instances 60 minutes before the end, and blocks cannot be moved, split, or placed in placement groups. AWS also states that "Reservation prices are updated regularly based on trends in supply and demand" and are locked at the moment of purchase, with the next scheduled update in October 2026.

Google Cloud is the loser here. It charges the highest price on the table, holds the same Silver tier as providers charging half as much, and offers no reliability, support or availability advantage that shows up in an independent rating. If you are not already committed to GCP for data-gravity reasons, there is no case for renting B200s there.


Egress Is a Small Tax. Exit Terms Are the Big One.

Egress fees are the most-discussed and least-important line item in this comparison. Exit terms are the reverse.

CoreWeave states that data transfer between its cloud and the internet is free; Lambda advertises "no ingress/egress fees"; Nebius charges nothing for network traffic. That is worth roughly $8,000 per 100 TB against AWS — about 2% of a single month's compute on our reference workload. It is a real saving and it is not a reason to choose anything.

What should drive the decision is what happens when you want to stop paying. Rank them by how expensive it is to be wrong:

  • Together: non-refundable, charged upfront, for terms up to 180 days.
  • AWS Capacity Blocks: cancellations are not allowed, at a price fixed at purchase.
  • Lambda: two-week minimum on clusters, no lock-in beyond the reservation.
  • Nebius: hourly on-demand, commitment discounts up to 35% if you want them.
  • CoreWeave: the shortest term you will likely be offered is measured in quarters, not weeks.

The correct posture for a first commitment is the same one that applies to multi-year compute deals: buy the shortest term whose price you can live with, and let measured utilisation — not a vendor's discount curve — earn the longer one.


Who Should Not Buy Each of These

This is the section a vendor comparison usually skips.

  • Do not buy CoreWeave if you need fewer than a few hundred GPUs, need them this week, need to autoscale, or need root-level profiling tools by default. Its economics and its operating model are both built for large committed contracts.
  • Do not buy Lambda's 1-Click Clusters without putting CoreWeave and Nebius quotes next to them — at 64 GPUs it is the third-most expensive option here. Do buy Lambda for single-node work, where it is the fastest, cheapest path in the market.
  • Do not buy Together's reserved tiers if there is any chance the project gets cancelled. The discount is real; so is the non-refundable clause. Its on-demand and 7-day tiers exist for exactly this reason.
  • Do not buy Nebius if a multi-year deal is on the table and your risk function will not accept a supplier with an unsettled credit rating and unresolved colocation supply — or if you cannot get the corporate-origin question closed with your security team.
  • Do not buy AWS or Google Cloud at list rates for training. If you are on a hyperscaler, buy Capacity Blocks or a committed-use discount; paying 2x list for elastic capacity you are going to run flat-out for 30 days is the most common way enterprises overspend on GPUs.
  • Do not buy Azure for GPU rental on price. It publishes no on-demand B200 rate and its H200 list price is over 3x Nebius for the same chip.

What Actually Predicts Regret at Renewal

Four criteria, in the order they will bite you.

1. Node-hours lost to failure, not dollars per GPU-hour. A cluster that is dead 17% of the month at $7.15 costs the same per useful GPU-hour as a flawless one at $8.60 — which is exactly the Nebius-to-CoreWeave gap, and 17% of a month is five days. That is the bar a reliability premium has to clear. Ask every vendor, in writing, for their mean time to detect a bad node, their mean time to replace it, and whether you are billed for the drained node. Get it in the contract, not the sales deck. The Llama 3 numbers are the reason: hardware fails on a clock set by your GPU count, and the only variable you control is who notices first.

2. Whether the price you were quoted is the price for the topology you need. Single-node and InfiniBand-connected are different products at the same vendor. Normalise every quote to the same node count, the same fabric, and the same term before you compare anything.

3. The shortest term you can exit. Every provider will offer you a better rate for a longer commitment. The discount is real and the option value you surrender is larger than the discount for any workload you have not yet run at steady state.

4. Independent quality rating, not the vendor's uptime page. ClusterMAX is the only third-party instrument in this market. Where it disagrees with a sales claim, believe it.

What changes the answer: if you are serving inference rather than training, most of this table stops mattering and the inference runtime you choose matters more than the cloud you rent. If you are buying for 2027 capacity rather than next month, contracted delivery dates beat headline capacity numbers every time. And if utilisation is your actual problem — as it is for most enterprises with idle GPUs — no provider on this list fixes it.


The Bottom Line

Rent from Nebius or Together, and treat CoreWeave's premium as a question of scale rather than nerve. The threshold is calculable: on 64 B200s, CoreWeave costs $66,816 more per 30 days — about $93 for every hour of the month, whether or not anything breaks. That asymmetry is the whole point. You pay the premium continuously and only collect on it while a node is down, so on rent alone it breaks even only if CoreWeave spares you around 121 hours of downtime a month. Meta's failure rate scaled to 64 GPUs predicts about one interruption in that month, not five days of them. Add blocked engineer time and the bar comes down, but not to anything most 64-GPU teams will clear.

At a few thousand GPUs the arithmetic inverts, because one bad node stalls every GPU you are renting and failures arrive daily rather than monthly. That is the cluster CoreWeave is built for — and its refusal to rent you 64 GPUs for a month is the same fact stated as a sales policy.

The hyperscalers are worth their premium only when the alternative is moving your data or re-running your compliance review. Google Cloud is not worth it at all right now. And Lambda's price, the one everybody quotes, is for a machine that only does half the job.

The GPU is a commodity. The gap between a node dying and someone telling you is not.

Continue Reading

Share:

Frequently Asked Questions

Why is Lambda's cluster price so much higher than its on-demand price?

Because they are different products. Lambda's on-demand $3.99 H100 and $6.69 B200 buy a single isolated 8-GPU node with no inter-node fabric. Its 1-Click Clusters add Quantum-2 InfiniBand and cost $5.54–$6.16 per H100 and $8.87–$9.86 per B200 depending on cluster size, with a 16-GPU and two-week minimum. Any workload spanning more than one node pays the cluster rate — 33% to 54% higher.

How much does it cost to move 100 TB of data out of AWS?

About $8,000. AWS charges $0.09/GB for the first 10 TB, $0.085 for the next 40 TB, $0.07 for the next 100 TB and $0.05 above 150 TB, after 100 GB free per month aggregated across services and regions. CoreWeave, Lambda and Nebius all charge nothing for network egress. On a 46,080 GPU-hour workload that saving is about 2% of compute spend — real, but not a reason to pick a provider.

Can you cancel an AWS EC2 Capacity Block reservation?

No. AWS documentation states plainly that Capacity Block cancellations aren't allowed, and the price is fixed at the moment of purchase regardless of later changes. Blocks can be reserved up to eight weeks ahead, hold up to 64 instances each (256 across all blocks), always end at 11:30 UTC with termination starting at 11:00 UTC, and cannot be moved, split, or used with placement groups.

Which GPU cloud is cheapest for training in 2026?

Nebius, at every current generation. As of 30 August 2026 its published on-demand rates are $3.85 per H100-hour, $4.50 per H200, $7.15 per B200 and $7.85 per B300, with preemptible instances about 45% cheaper and free network egress. It also holds a Gold ClusterMAX rating — above Lambda, Together, AWS and Google Cloud — so it is not a quality tradeoff. The open risks are its unsettled credit rating and colocation supply.

Newsletter

Stay Ahead of the Curve

Weekly enterprise AI insights for technology leaders. No spam, no vendor pitches—unsubscribe anytime.

Subscribe