Microsoft's GPUs Sit Unplugged. Buy Delivery, Not Capex.

A Guardian investigation put a measured number on the gap between Microsoft's announced AI capacity and its installed accelerators. Azure's own documentation already says a PTU reservation guarantees you a price, not capacity — which is why a capacity clause has to name region, SKU, quantity, delivery date and remedy.

By Rajesh Beri·August 17, 2026·13 min read
Share:
A row of sealed cardboard shipping crates stacked on a concrete loading dock inside a half-finished data hall, with empty unpowered server racks and bare hanging cable trays visible behind them.

Illustration generated using AI

The number your cloud contract is written against — announced capex, fleet-wide chip counts, gigawatts under construction — is not the number that serves your traffic. The only capacity you can use is an accelerator that is racked, powered, cooled, networked and reachable from the region your data lives in. Everything upstream of that is a press release.

On Monday, a Guardian investigation by Aisha Down and Ed Zitron put a measured figure on the distance between those two things at the largest AI infrastructure buyer on earth. Microsoft says the arithmetic is wrong. It does not dispute the mechanism — and neither should you, because the mechanism is the reason your provisioned throughput order got rejected last quarter.


What the Guardian Counted, and What Microsoft Denied

The investigation's core claim is that Microsoft has far fewer AI accelerators actually installed than its spending and its announced power footprint imply. Working from internal documents, the Guardian reports roughly 2.2 million AI chips installed globally — against an internal target of 1.8 million by the end of 2024, nearly two years earlier. It reports Microsoft claiming 5GW of datacentre capacity added over two years while UC Riverside professor Shaolei Ren, reading the company's sustainability reports, put its AI capacity in 2024 closer to 1.2GW — and noted that even that lower figure would imply Microsoft needs roughly 4 million chips if it added 5GW. The Guardian's own arithmetic runs higher still: dividing 8GW by 10kW per server gives 6.4 million chips, and 10GW divided by an H100's 700 watts gives 12 million. All of this sits on roughly $280bn of AI investment since 2022.

Microsoft's response was that "the estimates the Guardian has shared with us are inaccurate, drawing the wrong conclusions from incorrect assumptions," and that it does not disclose the volume of specific chips in its infrastructure.

Take that seriously, because it is probably right in the particulars. Chip-count-from-gigawatts is a division problem with three unstable variables: the power draw per accelerator generation, the power usage effectiveness of the building, and how much of a site's total capacity is AI versus general compute. Move any one and the answer swings by a factor of two. A range of 4 to 12 million is not a measurement; it is an admission that the measurement is hard.

But notice what Microsoft did not say. It did not say every chip it owns is installed and serving. It did not say announced capacity equals energized capacity. And on the substance of the gap, its own executives have been more candid than the denial suggests.

Microsoft Already Has a Name for This Gap

Microsoft measures the delay between a chip arriving and a chip serving traffic, and it reports improving that number to investors. On the FY26 Q4 call, Satya Nadella told analysts that "over the last fiscal year, we have reduced dock-to-live times for new GPUs in our largest regions by nearly 50%."

Dock-to-live is the interval between a GPU landing on a loading dock and that GPU accepting a customer request. You cannot cut a metric by half unless the metric was large, and you do not build a metric for something that does not cost you money. Microsoft has a metric for the mechanism the Guardian is arguing about — the delay, not the size of the installed base — and it publishes only the improvement, never the level.

The bottleneck is not silicon. On the BG2 podcast in late October 2025, Nadella said that "the biggest issue we are now having is not a compute glut, but it's power — it's sort of the ability to get the builds done fast enough close to power," and, more bluntly: "you may actually have a bunch of chips sitting in inventory that I can't plug in. In fact, that is my problem today. It's not a supply issue of chips; it's actually the fact that I don't have warm shells to plug into." A warm shell, in datacentre terms, is a finished building with power, cooling and network already energized — the thing you rack into. Nadella made the remarks alongside Sam Altman in a conversation hosted by Brad Gerstner.

The scale of the build is not in doubt. Nadella told the FY26 Q4 call Microsoft "added 31 new datacenters across 5 continents this quarter, bringing the total to 88 this year" and "another gigawatt of capacity this quarter." Microsoft's own FY26 Q4 results show $115,948m of additions to property and equipment for the year, $35,802m of it in the June quarter, against $90.0bn of quarterly revenue and Azure growth of 43%.

And still, CFO Amy Hood told the same call that "customer demand continues to exceed available capacity." That is the sentence that matters to you. It is not an admission of failure — it is a description of a market in which the seller allocates. When supply is rationed, the contract language decides who gets rationed, and we have watched Google ration Gemini compute to Meta when the same conditions applied.


Quota Is Not Capacity, and Azure Says So

Every enterprise AI capacity instrument sold today explicitly disclaims the thing buyers think they are buying. This is not a gotcha; it is written in the product documentation, and almost nobody in a procurement conversation has read it.

Microsoft's own provisioned throughput documentation for Azure AI Foundry states it three separate ways:

  • "Having PTU quota doesn't guarantee that capacity is available. If capacity in the region is insufficient for the requested PTU count, the deployment fails." Quota is a policy limit with no cost attached. Capacity is the physical resource, allocated at deployment time.
  • "Reservations don't guarantee capacity. First create deployments to confirm that capacity is available, then purchase the reservation to lock in the discounted rate." The reservation is a price instrument, not a supply instrument.
  • "Deleting or scaling down a deployment releases its capacity back to the region pool. There's no guarantee the same capacity is available if you re-create or scale the deployment up later."

Read that third one twice. It means a provisioned deployment is a position you hold, not a right you own. Scale down for a quiet quarter and you may not get back in. The documentation also notes that capacity "changes throughout the day based on customer demand across all regions and models," and that quota is scoped per subscription, per region, per deployment type — "quota in East US doesn't carry over to West Europe."

Now pair that with the exit. Azure's self-service exchange and refund policy caps total cancelled commitment at USD 50,000 in a rolling 12-month window per billing profile or enrolment, and warns that while Microsoft "isn't currently charging early termination fees," there "might be a 12% early termination fee for cancellations" in future. Microsoft's own worked example is instructive: on a three-year, $3,000-a-month reservation totalling $108,000, "you can't cancel the reservation until you've spent 58,000 USD of your commitment."

So the shape of the deal is: you commit for a term, the discount is guaranteed, the capacity is not, and the unwind is capped at fifty thousand dollars a year. That asymmetry is the entire negotiation.

This is not hypothetical scarcity. In late July 2025 a demand spike exhausted compute in Azure's East US region and customers hit AllocationFailed errors for over a week; by early 2026 Microsoft temporarily stopped accepting new VM deployments for GPU and AMD SKUs in UK South. Regions run out. Yours will.

What the Three Clouds Actually Promise You

All three hyperscalers sell region-and-model-specific capacity with a commitment term, and none of them sells guaranteed delivery of physical accelerators into a region on a date. The differences are in the shape of the option, not the strength of the promise.

Azure provisioned throughput AWS Google Vertex AI
Unit PTU (per subscription, per region, per deployment type) Model Units (Bedrock); instances (EC2 Capacity Blocks) GSU (per model, per region)
Terms Hourly, or 1-month / 1-year Azure Reservation Bedrock: none / 1 month / 6 months 1 week / 1 month / 3 months / 1 year
Forward booking None — deploy against live capacity Capacity Blocks: start date up to 8 weeks out None documented
Early exit Reservation refund capped at $50k / 12 months Bedrock commitments cannot be deleted mid-term; Capacity Block "cancellations aren't allowed" Committed term
Overflow Spillover to a standard deployment, or 429 Throttling Falls back to pay-as-you-go if available, else 429

For Amazon Bedrock, the provisioned throughput documentation is explicit that a one-month or six-month commitment means "you can't delete the Provisioned Throughput until the … commitment term is over," and that Model Unit sizing and limit increases go through your account manager rather than a self-service console.

EC2 Capacity Blocks for ML is the closest thing anyone sells to a dated delivery promise — and its limits tell you how tight physical supply is. You can book a start time up to eight weeks in the future, up to 64 instances per block and 256 across blocks, and "Capacity Block cancellations aren't allowed." More telling is the SKU-by-region table: the newest Blackwell part on the list, p6-b300.48xlarge, is bookable in four locations — N. Virginia, Oregon, GovCloud US-East, and the Atlanta local zone. If your workload has to run in Frankfurt or Sydney, the newest silicon is not an option you can buy at any price. That is the same constraint we found when AMD's inference discount depended on a GPU you could not actually rent.

Google's Vertex AI Provisioned Throughput sells Generative AI Scale Units against a named model in a named region, with 1-week, 1-month, 3-month and 1-year terms; traffic above your GSU allocation falls back to pay-as-you-go only "if capacity availability allows," and returns 429 if it does not — a vendor telling you, in writing, that even your overflow is conditional on capacity somebody else has not already taken.

Mistral's European Compute Units are structured the same way, which is why we argued in August that a multi-year commitment to a regional endpoint needs its own limits written down.

Five Things a Capacity Clause Must Name

A capacity commitment that does not name all five of these is a discount agreement with a capacity-shaped hole in it. Write them into the order form, not the MSA, because the order form is where the specifics live and where the vendor's deal desk has authority.

  1. Region. Not "US" and not "global." The named Azure region, AWS Region, or Google region your data residency obligation actually requires. Global routing is a real product with real availability advantages, but it is a different product from the one your regulator approved.
  2. SKU and model version. "Blackwell-class" is not a SKU. p6-b300.48xlarge is. "GPT-class frontier model" is not a model version. If you are buying PTUs, the per-model PTU-to-token ratio is what determines whether your quota actually serves your traffic.
  3. Quantity with a floor and a ceiling. The floor is what you can rely on for the term. The ceiling is what you can burst to without a new order. Both should be numbers.
  4. Delivery date, and what "delivered" means. Delivered is a successful deployment serving production traffic at the contracted throughput — not an approved quota request, not an accepted purchase order. Borrow Microsoft's own vocabulary: you are buying live, not docked.
  5. Remedy. What happens if the region comes up short. Service credits are the weakest form and are what you will be offered. Stronger: the right to redeploy the commitment to another region or another model without penalty; a suspension of the commitment clock while capacity is unavailable; a termination right that is not subject to the standard cancellation cap.

Point five is where the negotiation actually happens, and it is winnable, because the vendor's own documentation has already conceded that capacity may not be there. You are not asking them to promise something new. You are asking them to price the risk they have already disclosed.

Three Moves Before Your Next Renewal

This Week: Pull every AI capacity commitment you hold — PTU reservations, Bedrock commitments, GSU orders, committed-spend discounts — into one sheet with five columns: region, SKU/model, committed quantity, term end date, and stated remedy. Most organisations discover the remedy column is empty on every row. Then run one deployment test: try to create a provisioned deployment at your committed size in your primary region and record whether it succeeds. If it fails, you have a live problem and a documented one.

This Month: Instrument the gap yourself. Log every 429 and every allocation failure by region, model and hour, and keep a 90-day rolling record. When you sit down with the deal desk, "we were capacity-denied 340 times in Q3, here is the log" moves the conversation somewhere "we are worried about capacity" never will. In parallel, define your fallback path before you need it — a second region, a second model family, or an open-weight deployment you control, the way we argued for an open-weights inference hedge against DeepSeek's price ceiling.

Before Renewal: Take the five-item clause above to the order form and refuse to sign a commitment whose remedy is service credits alone. Ask one question directly: what is your dock-to-live time in the region I am buying? Microsoft publishes that it improved that number by nearly 50%; ask for the level. The answer, or the refusal to answer, tells you what you need to know about whether the capacity you are paying for exists yet.


The Bigger Picture

We have run this experiment before, underground. After the Telecommunications Act of 1996, carriers buried hundreds of thousands of miles of fibre across the continent, and by 2004 analysts estimated only about one-tenth of installed fibre was actually "lit". The fibre was real. The capital was real. The route miles in the press releases were real. What was not real was the assumption that a strand in a conduit is the same thing as a circuit you can order, and the companies that wrote contracts against route miles instead of lit capacity did not survive to see the glut turn useful.

The AI buildout is not the fibre bubble — demand here is arriving faster than supply, which is the opposite problem, and the capex debate has already reached the investors. But the contracting error is identical. Announced capacity is an intention. Energized capacity is an asset. Between them sits a substation queue, a transformer lead time, a water permit and a grid interconnect — the same constraints that are already delaying US data centre projects and showing up in your inference bill through electricity policy.

Whether Microsoft has 2.2 million chips installed or 6 million is, honestly, not your problem. You are not an equity analyst. Your problem is that you signed a multi-year commitment denominated in a number that describes your vendor's ambition, and you will be measured on a number that describes your users' latency.

Stop buying gigawatts. Buy the chip you can plug in, in the region you are allowed to use, on a date someone will sign.

Continue Reading

Mistral Wants Multi-Year Money. Its EU Endpoint Drops Agents. Google Just Rationed AI to Meta. Your Enterprise Is Next. AMD's Inference Discount Depends on a GPU You Can't Rent North Carolina Taxed the Power. Inference Pays. 50% of US Data Centers Delayed: The AI Power Crisis Is Here Snowflake Cortex vs Databricks Mosaic AI: Pick on Exit Cost Your AI Router Is Trading a 10x Discount for a 2.5x One

Share:

Frequently Asked Questions

Does an Azure PTU reservation guarantee I get capacity?

No. Microsoft's provisioned throughput documentation states plainly that "reservations don't guarantee capacity" and that "having PTU quota doesn't guarantee that capacity is available." A reservation is a financial discount applied to the PTU billing meter. Microsoft's own guidance is to create the deployment first to confirm capacity exists, then buy the reservation to lock in the rate.

What is dock-to-live time?

Dock-to-live is the interval between a GPU physically arriving at a data centre loading dock and that GPU accepting production customer requests. Satya Nadella told Microsoft's FY26 Q4 earnings call that the company had "reduced dock-to-live times for new GPUs in our largest regions by nearly 50%" over the fiscal year. It is the vendor's own measure of the gap between purchased hardware and usable capacity.

How far in advance can I reserve GPU capacity on AWS?

EC2 Capacity Blocks for ML let you reserve a start time up to eight weeks in the future, with up to 64 instances per block and 256 instances across all blocks. Cancellations aren't allowed once booked. Availability is SKU- and Region-specific: the p6-b300.48xlarge Blackwell instance is listed in only four locations, so newer silicon may not be bookable in your required region at all.

Can I cancel an Azure reservation if I can't get capacity?

Only partially. Azure's self-service refund policy caps total cancelled commitment at USD 50,000 in a rolling 12-month window per billing profile or enrolment, and Microsoft notes it may introduce a 12% early termination fee in future. On a $108,000 three-year commitment, Microsoft's own example says you cannot cancel until you have spent $58,000 of it.

What should an AI capacity commitment clause actually specify?

Five things: the named region (not "US" or "global"), the exact SKU and model version, a committed quantity with a floor and a ceiling, a delivery date where "delivered" means serving production traffic rather than an approved quota request, and a remedy if the region comes up short. The remedy is the negotiable part — push for redeployment rights or a suspended commitment clock rather than service credits alone.

Newsletter

Stay Ahead of the Curve

Weekly enterprise AI insights for technology leaders. No spam, no vendor pitches—unsubscribe anytime.

Subscribe