When the Colo's Chillers Fail, the GPU Loss Is Yours

S&P says there is no obvious insurance product for the GPUs inside AI data centers, and colo agreements often cap the operator's liability at three months of fees. If you own a GPU cluster in someone else's hall, the gap between that cap and the hardware is yours to close.

By Rajesh Beri·September 14, 2026·11 min read
Share:
A row of liquid-cooled GPU server racks in a dim colocation data hall, coolant pooling on the raised floor beneath one open rack, and a wheeled emergency cooling unit parked in the aisle.

Illustration generated using AI

The GPUs you own in a colocation hall are mostly your own risk. Colo agreements often cap the operator's liability at three months of fees, can exclude damage to your equipment outright, and require you to insure the hardware yourself. And on September 14, Bloomberg reported that S&P sees "no obvious coverage product" for the high-value GPUs housed in every data center.

So an enterprise running a cluster worth tens of millions of dollars in someone else's building is leaning on its own property policy. That policy may not respond to an outage that damages nothing, and where business interruption cover does apply, it can carry a 12-to-24-hour waiting period. Before the next colo signature or policy renewal, do three things: schedule the cluster at replacement value, test the waiting period and damage trigger against a realistic outage, and negotiate a power-and-cooling carve-out from the liability cap.

What S&P, Marsh and AM Best Are Telling the Market

The firms that rate, broker and analyze this risk are saying plainly that AI data centers have outgrown the standard insurance stack. Marsh's US and Canada captive solutions leader, Michael Serricchio, told Bloomberg to expect "explosive growth in the use of captives to take on the portfolio risks for data centers." S&P Global Ratings director Charles-Marie Delpuech called Meta's Hyperion campus "beyond fully-insurable" and said "there's an insurance gap that needs to be solved." Meta ended up providing a special guarantee for the bondholders financing part of the project, and S&P rated the senior secured notes issued by Beignet Investor, the Blue Owl-majority financing vehicle, A+.

A captive is an insurance company that a corporation owns to insure its own risks. There are more than 6,000 of them worldwide, writing about $240 billion in premiums, and Marsh alone manages about 1,900 captives writing $79 billion. Aon's Joe Peiser, chief executive of risk capital, gave the reason in the same report: when premiums go up, "clients look for ways to manage losses by taking the bottom layer" of the risk.

The capacity being built is aimed at the people who own the buildings. AM Best's June special report said the coverage data centers now need is "currently beyond what the traditional property/casualty industry has previously experienced." Aon lifted its Data Center Lifecycle Insurance Program to $5 billion in July, a program designed for "data center owners, developers and investors." The company whose GPUs sit in a rented cage is not in that sentence.

Why Your Colo Agreement Leaves the GPU Loss With You

Standard colocation paper measures the operator's exposure in rent and yours in hardware. Atlantic.Net's published colocation master services agreement caps its total liability at the fees paid "in the three (3) month period immediately preceding the event giving rise to the liability." For any breach, the customer's "sole and exclusive remedy" is re-performance or a refund or credit under the service level agreement. The customer must carry property insurance "at full replacement cost value," that insurer "waives all rights of subrogation" against the operator, and the customer releases the operator from liability for "damage, loss or injury to person and/or property, even if caused by Atlantic.Net's own negligence."

The subrogation waiver is the clause most readers skim. It means that after your carrier pays for the damaged racks, it cannot recover from the operator whose chiller or battery room caused the loss. The loss stays on your insurance program.

High-density hosting agreements follow the same pattern. A 2020 bitcoin-mining colocation agreement between Whinstone US and Jordan HPC, filed with the SEC, lists "damage to customer equipment" among the consequential damages neither party is liable for, requires the customer to insure enough "to provide for the complete replacement of the Customer Equipment," and removes the cap only for recklessness, gross negligence, fraud or willful misconduct. A routine equipment failure is unlikely to clear that bar.

AC Research's review of GPU colocation contracts says most of the contracts it has examined cap the provider's total liability at the fees paid in the three months before the incident. It does not say how many it reviewed, and its publisher, American Compute, brokers insurance for GPU hardware, so it has an interest in the gap. Its worked example: a $300,000-a-month deployment gets a $900,000 ceiling, against hardware that might be worth $25 million — under 4% of the exposure. Every contract it reviewed also caps service credits at 100% of the monthly fee for the affected service.

The operator's case deserves a fair hearing. Rent is priced for space, power and cooling, not for underwriting a customer's accelerators, and no single tenant's hardware is the operator's to value. Lawyers who draft these deals see the tension growing: Morgan Lewis's colocation guidance says insurance and subrogation rights "may be a critical component in the evaluation of risk with respect to damage to the colocation facility (for the service provider) and servers (for the customer)," and warns that "given the substantial value and scarcity of certain hardware, these risks are becoming more significant." Buyers spend their negotiating energy on commitment term and price, as with Equinix's inference exchange placements. The liability section deserves at least as much.

What Actually Takes GPUs Offline in a Data Hall

Fire and water cause the expensive losses, and power is the leading cause of outages. In FM's 15-year review of data center losses, cited in the Swiss Re Institute's sigma report, fire accounted for about 11% of loss events but more than 42% of loss costs, and liquid-related losses were nearly 24% of total loss costs. Uptime Institute ranks power as the leading cause of impactful data center outages, at 45% in the Uptime figures cited in the same coverage.

Both hazards grow with AI density. AI server racks can require more than 100 kilowatts, up from 5-15 kilowatts for traditional servers. A liquid-cooled GB200 NVL72 rack draws about 120 kilowatts and costs roughly $3 million, per Introl's deployment guide, which also notes that deviating from cooling specifications triggers throttling that can cut performance by 60%. Ten of those racks is about $30 million of hardware with coolant running through it. The liquid cooling buildout that makes those racks possible also puts a water path next to the most valuable equipment in the building.

Two recent incidents show how these losses unfold:

Where Your Own Policy Goes Quiet

Your property policy will probably respond to a fire that destroys your racks; it may not respond to the outage that costs you the most. Property policies generally limit business interruption coverage to situations where "some physical loss or damage to the property precipitated the closure," Morgan Lewis insurance lawyers wrote, and some carry "utility services" exclusions that can bar losses from a failure of power supply. A chiller failure in the operator's plant that shuts your cluster down without damaging it may never meet that trigger.

Where cover does apply, three more gaps open:

  1. The waiting period. S&P said tech firms that buy business interruption cover against power failure may face 12 to 24 hours before it kicks in. For a cluster training a model or serving production traffic, those are the hours that hurt.
  2. Dependent property limits. When the damage is to someone else's property — the operator's switchgear, a utility substation — you need contingent or dependent property cover, which Bracewell's insurance lawyers note is "frequently sublimited or subject to restrictive triggering conditions."
  3. The restoration window. Business interruption pays for the time it takes to restore. Bracewell points out that "client attrition, service-level agreement penalties and reputational damage can extend the financial impact well beyond the restoration window." Replacement cost is a dollar figure. It does not put an accelerator back in the rack.

Bracewell also flags the "silent cyber" problem, where a single intrusion "could simultaneously cause physical equipment damage, data loss and prolonged business interruption" and land between policies — the same seam we examined in AI cyber insurance riders.


Own, Rent or Self-Insure: The Choice You Are Already Making

Every GPU cluster has an insurer, and if you have not picked one deliberately, it is your balance sheet. There are three honest options.

Rent the capacity. When you buy GPU hours from a provider such as Crusoe, Nebius or Together AI, the hardware on the floor is theirs, so the property loss is theirs. You keep the downtime, compensated only through their service credits. Our GPU cloud comparison and GPU hour pricing breakdown cover what that flexibility costs.

Own it and insure it properly. The cluster is scheduled at replacement value at its facility address, the underwriter knows the racks are liquid-cooled, and the business interruption terms were tested against a realistic outage rather than inherited from the office property program. Specialist brokers such as American Compute now place all-risk property and offtake-interruption cover written for GPU hardware with surplus lines insurers. That is a bespoke placement, not the standard product S&P says is missing, and it is worth pricing against your existing program.

Own it and self-insure on purpose. If your company already runs a captive, the waiting-period layer and the gap between the colo cap and your property deductible are exactly the "bottom layer" Aon's Peiser described. Retaining it there is a decision a risk committee can make, document and fund. That beats discovering the same exposure after a fire.

What to Do Before the Next Colo Signature or Renewal

The work is mostly reading and a few calls, and it fits inside one quarter.

This Week:

  1. Put four numbers on one page. From the colo agreement: the liability cap in months of fees, the service-credit ceiling, and whether damage to customer equipment is excluded or released. From finance: the cluster's replacement value. The ratio between the cap and that value is the conversation.
  2. Ask your risk manager one question: is the cluster listed at replacement value at that facility's address, or folded into a blanket IT equipment limit written before you bought it?

This Month:

  1. Get three answers from your broker in writing: the business interruption waiting period, whether cover requires physical damage to your own property, and whether a utility-services or power-failure exclusion applies. Then walk a CHI1-style chiller failure through the policy and see what actually pays.
  2. Tell the underwriter how the racks are cooled. If the cluster moved to direct liquid cooling since the last submission, ask in writing whether that changes the terms.
  3. If your company has a captive, ask whether it can write the waiting-period layer and the gap between the colo cap and your deductible.

Before Renewal:

  1. Negotiate a carve-out from the cap for equipment damage caused by the operator's failure to deliver contracted power or cooling, at a limit tied to the hardware rather than to three months of rent.
  2. Read the release and subrogation clauses with your insurer on the call. A release that survives the operator's own negligence, combined with a subrogation waiver, means nobody recovers from the party that caused the loss. Decide whether that is a term you accept or one you trade for something.
  3. Ask for the operator's certificate of insurance and its property and liability limits for the building, so you know what stands behind the cap you did get.

The Bottom Line

Moving workloads to the cloud took hardware risk off enterprise balance sheets. Buying your own GPU cluster puts it back, in a building you do not control, under colocation terms that price space, power and cooling rather than what sits in the rack. The insurance market is telling you, through S&P's ratings analysts and Marsh's captive practice, that the product for that risk does not exist in standard form yet.

The colo sells you space, power and cooling. It does not sell you your GPUs back. Insure them as if it never will.

Continue Reading

Share:

Frequently Asked Questions

Does a colocation provider pay if its cooling failure damages my GPUs?

Usually only up to a cap. Colo agreements commonly limit the operator's total liability to about three months of fees, make service credits the main remedy, and some exclude damage to customer equipment outright. Caps typically lift only for gross negligence or willful misconduct, so the tenant's own property insurance carries most of the loss.

Is there an insurance product for GPUs in data centers?

Not a standard one. S&P told Bloomberg in September 2026 that there is no obvious coverage product for the high-value GPUs housed in data centers. Tenants typically insure their equipment under a property policy at replacement cost, which colo agreements require, and specialist brokers place bespoke GPU cover in the surplus lines market, but outage and business interruption cover has significant gaps.

Why might business interruption insurance not cover a data center outage?

Property policies generally require physical loss or damage before business interruption cover responds, some exclude utility or power-supply failures, and S&P says cover bought against power failure can carry a 12 to 24 hour waiting period. A cooling or power outage that damages nothing, or ends inside the waiting period, may pay nothing.

What should I negotiate in a colocation agreement for a GPU cluster?

Ask for a separate, higher liability limit for equipment damage caused by the operator's failure to deliver contracted power or cooling; review release and subrogation-waiver clauses with your insurer; confirm the service-credit ceiling; and request the operator's certificate of insurance and building property limits.

What is a captive insurer and why are data centers using them?

A captive is an insurance company a corporation owns to insure its own risks. Marsh expects explosive growth in captives taking on data center portfolio risks, and Aon says clients facing higher premiums manage losses by retaining the bottom layer of risk themselves.

Newsletter

Stay Ahead of the Curve

Weekly enterprise AI insights for technology leaders. No spam, no vendor pitches—unsubscribe anytime.

Subscribe

Latest Articles

View All →