Clockwork's GPU Failover Lands at LinkedIn With No Public Price

Clockwork raised $31M with LinkedIn running its GPU failover and Together AI set to resell it, while AWS bundles checkpointless recovery into HyperPod for free. Every recovery figure on offer was measured by the seller, so price failover on your own interruption rate.

By Rajesh Beri·October 8, 2026·9 min read
Share:
A row of open GPU server trays in a dim data centre aisle, one tray pulled halfway out with a red status light and an unplugged InfiniBand cable dangling, while a technician's hand slides a spare tray into the next slot.

Illustration generated using AI

If you run multi-node training, fine-tuning or RL jobs, GPU failover is now something you can buy from your GPU cloud. Whether to buy it depends on one number you probably have not measured: how often your own jobs get interrupted and what each restart costs. On October 5, Clockwork.io raised $31 million and named LinkedIn as a fleet-wide customer and Together AI as a reseller. AWS has bundled a competing approach into SageMaker HyperPod since December 2025 at no extra charge. Neither Clockwork's release nor Together AI's cluster pricing page lists a price, and every recovery figure on the market was measured by the company selling it.

The useful move this quarter is to pull your interruption log, put a dollar figure on each restart, and only then take the vendor call.

What Clockwork Announced and Who Is Running It

Clockwork's round puts fault-tolerance software in production at a named enterprise and in a GPU cloud's product catalog. The $31 million was co-led by Premji Invest, Wing Venture Capital and Seligman Ventures, with NEA and e& Capital participating, bringing total funding to $73 million, according to Pulse 2.0. The company's previous raise was a $21 million Series A in March 2022, RuntimeWire reports.

Two products matter. LinkPass reroutes network traffic around a failed link. TorchPass moves a job off a failing GPU onto healthy hardware instead of restarting it from the last checkpoint. Clockwork made TorchPass generally available on March 11, 2026, and its event material says the software sits between the training framework and the cluster scheduler, provisions a healthy spare and transfers the training state to it.

LinkedIn runs LinkPass across its AI fleet. Raghu Hiremagalur, its infrastructure CTO, credits Clockwork with preventing tens of thousands of GPU-hours of downtime each month across LinkedIn's AI infrastructure. He describes the failure it fixes: one InfiniBand NIC flap could take an eight-GPU server out of service, and a switch port flap could drain a second server, so one bad link idled 16 GPUs.

Together AI has a production deployment and plans to offer TorchPass as a service on its GPU Clusters. WhiteFiber is extending Clockwork across its GPU-as-a-service network and uses its fleet audit, which validates every link and node, for acceptance testing.

The release also adds multi-node snapshots that capture a distributed job's state without code changes, and background checkpoints that push updated weights to inference replicas during RL. RuntimeWire notes that LinkedIn's figure is a customer claim and that the release gives no contract values.


How Often Do GPU Training Jobs Actually Fail?

The best public answer is still Meta's. During a 54-day stretch of Llama 3 pretraining on 16,384 GPUs, Meta logged 466 job interruptions, 419 of them unexpected, and GPU issues accounted for 58.7% of the unexpected ones. Even so, Meta reported more than 90% effective training time, with significant manual intervention needed only three times.

That rate scales down to your cluster with simple arithmetic. 419 unexpected interruptions over 884,736 GPU-days is about one interruption per 2,100 GPU-days. If your hardware fails at Meta's rate (and remember that the hourly rate you pay is only part of what a GPU-hour costs you):

Job size Expected unexpected interruptions
64 GPUs about one a month
256 GPUs about one every 8 days
2,048 GPUs about one a day
16,384 GPUs about one every 3 hours

The table is derived from Meta's published count, not measured on your hardware. Newer GPUs, a different provider and a different job mix will move it in either direction. It still shows where the money is: the interruption rate grows with job size, and so does the number of GPUs idled by each restart.

What Does One Restart Cost?

AWS gives the cleanest worked example. A conventional checkpoint restart on a 2,304-GPU job took AWS 15 to 30 minutes because every process restarts, the process group rebuilds, the checkpoint reloads and the steps since that checkpoint get recomputed. On 256 P5 instances (2,048 GPUs) at $55 an hour with checkpoints every 20 minutes, AWS puts each disruption at about $4,693, and daily disruptions at about $141,000 a month plus roughly 10 hours of schedule slip.

A 256-GPU job (32 P5 instances, an eighth of AWS's example) failing at Meta's rate gets interrupted three or four times a month, which at an eighth of AWS's per-incident figure is under $2,500 of lost compute. A failover product for that job has to cost less than that, after its own overhead. At thousands of GPUs, Clockwork claims over $6 million a year in savings for a 2,048-GPU H200 deployment, an unverified vendor figure that is at least the right order of magnitude for a cluster interrupted daily.

SemiAnalysis modeled the same spread in April. For a 5,184-GPU GB300 pretraining run, it put goodput loss at 6.14% on a gold-rated neocloud, 10.53% on a hyperscaler and 20.91% on a silver-rated one. For a 2,048-GPU multimodal RL job the gap between providers shrank to 0.23% to 0.96%, and for inference the extra downtime touched about 0.5% of cluster cost. Goodput, in SemiAnalysis's definition, is the useful work a cluster delivers once failures, restarts and recomputed steps are subtracted from raw GPU-hours.


Buy, Bundle or Build: The Three Options on the Table

You now have three ways to cut restart losses, and the biggest difference between them is how much of your stack each one dictates.

The first is to buy it from a GPU cloud. Together AI will sell TorchPass as a service, though its GPU Clusters page lists H100s at $3.99 an hour on demand and $3.19 reserved with no TorchPass line, no SLA terms, and a general promise of "continuous health checks, automated remediation, and self-serve node repair." The quote in Clockwork's release from SemiAnalysis's Dylan Patel says TorchPass "cuts training goodput loss from 14% to under 3% for a gold-rated neocloud" in ClusterMAX testing. SemiAnalysis's own public modeling puts gold-tier loss at 6.14% for its large pretraining scenario, and the release does not say which job size or setup produced the 14% baseline. Ask for it.

The second is the bundled version. AWS shipped checkpointless and elastic training on HyperPod on December 3, 2025, at no additional cost. It recovers state from healthy peers instead of storage. AWS reports recovery on a 256-GPU Llama 3 70B job falling from 4 minutes 52 seconds to 47 seconds, and goodput above 95% on deployments over 2,300 GPUs. SemiAnalysis says its own test on a 4-node H200 cluster confirmed roughly 1 minute 45 seconds against about 15 minutes for a checkpoint restart. The catch is the stack: HyperPod on EKS, the HyperPod training operator v1.2 or later, at least two nodes, and NeMo, PyTorch or PyTorch Lightning.

The third is to build on open source. PyTorch's torchft isolates a failure to one replica group and lets the rest keep training, then refills the recovered group from a healthy peer. On 300 L40S GPUs lent by Crusoe, a failure injected every 60 seconds still left 81.2% training efficiency over about 1,100 failures. It is BSD-licensed and designed for replicated-weight setups such as DDP or HSDP. SemiAnalysis measured a performance difference of more than 10% against comparable HSDP jobs in initial testing.

The steel-man for buying: detection is half the problem, and most teams are bad at it. SemiAnalysis found that one provider's test system triggered no alert or node replacement over an 18-hour window. Software that watches every link, the way LinkPass does at LinkedIn, fixes a failure class that faster recovery never touches.

What Every Recovery Number Has in Common

Every figure above was produced by the party selling it or by its launch partner. LinkedIn's GPU-hours figure is a customer testimonial in a funding release. AWS's numbers come from AWS's internal studies. Clockwork's recovery time of under two minutes comes from vendor materials and has not been independently benchmarked. The only independent check in the public record is SemiAnalysis's single 4-node test of AWS.

Recovery time also leaves out what you pay to make recovery possible. TorchPass needs a spare GPU to move work onto, and SemiAnalysis reports that top-tier providers hold 2% to 6% of nodes as hot spares at 4,000-plus GPU scale. Someone pays for that idle capacity, and in a managed service it will show up in the rate whether or not it appears as a line item.


What to Do Before You Sign a GPU Contract

This Week:

  1. Pull 90 days of job logs for every multi-node training, fine-tuning and RL run. Count unexpected interruptions per 1,000 GPU-days and compare with Meta's rate of about one per 2,100 GPU-days.
  2. Time one real restart end to end: detection, process group init, checkpoint load, recomputed steps. Multiply by GPUs in the job and your hourly rate. That is your per-incident cost.

This Month:

  1. If your jobs run under a few hundred GPUs and interruptions are rare, write the decision down and stop. A faster checkpoint cadence is cheaper than any product.
  2. If you are already on SageMaker, run one job with checkpointless training enabled and record the recovery time on your own model. It costs nothing extra.
  3. Ask Together AI, or any neocloud you rent from, for TorchPass pricing and for goodput or recovery-time data measured on a job the size and shape of yours.

Before Renewal:

  1. Write recovery terms into the GPU contract: time to detect, time to replace a node, spare-pool commitment, and credits tied to goodput. In SemiAnalysis's example, a 95% uptime SLA lets a provider be down 5% of the time with no response and no credit.

The Bottom Line

GPU fault tolerance has moved from an in-house engineering project to something a provider bundles, as AWS does, or resells, as Together AI now will. For a buyer, the work shifts from building recovery to auditing a seller's claims about it, and nobody yet publishes independent goodput numbers across providers. If you already rent capacity, the commitment math in Reserved vs Spot GPUs and the provider comparison in GPU Clouds Compared should now carry a goodput line. Your own interruption log is the only number in that negotiation you did not get from the seller.

Count your interruptions before you price the fix.

Continue Reading

Share:

Frequently Asked Questions

What is Clockwork TorchPass?

TorchPass is Clockwork.io's fault-tolerance software for distributed training. When a GPU or node fails, it moves the affected part of the job onto a healthy spare instead of restarting the whole job from its last checkpoint. It became generally available in March 2026, and Together AI plans to offer it as a service on its GPU Clusters.

How often do large GPU training jobs fail?

Meta reported 419 unexpected interruptions during 54 days of Llama 3 pretraining on 16,384 GPUs, about one every three hours. Scaled linearly, that is roughly one interruption per 2,100 GPU-days: about one a month for a 64-GPU job and about one a day at 2,048 GPUs, though your hardware and provider will differ.

How much does a GPU training restart cost?

AWS estimates that on 256 P5 instances (2,048 GPUs) at $55 an hour with checkpoints every 20 minutes, each disruption costs about $4,693, because every process restarts, the checkpoint reloads and recent steps are recomputed. AWS says a conventional checkpoint restart on a 2,304-GPU job took 15 to 30 minutes.

Is SageMaker HyperPod checkpointless training free?

Yes. AWS launched checkpointless and elastic training on SageMaker HyperPod on December 3, 2025 at no additional cost. It requires HyperPod on EKS, the HyperPod training operator v1.2 or later, at least two nodes, and NeMo, PyTorch or PyTorch Lightning.

Should I buy GPU failover software?

Only after measuring your own interruption rate and per-restart cost. SemiAnalysis modeled goodput loss of 6% to 21% across providers for a 5,184-GPU pretraining run but under 1% of difference for a 2,048-GPU RL job, so small or rarely interrupted jobs may be better served by a faster checkpoint cadence.

Newsletter

Stay Ahead of the Curve

Weekly enterprise AI insights for technology leaders. No spam, no vendor pitches—unsubscribe anytime.

Subscribe

Latest Articles

View All →