Nebius bought Inferize to stop customers paying for GPUs that sit warm and idle, but your bill doesn't change yet. Inferize says it restores a loaded model in seconds, not minutes, so an endpoint could scale down without the usual penalty. Nebius has published no ship date, no price change and no capacity guarantee. Nebius announced the deal on October 1, 2026. Its own Token Factory billing policy still warns that once capacity is released, it "may not be immediately available again." Until restore time and capacity are both written into your contract, do not cut minimum replicas on the strength of this acquisition.
That is the stake for anyone running open-weight models on dedicated endpoints, whether on Nebius, Baseten, Fireworks AI or Together AI. The warm replica you keep running at 3 a.m. is an insurance premium. Inferize sells cheaper insurance. Nobody has told you yet what it covers.
What Did Nebius Actually Buy?
Nebius bought a 17-person, ten-month-old Tel Aviv startup that snapshots a running inference engine and restores it onto fresh GPUs. Per CTech's report, Inferize was founded in early 2026 by Guy Bortnikov and Lior Gorbonos, both from Granulate's founding team. It raised a seed round led by TLV Partners and was still in stealth. Nebius did not disclose terms. CTech estimates the price at $100-150 million, and RuntimeWire notes Globes put it at $100-130 million.
A cold start is the time a new model replica takes from launch until it can serve its first request: pulling weights, loading them into GPU memory, compiling kernels, capturing CUDA graphs. On a large model that takes minutes, which is why buyers keep spare replicas running.
Inferize works one layer down. As RuntimeWire describes it, the system captures a warmed serving engine's CPU and GPU state, compiled kernels and CUDA graphs included, then restores that state onto GPUs in seconds. Neither the serving engine nor the application has to change. The stated uses are scaling replicas up when traffic rises, releasing them when it falls, and recovering after a spot or preemptible node is reclaimed.
Bortnikov's pitch in the Nebius release is blunt: "Keeping spare GPUs running is the price of being ready for demand. Removing that cost is what we built Inferize to do."
Is the 20-Second Restore Real?
It is one vendor benchmark on one model, and there is no independent measurement of it. RuntimeWire reports that in an Inferize-published DeepSeek V4 Pro test, a snapshot started in 20 seconds against roughly 23 minutes for a cold start, about a 70x speedup. The same report quotes Inferize's site claiming 10-30% GPU cost reductions. The Nebius announcement itself publishes no benchmark figures. Inferize's own homepage is a waitlist page with no numbers on it.
Steel-man it first. The technique is real, and other vendors have shipped it. Modal's GPU memory snapshots, released in alpha on July 30, 2025, cut a vLLM cold start on Qwen2.5-0.5B from 45 seconds to 5 by restoring device memory, kernels and compiled state instead of rebuilding them. If a 0.5B model goes from 45 to 5 seconds, restoring a frontier-scale MoE in tens of seconds is plausible.
Here's what the headline number leaves out. Modal shipped as an alpha and said it was "still exploring the limitations of this feature." It depends on NVIDIA driver branches 570 and 575. A restore is only as fast as the GPUs you land on, and that is where the Token Factory documentation gets in the way.
Why Your Minimum Replicas Stay Where They Are
A 20-second restore is worthless if there's no GPU to restore onto, and Nebius's own docs say there might not be. The Token Factory billing policy is precise on three points:
- "As long as your endpoint remains active, minimum replicas remain allocated to you."
- "When zero replicas are running, the endpoint is not accessible and billing charges do not apply."
- "Once capacity is released, it may be reassigned and may not be immediately available again."
The best-practice guide goes further: "Max replicas scaling depends on available capacity and may not always be continuously available after scale-down." It recommends higher minimum replicas when "cold starts are unacceptable."
The arithmetic follows. Today, minimum replicas buy you two things: a warm engine and a reserved slot. Inferize addresses only the first. Scale down to save money and you give the slot back. When the morning spike arrives, the question is whether a GPU is free, and fast restore doesn't answer it.
Rivals frame the same trade the same way. Baseten's cold-start documentation tells production users to "set min_replica to 1 or higher so a replica is always running," and to plan surges from the p99 startup time, not the mean. Its tools for shortening startup are weight and image caching plus torch compile caching.
What Does a Warm Replica Cost This Week?
More than it did last week. Nebius's higher list GPU rates took effect on the same day it announced the deal. The Nebius price page lists on-demand H200 rising from $4.50 to $5.40 per GPU-hour, H100 from $3.85 to $4.50, and B200 from $7.15 to $8.50, all effective October 1, 2026. Those are AI Cloud compute rates. Token Factory dedicated-endpoint pricing is quoted separately, so treat the figures below as an illustration, not your invoice.
Take an eight-GPU H200 replica held as a minimum around the clock: 8 × $5.40 × 730 hours is about $31,500 a month at the new list rate, against about $26,300 at the old one. If Inferize's claimed 10-30% saving held for that replica, it would be worth roughly $3,150 to $9,450 a month, and that is a vendor's claim applied to a list price. The same page lists preemptible H200 capacity from $0.79 an hour, which is where fast restore matters most: recovering a reclaimed node in seconds is what makes cheap, interruptible capacity usable for serving.
So Nebius's incentive is straightforward. Faster restore lets it pack more customers onto the same GPUs, and that saving lands on Nebius's margin first. Whether any of it reaches your bill depends on whether you negotiate for it. That pattern already showed up when Anthropic walked from its Decart bid: an efficiency acquisition reprices the vendor's costs, not your rate card.
Third Inference Buy, Same Silence on the Contract
Inferize is Nebius's third inference acquisition in five months, and none of the three announcements changed a customer-facing term. Eigen AI was announced on May 1, 2026 at roughly $643 million in cash and stock, for model-level work like quantization, KV-cache optimisation and custom CUDA kernels. On May 12, Nebius took Clarifai's core team and licensed its inference and orchestration stack, with founder Matthew Zeiler joining as SVP of Research. The press release disclosed no terms. Nebius's half-year filing puts the Clarifai consideration at $97.4 million, $75.0 million of it cash at closing, and records Eigen AI's purchase consideration at $331.4 million when it closed on June 10.
Read together, the strategy is coherent: model optimisation from Eigen, system orchestration from Clarifai, elastic capacity from Inferize. Token Factory, which launched in November 2025, is being assembled into a full inference stack. What has not changed is the contract you sign. The Inferize release says only that its engineers will start "with the integration of their technology," with no completion date.
For customers, then, this is a roadmap item. It isn't a feature yet. The same applies to the deals we covered when Baseten bought Blaxel and LiveKit bought Loophole Labs, the latter also a cold-start play.
What to Do Before Your Next Inference Renewal
Treat fast restore as a term to negotiate, not a reason to cut capacity. If you run dedicated endpoints on Nebius or a rival, the work splits into three horizons.
This Week:
- Pull your p99 replica startup time per model from your provider's metrics, not the mean. That is your current cold-start exposure, and the number any vendor promise has to beat.
- Price your idle floor. Multiply minimum replicas × GPUs per replica × the October 1 rate × 730. Put the figure next to your peak-hour traffic so you can see what the insurance costs per covered request.
This Month:
- Ask your Nebius account team three written questions: when snapshot restore reaches Token Factory dedicated endpoints, whether it covers your model and GPU type, and whether a scaled-down endpoint gets any priority when it scales back up. Read the billing policy's "may not be immediately available again" back to them.
- Run your own restore test where one exists. If your stack already runs on a platform with GPU snapshots, time a restore on your largest model and compare it with the 23-minute cold-start baseline. A vendor's 70x is not your 70x. The guidance in our vLLM vs SGLang comparison applies: test on your model, at your batch size.
Before Renewal:
- Write two numbers into the order form: a maximum restore time from zero (or from your minimum) to serving, and a capacity-reacquisition commitment within that window. Without the second, the first is a benchmark, not a guarantee.
- Tie the saving to your rate. If Nebius's cost to serve you drops by the 10-30% Inferize claims, ask for a lower minimum-replica rate or a credit for scaled-down hours. Our GPU hour pricing breakdown shows how far a list rate drifts from what you actually pay.
The Bottom Line
Snapshot restore is the right idea arriving at the wrong layer of the contract. Every autoscaling feature teaches the same lesson: making a replica start faster is not a promise that there is a GPU to start it on. The buyers who come out ahead negotiate the speed and the capacity together. GPUs are scarce, and Nebius's list-price rise took effect the same day it bought the fix.
Inferize may well cut a 23-minute wait to 20 seconds. The 20 seconds only counts if a GPU is waiting at the end of it.
Continue Reading
- GPU Clouds Compared: Nebius on Price, CoreWeave at 3AM
- GPU Hour Pricing: A $3.44 Rate Cost $68.80 to Use
- LiveKit Buys Loophole Labs, Leaving Architect Buyers Without a Word
- Baseten Buys Blaxel, Whose 'Keeps Running' Pledge Has No End Date
- Run:ai vs Kubernetes vs Slurm: Static Quotas Strand Your GPUs
- DeepSeek Will Raise Prices. Your Ceiling Is Already 4x.
