DeepSeek said it would raise its API prices, and refused to say by how much or when. That was the entire content of the notice. Every run-rate you built on its price floor became an estimate with no error bars.
Here is the part the wire reports missed. Because the weights are open, you did not have to wait for the number. Other companies already serve the identical model and already publish what it costs them to do it. That was your ceiling, and it was worse than "3x to 4x" — on cache hits, which is where agentic workloads live, it was 40x.
Update — August 16, 2026: The number landed, and it takes effect today. DeepSeek's pricing page now states that new rates apply from 16:00 UTC on 16 August 2026, tiered by time of day: peak is 01:00–04:00 and 06:00–10:00 UTC, and off-peak is exactly half of peak. V4-Pro output goes from $0.87 to $3.96 at peak — 4.55x, just past the 4x ceiling this article computed from third-party hosts a week earlier. The cache-hit line was worse, as this article argued it would be: V4-Pro cache-hit input goes from $0.003625 to $0.044, a 12.1x increase, and that is the line agentic workloads pay on every step after the first. Most importantly, the hedge has inverted. At peak, DeepSeek's own $3.96 output rate is now higher than Fireworks at $3.48 and DeepInfra at $2.60 for the identical MIT-licensed weights. The sections below carry the landed rates.
What DeepSeek Actually Posted on August 6
DeepSeek published a warning, not a price. Its official API pricing page now carries the line: "We plan to raise the overall pricing for DeepSeek API services in the near future, with a significant increase expected," followed by "The specific pricing plan will be subject to official notice."
No magnitude. No effective date. No stated reason. TechNode reported it on 6 August 2026, and The Next Web's read is the obvious one — the low prices drew a surge of users, and serving them strains the compute behind the service. The South China Morning Post framed it as a company struggling to hold an aggressively low price against fierce competition. For ten days nobody had a number, because DeepSeek had not published one.
The rates in force at the time were the baseline that moved. DeepSeek V4-Flash was $0.14 per million input tokens on a cache miss, $0.28 output, and $0.0028 on a cache hit. V4-Pro was $0.435 in, $0.87 out, $0.003625 on a cache hit. Every multiple in this article is measured against those six numbers.
This was also the second pricing change in about five weeks. On 1 July DeepSeek announced peak-hour surcharging — every billing item at 2x during 9am–noon and 2pm–6pm Beijing time, with the stated goal of "better distribution of resources." At the time that surcharge had not reached the pricing page, and the posted rates were still flat. It has now landed, on exactly those windows and at exactly that ratio — off-peak is half of peak — but folded into a much larger overall increase rather than applied to the old rates. A vendor that announces time-of-day surcharging and then warns of an overall increase inside the same quarter is not making a one-time adjustment. It is repricing.
The Rates That Actually Landed
DeepSeek published the numbers last week and switched them on at 16:00 UTC on 16 August 2026 — midnight Beijing on the 17th, which is why some wires date it a day later. PyMNTS characterised the move as "quadrupling" the previous levels. On output that is right. On cache hits it is a considerable understatement.
Per million tokens, against the flat rates above:
| Line item | Old flat | New peak | New off-peak | Peak multiple |
|---|---|---|---|---|
| V4-Pro cache-hit input | $0.003625 | $0.044 | $0.022 | 12.1x |
| V4-Pro cache-miss input | $0.435 | $1.32 | $0.66 | 3.0x |
| V4-Pro output | $0.87 | $3.96 | $1.98 | 4.6x |
| V4-Flash cache-hit input | $0.0028 | $0.014 | $0.007 | 5.0x |
| V4-Flash cache-miss input | $0.14 | $0.44 | $0.22 | 3.1x |
| V4-Flash output | $0.28 | $1.32 | $0.66 | 4.7x |
Two things to take from that table before anything else. The worst line is cache-hit input on V4-Pro, at 12.1x — a 1,114% increase, and the only line that broke out of the 3-5x band. And the tiering is not a surcharge bolted onto the old prices; it is a new price list where off-peak is defined as half of peak, so even the discounted tier is 2.3x the old output rate and 6.1x the old cache-hit rate.
Open Weights Mean the Reference Price Is Public
You could price the worst case eight days before DeepSeek published anything, because the model DeepSeek sells is the same model anyone can download. DeepSeek-V4-Pro on Hugging Face is MIT-licensed — 1.6 trillion total parameters, 49 billion activated, 1M context. V4-Flash-0731 is MIT-licensed too, at 304B total parameters. MIT means commercial use, modification and redistribution with no restriction and no negotiation.
So look at what companies that must actually cover their own compute charge for it. Fireworks AI's serverless price list puts DeepSeek V4-Pro at $1.74 input / $0.145 cached / $3.48 output per million tokens. DeepInfra lists the same model at $1.30 / $0.10 / $2.60.
Against DeepSeek's old first-party $0.435 / $0.003625 / $0.87, that was 3x at DeepInfra and exactly 4x at Fireworks on both input and output.
That number was a model, not a guess — and it held. DeepSeek landed at 4.55x on V4-Pro output at peak, within 14% of the Fireworks-derived estimate, eight days before it published anything.
But read the same table again against the rates that actually landed, because the relationship reversed:
| V4-Pro, per 1M tokens | DeepSeek peak | DeepSeek off-peak | Fireworks | DeepInfra |
|---|---|---|---|---|
| Cache-hit input | $0.044 | $0.022 | $0.145 | $0.10 |
| Cache-miss input | $1.32 | $0.66 | $1.74 | $1.30 |
| Output | $3.96 | $1.98 | $3.48 | $2.60 |
On output tokens during peak hours, the first-party API is now the most expensive of the three. DeepSeek charges $3.96 where Fireworks charges $3.48 and DeepInfra $2.60 for weights anyone can download. On cache-miss input at peak it has drawn level with DeepInfra to within two cents. Only cache-hit input is still decisively cheaper first-party — $0.044 against DeepInfra's $0.10 — and off-peak DeepSeek is still the cheapest option on every line.
State that split precisely, because it is the whole decision: on V4-Pro, peak-hour output is the only place where porting off DeepSeek now saves money outright. Every other V4-Pro line is still cheaper first-party, and the gap is a schedule problem rather than a vendor problem. Flash is a different story, below.
One caveat on shopping the marketplaces: on 8 August OpenRouter listed V4-Pro at $0.338 / $0.676, below DeepSeek's own posted rate at the time, with an 80% discount flagged on the listing. A promotional routed rate is not a contracted rate. Do not build a run-rate on one, and re-check it now that the underlying first-party price has tripled.
The Cache-Hit Gap Is Where the Damage Is
The 4x is the headline number and it is the smaller problem. The cache-hit price is where a repricing actually lands, and that is exactly where this one landed hardest — 12.1x on V4-Pro, against 3.0x on the cache miss sitting next to it.
Compare cache-hit input, per million tokens:
| Model | DeepSeek, old flat | DeepSeek peak | DeepSeek off-peak | Fireworks |
|---|---|---|---|---|
| V4-Pro | $0.003625 | $0.044 | $0.022 | $0.145 |
| V4-Flash | $0.0028 | $0.014 | $0.007 | $0.028 |
The 40x gap to Fireworks that this article flagged on 8 August has closed to 3.3x at peak — not because Fireworks moved, but because DeepSeek climbed most of the way to it on the one line it had been discounting hardest. DeepSeek used to price a cache hit at roughly 1/120th of a cache miss on V4-Pro. It now prices one at 1/30th. Fireworks prices one at 1/12th. The economic models are converging, and the convergence is being paid for from one direction.
This matters more than it looks because agentic workloads are cache-heavy by construction. A multi-turn agent replays a large, stable prefix — system prompt, tool definitions, retrieved context — on every single step, and that prefix is a cache hit every time after the first. We have written before about how prompt caching quietly restructured agent economics and made per-token list prices a poor proxy for what you actually pay. This is the same trap, pointed the other way.
Run the number for your own traffic, and run it twice — once at peak and once off-peak. At a 70% cache-hit ratio, a move from DeepSeek to Fireworks on V4-Pro input used to cost you 4.7x blended; against the new peak rates it costs 1.5x, and against off-peak 2.9x. At a 90% hit ratio it was 6.5x, and is now 1.8x at peak and 3.6x off-peak. The lower your cache-miss share, the worse the port still looks — the direction of that gradient has not changed, and it is still the opposite of the intuition most FinOps models encode. What changed is the size of the penalty. For a heavily cached agentic workload running in peak hours, the third-party hedge went from unaffordable to a rounding error.
Your Flash Hedge Is Now Cheaper Than DeepSeek
Split your workloads by model before you do anything else, because the exposure is wildly asymmetric — and on Flash it has flipped outright.
On 8 August, Fireworks served V4-Flash-0731 at $0.14 / $0.028 / $0.28 — identical to DeepSeek's first-party cache-miss and output rates, to the cent. Flash is a 304B model with a small activated parameter count; it is cheap enough to serve that a commercial host matched DeepSeek without subsidy. That was the argument for calling Flash "already hedged."
Fireworks has not moved. DeepSeek has, and the par is gone: Flash cache-miss input is now $0.44 at peak and $0.22 off-peak against Fireworks' $0.14, and output is $1.32 / $0.66 against $0.28. For Flash cache-miss and output traffic, the third-party host is now unambiguously the cheaper seller — 4.7x cheaper on peak-hour output — for the same MIT-licensed weights on a base-URL change. Cache-hit input is the one exception, still 2x cheaper first-party at peak and 4x off-peak.
V4-Pro remains the harder problem, but for a different reason than a week ago. 1.6 trillion parameters with 49B activated is expensive to serve, and every independent host priced it well above the old first-party rate: 3x at DeepInfra, 4x at Fireworks and Baseten, and about 5x at Together AI, which lists $2.10 input / $0.20 cached / $4.40 output. DeepSeek has now climbed into the middle of that spread rather than sitting under it — above Fireworks on peak output, still below Together. It is also, in most shops, the workload doing the reasoning and the agent orchestration, the expensive and hard-to-substitute part.
The practical read has inverted with the prices: your Flash spend is now the easy win, and your Pro spend is a scheduling decision. If you have not separated them in your cost reporting — by model, by cache ratio, and now by hour — you cannot size this at all.
Where Your Team Sits Now Sets Your Bill
Every wire report covered the peak windows as a number. None of them covered the part that decides what you actually pay: the windows are pinned to Beijing business hours, and your team is not in Beijing.
Peak is 01:00–04:00 and 06:00–10:00 UTC — 09:00–noon and 14:00–18:00 in Beijing, the domestic working day DeepSeek is trying to flatten. That is seven hours of the clock, so 71% of every day bills at the off-peak half rate. Where those seven hours land for you is pure geography:
| Your team | Peak window one | Peak window two |
|---|---|---|
| US Eastern (EDT) | 21:00–00:00 | 02:00–06:00 |
| US Pacific (PDT) | 18:00–21:00 | 23:00–03:00 |
| UK (BST) | 02:00–05:00 | 07:00–11:00 |
| Central Europe (CEST) | 03:00–06:00 | 08:00–12:00 |
| India (IST) | 06:30–09:30 | 11:30–15:30 |
For a US team this is close to the best possible outcome and nobody has told them. The entire American working day is off-peak — 09:00 to 18:00 Eastern is 13:00 to 22:00 UTC, which does not touch either window. Interactive traffic, developer tooling, anything a human is waiting on: half price, no action required.
Then there is the trap. The conventional overnight window is the single worst hour to run anything. The habit of scheduling batch inference, nightly evals and reindexing for 02:00 local puts a US East Coast team dead-centre in the 06:00–10:00 UTC peak, and a Pacific team inside 23:00–03:00. The workload with the least need to run at a particular hour is the one most likely to be scheduled straight into the 2x tier — and unlike interactive traffic, it can be moved with a cron edit.
European and Indian teams get the reverse problem and cannot cron their way out of it. Peak sits inside the working morning — 08:00–12:00 CEST, 11:30–15:30 IST — so their interactive traffic pays double while their overnight batch runs free. If you are running a follow-the-sun engineering org, the same workload now costs a different amount depending on which office kicked it off, and none of your existing cost attribution captures that.
Self-Hosting Is Not the Hedge Most People Think
Renting your own GPUs is still a worse deal than the API for most people, and the arithmetic is short enough to check in your head — but the price rise moved the break-even by nearly 5x, so re-run it rather than trusting the old answer.
The V4-Flash-0731 model card documents deployment on a single 4×GB300 node under vLLM or SGLang. Fireworks publishes on-demand GPU rates of $12.00/hour for a B300 288GB — so call that node $48/hour. At DeepSeek's old flat V4-Flash output price of $0.28 per million tokens, $48 bought about 171 million output tokens, and breaking even on output alone meant sustaining roughly 47,600 output tokens per second for the entire hour. No four-GPU node does that. That was the whole argument.
At the new rates the bar drops hard. At the off-peak $0.66, $48 buys 72.7 million output tokens and the break-even is 20,200 tokens per second; at the peak $1.32 it is 10,100. Counting the input spend you would also stop paying lowers it further — at a heavy 10:1 input-to-output mix on cache misses, to roughly 4,700 tokens per second off-peak and 2,300 at peak. That last number is no longer obviously out of reach for a 304B model with a small activated parameter count, which is a different conclusion than the one this article reached on 8 August. On the cheapest reserved capacity — GB300 listings start around $5.60/hour per GPU on a one-year commitment — output-only break-even falls to about 9,400 tokens per second off-peak and 4,700 at peak.
Three caveats keep this from being a green light. The break-even still collapses back toward the output-only figure the moment your prefix is cached, because a cached input token is $0.007 off-peak. The same listing page notes most GB300 capacity is out of stock or waitlisted, and a hedge you cannot provision is not a hedge. And these are break-even points against list price, assuming you sustain that throughput every second of every hour you are paying for — real utilisation is what kills self-hosting, not peak throughput. Benchmark your own node before you believe any of it.
Self-hosting becomes rational at genuine scale, with high sustained utilisation, or when a compliance requirement makes the API unusable at any price. This is the same wall we hit when AMD's inference discount turned out to depend on a GPU nobody could rent, and the same reason NVIDIA's free 550B Nemotron did not make self-hosting free. Open weights remove the licence cost. They do not remove the GPUs.
For most readers the real hedge is a third-party host of the same weights. Price Fireworks AI, Together AI, Baseten and DeepInfra against your actual token mix, not against a list price. Check availability before you plan around a host: Amazon Bedrock currently lists DeepSeek V3.1, V3.2 and R1 and not V4, so it is not a like-for-like exit today.
The Port That Also Closes a Compliance Gap
Moving off DeepSeek's first-party API fixes a problem your security team probably already raised. DeepSeek's privacy policy states plainly: "To provide you with our services, we directly collect, process and store your Personal Data in People's Republic of China." There is no region selector.
For a regulated workload that is not a preference, it is a blocker — and for those teams the first-party endpoint was never available in the first place, which makes a Western host of the same weights the only version of DeepSeek they could ever have run. We covered the supply-chain and jurisdiction side of this when Chinese models crossed 46% of routed enterprise token volume, and the three risks that survive the 90% cost saving have not changed.
So the hedge you are being forced into for cost reasons is the one your governance function wanted anyway. That is unusual, and as of this week it is also nearly free. The premium for a Western host of these weights was 3-4x on 8 August; against the landed peak rates it is roughly par on V4-Pro cache-miss input, a discount on V4-Pro output, and a large discount on Flash output. The compliance argument no longer has to be paid for out of the infrastructure budget — which removes the last reason most teams had for deferring it.
The Steel Man: This Repricing Was Always Coming
DeepSeek's prices were not a market clearing price, and pretending otherwise was the modelling error.
Look at the volume-versus-spend split. Vercel's July 2026 AI Gateway index — one gateway's routed traffic, reporting June data, not a market-wide measure — put DeepSeek at 22.6% of gateway token volume, third behind Anthropic and Google, while its share of spend was negligible. Open-weight models overall took 29% of volume against just under 4% of spend. The June index, covering May, had DeepSeek at 17% of tokens and roughly 1% of spend, against Anthropic's 32% of tokens and 65% of spend.
Treat that as one well-instrumented sample rather than the market — Vercel's gateway skews toward its own developer base, and no comparable industry-wide split is published. But a provider carrying a fifth of a major gateway's tokens for approximately none of its revenue is running a land-grab, not a business model. Nobody sustains that indefinitely. The honest version of the criticism is that buyers who treated $0.435 as a durable input to a three-year plan made the mistake, not DeepSeek.
There is a wider consequence too. That same July index notes closed-weight frontier prices rose about 12% per token, and the industry's average price per token stayed flat only because cheap open-weight volume grew to offset it. DeepSeek's volume is a meaningful part of that offset. That was written as a conditional on 8 August: if DeepSeek reprices toward what independent hosts charge, the flat line stops being flat. It has now repriced past them on peak-hour output, so the condition is met and then some — and every multi-model routing strategy built on a cheap fallback tier lost its headroom on the same day.
Do This Before Your First Peak-Hour Invoice
The rates are live as of 16:00 UTC today. Your next invoice already reflects them, so the sizing work is no longer a scenario exercise.
This Week:
- Split your DeepSeek spend by model, by cache-hit ratio, and now by hour. Pro versus Flash, hit versus miss, peak-UTC versus off-peak, per workload. The hour dimension is new and no existing cost report has it; without all five numbers you cannot size the exposure, and every step below depends on them.
- Re-run your run-rate at the landed rates, not the multiples. $0.044 / $1.32 / $3.96 on V4-Pro at peak and half that off-peak; $0.014 / $0.44 / $1.32 on Flash. Take that number to whoever signed off on the original business case this week — a 12.1x line item does not survive being discovered in an invoice.
- Move every batch job out of 01:00–04:00 and 06:00–10:00 UTC. This is a cron edit that halves the bill on the affected workloads, and for a US team it costs nothing at all — the entire American working day is already off-peak. Do this before the harder work below, because it is the only free item on the list.
- Send one workload's traffic to a second host today. Fireworks, DeepInfra, Together or Baseten. Not a proof of concept — a real slice of production traffic, so you learn what actually breaks before you need the exit. On Flash output and peak-hour Pro output, the second host is now the cheaper seller, so this step pays for itself rather than buying insurance.
This Month:
- Confirm the weights are actually identical. Third-party hosts pin a checkpoint; DeepSeek's first-party endpoint has moved under people before. That is exactly how a silent post-training swap invalidated agent evals without a version bump. Pin the checkpoint by name in your config and re-run your eval suite against each candidate host.
- Price the cache economics, not the sticker. Ask each host directly: what is the cache-hit rate, what is the TTL, and is it per-tenant. Cache-hit input is now the only line where DeepSeek is decisively cheapest, so it is the line that determines whether you stay. A host with no prefix caching at any price is disqualified for agentic traffic regardless of its input rate.
- Put a repricing clause in anything you sign. Any host, any term. Notice period, cap on increase, right to terminate without penalty. Ten days between "we plan to raise prices" and a 12.1x line item is the argument; make it at the table. This is the same class of exposure as an acquirer repricing your inference credits mid-contract — you get the protection at signature or you never get it.
Before Renewal:
- Set a hard trigger. Decide now, in writing, at what blended cost per million tokens you move each workload, and who executes it. A threshold agreed under no time pressure is worth more than a review scheduled for after the invoice — and assume the next change arrives with the same notice this one did.
The Bottom Line
The valuable thing here is not that DeepSeek raised prices. Vendors raise prices. The valuable thing is that this was a frontier-tier repricing where the buyer could compute the worst case eight days before the vendor announced it — because the weights are MIT-licensed and the substitutes publish their rates. The estimate came in at 4x on V4-Pro output; the answer was 4.55x.
One honest correction to how this article put it on 8 August. Open weights do not give you a ceiling — DeepSeek priced its own peak-hour output above Fireworks and DeepInfra for the identical model, so the substitute price is not a cap on what the originator can charge. What open weights give you is a published reference price and a live second seller, which is the more useful of the two: a ceiling tells you how bad it can get, a second seller lets you leave. This week that distinction stopped being academic, because leaving is now the cheaper option on several lines rather than the expensive insurance policy.
Every closed-model dependency in your stack has the same exposure and neither property. When OpenAI or Anthropic sends the next pricing email, there is no second seller of those weights to check against, no reference price to have computed in advance, and the number you get is the number you pay.
DeepSeek gave you ten days' notice and no figure. It also, by publishing its weights, gave you somewhere to go. Use the second one.
Continue Reading
- DeepSeek Swapped the Model. Your Eval Didn't Notice.
- Your AI Router Is Trading a 10x Discount for a 2.5x One
- AMD's Inference Discount Depends on a GPU You Can't Rent
- 46% of Your AI Now Runs on Chinese Models
- Claude vs GPT vs Gemini: Stop Comparing Per-Token Prices
- Free 550B Model: NVIDIA Ends Self-Hosted AI Quality Gap
- Bending Spoons Bought Airtable. Reprice Before It Closes.
- Chinese AI: 90% Cheaper—3 Risks Every CIO Can't Ignore
