Nvidia holds $279 billion in supply commitments, an increase it attributes primarily to memory, and can raise its own list prices. It is still guiding away three to four points of gross margin because of RAM. That is what the memory crunch costs the buyer with the most leverage on earth. Whatever number is on the server quote sitting in your procurement queue — placed through a channel partner, backed by no long-term agreement — is worse than Nvidia's, and it is probably already out of date.
If you built a self-hosted inference or vector-database business case on hardware prices from 2025, that model is understated. Not by a rounding error. On a memory-heavy build, the memory line has moved by a multiple.
What Nvidia Told the Market on August 26
Nvidia reported a record quarter and then spent the call explaining why its margins are going down anyway. For the quarter ended July 26, 2026, Nvidia reported revenue of $96.2 billion, up 106% year over year, with data center revenue of $89.0 billion and both GAAP and non-GAAP gross margin at 75.0%.
Then the guidance. Q3 comes in at 74.0% plus or minus 50 basis points. CFO Colette Kress told the call that margins bottom in Q4 at 71% to 72% and settle at 72% to 73% in fiscal 2028 — and that recovery depends on Nvidia's own price increases taking effect. Her words on the input cost were not hedged: the company is "experiencing extreme pricing conditions in memory," and "the magnitude of the price increase has exceeded our prior expectations and are headed even higher into next year."
The balance sheet says the same thing louder. Nvidia's supply commitments more than doubled from $119 billion to $279 billion in a single quarter, an increase the company attributed primarily to procuring memory. Inventory went from $25.8 billion to $31.6 billion. Jensen Huang's summary of the supply chain was that "everybody is really running flat out."
This matters to you for one reason. It is a rare case of the cost pressure being quantified by a buyer with published financials, contractual supply, and no incentive to exaggerate — Dell and HPE have both flagged memory inflation on recent calls without putting a margin number on it.
The Long-Term Agreement Is the Entire Story
The reason Nvidia's number is a floor and not a ceiling is that Nvidia has contracts and you do not. A long-term agreement, or LTA, is a multi-year supply contract that fixes both volume and price — and in this market it is the only instrument that stops a supplier repricing you every ninety days.
TrendForce is explicit about who that protects. Several US cloud service providers have signed multi-year LTAs that "restrict suppliers from raising prices for these clients." The consequence, stated in the same note, is that from Q3 2026 the primary source of server DRAM price increases shifts toward customers without LTAs, plus incremental volume sold outside an LTA to customers who have one.
Read that again as a buyer. Supply is fixed, demand is not, and a growing share of the buyers is contractually shielded. The increase has to land somewhere. It lands on the mid-market enterprise placing a 40-node order through a reseller.
The compounding is the part people underestimate. Conventional DRAM contract prices rose 93% to 98% quarter over quarter in Q1 2026, with TrendForce forecasting 58% to 63% in Q2 and a further 13% to 18% in Q3. Stack three quarters of that and the multiplier is roughly three and a half times. NAND ran a similar path — up 70% to 75% in Q2 2026, another 10% to 15% in Q3 — so your NVMe line moved too.
Memory Went From an Eighth of the Server to Half
The share of a server's bill of materials that is memory has roughly tripled in eighteen months, which is why re-pricing the memory line is not a line-item adjustment but a re-quote. One published index that models a reference dual-socket 2U build — two mid-range CPUs, 512GB of DDR5 as 8 × 64GB RDIMM, four NVMe SSDs — puts memory at about 18% of BOM in January 2025 and about 53% by July 2026, with the memory line moving from roughly £1.6k to £9.9k and the whole build index going from 100 to 205. Two caveats, both load-bearing. That index is a model built on published TrendForce contract moves, not a measured price feed, and its authors say so. And it is published by Servnet, a UK server reseller rather than an independent analyst house — a firm whose business improves when buyers treat a quote as urgent. Treat it as a planning shape from an interested party, not a quotation.
For a sanity check against live data: a public DRAM tracker derived from the DRAMeXchange spot board put a 64GB DDR5 RDIMM at $1,750, or $27.34 per GB, on August 4, 2026 — up 455% year over year. Spot runs hotter than contract and that figure is a derivation rather than a vendor quote, so do not put it in a business case. Put it in the row that tells you which direction contract is being dragged.
Gartner's read is consistent. Director analyst Shrish Pant told CRN that spot memory pricing is "just crazy", with some buyers paying three to four times prior-year rates, that combined DRAM and SSD prices rise about 130% by the end of 2026, and that stabilisation does not arrive until roughly Q1 2027. He also made a prediction worth writing down: businesses will extend hardware end-of-life by about 15% to avoid the refresh — which converts a procurement problem into a patching problem.
Your GPU Line Is Repriced. Your Everything-Else Line Isn't.
The most common mistake right now is checking the accelerator quote, seeing it unchanged, and concluding you are fine. The HBM inside a GPU is bundled into the accelerator's price, and Nvidia has told you plainly it intends to raise that price in fiscal 2028 to recover margin. That increase will arrive as a new SKU price, visibly, on a schedule.
The commodity DDR5 and NVMe in every other node arrives differently — silently, at re-quote, on hardware nobody thinks of as AI infrastructure. And a production retrieval stack is mostly those nodes. The vector database tier, the ingest and parsing workers, the metadata store, the observability and trace retention, the Kubernetes control plane, the CPU RAM holding offloaded KV cache. If you priced a self-hosted vector database against a managed one, the entire cost advantage of self-hosting lived in commodity RAM and disk. That is the exact line carrying the three-and-a-half-times move — one quarter of it realised, two still forecast.
The same applies to any end-to-end RAG cost model and to capacity planning for self-hosted inference runtimes, where the memory-per-replica assumption is what sets node count. None of this makes self-hosting wrong. It makes the crossover point move — and if you have not recomputed it since 2025, you do not currently know where it is.
Managed endpoints are not automatically the escape hatch either. Providers buy from the same three suppliers, and the pass-through shows up as instance and managed-service pricing rather than a labelled surcharge. If your fallback plan is "we'll just run it on Bedrock" or a neocloud, price that plan now rather than assuming it holds. The capacity-versus-capex question has a different answer this quarter than last.
The Case That This Is Already Over
The honest counter-argument is that the worst is behind us, and the deceleration is real. Contract increases went from 93-98% to 58-63% to 13-18% across three quarters. Gartner expects stabilisation around Q1 2027 and modest declines in the second half of 2027. Counterpoint Research projected in November 2025 that DRAM output would grow more than 20% in 2026. Korea has committed enormous capital to new fabs. If AI infrastructure spend slows before that capacity lands, today's shortage becomes tomorrow's glut — that cycle has run several times in this industry.
Two things break the comfort. First, decelerating increases still compound: 13-18% on a base that already tripled is a larger absolute dollar move than 90% was in early 2025. Second, the capacity is not close. Significant new manufacturing capacity is not expected online until late 2027 or 2028 — SK Hynix's Indiana facility is scheduled for mass production in late 2028 — and Micron is currently meeting only about half to two-thirds of core customer demand. Gartner's own numbers say DRAM revenue rises 246.6% in 2026 and NAND 371.9%.
So the deceleration is genuine and the relief is not, at least not inside a fiscal 2027 budget cycle. Plan for elevated pricing through your next refresh and treat a 2027 correction as upside, not as an assumption.
What to Do Before Your Next Hardware Quote Expires
This Week:
- Pull every open hardware quote and read the expiry date and the repricing clause. Quote validity has been the quiet variable all year — vendors have moved it in both directions. HPE went from 14 days to 30 in June and then, on August 6, to honouring quoted prices until the product ships for Compute, Storage and GreenLake Flex orders up to $1 million. Others have run as short as a week. Whether you have a price hold is a factual question with an answer in the document.
- Re-run your self-hosted inference and vector-database TCO with today's memory and NVMe line, not last year's. Report the crossover point against managed pricing as a number, not a conclusion.
- Ask your OEM to break memory and storage out as separately-quoted lines. If they are buried in a configured SKU price you cannot see the delta, and you cannot negotiate what you cannot see.
This Month:
- Rank your fleet by memory intensity, not by age. The 512GB virtualisation hosts and the vector-DB nodes are where the increase concentrates; a CPU-bound node barely moved. That ranking, not the refresh calendar, should set the order you buy in.
- Get a written answer on price protection through delivery. Counterpoint's Neil Shah put the split plainly: buyers who control the BOM should negotiate and lock in supply and cost in advance, while smaller buyers should spread the rollout to average out the spikes. Decide which of those you are before the conversation, not during it.
- Right-size the configuration. Distributors are already reporting customers downgrading from 96GB and 128GB modules to 64GB and 32GB on price. If a workload was specified with headroom nobody measured, that headroom now has a real price.
Before Renewal:
- Model the extend-versus-refresh decision explicitly, including the security cost. Gartner expects a 15% extension in hardware life across the market. Extending is defensible; extending by accident, onto hardware past vendor support, is not.
- Re-test the make-or-buy assumption on every AI service you decided to self-host in 2025. Some of those decisions were correct on GPU economics and are now wrong on host memory. Deciding again is cheap. Discovering it at the invoice is not.
- Write your pass-through expectation into the contract. Nvidia said it will recover margin through price increases starting in fiscal 2028. Your suppliers will do the same, and consumption-priced AI services are the easiest place to hide it.
The Bottom Line
The last time enterprise IT was repriced this hard by a component nobody owned the roadmap for, it was flash in 2017, and the organisations that came out ahead were the ones that treated a quote as a perishable good rather than a fact. This is the same shape, with more leverage concentrated in fewer hands: three suppliers, one demand curve, and a small set of buyers holding contracts that push the increase onto everyone else.
Nvidia has $279 billion of committed supply and it is still absorbing three to four points. It will get most of that back by raising prices — to its customers, which in the broader accelerator supply chain eventually means you. The company with maximum leverage is not eating this cost. It is holding it for a quarter and then handing it over.
The memory did not get more expensive last week. Your spreadsheet just got audited by somebody else's earnings call.
Continue Reading
- Self-Host the Vector DB for Residency. Not for the Bill.
- What RAG Actually Costs: $1,308 a Month at 10M Tokens/Day
- vLLM vs TensorRT-LLM vs SGLang: Default to vLLM
- Microsoft's GPUs Sit Unplugged. Buy Delivery, Not Capex.
- Why Samsung's $648B AI Bet Won't Help Your GPU Shortage Yet
- Vector Database Pricing: Only pgvector Publishes a Rate
- Groq Runs Nvidia Now. Recount Your Non-Nvidia Capacity.
