Groq Runs Nvidia Now. Recount Your Non-Nvidia Capacity.

Groq certified as an NVIDIA Cloud Partner on August 12 and closed a $350M down round on August 17 with Nvidia planning to participate. If your capacity plan lists Groq as the non-Nvidia leg, that row no longer exists — and the API was built so you would never notice.

By Rajesh Beri·August 18, 2026·10 min read
Share:
A single liquid-cooled inference rack on a data-centre floor, front panel open, with a blank brushed-metal nameplate freshly riveted over an older plate underneath, coolant hoses running up to the ceiling.

Illustration generated using AI

The non-Nvidia leg of your inference strategy just became an Nvidia channel. Groq — the vendor most enterprise architecture diagrams name as the alternative to buying Nvidia — certified as an NVIDIA Cloud Partner on August 12, and five days later closed a $350 million Series A with planned participation from NVIDIA itself.

If your capacity plan carries a row labelled "non-Nvidia inference" and the vendor in that row is Groq, the number is no longer approximately right. The row is gone. And because Groq's own promise to customers is that they "will get more capacity and the latest inference technology from Groq — without changing a line of code," nobody on your team will notice the day it happens.

What Groq Announced in August, in Order

Two announcements five days apart closed a loop that opened last December. Read together, they say Groq is now an operator of Nvidia hardware rather than a builder of an alternative to it.

On August 12, Groq announced NVIDIA Cloud Partner status — "certified to design, deploy, and operate NVIDIA accelerated computing to the NVIDIA's reference architecture and operational standards." CEO Adam Winter framed it as continuity: "Becoming an NVIDIA Cloud Partner is a natural next step in our collaboration with NVIDIA."

On August 17, Groq closed $350 million led by Disruptive, "with planned participation from NVIDIA." The release says Groq operates 13 data centers and "expects to scale from 54 megawatts to 200+ megawatts in 2027." It is explicit about what fills those megawatts: "The injection of capital will support those seeking usage of medium and larger sized clusters of NVIDIA accelerated computing for training and inference."

That is a hardware company describing its expansion in the vocabulary of a reseller. TechCrunch reported the round at a $3.5 billion valuation, down from $6.9 billion in September 2025 — a Groq spokesperson disputed the "down round" framing, calling it a new valuation for "the post-Nvidia-licensing-deal version of Groq." That phrasing is the most honest thing anyone has said about this transition. It is a different company. SiliconANGLE noted the raise landed less than three months after a $650 million round.

An NVIDIA Cloud Partner certification is not a logo. Nvidia's reference architecture for AI cloud providers specifies the stack end to end: GPU servers on Nvidia's latest architectures, "NVIDIA Quantum-2 InfiniBand and Spectrum-X Ethernet networking," "NVIDIA BlueField-3 DPUs," and storage from certified partners. Certifying means committing to compute, fabric, DPU and storage choices Nvidia sets. There is no partial version.


The LPU Is Now Nvidia's, and It Ships Beside a Rubin Rack

The technical argument for Groq as a hedge weakened before the corporate one did. The LPU is now an Nvidia product too, and Nvidia sells the current generation as a companion to its GPU racks rather than as a replacement for them.

In December 2025, Groq entered a non-exclusive licensing agreement with Nvidia for its inference technology. Founder Jonathan Ross and president Sunny Madra joined Nvidia; Simon Edwards stepped in as CEO. Constellation Research read the structure plainly: "Nvidia will get Groq's AI accelerator chip technology no matter what you call the deal." TechCrunch put the payout to investors at $20 billion.

At GTC in March 2026, Nvidia shipped it as its own product. The Decoder's account of the launch is the sentence to circulate internally: "With Groq 3 LPX, customers can now buy comparable hardware directly from Nvidia, letting the company leverage its platform advantage."

Then look at what the product actually is. Nvidia's engineering write-up on the NVIDIA Groq 3 LPX describes a rack "built around 256 interconnected NVIDIA Groq 3 LPU accelerators" — and nowhere describes it as a standalone system. "Deployed alongside Vera Rubin NVL72, LPX accelerates the latency-sensitive portions of the decode loop, including FFN and MoE expert execution, while Rubin GPUs continue to handle prefill and decode attention." That split is attention–FFN disaggregation: within the decode loop, attention runs on one engine and the feed-forward network on another, with intermediate activations exchanged for each token.

Say that in procurement terms. The LPU is now a co-processor to an Nvidia GPU rack. The Register's teardown explains why: each LPU "only has enough die space for 500 MB of on-chip memory," so serving a trillion-parameter model takes "between four and eight LPX racks, or 1,024 to 2,048 LPUs," and "one or more LPX racks is paired with a Vera-Rubin NVL72." Be precise about what that means, because the same teardown is explicit that it is "entirely possible to run large language models (LLMs) entirely on an LPX cluster, [but] that's not how Nvidia is positioning the product." The constraint is commercial, not physical — which for a buyer is the same constraint, since you can only deploy what the vendor sells and supports. The headline claim — Nvidia's "up to 35x higher inference throughput per megawatt" — is a claim about the combined Rubin-plus-LPX system, not about LPUs standing alone against GPUs.

A hedge the incumbent packages beside its own flagship rack is a weak hedge. It is closer to a line item inside the incumbent's platform than an exit from it.

Why Nobody on Your Team Will Notice

The substitution is invisible because Groq's API was deliberately built to hide the hardware, and that abstraction is the product.

Groq's OpenAI compatibility documentation tells you to point an existing OpenAI client at https://api.groq.com/openai/v1 and keep going. It enumerates precisely what differs — logprobs, logit_bias and top_logprobs are unsupported, n must be 1, a temperature of 0 is converted to 1e-8. It says nothing about what silicon answers the call. Neither does the production model list, which identifies each endpoint by model string and tokens per second. Models are named. Accelerators are not.

This is good API design and terrible governance input. Your latency SLO was benchmarked on a specific accelerator. Your cost model assumed a specific throughput-per-dollar curve. Your concentration register recorded a specific supplier of silicon. All three were derived from a hardware fact that your integration was explicitly built not to expose — and the only signal you would get from a swap is a change in the numbers you are already tolerant of.

We have watched this exact failure mode in the model layer. When DeepSeek swapped a model behind a stable endpoint, the evals that were supposed to catch it did not, because they were written against the endpoint rather than the artifact. The hardware layer is worse: there is no version string to pin.

The Strongest Case That Nothing Changed

Steel-man it properly, because the counter-argument is not weak. If you buy tokens per second at a price, you never bought silicon in the first place.

GroqCloud serves "more than six million developers and thousands of AI-native companies," per the August 12 release, and Groq says it remains "the only team in the world with hands-on experience operating LPUs in production at scale." Nvidia's licence was explicitly non-exclusive. The December announcement said GroqCloud "will continue to operate without interruption," and it has. Your p99 today is what it was in July. Nothing in either August announcement obliges Groq to retire a single deployed LPU.

There is also a defensible reading where this is better for you. A vendor that ships Nvidia reference architecture is a vendor with supply, a hardware roadmap it does not have to fund alone, and a lower chance of becoming the alternative-silicon company that quietly stops shipping. TechCrunch's neocloud caveat cuts the same way — "high capital expenditures, heavy reliance on debt, and exposure to rapidly depreciating hardware" is a description of Groq's new business, and Nvidia on the cap table is precisely what makes that business financeable.

Both things are true. Groq is probably a more durable supplier this week than it was in November, and it is no longer a diversification instrument. The mistake is letting the first fact quietly update a spreadsheet cell that was justified by the second. Two CEOs in eight months — Simon Edwards in December, Adam Winter in August — is a reasonable prompt to re-read what you actually signed.


Recount the Number, Then Fix the Contract

This Week:

  1. Reclassify Groq in your concentration register. Move whatever percentage sat under "alternative silicon" into your Nvidia exposure and look at the new total. That single number is the deliverable — take it to whoever owns supplier concentration, not to your platform team.
  2. Pull the actual contract language. Find whether you committed to LPUs, to a named accelerator, to a throughput floor, or to "inference." If the word is "inference," you have already consented to the swap. Note which it is before you talk to anyone.
  3. Grep your architecture decision records for "non-Nvidia," "silicon diversity" and "vendor-neutral." Every ADR that named Groq as the mitigation now has an unmitigated risk. Reopen them; do not silently edit them.

This Month:

  1. Re-baseline latency against production, not against a 2025 benchmark. Record p50, p95 and p99 per model per week and store the series. You cannot detect a hardware change you never measured before it happened. This is the same discipline that catches a model swapped behind a stable endpoint.
  2. Name a second inference supplier and send it real traffic. Not a proof of concept — a routed percentage. Be deliberate about which risk you are buying down: Cerebras and SambaNova build their own accelerators, while Together AI, Fireworks AI and Baseten serve open-weight models behind comparable APIs but run them on Nvidia GPUs — they diversify your supplier, not your silicon. A hedge you have never actually run under load is not a hedge either.
  3. Put a gateway in front of both. If your application code holds a base URL, your switching cost is an engineering project during an incident. A self-hosted gateway makes the second supplier a config change — and the same layer is where you enforce per-provider cost and residency rules.

Before Renewal:

  1. Ask, in writing, what accelerator serves your endpoints, and require notice before it changes. Groq will not put an LPU guarantee in a contract — no operator scaling to 200+ megawatts would. Ask anyway, and record the answer. The refusal is itself the finding, and it is the same "get it in writing" test that applies whenever a vendor promises hardware-agnostic inference.
  2. Convert any hardware-based pricing assumption into a throughput-and-latency SLO. If your unit economics rest on the price-performance of a specific chip, they rest on something your supplier is now free to change. Contract for the outcome instead.
  3. Check whether your commitment survives a change of control or a technology substitution. This is the same clause that matters when an inference vendor is acquired mid-term, and when a supplier asks for multi-year money against capacity it has not built yet.

The Bottom Line

Alternative silicon has not been a uniform story of failure — Cerebras raised $5.5 billion going public on Nasdaq in May, and SambaNova took $1 billion at an $11 billion valuation in July, both still building their own accelerators. Which is exactly why Groq's path matters rather than being written off as inevitable: it kept the brand, the customers and the API, sold the differentiator, and now buys the incumbent's reference architecture with the incumbent's money. From your API's point of view nothing happened, which is the entire problem. The comparison you were running on cost per token and the capacity you thought you were reserving were both premised on a supplier that has quietly changed what it sells.

Diversification you cannot observe is not diversification. Go count it again.

Continue Reading

Share:

Frequently Asked Questions

Is Groq still a non-Nvidia inference option?

No, not in the sense a concentration register means. Groq certified as an NVIDIA Cloud Partner on August 12, 2026, its Series A release describes scaling from 54MW to 200+MW of NVIDIA accelerated computing, and Nvidia planned to participate in the round. Groq is now an operator of Nvidia reference architecture rather than an alternative to it.

Will my Groq API calls break because of this?

No. That is the risk, not the reassurance. Groq's own announcement promises customers more capacity 'without changing a line of code,' and its OpenAI-compatible API identifies endpoints by model string, never by accelerator. A hardware substitution surfaces only as a drift in latency and throughput you may already tolerate.

Does the NVIDIA Groq 3 LPX still run without Nvidia GPUs?

Not as Nvidia sells it. The Register's teardown notes it is technically possible to run a model entirely on an LPX cluster, but that is not how Nvidia positions the product: its engineering write-up describes the 256-accelerator rack as deployed alongside a Vera Rubin NVL72, with LPUs handling FFN and MoE expert execution while Rubin GPUs handle prefill and decode attention. Each LPU carries roughly 500 MB of on-chip memory, so trillion-parameter serving needs four to eight LPX racks. For a buyer the distinction is academic — you deploy what the vendor supports.

What should I check in my Groq contract this week?

Whether you committed to LPUs, to a named accelerator, to a throughput floor, or simply to 'inference.' If the contracted word is inference, you have already consented to a hardware substitution. Also check whether change-of-control and technology-substitution clauses let you exit or reprice.

Which inference providers are still genuinely non-Nvidia?

Cerebras and SambaNova still design and build their own accelerators; Cerebras went public on Nasdaq in May 2026 and SambaNova raised $1 billion at an $11 billion valuation in July 2026. Together AI, Fireworks AI and Baseten serve open-weight models behind comparable APIs but run them on Nvidia GPUs, so they diversify your supplier rather than your silicon. Treat any of them as a hedge only once you have routed real production traffic through it under load — an untested second supplier is not a hedge.

Newsletter

Stay Ahead of the Curve

Weekly enterprise AI insights for technology leaders. No spam, no vendor pitches—unsubscribe anytime.

Subscribe

Latest Articles

View All →