The press release called it a collaboration. The 8-K called it a warrant. On September 8, 2026, Qualcomm announced a multi-generation product collaboration with Amazon to build customized AI inference silicon and optical connectivity for AWS data centers. That release contains no dollar figure, no name for the silicon and no timeline. The same day, Qualcomm filed an Item 3.02 8-K, event-dated September 3, disclosing the part that has a price: a warrant issued to Amazon.com NV Investment Holdings LLC on 25,000,000 Qualcomm shares at $161.26, expiring September 3, 2036, of which 3,750,000 vested on issuance and the rest vest as Amazon buys — up to $60 billion in payments.
So the practical question is not whether Qualcomm silicon is good. Nobody outside the two companies knows yet — but it is further along than the press release implies. Speaking at a Goldman Sachs conference the same day, Qualcomm CFO Akash Palkhiwala said the company is "already in production" with Amazon and books revenue from the engagement starting in the December quarter. What neither company has disclosed is the thing you would actually need: when an AWS customer can rent the result. The question is what you sign in the meantime. Your cloud provider now holds a disclosed, decade-long financial position that pays off in proportion to how much of one supplier's silicon it purchases. That is legal, transparent, and increasingly ordinary. It is also a reason to make sure that the multi-year AWS compute commitment sitting in your procurement queue this quarter is not denominated in an instance family.
What the Press Release Left Out
Every number that matters is in the filing, not the announcement. The 8-K states that the warrant shares "vest in tranches tied to the execution of certain commercial arrangements, the placement of binding purchase orders and actual purchases of QTI's server chip products, technology, systems and manufacturing services by Amazon during the term of the Warrant, up to a maximum amount of $60 billion in payments, with 3,750,000 shares being vested upon issuance of the Warrant based on initial purchase commitments."
Read that sentence twice. Three triggers are named — commercial arrangements, binding purchase orders, actual purchases — and none of them is a performance benchmark, a utilization threshold, or a customer price. The instrument is cashless-exercisable, carries no voting rights before exercise, comes with registration rights and a resale prospectus supplement Qualcomm expects to file, and was issued under the Section 4(a)(2) private-placement exemption.
Only 15% of it is vested. The other 21.25 million shares are a ten-year forward contract on AWS's own purchasing behaviour.
The announcement, by contrast, gives you SerDes, optical DSP, "up to 1.6T," and two quotes. AWS Vice President Prasad Kalyanaraman said the companies are "delivering more performant, efficient, and cost-effective infrastructure for our customers." As Unite.AI noted, the announcement omitted financial terms, deployment timelines and volume commitments entirely. Markets read the filing anyway: Qualcomm shares rallied as much as 9.5% on the day, settling around 3.9% up in afternoon trading.
Amazon Is Getting Paid to Be a Customer
This is the third variant of a structure that became standard in eleven months, and it is the most consequential one for buyers. In the first variant, the supplier invests in the customer. In the second, the customer gets equity for consuming. In this one, the customer gets equity for consuming and then resells the result to you.
Compare the terms directly, because they are not the same instrument:
| AMD → OpenAI | Qualcomm → Amazon | |
|---|---|---|
| Shares | 160,000,000 | 25,000,000 |
| Exercise price | $0.01 | $161.26 |
| Expires | October 5, 2030 | September 3, 2036 |
| Vests on | GPU purchases, 1 GW to 6 GW | Purchase orders + purchases, up to $60B |
| Share-price condition | Escalating to $600 | None disclosed |
AMD's 8-K and its October 2025 announcement describe a warrant at a penny a share whose final tranche requires AMD stock to reach $600. That is a very generous grant with a very hard hurdle. Qualcomm's is the inverse: Amazon must pay $161.26 a share — real money, not a token — but the disclosed vesting conditions are purchase milestones, with no share-price target named in the filing. Amazon's option is worth less per share and easier to earn, and it runs five years longer.
And Amazon is not OpenAI. OpenAI consumes its GPUs internally. Amazon rents its accelerators to you, by the hour, from a console that recommends instance types. That is the distinction worth carrying into a negotiation.
Steel-manning the other side properly: Amazon already pays for silicon without getting equity, and other people pay Amazon without getting any. Anthropic committed more than $100 billion over ten years to AWS in April 2026 for up to 5 gigawatts of Trainium2, Trainium3 and Trainium4 capacity. Nobody handed Anthropic a warrant for that. A supplier paying a hyperscaler to take a risk on unproven silicon is how you get a third merchant inference vendor at scale — which is exactly what the cost-per-token comparison across NVIDIA, AMD, TPU and Trainium says the market still lacks.
Vesting Tracks Purchase Orders, Not Your Cost per Token
The alignment created by this warrant is between Amazon and Qualcomm's order book. It is not between Amazon and your bill.
Re-read the trigger language: "the placement of binding purchase orders and actual purchases." Amazon's tranches vest when Amazon buys. They do not vest when the fleet is well utilized, when inference latency improves, or when a customer's cost per million tokens falls. Those outcomes may well follow — but they are not what the instrument pays for, and a contract's incentives live in its triggers, not in its press release.
This matters because AWS is the party that tells you which instance to run. When a solutions architect recommends an inference family on cost-efficiency grounds in 2028, that recommendation will be sound or unsound on its technical merits, and it will also sit inside a company holding an unvested position that grows with Qualcomm purchase volume. You do not need to assume bad faith to want a benchmark clause. You need only notice that the disclosure exists and that it runs to 2036.
The honest framing for a vendor-risk register: this is a disclosed conflict, not a hidden one. Qualcomm filed it. Treat it the way you already treat a reseller that also bills you — the same problem that made cloud-cost tooling ownership worth checking when the layer that measures your spend gets bought by a layer that generates it.
$60 Billion Is Attached to a Business That Starts in December
The ceiling in the warrant is not a forecast. It is a cap, and the gap between it and Qualcomm's actual data center business is the clearest measure of how long this arrangement is designed to run.
Qualcomm's Q3 fiscal 2026 results show total revenues of $9.9 billion and QCT revenues of $8,504 million, broken into three streams: handsets at $5,086 million (down 20%), automotive at $1,588 million (up 61%) and IoT at $1,830 million (up 9%). There is no data center line item at all. On the same earnings call, management said data center revenue from its custom silicon engagements begins in the December quarter, with wafers already in production, and that it will be dilutive to QCT gross margins by 1.5% to 2%.
So the maximum trigger is $60 billion in cumulative payments against a business line that has not yet reported a dollar and starts below the segment margin. For scale, Bank of America's estimate cited by TNW puts the entire server processor market at $27 billion in 2025, growing to $60 billion by 2030. The warrant's ceiling is roughly the size of the whole market at the midpoint of the warrant's life. That is why it expires in 2036 and not 2029.
The "$4 billion" figure in most coverage deserves the same scrutiny. Twenty-five million shares at $161.26 is $4.03 billion — that is the cost of exercising the warrant, not its value. Struck at market rather than at a penny, the warrant is worth nothing unless Qualcomm's stock rises, which requires the silicon to work and other buyers to want it. That is a real alignment, and it is the strongest argument in this deal's favour.
The Silicon Is in Production. The Instance Type Is Not Announced.
The migration decision is not available to you in 2026 — but not because nothing is being built. It is because AWS has not said when you can buy it.
Do not date the AWS parts off Qualcomm's public roadmap; they are on a different track. The roadmap published June 24, 2026 commits to "a multi-generation data center roadmap with an annual cadence" and gives dates for the merchant parts: commercial sampling of HBC Gen 1 with the AI250 in mid-2027; the AI300 with HBC Gen 2 in 2028; the Dragonfly C1000 CPU at commercial availability in 2028. Qualcomm projects four to eight times better memory bandwidth per watt per card for the AI300 against existing GPU-based architectures — a vendor claim, on a metric of the vendor's choosing, unbenchmarked by anyone independent, about a part that does not exist yet.
The Amazon silicon is custom, and it is ahead of that. On the July earnings call Qualcomm said it already held purchase orders and had started wafer production on two hyperscaler engagements, with revenue beginning in the December quarter; on September 8 the CFO confirmed Amazon is one of them. So the honest statement of the timeline is narrower than "nothing ships": Qualcomm is shipping. AWS has announced no instance type, no region, no preview and no price. Nor is Qualcomm silicon on EC2 novel — AWS has offered DL2q instances built on the Qualcomm Cloud AI 100 since November 2023, in one size and two regions.
That gap is the whole problem. A part in production with no announced instance type can become bookable faster than a roadmap date implies, and you will not get advance notice of the console entry. Meanwhile AWS's own hybrid silicon strategy already spans NVIDIA and in-house Trainium, and Qualcomm has been building toward this since it bought Modular for its CUDA-portability layer in June.
If you are signing a one-year commitment, this is a low-stakes bet either way. If you are signing three, you are signing across the window in which custom Qualcomm silicon reaches AWS racks and the merchant AI250 and AI300 parts arrive behind it — and the terms you accept now are the terms you live with while the accelerator mix underneath you changes.
Price the Optionality Yourself — AWS Will Not Quote It
AWS already sells you the choice explicitly, and it is cheaper than repapering later. Most teams have never priced it because nobody made them.
Per AWS's own Savings Plans pricing page, there are two shapes on a 1- or 3-year term:
- Compute Savings Plans — savings up to 66%. They "automatically apply to EC2 instance usage regardless of instance family, size, AZ, Region, OS or tenancy," and also cover Fargate and Lambda.
- EC2 Instance Savings Plans — savings up to 72%. They commit you to "individual instance families in a Region (e.g. M5 usage in N. Virginia)."
Six percentage points — with a caveat that matters more than the number. Both figures are "up to" ceilings, and AWS publishes neither the instance families that reach them nor a rate card behind them. On general-purpose compute the gap is real and roughly this size; cost-optimisation vendors independently put the flexibility premium at five to seven points. On accelerated instances it is neither six points nor 66% to begin with. Published third-party rates for a P5 sit near 57% off on-demand for a three-year commitment, well under the 72% headline, and the spread between the flexible and family-locked shapes runs wider on GPU families than on an M5. The direction of the argument survives; the specific number does not travel from a marketing page to your accelerator bill.
So price it on your own rates, not on AWS's ceilings — and price it before you decide, because on accelerator families the premium is large enough to be worth arguing about rather than waving through. That cuts both ways, and it is the point: locking the family is a defensible trade for a stable, well-understood workload — a database fleet, a batch pipeline — and a bad one for inference capacity specifically, in a three-year window during which the provider has a disclosed financial interest in changing what it recommends.
One exception worth knowing before you build the spreadsheet: Savings Plans do not apply to Capacity Blocks at all. AWS's own documentation states that "Savings Plans and Reserved Instance discounts don't apply to Capacity Blocks" — so if you reserve GPU capacity that way, this lever does not exist and the term of the block is your exposure. Reserved Instances carry the same family-versus-convertible choice as Savings Plans; the full commitment-structure comparison for GPU hours covers where each shape actually pays. And it is the same lesson as capping the term on an inference placement commitment when the vendor has set a date but not a price.
What to Do
This Week:
- Pull every active AWS Savings Plan and Reserved Instance and sort them into two buckets: family-locked and family-flexible. Put the dollar value and the end date next to each. Most teams cannot produce this list in under an hour, which is the finding.
- Identify which of those cover inference or accelerator workloads specifically, versus general compute. The argument in this piece applies only to the first bucket. Do not let it become a reason to repaper your whole estate.
- Add one line to the vendor-risk register for AWS: supplier equity position disclosed, Qualcomm 8-K filed 2026-09-08, warrant expiry 2036-09-03, vesting tied to purchase volume. Link the filing. This is a fact, not an accusation, and it belongs on the record with a citation.
This Month:
- Price the delta yourself, off quoted rates rather than the 66%/72% headlines — those ceilings are not what accelerator families actually clear. Pull the Compute and EC2 Instance Savings Plan rates for the specific instance types your inference runs on, and put the annual difference in front of whoever signs. Then decide deliberately. A premium in dollars is a number a CFO can approve or reject; "flexibility" is not.
- Ask your AWS account team, in writing, one question: which accelerator families will carry inference workloads in this region through the commitment term, and what happens to the discount if we move between them? File the answer.
- Establish a benchmark baseline on your current instance family now — tokens per second, cost per million tokens, p95 latency, at your real batch sizes. You cannot exercise a benchmark-before-commit right against a new family if you have no current number to compare it against, and the announcement will not come with advance warning. If you run on Amazon Bedrock, capture the provisioned-throughput equivalents too.
Before You Sign Anything Multi-Year:
- Insist the commitment be denominated in dollars of compute, not in an instance family, an accelerator generation, or a product name. This is the whole play. It is usually available, its cost is a single-digit discount give-up on general compute and a larger one on GPU families, and it survives every roadmap change either party makes.
- If the vendor insists on family denomination, ask for a migration right: the ability to reallocate committed spend to a successor family at the same discount, without a term reset. Get it in the agreement, not in an email.
- Do not try to time the term around a launch date, because AWS has not published one and the silicon is already in production. Either keep the term short enough to re-decide (one year), or make the flexibility contractual so the launch date stops mattering. A three-year family-locked commitment signed blind into an unannounced transition gives you the worst renewal leverage of any option available.
The Bottom Line
The last industry cycle that looked like this was carrier handset subsidies: the party recommending the device was compensated by the party making it, disclosure was technically complete, and the economics were invisible to the buyer at the point of sale for the better part of a decade. It resolved when buyers started pricing the phone and the plan separately. Nobody legislated it; people just learned to ask for the two numbers.
This is the same request, made early. Qualcomm's own filing tells you Amazon's upside is measured in purchase orders through 2036. Your upside is measured in cost per million tokens next quarter. Those are different numbers, and the repricing lesson from the last inference-efficiency scare applies exactly: do not fix your rate to something the vendor can change underneath you.
The warrant is disclosed. The launch date is not. The only thing still undecided is whether your next commitment names a family.
Buy dollars of compute. Let them pick the silicon.
Continue Reading
- NVIDIA Alternatives for Inference: Only Trainium Pays Off
- GPU Hour Pricing: A $3.44 Rate Cost $68.80 to Use
- AWS Orders 1 Million Nvidia GPUs—Then Bets Half on Custom Chips
- Groq Runs Nvidia Now. Recount Your Non-Nvidia Capacity.
- AMD Bought Taalas. Now Name the Model You'd Freeze.
- Microsoft's GPUs Sit Unplugged. Buy Delivery, Not Capex.
- d-Matrix Bought Wallaroo. Get 'Any Hardware' in Writing.
