Walmart and Amazon Both Said 40%. Neither Ran a Holdout.

Walmart and Amazon both told analysts that AI-assistant users spend 40% more per order than customers who don't use one. Neither disclosed a holdout, and the only figure in the whole earnings cycle a retail team can independently instrument is Target's external-AI referral channel.

By Rajesh Beri·August 20, 2026·14 min read
Share:
Two identical metal shopping carts standing side by side in a bright supermarket aisle — one heaped with groceries, the other holding a single item — with a long printed till receipt spooling from the full cart onto the

Illustration generated using AI

Every AI shopping number your board saw this month compares customers who chose the assistant against customers who didn't. That is a measurement of who uses the feature. It is not a measurement of what the feature caused.

On August 20, Walmart CEO John Furner told analysts that "the number of customers using Sparky is up 70% from last year, and the customers and members who use Sparky for shopping spend 40% more per order than others who do not," according to the Q2 FY27 earnings call transcript. Three weeks earlier, Amazon said U.S. customers who use Alexa for Shopping spend 40% more per order than shoppers who don't use the assistant. Same figure, same construction, two different companies, three weeks apart. Neither disclosed a control group. On the evidence available, neither has one.

If you are being asked this quarter to fund or expand a shopping-assistant program on a slide carrying these numbers, the question to ask is not whether the number is true. It almost certainly is. The question is what it is a number about.

The Same 40% Showed Up Twice in Three Weeks

Three large retailers reported AI-assistant metrics in the same earnings window, and every one of them is a comparison between adopters and non-adopters. PYMNTS collected them: Walmart's 40% per order with users up 70% year over year; Amazon's 40% per order with active users nearly doubling; Albertsons' 10% average order value on conversational search and 26% on its fuller assistants.

The Albertsons figures came from Jill Pavlovich, senior vice president of digital shopping experiences, in a Wall Street Journal report published August 17, and she volunteered the mechanism: "When they're using a more comprehensive experience, they're adding even more because they're not forgetting items." Hold that sentence. It matters later.

The construction is a house style in retail reporting, not a quirk of AI. On Albertsons' July 23 earnings call, CEO Susan Morris used the identical shape for the loyalty program — "engaged members shop more frequently and with higher average basket than nonmembers," per the call transcript. Nobody thinks the loyalty card causes people to want groceries. It sorts the people who already do.

There is one more tell worth noting. Walmart's Q2 FY27 earnings release leads with global eCommerce up 23% and U.S. comparable sales up 2.6%. Sparky is not in it. The 40% exists only in spoken remarks — a category of disclosure that carries no reconciliation requirement and no footnote. CFO John David Rainey told the same call that the advertising business "continues to grow at a 40% clip," per the transcript, while that call put global advertising at 38% and Walmart Connect at 43%. Round numbers travel; audited ones don't.

What "Spend 40% More Per Order" Cannot Tell You

A cross-sectional comparison between users and non-users of an opt-in feature cannot separate the effect of the feature from the characteristics of the people who opted in. That is the entire problem, and no amount of sample size fixes it.

Think about who reaches for a shopping assistant. Someone planning a week of meals, not someone buying batteries. Someone stocking a dorm room in August. Someone with a large, complex, high-intent basket already forming in their head — the exact shopper whose order would have been bigger anyway. The assistant is a filter on intent before it is a lever on intent.

There is a second, quieter problem, and Pavlovich named it without meaning to. If the assistant stops customers forgetting items, it consolidates two trips into one. Average order value rises. Spend per customer per month may not move at all. Average order value is a per-order metric; a program's business case is a per-customer, per-period metric. Those two things come apart precisely when a tool improves list completeness, which is the main thing these assistants actually do well.

Then there is the drift. In May, Walmart reported Sparky users had an average order value about 35% higher than non-users, with weekly active users up over 100% in a single quarter. By August the lift was 40% and the user metric had changed shape, to total customers up 70% year over year. Under a pure selection story you would expect the opposite: as a feature recruits past its enthusiast core, the gap between users and non-users should compress toward zero. It widened. That is either genuine product improvement — Walmart also said in May that AI investment had raised Sparky's "intelligence and response quality by 40%" — or it is a ratio computed on a moving denominator. From outside, the two are indistinguishable, and that is the point.


Walmart Already Published the Number It Could Measure

The most instructive AI-commerce number Walmart has disclosed this year is negative, and it came from the one experience where the company could observe both arms of the funnel. In March, EVP of product and design Daniel Danker told WIRED that in-chat purchases through OpenAI's Instant Checkout converted at one-third the rate of click-out transactions — roughly 200,000 products, live since November 2025. He called the experience "unsatisfying." Walmart moved off it, and is embedding Sparky inside ChatGPT instead, with account linking, loyalty and its own payment options.

Same company. Same product surface. Five months apart. When the comparison was structurally clean — two checkout paths for the same shopper population — the number was bad and it got published. When the comparison is between self-selected adopters and everyone else, the number is 40% and it leads the call.

That contrast is not an accusation of bad faith. It is a description of what each measurement setup is capable of producing. OpenAI drew the same conclusion from its side, telling merchants that "the initial version of Instant Checkout did not offer the level of flexibility that we aspire to provide" and refocusing on product discovery — the shift we covered when ChatGPT first became a sales channel for retailers and again when Google's rival Universal Commerce Protocol picked up Microsoft, Amazon and Meta.

Target Reported a Channel, Not a Comparison

Target disclosed the only figure in this earnings cycle that a retail team can independently instrument: traffic arriving from outside AI platforms. On the August 19 call, CEO Michael Fiddelke said that "while still small in total today, as more consumers begin to explore the benefits of agentic shopping, Target's digital traffic sourced from external AI platforms is growing" more than three and a half times the industry rate, per the call transcript. Target's own release repeats it: digital traffic from external AI platforms like OpenAI and Google growing more than 3.5x the industry, against comparable sales up 3.8% and digital comps up 8.7%.

Two things separate this from the 40%. First, "small in total today" is an honest denominator disclosure — the only one anybody offered. Second, referral traffic is a count, not a comparison. A visit arriving with a ChatGPT or Gemini referrer is a row in your own logs. You do not need a control group to count rows.

The size of that channel is knowable from outside, too. Adobe Analytics, working from more than a trillion visits to U.S. retail sites, found AI-referred traffic up 138% year over year in May 2026, with those visitors spending 53% more time on site and viewing 23% more pages per visit. Growth is real and fast. The base is still tiny: Etsy told the same earnings cycle that its agentic traffic grew roughly 15x year over year and remains under 1% of total traffic.

Be careful with the rest of that dataset. Adobe also reports AI-referred traffic converting 54% better than non-AI traffic, and that comparison has exactly the same shape as Walmart's — people who arrive after asking an assistant for a product recommendation are further down the funnel than people who arrive from a banner. The traffic count is instrumentable. The conversion delta is not causal evidence, and no vendor deck will tell you which half you are reading.

Target's incoming chief AI officer, Chandhu Nair, starts August 24 and inherits this measurement problem in his first week.

The Case That the Lift Is Real

The strongest version of the other side deserves stating plainly, because it is not weak. These assistants do things that mechanically enlarge a basket: they assemble a meal plan and add every ingredient in one click, they surface complements a search box never would, and — per Walmart's own example — they check what you already bought so you don't duplicate it. A tool that turns "chicken" into a twelve-item recipe basket does not need selection bias to raise order value.

Adoption curves support it too. Amazon said active Alexa for Shopping users nearly doubled year over year with interactions up more than 5x, and that over 350 million customers used it in the last 12 months. At that scale you are well past the enthusiast cohort, and the population using the assistant starts to look like the population, which shrinks the selection gap. Note that the two figures describe different groups, though: the 350 million counts anyone who used it once in a year, while the 40% is computed on U.S. customers who use it to shop. A broad adoption number tells you nothing about how self-selected the narrow one is.

Both of those are good arguments that the true effect is greater than zero. Neither is an argument that it is 40%. And where the experiment has actually been run, the answer comes back far smaller — and arrives through a different mechanism than the basket story predicts.

Two published randomized field experiments have tested an on-site AI shopping assistant against a control arm. On a livestream selling platform, an AI streaming assistant raised sales 3.00% and cut product return rates 12.55%, in Information Systems Research. Larger and closer to the case at hand: across seven GenAI workflows at a cross-border retail platform, randomizing millions of consumers, a pre-sale service chatbot produced a 16.3% sales lift and a 21.7% conversion lift — the biggest of the seven, with the others running from 2.9% down to no detectable effect at all.

That second study lands squarely on the basket argument. Its authors report that sales gains "are accompanied by higher conversion rates ... while average cart values remain largely unchanged," and find "no evidence of effects along the intensive margin." The randomized evidence says these tools get more people to buy. It does not say they make the order bigger — and order size is the only thing "spend 40% more per order" measures.

Neither study is Sparky and neither is 2026, so treat them as calibration rather than verdict. What nobody has published is a retailer testing its own assistant. An AI shopping-agent vendor, Alhena, stated flatly in July that "no brand and no vendor — Alhena included — has published a completed holdout or randomized test reporting the incremental revenue lift of an on-site AI shopping agent," and that practitioners who do run them "consistently report that true incremental lift lands well below the headline attributed figure — often in the single digits once self-selection is stripped out." That is a vendor's characterization rather than a study — but 3.00%, from a peer-reviewed experiment, is exactly the order of magnitude it describes.

Retail Ran This Experiment in 2014 and Lost

Digital commerce has already been through one full cycle of confusing correlation with causation at enormous scale, and the correction was brutal. Tom Blake, Chris Nosko and Steven Tadelis ran large-scale field experiments at eBay and found that returns from paid search were a fraction of conventional non-experimental estimates — and that in the extreme case, brand-keyword ads had "no measurable short-term benefits." When eBay switched them off, most of the traffic showed up anyway. The clicks had been real. The incrementality had not.

The follow-up work is worse for anyone hoping better data solves it. Researchers at Northwestern and Facebook compared observational estimates against randomized experiments across 15 U.S. advertising experiments, 500 million user-experiment observations and 1.6 billion ad impressions. Florian Zettelmeyer's summary: "Generally, the current and more common methods overestimate ad effectiveness relative to what we found in our randomized tests."

Then they scaled it up. Across 663 large-scale Facebook experiments with access to over 5,000 user-level features, median true experimental lifts were 29%, 18% and 5% for upper-, mid- and lower-funnel outcomes. Double/debiased machine learning — the good method — returned 83%, 58% and 24%. Propensity-score matching returned 173%, 176% and 64%. The authors' conclusion: "despite having access to large-scale experiments and rich user-level data, we are unable to reliably estimate an ad campaign's causal effect."

Those platforms had every behavioral signal about their users that exists. The bias survived. A quarterly slide comparing Sparky users to non-Sparky users is a far cruder instrument than double/debiased ML on 5,000 features, and it is being read with far more confidence.

What to Instrument Before Peak Season

You cannot randomize who chooses to use an assistant. You can randomize who is offered one, and that is enough — it is the standard fix, and it is cheap.

This Week:

  1. Get the denominator in writing. Ask your digital team for the exact definition behind your own version of the 40%: is a "user" any session that opened the assistant, or only an order attributed to it? Ask for the same cut expressed as spend per customer per 30 days, not average order value per order. If the second number is much smaller than the first, you are looking at basket consolidation, not growth.
  2. Count the AI referral channel. Pull sessions and revenue by referrer for chatgpt.com, gemini.google.com, perplexity.ai and copilot.microsoft.com against total sessions, for the last four quarters. This is a count from your own logs and needs no vendor's cooperation. Target's disclosure says it is growing fast and is still small; verify both halves for yourself.

This Month:

  1. Randomize the surface, not the usage. Hold the assistant's entry point back from a random 5–10% of app traffic for four weeks and measure intent-to-treat: revenue per assigned user, not revenue per adopter. This is the single experiment that separates the two stories, and it costs you a rounding error of exposure.
  2. Split click-out from in-chat. Report external-platform sessions that land on your site separately from transactions completed inside an assistant. Walmart's own 3x conversion gap says these are not the same funnel and must not share a line on the dashboard.

Before Peak Season:

  1. Put the measurement in the contract. If a vendor's business case cites a lift number, require a holdout-capable configuration as a deliverable — a documented way to withhold the feature from a random cohort and export the arms. A vendor who cannot supply that is selling you an attribution report, and you should price it as one.
  2. Fix what the number is used for. An adoption metric is fine for a product roadmap and useless for a capital request. Decide now, in writing, which of your AI numbers are activity and which are outcome — the same distinction that separates an auditable earnings-call AI metric from a decorative one, and the reason Meta killed its own AI leaderboard inside 48 hours.

The Bottom Line

Enterprise AI keeps rediscovering that throughput is easier to measure than value. Agent teams ship 65 pull requests a week and nobody gets time back. Two-thirds of the Fortune 500 run AI coding agents and a third of them measure the return. Walmart, Uber and Microsoft spent a year discovering what their AI actually cost before they could say what it earned. Retail's shopping assistants are the same story with a friendlier number attached, and the friendliness is the danger: a figure this clean invites no scrutiny, which is exactly how a confident explanation earns trust it has not earned.

The assistants are probably working. Adobe's traffic curve is real, Amazon's 350 million users are real, and a tool that assembles a twelve-item basket from one request does something a search box cannot. But "probably working" and "40% more per order" are separated by an experiment no retailer has run on its own assistant — and where researchers have run it elsewhere, the lift came in between 3% and 16%, through conversion rather than basket size. It takes four weeks and a randomized 5% holdout to find out which one you are funding.

A number without a control group describes your customers. Your board is reading it as a description of your software.

Continue Reading

OpenAI Turns ChatGPT Into a Sales Channel for Retailers

Google UCP Beats OpenAI Protocol: Microsoft, Amazon, Meta Adopt

Airbnb Shipped 80% More Features. One Number Is Auditable.

Why Meta Killed Its AI Leaderboard in 48 Hours

Agent Teams Hit 65 PRs a Week. Nobody Got Time Back.

Enterprise AI Costs Explode: Uber, Walmart, Microsoft Rein In

The Vaguer the AI Explanation, the More Novices Trusted It

Share:

Frequently Asked Questions

Why isn't "Sparky users spend 40% more per order" proof that the assistant works?

Because it compares customers who chose the assistant against customers who didn't. Shoppers who open an AI assistant are already planning larger, more complex baskets — meal plans, back-to-school lists — so their orders would likely have been bigger regardless. Without a randomized holdout, the figure measures who adopts the feature, not what the feature causes.

How do you run a holdout test on an opt-in AI shopping assistant?

You randomize access rather than usage. Withhold the assistant's entry point from a random 5-10% of app or site traffic for about four weeks, then compare revenue per assigned user across both arms — an intent-to-treat measurement. That removes self-selection because assignment, not the shopper's choice, determines the group.

What AI shopping metric can a retailer actually verify on its own?

Referral traffic from external AI platforms. Sessions arriving with a chatgpt.com, gemini.google.com, perplexity.ai or copilot.microsoft.com referrer are rows in your own web logs — a count, not a comparison, so no control group is needed. Target disclosed this channel is growing more than 3.5x the industry rate while still being small in total today.

Does average order value overstate the value of an AI shopping assistant?

It can. If the assistant stops customers forgetting items, it consolidates two trips into one: average order value rises while spend per customer per month stays flat. Albertsons' own explanation for its 26% lift was that shoppers 'are not forgetting items.' Always request the per-customer, per-period figure alongside the per-order one.

Newsletter

Stay Ahead of the Curve

Weekly enterprise AI insights for technology leaders. No spam, no vendor pitches—unsubscribe anytime.

Subscribe