For the first time, enterprise AI has receipts. Not predictions. Not vendor case studies. Named companies, real deployments, documented outcomes — including the six that got pulled back. Those six might be worth more than the other 153.
AI Weekly published the AI Use-Case Library last week: 159 real, named AI deployments across 21 industries. Reported outcomes on 77 of them. Six entries marked as halted or reversed.
I've spent the past 48 hours going through it. Here's what enterprise leaders need to know — and what the failures teach better than any success story.
Why This Data Set Is Different
Most AI case studies are written by vendors. They're polished, selective, and skewed toward the wins. The AI Weekly library is different: every entry links to its original source, whether that's Bloomberg, TechCrunch, a government press release, or a SEC filing.
Six entries are marked halted or reversed. A vendor would never publish that list. A neutral observer would.
For a CIO walking into a budget meeting, the question isn't "can AI do this." The question is: has anyone in my industry actually done this, with what tools, and what happened? That's what this library answers.
What's Actually Working
The wins share a pattern: narrow scope, measurable outcome, no judgment calls required from the AI.
A few entries from the library that stand out:
OpenAI went all-in on its own agents. [97.9% of OpenAI employees now use its Codex agent](https://www.theregister.com/ai-and-ml/2026/06/25/openai-says-employees-moving-beyond-chat-to-agents/5262499) — up from roughly 40% last August. Legal and Recruiting are on it. When an AI company converts its own internal operations at that rate, it's a signal worth noting.
The NHS is reading chest X-rays at national scale. A £20M expansion has helped over 4 million patients get faster lung-cancer diagnoses. Analysis on complex cases dropped from eight days to four. The task is narrow (read this scan, flag anomalies), the outcome is measurable (diagnosis time), and a human still makes the final call.
Pinterest's AI ad engine hit its first $1 billion quarter. Its AI-powered Performance+ campaigns now drive about 30% of lower-funnel revenue. Advertisers who adopted them grew that spend at nearly twice the rate of non-adopters. This is the cleanest ROI story in the library: AI managing bidding and targeting at scale, with revenue as the direct measurement.
A robot is making 500 bowls an hour. Wonder put the bowl-making system it bought from Sweetgreen into its first kitchen. A human line cook makes about 45 bowls per hour. 11x throughput improvement. The task: make the same bowl repeatedly at speed. Zero judgment required.
Momenta's ADAS shipped in nearly 900,000 cars. Its urban Navigate-on-Autopilot software is now in production vehicles from Toyota, Mercedes-Benz, BYD, GM, and Audi. This is AI at production scale in a safety-critical industry. It works because the input space (road, lanes, traffic signals) is structured, and the output (steer, brake, accelerate) is directly measurable.
The pattern: well-defined input, well-defined output, no institutional ambiguity in the middle.
The 6 That Got Pulled Back
These are the most expensive lessons in the library. They show where AI confidently fails — and where enterprises are still learning the limits.
Ford rehired 350 quality inspectors. The company had deployed automated quality inspection systems in manufacturing. Those systems produced defects that only experienced human inspectors reliably caught. The edge cases — the ones that look "close enough" to a model but aren't — accumulated. At a carmaker, defect accumulation isn't an acceptable error rate.
This one should make every operations leader pause. The mistake wasn't deploying AI for quality inspection. The mistake was removing the human checkpoint before validating edge-case coverage at production scale.
Waymo paused robotaxis in four cities. A robotaxi drove into an Atlanta flood and sat stuck for about an hour. Waymo then paused service in Atlanta, San Antonio, Dallas, and Houston — the last two as a precaution. Weather edge cases surfaced a limitation that wasn't apparent under normal operating conditions.
Meta dropped automated hate-speech moderation. After switching to a Community Notes model, abusive and racist posts targeting US legislators tripled within six months. The task — identify policy violations in natural language — requires judgment about context, intent, and evolving norms. Models handle classification; they struggle with evolving norms.
A federal judge threw out DOGE's AI grant cuts. DOGE used ChatGPT to flag approximately $100M in National Endowment for the Humanities grants as "DEI-related." The model labeled a Holocaust literature anthology that way. A 143-page ruling found unconstitutional viewpoint discrimination. This is what happens when AI output is treated as a final decision rather than a flagging mechanism.
Wake County schools banned AI detection tools. A student given a zero based on an AI detector appealed. A second teacher found no AI use and changed the grade to 100. The false positive rate was high enough to end the program. AI detectors are probabilistic tools being used as enforcement mechanisms — a mismatch in function.
The common thread: these deployments failed where AI was trusted to make categorical judgments in ambiguous situations without a human backstop.
Ford's defects. Meta's hate speech. DOGE's grants. Each required the kind of nuanced judgment that a model will get confidently wrong in a meaningful percentage of cases.
The Audit Era Is Here
Here's the broader context for why a library like this matters right now.
Amazon CTO Werner Vogels told Fortune that enterprises are shifting to cheaper open-source models because the bills got real. He cited Uber burning through its entire 2026 AI budget in four months. Another company ran through half a billion dollars in a single month before capping employee usage.
404 Media documented Amazon, Adobe, Atlassian, and Citi throttling employee AI use. One firm's monthly spend tripled past $15 million before controls went in.
PwC's 2026 Global CEO Survey found that 56% of global CEOs cannot point to measurable business impact from their AI spend. Gartner and CFO Dive have flagged the same pattern: companies are acquiring AI licenses the way they once bought software — volume purchase, figure out ROI later.
MIT Sloan found that 73% of failed AI projects had no agreed definition of success before they started. Another finding: 61% were approved with a projected ROI that was never actually measured after launch.
The AI market has moved from "can it do this" to "what did it cost, and did it work." Procurement teams that once accepted a slide deck as justification now want precedent. That's why a library of 159 named deployments — including the six that failed — is more useful than any vendor pitch.
What Enterprise Leaders Should Take From This
Before the next AI budget meeting, apply three filters from this data:
1. Is the task narrow and the outcome measurable?
The wins in this library share a common structure: defined input, defined output, direct measurement. The NHS reads X-rays and measures time to diagnosis. Pinterest runs ad campaigns and measures revenue. OpenAI's Codex agent handles code tasks and measures adoption and output quality.
If you can't state the measurement in one sentence, the project isn't ready to scale.
2. Where does a human need to remain in the loop?
Every reversal in this library involved removing the human checkpoint from a judgment-heavy decision. Ford's inspectors. DOGE's grant reviewers. Meta's content moderators.
The rule: AI as flagging mechanism, human as decision-maker. Not the other way around. The moment AI output becomes the final decision in an ambiguous context, you're in reversal territory.
3. Have you defined success before you commit budget?
MIT Sloan's 73% figure is not abstract. In conversations with enterprise technology and operations leaders, I hear a version of this regularly: an AI project gets funded because someone in the C-suite saw a demo, the business case was built backward from the purchase decision, and twelve months later, no one can point to what changed.
Success definition precedes procurement. Always.
For CFOs and Finance Leaders
The financial picture in 2026 is sharper than it was twelve months ago.
Companies plan to spend an average of 1.7% of revenue on AI this year. For a $5B enterprise, that's $85M. The data from this library suggests that spend is unevenly productive: high returns where the task is structured and measurable, visible reversals where it isn't.
The audit framework that's emerging across enterprises focuses on three questions before any AI investment is approved: What is the specific measurable outcome? What is the cost per outcome? What is the rollback plan if edge cases surface at scale?
Ford didn't have a credible answer to question three. DOGE didn't have a credible answer to question one.
Bottom Line
159 deployments. 21 industries. Six pulled back.
The pattern in the wins: narrow scope, measurable output, human judgment preserved at the decision boundary.
The pattern in the failures: AI trusted to make categorical decisions in ambiguous contexts without a fallback.
The AI Use-Case Library is free, searchable, and linked to primary sources. Before your next AI pitch — or your next AI budget review — it's the first document worth opening.
The audit era isn't coming. It's here. The companies that win it will be the ones who already built their precedent file.
Want more analysis like this? Follow THE DAILY BRIEF for enterprise AI insights twice a week — no hype, just what's working and what isn't.
