If your health system bought an AI coding or documentation tool on the strength of the revenue it "finds," the payers have just told you how they plan to take some of it back. The Blue Cross Blue Shield Association says hospital AI coding added $942 million in costs to Blue plans over 2024 and 2025, about $653 million of it from secondary diagnoses — and its evidence is a diagnosis the chart gains while the treatment never changes. That gap is the test your coding tool's lift now has to survive. Until it does, the revenue on the vendor's dashboard is a receivable, not a result.
This is not a story about whether AI coding is fraud. Nobody can prove that from claims data, and BCBSA does not claim to. It is a story about which half of your coding tool's ROI is durable, and how you find out before a payer's auditor does it for you.
What Did Blue Cross Actually Measure?
Blue Cross measured coding intensity — the share of inpatient stays billed at the most complex, highest-paying severity levels — and found it rising faster than the care delivered. According to Medical Daily's account of the white paper, complex cases rose from 37% to 40% of Blue plan inpatient claims between early 2023 and late 2025, across more than 55,000 affected claims averaging roughly $11,000 each. Major bowel procedures coded at the highest complexity went from 10.2% to 22.7%. Note what the paper does not observe: which hospitals actually run AI coding tools. It sorts hospitals by how fast their coding complexity grew and attributes the excess to AI.
The claim that matters is not the dollar figure. It is the mismatch. The association found a sharp increase in patients documented as having complex conditions and "no evidence of corresponding change in care delivered," as TechCrunch reported. Its showcase is anemia: hospitals with the highest anemia diagnosis rates showed lower transfusion rates, 16.9% against 19.3% elsewhere.
"The disconnect between diagnoses and treatment suggests that AI is identifying more billable conditions, not sicker patients," BCBSA senior vice president Luke Chalker told Medical Daily. The association's clinical VP, Dr. Razia Hashmi, was more careful: "There may be an element of correct coding there, but the likelihood that this is technology-enabled upcoding is higher, in my view."
This is the second installment. An earlier Blue Health Intelligence analysis, covering Q2 2022 through Q1 2025, put the total at $2.3 billion — $663 million inpatient, more than $1.67 billion outpatient. It found the top 10% of hospitals raised complex-coded admissions by 13.1 percentage points, to 59.8%, while everyone else rose 4.2 points. In maternity, postpartum anemia coding at the high-growth hospitals went from 4% to 12.3% while low-growth hospitals went from 7.9% to 8.2%, and transfusion rates stayed relatively flat, per HFMA's summary. That study pinned $22 million in added maternity costs in one year on posthemorrhagic anemia coding alone.
Is This Upcoding or Better Documentation?
Honestly, some of both, and the white paper cannot tell you the proportion. Take the hospitals' side seriously, because it is strong.
The American Hospital Association's fact sheet says hospital case-mix index rose about 5% between 2019 and 2024, and 19% of hospital expense growth over that period reflects sicker, more complex patients. It points to an aging population, more chronic disease, lower-acuity care moving to outpatient settings, and changes to coding guidelines. It calls payer "downcoding programs" that cut reimbursement with automated edits "without reviewing medical documentation" the real abuse.
HFMA's Shawn Stack made the clinical point that undercuts BCBSA's showcase: capturing these diagnoses more accurately "does not necessarily represent upcoding," and postpartum anemia managed conservatively, without a transfusion, is still a clinically valid diagnosis. Not every real condition gets a procedure. And the source is a payer trade group whose members pay the claims — a party with a financial stake in the conclusion, publishing a claims-only analysis rather than a chart review.
University of Minnesota economist Hannah Neprash put the honest answer on the record: "It's probably somewhere in between."
Here is why that does not let you off the hook. "Somewhere in between" is exactly the zone where money moves on audit. A payer does not need to prove fraud to recover a payment. It needs to argue the diagnosis lacks clinical support.
Why Clinical Validation Is the Test That Matters
A clinical validation denial is a payer's refusal to pay for a correctly coded diagnosis on the grounds that the patient's record does not clinically support it. Health-care lawyers at Davis Wright Tremaine draw the line sharply: a DRG downgrade challenges code assignment, while a clinical validation denial "do[es] not challenge the accuracy of ICD-10 code assignment" — the payer argues the code lacks clinical support, often under the payer's own medical policy.
That distinction is what makes BCBSA's anemia chart dangerous to you. An AI coding tool is built to be right about coding. It reads 100% of charts, finds a lab value and a physician note, and queries for the secondary diagnosis that bumps the stay to a higher-paying tier. The coding can be impeccable and still lose a clinical validation fight, because the argument is about whether the condition was clinically significant enough to count — and "the treatment never changed" is the payer's opening exhibit.
The same firm lists the DRG families where clinical validation denials cluster: sepsis (871-873), acute kidney injury (682-684), malnutrition (951-953) and encephalopathy/stroke (064-066). These are the secondary-diagnosis bumps an AI query engine is designed to find. The earlier BCBSA study had already flagged patients billed for sepsis without treatment consistent with that diagnosis.
Why the Payer Counter-Move Is Coming Now
Because the payers have put it in their pricing. A PwC survey of actuaries at 27 health plans found nearly 70% ranked providers' AI documentation and coding tools among their top three cost inflators for 2027, and about 20% called it the single biggest, inside a projected 9% commercial trend. When a cost driver is in the actuarial assumptions, the program to claw it back is in the budget.
The regulator is not arguing either. CMS Administrator Dr. Mehmet Oz said on September 24 that "short term, AI is going to be inflationary because it's going to turbocharge the ability of the current billing systems to work more effectively," pointing to accountable care as the longer-run fix.
And the language has turned hostile. Chalker described the insurer position as "a completely one-sided blood bath". Abridge founder Dr. Shiv Rao, per the same TechCrunch piece, warned of "bots fighting bots, agents fighting agents." That is the likely equilibrium: your AI adds the diagnosis, theirs flags it, and the dispute lands in a queue your CDI team staffs.
What the Vendor's ROI Number Doesn't Count
The vendor's ROI number counts revenue billed, not revenue kept. SmarterDx's McLaren Health Care case study reports $11.3 million in "realized annual net new revenue" from reviewing 100% of charts before billing. The visible case study does not describe how that figure nets out later denials, appeals or recoveries. SmarterDx sells denials prevention as a separate product, SmarterDenials.
That is not a knock on one vendor; it is how the category reports. The same pattern ran through contact-center AI, where containment rates were reported as if they were resolutions. The honest metric for a coding tool is retained lift: incremental reimbursement from AI-prompted diagnoses, minus what is denied, downgraded or recovered on audit, minus the CDI and appeals labor to defend it — measured on claims old enough for the audit window to have closed.
If your contract pays the vendor a percentage of "identified" revenue, you are paying on the gross and absorbing the net. Our guide to agentic AI pricing makes the general case: an outcome fee is only as good as the definition of the outcome.
What to Do About It
The work is to separate the lift that will hold from the lift that will not, before a payer does it on their terms.
This Week:
- Pull the AI-prompted secondary diagnoses by DRG family. Ask your coding or CDI lead for every CC/MCC added via an AI query in the last 12 months, grouped by sepsis, AKI, malnutrition, encephalopathy and postpartum anemia. That list is the payers' target list.
- Run the treatment test yourself. For each family, check what share of added diagnoses had a documented treatment change — a transfusion, an antibiotic escalation, a dietitian consult. Low shares are not proof of anything, but they are exactly where you will be challenged first.
This Month:
- Demand the evidence trail from the vendor. For every query the tool raises, you need the clinical indicators it relied on, stored with the claim and retrievable for appeal. If it cannot produce that per claim, your appeals team is rebuilding it by hand.
- Re-baseline ROI on retained revenue. Have finance restate the tool's value as billed lift minus clinical validation denials, downgrades and recoveries, on claims past the payer's audit and appeal window — which ranges from 30 days to a year by payer.
Before Renewal:
- Hold a reserve against the gain. Book a denial and recovery reserve against AI-attributed lift, sized from your own step-2 numbers, rather than spending the gross.
- Rewrite the fee basis. Move any contingency fee from "identified" to "paid and not recovered after the audit window," with a clawback when a payer recovers. A vendor confident in its clinical support should accept that.
The Bottom Line
Every coding cycle has run this loop. Medicare's DRG system arrived in the 1980s and hospitals learned to code for it; Medicare Advantage risk adjustment did the same, to the point that the AHA's own fact sheet cites MedPAC's finding that upcoding contributed to $40 billion in Medicare Advantage overpayments. Each time the payer response came a few years later, and it was retroactive.
AI compresses that loop. Hospitals found the revenue in months; the payers have put a price on it in two white papers and an actuarial survey. The audit phase is next, and it will be fought on clinical support, not coding accuracy. A health system that measures its coding tool on what survives that fight will keep most of the gain. One that measures it on the vendor's dashboard is booking revenue a payer has already put on its list.
The diagnosis the treatment never followed is the payer's first question. Make sure it is also yours.
Continue Reading
- Best Ambient AI Scribes Save 16 Minutes a Day, Not an Hour
- NHS Scribes Dropped 'Null.' Patients Caught It, Not GPs.
- Hospitals Test Vendor AI. Fewer Than Half Have a Sandbox.
- FDA Cleared 1,357 AI Devices. Three Were Tested on Patients.
- Lung Cancer AI: Chart Data Held Up. Scans and Slides Didn't.
- Agentic AI Pricing: Don't Buy Consumption Without a Cap
