FDA clearance tells you a device matched its maker's test set. It does not tell you a patient got better. Of the 1,357 AI and machine-learning devices the agency had cleared through 5 December 2025, exactly three — 0.2% — have any publicly documented evaluation against an outcome a patient would recognize: death, stroke, readmission, or quality of life. That analysis published on 19 August, one day after the FDA opened a public docket asking how it should regulate the next generation of these devices. If you sign clinical AI contracts, the evidence you assumed the regulator had already collected is evidence you now have to buy yourself.
What the 1,357 Number Actually Counts
It counts market authorizations, not demonstrated benefit — and the attrition between the two is brutal. Rawan Abulibdeh of the University of Toronto and colleagues cross-matched the FDA device database, the ACR Data Science Institute catalogue, ClinicalTrials.gov and PubMed, and published the result in PLOS Digital Health on 19 August 2026. Of 1,357 cleared devices, 34 (2.5%) had a registered prospective trial. Twelve (0.9%) posted results. Twelve (0.9%) produced a peer-reviewed publication. Three (0.2%) were measured on mortality, major morbidity, hospitalization as a primary endpoint, or a validated quality-of-life instrument.
Read that as a floor, not a census. The method is registry-based, and the authors say so: it "may undercount proprietary validation studies conducted internally by manufacturers but not publicly disclosed," some 510(k) summaries were not accessible in full text, and unregistered validation studies sitting in the literature were not systematically searched. For a buyer the caveat cuts the wrong way. Evidence a vendor holds and will not show you is evidence you cannot act on, and the fix for both the real gap and the invisible one is the same: make the vendor produce it.
The 34 trials that do exist are narrow. Thirty-two of them — 94% — were industry-led. Nearly three-quarters (73.5%) enrolled fewer than 500 participants and a quarter enrolled fewer than 100. Sixty-eight percent ran only in the United States. Nine of the 34 (27%) reported any demographic subgroup analysis at all.
Then look at who was left out. Pregnancy was an exclusion criterion in 42% of the cardiovascular trials and 33% of the radiology trials. Non-English speakers and patients with mental impairment were routinely omitted; pediatric populations were excluded almost universally. "We expected the evidence base to be thin, but not this thin," the authors said in a statement released with the study. Their conclusion is blunter: regulatory approval has outpaced clinical validation.
The concentration matters too. Radiology accounts for 1,059 of the 1,357 devices — 78% — and its trial registration rate is under 1%, against 9.5% in cardiovascular and 9.7% in neurology. The busiest corner of clinical AI is also the least studied. That aligns with the FDA's own running tally: the AI-Enabled Medical Device List held 1,451 entries through 31 December 2025, of which 1,104 (76%) were radiology.
Why a 510(k) Clearance Cannot Carry Your Risk
Because the 510(k) pathway was never designed to answer the question you are asking it. The Congressional Research Service's primer on FDA regulation of AI-enabled devices sets out the mechanics: roughly 1,450 devices have been authorized since 1995, most through 510(k) premarket notification, and substantial equivalence requires a manufacturer to show parity with a legally marketed predicate on intended use, technological characteristics and performance testing. It does not require new clinical evidence. A vendor can clear a device by demonstrating it behaves like a device that itself cleared by resembling an earlier one.
"Substantial equivalence" is a comparison to a predicate product, not a demonstration that patients benefit. That single sentence is the whole gap, and it is the sentence your clinical governance committee needs on a slide.
The consequences are not hypothetical. The Epic Sepsis Model was never an FDA-cleared device — it shipped as clinical decision support inside an EHR, which is a different regulatory lane — but its external validation is the clearest picture in the literature of what happens when a widely deployed model is finally tested outside its developer's data. Wong and colleagues ran it against 27,697 patients and 38,455 hospitalizations at Michigan Medicine and reported in JAMA Internal Medicine an AUC of 0.63, sensitivity of 33%, and positive predictive value of 12%. It missed 1,709 of the patients who developed sepsis — 67% — while firing alerts on 18% of all admissions. Hundreds of hospitals had it switched on.
If your organization is one of the many that still lacks a place to run that kind of test before go-live, that is the more urgent problem: fewer than half of health systems have a sandbox for third-party clinical AI.
Devices With No Clinical Evidence Get Recalled More
The absence of evidence is not neutral — it is independently associated with failure in the field. Researchers at Johns Hopkins and Yale examined 950 FDA-authorized AI devices through November 2024 in a cross-sectional analysis and found 60 of them tied to 182 recall events, with devices lacking clinical validation recalled at substantially higher rates than those backed by retrospective or prospective study. The most common causes were diagnostic or measurement errors, then loss or delay of functionality.
Two findings from that work belong in your risk register. First, 43% of all recalls happened within a year of authorization — the failure mode is early, not late, so the "we'll watch it for a while" posture buys less protection than it feels like it does. Second, the pattern tracks the vendor, not just the product: publicly traded companies accounted for about 53% of the devices but more than 90% of recall events and 98.7% of recalled units, and their validation rates were markedly worse than private manufacturers'. Lead author Tinglong Dai's explanation for why so many vendors skip clinical study was that it is not required, so it does not get done — and he traces that directly to the 510(k) pathway.
The American Hospital Association pushed the same numbers to its members a year ago, telling hospital leaders to keep a closer eye on clinical validation because 510(k) clearance does not require prospective human testing. A year on, the PLOS numbers say the base rate has not moved.
The Strongest Argument Against Demanding Trials
There is a real case on the other side, and it is worth stating properly before dismissing it. A randomized outcome trial for a triage algorithm is expensive, slow, and often structurally unfair: the model does not treat anyone, the clinician does, so a mortality endpoint measures the whole care pathway rather than the software. Diagnostic accuracy is a legitimate intermediate endpoint for a tool whose job is detection. AdvaMed, the medical device industry association, argued in its February 2026 submission on federal AI policy that oversight should stay inside the FDA's risk-based and least burdensome principles, warning that blanket post-market and documentation requirements would create substantial burden without proportionate clinical benefit — and that a deep learning system and a more traditional, less complex machine learning model do not warrant identical scrutiny.
Take that seriously, and the conclusion still holds. The argument is about what the regulator should require before a device may be sold. It says nothing about what you should require before you deploy it into a clinical workflow and put your license and your malpractice exposure behind its output. A floor set by least-burdensome principles is a floor, and a buyer is entitled to a higher one. The same reasoning applies anywhere a certification is being treated as proof of performance — it is why inheriting a model vendor's own safety assurance is a governance decision rather than a technical fact, and why concentration in third-party evaluation is worth naming in a contract.
What the FDA Is Actually Proposing for Generative AI
A physician-style competency exam, not an outcome trial — which is why the comment window matters. On 18 August the FDA's Digital Health Center of Excellence, inside the Center for Devices and Radiological Health, published a roughly 30-page discussion paper on regulating generative AI-enabled medical devices. It sketches a two-axis framework for assessing risk, a premarket path built on "competency assessment" — non-clinical device benchmarking plus clinical confirmation, modelled at a high level on how physicians are trained and examined — and several options for risk-proportionate postmarket monitoring. Foundation models and agentic systems are explicitly in scope.
This is a discussion paper. It is not guidance and it changes no rule. Acting Commissioner Kyle Diamantas, CDRH director Michelle Tarver and DHCoE director Rick Abramson all framed it as an opening bid rather than a decision, with Tarver saying patients and clinicians deserve a regulatory approach that keeps pace with digital health innovation. Feedback goes to docket FDA-2026-N-7874 on Regulations.gov, and the window closes on 19 October 2026.
Note what "competency assessment" implies. A physician who passes boards is then supervised, credentialed locally, and reviewed on the job. If the FDA borrows the metaphor, the postmarket half is where patient-outcome evidence would actually live — and the postmarket half is the part still written as open questions. Congress has already told the agency to assess whether it has sufficient authority for post-deployment performance monitoring, and final guidance on predetermined change control plans landed in August 2025. Deciding how a device that rewrites its own outputs gets watched in production is precisely the thing a health system, not a vendor, should be filing comments about. Federal buyers have seen how much leverage a single procurement clause carries — one GSA AI clause reshaped a $91.8 billion market.
What Your Clinical AI RFP Has to Say Instead
Stop treating "FDA cleared" as a validation answer and start treating it as a threshold question. Clearance tells you the device may legally be sold. Everything you care about after that has to be extracted from the vendor in writing.
This Week: Pull your inventory of AI-enabled devices and clinical decision support in production and mark each one with its authorization pathway — 510(k), De Novo, PMA, or none, because CDS and EHR-embedded models are frequently unregulated. Then add one column: does a registered trial, posted result or peer-reviewed publication exist for this exact device and version? On the base rate, expect roughly 1 in 40 to have a registered trial and 1 in 100 to have published results. An inventory you can filter is the prerequisite for everything else, and it is the one governance artifact worth buying rather than building.
This Month: Rewrite the evidence section of your clinical AI RFP to ask five questions that "cleared" does not answer. What is the primary endpoint of your strongest study, and is it accuracy or an outcome? What was the population — age, pregnancy status, language, comorbidity — and how does it compare to our panel? Show subgroup performance, not aggregate AUC. What is your predetermined change control plan, and what triggers a re-validation on our data? What performance telemetry do you expose so we can detect drift ourselves, and what is our contractual remedy when it degrades? Require a named clinical owner on your side who signs off before go-live, and remember that a plausible-sounding explanation from the tool makes reviewers more deferential, not less — vague explanations increased novice trust in exactly the reviewers least equipped to override.
Before Renewal: Convert the answers into contract terms. Local validation on your own retrospective data before go-live, with an agreed accuracy floor and an exit right if it is missed. Shadow-mode operation for a defined period. A drift SLA with monitoring you control rather than a dashboard the vendor hosts — the same monitoring-and-kill-switch discipline any production AI system needs. Notification obligations on recall, model update, or FDA action within a fixed number of days.
The governance scaffolding for this now exists and you do not have to invent it. The Joint Commission and the Coalition for Health AI published joint guidance on responsible AI use in September 2025 covering local validation and ongoing monitoring, and the Joint Commission announced a voluntary Responsible Use of AI in Healthcare certification on 1 June 2026, built on five areas: governance, data management, risk and bias reduction, monitoring and validation, and transparency and training. Read the fine print, though — that certification assesses your organization's governance. It does not evaluate or certify individual AI products. It is one more credential that is not a substitute for evidence, which is the entire lesson of the 1,357 number.
The Bottom Line
Every regulated industry eventually discovers that its certification means something narrower than the market assumed. Financial firms learned it about credit ratings. Security teams learned it about compliance attestations, which is why pre-deployment testing of frontier models became a procurement question rather than a vendor claim. Healthcare is learning it about clearance, with an unusually clean number attached: 1,357 devices, three tested on whether anyone lived longer or better.
The FDA is asking, until 19 October, what the bar should be for the next generation. If health systems do not answer, device manufacturers will — and they have already filed.
A clearance is permission to sell. It was never a promise that it works.
Continue Reading
- Hospitals Test Vendor AI. Fewer Than Half Have a Sandbox.
- The Vaguer the AI Explanation, the More Novices Trusted It
- Anthropic's Evals Maxed Out. Stop Inheriting Its Assurance.
- One Firm Ran Three Labs' Cyber Evals. Name Yours.
- EU AI Act Governance Tools: Buy Inventory, Not Policy Packs
- Feds Now Pre-Test Every Frontier AI Model: Vendor Playbook
