The self-described largest real-world multimodal oncology AI study found that chart and blood data carried the prediction, and adding CT scans and pathology slides produced no gain that survived held-out testing. That result, from the I3LUNG consortium's Nature Medicine paper published 13 September, changes the order in which a health system should spend: validate a decision-support model on your own patients, and watch how your clinicians respond to it, before you pay to feed it more data types.
This matters because the opposite story is already circulating. Press coverage has led with an AUC of 0.88 for a model combining clinical, blood, imaging and digital pathology data. The paper's own abstract says the incremental benefit of that integration "remains uncertain, not translated in TEST and EXVAL." If you are a CMIO or head of oncology informatics building a business case for a radiomics-plus-pathology pipeline, those are two very different inputs.
What Did I3LUNG Actually Test?
I3LUNG tested whether AI could predict which patients with advanced non-small cell lung cancer (NSCLC) benefit from immunotherapy, using data hospitals already hold. The study enrolled 2,396 patients and fused four data types — clinical and blood data, CT images, digital pathology and genomics — into machine learning and deep learning models. Its authors describe it as currently the largest international, real-world, multimodal AI study of its kind.
The clinical problem is real. Immunotherapy selection in NSCLC, the paper notes, still leans on subgroup analyses and an imperfect PD-L1 biomarker. Six centres contributed patients, in Italy, Germany, Greece, Israel, Spain and the United States.
The headline result is good news for the cheap version. Models built only from clinical and blood data reached an AUC of up to 0.77 on the independent test set and significantly outperformed PD-L1, ECOG performance status, neutrophil-to-lymphocyte ratio, LDH and the LIPI score. AUC (area under the ROC curve) measures how well a model ranks patients who have an outcome above patients who don't: 0.5 is a coin flip, 1.0 is perfect.
"Decision support tools built even from routinely available clinical data can outperform the biomarkers we rely on today," senior author Marina Garassino of UChicago Medicine said. Read that sentence carefully. The claim that holds up is about routinely available data.
The 0.88 Everyone Quotes Came From Cross-Validation
The 0.88 figure is a cross-validation result in one subgroup, not a test-set or external-validation result. According to the paper, the multimodal models were compared in cross-validation "due to the small TEST size", and the 0.88 applies to first-line patients, predicting 24-month overall survival. Cross-validation reuses the development data, which makes it the step most likely to flatter a model. Inside Precision Medicine's report put it plainly: adding CT, pathology and genomic data showed no consistent benefit in independent testing.
Why the test set was small is the whole story for a buyer. Of 2,396 patients, only 339 had all four modalities, and only 936 had digital pathology. Even in a multi-centre research consortium built around collecting this data, fewer than one patient in seven arrived with the complete set.
The published paper is more cautious than the preprint. The January preprint said multimodal models outperformed clinical data in clinical-specific subgroups, with AUCs up to 0.86. The Nature Medicine version calls the benefit uncertain. Its listed limitations land squarely on the imaging pipeline: roughly 15% of CT scans were not taken at immunotherapy baseline, radiomic features were restricted to primary lesions, and genomic data was sparse.
Steel-man the other side before dismissing it. Uncertain is not the same as absent. With 339 complete cases, the study could not have proven a modest multimodal gain even if one exists, and the consortium's prospective validation in more than 2,000 patients covers both the clinical-only and the multimodal systems. The right reading is not "imaging is useless." It is "nobody has yet shown that paying for it improves this decision."
Why Local Validation Matters More Than Modalities
The larger risk in this study is not the missing modality — it is the drop when the model left home. External validation means testing a model on patients from a site or population it was never trained or tuned on. In I3LUNG's external validation, AUCs fell to a range of 0.55 to 0.72, which the authors attribute to population differences. At the bottom of that range, the model is barely better than chance. The authors also note that external validation was limited to a single cohort with different baseline characteristics.
Health systems have seen this before. When researchers externally validated the widely implemented Epic Sepsis Model across 38,455 hospitalizations, it produced an AUC of 0.63, against 0.76 to 0.83 in the developer's internal documentation. At its alert threshold it missed 67% of sepsis cases while alerting on 18% of all hospitalizations. No extra data type would have fixed that. Local testing would have caught it.
Accreditation guidance already expects this. The Joint Commission and Coalition for Health AI guidance, released on 17 September 2025, tells organisations to ask vendors during procurement "whether they are willing to tune and/or validate a sample that is representative of the deployment context," and to "ensure they are appropriately tuned and/or tested on local data."
For certified EHR software, ONC's HTI-1 rule defines 31 source attributes for predictive decision support interventions, including the external validation process and the validity of the intervention in local data, with developer obligations starting 1 January 2025. ONC's HTI-5 proposed rule, not yet final, would remove those model-card requirements. It covers predictive tools the certified developer supplies, not every third-party model you bolt on. The paperwork for the question exists. Most organisations have not yet made the answer a go-live gate.
The testing environment is its own gap — see why fewer than half of hospitals testing vendor AI have a sandbox, and how thin the patient-outcome evidence is behind FDA-cleared AI devices.
Experts Followed Wrong AI Advice More Often
In I3LUNG's reader study, lung cancer experts accepted incorrect AI suggestions more often than non-experts did. Twenty physicians, ten experts and ten non-experts, predicted disease control with and without the explainable AI tool. With the tool, sensitivity rose from 0.72 to 0.87 and accuracy from 0.57 to 0.65. But when the AI was wrong, experts followed it 72.2% of the time and non-experts 63.6% — a gap between two groups of ten, for which the paper reports no significance test.
Automation bias is the tendency to accept a system's recommendation even when it is wrong. The usual assumption is that seniority protects against it, and a 2023 mammography study in Radiology partly supports that: when the AI suggested the wrong category, inexperienced radiologists' accuracy fell below 20%, while very experienced readers fell to 45.5% — better, but still badly degraded. I3LUNG points the other way on seniority, in a small sample. Treat it as a warning rather than a law: twenty physicians cannot settle the question, but they are enough to stop you assuming your most senior oncologists are the safety net.
Some coverage also blurred the reader study, describing the 0.72-to-0.87 improvement as an AUC gain. The paper reports it as sensitivity. Rising sensitivity alongside experts following wrong suggestions is the pattern you would expect when clinicians defer to the tool.
Regulators are watching the same failure. The FDA's clinical decision support guidance, revised 6 January 2026, ties its automation-bias concern to the statutory test that clinicians must be able to independently review the basis for a recommendation and "not rely primarily on" it. You cannot show that without data on when clinicians overrode the tool and when they should have. The same dynamic shows up in how vague AI explanations raised novice trust.
What Imaging and Pathology Pipelines Really Cost
A digital pathology pipeline is a large, mostly-IT fixed cost, and I3LUNG gives you no AI-accuracy reason to add one for this decision. Memorial Sloan Kettering's published accounting of its 2021 operations is the most detailed public reference: it ran 25 whole-slide scanners, allocated 9 PB for slide storage with similar storage again for redundancy, and put the annualized cost of that storage at $1.6 million. IT hardware and software made up 33% of annualized costs; scanner acquisition only 21%. Cost per scan ranged from $0.55 to $19.53, driven by how many slides each scanner processed.
That does not make digital pathology a bad investment. It earns its keep in other ways — slide retrieval, remote review, image sharing — and those cases can stand on their own. The point is narrower: don't let a decision-support vendor's multimodal accuracy claim carry that business case. On the best evidence published this week, it cannot.
The same caution applies to any model validated at someone else's site on a small subset of complete records. It is the sample-size trap that also distorts enterprise model benchmarks.
What to Do Before You Fund the Pipeline
The order of spend is local validation, then override monitoring, then — only if both justify it — more data types.
This Week:
- Pull the HTI-1 source attributes for every predictive model live in your EHR. Ask your EHR team for the "external validation process" and "validity of intervention in local data" entries. Blank or generic answers go on a list for your AI governance committee.
- Count your complete records. For any proposed oncology decision-support tool, have your informatics lead count how many of last year's patients had every required input available at the moment the decision was made. I3LUNG's consortium managed 339 of 2,396.
This Month:
- Run a silent retrospective validation before go-live. Score 12 to 24 months of your own patients, compare predictions with outcomes, and agree in writing on the AUC and calibration floor the model must clear on your population.
- Instrument overrides. Log the AI suggestion, the clinician's decision and the eventual outcome. Report "agreed with the AI when the AI was wrong" by clinician experience level every quarter, and do not exempt senior staff.
Before Budget Sign-off:
- Split the business cases. Justify digital pathology and radiomics on their diagnostic and operational merits, separately from any AI decision-support claim.
- Put local validation in the contract. Require the vendor to validate on a representative sample of your patients, and write rights to local performance data into the data use agreement — both terms the Joint Commission guidance tells you to raise.
The Bottom Line
Health IT has made this mistake before: collect everything, integrate everything, and assume the value follows. The sepsis-alert era taught a simpler lesson — a model that works at its home site can collapse at yours, and the clinicians using it will follow it anyway. I3LUNG is a genuinely useful study, and its strongest finding is the unglamorous one: chart and blood data your system already has beat the biomarker oncologists use today.
That is what to build on. Prove the model works on your patients, prove your clinicians catch its mistakes, and only then buy it more pixels.
Validate first. Then decide what to scan.
Continue Reading
- Hospitals Test Vendor AI. Fewer Than Half Have a Sandbox.
- FDA Cleared 1,357 AI Devices. Three Were Tested on Patients.
- The Vaguer the AI Explanation, the More Novices Trusted It
- NHS Scribes Dropped 'Null.' Patients Caught It, Not GPs.
- ChatGPT Made Spec Reviewers Worse. Save It for Round 2.
- Only 3 of 36 Model Gaps Were Real. Size Your Eval Set.
