950 cleared devices, three tested on outcomes
A team at MIT Critical Care just published the first honest accounting of the radiology AI evidence base. Out of 1,357 FDA-authorized AI medical devices, three have been tested on whether they actually improve patient outcomes. Radiology is the thinnest specialty. Procurement teams should read this before the next vendor demo.
The FDA has authorized 1,357 AI-enabled medical devices. Three of them have been evaluated for their impact on mortality, morbidity, or readmissions. Radiology is the specialty with the thinnest evidence base of any major category. We should stop calling this an evidence gap and start calling it what it is — a procurement problem.
The study landed in PLOS Digital Health on August 19, 2026. Lead author is the MIT Critical Care team. The headline is the headline: of 1,357 FDA-authorized AI medical devices cleared or approved through the end of 2025, only three were tested on patient-centered outcomes. In radiology specifically, three of 1,059 cleared devices had a registered prospective trial. That is 0.3 percent of the tools sitting in our reading rooms.
If you have been in a procurement meeting in the last two years, you already felt this. The vendor's slide deck shows a benchmark, the benchmark looks impressive, and somewhere in the back of your mind is the question you never quite get to ask. The MIT team just put that question on paper.
What the study actually measured
The methodology is what makes this paper land. The team did not count press releases. They linked the FDA device database against the ACR Data Science Institute catalog, then cross-referenced every device against registered trials on ClinicalTrials.gov and peer-reviewed publications on PubMed. They did the work the way a procurement team would — pulling the public record for every cleared device and asking, "what evidence sits behind this clearance?"
The numbers, layer by layer:
- 1,357 total FDA-authorized AI/ML medical devices, through December 5, 2025.
- 76% of those are radiology devices — 1,059 in our specialty.
- 34 devices (2.5%) had a registered prospective trial of any kind.
- 12 devices (0.9%) had posted results.
- 12 devices (0.9%) had peer-reviewed publications.
- 3 devices (0.2%) evaluated patient-centered outcomes — mortality, morbidity, readmissions.
- 3 of the 1,059 radiology devices had a registered trial.
Read that again. Out of 1,059 cleared radiology AI tools, three have a registered prospective trial. The other 1,056 cleared on the basis of substantial equivalence to a predicate device — which is the FDA's polite way of saying "this is similar enough to something already on the market to clear without new clinical evidence."
What "510(k) substantial equivalence" actually means
This is the part of the conversation where vendors get careful. The 510(k) pathway is the workhorse clearance route for AI medical devices — over 90% of cleared devices went through it. The pathway is designed to confirm a new device is substantially equivalent to an already-cleared predicate. It is not designed to confirm the new device improves patient outcomes. It never was.
The MIT lead, Dr. Oscar Ordóñez, said the right thing out loud to AuntMinnie: "Under the 510(k) pathway, clearance means a device is substantially equivalent to something already on the market, not that using it makes patients better off. Clearance is the beginning of due diligence, not the end." The full device list and trial identifiers are published with the paper. A radiologist can look up the tool in their own reading room and see exactly what public evidence stands behind it.
This is not an indictment of the FDA. The 510(k) pathway exists for a reason — it routes low-to-moderate-risk devices efficiently so higher-risk devices get more scrutiny. The problem is that the radiology AI ecosystem has quietly built an entire market on top of a clearance pathway that was never designed to validate clinical benefit. When a hospital procurement officer signs a seven-figure contract for a "FDA-cleared" AI tool, the clearance does not mean what they think it means.
The exclusions are worse than the headline
The number that worries me more than 0.3% is the exclusions. The MIT team found that 62% of the underlying studies used observational designs with small, homogenous cohorts and frequent exclusion of vulnerable populations. Specifically:
- Pregnancy excluded in one-third of radiology trials.
- Pediatric patients excluded almost universally.
- Non-English speakers frequently omitted.
Those patients still show up in your emergency department. They show up in your outpatient clinic. They show up in your trauma bay at 2 AM. And the tool you bought to help you read their study was never validated on anyone like them.
This is not a vendor-specific failure. It is a structural failure of how the evidence gets generated. But it lands in the reading room.
The three questions every procurement team should be asking
Dr. Ordóñez distilled the procurement question into three lines that should be tattooed on every CMIO and radiology chairman:
- What was the trial registration number and primary endpoint? If the vendor cannot hand you a ClinicalTrials.gov ID and a defined primary endpoint, the answer is they did not run one.
- Does the validation population resemble ours? Age distribution, disease prevalence, modality mix, acquisition protocol. If your site reads 40% inpatient and 60% outpatient but the validation cohort was 90% outpatient screening, the model's operating characteristics on your population are unknown.
- Who was excluded? Pregnancy, pediatrics, non-English speakers, patients with prior imaging at outside institutions. If the exclusions match your patient population, the device is unvalidated for your reading room.
None of these questions are unreasonable. None of them require the vendor to disclose proprietary IP. All of them are answerable from public records in five minutes. If a vendor balks at any of these three questions, the conversation should end.
What the authors propose — and what is actually actionable
The MIT team proposes a three-phase evidence ladder:
- Phase 0 (pre-clearance): retrospective validation on diverse datasets with mandatory demographic reporting.
- Phase 1 (around clearance): prospective studies of at least 500 patients embedded in real workflows.
- Phase 2 (post-clearance): multi-center outcome trials of at least 2,000 patients with pre-specified subgroup analyses.
The 500-patient bar is where 73.5% of today's trials fall short. That is the structural gap. The agency's most immediate lever is requiring pre-registration of prospective trials for Class II and III AI devices — a regulatory ask, not an aspirational one.
But the FDA is not the only lever. The authors point to journals (enforce SPIRIT-AI and CONSORT-AI reporting standards) and to payers (tie reimbursement to demonstrated benefit rather than regulatory status). The CMS NTAP mechanism for Aidoc's CT triage device is the first concrete example of reimbursement tied to a specific AI tool — and it will be the test case for whether payment follows clearance or evidence.
What this means for the people writing the checks
Health system procurement is where this lands. The next time a vendor walks in with a benchmark slide, the right opening question is not "what is your AUC." It is "what is your ClinicalTrials.gov ID, what is your primary endpoint, and what was your exclusion criteria." The first three minutes of that meeting will tell you whether the rest of the meeting is worth having.
If the answer is "we have a retrospective validation on 10,000 studies from three academic sites," the answer is not wrong. It is just incomplete. Retrospective validation is Phase 0. It tells you the model can run. It does not tell you it helps patients. Phase 1 evidence — prospective, in your workflow, on your population — is the question. Phase 2 — outcomes — is the only one that matters for value-based-care contracts.
For the people reading this who are radiologists: your job is not to memorize the regulatory framework. Your job is to refuse to sign off on a tool whose evidence you cannot summarize in one sentence. If the vendor cannot summarize theirs, neither can you.
What it means for the people building the tools
I sit on the building side. The PLOS paper is uncomfortable reading for anyone shipping radiology AI in 2026. The temptation is to defend the category — to point out that radiology has the most devices because the imaging data is the most structured, that the 510(k) pathway is doing what it was designed to do, that prospective trials are coming.
The defense is true and it is also beside the point. The point is that a market built on 510(k) substantial equivalence is a market that has outsourced clinical validation to the buyer's reading room. That is not sustainable. The buyer will eventually ask for evidence. The vendor who does not have it loses.
The vendors who will win the 2027–2028 procurement cycle are the ones who started running their Phase 1 prospective trials in 2024–2025. The window for "we are working on it" is closing. By the time the next PLOS Digital Health paper lands — and there will be a next one — the procurement teams will be asking not whether the trial exists, but whether the trial enrolled patients who look like theirs.
What to do on Monday morning
If you are a radiology chairman or CMIO reading this with a stack of vendor contracts on your desk:
- Pull the FDA 510(k) summary for every AI device running in your reading room. Read the "substantial equivalence" claim. Understand what the predicate was.
- For each device, search ClinicalTrials.gov for the manufacturer's name and the device name. Count the registered prospective trials. Count the results that have posted.
- For each device, ask the three questions. Put the answers in writing. Vendor "we don't have that data" is an answer — it tells you the answer is no.
- For any device where the answer to question 1 is "we don't have a registered trial," build a Phase 1 evaluation into your next contract. Embed the AI in your workflow for 90 days, measure prospectively, and decide on your own data.
If you are a radiologist who reads the studies the AI produces: your signature is on the report. The audit trail records that you signed. The PLOS paper is a preview of the deposition question — "doctor, what evidence did you review before agreeing to use this tool?" — that is coming. Build the answer now, while you still have time.
The PLOS Digital Health paper is open access. Read it. Cite it. Send it to your procurement counterpart. The era of "FDA-cleared" as a sufficient answer is over. The era of evidence-first procurement is here.
— Aldo Ruffolo, DO, MBA
Founder, ApertureAI
Connect on LinkedIn
Source: Ordóñez O, et al. "Evaluation of prospective evidence supporting FDA-authorized AI/ML-enabled medical devices." PLOS Digital Health. August 19, 2026. Full text and device-trial identifiers list (open access). Coverage: "Most FDA-cleared AI medical devices not tested on patient outcomes," AuntMinnie, August 19, 2026.