Radiology was one of medicine’s earliest testing grounds for artificial intelligence because the specialty already produces the kind of structured digital information machines can process: images, measurements, reports and increasingly large archives of clinical data.

What began as computer-assisted detection has evolved into systems that can flag suspected strokes, prioritize scans, quantify lesions, identify abnormalities, generate report drafts and, in some workflows, determine which examinations may need to be reviewed by a radiologist at all.
That progression is changing the nature of the problem.
The question is no longer simply whether engineers can build an AI model that performs a task accurately. It is whether a model developed in a laboratory can continue to perform safely when it becomes part of a clinical system, encounters a different population, influences a physician and changes as its underlying software evolves.
That distinction matters because radiology AI is already moving from experimentation into routine care.
The U.S. Food and Drug Administration had authorized more than 1,600 AI-enabled medical devices by September 2026. Radiology accounts for the overwhelming majority of those devices. A 2025 systematic review of FDA authorizations found that, among 950 AI/ML devices authorized through June 2024, 723, or 76%, were radiology devices. But the same review found a striking gap between regulatory authorization and clinical testing: only 29% of radiology devices with available submission documentation had incorporated clinical testing, while just 5% had undergone prospective testing and 8% included a human-in-the-loop component.
The authors concluded that those testing gaps underscored the need for clinical oversight.
The numbers expose the central tension in medical AI. A product can satisfy a regulatory pathway designed to establish safety and effectiveness for a specified use without answering every question that arises once that product becomes part of the messy, variable environment of clinical practice.
In radiology, that environment is unusually complex.
A model may be trained using images from one group of hospitals, scanners and patients and then deployed somewhere with different imaging protocols, disease prevalence, patient demographics and reporting practices. A model that performs well in development may therefore behave differently once it leaves the controlled environment in which it was created.
A systematic review published in 2025 examined studies comparing internal and external validation of AI models used on CT, MRI and X-rays. The researchers found that performance generally declined on external data. Median area under the curve dropped by about 0.03, while specificity in some studies fell by as much as 24 percentage points.
The implication is straightforward but consequential: a model is not finished when the developer finishes building it.
That is increasingly reflected in how radiology organizations are approaching AI.
The American College of Radiology approved its first practice parameter for imaging AI in May 2026. The framework calls for imaging facilities to maintain inventories of AI tools, conduct local acceptance testing, monitor real-world performance for drift and safety problems, establish stop rules and maintain appropriate privacy and security controls.
Tessa Cook, chair of the ACR practice-parameter writing committee, described the approach as covering “everything from selection, to monitoring, to continuous quality improvement.”
That language is important because it moves responsibility beyond the developer.
For conventional software, a hospital can largely assume that a particular button will produce a predictable response. AI systems are different. Christoph Wald and Cook wrote in an ACR analysis that AI is “probabilistic, not deterministic,” meaning its behavior can vary in ways that may be subtle or difficult to anticipate.
The clinical consequence is that AI becomes part of the practice environment itself. A radiologist is no longer simply reading an image. The radiologist may be reading an image alongside an algorithmic probability, a highlighted region, a triage score, an automatically generated measurement or a draft report.
The machine becomes another participant in the diagnostic process. And that creates a new form of clinical risk: the human being may be influenced by the machine even when the machine is wrong.
A 2023 study published in Radiology demonstrated this problem in mammography. Twenty-seven radiologists were asked to assess mammograms while receiving AI-generated BI-RADS suggestions. When the AI recommendation was incorrect, radiologists of varying experience were significantly influenced by it. Even very experienced readers were affected.
More recent research has shown that the effect of AI assistance is not uniform.
A 2024 Nature Medicine study examined 140 radiologists performing 15 chest X-ray diagnostic tasks with and without AI assistance. The researchers found substantial variation in how individual radiologists benefited from the technology. Experience, subspecialty and familiarity with AI did not reliably predict who would benefit.
More revealing was the role of AI errors. Incorrect AI predictions could adversely affect radiologist performance across the set of pathologies studied.
That complicates one of the most persistent assumptions about medical AI: that adding a highly accurate algorithm to a human expert will automatically produce a better combined system.
It may. But the interaction itself has to be studied. The question therefore shifts from “How accurate is the AI?” to “What happens when this AI meets this clinician, this workflow and this patient population?”
That is a much harder question. It is also why explainability matters.
A 2024 randomized study involving 220 physicians found that the way an AI system explained its recommendation affected diagnostic performance and physicians’ trust in the system. Local, feature-based explanations produced better diagnostic accuracy when the AI advice was correct and reduced the time physicians spent considering the recommendation compared with global, prototype-based explanations.
The researchers also found that physicians placed greater trust in local explanations regardless of whether the AI advice was correct.
That finding creates a paradox. An explanation can make an AI system more useful, but it can also make a wrong system more persuasive. Trust, in other words, is not the same thing as accuracy.
The evidence is becoming more clinical
Yet the risks should not obscure evidence that AI can produce meaningful improvements in radiology. One of the strongest examples comes from breast cancer screening.
A randomized controlled trial involving 105,934 women in Sweden found that AI-supported mammography detected 6.4 cancers per 1,000 women compared with 5.0 per 1,000 under standard screening. The study also reported a 44.2% reduction in screen-reading workload without a significant increase in recall or false-positive rates.
A separate nationwide real-world study in Germany involving 463,094 women found that AI-supported screening produced a breast cancer detection rate of 6.7 per 1,000 compared with 5.7 per 1,000 under standard double reading, while the recall rate was slightly lower in the AI-supported group.
These are not laboratory demonstrations. They are evidence from clinical environments. But they also illustrate why the debate has moved beyond whether AI “works.”
In the Swedish trial, AI was integrated into a defined screening workflow, with radiologists remaining part of the process. The system did not simply replace the clinical infrastructure around mammography. It changed how that infrastructure operated.
That distinction may define the next phase of medical AI.
The developer becomes part of the clinical ecosystem
Traditional medical-device development has a relatively clear boundary: manufacturers build products and clinicians use them.
AI weakens that boundary because software can be updated, retrained, recalibrated and modified after deployment.
The FDA’s August 2025 guidance on predetermined change control plans explicitly recognizes this problem. It allows manufacturers to specify planned modifications to AI-enabled devices, together with methods for developing, validating and implementing those changes.
The underlying idea is that AI development can become iterative rather than ending at market authorization. That makes sense technologically.
But clinically, it raises a difficult question: when does an improved model become a new clinical intervention?
If a developer changes an algorithm’s training data, modifies its architecture or updates its thresholds, the software may become technically better. But the change could also alter how often it flags disease, how many false positives it generates or how radiologists respond to its recommendations.
The FDA’s lifecycle approach therefore reflects a reality that radiology is increasingly confronting: AI must be governed not just at launch, but throughout its existence.
Keith Dreyer, chief data science officer at the ACR Data Science Institute, has framed the issue in similarly direct terms. “Since the U.S. Food and Drug Administration has not developed post-market surveillance for these algorithms, it is up to radiologists to determine how accurate these models are in practice,” he said. “It is beholden to us to be able to go through a process to validate AI to ensure its accuracy.”
The ACR has responded by developing Assess-AI, a registry intended to monitor AI performance in clinical environments.
That represents a fundamental shift in responsibility. Hospitals and radiologists are becoming evaluators of commercial technology after deployment, rather than passive consumers of finished products.
The data problem is also a clinical problem
There is another reason the distinction between technology and medicine is becoming harder to sustain: data.
AI systems inherit the characteristics of the datasets used to build them. A review of radiology AI datasets found persistent problems involving demographic, geographic, genetic and disease representation. The authors argued that diverse, high-quality datasets are essential to maintaining validity, particularly for underserved populations.
Research has also shown that demographic shortcuts can enter medical imaging models. One study examining AI across radiology, dermatology and ophthalmology found that medical imaging systems could use demographic information as a shortcut for disease classification and that attempts to correct these shortcuts within one dataset did not necessarily produce fairness in new testing environments.
In another study of a commercial AI tool for detecting intracranial hemorrhage on 9,736 head CT scans, researchers found an aggregate accuracy of 93%, but also identified differences in performance across demographic and socioeconomic groups.
This matters because the clinical question is never simply whether a model has a high average accuracy. A hospital does not treat an average patient. It treats individuals.
A model can therefore achieve impressive aggregate performance while performing less reliably for a particular group, scanner, institution or disease presentation.
That is particularly relevant for countries that are not well represented in the datasets used to develop commercial systems.
The World Health Organization has warned that health AI can reproduce bias in training data and has called for patients, healthcare professionals and other stakeholders to participate in AI development from the earliest stages. Its guidance also recommends post-release auditing and impact assessments for large-scale deployment.
For radiology, this means that data governance is not simply an engineering concern. It is clinical governance.
From software procurement to medical oversight
Radiology’s response is increasingly to treat AI implementation as a clinical quality-improvement process.
A 2026 framework published in the Journal of the American College of Radiology proposed four stages: validation, deployment, value assessment and post-deployment surveillance. The basic questions are deliberately practical: Does the system work? Does it help? Does it continue to work?
That last question may ultimately be the most important. AI can drift. Hospitals change scanners. Imaging protocols change. Patient populations change. Disease prevalence changes. Software gets upgraded. Radiologists learn how to interact with systems. Vendors retrain models.
The environment in which the algorithm operates is therefore not static. Radiologists interviewed in a 2025 study described AI monitoring as still relatively immature. Most of the projects studied had not reached fully live clinical deployment, and monitoring was often performed manually by comparing AI results with radiology reports. The researchers identified a lack of resources and uncertainty about how to create scalable monitoring systems as major barriers.
That is a remarkable development in the history of clinical technology. The industry spent years asking whether machines could interpret medical images. It is now having to build infrastructure to determine whether those machines continue to interpret them correctly after they have entered hospitals.
The result is a blurring of roles. Software engineers need to understand clinical workflows. Radiologists increasingly need to understand model limitations, validation and data drift. Hospitals need AI governance structures. Regulators are considering how software can evolve after approval. Patients are becoming stakeholders in systems that may influence diagnoses without ever interacting with the technology directly.
The most consequential development may therefore not be an AI model that matches a radiologist on a benchmark.
It may be the emergence of an entirely new clinical layer around the model: monitoring, validation, accountability, cybersecurity, bias testing, human oversight and rules governing when an algorithm should be ignored or switched off.
AI is making radiology more computational, but that does not make it less clinical. If anything, the opposite is happening.
The closer AI gets to diagnosis, triage and clinical communication, the less meaningful the distinction becomes between building the technology and practicing medicine.
The future of radiology AI will therefore depend on who is responsible for what happens when the model encounters a patient it has never seen before, a workflow it was never tested in, a population underrepresented in its training data or a clinical decision that cannot be reduced to a probability.
The technology may produce the answer. But medicine still has to decide what the answer means.
Stay ahead of the Stories shaping our world. Subscribe to Impact Newswire and join our
WhatsApp Channel for updates on global tech, business, and innovation—all in one place.
Dive deeper into the future with the Cause Effect 4.0 Podcast, where we explore the ideas, trends, and technologies driving the global AI conversation.
Got a story to share? Contact Us to reach a global audience with Impact Newswire.
Discover more from Impact AI News
Subscribe to get the latest posts sent to your email.


