Although AI tools are becoming increasingly integrated into medical imaging workflows, distrust in AI predictions often remains in the form of one key question: how can physicians know when to trust the prediction made by a black-box AI tool?
A new study from the Stanford Radiology AI Development and Evaluation (AIDE) Lab, published October 16 in npj Digital Medicine, illustrates how the Ensembled Monitoring Model (EMM) framework can act like a real-time second opinion system for deployed AI tools. EMM evaluates how much confidence can be placed in the the AI prediction, helping physicians decide whether to rely on the result or take a closer look.
Why This Matters
Radiology AI tools are often deployed without mechanisms to monitor their reliability in real-world use. Because it’s generally not possible to see how black-box AI tools make their decisions or what kinds of data they were trained on, the task of determining whether the AI prediction is trustworthy falls onto the physician.
“This extra burden on physicians can lead to reduced efficiency, misdiagnoses, and ultimately hesitancy to use AI tools,” said Zhongnan Fang, PhD, Principal Machine Learning Scientist and lead author of the study. “EMM is like a real-time quality check that tells you if you should be confident about what the AI tool is telling you.”

