Debate on AI’s Role in Chest X-Ray Interpretation at RSNA 2025
February 9, 2026 · News Release
A recent session at the RSNA 2025 conference highlighted a significant debate regarding the role of artificial intelligence (AI) in interpreting chest X-rays (CXR). As AI technology rapidly advances, questions arise about its readiness to autonomously interpret these images or if human oversight remains essential to ensure safety and accuracy in diagnostics.
An informal poll by the RSNA Daily Bulletin revealed that 68.4% of participants share the view of Dr. Warren Gefter from Penn Medicine. Dr. Gefter argues that current AI models lack the necessary accuracy, reliability, and comprehensive capability for independent interpretation of CXRs. In contrast, 31.6% support Dr. Saurabh Jha of the Hospital of the University of Pennsylvania. Dr. Jha contends that CXRs no longer require medical degrees for interpretation, a task AI is competent to handle independently.
Dr. Eun Kyoung (Amy) Hong from Stanford Medicine provided insights from her experience with AI in chest X-ray tasks. She delineated between conventional vision-based AI systems and generative AI models. While vision-based models aid in classifying and detecting abnormalities, generative models, trained on image-data pairs, create preliminary reports for radiologist review.
Dr. Hong highlighted the issue of "hallucinations" in AI-generated reports, defined as statements not supported by the input given to the AI model. She noted that these unsupported statements appear in 15-20% of reports. Despite this challenge, Dr. Hong believes that technological advancements will soon address this problem.
Generative AI has shown potential to enhance efficiency. A large trial involving nearly 12,000 CXRs demonstrated that an AI assistant reduced interpretation time by approximately 15%, from 189 to 160 seconds. However, this study did not examine the feasibility of AI operating without human oversight.
When testing different generative AI models, Dr. Hong found variability in their outputs, with acceptability ranging from 28% to 67% and hallucinations from 5% to 55%. This variability underscores the importance of rigorous model evaluation before incorporating AI for autonomous CXR interpretation. Dr. Hong also pointed out issues such as automation bias and a lack of legal frameworks for accountability in AI errors.
In conclusion, Dr. Hong suggests that AI serves best as an adjunct tool rather than as a replacement for human expertise in CXR interpretation. While AI can aid in drafting reports and enhancing efficiency, it remains inadequate for independent operation until it can provide consistent results and assume responsibility for its analyses.


