Industry News · Artificial Intelligence
Innovative AI System Developed to Identify Authentic Radiology Reports
March 24, 2026 · News Release

Researchers from the University at Buffalo (UB) have developed an advanced artificial intelligence (AI) system that can identify discrepancies between radiology reports composed by clinicians and those generated by AI. Nalini Ratha, the lead investigator and SUNY Empire Innovation Professor in the Department of Computer Science and Engineering, underscores the importance of such detection tools in health care, where the stakes of fraudulent reports are far greater.
The misuse of AI to fabricate radiology reports could potentially facilitate insurance fraud and other malicious activities. In response, the UB research team, led by Ratha and including PhD students Arjun Ramesh Kaushik and Tanvi Ranga, introduced their study titled “Detecting Synthetic Radiology Reports Using Style Disentanglement” at the GenAI4Health workshop. This took place during the 2025 Conference on Neural Information Processing Systems.
The research emphasized the creation of a pioneering dataset comprising 14,000 pairs of chest x-ray reports, comparing reports authored by radiologists to those generated by AI. The synthetic reports were created through two primary methods: paraphrasing existing radiologist reports using large language models (LLMs) and generating reports from chest radiographs through vision-language models (VLMs). The findings section, a critical component of any radiology report, was central to the study’s focus. It contains the radiologist’s detailed analysis and specific terminology, making it particularly vulnerable to exploitation.
UB’s detection framework, based on a BERT-Mamba model, processes this dataset. It is designed to recognize the stylistic differences between human-written and AI-generated reports. Although LLMs replicate clinical terminology, they fall short in mimicking the nuanced stylistic attributes of reports created by radiologists. The detection model successfully identified AI-generated reports with a Matthews correlation coefficient (MCC) score ranging from 92% to 100%.
The study also highlights subtle stylistic differences between AI and clinician-authored reports, with AI often using more verbose language. For instance, AI might use terms such as "pulmonary vasculature" in place of simpler terms such as "heart" or "lung." These discrepancies serve as signals for the model to distinguish between authentic and fabricated reports.
Looking ahead, the researchers intend to enhance the AI model, expand the dataset to include more categories, and tailor it for public use. While the immediate focus is on radiology, Ratha suggests that the methodology could be beneficial across various sectors, such as insurance, journalism, and law, where the authenticity of written records is paramount.





