Industry News · Artificial Intelligence · Thoracic Imaging
AI Algorithm Achieves Clinician Accuracy with Limited Data in Lung Nodule Risk Assessment
July 8, 2026 · News Release

A recent study published in Radiology: Artificial Intelligence demonstrates that an AI-based algorithm can accurately estimate the malignancy risk of lung nodules using significantly less training data than traditionally required. The research shows that the model achieved clinician-level performance using just 20% of the usual data input.
The study, conducted by a team led by Bogdan Obreja, MSc, a PhD candidate at Radboud University Medical Center in the Netherlands, has significant implications for the development of AI tools in medicine, particularly for lung cancer screening. Lung nodules detected through low-dose CT screening often do not progress to cancer. An efficient AI tool could help prioritize patients for follow-up interventions while reducing unnecessary procedures.
"A key consideration when developing these algorithms is the amount of data needed to effectively train the system while maintaining robust performance," Obreja stated. The research team explored whether a reduction in data volume would affect the accuracy of their AI model, specifically developed to assess risk in pulmonary nodules from lung cancer screening CT.
Utilizing 16,077 annotated nodules from the National Lung Screening Trial (NLST), the model's performance was tested externally with data from the Danish Lung Cancer Screening Trial. Remarkably, the algorithm matched clinician accuracy even when trained with just 20% of the available data. This finding challenges the common assumption that larger datasets invariably result in better AI performance.
The unique capabilities of the model were largely due to its advanced architecture, which integrates both 2D and 3D convolutional neural networks. This allows for comprehensive analysis by capturing intricate slice-level and volumetric details. "Combining 2D and 3D models improves overall performance and enhances robustness, reducing variability across predictions," Obreja noted.
Researchers stress that this breakthrough presents new opportunities for smaller teams and organizations with limited access to vast datasets. Instead of focusing on sheer volume, efforts could be better allocated to collecting diverse datasets, including varied demographics and medical cases. "Resources may be better spent on collecting and annotating diverse datasets, such as scans from different populations, scanners, and rare or challenging cases," Obreja added.
Study senior author Colin Jacobs, PhD, emphasized the importance of diversity over volume, saying, "We believe that beyond a certain point, gains in performance from additional data, especially if they consist mainly of common cases, become increasingly limited." However, including data from atypical or rare conditions can still significantly enhance AI models.
The accompanying commentary by Kellie J. Archer, PhD, at the Ohio State University, highlights the generalizability of the model's results due to its quality and variability. Though the training dataset did not fully reflect U.S. racial demographics, it included diverse representation across major groups and achieved gender balance. Archer also stressed the necessity of reflecting clinical applications' full variability in training data to optimize AI performance.
As the study indicates promising results for limited-data AI models, future research will be essential in validating this approach across different populations and exploring diversification of datasets. For more information, the full study and related commentary are accessible in Radiology: Artificial Intelligence.





