Applied Radiology

Industry News · Artificial Intelligence · CT

LLMs Could Automate CT Protocoling with High Accuracy, Study Finds

January 6, 2026 · News Release

LLMs Could Automate CT Protocoling with High Accuracy, Study Finds

Large language models may soon help radiologists streamline one of the more tedious and time-consuming tasks in clinical imaging: selecting the appropriate protocol for CT exams. With the right context, a new study suggests, models like GPT-4o can outperform human readers in matching abdominal and pelvic CT scans to the most suitable imaging protocol.

The findings, published in Radiology, show that prompt engineering—specifically, supplying the model with relevant contextual information—significantly enhances performance. Researchers used GPT-4o to assign protocols for all abdomen and pelvis CT scans performed at their institution over a six-month period in 2024, comparing its results to those selected by radiologists, residents and fellows.

“Accuracy is critical, as incorrect protocols can lead to nondiagnostic examinations, repeat imaging, and delayed diagnoses,” said lead author Rajesh Bhayana, MD, of University Medical Imaging Toronto. “However, protocoling is a manual and time-consuming task at most institutions that accounts for up to 6% of radiologists’ clinical time, competing with core interpretive responsibilities. Protocoling is also a recognized source of interruptions for radiologists, which can lead to increased diagnostic errors.”

To test how well large language models could handle this task, the team fed GPT-4o structured prompts that included clinical context such as patient medical history, body mass index, laboratory data, referring provider notes, and scanner-specific information. The model was asked to choose from among 46 institutional CT protocols, each with detailed selection criteria.

Using this approach, the model achieved an overall accuracy of 96.2%, well above the 88.3% accuracy rate recorded by human providers. Interestingly, there was no observed difference in the rate of inappropriate protocol selections between the AI and its human counterparts. Radiologists, fellows and residents all performed similarly in terms of matching the reference standard.

The authors noted that GPT-4o’s performance relied heavily on detailed and structured context engineering, a strategy increasingly recognized as essential for adapting LLMs to clinical use.

“Our results suggest that for protocoling, state-of-the-art LLMs can be efficiently adapted with detailed prompt instructions to select optimal protocols more frequently than standard-of-care manual protocoling, without an increase in inappropriate protocols,” the team wrote. “Thus, LLMs could efficiently enable a more widespread use of automated protocol selection in supervised settings.”

The study adds to a growing body of evidence that generative AI, when deployed with attention to clinical nuance, could ease administrative burdens on radiologists without compromising accuracy. While the authors emphasized that LLM-based protocoling should be supervised, their findings hint at a future where such tools could become a standard part of radiology workflows.

More in Industry News