epocrates logo
epocrates logo
epocrates logo
  • 0

Journal Article Synopsis

Science

Generalist AI helped radiologists catch 10% more on abdominal CT

September 18, 2026

card-image

Clinical takeaway: AI assistance raised radiologists' sensitivity across 146 abdominal CT findings without a meaningful specificity cost. No generalist model of this kind yet holds FDA clearance. 

An abdominal CT asks more of its reader than almost any other imaging study. The read spans dozens of organs and hundreds of possible diagnoses, many of them low-contrast findings that an experienced eye catches, but a less practiced one may miss. That expertise concentrates in urban and academic centers. There is little room for error, which might result in missed cancers or misread acute scans under emergency pressure. 

Radiology holds more FDA-cleared AI than any other specialty, with more than 1,100 algorithms. But nearly all are assigned to perform a single task. A model that reads the whole study has been the field's long-stated ambition. Expert annotation cannot label training data at that scale, and models that learned from radiology reports instead covered a few dozen diagnoses at accuracy well short of expert. A new study reports a generalist vision-language model for abdominal CT that closes much of that distance. 

The model, dubbed RADAR, was tested as a second reader: with its predicted findings and attention maps in view, radiologists' diagnostic sensitivity rose 10.0% while specificity slipped from 98.8% to 98.2%, and reading time per case fell 30.7%. Assisted junior radiologists at top-tier hospitals performed comparably to unassisted seniors at the same hospitals, and juniors at other hospitals raised their sensitivity 10.1%, drawing level with unassisted seniors in their own tier. Detection improved where misses cost the most: sensitivity for malignancies rose 7.3%, for emergency conditions 8.5%, and for hollow-organ disease 9.4%. 

Reading on its own, RADAR scored 91.3% on average by the study's accuracy measure across the 146 findings in 18 organ systems. The best of three prior vision-language models, retrained by the authors on their own dataset for comparison, managed 77.6%. That accuracy held at 89.5% across eight external centers. It stayed at 90.4% on 17 acute abdominal conditions excluded from its training. Against pathology-confirmed cancers of the liver, pancreas, stomach, and colorectum, RADAR scored 89.1% to 98.4%. 

The lone US test came against Merlin, a Stanford-built abdominal CT model with a public benchmark drawn from Stanford Hospital patients. RADAR scored 88.3% on 21 of that benchmark's 30 findings without any tuning, exceeding Merlin's performance on Merlin's own data. 

RADAR was trained on 424,911 contrast-enhanced abdominal CT examinations and their radiology reports, collected from 2010 through 2023 at a single Chinese academic center, with no manual annotation. Testing ran on 39,160 later examinations at that center, another 24,239 across eight external Chinese hospitals, and two pathology-confirmed cohorts; every validation site except the 5,137-scan Stanford benchmark was in China. The reader study put 300 internal cases spanning 61 findings before 26 radiologists from 14 institutions, who read each case unassisted and again with RADAR after a washout of at least a month. 

The authors have released RADAR's source code and model checkpoints for public use. No generative vision-language model for diagnostic image interpretation has received FDA marketing authorization. In August, Mosaic Clinical Technologies, the technology services division of Radiology Partners, petitioned the agency to clarify when commercially distributed vision-language models for diagnostic image analysis are medical devices. FDA's response could shape the timing and pathway for marketing models such as RADAR for US diagnostic use. 

"This work suggests that large-scale vision-language learning may provide a practical path toward efficiently developing generalist AI systems for complex and diverse radiologic interpretation tasks," the authors conclude. 

Source: Zhang Q, et al. (2026 Sep 17) Science. An expert-level generalist AI for abdominal CT diagnosis 

 

learn more about epocrates plus

Clinical FAQs

Check out the answers to frequently asked questions about our clinical content.

Download Epocrates from the App StoreDownload Epocrates from the Play Store
About UsFeaturesBusiness SolutionsHelp & FeedbackCookie Preferences
© 2026 epocrates, Inc.   Terms of UsePrivacy PolicyEditorial PolicyDo Not Sell or Share My Information