AI-Assisted Pediatric Fracture Detection and Diagnostic Revisions

Key Takeaways
- Among children and adolescents undergoing out-of-hours appendicular radiography in a tertiary pediatric emergency setting, next-day recall-triggering diagnostic revisions were less frequent with AI support but not significantly different.
- The fracture-detection system showed high diagnostic accuracy against the reference standard.
- Therapy-altering corrections were rare and numerically lower with AI, while senior consultation, subjective diagnostic confidence, and emergency department length of stay did not differ significantly.
- Most families viewed AI positively as a clinician-support tool, even as anxiety about AI use in medicine remained common.
Diagnostic revisions requiring next-day recall occurred in 35/405 (8.6%) examinations without AI and 15/262 (5.7%) with AI, for a risk ratio (RR) of 0.66 (95% confidence interval (CI) 0.37–1.19; p=0.24). Therapy-altering corrections occurred in 8/405 (2%) examinations without AI and 1/262 (0.4%) with AI, with RR 0.19 (95% CI 0.02–1.53; p=0.10). The direction favored AI, but the main clinical endpoints were not statistically different.
Against the reference diagnosis, AI achieved 95.1% accuracy, 93% sensitivity, 95.1% specificity, 92.2% positive predictive value (PPV), and 95.6% negative predictive value (NPV). Senior consultation, subjective diagnostic confidence, and emergency department length of stay did not differ significantly between study conditions. Among 650 responding families, 82% viewed AI positively as a supportive tool, 64% believed it could improve care, 13% thought it could replace physicians in the medium term, and 61% reported anxiety about AI use in medicine. Exploratory Bayesian and age- and sex-adjusted analyses pointed in the same direction as the primary analysis but remained inconclusive.
This exploratory study was not primarily powered to detect small differences in recall rates. Unequal group sizes were influenced by intermittent hospital infrastructure failures that prevented AI activation on about 16 study days, and those examinations were analyzed in the non-AI group. In this non-U.S. high-performing academic setting, daily multidisciplinary review and the fact that only 9 diagnostic revisions led to a relevant therapeutic change likely constrained measurable effect size. Workload and reporting-time effects were not evaluated, and the findings may not generalize to non-academic or resource-limited settings.
Clinician Questions
How was a recall-triggering diagnostic revision defined in pediatric fracture AI evaluation?
In pediatric appendicular fracture imaging, a recall-triggering diagnostic revision was a discrepancy between the initial emergency department interpretation and the next-day reference diagnosis that required correction with patient recall the following day. The reference diagnosis came from consensus review by two senior pediatric radiologists and was supplemented by follow-up imaging when available.
What counted as a therapy-altering correction after overnight pediatric fracture radiography?
For overnight pediatric appendicular fracture radiography, a therapy-altering correction meant a relevant therapeutic change consisting of surgical intervention, initiation or discontinuation of medication other than analgesics, or modification of cast or splint immobilization lasting longer than one week.
How was fracture-detection AI integrated into the overnight pediatric emergency workflow?
In the overnight pediatric emergency workflow, the commercial Milvue/TechCare Kids system was available on alternating study days and operated as an on-premises decision-support tool within PACS. It categorized images as definite fracture, possible fracture, or no fracture and returned annotated images in real time while junior pediatric surgery residents retained diagnostic and management autonomy, with no automated decisions made without physician oversight.