AI-Assisted Anesthesia Decisions Show Mixed Concordance Across 6 Centers

Key Takeaways
- Overall concordance during anesthesia maintenance was 73.3%, with anesthesiologists' decisions serving as the reference standard.
- Agreement was strongest for propofol dosage adjustments and was lower across hemodynamic medication decisions.
- The AI system reached propofol decisions faster, and the authors said additional optimization and validation remain necessary.
Data came from six medical centers in China between March and December 2024, and the cohort included 619 women with a median age of 50 years. All data were drawn from existing perioperative records rather than prospective bedside testing. Eligible patients had ASA physical status I to III and underwent elective noncardiac procedures under total intravenous anesthesia. Anesthesiologists' decisions were the reference standard, and concordance required a matching intervention or recommendation within a 300-second temporal window. Agreement was quantified across decision categories with percentage agreement, prevalence-adjusted and bias-adjusted kappa, and Gwet's AC1. Propofol was managed as a continuous infusion, whereas vasoactive drugs were single-dose boluses, making this an interrater-style comparison of medication adjustments rather than outcomes.
Within that framework, investigators identified 6060 concordant decisions, with PABAK of 0.467, 95% CI 0.448-0.486, and AC1 of 0.653, 95% CI 0.638-0.669. Both agreement coefficients were significantly different from zero, with P values below .001 for each measure. Propofol showed the strongest alignment, reaching 91.1% dosage-adjustment concordance, 95% CI 90.4%-91.8%, and 68.6% direction concordance, 95% CI 67.3%-69.8%. This split indicated closer agreement on whether a propofol change was needed than on the direction of that change. The remaining medication categories showed considerably less agreement, especially in hemodynamic management decisions during maintenance. The overall 73.3% concordance was heavily influenced by propofol; excluding propofol, concordance for the remaining vasoactive medications was 26.3%.
For propofol, the AI system reached decisions sooner than anesthesiologists, with a pseudomedian latency difference of -77.5 seconds, 95% CI -79.5 to -75.5, P<.001. Agreement for hemodynamic drugs was lower overall, with esmolol reaching a numerically higher 71.4% concordance that was not statistically significant and atropine, ephedrine, and urapidil ranging from 17.4% to 29.8%. Researchers also observed that the AI recommended hemodynamic interventions about 1.74 to 3.64 times more often than anesthesiologists, while performance remained generally consistent across centers.
The authors cautioned that the evaluation was retrospective and simulation-based, did not assess adverse events, hemodynamic stability, or drug consumption, and did not fully capture true negatives. They described ZW-AA-001 as decision support rather than autonomous decision-making and said further optimization and prospective validation are still needed.