Comparative Analysis of Anesthesiologists Judgment and ChatGPT in Predicting Intra Operative and Emergency Related Anesthesia Complications
Main Article Content
Abstract
Background: Artificial intelligence and large language models are increasingly being investigated for perioperative decision support, but their agreement with anesthesiologist judgment in predicting anesthesia-related complications remains insufficiently characterized. Objective: To compare anesthesiologist and ChatGPT predictions of intraoperative and emergency-related anesthesia complications using the same clinical information. Methods: This observational comparative study included 140 patients undergoing surgical procedures at Lady Reading Hospital and Khyber Teaching Hospital, Peshawar. Anesthesiologists and ChatGPT independently assessed standardized anonymized clinical information. Inter-method associations were assessed using chi-square analysis, and agreement was quantified using Cohen's kappa. Results: Of 140 participants, 73 (52.1%) were aged 41–70 years, 88 (62.9%) were male, and 124 (88.6%) underwent elective surgery. Agreement was perfect for bradycardia and hypo-/hyperthermia (κ=1.000 each), and remained high for hypotension (κ=0.959), bleeding (κ=0.867), hypertension (κ=0.826), and hypoventilation/hypoxia (κ=0.793). Lower concordance was observed for shivering (κ=0.659) and tachycardia (κ=0.646). All analyzable inter-method associations had p<0.001. Conclusion: ChatGPT classifications showed substantial concordance with anesthesiologist judgments for several common anesthesia-related complications. These findings demonstrate inter-method agreement but do not independently establish predictive accuracy against actual clinical outcomes
Article Details
Issue
Section

This work is licensed under a Creative Commons Attribution 4.0 International License.
How to Cite
References
1. Irita K, Tsuzaki K, Sanuki M, Sawa T, Nakatsuka H, Makita K, et al. Recent changes in the incidence of life-threatening events in the operating room: JSA surveys between 2001 and 2005. Masui. 2007;56:1433-46.
2. Cook TM, Woodall N, Frerk C; Fourth National Audit Project. Major complications of airway management in the UK: results of the Fourth National Audit Project of the Royal College of Anaesthetists and the Difficult Airway Society. Part 1: Anaesthesia. Br J Anaesth. 2011;106(5):617-31. doi:10.1093/bja/aer058.
3. Almghairbi DS, Marufu TC, Moppett IK. Anaesthesia workload measurement devices: qualitative systematic review. BMJ Simul Technol Enhanc Learn. 2018;4(3):112-6.
4. Yoon HK, Yang HL, Jung CW, Lee HC. Artificial intelligence in perioperative medicine: a narrative review. Korean J Anesthesiol. 2022;75(3):202-15. doi:10.4097/kja.22157.
5. Jo YY, Jang JH, Kwon JM, Lee HC, Jung CW, Byun S, et al. Predicting intraoperative hypotension using deep learning with waveforms of arterial blood pressure, electroencephalogram, and electrocardiogram: retrospective study. PLoS One. 2022;17(8). doi:10.1371/journal.pone.0272055.
6. Daccache N, Zako J, Morisson L, Laferrière-Langlois P. The applications of ChatGPT and other large language models in anesthesiology and critical care: a systematic review. Can J Anaesth. 2025;72(6):904-22. doi:10.1007/s12630-025-02973-9.
7. Barroso A, Casans R. Application of generative artificial intelligence chatbots in the field of anesthesia. Rev Esp Anestesiol Reanim. 2025;72:501667.
8. Angel MC, Rinehart JB, Cannesson MP, Baldi P. Clinical knowledge and reasoning abilities of AI large language models in anesthesiology: a comparative study on the American Board of Anesthesiology examination. Anesth Analg. 2024;139(2):349-56. doi:10.1213/ANE.0000000000006892.
9. Cheng T, Li Y, Gu J, He Y, He G, Zhou P, et al. The performance of ChatGPT in day surgery and pre-anesthesia risk assessment: a case-control study of 150 simulated patient presentations. Perioper Med (Lond). 2024;13:111. doi:10.1186/s13741-024-00469-6.
10. Wang B, Tian Y, Wang XT. An exploratory comparison of AI models for preoperative anesthesia planning. J Med Syst. 2025;49:104. doi:10.1007/s10916-025-02243-7.
11. Çamkıran V, Tunç H, Achmar B, Ürker TS, Kutlu İ, Torun A. Artificial intelligence (ChatGPT) ready to evaluate ECG in real life? Not yet! Digit Health. 2025;11:20552076251325279. doi:10.1177/20552076251325279.
12. Wang C, Liu S, Yang H, Guo J, Wu Y, Liu J. Ethical considerations of using ChatGPT in health care. J Med Internet Res. 2023;25. doi:10.2196/48009.
13. Gupta B, Ahluwalia P, Gupta A, Mahaseth R. ChatGPT in anesthesiology practice: a friend or a foe. Saudi J Anaesth. 2024;18(2):150-3. doi:10.4103/sja.sja_336_23.
14. Patel S, Ngo V, Wilhelmi B. Evaluating large language models on American Board of Anesthesiology-style anesthesiology questions: accuracy, domain consistency, and clinical implications. J Cardiothorac Vasc Anesth. 2025;39(9):2511-5. doi:10.1053/j.jvca.2025.05.033.
15. Sallam M, Barakat M, Sallam M. Pilot testing of a tool to standardize the assessment of the quality of health information generated by artificial intelligence-based models. Cureus. 2023;15(11). doi:10.7759/cureus.49373.
16. Brügge E, Ricchizzi S, Arenbeck M, Keller MN, Schur L, Stummer W, et al. Large language models improve clinical decision making of medical students through patient simulation and structured feedback: a randomized controlled trial. BMC Med Educ. 2024;24:1391. doi:10.1186/s12909-024-06399-7.