Transactions on Machine Intelligence

Transactions on Machine Intelligence

Evidence-Gated Autonomy in Healthcare AI Agents: A Systematic Review of Action Scope, Human-Oversight Effectiveness, and Multi-Step Safety Failures

Document Type : Original Article

Authors
1 Associate Professor of Biomedical Engineering, Vali-e-Asr University of Rafsanjan
2 Department of Biomedical Engineering, Meybod University, Meybod, Iran
3 Department of Engineering.Vali-e-Asr University of Rafsanjan,Iran
4 Assistant Professor,Faculty of Electrical and Computer Engineering, Shams Gonbad Higher Education Institute, Gorgan, Iran
Abstract
Background: Healthcare AI agents increasingly perform multi-step workflows and tool use, but evidence on their operational autonomy, safety, and human oversight remains limited. This review examines whether greater autonomy is supported by empirical safety and oversight evidence.
Methods: PubMed/MEDLINE, Scopus, and Web of Science were searched from January 2022 to July 2026. Of 4,096 records, 2,557 remained after deduplication. Fifty priority reports were selected for full-text review; 34 were assessed and 28 met the inclusion criteria. Data on workflow, architecture, autonomy, safety, oversight, and evaluation were narratively synthesized.
Results: Among 28 studies, 18 were published in 2026. Thirteen addressed diagnostic or clinical decision support, nine EHR or administrative workflows, four treatment planning or prescribing, and two patient-facing or cross-domain applications. Autonomy levels were L0 in five studies, L1 in 13, L2 in three, L3 in one, and L4 in six; none reached L5. Most L4 systems were evaluated only in sandbox or retrospective settings. Only three studies empirically assessed oversight effectiveness.
Conclusions: Current evidence supports tool-using and workflow-capable healthcare agents, but not unrestricted clinical autonomy. Greater autonomy should require stronger evidence of safety, failure containment, and effective human oversight.
Keywords

[1]     Topol, E. J. (2019). High-performance medicine: The convergence of human and artificial intelligence. Nature Medicine, 25(1), 44–56. https://doi.org/10.1038/s41591-018-0300-7
[2]     Rajpurkar, P., Chen, E., Banerjee, O., & Topol, E. J. (2022). AI in health and medicine. Nature Medicine, 28(1), 31–38. https://doi.org/10.1038/s41591-021-01614-0
[3]     Liu, X., Cruz Rivera, S., Moher, D., Calvert, M. J., Denniston, A. K., & SPIRIT-AI and CONSORT-AI Working Group. (2020). Reporting guidelines for clinical trial reports for interventions involving artificial intelligence: The CONSORT-AI extension. Nature Medicine, 26(9), 1364–1374. https://doi.org/10.1038/s41591-020-1034-x
[4]     van de Sande, D., et al. (2022). Developing, implementing and governing artificial intelligence in medicine: A step-by-step approach to prevent an artificial intelligence winter. BMJ Health Care Informatics, 29(1), e100495. https://doi.org/10.1136/bmjhci-2021-100495
[5]     Singhal, K., Azizi, S., Tu, T., et al. (2023). Large language models encode clinical knowledge. Nature, 620(7972), 172–180. https://doi.org/10.1038/s41586-023-06291-2
[6]     Nori, H., King, N., McKinney, S. M., Carignan, D., & Horvitz, E. (2023). Capabilities of GPT-4 on medical challenge problems [Preprint]. arXiv. https://arxiv.org/abs/2303.13375
[7]     Moor, M., Banerjee, O., Abad, Z. S. H., et al. (2023). Foundation models for generalist medical artificial intelligence. Nature, 616(7956), 259–265. https://doi.org/10.1038/s41586-023-05881-4
[8]     Zhao, L., Liu, S., Xin, T., et al. (2026). AI agent in healthcare: Applications, evaluations, and future directions. npj Artificial Intelligence, 2, 31. https://doi.org/10.1038/s44387-026-00076-4
[9]     Xu, X., & Sankar, R. (2025). Large language model agents for biomedicine: A comprehensive review of methods, evaluations, challenges, and future directions. Information, 16(10), 894. https://doi.org/10.3390/info16100894
[10]   Masterman, T., Besen, S., Sawtell, M., & Chao, A. (2024). The landscape of emerging AI agent architectures for reasoning, planning, and tool calling: A survey [Preprint]. arXiv. https://arxiv.org/abs/2404.11584
[11]   Tierney, A. A., Gayre, G., Hoberman, B., et al. (2025). Ambient artificial intelligence scribes: Learnings after 1 year and over 2.5 million uses. NEJM Catalyst Innovations in Care Delivery, 6(5), CAT.25.0040. https://doi.org/10.1056/CAT.25.0040
[12]   Lukac, P. J., Turner, W., Vangala, S., et al. (2025). Ambient AI scribes in clinical practice: A randomized trial. NEJM AI, 2(12), AIoa2501000. https://doi.org/10.1056/AIoa2501000
[13]   Olson, K. D., Meeker, D., Troup, M., et al. (2025). Use of ambient AI scribes to reduce administrative burden and professional burnout. JAMA Network Open, 8(10), e2534976. https://doi.org/10.1001/jamanetworkopen.2025.34976
[14]   Shi, W., et al. (2024). EHRAgent: Code empowers large language models for few-shot complex tabular reasoning on electronic health records. In Proceedings of EMNLP 2024 (pp. 22315–22339). https://doi.org/10.18653/v1/2024.emnlp-main.1245
[15]   Zakka, C., Cho, J., Fahed, G., et al. (2025). Almanac Copilot: Towards autonomous electronic health record navigation [Preprint]. Research Square. https://doi.org/10.21203/rs.3.rs-6102516/v1
[16]   Gebreab, S. A., et al. (2024). LLM-based framework for administrative task automation in healthcare. In Proceedings of the 12th International Symposium on Digital Forensics and Security (ISDFS) (pp. 1–7). IEEE. https://doi.org/10.1109/ISDFS60797.2024.10527275
[17]   Pandey, H. G., Amod, A., & Kumar, S. (2024). Advancing healthcare automation: Multi-agent system for medical necessity justification. Proceedings of the BioNLP Workshop. https://doi.org/10.18653/v1/2024.bionlp-1.4
[18]   Tu, T., Palepu, A., Schaekermann, M., et al. (2025). Towards conversational diagnostic artificial intelligence. Nature. https://doi.org/10.1038/s41586-025-08866-7
[19]   Ferber, D., El Nahhas, O. S. M., Wölflein, G., et al. (2025). Development and validation of an autonomous artificial intelligence agent for clinical decision-making in oncology. Nature Cancer, 6(8), 1337–1349. https://doi.org/10.1038/s43018-025-00991-6
[20]   Kim, Y., et al. (2024). MDAgents: An adaptive collaboration of LLMs for medical decision-making. In Advances in Neural Information Processing Systems (NeurIPS). https://doi.org/10.52202/079017-2522
[21]   Wang, Z., et al. (2025). ColaCare: Enhancing electronic health record modeling through large language model-driven multi-agent collaboration. In Proceedings of the ACM Web Conference 2025 (pp. 2250–2261). https://doi.org/10.1145/3696410.3714877
[22]   Li, J., et al. (2024). Agent Hospital: A simulacrum of hospital with evolvable medical agents [Preprint]. arXiv. https://arxiv.org/abs/2405.02957
[23]   Nie, M., Chung, W., Waxler, J., et al. (2026). Hard to halt: Automation bias in agent-driven sequencing prior authorization workflows [Preprint]. medRxiv. https://doi.org/10.64898/2026.06.16.26355782
[24]   Yang, G. Z., Cambias, J., Cleary, K., et al. (2017). Medical robotics—Regulatory, ethical, and legal considerations for increasing levels of autonomy. Science Robotics, 2(4), eaam8638. https://doi.org/10.1126/scirobotics.aam8638
[25]   European Parliament & Council of the European Union. (2024). Regulation (EU) 2024/1689 laying down harmonised rules on artificial intelligence (Artificial Intelligence Act). Official Journal of the European Union.
[26]   U.S. Food and Drug Administration. (2025). Artificial intelligence-enabled device software functions: Lifecycle management and marketing submission recommendations—Draft guidance. U.S. Food and Drug Administration.
[27]   Amodei, D., Olah, C., Steinhardt, J., Christiano, P., Schulman, J., & Mane, D. (2016). Concrete problems in AI safety [Preprint]. arXiv. https://arxiv.org/abs/1606.06565
[28]   Cemri, M., et al. (2025). Why do multi-agent LLM systems fail? [Preprint]. arXiv. https://doi.org/10.52202/085713-4082
[29]   Jiang, Y., Black, K. C., Geng, G., et al. (2025). MedAgentBench: A virtual EHR environment to benchmark medical LLM agents. NEJM AI, 2(9), AIdbp2500144. https://doi.org/10.1056/AIdbp2500144
[30]   Alkaissi, H., & McFarlane, S. I. (2023). Artificial hallucinations in ChatGPT: Implications in scientific writing. Cureus, 15(2), e35179. https://doi.org/10.7759/cureus.35179
[31]   Goddard, K., Roudsari, A., & Wyatt, J. C. (2012). Automation bias: A systematic review of frequency, effect mediators, and mitigators. Journal of the American Medical Informatics Association, 19(1), 121–127. https://doi.org/10.1136/amiajnl-2011-000089
[32]   Parasuraman, R., & Manzey, D. H. (2010). Complacency and bias in human use of automation: An attentional integration. Human Factors, 52(3), 381–410. https://doi.org/10.1177/0018720810376055
[33]   Ancker, J. S., Edwards, A., Nosal, S., Hauser, D., Mauer, E., & Kaushal, R. (2017). Effects of workload, work complexity, and repeated alerts on alert fatigue in a clinical decision support system. BMC Medical Informatics and Decision Making, 17(1), 36. https://doi.org/10.1186/s12911-017-0430-8
[34]   Guo, Y., Hu, D., Zhou, Y., et al. (2026). From conversation to chart: An analysis of clinician edits to ambient AI draft notes [Preprint]. medRxiv. https://doi.org/10.64898/2026.01.05.26343471
[35]   Price, W. N., II, Gerke, S., & Cohen, I. G. (2019). Potential liability for physicians using artificial intelligence. JAMA, 322(18), 1765–1766. https://doi.org/10.1001/jama.2019.15064
[36]   Habli, I., Lawton, T., & Porter, Z. (2020). Artificial intelligence in health care: Accountability and safety. Bulletin of the World Health Organization, 98(4), 251–256. https://doi.org/10.2471/BLT.19.237487
[37]   Gerke, S., Minssen, T., & Cohen, G. (2020). Ethical and legal challenges of artificial intelligence-driven healthcare. In Artificial intelligence in healthcare (pp. 295–336). Academic Press. https://doi.org/10.1016/B978-0-12-818438-7.00012-5
[38]   World Health Organization. (2024). Ethics and governance of artificial intelligence for health: Guidance on large multi-modal models. World Health Organization.
[39]   Gilbert, S., Harvey, H., Melvin, T., Vollebregt, E., & Wicks, P. (2023). Large language model AI chatbots require approval as medical devices. Nature Medicine, 29(10), 2396–2398. https://doi.org/10.1038/s41591-023-02412-6
[40]   Thirunavukarasu, A. J., Ting, D. S. J., Elangovan, K., Gutierrez, L., Tan, T. F., & Ting, D. S. W. (2023). Large language models in medicine. Nature Medicine, 29(8), 1930–1940. https://doi.org/10.1038/s41591-023-02448-8
[41]   Collaco, B. G., Haider, S. A., Prabha, S., et al. (2026). The role of agentic artificial intelligence in healthcare: A scoping review. npj Digital Medicine, 9, 345. https://doi.org/10.1038/s41746-026-02517-5
[42]   Njei, B., Al-Ajlouni, Y. A., Kanmounye, U. S., Boateng, S., Nguefang, G. L., Njei, N., et al. (2026). Artificial intelligence agents in healthcare research: A scoping review. PLOS ONE, 21(2), e0342182. https://doi.org/10.1371/journal.pone.0342182
[43]   Gorenshtein, A., Omar, M., Glicksberg, B. S., Nadkarni, G. N., & Klang, E. (2025). AI agents in clinical medicine: A systematic review [Preprint]. medRxiv. https://doi.org/10.1101/2025.08.22.25334232
[44]   Xu, J., Ko, J. M., & Kvedar, J. C. (2026). AI agents in clinical practice: An evidence map. npj Digital Medicine. https://doi.org/10.1038/s41746-026-02960-4
[45]   SAE International. (2021). Taxonomy and definitions for terms related to driving automation systems for on-road motor vehicles (SAE Standard J3016_202104).
[46]   Wooldridge, M., & Jennings, N. R. (1995). Intelligent agents: Theory and practice. The Knowledge Engineering Review, 10(2), 115–152. https://doi.org/10.1017/S0269888900008122
[47]   Russell, S., & Norvig, P. (2021). Artificial intelligence: A modern approach (4th ed.). Pearson.
[48]   Xi, Z., Chen, W., Guo, X., et al. (2025). The rise and potential of large language model based agents: A survey. Science China Information Sciences, 68(2), 121101. https://doi.org/10.1007/s11432-024-4222-0
[49]   Wang, L., Ma, C., Feng, X., et al. (2024). A survey on large language model based autonomous agents. Frontiers of Computer Science, 18(6), 186345. https://doi.org/10.1007/s11704-024-40231-1
[50]   Sheridan, T. B., & Verplank, W. L. (1978). Human and computer control of undersea teleoperators. MIT Man-Machine Systems Laboratory. https://doi.org/10.21236/ADA057655
[51]   Parasuraman, R., Sheridan, T. B., & Wickens, C. D. (2000). A model for types and levels of human interaction with automation. IEEE Transactions on Systems, Man, and Cybernetics—Part A: Systems and Humans, 30(3), 286–297. https://doi.org/10.1109/3468.844354
[52]   Endsley, M. R. (2017). From here to autonomy: Lessons learned from human-automation research. Human Factors, 59(1), 5–27. https://doi.org/10.1177/0018720816681350
[53]   Santoni de Sio, F., & van den Hoven, J. (2018). Meaningful human control over autonomous systems: A philosophical account. Frontiers in Robotics and AI, 5, 15. https://doi.org/10.3389/frobt.2018.00015
[54]   Page, M. J., McKenzie, J. E., Bossuyt, P. M., et al. (2021). The PRISMA 2020 statement: An updated guideline for reporting systematic reviews. BMJ, 372, n71. https://doi.org/10.1136/bmj.n71
[55]   Rethlefsen, M. L., Kirtley, S., Waffenschmidt, S., et al., & PRISMA-S Group. (2021). PRISMA-S: An extension to the PRISMA Statement for reporting literature searches in systematic reviews. Systematic Reviews, 10(1), 39. https://doi.org/10.1186/s13643-020-01542-z
[56]   McGowan, J., Sampson, M., Salzwedel, D. M., Cogo, E., Foerster, V., & Lefebvre, C. (2016). PRESS peer review of electronic search strategies: 2015 guideline statement. Journal of Clinical Epidemiology, 75, 40–46. https://doi.org/10.1016/j.jclinepi.2016.01.021
[57]   Bramer, W. M., Giustini, D., de Jonge, G. B., Holland, L., & Bekhuis, T. (2016). De-duplication of database search results for systematic reviews in EndNote. Journal of the Medical Library Association, 104(3), 240–243. https://doi.org/10.3163/1536-5050.104.3.014
[58]   Ouzzani, M., Hammady, H., Fedorowicz, Z., & Elmagarmid, A. (2016). Rayyan—A web and mobile app for systematic reviews. Systematic Reviews, 5, 210. https://doi.org/10.1186/s13643-016-0384-4
[59]   McHugh, M. L. (2012). Interrater reliability: The kappa statistic. Biochemia Medica, 22(3), 276–282. https://doi.org/10.11613/BM.2012.031
[60]   Sterne, J. A. C., Savovic, J., Page, M. J., et al. (2019). RoB 2: A revised tool for assessing risk of bias in randomised trials. BMJ, 366, l4898. https://doi.org/10.1136/bmj.l4898
[61]   Sounderajah, V., Ashrafian, H., Rose, S., et al. (2021). A quality assessment tool for artificial intelligence-centered diagnostic test accuracy studies: QUADAS-AI. Nature Medicine, 27(10), 1663–1665. https://doi.org/10.1038/s41591-021-01517-0
[62]   Wolff, R. F., Moons, K. G. M., Riley, R. D., et al. (2019). PROBAST: A tool to assess the risk of bias and applicability of prediction model studies. Annals of Internal Medicine, 170(1), 51–58. https://doi.org/10.7326/M18-1376
[63]   Gallifant, J., Afshar, M., Ameen, S., et al. (2025). The TRIPOD-LLM reporting guideline for studies using large language models. Nature Medicine, 31(1), 60–69. https://doi.org/10.1038/s41591-024-03425-5
[64]   Vasey, B., Nagendran, M., Campbell, B., et al. (2022). Reporting guideline for the early-stage clinical evaluation of decision support systems driven by artificial intelligence: DECIDE-AI. Nature Medicine, 28(5), 924–933. https://doi.org/10.1038/s41591-022-01772-9
[65]   Campbell, M., McKenzie, J. E., Sowden, A., et al. (2020). Synthesis without meta-analysis (SWiM) in systematic reviews: Reporting guideline. BMJ, 368, l6890. https://doi.org/10.1136/bmj.l6890
[66]   Miake-Lye, I. M., Hempel, S., Shanman, R., & Shekelle, P. G. (2016). What is an evidence map? A systematic review of published evidence maps and their definitions, methods, and products. Systematic Reviews, 5, 28. https://doi.org/10.1186/s13643-016-0204-x
[67]   Cabitza, F., Rasoini, R., & Gensini, G. F. (2017). Unintended consequences of machine learning in medicine. JAMA, 318(6), 517–518. https://doi.org/10.1001/jama.2017.7797
[68]   Helmreich, R. L. (2000). On error management: Lessons from aviation. BMJ, 320(7237), 781–785. https://doi.org/10.1136/bmj.320.7237.781
[69]   Barach, P., & Small, S. D. (2000). Reporting and preventing medical mishaps: Lessons from non-medical near miss reporting systems. BMJ, 320(7237), 759–763. https://doi.org/10.1136/bmj.320.7237.759
[70]   Schmidgall, S., Ziaei, R., Harris, C., et al. (2026). AgentClinic: A multimodal benchmark for tool-using clinical AI agents. npj Digital Medicine. https://doi.org/10.1038/s41746-026-02674-7
[71]   Chen, E., Postelnik, S., Black, K., et al. (2026). MedAgentBench v2: Improving medical LLM agent design. In Pacific Symposium on Biocomputing. https://doi.org/10.1142/9789819824755_0025
[72]   Liu, Y., Carrero, Z. I., Jiang, X., et al. (2026). Benchmarking large language model-based agent systems for clinical decision tasks. npj Digital Medicine. https://doi.org/10.1038/s41746-026-02443-6
[73]   Mokssit, Y., Ravi, K., Nie, M., et al. (2026). FHIR-AgentEval: A modular sandbox for benchmarking clinical LLM agents with an evaluation of memory-augmented configurations [Preprint]. Research Square. https://doi.org/10.21203/rs.3.rs-8746188/v1
[74]   Hou, R., Xue, D., Sun, H., et al. (2026). CDAFlow: Enhancing LLM clinical decision-making through agentic workflow. Expert Systems with Applications. https://doi.org/10.1016/j.eswa.2026.131806
[75]   Xu, S., Huang, X., Wei, Z., et al. (2026). DxDirector: An agentic large language model driving the full-process clinical diagnosis. Nature Communications. https://doi.org/10.1038/s41467-026-71928-5
[76]   Han, S., & Choi, W. (2025). Development of a large language model-based multi-agent clinical decision support system for Korean Triage and Acuity Scale (KTAS)-based triage and treatment planning in emergency departments. Advances in Artificial Intelligence and Machine Learning. https://doi.org/10.54364/AAIML.2025.51187
[77]   Wu, X., Zhang, H., Garduno-Rapp, N. E., et al. (2026). Orchestrator multi-agent clinical decision support system for secondary headache diagnosis in primary care. Journal of the American Medical Informatics Association. https://doi.org/10.1093/jamia/ocag111
[78]   Klang, E., Omar, M., Raut, G., et al. (2026). Orchestrated multi agents sustain accuracy under clinical-scale workloads compared to a single agent. npj Health Systems. https://doi.org/10.1038/s44401-026-00077-0
[79]   Ucdal, M., & Ekingen, E. (2026). Performance comparison of a neuro-symbolic large language model system versus conventional AI models and human experts in cholangitis management. BMC Medical Informatics and Decision Making. https://doi.org/10.1186/s12911-026-03593-z
[80]   Ekingen, E., & Ucdal, M. (2026). Performance comparison of a neuro-symbolic large language model system versus human experts in acute cholecystitis management. Journal of Clinical Medicine, 15(5), 1730. https://doi.org/10.3390/jcm15051730
[81]   Zhai, G., Bar, M., Cowan, A. J., et al. (2025). AI for evidence-based treatment recommendation in oncology: A blinded evaluation of large language models and agentic workflows. Frontiers in Artificial Intelligence. https://doi.org/10.3389/frai.2025.1683322
[82]   Wang, J., Mullick Chowdhury, S., & Nazha, A. (2025). Virtual oncology collaborative tumor board using multiple artificial intelligence agents. Journal of Clinical Oncology, 43(16_suppl), 1563. https://doi.org/10.1200/JCO.2025.43.16_suppl.1563
[83]   Kochuiev, E., Kaliuzhka, V., Markevych, M., et al. (2026). An agentic AI framework for integrated decision support and surgical planning in intracerebral hemorrhage. Acta Neurochirurgica. https://doi.org/10.1007/s00701-026-06954-9
[84]   Liu, C., Geltzeiler, A., Afyouni, A., et al. (2026). RESCUE: An end-to-end multi-agent LLM system for proactive rare-disease patient screening in the EHR [Preprint]. medRxiv. https://doi.org/10.64898/2026.06.24.26356357
[85]   Maniscalco, M. A., Park, Y. K., Domal, S. J., et al. (2026). Autonomous radiotherapy planning via agentic orchestration using a multimodal TPS-integrated compound AI platform. Machine Learning: Health. https://doi.org/10.1088/3049-477X/ae7978
[86]   Wang, Q., Wang, Z., Li, M., et al. (2025). A feasibility study of automating radiotherapy planning with large language model agents. Physics in Medicine & Biology. https://doi.org/10.1088/1361-6560/adbff1
[87]   Choi, H., Bae, S., & Na, K. J. (2026). End-to-end PET/CT interpretation and quantification with an LLM-orchestrated AI agent: A real-world pilot study. Journal of Nuclear Medicine. https://doi.org/10.2967/jnumed.126.272362
[88]   Vashistha, R., Brosda, S., Aoude, L. G., et al. (2026). Agent-MIRA: AI-orchestrated medical imaging agent for PET image retrieval and assistance. Computerized Medical Imaging and Graphics. https://doi.org/10.1016/j.compmedimag.2026.102725
[89]   Solages, N., Scherer, R., Samico, G. A., et al. (2026). GLLaucoMed: A secure LLM-powered agentic workflow for automated medication extraction from free-text glaucoma clinical notes [Preprint]. medRxiv. https://doi.org/10.64898/2026.06.12.26355525
[90]   Yang, E., Garcia, T., Williams, H. G., et al. (2025). A behavioral science-informed agentic workflow for personalized nutrition coaching: Development and validation study. JMIR Formative Research. https://doi.org/10.2196/75421
[91]   Li, C., Lai, P., Zhang, N., et al. (2026). EcoRxAgent: An AI agent for generating economically substitutable prescriptions. npj Digital Medicine. https://doi.org/10.1038/s41746-026-02612-7
[92]   Chen, J., Cao, Y., Feng, Y., et al. (2025). Autonomous artificial intelligence prescribing a drug to prevent severe acute graft-versus-host disease in HLA-haploidentical transplants. Nature Communications. https://doi.org/10.1038/s41467-025-62926-0
[93]   Lanzola, G., Polce, F., & Parimbelli, E. (2023). The Case Manager: An agent controlling the activation of knowledge sources in a FHIR-based distributed reasoning environment. Applied Clinical Informatics. https://doi.org/10.1055/a-2113-4443
Volume 9, Issue 2
Spring 2026
Pages 84-111

  • Receive Date 23 November 2025
  • Revise Date 29 January 2026
  • Accept Date 04 March 2026