[1] Topol, E. J. (2019). High-performance medicine: The convergence of human and artificial intelligence.
Nature Medicine, 25(1), 44–56.
https://doi.org/10.1038/s41591-018-0300-7
[2] Rajpurkar, P., Chen, E., Banerjee, O., & Topol, E. J. (2022). AI in health and medicine.
Nature Medicine, 28(1), 31–38.
https://doi.org/10.1038/s41591-021-01614-0
[3] Liu, X., Cruz Rivera, S., Moher, D., Calvert, M. J., Denniston, A. K., & SPIRIT-AI and CONSORT-AI Working Group. (2020). Reporting guidelines for clinical trial reports for interventions involving artificial intelligence: The CONSORT-AI extension.
Nature Medicine, 26(9), 1364–1374.
https://doi.org/10.1038/s41591-020-1034-x
[4] van de Sande, D., et al. (2022). Developing, implementing and governing artificial intelligence in medicine: A step-by-step approach to prevent an artificial intelligence winter.
BMJ Health Care Informatics, 29(1), e100495.
https://doi.org/10.1136/bmjhci-2021-100495
[5] Singhal, K., Azizi, S., Tu, T., et al. (2023). Large language models encode clinical knowledge.
Nature, 620(7972), 172–180.
https://doi.org/10.1038/s41586-023-06291-2
[6] Nori, H., King, N., McKinney, S. M., Carignan, D., & Horvitz, E. (2023).
Capabilities of GPT-4 on medical challenge problems [Preprint]. arXiv.
https://arxiv.org/abs/2303.13375
[7] Moor, M., Banerjee, O., Abad, Z. S. H., et al. (2023). Foundation models for generalist medical artificial intelligence.
Nature, 616(7956), 259–265.
https://doi.org/10.1038/s41586-023-05881-4
[8] Zhao, L., Liu, S., Xin, T., et al. (2026). AI agent in healthcare: Applications, evaluations, and future directions.
npj Artificial Intelligence, 2, 31.
https://doi.org/10.1038/s44387-026-00076-4
[9] Xu, X., & Sankar, R. (2025). Large language model agents for biomedicine: A comprehensive review of methods, evaluations, challenges, and future directions.
Information, 16(10), 894.
https://doi.org/10.3390/info16100894
[10] Masterman, T., Besen, S., Sawtell, M., & Chao, A. (2024).
The landscape of emerging AI agent architectures for reasoning, planning, and tool calling: A survey [Preprint]. arXiv.
https://arxiv.org/abs/2404.11584
[11] Tierney, A. A., Gayre, G., Hoberman, B., et al. (2025). Ambient artificial intelligence scribes: Learnings after 1 year and over 2.5 million uses.
NEJM Catalyst Innovations in Care Delivery, 6(5), CAT.25.0040.
https://doi.org/10.1056/CAT.25.0040
[12] Lukac, P. J., Turner, W., Vangala, S., et al. (2025). Ambient AI scribes in clinical practice: A randomized trial.
NEJM AI, 2(12), AIoa2501000.
https://doi.org/10.1056/AIoa2501000
[13] Olson, K. D., Meeker, D., Troup, M., et al. (2025). Use of ambient AI scribes to reduce administrative burden and professional burnout.
JAMA Network Open, 8(10), e2534976.
https://doi.org/10.1001/jamanetworkopen.2025.34976
[14] Shi, W., et al. (2024). EHRAgent: Code empowers large language models for few-shot complex tabular reasoning on electronic health records. In
Proceedings of EMNLP 2024 (pp. 22315–22339).
https://doi.org/10.18653/v1/2024.emnlp-main.1245
[15] Zakka, C., Cho, J., Fahed, G., et al. (2025).
Almanac Copilot: Towards autonomous electronic health record navigation [Preprint]. Research Square.
https://doi.org/10.21203/rs.3.rs-6102516/v1
[16] Gebreab, S. A., et al. (2024). LLM-based framework for administrative task automation in healthcare. In
Proceedings of the 12th International Symposium on Digital Forensics and Security (ISDFS) (pp. 1–7). IEEE.
https://doi.org/10.1109/ISDFS60797.2024.10527275
[17] Pandey, H. G., Amod, A., & Kumar, S. (2024). Advancing healthcare automation: Multi-agent system for medical necessity justification.
Proceedings of the BioNLP Workshop.
https://doi.org/10.18653/v1/2024.bionlp-1.4
[18] Tu, T., Palepu, A., Schaekermann, M., et al. (2025). Towards conversational diagnostic artificial intelligence.
Nature.
https://doi.org/10.1038/s41586-025-08866-7
[19] Ferber, D., El Nahhas, O. S. M., Wölflein, G., et al. (2025). Development and validation of an autonomous artificial intelligence agent for clinical decision-making in oncology.
Nature Cancer, 6(8), 1337–1349.
https://doi.org/10.1038/s43018-025-00991-6
[20] Kim, Y., et al. (2024). MDAgents: An adaptive collaboration of LLMs for medical decision-making. In
Advances in Neural Information Processing Systems (NeurIPS).
https://doi.org/10.52202/079017-2522
[21] Wang, Z., et al. (2025). ColaCare: Enhancing electronic health record modeling through large language model-driven multi-agent collaboration. In
Proceedings of the ACM Web Conference 2025 (pp. 2250–2261).
https://doi.org/10.1145/3696410.3714877
[22] Li, J., et al. (2024).
Agent Hospital: A simulacrum of hospital with evolvable medical agents [Preprint]. arXiv.
https://arxiv.org/abs/2405.02957
[23] Nie, M., Chung, W., Waxler, J., et al. (2026). Hard to halt: Automation bias in agent-driven sequencing prior authorization workflows [Preprint].
medRxiv.
https://doi.org/10.64898/2026.06.16.26355782
[24] Yang, G. Z., Cambias, J., Cleary, K., et al. (2017). Medical robotics—Regulatory, ethical, and legal considerations for increasing levels of autonomy.
Science Robotics, 2(4), eaam8638.
https://doi.org/10.1126/scirobotics.aam8638
[25] European Parliament & Council of the European Union. (2024). Regulation (EU) 2024/1689 laying down harmonised rules on artificial intelligence (Artificial Intelligence Act). Official Journal of the European Union.
[26] U.S. Food and Drug Administration. (2025). Artificial intelligence-enabled device software functions: Lifecycle management and marketing submission recommendations—Draft guidance. U.S. Food and Drug Administration.
[27] Amodei, D., Olah, C., Steinhardt, J., Christiano, P., Schulman, J., & Mane, D. (2016).
Concrete problems in AI safety [Preprint]. arXiv.
https://arxiv.org/abs/1606.06565
[28] Cemri, M., et al. (2025).
Why do multi-agent LLM systems fail? [Preprint]. arXiv.
https://doi.org/10.52202/085713-4082
[29] Jiang, Y., Black, K. C., Geng, G., et al. (2025). MedAgentBench: A virtual EHR environment to benchmark medical LLM agents.
NEJM AI, 2(9), AIdbp2500144.
https://doi.org/10.1056/AIdbp2500144
[30] Alkaissi, H., & McFarlane, S. I. (2023). Artificial hallucinations in ChatGPT: Implications in scientific writing.
Cureus, 15(2), e35179.
https://doi.org/10.7759/cureus.35179
[31] Goddard, K., Roudsari, A., & Wyatt, J. C. (2012). Automation bias: A systematic review of frequency, effect mediators, and mitigators.
Journal of the American Medical Informatics Association, 19(1), 121–127.
https://doi.org/10.1136/amiajnl-2011-000089
[32] Parasuraman, R., & Manzey, D. H. (2010). Complacency and bias in human use of automation: An attentional integration.
Human Factors, 52(3), 381–410.
https://doi.org/10.1177/0018720810376055
[33] Ancker, J. S., Edwards, A., Nosal, S., Hauser, D., Mauer, E., & Kaushal, R. (2017). Effects of workload, work complexity, and repeated alerts on alert fatigue in a clinical decision support system.
BMC Medical Informatics and Decision Making, 17(1), 36.
https://doi.org/10.1186/s12911-017-0430-8
[34] Guo, Y., Hu, D., Zhou, Y., et al. (2026). From conversation to chart: An analysis of clinician edits to ambient AI draft notes [Preprint].
medRxiv.
https://doi.org/10.64898/2026.01.05.26343471
[35] Price, W. N., II, Gerke, S., & Cohen, I. G. (2019). Potential liability for physicians using artificial intelligence.
JAMA, 322(18), 1765–1766.
https://doi.org/10.1001/jama.2019.15064
[36] Habli, I., Lawton, T., & Porter, Z. (2020). Artificial intelligence in health care: Accountability and safety.
Bulletin of the World Health Organization, 98(4), 251–256.
https://doi.org/10.2471/BLT.19.237487
[37] Gerke, S., Minssen, T., & Cohen, G. (2020). Ethical and legal challenges of artificial intelligence-driven healthcare. In
Artificial intelligence in healthcare (pp. 295–336). Academic Press.
https://doi.org/10.1016/B978-0-12-818438-7.00012-5
[38] World Health Organization. (2024). Ethics and governance of artificial intelligence for health: Guidance on large multi-modal models. World Health Organization.
[39] Gilbert, S., Harvey, H., Melvin, T., Vollebregt, E., & Wicks, P. (2023). Large language model AI chatbots require approval as medical devices.
Nature Medicine, 29(10), 2396–2398.
https://doi.org/10.1038/s41591-023-02412-6
[40] Thirunavukarasu, A. J., Ting, D. S. J., Elangovan, K., Gutierrez, L., Tan, T. F., & Ting, D. S. W. (2023). Large language models in medicine.
Nature Medicine, 29(8), 1930–1940.
https://doi.org/10.1038/s41591-023-02448-8
[41] Collaco, B. G., Haider, S. A., Prabha, S., et al. (2026). The role of agentic artificial intelligence in healthcare: A scoping review.
npj Digital Medicine, 9, 345.
https://doi.org/10.1038/s41746-026-02517-5
[42] Njei, B., Al-Ajlouni, Y. A., Kanmounye, U. S., Boateng, S., Nguefang, G. L., Njei, N., et al. (2026). Artificial intelligence agents in healthcare research: A scoping review.
PLOS ONE, 21(2), e0342182.
https://doi.org/10.1371/journal.pone.0342182
[43] Gorenshtein, A., Omar, M., Glicksberg, B. S., Nadkarni, G. N., & Klang, E. (2025). AI agents in clinical medicine: A systematic review [Preprint].
medRxiv.
https://doi.org/10.1101/2025.08.22.25334232
[44] Xu, J., Ko, J. M., & Kvedar, J. C. (2026). AI agents in clinical practice: An evidence map.
npj Digital Medicine.
https://doi.org/10.1038/s41746-026-02960-4
[45] SAE International. (2021). Taxonomy and definitions for terms related to driving automation systems for on-road motor vehicles (SAE Standard J3016_202104).
[46] Wooldridge, M., & Jennings, N. R. (1995). Intelligent agents: Theory and practice.
The Knowledge Engineering Review, 10(2), 115–152.
https://doi.org/10.1017/S0269888900008122
[47] Russell, S., & Norvig, P. (2021). Artificial intelligence: A modern approach (4th ed.). Pearson.
[48] Xi, Z., Chen, W., Guo, X., et al. (2025). The rise and potential of large language model based agents: A survey.
Science China Information Sciences, 68(2), 121101.
https://doi.org/10.1007/s11432-024-4222-0
[49] Wang, L., Ma, C., Feng, X., et al. (2024). A survey on large language model based autonomous agents.
Frontiers of Computer Science, 18(6), 186345.
https://doi.org/10.1007/s11704-024-40231-1
[50] Sheridan, T. B., & Verplank, W. L. (1978).
Human and computer control of undersea teleoperators. MIT Man-Machine Systems Laboratory.
https://doi.org/10.21236/ADA057655
[51] Parasuraman, R., Sheridan, T. B., & Wickens, C. D. (2000). A model for types and levels of human interaction with automation.
IEEE Transactions on Systems, Man, and Cybernetics—Part A: Systems and Humans, 30(3), 286–297.
https://doi.org/10.1109/3468.844354
[52] Endsley, M. R. (2017). From here to autonomy: Lessons learned from human-automation research.
Human Factors, 59(1), 5–27.
https://doi.org/10.1177/0018720816681350
[53] Santoni de Sio, F., & van den Hoven, J. (2018). Meaningful human control over autonomous systems: A philosophical account.
Frontiers in Robotics and AI, 5, 15.
https://doi.org/10.3389/frobt.2018.00015
[54] Page, M. J., McKenzie, J. E., Bossuyt, P. M., et al. (2021). The PRISMA 2020 statement: An updated guideline for reporting systematic reviews.
BMJ, 372, n71.
https://doi.org/10.1136/bmj.n71
[55] Rethlefsen, M. L., Kirtley, S., Waffenschmidt, S., et al., & PRISMA-S Group. (2021). PRISMA-S: An extension to the PRISMA Statement for reporting literature searches in systematic reviews.
Systematic Reviews, 10(1), 39.
https://doi.org/10.1186/s13643-020-01542-z
[56] McGowan, J., Sampson, M., Salzwedel, D. M., Cogo, E., Foerster, V., & Lefebvre, C. (2016). PRESS peer review of electronic search strategies: 2015 guideline statement.
Journal of Clinical Epidemiology, 75, 40–46.
https://doi.org/10.1016/j.jclinepi.2016.01.021
[57] Bramer, W. M., Giustini, D., de Jonge, G. B., Holland, L., & Bekhuis, T. (2016). De-duplication of database search results for systematic reviews in EndNote.
Journal of the Medical Library Association, 104(3), 240–243.
https://doi.org/10.3163/1536-5050.104.3.014
[58] Ouzzani, M., Hammady, H., Fedorowicz, Z., & Elmagarmid, A. (2016). Rayyan—A web and mobile app for systematic reviews.
Systematic Reviews, 5, 210.
https://doi.org/10.1186/s13643-016-0384-4
[59] McHugh, M. L. (2012). Interrater reliability: The kappa statistic.
Biochemia Medica, 22(3), 276–282.
https://doi.org/10.11613/BM.2012.031
[60] Sterne, J. A. C., Savovic, J., Page, M. J., et al. (2019). RoB 2: A revised tool for assessing risk of bias in randomised trials.
BMJ, 366, l4898.
https://doi.org/10.1136/bmj.l4898
[61] Sounderajah, V., Ashrafian, H., Rose, S., et al. (2021). A quality assessment tool for artificial intelligence-centered diagnostic test accuracy studies: QUADAS-AI.
Nature Medicine, 27(10), 1663–1665.
https://doi.org/10.1038/s41591-021-01517-0
[62] Wolff, R. F., Moons, K. G. M., Riley, R. D., et al. (2019). PROBAST: A tool to assess the risk of bias and applicability of prediction model studies.
Annals of Internal Medicine, 170(1), 51–58.
https://doi.org/10.7326/M18-1376
[63] Gallifant, J., Afshar, M., Ameen, S., et al. (2025). The TRIPOD-LLM reporting guideline for studies using large language models.
Nature Medicine, 31(1), 60–69.
https://doi.org/10.1038/s41591-024-03425-5
[64] Vasey, B., Nagendran, M., Campbell, B., et al. (2022). Reporting guideline for the early-stage clinical evaluation of decision support systems driven by artificial intelligence: DECIDE-AI.
Nature Medicine, 28(5), 924–933.
https://doi.org/10.1038/s41591-022-01772-9
[65] Campbell, M., McKenzie, J. E., Sowden, A., et al. (2020). Synthesis without meta-analysis (SWiM) in systematic reviews: Reporting guideline.
BMJ, 368, l6890.
https://doi.org/10.1136/bmj.l6890
[66] Miake-Lye, I. M., Hempel, S., Shanman, R., & Shekelle, P. G. (2016). What is an evidence map? A systematic review of published evidence maps and their definitions, methods, and products.
Systematic Reviews, 5, 28.
https://doi.org/10.1186/s13643-016-0204-x
[67] Cabitza, F., Rasoini, R., & Gensini, G. F. (2017). Unintended consequences of machine learning in medicine.
JAMA, 318(6), 517–518.
https://doi.org/10.1001/jama.2017.7797
[68] Helmreich, R. L. (2000). On error management: Lessons from aviation.
BMJ, 320(7237), 781–785.
https://doi.org/10.1136/bmj.320.7237.781
[69] Barach, P., & Small, S. D. (2000). Reporting and preventing medical mishaps: Lessons from non-medical near miss reporting systems.
BMJ, 320(7237), 759–763.
https://doi.org/10.1136/bmj.320.7237.759
[70] Schmidgall, S., Ziaei, R., Harris, C., et al. (2026). AgentClinic: A multimodal benchmark for tool-using clinical AI agents.
npj Digital Medicine.
https://doi.org/10.1038/s41746-026-02674-7
[71] Chen, E., Postelnik, S., Black, K., et al. (2026). MedAgentBench v2: Improving medical LLM agent design. In
Pacific Symposium on Biocomputing.
https://doi.org/10.1142/9789819824755_0025
[72] Liu, Y., Carrero, Z. I., Jiang, X., et al. (2026). Benchmarking large language model-based agent systems for clinical decision tasks.
npj Digital Medicine.
https://doi.org/10.1038/s41746-026-02443-6
[73] Mokssit, Y., Ravi, K., Nie, M., et al. (2026). FHIR-AgentEval: A modular sandbox for benchmarking clinical LLM agents with an evaluation of memory-augmented configurations [Preprint].
Research Square.
https://doi.org/10.21203/rs.3.rs-8746188/v1
[74] Hou, R., Xue, D., Sun, H., et al. (2026). CDAFlow: Enhancing LLM clinical decision-making through agentic workflow.
Expert Systems with Applications.
https://doi.org/10.1016/j.eswa.2026.131806
[75] Xu, S., Huang, X., Wei, Z., et al. (2026). DxDirector: An agentic large language model driving the full-process clinical diagnosis.
Nature Communications.
https://doi.org/10.1038/s41467-026-71928-5
[76] Han, S., & Choi, W. (2025). Development of a large language model-based multi-agent clinical decision support system for Korean Triage and Acuity Scale (KTAS)-based triage and treatment planning in emergency departments.
Advances in Artificial Intelligence and Machine Learning.
https://doi.org/10.54364/AAIML.2025.51187
[77] Wu, X., Zhang, H., Garduno-Rapp, N. E., et al. (2026). Orchestrator multi-agent clinical decision support system for secondary headache diagnosis in primary care.
Journal of the American Medical Informatics Association.
https://doi.org/10.1093/jamia/ocag111
[78] Klang, E., Omar, M., Raut, G., et al. (2026). Orchestrated multi agents sustain accuracy under clinical-scale workloads compared to a single agent.
npj Health Systems.
https://doi.org/10.1038/s44401-026-00077-0
[79] Ucdal, M., & Ekingen, E. (2026). Performance comparison of a neuro-symbolic large language model system versus conventional AI models and human experts in cholangitis management.
BMC Medical Informatics and Decision Making.
https://doi.org/10.1186/s12911-026-03593-z
[80] Ekingen, E., & Ucdal, M. (2026). Performance comparison of a neuro-symbolic large language model system versus human experts in acute cholecystitis management.
Journal of Clinical Medicine, 15(5), 1730.
https://doi.org/10.3390/jcm15051730
[81] Zhai, G., Bar, M., Cowan, A. J., et al. (2025). AI for evidence-based treatment recommendation in oncology: A blinded evaluation of large language models and agentic workflows.
Frontiers in Artificial Intelligence.
https://doi.org/10.3389/frai.2025.1683322
[82] Wang, J., Mullick Chowdhury, S., & Nazha, A. (2025). Virtual oncology collaborative tumor board using multiple artificial intelligence agents.
Journal of Clinical Oncology, 43(16_suppl), 1563.
https://doi.org/10.1200/JCO.2025.43.16_suppl.1563
[83] Kochuiev, E., Kaliuzhka, V., Markevych, M., et al. (2026). An agentic AI framework for integrated decision support and surgical planning in intracerebral hemorrhage.
Acta Neurochirurgica.
https://doi.org/10.1007/s00701-026-06954-9
[84] Liu, C., Geltzeiler, A., Afyouni, A., et al. (2026). RESCUE: An end-to-end multi-agent LLM system for proactive rare-disease patient screening in the EHR [Preprint].
medRxiv.
https://doi.org/10.64898/2026.06.24.26356357
[85] Maniscalco, M. A., Park, Y. K., Domal, S. J., et al. (2026). Autonomous radiotherapy planning via agentic orchestration using a multimodal TPS-integrated compound AI platform.
Machine Learning: Health.
https://doi.org/10.1088/3049-477X/ae7978
[86] Wang, Q., Wang, Z., Li, M., et al. (2025). A feasibility study of automating radiotherapy planning with large language model agents.
Physics in Medicine & Biology.
https://doi.org/10.1088/1361-6560/adbff1
[87] Choi, H., Bae, S., & Na, K. J. (2026). End-to-end PET/CT interpretation and quantification with an LLM-orchestrated AI agent: A real-world pilot study.
Journal of Nuclear Medicine.
https://doi.org/10.2967/jnumed.126.272362
[88] Vashistha, R., Brosda, S., Aoude, L. G., et al. (2026). Agent-MIRA: AI-orchestrated medical imaging agent for PET image retrieval and assistance.
Computerized Medical Imaging and Graphics.
https://doi.org/10.1016/j.compmedimag.2026.102725
[89] Solages, N., Scherer, R., Samico, G. A., et al. (2026). GLLaucoMed: A secure LLM-powered agentic workflow for automated medication extraction from free-text glaucoma clinical notes [Preprint].
medRxiv.
https://doi.org/10.64898/2026.06.12.26355525
[90] Yang, E., Garcia, T., Williams, H. G., et al. (2025). A behavioral science-informed agentic workflow for personalized nutrition coaching: Development and validation study.
JMIR Formative Research.
https://doi.org/10.2196/75421
[91] Li, C., Lai, P., Zhang, N., et al. (2026). EcoRxAgent: An AI agent for generating economically substitutable prescriptions.
npj Digital Medicine.
https://doi.org/10.1038/s41746-026-02612-7
[92] Chen, J., Cao, Y., Feng, Y., et al. (2025). Autonomous artificial intelligence prescribing a drug to prevent severe acute graft-versus-host disease in HLA-haploidentical transplants.
Nature Communications.
https://doi.org/10.1038/s41467-025-62926-0
[93] Lanzola, G., Polce, F., & Parimbelli, E. (2023). The Case Manager: An agent controlling the activation of knowledge sources in a FHIR-based distributed reasoning environment.
Applied Clinical Informatics.
https://doi.org/10.1055/a-2113-4443