Document Type : Original Article
Authors
1
Associate Professor of Biomedical Engineering, Vali-e-Asr University of Rafsanjan
2
Department of Biomedical Engineering, Meybod University, Meybod, Iran
3
Department of Engineering.Vali-e-Asr University of Rafsanjan,Iran
4
Assistant Professor,Faculty of Electrical and Computer Engineering, Shams Gonbad Higher Education Institute, Gorgan, Iran
10.22034/tmi.2026.248202
Abstract
Background: Healthcare AI agents increasingly plan, invoke tools, and execute multi-step workflows. Existing reviews have mapped applications and architectures, but have not jointly classified consequential action authority, multi-step failure expression, and the empirically measured effectiveness of human oversight. This review evaluates whether greater operational autonomy is supported by evidence on safety and human control.
Methods: PubMed/MEDLINE, Scopus, and Web of Science Core Collection were searched from 1 January 2022 through 27 July 2026. The searches retrieved 4,096 records, of which 2,557 unique report-level records remained after deduplication. For the current synthesis, 50 priority reports underwent full-text retrieval; 34 were assessed, 28 met the operational agent criteria, six were excluded, and one was unavailable. One reviewer performed AI-assisted structured extraction of workflow, architecture, action scope, autonomy, safety, oversight, evaluation, and governance data. Heterogeneous evidence was synthesized narratively following SWiM principles.
Results: The 28 studies were published from 2023 to 2026; 18 (64.3%) appeared in 2026. Diagnostic and clinical decision-support workflows accounted for 13 studies, EHR/information or administrative workflows for nine, treatment planning or prescribing for four, and patient-facing or cross-domain workflows for two. Twelve systems were single-agent, 12 multi-agent, three hybrid, and one incompletely specified. Operational autonomy was L0 in five studies, L1 in 13, L2 in three, L3 in one, and L4 in six; none reached L5. Five of the six L4 systems operated only in sandboxes or retrospective research environments, whereas one prospective clinical trial implemented live conditional prescribing. Safety-evidence grades were high in one study, moderate in 15, low in 11, and very low in one. Only three studies empirically characterized an operating property of oversight, and none measured approval-gate catch rate under sustained clinical workload.
Conclusions: The current evidence supports tool-using and workflow-capable agents, but not unrestricted clinical autonomy. Operational capability frequently exceeded the maturity of safety evaluation, and most oversight mechanisms were asserted by design rather than validated in use. Autonomy should therefore be promoted only within a bounded operational design domain and only after enacted and propagated failures, escalation sensitivity, override behavior, and reviewer workload have been measured.
Keywords