Abstract
As Generative Artificial Intelligence (AI) and Robotic Process Automation (RPA) tools become increasingly integrated into digital workflows, ensuring the usability of such automation before deployment on live systems is critical. This paper introduces a novel sandbox environment implemented within a virtual machine designed to safely test AI-driven task automation in isolation. The study evaluates user interactions with automated systems through two distinct feedback modalities: direct mouse control and text-based input, across tasks of varying difficulty levels (easy, medium, and hard) through measuring System Usability Scale, task completion rate and NASA TLX. The experiment introduces two primary independent variables: interaction modality and task difficulty. Interaction modality is categorized into direct mouse control by the user versus providing guidance to the AI through a chat interface. Task difficulty is divided into three levels-easy, medium, and hard; each presented sequentially to participants within their assigned interaction modality. A field experiment with 28 participants revealed that direct mouse control outperformed text-based feedback in task completion rates (83.3% vs. 61.9%) and usability. However, as task difficulty increased, user workload also rose significantly, regardless of the feedback modality. Qualitative analysis highlighted common barriers to effective interaction, such as delayed AI responses and frustration with error correction responsiveness. The study aims to identify patterns in user-AI collaboration dynamics, pinpoint challenges in the AI's autonomous decision-making, and assess the efficacy of the intervention methods. The findings are expected to inform the design of future human-centered AI systems that can effectively balance autonomy with user oversight in complex environments.
Cite this work
Wangfan Li, Athish Venkatachalam, Jackson Bowen, and Carlos Toxtli-Hernández. 2025. Human Oversight Over Autonomous Task Execution in Sandbox Environments. 2025 IEEE Conference on Artificial Intelligence (CAI). https://doi.org/10.1109/cai64502.2025.00121
@inproceedings{Li2025Human,
title = {Human Oversight Over Autonomous Task Execution in Sandbox Environments},
url = {http://dx.doi.org/10.1109/CAI64502.2025.00121},
doi = {10.1109/cai64502.2025.00121},
booktitle = {2025 IEEE Conference on Artificial Intelligence (CAI)},
publisher = {IEEE},
author = {Li, Wangfan and Venkatachalam, Athish and Bowen, Jackson and Toxtli-Hernández, Carlos},
year = {2025},
month = May,
pages = {663--668}
}