Abstract
In an era where the automation of tasks via Generative AI and Robotic Process Automation (RPA) tools has become increasingly prevalent, ensuring the safety and reliability of these automations before their deployment on live systems is paramount. This paper introduces a novel sandbox environment, implemented within a virtual machine, designed to serve as a testing ground for autonomous task executions powered by Large Language Models (LLMs) and Large Multimodal Models (LMMs). Our solution leverages a human-in-the-loop approach, wherein users validate and provide feedback on tasks executed within the sandbox, thereby informing iterative refinements to the models' internal prompts. This feedback loop facilitates the in-context learning of models, allowing them to adapt and prevent recurring errors in future task executions. We explore the effectiveness of this approach through user studies conducted in a controlled virtual environment, utilizing proprietary and open-source multimodal models. The study is structured around three distinct scenarios, each designed to evaluate the models' performance across different tasks and user interactions. Through this research, we aim to not only enhance the safety and efficacy of task automation technologies but also to contribute to the broader discourse on human oversight mechanisms in AI-driven systems. This work underscores the importance of sandbox environments in mitigating potential risks associated with deploying automated tasks in live environments, and highlights the role of human feedback in refining AI behaviors for optimal performance.
Cite this work
Wangfan Li and Carlos Toxtli-Hernández. 2024. Human Oversight Mechanisms over Autonomous Task Execution in Sandbox Environments. CHIWORK '24: 3rd Annual Meeting of the Symposium on Human-Computer Interaction for Work, Demo Track. https://doi.org/10.13140/RG.2.2.26571.20003
@misc{Li2024Human,
doi = {10.13140/RG.2.2.26571.20003},
url = {https://www.researchgate.net/doi/10.13140/RG.2.2.26571.20003},
author = {{Wangfan Li} and Toxtli, Carlos},
language = {en},
title = {Human Oversight Mechanisms over Autonomous Task Execution in Sandbox Environments},
publisher = {Unpublished},
year = {2024},
howpublished = {Preprint, ResearchGate},
note = {Preprint}
}