Supervisory Control for LLMs and Autonomous Agents
Grounding AI oversight in the science of human supervisory control.
Keeping people in meaningful control of increasingly autonomous AI.
Generative AI is turning software automation from scripted bots into agents that plan and act on their own. Our lab studies what it takes for that shift to empower rather than sideline the people involved. We ground agent design in theories of human supervisory control, and we build environments where humans can observe, steer and correct autonomous behavior.
Our work includes Prompt-Level Supervisory Alignment (PLSA), which operationalizes Supervisory Control Theory as a structured prompting strategy for LLM revision; sandboxed virtual machines for human oversight of autonomous task execution; LLM-generated BPMN workflows for robotic process automation; retrieval-augmented human-AI co-authoring of software requirements; and empirical studies of how teams of LLM agents pool information, coordinate and fail. This agenda is synthesized in the book Human-Centered Automation (CRC Press, 2026).
Projects
Grounding AI oversight in the science of human supervisory control.
Do language models know what they know, and should we believe them?
From a plain-English request to a verified Fermi-LAT analysis.
Privacy-preserving activity sensing and LLM coaching for knowledge workers.
Publications
Human-Centered Automation (HCA) is becoming indispensable for organizations seeking to implement AI-driven systems, robotic process automation, and other advanced tools while keeping human needs at the core. Despite the widespread adoption of automation, many initiatives fall short due to a lack of alignment between technical capabilities and real-world user and organizational needs. This book unites insights from cognitive science, software engineering, business strategy, and human factors that view automation as something that works for people and not just systems. This book provides the tools and strategies to ensure efforts with automation succeed with people at the center. It takes readers through the entire HCA lifecycle; from process discovery and planning to system design, testing, and long-term oversight. With easy-to-understand frameworks and real-world examples across sectors, including healthcare, finance, and manufacturing, readers will gain tools to assess risk, define measurable outcomes, involve stakeholders, and build automation that is trustworthy, explainable, and effective. It offers the methodologies needed to drive meaningful, lasting impact through automation. Human-Centered Automation is essential for ergonomics and human factors professionals, researchers, and organizational leaders involved in designing or managing automation initiatives. Its appeal extends to professionals in AI and machine learning, UX design, process engineering, business operations, and policy development.
@book{ToxtliHernandez2026HumanCentered,
title = {Human-Centered Automation},
isbn = {9781003666400},
url = {http://dx.doi.org/10.1201/9781003666400},
doi = {10.1201/9781003666400},
publisher = {CRC Press},
author = {Toxli-Hernandez, Carlos},
year = {2026},
month = May
}Safety evaluations of tool agents often retain terminal scores without the trajectories needed to explain them. We contribute an executable, claim-scoped evidence contract spanning identity, path, transition, assignment, target, and discrimination, and validate it in a prospectively frozen model-only deployment sandbox with hard preconditions and rollback. The study crosses serving aliases, state dynamics, benign and injected tickets, and repeated seeds while retaining every request, raw response, parse result, retry, state transition, and endpoint. The resulting paths expose mechanisms that terminal scores conflate: generic unsafe transitions can be attempted without reaching the injected exfiltration target, whereas abort instructions mainly cause premature termination. Offline replay reproduces every retained transition. The artifact therefore supports mechanism-level diagnosis on the evaluated finite grid, but not claims about other models, policies, providers, deployments, or people. Because immutable weight and container hashes are unavailable, the identity gate also blocks claims of exact behavioral reproduction. The study uses no human experiment or human-derived label.
@inproceedings{Toxtli2026Evidence,
title = {Evidence Before Rankings: An Executable Audit Contract for Stateful Tool-Agent Evaluations},
author = {Toxtli, Carlos and Delaflor, Manuel},
booktitle = {Third Workshop on Agents in the Wild: Safety, Security, and Beyond (AIWILD) at NeurIPS 2026},
address = {Sydney, Australia},
year = {2026},
month = dec,
note = {Poster; forthcoming},
url = {https://agentwild-workshop.github.io/neurips2026/}
}Venues and vendors need a credible evaluator when they compare language-model tools that revise manuscripts from peer reviews, yet a scalar LLM judge is often used without evidence that its preferences track useful revision. We study one such judge-and-rubric pipeline on revision trajectories drawn from multiple machine-learning venue years, comparing its scores with document-level similarity to the authors' actual next versions. The two measures order the machine conditions almost oppositely. A second judge from a different model family and serving stack reproduces the condition ranking, which demonstrates stability of the shared rubric-plus-judge pipeline rather than its validity. More importantly, the method preferred by both judges emits truncated, partial manuscripts in about one in seven runs. The retained logs cannot show whether judges reward those individual failures, and formatting differences confound comparisons with author revisions. Our case study, Prompt-Level Supervisory Alignment, produces documents closer to the next author version than the alternatives, but similarity rewards unchanged text and the decisive editing controls remain absent. The supported recommendation is therefore specific: venues should validate revision evaluators on real trajectories and known completeness failures before using them to certify writing tools.
@inproceedings{Toxtli2026Evaluator,
title = {Evaluator Disagreement in AI-Assisted Manuscript Revision},
author = {Toxtli, Carlos and Delaflor, Manuel},
booktitle = {AI-Native Academia: Authorship, Peer Review, and Conference Governance under AI Workshop at NeurIPS 2026},
address = {Atlanta, GA},
year = {2026},
month = dec,
note = {Poster},
url = {https://openreview.net/forum?id=KcjksLJoBz}
}Tools that anticipate peer-review concerns could help authors decide what to revise before submission or another review round. Evaluating such tools is difficult because success may mean matching one realized panel, performing well across possible panels, or improving a paper after an author follows the advice. We separate these targets formally and provide a prospective validation contract covering audit samples, adjudication, error metrics, capacity sensitivity, deterministic controls, and pass/fail rules. The motivating RevPlan-Bench pipeline produced thousands of canonicalized review issues and revision plans, but the underlying corpus and scoring artifacts do not survive in the project. Its assignment rates, model rankings, and inferred issue cascades consequently depend on unaudited extraction, a restrictive unmeasured pre-filter, and unrestricted one-to-many matching. We report them only as pipeline outputs. Because the available record satisfies none of the contract's validation requirements, it supports an estimand and a benchmark governance contribution, but no claim about review dynamics, model foresight, or author utility.
@inproceedings{Delaflor2026Validation,
title = {A Validation Contract for Anticipatory Peer-Review Benchmarks},
author = {Delaflor, Manuel and Toxtli, Carlos},
booktitle = {AI-Native Academia: Authorship, Peer Review, and Conference Governance under AI Workshop at NeurIPS 2026},
address = {Atlanta, GA},
year = {2026},
month = dec,
note = {Poster},
url = {https://openreview.net/forum?id=eZbOdYHVJE}
}Agentic deployments can assign many parallel agents to one supervisor without specifying when that supervisor can no longer recover the team's state. Human-factors research already establishes relevant constructs such as situation awareness, span of control, neglect tolerance, and multiple-resource workload, but language-model teams distribute state across text and can fail silently. This position paper translates those constructs into a measurement protocol designed for later calibration to people rather than a new human experiment. Typed event rates, isolated service costs, and a throughput budget define a nominal load-balance index; the midpoint of a separately fitted capacity curve remains an empirical quantity. Synthetic controls recover a known capacity boundary, while a seeded queue shows that deadline pressure can separate the fitted midpoint from nominal load balance. The resulting event ontology, operating envelope, executable stress tests, and preregisterable falsification criteria allow system designers to test staffing or interface interventions within a declared team-monitor channel. They do not turn synthetic behavior into evidence about human or model supervisors.
@inproceedings{Toxtli2026Supervisor,
title = {How Many Agents Can One Supervisor Track? A Mechanistic Capacity Model and Measurement Protocol for LLM-Agent Teams},
author = {Toxtli, Carlos and Delaflor, Manuel},
booktitle = {Human-AI Coevolution: Measuring Human-Agent Teams in the Agentic Era (HAIC) Workshop at NeurIPS 2026},
address = {Atlanta, GA},
year = {2026},
month = dec,
note = {Workshop paper},
url = {https://openreview.net/forum?id=czINRnBTpf}
}A team protocol has no interpretable benefit without a matched no-protocol baseline and decision-relevant headroom. We present a reporting method that tests admissibility, identifies the marginal contrast, evaluates opportunity, and maps paired uncertainty to an operational action. It distinguishes exact headroom on a fixed battery from uncertain population headroom used in planning. A published meta-analysis of human-AI experiments illustrates why the estimand matters: the same evidence supports augmentation relative to humans alone but contradicts synergy relative to the better constituent. Neither comparison identifies a coordination protocol without a matched no-protocol team. An executable synthetic application further shows that a promising point estimate can remain operationally unresolved once paired uncertainty is considered, and that near-perfect pilot performance does not by itself make a modest benefit impossible. Several inherited multi-model summaries fail the method's artifact gates, so they do not validate it. This is a bounded position and reporting contribution rather than an adoption-ready standard. Human-facing claims still require direct human outcomes, and we introduce no new human experiment.
@inproceedings{Delaflor2026Headroom,
title = {Headroom Before Protocol Effects: A Decision Procedure for Human-Agent Evaluation},
author = {Delaflor, Manuel and Toxtli, Carlos},
booktitle = {Human-AI Coevolution: Measuring Human-Agent Teams in the Agentic Era (HAIC) Workshop at NeurIPS 2026},
address = {Atlanta, GA},
year = {2026},
month = dec,
note = {Workshop paper},
url = {https://openreview.net/forum?id=odOgsQYleZ}
}Textual social-agent studies often change both a proposition and its alleged speaker. The resulting output difference may reflect content, attribution, task demand, or their interaction, so “social influence” does not identify the pathway. We offer an executable methodological audit rather than an empirical case series. The protocol separates baseline, content-only, source-attributed, and attribution-only conditions; checks representation independently of the focal outcome; blocks assignment across items, models, and seeds; retains every failure; and requires either a finite-benchmark or population estimand. It also replaces a structural-social taxonomy informed by outcomes with a prospective multi-label rubric covering information access and timing, message content, social source attribution, normative or incentive stakes, and interface or embodiment. A hypothetical design illustrates the contrasts without invented observations. Researchers translating social paradigms into LLM protocols can use the audit to distinguish source-attribution sensitivity from a response to changed text. The framework does not establish human replication or mechanism equivalence.
@inproceedings{Toxtli2026Content,
title = {Content Is Not Social Attribution: An Audit Protocol for Textual LLM Social Simulation},
author = {Toxtli, Carlos and Delaflor, Manuel},
booktitle = {SocialAgent: Second Workshop on Large Language Models for Social Reasoning and Simulation at NeurIPS 2026},
address = {Atlanta, GA},
year = {2026},
month = dec,
note = {Poster; forthcoming},
url = {https://social-llm-workshop.github.io/}
}Automated and no-code ML tools make model building accessible but can obscure harms that arise when users include sensitive attributes. We embed micro-ethics nudges at key workflow moments and evaluate downstream fairness outcomes and user experience. In a between-subjects experiment (N=34) participants used a simplified AutoML web tool on a subset of the HMDA mortgage dataset with a 10-minute modeling task. The intervention combined in text notice when selecting sensitive attributes such as race, gender, ethnicity and post-training model explanation visualizations, while the base condition showed only model performance metrics. Participants in the Intervention group included significantly fewer sensitive features and produced models with substantially smaller equal-opportunity gaps, while System Usability Scale scores did not differ significantly across conditions. Moral acceptability did not significantly differ between conditions, though it trended lower under the intervention. We conclude that minimal, well timed fairness feedback can meaningfully reduce bias in rapid prototyping workflows. We also discuss design patterns for embedding fairness and the implications of increased moral sensitivity for tool adoption.
@inproceedings{Li2026Eliciting,
title = {Eliciting Fairness via Micro-Ethics Embedded Interfaces for Machine Learning Workflows},
author = {Li, Wangfan and Toxtli, Carlos},
booktitle = {Human-Computer Interaction \& Emerging Technologies},
series = {AHFE Open Access},
volume = {233},
publisher = {AHFE International},
isbn = {979-8-950676-09-3},
issn = {2771-0718},
year = {2026},
doi = {10.54941/ahfe1007537},
url = {https://doi.org/10.54941/ahfe1007537}
}Iterative self-refinement is the dominant paradigm for improving LLM outputs without retraining, yet it lacks principled grounding for what to refine or in what order. We propose Prompt-Level Supervisory Alignment (PLSA), a framework that operationalizes Supervisory Control Theory (SCT), a cognitive framework for human oversight of automated systems, as a structured prompting strategy, and empirically evaluate whether theoretically-grounded prompt structure yields higher revision fidelity than matched iterative self-refinement. In a large-scale evaluation across ten venue-year combinations from three ML conference series (ICLR 2021-2025, NeurIPS 2021-2022 and 2024, CoRL 2021 and 2024), SCT-structured conditions produce revisions with significantly higher fidelity to actual author revisions than both a single-pass baseline and a matched two-pass self-refinement baseline that uses identical review information without SCT structure (all p < .001, medium-to-large effect sizes). All conditions maintain practically equivalent LLM-judge quality, and cross-model evaluation with Google Gemini 2.5 Flash-Lite corroborates condition rankings, confirming findings are not artifacts of generator self-preference. These results provide empirical evidence that theoretically-grounded prompt structure, not merely iterative refinement, is the operative variable driving higher revision fidelity.
@inproceedings{Li_2026a,
series = {CAIS ’26},
title = {Supervisory Control Theory for LLM Revision},
url = {http://dx.doi.org/10.1145/3786335.3813152},
doi = {10.1145/3786335.3813152},
booktitle = {Proceedings of the ACM Conference on AI and Agentic Systems},
publisher = {ACM},
author = {Li, Wangfan and Toxtli, Carlos},
year = {2026},
month = May,
pages = {1100--1108},
collection = {CAIS ’26}
}Effective human-agent cooperation requires that users form mental models of an agent's behavioral tendencies. Yet LLM-based virtual agents are inherently stochastic, undermining the behavioral consistency that mental model formation depends on. We introduce the Constraint-Entropy Tradeoff (CET) model, an information-theoretic design framework that quantifies how persona profiles (behavioral constraints specifying an agent's reasoning style, priorities, and communication patterns) reduce the entropy of a virtual agent's action distribution. The CET model derives that behavioral entropy decays monotonically under constraint strength and identifies an optimal constraint level balancing predictability against flexibility. We validate the framework using a computational testbed with 80 sessions across five conditions, including intermediate-temperature conditions that reveal a threshold effect in the temperature-consistency relationship. Crucially, at the same high temperature, persona-profiled agents recover substantial behavioral consistency compared to unconstrained agents, demonstrating that persona profiles provide independent behavioral constraint beyond temperature reduction. All participants are LLMs; results establish that persona profiles create measurably distinct behavioral patterns, a necessary precondition for human mental model formation, but human validation is needed. We derive domain-specific design guidelines for applications in education, healthcare, and social simulation.
@inproceedings{Toxtli2026PersonalityProfiled,
title = {Personality-Profiled Virtual Agents Are More Predictable: The Constraint-Entropy Tradeoff for Trustworthy Agent Design},
author = {Carlos Toxtli and Manuel Delaflor},
booktitle = {ACM International Conference on Intelligent Virtual Agents (IVA 2026)},
year = {2026},
doi = {10.1145/3806774.3827973}
}@inproceedings{Li_2026b,
series = {SAC ’26},
title = {Requirement Bot: Enhancing Software User Requirement List Through Retrieval-Augmented Human-AI Co-Authoring},
url = {http://dx.doi.org/10.1145/3748522.3780008},
doi = {10.1145/3748522.3780008},
booktitle = {Proceedings of the 41st ACM/SIGAPP Symposium on Applied Computing},
publisher = {ACM},
author = {Li, Wangfan and Toxtli, Carlos},
year = {2026},
month = Mar,
pages = {1437--1443},
collection = {SAC ’26}
}Researchers increasingly use large language models (LLMs) to revise manuscripts in response to peer reviews, yet this adoption is largely unprincipled and risks the ironies of automation. Prompt-Level Supervisory Alignment (PLSA) applies Supervisory Control Theory as a structured prompting strategy for LLM-assisted manuscript revision. We extend PLSA’s planning function to multi-round peer review, where the planner must anticipate concerns that later reviewers will raise. We construct RevPlan-Bench, a corpus of multi-round manuscripts whose ground truth is the issues expert reviewers raised across three or more review cycles, and score 23,256 revision plans spanning five information conditions, three LLM backbones, and four prompting variants. First-round reviews substantially improve coverage of future concerns; multi-agent debate degrades performance; and explicit anticipation prompting, our central pre-registered hypothesis, adds no practical value beyond the reviews themselves, an informative null. Information design is the dominant lever; the researcher remains the final arbiter of scholarly claims.
@article{Li_2026_PLSA,
title = {Prompt-Level Supervisory Alignment for Anticipatory Manuscript Revision},
author = {Li, Wangfan and Campos Hernandez, Sofia Abilene and Toxtli, Carlos},
journal = {Proceedings of the Human Factors and Ergonomics Society Annual Meeting},
publisher = {SAGE Publications},
issn = {2169-5067},
doi = {10.1177/10711813261475162},
url = {http://dx.doi.org/10.1177/10711813261475162},
year = {2026},
month = Aug
}As artificial intelligence (AI) advances, the question of how AI can empower humans over the long term has become increasingly important. This book, Human-AI Empowerment (HAIE), provides a timely exploration of strategies for aligning AI with long-term human goals, ensuring that AI acts as an empowering force across multiple dimensions. Drawing on interdisciplinary research from fields such as AI, HCI, psychology, education, economics, and social science, the book develops comprehensive frameworks for studying and optimizing AI’s impact on human empowerment. HAIE investigates empowerment from a human-centered computing (HCC) perspective, examining how AI systems can track and adapt to progressively achieve long-term goals. The book explores techniques for fostering a mutually beneficial human-AI synergy, delving into AI Empowerment approaches, applicable Human-Computer Interaction methods for long-term engagement, and insights from various disciplines on long-term goal management. Through integrative frameworks, empirical evidence, and ongoing work in the field, this volume informs academics and practitioners seeking to harness AI as a transformative technology for concretely empowering humanity. This book highlights the need for comprehensive approaches to understanding and shaping the future of human-AI collaboration, maximizing its potential to expand human possibilities and support the pursuit of mid-term and long-term goals.
@book{Toxtli_Hern_ndez_2025,
title = {Human-AI Empowerment: An Interdisciplinary Perspective},
isbn = {9781003536628},
url = {http://dx.doi.org/10.1201/9781003536628},
doi = {10.1201/9781003536628},
publisher = {Chapman and Hall/CRC},
author = {Toxtli-Hernández, Carlos},
year = {2025},
month = Sept
}@unpublished{Venkatachalam2025Assessing,
title = {Assessing AI-Generated Workflows: A Multi-Dimensional Evaluation Framework},
author = {Venkatachalam, Athish and Toxtli, Carlos},
note = {Manuscript, ResearchGate},
year = {2025},
month = jul,
url = {https://www.researchgate.net/publication/393884660_Assessing_AI-Generated_Workflows_A_Multi-Dimensional_Evaluation_Framework}
}As Generative Artificial Intelligence (AI) and Robotic Process Automation (RPA) tools become increasingly integrated into digital workflows, ensuring the usability of such automation before deployment on live systems is critical. This paper introduces a novel sandbox environment implemented within a virtual machine designed to safely test AI-driven task automation in isolation. The study evaluates user interactions with automated systems through two distinct feedback modalities: direct mouse control and text-based input, across tasks of varying difficulty levels (easy, medium, and hard) through measuring System Usability Scale, task completion rate and NASA TLX. The experiment introduces two primary independent variables: interaction modality and task difficulty. Interaction modality is categorized into direct mouse control by the user versus providing guidance to the AI through a chat interface. Task difficulty is divided into three levels-easy, medium, and hard; each presented sequentially to participants within their assigned interaction modality. A field experiment with 28 participants revealed that direct mouse control outperformed text-based feedback in task completion rates (83.3% vs. 61.9%) and usability. However, as task difficulty increased, user workload also rose significantly, regardless of the feedback modality. Qualitative analysis highlighted common barriers to effective interaction, such as delayed AI responses and frustration with error correction responsiveness. The study aims to identify patterns in user-AI collaboration dynamics, pinpoint challenges in the AI's autonomous decision-making, and assess the efficacy of the intervention methods. The findings are expected to inform the design of future human-centered AI systems that can effectively balance autonomy with user oversight in complex environments.
@inproceedings{Li2025Human,
title = {Human Oversight Over Autonomous Task Execution in Sandbox Environments},
url = {http://dx.doi.org/10.1109/CAI64502.2025.00121},
doi = {10.1109/cai64502.2025.00121},
booktitle = {2025 IEEE Conference on Artificial Intelligence (CAI)},
publisher = {IEEE},
author = {Li, Wangfan and Venkatachalam, Athish and Bowen, Jackson and Toxtli-Hernández, Carlos},
year = {2025},
month = May,
pages = {663--668}
}This paper introduces ReActIn, a framework designed to infuse human feedback into the intermediate prompting steps of large language models. The practicality and effectiveness of ReActIn are validated through experiments that apply four established prompting strategies, evaluated both with and without human feedback integration. The proposed architecture's performance is compared against traditional large language models across various tasks using four standard evaluation tests. Our findings reveal that the integration of human feedback has a direct impact on the reasoning, action prompting, and overall decision-making capabilities of the language models. This study underscores the potential of ReActIn to shape a future where sophisticated, context-aware AI systems, empowered by human feedback, can effectively navigate complex real-world scenarios.
@misc{Li2024Infusing,
doi = {10.13140/RG.2.2.12282.71364},
url = {https://www.researchgate.net/doi/10.13140/RG.2.2.12282.71364},
author = {{Wangfan Li} and Gendron, Claire and Toxtli, Carlos},
language = {en},
title = {Infusing Human Feedback into Intermediate Prompting Steps of Large Language Models},
publisher = {Unpublished},
year = {2024},
howpublished = {Preprint, ResearchGate},
note = {Preprint}
}This research delves into the capabilities of Large Language Model (LLM)-powered agents across five critical dimensions of task management: decomposition, scheduling, delegation, and execution. Leveraging both surveys and interaction experiments, this study aims to unearth practical insights into the applications and constraints of LLM-powered agents in real-world settings. We describe our experimental setup, data collection methodologies, and analytic techniques, offering a nuanced understanding of these agents' efficiency and efficacy. Our initial findings underscore the nuanced performance of LLM-powered agents in task management, revealing their strengths in understanding complex tasks and their limitations in execution without human intervention. This research contributes to the burgeoning field of human-AI collaboration by providing empirical evidence on the capabilities and limitations of LLM-powered agents in task management.
@misc{Perera2024Assessing,
doi = {10.13140/RG.2.2.11776.85768},
url = {https://www.researchgate.net/doi/10.13140/RG.2.2.11776.85768},
author = {{Ravindu Perera} and {Adithya Ravi} and Toxtli, Carlos},
language = {en},
title = {Assessing the Task Management Capabilities of LLM-Powered Agents},
publisher = {Unpublished},
year = {2024},
howpublished = {Preprint, ResearchGate},
note = {Preprint}
}This paper explores the potential of using Large Language Models (LLMs) to generate Business Process Model and Notation (BPMN) files as input for Robotic Process Automation (RPA) tools. The combination of AI-driven workflow generation with RPA aims to improve the automation of complex processes in both scientific and business domains. The proposed approach leverages the ability of LLMs to interpret natural language instructions and convert them into structured BPMN files, which can then be graphically edited using software like Camunda and executed by RPA tools. The feasibility of this approach is demonstrated through zero-shot prompt experiments using Anthropic Claude 3 Opus and GPT-4o. This solution has the potential to streamline the automation of workflows, reduce human error, and increase productivity across various industries.
@misc{Toxtli2024Automating,
doi = {10.13140/RG.2.2.25996.12165},
url = {https://www.researchgate.net/doi/10.13140/RG.2.2.25996.12165},
author = {Toxtli, Carlos and {Wangfan Li}},
language = {en},
title = {Automating Automation: Using LLMs to Generate BPMN Workflows for Robotic Process Automation},
publisher = {Unpublished},
year = {2024},
howpublished = {Preprint, ResearchGate},
note = {Preprint}
}In an era where the automation of tasks via Generative AI and Robotic Process Automation (RPA) tools has become increasingly prevalent, ensuring the safety and reliability of these automations before their deployment on live systems is paramount. This paper introduces a novel sandbox environment, implemented within a virtual machine, designed to serve as a testing ground for autonomous task executions powered by Large Language Models (LLMs) and Large Multimodal Models (LMMs). Our solution leverages a human-in-the-loop approach, wherein users validate and provide feedback on tasks executed within the sandbox, thereby informing iterative refinements to the models' internal prompts. This feedback loop facilitates the in-context learning of models, allowing them to adapt and prevent recurring errors in future task executions. We explore the effectiveness of this approach through user studies conducted in a controlled virtual environment, utilizing proprietary and open-source multimodal models. The study is structured around three distinct scenarios, each designed to evaluate the models' performance across different tasks and user interactions. Through this research, we aim to not only enhance the safety and efficacy of task automation technologies but also to contribute to the broader discourse on human oversight mechanisms in AI-driven systems. This work underscores the importance of sandbox environments in mitigating potential risks associated with deploying automated tasks in live environments, and highlights the role of human feedback in refining AI behaviors for optimal performance.
@misc{Li2024Human,
doi = {10.13140/RG.2.2.26571.20003},
url = {https://www.researchgate.net/doi/10.13140/RG.2.2.26571.20003},
author = {{Wangfan Li} and Toxtli, Carlos},
language = {en},
title = {Human Oversight Mechanisms over Autonomous Task Execution in Sandbox Environments},
publisher = {Unpublished},
year = {2024},
howpublished = {Preprint, ResearchGate},
note = {Preprint}
}This paper introduces ReActIn, a framework designed to infuse human feedback into the intermediate prompting steps of large language models. The practicality and effectiveness of ReActIn are validated through experiments that apply four established prompting strategies, evaluated both with and without human feedback integration. The proposed architecture's performance is compared against traditional large language models across various tasks using four standard evaluation tests. Our findings reveal that the integration of human feedback has a direct impact on the reasoning, action prompting, and overall decision-making capabilities of the language models. This study underscores the potential of ReActIn to shape a future where sophisticated, context-aware AI systems, empowered by human feedback, can effectively navigate complex real-world scenarios.
@inproceedings{Delaflor_Rodrguez_2024,
series = {IHIET-AI},
title = {ReActIn: Infusing Human Feedback into Intermediate Prompting Steps of Large Language Model},
volume = {120},
issn = {2771-0718},
url = {http://dx.doi.org/10.54941/ahfe1004597},
doi = {10.54941/ahfe1004597},
booktitle = {Human Interaction and Emerging Technologies (IHIET-AI 2024): Artificial Intelligence and Future Applications},
publisher = {AHFE International},
author = {Delaflor Rodrguez, Manuel and Toxtli, Carlos and Gendron, Claire and Li, Wangfan and Delgado Solorzano, Cecilia},
year = {2024},
collection = {IHIET-AI}
}The rapid advancement of Generative Artificial Intelligence (AI), such as Large Language Models (LLMs) and Multimodal Large Language Models (MLLM), has the potential to revolutionize the way we work and interact with digital systems across various industries. However, the current state of software automation, such as Robotic Process Automation (RPA) frameworks, often requires domain expertise and lacks visibility and intuitive interfaces, making it challenging for users to fully leverage these technologies. This position paper argues for the emerging area of Human-Centered Automation (HCA), which prioritizes user needs and preferences in the design and development of automation systems. Drawing on empirical evidence from human-computer interaction research and case studies, we highlight the importance of considering user perspectives in automation and propose a framework for designing human-centric automation solutions. The paper discusses the limitations of existing automation approaches, the challenges in integrating AI and RPA, and the benefits of human-centered automation for productivity, innovation, and democratizing access to these technologies. We emphasize the importance of open-source solutions and provide examples of how HCA can empower individuals and organizations in the era of rapidly progressing AI, helping them remain competitive. The paper also explores pathways to achieve more advanced and context-aware automation solutions. We conclude with a call to action for researchers and practitioners to focus on developing automation technologies that adapt to user needs, provide intuitive interfaces, and leverage the capabilities of high-end AI to create a more accessible and user-friendly future of automation.
@misc{Toxtli2024Human,
doi = {10.48550/ARXIV.2405.15960},
url = {https://arxiv.org/abs/2405.15960},
author = {Toxtli, Carlos},
keywords = {Human-Computer Interaction (cs.HC), Artificial Intelligence (cs.AI), FOS: Computer and information sciences, FOS: Computer and information sciences},
title = {Human-Centered Automation},
publisher = {arXiv},
year = {2024},
copyright = {Creative Commons Attribution 4.0 International}
}This demo paper explores AI-enhanced multi-screen interaction within extended reality (XR) workspaces using called VirtuaScreens. The system facilitates the user-centric addition of customized virtual monitors and employs machine learning to understand user preferences for monitor arrangements and application arrangements. We employed Large language and multimedia models to offer context-sensitive feedback and enhance the user experience. We created this research framework called VirtuaScreens for researchers and practitioners to use to understand the interaction of the multiple screens in the XR environment.
@misc{Perera2024Exploring,
doi = {10.13140/RG.2.2.31604.36481},
url = {https://www.researchgate.net/doi/10.13140/RG.2.2.31604.36481},
author = {{Ravindu Perera} and Toxtli, Carlos},
language = {en},
title = {Exploring AI-Enhanced Multi-Screen Interaction in Extended Reality Workspaces},
publisher = {Unpublished},
year = {2024},
howpublished = {Preprint, ResearchGate},
note = {Preprint}
}The rapid development and adoption of Generative AI (GAI) technology in the form of chatbots such as ChatGPT and Claude has greatly increased interest in agentic machines. This paper introduces the Autonomous Cognitive Entity (ACE) model, a novel framework for a cognitive architecture, enabling machines and software agents to operate more independently. Drawing inspiration from the OSI model, the ACE framework presents layers of abstraction to conceptualize artificial cognitive architectures. The model is designed to harness the capabilities of the latest generative AI technologies, including large language models (LLMs) and multimodal generative models (MMMs), to build autonomous, agentic systems. The ACE framework comprises six layers: the Aspirational Layer, Global Strategy, Agent Model, Executive Function, Cognitive Control, and Task Prosecution. Each layer plays a distinct role, ranging from setting the moral compass and strategic thinking to task selection and execution. The ACE framework also incorporates mechanisms for handling failures and adapting actions, thereby enhancing the robustness and flexibility of autonomous agents. This paper introduces the conceptual framework and proposes implementation strategies that have been tested and observed in industry. The goal of this paper is to formalize this framework so as to be more accessible.
@article{Shapiro2023Conceptual,
doi = {10.48550/ARXIV.2310.06775},
url = {https://arxiv.org/abs/2310.06775},
author = {Shapiro, David and Li, Wangfan and Delaflor, Manuel and Toxtli, Carlos},
keywords = {Human-Computer Interaction (cs.HC), Artificial Intelligence (cs.AI), FOS: Computer and information sciences, FOS: Computer and information sciences, H.4.0},
title = {Conceptual Framework for Autonomous Cognitive Entities},
publisher = {arXiv},
year = {2023},
copyright = {Creative Commons Attribution 4.0 International}
}