Automation & AgentsTrustworthy AIActive · 2023-present

Supervisory Control for LLMs and Autonomous Agents

Grounding AI oversight in the science of human supervisory control.

23,256
revision plans scored in RevPlan-Bench
10
venue-years of real peer reviews analyzed

As LLMs revise documents, operate computers and execute workflows, users need principled ways to plan, monitor and intervene. This line of work, led by Ph.D. student Wangfan Li, imports Supervisory Control Theory, which describes how humans oversee automated systems, into the design of LLM prompting and agent interfaces.

Prompt-Level Supervisory Alignment (PLSA) structures LLM revision as planning, monitoring and intervention, producing revisions with higher fidelity to real author revisions across ICLR, NeurIPS and CoRL review cycles (ACM CAIS 2026), and extends to anticipating concerns raised in later review rounds with RevPlan-Bench (HFES 2026). Earlier systems include a sandboxed virtual machine for human oversight of autonomous task execution (IEEE CAI 2025, ACM CHIWORK 2024), ReActIn for infusing human feedback into intermediate reasoning steps, LLM-generated BPMN workflows for robotic process automation, Requirement Bot for human-AI co-authoring of software requirements (ACM SAC 2026), and micro-ethics nudges that reduce unfair models in no-code ML tools.

Output

Publications

NeurIPS 2026 AI-Native Academia Workshop

Evaluator Disagreement in AI-Assisted Manuscript Revision

Carlos Toxtli-Hernández, Manuel Delaflor

PDF Website
NeurIPS 2026 AI-Native Academia Workshop

A Validation Contract for Anticipatory Peer-Review Benchmarks

Manuel Delaflor, Carlos Toxtli-Hernández

PDF Website
NeurIPS 2026 HAIC Workshop

How Many Agents Can One Supervisor Track? A Mechanistic Capacity Model and Measurement Protocol for LLM-Agent Teams

Carlos Toxtli-Hernández, Manuel Delaflor

PDF Website
NeurIPS 2026 HAIC Workshop

Headroom Before Protocol Effects: A Decision Procedure for Human-Agent Evaluation

Manuel Delaflor, Carlos Toxtli-Hernández

PDF Website
AHFE 2026

Eliciting Fairness via Micro-Ethics Embedded Interfaces for Machine Learning Workflows

Wangfan Li, Carlos Toxtli-Hernández

DOI
ACM CAIS 2026

Supervisory Control Theory for LLM Revision

Wangfan Li, Carlos Toxtli-Hernández

DOI
ACM SAC 2026

Requirement Bot: Enhancing Software User Requirement List Through Retrieval-Augmented Human-AI Co-Authoring

Wangfan Li, Carlos Toxtli-Hernández

DOI
HFES 2026

Prompt-Level Supervisory Alignment for Anticipatory Manuscript Revision

Wangfan Li, Sofia Abilene Campos Hernandez, Carlos Toxtli-Hernández

DOI Website
Manuscript 2025

Assessing AI-Generated Workflows: A Multi-Dimensional Evaluation Framework

Athish Venkatachalam, Carlos Toxtli-Hernández

Website
IEEE CAI 2025

Human Oversight Over Autonomous Task Execution in Sandbox Environments

Wangfan Li, Athish Venkatachalam, Jackson Bowen, Carlos Toxtli-Hernández

DOI Website
CSCE 2024

Infusing Human Feedback into Intermediate Prompting Steps of Large Language Models

Wangfan Li, Claire Gendron, Carlos Toxtli-Hernández

DOI Website
CSCE 2024

Automating Automation: Using LLMs to Generate BPMN Workflows for Robotic Process Automation

Carlos Toxtli-Hernández, Wangfan Li

DOI Website
ACM CHIWORK 2024 Demo

Human Oversight Mechanisms over Autonomous Task Execution in Sandbox Environments

Wangfan Li, Carlos Toxtli-Hernández

DOI Website
AHFE IHIET-AI 2024

ReActIn: Infusing Human Feedback into Intermediate Prompting Steps of Large Language Model

Manuel Delaflor, Carlos Toxtli-Hernández, Claire Gendron, Wangfan Li, Cecilia Delgado Solorzano

DOI
arXiv 2023

Conceptual Framework for Autonomous Cognitive Entities

David Shapiro, Wangfan Li, Manuel Delaflor, Carlos Toxtli-Hernández

PDF DOI