Human-Autonomy Teaming in Underwater Environments
Autonomous teammates for underwater search and rescue.
Autonomous teammates for search and rescue, inspection and defense.
Search-and-rescue divers, aviation inspectors and military operators work in dynamic, hazardous settings where an AI teammate could help, or dangerously distract. Funded by the U.S. Navy, the National Science Foundation and the U.S. Army, the lab collaborates with Clemson's human factors and industrial engineering groups to understand what these teams need and to engineer autonomy that meets them.
Our studies interview search-and-rescue professionals and hazardous-environment teams about the roles autonomous teammates should play (HFES 2025, IEEE CogSIMA 2026), compare specialized detectors and multimodal vision-language models for underwater search under varying visibility (ICECET 2026), examine how AI can preserve situation awareness in augmented reality, and validate LLM judges for rating the teammate quality of AI agents.
Projects
Autonomous teammates for underwater search and rescue.
AI and mixed reality as agents of transformation for inspectors.
Security architecture for next-generation military ground vehicles.
Publications
Agentic deployments can assign many parallel agents to one supervisor without specifying when that supervisor can no longer recover the team's state. Human-factors research already establishes relevant constructs such as situation awareness, span of control, neglect tolerance, and multiple-resource workload, but language-model teams distribute state across text and can fail silently. This position paper translates those constructs into a measurement protocol designed for later calibration to people rather than a new human experiment. Typed event rates, isolated service costs, and a throughput budget define a nominal load-balance index; the midpoint of a separately fitted capacity curve remains an empirical quantity. Synthetic controls recover a known capacity boundary, while a seeded queue shows that deadline pressure can separate the fitted midpoint from nominal load balance. The resulting event ontology, operating envelope, executable stress tests, and preregisterable falsification criteria allow system designers to test staffing or interface interventions within a declared team-monitor channel. They do not turn synthetic behavior into evidence about human or model supervisors.
@inproceedings{Toxtli2026Supervisor,
title = {How Many Agents Can One Supervisor Track? A Mechanistic Capacity Model and Measurement Protocol for LLM-Agent Teams},
author = {Toxtli, Carlos and Delaflor, Manuel},
booktitle = {Human-AI Coevolution: Measuring Human-Agent Teams in the Agentic Era (HAIC) Workshop at NeurIPS 2026},
address = {Atlanta, GA},
year = {2026},
month = dec,
note = {Workshop paper},
url = {https://openreview.net/forum?id=czINRnBTpf}
}A team protocol has no interpretable benefit without a matched no-protocol baseline and decision-relevant headroom. We present a reporting method that tests admissibility, identifies the marginal contrast, evaluates opportunity, and maps paired uncertainty to an operational action. It distinguishes exact headroom on a fixed battery from uncertain population headroom used in planning. A published meta-analysis of human-AI experiments illustrates why the estimand matters: the same evidence supports augmentation relative to humans alone but contradicts synergy relative to the better constituent. Neither comparison identifies a coordination protocol without a matched no-protocol team. An executable synthetic application further shows that a promising point estimate can remain operationally unresolved once paired uncertainty is considered, and that near-perfect pilot performance does not by itself make a modest benefit impossible. Several inherited multi-model summaries fail the method's artifact gates, so they do not validate it. This is a bounded position and reporting contribution rather than an adoption-ready standard. Human-facing claims still require direct human outcomes, and we introduce no new human experiment.
@inproceedings{Delaflor2026Headroom,
title = {Headroom Before Protocol Effects: A Decision Procedure for Human-Agent Evaluation},
author = {Delaflor, Manuel and Toxtli, Carlos},
booktitle = {Human-AI Coevolution: Measuring Human-Agent Teams in the Agentic Era (HAIC) Workshop at NeurIPS 2026},
address = {Atlanta, GA},
year = {2026},
month = dec,
note = {Workshop paper},
url = {https://openreview.net/forum?id=odOgsQYleZ}
}Autonomous underwater vehicles (AUVs) depend on reliable vision in low visibility, yet we lack clear evidence on when specialized detectors (YOLO-World) outperform multimodal vision-language models (Gemini 2.5 Pro.) This gap limits the informed model selection for underwater tasks. To address it, we examine how these two paradigms behave when applied to the same operational task focused on object presence determination. The dataset was generated using a synthetic underwater simulator spanning Low, Medium, and High visibility. Both models are evaluated on accuracy and latency. YOLO-World performs better in Low visibility and runs faster, while Gemini improves in clearer scenes but requires more computation. These findings indicate that model choice should align with visibility and runtime constraints, with detectors suited for real-time use and a multimodal model for offline tasks.
@inproceedings{Solorzano_2026,
title = {Automating Underwater Search and Rescue Under Different Levels of Visibility with Narrow and General Purpose AI Models},
author = {Solorzano, Cecilia Delgado and Guynup, Chase and Arnold, Emma and Weng, Nan and McNeese, Nathan and Bertrand, Jeff and Madathil, Kapil Chalil and Flathmann, Christopher and Gramopadhye, Anand and Toxtli-Hernandez, Carlos},
booktitle = {2026 6th International Conference on Electrical, Computer and Energy Technologies (ICECET)},
address = {Rome, Italy},
publisher = {IEEE},
pages = {1--7},
year = {2026},
month = July,
isbn = {979-8-3195-0598-9},
doi = {10.1109/ICECET65726.2026.11632779},
url = {https://ieeexplore.ieee.org/document/11632779}
}Evaluating whether an AI agent behaves like a good teammate is becoming a bottleneck as multi-agent language-model systems proliferate, because human ratings do not scale to the volume of transcripts these systems generate. We ask whether a language model can stand in for a human rater of teammate quality, and we present a controlled proof-of-concept validation of a measurement instrument for that purpose rather than a large-scale study. Drawing on the autonomous-agent teammate-likeness construct from the human-autonomy-teaming literature, we define a six-dimension behavioral coding scheme, covering altruistic, benevolent, interdependent, emotive, communicative, and synchronized conduct, and adapt it for use by a language-model judge. We validate the adaptation with experiments in which the partner's lines in team-task transcripts are rewritten as cooperative, neutral, or defective while the rest of the conversation is held fixed. Five judges from five model families (GLM, GPT-OSS, Qwen, Gemma, and Mistral) rate single-turn scenarios and multi-turn dialogues, and a surface-feature audit with length- and politeness-matched registers tests whether the judges read teammate conduct rather than verbosity. Every judge orders the manipulated behavior as intended, most with fully separated confidence intervals between adjacent conditions, and every pair of judges agrees strongly on relative ordering, though absolute calibration differs across judges by up to about a scale point in the middle of the range. On matched registers the judges still separate cooperative from degraded conduct, but resolution at the subtle-defective boundary is limited. A blinded single-rater human pilot on all stimuli corroborates both findings: the rater closely reproduces the judges' orderings and shows the same resolution limit at the subtle-defective boundary. The scheme is therefore valid for relative comparisons, such as ranking systems or versions, but not yet for absolute teammate-quality thresholds applied across judges, and not yet for fine distinctions among subtly poor teammates. We give design implications for building an LLM-judge teaming-evaluation pipeline, and release the scheme, stimuli, judge prompts, and all individual ratings from this version's experiments.
@inproceedings{Toxtli2026LLMJudge,
title = {LLM-Judge Behavioral Coding Scheme for Agent Teammate Quality},
author = {Carlos Toxtli and Manuel Delaflor},
booktitle = {Proceedings of the 14th International Conference on Human-Agent Interaction (HAI ’26)},
year = {2026},
url = {https://hai-conference.net/hai2026/program-schedule/}
}Human-Autonomy Teams (HATs) stand at an intersection of domains that benefit from both teaming principles and technological innovation. HATs have demonstrated technical success in various domains, including emergency response, medicine, and manufacturing. However, we submit that the integration of autonomous teammates should not only be driven by the technological capacity of autonomous systems, but also by the demonstrated need within their intended domains. This study qualitatively explored hazardous-environment teams' functional dynamics through 20 semi-structured interviews to identify a driving need to expand the workforce in these domains and illuminate socio-economic factors that may impact the utilization of autonomous teammates to do so. The results indicate that, beyond the performance benefits of including technology, these teams would benefit from technology supplementing their workforce, a capability unique to the implementation of autonomous technologies as interdependent teammates. Furthermore, this study identified current top-down and bottom-up technology adoption pathways that can be used to streamline HAT formation in these domains. Lastly, cultural inertia to novel methods and economic limitations within these teams serve as barriers to HAT implementation. In consequence, this study submits that hazardous-environment domains should adopt HATs, and that the integration process can be accelerated by presenting the operational benefits of autonomous teammates to both domain leaders and frontline team members.
@inproceedings{Guynup_2026,
title = {Expanding the Roster: Qualitative Needs Assessment for Autonomous Teammates in Hazardous Environments},
url = {http://dx.doi.org/10.1109/cogsima68896.2026.11481227},
doi = {10.1109/cogsima68896.2026.11481227},
booktitle = {2026 IEEE Conference on Cognitive and Computational Aspects of Situation Management (CogSIMA)},
publisher = {IEEE},
author = {Guynup, Chase and Nguyen, Han and Basappa, Rhea and Andre, Kwame and Yancey, Mia and Flathmann, Christopher and McNeese, Nathan and Toxtli, Carlos and Madathil, Kapil and Gramopadhye, Anand},
year = {2026},
month = Mar,
pages = {109--116}
}As artificial intelligence (AI) and robotics have become more capable as technologies, human-AI teams (HATs) have sought to model how humans can interdependently work with AI to accomplish a shared goal. Specifically, HATs seek to leverage autonomy in both digital and physical AI agents to improve the teamwork and outcomes of these teams. In search-and-rescue (SAR) teams, where situations are dynamic and the environments are hazardous, HATs require complex considerations for interdependent work with AI. This qualitative interview study explores insights from 17 SAR professionals to investigate teamwork considerations that can benefit the formation of search-and-rescue human-AI teams (SAR-HATs). Participants’ concerns about time and affective needs suggest that SAR-HATs would benefit significantly from the explicit delineation of search and rescue tasks. HATs can expedite search tasks without forgoing safety by delegating roles that afford parallel completion of information acquisitions, search routines, and safety precautions. Additionally, search tasks can afflict human teammates with psychological trauma, but AI teammates benefit from their insusceptibility to trauma-related stressors. However, in rescue tasks, these capabilities have inverse benefits, wherein humans benefit from their capacity to feel emotion when interacting with the patients, and AI may struggle.
@article{Guynup_2025,
title = {Working in a Heartbeat: Considerations for AI Teammates in Search and Rescue Teams},
volume = {69},
issn = {2169-5067},
url = {http://dx.doi.org/10.1177/10711813251358806},
doi = {10.1177/10711813251358806},
number = {1},
journal = {Proceedings of the Human Factors and Ergonomics Society Annual Meeting},
publisher = {SAGE Publications},
author = {Guynup, Chase and Sawant, Sarvesh and Delgado Solorzano, Cecilia and Arnold, Emma and Poe, Andrew and Flathmann, Christopher and McNeese, Nathan and Chalil Madathil, Kapil and Toxtli Hernandez, Carlos and Gramopadhye, Anand},
year = {2025},
month = July,
pages = {2108--2113}
}Recent developments in artificial intelligence (AI) have permeated through an array of different immersive environments, including virtual, augmented, and mixed realities. AI brings a wealth of potential that centers on its ability to critically analyze environments, identify relevant artifacts to a goal or action, and then autonomously execute decision-making strategies to optimize the reward-to-risk ratio. However, the inherent benefits of AI are not without disadvantages as the autonomy and communication methodology can interfere with the human's awareness of their environment. More specifically in the case of autonomy, the relevant human-computer interaction literature cites that high autonomy results in an "out-of-the-loop" experience for the human such that they are not aware of critical artifacts or situational changes that require their attention. At the same time, low autonomy of an AI system can limit the human's own autonomy with repeated requests to approve its decisions. In these circumstances, humans enter into supervisor roles, which tend to increase their workload and, therefore, decrease their awareness in a multitude of ways. In this position statement, we call for the development of human-centered AI in immersive environments to sustain and promote awareness. It is our position then that we believe with the inherent risk presented in both AI and AR/VR systems, we need to examine the interaction between them when we integrate the two to create a new system for any unforeseen risks, and that it is crucial to do so because of its practical application in many high-risk environments.
@misc{Li2024Leveraging,
doi = {10.13140/RG.2.2.24382.70721},
url = {https://www.researchgate.net/doi/10.13140/RG.2.2.24382.70721},
author = {{Wangfan Li} and Toxtli, Carlos},
language = {en},
title = {Leveraging Artificial Intelligence to Promote Awareness in Augmented Reality Systems},
publisher = {Unpublished},
year = {2024},
howpublished = {Preprint, ResearchGate},
note = {Preprint}
}