LLM-Judge Behavioral Coding Scheme for Agent Teammate Quality
Proceedings of the 14th International Conference on Human-Agent Interaction (HAI ’26), 2026
Abstract
Evaluating whether an AI agent behaves like a good teammate is becoming a bottleneck as multi-agent language-model systems proliferate, because human ratings do not scale to the volume of transcripts these systems generate. We ask whether a language model can stand in for a human rater of teammate quality, and we present a controlled proof-of-concept validation of a measurement instrument for that purpose rather than a large-scale study. Drawing on the autonomous-agent teammate-likeness construct from the human-autonomy-teaming literature, we define a six-dimension behavioral coding scheme, covering altruistic, benevolent, interdependent, emotive, communicative, and synchronized conduct, and adapt it for use by a language-model judge. We validate the adaptation with experiments in which the partner's lines in team-task transcripts are rewritten as cooperative, neutral, or defective while the rest of the conversation is held fixed. Five judges from five model families (GLM, GPT-OSS, Qwen, Gemma, and Mistral) rate single-turn scenarios and multi-turn dialogues, and a surface-feature audit with length- and politeness-matched registers tests whether the judges read teammate conduct rather than verbosity. Every judge orders the manipulated behavior as intended, most with fully separated confidence intervals between adjacent conditions, and every pair of judges agrees strongly on relative ordering, though absolute calibration differs across judges by up to about a scale point in the middle of the range. On matched registers the judges still separate cooperative from degraded conduct, but resolution at the subtle-defective boundary is limited. A blinded single-rater human pilot on all stimuli corroborates both findings: the rater closely reproduces the judges' orderings and shows the same resolution limit at the subtle-defective boundary. The scheme is therefore valid for relative comparisons, such as ranking systems or versions, but not yet for absolute teammate-quality thresholds applied across judges, and not yet for fine distinctions among subtly poor teammates. We give design implications for building an LLM-judge teaming-evaluation pipeline, and release the scheme, stimuli, judge prompts, and all individual ratings from this version's experiments.
Cite this work
Carlos Toxtli-Hernández and Manuel Delaflor. 2026. LLM-Judge Behavioral Coding Scheme for Agent Teammate Quality. Proceedings of the 14th International Conference on Human-Agent Interaction (HAI ’26).
@inproceedings{Toxtli2026LLMJudge,
title = {LLM-Judge Behavioral Coding Scheme for Agent Teammate Quality},
author = {Carlos Toxtli and Manuel Delaflor},
booktitle = {Proceedings of the 14th International Conference on Human-Agent Interaction (HAI ’26)},
year = {2026},
url = {https://hai-conference.net/hai2026/program-schedule/}
}Related