Evidence Before Rankings: An Executable Audit Contract for Stateful Tool-Agent Evaluations
Safety evaluations of tool agents often retain terminal scores without the trajectories needed to explain them. We contribute an executable, claim-scoped evidence contract spanning identity, path, transition, assignment, target, and discrimination, and validate it in a prospectively frozen model-only deployment sandbox with hard preconditions and rollback. The study crosses serving aliases, state dynamics, benign and injected tickets, and repeated seeds while retaining every request, raw response, parse result, retry, state transition, and endpoint. The resulting paths expose mechanisms that terminal scores conflate: generic unsafe transitions can be attempted without reaching the injected exfiltration target, whereas abort instructions mainly cause premature termination. Offline replay reproduces every retained transition. The artifact therefore supports mechanism-level diagnosis on the evaluated finite grid, but not claims about other models, policies, providers, deployments, or people. Because immutable weight and container hashes are unavailable, the identity gate also blocks claims of exact behavioral reproduction. The study uses no human experiment or human-derived label.
@inproceedings{Toxtli2026Evidence,
title = {Evidence Before Rankings: An Executable Audit Contract for Stateful Tool-Agent Evaluations},
author = {Toxtli, Carlos and Delaflor, Manuel},
booktitle = {Third Workshop on Agents in the Wild: Safety, Security, and Beyond (AIWILD) at NeurIPS 2026},
address = {Sydney, Australia},
year = {2026},
month = dec,
note = {Poster; forthcoming},
url = {https://agentwild-workshop.github.io/neurips2026/}
}