Abstract
When a language model drafts an incident audit, it may conflate the party that caused a failure with the party a governance charter makes answerable for it. We test that distinction in a construct-separated, artifact-complete, model-only audit spanning fixed incidents, several governance descriptions, and three model families. The models consistently track the stipulated source of operational fault while assigning greater governance accountability to the explicitly designated bearer. Changing that bearer moves governance ratings in the prespecified direction without collapsing them into causal contribution or operational fault. Absolute AI designations reach a ceiling, however, so a frozen shared-versus-sole extension distinguishes a dose response for two model families but not for the third. Every request, response, failure, retry, parse, and protocol deviation is retained without repair or renormalization. The audit therefore shows how to test whether model-generated reports preserve a governance distinction under a declared rubric. It does not establish moral judgment, human agreement, legal validity, or deployment benefit.
Cite this work
Carlos Toxtli-Hernández and Manuel Delaflor. 2026. Separating Governance Designation from Operational Fault in Model-Generated Incident Audits. Trustworthy AI for Good Workshop (AI4GOOD) at NeurIPS 2026.
@inproceedings{Toxtli2026Governance,
title = {Separating Governance Designation from Operational Fault in Model-Generated Incident Audits},
author = {Toxtli, Carlos and Delaflor, Manuel},
booktitle = {Trustworthy AI for Good Workshop (AI4GOOD) at NeurIPS 2026},
address = {Paris, France},
year = {2026},
month = dec,
note = {Poster},
url = {https://openreview.net/forum?id=oZDf3mJY4F}
}Related