Abstract
Low calibration error does not establish that verbalized confidence can support routing. We audit the evidence needed for that decision in two retained records. The first is nearly perfectly accurate and has low aggregate calibration error, but it contains only one observed error and its joint answer-confidence coverage varies sharply by endpoint. It can describe average confidence on this finite record, but cannot establish discrimination or selective risk. The second record omits item text, answers, gold labels, raw outputs, and verification evidence. Reproducible arithmetic on its supplied correctness field therefore does not become a verified behavioral estimate. Together, the cases change the operational decision from choosing a confidence threshold to redesigning the items, interface, and evidence record before any routing evaluation. We provide a compact machine-readable audit covering denominators, artifact status, calibration, discrimination, selective risk, uncertainty, and unidentified outcomes. The contribution is a bounded workflow for deciding what a calibration record can support, not a validated benchmark or a general claim about verbalized confidence.
Cite this work
Carlos Toxtli-Hernández and Manuel Delaflor. 2026. When Calibration Records Cannot Support Routing: A Denominator and Artifact Audit. TAE (Trust-AI-Eval): Can We Trust AI Evaluation? Workshop at NeurIPS 2026.
@inproceedings{Toxtli2026Calibration,
title = {When Calibration Records Cannot Support Routing: A Denominator and Artifact Audit},
author = {Toxtli, Carlos and Delaflor, Manuel},
booktitle = {TAE (Trust-AI-Eval): Can We Trust AI Evaluation? Workshop at NeurIPS 2026},
address = {Sydney, Australia},
year = {2026},
month = dec,
note = {Poster},
url = {https://openreview.net/forum?id=uPtp2butrm}
}Related