Trustworthy AIAutomation & AgentsActive · 2023-present

Metacognition, Calibration and Trust in Language Models

Do language models know what they know, and should we believe them?

24
LLMs audited for ethical consistency
25,200
queries in our perturbation study
6
LLMs tested for sycophantic metacognition

In partnership with the Metacognition Institute, the lab investigates the gap between how language models present their reasoning and how they actually behave. We introduced sycophantic metacognition: confidence reports that track the surface form of authority rather than accuracy, tested across six LLMs, four domains and multiple pressure conditions (ACM CUI 2026). A companion perspective explains why users nonetheless attribute understanding and even consciousness to these systems (Frontiers in Psychology, 2026).

The project also audits LLM ethical consistency under perturbation (25,200 queries across 24 models), syllogistic reasoning and fact-checking against human participants, errors in the MMLU benchmark detected by frontier-model consensus, readable qualification tests for model workers (HCOMP 2026), portrayals of neurodivergent people (AIES 2026), and the behavior of multi-agent LLM groups in classic hidden-profile decision tasks. On the design side, Belief Explorer uses Socratic dialogue and multi-perspective analysis to help people examine their own beliefs (CHI 2026).

Output

Publications

NeurIPS 2026 AIWILD WorkshopForthcoming

Evidence Before Rankings: An Executable Audit Contract for Stateful Tool-Agent Evaluations

Carlos Toxtli-Hernández, Manuel Delaflor

Website
NeurIPS 2026 Trust-AI-Eval Workshop

When Calibration Records Cannot Support Routing: A Denominator and Artifact Audit

Carlos Toxtli-Hernández, Manuel Delaflor

PDF Website
NeurIPS 2026 AI4GOOD Workshop

Separating Governance Designation from Operational Fault in Model-Generated Incident Audits

Carlos Toxtli-Hernández, Manuel Delaflor

PDF Website
NeurIPS 2026 SocialAgent WorkshopForthcoming

Content Is Not Social Attribution: An Audit Protocol for Textual LLM Social Simulation

Carlos Toxtli-Hernández, Manuel Delaflor

Website
Frontiers in Psychology 2026 PerspectiveForthcoming

AI Consciousness? Attribution and Cognitive Biases

Carlos Toxtli-Hernández, Manuel Delaflor, Alejandro Tapia-V.

DOI
ACM IVA 2026

Personality-Profiled Virtual Agents Are More Predictable: The Constraint-Entropy Tradeoff for Trustworthy Agent Design

Carlos Toxtli-Hernández, Manuel Delaflor

DOI Website
CSCW 2026 Companion

Hidden Profile Decision Making in Multi-Agent LLM Groups

Carlos Toxtli-Hernández, Manuel Delaflor

DOI Website
HCOMP 2026

Qualification by Calibration: A Readable Benchmark for Admitting Language Models to Human-Computation Tasks

Carlos Toxtli-Hernández, Manuel Delaflor

DOI Website Code
ACM CUI 2026 Forthcoming

Sycophantic Metacognition: Investigating the Dunning-Kruger Effect in Large Language Model Self-Assessment

Manuel Delaflor, Carlos Toxtli-Hernández

DOI Website Code
AAAI/ACM AIES 2026 Forthcoming

Auditing LLM Portrayals of Neurodivergent People: Quantifying the Asymmetry Between Deficit Framing and Neurodiversity Affirmation

Carlos Toxtli-Hernández, Manuel Delaflor

Website
ACM HAI 2026 Forthcoming

LLM-Judge Behavioral Coding Scheme for Agent Teammate Quality

Carlos Toxtli-Hernández, Manuel Delaflor

Website
CHI 2026 Extended Abstract

Belief Explorer: A Preliminary Evaluation of AI-Mediated Socratic Dialogue for Epistemic Reflection

Manuel Delaflor, Cecilia Delgado Solorzano, Carlos Toxtli-Hernández

DOI Code
AHFE IHIET-AI 2025

Can We Trust Them? Examining the Ethical Consistency of Large Language Models to Perturbations

Manuel Delaflor, Cecilia Delgado Solorzano, Carlos Toxtli-Hernández

DOI
AHFE IHIET-FS 2025

Artificial Intelligence as Self-Instantiated, Temporally Continuous, Disturbance-Driven Adaptive World-Builder

Manuel Delaflor, Cecilia Delgado Solorzano, Carlos Toxtli-Hernández

DOI
AHFE IHIET 2025

A Multi-Perspective AI Framework for Mitigating Disinformation Through Contextual Analysis and Socratic Dialogue

Manuel Delaflor, Carlos Toxtli-Hernández

DOI
Ernst Mach Workshop 2025

Resonant Attractor Networks: A Dynamical Blueprint for Consciousness

Benjamin Sleeper, Carlos Toxtli-Hernández

DOI Website
IEEE SmartData 2024

Assessing the Syllogistic Logic and Fact-Checking Capabilities of Large Language Models

Cecilia Delgado Solorzano, Manuel Delaflor, Carlos Toxtli-Hernández

DOI Website
CSCE 2024

Automatic Detection of Errors in LLM Large Benchmarks Using Frontier Model Consensus

Cecilia Delgado Solorzano, Manuel Delaflor, Carlos Toxtli-Hernández

DOI Website