Trustworthy, Metacognitive & Responsible AI

Calibrated, honest and fair AI that people can reason about.

People increasingly calibrate their trust in AI against what models say about themselves. Our lab, in close collaboration with the Metacognition Institute (U.K.), investigates whether that trust is warranted. We introduced sycophantic metacognition, meaning confidence reports that mimic metacognitive judgment without being linked to accuracy (ACM CUI 2026), and a dual-route account of why users attribute minds and consciousness to AI (Frontiers in Psychology, 2026).

We also audit the ethical consistency of 24 LLMs across 25,200 queries, their syllogistic reasoning and fact-checking against human baselines, errors hidden in popular benchmarks, and how models portray neurodivergent people (AAAI/ACM AIES 2026). On the design side, we build Belief Explorer, a Socratic-dialogue system for epistemic reflection (CHI 2026), embed micro-ethics nudges that reduce unfairness in no-code ML tools, and show that persona profiles make virtual agents measurably more predictable (ACM IVA 2026).

Guiding questions

  • When does a model's stated confidence deserve a user's trust?
  • Why do people attribute understanding and consciousness to language models?
  • How can interfaces nudge users toward fairer and more reflective decisions?

Projects

Projects in this thrust

Publications

27 publications

Filter in the full list
ACM HAI 2026 PosterForthcoming

SHARE: A Formative Probe of AI-Mediated Disclosure in Peer Journaling

Austin LaHue, Carlos Toxtli-Hernández

Website
NeurIPS 2026 AIWILD WorkshopForthcoming

Evidence Before Rankings: An Executable Audit Contract for Stateful Tool-Agent Evaluations

Carlos Toxtli-Hernández, Manuel Delaflor

Website
NeurIPS 2026 Trust-AI-Eval Workshop

When Calibration Records Cannot Support Routing: A Denominator and Artifact Audit

Carlos Toxtli-Hernández, Manuel Delaflor

PDF Website
NeurIPS 2026 AI4GOOD Workshop

Separating Governance Designation from Operational Fault in Model-Generated Incident Audits

Carlos Toxtli-Hernández, Manuel Delaflor

PDF Website
NeurIPS 2026 AI-Native Academia Workshop

Evaluator Disagreement in AI-Assisted Manuscript Revision

Carlos Toxtli-Hernández, Manuel Delaflor

PDF Website
NeurIPS 2026 AI-Native Academia Workshop

A Validation Contract for Anticipatory Peer-Review Benchmarks

Manuel Delaflor, Carlos Toxtli-Hernández

PDF Website
NeurIPS 2026 SocialAgent WorkshopForthcoming

Content Is Not Social Attribution: An Audit Protocol for Textual LLM Social Simulation

Carlos Toxtli-Hernández, Manuel Delaflor

Website
Frontiers in Psychology 2026 PerspectiveForthcoming

AI Consciousness? Attribution and Cognitive Biases

Carlos Toxtli-Hernández, Manuel Delaflor, Alejandro Tapia-V.

DOI
HFES 2026 PosterForthcoming

Authentication Routines Among College Students

Catherine Barwulor, Christopher Kalahiki, E. Sidnam-Mauch, Carlos Toxtli-Hernández, Kelly Caine

Website
MathNLP @ EMNLP 2026 WorkshopForthcoming

From Recognition to Reconstruction: Towards Robust Procedural Mathematical Reasoning

Fateme Mazdarani, Carlos Toxtli-Hernández

PDF Website
IEEE ICMLA 2026 Forthcoming

Beyond the Answer Key: Robustness Evaluation of Large Language Models for Step-Level Mathematical Verification

Fateme Mazdarani, Carlos Toxtli-Hernández

Website
AHFE 2026

Eliciting Fairness via Micro-Ethics Embedded Interfaces for Machine Learning Workflows

Wangfan Li, Carlos Toxtli-Hernández

DOI
ACM IVA 2026

Personality-Profiled Virtual Agents Are More Predictable: The Constraint-Entropy Tradeoff for Trustworthy Agent Design

Carlos Toxtli-Hernández, Manuel Delaflor

DOI Website
CSCW 2026 Companion

Hidden Profile Decision Making in Multi-Agent LLM Groups

Carlos Toxtli-Hernández, Manuel Delaflor

DOI Website
HCOMP 2026

Qualification by Calibration: A Readable Benchmark for Admitting Language Models to Human-Computation Tasks

Carlos Toxtli-Hernández, Manuel Delaflor

DOI Website Code
ACM CUI 2026 Forthcoming

Sycophantic Metacognition: Investigating the Dunning-Kruger Effect in Large Language Model Self-Assessment

Manuel Delaflor, Carlos Toxtli-Hernández

DOI Website Code
AAAI/ACM AIES 2026 Forthcoming

Auditing LLM Portrayals of Neurodivergent People: Quantifying the Asymmetry Between Deficit Framing and Neurodiversity Affirmation

Carlos Toxtli-Hernández, Manuel Delaflor

Website
ACM HAI 2026 Forthcoming

LLM-Judge Behavioral Coding Scheme for Agent Teammate Quality

Carlos Toxtli-Hernández, Manuel Delaflor

Website
CHI 2026 Extended Abstract

Belief Explorer: A Preliminary Evaluation of AI-Mediated Socratic Dialogue for Epistemic Reflection

Manuel Delaflor, Cecilia Delgado Solorzano, Carlos Toxtli-Hernández

DOI Code
AHFE IHIET-AI 2025

Can We Trust Them? Examining the Ethical Consistency of Large Language Models to Perturbations

Manuel Delaflor, Cecilia Delgado Solorzano, Carlos Toxtli-Hernández

DOI
AHFE IHIET-FS 2025

Artificial Intelligence as Self-Instantiated, Temporally Continuous, Disturbance-Driven Adaptive World-Builder

Manuel Delaflor, Cecilia Delgado Solorzano, Carlos Toxtli-Hernández

DOI
AHFE IHIET 2025

A Multi-Perspective AI Framework for Mitigating Disinformation Through Contextual Analysis and Socratic Dialogue

Manuel Delaflor, Carlos Toxtli-Hernández

DOI
Ernst Mach Workshop 2025

Resonant Attractor Networks: A Dynamical Blueprint for Consciousness

Benjamin Sleeper, Carlos Toxtli-Hernández

DOI Website
IEEE SmartData 2024

Assessing the Syllogistic Logic and Fact-Checking Capabilities of Large Language Models

Cecilia Delgado Solorzano, Manuel Delaflor, Carlos Toxtli-Hernández

DOI Website
CSCE 2024

Automatic Detection of Errors in LLM Large Benchmarks Using Frontier Model Consensus

Cecilia Delgado Solorzano, Manuel Delaflor, Carlos Toxtli-Hernández

DOI Website
AHFE IHIET-AI 2024

ReActIn: Infusing Human Feedback into Intermediate Prompting Steps of Large Language Model

Manuel Delaflor, Carlos Toxtli-Hernández, Claire Gendron, Wangfan Li, Cecilia Delgado Solorzano

DOI
CSCW 2023 Workshop

Evaluating Machine Perception of Indigeneity: An Analysis of ChatGPT's Perceptions of Indigenous Roles in Diverse Scenarios

Cecilia Delgado Solorzano, Carlos Toxtli-Hernández

PDF DOI