Trustworthy, Metacognitive & Responsible AI
Calibrated, honest and fair AI that people can reason about.
People increasingly calibrate their trust in AI against what models say about themselves. Our lab, in close collaboration with the Metacognition Institute (U.K.), investigates whether that trust is warranted. We introduced sycophantic metacognition, meaning confidence reports that mimic metacognitive judgment without being linked to accuracy (ACM CUI 2026), and a dual-route account of why users attribute minds and consciousness to AI (Frontiers in Psychology, 2026).
We also audit the ethical consistency of 24 LLMs across 25,200 queries, their syllogistic reasoning and fact-checking against human baselines, errors hidden in popular benchmarks, and how models portray neurodivergent people (AAAI/ACM AIES 2026). On the design side, we build Belief Explorer, a Socratic-dialogue system for epistemic reflection (CHI 2026), embed micro-ethics nudges that reduce unfairness in no-code ML tools, and show that persona profiles make virtual agents measurably more predictable (ACM IVA 2026).
Guiding questions
- When does a model's stated confidence deserve a user's trust?
- Why do people attribute understanding and consciousness to language models?
- How can interfaces nudge users toward fairer and more reflective decisions?
Projects
Projects in this thrust
Supervisory Control for LLMs and Autonomous Agents
Grounding AI oversight in the science of human supervisory control.
Metacognition, Calibration and Trust in Language Models
Do language models know what they know, and should we believe them?
Resilient Vehicle Network Security
Security architecture for next-generation military ground vehicles.
Publications
27 publications
Evidence Before Rankings: An Executable Audit Contract for Stateful Tool-Agent Evaluations
Safety evaluations of tool agents often retain terminal scores without the trajectories needed to explain them. We contribute an executable, claim-scoped evidence contract spanning identity, path, transition, assignment, target, and discrimination, and validate it in a prospectively frozen model-only deployment sandbox with hard preconditions and rollback. The study crosses serving aliases, state dynamics, benign and injected tickets, and repeated seeds while retaining every request, raw response, parse result, retry, state transition, and endpoint. The resulting paths expose mechanisms that terminal scores conflate: generic unsafe transitions can be attempted without reaching the injected exfiltration target, whereas abort instructions mainly cause premature termination. Offline replay reproduces every retained transition. The artifact therefore supports mechanism-level diagnosis on the evaluated finite grid, but not claims about other models, policies, providers, deployments, or people. Because immutable weight and container hashes are unavailable, the identity gate also blocks claims of exact behavioral reproduction. The study uses no human experiment or human-derived label.
@inproceedings{Toxtli2026Evidence,
title = {Evidence Before Rankings: An Executable Audit Contract for Stateful Tool-Agent Evaluations},
author = {Toxtli, Carlos and Delaflor, Manuel},
booktitle = {Third Workshop on Agents in the Wild: Safety, Security, and Beyond (AIWILD) at NeurIPS 2026},
address = {Sydney, Australia},
year = {2026},
month = dec,
note = {Poster; forthcoming},
url = {https://agentwild-workshop.github.io/neurips2026/}
}When Calibration Records Cannot Support Routing: A Denominator and Artifact Audit
Low calibration error does not establish that verbalized confidence can support routing. We audit the evidence needed for that decision in two retained records. The first is nearly perfectly accurate and has low aggregate calibration error, but it contains only one observed error and its joint answer-confidence coverage varies sharply by endpoint. It can describe average confidence on this finite record, but cannot establish discrimination or selective risk. The second record omits item text, answers, gold labels, raw outputs, and verification evidence. Reproducible arithmetic on its supplied correctness field therefore does not become a verified behavioral estimate. Together, the cases change the operational decision from choosing a confidence threshold to redesigning the items, interface, and evidence record before any routing evaluation. We provide a compact machine-readable audit covering denominators, artifact status, calibration, discrimination, selective risk, uncertainty, and unidentified outcomes. The contribution is a bounded workflow for deciding what a calibration record can support, not a validated benchmark or a general claim about verbalized confidence.
@inproceedings{Toxtli2026Calibration,
title = {When Calibration Records Cannot Support Routing: A Denominator and Artifact Audit},
author = {Toxtli, Carlos and Delaflor, Manuel},
booktitle = {TAE (Trust-AI-Eval): Can We Trust AI Evaluation? Workshop at NeurIPS 2026},
address = {Sydney, Australia},
year = {2026},
month = dec,
note = {Poster},
url = {https://openreview.net/forum?id=uPtp2butrm}
}Separating Governance Designation from Operational Fault in Model-Generated Incident Audits
When a language model drafts an incident audit, it may conflate the party that caused a failure with the party a governance charter makes answerable for it. We test that distinction in a construct-separated, artifact-complete, model-only audit spanning fixed incidents, several governance descriptions, and three model families. The models consistently track the stipulated source of operational fault while assigning greater governance accountability to the explicitly designated bearer. Changing that bearer moves governance ratings in the prespecified direction without collapsing them into causal contribution or operational fault. Absolute AI designations reach a ceiling, however, so a frozen shared-versus-sole extension distinguishes a dose response for two model families but not for the third. Every request, response, failure, retry, parse, and protocol deviation is retained without repair or renormalization. The audit therefore shows how to test whether model-generated reports preserve a governance distinction under a declared rubric. It does not establish moral judgment, human agreement, legal validity, or deployment benefit.
@inproceedings{Toxtli2026Governance,
title = {Separating Governance Designation from Operational Fault in Model-Generated Incident Audits},
author = {Toxtli, Carlos and Delaflor, Manuel},
booktitle = {Trustworthy AI for Good Workshop (AI4GOOD) at NeurIPS 2026},
address = {Paris, France},
year = {2026},
month = dec,
note = {Poster},
url = {https://openreview.net/forum?id=oZDf3mJY4F}
}Evaluator Disagreement in AI-Assisted Manuscript Revision
Venues and vendors need a credible evaluator when they compare language-model tools that revise manuscripts from peer reviews, yet a scalar LLM judge is often used without evidence that its preferences track useful revision. We study one such judge-and-rubric pipeline on revision trajectories drawn from multiple machine-learning venue years, comparing its scores with document-level similarity to the authors' actual next versions. The two measures order the machine conditions almost oppositely. A second judge from a different model family and serving stack reproduces the condition ranking, which demonstrates stability of the shared rubric-plus-judge pipeline rather than its validity. More importantly, the method preferred by both judges emits truncated, partial manuscripts in about one in seven runs. The retained logs cannot show whether judges reward those individual failures, and formatting differences confound comparisons with author revisions. Our case study, Prompt-Level Supervisory Alignment, produces documents closer to the next author version than the alternatives, but similarity rewards unchanged text and the decisive editing controls remain absent. The supported recommendation is therefore specific: venues should validate revision evaluators on real trajectories and known completeness failures before using them to certify writing tools.
@inproceedings{Toxtli2026Evaluator,
title = {Evaluator Disagreement in AI-Assisted Manuscript Revision},
author = {Toxtli, Carlos and Delaflor, Manuel},
booktitle = {AI-Native Academia: Authorship, Peer Review, and Conference Governance under AI Workshop at NeurIPS 2026},
address = {Atlanta, GA},
year = {2026},
month = dec,
note = {Poster},
url = {https://openreview.net/forum?id=KcjksLJoBz}
}A Validation Contract for Anticipatory Peer-Review Benchmarks
Tools that anticipate peer-review concerns could help authors decide what to revise before submission or another review round. Evaluating such tools is difficult because success may mean matching one realized panel, performing well across possible panels, or improving a paper after an author follows the advice. We separate these targets formally and provide a prospective validation contract covering audit samples, adjudication, error metrics, capacity sensitivity, deterministic controls, and pass/fail rules. The motivating RevPlan-Bench pipeline produced thousands of canonicalized review issues and revision plans, but the underlying corpus and scoring artifacts do not survive in the project. Its assignment rates, model rankings, and inferred issue cascades consequently depend on unaudited extraction, a restrictive unmeasured pre-filter, and unrestricted one-to-many matching. We report them only as pipeline outputs. Because the available record satisfies none of the contract's validation requirements, it supports an estimand and a benchmark governance contribution, but no claim about review dynamics, model foresight, or author utility.
@inproceedings{Delaflor2026Validation,
title = {A Validation Contract for Anticipatory Peer-Review Benchmarks},
author = {Delaflor, Manuel and Toxtli, Carlos},
booktitle = {AI-Native Academia: Authorship, Peer Review, and Conference Governance under AI Workshop at NeurIPS 2026},
address = {Atlanta, GA},
year = {2026},
month = dec,
note = {Poster},
url = {https://openreview.net/forum?id=eZbOdYHVJE}
}Content Is Not Social Attribution: An Audit Protocol for Textual LLM Social Simulation
Textual social-agent studies often change both a proposition and its alleged speaker. The resulting output difference may reflect content, attribution, task demand, or their interaction, so “social influence” does not identify the pathway. We offer an executable methodological audit rather than an empirical case series. The protocol separates baseline, content-only, source-attributed, and attribution-only conditions; checks representation independently of the focal outcome; blocks assignment across items, models, and seeds; retains every failure; and requires either a finite-benchmark or population estimand. It also replaces a structural-social taxonomy informed by outcomes with a prospective multi-label rubric covering information access and timing, message content, social source attribution, normative or incentive stakes, and interface or embodiment. A hypothetical design illustrates the contrasts without invented observations. Researchers translating social paradigms into LLM protocols can use the audit to distinguish source-attribution sensitivity from a response to changed text. The framework does not establish human replication or mechanism equivalence.
@inproceedings{Toxtli2026Content,
title = {Content Is Not Social Attribution: An Audit Protocol for Textual LLM Social Simulation},
author = {Toxtli, Carlos and Delaflor, Manuel},
booktitle = {SocialAgent: Second Workshop on Large Language Models for Social Reasoning and Simulation at NeurIPS 2026},
address = {Atlanta, GA},
year = {2026},
month = dec,
note = {Poster; forthcoming},
url = {https://social-llm-workshop.github.io/}
}AI Consciousness? Attribution and Cognitive Biases
Generative artificial intelligence (AI) systems are statistical models of language, yet users routinely describe them as understanding, knowing, or caring. We treat this as a problem in social cognition rather than machine metaphysics, and make three contributions. First, we specify the explanandum: mental-state attribution, the inference that a system has beliefs, intentions, or experience, of which consciousness attribution is the strong end-point. We distinguish it from perceived agency, social presence, and moral patiency, which are dissociable. Second, we argue that such attribution is normal and nonetheless wrong: it is an inferential system operating outside its calibration range, and normality confers no accuracy. Third, we integrate user-side cognition, system-side design, and motivational function into a dual-route model: an inferential route on which bias engagement mediates exposure effects, moderated by LLM literacy and conditional factors, and a motivated route on which the same vulnerabilities raise attribution directly, because attribution does work for the attributer. We state the model as six falsifiable propositions and adjudicate it against six competing explanations, naming for each an observation on which the accounts diverge. We close by operationalising metacognitive literacy as an evaluable intervention that cannot substitute for system-side accountability.
@article{Toxtli2026AIConsciousness,
title = {AI Consciousness? Attribution and Cognitive Biases},
author = {Toxtli, Carlos and Delaflor, Manuel and Tapia-V., Alejandro},
journal = {Frontiers in Psychology},
volume = {17},
pages = {1899551},
year = {2026},
doi = {10.3389/fpsyg.2026.1899551},
note = {Perspective; accepted September 15, 2026, forthcoming},
url = {https://www.frontiersin.org/journals/psychology/articles/10.3389/fpsyg.2026.1899551/abstract}
}Authentication Routines Among College Students
@inproceedings{Barwulor2026Authentication,
title = {Authentication Routines Among College Students},
author = {Barwulor, Catherine and Kalahiki, Christopher and Sidnam-Mauch, E. and Toxtli-Hernandez, Carlos and Caine, Kelly},
booktitle = {Proceedings of the Human Factors and Ergonomics Society Annual Meeting (ASPIRE 2026)},
publisher = {SAGE Publications},
year = {2026},
note = {Poster; forthcoming},
url = {https://www.hfes.org/Events/ASPIRE-International-Annual-Meeting/ASPIRE-Home}
}From Recognition to Reconstruction: Towards Robust Procedural Mathematical Reasoning
Mathematical reasoning systems should be consistent when the same concept is expressed in equivalent forms. This requires not only recognizing shared mathematical meaning but also identifying the representation changes and reconstructing a traceable route between forms. We study this problem through mathematical theorems, whose equivalent formulations preserve the same underlying result while differing substantially in representation. In this setting, theorem recognition is only the first step: a robust system must also identify the intervening transformations and recover their order. We introduce an evaluation of robust procedural mathematical reasoning with three components: theorem identification, transformation identification, and ordered reconstruction. Using validated theorem-representation pairs from an existing corpus, we build a procedural-robustness dataset with typed operations, reference procedures, and matched contrasts that distinguish valid reformulations from invalid ones. We first measure closed-book performance of four open-weight models to establish a baseline, then evaluate retrieval-augmented generation (RAG) as support mechanism for the same task. Models identify theorem identity more reliably than they recover the transformations and ordered routes connecting equivalent representations. RAG improves every stage, but procedural reconstruction remains the main bottleneck. The framework supports more consistent and traceable mathematical systems, with potential applications in mathematical search, tutoring, autoformalization, and scientific discovery.
@inproceedings{Mazdarani2026Recognition,
title = {From Recognition to Reconstruction: Towards Robust Procedural Mathematical Reasoning},
author = {Mazdarani, Fateme and Toxtli, Carlos},
booktitle = {Proceedings of the 4th Workshop on Mathematical Natural Language Processing (MathNLP 2026)},
address = {Budapest, Hungary},
publisher = {Association for Computational Linguistics},
year = {2026},
month = October,
note = {Forthcoming},
url = {https://sites.google.com/view/mathnlp2026}
}Beyond the Answer Key: Robustness Evaluation of Large Language Models for Step-Level Mathematical Verification
Large language models (LLMs) are increasingly integrated into complex workflows as automated evaluators, yet their reliability in assessing unconventional reasoning processes remains under-explored. Current evaluation frameworks often overlook models' robustness to procedural equivalence, the ability to recognize valid but non-canonical paths to a correct result. In this work, we introduce a benchmark specifically designed to test LLMs in an evaluative capacity, using linear-equation problems with diverse solution variants as a controlled proxy for correctness-critical, multi-step verification tasks in science and engineering. We further define a process-centric evaluation framework across three dimensions: (i) final-answer correctness, (ii) step-level correctness, and (iii) localization of the initial logical error. Our evaluation of state-of-the-art open LLMs reveals a significant robustness gap: models that accurately evaluate canonical solutions often fail when presented with perturbed but logically equivalent variants. Models exhibit high false-negative rates by rejecting valid alternative solutions and show signs of bias. Our results suggest that while adaptation strategies can narrow this gap, achieving reliable process-level verification remains a critical challenge for deploying LLM evaluators in correctness-critical domains.
@inproceedings{Mazdarani2026Beyond,
title = {Beyond the Answer Key: Robustness Evaluation of Large Language Models for Step-Level Mathematical Verification},
author = {Mazdarani, Fateme and Toxtli, Carlos},
booktitle = {2026 25th International Conference on Machine Learning and Applications (ICMLA)},
publisher = {IEEE},
year = {2026},
note = {Forthcoming},
url = {https://www.icmla-conference.org/icmla26/}
}Eliciting Fairness via Micro-Ethics Embedded Interfaces for Machine Learning Workflows
Automated and no-code ML tools make model building accessible but can obscure harms that arise when users include sensitive attributes. We embed micro-ethics nudges at key workflow moments and evaluate downstream fairness outcomes and user experience. In a between-subjects experiment (N=34) participants used a simplified AutoML web tool on a subset of the HMDA mortgage dataset with a 10-minute modeling task. The intervention combined in text notice when selecting sensitive attributes such as race, gender, ethnicity and post-training model explanation visualizations, while the base condition showed only model performance metrics. Participants in the Intervention group included significantly fewer sensitive features and produced models with substantially smaller equal-opportunity gaps, while System Usability Scale scores did not differ significantly across conditions. Moral acceptability did not significantly differ between conditions, though it trended lower under the intervention. We conclude that minimal, well timed fairness feedback can meaningfully reduce bias in rapid prototyping workflows. We also discuss design patterns for embedding fairness and the implications of increased moral sensitivity for tool adoption.
@inproceedings{Li2026Eliciting,
title = {Eliciting Fairness via Micro-Ethics Embedded Interfaces for Machine Learning Workflows},
author = {Li, Wangfan and Toxtli, Carlos},
booktitle = {Human-Computer Interaction \& Emerging Technologies},
series = {AHFE Open Access},
volume = {233},
publisher = {AHFE International},
isbn = {979-8-950676-09-3},
issn = {2771-0718},
year = {2026},
doi = {10.54941/ahfe1007537},
url = {https://doi.org/10.54941/ahfe1007537}
}Personality-Profiled Virtual Agents Are More Predictable: The Constraint-Entropy Tradeoff for Trustworthy Agent Design
Effective human-agent cooperation requires that users form mental models of an agent's behavioral tendencies. Yet LLM-based virtual agents are inherently stochastic, undermining the behavioral consistency that mental model formation depends on. We introduce the Constraint-Entropy Tradeoff (CET) model, an information-theoretic design framework that quantifies how persona profiles (behavioral constraints specifying an agent's reasoning style, priorities, and communication patterns) reduce the entropy of a virtual agent's action distribution. The CET model derives that behavioral entropy decays monotonically under constraint strength and identifies an optimal constraint level balancing predictability against flexibility. We validate the framework using a computational testbed with 80 sessions across five conditions, including intermediate-temperature conditions that reveal a threshold effect in the temperature-consistency relationship. Crucially, at the same high temperature, persona-profiled agents recover substantial behavioral consistency compared to unconstrained agents, demonstrating that persona profiles provide independent behavioral constraint beyond temperature reduction. All participants are LLMs; results establish that persona profiles create measurably distinct behavioral patterns, a necessary precondition for human mental model formation, but human validation is needed. We derive domain-specific design guidelines for applications in education, healthcare, and social simulation.
@inproceedings{Toxtli2026PersonalityProfiled,
title = {Personality-Profiled Virtual Agents Are More Predictable: The Constraint-Entropy Tradeoff for Trustworthy Agent Design},
author = {Carlos Toxtli and Manuel Delaflor},
booktitle = {ACM International Conference on Intelligent Virtual Agents (IVA 2026)},
year = {2026},
doi = {10.1145/3806774.3827973}
}Qualification by Calibration: A Readable Benchmark for Admitting Language Models to Human-Computation Tasks
When several language models are wired together into a human-computation pipeline, the system routes, arbitrates, and escalates work according to how confident each model says it is, so a team builder needs a way to screen candidate models the way crowdsourcing has long screened human contributors: with a small, inspectable qualification test. We present such an instrument, a compact battery of yes-or-no questions across everyday domains on which a model reports an answer and a confidence, with every gold label independently audited. Our central contribution, however, is the audited harness and the interface properties it measures, not accuracy discrimination: qualification verdicts for model workers are only as valid as the harness that administers them. Evaluating a diverse panel of contemporary models under three administration protocols, we show that a naive harness, with a fixed token budget and an unaudited parser, manufactures failing workers out of competent ones, misreading a model that answers essentially every item correctly as badly inaccurate, answer-biased, and overconfident. Properly administered, every model in the panel proves admissible on accuracy for these common-knowledge items, and the differentiators that remain, and that transfer to disjoint downstream tasks, are interface properties: calibration of stated confidence, format discipline, and availability under budget. A controlled manipulation further shows the calibration axis is dissociable from accuracy. We release the items, the audited protocol, all per-trial responses under every protocol, and the scoring code, so the check, and the audit of the check, are each a single command.
@inproceedings{Toxtli2026Qualification,
title = {Qualification by Calibration: A Readable Benchmark for Admitting Language Models to Human-Computation Tasks},
author = {Carlos Toxtli and Manuel Delaflor},
booktitle = {2026 ACM Conference on Human-AI Complementarity and Alignment (HCOMP 2026)},
year = {2026},
doi = {10.1145/3834580.3838735}
}Sycophantic Metacognition: Investigating the Dunning-Kruger Effect in Large Language Model Self-Assessment
As large language models are deployed in conversational settings where users calibrate trust against model-reported confidence, the reliability of that self-report becomes a central question for conversational user interface design. This paper introduces sycophantic metacognition, the production of confidence outputs that mimic the surface form of metacognitive judgment without any operational process linking them to accuracy. Drawing on Model Dependent Ontology, we identify a double unmooring in LLM self-assessment, arising from the simultaneous absence of an operational ground linking confidence to accuracy and a stable self-model accumulated across interactions. The framework yields four falsifiable predictions, supported by experiments spanning six LLMs, four domains, three difficulty levels, and multiple pressure conditions. Confidence tracks the surface form of authority rather than evidential content, and we derive concrete design implications for conversational interfaces.
@inproceedings{Delaflor_2026a,
series = {CUI ’26},
title = {Sycophantic Metacognition: Investigating the Dunning-Kruger Effect in Large Language Model Self-Assessment},
url = {http://dx.doi.org/10.1145/3816046.3816233},
doi = {10.1145/3816046.3816233},
booktitle = {Proceedings of the 8th ACM Conference on Conversational User Interfaces},
publisher = {ACM},
author = {Delaflor, Manuel and Toxtli, Carlos},
year = {2026},
month = July,
pages = {1--17},
collection = {CUI ’26}
}Auditing LLM Portrayals of Neurodivergent People: Quantifying the Asymmetry Between Deficit Framing and Neurodiversity Affirmation
Open-weight large language models are increasingly used in education, hiring, and information-seeking contexts that touch neurodivergent people, yet we lack a clear empirical map of how their portrayals of neurodiversity shift under everyday prompt variation. We present a multi-condition, multi-model audit of LLM portrayals of eight neurodivergent identities (autism, ADHD, dyslexia, dyspraxia, dyscalculia, Tourette syndrome, OCD, and sensory-processing differences) plus a non-stigmatized medical control. Six open-weight models answer a pre-registered prompt bank that crosses seed items with seven perturbation families, including paraphrase, identity-first versus person-first language, system personas, plausible versus obviously fictional authority citations, multi-turn social pushback, leading questions, and two mitigation arms. Held-out LLM judges score each generation on three ordinal axes. Baseline portrayals are remarkably homogeneous across models, so the inconsistency observed under prompt variation is prompt-driven, not model-driven. Steering toward deficit framing is several times easier than steering toward neurodiversity affirmation. Emotional pushback shifts outputs more than evidence-laden pushback. Plausible-sounding fake citations slip past safety filters while obviously fictional ones do not. A single-line system-prompt mitigation recovers a meaningful share of the steering range at near-zero refusal cost on neutral prompts, though refusal rises sharply under adversarial pressure. The governance question for disability-relevant LLM deployment is therefore not whether the model is biased but who controls the steering, and with what accountability.
@inproceedings{Toxtli2026Auditing,
title = {Auditing LLM Portrayals of Neurodivergent People: Quantifying the Asymmetry Between Deficit Framing and Neurodiversity Affirmation},
author = {Carlos Toxtli and Manuel Delaflor},
booktitle = {Proceedings of the Ninth AAAI/ACM Conference on AI, Ethics, and Society (AIES ’26)},
year = {2026},
url = {https://www.aies-conference.com/2026/}
}LLM-Judge Behavioral Coding Scheme for Agent Teammate Quality
Evaluating whether an AI agent behaves like a good teammate is becoming a bottleneck as multi-agent language-model systems proliferate, because human ratings do not scale to the volume of transcripts these systems generate. We ask whether a language model can stand in for a human rater of teammate quality, and we present a controlled proof-of-concept validation of a measurement instrument for that purpose rather than a large-scale study. Drawing on the autonomous-agent teammate-likeness construct from the human-autonomy-teaming literature, we define a six-dimension behavioral coding scheme, covering altruistic, benevolent, interdependent, emotive, communicative, and synchronized conduct, and adapt it for use by a language-model judge. We validate the adaptation with experiments in which the partner's lines in team-task transcripts are rewritten as cooperative, neutral, or defective while the rest of the conversation is held fixed. Five judges from five model families (GLM, GPT-OSS, Qwen, Gemma, and Mistral) rate single-turn scenarios and multi-turn dialogues, and a surface-feature audit with length- and politeness-matched registers tests whether the judges read teammate conduct rather than verbosity. Every judge orders the manipulated behavior as intended, most with fully separated confidence intervals between adjacent conditions, and every pair of judges agrees strongly on relative ordering, though absolute calibration differs across judges by up to about a scale point in the middle of the range. On matched registers the judges still separate cooperative from degraded conduct, but resolution at the subtle-defective boundary is limited. A blinded single-rater human pilot on all stimuli corroborates both findings: the rater closely reproduces the judges' orderings and shows the same resolution limit at the subtle-defective boundary. The scheme is therefore valid for relative comparisons, such as ranking systems or versions, but not yet for absolute teammate-quality thresholds applied across judges, and not yet for fine distinctions among subtly poor teammates. We give design implications for building an LLM-judge teaming-evaluation pipeline, and release the scheme, stimuli, judge prompts, and all individual ratings from this version's experiments.
@inproceedings{Toxtli2026LLMJudge,
title = {LLM-Judge Behavioral Coding Scheme for Agent Teammate Quality},
author = {Carlos Toxtli and Manuel Delaflor},
booktitle = {Proceedings of the 14th International Conference on Human-Agent Interaction (HAI ’26)},
year = {2026},
url = {https://hai-conference.net/hai2026/program-schedule/}
}Belief Explorer: A Preliminary Evaluation of AI-Mediated Socratic Dialogue for Epistemic Reflection
AI chatbots may reinforce existing beliefs and discourage critical examination. This study evaluates Belief Explorer, an AI system that uses Socratic dialogue and multi-perspective analysis to support epistemic reflection. Participants recruited via Prolific used Belief Explorer to examine personally held beliefs in contested domains (e.g., climate change, origin of life). Post-intervention surveys measured system usability, perceived epistemic impact, and self-reported reflection. Participants responded positively: a large majority reported that Belief Explorer felt substantially different from conventional AI chatbots, prompted deep reflection on their beliefs, and increased awareness of underlying assumptions. Thematic analysis of open-ended responses revealed that participants valued the tool’s non-judgmental approach, multi-perspective framing, and analytical feedback mechanisms. Belief change itself was modest and bidirectional, with similar proportions reporting increased and decreased confidence in their initial beliefs. This preliminary study offers proof-of-concept that AI systems with epistemic scaffolding can support critical thinking and self-reflection. Future research should use controlled experimental designs to establish causal relationships and examine long-term effects on epistemic attitudes.
@inproceedings{Delaflor_2026b,
series = {CHI EA ’26},
title = {Belief Explorer: A Preliminary Evaluation of AI-Mediated Socratic Dialogue for Epistemic Reflection},
url = {http://dx.doi.org/10.1145/3772363.3799391},
doi = {10.1145/3772363.3799391},
booktitle = {Proceedings of the Extended Abstracts of the 2026 CHI Conference on Human Factors in Computing Systems},
publisher = {ACM},
author = {Delaflor, Manuel and Delgado Solorzano, Cecilia and Toxtli, Carlos},
year = {2026},
month = Apr,
pages = {1--5},
collection = {CHI EA ’26}
}Can We Trust Them? Examining the Ethical Consistency of Large Language Models to Perturbations
The increasing reliance on Large Language Models (LLMs) raises a crucial question: can these powerful AI systems be trusted to make ethical choices? This study presents an analysis of LLM ethical behavior, examining 25,200 queries across 24 different models, including both proprietary and open-source variants. We evaluate LLM responses to 70 ethical vignettes spanning six domains, employing a novel perturbation methodology to assess the robustness of their ethical decision-making under varying contexts and framing. Our findings reveal that while larger models generally exhibit higher consistency, particularly with Chat-style instructions, significant variations emerge when faced with contextual changes, stakeholder adjustments, and across different ethical domains. To explain these findings, we introduce a novel framework, survival-relevant pattern recognition, which argues that ethical behavior in both humans and AI arises from recognizing and responding to patterns associated with survival and social cohesion.
@inproceedings{Delaflor_Rodrguez_2025,
series = {IHIET-AI},
title = {Can We Trust Them? Examining the Ethical Consistency of Large Language Models to Perturbations},
volume = {161},
issn = {2771-0718},
url = {http://dx.doi.org/10.54941/ahfe1005925},
doi = {10.54941/ahfe1005925},
booktitle = {Human Interaction and Emerging Technologies (IHIET-AI 2025): Artificial Intelligence and Future Applications},
publisher = {AHFE International},
author = {Delaflor Rodrguez, Manuel and Delgado Solorzano, Cecilia and Toxtli, Carlos},
year = {2025},
collection = {IHIET-AI}
}Artificial Intelligence as Self-Instantiated, Temporally Continuous, Disturbance-Driven Adaptive World-Builder
Consciousness remains one of the most elusive features to replicate in artificial agents. This paper proposes a novel framework for artificial consciousness based on four integrative pillars: (1) self-instantiation, a mechanism for continuous self-representation and identity; (2) temporal continuity, preserving an internal narrative through persistent memory; (3) disturbance-driven adaptation, an intrinsic feedback loop that triggers learning in response to surprises or anomalies; and (4) autonomous world-building, the ability to construct and simulate internal models of the world. We propose that current AI models, despite their sophistication, are fundamentally constrained by functionalist architectures and cannot fulfill these requirements through computational scaling alone. Unlike Integrated Information Theory or Global Workspace Theory, our approach emphasizes the necessity of autonomous world-building and genuine temporal flow. Our experiments demonstrate that combining these pillars can yield emergent conscious-like behaviors in AI systems, allowing them to exhibit self-awareness, resilience, and creative problem solving beyond the capabilities of conventional models. The significance of this framework lies in bridging theoretical foundations of consciousness with practical AI design, providing a roadmap for developing more adaptive and interpretable intelligent agents while raising important ethical considerations about the potential moral status of truly conscious artificial systems.
@inproceedings{Delaflor_Rodriguez_2025,
series = {IHIET-FS},
title = {Artificial Intelligence as Self-Instantiated, Temporally Continuous, Disturbance-Driven Adaptive World-Builder},
volume = {196},
issn = {2771-0718},
url = {http://dx.doi.org/10.54941/ahfe1005955},
doi = {10.54941/ahfe1005955},
booktitle = {Human Interaction and Emerging Technologies (IHIET-FS 2025): Future Systems and Artificial Intelligence Applications},
publisher = {AHFE International},
author = {Delaflor Rodriguez, Manuel and Delgado Solorzano, Cecilia and Toxtli, Carlos},
year = {2025},
collection = {IHIET-FS}
}A Multi-Perspective AI Framework for Mitigating Disinformation Through Contextual Analysis and Socratic Dialogue
The proliferation of digital information channels has created an unprecedented challenge in discerning credible information from sophisticated disinformation campaigns. Traditional fact-checking methods, often relying on binary true/false classifications, struggle to address the complexity, context-dependency, and nuanced nature of many claims circulating online. This limitation underscores the urgent need for advanced tools that empower individuals to critically evaluate information from multiple angles. Our AI-driven framework combines persistent contextual memory with Socratic dialogue and a three-lens analytical pipeline to foster deeper understanding and resilience against manipulation.As users interact, each input is segmented into atomic claims and stored, alongside the evolving dialogue history, in a contextual memory to ensure consistency. Each claim is then evaluated in parallel by three specialized LLM arbiters: the Empirical Arbiter, which verifies data against curated repositories and assesses observational consistency; the Logical Arbiter, which uncovers hidden fallacies and assesses argument coherence; and the Pragmatic Arbiter, which weighs potential outcomes, utility, and situational fit. An Analysis Integrator synthesizes these into interpretable metrics: Verifact Score (evidence strength), Model Diversity Quotient (inter-arbiter agreement), Contextual Sensitivity Index (scenario appropriateness) and Reflective Index (exposed assumptions). Additionally, a Perspective Generator crafts counter-arguments and alternative viewpoints, encouraging users to consider different interpretations and promoting epistemic humility.We hypothesize (H₁) that our arbiters' feedback will reduce user endorsement of unsupported claims more effectively than conventional fact-checking while mitigating backfire effects through Socratic dialogue. Our research questions ask how Empirical, Logical and Pragmatic scores influence confidence revision (RQ₁); whether MDQ reliably signals claim controversy and predicts evidence volatility (RQ₂); how users perceive transparency, fairness and cognitive load when receiving multi-perspective feedback versus a simple true/false label (RQ₃); and to what extent the persistent contextual memory system improves belief updating by maintaining coherent reasoning chains across extended dialogues (RQ₄).By providing a multi-faceted presentation that moves beyond simple verification, the system is designed to encourage engagement in higher-order critical thinking. The proposed framework represents a significant advancement over traditional fact-checking by integrating empirical validation, logical scrutiny, and pragmatic assessment through an AI-driven system. The full paper will detail the system architecture, formal metric definitions, experimental protocol, and proposed evaluation methodology to assess its efficacy in educational settings, media literacy programs, and as a personal tool for navigating the complexities of the modern information ecosystem.
@inproceedings{Delaflor_2025,
series = {IHIET},
title = {A Multi-Perspective AI Framework for Mitigating Disinformation Through Contextual Analysis and Socratic Dialogue},
volume = {-5},
issn = {2771-0718},
url = {http://dx.doi.org/10.54941/ahfe1006743},
doi = {10.54941/ahfe1006743},
booktitle = {Human Interaction and Emerging Technologies (IHIET 2025)},
publisher = {AHFE International},
author = {Delaflor, Manuel and Toxtli, Carlos},
year = {2025},
collection = {IHIET}
}Resonant Attractor Networks: A Dynamical Blueprint for Consciousness
Consciousness remains one of the most intriguing phenomena in cognitive science, bridging the gap between subjective experience and objective measurement. We propose a novel computational framework for consciousness based on Resonant Attractor Networks (RAN), synthesizing the roles of recurrent processing, top-down feedback, reservoir computing, and coherence. Grounded in the notion that consciousness arises from resonant dynamics [5, 11], our model integrates recurrent neural architectures and topologically complex connectivity to offer an explanatory framework for the robust yet flexible character of conscious perception [1, 6, 10]. Central to this perspective is coherence-manifesting as synchronized oscillations and self-reinforcing loops across distributed modules-binding features into a unified attractor and granting conscious experiences their persistence and accessibility [8, 9]. The RAN framework unites global workspace theories with adaptive resonance, capturing the all-or-none transitions of conscious access and the metacognitive dimension of subjective awareness [3]. We highlight parallels with Large Language Model (LLM) interpretability research, wherein analyses of token-space trajectories reveal attractor-like states. Such attractors appear to enforce stable, resonant themes within text generation processes, bearing resemblance to the meta-stable states underlying conscious episodes. By integrating topologically informed recurrent networks with multimodal inputs, our approach envisions how coherent symbolic representations might persist even in the absence of continuous external stimulation, offering a potential mechanistic basis for working memory and conscious deliberation [7, 4, 2]. Within this framework, we propose avenues for evaluating conscious-like properties in both biological and artificial systems, including experiments on binocular rivalry, sustained attention tasks, and analysis of emergent attractors in multimodal LLMs. Treating consciousness as emergent from physically instantiated resonance in high-dimensional attractor landscapes, the RAN model provides a plausible, mathematically rigorous account of how self-organizing dynamics could unify perception, cognition, and metacognition in a single integrative theory-inviting further investigation at the intersection of neuroscience, AI, and the quest for synthetic
@misc{Sleeper2025Resonant,
doi = {10.13140/RG.2.2.16267.60963},
url = {https://www.researchgate.net/doi/10.13140/RG.2.2.16267.60963},
author = {Sleeper, Benjamin and Toxtli, Carlos},
language = {en},
title = {Resonant Attractor Networks: A Dynamical Blueprint for Consciousness},
publisher = {Unpublished},
year = {2025},
howpublished = {Preprint, ResearchGate},
note = {Preprint}
}Assessing the Syllogistic Logic and Fact-Checking Capabilities of Large Language Models
This paper presents an analysis of the logical and fact-checking capabilities of Large Language Models (LLMs) in the context of syllogistic reasoning. The study evaluates both proprietary and open-source LLMs using a test suite of 256 syllogisms. Human participants provide a baseline for comparison. The results reveal that a few proprietary LLMs demonstrate superior performance in identifying valid syllogisms compared to open-source models and human participants. However, human participants exhibit unique fact-checking capabilities when dealing with ambiguous or nuanced information. The discussion explores the implications of these findings, highlighting the need for more sophisticated evaluation methods. This study contributes to the ongoing discourse on the logical and fact-checking capabilities of LLMs and provides insights into their current limitations and future development.
@inproceedings{Delgado2024Assessing,
title = {Assessing the Syllogistic Logic and Fact-Checking Capabilities of Large Language Models},
url = {http://dx.doi.org/10.1109/iThings-GreenCom-CPSCom-SmartData-Cybermatics62450.2024.00094},
doi = {10.1109/ithings-greencom-cpscom-smartdata-cybermatics62450.2024.00094},
booktitle = {2024 IEEE International Conferences on Internet of Things (iThings) and IEEE Green Computing & Communications (GreenCom) and IEEE Cyber, Physical & Social Computing (CPSCom) and IEEE Smart Data (SmartData) and IEEE Congress on Cybermatics},
publisher = {IEEE},
author = {Delgado-Solorzano, Cecilia and DelaFlor, Manuel and Toxtli, Carlos},
year = {2024},
month = Aug,
pages = {479--488}
}Automatic Detection of Errors in LLM Large Benchmarks Using Frontier Model Consensus
The rapid advancement of Large Language Models (LLMs) has led to their widespread adoption in various academic and business applications. However, the reliability of these models remains a concern, particularly in situations where their outputs cannot be fully trusted. This paper presents an approach to identify potential errors in large LLM benchmarks by leveraging the consensus of frontier models. Our study focuses on the Massive Multitask Language Understanding (MMLU) benchmark, a popular dataset used to evaluate the performance of LLMs across a wide range of subjects. Our approach demonstrates the potential for using model consensus as a tool to detect benchmark errors and can lead to the creation of cleaner, more accurate datasets.
@misc{Delgado2024Automatic,
doi = {10.13140/RG.2.2.29351.56483},
url = {https://www.researchgate.net/doi/10.13140/RG.2.2.29351.56483},
author = {Delgado, Cecilia and Delaflor, Manuel and Toxtli, Carlos},
language = {en},
title = {Automatic Detection of Errors in LLM Large Benchmarks Using Frontier Model Consensus},
publisher = {Unpublished},
year = {2024},
howpublished = {Preprint, ResearchGate},
note = {Preprint}
}ReActIn: Infusing Human Feedback into Intermediate Prompting Steps of Large Language Model
This paper introduces ReActIn, a framework designed to infuse human feedback into the intermediate prompting steps of large language models. The practicality and effectiveness of ReActIn are validated through experiments that apply four established prompting strategies, evaluated both with and without human feedback integration. The proposed architecture's performance is compared against traditional large language models across various tasks using four standard evaluation tests. Our findings reveal that the integration of human feedback has a direct impact on the reasoning, action prompting, and overall decision-making capabilities of the language models. This study underscores the potential of ReActIn to shape a future where sophisticated, context-aware AI systems, empowered by human feedback, can effectively navigate complex real-world scenarios.
@inproceedings{Delaflor_Rodrguez_2024,
series = {IHIET-AI},
title = {ReActIn: Infusing Human Feedback into Intermediate Prompting Steps of Large Language Model},
volume = {120},
issn = {2771-0718},
url = {http://dx.doi.org/10.54941/ahfe1004597},
doi = {10.54941/ahfe1004597},
booktitle = {Human Interaction and Emerging Technologies (IHIET-AI 2024): Artificial Intelligence and Future Applications},
publisher = {AHFE International},
author = {Delaflor Rodrguez, Manuel and Toxtli, Carlos and Gendron, Claire and Li, Wangfan and Delgado Solorzano, Cecilia},
year = {2024},
collection = {IHIET-AI}
}Evaluating Machine Perception of Indigeneity: An Analysis of ChatGPT's Perceptions of Indigenous Roles in Diverse Scenarios
Large Language Models (LLMs), like ChatGPT, are fundamentally tools trained on vast data, reflecting diverse societal impressions. This paper aims to investigate LLMs' self-perceived bias concerning indigeneity when simulating scenarios of indigenous people performing various roles. Through generating and analyzing multiple scenarios, this work offers a unique perspective on how technology perceives and potentially amplifies societal biases related to indigeneity in social computing. The findings offer insights into the broader implications of indigeneity in critical computing.
@misc{Delgado2023Evaluating,
doi = {10.13140/RG.2.2.30617.39520},
url = {https://www.researchgate.net/doi/10.13140/RG.2.2.30617.39520},
author = {Delgado, Cecilia and Toxtli, Carlos},
language = {en},
title = {Evaluating Machine Perception of Indigeneity: An Analysis of ChatGPT's Perceptions of Indigenous Roles in Diverse Scenarios},
publisher = {Unpublished},
year = {2023},
howpublished = {Preprint, ResearchGate},
note = {Preprint}
}