Abstract
As large language models are deployed in conversational settings where users calibrate trust against model-reported confidence, the reliability of that self-report becomes a central question for conversational user interface design. This paper introduces sycophantic metacognition, the production of confidence outputs that mimic the surface form of metacognitive judgment without any operational process linking them to accuracy. Drawing on Model Dependent Ontology, we identify a double unmooring in LLM self-assessment, arising from the simultaneous absence of an operational ground linking confidence to accuracy and a stable self-model accumulated across interactions. The framework yields four falsifiable predictions, supported by experiments spanning six LLMs, four domains, three difficulty levels, and multiple pressure conditions. Confidence tracks the surface form of authority rather than evidential content, and we derive concrete design implications for conversational interfaces.
Cite this work
Manuel Delaflor and Carlos Toxtli-Hernández. 2026. Sycophantic Metacognition: Investigating the Dunning-Kruger Effect in Large Language Model Self-Assessment. Proceedings of the 8th ACM Conference on Conversational User Interfaces. https://doi.org/10.1145/3816046.3816233
@inproceedings{Delaflor_2026a,
series = {CUI ’26},
title = {Sycophantic Metacognition: Investigating the Dunning-Kruger Effect in Large Language Model Self-Assessment},
url = {http://dx.doi.org/10.1145/3816046.3816233},
doi = {10.1145/3816046.3816233},
booktitle = {Proceedings of the 8th ACM Conference on Conversational User Interfaces},
publisher = {ACM},
author = {Delaflor, Manuel and Toxtli, Carlos},
year = {2026},
month = July,
pages = {1--17},
collection = {CUI ’26}
}Related