Can We Trust Them? Examining the Ethical Consistency of Large Language Models to Perturbations
Human Interaction and Emerging Technologies (IHIET-AI 2025): Artificial Intelligence and Future Applications, 2025
Abstract
The increasing reliance on Large Language Models (LLMs) raises a crucial question: can these powerful AI systems be trusted to make ethical choices? This study presents an analysis of LLM ethical behavior, examining 25,200 queries across 24 different models, including both proprietary and open-source variants. We evaluate LLM responses to 70 ethical vignettes spanning six domains, employing a novel perturbation methodology to assess the robustness of their ethical decision-making under varying contexts and framing. Our findings reveal that while larger models generally exhibit higher consistency, particularly with Chat-style instructions, significant variations emerge when faced with contextual changes, stakeholder adjustments, and across different ethical domains. To explain these findings, we introduce a novel framework, survival-relevant pattern recognition, which argues that ethical behavior in both humans and AI arises from recognizing and responding to patterns associated with survival and social cohesion.
Cite this work
Manuel Delaflor, Cecilia Delgado Solorzano, and Carlos Toxtli-Hernández. 2025. Can We Trust Them? Examining the Ethical Consistency of Large Language Models to Perturbations. Human Interaction and Emerging Technologies (IHIET-AI 2025): Artificial Intelligence and Future Applications. https://doi.org/10.54941/ahfe1005925
@inproceedings{Delaflor_Rodrguez_2025,
series = {IHIET-AI},
title = {Can We Trust Them? Examining the Ethical Consistency of Large Language Models to Perturbations},
volume = {161},
issn = {2771-0718},
url = {http://dx.doi.org/10.54941/ahfe1005925},
doi = {10.54941/ahfe1005925},
booktitle = {Human Interaction and Emerging Technologies (IHIET-AI 2025): Artificial Intelligence and Future Applications},
publisher = {AHFE International},
author = {Delaflor Rodrguez, Manuel and Delgado Solorzano, Cecilia and Toxtli, Carlos},
year = {2025},
collection = {IHIET-AI}
}Related