AHFE IHIET-AI 2025 Conference paper

Can We Trust Them? Examining the Ethical Consistency of Large Language Models to Perturbations

Manuel Delaflor, Cecilia Delgado Solorzano, Carlos Toxtli-Hernández

Human Interaction and Emerging Technologies (IHIET-AI 2025): Artificial Intelligence and Future Applications, 2025

Abstract

The increasing reliance on Large Language Models (LLMs) raises a crucial question: can these powerful AI systems be trusted to make ethical choices? This study presents an analysis of LLM ethical behavior, examining 25,200 queries across 24 different models, including both proprietary and open-source variants. We evaluate LLM responses to 70 ethical vignettes spanning six domains, employing a novel perturbation methodology to assess the robustness of their ethical decision-making under varying contexts and framing. Our findings reveal that while larger models generally exhibit higher consistency, particularly with Chat-style instructions, significant variations emerge when faced with contextual changes, stakeholder adjustments, and across different ethical domains. To explain these findings, we introduce a novel framework, survival-relevant pattern recognition, which argues that ethical behavior in both humans and AI arises from recognizing and responding to patterns associated with survival and social cohesion.

Cite this work

Manuel Delaflor, Cecilia Delgado Solorzano, and Carlos Toxtli-Hernández. 2025. Can We Trust Them? Examining the Ethical Consistency of Large Language Models to Perturbations. Human Interaction and Emerging Technologies (IHIET-AI 2025): Artificial Intelligence and Future Applications. https://doi.org/10.54941/ahfe1005925

@inproceedings{Delaflor_Rodrguez_2025,
  series = {IHIET-AI},
  title = {Can We Trust Them? Examining the Ethical Consistency of Large Language Models to Perturbations},
  volume = {161},
  issn = {2771-0718},
  url = {http://dx.doi.org/10.54941/ahfe1005925},
  doi = {10.54941/ahfe1005925},
  booktitle = {Human Interaction and Emerging Technologies (IHIET-AI 2025): Artificial Intelligence and Future Applications},
  publisher = {AHFE International},
  author = {Delaflor Rodrguez, Manuel and Delgado Solorzano, Cecilia and Toxtli, Carlos},
  year = {2025},
  collection = {IHIET-AI}
}