Abstract
The rapid advancement of Large Language Models (LLMs) has led to their widespread adoption in various academic and business applications. However, the reliability of these models remains a concern, particularly in situations where their outputs cannot be fully trusted. This paper presents an approach to identify potential errors in large LLM benchmarks by leveraging the consensus of frontier models. Our study focuses on the Massive Multitask Language Understanding (MMLU) benchmark, a popular dataset used to evaluate the performance of LLMs across a wide range of subjects. Our approach demonstrates the potential for using model consensus as a tool to detect benchmark errors and can lead to the creation of cleaner, more accurate datasets.
Cite this work
Cecilia Delgado Solorzano, Manuel Delaflor, and Carlos Toxtli-Hernández. 2024. Automatic Detection of Errors in LLM Large Benchmarks Using Frontier Model Consensus. The 26th International Conference on Artificial Intelligence (ICAI'24), World Congress in Computer Science, Computer Engineering & Applied Computing (CSCE 2024). Springer Nature, Cham.. https://doi.org/10.1007/978-3-031-86623-4_16
@misc{Delgado2024Automatic,
doi = {10.13140/RG.2.2.29351.56483},
url = {https://www.researchgate.net/doi/10.13140/RG.2.2.29351.56483},
author = {Delgado, Cecilia and Delaflor, Manuel and Toxtli, Carlos},
language = {en},
title = {Automatic Detection of Errors in LLM Large Benchmarks Using Frontier Model Consensus},
publisher = {Unpublished},
year = {2024},
howpublished = {Preprint, ResearchGate},
note = {Preprint}
}Related