Document Type
Article
Publication Date
7-14-2026
Abstract
BACKGROUND: Large language models (LLMs) are increasingly used to generate health information, yet their reliability as evaluators remains unclear. This study investigated the feasibility of an LLM-as-a-judge methodology in the context of infection prevention and antimicrobial resistance (AMR), comparing automated ratings with human expert benchmarks.
METHODS: We performed a secondary analysis of an expert-annotated dataset of health messages. Three leading LLMs (ChatGPT, Claude, Gemini) independently evaluated the same messages using an adapted DISCERN tool across five domains: information reliability, quality, AMR impact, persuasiveness, and overall score. We utilized descriptive statistics, intra-rater reliability tests, and mixed-effects ordinal regression to analyze divergence between automated and human assessments, adhering to CHART reporting guidelines.
RESULTS: Analysis of 404 evaluations revealed a systematic upward divergence: all LLMs consistently assigned higher scores than human experts. This optimism bias persisted after adjusting for domain-specific differences and clustering effects. The gap was particularly pronounced in domains of persuasiveness and AMR impact, while information quality showed more heterogeneous results. Intra-rater reliability assessments demonstrated that LLMs maintained stable scoring patterns under identical prompting conditions.
CONCLUSIONS: LLMs exhibit a consistent leniency bias, systematically overestimating the quality of AMR-related health communication compared to human evaluators. These results do not support the use of LLMs for autonomous evaluation in high-stakes public health contexts. Rather, LLM-based judging is best suited as a scalable screening tool within supervised human-in-the-loop workflows, where expert oversight serves as a necessary safeguard for evidence-based accuracy.
Recommended Citation
Di Pumpo, Marcello; Villani, Leonardo; Gualano, Maria Rosaria; Buonsenso, Danilo; Raffaelli, Francesca; Donà, Daniele; Laurenti, Patrizia; Maio, Vittorio; Boccia, Stefania; and Ricciardi, Walter, "LLM-as-a-judge for Infection Prevention and Control and Antimicrobial Resistance Impact: Comparing Three Main LLMs vs. Human Experts' Assessment" (2026). College of Population Health Faculty Papers. Paper 258.
https://jdc.jefferson.edu/healthpolicyfaculty/258
Creative Commons License

This work is licensed under a Creative Commons Attribution 4.0 License.
Supplementary_File_2.docx (31 kB)
Supplementary_File_3.docx (21 kB)
Language
English
Included in
Artificial Intelligence and Robotics Commons, Chemical and Pharmacologic Phenomena Commons, Diagnosis Commons, Public Health Commons

Comments
This article, first published by Frontiers Media, is the author’s final published version in Frontiers in Public Health, Volume 14, 2026, Article number 1874389.
The published version is available at https://doi.org/10.3389/fpubh.2026.1874389. Copyright © 2026 DiPumpo, Villani, Gualano, Buonsenso, Raaelli, Donà, Laurenti, Maio, Boccia and Ricciardi.