Document Type

Article

Publication Date

7-14-2026

Comments

This article, first published by Frontiers Media, is the author’s final published version in Frontiers in Public Health, Volume 14, 2026, Article number 1874389.

The published version is available at https://doi.org/10.3389/fpubh.2026.1874389. Copyright © 2026 DiPumpo, Villani, Gualano, Buonsenso, Raaelli, Donà, Laurenti, Maio, Boccia and Ricciardi.

 

Abstract

BACKGROUND: Large language models (LLMs) are increasingly used to generate health information, yet their reliability as evaluators remains unclear. This study investigated the feasibility of an LLM-as-a-judge methodology in the context of infection prevention and antimicrobial resistance (AMR), comparing automated ratings with human expert benchmarks.

METHODS: We performed a secondary analysis of an expert-annotated dataset of health messages. Three leading LLMs (ChatGPT, Claude, Gemini) independently evaluated the same messages using an adapted DISCERN tool across five domains: information reliability, quality, AMR impact, persuasiveness, and overall score. We utilized descriptive statistics, intra-rater reliability tests, and mixed-effects ordinal regression to analyze divergence between automated and human assessments, adhering to CHART reporting guidelines.

RESULTS: Analysis of 404 evaluations revealed a systematic upward divergence: all LLMs consistently assigned higher scores than human experts. This optimism bias persisted after adjusting for domain-specific differences and clustering effects. The gap was particularly pronounced in domains of persuasiveness and AMR impact, while information quality showed more heterogeneous results. Intra-rater reliability assessments demonstrated that LLMs maintained stable scoring patterns under identical prompting conditions.

CONCLUSIONS: LLMs exhibit a consistent leniency bias, systematically overestimating the quality of AMR-related health communication compared to human evaluators. These results do not support the use of LLMs for autonomous evaluation in high-stakes public health contexts. Rather, LLM-based judging is best suited as a scalable screening tool within supervised human-in-the-loop workflows, where expert oversight serves as a necessary safeguard for evidence-based accuracy.

Creative Commons License

Creative Commons License
This work is licensed under a Creative Commons Attribution 4.0 License.

Language

English

Share

COinS
 
 

To view the content in your browser, please download Adobe Reader or, alternately,
you may Download the file to your hard drive.

NOTE: The latest versions of Adobe Reader do not support viewing PDF files within Firefox on Mac OS and if you are using a modern (Intel) Mac, there is no official plugin for viewing PDF files within the browser window.