Document Type
Article
Publication Date
6-6-2026
Abstract
OBJECTIVES: Vaccination is one of the most effective public health interventions. Large language models (LLMs) show great capability for providing health information. This study compared the performance of ChatGPT 3.5 (G3E), ChatGPT 4o in English (G4E), Claude 3.0 (CDE), Gemini 1.5 (GME), and ChatGPT 4o in Italian (G4I) in delivering information on vaccination and preventive medicine.
STUDY DESIGN: Mixed-method analysis evaluating large language models.
METHODS: Twenty-six expert-designed healthcare scenarios were used to evaluate each model through an adapted DISCERN-based instrument. Four experts independently rated outputs across six domains: Information Reliability, Information Quality, Medical Appropriateness, Impact on Vaccine Hesitancy, Potential for Behavioral Influence, and Overall Rating. Medians and interquartile ranges were calculated, and mixed-effects ordinal logistic regression models were applied to account for inter-rater variability.
RESULTS: G4E showed the highest performance, with significant advantages in Overall Rating (OR = 2.17, 95% CI: 1.20-3.92, p = 0.010) and Medical Appropriateness (OR = 1.86, 95% CI: 1.04-3.33, p = 0.036). G4I outperformed others in Information Quality (OR = 1.76, 95% CI: 1.06-2.93, p = 0.030) but scored lower for vaccine hesitancy and behavioral influence. GME performed weaker across qualitative domains, with occasional generation issues, while CDE and G3E yielded intermediate, consistent results.
CONCLUSIONS: Differences among LLMs reflect model architecture, training data, and language adaptation, influencing clarity, accuracy, and persuasive tone. These disparities highlight the need for domain-specific fine-tuning and language-sensitive optimization to enhance public health communication. LLMs show uneven performance in providing accurate and behaviorally effective vaccine information, underscoring the importance of evaluation and cautious integration into health communication strategies.
Recommended Citation
Di Pumpo, Marcello; Villani, Leonardo; Maio, Vittorio; Gualano, Maria Rosaria; Marziali, Eleonora; De Maio, Lucia; Mancini, Rossella; Boccia, Stefania; Ricciardi, Walter; and Laurenti, Patrizia, "Multidomain Expert Evaluation of Leading Large Language Models as Providers of Vaccination and Preventive Medicine Information" (2026). College of Population Health Faculty Papers. Paper 253.
https://jdc.jefferson.edu/healthpolicyfaculty/253
Creative Commons License

This work is licensed under a Creative Commons Attribution-Noncommercial-No Derivative Works 4.0 License.
Multimedia component 2.pdf (87 kB)
Multimedia component 3.pdf (87 kB)
Language
English

Comments
This article is the author’s final published version in Public Health, Volume 257, 2026, Article number 106359.
The published version is available at https://doi.org/10.1016/j.puhe.2026.106359. Copyright © 2026 The Authors.