Document Type
Article
Publication Date
4-30-2026
Abstract
Accurate disease classification from radiology reports is essential for many applications. While supervised fine-tuning (SFT) of lightweight LLMs improves accuracy, it can degrade reasoning. We propose a two-stage approach: SFT on disease labels followed by Group Relative Policy Optimization (GRPO) to refine predictions by optimizing accuracy and format without reasoning supervision. Across three radiologist-annotated datasets, SFT outperformed baselines and GRPO further improved classification and enhanced reasoning recall and comprehensiveness.
Recommended Citation
Wei, Yishu; Lin, Yi; Flanders, Adam; Shih, George; and Peng, Yifan, "Reinforcement Learning Improves LLM Accuracy and Reasoning in Disease Classification From Radiology Reports" (2026). Department of Radiology Faculty Papers. Paper 197.
https://jdc.jefferson.edu/radiologyfp/197
PubMed ID
42062541
Language
English

Comments
This article is the author’s final published version in npj Digital Medicine, Volume 9, Issue 1, 2026, Article number 515.
The published version is available at https://doi.org/10.1038/s41746-026-02685-4. Copyright © The Author(s) 2026.