Document Type
Abstract
Publication Date
2-11-2026
Abstract
Large Language Models (LLMs) have been previously utilized in the medical field to successfully generate information. However, LLMs may hallucinate, which generates a response based on false or misleading information. This can negatively impact healthcare outcomes, thus, there is an increasing need to develop a model free of hallucinations.
A multi-target voting ensemble (Random Forest with HistGradientBoosting) algorithm was utilized in the analysis of 454 training data and 114 test data. Raw accuracy and macro-F1 statistics determined that the model is capable of identifying clinically meaningful data for majority classes, while minority-class detection was insignificant. This demonstrates moderate predictive value for some features of ordinal and categorical data, while continuous variables did not report any significant predictive capability.
A Retrieval-Augmented Generation (RAG) model was then utilized in conjunction with an Ollama LLM to create a user-friendly interface to answer questions and predict patient outcomes. Results from user interface testing demonstrate that the RAG-LLM model did not hallucinate, and only provided information derived from the dataset.
The results from this model suggest that the dataset used was insufficient to train a model capable of current implementation. However, it serves as a proof of concept that a RAG-LLM model can learn clinically significant relationships with predictive capability. Thus, with a greater number of balanced data, a more extensive model has the potential to become a feasible tool for predicting health outcomes.
Recommended Citation
Hone, Alexander; Musmar, Basel; Roy, Joanna; and Jabbour, Pascal, "Utility of Neurosurgery-specific RAG-LLM in Accurate Patient Outcome Prediction" (2026). Phase 1. Paper 3.
https://jdc.jefferson.edu/si_dh_2028_phase1/3
Language
English

Comments
Presented at the 2026 Scholarly Inquiry (SI) Research Project Symposium.