Document Type
Abstract
Publication Date
2-11-2026
Abstract
Introduction: Accurate grading of cartilage pathology is essential for diagnosis and operative planning, yet physician-reported grading systems remain difficult to standardize. Artificial intelligence (AI)has the potential to automate cartilage assessment, but its methodology and clinical applicability remain unclear. This scoping review evaluates current applications of AI for cartilage grading and assesses their translational readiness.
Methods: A scoping review was conducted in accordance with PRISMA-ScR guidelines, including a systematic literature search, predefined eligibility criteria, transparent study selection, and structured data extraction. Eligible studies included English-language human research describing AI models for cartilage evaluation that were assessed against human or clinical grading standards. Extracted variables included AI task, reference standard (ground truth), external validation, and reported clinical application.
Results: Of 1,730 identified studies, 12 met inclusion criteria. Ten studies evaluated knee cartilage, and two evaluated hip cartilage. Sample sizes ranged from five to 46,966 patients. Ground-truth methods varied, with six studies using manual MRI labeling, four using arthroscopy, and two using histology. Model outputs included binary lesion detection (n=5) and multiclass grading systems such as ICRS, Outerbridge, or WORMS (n=7). Commonly reported performance metrics included accuracy, area under the curve (AUC), sensitivity, specificity, and Dice coefficient. Multiclass grading models reported a mean accuracy of 91.3%, compared with81.3% for binary detection models. Interpretation of these results was limited by methodological heterogeneity and the fact that only one study performed external validation.
Conclusions: AI-based cartilage assessment demonstrates high internal performance across multiple task types. However, heterogeneity in reference standards, limited external validation, and insufficient clinical correlation remain major barriers to translation. Future studies should prioritize standardized ground-truth definitions, external validation, and clinically meaningful endpoints.
Recommended Citation
Charlton, BS, Alex R.; Markmann, BS, Caroline A.; Wu, MD, MPH, Isabella T.; and Freedman, MD, Kevin B., "Evaluating Concordance Between AI-Generated Cartilage Grading and Physician-Reported Systems: A Scoping Review" (2026). Phase 1. Paper 1.
https://jdc.jefferson.edu/si_dh_2028_phase1/1
Language
English

Comments
Presented at the 2026 Scholarly Inquiry (SI) Research Project Symposium.