📝 Clinical notes yielded scalable CGI-S scoring in MDD
📝 Clinical notes yielded scalable CGI-S scoring in MDD
In a study of 77 psychiatrist-authored notes for patients with major depressive disorder (MDD), GPT-4o estimated Clinical Global Impression-Severity (CGI-S) scores from routine documentation with weighted agreement of κ=0.85 versus average human ratings — slightly above psychiatrist interrater reliability of κ=0.77-0.78. The Johns Hopkins electronic health record study suggests scalable severity scoring may be feasible from existing notes, though performance dropped for shorter notes (κ=0.72 vs 0.92; P=.003).
Why It Matters To Your Practice
CGI-S is widely used in psychiatric research but is rarely captured in routine care, limiting measurement-based practice and real-world outcomes tracking.
If validated further, LLMs could convert narrative notes already being written into structured severity data without adding clinician documentation burden.
Performance was stable across age, sex, race, treatment setting, and copy-forwarded text, supporting potential generalizability within typical clinical workflows.
Clinical Implications
Zero-shot GPT-4o performed best, and few-shot prompting did not improve results, suggesting complex prompt engineering may not be necessary for this task.
Llama-4 showed lower agreement than GPT-4o (κ=0.70 vs 0.85), highlighting that model choice materially affects clinical NLP performance.
Shorter notes were a limitation, so sparse documentation may reduce reliability of automated severity scoring.
Insights
Three board-certified psychiatrists independently rated the notes using a validated depression-specific CGI rubric.
Model-human agreement for GPT-4o was comparable to, and numerically exceeded, expert-human agreement, which is notable for a clinician-rated global severity scale.
The study focused on note-based estimation of illness severity, not diagnosis, treatment selection, or prediction of outcomes.
The Bottom Line
For clinicians interested in AI, this is a practical use case: extracting research-grade severity measures from routine MDD notes at near-expert agreement.
Near-term value is most likely in retrospective research, registry building, and quality measurement rather than autonomous clinical decision-making.
Before deployment, practices would need external validation, governance, and attention to documentation quality — especially for brief notes.