🧠 Speech clusters linked to depression in 3 cohorts
🧠 Speech clusters linked to depression in 3 cohorts
LLM-clustered spontaneous speech was linked to depression and related symptom burden across 3 multilingual cohorts—French (n=1338), Italian (n=116), and Chinese (n=52)—with the strongest signal in the French sample, where PHQ-9 scores differed across clusters with a large effect (η²=0.19, 95% CI 0.17-0.24). In this study of 4 cohorts, the “feelings and sleep” prompt produced the strongest clinical associations, while past- or future-focused questions had small effects.
Why It Matters To Your Practice
Depression remains underdiagnosed, and clinicians often depend on subjective interpretation of patient speech.
This approach suggests AI may help scale speech-based phenotyping by turning open-ended responses into human-readable clusters tied to validated symptom scales.
Because the signal appeared across multiple languages, the method may be relevant to diverse practice settings.
Clinical Implications
Not all interview questions are equally informative: prompts about feelings and sleep showed the strongest associations with clinical status.
In the Italian cohort, cluster membership aligned with clinician-assigned depression diagnosis (Cramér V=0.55, 95% CI 0.36-0.78; P=.001).
Potential near-term uses include screening support and structured intake augmentation, but prospective validation for treatment planning or decision support is still needed.
Insights
The model used unsupervised clustering of multilingual speech transcripts, then generated natural-language summaries of each cluster with an LLM.
Associations extended beyond depression to anxiety, insomnia, fatigue, and in one smaller cohort an exploratory suicide-risk signal.
Sociodemographic factors also shaped cluster membership—especially age (η²=0.27) and sex (Cramér V=0.22) for the “last 24 hours” question—highlighting a key confound for clinical AI tools.
The Bottom Line
AI-based clustering of spontaneous speech may help clinicians detect depression-related patterns at scale, but question design matters and demographic confounding is real.
For now, this looks more like a promising workflow aid than a practice-ready diagnostic tool.