NTH

EmoTrack: Robust Depression Tracking from Counseling Transcripts across Session Regimes

AuthorsZhaomin Wu, Jiayi Li, Bingsheng He

May 30, 2026 2 min read
Watch on YouTube
The one-line take

EmoTrack uses LLM-derived clinical signals plus structured transcript embeddings to track depression severity across counseling sessions more robustly, especially when longitudinal context is available.

Key results

2.4434
DAIC-WOZ MAE

EmoTrack's overall mean absolute error on the real single-session DAIC-WOZ benchmark.

13.5%
DAIC-WOZ relative MAE reduction

Relative reduction over the strongest DAIC-WOZ baseline (AIDA at 2.8234 MAE).

2.6478
LONGCOUNSEL-8 MAE

EmoTrack's overall mean absolute error on the synthetic longitudinal LONGCOUNSEL-8 benchmark.

3599
LONGCOUNSEL-8 trajectories

Number of usable five-visit trajectories in LONGCOUNSEL-8.

17122
LONGCOUNSEL-8 labeled sessions

Total labeled sessions in LONGCOUNSEL-8 after validity filtering.

23
Clinical features extracted

Number of AIDA-style clinician feature scores extracted by the model.

What the paper found

EmoTrack, from researchers at the National University of Singapore, tackles PHQ-8 depression tracking in counseling transcripts across both single-session and multi-session regimes. The core idea is hybrid: instead of relying on end-to-end fine-tuning or holistic LLM prompting, it uses Qwen3.5-35B-A3B to extract 23 structured clinical features from an AIDA-style schema, combines them with frozen turn-level semantic embeddings from Qwen3-Embedding-8B, and feeds both through a lightweight Transformer encoder-decoder with eight PHQ symptom queries plus an optional previous-session memory module. The authors also introduce LONGCOUNSEL-8, a synthetic longitudinal benchmark with 3,599 five-visit trajectories and 17,122 labeled sessions, built from PsychEval cases, PSYCHE-D severity trajectories, NHANES symptom co-occurrence patterns, and RealCBT dialogue calibration; it provides session-level PHQ-8 supervision for controlled testing of partial disclosure and cross-session continuity. On DAIC-WOZ, EmoTrack reduces MAE from the strongest AIDA baseline’s 2.8234 to 2.4434, a 13.5 percent relative improvement, and on LONGCOUNSEL-8 it reaches 2.6478 MAE, matching the strongest longitudinal baseline closely at 2.65. Ablations show that client-only turns, combined clinical features plus embeddings, symptom-wise supervision, and compact memory retrieval each contribute materially, while direct history concatenation offers little benefit.

Original abstract

Text-based counseling is an important interface for AI mental-health support, where transcripts may be used to monitor depression severity and flag sessions requiring timely human review. However, robust PHQ-8 prediction across session regimes remains challenging: fine-tuning-based methods can exploit richer supervision but may generalize poorly under data scarcity, while prompt-based LLM methods are data-efficient but usually treat each transcript holistically and provide limited support for longitudinal context. We study robust depression tracking from counseling transcripts across single-session and multi-session regimes. We introduce LongCounsel, a multi-session counseling dataset with session-level PHQ-8 supervision for evaluating repeated-session tracking under partial symptom disclosure and cross-session continuity. We further propose EmoTrack, a PHQ-8 prediction framework that combines LLM-extracted clinical signals with frozen turn-level semantic embeddings and trains symptom-specific predictors over the resulting transcript representation. When prior sessions are available, EmoTrack can further incorporate them through compact cross-session memory. Experiments on LongCounsel and DAIC-WOZ show that EmoTrack achieves a clear gain on the real single-session benchmark, including a 13.5% relative MAE reduction over the strongest DAIC-WOZ baseline, and remains competitive with the strongest longitudinal baseline on LongCounsel.

Read the original paper

More in Natural Language Processing

Browse all 26 papers →