EmoTrack: Robust Depression Tracking from Counseling Transcripts across Session Regimes
AuthorsZhaomin Wu, Jiayi Li, Bingsheng He
Resources
EmoTrack uses LLM-derived clinical signals plus structured transcript embeddings to track depression severity across counseling sessions more robustly, especially when longitudinal context is available.
Key results
EmoTrack's overall mean absolute error on the real single-session DAIC-WOZ benchmark.
Relative reduction over the strongest DAIC-WOZ baseline (AIDA at 2.8234 MAE).
EmoTrack's overall mean absolute error on the synthetic longitudinal LONGCOUNSEL-8 benchmark.
Number of usable five-visit trajectories in LONGCOUNSEL-8.
Total labeled sessions in LONGCOUNSEL-8 after validity filtering.
Number of AIDA-style clinician feature scores extracted by the model.
What the paper found
EmoTrack, from researchers at the National University of Singapore, tackles PHQ-8 depression tracking in counseling transcripts across both single-session and multi-session regimes. The core idea is hybrid: instead of relying on end-to-end fine-tuning or holistic LLM prompting, it uses Qwen3.5-35B-A3B to extract 23 structured clinical features from an AIDA-style schema, combines them with frozen turn-level semantic embeddings from Qwen3-Embedding-8B, and feeds both through a lightweight Transformer encoder-decoder with eight PHQ symptom queries plus an optional previous-session memory module. The authors also introduce LONGCOUNSEL-8, a synthetic longitudinal benchmark with 3,599 five-visit trajectories and 17,122 labeled sessions, built from PsychEval cases, PSYCHE-D severity trajectories, NHANES symptom co-occurrence patterns, and RealCBT dialogue calibration; it provides session-level PHQ-8 supervision for controlled testing of partial disclosure and cross-session continuity. On DAIC-WOZ, EmoTrack reduces MAE from the strongest AIDA baseline’s 2.8234 to 2.4434, a 13.5 percent relative improvement, and on LONGCOUNSEL-8 it reaches 2.6478 MAE, matching the strongest longitudinal baseline closely at 2.65. Ablations show that client-only turns, combined clinical features plus embeddings, symptom-wise supervision, and compact memory retrieval each contribute materially, while direct history concatenation offers little benefit.
Original abstract
Text-based counseling is an important interface for AI mental-health support, where transcripts may be used to monitor depression severity and flag sessions requiring timely human review. However, robust PHQ-8 prediction across session regimes remains challenging: fine-tuning-based methods can exploit richer supervision but may generalize poorly under data scarcity, while prompt-based LLM methods are data-efficient but usually treat each transcript holistically and provide limited support for longitudinal context. We study robust depression tracking from counseling transcripts across single-session and multi-session regimes. We introduce LongCounsel, a multi-session counseling dataset with session-level PHQ-8 supervision for evaluating repeated-session tracking under partial symptom disclosure and cross-session continuity. We further propose EmoTrack, a PHQ-8 prediction framework that combines LLM-extracted clinical signals with frozen turn-level semantic embeddings and trains symptom-specific predictors over the resulting transcript representation. When prior sessions are available, EmoTrack can further incorporate them through compact cross-session memory. Experiments on LongCounsel and DAIC-WOZ show that EmoTrack achieves a clear gain on the real single-session benchmark, including a 13.5% relative MAE reduction over the strongest DAIC-WOZ baseline, and remains competitive with the strongest longitudinal baseline on LongCounsel.
Read the original paperMore in Natural Language Processing
Browse all 26 papers →AdaTutoRank: Learning to Rerank Document Sets via Adaptive Tutoring Optimization for RAG and Deep Research
Kailin Jiang, Lei Liu, Jian Xi, Yangqi Chen, Hui Xu, Hongwei Zhao, Bin Li, Yu Lu, Haibo Shi
AdaTutoRank teaches rerankers to assemble complementary evidence sets rather than merely picking individually relevant documents, improving RAG and deep-research retrieval with fewer calls.
Computation Over Geometry: Meaning Identity Is Computed, Not Shipped in the Embeddings
Jiaqi Deng
The paper argues that sentence meaning equivalence is not reliably stored in separate embeddings but is computed when models process both sentences together.
SlopShape: Identifying AI-Generated Commercial Web Content
Jochen Madler
SlopShape detects and identifies AI-written commercial content by recognizing its underlying organizational style, even after the text has been reworded.