NTH

ESC-Skills: Discovering and Self-Evolving Skills for Emotional Support Conversations

AuthorsJie Zhu, Huaixia Dou, Shuo Jiang, Junhui Li, Lifan Guo, Feng Chen, Chi Zhang, Fang Kong

June 13, 2026 2 min read
Watch on YouTube
The one-line take

ESC-Skills turns emotional support chat into a set of explicit, self-improving skills so the assistant can respond more safely, consistently, and helpfully.

Key results

910
ESConv training conversations

Successful dialogues used to extract intervention patterns

196
FailedESConv conversations

Unsuccessful dialogues used to extract intervention patterns

17858
Intervention Units

Total localized state-action-outcome instances extracted

258
Skill prototypes

Recurring (seeker state, support action) groups after filtering

27
Initial skill bank

Executable emotional support skills in B0

34
Final skill bank

Refined executable emotional support skills in B⋆

What the paper found

ESC-Skills, developed by Soochow University and Alibaba Cloud’s Qwen DianJin Team, reframes emotional support conversation as an intervention-driven process rather than a pure response-generation task. The paper introduces Intervention Units, or IUs, to model localized state–action–outcome dynamics from both successful and failed dialogues, then converts recurring patterns into an executable ESC-Skills Bank with structured SKILL.md packages. Starting from ESConv’s 910 training conversations and FailedESConv’s 196 unsuccessful conversations, the authors extract 17,858 IUs, distill 258 skill prototypes, and build an initial bank of 27 skills, later refined to 34 skills through a multi-profile self-evolution loop using 500 RLVER seeker profiles and SAGE evaluation. On ESConv and SAGE, the final bank improves all tested LLM backbones, with Qwen3.6-Plus rising from 11.50 to 23.56 ACC on ESConv and from 66.4 to 72.1 average sentient score on SAGE, while successful dialogues increase from 13 to 31. The strongest baseline contrast is important: static or one-pass skill generation yields marginal gains or can even hurt long-horizon performance, but the generation–verification refinement loop filters ineffective interventions and produces more robust, controllable, and interpretable support behavior. The work also reports GPT-5.4-based judge scores and human ratings that align with the automatic metrics, supporting the claim that explicit, editable emotional-support skills are more reliable than implicit prompting for sustained counseling-style dialogue.

Original abstract

Existing emotional support conversation (ESC) systems mainly rely on end-to-end response generation or coarse strategy supervision, offering limited interpretability and little support for systematic skill improvement. We propose ESC-Skills, a skill-centric framework that discovers and self-evolves executable emotional support skills. We first model localized support interactions as Intervention Units (IUs), which capture state--action--outcome dynamics between seeker states, support interventions, and post-response emotional changes. Based on IUs extracted from both successful and failed ESC dialogues, we construct the ESC-Skills Bank, a repository of executable emotional support skills containing intervention guidance, applicability conditions, expected outcomes, and potential risks. To further improve robustness, we introduce a multi-profile self-evolutionary refinement framework in which an ESC agent interacts with diverse simulated seeker profiles under SAGE evaluation. The resulting interaction traces are analyzed to identify missing skills, unsafe interventions, and profile-specific failure patterns, which are then used to refine the Skills Bank through simulation-based verification. Experimental results demonstrate that ESC-Skills improves both response-level quality and dialogue-level emotional outcomes while providing more interpretable and controllable support behaviors. We will release the code, prompts, and ESC-Skills Bank at https://github.com/aliyun/qwen-dianjin.

Read the original paper

More in Natural Language Processing

Browse all 26 papers →