ESC-Skills: Discovering and Self-Evolving Skills for Emotional Support Conversations
AuthorsJie Zhu, Huaixia Dou, Shuo Jiang, Junhui Li, Lifan Guo, Feng Chen, Chi Zhang, Fang Kong
ESC-Skills turns emotional support chat into a set of explicit, self-improving skills so the assistant can respond more safely, consistently, and helpfully.
Key results
Successful dialogues used to extract intervention patterns
Unsuccessful dialogues used to extract intervention patterns
Total localized state-action-outcome instances extracted
Recurring (seeker state, support action) groups after filtering
Executable emotional support skills in B0
Refined executable emotional support skills in B⋆
What the paper found
ESC-Skills, developed by Soochow University and Alibaba Cloud’s Qwen DianJin Team, reframes emotional support conversation as an intervention-driven process rather than a pure response-generation task. The paper introduces Intervention Units, or IUs, to model localized state–action–outcome dynamics from both successful and failed dialogues, then converts recurring patterns into an executable ESC-Skills Bank with structured SKILL.md packages. Starting from ESConv’s 910 training conversations and FailedESConv’s 196 unsuccessful conversations, the authors extract 17,858 IUs, distill 258 skill prototypes, and build an initial bank of 27 skills, later refined to 34 skills through a multi-profile self-evolution loop using 500 RLVER seeker profiles and SAGE evaluation. On ESConv and SAGE, the final bank improves all tested LLM backbones, with Qwen3.6-Plus rising from 11.50 to 23.56 ACC on ESConv and from 66.4 to 72.1 average sentient score on SAGE, while successful dialogues increase from 13 to 31. The strongest baseline contrast is important: static or one-pass skill generation yields marginal gains or can even hurt long-horizon performance, but the generation–verification refinement loop filters ineffective interventions and produces more robust, controllable, and interpretable support behavior. The work also reports GPT-5.4-based judge scores and human ratings that align with the automatic metrics, supporting the claim that explicit, editable emotional-support skills are more reliable than implicit prompting for sustained counseling-style dialogue.
Original abstract
Existing emotional support conversation (ESC) systems mainly rely on end-to-end response generation or coarse strategy supervision, offering limited interpretability and little support for systematic skill improvement. We propose ESC-Skills, a skill-centric framework that discovers and self-evolves executable emotional support skills. We first model localized support interactions as Intervention Units (IUs), which capture state--action--outcome dynamics between seeker states, support interventions, and post-response emotional changes. Based on IUs extracted from both successful and failed ESC dialogues, we construct the ESC-Skills Bank, a repository of executable emotional support skills containing intervention guidance, applicability conditions, expected outcomes, and potential risks. To further improve robustness, we introduce a multi-profile self-evolutionary refinement framework in which an ESC agent interacts with diverse simulated seeker profiles under SAGE evaluation. The resulting interaction traces are analyzed to identify missing skills, unsafe interventions, and profile-specific failure patterns, which are then used to refine the Skills Bank through simulation-based verification. Experimental results demonstrate that ESC-Skills improves both response-level quality and dialogue-level emotional outcomes while providing more interpretable and controllable support behaviors. We will release the code, prompts, and ESC-Skills Bank at https://github.com/aliyun/qwen-dianjin.
Read the original paperMore in Natural Language Processing
Browse all 26 papers →AdaTutoRank: Learning to Rerank Document Sets via Adaptive Tutoring Optimization for RAG and Deep Research
Kailin Jiang, Lei Liu, Jian Xi, Yangqi Chen, Hui Xu, Hongwei Zhao, Bin Li, Yu Lu, Haibo Shi
AdaTutoRank teaches rerankers to assemble complementary evidence sets rather than merely picking individually relevant documents, improving RAG and deep-research retrieval with fewer calls.
Computation Over Geometry: Meaning Identity Is Computed, Not Shipped in the Embeddings
Jiaqi Deng
The paper argues that sentence meaning equivalence is not reliably stored in separate embeddings but is computed when models process both sentences together.
SlopShape: Identifying AI-Generated Commercial Web Content
Jochen Madler
SlopShape detects and identifies AI-written commercial content by recognizing its underlying organizational style, even after the text has been reworded.