Zing: Social Mind for LLMs
AuthorsZing Team, Ao Xiang, Bi Jingping, Chen Jiahui, Chen Lehan, Chen Yilin, Cheng Xueqi, Fan Yixing, Gan Kairong, Gao Haowen, Gao Jinhua, Gao Shuxuan, Gong Chang, Guo Jiafeng, Guo Ruijie, Han Zhouyu, He Guangfu, He Yichun, Jiang Shuo, Jing Shaoling, Jing Ya, Lei Chenhao, Lei Yan, Li Anqi, Li Chengao, Li Haoyu, Li Shitian, Liang Xinjian, Liu Zhaoge, Lyu Xingyu, Nie Zhuwei, Pang Liang, Quan Zeping, Shan Shiguang, Shen Huawei, Tang Xinran, Tian Feng, Wang Qian, Wang Ruiping, Wang Xiaohong, Xia Zaiyu, Xiao Yi, Xu Jiayuan, Xu Kehan, Xu Qianqian, Xu Tianyu, Xu Yongjun, Yang Haoming, Yang Jun, Yao Di, Yu Xiaoming, Zhang Futong, Zhang Jie, Zhang Shixuan, Zhang Yuxuan, Zhao Xinyu, Zhao Zhuoran, Zhong Yunfei, Zhu Shengyu
Resources
Zing gives LLMs a way to measure, learn, and apply social intelligence through specialized benchmarks, training, and runtime support.
Key results
The benchmark contains 3481 expert-verified instances across 284 scenarios.
Anthropic’s Claude Opus 4.8 achieves the highest overall accuracy among 20 evaluated LLMs.
Average score across HiToM, ToMBench, EmoBench, FANToM, and SoMBench.
The full Actio harness improves 14 of 15 model–benchmark pairs.
Average improvement from deployment-time grounding across the evaluated pairs, reported as percentage points.
What the paper found
“Zing: Social Mind for LLMs” argues that socially capable language models need more than task reasoning: they must track beliefs, intentions, emotions, relationships, power, and context-sensitive norms. Its evaluation component, SoMBench, organizes social intelligence into 3 primary dimensions, 17 secondary dimensions, and 71 task paradigms across 284 scenarios and 3481 expert-verified instances. Testing 20 LLMs, including Anthropic’s Claude Opus 4.8, OpenAI’s GPT-5.5, Google’s Gemini 3.1 Pro, and DeepSeek-V4-Pro, finds that Claude Opus 4.8 leads at only 72.08% overall accuracy, with no secondary dimension reaching the 90% near-ceiling band. The Zing training recipe uses diagnosis-driven data iteration through FLARE, supervised fine-tuning, on-policy distillation, and rubric-based GRPO reinforcement learning on Qwen backbones; Zing-27B-Stage2 achieves a 79.80% average across HiToM, ToMBench, EmoBench, FANToM, and SoMBench, while Zing-32B-Stage2 is competitive with DeepSeek-V4-Pro. For deployment, Actio wraps frozen models with selectively routed supports: PRISM for procedural skills, Starling for explicit mental-state tracking, SAGE for reusable experience, and gated social-mental RAG for normative knowledge. Across five base models and three benchmarks, Actio improves 14 of 15 model–benchmark pairs, with an average gain of 3.70 percentage points. The central contribution is treating social intelligence as a coordinated pipeline of measurement, parametric internalization, and controlled runtime grounding rather than as a single prompting capability.
Original abstract
As large language models move from isolated task solving toward long-term service in human environments, they require social intelligence: the ability to infer mental states, track social relations, reason over norms, and adapt behavior under context. This report presents Zhijing, an integrated framework for measuring, internalizing, and grounding social intelligence. For measurement, we introduce SoMBench, a psychology-grounded benchmark spanning 3 primary dimensions, 17 secondary dimensions, and 71 task paradigms. It controls question format, narrative perspective, and context length across 284 shared scenarios and 3,481 expert-verified instances. Evaluation of 20 representative LLMs reveals substantial headroom: the best model achieves only 72.08% overall accuracy, and none of the 17 secondary dimensions reaches the 90% near-ceiling band. For internalization, we develop Zing, a diagnosis-driven training recipe combining supervised fine-tuning, on-policy distillation, and rubric-based reinforcement learning. Across five social-cognition benchmarks, Zing consistently outperforms its base models, with Zing-27B-Stage2 achieving the best average score and Zing-32B-Stage2 remaining competitive with DeepSeek-V4-Pro. For deployment-time grounding, we build Actio, a harness-controlled inference architecture that routes four typed supports into reasoning: PRISM for procedural guidance, Starling for runtime mental-state representation, SAGE for reusable experience, and gated RAG for external social and normative knowledge. Across five base models and three benchmarks, the full harness improves 14 of 15 model-benchmark pairs and is best or tied for best in 8, demonstrating the effectiveness of typed runtime support. Together, these results show that socially intelligent LLMs require coordinated advances in evaluation, parametric internalization, and deployment-time grounding.
Read the original paperMore in Large Language Models
Browse all 81 papers →Finetuning with Sampling: SFT Learns Better Than You Think
Aayush Karan, Sitan Chen, Yilun Du
By sampling and reshaping expert data before training, this work argues that supervised finetuning can match RL while generalizing better and forgetting less.
Generalization Dynamics of LM Pre-training
Jiaxin Wen, Zhengxuan Wu, Dawn Song, Lijie Chen
Language models may repeatedly switch between shallow memorization and genuine reasoning during training, and the paper shows how to detect and potentially control these swings.
Rethinking Self-Distillation for Multi-Teacher Capability Merging
Roy Xie, Dan Friedman, Feng Nan, Yukun Huang, Zhichao Xu, Chengjiu Zhang, Jun Xu, Manaal Faruqui, Vivek Rathod, Bhuwan Dhingra
The study finds that expensive multi-teacher on-policy distillation may offer little advantage over carefully tuned, cheaper alternatives such as SFT and weight merging.