NTH

Zing: Social Mind for LLMs

AuthorsZing Team, Ao Xiang, Bi Jingping, Chen Jiahui, Chen Lehan, Chen Yilin, Cheng Xueqi, Fan Yixing, Gan Kairong, Gao Haowen, Gao Jinhua, Gao Shuxuan, Gong Chang, Guo Jiafeng, Guo Ruijie, Han Zhouyu, He Guangfu, He Yichun, Jiang Shuo, Jing Shaoling, Jing Ya, Lei Chenhao, Lei Yan, Li Anqi, Li Chengao, Li Haoyu, Li Shitian, Liang Xinjian, Liu Zhaoge, Lyu Xingyu, Nie Zhuwei, Pang Liang, Quan Zeping, Shan Shiguang, Shen Huawei, Tang Xinran, Tian Feng, Wang Qian, Wang Ruiping, Wang Xiaohong, Xia Zaiyu, Xiao Yi, Xu Jiayuan, Xu Kehan, Xu Qianqian, Xu Tianyu, Xu Yongjun, Yang Haoming, Yang Jun, Yao Di, Yu Xiaoming, Zhang Futong, Zhang Jie, Zhang Shixuan, Zhang Yuxuan, Zhao Xinyu, Zhao Zhuoran, Zhong Yunfei, Zhu Shengyu

July 31, 2026 3 min read
Watch on YouTube
The one-line take

Zing gives LLMs a way to measure, learn, and apply social intelligence through specialized benchmarks, training, and runtime support.

Key results

3481
SoMBench expert-verified instances

The benchmark contains 3481 expert-verified instances across 284 scenarios.

72.08%
Best SoMBench overall accuracy

Anthropic’s Claude Opus 4.8 achieves the highest overall accuracy among 20 evaluated LLMs.

79.80%
Zing-27B-Stage2 five-benchmark average

Average score across HiToM, ToMBench, EmoBench, FANToM, and SoMBench.

14/15
Actio improved model-benchmark pairs

The full Actio harness improves 14 of 15 model–benchmark pairs.

3.70%
Actio average gain

Average improvement from deployment-time grounding across the evaluated pairs, reported as percentage points.

What the paper found

“Zing: Social Mind for LLMs” argues that socially capable language models need more than task reasoning: they must track beliefs, intentions, emotions, relationships, power, and context-sensitive norms. Its evaluation component, SoMBench, organizes social intelligence into 3 primary dimensions, 17 secondary dimensions, and 71 task paradigms across 284 scenarios and 3481 expert-verified instances. Testing 20 LLMs, including Anthropic’s Claude Opus 4.8, OpenAI’s GPT-5.5, Google’s Gemini 3.1 Pro, and DeepSeek-V4-Pro, finds that Claude Opus 4.8 leads at only 72.08% overall accuracy, with no secondary dimension reaching the 90% near-ceiling band. The Zing training recipe uses diagnosis-driven data iteration through FLARE, supervised fine-tuning, on-policy distillation, and rubric-based GRPO reinforcement learning on Qwen backbones; Zing-27B-Stage2 achieves a 79.80% average across HiToM, ToMBench, EmoBench, FANToM, and SoMBench, while Zing-32B-Stage2 is competitive with DeepSeek-V4-Pro. For deployment, Actio wraps frozen models with selectively routed supports: PRISM for procedural skills, Starling for explicit mental-state tracking, SAGE for reusable experience, and gated social-mental RAG for normative knowledge. Across five base models and three benchmarks, Actio improves 14 of 15 model–benchmark pairs, with an average gain of 3.70 percentage points. The central contribution is treating social intelligence as a coordinated pipeline of measurement, parametric internalization, and controlled runtime grounding rather than as a single prompting capability.

Original abstract

As large language models move from isolated task solving toward long-term service in human environments, they require social intelligence: the ability to infer mental states, track social relations, reason over norms, and adapt behavior under context. This report presents Zhijing, an integrated framework for measuring, internalizing, and grounding social intelligence. For measurement, we introduce SoMBench, a psychology-grounded benchmark spanning 3 primary dimensions, 17 secondary dimensions, and 71 task paradigms. It controls question format, narrative perspective, and context length across 284 shared scenarios and 3,481 expert-verified instances. Evaluation of 20 representative LLMs reveals substantial headroom: the best model achieves only 72.08% overall accuracy, and none of the 17 secondary dimensions reaches the 90% near-ceiling band. For internalization, we develop Zing, a diagnosis-driven training recipe combining supervised fine-tuning, on-policy distillation, and rubric-based reinforcement learning. Across five social-cognition benchmarks, Zing consistently outperforms its base models, with Zing-27B-Stage2 achieving the best average score and Zing-32B-Stage2 remaining competitive with DeepSeek-V4-Pro. For deployment-time grounding, we build Actio, a harness-controlled inference architecture that routes four typed supports into reasoning: PRISM for procedural guidance, Starling for runtime mental-state representation, SAGE for reusable experience, and gated RAG for external social and normative knowledge. Across five base models and three benchmarks, the full harness improves 14 of 15 model-benchmark pairs and is best or tied for best in 8, demonstrating the effectiveness of typed runtime support. Together, these results show that socially intelligent LLMs require coordinated advances in evaluation, parametric internalization, and deployment-time grounding.

Read the original paper

More in Large Language Models

Browse all 81 papers →
02Llm

Generalization Dynamics of LM Pre-training

Jiaxin Wen, Zhengxuan Wu, Dawn Song, Lijie Chen

Language models may repeatedly switch between shallow memorization and genuine reasoning during training, and the paper shows how to detect and potentially control these swings.

Read analysis
03Llm

Rethinking Self-Distillation for Multi-Teacher Capability Merging

Roy Xie, Dan Friedman, Feng Nan, Yukun Huang, Zhichao Xu, Chengjiu Zhang, Jun Xu, Manaal Faruqui, Vivek Rathod, Bhuwan Dhingra

The study finds that expensive multi-teacher on-policy distillation may offer little advantage over carefully tuned, cheaper alternatives such as SFT and weight merging.

Read analysis