StudentSim: Training LLM-based Student Simulators
AuthorsKe Yang, Chenglong Wang, Michel Galley, Chandan Singh, Jeevana Priya Inala, ChengXiang Zhai, Jianfeng Gao
StudentSim teaches LLMs to imitate individual learners and respond realistically to tutoring, enabling more personalized AI tutors.
Key results
Standardized evaluation across chess, second-language English writing, and mathematics.
StudentSim’s held-out chess move-prediction score using Qwen3-4B-Instruct with per-student LoRA specialization.
Rate at which StudentSim reaches the guided canonical move after tutor feedback.
StudentSim’s held-out answer-correction rate after mathematical guidance.
Expert-rated chess-tutor responses without actively misleading factual errors after GRPO reinforcement learning.
What the paper found
StudentSim addresses a central weakness in AI tutoring: prompted role-play models such as OpenAI’s GPT-5.4 can respond fluently to guidance but often fail to reproduce a specific learner’s mistakes, while state-tracking systems such as Maia2 imitate behavior but cannot process natural-language explanations. The proposed framework uses two-stage training: pooled domain training first learns shared response and correction patterns, then per-student LoRA specialization adapts Qwen3-4B-Instruct to sparse records from an individual learner. StudentSim Eval standardizes held-out testing for 60 students across chess, second-language English writing, and mathematics, separating behavioral fidelity, F, from guidance responsiveness, R. On chess, StudentSim achieves F = 0.5150 and R = 0.9067, versus GPT-5.4 at F = 0.2316 and R = 0.7186, and Maia2 at F = 0.4535 and R = 0.2721. The method also leads in L2 writing and math, reaching math R = 0.9181, showing that multi-turn training teaches the simulator to revise answers under direct, comparative, Socratic, or conceptual guidance. As a downstream test, the researchers use the frozen simulator as a reward source for GRPO-based chess-tutor reinforcement learning; in an expert study, that tutor reaches 90.5% factual accuracy, compared with 75.7% without reinforcement learning and 71.6% using a GPT-5.4 simulator reward. The result is a practical, learner-grounded proxy for optimizing adaptive tutors, although modeling long-term learning, retention, and forgetting remains unresolved.
Original abstract
AI tutors are most useful when they adapt to each student's strengths, weaknesses, and preferred guidance, but evidence about which guidance works for which student is sparse, slow, and costly to collect from real learners. Student simulators can provide this signal as a proxy, yet existing approaches are limited: state-tracking models fit student behavior but struggle to process explanations or corrections, while LLM role-play follows guidance fluently but does not reliably match the competence of the student being imitated. We present StudentSim, a training framework that turns sparse per-student data into individualized simulators through pooled training followed by per-student specialization. The resulting simulators both mirror a student's own responses and update them under tutor guidance. We also introduce StudentSimEval, a standardized protocol covering 60 students across chess, second-language English writing, and mathematics, using public learner datasets with de-identified records shared for research. StudentSimEval measures behavioral fidelity (F), or how well a simulator matches a student's responses, and guidance responsiveness (R), or how readily it updates under tutor guidance, with all methods fit and evaluated on the same records. Across all three domains, StudentSim outperforms GPT-5.4 on both metrics. In chess, StudentSim reaches F=0.51 and R=0.91, compared with 0.23 and 0.72 for GPT-5.4 and 0.45 and 0.27 for Maia2. As a proof of concept, using StudentSim as a reward model for tutor reinforcement learning produces a chess tutor that expert humans rate as more accurate, better-guided, and more personalized than a no-RL baseline and a tutor trained against a GPT-5.4 simulator reward. Code is available at https://github.com/microsoft/StudentSim.
Read the original paperMore in Large Language Models
Browse all 81 papers →Finetuning with Sampling: SFT Learns Better Than You Think
Aayush Karan, Sitan Chen, Yilun Du
By sampling and reshaping expert data before training, this work argues that supervised finetuning can match RL while generalizing better and forgetting less.
Generalization Dynamics of LM Pre-training
Jiaxin Wen, Zhengxuan Wu, Dawn Song, Lijie Chen
Language models may repeatedly switch between shallow memorization and genuine reasoning during training, and the paper shows how to detect and potentially control these swings.
Rethinking Self-Distillation for Multi-Teacher Capability Merging
Roy Xie, Dan Friedman, Feng Nan, Yukun Huang, Zhichao Xu, Chengjiu Zhang, Jun Xu, Manaal Faruqui, Vivek Rathod, Bhuwan Dhingra
The study finds that expensive multi-teacher on-policy distillation may offer little advantage over carefully tuned, cheaper alternatives such as SFT and weight merging.