NTH

LLM-as-a-Tutor: Policy-Aware Prompt Adaptation for Non-Verifiable RL

AuthorsYujin Kim, Namgyu Ho, Sangmin Hwang, Joonkee Kim, Yongjin Yang, Sangmin Bae, Seungone Kim, Jaehun Jung, Se-Young Yun, Hwanjun Song

July 12, 2026 2 min read
Watch on YouTube
The one-line take

This paper turns an LLM from a static grader into a dynamic tutor that makes training prompts progressively harder so reinforcement learning stays informative as the policy improves.

Key results

40.91
FollowBench HSR

Best reported HSR score with LLM-as-a-Tutor

62.28
FollowBench SSR

Best reported SSR score with LLM-as-a-Tutor

15.07
AdvancedIF Overall

Best reported Overall score with LLM-as-a-Tutor

73.59
InfoBench DRFR

Best reported DRFR score with LLM-as-a-Tutor

51.96
Average score

Mean across the three benchmarks for LLM-as-a-Tutor

8.1% to 40.5%
Constraint-added ratio

Constraint augmentation rises as policy scale increases from 0.6B to 4B

What the paper found

LLM-as-a-Tutor is a policy-aware RL framework for non-verifiable instruction following that uses one Qwen3-8B model as both examiner and generator to detect when a prompt is too easy for the current policy and then append a single atomic constraint that makes the task harder without rewriting the original instruction. Trained with GRPO for 3 epochs on 4K WildChat prompts, using Qwen3-1.7B-Thinking as the policy and evaluating on FollowBench, AdvancedIF, and InfoBench, it achieves the strongest overall instruction-following performance while preserving the seed distribution. On the reported benchmarks, it reaches 40.91 on FollowBench HSR, 62.28 on SSR, 15.07 on AdvancedIF Overall, 67.97 on AdvancedIF Micro, 73.59 on InfoBench DRFR, and 51.96 average, beating policy-unaware baselines, prompt-rewriting methods like Evol-Instruct, and the policy-adaptive EVA variant. The key technical insight is that rubric refinement alone cannot recover reward variance when a prompt saturates the policy, so the tutor performs pairwise rollout comparisons to trigger prompt augmentation only when quality is indistinguishable. Ablations show that the targeted append strategy matters: Always, Random, and Wrong fall to 42.13, 42.82, and 42.79 average, while the adaptive append design reaches 43.19 in the ablation setting. Analysis further shows constraint-added prompts rise from 8.1% to 40.5% as policy scale grows from 0.6B to 4B, confirming that adaptation tracks policy capability rather than a fixed schedule.

Original abstract

Reinforcement learning (RL) for non-verifiable instruction following increasingly relies on LLM judges with prompt-specific rubrics as reward signals. While recent methods adapt these rubrics to the evolving policy during training, the training prompts themselves remain static, drawn from fixed corpora. This static approach often results in a critical misalignment between prompt difficulty and policy capability, leaving the judge unable to recover a discriminative reward signal when prompts fail to elicit quality variance among rollouts. To address this misalignment, we introduce LLM-as-a-Tutor, a framework that extends the LLM's role from judge to tutor: a single model serves as an examiner that pairwise-compares policy rollouts to detect non-challenging prompts, and as a generator that appends atomic constraints to them. This append-only design monotonically raises difficulty in step with the policy's capability, producing a self-calibrating training signal without external difficulty schedules. On three complex instruction-following benchmarks, our method consistently outperforms both policy-unaware baselines and prior policy-adaptive methods that adapt rubrics or rewrite prompts, suggesting prompt adaptation as a missing axis of policy-awareness in non-verifiable RL.

Read the original paper

More in Reinforcement Learning

Browse all 54 papers →