NTH

CreativeInstruct: Scalably Teaching LLMs to Balance Quality, Creativity, and Diversity

AuthorsAnanya Sahu, Mohit Bansal, Elias Stengel-Eskin

August 14, 2026 2 min read
Watch on YouTube
The one-line take

CreativeInstruct teaches LLMs to produce more diverse and creative outputs without sacrificing quality, and shows that this creativity can also improve downstream reinforcement learning.

Key results

4000
Training prompts

English writing prompts drawn from Tülu V3 SFT for data construction

12000
Training samples

Synthetic instruction-tuning samples generated from 3 outputs per prompt

48%
LLaMA-3.1 8B semantic diversity gain

Relative improvement over the base instruct model

63%
LLaMA-3.1 8B structural diversity gain

Relative improvement measured with narrative-level structural diversity

70.3%
Human creativity preference

Comparisons in which evaluators rated CreativeInstruct more creative than instruct outputs

4%
AMC gain after GRPO

Improvement from starting GRPO with CreativeInstruct Qwen3 8B instead of the instruct checkpoint

What the paper found

CreativeInstruct addresses a central post-training problem: instruction tuning improves compliance but often makes LLM outputs repetitive and less exploratory. It uses BACo, an entropy- and token-type-based router between base and aligned models, to generate training responses, then marks base-model spans with [StartCreativity] and [EndCreativity] tokens. LoRA fine-tuning teaches one aligned checkpoint to insert these spans selectively at inference time, recovering creativity without loading a second model. The method is evaluated on Meta’s LLaMA-3.1 and Alibaba’s Qwen families from 7B to 32B parameters, using 4,000 English writing prompts expanded into 12,000 training samples. To measure narrative variation beyond lexical overlap, the paper introduces LLM Graph-Edit-Distance, which compares canonicalized event graphs containing entities, events, semantic roles, and temporal edges. On LLaMA-3.1 8B, CreativeInstruct produces approximately 48% relative gains in semantic diversity and 63% gains in structural diversity over the instruct model, while maintaining competitive coherence, fluency, relevance, and writing-reward scores. Human evaluators preferred its generations as more creative in 70.3% of comparisons. The approach also improves reinforcement learning: GRPO on the CreativeInstruct Qwen3 8B checkpoint gains 4% on AMC and 5 percentage points on MATH compared with GRPO starting from the standard instruct checkpoint. GPT-5-mini-based evaluation and cross-family results suggest that creativity tags learn an internal switching policy that generalizes beyond the routing data and avoids dual-model inference cost.

Original abstract

While post-training improves the capabilities of large language models (LLMs), it generally lowers their output diversity and creativity, negatively impacting tasks that explicitly require creativity (e.g., story generation) as well as those that require it implicitly, e.g., reinforcement learning (RL). We instead propose CreativeInstruct, a scalable instruction-tuning method that teaches LLMs to balance creative, base-model-like generations with the quality of post-trained models, by learning to inject special [StartCreativity] spans that bias generation toward creativity. Furthermore, we introduce a structural diversity metric based on graph edit distance, which captures narrative level variation missed by purely lexical and semantic metrics. On narrative generation, CreativeInstruct matches or exceeds the diversity of both multi-model baselines and distilled variants of their outputs, without sacrificing quality or requiring multiple models at inference time. These results are mirrored in our human evaluation, where we find that annotators rate CreativeInstruct generations as more creative than the post-trained LLMs' generations in 70.3% of cases. We also show the benefits of creative models as a substrate for RL: GRPO applied to a CreativeInstruct checkpoint improves by ~4% on AMC and ~5% points on MATH over the same training applied to the post-trained checkpoint.

Read the original paper

More in Large Language Models

Browse all 81 papers →
02Llm

Generalization Dynamics of LM Pre-training

Jiaxin Wen, Zhengxuan Wu, Dawn Song, Lijie Chen

Language models may repeatedly switch between shallow memorization and genuine reasoning during training, and the paper shows how to detect and potentially control these swings.

Read analysis
03Llm

Rethinking Self-Distillation for Multi-Teacher Capability Merging

Roy Xie, Dan Friedman, Feng Nan, Yukun Huang, Zhichao Xu, Chengjiu Zhang, Jun Xu, Manaal Faruqui, Vivek Rathod, Bhuwan Dhingra

The study finds that expensive multi-teacher on-policy distillation may offer little advantage over carefully tuned, cheaper alternatives such as SFT and weight merging.

Read analysis