CreativeInstruct: Scalably Teaching LLMs to Balance Quality, Creativity, and Diversity
AuthorsAnanya Sahu, Mohit Bansal, Elias Stengel-Eskin
Resources
CreativeInstruct teaches LLMs to produce more diverse and creative outputs without sacrificing quality, and shows that this creativity can also improve downstream reinforcement learning.
Key results
English writing prompts drawn from Tülu V3 SFT for data construction
Synthetic instruction-tuning samples generated from 3 outputs per prompt
Relative improvement over the base instruct model
Relative improvement measured with narrative-level structural diversity
Comparisons in which evaluators rated CreativeInstruct more creative than instruct outputs
Improvement from starting GRPO with CreativeInstruct Qwen3 8B instead of the instruct checkpoint
What the paper found
CreativeInstruct addresses a central post-training problem: instruction tuning improves compliance but often makes LLM outputs repetitive and less exploratory. It uses BACo, an entropy- and token-type-based router between base and aligned models, to generate training responses, then marks base-model spans with [StartCreativity] and [EndCreativity] tokens. LoRA fine-tuning teaches one aligned checkpoint to insert these spans selectively at inference time, recovering creativity without loading a second model. The method is evaluated on Meta’s LLaMA-3.1 and Alibaba’s Qwen families from 7B to 32B parameters, using 4,000 English writing prompts expanded into 12,000 training samples. To measure narrative variation beyond lexical overlap, the paper introduces LLM Graph-Edit-Distance, which compares canonicalized event graphs containing entities, events, semantic roles, and temporal edges. On LLaMA-3.1 8B, CreativeInstruct produces approximately 48% relative gains in semantic diversity and 63% gains in structural diversity over the instruct model, while maintaining competitive coherence, fluency, relevance, and writing-reward scores. Human evaluators preferred its generations as more creative in 70.3% of comparisons. The approach also improves reinforcement learning: GRPO on the CreativeInstruct Qwen3 8B checkpoint gains 4% on AMC and 5 percentage points on MATH compared with GRPO starting from the standard instruct checkpoint. GPT-5-mini-based evaluation and cross-family results suggest that creativity tags learn an internal switching policy that generalizes beyond the routing data and avoids dual-model inference cost.
Original abstract
While post-training improves the capabilities of large language models (LLMs), it generally lowers their output diversity and creativity, negatively impacting tasks that explicitly require creativity (e.g., story generation) as well as those that require it implicitly, e.g., reinforcement learning (RL). We instead propose CreativeInstruct, a scalable instruction-tuning method that teaches LLMs to balance creative, base-model-like generations with the quality of post-trained models, by learning to inject special [StartCreativity] spans that bias generation toward creativity. Furthermore, we introduce a structural diversity metric based on graph edit distance, which captures narrative level variation missed by purely lexical and semantic metrics. On narrative generation, CreativeInstruct matches or exceeds the diversity of both multi-model baselines and distilled variants of their outputs, without sacrificing quality or requiring multiple models at inference time. These results are mirrored in our human evaluation, where we find that annotators rate CreativeInstruct generations as more creative than the post-trained LLMs' generations in 70.3% of cases. We also show the benefits of creative models as a substrate for RL: GRPO applied to a CreativeInstruct checkpoint improves by ~4% on AMC and ~5% points on MATH over the same training applied to the post-trained checkpoint.
Read the original paperMore in Large Language Models
Browse all 81 papers →Finetuning with Sampling: SFT Learns Better Than You Think
Aayush Karan, Sitan Chen, Yilun Du
By sampling and reshaping expert data before training, this work argues that supervised finetuning can match RL while generalizing better and forgetting less.
Generalization Dynamics of LM Pre-training
Jiaxin Wen, Zhengxuan Wu, Dawn Song, Lijie Chen
Language models may repeatedly switch between shallow memorization and genuine reasoning during training, and the paper shows how to detect and potentially control these swings.
Rethinking Self-Distillation for Multi-Teacher Capability Merging
Roy Xie, Dan Friedman, Feng Nan, Yukun Huang, Zhichao Xu, Chengjiu Zhang, Jun Xu, Manaal Faruqui, Vivek Rathod, Bhuwan Dhingra
The study finds that expensive multi-teacher on-policy distillation may offer little advantage over carefully tuned, cheaper alternatives such as SFT and weight merging.