Beyond Starry Night: Shortcut-Aware Control-State Planning for Artist-Grounded Text to Image Generation
AuthorsKuan Xing, Ye Wang, Changyi Gan, Yuheng Li, Thao Nguyen, Yi Chang, Yilin Wang
Resources
Atelier aims to make artist-inspired image generation more faithful by planning around artistic intent while avoiding the visual shortcuts that models commonly associate with famous artists.
Key results
Atelier’s closed-source configuration achieved the best artist-level style proximity on ArtIntentBench.
Share of outputs containing at least one unrequested canonical shortcut.
Best open-weight direct baseline rate, reported for Hunyuan Image 3.0.
Patch examples used to train the Gemma-4 LoRA authenticity critic.
Accuracy on 1,485 held-out patches for distinguishing real Van Gogh patches from others.
What the paper found
The paper introduces Atelier, a shortcut-aware control-state planning framework for artist-grounded text-to-image generation. Its central claim is that naming an artist is not enough: models such as FLUX, Qwen-Image, LongCat-Image, Hunyuan Image 3.0, OpenAI’s GPT-Image-2, and Google’s Nano Banana Pro often replace the requested scene with canonical cues—for example, a Starry Night-like sky or unrequested Qi Baishi ink landscapes. Atelier instead converts an underspecified request into a structured state covering scene anchors, preserve-versus-transform decisions, period or style-regime hypotheses, role-bound artwork patches, backend controls, and anti-shortcut constraints. It retrieves global artist knowledge and Grounded-SAM-extracted local patches, compiles backend-specific plans, then uses Kimi K2.6 as a global critic and a LoRA-adapted Gemma-4 AuthCritic for patch-level authenticity feedback across iterative rounds. On ArtIntentBench, which evaluates Van Gogh and Qi Baishi through re-rendering, period control, unseen subjects, and shortcut auditing, the closed-source Atelier configuration reaches an IntroStyle W2 of 64.22 for Van Gogh re-rendering, while the open-weight configuration reduces Van Gogh shortcut substitution to 31.56%, versus 48.04% for the strongest direct baseline. AuthCritic is trained on 15,000 patches and reaches 91.65% binary accuracy on 1,485 held-out patches. The framework also transfers to Qi Baishi, although sparse patch-bank coverage remains a limitation; results show that structured artistic control, not merely longer prompts or stronger generators, is the main novelty.
Original abstract
Artist-grounded image generation requires more than appending an artist name to a prompt. Image models often respond to artist names through canonical shortcuts, such as recurring motifs, generic palettes, or overrepresented period signatures, rather than preserving the user's intended scene. We introduce Atelier, a shortcut-aware control-state planning framework for artist-grounded image generation. Atelier translates underspecified artistic intent into an explicit control state that separates scene anchors, preserve/transform decisions, style-regime hypotheses, role-bound artist evidence, and shortcut-avoidance constraints. It grounds this state using artist-level knowledge and local patch references, compiles backend-aware generation plans, and iteratively refines candidates through global and local authenticity feedback. We further introduce ArtIntentBench, a benchmark covering Van Gogh and Qi Baishi across artwork re-rendering, period/style-controlled generation, historically unseen subjects, shortcut auditing, and human preference evaluation. Across open-weight and closed-source generators, Atelier improves artist-level style fidelity, preserves source structure more faithfully, and substantially reduces shortcut substitution compared with prompt-engineered, retrieval-augmented, and general-purpose agent baselines. These results suggest that artist-grounded generation is bottlenecked not only by image synthesis, but by the upstream inference of explicit, evidence-grounded artistic controls.
Read the original paperMore in Generative Models
Browse all 63 papers →RULER: Instance-aware Rubric Rewards for SVG Generation
Hangyu Ran, Yuhao Zheng, Yingying Zhang, Kevin Qinghong Lin, Han Peng
RULER uses instruction-specific visual rubrics as reinforcement-learning rewards to make SVG generation more faithful, stylish, and resistant to reward hacking.
Think Before You Score: Thinking Reward Model for Visual Generation
Xuehai Bai, Zhenchen Tang, Yang Shi, Dianyi Wang, Tengfei Liu, Wanshun Su, Xuanyu Zhu, Ruohui Wang, Haiwen Diao, Haotian Wang, Xiaoling Gu, Yuanxing Zhang
A visual reward model that first decides what matters in each image-generation case, then scores outputs with detailed rubrics to provide better training signals.
WanPE: Towards Cinematic Prompt Enhancement for Modern Text-to-Video Generation
Yubo Zhu, Yawen Shao, Ziyun Dai, Zixun Fang, Kai Zhu, Siyang Sun, Haolan Xue, Chuxin Wang, Tingyu Weng, Jingming Luo, Chen Shi, Lianghua Huang, Yufeng Ai, Yuzheng Wang, Wenyuan Zhang, Yu Shang, Yuxiang Bao, Zoubin Bi, Jie Xiao, Jinbo Xing, Jiaxing Zhao, Chongyang Zhong, Hengjian Chen, Chenwei Xie, Akide Liu, Zhehan Kan, Yu Liu, Wei Zhai, Sheng Zhong, Wei Tong
WanPE turns ordinary text prompts into director-level cinematic plans, substantially improving the quality and consistency of long-form AI-generated videos.