What to Edit Next: Visually Aligned Image-Editing Follow-Up Suggestions in Conversational Systems
AuthorsZhijing Zhang, Jinpeng Yu, Xin Song, Bingnan Li, Chuyue Li, Changhui Du, Xiaolin Fang, Jiaming Liu, Ruihua Huang
Resources
A multimodal assistant learns to suggest visually consistent next edits, substantially improving user engagement in large-scale image-creation conversations.
Key results
Real multi-turn image-creation interactions used to assess visual dependence.
Share of follow-up editing queries requiring grounding in the latest image.
Final inconsistency rate after adding the source-target verifier, down from 3.7% after click-based RL.
Relative improvement over the prompt-engineered policy in the live A/B test.
Relative increase in users keeping the edited image.
Relative increase in average conversation turns per user.
What the paper found
This Alibaba Qwen App study addresses a gap in conversational image creation: follow-up edit suggestions must reflect user preferences while remaining executable on the latest image. An audit of 100,000 real multi-turn interactions found that 80.1% of follow-up edits depend on visual context. The proposed three-stage pipeline first uses a human-reviewed catalog of 61 editing intents, validated Gemini 3 Flash candidates, and supervised fine-tuning of a Qwen3-VL-8B policy with rank-4 LoRA. Next, position-aware click pairs train an 8B vision-language reward model with the Bradley–Terry objective, and multi-objective GRPO optimizes click preference, formatting, perplexity, length, and diversity. Finally, a Qwen3-VL-30B-A3B image-first verifier decomposes each suggestion into required sources and target states, adding visual grounding as a sixth reinforcement-learning reward without increasing serving latency. Click optimization alone worsened visual inconsistency from 3.7% to 0.9% after the verifier was added, while expert-rated quality and slate diversity improved. In a 14-day randomized Qwen App A/B test involving millions of users, the full framework increased recommendation CTR by 32.70%, image take-away rate by 16.32%, and average conversation turns per user by 39.90%, with all lifts statistically significant. Gemini 3.1 Pro independently audited visual consistency, separating evaluation from the Qwen-based training verifier.
Original abstract
Conversational assistants increasingly recommend follow-up edits to help users continue a task. Existing systems primarily target text-only interactions, leaving image-creation conversations underexplored. In image-creation tasks, useful follow-up edit suggestions must reflect user preferences, offer diverse directions, and remain executable on the current image. We collected 100,000 real multi-turn image-creation conversation samples from Qwen App and found that 80.1% are image-dependent, underscoring the need for multimodal recommendation. We address this setting with a three-stage framework. In Stage 1, we use real online data to build a human-reviewed table of appropriate follow-up editing intents, then create SFT targets and fine-tune a multimodal policy. In Stage 2, to align rule-guided SFT suggestions with actual user choices, we use user click feedback to optimize the policy through multi-objective reinforcement learning. In Stage 3, to reduce visual inconsistencies between suggested edits and the current image, we introduce a visual verifier as additional training supervision. Extensive experiments demonstrate that our framework significantly outperforms baselines on both automatic and human evaluations. In a live user-randomized A/B test with millions of users, our final framework reduces visual inconsistency from 3.7% to 0.9%. Furthermore, it significantly improves recommendation CTR by 32.70%, image take-away rate by 16.32%, and average conversation turns per user by 39.90% (all p<0.05).
Read the original paperMore in Multimodal AI
Browse all 61 papers →Omni-Embed-Mini: Binding Modalities Without Forgetting via Dense Distillation
Mohammed Irfan Kurpath, Jaseel Muhammad Kaithakkodan, Sahal Shaji Mullappilly, Ivan Laptev, Hisham Cholakkal
A compact embedding model unifies text, speech, audio, images, video, and documents in one search space without sacrificing the original text capabilities.
Qwen3.8-Omni: Towards Native Omni-Modal Agents
Qwen Team
Qwen3.8-Omni-Flash combines native text, audio, and video reasoning with long-context agentic planning and open-source tools for practical multimodal workflows.
YuE2: Unifying Symbolic and Audio Music Generation at Frontier Quality
Ruibin Yuan, Jiahao Pan, Junyan Jiang, Zhiyue Wu, Ziya Zhou, Jiankai Sun, Yizhi Li, Ge Zhang, Yicheng Gu, Zeyue Tian, Junyu Dai, Hanfeng Lin, Kai Li, Shangda Wu, Xuanjie Liu, Jiaming Wang, Zihan Liu, Yue Wang, Yinghao Ma, Hanzhi Yin, Kangrui Chen, Xinyue Zhang, Ziyang Ma, Mengqi Liao, Hejia Zhao, Guowei Huang, Chao Yan, Lei Ke, Jianwei Yu, Bei Liu, Joe Guo, Liumeng Xue, Gus Xia, Wei Xue, Yike Guo
YuE2 turns readable musical scores into high-quality full songs, combining controllable composition with end-to-end audio generation.