NTH

What to Edit Next: Visually Aligned Image-Editing Follow-Up Suggestions in Conversational Systems

AuthorsZhijing Zhang, Jinpeng Yu, Xin Song, Bingnan Li, Chuyue Li, Changhui Du, Xiaolin Fang, Jiaming Liu, Ruihua Huang

August 17, 2026 2 min read
Watch on YouTube
The one-line take

A multimodal assistant learns to suggest visually consistent next edits, substantially improving user engagement in large-scale image-creation conversations.

Key results

100000
Audited conversations

Real multi-turn image-creation interactions used to assess visual dependence.

80.1%
Image-dependent follow-ups

Share of follow-up editing queries requiring grounding in the latest image.

0.9%
Visual inconsistency reduction

Final inconsistency rate after adding the source-target verifier, down from 3.7% after click-based RL.

32.70%
Recommendation CTR lift

Relative improvement over the prompt-engineered policy in the live A/B test.

16.32%
Image take-away lift

Relative increase in users keeping the edited image.

39.90%
Conversation-turn lift

Relative increase in average conversation turns per user.

What the paper found

This Alibaba Qwen App study addresses a gap in conversational image creation: follow-up edit suggestions must reflect user preferences while remaining executable on the latest image. An audit of 100,000 real multi-turn interactions found that 80.1% of follow-up edits depend on visual context. The proposed three-stage pipeline first uses a human-reviewed catalog of 61 editing intents, validated Gemini 3 Flash candidates, and supervised fine-tuning of a Qwen3-VL-8B policy with rank-4 LoRA. Next, position-aware click pairs train an 8B vision-language reward model with the Bradley–Terry objective, and multi-objective GRPO optimizes click preference, formatting, perplexity, length, and diversity. Finally, a Qwen3-VL-30B-A3B image-first verifier decomposes each suggestion into required sources and target states, adding visual grounding as a sixth reinforcement-learning reward without increasing serving latency. Click optimization alone worsened visual inconsistency from 3.7% to 0.9% after the verifier was added, while expert-rated quality and slate diversity improved. In a 14-day randomized Qwen App A/B test involving millions of users, the full framework increased recommendation CTR by 32.70%, image take-away rate by 16.32%, and average conversation turns per user by 39.90%, with all lifts statistically significant. Gemini 3.1 Pro independently audited visual consistency, separating evaluation from the Qwen-based training verifier.

Original abstract

Conversational assistants increasingly recommend follow-up edits to help users continue a task. Existing systems primarily target text-only interactions, leaving image-creation conversations underexplored. In image-creation tasks, useful follow-up edit suggestions must reflect user preferences, offer diverse directions, and remain executable on the current image. We collected 100,000 real multi-turn image-creation conversation samples from Qwen App and found that 80.1% are image-dependent, underscoring the need for multimodal recommendation. We address this setting with a three-stage framework. In Stage 1, we use real online data to build a human-reviewed table of appropriate follow-up editing intents, then create SFT targets and fine-tune a multimodal policy. In Stage 2, to align rule-guided SFT suggestions with actual user choices, we use user click feedback to optimize the policy through multi-objective reinforcement learning. In Stage 3, to reduce visual inconsistencies between suggested edits and the current image, we introduce a visual verifier as additional training supervision. Extensive experiments demonstrate that our framework significantly outperforms baselines on both automatic and human evaluations. In a live user-randomized A/B test with millions of users, our final framework reduces visual inconsistency from 3.7% to 0.9%. Furthermore, it significantly improves recommendation CTR by 32.70%, image take-away rate by 16.32%, and average conversation turns per user by 39.90% (all p<0.05).

Read the original paper

More in Multimodal AI

Browse all 61 papers →
02Multimodal

Qwen3.8-Omni: Towards Native Omni-Modal Agents

Qwen Team

Qwen3.8-Omni-Flash combines native text, audio, and video reasoning with long-context agentic planning and open-source tools for practical multimodal workflows.

Read analysis
03Multimodal

YuE2: Unifying Symbolic and Audio Music Generation at Frontier Quality

Ruibin Yuan, Jiahao Pan, Junyan Jiang, Zhiyue Wu, Ziya Zhou, Jiankai Sun, Yizhi Li, Ge Zhang, Yicheng Gu, Zeyue Tian, Junyu Dai, Hanfeng Lin, Kai Li, Shangda Wu, Xuanjie Liu, Jiaming Wang, Zihan Liu, Yue Wang, Yinghao Ma, Hanzhi Yin, Kangrui Chen, Xinyue Zhang, Ziyang Ma, Mengqi Liao, Hejia Zhao, Guowei Huang, Chao Yan, Lei Ke, Jianwei Yu, Bei Liu, Joe Guo, Liumeng Xue, Gus Xia, Wei Xue, Yike Guo

YuE2 turns readable musical scores into high-quality full songs, combining controllable composition with end-to-end audio generation.

Read analysis