What to Edit Next: Visually Aligned Image-Editing Follow-Up Suggestions in Conversational Systems
summary
The gist
This paper introduces a three-stage framework for multimodal follow-up edit recommendation in image-creation conversations, addressing the need for suggestions that align with user preferences while
In short
The episode discusses a paper titled "What to Edit Next: Visually Aligned Image-Editing Follow-Up Suggestions in Conversational Systems." The hosts detail a three-stage framework used by researchers to improve image editing suggestions by combining supervised fine-tuning, reinforcement learning based on user clicks, and a visual verifier. The system achieved significant improvements in click-through rate and reduced visual inconsistency.
Key concepts
- Visual Consistency
- This refers to the problem where AI suggestions do not make sense relative to the current image. The paper addresses this by building a visual verifier that checks if suggestions rely on absent sources or request states that are already satisfied, aiming to reduce inconsistencies in image editing recommendations.
- Three-Stage Framework
- This is the system's training pipeline. Stage one uses supervised fine-tuning with human-reviewed intents. Stage two uses reinforcement learning from user clicks to align suggestions with user behavior. Stage three introduces a visual verifier to ensure suggestions are visually grounded and make sense.
- Source-Target Asymmetry
- This concept highlights that an edit can introduce a new target state that does not exist in the image, unlike simple corrections. The visual verifier is designed to understand this distinction, treating edits that require existing elements differently from those that introduce new desired states.
Terminology used across episodes
This episode discusses
- What to Edit Next: Visually Aligned Image-Editing Follow-Up Suggestions in Conversational Systems · Paper Radio
- Qwen3-VL Technical Report
- OnePred: Next-Query Prediction via Recursive Intent Memory in Multi-Turn Conversations
- Post-training Large Language Models for Diverse High-Quality Responses
- MO-GRPO: Mitigating Reward Hacking of Group Relative Policy Optimization on Multi-Objective Problems
- DVAO: Dynamic Variance-adaptive Advantage Optimization for Multi-reward Reinforcement Learning
- GDPO: Group reward-Decoupled Normalization Policy Optimization for Multi-reward RL Optimization
- Uncovering Cross-Objective Interference in Multi-Objective Alignment · Paper Radio
- Learning to Optimize Multi-Objective Alignment Through Dynamic Reward Weighting
- From Prompting to Alignment: A Generative Framework for Query Recommendation
- DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models
- Empowering Visual Creativity: A Vision-Language Assistant to Image Editing Recommendations
- Qwen3 Embedding: Advancing Text Embedding and Reranking Through Foundation Models
The paper
What to Edit Next: Visually Aligned Image-Editing Follow-Up Suggestions in Conversational Systems · Read on arXiv
Zhijing Zhang, Jinpeng Yu, Xin Song, Bingnan Li, Chuyue Li, Changhui Du, Xiaolin Fang, Jiaming Liu, Ruihua Huang
Qwen Business Unit of Alibaba · Southeast University
Conversational assistants increasingly recommend follow-up edits to help users continue a task. Existing systems primarily target text-only interactions, leaving image-creation conversations underexplored. In image-creation tasks, useful follow-up edit suggestions must reflect user preferences, offer diverse directions, and remain executable on the current image. We collected 100,000 real multi-turn image-creation conversation samples from Qwen App and found that 80.1% are image-dependent, underscoring the need for multimodal recommendation. We address this setting with a three-stage framework. In Stage 1, we use real online data to build a human-reviewed table of appropriate follow-up editing intents, then create SFT targets and fine-tune a multimodal policy. In Stage 2, to align rule-guided SFT suggestions with actual user choices, we use user click feedback to optimize the policy through multi-objective reinforcement learning. In Stage 3, to reduce visual inconsistencies between suggested edits and the current image, we introduce a visual verifier as additional training supervision. Extensive experiments demonstrate that our framework significantly outperforms baselines on both automatic and human evaluations. In a live user-randomized A/B test with millions of users, our final framework reduces visual inconsistency from 3.7% to 0.9%. Furthermore, it significantly improves recommendation CTR by 32.70%, image take-away rate by 16.32%, and average conversation turns per user by 39.90% (all p<0.05).
Transcript
Introduction to the show: ident: AI Radio. Generated commentary on the latest Artificial Intelligence papers.
Tom: Next we'll be talking about the paper "What to Edit Next: Visually Aligned Image-Editing Follow-Up Suggestions in Conversational Systems".
Jane: The paper was written by Zhijing Zhang, Jinpeng Yu, Xin Song, Bingnan Li, Chuyue Li et al. from Qwen Business Unit of Alibaba and Southeast University.
Tom: Stay tuned as we take you through the paper and discuss its implications.
Title: Tom: Welcome back to the show, everyone. Today we’re digging into a paper that’s got a mouthful of a title: “What to Edit Next: Visually Aligned Image-Editing Follow-Up Suggestions in Conversational Systems.” Jane, when you first read that title, what jumped out at you?
Jane: Oh, Tom, the phrase “What to Edit Next” is the whole story right there. It’s not about generating an image from scratch. It’s about that moment after you’ve made an edit, and the system has to figure out what you might want to do next. Like, you just turned your photo into a watercolor painting—what’s the natural follow-up? Maybe add a border, maybe change the lighting, maybe zoom in on a face. The system has to guess which of those you’d actually pick.
Tom: And that’s harder than it sounds, right? Because the suggestion has to be grounded in what’s actually in the image. You can’t suggest “remove the hat” if there’s no hat. And you can’t suggest “make the sky bluer” if the sky’s already perfectly blue. The paper calls that visual consistency, and it turns out to be a real problem.
Jane: Exactly. And the authors are from Alibaba’s Qwen team and Southeast University. They looked at a hundred thousand real conversations from the Qwen App, and they found that over eighty percent of follow-up edit queries depend on the image itself, not just the text of the conversation. So if you ignore the pixels, you’re guessing blind.
Tom: That’s a huge number. It means the old text-only approach to suggestions just doesn’t cut it for image editing. You need a system that actually looks at the picture before it opens its mouth.
Jane: And that’s what they built. A three-stage framework that starts with supervised fine-tuning, then adds reinforcement learning from real user clicks, and finally brings in a visual verifier to make sure the suggestions actually make sense for the current image. We’re going to unpack each of those stages as we go.
Tom: Can’t wait. So the title tells us the problem, but the method is where the real meat is. Stick around, because we’re about to see how they trained this thing.
Summary: Jane: So, Tom, let’s talk about what this paper actually accomplishes. The summary is pretty dense, but the core idea is that they built a system that recommends follow-up edits in a conversational image-editing app, and they made it work by combining three different kinds of supervision.
Tom: Right, and I love that they start with a real-world observation. They audited a hundred thousand adjacent user-turn pairs from the Qwen App, and they found that eighty point one percent of follow-up editing queries are image-dependent. That means you can’t figure out what the user wants just from the text. You have to see the picture.
Jane: Exactly. So they built a three-stage pipeline. Stage one is supervised fine-tuning. They took real user contexts—the latest image, the current query, and the editing intent—and they paired those with a human-reviewed table of appropriate follow-up intents. Then they used a vision-language teacher model to generate candidate suggestions, validated them, and formed six-suggestion training targets.
Tom: And that gives the model a solid foundation. It learns the structure of a good follow-up suggestion. But here’s the catch: just because a suggestion follows the rules doesn’t mean users will click it. So stage two brings in real user behavior.
Jane: Right. They collected click data from the app and trained a reward model that predicts whether a user would pick a given suggestion. Then they used reinforcement learning—specifically GRPO—to optimize the policy toward those clicks. But they didn’t just optimize clicks. They also added rewards for format validity, staying close to the SFT policy, keeping suggestions at a reasonable length, and ensuring diversity within the slate.
Tom: And that’s where things get interesting, because optimizing for clicks alone actually made the visual consistency worse. The paper shows that stage two increased the visual inconsistency rate from three point zero percent to three point seven percent. Users were clicking on appealing suggestions that sometimes didn’t make sense for the image.
Jane: Which is why they added stage three. They built a visual verifier that checks whether each suggestion relies on an absent source or requests a target state that’s already satisfied. That grounding signal became a sixth reward in the reinforcement learning, and it brought the inconsistency rate down to zero point nine percent.
Tom: And the online results are pretty dramatic. In a fourteen-day A/B test with millions of users, they saw a thirty-two point seven percent improvement in click-through rate, a sixteen point three percent improvement in image take-away rate, and a thirty-nine point nine percent improvement in average conversation turns per user. All statistically significant.
Jane: So the summary is really about the journey from rule-based suggestions to behaviorally aligned, visually verified suggestions. And the numbers show it works.
Tom: Now let’s dig into the improvements they propose. Because the three-stage framework is clever, but the details matter.
Improvements: Tom: So Jane, the improvements here are really about the architecture of the training pipeline. Let’s break down what they actually changed compared to a naive approach.
Jane: Sure. The first improvement is in stage one. They didn’t just randomly generate suggestions. They built a human-reviewed table of sixty-one editing intents and a mapping from each current intent to appropriate follow-up intents. So if the user just asked to change the background, the system knows the allowed next steps might be adjusting lighting, adding objects, or changing the style—but not, say, removing a person who isn’t there.
Tom: That’s a smart constraint. It defines the action space before the model even starts generating. But the second improvement is about learning from user behavior without falling into the trap of position bias.
Jane: Right. In the click data, a suggestion that appears higher on the screen is more likely to be clicked just because it’s more visible. So they only compared a clicked suggestion with unclicked suggestions that were displayed above it. That way, the user was more likely to have actually seen both options before making a choice.
Tom: And that simple rule improved the reward model’s accuracy from zero point six one nine to zero point six nine zero on held-out pairs. It’s a clean way to reduce bias without needing complex propensity models.
Jane: The third improvement is the visual verifier in stage three. This is the part I find most clever. They split each suggestion into required sources and a target state. A source is something that must already exist in the image, like a hat if you’re suggesting “remove the hat.” A target is the desired state after editing, like “the hat is red.”
Tom: And the verifier checks both. It looks at the image first, before reading the suggestion, to avoid being influenced by the wording. Then it checks whether each source actually exists and whether the target state is already satisfied. If the hat is already red, suggesting “make the hat red” is redundant.
Jane: Exactly. And they found that this structured approach dramatically outperforms a single-pass baseline. It recalls seventy-eight point seven percent of visual inconsistencies versus forty-seven point five percent, and its false rejection rate is just zero point six percent versus twenty-two point two percent. That last number is huge because false rejections would penalize valid creative edits and push the model toward boring, generic suggestions.
Tom: So the improvements are really about adding the right supervision at the right time. Rules first, then behavior, then visual grounding. Each stage fixes a problem the previous stage couldn’t see.
Jane: And that’s the story of the paper. Now let’s look at the first page in detail, because that’s where they set up the problem and the key statistics.
First Page: Jane: Alright, Tom, let’s go back to the very beginning of the paper. The first page sets up the problem with that striking statistic: eighty point one percent of follow-up editing queries are image-dependent. Only nineteen point nine percent can be handled from text alone.
Tom: And they’re careful to define what that means. An image-dependent query introduces concrete content that isn’t in the previous text. Like, “make the dog bigger” when the dog was never mentioned—you have to see the image to know there’s a dog.
Jane: Right. And they contrast their work with existing systems. Query suggestion systems work on text. Instruction-guided image editing executes user-specified edits. Image-editing recommendation generates diverse candidate instructions. But none of them jointly learns a behaviorally aligned follow-up slate and verifies it against the latest image.
Tom: So they’re filling a gap. And the first page also introduces the three-stage framework in a nutshell. Stage one uses real online data to build a human-reviewed table of follow-up intents and fine-tunes a multimodal policy. Stage two uses click feedback to optimize through multi-objective reinforcement learning. Stage three introduces the visual verifier.
Jane: And there’s a really interesting detail in the introduction. They mention that the click reward model receives the image as input, but its supervision comes solely from user choices. Those choices reflect appeal, not executability. So a user might click on “make the background tropical” even if the background is already tropical, because it sounds nice.
Tom: That’s the key insight. Clicks tell you what users like, but not whether the edit makes sense. That’s why stage three is necessary. And the numbers back it up: stage two raised visual inconsistency from three point zero percent to three point seven percent, and stage three brought it down to zero point nine percent.
Jane: The first page also previews the online results. A thirty-two point seven percent lift in CTR, sixteen point three percent in take-away rate, and thirty-nine point nine percent in conversation turns per user. Those are big numbers for a production system.
Tom: And they emphasize that the deployed system retains a single 8B policy without additional serving latency. The verifier is only used during training. That’s a practical detail that engineers will appreciate.
Jane: Absolutely. So the first page really sets the stage. It defines the problem, explains why existing work falls short, and previews the solution and the results. Now let’s bring in the rest of the team to get their take.
Lu: I’m really struck by the source–target asymmetry. Most visual grounding work checks whether a statement about an image is true. But here, an edit can introduce a new target that isn’t in the image yet, and that’s fine. The verifier has to understand that distinction, and the fact that they built it explicitly is a real contribution.
Meng: And from an engineering standpoint, the fact that they kept the verifier out of the serving path is huge. You get the benefit of visual consistency without paying the latency cost. That’s the kind of design decision that makes a paper actually deployable.
Lalam: I think the cultural impact is interesting too. When image-editing assistants can suggest meaningful next steps, they become more like creative partners than tools. Users stay engaged longer, they explore more directions, and the technology feels more collaborative.
Tom: Great points from everyone. Let’s wrap this up in the conclusion.
Conclusion: Tom: Alright, let’s bring it home. We’ve been talking about “What to Edit Next: Visually Aligned Image-Editing Follow-Up Suggestions in Conversational Systems,” and honestly, this paper is a masterclass in combining supervision sources.
Jane: It really is. They started with the observation that eighty point one percent of follow-up edits depend on the image, not just the text. Then they built a three-stage pipeline: rules and human-reviewed intents for SFT, click-based reinforcement learning for behavioral alignment, and a visual verifier for grounding.
Tom: And the results speak for themselves. Visual inconsistency dropped from three point seven percent to zero point nine percent. CTR went up thirty-two point seven percent. Take-away rate up sixteen point three percent. Conversation turns per user up thirty-nine point nine percent. All significant, all in a live production system with millions of users.
Jane: What I love most is that they didn’t just optimize clicks. They recognized that clicks measure appeal, not executability. So they added a separate signal for visual consistency. That’s the kind of nuance that separates a good system from a great one.
Lu: And the source–target verifier is a genuinely novel idea. It treats “remove the hat” and “add a hat” differently, because one requires the hat to exist and the other doesn’t. That asymmetry is easy to miss, and they built an entire verification pipeline around it.
Meng: From an engineering perspective, the fact that the verifier only runs during training is a big win. You get the quality improvement without any added latency at serving time. That’s how you ship something like this.
Lalam: And the cultural angle is real. When an assistant can suggest meaningful next steps, it becomes a creative collaborator. Users stay engaged, they explore more, and the technology feels more human. That’s a meaningful shift in how we interact with AI.
Tom: Well said, everyone. So what’s next for this line of research? I’d love to see the verifier handle more complex edits, like multi-step transformations or edits that depend on spatial relationships.
Jane: And I’d love to see how this generalizes to other domains—video editing, document design, maybe even three dee modeling. The core idea of aligning behavioral preference with visual validity is pretty universal.
Tom: Agreed. For now, let’s give a round of applause to the authors for a fantastic paper. We’ll be back next time with another exciting piece of research. Until then, keep creating, keep editing, and keep asking “what next?”
Jane: See you all on the next episode!
More episodes
- 2610.10857-Self-Supervised Keyframe Discovery for Horizon-Invariant Behavior Cloning
- 2610.10768-Strategic Investment Decision Making for Value Creation in Energy Transition: A Reinforcement Learning Approach
- 2610.10858-RFChipAgent: Multi-Agentic AI Flow for Analog/RF Chip Design
- 2610.10613-Temporal transformer CAN encoder with federated lightweight heads for anomaly detection
- 2610.10616-When Routing Reveals Membership: Privacy Leakage from MoE Router Telemetry
- 2610.10655-Nullify: Null-Space Activation Steering for Training-Free LLM Unlearning
- 2610.11031-Language Modeling is Monotone Compression
- 2610.01253-Context-Aware Error Mitigation Orchestration for Hybrid Quantum Reinforcement Learning on NISQ Systems
- 2604.24201-CMGL: Confidence-guided Multi-omics Graph Learning for Cancer Subtype Classification
- 2609.34069-Towards Certificate-Driven Software Porting: A Self-Improving Agentic Harness for Scientific Program Optimization