COMEX: A Composition-Grounded Benchmark and Learning Framework for Explainable Aesthetic Image Cropping
summary
In short
The episode discusses COMEX, a paper presenting a benchmark and learning framework for explainable aesthetic image cropping. The authors propose outputting a crop box, composition category, and explanation grounded in photographic principles. They use an IO-reversal trick to create a large dataset and employ an SFT plus GRPO training method that significantly improves performance on real photos.
Key concepts
- Aesthetic Image Cropping
- This is the task of automatically finding the best rectangular cut from a photo to make it look professional and pleasing, similar to how a photographer trims an image. The goal is to find the crop that makes the image look most professional.
- Explainable Cropping
- This means a cropping system must not only provide the crop box but also explain its reasoning. The explanation should be tied to actual photographic composition principles, such as symmetry or leading lines, moving beyond vague statements like 'this crop focuses on the subject'.
- COMEX Dataset
- This is a new dataset created using an IO-reversal trick. They started with existing images labeled by composition experts and used a generative model to expand them outward, creating a large canvas with the original photo embedded inside it, providing ground truth for training.
- SFT plus GRPO Framework
- This is the two-stage training method. Stage one is Supervised Fine-Tuning (SFT) to learn the basic pattern. Stage two uses Group Relative Policy Optimization (GRPO), a reinforcement learning technique, which rewards the model based on crop accuracy, explanation quality, and composition consistency.
Terminology used across episodes
This episode discusses
- COMEX: A Composition-Grounded Benchmark and Learning Framework for Explainable Aesthetic Image Cropping · Paper Radio
- Qwen3-VL Technical Report
- Qwen2.5-VL Technical Report
- Gemini 2.5: Pushing the Frontier with Advanced Reasoning, Multimodality, Long Context, and Next Generation Agentic Capabilities
- Venus: Benchmarking and Empowering Multimodal Large Language Models for Aesthetic Guidance and Cropping
- GPT-4o System Card
- FLUX.1 Kontext: Flow Matching for In-Context Image Generation and Editing in Latent Space
- Q-Insight: Understanding Image Quality via Visual Reinforcement Learning
- Kimi-VL Technical Report
- InternVL3.5: Advancing Open-Source Multimodal Models in Versatility, Reasoning, and Efficiency
- Qwen3 Technical Report
The paper
COMEX: A Composition-Grounded Benchmark and Learning Framework for Explainable Aesthetic Image Cropping · Read on arXiv
Rui Yang, Wei Zhou, Dingyong Gou, Xiaohui Cui, Cong Li, Yinyin Gong, Yipo Huang, Jiliang Zhao
State Key Laboratory of Mobile Network and Mobile Multimedia Technology, ZTE Corporation · School of Data Science and Institute of Artificial Intelligence, Chang'an University
Explainable aesthetic image cropping requires not only localizing a visually pleasing crop but also explaining why it is preferred. Existing crop-and-explain methods largely treat explanation as post-hoc text generation and overlook composition, a key aesthetic factor that links crop decisions with interpretable reasoning. In this paper, we reformulate explainable aesthetic image cropping as a structured crop-composition-explanation problem. To support this setting, we introduce COMEX, a new benchmark built through image expansion and an IO-reversal pipeline. COMEX contains 33,161 quadruples, each consisting of an expanded image, a crop box, a composition category, and a composition-grounded explanation, enabling joint learning of crop localization, composition understanding, and explanation generation. We further propose a two-stage SFT+GRPO framework, where supervised fine-tuning establishes the structured output protocol and basic cropping ability, and GRPO further improves crop quality, composition prediction, and explanation faithfulness. We benchmark 15 large vision-language models and existing cropping methods on COMEX, establishing a comprehensive testbed for composition-grounded explainable aesthetic cropping. Experiments on both COMEX and prior benchmarks demonstrate the effectiveness and transferability of our framework, with strong performance across evaluation metrics.
Transcript
Introduction to the show: ident: AI Radio. Generated commentary on the latest Artificial Intelligence papers.
Tom: Next we'll be talking about the paper "COMEX: A Composition-Grounded Benchmark and Learning Framework for Explainable Aesthetic Image Cropping".
Jane: The paper was written by Rui Yang, Wei Zhou, Dingyong Gou, Xiaohui Cui, Cong Li et al. from State Key Laboratory of Mobile Network and Mobile Multimedia Technology, ZTE Corporation and School of Data Science and Institute of Artificial Intelligence, Chang'an University.
Tom: Stay tuned as we take you through the paper and discuss its implications.
Title: Tom: Welcome back to the show, everybody. Today we're looking at a paper that's got a mouthful of a title — "COMEX: A Composition-Grounded Benchmark and Learning Framework for Explainable Aesthetic Image Cropping." Jane, I'm going to be honest, when I first read that title I had to break it into pieces.
Jane: You and me both, Tom. But once you unpack it, it's actually a really intuitive idea. So "aesthetic image cropping" is the task of automatically finding the best rectangle to cut out of a photo — the crop that makes the image look most professional and pleasing. Think of how a photographer might trim a photo to make the subject pop.
Tom: Right, and the "explainable" part is the twist. Most cropping systems just give you the box — they say "crop here" but they don't tell you why. This paper wants the system to also explain its reasoning, like "I cropped here because it follows the rule of thirds and centers the subject."
Jane: Exactly. And "composition-grounded" means the explanation is tied to actual photographic composition principles — things like symmetry, leading lines, or how the subject is arranged in the frame. That's the piece that's been missing.
Tom: So the title is basically promising three things: a new dataset, a new training method, and a way to make the whole thing explainable. And the authors are from ZTE Corporation's State Key Lab, plus a researcher from Chang'an University. I love seeing industry labs publishing this kind of work.
Jane: It makes sense though, because this has real product applications. Think about photo editing apps on your phone — if the app can tell you *why* it's suggesting a crop, that's a much better user experience than just silently moving the box around.
Tom: And that's what gets me excited. We're moving from "black box tells you what to do" to "system explains its reasoning." That's a big deal for trust and usability.
Jane: It really is. And the authors are betting that composition is the bridge — the thing that connects the crop decision to the explanation. We'll get into how they built that bridge in a moment.
Tom: Stay with us, because this paper has some numbers that are going to surprise you.
Summary: Jane: So Tom, let's get into what this paper actually does. The core claim is that previous explainable cropping methods treated the explanation as an afterthought — they'd predict the crop box, then generate text to describe it. But that text wasn't grounded in any real photographic principle.
Tom: And that's the problem, right? The explanations end up being vague. Like, "this crop focuses on the main subject" — which could describe almost any crop of almost any photo. It's technically true but completely useless.
Jane: Exactly. So the authors reformulate the whole task. Instead of crop-and-explain, they propose crop-composition-explanation. The model has to output three things: the crop box, the composition category, and then the explanation that ties them together.
Tom: And to support that, they built a new dataset called COMEX. This is where it gets clever. They didn't go out and hire photographers to take thousands of crops. Instead, they started with an existing dataset called PICD, which has images labeled with composition categories by experts.
Jane: Right, and then they did something really smart. They took each original image and treated it as the *ideal crop* — the target. Then they used a generative model to expand the image outward, creating a larger canvas with the original photo embedded inside it. So now they have a big image and a known good crop box, with a known composition label.
Tom: That's the IO-reversal trick — input-output reversal. Instead of finding the crop for an image, they start with the crop and generate the image around it. That gives them ground truth for free, essentially.
Jane: And then they used a multimodal language model to generate explanations for each crop, given the composition category. After quality control, they ended up with thirty-three thousand one hundred sixty-one quadruples — image, crop box, composition category, and explanation.
Tom: That's more than three times the size of the previous largest dataset for this task. And it's the first one that has all three annotations aligned — box, composition, and explanation.
Jane: The size matters because it lets them train models properly. And the composition annotation is the key innovation — it's the missing link that makes explanations actually meaningful.
Tom: So they built the data, but then they also had to build the training method. And that's where things get really interesting, because they're not just doing supervised learning.
Jane: Right, and I think that's the part that's going to get the engineers in our audience excited.
Improvements: Tom: So Jane, the dataset is impressive, but the training framework is where the real innovation lives. They propose a two-stage approach called SFT plus GRPO. Let's break that down for our listeners.
Jane: Sure. Stage one is supervised fine-tuning — SFT. That's the standard approach: you show the model thousands of examples and train it to imitate the correct output. In this case, the correct output is the structured triplet — box, composition, explanation.
Tom: And that gets you a decent baseline. But here's the problem — SFT has a ceiling. The model learns to copy the pattern, but it's just maximizing the likelihood of the next token. It has no idea whether the crop it's predicting is actually *good*.
Jane: Right, and the paper shows this empirically. They trained Stage I for ten epochs, then kept training for five more epochs. The improvement was marginal — IoU went from zero point seven five three two to zero point seven five four eight. Basically flat.
Tom: So they introduce Stage II — GRPO, which stands for Group Relative Policy Optimization. This is a reinforcement learning technique that came out of DeepSeek-R1. Instead of imitating reference text, the model samples multiple candidate outputs and gets rewarded based on how good they actually are.
Jane: And the key insight is that they designed rewards that directly measure what they care about. There's a geometric reward for crop accuracy — how well does the predicted box overlap with the ground truth. There's a semantic reward for explanation quality and composition consistency. And there's a format reward to keep the output structured.
Tom: And the results are striking. Same checkpoint, same amount of additional training — but with GRPO instead of continued SFT, IoU jumps from zero point seven five three two to zero point seven seven six five. That's a much bigger gain.
Jane: And it's not just the box. Composition accuracy goes up, explanation quality goes up, and the number of high-quality crops — those with IoU above zero point nine — increases substantially.
Tom: The other thing I love is that they tested whether composition information actually matters. They ran an ablation where they trained without composition supervision, and the explanations were noticeably worse. When they asked human evaluators to compare, seventy percent of general users and seventy-six percent of experts preferred the explanations from the composition-grounded model.
Jane: That's a strong signal. It means composition isn't just a nice extra — it's actually making the explanations more faithful and more useful.
Tom: And the framework works across different model sizes. They tested it on a 0 point 8B model and a 2B model, and both improved with Stage II. That's important because it means you don't need a massive model to get these benefits.
Jane: So the improvements are threefold: better crops, better composition understanding, and better explanations. All from the combination of better data and better training.
Tom: And that's the story so far. But I want to talk about what this means in practice — how well does it transfer to real photos?
First Page: Jane: So Tom, we've talked about the dataset and the training method. But there's a question that's been nagging me — the COMEX images are created with generative outpainting. The original photo is real, but the surrounding context is synthetic. Does that actually transfer to real-world photos?
Tom: That's exactly the question the authors tackle in their experiments. They took models trained on COMEX and evaluated them on FCDB, which is a benchmark of real photos with human-annotated crops. And the results are genuinely encouraging.
Jane: The zero-shot transfer — meaning no fine-tuning on FCDB at all — gets an IoU of zero point six two five one after Stage II. That's approaching some of the older supervised methods that were trained directly on FCDB. And it's way better than general-purpose vision-language models, which struggle to get above zero point four.
Tom: And here's the thing — the GRPO stage helps with transfer too. Stage I alone gets zero point five nine four six on FCDB, but Stage I plus Stage II gets zero point six two five one. So reinforcement learning on COMEX is teaching the model something general about composition, not just memorizing the synthetic data.
Jane: That's a really important finding. It suggests that the compositional knowledge learned from the outpainted images — things like the rule of thirds, symmetry, leading lines — transfers to real photographs because those principles are universal.
Tom: And they also tested the framework directly on FCDB, training from scratch. Stage I alone gets zero point seven zero zero seven IoU, which is competitive with state-of-the-art methods like CACNet. But Stage I plus Stage II gets zero point seven two two five, which actually beats the previous best MLLM-based method, InstructCrop, by a meaningful margin.
Jane: So the framework is broadly effective, not just for the COMEX setting. That's a strong validation of the approach.
Tom: And I think the first page of the paper sets this up really well. It has this great figure showing the paradigm shift — from crop-and-explain to crop-composition-explanation. And it shows qualitative examples where the old style of explanation is vague, while the new style is specific and grounded.
Jane: The example that stuck with me was the heron photo. The old explanation says something like "this crop centers the heron, trimming redundant sky and grass." The new explanation says "this vertical third-line crop places the heron along the left third line." You can see the difference — one is describing what happened, the other is explaining the *principle* behind it.
Tom: And that's the whole point. The composition category gives the explanation a structure, a vocabulary, a way to be precise.
Jane: It also makes the explanations more useful for learning. If you're a novice photographer and the system tells you "this uses the rule of thirds," you can actually learn from that. You can apply it to your own photos.
Tom: That's a great point. This isn't just about automating a task — it's about teaching aesthetic principles. And that has implications beyond cropping.
Jane: We'll get into those implications in our conclusion. But first, let's bring in the rest of the team to get their take.
Conclusion: Tom: So we've covered a lot of ground on "COMEX: A Composition-Grounded Benchmark and Learning Framework for Explainable Aesthetic Image Cropping." Let's bring in Lu, Meng, and Lalam to get their final thoughts.
Jane: Lu, you're the researcher — what excites you most about this work?
Lu: I think the most exciting implication is that this framework could generalize beyond cropping. The idea of using composition as an intermediate layer between a visual decision and a language explanation — that could apply to other aesthetic tasks. Image retouching, layout design, even video editing. If you can articulate *why* a visual choice is good, you can teach that principle to a model.
Meng: From an engineering standpoint, I'm impressed by the efficiency. They're using a 2B parameter model and getting state-of-the-art results. That means this could run on-device — on a phone, on a camera. And the GRPO stage doesn't require a separate value model, which keeps training costs down. That's practical.
Jane: And Lalam, you think about the cultural impact. What does this mean for how people interact with visual media?
Lalam: I think the biggest impact is on accessibility. Professional photography has always had this barrier — the knowledge of composition was something you had to learn through years of practice or expensive courses. This paper is a step toward making that knowledge explicit and teachable. A novice photographer could get not just a suggested crop, but an explanation of the principle behind it. That's democratizing aesthetic education.
Tom: That's a beautiful way to put it. And it connects back to what Jane said earlier about learning.
Jane: Right. And I think that's the lasting contribution of this paper — it's not just a better cropping system. It's a framework for making aesthetic decisions interpretable and teachable.
Tom: So let's wrap up. COMEX gives us the first large-scale benchmark with aligned crop, composition, and explanation annotations. The two-stage SFT plus GRPO framework shows that reinforcement learning can meaningfully improve beyond supervised fine-tuning. And the transfer results show that the compositional knowledge generalizes to real photos.
Jane: And the human evaluations confirm that composition-grounded explanations are genuinely better — more specific, more faithful, more useful.
Lu: I'd add that the ablation studies are really clean. They show that composition supervision helps, that GRPO helps, and that the semantic reward is necessary for maintaining explanation quality.
Meng: And from a deployment perspective, the fact that it works on small models makes it viable for real products.
Lalam: The cultural takeaway is that we're moving toward AI systems that can not only make aesthetic judgments but explain them in human terms. That builds trust and enables learning.
Tom: Well said, everyone. That's our discussion of COMEX. Jane, what's next on the docket?
Jane: We've got a paper on multimodal reasoning in medical imaging coming up — should be a good one.
Tom: Looking forward to it. Thanks for listening, everybody. We'll catch you on the next episode.
More episodes
- 2610.10857-Self-Supervised Keyframe Discovery for Horizon-Invariant Behavior Cloning
- 2610.10768-Strategic Investment Decision Making for Value Creation in Energy Transition: A Reinforcement Learning Approach
- 2610.10858-RFChipAgent: Multi-Agentic AI Flow for Analog/RF Chip Design
- 2610.10613-Temporal transformer CAN encoder with federated lightweight heads for anomaly detection
- 2610.10616-When Routing Reveals Membership: Privacy Leakage from MoE Router Telemetry
- 2610.10655-Nullify: Null-Space Activation Steering for Training-Free LLM Unlearning
- 2610.11031-Language Modeling is Monotone Compression
- 2610.01253-Context-Aware Error Mitigation Orchestration for Hybrid Quantum Reinforcement Learning on NISQ Systems
- 2604.24201-CMGL: Confidence-guided Multi-omics Graph Learning for Cancer Subtype Classification
- 2609.34069-Towards Certificate-Driven Software Porting: A Self-Improving Agentic Harness for Scientific Program Optimization