Rethinking Multi-Image Re-Representation in Multi-Image Understanding

arXiv:2609.39363 · cs.CV, cs.AI · Submitted 2026-09-30 · Read on arXiv

Listen

Radio episode about this paper

Transcript

Introduction to the show: ident: AI Radio. Generated commentary on the latest Artificial Intelligence papers.

Tom: Today's paper: "Rethinking Multi-Image Re-Representation in Multi-Image Understanding".

Jane: Multi-image understanding requires MLLMs not only to recognise the content of individual images, but also to organise visual evidence distributed across them.

Tom: First, who's behind it and why it matters.

Paper summary: Tom: Moving on to a bit more detail about what this paper is actually proposing, the "Rethinking Multi-Image Re-Representation in Multi-Image Understanding." Essentially, the thesis is that multi-image understanding requires models not only to recognize individual images but also to organize that visual evidence scattered across them.

Jane: They claim that they study this problem through multi-image re-representation, which they define as constructing a variablelength intermediate record from the input specifically to organize all that visual evidence for answering a question.

Lu: The paper introduces two ways of doing this: textual re-representation uses a text sequence as its record, while visual re-representation constructs these intermediates through image operations in an online setting, resulting in an interaction trace.

Meng: So the goal is to define a rerepresenter function p theta(z x) that creates this intermediate record z, and then a solver p phi(y x, z) uses that record along with the original input to produce an answer distribution.

Lalam: This whole structure covers both textual and visual forms of re-representation, which is a big step because it formalizes how we can systematically organize information for reasoning tasks.

Tom: They then introduce Mosaic as a general-purpose multiimage visual harness designed to let an MLLM actively construct these visual intermediates using ten composable image operations.

Jane: These operations include things like cropping, geometric transformations, and image composition, giving the model direct control over how it manipulates the visual evidence during its reasoning process.

Lu: They compare five different re-representation settings—from no explicit re-representation to various textual and visual approaches—across existing benchmarks and on their new grounding-focused benchmark called MosaicBench.

Meng: The key finding they highlight is that the relative benefits of textual versus visual re-representation are strongly task-dependent, meaning one isn't universally superior for every multi-image question.

Lalam: This comparison using MosaicBench is really important because it provides a focused evaluation environment specifically designed to test where each re-representation style shines.

Tom: So, the main takeaway from this overview is that visual re-representation shows particular effectiveness for tasks that demand precise visual evidence or fine-grained relations across images.

Jane: That means if your task is about exact spatial relationships or subtle differences, focusing on building these visual intermediates could yield better results than just relying on text.

Conclusion: Tom: So, wrapping up our discussion on "Rethinking Multi-Image Re-Representation in Multi-Image Understanding," the authors are really emphasizing the importance of understanding that multi-image reasoning isn't a single approach; it depends heavily on what kind of evidence you need.

Jane: They’re pointing out that visual re-representation is particularly useful for tasks where precision and fine-grained relations are critical, like hypothesis testing or comparing orientations.

Lu: The authors suggest that the way we structure the evidence through these intermediate records fundamentally changes how the model approaches complex multi-image questions.

Meng: It means future development needs to be flexible systems that can dynamically select between textual description and active visual manipulation based on the specific demands of a reasoning task at hand.

Lalam: If we can teach our AI to choose which visual assets to operate on and which operations to apply, it could really enhance the culture by enabling us to build tools that handle intricate visual tasks with greater accuracy.

Tom: Ultimately, this work provides a solid framework for how we can think about organizing scattered evidence into a structured record for reasoning, whether that record is textual or visual.

Jane: It gives us the language to better understand why certain methods succeed on specific benchmarks while others fall short when facing fine-grained visual challenges.

Lu: The implication is that we need to move away from monolithic solutions and toward more adaptable systems that can choose their internal representation strategy intelligently during reasoning.

Meng: This suggests a future where AI agents are not just passive observers but active manipulators of the visual scene when necessary for high-stakes visual tasks.

Lalam: And if we can achieve that level of flexible, task-aware structuring, it opens up possibilities for building more sophisticated AI systems capable of handling nuanced visual contexts across multiple inputs.

Gengyuan Zhang, Xiao Han, Xinyu Xie, Tong Liu, Volker Tresp

LMU Munich

cs.CV, cs.AI

Submitted: 2026-09-30

Updated: 2026-09-30

Code: https://github.com/gengyuanmax/Mosaic

Importance score: 79/100

The gist: Multi-image understanding requires MLLMs not only to recognise the content of individual images, but also to organise visual evidence distributed across them.

Key concepts

Multi-image Re-representation
This is the process of creating a variable-length intermediate record from input images to structure the visual evidence for answering a question. It can be textual (a sequence of text) or visual (a trace of image operations performed online).
Mosaic
Mosaic is a general-purpose harness that lets an MLLM use ten composable image operations, including cropping, geometric transformations, and image composition. This allows the model to actively construct its own visual intermediates during reasoning.
Visual Re-representation (V-Re2)
This setting involves using Mosaic to build visual intermediates. The paper finds this method is particularly effective for tasks that require precise visual evidence or fine-grained relationships across different images, showing greater gains than textual re-representation in these specific scenarios.

Terminology

Summary

Multi-image understanding requires MLLMs not only to recognise the content of individual images, but also to organise visual evidence distributed across them. The gist: Visual re-representation is particularly effective for tasks requiring precise visual evidence, including hypothesis testing, precision comparison, and orientation-sensitive reasoning.

Defining Multi-Image Re-representation

Multi-image re-representation constructs a variablelength intermediate record from the input to organise visual evidence for answering a question. A rerepresenter is specified by the conditional distribution pθ(z x), where m ∈ Zm denotes the representation form, and a solver pϕ(y x, z) uses this record alongside the original input to give an answer distribution. This process covers both textual and visual forms of re-representation. Textual re-representation uses a text sequence z = d = (d1,..., dL) as its intermediate record, while visual re-representation constructs visual intermediates through image operations in an online setting with a multimodal interaction trace z = τ = (a1, o1,..., aT, oT).

The Mosaic Harness and Re-representation Settings

Mosaic is introduced as a general-purpose multiimage visual harness that enables an MLLM to actively construct visual intermediates with ten composable image operations. These operations include cropping, geometric transformation, and image composition. The paper compares five re-representation settings: A. No explicit re-representation (No-Re2), B. Free-form textual re-representation (T-Re2), C. Prompt-guided textual re-description (PG-Re2), D. Visual re-representation (V-Re2) using Mosaic, and E. Prefabricated visual re-representation (PV-Re2).

Evaluating Re-representation Benefits

The relative benefits of textual and visual re-representation are strongly task-dependent. Visual re-representation is particularly effective for tasks requiring precise visual evidence or finegrained relations across images, while tasks dominated by higher-level semantic content often show smaller or less consistent gains. This dependency is examined using MosaicBench, a new groundingfocused benchmark for fine-grained multi-image understanding. For example, on M4Bench’s Detailed Difference task, Qwen3-VL-8B improves from 7.3% under No-Re2 to 46.1% under T-Re2 and 59.9% under V-Re2 for Qwen3-VL-8B.

Learning to Construct Visual Representations

The second research question is how an agent learns to construct useful visual re-representations, as this requires the agent to choose which image assets to operate on, which operations to apply, and how to proceed from the resulting views. To address this, MosaicAgent-8B is trained using reinforcement learning with only accuracy and format rewards, without demonstration trajectories or rewards for specific tool-use. The results show that this is sufficient for the agent to learn multi-step compositions of visual operations and exhibit diverse problem-solving patterns unpromptedly.

Impact of Training on Reasoning Dynamics

Training changes the mixture of problem-solving modes. While absolute counts increase, their share decreases from 41.9% to 35.5%, while self-correction and post-stabilisation continuation become substantially more frequent. The learned policy exhibits a different mixture of progression, revision, and continued processing rather than converging to a single strategy. Furthermore, the harness-enabled visual re-representation provides a substantial additional benefit beyond CoT RL on the same training data. Post-training trajectories contain more tool steps on average (5.41 vs. 3.16).

MosaicBench and Fine-Grained Evaluation

MosaicBench was introduced to evaluate re-representation on tasks requiring precise visual evidence and relations, covering 32 task types grouped by core visual challenge: resolution, orientation, precision comparison, hypothesis testing, context interference, and spatial reference. The benchmark is constructed from 12 public datasets using annotation augmentation to create task-specific questions and ground-truth relations. This allows for a focused evaluation of tasks such as hypothesis testing, where MosaicAgent-8B reaches 62.9% compared to 31.7% for the strongest open-weight baseline on MosaicBench.

Tool Ablation and Failure Profiles

Ablation studies show that the full visual toolset increases accuracy by 9.4 percentage points on MosaicBench and 3.0 points on M4Bench, indicating that operations beyond cropping is therefore most pronounced on MosaicBench. Failure profiles show that training reshapes failures: fewer trajectories destroy an answer the model already had, while more find an answer they cannot hold, and more settle into a wrong answer several steps before stopping. This suggests that successful visual reasoning can involve revising intermediate rather than only accumulating evidence.

Improvements for AI systems

Based on the provided scientific paper, here are specific improvements that can be made to AI systems by integrating the concepts from Mosaic and its surrounding research:


)1. Enhanced Multi-Image Evidence Organization (Multi-Image Re-Representation):

The system should move beyond simple interleaved text/image context. It must be equipped with a mechanism to actively construct intermediate visual representations.

The improved system can perform visual re-representation by using a harness (like Mosaic) to actively manipulate source images—cropping, rotating, resizing, compositing, and applying affine/homography transformations—to create new image views that serve as explicit reasoning steps.

)2. Task-Dependent Re-Representation Strategy Selection:

The system should not use a single re-representation method universally. It needs a decision-making layer to choose the most appropriate evidence organization strategy based on the nature of the query.

The system can dynamically switch between:

  • Textual Re-representation (for high-level semantic content).
  • Visual Re-representation (specifically for tasks requiring precise visual evidence, such as orientation-sensitive reasoning, hypothesis testing, and precision comparison).

)3. Agentic Visual Tool Use with Reinforcement Learning:

The system should be trained not just to follow instructions but to autonomously decide which tool to use, when to use it, and how to chain operations together.

The improved agent can be trained using reinforcement learning (with simple rewards like accuracy and format compliance) without explicit demonstration trajectories or reward sequences. This enables it to learn complex, multi-step compositions of visual operations unpromptedly.

)4. Grounding-Focused Evaluation and Fine-Grained Understanding:

The AI system's capability should be rigorously tested on benchmarks designed for fine-grained, precise visual grounding rather than just general VQA.

The system can achieve superior performance on tasks requiring:

  • Hypothesis testing (competing transformations or assemblies).
  • Precision comparison (subtle differences between close views).
  • Spatial reference and orientation-sensitive reasoning.

)5. Robustness Against Context Interference and Ambiguity:

The system must be specifically trained to distinguish relevant visual evidence from confounding factors like lighting changes, background structures, or illusion patterns.

The improved system can accurately locate defects by explicitly disregarding photometric variation (illumination vs. defect) and correctly identify target properties in the presence of simultaneous contrast confusion or illusion patches.

)6. Improved Answer Retention and Stopping Policies:

The learned policy should be refined to ensure it doesn't prematurely stop reasoning when a correct answer is found, but rather continues processing to verify or explore further evidence.

The agent can exhibit diverse problem-solving patterns, including self-correction (revising intermediate representations) and post-stabilization continuation (verifying an answer by inspecting additional evidence), leading to more robust final answers.

)7. Scalable Tool Integration via a General Harness:

The underlying architecture should incorporate a general-purpose visual harness that allows for the dynamic construction, retention, and reuse of visual intermediates across various reasoning steps.

The improved system can maintain a persistent workspace where derived views (e.g., cropped regions or collages) are stored as reusable assets for subsequent operations, allowing complex reasoning chains to build incrementally on previously derived evidence.

Sources

Related papers