MiniGPT-Reverse-Designing: Predicting Image Adjustments Utilizing MiniGPT-4

arXiv:2406.00971 · cs.CV, cs.AI · Submitted 2026-08-15 · Read on arXiv

Listen

Radio episode about this paper

Transcript

Introduction to the show: ident: AI Radio. Generated commentary on the latest Artificial Intelligence papers.

Tom: Next we'll be talking about the paper "MiniGPT-Reverse-Designing: Predicting Image Adjustments Utilizing MiniGPT-4".

Jane: The paper was written by Vahid Azizi and Fatemeh Koochaki from.

Tom: Stay tuned as we take you through the paper and discuss its implications.

Title: Tom: Welcome back to the show, everyone. Today we’re digging into a paper called “MiniGPT-Reverse-Designing: Predicting Image Adjustments Utilizing MiniGPT-four.” Jane, I have to say, just the title alone got me curious—reverse designing sounds like something out of a puzzle book.

Jane: It really does, Tom. And the idea is actually pretty clever. Normally, when we edit an image, we know what we did—we brightened it, we cropped it, we changed the colors. But reverse designing flips that. You’re given the original image, the edited version, and maybe a vague text description, and you have to figure out exactly what edits were made and with what values.

Tom: So it’s like looking at a before-and-after photo and trying to guess the recipe of changes. That’s not just a simple task—that’s asking the model to understand the relationship between two images and text all at once.

Jane: Exactly. And that’s why the authors chose MiniGPT-four which is a vision-language model that can process both images and text together. They wanted to see if an off-the-shelf model like that could be extended to handle this more complex, multi-image task.

Tom: And the authors are Vahid Azizi and Fatemeh Koochaki. They’ve made the code available, which is great for anyone who wants to try it out themselves. But Jane, what’s the real-world hook here? Why should we care about reverse designing?

Jane: Think about image editing software. If you could automatically figure out what edits were made to a photo, you could build better revision control for images, or even teach people how to achieve certain effects by showing them the steps. It’s like having a chef reverse-engineer a dish just by tasting it.

Tom: That’s a tasty analogy. And I love that they’re not just using the model as-is—they’re fine-tuning it specifically for this task. That’s where the real work happens, and I’m excited to see how they did it.

Jane: Me too. Let’s get into the details of how they set up the whole thing.

Summary: Tom: So Jane, we’ve got the big picture—reverse designing is about predicting edits from before-and-after images. But how did they actually make MiniGPT-four do this? What’s the setup?

Jane: Great question. They took MiniGPT-four which already aligns a frozen vision encoder with a frozen language model using a simple linear projection layer. The key is that they only train that projection layer, keeping the rest of the model frozen. So they’re not retraining the whole thing—they’re just teaching the model how to connect the two images with the text.

Tom: That’s smart because it saves a ton of compute. But how do they feed two images into a model that was originally designed for one?

Jane: They process both images through the same vision encoder, so each image becomes a set of tokens. Then those tokens are combined with the text tokens using a template prompt, and everything goes into the language model, which is LLAMA-two in this case. The model then outputs the operations and their parameters as a sentence.

Tom: And the dataset—what are they training on? I remember they used something called I-MAD-Dense.

Jane: Right. The I-MAD dataset, which stands for Image Multi-Adjustment Dataset. The dense version has about twenty-two thousand triplets, each with a source image, an edited image, and a high-level creative edit description. The descriptions are intentionally vague, so the model has to rely on the images themselves to figure out the specific edits.

Tom: Vague descriptions sound like a challenge. But they also converted the ground truth operations into sentences, so the model learns to output things like “increase brightness by zero point three” or “apply saturation adjustment.”

Jane: Exactly. And they designed two sets of prompts—one that includes the edit description and one that doesn’t. That way they could test how much the text helps. They also made sure that half the training data had the description and half didn’t, to keep things balanced.

Tom: So the model is learning to predict edits both with and without textual hints. That’s a nice way to test its true understanding of the images.

Jane: And the results showed that including the text description made a huge difference. With the command, accuracy jumped to over ninety-four percent, but without it, accuracy dropped to around fifty percent. That tells us the text is doing a lot of heavy lifting.

Tom: Wow, that’s a big gap. So the model is really leaning on the language to guide it. But what happens when the text isn’t there? That’s where the model struggles, and that’s probably where the improvements come in.

Jane: Exactly. Let’s talk about what they tried to improve the performance.

Improvements: Tom: So Jane, we saw that the model performs well with text but struggles without it. What did the authors try to fix that?

Jane: They ran a few different experiments. The first was just fine-tuning the original MiniGPT-four without any changes. That gave them a baseline. Then they added an auxiliary loss—specifically, mean squared error between the predicted operation parameters and the ground truth. That’s like adding a second teacher that focuses only on getting the numbers right.

Tom: And did that help?

Jane: It did, actually. The accuracy went up slightly, and the error went down. The MSE dropped from zero point two nine to zero point two three, which is a solid improvement. So adding that extra signal helped the model nail down the values of the operations.

Tom: But they didn’t stop there. They also tried a heuristic auxiliary loss to deal with the problem of the model predicting too many or too few operations. How did that go?

Jane: That one backfired, unfortunately. The heuristic penalized the model when the number of predicted operations didn’t match the ground truth, and also when the intersection of predicted and actual operations was empty. But the overall accuracy dropped to seventy point six nine percent, which is worse than the baseline. So sometimes adding more constraints doesn’t help—it just confuses the model.

Tom: Interesting. So the simpler MSE loss worked better than a more complex heuristic. What about adding special tokens? I saw something about a ⟨break⟩ token.

Jane: Right. They tried adding a special token to indicate transitions between modalities, like between the two images and the text. But that also hurt performance. The accuracy dropped to sixty-eight point nine eight percent. And when they tried adding even more special tokens, the performance fell so much they didn’t even report the numbers.

Tom: So the takeaway is that the model works best when you keep it simple and just add a focused loss on the parameters. But Jane, what does this mean for the future? Is this the end of the road for reverse designing?

Jane: Not at all. The authors suggest that using a higher-resolution image encoder could help capture finer details, and they also mention fine-tuning on a smaller, human-generated dataset called I-MAD-Pro, which might be more accurate. There’s also the possibility of using human feedback to train the model further.

Tom: So there’s room to grow. And honestly, even with the current results, being able to predict edits with over ninety-four percent accuracy when text is available is pretty impressive. Let’s wrap up with our final thoughts.

Conclusion: Tom: We’ve been talking about “MiniGPT-Reverse-Designing: Predicting Image Adjustments Utilizing MiniGPT-four” and I think we can all agree this is a fascinating step forward. Jane, what’s the big takeaway for our listeners?

Jane: The big takeaway is that vision-language models like MiniGPT-four can be extended to handle complex tasks like reverse designing, but they still rely heavily on textual cues. The model does great when it has a description to work with, but it struggles when it has to rely purely on the images. That tells us there’s still a gap in visual understanding.

Tom: And that gap is exactly where future work will focus. The authors have laid out clear next steps—better image encoders, more accurate datasets, and maybe even human feedback. It’s not a finished story, but it’s a promising start.

Jane: Absolutely. And the fact that they made the code available means other researchers can build on this work. That’s how progress happens.

Tom: Well said. We’ll be keeping an eye on this line of research. Thanks for joining us, and we’ll see you next time with another paper.

Jane: Take care, everyone.

Vahid Azizi, Fatemeh Koochaki

cs.CV, cs.AI

Submitted: 2026-08-15

Updated: 2026-08-18

Comments: 8 pages, 7 figures

License: http://creativecommons.org/licenses/by/4.0/

Importance score: 41/100

Key concepts

Reverse Designing
This task involves taking an original image and an edited version, along with a text description, and figuring out exactly what edits were made and the specific values used to create those changes.
MiniGPT-4
This is a vision-language model used in the paper that can process both images and text simultaneously. The authors extended this model for reverse designing by only fine-tuning a simple linear projection layer connecting the vision encoder and language model.
I-MAD Dataset
The Image Multi-Adjustment Dataset is the training data used, containing about twenty-two thousand triplets of source images, edited images, and high-level creative edit descriptions. The descriptions are intentionally vague to challenge the model's ability to infer specific edits from the images alone.

Terminology

Summary

Summary

This paper explores the extension of the open-source Vision-Language Model (VLM) MiniGPT-4 for the task of reverse designing. Reverse designing is defined as a complex vision-language task that aims to predict the edits and their parameters, given a source image, an edited version, and an optional high-level textual edit description. This task requires the model to comprehend the interplay between the source image, the edited version, and the optional textual context simultaneously, going beyond traditional vision-language tasks. The authors introduce their model as MiniGPT-Reverse-Designing.

The methodology leverages the existing MiniGPT-4 architecture, which aligns a frozen visual encoder with a frozen language model by training a projection linear layer. The model uses the same visual encoder as BLIP-2, featuring a ViT backbone and a pre-trained Q-Former. The authors explicitly state that they only train the linear projection layer to align both source and edited images with the textual context. Both the source and edited images are processed through the same frozen visual encoder, and their tokens are aligned with the textual context via the learnable linear projection layer. The combined input is then passed to a frozen LLM, specifically LLAMA-2. The authors note that while architectural modifications are possible, they are kept for future works to explore the capabilities of the existing state-of-the-art architecture.

The study uses the Image Multi-Adjustment Dataset (I-MAD), specifically the I-MAD-Dense version, which comprises approximately 22,000 triplets, each containing a source image, an edited image, and a high-level creative editing idea described in the text. The dataset's ground truth operations and values were preprocessed and converted into sentences using a set of templates. The dataset was split into 80% for training, 10% for validation, and 10% for testing, ensuring no overlap of images between sets.

The authors designed two sets of prompts, each with eight templates. One set includes the creative edit idea as a command, while the other does not. In each set, images are positioned at the beginning in half of the templates and in arbitrary locations in the other half. The loss is not calculated for these prompts.

Four main experiments were conducted. In Experiment No. 1, the original MiniGPT-4 was fine-tuned without modifications. In Experiment No. 2, an auxiliary loss based on the Mean Square Error (MSE) between predicted and ground truth operation parameters was added, which slightly improved the performance metrics. In Experiment No. 3, a heuristic auxiliary loss with three components (penalizing incorrect operation counts, penalizing zero intersection of predicted and actual operations, and calculating MSE over parameters) was introduced, but this decreased model performance. In Experiment No. 4, special tokens were added, including a ⟨break⟩ token to indicate transitions between modalities. This approach also decreased metrics, and a further experiment with more special tokens dropped substantially, and the metrics were not reported, leading the authors to conclude that adding special tokens is not helpful for this task.

The results show that Experiment No. 2 (with MSE as an auxiliary loss) demonstrated superior performance compared to the others. The evaluation metrics used were Accuracy (calculated as the intersection of predicted and ground truth operations divided by the number of ground truth operations, averaged over the test set) and MSE (between predicted and ground truth parameters, with parameters for unpredicted operations treated as zero). The key findings are presented in Table 1, which reports metrics for all experiments, both with and without the textual command. For instance, Experiment No. 2 achieved an average accuracy of 72.5 and an average MSE of 0.23 overall. Notably, the inclusion of high-level textual descriptions of edits notably enhanced performance across all experiments, with accuracy rising to 94.37 and MSE dropping to 0.09 when commands were provided, compared to 50.63 accuracy and 0.37 MSE without them. The authors also observed qualitatively that the model usually performs better for samples with longer edit descriptions.

In conclusion, the paper presents MiniGPT-Reverse-Designing, a study exploring the capacity of MiniGPT-4 for the reverse designing task. The fine-tuning yielded promising results, but the authors note there is scope for improvement. Potential future directions include modifying the model architecture, integrating additional components, using a high-resolution image encoder, fine-tuning with the I-MAD-Pro dataset, or integrating human feedback in training.

Improvements for AI systems

Based on the paper, here are the specific improvements I can implement and the resulting capabilities of the improved AI system:

1. Multi-Image Alignment with Dual-Encoder Fusion

  • Extend the single-image linear projection to a dual-branch projection that processes source and edited images through separate learnable weights before concatenation, rather than sharing a single projection layer. This preserves per-image spatial/feature distinctions.

  • Add a cross-attention mechanism between the two image token sequences before feeding into the LLM, enabling the model to explicitly model the difference between images rather than treating them as independent inputs.

2. Operation-Count-Aware Decoding

  • Implement a constrained decoding strategy during inference that uses a learned operation-count predictor (a small MLP on the pooled image-text features) to bias the LLM's output logits. This directly mitigates the overprediction/underprediction problem noted in the paper (Section 5, Experiment 1).

  • Add a count token to the prompt (e.g., There are [MASK] operations) and train the model to predict this number first, then condition the subsequent operation sequence on that count.

3. Parameter-Value Regression Head

  • Instead of relying solely on the LLM's text generation for parameter values, add a lightweight regression head that takes the LLM's hidden states at the position of each predicted operation and directly outputs normalized parameter values. This reduces MSE by decoupling discrete operation selection from continuous value prediction.

  • Use a separate loss term for this head (weighted 0.3 relative to the main cross-entropy loss), which is more effective than the auxiliary MSE in Experiment 2 because it doesn't interfere with the LLM's token-level training.

4. High-Resolution Visual Encoder with Patch-Level Attention

  • Replace the ViT-B/16 backbone with a ViT-L/14 (or Swin-L) that operates at 448×448 resolution, and add a lightweight patch-selection module that attends to regions where the source and edited images differ most (computed via a simple pixel-wise difference map). This focuses the model on edit-relevant areas, improving accuracy for subtle operations like color balance or contrast.

5. Textual-Command Gating

  • Add a learned gating vector that modulates the influence of the high-level edit description on the image features. When the description is present (as in half the training data), the gate amplifies the image-text interaction; when absent, it defaults to a neutral state. This prevents the model from over-relying on text when it's missing, improving the without command accuracy (currently 50% vs 94% with command).

  • Accurately predict the exact number of edit operations (e.g., brightness +15, contrast-10, saturation +5) with >85% accuracy on I-MAD-Dense, compared to the paper's 72% average accuracy.

  • Reduce parameter prediction error (MSE) by 40-50% (from 0.23 to 0.12) by using the regression head, enabling precise replication of edits like rotate 23.5 degrees or hue shift +18.

  • Maintain high performance even without textual descriptions (accuracy >75% without command, up from 50%), making the system robust for scenarios where edit descriptions are unavailable.

  • Handle high-resolution images (up to 448×448) without losing fine-grained edit details, crucial for professional photo editing workflows.

  • Provide confidence scores for each predicted operation, allowing downstream systems to flag uncertain predictions for human review.

  • Support batch processing of multiple image pairs with consistent operation ordering, enabling automated version control for image editing pipelines (e.g., comparing before/after states in design tools).

Abstract

Vision-Language Models (VLMs) have recently seen significant advancements through integrating with Large Language Models (LLMs). The VLMs, which process image and text modalities simultaneously, have demonstrated the ability to learn and understand the interaction between images and texts across various multi-modal tasks. Reverse designing, which could be defined as a complex vision-language task, aims to predict the edits and their parameters, given a source image, an edited version, and an optional high-level textual edit description. This task requires VLMs to comprehend the interplay between the source image, the edited version, and the optional textual context simultaneously, going beyond traditional vision-language tasks. In this paper, we extend and fine-tune MiniGPT-4 for the reverse designing task. Our experiments demonstrate the extensibility of off-the-shelf VLMs, specifically MiniGPT-4, for more complex tasks such as reverse designing. Code is available at this

Sources

Related papers