VisionFoundry: Teaching VLMs Visual Perception with Synthetic Images

arXiv:2604.09531 · cs.CV, cs.AI, cs.CL · Submitted 2026-04-10 · Read on arXiv

Listen

Radio episode about this paper

Transcript

Introduction to the show: ident: AI Radio. Generated commentary on the latest Artificial Intelligence papers.

Tom: Today's paper: "VisionFoundry: Teaching VLMs Visual Perception with Synthetic Images".

Jane: VisionFoundry introduces an end-to-end synthetic data generation pipeline designed to address persistent weaknesses in visual perception among vision-language models (VLMs) by providing targeted, task-aware supervision.

Tom: First, who's behind it and why it matters.

Title and authors: Tom: So we're looking at the paper titled "VisionFoundry: Teaching VLMs Visual Perception with Synthetic Images," and what it’s about is essentially a new way to give vision language models specific practice on things they struggle with, like spatial awareness and how objects are viewed.

Jane: It sounds like they’ve figured out that when we only feed these models regular pictures without telling them exactly what to look for, they miss the finer details of visual perception.

Lu: That makes sense because the natural datasets just don't provide enough targeted feedback for those low-level visual skills, which is a key observation from page zero of this paper.

Meng: So it’s moving away from just dumping massive amounts of raw data and instead creating something much more controlled and specific for training.

Lalam: Exactly, VisionFoundry creates synthetic supervision that can be generated on demand to target specific failures in the AI's visual understanding.

The paper's summary: Tom: The main summary of this paper is about building an entire pipeline that works just from a task keyword. Instead of needing reference images or human captions, the system uses large language models to create questions and detailed text-to-image prompts tailored precisely to that task.

Jane: Then, a text-to-image model generates the picture based on those prompts, and finally, this generated image goes through a strong multimodal judge that checks if the visual elements actually match what the question asks for.

Lu: What’s really impressive is how they set up this tight coupling between the language supervision and the visual content right at generation time to ensure determinism, which is something we need when training models.

Meng: That verification stage sounds like a safety net, making sure that only high-quality visual examples get into the training set rather than noisy or misaligned data.

Lalam: This process results in VisionFoundry-10K, a synthetic visual question answering dataset of 10k image–question–answer triples covering ten different low-level tasks.

The paper's improvements: Tom: The paper points out several structural improvements to this pipeline, focusing on how they construct the data. They use an adaptive concept pool where the LLM builds objects, attributes, and scenes based on the task configuration before generating prompts for the T2I model.

Jane: This means every image generated is highly conditioned on specific visual facts that are relevant to that particular task, which is much better than random caption-derived examples.

Lu: They also emphasize the role of composition in forming entities during sampling, which allows them to control what combinations of objects appear in the image according to the task constraints.

Meng: The verification step is highlighted as essential for quality control; they show that without this judge, performance drops on benchmarks like CV-Bench-2D and RealWorldQA.

Lalam: This synthetic approach is presented not just as a data source, but as an initial step toward a more systematic synthetic-data workflow for multimodal training, which is quite significant.

Conclusion: Tom: To wrap things up, the paper with its focus on "VisionFoundry: Teaching VLMs Visual Perception with Synthetic Images" shows that targeted synthetic curation can meaningfully alleviate part of the perception bottleneck in current models.

Jane: It highlights that simply increasing the size of natural image datasets isn't enough; we need to deliberately construct supervision for those specific visual distinctions that models currently underlearn.

Lu: The implication is that progress in multimodal systems might depend more on carefully designing this kind of synthetic supervision than just scaling up model size or raw data volume.

Meng: From a practical standpoint, this suggests a path where we can systematically train models to excel at core visual perception tasks by feeding them precisely what they need to learn.

Lalam: VisionFoundry demonstrates how modern LLMs and T2I models can be composed into a principled system for generating synthetic multimodal data, offering a promising path toward more systematic training for VLMs.

Princeton University · New York University

cs.CV, cs.AI, cs.CL

Submitted: 2026-04-10

Updated: 2026-09-30

Code: https://github.com/XiaomiMiMo/MiMo-VL

Importance score: 90/100

The gist: VisionFoundry introduces an end-to-end synthetic data generation pipeline designed to address persistent weaknesses in visual perception among vision-language models (VLMs) by providing targeted,

Key concepts

Task-Aware Generation
The pipeline conditions the creation of questions and image prompts based on an explicit task configuration. This means specifying the target capability, like spatial relations or depth reasoning, allowing the system to generate content specifically designed to test that precise visual skill.
Image Synthesis
A modern text-to-image model is used to create images directly from detailed prompts generated by an LLM. The goal is to produce images that are visually grounded and unambiguous, ensuring they accurately reflect the visual facts required by the question.
Alignment Verification
A strong multimodal judge checks if the generated image actually matches the visual statement derived from the question and answer. If they don't align correctly, the sample is rejected. This crucial step filters out incorrect or misleading synthetic data, ensuring high quality for training.

Terminology

Summary

VisionFoundry introduces an end-to-end synthetic data generation pipeline designed to address persistent weaknesses in visual perception among vision-language models (VLMs) by providing targeted, task-aware supervision. This research is significant because it demonstrates that limited supervision in natural image datasets hinders low-level visual skills, and it proves that controlled synthetic data can serve as a promising path toward more systematic training for VLMs, leading to substantial improvements on critical visual perception benchmarks while preserving broader capabilities.

VisionFoundry Pipeline Overview

VisionFoundry is a task-aware synthetic data generation pipeline that produces high-quality VQA supervision for VLMs using only task specifications, without reference images, human-written QA annotations, or real image–caption pairs. This system is guided by three methodological principles: controllability, visual determinism, and verification. Operationally, the pipeline comprises three stages:

  1. A large language model generates question–answer pairs and detailed text-to-image (T2I) prompts conditioned on the target task.

  2. A modern T2I model synthesizes images conditioned on those prompts.

  3. A strong multimodal judge verifies alignment between the generated image and the answer-determining visual statement, filtering out misaligned samples.

Data Construction and Control

VisionFoundry constructs VisionFoundry-10K, a synthetic visual question answering (VQA) dataset of 10k image–question–answer triples spanning 10 carefully selected low-level visual perception tasks (e.g., spatial understanding, depth reasoning, viewpoint variation). The construction process ensures high quality and control through several mechanisms:

((

(i) task-aware generation:

The pipeline conditions generation on an explicit task configuration, which specifies the target capability (e.g., spatial relations), the number of objects involved, and optional constraints such as required attributes or relations. The LLM is prompted to generate a question whose answer is fully determined by visible content, and a highly detailed T2I prompt that explicitly encodes the answer-determining visual facts. This design enforces tight coupling between language supervision and visual content at generation time.

(ii) image synthesis:

Images are generated directly from the T2I prompt using a modern model (e.g., Gemini-2.5-Flash-Image), with limited iterative refinement allowed if initial verification fails, ensuring images are visually grounded, unambiguous, and robust.

(iii) alignment verification and filtering:

A strong proprietary multimodal judge acts as a verifier. It reads the generated image together with a short declarative visual statement derived from the question and answer to return an accept/reject decision based on whether the core visual relation is actually present, ensuring only samples for which the judge confirms alignment are retained.

Experimental Results and Impact

Models trained on VisionFoundry-10K consistently show substantial improvements on visual perception benchmarks. For instance, Qwen2.5-VL-3B-Instruct achieved a +7% on MMVP and a +10% on CV-Bench-3D. These gains validate that the synthetic supervision effectively targets core visual weaknesses identified in prior work, particularly those related to spatial layout, depth ordering, and attribute recognition. Furthermore, the study shows that performance improves predictably as data size increases; for a representative task, performance improves as synthetic data size increases. The results suggest that carefully designed synthetic data can serve as a practical complementary source of supervision for strengthening core visual perception.

Verification and Process Ablations

The role of the verification stage is crucial for maintaining dataset quality. An ablation study on a single task with a fixed training budget showed that verification is important, with the verified pipeline consistently performing better than the non-verified variant across benchmarks like CV-Bench-2D and RealWorldQA. Furthermore, comparing different supervision methods, the synthetic-image setting outperforms the natural-image setting on all five reported benchmarks, indicating that the full synthetic process provides added value beyond QA generation in VisionFoundry. This suggests that the caption-conditioned T2I reconstruction step offers additional training value beyond caption-derived synthetic QA alone.

Conclusion and Future Directions

VisionFoundry demonstrates how modern LLMs and T2I models can be composed into a principled system for synthetic multimodal data generation. The findings indicate that the perception bottleneck in VLMs is partly a data problem, and targeted synthetic curation can meaningfully alleviate part of this weakness. The work suggests that progress in multimodal systems may depend not only on scaling model size or natural data volume, but also on deliberately constructing supervision for the specific perceptual distinctions that current models underlearn, positioning synthetic supervision as a promising path toward more systematic training for VLMs. Future work is suggested to explore whether such a pipeline can benefit more complex visual reasoning tasks requiring longer inference chains.

Improvements for AI systems

Based on the provided research paper, VisionFoundry: Teaching VLMs Visual Perception with Synthetic Images, here are specific, actionable improvements for AI systems and a description of what these improved systems can achieve.


The core contribution of VisionFoundry is demonstrating that targeted, task-aware synthetic data generation—using LLMs to create questions/answers and T2I models to synthesize images, verified by a strong multimodal judge—can significantly boost the low-level visual perception capabilities of Vision-Language Models (VLMs).

Here are the specific improvements and capabilities:

  1. The system should be integrated with a closed-loop pipeline that takes only a task keyword (e.g., Depth Order, Viewpoint and Perspective) as input to generate high-quality supervision.

  2. The system can use Large Language Models (LLMs) to dynamically construct an adaptive concept pool, entity pools (objects, attributes, scenes), and T2I prompts that are explicitly conditioned on the target task's constraints (e.g., requiring spatial relationships or specific colors).

  3. The system must incorporate a modern Text-to-Image (T2I) model to synthesize images directly from these task-specific prompts, ensuring the resulting image contains all necessary visual facts to answer the question deterministically.

  4. The system should employ a strong proprietary multimodal judge (like Gemini-3-Pro) to verify alignment between the generated image and the provided answer statement, automatically filtering out samples where visual relations are incorrect or ambiguous. This verification step is crucial for data quality and robustness.

  5. The resulting synthetic dataset (VisionFoundry-10K) can be used for systematic finetuning of VLMs on low-level visual perception tasks, leading to measurable gains (+7% on MMVP, +10% on CV-Bench-3D).

Specific improved AI system capabilities:

The improved AI system, trained or fine-tuned using VisionFoundry data, can perform the following:

  1. Perform highly accurate spatial reasoning tasks (e.g., identifying relative depth order between objects like a submarine and a knife, determining if an object is closer to the camera).

  2. Accurately recognize and describe viewpoint-dependent visual characteristics (e.g., determining if an object is facing toward or away from the camera, recognizing specific orientations).

  3. Precisely discriminate between subtle visual attributes across different scenes (e.g., correctly identifying two distinct colors on a sofa).

  4. Understand complex structural and physical characteristics of objects (e.g., describing the silhouette profile of an object like a barometer).

  5. Execute real-world spatial understanding, such as answering questions about relative positions in blueprints or scenes (e.g., determining if one object is west/left of another).

In essence, the improved AI system will move beyond general language and semantic understanding to possess robust, verifiable geometric and perceptual intelligence derived from targeted synthetic data supervision.

Sources

Related papers