VisionFoundry: Teaching VLMs Visual Perception with Synthetic Images

summary

Video file (mp4)

The gist

VisionFoundry introduces an end-to-end synthetic data generation pipeline designed to address persistent weaknesses in visual perception among vision-language models (VLMs) by providing targeted,

In short

VisionFoundry creates a synthetic data pipeline to improve vision-language models' visual perception by generating 10k high-quality, task-aware image-question pairs. The system uses an LLM to create prompts, a T2I model to generate images, and a multimodal judge to verify alignment. Results show this synthetic supervision significantly boosts performance on visual benchmarks.

Key concepts

Task-Aware Generation
The pipeline conditions the creation of questions and image prompts based on an explicit task configuration. This means specifying the target capability, like spatial relations or depth reasoning, allowing the system to generate content specifically designed to test that precise visual skill.
Image Synthesis
A modern text-to-image model is used to create images directly from detailed prompts generated by an LLM. The goal is to produce images that are visually grounded and unambiguous, ensuring they accurately reflect the visual facts required by the question.
Alignment Verification
A strong multimodal judge checks if the generated image actually matches the visual statement derived from the question and answer. If they don't align correctly, the sample is rejected. This crucial step filters out incorrect or misleading synthetic data, ensuring high quality for training.

Terminology used across episodes

This episode discusses

The paper

VisionFoundry: Teaching VLMs Visual Perception with Synthetic Images · Read on arXiv

Princeton University · New York University

Vision-language models (VLMs) still struggle with visual perception tasks such as spatial understanding and viewpoint recognition, largely because natural image datasets provide limited supervision for low-level visual skills. Can targeted synthetic supervision address these weaknesses without reference images or manual annotation? To investigate this, we introduce VisionFoundry, an automated pipeline that takes only a task name as input, uses LLMs to synthesize paired questions, answers, and text-to-image (T2I) prompts, generates images with T2I models, and filters samples via multimodal verification. With VisionFoundry, we construct VisionFoundry-10k, a synthetic VQA dataset spanning 10 perception tasks. Finetuning on VisionFoundry-10k consistently improves perception benchmarks across three open-source backbones (e.g., +6.7% on MMVP-pair and +10.5% on CV-Bench-3D for Qwen2.5-VL-3B-Instruct) while preserving broader capabilities and showing positive data scaling. The same synthetic supervision also yields consistent gains under reinforcement learning (RL) across all three backbones, and the framework remains effective under open-source synthesis and self-verification. Our findings demonstrate that automated synthetic supervision offers an effective and scalable path toward systematic VLM training.

Transcript

Introduction to the show: ident: AI Radio. Generated commentary on the latest Artificial Intelligence papers.

Tom: Today's paper: "VisionFoundry: Teaching VLMs Visual Perception with Synthetic Images".

Jane: VisionFoundry introduces an end-to-end synthetic data generation pipeline designed to address persistent weaknesses in visual perception among vision-language models (VLMs) by providing targeted, task-aware supervision.

Tom: First, who's behind it and why it matters.

Title and authors: Tom: So we're looking at the paper titled "VisionFoundry: Teaching VLMs Visual Perception with Synthetic Images," and what it’s about is essentially a new way to give vision language models specific practice on things they struggle with, like spatial awareness and how objects are viewed.

Jane: It sounds like they’ve figured out that when we only feed these models regular pictures without telling them exactly what to look for, they miss the finer details of visual perception.

Lu: That makes sense because the natural datasets just don't provide enough targeted feedback for those low-level visual skills, which is a key observation from page zero of this paper.

Meng: So it’s moving away from just dumping massive amounts of raw data and instead creating something much more controlled and specific for training.

Lalam: Exactly, VisionFoundry creates synthetic supervision that can be generated on demand to target specific failures in the AI's visual understanding.

The paper's summary: Tom: The main summary of this paper is about building an entire pipeline that works just from a task keyword. Instead of needing reference images or human captions, the system uses large language models to create questions and detailed text-to-image prompts tailored precisely to that task.

Jane: Then, a text-to-image model generates the picture based on those prompts, and finally, this generated image goes through a strong multimodal judge that checks if the visual elements actually match what the question asks for.

Lu: What’s really impressive is how they set up this tight coupling between the language supervision and the visual content right at generation time to ensure determinism, which is something we need when training models.

Meng: That verification stage sounds like a safety net, making sure that only high-quality visual examples get into the training set rather than noisy or misaligned data.

Lalam: This process results in VisionFoundry-10K, a synthetic visual question answering dataset of 10k image–question–answer triples covering ten different low-level tasks.

The paper's improvements: Tom: The paper points out several structural improvements to this pipeline, focusing on how they construct the data. They use an adaptive concept pool where the LLM builds objects, attributes, and scenes based on the task configuration before generating prompts for the T2I model.

Jane: This means every image generated is highly conditioned on specific visual facts that are relevant to that particular task, which is much better than random caption-derived examples.

Lu: They also emphasize the role of composition in forming entities during sampling, which allows them to control what combinations of objects appear in the image according to the task constraints.

Meng: The verification step is highlighted as essential for quality control; they show that without this judge, performance drops on benchmarks like CV-Bench-2D and RealWorldQA.

Lalam: This synthetic approach is presented not just as a data source, but as an initial step toward a more systematic synthetic-data workflow for multimodal training, which is quite significant.

Conclusion: Tom: To wrap things up, the paper with its focus on "VisionFoundry: Teaching VLMs Visual Perception with Synthetic Images" shows that targeted synthetic curation can meaningfully alleviate part of the perception bottleneck in current models.

Jane: It highlights that simply increasing the size of natural image datasets isn't enough; we need to deliberately construct supervision for those specific visual distinctions that models currently underlearn.

Lu: The implication is that progress in multimodal systems might depend more on carefully designing this kind of synthetic supervision than just scaling up model size or raw data volume.

Meng: From a practical standpoint, this suggests a path where we can systematically train models to excel at core visual perception tasks by feeding them precisely what they need to learn.

Lalam: VisionFoundry demonstrates how modern LLMs and T2I models can be composed into a principled system for generating synthetic multimodal data, offering a promising path toward more systematic training for VLMs.

More episodes

← Home