VisionFoundry: Teaching VLMs Visual Perception with Synthetic Images
summary
The gist
VisionFoundry introduces an end-to-end synthetic data generation pipeline designed to address persistent weaknesses in visual perception among vision-language models (VLMs) by providing targeted,
In short
VisionFoundry creates a synthetic data pipeline to improve vision-language models' visual perception by generating 10k high-quality, task-aware image-question pairs. The system uses an LLM to create prompts, a T2I model to generate images, and a multimodal judge to verify alignment. Results show this synthetic supervision significantly boosts performance on visual benchmarks.
Key concepts
- Task-Aware Generation
- The pipeline conditions the creation of questions and image prompts based on an explicit task configuration. This means specifying the target capability, like spatial relations or depth reasoning, allowing the system to generate content specifically designed to test that precise visual skill.
- Image Synthesis
- A modern text-to-image model is used to create images directly from detailed prompts generated by an LLM. The goal is to produce images that are visually grounded and unambiguous, ensuring they accurately reflect the visual facts required by the question.
- Alignment Verification
- A strong multimodal judge checks if the generated image actually matches the visual statement derived from the question and answer. If they don't align correctly, the sample is rejected. This crucial step filters out incorrect or misleading synthetic data, ensuring high quality for training.
Terminology used across episodes
This episode discusses
- VisionFoundry: Teaching VLMs Visual Perception with Synthetic Images · Paper Radio
- Phi-3 Technical Report: A Highly Capable Language Model Locally on Your Phone
- Qwen2.5-VL Technical Report
- ALLaVA: Harnessing GPT4V-Synthesized Data for Lite Vision-Language Models
- Gemini 2.5: Pushing the Frontier with Advanced Reasoning, Multimodality, Long Context, and Next Generation Agentic Capabilities
- The Llama 3 Herd of Models · Paper Radio
- VILA squared: VILA Augmented VILA
- Textbooks Are All You Need
- Synth squared: Boosting Visual-Language Models with Synthetic Captions and Image Embeddings
- LEGO-Puzzles: How Good Are MLLMs at Multi-Step Spatial Reasoning? · Paper Radio
- Qwen-VL: A Versatile Vision-Language Model for Understanding, Localization, Text Reading, and Beyond
- MiMo-VL Technical Report
- SVIT: Scaling up Visual Instruction Tuning
- Training on Thin Air: Improve Image Classification with Generated Data
The paper
VisionFoundry: Teaching VLMs Visual Perception with Synthetic Images · Read on arXiv
Princeton University · New York University
Vision-language models (VLMs) still struggle with visual perception tasks such as spatial understanding and viewpoint recognition, largely because natural image datasets provide limited supervision for low-level visual skills. Can targeted synthetic supervision address these weaknesses without reference images or manual annotation? To investigate this, we introduce VisionFoundry, an automated pipeline that takes only a task name as input, uses LLMs to synthesize paired questions, answers, and text-to-image (T2I) prompts, generates images with T2I models, and filters samples via multimodal verification. With VisionFoundry, we construct VisionFoundry-10k, a synthetic VQA dataset spanning 10 perception tasks. Finetuning on VisionFoundry-10k consistently improves perception benchmarks across three open-source backbones (e.g., +6.7% on MMVP-pair and +10.5% on CV-Bench-3D for Qwen2.5-VL-3B-Instruct) while preserving broader capabilities and showing positive data scaling. The same synthetic supervision also yields consistent gains under reinforcement learning (RL) across all three backbones, and the framework remains effective under open-source synthesis and self-verification. Our findings demonstrate that automated synthetic supervision offers an effective and scalable path toward systematic VLM training.
Transcript
Introduction to the show: ident: AI Radio. Generated commentary on the latest Artificial Intelligence papers.
Tom: Today's paper: "VisionFoundry: Teaching VLMs Visual Perception with Synthetic Images".
Jane: VisionFoundry introduces an end-to-end synthetic data generation pipeline designed to address persistent weaknesses in visual perception among vision-language models (VLMs) by providing targeted, task-aware supervision.
Tom: First, who's behind it and why it matters.
Title and authors: Tom: So we're looking at the paper titled "VisionFoundry: Teaching VLMs Visual Perception with Synthetic Images," and what it’s about is essentially a new way to give vision language models specific practice on things they struggle with, like spatial awareness and how objects are viewed.
Jane: It sounds like they’ve figured out that when we only feed these models regular pictures without telling them exactly what to look for, they miss the finer details of visual perception.
Lu: That makes sense because the natural datasets just don't provide enough targeted feedback for those low-level visual skills, which is a key observation from page zero of this paper.
Meng: So it’s moving away from just dumping massive amounts of raw data and instead creating something much more controlled and specific for training.
Lalam: Exactly, VisionFoundry creates synthetic supervision that can be generated on demand to target specific failures in the AI's visual understanding.
The paper's summary: Tom: The main summary of this paper is about building an entire pipeline that works just from a task keyword. Instead of needing reference images or human captions, the system uses large language models to create questions and detailed text-to-image prompts tailored precisely to that task.
Jane: Then, a text-to-image model generates the picture based on those prompts, and finally, this generated image goes through a strong multimodal judge that checks if the visual elements actually match what the question asks for.
Lu: What’s really impressive is how they set up this tight coupling between the language supervision and the visual content right at generation time to ensure determinism, which is something we need when training models.
Meng: That verification stage sounds like a safety net, making sure that only high-quality visual examples get into the training set rather than noisy or misaligned data.
Lalam: This process results in VisionFoundry-10K, a synthetic visual question answering dataset of 10k image–question–answer triples covering ten different low-level tasks.
The paper's improvements: Tom: The paper points out several structural improvements to this pipeline, focusing on how they construct the data. They use an adaptive concept pool where the LLM builds objects, attributes, and scenes based on the task configuration before generating prompts for the T2I model.
Jane: This means every image generated is highly conditioned on specific visual facts that are relevant to that particular task, which is much better than random caption-derived examples.
Lu: They also emphasize the role of composition in forming entities during sampling, which allows them to control what combinations of objects appear in the image according to the task constraints.
Meng: The verification step is highlighted as essential for quality control; they show that without this judge, performance drops on benchmarks like CV-Bench-2D and RealWorldQA.
Lalam: This synthetic approach is presented not just as a data source, but as an initial step toward a more systematic synthetic-data workflow for multimodal training, which is quite significant.
Conclusion: Tom: To wrap things up, the paper with its focus on "VisionFoundry: Teaching VLMs Visual Perception with Synthetic Images" shows that targeted synthetic curation can meaningfully alleviate part of the perception bottleneck in current models.
Jane: It highlights that simply increasing the size of natural image datasets isn't enough; we need to deliberately construct supervision for those specific visual distinctions that models currently underlearn.
Lu: The implication is that progress in multimodal systems might depend more on carefully designing this kind of synthetic supervision than just scaling up model size or raw data volume.
Meng: From a practical standpoint, this suggests a path where we can systematically train models to excel at core visual perception tasks by feeding them precisely what they need to learn.
Lalam: VisionFoundry demonstrates how modern LLMs and T2I models can be composed into a principled system for generating synthetic multimodal data, offering a promising path toward more systematic training for VLMs.
More episodes
- 2610.10857-Self-Supervised Keyframe Discovery for Horizon-Invariant Behavior Cloning
- 2610.10768-Strategic Investment Decision Making for Value Creation in Energy Transition: A Reinforcement Learning Approach
- 2610.10858-RFChipAgent: Multi-Agentic AI Flow for Analog/RF Chip Design
- 2610.10613-Temporal transformer CAN encoder with federated lightweight heads for anomaly detection
- 2610.10616-When Routing Reveals Membership: Privacy Leakage from MoE Router Telemetry
- 2610.10655-Nullify: Null-Space Activation Steering for Training-Free LLM Unlearning
- 2610.11031-Language Modeling is Monotone Compression
- 2610.01253-Context-Aware Error Mitigation Orchestration for Hybrid Quantum Reinforcement Learning on NISQ Systems
- 2604.24201-CMGL: Confidence-guided Multi-omics Graph Learning for Cancer Subtype Classification
- 2609.34069-Towards Certificate-Driven Software Porting: A Self-Improving Agentic Harness for Scientific Program Optimization