Exemplar2VQA: A Scalable Exemplar-Driven Visual Question Answering Generation Framework via Multi-Agent Coding
cs.CV
Submitted: 2026-09-29
Updated: 2026-09-29
Code: https://github.com/yingjiayu12/Exemplar2VQA
Terminology
Sources
- AutoGen: Enabling Next-Gen LLM Applications via Multi-Agent Conversation
- AI2-THOR: An Interactive 3D Environment for Visual AI
- CodeACT: Code Adaptive Compute-efficient Tuning Framework for Code LLMs
- Qwen2.5-VL Technical Report
- ViewSpatial-Bench: Evaluating Multi-perspective Spatial Localization in Vision-Language Models
- Multi-SpatialMLLM: Multi-Frame Spatial Understanding with Multi-Modal Large Language Models
- MMSI-Bench: A Benchmark for Multi-Image Spatial Intelligence
- Qwen3 Technical Report
- InternVL3: Exploring Advanced Training and Test-Time Recipes for Open-Source Multimodal Models
- Expanding Performance Boundaries of Open-Source Multimodal Models with Model, Data, and Test-Time Scaling
- DeepSeek-VL2: Mixture-of-Experts Vision-Language Models for Advanced Multimodal Understanding
- SpaCE-10: A Comprehensive Benchmark for Multimodal Large Language Models in Compositional Spatial Intelligence
- GPT-4o System Card
- Gemini 1.5: Unlocking multimodal understanding across millions of tokens of context
- Are Multimodal Large Language Models Ready for Omnidirectional Spatial Reasoning?
Related papers
- Loss Knows Best: Detecting Annotation Errors in Videos via Loss Trajectories
- AnchorWeave: World-Consistent Video Generation with Retrieved Local Spatial Memories
- Benchmarking the Robustness of Foundation Models for Mammography under Domain Shift
- MambaX-Net: Dual-Input Mamba-Enhanced Cross-Attention Network for Longitudinal MRI Segmentation
- TeleOCR: Navigating Document Parsing Across Digital and Camera-Captured Documents
- A Survey on Efficient Vision-Language-Action Models