V-Gym: Enhancing Agentic Visual Reasoning via Skill-Data Co-Evolution
cs.CV
Submitted: 2026-09-28
Updated: 2026-09-28
Terminology
Sources
- DecoEvo: Score-Decoupled Co-Evolution of Solver and Rubric-Generator Skills in Text Space
- Self-Play Meets Skill Evolution: Self-Evolving Search Agents that Pose, Solve, and Remember
- Beyond Seeing: Evaluating Multimodal LLMs on Tool-Enabled Image Perception, Transformation, and Reasoning
- From Failure to Mastery: Generating Hard Samples for Tool-use Agents
- Active Zero: Self-Evolving Vision-Language Models through Active Environment Exploration
- Towards On-Policy Data Evolution for Visual-Native Multimodal Deep Search Agents
- Skill Self-Play: Pushing the Frontier of LLM Capability with Co-Evolving Skills
- MM-BrowseComp: A Comprehensive Benchmark for Multimodal Browsing Agents
- SkillLens: Visual Skill Cards for Retrieval-Augmented GUI Action Prediction and On-Policy Distillation
- SKILL0: In-Context Agentic Reinforcement Learning for Skill Internalization
- Trace2Skill: Distill Trajectory-Local Lessons into Transferable Agent Skills
- SKILL-KD: Contrastive Skill Distillation for LLM Agents
- AgentVista: Evaluating Multimodal Agents in Ultra-Challenging Realistic Visual Scenarios
- DiagEvo: Diagnosis-Guided Self-Evolution via Hierarchical Error Memory
- SkillSmith: Co-Evolving Skills and Tools for Self-Improving Agent Systems
- Ace-Skill: Bootstrapping Multimodal Agents with Prioritized and Clustered Evolution
- SkillOpt: Executive Strategy for Self-Evolving Agent Skills
- VISA: Agentic Self-Evolving Data Synthesis for Multimodal Instruction Following
- What Makes Good Agentic Data? An ACE Lens on Data Generation for LLM Agents
Related papers
- Loss Knows Best: Detecting Annotation Errors in Videos via Loss Trajectories
- AnchorWeave: World-Consistent Video Generation with Retrieved Local Spatial Memories
- Benchmarking the Robustness of Foundation Models for Mammography under Domain Shift
- MambaX-Net: Dual-Input Mamba-Enhanced Cross-Attention Network for Longitudinal MRI Segmentation
- TeleOCR: Navigating Document Parsing Across Digital and Camera-Captured Documents
- A Survey on Efficient Vision-Language-Action Models