BVB: Benchmarking Agentic Video Understanding via Programmatic Reconstruction in Blender
cs.CV
Submitted: 2026-09-14
Updated: 2026-09-26
Code: https://github.com/ahujasid/blender-mcp
Terminology
Sources
- V-JEPA 2: Self-Supervised Video Models Enable Understanding, Prediction and Planning
- Seed1.5-VL Technical Report
- Thinking with Spatial Code for Physical-World Video Reasoning
- Gemini 2.5: Pushing the Frontier with Advanced Reasoning, Multimodality, Long Context, and Next Generation Agentic Capabilities
- GLM-4.5V and GLM-4.1V-Thinking: Towards Versatile Multimodal Reasoning with Scalable Reinforcement Learning
- Kimi K2: Open Agentic Intelligence
- VisPhyWorld: Probing Physical Reasoning via Code-Driven Video Reconstruction
- MiniMax-01: Scaling Foundation Models with Lightning Attention
- V-JEPA 2.1: Unlocking Dense Features in Video Self-Supervised Learning
- Qwen2.5-VL Technical Report
- Code-as-Room: Generating 3D Rooms from Top-Down View Images via Agentic Code Synthesis
- Vision-as-Inverse-Graphics Agent via Interleaved Multimodal Reasoning
- SceneActBench: Can Agents Act on the 3D Scenes They See?
Related papers
- Loss Knows Best: Detecting Annotation Errors in Videos via Loss Trajectories
- AnchorWeave: World-Consistent Video Generation with Retrieved Local Spatial Memories
- Benchmarking the Robustness of Foundation Models for Mammography under Domain Shift
- MambaX-Net: Dual-Input Mamba-Enhanced Cross-Attention Network for Longitudinal MRI Segmentation
- TeleOCR: Navigating Document Parsing Across Digital and Camera-Captured Documents
- A Survey on Efficient Vision-Language-Action Models