VOR-Bench: A Human Perception-Driven Benchmark for Video Object Removal
cs.CV, cs.AI
Submitted: 2026-09-15
Updated: 2026-09-15
Comments: BMVC-2026
License: http://creativecommons.org/licenses/by/4.0/
The gist: Despite its crucial role in video object removal (VOR), existing evaluation paradigms face two critical limitations: questionable references and a misalignment between tradi- tional metrics and human
Terminology
Abstract
Despite its crucial role in video object removal (VOR), existing evaluation paradigms face two critical limitations: questionable references and a misalignment between tradi- tional metrics and human preference. To address these challenges, we introduce VOR- Bench, which advances VOR evaluation through three integrated components. First, we present the VOR Dataset (VORD), the first benchmark dataset providing both paired edited videos and graffiti masks. Its unique strength lies in a diverse data spectrum, which encompasses model-generated, tool-rendered, and camera-captured data, ensuring robust assessment across real-world scenarios. Second, we develop rMPAF, a realistic Motion- capable Paired-video Acquisition Framework. By combining the strengths of image- based object removal and fine-tuned video generation models, rMPAF automatically generates realistic, motion-coherent paired videos. Finally, we propose three evaluation dimensions and introduce VOR-MDSM, the first perception-driven VLM-based scoring model specifically designed for mask-guided VOR. It bridges the gap between arithmetic metrics and human perception by covering the essential visual attributes and matching nuanced human judgment. Extensive experiments demonstrate that VOR-Bench yields evaluation results that align closely with human perception, achieving a remarkable cor- relation (> 0.9) with subjective assessments. We will release VOR-Bench along with its documentation to ensure full reproducibility.
Sources
- Qwen2.5-VL Technical Report
- O-DisCo-Edit: Object Distortion Control for Unified Realistic Video Editing
- Vchitect-2.0: Parallel Transformer for Scaling Up Video Diffusion Models
- EffectErase: Joint Video Object Removal and Insertion for High-Quality Effect Erasing
- Coherent Video Inpainting Using Optical Flow-Guided Efficient Diffusion
- EditVerse: Unifying Image and Video Editing and Generation with In-Context Learning
- DiffuEraser: A Diffusion Model for Video Inpainting
- Flow Matching for Generative Modeling
- EraserDiT: Fast Video Inpainting with Diffusion Transformer Model
- VOID: Video Object and Interaction Deletion
- The 2017 DAVIS Challenge on Video Object Segmentation
- SAM 2: Segment Anything in Images and Videos
- Denoising Diffusion Implicit Models
- Wan: Open and Advanced Large-Scale Video Generative Models
- TDVE-Assessor: Benchmarking and Evaluating the Quality of Text-Driven Video Editing with LMMs
- CineTechBench: A Benchmark for Cinematographic Technique Understanding and Generation
- DiTPainter: Efficient Video Inpainting with Diffusion Transformers
- Qwen3 Technical Report
- MTV-Inpaint: Multi-Task Long Video Inpainting
- ImgEdit: A Unified Image Editing Dataset and Benchmark
Related papers
- Loss Knows Best: Detecting Annotation Errors in Videos via Loss Trajectories
- AnchorWeave: World-Consistent Video Generation with Retrieved Local Spatial Memories
- Benchmarking the Robustness of Foundation Models for Mammography under Domain Shift
- MambaX-Net: Dual-Input Mamba-Enhanced Cross-Attention Network for Longitudinal MRI Segmentation
- TeleOCR: Navigating Document Parsing Across Digital and Camera-Captured Documents
- A Survey on Efficient Vision-Language-Action Models