Timeline-Bench: Evaluating Agents on Realistic Video-Editing Tasks, from Raw Footage to Final Cut
cs.CV, cs.AI, cs.MM
Submitted: 2026-09-28
Updated: 2026-09-28
Code: https://github.com/Breakthrough/PySceneDetect
Project page: https://timelinebench.tensortest.com
Terminology
Sources
- AgenticVBench: Can AI Agents Complete Real-World Post-Production Tasks?
- PP-OCR: A Practical Ultra Lightweight OCR System
- MMTB: Evaluating Terminal Agents on Multimedia-File Tasks
- WorldSense: Evaluating Real-world Omnimodal Understanding for Multimodal LLMs
- CutVerse: A Compositional GUI Agents Benchmark for Media Post-Production Editing
- Terminal-Bench: Benchmarking Agents on Hard, Realistic Tasks in Command Line Interfaces
- MEDit-Bench: A Dataset for Evaluating Message-Driven Narrative Video Editing
- AutomationBench
- Agents' Last Exam
- VideoAgent: All-in-One Framework for Video Understanding and Editing
Related papers
- Loss Knows Best: Detecting Annotation Errors in Videos via Loss Trajectories
- AnchorWeave: World-Consistent Video Generation with Retrieved Local Spatial Memories
- Benchmarking the Robustness of Foundation Models for Mammography under Domain Shift
- MambaX-Net: Dual-Input Mamba-Enhanced Cross-Attention Network for Longitudinal MRI Segmentation
- TeleOCR: Navigating Document Parsing Across Digital and Camera-Captured Documents
- A Survey on Efficient Vision-Language-Action Models