CausalChapter: Improving Long-Video Chaptering with Interventional Dependency Modeling

arXiv:2609.08686 · cs.CV, cs.AI · Submitted 2026-09-08 · Read on arXiv

cs.CV, cs.AI

Submitted: 2026-09-08

Updated: 2026-09-08

Comments: Accepted by EMNLP 2026 conference

License: http://creativecommons.org/licenses/by/4.0/

The gist: Long-form instructional videos require automatic chaptering to support browsing, navigation, and knowledge access.

Terminology

Abstract

Long-form instructional videos require automatic chaptering to support browsing, navigation, and knowledge access. Recent long-context language models can perform chaptering from textualized video inputs, but they remain costly and brittle for content-dense lecture videos with long transcripts, smooth topic transitions, and detailed chapter outputs. A scalable segment-then-caption paradigm reduces this cost, but introduces two new challenges: boundary error propagation and fragmented cross-chapter context. We propose CausalChapter, an intervention-inspired framework for long-video chaptering that estimates prediction-level influence through lightweight masking and removal interventions. For boundary localization, our Local Dependency Shift module detects drops in predictive dependency between adjacent temporal windows; for chapter description generation, our Cross-Segment Support Selection module reranks historical contexts according to their support for the current prediction. Experiments on long-video chaptering benchmarks show that CausalChapter improves boundary localization, chapter description quality, and cross-chapter coherence.

Related papers