NEST: Narrative Event Structures in Time for Long Video Understanding
cs.CV, cs.CL
Submitted: 2026-06-18
Updated: 2026-09-05
Comments: EMNLP 2026 (Main)
License: http://arxiv.org/licenses/nonexclusive-distrib/1.0/
Terminology
Sources
- Phi-3 Technical Report: A Highly Capable Language Model Locally on Your Phone
- InfiniBench: A Benchmark for Large Multi-Modal Models in Long-Form Movies and TV Shows
- Qwen-VL: A Versatile Vision-Language Model for Understanding, Localization, Text Reading, and Beyond
- Qwen2.5-VL Technical Report
- LLMs instead of Human Judges? A Large Scale Empirical Study across 20 NLP Evaluation Tasks
- Molmo and PixMo: Open Weights and Open Data for State-of-the-Art Vision-Language Models
- Event Extraction by Answering (Almost) Natural Questions
- CLAP: Learning Audio Concepts From Natural Language Supervision
- Joint Multimedia Event Extraction from Video and Article
- Video-MME: The First-Ever Comprehensive Evaluation Benchmark of Multi-modal LLMs in Video Analysis
- ReXTime: A Benchmark Suite for Reasoning-Across-Time in Videos
- Grounded Multi-Hop VideoQA in Long-Form Egocentric Videos
- Extending Context Window of Large Language Models via Positional Interpolation
- The Llama 3 Herd of Models
- Gemini 2.5: Pushing the Frontier with Advanced Reasoning, Multimodality, Long Context, and Next Generation Agentic Capabilities
- A Survey on LLM-as-a-Judge
- FlashAttention: Fast and Memory-Efficient Exact Attention with IO-Awareness
- RULER: What's the Real Context Size of Your Long-Context Language Models?
- LoRA: Low-Rank Adaptation of Large Language Models
- Visual Instruction Tuning
Related papers
- Loss Knows Best: Detecting Annotation Errors in Videos via Loss Trajectories
- AnchorWeave: World-Consistent Video Generation with Retrieved Local Spatial Memories
- Benchmarking the Robustness of Foundation Models for Mammography under Domain Shift
- MambaX-Net: Dual-Input Mamba-Enhanced Cross-Attention Network for Longitudinal MRI Segmentation
- TeleOCR: Navigating Document Parsing Across Digital and Camera-Captured Documents
- A Survey on Efficient Vision-Language-Action Models