NARU: A Benchmark for NARrative Evolution and Cultural Nuance Understanding in Japanese Extreme Long Video
Yuheng Huang, Jianlang Chen, Jiayang Song, Hua Qi, Aza Kai, Vincent Markert, Edison Marrese-Taylor, Jianjun Zhao, Lei Ma
The University of Tokyo · Kyushu University · Macau University of Science and Technology · Infinimind Japan Inc. · University of Alberta
cs.CV, cs.AI, cs.MM
Submitted: 2026-08-13
Updated: 2026-08-14
Comments: Yuheng Huang and Jianlang Chen contributed equally to this work. More details available on the project's website https://ma-labo.github.io/naru/ and https://infinimind.io/en/company/news/2026/narubench-release
Project page: https://ma-labo.github.io/naru
License: http://arxiv.org/licenses/nonexclusive-distrib/1.0/
Importance score: 75/100
The gist: NARU is a benchmark designed to evaluate narrative evolution and cultural understanding in Japanese extreme long-form video.
Terminology
Summary
NARU is a benchmark designed to evaluate narrative evolution and cultural understanding in Japanese extreme long-form video. It consists of 1,481 multiple-choice questions grounded in 155 videos totaling 146.8 hours, spanning four narrative and five cultural dimensions. The benchmark was constructed using a hierarchical memory-based annotation pipeline that transforms raw video into structured event, narrative, and cultural annotations, followed by task-oriented question synthesis and iterative shortcut removal. The construction process included two native-speaker verification stages involving 68 annotators.
The paper distinguishes context-rich long-form video
from video that is merely long in duration, formulating its understanding as the joint problem of maintaining narrative state and interpreting culturally situated implicit meaning. The benchmark evaluates two core competencies: Narrative Intelligence (tracking story evolution across videos of 30 to 240 minutes) and Native Cultural Understanding (comprehending Japanese conversational nuances, social dynamics, and implicit communicative signals).
The capability taxonomy includes four narrative subcategories: N.1 Character/Entity Evolution (local coherence), N.2 Sequential/Topical Flow (local coherence), N.3 Plot/Conflict Progression (global coherence), and N.4 Idea/Thematic Development (global coherence). It also includes five cultural subcategories: C.1 Aizuchi (interactional signalling), C.2 Kuuki wo Yomu (shared situational understanding), C.3 Subtext Interpretation (speaker intent), C.4 Cultural Context Recognition (cultural grounding), and C.5 Sentiment Analysis (affective and interpersonal dynamics).
The dataset construction involved acquiring over 100,000 unique videos from YouTube, filtering for Japanese-language content (51,643 videos), and retaining videos of at least 30 minutes duration, resulting in 8,018 candidates. Videos were then manually screened against criteria including visual integrity and temporal semantic progression, followed by incremental semantic diversity sampling to select the final 155 videos.
The MLLM-driven QA generation pipeline consists of two stages. The annotation stage uses chunking (five-minute units), semantic segmentation (chapter-level segments), and task-oriented annotation (narrative and cultural modules) to create a structured representation. The question-answer pair generation stage uses capability-conditioned formulation and iterative debiasing refinement, including a Solver-Critic Loop with a Blind Solver Agent, Diagnostic Agent, and Question Revision Agent to eliminate text-only shortcuts.
Quality control involved two human validation stages: 40 native Japanese experts verified initial QA items (178 flagged as invalid, 161 repaired, 17 removed), and 28 experts verified post-refinement items (949 accepted unchanged, 532 revised, 2 removed). The final benchmark contains 1,481 items.
Evaluation across eight model configurations revealed substantial limitations in both long-range narrative integration and culturally grounded reasoning. Proprietary models (Gemini-3-Flash at 76.2%, Gemini-3-Pro at 70.0%, Gemini-2.5-Flash at 51.4%) significantly outperformed open-source models (29.6–39.8%). Gemini models achieved approximately 11 percentage points higher accuracy on narrative tasks than cultural ones, while open-source models showed no difference. Weaker models failed primarily at tracking low-level entity continuity (N.1), while stronger models struggled most with high-level abstraction (N.4). C.3 (Subtext Interpretation) proved exceptionally difficult even for frontier models.
Frame-budget experiments showed that increasing sampled frames from 8 to 128 improved narrative understanding far more than cultural understanding, with narrative accuracy rising by 1.4 to 20.5 percentage points while cultural accuracy fluctuated between a 3.9-point decline and a 10.0-point gain. Open-ended evaluation using FActScore recall revealed that N.2 (Sequential/Topical Flow) degraded from the strongest narrative category in multiple-choice settings to the weakest in open-ended settings, and the relative difficulty between narrative and cultural understanding reversed across formats.
Improvements for AI systems
Improvements to AI Systems:
- Hierarchical Memory Architecture for Long-Form Video
-
Implement a multi-scale memory module that maintains event-level (5-minute chunks), chapter-level (semantic segments), and narrative-level (global story arc) representations simultaneously.
-
The system can track character/entity state changes across 30–240 minute videos without losing local coherence (N.1) or global plot progression (N.3), addressing the failure of weaker models on entity continuity and stronger models on thematic abstraction (N.4).
- Dual-Stream Reasoning: Narrative State Tracking + Cultural Grounding
-
Add a separate cultural inference pathway that processes implicit conversational signals (aizuchi, kuuki wo yomu, subtext) in parallel with factual narrative tracking.
-
The system can distinguish between
what happened
(narrative) andwhat is meant
(cultural subtext), improving performance on C.3 (Subtext Interpretation) which frontier models currently fail, and closing the 11-point gap between narrative and cultural accuracy seen in Gemini models.
- Adaptive Frame Sampling with Task-Aware Allocation
-
Dynamically allocate more frames to narrative-critical segments (e.g., scene transitions, character introductions) and fewer to culturally redundant dialogue scenes, rather than uniform sampling.
-
The system can achieve higher narrative accuracy (up to +20.5 points) without sacrificing cultural understanding, avoiding the trade-off where increased frames improved narrative but left cultural accuracy stagnant or declining.
- Blind Solver-Critic Debiasing for Robust QA
-
Integrate a self-refinement loop during inference where a
blind
text-only solver identifies questions answerable without video, and a diagnostic agent flags such shortcuts for revision. -
The system becomes resistant to dataset artifacts and can generalize to open-ended generation tasks, preventing the format-dependent reversal seen where N.2 (Sequential/Topical Flow) dropped from strongest in multiple-choice to weakest in open-ended settings.
- Cross-Format Consistency Training
-
Train models jointly on multiple-choice and open-ended generation objectives with contrastive loss to align their internal representations across formats.
-
The system maintains consistent relative difficulty between narrative and cultural understanding regardless of output format, avoiding the reversal where narrative outperformed cultural in multiple-choice but underperformed in open-ended recall (FActScore).
- Capability-Specific Fine-Tuning with Hierarchical Taxonomy
-
Use the 9 subcategories (N.1–N.4, C.1–C.5) as explicit training targets, with separate loss heads for local coherence, global coherence, interactional signaling, and subtext interpretation.
-
The system can identify its own weakness profile (e.g., low-level entity tracking vs. high-level abstraction) and allocate reasoning depth accordingly, improving both weak models (which fail on N.1) and strong models (which fail on N.4).
What the Improved AI System Can Do:
-
Watch a 4-hour Japanese drama and correctly answer questions about a minor character's motivation shift across episodes 2 and 7, while also inferring that a polite refusal in a business meeting actually signals strong disagreement based on kuuki wo yomu.
-
Generate a coherent summary of a long-form video that preserves both the causal chain of events and the culturally implicit emotional arc, without degrading when asked to produce free-form text instead of selecting from options.
-
Automatically decide when to sample more frames (e.g., during a tense negotiation scene) versus fewer (e.g., during a static monologue), optimizing compute while maintaining high accuracy on both narrative and cultural questions.
-
Detect and reject questions that can be answered from subtitles alone, forcing the model to rely on visual and temporal cues, thus improving robustness to shortcut exploitation in real-world video understanding tasks.
Sources
- InternVL3.5: Advancing Open-Source Multimodal Models in Versatility, Reasoning, and Efficiency
- Qwen2.5-VL Technical Report
- VideoLLaMA 3: Frontier Multimodal Foundation Models for Image and Video Understanding
- From Seconds to Hours: Reviewing MultiModal Large Language Models on Comprehensive Long Video Understanding
- CinePile: A Long Video Question Answering Dataset and Benchmark
- NarrativeTrack: Evaluating Entity-Centric Reasoning for Narrative Understanding
- MIRAGE: The Illusion of Visual Understanding
- Qwen3-VL Technical Report
Related papers
- Loss Knows Best: Detecting Annotation Errors in Videos via Loss Trajectories
- AnchorWeave: World-Consistent Video Generation with Retrieved Local Spatial Memories
- Benchmarking the Robustness of Foundation Models for Mammography under Domain Shift
- MambaX-Net: Dual-Input Mamba-Enhanced Cross-Attention Network for Longitudinal MRI Segmentation
- TeleOCR: Navigating Document Parsing Across Digital and Camera-Captured Documents
- A Survey on Efficient Vision-Language-Action Models