AVA-Encoder: Towards Agent-Native Video Representation Learning

arXiv:2608.12313 · cs.CV, cs.CL · Submitted 2026-08-18 · Read on arXiv

Chuyue Li, Jinpeng Yu, Haozhe Wang, Tian Xueyun, Zhijing Zhang, Bingnan Li, Shuqi Gu, Kan Ren, Jiaming Liu, Ruihua Huang

Alibaba · ShanghaiTech University · The Hong Kong University of Science and Technology · Institute of Computing Technology · Southeast University

cs.CV, cs.CL

Submitted: 2026-08-18

Updated: 2026-08-19

Code: https://github.com/BroderQi/Storyboard

License: http://creativecommons.org/licenses/by-nc-nd/4.0/

Importance score: 100/100

The gist: AVA-Encoder: Towards Agent-Native Video Representation Learning Summary This paper introduces the Agentic Video Auto-Encoder (AVA-Encoder), a framework for learning agent-native video representations.

Terminology

Summary

AVA-Encoder: Towards Agent-Native Video Representation Learning

Summary

This paper introduces the Agentic Video Auto-Encoder (AVA-Encoder), a framework for learning agent-native video representations. The core problem addressed is that creative agents cannot effectively learn from high-quality human films due to a fundamental mismatch between film space and agent space. Films are tightly connected multimodal content, while agents operate through structured representations like text, code, plans, and graphs. The paper proposes a solution that transforms videos into a structured knowledge graph (KG) representation that is faithful to film content and directly usable for agentic reasoning and manipulation.

1. Proposed Representation: Film Knowledge Graph

The paper proposes a film-creation knowledge-graph (KG) representation to meet three requirements: it must be understandable to agents, easy for agents to reason over and edit, and faithful enough to preserve cinematic information for future generation. The representation is text-centered, with a Story–Event–Shot hierarchy and Character, Scene, Object, Style, Camera, and Audio states storing structured text. Generated images, audio, and video are kept in a linked asset layer. The graph uses typed edges to preserve relations between text descriptions and assets, allowing agents to easily understand, query, and edit the content. The representation is built through film-, shot-, and keyframe-level understanding, with each finer level using context from the level above it to reduce information loss.

2. AVA-Encoder Framework

AVA-Encoder is described as the first agentic video auto-encoding framework for self-evolving, cinematically faithful, agent-native video representation learning. It encodes an input film into the proposed KG and reconstructs the film from that representation. Reconstruction quality is treated as a direct measure of representational faithfulness. The framework addresses four major challenges with four components:

  • Multi-level Agentic Video Encoder: This component analyzes the film, its shots, and its keyframes in order, passing high-level context to each finer level to retain information needed for reconstruction.

  • Film Knowledge Graph: This component separates information into structured-text nodes and linked assets while using typed edges to preserve hierarchy, temporal order, and cross-shot dependencies.

  • Dual-Loop Textual-Gradient Framework: This component combines Data-Independent Encoding Policy Pseudo-Training, which improves the shared encoding policy before deployment, with Data-Dependent KG Representation Refinement, which improves the current video’s KG representation at test time.

  • Reconstruction-Error Design: This component uses the loop-facing reward to diagnose reconstruction failures and verify revisions, while a separate evaluation metric measures all representation systems in four reconstruction directions and eight film dimensions.

3. Dual-Loop Textual-Gradient Evolution

The framework improves its shared encoding policy and input-specific KG representations in two separate textual-gradient stages:

  • Outer Loop (Data-Independent Encoding Policy Pseudo-Training): This stage learns the shared shot-level encoding policy across a collection of videos before deployment. It uses a reconstruction-derived reward to evaluate candidate policies and an Anti-Forgetting Gate to prevent performance regression on previously seen data.

  • Inner Loop (Data-Dependent KG Representation Refinement): This optional stage refines only the input-specific representation of the current video at test time. It uses an Anti-Degradation Gate to filter out sampling variance and ensure off-target stability.

4. Key Results

The paper reports extensive experiments demonstrating the effectiveness of AVA-Encoder:

  • Reconstruction Fidelity (RQ1): AVA-Encoder achieves an Overall reconstruction score of 49.0%, a 20.7-percentage-point absolute improvement over the strongest external baseline (soap2soap at 28.3%). It outperforms all baselines in all four comparison directions: Video (57.8% vs. 36.7%), Keyframe (73.7% vs. 39.5%), Video Back-Captioning (29.7% vs. 15.8%), and Keyframe Back-Captioning (34.6% vs. 21.3%).

  • Optimization Stage Contributions (RQ2): The two optimization stages provide separate gains and perform best together. The complete configuration (49.0%) outperforms configurations with only policy pseudo-training (45.8%), only KG refinement (45.4%), and neither (42.4%). The pseudo-trained shot-level policy also outperforms a carefully human-tuned policy (45.8% vs. 44.4%) while using 74.3% fewer system-prompt tokens (8,052 vs. 31,336).

  • Knowledge-Graph Operability (RQ3): The graph supports controlled edits that remain consistent across linked shots, including identity replacement, visual-treatment replacement, plot modification, and fine-grained adjustment of cinematographic language.

  • Downstream Reuse (RQ4): The same representation improves every tested downstream generation system (MovieAgent, FilmAgent, Anim-Director, VideoStudio) when supplied as a single textual input.

5. Contributions

The paper's contributions are threefold:

  1. Agentic Video Auto-Encoder: Introduces the paradigm and releases the complete AVA-Encoder framework.

  2. Agentic Video Representation Benchmark: Establishes the first benchmark for evaluating agentic video representations through reconstruction faithfulness, with automatic metrics agreeing with human judgments on 710 of 730 blinded triples (97.3%).

  3. Film Knowledge Graph Dataset and Editing Framework: Releases the first dataset of high-quality film knowledge-graph representations together with a graph-based editing framework.

Improvements for AI systems

Improvements to AI systems:

  1. Agent-native video understanding module: Add a multi-level video encoder that decomposes input video into Story–Event–Shot hierarchy, extracting structured text for characters, scenes, objects, styles, camera, and audio at each level. The improved system can parse any video into an editable, queryable knowledge graph, enabling agents to reason about film content as structured data rather than raw pixels.

  2. Self-evolving encoding policy optimizer: Implement a dual-loop textual-gradient framework where the AI system pseudo-trains its shared encoding policy on a corpus of videos (outer loop) and refines per-video representations at test time (inner loop), using reconstruction error as reward and anti-forgetting/anti-degradation gates for stability. The improved system continuously improves its video-to-graph conversion accuracy without human retuning, adapting to new film styles and genres automatically.

  3. Reconstruction-faithful representation validator: Add a reconstruction-error design that measures how well the system can regenerate the original video (and its keyframes, captions, and back-captions) from the knowledge graph. The improved system can self-diagnose information loss in its representations, verify whether edits preserve cinematic fidelity, and flag when a representation is insufficient for downstream generation tasks.

  4. Graph-based film editing engine: Integrate the typed-edge knowledge graph with an editing framework that supports identity replacement, visual-treatment changes, plot modification, and cinematographic adjustments. The improved system can perform controlled edits on a video’s semantic structure (e.g., change a character’s appearance, alter lighting style, or modify a scene’s plot point) and propagate those changes consistently across all linked shots, keyframes, and assets.

  5. Cross-system representation injection layer: Build an interface that converts the film knowledge graph into a single textual input compatible with multiple downstream generation systems (e.g., MovieAgent, FilmAgent, Anim-Director, VideoStudio). The improved system can act as a universal video-understanding front-end, boosting the performance of any existing video-generation or agentic filmmaking tool by providing it with a richer, structured understanding of source content.

  6. Agentic video representation benchmark evaluator: Implement an automatic evaluation suite that scores any representation system across four reconstruction directions (video, keyframe, video back-captioning, keyframe back-captioning) and eight film dimensions, with metrics aligned to human judgment (97.3% agreement). The improved system can objectively compare different video-understanding approaches, track representation quality over time, and guide further optimization without expensive human evaluation.

Abstract

Creative agents still lack an effective way to learn from high-quality human films, limiting their ability to produce cinematic-grade videos. A key challenge is the absence of a structured video representation that is both faithful to film content and directly usable for agentic reasoning and manipulation. To address the challenge, we propose the Agentic Video Auto-Encoder (AVA-Encoder), a framework for learning agent-native video representations via agentic auto-encoding. AVA-Encoder transforms a video into a knowledge graph (KG) representation and then reconstructs it back into video. Its hierarchy and state nodes store structured text, while a linked asset layer holds generated images, audio, and video. Typed edges preserve the relations between these text descriptions and assets in a form that agents can easily understand, query, and edit. The video reconstruction differences drive a textual-gradient optimization framework, which expresses evaluation feedback as natural-language update directions for Data-Independent Encoding Policy Pseudo-Training in the outer loop and optional Data-Dependent KG Representation Refinement in the test-time inner loop. Extensive experiments show that AVA-Encoder improves by 20.7 percentage points over the strongest external baseline. In the controlled policy-only setting, its pseudo-trained shot-level Agentic Video Encoder policy also outperforms a carefully human-tuned policy while using 74.3% fewer system-prompt tokens. We release the complete AVA-Encoder framework, a reliable agentic video reconstruction benchmark, and the first dataset of high-quality film KG representations.

Sources

Related papers