EverAnimate: Minute-Scale Human Animation via Latent Flow Restoration

arXiv:2605.15042 · cs.CV, cs.AI · Submitted 2026-05-14 · Read on arXiv

Listen

Radio episode about this paper

Transcript

Introduction to the show: ident: AI Radio. Generated commentary on the latest Artificial Intelligence papers.

Tom: I'm Tom, and with me are Jane, Lu, senior AI researcher at Tsinghua, Meng, lead engineer at a mysterious AI startup and Lalam, the in-house Large Language Model.

Jane: Today's paper: "EverAnimate: Minute-Scale Human Animation via Latent Flow Restoration".

Tom: EverAnimate is an efficient post-training method designed for long-horizon animated video generation that preserves visual quality and character identity by restoring drifted flow trajectories through persistent latent context memory and…

Jane: First, who's behind it and why it matters.

Paper summary: Tom: We’ve covered a lot about EverAnimate: Minute-Scale Human Animation via Latent Flow Restoration. To wrap things up, we looked at how this method tackles the two major forms of drift in long-form animation and the mechanisms it uses to maintain stability.

Jane: It really boils down to using persistent latent context memory for identity and motion propagation, combined with Restorative Flow Matching for implicit restoration during sampling. The authors are showing that this combination significantly improves stability over time compared to methods that rely on simpler anti-drifting techniques.

Lu: From a research perspective, the way they structure the memory—separating short-term motion continuity from the global identity memory—is a very smart architectural decision that addresses the core challenge of motion heterogeneity in long sequences.

Meng: While I’m still thinking about deployment and real-world robustness to unexpected changes, the reported performance gains across those longer horizons are undeniable evidence that this approach is effective at maintaining quality over extended generation lengths.

Lalam: For our culture initiatives, the implication is that we can move toward creating AI tools capable of supporting much longer, more complex artistic and narrative projects where consistency isn't just about a few seconds, but about an entire sustained sequence.

Tom: So, in simple terms, EverAnimate is an efficient post-training method that keeps the visual quality and character identity steady while generating very long animated videos by anchoring the generation process to a persistent context memory.

Jane: That’s the essence of it, Tom. It focuses on solving those issues of low-level quality drift and high-level semantic drift that plague current chunk-based generation methods.

Lu: The paper does show its limitations, though they point out that they rely on the model being adapted to this memory condition during training, which means the method’s success is tied to how well it handles the initial adaptation phase.

Meng: So, while it’s a strong technical advancement for fidelity over long sequences, we’ll need to see how easily this can be integrated into existing pipelines without requiring massive retraining cycles every time we want to generate something new.

Lalam: I think the main impact is enabling richer AI-assisted creative works where the visual consistency allows creators to focus more on the narrative and less on worrying about the video quality degrading over time.

Conclusion: Tom: So, we've been looking at EverAnimate, which is all about using persistent memory and flow restoration to handle those tricky long-form animation problems.

Jane: Exactly, Tom; it really tackles the issue of keeping a character looking consistent across an entire video without that drift we see in other methods.

Lu: The authors have put together a sophisticated system where they decouple the motion memory from the global identity memory, which is fascinating because it treats motion and identity as separate but connected flows in latent space.

Meng: From an engineering standpoint, I'm really interested in how they manage that propagation across chunks so we can actually build something scalable for real-world applications.

Lalam: I see the potential here for narrative generation that lasts much longer than what we’ve seen before, allowing us to create truly immersive and sustained virtual worlds.

Tom: It sounds like the title, EverAnimate: Minute-Scale Human Animation via Latent Flow Restoration, really sums up this idea of maintaining high quality over a long sequence through flow correction.

Jane: That title is perfect because it tells us exactly what the core mechanism is: restoring the flow while keeping it minute by minute consistent.

Lu: What I find compelling is that they use a velocity adjustment technique within Restorative Flow Matching, which allows the model to correct drift implicitly without having to mess with the conditional images directly.

Meng: That implicit correction is what I need to hear; if it does that without explicit perturbations, it suggests a much more stable and less brittle generation process in practice.

Lalam: If we can achieve minute-scale consistency reliably, imagine the cultural impact on interactive storytelling where character relationships and physical appearances remain perfectly stable throughout an extended narrative.

Tom: And those stability gains are what they report: significant improvements in both visual quality and character identity across very long rollouts, which is a huge deal for the industry.

Jane: So, essentially, EverAnimate shows a way to solve the drift problem by giving the generation process an internal feedback loop that fixes errors as it goes.

Lu: The architectural separation of motion memory and identity memory is what I think opens up new avenues for manipulating character traits or environmental details in future iterations.

Meng: I'm curious about the training phase; how difficult was it to get the model to learn that specific velocity adjustment for the Restorative Flow Matching objective?

Lalam: That level of control over temporal consistency could fundamentally improve how we design our large language models for video, allowing them to maintain complex, long-term contextual understanding.

Tom: It really shows that anchoring generation through persistent memory isn't just a trick; it’s a robust way to handle the inherent instability of synthesizing human motion against static scenes.

Jane: And the authors’ conclusion emphasizes that this approach provides stable quality in the background and character identity without noticeable artifacts, which is exactly what we want to see.

Lu: The paper lays out a solid foundation for how we might approach multi-modal generation where temporal coherence is paramount across very long sequences.

Meng: We'll have to see if the training adaptation stage adds too much complexity for our current infrastructure, but the potential payoff seems worth exploring further.

Lalam: This work could profoundly influence how AI generates cinematic content, moving us toward a future where we can produce incredibly detailed and visually consistent animated narratives at scale.

Tom: That's a lot to chew on for our listeners; we’ll be diving into the specifics of those results right after this.

VITA@EPFL

cs.CV, cs.AI

Submitted: 2026-05-14

Updated: 2026-09-28

Comments: NeurIPS 2026; Project Page: https://everanimate.github.io/homepage/

Code: https://github.com/aigc-apps/VideoX-Fun

Project page: https://everanimate.github.io/homepage/Abstract

License: http://creativecommons.org/licenses/by/4.0/

Importance score: 83/100

The gist: EverAnimate is an efficient post-training method designed for long-horizon animated video generation that preserves visual quality and character identity by restoring drifted flow trajectories

Key concepts

Persistent Latent Propagation
This mechanism maintains consistency across generated video segments by storing two types of memory: short-term motion memory tracking recent frames, and a global identity memory encoding the character's look. This context is injected into the generation process to ensure smooth continuity between different parts of the animation.
Restorative Flow Matching (RFM)
This technique adds an implicit correction step during sampling. Instead of just following a path, RFM simulates endpoint drift by calculating a unique constant velocity needed to transport a perturbed state back to the clean data manifold. This actively corrects visual errors and improves fidelity within each generated chunk.
Drift Mitigation
Long animations suffer from drift where details degrade over time due to repeated reconstruction. EverAnimate tackles this by using memory propagation for high-level semantic consistency (identity) and RFM for low-level visual correction, ensuring that the character remains consistent and the background quality stays stable throughout the entire video.

Terminology

Summary

EverAnimate is an efficient post-training method designed for long-horizon animated video generation that preserves visual quality and character identity by restoring drifted flow trajectories through persistent latent context memory and intrinsic restoration capabilities.

The gist

EverAnimate restores drifted flow trajectories by anchoring generation to a persistent latent context memory, consisting of two complementary mechanisms: Persistent Latent Propagation maintains a context memory across chunks to propagate identity and motion in latent space while mitigating temporal forgetting, and Restorative Flow Matching introduces an implicit restoration objective during sampling through velocity adjustment, improving within-chunk fidelity.

Problem Analysis

Long-form animation is challenging because highly dynamic human motion must be synthesized against relatively static environments, making chunk-based generation prone to accumulated drift. This drift manifests in two primary forms: (i) low-level quality drift, such as progressive degradation of static backgrounds due to repeated latent-to-pixel reconstruction during cross-chunk propagation; and (ii) high-level semantic drift, such as inconsistent character identity and view-dependent attributes due to limited semantic memory that acts only as a positive signal without correcting drift. Empirical analysis revealed that image-space continuation is fundamentally ill-suited for long-horizon animation, and attention sinks alone are insufficient for reliably anchoring long horizons. The core challenge lies in the motion heterogeneity between the rapidly evolving human motion and the comparatively stable scene, necessitating a latent-space principle where semantic memory is propagated autoregressively across chunks, while intrinsic restoration corrects within-chunk drift.

Persistent Latent Propagation

This mechanism maintains semantic consistency across generated chunks via multi-view latent memory, thereby avoiding repeated destructive reconstruction and strengthening cross-chunk continuity. Memory construction involves extracting two components from the context chunk V(1): a motion memory Mmot that preserves short-term temporal continuity across adjacent chunks (keeping the last 'r' latent slices), and a global identity memory Mid encoded by sampling frames: Mid = n E Tid(I(1) k) o K k=1. To solve context bias, a simple augmentation called Tid applies mild identity-preserving spatial augmentation, e.g., random translation and rescaling, in training to prevent spatial biases of memory context. The memories are then injected into the DiT input by concatenating them into Mctx: Mctx = Concatt(Mmot, Mid, Xpad). This context token is then concatenated with the pose-injected target latent to form the final DiT input H(2)t.

Restorative Flow Matching

This component enables a built-in restorative ability to actively correct emerging drift implicitly without explicitly perturbing conditional images, thereby improving within-chunk visual fidelity. The method first establishes a standard Flow Matching (FM) objective, LFM = E vθ H(2)t, t C(2), which trains the vector field to transport Gaussian noise toward the clean data manifold under the memory/control pathway. To address drift, EverAnimate introduces Restorative Flow Matching (RFM) with Velocity Adjustment. Instead of perturbing the transmitted context, it simulates endpoint drift by defining a perturbed state Ve and then asking for a unique constant velocity that transports this perturbed state to the clean endpoint X1 over the remaining interval [t, 1]. The exact coefficient for this velocity is derived as Uet,exact = X1 − Xet / (1 − t).

Training and Inference

Training is conducted in two stages: (i) Memory adaptation, where the model is adapted to the memory condition by perturbing motion memory and optimizing with standard FM loss LFM; and (ii) Anti-drift adaptation, where the model is optimized with Restorative FM loss LRFM to improve long-range stability. The LRFM objective replaces the standard flow matching velocity with a corrected velocity Uet: Uet = Ut + λ(t) X t − X et, where λ(t) follows a bounded bell-shaped schedule to avoid overconstraining the clean endpoint. During inference, the method reuses the video latent to guide subsequent chunk generation without autoregressively decoding and encoding frames between chunks, ensuring that the last r latent slices from the previous chunk are propagated as short-term motion memory, while identity memory remains fixed and shared across chunks. This process allows for flexible user input of reference frames to specify the target identity.

Main Results

EverAnimate consistently achieves the best performance across rollout horizons compared to state-of-the-art methods. At 10 seconds, it improves PSNR/SSIM by 8%/7% and reduces LPIPS/FID by 22%/11%; at 90 seconds, these gains increase to 15%/15% and 32%/27%, respectively. Qualitative comparison shows that while most models deteriorate over time, EverAnimate can maintain stable quality in the background and human identity without obvious artifacts, demonstrating robustness for long-range generation.

Improvements for AI systems

Here are specific improvements for AI systems based on EverAnimate, categorized by the technical capabilities they enable:


The improved AI system (EverAnimate) is a lightweight post-training framework designed for long-horizon pose-guided human animation, specifically targeting minute-scale video generation while maintaining high visual quality and character identity.

Here are the specific improvements and capabilities:

I am ready to proceed with the next step or further analysis of the paper if you have a specific task in mind (e.g., proposing new architectures, analyzing limitations, or simulating experimental results).

Abstract

We propose EverAnimate, an efficient post-training method for long-horizon animated video generation that preserves visual quality and character identity. Long-form animation remains challenging because highly dynamic human motion must be synthesized against relatively static environments, making chunk-based generation prone to accumulated drift: (i) low-level quality drift, such as progressive degradation of static backgrounds, and (ii) high-level semantic drift, such as inconsistent character identity and view-dependent attributes. To address this issue, EverAnimate restores drifted flow trajectories by anchoring generation to a persistent latent context memory, consisting of two complementary mechanisms. (i) Persistent Latent Propagation maintains a context memory across chunks to propagate identity and motion in latent space while mitigating temporal forgetting. (ii) Restorative Flow Matching introduces an implicit restoration objective during sampling through velocity adjustment, improving within-chunk fidelity. With only lightweight LoRA tuning, EverAnimate outperforms state-of-the-art long-animation methods in both short- and long-horizon settings: at 10 seconds, it improves PSNR/SSIM by 8%/7% and reduces LPIPS/FID by 22%/11%; at 90 seconds, the gains increase to 15%/15% and 32%/27%, respectively.

Sources

Related papers