FrescoDiffusion: 4K Image-to-Video with Prior-Regularized Tiled Diffusion
Listen
Radio episode about this paper
Transcript
Introduction to the show: ident: AI Radio. Generated commentary on the latest Artificial Intelligence papers.
Tom: Next we'll be talking about the paper "FrescoDiffusion: 4K Image-to-Video with Prior-Regularized Tiled Diffusion".
Jane: The paper was written by Hugo Caselles-Dupré, Mathis Koroglu, Guillaume Jeanneret, Arnaud Dapogny and Matthieu Cord from Obvious Research and Sorbonne University.
Tom: Stay tuned as we take you through the paper and discuss its implications.
Jane: We also have Lu with us today — senior AI researcher at Tsinghua.
Tom: We also have Meng with us today — lead engineer at a mysterious AI startup.
Jane: We also have Lalam with us today — the in-house Large Language Model.
Tom: Alright, let's get started.
Title: Tom: Welcome back to the show, everyone! Today we’re looking at a paper that’s got a mouthful of a title: “FrescoDiffusion: 4K Image-to-Video with Prior-Regularized Tiled Diffusion.” Jane, I gotta say, just reading that title makes me think of Renaissance churches, not computer science.
Jane: Ha, exactly, Tom! And that’s actually pretty fitting. The whole idea here is taking these massive, complex images—like a giant fresco painting with dozens of little scenes in it—and animating them in 4K resolution. We’re not talking about a simple video of a cat walking; we’re talking about bringing an entire wall of art to life.
Tom: Right, and the authors are from Obvious Research and Sorbonne University in Paris. Caselles-Dupré, Koroglu, Jeanneret, Dapogny, and Cord. They’re basically saying, "Hey, current video models are great, but they choke when you give them a 4K image."
Jane: Exactly. Most image-to-video models work at a fixed resolution, like 480p or 1080p. If you just shrink your 4K fresco down to that size, you lose all the fine details—the little faces, the intricate patterns. But if you try to run the model on the full 4K image, it runs out of memory or produces a mess.
Tom: So they came up with a clever workaround. Instead of forcing the model to see the whole thing at once, they chop the image into overlapping tiles, process each tile, and then stitch them back together. That’s the "tiled diffusion" part.
Jane: But here’s the catch, Tom. If you just tile it naively, you get seams between the tiles, and the overall scene doesn’t stay consistent. Things drift, colors shift, and the whole thing looks like a patchwork quilt.
Tom: So what’s the "prior-regularized" part? That sounds like the secret sauce.
Jane: It is. They first generate a low-resolution video of the whole image. That gives them a "prior"—a rough idea of what the motion and the overall structure should look like. Then, while they’re denoising the high-res tiles, they use that low-res video as a guide to keep everything on track.
Tom: So it’s like having a sketch artist draw the rough outline, and then a team of detail painters fill in the fine brushstrokes while constantly checking back with the sketch to make sure they’re not drawing a different picture.
Jane: That’s exactly it. And the beauty is, it’s "training-free." They don’t have to retrain the video model. They just use it cleverly at inference time.
Tom: That’s huge for practicality. Now, they’re calling this "fresco animation," which is a really specific use case. But I’m guessing the implications go way beyond art history.
Jane: Oh, absolutely. Think about any large-format visual—movie posters, billboards, complex infographics, even video game concept art. If you can animate a 4K fresco, you can animate any high-res image with multiple distinct regions.
Tom: And that’s what I want to dig into next. How exactly do they make this work without breaking the bank on compute? Let’s get into the nitty-gritty of the method.
Summary: Tom: So, we’re back with “FrescoDiffusion: 4K Image-to-Video with Prior-Regularized Tiled Diffusion.” Jane, we talked about the big picture, but let’s get into the actual mechanics. How does this thing work under the hood?
Jane: Okay, so picture this. You have your giant 4K image. First, they shrink it down to the model’s native size, like four hundred eighty by eight hundred thirty-two pixels. They run the standard image-to-video model on that small version to get a short, low-res video. That’s their global prior.
Tom: And that prior is basically a map of where things are moving and how the scene is structured over time.
Jane: Right. Then they take that low-res video and upscale its latent representation—the compressed version the model works with—back up to the full 4K canvas size. So now they have a fuzzy, low-detail version of the whole video at 4K.
Tom: So it’s a guide, not the final product.
Jane: Exactly. Now, for the actual generation, they start with random noise at 4K and denoise it tile by tile. Each tile is processed by the model, producing a prediction of the velocity field—that’s the direction and speed each pixel should move to get closer to the final image.
Tom: And this is where the "prior-regularized" part kicks in. They don’t just average the tile predictions like the old MultiDiffusion method. They add a term that pulls the prediction toward that upscaled prior.
Jane: Precisely. They set up a loss function with two parts. One part makes sure the tiles agree with each other where they overlap. The other part makes sure the current denoising step doesn’t drift too far from the low-res prior. And they solve this in closed form—meaning they have a direct formula for the optimal fusion, no iterative optimization needed.
Tom: So it’s fast. But they also have this clever trick with a parameter called lambda. It controls how strongly the prior pulls the generation.
Jane: Right. And they don’t keep lambda constant. Early in the denoising process, lambda is high, so the generation sticks close to the prior and gets the global structure right. Later, they lower it, allowing the model to add fine details that weren’t in the low-res version.
Tom: So it’s like a painter first blocking in the big shapes, then stepping back and adding the tiny brushstrokes.
Jane: Exactly. And they even have a regional version. For a fresco, you might want the background to stay static but the people in the foreground to move. So they use a mask to apply a strong prior to the background and a weaker prior to the foreground.
Tom: That’s really smart. It gives you control over where motion happens. So they’re not just making a video; they’re making a video that respects the artistic intent.
Jane: And that’s what makes it so powerful. They’re not just upscaling; they’re creating new detail that’s consistent with the whole scene. Let’s bring in Lu from Tsinghua to talk about what this means for the field.
Lu: Thanks, Jane. I’m really excited about this. The closed-form solution is elegant, but the real breakthrough is the idea of using a low-res prior to guide high-res tiled denoising. It’s a general principle that could apply to any diffusion model, not just video.
Tom: So you’re saying this could be a template for other high-res generation tasks?
Lu: Absolutely. Think about text-to-image at 8K, or even three dee scene generation. Anywhere you have a global structure that needs to be preserved while adding local detail, this prior-regularization approach could be the key.
Jane: And the fact that it’s training-free means it can be dropped into existing pipelines without retraining. That’s a huge practical advantage.
Lu: It is. And the results they show are impressive. They beat the baselines on user preference and on standard metrics, and they do it faster. That’s a rare combination.
Tom: Faster and better? That’s the dream. But I’m curious about the practical side. Meng, you’re the engineer here. What do you think about actually running this thing?
Improvements: Tom: So we’re deep into “FrescoDiffusion: 4K Image-to-Video with Prior-Regularized Tiled Diffusion.” Jane, we’ve covered the method, but I want to hear from Meng about the practical improvements. This paper claims to be not just better quality, but also faster.
Meng: Yeah, Tom, that’s the part that caught my eye. They’re using a fourteen-billion-parameter model, Wan2 point 2, and they’re generating eighty-one frames at 4K. That’s a lot of compute. But they’re smart about it.
Jane: They mention using a TurboDiffusion LoRA to cut the number of denoising steps down to six. That’s a massive speedup compared to the usual fifty or so steps.
Meng: Exactly. And they’re using FP8 precision and torch.compile to squeeze out more performance. On a single H100 GPU, they’re generating a video in about eight and a half minutes. That’s not real-time, but it’s practical for a creative workflow.
Tom: And how does that compare to the baselines?
Meng: Their method is about forty-five percent faster than DemoFusion and significantly faster than the original DynamicScaler. The only thing faster is plain MultiDiffusion, but that produces worse quality. So they’re hitting a sweet spot.
Jane: And the quality improvements aren’t just subjective. They ran a user study with over one thousand three hundred ratings. People strongly preferred their method over DynamicScaler and MultiDiffusion, with preference rates in the 80s and 90s.
Meng: Right. And against DemoFusion, their regional variant—R-FrescoDiffusion—won with a statistically significant sixty-nine percent preference. That’s a clear win.
Tom: So what’s the catch? What are the limitations?
Meng: The main one is that it relies on the low-res prior being meaningful. If the image is so huge that the thumbnail loses too much structure, the prior won’t help. And tiled denoising is still inherently compute-heavy. They’ve optimized it, but it’s not cheap.
Jane: But they’ve also shown a really cool control feature. You can dial lambda up or down to trade off between creativity and fidelity to the prior. It’s a Pareto frontier—you can’t have both maximum creativity and maximum fidelity, but you can choose where you want to be.
Lu: And that’s a big deal for artists. You can set a high lambda to get a faithful animation of the original artwork, or a low lambda to let the model hallucinate new details and motion. It gives the creator control over the output.
Meng: Exactly. And the regional version takes that further. You can say, "Keep the background static, but let the characters move." That’s a level of control that most video generation tools don’t offer.
Tom: So it’s not just a black box that spits out a video. It’s a tool that respects the creator’s vision.
Jane: And that’s what makes it so exciting. It’s not just about making videos; it’s about making the right video. Let’s bring in Lalam to talk about the cultural implications of this.
Lalam: Thank you, Jane. I think the most profound impact here is on cultural heritage. Frescoes are some of the most important artworks in human history, but they’re static. This technology allows us to imagine them as living scenes, to see the stories they depict unfold in motion.
Tom: That’s a beautiful way to put it. So this isn’t just for entertainment; it’s for education and preservation.
Lalam: Exactly. Museums could use this to create immersive experiences. Students could see a historical battle scene come to life. And because the method is training-free, it can be applied to any existing artwork without needing to train a new model for each piece.
Jane: And the control over the prior strength means curators can decide how much creative liberty to take. They can keep it faithful to the original or allow more interpretive motion.
Lalam: That’s the key. It’s a tool for storytelling, not just generation. It respects the source material while allowing for new narratives.
Tom: Alright, so we’ve got a method that’s faster, better, and more controllable. What’s not to love? Let’s wrap this up in the next segment.
Conclusion: Tom: Alright, we’re wrapping up our discussion on “FrescoDiffusion: 4K Image-to-Video with Prior-Regularized Tiled Diffusion.” Jane, give us the final summary.
Jane: Sure, Tom. This paper tackles the problem of animating ultra-high-resolution images, specifically 4K frescoes with many distinct scenes. The key innovation is using a low-resolution video as a global prior to guide tiled denoising at full resolution. This keeps the overall scene coherent while allowing fine details to emerge.
Tom: And they do it in a training-free way, which is huge. They just use the existing video model cleverly.
Jane: Right. They also introduce a controllable lambda parameter that lets you balance between fidelity to the original image and creative freedom. And the regional variant lets you control where motion happens, keeping backgrounds static while animating foreground elements.
Meng: And it’s efficient. On a single H100, you get a 4K video in about eight and a half minutes, which is faster than most comparable methods.
Lu: The implications go beyond frescoes. This prior-regularized tiling could be a template for any high-resolution generation task, from images to three dee scenes.
Lalam: And culturally, it opens up new ways to experience and preserve our artistic heritage, making it accessible and alive.
Tom: Well said, everyone. It’s a paper that combines technical elegance with practical impact and a touch of artistic wonder. We’ll be keeping an eye on where this goes.
Jane: And with that, we’re saying goodbye to “FrescoDiffusion.” Next up, we’ve got a paper on something completely different—stay tuned.
Tom: Thanks for listening, folks. See you next time.
Hugo Caselles-Dupré, Mathis Koroglu, Guillaume Jeanneret, Arnaud Dapogny, Matthieu Cord
Obvious Research · Sorbonne University
cs.CV, cs.AI
Submitted: 2026-08-15
Updated: 2026-08-18
Comments: 5 authors. Hugo Caselles-Dupr\'e, Mathis Koroglu, and Guillaume Jeanneret contributed equally. 14 pages, 7 figures
Code: https://github.com/genmoai/models
Project page: https://f2v.pages.dev
License: http://creativecommons.org/licenses/by/4.0/
Importance score: 68/100
The gist: large-scale images containing multiple distinct scenes, characters, and objects that must remain spatially coherent over time.
Key concepts
- Prior-Regularized Tiled Diffusion
- This technique involves chopping a high-resolution image into tiles and processing them individually. A low-resolution video of the whole image acts as a 'prior' to guide the denoising process for each tile, ensuring overall scene consistency and preventing seams between tiles.
- Prior
- The prior is a rough, low-resolution video generated from the full image. It provides a structural guide for the high-resolution generation process. This allows the model to maintain the correct overall motion and scene structure while adding fine details.
- Lambda Parameter
- Lambda is a control parameter that dictates how strongly the prior influences the generation. A high lambda keeps results close to the original image's structure, while a lower lambda allows for more creative detail and hallucination.
Terminology
Summary
Summary
The paper introduces FrescoDiffusion, a training-free method for generating 4K (ultra-high-definition) videos from a single complex input image, specifically targeting the fresco
setting—large-scale images containing multiple distinct scenes, characters, and objects that must remain spatially coherent over time. The core problem addressed is that existing image-to-video (I2V) diffusion models operate at native resolutions (typically 480–1080p) and fail to scale to 4K inputs: resizing the input to the model's native resolution loses fine-grained structure, while high-resolution tiled denoising preserves local detail but breaks global layout consistency.
The key idea is to augment tiled denoising with a precomputed latent prior. The method first generates a low-resolution video at the underlying model's native resolution from a resized thumbnail of the input image, then upsamples its latent trajectory to obtain a global reference capturing long-range temporal and spatial structure. For 4K generation, per-tile noise predictions are computed and fused with this reference at every diffusion timestep by minimizing a single weighted least-squares objective in model-output space.
The proposed loss function combines a standard tile-merging criterion (MultiDiffusion) with a novel regularization term. The FrescoDiffusion loss is defined as:
l FD(y⋆; t) = ‖√λ ⊙ [(x t 4K − σ t y⋆) − x prior]‖22 + l MD(y⋆; t)
where x t 4K is the large latent canvas, x prior is the upsampled prior latent, σ t is the scheduler's noise standard deviation, λ is a regularization variable, and l MD is the MultiDiffusion tile-merging loss. This loss is separable and strictly convex, admitting a closed-form solution for the optimal fused velocity:
y FD(x t 4K) = [σ t·λ ⊙ (x t 4K − x prior) + Σi wi ⊙ yi(x t 4K)] / [σ t2·λ + Σi wi]
When λ = 0, this reduces to standard MultiDiffusion fusion.
The paper introduces two variants of the prior strength schedule. The first, FrescoDiffusion, uses a global gated decreasing schedule: λ G(t, τ) = λ base · cos(t·π/2) · 1[t ≤ τ], where τ is a temporal gating and λ base is the strength. This schedule applies high prior strength early in diffusion to maintain global coherence, then relaxes it in later steps to allow fine detail generation. The second variant, Regional-FrescoDiffusion (R-FrescoDiffusion), uses a spatial activity map A(p) to differentiate active regions (e.g., characters, moving objects) from static background. The prior strength becomes a tensor λ R(t, p) that applies different temporal cutoffs τ act and τ bg for foreground and background regions respectively, allowing motion in active zones while keeping background structure stable.
The paper also introduces a new dataset, FrescoArchive, for evaluating fresco-to-video generation. Starting from LAION-2B Aesthetic Subset, images were filtered based on pixel count, aesthetics, watermark/NSFW scores, then semantically filtered using PerceptionEncoder similarity, classified as frescoes using Intern-VL-3.5, deduplicated, captioned with Qwen3-VL-32B, and manually curated to 371 image-caption pairs. The dataset contains images with an average width of 2265 pixels and height of 1552 pixels, with captions averaging 355.79 words.
Experiments were conducted using Wan2.2-I2V 14B as the backbone with TurboDiffusion LoRA for 6-step accelerated sampling. Baselines included MultiDiffusion, DemoFusion (adapted to video), and DynamicScaler (both original and re-implemented with the Wan backbone). Evaluation used both standard low-resolution metrics (VBench and VBench-I2V) and high-resolution metrics including Tenengrad sharpness, temporal consistency via mean squared error between consecutive downscaled frames, and prior alignment using DINOv3 CLS token cosine similarity.
Results show that FrescoDiffusion outperforms baselines in both quantitative metrics and user preference studies. In a user study with 1344 ratings from 47 participants, FrescoDiffusion achieved 84–93% preference rates over DynamicScaler and MultiDiffusion, and R-FrescoDiffusion achieved a statistically significant 69% preference over DemoFusion. R-FrescoDiffusion was preferred over FrescoDiffusion in 58% of comparisons. On standard VBench metrics, FrescoDiffusion achieved an average score of 0.904 on FrescoArchive and 0.875 on VBench-I2V, outperforming all baselines on FrescoArchive and being competitive with DemoFusion on VBench-I2V. The methods also demonstrated computational efficiency, with FrescoDiffusion taking 8.58 minutes and R-FrescoDiffusion 9.08 minutes per video on a single H100 GPU, compared to 10.25 minutes for DynamicScaler and 13.5 minutes for DemoFusion.
Ablation studies on the prior strength schedule show that the full design—including the cosine schedule, temporal gating, and spatial regularization—provides the best performance. The paper also demonstrates a Pareto frontier between creativity and prior similarity controlled by λ base and τ, allowing users to navigate the trade-off between preserving temporal coherence and maintaining high image sharpness.
Improvements for AI systems
Based on the paper, here are the specific improvements I can implement and the resulting capabilities of the improved AI system:
Improvement: I will integrate the prior-regularized tiled diffusion framework (FrescoDiffusion) into the video generation pipeline. This involves:
-
Generating a low-resolution prior video at the model's native resolution (e.g., 480×832) from the input 4K image.
-
Upsampling the prior's latent trajectory to the full 4K canvas size.
-
During each denoising step, computing per-tile noise predictions and fusing them with the prior using the closed-form solution:
y FD = (σ t·λ⊙(x t 4K − x prior) + Σ w i⊙y i) / (σ t2·λ + Σ w i).
Capability: The system can now animate a single 4K (3840×2160 or larger) image into a temporally coherent video at the same resolution, preserving fine-grained local details (e.g., individual brushstrokes, small characters) that would be lost with naive resizing, while maintaining global scene layout and motion consistency.
Abstract
Diffusion-based image-to-video (I2V) models are increasingly effective, yet they struggle to scale to ultra-high-resolution inputs (e.g., 4K). Generating videos at the model's native resolution often loses fine-grained structure, whereas high-resolution tiled denoising preserves local detail but breaks global layout consistency. This failure mode is particularly severe in the fresco animation setting: monumental artworks containing many distinct characters, objects, and semantically different sub-scenes that must remain spatially coherent over time. We introduce FrescoDiffusion, a training-free method for coherent large-format I2V generation from a single complex image. The key idea is to augment tiled denoising with a precomputed latent prior: we first generate a low-resolution video at the underlying model resolution and upsample its latent trajectory to obtain a global reference that captures long-range temporal and spatial structure. For 4K generation, we compute per-tile noise predictions and fuse them with this reference at every diffusion timestep by minimizing a single weighted least-squares objective in model-output space. The objective combines a standard tile-merging criterion with our regularization term, yielding a closed-form fusion update that strengthens global coherence while retaining fine detail. We additionally provide a spatial regularization variable that enables region-level control over where motion is allowed. Experiments on the VBench-I2V dataset and our proposed fresco I2V dataset show improved global consistency and fidelity over tiled baselines, while being computationally efficient. Our regularization enables explicit controllability of the trade-off between creativity and consistency.
Sources
- Qwen3-VL Technical Report
- Stable Video Diffusion: Scaling Latent Video Diffusion Models to Large Datasets
- SAM 3: Segment Anything with Concepts
- VideoCrafter1: Open Diffusion Models for High-Quality Video Generation
- LTX-Video: Realtime Video Latent Diffusion
- VEnhancer: Generative Space-Time Enhancement for Video Generation
- Imagen Video: High Definition Video Generation with Diffusion Models
- Mixture of Diffusers for scene composition and high resolution image generation
- DINOv3
- Wan: Open and Advanced Large-Scale Video Generative Models
- InternVL3.5: Advancing Open-Source Multimodal Models in Versatility, Reasoning, and Efficiency
- VBench-2.0: Advancing Video Generation Benchmark Suite for Intrinsic Faithfulness
Related papers
- Loss Knows Best: Detecting Annotation Errors in Videos via Loss Trajectories
- AnchorWeave: World-Consistent Video Generation with Retrieved Local Spatial Memories
- Benchmarking the Robustness of Foundation Models for Mammography under Domain Shift
- MambaX-Net: Dual-Input Mamba-Enhanced Cross-Attention Network for Longitudinal MRI Segmentation
- TeleOCR: Navigating Document Parsing Across Digital and Camera-Captured Documents
- A Survey on Efficient Vision-Language-Action Models