Bake It Till You Make It: Ultrafast Spatial Texture-Atlas Splatting
Listen
Radio episode about this paper
Transcript
Introduction to the show: ident: AI Radio. Generated commentary on the latest Artificial Intelligence papers.
Tom: Today's paper: "Bake It Till You Make It".
Jane: This paper introduces "Ultrafast Spatial Texture-Atlas Splatting," a novel method for novel view synthesis that aims to achieve real-time 4K rendering by decoupling high-frequency texture details from geometry and view-dependent…
Tom: First, who's behind it and why it matters.
Title and authors: Tom: So, we're diving into "Bake It Till You Make It: Ultrafast Spatial Texture-Atlas Splatting" today, and the team is buzzing because this paper tackles a major hurdle in real-time rendering. Jane, can you give us the basic flavor of what this whole thing is about?
Jane: Absolutely. Basically, this paper proposes a new way to make three dee Gaussian Splatting much faster while keeping the visual quality high. They achieve this by splitting the scene's look into two parts: a view-independent high-frequency texture detail and a view-dependent color that changes depending on where you're looking from. It’s like separating the fine surface scratches from the basic shape of an object.
Lu: That separation is what interests me most; it sounds like they're moving away from purely geometric or purely texture-based methods to get a better result. This decoupling allows for much more sophisticated modeling of complex surface variations, which is something I think has huge potential for future scene reconstruction.
Meng: From an engineering standpoint, the goal here seems to be achieving real-time 4K rendering on consumer hardware, which is a tough target for any AI system right now. How does this paper actually manage to get that speed increase over existing methods?
Tom: That's the million-dollar question, Meng. The core innovation involves replacing slow neural queries with super fast GPU texture lookups after a clever baking process. They represent the low-frequency geometry using 2D Beta surfels and store those high-frequency details in a baked RGB texture atlas, which lets them do fast per-fragment queries during rendering.
Jane: So, think of it like this: instead of asking a complicated network for every little surface detail every frame, the system pre-calculates those details and just looks up a picture from an atlas. This is what allows them to get up to fivefold speedup compared to three dee Gaussian Splatting.
Lalam: If I consider the implications for our culture, this means we can generate photorealistic 4K scenes in real-time. This moves us closer to interactive experiences where visual fidelity isn't limited by processing power, which is really exciting for how we create and consume media.
Lu: I think the idea of using a spatial multi-resolution hash grid to capture that texture residual is really clever, especially when combined with the per-primitive view-dependent color learned via a soft spherical Voronoi representation. It suggests a very structured way to organize scene information.
Title and authors: Tom: Exactly! And they aren't just stopping there; they introduce some specific optimizations to prune the geometry down. They use a falloff-reduction regularizer that flattens the kernel for sparser geometry, which helps control overdraw during rendering.
Jane: That regularization term is interesting because it actively tries to pull the kernels of accurate surfels toward a hard-edged uniform disk, which limits how much spatial influence each primitive has. It’s a way to make sure we aren't wasting computation on redundant parts of the scene.
Meng: But I have to ask about the training side; managing millions of primitives sounds computationally expensive during the learning phase. What did they do to handle that bottleneck?
Tom: They employed a hybrid training strategy, which includes a periodic freezing schedule after five thousand iterations, letting the geometric surfels adjust their positions while keeping the texture output stable at a lower cost. It’s a smart way to manage the computational load over time.
Lalam: For us, that suggests that future AI models don't have to be purely brute-force during training; we can use smarter scheduling to focus computational power where it matters most for stability. It shows a path toward more efficient model development overall.
Lu: The combination of the Deformable Beta kernels and the spherical color models captures both broad diffuse shifts and sharp specular reflections simultaneously, which is a significant advancement over simpler methods. It gives us a richer representation of how light interacts with surfaces.
Jane: That ability to model those sharp reflections accurately through the spherical Voronoi representation really boosts the visual quality, even though we’re using an atlas for the fast lookup. It shows that speed and detail aren't always in direct conflict.
Tom: So, to wrap up on the core mechanism of this "Bake It Till You Make It: Ultrafast Spatial Texture-Atlas Splatting," we’ve seen how they disentangle geometry from texture and then use a baked atlas for inference. This leads us perfectly into what specific improvements they claim in terms of performance and fidelity.
Meng: Beyond the speedup, I'm focused on the efficiency gains they report regarding the primitive count. They mention achieving up to twenty-four times faster performance than BBSplat specifically in low-primitive regimes. That suggests a real impact on memory usage for complex scenes.
Title and authors: Jane: And that makes sense because they also report achieving up to six times faster performance than three deeGS, which is a substantial improvement in the speed of generating views. It’s not just about being fast; it’s about being fast without sacrificing the visual quality that we expect from these methods.
Lu: The structure of how they store the data in that BC7 compressed texture atlas, achieving a memory reduction while maintaining visual quality, is a practical engineering win. This shows that efficient storage techniques are just as important as the rendering speed itself.
Lalam: I see this impacting our development culture by proving that we can push the limits of efficiency without needing exponentially more compute, which encourages us to think about data representation in a way that is fundamentally smarter. It’s about finding the right structure for information.
Tom: So, to summarize these improvements, we have a method that gives us real-time 4K rendering on hardware like the MacBook M3 Pro at sixty FPS. It also produces a significantly lower primitive count than some competitors, like achieving nine hundred fifty-one point eight PSNR with only one hundred fifty-six thousand points compared to BBSplat's three hundred which is a major point.
Jane: That low primitive count combined with high fidelity means we can reconstruct much more detailed scenes than before. The system manages to maintain state-of-the-art visual fidelity while being extremely fast and memory efficient.
Meng: For practical deployment, the fact that they can achieve real-time 4K rendering suggests this is viable for applications that need immediate visual feedback, like interactive design or virtual reality environments. It moves these techniques out of purely research settings and into usable product space.
Lu: Looking ahead, the paper mentions optimizing sparsity further using multiview error-based optimization, which lets them gate densification and pruning based on photometric error across multiple views. That level of iterative refinement suggests a robust path toward even more optimized scene representations.
Lalam: For me, that focus on error-based optimization points toward the future of AI systems being incredibly self-correcting and adaptive during their creation process. We can build systems that learn exactly what geometry is truly important to represent.
Tom: So, to wrap up our discussion on "Bake It Till You Make It: Ultrafast Spatial Texture-Atlas Splatting," we've seen how this paper successfully decouples geometry and texture using a hash grid atlas for blazing fast rendering. This results in real-time 4K rendering on consumer hardware and a superior quality-to-primitive ratio compared to methods like BBSplat.
Title and authors: Jane: It really shows how combining smart data structures, like the spatial hash grid, with clever baking techniques can unlock performance that was previously unattainable in this area. The final results confirm high inference speed on consumer hardware at sixty FPS on a MacBook M3 Pro.
Lu: The overall implication is that we can have scenes that are simultaneously extremely fast to render, memory efficient, and photorealistically detailed. This opens up a lot of possibilities for complex scene representation in the future.
Meng: I think the most practical impact is moving these techniques from research papers into products where performance metrics matter immediately, like real-time visualization tools. It validates that optimizing for speed and fidelity together is a viable engineering path.
Lalam: For our AI culture, this paper reinforces the idea that there are often more effective ways to structure data than we initially assume, especially when dealing with complex visual information. It’s a great reminder to look for those structural efficiencies in our own models and applications.
Tom: So, "Bake It Till You Make It: Ultrafast Spatial Texture-Atlas Splatting" is a major step forward in how we handle high-frequency visual data in three dee reconstruction. We’ve explored the title, the summary, the specific improvements, and what this means for future work.
Jane: It’s been fascinating seeing how they used that texture atlas to eliminate neural query overhead during rendering. We're definitely looking forward to seeing how these ideas evolve in the next generation of three dee scene synthesis papers.
Lu: I think the combination of view-independent neural fields and baked residuals sets a strong foundation for future work in disentangling appearance features even further. It’s a solid blueprint for how to structure such complex models.
Meng: For me, the real question is how much more aggressive this sparsity optimization can get before the visual fidelity starts dropping, which is where I need to see concrete data for commercial viability.
Lalam: I think we should keep an eye on this space because it shows that even complex visual problems can be solved by breaking them down into structured components, which is a valuable lesson for our broader AI goals.
Tom: That’s all the time for this deep dive into "Bake It Till You Make It: Ultrafast Spatial Texture-Atlas Splatting." Thanks to Jane, Lu, Meng, and Lalam for bringing such insightful perspectives on this paper.
The paper's summary: Tom: So, we've been talking about how this new method tackles separating geometry from texture detail using that baked texture atlas, and now Jane, can you lay out the main gist of what the paper actually proposes?
Jane: Well, essentially, this work introduces a novel approach to three dee scene reconstruction called Ultrafast Spatial Texture-Atlas Splatting. The core idea is to divide the scene's appearance into two distinct pieces: one part that doesn't change based on where you are looking, and another part that does depend on your view. They handle the geometry using these surfels, and they capture that high-frequency texture detail by baking it into a compact RGB texture atlas.
Tom: That sounds like they’re trying to solve a massive bottleneck in rendering speed by moving away from constant complex neural calculations during the final viewing pass, right?
Jane: Exactly. They replace those slow queries with very fast 2D lookups into that atlas, which is what lets them aim for real-time 4K rendering on consumer hardware. The authors show that this method can achieve up to fivefold speedup over standard three dee Gaussian Splatting while still maintaining a top-tier visual quality.
Lu: I’m really thinking about the structural implications here, Tom; decoupling the view-independent geometry from the view-dependent texture is a very clean way to organize scene data. It opens up possibilities for much more flexible scene representations down the line, something we can build on.
Meng: From an engineering standpoint, that speedup is huge because it means real applications could actually use these systems immediately instead of waiting months for training or slow inference times. I’m focusing on how practical this is for deployment.
Lalam: For me, the cultural implication is that this demonstrates a path where high fidelity and extreme efficiency aren't mutually exclusive; it shows that we can build tools that are both visually stunning and incredibly responsive in real-time applications.
Tom: It really comes down to how they manage the training side too. They’ve got some clever regularization techniques to keep the geometry sparse and efficient during the learning process, which is a big practical step toward making these models easier to train effectively.
Jane: That hybrid training strategy, involving a periodic freezing schedule after five thousand iterations, seems like a smart compromise for balancing computational cost and model refinement. It suggests that we don't need to run everything at maximum intensity all the time.
Lu: And their use of Deformable Beta kernels to smooth out geometry while keeping the high-frequency texture separate is a really creative way to handle the details without overwhelming the system with unnecessary complexity during rendering.
Meng: I’m interested in those performance metrics they report; when they mention achieving up to twenty-four times faster performance than BBSplat in low-primitive regimes, that’s what tells me this has real weight for memory efficiency, which is a huge factor for deploying these models on mobile devices or edge hardware.
Lalam: I see the cultural impact as enabling a new generation of interactive content creation where the barrier to entry for achieving photorealistic results is significantly lowered because the system can handle complex scenes in real-time without massive computational overhead.
Tom: So, we've seen that this method successfully decouples geometry and texture, uses a baked atlas for speed, and manages training bottlenecks smartly; what do you think these findings mean for how we think about three dee scene synthesis overall?
The paper's improvements: Tom: So, we've broken down how they use that baked texture atlas for rendering speed, and now Jane, can you walk us through the specific improvements they highlight in terms of performance gains and visual quality?
Jane: They really emphasize a few key areas where this method shines. First is the geometric sparsity; they show that by using techniques like multi-view error optimization, you can aggressively prune insignificant elements to get a much smaller primitive count while keeping the visual fidelity high.
Tom: That’s impressive, because reducing the number of points directly translates to lower memory usage and faster rendering times, which is exactly what engineers care about.
Lu: I think that geometric pruning combined with the falloff-reduction regularizer is what really sets this apart; it actively pulls the kernels of accurate surfels toward a uniform disk shape, which helps bound their spatial influence and reduces overdraw during rendering. That’s a very sophisticated way to control complexity.
Meng: From my side, that focus on memory efficiency and reduced overdraw is crucial because it makes the system viable for deployment on less powerful hardware where VRAM is limited. It moves this past just being a cool research demo into something that can actually run in production environments.
Lalam: For me, the biggest implication here is how we can build AI systems that are both highly detailed and extremely resource-efficient; it proves we don't have to sacrifice one for the other when structuring our scene data.
Tom: And visually, they also stress that this approach handles view-dependent appearance better than previous methods by using a soft spherical Voronoi representation for color features, which captures those sharp specular highlights more accurately.
Jane: That’s a great point because capturing both broad diffuse shifts and sharp reflections simultaneously in that way really boosts the realism of what the system produces. It shows that speed and detail aren't competing against each other when you use these specialized kernels.
Lu: The texture baking process itself is also an improvement; they store everything in a BC7 compressed atlas, which achieves a significant memory reduction compared to using full-resolution neural representations while introducing no perceivable quality loss. That’s a smart data storage solution.
Meng: I need to ask about the limitations here; the authors mention that while they achieve high FPS, the method still relies on those 2D texture lookups, meaning if you needed something completely novel outside what's baked into that atlas, it would be challenging. That’s a constraint we should keep in mind for future scaling.
Lalam: That constraint actually points toward an interesting future direction; it suggests that the next evolution of this field might involve dynamic texture updates or more adaptive atlas generation to handle truly novel visual information on the fly.
Tom: So, to wrap up these improvements, we have a system that aggressively prunes geometry for efficiency while maintaining high fidelity through clever kernel regularization and improved view-dependent color modeling.
Jane: It really shows how combining structured data representation with optimized baking can lead to a much more practical rendering pipeline for three dee synthesis.
Lu: I think the path forward involves pushing that sparsity optimization even further using error-based methods, which promises even cleaner scene representations down the line.
Conclusion: Tom: So, to wrap up on "Bake It Till You Make It: Ultrafast Spatial Texture-Atlas Splatting," we've seen how this method successfully decouples geometry and texture using a hash grid atlas for blazing fast rendering.
Jane: It really shows how combining smart data structures, like the spatial hash grid, with clever baking techniques can unlock performance that was previously unattainable in this area.
Tom: And the final results confirm high inference speed on consumer hardware at sixty FPS on a MacBook M3 Pro.
Lu: The overall implication is that we can have scenes that are simultaneously extremely fast to render, memory efficient, and photorealistically detailed. It opens up a lot of possibilities for complex scene representation in the future.
Meng: For practical deployment, this validates that optimizing for speed and fidelity together is a viable engineering path for real-time visualization tools.
Lalam: I see this impacting our development culture by proving that we can push the limits of efficiency without needing exponentially more compute, which encourages us to think about data representation in a way that is fundamentally smarter.
Tom: We've explored the title, the summary, the specific improvements, and what this means for future work.
Jane: It’s been fascinating seeing how they used that texture atlas to eliminate neural query overhead during rendering.
Lu: I think the combination of view-independent neural fields and baked residuals sets a strong foundation for future work in disentangling appearance features even further.
Meng: For me, the real question is how much more aggressive this sparsity optimization can get before the visual fidelity starts dropping, which is where I need to see concrete data for commercial viability.
Lalam: I think we should keep an eye on this space because it shows that even complex visual problems can be solved by breaking them down into structured components, which is a valuable lesson for our broader AI goals.
Tom: So "Bake It Till You Make It: Ultrafast Spatial Texture-Atlas Splatting" is a major step forward in how we handle high-frequency visual data in three dee scene reconstruction.
Jane: We're definitely looking forward to seeing how these ideas evolve in the next generation of three dee scene synthesis papers.
Lu: I think the combination of view-independent neural fields and baked residuals sets a strong foundation for future work in disentangling appearance features even further.
Meng: I still see it as an important step toward making these models more deployable by focusing on those hardware-friendly data structures.
Lalam: This work reinforces the idea that there are often more effective ways to structure data than we initially assume, especially when dealing with complex visual information.
Neel Kelkar, Simon Niedermayr, Kaloian Petkov, Klaus Engel, Rüdiger Westermann
Technical University of Munich · Siemens Healthineers
cs.CV, cs.GR
Submitted: 2026-07-15
Updated: 2026-09-29
Comments: 15 main pages, 8 figures. Extra information in Appendix. Project page with videos and interactive demos: https://nilkel.github.io/bitymi/
Project page: https://nilkel.github.io/bitymi
License: http://arxiv.org/licenses/nonexclusive-distrib/1.0/
Importance score: 79/100
The gist: This paper introduces "Ultrafast Spatial Texture-Atlas Splatting," a novel method for novel view synthesis that aims to achieve real-time 4K rendering by decoupling high-frequency texture details
Key concepts
- 2D Beta Surfels
- These are the fundamental building blocks used to model the scene's geometry. They are parameterized by position, rotation, scale, and opacity. The authors use 'Deformable Beta kernels' which morph these shapes between a standard Gaussian distribution and an opaque disk to reduce rendering overdraw and improve speed.
- Texture Baking
- Instead of querying a neural network for every fragment texture detail during rendering, the high-frequency texture information is pre-calculated and stored in a global RGB atlas. This allows the system to use fast GPU texture lookups during inference, significantly boosting rendering speed.
- Per-Primitive View-Dependent Color
- This component captures how light interacts with each specific surface based on its view. It is represented using a soft spherical Voronoi representation (SV), which divides the sphere among learnable points. This allows for smooth and differentiable transitions between different appearance regions across the scene.
- BC7 Quantization
- To reduce memory usage, the pre-baked texture atlas is compressed using Block Compression (BC7). This converts floating-point RGB data into 8-bit RGBA data, resulting in a significant memory reduction while introducing no noticeable loss in visual quality.
Terminology
Summary
This paper introduces Ultrafast Spatial Texture-Atlas Splatting,
a novel method for novel view synthesis that aims to achieve real-time 4K rendering by decoupling high-frequency texture details from geometry and view-dependent appearance features. By representing low-frequency geometry and view-dependent appearance with 2D surfels and storing high-frequency textures in a compact, baked RGB texture atlas, the authors propose a pipeline that leverages fast GPU texture lookups to achieve up to fivefold speedup over 3D Gaussian Splatting while preserving state-of-the-art visual fidelity.
Method Overview
The core innovation lies in separating scene appearance into two distinct components: a view-independent high-frequency texture residual per fragment and a per-primitive view-dependent color. During training, the geometry is modeled using 2D Beta surfels, while the texture residual is captured by a spatial multi-resolution hash grid and a world-space multilayer perceptron (MLP) decoder. The final fragment color is computed as:
c(x, d) = ReLU fSV (d) + bSV + f phi Etheta (x).
Key Components and Techniques
The method relies on several key technical components to achieve its performance gains:
-
A view-independent neural field with sparsity optimization that produces a compact set of surfels with per-primitive appearance variations.
-
A baking procedure that converts the optimized hash grid into an RGB texture atlas, leveraging native GPU texture sampling for high inference speeds. This allows for
fast, per-fragment queries during rendering.
-
A falloff-reduction regularizer that flattens the per-primitive kernel for sparser geometry and lower fragment overdraw. This is implemented by adding a loss term: L β = 1/P ∑ p∈P w(p) ¯β (p), where w(p) scales with reconstruction quality.
-
A hybrid training strategy to mitigate training-time bottlenecks, involving a
periodic freezing schedule
after 5000 iterations to allow geometric surfels to adjust their positions and fit the stable texture output at a lower computational cost.
Geometry and Appearance Modeling
The geometry is modeled using 2D Beta surfels, which are parameterized by position, rotation, scale, and opacity. These kernels are replaced with Deformable Beta kernels
that morph between a Gaussian distribution and an opaque disk to reduce overdraw and increase rendering speed. The per-primitive view-dependent color is captured using a soft spherical Voronoi representation (SV), which partitions the sphere among learnable generator points, allowing for smooth, differentiable transitions between regions.
Texture Baking and Inference Speedup
To eliminate neural representation query overhead during inference, the view-independent hash grid output is baked into a global RGB texture atlas. This process involves:
-
Assigning an anisotropic UV-grid resolution r g to each surfel based on its scale and tangent axis to match the Nyquist rate of the hash grid along each tangent axis.
-
Storing the per-surfel grids in an atlas using a
Shelf-First-FitDecreasing algorithm.
-
Applying Block Compression (BC7) quantization, resulting in a memory reduction from FP16 RGB to uint8 RGBA, achieving
a significant memory reduction is achieved
while introducingno perceivable quality loss.
Optimization and Results
Sparsity is further enhanced by the multiview error-based optimization of FastGS, which computes an importance score based on photometric error across multiple views to strictly gate densification and pruning. The falloff regularizer actively pulls the kernel of accurate surfels toward a hard-edged uniform disk,
bounding its spatial influence. Experiments show that the method achieves up to 6× faster performance than 3DGS and up to 24× faster than BBSplat in low-primitive regimes,
demonstrating a superior quality-to-point ratio. The results confirm high inference speed on consumer hardware, achieving real-time 4K rendering at 60 FPS on a MacBook M3 Pro.
Conclusion
The paper concludes that the proposed texture-based method successfully achieves high inference speed and high-quality reconstruction with a low number of surfels. It demonstrates a superior quality-to-primitive ratio and rendering performance on par with the fastest existing methods, but with significantly higher fidelity.
The final results show that post-baking into a BC7 atlas significantly reduces memory footprint while maintaining visual quality.
Table 1: Atlas storage formats.
Format Bytes / texel Mean atlas size ΔPSNR vs. neural FP16 RGB (ref.)
:---:---:---:---
uint8 RGBA (ours) 4 ∼2.6 GB ±0.
Improvements for AI systems
As a fastidious researcher, I have analyzed Bake It Till You Make It: Ultrafast Spatial Texture-Atlas Splatting.
This work introduces a novel decoupled representation for 3D Gaussian Splatting (3DGS) that separates low-frequency geometry/view-dependent appearance from high-frequency textures via a view-independent spatial hash grid baked into an RGB texture atlas.
Here are the specific improvements and capabilities this method enables in AI systems:
)1. Improved Real-Time 3D Scene Reconstruction and Novel View Synthesis (NVS):
The system can perform photorealistic 4K rendering at 60 FPS on consumer hardware (as demonstrated by achieving high FPS on MacBook M3 Pro), a significant leap over standard 3DGS.
-
It achieves up to a fivefold speedup over standard 3DGS and up to twenty-four times faster than BBSplat in low-primitive regimes.
-
It maintains state-of-the-art visual fidelity (PSNR/SSIM) by decoupling geometry from appearance, allowing for the reconstruction of complex scenes with high texture detail that were previously impossible with purely geometric or purely texture-based methods.
)2. Enhanced Geometric Sparsity and Efficiency:
The method produces a compact set of surfels by aggressively pruning insignificant elements through sparsity-enhancing optimizations (like multi-view error optimization based on FastGS principles).
- The system can achieve higher geometric sparsity than prior work, leading to a reduced primitive count (e.g., 648K points vs. 156K for BBSplat in some comparisons) while maintaining or exceeding fidelity, resulting in lower memory footprints and faster inference times.
)3. Superior Texture Representation Handling:
The use of a view-independent spatial hash grid baked into an RGB texture atlas fundamentally changes how high-frequency details are stored and accessed.
-
The system can represent intricate, high-frequency textures (like fine surface patterns or complex material variations) without requiring a prohibitive number of primitives or per-primitive textures, which typically scale poorly with scene resolution.
-
This enables the reconstruction of scenes where texture fidelity is paramount, as the hash grid provides a dense, view-independent spatial feature map for texture residuals.
)4. Optimized Rendering Pipeline:
The pipeline transitions from expensive neural field queries (during training/inference) to fast, hardware-accelerated 2D texture lookups during rendering.
-
Inference relies solely on fast 2D texture sampling (BC7 compressed atlas), eliminating the bottleneck of querying a complex MLP or hash grid for every fragment during the final rendering pass.
-
The system benefits from architectural optimizations like SnugBox and AccuTile to maximize hardware utilization, leading to high framerates even under heavy overdraw conditions.
)5. Robust Handling of View-Dependent Appearance:
By using a soft spherical Voronoi representation for per-primitive color features, the system can model complex view-dependent effects (like sharp specular highlights) more accurately than standard Spherical Harmonics (SH).
- The system can capture both broad diffuse shifts and sharp reflections simultaneously on each primitive, leading to a higher LPIPS score improvement compared to methods relying solely on diffuse representations.
In summary, this research allows for the creation of 3D scenes that are simultaneously:
-
Extremely fast to render (real-time capability).
-
Highly memory efficient (low primitive count).
-
Photorealistically detailed (high texture fidelity and view-dependent effects).
Sources
- Spherical Voronoi: Directional Appearance as a Differentiable Partition of the Sphere
- Hybrid Latents: Geometry-Appearance-Aware Surfel Splatting
- FastGS: Training 3D Gaussian Splatting in 100 Seconds
- Nexels: Neurally-Textured Surfels for Real-Time Novel View Synthesis with Sparse Geometries
Related papers
- Loss Knows Best: Detecting Annotation Errors in Videos via Loss Trajectories
- AnchorWeave: World-Consistent Video Generation with Retrieved Local Spatial Memories
- Benchmarking the Robustness of Foundation Models for Mammography under Domain Shift
- MambaX-Net: Dual-Input Mamba-Enhanced Cross-Attention Network for Longitudinal MRI Segmentation
- TeleOCR: Navigating Document Parsing Across Digital and Camera-Captured Documents
- A Survey on Efficient Vision-Language-Action Models