GeoCache: Training-Free Acceleration of Multi-View Texture Diffusion via Geometric Delta Transport

arXiv:2608.13255 · cs.CV, cs.AI · Submitted 2026-08-13 · Read on arXiv

Haotang Li, Zhenyu Qi, Shaohan Henry Wang, Kebin Peng, Yutong Zhao, Zi Wang, Bo Liu, Huanrui Yang, Sen He

University of Arizona · East Carolina University · California State University, Long Beach · Augusta University

cs.CV, cs.AI

Submitted: 2026-08-13

Updated: 2026-08-14

License: http://arxiv.org/licenses/nonexclusive-distrib/1.0/

Importance score: 75/100

The gist: GeoCache is a training-free acceleration method for multi-view texture diffusion models.

Terminology

Summary

GeoCache is a training-free acceleration method for multi-view texture diffusion models. The paper identifies that geometry-conditioned multi-view diffusion enables high-quality 3D texture generation, but its repeated per-view denoiser evaluations introduce substantial computational cost. Existing training-free accelerators primarily exploit temporal redundancy by reusing computation across denoising steps, but in multi-view texturing, skipping a step removes the cross-view interaction that continually aligns different observations of the same surface, leading to rapidly degraded consistency and fidelity.

The paper's analysis identifies a complementary source of redundancy: although intermediate features remain view-specific, geometrically corresponding surface points exhibit transferable evolution in their predicted clean signals. Based on this observation, the authors introduce GeoCache, a training-free plugin that evaluates a rotating subset of anchor views and transports their geometry-aligned per-step x0 updates to the remaining views. Periodic full-view computation controls accumulated error, while sampler-consistent reconstruction preserves the denoising trajectory. GeoCache requires neither retraining nor architectural modification and uses the position maps already available in geometry-conditioned texturing pipelines.

The motivating study analyzes Hunyuan3D-2.1 Paint, which jointly denoises six geometry-conditioned views with a 15-step UniPC sampler. The study finds that temporal output reuse (step caches) introduces a consistency failure beyond the quality loss produced by shortening the trajectory. At a matched speedup of approximately 2.1×, MagCache raises SeamErr to 1.21× the stock value. The study also finds that deep features at matched surface tokens have mean cosine similarity 0.362, only 2.9% of block-step cells exceed 0.6, and the mean advantage over randomly paired tokens is 0.10, meaning intermediate features remain view-specific. Oracle substitutions show that copying an anchor view's x0 value into another view produces MV-LPIPS 0.090, while transporting only the anchor's per-step change in x0 and adding it to the target view's own previous state reduces the error to 0.025 at identical compute.

The methodology works as follows: From position maps, GeoCache precomputes a sparse linear gather operator Gu→v that for each target token takes K nearest source taps within a tolerance of 1% of the bounding-box diagonal, area-weights them, and sums them. At a cached step, only a of the N views run the denoiser (the anchors A, which rotate every step). The remaining views keep their own state and integrate the anchors' per-step change in x0, transported through correspondence: x0(v)(t) = x0(v)(t−1) + GA→v[Δx0(A)(t)]. The transported x0 is converted back to the sampler's native prediction parameterization (epsilon, v, or velocity) to keep the solver's history buffer consistent. Four full steps bound error: a two-step head, one mid-trajectory refresh, and a tail step before decoding.

Results across three backbones show GeoCache achieves a stronger speed–fidelity trade-off than temporal caches and step reduction at operating points above 2×. On Hunyuan3D-2.1, GeoCache reaches 2.21× denoiser-loop speedup with MV-LPIPS of 0.0293 and MV-PSNR of 33.60 dB, providing the best fidelity among all tested methods above 2×. The same transferred configuration reaches the highest speedup (2.60×) and lowest FLOPs (16.24 T) on SyncMVD, while GeoCache achieves the lowest FLOPs (102.30 T) and best fidelity (MV-LPIPS 0.0282) among the accelerated methods on MVPainter at 4.04× speedup. Across three backbones and four asset pools, each additional 0.1× of speed costs GeoCache +3.1% MV-LPIPS against +12.4 to +33.5% for the step caches.

Ablations show delta transport is the most load-bearing choice: a value copy costs 2.8× the LPIPS at equal speed. Refresh placement beats refresh count, and the anchor count sits at the knee of the trade-off curve. Shuffling correspondence alone leaves LPIPS within 3% of the base but raises seam error 11%, showing correspondence expresses itself in cross-view consistency.

Limitations include dependence on geometric correspondence quality, the end-to-end gain following the fraction of pipeline time spent denoising (1.11× on the paint stage and 1.07× end-to-end at 6×5122), and reduced effectiveness on long denoising trajectories where TaylorSeer reaches higher fidelity on MV-Adapter's 50-step Euler sampler.

Improvements for AI systems

Based on the paper, here are the specific improvements I can make to AI systems and what the improved systems can do:

1. Cross-view feature transport for multi-view consistency

  • I can implement a sparse, geometry-aligned gather operator that maps per-step denoising updates (Δx0) from anchor views to non-anchor views using precomputed position maps, rather than copying full features or outputs.

  • The improved system can maintain cross-view consistency during accelerated inference without retraining, reducing seam errors by 11% compared to random correspondence and achieving MV-LPIPS of 0.025 (vs. 0.090 for value copying) at identical compute.

2. Sampler-consistent state injection

  • I can convert transported x0 predictions back into the sampler's native parameterization (epsilon, v, or velocity) before updating the solver's history buffer, preserving the denoising trajectory.

  • The improved system can integrate acceleration into any existing diffusion sampler (UniPC, Euler, etc.) without breaking convergence, enabling seamless drop-in acceleration for multi-view texturing pipelines.

3. Rotating anchor selection with bounded error control

  • I can implement a schedule where anchor views rotate every step, with four mandatory full-view computations (two-step head, one mid-trajectory refresh, one tail step) to bound accumulated error.

  • The improved system can achieve 2.21×–4.04× denoiser-loop speedups across backbones while keeping MV-LPIPS below 0.03, with each additional 0.1× speed costing only +3.1% LPIPS (vs. +12.4–33.5% for temporal caches).

4. Adaptive cache triggering based on geometric redundancy

  • I can precompute a per-token correspondence quality metric (e.g., nearest-neighbor distance within 1% of bounding-box diagonal) to decide when transport is reliable, falling back to full computation for low-confidence regions.

  • The improved system can handle objects with complex topology or occlusions more robustly, avoiding quality degradation where geometric correspondence is ambiguous.

5. Cross-backbone generalization without architectural changes

  • I can package GeoCache as a plugin that reads existing position maps from geometry-conditioned pipelines, requiring no retraining or model modification.

  • The improved system can accelerate any multi-view diffusion model (Hunyuan3D-2.1, SyncMVD, MVPainter) with a single configuration, reducing FLOPs by up to 60% while maintaining state-of-the-art fidelity among accelerated methods.

6. Error-aware refresh placement

  • I can optimize the placement of full-view refreshes based on trajectory position (e.g., mid-trajectory refresh outperforms adding more refreshes at arbitrary points), rather than using uniform or count-based schedules.

  • The improved system can achieve better speed-fidelity trade-offs with fewer full computations, as demonstrated by ablations showing refresh placement beats refresh count.

7. Delta-transport prioritization over value copying

  • I can enforce that only per-step changes (Δx0) are transported, never absolute values, since delta transport reduces MV-LPIPS by 2.8× compared to value copying at equal speed.

  • The improved system can maintain higher fidelity during long denoising runs, making it suitable for high-resolution texture generation (e.g., 6×5122) where accumulated errors would otherwise dominate.

Sources

Related papers