CLONE: Continuous Latent Optimization for Normal Estimation via 3D Gaussian Splatting

arXiv:2508.05950 · cs.CV, cs.AI · Submitted 2026-08-15 · Read on arXiv

Listen

Radio episode about this paper

Transcript

Introduction to the show: ident: AI Radio. Generated commentary on the latest Artificial Intelligence papers.

Tom: Next we'll be talking about the paper "CLONE: Continuous Latent Optimization for Normal Estimation via 3D Gaussian Splatting".

Jane: The paper was written by Yanxing Liang, Yinghui Wang, Wei Li, Tao Yan and Jiaxing Shen from Jiangnan University and Lingnan University.

Tom: Stay tuned as we take you through the paper and discuss its implications.

Title: Tom: Alright, welcome back to the show, everybody. I'm Tom, and as always, I'm here with my co-host, the brilliant Jane. And we are looking at a new paper that just hit arXiv, and it is called "CLONE: Continuous Latent Optimization for Normal Estimation via three dee Gaussian Splatting."

Jane: Tom, I have to say, that title is a mouthful, but the idea behind it is actually pretty simple once you break it down. We're looking at a way to teach a computer to figure out the three dee shape of an object from just a single photo. Specifically, we want to know the direction every surface is pointing, which we call the "normal."

Tom: Right, and the key word in that title is "weakly supervised." In the old days, you needed millions of images where every single pixel had a perfectly labeled three dee direction. That's expensive and hard to get. This paper says, hey, we can do it without those labels, just by using the object's three dee model and the image together.

Jane: Exactly. And the "three dee Gaussian Splatting" part is the clever trick. Instead of trying to guess the shape pixel by pixel, they represent the object as a bunch of tiny, squishy three dee blobs—Gaussians. By looking at how those blobs are shaped and oriented, you can mathematically derive which way the surface is pointing.

Tom: It’s like looking at a pile of jellybeans. If a jellybean is flat and wide, you know the surface there is flat. If it’s tall and skinny, you know the surface is curving away. That’s the intuition, right?

Jane: That’s the intuition. And the authors—Liang, Wang, Wu, Li, Yan, and Shen—they’ve built a whole pipeline around that. They call it an "image-geometry-image" loop. You look at the image, you predict the blobs, you render the blobs back into an image, and you check if it matches the original photo. If it doesn’t, you adjust the blobs.

Tom: So it’s self-correcting. The computer is essentially checking its own homework using the light and shadows in the photo, without ever being told the "right answer" for the normals.

Jane: Precisely. And that’s what makes this so exciting. It means we could train these models on massive libraries of three dee objects—like the Objaverse dataset they used—without needing a human to manually annotate every surface. That’s a huge unlock for scalability.

Tom: And the implications? Think about robotics. A robot arm needs to know the shape of the object it’s about to pick up. Or augmented reality, where you want to place a virtual chair on a real floor and have the lighting look right. This kind of technology makes those systems more robust and cheaper to build.

Jane: And it’s not just about robots. It’s about making three dee understanding accessible. If you can get high-quality geometry from a single photo, you can digitize cultural artifacts, preserve historical objects, or create three dee content for games and movies much faster.

Tom: So we’ve got the title and the big idea. But how do they actually pull this off? The paper has a few clever components, and I’m curious to see how they stack up. Let’s dig into the summary next.

Jane: Sounds good. Let’s see how they make this magic happen.

Summary: Tom: So, Jane, we’ve set the stage with the title. Now let’s get into the meat of the paper, "CLONE: Continuous Latent Optimization for Normal Estimation via three dee Gaussian Splatting." The authors don’t just use one method; they stack three different ideas together to get the best result.

Jane: Right. It’s like building a car. You need an engine, a steering wheel, and a suspension. Each part does a different job, and they have to work together. The first part is the three dee Gaussian Splatting itself, which gives you a rough, smooth estimate of the shape. It’s physically grounded, but it’s a bit blurry.

Tom: Blurry is the right word. Because those Gaussian blobs are smooth, they miss the fine details—the edges of a leaf, the grooves on a keyboard, the feathers on a bird. So the paper introduces a second part: a diffusion-based refinement network. This is like a detail restorer. It takes the smooth estimate and sharpens it up.

Jane: But here’s the catch. Diffusion models are usually trained with labels, and we don’t have labels here. So they reformulate it as a "single-step deterministic" process. Instead of iterating many times, it does one clean pass, and it’s guided by the photometric loss—the difference between the rendered image and the real image. That keeps it trainable without ground truth.

Tom: And then there’s the third part, the gating fusion. Because you have two estimates now—the smooth geometric one and the sharp but possibly noisy one—you need to decide which to trust at each pixel. The gating mechanism learns to do that. In flat areas, it trusts the geometry. In textured areas, it trusts the detail.

Jane: And they prove it works. On the GSO and Omniobjectthree dee benchmarks, they get a mean angle error of thirteen point two degrees. That’s better than the fully supervised baseline, Metricthree dee v2, which gets fourteen point one degrees. So they’re beating methods that had access to perfect labels, using none.

Tom: That’s the headline result. It’s not just "good for a weakly supervised method." It’s actually state-of-the-art, period. And they show it’s robust, too. They tested it with added noise, different lighting, and different viewpoints, and it held up much better than the competition.

Jane: The robustness part is key for real-world use. In a lab, you control the lighting. In the real world, you don’t. The fact that their method degrades gracefully when you move the light source or add noise suggests it’s not just memorizing training conditions.

Tom: So they’ve got the smooth base, the sharp detail, and the smart fusion. But how do they actually train this thing without labels? That’s the part I want to dig into. The loss functions and the optimization loop seem pretty intricate.

Jane: It is intricate, but it’s also elegant. They use a photometric loss, which is just "does my rendered image match the real one?" And then they add a few regularizers to keep the geometry sane—like preventing the Gaussian blobs from becoming infinitely thin or pointing in weird directions.

Tom: So it’s a balancing act. You want to match the image, but you also want the three dee structure to be physically plausible. And they’ve tuned those weights to make it work. I’m excited to see what happens when they push this further. What improvements are they suggesting next?

Jane: That’s the next segment. Let’s get into it.

Improvements: Tom: Alright, we’ve covered the basics and the results. Now, Jane, let’s talk about where this paper says it falls short and what they want to do next. Because no paper is perfect, and the authors are pretty honest about the limitations.

Jane: They are. The biggest failure mode they identify is "texture-geometry confusion." That’s when the model sees a pattern on the surface—like a checkerboard or a printed label—and mistakes it for actual three dee bumps. It’s a classic problem in normal estimation, and they say it accounts for almost half of their errors.

Tom: And they trace that back to their lighting model. They use a simplified Lambertian model, which assumes surfaces are matte. But real objects have specular highlights, metal, glass. Their learnable modulation kernel helps, but it’s not a full physical model. So they want to upgrade to a BRDF model with learnable roughness.

Jane: A BRDF is basically a function that describes how light reflects off a surface. It’s much more realistic. If they can make that differentiable, they could handle shiny and metallic objects much better. That would directly attack that twenty-three percent of failures they attribute to non-Lambertian materials.

Tom: And then there’s the issue of slender structures and self-occlusion. If you have a thin stick or a complex shape where parts hide behind other parts, their method struggles. They mention that explicitly. They don’t have explicit occlusion modeling, so the Gaussian blending can create "ghost geometry."

Jane: Right. And they also mention the training data bias. They train on Objaverse, which is full of man-made objects. So they’re great at chairs and mugs, but they struggle with organic shapes or rare materials. The model is only as good as the data it sees.

Tom: So the improvements they’re suggesting are pretty clear. First, a better BRDF model. Second, a self-supervised pre-training pipeline on egocentric video data, which would give them multi-view consistency without needing synthetic three dee models. That would help with the data bias.

Jane: And third, they want to make it faster. Right now, it takes about zero point six eight seconds per image on a four thousand ninety GPU. That’s fine for offline processing, but not for real-time robotics. They mention using lighter backbones like MobileNetV3 to cut that down.

Tom: So they’re thinking about deployment. That’s good to hear. It’s one thing to have a great result on a benchmark; it’s another to have something you can actually ship. And I think that’s where the real impact will be.

Jane: Absolutely. If they can make it robust to lighting, fast enough for real-time, and handle a wider range of materials, this could be the default way we teach machines to see three dee structure. It’s a roadmap, not just a one-off result.

Tom: And that roadmap leads us to the big picture. Let’s wrap this up and talk about what this means for the world.

Conclusion: Tom: Well, we’ve reached the end of our discussion on "CLONE: Continuous Latent Optimization for Normal Estimation via three dee Gaussian Splatting." And I’ve got to say, Jane, this paper feels like a turning point.

Jane: It really does, Tom. The core achievement is that they’ve shown you can get fully supervised-level accuracy without any pixel-level labels. That’s a massive deal for scalability. It means we can leverage the huge libraries of three dee models that already exist, instead of being bottlenecked by manual annotation.

Tom: And the implications go beyond just making better algorithms. Think about cultural heritage. Museums have thousands of artifacts, and they can’t afford to three dee scan every single one. But they might have photos. With this technology, you could generate high-quality three dee models from those photos, preserving history digitally.

Jane: Or think about e-commerce. You want to show a customer a shoe from every angle, but you only have a single product photo. This method could generate the three dee shape, letting customers rotate the shoe and see it in different lighting. That’s a direct commercial application.

Tom: And for the AI community, this is a blueprint. The idea of closing the loop between image and geometry, using differentiable rendering as a self-supervision signal, is going to influence a lot of future work. It’s not just about normals; it’s about a philosophy of learning.

Jane: Exactly. They’ve shown that you don’t always need ground truth. Sometimes, you can let physics and geometry be the teacher. And that’s a powerful lesson.

Tom: So, as we say goodbye to this paper, I want to thank the authors—Liang, Wang, Wu, Li, Yan, and Shen—for their work. It’s a solid contribution, and we’re excited to see where they take it next.

Jane: And we’re excited to see what the next paper on our list has to offer. But for now, that’s a wrap on CLONE. Thanks for listening, everyone. We’ll see you on the next episode.

Tom: Take care, folks. And keep looking at the world from new angles.

Yanxing Liang, Yinghui Wang, Wei Li, Tao Yan, Jiaxing Shen

Jiangnan University · Lingnan University

cs.CV, cs.AI

Submitted: 2026-08-15

Updated: 2026-08-18

License: http://creativecommons.org/licenses/by/4.0/

Importance score: 59/100

The gist: "The core idea is to construct an image-geometry-image consistency loop that unifies explicit geometric representation with differentiable rendering, thereby enabling weakly supervised learning

Key concepts

Three Dee Gaussian Splatting
This technique represents an object as many tiny, squishy three-dimensional blobs called Gaussians. By examining the shape and orientation of these blobs, the method mathematically derives which way a surface is pointing (the normal). This provides a rough, smooth estimate of the object's shape.
Weakly Supervised Learning
This refers to training a model without needing millions of perfectly labeled data points where every pixel has a correct three-dimensional direction. CLONE achieves this by using the object's 3D model and the image together, allowing it to learn normals from just one photo.
Photometric Loss
This is a loss function used during training that measures how well the computer-rendered image of an object matches the original real photograph. It guides the optimization process by minimizing the difference between what is rendered and what was actually captured in the image.

Terminology

Summary

Summary

The paper proposes CLONE, a weakly supervised framework for single-image object surface normal estimation that does not require pixel-level normal ground truth. The core idea is to construct an image-geometry-image consistency loop that unifies explicit geometric representation with differentiable rendering. The paper states: "The core idea is to construct an image-geometry-image consistency loop that unifies explicit geometric representation with differentiable rendering, thereby enabling weakly supervised learning without normal ground truth."

The framework comprises four main components. First, a differentiable light interaction model with a learnable modulation kernel performs "a unified reparameterization of the 3DGS parameter space, establishing an explicit and stable mapping between 3DGS geometric parameters and surface normals and turning the photometric loss into an internal supervision signal. Second, a conditional single-step deterministic refinement network integrates denoising architectures with differentiable reprojection constraints to refine the initial normals, thereby adaptively recovering the high-frequency details erased by the inherently smooth Gaussian primitives. Third, a cross-domain gating fusion mechanism adaptively combines the two complementary normal estimates while imposing multi-view reprojection consistency and implicit geometric regularization, reconciling the geometrically consistent yet over-smooth 3DGS estimate with the detailed yet potentially geometry-inconsistent refinement. Finally, all components are jointly optimized under a unified photometric reprojection objective with geometric consistency regularizations, and a directional regularization aligns the learnable principal directions with the geometric normals, achieving an end-to-end optimization closed loop without relying on external normal labels."

The method is trained only on images paired with their corresponding 3D models, without any pixel-level normal annotations. The paper reports: "Trained only on images paired with their corresponding 3D models, without any pixel-level normal annotations, our method achieves state-of-the-art performance among weakly supervised methods and delivers accuracy competitive with fully supervised baselines."

The paper identifies limitations of existing weakly supervised methods: "Discriminative methods learn an image-to-normal mapping under proxy supervision signals built from pseudo-labels, cross-domain transfer or auxiliary alignment... the estimate's geometric fidelity is capped by the signals' own accuracy. Generative methods reconstruct 3D shape from a learned prior and read out normals as a by-product... the result reflects the prior's statistics, and when the prior and the observed image disagree, the estimate follows the prior. Each route therefore supplies only one of the two things the task requires: a geometry prior that provides the initial estimate or an image-derived signal that verifies it, and neither supplies both."

The main contributions are summarized as: "A unified framework that combines feedforward geometric priors with differentiable photometric verification through an image-geometry-image consistency loop, providing a self-supervising mechanism for weakly supervised normal estimation that corrects both proxy inaccuracy and prior deviation without external normal labels"; A high-fidelity differentiable illumination rendering pipeline with a learnable modulation kernel that turns the photometric loss into an internal supervision signal under weak supervision; A single-step deterministic refinement network that recovers high-frequency detail while preserving end-to-end differentiability, without external label data; and "A cross-domain gating fusion mechanism with multi-view reprojection consistency and implicit geometric regularization that reconciles geometric consistency with recovered detail, without relying on external normal labels."

The paper addresses four obstacles when adapting 3DGS to single-view weak supervision: "standard parameterization lacks a stable analytical mapping from covariances to normals without normal ground truth; native rendering pipelines sacrifice gradient accuracy for rasterization efficiency, providing insufficient fidelity for normal-level optimization; the intrinsic smoothness of Gaussian primitives erases high-frequency details, yet mainstream detail enhancement depends on normal ground-truth supervision; and naively fusing normals from geometric priors with those from detail enhancement breaks global consistency without ground-truth constraints on fusion boundaries."

The method defines each Gaussian primitive as G = µ, Σ, σ, kD, ωg, ξ, ψ, d, where µ is the Gaussian center, Σ is the covariance matrix, σ is opacity, kD is RGB diffuse albedo, ωg, ξ, ψ control frequency, spatial decay and phase shift of local light interaction, and d is a learnable local principal direction. The covariance matrix is parameterized using rotation-scale decomposition Σ = RSR T to ensure symmetric positive definiteness. Surface normals are derived via eigen-decomposition of the covariance matrix, where the eigenvector umin corresponding to the smallest eigenvalue gives the principal normal direction of the local surface, with sign disambiguation enforced by n3dgs T v > 0.

The illumination model is grounded in the Lambertian reflectance framework, with a learnable oscillatory modulation kernel K(p) = exp(−ξ∥p − µ∥22) cos((2π/ωg)d T(p − µ) + ψ) that captures high-frequency residuals from micro-geometric variations. The paper explains: "the explicit oscillatory form prevents the kernel from degenerating into a texture-fitting function, since it cannot reproduce arbitrary image patterns, and the Gaussian envelope ensures the modulation is localized, preserving the global dominance of the Lambertian term."

The radiance contribution of a single Gaussian primitive is formulated as Ci(x) = kD,i ⊙ max(0, n T fuse,i l) · Ki(p(x)), where l is the light direction fixed to the camera optical axis for training. The final pixel radiance is obtained by blending contributions from all Gaussians through the 3DGS rasterization process with cumulative transmittance.

The deterministic refinement network reformulates diffusion models from probabilistic generators to structural refinement operators. The paper states: Given n3dgs, Gaussian noise is first injected at a random time step t... One-step denoising is then performed using a noise prediction network ϵθ with time embedding to obtain the refined normal estimate. The one-step mapping is described as a noise-conditioned residual refinement, where the network learns to correct the initial normal using local structure information.

The cross-domain fusion mechanism introduces complementary feature representations: "The geometric feature Fgeo denotes the geometry-aware representation derived from the 3DGS-based prediction module... The texture feature Ftex denotes the appearance-aware intermediate feature extracted from the diffusion refinement network." The fused feature is expressed as Ffuse = Wg ⊙ Fgeo + Wt ⊙ Ftex, with weights obtained via softmax normalization ensuring Wg + Wt = 1 element-wise. The final fused normal is defined as nfuse = g ⊙ sg(n3dgs) + (1 − g) ⊙ ndiff, where g is a spatially adaptive gating weight predicted by a lightweight convolutional network.

The total loss is Ltotal = Lphoto + λscale Lscale + λnormal Lnormal + λdir Ldir, where Lphoto is the photometric consistency loss, Lscale is a scale regularization term preventing Gaussian primitives from degenerating, Lnormal is a normal consistency constraint encouraging the refined normal to stay close to the geometric normal in regions where the gating weight is small, and Ldir is a directional regularization aligning the learnable principal direction with the geometric normal.

Experiments are conducted on the Objaverse dataset for training (80k diverse 3D object models) and evaluated on Google Scanned Objects, Omniobject3D, and Wonder3D benchmarks. The paper reports: "under the weakly-supervised paradigm (no normal labels), CLONE achieves the best performance across all metrics. CLONE's MAE of 13.2◦ exceeds the strongest fully-supervised baseline Metric3D v2 (14.1◦, trained with ground-truth normals) by 0.9◦, and the strongest weakly-supervised baseline FE2E (14.3◦) by 1.1◦."

Ablation studies confirm the necessity of each core component: Eliminating the 3DGS geometric initialization module results in the sharpest performance decline, where MAE rises from 13.2◦ to 15.7◦ and Acc@11.25◦ drops by 6.9%. Removing the learnable modulation kernel elevates MAE to 13.9◦, removing cross-domain alignment also increases MAE to 13.9◦, removing gating fusion increases MAE to 13.6◦, and removing diffusion refinement increases MAE to 13.4◦.

Robustness tests show the method maintains stability under additive Gaussian noise, lighting variations, and viewpoint changes. The paper states: our method maintains exceptional stability: MAE increases by only 6.0◦ from σnoise = 0 to σnoise = 15, whereas the most affected baseline, Marigold-Normals, suffers a 13.3◦ MAE surge over the same noise range.

The paper acknowledges limitations including texture-geometry confusion, occlusion and slender structures, and non-Lambertian materials. The paper states: high texture gradients can mislead the diffusion branch, and specular reflections give rise to systematic errors. Furthermore, severe self-occlusion remains intractable due to the absence of explicit occlusion modeling. The paper concludes: "CLONE provides a new weakly-supervised paradigm for single-image single-object normal estimation by closing the loop between image observations and 3D geometry. It achieves competitive performance against fully supervised methods on synthetic benchmarks and demonstrates strong robustness to noise, lighting and viewpoint changes."

Improvements for AI systems

Based on the paper, here are the specific improvements I can make to AI systems and what the improved system can do:

1. Implement a self-supervising normal estimation pipeline using an image-geometry-image consistency loop

  • Integrate 3D Gaussian Splatting (3DGS) as an explicit geometric representation with differentiable rendering to create a closed-loop optimization that verifies predicted geometry against the input image

  • Replace the need for pixel-level normal ground truth with photometric reprojection loss as the primary supervision signal

2. Add a differentiable light interaction model with a learnable oscillatory modulation kernel

  • Reparameterize the 3DGS parameter space to establish an explicit, stable mapping from Gaussian covariances to surface normals via eigen-decomposition

  • Use the kernel K(p) = exp(-ξp-μ2)cos((2π/ω g)dT(p-μ)+ψ) to capture high-frequency micro-geometry variations beyond Lambertian reflectance, turning photometric loss into an internal geometric supervision signal

3. Integrate a single-step deterministic refinement network (conditional denoising U-Net)

  • Reformulate diffusion models from probabilistic generators to structural refinement operators, avoiding gradient truncation inherent in multi-step DDPM

  • Inject noise into the initial 3DGS normal, perform one-step denoising conditioned on fused geometric and texture features, and backpropagate photometric loss directly through the network

4. Implement a cross-domain gating fusion mechanism

  • Fuse geometric normals (globally consistent but over-smooth) with refined normals (detail-rich but potentially geometry-inconsistent) using a spatially adaptive gating weight g ∈ [0,1]

  • Apply joint channel-spatial attention with softmax normalization to combine 3D geometric priors and 2D appearance features, with a stop-gradient operation on the geometric branch to preserve end-to-end differentiability

5. Add multi-view reprojection consistency and geometric regularizations

  • Use a scale regularization term L scale = Σ(λ min + γλ max) to prevent Gaussian primitive degeneration

  • Add a normal consistency loss L normal = n 3dgs - n diff1 ⊙ (1-g) that adaptively relaxes constraints in texture-rich regions

  • Include directional regularization L dir to align learnable principal directions with geometric normals

1. Estimate surface normals from a single RGB image without any pixel-level normal annotations

  • Achieve state-of-the-art performance among weakly supervised methods (MAE 13.2° vs. 14.3° for the best competing weakly supervised baseline)

  • Match or exceed fully supervised baselines (e.g., Metric3D v2 at 14.1° MAE) while using only RGB-3D model pairs for training

2. Recover high-frequency geometric details (thin structures, sharp edges, complex topologies) that pure 3DGS or generative methods erase

  • Maintain consistent normals on slender objects (shoes, cactus branches, bamboo) where competing methods produce broken or discontinuous normals

  • Accurately resolve part boundaries and junctions on complex objects (excavators, house models) without artifacts

3. Maintain robustness under challenging conditions

  • Degrade gracefully under additive Gaussian noise (MAE increases only 6.0° from σ=0 to σ=15, vs. 13.3° for the worst baseline)

  • Handle lighting variations (front, side, top) and viewpoint changes (0°, 45°, 70°) with minimal performance loss (MAE stays below 14.5° in all cases)

  • Tolerate common image degradations like motion blur and JPEG compression

4. Operate with label efficiency

  • Train on large-scale 3D asset libraries (e.g., Objaverse) without expensive per-pixel labeling

  • Achieve competitive accuracy with only image-3D model pairs, reducing annotation costs by orders of magnitude compared to fully supervised approaches

5. Provide a unified framework that corrects both proxy inaccuracy and prior deviation

  • The image-geometry-image loop ensures that errors in initial geometric estimates are corrected through photometric verification, avoiding the limitations of both pure discriminative and pure generative approaches

6. Achieve end-to-end differentiability with stable convergence

  • Maintain a Gradient Transfer Ratio of 0.98 (vs. 0.12 for multi-step DDPM), ensuring gradients flow effectively through the entire pipeline

  • Avoid training divergence (0% failure rate) through rotation-scale decomposition of covariance matrices, unlike direct covariance optimization which has 30% divergence risk

Abstract

We propose CLONE, a Continuous Latent Optimization framework for Normal Estimation via 3D Gaussian splatting. The core idea is to construct an image-geometry-image consistency loop that unifies explicit geometric representation with differentiable rendering, thereby enabling weakly supervised learning without normal ground truth. Specifically, CLONE comprises four components. First, by introducing a differentiable light interaction model with a learnable modulation kernel, we perform a unified reparameterization of the 3DGS parameter space, establishing an explicit and stable mapping between 3DGS geometric parameters and surface normals and turning the photometric loss into an internal supervision signal. Second, the conditional single-step deterministic refinement network integrates denoising architectures with differentiable reprojection constraints to refine the initial normals, thereby adaptively recovering the high-frequency details erased by the inherently smooth Gaussian primitives. Third, the cross-domain gating fusion mechanism adaptively combines the two complementary normal estimates while imposing multi-view reprojection consistency and implicit geometric regularization, reconciling the geometrically consistent yet over-smooth 3DGS estimate with the detailed yet potentially geometry-inconsistent refinement. Finally, all components are jointly optimized under a unified photometric reprojection objective with geometric consistency regularizations in a fully differentiable pathway, and the directional regularization aligns the learnable principal directions with the geometric normals, achieving an end-to-end optimization closed loop without relying on external normal labels.

Sources

Related papers