FujinSplat: Seeing Through Smoke with RAW-Domain Gaussian Splatting

summary

Video file (mp4)

The gist

FujinSplat introduces a RAW-domain Gaussian Splatting framework that enables novel-view synthesis through dense smoke from hazy observations, addressing the problem where participating media and

In short

FujinSplat is a RAW-domain Gaussian Splatting framework that recovers clean 3D scenes from hazy smoke observations. It separates medium attenuation from image processing effects by fitting a calibrated Base ISP and training a single scene-agnostic controller to predict view-dependent dehazing parameters. This outperforms other methods on real-world smoke benchmarks.

Key concepts

Base ISP (Bs)
This is a calibrated Image Signal Processor fitted to the scene's hazy RAW captures. It faithfully reproduces the camera coordinates and provides a fixed photometric anchor, effectively removing basic sensor and camera effects before dehazing begins.
Monotone Color Flow (MCF)
The color operator used to correct haze is modeled as an MCF. This combines channelwise monotone tone curves with volume-preserving color couplings. It allows the system to model how smoke affects different color channels while preserving the underlying scene structure.
Reverse ISP-Action Synthesis
Since clean reference views are unavailable, this technique generates synthetic training data. It inverts the calibrated Base ISP and then inverts the learned color operator to create a labeled synthetic RAW input, which is used to train the controller.

Terminology used across episodes

This episode discusses

The paper

FujinSplat: Seeing Through Smoke with RAW-Domain Gaussian Splatting · Read on arXiv

Gengjia Chang, Ziteng Cui, Shuhong Liu

Hefei University of Technology · The Hong Kong University of Science and Technology, Guangzhou · The University of Tokyo

Transcript

Introduction to the show: ident: AI Radio. Generated commentary on the latest Artificial Intelligence papers.

Tom: Today's paper: "FujinSplat: Seeing Through Smoke with RAW-Domain Gaussian Splatting".

Jane: FujinSplat introduces a RAW-domain Gaussian Splatting framework that enables novel-view synthesis through dense smoke from hazy observations,

Tom: First, who's behind it and why it matters.

Title and authors: Tom: So we're diving into the paper titled "FujinSplat: Seeing Through Smoke with RAW-Domain Gaussian Splatting." It looks like this work tackles a really tough problem in computer vision, which is recovering clean three dee scenes when they are filled with smoke or haze <ref:2609.06017#pg0>.

Jane: That sounds intense, Tom. I'm curious what exactly makes this approach different from the methods we usually see when dealing with smoky images. It seems like they're trying to untangle something that's been tangled up for years in standard image processing pipelines.

Lu: What really strikes me about this paper is their decision to operate entirely within the RAW domain instead of starting with sRGB images, which is a significant methodological choice, and this is what sets FujinSplat apart <ref:2609.06017#pg0>.

Meng: From an engineering standpoint, operating in the RAW domain suggests they are trying to get closer to the raw sensor data before any of the camera's color transformations happen, which sounds promising for fidelity.

Lalam: If this works well, Lu and Meng, I think it could fundamentally improve how we build cultural artifacts or simulations involving complex environmental degradation because we get a cleaner base signal <ref:2609.06017#pg1>.

Tom: Exactly, Lalam. The core idea is that the smoke effect is created by two things—the medium itself changing the scene radiance, and then the camera's image signal processor remapping that result through tone and color changes <ref:2609.06017#pg0>.

Jane: So, instead of applying a general sRGB dehazing filter after everything is already mixed together by the ISP, they are trying to separate those two processes before reconstruction starts <ref:2609.06017#pg1>.

Lu: Precisely, and the paper describes fitting a per-scene Base ISP first, which acts like a fixed photometric anchor that performs no dehazing because it accurately reproduces the camera coordinates <ref:2609.06017#pg2>.

Meng: That calibration step sounds crucial because if that base ISP is good at mapping coordinates, it gives us a reliable starting point before we even deal with the complex view-dependent color actions <ref:2609.06017#pg1>.

Lalam: And then they introduce this scene-agnostic controller to predict a parameter vector that characterizes a Monotone Color Flow, which is how they model those view-dependent transformations <ref:2609.06017#pg1>.

Tom: That controller predicting the color operator via a Monotone Color Flow, decomposed into tone and coupling parameters, is where they handle the variation across different views in a clever way <ref:2609.06017#pg2>.

Jane: It sounds like they are not fitting a new correction for every single view but instead learning one general rule that applies to the whole scene's color effects, which simplifies things immensely for training <ref:2609.06017#pg1>.

Lu: And the training process uses something called "Reverse ISP-Action Synthesis" to generate synthetic RAW inputs, which is a smart way to get those exact labels needed to train the controller effectively <ref:2609.06017#pg2>.

Meng: That reverse synthesis sounds computationally heavy, so how do they manage that process efficiently enough for training without it taking forever?

Title and authors: Lalam: Lalam thinks the efficiency comes from using a lightweight convolutional encoder and MLP for the prediction step, which keeps the model structure manageable while still learning complex transformations <ref:2609.06017#pg2>.

Tom: Speaking of complexity, they also introduce a "bounded per-view residual" handled by a-ISP to manage any remaining photometric disagreement between the view actions and the static three dee Gaussian representation <ref:2609.06017#pg2>.

Jane: So, if the main controller handles most of it, this residual term is just there to clean up any small errors that still exist between views, ensuring the final three deeGS stays consistent <ref:2609.06017#pg2>.

Lu: That decomposition into a scene-shared action and a per-view displacement is how they reconcile those cross-view inconsistencies without messing up the actual geometry of the static scene <ref:2609.06017#pg2>.

Meng: That sounds like a necessary step for practical application, because if the three dee reconstruction drifts from one view to another, it's useless for building anything solid <ref:2609.06017#pg2>.

Lalam: For me, this consistency is vital because it means we can trust the geometry of the scene even when we are synthesizing a novel view that hasn't been explicitly seen during training <ref:2609.06017#pg2>.

Tom: The results are quite compelling; they claim FujinSplat clearly outperforms both physics-based reconstruction and restoration-then-three deeGS pipelines on the RealXthree dee smoke benchmark <ref:2609.06017#pg0>.

Jane: That is a big claim, Tom. They reported a PSNR of eighteen point four two dB averaged over eight scenes, which they said was two point five seven dB better than the strongest comparable baseline <ref:2609.06017#pg0>.

Lu: The ablation studies also show that using RAW measurements provides a tangible benefit; for instance, replacing RAW input with standard RGB input causes the novel-view PSNR to drop from eighteen point four two dB down to eighteen point two nine dB <ref:2609.06017#pg0>.

Meng: That quantitative proof regarding the RAW representation is what I'm paying attention to, because it shows that the data itself is fundamentally better for this type of restoration task <ref:2609.06017#pg2>.

Lalam: And they also found that the color operator is already quite expressive, noting that it requires only five hundred seventy-three coefficients per view to capture what's necessary <ref:2609.06017#pg2>.

Tom: That coefficient count seems surprisingly manageable, especially when you consider the complexity of modeling a whole color transformation through a Monotone Color Flow <ref:2609.06017#pg1>.

Jane: It suggests that the underlying physics of how smoke interacts with the camera sensor can be captured by a relatively compact set of mathematical operations rather than needing an enormous, complex model <ref:2609.06017#pg2>.

Lu: And on a broader level, this approach moves away from relying on external priors or massive generative models for every training view, which is a significant step in making the system more self-contained <ref:2609.06017#pg2>.

Meng: From a practical deployment viewpoint, if we can train this controller once and then use it for many scenes, that drastically reduces the need to run complex inference pipelines every time we want a new view <ref:2609.06017#pg2>.

Title and authors: Lalam: I think the cultural implication here is that AI tools designed for visualization and scene understanding can become much more robust in real-world, messy environments, moving beyond clean datasets into scenarios that look like actual accidents or industrial settings <ref:2609.06017#pg1>.

Tom: So we've seen how they tackle the entanglement of medium attenuation and ISP effects by separating them into a calibrated base ISP and a learned view-dependent controller in this FujinSplat work.

Jane: And the paper suggests that the method is highly effective because it leverages RAW data to learn complex, view-dependent color transformations efficiently through synthesis <ref:2609.06017#pg2>.

Lu: The structure involving a fixed Base ISP and a scene-agnostic controller trained via reverse synthesis is the main technical innovation described in this paper <ref:2609.06017#pg2>.

Meng: I'm interested in the limitation they mentioned regarding texture recovery; they explicitly state that the correction is a global color action, so textures destroyed by smoke simply cannot be recovered <ref:2609.06017#pg2>.

Lalam: That is a fair caveat, Meng. It means if we are reconstructing something highly detailed where the smoke has completely obliterated fine surface details, this method won't bring them back into view <ref:2609.06017#pg2>.

Tom: So, to wrap up the core of what FujinSplat does, it uses a RAW-domain Gaussian Splatting framework to recover clean three dee scenes from hazy observations by separating medium attenuation from ISP tone and color transformations <ref:2609.06017#pg0>.

Jane: And the overall implication is that this method offers superior quantitative performance on real-world smoke benchmarks, outperforming existing physics-based and restoration pipelines <ref:2609.06017#pg0>.

Lu: This work demonstrates a path toward creating more robust three dee scene reconstruction methods that don't get lost in the complexities of camera sensors and atmospheric effects <ref:2609.06017#pg1>.

Meng: For practical applications, the ability to generate novel views reliably from challenging, real-world degraded data is where I see the most immediate utility <ref:2609.06017#pg2>.

Lalam: Ultimately, FujinSplat shows that separating the camera's rendering process from the physical medium effects using RAW inputs allows for a much cleaner three dee representation of reality <ref:2609.06017#pg1>.

Tom: And that brings us to the end of our discussion on this paper. It’s been really fascinating to see how they've managed the challenge of entangled medium and ISP effects with this RAW-domain Gaussian Splatting approach <ref:2609.06017#pg0>.

Jane: It certainly was a deep dive into how we can disentangle those two major factors in image reconstruction, moving beyond standard sRGB pipelines <ref:2609.06017#pg1>.

Lu: I think the combination of the scene-agnostic controller and the bounded residual mechanism provides a really solid framework for handling view-dependent variations reliably <ref:2609.06017#pg2>.

Meng: We’ve discussed how promising it is in terms of performance, but we also noted its limitation regarding texture recovery, which means we have to be careful about the kind of data we feed it <ref:2609.06017#pg2>.

Lalam: So, FujinSplat isn't a complete solution for every single visual fidelity problem, but it is a strong step forward in building more resilient AI tools for three dee scene understanding and novel view synthesis <ref:2609.06017#pg1>.

The paper's summary: Tom: So, to recap, FujinSplat is this new Gaussian Splatting method that takes raw sensor data to reconstruct clean scenes by cleverly separating the camera's imaging distortions from the physical smoke in a way that standard sRGB pipelines just can't handle.

Jane: That’s a really cool way to put it, Tom. Basically, instead of trying to fix the final picture after everything has been mixed up by the camera's processing, they are working at the source level with RAW data and learning how to undo those specific distortions before building the three dee model.

Lu: The real magic here is in how they’ve structured that correction; fitting a fixed base ISP for every scene and then training a single controller that learns the view-dependent color adjustments through synthetic examples. It's like teaching an AI to be an expert at undoing a specific camera's messy lens and sensor behavior, scene by scene.

Meng: From an engineering standpoint, that separation into a fixed component and a learned variable makes it much more manageable for deployment because you only need to handle the fixed part once and then use the controller parameters for each new view. I’m interested in how stable that controller learning is when dealing with wildly different smoke conditions.

Lalam: And from an information perspective, this moves us toward a future where AI can reliably interpret visual data under extreme environmental degradation, like heavy smog or dense fog, which could have huge implications for disaster response visualization. It's about giving AI the ability to see through real-world noise instead of just idealized digital scenes.

Tom: Exactly! And the quantitative results are pretty impressive; they showed it actually beat established methods on a challenging smoke benchmark, proving that this RAW-domain approach is objectively better than physics-based reconstruction techniques right now.

Jane: That’s a strong result, Tom. It shows that by focusing on the raw physical signal and using a compact mathematical model like the Monotone Color Flow to describe transformations, they found a way to achieve better fidelity than what we were seeing before.

Lu: I think the efficiency comes from that five hundred seventy-three coefficient limit they found; it means you don't need a massive, computationally expensive model to capture all the necessary view-dependent changes in smoke. That compactness is what makes it feasible for real-world use.

Meng: That compactness is key for my startup because if we can get a robust correction mechanism that doesn't require training on millions of specific synthetic datasets, the cost of adapting this technology across different camera types drops significantly. I just need to know if that forty-two minutes per scene reconstruction time is fast enough for our application needs.

Lalam: The cultural impact here is huge because it means visualization tools won't be limited to clean, studio-quality renders; they can accurately model and understand the visual reality of messy, real-world environments where things get obscured by smoke.

Tom: We’ve definitely seen that potential, Lalam. It’s moving us past just pretty pictures into actual scene understanding in difficult conditions. Now that we’ve covered the core mechanics and those exciting initial results, we should probably talk about what this means for the next phase of development and where these researchers are heading with this work.

The paper's improvements: Tom: So, we’re looking at how FujinSplat suggests ways to make this framework even better than it already is, and it’s all about boosting fidelity and making things faster for real use.

Jane: It seems they are focusing on refining the way the system handles view-dependent inconsistencies and optimizing its performance during the rendering stage to get that real-time speed we talked about earlier.

Lu: One big suggestion is to focus on how they handle those residual errors between views, which they call a bounded per-view displacement; it’s about making sure that even if two different training views don't align perfectly, the final three dee reconstruction remains geometrically stable.

Meng: That makes sense from an engineering standpoint because stability is everything for three dee assets. If we can ensure the geometry doesn't drift across views, we can use it reliably in simulations or AR applications without worrying about artifacts creeping in.

Lalam: For me, the emphasis on reconciling cross-view inconsistencies means that AI-generated environments will become much more believable because they won’t have those weird "flickering" or inconsistent geometry issues when you look at them from different angles. It improves the sense of reality in digital spaces.

Tom: And another important improvement mentioned is how they keep the system light for inference; they want to ensure that all these complex calculations are done during training and then frozen, so we can render novel views super quickly without any extra prediction steps at runtime.

Jane: That’s a practical win because rendering speed directly impacts user experience; if the process takes too long, people won't use it often, so making the inference part lean is a smart move for adoption.

Lu: They also talked about how they can systematically analyze the parameters themselves to find out which parts of that color correction are actually doing most of the heavy lifting across different scenes; that’s giving us a better understanding of what those five hundred seventy-three coefficients are actually controlling.

Meng: That analytical capability is valuable because it lets us prune unnecessary complexity later, perhaps leading to even leaner versions for specific deployment targets or lower-power hardware constraints.

Lalam: On a broader level, this systematic analysis helps the AI understand the underlying physics of how smoke affects color in different contexts; this knowledge could eventually be applied to other complex visual problems outside of just smoke scenes.

Tom: So, we’ve seen ways to stabilize geometry and speed up rendering, which are huge steps toward making this a viable tool for professional visualization rather than just a research curiosity.

Conclusion: Tom: Alright team, we've covered a lot about FujinSplat: Seeing Through Smoke with RAW-Domain Gaussian Splatting, and to wrap up, this paper shows how separating the scene's physical medium from the camera’s image processing allows for much cleaner three dee reconstruction when dealing with hazy observations.

Jane: It really is impressive how they tackle that entanglement by using a calibrated Base ISP alongside a learned controller to handle the view-dependent color changes directly from RAW data.

Lu: The structural elegance of decomposing the correction into a fixed component and a scene-agnostic parameter vector is what makes this approach so flexible for novel applications in three dee reconstruction.

Meng: From an engineering standpoint, the ability to freeze those parameters after training means we can deploy this on hardware without needing constant re-optimization, which drastically cuts down on our operational costs for large-scale scene processing.

Lalam: I think the biggest impact here is that it gives AI tools a much more accurate way to perceive and reconstruct reality, whether that's for cultural preservation or understanding complex physical scenarios. It’s about building smarter visual models for the real world.

Tom: Exactly, Lalam. The quantitative results on the RealXthree dee benchmark really back up the claim that this method is a solid performer compared to existing pipelines in this area.

Jane: And we saw how they are already making it more robust by introducing that bounded residual term to keep the static three dee Gaussian representation consistent across different training views.

Lu: That reconciliation mechanism is what gives them confidence in the final output, ensuring that the geometry stays true even when the view-dependent actions have some noise.

Meng: I’m still looking at how they handle generalization; if we can show that this learned controller works well on cameras we haven't seen during training, then the practical utility skyrockets.

Lalam: That future work on generalization is critical because it means these systems won't just be good for the specific scenes they were trained on, but will actually be useful in diverse, messy real-world situations.

Tom: So we’ve seen how FujinSplat moves toward more reliable scene understanding by leveraging RAW domain knowledge to disentangle sensor noise from physical effects.

Jane: It’s been a deep dive into separating the medium attenuation from the ISP tone and color transformations using raw data, which is a really sophisticated method.

Lu: The core innovation of this work lies in that clever decomposition strategy, fitting the Base ISP first and then learning the view-dependent operator through synthesis.

Meng: For us at our startup, it means we have a clear path forward for building more robust AI systems for three dee reconstruction under challenging visual conditions.

Lalam: Ultimately, FujinSplat shows that by treating raw sensor data as the source of truth, we can build AI that is far better equipped to model and interpret complex visual information in our world.

Tom: That concludes our discussion on FujinSplat: Seeing Through Smoke with RAW-Domain Gaussian Splatting. It’s been a fascinating look at how we can push the boundaries of three dee scene reconstruction with this technique.

More episodes

← Home