Point-MF: Stabilizing One-Step Mean Flows for Single-Image Point Cloud Reconstruction

arXiv:2604.24586 · cs.CV · Submitted 2026-04-27 · Read on arXiv

Listen

Radio episode about this paper

Transcript

Introduction to the show: ident: AI Radio. Generated commentary on the latest Artificial Intelligence papers.

Tom: Today's paper: "Point-MF: Stabilizing One-Step Mean Flows for Single-Image Point Cloud Reconstruction".

Jane: Single-image point cloud reconstruction requires inferring complete 3D geometry, including occluded parts, from only visible content in a single RGB image, and this paper introduces Point-MF,

Tom: First, who's behind it and why it matters.

Title and authors: Tom: Moving on to the core concept of the paper, "Point-MF: Stabilizing One-Step Mean Flows for Single-Image Point Cloud Reconstruction," they introduce a novel way to generate complete three dee geometry from a single RGB image without needing those many iterative denoising steps typical in diffusion methods <ref:2604.24586#pg0,Single-Image Point Cloud Reconstruction>.

Jane: They are essentially proposing that instead of slowly refining the shape over many small steps, you can learn the average movement, or mean velocity field, between two points in time directly.

Lu: This mean velocity field formulation is what lets them achieve one-step reconstruction with a single network function evaluation, which is a big improvement over existing methods that require numerous iterations.

Meng: So they're skipping the long sequence of denoising operations and jumping straight to the final shape estimate using this learned flow, which cuts down on inference time considerably, I assume?

Lalam: It’s about making the generation process direct; they skip those intermediate steps that usually consume a lot of computation when you're trying to build something complex like a three dee scene <ref:2604.24586#pg0>.

Tom: Right; and to address the potential instability that comes with these large jumps, they introduce an auxiliary loss based on a set distance defined on the implied denoised estimate in data space.

Jane: That auxiliary loss is crucial because it’s designed not to look at the velocity fields themselves, but rather enforces geometric consistency directly on the predicted point cloud during generation.

Lu: This stability mechanism is what allows Point-MF to operate effectively in low-step generation over point cloud space, even when making those large interval jumps that could otherwise lead to outliers.

Meng: So instead of relying solely on the network learning a smooth flow, they’ve added an extra constraint that pulls the resulting geometry back toward what should be geometrically sound.

The paper's summary: Tom: To summarize, Point-MF is introducing a Mean-Flow-based framework specifically designed for low Network Function Evaluations for single-image point cloud reconstruction, which operates directly in point cloud space.

Jane: Essentially, they’ve combined a specialized Diffusion Transformer architecture with a specific conditioning strategy based on frozen DINOv3 features and explicit time and interval embeddings.

Lu: The paper proposes that by formulating the generation process this way, they can learn the mean velocity field over an interval

r, t: , which leads to one-step reconstruction with a single network function evaluation.

Meng: It’s about replacing the complex iterative denoising pipeline with a direct regression of this learned flow, which is what makes it faster than traditional diffusion-based baselines like BDM.

Lalam: This directly implies that we can build three dee scenes from images much more quickly, potentially enabling applications where speed and efficiency are paramount <ref:2604.24586#pg0>.

Tom: And they’re not just stopping there; they’ve also formulated the training objective as LMF-CFG to incorporate Classifier-Free Guidance without increasing the sampling Network Function Evaluations.

Jane: That means we can still leverage conditional generation—guiding the output based on some input condition—while keeping that speed advantage intact during inference.

Lu: The regression target they define is specific: "utgt:= ˜vt − (t − r) d/dt uθ(xt, r, t c)." This mathematical formulation is central to how they train the model to predict the correct velocity field.

The paper's improvements: Tom: Now let’s talk about what they actually improved in terms of methodology; their main technical contribution is introducing that auxiliary set-distance-based loss to suppress shape deviations during large interval jumps.

Jane: That loss works by defining a set distance on the implied denoised estimate, which promotes geometric consistency right in data space, rather than just penalizing errors in the velocity fields.

Lu: This stabilizes generation when dealing with those large interval jumps, which is something that has historically caused deviations from the underlying shape manifold and created outliers during few-step sampling.

Meng: The paper also introduced a lightweight Post-MHSA Adapter for image patch features fed into the Cross-Attention to adapt them specifically for point cloud generation tasks.

Lalam: That adapter is smart because it helps take the general visual features from DINOv3 and make them more relevant and effective when they are being used to predict three dee points <ref:2604.24586#pg0>.

Tom: And they’ve also adopted the Adaptive Probabilistic Matching Loss, which they use as a set distance metric for DSA because it provides a soft one-to-one correspondence between the predicted and ground-truth point clouds.

Jane: That matching loss is chosen over simpler metrics like Chamfer Distance because it promotes a more stable global shape alignment when fitting the reconstructed point cloud to the ground truth geometry.

Conclusion: Tom: So, wrapping up this discussion on "Point-MF: Stabilizing One-Step Mean Flows for Single-Image Point Cloud Reconstruction," the main implication is that we have a method for reconstruction that is significantly faster than multi-step diffusion methods while maintaining good quality.

Jane: They achieve this by directly learning the mean velocity field and using a specific auxiliary loss to keep the geometry stable during those single, large updates.

Lu: The work shows that it’s possible to balance reconstruction quality with geometric stability when dealing with point cloud generation from a single image without needing complex VAE latent representations for conditioning.

Meng: Practically, this means we could see faster three dee scene generation in areas like robotics or AR where low latency is critical, provided the computational requirements are manageable on consumer hardware <ref:2604.24586#pg0>.

Lalam: For the AI culture, this points toward a future where high-fidelity three dee content creation becomes much more accessible because it reduces the reliance on slow and expensive iterative processes <ref:2604.24586#pg0>.

Tom: It’s definitely a solid piece of work that shows how focusing on flow dynamics in point cloud space can yield efficient results when paired with the right regularization techniques.

Jane: I agree; Point-MF offers a clear path forward for anyone looking to build generative models that handle three dee structure from 2D images more efficiently <ref:2604.24586#pg1>.

Lu: Indeed, the synergy between the Mean Flow formulation and the data space supervision is what makes this approach so interesting for future work.

Meng: I'm just thinking about scaling this; if we can get it to run reliably on more complex scenes, that’s where we’ll see its real practical value.

The University of Electro-Communications

cs.CV

Submitted: 2026-04-27

Updated: 2026-10-02

Importance score: 90/100

The gist: Single-image point cloud reconstruction requires inferring complete 3D geometry, including occluded parts, from only visible content in a single RGB image, and this paper introduces Point-MF, a

Key concepts

Mean Flow (MF)
This is the core mechanism that replaces traditional iterative denoising. It directly learns the average velocity field between two points in time. By parameterizing this field over an interval, Point-MF can perform a full reconstruction in just one network evaluation, significantly reducing computational cost compared to multi-step methods.
LMF-CFG
This is the training objective used to train the model. It incorporates Classifier-Free Guidance (CFG), which helps guide the generation towards the desired output quality, without requiring extra sampling steps during inference. It averages a CFG-guided velocity field over time to define what the network should learn.
Denoised Space Anchor (DSA)
This is an auxiliary loss function designed to keep the reconstructed geometry stable, especially when making large jumps in time. It enforces consistency by matching the predicted point cloud's extrapolation back to the true ground-truth shape, ensuring the generated points stay on a valid geometric manifold.
Diffusion Transformer (DiT) Backbone
The network architecture uses a Diffusion Transformer structure. This consists of Self-Attention for capturing local and global point geometry and Cross-Attention to align the point features with visual features extracted from the input image patches, ensuring rich contextual understanding.

Terminology

Summary

Single-image point cloud reconstruction requires inferring complete 3D geometry, including occluded parts, from only visible content in a single RGB image, and this paper introduces Point-MF, a Mean-Flow-based framework that enables lowNFE single-image point cloud reconstruction by formulating the generation process directly in point cloud space without relying on VAE-based latent representations.

How it works

The core of Point-MF is a Mean Flow (MF) formulation designed to reduce the number of Network Function Evaluations (NFEs) required for generation. Unlike traditional diffusion-based methods that require many denoising iterations, Point-MF directly learns the mean velocity field between two time points, allowing for one-step reconstruction with a single network function evaluation (1-NFE). This is achieved by parameterizing the mean velocity field over an interval [r, t] using a neural network, which enables large temporal jumps in a single update.

The training objective is formulated as LMF-CFG to incorporate Classifier-Free Guidance (CFG) without increasing sampling NFEs. The conditional mean velocity field, denoted as ucfg(xt, r, t c), is derived by averaging the CFG-guided instantaneous velocity field over the interval [r, t]. The regression target is then defined as:

utgt:= ˜vt − (t − r) d/dt uθ(xt, r, t c).

The network architecture employs a Diffusion Transformer (DiT) backbone tailored for this setting. This includes:

  1. Image Conditioning with DINOv3: A frozen DINOv3 encoder extracts a global feature vector and patch-sequence features. The global feature is combined with the time embedding t and interval-length embedding dt = t − r, and this conditioning vector is used via AdaLN-Zero to modulate features in each DiT block.

  2. Diffusion Transformer Architecture: Each block consists of Self-Attention (to capture local geometry and global structure) and Cross-Attention (to align point features with image patch features).

  3. Time and Image Condition: Both the current time t and the update interval length dt are explicitly provided as inputs, embedded using sinusoidal embeddings. The final conditioning vector is defined as c = et(t) + edt(dt) + eimg.

Key Components and Stabilizing Losses

Point-MF incorporates several mechanisms to ensure stability and geometric consistency during training:

  1. Image Conditioning with DINOv3: This provides robust visual representations, using the frozen DINO encoder to extract a global feature (zimg) and patch-sequence features (Zctx). A lightweight Post-MHSA Adapter is introduced for the image patch features fed into Cross-Attention to adapt them for point cloud generation.

  2. Denoised Space Anchor (DSA): To mitigate deviations from the underlying shape manifold during large interval jumps, an auxiliary loss is introduced in data space. The reconstructed point cloud xθ(xt, r, t c) is extrapolated using the Mean Flow update: xθ(xt, r, t c) = xt − t uθ(xt, r, t c). The DSA loss enforces geometric consistency by matching this estimate to the ground-truth point cloud xGT0.

  3. Adaptive Probabilistic Matching Loss (APML): As the set distance metric for DSA, APML is adopted because it provides a soft one-to-one correspondence between the predicted and ground-truth point clouds, promoting more stable global shape alignment than Chamfer Distance.

Experimental Validation and Results

The method was evaluated on ShapeNet [5] (ShapeNet-R2N2 split) and Pix3D [61]. Reconstruction quality was assessed using L1 Chamfer Distance (CD) and Earth Mover’s Distance (EMD).

Point-MF outperforms diffusion-based methods that require multi-step sampling, despite using only 1-NFE.

The results on ShapeNetR2N2 showed that Point-MF achieved the best overall F-Score, demonstrating improved coverage and density consistency. On the real-world Pix3D dataset, Point-MF achieved consistently strong EMD scores for challenging structures like the underside of tables and complex chair legs.

Practical Performance

The efficiency of Point-MF is highlighted by its inference speed.

Inference is performed on a single NVIDIA RTX A4000 GPU; after warm-up, each method is run 100 times on the same input, and the mean and standard deviation are reported.

Point-MF achieves a runtime of 63.45 ms/sample with only 1-NFE, which is substantially faster than multi-step diffusion baselines (e.g., BDM at 28000 ms/sample). Peak VRAM usage was reported at 1.

Improvements for AI systems

Here are specific improvements that can be made to AI systems by implementing the Point-MF framework, along with what those improved systems will be capable of:


Improvement 1: Development of a Single-Step, Latent-Free Point Cloud Generator (Point-MF)

By implementing the Point-MF architecture, AI systems can move from slow, iterative reconstruction methods to highly efficient, single-step generation. This system will eliminate the need for complex Variational Autoencoder (VAE) latent representations and explicit camera pose estimation during inference.

Specific capabilities:

  1. Reconstruct complete 3D geometry (including occluded parts) directly from a single RGB image in one network function evaluation (1-NFE).

  2. Achieve millisecond-level inference latency, making it suitable for real-time applications like autonomous navigation and interactive AR/VR environments.

Improvement 2: Enhanced Geometric Consistency via Data Space Supervision (Denoised Space Anchor - DSA)

The introduction of the set-distance auxiliary loss (DSA), which enforces geometric consistency on the denoised estimate in data space, directly addresses the common issue of shape deviation during fast generation steps.

Specific capabilities:

  1. Mitigate large interval jump artifacts and outliers that plague few-step diffusion models, resulting in reconstructions with superior global shape alignment.

  2. Improve the fidelity of complex structures (e.g., undersides of tables, intricate chair legs) by pulling the predicted point cloud toward the ground-truth geometry using an Adaptive Probabilistic Matching Loss (APML).

Improvement 3: Robust Conditional Generation via Mean Flow and Classifier-Free Guidance (CFG)

By leveraging Mean Flow to predict the interval-averaged mean velocity field, and incorporating CFG into the training objective, the system gains a powerful mechanism to steer generation based on specific image conditions.

Specific capabilities:

  1. Enable high-quality conditional reconstruction where the output point cloud is accurately guided by an input image condition (e.g., text prompts or object classes) without relying solely on iterative denoising steps.

  2. Maintain strong performance even when sampling is constrained to a single step, ensuring that the generated geometry adheres precisely to the desired conditioning signal.

Improvement 4: Efficient Feature Conditioning and Adaptation (DINOv3 + Post-MHSA Adapter)

The integration of frozen DINOv3 features with a lightweight Post-MHSA Adapter allows for robust, scalable, and condition-aware feature extraction without incurring the computational overhead of full transformer processing at every layer.

Specific capabilities:

  1. Ensure that the point cloud generation process is highly sensitive to visual cues in the input image (viewpoint, texture) while maintaining optimization stability through AdaLN-Zero modulation.

  2. Allow for efficient adaptation of pre-trained image features specifically for point cloud tasks, improving generalization across different object categories in ShapeNet and Pix3D datasets.

Improvement 5: Scalable and Versatile Training Objective Optimization

The training objective is designed to balance the learned mean-velocity mechanism (LMF-CFG) with geometric regularization (LDSA), using a carefully scaled loss coefficient that adapts based on the error magnitude.

Specific capabilities:

  1. Ensure that the model learns both accurate flow dynamics and stable point geometry simultaneously, preventing either from dominating during training.

  2. Provide a mechanism to tune the trade-off between reconstruction accuracy (CD/EMD) and geometric stability (DSA weight), allowing researchers to optimize for specific downstream requirements (e.g., prioritizing global correspondence vs. local surface alignment).

Sources

Related papers