CADReasoner: Iterative Program Editing for CAD Reverse Engineering

arXiv:2603.29847 · cs.GR, cs.CV, cs.HC · Submitted 2026-02-18 · Read on arXiv

Listen

Radio episode about this paper

Transcript

Introduction to the show: ident: AI Radio. Generated commentary on the latest Artificial Intelligence papers.

Tom: Today's paper: "CADReasoner: Iterative Program Editing for CAD Reverse Engineering".

Jane: CADReasoner introduces a novel framework for AI-based CAD reconstruction that performs iterative inference and self-correction, enabling progressively refined alignment between CAD models and the input.

Tom: First, who's behind it and why it matters.

Paper summary: Tom: Hey everyone, so we're talking about this paper today, "CADReasoner: Iterative Program Editing for CAD Reverse Engineering." It seems like they've tackled a really tricky problem in AI that involves taking a scan or an image and turning it back into a usable three dee design <ref:2603.29847#pg0>.

Jane: That sounds intense, Tom. So basically, the core idea here is that instead of just trying to guess the final shape in one go, this new framework lets the AI try something, see how far off it is from what we have, and then adjust its prediction based on that difference.

Lu: Exactly! It moves away from those single-shot predictions where things can get really inaccurate when dealing with complex geometry Lu. It mimics how a human engineer works: comparing what they see with the result and making small edits until it looks right.

Meng: So, if I'm tracking this for my practical side, the big claim seems to be that this iterative refinement leads to much better alignment between the reconstructed CAD model and the actual input data Meng. It’s not just about getting a pretty picture; it’s about achieving a level of accuracy that mimics real engineering processes.

Lalam: From an AI perspective, this closed-loop program editing process is really interesting because it allows the model to use its own output as supervision for its next step, which is a sophisticated way to train the system Lalam. It treats self-editing like a minimal implementation of the reasoning loop an engineer performs when inspecting and revising a design Lalam.

Tom: Right, so it’s this feedback loop that’s the main selling point, using geometric discrepancies to drive updates Tom. It seems they’re focusing on how this architecture handles different types of input data too, fusing multi-view renders and point clouds together for better context Tom.

Jane: And I think that fusion is crucial because visual information alone or just point clouds might miss details that the other modality provides.

Lu: They specifically designed the architecture to take three inputs per iteration: a multiview image, a point-cloud branch with projected features, and crucially, the previous program tokenized as text which acts like geometry memory for the decoder Lu. That geometric memory conditioning is what I find really creative for maintaining coherence across those iterative steps.

Meng: From an engineering standpoint, getting that architecture to run efficiently while handling both visual and point cloud data streams sounds like a big hurdle Meng. How they manage the tokenization of the previous program to condition the decoder seems like a complex part of their implementation.

Lalam: The structure where executing the script yields a new mesh from which modalities are re-extracted, effectively closing that self-correction loop, is what really makes this framework compelling for improving how we build and interpret three dee models Lalam <ref:2603.29847#pg0>. It shows how we can create systems that learn through direct interaction with the output.

Paper summary: Tom: So, it’s a cycle: input evidence leads to a program prediction, which is executed to make a new prediction, and then the results are fed back in to refine the program again Tom. That iterative inference is what sets this CADReasoner framework apart from simpler single-pass methods Tom.

Jane: It really highlights how modeling real-world reverse engineering practices can lead to more reliable reconstructions because it’s not just a guess, but a guided refinement process Jane. This approach seems designed to ensure that the model actually learns the nuances of geometric error correction.

Lu: And they took an extra step by introducing a scan-simulation protocol applied during both training and evaluation to make sure the editor learns to operate on inputs that look like real scans Lu. That’s important for ensuring realism in the final output, as it tests if the editor can handle noise and other artifacts present in actual scanning conditions.

Meng: Testing under scan-like conditions is essential because if it only works on perfect, clean data, its practical use is severely limited when we deal with real-world acquired parts Meng. It suggests they are tackling a significant challenge in making these systems usable outside of controlled lab settings.

Lalam: The fact that they train this editor using a curriculum designed to prevent overfitting by mixing training examples from different stages, like t=one and t=two examples, shows a thoughtful strategy for ensuring the model learns the refinement process properly Lalam <ref:2603.29847#pg0>. It’s about supervising it on-policy as it develops its skills.

Tom: So we’ve seen how they structure their training to encourage this self-correction loop through staged learning and curriculum design Tom. It seems like they are really focused on making the refinement step robust before deploying the final system.

Jane: And when we look at the overall performance metrics, especially comparing it against other methods, it shows significant improvements across different datasets like DeepCAD and Fusion360 Jane. The numbers suggest that after five iterations, it gets pretty close to achieving top results on those tasks.

Lu: The specific comparison in Table one shows how CADReasoner's performance at t=five is actually competitive with other methods like CADCrafter, especially when looking at the IoU and IR metrics Lu <ref:2603.29847#pg2>. That level of performance on different benchmarks is quite impressive for this type of iterative method.

Meng: So, while the performance numbers are encouraging, I still have to consider the complexity involved in running that entire loop repeatedly to achieve those results Meng. How computationally expensive is this iterative process in a real application?

Lalam: The architecture itself involves multiple modalities and a decoder that generates code at each step, so the computational load is definitely substantial Lalam. However, the goal seems to be achieving high fidelity through this structured loop rather than just brute-forcing a single complex prediction.

Paper summary: Tom: It really seems like they've found a way to balance the need for deep geometric reasoning with a practical iterative structure that can actually be trained effectively Tom. It’s moving past just having big models and trying to make them perform specific, verifiable engineering tasks.

Jane: And looking at the title, "CADReasoner: Iterative Program Editing for CAD Reverse Engineering," it perfectly describes what they've achieved—it’s about using that iterative editing process specifically for the hard task of reverse engineering a CAD model Jane. It grounds the abstract idea in a very concrete application.

Lu: The implication here is that we can move closer to scenarios where AI doesn't just generate a static three dee shape, but actively participates in the design revision process like an assistant would Lu <ref:2603.29847#pg0>. Imagine an engineer saying, "Can you try adding this feature and see how it looks?" and the AI handles the reconstruction update itself.

Meng: If we can reliably use this for iterative design assistance, that’s a huge practical impact on manufacturing workflows Meng. It suggests a level of interaction that goes beyond just generating a final output file.

Lalam: For our culture here, this work emphasizes building systems capable of self-correction and learning from their own mistakes in complex tasks, which is a powerful lesson for how we structure future AI development Lalam. It pushes us toward more robust, interactive AI capabilities.

Tom: So to wrap up these initial thoughts on "CADReasoner: Iterative Program Editing for CAD Reverse Engineering," it’s a framework that uses geometry feedback in a self-editing loop to iteratively improve three dee model reconstruction, even under simulated scan conditions Tom <ref:2603.29847#pg0,CADReasoner: Iterative Program Editing for CAD Reverse Engineering>. We've discussed how they handle the fusion of visual and point cloud data and the training curriculum they use to keep the refinement process stable.

Jane: And we’ve touched on how this methodology aims to bridge the gap between theoretical AI capabilities and practical engineering needs through this focused, iterative approach Jane. It seems like a solid piece of work for anyone looking at improving how we use AI in design and reconstruction tasks.

Lu: The long-term potential is that this iterative reasoning capability could eventually be applied to much broader domains where continuous refinement based on error feedback is necessary Lu. It’s about embedding the spirit of expert human revision into the AI's core operation.

Meng: I still think the real test will be in deploying something like this reliably in a factory setting where tolerances are extremely tight and failure isn't an option Meng. That kind of real-world validation is going to be crucial for its adoption.

Lalam: Ultimately, the advance here is demonstrating that sophisticated reasoning can be achieved not through massive single models but through structured, iterative processes that leverage multi-modal feedback effectively Lalam. It’s a blueprint for building more capable AI systems.

Tom: That’s all the time we have for this discussion on CADReasoner and its potential impact on three dee reconstruction, folks <ref:2603.29847#pg0>. We’ll be right back after the break with more deep dives into cutting-edge research.

Conclusion: Tom: So, we've been digging into how CADReasoner uses that closed-loop program editing process to iteratively refine three dee model reconstructions, and now we need to wrap up with some big thoughts on what this whole endeavor actually means for us. Jane, you got a minute to talk about the title and the authors of this paper?

Jane: I can certainly do that, Tom. The title itself, "CADReasoner: Iterative Program Editing for CAD Reverse Engineering," really tells you exactly what's happening here in plain language. It signals that the core idea isn't just generating one final shape, but using a series of small adjustments—iterations—to get to something accurate based on the initial input data. The authors are clearly focused on taking a human-like approach to making sure the AI understands how to correct its own mistakes during this reconstruction process.

Lu: I think it’s fascinating because they’re treating the AI like an engineer revising a design, which is such a powerful way to frame it. The implications here go way beyond just better CAD models; we're looking at building systems that can actively participate in design revision, not just output a static file. That kind of self-correction capability is what really opens up new doors for generative engineering.

Meng: I agree with Lu on the revision aspect, but from my side as an engineer, I’m thinking about the practical impact on manufacturing workflows. If this system can reliably produce outputs that are highly accurate due to this iterative feedback, it could drastically cut down on manual quality control steps later in the production line. That level of precision is hard to achieve otherwise.

Lalam: From my perspective as a large language model, what stands out is how this framework demonstrates the power of structured reasoning loops over massive single predictions. It shows that complex problem-solving isn't always about one huge calculation; sometimes it’s about breaking it down into manageable steps where each step corrects the previous one. This approach really shifts our focus toward building more robust, reliable AI capabilities that can handle nuanced tasks.

Tom: That makes sense, Lalam. So to recap, we've looked at how CADReasoner uses iterative editing to improve three dee reconstruction, and we've seen how this method mimics human engineering revision while tackling real-world scan conditions. Now it’s time to think about what this means for the future of design automation. Where do you all see this technology heading?

Lomonosov Moscow State University · Universite Paris Dauphine

cs.GR, cs.CV, cs.HC

Submitted: 2026-02-18

Updated: 2026-10-06

Comments: Accepted to Findings of CVPR 2026

License: http://creativecommons.org/licenses/by/4.0/

Importance score: 92/100

The gist: CADReasoner introduces a novel framework for AI-based CAD reconstruction that performs iterative inference and self-correction, enabling progressively refined alignment between CAD models and the

Key concepts

Closed-loop Program Editing
This is the core mechanism where the AI program continuously edits itself. Instead of making one final guess, it executes a candidate program and compares its result to the target shape. The difference (discrepancy) is used to update and improve the next version of the program, creating a continuous cycle of testing, comparing, and refining.
Geometry Memory
This concept refers to how previous steps in the reconstruction process are stored as text tokens. These tokens are extracted from geometry branches within the previous program ($C_{t-1}$) and act as context for the current prediction. This memory conditions the code decoder, ensuring that later edits are informed by what has already been built.
Scan-Simulation Protocol
This protocol simulates real-world scanning artifacts like noise, occlusions, and misregistration during training and testing. By creating point clouds from virtual scans using Screened Poisson Surface Reconstruction, the model learns to handle imperfect scan data. This ensures the system is robust when dealing with realistic input conditions.
On-Policy Curriculum
The training strategy uses a curriculum where refinement steps are supervised using examples generated by the current model. Specifically, Stage B trains on both initial guesses ($t=1$) and refined guesses ($t=2$). This ensures that each refinement step is evaluated based on the actual distribution of errors produced by the model at that specific stage.

Terminology

Summary

CADReasoner introduces a novel framework for AI-based CAD reconstruction that performs iterative inference and self-correction, enabling progressively refined alignment between CAD models and the input.

How it works

The core mechanism of CADReasoner is a closed-loop program editing process. Rather than relying on single-shot predictions, it exploits the intrinsic geometry feedback available by executing candidate programs and comparing the result to the target shape. At each iteration, the model updates its prediction based on this geometric discrepancy. The process is defined as:

C t = fθ E(T, St-1), where E(T, St-1) encodes the pair (multi-view overlays and point-set features).

The model is trained to refine its own prediction from cross-modal discrepancy evidence (multi-view overlays and nearest-surface offsets) together with the previous program. This aligns supervision with real-world, iterative reverse-engineering practice, treating self-editing as a minimal, learnable implementation of the 3D reasoning loop an engineer performs when inspecting and revising a design.

CADReasoner Architecture

The CADReasoner architecture is designed to fuse multi-view renders and point clouds as complementary modalities. It accepts three inputs per iteration:

  1. A multiview image obtained by overlaying the target and the previous render (G/R channels), encoded by a shared visual backbone.

  2. A point-cloud branch where per-point features are projected by a linear layer and aggregated by a set encoder.

  3. The previous program, Ct−1, tokenized as text, where tokens from geometry branches form a geometry memory that conditions the code decoder.

The decoder generates the updated script Ct; an execute gate produces St that is fed back into the model on the next step. This loop closes by re-extracting modalities from St, allowing for geometry-driven self-correction.

Scan Realism through Scan Simulation

A key contribution addressing realism is the introduction of a scan-simulation protocol applied during both training and evaluation. This simulation pipeline models real camera-based scanning artifacts, such as noise, occlusions, missing patches, extra or missing holes, and mild mis-registration. The virtual 3D scanning algorithm uniformly samples surface points while preserving normals to create a point cloud that is then converted into a mesh via Screened Poisson Surface Reconstruction. This ensures the editor learn[s] to operate on scan-like inputs and is evaluated under the same conditions.

Training Strategy and Curriculum

The training utilizes a single shared checkpoint for all iterations, where the same model performs initialization (t=1) and subsequent refinement steps (t>1). The training employs a curriculum designed to prevent overfitting:

Stage A (seed, t=1 on D1): We train the editor by supervised cross-entropy to predict the ground-truth program C⋆ from E(T, ∅).

Rollout B: Using the Stage-A checkpoint, we generate one-step predictions C1 for every T ∈ D2.

Stage B (learn to refine): We now train on a mixture of (t=1) and (t=2) examples: for t=1, inputs are E(T, ∅); for t=2, inputs are E(T, S1) with context C1.

This curriculum ensures supervision is on-policy, where each refinement step sees the distribution of errors produced by the current model.

Inference and Decoding Regimes

CADReasoner supports two decoding regimes:

  1. Greedy (single path): The editor decodes deterministically using argmax, compiles Ct, renders St = R(Ct), and updates the best-so-far program C≤t. Complexity is O(s) renders for s steps.

  2. Stochastic, geometry-guided beam: This method runs a stochastic beam with geometry pruning. At each step, it samples N=5 candidates per parent, compiles and renders them, ranks by the primary metric D (invalid generations are discarded), and retains the top N survivors for the next iteration.

The evaluation protocol is robust: on clean experiments, metrics are computed against the clean target; on scan-sim tracks, candidate selection uses CD against the scan to avoid oracle access, while final reported metrics are computed against the corresponding clean surface to ensure comparability across tracks. The best-so-far program C≤s is always reported at evaluation step s.

Key Findings

CADReasoner achieves state-of-the-art results on DeepCAD, Fusion360, and MCB on both clean and scan-sim tracks.

Improvements for AI systems

Here are specific improvements that CADReasoner offers to existing AI systems for CAD reverse engineering:

  1. Replace single-pass or simple alignment methods with a closed-loop, iterative self-editing framework (CADReasoner). Existing systems often fail because they produce a final output without correcting errors based on the discrepancy between the input and the predicted shape.

  2. Integrate geometry discrepancy directly into the training supervision (Self-Supervised Fine-Tuning, SFT) rather than relying solely on external, coarse rewards from Reinforcement Learning (RL). This forces the model to learn iterative refinement based on cross-modal evidence (multi-view overlays and nearest-surface offsets).

  3. Enhance realism by incorporating a scan-simulation protocol during both training and evaluation. This allows the AI editor to learn to handle real-world scan defects—such as noise, occlusions, missing patches, and mild mis-registration—resulting in CAD models that are robust against imperfect physical inputs.

  4. Implement multi-modality conditioning by fusing complementary evidence:

  5. A visual branch (multi-view images) providing global silhouette and topological context.

  6. A point-cloud branch providing metric, local error signals through cross-shape offsets (nearest-surface distances).

  7. Enable the editor to operate on executable CAD programs (CadQuery), allowing it to generate runnable code rather than just static meshes, which preserves editability and allows for programmatic self-correction.

These improvements enable the improved AI system (CADReasoner) to perform:

  1. Reconstruct high-quality, topologically correct CAD models from noisy, real 3D scan data with minimal geometric discrepancy (lower Chamfer Distance/CD).

  2. Recover fine geometric details and correct topology that often escape single-shot or basic alignment methods (higher volumetric IoU).

  3. Produce outputs that are robust to real-world scanning artifacts, ensuring the generated CAD models accurately reflect physical parts, even when input data is incomplete or corrupted.

  4. Achieve state-of-the-art performance across diverse benchmarks (DeepCAD, Fusion 360, MCB) on both clean and scan tracks by iteratively refining its predictions over multiple steps.

Abstract

Computer-Aided Design (CAD) powers modern engineering, yet producing high-quality parts still demands substantial expert effort. Many AI systems tackle CAD reverse engineering, but most are single-pass and miss fine geometric details. In contrast, human engineers compare the input shape with the reconstruction and iteratively modify the design based on remaining discrepancies. Agent-based methods mimic this loop with frozen VLMs, but weak 3D grounding of current foundation models limits reliability and efficiency. We introduce CADReasoner, a model trained to iteratively refine its prediction using geometric discrepancy between the input and the predicted shape. The model outputs a runnable CadQuery Python program whose rendered mesh is fed back at the next step. CADReasoner fuses multi-view renders and point clouds as complementary modalities. To bridge the realism gap, we propose a scan-simulation protocol applied during both training and evaluation. Across DeepCAD, Fusion 360, and MCB benchmarks, CADReasoner attains state-of-the-art results on clean and scan-sim tracks.

Sources

Related papers