RA-CAD: Learning Post-Execution Critique for State-Aware Text-to-CAD Generation
Listen
Radio episode about this paper
Transcript
Introduction to the show: ident: AI Radio. Generated commentary on the latest Artificial Intelligence papers.
Tom: Next we'll be talking about the paper "RA-CAD: Learning Post-Execution Critique for State-Aware Text-to-CAD Generation".
Jane: The paper was written by Shuhao Yan, Changhao He, Peng Hu and Xi Peng from Sichuan University and National Key Laboratory of Fundamental Algorithms and Models for Engineering Numerical Simulation, Sichuan University.
Tom: Stay tuned as we take you through the paper and discuss its implications.
Title and Authors: Tom: Welcome back to the show, everyone. Today we’re digging into a fresh arXiv paper called "RA-CAD: Learning Post-Execution Critique for State-Aware Text-to-CAD Generation." And Jane, I gotta say, just the title alone has me hooked.
Jane: Oh, absolutely, Tom. And it’s from a team at Sichuan University, with Shuhao Yan and Changhao He as the lead authors, plus Xi Peng and Peng Hu. The idea of teaching a model to critique its own work after it runs the code? That’s a big leap from just generating something and hoping it’s right.
Tom: Right, because in the world of CAD, or computer-aided design, you’re not just writing text. You’re writing code that has to actually build a three dee model. And if the code is even slightly off, the whole thing falls apart.
Jane: Exactly. And the title mentions "state-aware," which I think is the key. The model isn’t just looking at the original prompt. It’s looking at the current state of the code, what the execution environment said, and then deciding what to do next. It’s like having a conversation with the software.
Tom: So instead of a one-shot guess, it’s a loop. Generate, run, critique, rewrite. That’s the core of what they call the ReAct Agent for CAD. And honestly, that feels like how a human designer would actually work.
Jane: For sure. You don’t just draw a part once and call it done. You look at it, you measure it, you see if it matches the spec, and you fix it. This paper is trying to give the AI that same kind of iterative, self-correcting workflow.
Tom: And the implications are huge. Think about manufacturing, prototyping, even education. If you can describe a part in plain English and get a working CAD file back, you’ve just lowered the barrier to entry for a whole lot of people.
Jane: But it’s not just about getting *a* file. It’s about getting a file that’s actually executable. The paper spends a lot of time on that "invalidity ratio," which is basically how often the generated code just crashes or can’t be turned into a solid object. And that’s where the critique part becomes so important.
Tom: So the model is learning to catch its own mistakes before it hands you the final product. That’s a pretty powerful idea. I can’t wait to see how they actually trained it to do that.
Jane: Me neither. Because teaching a model to critique itself is tricky. You need the right feedback signals, and you need to make sure the critique is actually useful for the next rewrite, not just a bunch of generic complaints.
Tom: Well, we’re about to get into exactly that. Stick around, because next we’re going to break down the summary of the paper and how they made this whole loop work.
Summary: Tom: So, Jane, we’re back with "RA-CAD: Learning Post-Execution Critique for State-Aware Text-to-CAD Generation," and I want to get into the meat of the summary. The paper isn’t just about adding a critique step; it’s about making that critique a learnable part of the policy.
Jane: Right. And that’s a subtle but important distinction. A lot of previous work might use an external reviewer or a fixed prompt to give feedback. But here, the critique itself is generated by the model, and then the model is trained to make that critique better over time.
Tom: So it’s not just a tool the model uses; it’s a skill the model is learning. And they’re doing this through a two-stage training process. The first stage is what they call CAD Code Bootstrapping, or CCB.
Jane: That’s the supervised fine-tuning stage. They’re teaching the model the basic syntax and structure of CAD code by showing it tons of examples of descriptions paired with the correct code. It’s like teaching someone the alphabet before you ask them to write a novel.
Tom: Makes sense. You can’t expect a model to critique code if it doesn’t even know how to write valid code in the first place. But then comes the second stage, which is where the magic happens. They call it Feedback-Driven Agent Optimization, or FAO.
Jane: And this is where they use reinforcement learning. Specifically, they’re using something called GRPO, which is a way to train the model by comparing different attempts and rewarding the ones that produce better final results.
Tom: So the model generates a whole trajectory of actions—generate code, run it, critique it, rewrite it—and then at the very end, they check how good the final CAD model is. If it’s good, the whole trajectory gets a high reward.
Jane: And the reward isn’t just about whether the code runs. They’re also measuring geometric quality using something called Chamfer Distance, which basically measures how close the generated three dee shape is to the ground truth shape. So the model is learning to critique and rewrite in a way that actually improves the final geometry.
Tom: That’s the part that really excites me. The critique isn’t just a formality. It’s being optimized to produce actionable feedback that leads to a better final product. It’s like the model is learning to be a better engineer.
Jane: Exactly. And the results seem to back that up. They tested this on two datasets, CADFusion and Text2CAD, and they saw significant improvements in both execution validity and geometric accuracy compared to other methods.
Tom: I think the most striking number was the invalidity ratio. On the CADFusion dataset, they got it down to six point two percent, which is way lower than the baselines. That means the model is producing code that actually works almost all the time.
Jane: And that’s a huge deal for practical use. If you’re a designer and you’re using this tool, you don’t want to spend half your time debugging the AI’s output. You want it to work the first time, or at least get it right after a couple of revisions.
Tom: So the summary is basically: teach the model the basics, then let it learn from its own mistakes through trial and error, and make sure the critique it generates is actually driving those improvements.
Jane: That’s the gist of it. But I’m curious about the specific improvements they made over existing methods. I think that’s where we should go next.
Improvements: Tom: Alright, Jane, so we’ve covered the basics of "RA-CAD: Learning Post-Execution Critique for State-Aware Text-to-CAD Generation." Now let’s talk about what actually makes it better than what came before. Because it’s not just about having a critique step; it’s about how that critique is integrated.
Jane: Right. And I think the biggest improvement is that they treat the critique as a policy action, not just an auxiliary output. In older systems, the critique might be generated by a separate model or a fixed set of rules. Here, the critique is part of the same model that generates the code, and it’s trained jointly.
Tom: So the model is learning to generate code and critique it at the same time, and both of those skills are being optimized together. That’s a pretty elegant way to think about it. And it means the critique is directly tied to the final reward.
Jane: Exactly. And they also made a point about the state. The model isn’t just looking at the original description. It’s looking at the current code, the execution results, and the previous critique. So it has a full picture of where things stand before it decides what to do next.
Tom: That’s the "state-aware" part of the title. And it’s a big deal because it lets the model make more informed decisions. If the execution failed, the model knows exactly why it failed and can target that specific issue in the rewrite.
Jane: And they have this really nice breakdown of the action into four modules: Generation, Execution, Critique, and Rewriting. Execution is the environment, but the other three are all part of the agent’s policy. So the model is deciding what to generate, how to critique it, and how to rewrite it.
Tom: I like how they frame the initial generation as just a hypothesis. It’s not the final answer; it’s a starting point that will be verified and refined. That’s a much more realistic way to approach complex tasks like CAD modeling.
Jane: And the improvements show up in the numbers. They compared against strong baselines like Text2CAD and CADFusion, and even against proprietary models like GPT-4o and DeepSeek. RA-CAD consistently came out ahead on geometric quality and execution validity.
Tom: One thing that stood out to me was how much better they did on the invalidity ratio. On the Text2CAD dataset, they got it down to nine point four four percent, while some of the proprietary models were above sixty percent. That’s a massive difference in reliability.
Jane: And that reliability is what makes this practical. If you’re an engineer using this in your workflow, you need to trust that the output is going to be usable. RA-CAD is building that trust by making the critique loop actually work.
Tom: So the improvements are really about integration and learning. It’s not just bolting on a critique step; it’s making the critique a core part of the learning process. And that’s what leads to the better results.
Jane: And I think that’s a lesson that could apply beyond CAD. Any task where you can execute your output and get feedback could benefit from this kind of closed-loop, self-critiquing approach.
Tom: That’s a great point. But before we get too philosophical, let’s wrap up and think about what this all means for the future.
Conclusion: Tom: Well, Jane, we’ve had a great time digging into "RA-CAD: Learning Post-Execution Critique for State-Aware Text-to-CAD Generation." Let’s try to pull it all together for our listeners.
Jane: Absolutely. At its heart, this paper is about making text-to-CAD generation more reliable by teaching the model to critique its own work. They do that with a two-stage process: first, supervised learning to teach the basics, and then reinforcement learning to refine the whole generate-execute-critique-rewrite loop.
Tom: And the key insight is that the critique isn’t just a formality. It’s a learnable policy action that’s optimized to produce better final results. The model learns to identify its own mistakes and fix them, which leads to much higher execution validity and better geometric accuracy.
Jane: The results on CADFusion and Text2CAD really speak for themselves. They beat out strong baselines and even proprietary models, especially when it comes to producing code that actually runs and creates the right shape.
Tom: And for me, the biggest takeaway is the shift from one-shot generation to iterative refinement. That’s how humans work, and it’s exciting to see AI models starting to work that way too.
Jane: Definitely. There are still limitations, like the fact that it only works with a specific set of CAD operations and doesn’t support visual inputs like images or point clouds yet. But the framework they’ve built is solid.
Tom: And it opens up a lot of possibilities. Imagine being able to describe a part in plain English and getting a working CAD file back, or even being able to edit existing designs by just describing what you want to change.
Jane: It could really democratize design and manufacturing, making it accessible to people who don’t have years of CAD training. And that’s a pretty exciting vision for the future.
Tom: Couldn’t agree more. So let’s say goodbye to "RA-CAD: Learning Post-Execution Critique for State-Aware Text-to-CAD Generation." It’s been a fascinating look at how AI can learn to check its own work.
Jane: Thanks for joining us, everyone. We’ll be back next time with another paper to break down. Until then, keep building and keep questioning.
Shuhao Yan, Changhao He, Peng Hu, Xi Peng
Sichuan University · National Key Laboratory of Fundamental Algorithms and Models for Engineering Numerical Simulation, Sichuan University
cs.AI
Submitted: 2026-08-07
Updated: 2026-08-18
Comments: 17 pages, 7 figures, 5tables, 8 listings
License: http://arxiv.org/licenses/nonexclusive-distrib/1.0/
Importance score: 40/100
The gist: The paper introduces RA-CAD (ReAct Agent for CAD), a state-aware agent for text-to-CAD generation that operates through a Generate–Execute–Critique–Rewrite loop.
Key concepts
- State-Aware
- The model does not just use the original text prompt; it considers the current status of the code and execution environment. This allows it to make more informed decisions by knowing exactly where things stand before deciding what to do next.
- ReAct Agent for CAD
- This refers to the core workflow: Generate, run, Critique, and Rewrite. Instead of a single guess, the AI operates in a loop, mimicking how human designers refine work by checking it and fixing mistakes iteratively.
- Post-Execution Critique
- The model is trained to critique its own output after running the code. This critique is not merely an afterthought but a learnable action that guides subsequent rewrites, improving the final product's quality.
- Invalidity Ratio
- This metric measures how often the generated code fails or cannot be turned into a solid object. RA-CAD significantly improves this ratio, meaning the model produces working code much more reliably.
Terminology
Summary
The paper introduces RA-CAD (ReAct Agent for CAD), a state-aware agent for text-to-CAD generation that operates through a Generate–Execute–Critique–Rewrite loop. The core problem addressed is that existing text-to-CAD methods incorporate fixed, externally supplied, prompt-induced, or separately optimized critique mechanisms
which do not necessarily optimize how feedback is interpreted and translated into effective corrective actions throughout the generation process.
The authors identify a feedback-utilization gap
and propose bridging it by modeling post-execution critique as an explicit, learnable policy decision.
The central formulation treats CAD generation as a Markov Decision Process (MDP) where, at each iteration, the agent executes the current code, observes its outcome, and generates an explicit post-execution critique as an intermediate policy action. This critique either validates the current result for termination or provides revision-oriented guidance that conditions the next rewrite.
The complete trajectory is defined as τ = (s0, a1, s1,..., a T, s T), where s0 is the initial state containing only the description, a>0 indicates each composite action, and s>0 represents each complete state.
The framework consists of four functional modules: Generation, Execution, Critique, and Rewriting. Generation and Rewriting produce CAD code proposals, Execution belongs to the environment and returns execution status with error messages, and Critique evaluates whether the current code satisfies design requirements. The agent's decision at each round is a composite action a t = (c̃ t, f̃ t), where c̃ t is the proposed CAD code and f̃ t is the critique output. The policy is factorized as π θ(a t s t−1) = π θ g&r(c̃ t s t−1)π θ crit(f̃ t s̄ t).
Training proceeds in two stages. CAD Code Bootstrapping (CCB) first performs supervised fine-tuning on paired text-CAD data using cross-entropy loss: L SFT = −(1/L)Σ log π θ(c⋆ l d, c⋆<l), establishing fundamental parametric CAD coding capabilities.
Feedback-Driven Agent Optimization (FAO) then applies trajectory-level Group Relative Policy Optimization (GRPO) to both policy-generated code and critique sequences, assigning terminal rewards to the complete interaction trajectory. The reward function is R = λ1R F1 + λ2R CD, where R F1 measures F1 scores of basic sketch topological tokens line, arc, circle, and R CD is defined as e−γCD if the CAD code is executable and 0 otherwise, with CD being the Chamfer distance between generated and ground-truth point clouds. GRPO normalizes rewards within groups of trajectories, computes advantages as A k[p] = (R k − μ i)/σ i, and optimizes a clipped surrogate objective with KL regularization.
Experiments were conducted on the CADFusion and Text2CAD datasets, using Meta-Llama-3-8B-Instruct as the base model with LoRA for efficient tuning. On CADFusion, RA-CAD achieved Avg F1 of 85.30, Avg CD of 35.44, and IR of 6.20, outperforming Text2CAD (Avg F1 76.05, Avg CD 41.50, IR 34.58) and CADFusion (Avg F1 83.59, Avg CD 42.08, IR 21.98). On Text2CAD, RA-CAD achieved Avg F1 of 77.51, Avg CD of 38.86, and IR of 9.44, compared to Text2CAD (Avg F1 70.86, Avg CD 42.98, IR 35.57) and CADFusion (Avg F1 76.69, Avg CD 49.78, IR 40.96).
Comparisons with proprietary LLMs (DeepSeek-V4-Flash, Qwen3.7-Flash, GLM-5.2, Kimi-K2.6-Pro, GPT-4o-mini, GPT-4o) under a uniform 8-shot prompting strategy showed RA-CAD significantly outperforming all of them. On CADFusion, RA-CAD achieved Avg F1 85.30, Avg CD 35.44, and IR 6.20, while the best proprietary model (Kimi-K2.6-Pro) achieved Avg F1 80.36, Avg CD 41.47, and IR 66.25. On Text2CAD, RA-CAD achieved Avg F1 77.51, Avg CD 38.86, and IR 9.44, while the best proprietary model (Kimi-K2.6-Pro) achieved Avg F1 78.78, Avg CD 47.69, and IR 70.27.
Ablation studies isolated the contributions of each component. Without CCB, the model could not generate valid models (IR 100.00). Removing FAO and keeping only CCB resulted in Avg F1 76.78, Avg CD 54.37, and IR 90.01. Removing the Critique module (CCB + GRPO without critique) yielded Avg F1 81.92, Avg CD 38.00, and IR 18.82. The complete RA-CAD achieved Avg F1 85.30, Avg CD 35.44, and IR 6.20, demonstrating that performance gains are not only brought about by GRPO itself but also largely benefit from the post-execution critique rewriting mechanism.
The paper concludes that RA-CAD learned not only to generate CAD codes but also to critique and revise them after execution,
with the central formulation factorized a complete code proposal/rewrite and an explicit post-execution critique within one state-aware policy, then jointly optimized their token sequences over complete trajectories.
The authors note limitations including lack of direct visual conditioning, restriction to the sketch-and-extrude DSL with 6-bit quantized parameters, and limited generalization to broader code spaces, with future work planned on multimodal inputs, interactive user feedback, and broader CAD code spaces.
Improvements for AI systems
Based on the RA-CAD paper, here are the specific improvements I can implement and the resulting capabilities of the improved AI system:
-
Implementation: Replace the one-shot text-to-CAD decoder with a ReAct-style agent that operates in a Generate–Execute–Critique–Rewrite loop. The critique module is not a fixed prompt or external reviewer but a learnable policy component that produces explicit, structured feedback (e.g.,
Missing holes
,Arc collinear error at SE1 F1 L1 C1
) conditioned on the current code and its execution result. -
Benefit: The system can detect and correct its own errors (e.g., missing features, invalid parameters, geometric inconsistencies) through multiple iterations, rather than producing a single unrepairable output.
-
Implementation:
-
Stage 1 (CCB): Supervised fine-tuning on paired text-CAD data using cross-entropy loss to learn syntax and geometric constraints.
-
Stage 2 (FAO): Trajectory-level Group Relative Policy Optimization (GRPO) where the terminal reward combines F1 score on sketch topology tokens (line, arc, circle) and Chamfer Distance on the instantiated 3D model. The reward is assigned to the entire interaction trajectory, jointly optimizing code generation, critique, and rewriting.
-
Benefit: The system learns not just to generate code but to make optimal decisions about when to accept, critique, and revise, leading to higher executability and geometric fidelity.
-
Implementation: Decompose each round's action into a composite
(c̃ t, f̃ t)wherec̃ tis the code proposal/rewrite andf̃ tis the critique output. The policy is factorized asπ(a t s t-1) = π g&r(c̃ t s t-1) · π crit(f̃ t s̄ t), with the state including description, current code, execution result, and previous critique. -
Benefit: The system can explicitly separate
what to generate
fromhow to evaluate it,
enabling more targeted optimization and better interpretability of the decision process. -
Implementation: Use a reward function
R = λ1·RF1 + λ2·RCDwhereRCD = e-γ·CDif the code is executable and 0 otherwise. This penalizes non-executable outputs heavily while rewarding geometric closeness to the ground truth. -
Benefit: The system is strongly incentivized to produce valid, executable code first, then optimize geometric accuracy, leading to dramatically lower invalidity ratios (IR: 6.20% vs. 21.98% for CADFusion baseline).
-
Implementation: Encode CAD models using SkexGen's sketch-and-extrude DSL with topological tokens (line, arc, circle), extrusion tokens (add, cut, intersect), and hierarchical end tokens (
,,,, ``). All continuous parameters are quantized to 6-bit discrete tokens (0–63). -
Benefit: The system can process CAD models as unified text sequences, making them compatible with LLM architectures while preserving geometric and topological validity.
-
Generate Executable CAD Code from Natural Language: Given a text description like
The 3D shape consists of two coaxial overlapping parts with eight evenly spaced holes,
the system produces a valid parametric CAD code sequence that executes without errors. -
Self-Correct Through Multi-Round Interaction: If the initial code has errors (e.g., missing holes, invalid arcs, wrong Boolean operations), the system executes the code, receives diagnostic feedback, critiques the result, and rewrites the code—repeating until the model is correct or the iteration limit is reached.
-
Achieve State-of-the-Art Geometric Accuracy: On CADFusion and Text2CAD benchmarks, the system reduces average Chamfer Distance to 35.44 and 38.86 (×102) respectively, outperforming strong baselines like Text2CAD (41.50, 42.98) and CADFusion (42.08, 49.78), and proprietary LLMs like GPT-4o (61.53, 50.16).
-
Maintain High Execution Validity: The system achieves an invalidity ratio of only 6.20% on CADFusion and 9.44% on Text2CAD, compared to 21.98% and 21.85% for the best existing methods, meaning nearly all generated models can be instantiated and rendered.
-
Handle Complex Multi-Feature Designs: The system correctly generates models with multiple sketches, Boolean operations, arrays of holes, and intricate geometric relationships (e.g., coaxial parts, evenly spaced holes, curved segments), as demonstrated in qualitative case studies.
-
Provide Explainable Critique: The system outputs structured, human-readable critiques (e.g.,
The required cylinder feature is missing: no circle primitive is defined...
) that explain why the current code is insufficient and guide the next revision, making the generation process transparent and debuggable. -
Optimize for Both Sequence and Geometry Quality: By jointly optimizing F1 scores on sketch topology tokens and Chamfer Distance on instantiated models, the system balances parameter-level accuracy with final 3D shape fidelity, avoiding the trade-off seen in methods that optimize only one aspect.
Sources
- PR-CAD: Progressive Refinement for Unified Controllable and Faithful Text-to-CAD Generation with Large Language Models
- CADDesigner: Conceptual CAD Model Generation with a General-Purpose Agent
- CADmium: Fine-Tuning Code Language Models for Text-Driven Sequential CAD Design
- CAD-Coder: Text-to-CAD Generation with Chain-of-Thought and Geometric Reward
- CAD-Coder:Text-Guided CAD Files Code Generation
- IterCAD: An Iterative Multimodal Agent for Visually-Grounded CAD Generation and Editing
- An Overview for Markov Decision Processes in Queues and Networks
- Agent Lightning: Train ANY AI Agents with Reinforcement Learning
- DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models
- Text-to-CadQuery: A New Paradigm for CAD Generation with Scalable Large Model Capabilities
- ReAct: Synergizing Reasoning and Acting in Language Models
- Text2CAD: Text to 3D CAD Generation via Technical Drawings
- Clarify Before You Draw: Proactive Agents for Robust Text-to-CAD Generation
Related papers
- MAVEN-T: Reinforced Heterogeneous Distillation for Real-Time Multi-Agent Trajectory Prediction
- Model Discovery Agent: LLM-assisted Bayesian experiment design for data-efficient discovery of mechanistic world models
- The Clinician's Veto: Navigating Trust, Liability, and Uncertainty in Autonomous AI Prescribing
- MindHelper: Closed-Loop Embodied Mental-State Reasoning for Precision Intervention
- Incumbent Advantage: Brand Bias and Cognitive Manipulation Dynamics in LLM Recommendation Systems
- VSAL: A Vision Solver with Adaptive Layouts for Graph Property Detection