CADEngBench: It Looks Like CAD, but Does It Work? Evaluating Parametric Design, Assembly Reasoning, and Physics Simulation

arXiv:2608.09296 · cs.AI, cs.CV, cs.LG, cs.RO · Submitted 2026-08-10 · Read on arXiv

Harmanjot Singh, Abhra Dubey, Jorge Alejandro Amador Herrera

Mohamed bin Zayed University of Artificial Intelligence

cs.AI, cs.CV, cs.LG, cs.RO

Submitted: 2026-08-10

Updated: 2026-08-11

License: http://arxiv.org/licenses/nonexclusive-distrib/1.0/

Importance score: 75/100

The gist: CADEngBench is a two-track benchmark for evaluating engineering-grade CAD capabilities beyond appearance and executability.

Terminology

Summary

CADEngBench is a two-track benchmark for evaluating engineering-grade CAD capabilities beyond appearance and executability. The paper argues that "A CAD model is not engineering-grade merely because it looks correct. It must satisfy design requirements, respond predictably to parameter changes, support controlled edits, match a reference structural response under a declared analysis, and connect to other parts through valid joints."

The benchmark has two tracks. "CADEngBench-P evaluates 300 parametric parts, each used for one zero-to-CAD task and one functional-editing task (600 tasks in total), through boundary-representation (B-Rep) validity, engineering and DFM checks, parameter-family perturbations, functional editing, and matched linear-static FEA in CalculiX. CADEngBench-A evaluates 150 body pairs through ranked joint retrieval, exact face-and-edge grounding, joint-frame prediction, and kinematic verification."

The authors make four key contributions: "We develop a two-track benchmark for engineering-grade CAD: CADEngBench-P provides 600 zero-to-CAD and functional-editing tasks, while CADEngBench-A provides 150 assembly pairs. We develop an evaluation hierarchy that moves beyond nominal shape similarity to test B-Rep validity, engineering and DFM requirements, operational parameterization, functional edits, and preservation of non-target design properties. We introduce assembly evaluation in which models must retrieve the exact mating faces or edges, predict the joint, and demonstrate motion through kinematic simulation. We introduce FEA-based physics verification across parameter changes using CalculiX to compare stress, deformation, compliance, and stress concentration across parametric design families."

For CADEngBench-P, the evaluation has layers. L0 tests whether the LLM-generated output is usable CAD. The CadQuery must execute, create at least one valid B-Rep solid, export to STEP, and successfully re-import the exported STEP. L1 tests whether the valid solid satisfies the engineering zero-to-CAD task with checks for dimensions, feature counts, interfaces, and DFM screens (1.0 mm minimum wall, 2.0 mm minimum hole diameter, hole depth/diameter ≤ 8). L2-Z tests parametric integrity: whether a declared parameter controls the intended geometry by rebuilding at multiple parameter values. L2-E tests controlled modification of existing CAD rather than generation from scratch with hidden preservation checks. L3 compares generated and reference CAD under the same linear-static material, supports, and loads using Gmsh meshing and CalculiX solving, comparing 95th-percentile von Mises stress, maximum displacement, and compliance or stress concentration.

For CADEngBench-A, A1 receives two bodies, multiview images, and indexed B-Rep face/edge candidates with stable IDs and returns up to three hypotheses with joint family and one entity from each body. A2 conditions on the rank-1 A1 hypothesis and predicts its joint origin and family-specific axes or directions in Body A’s coordinate frame which becomes a URDF joint executed in PyBullet.

The paper evaluates eight multimodal, code-capable models: GPT-5.2, Claude Sonnet 4.5, Gemini 3 Flash, GLM-4.6V, Kimi K2.5, Mistral Medium 3.5, Llama 4 Maverick, and Qwen3.5-35B-A3B.

Key results show substantial variation across models and stages. Gemini leads engineering-requirement satisfaction, functional editing, matched FEA, and assembly retrieval, whereas Claude leads executability, parametric behavior, and L3 simulation reach. Across all eight systems, 1,030/2,400 generated programs pass L0, but only 432 of those also pass L1. Thus, 58.1% of executable outputs violate at least one stated engineering or DFM requirement.

For editing versus generation: "Among 2,330 item–model pairs scoreable under both protocols, 1,071 pass only editing, compared with 231 that pass only generation; 468 pass both and 560 pass neither. The 4.64× edit-only asymmetry explains the 25.5–55.4 point gaps. Editing difficulty depends on structure: A single independent-body addition passes in 247/248 cases, whereas histories involving joins, cuts, or multiple bodies pass in only 40.2–45.8% of cases."

For physics: "a successful solve does not establish correct physics. Models that produce more valid CAD by passing all previous evaluation layers reach more L3 cases (ρs = 0.976), but reach is unrelated to agreement with the reference FEA (ρs = 0.024). CalculiX itself fails in only 5 of 1,809 attempts. In contrast, 573 attempts fail before the solve because meshing or boundary-condition assignment fails, and 554 solve but violate engineering limits or disagree with the reference stress or deformation. Also, 22/435 matched cases change verdict away from the default parameter value, so testing only one design state can hide physics failures."

For assembly: A correct B-Rep entity pair is ranked first in 542/960 requests, but only 143 also identify the joint family and body ordering. Thus, just 26.4% of correct rank-1 entity localizations specify the full joint. Joint family matters: Revolute and cylindrical joints are easier than slider, planar, and pin-slot joints. The joint frame is a further bottleneck: 110 of 142 correct rank-1 joints place it within tolerance, and only 11.5% of all requests pass end to end.

The paper concludes: "CADEngBench evaluates CAD beyond appearance and executability through engineering requirements, parameter behavior, controlled edits, matched FEA, and assembly grounding. Across the eight evaluated systems, performance is only weakly aligned across these capabilities; no single score describes the full profile. The benchmark exposes failures that surface checks miss: executable code can violate design intent, generated parameter families can diverge from the reference structural response, and plausible assembly predictions can fail to recover the recorded mating relation. CADEngBench therefore treats CAD as an engineering artifact rather than a plausible shape."

Improvements for AI systems

Improvements to AI systems:

  1. Add a multi-layer validation pipeline for CAD code generation. The AI system should not stop at code executability or B-Rep validity (L0). It must automatically run engineering and DFM checks (minimum wall thickness, hole diameter/depth ratio, feature counts, interface dimensions) and reject outputs that violate them, reducing the 58.1% failure rate where executable code still fails engineering requirements.

  2. Implement parametric integrity verification. The AI should generate CAD programs that declare parameters and then rebuild the model at multiple parameter values (e.g., ±20%, ±50%) to confirm the parameter controls the intended geometry. This prevents fake parameters that only change a label without altering the actual shape.

  3. Add functional-editing capability with preservation checks. Instead of generating from scratch, the AI must learn to modify existing CAD histories—specifically handling joins, cuts, and multi-body operations. The system should verify that non-target design properties (e.g., other dimensions, fillets, or features) remain unchanged after an edit, addressing the 4.64× edit-only asymmetry where edits are far harder than generation.

  4. Integrate physics-based verification via FEA. The AI should automatically mesh its generated CAD (using Gmsh) and run linear-static FEA (CalculiX) under the same material, supports, and loads as the reference. It must compare 95th-percentile von Mises stress, max displacement, compliance, and stress concentration. Critically, the system should test multiple design states (parameter values), not just one, because 22/435 cases changed verdicts across states—single-state testing hides physics failures.

  5. Add assembly grounding with joint-frame prediction. For assembly tasks, the AI must not only retrieve the correct mating faces/edges (ranked) but also predict the joint family (revolute, prismatic, cylindrical, etc.) and the exact joint origin and axes in the body coordinate frame. The system should then execute the joint in a kinematic simulator (e.g., PyBullet) to verify motion, not just static placement.

  6. Create a capability-profile-aware routing system. Since performance is only weakly correlated across capabilities (e.g., Gemini leads in engineering/FEA/assembly, Claude leads in executability/parametric), the AI should maintain a per-capability scorecard and route tasks to the best-suited model or use ensemble voting for each layer (L0–L3, A1–A2) rather than relying on a single global score.

  7. Implement failure-aware meshing and boundary-condition handling. The AI should pre-check meshing feasibility and boundary-condition assignment before solving, since 573/1,809 attempts failed at meshing or BC assignment. It should generate CAD that is mesh-friendly (e.g., avoiding tiny features or sliver faces) and auto-correct BC definitions to reduce pre-solve failures.

  8. Add joint-family-specific training and verification. For assembly, the AI should be trained to handle harder joint types (slider, planar, pin-slot) with more examples, and it should verify joint frame tolerance (within 5 mm / 5 degrees, for example) before declaring success, since only 11.5% of all assembly requests pass end to end.

What the improved AI system can do:

  • Generate CAD code that is not only executable and valid but also passes engineering, DFM, and parametric integrity checks, reducing hidden design-rule violations.

  • Perform controlled functional edits on existing CAD histories (including joins, cuts, multi-body) while preserving non-target properties, with a success rate comparable to generation.

  • Produce CAD models whose structural response (stress, deformation, compliance) matches a reference FEA across multiple parameter values, catching physics failures that single-state checks miss.

  • Assemble parts by correctly identifying mating faces/edges, predicting the joint family and exact joint frame, and verifying motion through kinematic simulation—achieving end-to-end assembly success beyond the current 11.5%.

  • Adaptively select or combine models per capability layer, maximizing overall engineering-grade CAD quality rather than optimizing a single average score.

Sources

Related papers