DepthBenchCAD: When Does Deeper Auditing Yield More Reliable Conclusions?
cs.SE, cs.AI
Submitted: 2026-09-14
Updated: 2026-09-14
Comments: 29 pages, 3 figures, 20 tables. Preprint
License: http://arxiv.org/licenses/nonexclusive-distrib/1.0/
Terminology
Sources
- Text2CAD-Bench: A Benchmark for LLM-based Text-to-Parametric CAD Generation
- BenchCAD: A Comprehensive, Industry-Standard Benchmark for Programmatic CAD
Related papers
- Falsification-Based Verification of LLM-Generated Optimization Models: Sound Test Batteries and Their Detection Limits
- GitSkills: A Dataset of Agent Skills on GitHub
- SABER: Benchmarking Operational Safety of LLM Coding Agents in Stateful Project Workspaces
- PackMonitor: Enabling Zero Package Hallucinations Through Decoding-Time Monitoring
- IntentCoding: Amplifying User Intent in Code Generation
- Incentives and Outcomes in Bug Bounties