Text-to-CAD Evaluation with CADTests
AuthorsDimitrios Mallis, Marco Wang, Ahmet Serdar Karadeniz, Elisa Ricci, Anis Kacem, Djamila Aouada
Resources
This paper brings software-style unit testing to Text-to-CAD, creating a benchmark that checks whether generated CAD models actually satisfy the prompt and even helps improve generation methods.
Key results
CADTestBench is built from the CADPrompt benchmark and evaluates prompt-to-CAD correspondence on 200 CAD programs.
The benchmark contains 5,937 executable CADTests across abstract and detailed prompts.
Mutation analysis generates 1,275 CAD mutants used to refine and validate CADTest suites.
After four rounds of iterative refinement, the Claude-4.6-Sonnet planner raises mutant detection to over 90% for both prompt types.
CADTests+Log with Claude-4.6-Sonnet achieves a 0.810 requirement score on detailed prompts.
CADTests+Log with Claude-4.6-Sonnet achieves a 0.962 pass rate on detailed prompts.
What the paper found
Text-to-CAD Evaluation with CADTests reframes CAD generation as executable program synthesis and replaces reference-based similarity metrics with property-based software tests over Boundary Representation, or B-rep, geometry. The paper introduces CADTests, Python/CadQuery snippets that verify prompt constraints directly through topology counts, bounding-box dimensions, face types, volume, and spatial relations, and CADTestBench, the first test-driven Text-to-CAD benchmark built from 200 CADPrompt samples with 5,937 tests and 1,275 mutated CAD programs. A Claude-4.6-Sonnet planner, guided by mutation analysis and iterative refinement, raises mutant detection from about 65 percent to over 90 percent after four rounds, showing that automated test synthesis can produce discriminative suites without a fixed reference shape. On evaluation, recent methods such as Text2CAD, CADCodeVerify, ReAct, and ReAct+Image are judged with pass rate, requirement score, and invalidity rate; the strongest test-based baseline, CADTests+Log with Claude-4.6-Sonnet, reaches 0.810 requirement score and 0.962 pass rate on detailed prompts, outperforming prior methods. The authors also report that CADTests align better with human judgments than Chamfer Distance, CLIP score, or LVM-based metrics, achieving 0.938 accuracy and 0.962 F1 against expert consensus, which is a substantial gain over LVM’s 0.474 accuracy and 0.549 F1. The key novelty is not a new generator but a new verification layer that makes prompt compliance measurable, interpretable, and capable of guiding generation.
Original abstract
Text-to-CAD has recently emerged as an important task with the potential to substantially accelerate design workflows. Despite its significance, there has been surprisingly little work on Text-to-CAD evaluation, and assessing CAD model generation performance remains a considerable challenge. In this work, we introduce a new evaluation perspective for Text-to-CAD based on automated testing. We propose CADTestBench, the first test-based benchmark for Text-to-CAD, based on CADTests, executable software tests that verify whether a generated CAD model satisfies the geometric and topological requirements of the input prompt. Using CADTestBench, we conduct comprehensive benchmarking of recent Text-to-CAD methods and further demonstrate that CADTests can also guide CAD model generation, yielding simple baselines that surpass performance of current methods. CADTestBench code and data are available at GitHub and Hugging Face dataset.
Read the original paper