NTH
AI research

Text-to-CAD Evaluation with CADTests

AuthorsDimitrios Mallis, Marco Wang, Ahmet Serdar Karadeniz, Elisa Ricci, Anis Kacem, Djamila Aouada

May 18, 2026 2 min read
Watch on YouTube
The one-line take

This paper brings software-style unit testing to Text-to-CAD, creating a benchmark that checks whether generated CAD models actually satisfy the prompt and even helps improve generation methods.

Key results

200 CADPrompt samples
CADTestBench samples

CADTestBench is built from the CADPrompt benchmark and evaluates prompt-to-CAD correspondence on 200 CAD programs.

5,937 tests
CADTestBench tests

The benchmark contains 5,937 executable CADTests across abstract and detailed prompts.

1,275 mutated CAD programs
CAD mutants

Mutation analysis generates 1,275 CAD mutants used to refine and validate CADTest suites.

over 90%
Mutation score after refinement

After four rounds of iterative refinement, the Claude-4.6-Sonnet planner raises mutant detection to over 90% for both prompt types.

0.810
Best baseline requirement score

CADTests+Log with Claude-4.6-Sonnet achieves a 0.810 requirement score on detailed prompts.

0.962
Best baseline pass rate

CADTests+Log with Claude-4.6-Sonnet achieves a 0.962 pass rate on detailed prompts.

What the paper found

Text-to-CAD Evaluation with CADTests reframes CAD generation as executable program synthesis and replaces reference-based similarity metrics with property-based software tests over Boundary Representation, or B-rep, geometry. The paper introduces CADTests, Python/CadQuery snippets that verify prompt constraints directly through topology counts, bounding-box dimensions, face types, volume, and spatial relations, and CADTestBench, the first test-driven Text-to-CAD benchmark built from 200 CADPrompt samples with 5,937 tests and 1,275 mutated CAD programs. A Claude-4.6-Sonnet planner, guided by mutation analysis and iterative refinement, raises mutant detection from about 65 percent to over 90 percent after four rounds, showing that automated test synthesis can produce discriminative suites without a fixed reference shape. On evaluation, recent methods such as Text2CAD, CADCodeVerify, ReAct, and ReAct+Image are judged with pass rate, requirement score, and invalidity rate; the strongest test-based baseline, CADTests+Log with Claude-4.6-Sonnet, reaches 0.810 requirement score and 0.962 pass rate on detailed prompts, outperforming prior methods. The authors also report that CADTests align better with human judgments than Chamfer Distance, CLIP score, or LVM-based metrics, achieving 0.938 accuracy and 0.962 F1 against expert consensus, which is a substantial gain over LVM’s 0.474 accuracy and 0.549 F1. The key novelty is not a new generator but a new verification layer that makes prompt compliance measurable, interpretable, and capable of guiding generation.

Original abstract

Text-to-CAD has recently emerged as an important task with the potential to substantially accelerate design workflows. Despite its significance, there has been surprisingly little work on Text-to-CAD evaluation, and assessing CAD model generation performance remains a considerable challenge. In this work, we introduce a new evaluation perspective for Text-to-CAD based on automated testing. We propose CADTestBench, the first test-based benchmark for Text-to-CAD, based on CADTests, executable software tests that verify whether a generated CAD model satisfies the geometric and topological requirements of the input prompt. Using CADTestBench, we conduct comprehensive benchmarking of recent Text-to-CAD methods and further demonstrate that CADTests can also guide CAD model generation, yielding simple baselines that surpass performance of current methods. CADTestBench code and data are available at GitHub and Hugging Face dataset.

Read the original paper