| Rank | Model | Score |
|---|---|---|
| 1 | gpt-6-astra | 95.9 |
1 phaseActive
CAD/parametric-design benchmark testing whether models can produce and edit computer-aided designs against a spec.
Quick answer: BenchCAD is a benchmark measuring whether AI models can produce and edit computer-aided design (CAD) artifacts — from parametric modeling to spec-following geometry edits. It first appeared in OpenAI's GPT-6 Astra announcement (September 2026), where Astra scored 95.9%, ahead of Claude Fable 5.1 (84.3%) and Claude Opus 5 (82.1%).
What it tests: Whether a model can generate, modify, and reason about CAD designs — parametric models, dimensional constraints, and spec compliance — a professional engineering-adjacent skill distinct from general code generation.
Why it matters: CAD work sits at the intersection of spatial reasoning, precise numerical constraint-following, and domain-specific tooling. Strong BenchCAD performance signals a model is viable for engineering, product design, and manufacturing-adjacent agentic workflows, not just software engineering.
Known limitations: Full methodology (task count, dataset composition, grading rubric) has not yet been publicly documented in detail as of this writing — the benchmark is new and primarily known through frontier lab technical reports rather than an independent paper or public dataset.
BenchCAD evaluates models on computer-aided design tasks — the kind of parametric modeling and spec-following work done in professional CAD software. This differentiates it from general-purpose coding benchmarks, since CAD tasks require reasoning about physical geometry, dimensional constraints, and manufacturability alongside standard tool-use and instruction-following.
As reported in OpenAI's GPT-6 Astra launch materials, results for Claude models on this benchmark reflect three modifications made to the evaluation (per Anthropic's own Claude Fable 5.1 system card), so direct comparisons across labs should be read with that caveat in mind — differing eval harnesses or task modifications can shift scores independent of underlying model capability.
| Field | Value |
|---|---|
| Task category | Agent / Professional |
| Metric | % task success |
| Saturation | Medium — top models cluster in the 80-96% range |
| Created by | OpenAI (introduced via GPT-6 Astra announcement) |
| Public dataset | Not yet publicly released as of Sep 2026 |
Scores are reported as a percentage task-success rate. As with many frontier-lab-reported benchmarks, exact grading criteria (partial credit for near-correct geometry, tolerance thresholds for dimensional accuracy) have not been fully disclosed publicly.
| Rank | Model | Score | Source | Date |
|---|---|---|---|---|
| 1 | GPT-6 Astra | 95.9% | OpenAI: GPT-6 Astra | 2026-09 |
| 2 | Claude Fable 5.1 | 84.3% | OpenAI: GPT-6 Astra (via Anthropic system card) | 2026-09 |
| 3 | Claude Opus 5 | 82.1% | OpenAI: GPT-6 Astra | 2026-09 |
| 4 | GPT-5.6 Sol | 83.3% | OpenAI: GPT-6 Astra | 2026-09 |
| 5 | Claude Fable 5 | 67.5% | OpenAI: GPT-6 Astra (via Anthropic system card) | 2026-09 |
Scores sourced from OpenAI's GPT-6 Astra announcement (September 2026). Claude scores reflect 3 modifications to the eval per Anthropic's Claude Fable 5.1 system card — see source for methodology differences.
No Benchgen results yet — be the first to run BenchCAD.
| Benchmark | What it tests | Saturation |
|---|---|---|
| BenchCAD | CAD/parametric design generation and editing | Medium |
| FrontierCode | General agentic coding, rubric-weighted | Low |
| AutomationBench | SaaS/REST-API workflow automation | Low |
BenchCAD is narrower and more domain-specific than general coding benchmarks — useful specifically for evaluating engineering/manufacturing-adjacent agentic use cases.
Benchgen lets teams track BenchCAD-style CAD/design task performance across model versions over time, catching regressions that a single vendor-reported snapshot would miss.
Benchmark definition based on OpenAI's GPT-6 Astra announcement. State-of-the-art scores sourced from the same and attributed inline. Last updated 2026-09-07.