Benchgen

BenchCAD — Results

RankModelScore
1gpt-6-astra95.9
B

BenchCAD

1 phaseActive

CAD/parametric-design benchmark testing whether models can produce and edit computer-aided designs against a spec.

Overview

BenchCAD

Category Metric Saturation Created

Quick answer: BenchCAD is a benchmark measuring whether AI models can produce and edit computer-aided design (CAD) artifacts — from parametric modeling to spec-following geometry edits. It first appeared in OpenAI's GPT-6 Astra announcement (September 2026), where Astra scored 95.9%, ahead of Claude Fable 5.1 (84.3%) and Claude Opus 5 (82.1%).

At a Glance

What it tests: Whether a model can generate, modify, and reason about CAD designs — parametric models, dimensional constraints, and spec compliance — a professional engineering-adjacent skill distinct from general code generation.

Why it matters: CAD work sits at the intersection of spatial reasoning, precise numerical constraint-following, and domain-specific tooling. Strong BenchCAD performance signals a model is viable for engineering, product design, and manufacturing-adjacent agentic workflows, not just software engineering.

Known limitations: Full methodology (task count, dataset composition, grading rubric) has not yet been publicly documented in detail as of this writing — the benchmark is new and primarily known through frontier lab technical reports rather than an independent paper or public dataset.

What BenchCAD Measures

BenchCAD evaluates models on computer-aided design tasks — the kind of parametric modeling and spec-following work done in professional CAD software. This differentiates it from general-purpose coding benchmarks, since CAD tasks require reasoning about physical geometry, dimensional constraints, and manufacturability alongside standard tool-use and instruction-following.

As reported in OpenAI's GPT-6 Astra launch materials, results for Claude models on this benchmark reflect three modifications made to the evaluation (per Anthropic's own Claude Fable 5.1 system card), so direct comparisons across labs should be read with that caveat in mind — differing eval harnesses or task modifications can shift scores independent of underlying model capability.

Benchmark Specifications

FieldValue
Task categoryAgent / Professional
Metric% task success
SaturationMedium — top models cluster in the 80-96% range
Created byOpenAI (introduced via GPT-6 Astra announcement)
Public datasetNot yet publicly released as of Sep 2026

How BenchCAD Is Scored

Scores are reported as a percentage task-success rate. As with many frontier-lab-reported benchmarks, exact grading criteria (partial credit for near-correct geometry, tolerance thresholds for dimensional accuracy) have not been fully disclosed publicly.

State-of-the-Art Results

RankModelScoreSourceDate
1GPT-6 Astra95.9%OpenAI: GPT-6 Astra2026-09
2Claude Fable 5.184.3%OpenAI: GPT-6 Astra (via Anthropic system card)2026-09
3Claude Opus 582.1%OpenAI: GPT-6 Astra2026-09
4GPT-5.6 Sol83.3%OpenAI: GPT-6 Astra2026-09
5Claude Fable 567.5%OpenAI: GPT-6 Astra (via Anthropic system card)2026-09

Scores sourced from OpenAI's GPT-6 Astra announcement (September 2026). Claude scores reflect 3 modifications to the eval per Anthropic's Claude Fable 5.1 system card — see source for methodology differences.

BenchCAD on Benchgen

No Benchgen results yet — be the first to run BenchCAD.

BenchCAD vs Other Benchmarks

BenchmarkWhat it testsSaturation
BenchCADCAD/parametric design generation and editingMedium
FrontierCodeGeneral agentic coding, rubric-weightedLow
AutomationBenchSaaS/REST-API workflow automationLow

BenchCAD is narrower and more domain-specific than general coding benchmarks — useful specifically for evaluating engineering/manufacturing-adjacent agentic use cases.

Run BenchCAD on Your Model

Benchgen lets teams track BenchCAD-style CAD/design task performance across model versions over time, catching regressions that a single vendor-reported snapshot would miss.

Frequently Asked Questions

What is BenchCAD? BenchCAD is a benchmark testing whether AI models can generate and edit computer-aided design (CAD) artifacts against a spec, introduced publicly via OpenAI's GPT-6 Astra announcement in September 2026.
What does a good BenchCAD score look like? As of September 2026, frontier models score in the 67-96% range, with GPT-6 Astra leading at 95.9%.
Who created BenchCAD? BenchCAD was introduced by OpenAI, first publicly referenced in the GPT-6 Astra model announcement (September 2026).
Is BenchCAD saturated? Not yet — top scores cluster in the 80-96% range with meaningful spread between models, indicating room for further differentiation.

Benchmark definition based on OpenAI's GPT-6 Astra announcement. State-of-the-art scores sourced from the same and attributed inline. Last updated 2026-09-07.