← index
The CAD Benchmark
a reference benchmark for whether a model actually reasons about a mechanical part — or is pattern-matching on its shape. built from a dataset i'm making myself. results publish november 2026.
every lab training on mechanical design evaluates against whatever data it happens to hold. one team scores on its own scraped step files, another on a supplier's parts library, a third on synthetic geometry it generated last quarter. all three publish a number. none of the three numbers can be compared with either of the others, and nobody can say which model is better at the thing they all claim to do.
that's the gap. not more data — shared data, held out, with a scoring method that everyone can run and nobody controls. that's what i'm building at datak, on top of the cad and simulation datasets i already sell.
| # | Model | Geometry | Edit | Mfg | Sim | Score |
|---|---|---|---|---|---|---|
| 01 | not yet published | — | — | — | — | — |
| 02 | not yet published | — | — | — | — | — |
| 03 | not yet published | — | — | — | — | — |
no scores here yet, and i'm not going to put placeholder numbers on a page that looks like a leaderboard. the board fills when the first run completes — targeting before november 2026.
What it measures
a model can describe a bracket convincingly and still have no idea what happens when you load it. the benchmark separates the two by asking questions that a wrong answer fails visibly — where there is a known-correct result from the underlying geometry or a solver run, not a human opinion.
- geometry comprehension. given a part, identify its features — holes, fillets, ribs, draft, wall thickness, tolerances. answers check against the feature tree in the native file, so there's no grading by vibes.
- parametric edit. change a dimension and keep the constraints valid. this is where shape-matching falls over: a model that has only seen pictures of brackets cannot tell you what else moves when this hole shifts 4 mm.
- manufacturability. can this be machined, cast or moulded as drawn — and if not, which feature is the reason. a real answer names the feature.
- simulation outcome. given a load case, predict where it fails and roughly how much it deflects. scored against the actual solver run — ansys, abaqus or openfoam — that shipped with the part.
Where the data comes from
this is the part that's hard to fake and the reason i think i can build it. datak already sources original cad and simulation data through an exclusive network of design partners — parts authored for us, not scraped, with the manifests and provenance that let a buyer trust geometry they can't eyeball. the benchmark's held-out set is drawn from the same pipeline, which means:
- it's not in anyone's training data, because it was made for this and never published.
- every task has a known-correct answer from the native file or the solver run, not a human rater's judgement.
- i can keep making more of it, so a contaminated set can be retired rather than defended.
Why i'm building it
partly because the category needs it. mostly because of where it sits: sell the data, own the benchmark, then train the model. whoever defines how these systems get measured is best placed to build the one that wins — and the measuring is the cheapest of the three to do first.
the model doesn't exist yet. i'd rather say that plainly here than imply otherwise.