Video yükleniyor...
Video Yüklenemedi
Can frontier language models turn a technical drawing into real CAD? We ran the experiment. Gemini 3.1 Pro, Claude Opus 4.7, and GPT-5.4 got a set of technical drawings and were asked to reproduce a part as a valid 3D CAD model. Still a far way to go, but... show more
173,014 görüntüleme • 5 ay önce •via X (Twitter)
39 Yorum

This is the start of a bigger push. Over the coming weeks we'll be releasing a proper benchmark for AI-driven CAD generation, along with tooling to support it. DM me if you want to be part of it.

Changed my settings here. DMs should work.

Each model had access to the build123d Python library, with rendered feedback and a B-rep validation routine running after every iteration, all inside a 1M-token context window to iterate in a loop. We ran all three on default reasoning settings.

Have you seen what @Reshef_ is doing in this space?

I am rendering uavs, with Claude code, first in Autocad but now moving to onshape to do parts and pieces..

Hi there! Would love to chat if there is anything I can help with. I’m building this:

I'd prefer a specialized CAD model than an LLM solving my CAD

What would make this a real benchmark: full NIST PMI suite (dozens of parts across difficulty tiers), n≥5 per model per part with variance, open-source identical agent scaffold, parseable/manufacturable validation criteria (not rendered thumbnails), and published alongside the failure cases. Until then this is a vibes-check with production values.

ChatGPT be like..

Really interesting benchmark! I'm curious, inside the feedback loop, can it read the image again when needed or only once? And what tools do they have access to, besides the main build123d library?

Incredible progress. The real value here isn't just "automated CAD"—it’s the B-rep validation loop. In early-stage structural design, the "interpretation gap" between a concept and a buildable CAD part is where most errors are born. By using frontier models as "Digital Apprentices" that check their own geometry, we achieve a level of Absolute Structural Certainty that manual drafting simply can't match. Great work on the CAD Benchmark.

Great benchmark! I have been trying to make scale drawings in SWG and then re-skin them with image gen models with mixed results (including new OAI image v2)

Try Chinese ones

Try freecad and python too.

@nithinkd I also wonder if this could benefit from a detailed skill.md file to break down the process and help it along.

Awesome work, the NIST dataset is hard! If you haven't yet, try Model Mania, I found that it's a bit easier for the agent. What kind of info is in the rendered validation?

I still believe this could be handled programmatically better than having an oversized neural try and figure it out manually.

This is so cool. We're prototyping a similar tool at Autonomous. The biggest pain point in having more hardware in the world is 3D design. The easier we can make 3D design, the more hardware can be created. Why did you guys choose build123d vs others like CAD Query?

Build123d is great and seems to have a lot of momentum in development right now. Having said that: 1. My focus is on a benchmark that will be out soon. Having a simple harness as a baseline. 2. Curious to see performance difference on that baseline when switching to CadQuery

This is super interesting and relevant to what we're doing. No official benchmarks yet. More like just see how it goes. But so far, we find that models don't make a big difference. We are using open weight models and they achieve similar result to frontier closed models. DeepSeek and Qwen give similar performance to Opus. At least for simple objects for hobbyist 3d printing. We also find that the agentic frameworks help a lot. Claude Code and OpenCode work out pretty well. I haven't tried with Codex yet but I assume it will give similar harness performance. We're also working on a few different subagents and skills to help with the entire pipeline. And a few tools like slicer to make the entire process end to end. So hopefully if we get this working, people can just "vibe 3d design" and going from prompt to 3d printer to having an object in their living room in an hour. Where can we learn more about what you guys are doing? We'll open source our work soon like in a couple of weeks!

Sounds amazing. I’ll update here. We will have something out soon which will allow you and others to test their toolings.

How do the FOSS models compare?

You can upload the experiment dataset as an eval to @huggingface

I want to be able to provide a CAD to an LLM, and 200 images of all the components, and have it use a program like Cinema4D to create materials and textures automatically. This is like 70% of my day job and I hate doing it.

the gap between looking right and actually manufacturable is still massive. not holding my breath

Someone should do an eCAD benchmark for circuits/PCBs too.

gemini in orange and claude in blue feels like a crime

@Thom_Wolf Our results align with yours— Gemini 3.1 Pro is the clear leader for CAD, but it isn't quite perfect yet in terms of fine details and complex topological structures.

Gemini 41m.. 💀💀

간절히 바라는 기능입니다!🙏

b123 needs extra interfacing. CadQuery and Onshape are better with each having quirks and use cases. there is still lots of rules and plumbing required. and that before optimizing OCC decades old kernel logic into rust for faster typesafe processing.

👍

Single prompt or enriched?

Fun fact, the team from Cursor pivoted away from CAD to code early on. Had they not, we might’ve had ‘vibe CAD’ before we had ‘vibe code’ 🤔. Counterfactuals aside, it seems to be an idea whose time has come.

ChatGPT ist like: Here it is, whatever 🤷🏻♂️

まずい、

Curious to see how each model can deal with PMI/MBD - like datums or some surface callout - when reverse engineering models or prints.

Question: to make that technical drawing, do you need more or less time than to make it in cad directly?

It’s far too early for success, AI making solid manifolded body with right bevel right now - impossible. Maybe in few months.
