Video yükleniyor...

Video Yüklenemedi

Ana Sayfaya Dön

Can frontier language models turn a technical drawing into real CAD? We ran the experiment. Gemini 3.1 Pro, Claude Opus 4.7, and GPT-5.4 got a set of technical drawings and were asked to reproduce a part as a valid 3D CAD model. Still a far way to go, but...

173,014 görüntüleme • 5 ay önce •via X (Twitter)

39 Yorum

Michael Rabinovich profil fotoğrafı
Michael Rabinovich5 ay önce

This is the start of a bigger push. Over the coming weeks we'll be releasing a proper benchmark for AI-driven CAD generation, along with tooling to support it. DM me if you want to be part of it.

Michael Rabinovich profil fotoğrafı
Michael Rabinovich5 ay önce

Changed my settings here. DMs should work.

Michael Rabinovich profil fotoğrafı
Michael Rabinovich5 ay önce

Each model had access to the build123d Python library, with rendered feedback and a B-rep validation routine running after every iteration, all inside a 1M-token context window to iterate in a loop. We ran all three on default reasoning settings.

Corey Ward profil fotoğrafı
Corey Ward5 ay önce

Have you seen what @Reshef_ is doing in this space?

Chris Zazula profil fotoğrafı
Chris Zazula5 ay önce

I am rendering uavs, with Claude code, first in Autocad but now moving to onshape to do parts and pieces..

Stepan Kukharskiy profil fotoğrafı
Stepan Kukharskiy5 ay önce

Hi there! Would love to chat if there is anything I can help with. I’m building this:

Alvaro L profil fotoğrafı
Alvaro L5 ay önce

I'd prefer a specialized CAD model than an LLM solving my CAD

Dr. Disclosure profil fotoğrafı
Dr. Disclosure5 ay önce

What would make this a real benchmark: full NIST PMI suite (dozens of parts across difficulty tiers), n≥5 per model per part with variance, open-source identical agent scaffold, parseable/manufacturable validation criteria (not rendered thumbnails), and published alongside the failure cases. Until then this is a vibes-check with production values.

MnM Fehler profil fotoğrafı
MnM Fehler5 ay önce

ChatGPT be like..

Rafa profil fotoğrafı
Rafa5 ay önce

Really interesting benchmark! I'm curious, inside the feedback loop, can it read the image again when needed or only once? And what tools do they have access to, besides the main build123d library?

Marcin Kasiak profil fotoğrafı
Marcin Kasiak5 ay önce

Incredible progress. The real value here isn't just "automated CAD"—it’s the B-rep validation loop. In early-stage structural design, the "interpretation gap" between a concept and a buildable CAD part is where most errors are born. By using frontier models as "Digital Apprentices" that check their own geometry, we achieve a level of Absolute Structural Certainty that manual drafting simply can't match. Great work on the CAD Benchmark.

James Sanders ⚛🌍 profil fotoğrafı
James Sanders ⚛🌍5 ay önce

Great benchmark! I have been trying to make scale drawings in SWG and then re-skin them with image gen models with mixed results (including new OAI image v2)

Nero🍀 profil fotoğrafı
Nero🍀5 ay önce

Try Chinese ones

Listentotheradio profil fotoğrafı
Listentotheradio5 ay önce

Try freecad and python too.

Blaine Wilson profil fotoğrafı
Blaine Wilson5 ay önce

@nithinkd I also wonder if this could benefit from a detailed skill.md file to break down the process and help it along.

Reshef profil fotoğrafı
Reshef5 ay önce

Awesome work, the NIST dataset is hard! If you haven't yet, try Model Mania, I found that it's a bit easier for the agent. What kind of info is in the rendered validation?

Militant Hitchhiker ♥ profil fotoğrafı
Militant Hitchhiker ♥5 ay önce

I still believe this could be handled programmatically better than having an oversized neural try and figure it out manually.

Dee profil fotoğrafı
Dee4 ay önce

This is so cool. We're prototyping a similar tool at Autonomous. The biggest pain point in having more hardware in the world is 3D design. The easier we can make 3D design, the more hardware can be created. Why did you guys choose build123d vs others like CAD Query?

Michael Rabinovich profil fotoğrafı
Michael Rabinovich4 ay önce

Build123d is great and seems to have a lot of momentum in development right now. Having said that: 1. My focus is on a benchmark that will be out soon. Having a simple harness as a baseline. 2. Curious to see performance difference on that baseline when switching to CadQuery

Dee profil fotoğrafı
Dee4 ay önce

This is super interesting and relevant to what we're doing. No official benchmarks yet. More like just see how it goes. But so far, we find that models don't make a big difference. We are using open weight models and they achieve similar result to frontier closed models. DeepSeek and Qwen give similar performance to Opus. At least for simple objects for hobbyist 3d printing. We also find that the agentic frameworks help a lot. Claude Code and OpenCode work out pretty well. I haven't tried with Codex yet but I assume it will give similar harness performance. We're also working on a few different subagents and skills to help with the entire pipeline. And a few tools like slicer to make the entire process end to end. So hopefully if we get this working, people can just "vibe 3d design" and going from prompt to 3d printer to having an object in their living room in an hour. Where can we learn more about what you guys are doing? We'll open source our work soon like in a couple of weeks!

Michael Rabinovich profil fotoğrafı
Michael Rabinovich4 ay önce

Sounds amazing. I’ll update here. We will have something out soon which will allow you and others to test their toolings.

A Concerned Human profil fotoğrafı
A Concerned Human5 ay önce

How do the FOSS models compare?

Kirito (e/acc) 🏴‍☠️ profil fotoğrafı
Kirito (e/acc) 🏴‍☠️5 ay önce

You can upload the experiment dataset as an eval to @huggingface

Brandon Smith profil fotoğrafı
Brandon Smith5 ay önce

I want to be able to provide a CAD to an LLM, and 200 images of all the components, and have it use a program like Cinema4D to create materials and textures automatically. This is like 70% of my day job and I hate doing it.

nick profil fotoğrafı
nick5 ay önce

the gap between looking right and actually manufacturable is still massive. not holding my breath

gener8or profil fotoğrafı
gener8or5 ay önce

Someone should do an eCAD benchmark for circuits/PCBs too.

fhub profil fotoğrafı
fhub5 ay önce

gemini in orange and claude in blue feels like a crime

C3ll256 profil fotoğrafı
C3ll2565 ay önce

@Thom_Wolf Our results align with yours— Gemini 3.1 Pro is the clear leader for CAD, but it isn't quite perfect yet in terms of fine details and complex topological structures.

세카르 | Sekhar profil fotoğrafı
세카르 | Sekhar5 ay önce

Gemini 41m.. 💀💀

Alltaku profil fotoğrafı
Alltaku5 ay önce

간절히 바라는 기능입니다!🙏

Draja44 profil fotoğrafı
Draja445 ay önce

b123 needs extra interfacing. CadQuery and Onshape are better with each having quirks and use cases. there is still lots of rules and plumbing required. and that before optimizing OCC decades old kernel logic into rust for faster typesafe processing.

Grit( Latest AI NEWS ) profil fotoğrafı
Grit( Latest AI NEWS )5 ay önce

👍

Aaron Green profil fotoğrafı
Aaron Green5 ay önce

Single prompt or enriched?

gener8or profil fotoğrafı
gener8or5 ay önce

Fun fact, the team from Cursor pivoted away from CAD to code early on. Had they not, we might’ve had ‘vibe CAD’ before we had ‘vibe code’ 🤔. Counterfactuals aside, it seems to be an idea whose time has come.

Nico Schulz profil fotoğrafı
Nico Schulz5 ay önce

ChatGPT ist like: Here it is, whatever 🤷🏻‍♂️

フレッシュボーイ profil fotoğrafı
フレッシュボーイ5 ay önce

まずい、

Subcrypt profil fotoğrafı
Subcrypt5 ay önce

Curious to see how each model can deal with PMI/MBD - like datums or some surface callout - when reverse engineering models or prints.

Emanuele profil fotoğrafı
Emanuele5 ay önce

Question: to make that technical drawing, do you need more or less time than to make it in cad directly?

Siliconway profil fotoğrafı
Siliconway5 ay önce

It’s far too early for success, AI making solid manifolded body with right bevel right now - impossible. Maybe in few months.

Benzer Videolar