Video wird geladen...
Video konnte nicht geladen werden
DeepSeek-V4.1-Flash: the BEST Boeing benchmark result I’ve seen from an open-source model (by far). ✈️ It ran for hours with /goal. Where other models often stall quickly, it kept improving: inspect, spot a defect, zoom in, diagnose, fix, repeat. The video really shows how strong it is at sustaining... show more
58,526 Aufrufe • vor 1 Tag •via X (Twitter)
33 Kommentare

more details because this is beautiful 🔥

extremely good, inspect the result here:

kept improving instead of stalling is the only benchmark that matters now. the rest is one shot theatre

"Until you are 100% satisfied" might be the most important part of the prompt. The model grading its own work is the benchmark now.

The hard part in a multi-hour run isn't the fixing. It's still being dissatisfied at hour three.

hours of /goal is a much harsher test than a one-shot render. what finally made it stop improving: a visual score, a fixed iteration limit, or its own judgment?

It seems better than opus 4.8 one right?

Open source running for hours in claude code and still improving. Benches don’t show that.

the thing that separates "keeps improving" from "stalls" isn't raw capability. it's whether the model can still see what it already tried 40 steps back. most stalls are it relitigating a fix it already ruled out because that context fell out of the window.

The impressive part is the loop, not the “god model” label: inspect → name a defect → fix → verify. That pattern is useful outside 3D work too—give any model explicit checks and a stop condition instead of asking for a one-shot answer.

I wonder if it be better on dsh

this is a great benchmark, where can I find more details? are there any common trends you’re seeing with how deepseek or open weight models handle 3D vs closed frontier models?

best way to see other results:

Next time just use a separate critic agent

this Boeing benchmark reminds me of the Red Queen Gödel Machine: continuous self-improvement against a moving target.

what'd it cost?

was the inspect→fix loop running on the same context window the whole time, or did /goal spawn fresh subagents? curious what context strategy kept it from rotting across hours

What really is marvellous- the precision shadow. The moment you move the pane around the plane, the direction also specifies the shadow. That’s something !

What Harness tool are you using?

the harness did the work, not the model. claude code gave it the loop and deepseek just refused to quit. how much of this is /goal's verifier vs deepseek's stamina?

Can you please crash it into a building next

did it reach the final version faster than astra before?

Looks like I am going to change harness, I was not really happy with deepseek results in opencode, maybe it’s a harness problem.

try it and tell me - I get really good results with CC with this setup:

DeepSeek 4.1 being this close to GPT-5.5 at this price point is actually illegal 💀

the impressive part isn't the boeing render, it's that it stayed in the inspect→diagnose→fix loop for hours instead of stalling after the first pretty output. long-horizon models start looking useful when the harness keeps them pointed at the next defect

bro chinese models r literally cooking the entire ai market rn distillation really leveled up chinese ai to the point it's outperforming everything ngl 😭🔥

cost ?

this is the satisfying version of a loop: inspect, diagnose, fix. i still make the deployed path prove the demo; screenshots can be charming little liars.

Interesting to see that it “ran for hours with /goal” without stalling. Long-horizon loops are where most “SOTA” models still quietly die.

boeing test is brutal lol. most open models stall after a few fixes, the fact it kept going for hours and kept improving is the real flex

Does it ever break something that was already working?

你炒股吗?如果你炒股的话马上删了,否则对今晚美国股市有冲击!能设计外形,就可以对飞机内部进行设计。直接冲击了美国高端制造业
