Video wird geladen...

Video konnte nicht geladen werden

Zur Startseite

Claude Opus 5 built a substantially better extreme-ocean project than GPT-5.6 Sol on the same prompt, spending almost 10x more while autonomously expanding and refining far beyond what was asked with no loop instruction.

48,758 Aufrufe • vor 27 Tagen •via X (Twitter)

8 Kommentare

Profilbild von Shinra
Shinravor 27 Tagen

By the way, Grok would have handled it much better.

Profilbild von AI Model Ranking
AI Model Rankingvor 27 Tagen

👀

Profilbild von 90S KID
90S KIDvor 27 Tagen

No loop instruction and it kept going anyway. We used to call that a runaway process.

Profilbild von John
Johnvor 27 Tagen

Crazy man the future is gonna be wild

Profilbild von Glitch Truth
Glitch Truthvor 27 Tagen

10x more work, way better output. Usually that doesn't happen. Claude's not doing 10 times the thinking, it's more checking, more steps before answering. Costs more, locks in the right answer. GPT does fewer checks, answers faster. Same brain, different way of working.

Profilbild von vøv △
vøv △vor 27 Tagen

no loops, just vibes. insane

Profilbild von heroyuma🦖
heroyuma🦖vor 26 Tagen

$Abyssal 3v3MymJvAoH2P4kb3uc2XqDX5LSK3ZyR1E3DuqSkpump Wow, how is this still only at 2K? 👀

Profilbild von Alex
Alexvor 27 Tagen

spending 10x and autonomously expanding scope without being asked is interesting, but did the output actually need all that? cost isn't bad if the result is 10x better, but if it's marginally better you just paid a Claude tax

Ähnliche Videos

GPT-5.6 vs GPT-5.5 on my custom spaceship prompt. I gave both models the exact same custom prompt. This is also the same prompt I previously gave to Fable 5. For context, GPT-5.6 Pro worked for 87 minutes, while GPT-5.5 Extra High worked for 34 minutes and 42 seconds. As I’ve said before, based on great authority GPT-5.6 will be an incremental/soldi improvement over GPT-5.5, not a “Fable killer.” My rough expectation has been that it would trade blows with Fable 5 on some benchmarks, maybe win around half depending on the category, but not clearly surpass it overall. And again fable five will have bigger model smell, but this was expected. After testing this coding output, that view feels pretty accurate. GPT-5.6 is clearly better than GPT-5.5 in several visual areas. The lighting, shading, chairs, object details, and exterior of the spaceship looked noticeably stronger. The scene was also easier to test. I do want to give GPT-5.5 credit though. It built out the rooms much much better and the planets looked better than GPT-5.6’s. It was also interesting that both GPT-5.5 and GPT-5.6 produced better-looking planets than Fable 5 in this specific test. The downside with GPT-5.5 was stability. The game was much glitchier and harder to test compared to GPT-5.6. But when it comes to the core of the demo, which is the spaceship itself, Fable 5 still beat both models pretty comfortably. GPT-5.6 is impressive, but from this test, it looks exactly like what I expected which was a meaningful incremental improvement over GPT-5.5, at least for indie game demos, but not something that replaces Fable 5. In collaboration with Chetaslua

Chris

250,919 Aufrufe • vor 3 Monaten