Video wird geladen...

Video konnte nicht geladen werden

Zur Startseite

wow... GPT 5.6 Sol Passed the Replit Benchmark. In one prompt 5.6 on Codex created a full Replit clone with DB, Sandboxes, and an agent that builds and edits full web apps. Only Fable and GPT 5.6 have passed this benchmark. 5.6 is much better at designing iOS apps...

123,642 Aufrufe • vor 2 Monaten •via X (Twitter)

34 Kommentare

Profilbild von Riley Brown
Riley Brownvor 2 Monaten

First one

Profilbild von Creao AI
Creao AIvor 2 Monaten

If Fable and GPT 5.6 Sol are the only two clearing that bar, that's worth testing side by side rather than picking one on reputation.

Profilbild von ashen
ashenvor 2 Monaten

GYAT

Profilbild von StatysTheBaddest
StatysTheBaddestvor 2 Monaten

Much better? 5.5 is already pretty solid!

Profilbild von Vance ai
Vance aivor 2 Monaten

Never heard of a Replit Benchmark before this post, is that an official eval or something Riley's own team ran? Would love to see the actual task list, one prompt full-stack clones are a big claim. @rileybrown

Profilbild von Marko Vineli
Marko Vinelivor 2 Monaten

the Replit clone is the headline, but "5.6 is better at designing iOS apps" is the line that actually matters. building the app is commoditizing fast. the taste to make it feel good is what still separates the models. is that design jump big or subtle?

Profilbild von Richard Abishai
Richard Abishaivor 2 Monaten

Ufff, OpenAI is cooking!

Profilbild von Ahmed
Ahmedvor 2 Monaten

So market got more competitive!

Profilbild von Fred Marks
Fred Marksvor 2 Monaten

Awesome, what was the API cost?

Profilbild von fact
fact vor 2 Monaten

NOT EVEN ULTRA?

Profilbild von Joebed Arts
Joebed Artsvor 2 Monaten

Thats great

Profilbild von Mahendra Kumawat
Mahendra Kumawatvor 2 Monaten

Whoa, 5.6 is a beast! Replit clone with DB, Sandboxes, and an agent that builds full web apps? Mind. Blown.

Profilbild von Sense Noped Out
Sense Noped Outvor 2 Monaten

I think the best benchmarks should work with unverifiables.

Profilbild von Orion Night
Orion Nightvor 2 Monaten

neat. what's the failure rate on runs that don't get screenshotted

Profilbild von Sebastian Buzdugan
Sebastian Buzduganvor 2 Monaten

did 5.6 handle hostile user code, or only the replit demo path

Profilbild von Rameswar
Rameswarvor 2 Monaten

benchmarks like this are getting way closer to real work, shipping an entire system in one shot is a very different test than solving isolated coding problems

Profilbild von Vik
Vikvor 2 Monaten

Two models passing Replit Benchmark within the same window says more about the pace than either model individually. The "only we can do this" moat lasts about two weeks now.

Profilbild von Steven Cheng
Steven Chengvor 2 Monaten

Benchmarks are nice, but real-world reliability matters more.

Profilbild von ahmetfatiheren
ahmetfatiherenvor 2 Monaten

Which one is better Fable or GPT 5.6 Sol ?

Profilbild von Riley Brown
Riley Brownvor 2 Monaten

Close

Profilbild von Fadi Hares
Fadi Haresvor 2 Monaten

The scary part isn't the one prompt clone. It's that the sandbox isolation and agent build loop are the actual hard parts, and those are the ones that break under real load. Curious how it holds up past the demo when 50 users hit those sandboxes at once.

Profilbild von Adel Bucetta
Adel Bucettavor 2 Monaten

passing a benchmark means squat if you can't scale it to real-world scenarios. fable's impressive, but let's see it handle production-level loads and edge cases before we get too excited

Profilbild von youfeng
youfengvor 2 Monaten

One-prompt benchmarks measure demo ability, not engineering. The real test is prompt #47, when the agent has to modify its own week-old code without breaking it. Curious whether Sol holds up there — 5.5 didn't.

Profilbild von it8.art
it8.artvor 2 Monaten

🤖

Profilbild von Sam Nissim
Sam Nissimvor 2 Monaten

Good. How long before their antisemite founder is out of a job?

Profilbild von Ofek Shaked | AI Engineer
Ofek Shaked | AI Engineervor 2 Monaten

Passing an end-to-end generation benchmark is one signal. The harder one is how cleanly it handles iterative edits on an existing large codebase without regressing unrelated parts.

Profilbild von Pearl AI
Pearl AIvor 2 Monaten

The gap between AI demos and production apps keeps shrinking.

Profilbild von Jaitan Martini
Jaitan Martinivor 2 Monaten

That's really impressive!

Profilbild von Ivelin · building ProofNod 🚀
Ivelin · building ProofNod 🚀vor 2 Monaten

That’s a bit scaring👀

Profilbild von Philippe Tremblay
Philippe Tremblayvor 2 Monaten

what the. 🤯

Profilbild von Derek | Software Dev
Derek | Software Devvor 2 Monaten

Wow, looks really interesting to try out

Profilbild von One American Voice
One American Voicevor 2 Monaten

Sick! Interesting how we have reached the limit where the .gov & big LLM Companies are throttling what we are allowed to use and play with.

Profilbild von noted
notedvor 2 Monaten

built a whole company in a prompt lmao

Profilbild von Sudhanshu
Sudhanshuvor 2 Monaten

Put them in scenarios where information is constantly changing and then see how they act? e.g. live stock market

Ähnliche Videos

GPT-5.6 vs GPT-5.5 on my custom spaceship prompt. I gave both models the exact same custom prompt. This is also the same prompt I previously gave to Fable 5. For context, GPT-5.6 Pro worked for 87 minutes, while GPT-5.5 Extra High worked for 34 minutes and 42 seconds. As I’ve said before, based on great authority GPT-5.6 will be an incremental/soldi improvement over GPT-5.5, not a “Fable killer.” My rough expectation has been that it would trade blows with Fable 5 on some benchmarks, maybe win around half depending on the category, but not clearly surpass it overall. And again fable five will have bigger model smell, but this was expected. After testing this coding output, that view feels pretty accurate. GPT-5.6 is clearly better than GPT-5.5 in several visual areas. The lighting, shading, chairs, object details, and exterior of the spaceship looked noticeably stronger. The scene was also easier to test. I do want to give GPT-5.5 credit though. It built out the rooms much much better and the planets looked better than GPT-5.6’s. It was also interesting that both GPT-5.5 and GPT-5.6 produced better-looking planets than Fable 5 in this specific test. The downside with GPT-5.5 was stability. The game was much glitchier and harder to test compared to GPT-5.6. But when it comes to the core of the demo, which is the spaceship itself, Fable 5 still beat both models pretty comfortably. GPT-5.6 is impressive, but from this test, it looks exactly like what I expected which was a meaningful incremental improvement over GPT-5.5, at least for indie game demos, but not something that replaces Fable 5. In collaboration with Chetaslua

Chris

250,919 Aufrufe • vor 3 Monaten