Video wird geladen...
Video konnte nicht geladen werden
wow... GPT 5.6 Sol Passed the Replit Benchmark. In one prompt 5.6 on Codex created a full Replit clone with DB, Sandboxes, and an agent that builds and edits full web apps. Only Fable and GPT 5.6 have passed this benchmark. 5.6 is much better at designing iOS apps... show more
123,642 Aufrufe • vor 2 Monaten •via X (Twitter)
34 Kommentare

First one

If Fable and GPT 5.6 Sol are the only two clearing that bar, that's worth testing side by side rather than picking one on reputation.

GYAT

Much better? 5.5 is already pretty solid!

Never heard of a Replit Benchmark before this post, is that an official eval or something Riley's own team ran? Would love to see the actual task list, one prompt full-stack clones are a big claim. @rileybrown

the Replit clone is the headline, but "5.6 is better at designing iOS apps" is the line that actually matters. building the app is commoditizing fast. the taste to make it feel good is what still separates the models. is that design jump big or subtle?

Ufff, OpenAI is cooking!

So market got more competitive!

Awesome, what was the API cost?

NOT EVEN ULTRA?

Thats great

Whoa, 5.6 is a beast! Replit clone with DB, Sandboxes, and an agent that builds full web apps? Mind. Blown.

I think the best benchmarks should work with unverifiables.

neat. what's the failure rate on runs that don't get screenshotted

did 5.6 handle hostile user code, or only the replit demo path

benchmarks like this are getting way closer to real work, shipping an entire system in one shot is a very different test than solving isolated coding problems

Two models passing Replit Benchmark within the same window says more about the pace than either model individually. The "only we can do this" moat lasts about two weeks now.

Benchmarks are nice, but real-world reliability matters more.

Which one is better Fable or GPT 5.6 Sol ?

Close

The scary part isn't the one prompt clone. It's that the sandbox isolation and agent build loop are the actual hard parts, and those are the ones that break under real load. Curious how it holds up past the demo when 50 users hit those sandboxes at once.

passing a benchmark means squat if you can't scale it to real-world scenarios. fable's impressive, but let's see it handle production-level loads and edge cases before we get too excited

One-prompt benchmarks measure demo ability, not engineering. The real test is prompt #47, when the agent has to modify its own week-old code without breaking it. Curious whether Sol holds up there — 5.5 didn't.

🤖

Good. How long before their antisemite founder is out of a job?

Passing an end-to-end generation benchmark is one signal. The harder one is how cleanly it handles iterative edits on an existing large codebase without regressing unrelated parts.

The gap between AI demos and production apps keeps shrinking.

That's really impressive!

That’s a bit scaring👀

what the. 🤯

Wow, looks really interesting to try out

Sick! Interesting how we have reached the limit where the .gov & big LLM Companies are throttling what we are allowed to use and play with.

built a whole company in a prompt lmao

Put them in scenarios where information is constantly changing and then see how they act? e.g. live stock market

