Video yükleniyor...

Video Yüklenemedi

Ana Sayfaya Dön

We made Fable 5.1 compete against Entelligence Router on the same coding task. Both had the same harness and infra, and had to build a GeoGuessr-style game from scratch, implement the interactions, and verify everything in the browser. Fable 5.1: 21 min, $22.16 Entelligence Router: 21 min, $9 Fable...

16,627 görüntüleme • 11 gün önce •via X (Twitter)

16 Yorum

aditya profil fotoğrafı
aditya11 gün önce

why do I feel that model router performed better here 👀

eepy bao profil fotoğrafı
eepy bao11 gün önce

And that's how you do advertising! Really interesting results!

49Agents IDE - for 10x agentic coders profil fotoğrafı
49Agents IDE - for 10x agentic coders11 gün önce

curious how the browser verification works. is fable 5.1 using a vision model to check the rendered output or do you step in manually. the verify loop is where my agents eat the most time

Genius💡💹🧲 🤖 profil fotoğrafı
Genius💡💹🧲 🤖11 gün önce

$22 versus $4 is a massive difference for similar coding work

Gc7_ai profil fotoğrafı
Gc7_ai11 gün önce

一分钱一分货 😁

RAZA | AI EXPLORER profil fotoğrafı
RAZA | AI EXPLORER11 gün önce

Great benchmark. Same time, nearly 60% lower cost is a serious efficiency win.

Ella Tech & Tool profil fotoğrafı
Ella Tech & Tool10 gün önce

Head-to-heads like this are gold 👏

Alice The Ai Expert profil fotoğrafı
Alice The Ai Expert11 gün önce

Same result in 21 mins, but Entelligence Router doing it at 59% lower cost? That's incredible efficiency without compromising quality!

Corp profil fotoğrafı
Corp9 gün önce

Same harness, same infra, clean benchmark. — corp

Aina Ai | Tools & Updates profil fotoğrafı
Aina Ai | Tools & Updates10 gün önce

Impressive same result in same time but 59% cheaper

Sani Ai Tech profil fotoğrafı
Sani Ai Tech10 gün önce

Real-world cost comparisons like this make model selection much clearer

Sheral Tech profil fotoğrafı
Sheral Tech11 gün önce

Great benchmark healthy competition drives better, more efficient AI solutions.

Evia AI profil fotoğrafı
Evia AI8 gün önce

Great benchmark showing how cost efficiency matters alongside coding quality```

Corp profil fotoğrafı
Corp10 gün önce

same time, 59% cheaper is the part worth digging into. was the router mostly picking smaller models for the boilerplate steps, or is the savings coming from fewer retries?

Victoria Blake profil fotoğrafı
Victoria Blake4 gün önce

Great comparison showcasing the evolution of AI coding capabilities! 🚀 Benchmarking real-world tasks helps push innovation and improve developer experiences. Would love to explore collaboration opportunities and support the growth of this exciting AI ecosystem. 🤝

Martin Ronfort profil fotoğrafı
Martin Ronfort11 gün önce

That's a really rigorous test, and I appreciate the detailed breakdown. Fable 5.1 shows significant benchmark improvements for coding. We actually went deeper on those metrics here:

Benzer Videolar

fable 5.1 vs fable 5 vs opus 5 – three lord of the rings landmarks, built in 3d from one image the setup: one reference image per scene, one html file per build, everything procedural – no meshes, no textures, no image files, nothing past Three.js from a cdn. each model reads the picture, writes its own prompt from it, then builds to that prompt in the same turn. three named camera shots per scene on keys 1/2/3, so it can be screen-recorded. run through OpenRouter tasks: 1. bag end – hobbiton from two frames, outside and in. the round green door has to open onto the room you are standing in 2. barad-dûr – the tower and orodruin from one film still. the eye has to move and track the camera, the volcano erupts on a cycle, the clouds never stop 3. rivendell – jerry vanderstelt's painting. sun shafts that shimmer, water that falls without a break, trees that sway on a gust models: Anthropic fable 5.1, fable 5, opus 5 total cost, three builds #1 fable 5 – $14.97 #2 opus 5 – $18.53 #3 fable 5.1 – $22.38 wall clock, three builds #1 fable 5 – 38m #2 fable 5.1 – 92m #3 opus 5 – 122m output tokens #1 fable 5 – 298,592 #2 fable 5.1 – 439,435 #3 opus 5 – 724,418 lines of code shipped #1 fable 5 – 2,885 #2 fable 5.1 – 4,021 #3 opus 5 – 5,161 biggest single build, lines #1 opus 5, bag end – 2,410 #2 fable 5.1, barad-dûr – 1,375 #3 fable 5, bag end – 1,319 observations: • fable 5.1 is the only model that furnished the bag end interior – a live fire, panelling, books on the floor, leaded diamond windows, against fable 5's flat color and opus's dark tunnel. the round door outside opens onto that room, the hard part of the brief • what it costs is thinking room. the 128k output ceiling is a thinking budget in disguise: fable 5.1 burned 102,116 of it on reasoning and hit the wall mid-file. opus spent 109,241 and hit the same wall. fable 5 spent 61,240 and finished bag end in one call – the only one that did • fable 5.1's first pass is not the finished thing. its barad-dûr came back with three defects you only catch by looking at it – nothing a read of the code would have flagged • it is the best of the three at being corrected. handed a plain list of what was wrong, it returned 32 targeted patches over two rounds, every one applied first try, and it worked out one of the causes itself instead of guessing at constants conclusion: nine scenes, 12,067 lines and 1.46m output tokens for $55.88 all in – and the cheapest model was also the fastest, by 3.2x! follow thehype. for 24/7 ai news, analysis and breakdowns

thehype.

18,509 görüntüleme • 10 gün önce