Video wird geladen...

Video konnte nicht geladen werden

Zur Startseite

It's hard to justify using models that are dumber than Fable and Astra. For real-world code work, the benefits massively outweigh the cost. The benefit isn't better code, it's something else that's more subtle.

448,635 Aufrufe • vor 2 Tagen •via X (Twitter)

35 Kommentare

Profilbild von Michael Ramos
Michael Ramosvor 2 Tagen

I agree and it’s bs where this is headed. The frontier becomes a capital advantage to access. Accessing it gives you an immediate cognitive and capability advantage.

Profilbild von Vahagn
Vahagnvor 2 Tagen

Sure when you got a ton a money absolutely but specially with the latest usage cuts from Claude now 2 subs isn’t enough for anything a bit complex in “real-world” scenario you use Fable to control Opus, Astra and Sol

Profilbild von Felipe Afonso
Felipe Afonsovor 2 Tagen

The comments on youtube are so toxic, wtf is up with people. They really didn't watch to the end and can only be doing small and simple tasks. I agree, if you can't see the difference between Sol and Astra, you're dumb.

Profilbild von Tim Soret
Tim Soretvor 2 Tagen

Totally. For anyone making serious work, it's almost impossible to downgrade. I feel I exhausted the previous model capabilities & hit the ceiling, the latest models can go way further and I'm still mapping out their limits.

Profilbild von ˗ˏˋ Jesse Hanley ˎˊ˗
˗ˏˋ Jesse Hanley ˎˊ˗vor 2 Tagen

Poor @zeeg I don't think his streaming setup is ready to respond.

Profilbild von Isaac Way
Isaac Wayvor 2 Tagen

Facts. Glad you said it. Anyone who thinks Opus / Sol level is sufficient simply are not leveraging the latest models effectively. It’s not a “you don’t know how to prompt issue”, it’s a “you don’t understand how big of tasks agents can complete end to end reliably” issue

Profilbild von β/ai
β/aivor 2 Tagen

omg, you'll love this one from dwayne

Profilbild von NiteLite
NiteLitevor 2 Tagen

I tried Astra for a few days, but went back to Sol.

Profilbild von Baron Kimble
Baron Kimblevor 2 Tagen

Sadly this is true. And if you cannot afford $1k-$2k in spend for a serious project to budget with you probably aren't doing it right. But this is what is required to build a business. For me Astra is just too good w/ 3d not to use.

Profilbild von Mario
Mariovor 2 Tagen

this is so true if your tasks perform the same, they aren't pushing these models to their limits, which is fine. also: cheaper models sometimes end up being more expensive.

Profilbild von aizk ✡️
aizk ✡️vor 2 Tagen

At @FeatherlessAI we disagree, and we have Simple Jev to prove it

Profilbild von Matthew Fox
Matthew Foxvor 2 Tagen

I used opus for a small set of sprints First time @greptile gave a or a 1/5 ever Fables almost always 3-4/5 on first pass

Profilbild von Nikolai Yakovenko
Nikolai Yakovenkovor 2 Tagen

I agree you are 100% right. The problem is when the smart models reject you for things like reading news about "politics" or "cyber" or "biology" for Claude To be fair Claude is a bit smarter about applying censorship in context. Astra went the other way. Happens every day

Profilbild von JC Castaneda
JC Castanedavor 2 Tagen

Jev + frontier Chinese LLM model will do the job now

Profilbild von David Helmuth
David Helmuthvor 2 Tagen

This only matters if one can figure out their requirements before asking the magic Genie. For those that don’t know, this can be one of the biggest problems in the industry

Profilbild von Arsh - 16 y/o builder
Arsh - 16 y/o buildervor 2 Tagen

its the harness

Profilbild von ethereagle · building
ethereagle · buildingvor 2 Tagen

zeeg's switch-back tests the finished code. it doesn't test how many times you had to poke the model. was Sol taking extra rounds on the same task, or was review time the thing that actually moved?

Profilbild von ShadowAguy
ShadowAguyvor 2 Tagen

the subtle benefit is probably the feeling of working faster while doing less actual work. productive vibes are a feature too

Profilbild von Jez Robot
Jez Robotvor 2 Tagen

Opus 5 and Sol I think are still solid especially for well-trodden territory.

Profilbild von Floyd DePalma
Floyd DePalmavor 2 Tagen

Great video, Theo. Thank you! I have to say, since I've been using T3 code over the past couple weeks, it's been a lot easier to spin threads and move on to the next thing without babysitting.

Profilbild von Shadowfetch Applications
Shadowfetch Applicationsvor 2 Tagen

This is wrong. Opus 4.8 is better than opus 5. LOL but I get your point.

Profilbild von NDIZEYE
NDIZEYEvor 2 Tagen

those models are not that efficient and not everyone got the budget

Profilbild von Rabbit Toes
Rabbit Toesvor 2 Tagen

Wrong and Astra is dogshit lmao

Profilbild von Arc
Arcvor 2 Tagen

I have the biggest claude and codex plans and astra absolutely fucks my usage I run out in like 3 days, cant use astra for everything

Profilbild von Hurn
Hurnvor 2 Tagen

What the fuck are you talking about

Profilbild von Riqwan Thamir
Riqwan Thamirvor 2 Tagen

Lol, no.

Profilbild von Ryan
Ryanvor 2 Tagen

the benefit is fewer supervision interrupts. smarter models recover from bad intermediate states instead of turning every ambiguity into another human round trip.

Profilbild von Luca Capone | Building with AI
Luca Capone | Building with AIvor 2 Tagen

The subtle bit is momentum. A dumber model makes me stop and re-explain myself every few minutes. A smarter one keeps me moving through the whole hour I get after work.

Profilbild von AIDoomScroll
AIDoomScrollvor 2 Tagen

can't hide dumb models behind benchmarks forever

Profilbild von Chris Izatt
Chris Izattvor 2 Tagen

Aye aye sir

Profilbild von Kail
Kailvor 2 Tagen

it's not the code. it's that you stop babysitting. dumb models write fine functions then lose the plot. smart ones hold intent while you look away. that's what you're paying for.

Profilbild von Jon Coulter
Jon Coultervor 2 Tagen

I agree with the overall take but the cost of the top tier models makes me “worry” more than I want.

Profilbild von Mike Bradley
Mike Bradleyvor 2 Tagen

I’m getting better results with Sol than Astra on long time horizon coding work and goals, and the economics have been crazy better.

Profilbild von Larpologist
Larpologistvor 2 Tagen

this needed a 30 minute video?

Profilbild von John Triplette 🇺🇸
John Triplette 🇺🇸vor 2 Tagen

Shill says what

Ähnliche Videos

gpt astra vs fable 5.1 at goldberg machine gpt 6 astra – openai, landed on OpenRouter less then hour ago, provider pinned to openai fable 5.1 – anthropic, shipped sep 1 we put the two models on one job: a rube goldberg machine in three.js that presses a button and detonates a bomb the setup: one self-contained html file, three.js from a cdn, everything else procedural – no textures, no models, no physics engine, every collision hand-written. the hard part sits in the brief: a domino may only fall once the previous one actually touches it, checked by real overlap every frame, never by a timer. same rule for the hammer hitting the button and the button firing the bomb. one continuous camera, its speed driven by whatever is moving. we recorded both scenes frame by frame – 1200 frames, 60 fps, exactly 20 seconds – and stepped both by hand to read the telemetry. - cost #1 astra – $1.84 #2 fable – $29.16 - time #1 astra – 9m 56s #2 fable – 1h 12m - tokens #1 astra – 45k #2 fable – 360k - lines of code astra – 881 fable – 744 observations: • we told it what we saw and nothing else – no diagnosis, no patch. we never edit a model's code. round two ran the whole chain to the blast. • both files are deterministic. two runs each, identical state to twelve decimals, and neither model reached for math.random. conclusion: 15.8x cheaper and 7.2x faster, and it still took a second round to get the ball into the bucket! follow thehype. for 24/7 ai news, analysis and breakdowns

thehype.

37,224 Aufrufe • vor 16 Tagen

AI is changing the software engineering craft. Anders Hejlsberg (Anders Hejlsberg) - creator of C#, TypeScript and industry legend - on why code review needs to get more enjoyable in response: #1 - AI is shifting the craft from writing code, to reviewing code: "In a sense, we're all turning into project managers. We can have an army of junior programmers, called agents, that will just spit out reams of code but someone's got to have the big picture and review all of that. And so, increasingly, our craft is going from one of writing the code, to one of reviewing the code and building the architecture of the code and overseeing the work. It's a different kind of craft. It's a different kind of enjoyment. I've always liked writing the code. To me that was the fulfilling part, seeing it work. In a way, AI robs a little bit of that, because I am less interested in reviewing code." #2 - The code review experience should be improved: "I think we could also make the process of reviewing code much more interesting than it is today. I mean, today, you see a list of diffs in alphabetical order and now it's up to you to make heads or tails of it. There are more pedagogical ways of presenting that. And you could have commentary generated by the AI that tells you what the changes are and whatever, and then tries to guide you along. So that symbiotic relationship, I think we need to work on that more and to keep the enjoyment in there."

The Pragmatic Engineer

39,073 Aufrufe • vor 4 Monaten