Video wird geladen...
Video konnte nicht geladen werden
It's hard to justify using models that are dumber than Fable and Astra. For real-world code work, the benefits massively outweigh the cost. The benefit isn't better code, it's something else that's more subtle.
448,635 Aufrufe • vor 2 Tagen •via X (Twitter)
35 Kommentare

I agree and it’s bs where this is headed. The frontier becomes a capital advantage to access. Accessing it gives you an immediate cognitive and capability advantage.

Sure when you got a ton a money absolutely but specially with the latest usage cuts from Claude now 2 subs isn’t enough for anything a bit complex in “real-world” scenario you use Fable to control Opus, Astra and Sol

The comments on youtube are so toxic, wtf is up with people. They really didn't watch to the end and can only be doing small and simple tasks. I agree, if you can't see the difference between Sol and Astra, you're dumb.

Totally. For anyone making serious work, it's almost impossible to downgrade. I feel I exhausted the previous model capabilities & hit the ceiling, the latest models can go way further and I'm still mapping out their limits.

Poor @zeeg I don't think his streaming setup is ready to respond.

Facts. Glad you said it. Anyone who thinks Opus / Sol level is sufficient simply are not leveraging the latest models effectively. It’s not a “you don’t know how to prompt issue”, it’s a “you don’t understand how big of tasks agents can complete end to end reliably” issue

omg, you'll love this one from dwayne

I tried Astra for a few days, but went back to Sol.

Sadly this is true. And if you cannot afford $1k-$2k in spend for a serious project to budget with you probably aren't doing it right. But this is what is required to build a business. For me Astra is just too good w/ 3d not to use.

this is so true if your tasks perform the same, they aren't pushing these models to their limits, which is fine. also: cheaper models sometimes end up being more expensive.

At @FeatherlessAI we disagree, and we have Simple Jev to prove it

I used opus for a small set of sprints First time @greptile gave a or a 1/5 ever Fables almost always 3-4/5 on first pass

I agree you are 100% right. The problem is when the smart models reject you for things like reading news about "politics" or "cyber" or "biology" for Claude To be fair Claude is a bit smarter about applying censorship in context. Astra went the other way. Happens every day

Jev + frontier Chinese LLM model will do the job now

This only matters if one can figure out their requirements before asking the magic Genie. For those that don’t know, this can be one of the biggest problems in the industry

its the harness

zeeg's switch-back tests the finished code. it doesn't test how many times you had to poke the model. was Sol taking extra rounds on the same task, or was review time the thing that actually moved?

the subtle benefit is probably the feeling of working faster while doing less actual work. productive vibes are a feature too

Opus 5 and Sol I think are still solid especially for well-trodden territory.

Great video, Theo. Thank you! I have to say, since I've been using T3 code over the past couple weeks, it's been a lot easier to spin threads and move on to the next thing without babysitting.

This is wrong. Opus 4.8 is better than opus 5. LOL but I get your point.

those models are not that efficient and not everyone got the budget

Wrong and Astra is dogshit lmao

I have the biggest claude and codex plans and astra absolutely fucks my usage I run out in like 3 days, cant use astra for everything

What the fuck are you talking about

Lol, no.

the benefit is fewer supervision interrupts. smarter models recover from bad intermediate states instead of turning every ambiguity into another human round trip.

The subtle bit is momentum. A dumber model makes me stop and re-explain myself every few minutes. A smarter one keeps me moving through the whole hour I get after work.

can't hide dumb models behind benchmarks forever

Aye aye sir

it's not the code. it's that you stop babysitting. dumb models write fine functions then lose the plot. smart ones hold intent while you look away. that's what you're paying for.

I agree with the overall take but the cost of the top tier models makes me “worry” more than I want.

I’m getting better results with Sol than Astra on long time horizon coding work and goals, and the economics have been crazy better.

this needed a 30 minute video?

Shill says what

