Video wird geladen...
Video konnte nicht geladen werden
Boris Cherny, Claude Code creator at Anthrpic: "The more general model will always outperform the more specific model. Don't try to use tiny models for stuff. Don't try to fine-tune. Don't try to do any of this stuff. There are some applications, there are some reasons to do this,... show more
293,396 Aufrufe • vor 3 Tagen •via X (Twitter)
33 Kommentare

easy to say when the big model is free for you :) my token bill might have a slightly diferent opinion...

It's purely coincidental that Anthropic's entire business model is general purpose models

That categorically incorrect. The more studies conducted the otherwise is proven - many small models fine tuned on a task outperform by far the huge generalists. But I understand the commercial and hype motivated claims here. That is otherwise misunderstanding how these NNs work

He is right. Do not try to pay 10x less for the same job, please, my bonus depends on users not optimizing their token use. Will be difficult for you, you will need to learn stuff, it is too complex. Just pay.

Convenient position for the people selling the general model )

Lmao selling shovels. God I hate this retard. It's been proven many times that fine tuned models perform much better than frontier on specific tasks

Very interesting take. I'd be curious to see data on this.

Mostly agree, with one production caveat: scaffolding gains vanish with the next model, but the evals, permissions and logs you built around it carry over. Invest there.

"wait for the next model" assumes the next model helps your use case. For narrow tasks like booking or collections, it often doesn't. And if you're in a regulated industry, you can't just swap it in anyway.

Interesting perspective. The idea that model improvements can eventually outweigh complex scaffolding is definitely worth thinking about.

never fine-tuned anything, just plain claude code the whole way through yaka. has a scaffolding trick you built ever gotten wiped out by a model update like he's describing?

Lol did you expect hes going to say something else. None of lab voices have ever talked about how they deal with slop or other issues, do you think they have none?

easy to say from the lab that sells the general model tho, my Qwen 3.8 27B handles a lot of my daily stuff without a token bill

LOL

yeah scaffolding wins fade. the general model keeps eating the special cases

Stupid take. Firstly, it depends on the task at hand, the desired quality and the costs. Also: a swiss army knife is never better than a specific tool.

Unit economics are everything in the real world. Boris is too abstracted from this reality

The interesting takeaway is that better general models can make complex scaffolding obsolete surprisingly quickly.

It’s their business model and it makes sense for uncosntrained budgets at Anthropic. On the contrary, a model like Jev opens so many business cases for things that were easily feasible with fable but just too expensive.

Yes, but make your “general” model a specialists in every field you have it learn in its digital data codex. Then you get the hybrid best of both worlds, without having to sacrifice anything in any one area, and you’ll end up with monster Super Intelligence.

This aged like milk.

Makes sense. For example, what does "good at coding" mean ? Need general intelligence and knowledge to actually implement things that aren't simple web development tasks. The best programmers aren't just people that learned languages and frameworks.

I Bet he would say that 🤣🤣🤣🤣

waiting for the next model is a tough sell when a customer needs it working this month

yeah, sure...

speed/price/quality - choose any 2. the bigger models will always be slower than smaller models. and you can squeeze good or even better quality from fine-tuned models ... BUT the real reason to use the latest most general models is about surfing the wave of progress. By the time you've tuned (and spent moeny/time) on tuning some smaller model, newer models emerge that invalidate your work. now you need to tune a bigger faster more capable model again. I personally wont be fine tuning any small models until the avalanche of ever improving models stops... and I don't see it slowing any time soon. so there's that.

(interview 7 months ago)

Wrong. Because efficiency. The most general solution may be "best", but if you care about cost (= energy usage) and speed, you don't want Astra for zip code digit recognition. An LLM can multiply two numbers with billions of ops, or use one assembly instruction. cf. Jev, etc.

That’s true, unless you have proprietary data, or there are specific traps that any LLM falls into regarding the process you’re asking about
![Profilbild von (identity '[:pankaj :λ])](https://image.24vids.com/tw/profile_images/1748379845397184512/BXfj6M6z_x96.jpg)
In the human world, it is completely true, generalists generally outperform experts in every field. MBAs don't create great billion dollar organizations. It is generalists who do.

fine-tuning is now called waiting three months
![Profilbild von (identity '[:pankaj :λ])](https://image.24vids.com/tw/profile_images/1748379845397184512/BXfj6M6z_x96.jpg)
That general advice is too. genaral. You can't use a transformer to generate a image that is best done by a diffusion model. You can't generate audio with a LLM because That is served better by a STS or TTS model.

Corollary of The Bitter Lesson.
