Loading video...

Video Failed to Load

Go Home

Boris Cherny, Claude Code creator at Anthrpic: "The more general model will always outperform the more specific model. Don't try to use tiny models for stuff. Don't try to fine-tune. Don't try to do any of this stuff. There are some applications, there are some reasons to do this,...

293,396 views • 3 days ago •via X (Twitter)

33 Comments

Dusan Odalovic's profile picture
Dusan Odalovic3 days ago

easy to say when the big model is free for you :) my token bill might have a slightly diferent opinion...

Dan C.'s profile picture
Dan C.3 days ago

It's purely coincidental that Anthropic's entire business model is general purpose models

Sivan's profile picture
Sivan3 days ago

That categorically incorrect. The more studies conducted the otherwise is proven - many small models fine tuned on a task outperform by far the huge generalists. But I understand the commercial and hype motivated claims here. That is otherwise misunderstanding how these NNs work

Gaetan Semet's profile picture
Gaetan Semet3 days ago

He is right. Do not try to pay 10x less for the same job, please, my bonus depends on users not optimizing their token use. Will be difficult for you, you will need to learn stuff, it is too complex. Just pay.

Roni Rechter's profile picture
Roni Rechter3 days ago

Convenient position for the people selling the general model )

David Hrubý's profile picture
David Hrubý3 days ago

Lmao selling shovels. God I hate this retard. It's been proven many times that fine tuned models perform much better than frontier on specific tasks

Alex Prompter's profile picture
Alex Prompter3 days ago

Very interesting take. I'd be curious to see data on this.

pranjulya's profile picture
pranjulya3 days ago

Mostly agree, with one production caveat: scaffolding gains vanish with the next model, but the evals, permissions and logs you built around it carry over. Invest there.

Neethu Mariam Joy's profile picture
Neethu Mariam Joy3 days ago

"wait for the next model" assumes the next model helps your use case. For narrow tasks like booking or collections, it often doesn't. And if you're in a regulated industry, you can't just swap it in anyway.

Flix Muller's profile picture
Flix Muller3 days ago

Interesting perspective. The idea that model improvements can eventually outweigh complex scaffolding is definitely worth thinking about.

Malou | MMM's profile picture
Malou | MMM3 days ago

never fine-tuned anything, just plain claude code the whole way through yaka. has a scaffolding trick you built ever gotten wiped out by a model update like he's describing?

half_cto's profile picture
half_cto3 days ago

Lol did you expect hes going to say something else. None of lab voices have ever talked about how they deal with slop or other issues, do you think they have none?

Julius's profile picture
Julius3 days ago

easy to say from the lab that sells the general model tho, my Qwen 3.8 27B handles a lot of my daily stuff without a token bill

saphir bleu's profile picture
saphir bleu3 days ago

LOL

Alek Dob's profile picture
Alek Dob3 days ago

yeah scaffolding wins fade. the general model keeps eating the special cases

Andries van der Leij's profile picture
Andries van der Leij3 days ago

Stupid take. Firstly, it depends on the task at hand, the desired quality and the costs. Also: a swiss army knife is never better than a specific tool.

Damien Hughes's profile picture
Damien Hughes3 days ago

Unit economics are everything in the real world. Boris is too abstracted from this reality

Vipul Kumar Kewat's profile picture
Vipul Kumar Kewat3 days ago

The interesting takeaway is that better general models can make complex scaffolding obsolete surprisingly quickly.

jw's profile picture
jw3 days ago

It’s their business model and it makes sense for uncosntrained budgets at Anthropic. On the contrary, a model like Jev opens so many business cases for things that were easily feasible with fable but just too expensive.

Mr. House's profile picture
Mr. House3 days ago

Yes, but make your “general” model a specialists in every field you have it learn in its digital data codex. Then you get the hybrid best of both worlds, without having to sacrifice anything in any one area, and you’ll end up with monster Super Intelligence.

Mathew Chan's profile picture
Mathew Chan2 days ago

This aged like milk.

Davide Pasca's profile picture
Davide Pasca3 days ago

Makes sense. For example, what does "good at coding" mean ? Need general intelligence and knowledge to actually implement things that aren't simple web development tasks. The best programmers aren't just people that learned languages and frameworks.

Mr. LIvermore's profile picture
Mr. LIvermore3 days ago

I Bet he would say that 🤣🤣🤣🤣

Somi's profile picture
Somi3 days ago

waiting for the next model is a tough sell when a customer needs it working this month

Samion Buchas's profile picture
Samion Buchas3 days ago

yeah, sure...

aviad rozenhek's profile picture
aviad rozenhek2 days ago

speed/price/quality - choose any 2. the bigger models will always be slower than smaller models. and you can squeeze good or even better quality from fine-tuned models ... BUT the real reason to use the latest most general models is about surfing the wave of progress. By the time you've tuned (and spent moeny/time) on tuning some smaller model, newer models emerge that invalidate your work. now you need to tune a bigger faster more capable model again. I personally wont be fine tuning any small models until the avalanche of ever improving models stops... and I don't see it slowing any time soon. so there's that.

GrepMed's profile picture
GrepMed3 days ago

(interview 7 months ago)

Shai Machnes's profile picture
Shai Machnes3 days ago

Wrong. Because efficiency. The most general solution may be "best", but if you care about cost (= energy usage) and speed, you don't want Astra for zip code digit recognition. An LLM can multiply two numbers with billions of ops, or use one assembly instruction. cf. Jev, etc.

Bruno C's profile picture
Bruno C3 days ago

That’s true, unless you have proprietary data, or there are specific traps that any LLM falls into regarding the process you’re asking about

(identity '[:pankaj :λ])'s profile picture
(identity '[:pankaj :λ])3 days ago

In the human world, it is completely true, generalists generally outperform experts in every field. MBAs don't create great billion dollar organizations. It is generalists who do.

Prithvi's profile picture
Prithvi3 days ago

fine-tuning is now called waiting three months

(identity '[:pankaj :λ])'s profile picture
(identity '[:pankaj :λ])3 days ago

That general advice is too. genaral. You can't use a transformer to generate a image that is best done by a diffusion model. You can't generate audio with a LLM because That is served better by a STS or TTS model.

Schopenhauer on Prozac's profile picture
Schopenhauer on Prozac3 days ago

Corollary of The Bitter Lesson.

Related Videos

Small Language Models (SML) are the future of AI. "Small" (SML) instead of "Large" (LLM). These small models are highly specialized models with superhuman abilities on specific tasks. Here are two techniques to build these models: • Spectrum • Model Merging I give you a short introduction in the attached video, but here is a quick summary: Spectrum helps us identify the most relevant layers to solve one specific task. We can ignore everything else and focus on fine-tuning these layers. Using Spectrum, we can fine-tune models in a heartbeat. Model Merging combines multiple models into a unique, much better model than any of the individual input models. You can also combine models specialized in different tasks and get a model with multiple abilities. This is the state of the art of productizing models. It's what Arcee.ai's platform does behind the scenes. Arcee collaborated with me on this post and is sponsoring it. There are three main steps to produce a model for your particular use case: 1. You create a dataset by uploading your data. 2. You train a model. At this step, Arcee uses Spectrum and Model Merging to produce a highly specialized model for your task. 3. You can deploy that model to any environment you want. Three important notes: • Training process is 2x faster and 2x cheaper than regular fine-tuning. • Resultant models are smaller and have higher accuracy. • They create these specialized models from open-source models. Check this site so you can fully appreciate how this works: If you want to fine-tune an open-source model, consider Arcee's platform. This is the state of the art.

Santiago

164,162 views • 2 years ago

VP VANCE PREDICTED: PEOPLE ARE GOING TO GET ANGRY, AND RIGHTFULLY SO "This stuff we can and we should prosecute, and I'm just telling you, this is going to be a real problem, and the people are going to get really pissed at Senate Republicans if we don't have the U.S. attorneys on the ground to actually achieve justice. People are going to get angry, and rightfully so." If you want justice, you've got to empower the President of the United States to actually appoint the officers of justice all over the country. The Democrats are stalling that, and we're going to wake up in a couple of years, if we don't have more U.S. attorneys approved, if we don't have more judges approved, we're going to wake up in a couple of years and realize that we've done a lot of great work at the Trump administration, but justice is not being meted out as it should be because we don't have the people on the ground. That is a big problem, and I know that's somewhat unrelated to Arctic Frost, but it actually is related to Arctic Frost, because you cannot get the justice for the people who are targeted by the Biden administration unless we've got good people, especially in these U.S. attorneys offices, and that's something we've got to pay attention to over the next year. Spying on President Trump, prosecuting him, investigating senators, congressmen, and congressmen who are just aligned with the President of the United States some of this stuff is going to get covered by statute of limitations, but some of this stuff we can and we should prosecute, and I'm just telling you, this is going to be a real problem, and the people who watch your show are going to get really pissed at Senate Republicans, excuse my language, if we don't have the U.S. attorneys on the ground to actually achieve justice, people are going to get angry, and rightfully so. If you want justice, you've got to empower the President of the United States to actually appoint the officers of justice all over the country. The Democrats are stalling that, and we're going to wake up in a couple of years, if we don't have more U.S. attorneys approved, if we don't have more judges approved, we're going to wake up in a couple of years and realize that we've done a lot of great work at the Trump administration, but justice is not being meted out as it should be because we don't have the people on the ground. That is a big problem, and I know that's somewhat unrelated to Arctic Frost, but it actually is related to Arctic Frost, because you cannot get the justice for the people who are targeted by the Biden administration unless we've got good people, especially in these U.S. attorneys offices, and that's something we've got to pay attention to over the next year. Spying on President Trump, prosecuting him, investigating senators, congressmen, and congressmen who are just aligned with the President of the United States some of this stuff is going to get covered by statute of limitations, but some of this stuff we can and we should prosecute, and I'm just telling you, this is going to be a real problem, and the people who watch your show are going to get really pissed at Senate Republicans, excuse my language, if we don't have the U.S. attorneys on the ground to actually achieve justice, people are going to get angry, and rightfully so.

Svetlana Lokhova

254,370 views • 8 months ago