Video wird geladen...

Video konnte nicht geladen werden

Zur Startseite

Boris Cherny, Claude Code creator at Anthrpic: "The more general model will always outperform the more specific model. Don't try to use tiny models for stuff. Don't try to fine-tune. Don't try to do any of this stuff. There are some applications, there are some reasons to do this,...

293,396 Aufrufe • vor 3 Tagen •via X (Twitter)

33 Kommentare

Profilbild von Dusan Odalovic
Dusan Odalovicvor 3 Tagen

easy to say when the big model is free for you :) my token bill might have a slightly diferent opinion...

Profilbild von Dan C.
Dan C.vor 2 Tagen

It's purely coincidental that Anthropic's entire business model is general purpose models

Profilbild von Sivan
Sivanvor 3 Tagen

That categorically incorrect. The more studies conducted the otherwise is proven - many small models fine tuned on a task outperform by far the huge generalists. But I understand the commercial and hype motivated claims here. That is otherwise misunderstanding how these NNs work

Profilbild von Gaetan Semet
Gaetan Semetvor 2 Tagen

He is right. Do not try to pay 10x less for the same job, please, my bonus depends on users not optimizing their token use. Will be difficult for you, you will need to learn stuff, it is too complex. Just pay.

Profilbild von Roni Rechter
Roni Rechtervor 2 Tagen

Convenient position for the people selling the general model )

Profilbild von David Hrubý
David Hrubývor 3 Tagen

Lmao selling shovels. God I hate this retard. It's been proven many times that fine tuned models perform much better than frontier on specific tasks

Profilbild von Alex Prompter
Alex Promptervor 3 Tagen

Very interesting take. I'd be curious to see data on this.

Profilbild von pranjulya
pranjulyavor 3 Tagen

Mostly agree, with one production caveat: scaffolding gains vanish with the next model, but the evals, permissions and logs you built around it carry over. Invest there.

Profilbild von Neethu Mariam Joy
Neethu Mariam Joyvor 2 Tagen

"wait for the next model" assumes the next model helps your use case. For narrow tasks like booking or collections, it often doesn't. And if you're in a regulated industry, you can't just swap it in anyway.

Profilbild von Flix Muller
Flix Mullervor 3 Tagen

Interesting perspective. The idea that model improvements can eventually outweigh complex scaffolding is definitely worth thinking about.

Profilbild von Malou | MMM
Malou | MMMvor 3 Tagen

never fine-tuned anything, just plain claude code the whole way through yaka. has a scaffolding trick you built ever gotten wiped out by a model update like he's describing?

Profilbild von half_cto
half_ctovor 3 Tagen

Lol did you expect hes going to say something else. None of lab voices have ever talked about how they deal with slop or other issues, do you think they have none?

Profilbild von Julius
Juliusvor 3 Tagen

easy to say from the lab that sells the general model tho, my Qwen 3.8 27B handles a lot of my daily stuff without a token bill

Profilbild von saphir bleu
saphir bleuvor 3 Tagen

LOL

Profilbild von Alek Dob
Alek Dobvor 3 Tagen

yeah scaffolding wins fade. the general model keeps eating the special cases

Profilbild von Andries van der Leij
Andries van der Leijvor 2 Tagen

Stupid take. Firstly, it depends on the task at hand, the desired quality and the costs. Also: a swiss army knife is never better than a specific tool.

Profilbild von Damien Hughes
Damien Hughesvor 3 Tagen

Unit economics are everything in the real world. Boris is too abstracted from this reality

Profilbild von Vipul Kumar Kewat
Vipul Kumar Kewatvor 3 Tagen

The interesting takeaway is that better general models can make complex scaffolding obsolete surprisingly quickly.

Profilbild von jw
jwvor 2 Tagen

It’s their business model and it makes sense for uncosntrained budgets at Anthropic. On the contrary, a model like Jev opens so many business cases for things that were easily feasible with fable but just too expensive.

Profilbild von Mr. House
Mr. Housevor 3 Tagen

Yes, but make your “general” model a specialists in every field you have it learn in its digital data codex. Then you get the hybrid best of both worlds, without having to sacrifice anything in any one area, and you’ll end up with monster Super Intelligence.

Profilbild von Mathew Chan
Mathew Chanvor 2 Tagen

This aged like milk.

Profilbild von Davide Pasca
Davide Pascavor 2 Tagen

Makes sense. For example, what does "good at coding" mean ? Need general intelligence and knowledge to actually implement things that aren't simple web development tasks. The best programmers aren't just people that learned languages and frameworks.

Profilbild von Mr. LIvermore
Mr. LIvermorevor 2 Tagen

I Bet he would say that 🤣🤣🤣🤣

Profilbild von Somi
Somivor 2 Tagen

waiting for the next model is a tough sell when a customer needs it working this month

Profilbild von Samion Buchas
Samion Buchasvor 2 Tagen

yeah, sure...

Profilbild von aviad rozenhek
aviad rozenhekvor 2 Tagen

speed/price/quality - choose any 2. the bigger models will always be slower than smaller models. and you can squeeze good or even better quality from fine-tuned models ... BUT the real reason to use the latest most general models is about surfing the wave of progress. By the time you've tuned (and spent moeny/time) on tuning some smaller model, newer models emerge that invalidate your work. now you need to tune a bigger faster more capable model again. I personally wont be fine tuning any small models until the avalanche of ever improving models stops... and I don't see it slowing any time soon. so there's that.

Profilbild von GrepMed
GrepMedvor 2 Tagen

(interview 7 months ago)

Profilbild von Shai Machnes
Shai Machnesvor 2 Tagen

Wrong. Because efficiency. The most general solution may be "best", but if you care about cost (= energy usage) and speed, you don't want Astra for zip code digit recognition. An LLM can multiply two numbers with billions of ops, or use one assembly instruction. cf. Jev, etc.

Profilbild von Bruno C
Bruno Cvor 3 Tagen

That’s true, unless you have proprietary data, or there are specific traps that any LLM falls into regarding the process you’re asking about

Profilbild von (identity '[:pankaj :λ])
(identity '[:pankaj :λ])vor 2 Tagen

In the human world, it is completely true, generalists generally outperform experts in every field. MBAs don't create great billion dollar organizations. It is generalists who do.

Profilbild von Prithvi
Prithvivor 2 Tagen

fine-tuning is now called waiting three months

Profilbild von (identity '[:pankaj :λ])
(identity '[:pankaj :λ])vor 2 Tagen

That general advice is too. genaral. You can't use a transformer to generate a image that is best done by a diffusion model. You can't generate audio with a LLM because That is served better by a STS or TTS model.

Profilbild von Schopenhauer on Prozac
Schopenhauer on Prozacvor 2 Tagen

Corollary of The Bitter Lesson.

Ähnliche Videos

Small Language Models (SML) are the future of AI. "Small" (SML) instead of "Large" (LLM). These small models are highly specialized models with superhuman abilities on specific tasks. Here are two techniques to build these models: • Spectrum • Model Merging I give you a short introduction in the attached video, but here is a quick summary: Spectrum helps us identify the most relevant layers to solve one specific task. We can ignore everything else and focus on fine-tuning these layers. Using Spectrum, we can fine-tune models in a heartbeat. Model Merging combines multiple models into a unique, much better model than any of the individual input models. You can also combine models specialized in different tasks and get a model with multiple abilities. This is the state of the art of productizing models. It's what Arcee.ai's platform does behind the scenes. Arcee collaborated with me on this post and is sponsoring it. There are three main steps to produce a model for your particular use case: 1. You create a dataset by uploading your data. 2. You train a model. At this step, Arcee uses Spectrum and Model Merging to produce a highly specialized model for your task. 3. You can deploy that model to any environment you want. Three important notes: • Training process is 2x faster and 2x cheaper than regular fine-tuning. • Resultant models are smaller and have higher accuracy. • They create these specialized models from open-source models. Check this site so you can fully appreciate how this works: If you want to fine-tune an open-source model, consider Arcee's platform. This is the state of the art.

Santiago

164,162 Aufrufe • vor 2 Jahren

VP VANCE PREDICTED: PEOPLE ARE GOING TO GET ANGRY, AND RIGHTFULLY SO "This stuff we can and we should prosecute, and I'm just telling you, this is going to be a real problem, and the people are going to get really pissed at Senate Republicans if we don't have the U.S. attorneys on the ground to actually achieve justice. People are going to get angry, and rightfully so." If you want justice, you've got to empower the President of the United States to actually appoint the officers of justice all over the country. The Democrats are stalling that, and we're going to wake up in a couple of years, if we don't have more U.S. attorneys approved, if we don't have more judges approved, we're going to wake up in a couple of years and realize that we've done a lot of great work at the Trump administration, but justice is not being meted out as it should be because we don't have the people on the ground. That is a big problem, and I know that's somewhat unrelated to Arctic Frost, but it actually is related to Arctic Frost, because you cannot get the justice for the people who are targeted by the Biden administration unless we've got good people, especially in these U.S. attorneys offices, and that's something we've got to pay attention to over the next year. Spying on President Trump, prosecuting him, investigating senators, congressmen, and congressmen who are just aligned with the President of the United States some of this stuff is going to get covered by statute of limitations, but some of this stuff we can and we should prosecute, and I'm just telling you, this is going to be a real problem, and the people who watch your show are going to get really pissed at Senate Republicans, excuse my language, if we don't have the U.S. attorneys on the ground to actually achieve justice, people are going to get angry, and rightfully so. If you want justice, you've got to empower the President of the United States to actually appoint the officers of justice all over the country. The Democrats are stalling that, and we're going to wake up in a couple of years, if we don't have more U.S. attorneys approved, if we don't have more judges approved, we're going to wake up in a couple of years and realize that we've done a lot of great work at the Trump administration, but justice is not being meted out as it should be because we don't have the people on the ground. That is a big problem, and I know that's somewhat unrelated to Arctic Frost, but it actually is related to Arctic Frost, because you cannot get the justice for the people who are targeted by the Biden administration unless we've got good people, especially in these U.S. attorneys offices, and that's something we've got to pay attention to over the next year. Spying on President Trump, prosecuting him, investigating senators, congressmen, and congressmen who are just aligned with the President of the United States some of this stuff is going to get covered by statute of limitations, but some of this stuff we can and we should prosecute, and I'm just telling you, this is going to be a real problem, and the people who watch your show are going to get really pissed at Senate Republicans, excuse my language, if we don't have the U.S. attorneys on the ground to actually achieve justice, people are going to get angry, and rightfully so.

Svetlana Lokhova

254,370 Aufrufe • vor 8 Monaten