Загрузка видео...

Не удалось загрузить видео

На главную

Boris Cherny, Claude Code creator at Anthrpic: "The more general model will always outperform the more specific model. Don't try to use tiny models for stuff. Don't try to fine-tune. Don't try to do any of this stuff. There are some applications, there are some reasons to do this,...

293,396 просмотров • 3 дней назад •via X (Twitter)

Комментарии: 33

Фото профиля Dusan Odalovic
Dusan Odalovic3 дней назад

easy to say when the big model is free for you :) my token bill might have a slightly diferent opinion...

Фото профиля Dan C.
Dan C.2 дней назад

It's purely coincidental that Anthropic's entire business model is general purpose models

Фото профиля Sivan
Sivan3 дней назад

That categorically incorrect. The more studies conducted the otherwise is proven - many small models fine tuned on a task outperform by far the huge generalists. But I understand the commercial and hype motivated claims here. That is otherwise misunderstanding how these NNs work

Фото профиля Gaetan Semet
Gaetan Semet3 дней назад

He is right. Do not try to pay 10x less for the same job, please, my bonus depends on users not optimizing their token use. Will be difficult for you, you will need to learn stuff, it is too complex. Just pay.

Фото профиля Roni Rechter
Roni Rechter3 дней назад

Convenient position for the people selling the general model )

Фото профиля David Hrubý
David Hrubý3 дней назад

Lmao selling shovels. God I hate this retard. It's been proven many times that fine tuned models perform much better than frontier on specific tasks

Фото профиля Alex Prompter
Alex Prompter3 дней назад

Very interesting take. I'd be curious to see data on this.

Фото профиля pranjulya
pranjulya3 дней назад

Mostly agree, with one production caveat: scaffolding gains vanish with the next model, but the evals, permissions and logs you built around it carry over. Invest there.

Фото профиля Neethu Mariam Joy
Neethu Mariam Joy3 дней назад

"wait for the next model" assumes the next model helps your use case. For narrow tasks like booking or collections, it often doesn't. And if you're in a regulated industry, you can't just swap it in anyway.

Фото профиля Flix Muller
Flix Muller3 дней назад

Interesting perspective. The idea that model improvements can eventually outweigh complex scaffolding is definitely worth thinking about.

Фото профиля Malou | MMM
Malou | MMM3 дней назад

never fine-tuned anything, just plain claude code the whole way through yaka. has a scaffolding trick you built ever gotten wiped out by a model update like he's describing?

Фото профиля half_cto
half_cto3 дней назад

Lol did you expect hes going to say something else. None of lab voices have ever talked about how they deal with slop or other issues, do you think they have none?

Фото профиля Julius
Julius3 дней назад

easy to say from the lab that sells the general model tho, my Qwen 3.8 27B handles a lot of my daily stuff without a token bill

Фото профиля saphir bleu
saphir bleu3 дней назад

LOL

Фото профиля Alek Dob
Alek Dob3 дней назад

yeah scaffolding wins fade. the general model keeps eating the special cases

Фото профиля Andries van der Leij
Andries van der Leij3 дней назад

Stupid take. Firstly, it depends on the task at hand, the desired quality and the costs. Also: a swiss army knife is never better than a specific tool.

Фото профиля Damien Hughes
Damien Hughes3 дней назад

Unit economics are everything in the real world. Boris is too abstracted from this reality

Фото профиля Vipul Kumar Kewat
Vipul Kumar Kewat3 дней назад

The interesting takeaway is that better general models can make complex scaffolding obsolete surprisingly quickly.

Фото профиля jw
jw3 дней назад

It’s their business model and it makes sense for uncosntrained budgets at Anthropic. On the contrary, a model like Jev opens so many business cases for things that were easily feasible with fable but just too expensive.

Фото профиля Mr. House
Mr. House3 дней назад

Yes, but make your “general” model a specialists in every field you have it learn in its digital data codex. Then you get the hybrid best of both worlds, without having to sacrifice anything in any one area, and you’ll end up with monster Super Intelligence.

Фото профиля Mathew Chan
Mathew Chan2 дней назад

This aged like milk.

Фото профиля Davide Pasca
Davide Pasca2 дней назад

Makes sense. For example, what does "good at coding" mean ? Need general intelligence and knowledge to actually implement things that aren't simple web development tasks. The best programmers aren't just people that learned languages and frameworks.

Фото профиля Mr. LIvermore
Mr. LIvermore2 дней назад

I Bet he would say that 🤣🤣🤣🤣

Фото профиля Somi
Somi3 дней назад

waiting for the next model is a tough sell when a customer needs it working this month

Фото профиля Samion Buchas
Samion Buchas2 дней назад

yeah, sure...

Фото профиля aviad rozenhek
aviad rozenhek2 дней назад

speed/price/quality - choose any 2. the bigger models will always be slower than smaller models. and you can squeeze good or even better quality from fine-tuned models ... BUT the real reason to use the latest most general models is about surfing the wave of progress. By the time you've tuned (and spent moeny/time) on tuning some smaller model, newer models emerge that invalidate your work. now you need to tune a bigger faster more capable model again. I personally wont be fine tuning any small models until the avalanche of ever improving models stops... and I don't see it slowing any time soon. so there's that.

Фото профиля GrepMed
GrepMed3 дней назад

(interview 7 months ago)

Фото профиля Shai Machnes
Shai Machnes3 дней назад

Wrong. Because efficiency. The most general solution may be "best", but if you care about cost (= energy usage) and speed, you don't want Astra for zip code digit recognition. An LLM can multiply two numbers with billions of ops, or use one assembly instruction. cf. Jev, etc.

Фото профиля Bruno C
Bruno C3 дней назад

That’s true, unless you have proprietary data, or there are specific traps that any LLM falls into regarding the process you’re asking about

Фото профиля (identity '[:pankaj :λ])
(identity '[:pankaj :λ])2 дней назад

In the human world, it is completely true, generalists generally outperform experts in every field. MBAs don't create great billion dollar organizations. It is generalists who do.

Фото профиля Prithvi
Prithvi3 дней назад

fine-tuning is now called waiting three months

Фото профиля (identity '[:pankaj :λ])
(identity '[:pankaj :λ])2 дней назад

That general advice is too. genaral. You can't use a transformer to generate a image that is best done by a diffusion model. You can't generate audio with a LLM because That is served better by a STS or TTS model.

Фото профиля Schopenhauer on Prozac
Schopenhauer on Prozac3 дней назад

Corollary of The Bitter Lesson.

Похожие видео

Small Language Models (SML) are the future of AI. "Small" (SML) instead of "Large" (LLM). These small models are highly specialized models with superhuman abilities on specific tasks. Here are two techniques to build these models: • Spectrum • Model Merging I give you a short introduction in the attached video, but here is a quick summary: Spectrum helps us identify the most relevant layers to solve one specific task. We can ignore everything else and focus on fine-tuning these layers. Using Spectrum, we can fine-tune models in a heartbeat. Model Merging combines multiple models into a unique, much better model than any of the individual input models. You can also combine models specialized in different tasks and get a model with multiple abilities. This is the state of the art of productizing models. It's what Arcee.ai's platform does behind the scenes. Arcee collaborated with me on this post and is sponsoring it. There are three main steps to produce a model for your particular use case: 1. You create a dataset by uploading your data. 2. You train a model. At this step, Arcee uses Spectrum and Model Merging to produce a highly specialized model for your task. 3. You can deploy that model to any environment you want. Three important notes: • Training process is 2x faster and 2x cheaper than regular fine-tuning. • Resultant models are smaller and have higher accuracy. • They create these specialized models from open-source models. Check this site so you can fully appreciate how this works: If you want to fine-tune an open-source model, consider Arcee's platform. This is the state of the art.

Santiago

164,162 просмотров • 2 лет назад

VP VANCE PREDICTED: PEOPLE ARE GOING TO GET ANGRY, AND RIGHTFULLY SO "This stuff we can and we should prosecute, and I'm just telling you, this is going to be a real problem, and the people are going to get really pissed at Senate Republicans if we don't have the U.S. attorneys on the ground to actually achieve justice. People are going to get angry, and rightfully so." If you want justice, you've got to empower the President of the United States to actually appoint the officers of justice all over the country. The Democrats are stalling that, and we're going to wake up in a couple of years, if we don't have more U.S. attorneys approved, if we don't have more judges approved, we're going to wake up in a couple of years and realize that we've done a lot of great work at the Trump administration, but justice is not being meted out as it should be because we don't have the people on the ground. That is a big problem, and I know that's somewhat unrelated to Arctic Frost, but it actually is related to Arctic Frost, because you cannot get the justice for the people who are targeted by the Biden administration unless we've got good people, especially in these U.S. attorneys offices, and that's something we've got to pay attention to over the next year. Spying on President Trump, prosecuting him, investigating senators, congressmen, and congressmen who are just aligned with the President of the United States some of this stuff is going to get covered by statute of limitations, but some of this stuff we can and we should prosecute, and I'm just telling you, this is going to be a real problem, and the people who watch your show are going to get really pissed at Senate Republicans, excuse my language, if we don't have the U.S. attorneys on the ground to actually achieve justice, people are going to get angry, and rightfully so. If you want justice, you've got to empower the President of the United States to actually appoint the officers of justice all over the country. The Democrats are stalling that, and we're going to wake up in a couple of years, if we don't have more U.S. attorneys approved, if we don't have more judges approved, we're going to wake up in a couple of years and realize that we've done a lot of great work at the Trump administration, but justice is not being meted out as it should be because we don't have the people on the ground. That is a big problem, and I know that's somewhat unrelated to Arctic Frost, but it actually is related to Arctic Frost, because you cannot get the justice for the people who are targeted by the Biden administration unless we've got good people, especially in these U.S. attorneys offices, and that's something we've got to pay attention to over the next year. Spying on President Trump, prosecuting him, investigating senators, congressmen, and congressmen who are just aligned with the President of the United States some of this stuff is going to get covered by statute of limitations, but some of this stuff we can and we should prosecute, and I'm just telling you, this is going to be a real problem, and the people who watch your show are going to get really pissed at Senate Republicans, excuse my language, if we don't have the U.S. attorneys on the ground to actually achieve justice, people are going to get angry, and rightfully so.

Svetlana Lokhova

254,370 просмотров • 8 месяцев назад