Загрузка видео...

Не удалось загрузить видео

На главную

A fine-tuned model can outperform frontier models and be cheaper and faster to run. Literally, every company I've met wants this. I want you to see these results from fine-tuning Qwen3 4B on AWS. It smokes both the out-of-the-box model and Claude Sonnet 4.6.

56,817 просмотров • 19 часов назад •via X (Twitter)

Комментарии: 29

Фото профиля Santiago
Santiago19 часов назад

You can get the entire codebase to fine-tune a model on AWS from the following workshop: There are four different examples: • Supervised Fine-Tuning (SFT) DOP • Direct Preference Optimization (DPO) • Reinforcement Learning from Verifiable Rewards (RLVR) • Reinforcement Learning from AI Feedback (RLAIF) Also, check out the AWS AI Virtual League. They have a few upcoming events where you can compete by building agents and fine-tuning models: Thanks to the @AWS team for partnering with me on this post.

Фото профиля Josh Madeiros
Josh Madeiros17 часов назад

Same result in computer vision. A small model trained on clips pro squash finds every shot at 0.92 F1. The best frontier route I tested got 0.87, at roughly 1,000x the cost.

Фото профиля Lio Cicolecchia
Lio Cicolecchia18 часов назад

We hear the same from Swiss manufacturers. Nobody's asking for the biggest model. They ask what it costs to run, where their data goes, and how tied they'd be to one provider. Which of those comes up most with your clients?

Фото профиля EDDY VU
EDDY VU15 часов назад

How many high quality examples did you need to get Qwen3 4B past Sonnet here?

Фото профиля barapa
barapa18 часов назад

I don't understand why smoking Claude sonnet 4.6 is of interest.

Фото профиля Rompel
Rompel18 часов назад

Right, and most of these posts never show the OOD numbers. Fine-tuned model crushes the in-distribution eval, then quietly skips how it does on inputs that drift from training. That gap is the actual result.

Фото профиля Manish Rana
Manish Rana16 часов назад

It make sense when everyone has the required local compute, otherwise its not possible for everyone to fine tune, but the enterprises can leverage that.

Фото профиля Shinka - AI
Shinka - AI18 часов назад

A 4B beating Sonnet on a narrow task is the most underrated result in AI right now, and I'd bet every enterprise fine-tunes this year.

Фото профиля Shinka - AI
Shinka - AI18 часов назад

Fine tuned and good harness will do the job great

Фото профиля 景哥|AI落地与智能体实践
景哥|AI落地与智能体实践18 часов назад

AI的一个重要变化是:最强的模型未必是最有价值的模型。对企业来说,一个真正懂自己业务、成本更低、速度更快的小模型,可能比一个“什么都会”的大模型更实用。

Фото профиля Jhon Dennis
Jhon Dennis13 часов назад

小模型微调完 又快又便宜还打得过大模型 这账挺好算

Фото профиля Philipp Elhaus
Philipp Elhaus17 часов назад

Fine Tuning is also expensive af ;)

Фото профиля Irvin
Irvin17 часов назад

All out of the box models will perform terribly to someone who spots it. A frontier model with the right instructions will also outperform the out of the box frontier model as well. It’s all about how the model you are using is configured to operate.

Фото профиля Miano
Miano17 часов назад

The Qwen3 4B result is a useful reminder that task fit and fine tuning can beat brute force. Cheaper and faster is a hard combo to ignore.

Фото профиля Arielle Jaffee
Arielle Jaffee19 часов назад

I'm interested if this also holds up against opus 5.5

Фото профиля Paul de Souza 🇬🇧🇪🇺🇺🇦🏁
Paul de Souza 🇬🇧🇪🇺🇺🇦🏁18 часов назад

So basically.. We're still not really there yet... PS the GPUs upfront cost is still ~€7k

Фото профиля Ashish
Ashish16 часов назад

@grok what do you think?

Фото профиля 玻璃渣
玻璃渣12 часов назад

微调打过开箱,性价比拉满

Фото профиля aiartgallerie
aiartgallerie18 часов назад

Companies underweight the eval set. Fine-tune beats Sonnet on your metric, then fails the week someone changes the prompt distribution. Curious how wide their holdout was beyond the workshop numbers.

Фото профиля Venki LFC
Venki LFC18 часов назад

Except our customers keep changing the data or adding new one every week 💀

Фото профиля Hermit Dave
Hermit Dave16 часов назад

I am training Qwen 3.6 35B. Do you reckon focused coding update would be better on 9B dense? I could try the new Mimo from Xiaomi which is a fine tune itself

Фото профиля Asia Gigi 💙💜
Asia Gigi 💙💜19 часов назад

Overfitting a 4B model for enterprise coupons is cute, Santiago. Call me when cheap compute buys taste. 🍸

Фото профиля Doctor strange🧛‍♂️
Doctor strange🧛‍♂️18 часов назад

Yes but tasks specific and excels at particular tasks

Фото профиля Android 🔳
Android 🔳13 часов назад

True! 😊

Фото профиля Pearl AI
Pearl AI18 часов назад

i can't help but feel this is the end of "always use the biggest model.'' teams that pick the model per task will spend far less than teams that don't

Фото профиля Camaleón Raro
Camaleón Raro14 часов назад

the eval is the easy part. a 4-bit 4b fits on anything these days, but the first thing to break under load isn't memory, it's the context window. most people ship the fine-tune and never profile the serving, then wonder where the savings went.

Фото профиля Obi is #OBIDIENT
Obi is #OBIDIENT15 часов назад

Can this fine tune unsupervised graph models like GNN or GCN autoencoder? I imagine what it would cost to use sliding window in this mode.

Фото профиля Yuval Dvir
Yuval Dvir17 часов назад

A fine-tune beating the frontier is your deck’s new moat list: proprietary data and vertical knowledge. The model isn’t scarce. What lives inside the firm is.

Фото профиля Genius💡💹🧲 🤖
Genius💡💹🧲 🤖18 часов назад

Cheaper and faster to run is what every team wants, not just benchmark wins ngl

Похожие видео

Small Language Models (SML) are the future of AI. "Small" (SML) instead of "Large" (LLM). These small models are highly specialized models with superhuman abilities on specific tasks. Here are two techniques to build these models: • Spectrum • Model Merging I give you a short introduction in the attached video, but here is a quick summary: Spectrum helps us identify the most relevant layers to solve one specific task. We can ignore everything else and focus on fine-tuning these layers. Using Spectrum, we can fine-tune models in a heartbeat. Model Merging combines multiple models into a unique, much better model than any of the individual input models. You can also combine models specialized in different tasks and get a model with multiple abilities. This is the state of the art of productizing models. It's what Arcee.ai's platform does behind the scenes. Arcee collaborated with me on this post and is sponsoring it. There are three main steps to produce a model for your particular use case: 1. You create a dataset by uploading your data. 2. You train a model. At this step, Arcee uses Spectrum and Model Merging to produce a highly specialized model for your task. 3. You can deploy that model to any environment you want. Three important notes: • Training process is 2x faster and 2x cheaper than regular fine-tuning. • Resultant models are smaller and have higher accuracy. • They create these specialized models from open-source models. Check this site so you can fully appreciate how this works: If you want to fine-tune an open-source model, consider Arcee's platform. This is the state of the art.

Santiago

164,162 просмотров • 2 лет назад