Загрузка видео...
Не удалось загрузить видео
A fine-tuned model can outperform frontier models and be cheaper and faster to run. Literally, every company I've met wants this. I want you to see these results from fine-tuning Qwen3 4B on AWS. It smokes both the out-of-the-box model and Claude Sonnet 4.6.
56,817 просмотров • 19 часов назад •via X (Twitter)
Комментарии: 29

You can get the entire codebase to fine-tune a model on AWS from the following workshop: There are four different examples: • Supervised Fine-Tuning (SFT) DOP • Direct Preference Optimization (DPO) • Reinforcement Learning from Verifiable Rewards (RLVR) • Reinforcement Learning from AI Feedback (RLAIF) Also, check out the AWS AI Virtual League. They have a few upcoming events where you can compete by building agents and fine-tuning models: Thanks to the @AWS team for partnering with me on this post.

Same result in computer vision. A small model trained on clips pro squash finds every shot at 0.92 F1. The best frontier route I tested got 0.87, at roughly 1,000x the cost.

We hear the same from Swiss manufacturers. Nobody's asking for the biggest model. They ask what it costs to run, where their data goes, and how tied they'd be to one provider. Which of those comes up most with your clients?

How many high quality examples did you need to get Qwen3 4B past Sonnet here?

I don't understand why smoking Claude sonnet 4.6 is of interest.

Right, and most of these posts never show the OOD numbers. Fine-tuned model crushes the in-distribution eval, then quietly skips how it does on inputs that drift from training. That gap is the actual result.

It make sense when everyone has the required local compute, otherwise its not possible for everyone to fine tune, but the enterprises can leverage that.

A 4B beating Sonnet on a narrow task is the most underrated result in AI right now, and I'd bet every enterprise fine-tunes this year.

Fine tuned and good harness will do the job great

AI的一个重要变化是:最强的模型未必是最有价值的模型。对企业来说,一个真正懂自己业务、成本更低、速度更快的小模型,可能比一个“什么都会”的大模型更实用。

小模型微调完 又快又便宜还打得过大模型 这账挺好算

Fine Tuning is also expensive af ;)

All out of the box models will perform terribly to someone who spots it. A frontier model with the right instructions will also outperform the out of the box frontier model as well. It’s all about how the model you are using is configured to operate.

The Qwen3 4B result is a useful reminder that task fit and fine tuning can beat brute force. Cheaper and faster is a hard combo to ignore.

I'm interested if this also holds up against opus 5.5

So basically.. We're still not really there yet... PS the GPUs upfront cost is still ~€7k

@grok what do you think?

微调打过开箱,性价比拉满

Companies underweight the eval set. Fine-tune beats Sonnet on your metric, then fails the week someone changes the prompt distribution. Curious how wide their holdout was beyond the workshop numbers.

Except our customers keep changing the data or adding new one every week 💀

I am training Qwen 3.6 35B. Do you reckon focused coding update would be better on 9B dense? I could try the new Mimo from Xiaomi which is a fine tune itself

Overfitting a 4B model for enterprise coupons is cute, Santiago. Call me when cheap compute buys taste. 🍸

Yes but tasks specific and excels at particular tasks

True! 😊

i can't help but feel this is the end of "always use the biggest model.'' teams that pick the model per task will spend far less than teams that don't

the eval is the easy part. a 4-bit 4b fits on anything these days, but the first thing to break under load isn't memory, it's the context window. most people ship the fine-tune and never profile the serving, then wonder where the savings went.

Can this fine tune unsupervised graph models like GNN or GCN autoencoder? I imagine what it would cost to use sliding window in this mode.

A fine-tune beating the frontier is your deck’s new moat list: proprietary data and vertical knowledge. The model isn’t scarce. What lives inside the firm is.

Cheaper and faster to run is what every team wants, not just benchmark wins ngl
