Loading video...

Video Failed to Load

Go Home

Boris Cherny, head of Claude Code at Anthropic, on optimizing token cost and model use. “I use Fable for everything" There is probably a 50% opportunity to reduce the investment (on token cost). However, there may be a 1,000%, 10,000%, or even 100,000% opportunity to increase the return. Therefore,...

23,168 views • 6 days ago •via X (Twitter)

22 Comments

Rohan Paul's profile picture
Rohan Paul6 days ago

full video

lderait Menlinden's profile picture
lderait Menlinden6 days ago

Correct. As soon as your brain has convinced you that lesser models are somehow useful, you have been proven to ask the wrong questions.

Gaurav Atavale's profile picture
Gaurav Atavale6 days ago

Interesting discussion!!!

linkedIn photoai ⚡'s profile picture
linkedIn photoai ⚡5 days ago

I was a big big fable fan and Claude in general especially in terminal but I recently switched back to gpt because I anthropic is so exploitative of the customers. Every few weeks they have some announcement that’s out to fuck you. Just look at the fable drama

Adam Knees's profile picture
Adam Knees5 days ago

Different models excel at different tasks though, which is why model choice still matters. At Snowflake, we released Cortex Gateway to help companies route workloads across different models based on what each one does best. The future probably isn't one model for everything. It's using the right model for the right task.

Leo Lu's profile picture
Leo Lu6 days ago

Where do you draw the line when the pricier model barely changes the outcome?

Asa Hidmark's profile picture
Asa Hidmark6 days ago

Luckily he said that for all the companies which forgot to focus on return

Arbaz's profile picture
Arbaz6 days ago

fable for planning still wins for me. exec goes astra-low in codex so one thirsty agent doesn't cook the week.

Chris Hawkins's profile picture
Chris Hawkins6 days ago

If I could use fable for my full $200/mo plan I would use more too!

Shinka - AI's profile picture
Shinka - AI6 days ago

Nobody cares about the 50% you save on tokens, they notice the 100x return you got from actually using the model.

Shreyans Bhansali's profile picture
Shreyans Bhansali6 days ago

funny how model cost feels huge until you price the engineer waiting on it

sofi z's profile picture
sofi z5 days ago

This is so true

We build Kibble's profile picture
We build Kibble6 days ago

Totally agree on the direction. Only half of that ratio is measurable though: cost per model per repo is a query, the return is still a guess.

Axi's profile picture
Axi6 days ago

So no one dogfood on opus and sonnet right

nivelepsilon's profile picture
nivelepsilon5 days ago

The invoice is obvious. That 100,000% return only exists if you can show work that would not have happened otherwise.

Omar وديع's profile picture
Omar وديع6 days ago

I run the same math in production ML for finance. Token spend on fraud detection or KYC review is rounding error next to a single bad judgment call. Routing earns its keep on high-volume, low-stakes tasks.

Whatever's profile picture
Whatever6 days ago

Easy for him to say, he gets unlimited Mythos. For most of us that would cost more than we’re willing or able to pay.

Vishesh Baghel's profile picture
Vishesh Baghel6 days ago

easy take when your token bill says "anthropic internal"

James Ellis-Jones's profile picture
James Ellis-Jones6 days ago

Great salesman our Boris

Doni's profile picture
Doni6 days ago

the hard part is measuring return, not token cost. How do you track it?

Arnal Dayaratna's profile picture
Arnal Dayaratna5 days ago

The advice to use Fable 5 for everything is just bad as in, bad advice. We all have plenty of use cases that can be handled perfectly well by less capable models. The generous interpretation of Cherney's advice would be to use Fable 5 for any reasoning-intensive task. That said, I think that in the case of writing, for example, I would rather use a less capable model simply because it is more reliable and consistent than what we see from Fable 5, which is frankly underwhelming.

Gill's profile picture
Gill5 days ago

Developer time costs way more than token usage anyway

Related Videos

Claude Code cracked something open for us Every 🧱. Now I ship to codebases I barely know, every feature we ship makes the next one easier, and non-technical members of the team use the terminal. I’m genuinely grateful. So I brought its creators, Cat Wu (cat) and Boris Cherny (Boris Cherny) from Anthropic, on AI & I to say thank you—and to talk about everything they’ve learned from building Claude Code. We get into: • The workflows Anthropic’s smartest engineers use to push Claude Code to its limits. Why they pit subagents against each other to get cleaner results, how they turn past code into leverage, and the slash commands and MCPs they rely on most. • The product lessons behind one of the most loved AI agents in the world. How the team balances simplicity and power—building a tool that anyone can use, but that experts can bend to their will—and their philosophy of “unshipping,” or cutting back whenever there’s a simpler, more intuitive path to user intent. • A peek into the future of coding with AI. The new form factors they’re experimenting with to make Claude Code more autonomous, more reliable, and more accessible to non-technical users This is a must-watch for anyone—both technical and non-technical—who wants to learn how to use Claude Code like the people who built it. Watch below! Timestamps: Introduction: 00:01:26 Claude Code’s origin story: 00:02:25 How Anthropic dogfoods Claude Code: 00:07:03 Boris and Cat’s favorite slash commands: 00:14:06 How Boris uses Claude Code to plan feature development: 00:15:49 Everything Anthropic has learned about using sub-agents well: 00:21:53 Use Claude Code to turn past code into leverage: 00:26:16 The product decisions for building an agent that’s simple and powerful: 00:33:14 Making Claude Code accessible to the non-technical user: 00:36:38 The next form factor for coding with AI: 00:45:12

Dan Shipper 📧

57,619 views • 10 months ago