Loading video...
Video Failed to Load
Boris Cherny, head of Claude Code at Anthropic, on optimizing token cost and model use. “I use Fable for everything" There is probably a 50% opportunity to reduce the investment (on token cost). However, there may be a 1,000%, 10,000%, or even 100,000% opportunity to increase the return. Therefore,... show more
23,168 views • 6 days ago •via X (Twitter)
22 Comments

full video

Correct. As soon as your brain has convinced you that lesser models are somehow useful, you have been proven to ask the wrong questions.

Interesting discussion!!!

I was a big big fable fan and Claude in general especially in terminal but I recently switched back to gpt because I anthropic is so exploitative of the customers. Every few weeks they have some announcement that’s out to fuck you. Just look at the fable drama

Different models excel at different tasks though, which is why model choice still matters. At Snowflake, we released Cortex Gateway to help companies route workloads across different models based on what each one does best. The future probably isn't one model for everything. It's using the right model for the right task.

Where do you draw the line when the pricier model barely changes the outcome?

Luckily he said that for all the companies which forgot to focus on return

fable for planning still wins for me. exec goes astra-low in codex so one thirsty agent doesn't cook the week.

If I could use fable for my full $200/mo plan I would use more too!

Nobody cares about the 50% you save on tokens, they notice the 100x return you got from actually using the model.

funny how model cost feels huge until you price the engineer waiting on it

This is so true

Totally agree on the direction. Only half of that ratio is measurable though: cost per model per repo is a query, the return is still a guess.

So no one dogfood on opus and sonnet right

The invoice is obvious. That 100,000% return only exists if you can show work that would not have happened otherwise.

I run the same math in production ML for finance. Token spend on fraud detection or KYC review is rounding error next to a single bad judgment call. Routing earns its keep on high-volume, low-stakes tasks.

Easy for him to say, he gets unlimited Mythos. For most of us that would cost more than we’re willing or able to pay.

easy take when your token bill says "anthropic internal"

Great salesman our Boris

the hard part is measuring return, not token cost. How do you track it?

The advice to use Fable 5 for everything is just bad as in, bad advice. We all have plenty of use cases that can be handled perfectly well by less capable models. The generous interpretation of Cherney's advice would be to use Fable 5 for any reasoning-intensive task. That said, I think that in the case of writing, for example, I would rather use a less capable model simply because it is more reliable and consistent than what we see from Fable 5, which is frankly underwhelming.

Developer time costs way more than token usage anyway



