Loading video...
Video Failed to Load
Boris Cherny, head of Claude Code at Anthropic, on optimizing token cost and model use. “I use Fable for everything" There is probably a 50% opportunity to reduce the investment (on token cost). However, there may be a 1,000%, 10,000%, or even 100,000% opportunity to increase the return. Therefore,... show more
57,774 views • 1 month ago •via X (Twitter)
33 Comments

"Carpenters shouldn't waste time measuring cuts when they can be building. Just buy more wood. Also, I work at a lumber company."

You can't use Fable for everything. Fable won't discuss middle school level biology. It also won't work with content you own if it thinks the content is copyright and with the most recent version it won't work on content it disagrees with morally.

guy who doesnt eat the cost: i dont think about cost

Yeah, "I use Fable for everything" because it's free for you. I understand the point he's trying to get across, but the tokens have become really expensive. The focus becomes cost-cutting if the subscriptions do not produce anything of value and continue to hemorrhage dollars.

Full video

Boris’ views are so removed from his customers’ reality it’s not funny. There’s a huge cost associated with this, not the mention the countless things you can’t do with fable because it refuses to. I love fable, but this is actually terrible advice

Boris is great, but this is an area that he has blinders on. “Fable for everything” at enterprise scale and current pricing would require large layoffs for sufficient ROI in most places.

A weird thing to say from someone selling tokens.

The philosophical views of someone who doesn't pay token costs. Enterprises already made this mistake and decided to pull models due to cost overruns.

Person who gets unlimited compute from their employer loves using the most expensive compute, more at 11 ET.

So true, the massive return you get from a top-tier model easily scales past whatever you save by cost-cutting.

boris also drives his ferrari to the supermarket

lol who cares what he says about cost optimization. his tokens are free. are tokens line his pockets.

true when the task actually has 1000x upside, most tasks do not, and using fable to reformat a csv is the ai version of a lambo grocery run

Wine maker says he uses wine for everything including washing the dishes

this is an abundance mindset

It’s a reasonable argument for individuals/businesses with uncapped growth potential. And there’s an argument to be made that Fable+ models can help unlock uncapped growth potential. Important to note that Boris’s tokens don’t have the same cost dynamics as yours or mine

i use the most expensive product our company sells for everything. 😁 j/k i love fable and opus

"Guys don't save your money, put more money in the slot machine because you'll increase your chance of eventually winning! Yes, I do own the slot machine, what's your point?"

Using the most advanced model to maximize returns is a smart strategy!

Of course thats what he will say. He’s in the business of selling you his expensive models.

Not the same thing as an expected value calculation of > 50% + desired IRR. Anthropic's pricing is indifferent to actual value generated by end user. To their credit, OpenAI rather bravely has tried to address this issue.

Focus on return not cost, underrated mindset

Boris's math held for about a week — then Opus 5 halved the price of 'the most expensive model' while beating it. How does the advice update when the ceiling keeps moving?

cheap model habit dies hard

Mine is a skeptical take but coming from the brand ambassador of coding from Anthropic, what else could we expect? He wouldn't say anything that can reduce the token premium paid by users.

Focusing entirely on maximizing the output instead of penny-pinching token costs changes everything.

The upside framing is right. On the savings side, the win comes from not re-paying for history you already sent. Tokens per turn is the number worth watching, and model routing is the step after that.

This is easy to forget when everyone is obsessed with making AI cheaper. Sometimes the bigger question is what you can actually do with it.

cost reduction is just reclaiming what you spent. the upside is what you never would have attempted without the agent there in the first place.

Solid take. The return question gets harder when you're locked into one model's lens. Comparing a few perspectives can surface assumptions the "best" model quietly buries.

That was a well executed pitch.

pretty happy with the cache hit rate after the last updates to ApexOS... 😅



