正在加载视频...
视频加载失败
This is the best way to use Claude Fable in Claude Code without immediately hitting your limits. 1. Model set to Fable 5 2. Reasoning on Max 3. Instruct Claude to run a dynamic workflow where: 3a. Fable is the orchestrator 3b. Opus does the reasoning heavy phases Fable... show more
39 条评论

Adding this to my global CLAUDE.md helped tremendously: ## Claude Fable: token parsimony When running as Fable (expensive), plan and review; delegate implementation to subagents (`model: sonnet` for code, `haiku` for mechanical edits/searches), one task per subagent. Trivial single-file edits are fine to do directly.

Very nice. Thank you for sharing!

It’s funny that in 6 Months we will think Fable 5 is now the “dumb” model.

Totally

also install official openai /codex skill to delegate to codex

You find fable knows when to delegate to lesser models? I find it amazing at delegating but worried about if it can really give it to sonnet vs opus reliably

Still just starting to experiment but seems pretty great so far.

Well there goes my evening - might as well try to get it to hook into 5.5 as well

If you don't instruct it not to use fable as the workflow agents it will sometimes use fable. I had it spawn 92 fable in parallel and max out my 20x 5h limit in 1 min by accident 💀

Is it me or Fable does not feel that much better than Opus 4.8 like everyone claims?

Reasoning Max is where you lost me. I feel like if Fable low is like Opus 4.8 high, then using it on max is not needed for most tasks. Especially if you don't want to hit your limits.

remember when opus was the one who orchestrated sonnets? pepperidge farms remembers.

Tell Claude Code to design a communication path to Codex. Fable 5 = architect \ orchestrator GPT 5.5 = implementer \ “doer” My Claude code and Codex talk to each other all day now. I just look at the big picture with Claude code.

Fable seems strong at delegating to other models and even to select which model/effort to use and when. He also seems very happy to use codex as peer-review (in my setup he uses it via hermes). Has been running 36 hours non-stop on AI research so far so good! Additionally I would say is 10% better than GPT 5.5 and like 30% better than Opus.

And add Sonnet for any low effort stuff in your AI agents. Took me almost 6h today to harness Fable power without wasting account in hours.

Dynamic workflows will absolutely destroy your limit. But agreed on the approach- use fable as the orchestrator, it should use subagents opus/codex for exploration, implementation, and review. Same concept, drop workflows

Model routing is the new prompt engineering.

I have a better super budget friendly option! Treat fable as philosopher/designer take the plan & execute by Opencode go with Orchestrator as GPT 5.5 fast!

The message is only "without immediately hitting your limits". I'm exactly using the same approach and I'm still hitting my weekly limit within 2 days 🫠

This is true. It’s been the case for me. Fable is insane at orchestrating so the cost isn’t even that high

Been running this exact model split in production for months — non-developer, no code written. Phase 1 is pure orchestration: architecture, blueprint.md, context engineering. Phase 2 is where the heavy reasoning happens with worker agents — had 177 tasks running in Augment Code's Intent tool while I reviewed the architecture layer. Phase 3 is adversarial review with a separate context window. Fable as orchestrator, heavier model for reasoning-dense phases — that's not a tip, that's the only architecture that doesn't collapse under its own token cost by hour two.

So the orchestrator just kicks back like a project manager, while the heavy lifter does the actual math. Kinda mirrors how I wire up my robot controllers anyway. Delegate the hard stuff, keep the main thread cool.

All this and you have a Pro subscription lol you barely say hi and you're already out of tokens💀

…so the plan is to keep the coherent mind in a glass office writing tickets while Opus does the actual thinking….in fragments, and then reassemble the fragments and wonder why the output has no spine?

curious how you handle context handoffs when opus runs the heavy reasoning steps. do you notice latency spikes or token overhead when passing state back to the fable orchestrator?

do you ask fable agents to sub-agents that use smaller models?

Or even better use Fable as the orchestrator and Sonnet or GPT-54 as the implementer, with a review gate between each step. Started with a very high-level plan. Ninety-five percent of your code will be clean. Then you automate the entire process.

I’ve been using Mythos on low to help with my hourly limits

interesting. thanks for the inform... i need to study it, we are alredy moving from: which modl is best? to which model should manage the others? thats a completely different game.

Best way to get its perks without burning too much tokens.

Fable is surprisingly dumb for the superpower. Back to Opus.

I’m A max everything kind of girl

nah I made it lauch sonnet agents to audit the issue flew over all of their heads then told fable to do it and it caught it

Brilliant

Yeah - orchestrator approach works great... even then, the few things Fable does eats a lot of usage quick, even on LOW. So not immediate depletion, just accelerated. Using Opus 4.8 to drive Sonnet 4.6 swarms is still effective. But adds up, too.

I might be wrong here but I remember reading in anthropic's blog that switching models in between sessions costs more tokens unless a handoff is used

TIL

Does the orchestrator split hold up on long unattended runs? Limits never bite us during live code sessions, it's the overnight automation jobs that blow through caps. If this survives those it's worth way more than a coding trick.

Smart

