Loading video...
Video Failed to Load
1m context windows are a nice gimmick But you might be better off sticking to only the first 150K tokens:
134,525 views • 2 months ago •via X (Twitter)
49 Comments

Learn about this stuff here:

Fortunately at least the smart zone is expanding compared to the last time I saw it

Yep, I mention this

How did you decide to change your rule of thumb from 100k tokens to 150k tokens? Do you have any metrics or evals you keep track of to know where the smart and dumb zone are for your workflows?

No, just personal vibes

Yes. If your project is on the smaller side but you have little choice if you have a larger code base. I go up to 300k context window and then wrap things up..

This is a false dichotomy. Larger code bases don't require more context window. You just need to structure your codebase in a way that the agent can navigate it easily. Smaller files, more descriptive file system, better context pointers in AGENTS.md files.

If we indexed our codebases in a disciplined way, absolutely - it would be an easy lookup. But that rarely is the case. Importantly, too short a session for large feature builds, means you end up recreating the context for the next session quite often i.e. more cost. Agree with the overall point but for me, 300k works well.

A bigger window can hide bad context hygiene. Scoped sessions and explicit handoffs are easier to debug than one giant transcript.

Note: Before we design agents, tune prompts, or pick a model, understanding how to structure what goes into the context window is the most impactful skill in practical AI engineering. Most quality problems in production AI systems trace back to poor context design, not the wrong model or missing features.

Need a skill to auto-compact when reaching 150-200k 🥹

I don't think so - autocompacting can be disastrous when it happens mid-phase.

Yeap, I agree, that’s why I put 150-200, not always 150, so it actually can happen after some milestone like finished some part, not mid-phase

Possibly, though some phases spiral out of control and end up needing 300K tokens

It does make sense! If only models were smart enough to determine that automatically, I think we’ll got there at some point

Do you have any proof that for new models such as Sol and Fable the dumb zone start at 150k tokens. My main experience is that, the more context I give them, the better the results are. Things like discussion transcripts, slack threads etc.

You are joking though? 150k context? Ok, i understand 256k or similar, but someone advocating 150k as maximum needed context is clearly funny about it

Did you watch the video?

No, to be honest, but it is the second sentence that set me off. Will watch it later!

Thanks for your contribution to the discourse

I watched your video and while I agree with most of this stuff as it is pretty basic knowledge at this point, I do believe 250k sessions are pretty much fine and you will not feel that dumbness. 1M context sessions are also fine if you handle that context well and use tooling that steers the models correctly. Read the open papers on model attention spans and context management.

Funny how the AI narrative claims models are emulating human brains, yet they obsess over massive context windows that biology never uses. Neurons don’t hold massive context windows — they focus on what’s relevant to the current task, with a few rogue connections here and there. LLMs, running on raw compute and probability, often overdo context. At some point it just becomes noise: higher costs, more hallucinations, and diminishing returns. Less can be more.

Content window management is a solved issue if you're using Mastra Code. It uses Observational Memory, which keeps what the agent needs to remember within the small limits and is also cache friendly.

i love the short form content

It’s definitely important to be aware of the dumbzone but I personally prefer being able to use part of that expanded window for tail end tasks where smarts isn’t as important. Beats hitting a hard limit and having to roll the dice with compaction. That said, I’ve never hit 1M

I agree, this matches how I use it. I'm not scared of the dumb zone, it's fine sometimes

openai got it right with the 250k-ish token limit and crazy good compaction, I despite the 1M context window because it eats my usage limits like anything....sigh!

A silly question maybe, but what do you think needs to change to "expand" this smart zone? Will it always be a fraction of the complete size?

Second this. I have my Claude status line set to show me context usage out of 200k instead of 1M. It spills over which is neat. At this point I /handoff

I like these short videos! It's easy to share things, I am trying to say when introducing your skills repo to people. Also just for me to keep myself updated!

The context is only useful if the important stuff survives the noise I’d rather give an agent a smaller, clean project brief than a million tokens of old assumptions and random logs

yet again keep the context window small!!!! thx Matt

Could you help me with startups?

Why over complicate. Keep going in the same chat and it has all the context. Most tasks much easier to just complete in the regular context

Care to give any data on this? From my experience I don’t notice any degradation in Claude code even though the context window closes in on 1M tokens. I would guess they train the models explicitly at these large contexts as well

I guess this makes sense why Codex dropped their context back to 272k

What if you give your agent Vyvanse?

怎么让 coding Agent写到接近 15的时候自动 /compact 是一个问题, 我就是因为这个原因不使用 /goal 的 ,如果 goal 能配置一个范围让 agent 主动 /compact , 那将会非常好 @thsottiaux

Thanks for this. Even though it's fairly known, i like how clearly you present it here. Somehow using the whole 1m context window reminds me of the "EV cash out" option in online poker all-ins. You lose EV by accepting that option, so there should be no case that it makes sense to do so. Same here, entering the dumb zone is -EV and we should work around it ro avoid it!

Seriously, I don't really understand how people make 1M context even work. Models really become much less capable beyond 200-300k tokens.

yeah, quality falls off a cliff way before 1m.

I feel this days, just 3 prompts and you already at 400k for large codebase

Check out these great visualizations of context rot from Jake Minns, I really like them:

Great! How do you deal with Attention Degradation so it doesn't start hallucinate? What is the best way to keep - Relationship - Context Aware incase working with LARGE repo. To Avoid Dumbzone. - Does Saving Progress in "progress.md" works? - Thoughts on "/compact" command? Any better recommendation?

yep context engineering in a nutshell! thanks for the video.

How do you do this in autonomous loops with larger slices?

You run /to-tickets and read the slices first
another possibility: frequently use `/compact` to compact the context

Chasing million-token windows is purely a marketing race at this stage
