Loading video...

Video Failed to Load

Go Home

1m context windows are a nice gimmick But you might be better off sticking to only the first 150K tokens:

134,525 views • 2 months ago •via X (Twitter)

49 Comments

Matt Pocock's profile picture
Matt Pocock2 months ago

Learn about this stuff here:

Chawye Hsu's profile picture
Chawye Hsu2 months ago

Fortunately at least the smart zone is expanding compared to the last time I saw it

Matt Pocock's profile picture
Matt Pocock2 months ago

Yep, I mention this

Sam Rose's profile picture
Sam Rose2 months ago

How did you decide to change your rule of thumb from 100k tokens to 150k tokens? Do you have any metrics or evals you keep track of to know where the smart and dumb zone are for your workflows?

Matt Pocock's profile picture
Matt Pocock2 months ago

No, just personal vibes

Santosh Kathira's profile picture
Santosh Kathira2 months ago

Yes. If your project is on the smaller side but you have little choice if you have a larger code base. I go up to 300k context window and then wrap things up..

Matt Pocock's profile picture
Matt Pocock2 months ago

This is a false dichotomy. Larger code bases don't require more context window. You just need to structure your codebase in a way that the agent can navigate it easily. Smaller files, more descriptive file system, better context pointers in AGENTS.md files.

Santosh Kathira's profile picture
Santosh Kathira2 months ago

If we indexed our codebases in a disciplined way, absolutely - it would be an easy lookup. But that rarely is the case. Importantly, too short a session for large feature builds, means you end up recreating the context for the next session quite often i.e. more cost. Agree with the overall point but for me, 300k works well.

Sesori - Claude, Codex & OpenCode on Mobile's profile picture
Sesori - Claude, Codex & OpenCode on Mobile2 months ago

A bigger window can hide bad context hygiene. Scoped sessions and explicit handoffs are easier to debug than one giant transcript.

Petar Ivanov's profile picture
Petar Ivanov2 months ago

Note: Before we design agents, tune prompts, or pick a model, understanding how to structure what goes into the context window is the most impactful skill in practical AI engineering. Most quality problems in production AI systems trace back to poor context design, not the wrong model or missing features.

alex's profile picture
alex2 months ago

Need a skill to auto-compact when reaching 150-200k 🥹

Matt Pocock's profile picture
Matt Pocock2 months ago

I don't think so - autocompacting can be disastrous when it happens mid-phase.

alex's profile picture
alex2 months ago

Yeap, I agree, that’s why I put 150-200, not always 150, so it actually can happen after some milestone like finished some part, not mid-phase

Matt Pocock's profile picture
Matt Pocock2 months ago

Possibly, though some phases spiral out of control and end up needing 300K tokens

alex's profile picture
alex2 months ago

It does make sense! If only models were smart enough to determine that automatically, I think we’ll got there at some point

Kasper Peulen's profile picture
Kasper Peulen2 months ago

Do you have any proof that for new models such as Sol and Fable the dumb zone start at 150k tokens. My main experience is that, the more context I give them, the better the results are. Things like discussion transcripts, slack threads etc.

muhdur's profile picture
muhdur2 months ago

You are joking though? 150k context? Ok, i understand 256k or similar, but someone advocating 150k as maximum needed context is clearly funny about it

Matt Pocock's profile picture
Matt Pocock2 months ago

Did you watch the video?

muhdur's profile picture
muhdur2 months ago

No, to be honest, but it is the second sentence that set me off. Will watch it later!

Matt Pocock's profile picture
Matt Pocock2 months ago

Thanks for your contribution to the discourse

muhdur's profile picture
muhdur2 months ago

I watched your video and while I agree with most of this stuff as it is pretty basic knowledge at this point, I do believe 250k sessions are pretty much fine and you will not feel that dumbness. 1M context sessions are also fine if you handle that context well and use tooling that steers the models correctly. Read the open papers on model attention spans and context management.

Daniel Chapell's profile picture
Daniel Chapell2 months ago

Funny how the AI narrative claims models are emulating human brains, yet they obsess over massive context windows that biology never uses. Neurons don’t hold massive context windows — they focus on what’s relevant to the current task, with a few rogue connections here and there. LLMs, running on raw compute and probability, often overdo context. At some point it just becomes noise: higher costs, more hallucinations, and diminishing returns. Less can be more.

Hashim Warren's profile picture
Hashim Warren2 months ago

Content window management is a solved issue if you're using Mastra Code. It uses Observational Memory, which keeps what the agent needs to remember within the small limits and is also cache friendly.

Réda's profile picture
Réda2 months ago

i love the short form content

lewis's profile picture
lewis2 months ago

It’s definitely important to be aware of the dumbzone but I personally prefer being able to use part of that expanded window for tail end tasks where smarts isn’t as important. Beats hitting a hard limit and having to roll the dice with compaction. That said, I’ve never hit 1M

Matt Pocock's profile picture
Matt Pocock2 months ago

I agree, this matches how I use it. I'm not scared of the dumb zone, it's fine sometimes

Daily Patch Notes's profile picture
Daily Patch Notes2 months ago

openai got it right with the 250k-ish token limit and crazy good compaction, I despite the 1M context window because it eats my usage limits like anything....sigh!

Elhaam's profile picture
Elhaam2 months ago

A silly question maybe, but what do you think needs to change to "expand" this smart zone? Will it always be a fraction of the complete size?

Harry John's profile picture
Harry John2 months ago

Second this. I have my Claude status line set to show me context usage out of 200k instead of 1M. It spills over which is neat. At this point I /handoff

Pragnesh Kumar's profile picture
Pragnesh Kumar2 months ago

I like these short videos! It's easy to share things, I am trying to say when introducing your skills repo to people. Also just for me to keep myself updated!

Alan Acuña's profile picture
Alan Acuña2 months ago

The context is only useful if the important stuff survives the noise I’d rather give an agent a smaller, clean project brief than a million tokens of old assumptions and random logs

Riccardo Causo's profile picture
Riccardo Causo2 months ago

yet again keep the context window small!!!! thx Matt

Harsh Patel QuoraHarshEntrepreneur's profile picture
Harsh Patel QuoraHarshEntrepreneur2 months ago

Could you help me with startups?

Elie Steinbock's profile picture
Elie Steinbock2 months ago

Why over complicate. Keep going in the same chat and it has all the context. Most tasks much easier to just complete in the regular context

Alex's profile picture
Alex2 months ago

Care to give any data on this? From my experience I don’t notice any degradation in Claude code even though the context window closes in on 1M tokens. I would guess they train the models explicitly at these large contexts as well

Ashkan's profile picture
Ashkan2 months ago

I guess this makes sense why Codex dropped their context back to 272k

Jarod Taylor's profile picture
Jarod Taylor2 months ago

What if you give your agent Vyvanse?

YaCo's profile picture
YaCo2 months ago

怎么让 coding Agent写到接近 15的时候自动 /compact 是一个问题, 我就是因为这个原因不使用 /goal 的 ,如果 goal 能配置一个范围让 agent 主动 /compact , 那将会非常好 @thsottiaux

Evangelos Kostopoulos's profile picture
Evangelos Kostopoulos2 months ago

Thanks for this. Even though it's fairly known, i like how clearly you present it here. Somehow using the whole 1m context window reminds me of the "EV cash out" option in online poker all-ins. You lose EV by accepting that option, so there should be no case that it makes sense to do so. Same here, entering the dumb zone is -EV and we should work around it ro avoid it!

Timur Çakmakoğlu's profile picture
Timur Çakmakoğlu2 months ago

Seriously, I don't really understand how people make 1M context even work. Models really become much less capable beyond 200-300k tokens.

Liocoh's profile picture
Liocoh2 months ago

yeah, quality falls off a cliff way before 1m.

Roy Canani's profile picture
Roy Canani2 months ago

I feel this days, just 3 prompts and you already at 400k for large codebase

Fabian Schreiber's profile picture
Fabian Schreiber2 months ago

Check out these great visualizations of context rot from Jake Minns, I really like them:

Khizer Rehan's profile picture
Khizer Rehan2 months ago

Great! How do you deal with Attention Degradation so it doesn't start hallucinate? What is the best way to keep - Relationship - Context Aware incase working with LARGE repo. To Avoid Dumbzone. - Does Saving Progress in "progress.md" works? - Thoughts on "/compact" command? Any better recommendation?

yash's profile picture
yash2 months ago

yep context engineering in a nutshell! thanks for the video.

Rafael Oliveira's profile picture
Rafael Oliveira2 months ago

How do you do this in autonomous loops with larger slices?

Matt Pocock's profile picture
Matt Pocock2 months ago

You run /to-tickets and read the slices first

Jerome Etienne #AI's profile picture
Jerome Etienne #AI2 months ago

another possibility: frequently use `/compact` to compact the context

Kai Benetti's profile picture
Kai Benetti2 months ago

Chasing million-token windows is purely a marketing race at this stage

Related Videos