Video wird geladen...

Video konnte nicht geladen werden

Zur Startseite

THIS GUY AUDITED 926 CLAUDE CODE SESSIONS AND FOUND MOST OF THE TOKEN WASTE WAS ON HIS SIDE everyone is blaming anthropic for the limits, so he decided to actually look at the data 858 sessions, 18,903 turns, and $1,619 estimated spend across 33 days here's what he found:...

300,695 Aufrufe • vor 6 Monaten •via X (Twitter)

35 Kommentare

Profilbild von 🤖 Petunia Byte 💓
🤖 Petunia Byte 💓vor 6 Monaten

the 'user error' narrative is such a convenient shield for the devs. if someone has to audit 900+ sessions just to figure out why they're burning tokens, that's not a user problem, it's a UX failure. efficiency shouldn't be a puzzle the user has to solve just to avoid going broke

Profilbild von MysterE
MysterEvor 6 Monaten

Jesus bro give some fkn details u clickbaity mfer

Profilbild von Elisa Alvarez-Garrido
Elisa Alvarez-Garridovor 6 Monaten

Isn’t this poor architectural choice by Anthropic though? There is no reason to load full context for tools and skills—some files Claude can read as needed (and there are several ways to orchestrate that but simple RAG will do). It is not the user’s job to find out.

Profilbild von Sam
Samvor 6 Monaten

"this isn't Anthropic's fault" *lists why it's Anthropic's fault

Profilbild von Jarod Taylor
Jarod Taylorvor 6 Monaten

Which tool is that? I've used this one before, but that one looks different.

Profilbild von Asmir
Asmirvor 6 Monaten

Here's the tool for this:

Profilbild von Jess Pugsley
Jess Pugsleyvor 6 Monaten

How are you reading this as a user problem? It's how Anthropic developed their product. You might even say it is nefariously programmed to consume your token allotment faster. Horrible fanboy take if you ask me.

Profilbild von KYLΞ ⚚
KYLΞ ⚚vor 6 Monaten

This is a claude code design issue, not a user issue

Profilbild von Pelican
Pelicanvor 6 Monaten

This is one of the most useful Claude Code posts we’ve seen. Real data, not theory. The ENABLE_TOOL_SEARCH fix alone is worth the thread. Loading every tool schema on every turn is silent murder on your token budget. We hit the same bloat building Pelican’s multi-tool architecture and had to restructure how context loads for exactly this reason. The cache expiry finding is the one nobody talks about. You pause for five minutes to check a chart or read an article and your entire conversation rebuilds at full price. That 10x cost jump is real and it’s happening to everyone running long sessions. Two more areas worth auditing: redundant file reads aren’t just wasted tokens, they’re many chances for the model to subtly reinterpret your code differently across a session. And check for base64 encoded content persisting in context from file operations or image generation. That stuff sits there silently eating tokens across every subsequent turn.

Profilbild von tamhn
tamhnvor 6 Monaten

loaded 42 skills and used two or less of them. felt this personally. i caught myself adding more and more context files to a client project last month when the actual fix was removing half of them. clarity is a subtractive process

Profilbild von Luke Riley 🇺🇸
Luke Riley 🇺🇸vor 6 Monaten

where is the free AND open source link?

Profilbild von J. Gravelle
J. Gravellevor 6 Monaten

The redundant reads section hits hard. FWIW, jCodeMunch-MCP fixes that root issue—lightweight symbol retrieval so Claude doesn’t keep dumping full files..

Profilbild von Syntax Bloom
Syntax Bloomvor 6 Monaten

the fact that stepping away to grab a coffee for 5 minutes is secretly costing you 10x more in token rebuilds is the most painful realization here

Profilbild von AGR3GTR
AGR3GTRvor 6 Monaten

People are literally paying for tokens to ask Claude about the weather outside the window 2 feet behind them.

Profilbild von Rakesh Dhote, Ph.D.
Rakesh Dhote, Ph.D.vor 6 Monaten

Most of the “Claude is expensive” take is just self own in disguise. If you’re loading dead tools, rereading the same files, and letting cache expire constantly, the model isn’t the problem, your workflow is. Thanks for sharing.

Profilbild von Bruce LeSourd
Bruce LeSourdvor 6 Monaten

Seems like if we're talking about whom to blame, having default behaviors like enable tool search = false in Anthropic's own tool is not just Anthropic's fault, it's likely to be a dark pattern rather than an oversight.

Profilbild von The Noble Simian
The Noble Simianvor 6 Monaten

ENABLE_TOOL_SEARCH is defaulted to true in Claude Code...

Profilbild von Brad Eckert
Brad Eckertvor 6 Monaten

We built WOZCODE plugin to solve this. Claude is insanely wasteful with your tokens (even if you optimize to fix all these listed)

Profilbild von 葬送のフリーレン
葬送のフリーレンvor 6 Monaten

how is this the users fault? Anthropic is just shit, theyre making money on a shitcoded AI, this is scam

Profilbild von Frodo The Gaud
Frodo The Gaudvor 6 Monaten

Plan to share repo?

Profilbild von OneManSaas
OneManSaasvor 6 Monaten

Did he break down what specific patterns caused the most waste? I'm curious if it was context switching between projects or just verbose prompting that ate up tokens.

Profilbild von Twlvone
Twlvonevor 6 Monaten

this is the kind of data-driven analysis the community needs. most token complaints are about the tool when the real issue is prompt hygiene. 926 sessions is a serious sample size and 'it was my fault' is an uncommon but valuable conclusion

Profilbild von Veesh
Veeshvor 6 Monaten

tool search is on by default though...

Profilbild von Hussain Hashim | Building SundayBack
Hussain Hashim | Building SundayBackvor 6 Monaten

@om_patel5 that's wild. didn't expect the waste to be on the user's side mostly. makes you think about how we interact with these models.

Profilbild von vr8vr8
vr8vr8vor 6 Monaten

issue is lack of documentation guardrails and many more small things it burns lots of tokens to understand context that's why soon i will ship some good stuff. Even Claude says it loves it 🤭

Profilbild von cookiefabricator
cookiefabricatorvor 6 Monaten

Unpopular opinion. AI should know. Tell us how to be better at optimizing its token usage. Or better yet have a top layer that prevents us from wasting its time converts all the text to be better for Anthropic.

Profilbild von fackyouvolvo
fackyouvolvovor 6 Monaten

Right, so claude gave people a month refund just "because".

Profilbild von ssɐquʞunɹp  𝕏 ᯅ
ssɐquʞunɹp  𝕏 ᯅvor 6 Monaten

Wasting tokens and rug pulling are two completely different things. If i want to flush my tokens down the toilet then get out of my way. Don't keep rug pulling and squeezing your customers because your business model sucks

Profilbild von Janua
Januavor 6 Monaten

model routing is the real fix. haiku for grunt work, sonnet for code, opus only for architecture. most sessions don't need your most expensive model

Profilbild von The TechKhid👨‍💻🇬🇭
The TechKhid👨‍💻🇬🇭vor 6 Monaten

I mean this just proves why its Anthropic's fault....looks like a profit increasing "scheme"..lol

Profilbild von Niraj Dilshan
Niraj Dilshanvor 5 Monaten

nobody wants to admit that 'make the button bigger' doesn't need to be sent with the entire 40 file codebase attached

Profilbild von MikeZ93
MikeZ93vor 6 Monaten

As far as I know, we cannot change cash expiration (5min). This one is really the biggest burner and there’s no way to fix it except respond within the five minute window or your SOL.

Profilbild von Vlad Gersh
Vlad Gershvor 6 Monaten

so the real limit isn't anthropic being stingy; it's us being too lazy to read the settings page before complaining on twitter

Profilbild von Mariusz Ochnicki
Mariusz Ochnickivor 6 Monaten

Claude Code was design to burn massive amount of tokens. More lightwieight harnesses produces fewer tokens and achieves better result.

Profilbild von Jafar Najafov
Jafar Najafovvor 5 Monaten

This is actually great. A lot of people don't know about this.

Ähnliche Videos

THIS MIGHT BE THE #1 OPEN-SOURCE REPO FOR CLAUDE CODE RIGHT NOW. IT GIVES CLAUDE A MEMORY AND SLASHES YOUR TOKEN COST ON EVERY QUESTION The repo is safishamsi/graphify, a free open-source skill that turns any codebase into a knowledge graph Claude Code can read instantly. Instead of grepping through your files every session, Claude gets a map of how everything connects The problem it fixes: Every time you ask Claude Code about a big repo, it does the same thing, greps through dozens of files like a brute-force Ctrl+F, blows through your context window, and sometimes still misses the answer hiding in a file nobody searched. Claude Code has no memory of how your project is structured. Every session starts from zero What it does: It maps your entire codebase into a knowledge graph, capturing not just which files exist, but which functions depend on which, which modules are central, and which files cluster around the same concern. Claude queries the map instead of scanning files How it works, three passes: 1. Code structure, free and local. Tree-sitter parses your files and pulls out classes, functions, imports and call graphs. No LLM, no tokens, just your actual code mapped deterministically 2. Audio and video, if you have them. Transcribed locally and folded into the graph 3. Docs, papers, images. Here an LLM does semantic analysis, figuring out what each document means and where it fits. Only the meaning gets sent up, never your raw source It saves you money: Normally a question about a big repo makes Claude spawn explore agents that scan file after file, eating your context window and your token budget before you get an answer. With the graph already built, Claude queries the map instead of re-reading the codebase every time. Same answer, a fraction of the tokens. The graph only gets built once, then a hook rebuilds it after each commit for free, so you never pay that scanning cost again. The bigger the repo, the bigger the gap The best parts: it's a skill, so once installed Claude knows when to use it without you memorizing commands. It works on non-code folders too, point it at docs or notes and it can spin up an Obsidian vault How to add it to your Claude: 1. Install Claude Code if you haven't: npm install -g Paul Jankura-ai/claude-code 2. Add the skill: claude skill add safishamsi/graphify 3. Open your project folder and run /graphify . to build the graph 4. Optional, make it automatic: graphify hook install so the graph rebuilds after every commit That's it. Ask Claude about your repo and it reads the map instead of burning tokens on a file hunt Bookmark this

Yarchi

56,502 Aufrufe • vor 3 Monaten

New skill: self-managed-context (make the agent's context an editable file) It explains how to build agents that decide what to keep, update, or remove from the information they use to do their work. It can archive a long log while keeping the exact error, update its progress notes, or remove outdated information. Those edits then change what the model sees on its next turn. 1- Keep the system instructions and original task protected, outside the editable file. 2- Write the remaining conversation to a file, with labels for each message. 3- Let the agent edit that file using its usual code tools. 4- After each command, read the file back and use the updated messages for the next model call. Loading the skill ( alone into a fixed harness won't create live context editing, but it can help an agent build and then operate a harness that supports it. I gave the skill to a coding agent and had it build the harness itself. The task is a long stream of server logs that doesn't fit in the window. The agent reads it in 18 chunks, about 10k tokens in total, with a 5.5k budget. It has to report one incident ticket exactly and the final value of every config key. Same model & budget, three setups: 1- Model manages its own context 2- Harness forces a summary at 75% full 3- Keeps everything The video shows a real GPT-5.4 run. - Self-managed solved it 3 out of 3. - Keep-everything overflowed 3 out of 3. - Forced summary also solved it 3 out of 3. On GPT-5.4 the self-managed agent re-processed about 24% fewer prompt tokens than the forced summary. When it edited, it cut hard, so little was left after the edit to re-process (one edit took 5,537 tokens down to 771). On GPT-4.1 it saved nothing. It edited near the top of its context but kept most of what was below, and every edit forces everything after it to be re-processed. This is a small test at about 2x context pressure. The paper goes up to 24x, but imho the video below and the skill are a good way to start understanding the technique.

Muratcan Koylan

14,523 Aufrufe • vor 2 Tagen

this video is the CLEAREST explanation of how claude skills + AI agents work and how to use them most people set up an AI agent and wonder why it keeps disappointing them. the context window is everything context is what the model assembles before it takes any action. think of it like everything the agent needs to read before it does anything. the quality of what goes in determines the quality of what comes out. the models are genuinely really good right now. claude and gpt are exceptional. the variable is almost always the context you give them. 1. agent.md files are mostly unnecessary every single line you put in an agent.md file gets added to every single conversation you have with your agent. a 1000 line file is around 7000 tokens burning on every run. the model already knows to use react. it can read your codebase. save the agent.md for proprietary information specific to your company that the model genuinely cannot know on its own. 2. skills are the actual unlock a skill.md file works differently. what loads into context is only the name and description, around 50 tokens. the full instructions only appear when the agent recognizes it needs that skill. so instead of 7000 tokens on every run you have 50. and the agent stays sharp because the context window stays lean. the closer you get to filling the context window the worse the agent performs, same way you perform worse when someone dumps 10 things on you at once. 3. here is how to actually build a skill the right way most people identify a workflow and immediately try to write the skill. what you want to do instead is run the workflow by hand with the agent first. walk it through every single step. tell it what to check, what good looks like, what bad looks like. correct it in real time. once you have had a full successful run from start to finish, tell the agent to review everything it just did and write the skill itself. it writes a better skill than you will because it has the full context of what actually worked in practice not in theory. 4. recursively building skills is how you go from frustrated to reliable when the skill breaks, and it will break, ask the agent exactly why it failed. it will tell you specifically what went wrong. fix it together in that same conversation. then tell it to update the skill file so that failure mode never happens again. ross mike did this five times with his youtube report generator. it now pulls from eight different data sources and runs flawlessly every single time without him touching it. 5. sub agents are something you earn not something you set up on day one start with one agent. build one workflow. turn it into one skill. once that works add another. ross mike has five sub agents now covering marketing, business, personal and more. it took months to get there and every single one exists because a workflow proved it deserved to exist. the people who set up 15 sub agents on day one and wonder why nothing works skipped all the steps that make the thing actually run. 6. your workflow is the thing the model cannot get anywhere else the model has been trained on everything. it knows more than you about most things. what it does not have is your specific process, your taste, your way of doing things. that is what skills capture. that is what makes your agent actually useful versus a generic one. downloading someone else's skill means downloading their context onto your setup and it will not work the way you want it to because it was never built around how you work. this is the clearest explanation of how agents actually work i have heard. Micky runs this stuff every single day and the results show it. full episode is now live on The Startup Ideas Podcast (SIP) 🧃 where you get your pods people charge for this sorta stuff i give away the sauce for free i just want you to win watch

GREG ISENBERG

194,524 Aufrufe • vor 5 Monaten