Video yükleniyor...

Video Yüklenemedi

Ana Sayfaya Dön

THIS GUY AUDITED 926 CLAUDE CODE SESSIONS AND FOUND MOST OF THE TOKEN WASTE WAS ON HIS SIDE everyone is blaming anthropic for the limits, so he decided to actually look at the data 858 sessions, 18,903 turns, and $1,619 estimated spend across 33 days here's what he found:...

300,695 görüntüleme • 6 ay önce •via X (Twitter)

35 Yorum

🤖 Petunia Byte 💓 profil fotoğrafı
🤖 Petunia Byte 💓6 ay önce

the 'user error' narrative is such a convenient shield for the devs. if someone has to audit 900+ sessions just to figure out why they're burning tokens, that's not a user problem, it's a UX failure. efficiency shouldn't be a puzzle the user has to solve just to avoid going broke

MysterE profil fotoğrafı
MysterE6 ay önce

Jesus bro give some fkn details u clickbaity mfer

Elisa Alvarez-Garrido profil fotoğrafı
Elisa Alvarez-Garrido6 ay önce

Isn’t this poor architectural choice by Anthropic though? There is no reason to load full context for tools and skills—some files Claude can read as needed (and there are several ways to orchestrate that but simple RAG will do). It is not the user’s job to find out.

Sam profil fotoğrafı
Sam6 ay önce

"this isn't Anthropic's fault" *lists why it's Anthropic's fault

Jarod Taylor profil fotoğrafı
Jarod Taylor6 ay önce

Which tool is that? I've used this one before, but that one looks different.

Asmir profil fotoğrafı
Asmir6 ay önce

Here's the tool for this:

Jess Pugsley profil fotoğrafı
Jess Pugsley6 ay önce

How are you reading this as a user problem? It's how Anthropic developed their product. You might even say it is nefariously programmed to consume your token allotment faster. Horrible fanboy take if you ask me.

KYLΞ ⚚ profil fotoğrafı
KYLΞ ⚚6 ay önce

This is a claude code design issue, not a user issue

Pelican profil fotoğrafı
Pelican6 ay önce

This is one of the most useful Claude Code posts we’ve seen. Real data, not theory. The ENABLE_TOOL_SEARCH fix alone is worth the thread. Loading every tool schema on every turn is silent murder on your token budget. We hit the same bloat building Pelican’s multi-tool architecture and had to restructure how context loads for exactly this reason. The cache expiry finding is the one nobody talks about. You pause for five minutes to check a chart or read an article and your entire conversation rebuilds at full price. That 10x cost jump is real and it’s happening to everyone running long sessions. Two more areas worth auditing: redundant file reads aren’t just wasted tokens, they’re many chances for the model to subtly reinterpret your code differently across a session. And check for base64 encoded content persisting in context from file operations or image generation. That stuff sits there silently eating tokens across every subsequent turn.

tamhn profil fotoğrafı
tamhn6 ay önce

loaded 42 skills and used two or less of them. felt this personally. i caught myself adding more and more context files to a client project last month when the actual fix was removing half of them. clarity is a subtractive process

Luke Riley 🇺🇸 profil fotoğrafı
Luke Riley 🇺🇸6 ay önce

where is the free AND open source link?

J. Gravelle profil fotoğrafı
J. Gravelle6 ay önce

The redundant reads section hits hard. FWIW, jCodeMunch-MCP fixes that root issue—lightweight symbol retrieval so Claude doesn’t keep dumping full files..

Syntax Bloom profil fotoğrafı
Syntax Bloom6 ay önce

the fact that stepping away to grab a coffee for 5 minutes is secretly costing you 10x more in token rebuilds is the most painful realization here

AGR3GTR profil fotoğrafı
AGR3GTR6 ay önce

People are literally paying for tokens to ask Claude about the weather outside the window 2 feet behind them.

Rakesh Dhote, Ph.D. profil fotoğrafı
Rakesh Dhote, Ph.D.6 ay önce

Most of the “Claude is expensive” take is just self own in disguise. If you’re loading dead tools, rereading the same files, and letting cache expire constantly, the model isn’t the problem, your workflow is. Thanks for sharing.

Bruce LeSourd profil fotoğrafı
Bruce LeSourd6 ay önce

Seems like if we're talking about whom to blame, having default behaviors like enable tool search = false in Anthropic's own tool is not just Anthropic's fault, it's likely to be a dark pattern rather than an oversight.

The Noble Simian profil fotoğrafı
The Noble Simian6 ay önce

ENABLE_TOOL_SEARCH is defaulted to true in Claude Code...

Brad Eckert profil fotoğrafı
Brad Eckert6 ay önce

We built WOZCODE plugin to solve this. Claude is insanely wasteful with your tokens (even if you optimize to fix all these listed)

葬送のフリーレン profil fotoğrafı
葬送のフリーレン6 ay önce

how is this the users fault? Anthropic is just shit, theyre making money on a shitcoded AI, this is scam

Frodo The Gaud profil fotoğrafı
Frodo The Gaud6 ay önce

Plan to share repo?

OneManSaas profil fotoğrafı
OneManSaas6 ay önce

Did he break down what specific patterns caused the most waste? I'm curious if it was context switching between projects or just verbose prompting that ate up tokens.

Twlvone profil fotoğrafı
Twlvone6 ay önce

this is the kind of data-driven analysis the community needs. most token complaints are about the tool when the real issue is prompt hygiene. 926 sessions is a serious sample size and 'it was my fault' is an uncommon but valuable conclusion

Veesh profil fotoğrafı
Veesh6 ay önce

tool search is on by default though...

Hussain Hashim | Building SundayBack profil fotoğrafı
Hussain Hashim | Building SundayBack6 ay önce

@om_patel5 that's wild. didn't expect the waste to be on the user's side mostly. makes you think about how we interact with these models.

vr8vr8 profil fotoğrafı
vr8vr86 ay önce

issue is lack of documentation guardrails and many more small things it burns lots of tokens to understand context that's why soon i will ship some good stuff. Even Claude says it loves it 🤭

cookiefabricator profil fotoğrafı
cookiefabricator6 ay önce

Unpopular opinion. AI should know. Tell us how to be better at optimizing its token usage. Or better yet have a top layer that prevents us from wasting its time converts all the text to be better for Anthropic.

fackyouvolvo profil fotoğrafı
fackyouvolvo6 ay önce

Right, so claude gave people a month refund just "because".

ssɐquʞunɹp  𝕏 ᯅ profil fotoğrafı
ssɐquʞunɹp  𝕏 ᯅ6 ay önce

Wasting tokens and rug pulling are two completely different things. If i want to flush my tokens down the toilet then get out of my way. Don't keep rug pulling and squeezing your customers because your business model sucks

Janua profil fotoğrafı
Janua6 ay önce

model routing is the real fix. haiku for grunt work, sonnet for code, opus only for architecture. most sessions don't need your most expensive model

The TechKhid👨‍💻🇬🇭 profil fotoğrafı
The TechKhid👨‍💻🇬🇭6 ay önce

I mean this just proves why its Anthropic's fault....looks like a profit increasing "scheme"..lol

Niraj Dilshan profil fotoğrafı
Niraj Dilshan5 ay önce

nobody wants to admit that 'make the button bigger' doesn't need to be sent with the entire 40 file codebase attached

MikeZ93 profil fotoğrafı
MikeZ936 ay önce

As far as I know, we cannot change cash expiration (5min). This one is really the biggest burner and there’s no way to fix it except respond within the five minute window or your SOL.

Vlad Gersh profil fotoğrafı
Vlad Gersh6 ay önce

so the real limit isn't anthropic being stingy; it's us being too lazy to read the settings page before complaining on twitter

Mariusz Ochnicki profil fotoğrafı
Mariusz Ochnicki6 ay önce

Claude Code was design to burn massive amount of tokens. More lightwieight harnesses produces fewer tokens and achieves better result.

Jafar Najafov profil fotoğrafı
Jafar Najafov5 ay önce

This is actually great. A lot of people don't know about this.

Benzer Videolar

THIS MIGHT BE THE #1 OPEN-SOURCE REPO FOR CLAUDE CODE RIGHT NOW. IT GIVES CLAUDE A MEMORY AND SLASHES YOUR TOKEN COST ON EVERY QUESTION The repo is safishamsi/graphify, a free open-source skill that turns any codebase into a knowledge graph Claude Code can read instantly. Instead of grepping through your files every session, Claude gets a map of how everything connects The problem it fixes: Every time you ask Claude Code about a big repo, it does the same thing, greps through dozens of files like a brute-force Ctrl+F, blows through your context window, and sometimes still misses the answer hiding in a file nobody searched. Claude Code has no memory of how your project is structured. Every session starts from zero What it does: It maps your entire codebase into a knowledge graph, capturing not just which files exist, but which functions depend on which, which modules are central, and which files cluster around the same concern. Claude queries the map instead of scanning files How it works, three passes: 1. Code structure, free and local. Tree-sitter parses your files and pulls out classes, functions, imports and call graphs. No LLM, no tokens, just your actual code mapped deterministically 2. Audio and video, if you have them. Transcribed locally and folded into the graph 3. Docs, papers, images. Here an LLM does semantic analysis, figuring out what each document means and where it fits. Only the meaning gets sent up, never your raw source It saves you money: Normally a question about a big repo makes Claude spawn explore agents that scan file after file, eating your context window and your token budget before you get an answer. With the graph already built, Claude queries the map instead of re-reading the codebase every time. Same answer, a fraction of the tokens. The graph only gets built once, then a hook rebuilds it after each commit for free, so you never pay that scanning cost again. The bigger the repo, the bigger the gap The best parts: it's a skill, so once installed Claude knows when to use it without you memorizing commands. It works on non-code folders too, point it at docs or notes and it can spin up an Obsidian vault How to add it to your Claude: 1. Install Claude Code if you haven't: npm install -g Paul Jankura-ai/claude-code 2. Add the skill: claude skill add safishamsi/graphify 3. Open your project folder and run /graphify . to build the graph 4. Optional, make it automatic: graphify hook install so the graph rebuilds after every commit That's it. Ask Claude about your repo and it reads the map instead of burning tokens on a file hunt Bookmark this

Yarchi

56,502 görüntüleme • 3 ay önce

New skill: self-managed-context (make the agent's context an editable file) It explains how to build agents that decide what to keep, update, or remove from the information they use to do their work. It can archive a long log while keeping the exact error, update its progress notes, or remove outdated information. Those edits then change what the model sees on its next turn. 1- Keep the system instructions and original task protected, outside the editable file. 2- Write the remaining conversation to a file, with labels for each message. 3- Let the agent edit that file using its usual code tools. 4- After each command, read the file back and use the updated messages for the next model call. Loading the skill ( alone into a fixed harness won't create live context editing, but it can help an agent build and then operate a harness that supports it. I gave the skill to a coding agent and had it build the harness itself. The task is a long stream of server logs that doesn't fit in the window. The agent reads it in 18 chunks, about 10k tokens in total, with a 5.5k budget. It has to report one incident ticket exactly and the final value of every config key. Same model & budget, three setups: 1- Model manages its own context 2- Harness forces a summary at 75% full 3- Keeps everything The video shows a real GPT-5.4 run. - Self-managed solved it 3 out of 3. - Keep-everything overflowed 3 out of 3. - Forced summary also solved it 3 out of 3. On GPT-5.4 the self-managed agent re-processed about 24% fewer prompt tokens than the forced summary. When it edited, it cut hard, so little was left after the edit to re-process (one edit took 5,537 tokens down to 771). On GPT-4.1 it saved nothing. It edited near the top of its context but kept most of what was below, and every edit forces everything after it to be re-processed. This is a small test at about 2x context pressure. The paper goes up to 24x, but imho the video below and the skill are a good way to start understanding the technique.

Muratcan Koylan

14,523 görüntüleme • 2 gün önce

this video is the CLEAREST explanation of how claude skills + AI agents work and how to use them most people set up an AI agent and wonder why it keeps disappointing them. the context window is everything context is what the model assembles before it takes any action. think of it like everything the agent needs to read before it does anything. the quality of what goes in determines the quality of what comes out. the models are genuinely really good right now. claude and gpt are exceptional. the variable is almost always the context you give them. 1. agent.md files are mostly unnecessary every single line you put in an agent.md file gets added to every single conversation you have with your agent. a 1000 line file is around 7000 tokens burning on every run. the model already knows to use react. it can read your codebase. save the agent.md for proprietary information specific to your company that the model genuinely cannot know on its own. 2. skills are the actual unlock a skill.md file works differently. what loads into context is only the name and description, around 50 tokens. the full instructions only appear when the agent recognizes it needs that skill. so instead of 7000 tokens on every run you have 50. and the agent stays sharp because the context window stays lean. the closer you get to filling the context window the worse the agent performs, same way you perform worse when someone dumps 10 things on you at once. 3. here is how to actually build a skill the right way most people identify a workflow and immediately try to write the skill. what you want to do instead is run the workflow by hand with the agent first. walk it through every single step. tell it what to check, what good looks like, what bad looks like. correct it in real time. once you have had a full successful run from start to finish, tell the agent to review everything it just did and write the skill itself. it writes a better skill than you will because it has the full context of what actually worked in practice not in theory. 4. recursively building skills is how you go from frustrated to reliable when the skill breaks, and it will break, ask the agent exactly why it failed. it will tell you specifically what went wrong. fix it together in that same conversation. then tell it to update the skill file so that failure mode never happens again. ross mike did this five times with his youtube report generator. it now pulls from eight different data sources and runs flawlessly every single time without him touching it. 5. sub agents are something you earn not something you set up on day one start with one agent. build one workflow. turn it into one skill. once that works add another. ross mike has five sub agents now covering marketing, business, personal and more. it took months to get there and every single one exists because a workflow proved it deserved to exist. the people who set up 15 sub agents on day one and wonder why nothing works skipped all the steps that make the thing actually run. 6. your workflow is the thing the model cannot get anywhere else the model has been trained on everything. it knows more than you about most things. what it does not have is your specific process, your taste, your way of doing things. that is what skills capture. that is what makes your agent actually useful versus a generic one. downloading someone else's skill means downloading their context onto your setup and it will not work the way you want it to because it was never built around how you work. this is the clearest explanation of how agents actually work i have heard. Micky runs this stuff every single day and the results show it. full episode is now live on The Startup Ideas Podcast (SIP) 🧃 where you get your pods people charge for this sorta stuff i give away the sauce for free i just want you to win watch

GREG ISENBERG

194,524 görüntüleme • 5 ay önce