Loading video...

Video Failed to Load

Go Home

Anthropic shipped 125 settings for Claude The official docs cover 40 One developer found the other 85 and his API bill dropped from $340 to $87 - not by using a cheaper model - not by writing shorter prompts just by moving one line in a config file to...

380,456 views • 4 months ago •via X (Twitter)

0 Comments

No comments available

Comments from the original post will appear here

Related Videos

Jev + Opus 5.5: Anthropic's new model beats GPT-6 Astra for 1/5 the cost, and 4 API changes will 400 your agent before it writes a single line I pulled these 10 steps from the migration docs so you don't learn them in production step 1 → $4 / $20 per 1M. Opus 5 was $5 / $25. cache reads dropped from $0.50 to $0.20 step 2 → 66.4% on Terminal-Bench 4.0 vs GPT-6 Astra 57.9% and Opus 5 52.3%. +14.1 points in one release, and on FrontierCode it beats Astra at default effort for 1/5 the cost step 3 → thinking can't be turned off anymore. send thinking: disabled and you get a 400. drop the field, set effort step 4 → tool_choice any and tool are gone. 400. switch to auto + strict step 5 → edit anything above a thinking block and the request dies. append only, or opt into drop_block step 6 → computer_20251124 is dead on the API. 400. move to computer_toolset_20260801 step 7 → the quiet one: default effort fell from high to medium. your agent thinks less than you set it up to and nothing tells you step 8 → hop Opus 5.5 → Sonnet 5 → Opus 5.5 and you pay 4.36 instead of 3.32. +31%, the cache dies and Sonnet can't read Opus's reasoning step 9 → change effort at the top of the request and the cache is gone. Jev sets it per message and the cache stays step 10 → switch fast - standard mid-session and it's a full cache miss. Jev picks speed once, on turn one one model, three knobs, zero 400s. that is Jev + Opus 5.5 send this to your Claude Code before you touch the model ID, then read my full Jev deep dive in the article below ↓

Carnage

16,674 views • 5 days ago

A developer in Hangzhou runs an AI that remembers everything about him for $0.40 a year. No vector database. One file that never grows past 4,000 tokens. He published the whole schema. His version starts from the opposite idea. Memory is not storage. It's a write policy. Six fields. Rewritten every time, never appended: > IDENTITY - who you are, what you build. 300 tokens. Changes monthly at most > STATE - what you're on right now. 400 tokens. Rewritten daily > DECISIONS - what's already settled, so nothing gets re-argued. 800 tokens > CORRECTIONS - every time you said "no, not like that." 600 tokens > PEOPLE - names, roles, who's waiting on what. 500 tokens > DEAD - tried and abandoned, so it never comes back as a suggestion. 400 tokens Three thousand tokens. Ceiling of four. When a section fills, the model rewrites it shorter. Nothing is ever added. Only replaced. Kimi K2.5 bills $0.10 per million cached input tokens. Four thousand tokens a turn is $0.0004. That's 2,500 turns for a dollar. The free tier hands you 1.5 million tokens a day. 375 turns before you pay anything at all. CORRECTIONS is the field nobody builds, and it's the one that does the work. A model that remembers being wrong stops repeating it. Everyone else is paying to search their own history. He pays to keep it short. The bill stopped growing when the file did. Your memory system isn't defined by what it stores. It's defined by what it agrees to delete. The article below is the full build - schema, rewrite prompts, the compaction rule that keeps it under the cap. Save it. You'll want it open in the other tab.

wast3

15,862 views • 1 month ago

How to cut your AI bill by 60% switching to Opus 5.5: Most people will swap the model name, save 24%, and stop there. The other 37% is sitting in your settings. Here's the real before and after on an example agent workload. One month: 100M input tokens (80M cached reads, 5M cache writes, 15M uncached) 10M output tokens BEFORE: Opus 5 Cache reads: 80M × $0.50 = $40 Cache writes: 5M × $6.25 = $31.25 Uncached input: 15M × $5 = $75 Output: 10M × $25 = $250 Total: $396.25 STEP 1. Just switch the model. Same tokens, new prices. Cache reads $0.20. Writes $5. Input $4. Output $20. $16 + $25 + $60 + $200 = $301 24% cheaper. That's what everyone's screenshotting. STEP 2. Let it write less. Box measured Opus 5.5 at 40% less verbose with no drop in accuracy. 10M output tokens becomes 6M. Output drops from $200 to $120. Total: $221. Now you're at 44% off. STEP 3. Stop running every call at max effort. Thinking can't be switched off on 5.5 anymore, so effort is your lever. Classifying, routing, formatting, summarizing? Drop the effort. Save the high settings for the calls that actually reason. Say that trims output another 25%, 6M down to 4.5M. Output: $90. Total: $191. 52% off. STEP 4. Cache the stuff you keep resending. This is the one nobody does. Cache reads went from $0.50 to $0.20. That's 60% off the cheapest line on your bill. Move your system prompt, tool definitions and repo context into the cache. Uncached input drops from 15M to 5M, cache reads go up to 90M. $18 + $25 + $20 + $90 = $153 AFTER: $153 From $396.25. 61% cheaper. Same work. The model gave you 24%. You gave yourself the other 37%. Anthropic's own number is 40% cheaper than Opus 5 on typical workloads. Your number depends on how much of your bill is output and how much of your prompt you're resending uncached every call. So check those two first. Bookmark this for when you migrate. follow CyrilXBT

CyrilXBT

15,803 views • 4 days ago

10 repos that cut your ai agent token bill by up to 80% 1. microsoft/LLMLingua → cuts prompt size by up to 95% compresses prompts before the api call. 20x compression. published at EMNLP + ACL. near-zero quality loss. 6,100 stars 2. mem0ai/mem0 → replaces full conversation history in context stores what matters. retrieves only what's needed. 10,000 token history → 200 token memory. per agent. 54,800 stars 3. BerriAI/litellm → routes each call to the cheapest model simple task → haiku. complex task → sonnet. tracks cost per agent, per call, per day. 45,700 stars 4. run-llama/llama_index → replaces sending full documents rag: 100-page doc → 3 relevant chunks → same answer. 98% fewer tokens per query. 49,100 stars 5. chroma-core/chroma → replaces keyword search in full context vector store. finds the closest match. feeds only that. 50-200 tokens per query instead of thousands. 27,800 stars 6. letta-ai/letta → replaces infinite context window crashes paged memory for agents. loads only relevant memory. stops your agent from hitting limits and retrying. 22,400 stars 7. guidance-ai/guidance → cuts output token bloat by 30-50% structured generation. constrains model output natively. no more 100-token prompts to get json back. 21,400 stars 8. Aider-AI/aider → replaces pasting entire codebases builds a repo map. sends only files relevant to the task. not your whole project. just what the agent needs. 44,300 stars 9. openai/tiktoken → count tokens before you send know the exact cost before the api call happens. not after the bill arrives. 18,100 stars 10. simonw/ttok → hard cap on what gets sent cli tool: count tokens, truncate to budget limit. pipe any text in. get truncated output back. 389 stars most agents are expensive not because the model is expensive. because nobody checked what was being sent to it.

self.dll

39,554 views • 4 months ago

This guy replaced an $8,000 survey crew with one drone and a pipeline on Claude Code that digitizes a whole site in a single trip. Inside it is not one drone but a whole pipeline of 5 modules on Claude Code each with its own job all answering to a single orchestrator. And he built them himself with no team. He draws one line on the map from the controller and the whole site starts turning into data at once. Pilot flies the DJI Matrice 350 RTK along the route and reaches where a crew drags tripods: the parking lot and the roofs and the grading. Scanner on the Zenmuse L2 records the geometry with a laser down to the centimeter. Builder in DJI Terra stitches the point cloud into a digital twin. Analyst on Claude builds what the client needs out of the finished model: a volume report and a progress diff and a tour behind a link. Mobile lives in his iPhone and hands the developer a link to the model while he drives to the next site. The site is ready as a file in about an hour and the developer rotates it in the browser himself and measures distances. No crew. No tripods. No week on site. Just him and a pickup and a drone and one API key. The whole processing pipeline lives in a folder at /Users/dev/site-capture. But he did not stop at a one time scan. The pipeline raises him by voice only when a flight misses a patch or when a sag on site goes past tolerance. And this is not made up: the DJI Matrice 350 RTK and the Zenmuse L2 are off the shelf enterprise hardware. The specs are one search away and the L2 really does write around 240,000 points per second. A crew for the same job charges $8,000 and three trips plus a week of desk work. His cost is tokens and subscriptions: one battery charge per flight and around $300 a month to host the finished models. In the end he draws one line on the map and an operator that does not exist scans the site and calculates the volumes and builds the report while he never leaves his truck. There is a huge market of everyone who needs to measure construction sites and warehouses and roofs and parking lots over and over while they still send out a photographer with a camera and wait a week. And he built this whole pipeline himself: one drone and one scanner and Claude Code that turns a flight into a file the developer pays for.

Blaze

10,772 views • 3 months ago