Loading video...

Video Failed to Load

Go Home

this is what a real AGI loop looks like: 6+ hours agent runtime 240M+ tokens 86%+ cache hit 500+ tool calls $125 spent = 1 output that would have taken an expert team months to complete and cost tens of thousands

58,330 views • 3 months ago •via X (Twitter)

0 Comments

No comments available

Comments from the original post will appear here

Related Videos

How to cut your AI bill by 60% switching to Opus 5.5: Most people will swap the model name, save 24%, and stop there. The other 37% is sitting in your settings. Here's the real before and after on an example agent workload. One month: 100M input tokens (80M cached reads, 5M cache writes, 15M uncached) 10M output tokens BEFORE: Opus 5 Cache reads: 80M × $0.50 = $40 Cache writes: 5M × $6.25 = $31.25 Uncached input: 15M × $5 = $75 Output: 10M × $25 = $250 Total: $396.25 STEP 1. Just switch the model. Same tokens, new prices. Cache reads $0.20. Writes $5. Input $4. Output $20. $16 + $25 + $60 + $200 = $301 24% cheaper. That's what everyone's screenshotting. STEP 2. Let it write less. Box measured Opus 5.5 at 40% less verbose with no drop in accuracy. 10M output tokens becomes 6M. Output drops from $200 to $120. Total: $221. Now you're at 44% off. STEP 3. Stop running every call at max effort. Thinking can't be switched off on 5.5 anymore, so effort is your lever. Classifying, routing, formatting, summarizing? Drop the effort. Save the high settings for the calls that actually reason. Say that trims output another 25%, 6M down to 4.5M. Output: $90. Total: $191. 52% off. STEP 4. Cache the stuff you keep resending. This is the one nobody does. Cache reads went from $0.50 to $0.20. That's 60% off the cheapest line on your bill. Move your system prompt, tool definitions and repo context into the cache. Uncached input drops from 15M to 5M, cache reads go up to 90M. $18 + $25 + $20 + $90 = $153 AFTER: $153 From $396.25. 61% cheaper. Same work. The model gave you 24%. You gave yourself the other 37%. Anthropic's own number is 40% cheaper than Opus 5 on typical workloads. Your number depends on how much of your bill is output and how much of your prompt you're resending uncached every call. So check those two first. Bookmark this for when you migrate. follow CyrilXBT

CyrilXBT

15,803 views • 3 days ago