Загрузка видео...

Не удалось загрузить видео

На главную

Anthropic shipped 125 settings for Claude The official docs cover 40 One developer found the other 85 and his API bill dropped from $340 to $87 - not by using a cheaper model - not by writing shorter prompts just by moving one line in a config file to...

379,914 просмотров • 2 месяцев назад •via X (Twitter)

Комментарии: 0

Нет доступных комментариев

Здесь появятся комментарии из оригинального поста

Похожие видео

10 repos that cut your ai agent token bill by up to 80% 1. microsoft/LLMLingua → cuts prompt size by up to 95% compresses prompts before the api call. 20x compression. published at EMNLP + ACL. near-zero quality loss. 6,100 stars 2. mem0ai/mem0 → replaces full conversation history in context stores what matters. retrieves only what's needed. 10,000 token history → 200 token memory. per agent. 54,800 stars 3. BerriAI/litellm → routes each call to the cheapest model simple task → haiku. complex task → sonnet. tracks cost per agent, per call, per day. 45,700 stars 4. run-llama/llama_index → replaces sending full documents rag: 100-page doc → 3 relevant chunks → same answer. 98% fewer tokens per query. 49,100 stars 5. chroma-core/chroma → replaces keyword search in full context vector store. finds the closest match. feeds only that. 50-200 tokens per query instead of thousands. 27,800 stars 6. letta-ai/letta → replaces infinite context window crashes paged memory for agents. loads only relevant memory. stops your agent from hitting limits and retrying. 22,400 stars 7. guidance-ai/guidance → cuts output token bloat by 30-50% structured generation. constrains model output natively. no more 100-token prompts to get json back. 21,400 stars 8. Aider-AI/aider → replaces pasting entire codebases builds a repo map. sends only files relevant to the task. not your whole project. just what the agent needs. 44,300 stars 9. openai/tiktoken → count tokens before you send know the exact cost before the api call happens. not after the bill arrives. 18,100 stars 10. simonw/ttok → hard cap on what gets sent cli tool: count tokens, truncate to budget limit. pipe any text in. get truncated output back. 389 stars most agents are expensive not because the model is expensive. because nobody checked what was being sent to it.

self.dll

39,475 просмотров • 2 месяцев назад

This guy replaced an $8,000 survey crew with one drone and a pipeline on Claude Code that digitizes a whole site in a single trip. Inside it is not one drone but a whole pipeline of 5 modules on Claude Code each with its own job all answering to a single orchestrator. And he built them himself with no team. He draws one line on the map from the controller and the whole site starts turning into data at once. Pilot flies the DJI Matrice 350 RTK along the route and reaches where a crew drags tripods: the parking lot and the roofs and the grading. Scanner on the Zenmuse L2 records the geometry with a laser down to the centimeter. Builder in DJI Terra stitches the point cloud into a digital twin. Analyst on Claude builds what the client needs out of the finished model: a volume report and a progress diff and a tour behind a link. Mobile lives in his iPhone and hands the developer a link to the model while he drives to the next site. The site is ready as a file in about an hour and the developer rotates it in the browser himself and measures distances. No crew. No tripods. No week on site. Just him and a pickup and a drone and one API key. The whole processing pipeline lives in a folder at /Users/dev/site-capture. But he did not stop at a one time scan. The pipeline raises him by voice only when a flight misses a patch or when a sag on site goes past tolerance. And this is not made up: the DJI Matrice 350 RTK and the Zenmuse L2 are off the shelf enterprise hardware. The specs are one search away and the L2 really does write around 240,000 points per second. A crew for the same job charges $8,000 and three trips plus a week of desk work. His cost is tokens and subscriptions: one battery charge per flight and around $300 a month to host the finished models. In the end he draws one line on the map and an operator that does not exist scans the site and calculates the volumes and builds the report while he never leaves his truck. There is a huge market of everyone who needs to measure construction sites and warehouses and roofs and parking lots over and over while they still send out a photographer with a camera and wait a week. And he built this whole pipeline himself: one drone and one scanner and Claude Code that turns a flight into a file the developer pays for.

Blaze

10,681 просмотров • 1 месяц назад

andrej karpathy spent two hours teaching one thing: tokens are the atom of llms. tokenization is at the heart of every llm weirdness you've ever debugged. [watch the 15-min clip below. then run the 7-day playbook] ↓ save this before everyone copies it learn how the tokenizer works. understand how your llm actually consumes input. then run the engineering roadmap that took one production agent from $4,800/mo to $620/mo in 7 days. 87% reduction. no model swap. no framework migration. no quality drop on the eval set. token cost in 2026 is an engineering discipline. every line of your system prompt is rent you pay forever. what was eating the budget: → a single forgotten cron job ate 47% of one team's bill. they turned it off on a tuesday and the bill dropped before they wrote any optimization code. → anthropic ships a 90% discount on cache reads. one config line, cache_control ephemeral, break-even after one hit. most teams cache the volatile parts of the prompt and watch their hit rate sit at 12%. → one production agent went from 14,500 tokens of context overhead per turn to 850. a 94% drop. output quality held within 2% of the uncompressed baseline. → 60% of agent calls are haiku-tier work running on opus rates. classify the task first. pick the model second. → retry loops are the silent killer. no MAX_STEPS bound, one bad search query, $14 burned in a single session. one team traced 38% of their bill to this single pattern. karpathy gave you the atom. the playbook below gives you the harness. watch the lecture. read the playbook ↓

Rohit

73,258 просмотров • 2 месяцев назад

Someone ran Claude Code on a beach where any device overheats and that spot suddenly turned out to be the best home for the most powerful AI in the world. This is the reMarkable Paper Pro. A paper tablet for notes with no browser and no social media and not a single app. He sat down right on the sand in the open sun and brought up Claude Code on Opus 4.6 over the Claude API on the paper screen and opened his project ~/repos/webs while the waves broke a few steps away. For years every device had the same trouble outside. In direct sun the screen glares and washes out and heats up and instead of your work you see your own reflection. But e-ink does not blast its own light into your face. It reflects the sunlight like the page of a book. And here is what came out of it. The very thing that kills any normal screen outside turned into fuel for this one. The brighter the sun the sharper the picture because it has nothing to glare with and nothing to wash out. And then comes the thing no laptop on a beach will give you. Your eyes do not get tired. You can watch Opus think on max effort for an hour and it reads like a book in the sun and not a backlight you squint into. The picture only comes alive. In bright light it does not fade but turns sharper and higher in contrast than it ever was in a room. The charge lasts for days. E-ink barely touches the battery so there is no outlet anywhere on the sand and the tablet does not care. It weighs as much as a notebook. The whole setup folds into a beach bag like a pad with a pen on top. Everything on the screen is for real. Claude Code v2.1.110 and Opus 4.6 on the Claude API and the project ~/repos/webs open right on the e-ink in the middle of the sand. In my opinion this is the most unexpected home for an AI this year. Not an office with the blinds drawn and not a monitor cranked to full brightness but a quiet sheet of paper on the sand that open sun only makes better and on it the most powerful Claude writes code right on the page like a pen.

Blaze

89,297 просмотров • 1 месяц назад

THIS GUY AUDITED 926 CLAUDE CODE SESSIONS AND FOUND MOST OF THE TOKEN WASTE WAS ON HIS SIDE everyone is blaming anthropic for the limits, so he decided to actually look at the data 858 sessions, 18,903 turns, and $1,619 estimated spend across 33 days here's what he found: 1\ one default setting was burning 14,000 tokens per turn Claude Code loads the full JSON schema for every tool into context at session start. whether you use them or not. 20,000 tokens of tool definitions sitting there on every single turn. the fix: one line in your settings.json "ENABLE_TOOL_SEARCH": "true" context dropped from 45K to 20K instantly. across 858 sessions that one setting was wasting an estimated 264 million tokens 2\ cache expiry is the single biggest waste 54% of his turns came after a 5+ minute idle gap. every one of those turns re-processed the entire conversation at full price which caused a 10x cost jump you go grab coffee. come back 5 minutes later. type your next message. everything rebuilds from scratch. the context didn't change. you didn't change. the cache just expired. 12.3 million tokens wasted on idle gaps alone 3\ 42 skills loaded. 19 of them used twice or less across 858 sessions. every one of those skill schemas sat in context on every turn eating tokens for nothing. 4\ 1,122 redundant file reads where the same file was read 3+ times one session read the same file 33 times. he ALSO built a full token auditor dashboard that shows you exactly where your waste is coming from 19 charts, opens in your browser, free AND open source

Om Patel

298,746 просмотров • 3 месяцев назад