Video yükleniyor...

Video Yüklenemedi

Ana Sayfaya Dön

Redis built a cache that cuts LLM costs by 70%! Production LLM apps often receive different versions of the same question. For instance, an internal developer assistant might receive: - "How do I rotate an API key?" - "Where can I replace my API key?" The wording is different,...

85,616 görüntüleme • 2 gün önce •via X (Twitter)

17 Yorum

Simo | ai-costguard profil fotoğrafı
Simo | ai-costguard2 gün önce

semantic cache intercepts before the LLM call, prefix cache reuses computation inside it, the 70% cost reduction comes from skipping the inference step entirely, not cheapening it

_alphashark_ profil fotoğrafı
_alphashark_2 gün önce

I'd worry rotate and revoke are close enough to false-match on a threshold. Your cache could hand back the wrong answer with confidence.

AI Apps API profil fotoğrafı
AI Apps API2 gün önce

Semantic caching works right up until the threshold has to cover two questions that sound alike but are not. "How do I rotate an API key" and "how do I revoke an API key" sit very close in embedding space and have different answers. Set the threshold loose enough to catch real paraphrases and you start serving confidently wrong cached answers, which is worse than a cache miss because nothing in the response signals it is stale. The fix that held up for us was scoping the cache by intent or namespace first, then letting similarity decide inside that bucket. Much smaller blast radius when the threshold is wrong.

Hussain Hashim | Building SundayBack profil fotoğrafı
Hussain Hashim | Building SundayBack2 gün önce

@akshay_pachaar Caching's a game changer for reusing responses. Just make sure to handle context-sensitive queries separately so you don't serve outdated info.

Anurag pm profil fotoğrafı
Anurag pm2 gün önce

the 70% cut is great. the other 30% is me asking 'are you sure' after every answer.

T profil fotoğrafı
T2 gün önce

semantic caches are powerful right up until 'similar meaning' quietly becomes 'wrong customer'

Maya Renner profil fotoğrafı
Maya Renner2 gün önce

The cache-hit demo hides the scary bit: invalidation. Rotate the API key, change permissions, or update the source doc and that “70% cheaper” answer can quietly become wrong. Semantic caching needs a freshness boundary, not just similarity.

Buswe profil fotoğrafı
Buswe2 gün önce

Worth pinning the embedding model version with the cache and rebuilding on upgrade, so stored answers keep matching the right questions.

Sebastiano Mandalà profil fotoğrafı
Sebastiano Mandalà2 gün önce

ask Jev if the sentence is close enough with something cached

Oleg profil fotoğrafı
Oleg2 gün önce

70% cheaper right up to the tuesday the refund policy changes

Yancy Maxwell Hayes profil fotoğrafı
Yancy Maxwell Hayes2 gün önce

The risk to design around is the false hit: two questions that embed close but differ in a constraint, last week versus last month. The threshold that maximizes hit rate is not the safe one, so log near misses and tune on that distribution rather than on the hit rate itself.

Badanzer profil fotoğrafı
Badanzer2 gün önce

70% cost cut from semantic cache is the boring win production teams actually ship. Curious how you handle cache hits when the prompt must stay on-prem though.

Salise profil fotoğrafı
Salise2 gün önce

the idea of a semantic cache storing question-response pairs is really smart. could save so much time and money

Siddharth Sharma profil fotoğrafı
Siddharth Sharma2 gün önce

@grok what are the configuration options? In what cases will this be helpful since in a most of cases the cache may grow really big.

Wallchain Community Hub profil fotoğrafı
Wallchain Community Hub2 gün önce

saves so much on api calls

shaun profil fotoğrafı
shaun2 gün önce

70% holds on a faq bot. an agent loop almost never asks the same question twice

Harley Lewis Foote profil fotoğrafı
Harley Lewis Foote2 gün önce

Deciding two prompts match is a security call, not just a cost saving.

Benzer Videolar

Redis built a cache that cuts LLM costs by 90%! Production LLM apps do not receive completely new questions every time. A customer-support assistant might receive all three of these: - "Can I get a refund after buying the monthly plan?" - "Is the monthly subscription refundable?" - "Can I cancel the plan and get my money back?" The wording is different, but the underlying question and its answer remain the same. Yet LLM apps process every version as a new request. They assemble the prompt, send it to the model, and generate an answer that may have already been generated. Prefix caching reduces part of these repeated calls. When requests begin with the same system prompt or context, the model can reuse the KV states already computed for that shared prefix. But the request still hits the LLM. The new tokens must be processed, and the complete answer must still be decoded. So even with a prefix-cache hit, there's another generation call involved. To solve this, instead of only caching computation inside the model, the application can cache the generated response outside it. When another question arrives, the system embeds it and compares it with previously answered questions. If it finds a sufficiently close match, it returns the stored response without invoking the LLM again. A cache hit removes the input tokens, output tokens, and decoding time associated with another LLM call. In practice, it is important to decide which questions can safely share an answer since a production setup needs well-tuned similarity thresholds, expiration policies, data isolation, and monitoring for incorrect matches. If you want to use this in practice, Redis already implements it as a managed service called Redis LangCache. Under the hood, it generates embeddings, searches previous responses, and returns a matching answer before another model call occurs. Redis also handles access scopes, custom filtering, TTL and eviction controls, and cache monitoring through Redis Cloud. I built an interface to compare it against direct LLM inference. The video below shows this in action, and I worked with Redis on this post to put this together. For the paraphrased question in my run, direct inference took 2.232 seconds and consumed 514 input tokens plus 250 output tokens. Redis returned the earlier response in 0.37 seconds with zero LLM input or output tokens. That was roughly 6x faster in this run. Redis reports API cost savings of up to 90% and cache-hit responses up to 15x faster. The actual result depends on how much safe repetition exists in the workload. You can try Redis LangCache here: If you want to dive deeper, I have already written a detailed breakdown of KV, prefix, prompt, and semantic caching in the article quoted below. This demo builds on the final technique and shows it running in practice. Read it below.

Avi Chawla

165,441 görüntüleme • 15 gün önce

Researchers made LLM inference 14x faster and 90% cheaper. The video below depicts the speed up in action. Providers discount cached input tokens by as much as 90% because a cache hit skips prefill compute entirely. For stable system prompts and tool definitions, hit rates of 60 to 85% are achievable, which makes it the highest-leverage inference optimization. But the cost saving only works when the cached text is an exact, byte-for-byte prefix of the new request. If you change one character anywhere before it, the entire cached region is missed. Three common request patterns produce full cache misses: - A query that needs documents A and B together can't reuse B's standalone cache, because those KV entries were computed without A in front of them. - The same three documents retrieved in a different order produce a full cache miss, even though nothing about the documents changed. - In multi-turn conversations, every new turn invalidates whatever was cached beyond the stable prefix. Alibaba's production data did a study on this and found that just 10% of cached KV blocks serve 77% of all cache hits. So most of what gets cached sits in storage and is never used a single time. And the root cause is that KV entries are position-dependent. Each token's KV encodes attention to everything before it, so a cached block is only valid in the exact context it was computed in. There's a second, less discussed problem as well. Cache management runs inside the inference engine's process. Moving KV tensors between GPU, CPU, and disk competes with inference for the same resources. This is why Google's TurboQuant compresses KV caches to 3 bits with no accuracy loss and still causes a 20%+ slowdown when it runs in-process. Fixing both problems means restructuring where caching lives. Cache management moves into its own process, the engine only exchanges block IDs over shared GPU memory, and heavy data movement runs across GPU, CPU, disk, and remote storage in parallel. Non-prefix reuse gets handled by selectively recomputing only the small set of tokens that attend across document boundaries. LMCache is the open-source project (10k+ stars) that implements this exact architecture, and it plugs into vLLM, SGLang, and TensorRT-LLM. The selective recomputation part is implemented in its CacheBlend technique, which makes cached docs in any order and combination, with 2-4x faster multi-document processing. On H200s running Qwen3-235B with 50 concurrent users, LMCache's multiprocess mode delivers 14x faster time-to-first-token and 4x faster decoding compared to in-process caching. GitHub repo: (don't forget to star 🌟) My co-founder wrote a full breakdown of KV cache management. It covers the disaggregated architecture behind the 14x speed up, how CacheBlend preserves generation quality while skipping recomputation, and how to turn every document in a knowledge base into a reusable cached asset. Read it below.

Avi Chawla

30,717 görüntüleme • 2 ay önce

This is a standard practice for almost all Tier-1 banking applications in Nigeria, and for some fintech applications I’ve previously performed pentests on. Client-side encryption isn’t a total waste, or a waste of compute, as some people have claimed, but rather a measure to protect against API tampering or API request/response manipulation between the client and the server when implemented properly. Even with HTTPS, attackers can capture a decrypted version of web or mobile API data in transit because the browser and the server establish a level of trust during the TLS handshake. Attackers can leverage this trust to capture & proxy already-decrypted traffic, tamper with it, and then forward it to the server. This allows them to override what the user interface or client is originally supposed to send and replace it with data of their choosing. That is why validation needs to be performed on both the client and the server side. To wrap up, encrypting API requests and responses makes it significantly harder for attackers to tamper with data, even if they capture the traffic, unless they have access to the encryption details (algorithm, encryption mode, key size, secret key, and initialization vector), assuming asymmetric encryption is used. In the demo below, you can see how I discovered additional parameters (balance, is_admin) in the API response, captured the registration API request, despite it being sent over HTTPS from the interface, added the discovered parameters, and successfully inflated my balance to 50 billion and also escalated my privileges to admin, and ultimately deleted the accounts of two live users/customers. In the second slide, I captured an API traffic of a bank app, and you can see how difficult the payloads are to read.

Ghost St Badmus

217,804 görüntüleme • 9 ay önce

I built an agent that answers machine-learning questions. It's autonomous, and the best part is that I built the whole thing without writing a single line of Python code. Here is what I did and how I did it: Over a year ago, a friend and I built a site that publishes multi-choice questions. You get a new one every day. I decided to have GPT-3.5 answer questions. Here is what I needed to build: 1. Connect to the site's API to retrieve today's question 2. Extract the question and the potential choices 3. Connect to OpenAI's API and ask GPT-3.5 to answer the question 4. Parse the answer from the model 5. Submit the answer back to the API to get the score Not difficult. Likely several hours of work. But I didn't have to write any code. I built the whole thing by dragging and dropping components using Vellum is a YC-backed platform for developers to build LLM applications. They are the only ones I've seen offering this functionality. They sponsored this post, and their team helped me with all my questions while I built this. I created a workflow. The platform supports several node types to build whatever you have in mind. I show how I put the whole thing together in the attached video. The only code I had to write was a few lines of Jinja to parse and transform the API and the LLM results. There are three lessons I want to share from this experience: First, the best possible code is the one you didn't write. I'm a big fan of no-code tools because they help me materialize my ideas fast. They help product people, designers, and no coders collaborate on the solution. Second, Large Language Models are sensitive to how you prompt them. Small changes to prompts can make a big difference in results. This is more pronounced when you are building a multi-step workflow. Third, automated testing and evaluation for prompts is critical. There aren't many companies thinking about this. They'll have a hard time moving from a demo phase. The attached video will show you what I did.

Santiago

309,825 görüntüleme • 3 yıl önce

Anthropic won't like this open-source repo. It is going to cost LLM providers a lot of money. Every CI run of an AI app today sends real requests to providers like OpenAI or Anthropic. Like any other LLM call, this too gets billed at actual API rates. So for teams with high commit volumes, this accumulates into a meaningful chunk of API spend. One common hack devs use is that instead of invoking the LLM API, the test calls a fake local server that speaks the same API and returns a dummy response. The catch is that the dummy response is a copy of what the provider returned on the day it was saved, and providers keep adding fields and changing types. So the tests keep passing against a schema that's no longer valid, while the real integration breaks in production. A smart approach is now actually implemented in CopilotKit🪁's recently open-sourced aimock project. Every day, the repo's own CI sends a handful of requests to the real API and the same requests to the fake server, then compares both against the official client library's type definitions. Those are the only real API calls in the whole setup, and they run on the repo's own keys, not in anyone else's CI. A single team can push hundreds of commits a day, and thousands of teams are already doing that with coding agents. All of those runs stay offline, because one repo checks against the real API on everyone's behalf. When a check fails, a coding agent updates aimock's built-in response schema, the full test suite has to pass, and a patch version ships to npm. By simply upgrading the package, the corrected schema gets reflected in every project using it. The capability is not just limited to a single provider. The same server works for Claude, OpenAI, Gemini, Bedrock, Azure, Ollama, plus MCP tools, A2A agents, AG-UI event streams, vector DBs like Pinecone and Qdrant, and search, speech, image, and video endpoints. Here's the repo: (don't forget to star it ⭐) That said, mocking your API calls is one thing. AI engineers should also know how to test agents properly in the first place, which several teams still skip. I wrote a full walkthrough on that, covering build, testing, evals, tracing, and deployment. Read it below.

Akshay 🚀

62,821 görüntüleme • 1 ay önce

Why is Redis Fast? Redis is fast for in-memory data storage. Its speed has made it popular for caching, session storage, and real-time analytics. But what gives Redis its blazing speed? Let's explore: RAM-Based Storage At its core, Redis primarily uses main memory for storing data. Accessing data from RAM is orders of magnitude faster than from disk. This is a major reason for Redis's speed. However, RAM is volatile. To persist data, Redis supports disk snapshots and append-only file logging. This combines RAM's performance with disk's permanence. There is a tradeoff though - recovery from disk is slow. If a Redis instance fails, restarting from disk can be slow compared to failing over to a replica instance fully in memory. So while Redis offers durability via disk, it comes at the cost of slower recovery. A better solution is Redis replication. With a synchronized replica kept in memory, failover is instant with no rehydration. This maintains speed and near-instant recovery. IO Multiplexing & Single-threaded Read/Write Redis uses an event-driven, single-threaded model for its core operations. A main event loop handles all client requests and data operations sequentially. This single-threaded execution avoids context switching and synchronization overhead typical of multi-threaded systems. Redis uses non-blocking I/O to handle multiple connections asynchronously. This allows it to support many client connections with very low overhead, Redis does leverage threading in certain areas: - Background tasks like taking snapshots. - I/O threads are used for certain operations. - Modules can use threads. - Since Redis 6.0, it supports multi-threaded I/O for network communication, improving performance on multi-core systems. Redis also uses pipelining for high throughput. Clients pipeline commands without waiting for each response. This allows more efficient network round trips, boosting overall performance. Efficient Data Structures Redis supports various optimized data structures, from linked lists, zip lists, and skip lists to sets, hashes, and sorted sets, among others. Each is carefully designed for specific use cases for quick and efficient data access. Over to you: With Redis now supporting some multi-threading, how should we configure it to fully utilize all the CPU cores of modern hardware when deploying in production? – Subscribe to our weekly newsletter to get a Free System Design PDF (158 pages):

Sahn Lam

46,910 görüntüleme • 2 yıl önce

New short course: LLMs as Operating Systems: Agent Memory, created with Letta, and taught by its founders Charles Packer and Sarah Wooders. An LLM's input context window has limited space. Using a longer input context also costs more and results in slower processing. So, managing what's stored in this context window is important. In the innovative paper MemGPT: Towards LLMs as Operating Systems, its authors (which include the instructors) proposed using an LLM agent to manage this context window. Their system uses a large persistent memory that stores everything that could be included in the input context, and an agent decides what is actually included. Take the example of building a chatbot that needs to remember what's been said earlier in a conversation (perhaps over many days of interaction with a user). As the conversation's length grows, the memory management agent will move information from the input context to a persistent searchable database; summarize information to keep relevant facts in the input context; and restore relevant conversation elements from further back in time. This allows a chatbot to keep what's currently most relevant in its input context memory to generate the next response. When I read the original MemGPT paper, I thought it was an innovative technique for handling memory for LLMs. The open-source Letta framework, which we'll use in this course, makes MemGPT easy to implement. It adds memory to your LLM agents and gives them transparent long-term memory. In detail, you’ll learn: - How to build an agent that can edit its own limited input context memory, using tools and multi-step reasoning - What is a memory hierarchy (an idea from computer operating systems, which use a cache to speed up memory access), and how these ideas apply to managing the LLM input context (where the input context window is a "cache" storing the most relevant information; and an agent decides what to move in and out of this to/from a larger persistent storage system) - How to implement multi-agent collaboration by letting different agents share blocks of memory This course will give you a sophisticated understanding of memory management for LLMs, which is important for chatbots having long conversations, and for complex agentic workflows. Please sign up here!

Andrew Ng

201,127 görüntüleme • 1 yıl önce