Video wird geladen...

Video konnte nicht geladen werden

Zur Startseite

@TryTrustAI is building the fastest, cheapest lightweight inference by optimizing KV cache at Y Combinator. Our group of MIT grads and olympiad medalists from Google DeepMind, Jane Street, and NeurIPS/ICML publishers are setting a new standard for cost and latency efficiency.

33,169 Aufrufe • vor 1 Monat •via X (Twitter)

36 Kommentare

Profilbild von Claire Mao
Claire Maovor 1 Monat

@TryTrustAI @ycombinator @GoogleDeepMind @JaneStreetGroup @medha_rv WOWW so impressive congrats!!

Profilbild von Justin Hong
Justin Hongvor 1 Monat

@TryTrustAI @ycombinator @GoogleDeepMind @JaneStreetGroup @medha_rv been waiting for this

Profilbild von Hannah Chung
Hannah Chungvor 1 Monat

@TryTrustAI @ycombinator @GoogleDeepMind @JaneStreetGroup @medha_rv me too 🥲

Profilbild von medha
medhavor 1 Monat

@TryTrustAI @ycombinator @GoogleDeepMind @JaneStreetGroup Exciting!

Profilbild von Hannah Chung
Hannah Chungvor 1 Monat

@TryTrustAI @ycombinator @GoogleDeepMind @JaneStreetGroup I hope you're hyped

Profilbild von Jeremy
Jeremyvor 1 Monat

@TryTrustAI @ycombinator @GoogleDeepMind @JaneStreetGroup @medha_rv So gmi 🚀

Profilbild von Nour Zahzah
Nour Zahzahvor 1 Monat

@TryTrustAI @ycombinator @GoogleDeepMind @JaneStreetGroup @medha_rv Let’s go

Profilbild von Shaurya Aggarwal
Shaurya Aggarwalvor 1 Monat

@TryTrustAI @ycombinator @GoogleDeepMind @JaneStreetGroup @medha_rv cracked team! congrats guys

Profilbild von Hannah Chung
Hannah Chungvor 1 Monat

@TryTrustAI @ycombinator @GoogleDeepMind @JaneStreetGroup @medha_rv thank you!!!

Profilbild von aryan mahajan
aryan mahajanvor 1 Monat

@TryTrustAI @ycombinator @GoogleDeepMind @JaneStreetGroup @medha_rv 🙌

Profilbild von Hannah Chung
Hannah Chungvor 1 Monat

@TryTrustAI @ycombinator @GoogleDeepMind @JaneStreetGroup @medha_rv 🙏

Profilbild von Sumit Singh🇮🇳
Sumit Singh🇮🇳vor 1 Monat

@TryTrustAI @ycombinator @GoogleDeepMind @JaneStreetGroup @medha_rv Soo great 🤝👨🏻‍💻🌐 looking to know more . Let's connect

Profilbild von Nikhil Gupta
Nikhil Guptavor 1 Monat

@TryTrustAI @ycombinator @GoogleDeepMind @JaneStreetGroup @medha_rv wait IPhO camp does not the mean medalist

Profilbild von Rick A.F.
Rick A.F.vor 1 Monat

@TryTrustAI @ycombinator @GoogleDeepMind @JaneStreetGroup @medha_rv Amazing, cache hits different when its optimised! Lets go, lets improve the efficiency of interference!

Profilbild von Akira
Akiravor 1 Monat

@TryTrustAI @ycombinator @GoogleDeepMind @JaneStreetGroup @medha_rv Nice

Profilbild von Ruiyang Chen
Ruiyang Chenvor 1 Monat

@TryTrustAI @ycombinator @GoogleDeepMind @JaneStreetGroup @medha_rv Curious what’s the use case?

Profilbild von Evan Dickinson
Evan Dickinsonvor 1 Monat

@TryTrustAI @ycombinator @GoogleDeepMind @JaneStreetGroup @medha_rv Mit!

Profilbild von Luke Gainey
Luke Gaineyvor 1 Monat

@TryTrustAI @ycombinator @GoogleDeepMind @JaneStreetGroup @medha_rv FANTASTIC

Profilbild von Dheeraj Kaim
Dheeraj Kaimvor 1 Monat

@TryTrustAI @ycombinator @GoogleDeepMind @JaneStreetGroup @medha_rv congrats, hope it works out for you guys

Profilbild von An Engineer's Log
An Engineer's Logvor 1 Monat

@TryTrustAI @ycombinator @GoogleDeepMind @JaneStreetGroup @medha_rv what's the largest model you can host, and the price per million token for it ?

Profilbild von Rahul Wagh
Rahul Waghvor 1 Monat

@TryTrustAI @ycombinator @GoogleDeepMind @JaneStreetGroup @medha_rv interesting!!

Profilbild von om_lathiya_
om_lathiya_vor 1 Monat

@TryTrustAI @ycombinator @GoogleDeepMind @JaneStreetGroup @medha_rv Interesting

Profilbild von Evï Anton
Evï Antonvor 1 Monat

@TryTrustAI @ycombinator @GoogleDeepMind @JaneStreetGroup @medha_rv is this any different from

Profilbild von Niraj Adhikari
Niraj Adhikarivor 1 Monat

@TryTrustAI @ycombinator @GoogleDeepMind @JaneStreetGroup @medha_rv Looks good!

Profilbild von just a monolith
just a monolithvor 1 Monat

@TryTrustAI @ycombinator @GoogleDeepMind @JaneStreetGroup @medha_rv I have nothing to add. But goddamn, you're cute

Profilbild von Marco H
Marco Hvor 1 Monat

@TryTrustAI @ycombinator @GoogleDeepMind @JaneStreetGroup @medha_rv So interesting.. i was just running research on this as I just heard KV cache for first time today and algorithm is showing it to me now.

Profilbild von Yoonseok Yang
Yoonseok Yangvor 1 Monat

@TryTrustAI @ycombinator @GoogleDeepMind @JaneStreetGroup @medha_rv lets go hannah!!

Profilbild von Santz
Santzvor 1 Monat

@TryTrustAI @ycombinator @GoogleDeepMind @JaneStreetGroup @medha_rv Congrats on the launch!

Profilbild von John
Johnvor 1 Monat

@TryTrustAI @ycombinator @GoogleDeepMind @JaneStreetGroup @medha_rv Dope

Profilbild von AkinAyo | UI/UX Designer
AkinAyo | UI/UX Designervor 1 Monat

@TryTrustAI @ycombinator @GoogleDeepMind @JaneStreetGroup @medha_rv Impressive

Profilbild von 𝕐𝕖𝕤𝕙𝕦𝕒
𝕐𝕖𝕤𝕙𝕦𝕒vor 1 Monat

@TryTrustAI @ycombinator @GoogleDeepMind @JaneStreetGroup @medha_rv Great team 👏🏻💯

Profilbild von E
Evor 1 Monat

@TryTrustAI @ycombinator @GoogleDeepMind @JaneStreetGroup @medha_rv Bubble

Profilbild von mr ceo
mr ceovor 1 Monat

@TryTrustAI @ycombinator @GoogleDeepMind @JaneStreetGroup @medha_rv Good luck

Profilbild von Chandan kumar
Chandan kumarvor 1 Monat

@TryTrustAI @ycombinator @GoogleDeepMind @JaneStreetGroup @medha_rv Hii

Profilbild von Lokesh
Lokeshvor 26 Tagen

@TryTrustAI @ycombinator @GoogleDeepMind @JaneStreetGroup @medha_rv Any opening role in product management with visa sponsorship?

Profilbild von Amritesh Anand
Amritesh Anandvor 1 Monat

@TryTrustAI @ycombinator @GoogleDeepMind @JaneStreetGroup @medha_rv Is this something we can try out on our GPU's or you serve endpoints to end clients?

Ähnliche Videos

The Cost of Intelligence is Heading to Zero | Hyperspace P2P Distributed Cache We present to you our breakthrough cross-domain work across AI, distributed systems, cryptography, game theory to solve the primary structural inefficiency at the heart of AI infrastructure: most inference is redundant. Google has reported that only 15% of daily searches are truly novel. The rest are repeats or close variants. LLM inference inherits this same power-law distribution. Enterprise chatbots see 70-80% of queries fall into a handful of intent categories. System prompts are identical across 100% of requests within an application. The KV attention state for "You are a helpful assistant" has been computed billions of times, on millions of GPUs, identically. And yet every AI lab, every startup, every self-hosted deployment - computes and caches these results independently. There is no shared layer. No global memory. Every provider pays the full compute cost for every query, even when the answer already exists somewhere in the network. This is the problem Hyperspace solves where distributed cache operates at three levels, each catching a different class of redundancy: 1. Response cache Same prompt, same model, same parameters - instant cached response from any node in the network. SHA-256 hash lookup via DHT, with cryptographic cache proofs linking every response to its original inference execution. No trust required. Fetchers re-announce as providers, so popular responses replicate naturally across more nodes. 2. KV prefix cache Same system prompt tokens - skip the most expensive part of inference entirely. Prefill (computing Key-Value attention states) is deterministic: same model plus same tokens always produces identical KV state. The network caches these states using erasure coding and distributes them via the routing network. New questions that share a common prefix resume generation from cached state instead of recomputing from scratch. 3. Routing to cached nodes Instead of transferring KV state across the network for every request, Hyperspace routes the request to the node that already has the state loaded in VRAM. The request goes to the cache, not the cache to the request. Together, these three layers mean that 70-90% of inference requests at network scale never require full GPU computation. This work doesn't exist in isolation. It builds on research from across the industry: SGLang's RadixAttention demonstrated that automatic prefix sharing can yield up to 5x speedup on structured LLM workloads. Moonshot AI's Mooncake built an entire KV-cache-centric disaggregated architecture for production serving at Kimi. Anthropic, OpenAI, and Google all launched prompt caching products in 2024 - priced at 50-90% discounts - because system prompt reuse is so pervasive that it changes the economics of inference. What all of these systems share is a common limitation: they operate within a single organization's infrastructure. SGLang caches prefixes within one server. Mooncake disaggregates KV cache within one datacenter. Anthropic's prompt caching works within one API provider's fleet. None of them can share cached state across organizational boundaries. Hyperspace removes this boundary. The cache is global. A response computed by a node in Tokyo is immediately available to a node in Berlin. A KV prefix state generated for Qwen-32B on one machine is verifiable and reusable by any other machine running the same model. The routing network provides the delivery guarantees, the erasure coding provides the redundancy, and the cache proofs provide the trust. What this means for the cost of intelligence Big AI labs scale linearly: twice the users means twice the GPU spend. Every query is a cost center. Their internal caching helps, but it's siloed - Lab A's cache can't serve Lab B's users, and neither can serve a self-hosted Llama deployment. Hyperspace scales sub-linearly. Every new node that joins the network adds to the global cache. Every inference result enriches the cache for all future requests. The cache hit rate rises with network size because query distributions follow a power law - the most common questions are asked exponentially more often than rare ones. The implication is simple: as the network grows, the effective cost per inference drops. Not linearly. Logarithmically. At 10 million nodes, we estimate 75-90% of all inference requests can be served from cache, eliminating 400,000+ MWh of energy consumption per year and avoiding over 200,000 tons of CO2 emissions. The first person to ask a question pays the compute cost. Everyone after them gets the answer for free, with cryptographic proof that it's authentic. Training is competitive. Inference is shared Open-weight models are converging on quality with closed models. Labs will continue to differentiate on training - data curation, architecture innovation, RLHF tuning. That's where the real intellectual property lives. But inference is a commodity. Two copies of Qwen-32B running the same prompt produce the same KV state and the same response, byte for byte, regardless of whose GPU runs the matrix multiplication. There is no moat in multiplying matrices. The moat is in training the weights. A global distributed cache makes this separation explicit. It doesn't matter who trained the model. Once the weights are open, the inference cost approaches zero at scale - because the network remembers every answer and can prove it's correct. No lab, no matter how well-funded, can match this. They cannot share caches across competitors. They scale linearly. The network scales logarithmically. The marginal cost of intelligence approaches zero. That's the endgame.

Varun

37,555 Aufrufe • vor 6 Monaten