Загрузка видео...

Не удалось загрузить видео

На главную

@TryTrustAI is building the fastest, cheapest lightweight inference by optimizing KV cache at Y Combinator. Our group of MIT grads and olympiad medalists from Google DeepMind, Jane Street, and NeurIPS/ICML publishers are setting a new standard for cost and latency efficiency.

33,169 просмотров • 25 дней назад •via X (Twitter)

Комментарии: 36

Фото профиля Claire Mao
Claire Mao25 дней назад

@TryTrustAI @ycombinator @GoogleDeepMind @JaneStreetGroup @medha_rv WOWW so impressive congrats!!

Фото профиля Justin Hong
Justin Hong25 дней назад

@TryTrustAI @ycombinator @GoogleDeepMind @JaneStreetGroup @medha_rv been waiting for this

Фото профиля Hannah Chung
Hannah Chung25 дней назад

@TryTrustAI @ycombinator @GoogleDeepMind @JaneStreetGroup @medha_rv me too 🥲

Фото профиля medha
medha25 дней назад

@TryTrustAI @ycombinator @GoogleDeepMind @JaneStreetGroup Exciting!

Фото профиля Hannah Chung
Hannah Chung25 дней назад

@TryTrustAI @ycombinator @GoogleDeepMind @JaneStreetGroup I hope you're hyped

Фото профиля Jeremy
Jeremy25 дней назад

@TryTrustAI @ycombinator @GoogleDeepMind @JaneStreetGroup @medha_rv So gmi 🚀

Фото профиля Nour Zahzah
Nour Zahzah25 дней назад

@TryTrustAI @ycombinator @GoogleDeepMind @JaneStreetGroup @medha_rv Let’s go

Фото профиля Shaurya Aggarwal
Shaurya Aggarwal25 дней назад

@TryTrustAI @ycombinator @GoogleDeepMind @JaneStreetGroup @medha_rv cracked team! congrats guys

Фото профиля Hannah Chung
Hannah Chung25 дней назад

@TryTrustAI @ycombinator @GoogleDeepMind @JaneStreetGroup @medha_rv thank you!!!

Фото профиля aryan mahajan
aryan mahajan25 дней назад

@TryTrustAI @ycombinator @GoogleDeepMind @JaneStreetGroup @medha_rv 🙌

Фото профиля Hannah Chung
Hannah Chung25 дней назад

@TryTrustAI @ycombinator @GoogleDeepMind @JaneStreetGroup @medha_rv 🙏

Фото профиля Sumit Singh🇮🇳
Sumit Singh🇮🇳25 дней назад

@TryTrustAI @ycombinator @GoogleDeepMind @JaneStreetGroup @medha_rv Soo great 🤝👨🏻‍💻🌐 looking to know more . Let's connect

Фото профиля Nikhil Gupta
Nikhil Gupta24 дней назад

@TryTrustAI @ycombinator @GoogleDeepMind @JaneStreetGroup @medha_rv wait IPhO camp does not the mean medalist

Фото профиля Rick A.F.
Rick A.F.25 дней назад

@TryTrustAI @ycombinator @GoogleDeepMind @JaneStreetGroup @medha_rv Amazing, cache hits different when its optimised! Lets go, lets improve the efficiency of interference!

Фото профиля Akira
Akira25 дней назад

@TryTrustAI @ycombinator @GoogleDeepMind @JaneStreetGroup @medha_rv Nice

Фото профиля Ruiyang Chen
Ruiyang Chen25 дней назад

@TryTrustAI @ycombinator @GoogleDeepMind @JaneStreetGroup @medha_rv Curious what’s the use case?

Фото профиля Evan Dickinson
Evan Dickinson25 дней назад

@TryTrustAI @ycombinator @GoogleDeepMind @JaneStreetGroup @medha_rv Mit!

Фото профиля Luke Gainey
Luke Gainey24 дней назад

@TryTrustAI @ycombinator @GoogleDeepMind @JaneStreetGroup @medha_rv FANTASTIC

Фото профиля Dheeraj Kaim
Dheeraj Kaim25 дней назад

@TryTrustAI @ycombinator @GoogleDeepMind @JaneStreetGroup @medha_rv congrats, hope it works out for you guys

Фото профиля An Engineer's Log
An Engineer's Log25 дней назад

@TryTrustAI @ycombinator @GoogleDeepMind @JaneStreetGroup @medha_rv what's the largest model you can host, and the price per million token for it ?

Фото профиля Rahul Wagh
Rahul Wagh25 дней назад

@TryTrustAI @ycombinator @GoogleDeepMind @JaneStreetGroup @medha_rv interesting!!

Фото профиля om_lathiya_
om_lathiya_25 дней назад

@TryTrustAI @ycombinator @GoogleDeepMind @JaneStreetGroup @medha_rv Interesting

Фото профиля Evï Anton
Evï Anton25 дней назад

@TryTrustAI @ycombinator @GoogleDeepMind @JaneStreetGroup @medha_rv is this any different from

Фото профиля Niraj Adhikari
Niraj Adhikari25 дней назад

@TryTrustAI @ycombinator @GoogleDeepMind @JaneStreetGroup @medha_rv Looks good!

Фото профиля just a monolith
just a monolith25 дней назад

@TryTrustAI @ycombinator @GoogleDeepMind @JaneStreetGroup @medha_rv I have nothing to add. But goddamn, you're cute

Фото профиля Marco H
Marco H25 дней назад

@TryTrustAI @ycombinator @GoogleDeepMind @JaneStreetGroup @medha_rv So interesting.. i was just running research on this as I just heard KV cache for first time today and algorithm is showing it to me now.

Фото профиля Yoonseok Yang
Yoonseok Yang25 дней назад

@TryTrustAI @ycombinator @GoogleDeepMind @JaneStreetGroup @medha_rv lets go hannah!!

Фото профиля Santz
Santz24 дней назад

@TryTrustAI @ycombinator @GoogleDeepMind @JaneStreetGroup @medha_rv Congrats on the launch!

Фото профиля John
John25 дней назад

@TryTrustAI @ycombinator @GoogleDeepMind @JaneStreetGroup @medha_rv Dope

Фото профиля AkinAyo | UI/UX Designer
AkinAyo | UI/UX Designer25 дней назад

@TryTrustAI @ycombinator @GoogleDeepMind @JaneStreetGroup @medha_rv Impressive

Фото профиля 𝕐𝕖𝕤𝕙𝕦𝕒
𝕐𝕖𝕤𝕙𝕦𝕒24 дней назад

@TryTrustAI @ycombinator @GoogleDeepMind @JaneStreetGroup @medha_rv Great team 👏🏻💯

Фото профиля E
E25 дней назад

@TryTrustAI @ycombinator @GoogleDeepMind @JaneStreetGroup @medha_rv Bubble

Фото профиля mr ceo
mr ceo25 дней назад

@TryTrustAI @ycombinator @GoogleDeepMind @JaneStreetGroup @medha_rv Good luck

Фото профиля Chandan kumar
Chandan kumar25 дней назад

@TryTrustAI @ycombinator @GoogleDeepMind @JaneStreetGroup @medha_rv Hii

Фото профиля Lokesh
Lokesh21 дней назад

@TryTrustAI @ycombinator @GoogleDeepMind @JaneStreetGroup @medha_rv Any opening role in product management with visa sponsorship?

Фото профиля Amritesh Anand
Amritesh Anand25 дней назад

@TryTrustAI @ycombinator @GoogleDeepMind @JaneStreetGroup @medha_rv Is this something we can try out on our GPU's or you serve endpoints to end clients?

Похожие видео

The Cost of Intelligence is Heading to Zero | Hyperspace P2P Distributed Cache We present to you our breakthrough cross-domain work across AI, distributed systems, cryptography, game theory to solve the primary structural inefficiency at the heart of AI infrastructure: most inference is redundant. Google has reported that only 15% of daily searches are truly novel. The rest are repeats or close variants. LLM inference inherits this same power-law distribution. Enterprise chatbots see 70-80% of queries fall into a handful of intent categories. System prompts are identical across 100% of requests within an application. The KV attention state for "You are a helpful assistant" has been computed billions of times, on millions of GPUs, identically. And yet every AI lab, every startup, every self-hosted deployment - computes and caches these results independently. There is no shared layer. No global memory. Every provider pays the full compute cost for every query, even when the answer already exists somewhere in the network. This is the problem Hyperspace solves where distributed cache operates at three levels, each catching a different class of redundancy: 1. Response cache Same prompt, same model, same parameters - instant cached response from any node in the network. SHA-256 hash lookup via DHT, with cryptographic cache proofs linking every response to its original inference execution. No trust required. Fetchers re-announce as providers, so popular responses replicate naturally across more nodes. 2. KV prefix cache Same system prompt tokens - skip the most expensive part of inference entirely. Prefill (computing Key-Value attention states) is deterministic: same model plus same tokens always produces identical KV state. The network caches these states using erasure coding and distributes them via the routing network. New questions that share a common prefix resume generation from cached state instead of recomputing from scratch. 3. Routing to cached nodes Instead of transferring KV state across the network for every request, Hyperspace routes the request to the node that already has the state loaded in VRAM. The request goes to the cache, not the cache to the request. Together, these three layers mean that 70-90% of inference requests at network scale never require full GPU computation. This work doesn't exist in isolation. It builds on research from across the industry: SGLang's RadixAttention demonstrated that automatic prefix sharing can yield up to 5x speedup on structured LLM workloads. Moonshot AI's Mooncake built an entire KV-cache-centric disaggregated architecture for production serving at Kimi. Anthropic, OpenAI, and Google all launched prompt caching products in 2024 - priced at 50-90% discounts - because system prompt reuse is so pervasive that it changes the economics of inference. What all of these systems share is a common limitation: they operate within a single organization's infrastructure. SGLang caches prefixes within one server. Mooncake disaggregates KV cache within one datacenter. Anthropic's prompt caching works within one API provider's fleet. None of them can share cached state across organizational boundaries. Hyperspace removes this boundary. The cache is global. A response computed by a node in Tokyo is immediately available to a node in Berlin. A KV prefix state generated for Qwen-32B on one machine is verifiable and reusable by any other machine running the same model. The routing network provides the delivery guarantees, the erasure coding provides the redundancy, and the cache proofs provide the trust. What this means for the cost of intelligence Big AI labs scale linearly: twice the users means twice the GPU spend. Every query is a cost center. Their internal caching helps, but it's siloed - Lab A's cache can't serve Lab B's users, and neither can serve a self-hosted Llama deployment. Hyperspace scales sub-linearly. Every new node that joins the network adds to the global cache. Every inference result enriches the cache for all future requests. The cache hit rate rises with network size because query distributions follow a power law - the most common questions are asked exponentially more often than rare ones. The implication is simple: as the network grows, the effective cost per inference drops. Not linearly. Logarithmically. At 10 million nodes, we estimate 75-90% of all inference requests can be served from cache, eliminating 400,000+ MWh of energy consumption per year and avoiding over 200,000 tons of CO2 emissions. The first person to ask a question pays the compute cost. Everyone after them gets the answer for free, with cryptographic proof that it's authentic. Training is competitive. Inference is shared Open-weight models are converging on quality with closed models. Labs will continue to differentiate on training - data curation, architecture innovation, RLHF tuning. That's where the real intellectual property lives. But inference is a commodity. Two copies of Qwen-32B running the same prompt produce the same KV state and the same response, byte for byte, regardless of whose GPU runs the matrix multiplication. There is no moat in multiplying matrices. The moat is in training the weights. A global distributed cache makes this separation explicit. It doesn't matter who trained the model. Once the weights are open, the inference cost approaches zero at scale - because the network remembers every answer and can prove it's correct. No lab, no matter how well-funded, can match this. They cannot share caches across competitors. They scale linearly. The network scales logarithmically. The marginal cost of intelligence approaches zero. That's the endgame.

Varun

37,555 просмотров • 5 месяцев назад