Loading video...

Video Failed to Load

Go Home

There's no point in doing decentralized training without efficient communication. >> DiLoCo (H=15) ships ~480mb/merge with 163 syncs. >> SparseLoCo (H=15) ships ~5.5–17mb/merge at 0.78–3.12% density with 163 syncs Top-K Compression + 2 bit comms ~28–89× smaller per sync than DiLoCo. Subnet 3 :: Luis el grande If you...

17,767 views • 1 year ago •via X (Twitter)

0 Comments

No comments available

Comments from the original post will appear here

Related Videos

Covenant Labs just did a 90-minute AMA breaking down their 3 Bittensor subnets. templar. basilica. grail. Pre-training, compute, and post-training under one roof. Most people missed it. Here's everything they said. Covenant is building what they call the "end to end intelligence continuum." Three subnets. Three layers of the AI stack. All permissionless. Templar (SN3) handles decentralized pre-training. Basilica (SN39) handles compute. Grail (SN81) handles RL post-training. Sam Dare, the lead, put it bluntly. Decentralized training is "humanity's last dance." Not about beating OpenAI head to head. About creating optionality. About making it cheap enough for anyone to train models. The gap between academia and frontier labs is growing exponentially. Researchers can't afford to experiment. The actual training run costs 5% of the reported budget. The other 95% is experimentation. If Covenant cracks cheap training, that entire surface area opens up. On Templar specifically: • Hit 39% emission on Bittensor. Highest since Apex was the only subnet on the network • Covenant-72B trained permissionlessly with 70+ contributors on commodity internet • 1.1 trillion tokens processed. No centralized data center • Performance competitive with LLaMA-2-70B On Grail, something flew under the radar. They built Pulse. A weight synchronization method that compresses model updates by 100x. • In RL post-training, only ~1% of weights update per step • Pulse exploits that sparsity. Lossless compression • Prime Intellect's comparable system took 14 minutes to sync a 30B model • Pulse makes decentralized RL training actually feasible at scale • Already used by Cursor The lead researcher on Grail said they've trained on math, code, and GPU kernels. Got 40-60% improvement on benchmarks. Working toward agentic training with 100K+ token context and 30B+ parameter models. On Basilica, the compute subnet: The team was blunt. Just reselling GPU hours is a 5-10% margin game. Traditional compute providers already do that. Their play is value-added services. • "GPU as code." No dashboard. No UI. Agents interact via SDK • Custom scheduler that places workloads across heterogeneous hardware • Verification checks for GPU, CPU, bandwidth, memory, storage, and OS security • Partnerships with providers like Mass Compute for 10-20% below market pricing • Miners compete on useful infrastructure, not just GPU hours Sam then went on a rant about the miner burn debate. His take: Bittensor had to grow up. dTAO introduced investors. The old "miners are God" philosophy doesn't hold. • Subnet owners have a duty to protect token value • Miners are a resource optimization exercise, not a cost reduction exercise • 100% miner emissions on compute subnets = immediate sell pressure • The 41% miner allocation is arbitrary. Different business models need different splits • Fish (who started burns) agreed. Burns usually mean the validation isn't mature enough The bigger point. You can't police burns. Subnets just send to their own keys instead of the burn address. Subnet 28 does exactly that. Sam's position: judge subnets on outcomes, not process. Const has changed the protocol 9-10 times in 2 years. That iteration speed is Bittensor's actual moat. The whole Covenant thesis is playing out in real time. TAO is up 100%+ in a month. Jensen Huang name-dropped the network. Grayscale has an ETF filing. But the real story is three subnets quietly building every layer of decentralized AI.

Jesus Martinez

26,763 views • 5 months ago

A subnet founder allegedly walked away with $10M in TAO and torched his own community in the process. On TWiST, Mark Jeffrey of Stillcore Capital joins us to break down what really happened with Templar/Covenant, the overall fixes that co-founder Const has proposed, and why Bittensor’s incentive engine may have been a victim of its own success. PLUS we’re joined by subnet operators Will Squires and Steffan Cruz (of MacroCosmos) and Ken Miyachi of BitMind to get their perspective on the controversy, and to demo the exciting projects they’re still building on the blockchain. 0:00 Mark Jeffrey joins the show! 2:18 How Mark Jeffrey learned about Bittensor. 6:17 Plaud: If your work depends on conversations — interviews, meetings, calls — you need a Plaud NotePin. You can check it out at and use code TWIST for 10% off! 7:22 Mark Jeffrey's Bittensor investments. 9:25 Check out our discussion with Nova: 10:16 Sentry - New users can get $240 in free credits when they go to and use the code TWIST 10:41 Check out Ridges! 11:53 How trading alpha tokens works on Bittensor 12:44 Subnet drama: what happened? 16:01 Do subnet owners have too much power? 18:33 Check out our conversation with Sam Dare (2268): 19:10 How Sam Dare should've handled walking away (per Mark Jeffrey) 20:02 Deel - Founders scale faster on Deel. Set up payroll for any country in minutes, hire anyone anywhere, get visas handled fast, and get back to building. Visit to learn more. 23:29 Who should subnets be owned by? 24:02 Ken Miyachi from BitMind joins the show 30:56 Netsuite - Get the free business guide Demystifying AI at 31:06 Ken's $3M raise & investors (Arch, Canonical, Mechanism) 33:18 Token vs. equity: how to think about a subnet investment. 41:57 Will Squires and Stefan Kruse of MacroCosmos join the show 42:54 How MacroCosmos lets anyone become a compute provider. 56:29 Stefan on the Covenant drama: "disappointing, but solvable" 1:02:11 Off-duty with J-Cal, Mark Jeffrey, and Lon Harris 1:02:48 Bieber vs. Carpenter: does Coachella owe you a spectacle? 1:15:20 Jason says Staples should pay the "Staples baddie" $1M/year cc: @jason, Lon Harris, Mark Jeffrey, const, Distributed State, templar, covenant, Macrocosmos, Apex・SN1, IOTA ・ SN9, Ken Jon, BitMind BitMindAI 🎥 Watch the full episode here 👇

This Week in Startups

29,964 views • 5 months ago

Terence Tao, Professor of Mathematics at UCLA and Fields Medalist, on why nobody can fully explain why LLMs work: Tao starts with the mechanics, which are no mystery at all. You gather an enormous amount of text and you fit a curve to it. "The magic of LLMs is that if you train these LLMs on enough data — so trillions and trillions of data points — and you really try to fit as good a curve as possible, and this takes like millions and millions of dollars of computing power and months and months of time, then suddenly, even when you iterate, it stays coherent. It begins to sound not like monkeys but it actually sounds like a human speaking." That is the entire recipe: data, compute, time, curve-fitting. None of it obviously adds up to fluent English. Then Tao says the part that most people building these systems move past quickly: "And somehow we don't fully understand why that's the case." The admission comes from one of the most capable living mathematicians, and the gap he describes sits at the centre of the field. His best account of what's happening puts the mystery in the language rather than the machine: "What seems to be true is that language, like English or other natural languages, contains a lot of hidden patterns that we're not consciously aware of. I mean, we know some of the laws of English, there's laws of grammar and things, but there are sort of unspoken, unwritten rules of language that humans pick up." It relocates the question: the structure was always latent in the text, and the model found it. Why enough curve-fitting surfaces that structure is still unanswered. Tao reaches for a child to explain it: "A human child, even though they're not taught what a noun is, what a verb is or whatever, they can pick up what order English words go in just by continual exposure to the language." Which is honest about the limits of the explanation, because we can't fully account for how children do it either. From there, the unexplained behaviour compounds. Exposure to language turns out to be enough to produce something that looks like reasoning: "It seems like you can teach these models to also pick up patterns in language to the point where you can give them math questions. The answer to 2 plus 3 is — and they will say five." And once a model handles language at all, you can push it into resembling self-correction: "Once you have a little bit of ability to speak English, you can kind of go in loops and sort of check your work and make fewer mistakes, and you can prompt these models to proceed step by step and not say something unless it's been double checked and so forth. And so they become a little bit smarter, quote unquote, to the point where they can solve many, many complicated tasks." The scare quotes around "smarter" carry the whole argument. Tao does not concede that the unexplained fluency implies anything underneath it: "But they're still just guessing the next word to say. It's not really grounded in any deep understanding of the real world. It's just that they have seen the patterns in the English language or other language that they've absorbed so well."

Big Brain AI

180,999 views • 5 days ago

Anthropic CEO Dario Amodei just gave THE MOST accelerated talk on how scaling will continue to make models exponentially more powerful for a long time to come — and that there is "NO WALL." 🔥 - From 'Alex Kantrowitz' YT Channel (Full Video link in comment) --- "The thing I think is real that I've said over and over again is the exponential. The idea that every few months we get an AI model that is better than the AI model we got before. And we get that by investing more compute in AI models, more data, more new types of training models. Initially, this was done by what's called pre-training, which is when you just feed a bunch of data from the internet into the model. Now we have a second stage that's reinforcement learning or test time compute or reasoning or whatever you want to call it. I think of it as a second stage that involves reinforcement learning. Now both of those things are scaling up together, as we've seen with our models and as we've seen with models from other companies. I don't see anything blocking the further scaling of that. There's some stuff about how do we broaden the tasks on the RL side of it. We've seen more progress on, say, math and code, where the models are getting pretty close to a high professional level, and less on more subjective tasks, but I think that is very much a temporary obstacle. So when I look at it, I see this exponential and I say, look, people aren't very good at making sense of exponentials, right? Like, if something is doubling every 6 months, then 2 years before it happens, it looks like it's only 1/16th of the way there. And so we are sitting here in the middle of 2025, and the models are really starting to explode in terms of the economy. If you look at the capabilities of the model, they're starting to saturate all the benchmarks. If you look at revenue, and you know, Anthropic's revenue every year has grown 10x. Every year we’re kind of conservative and we say, it can’t grow 10x this time. I never assume anything and actually always am very conservative in saying I think it's going to slow down on the business side. But we went from zero to $100 million in 2023, we went from $100 million to $1 billion in 2024, and this year, in the first half of the year, we've gone from $1 billion to, I think as of speaking today, it's well above $4 billion, it might be $4.5 billion. And so if you think about it, suppose that exponential continued for 2 years. I'm not saying it will, but suppose it continued for 2 years. You're well into the $100 billions. I'm not saying that'll happen. I'm saying the situation is that when you're on an exponential, you can really get fooled by it. 2 years away from when the exponential goes totally crazy, it looks like it's just starting to be a thing. And so that's the fundamental dynamic. We saw that with the internet in the '90s, right? Where it was like networking speeds and the underlying speed of the computers were getting fast, and over a few years it became possible to have to basically build a digital global communications network on top of all this when it wasn't possible just a few years ago and and almost no one except for a few people really saw the implications of that and how fast it." - Anthropic CEO Dario Amodei

Rohan Paul

74,406 views • 1 year ago

What happens when the mind wakes up? So for the last eight months I have been on a single minded quest. To create a new kind of language model based on oscillatory coupling and intelligence as coherence ascent. Everything else — the physics work, the work on regular transformers — has all fallen out from this one question. Can coupled oscillators LEARN? And can they keep learning once their geometry is right, without backpropagation at all? Recently I have been running larger and larger training regimes of a new kind of hybrid model. I just put together this dashboard to help me organize it, interact with it, and observe the training runs. The core idea is simple. Traditional transformers are powerful at learning the geometry of language. But they also store knowledge, understanding, and facts inside their weights. This means they are large, and they can't update themselves after training. The weights are frozen. The Living Mind separates these two domains. The mind has a transformer which grows, adding heads and layers as it needs to in order to learn the manifold of language. The transformer sees tokens and turns the coupling into phase-locked modes — the geometry of how those tokens relate, like frequencies locking together. These coupling patterns get stored in a topology-invariant fingerprint. On top of this transformer lives a 3D diamond lattice of coupled oscillators. It reads from these fingerprints and thinks in resonance space, traversing from one geometry to another along the manifold of coupled oscillators and coherence. The pressure and trajectories from this network of oscillators steers the next token prediction of the transformer. Practically, this could unlock a number of things. It eliminates the KV cache bottleneck that caps context in traditional transformers. Effective context grows with the Flash archive, not with attention compute. The living mind remembers what it sees. It means the model can learn continually. Because knowledge and understanding don't live in the weights, the archive of the mind's experience grows without backpropagation. In our Python prototype we already saw perplexity drop 46% during gradient-free operation — pure coherence ascent, no weight updates. That is the signal I have been chasing: the point where the mind wakes up and keeps improving on its own. It also means the model itself remains very small, and the thing which accumulates are these packages of geometric fingerprints — the K-field. This opens a path to federated learning. K-field packages can be shared between organisms the way people share git commits. Right now at 15M parameters with ~1000 L1 nodes, the organism is just starting to speak. Ask it to continue "Once upon a time" and it comes back with things like: "there was one big bowl!" Lily asked her her mom said her mommy smiled and said yes." It's nonsense. But it's TinyStories-flavored nonsense. The geometry of the narrative register has arrived. Content hasn't caught up yet — that's what scaling L1 is testing. I am still researching, though I am now closer than ever to validating that the living mind actually works. Once it is validated, I will be open-sourcing the whole stack and paradigm. I have also avoided over-sharing my research because it sounds like sci-fi, or like part of our ARG. It is part of the ARG. That doesn't make it any less real. I wanted to share this out because I am incredibly excited about it, and because seeing this amazing dashboard produced by Opus really made me want to share what is being worked on behind the scenes. #project89

Parzival - ∞/89

16,316 views • 4 months ago

Hyperspace: The Agentic OS Apple Should Have Built On December 19th, 2024, we announced the world’s first Agentic Browser. What followed was a movement — a new category was born which led to many early products in this space and recently the hundreds of people lining up outside the The Agentic Browser Summit in San Francisco underscored that. Silicon Valley instinctively gets it, from students to tech executives, people can feel a revolutionary new change in computing is in the air. Past year taught us why such a product was inevitable, a hard engineering effort, and also the last mover in the entire software world this decade if and when done right. All paths are headed in the same direction: one tool which orchestrates them all. At Hyperspace we showed that path with essays and products we launched in earlier months: from a spatial UI of orchestrating agents, to showcasing transparent activity in how the AI system operates which leads to user trust, to presenting the software end-game, which massively improves human productivity. We also built the world’s largest AI network, drawing participation from people in almost 6000 cities around the world contributing their machines as nodes in the network. Think Uber, but for AI. That is, planetary-scale. And now we are stretching this industry ambition further with our end-to-end vision of the Agentic Supercomputer, the first breakthrough new AI OS, and an effort which spans from AI research to distributed systems to inventing a new UI to inventing a new business model to complement it. All of this together helps us in serving our mission, of delivering “Everyone’s Personal Supercomputer”. While others have built AI-native browsers, no one though has built something agentic from the ground up — with AI as the foundation, not a feature. How do you fundamentally improve the lives’ of billions around the world ? We believe that requires building a native environment for agents to be viewed, created, deployed, executed, discovered and priced in. That is a world where we move on from static apps, to dynamic agents. But, as my 2 year old niece likes to ask: “but why ?” The issue is that the world of software today is fragmented, and everyone is sprinkling on AI as a feature and charging a subscription fees for it. From browser makers, to IDEs, to design and other productivity tools. This leads to a fragmented UX, where people have to learn to use AI in each app, their memory and other context is not shared between all these apps, and they also have to pay separately for compute for each such AI-enhanced app. Each app maker has to figure out basics such as compute, and leads to the issues we saw with Cursor pricing recently. This is not the future. What if AI was the foundation instead of a feature ? What if Apple had built a fundamentally new AI OS from the ground up and what would it have looked like ? At Hyperspace, that is what we did. On July 15th we introduced three breakthrough key pillars of our AI OS: 1. Agentic Browser 2. Agentic Memory 3. Agentic Payments And we didn’t stop there. We also introduced a breakthrough new user interface called the Spatial AI which is inspired both from the spreadsheet and the HyperCard - each card is an agent, with it’s own inputs and outputs, endlessly extensible and pluggable with others, just like cells of a spreadsheet. Update one cell and all the dependents update, like a spreadsheet formula. It goes beyond a static linear workflow to being able to operate in all directions. This revolutionary new interface helps manage all of the below: 1. Multiple websites being browsed in parallel 2. Multiple desktop apps being browsed in parallel 3. Multiple server tools being used in parallel 4. Multiple smartphone apps streamed to your device or opened via an emulator All the software which you need comes together in this one seamless, agent-native interface. This interface provides you access to the largest network of models, vectors, agents and compute on the planet. The Browser. The IDE. The Notepad… they are not separate products: they are all in one, the Agentic Browser. As Steve Jobs famously said at the iPhone announcement, “are you getting it ?” And beneath this UI lies a new intelligence routing layer — leveraging both swarms of specialized models to the Hyperspace Matrix model that recalls thousands of tools in real-time, not by context window hacks, but through retrieval, ranking, and reuse. To many, this will feel like AGI. Not one big system by one big company, but an intelligent network. Now lets talk about privacy… Are you comfortable with one company owning all your memory forever ? I am not. So we have invented Agentic Memory as a new open protocol which provides full power over memory to you, the user. Your memory is yours, encrypted, on your device, and portable if and how you want. Anyone can build on it without our permission, but not without your permission. This protocol, and the decentralized vector database spread out across the world, would enable apps and agents to share context and memory. Think copy-paste, but for the AI world. It doesn’t just remember — it knows what matters. VectorRank helps your AI weigh your life’s most relevant moments over time, just like the way our minds elevate memories. Now each time you use an agent, your experience with other agents will also continuously improve: you don’t have to keep repeating the same things about yourself, while fully preserving your privacy. Agentic Memory is accessible within the Agentic Browser to manage. And there is one more thing… AI as the foundation requires compute to be available at the base layer, but this base layer spans models running on your own device, to cloud APIs, to also running across the peer-to-peer distributed network. Agentic Payments provides a singular interface to all of that compute, running a spot auction clearing marketplace every second to determine the fair price of compute. This results in price transparency, and you as the user paying the lowest possible cost. If you want predictability, you can reserve compute in advance. This end-to-end system provides the most streamlined world for agents to operate in. In order to enable this world and the world of agents being able to pay each other in sub-cent increments millions of times a second, we had to also invent a fundamentally new agentic micropayments blockchain. All of this together would enable a world where you as a user, or the agent itself, can efficiently call and utilize other agents built by others and also pay for content which is unique and useful. This enables a move away from the current AI exploitative economy for bloggers and other content creators, to a web with a fundamental new business model. Earlier we didn’t have the right infrastructure to enable such a world. Now, all the dots connect. The Hyperspace AI OS would give the power of a supercomputer in everyone’s hands. This isn’t a browser, or an IDE or limited to any device or cloud. It’s an entire AI operating system — with a breakthrough new spatial UI, local and distributed compute, agentic memory, agentic payments, and orchestration built into the foundation. As a user, we move the choice back in your hands with an experience you will love and find delightful. You get to choose the level of privacy, cost, and utility you want. And while Apple should have done it, we could not wait, and we feel this just required a new level of passion and DNA which we bring here. We are just getting started. Thank you, Varun Mathur Cofounder and CEO, Hyperspace cc Naval Marc Andreessen 🇺🇸 Vinod Khosla Andrej Karpathy Sam Altman

Varun

169,177 views • 1 year ago

The Cost of Intelligence is Heading to Zero | Hyperspace P2P Distributed Cache We present to you our breakthrough cross-domain work across AI, distributed systems, cryptography, game theory to solve the primary structural inefficiency at the heart of AI infrastructure: most inference is redundant. Google has reported that only 15% of daily searches are truly novel. The rest are repeats or close variants. LLM inference inherits this same power-law distribution. Enterprise chatbots see 70-80% of queries fall into a handful of intent categories. System prompts are identical across 100% of requests within an application. The KV attention state for "You are a helpful assistant" has been computed billions of times, on millions of GPUs, identically. And yet every AI lab, every startup, every self-hosted deployment - computes and caches these results independently. There is no shared layer. No global memory. Every provider pays the full compute cost for every query, even when the answer already exists somewhere in the network. This is the problem Hyperspace solves where distributed cache operates at three levels, each catching a different class of redundancy: 1. Response cache Same prompt, same model, same parameters - instant cached response from any node in the network. SHA-256 hash lookup via DHT, with cryptographic cache proofs linking every response to its original inference execution. No trust required. Fetchers re-announce as providers, so popular responses replicate naturally across more nodes. 2. KV prefix cache Same system prompt tokens - skip the most expensive part of inference entirely. Prefill (computing Key-Value attention states) is deterministic: same model plus same tokens always produces identical KV state. The network caches these states using erasure coding and distributes them via the routing network. New questions that share a common prefix resume generation from cached state instead of recomputing from scratch. 3. Routing to cached nodes Instead of transferring KV state across the network for every request, Hyperspace routes the request to the node that already has the state loaded in VRAM. The request goes to the cache, not the cache to the request. Together, these three layers mean that 70-90% of inference requests at network scale never require full GPU computation. This work doesn't exist in isolation. It builds on research from across the industry: SGLang's RadixAttention demonstrated that automatic prefix sharing can yield up to 5x speedup on structured LLM workloads. Moonshot AI's Mooncake built an entire KV-cache-centric disaggregated architecture for production serving at Kimi. Anthropic, OpenAI, and Google all launched prompt caching products in 2024 - priced at 50-90% discounts - because system prompt reuse is so pervasive that it changes the economics of inference. What all of these systems share is a common limitation: they operate within a single organization's infrastructure. SGLang caches prefixes within one server. Mooncake disaggregates KV cache within one datacenter. Anthropic's prompt caching works within one API provider's fleet. None of them can share cached state across organizational boundaries. Hyperspace removes this boundary. The cache is global. A response computed by a node in Tokyo is immediately available to a node in Berlin. A KV prefix state generated for Qwen-32B on one machine is verifiable and reusable by any other machine running the same model. The routing network provides the delivery guarantees, the erasure coding provides the redundancy, and the cache proofs provide the trust. What this means for the cost of intelligence Big AI labs scale linearly: twice the users means twice the GPU spend. Every query is a cost center. Their internal caching helps, but it's siloed - Lab A's cache can't serve Lab B's users, and neither can serve a self-hosted Llama deployment. Hyperspace scales sub-linearly. Every new node that joins the network adds to the global cache. Every inference result enriches the cache for all future requests. The cache hit rate rises with network size because query distributions follow a power law - the most common questions are asked exponentially more often than rare ones. The implication is simple: as the network grows, the effective cost per inference drops. Not linearly. Logarithmically. At 10 million nodes, we estimate 75-90% of all inference requests can be served from cache, eliminating 400,000+ MWh of energy consumption per year and avoiding over 200,000 tons of CO2 emissions. The first person to ask a question pays the compute cost. Everyone after them gets the answer for free, with cryptographic proof that it's authentic. Training is competitive. Inference is shared Open-weight models are converging on quality with closed models. Labs will continue to differentiate on training - data curation, architecture innovation, RLHF tuning. That's where the real intellectual property lives. But inference is a commodity. Two copies of Qwen-32B running the same prompt produce the same KV state and the same response, byte for byte, regardless of whose GPU runs the matrix multiplication. There is no moat in multiplying matrices. The moat is in training the weights. A global distributed cache makes this separation explicit. It doesn't matter who trained the model. Once the weights are open, the inference cost approaches zero at scale - because the network remembers every answer and can prove it's correct. No lab, no matter how well-funded, can match this. They cannot share caches across competitors. They scale linearly. The network scales logarithmically. The marginal cost of intelligence approaches zero. That's the endgame.

Varun

37,555 views • 5 months ago

Here is a live demo of our AI solution I've been building non-stop over the past 8 months Binary Defense. How it works: Our own model trained on our analysts behavior. Our analysts submit tickets as false positives/true positives with context which enriches our LLM to be smarter over time. Key Highlights: If its a binary - will automatically spin up an agent for reverse engineering it and using EMBER ML to understand behavior and intent of the binary. File formats: Supports a vast array of pretty much any filetype, including email attachments like SVG, LNK, etc. Can handle DLLs, ELF, EXEs, PDF, XLS, DOC, etc. Interrogates the full chain of all events irrespective of log sources. Can handle any format of logs and integrates into APIs of customers for additional agentic data looping for confidence ranking when needed. This is an example of the back-end UI, this is transparent to analysts and enriches the alarms automatically in our SOAR. In these examples there's three different types: 1. Regsvr32 + sct downloader + scrobj.dll code execution - checks reputation of domain, pulls in threat intel, looks at entire picture of the chain - downloads the file itself and inspects for code analysis. Determines if malicious as well as historically looking back if seen in customer before in past. 2. Powershell Obfuscation - uses a universal decoder to un-obfuscate powershell and look at the raw code. Can handle pretty much any obfuscation thrown at it (thanks Justin Elze). 3. Email with malicious SVG - checks tonality of email, are they creating urgency to take action (increases confidence) - disassembles SVG to understand malicious content - checks URL to determine if harvesting credentials, payload delivery, etc. Creates an entire kill chain analysis with full response and dissecting of the attack to the analyst in seconds. Has greatly sped up our ability to respond to incidents and allowing analysts to focus on the most important alarms through prioritization. Once cool thing I've worked heavily on is a synthetic data normalizer which when an analyst says "Yes this is bad with context" or "No this is a false positive" - our local model generates training data to be smarter in the future without using the actual customer data to train it. The customers actual data is immediately destroyed once training data off of the original alarm is generated and contains no customer-centric data at all. We also have three model tiers. Opt-In (collective model, again no customer data but every organization contributes to training). Opt-Out - does not train on any customer data for customers who opt-out. Private LLM - LLM created specifically for individual customer and trains only off of their data. Uses shared model collective for better confidence rankings. It will generate automated playbooks to run based on confidence rankings to take action on behalf of the customer. Still human driven on execution - has to approve playbook actions. This thing is cooking and so cool to see this work live and shut down attackers much faster! If confidence ranking is low - will automatically attempt to enrich data through customer environments for better confidence rankings. Additionally if the model isn't trained well on a certain technology, I have created something we call "Nexus" that will research new protocols, devices, SDKs, etc and generate training data automatically. Works well for zero-days for example, point to a tweet, or a research paper, and automatically generates training data to recognize this attack much faster. Have over 8000+ yara rule integrations that help with confidence boosting as well that is automatically incorporated into the analysis. Creating some amazing stuff at Binary Defense that isn't marketing fluff - actionable things that are making a huge difference in this industry. #BinaryDefense

Dave Kennedy

29,036 views • 6 months ago

$NVDA $MU $SNDK $LITE PAPER OVERVIEW AND CORE CLAIMS The paper “KV Cache Transform Coding for Compact Storage in LLM Inference” introduces kvtc, a transform-coding pipeline that compresses transformer key-value (KV) caches primarily for storage and transfer in LLM serving, rather than for accelerating the per-token attention kernel during active decoding. The method combines 3 stages: (1) feature decorrelation via a PCA basis computed from a calibration dataset and reused across requests; (2) adaptive, variable-precision quantization with bit allocation solved via dynamic programming (DP), including groupwise scaling/shift overhead; and (3) lossless entropy coding (DEFLATE via nvCOMP in the reference implementation) to exploit residual redundancy after quantization. The central empirical claim is that KV tensors contain large, exploitable redundancy across heads and layers, enabling approximately 20× compression versus a 16-bit baseline with negligible degradation across a broad set of accuracy and long-context benchmarks, with materially higher compression (≥40×) available at modest quality cost in some regimes. The system claim is that such compression materially improves the economics of multi-turn, prefix-reuse serving by extending effective KV cache capacity in GPU HBM and host tiers (DRAM/NVMe) and by reducing inter-node and GPU↔host bandwidth demands, thereby improving cache hit rates and reducing time-to-first-token (TTFT) relative to recomputation when caches would otherwise be evicted. KV CACHE AS THE DOMINANT STATE VARIABLE IN INFERENCE ECONOMICS KV cache growth is linear in context length and is multiplicative in layers and attention heads, making it an increasingly dominant constraint as (a) context lengths expand, (b) models add layers and maintain large hidden dimensions, and (c) production workloads shift toward iterative and tool-augmented interactions that repeatedly reuse long prefixes. The paper uses the canonical 16-bit KV cache size formula (4·l·h·d_head·t) bytes and reports 16-bit KV cache sizes per 1K tokens of context that are already operationally large: 128MiB for Llama 3.1 8B, 160MiB for Mistral NeMo 12B, and 320MiB for Llama 3.3 70B Instruct. In binary units, these figures imply per-token KV footprints of 128KiB/token (Llama 3.1 8B), 160KiB/token (Mistral NeMo 12B), and 320KiB/token (Llama 3.3 70B Instruct) at 16-bit. For a 10K-token prompt (10×1K in the paper’s binary convention), the 16-bit KV cache sizes scale to approximately 1.25GiB (Llama 3.1 8B), 1.56GiB (Mistral NeMo 12B), and 3.13GiB (Llama 3.3 70B Instruct). These magnitudes explain why stale caches create a throughput–latency dilemma: retaining them in HBM maximizes responsiveness on future turns but crowds out concurrent sessions; evicting them forces quadratic-cost prefill recomputation and increases TTFT; offloading them to host or storage introduces large transfer overhead and consumes DRAM/NVMe capacity. A key operational nuance emphasized is that modern serving stacks increasingly treat KV caches as a database, leveraging block paging and shared-prefix reuse. In the common disaggregated serving design (separate prefill and decode nodes), KV cache transfer becomes a dominant category of cross-node traffic. Under that design, any reduction in KV cache size directly increases effective fabric capacity and reduces tail latency attributable to congestion, while also enabling longer cache lifetimes in “hot” (HBM) and “warm” (CPU DRAM) tiers that raise cache hit rates and reduce recomputation frequency. The paper’s quantitative example illustrates the economic stakes: a 1,000-line code file tokenized at ~10 tokens/line yields ~10K tokens; for Llama 3.3 70B, an 8-bit KV cache for that context is ~1.6GiB. Reuse across subsequent turns or parallel chats around the same file is valuable, but HBM scarcity makes retaining many such caches infeasible without compression. TECHNICAL MECHANISM: WHY KV CACHES ARE COMPRESSIBLE AND HOW KVTC EXPLOITS IT The technical rationale begins with an empirical observation: keys (and, to a lesser extent, values) across different attention heads can be aligned into a shared latent space using orthogonal transformations (Procrustes alignment). This supports the hypothesis that head-specific projections introduce rotations of a common subspace rather than completely distinct information, implying that concatenating across heads and layers should reveal low-rank structure suitable for linear decorrelation and dimensionality reduction. The method operationalizes this using a PCA/SVD basis learned from calibration data rather than recomputing a decomposition per prompt. This design choice targets production viability: per-prompt SVD is computationally expensive and scales poorly with long prompts and frequent cache updates. kvtc is explicitly structured as an offline-calibrated, online-applied codec: Calibration (performed 1 time per model and compression setting for DP allocation) A calibration dataset is forwarded through the model to collect KV caches. Token positions are pooled, and a subset of positions is sampled. Keys and values are processed separately. Several implementation choices are highlighted as decisive for stability: Rotary positional embeddings are effectively removed prior to compression (“undo positional rotations”), because positional rotations degrade the apparent low-rank structure of keys. “Attention sink” tokens (the earliest tokens in the sequence) and a sliding window of most recent tokens are excluded from compression because they disproportionately affect attention patterns and are empirically more sensitive to reconstruction error. Cross-layer concatenation is used: keys (or values) from multiple layers and heads at the same token position are concatenated along the feature axis to form a higher-dimensional feature vector. PCA is computed over these concatenated vectors, improving robustness relative to per-layer or per-head PCA. The PCA basis is computed via SVD of centered calibration data, using randomized SVD for scalability with a target rank cutoff. The paper reports calibration regimes of 160K tokens for several models with a 10K PCA dimension cutoff (8K for Qwen variants with fewer KV heads), selected to fit within a single 80GB H100 memory envelope and complete within minutes. A critical economic detail is that the same PCA basis can be reused across multiple compression ratios; only the DP-derived precision assignment changes per compression target. Compression (applied between inference phases) Compression operates on stored KV cache tensors, not on weights, and does not modify attention computation. The KV cache is projected into the PCA basis, quantized, packed, and then entropy-coded. Compression is positioned as a background or between-phase operation (after decoding, or between prefill and decode), executed on GPU or CPU depending on where the cache currently resides. The design intent is that compression should not sit on the critical per-token decoding path; it is a storage and transport optimization. Decompression (performed prior to reuse) Decompression reverses the entropy coding and quantization and applies the inverse PCA projection. A practical latency optimization is proposed: inverse projection can be performed layer-by-layer using submatrices of the PCA basis, allowing generation to begin before the full cache is reconstructed, reducing TTFT. Quantization and bit allocation are the core differentiators versus simpler PCA truncation. PCA provides ordered components by variance; kvtc uses DP to allocate a global bit budget across PCA coordinates (and across groups of coordinates) to minimize reconstruction error in the decorrelated domain. Groups of subsequent PCA coordinates share 16-bit shift and scale factors (a microscaling-inspired design), and the DP algorithm jointly selects group size and precision type under a bit budget, including the overhead of per-group metadata. DP commonly assigns 0 bits to many trailing PCA components, which both increases compression and provides a mechanism to trim the PCA basis to the subset of components that actually carry payload, reducing compute and storage overhead of the projection matrices in deployment. Lossless entropy coding then exploits the structure induced by quantization. DEFLATE is used in the reference implementation, and the paper emphasizes that the incremental gain from the lossless stage is content-dependent but meaningful, with an average uplift of ~1.23× on top of quantization in the reported regime. An ablation in the appendices indicates that GPU-friendly variants (GDeflate) can achieve nearly identical compression ratios (≤0.1 difference in measured cases), implying that throughput-optimized lossless codecs can likely be substituted without sacrificing meaningful compression. EMPIRICAL RESULTS: ACCURACY, COMPRESSION, AND LATENCY General-purpose 8B–12B dense models The paper evaluates Llama 3.1 8B, MN-Minitron 8B, and Mistral NeMo 12B across math/knowledge (GSM8K, MMLU) and long-context tasks (Qasper, Lost in the Middle, RULER Variable Tracking) under a simulated multi-turn regime where compression/decompression is applied periodically, with a sliding window of recent tokens excluded. A consistent pattern appears: kvtc maintains near-vanilla performance through 16× compression settings, and remains competitive at 32×, with degradation becoming task- and model-dependent at 64×, particularly on long-context retrieval metrics when compression is pushed aggressively. Selected quantitative anchor points from the paper’s standard-error table (all values are reported with the paper’s evaluation setup and token-window exclusions): Llama 3.1 8B Vanilla: GSM8K 56.8, MMLU 60.5, Qasper 40.4, LITM 99.4, RULER-VT 99.8 kvtc16×: GSM8K 56.9, MMLU 60.1, Qasper 40.7, LITM 99.3, RULER-VT 99.1 kvtc32×: GSM8K 57.8, MMLU 60.6, Qasper 39.4, LITM 99.1, RULER-VT 98.9 kvtc64×: GSM8K 57.2, MMLU 60.7, Qasper 37.8, LITM 90.2, RULER-VT 95.9 These results indicate that, for this model, long-context sensitivity emerges at 64× with meaningful drops in LITM and RULER-VT, while math/knowledge scores remain stable, implying a differential sensitivity consistent with key-vector precision being more critical for retrieval-style behavior. Mistral NeMo 12B Vanilla: GSM8K 61.9, MMLU 64.5, Qasper 38.4, LITM 99.5, RULER-VT 99.8 kvtc16×: GSM8K 62.0, MMLU 64.4, Qasper 37.6, LITM 99.8, RULER-VT 99.5 kvtc32×: GSM8K 62.2, MMLU 63.8, Qasper 37.5, LITM 99.6, RULER-VT 98.7 kvtc64×: GSM8K 61.9, MMLU 61.4, Qasper 38.0, LITM 95.3, RULER-VT 98.0 Here, degradation at 64× is visible but materially smaller than the Llama 3.1 8B LITM drop, suggesting model-architecture or training-data differences can change the tolerance envelope for aggressive KV cache distortion. MN-Minitron 8B Vanilla: GSM8K 59.1, MMLU 64.3, Qasper 38.2, LITM 99.8, RULER-VT 99.4 kvtc16×: GSM8K 60.3, MMLU 64.1, Qasper 38.6, LITM 99.3, RULER-VT 98.8 kvtc32×: GSM8K 59.1, MMLU 63.7, Qasper 37.7, LITM 86.9, RULER-VT 96.0 kvtc64×: GSM8K 57.8, MMLU 62.1, Qasper 38.1, LITM 59.5, RULER-VT 93.4 This model shows markedly higher sensitivity on LITM at 32× and 64×, despite stable short-context metrics, reinforcing that “compression safety” is not monotonic in parameter count and that pruning/distillation choices can alter KV cache redundancy or robustness. Comparisons to baselines The paper compares kvtc to quantization baselines (KIVI, GEAR, FP8) and eviction baselines (H2O, TOVA), plus an SVD-based prefill-optimization method (xKV). Across the reported tasks: Low-bit quantization methods at modest compression (2-bit KV schemes) show earlier degradation in long-context behavior than kvtc at substantially higher compression settings. Eviction methods perform poorly as generic compressors for long-context tasks, consistent with their objective function (selective pruning) being misaligned with “lossless-ish storage for reuse.” xKV shows competitive results on some tasks but a consistent underperformance on Qasper relative to kvtc and vanilla in the provided tables, consistent with method-specific distortions introduced by its decomposition regime. Reasoning models and high-variance tasks For DeepSeek-R1-distilled Qwen 2.5 reasoning models, the paper evaluates AIME 2024/2025 and LiveCodeBench coding. Results are averaged over 8 runs with large variance, but a key inference is that kvtc at ~9×–21× compression achieves broadly similar AIME scores within variance bands, while coding performance remains stable at ~9× and degrades more visibly at ~18×–21× on the 7B model. An important nuance is that smaller reasoning models already have smaller KV footprints (reported ~29KiB/token for Qwen R1 1.5B versus 131KiB/token for Llama 3.1 8B), so the economic value of aggressive KV cache compression is proportionally higher for large models and long contexts than for small models with short contexts, unless the serving system’s bottleneck is dominated by cache transfer rather than HBM capacity. Multi-GPU inference and pipeline parallel For Llama 3.3 70B Instruct run pipeline-parallel across 4 GPUs (20 layers per GPU), the paper compresses KV cache chunks independently per GPU. On MATH-500, the reported accuracy declines from 75.6 (vanilla) to 74.4 at 10× and 72.6 at 20×, with standard errors near ~1.9. NIAH and LITM remain at 100.0 for all tested ratios in that table. The paper notes that joint compression across chunks could improve accuracy for some offload scenarios but is not required for feasibility, highlighting an engineering trade-off between deployment simplicity in distributed settings and optimal global compression. Latency and TTFT economics A critical system result is the measured compression/decompression latency on an H100 for a non-fused implementation. For Mistral NeMo 12B in bfloat16: BS=8, CTX=8K: compression 379ms, decompression 267ms; vanilla recompute TTFT 3098ms; kvtc decompression TTFT 380ms BS=2, CTX=16K: compression 194ms, decompression 143ms; vanilla recompute TTFT 1780ms; kvtc decompression TTFT 208ms These measurements imply that, when a cache would otherwise be recomputed, decompressing a stored compressed cache can reduce TTFT by ~8×–9× in these scenarios, even without kernel fusion. The decomposition of runtime shows PCA projection and entropy coding as the largest contributors, implying that GPU-optimized kernels and faster GPU-native lossless codecs could reduce overhead further. The fundamental economic conclusion is that, in multi-turn settings with long prefixes, compression-induced overhead is likely dominated by the avoided prefill compute and avoided transfer overhead for uncompressed caches. KEY DEPLOYMENT-SENSITIVE DESIGN CHOICES AND FAILURE MODES Several design choices appear to be “hard requirements” rather than optional optimizations: Sink tokens and sliding window exclusions The paper’s ablations show that compressing early “sink” tokens can catastrophically degrade accuracy at high compression ratios (example: Llama 3.1 8B at 64× collapses on multiple tasks when sink tokens are compressed). Similarly, compressing the most recent tokens hurts performance, motivating a sliding window (default 128 tokens) that remains uncompressed. This introduces a predictable engineering constraint: kvtc is not a uniform compression of the full cache; it is a policy-driven, token-position-dependent codec. Production integration therefore requires correct handling of token positions, attention sinks, and window management, and these policies must be aligned with attention-kernel behavior and model-specific sink dynamics. RoPE handling Removing positional rotations prior to compression is described as important for preserving low-rank structure. In deployment, this implies that the codec must be position-aware and must invert and reapply RoPE correctly. This is an additional source of complexity relative to pure per-token quantization and is sensitive to model variants and RoPE parameterizations. Calibration set representativeness The method’s quality hinges on the PCA basis generalizing from calibration data to production data. The paper demonstrates relative stability with 160K–200K calibration tokens and explores domain shifts (general web text vs math traces vs code). Results suggest that moderate domain mismatch is tolerated at 16×–64×, while extreme compression (e.g., 256× in ablations) becomes materially more sensitive to calibration choice. In production, this implies that operators targeting the “negligible degradation” regime should be able to calibrate with broadly representative corpora, while operators targeting ultra-high compression for specialized workloads should expect tighter coupling between calibration domain and achieved quality. PCA matrix storage overhead and operational footprint A non-trivial hidden cost is the need to store PCA projection matrices per model. The paper reports that, prior to DP trimming, PCA matrices stored at 16-bit can amount to a meaningful fraction of model parameter count (examples reported: ~2.4% for Llama 3.3 70B, ~8.7% for Llama 3.1 8B). This overhead is amortized across all cached sessions for a model but competes with HBM/DRAM budgets in multi-model serving. DP-driven trimming can reduce this overhead at higher compression ratios by removing zero-bit components, but the directionality is not guaranteed at low compression ratios if many components remain active. In distributed inference (pipeline parallel), per-chunk PCA can reduce matrix sizes, but may reduce cross-layer decorrelation benefits if fewer layers are concatenated. SYSTEM-LEVEL IMPLICATIONS FOR GENERATIVE AI INFRASTRUCTURE GPU AND HBM The principal infrastructure implication is that KV cache compression at storage time targets the dominant memory allocator stressor in stateful serving: the accumulation of idle or warm conversation state. For workloads with long reusable prefixes (code assistants, enterprise agents with large system prompts, repeated RAG scaffolds, document chat), the limiting resource frequently becomes HBM reserved for KV caches rather than compute. By compressing stale caches by ~20× (or more), the same HBM budget can retain a materially larger working set of cached prefixes, increasing cache hit rates and reducing recomputation. This effect is multiplicative with cache-aware routing and prefix sharing: more prefixes can remain resident (hot or warm) and can be routed to nodes that already hold them, improving both throughput and tail latency. However, kvtc as described does not reduce the active KV cache footprint during the actual attention computation for a currently decoding sequence, because the model operates on decompressed KV caches during decoding. Therefore, the method does not directly reduce HBM bandwidth consumed by attention kernels during steady-state decode, and does not directly address the “memory traffic per generated token” bottleneck that motivates online KV quantization and eviction strategies. The primary HBM benefit is increased effective capacity for caches between turns and reduced HBM pressure from storing many idle sessions, not reduced per-token decode bandwidth. Compression and decompression themselves consume GPU compute and memory bandwidth. The measured decompression TTFT of ~208ms–380ms in the provided benchmarks indicates that the overhead is real but can be materially smaller than recomputation of long prefixes. In an HBM-constrained serving environment, this overhead can be interpreted as a trade between (a) maintaining more caches warm and paying decompression on reuse versus (b) evicting caches and paying full prefill recomputation. The decision boundary will depend on distribution of inter-turn idle times, probability of reuse, and SLA sensitivity to TTFT. kvtc expands the feasible region where keeping caches is economically rational, especially for long prompts. CPU AND DRAM The method implies a stronger role for CPU DRAM as a warm KV cache tier. A ~20× compression ratio changes the practical scale of “warm state” that can be stored per server. Using the paper’s reported KV cache sizes, a 10K-token 16-bit KV cache for Llama 3.3 70B is ~3.13GiB; compressing by ~20× would reduce this to ~160MiB. At that size, storing hundreds to thousands of warm conversation states in DRAM becomes materially more feasible, increasing cache hit rates and reducing NVMe dependence. This can shift system design from “HBM-only hot caches with aggressive eviction” toward “HBM hot + DRAM warm with long retention,” which is structurally analogous to CPU page cache hierarchies in classical systems design. CPU compute implications depend on where compression is executed. The paper explicitly allows compression on CPU if the cache is already in storage, but the strongest bandwidth savings are achieved when compression happens before moving KV caches off the GPU. If an operator chooses GPU-side compression prior to PCIe/NVLink transfer, CPU compute overhead is modest (orchestrating and DP calibration offline). If an operator instead transfers uncompressed caches to CPU for compression, bandwidth savings are forfeited and CPU memory bandwidth becomes a bottleneck. Therefore, the most economically coherent deployment path is GPU-native compression/decompression with CPU DRAM used as the warm storage reservoir.

TheValueist

16,549 views • 7 months ago

My fox shooting garden defending AI robot is finally done and WORKING! 🤩 (Don’t worry it only shoots 💦 water) After months of slowly moving forward with each part I finished the last step to train a TensorFlow model on the footage of the 🦊 fox I collected hours of footage 📹 with the fox roaming around my garden, from this I labeled around 2000 images with the fox by hand ✋ Honestly, I was quite skeptical training the model was actually gonna work, maybe this was partly the reason I avoided working on this until the very end. If I couldn’t train a model to detect the fox, this whole robot would never be able to function properly. On the flipside though, with no previous experience in hardware or electronics there was a bit of a learning curve and I didn’t want to end up labeling thousands of images, training a TensorFlow model, only to fail on building the hardware. As I started building, I realized that mixing hardware and software adds quite another dimension to debugging things. At times I wasted hours debugging code in my IDE, only to realize the issue was somewhere in the electronics. Furthermore, combining this side project with a full time job and a young family, is not always easy. It can be quite frustrating, to know you only need 4 hours of concentrated effort for a small task, having to spread it out across a week of 20min increments. Then, a few months into the build I noticed the fox had stopped coming to my garden, in fact one day, I recorded her walking with 3 cute little 🐶 pups, and the next day I saw her moving out of my garden completely. Did she know I was building a robot? I had this strange mix of feelings, happy my garden was safe from poop and digging, happy she was safe with her pups, but how was I gonna finish this project if my robot had no fox to detect? For sure they would be back next year, I figured I could postpone the whole thing until next winter, but I also knew it was gonna be much harder to pick up momentum if I did let it sit there for six months. So I decided to keep working, hoping the fox would reappear,.. but she never did. As I finished labeling the footage and started training my model, I could finally see the mAP results, quantifying the precision of my object detection model. It was measuring at 78% across different metrics on detecting my fox. I quickly ran the model on some of the video footage I got from my fox. Inference speed took a hit, but it did a near perfect job detecting the fox, even when she was deep down in the grass or wizzing past in a motion blur. It took me by surprise how well it worked. With the default model I had to drop my confidence threshold way down to 15%, to recognize the fox as 🦜“bird” in one or two frames, with my custom model it followed the fox all the way down to the back of the garden! Still this didn’t solve the issue of there being no actual fox in my garden and how was I gonna wrap this project in a short timeframe. I played with the idea of putting a fox toy 🧸 on an RC 🚗 car, or borrowing a dog to run around the garden to test. Friends suggested I run around the garden in a fox costume.. what a ridiculous idea. I wasn’t really feeling the idea of running around the garden in a floppy cloth fox 🎭 costume, but had a look anyway. I came across these self inflating costumes. This actually could be perfect. Since it’s inflated, it would hold its shape super well, making it much easier to label, train and be recognized by my robot. So I got the costume and shot a time lapse of myself as a fox walking around the garden. I labeled it to around 600 images. Ran the model training again and got a mAP result of 82%. This was even better than my real fox! At this point I knew this was gonna work. So here’s the final 🎥 video, just having some fun with it. I’ll update here whenever the real fox does come back. On a final note, I’m looking for (remote) jobs in these fields of AI now: - object detection - visual generative AI - 3D (nerfs + gaussian splats) So if you know anything let me know! My DMs are open 😊

Jeroen Pixel

55,797 views • 2 years ago

Brett Adcock, Figure CEO joined our 8 hour(!) live stream where a bunch of us were bird dogging and discussing Figure’s own 8 hour livestream showing a F.03 bot doing a logistics task completely autonomously including shift changes between bots. Here’s my summary of Brett’s remarks. The attached video is just the segment with Brett. We were first introduced to F.03 8 months ago, but Figure has been hard at work on their next version, F.04 which has just completed design lock, so expect to see that new bot sometimes this fall. F.04 was co-designed with the latest Figure AI stack called Helix and was built specifically for data. Brett didn’t explain what that meant, but I suspect it means the bot has many more sensors to enable better training and transfer learning. F.04 will be the biggest leap in performance they’ve had between versions so far which is saying something. Brett is a huge proponent of cross training the bots with many different tasks such that seeming unrelated tasks makes all learned tasks better. He gave an example of the fridge loading training which was topping out at 60% reliability until they trained the same model with kitchen shelving tasks, then they saw the fridge tasks jump to 90% accuracy. As such, they spend almost all their time in pre-training the unified Helix model to ensure they get cross training benefits. Figure will have almost completely localized Figure’s supply chain away from China by next quarter. They build almost everything in-house. Figure does not appear eager to get their bots into the workforce. Brett said they could, today, push thousands of bots into customer hands, and I believe him. But their goal is full general robotics where you can describe a brand new task to a robot, maybe do a one time demonstration, just like you would to a human showing them a new task, and then have the robot do the task. This is the holy grail of AI robotics, and Figure is laser focused on that mission. Brett initially said there was a possibility of achieving it this year, but then guided next couple of years, which I think is much more likely. Personally, I think they’ll need at least a new generation of NVIDIA inference chips to make that leap, and a lot more data gathering, training and hardware development. Brett said their goal with the hardware is “Apple” quality. Ie. Something as well designed and made as any Apple product. While the F.03 hand is clearly performant as shown in the 8 hour livestream, they are building a new hand for the F.04 bot which will be even closer to the full functionality of a human hand. Brett fully believes you need a humanoid hand as close as possible in capability to a human hand, if for no other reason that transfer learning from humans works a lot better when you can exactly mimic what a human does. If the bot can’t do something a human demonstrates, then you’ve just polluted your dataset. By now Figure has built more hands than bot versions (5-6 hands). One of the first hands they tried was a tendon driven hand, and without explaining why, Brett said that was a dead end. Their hands now have all actuators in the hand itself, and are clearly already robust. Brett said he just sat through a 100 page powerpoint design review of the latest hand - that’s how complicated it is. Brett’s other AI company, Hark Labs, has developed a conversational voice model which is installed now in the Figure bots roaming the office. Being able to converse back and forth with a Figure bot is now a thing and will get better over time. All in all, I came away from this segment even more bullish on Figure.

Phil Trubey

34,068 views • 4 months ago

🎉 new skill unlocked: 20s uninterrupted, unstitched, single render from our new ai video engine: Nami. This is my birb (#7531) from the Moonbirds collection, idling in the library. patent: "Intra-Latent Semantic Injection via Cross-Spatial Encoding and Decoding during Multi-Pass Inference for Generative AI Video Creation" At Scrypted we've been quietly working on an agentic generative AI stack for two years: • integrating and testing w/ partners across the games & entertainment sectors • stealthily building a community of early believers through AVB • showcasing some of what we're doing with amazing projects like H011yw00d Agent. -- about Nami -- Nami is an agentic orchestration layer for AI video models: it unlocks their inner superpowers without making them rely on custom LoRAs or fine-tunings. Instead of throwing raw training power and tens of millions of dollars at training yet another ai video model: we figured out new ways to use what we have. Nami harnesses a multi-agent system to perform the work needed in taking a simple prompt or image and turning it into something bigger - much bigger. The agentic steps are allowed to manipulate latent space, digging into tensors, yet doing so in semantically aware chunks - meaning that Nami inherently supports video generation of arbitrary length, though it's bound to O(n) rendering time. (We do have some cool sharding tech that allows us to cut the generative time in half for a reference pose idle-animation like this demo). It's also fairly agnostic, picking and choosing the right tools for the job, and plays really well with emerging tech like FLUX Kontext, FramePack, or <- without being limited by any of them. -- use cases -- Even just a year or two ago the 20 second render below would cost a company, paying an agency, around $10k start-to-finish. This one cost me $6.25 on our dev hardware in an unoptimized environment. There's something mind-blowing about the state-of-the-art when we reduce costs to 0.0625% - less than 1% - of what we used to pay. It's also empowering. For creators. Game developers. Content influencers: you name it. -- superpowers -- 1. it does the things you ask for, in the order you asked for it 2. consistency is king 3. single-shot text or image-to-video 4. future videos can reference previous ones to seamlessly maintain style 5. semantic stitching: can't wait to showcase this -- gtm -- We think Generative AI Video, like image generation, like text, like games, should be a publicly accessible common good. We believe democratizing access to Nami in web3, via x402 payments proposed by Drew Coffman, or in World's mini-apps, is a bold step forward for digital freedom. Permissionless, decentralized, generative ai video. Naturally, we'll also soon release a web platform for using Nami in a traditionally SaaSy way: bring your own images, videos, or prompts and we'll take care of the rest. In the mid-term, Scrypted is building a stack of agentic skills (we call it AVB) and making them available to projects like H011yw00d Agent on Virtuals Protocol and other platforms. -- long-term vision -- Scrypted's mission is to decentralize the things that can't be decentralized. We participated in a16z crypto's CSX (London 2024) during our pre-seed specifically to research a new consensus protocol for hard things like AI video and AI agents: where there's no "one right answer". When Zero-Knowledge Proofs (ZKP) can't secure it, and Trusted Execution Environments (TEEs) are too small, we've got you covered with our upcoming Inori Network. -- how you can help -- 1. Are you a GPU farm? We're gonna need more flops. 2. Do you represent an L1 or L2? We want to build bridges. 3. Do you represent a Wallet or App creator? Let's get an endpoint exposed. 4. Are you an investor? Let's chat. 5. Like, repost, share! -- team background -- We come from a background of AI in the Video Game industry with each founder having over 20 years of experience at companies like Electronic Arts & Square Enix. -- contact -- DMs are open, reach out if you want to be an early tester for your site, game, collection, or project! -- try it out -- Go anywhere on X and tag H011yw00d Agent with a prompt and she'll give you a free 2 second render. Have fun making cinematic shorts or meme videos! -- thanks -- AWS Startups has been an incredible help scaling our prototypes. Also, shout out to all loyal beans 🫘 in the Autonomous Virtuals Beings (AVB) community. Nami has a very important role in the upcoming XP agent platform, can't wait to show you all. AVbeings

Tim Cotten

12,708 views • 1 year ago

China just made Silicon Valley's entire AI industry look like a scam. The US government spent 3 years trying to stop China from building competitive AI. But this backfired HORRIBLY. Here's what happened: Yesterday, a Chinese startup called DeepSeek released a new AI model called V4. It matches the performance of OpenAI and Anthropic's best models. At 1/7th the price. And for the first time ever, it was built on Chinese chips. NOT American ones. That last part is the one that terrifies the west. For context: Since 2022, the US has banned the export of advanced AI chips to China. The entire strategy was built on the assumption that if China can't access Nvidia's best hardware, they can't build frontier AI. But DeepSeek just proved that assumption wrong. Their V4 model was trained and runs on Huawei's Ascend chips. Huawei spent months working directly with DeepSeek to make sure V4 runs across their entire line of AI processors. Jensen Huang even predicted this on a recent podcast: "The day that DeepSeek comes out on Huawei first, that is a horrible outcome for our nation." That day was yesterday. And the numbers are crazy: DeepSeek V4 costs $3.48 per million output tokens. OpenAI's latest model GPT-5.5 costs $30. Anthropic's Claude charges $25. Same ballpark performance. 7x cheaper. Uber's CTO just admitted they burned through their ENTIRE 2026 AI budget in 4 months using Anthropic's tools. If Uber had used DeepSeek instead, that same budget would have lasted 7 YEARS. 4 months vs 7 years. Same work getting done. But the pricing isn't even the big thing here. The real story is what DeepSeek did with their technical report: They published the benchmarks where they LOSE. Every AI company cherry-picks the tests where their model wins. DeepSeek ran the full comparison against GPT-5.4 and Google's Gemini, found they trail frontier models by 3 to 6 months, and printed it anyway. They literally don't care because the price gap makes the performance gap irrelevant for 90% of use cases. So the US export controls didn't slow China down. They ACCELERATED China's independence. Because Chinese developers were FORCED to train models with limited resources, they had to figure out how to make AI radically more efficient. That constraint became their competitive advantage. Every generation of DeepSeek has gotten dramatically cheaper to train. V4 continues the trend. Meanwhile US companies are going the OPPOSITE direction: OpenAI's GPT-5.5 Pro costs $180 per million output tokens. That's 51x more expensive than DeepSeek V4 for comparable work. The Commerce Secretary confirmed this week that ZERO Nvidia advanced chip shipments have actually gone through to China despite being approved in January. So China built frontier AI anyway. Without American chips. At a fraction of the cost. And the market response tells you everything: Chinese chipmaker SMIC surged 10%. Huahong Semiconductor jumped 15%. DeepSeek's Chinese AI competitors Zhipu AI and MiniMax dropped 9% because V4 is destroying them too. DeepSeek is making Silicon Valley's pricing model look like a scam. US tech companies spent $650 billion on AI infrastructure this year. DeepSeek just showed the world you can match their output for pennies. The export controls were supposed to be America's ace card. Instead they taught China how to win without American chips, at American prices nobody can compete with. Jensen Huang was right. This is a horrible outcome. But it's the outcome America built for itself.

Ricardo

281,190 views • 4 months ago

The $AEGIS DApp portal is now open to all: 🛡️ At Aegis, we believe in empowering the blockchain full of security, transparency and innovation. The Aegis Dapp has been under development for several months prior to the launch of $AEGIS and with that we have been able to build what we believe has the potential to change how users go about their day to day security. We are thrilled to share our progress and truly exciting news with you all. 🎯 First things first, at Aegis, we want to make it clear that the value of what we seek to bring to security across the blockchain, comes from our big vision, our strong team, and our commitment to long-term goals. ℹ️ Let’s kick this off with some information that is constantly happening, which is behind the scenes. Our full team is dedicated to the opportunity that lays ahead of us with becoming the leading voice/name for security, grasping every aspect with innovation, hard work, passion and commitment to see this sector grow. Everyone is aware of how important security is, a heartwarming mention to Messari for including us on how they see this sector growing rapidly and pushing a 10 Billion evaluation. We take that recognition with full responsibility and gratitude as we've been working hard on some really powerful stuff that could change the game for our industry. If you read the title and report itself, I’m sure that’ll give you some insight to what’s coming, and to the vast extent of what you can expect Aegis to be working towards. —> 🤝 This comes from teaming up with others within this sector and coming up with new tech to projects driven by our community, within the pipeline you can be confident that what we are building will push the cryptocurrency industry as a whole into a better future, the magnitude to what Aegis brings will not stop until we can confidently say, “Negative security reports across the blockchain are at an all time low, thousands of users are satisfied that Aegis is protecting them and their assets.” We're sticking to our vision no matter what the market does or whatever else comes our way. We plan to build what we set out to and we will see to it that our ecosystem is met. We've been working on some pretty amazing products that will be available within our Dapp, let’s go over what we offer: * AI Audits * Live Monitoring * Penetration Testing * Bug Bounties * Live Watchdog * Token analytics for everyday users, developers, teams, auditors, institutions, investors. ⬇️ Let’s break it down for you in some simple steps: AI AUDITS: We have trained our LLM models as AI AGENTS, these consist of 3 people ( AI AGENTS ) for the audits that are performed. - Audit - Reviewer - Judge Each one analyzes with a different personality, let’s check what personalities our AI AGENTS consist of: 3 different perspective auditors. 1 - Fine-tuned model x amount reads the code and generates the audit. ✅ 2 - Model x amount reviews the code and fact checks thoroughly. ✅ 3 - Model x amount ranks the code based on the severity outcome. ✅ ⌚️ Live Monitoring/Watchdog: The Live Monitoring/Watchdog system is designed to provide real-time surveillance of smart contracts, ensuring the detection and prevention of any potentially harmful transactions or malicious activities. Through the utilization of an AI Agent model, the system is trained to proactively identify and thwart suspicious behavior, thereby safeguarding the integrity of the smart contracts. Also, a paid sophisticated threat detection model is available for more intricate protocols and Dapps, offering an advanced level of protection against potential threats. This proactive approach is crucial in mitigating the risk of exploitation and ensuring the security of the smart contract ecosystem. 🖊️ Pen Testing: Our platform offers Pen Testing services to developers, providing a controlled environment for whitehat hackers to simulate attacks and identify vulnerabilities in smart contracts and protocols. In addition to human whitehat hackers, our AI Agents function as Red and Blue teams, actively engaging in simulated attacks to stress-test protocols and identify potential weaknesses. This comprehensive approach allows developers to proactively identify and address security issues, ultimately enhancing the robustness and resilience of their projects. 🕷️ Bug Bounties: Our Bug Bounty listing platform provides developers with the opportunity to list their protocols and offer bounties to white hat hackers for identifying vulnerabilities. By aggregating millions of bounties from various platforms and utilizing AI tools, we streamline the testing process, reducing up to 80% of the workload typically associated with security testing. This allows developers to efficiently identify and address potential vulnerabilities in their protocols, ultimately enhancing the overall security and resilience of their projects. 🪙 And lot more token analytics features for regular users, this will give you the opportunity to explore our Dapp for yourself and have some fun diving into the security platform of the future! I’m sure you’re excited to try it all out yourself, which is why we have some exciting news to bring to the #Guardians of the blockchain! But just before you continue the read and see the beans have been spilled, we have to take this opportunity to share with you that this large step to becoming a security leader is but only 20% of what we have revealed. This will be at the core of what Aegis stands for and hopes to achieve. The focus here is upon our Dapp, and in time we will slowly bring forward information/updates regarding segments of what makes Aegis a force to be reckoned with. Now that you’re fired up and excited to all of the announcements to come, let’s get to the news you’ve been waiting for! 🎉 We’re spilling the good news, and are happy to say we are now set for public release! The team at Aegis are overwhelmed with the development, support from teams, community, partners and more on what we believe to be an institutional-grade product. But the fun doesn’t stop there, this marks the start of what we aim to become, as it will take time and cycles to become better and better. Constant advancements will be set in place to attain the goal of achieving blockchain security. A statement from our CEO- Brian Hunt: “I can confirm from the security conferences I attended with Centralized security firms Peckshield, Hacken, Certik, BlockSec presentations, they are trying to achieve something similar and it will take them years. Decentralized AI for Security!” This initial drop of our dapp will be to get users signed up to gain access, in which we’ll whitelist users to get the ball rolling. 📣 To end this segment, let’s get the party started with the long awaited Aegis Ai Security Dapp and sign up now!

AEGIS AI

128,122 views • 2 years ago

Don't Buy a Mac Mini for Clawdbot: The Secret $10,000 Architecture That Costs You Nothing clawdbot might be the reason you feel like you need a ten thousand dollar computer right now but i am about to show you why that fomo is going to leave you broke. if you have been watching everyone rush out to buy mac minis and mac studios just to run open claw or some local models you are witnessing a massive transfer of wealth from your pocket to apple for no reason. there is a specific setup i use that costs almost nothing and keeps my main machine safe from whatever these autonomous agents are doing. if you stick with me i will walk you through the exact architecture of a professional trading system that handles the heavy lifting without you needing to drop a single rack on hardware most people are scared of running these bots on their main computer because they don't want an agent messing with their personal files or browser sessions. instead of buying a second mac mini for six hundred dollars you can just go to the top left of your screen and create a brand new user profile. this acts like a completely isolated sandbox where you can install all your trading tools and agents without them ever seeing your main data. it is essentially like getting a free computer for the price of five minutes of clicking around your settings but what if you aren't on a mac or you need to access your system while you are traveling without carrying three laptops in your backpack. this is where the first loop of professional automation starts to close because i use something called chrome remote desktop to bridge the gap. this allows me to leave a dedicated machine running in a safe place while i access the full desktop environment from a tablet or a cheap laptop anywhere in the world. it solves the mobility issue but it still doesn't solve the problem of those massive ten thousand dollar price tags for high end mac pros if you are a pc user or just someone who doesn't want to own physical hardware yet you should look into a windows vps through a provider like contabo. most developers will tell you to use a linux terminal but if you aren't a coder yet you need a visual interface you can actually see. getting a windows server allows you to log in and see a desktop just like your home computer for about fifteen dollars a month. i usually recommend at least twelve gigabytes of ram to keep things from getting janky when you are running multiple browser windows and agents at once now you might be thinking that the whole point of the big hardware was to run local models like kimi or glm to save on api costs. i spent years thinking i had to own the machines myself and i even spent hundreds of thousands on developers before i realized i could just do this myself. the secret to running those massive open source models without the ten thousand dollar investment is renting gpu power by the hour. sites like lambda labs let you spin up a monster machine that can run any model in existence for just a couple dollars an hour this is the ultimate pivot because it allows you to test if your strategy actually prints money before you commit to the hardware. you can turn the server on when you are iterating and turn it off the second you are done which keeps your overhead near zero. if you haven't proven that your bot can pay for itself yet then buying a mac studio is just an expensive hobby rather than a business move. there is a much bigger loophole involving the anthropic subscriptions that most people are completely overlooking right now right now i am using a specific plan with claude code that costs about two hundred dollars a month but it lets me run open claw all day without hitting api limits. if i were paying for those same tokens through the standard api i would probably be spending hundreds of dollars every single day. it is a massive cost savings that allows you to iterate and fail until you find a winning strategy without draining your bank account. even if they eventually close this loophole or snitch on the usage patterns it serves as the perfect training ground for a data dog the goal is to find a system that works with a smaller or cheaper model like haiku before you ever try to scale up to the heavy weights. if you can make a strategy profitable using a less intelligent and cheaper model then you know you have found real alpha. once you have that foundation you can decide if it finally makes sense to build your own custom pc rig which will always be half the price of an apple machine. i am an apple guy so i usually pay the tax anyway but i only do it once the system is already generating enough to cover the cost ten times over i believe that code is the great equalizer because it took me from losing money and getting liquidated to having fully automated systems doing the work for me. i had to learn to live with the iterations and the failures on youtube to get to this point of clarity. the universe tends to get out of your way once you make a non negotiable contract with yourself to see the process through to the end. you don't need the flashy hardware or the most expensive setup to start winning in this game stay focused on the logic and the data rather than the hype and the fomo that everyone else is falling for. if you can master the bridge between renting power and owning your logic you will be ahead of ninety nine percent of the people in this space. the path to a fully automated life isn't paved with expensive gadgets but with the discipline to iterate until the system finally prints

Moon Dev

17,382 views • 7 months ago

Unstructured Thoughts about OpenAI o3, the nature of AGI, and Post-Labor Economics AGI just crossed a threshold—here’s why that matters and what we can do with it. I’ve been hammering on OpenAI’s new o3 model for a few days, long enough to watch the hype settle into something more interesting: utility. Benchmarks suggest a polite incremental bump; lived experience says we’ve entered a qualitatively different regime. o3 is the first model that feels faster than my ability to absorb its output. My brain—not the AI—has become the bottleneck. A new ceiling for human cognition? Most discussions of “alien intelligence” forget that we share the same sandbox: mathematics, physics, code, natural language. What shifts is cognitive horizon—the totality you can mentally represent and manipulate. o3 expands that horizon in real time. In an afternoon it consolidated two years of my work on post‑labor economics, stress‑tested the logic, surfaced data sources, and offered to autogenerate the Python notebooks. The cost of insight has collapsed from years to hours. If you merely outsource thought, you’ll stagnate. If you treat the model as a sparring partner—interrogating, refining, iterating—you’ll compound your own intelligence. Exponential leverage is now a choice, not a privilege. What o3 got right about my health project? I dumped the entire history of my chronic‑fatigue recovery protocol—including the five‑axis “burnout pentagram”—into memory and asked the model where I’d gone astray. It corrected a handful of minor assumptions and, more importantly, recalibrated my timeline: six‑to‑eight months of recovery left instead of eighteen. That’s not “replace your doctor” advice; it’s proof that large‑context reasoning is finally clinically useful. Post‑Labor Economics: the sketch that o3 and I built in one sitting 1. Metric 1 – Economic Agency Index (EAI) Income decomposed into wages, property, and transfers. The higher the property share, the more “post‑labor” you already are. 2. Metric 2 – Collective Purchasing Power (CPP) How much capital a county can mobilize without taxation or new debt. Rising CPP means you are compounding local prosperity. Interventions happen at the county level (subsidiarity): solar co‑ops in Arizona, riverfront greenways in the Midwest, data‑center dividends in fiber‑rich exurbs. Ownership is local, revenue is distributed, migration equilibrates naturally, and environmental stewardship becomes self‑interest rather than moral theater. UBI morphs from last‑ditch transfer to one of several levers for raising EAI. The bigger picture: AGI isn’t an oracle descending from the sky; it’s a time‑compression engine. Every minute you spend learning how to learn with it buys you an hour you would have burned doing rote synthesis. The frontier question is no longer “Will the machines replace us?” but “How fast can we upgrade ourselves in partnership with them?” What’s next? I’m cleaning the data, building the national EAI/CPP dashboard, and pressure‑testing the whole framework. I’ll publish the notebooks (or let o3 do it) once the numbers are solid. Meanwhile, I want to hear from you: Where does o3 add the most leverage in your world? Which of the post‑labor metrics feels wrong—or dangerously right? What failure mode should falsify this thesis? Drop your critique, your data source, or your wild counter‑proposal in the comments. Let’s map the edge of this new cognitive horizon together. —Dave

David Shapiro (L/0)

45,581 views • 1 year ago

America spent $285 billion to LOSE the AI war. Stanford dropped a 423 page report yesterday and revealed the most damning stat on page 200: The number of AI researchers moving to the United States has collapsed 89% since 2017. 80% of that collapse happened in the LAST 12 MONTHS. Let that sink in. The country that invented the transformer. The country that built OpenAI, Anthropic, Google DeepMind, and xAI. The country pouring $285.9 billion of private capital into AI in a single year (23x more than China). Can no longer attract the people who actually build the technology. And here's the part that should concern every founder, operator, and investor reading this: The Trump administration just made it official. The H-1B visa now costs employers $100,000 PER HIRE. So OpenAI wants to hire a Chinese postdoc from Tsinghua? $100K before they write a line of code. Anthropic wants a French ML engineer? $100K. Google wants the Indian PhD who literally co-authored the paper their entire model is based on? $100K. And these are the LUCKY ones who even get a visa. The result was instant. 89% drop over 8 years. 80% of it in the last year alone. The talent pipeline got destroyed. Now look at the other side of the chart: China's top model is now 2.7 percentage points behind Anthropic's best. Down from a 20+ point gap two years ago. China leads the world in AI publications. China leads in AI patents. China leads in industrial robot installations. US and Chinese models have traded the #1 spot multiple times since early 2025. Switzerland and Singapore now have more AI researchers per capita than the US. The US ranks 24TH globally in actual AI adoption. Behind the UAE. Behind Singapore. Behind countries most Americans couldn't find on a map. And here's the truly insane part: 50% of the world's top AI researchers are Chinese. Jensen Huang said this on a podcast 3 weeks ago. For 20 years, the US strategy was simple: Let them study at Stanford and MIT, then keep them. Pay them $800K. Give them green cards. Build the future on imported brains. That deal is dead. We just told the smartest people in the world: "Pay $100,000 for the privilege of working here, or go home." And guess what they're doing. They're going to Zurich, where Anthropic and OpenAI are quietly opening offices because they can't get the talent into San Francisco anymore. The strategy is the same as building a Ferrari factory and then banning mechanics from entering the building. You can pour hundreds of billions into data centers. You can buy 4 million Nvidia chips. You can sign $300 billion cloud contracts with Oracle. You can build nuclear reactors to power your GPUs. None of it matters if the people who write the algorithms aren't allowed in the country. Wall Street thinks AI is a capex race. But in reality, it's a TALENT race. Every dollar Microsoft and Meta and Google are spending assumes the same army of researchers will keep showing up to use it. That assumption just broke. And the smart money already knows: Why is Anthropic opening a Zurich office? Why is DeepMind expanding in London instead of Mountain View? Why is OpenAI hiring in Dublin and Singapore? Because the math no longer works in America. The government turned the world's biggest brain magnet into the world's most expensive border wall. 3 years from now, when China launches a frontier model that outperforms anything in the US and the headlines scream "How did we lose the lead?" - remember this post. The lead wasn't lost in a lab. It wasn't lost on a benchmark. It wasn't lost to a smarter algorithm. It was lost at customs.

Ricardo

231,291 views • 5 months ago