Loading video...

Video Failed to Load

Go Home

Robots need strong visuo-motor representations to manipulate objects, but it’s hard to learn these using demo data alone. Our #RSS2024 project vastly improves robotic representations, using human affordances mined from Ego4D! w/ Mohan Kumar Srirama Shikhar Bahl Abhinav Gupta

11,041 views • 2 years ago •via X (Twitter)

2 Comments

Sudeep Dasari's profile picture
Sudeep Dasari2 years ago

If you’re at @RoboticsSciSys in Delft, drop by our presentation and poster this Wednesday (session 10) at 1:30 PM! Website: Paper:

Yang's profile picture
Yang1 year ago

Want to learn how practical AI skills and automations for your business and work? Check out our 50+ step-by-step video tutorials 100% FREE 20+ hours of Ai and Automation goodness absolutely free 🥳

Related Videos

Check out our #PAMI paper with code "Dense Continuous-Time Optical Flow from Event Cameras," where we show how to regress *continuous-time* trajectories of every pixel from event cameras alone or events plus frames! The key idea is to iteratively estimate per-pixel polynomials using a recurrent lookup and update scheme. Paper: Code: DOI: We present a method for estimating dense continuous-time optical flow from event data. Traditional dense optical flow methods compute the pixel displacement between two images. Due to missing information, these approaches cannot recover the pixel trajectories in the blind time between two images. We show that it is possible to compute per-pixel, continuous-time optical flow using events from an event camera. Events provide temporally fine-grained information about movement in pixel space due to their asynchronous nature and microsecond response time. We leverage these benefits to predict pixel trajectories densely in continuous time via parameterized Bézier curves. To achieve this, we build a neural network with strong inductive biases for this task: First, we build multiple sequential correlation volumes in time using event data. Second, we use Bézier curves to index these correlation volumes at multiple timestamps along the trajectory. Third, we use the retrieved correlation to update the Bézier curve representations iteratively. Our method can optionally include image pairs to boost performance further. To train and evaluate our model, we introduce a synthetic dataset (MultiFlow) that features moving objects and ground truth trajectories for every pixel. Our quantitative experiments suggest that our method successfully predicts pixel trajectories in continuous time and is competitive in the traditional two-view pixel displacement metric on MultiFlow and DSEC-Flow. Open source code and datasets are released to the public. Kudos to Mathias Gehrig Manasi Muglikar

Davide Scaramuzza

12,682 views • 2 years ago

Imagine controlling a real robot from your home… no money, no experience needed. Sounds crazy, right? But it’s already possible. BitRobot 🦾 is building the world’s first open robotics lab powered by crypto incentives. Instead of one company doing everything, it connects people from all over the world to work together on real robotics and AI tasks. The network is made up of specialized subnets, each focused on different missions from collecting real-world data with robots to developing humanoid robots for everyday use. What makes it powerful? It uses crypto rewards to coordinate global resources like compute power, robot fleets, teleoperation time, and even human effort. This allows BitRobot to scale much faster than traditional labs. Now here’s the best part 👇 The easiest way to get involved right now is through TeleArms. You don’t need: – a robot – engineering skills – or any investment – Hardware All you need is a laptop and an internet connection. From your home, you can remotely control a real robotic arm inside BitRobot’s lab using your keyboard or mouse to pick up, move, and place objects. Every action you take helps generate real-world data that trains the next generation of AI to perform useful physical tasks. So you’re not just playing with a robot… You’re actually helping build the future of AI. I’ve been talking about BitRobot for a while, and now TeleArms is live! You can control a real robotic arm from home, but it’s in a private beta with limited access. I’m now an ambassador for BitRobot Network. I’m giving 4 exclusive access codes to my community so they can experience it too. A lot of people want to experience this, but since it’s limited, I decided to do a random giveaway. To participate in this giveaway : 1. Join the BitRobot Network Discord (Link in comments) 2. Come back to this post and comment below, explaining why you want to join TeleArms and how you plan to contribute. Note : Winner will be announced in the last 7 days. Once you do that, you’ll be in the running for one of the codes! Good luck, and I can’t wait to see your ideas!

Apurba.Eth

36,318 views • 5 months ago

Yann LeCun (Yann LeCun ) beautifully explains how the architecture and principles used to train LLMs can not be extended to teach AI the real-world intelligence. In 1 line: LLMs excel where intelligence equals sequence prediction over symbols. Real-world intelligence requires learned world models, abstraction, causality, and action planning under uncertainty, which current next-token training does not provide. He says current LLMs learn by predicting the next token. That objective works very well when the task itself can be reduced to manipulating discrete symbols and sequences. Math, physics problem solving on paper, and coding fit this pattern because success largely comes from searching and composing the right sequences of symbols, equations, or program tokens. With enough data and scale, these models get very good at that kind of structured sequence prediction. Real-world intelligence is different. The physical world is continuous, noisy, uncertain, and high dimensional. To act in it, a system needs internal models that capture objects, dynamics, causality, constraints from the body, and the outcomes of actions over time. Humans and animals build abstract representations from rich sensory streams, then make predictions in that abstract space, not at the raw pixel level. That is why a child can learn intuitive physics, plan multi-step actions, and adapt quickly in new situations with little data. His claim about saturation follows from this gap. Scaling token prediction keeps improving symbol manipulation tasks like math and code, but it hits limits on embodied reasoning and common sense because text alone does not provide the right learning signals for world models. Predicting the next word cannot efficiently teach contact forces, affordances, occlusion, friction, or how actions change the state of the environment. For that, he argues we need architectures that learn abstractions from sensory data and predict futures in abstract latent spaces, then use those predictions to plan actions toward goals with built-in guardrails. --- From 'Pioneer Works' YT Channel (link in comment)

Rohan Paul

104,460 views • 9 months ago

Welcome to the Lab of the Future! 🧬🤖 Excited to share LUMI-lab, out today in Cell — a self-driving platform that pairs an AI foundation model with a robotic lab to autonomously discover ionizable lipids (LNPs) for mRNA delivery. The core problem: Designing lipid nanoparticles (LNPs) is hard. The chemical space of ionizable lipids is vast, experimental cycles are slow, and — critically — historical LNP datasets are far too small to train a predictive model from scratch. Most AI approaches in this space hit a wall immediately: not enough data to learn from. Our solution: lab-in-the-loop foundation model learning. Instead of training on LNP data alone, LUMI starts as a transformer-based foundation model pretrained across broad chemical space, building rich molecular representations before it ever sees a single LNP experiment. Then it enters a closed loop with a robotic synthesis platform: predict → synthesize → assay → update. Each round of real wet-lab experiments fine-tunes the model, which then proposes smarter candidates for the next round. The lab isn't just validating AI predictions — it's actively teaching the model, continuously. What happened when we let it run: LUMI-lab autonomously synthesized and screened 1,700+ ionizable lipids in human bronchial epithelial cells. The top candidate — LUMI-6 — features a brominated lipid tail, a structural motif that had been largely overlooked in LNP design. LUMI found it without being told where to look. When formulated into LNPs and delivered intratracheally to mice, LUMI-6 achieved 20.3% gene editing efficiency in lung epithelial cells — a compelling result for one of the hardest-to-reach therapeutic targets, directly relevant to diseases like cystic fibrosis and alpha-1 antitrypsin deficiency. Why this matters beyond LNPs: This is a proof of concept for a broader thesis — that foundation model pretraining + active learning + robotic experimentation can overcome the data scarcity bottleneck that plagues AI-driven discovery in biology. You don't need a massive domain-specific dataset to start. You need a model that can generalize, a lab that can generate the right data, and a loop that connects them. Huge congratulations to first authors Yue Xu, Haotian Cui, and Kuan Pang, and to the entire Bowen LI team. Grateful to our collaborators at University Health Network and Leslie Dan Faculty of Pharmacy, and to Princess Margaret Cancer Centre Research Princess Margaret Cancer Centre Research. 📄 Paper:

Bo Wang

57,625 views • 7 months ago

ARE WE ALONE? Harvard’s Avi Loeb is running the UAP Science Advisory Council, and he’s not here for the hype. Lara Trump sat down with him at Harvard to separate fact from fiction. Five major tranches of Pentagon data have now been released. Objects are still showing up over military installations and strategic sites. Loeb’s take is straightforward: “If we see objects hovering over strategic assets and we don’t know what they are, we need to figure them out. That’s common sense.” He reviewed the newly declassified videos with us. Some show objects moving in ways that don’t match known drones or U.S. missiles. In one case where they could calculate speed using background terrain, it tracked like a missile, but no U.S. system matched the shape or performance exactly. The council’s recommendation? Deploy better sensors. Multiple cameras for distance and tracking. Spectroscopy to analyze what these orbs are made of. Instruments that can tell us composition if any material is recovered. Loeb’s team requested more than 50 specific data items. Some remain classified because of the sensors used — not necessarily the objects themselves. On the big question — are these alien? “We haven’t seen clear evidence for alien technology yet.” But he refuses to dismiss the possibility. The sun is relatively young. Most stars formed long before ours. It’s completely reasonable, he says, to assume more advanced civilizations exist. Loeb co-founded the Galileo Project, which is actively scanning the skies with telescopes (including one near the Sphere in Las Vegas). So far, nothing unexplained. His standing order to the team: “If you see anything, wake me up in the middle of the night.” His bottom line cuts through both the skeptics and the true believers: “The foundation of science is curiosity. Not claiming we already know the answer.” What are we actually seeing in our skies? The Trump administration opened the door. Loeb’s council is trying to answer it with data instead of dogma. The search is on.

Gunther Eagleman™

47,851 views • 1 month ago

We are thrilled to share our breakthrough research on "Agile Flight from Pixels without State Estimation," to be presented and live-demonstrated at #RSS2024 next week! You heard well: no state estimation means no explicit visual localization, no SLAM, no VIO, and no IMU! Paper: Video (Narrated): Last year, we demonstrated that #ReinforcementLearning (RL) policies could outperform world-champion drone-racing pilots using the same quadrotor hardware; however, unlike human pilots, these policies continuously estimated an explicit state from known gate positions, the camera feed, and inertial measurements (IMU). In this new work, we tackle the challenge of learning vision-based drone racing using an end-to-end reinforcement learning approach that eliminates the need for IMU data or explicit state estimation. Like professional pilots, we go directly from images to control commands. The training is facilitated by an asymmetric actor-critic with access to privileged information. To overcome the computational complexity during image-based RL training, we use an appropriate sensor representation, which can be efficiently simulated during training without rendering images. We achieve agile flight at speeds up to 40 km/h with accelerations up to 2 g's. Although our demonstration focuses on drone racing, we believe that our method has an impact beyond drone racing and can serve as a foundation for future research into real-world applications in structured environments. Besides the paper presentation, we will also give a live demo next Tuesday and Wednesday between and hrs at TU Delft: Reference: Ismail Geles*, Leonard Bauersfeld*, Angel Romero, Jiaxu Xing, Davide Scaramuzza "Demonstrating Agile Flight from Pixels without State Estimation" Robotics: Science and Systems (RSS), 2024. Kudos to Ismail Geles Leonard Bauersfeld Ángel Romero Jiaxu Xing! University of Zurich UZH Science UZH Space Hub Aerial Core AUTOASSESS European Research Council (ERC)

Davide Scaramuzza

27,959 views • 2 years ago

A Petri Dish Of HUMAN Brain Cells LEARN TO PLAY THE GAME DOOM! In a groundbreaking fusion of biology and silicon, scientists at Cortical Labs have taught a cluster of lab-grown human neurons to play the iconic video game Doom. Not your typical AI triumph, it’s a petri dish of actual human brain cells, reprogrammed from adult donor skin or blood samples, wired into a $35,000 biological computer called the CL1. Building on their earlier Pong demo, this new feat sees the neurons navigating hellish levels, dodging demons, and even firing shots with surprising efficiency. Programmer Sean Cole pulled it off in just a week using a Python API on GitHub, a stark contrast to the year-plus effort for Pong. Astonishingly, these organic gamers outperform GPT-4 in speed and latency, proving that even a tiny blob of human intelligence can adapt and learn in ways silicon struggles to match. The excitement is palpable: this isn’t just a gimmick; it’s a window into revolutionary medical advancements. Imagine using such bio-computers to model brain diseases, test drugs, or even restore neural functions in patients. With cloud access to CL1 rentals, developers worldwide can experiment, accelerating discoveries that could redefine neuroscience. We’re witnessing the dawn of hybrid intelligence, human biology augmented by tech, evolving beyond our wildest dreams. Yet, amid the thrill, a chill runs down my spine. What are we building here? These neurons aren’t conscious (we hope), but they’re derived from humans and exhibit learning behaviors that echo our own cognition. Echoes of The Matrix or dystopian sci-fi like the “torment nexus” from Doom novels loom large. Could this lead to ethical nightmares—exploiting bio-intelligence for warfare simulations, or worse, creating sentient systems trapped in digital hells? And the philosophical rabbit hole deepens: Is life merely nested Russian dolls (matryoshka, if you prefer) of biological smarts? We, as evolved intelligences, are now crafting our own mini-brains, layering complexity upon complexity. Are we “gods” in the making, or just the next doll in an infinite regress, destined to birth something that surpasses—and perhaps supplants, us? This experiment, detailed in HotHardware’s coverage, pushes boundaries we might not be ready to cross. It’s exhilarating proof of human ingenuity, but let’s proceed with caution lest we summon demons we can’t control and we wind up in the Petri dish?

Brian Roemmele

660,857 views • 6 months ago

The tide has turned. In the last week, this is what Sandra Bullock, Reese Witherspoon, Steven Soderbergh, and Christopher Nolan said about AI: "It’s here. We have to observe it. We have to understand it. We have to lean into it. We have to use it in a really constructive and creative way, make it our friend rather than — I mean, we have to be incredibly cautious and aware of it because there are people who will use it for evil and not good. But I do feel that there’s a place for it… it’s here. We have to just be friends in some dark way." - Sandra bullock “The AI revolution has begun, and I need to learn as much as I possibly can about AI and share it with all of you. Also, FYI: the jobs women hold are 3x more likely to be automated by AI, yet women are using AI at a rate 25% lower than men on average. We don’t want to be left behind. So…do you want to learn with me?” - Reese Witherspoon “Five years from now, we all may be going, ‘That was a fun phase.’ We may end up not using it as much as we thought we were going to. There are some people that I have absolute love and respect for that refuse to engage with it. That’s their privilege. But I’m not built that way. You show me a new tool. I want to get my hands on it and see what’s going on.” - Steven Soderbergh “I think any tool, whether AI generated, whether it's computer based, or whatever, it's all another tool for filmmakers to create with. And so as long as we have faith in our in our human beings, creating these tools now, I think the medium of film will continue to develop in exciting ways.” - Christopher Nolan

Minh Do

23,410 views • 5 months ago

A mysterious embodied AI demo has recently sparked a lot of discussion. In the video, two robots from different manufacturers with significantly different hardware architectures, the Unitree G1 and AgiBot Yuanzheng A3, are reportedly running on the same “brain.” In a complex indoor environment, they work continuously for around 10 minutes in a single uncut take, performing a series of long-horizon tasks including cleaning windows, organizing objects, resuming interrupted tasks, and cooperating with each other. What makes it even more interesting is that when one robot cannot reach a high shelf, it attempts to use a box to solve the problem. The two robots also appear capable of cooperating based on each other’s physical capabilities. If the claimed level of autonomy and the use of the same model across different embodiments are eventually verified, I think there are three things that really deserve attention: 1. Cross-embodiment generalization. If the same foundation model can operate two substantially different robotic platforms, it could mean that robotic “intelligence” is gradually becoming decoupled from a specific physical body. 2. Long-horizon closed-loop execution. Continuously performing complex tasks for 10 minutes, while being able to pause, switch tasks, and later resume previous ones, is much more meaningful than completing a single 10-second demo. 3. Collaboration and dynamic planning. The two robots appear able to adjust their behavior according to environmental changes and each other’s physical capabilities. These are some of the characteristics that truly general-purpose embodied intelligence will eventually need. That said, I would remain cautious for now. The video demonstrates extremely impressive behavior, but stronger claims such as “self-evolution,” “true understanding of the physical world,” or overturning the Scaling Law with only dozens of hours of training data still require much stronger evidence. Failure recovery and replanning during a task also do not automatically demonstrate that the model is learning by itself. So I wouldn’t call this the “ChatGPT moment” of embodied AI yet. But if the team later discloses the model architecture, training data scale, level of human intervention, and can repeatedly reproduce these capabilities in completely unfamiliar environments, this seemingly rough 10-minute video could become one of the most memorable embodied AI demos of 2026. For now, my biggest question is simple: Who is the mysterious team behind it?

Ice Universe

28,209 views • 27 days ago

Data has always been the bottleneck for physical AI in self driving and robotics. Tesla is taking two very different approaches for FSD and Optimus. Tesla’s Optimus Training Playbook: 1. Build 30k Optimus Gen 3 robots 2. Operate them in a mock environment where they can perform self-play “Optimus Academy” 3. Train in sim using the real robot data to close sim2real gap. Tesla FSD Training Playbook: 1. Sell millions of cars outfitted with cheap cameras 2. Collect diverse real world driving data (especially intervention and failure recovery data) for free as a byproduct of customers driving the cars. 3. Use driving data to train Autopilot/FSD and deploy policies incrementally as a supervised FSD product 4. Repeat until policy reaches robust unsupervised full self driving for robotaxi launch. The Tesla FSD playbook is a beautiful self-funding, customer subsidized, diverse real world data flywheel. The Optimus playbook is the opposite and shares none of the beautiful attributes of the FSD training flywheel that made FSD successful. The key differences: 1. Instead of having customers pay you for vehicles, Tesla will need to fund 30,000 Optimus robots. Assuming the current landed cost per unit is $100k, that will be $3B to build plus another ~30% per year for maintenance labor and spare parts given it’s still an unhardened pre-production prototype is another $900M per year. For reference, Tesla’s GAAP net income in 2025 was $3.8B. 2. Instead of having customers drive their Teslas on roads all across the world giving Tesla an insanely rich and diverse dataset that Waymo and other AV companies could never collect, the Optimus Academy is doing the equivalent of building a fake town in a parking lot and driving their car in that parking lot. No matter how real you try to make the environments for self-play you can never replicate the diversity, complexity and failure modes of the real world. Data collected in staged environments produces demo-grade policies and will not be rich enough to generalize to the vast diversity of environments, tasks, objects, etc. out of distribution. 3. Instead of having customers collect real world failure recovery data (DAgger style) for free every time FSD disengages, the Optimus Academy will need paid teleoperators or onsite operators to collect the recovery data. Assuming 1 person can manage 2 robots to start that would cost $3.5B in labor per year (30,000 robots, $40/hr fully loaded, 16 hrs/day, 365 days per year, 2:1 robot:operator). Tesla can come up with the money to do this but money doesn’t solve the “mock data” problem. Given the higher degrees of freedom in humanoids vs. cars, training a generalized humanoid will be harder and require more data than a self-driving vehicle. The best way to train your robot is by deploying them in the diverse real world, subsidized by real customer operations. Humanoids face a chicken and egg where it’s very hard to bootstrap your way to a first policy that’s good enough to deploy in real production environments. This is an extremely capital intensive playbook (which doesn’t even include cost of training). Time will tell if it works but a better playbook would be finding a way to copy the FSD playbook.

Simon Kalouche

34,671 views • 7 months ago

Bret Weinstein on the Melania Trump AI teachers: "I get it. And it’s not that it is impossible to imagine robotic teachers doing an excellent job, but it is stunning to watch a sophisticated person fail to recognize what happens when you think that that’s what you’re going to produce, and you set it in motion. Let me point out that Wikipedia has many of the advantages that Melania is describing in this video. It is completely democratizing of knowledge, such that it doesn’t matter where on e arth you are. If you have an internet connection, you’ve got Wikipedia. It’s like an extension of your own mind, and it will make us all brilliant. Now, of course, that didn’t happen, did it? Wikipedia is a hellscape of misinformation, much of it targeted based on a political agenda. We are less certain of what we know, and less capable of reasoning on our own. Now, that doesn’t all come from Wikipedia, but my point is the promise of Wikipedia was not realized. And what we got instead is arguably worse than what we had before it was invented. The same thing is virtually guaranteed here, because you’re talking about not only the capability of educating students using a robot that has vastly more knowledge than a human teacher would, but you’re talking about the irresistible opportunity to capture those minds and steer them in one direction or another, whether that’s political or economic. The idea that these robotic teachers are going to be immune to the kind of flights of fancy that have ruined teaching in the modern era is preposterous. In fact, they will likely be even more easily steered. I would caution everyone to simply realize the distinction between complicated systems and complex systems. AI is a complex system. Human beings are complex systems. And any time you intervene in these systems, thinking you know what’s going to happen, you’re going to be embarrassed by the discovery of the unintended consequences that will come to dominate your project. As much as I like the idea of smarter, wiser, more empathic teachers, and as much as those possibilities do exist in the space of AI, we are still at a very early point in this revolution, and anybody who thinks they can predict it with this kind of precision is actually a hazard."

The DarkHorse Podcast

48,380 views • 5 months ago

My crew exposed $100,000,000 in California fraud in one day. But it’s so much darker than stolen taxpayer cash. California is using that money to rig American elections. Here’s How: California leverages their homeless crisis to funnel billions of federal dollars into the state to help homeless Americans. But Whistleblowers tell us +60% of the ‘homeless services’ are actually going to criminal aliens. These illegals live rent-free in long term homeless housing, drive luxury cars and live lavish lifestyles. The shelters guard the criminal aliens from ICE deportations. But why pack your state with homeless illegals? One word: Power. California population loss is set to cost the state major electoral votes in the next election. The only remedy is to pack more bodies into the state ahead of the 2030 census. The census counts persons, not citizens. The federally funded ‘free homeless services’ act as magnets for criminal aliens who fraudulently inflate population numbers. There are 2.5 million illegals in California alone. Finally, California makes it a criminal act to conduct citizenship tests at the shelters, hospitals or the voting booth — and now the fraud is complete. California leaders are trafficking in human misery, intentionally destroying the once great state for power — and using federal tax dollars to do it. This must be stopped. Now. Share this video like wildfire! Americans are getting really sick and tired of watching our hard work and taxes go to pay foreign fraudsters to rig elections for corrupt politicians. If the Trump Administration cut off the funds, the California homeless industrial complex would collapse tomorrow. Time for them to pay.

Benny Johnson

2,050,482 views • 7 months ago

Chinese robotics company Astribot released their latest World-Action Model (WAM), Lumo-2. Technical breakdown: - based on a frozen 🥶 Qwen-3.5 4B VLM - trained in 3 progressive stages: 1. Action is aligned with latent world dynamics (an abstract representation of action). Real-world actions are anchored to physical constraints, while the latent space is guided to focus on motion-relevant changes. This bidirectional relationship makes the model physically grounded -> critical for a world model. 2. Action is aligned with vision and language. Reusing the vision backbone and action encoder from the frozen VLM, the authors add a custom vocabulary (for new actions), a semantic module, an action decoder, and an action projector. This aligns the (new) action representations with the (existing) vision-language semantic space. Most importantly: it builds a direct mapping from natural-language instructions to motor execution. 3. End-to-end training on language, video, and robot data. Only the new modules (everything outside the frozen backbone) are trained end-to-end across temporal reasoning, physical understanding, long-horizon, and dexterous manipulation. At the end of the day, Lumo-2 is not the best on benchmarks, but that's not the point. What's genuinely new: - a way to combine latent world modeling and action generation through progressive alignment - a physically-grounded latent dynamics space - it lifts performance on unseen objects using un-annotated human egocentric video + Vision Pro captures, no special transfer algorithm needed Why it matters: - the whole model is thin trainable adapters (semantic module, action decoder/projector) on a frozen 4B backbone (cheap) - that scale is suited for real-time embedded inference (~2.71× decode speedup, no accuracy loss) - its real moat is long-horizon execution, where the added temporal memory pays off far more than on any other task As a result, this robot can now make your latte (5x sped up video):

Léo

32,331 views • 2 months ago

Sharing a super simple, user-owned memory module we've been playing around: nanomem The basic idea is to treat memory as a pure intelligence problem: ingestion, structuring, and (selective) retrieval are all just LLM calls & agent loops on a on-device markdown file tree. Each file lists a set of facts w/ metadata (timestamp, confidence, source, etc.); no embeddings/RAG/training of any kind. For example: - `nanomem add ` starts an agent loop to walk the tree, read relevant files, and edit. - `nanomem retrieve ` walks the tree and returns a single summary string (possibly assembled from many subtrees) related to the query. What’s nice about this approach is that the memory system is, by construction: 1. partitionable (human/agents can easily separate `hobbies/snowboard.md` from `tax/residency.md` for data minimization + relevance) 2. portable and user-owned (it’s just text files) 3. interpretable (you know exactly what’s written and you can manually edit) 4. forward-compatible (future models can read memory files just the same, and memory quality/speed improves as models get better) 5. modularized (you can optimize ingestion/retrieval/compaction prompts separately) Privacy & utility. I'm most excited about the ability to partition + selectively disclose memory at inference-time. Selective disclosure helps with both privacy (principle of least privilege & “need-to-know”) and utility (as too much context for a query can harm answer quality). Composability. An inference-time memory module means: (1) you can run such a module with confidential inference (LLMs on TEEs) for provable privacy, and (2) you can selectively disclose context over unlinkable inference of remote models (demo below). We built nanomem as part of the Open Anonymity project ( but it’s meant to be a standalone module for humans and agents (e.g., you can write a SKILL for using the CLI tool). Still polishing the rough edges! - GitHub (MIT): - Blog: - Beta implementation in chat client soon: Work done with amazing project co-leads Amelia Kuang Coco Xu Erik Chi !!

Ken Liu

75,023 views • 5 months ago