🚀 Excited to share our #ICLR2025 work on planning... with neural dynamics models! While our lab has developed diverse neural dynamics models for manipulating rigid, deformable, and granular objects, having the model alone doesn’t solve the problem—planning with it remains a challenge. 💡 Enter BaB-ND, led by Keyi and Jiangwei! We propose a scalable, GPU-accelerated branch-and-bound algorithm, inspired by neural network verification, to enable effective planning for diverse objects modeled with neural dynamics. 🔗 Project page (open-source + detailed docs!): 🎥 Watch the video to see T being pushed around obstacles, and check out Keyi’s thread for more details!show more

Yunzhu Li
10,561 просмотров • 1 год назад
NeuRBF: A Neural Fields Representation with Adaptive Radial Basis... Functions paper page: present a novel type of neural fields that uses general radial bases for signal representation. State-of-the-art neural fields typically rely on grid-based representations for storing local neural features and N-dimensional linear kernels for interpolating features at continuous query points. The spatial positions of their neural features are fixed on grid nodes and cannot well adapt to target signals. Our method instead builds upon general radial bases with flexible kernel position and shape, which have higher spatial adaptivity and can more closely fit target signals. To further improve the channel-wise capacity of radial basis functions, we propose to compose them with multi-frequency sinusoid functions. This technique extends a radial basis to multiple Fourier radial bases of different frequency bands without requiring extra parameters, facilitating the representation of details. Moreover, by marrying adaptive radial bases with grid-based ones, our hybrid combination inherits both adaptivity and interpolation smoothness. We carefully designed weighting schemes to let radial bases adapt to different types of signals effectively. Our experiments on 2D image and 3D signed distance field representation demonstrate the higher accuracy and compactness of our method than prior arts. When applied to neural radiance field reconstruction, our method achieves state-of-the-art rendering quality, with small model size and comparable training speed.show more

AK
194,469 просмотров • 2 лет назад
Here are more results from #RigidFormer: predicting physical dynamics... with purely neural simulators — an attempt to learn physical dynamics in a scalable manner. 🤖 1) Controllable Articulated Body Simulation — More Results Additional Unitree G1 humanoid rollouts under controlled motion. Each sample uses a different initial state and control signal (direction and velocity). 🏺 2) Object Fragmentation Simulating the cracking and fragmentation process of objects. Thanks Žiga Kovačič for suggesting this experiment! 🎬 3) Combining Rigidformer with Diffusion-as-Shader for controllable video generation. Note: the meshes shown here are only for visualization — the network takes point clouds as input and predicts the updated state of each point.show more

Zhiyang (Frank) Dou
20,955 просмотров • 2 месяцев назад
🚀 First demonstration of learning-accelerated trajectory optimization in space!... Our team used a neural network to warm-start trajectory optimization for the NASA's Astrobee free-flying robot on-board the International Space Station, cutting solver iterations by up to 60% while maintaining safety constraints. 🎥 Video: 📄 Paper: 📅 Submitted to iSpaRo 2025 With Somrita Banerjee and Abhishek Cauligishow more

Marco Pavone
11,154 просмотров • 1 год назад
Excited to share a few presentations, demos, and workshop... talks from our group and collaborators at #ICRA2026! We will present recent work on real-to-sim-to-real robot policy evaluation, model-based planning with learned dynamics, and multi-modal manipulation. We will also have a joint live demo between SceniX and Analog Devices, Inc. on real-to-sim-to-real cable manipulation at the ICRA exhibition. This is a small teaser of what we have been building, with more to come soon! If you are at ICRA, please stop by the sessions or the demo booth. Happy to chat about robot learning, simulation, world models, and sim-to-real!show more

Yunzhu Li
11,026 просмотров • 2 месяцев назад
World modeling and imitation learning have largely been considered... two disparate worlds. In our recent work, Unified World Models, just accepted to #RSS2025, Chuning Zhu provides a dead-simple unifying solution: just train a joint diffusion model over actions and future states, but with *decoupled* diffusion time steps across these modalities. Manipulating these decoupled time steps then allows for marginalization or conditioning on actions or states; a single model can serve as a policy, forward dynamics model, video prediction model, or inverse dynamics model by simply setting diffusion timesteps carefully. The resulting model can leverage video datasets along with robot training data much more effectively, and shows improved robustness, generalization, and flexibility. This is exciting because it is frustratingly simple, scalable, and shows strong improvement on real-world robotics problems. Please refer to Chuning Zhu 's excellent thread for more details! More details/code can be found on our website and in the paper -show more

Abhishek Gupta
11,430 просмотров • 1 год назад
Introducing Attio Objects 🚀 We know how hard... it is to find a CRM that fits your unique business model. That's why we built Attio Objects – our powerful data model with custom objects that gives you complete flexibility to structure your CRM exactly how you need it. Along with custom objects, we've also introduced new standard objects: - Workspaces and Users objects for PLG businesses. - A robust Deals object for sales-driven companies. This is the culmination of a 4-year effort, with 3 years of work put in even before launching Attio. Since day one, we've been determined to solve the fundamental problem in the CRM space: the trade-off between power and time-to-value. If you wanted power and flexibility, your CRM would take forever to build and not work well with your stack. If you wanted speed, you'd need to use highly opinionated, inflexible software that doesn't really work for your business. That ends today. With Attio, you no longer have to compromise. Build your CRM your way, fast. Iterate as you grow. High-growth startups like Replicate, , and Modal and more are already using Attio's object architecture to perfectly match their businesses and accelerate their growth. To get all the details, check out our blog post 👇 show more

Attio
26,821 просмотров • 2 лет назад
Depth Any Video with Scalable Synthetic Data AI physicists... and chemists continue to make strides in depth estimation from video. Check out this new paper featuring some impressive examples. See the thread for more details (unfortunately no code yet). Abstract: Video depth estimation has long been hindered by the scarcity of consistent and scalable ground truth data, leading to inconsistent and unreliable results. In this paper, we introduce Depth Any Video, a model that tackles the challenge through two key innovations. First, we develop a scalable synthetic data pipeline, capturing real-time video depth data from diverse game environments, yielding 40,000 video clips of 5-second duration, each with precise depth annotations. Second, we leverage the powerful priors of generative video diffusion models to handle real-world videos effectively, integrating advanced techniques such as rotary position encoding and flow matching to further enhance flexibility and efficiency. Unlike previous models, which are limited to fixed-length video sequences, our approach introduces a novel mixed-duration training strategy that handles videos of varying lengths and performs robustly across different frame rates 0 - even on single frames. At inference, we propose a depth interpolation method that enables our model to infer high-resolution video depth across sequences of up to 150 frames. Our model outperforms all previous generative depth models in terms of spatial accuracy and temporal consistency.show more

MrNeRF
27,428 просмотров • 1 год назад
Introducing Kaleido💮 from AI at Meta — a universal... generative neural rendering engine for photorealistic, unified object and scene view synthesis. Kaleido is built on a simple but powerful design philosophy: 3D perception is a form of visual common sense. Following this idea, we formulate rendering purely as a sequence-to-sequence generation problem, successfully unifying neural rendering with the architecture principles behind modern language and video models. Unlike traditional neural rendering methods, Kaleido learns 3D purely in a data-driven way, without explicit 3D representations or structures. It acquires spatial understanding directly through large-scale video pretraining, then multi-view 3D data finetuning, inspired by how LLMs acquire textual common sense from large corpora before specialising in domains like coding. Through extensive ablations, we progressively modernised the architecture design and training strategies and tackled key scaling challenges in sequence-to-sequence generative rendering, arriving at a design that’s simple, versatile, and scalable. Kaleido significantly outperforms prior generative models in few-view settings, and remarkably is the first zero-shot generative method matches InstantNGP-level rendering quality in multi-view settings. We view Kaleido also as an alternative step towards world modeling that flexibly spans a spectrum of “realities": with many views, it faithfully reconstructs grounded reality; with fewer views, it imagines plausible unseen details. 🔗 Explore more results and paper:show more

Shikun Liu
22,389 просмотров • 10 месяцев назад
🚀Thrilled to share what we’ve been building at TRI... over the past several months: our first Large Behavior Models (LBMs) are here! I’m proud to have been a core contributor to the multi-task policy learning and post-training efforts. At TRI, we’ve been researching how LBMs can help robots learn faster, better, and more efficiently. The key takeaways: ✅ We built an evaluation pipeline to benchmark LBM performance with real 𝐬𝐭𝐚𝐭𝐢𝐬𝐭𝐢𝐜𝐚𝐥 𝐜𝐨𝐧𝐟𝐢𝐝𝐞𝐧𝐜𝐞 ✅ Pre-training on hundreds of tasks makes models more robust—plus, we can teach new, complex tasks with 80% 𝐥𝐞𝐬𝐬 𝐝𝐚𝐭𝐚 ✅ The bigger and more diverse the pre-training, the better the results Check out our overview video, webpage and paper for more details: ✨ 🌎 📄 We hope this work helps move the field of robotics forward!show more

Zubair Irshad
20,377 просмотров • 1 год назад
🚀 To our valued investors, thank you for your... patience and continued support! Several months ago, the Virexit Team identified a shift in market dynamics, which we believe has opened up several opportunities that required us to act fast and decisively, and we’re thrilled with the steps we’ve taken to date. 💥 Over the past few months, we've been evaluating, planning, and executing on exciting opportunities that we are confident has positioned us for a very bright future.🌟 This includes restructuring, optimizing team roles, adding seasoned consultants, and expanding the scope of our products and markets to secure what we believe is a prosperous path forward both in the short, and long term. 📈 We’re excited to share more updates, PR's, 8k's when appropriate in the coming weeks—stay tuned! 📰 #StrategicGrowth #DecisiveAction #InvestInSuccessshow more

Lavish Enterprises, Inc
13,720 просмотров • 1 год назад
After 5 months of planning, our first Live2D Jakarta... Chapter event was a huge success! We brought an interactive Live2D showcase to the biggest anime con in Indonesia, featuring over 20 finest models created by our local artists and riggers. This showcase also included a custom-made plugin for VTube Studio that adds touch interactivity to Live2D models, allowing visitors to interact with VTuber models in a more engaging way! We hope this demo could open up new potential for Vtubers in IRL events. Special thanks to Live2D Inc. for trusting me to host this event! The showcase attracted many curious visitors and gave us the opportunity to introduce this tech to a wider audience, and we hope we are opening up new opportunities for everyone. Look forward to our next one! (Model: Key Oriesa by karamomo🍑 | Live2D Rigging | Food Illustration)show more

Ran 🐲 Live2D Animator
165,673 просмотров • 2 месяцев назад
Introducing ASAL: Automating the Search for Artificial Life with... Foundation Models Artificial Life (ALife) research holds key insights that can transform and accelerate progress in AI. By speeding up ALife discovery with AI, we accelerate our understanding of emergence, evolution, and intelligence–core principles that can inspire the next generation of AI systems! We proudly collaborated with MIT, OpenAI, Swiss AI Lab IDSIA, and Ken Stanley on this exciting project. Full Paper (Website): Full Paper (arxiv): Code: In this work, we propose a new algorithm called Automated Search for Artificial Life (“ASAL”) to automate the discovery of artificial life using vision-language foundation models. Instead of tediously hand-designing every tiny rule of an Alife simulation, simply describe the space of simulations to search over, and ASAL will automatically discover the most interesting and open-ended artificial lifeforms! Because of the generality of foundation models, ASAL can discover new lifeforms across a diverse range of seminal ALife simulations, including Boids, Particle Life, Game of Life, Lenia, and Neural Cellular Automata. ASAL even discovered novel cellular automata rules that are more open-ended and expressive than the original Conway’s Game of Life. We believe this new paradigm may reignite ALife research by overcoming the bottleneck of manually designed simulations, thus advancing beyond the limits of human ingenuity.show more

Sakana AI
751,029 просмотров • 1 год назад
✨ Excited to share QVQ-Max, our visual reasoning model... that's still evolving We've been experimenting with this approach for a while - try it out on Qwen Chat! ( 🚀 Just upload any image or video, ask away, and hit the "Thinking" button to see how it processes visual information step-by-step. It's a work-in-progress but fascinating to watch! Your early feedback will be super helpful as we continue developing! 🙏 Blog:show more

Qwen
147,074 просмотров • 1 год назад
Predicting the next word "only" is sufficient for language... models to learn a large body of knowledge that enables then to code, answer questions, understand many topics, chat, and so on. This is clear to many researchers now, and there are nice tutorials on why this works by Ilya Sutskever resorting to compression ( ) and by Geoffrey Hinton ( ). However, the emergence of types of understanding is not unique to language models. In by Misha Denil and Brandon Amos the authors trained models to predict the next few time stems of over a hundred robot hand sensors (Touch, Gyro, Accelerometer, Joint Info, Actuator Info, etc.). They ten found out that they could regress the shape of the thing the hand was touching from the activations of the neural networks using probes. That is, the model developed an internal representation of shapes even though it was simply used to predict "only" the next few senses. Awareness follows from simple predictions and interaction with the world.show more

Nando de Freitas
134,252 просмотров • 2 лет назад
🎉 How D-lightful! 🎉 Chainguard has raised a $356M... Series D round co-led by new investor Kleiner Perkins and existing investor IVP, with participation from Salesforce Ventures, Datadog, Inc. Ventures, and every single other existing investor. This brings our total funds raised to over $600M, and puts us at a $3.5 billion valuation. We’re on a mission to be the safe source for open source and build a future where security and innovation can move in lockstep. This new investment allows us to continue creating secure, easy to use products and solutions that enable that future for our users. The momentum we’ve built over the past few years has been incredible to experience, and we’re grateful to all our customers, investors, and community for supporting us and being along for the ride. 🫶 Reach out if you are interested in joining us on this journey—we are hiring in every department! Check out our CEO and co-founder Dan Lorenc's blog:show more

Chainguard ⛓️
41,253 просмотров • 1 год назад
A Letter to Our Community: The Road Ahead for... Robotics To our Community and Partners, As we step into 2026, our mission at Axis is clearer than ever: Constructing the definitive End-to-End Scaling Layer for Robotics. Our goal is to accelerate the transfer of diverse human intelligence into Robotics General Intelligence (RGI). By owning the critical path of intelligence creation, we are turning the physical limitations of robotics into a scalable, software-driven future. Here is our strategic outlook and roadmap for the year ahead. The Core Thesis: Simulation is the Only Way Out The path to RGI is currently blocked by Data Scarcity, Generalization Fragility, and Hardware Fragmentation. At Axis, we believe Simulation is the only way out. Our Simulation Data Platform and Data Augmentation Engine transform raw data into "Synthetic Gold". Backed by academic milestones like Roboverse, Skill Blending, and GraspVLA, we have proven that pure simulation can achieve the generalization required for the real world. We don’t just collect data; we architect it. The Engine: Why Crypto? We believe RGI should come from all, not a few. Crypto is not just a feature; it is the primitive that powers our entire ecosystem flywheel: - Incentive Mechanism: Democratizing contribution and rewarding the trainers and developers. - Assetization: Turning proprietary data and refined models into liquid, ownable assets. - Verifiable Workflow: We are opening the "Black Box" of AI. By bringing total transparency to the Task Generation → Data Collection → Model Training pipeline, we ensure every byte of intelligence is verifiable, traceable, and secure. 2026 Strategic Deliverables This year, we are committed to delivering three foundational pillars: - The World's Largest Training Dataset for Robots: A robot training set—diverse, high-quality interaction data at an unprecedented scale. - A Robotics Foundation Model: A universal robotic brain trained on our pure simulation and synthetic data, capable of robust cross-embodiment transfer and open-world adaptability. - Evolvable Robot Hardware: Robots deployed with Axis models that autonomously evolve through continuous interaction, turning every deployment into a self-improving node within our RGI network. The Ultimate Vision We are building more than models; we are architecting the Distributed Machine Economy. A future where every dataset, model, and robotic embodiment is a verifiable asset in a global, autonomous network. Thank you for building the future of intelligence with us✌️📷show more

Axis Robotics
27,858 просмотров • 7 месяцев назад
🚀 We introduce Neural Theorizer (NEO) — a new... type of world model that learns to theorize the world from observation, without language or LLM supervision. Selected as an ICML 2026 oral presentation — 0.7% of submitted papers. The paper asks: "What does it mean to understand the world and build a world model?" Today’s world models are often trained to predict the future: the next frame, next latent state, or next observation. But is prediction enough? We argue that a world model should be a theory-building system: one that discovers reusable primitives, composes them into executable explanations, and transfers those explanations to novel phenomena. NEO is our first step toward this vision — a World Theory Model that learns explicit, compositional theories from raw observation. This work was led by my wonderful students: Doojin Baek*(Doojin Baek), Gyubin Lee* (GyuBin Lee), Junyeob Baek (Junyeob Baek), and Hosung Lee (Hosung Lee). For more details, take a look at the paper — and if you’re attending ICML, let’s talk there! 📄 arXiv: 🌐 Project page:show more

Sungjin Ahn
99,136 просмотров • 1 месяц назад
Chop the gradients ✂️! We found that truncating decoder... gradients in latent video diffusion to a fixed window allows us to finetune on videos with pixel-wise perceptual losses without running out of memory. Pixel losses have been essential for image generation and reconstruction, but until now, they haven't scaled to long-duration, high-resolution video diffusion due to recursive activation accumulation in causal decoders, leading to OOM during training 💥📉. Project: Video diffusion models can do a lot more 🚀 when you can backprop the decoder! Post-process neural rendered scenes, super-resolve videos, harmonize lighting in controlled synthetic driving scenes, and inpaint videos — all in a single step ⚡ with a quick finetune from a standard diffusion model.show more

Felix Heide
28,399 просмотров • 4 месяцев назад