Загрузка видео...

Не удалось загрузить видео

На главную

Generalist AI has just broken the "impossible triangle" of speed, reliability, and intelligence in robotics. They’ve pushed their embodied foundation model forward in a big way: GEN-1 hits “Mastery” on simple physical tasks,99% average success rate, and it can improvise when things go wrong. Smooth, natural, fast(factory owners must...

25,519 просмотров • 5 месяцев назад •via X (Twitter)

Комментарии: 0

Нет доступных комментариев

Здесь появятся комментарии из оригинального поста

Похожие видео

My conversation with Sergey Levine (Sergey Levine). Sergey is the co-founder of Physical Intelligence -- a company building foundation models that can control any robot to do any task in any environment. The company's thesis is that generality is more scalable than specialization, meaning that a model trained across many different robots and tasks will ultimately outperform any system built to do one thing well (eg, just wash dishes). Sergey is a researcher by background, but I think you will appreciate how practical and commercially grounded this conversation is. We discuss: - Why changing a diaper will be the last task a robot masters - The simulation v. real-world data debate - How multimodal LLMs give robots common sense - Moravec's Paradox + Robot Olympics - Why robots can do long-horizon tasks now - A realistic timeline for robots in our homes I should note that I am an investor in Physical Intelligence -- I made the investment because I believe it is one of the most important companies tackling the problem of robotics. Enjoy! Timestamps: 0:00 Intro 2:39 Defining Physical Intelligence 5:19 The Challenge of Building General Models 6:34 The Stakes and Future of General Purpose Robotics 8:15 Pros and Cons of Humanoid Robots 10:12 Historical Milestones in Robotics Research 15:31 Combining Generative AI and Deep RL 21:24 Moravec's Paradox 25:33 Kitchen Robots 29:30 Simulation vs. Real-World Data 30:48 The Robot Olympics 36:31 The Physiological Reality of Embodiment 38:56 Controversies in the Robotics Community 44:18 What Makes a Great Researcher 48:27 How Businesses Should Prepare for Robotics 54:09 Tracking Progress Through Research Papers 57:02 The Next Step: Mid-Level Reasoning 1:02:00 The Kindest Thing

Patrick OShaughnessy

134,397 просмотров • 5 месяцев назад

Karol Hausman is the co-founder and CEO of Physical Intelligence, a robotics company building a general-purpose “AI brain for the physical world.” The company has raised more than $1 billion in funding to develop foundation models that allow robots to operate across many machines, environments, and tasks rather than being programmed for a single purpose. In our conversation, we explore: • The moment a lecture from Sergey Levine convinced him to abandon his PhD research direction and pivot fully to deep learning • The case for building a general “AI brain” for the physical world rather than a single specialized robot • The role of real-world data in training robots, the limits of simulation, and how deployment could create a powerful data flywheel • The unique challenges of physical intelligence and why robots must operate with far higher reliability than language models Thank you to the partners who make this possible - Brex: The intelligent finance platform: - Granola: The app that might actually make you love meetings: Timestamps (00:00) Intro (04:05) Karol’s early fascination with robots (18:21) Karol’s entry point to robotics and PhD program (25:49) Combining robotics with LLMs: The Taylor Swift demo (30:48) The 1970s SHRDLU AI experiment (39:40) How research shapes what Physical Intelligence builds (49:07) The return of reinforcement learning in robotics (1:00:00) NVIDIA’s simulation engines (1:07:31) Compensating for missing senses

Mario Gabriele 🦊

27,871 просмотров • 6 месяцев назад

Excited to announce GR00T N1, the world’s first open foundation model for humanoid robots! We are on a mission to democratize Physical AI. The power of general robot brain, in the palm of your hand - with only 2B parameters, N1 learns from the most diverse physical action dataset ever compiled and punches above its weight: - Real humanoid teleoperation data. - Large-scale simulation data: we are open-sourcing 300K+ trajectories! - Neural trajectories: we apply SOTA video generation models to “hallucinate” new synthetic data that features accurate physics in pixels. Using Jensen’s words, “systematically infinite data”! - Latent actions: we develop novel algorithms to extract action tokens from in-the-wild human videos and neural generated videos. GR00T N1 is a single end-to-end neural net, from photons to actions: - Vision-Language Model (System 2) that interprets the physical world through vision and language instructions, enabling robots to reason about their environment and instructions, and plan the right actions. - Diffusion Transformer (System 1) that “renders” smooth and precise motor actions at 120 Hz, executing the latent plan made by System 2. We deploy N1 on GR1 robot, 1X Neo robot, and a large collection of simulation benchmarks. N1 achieves up to +30% boost in diverse manipulation tasks for household and industrial settings. While humanoid robots are the main focus of N1, our model also supports cross-embodiment. We finetune it to work on the $110 HuggingFace LeRobot SO100 robot arm! Open robot brain runs on open hardware. Sounds just right. Let’s solve robotics, together, one token at a time. Links to our Whitepaper, Github repo, HuggingFace model, and open dataset page in the thread: 🧵

Jim Fan

467,237 просмотров • 1 год назад

A policy that teaches robot hands to touch things the way humans do... not just grab and move, but feel and adjust in real time. Robot manipulation research often stops at picking up objects and placing them. CGP goes further: it handles tasks like opening jars, flipping objects in-hand, wiping dishes, and grasping fragile eggs, the kind of dexterous, contact-rich skills that require constant micro-adjustments based on what the fingers are actually feeling. The robot doesn't just see what it's doing; it predicts what contact should feel like at each step, then checks whether reality matches the prediction. If a finger is slipping, the policy knows before the object drops. Works on real robot hands (both 4-finger and 5-finger designs) with tactile sensors embedded in the fingertips Robust to visual distractions! The robot keeps flipping a box correctly even when the camera view is disrupted, because it's grounding decisions in touch, not just vision. Baseline policies without contact grounding fail in predictable ways: slipping mid-task, incomplete motions, loss of grasp, CGP avoids these This is a meaningful step toward robots that can handle the physical world with the kind of reliable, adaptive grip that humans take for granted. Relevant for manufacturing, logistics, assistive robotics, and anywhere fragile or irregular objects need to be handled carefully. Published at RSS 2026, developed with Meta Reality Labs Research. Thanks for sharing, Zhengtong Xu / Zhengtong Xu ——- Weekly robotics and AI insights. Subscribe free:

Ilir Aliu

12,854 просмотров • 3 месяцев назад