Video wird geladen...

Video konnte nicht geladen werden

Zur Startseite

๐Ÿš€Top embodied intelligence competition is back! AGIBOT WORLD CHALLENGE @ ICRA 2026 Reasoning-Action & World Model tracks $530K prize pool Robot finals at ICRA 2026 Win AGIBOT robot purchase vouchers Global teams competing Apply now!

12,154,864 Aufrufe โ€ข vor 5 Monaten โ€ขvia X (Twitter)

0 Kommentare

Keine Kommentare verfรผgbar

Kommentare vom Original-Post werden hier angezeigt

ร„hnliche Videos

๐Ÿ”ฅ JUST IN: Open-source robotics dataset from 100% real-world scenarios! ๐Ÿคฏ Chinese robotics company AGIBOT just released AGIBOT WORLD 2026, an open-source dataset systematically covering key embodied AI research directions. Built entirely from real-world environments: commercial spaces, and homes. Collected using AGIBOT G2 robots in free-form collection mode, providing structured, accurately annotated, high-quality data. Digital twin technology creates 1:1 scale replicas in simulation matching the real environments. Both real-world and simulation data are open-sourced. The AGIBOT G2 platform collects multiple data types simultaneously: RGB(D) cameras, tactile sensors, force sensors, LiDAR, IMU, and full-body joint states. Whole-body control coordinates arms, waist, and hands for complex tasks. First-person teleoperation lets operators control the robot from its perspective. The tasks covered are fine-grained manipulation, ultra-long-horizon tasks, spatial navigation, dual-arm coordination, and multi-agent/human-robot collaboration. The dataset includes error-recovery trajectories with annotations. Most datasets only show successful demonstrations. AGIBOT includes failures and how the robot recovers, teaching models how to handle mistakes. After collection, data is tested through policy training and real-robot deployment to ensure quality. Then processed through industrial quality control with multiple screening and cleaning rounds. Making it open-source accelerates embodied AI research by giving researchers access to high-quality real-world robot data at scale. ๐Ÿ‡จ๐Ÿ‡ณ Learn more here: ~~ โ™ป๏ธ Join the weekly robotics newsletter, and never miss any news โ†’

Lukas Ziegler

40,583 Aufrufe โ€ข vor 4 Monaten

๐—–๐—ต๐—ถ๐—ป๐—ฎ ๐—ถ๐˜€ ๐—ณ๐—ถ๐—ป๐—ถ๐˜€๐—ต๐—ถ๐—ป๐—ด ๐˜๐—ต๐—ฒ ๐—ต๐˜‚๐—บ๐—ฎ๐—ป๐—ผ๐—ถ๐—ฑ ๐—ฟ๐—ผ๐—ฏ๐—ผ๐˜ ๐—ฟ๐—ฎ๐—ฐ๐—ฒ ๐—ฏ๐—ฒ๐—ณ๐—ผ๐—ฟ๐—ฒ ๐—บ๐—ผ๐˜€๐˜ ๐—ผ๐—ณ ๐˜๐—ต๐—ฒ ๐—ช๐—ฒ๐˜€๐˜ ๐—ฟ๐—ฒ๐—ฎ๐—น๐—ถ๐˜‡๐—ฒ๐˜€ ๐—ถ๐˜ ๐—ต๐—ฎ๐˜€ ๐˜€๐˜๐—ฎ๐—ฟ๐˜๐—ฒ๐—ฑ. AGIBOT held its Partner Conference in Shanghai last week. The real headline wasn't the new hardware. It was their CTO standing on stage, telling investors that humanoid R&D season is over. 2026, he said, is "Deployment Year One." Not research. Not demos. Deployment into real factories, real warehouses, real stores. The manufacturing ramp is getting faster. 1,000 humanoid robots in the first 2 years. Another 4,000 in the next 12 months. Another 5,000 in just 3 months after that. AGIBOT is now shipping more humanoids per quarter than most US robotics companies have built in their entire existence. Then came the announcements the industry will spend the rest of the year reacting to. AIMA. The first full-stack open architecture for embodied AI. A unified robot operating system called Link-U, three dev platforms for motion, interaction, and task creation, plus an open agent framework. Any developer can build on top of it. This is the Android play for humanoids. GO-2. A vision-language-action foundation model with Action Chain-of-Thought reasoning. Planning and execution collapsed into one model. GE-2. A world model for simulation, strategy testing, and sim-to-real transfer. AGIBOT WORLD 2026. An open-source, production-grade real-world dataset pulled from actual industrial, logistics, hotel, and commercial sites. Seven standardized "productivity packages" covering logistics sorting, retail service, security patrol, commercial cleaning, and more. Plug, deploy, bill. A 5-year, $280 million commitment to seed a global developer and partner ecosystem. Now look at the competition. Boston Dynamics has been building humanoids since 1992. Tesla's Optimus is still climbing its own hype curve. Apptronik and Agility are well-funded but pre-scale on real deployments. AGIBOT has pulled all of this off in three years, with no acquisitions, no legacy platform, and no IPO distractions. While the West is still asking when humanoids will scale, China is already shipping them by the thousand.

Shruti

214,939 Aufrufe โ€ข vor 4 Monaten

X Square Robot just closed its Series C at a valuation above RMB 20 billion, about $2.8 billion ๐Ÿค– IDG came into this round. The bigger signal is the cap table. HongShan and Xiaomi were already in across earlier rounds, while Meituan, Alibaba, ByteDance, and Xiaomi have each led rounds at different stages. That puts X Square in a rare position for an embodied AI company: top-tier financial capital on one side, and four of Chinaโ€™s biggest tech platforms on the other. This is not just a money story. Meituan, Alibaba, ByteDance, and Xiaomi bring very different strategic assets: real-world scenarios, cloud infrastructure, consumer traffic, supply chains, and hardware ecosystems. The deployment side is already moving: robot home-cleaning services first, then a โ€œRobots Into Homesโ€ program with the first batch entering real households. The model stack is worth watching too. X Square has open-sourced WALL-OSS-0.5 for robot manipulation and WALL-WM for world modeling. WALL-OSS-0.5 showed strong real-robot performance without post-training, while WALL-WM uses event-level prediction to align language, vision, and action around meaningful physical-world events. They are also building a model-driven data pipeline for large-scale collection, cleaning, annotation, quality control, and augmentation. That matters because home robotics dies in the long tail: weird rooms, messy objects, bad lighting, and tasks that never look the same twice. Founded in 2023, X Square is building general-purpose embodied AI robots and foundation models for real-world environments, tying models, robot hardware, high-precision manipulation, data, and deployment into one system.

RoboHub๐Ÿค–

12,975 Aufrufe โ€ข vor 1 Monat

Beijing hosts worldโ€™s first humanoid robot games | Ariana News The worldโ€™s inaugural humanoid robot competition is underway in Beijing, drawing more than 500 robots from 280 teams across 16 countries to compete in a uniquely futuristic sporting spectacle. The three-day event, held at the National Speed Skating Ovalโ€”once the โ€œIce Ribbonโ€ of the 2022 Winter Olympicsโ€”kicked off on August 15 and runs through to Sunday August 17. The tournament features 26 events spread across athletic, performance, and scenario-based categories. Athletic challenges include sprinting, soccer, and kickboxing, while performance segments showcase robot dance routines and musical instrument displays. Real-world scenarios, such as medication sorting, cleaning tasks, and industrial material handling, are also on the agenda to test practical functionality. Organizers meanwhile emphasize the eventโ€™s role in accelerating the integration of humanoid robots into everyday life, from manufacturing and hospitality to healthcare. One Chinese official summed it up: โ€œEvery robot that participates is creating history.โ€ The competition has yielded both triumphant strides and technical stumbling blocks. In running events, the robot H1 from Unitree Robotics claimed top honors in the 1,500-meter race, demonstrating promising agility. Yet, many robots struggled with balance, coordination, and task execution, including some collapsing mid-sprint or requiring human help to standโ€”underscoring the still-developing nature of embodied artificial intelligence.

Owen Gregorian

43,158 Aufrufe โ€ข vor 1 Jahr

Chinese robotics company Astribot released their latest World-Action Model (WAM), Lumo-2. Technical breakdown: - based on a frozen ๐Ÿฅถ Qwen-3.5 4B VLM - trained in 3 progressive stages: 1. Action is aligned with latent world dynamics (an abstract representation of action). Real-world actions are anchored to physical constraints, while the latent space is guided to focus on motion-relevant changes. This bidirectional relationship makes the model physically grounded -> critical for a world model. 2. Action is aligned with vision and language. Reusing the vision backbone and action encoder from the frozen VLM, the authors add a custom vocabulary (for new actions), a semantic module, an action decoder, and an action projector. This aligns the (new) action representations with the (existing) vision-language semantic space. Most importantly: it builds a direct mapping from natural-language instructions to motor execution. 3. End-to-end training on language, video, and robot data. Only the new modules (everything outside the frozen backbone) are trained end-to-end across temporal reasoning, physical understanding, long-horizon, and dexterous manipulation. At the end of the day, Lumo-2 is not the best on benchmarks, but that's not the point. What's genuinely new: - a way to combine latent world modeling and action generation through progressive alignment - a physically-grounded latent dynamics space - it lifts performance on unseen objects using un-annotated human egocentric video + Vision Pro captures, no special transfer algorithm needed Why it matters: - the whole model is thin trainable adapters (semantic module, action decoder/projector) on a frozen 4B backbone (cheap) - that scale is suited for real-time embedded inference (~2.71ร— decode speedup, no accuracy loss) - its real moat is long-horizon execution, where the added temporal memory pays off far more than on any other task As a result, this robot can now make your latte (5x sped up video):

Lรฉo

32,296 Aufrufe โ€ข vor 1 Monat

CHINA JUST SOLVED THE PROBLEM THAT'S BEEN BREAKING ROBOT AI FOR A DECADE. and the fix wasn't a smarter model. for years, every robot AI failure got the same diagnosis. the model isn't smart enough. so everyone scaled intelligence. bigger models. more parameters. better reasoning. AGIBOT asked a different question: what if the reasoning was never the problem? there's a gap that runs through every traditional robot AI system. reasoning on one side & motor commands on the other. the brain decides but the body executes something different, because thinking and moving were never actually connected. GO-2 fixes this by reasoning INSIDE the action space, not above it. before moving, it runs a complete mental simulation of every step - like a basketball player mentally tracing the arc of a shot before releasing the ball. watch the demo and you'll see exactly what this means. the robot works through a task queue autonomously. classify toiletries. upright the drink bottle. place headphones in the leather box. mid-execution, a new instruction drops: "my phone's missing. help me find it." it doesn't pause. doesn't reset. it processes the new task and keeps moving. that's not a scripted sequence. that's real-time instruction following on top of an active task queue. that one architectural change is where the numbers come from. > #1 on LIBERO across Spatial, Object, Goal, and Long tasks โ†’ 98.5% average success > 86.6% zero-shot accuracy in active disturbance environments > 47.4 on VLABench โ†’ best-in-class on objects and textures it's never seen before > 82.9% success trained on simulation only, tested on real hardware sim-to-real is the graveyard of robotics research. models trained in simulation collapse the moment they touch the real world. 82.9% means that graveyard just got a lot smaller. it holds because of how GO-2 trains. deliberately fed imperfect reasoning conditions, then trained to execute robustly anyway. not a researcher assumption. a design decision from a team that ships hardware and knows exactly what breaks. then there's the infrastructure layer. Genie Studio. fleet-wide data collection. cloud training. online post-training in live environments. 10x improvement in training efficiency. task startup reduced to minutes. 2-4x better success rates with 50%+ less data. the model gets smarter every time a robot fails in the field. this isn't a benchmark story. it's a compounding moat. dual CVPR 2026 + ACL 2026 acceptance. computer vision AND natural language processing. top conferences. simultaneously. that doesn't happen with incremental research. the US-China robotics race has been framed as a compute race. a model quality race. it was always an execution race. the robot that wins won't be the smartest one in the lab. it'll be the most reliable one on the floor. full breakdown: is execution reliability the real bottleneck, or are we still underestimating how far reasoning needs to go?

Shruti

18,622 Aufrufe โ€ข vor 4 Monaten

Most video-action robot models are a content-creation video generator with an action module attached. LingBot-VA 2.0 from Robbyant, a video-action foundation model, throws that starting point out and trains the whole stack natively for control. And it runs closed-loop at a peak 225 Hz. It's so important because A robot cannot move responsively when its controller pauses to imagine the next few frames. LingBot-VA 2.0 predicts during execution, then corrects using each real observation. And it carries only about 13B video parameters while activating roughly 1.9B per token. Bigger robot models usually mean slower reactions, creating a direct conflict between intelligence and control. LingBot-VA 2.0 is trained from scratch for robot control rather than adapted from a video generator built for content creation. Robbyant, an embodied AI company under Ant Group, built it to learn how scenes change under actions, predict what should happen next, and turn those predictions into real-time robot movements. Most video-action systems inherit a tokenizer and video backbone trained mainly to reproduce visual appearance. LingBot-VA 2.0 rebuilds both parts around physical control. Its semantic visual-action tokenizer maps observations toward features from a frozen vision foundation model and learns compact latent actions from frame-to-frame changes using self-supervised inverse and forward dynamics. Unlabeled web video can therefore carry action-relevant training signals without robot action labels. The policy is causal from the start, so every prediction can use only past observations. Its sparse Mixture-of-Experts video backbone has about 13B total parameters, while about 1.9B are active per token, keeping the compute lower during each step. A high-level vision-language planner breaks long tasks into smaller instructions, while the low-level video-action policy handles continuous movement. Foresight Reasoning predicts future visual states while the robot is already acting, then replaces imagined states with every new real observation. Combined with few-step distillation and systems acceleration, the paper reports a peak asynchronous execution frequency of 225 Hz. The model adapts from 10โ€“15 demonstrations, transfers across robot embodiments, and handles some new tasks zero-shot. In the paperโ€™s own evaluations, it reaches 93.6 average on RoboTwin 2.0 and reports stronger real-world results than LingBot-VA and ฯ€0.5 across the tested tasks. ๐Ÿงต 1.

Rohan Paul

11,253 Aufrufe โ€ข vor 1 Monat