ๆญฃๅœจๅŠ ่ฝฝ่ง†้ข‘...

่ง†้ข‘ๅŠ ่ฝฝๅคฑ่ดฅ

How can robots reliably place objects in diverse real-world tasks? ๐Ÿค–๐Ÿ” Placement is toughโ€”objects vary in shape and placement modes (such as stacking, hanging, and insertion), making it a challenging problem. We introduce AnyPlace, a two-stage method trained purely on synthetic data to predict diverse placement poses of unseen...

25,378 ๆฌก่ง‚็œ‹ โ€ข 1 ๅนดๅ‰ โ€ขvia X (Twitter)

0 ๆก่ฏ„่ฎบ

ๆš‚ๆ— ่ฏ„่ฎบ

ๅŽŸๅง‹ๅธ–ๅญ็š„่ฏ„่ฎบๅฐ†ๆ˜พ็คบๅœจ่ฟ™้‡Œ

็›ธๅ…ณ่ง†้ข‘

You can't 3D reconstruct glass from images... ...WRONG! Thanks for video diffusion, now just about anything is possible! Introducing...Diffusion Knows Transparency (DKT) Transparent and reflective objects usually break robot vision and photogrammetry pipelines because they don't follow the "solid object" rules standard cameras expect. DKT is a new AI model that repurposes the "internal physics engine" found in video generation models to solve this problem. Researchers took a massive video diffusion model (WAN) and fine-tuned it using a custom-built synthetic dataset to turn it into a high-precision depth sensor. To train the AI, they built the first massive synthetic video library of transparent objects, 1.32 million frames of perfectly labeled glass and metal objects in motion. Without ever seeing a "real" labeled video of glass during training, the model (DKT) outperformed all previous specialized systems on real-world benchmarks (ClearPose, DREDS). They created a "lightweight" 1.3B parameter version that runs fast enough (0.17s per frame) to be used on actual robot hardware. Two reasons I find this project important: 1. It further proves that synthetic data will be essential for training the next generation vision models. 2. In real-world robotic tests, using DKT's depth maps nearly doubled the success rate of robot arms trying to pick up objects on tricky reflective or translucent surfaces. At home robots will need to interact with these types of objects on a daily basis. Check out the project page here: Code is LIVE! #Computervision #Robotics #AI

Jonathan Stephens

17,712 ๆฌก่ง‚็œ‹ โ€ข 8 ไธชๆœˆๅ‰

Trained on zero real-world data. Learned to walk, pick up boxes, and follow multi-step instructions... in the REAL world. ( ๐Ÿ“Œ Paper below) Researchers from Amazon FAR, Berkeley, Stanford, and CMU scanned real rooms with an iPhone, rebuilt them as 3D Gaussian Splatting scenes, then generated 48,000 synthetic trajectories of a Unitree G1 walking, grasping, and placing objects inside those virtual replicas. They rendered the robot's first-person camera view from each run and paired it with the matching language instruction and motion data. That's the dataset every humanoid team needs and nobody has: synced egocentric video + language + kinematics, at scale. Instead of collecting it in the real world, they manufactured it. They trained a vision-language-kinematics policy on that synthetic data alone, then deployed it on the physical G1 across five task types: navigation to a named object, lifting boxes of three different sizes with no per-size tuning, chained multi-step tasks, robustness to mid-task layout changes and flickering lights, and multi-minute long-horizon runs. No real-world fine-tuning at any point. Real-world interaction data has been the hard limit on humanoid learning... slow, expensive, and small. If scanning a room once and synthesizing thousands of labeled interactions holds up as a general recipe, that limit moves. Data stops being the bottleneck robotics teams have to solve for. ๐Ÿ“Œ Paper: Project: โ€”โ€”- Weekly robotics and AI insights. Subscribe free:

Ilir Aliu

12,950 ๆฌก่ง‚็œ‹ โ€ข 1 ไธชๆœˆๅ‰

๐—˜๐˜ƒ๐—ฒ๐—ฟ๐˜†๐—ผ๐—ป๐—ฒโ€™๐˜€ ๐˜๐—ฎ๐—น๐—ธ๐—ถ๐—ป๐—ด ๐—ฎ๐—ฏ๐—ผ๐˜‚๐˜ โ€œ๐—ฃ๐—ต๐˜†๐˜€๐—ถ๐—ฐ๐—ฎ๐—น ๐—”๐—œ" - the idea that we can simulate real-world environments so well that robots trained in simulation will work perfectly in reality. ๐—ง๐—ต๐—ฒ ๐—ฝ๐—ฟ๐—ผ๐—บ๐—ถ๐˜€๐—ฒ: Train in virtual worlds โ†’ deploy anywhere. ๐—ง๐—ต๐—ฒ ๐—ฟ๐—ฒ๐—ฎ๐—น๐—ถ๐˜๐˜†: Iโ€™ve seen too many teams fall into this trap. After working with manipulation teams at Berkeley, Imperial, and Dyson, hereโ€™s the pattern: โ€ข ๐—ช๐—ฒ๐—ฒ๐—ธ ๐Ÿญ: โ€œOur policy works perfectly in simulation!โ€ โ€ข ๐—ช๐—ฒ๐—ฒ๐—ธ ๐Ÿฐ: โ€œWhy doesnโ€™t this work on real objects?โ€ โ€ข ๐— ๐—ผ๐—ป๐˜๐—ต ๐Ÿฎ: โ€œWe basically need to retrain from scratch with real data.โ€ ๐—ง๐—ต๐—ฒ ๐—ด๐—ฎ๐—ฝ ๐˜€๐—ถ๐—บ๐˜‚๐—น๐—ฎ๐˜๐—ถ๐—ผ๐—ป๐˜€ ๐—ฐ๐—ฎ๐—ปโ€™๐˜ ๐—ฏ๐—ฟ๐—ถ๐—ฑ๐—ด๐—ฒ: Unlike blind locomotion policies that can get away with sim-to-real transfer because they rely mainly on proprioception and contact forces, ๐˜ƒ๐—ถ๐˜€๐—ถ๐—ผ๐—ป-๐—ด๐˜‚๐—ถ๐—ฑ๐—ฒ๐—ฑ ๐—บ๐—ฎ๐—ป๐—ถ๐—ฝ๐˜‚๐—น๐—ฎ๐˜๐—ถ๐—ผ๐—ป ๐—ถ๐˜€ ๐—ฒ๐˜…๐˜๐—ฟ๐—ฒ๐—บ๐—ฒ๐—น๐˜† ๐˜€๐—ฒ๐—ป๐˜€๐—ถ๐˜๐—ถ๐˜ƒ๐—ฒ ๐˜๐—ผ ๐˜ƒ๐—ถ๐˜€๐˜‚๐—ฎ๐—น ๐—ฑ๐—ผ๐—บ๐—ฎ๐—ถ๐—ป ๐—ด๐—ฎ๐—ฝ๐˜€. โ€ข Real friction vs simulated surface textures โ€ข Manufacturing tolerances vs perfect CAD models โ€ข Dynamic lighting vs controlled virtual environments โ€ข Sensor noise vs instantaneous virtual readings ๐—›๐—ฒ๐—ฟ๐—ฒ'๐˜€ ๐˜„๐—ต๐—ฎ๐˜ ๐—ฝ๐—ฒ๐—ผ๐—ฝ๐—น๐—ฒ ๐—ฑ๐—ผ๐—ป'๐˜ ๐˜๐—ฎ๐—น๐—ธ ๐—ฎ๐—ฏ๐—ผ๐˜‚๐˜: Building these detailed simulated environments takes forever. If it takes 7 days to build a simulated kitchen in simulation, wouldn't it be better to just collect real-world data in a real kitchen instead? ๐——๐—ผ๐—ป'๐˜ ๐—ด๐—ฒ๐˜ ๐—บ๐—ฒ ๐˜„๐—ฟ๐—ผ๐—ป๐—ด - simulation is incredible for debugging, safety testing, and exploring edge cases. But it's not a magic solution to real-world deployment. ๐—ช๐—ต๐—ฎ๐˜ ๐—ฎ๐—ฐ๐˜๐˜‚๐—ฎ๐—น๐—น๐˜† ๐˜„๐—ผ๐—ฟ๐—ธ๐˜€: Use simulation strategically while making real-world data collection as efficient and flexible as possible. This is why Neuracore focuses on streamlined real-world data infrastructure. Because no amount of virtual training can replace understanding how your robot actually behaves in actual environments. ๐—ง๐—ต๐—ฒ ๐—ฝ๐—ต๐˜†๐˜€๐—ถ๐—ฐ๐˜€ ๐—ผ๐—ณ ๐˜†๐—ผ๐˜‚๐—ฟ ๐—ฑ๐—ฒ๐—ฝ๐—น๐—ผ๐˜†๐—บ๐—ฒ๐—ป๐˜ ๐—ฒ๐—ป๐˜ƒ๐—ถ๐—ฟ๐—ผ๐—ป๐—บ๐—ฒ๐—ป๐˜ ๐—ฐ๐—ฎ๐—ป'๐˜ ๐—ฏ๐—ฒ ๐˜€๐—ถ๐—บ๐˜‚๐—น๐—ฎ๐˜๐—ฒ๐—ฑ ๐—ฎ๐˜„๐—ฎ๐˜†. Whatโ€™s been your experience with sim-to-real transfer?

Stephen James

25,347 ๆฌก่ง‚็œ‹ โ€ข 1 ๅนดๅ‰