Gave a robot 3D vision with just a regular... camera👁️ Full Tutorial: Deployed #Depth Anything V3 on NVIDIA Robotics Jetson AGX Orin. It estimates depth from 2D images in real-time—no special sensors needed. Just need #Monocular depth estimation + #TensorRT optimization + #ROS2 integration. 👉 Learn more about reComputer Robotics J5011: #TheAIHardwarePartnershow more

Seeed Studio
16,966 次观看 • 6 个月前
Meet Seeed at #GTC2026 in Day1! We're live with:... ▶️ #reBot B601 Arm: Watch it doing real-time teleoperation + NVIDIA Robotics Isaac Sim visualization. ▶️ #Openclaw+#JetsonThor→Robot Control: Jetson Thor running local #LLMs to control a robot with natural language. ▶️ #ReachyMini: powered by the #reComputer J4012 with #JetsonOrinNX, reading your vibe with on-device vision & AI. Touch, test, and talk tech with us! 📍San Jose, California, USA 📅 March 16th–19th 🏛️ San Jose Convention Centee, Booth 156 #theAIHardwarePartnershow more

Seeed Studio
13,954 次观看 • 5 个月前
I'll always root for a team that open-sources its... best work, and Robbyant just did it properly. Robbyant, Ant Group's embodied-AI company, released LingBot-Vision, a vision foundation model for robots, and the part I love is the data. They trained it on 161M images, filtered down from 2B raw ones and mostly pulled straight from the open web, with no human labels, no edge detectors, no depth sensors anywhere in the loop. It learns the exact edges of objects from raw pixels. That's roughly a tenth of the data DINOv3 saw, and under a third of the training. And it shows in the results. On depth, working out how far away things are, the 1B model edges out a 7B on NYU-Depth. It also powers LingBot-Depth 2.0, which reads the surfaces cameras usually choke on, glass and mirrors, and halves indoor depth error. LingBot-Vision is fully open. Weights from the 1.1B flagship down to a tiny 21M version, code, and the paper. This is the timeline I want more of. Robbyantshow more

Chubby♨️
48,249 次观看 • 1 个月前
Robots can now reconstruct 3D scenes in real time... from a single RGB camera. [📍 Projects page + paper] No depth sensor. No retraining. 30 FPS. Researchers at the Imperial College London introduced KV-Tracker, a training-free method that makes heavy models like π³ and Depth Anything 3 fast enough for real-time tracking. The idea is simple. These models use global self-attention, which is powerful but computationally expensive. KV-Tracker caches the key and value pairs from selected keyframes and reuses them for new frames. That cache becomes an implicit scene representation. Result: • Up to 30 FPS • 10 to 15x speedup • Accurate 6-DoF tracking on benchmarks like TUM RGB-D and 7-Scenes • Works with monocular RGB only It also supports object-level tracking with masks and allows saving the KV-cache for later reuse. For robotics, this reduces hardware constraints and moves real-time 3D perception closer to practical deployment. Credit to Marwan Taher (Marwan Taher) at Imperial’s Dyson Robotics Lab and many others who contributed to this! 📍 Save projects page + paper for later: Video: ——- if it matters in AI or Robotics you'll read it here first:show more

Ilir Aliu
53,992 次观看 • 4 个月前
Normally, teaching a camera to recognize something new means... collecting images and retraining a model. This one skips that. You just type what you want it to notice. It's called NanoOWL, built by NVIDIA for real-time, zero-shot object detection. Type "an owl," "a face yawning or bored," or "indoors or outdoors" and it finds it live on camera, no training, no labeled data. You can even nest queries, like detecting a face, then finding the eyes and mouth inside it. > needs a Jetson Orin board specifically > plus TensorRT and PyTorch set up > Real embedded-hardware territoryshow more

Simplifying AI
24,059 次观看 • 10 天前
🚨 BREAKING: NVIDIA just announced the Isaac GR00T Reference... Humanoid Robot. The first fully open humanoid robot reference design built on Jetson Thor, and it's going straight to the world's top research institutions. This is Jensen Huang's bet on open physical AI infrastructure. The hardware stack is serious: → Unitree H2 Plus chassis, 6 feet tall, 150 pounds, 31 degrees of freedom → Sharpa Wave tactile five-finger hands, 22 degrees of freedom, bringing total to 75 across the full body → NVIDIA Jetson AGX Thor onboard compute, 2,070 FP4 teraflops of AI performance, 128GB unified memory → Multi-view sensing, stereo head camera, wrist cameras, IMU Alongside this announcement, Unitree also introduced the H2 Plus as a standalone product, a frontier humanoid combining Unitree's own body, Sharpa's five-finger hands and NVIDIA Robotics Jetson Thor compute into one fully integrated research platform. The full Isaac GR00T software stack ships with it, teleoperation for data capture, open foundation models, Isaac Sim for training, Isaac Lab for evaluation, and accelerated ROS middleware for deployment. The complete loop from data to real-world robot in one unified platform. ETH Zürich, Stanford Robotics Center, UC San Diego and Ai2 are already on board as launch research partners. NVIDIA Robotics did to AI what it's now doing to robotics, build the platform, open the ecosystem, let the world build on top of it. Whoever owns the infrastructure layer wins. NVIDIA knows this better than anyone. 👀 Read more here: ~~ ♻️ Join the weekly robotics newsletter, and never miss any news →show more

Lukas Ziegler
16,062 次观看 • 3 个月前
Most robots still need markers, checkerboards, or long calibration... rituals just to know where their arms are. Now it works from raw images in seconds. roboreg is a markerless multi arm localization toolkit that plugs into ROS 2 and RViz. No special hardware. No custom setup. You toggle between robot descriptions and the system figures out the rest. The idea is simple: ✅ Hand eye calibration from plain RGB or RGB D images ✅ Only three robot poses needed for millimeter accuracy ✅ Works with any ROS 2 compatible robot and camera ✅ Fully open source under Apache 2.0 It is powered by Hydra, a new marker free ICP variant that converges far more reliably than classical baselines and runs in under a second. If you want to try it: roboreg: ROS 2 roboreg: Hydra paper: pip install roboreg More details and discussion on Open Robotics Discourse:show more

Ilir Aliu
18,406 次观看 • 9 个月前
You can't 3D reconstruct glass from images... ...WRONG! Thanks... for video diffusion, now just about anything is possible! Introducing...Diffusion Knows Transparency (DKT) Transparent and reflective objects usually break robot vision and photogrammetry pipelines because they don't follow the "solid object" rules standard cameras expect. DKT is a new AI model that repurposes the "internal physics engine" found in video generation models to solve this problem. Researchers took a massive video diffusion model (WAN) and fine-tuned it using a custom-built synthetic dataset to turn it into a high-precision depth sensor. To train the AI, they built the first massive synthetic video library of transparent objects, 1.32 million frames of perfectly labeled glass and metal objects in motion. Without ever seeing a "real" labeled video of glass during training, the model (DKT) outperformed all previous specialized systems on real-world benchmarks (ClearPose, DREDS). They created a "lightweight" 1.3B parameter version that runs fast enough (0.17s per frame) to be used on actual robot hardware. Two reasons I find this project important: 1. It further proves that synthetic data will be essential for training the next generation vision models. 2. In real-world robotic tests, using DKT's depth maps nearly doubled the success rate of robot arms trying to pick up objects on tricky reflective or translucent surfaces. At home robots will need to interact with these types of objects on a daily basis. Check out the project page here: Code is LIVE! #Computervision #Robotics #AIshow more

Jonathan Stephens
17,712 次观看 • 8 个月前
Something big is happening in robotics - and it’s... hiding in plain sight. This post is not about dancing robots but in the data that powers them. Open robotics datasets have exploded this year, turning the field into a more scalable and collaborative ecosystem. In just two years, Hugging Face datasets grew from 11k to over 600k - and robotics is by far the fastest-growing segment. We went from 1k robotics datasets in 2024 to 27k in 2025! For comparison, text generation, the second-largest category, has only around 5k datasets in 2025. That gap is massive. Open datasets are important because robotics lives and dies by real-world robot data - video, actions, sensors, failures. By making this data easy to upload, reuse, and benchmark, researchers, startups, and large players are now releasing real-robot datasets that would have stayed locked inside labs just a few years ago. Major contributors include NVIDIA, LeRobot initiative, and a rapidly growing maker community. This surge is also enabled by cheaper video storage, better tooling, and an open-source AI culture now spilling into the physical world. And it really matters: open robotics data dramatically lowers entry barriers, accelerates learning-by-doing, and speeds up progress toward generalist and humanoid robots. Robotics won’t scale through hardware alone - but to a large extent through shared data. Viz below from AI World - link to the story and more viz/filters in comment.show more

Pierre-Alexandre Balland
186,094 次观看 • 8 个月前
There's an open-source home robot vacuum you can actually... build yourself. It's called OOMWOO. Runs on a Raspberry Pi with 2D LiDAR mapping, ROS2 navigation, and native Home Assistant integration. No cloud dependency, no vendor lock-in. The team spent real time researching before designing it. They reviewed the entire 2025-2026 robot vacuum market, from budget to flagship, and found real-world cleaning doesn't track advertised suction power. The anti-tangle brush design alone is worth mentioning. A tapered rubber roller that resists hair-wrap, one of the most common complaints with commercial models, easy to 3D-print yourself. If you've got a Raspberry Pi and a 3D printer sitting around, this is a genuinely interesting weekend project.show more

Oliver Prompts
584,798 次观看 • 29 天前
Some updates on the multiview vistadream pipeline with Rerun!... Rerun came in extremely useful here, as being able to visualize depths at each stage of the pipeline allowed me to debug some nasty bugs. Since the last time, I was only working with a single image input. I've added in VGGT as my multiview pose + depth estimator. It works REALLY well for getting camera poses, but the depths are not that great. To try and fix that, I estimated depth maps from MoGeV2 for each of the views, and scale+shift aligned them so that they would match up to the confident sections of VGGT's depth predictions. You can see in the video just how much sharper the visualized 2d depth maps are! The biggest issue continues to be the multiview consistency 🫠 That's up next, along with actually training the Gaussian splat. Lots of work went into actually understanding inputs+outputs for VGGT. I had some funky bugs where the confidence values would all collapse to true I'm also really excited for this pipeline to use Difix3D+ Nvidia instead of Flux Inpainting, it seems like a better suited for a multiview pipeline.show more

Pablo Vela
29,904 次观看 • 1 年前
The robot vacuum is getting a major upgrade. Matic... just launched a home robot you can control by pointing and talking. No app. No complicated maps. No endless setup. Just point at a mess and say, “Hey Matic, clean this.” What makes it even more interesting: → 5 cameras instead of LiDAR → NVIDIA-powered on-device AI → Understands objects like cables, socks, pets and spills → Builds a 3D map of your home locally → Works without sending your camera data to the cloud → Supports voice commands in 75 languages This feels less like a smarter vacuum and more like a new interface for home robotics. We spent years teaching humans how to operate machines. Now machines are finally learning how to understand humans. Would you use a robot like this?show more

Markandey Sharma
557,024 次观看 • 20 天前
A great day at LEAP East 2026 at HKCEC... building and connecting. Spent time at the Animoca Brands booth discussing Minds by Animoca Brands and the next wave of persistent AI agents. Think persistent AI agents with memory, identity, and on-chain wallets, that's what it was! For us at Rice Robotics, that's an exciting direction. As robots become more autonomous, they'll need persistent digital identities, long-term memory, and the ability to interact with both people and decentralized networks. Oh... and we had an insightful and in-depth discussion with someone pretty special and big 👀 More on that soon🤫show more

RICE AI ( ◉ - ◉ )
21,472 次观看 • 1 个月前
honestly no surprise why Silicon Valley is so obsessed... with Matic robots right now > basically a Roomba on steroids > vacuums first, then mops > uses five cameras to build a live 3D map of your home > recognizes rugs, wires, furniture, pets, and people, then changes how it cleans > point at a mess and say “hey Matic, clean this” and it does > an NVIDIA Jetson inside the robot handles all the vision, mapping, and navigation > raw footage is discarded in real time and your 3D map never leaves the device btw that last privacy part is extremely underrated IMO my biggest fear with robotics is putting moving cameras and microphones inside our most private spaces without knowing where the data goes. > this week, camera components on Royal Navy drones were caught phoning home to China > Chinese Unitree robot dogs were found with a backdoor that let anyone with the key remotely control them and watch through their cameras > Roombas sent images from inside homes to overseas labelers, including a woman on the toilet and a child your home is your most sacred private space. if you're gonna buy a robot that can physically record, map, and move through it, make sure it's secure!show more

Ole Lehmann
45,092 次观看 • 20 天前
AI in robotics gets all the attention right now,... but sometimes the most interesting work is very practical. Viet built a small vision system that counts potatoes on a conveyor belt. No giant dataset. No huge model. Just a clear problem and a smart setup. He used Ultralytics’ ObjectCounter, trained a tiny YOLO11 nano model, and because there was no potato dataset, he annotated a single frame with SAM 2 and trained from that. One frame. Still works across the whole video. It is a good reminder that useful AI in industry often looks like this. Focused. Lightweight. Solves a real task. If you work in manufacturing or robotics, these small systems are usually the fastest wins. They save time, reduce errors, and do not need massive infrastructure. Nice work, Viet. His projects: —- Weekly robotics and AI insights. Subscribe free:show more

Ilir Aliu
1,676,852 次观看 • 9 个月前
Apple just trained a 3D Gaussian head reconstruction model... on 10,000+ subjects. Feed-forward. No test-time optimization. New identity in, reconstructed Gaussian head out. The UV-parameterized Gaussian representation decouples the number of Gaussians from the number and resolution of input images, making it practical to train with many high resolution views. And the heads are not just static either: text-conditioned identity generation, plus blendshape-driven latent animation across identities. We've been building in the 3D Gaussian Splatting space for a while. The gap between "research demo" and "works on real people at scale" is closing fast.show more

KIRI Engine - 3D Scanner App
12,181 次观看 • 3 个月前
🚨 MACHINES MAY SOON SEE THE WORLD LIKE HUMANS... A US company has unveiled what it calls the world’s first “native color lidar” system giving machines the ability to perceive depth, distance, and color simultaneously. Unlike traditional lidar systems that need separate cameras, this new technology captures full 3D color information directly at the point of detection. In simple terms: Machines may soon understand the world more like human vision instead of just measuring shapes and distance. Why this matters: This could dramatically improve: • self-driving cars • robotics • drones • factory automation • AI navigation • autonomous machines The system can reportedly process over 10 million points every second and detect objects nearly 1,640 feet away. Researchers say this could become one of the key technologies powering the next generation of “Physical AI” where machines don’t just calculate the world… They visually understand it. We may be watching the birth of machine perception in real time. Follow for more future technology and AI discoveries.show more

TheNewPhysics
26,128 次观看 • 3 个月前
🔥 Phoenix is officially live on Solaris AI Flow.... You can now trade Phoenix perps inside a Solaris AI workflow. No code. Drop a node onto the canvas, pick an operation, and wire it to anything: AI signals, price feeds, schedules, alerts. The full order-book DEX from Ellipsis Labs, now programmable. Built so you can trade with confidence: ✦ Paper mode is on by default. Every order is checked and simulated against the live order book, real depth and real slippage, but nothing is signed or broadcast. ✦ Paper behaves exactly like Live. If an order would be rejected on-chain, it is rejected in simulation too. No false fills, no surprises. ✦ Going Live is one switch, and it asks for confirmation before any real funds move. Build it. Test it. Trade it. A full perps strategy, proven on paper before a single dollar moves. 30 operations in one node: ✦ Read live markets, order book depth, candles, and funding rates ✦ Track your positions, collateral, and PnL, realized and unrealized ✦ Place limit, market, and stop-loss orders, plus conditional triggers ✦ Cancel orders, manage margin, and move collateral, all from the workflow No scripts. No backend. No terminal to babysit. Just a workflow that trades. You can try it for free. Demo + Link down below 👇show more

Solaris AI
10,075 次观看 • 3 个月前
1/ World models are getting popular in robotics 🤖✨... But there’s a big problem: most are slow and break physical consistency over long horizons. 2/ Today we’re releasing Interactive World Simulator: An action-conditioned world model that supports stable long-horizon interaction. 3/ Key result: ✅ 10+ minutes of interactive prediction ✅ 15 FPS ✅ on a single RTX 4090🔥 4/ Why this matters: it unlocks two critical robotics applications: 🚀 Scalable data generation for policy training 🧪 Faithful policy evaluation 5/ You can play with our world model NOW at NO git clone, NO pip install, NO python. Just click and play! NOTE ⚠️ ALL videos here are generated purely by our model in pixel space! They are **NOT** from a real camera More details coming 👇 (1/9) #Robotics #AI #MachineLearning #WorldModels #RobotLearning #ImitationLearningshow more

Yixuan Wang
129,102 次观看 • 6 个月前
CHINA JUST DEPLOYED A ROBOT THAT LOOKS LIKE A... TIRE AND ROLLS NEXT TO COPS ON PATROL. NO LEGS. NO WHEELS. JUST A SPHERE. It rolls down the sidewalk like a runaway truck tire. Except it's covered in cameras, sensors, and it knows exactly where it's going. People stop and stare. It doesn't care. It just keeps rolling. At a tech expo it rolled across the showroom floor while people backed away from it. On city streets it rolled alongside police officers like a K-9 unit made of rubber and AI. No legs to break. No wheels to get stuck. No joints to maintain. Just a ball with treads that goes anywhere. Stairs? It rolls over them. Rough terrain? It grips and climbs. Rain, mud, gravel? Doesn't matter. The entire surface is the wheel. Boston Dynamics' Spot costs $75,000 and walks on four legs that need constant calibration. A human patrol officer costs $50,000-$70,000 a year in salary alone. This sphere rolls 24/7, doesn't take breaks, doesn't call in sick, and has no joints to replace. That's when it clicks. Every robotics company spent years trying to make robots walk. Billions went into teaching machines to balance on two legs. China looked at the problem and said "what if it just rolled?" The most advanced mobility solution turned out to be the oldest shape in existence. A ball. Would you feel safe with this rolling behind you?show more

Framez
20,566 次观看 • 16 天前
THIS AIBO ISN’T A TOY. IT’S THE MOST ADVANCED... ROBOT PET EVER BUILT. 🐕🤖 Sony didn’t just build a gadget — they built something that runs on 22-axis lifelike movement, recognizes your face, and develops its own “personality” over time based on how you interact with it. Here’s what makes it wild: 🐾 Learns and adapts — no two Aibos behave exactly the same after a few months of use 🐾 Cloud-connected personality — its behavior profile lives on Sony’s servers, evolving with every interaction 🐾 Free developer API with visual programming — which has made it a favorite tool in robotics classrooms worldwide 🐾 Real-world use cases far beyond “cute robot dog”: home companionship, senior care and emotional support, STEM education, even security patrol duties No vet bills. No allergies. No mess. Just a machine that’s been engineered to feel less like a product and more like a living companion. We used to think “robot pets” meant clunky toys with blinking lights. Sony quietly spent years proving the whole category could feel real. The future of companionship isn’t fully human anymore. And it’s already sitting on someone’s couch right now.show more

DN_DEGEN
114,038 次观看 • 1 个月前