正在加载视频...

视频加载失败

The robot’s “eyes” just received a big upgrade. LingBot-Depth 2.0, a depth-completion model with half the depth error just dropped. 12/16 benchmarks topped. Glass, mirrors, and transparent objects are so easy for us humans, but so hard for robots, because they do not behave like ordinary surfaces in a...

25,312 次观看 • 1 个月前 •via X (Twitter)

0 条评论

暂无评论

原始帖子的评论将显示在这里

相关视频

🚨 BREAKING: Big news in the computer vision world! 🎥 Luxonis | Robotic Vision just dropped its new OAK 4 line, and it’s a big upgrade for edge computer vision. Instead of being “just a stereo camera,” OAK 4 is a fully standalone vision computer with 52 TOPS of on-device AI. Models run locally, depth is computed locally, and no external PC or cloud pipeline is required. This is why robotics teams love it: lower latency, lower cost, fewer failure points in the field. The hardware is built for the real-world. IP67, shock-resistant, wide-FOV RGB + stereo pair, IR projection, IMU, audio, and a patent-pending calibration system that keeps depth accurate even when conditions change. But the real move is the platform. With Luxonis Hub, you can deploy models, grab telemetry, push OTA updates, or collect data when performance drifts, all from a unified interface. It turns a single device into an end-to-end edge CV system. Most customers today in robotics are groups who just want something that works: AMRs, bin-picking systems, trailer-loading robots, and ag-tech. 🤖 And they all say the same thing, the appeal isn’t raw TOPS, it’s the all-in-one simplicity that lets them scale without building custom infrastructure. Feels like the direction edge vision has been waiting for: rugged hardware + high-throughput on-device compute + a real management layer. A next step toward “plug-and-deploy” perception for robots. 🔗 Find out more here: ~~ ♻️ Join the weekly robotics newsletter, and never miss any news →

Lukas Ziegler

41,955 次观看 • 8 个月前

Most video-action robot models are a content-creation video generator with an action module attached. LingBot-VA 2.0 from Robbyant, a video-action foundation model, throws that starting point out and trains the whole stack natively for control. And it runs closed-loop at a peak 225 Hz. It's so important because A robot cannot move responsively when its controller pauses to imagine the next few frames. LingBot-VA 2.0 predicts during execution, then corrects using each real observation. And it carries only about 13B video parameters while activating roughly 1.9B per token. Bigger robot models usually mean slower reactions, creating a direct conflict between intelligence and control. LingBot-VA 2.0 is trained from scratch for robot control rather than adapted from a video generator built for content creation. Robbyant, an embodied AI company under Ant Group, built it to learn how scenes change under actions, predict what should happen next, and turn those predictions into real-time robot movements. Most video-action systems inherit a tokenizer and video backbone trained mainly to reproduce visual appearance. LingBot-VA 2.0 rebuilds both parts around physical control. Its semantic visual-action tokenizer maps observations toward features from a frozen vision foundation model and learns compact latent actions from frame-to-frame changes using self-supervised inverse and forward dynamics. Unlabeled web video can therefore carry action-relevant training signals without robot action labels. The policy is causal from the start, so every prediction can use only past observations. Its sparse Mixture-of-Experts video backbone has about 13B total parameters, while about 1.9B are active per token, keeping the compute lower during each step. A high-level vision-language planner breaks long tasks into smaller instructions, while the low-level video-action policy handles continuous movement. Foresight Reasoning predicts future visual states while the robot is already acting, then replaces imagined states with every new real observation. Combined with few-step distillation and systems acceleration, the paper reports a peak asynchronous execution frequency of 225 Hz. The model adapts from 10–15 demonstrations, transfers across robot embodiments, and handles some new tasks zero-shot. In the paper’s own evaluations, it reaches 93.6 average on RoboTwin 2.0 and reports stronger real-world results than LingBot-VA and π0.5 across the tested tasks. 🧵 1.

Rohan Paul

11,253 次观看 • 1 个月前