Video wird geladen...

Video konnte nicht geladen werden

Zur Startseite

๐Ÿค– Ever wished robots could learn new manipulation tasks with just a few demos โ€” and still generalize? ๐Ÿ”ฅ Introducing **ControlVLA**: Few-shot Object-centric Adaptation for Pre-trained Vision-Language-Action Models. ๐Ÿฆพ๐ŸŽฏ From opening cabinets to folding clothes, pouring cubes, organizing toys โ€” even long-horizon tasks! ๐Ÿ’ก Adapts object-centric cues via ControlNet-style...

19,747 Aufrufe โ€ข vor 1 Jahr โ€ขvia X (Twitter)

1 Kommentare

Profilbild von Raul Verdusco
Raul Verduscovor 1 Jahr

๐Ÿ’กImpressive leap in robot learning! ControlVLA shows how far few-shot adaptation has come โ€” enabling robots to handle complex, long-horizon tasks with minimal demos. ๐Ÿš€๐Ÿ”ง This could be a game-changer for real-world deployment. #Automation #Robotics #AI #FewShotLearning ๐Ÿค–โœจ

ร„hnliche Videos

A team tested Pi0, Pi0 Fast, Gr00t, and ACT on real robot arms in manufacturing tasks. (๐Ÿ”– Bookmark this for later!) The task was precise: place thin rectangular frames from a messy stack into a holder. The team fine-tuned each model on 100 real trajectories and compared training time, inference speed, motion quality, and success rates. โฌ‡๏ธ Hereโ€™s a breakdown of what they found Pi0 (Original) โœ… Strongest overall performance in precise pick-and-place โœ… High success rate even in edge cases โœ… Longest training time (~11 hours, ~$30 per run) โœ… Inference time of 80 ms causes short pauses between actions Despite delays, it handles complex scenarios wellโ€ฆ solid for high-precision tasks, but slow to train. Gr00t โœ… Trains fast (~2 hours, ~$5 per run) โœ… Performs almost as well as Pi0 on large-object tasks โœ… Struggles with fine precision; random movement in some trials โœ… More training didnโ€™t fix jitter or random offsets Best suited for tasks where exact precision isnโ€™t critical. Not ready for manufacturing-grade accuracy without more tuning. Pi0 Fast โœ… Promised faster training, but results were underwhelming โœ… Training at 6 hours still showed low success rates โœ… Inference was slower than expected โœ… Not reliable for generalizing even slightly new tasks Currently too unstable for real-world deployment. Doesnโ€™t live up to the โ€œFastโ€ name yet. ACT (Baseline) โœ… 200MB modelโ€”lightweight, but limited โœ… Struggles with stacked objects or ambiguous scenes โœ… Success rates around 70% in best-case setups โœ… Canโ€™t match newer models on precision or generalization Still a solid baseline, but clearly a generation behind in robustness. ๐Ÿšจ Extra Notes All newer models share a common issue: โ€ขInference takes longer than a frame (80 ms vs 33 ms), so robots โ€œpauseโ€ between chunks. โ€ขThis results in jittery movements, but not a dealbreaker unless tasks are time-sensitive. Language-conditioned tasks also fell short: after training on two labeled tasks, the model couldnโ€™t generalize to a third unseen combination using only text prompts. โœ… The good news? These models adapt well to new robot arms with quick fine-tuning. โŒ The bad news? Thereโ€™s still no plug-and-play solution for improving performance after deployment. Reinforcement learning or DAgger-style data collection during real-world operation may be the next big step, something many teams in robotics are actively working on.

Ilir Aliu

21,844 Aufrufe โ€ข vor 1 Jahr

Most video-action robot models are a content-creation video generator with an action module attached. LingBot-VA 2.0 from Robbyant, a video-action foundation model, throws that starting point out and trains the whole stack natively for control. And it runs closed-loop at a peak 225 Hz. It's so important because A robot cannot move responsively when its controller pauses to imagine the next few frames. LingBot-VA 2.0 predicts during execution, then corrects using each real observation. And it carries only about 13B video parameters while activating roughly 1.9B per token. Bigger robot models usually mean slower reactions, creating a direct conflict between intelligence and control. LingBot-VA 2.0 is trained from scratch for robot control rather than adapted from a video generator built for content creation. Robbyant, an embodied AI company under Ant Group, built it to learn how scenes change under actions, predict what should happen next, and turn those predictions into real-time robot movements. Most video-action systems inherit a tokenizer and video backbone trained mainly to reproduce visual appearance. LingBot-VA 2.0 rebuilds both parts around physical control. Its semantic visual-action tokenizer maps observations toward features from a frozen vision foundation model and learns compact latent actions from frame-to-frame changes using self-supervised inverse and forward dynamics. Unlabeled web video can therefore carry action-relevant training signals without robot action labels. The policy is causal from the start, so every prediction can use only past observations. Its sparse Mixture-of-Experts video backbone has about 13B total parameters, while about 1.9B are active per token, keeping the compute lower during each step. A high-level vision-language planner breaks long tasks into smaller instructions, while the low-level video-action policy handles continuous movement. Foresight Reasoning predicts future visual states while the robot is already acting, then replaces imagined states with every new real observation. Combined with few-step distillation and systems acceleration, the paper reports a peak asynchronous execution frequency of 225 Hz. The model adapts from 10โ€“15 demonstrations, transfers across robot embodiments, and handles some new tasks zero-shot. In the paperโ€™s own evaluations, it reaches 93.6 average on RoboTwin 2.0 and reports stronger real-world results than LingBot-VA and ฯ€0.5 across the tested tasks. ๐Ÿงต 1.

Rohan Paul

11,253 Aufrufe โ€ข vor 2 Monaten

๐Ÿšจ BREAKING: Microsoft's first robotics foundation model! ๐Ÿคฏ Microsoft just announced Rho-alpha (ฯฮฑ), their first robotics model derived from the Phi series of vision-language models. Rho-alpha translates natural language commands into control signals for robotic systems performing bimanual manipulation tasks. Commands like "push the green button with the right gripper," "pull out the red wire," "flip the top switch on," or "turn the knob to position 5" get executed directly by dual-arm robots. What makes this different from standard vision-language-action (VLA) models is the additional modalities. Rho-alpha is a VLA+ model that adds tactile sensing to the perceptual mix, with plans to incorporate force feedback. On the learning side, the model is designed to continually improve during deployment by learning from human feedback. The training approach combines trajectories from physical demonstrations and simulated tasks with web-scale visual question answering data. Since teleoperation data is scarce and expensive, Microsoft is using NVIDIA Isaac Sim on Azure to generate physically accurate synthetic datasets via reinforcement learning. These simulated trajectories get combined with commercial and open physical demonstration datasets. The model is currently under evaluation on dual-arm setups and humanoid robots. Microsoft is opening an Early Access Program for organizations interested in evaluating Rho-alpha. Robots that can adapt to dynamic situations and human preferences are more useful in real environments and more trusted by the people operating them. Read more here: ~~ โ™ป๏ธ Join the weekly robotics newsletter, and never miss any news โ†’

Lukas Ziegler

61,071 Aufrufe โ€ข vor 8 Monaten

Don't waste 2 years figuring out AI on your own in 2026. a free GitHub repo from Microsoft with 56,200 stars just gave out the entire framework to master "AI Engineering in 2026" In only 12 weeks. You can go from 0 % โ†’100 % [ here is how to do that] 4% โ†’ Introduction and History of AI 8% โ†’ Knowledge Representation and Expert Systems 13% โ†’ Perceptron 17% โ†’ Multi-Layered Perceptron and Creating Your Own Framework 21% โ†’ Intro to Frameworks (PyTorch/TensorFlow) and Overfitting 25% โ†’ Intro to Computer Vision, OpenCV 29% โ†’ Convolutional Neural Networks and CNN Architectures 33% โ†’ Pre-trained Networks, Transfer Learning and Training Tricks 38% โ†’ Autoencoders and VAEs 42% โ†’ Generative Adversarial Networks and Style Transfer 46% โ†’ Object Detection 50% โ†’ Semantic Segmentation, U-Net 54% โ†’ Text Representation, BoW/TF-IDF 58% โ†’ Semantic Word Embeddings, Word2Vec & GloVe 63% โ†’ Language Modeling, Training Your Own Embeddings 67% โ†’ Recurrent Neural Networks 71% โ†’ Generative Recurrent Networks 75% โ†’ Transformers and BERT 79% โ†’ Named Entity Recognition 83% โ†’ Large Language Models, Prompt Programming and Few-Shot Tasks 88% โ†’ Genetic Algorithms 92% โ†’ Deep Reinforcement Learning 96% โ†’ Multi-Agent Systems 100% โ†’ AI Ethics and Responsible AI it replaces 10 paid courses on agentic engineering. watch it today, then read the article below on how to generate alpha using graph engineering for hedge fund โ†“

Avid

14,619 Aufrufe โ€ข vor 2 Monaten