MAKING AXIS TASKS FEEL EASY 👀🤖 I made a... quick video showing how to complete an Axis Robotics task step by step. The task gives you a clear goal, steps, controls and camera views to guide you. In this one, the goal is simple, Open the trash bin, pick up the onion, and place it inside. 🧅 Just follow the instructions, control the robot arm and complete the task. What I like is that even beginners can understand what they need to do. But the interesting part is what happens with these interactions after we complete them. Each completed task can contribute valuable robot training data, helping Physical Ai systems learn how to handle different objects and environments. Small tasks from the community can become useful data for teaching robots how to interact with the real world. That’s what makes Axis interesting to me.show more

SufianXFN
14,419 views • 7 days ago
If you’ve been ignoring Axis Robotics because it looks... like another random points farm, read this. In the last 48 hours, Unitree said robotics is approaching its "ChatGPT moment", while the chairman of ACE Robotics believes it could happen by the end of 2027. But for robots to reach that level, they need a crazy amount of training data. That’s exactly what Axis is building. You control simulated robots directly from your browser, complete simple tasks and earn points. Those movements also help create data that can be used to train real robots. And this isn’t some tiny experiment anymore: - $12M raised - 123K+ contributors - 3M+ robot trajectories - Community data already used to train a real robot So yeah, we’re basically farming a potential airdrop while teaching our future robot servants how to work 😂 If you haven’t started yet, it’s completely free. You only need a tiny amount of gas on Base to sign your completed tasks. ✅ Start farming Axis points: Important: Sign every completed task from the History page, otherwise you won’t receive the points.show more

Pranjal Bora 🧭
29,321 views • 22 days ago
From seeing → to understanding → to following →... now to finding. In the last update, our robot learned how to recognize and follow people SR Agentic now evolves from human-following → to goal-driven object search. Give it a task like: → “find the white bottle on the table ” It will: • break down the instruction via LLM task planning • scan the environment in real-time • localize the object in 3D space • navigate and approach autonomously Step by step, capability by capability — from perception → to tracking → to task execution in the physical world. This is how real-world agentic robotics is builtshow more

Strike Robot
22,038 views • 5 months ago
Alright, now that we know *what* an agent is,... how does it actually work? When you ask for help on a task, the agent plans a series of steps and executes them directly in the application on your behalf, using the tools it has access to. Say you are booking a local service or trying to organize your inbox (which typically takes multiple steps): the AI model first plans how to achieve the task using its existing knowledge and then interacts with your inbox to execute the task. The agent will continue until it is confident the task has been successfully completed.show more

Google AI
22,487 views • 9 months ago
Robotics keeps hitting the same wall. Single task RL... works, but... it does not scale to hundreds of tasks or new embodiments. This new paper looks like a real step toward fixing that. The team introduces MMBench, a benchmark with 200 tasks across many domains and robots, and Newt, a language conditioned world model trained online across all 200 tasks at once. The simple idea behind Newt: The model learns from demos to get the right priors It trains across many tasks through online interaction It uses language to ground the goal It adapts fast when a new task shows up What stood out to me: ✅ One model trained on 200 tasks at the same time ✅ Language conditioned control for both states and RGB ✅ Better data efficiency than strong baselines ✅ Strong open loop control ✅ Fast adaptation to new tasks and embodiments ✅ Full release of 200 checkpoints, 4000 demos, code, and benchmark This is a good push toward general control instead of one model per task. If you want the full paper: Project page: —- Weekly robotics and AI insights. Subscribe free:show more

Ilir Aliu
70,090 views • 9 months ago
Rare-event discovery for robots training is what we’ve been... working on lately How would a robot react to an unusual event? What if a firefighter drone won’t be able to choke a fire (like on the video below)? Robo-doctors use cases? How to mass-produce these situations to train robots to see & react the best way? I think that open-sourced, world models like LTX can become the solution for such training process. And they might become the differentiator for the future of robotics. Of course, I’m not an expert in robotics and ML, but this topic makes me curious - and I’m curious about your thoughtsshow more

AmirMušić
41,839 views • 1 month ago
AI has had exactly two scaling axes that worked... so far, and the second one is starting to look finite too the first one was pretraining: with scaling parameters and data, we got world knowledge (i.e. ChatGPT had read enough to know things), but it started saturating a while ago the second one was RL, and people had been doing RL the whole time before that: RLHF is RL but it never scaled far because it was trying to control the exact output, which tokens come out, how the text reads, but you can only push that so far before you’re just polishing RLVR dropped that constraint: giving the model a task, then checking whether the final answer is right, and ignoring everything in between -- so the model does whatever it wants in the middle and only the endpoint gets graded, and that’s much closer to actual RL and it’s what bought us planning and reasoning (arguably, tool use sits around 2.5 on this list -- while useful, it's not a different kind of thing) so one axis gave knowledge, the other gave reasoning, and both of them are one model working alone the next axis is how many models you can get working on the same problem, which is a different kind of axis than the previous two we know that multi-agent RL has always been the harder problem: I spent years in that literature and the gap between single-agent and multi-agent is definitely not incremental -- it’s a whole different class of difficulty! which is also why the derivatives are steep at the start, nobody has picked the easy wins yet... and the thing that gates this multi-agent coordination is communication: models can only coordinate as well as they can exchange information, and right now they do that by writing sentences to each other imagine what could we possibly achieve if we properly open that third axis development by letting models to exchange information in their native "language" without loosing any computational data that they produce during inferenceshow more

Sasha Malysheva
14,217 views • 1 month ago
A Gaussian Splat can become a world where Robots... and AI agents can act. In our latest OVER Research experiment, we placed a robot inside a real-world 3D capture, with a VLM making decisions based on what it sees. At every step, the robot holds a pose in the reconstruction, gets a newly rendered view of the environment, takes an action, moves, and sees the world again from its new position. Why does this matter? Because 3D captures can become more than reconstructions to explore. They can become environments where embodied AI and robots can navigate, act, be evaluated and eventually train across real-world spaces at scale. Capture a place once. Then turn it into a world where AI and Robots can act. The full experiment, including what we discovered once we actually put the loop to the test:show more

Over the Reality 🌐
14,975 views • 12 days ago
Boom! Grok Tasks Make It One Of The Most... POWERFUL Real-Time AI Systems In The World. — My How to Use Grok Tasks With Hidden Tools For Powerful Daily Output. Grok Tasks are customizable AI workflows that integrate a variety of tools to streamline daily activities, from research and analysis to creative planning and problem-solving. I have been using them for quite sometime and because of the vital heartbeat of news and first person data on X, it is the most powerful AI platform available. By combining Tasks with tools like web searches, X platform interactions, code execution, and media viewers, you can build efficient, automated processes. These tasks work by prompting Grok with a clear description of what you want to achieve, and Grok will intelligently call the necessary tools in sequence or parallel to deliver results. Here's a step-by-step guide to creating and using Grok Tasks: Step 1: Define Your Task Start by clearly outlining the daily activity or goal. Consider what inputs you have (e.g., a URL, a query, or an attachment) and what output you need (e.g., a summary, calculation, or visual analysis). Break it down into subtasks to identify tool needs. For example, if your task involves researching current events, note that you'll need search and browsing capabilities. Step 2: Review Available Tools Familiarize yourself with the tools Grok can access. Here's a quick overview: - Code Execution: Run Python code for calculations, data processing, or simulations using libraries like numpy, pandas, or sympy. - Browse Page: Fetch and summarize content from any website URL with custom instructions. - Web Search: Perform general internet searches, returning results with optional operators like site:. - Web Search With Snippets: Get quick, detailed excerpts from search results for fact-checking. - X Keyword Search: Advanced search for X posts using operators like from:, since:, or filter:. - X Semantic Search: Find semantically related X posts based on a query, with filters for dates or users. - X User Search: Locate X users by name or handle. - X Thread Fetch: Retrieve a full X post thread, including context like replies and parents. - View Image: Analyze an image from a URL or conversation ID. - View X Video: Extract frames and subtitles from an X-hosted video. - Search PDF Attachment: Query a PDF file for relevant pages using keyword or regex modes. - Browse PDF Attachment: View specific pages of a PDF with text and screenshots. Select tools that align with your task. Aim for a mix to handle data gathering, processing, and visualization. Step 3: Craft Your Prompt Write a detailed prompt to Grok describing the task. Include: - The overall goal. - Specific steps or subtasks. - References to tools if you want to guide the process (e.g., "Use web_search to find sources, then code_execution to analyze data"). - Any constraints, like dates or limits. Example prompt: "Create a Grok Task for my morning routine: Search recent X posts about tech news using x_keyword_search, fetch a key thread with x_thread_fetch, and summarize with browse_page on linked articles." Step 4: Submit and Interact Send your prompt to Grok. It will process the task by calling tools as needed, often in parallel for efficiency. Review the output and refine with follow-up prompts if required (e.g., "Expand on that using view_image for visuals"). Iterate to fine-tune the workflow for reuse. Step 5: Save and Reuse Once refined, note the prompt as a template for future use. You can adapt it for similar tasks, making Grok Tasks a habitual part of your day. Finding Grok Tasks To discover existing Grok Tasks or inspiration for new ones, use X searches with tools like x_keyword_search or x_semantic_search (e.g., query: "Grok Tasks examples" with mode: Latest). Browse community-shared threads via x_thread_fetch, or web_search for tutorials on xAI features. Prompt Grok directly: "Show me popular Grok Tasks for productivity." 1 of 3show more

Brian Roemmele
152,242 views • 8 months ago
Here is a drill to help you learn what... it feels like to use the ground and start to sequence effectively. Notice how the club is setup in the video between my feet. Get to your back swing and jump favoring your lead foot left and laterally like I am. With that feeling step in and hit one. Not where you are pushing off from when you swing and try and get that into a feeling that makes sense to you. This is how power is created in the golf swing. This is also where physical limitations can show up.show more

Drake Smith
30,876 views • 4 months ago
Pi Ventures Joins Hack VC To Back AI Robotics... Startup On Base Axis Robotics (Axis Robotics) has raised a $12 million seed round led by Hack VC, with participation from Pi Network (Pi Network) Ventures, Nomad Capital, 10K Ventures, and other angel investors. The company is building a data engine for Physical AI that combines simulation, real world data capture, and human feedback to generate scalable robotics datasets. Axis said the funding will accelerate development of its global human in the loop data engine. Base (Base APAC) congratulated the team, calling Axis one of the leading scalable Physical AI and robotics platforms building on the network.show more

BSCN
51,740 views • 1 month ago
Claude Code can now go find the right skill... itself, instead of you searching for one. it's called find-skills. a small package that plugs into Claude Code, and instead of you hunting for rest, you just describe the task and it searches the whole open skills ecosystem, finds the ones that fit, and installs them for you. > tell it what you're trying to do, in plain english > it scans the skills registry and maps your task to real skills > it pulls the right ones in and sets them up half the time you don't even know a skill exists for what you're doing. now you don't have to.show more

Alvaro Cintas
38,829 views • 29 days ago
Super excited about Hydra-0 from Hongyu Li and team!... The key idea is to use flow as a shared visual interface across embodiments/objects for controllable video generation, allowing a single generalist world model to learn from human, handheld-gripper, and robot interaction data. My favorite result is the video below: start from a real video of a human doing the task (left), extract the desired object flow, and condition the model on that flow (right). The model then hallucinates a plausible robot motion that could produce the same object motion. Very cool glimpse of how a generalist world model can bridge human demonstrations and robot control.show more

Yunzhu Li
10,613 views • 23 days ago
Pi0 vs. ACT with BBox conditioning 🟦 Not many... know you can push ACT to *almost-Pi0* generalisation by conditioning on bounding boxes (BBoxes). How is the training data collected? • Generate BBoxes for all pick-and-place objects in the scene.(I used Gemini) • Pick-and-place targets are selected randomly. • Add the BBox coordinates to the robot’s state. • Overlay the BBoxes in the visualisation so you know what to grab and where to drop. During inference: • Generate BBoxes for every object again. • Click the object you want to pick and its target spot; those BBoxes get added to the robot state. • Let the robot do the work for you 😃 Setup: - Trained ACT for 100k steps and fine-tuned Pi0 for only 20k. - Training data is 60 episodes and had *only* LEGO bricks. - Using single front camera (Laptop in this case) Got the idea from xun in LeRobot discord. Here’s ACT vs Pi0 on a toy car that isn’t in the dataset. 1/3show more

Shreyas Gite
34,863 views • 1 year ago
What happens when an AI stops telling you how... to build something and actually builds it? Nex-N2.5 Pro took an idea from concept to playable game: → Rough concept → 2D design → Blender 3D environment → Working game The interesting part wasn’t the final result. It was the process. The model interacted directly with the tools, observed the changes on screen, made adjustments, and continued working instead of simply generating instructions. That’s the real computer-use test. Can an AI: • Open Blender and build a scene? • Inspect the output and iterate? • Write a script, tweak parameters, and troubleshoot? • Hit a limitation, switch to the GUI, and finish the task? Nex-N2.5 Pro is starting to make that workflow feel surprisingly real. You can try it on OpenRouter for a limited time: Try Nex-N2.5 Pro on OpenRouter #NexN2 #Blender #ComputerUse #OpenRoutershow more

IMOLEAYOMIDE
108,494 views • 3 days ago
Ready to bring AI into the physical world? 🤖... Join our livestream on Tuesday, August 11th at 9:00am PT to learn how to: 🧠 Run and fine-tune NVIDIA GR00T 1.7 🔗 Connect models, sensors, and actuators with ROS 2 🦾 Deploy a model on the SO-101 robot arm to execute a real-world task 🎙️Guest speaker Nikodem Bartnik will also share the latest advances and highlights from the open-source robotics community. Add to calendar:show more

NVIDIA Robotics
23,616 views • 1 month ago
It's 2030 and you are reviewing humanoid robots. A... Tesla. A Google. An Apple. An OpenAI. A Meta. A Figure. And a bunch of Chinese-made ones. Which one is best, and why? I think the Tesla understands the world much better. Why? There were eight Teslas around me on the freeway today. Start there. No other robot company has that data. But my robot is parked at the local high school twice a day. Its cameras see humans in all of our weirdness. How we move. Where we go. Where we walk. Who we talk with. What you are wearing. Whether your hair was combed this morning. That data will lead to robotics breakthroughs. Apple might keep up with its Vision Pro data, but it is too freaked out by the privacy implications of using said data. (On the front are six cameras and a couple of TOF -- Time Of Flight -- sensors that can see everything in your home in great detail). Google has a lot of data, for sure. All my: 1. Email. 2. Calendars. 3. Photos. 4. TV watching behavior. 5. Contacts. 6. Documents and spreadsheets. 7. Files. 8. Location data. So I expect Google's robot will be attractive to many. But how do you see the others shake out over the next five years? Make some guesses. But remember what an AI pioneer told me years ago about AI: it's all about the data. The Chinese ones have huge advantages: the Chinese have more data on their citizens, and many more citizens to boot AND they can make robots cheaper than we can. But now that you know OpenAI is building its own robot you have caught wind of what I've heard from many in San Francisco and Silicon Valley: that humanoid robots are the real prize of AI and will be highly profitable for those that can make them and find customers willing to buy them. Here, too, I learned long ago never to bet against Elon Musk. Will you?show more

Robert Scoble
33,804 views • 1 year ago
Trained on zero real-world data. Learned to walk, pick... up boxes, and follow multi-step instructions... in the REAL world. ( 📌 Paper below) Researchers from Amazon FAR, Berkeley, Stanford, and CMU scanned real rooms with an iPhone, rebuilt them as 3D Gaussian Splatting scenes, then generated 48,000 synthetic trajectories of a Unitree G1 walking, grasping, and placing objects inside those virtual replicas. They rendered the robot's first-person camera view from each run and paired it with the matching language instruction and motion data. That's the dataset every humanoid team needs and nobody has: synced egocentric video + language + kinematics, at scale. Instead of collecting it in the real world, they manufactured it. They trained a vision-language-kinematics policy on that synthetic data alone, then deployed it on the physical G1 across five task types: navigation to a named object, lifting boxes of three different sizes with no per-size tuning, chained multi-step tasks, robustness to mid-task layout changes and flickering lights, and multi-minute long-horizon runs. No real-world fine-tuning at any point. Real-world interaction data has been the hard limit on humanoid learning... slow, expensive, and small. If scanning a room once and synthesizing thousands of labeled interactions holds up as a general recipe, that limit moves. Data stops being the bottleneck robotics teams have to solve for. 📌 Paper: Project: ——- Weekly robotics and AI insights. Subscribe free:show more

Ilir Aliu
12,950 views • 1 month ago