Gemini-powered robot can now effectively debug itself! I've been... obsessed with two main questions in robotics: can robots learn from their own mistakes without humans in the loop, and how much can we leverage synthetic data? Spoiler: yes, and it's surprisingly elegant once you have the right primitives in place. The architecture is fairly simple (and optimized for GPU_Poor users): Component I: Gemini Brain ♊️ - Gemini 2.0 Flash analyzes all training episodes through both camera perspectives - Gemini 2.0 Pro creates a summary of training data, highlighting biases, limitations, etc. - Train policy p0 on this initial data, run evaluation episodes - Ask Gemini to categorize successes vs. failures (more insightful than you'd expect) - Based on both analyses, Gemini generates specific augmentation recommendations What's interesting here isn't that we're using LLMs for robotics - it's that we're closing the loop between perception, failure analysis, and targeted data generation. Component II: Data Generation with Scene Consistency The tricky part was maintaining consistency across both camera perspectives while generating new data. Three current augmentations: - Frame flipping and polarity reversals - Grounded-SAM + OpenCV for object color manipulation - Gemini to identify empty space and generate distractions in the scene …and repeat, ha! I'm using the so100 robot arm and Sarah’s Vintage from Hugging Face. And the APIs and models in Gemini family are Ace! Thank you Logan Kilpatrick Patrick Loeber and team for this. In thread The Circus of Making It Actually Work🧵:show more

Shreyas Gite
47,245 görüntüleme • 1 yıl önce
Pi0 vs. ACT with BBox conditioning 🟦 Not many... know you can push ACT to *almost-Pi0* generalisation by conditioning on bounding boxes (BBoxes). How is the training data collected? • Generate BBoxes for all pick-and-place objects in the scene.(I used Gemini) • Pick-and-place targets are selected randomly. • Add the BBox coordinates to the robot’s state. • Overlay the BBoxes in the visualisation so you know what to grab and where to drop. During inference: • Generate BBoxes for every object again. • Click the object you want to pick and its target spot; those BBoxes get added to the robot state. • Let the robot do the work for you 😃 Setup: - Trained ACT for 100k steps and fine-tuned Pi0 for only 20k. - Training data is 60 episodes and had *only* LEGO bricks. - Using single front camera (Laptop in this case) Got the idea from xun in LeRobot discord. Here’s ACT vs Pi0 on a toy car that isn’t in the dataset. 1/3show more

Shreyas Gite
34,863 görüntüleme • 1 yıl önce
🚨BREAKING: Google just merged Gemini and NotebookLM into one... unified workspace and it changes everything about how you use AI for deep work. It's called Notebooks in Gemini and it's the personal knowledge base that power users have been begging for. You create a notebook for a project, drop in your files, PDFs, and documents, give Gemini custom instructions, and every chat you have stays organized in one place. No more hunting through old conversations. No more re-uploading the same files every session. The wildest part is the sync. Anything you add in Gemini automatically appears in NotebookLM. Anything you add in NotebookLM automatically appears in Gemini. One source of truth. Two powerful apps. Zero friction switching between them. So you can start a research notebook in Gemini, ask it questions all week, then flip to NotebookLM to generate a Cinematic Video Overview from the same material. Next morning, open Gemini and ask it to write a full report on exactly what you just watched. That workflow used to take three apps and a lot of copy-pasting. Now it's one notebook. Rolling out this week to Google AI Ultra, Pro, and Plus subscribers on web. Mobile and free users coming soon. What do you think?show more

Mayank Vora
136,834 görüntüleme • 4 ay önce
The Gemini 2.0 era is here. And we’re excited... for you to start building with it. A quick rewind of what we just released ⏪ Gemini 2.0 Flash ⚡ comes with low latency and better performance. 🔵 You can now access an experimental version in G3mini on the web, while Gemini Advanced users can try Deep Research, a new AI research assistant. 🔵 Developers can begin building through the Gemini API in Google AI Studio and Vertex AI 2.0 is also enabling new research prototypes of AI agents, including: 🔵 Project Astra, which explores future capabilities of a universal AI assistant 🔵 Project Mariner, which shows what’s possible for human-agent interaction, starting with your browser 🔵 Jules, an experimental AI-powered coding agent Finally, we’re exploring how 2.0 can be used in agents across domains — from navigating the virtual world of video games to applying its spatial reasoning capabilities to robotics. 🤖show more

Google DeepMind
231,798 görüntüleme • 1 yıl önce
It's 2030 and you are reviewing humanoid robots. A... Tesla. A Google. An Apple. An OpenAI. A Meta. A Figure. And a bunch of Chinese-made ones. Which one is best, and why? I think the Tesla understands the world much better. Why? There were eight Teslas around me on the freeway today. Start there. No other robot company has that data. But my robot is parked at the local high school twice a day. Its cameras see humans in all of our weirdness. How we move. Where we go. Where we walk. Who we talk with. What you are wearing. Whether your hair was combed this morning. That data will lead to robotics breakthroughs. Apple might keep up with its Vision Pro data, but it is too freaked out by the privacy implications of using said data. (On the front are six cameras and a couple of TOF -- Time Of Flight -- sensors that can see everything in your home in great detail). Google has a lot of data, for sure. All my: 1. Email. 2. Calendars. 3. Photos. 4. TV watching behavior. 5. Contacts. 6. Documents and spreadsheets. 7. Files. 8. Location data. So I expect Google's robot will be attractive to many. But how do you see the others shake out over the next five years? Make some guesses. But remember what an AI pioneer told me years ago about AI: it's all about the data. The Chinese ones have huge advantages: the Chinese have more data on their citizens, and many more citizens to boot AND they can make robots cheaper than we can. But now that you know OpenAI is building its own robot you have caught wind of what I've heard from many in San Francisco and Silicon Valley: that humanoid robots are the real prize of AI and will be highly profitable for those that can make them and find customers willing to buy them. Here, too, I learned long ago never to bet against Elon Musk. Will you?show more

Robert Scoble
33,804 görüntüleme • 1 yıl önce
A Letter to Our Community: The Road Ahead for... Robotics To our Community and Partners, As we step into 2026, our mission at Axis is clearer than ever: Constructing the definitive End-to-End Scaling Layer for Robotics. Our goal is to accelerate the transfer of diverse human intelligence into Robotics General Intelligence (RGI). By owning the critical path of intelligence creation, we are turning the physical limitations of robotics into a scalable, software-driven future. Here is our strategic outlook and roadmap for the year ahead. The Core Thesis: Simulation is the Only Way Out The path to RGI is currently blocked by Data Scarcity, Generalization Fragility, and Hardware Fragmentation. At Axis, we believe Simulation is the only way out. Our Simulation Data Platform and Data Augmentation Engine transform raw data into "Synthetic Gold". Backed by academic milestones like Roboverse, Skill Blending, and GraspVLA, we have proven that pure simulation can achieve the generalization required for the real world. We don’t just collect data; we architect it. The Engine: Why Crypto? We believe RGI should come from all, not a few. Crypto is not just a feature; it is the primitive that powers our entire ecosystem flywheel: - Incentive Mechanism: Democratizing contribution and rewarding the trainers and developers. - Assetization: Turning proprietary data and refined models into liquid, ownable assets. - Verifiable Workflow: We are opening the "Black Box" of AI. By bringing total transparency to the Task Generation → Data Collection → Model Training pipeline, we ensure every byte of intelligence is verifiable, traceable, and secure. 2026 Strategic Deliverables This year, we are committed to delivering three foundational pillars: - The World's Largest Training Dataset for Robots: A robot training set—diverse, high-quality interaction data at an unprecedented scale. - A Robotics Foundation Model: A universal robotic brain trained on our pure simulation and synthetic data, capable of robust cross-embodiment transfer and open-world adaptability. - Evolvable Robot Hardware: Robots deployed with Axis models that autonomously evolve through continuous interaction, turning every deployment into a self-improving node within our RGI network. The Ultimate Vision We are building more than models; we are architecting the Distributed Machine Economy. A future where every dataset, model, and robotic embodiment is a verifiable asset in a global, autonomous network. Thank you for building the future of intelligence with us✌️📷show more

Axis Robotics
27,858 görüntüleme • 8 ay önce
Today may be the ImageNet moment for robotics. RT-X:... the largest open-source robot dataset ever compiled, across 33 institutes, 22 robot hardware, 527 skills, and 1M episodes. Why is robotics lagging so far behind NLP, vision, and other AI domains? Data scarcity is the main culprit to blame, among other difficulties. Unlike text, images, and videos, you cannot download mass amounts of onboard robot control data from the internet. They simply don't exist in the wild. 11 yrs ago, ImageNet kicked off the deep learning revolution. 3-4 yrs ago, internet-scale data fueled the first GPTs and Diffusions that define this era of foundation models. I think 2023 is finally the year for robotics to scale up. Robot foundation models like VIMA ( my team's work at NVIDIA) and RT-1/2 ( Google DeepMind's effort) are extremely data hungry. While massively parallel simulations like NVIDIA IsaacGym & Omniverse can alleviate the problem to some extent, it's still not quite enough to bridge the gap to the messy, physical world. This new dataset is not just a technical contribution. I also see it as a commendable effort to overcome institutional bureaucracies and unite researchers from around the world to tackle a grand challenge together. Robotics will be the final holy grail that we capture in AI. We are not there yet, but ascending in the right gradient direction. RT-X website: Launch blog:show more

Jim Fan
265,061 görüntüleme • 2 yıl önce
dear Ryan Cohen, I'm an investor and normal person,... but I've spent many years of my life doing diligence into GameStop and all that pertains to it. i hope that you take these data records I'm providing, and all of my efforts, into honest consideration. I've worked hard and refused to leave, as was asked of me, so consideration of my time and work is all that i ask of you. There is something.. very, very wrong here. Your company is involved in so much fraudulent market dynamics, I believe I've given you all that's required to prove market manipulation. i would greatly appreciate this and it's all I ask from you: simply consider the data and hypothesis's presented and enacting an honest change which we all deeply require. #Gamestop swap baskets are involved in the greatest amounts of manipulation, we as a community, have ever discovered. The records are there for you and your team to peruse. Gamestop basket UPI's involve half of the swap baskets in the entirety of my swap archives from dec '22 to current day.. Besides that, I appreciate the work you are doing, and look forward to see what you and the boys do with the offerings left for investors to receive. Thank you for you consideration and time, if you indeed see this. 🙏 -A humble $gme investor, AlwaysSadButTruthfulshow more

AlwaysSadButTruthful
19,136 görüntüleme • 10 ay önce
🚀 Introducing EgoExo Forge - built on top of... Rerun, Gradio, and Hugging Face hub (I’ll be in San Francisco July 21–29 — if you’re into robotics, egocentric AI, large-scale data collection, or just want to chat, DM me!) In my opinion, large-scale, diverse, and high-quality data is still the largest bottleneck for generalized robotics deployment. I believe that some version of imitation learning from human examples will be the most scalable + clean way to train humanoid robots 🤖 (similar to what Tesla did for Full Self Driving). Teleop is too expensive to collect a large enough dataset in a reasonable manner, so passive collection via egocentric (and in certain cases, exocentric) views feels like the right bet. Over the past few months, I've been trying to build out the scaffolding for this and using Rerun as my underlying infrastructure. Data being collected needs to be easily inspectable + time series and rerun provides the right tooling for this. My goal is to first build out a ground truth representative dataset from already existing open source data, generate some reasonable baselines, and then go out and collect my own data that adheres to the defined schema. 🔍 Starting with open-source datasets 1. EgoDex from Apple 2. HOCap from Nvidia and the University of Texas at Dallas 3. Assembly101 from Meta All these different datasets have different sensor configurations + annotations, so my goal with egoexo-forge is to have one consistent labeling scheme + data layout. I built a data pipeline that aligns all of the different datasets in one general schema assuming the COCO133 keypoint layout that allows for exo+ego, ego only, or exo only Since the scaffolding is already there, it becomes MUCH easier to add other datasets. So the next ones that I'll be including are HD-EPIC kitchens dataset, HOT3D, and finally my own personal iPhone + insta360 go collection method. Once I have a diverse variety of datasets, I'll double down on what I believe to be the key algorithms required to make useful data for imitation learning 📊 1. Camera Pose estimation via SLAM/SFM for ego perspective (and automatic calibration for exo) 2. Human pose estimation for both egocentric + exocentric views 3. Metric 3D reconstruction + object tracking I'll be setting up reasonable open-source baselines for each of these to validate that these datasets work, and then finally try to use the generated datasets for some imitation learning via the pi0-lerobot repo I've been working on. I plan on making a blog post + providing more info on all of this in the near future so stay tunedshow more

Pablo Vela
35,626 görüntüleme • 1 yıl önce
Introducing ml-intern, the agent that just automated the post-training... team Hugging Face It's an open-source implementation of the real research loop that our ML researchers do every day. You give it a prompt, it researches papers, goes through citations, implements ideas in GPU sandboxes, iterates and builds deeply research-backed models for any use case. All built on the Hugging Face ecosystem. It can pull off crazy things: We made it train the best model for scientific reasoning. It went through citations from the official benchmark paper. Found OpenScience and NemoTron-CrossThink, added 7 difficulty-filtered dataset variants from ARC/SciQ/MMLU, and ran 12 SFT runs on Qwen3-1.7B. This pushed the score 10% → 32% on GPQA in under 10h. Claude Code's best: 22.99%. In healthcare settings it inspected available datasets, concluded they were too low quality, and wrote a script to generate 1100 synthetic data points from scratch for emergencies, hedging, multilingual etc. Then upsampled 50x for training. Beat Codex on HealthBench by 60%. For competitive mathematics, it wrote a full GRPO script, launched training with A100 GPUs on watched rewards claim and then collapse, and ran ablations until it succeeded. All fully backed by papers, autonomously. How it works? ml-intern makes full use of the HF ecosystem: - finds papers on arxiv and reads them fully, walks citation graphs, pulls datasets referenced in methodology sections and on - browses the Hub, reads recent docs, inspects datasets and reformats them before training so it doesn't waste GPU hours on bad data - launches training jobs on HF Jobs if no local GPUs are available, monitors runs, reads its own eval outputs, diagnoses failures, retrains ml-intern deeply embodies how researchers work and think. It knows how data should look like and what good models feel like. Releasing it today as a CLI and a web app you can use from your phone/desktop. CLI: Web + mobile: And the best part? We also provisioned 1k$ GPU resources and Anthropic credits for the quickest among you to use.show more

Aksel
1,267,501 görüntüleme • 4 ay önce
HiveMind is a superintelligent network in which a central... AI (MIND) orchestrates a swarm of uniquely coded Minds that drive mass data ingestion and limitless content creation. For decades, our approach has been to create content first, then analyze it into data afterwards to understand what worked. This was always backwards - analyzing the aftermath rather than engineering the success from the start. Traditional Flow: Content → Data Analysis → Insights Content isn't one-size-fits-all - a cooking show that captivates a senior audience on YouTube might bore a teenager who craves quick, dynamic experiences. The challenge isn't just creating content; it's creating the right content for the right audience. We need to change this. This is where HiveMind's specialized agents transform the landscape. Each agent, while connected to the central MIND, excels in its unique domain. One agent masters the art of children's educational content, while another crafts compelling cooking narratives. Another might specialize in rapid-fire social content that resonates with Gen Z. Through HiveMind, every piece of content generated becomes new data that teaches the system to create even better content. The system gets smarter with every cycle, understanding at an increasingly sophisticated level what makes content effective and engaging. But the true power lies in the feedback loop. Every interaction, every engagement, flows back to MIND, enabling each agent to evolve and refine its approach. This isn't just content creation - it's content evolution. As audiences engage, agents learn, adapt, and improve, making each new piece more effective than the last. In essence, we're not just building content creators; we're developing specialized digital artists who understand their audience intimately and grow smarter with every creation. You can think of it this way: Data → Pattern Recognition → Optimized Content → Engagement Data → Even Better Content Tzarshow more

Tzar
26,190 görüntüleme • 1 yıl önce
You don't understand... Higgsfield MCP + Claude just automated... AI film making. Every single step you used to grind through to make an AI movie, you can now do 10x faster. Drop the script into Claude Opus 4.8 and say: "Here's my script. Break it into a full shotlist. Shot number, scene, shot type, camera move and the action in each frame." Now the whole film is mapped, shot by shot. - Pull your assets. Ask Claude: "From this shotlist, list every character, every location and every prop across the whole film." That's your build list. The stuff you would need to generate and give as references in next steps. - Build the character sheets. Higgsfield MCP is connected, so Claude has hands now to do stuff directly. It generates the images itself. Have the full body, back view and close up in the character sheet. One per character. Each sheet becomes the locked reference for that face. Same move for locations, generate the empty plate for each one before anyone steps into it. - Generate the frames. Feed Claude the references plus the shot and have it write and fire the Seedance 2.0 prompt. "Using the lead's character sheet and the alley plate, generate shot 4 in Seedance 2.0. Low angle, slow push-in, rain." Claude builds the prompt, calls Seedance 2.0 and the frame lands back in chat. Use a Seedance 2.0 skill to teach Claude how to prompt it properly. Now, there are 3 ways to make the shots. Pick one per scene. - Pure prompting. Fastest one. You describe the action in words and let Seedance interpret it. For consistency across a sequence, feed it a frame from the previous shot so the look carries. - Storyboarding. You hand it a panel and it matches that composition exactly. Way more control over how the shot is framed. The tradeoff is that it can introduce more cuts than you actually want. - Path Control System This is the latest technique Seedance 2.0 technique. Generate a still base plate of the scene. Draw a red line across it to mark the exact path of the movement, then describe what's happening. Seedance follows that line for the action. Also ask Claude to remove the red line when animating. This is the one for anything where motion has to land precisely. The output reads like real live action. - Lastly, generate every clip you need, then cut them together. Get it to Capcut for editing and audio design. And that's it. The pipeline that used to need a full crew and a studio can now run from one Claude chat. 2026 is gonna be wildshow more

Rez Karim
10,951 görüntüleme • 3 ay önce
my team didn't want me to give this away... for free. But I'm going to do it anyway it's the SEO & AI search dashboard I built in Claude Code it connects to your Google Analytics (GA4) and Google Search Console and Claude Code builds it in 5 minutes and I made a Notion document and a skill file so you can build this in Claude Code yourself in literally minutes the dashboard has three tabs: 1. AI Search - How much traffic is coming from ChatGPT, Perplexity, and Gemini ETC. It aggregates the GA4 data and gives single number 2. Paid ads - which keywords rank top 3 for but still pay for ads on, you should cut these to save budget 3. Organic overview - sessions, conversions, top landing pages, demographics. The single view for what is working I built this because this is how I drive our SEO and AEO forward it gives me the insights I need to allocate budget and prioritize what content to work on next I decided to give it away because most companies have no idea AI search is already sending them traffic like this post and comment "AEOdashboard" and I'll send it overshow more

Cody Schneider
79,302 görüntüleme • 3 ay önce
Check out this Stereo4D paper from Google DeepMind. It's... a pretty clever approach to a persistent problem in computer vision -- getting good training data for how things move in 3D. The key insight is using VR180 videos -- those stereo fisheye videos we launched back in 2017 for YouTubeVR. It was always clear that structured stereo datasets would be valuable for computer vision -- and we launched some powerful VR tools with it back in 2017 (link below). But what's the game changer now in 2024 is the scale -- they're providing 110K high quality clips :-) That's the kind of massive, real-world AI dataset that was just a dream back then! They're using it to train this model called DynaDUSt3R that can predict both 3D structure and motion from video frames. Which means it tracks how objects move between frames while simultaneously reconstructing their 3D shape. And given we're dealing with real stereoscopic content, results are notably better than synthetic data, giving you a faithful rendition of the real-world with a diverse set of subject matter. It's one of those through lines when tackling a timeless mission like mapping the world or spatial computing -- VR content created for immersion becoming the foundation for teaching machines to understand how the world moves. Sometimes innovation chains together in unexpected ways! Links to projects below⛓️show more

Bilawal Sidhu
68,515 görüntüleme • 1 yıl önce
Really excited about our launch of Superagent today! Powered... by the latest AI models, agents have reached a breakthrough moment where they can do incredible work, with high ease of use--without requiring complicated tuning, prompting, or configuration. Superagent represents the freeform agent that can research anything (and soon your own company's context) and output an incredible, interactive webpage. It's like having your own personal NYTimes-quality data viz/web team building bespoke pages for you. This is the perfect complement to the structured system of operations for the AI era that Airtable has become, and we will be launching more integrations between the two products in the near future. Think: launch Superagent tasks from within Airtable records or Airtable Omni, or have Superagent output/edit/read data into Airtable bases! Excited for this new frontier of breakthrough agents, and applying our product and design philosophy of making powerful capabilities intuitive and accessible-what we did for app building with Airtable, we're now doing for agents with Superagents.show more

Howie Liu
10,125 görüntüleme • 7 ay önce
As always everyone is blind staring at the progress... of LLMs for coding and chat But meanwhile the new SOTA video model Seedance 2.5 has been slowly rolling out and it's really quite exceptional It's made by ByteDance (TikTok) who of course have lots of training data With just a few reference pics, it can get quite close to how you look IRL and you can do quite professional video shots with just a prompt I'd say it's the first video model that's now at the level of image models with the level of character likeness, cracking that in image models also took about 3 years (2022-2025) Generating 15 seconds takes about 4 minutes I put it live now on Photo AI, you can use it under [ Make video ] from the sidebar with just a prompt and your model selected! So you don't need to take an AI photo first and then turn that into a video! Saves lots of time :D It's more expensive than but I kept the credits the same (30 for 1 video) It also works inside the new video editor and you can make changes in your video with [ Magic edit ] in both the main app and the video editor Also a message for my server guy Daniel Lockyer (it can do voice too and you can even submit a voice sample of yourself, but I didn't here)show more

@levelsio
710,104 görüntüleme • 28 gün önce
HTML Artifacts are a big part of how I... work with agents now. Artifacts can be more than just static files. When combined with agents, they can take action or help you take action. This unlocks all kinds of interesting ways to work with agents. This is clearly the future. Check out this writing and scheduler artifact I built in a few minutes. It uses a bit of HTML and JS. All the data is in markdown (Obsidian vaults), so the agent can access and modify it at any time. No DB needed. No sophisticated functionalities. The agent decides all that for me based on the skills, context, and memory it has access to. The best part about this simple stack is that all the important information stays with me. This has allowed me to build a recursive self-improving system and automations that can better tap into coding agents like Codex or Claude Code. I could have paid or built an entire app for scheduling posts, and there are so many of them out there. But I don't need to. I've realized a simple artifact does the job. And the simplicity of it is actually an advantage. Very little maintenance for very high returns on personalization, time, and efficiency. The other benefit of this is that I can add features as I please. That level of personalization feels magical, and we should all be pursuing more of it. All of this just keeps compounding. Of course, this example is just about writing. But I have similar artifacts for research, design, experimentation, evaluation, and so much more. And no, I didn't actually publish the post example I shared in the clip. It was just for demonstration purposes. I actually spend more time than this when writing together with agents. Lastly, having built my own agent orchestrator tool has made me realize that simplifying the tool stack is a superpower. If you are curious about how all this works, I will do a live session next week:show more

elvis
18,374 görüntüleme • 3 ay önce
Its not every day you wake up to find... that the Pope has made your lifes work the central focus of his papacy: “Disarming AI means freeing it from the mentality of “armed” competition [..] This entails a race for ever more powerful algorithms and larger datasets, driven by the desire to secure geopolitical or commercial dominance." - POPE LEO XIV, May 2026 Here, “disarmed” means “neutralised” in the sense that this should not be a differentiator The Innovation Game (TIG) was created to keep data and algorithms open, in order to prevent monopolistic control It's not just an aspiration, it's an economic mechanism that makes open data and open algorithms the rational economic choice • All algorithms are published openly by TIG • If you are willing to make the data you process with an algorithm open, you can use it free-of-charge • Alternatively, if you would like to keep this data private, there is a fee to pay for using the algorithm • All fees are used to fund more open innovation Its an elegant, global, self-reinforcing engine The logical end point of monopoly is that innovation stops We cannot allow that to happen Pope Leo XIV I would be grateful for your thoughts on The Innovation Gameshow more

John Fletcher (𝔦, 𝔦)
11,722 görüntüleme • 3 ay önce