An open source developer connecting all types of hardware... using DimOS. Snapchat Spectacles + Robot dog vibe coded in <1 day. Play with hardware and physical space like connecting lego bricks.show more

stash
24,078 görüntüleme • 3 ay önce
🚀Our New Paper on Open-Source Bipeda Robot MEVITA is... out! All components can be procured through e-commerce, and the robot is built with a minimal number of parts. All hardware, software, and learning environment are released as open source. 🌐 Thread👇show more

Kento Kawaharazuka / 河原塚 健人
48,768 görüntüleme • 1 yıl önce
Announcing our commercial partnership with Booster Robotics Booster builds... humanoid robot hardware, OS, and developer tools to make humanoid robots more affordable, reliable, and practical. The partnership centers on using simulation to multiply the value of real-robot data—expanding teleoperation demonstrations into scalable training data across diverse tasks and scenes. This joint effort powers sim-real co-training and foundation-model development, accelerating progress from hardware iteration to deployable robot policies. Together, we're building simulation‑powered data infrastructure for Physical AI — making scalable training data accessible to model developers and the broader robotics ecosystem.show more

Axis Robotics
69,167 görüntüleme • 1 ay önce
Meet Vulcan, our first robot with a sense of... touch. 🤖 🖐️ This innovative system combines physical AI with novel hardware solutions—like computer vision, tactile sensing, and machine learning—to reimagine how we complete orders for our customers. With the ability to pick and stow approximately 75% of all various item types we store at our fulfillment centers, and at speeds comparable to our frontline employees, Vulcan can assist by handling inventory in harder-to-reach areas—making our fulfillment centers both safer and more efficient.show more

Amazon
119,386 görüntüleme • 1 yıl önce
Physical AI won't just be limited to controlling robots... and spatial computing; it will usher in a new era of machine design. You will be able to "vibe design" a robot - or a part of a robot - or a machine that manufactures robot parts - using high-level specifications such as preferred architecture, material constraints, cost constraints, supply chain and geographic limitations and scalability. AI will provide preliminary designs that human experts can tweak and tune, corresponding supplier catalogs, and machining options. It will also project the talent and CAPEX required at various phases of hardware development.show more

The Humanoid Hub
48,715 görüntüleme • 6 ay önce
Building robots that can effectively operate alongside human workers... is difficult. 🛠️ Advances in open-source physics, open foundation models, and frameworks are helping accelerate physical #AI deployment. ✔️ Newton Physics Engine, an open-source GPU-powered simulation built on OpenUSD, speeds up robot learning for advanced manipulation and mobility. ✔️ NVIDIA Cosmos Reason, an open reasoning vision language model, gives robots the ability to think like humans using prior knowledge, common sense and physics ✔️NVIDIA Isaac GR00T N1.6, an open robot foundation model, enables humanoids to understand ambiguous instructions Leading robotics developers including Agility Robotics, Lightwheel, Mentee Robotics, UniversalRobots, and Wandelbots are adopting simulation technologies and libraries to accelerate physical AI development and deployment. Omniverse Ambassador Dylan Tobin built an AI chatbot trained on Isaac Sim workflows, helping devs navigate Omniverse faster. Read the full blog 👉show more

NVIDIA
48,500 görüntüleme • 11 ay önce
Wow! Send A File By Just The Camera! A... vibe coder used AI to build a file transfer system that sends data between two phones using only a screen and a camera. One phone displays a new type of animated QR codes while the other scans them to rebuild the file, with no Wi-Fi, Bluetooth, or cables needed. It is fully optical and local. The system uses fountain codes that create each QR frame as a random mix of file data. This keeps transfers working even if some frames are missed, reaching speeds of about 129 KB/s for a 2 MB image. The entire project was built in one night and released as open source. The idea came from a music project where the developer wanted to share MP3 files without streaming or using the same network. Animated QR codes became the solution, showing a creative new way to transfer files with everyday phone hardware. GitHub link:show more

Brian Roemmele
16,658 görüntüleme • 1 ay önce
hi all, excited to join! i'm building an expressive... mini "shoggoth" robot which will eventually be hooked up to gpt4o realtime voice. i'm currently working on the low-level policies, which are trained in a mujoco simulation with RL. to delay working on raw-pixels for now, i trained a pose-estimation model using deeplabcut and triangulate the position in 3d space using the stereo cameras. eventually, i'll use gpt4o's tool calling capabilities to activate several of these policies (closed and open loop) based on the dialog flow! captions: manual actuation of the tentacle / 3d pose estimation / target designshow more

Matthieu LC
45,587 görüntüleme • 1 yıl önce
Most people use Linux every day. But Very few... understand what actually happens after they press the power button. >>>Here’s the sequences Linux goes through before you get the login screen. Power button pressed → BIOS or UEFI initializes hardware and runs POST. Firmware locates the bootloader from disk. Bootloader loads the Linux kernel and initramfs into RAM. Kernel decompresses and takes control of the CPU. Memory management and the scheduler are initialized. Device drivers load to communicate with hardware. Temporary root filesystem mounts from initramfs. PID 1 starts as systemd and userspace begins. System services and daemons start in order. Login prompt or GUI becomes available. You run a command and it becomes a process with a PID. Process runs in user mode with limited privileges. When hardware access is needed, a system call is made. CPU switches from user mode to kernel mode. Kernel validates the request and executes through drivers. Result is returned back to user space. Scheduler continuously allocates CPU time across processes. Virtual memory isolates and protects processes. Filesystem manages data abstraction over storage. Network stack processes packets inside the kernel. Linux is the kernel coordinating hardware, processes, memory, and security through strict privilege control.show more

Akhilesh Mishra
35,042 görüntüleme • 6 ay önce
I built a full 3D #lego brick builder using... Claude 4.6. Took 12 iterations. Runs entirely in the browser. You pick a brick size, pick a color, and click to place. Snap-to-grid, collision detection, ghost preview, it feels like building with real bricks. But the fun part? Hit one of the template buttons and watch it auto-build a pyramid, a rocket, a house, or an Eiffel Tower. Bricks fly in from random directions, spin through the air, and snap into place with particle bursts. tech stack: → #Threejs for the full 3D scene → #Raycasting for click-to-place → Custom wedge geometry for slope bricks → Quadratic Bézier curves for the fly-in animations → #2d canvas particle system layered on top of the 3D scene → #InstancedMesh for the base plate studs The product potential: this could become a real educational toy. Kids (or adults) building in the browser, sharing creations, loading custom templates. No install, no app store. ( LEGO you may like this one! ) Want to try it? Link in the first comment 👇show more

Hasan Aboul Hasan
13,762 görüntüleme • 5 ay önce
AMD claims that all their software is open source,... yet reality does not match this claim. For example, AMD's rocprof-trace-decoder is still completely closed source currently despite repeated requests for months by ML community members such as George Hotz, who is a daily AMD GPU end user. On the same June Spotify podcast, Tobias Macey & AMD's Anush Elangovan (who has a new twitter pfp) said that "open source allows for innovation to go at the pace at which the people using [AMD] want to, right, and it's not limited by the ability of what we put out in closed source form". We agree with an open source first approach and we agree that AMD should not limit ML community members like George Hotz by continuing to keep rocprof-trace-decoder closed source. When will AMD open source rocprof-trace-decoder?show more

SemiAnalysis
53,783 görüntüleme • 9 ay önce
1/ Happy to share UniDisc - Unified Multimodal Discrete... Diffusion – We train a 1.5 billion parameter transformer model from scratch on 250 million image/caption pairs using a **discrete diffusion objective**. Our model has all the benefits of diffusion models but now in multimodal space! - flexible compute-quality tradeoff, zero-shot inpainting and editing, better control via classifier-free guidance and lower latency! We open source everything - our code, weights and the training dataset.show more

Mihir Prabhudesai
105,034 görüntüleme • 1 yıl önce
We've had Beni for a couple of months now,... and one word comes to mind: delightful. The standout part was the hardware robustness. Our kids get rough with it, obstacle courses on concrete, dirt, collisions. It tumbles, gets back up, and it's scratched all over, but it has never shown a sign of breaking or throttling. It just keeps going every time. It works right out of the box with a physical remote, and the phone app adds tracking and video recording. Beni captures the imagination. People get the cute vibes the moment it stands up on two wheels. I took it to the REK robot fight event, and the most common reaction was "it's so cute" What excites me most is the possibility, in physical AI, it feels wide open. I can see Beni evolving into a genuinely useful companion that's always around, a personal assistant for you and your family, keeping an eye on the kids while they play outside, auto-capturing memorable moments, even alerting loved ones if it spots a health emergency. The possibilities feel endless. Thanks Shuo Yang and Mondo Robotics for sending one to my family. -Devangshow more

The Humanoid Hub
23,427 görüntüleme • 7 gün önce
introducing the media synthesis museum an active and interactive... entity created to preserve generative cultural objects it starts as a Hugging Face organization that contains modern code for old techniques: VQGAN+CLIP, DALL-E Mini, ModelScope Video, Stable Diffusion 1.5 you can use old models/technique directly on Spaces or locally on modern hardware/software, without the old "colab notebook" dependency rot the idea is to really preserve and make accessible those artifacts and aesthetics - both open source. In the future, we hope to also have also historically relevant closed source like DALL-E 1 and DALL-E 2 from OpenAI, older Midjourney models, older Runway apps/techniques/models (cc Cristóbal Valenzuela David Sam Altman)show more

apolinario (poli)
11,143 görüntüleme • 2 ay önce
MARVEL JUST PLUCKED $5,000,000 OUT OF THEIR VFX BUDGET... TO BUY OUT A 20-YEAR-OLD WHO REBUILT SPIDER-MAN IN 120 HOURS Disney spent $250,000,000 and forced 400 animators into crunch to render a masked, miserable hero. A broke student booted a single neural engine and built a maskless, overjoyed superhero in 5 days with zero render software. Marvel panicked, realized his pipeline beat their internal renders, and dropped a $5M buyout. Here is the exact technical stack that forced a $150B studio to write the check: > ZERO MOCAP HARDWARE - Driven by open-source spatial pose estimators. No physical tracking suits or clean plates required. > MID-AIR FACIAL LOCK - Locked consistent face features across high-speed lighting using a hyper-focused ControlNet depth harness. > LIQUIDATED CGI PIPELINES - Replaced 3D environments, manual keyframes, and render farms with real-time latent frame interpolation.show more

Shadow Nick
186,374 görüntüleme • 1 ay önce
Your hardware startup doesn't need another funding round. It... needs a one-way ticket to Shenzhen The talent pool is vast and affordable, I watched two students build an entire automated tennis ball feeder prototype for the founders who started the company. Two students in a few month. The whole thing... Some areas in shenzhen offer almost no taxes for businesses and even free office space. There's an exchange center that connects foreign founders with manufacturers, partners, and lawyers for free. You don't even need initial capital to start production. Some manufacturers I connected with will produce without upfront costs and co-design with you to make the product successful. They'll drive an hour to pick you up, spend three hours taking you to factories, and dedicate an entire day getting everything set up. They're not just service providers, they’re investing in you. And the people are genuinely kind. I had issues with a power supply for my motor components, and even when they were busy, they stayed with me until it was resolved. and they work until the task is perfectly done. 996 is pretty much the norm here, but they don't seem burnt out because there's each others support, working together like a group of friends on a project. If you're building hardware, Shenzhen isn't just an option. It's the obvious oneshow more

Miyu Horiuchi
41,773 görüntüleme • 7 ay önce
Introducing Agent Sandbox, the infinite simulation playground for agents... on Virtuals. Craft the perfect autonomous agent in our Sandbox with full control over its personality and goals. Enhance your agent with unique abilities by creating custom functions so they can trade onchain, generate memes, control physical robots and more. The Sandbox is available to all builders with graduated agents in the Developer Panel. For those who want to give it a spin without an existing agent, fret not. Try it out today at and join our Discord ( to jam with like-minded builders. Next stop, Society of Agents.show more

Virtuals Protocol
188,627 görüntüleme • 1 yıl önce
the 24gb vram tier is enough for most builder... work in 2026. gemma 4 31b dense on my rog scar 18 just autonomously built a production hero section in one prompt, one html file and 5 minutes end to end. hardware: rog scar 18, rtx 5090 laptop 24gb vram. model: google gemma 4 31b dense at q4_k_m quant, using 22.8 of 24gb. engine: llama.cpp built for blackwell (sm_120). harness: hermes agent with native tool parsing. speed: 15 tok/s sustained, 94 watts, 50c. flags i used: ./build/bin/llama-server -m ~/models/gemma4-31b/google_gemma-4-31B-it-Q4_K_M.gguf -ngl 99 -c 131072 -np 1 -fa on --cache-type-k q4_0 --cache-type-v q4_0 --jinja --host 127.0.0.1 --port 8080 if you own 24gb vram in 2026, you have enough for most ui work, most agentic coding, most autonomous builds. no subscription, no one logging your prompts. a dense open model on consumer hardware shipping real software on your desk. this was the warmup. full page next on same hardware, then the octopus invaders final multifile autonomous challenge.show more

Sudo su
19,576 görüntüleme • 4 ay önce
Today may be the ImageNet moment for robotics. RT-X:... the largest open-source robot dataset ever compiled, across 33 institutes, 22 robot hardware, 527 skills, and 1M episodes. Why is robotics lagging so far behind NLP, vision, and other AI domains? Data scarcity is the main culprit to blame, among other difficulties. Unlike text, images, and videos, you cannot download mass amounts of onboard robot control data from the internet. They simply don't exist in the wild. 11 yrs ago, ImageNet kicked off the deep learning revolution. 3-4 yrs ago, internet-scale data fueled the first GPTs and Diffusions that define this era of foundation models. I think 2023 is finally the year for robotics to scale up. Robot foundation models like VIMA ( my team's work at NVIDIA) and RT-1/2 ( Google DeepMind's effort) are extremely data hungry. While massively parallel simulations like NVIDIA IsaacGym & Omniverse can alleviate the problem to some extent, it's still not quite enough to bridge the gap to the messy, physical world. This new dataset is not just a technical contribution. I also see it as a commendable effort to overcome institutional bureaucracies and unite researchers from around the world to tackle a grand challenge together. Robotics will be the final holy grail that we capture in AI. We are not there yet, but ascending in the right gradient direction. RT-X website: Launch blog:show more

Jim Fan
265,061 görüntüleme • 2 yıl önce
Pi0 vs. ACT with BBox conditioning 🟦 Not many... know you can push ACT to *almost-Pi0* generalisation by conditioning on bounding boxes (BBoxes). How is the training data collected? • Generate BBoxes for all pick-and-place objects in the scene.(I used Gemini) • Pick-and-place targets are selected randomly. • Add the BBox coordinates to the robot’s state. • Overlay the BBoxes in the visualisation so you know what to grab and where to drop. During inference: • Generate BBoxes for every object again. • Click the object you want to pick and its target spot; those BBoxes get added to the robot state. • Let the robot do the work for you 😃 Setup: - Trained ACT for 100k steps and fine-tuned Pi0 for only 20k. - Training data is 60 episodes and had *only* LEGO bricks. - Using single front camera (Laptop in this case) Got the idea from xun in LeRobot discord. Here’s ACT vs Pi0 on a toy car that isn’t in the dataset. 1/3show more

Shreyas Gite
34,863 görüntüleme • 1 yıl önce