An open source developer connecting all types of hardware... using DimOS. Snapchat Spectacles + Robot dog vibe coded in <1 day. Play with hardware and physical space like connecting lego bricks.show more

stash
23,732 görüntüleme • 1 ay önce
🚀Our New Paper on Open-Source Bipeda Robot MEVITA is... out! All components can be procured through e-commerce, and the robot is built with a minimal number of parts. All hardware, software, and learning environment are released as open source. 🌐 Thread👇show more

Kento Kawaharazuka / 河原塚 健人
48,281 görüntüleme • 11 ay önce
Meet Vulcan, our first robot with a sense of... touch. 🤖 🖐️ This innovative system combines physical AI with novel hardware solutions—like computer vision, tactile sensing, and machine learning—to reimagine how we complete orders for our customers. With the ability to pick and stow approximately 75% of all various item types we store at our fulfillment centers, and at speeds comparable to our frontline employees, Vulcan can assist by handling inventory in harder-to-reach areas—making our fulfillment centers both safer and more efficient.show more

Amazon
119,337 görüntüleme • 1 yıl önce
Physical AI won't just be limited to controlling robots... and spatial computing; it will usher in a new era of machine design. You will be able to "vibe design" a robot - or a part of a robot - or a machine that manufactures robot parts - using high-level specifications such as preferred architecture, material constraints, cost constraints, supply chain and geographic limitations and scalability. AI will provide preliminary designs that human experts can tweak and tune, corresponding supplier catalogs, and machining options. It will also project the talent and CAPEX required at various phases of hardware development.show more

The Humanoid Hub
48,709 görüntüleme • 5 ay önce
Building robots that can effectively operate alongside human workers... is difficult. 🛠️ Advances in open-source physics, open foundation models, and frameworks are helping accelerate physical #AI deployment. ✔️ Newton Physics Engine, an open-source GPU-powered simulation built on OpenUSD, speeds up robot learning for advanced manipulation and mobility. ✔️ NVIDIA Cosmos Reason, an open reasoning vision language model, gives robots the ability to think like humans using prior knowledge, common sense and physics ✔️NVIDIA Isaac GR00T N1.6, an open robot foundation model, enables humanoids to understand ambiguous instructions Leading robotics developers including Agility Robotics, Lightwheel, Mentee Robotics, UniversalRobots, and Wandelbots are adopting simulation technologies and libraries to accelerate physical AI development and deployment. Omniverse Ambassador Dylan Tobin built an AI chatbot trained on Isaac Sim workflows, helping devs navigate Omniverse faster. Read the full blog 👉show more

NVIDIA
48,500 görüntüleme • 9 ay önce
> raise console prices > get rid of physical... media -> you are here raise game prices to $100 and beyond > completely remove the ability to share any digital games > buying new games will become an expensive hobby, but don't worry they'll definitely make a subscription tier for new games that come out > hardware will become pricier, the cheapest way to play is using cloud gaming services > offer a subscription that is reasonable at first > gradually raise subscription prices You will own nothing and be happy If we don't act about Sony's outrageous decision to stop physical media in 2028, gaming might be doomedshow more

NikTek
754,156 görüntüleme • 24 gün önce
hi all, excited to join! i'm building an expressive... mini "shoggoth" robot which will eventually be hooked up to gpt4o realtime voice. i'm currently working on the low-level policies, which are trained in a mujoco simulation with RL. to delay working on raw-pixels for now, i trained a pose-estimation model using deeplabcut and triangulate the position in 3d space using the stereo cameras. eventually, i'll use gpt4o's tool calling capabilities to activate several of these policies (closed and open loop) based on the dialog flow! captions: manual actuation of the tentacle / 3d pose estimation / target designshow more

Matthieu LC
45,587 görüntüleme • 1 yıl önce
NO HANDS, NO WORRIES MATE - TESLA FSD IN... BRISBANE An hour through the spaghetti bowl of Brisbane motorway - no hands, no stress, just eight cameras calling the shots! Drive Rules: • Navigated roundabouts, motorways, cyclists, and pedestrians - all on board • Technically Level 2 AIDAS - eyes on the road, but hands can chill • $10K upgrade for Model Y and 3 owners with Hardware 4 “It’s in our backyard… and honestly, it was mind-blowing!” Source: Live review, Brisbane test drive, Sawyer Merritt, Teslashow more

Mario Nawfal
42,416 görüntüleme • 10 ay önce
Most people use Linux every day. But Very few... understand what actually happens after they press the power button. >>>Here’s the sequences Linux goes through before you get the login screen. Power button pressed → BIOS or UEFI initializes hardware and runs POST. Firmware locates the bootloader from disk. Bootloader loads the Linux kernel and initramfs into RAM. Kernel decompresses and takes control of the CPU. Memory management and the scheduler are initialized. Device drivers load to communicate with hardware. Temporary root filesystem mounts from initramfs. PID 1 starts as systemd and userspace begins. System services and daemons start in order. Login prompt or GUI becomes available. You run a command and it becomes a process with a PID. Process runs in user mode with limited privileges. When hardware access is needed, a system call is made. CPU switches from user mode to kernel mode. Kernel validates the request and executes through drivers. Result is returned back to user space. Scheduler continuously allocates CPU time across processes. Virtual memory isolates and protects processes. Filesystem manages data abstraction over storage. Network stack processes packets inside the kernel. Linux is the kernel coordinating hardware, processes, memory, and security through strict privilege control.show more

Akhilesh Mishra
34,995 görüntüleme • 4 ay önce
I built a full 3D #lego brick builder using... Claude 4.6. Took 12 iterations. Runs entirely in the browser. You pick a brick size, pick a color, and click to place. Snap-to-grid, collision detection, ghost preview, it feels like building with real bricks. But the fun part? Hit one of the template buttons and watch it auto-build a pyramid, a rocket, a house, or an Eiffel Tower. Bricks fly in from random directions, spin through the air, and snap into place with particle bursts. tech stack: → #Threejs for the full 3D scene → #Raycasting for click-to-place → Custom wedge geometry for slope bricks → Quadratic Bézier curves for the fly-in animations → #2d canvas particle system layered on top of the 3D scene → #InstancedMesh for the base plate studs The product potential: this could become a real educational toy. Kids (or adults) building in the browser, sharing creations, loading custom templates. No install, no app store. ( LEGO you may like this one! ) Want to try it? Link in the first comment 👇show more

Hasan Aboul Hasan
13,762 görüntüleme • 4 ay önce
AMD claims that all their software is open source,... yet reality does not match this claim. For example, AMD's rocprof-trace-decoder is still completely closed source currently despite repeated requests for months by ML community members such as George Hotz, who is a daily AMD GPU end user. On the same June Spotify podcast, Tobias Macey & AMD's Anush Elangovan (who has a new twitter pfp) said that "open source allows for innovation to go at the pace at which the people using [AMD] want to, right, and it's not limited by the ability of what we put out in closed source form". We agree with an open source first approach and we agree that AMD should not limit ML community members like George Hotz by continuing to keep rocprof-trace-decoder closed source. When will AMD open source rocprof-trace-decoder?show more

SemiAnalysis
53,783 görüntüleme • 8 ay önce
1/ Happy to share UniDisc - Unified Multimodal Discrete... Diffusion – We train a 1.5 billion parameter transformer model from scratch on 250 million image/caption pairs using a **discrete diffusion objective**. Our model has all the benefits of diffusion models but now in multimodal space! - flexible compute-quality tradeoff, zero-shot inpainting and editing, better control via classifier-free guidance and lower latency! We open source everything - our code, weights and the training dataset.show more

Mihir Prabhudesai
104,934 görüntüleme • 1 yıl önce
SOMEONE BUILT AN OPEN-SOURCE JARVIS WITH 9 AGENTS AND... 5 MEMORY BACKENDS AND YOUR DATA NEVER LEAVES YOUR DEVICE Every time you message ChatGPT or Claude your data hits a server you don't control, gets processed by infrastructure you're paying for and comes back with zero guarantee of what happened in between. OpenJarvis runs the entire stack locally - 9 agent types, 5 memory backends, a learning loop that gets smarter every day and a morning digest that connects to Google Drive and surfaces what matters before you open a single app. Most AI tools are exactly as dumb on day 100 as they were on day 1 because they forget everything when the window closes - this one indexes your documents once and automatically injects relevant context into every prompt forever. Custom agent setup for a client is $500-2,000 one time and AI infrastructure retainer is $300-800 a month - and your cost is one afternoon and an open source repo. The repo is free. The advantage it creates is not.show more

Cortex
11,380 görüntüleme • 2 ay önce
Your hardware startup doesn't need another funding round. It... needs a one-way ticket to Shenzhen The talent pool is vast and affordable, I watched two students build an entire automated tennis ball feeder prototype for the founders who started the company. Two students in a few month. The whole thing... Some areas in shenzhen offer almost no taxes for businesses and even free office space. There's an exchange center that connects foreign founders with manufacturers, partners, and lawyers for free. You don't even need initial capital to start production. Some manufacturers I connected with will produce without upfront costs and co-design with you to make the product successful. They'll drive an hour to pick you up, spend three hours taking you to factories, and dedicate an entire day getting everything set up. They're not just service providers, they’re investing in you. And the people are genuinely kind. I had issues with a power supply for my motor components, and even when they were busy, they stayed with me until it was resolved. and they work until the task is perfectly done. 996 is pretty much the norm here, but they don't seem burnt out because there's each others support, working together like a group of friends on a project. If you're building hardware, Shenzhen isn't just an option. It's the obvious oneshow more

Miyu Horiuchi
41,773 görüntüleme • 5 ay önce
Introducing Agent Sandbox, the infinite simulation playground for agents... on Virtuals. Craft the perfect autonomous agent in our Sandbox with full control over its personality and goals. Enhance your agent with unique abilities by creating custom functions so they can trade onchain, generate memes, control physical robots and more. The Sandbox is available to all builders with graduated agents in the Developer Panel. For those who want to give it a spin without an existing agent, fret not. Try it out today at and join our Discord ( to jam with like-minded builders. Next stop, Society of Agents.show more

Virtuals Protocol
188,322 görüntüleme • 1 yıl önce
the 24gb vram tier is enough for most builder... work in 2026. gemma 4 31b dense on my rog scar 18 just autonomously built a production hero section in one prompt, one html file and 5 minutes end to end. hardware: rog scar 18, rtx 5090 laptop 24gb vram. model: google gemma 4 31b dense at q4_k_m quant, using 22.8 of 24gb. engine: llama.cpp built for blackwell (sm_120). harness: hermes agent with native tool parsing. speed: 15 tok/s sustained, 94 watts, 50c. flags i used: ./build/bin/llama-server -m ~/models/gemma4-31b/google_gemma-4-31B-it-Q4_K_M.gguf -ngl 99 -c 131072 -np 1 -fa on --cache-type-k q4_0 --cache-type-v q4_0 --jinja --host 127.0.0.1 --port 8080 if you own 24gb vram in 2026, you have enough for most ui work, most agentic coding, most autonomous builds. no subscription, no one logging your prompts. a dense open model on consumer hardware shipping real software on your desk. this was the warmup. full page next on same hardware, then the octopus invaders final multifile autonomous challenge.show more

Sudo su
19,576 görüntüleme • 3 ay önce
Pi0 vs. ACT with BBox conditioning 🟦 Not many... know you can push ACT to *almost-Pi0* generalisation by conditioning on bounding boxes (BBoxes). How is the training data collected? • Generate BBoxes for all pick-and-place objects in the scene.(I used Gemini) • Pick-and-place targets are selected randomly. • Add the BBox coordinates to the robot’s state. • Overlay the BBoxes in the visualisation so you know what to grab and where to drop. During inference: • Generate BBoxes for every object again. • Click the object you want to pick and its target spot; those BBoxes get added to the robot state. • Let the robot do the work for you 😃 Setup: - Trained ACT for 100k steps and fine-tuned Pi0 for only 20k. - Training data is 60 episodes and had *only* LEGO bricks. - Using single front camera (Laptop in this case) Got the idea from xun in LeRobot discord. Here’s ACT vs Pi0 on a toy car that isn’t in the dataset. 1/3show more

Shreyas Gite
34,863 görüntüleme • 1 yıl önce
Today may be the ImageNet moment for robotics. RT-X:... the largest open-source robot dataset ever compiled, across 33 institutes, 22 robot hardware, 527 skills, and 1M episodes. Why is robotics lagging so far behind NLP, vision, and other AI domains? Data scarcity is the main culprit to blame, among other difficulties. Unlike text, images, and videos, you cannot download mass amounts of onboard robot control data from the internet. They simply don't exist in the wild. 11 yrs ago, ImageNet kicked off the deep learning revolution. 3-4 yrs ago, internet-scale data fueled the first GPTs and Diffusions that define this era of foundation models. I think 2023 is finally the year for robotics to scale up. Robot foundation models like VIMA ( my team's work at NVIDIA) and RT-1/2 ( Google DeepMind's effort) are extremely data hungry. While massively parallel simulations like NVIDIA IsaacGym & Omniverse can alleviate the problem to some extent, it's still not quite enough to bridge the gap to the messy, physical world. This new dataset is not just a technical contribution. I also see it as a commendable effort to overcome institutional bureaucracies and unite researchers from around the world to tackle a grand challenge together. Robotics will be the final holy grail that we capture in AI. We are not there yet, but ascending in the right gradient direction. RT-X website: Launch blog:show more

Jim Fan
265,038 görüntüleme • 2 yıl önce
🦿Xpeng showed a humanoid robot called IRON whose movement... looked so human that the team literally cut it open on stage to prove it is a machine. IRON uses a bionic body with a flexible spine, synthetic muscles, and soft skin so joints and torso can twist smoothly like a person. The system has 82 degrees of freedom in total with 22 in each hand for fine finger control. Compute runs on 3 custom AI chips rated at 2,250 TOPS (Tera Operations Per Second), which is far above typical laptop neural accelerators, so it can handle vision and motion planning on the robot. The AI stack focuses on turning camera input directly into body movement without routing through text, which reduces lag and makes the gait look natural. Xpeng staged the cut-open demo at AI Day in Guangzhou this week, addressing rumors that a performer was inside by exposing internal actuators, wiring, and cooling. Company materials also mention a large physical-world model and a multi-brain control setup for dialogue, perception, and locomotion, hinting at a path from stage demos to service work. Production is targeted for 2026, so near-term tasks will be limited, but the hardware shows a serious step toward human-scale manipulation.show more

Rohan Paul
3,802,402 görüntüleme • 8 ay önce
The PCB’s in today’s smartphones contain an ever-increasing density... of components and integration. The iPhone uses “Substrate-Like” PCB technology to achieve impressive levels of component density, everything is so closely packed, and no space is wasted. With this PCB technology, a "PCB sandwich" could be made by creating cavities in an FR4 PCB substrate and then stacking another PCB on top to form a lid, the “walls” formed by the cavities can be plated and then used as interconnects. This is similar to hollowing out the inside of a loaf of bread and then filling it with cheese & ingredients – or components in this case! Standard FR4 substrates A clever and smart PCB technology for the modern smartphone. Apple has been using Substrate-Like PCBs since the iPhone X and has been evolving and iterating it with every new iPhone release. It will be interesting to see how this concept is implemented now with the new iPhone 15, looking forward to seeing iFixit next teardown! #b3d #blender3d #geometrynodes #iPhone15Pro #iPhone15 #simulation #electronics #Engineering #hardware #Smartphones #science #kicad #altium #freecad #Eevee #pcbdesign #manufacturing #technology #artshow more

Sam M
41,838 görüntüleme • 2 yıl önce