Video wird geladen...

Video konnte nicht geladen werden

Zur Startseite

Starting the new year without human labeling ๐ŸŽ‰!! Multimodal lidar-camera data is a gold mine of dense 3D geometry hiding in plain sight. For supervised pretraining and validation at scale at Torc-Robotics, we rely on fully automated pseudo-labeling pipelines. Exploiting geometric priors from temporally accumulated LiDAR maps and an...

18,210 Aufrufe โ€ข vor 7 Monaten โ€ขvia X (Twitter)

0 Kommentare

Keine Kommentare verfรผgbar

Kommentare vom Original-Post werden hier angezeigt

ร„hnliche Videos

๐Ÿ“ข๐Ÿ“ข ๐๐ž๐ซ๐œ๐‡๐ž๐š๐: ๐๐ž๐ซ๐œ๐ž๐ฉ๐ญ๐ฎ๐š๐ฅ ๐‡๐ž๐š๐ ๐Œ๐จ๐๐ž๐ฅ ๐Ÿ๐จ๐ซ ๐’๐ข๐ง๐ ๐ฅ๐ž-๐ˆ๐ฆ๐š๐ ๐ž ๐Ÿ‘๐ƒ ๐‡๐ž๐š๐ ๐‘๐ž๐œ๐จ๐ง๐ฌ๐ญ๐ซ๐ฎ๐œ๐ญ๐ข๐จ๐ง & ๐„๐๐ข๐ญ๐ข๐ง๐ ๐Ÿ“ข๐Ÿ“ข PercHead reconstructs realistic 3D heads from a single image and enables disentangled 3D editing via geometric controls and style inputs from images or text. At its core is a generalized 3D head decoder trained with perceptual supervision from DINOv2 and SAM 2.1. We find that our new perceptual loss formulation improves reconstruction fidelity compared to commonly-used methods such as LPIPS. Our trained reconstruction model is able to generate 3D-consistent heads from a single input image. Even with challenging side-view inputs, the model robustly infers missing regions for a coherent, high-fidelity output. In addition, our architecture seamlessly adapts to downstream tasks: by swapping the encoder, we can transform the model into a disentangled 3D editing pipeline. In this scenario, we can control geometry through - potentially hand-drawn - segmentation maps, and condition style via image or text prompt. We also provide an interactive GUI to enable the exploration of our editing pipeline. ๐ŸŒ ๐Ÿ“ฝ๏ธ Great work by Antonio Oroz and Tobias Kirschstein

Matthias Niessner

18,855 Aufrufe โ€ข vor 9 Monaten

๐Ÿš€ Announcing Echo โ€” our new frontier model for 3D world generation. Echo turns a simple text prompt or image into a fully explorable, 3D-consistent world. Instead of disconnected views, the result is a single, coherent spatial representation you can move through freely. This is part of a bigger shift in AI: from generating pixels and tokens to generating spaces. Echo predicts a geometry-grounded 3D scene at metric scale, meaning every novel view, depth map, and interaction comes from the same underlying world โ€” not independent hallucinations. Once generated, the world is interactive in real time. You control the camera, explore from any angle, and render instantly โ€” even on low-end hardware, directly in the browser. High-quality 3D world exploration is no longer gated by expensive equipment. Under the hood, Echo infers a physically grounded 3D representation and converts it into a renderable format. For our web demo, we use 3D Gaussian Splatting (3DGS) for fast, GPU-friendly rendering โ€” but the representation itself is flexible and can be easily adapted. Why this matters: consistent 3D worlds unlock real workflows โ€” digital twins, 3D design, game environments, robotics simulation, and more. From a single photo or a line of text, Echo builds worlds that are reliable, editable, and spatially faithful. Echo also enables scene editing and restyling. Change materials, remove or add objects, explore design variations โ€” all while preserving global 3D consistency. Editing no longer breaks the world. This is only the beginning. Echo is the foundation for future world models with dynamics, physical reasoning, and richer interaction โ€” environments that donโ€™t just look right, but behave right. Explore the generated worlds on our website and sign up for the closed beta. The era of spatial intelligence starts here. ๐ŸŒ #Echo #WorldModels #SpatialAI #3DFoundationModels Check it out:

SpAItial AI

176,105 Aufrufe โ€ข vor 8 Monaten

MAGS-SLAM: Monocular Multi-Agent Gaussian Splatting SLAM for Geometrically and Photometrically Consistent Reconstruction TL;DR: The first RGB-only multi-agent 3D Gaussian Splatting SLAM for collaborative photorealistic scene reconstruction. Contributions: (1) We propose the first monocular RGB-only multi-agent 3D Gaussian Splatting SLAM system. It integrates Gaussian front-ends, compact submap summaries, inter-agent verification, Sim(3) submap pose graph, and occupancy-aware fusion into a unified framework, achieving accurate tracking and photorealistic reconstruction without depth sensors. (2) We propose a Pose-Graph Bundle Adjustment (PGBA)-consistent Sim(3) loop closure mechanism for multi-agent systems, which jointly resolves intra- and inter-agent scale drift through a submap-level Sim(3) pose graph coupling geometric and photometric residuals. Robustness is ensured by a spatial-extent gate that rejects degenerate loops and an adaptive edge invalidation scheme consistent with evolving PGBA corrections. (3) We propose an occupancy-aware fusion framework for coherent multi-agent Gaussian maps. It combines occupancy-grid deduplication, decoupled coordinator, and joint pose-Gaussian photometric refinement to eliminate duplicated Gaussians, residual misalignment, and photometric seams across agents. (4) We introduce ReplicaMultiagent Plus dataset. While existing multi-agent datasets are typically limited to 2-3 agents with short trajectories, our dataset scales to 4 agents with long-horizon trajectories. In addition, we provide ground-truth geometry and semantic annotations, supporting the evaluation of monocular, RGB-D, and semantic multi-agent SLAM for collaborative dense reconstruction.

MrNeRF

19,499 Aufrufe โ€ข vor 3 Monaten

โšก๏ธ๐Ÿ“ฃ๐Ÿ‘‡Tremendously excited to share our new Cell article, where we develop TriPath, a method for analyzing 3D pathology samples using weakly supervised AI. Article: TriPath enables 3D computational pathology via 3D multiple instance learning allowing AI models to capture intricate morphological details from pathology volumes. Code: Blog post: Tested on two different imaging modalities, and patient cohorts from two institutions. Our superstar Andrew H. Song put in a monumental effort of leading the study, in a fantastic collaboration with Jonathan Liu at University of Washington . Interesting aspects: - Utilizing the whole tissue volume and leveraging 3D deep learning enable superior risk prediction performance compared to 2D deep learning baselines based on a few sampled tissue sections that emulate standard clinical practice. This indicates TriPath can harness additional information provided by 3D tissue morphology. - The performance is also superior to clinical baselines from a reader study that involved six expert pathologists. - The morphologically heterogeneous tissue volume could lead to opposing patient-level outcome predictions, dependent on which portion of the tissue volume is used. This concurs with current clinical literature warning that tissue sampling bias can lead to misdiagnosis. Some limitations: - While the 3D pathology cohort size is unprecedented, it is smaller than typical 2D pathology cohorts. Further large-scale studies will be required for validation. Nevertheless, we believe that this study will initiate a positive cycle, encouraging academic institutions and pharmaceutical companies to contribute large banks of human tissue blocks with paired clinical outcomes, thus speeding up advancements in 3D computational pathology. Concluding insights: We believe that 3D pathology is just around the corner - It has the huge potential to not only augment/improve the current clinical practice centered around 2D examination of human tissue, but also help reveal novel biomarkers for prognosis and therapeutic response.. Harvard Medical School Harvard Data Science Initiative Mass General Brigham Broad Institute

Faisal Mahmood

65,541 Aufrufe โ€ข vor 2 Jahren

Weโ€™re thrilled to share that our MERFISH+ preprint is now live on bioRxiv!๐Ÿ‘‰ In this work, the Bintu and Zhu labs (UCSD) developed MERFISH+, a next-generation spatial genomics platform that combines genome-wide RNA and epigenetic imaging over a large field of view. By introducing acrydite-modified probes covalently anchored to hydrogels, MERFISH+ achieves remarkable imaging stability and enables >1,800-gene, multi-modal, and multi-month experiments. With this platform, they, together with the Chi lab at UCSD, profiled a whole developing human heart at 12 post-conception week with merely two slides, resulting in a total of 53 slides, 3.1 million single cells and more than 30 cell types. Building upon our previous 3D reconstruction and modeling framework, Spateo ( we reconstruct the 3D human heart that nicely captures the anatomical structure of the heart, including the intricate vasculature network. Sophisticated analyses provide a holistic view of an entire organ and enable systematic characterization of 3D cellular neighborhoods and transcriptional gradients of substructures such as the descending arteries. Furthermore, using a generative integration framework for spatial multimodal data (Spateo-VI), we harmonized these MERFISH+ transcriptomic and chromatin data to reconstruct a 3D spatially-resolved multi-omics atlas of the developing human heart, shared at and MERFISH+ thus sets a new standard for large-format, multi-omic spatial profiling, enabling holistic, 3D characterization of organs at subcellular resolution. Huge congratulations to first authors Colin Kern, qingquan Zhang, @YifanLu2024 , and Jacqueline Eschbach, and to all collaborators from the Bintu, Zhu, Chi, and Qiu labs for this amazing team effort. Thanks for your diligence, creativity, and hard work on this project. Weโ€™re grateful for support from Arc Institute and our generous donors. Our lab is expandingโ€”if youโ€™re excited about building the next generation of single-cell and spatial genomics techniques and predictive single cell and spatial foundation models, weโ€™re hiring! If you are interested, please reach out to me via direct message or email at [email protected]. We are excited for any potential collaborations along this line of research in Stanford, UCSF and Berkeley and other labs as well.

evo-devo

42,268 Aufrufe โ€ข vor 9 Monaten

๐Ÿ“ข๐Ÿ“ข ๐€๐ฏ๐š๐ญ๐Ÿ‘๐ซ ๐Ÿ“ข๐Ÿ“ข Avat3r creates high-quality 3D head avatars from just a few input images in a single forward pass with a new dynamic 3DGS reconstruction model. Video: Project: Our core idea is to make Gaussian Reconstruction Models animatable. We find that a simple cross-attention to an expression code sequence is already sufficient to model complex facial expressions. We then incorporate position maps from DUSt3R and feature maps from Sapiens to facilitate the prediction task. While DUSt3R's position maps act as a pixel-aligned initialization for the Gaussians' positions, the Sapiens feature maps help the cross-view transformer to match corresponding image tokens in the 4 input images. One major challenge in creating a 3D head avatar from smartphone images comes from inconsistent facial expressions when the subject could not remain perfectly static during the capture. We eliminate this static requirement by simply showing our model input images with different facial expressions during training. This technique makes our model robust to inconsistent input images later on. Finally, we show that despite the model has been trained with 4 input images, one can even create a 3D head avatar when only a single image is available. To achieve this, we employ a pre-trained 3D GAN to lift the single image to 3D and then render the 4 input images for our model. This allows us to create 3D head avatars from single images and even highly out-of-distribution examples like AI generated faces, paintings or statues. Great work by Tobias Kirschstein from his internship at Meta with Javier Romero, Artem Sevastopolsky, and Shunsuke Saito

Matthias Niessner

74,763 Aufrufe โ€ข vor 1 Jahr

We believe weโ€™re the first robotics company to demonstrate a robot peeling an apple with dual dexterous human-like hands. This breakthrough closes a key gap in robotics, achieving bimanual, contact-rich manipulation and moving far beyond the limits of simple grippers. ๐Ÿงตโ†“ Todayโ€™s AI models (VLMs) are excellent at perception but struggle with action. Controlling high-degree-of-freedom hands for tasks like this is incredibly complex, and precise finger-level teleoperation is nearly impossible for humans. Our first step was a shared-autonomy system: rather than controlling every finger, the operator triggers pre-learned skills like a โ€œrotate apple or tennis ballโ€ primitive via a keyboard press or pedal. This makes scalable data collection and RL training possible. How does the AI manage this? We created "MoDE-VLA" (Mixture of Dexterous Experts). It fuses vision, language, force, and touch data by using a team of specialist "experts," making control in high-dimensional spaces stable and effective. The combination of these two innovations allows for seamless, contact-rich manipulation. The human provides high-level guidance, and the robot executes the complex in-hand coordination required. This work paves the way for robots that can safely handle delicate tasks in human environments. Want the full technical details? ๐Ÿ“„ Read the full research paper: Visit us at NVIDIA GTC Booth #1838, Hall 3 to learn more! #Robotics #AI #DexterousManipulation #VLA #NVIDIAGTC Nancy Villicaรฑa NVIDIA GTC

Sharpa

20,429 Aufrufe โ€ข vor 5 Monaten

After 8 months of building in stealth and testing our infrastructure on 10000+ hours of real-world data and hundreds of unique environments, we're bringing FPV Labs into the open today. FPV Labs started with the following bet - if human data proves to be the underlying factor that determines scaling laws in general-purpose robotics, it will trigger the largest economic transformation in human history, and the underlying infrastructure that captures that data will determine how fast we get there. We will achieve this by building the full-stack infrastructure for capturing, processing, transferring, and evaluating human experience into spatial, temporal, and semantic knowledge for machines. Despite all the research novelty behind ChatGPT, its success can be attributed to one foundational fact - the scaling law of transformers. We believe the same dynamics have made their way into robotics. Recent studies showed task completion rates jumping from 30% to 70% when human demonstration data scaled from 1,000 to 20,000 hours, a log-linear trend that mirrors exactly what we saw in language and vision. Seeing these emergent signs of scaling law curves in robotics, we believe we are entering the era of general-purpose robotics policies, which makes the next few years the most exciting time in the history of this field. But the library of physical interactions required to train general-purpose robot policies does not exist yet. Over the last 8 months, we've seen dozens of companies emerge in this space. We were really happy to see new companies pushing this space forward, but we also saw the same pattern repeat: every egocentric data company was making some tradeoffs between quality, scale, and diversity. We have built FPV labs on the core principle that high-quality data is orders of magnitude more valuable than sheer volume. Case in point, self-driving cars collect thousands of hours of data per day, but only a small fraction of that data is actually useful for training better models. Several studies, like RT-2, have shown that as little as 1% of data improves as much as 25% on task success. The quality and diversity of data matter a lot more than scale, so there is clearly a power law curve in the downstream impact of data. We've spent months obsessing over data quality by building our stack, discarding it, rebuilding it, and iterating until we found a formula that doesn't compromise downstream quality at scale. We believe the downstream impact here is far more profound than most people realize. Workers globally are paid around $60 trillion per year in aggregate, and a lion's share of that compensation goes to physical labor - tasks that require navigating real spaces, manipulating real objects, and negotiating the infinite variability of the physical world. Human-to-robot transfer will be one of the most important infrastructures that will shape our society in the near future, and if it works, the economic impact will dwarf every technology transition that came before it in an exponential manner and lead to the creation of goods and services we canโ€™t imagine today. Our mission is to lay the groundwork for us to transition into this future - the future of abundance. We are deeply grateful to our earliest believers, Paras Chopra and Lossfunk, who played a critical role in shaping our thinking.

Abhishek Anand

81,839 Aufrufe โ€ข vor 4 Monaten

Want to create an avatar from a single image? FlexAvatar is a transformer model that creates full 360ยฐ, high-quality, and expressive 3D head avatar from just a single portrait image in minutes. Real-time Demo: FlexAvatar's lightweight architecture allows both animation and rendering in real-time, enabling interactive user experiences. To create a new 3D head avatar, only one image is required, e.g., from a webcam. The final avatar is ready after 2 minutes. Architecture: Under the hood, FlexAvatar adopts a transformer-based encoder-decoder design. The encoder maps the input image onto a latent avatar space, while the decoder produces 3D Gaussian attribute maps by incorporating the animation signal via cross-attention. The model learns all facial animations directly from the data without relying on pre-built 3D face models. This equips the avatars with realistic facial expressions. The internal avatar latent space can be conveniently used to integrate additional observations of a person via fitting. This enables use-cases where more than one image of a person is available, e.g., from a phone scan of the person. We train jointly on 2D monocular videos and multi-view data. However, in monocular videos, the animation signal leaks the target viewpoint, causing the model to produce incomplete 3D heads. We call this phenomenon entanglement of driving signal and target viewpoint. To prevent entanglement, we introduce bias sinks. These are learnable tokens that indicate whether a training sample stems from a monocular or a multi-view dataset. During training, the model learns to produce incomplete 3D heads only when the monocular token is present. During inference, FlexAvatar then always uses the multi-view token for which the model has learned to produce complete 3D heads. This simple design allows to combine the generalizability from monocular data with the quality of multi-view data. FlexAvatar summary: - Input: Single-image, phone scan, or monocular video - Output: Full 360ยฐ head avatar - Expressive animations - Real-time rendering and animation - Generalization to any portrait - Create a new avatar in 2 minutes - Use bias sinks to combine 2D and 3D data ๐Ÿ  ๐ŸŒ ๐ŸŽฅ Great work by Tobias Kirschstein and Simon Giebenhain!

Matthias Niessner

96,186 Aufrufe โ€ข vor 7 Monaten

๐Ÿšจ SIGGRAPH Asia 2025 Paper Alert ๐Ÿšจ โžก๏ธPaper Title: WorldExplorer: Towards Generating Fully Navigable 3D Scenes ๐ŸŒŸFew pointers from the paper ๐ŸŽฏGenerating 3D worlds from text is a highly anticipated goal in computer vision. Existing works are limited by the degree of exploration they allow inside of a scene, i.e., produce stretched-out and noisy artifacts when moving beyond central or panoramic perspectives. ๐ŸŽฏ To this end, authors of this paper proposed โ€œWorldExplorerโ€, a novel method based on autoregressive video trajectory generation, which builds fully navigable 3D scenes with consistent visual quality across a wide range of viewpoints. ๐ŸŽฏThey initialize their scenes by creating multi-view consistent images corresponding to a 360 degree panorama. ๐ŸŽฏThen, they expanded it by leveraging video diffusion models in an iterative scene generation pipeline. ๐ŸŽฏConcretely, they generated multiple videos along short, pre-defined trajectories, that explore the scene in depth, including motion around objects. ๐ŸŽฏTheir novel scene memory conditions each video on the most relevant prior views, while a collision-detection mechanism prevents degenerate results, like moving into objects. ๐ŸŽฏFinally,they fuse all generated views into a unified 3D representation via 3D Gaussian Splatting optimization. ๐ŸŽฏCompared to prior approaches, WorldExplorer produces high-quality scenes that remain stable under large camera motion, enabling for the first time realistic and unrestricted exploration. ๐ŸŽฏThey believe this marks a significant step toward generating immersive and truly explorable virtual 3D environments. ๐ŸขOrganization: TU Mรผnchen ๐Ÿง™Paper Authors: Manuel-Andreas Schneider, Lukas Hรถllein , Matthias Niessner ๐Ÿ“ Read the Full Paper here: ๐Ÿ—‚๏ธ Project Page: ๐Ÿง‘โ€๐Ÿ’ป Code: ๐ŸŽฅ Be sure to watch the attached Technical Summary Video - Sound on ๐Ÿ”Š๐Ÿ”Š Find this Valuable ๐Ÿ’Ž ? โ™ป๏ธQT and teach your network something new Follow me ๐Ÿ‘ฃ, naveen manwani , for the latest updates on Tech and AI-related news, insightful research papers, and exciting announcements. #SIGGRAPHAsia2025

naveen manwani

10,578 Aufrufe โ€ข vor 10 Monaten

Exciting updates on Project GR00T! We discover a systematic way to scale up robot data, tackling the most painful pain point in robotics. The idea is simple: human collects demonstration on a real robot, and we multiply that data 1000x or more in simulation. Letโ€™s break it down: 1. We use Apple Vision Pro (yes!!) to give the human operator first person control of the humanoid. Vision Pro parses human hand pose and retargets the motion to the robot hand, all in real time. From the humanโ€™s point of view, they are immersed in another body like the Avatar. Teleoperation is slow and time-consuming, but we can afford to collect a small amount of data. 2. We use RoboCasa, a generative simulation framework, to multiply the demonstration data by varying the visual appearance and layout of the environment. In Jensenโ€™s keynote video below, the humanoid is now placing the cup in hundreds of kitchens with a huge diversity of textures, furniture, and object placement. We only have 1 physical kitchen at the GEAR Lab in NVIDIA HQ, but we can conjure up infinite ones in simulation. 3. Finally, we apply MimicGen, a technique to multiply the above data even more by varying the *motion* of the robot. MimicGen generates vast number of new action trajectories based on the original human data, and filters out failed ones (e.g. those that drop the cup) to form a much larger dataset. To sum up, given 1 human trajectory with Vision Pro -> RoboCasa produces N (varying visuals) -> MimicGen further augments to NxM (varying motions). This is the way to trade compute for expensive human data by GPU-accelerated simulation. A while ago, I mentioned that teleoperation is fundamentally not scalable, because we are always limited by 24 hrs/robot/day in the world of atoms. Our new GR00T synthetic data pipeline breaks this barrier in the world of bits. Scaling has been so much fun for LLMs, and it's finally our turn to have fun in robotics! We are building tools to enable everyone in the ecosystem to scale up with us. Links in thread:

Jim Fan

364,514 Aufrufe โ€ข vor 2 Jahren

China unveils humanoid robot worker with brain that runs 275 trillion ops/sec | Jijo Malayil, Interesting Engineering In tests, SUYUAN used vision and joint control to sort and move crates of various sizes, greatly improving warehouse productivity. Chinese manufacturing firm Shanghai Electric has unveiled its first self-developed industrial humanoid robot, โ€œSUYUAN,โ€ marking a major milestone in its robotics journey. Debuting at the World Artificial Intelligence Conference (WAIC 2025) on July 26 in Shanghai, SUYUAN boasts 38 degrees of freedom and 275 TOPS of on-device computing power, enabling precise operations and fluid movements. According to the firm, designed for diverse industrial use, the robot showcases Shanghai Electricโ€™s end-to-end capabilitiesโ€”from core tech to integrated solutionsโ€”and reinforces its commitment to next-gen industrial automation through a full industry chain strategy. At WAIC 2025, Shanghai Electric also unveiled a new joint venture with Johnson Electric for next-gen humanoid robotics and showcased its โ€œLINGKEโ€ dual-arm robot. Recently, Hangzhou-based Unitree Robotics launched the R1 humanoid with 26 joints for $5,900, showcasing athletic feats like cartwheels, running, and quick recovery. Smart factory assistant Shanghai Electric claims SUYUAN, equipped with 38 degrees of freedom (DoF) and a powerful 275 TOPS on-device computing processor, delivers fluid, human-like movements and high-precision operations across various industrial scenarios. Its advanced articulation and real-time processing capabilities make it highly adaptable, enabling smooth execution of complex tasks in dynamic work environments. SUYUAN, who weighs 110 pounds (50 kilograms) and is 5 feet 6 inches (167 cm) tall, was designed to have human-like proportions. Its 38-DoF articulation offers dexterity, allowing for both wide-range motion and sensitive manipulation. With a single arm, the robot can lift objects up to 4.4 pounds (2 kilograms) in weight and carry a total payload of up to 22 pounds (10 kilograms). With a walking pace of 3.1 miles per hour (5 km/h), SUYUAN is ideal for environments including assembly lines, warehousing, and logistics, according to a statement. To navigate complex industrial settings, SUYUAN combines LiDAR and binocular vision for self-guided mobility. Its 275-TOPS AI processor enables rapid data analysis and integration with large language models, allowing it to understand tasks in natural language and handle objects adaptively, reports Fox 44 News. In pilot demonstrations, the robot successfully identified, picked, and relocated crates of varying sizes using advanced computer vision and coordinated joint controlโ€”delivering measurable gains in warehouse efficiency. The company claims that SUYUANโ€™s launch represents a major turning point in Shanghai Electricโ€™s foray into humanoid robotics and strengthens its vertically integrated approach to industrial automation solutions. Intelligent task handling Shanghai Electric also demonstrated its most recent developments in intelligent manufacturing at WAIC 2025, introducing a new joint venture with Johnson Electric centered on next-generation humanoid robotics and showcasing the โ€œLINGKEโ€ dual-arm robot. With its high-precision operations, adaptive teamwork, and closed-loop data capabilities, the LINGKE robot demonstrated live talents in handling complicated production jobs. LINGKE is made to do more than just replace human labor; it uses compliant force control and bimanual coordination to relieve workers of high-intensity, repetitive jobs. According to the company, the robot enhances operational efficiency by up to five times. Its core strength lies in a Data-Model-Deployment closed-loop system that starts with operational data, followed by data cleansing, model training, live deployment, and feedback-driven optimizationโ€”enabling autonomous learning and workflow improvement. Also at the event, Shanghai Electric and Johnson Electric introduced advanced hardware modules for humanoid robots, including rotary joints, linear joints, and dexterous finger joints. These components are designed to support smooth, precise, and quiet motion performance across robotics systems, reports Stock Titan. The joint venture announced two strategic agreements: a first-unit supply deal with the National and Local Co-Built Humanoid Robotics Innovation Center (Qinglong Project) and a cooperation memorandum with Fourier Robotics. Read more:

Owen Gregorian

51,638 Aufrufe โ€ข vor 1 Jahr

Iโ€™m thrilled to announce that we just released GraspGen, a multi-year project we have been cooking at NVIDIA Robotics ๐Ÿš€ GraspGen: A Diffusion-Based Framework for 6-DOF Grasping Grasping is a foundational challenge in robotics ๐Ÿค– โ€” whether for industrial picking or general-purpose humanoids. VLA + real data collection is all the rage now but is expensive and scales poorly for this task. For every new gripper and/or scene, youโ€™ll have to recollect the dataset in this paradigm for the best perf. ๐Ÿ’กKey Idea: Since grasping is such a well-defined task in simulation - why canโ€™t we just scale synthetic data generation and train a generative model for grasping? By embracing modularity and standardized grasp formats, we can make this a turnkey technology that works zero-shot for multiple settings. GraspGen is a modular framework for diffusion-based 6-DOF grasp generation that scales across embodiment types, observability conditions, clutter, task complexity. Key Features: โœ… Multi-embodiment support: suction, parallel-jaw, and multi-fingered grippers โœ… Generalization to partial + complete 3D point clouds โœ… Generalization to single-objects + cluttered scenes โœ… Modular design uses other robotics modules and foundation models (SAM2, cuRobo, FoundationStereo, FoundationPose). This allows GraspGen to focus on only one thing - grasp generation โœ… Training recipe: grasp discriminator is trained with On-Generator data from the diffusion model - so that it learns to correct the mistakes (if any) of the diffusion generator โœ… Real-time performance (~20 Hz) before any GPU acceleration; low memory footprint ๐Ÿ“Š Results: โ€ข SOTA on the FetchBench [Han et al. CoRL 2024] benchmark โ€ข Zero-shot sim-to-real transfer on unknown objects and cluttered scenes โ€ข Dataset of 53M simulated grasps across 8K objects from Objaverse ๐Ÿ“„ arXiv: ๐ŸŒ Website: ๐Ÿ’ป Code: A huge thank you to everyone involved in this journey โ€” excited to see what the community builds on top of it! Joint work with Clemens Eppner , Balakumar Sundaralingam , Yu-Wei, Jun Yamada Wentao Yuan and other collaborators #robotics #diffusionmodels #physicalAI #simtoreal

Adithya Murali

24,106 Aufrufe โ€ข vor 1 Jahr

Kled Version 3 is coming. Over $20M+ in rewards will be paid directly to users from leading AI labs across robotics, legal services, image and video generation, world modeling, and more. In the last seven days, weโ€™ve received inbound data requests from several decacorn AI labs and enterprises for datasets our human data marketplace is uniquely positioned to provide. Since receiving the specs for these requests, we now have a much better picture and understanding of how to reshape the systems that collect this data, so hereโ€™s whatโ€™s coming: 1. A fully redesigned home experience: The home feed is being rebuilt to surface the highest-value, most relevant tasks for each user, similar to how Uber Eats surfaces top restaurants. The goal is to turn every user into their most effective version as a data contributor. 2. Automated quality enforcement at scale: New ML systems are being built to evaluate task-specific requirements in real time. For example, if a task requires โ€œtwo hands visible on camera at all times,โ€ any video that fails that spec will be automatically rejected. This logic will apply across thousands of tasks and specifications using a general ML. 3. Kled Shop: Some tasks require better capture hardware. Weโ€™re introducing Kled Shop, where users can redeem points or tokens for equipment like Meta glasses, drones, and other tools. Points and tokens can be converted directly from payouts. 4. Partner-run data labeling and evaluation work: Some of our partners operate high-paying data labeling and model evaluation programs. Weโ€™re integrating their workflows directly into Kled so qualified users can access these roles in one place. These jobs are owned and managed by our partners. Kledโ€™s role is to route the right people to the right work. Some opportunities pay $50โ€“$1,000 per hour depending on expertise. 5. Global payouts and localization: Weโ€™re partnering with a major payment processor to enable cashouts in usersโ€™ native currencies. This unlocks broader global participation. Multi-language support is also coming to accelerate user growth. This full suite of tools will be rolling out soon, directly to Kled users. Top earners are currently making ~$7,000 per month. With this update, we should see the first ~$10,000 per month earner.

Avi Patel

124,728 Aufrufe โ€ข vor 6 Monaten

Everything you love about generative models โ€” now powered by real physics! Announcing the Genesis project โ€” after a 24-month large-scale research collaboration involving over 20 research labs โ€” a generative physics engine able to generate 4D dynamical worlds powered by a physics simulation platform designed for general-purpose robotics and physical AI applications. Genesis's physics engine is developed in pure Python, while being 10-80x faster than existing GPU-accelerated stacks like Isaac Gym and MJX. It delivers a simulation speed ~430,000 faster than in real-time, and takes only 26 seconds to train a robotic locomotion policy transferrable to the real world on a single RTX4090 (see tutorial: The Genesis physics engine and simulation platform is fully open source at We'll gradually roll out access to our generative framework in the near future. Genesis implements a unified simulation framework all from scratch, integrating a wide spectrum of state-of-the-art physics solvers, allowing simulation of the whole physical world in a virtual realm with the highest realism. We aim to build a universal data engine that leverages an upper-level generative framework to autonomously create physical worlds, together with various modes of data, including environments, camera motions, robotic task proposals, reward functions, robot policies, character motions, fully interactive 3D scenes, open-world articulated assets, and more, aiming towards fully automated data generation for robotics, physical AI and other applications. Open Source Code: Project webpage: Documentation: 1/n

Zhou Xian

3,818,940 Aufrufe โ€ข vor 1 Jahr