Video wird geladen...

Video konnte nicht geladen werden

Zur Startseite

Capturing every combination of human-object interaction is not feasible. We need compositionally. Most interactions are local. A hand holds a cup, a chair supports the pelvis, while much of the body remains free. COSMI, by Daniel Escandar Ilya Petrov builds on this observation to compose existing single-object datasets into...

20,393 Aufrufe • vor 2 Tagen •via X (Twitter)

8 Kommentare

Profilbild von Gerard Pons-Moll
Gerard Pons-Mollvor 2 Tagen

We need *compositionality*

Profilbild von Diakonos
Diakonosvor 2 Tagen

@ptrvilya Any hint on the release date?🥲

Profilbild von Sani Ai Tech
Sani Ai Techvor 1 Tag

@ptrvilya Compositional interaction modeling could unlock much stronger generalization to unseen object combinations

Profilbild von Eplurubusnullus
Eplurubusnullusvor 1 Tag

@ptrvilya This is huge, and exactly what I've been looking for.

Profilbild von ζ Pedram ζ
ζ Pedram ζvor 2 Tagen

@ptrvilya Need more then that but this is a good start Morphological intelligence is foundational to Kung Fu and Cybernetics imho.. The next level is something I study, how to keep the coffee in your mug.. while being pushed around.

Profilbild von Diakonos
Diakonosvor 2 Tagen

@ptrvilya God sent lolll 🔥🔥🔥great work can wait to use it

Profilbild von Jo
Jovor 1 Tag

@ptrvilya Perhaps we should borrow a concept from video compression: model only the interactions that change, while preserving and reusing everything that remains unchanged. I hear you, it’s like, why reconstruct the entire scene when only a small part of it is interacting?

Profilbild von Vantix AI Agency
Vantix AI Agencyvor 1 Tag

@ptrvilya 222k sequences and up to 5 objects composed so elegantly building on local interactions to scale compositionally is such a smart breakthrough

Ähnliche Videos

Blended-NeRF: Zero-Shot Object Generation and Blending in Existing Neural Radiance Fields paper page: Editing a local region or a specific object in a 3D scene represented by a NeRF is challenging, mainly due to the implicit nature of the scene representation. Consistently blending a new realistic object into the scene adds an additional level of difficulty. We present Blended-NeRF, a robust and flexible framework for editing a specific region of interest in an existing NeRF scene, based on text prompts or image patches, along with a 3D ROI box. Our method leverages a pretrained language-image model to steer the synthesis towards a user-provided text prompt or image patch, along with a 3D MLP model initialized on an existing NeRF scene to generate the object and blend it into a specified region in the original scene. We allow local editing by localizing a 3D ROI box in the input scene, and seamlessly blend the content synthesized inside the ROI with the existing scene using a novel volumetric blending technique. To obtain natural looking and view-consistent results, we leverage existing and new geometric priors and 3D augmentations for improving the visual fidelity of the final result. We test our framework both qualitatively and quantitatively on a variety of real 3D scenes and text prompts, demonstrating realistic multi-view consistent results with much flexibility and diversity compared to the baselines. Finally, we show the applicability of our framework for several 3D editing applications, including adding new objects to a scene, removing/replacing/altering existing objects, and texture conversion.

AK

62,768 Aufrufe • vor 3 Jahren

MaterialFusion Enhancing Inverse Rendering with Material Diffusion Priors discuss: Recent works in inverse rendering have shown promise in using multi-view images of an object to recover shape, albedo, and materials. However, the recovered components often fail to render accurately under new lighting conditions due to the intrinsic challenge of disentangling albedo and material properties from input images. To address this challenge, we introduce MaterialFusion, an enhanced conventional 3D inverse rendering pipeline that incorporates a 2D prior on texture and material properties. We present StableMaterial, a 2D diffusion model prior that refines multi-lit data to estimate the most likely albedo and material from given input appearances. This model is trained on albedo, material, and relit image data derived from a curated dataset of approximately ~12K artist-designed synthetic Blender objects called BlenderVault. we incorporate this diffusion prior with an inverse rendering framework where we use score distillation sampling (SDS) to guide the optimization of the albedo and materials, improving relighting performance in comparison with previous work. We validate MaterialFusion's relighting performance on 4 datasets of synthetic and real objects under diverse illumination conditions, showing our diffusion-aided approach significantly improves the appearance of reconstructed objects under novel lighting conditions. We intend to publicly release our BlenderVault dataset to support further research in this field.

AK

22,959 Aufrufe • vor 2 Jahren

Exciting updates on Project GR00T! We discover a systematic way to scale up robot data, tackling the most painful pain point in robotics. The idea is simple: human collects demonstration on a real robot, and we multiply that data 1000x or more in simulation. Let’s break it down: 1. We use Apple Vision Pro (yes!!) to give the human operator first person control of the humanoid. Vision Pro parses human hand pose and retargets the motion to the robot hand, all in real time. From the human’s point of view, they are immersed in another body like the Avatar. Teleoperation is slow and time-consuming, but we can afford to collect a small amount of data. 2. We use RoboCasa, a generative simulation framework, to multiply the demonstration data by varying the visual appearance and layout of the environment. In Jensen’s keynote video below, the humanoid is now placing the cup in hundreds of kitchens with a huge diversity of textures, furniture, and object placement. We only have 1 physical kitchen at the GEAR Lab in NVIDIA HQ, but we can conjure up infinite ones in simulation. 3. Finally, we apply MimicGen, a technique to multiply the above data even more by varying the *motion* of the robot. MimicGen generates vast number of new action trajectories based on the original human data, and filters out failed ones (e.g. those that drop the cup) to form a much larger dataset. To sum up, given 1 human trajectory with Vision Pro -> RoboCasa produces N (varying visuals) -> MimicGen further augments to NxM (varying motions). This is the way to trade compute for expensive human data by GPU-accelerated simulation. A while ago, I mentioned that teleoperation is fundamentally not scalable, because we are always limited by 24 hrs/robot/day in the world of atoms. Our new GR00T synthetic data pipeline breaks this barrier in the world of bits. Scaling has been so much fun for LLMs, and it's finally our turn to have fun in robotics! We are building tools to enable everyone in the ecosystem to scale up with us. Links in thread:

Jim Fan

364,721 Aufrufe • vor 2 Jahren

[] = edited info for privacy "I am a [Major US Carrier] A320/321 Captain, the following sighting occurred during one of my flights recently. Since I have shared my story, several other [Major US Carrier] pilots have reached out to me and shared their similar experiences, including sharing their video recordings of these objects from the Flight Levels. All seen at the base of the big dipper. [Last week of July], 2023, I departed Santo Domingo DR at 2305 destined for New York JFK. My route of flight was L453 in NY Oceanic airspace, non radar hundreds of miles offshore. At approximately 1 hour into the flight as we were approaching the southern boundary of the NY oceanic airspace, and at 32,000 feet, I called out a visual on traffic that was excessively bright and looked like about 80 miles range... and then disappeared visually. I never saw the traffic on TCAS. Then a few minutes later I saw two objects round in shape, one lighted and one not flying in a formation just above the horizon, at a range I guessed of 120-200 NM. The object/s would illuminate to be as bright as a star for several seconds, then go dark for a few minutes, only to illuminate again. The brightness would vary from bright to very bright to dark. The color of the lighted object was white light. There was a second object you can clearly see in the photos that would follow the illuminated object, but it would not illuminate itself. Except for the 3:10-3:30 second point of my video I think you can see the second object illuminate. This went on for the remaining 2-2.5 hours of my flight to NY on L453 in Oceanic airspace. After reviewing the photos, I think the objects might have been a bit further away, but distance is very difficult to gauge at night. I have the brand new Samsung S23 phone which has the best camera on the market for a cell phone and I started recording this object in video. I have a great 7 min video of it appearing and disappearing while I was talking to other airliners on 123.45 vhf about it. You can hear that conversation in the video! Another airliner approximately 400 NM ahead of us at 36,000 feet stated they saw the same thing. I also took about 30 photos of these objects in "night mode" on the phone and they came out really good... in one of them you can actually see the lighted object and the unlit object very clearly as round metallic objects. All of the photos were taken with some sort of long exposure setting to be able to get as much light as possible to the sensor... You will be able to see in the long exposure photos the stars are pins of light but the UFO's are streaks of light because they are moving! It is actually amazing! All of this happened over about a 2.5 hour flight and continued for the entirety of our flight. The light seemed to be on or just above the horizon until we got closer to our destination of NY. Just prior to beginning our descent the objects appeared much higher in the sky 80-90 degrees above the horizon and much further away, actually out of the atmosphere."

Ryan Graves

694,787 Aufrufe • vor 3 Jahren