Video wird geladen...

Video konnte nicht geladen werden

Zur Startseite

TidyBot: Personalized Robot Assistance with Large Language Models approach enables fast adaptation and achieves 91.2% accuracy on unseen objects in our benchmark dataset. We also demonstrate our approach on a real-world mobile manipulator called TidyBot, which successfully puts away 85.0% of objects in real-world test scenarios abs: project page:...

326,009 Aufrufe • vor 3 Jahren •via X (Twitter)

7 Kommentare

Profilbild von Jimmy Wu
Jimmy Wuvor 3 Jahren

Thanks @_akhaliq for sharing our work! I wrote a thread with more details here:

Profilbild von hardmaru
hardmaruvor 3 Jahren

This is the easy part :) I need a robot that can clean the bits of pieces of jam, bread crumbs, diapers, and occasionally pieces of poo around various corners of the room, under the tables, and hidden in the kids play area of the house.

Profilbild von hgtp:// Alkimi $ADS $QNT Cat 🐈‍⬛
hgtp:// Alkimi $ADS $QNT Cat 🐈‍⬛vor 3 Jahren

How do I order one??

Profilbild von Defend Intelligence (Anis Ayari)
Defend Intelligence (Anis Ayari)vor 3 Jahren

Really nice ! thank you for demonstrating this capability. LLM could then indeed be used as the "reasoning" block to achieves unseen world reasonning and allow tasks to get a better generalization to accomply them.

Profilbild von Dan Rockwell
Dan Rockwellvor 3 Jahren

I felt like it was missing something..

Profilbild von St. Clair Newbern IV
St. Clair Newbern IVvor 3 Jahren

Should be the standard upsell on all children. 😂

Profilbild von Astral Turf
Astral Turfvor 3 Jahren

@ericjang11 Not very impressive really.

Ähnliche Videos

DisCo: Disentangled Control for Referring Human Dance Generation in Real World paper page: Generative AI has made significant strides in computer vision, particularly in image/video synthesis conditioned on text descriptions. Despite the advancements, it remains challenging especially in the generation of human-centric content such as dance synthesis. Existing dance synthesis methods struggle with the gap between synthesized content and real-world dance scenarios. In this paper, we define a new problem setting: Referring Human Dance Generation, which focuses on real-world dance scenarios with three important properties: (i) Faithfulness: the synthesis should retain the appearance of both human subject foreground and background from the reference image, and precisely follow the target pose; (ii) Generalizability: the model should generalize to unseen human subjects, backgrounds, and poses; (iii) Compositionality: it should allow for composition of seen/unseen subjects, backgrounds, and poses from different sources. To address these challenges, we introduce a novel approach, DISCO, which includes a novel model architecture with disentangled control to improve the faithfulness and compositionality of dance synthesis, and an effective human attribute pre-training for better generalizability to unseen humans. Extensive qualitative and quantitative results demonstrate that DISCO can generate high-quality human dance images and videos with diverse appearances and flexible motions.

AK

161,453 Aufrufe • vor 3 Jahren

Differentiable Blocks World: Qualitative 3D Decomposition by Rendering Primitives paper page: Given a set of calibrated images of a scene, we present an approach that produces a simple, compact, and actionable 3D world representation by means of 3D primitives. While many approaches focus on recovering high-fidelity 3D scenes, we focus on parsing a scene into mid-level 3D representations made of a small set of textured primitives. Such representations are interpretable, easy to manipulate and suited for physics-based simulations. Moreover, unlike existing primitive decomposition methods that rely on 3D input data, our approach operates directly on images through differentiable rendering. Specifically, we model primitives as textured superquadric meshes and optimize their parameters from scratch with an image rendering loss. We highlight the importance of modeling transparency for each primitive, which is critical for optimization and also enables handling varying numbers of primitives. We show that the resulting textured primitives faithfully reconstruct the input images and accurately model the visible 3D points, while providing amodal shape completions of unseen object regions. We compare our approach to the state of the art on diverse scenes from DTU, and demonstrate its robustness on real-life captures from BlendedMVS and Nerfstudio. We also showcase how our results can be used to effortlessly edit a scene or perform physical simulations.

AK

38,571 Aufrufe • vor 3 Jahren