Video wird geladen...
Video konnte nicht geladen werden
Capturing every combination of human-object interaction is not feasible. We need compositionally. Most interactions are local. A hand holds a cup, a chair supports the pelvis, while much of the body remains free. COSMI, by Daniel Escandar Ilya Petrov builds on this observation to compose existing single-object datasets into... show more
20,393 Aufrufe • vor 2 Tagen •via X (Twitter)
8 Kommentare

We need *compositionality*

@ptrvilya Any hint on the release date?🥲

@ptrvilya Compositional interaction modeling could unlock much stronger generalization to unseen object combinations

@ptrvilya This is huge, and exactly what I've been looking for.

@ptrvilya Need more then that but this is a good start Morphological intelligence is foundational to Kung Fu and Cybernetics imho.. The next level is something I study, how to keep the coffee in your mug.. while being pushed around.

@ptrvilya God sent lolll 🔥🔥🔥great work can wait to use it

@ptrvilya Perhaps we should borrow a concept from video compression: model only the interactions that change, while preserving and reusing everything that remains unchanged. I hear you, it’s like, why reconstruct the entire scene when only a small part of it is interacting?

@ptrvilya 222k sequences and up to 5 objects composed so elegantly building on local interactions to scale compositionally is such a smart breakthrough

