Загрузка видео...

Не удалось загрузить видео

На главную

Our preprint on the Perception Test is now on arXiv: We will be benchmarking the cream of the crop multimodal video models in the Perception Test Challenge, happening in ICCV 2023. (1 / 2)

46,218 просмотров • 3 лет назад •via X (Twitter)

Комментарии: 3

Фото профиля joao carreira
joao carreira3 лет назад

PT from @Deepmind can diagnose video models, assessing their abilities around memory, abstraction, physics, semantics. It has 6 types of annotations (object & point tracks, action & sound segments, multiple-choice & grounded VQA) and 6 tasks. It is hard!

Фото профиля Omar Sanseviero
Omar Sanseviero3 лет назад

Very cool! Congrats! Could be cool to have this on @huggingface datasets to make it very easy to use and explore the benchmark 🚀 also looks 🔥Let us know if we can collaborate/help in some way!

Фото профиля co-varun
co-varun3 лет назад

This video incites both an overwhelming sensation of elation and a distinct sense of FOMO, given the alignment of their solution with my own line of thought. It's an emotional paradox, indeed. Hat's off and team. This represents the tangible trajectory towards true multimodality. Undoubtedly an epoch-making development! 🔥🔥🔥🔥

Похожие видео