Video yükleniyor...

Video Yüklenemedi

Ana Sayfaya Dön

mapping out the visual language of film using a multimodal llm: i fed frames of a short film to a vision-language model and mapped out its ratings of surrealism and presence of human figure in each moment along the timeline. the result is an interactive playback interface based on...

94,879 görüntüleme • 1 yıl önce •via X (Twitter)

11 Yorum

Kat ⊷ the Poet Engineer profil fotoğrafı
Kat ⊷ the Poet Engineer1 yıl önce

frames with high visibility of face, hands, human figure are all rated pretty high on the gestural index (2nd row). the text underneath each moving frame is llm's explanation for its rating of level of human figure in that frame.

Kat ⊷ the Poet Engineer profil fotoğrafı
Kat ⊷ the Poet Engineer1 yıl önce

the short film used: maya deren's meshes of the afternoon, an experimental classic <3 i think the surrealism score also works pretty well, rating those frames with unusual imagery higher than others.

AssemblyAI profil fotoğrafı
AssemblyAI1 yıl önce

Announcing: Our most advanced speech-to-text model goes beyond accuracy to capture the real-world complexity of human conversation and deliver reliable, source-of-truth audio data. Explore Universal-2 updates 👇

David Marjanidze profil fotoğrafı
David Marjanidze1 yıl önce

I've never seen anything like the perspectives you share. It's truly original

Kieran profil fotoğrafı
Kieran1 yıl önce

I would honestly have to try it out, but everytime i see these types of interfaces, I get excited, then i try them and our human micromovments are just too much for most with out a lot of lag added on.

Kat ⊷ the Poet Engineer profil fotoğrafı
Kat ⊷ the Poet Engineer1 yıl önce

yea the model (media pipe) is definitely not perfect and it's been out a few years. i feel there's probably much better proprietary ones out there.

アメド オシナイケ profil fotoğrafı
アメド オシナイケ1 yıl önce

Your work is enchanting. You should teach a course

glia profil fotoğrafı
glia1 yıl önce

so cool!

Pixel Girl profil fotoğrafı
Pixel Girl1 yıl önce

😍

Ian profil fotoğrafı
Ian1 yıl önce

Minority Report UI vibes. Rad.

Mike Marinos profil fotoğrafı
Mike Marinos1 yıl önce

This was the post below yours :)

Benzer Videolar

New short course Multimodal RAG: Chat with Videos, developed with Intel and taught by vasudevlal! In this course, you’ll work with LLaVA (Large Language and Vision Assistant), a Large Vision Language Model (LVLM) that can process both images and text. For example, given an image of a person doing a handstand on a skateboard at the beach, LLaVA doesn't just caption the scene, it’s able to predict possible outcomes, like the person losing balance or falling off. By understanding not just what's in a video frame, but what might happen next, your application can provide more insightful answers to questions about video. You'll build a full multimodal RAG pipeline that can chat about video content: - Use the BridgeTower model to create joint text-image embeddings in a 512-dimensional multimodal semantic space. - Learn video processing techniques to extract keyframes, generate transcripts using Whisper, and create captions. - Use the LanceDB vector database to store and retrieve high-dimensional multimodal embeddings. - Integrate the LLaVA model, combining CLIP's (Contrastive Language Image Pretraining) vision transformer with Llama, for advanced visual-textual reasoning. Your final system will ingest video data, generate embeddings for frames and text, perform similarity searches for relevant content, and use the retrieved multimodal context to inform LVLM-based response generation. The result is a system capable of answering nuanced questions about video content, effectively chatting about the video it has processed. Please sign up here!

Andrew Ng

107,825 görüntüleme • 1 yıl önce