Loading video...

Video Failed to Load

Go Home

🚀Introducing LLaVA-NeXT Interleave: Now AI can understand and reason with multiple images at once - This opens up multi-image scenarios like multi-frame videos, multi-view 3D, and multiple inter-leaved images. - An all round LMM that can understand videos, images, and 3D More⬇️

27,665 views • 2 years ago •via X (Twitter)

8 Comments

Gradio's profile picture
Gradio2 years ago

LLaVA-NeXT-Interleave🔥 - Interleave data format unifies different tasks. - New datasets on 🤗Hub: 1️⃣M4-Instruct, high-quality dataset, 1.1M samples from domains: multi-image, video, 3D & single-image 2️⃣LLaVA-Interleave Bench - Set of tasks to evaluate multi-image capabilities

Gradio's profile picture
Gradio2 years ago

LLaVA-NeXT-Interleave💪 - Attached videos show how it can explain jokes and understand content spread in multiple images and videos 🤯 - SoTA Performance, both, in multi and single images - Matches in perf with LLaVA-NeXT - Improved performance in video tasks

Gradio's profile picture
Gradio2 years ago

Gradio Multimodal Demo for LLaVA-NeXT-Interleave😍 : Models and Datasets are on 🤗 Hub:

Stark's profile picture
Stark2 years ago

how to finetune?

Omri Kaduri's profile picture
Omri Kaduri2 years ago

How can you refer to the order of the images in the prompt? Simply saying "first image" is enough? Like -"is the object in the first image shown in the second image"

Lily Zhang's profile picture
Lily Zhang2 years ago

How does it understand 3D?

Gradio's profile picture
Gradio2 years ago

Different views as multiple image input

Gradio's profile picture
Gradio2 years ago

Love this! 💡

Related Videos