Loading video...

Video Failed to Load

Go Home

🚀Introducing LLaVA Lightning: Train a lite, multimodal GPT-4 with just $40 in 3 hours! With our newly introduced datasets and the efficient design of LLaVA, you can now turbocharge your language model with image reasoning capabilities, in an incredibly affordable way.🧵

302,319 views • 3 years ago •via X (Twitter)

10 Comments

Haotian Liu's profile picture
Haotian Liu3 years ago

(2/5) Excited to release a 558K concept-balanced subset of LAION/CC/SBU & an 80K high-quality subset of LLaVA-Instruct-158K. The concept-balanced subset ensures a broad concept coverage, and the high-quality visual instruct tuning data enables models' visual reasoning capability.

Haotian Liu's profile picture
Haotian Liu3 years ago

(3/5) Upgrade your Vicuna-7B to LLaVA-Lightning in just 3 hrs: 2 hrs pretraining + 1 hr visual instruct tuning. Train on 8x A100s using cloud spot instances for just $40. Let's make this research more accessible to researchers, academia, and millions of AI enthusiasts today!

Haotian Liu's profile picture
Haotian Liu3 years ago

(4/5) We're also upgrading LLaVA to support Vicuna v0 & v1 weights, with more checkpoints arriving this week! Plus, we're working to support more hardware – stay tuned!

Haotian Liu's profile picture
Haotian Liu3 years ago

(5/5) 🤗 Demo: 🌐 Project page: 📄 Paper: Embark on your LLaVA-Lightning journey today and stay tuned for more models and support for more hardwares in the following weeks!

iamrobotbear (bk)'s profile picture
iamrobotbear (bk)3 years ago

Any way to easily swap LLaMA out for OpenAI or Dolly 2.0?

Haotian Liu's profile picture
Haotian Liu3 years ago

Yes, it is definitely possible. And even easier with the introduction of LLaVA lightning. MPT-7B just joins the LLaVA family today!

Zongheng Yang's profile picture
Zongheng Yang3 years ago

Congrats on the work @imhaotian. Glad to see SkyPilot was of help!

Chris's profile picture
Chris3 years ago

Recently program of open source projects is super fast, definitely surpass my expectation.

web3工作坊's profile picture
web3工作坊3 years ago

@_akhaliq Thank you for sharing. @savetonotion #tweet #AI

Jake Harrison's profile picture
Jake Harrison3 years ago

Good job!

Related Videos

New short course Multimodal RAG: Chat with Videos, developed with Intel and taught by vasudevlal! In this course, you’ll work with LLaVA (Large Language and Vision Assistant), a Large Vision Language Model (LVLM) that can process both images and text. For example, given an image of a person doing a handstand on a skateboard at the beach, LLaVA doesn't just caption the scene, it’s able to predict possible outcomes, like the person losing balance or falling off. By understanding not just what's in a video frame, but what might happen next, your application can provide more insightful answers to questions about video. You'll build a full multimodal RAG pipeline that can chat about video content: - Use the BridgeTower model to create joint text-image embeddings in a 512-dimensional multimodal semantic space. - Learn video processing techniques to extract keyframes, generate transcripts using Whisper, and create captions. - Use the LanceDB vector database to store and retrieve high-dimensional multimodal embeddings. - Integrate the LLaVA model, combining CLIP's (Contrastive Language Image Pretraining) vision transformer with Llama, for advanced visual-textual reasoning. Your final system will ingest video data, generate embeddings for frames and text, perform similarity searches for relevant content, and use the retrieved multimodal context to inform LVLM-based response generation. The result is a system capable of answering nuanced questions about video content, effectively chatting about the video it has processed. Please sign up here!

Andrew Ng

107,825 views • 1 year ago