正在加载视频...

视频加载失败

🚀Introducing LLaVA Lightning: Train a lite, multimodal GPT-4 with just $40 in 3 hours! With our newly introduced datasets and the efficient design of LLaVA, you can now turbocharge your language model with image reasoning capabilities, in an incredibly affordable way.🧵

302,319 次观看 • 3 年前 •via X (Twitter)

10 条评论

Haotian Liu 的头像
Haotian Liu3 年前

(2/5) Excited to release a 558K concept-balanced subset of LAION/CC/SBU & an 80K high-quality subset of LLaVA-Instruct-158K. The concept-balanced subset ensures a broad concept coverage, and the high-quality visual instruct tuning data enables models' visual reasoning capability.

Haotian Liu 的头像
Haotian Liu3 年前

(3/5) Upgrade your Vicuna-7B to LLaVA-Lightning in just 3 hrs: 2 hrs pretraining + 1 hr visual instruct tuning. Train on 8x A100s using cloud spot instances for just $40. Let's make this research more accessible to researchers, academia, and millions of AI enthusiasts today!

Haotian Liu 的头像
Haotian Liu3 年前

(4/5) We're also upgrading LLaVA to support Vicuna v0 & v1 weights, with more checkpoints arriving this week! Plus, we're working to support more hardware – stay tuned!

Haotian Liu 的头像
Haotian Liu3 年前

(5/5) 🤗 Demo: 🌐 Project page: 📄 Paper: Embark on your LLaVA-Lightning journey today and stay tuned for more models and support for more hardwares in the following weeks!

iamrobotbear (bk) 的头像
iamrobotbear (bk)3 年前

Any way to easily swap LLaMA out for OpenAI or Dolly 2.0?

Haotian Liu 的头像
Haotian Liu3 年前

Yes, it is definitely possible. And even easier with the introduction of LLaVA lightning. MPT-7B just joins the LLaVA family today!

Zongheng Yang 的头像
Zongheng Yang3 年前

Congrats on the work @imhaotian. Glad to see SkyPilot was of help!

Chris 的头像
Chris3 年前

Recently program of open source projects is super fast, definitely surpass my expectation.

web3工作坊 的头像
web3工作坊3 年前

@_akhaliq Thank you for sharing. @savetonotion #tweet #AI

Jake Harrison 的头像
Jake Harrison3 年前

Good job!

相关视频

New short course Multimodal RAG: Chat with Videos, developed with Intel and taught by vasudevlal! In this course, you’ll work with LLaVA (Large Language and Vision Assistant), a Large Vision Language Model (LVLM) that can process both images and text. For example, given an image of a person doing a handstand on a skateboard at the beach, LLaVA doesn't just caption the scene, it’s able to predict possible outcomes, like the person losing balance or falling off. By understanding not just what's in a video frame, but what might happen next, your application can provide more insightful answers to questions about video. You'll build a full multimodal RAG pipeline that can chat about video content: - Use the BridgeTower model to create joint text-image embeddings in a 512-dimensional multimodal semantic space. - Learn video processing techniques to extract keyframes, generate transcripts using Whisper, and create captions. - Use the LanceDB vector database to store and retrieve high-dimensional multimodal embeddings. - Integrate the LLaVA model, combining CLIP's (Contrastive Language Image Pretraining) vision transformer with Llama, for advanced visual-textual reasoning. Your final system will ingest video data, generate embeddings for frames and text, perform similarity searches for relevant content, and use the retrieved multimodal context to inform LVLM-based response generation. The result is a system capable of answering nuanced questions about video content, effectively chatting about the video it has processed. Please sign up here!

Andrew Ng

107,825 次观看 • 1 年前