Video wird geladen...

Video konnte nicht geladen werden

Zur Startseite

Introducing 🔥BLIP-Diffusion🔥, a novel method for enabling Text-to-image Diffusion models with multimodal controllable generation/editing, powered by BLIP-2 pre-trained text-aligned subject representation. Paper: Project: (1/n)

40,615 Aufrufe • vor 3 Jahren •via X (Twitter)

10 Kommentare

Profilbild von Steven Hoi
Steven Hoivor 3 Jahren

BLIP-Diffusion learns pretrained subject representation to unlock a range of zero-shot/few-step-tuned image generation and editing capabilities, e.g., subject-driven generation, zero-shot subject-driven image manipulation, controllable subject-driven image editing, etc. (2/n)

Profilbild von Steven Hoi
Steven Hoivor 3 Jahren

Two-stage pretraining strategy: 1) multimodal representation learning with BLIP-2 to produce text-aligned visual features for an input image; 2) subject representation learning trains the Diffusion models to use the features by BLIP-2 to generate novel subject renditions. (3/n)

Profilbild von Steven Hoi
Steven Hoivor 3 Jahren

BLIP-Diffusion can be extended on-the-fly without retraining with other existing controllable generation techniques, such as “ControlNet” and “prompt-to-prompt image editing”, to achieve more advanced multimodal controllable image generation/editing capabilities. (4/n)

Profilbild von Steven Hoi
Steven Hoivor 3 Jahren

Demo-I: Subject-driven Text-to-Image Generation Given one or a few images of a subject, our model can generate novel renditions of the subject based on text prompts. This figure shows some such subject-driven text-to-image generation results on the DreamBooth dataset. (5/n)

Profilbild von Steven Hoi
Steven Hoivor 3 Jahren

Demo-II: Zero-shot Subject-driven Image Manipulation Our model can extract subject features to guide the generation and enable intriguing and useful applications of zero-shot image manipulation, including subject-driven style transfer and subject interpolation. (6/n)

Profilbild von Steven Hoi
Steven Hoivor 3 Jahren

Demo-III: Subject-driven Image Editing BLIP-Diffusion can perform “subject-driven image editing” in a multimodal editing fashion, which edits a source image by replacing one subject specified in text with another subject specified by an input reference image. (7/n)

Profilbild von Steven Hoi
Steven Hoivor 3 Jahren

BLIP-Diffusion demonstrates that BLIP2 is a rather generic multimodal representation learning framework, which not only has state-of-the-art multimodal-to-text generation capabilities, but also can unlock a range of impressive multimodal-to-image generation capabilities. (8/n)

Profilbild von Steven Hoi
Steven Hoivor 3 Jahren

Find out more details from our research paper here: BLIP-Diffusion: Pre-trained Subject Representation for Controllable Text-to-Image Generation and Editing Paper: Another great work with @DongxuLi_ @LiJunnan0409 from our AI team at @SFResearch (9/n)

Profilbild von Shivam Kumar
Shivam Kumarvor 3 Jahren

Awesome. I just woke up and there is already a new model. How am I supposed to keep up?

Profilbild von Bert Christiaens
Bert Christiaensvor 3 Jahren

Looks amazing! 🤯🤯 Can't wait to try it out!! Are there plans of putting it on HuggingFace hub?

Ähnliche Videos