Loading video...
Video Failed to Load
[1/5] Always wondered what people see when looking at a Rorschach test? SpaText - our recent #CVPR2023 paper from @MetaAI may give you a sneak peek! TL;DR: We extend text-to-image models with region-specific textual controllability. Project Page:
19,389 views • 3 years ago •via X (Twitter)
7 Comments

[2/5] Recent text-to-image diffusion models are able to generate convincing results of unprecedented quality. However, it is nearly impossible to control the shapes of different regions/objects or their layout in a fine-grained fashion.

[3/5] SpaText is a new method for text-to-image generation using open-vocabulary scene control. In addition to a global text prompt that describes the entire scene, the user provides a segmentation map where each region of interest is annotated by a free-form text description

[4/5] Due to lack of large-scale datasets that have a detailed textual description for each region in the image, we choose to leverage the current large-scale text-to-image datasets and base our approach on a novel CLIP-based spatio-textual representation

[5/5] Thanks to my great collaborators at @MetaAI: @THayes427, Oran Gafni, Sonal Gupta, Yaniv Taigman, @deviparikh, @xi_yin_, and to my supervisors @DaniLischinski and @ohadf! Special thanks to @m_mmandel for the Rorschach video art. #GenerativeAI #AIart #AIArtwork

@MetaAI Wow

@MetaAI Interesting read!

@MetaAI Thank you @OmriAvr! Please watch video with audio!


