Загрузка видео...

Не удалось загрузить видео

На главную

Scaling up GANs for Text-to-Image Synthesis present our 1B-parameter GigaGAN, achieving lower FID than Stable Diffusion v1.5, DALL·E 2, and Parti-750M. It generates 512px outputs at 0.13s, orders of magnitude faster than diffusion and autoregressive models, and inherits the disentangled, continuous, and controllable latent space of GANs abs: project page:

278,115 просмотров • 3 лет назад •via X (Twitter)

Комментарии: 10

Фото профиля Daniel Losey 🔀
Daniel Losey 🔀3 лет назад

amazing

Фото профиля David Marx (@digthatdata.bsky.social)
David Marx (@digthatdata.bsky.social)3 лет назад

GANs are back baybee

Фото профиля Nicolay Mausz
Nicolay Mausz3 лет назад

Adobe research - I guess this will be part of CC

Фото профиля Draz ⚛️
Draz ⚛️3 лет назад

The upscaling is quite insane on how it accurately fills in details

Фото профиля Nerdy Rodent 🐀🤓💻
Nerdy Rodent 🐀🤓💻3 лет назад

It’s been hours now, why isn’t it showing up? 😉

Фото профиля Asriel H
Asriel H3 лет назад

It has the same schema of injecting latent vector into every scaling layer as StyleGAN has

Фото профиля okaris
okaris3 лет назад

The examples provided don’t look as good as diffusion models. Some details obscured or looking weird.

Фото профиля Adhik Joshi
Adhik Joshi3 лет назад

Weights aren't open-source

Фото профиля Julien Genoud
Julien Genoud3 лет назад

The 4k upsampler 🤯

Фото профиля Clarence Hu
Clarence Hu3 лет назад

paging @gwern

Похожие видео

We've officially released and open-sourced HunyuanImage 2.1, our latest text-to-image model. The new model delivers on our commitment to balancing performance and quality. With native 2K image generation, HunyuanImage 2.1 is an advanced open-source text-to-image model.🎨 ✨ New in 2.1: 🔹Advanced Semantics: Supports ultra-long and complex prompts of up to 1000 tokens, and precisely controls the generation of multiple subjects in a single image. 🔹Precise Chinese and English Text Rendering with seamless image–text integration: The model naturally integrates text into images, making it suitable for a wide range of applications such as product covers, illustrations, and poster design to meet the needs of various fields. 🔹Rich Styles and High Aesthetic: Capable of generating images in various styles—including photorealistic portraits, comics, and vinyl figures—it delivers outstanding visual appeal and artistic quality. 🔹High-Quality Generation: Efficiently produces ultra-high-definition (2K) images in the same time other models take to generate a 1K image. HunyuanImage 2.1 uses two text encoders: a multimodal large language model (MLLM) to improve the model's image and text alignment capabilities, and a multi-language character-aware encoder to improve text rendering capabilities. The model is a single- and double-stream diffusion transformer with 17B parameters. We've also open-sourced the weights of the the accelerated version with meanflow which reduces inference steps from 100 to just 8, and PromptEnhancer, the first industrial-grade rewriting model that enhances your prompts for more nuanced and expressive image generation. Now, creators turn complex ideas—like posters with slogans or multi-panel comics—into visuals faster than ever. We’re just getting started. Stay tuned for our native multimodal image generation model coming soon. 🌐Website: 🔗Github: 🤗Hugging Face: ✨Hugging Face Demo:

Tencent Hy

89,257 просмотров • 11 месяцев назад

In collaboration with Intel, our Depth Fusion showcases the power of our LDM3D diffusion model in generating 360° views from text prompts provided by the user. The LDM3D diffusion model generates a 2D RGB image and its corresponding relative depth map providing a complete RGBD representation corresponding to the text prompt. The LDM 3D model is a specialized version of the stable diffusion V 1.4 model that has been modified to fit both image and depth map data.The model was then fine tuned on a subset of the Laion400M data set - large scale image caption data set. The depth maps used to fine tune our model were generated by the DPTBeiT large 512 depth estimation model that provides highly accurate relative depth estimates for each pixel. We take the generated 2D RGB image and depth map and use them to compute a 360° projection using touchdesigner. Touchdesigner is a versatile platform that allows for the creation of immersive and interactive multimedia experiences. Our application harnesses the power of touchdesigner to bring the generated 360° views to life, providing users with a unique and engaging way to experience their text prompts, whether it’s a description of a tranquil forest, a noisy cityscape or a futuristic sci fi world. Our depth fusion can bring these concepts to life in a vivid and immersive detail. - Scottie Fox, VP Engineering Blockade Labs ScottieFox #AI #VR #3D #gamedev #stablediffusion

Blockade Labs

11,439 просмотров • 3 лет назад