正在加载视频...

视频加载失败

Scaling up GANs for Text-to-Image Synthesis present our 1B-parameter GigaGAN, achieving lower FID than Stable Diffusion v1.5, DALL·E 2, and Parti-750M. It generates 512px outputs at 0.13s, orders of magnitude faster than diffusion and autoregressive models, and inherits the disentangled, continuous, and controllable latent space of GANs abs: project page:

278,115 次观看 • 3 年前 •via X (Twitter)

10 条评论

Daniel Losey 🔀 的头像
Daniel Losey 🔀3 年前

amazing

David Marx (@digthatdata.bsky.social) 的头像
David Marx (@digthatdata.bsky.social)3 年前

GANs are back baybee

Nicolay Mausz 的头像
Nicolay Mausz3 年前

Adobe research - I guess this will be part of CC

Draz ⚛️ 的头像
Draz ⚛️3 年前

The upscaling is quite insane on how it accurately fills in details

Nerdy Rodent 🐀🤓💻 的头像
Nerdy Rodent 🐀🤓💻3 年前

It’s been hours now, why isn’t it showing up? 😉

Asriel H 的头像
Asriel H3 年前

It has the same schema of injecting latent vector into every scaling layer as StyleGAN has

okaris 的头像
okaris3 年前

The examples provided don’t look as good as diffusion models. Some details obscured or looking weird.

Adhik Joshi 的头像
Adhik Joshi3 年前

Weights aren't open-source

Julien Genoud 的头像
Julien Genoud3 年前

The 4k upsampler 🤯

Clarence Hu 的头像
Clarence Hu3 年前

paging @gwern

相关视频

We've officially released and open-sourced HunyuanImage 2.1, our latest text-to-image model. The new model delivers on our commitment to balancing performance and quality. With native 2K image generation, HunyuanImage 2.1 is an advanced open-source text-to-image model.🎨 ✨ New in 2.1: 🔹Advanced Semantics: Supports ultra-long and complex prompts of up to 1000 tokens, and precisely controls the generation of multiple subjects in a single image. 🔹Precise Chinese and English Text Rendering with seamless image–text integration: The model naturally integrates text into images, making it suitable for a wide range of applications such as product covers, illustrations, and poster design to meet the needs of various fields. 🔹Rich Styles and High Aesthetic: Capable of generating images in various styles—including photorealistic portraits, comics, and vinyl figures—it delivers outstanding visual appeal and artistic quality. 🔹High-Quality Generation: Efficiently produces ultra-high-definition (2K) images in the same time other models take to generate a 1K image. HunyuanImage 2.1 uses two text encoders: a multimodal large language model (MLLM) to improve the model's image and text alignment capabilities, and a multi-language character-aware encoder to improve text rendering capabilities. The model is a single- and double-stream diffusion transformer with 17B parameters. We've also open-sourced the weights of the the accelerated version with meanflow which reduces inference steps from 100 to just 8, and PromptEnhancer, the first industrial-grade rewriting model that enhances your prompts for more nuanced and expressive image generation. Now, creators turn complex ideas—like posters with slogans or multi-panel comics—into visuals faster than ever. We’re just getting started. Stay tuned for our native multimodal image generation model coming soon. 🌐Website: 🔗Github: 🤗Hugging Face: ✨Hugging Face Demo:

Tencent Hy

89,257 次观看 • 11 个月前