Video wird geladen...

Video konnte nicht geladen werden

Zur Startseite

Scaling up GANs for Text-to-Image Synthesis present our 1B-parameter GigaGAN, achieving lower FID than Stable Diffusion v1.5, DALL·E 2, and Parti-750M. It generates 512px outputs at 0.13s, orders of magnitude faster than diffusion and autoregressive models, and inherits the disentangled, continuous, and controllable latent space of GANs abs: project page:

278,115 Aufrufe • vor 3 Jahren •via X (Twitter)

10 Kommentare

Profilbild von Daniel Losey 🔀
Daniel Losey 🔀vor 3 Jahren

amazing

Profilbild von David Marx (@digthatdata.bsky.social)
David Marx (@digthatdata.bsky.social)vor 3 Jahren

GANs are back baybee

Profilbild von Nicolay Mausz
Nicolay Mauszvor 3 Jahren

Adobe research - I guess this will be part of CC

Profilbild von Draz ⚛️
Draz ⚛️vor 3 Jahren

The upscaling is quite insane on how it accurately fills in details

Profilbild von Nerdy Rodent 🐀🤓💻
Nerdy Rodent 🐀🤓💻vor 3 Jahren

It’s been hours now, why isn’t it showing up? 😉

Profilbild von Asriel H
Asriel Hvor 3 Jahren

It has the same schema of injecting latent vector into every scaling layer as StyleGAN has

Profilbild von okaris
okarisvor 3 Jahren

The examples provided don’t look as good as diffusion models. Some details obscured or looking weird.

Profilbild von Adhik Joshi
Adhik Joshivor 3 Jahren

Weights aren't open-source

Profilbild von Julien Genoud
Julien Genoudvor 3 Jahren

The 4k upsampler 🤯

Profilbild von Clarence Hu
Clarence Huvor 3 Jahren

paging @gwern

Ähnliche Videos

Meta just announced FlowVid Taming Imperfect Optical Flows for Consistent Video-to-Video Synthesis paper page: Diffusion models have transformed the image-to-image (I2I) synthesis and are now permeating into videos. However, the advancement of video-to-video (V2V) synthesis has been hampered by the challenge of maintaining temporal consistency across video frames. This paper proposes a consistent V2V synthesis framework by jointly leveraging spatial conditions and temporal optical flow clues within the source video. Contrary to prior methods that strictly adhere to optical flow, our approach harnesses its benefits while handling the imperfection in flow estimation. We encode the optical flow via warping from the first frame and serve it as a supplementary reference in the diffusion model. This enables our model for video synthesis by editing the first frame with any prevalent I2I models and then propagating edits to successive frames. Our V2V model, FlowVid, demonstrates remarkable properties: (1) Flexibility: FlowVid works seamlessly with existing I2I models, facilitating various modifications, including stylization, object swaps, and local edits. (2) Efficiency: Generation of a 4-second video with 30 FPS and 512x512 resolution takes only 1.5 minutes, which is 3.1x, 7.2x, and 10.5x faster than CoDeF, Rerender, and TokenFlow, respectively. (3) High-quality: In user studies, our FlowVid is preferred 45.7% of the time, outperforming CoDeF (3.5%), Rerender (10.2%), and TokenFlow (40.4%).

AK

123,729 Aufrufe • vor 2 Jahren