正在加载视频...
视频加载失败
LLM-grounded Diffusion: Enhancing Prompt Understanding of Text-to-Image Diffusion Models with Large Language Models paper page: github: Recent advancements in text-to-image generation with diffusion models have yielded remarkable results synthesizing highly realistic and diverse images. However, these models still encounter difficulties when generating images from prompts that demand spatial or... show more
83,657 次观看 • 3 年前 •via X (Twitter)
6 条评论

Boyi Li3 年前
Thanks @_akhaliq for sharing our work!

zorr0 (ττ)3 年前
@replytensor

haareblond3 年前
cool but still feels hacky

Takomo AI3 年前
That's great progress!

Cavit Erginsoy3 年前
@yuliangxiu I saw this about a month ago and had played around with it, is the same or a parallel dev? Wish someone built an extension for A1111

VIJAY KUMAR REDDY BOMMIREDDY3 年前
Impressive work! Expanding the text-to-image domain with diffusion models showcases great potential. Looking forward to exploring the paper and GitHub repository. Keep up the great work! 👍
