Video yükleniyor...

Video Yüklenemedi

Ana Sayfaya Dön

Nvidia presents ConsiStory Training-Free Consistent Text-to-Image Generation paper page: enable Stable Diffusion XL (SDXL) to generate consistent subjects across a series of images, without additional training.

161,685 görüntüleme • 2 yıl önce •via X (Twitter)

10 Yorum

Furkan Gözükara profil fotoğrafı
Furkan Gözükara2 yıl önce

We already do this with Dreambooth what is the difference advantage? and still no code :

Cesar Silva profil fotoğrafı
Cesar Silva2 yıl önce

@Grigomesmo @omisil44

pressed tin profil fotoğrafı
pressed tin2 yıl önce

things are gonna get wild huh

ari profil fotoğrafı
ari2 yıl önce

The future of video games is gonna be wild

Vahi Güner profil fotoğrafı
Vahi Güner2 yıl önce

Important need, character consistency

𝕄𝕚𝕔𝕙𝕒𝕖𝕝 𝕁𝕒𝕞𝕖𝕤 🦅 🇺🇸 profil fotoğrafı
𝕄𝕚𝕔𝕙𝕒𝕖𝕝 𝕁𝕒𝕞𝕖𝕤 🦅 🇺🇸2 yıl önce

🔥

DKRacingFan profil fotoğrafı
DKRacingFan2 yıl önce

How does it learn without training?

rk⚡fg  profil fotoğrafı
rk⚡fg 2 yıl önce

What would the subject look like if he was...

maru profil fotoğrafı
maru2 yıl önce

Cool

cardoso profil fotoğrafı
cardoso2 yıl önce

@artificialguybr

Benzer Videolar

We’re excited to announce the release and open-source of HunyuanImage 3.0 — the largest and most powerful open-source text-to-image model to date, with over 80 billion total parameters, of which 13 billion are activated per token during inference.The effect is completely comparable to the industry’s flagship closed-source model.🚀🚀🚀 HunyuanImage 3.0 originates from our internally developed native multimodal large language model, with fine-tuning and post-training focused on text-to-image generation. This unique foundation gives the model a powerful set of capabilities: ✅Reason with world knowledge ✅Understand complex, thousand-word prompts ✅Generate precise text within images Different from traditional DiT architecture image generation models, HunyuanImage 3.0’s MoE architecture uses a Transfusion-based approach to deeply couple Diffusion and LLM training for a single, powerful system. Built on Hunyuan-A13B, HunyuanImage 3.0 was trained on a massive dataset: 5 billion image-text pairs, video frames, interleaved image-text data, and 6 trillion tokens of text corpora. This hybrid training across multimodal generation, understanding, and LLM capabilities allows the model to seamlessly integrate multiple tasks. Whether you're an illustrator, designer, or creator, this is built to slash your workflow from hours to minutes. HunyuanImage 3.0 can generate intricate text, detailed comics, expressive emojis, and lively, engaging illustrations for educational content. The current release focuses solely on text-to-image generation and future updates will include image-to-image, image editing, multi-turn interaction, and more. 👉🏻Try it now: 🔗GitHub: 🤗Hugging Face:

Tencent Hy

412,658 görüntüleme • 10 ay önce