Video wird geladen...

Video konnte nicht geladen werden

Zur Startseite

๐Ÿš€ Pixel Diffusion Decoder (PiD) v1.5 is out ๐ŸŽจ No color-shifting problem, better 4K visual quality ๐Ÿ“ฆ Undistilled checkpoint and full training code ๐Ÿงฉ Support FLUX, FLUX2, Qwen-Image, Z-Image, ... Feel free to use it, reproduce it, and build on top of it ๐Ÿ’ป Code: ๐Ÿ”— Demo & comparison:

51,051 Aufrufe โ€ข vor 1 Monat โ€ขvia X (Twitter)

0 Kommentare

Keine Kommentare verfรผgbar

Kommentare vom Original-Post werden hier angezeigt

ร„hnliche Videos

We've officially released and open-sourced HunyuanImage 2.1, our latest text-to-image model. The new model delivers on our commitment to balancing performance and quality. With native 2K image generation, HunyuanImage 2.1 is an advanced open-source text-to-image model.๐ŸŽจ โœจ New in 2.1: ๐Ÿ”นAdvanced Semantics: Supports ultra-long and complex prompts of up to 1000 tokens, and precisely controls the generation of multiple subjects in a single image. ๐Ÿ”นPrecise Chinese and English Text Rendering with seamless imageโ€“text integration: The model naturally integrates text into images, making it suitable for a wide range of applications such as product covers, illustrations, and poster design to meet the needs of various fields. ๐Ÿ”นRich Styles and High Aesthetic: Capable of generating images in various stylesโ€”including photorealistic portraits, comics, and vinyl figuresโ€”it delivers outstanding visual appeal and artistic quality. ๐Ÿ”นHigh-Quality Generation: Efficiently produces ultra-high-definition (2K) images in the same time other models take to generate a 1K image. HunyuanImage 2.1 uses two text encoders: a multimodal large language model (MLLM) to improve the model's image and text alignment capabilities, and a multi-language character-aware encoder to improve text rendering capabilities. The model is a single- and double-stream diffusion transformer with 17B parameters. We've also open-sourced the weights of the the accelerated version with meanflow which reduces inference steps from 100 to just 8, and PromptEnhancer, the first industrial-grade rewriting model that enhances your prompts for more nuanced and expressive image generation. Now, creators turn complex ideasโ€”like posters with slogans or multi-panel comicsโ€”into visuals faster than ever. Weโ€™re just getting started. Stay tuned for our native multimodal image generation model coming soon. ๐ŸŒWebsite: ๐Ÿ”—Github: ๐Ÿค—Hugging Face: โœจHugging Face Demo:

Tencent Hy

89,257 Aufrufe โ€ข vor 11 Monaten

Weโ€™re excited to announce the release and open-source of HunyuanImage 3.0 โ€” the largest and most powerful open-source text-to-image model to date, with over 80 billion total parameters, of which 13 billion are activated per token during inference.The effect is completely comparable to the industryโ€™s flagship closed-source model.๐Ÿš€๐Ÿš€๐Ÿš€ HunyuanImage 3.0 originates from our internally developed native multimodal large language model, with fine-tuning and post-training focused on text-to-image generation. This unique foundation gives the model a powerful set of capabilities: โœ…Reason with world knowledge โœ…Understand complex, thousand-word prompts โœ…Generate precise text within images Different from traditional DiT architecture image generation models, HunyuanImage 3.0โ€™s MoE architecture uses a Transfusion-based approach to deeply couple Diffusion and LLM training for a single, powerful system. Built on Hunyuan-A13B, HunyuanImage 3.0 was trained on a massive dataset: 5 billion image-text pairs, video frames, interleaved image-text data, and 6 trillion tokens of text corpora. This hybrid training across multimodal generation, understanding, and LLM capabilities allows the model to seamlessly integrate multiple tasks. Whether you're an illustrator, designer, or creator, this is built to slash your workflow from hours to minutes. HunyuanImage 3.0 can generate intricate text, detailed comics, expressive emojis, and lively, engaging illustrations for educational content. The current release focuses solely on text-to-image generation and future updates will include image-to-image, image editing, multi-turn interaction, and more. ๐Ÿ‘‰๐ŸปTry it now: ๐Ÿ”—GitHub: ๐Ÿค—Hugging Face:

Tencent Hy

413,002 Aufrufe โ€ข vor 11 Monaten