Video yükleniyor...

Video Yüklenemedi

Ana Sayfaya Dön

🚀 Introducing 𝐈𝐃𝐌-𝐕𝐓𝐎𝐍 : A novel diffusion model for image-based virtual try-on! 👗 😍 Improves garment fidelity and generates authentic visuals. 🔧 Uses two modules to encode garment semantics: visual encoder for high-level & parallel UNet for low-level. Links below👇

104,492 görüntüleme • 2 yıl önce •via X (Twitter)

8 Yorum

Gradio profil fotoğrafı
Gradio2 yıl önce

💡IDM-VTON provides good support for detailed textual prompts for garments & person images 🎨Customization method using person-garment image pairs significantly improves fidelity. 🚀 Ready to create your own virtual try-on apps? Build with Gradio today!

Gradio profil fotoğrafı
Gradio2 yıl önce

📊IDM-VTON outperforms other diffusion-based & GAN-based approaches in preserving garment details 🌟Looks effective in real-world scenarios! 🔍IDM-VTON Official Gradio demo has been released on @huggingface Spaces! 🔗Demo:

LatinoRevolution (rev/acc) profil fotoğrafı
LatinoRevolution (rev/acc)2 yıl önce

lol, this is great

Mahmoud Ghulman profil fotoğrafı
Mahmoud Ghulman2 yıl önce

The name is genius 😅

Yuchen Jin profil fotoğrafı
Yuchen Jin2 yıl önce

Very interesting work, but is this expected?

dinos profil fotoğrafı
dinos2 yıl önce

@yacineMTB killer @dingboard_ feature

dweedify profil fotoğrafı
dweedify2 yıl önce

Pretty cool!

Aswanth achoo'z profil fotoğrafı
Aswanth achoo'z2 yıl önce

😅lol

Benzer Videolar

MagicAnimate: Temporally Consistent Human Image Animation using Diffusion Model with Gradio demo local demo: This paper studies the human image animation task, which aims to generate a video of a certain reference identity following a particular motion sequence. Existing animation works typically employ the frame-warping technique to animate the reference image towards the target motion. Despite achieving reasonable results, these approaches face challenges in maintaining temporal consistency throughout the animation due to the lack of temporal modeling and poor preservation of reference identity. In this work, we introduce MagicAnimate, a diffusion-based framework that aims at enhancing temporal consistency, preserving reference image faithfully, and improving animation fidelity. To achieve this, we first develop a video diffusion model to encode temporal information. Second, to maintain the appearance coherence across frames, we introduce a novel appearance encoder to retain the intricate details of the reference image. Leveraging these two innovations, we further employ a simple video fusion technique to encourage smooth transitions for long video animation. Empirical results demonstrate the superiority of our method over baseline approaches on two benchmarks. Notably, our approach outperforms the strongest baseline by over 38% in terms of video fidelity on the challenging TikTok dancing dataset. Code and model will be made available.

AK

810,731 görüntüleme • 2 yıl önce

We've officially released and open-sourced HunyuanImage 2.1, our latest text-to-image model. The new model delivers on our commitment to balancing performance and quality. With native 2K image generation, HunyuanImage 2.1 is an advanced open-source text-to-image model.🎨 ✨ New in 2.1: 🔹Advanced Semantics: Supports ultra-long and complex prompts of up to 1000 tokens, and precisely controls the generation of multiple subjects in a single image. 🔹Precise Chinese and English Text Rendering with seamless image–text integration: The model naturally integrates text into images, making it suitable for a wide range of applications such as product covers, illustrations, and poster design to meet the needs of various fields. 🔹Rich Styles and High Aesthetic: Capable of generating images in various styles—including photorealistic portraits, comics, and vinyl figures—it delivers outstanding visual appeal and artistic quality. 🔹High-Quality Generation: Efficiently produces ultra-high-definition (2K) images in the same time other models take to generate a 1K image. HunyuanImage 2.1 uses two text encoders: a multimodal large language model (MLLM) to improve the model's image and text alignment capabilities, and a multi-language character-aware encoder to improve text rendering capabilities. The model is a single- and double-stream diffusion transformer with 17B parameters. We've also open-sourced the weights of the the accelerated version with meanflow which reduces inference steps from 100 to just 8, and PromptEnhancer, the first industrial-grade rewriting model that enhances your prompts for more nuanced and expressive image generation. Now, creators turn complex ideas—like posters with slogans or multi-panel comics—into visuals faster than ever. We’re just getting started. Stay tuned for our native multimodal image generation model coming soon. 🌐Website: 🔗Github: 🤗Hugging Face: ✨Hugging Face Demo:

Tencent Hy

89,392 görüntüleme • 1 yıl önce

MagicScroll: Nontypical Aspect-Ratio Image Generation for Visual Storytelling via Multi-Layered Semantic-Aware Denoising paper page: Visual storytelling often uses nontypical aspect-ratio images like scroll paintings, comic strips, and panoramas to create an expressive and compelling narrative. While generative AI has achieved great success and shown the potential to reshape the creative industry, it remains a challenge to generate coherent and engaging content with arbitrary size and controllable style, concept, and layout, all of which are essential for visual storytelling. To overcome the shortcomings of previous methods including repetitive content, style inconsistency, and lack of controllability, we propose MagicScroll, a multi-layered, progressive diffusion-based image generation framework with a novel semantic-aware denoising process. The model enables fine-grained control over the generated image on object, scene, and background levels with text, image, and layout conditions. We also establish the first benchmark for nontypical aspect-ratio image generation for visual storytelling including mediums like paintings, comics, and cinematic panoramas, with customized metrics for systematic evaluation. Through comparative and ablation studies, MagicScroll showcases promising results in aligning with the narrative text, improving visual coherence, and engaging the audience. We plan to release the code and benchmark in the hope of a better collaboration between AI researchers and creative practitioners involving visual storytelling.

AK

22,379 görüntüleme • 2 yıl önce