Загрузка видео...

Не удалось загрузить видео

На главную

🚀 Introducing 𝐈𝐃𝐌-𝐕𝐓𝐎𝐍 : A novel diffusion model for image-based virtual try-on! 👗 😍 Improves garment fidelity and generates authentic visuals. 🔧 Uses two modules to encode garment semantics: visual encoder for high-level & parallel UNet for low-level. Links below👇

104,492 просмотров • 2 лет назад •via X (Twitter)

Комментарии: 8

Фото профиля Gradio
Gradio2 лет назад

💡IDM-VTON provides good support for detailed textual prompts for garments & person images 🎨Customization method using person-garment image pairs significantly improves fidelity. 🚀 Ready to create your own virtual try-on apps? Build with Gradio today!

Фото профиля Gradio
Gradio2 лет назад

📊IDM-VTON outperforms other diffusion-based & GAN-based approaches in preserving garment details 🌟Looks effective in real-world scenarios! 🔍IDM-VTON Official Gradio demo has been released on @huggingface Spaces! 🔗Demo:

Фото профиля LatinoRevolution (rev/acc)
LatinoRevolution (rev/acc)2 лет назад

lol, this is great

Фото профиля Mahmoud Ghulman
Mahmoud Ghulman2 лет назад

The name is genius 😅

Фото профиля Yuchen Jin
Yuchen Jin2 лет назад

Very interesting work, but is this expected?

Фото профиля dinos
dinos2 лет назад

@yacineMTB killer @dingboard_ feature

Фото профиля dweedify
dweedify2 лет назад

Pretty cool!

Фото профиля Aswanth achoo'z
Aswanth achoo'z2 лет назад

😅lol

Похожие видео

MagicAnimate: Temporally Consistent Human Image Animation using Diffusion Model with Gradio demo local demo: This paper studies the human image animation task, which aims to generate a video of a certain reference identity following a particular motion sequence. Existing animation works typically employ the frame-warping technique to animate the reference image towards the target motion. Despite achieving reasonable results, these approaches face challenges in maintaining temporal consistency throughout the animation due to the lack of temporal modeling and poor preservation of reference identity. In this work, we introduce MagicAnimate, a diffusion-based framework that aims at enhancing temporal consistency, preserving reference image faithfully, and improving animation fidelity. To achieve this, we first develop a video diffusion model to encode temporal information. Second, to maintain the appearance coherence across frames, we introduce a novel appearance encoder to retain the intricate details of the reference image. Leveraging these two innovations, we further employ a simple video fusion technique to encourage smooth transitions for long video animation. Empirical results demonstrate the superiority of our method over baseline approaches on two benchmarks. Notably, our approach outperforms the strongest baseline by over 38% in terms of video fidelity on the challenging TikTok dancing dataset. Code and model will be made available.

AK

810,578 просмотров • 2 лет назад

We've officially released and open-sourced HunyuanImage 2.1, our latest text-to-image model. The new model delivers on our commitment to balancing performance and quality. With native 2K image generation, HunyuanImage 2.1 is an advanced open-source text-to-image model.🎨 ✨ New in 2.1: 🔹Advanced Semantics: Supports ultra-long and complex prompts of up to 1000 tokens, and precisely controls the generation of multiple subjects in a single image. 🔹Precise Chinese and English Text Rendering with seamless image–text integration: The model naturally integrates text into images, making it suitable for a wide range of applications such as product covers, illustrations, and poster design to meet the needs of various fields. 🔹Rich Styles and High Aesthetic: Capable of generating images in various styles—including photorealistic portraits, comics, and vinyl figures—it delivers outstanding visual appeal and artistic quality. 🔹High-Quality Generation: Efficiently produces ultra-high-definition (2K) images in the same time other models take to generate a 1K image. HunyuanImage 2.1 uses two text encoders: a multimodal large language model (MLLM) to improve the model's image and text alignment capabilities, and a multi-language character-aware encoder to improve text rendering capabilities. The model is a single- and double-stream diffusion transformer with 17B parameters. We've also open-sourced the weights of the the accelerated version with meanflow which reduces inference steps from 100 to just 8, and PromptEnhancer, the first industrial-grade rewriting model that enhances your prompts for more nuanced and expressive image generation. Now, creators turn complex ideas—like posters with slogans or multi-panel comics—into visuals faster than ever. We’re just getting started. Stay tuned for our native multimodal image generation model coming soon. 🌐Website: 🔗Github: 🤗Hugging Face: ✨Hugging Face Demo:

Tencent Hy

89,257 просмотров • 11 месяцев назад