apolinario (poli)'s banner
apolinario (poli)'s profile picture

apolinario (poli)

@multimodalart16,147 subscribers

ML Engineer for Art and Creativity @HuggingFace ([email protected])

Shorts

MiniMax H3 image+audio to video feed an audio into h3 instead of generating it. i was testing this flow, got shocked with how good the lipsync and movement is implemented it on the i2v version in diffusers, made a demo for it. what do you think? 🔊 ▶️

MiniMax H3 image+audio to video feed an audio into h3 instead of generating it. i was testing this flow, got shocked with how good the lipsync and movement is implemented it on the i2v version in diffusers, made a demo for it. what do you think? 🔊 ▶️

23,194 views

Boring Reality LoRA just dropped for HunyuanVideo 🏙️🏞️ A fine-tune that lead not to cinematic shots, but to something that could've come out of your phone 📱

Boring Reality LoRA just dropped for HunyuanVideo 🏙️🏞️ A fine-tune that lead not to cinematic shots, but to something that could've come out of your phone 📱

407,400 views

testing out the Diffusers Image Fill demo capabilities on a random image

testing out the Diffusers Image Fill demo capabilities on a random image

274,371 views

Qwen Image Multiple Angles LoRA is an exquisitely trained LoRA! 📐˚₊‧꒰ა Keep character and scenes consistent, and flies the camera around! Open source got there! One of the best LoRAs I've come across lately 🙌

Qwen Image Multiple Angles LoRA is an exquisitely trained LoRA! 📐˚₊‧꒰ა Keep character and scenes consistent, and flies the camera around! Open source got there! One of the best LoRAs I've come across lately 🙌

121,656 views

GPUs for all 🤗 Creating ZeroGPU demos or apps is now available to ALL Hugging Face users tell your agent: "Build a HF ZeroGPU demo for this model" New Space: Agent skill:

GPUs for all 🤗 Creating ZeroGPU demos or apps is now available to ALL Hugging Face users tell your agent: "Build a HF ZeroGPU demo for this model" New Space: Agent skill:

24,719 views

Stable Audio 3 by Stability AI is just out It mainly comes with 3 open source variants: - Stable Audio 3 Medium (2B) - Stable Audio 3 Small (0.6B) - Music - Stable Audio 3 Small (0.6B) - VFX (and a "large" closed variant) The open models are really fast and high quality

Stable Audio 3 by Stability AI is just out It mainly comes with 3 open source variants: - Stable Audio 3 Medium (2B) - Stable Audio 3 Small (0.6B) - Music - Stable Audio 3 Small (0.6B) - VFX (and a "large" closed variant) The open models are really fast and high quality

41,668 views

I just built a demo for this Light Migration LoRA on Hugging Face the quality surprises me on every output 🤯

I just built a demo for this Light Migration LoRA on Hugging Face the quality surprises me on every output 🤯

80,276 views

NVidia just released PiD: super resolution in pixel space directly from model latents 🔎 4X resolution for any generated image, FAST! 🏎️💨 FLUX.1, 2 and Z-Image (Qwen Image coming) of course, i built a demo: generate 4K images with Z-Image

NVidia just released PiD: super resolution in pixel space directly from model latents 🔎 4X resolution for any generated image, FAST! 🏎️💨 FLUX.1, 2 and Z-Image (Qwen Image coming) of course, i built a demo: generate 4K images with Z-Image

29,948 views

Apply Texture Qwen Image Edit LoRA by tarn59 works with EVERYTHING! 👉🪵🧶, this model trains so well I've built this demo so you can apply *any* texture to *any* object on Hugging Face

Apply Texture Qwen Image Edit LoRA by tarn59 works with EVERYTHING! 👉🪵🧶, this model trains so well I've built this demo so you can apply *any* texture to *any* object on Hugging Face

67,831 views

Introducing Kontext Relight! 💡 ✨ A FLUX Kontext Relight LoRA + demo trained for state-of-the art relighting for subjects & landscapes

Introducing Kontext Relight! 💡 ✨ A FLUX Kontext Relight LoRA + demo trained for state-of-the art relighting for subjects & landscapes

76,088 views

LLaDA (the first Large Language Diffusion Model) is *just* out 💥 and I've built a demo, try out now 👨‍💻 It's mesmerizing to watch the diffusion process 🌀, and it being a diffusion model gives you superpowers like "the 4th word has to be pineapple" 🦸 Demo and weights 👇

LLaDA (the first Large Language Diffusion Model) is *just* out 💥 and I've built a demo, try out now 👨‍💻 It's mesmerizing to watch the diffusion process 🌀, and it being a diffusion model gives you superpowers like "the 4th word has to be pineapple" 🦸 Demo and weights 👇

82,599 views

Excited to introduce LEDITS++, a novel way to edit real images with precision ✏️ - Multiple edits ✂️🔁 - Automagic free masking 🪄🎭 - 🆕 DPM-Solver fast inversion 🔀⚡ 🤗 Try it: 🔗 Project: 📝 Paper

Excited to introduce LEDITS++, a novel way to edit real images with precision ✏️ - Multiple edits ✂️🔁 - Automagic free masking 🪄🎭 - 🆕 DPM-Solver fast inversion 🔀⚡ 🤗 Try it: 🔗 Project: 📝 Paper

131,559 views

introducing the media synthesis museum an active and interactive entity created to preserve generative cultural objects it starts as a Hugging Face organization that contains modern code for old techniques: VQGAN+CLIP, DALL-E Mini, ModelScope Video, Stable Diffusion 1.5 you can use old models/technique directly on Spaces or locally on modern hardware/software, without the old "colab notebook" dependency rot the idea is to really preserve and make accessible those artifacts and aesthetics - both open source. In the future, we hope to also have also historically relevant closed source like DALL-E 1 and DALL-E 2 from OpenAI, older Midjourney models, older Runway apps/techniques/models (cc Cristóbal Valenzuela David Sam Altman)

introducing the media synthesis museum an active and interactive entity created to preserve generative cultural objects it starts as a Hugging Face organization that contains modern code for old techniques: VQGAN+CLIP, DALL-E Mini, ModelScope Video, Stable Diffusion 1.5 you can use old models/technique directly on Spaces or locally on modern hardware/software, without the old "colab notebook" dependency rot the idea is to really preserve and make accessible those artifacts and aesthetics - both open source. In the future, we hope to also have also historically relevant closed source like DALL-E 1 and DALL-E 2 from OpenAI, older Midjourney models, older Runway apps/techniques/models (cc Cristóbal Valenzuela David Sam Altman)

11,143 views

GANs are so back?! Scientists from Brown and Cornell have published a paper with a ✨ modern architecture GAN ✨ that is 🗿 stable to train 🗿 and competitive with SOTA GANs and even diffusion models Paper and demo 👇

GANs are so back?! Scientists from Brown and Cornell have published a paper with a ✨ modern architecture GAN ✨ that is 🗿 stable to train 🗿 and competitive with SOTA GANs and even diffusion models Paper and demo 👇

64,122 views

Editing facial expressions in real time now on Hugging Face Spaces 👨‍🎤🔀 A Grog converted Cog image to Gradio running a ComfyUI backend - magic of open source 🤝 ▶️

Editing facial expressions in real time now on Hugging Face Spaces 👨‍🎤🔀 A Grog converted Cog image to Gradio running a ComfyUI backend - magic of open source 🤝 ▶️

71,701 views

You can now finally create your own stock photo smiling while eating salad in seconds 👨‍🎤🥗 IP-Apdater-FaceID Plus was silently released last week - it's first inference technique time face really captures my likeness 🥸🦚 ▶️

You can now finally create your own stock photo smiling while eating salad in seconds 👨‍🎤🥗 IP-Apdater-FaceID Plus was silently released last week - it's first inference technique time face really captures my likeness 🥸🦚 ▶️

60,745 views

The Dream 7B (diffusion reasoning language model) is OUT! 🚨 I built a demo so you can test it out (and check the diffusion process live) 𖣯🔍

The Dream 7B (diffusion reasoning language model) is OUT! 🚨 I built a demo so you can test it out (and check the diffusion process live) 𖣯🔍

35,679 views

Videos

No more content to load