Video wird geladen...
Video konnte nicht geladen werden
Sharing something exciting we've been working on as a Thanksgiving gift: Diffusion Self-Distillation (DSD), which redefines zero-shot customized image generation using FLUX. DSD is like DreamBooth, but zero-shot/training-free. It works across any input subject and desired context—character consistency, item/asset adaptation, scene relighting, and more. It even enables the creation... show more
60,760 Aufrufe • vor 1 Jahr •via X (Twitter)
45 Kommentare

HF discussion page:

A simple 9-panel comic made with our model --- it took me less than 10 minutes

And another one

Character Adaptation

Item/Asset Adaptation --- could be useful for commercial Ads!

InstructPix2Pix-type Instruction Prompting (we can extract different subjects from an input image too!)

Scene Relighting

Will you be releasing the code? My team builds Dashtoon Studio to make comics with AI and we distribute and monetise them through the Dashtoon mobile app. Would love to try it out and give you feedback by trying it out in our large volume production.

yes, it will be open sourced in the future!

Great developments

Really cool work! 💥

Thx 😬

This is a great idea. Congratulations.

Thx! 😬

Results are looking super cool! 🎉!

Thanks Junlin! 😃

Awesome work!

Thx! 😬

Awesome)

great results! are you going to open it to test?

Thanks and yes, I'm training a more stable Grande version and as soon as that's finished, I'll try to open up a Gradio demo/merge to ComfyUI

that would be awesome, if you need testers let me know pls!

Wow cool results Shengqu!

Thanks Mohamad! 😬

Is the model available in HF?

Not yet. I am training a steadier "Grande" version, and will open source as soon as that's finished.

Can you support SD3.5 please?

I think this strongly depends on how the ecosystem of SD3.5 goes --- personally I'd love to try the method on a non-distilled vanilla model.

Amazing work

@_akhaliq $Trias for the win🚀🚀🚀

Super cool! A smart way to curate diverse paired data. I am curious what is the success rate.

Congratulations on your work and happy Thanksgiving! We've added it to The Next AI Tool.

Awesome, Thank you.

Amazing work.

What is this sorcery?! That’s amazing

Will write about that a bit later! But in short, we distill FLUX into a conditional two-frame video generation model, using its own outputs --- which in fact already contain large portions of identity-preserving contents!

Thanks for sharing the insights! I’d love to read more about it :)

Do you think there will be the possibility of the same thing but for multiple characters in a scene in the future with a similar approach but many images?

excellent work! Tried to reach out to you via email. Will try it again here :) It seems after generating grids you end-up with 4 256x256 images. Do you use upscaling before fine-tuning the parallel processing architecture? If yes, which one ? :)

We generate 1024×1024 grids so the pairs will be 512×512 --- thus all images here are also 512×512. However, we find the framework work well to higher resolutions, since the approach is scalable to what the base model can do. If we directly generative pairs instead of grids, the resolution can easily go up to 960×960, etc.

thanks. Makes sense

Amazing work

Thx! 😬

so much promise in this now - if you want it used, create nodes for @ComfyUI

@ComfyUI on my schedule!
