Loading video...
Video Failed to Load
Vanilla SD turbo runs at ~1.5fps on my M1 Max with the diffusers lib. I wonder how close to realtime it can get? Maybe CoreML compilation, 6bits palettisation, TAESD and a few more tricks like RCFG from StreamDiffusion can push it to ~20fps?
38,762 views • 2 years ago •via X (Twitter)
10 Comments

Didn't have much time to dig in during the holidays, but TAESD & CoreML compilation gets it to ~2.5fps. Still ~260ms for 1 forward pass of the 1.8k unet ops on M1. Anyway, that's it for now - Happy new year y'all! 🥂

A quick win for that specific use case is running face detection and only generating the head, at proportional resolution. Your head is a small part of the frame so you could get a big speedup

Makes sense! It would also reduce the amount of flicker

on a 4090 i got down to 37ms with TAESD and stable-fast, i suspect an m1 max could do 50-60ms if you are able to reimplement all the stable-fast tricks, and beyond if you can do RCFG

Oh nice! I was assuming that stable-fast gains were mostly due to cuda related optimizations but there’s a lot more in fact

100fps awaits

I did a small experiment with RTX 3090 - the speed is round 20fps, but result is a bit flaky

Wow oh wow! I wanna try this one on M2 Ultra!!! Any hints on where to start?

Sure, I'm just using the diffusers lib with the mps device: on this model: Curious about the perfs on M2 Ultra

With some coreML optimisation you will gain some FPS for sure
