
Alexey Fateev
@superalesha • 1,840 subscribers
⚡I benchmark local LLMs on 4x RTX 3090s. exact configs, tok/s, VRAM, and what broke. ❤️ https://t.co/tIPzthsdkv - 2xDGX Spark 🚀96GB VRAM | Local AI
Shorts
Videos

I sped up deepseek v4 flash by 29x on my 4x3090s !!! No, its not joke. 15 -> 443 t/s. a 23k prompt used to take 25 mins. Now it takes 53 secs. 284b in 2bit, 87gb, barely squeezes into 96gb. Me and Fable 5 spent 4 days in llama.cpp. Fixed everything that was broken
Alexey Fateev121,475 görüntüleme • 1 ay önce
Daha fazla içerik yok.