Kyle Hessling's banner
Kyle Hessling's profile picture

Kyle Hessling

@KyleHessling1 • 7,495 subscribers

Father | Local AI Infra Engineer | Striving to be like Christ

Shorts

Qwen 3.8 27B at 56tps; on 9 year old GPU btw Nvidia V100 32GB ~$650 on EBay right now! Using Dflash 2; disabling the ECC adds some more speed too! Thinking and prose is a bit slower, but 56-63 tps in code gen! MTP runs faster for prose vs DFlash2 but slower sustained code generation speed. MTP also runs much faster power limited than DFlash does. Working on a repo so you can get up and going quickly. Fun fact, the Nvidia v100 was $11,500 per card when they first launched. Price you pay for future proofing I guess; they’re still great cards. Pcie 3.0 and the older software/architecture are the only drawbacks, but also those aren’t as much of an issue as you’d think. Especially when you consider the price today!

Qwen 3.8 27B at 56tps; on 9 year old GPU btw Nvidia V100 32GB ~$650 on EBay right now! Using Dflash 2; disabling the ECC adds some more speed too! Thinking and prose is a bit slower, but 56-63 tps in code gen! MTP runs faster for prose vs DFlash2 but slower sustained code generation speed. MTP also runs much faster power limited than DFlash does. Working on a repo so you can get up and going quickly. Fun fact, the Nvidia v100 was $11,500 per card when they first launched. Price you pay for future proofing I guess; they’re still great cards. Pcie 3.0 and the older software/architecture are the only drawbacks, but also those aren’t as much of an issue as you’d think. Especially when you consider the price today!

175,788 görüntüleme

Qwen 3.8 27B feels SO GOOOOOD! Here’s the thing, it’s thinking a lot, but for the first time it may be that the verbose thinking is less of a mistake than a lot of other local models. It’s basically doing a full build, like a model would in a harness, but entirely within its reasoning trajectory. So the final result will be as close as possible to a final output. There’s much less second guessing itself in reasoning than with Qwen 3.6. It’s very confidently peacing this together so it can have the full picture for the output. This might be the play for smaller models. If the completeness of thought can overcome the increase in wall time, I’m pumped, we can always make it run faster! I think this is going to be really exceptional, I will post the output soon 😎

Qwen 3.8 27B feels SO GOOOOOD! Here’s the thing, it’s thinking a lot, but for the first time it may be that the verbose thinking is less of a mistake than a lot of other local models. It’s basically doing a full build, like a model would in a harness, but entirely within its reasoning trajectory. So the final result will be as close as possible to a final output. There’s much less second guessing itself in reasoning than with Qwen 3.6. It’s very confidently peacing this together so it can have the full picture for the output. This might be the play for smaller models. If the completeness of thought can overcome the increase in wall time, I’m pumped, we can always make it run faster! I think this is going to be really exceptional, I will post the output soon 😎

65,585 görüntüleme

Videos

KyleHessling1's profile picture

First impressions on Muse Glimmer! It's incredibly fast for a dense model, currently running an average of 208tps with a max of 274tps on a single 5090 with their DFLASH config. Comparatively, though, both using Open Code, Qwopus Coder (with thinking off) produced a much better shark survival game than the one I got from Glimmer. Meta's new dense model is currently just lacking some HTML canvas taste, but this is something that can be added via SFT as long as the model is stable and capable from a back-end programming perspective. And it seems to be, without a doubt. The big kicker here is that I ran this at extra high thinking, and it did not take long at all to run. Our current local leader, Qwen 27B 3.6, has a tendency to overthink, but with glimmer, that is not the case. Right now, my recommendation for general local programming (Apps, Games, Websites, Visual Tools) in this class is still Qwopus Coder with thinking disabled, or Qwopus Fusion with thinking enabled. Of course Shark Survival is a very basic domain-specific test, but I find that the result scales very well across many domains. If we're going to be shipping apps generated entirely locally, visual taste is somewhat of a bare minimum requirement, solely in my opinion, and Qwen's models in this class offer significantly more at the moment. That's actually why I initially started getting into finetuning with Qwen 3.5, they were the first base that was able to do really good front-end with some opus-trace fine-tuning. Qwen 3.6 has taste even in the base model, and we know Qwen 3.8 is going to blow us all away! Regardless, this looks like a very tempting new base model. As a first offering from Meta in this class for a long time, I am incredibly impressed and elated to have it. We now finally have a proper Single GPU frontier race, instead of us just begging Qwen for more releases. Single GPU open frontier model race is a VERY good thing. Please keep pushing Meta

Kyle Hessling

17,485 görüntüleme • 1 ay önce

Daha fazla içerik yok.