Video yükleniyor...
Video Yüklenemedi
"which quant should I download?" is a question you may never have to answer again the team Hamster Labs has figured out how to kill it with pMLX. download once at full precision (bf16) and the engine re-fits it to your machine on the fly, based on the job... show more
10,590 görüntüleme • 26 gün önce •via X (Twitter)
31 Yorum

@HamsterResearch I'm really looking forward to when pMLX will be released constantly having to calculate memory usage myself is a real pain.

@HamsterResearch Unbelievable! If you can do that, it's nothing less than the revolution of local AI. Then almost everyone can have their own frontirr at home. A huge potential. When and where can I download pMLX?

@HamsterResearch Not yet released, it’s working fully locally but I need to optimize decode speed a bit before releasing it My goal is q4 at 75% resident decoding at 100+ tok/s Hovering around mid 60’s right now which is 50%+ faster than the current release anyway Let me cook

@HamsterResearch Sounds great! What do you predict for a mac Studio M2Max 64gb?

@HamsterResearch This tweet mentions an "engine" that requires two inputs -- desired speed and context length. Where do I get it and how do I run it? Or is everyone waiting for you to publish

@HamsterResearch yep it’s not yet released

@HamsterResearch Can you hook a brother up with some early access? I need it[1] for Geer[2] to cover all Apple Silicon from 16GB and up. [1] [2]

@HamsterResearch I’m nearly there, will push a first version even if tos benchmark doesn’t pass as it can be improved from there

@HamsterResearch @EyalToledano please update on status App is down, website is down, email support is down. Lots of small biz owners in limbo that don't know how to get in touch with you

@HamsterResearch 👀 mind dm’ing me?

@HamsterResearch I assume this will eventually work with larger models like DSv4Flash or GLM 5.3 Flash as well? Very interesting... Trying to decide just how big a Mac to buy, will have to watch your project for a few weeks...

@HamsterResearch gln 5.3 flash is a target right after the qwen3.8 flash next model it’s already loading on the engine at about 55gb resident but the decode is dreadful (but fixable) so odds are contributors could have at it or i’ll get to it next i can’t wait

@HamsterResearch 🤯

@HamsterResearch

@HamsterResearch Brilliant

@HamsterResearch I love that - it is helpful especially for newbies in the game like me.

@HamsterResearch This doesn't work for image models, right?

@HamsterResearch sure it does, qwen3.8-flash-next is multimodal and the vision tower’s unaffected

@HamsterResearch is it loading another already downloaded quant on the fly or somehow applying quants in real time???

@HamsterResearch Quantizing on the fly :)

@HamsterResearch now if that works in reql time the holy cow dude your team of mad lads are off the charts!!!!!

@HamsterResearch Right now it’s a config flag, meaning server would restart and it would add a bit of latency But I think real time is possible and worth reaching for

@HamsterResearch Do you know what disk space costs nowadays?

@HamsterResearch compared to ram though, it’s night and day

@laurent_zw @HamsterResearch Any thoughts on ‘fast enough’ storage to not get in the way? I mean, TBs of latest internal SSD on board is great, but if streaming the experts from external drive is good enough, that helps save some $. TB5? TB4? USB3? What is ‘fast enough’?

right now the expectation is the ssd is the nvme that ships on the macbook, which has plenty of bandwidth but does differ across different mac models right now i’m building and testing on an m4 max which has very good ssd bandwidth. the streaming is math exact so only thing limiting decode is hardware and the engine is obviously very young and i’m the only contributor on it for now. it will improve way faster once the community swarms on it (making the tooling to make that possible)

@laurent_zw @HamsterResearch Correct to assume the reads are more random than sustained… IOPS more important than sustained read and latencies with external are the big question mark?

yep it’s random, not sustained, so latency matters more than bandwidth but its large-block random each expert is a few mb, not a 4k scattered read it prefetches the next token’s experts and keeps the hot ones resident so most reads are already there. external adds a protocol hop which is why nvme is baseline. so thunderbolt can prob keep up but usb is spicy at least that’s how i understand it

@HamsterResearch

@HamsterResearch Auto-fitting a model to your hardware beats agonizing over which quant to grab. I sat with a similar kind of indecision last year, weighing whether to leave Singapore for a cheaper cost of living elsewhere, re-running the comparison long after the answer was obvious.

@HamsterResearch Wouldn’t you run into naive quantization issues? Not optimized?
