Video yükleniyor...

Video Yüklenemedi

Ana Sayfaya Dön

"which quant should I download?" is a question you may never have to answer again the team Hamster Labs has figured out how to kill it with pMLX. download once at full precision (bf16) and the engine re-fits it to your machine on the fly, based on the job...

10,590 görüntüleme • 26 gün önce •via X (Twitter)

31 Yorum

雪瑜 profil fotoğrafı
雪瑜23 gün önce

@HamsterResearch I'm really looking forward to when pMLX will be released constantly having to calculate memory usage myself is a real pain.

Felix Hesse profil fotoğrafı
Felix Hesse26 gün önce

@HamsterResearch Unbelievable! If you can do that, it's nothing less than the revolution of local AI. Then almost everyone can have their own frontirr at home. A huge potential. When and where can I download pMLX?

Eyal Toledano profil fotoğrafı
Eyal Toledano25 gün önce

@HamsterResearch Not yet released, it’s working fully locally but I need to optimize decode speed a bit before releasing it My goal is q4 at 75% resident decoding at 100+ tok/s Hovering around mid 60’s right now which is 50%+ faster than the current release anyway Let me cook

Felix Hesse profil fotoğrafı
Felix Hesse25 gün önce

@HamsterResearch Sounds great! What do you predict for a mac Studio M2Max 64gb?

Ilya Kaminsky profil fotoğrafı
Ilya Kaminsky25 gün önce

@HamsterResearch This tweet mentions an "engine" that requires two inputs -- desired speed and context length. Where do I get it and how do I run it? Or is everyone waiting for you to publish

Eyal Toledano profil fotoğrafı
Eyal Toledano25 gün önce

@HamsterResearch yep it’s not yet released

Ilya Kaminsky profil fotoğrafı
Ilya Kaminsky25 gün önce

@HamsterResearch Can you hook a brother up with some early access? I need it[1] for Geer[2] to cover all Apple Silicon from 16GB and up. [1] [2]

Eyal Toledano profil fotoğrafı
Eyal Toledano25 gün önce

@HamsterResearch I’m nearly there, will push a first version even if tos benchmark doesn’t pass as it can be improved from there

techntrails profil fotoğrafı
techntrails25 gün önce

@HamsterResearch @EyalToledano please update on status App is down, website is down, email support is down. Lots of small biz owners in limbo that don't know how to get in touch with you

Eyal Toledano profil fotoğrafı
Eyal Toledano24 gün önce

@HamsterResearch 👀 mind dm’ing me?

Doug Hohner profil fotoğrafı
Doug Hohner24 gün önce

@HamsterResearch I assume this will eventually work with larger models like DSv4Flash or GLM 5.3 Flash as well? Very interesting... Trying to decide just how big a Mac to buy, will have to watch your project for a few weeks...

Eyal Toledano profil fotoğrafı
Eyal Toledano24 gün önce

@HamsterResearch gln 5.3 flash is a target right after the qwen3.8 flash next model it’s already loading on the engine at about 55gb resident but the decode is dreadful (but fixable) so odds are contributors could have at it or i’ll get to it next i can’t wait

David King profil fotoğrafı
David King26 gün önce

@HamsterResearch 🤯

Eyal Toledano profil fotoğrafı
Eyal Toledano25 gün önce

@HamsterResearch

Ethpunk profil fotoğrafı
Ethpunk25 gün önce

@HamsterResearch Brilliant

Chris W profil fotoğrafı
Chris W26 gün önce

@HamsterResearch I love that - it is helpful especially for newbies in the game like me.

Afterkind profil fotoğrafı
Afterkind26 gün önce

@HamsterResearch This doesn't work for image models, right?

Eyal Toledano profil fotoğrafı
Eyal Toledano26 gün önce

@HamsterResearch sure it does, qwen3.8-flash-next is multimodal and the vision tower’s unaffected

EddyLeeKhane profil fotoğrafı
EddyLeeKhane25 gün önce

@HamsterResearch is it loading another already downloaded quant on the fly or somehow applying quants in real time???

Eyal Toledano profil fotoğrafı
Eyal Toledano25 gün önce

@HamsterResearch Quantizing on the fly :)

EddyLeeKhane profil fotoğrafı
EddyLeeKhane25 gün önce

@HamsterResearch now if that works in reql time the holy cow dude your team of mad lads are off the charts!!!!!

Eyal Toledano profil fotoğrafı
Eyal Toledano25 gün önce

@HamsterResearch Right now it’s a config flag, meaning server would restart and it would add a bit of latency But I think real time is possible and worth reaching for

Laurent Zuijdwijk profil fotoğrafı
Laurent Zuijdwijk26 gün önce

@HamsterResearch Do you know what disk space costs nowadays?

Eyal Toledano profil fotoğrafı
Eyal Toledano25 gün önce

@HamsterResearch compared to ram though, it’s night and day

SettledScience™️ profil fotoğrafı
SettledScience™️24 gün önce

@laurent_zw @HamsterResearch Any thoughts on ‘fast enough’ storage to not get in the way? I mean, TBs of latest internal SSD on board is great, but if streaming the experts from external drive is good enough, that helps save some $. TB5? TB4? USB3? What is ‘fast enough’?

Eyal Toledano profil fotoğrafı
Eyal Toledano24 gün önce

right now the expectation is the ssd is the nvme that ships on the macbook, which has plenty of bandwidth but does differ across different mac models right now i’m building and testing on an m4 max which has very good ssd bandwidth. the streaming is math exact so only thing limiting decode is hardware and the engine is obviously very young and i’m the only contributor on it for now. it will improve way faster once the community swarms on it (making the tooling to make that possible)

SettledScience™️ profil fotoğrafı
SettledScience™️24 gün önce

@laurent_zw @HamsterResearch Correct to assume the reads are more random than sustained… IOPS more important than sustained read and latencies with external are the big question mark?

Eyal Toledano profil fotoğrafı
Eyal Toledano24 gün önce

yep it’s random, not sustained, so latency matters more than bandwidth but its large-block random each expert is a few mb, not a 4k scattered read it prefetches the next token’s experts and keeps the hot ones resident so most reads are already there. external adds a protocol hop which is why nvme is baseline. so thunderbolt can prob keep up but usb is spicy at least that’s how i understand it

taisei profil fotoğrafı
taisei25 gün önce

@HamsterResearch

Tony Tong | Founder | Ancient Systems x AI profil fotoğrafı
Tony Tong | Founder | Ancient Systems x AI24 gün önce

@HamsterResearch Auto-fitting a model to your hardware beats agonizing over which quant to grab. I sat with a similar kind of indecision last year, weighing whether to leave Singapore for a cheaper cost of living elsewhere, re-running the comparison long after the answer was obvious.

Dr. Cannoli profil fotoğrafı
Dr. Cannoli25 gün önce

@HamsterResearch Wouldn’t you run into naive quantization issues? Not optimized?

Benzer Videolar

i spent 3 hours finding the sweet spot for hermes 4.3 36B on a single RTX 3090. saving you the trouble, anon. the model is 21.8GB at Q4_K_M. that leaves 2.2GB free on 24GB VRAM. not much room for KV cache. here's what actually happened: 4K: 35.3 tok/s 8K: 35.2 tok/s 16K: 34.8 tok/s 32K: 34.6 tok/s 64K: 6.4 tok/s 128K: 1.9 tok/s flat from 4K to 32K. then it falls off a cliff at 64K. the trick is quantized KV cache. without it you OOM at 16K. with quantized KV cache you get 32K at full speed. all 64 layers on GPU. at 64K something weird happens. ngl 99 (all layers on GPU) = 3.96 tok/s. the KV cache silently spills to CPU. drop to ngl 55 and speed jumps to 6.37. drop to ngl 48 and it gets worse again (3.46). there's an offload sweet spot where you free just enough VRAM for the cache without losing too much compute to PCIe transfers. 128K works at ngl 32 but you're at 1.95 tok/s. half the model on CPU. usable for batch work, not for interactive. the sweet spot command: llama-server -m hermes-4.3-36b-Q4_K_M.gguf -ngl 99 -c 32768 --cache-type-k q4_0 --cache-type-v q4_0 32K context. 34.6 tok/s. all on GPU. this is where dense 36B lives on 24GB. for comparison, qwen 3.5 (35B MoE, 3B active) holds 112 tok/s from 4K all the way to 262K on the same GPU. no speed drop. same total params, completely different architecture. hybrid linear attention means flat context scaling. dense pays for every token in the KV cache. code and quality comparison coming next. fast vs slow generation side by side in the videos below.

Sudo su

52,268 görüntüleme • 6 ay önce

Qwen3.8-27B running at full BF16 on a free Kaggle TPU is kind of ridiculous. No quantization. No tiny context window. No expensive GPU instance. Just Qwen3.8-27B running on a Kaggle TPU v5e-8. The reported numbers: ~130 tok/s decode ~10,000 tok/s prefill 262K context That prefill number is especially wild. You can throw a huge amount of code or context at the model and ingest it extremely quickly, while still getting around 130 tokens per second during generation. And because it’s running in full BF16, you’re not relying on an aggressive quant just to make the model fit. But the really interesting part isn’t even the raw throughput. You can expose it as an OpenAI-compatible endpoint. That means you can plug the model into tools that already understand OpenAI-style APIs. Claude Code. Codex. OpenCode. And other compatible clients. So the workflow becomes pretty simple: Spin up the Qwen3.8-27B endpoint on Kaggle. Point your coding tool at the API. And suddenly you have a 27B coding model sitting behind the same interface you’d normally use for hosted models. The 262K context is also a huge deal for agentic coding. Large repositories can fit into a single context. Long conversations don’t need to be constantly trimmed. And tools can feed much more information back to the model without hitting a tiny context ceiling. The fact that this can be built around a free TPU environment is what makes this especially interesting. We’re getting to a point where experimenting with serious open models doesn’t always require owning a $2,000 GPU or paying for a large cloud instance. Free compute + open weights + an OpenAI-compatible API + existing coding agents. That’s a pretty powerful combination. Qwen3.8-27B is already an interesting model. Running the full BF16 version at ~130 tok/s with 262K context on free Kaggle TPU compute makes it a lot more interesting.

FHILY👑

35,794 görüntüleme • 19 gün önce

HOW TO DRIFT IN FORZA HORIZON 6 ON CONTROLLER FOR BEGINNERS! Disclaimer, I’m not exactly Drift King, some people make me look like Donkey Kong. I do it well enough for people who are completely new to drifting to ask me how I do it, so I’m going to explain how I do it, as someone who is “decent” at best at it. First of all, you’re going to want to change your transmission to Manual so that you can maintain a consistent drift, as the speed of your vehicle determines the angle of your drift, and with Automatic, you can’t control your speed. It doesn’t matter, just personal preference here, you can do your gear changes with the buttons, but if you own an Elite Controller, you can put the paddles to good use here. I like to equip only two paddles, preferably the larger ones, and map the left to X (Downshift) and map the right to B (Upshift). You also want to turn off your Traction and Stability control assists to allow your car to slide. My vehicle of preference is the Silvia Spec R, and the code for the tune I use is 146 936 201. I am using a RWD Vehicle for reference, but some people prefer AWD. For the most part, you’re going to be driving in third and fourth gear. Third gear is great for small roads and tight, closed corners where you need a sharper drift, and fourth gear is great for wide roads and corners because the extra speed is going to throw your car further forwards and make your drift angle more shallow. It’s important to look at the size of the road and the corner you have to take and judge what gear you need to be in as the speed of your car makes all of the difference. That being said, these are not real cars, if you’re going too fast, I much prefer to hold the gas and brake pedal at the same time whilst taking the drift as opposed to engine braking. It almost makes you float around the corner at a consistent speed and gives you more control. If you are on the inside edge of the corner you’re taking, you can hold your foot brake and handbrake at the same time and your car will almost push back a little giving you more space to angle yourself correctly. If this helped you, let me know 🥹

Moki

201,342 görüntüleme • 4 ay önce

hey here is the final result of octopus invaders on nvidia's flagship at full precision. nemotron super 120B on 2x H200 NVL. BF16 unquantized. 287GB of VRAM. hermes agent as the harness. 60 tok/s. first try it autonomously coded for 6 minutes straight. created 11 files. correct project structure. correct load order. started the server. i opened the browser and the result was a blank screen. i did not give up. second try i gave it a precise list of bugs and things to fix. it went back in for another 3 minutes. patched the code. served it again. still blank. so i did what any sane person would do. third try i just said the screen is blank, test it and fix it yourself. and this is where nemotron showed what it actually is. it became a debugger. you can see it in the video. realtime CSS test squares, red screen flashes, hermes agent browser tools, inspecting its own output. it built the parallax background with planets and comets. it rendered a rocket ship that tracks your mouse with fire and bullet physics. the aesthetic is real. but no enemies spawn. no collision. not playable. what surprised me is qwen 27B one shotted this exact game on a single RTX 3090 at Q4 quant. and here is nvidia's flagship at full precision on enterprise hardware needing 3 tries and still not getting there. that makes my hope high for the undisputed qwen 122B which is about to face the same test next. same hardware. same prompt and same harness. lets see if it one shots or not. full session in the video. no cuts. 5x speed.

Sudo su

11,011 görüntüleme • 5 ay önce

How can you solve complex tasks using a Large Language Model? Here is a 2-minute introduction to everything you need to know to 10x the quality of your results. Let's talk about three techniques, in order of complexity, starting with the easiest one: • In-Context Learning • Indexing + In-Context Learning • Fine-tuning In-Context Learning The team that trained GPT-3 found something they couldn't explain: You can condition a model using examples of how you want it to behave. I included an example prompt in the attached video. You can "teach" the model how you want it to interpret questions, select the correct answers, and format the results by giving a few examples. You can also give specific knowledge to the model that will be helpful when formulating answers. We call this approach "grounding the model." There's another example in the video. Indexing + In-Context Learning Unfortunately, there is a limit to how much data you can include in a prompt. We call this the "context size." One version of GPT-4 supports a context of approximately 6,000 words, while the other supports 25,000 words. Although this sounds like a lot, many applications need more than that. Imagine you wrote a book and want to build an application to answer any questions about your story. What happens if your book is longer than the context? That's where Indexing comes in. Using a model, you can turn every book passage into an embedding. These are vectors, numbers that "encode" the passage's text. You can then store these embeddings in a particular database that supports fast retrieval of these vectors. You can then turn any question into an embedding and search the database for the list of passages that are similar to that query. Instead of using the entire book to ask the model, you can now use the relevant passages as in-context information, effectively working around the context size limitation. Fine-tuning Fine-tuning can give you an extra boost to get reliable outputs from your LLM. It is, however, the most complex approach on the list. There are different approaches to fine-tuning a model with your data. A popular technique is to process your data with your LLM and use the outputs to train a new classifier that solves your specific task. Notice that here you aren't modifying the LLM. Instead, you are chaining it with your trained classifier. Another approach is to modify the parameters of the LLM using your data. Think of this as "rewiring" the model in a way that solves your particular task. The results and costs will vary depending on how many layers you want to fine-tune from the original model. Many companies think that fine-tuning is the solution to their problems. In my experience, many will benefit from exploring the other two approaches. I love explaining Machine Learning and Artificial Intelligence ideas. If you enjoy in-depth content like this, follow me Santiago so you don't miss what comes next.

Santiago

384,510 görüntüleme • 3 yıl önce