Killy's banner
Killy's profile picture

Killy

@net_termina • 3,414 subscribers

Hard Work and Guts

Shorts

As promised, instead of for $15,000 GPUs, we max optimized a model that can fit comfortably in any GPU with 8GB VRAM running Windows 200 tok/s, *128k context, text/image/audio. Anyone who wants to join the local AI community with 5 year old laptop can run this with recipe below, and this test was done on 3070ti laptop I forgot I had. We pulled every instruct model released since May that fits 8 or 12 GB, and ran the six that mattered through the same 20 tasks: JSON extraction, code, a table, arithmetic, a fake-treaty honesty check, tool calls, language, a persona, a fact planted deep in a long prompt, and three photos. Gemma 4 E4B, Gemma 4 12B, Qwen3.5-9B, Qwen3.5-4B, Ornith-1.5-9B, the Qwen3.8-9B distill. E4B is not the top scorer. Qwen3.5-9B and Ornith beat it on long context and on reading fine print in a photo. It won the pick for the things a newcomer on an old card actually needs: Apache 2.0, text + image + audio in, 3 GB of headroom on an 8 GB card, Weights Google trained to be 4-bit (QAT), and Google's own speculative drafter shipped next to them. Nothing else in the tier has all five, and the drafter alone is worth 2–3×.

As promised, instead of for $15,000 GPUs, we max optimized a model that can fit comfortably in any GPU with 8GB VRAM running Windows 200 tok/s, *128k context, text/image/audio. Anyone who wants to join the local AI community with 5 year old laptop can run this with recipe below, and this test was done on 3070ti laptop I forgot I had. We pulled every instruct model released since May that fits 8 or 12 GB, and ran the six that mattered through the same 20 tasks: JSON extraction, code, a table, arithmetic, a fake-treaty honesty check, tool calls, language, a persona, a fact planted deep in a long prompt, and three photos. Gemma 4 E4B, Gemma 4 12B, Qwen3.5-9B, Qwen3.5-4B, Ornith-1.5-9B, the Qwen3.8-9B distill. E4B is not the top scorer. Qwen3.5-9B and Ornith beat it on long context and on reading fine print in a photo. It won the pick for the things a newcomer on an old card actually needs: Apache 2.0, text + image + audio in, 3 GB of headroom on an 8 GB card, Weights Google trained to be 4-bit (QAT), and Google's own speculative drafter shipped next to them. Nothing else in the tier has all five, and the drafter alone is worth 2–3×.

127,634 Aufrufe

Okay, I guess I didn't waste all that money after all Deepseek V4.1 Flash 304 tok/s code, 135 tok/s prose Zero optimization, will have it max optimized by the morning.

Okay, I guess I didn't waste all that money after all Deepseek V4.1 Flash 304 tok/s code, 135 tok/s prose Zero optimization, will have it max optimized by the morning.

44,653 Aufrufe

454 tok/s on Qwen3.8 27B on single RTX 6000 with +6000 mem overclock Mia’s Qwen3.8-27B NVFP4 W4A4 • SGLang + DFlash2 block-16 drafter • FP8 KV

454 tok/s on Qwen3.8 27B on single RTX 6000 with +6000 mem overclock Mia’s Qwen3.8-27B NVFP4 W4A4 • SGLang + DFlash2 block-16 drafter • FP8 KV

50,467 Aufrufe

What the hell does Quant actually do? A model is trained in BF16, released in Q8, and we use Q4 because we're told 'its nearly lossless'. Fine, but I was curious about the 'nearly' part, what is actually being lost between BF16 all the way down to Q2? As always, its complicated. Needed a way to visualize the loss, so I had a Qwen 3.8 27B BF16 create a paragraph then had every quant below it do the same and marked the mutations. When a model writes a word, it assigns 'how sure' percentage to each. For example, 'The weather tomorrow is *windy*, a model is 90% sure all the words are correct except the windy part, since it can also be sunny, cold, hot, cloudy, etc. Any words a BF16 model was 90% sure of, every quant down to Q2 rarely changes it. On the other hand, any word that it was only 40~50% sure of, such as windy in our example, are the types of data that gets impacted by the quant process. Okay so what, some words changed, how does this impact anything? I set out to find out, and its all below if you are just as weird as I am about these things.

What the hell does Quant actually do? A model is trained in BF16, released in Q8, and we use Q4 because we're told 'its nearly lossless'. Fine, but I was curious about the 'nearly' part, what is actually being lost between BF16 all the way down to Q2? As always, its complicated. Needed a way to visualize the loss, so I had a Qwen 3.8 27B BF16 create a paragraph then had every quant below it do the same and marked the mutations. When a model writes a word, it assigns 'how sure' percentage to each. For example, 'The weather tomorrow is *windy*, a model is 90% sure all the words are correct except the windy part, since it can also be sunny, cold, hot, cloudy, etc. Any words a BF16 model was 90% sure of, every quant down to Q2 rarely changes it. On the other hand, any word that it was only 40~50% sure of, such as windy in our example, are the types of data that gets impacted by the quant process. Okay so what, some words changed, how does this impact anything? I set out to find out, and its all below if you are just as weird as I am about these things.

19,592 Aufrufe

Hey gang, heres the full build as promised. First of all, nothing went as planned, demo didn't work, and I got poop all over my keyboard when baby's diaper leaked when I was putting all this together so this is going to be it. Part list is below in the link.

Hey gang, heres the full build as promised. First of all, nothing went as planned, demo didn't work, and I got poop all over my keyboard when baby's diaper leaked when I was putting all this together so this is going to be it. Part list is below in the link.

17,815 Aufrufe

Videos

Keine weiteren Inhalte verfügbar