David Hendrickson's banner
David Hendrickson's profile picture

David Hendrickson

@TeksEdge10,143 subscribers

CEO & Founder | PhD | Startup Advisor | @Columbia | Author Generative Software Engineering https://t.co/9oqvHuTX5f | 🔔 Follow for AI & Vibe Coding Tips 👇

Shorts

🚀 AMD Ryzen AI Halo is now available for pre-order! A compact local AI developer platform powered by the Ryzen AI Max+ 395: 🧠 128GB unified LPDDR5x memory ⚡ 40 CU Radeon 8060S graphics (RDNA 3.5) 📦 Run models up to 200B parameters locally 🖥️ Windows + Linux support out of the box Build and deploy AI workflows without cloud dependency. Pre-order → @ amd

🚀 AMD Ryzen AI Halo is now available for pre-order! A compact local AI developer platform powered by the Ryzen AI Max+ 395: 🧠 128GB unified LPDDR5x memory ⚡ 40 CU Radeon 8060S graphics (RDNA 3.5) 📦 Run models up to 200B parameters locally 🖥️ Windows + Linux support out of the box Build and deploy AI workflows without cloud dependency. Pre-order → @ amd

100,972 views

Videos

TeksEdge's profile picture

🤯 A localmaxxer hit ~381 tok/s on a SINGLE RTX 3090 with Qwen3.8-27B. This developer has turned a 24GB RTX 3090 into a monster Qwen inference box - w/some creativity. Four days ago 👉 ⚡ ~82 tok/s single-user Then 👉 ⚡ ~114 tok/s with optimized MTP ⚡ ~138 tok/s with DFlash2 + lookup drafting Now 👉 🔥 ~381 tok/s on ONE request How? The recipe combines ... 🧠 Qwen3.8-27B 🎮 1× RTX 3090 24GB @ 250W ⚙️ heavily optimized vLLM ⚡ DFlash2 speculative decoding 🔎 lookup-augmented drafting 📚 prefix caching 🧮 16-token verification blocks 💾 quantized KV / heads / activations DFlash2 normally proposes 7 tokens. The developer realized the verification block doesn't have to stop there. If Qwen is answering from a document already sitting in the prompt, the system can fill the remaining draft positions using tokens found directly in that context. 🎯 So the target model can verify 16 tokens at once. On a ~25K-token document reproduction task: Previous DFlash2 👉 ~260 tok/s Longer verification + context lookup 👉 🔥 ~382 tok/s Acceptance: 🤯 15 of 16 tokens per verification step That is where the crazy number comes from. ⚠️ On ordinary real-world chat prompts, the same setup is around ~133 tok/s Still extremely fast for a dense 27B model on an RTX 3090. The 381 tok/s mode shines when the answer largely comes from material already in context so these are best use cases 📚 RAG / document Q&A 💻 Coding assistants applying edits 📝 Quoting or rewriting documents 🔎 Extracting information from long prompts And another optimization 👉 With prefix caching, a second question against the same 25K-token document reportedly goes from: 🐌 22.4 sec TTFT → ⚡ 0.56 sec TTFT Because the model doesn't need to process the whole document again. 🎯 It's specifically a mode for RAG front ends and coding agents. Follow iamMess on Reddit or syv-ai on GitHub 🔗 Reddit: r/LocalLLaMA/comments/1vtup5s/ 🔗 GitHub: /syv-ai/qwen38-27b-rtx3090

David Hendrickson

105,418 views • 4 days ago

No more content to load