My dual RTX PRO 6000 setup is currently training... a Draft model for Qwen 3.6 27B! 🔥 I'm taking the paper DeepSeek dropped on 6/26 and going for a super ambitious application to the 27B scale. Thanks to my homelab, I was able to dive straight in — I read the paper and immediately started experimenting. The amount I've learned has been insane: - How memory bandwidth bottlenecks speed and clever ways to hack around it - Methods to train the draft model and boost its accuracy - Mechanisms to reference tokens all the way back to the previous one to skyrocket draft acceptance rates - The impact of Attention vs. GateDeltaNet on speculative decoding performance and how to handle those differences - The unique approaches and trade-offs of MTP, Dflash, JetSpec, and DSpark I could go on forever, but just from speculative decoding alone I've learned so much. The 27B architecture feels way more DSpark-native than JetSpec, so once draft training finishes, I'm going all-in with DSpark! My goal is to beat existing speculative decoding speeds outright — no task-specific shortcuts or cheating, pure general improvement. If you're into this kind of research, I'd love to hear your thoughts, impressions, and any suggestions — please reply! 🚀show more

Hikari∣LocalLLM⚡
56,870 Aufrufe • vor 2 Monaten
we sped up distributed inference by up to 5x... with decentralized speculative decoding. many don't realize that AI models normally generate text one single word at a time, waiting for the network after every word. speculative decoding changes this by using a "guess & confirm" system, similar to autocomplete. how it's done: 1. draft locally (the guess) instead of waiting for the network, a tiny, fast model on your device guesses the next few words instantly, without waiting for the network. 2. confirm remotely (the check) the massive remote model doesn't generate from scratch; it just checks the draft. it looks at the guesses in a batch and says "yes, yes, no." you get multiple words in the time it usually takes to get one. 3. adaptive logic dsd is smart. if the topic is creative, it lets the draft flow loose. if the topic is math or code, it checks more strictly. it balances speed and precision automatically so your inference almost feel instant. find out more: paper: blog:show more

Parallax
45,584 Aufrufe • vor 7 Monaten
Qwen 3.8 27B Q4_K_M - 90 tokens/sec on a... single NVIDIA RTX 4090 (24 GB VRAM) with Dflash2! (MTP 60 tps -> 90 tps Dflash2!!!!) Local AI moves so fast (literally!) it’s terrifying. Z lab just dropped DFlash 2 for Qwen 3.8 27b and Muse Glimmer. I patched llama.cpp (PR #27342) and paired it with Unsloth’s Qwen 3.8 27B UD-Q4_K_XL quant. The result? Lossless 90 tokens/s decode. My last post highlighted native MTP hitting 60 t/s at 130,000 context. But DFlash 2 just completely shattered that ceiling. By using parallel block diffusion drafting (predicting whole blocks of tokens in a single pass using dynamic convolutions), DFlash achieves a massive 5.39 token acceptance rate. THE ALPHA TWEAK: `n-max 7` eats too much VRAM for draft states. But if you drop the draft limit to `--spec-draft-n-max 4`, you slash the VRAM overhead and actually increase the throughput. Here is the new 24GB VRAM Physics Matrix (DFlash 2 @ n-max 4): - 30k Context: 1,725 t/s prefill | 87.05 t/s decode | 22.2 GB VRAM - 80k Context: 1,789 t/s prefill | 84.20 t/s decode | 23.3 GB VRAM - 110k Context: 1,767 t/s prefill | 83.35 t/s decode | 23.96 GB VRAM (110k context at 83+ tokens a second sitting exactly on the 24GB hardware limit is absolute wizardry). How to compile the PR today: git clone cd llama.cpp git fetch origin pull/27342/head:pr-27342 git switch pr-27342 cmake -B build -DGGML_CUDA=ON && cmake --build build -j Llama.cpp flags for Dflash (110k Context Ceiling): ./build/bin/llama-server -m Qwen3.8-27B-UD-Q4_K_XL.gguf -md Qwen3.8-27B-DFlash2-Q4_K_M.gguf --spec-type draft-dflash --spec-draft-n-max 4 -c 110000 -ngl 99 --port 8080 -ctv q4_0 -ctk q4_0 The fact that the open source community is shipping block diffusion drafters so quickly that run entirely locally on a gaming GPU is unbelievable. If you own a single RTX 3090 or 4090, it is officially time to upgrade to qwen 3.8 27b with dflash 2 and cancel your API subscriptions and let local silicon eat the cloud. This model beats GPT 5.6 Terra, GLM 5.2 DeepSeek V4 Pro, Muse Spark 1.2 and Claude Opus 4.8 on the artificial analysis agentic index (details in the replies) Hugging Face GGUF links (Base + DFlash2) and the full visual VRAM scaling and Dflash2 vs MTP graphs are also in the replies below. are you sticking to native MTP for the 130k context, or sacrificing 20k context to redline your decode speed? How many tokens/sec are you pushing on your current local rig?show more

Alok
103,895 Aufrufe • vor 10 Tagen
MTP speedup Qwen by 2.5x in Atomic Chat Dense... vs MoE models on 2x RTX 5090 Qwen3.6 27B: 51 → 117 tps +137% Qwen3.6 35B-A3B: 218 → 267 tps +25% MTP drafts several tokens ahead and verifies them in one pass. The speedup depends on memory moved per pass. Dense 27B reads all 27B params per token, MoE 35B-A3B only reads 3B active. Dense had way more to save by batching. The baseline tps also differ (218 vs 51) for the same reason from the other side. Token generation is memory-bandwidth bound, and MoE moves ~8x less memory per token, so its baseline is already 4x ahead. ~80% draft acceptance. Zero accuracy loss. ~1 GB extra VRAM. Open-source code and local AI app – in the comments 👇show more

atomic.chat
171,011 Aufrufe • vor 3 Monaten
Added context to my tiny diffusion model to enable... sequential generation of longer outputs! Currently the context is a quarter of the sequence length (seq_len=256, context_len=64). I have a theory that the less semantic-value-per-token, the worse the “curse of parallel decoding” is. With parallel decoding, we independently predict multiple tokens in one step. With the sentence “My poker hand was a ___ ___”, two valid predictions are “two pair” and “straight flush”. Because each token prediction is independent though, we can end up with a nonsensical output like “two flush”. This seems to be exacerbated with low semantic-value-per-token, as now you need more tokens to express the same concept. Instead of needing to independently predict two tokens, we might need to predict 10 instead (which is of course much harder). The model currently has noticeably worse output compared to nanogpt (similar size) and I believe this is a main reason. I’ll try adding confidence-aware parallel decoding (from NVIDIA’s Fast-dLLM paper) and other tricks and see how much they improve generation quality.show more

Nathan Barry
89,040 Aufrufe • vor 10 Monaten
I was informed that the team will be seeking... a trade immediately and will be working with me and my family to find the right place to continue competing for championships. I don’t agree with the decision and always believed it was going to begin and end in LA. Still, if there’s one thing that I have learned over the years: there are so many things that are out of your control, but it is how you respond to these things that you will look back on and remember. I have taken so much pride in playing alongside my teammates for the LA community, so thank you for embracing my family and making this such a special place for us. 2024 began with one of the best training camps of my career. Preparations start now for 2025. Highly motivated, as healthy as ever, and looking forward to playing elite football for years to come. Love you guys.. But coming for it all.show more

Cooper Kupp
19,594,872 Aufrufe • vor 1 Jahr
UPDATE: 🚨 +$165,000 ON THE DAY THANK YOU GOD... 🙏 A little backstory... I'm currently going through a divorce. It's been one of the toughest seasons of my life mentally and emotionally. Through it all, I've never lost sight of my goals, my dreams, or my faith. For some reason, things have continued to fall into place when I least expected it. I'm also in the process of buying a home for myself and my kids. Not once did I think any of this would be possible on my own. This economy has been tough, but by the grace of God, I've been able to keep moving forward. To anyone who feels behind, lost, or unsure of how they're going to make it through: GOD'S GOT YOU. He will make a way. Stay patient. Stay faithful. Keep going. HE LOVES YOU. ❤️🙏 And NEVER GIVE UP!show more

Sergio Solis
38,094 Aufrufe • vor 2 Monaten
I told you to claim your free 16GB NVIDIA... GPU for learning Local LLMs. Now I’m going to show you how to double its inference speed without touching the hardware. Google Colab gives you an enterprise grade NVIDIA Tesla T4 GPU for free, roughly 4 hours every single day. It is the absolute perfect sandbox for learning AI engineering, testing inference flags, and pushing massive context windows. The local AI timeline is moving way too fast. If you aren't using Multi Token Prediction (MTP) yet, you are leaving massive performance on the table. I just pushed DeepMind’s Gemma 4 26B to 64.9 t/s on this exact free tier. Let's look at the raw benchmark data running on an Ubuntu Linux environment with the latest compiled llama.cpp binaries and quantized GGUFs from Unsloth via HuggingFace: # Qwen 3.5 9B (Dense): Base: [ Prompt: 626.7 t/s | Generation: 21.0 t/s ] With MTP: [ Prompt: 539.1 t/s | Generation: 24.8 t/s ] # Gemma 4 26B QAT (MoE): Base: [ Prompt: 634.2 t/s | Generation: 48.3 t/s ] With MTP: [ Prompt: 572.1 t/s | Generation: 64.9 t/s ] If you are paying attention, this single Colab notebook reveals 3 massive observations about the current state of local LLMs: # 1. The MTP Speedup (Software Overclocking) Standard autoregressive decoding guesses one token at a time. MTP acts like a highly optimized, built in speculative decoder. It predicts multiple future tokens at once and the main model verifies them in parallel. The result? Zero accuracy loss and a massive throughput increase. Gemma jumped from 48 to 65 t/s just by flipping a flag. # 2. The MoE Paradox (Bigger is Faster) How does a 26B parameter model absolutely destroy a 9B model in raw speed on the exact same hardware? Architecture. Qwen 3.5 9B is a dense model. it activates all 9 billion parameters for every single token. Gemma 4 26B is a Mixture of Experts (MoE) model. It routes data efficiently, activating only 4B parameters per token. You get the reasoning capabilities of a 26B model with the compute cost of a 4B model. 3. Thinking Efficiency When I ran the exact same complex prompt on both models, the larger MoE spent significantly fewer "thinking" tokens to arrive at the correct answer. A smarter model doesn't just give better answers; it gets to the point faster, saving you compute cycles and preserving your context window. # Want to run this yourself? Here are the exact llama.cpp CLI commands. For Qwen (MTP is baked into the main model): ./llama-cli -m Qwen3.5-9B-UD-Q4_K_XL.gguf -p "Explain quantum computing." -n 2000 -c 8000 -ngl 99 -fa on --spec-type draft-mtp --spec-draft-n-max 4 --spec-draft-p-min 0.7 For Gemma (Using a separate lightweight draft model): ./llama-cli -m gemma-4-26B-A4B-it-qat-UD-Q4_K_XL.gguf --model-draft mtp-gemma-4-26B-A4B-it.gguf -p "Explain quantum computing." -n 2000 -c 8000 -ngl 99 -fa on --spec-type draft-mtp --spec-draft-n-max 4 --spec-draft-p-min 0.7 Stop waiting for a $3,000 rig. Boot up Colab, pull these models, and start building your stack. I’ve put together a completely free, cell by cell Google Colab notebook that automates this entire workflow so you can test it yourself in 5 minutes and learn. Link to the notebook is in the comments below. Experiemt with different MTP parameters, context windows and post your results in the comments.show more

Alok
170,442 Aufrufe • vor 1 Monat
【🌷 3D REVEAL~! 🌷】 It was a dream (of... spring hehe) come true to join some of my most precious friends to celebrate Nimi's 1-year #nimiversary at her 3D concert and to sing and dance together~! 🎶 This is actually both my new 3D model and was my first time doing mocap. 💖 It means a lot to me, and I'm so happy it got to be with friends!!! I love love love how Ayn @ VGEN was able to bring my wonderful mama さえきやひろ💀's art to life so skillfully and beautifully. 🎨 This is the reveal of my proper main hairstyle too~! I'm able to express myself so well with it, from every soft and silly expression, so the way the flowy dress moves like petals in the wind, to the bounce and flutter of my wings! 🎀 And this model actually was made possible BECAUSE of Nimi 💚 I didn't have a 3D model at all at the time, but I'm happy that she wanted me to be able to stand by her side and dance together and hug her ;////; I will treasure this model always. It's another important addition to the fairytale we've been writing together! I'll keep putting all of my heart into blooming more and more as an idol, and making happy memories with you all like we did today!show more

Phoebe Chan 🐝🌸 Densetsu.EXE 🔜 Skyline Serenade
63,861 Aufrufe • vor 6 Monaten
Norris reveals what he was thinking during the last... two laps in Abu Dhabi: "I just started thinking of my first day in a go-kart. I was on this kind of mini tennis court, and we put some slicks on the go-kart, and I was just doing some doughnuts and just having some fun" "And slowly around the lap, I kind of pictured the years that followed: stepping up into the proper racing to going on to race in Europe, the World Championships 2014, Formula 4, Ginetta, New Zealand, Jerez, F3, F2 " "And that's how my last two laps pretty much went, until we go under the hotel. All of a sudden, I pictured my mum in the garage. That was the first time, the first moment the whole year, I just about started to realise what was happening, what was about to happen" "And all I did was picture the garage. Picture my parents there, my brother, my sisters, all in the garage for the final four corners, and I came around the last corner, and again, then that next step of emotion starts to kick in — and realisation — of what's happened" "The last 18 years all led to this one moment"show more

Holiness
36,939 Aufrufe • vor 7 Monaten
The past year has seen me have a renaissance,... in the truest sense… I won’t go into details now but will at some point before long. What has brought so much happiness to my life and those around me this past year has been my falling back in love with sport. Cycling has, and always will be, my number one. Yet I’d forgotten that I simply love sport, not for results but for the sheer joy of doing it, I’d completely forgotten that the health of my mind is intrinsically connected to the health of my body. I’ve rediscovered the love I had for sport that existed before the world of professional cycling took over in the way it did. I’ve been pushing myself and trying new things this past year, indifferent to the results, just out having fun and at times going deeper than I thought I was capable of anymore. Last week I got on a TT bike for the first time in a decade, Factor Bikes built me a bike, I’ve been looking at it for two years and decided it was time to get fitted, getting back on it felt like going home. Anyway, the long and the short of this is that it’s inspired me to create a club to inspire and be inspired. A community for us to share our love for getting out there and doing it, because I’ve realized that although I spend most of my sporting life on my own I derive the most pleasure when feeling part of something. It’s in its early days, I’ve called it Sporting Club CHPT3 aka SCC3, I’d love you to check it out and join. It’s still in its infancy, but I hope it’s going to grow into something that will inspire you as much as me.show more

David Millar
111,710 Aufrufe • vor 2 Jahren
One of the things I’ve struggled with this year... is my weight, due to a new medication. I’ve always been someone who loves food, but after going on birth control for the first time in 10+ years I gained 25 pounds in less than 3 months. The birth control was one attempt to help with my severe PMDD symptoms, but sadly it didn’t work AND caused my cravings to skyrocket. After doing a full 3 month test my doctor and I decided to try other treatment options for PMDD. A big reason for that decision was because my body fat percentage was quickly nearing 40% and was starting to cause other health issues. I’ll be completely honest, I lost a lot of motivation in the months after gaining the weight. It became seemingly impossible to get back to where I was just 6 months ago. How could it go downhill so quickly? I am proud to report, however, that after 3 months of changing my mindset and focusing on myself I have lost 7 of those 25 pounds. I am finally going to the gym consistently again, eating the correct amount of calories per meal (except on big film days 😅), and have been spending more time with people who lift me up. I still have a long way to go, but for the first time in over a year I’m really starting to feel like myself again. Thank you to all of you who’ve been on this wild and incredible journey with me all these years 😘show more

Rosanna Pansino
266,082 Aufrufe • vor 5 Monaten
🚨 UPDATE: Twin Falls, Idaho shooting hero Jordan Salinas... dropped this reaction to the nationwide praise he's receiving for his heroic actions "The last 24 hours have been overwhelming. I want to sincerely thank my family, friends, and everyone who has reached out with messages, prayers, and support. I haven't been able to respond to everyone yet, but please know I've seen your messages, and they mean more to me than I can put into words. In the coming days, I'll share a more complete account of what happened, to the best of my recollection, when I'm able. For now, my heart remains with the victims, their families, and everyone affected by yesterday's tragedy. Thank you for continuing to keep them in your thoughts and prayers, and thank you all for the incredible support you've shown me and my family." 👏🏻🙏🏻show more

Eric Daugherty
155,388 Aufrufe • vor 26 Tagen
【🌷 POST-OTAKON AND OFFKAI INTRO🌷】 During the almost 8-week... performance gauntlet I've been on, I've had the privilege of meeting YOU!! For those who are new to the FeebeeHive, I'm Phoebe, your 2.5D blooming flower idol, singer-songwriter, and voice actor who performs on both the real life and virtual stage 🐝🌸 I've been doing kaigai idol activities since 2017-18 and kept them going when I started VTubing in 2020 via mixed reality integration. I aim to bridge dimensional and cultural boundaries, bringing people together through songs full of hope and empathy! 🎶 To this day, I'm doing it as both my IRL self and recently have also started chasing my virtual dreams with my best friends Vicky and Mint in VTuber idol group Densetsu.EXE as well 🩷💚🩵 I believe that this is all my life's purpose and I'm striving to reach as many hearts all over the world with my voice as possible. So please come along on this journey with me and let's GO TO THE MOOOON!!!! ✨️show more

Phoebe Chan 🐝🌸 Densetsu.EXE 🔜 Skyline Serenade
12,770 Aufrufe • vor 24 Tagen
🚨 Brie Bella has spoken out on her injury... and revealed that she broke her scapula. "I wanted to thank you all for sending me all the love and warm wishes. When I was sitting in the hospital, I was reading everything you guys were sending my way, love you all!!! Unfortunately I’ll be out for a bit. I broke my scapula. If there’s one thing I know about my sister and I is that we don’t let broken bones stop us. Not sure how I finished that match, but I do believe between the adrenaline, passion for the business and the love for all the women I work with, it gave me the drive to finish. Now my new journey starts!! Buddy said he’s going to be my assistant, Birdie‘s going to be my nurse and they both told me that daddy is going to be my personal chef, so I am set!! if there is one thing I know about Brie Mode, she always finds a way to come back to her wrestling home. Let the recovery begin. 🫶🏽"show more

maisha ★
129,649 Aufrufe • vor 25 Tagen