正在加载视频...

视频加载失败

My dual RTX PRO 6000 setup is currently training a Draft model for Qwen 3.6 27B! 🔥 I'm taking the paper DeepSeek dropped on 6/26 and going for a super ambitious application to the 27B scale. Thanks to my homelab, I was able to dive straight in — I...

56,314 次观看 • 1 个月前 •via X (Twitter)

0 条评论

暂无评论

原始帖子的评论将显示在这里

相关视频

I told you to claim your free 16GB NVIDIA GPU for learning Local LLMs. Now I’m going to show you how to double its inference speed without touching the hardware. Google Colab gives you an enterprise grade NVIDIA Tesla T4 GPU for free, roughly 4 hours every single day. It is the absolute perfect sandbox for learning AI engineering, testing inference flags, and pushing massive context windows. The local AI timeline is moving way too fast. If you aren't using Multi Token Prediction (MTP) yet, you are leaving massive performance on the table. I just pushed DeepMind’s Gemma 4 26B to 64.9 t/s on this exact free tier. Let's look at the raw benchmark data running on an Ubuntu Linux environment with the latest compiled llama.cpp binaries and quantized GGUFs from Unsloth via HuggingFace: # Qwen 3.5 9B (Dense): Base: [ Prompt: 626.7 t/s | Generation: 21.0 t/s ] With MTP: [ Prompt: 539.1 t/s | Generation: 24.8 t/s ] # Gemma 4 26B QAT (MoE): Base: [ Prompt: 634.2 t/s | Generation: 48.3 t/s ] With MTP: [ Prompt: 572.1 t/s | Generation: 64.9 t/s ] If you are paying attention, this single Colab notebook reveals 3 massive observations about the current state of local LLMs: # 1. The MTP Speedup (Software Overclocking) Standard autoregressive decoding guesses one token at a time. MTP acts like a highly optimized, built in speculative decoder. It predicts multiple future tokens at once and the main model verifies them in parallel. The result? Zero accuracy loss and a massive throughput increase. Gemma jumped from 48 to 65 t/s just by flipping a flag. # 2. The MoE Paradox (Bigger is Faster) How does a 26B parameter model absolutely destroy a 9B model in raw speed on the exact same hardware? Architecture. Qwen 3.5 9B is a dense model. it activates all 9 billion parameters for every single token. Gemma 4 26B is a Mixture of Experts (MoE) model. It routes data efficiently, activating only 4B parameters per token. You get the reasoning capabilities of a 26B model with the compute cost of a 4B model. 3. Thinking Efficiency When I ran the exact same complex prompt on both models, the larger MoE spent significantly fewer "thinking" tokens to arrive at the correct answer. A smarter model doesn't just give better answers; it gets to the point faster, saving you compute cycles and preserving your context window. # Want to run this yourself? Here are the exact llama.cpp CLI commands. For Qwen (MTP is baked into the main model): ./llama-cli -m Qwen3.5-9B-UD-Q4_K_XL.gguf -p "Explain quantum computing." -n 2000 -c 8000 -ngl 99 -fa on --spec-type draft-mtp --spec-draft-n-max 4 --spec-draft-p-min 0.7 For Gemma (Using a separate lightweight draft model): ./llama-cli -m gemma-4-26B-A4B-it-qat-UD-Q4_K_XL.gguf --model-draft mtp-gemma-4-26B-A4B-it.gguf -p "Explain quantum computing." -n 2000 -c 8000 -ngl 99 -fa on --spec-type draft-mtp --spec-draft-n-max 4 --spec-draft-p-min 0.7 Stop waiting for a $3,000 rig. Boot up Colab, pull these models, and start building your stack. I’ve put together a completely free, cell by cell Google Colab notebook that automates this entire workflow so you can test it yourself in 5 minutes and learn. Link to the notebook is in the comments below. Experiemt with different MTP parameters, context windows and post your results in the comments.

Alok

170,442 次观看 • 21 天前

The first generation of DeePle is officially over! To those who minted—welcome! And to those who bought from the secondary market, you have my utmost respect. So... what’s next? I’d like to say this is just the beginning of an evolving project. There will be a second generation with fresh items and unique features, followed by a third, and more. I'm excited to dive deeper into the DeePleVerse and explore all its possibilities, and it would be an honor to have all of you along for the journey. Now, I’d like to share some thoughts and the vision behind this experimental yet fun art project. Over the past years, we’ve seen countless ups and downs, with beautiful and ugly moments unfolding simultaneously in web3. Here’s what I’ve learned along the way: 1/1s and Editions are fantastic for maintaining authenticity and scarcity, especially from incredible artists. However, they can feel too exclusive, making it difficult to build a larger community. On the other hand, PFPs are great for fostering a community, bringing people together around a shared belief. But they often lack the authenticity of fine art and can attract too much speculation. I’ve always wondered how we could bring the best of both worlds together. And that’s how DeePle was born—after much thought and exploration. I wanted to remove the gambling aspect, so there’s no reveal or rarity. I wanted collectors to have full control, allowing them to choose what to mint based on what they see. I also believe it’s crucial to strike the right balance between public and coordinated spots for the drop to avoid creating a speculative atmosphere. While I have some interesting ideas for the future, I’m never going to reveal what’s coming or use the word "roadmap" because I want to avoid creating unnecessary hype and making unrealistic promises. For now, I hope people find value in the experience and become part of the DeePleVerse. The goal is to build a genuine community, step by step, with each new round of drops, and to gradually bring in more people over time. None of this would have been possible without the Shape and Transient Labs. Shape created an impressive and efficient L2 chain that enables the complexity and functionality necessary for this project. And Transient Labs fully understood my vision, creating the perfect tool to bring DeePle to life. I’m incredibly excited for what’s ahead!

DeeKay

74,156 次观看 • 1 年前

The past year has seen me have a renaissance, in the truest sense… I won’t go into details now but will at some point before long. What has brought so much happiness to my life and those around me this past year has been my falling back in love with sport. Cycling has, and always will be, my number one. Yet I’d forgotten that I simply love sport, not for results but for the sheer joy of doing it, I’d completely forgotten that the health of my mind is intrinsically connected to the health of my body. I’ve rediscovered the love I had for sport that existed before the world of professional cycling took over in the way it did. I’ve been pushing myself and trying new things this past year, indifferent to the results, just out having fun and at times going deeper than I thought I was capable of anymore. Last week I got on a TT bike for the first time in a decade, Factor Bikes built me a bike, I’ve been looking at it for two years and decided it was time to get fitted, getting back on it felt like going home. Anyway, the long and the short of this is that it’s inspired me to create a club to inspire and be inspired. A community for us to share our love for getting out there and doing it, because I’ve realized that although I spend most of my sporting life on my own I derive the most pleasure when feeling part of something. It’s in its early days, I’ve called it Sporting Club CHPT3 aka SCC3, I’d love you to check it out and join. It’s still in its infancy, but I hope it’s going to grow into something that will inspire you as much as me.

David Millar

111,669 次观看 • 2 年前

How I transformed my physique from a dream to a goal achieved. 1. Clarify Vision & Strategy. I knew I wanted more muscle and more curves but I didn’t know how. So, I got out of my head and took the necessary step to hire a reputable coach to lead me. 2. Develop A Plan & Stick To It. I followed the plan and trusted the process. I followed the food plan to the gram. I learned how to train effectively and efficiently. Never skipped a workout because I didn’t feel like training. Never skipped a cardio session because I didn’t feel like doing it. Never starved myself. I trained hard consistently for the last 7 years. 3. Improve Communication Skills. I became a better communicator so that my coach and I were on the same page, but this overflowed into all areas of my life. I stopped being afraid to ask questions. 4. Get Better At Training Every Day. I learned. I was dialled in. I lifted heavy with good form consistently. I sent training videos to my coach so I could get better. I pushed myself and surrounded myself with other like-minded individuals to help push me along. 5. Cultivate A Support Network. I cultivated a support network to help me achieve my physique goal: I found the right coach, physician (BHRT), sports psychologist, massage therapist, physiotherapist, and plastic surgeon (if you are looking to surgically enhance). Ready to transform your physique and turn your dream body into a reality? DM me for coaching. I know the way and I’m here to help. #transformation #onlinecoach #fitmom #goals #fitmotivation

Karen Orlena Wall

15,130 次观看 • 25 天前