Video yükleniyor...

Video Yüklenemedi

Ana Sayfaya Dön

Very proud to share that we just release Luce KVFlash. Run your preferred model inside Lucebox at 256k context, without thinking about KVCache and OOM, up to 2.9x faster decoding at long context. Taking inspiration from OS paging and using our speculative prefill method (Luce PFlash), we managed to...

25,703 görüntüleme • 3 ay önce •via X (Twitter)

0 Yorum

Yorum bulunmuyor

Orijinal gönderinin yorumları burada görünecek

Benzer Videolar

50% more context unlocked for Qwen 3.8 27b Q4_K_XL dflash 2 on a single RTX 4090 (24 GB VRAM) I found a hidden VRAM tax in llama.cpp. By combining my custom 2 bit DFlash 2 drafter with one overlooked server flag, I just unlocked another +80,000 tokens of context. Qwen3.8-27B is now running a massive 250,000 context at 75 tokens/s on a single RTX 4090. Here is the secret: By default, `llama-server` reserves massive chunks of your VRAM to handle multiple concurrent users (batching). If you are running a single user session, you are bleeding memory for features you aren't using. By passing the `--parallel 1` flag, you force the engine to dedicate 100% of your 24GB VRAM buffer to a single user. When we combine the VRAM saved by our Q2_K 2-bit drafter with the VRAM saved by `--parallel 1`, the context ceilings absolutely explode: Note: all benchmarks carried out with a massive 28k prompt. Ubuntu 22. ### THE NEW 24GB PHYSICAL LIMITS (Single RTX 4090): # 1. The "Repo Swallower" (Q4 KV Cache): - Context: 250,000 tokens (Up from 170k!) - Speed: 73.66 t/s decode | 1,608 t/s prefill - Peak VRAM: 23.8 GB # 2. The "High-Precision SWE" (Q8 KV Cache): - Context: 150,000 tokens (Up from 100k!) - Speed: 75.01 t/s decode | 1,667 t/s prefill - Peak VRAM: 23.9 GB # 3. The "Pristine Attention" (Unquantized FP16 KV): - Context: 90,000 tokens - Speed: 80.58 t/s decode | 1,699 t/s prefill - Peak VRAM: 23.92 GB ### HOW TO RUN THE 250K GOD STACK TODAY: (Requires PR #27342 + my Q2_K Hugging Face drafter) llama.cpp flags: ./build/bin/llama-server -m Qwen3.8-27B-UD-Q4_K_XL.gguf -md Qwen3.8-27B-DFlash2-Q2_K.gguf --spec-type draft-dflash --spec-draft-n-max 3 -c 250000 -ngl 99 --parallel 1 --port 8080 -ctv q4_0 -ctk q4_0 We are pushing a quarter million tokens of context with speculative DFlash 2 decoding at 73 tokens/second on a single consumer gaming GPU. I dropped my custom 2 bit Hugging Face GGUF links, visual performance graphs, and the PR #27342 build instructions in the replies below. If you own a single RTX 3090 or 4090, it is officially time to cancel your API subscriptions and let local silicon eat the cloud. how much monthly API spend does an optimized 4090 rig like this actually replace for you?

Alok

39,189 görüntüleme • 1 ay önce

Muse Glimmer, A 30B parameter dense model swallowing a 130,000 token context window using only 19.3 GB of VRAM (extreme efficiency). No KV cache quantization required. I just benched the new Muse Glimmer 30B (dense) on a single RTX 4090. We are pulling 3,100+ t/s prefill and 75 tokens/second decode. The throughput is violent. Meta superintelligence lab just open sourced this agentic beast, explicitly engineered to dominate 24GB consumer cards. I pulled the latest llama.cpp source on Ubuntu 22 (CUDA 13) to see if the specs were real. Fed it a 28k token prompt. Here is the exact llama.cpp God Stack and benchmarking breakdown: # 1. The Deep Context Run (No Speculative Decoding) The architecture uses a massive 16:1 GQA (Grouped Query Attention) ratio. This means the KV cache footprint is practically non existent. ./build/bin/llama-server -m Muse-Glimmer-30B-UD-Q4_K_XL.gguf -c 130000 -b 4096 -ub 4096 -ngl 99 --port 8080 Prefill: 3134.95 t/s Decode: 50.00 t/s VRAM: 19.34 GB (I hit 130k context on pristine, unquantized f16 cache and still had 4.5 GB of VRAM left over. Absolute witchcraft). # 2. The DFlash Speculative Overdrive Meta shipped this with a DFlash block diffusion drafter. Let's trade that extra VRAM for pure speed. ./build/bin/llama-server -m Muse-Glimmer-30B-UD-Q4_K_XL.gguf -md dflash-kquant.gguf --spec-type draft-dflash --spec-draft-n-max 3 -c 80000 -b 4096 -ub 4096 -ngl 99 --port 8080 Prefill: 1293.69 t/s Decode: 75.00 t/s VRAM: 23.93 GB (Maxed out on card) the dflash gguf is additional 1.6 GBs # The Architecture Insight (Muse Glimmer vs. Gemma 4 31B) If you look at my Gemma 4 31B tests from last week, getting 140k context required heavily degrading the memory with Q4 KV quantization (gemma 31b q4 can do only about 40k context with unquantized kv on a 24gb card). That "unzipping" overhead bottlenecked Gemma's MTP decode speeds down to 65 t/s. Muse Glimmer completely sidesteps this bottleneck. By using aggressive 16:1 GQA, it keeps the KV cache in native f16 format at massive context lengths. Flash Attention gets to run at maximum uncompressed speed, letting the DFlash drafter push decode safely to 75 t/s without compute lag. With a 76% on SWE Bench Verified and seamless local tool calling, this model looks promising. Unsloth's Hugging Face GGUF links, intelligence/agentic benchmark details, and inference throughput performance graphs are posted in the replies. For 24GB rig, what’s your current go to model?

Alok

65,480 görüntüleme • 1 ay önce

I'm proud to share that Glean has surpassed $300M ARR, just five months after crossing $200M and growing ~3x over the past 15 months. This is an exciting milestone for Glean, and it's a signal about where the enterprise AI market is heading. We’ve long believed the real challenge in enterprise AI is not access to models. It is grounding AI in how a company actually works: its people, knowledge, workflows, permissions, and systems. That’s even clearer now. The companies creating real value with AI are not just adopting better models. They are building systems that understand their business well enough to deliver reliable outcomes at scale. That is the real moat, and it is what we’ve been building at Glean: an unrivaled context layer for enterprise AI. That context has to work across the business, not just inside a single team or use case. We see that in how customers adopt Glean: more than 85% use it across five or more job functions. It also has to meet the security and governance demands of complex enterprises. We see that in who is choosing Glean: our Fortune 500 customer count nearly doubled year over year. And it has to make economic sense as usage grows. In our recent benchmark with Claude Cowork, Glean was preferred roughly 2.5x as often as off-the-shelf MCP tools and used 30% fewer tokens on average. Better context improves both quality and efficiency. I enjoyed talking with CNBC's Deirdre Bosa about this broader shift. In enterprise AI, the winners will not be defined by better models alone. They will be defined by who builds the strongest foundation for enterprise context. Thank you to our customers, partners, and team for helping us build the future of enterprise AI.

Arvind Jain

281,328 görüntüleme • 3 ay önce

The "I don't have enough VRAM" excuse just died. I’m running Meta’s new 30B Muse Glimmer Q6_K_XL with a massive 130k context window on just 26GB VRAM FREE compute on Kaggle. Kaggle provides you free 2x Nvidia T4 GPUs. 30 hours usage each week! Yesterday, I showed you the violent throughput of Muse Glimmer on a single RTX 4090. Today, we are securing a Dual NVIDIA T4 GPU cluster with 32GB of total VRAM for exactly $0 and dropping the massive 24.5GB Q6_K_XL GGUF onto it. Here is the exact Kaggle workflow and benchmarking breakdown: # 1. The Storage Bypass & Setup I built a clean cell by cell script in the file. We dynamically fetch the CUDA accelerated llama.cpp binaries and use wget to stream the model directly into Kaggle's /kaggle/tmp scratch storage, which cleanly bypasses their 19.5GB output directory limit. # 2. The Multi GPU Performance With the -ngl 99 flag offloading all model layers across both T4 GPUs (32GB VRAM combined), we pushed a massive 131,072 token context window (-c 131072). The benchmark numbers: Prefill: 265.9 t/s Decode: 9.0 t/s VRAM Total: 26.5 GB # 3. The Architecture Insight The Q6_K_XL model itself is 24.5 GB. Because of Muse Glimmer's aggressive 16:1 GQA, the unquantized KV cache for a massive 130k context window only takes up 2 GB of memory. No heavily degraded Q4 KV quantization required. It just works. No compiling from source. No credit card. No OOM crashes. Zero excuses. If you’re running a single RTX 3090, 4090, or 5090, you need to experience this hyper efficient KV cache right now before the upcoming Qwen 3.8 27B drop completely steals your VRAM tomorrow. pick the Q4 or Q5 quants for 24 GB VRAM rigs. I'm dropping the Unsloth huggingface GGUF links and the free Kaggle notebook link in the replies. spin up your own instance, and show me your multi GPU benchmarks.

Alok

19,370 görüntüleme • 1 ay önce

Recently there was increased discussion about DMCA takedown requests being used in bad faith between competitors on Shopify. I can confirm that the volume of these has very much increased recently. Let me give you some examples of what we are doing about it. Our goal is to work within the DMCA’s framework and make a system in which the right things happen quickly. We want to give merchants the maximum time to respond with the minimum disruption. We also need legitimate claims to resolve in favor of the claimants quickly. The first thing we did recently is to productize this properly. We shipped a feature to manage these cases directly in the Shopify admin to make it a lot more obvious. Before that a lot was managed through form emails. This process is now much simpler and faster. The second thing we did is to make the identity system more robust so that we know more about the claims filed. We consider the claimants previous cases in order to decide what decisions to take. Similarly we consider documentation from previous cases that merchants had to deal with more strongly to make accurate decisions when repeat claims are received against products. If we detect bad faith claims from accounts we take action. We recently filed multiple lawsuits against claimants in court. This matches our aggressive stance when we see frivolous or bad faith patent action in which we had a lot of success invalidating bad patents. The end result of this is that we catch spammy or frivolous claims at much higher frequency already. We also make responding to cases much easier and this can be managed directly from the admin now and is less work for the merchants. Lots of additional improvements coming- but what I can guarantee you is that that Shopify will be the worst platform on which to make fraudulent DMCA claims on. “Be merchant obsessed” is a core value here, and we reflect that better in this area.

tobi lutke

142,265 görüntüleme • 2 yıl önce

If you thought the Gemma 4 31B (dense) model was fast, sit down. I just benched the updated Gemma 4 26B A4B MoE on a single RTX 4090 (24 GB VRAM) 9,200 t/s prefill. 160 t/s decode. 250,000 context window. All on a single consumer RTX 4090. The numbers are completely unhinged. The 31B is a dense behemoth. But the 26B is a Mixture of Experts (MoE), specifically an Active 4 Billion (A4B). It holds 26B parameters of knowledge but only activates 4B per token. Because its inference memory footprint is so light, I didn’t even need KV cache quantization to hit a quarter million context. Compiled the latest llama.cpp from source on Ubuntu 22 (CUDA 13). Fed it a 28k token prompt, and manually cranked the batch sizes (-b 2048 -ub 2048) to absolutely redline the Tensor Cores. Here is the benchmarking breakdown: # 1. The Baseline (No MTP) Even without speculative decoding, the A4B architecture flies. llama.cpp flags: ./build/bin/llama-server -m gemma-4-26B-A4B-it-qat-UD-Q4_K_XL.gguf -c 250000 -ngl 99 -fa on -b 2048 -ub 2048 --port 8080 -v Context Ceiling: 250,000 tokens (21.5 GB VRAM) Prefill: 9,200 t/s (Absurd) Decode: 124 t/s # 2. The MTP Overdrive Injected the new MTP draft model to enable Speculative Decoding. llama.cpp flags: ./build/bin/llama-server -m gemma-4-26B-A4B-it-qat-UD-Q4_K_XL.gguf --spec-type draft-mtp --spec-draft-model mtp-gemma-4-26B-A4B-it.gguf --spec-draft-n-max 4 --spec-draft-p-min 0.7 -c 250000 -ngl 99 -fa on -b 2048 -ub 2048 --port 8080 -v Context Ceiling: 250,000 tokens (22.96 GB VRAM) Prefill: 7,054 t/s (MTP draft overhead slightly caps prefill) Decode: 156 t/s # The Agentic Architecture Insight Why does this matter? Because you can now build a killer local agentic loop on a consumer desktop. Use the 31B dense model (from the previous post) as your heavy, deliberate Orchestrator / Verifier / Planner. Pass the actual execution tasks to this 26B MoE. At 160 t/s, this MoE can chew through code generation, tool calling, and massive RAG document retrieval over a 250k context window almost instantly, drastically speeding up your agentic loop. If you own a single RTX 3090 or 4090 and haven't tried this specific stack yet, you need to pull these latest updates and run it. Local inference just leveled up. Hugging Face links to the Unsloth 26B QAT quants and MTP drafters are in the replies. performance graphs also available in the replies.

Alok

40,993 görüntüleme • 1 ay önce

I’m incredibly excited (and very nervous!) about the launch of AMPLIFY. I have always felt enormous gratitude to have been born in Australia but, like many of you, I felt that we weren’t going in the right direction as a country. I started turning my mind to what we might do about the situation and got advice and encouragement from some amazing people. We decided to set up an organisation that is all about community at its core. So what is AMPLIFY? AMPLIFY is a community where Australians get to have their say and make a difference on the most important issues that we face as a country. We are non-partisan and completely independent of any political party. Anyone can become a member of AMPLIFY at no charge. We are seeking to find “uncommon ground” and identify the right solutions to the big issues facing Australia. We will do this by bringing our community together for events in all parts of the country, we will facilitate online conversations, share evidence, talk with experts and come up with solutions. The way in which we will make an impact is by taking our ideas to politicians and Amplifying the voice of our community to spark change. Our goal is to help Australia become a more prosperous, fairer, more cohesive and happier country. I will be serving as founder and Chair of AMPLIFY. My job at Square Peg is not changing in any way. Georgina Harrisson is our CEO and has been doing a remarkable job since joining last October. I am also really proud that my son Joel is an important part of the AMPLIFY team. I am grateful to the amazing group of people who have joined the board; Suzi Carp, Rona-Glynn-McDonald, Kate Jones, Gill McLachlan, Dom Perrottet, Kate Pounder, Mike Schneider and Zara Seidler. If we are going to be successful we need your support. You can sign up at our website and become a member in a few seconds. Please spread the word. Please share this Tweet or, better still, post your own Tweet and share with your audience.

Paul Bassat

25,975 görüntüleme • 2 yıl önce

Hills I will die on as someone who has coached high school football for over 29 years: 1. If you are not PASSIONATE about blessing, serving, and empowering those you are blessed to coach, this profession is not for you. 2. As much as we need to know our trade, getting to know (and to love), our players is far more important. 3. This is an INTENSE game, and it’ll never be “just a game”, but it IS a game. Remember that when you’re with your team, and more importantly, remember that when you’re with your family. 4. Just as we teach our athletes to “leave things better than they found them”, we need to leave our athletes better than they were when they first entered into our program. Never let a day pass without pouring into each and every individual. 5. Life is complicated enough, let’s not complicate the game in such a way that we take the joy of it away from others. In other words… Keep it simple. 6. Our words carry little (or NO), value, if we don’t practice what we preach. WE as coaches should be learning and growing each and every day, just as we expect our athletes to. 7. As much as we all want to win those championship rings for our athletes, make sure you don’t lose your wedding ring in the process. 8. The athlete that may be “difficult to reach/teach” (the one who may get on your last nerve more than you could ever imagine), is someone’s EVERYTHING. Get to know them as human beings, find out what motivates them, and do everything you can to help them to thrive. 9. Be where your feet are. Don’t fall into the trap of chasing logos and thinking that a higher division, a bigger school, or going from HS to college, or even college to the pros, is going to be more rewarding or fulfilling. 10. The legacy you leave as a coach will never be determined by your wins and losses, but by the lives you were able to change for the better!

Coach Hines 🇺🇸

64,993 görüntüleme • 3 ay önce

Hills I will die on as an elementary school teacher, who just wrapped up my 32nd year teaching! 1. If you are not PASSIONATE about blessing, serving, and empowering those you are blessed to teach, this profession is not for you. 2. Our students are not “ours”. They are their families, and we need to understand the magnitude of the calling, and responsibility we have to meet them where they are, and to help them to get to a place that they never thought possible. 3. As important as the curriculum is (and it IS important), the children in our classroom are what matter most. It’s our JOB to teach the curriculum to SERVE our students, NOT to use our students to push any sort of agenda. 4. Just as we teach our students to “leave things better than they found them”, we need to leave our students better than they were when they first entered into our classroom. Never let a day pass without pouring into each and every child. 5. We should not teach our children WHAT to think, but HOW to think, and how to use that knowledge to bless and serve not only themselves, but the world around them. 6. Our words carry little (or NO value), if we do not practice what we preach. 7. If we don’t make learning fun, children will view learning as a chore, and we we will be creating a generation of children who grow up to be young adults who don’t see the joy in learning new things. 8. The child that may be “difficult to reach/teach” (the one who may get on your last nerve more than you could ever imagine), is someone’s EVERYTHING. Get to know them as human beings, find out what motivates them, and do everything you can to help them to thrive. 9. Never tell a child they “can’t” do something. God has blessed each of them with far more strengths and talents than we may know, and it’s not our job to tell them what they can’t do, but to help them to realize all the things they CAN do. 10. The legacy you leave as a teacher will never be determined by your student’s test scores, but by the human beings you helped them become throughout their lives.

Coach Hines 🇺🇸

62,267 görüntüleme • 4 ay önce