Sensitive content

This media may contain sensitive content.

正在加载视频...

视频加载失败

❌ S P A R K L E S & S O L E S ❌ This should be your only view forever. Beneath me. Beneath my Muslim soles. Lick your screens 😵‍💫 💦 ⛓️‍💥 Let your brain to slush for Mommy ♥️♦️ What are you twitching over? 1)...

17,363 次观看 • 9 个月前 •via X (Twitter)

0 条评论

暂无评论

原始帖子的评论将显示在这里

相关视频

50% more context unlocked for Qwen 3.8 27b Q4_K_XL dflash 2 on a single RTX 4090 (24 GB VRAM) I found a hidden VRAM tax in llama.cpp. By combining my custom 2 bit DFlash 2 drafter with one overlooked server flag, I just unlocked another +80,000 tokens of context. Qwen3.8-27B is now running a massive 250,000 context at 75 tokens/s on a single RTX 4090. Here is the secret: By default, `llama-server` reserves massive chunks of your VRAM to handle multiple concurrent users (batching). If you are running a single user session, you are bleeding memory for features you aren't using. By passing the `--parallel 1` flag, you force the engine to dedicate 100% of your 24GB VRAM buffer to a single user. When we combine the VRAM saved by our Q2_K 2-bit drafter with the VRAM saved by `--parallel 1`, the context ceilings absolutely explode: Note: all benchmarks carried out with a massive 28k prompt. Ubuntu 22. ### THE NEW 24GB PHYSICAL LIMITS (Single RTX 4090): # 1. The "Repo Swallower" (Q4 KV Cache): - Context: 250,000 tokens (Up from 170k!) - Speed: 73.66 t/s decode | 1,608 t/s prefill - Peak VRAM: 23.8 GB # 2. The "High-Precision SWE" (Q8 KV Cache): - Context: 150,000 tokens (Up from 100k!) - Speed: 75.01 t/s decode | 1,667 t/s prefill - Peak VRAM: 23.9 GB # 3. The "Pristine Attention" (Unquantized FP16 KV): - Context: 90,000 tokens - Speed: 80.58 t/s decode | 1,699 t/s prefill - Peak VRAM: 23.92 GB ### HOW TO RUN THE 250K GOD STACK TODAY: (Requires PR #27342 + my Q2_K Hugging Face drafter) llama.cpp flags: ./build/bin/llama-server -m Qwen3.8-27B-UD-Q4_K_XL.gguf -md Qwen3.8-27B-DFlash2-Q2_K.gguf --spec-type draft-dflash --spec-draft-n-max 3 -c 250000 -ngl 99 --parallel 1 --port 8080 -ctv q4_0 -ctk q4_0 We are pushing a quarter million tokens of context with speculative DFlash 2 decoding at 73 tokens/second on a single consumer gaming GPU. I dropped my custom 2 bit Hugging Face GGUF links, visual performance graphs, and the PR #27342 build instructions in the replies below. If you own a single RTX 3090 or 4090, it is officially time to cancel your API subscriptions and let local silicon eat the cloud. how much monthly API spend does an optimized 4090 rig like this actually replace for you?

Alok

39,189 次观看 • 5 天前

I told you to claim your free 16GB NVIDIA GPU for learning Local LLMs. Now I’m going to show you how to double its inference speed without touching the hardware. Google Colab gives you an enterprise grade NVIDIA Tesla T4 GPU for free, roughly 4 hours every single day. It is the absolute perfect sandbox for learning AI engineering, testing inference flags, and pushing massive context windows. The local AI timeline is moving way too fast. If you aren't using Multi Token Prediction (MTP) yet, you are leaving massive performance on the table. I just pushed DeepMind’s Gemma 4 26B to 64.9 t/s on this exact free tier. Let's look at the raw benchmark data running on an Ubuntu Linux environment with the latest compiled llama.cpp binaries and quantized GGUFs from Unsloth via HuggingFace: # Qwen 3.5 9B (Dense): Base: [ Prompt: 626.7 t/s | Generation: 21.0 t/s ] With MTP: [ Prompt: 539.1 t/s | Generation: 24.8 t/s ] # Gemma 4 26B QAT (MoE): Base: [ Prompt: 634.2 t/s | Generation: 48.3 t/s ] With MTP: [ Prompt: 572.1 t/s | Generation: 64.9 t/s ] If you are paying attention, this single Colab notebook reveals 3 massive observations about the current state of local LLMs: # 1. The MTP Speedup (Software Overclocking) Standard autoregressive decoding guesses one token at a time. MTP acts like a highly optimized, built in speculative decoder. It predicts multiple future tokens at once and the main model verifies them in parallel. The result? Zero accuracy loss and a massive throughput increase. Gemma jumped from 48 to 65 t/s just by flipping a flag. # 2. The MoE Paradox (Bigger is Faster) How does a 26B parameter model absolutely destroy a 9B model in raw speed on the exact same hardware? Architecture. Qwen 3.5 9B is a dense model. it activates all 9 billion parameters for every single token. Gemma 4 26B is a Mixture of Experts (MoE) model. It routes data efficiently, activating only 4B parameters per token. You get the reasoning capabilities of a 26B model with the compute cost of a 4B model. 3. Thinking Efficiency When I ran the exact same complex prompt on both models, the larger MoE spent significantly fewer "thinking" tokens to arrive at the correct answer. A smarter model doesn't just give better answers; it gets to the point faster, saving you compute cycles and preserving your context window. # Want to run this yourself? Here are the exact llama.cpp CLI commands. For Qwen (MTP is baked into the main model): ./llama-cli -m Qwen3.5-9B-UD-Q4_K_XL.gguf -p "Explain quantum computing." -n 2000 -c 8000 -ngl 99 -fa on --spec-type draft-mtp --spec-draft-n-max 4 --spec-draft-p-min 0.7 For Gemma (Using a separate lightweight draft model): ./llama-cli -m gemma-4-26B-A4B-it-qat-UD-Q4_K_XL.gguf --model-draft mtp-gemma-4-26B-A4B-it.gguf -p "Explain quantum computing." -n 2000 -c 8000 -ngl 99 -fa on --spec-type draft-mtp --spec-draft-n-max 4 --spec-draft-p-min 0.7 Stop waiting for a $3,000 rig. Boot up Colab, pull these models, and start building your stack. I’ve put together a completely free, cell by cell Google Colab notebook that automates this entire workflow so you can test it yourself in 5 minutes and learn. Link to the notebook is in the comments below. Experiemt with different MTP parameters, context windows and post your results in the comments.

Alok

170,442 次观看 • 1 个月前

this is the worst local ai will ever be. it only gets better from here. if you are not expanding your mind with these small models you are missing what's happening right now 99 percent tool call success rate. when steered well with the right skills and a framework like hermes agent the node becomes a cognition layer. not a chatbot. not a toy. an extension of how you think. i was cranking this node at 35 to 50 tok/s all day on personal experiments and now after all the work is done qwen 3.5 9B is iterating on its own code. the game it created. fixing its own bugs autonomously. and the part you should probably not miss is that all of this is happening on a RTX 3060. not an H100. not an A100. the card most of you have sitting in a drawer right now. if you just open that drawer and put that intelligence to work every tensor core on that card should be running for you. your work. your experiments. your thinking. you all have it but because nobody told you what this hardware can actually do in 2026 you never tried. the day it unlocks is the day you test your workload, understand the tradeoffs, debug the loops, and then decide if you need to scale the hardware. there is no point buying 3 mac studios when things done well you can squeeze a similar level of intelligence from 9B compared to 70B. but only when you create the right environment for your model through the right harness. and let me tell you i have tried claude code as a local harness. i have tried opencode. i have tried various others. somehow i landed on hermes agent and never left. there is something magical going on at Nous Research. the tool call parsers, the skills system, the way it handles small models natively. nothing else comes close for local inference. own your cognition. your AI. your agent. your prompts. your experiments. why give them away for free. those are who you are and they don't belong on someone else's servers being monitored. just give it a shot with your existing hardware. you run into a problem the community will help you. and if you are migrating from openclaw to hermes i will personally help you make the switch.

Sudo su

58,717 次观看 • 5 个月前

To all of the people JUST learning about the HORRIFIC stuff that's been done to children for sick thrills, political, monetary and leverage gain, come in and grab a chair. My advice to you, would be to find your nearest conspiracy theorist and have a discussion, or two, or three, depending on what you can actually handle because, it is DEEP, DARK and VERY SICK. Let me further educate you before your journey. Trump IS NOT one of these people. He is the "outsider" that started the ball rolling on this in the early 2000's. It doesn't matter what the media, your followers, or peer pressure is making you think. They are driving a narrative because the "outsider", the guy not allowed in the club anymore, KNOWS THEIR SECRETS. They want him gone, so business can continue as normal. This spans from your local cities, to your states, to the federal government, to other countries, all the way up to the people you now call the elites, the 1%. The World Banks, The WEF, The W.H.O., the U.N. It's how it's all been run, for at minimum, decades. It's done for control and power. I will say trust some of those people you have called crazy conspiracy theorists about this stuff for years and years. So much more of what they say is true, rather than not. A LOT of what you learned and were taught, was a lie. The term "Conspiracy Theorists" originated, to HIDE THE ACTUAL TRUTH! 🤔 We're not always right, but our track record is phenomenal. You'll see. Jus' Sayin' ✌️ Pray, A LOT!🙏🙏🙏 YOU'RE GOING TO NEED IT FOR THIS ONE. GOD LOVES YOU AND BROUGHT YOU HERE FOR A REASON! Most of US, are at this stage,👇 ready for JUSTICE 👇 #WWG1WGA #NCSWIC

Dark Knight

73,980 次观看 • 6 个月前

My Uber driver asked what I do for work. Software.... Cool. Can you look at something? He handed me his phone at a red light. Terminal. Claude chat. Green P&L. +$6,200. He drives Uber 4 days a week. Makes $1,100. Has a 2-year-old daughter. Where did you find this? Your article. The 10,000 wallets one. He read it 3 months ago. Did not understand half of it. Asked Claude to explain it like he's five. 214 messages. All during breaks between rides. Parked at gas stations. Waiting for pings. 1st thing Claude told him: 87% of wallets lose money. Do not be the 87%. He installed poly_data. Fed it to Claude. Found 47 wallets with Sharpe above 2.0. Filtered crypto only. Quarter Kelly. $200 starting bankroll. From his tips. 93 messages later Claude helped him build the 20-line brain from the article. Bayesian updates. EV filter at 5%. Fully automated. Last 45 days: → 480 trades → 91.3% win rate → +$6,200 Best trade, whale convergence on Fed rate cut. 4 wallets entered in 2 minutes. Entry $0.12. Resolved $1.00. +$1,760. While dropping off a passenger at JFK. The passenger tipped him $5. The bot made $1,760. His wife found the Telegram alerts on his phone. Thought he was texting another woman. He showed her the P&L curve. Can you make me one? How long until you quit driving? He looked at me through the rearview mirror. I'm not stopping. Uber is my cover story. I wrote the article. He actually opened terminal. You only need Claude + laptop + 1 hour/day. Giving This Free for 24 hours. To get it: 1. Comment the word "Trade" 2. Like and Retweet this post 3. Follow me Himanshu Kumar (so i can DM you) Save this post. Deploy the bot this weekend. Start with $200. Scale on evidence.

Himanshu Kumar

44,803 次观看 • 2 个月前

My Uber driver asked what I do for work. "Software." "Cool. Can you look at something?" He handed me his phone at a red light. Terminal. Claude chat. Green P&L. +$6,200. He drives Uber 4 days a week. Makes $1,100. Has a 2-year-old daughter. "Where did you find this?" "Your article. The 14,000 wallets one." He read it three months ago. Didn't understand half of it. Asked Claude to explain it like he's five. 214 messages. All during breaks between rides. Parked at gas stations. Waiting for pings. First thing Claude told him: 87% of wallets lose money. Don't be the 87%. He installed poly_data. Fed it to Claude. Found 47 wallets with Sharpe above 2.0. Filtered crypto only. Quarter Kelly. $200 starting bankroll. From his tips. 93 messages later Claude helped him build the 20-line brain from the article. Bayesian updates. EV filter at 5%. Fully automated. Last 45 days: → 480 trades → 91.3% win rate → +$6,200 Best trade: whale convergence on Fed rate cut. 4 wallets entered in 2 minutes. Entry $0.12. Resolved $1.00. +$1,760. While dropping off a passenger at JFK. The passenger tipped him $5. The bot made $1,760. His wife found the Telegram alerts on his phone. Thought he was texting another woman. He showed her the P&L curve. "Can you make me one?" "How long until you quit driving?" He looked at me through the rearview mirror. "I'm not stopping. Uber is my cover story." I wrote the article. He actually opened terminal. You only need Claude + laptop + 1 hour/day. Giving This Free for 24 hours. To get it: 1. Comment the word 'AutoPilot' 2. Like and Retweet this post 3. Follow me Marry Evan

Marry Evan

220,142 次观看 • 3 个月前

Muse Glimmer, A 30B parameter dense model swallowing a 130,000 token context window using only 19.3 GB of VRAM (extreme efficiency). No KV cache quantization required. I just benched the new Muse Glimmer 30B (dense) on a single RTX 4090. We are pulling 3,100+ t/s prefill and 75 tokens/second decode. The throughput is violent. Meta superintelligence lab just open sourced this agentic beast, explicitly engineered to dominate 24GB consumer cards. I pulled the latest llama.cpp source on Ubuntu 22 (CUDA 13) to see if the specs were real. Fed it a 28k token prompt. Here is the exact llama.cpp God Stack and benchmarking breakdown: # 1. The Deep Context Run (No Speculative Decoding) The architecture uses a massive 16:1 GQA (Grouped Query Attention) ratio. This means the KV cache footprint is practically non existent. ./build/bin/llama-server -m Muse-Glimmer-30B-UD-Q4_K_XL.gguf -c 130000 -b 4096 -ub 4096 -ngl 99 --port 8080 Prefill: 3134.95 t/s Decode: 50.00 t/s VRAM: 19.34 GB (I hit 130k context on pristine, unquantized f16 cache and still had 4.5 GB of VRAM left over. Absolute witchcraft). # 2. The DFlash Speculative Overdrive Meta shipped this with a DFlash block diffusion drafter. Let's trade that extra VRAM for pure speed. ./build/bin/llama-server -m Muse-Glimmer-30B-UD-Q4_K_XL.gguf -md dflash-kquant.gguf --spec-type draft-dflash --spec-draft-n-max 3 -c 80000 -b 4096 -ub 4096 -ngl 99 --port 8080 Prefill: 1293.69 t/s Decode: 75.00 t/s VRAM: 23.93 GB (Maxed out on card) the dflash gguf is additional 1.6 GBs # The Architecture Insight (Muse Glimmer vs. Gemma 4 31B) If you look at my Gemma 4 31B tests from last week, getting 140k context required heavily degrading the memory with Q4 KV quantization (gemma 31b q4 can do only about 40k context with unquantized kv on a 24gb card). That "unzipping" overhead bottlenecked Gemma's MTP decode speeds down to 65 t/s. Muse Glimmer completely sidesteps this bottleneck. By using aggressive 16:1 GQA, it keeps the KV cache in native f16 format at massive context lengths. Flash Attention gets to run at maximum uncompressed speed, letting the DFlash drafter push decode safely to 75 t/s without compute lag. With a 76% on SWE Bench Verified and seamless local tool calling, this model looks promising. Unsloth's Hugging Face GGUF links, intelligence/agentic benchmark details, and inference throughput performance graphs are posted in the replies. For 24GB rig, what’s your current go to model?

Alok

65,480 次观看 • 16 天前

🚨 I TOLD YOU THIS WAS COMING! Absolute bloodbath: $2,880,000,000,000 wiped out from the US stock market in just 2 HOURS Let that number sink in. Two TRILLION dollars. Gone! In the time it takes to watch a movie. And if you've been following me, none of this surprises you. Because I told you that S&P 500 was walking into a trap. No panic. Just the setup, laid out in plain sight. Today the trap snapped shut. So let me remind you WHY it's happening, in 3 forces: Stacked on top of each other. Every 4 years 1. THE PRESIDENTIAL CYCLE Year 1-2 is always pain. Tax hikes, budget cuts, "cleaning up the last guy's mess." We're standing in the pain right now. 2. THE MIDTERM TRAP 15 out of the last 16 midterm years, the market bled from May to October. A 93.75% hit rate since 1962. Big money knows it. So they sell in May and hold cash through the summer. Guess what month we just walked out of. 3. THE FED'S TIMING PROBLEM The Fed always hikes early, so it has room to cut before re-election. Translation: midterm year always lands at peak rates. Exactly when the economy can't take it. Now stack 2026 on top: ➮ Inflation at a 3-year high ➮ 10Y yield breaking 4.50% ➮ Mortgage rates pushing 7% This wasn't a forecast. It was a setup hiding in plain sight, and today it paid out in $2.8 TRILLION of other people's money The next few days are going to get violent. But don't worry. That's exactly what I'm here for. Like it or not, I've called the exact market tops and bottoms for 11 years My system flags the moment the market shifts from caution to DANGER You will be warned before it hits, like always. Not following me will be your #1 mistake of 2026, soon you'll understand why

Reflection🪩

76,016 次观看 • 2 个月前

Universities and High Schools have not moved rapidly enough to guide students to have skills for the next decade. THEY HAVE FAILED. It is a massive crisis that can be averted by understanding what AI and Robotics will bring about. Solutions are knowing how to use these tools and new industries that will rise. But this situation is also on ALL OF US. No “job” is safe from founder to entry level in most industries. You and I, by what we do, will be “replaced” ultimately. What to do? AI and Robotics are tools, the next decade is owned by those who know how to use them expertly, but this is also temporary. We have to understand that what we do for “work” will change giving ultimately a greater value to those that are: Creative Flexible Always learning Willing to be wrong Love being human Love being alive Know history Covet wisdom Knowing all tech has downsides Building strong family and friends Realize many institutions have failed The first four are required for you to be able to live through this period with your sanity intact. The rest will allow you to thrive. There are no true careers at this point anymore. There are advocation and vocations which will either earn you money or give life meaning. We will learn that we are not “what we do”, just like we knew for 99% of human existence. Let that sink in. — You and I are far, far ahead of knowing this and we can do two things: 1) Laugh at the “clueless” 2) Help people understand with grace Go to Reddit if you are 1, in fact don’t follow me because you will not like this next decade and what I post. You are 2 and thank you. Even if you and I have not solved this issue, we can help people understand what is ahead and with determination and creativity bound together to solve it locally. Or human family has done this millions of times. The evidence is: you are here. The Neo Luddite movement has not even begun and it will potentially rip apart society even more than all the fashionable moment in the recent past has. These Luddites will have a good point with the wrong answers cooked up by dying academics that cling to labels, “virtues” and victim hood. It will be readymade for some governments to enter in as “big daddy” to “help us”. You will not like what they do, but you will only know when it is too late. It will include YOU “volunteering” to “leave” by 60, to “help out” CanadaPod style. “Brian, I’m 24 what do I do?”. I hope to do much more here to help. But I do know this: 1) Learn a trade or vocation because it’s valuable. It may also be free to low cost if you do it right. 2) Learn everything you can about USING AI and TRAINING YOUR AI. Your expertise will be in the top 1% for a decade. But not forever. 3) Understand Bitcoin and how it will rise while other things sink. This is a short list for now. We will know more moving forward. When you see videos like this posted below, know one thing: Many of these folks had no real family of mental and physical support. Maybe no parent or one parent. Maybe only a broke system to prepare them for—nothing. This was not their doing. Now it is not your “job” to help them, it is your survival to help them if that is what you need. See some day after the dust settles these 20 year olds will be 40 year olds and running YOUR world. And at some point you may need them more than you think you do. You will need them, as they need you now. THIS IS WHAT PAST WISDOM KNEW. The elders of the past never found the need to piss on the youth and hope for the best. THE YOUTH ARE OUR BEST, let us all find ways to change it, even if every aspect of “the system” wants us to berate them into the ground.

Brian Roemmele

37,681 次观看 • 1 年前