Loading video...

Video Failed to Load

Go Home

👨‍🍳The longest continuous shot I have ever done🐻‍❄️ 2 clips: First: 909 frames in 12 context windows. Prompt executed in 00:39:25 Second: 2337 frames in 31 context windows. Prompt executed in 01:32:02 Max allocated VRAM=27.458 GB 720p resolution Need 260gb of System RAM though

56,900 views • 10 months ago •via X (Twitter)

0 Comments

No comments available

Comments from the original post will appear here

Related Videos

A Physics That Could Finally Unify Reality (Part 2/2) - James Ellias, DemystifySci #385 A quiet tremor sounds beneath the floorboards of physics, the still-living heartbeat of the forgotten search for the hidden substance that carries every wave and whisper of the universe. We walk through the fog of equations and theories of fundamental physics with James Ellias of James Ellias, and ask if there is a material truth lies beneath the symbols, or if we have to be satisfied with the short-sighted vision of mathematics alone. 00:00 Go! 00:04:37 Central Equations in Physics 00:08:20 The Importance of a Medium in Physics 00:10:00 Empowerment Through Understanding Physics 00:12:36 Rationality and Truth in Society 00:16:31 Existence vs Consciousness 00:20:30 Discussion on Existence and Consciousness 00:24:25 Role of Imagination in Existence 00:28:10 Properties and Entities in Physics 00:30:15 The Nature of Aether and Physical Mediums 00:36:08 Clarity and Understanding in Physics 00:40:26 Discussion on Force and Aether 00:44:57 Relationship of Entities and Actions 00:49:30 Theoretical Framework for Aether 00:58:04 Exploration of Aether Theories 01:00:40 Importance of Conciseness in Communication 01:02:00 Understanding Physics Before Proposing Hypotheses 01:05:15 Reevaluation of Flawed Theories 01:09:35 The Evolution of Key Physics Concepts 01:15:19 Context-Specific Nature of Constants 01:19:43 Historical Context in Physics 01:21:36 The Nature of Electrons 01:25:53 J.J. Thomson’s Evolving Perspective 01:29:24 Upcoming Work and Philosophical Frameworks 01:32:11 Collaborations

Anastasia

13,618 views • 9 months ago

The VRAM barrier is officially dead. I just ran Qwen 3.8 Flash Next (MoE) 125B A6B with a 250,000 context window on a single 24GB RTX 4090. 21 tokens/sec decode. 364 t/s prefill. no mtp. no dflash. no kv cache quantization! We are running datacenter models on consumer hardware. Tested on Ubuntu 22 | CUDA 13.0 | PCIe 4.0 x16 | 110 GB DDR4 System RAM with a continuous 28k prompt across all runs. ### The Benchmarks & Scaling # 1. Hybrid Offload (-ncmoe 40 @ 80k Context) Offloaded 40 expert layers to the GPU, pushing VRAM to the ceiling. ./build/bin/llama-server -m Qwen3.8-Flash-Next-UD-Q4_K_XL-00001-of-00004.gguf -c 80000 --port 8080 -v --fit off -b 4096 -ub 4096 -ncmoe 40 Prefill: 383.85 t/s | Decode: 22.52 t/s Footprint: 23.85 GB VRAM | 97 GB RAM # 2. Full CPU MoE Offload (-cmoe @ 80k Context) Pinned all 512 expert layers to DDR4 RAM (-cmoe), keeping attention on the 4090. llama.cpp flags: (Same as above, replace -ncmoe 40 with -cmoe) Prefill: 355.72 t/s | Decode: 20.84 t/s Footprint: 11.66 GB VRAM (12GB+ VRAM freed up!) | 110 GB RAM # 3. The 180,000 Context Run Prefill: 357.75 t/s | Decode: 20.98 t/s | VRAM: 15.6 GB | RAM: 110 GB # 4. The 250,000 Context Absolute Ceiling ./build/bin/llama-server -m Qwen3.8-Flash-Next-UD-Q4_K_XL-00001-of-00004.gguf -c 250000 --port 8080 -v --fit off -b 4096 -ub 4096 -cmoe Prefill: 364.29 t/s | Decode: 20.97 t/s Footprint: 18.3 GB VRAM (Still ~5.7 GB of VRAM headroom!) | 110 GB RAM ### Key Insights: -b 4096 -ub 4096: doubles the prompt ingestion from ~150 to 364+ t/s. -cmoe Free Lunch: Shifting expert layers to DDR4 RAM slashes VRAM from 24GB to 11.6GB with virtually zero decode penalty (22.5 -> 20.9 t/s), enabling the 250k context ceiling. Qwen 3.8 Flash-Next (UD-Q4_K_XL) is a massive 111.4 GB model split across 4 shards. To run this architecture, you must build from the experimental PR branch (#27742) by Daniel Han: git clone && cd llama.cpp git fetch origin pull/27742/head:qwen-next && git checkout qwen-next cmake -B build -DGGML_CUDA=ON -DCMAKE_CUDA_ARCHITECTURES=native -DBUILD_SHARED_LIBS=OFF cmake --build build --config Release -j $(nproc) --target llama-server A single 4090 paired with 100 GB of cheap DDR4 RAM will comfortably serve production grade 125B inference. While Qwen 3.8 27B (dense) still holds the crown for single 3090/4090 rigs, Flash Next proves 125B hybrid models are officially viable on consumer hardware. Hugging Face GGUF link and complete performance telemetry graphs are dropped in the replies below. GLM 5.3 Flash VS Qwen 3.8 Flash Next, which one takes the open weights crown this week?

Alok

999,150 views • 8 days ago

I just crammed the updated Gemma 4 26B A4B QAT (MoE) with 180k context into an 8GB RTX 4060 (8 GB VRAM + 16 GB RAM only!!) and optimized the batch size. 23 tokens/sec decode, 300 tokens/sec prefill Yesterday I showed you a Gemma 4 31B dense model running flawlessly on an RTX 4090. Today, we're breaking the VRAM bank on a budget card using Unsloth’s new Gemma 4 26B (A4B) QAT quants. Following Google’s chat template update that boosted agentic benchmarks by +10%, I pushed this model to its absolute limits. Here is how you squeeze 250k context out of 8GB of VRAM. # The Setup & The Optimization - Hardware: Nvidia RTX 4060 (8GB VRAM) + 16GB System RAM - Environment: CUDA 13.0 build of llama.cpp - Model: gemma-4-26B-A4B-it-qat-UD-Q4_K_XL.gguf - Prompt: 28,000 tokens of prompt for each run If you read my L2 cache breakdown (attached in replies), you know the 4060’s 24MB cache maxes out at `-b 1024 -ub 1024`. Push past that, and prefill crashes. I locked those flags in for every test below to ensure maximum GEMM throughput. # 1. The Raw Context Push (Unquantized KV Cache) First, I wanted to see how far pure 8GB VRAM + 16GB RAM could stretch without touching the KV cache: - 80k Context: Prefill 385 t/s | Decode 25.5 t/s - 120k Context: Prefill 270 t/s | Decode 24 t/s llama.cpp flags: .\llama-server -m gemma-4-26B-A4B-it-qat-UD-Q4_K_XL.gguf -c 120000 --port 8080 -ub 1024 -b 1024 Without KV quantization, 120k is your hard ceiling. push past that prefill throughput drops off a cliff, making the model practically unusable for large agentic workloads. # 2. The Q8 KV Cache Lifeline To survive 250k context on a budget card, you have to quantize the KV cache. I enabled 8 bit KV cache (`-ctk q8_0 -ctv q8_0`) and re ran: - 180k Context: Prefill 280 t/s | Decode 22.8 t/s - 250k Context: Prefill 115 t/s | Decode 20 t/s llama.cpp flags: .\llama-server -m gemma-4-26B-A4B-it-qat-UD-Q4_K_XL.gguf -c 180000 --port 8080 -b 1024 -ub 1024 -ctk q8_0 -ctv q8_0 Result: Q8 KV cache brings 250k context back from the dead. Decode speed stabilizes at a highly usable 20 t/s. You are trading a very small bit amount of reasoning precision for an extra 130,000 tokens of context window. if you own a single rtx 3050, 3060, 3070, 4050, 4060, 5050 or 5060, you must try this model and optimize your batch size for higher prefill. Hugging Face links to the updated Unsloth's QAT quants and performance graph are in the replies below. What model are you running on your 6GB, 8GB or 12GB cards right now? Let's see your setups.

Alok

36,617 views • 1 month ago

Deepseek V4 Flash 0731 (Q2) - 12 tokens/sec - Single RTX 4090 - 650+ tokens/sec prefill - 250k context - no kv cache quantization! DeepSeek just dropped the official V4 Flash 0731 two days ago with a massive agent capabilities upgrade. The official benchmarks are literally crushing their own V4-Pro-Preview on agentic tasks like Terminal Bench 2.1 and DeepSWE. Unsloth AI said they couldn't wait to bring it to local devices, and they delivered. If you thought my 118B Poolside Laguna S 2.1 MoE run last week on a single GPU was wild, hold onto your hardware. I just successfully ran Unsloth’s brand new 91GB DeepSeek-V4-Flash-0731 (UD-IQ2_M) GGUF entirely locally. And I pushed it to a mind-bending 250,000 context window. The VRAM ceiling is an illusion if you know how to optimize llama.cpp. Here are the benchmarks and the cheat codes to run a local frontier class model yourself. For the hardware and setup, I used a single NVIDIA RTX 4090 (24GB VRAM) hooked up via a PCIe 4 bus, running Ubuntu 22.04 LTS and CUDA 13.0. You don't need a massive enterprise server for this, if you have more than 80 GB of standard DDR4 RAM and a 24GB card like an RTX 3090 or 4090, you can run this exact stack yourself. All benchmarks were run using a massive 28k token prompt to truly stress test the prefill limits. no kv cache quantization THE BENCHMARKS (Scaling Context): # 80k Context (Baseline: -b 2048 -ub 2048): Prefill: 465.43 t/s | Decode: 13.00 t/s | VRAM: 22.87 GB # 80k Context (Optimized: -b 4096 -ub 4096): Prefill: 643.15 t/s | Decode: 12.20 t/s | VRAM: 23.00 GB (Notice how doubling the batch flags spiked my prefill throughput by nearly 200 t/s with almost zero VRAM penalty) # 180k Context (-b 4096 -ub 4096): Prefill: 629.18 t/s | Decode: 11.92 t/s | VRAM: 23.40 GB # 250k Context MAXIMUM (-b 4096 -ub 4096): Prefill: 619.02 t/s | Decode: 11.54 t/s | VRAM: 23.40 GB # THE SECRET SAUCE (Why this works): Unsloth’s UD-IQ2_M quant is ~91GB across 3 files. Since I only have 24GB of VRAM, the PCIe 4 bus and system RAM have to do the heavy lifting. The magic bullet is the --no-mmap flag. By completely bypassing OS disk paging, I forced llama.cpp to load the massive model weights directly into the system RAM upfront. Combined with Flash Attention (-fa on) and exactly 12 CPU threads (--threads 12), I maintained an incredibly stable 11.5+ tokens/sec decode speed even at a quarter million token context. # THE EXACT COMMAND: ./build/bin/llama-server -m /workspace/models/DeepSeek-V4-Flash-0731-UD-IQ2_M-00001-of-00003.gguf -c 250000 -fa on --port 8080 --threads 12 -b 4096 -ub 4096 --no-mmap -v Local conversational and agentic coding AI is fully here. You don’t need an API or an H100 cluster. Qwen 3.8 27b drops next week making the 24GB VRAM tier even more worthwhile. What does your current local AI rig look like, and what's the craziest model you've managed to squeeze into it? Official huggingface GGUF links from Unsloth and performance graphs are dropped in the replies below!

Alok

45,967 views • 1 month ago

Uncovering Aboriginal Australia: A Conversation with Mungo Manic Detailed timestamps 👇 Origins and Early Human Migration 00:00 Introduction to Mungo Manic and early Australians 00:55 Who is Mungo Manic 03:41 Terminology and definitions of "Aboriginal" 04:47 Terminology preferences in scientific contexts 05:29 Origins of the word "Aboriginal" and its historical usage 07:32 The term "Aboriginal" losing a genetic or racial definition in the 1960s or 1970s 09:10 The three-part test of Aboriginal identity Early Human Settlement and Fossil Evidence 12:55 How and when humans first arrived in Australia 16:40 Early skeletons, including Mungo Man 21:00 The study of Mungo Man 23:35 Did Denisovans make it to Australia 25:20 Fossil degradation in Lake Mungo Isolation, Cultural Practices, and Environmental Adaptation 26:00 Isolation of pre-colonial Australia and the ancestry of dingoes 33:15 Aboriginal women breastfeeding dingoes and their importance as pets 36:35 The taboo against eating dog in Parsi culture Social Structures and Survival Practices 37:30 Infanticide in Aboriginal Australian culture 40:55 Cannibalism in Aboriginal Australian culture including endocannibalism and exocannibalism 42:14 Warfare and violence in early Australian life 42:48 The nature of warfare and violence in ancient Australian societies 44:20 Christophe Darmangeat's expertise on Aboriginal warfare 45:39 Capital and corporal punishment in Aboriginal Australia 49:28 The lack of rights for Aboriginal women 51:30 Sorcery and the practice of pointing the bone Technology, Environment, and Innovation 54:00 Technological development in Aboriginal Australia and Jared Diamond’s theory 56:23 Was Australia actually a harsh environment 58:43 Technological simplicity in Tasmania after separation from mainland Australia 01:02:00 Competition and innovation 01:03:45 Kimberley spear points, Sulawesi, and Macassan sailors 01:05:30 The oldest known rock art Cultural Practices and Initiation Rituals 01:06:45 Cultural practices and initiation rituals 01:07:30 Initiation practices including scarification, tooth evulsion, and circumcision 01:12:20 Subincision 01:14:00 Female circumcision and virginity in Aboriginal Australia 01:20:00 The Fatal Shore by Robert Hughes, wife-sharing, and sexual practices in Aboriginal culture 01:24:46 The practice of clubbing women 01:25:35 The resilience and immune system of Aboriginal Australians Impact of Humans on the Environment 01:27:08 Human impact on the Australian landscape and megafauna extinction 01:31:00 Boomerangs used in warfare 01:33:50 The film Ten Canoes Religious and Spiritual Beliefs 01:35:20 Religious beliefs and spiritual practices including the Dreaming 01:40:00 Aboriginal eschatology 01:42:00 The extinction of the people of Kangaroo Island 01:43:00 Mourning periods and mummification 01:45:42 The cultural practice of not mentioning the dead Modern Aboriginal Identity and Historical Debates 01:48:04 Are Tasmanian Aboriginals extinct Was it genocide The Henry Reynolds black armband theory vs Keith Windschuttle 01:53:53 Defining Aboriginal identity today racial definitions vs self-identification 01:56:55 Blue-eyed blonde-haired Aboriginal people vs Tasmanian Aboriginals 02:00:52 The importance of Tasmanian DNA research 02:03:22 A critique of Bruce Pascoe’s Dark Emu and its depiction of ancient Australians as agriculturalists Colonial Impact and Unanswered Questions 02:11:27 Kidnappings 02:15:40 The need for anonymity in controversial discussions 02:29:27 Biggest mysteries in Aboriginal history

Quillette

83,061 views • 1 year ago