Video yükleniyor...

Video Yüklenemedi

Ana Sayfaya Dön

Two models. Zero rivals.​ Total dominance. DBX707 or DBX S. DBX. No Equals.​ Discover: #AstonMartin #DBX #NoEquals #POWERDRIVEN

15,177 görüntüleme • 9 ay önce •via X (Twitter)

0 Yorum

Yorum bulunmuyor

Orijinal gönderinin yorumları burada görünecek

Benzer Videolar

$10,000 invested with Jim Cramer in 2001 turned into $28,997 by 2020. The same $10,000 in a free S&P 500 index fund? $44,957. Jim Cramer's CNBC Investing Club is TRASH... But they're still gassing him up. Wharton School researchers tracked Cramer's Action Alerts PLUS portfolio for over 17 years. The verdict: 3.38% annualized returns versus the S&P 500's 5.59%. More volatility. Lower returns. A Sharpe ratio of 0.11 versus the index's 0.24. You were taking on MORE risk for LESS reward. But it gets even worse. From inception through June 2023, cumulative total return: 230%. The S&P 500 with dividends reinvested: 453%. From 2016 to 2022, Action Alerts outperformed the market exactly ONCE. And you were paying $399 a year for the "privilege". Here's what really bothers me: CNBC sells this product to retail investors who trust the brand. Regular people. People without decades of experience to know better. They watch the flashing lights, the sound effects, the screaming - and assume there's actual substance behind this. There isn't. After adjusting for market risk, Cramer's excess return was "essentially zero." Not alpha. Noise. And SVB? First Republic? February 2023: Cramer told viewers to BUY Silicon Valley Bank. One month later it collapsed. He called First Republic "a very good bank." It failed two months after that. Catastrophic calls on systemically important institutions. Delivered to retail investors with total confidence. I've spent 45 years in this business. I worked for Peter Lynch. I know what genuine investment analysis looks like. What CNBC sells is infotainment dressed as financial advice. The index fund is free. It outperforms. Remember that the next time someone tries to sell you a $399 subscription to underperform the market. Say NO to Cramer and CNBC.

George Noble

50,661 görüntüleme • 5 ay önce

Deepseek V4 Flash 0731 (Q2) - 12 tokens/sec - Single RTX 4090 - 650+ tokens/sec prefill - 250k context - no kv cache quantization! DeepSeek just dropped the official V4 Flash 0731 two days ago with a massive agent capabilities upgrade. The official benchmarks are literally crushing their own V4-Pro-Preview on agentic tasks like Terminal Bench 2.1 and DeepSWE. Unsloth AI said they couldn't wait to bring it to local devices, and they delivered. If you thought my 118B Poolside Laguna S 2.1 MoE run last week on a single GPU was wild, hold onto your hardware. I just successfully ran Unsloth’s brand new 91GB DeepSeek-V4-Flash-0731 (UD-IQ2_M) GGUF entirely locally. And I pushed it to a mind-bending 250,000 context window. The VRAM ceiling is an illusion if you know how to optimize llama.cpp. Here are the benchmarks and the cheat codes to run a local frontier class model yourself. For the hardware and setup, I used a single NVIDIA RTX 4090 (24GB VRAM) hooked up via a PCIe 4 bus, running Ubuntu 22.04 LTS and CUDA 13.0. You don't need a massive enterprise server for this, if you have more than 80 GB of standard DDR4 RAM and a 24GB card like an RTX 3090 or 4090, you can run this exact stack yourself. All benchmarks were run using a massive 28k token prompt to truly stress test the prefill limits. no kv cache quantization THE BENCHMARKS (Scaling Context): # 80k Context (Baseline: -b 2048 -ub 2048): Prefill: 465.43 t/s | Decode: 13.00 t/s | VRAM: 22.87 GB # 80k Context (Optimized: -b 4096 -ub 4096): Prefill: 643.15 t/s | Decode: 12.20 t/s | VRAM: 23.00 GB (Notice how doubling the batch flags spiked my prefill throughput by nearly 200 t/s with almost zero VRAM penalty) # 180k Context (-b 4096 -ub 4096): Prefill: 629.18 t/s | Decode: 11.92 t/s | VRAM: 23.40 GB # 250k Context MAXIMUM (-b 4096 -ub 4096): Prefill: 619.02 t/s | Decode: 11.54 t/s | VRAM: 23.40 GB # THE SECRET SAUCE (Why this works): Unsloth’s UD-IQ2_M quant is ~91GB across 3 files. Since I only have 24GB of VRAM, the PCIe 4 bus and system RAM have to do the heavy lifting. The magic bullet is the --no-mmap flag. By completely bypassing OS disk paging, I forced llama.cpp to load the massive model weights directly into the system RAM upfront. Combined with Flash Attention (-fa on) and exactly 12 CPU threads (--threads 12), I maintained an incredibly stable 11.5+ tokens/sec decode speed even at a quarter million token context. # THE EXACT COMMAND: ./build/bin/llama-server -m /workspace/models/DeepSeek-V4-Flash-0731-UD-IQ2_M-00001-of-00003.gguf -c 250000 -fa on --port 8080 --threads 12 -b 4096 -ub 4096 --no-mmap -v Local conversational and agentic coding AI is fully here. You don’t need an API or an H100 cluster. Qwen 3.8 27b drops next week making the 24GB VRAM tier even more worthwhile. What does your current local AI rig look like, and what's the craziest model you've managed to squeeze into it? Official huggingface GGUF links from Unsloth and performance graphs are dropped in the replies below!

Alok

42,894 görüntüleme • 8 gün önce