Video wird geladen...

Video konnte nicht geladen werden

Zur Startseite

"Cameras as Relative Positional Encoding" TLDR: comparison for conditioning transformers on cameras: token-level raymap, attention-level relative pose encodings, a (new) relative encoding Projective Positional Encoding -> camera frustums, (int|ext)insics for relative pos encoding

17,833 Aufrufe • vor 1 Jahr •via X (Twitter)

0 Kommentare

Keine Kommentare verfügbar

Kommentare vom Original-Post werden hier angezeigt

Ähnliche Videos

I spent and hour of my Saturday reviewing hundreds of charts. These are the setups that stood out and what you should focus on this week Friday changed the tone of this market. The AI trade is under pressure. Software is pulling back. Relative strength is starting to stand out. $GOOGL held up. $AAPL barely cracked. $C continues to show strength while growth stocks unwind. Here’s the watchlist and recording: $SPX: One of the ugliest days we've seen in months. Closed near the lows after breaking the 20-day. 7330-7290 is the first support zone. Below that opens 7273 and potentially 7150. $QQQ: Nearly 5% down on Friday. AI leadership is under pressure. Watching 695 support closely. $IWM: Back to 280 support. Watching whether this becomes a swing low or just another bounce that gets sold. $BTC: Still under pressure. Failed reclaim of the 200-day. No clear setup here. $SMH: Nearly 9% down Friday. Semis finally cracked. Watching for either a relief bounce or continuation lower. $MSFT: Failed after briefly reclaiming the 200-day. Still holding trend support but needs buyers soon. $AAPL: One of the stronger mega caps. Technical damage is limited compared to the rest of the market. Worth watching. $GOOGL: One of the better-looking charts. Holding the earnings gap and showing relative strength. Above 373 could trigger a relief move. $AMZN: Broke the 50-day and looks vulnerable. Could see a move toward the 200-day near 232. $NVDA: Momentum has faded. Sitting on the 50-day near 203. Must hold. $TSLA: Significant technical damage. Lost the 200-day, 50-day, 20-day, and 9-day. Needs major repair work. $META: Still holding the lower end of its range. 600 remains the key level. $AMD: Looks like it wants to fill the gap lower. Semis remain under pressure. $AAPL: Relative strength remains notable. One of the few mega caps still acting well. $NFLX: Quiet relative strength. Not an easy trade, but worth noting. $LLY: Strong healthcare leadership. Above 1165 opens another attempt at highs. Must hold 1100. $JPM: Financials are starting to show relative strength. $C: One of the stronger bank charts. Pullback remains very controlled. $WFC: Held up well and continues to show relative strength. $GS: Large engulfing pullback. Watching for stabilization. $GE: Rotational strength worth monitoring. $CROX: Continues to hold the 9-day and trend higher. Relative strength stands out. $SNOW: Pulling back into the 9-day after earnings. Watching for support. $DDOG: Pulling back with software but still one of the stronger charts in the group. $PLTR: Rejected at the 200-day. Needs more work. $IBM: Back below the 9-day. Harder chart for now. $DELL: Pulling back into the 9-day after earnings. Watching for buyers to step in. $HOOD: Pulling back into range support. $CRWD: Watching 670 as a potential support area after earnings. $NET: Backtesting the 9-day. One of the better software recovery stories. $BE: Still consolidating near highs. No major damage yet. $MU: Sharp pullback. Watching for a bounce near current levels. $WDC: Big pullback after a huge run. $SNDK: Pulling back but no major technical damage yet. Watching closely. Overall theme: Friday changed the character of the market. The focus shifts from chasing momentum to identifying what held up during the selloff. $GOOGL, $AAPL, $C, $WFC, $CROX, and select software names are showing the best relative strength. For now, caution is warranted. Let the market prove it wants to bounce before getting aggressive.

spacemonkey

37,634 Aufrufe • vor 3 Monaten

.Luke Rosiak talks about a Medicaid fraudster who flaunts on social media a life of private jets, yachts, etc. "These are people that came here as refugees & now they’re living as millionaires [via fraud]." ROSIAK: "Whistleblowers in the [Columbus] area tell me companies knock on doors in ethnic neighborhoods, where most people are on Medicaid & most people live in multi-generational households, & they tell the older family member to go to particular doctors & claim particular symptoms, & then they will put the younger family member on their payroll with the parent as their only patient. "Now that triggers payments of up to $90,000 a year, which the company splits with the younger relative & the younger relative splits with the older relative. 81 percent of Bhutanese are on welfare, a fact that some of the remainder use to become obscenely rich. And I had a story just yesterday, a 29-year-old named Roshan Adhikari works at his dad’s home health-care company. His social media documents a life fit for a rapper or a movie star. He’s drinking champagne on private jets, he’s on his yacht, he’s making movies starring himself. This was funded — his dad — his dad has this company that was paid $17 million for taking care of old people. Well he’s got it apparently another full-time job, he’s doing this on the side. The dad hasn’t paid taxes in, like, five years, he’s got tax liens in Ohio at the state level in a Cuyahoga County. And the 25-year-old brother has his own home health care company that got $10 million. So these are people that came here as refugees & now they’re living as millionaires & they’re getting the money because the other refugees are all on Medicaid & then are claiming to be very sick."

Tom Elliott

72,258 Aufrufe • vor 3 Monaten

I coded a Speech-to-Text model from scratch. 𝐇𝐞𝐫𝐞 𝐢𝐬 𝐭𝐡𝐞 𝐛𝐥𝐨𝐠 𝐟𝐨𝐫 𝐭𝐡𝐞 𝐬𝐚𝐦𝐞: No APIs. No pre-trained models. Just PyTorch, an A100 GPU, and hours of debugging. This started months ago. I wanted to understand how machines hear. Not surface-level understanding. I wanted to build the whole thing myself. So I built it piece by piece: autoencoders, VAEs, VQ-VAEs, Residual Vector Quantization, and CTC loss. Each one took days to get right. Trained for 3 hours on 13,100 audio clips. Got complete garbage. Changed the tokenizer from BPE to character-level. Rechecked everything. Asked AVB who built STT models before. His answer: these models are tricky to train and need days of compute, not hours. Cut the dataset to 200 clips. After 2 hours, actual words appeared. Overfitted? Absolutely. But watching noise turn into recognizable English was satisfying. I have made a blog about this as well so you can learn about the same and my process - Audio fundamentals and waveform representation - Why attention breaks on raw audio - Convolutional downsampling - Transformer encoder with positional encoding - Vector Quantization, straight-through estimator, and RVQ - CTC loss and greedy decoding - Full training loop with VQ loss warmup - What went wrong and what finally worked Resources: - Blog: - Code: More Resoures CTC loss AVB videos SoundStream Paper LJ speech dataset wav2vec paper RVQ blog Next up: I've already trained two TTS architectures from scratch. Video post about those coming soon. But first, I'm dropping a visual breakdown of Vision Transformers, covering how they work and how to fine-tune them. Follow me Mayank Pratap Singh you're into audio deep learning. Repost so others can find this

Mayank Pratap Singh

51,382 Aufrufe • vor 6 Monaten

I spent my Saturday mapping the setups that matter most for next week. The market is still searching for direction 👇 Semiconductors are losing momentum. Software is holding up. Select mega caps are quietly starting to reclaim relative strength. Here’s the watchlist and recording: $SPX: Failed breakout above 7530 and reversed sharply. Still trading inside a tightening triangle. Holding 7400 keeps the range intact. Reclaiming 7500 would shift the bias back higher. Losing 7400 opens 7300, then 7240. $QQQ: Failed breakout and back inside its range. Watching for either a reclaim of the highs or a break lower from consolidation. $IWM: Still one of the stronger indices. Broke to fresh all-time highs this week. Extended short term, but the trend remains constructive. $BTC: Showing relative strength versus equities but still below the 200-day moving average. No trade for now. $AAPL: One of the strongest charts on the board. AI chip headlines and foldable iPhone reports fueled the move. Watching a break of Friday’s highs toward all-time highs. $MSFT: Quietly improving. Reclaimed the 9-day and 20-day moving averages. Above 400 becomes much more interesting. $NFLX: Failed breakdown continues to work. Last week’s 75 reclaim is now extending higher. Still one of the stronger mega-cap charts. $NVDA: Semiconductors continue to weaken. Rejected the 9-day and drifting back toward the 200-day. Better opportunities elsewhere for now. $TSLA: Ugly reversal despite beating delivery expectations. Rejected 432 and finished near 392. Needs 400 reclaimed before becoming interesting again. $AMZN: Quiet relative strength. Still a difficult trade, but worth monitoring if mega caps continue to rotate higher. $GOOGL: Holding up better than many peers but still lacking a clean technical trigger. $SMH: Losing momentum. Watching closely as semiconductors continue to weaken. $MSTR: Bitcoin strength is helping. Worth keeping on the radar. $ADBE: Quiet recovery underway. Watching 225. $COIN: Showing improving relative strength alongside crypto. $HOOD: Expansion into 30 European countries fueled the recent rally. Watching 120 closely after today’s rejection. $PLTR: Failed breakdown continues to recover. Holding above 125 keeps the recovery thesis alive. $MA: Strong recovery back above the 200-day moving average. $V: Similar setup to Mastercard. Quiet accumulation after months of weakness. $PYPL: Improving technically. Worth monitoring if payment stocks continue rotating higher. $XYZ: Last week’s breakout worked almost perfectly. Above 82 opens another leg higher. $LLY: Healthcare leadership remains intact. Watching the bull flag and eventual breakout above 1238. $CRWD: Holding up well after the split. Watching the 200 level closely. $BAC: One of the stronger financials. Watching for a breakout into new highs. $GS: Weak relative to peers. $C: Also lagging despite strength elsewhere in financials. $WFC: Pulling back while BAC leads. $PANW: Strong continuation this week. Watching 360 after a brief pause. $DDOG: Still constructive but needs a breakout from the current lower-high structure. $NET: Similar setup to DDOG. Watching for software continuation. $MDB: Holding 365. Could become interesting if software stays strong. $HIMS: Quietly grinding higher. $UNH: Strong recovery. Watching 430. $AVGO: Breaking below the 200-day moving average. Semis continue to deteriorate. $MU: Lost the 1000 level after an impressive run. Momentum cooling. $SNDK: Continued downside after losing 2000. 1850 remains the next major area. $LRCX: Another semiconductor under pressure. $MRVL: Watching for either a failed breakdown or continuation lower. Overall theme: Software continues to hold up better than semiconductors. Several mega-cap names are quietly showing improving relative strength. Financials are becoming more selective. The tape still rewards patience more than aggression. AAPL, MSFT, NFLX, BAC, XYZ, LLY, and CRWD are some of my favorite charts going into next week.

spacemonkey

18,726 Aufrufe • vor 2 Monaten

TWI - $MSTR Trade Recap 📈💡 Today’s trade on $MSTR was an incredible opportunity, prepared with “A” sizing. Members received premarket notes, audio analysis, and all executions were handled live on video stream. Here’s a recap of how we approached this trade today in Twinsight - Pro 🔥 Execution Highlights ✅ • 400C: Entry $17.00 → $58.65 (+245.00%) (Closed at $23.80) • 430C: Entry $10.90 → $37.32 (+242.20%) (Closed at $17.10) • 450C: Entry $9.52 → $27.58 (+189.72%) (Closed at $17.10) Key Factors Discussed: 1. Relative Strength on Higher Time Frames: $MSTR showed r/s compared to its peers 2. Short Interest: Elevated levels added fuel to the move. 3. Recent Offering News: Provided a catalyst for volatility and opportunity. 4. Critical Levels: • 400 Psychological Level • 383 Break/Trend Line Held Strong 5. Elevated Relative Volume (RVOL): Confirmed conviction in the move. Additional Notes • Expected Value (EV): Probability relationship with the turn off the 383 level. • EMA Trail: Used alongside 5-minute candle analysis to manage the trade as price established a slower pace over 400. Reflection: Studying “top opportunities” over the last few years has given Matae the confidence to identify unusual opportunities like this and adjust sizing appropriately. $MSTR was nothing short of amazing this week, and it may not be done yet. Check out the recording clip and executions from today! If you have any questions or feedback, feel free to reach out. Let’s keep crushing it! 🚀 Are you Ready To Trade With Insight? 💡 Join us ✅ $SPX $SPY $QQQ $IWM #Bitcoin #Btc

Trade With Insight | John

30,524 Aufrufe • vor 1 Jahr

🚀 Sol-H3: MiniMax (official) H3 Video Generation Faster Than Playback 🤩 Five seconds of world. 1.653 seconds to infer. We’re releasing Sol-H3, our fastest end-to-end MiniMax-H3 inference stack yet. On one 8× NVIDIA B300 Blackwell system, it generates five seconds of 1344×768 video with stereo audio in 1.653 seconds. Across 1×, 4×, and 8× B300, Sol-H3 reaches up to a 15.54× speedup versus Base H3. Compared with 50-step Base H3 Dense on the same 8× B300 system, the four-step Sol-H3 profile delivers: • 5s: 18.250s → 1.653s (11.04×) • 10s: 50.660s → 3.732s (13.57×) • 15s: 99.513s → 6.612s (15.05×) Sol-H3 also scales across GPU counts: • 4× B300: 2.918s / 6.993s / 12.542s for 5s / 10s / 15s (12.11–15.54×) • 1× B300: 13.745s / 37.813s / 52.260s for 5s / 10s / 15s (9.45–14.29×) All figures are medians of three measured runs after one warmup at 1344×768 and 24 FPS with stereo audio. Base H3 uses 50 scheduler points (49 DiT forwards); Sol-H3 uses four DiT forwards, so this is a full-profile comparison—not an attention-only runtime change. Sol-H3 uses Dense attention on 1× B300 and SOL with INT8 QKV / FP8 output transport on 4× / 8×. Timing includes text encoding, DiT denoising, and video/audio VAE decoding; model loading, compilation warmup, and final MP4 encoding are excluded. Sol-H3 brings Sol-Engine × Sol-Attn into one full-stack runtime: • dynamic sparse attention with no retraining • fused norm, RoPE, MLP, and sparse-attention setup • fused INT8 QKV / FP8 output communication across 8 GPUs • parallel, batched VAE decoding • precomputed AdaLN caching Inside the stack: • sparse-attention setup: 1.206 → 0.285 ms (−76.4%) • VAE decode: 7.55 → 0.602 s • ~24 GB memory freed per GPU Any MiniMax-H3 few-step LoRA can plug into the same engine, and the code is deployment-friendly under Apache 2.0. For us, the bigger milestone is crossing from “fast generation” into “faster than playback.” That opens the path toward continuous 24 FPS generation and truly interactive video systems. We’re excited to partner with reactor to release Sol-H3 and make it available as an API day-0. Try it now on Reactor: 🔗 Amazing team effort—full credits in the blog. LoveSy Junsong_Chen yitong li Haopeng Li Haocheng Xi Song Han

Enze Xie

202,121 Aufrufe • vor 8 Tagen

Working on Sparse Volumetric Light-maps. Thanks to CynicatPro🎃 for pointing me at Unreal's version. In a nutshell it's just another sparse voxel data structure. My implementation is, no doubt, different from Epic Games Store's own. I'm using 4x4x4 probe grid with intermediate nodes having very wide branching factor of 64 as well (4x4x4). I liked the parameters that Unreal is using, of limiting both total memory as well as the lowest level of detail, which is common in sparse grid implementations. Here's Bistro scene with just 1Mb limit. This is roughly equivalent to a 512x512 lightmap texture in 2d, except surface light maps require unique UVs and you typically get very little detail out of 512 resolution texture with a lot of light leaking. There is also no directional response. My implementation encodes second-order spherical harmonics for each probe (9 coefficients), encoding RGB channels as RGBE9995 (4 bytes). So far only worked on the structure, actual bake is yet to come. I've been eyeing sparse voxel structures for a while now, and have been studying them roughly since the GigaVoxel paper by Cyril Crassin but never really implemented anything for the GPU before. I was always the BVH-kind of guy. It's a fascinating topic. --- Stats for the scene: --- Total memory usage: 1.000 MB Node count: 609 Unique probe count: 24,025 Probe reuse: 38.36 % Unexpanded nodes: 15,714 --- Again, note that there is no GI going on here, only the structure of the probe tree and the algorithm for building it from a given scene.

Alex Goldring

11,519 Aufrufe • vor 7 Monaten

#GoProMAX2 is here 🚨 Industry-leading, true 8K 360 video delivers 21% more resolution than the competition, resulting in unmatched image quality that you only get from #GoPro. ✔️ Emmy® Award-Winning 360 technology ✔️ The only true 8K 360 camera. No misleading upscaling, no unusable black pixels, no AI-generated content ✔️ Twist + go replaceable lenses made from water-repelling optical glass. No tools or calibration required to swap ✔️ 5.6K60 + 4K100 for up to 4x slo-mo in 360 ✔️ 10-Bit color, GP-Log encoding, + with GoPro Labs, category-leading 300mbps bitrate ✔️ Seamless, invisible pole shots thanks to a new, back-to-back lens design + built-in 1/4-20 mount ✔️ Easy AI-powered editing tools with intuitive Reframe modes that anyone can use ✔️ Automatic POV + Selfie modes for minimal editing, while retaining full 360 flexibility ✔️ Most-in-class 6 microphones that unlock true-to-life spatial audio with innovative Audio Field of View ✔️ 23% larger, 1960mAh cold-weather Enduro battery ✔️ 360 Night Effects for creative capture in the dark ✔️ Sleek form factor for noninvasive mounting in action sports ✔️ Unbreakable Max #HyperSmooth with 360° Horizon Lock ✔️ Rugged + waterproof to 16ft (5m) ✔️ Quick-release magnetic mounting + compatibility with GoPro's entire accessory ecosystem ✔️ Single lens 4K60 video in Max HyperView at 180° FOV ✔️ 29MP 360 photos for cropping + zooming without quality loss ✔️ Bluetooth® audio connectivity for wireless microphones ✔️ Lightning-fast transfer speeds to the GoPro Quik App with Wifi 6 + BLE 5.3 🇺🇸 Designed in the USA Enhanced by a GoPro Subscription: ✔️ AI-edited highlight videos automatically sent to your phone ✔️ Unlimited cloud storage at 100% quality ✔️ No-questions-asked camera replacement 📦 Pre-order today, with free shipping and a free 1-year GoPro Subscription at Orders will ship on or before September 30th. *MAX2 delivers up to 21% more video resolution compared to competitive 360 cameras’ native maximum video resolution before they up-scale.

GoPro

139,416 Aufrufe • vor 11 Monaten

📺 IS $META STARTING A NEW NARRATIVE? + SEMIS WEAKNESS IS REAL + $AMZN THE NEXT ROTATION TRADE? One of the biggest questions right now: is $META latest AI announcement simply creating a short-term trading opportunity, or is it the beginning of an entirely new market narrative? Rather than focusing on headlines, watch price action. The initial reaction to the news is only the first step. What matters now is whether #Meta can hold key support, consolidate its gains, and begin a sustained bullish sequence. For swing traders, the $595 area is the key level that needs to hold. If #META continues to digest the move constructively, a breakout above $628 could signal the next leg higher. * While Meta has grabbed the spotlight, the bigger story may be the growing weakness across semis $SMH $SOXX. After leading the market for months, the leadership names are starting to show signs of fatigue. $MU lost momentum, broke important support, and closed near its lows. $SNDK appeared ready to break out before reversing sharply lower and triggering an active exit. These types of failed breakouts are often early signals that leadership within a sector may be changing. The 21-day moving average remains one of the most important technical levels to monitor. Throughout the rally since March, this moving average has consistently acted as support. Until leading stocks begin breaking and closing below it, the broader uptrend remains intact. A pullback toward the 50-day moving average can still be considered a normal correction rather than the start of a bear market. * Another important takeaway is the potential for sector rotation. If institutional money continues to leave semiconductors and AI infrastructure, it will likely seek opportunities elsewhere within large-cap technology. For example, $AMZN. Amazon is a value-oriented mega-cap technology stock that could benefit if investors rotate away from expensive AI winners. Despite a compelling long-term story and an attractive valuation relative to its history, Amazon has been frustrating to trade. Still, improving relative strength could attract fresh buying if this rotation continues. I purchased July 10 $250 call options, giving myself limited risk while allowing time for the rotation thesis to develop. * So, you should avoid making bold predictions during periods of uncertainty. Instead, define your risk, stay flexible, follow your trading rules, and let price action determine whether $META is launching a new AI narrative, whether semiconductor weakness deepens, and whether $AMZN becomes one of the next beneficiaries of market rotation. * If you found this helpful, please ❤️like and 🔁retweet

Scott Redler

13,656 Aufrufe • vor 2 Monaten

After 8+ years on the Tesla Autopilot team and 3 years at Intel, I started Apex Compute to design a new architecture for efficient AI inference. For the past 9 months, we’ve been building our custom inference accelerator. Today we’re releasing Unified Engine v1. Last June we raised our seed round with Maxitech , DeepFin Research, Soma Capital and an incredible group of angel investors. In less than 9 months, we completed our RTL architecture and brought our first pre-silicon prototype to life on FPGA. Our architecture combines systolic array and vector processing in a single compute engine with multiple architectural optimizations, achieving very high FLOPs utilization. A single engine is super lean and it uses less than 90K LUTs and 1 MB Block RAM. It may also be one of the smallest logic-footprint compute engines developed so far. Our Unified Engine v1 supports: -matrix-matrix multiplication (~95% FLOPs utilization) -softmax (~90% FLOPs utilization) -broadcast and element-wise operations -RMSNorm / LayerNorm -block quantization/dequantization (fp4, int4) -multi-engine synchronization and many other operations. We even implemented memory-efficient attention similar to FlashAttention, reaching ~90% FLOP utilization. Full benchmarks and the software stack are available on our GitHub: We have basic compiler written in Python and it supports PyTorch tensors directly to easily test and transfer tensors between the accelerator and host using bf16, fp4 and int4 formats. Our FPGA prototype can already run LLM inference and outperform NVIDIA Jetson Orin Nano, even on a mid-tier FPGA setup (6.4x lower memory bandwidth, 18% slower clock speed at 4.5 Watts). Check the side-by-side comparison video below. Our GitHub includes low-level operator implementations, examples for tiled matrix multiplication, operation chaining, tensor parallelism, attention kernel and a full Gemma 3 1B model implementation. Many more models(Vision Transformers and VLA) are coming soon. Our accelerator IP is AXI-ready for deployment on any AMD(Xilinx) FPGA platform today. Even better, our two-engine prototype runs on an entry-level AMD(Xilinx) FPGA as a PCIe accelerator card. You can purchase it here for $50 to experiment our pre-silicon prototype on your desktop PC or Raspberry Pi 5. We will be releasing hardware bitstream updates as the architecture gets new features. More to come soon! We are expanding our team and looking for compiler engineers and floating-point hardware design engineers. If you're interested, please send me a DM.

Hasan

37,748 Aufrufe • vor 6 Monaten

Bob McGwier (Science Bob McGwier) worked as a contractor in mathematics and signals for three letter agencies (including the CIA) for decades. Now he has some dumbfounding things to say about UFOs. They involve alien prophecies given to Presidents like Obama and human implants used as tracking devices. Episode full link below 👇🏻: 1. Human Implants as Tracking Devices: Bob thinks implants found in experiencers emit RF signals, which can be geolocated using satellite technology like his own company, HawkEye 360. The potential for advanced tracking was demonstrated with measurable evidence of RF bursts. Bob’s girlfriend has an implant and he tested this LIVE with her. Video in the episode. 2. CRISPR and DNA as Data Storage: The use of CRISPR-Cas9 for encoding data into DNA raises the possibility that humans or other organisms could serve as biological data carriers, potentially undetectable by conventional methods. We could be walking hard drives! 3. Advanced Physics and Exotic Propulsion: We discuss the manipulation of the stress-energy tensor (through various methods), polarizable vacuum fields, and negative energy to engineer warp drives with cutting-edge theoretical and experimental physics, sometimes inspired by reverse-engineered technology. 4. Carl Sagan’s Alleged Involvement: Sagan’s purported deep ties to intelligence and his possession of high-security clearances contradict his public anti-UFO stance, implying he may have been privy to classified knowledge about extraterrestrial phenomena. 5. Secrecy in Government Science: The classification of new physics under the Atomic Energy Act, which includes discoveries potentially related to UAPs, underscores how deeply science tied to national security can remain hidden. 6. Consciousness and Physics Intersection: Theories linking consciousness to quantum mechanics and exotic propulsion hint at an overarching framework where the mind could influence physical reality, potentially tied to phenomena like UAP navigation. 7. Chris Bledsoe’s Prophecies and Government Interest: The involvement of high-level government figures like Jim Semivan with experiencers like Bledsoe suggests a serious interest in understanding the intersection of metaphysical experiences and their implications for national security. Bob claims Bledsoe’s prophecy made it to the White House itself!

Jesse Michels

117,601 Aufrufe • vor 1 Jahr

DOUG CASEY'S NEXT BIG WIN: OIL STOCKS POISED FOR A RUNAWAY BULL MARKET Legendary investor Doug Casey has identified a sector that the market has almost completely forgotten. Oil stocks trade at valuations that would have seemed impossible just a few years ago, yet they come with solid cash flows and attractive yields. With political tensions in the Middle East refusing to cool, this forgotten corner of the market may be about to wake up in dramatic fashion. THE HISTORICAL COLLAPSE IN ATTENTION ➡️ In 1980, during the last major oil market peak, oil and natural gas stocks made up 30 percent of the S&P 500. ➡️ Today that weighting has fallen all the way to just 4 percent. ➡️ Investors have turned their backs on the entire sector. THE ATTRACTIVE FUNDAMENTALS TODAY ➡️ Oil has reached what Casey describes as a new equilibrium level around 95 dollars per barrel. ➡️ Production costs for the industry sit near 60 dollars. ➡️ This spread allows producers to generate strong returns at current prices. THE GEOPOLITICAL TAILWIND ➡️ The conflict between Iran and Israel shows no signs of ending. ➡️ Casey puts it bluntly: "This thing with Iran and Israel ain't going to go away." ➡️ He believes oil prices are going to go higher for political reasons. THE DIVIDEND AND VALUATION EDGE ➡️ Most oil stocks offer fat dividend yields that the market is completely ignoring. ➡️ The sector trades at deeply depressed valuations relative to almost everything else. ➡️ Doug Casey sees this as the setup for a runaway bull market in oil stocks. THE BOTTOM LINE Doug Casey sees the oil stocks sector as one of the most compelling opportunities available to investors right now. The extreme underrepresentation in major indices, reliable profitability at current prices, and building pressure from geopolitics create conditions for significant appreciation that the broader market has yet to price in. Smart money positions itself before the rest of the world wakes up to what is hiding in plain sight. #OilStocks #DougCasey #EnergyInvesting #OilPrices #Geopolitics #DividendStocks #ContrarianInvesting

Mark

44,770 Aufrufe • vor 3 Monaten

$NVDA $MU $SNDK $LITE PAPER OVERVIEW AND CORE CLAIMS The paper “KV Cache Transform Coding for Compact Storage in LLM Inference” introduces kvtc, a transform-coding pipeline that compresses transformer key-value (KV) caches primarily for storage and transfer in LLM serving, rather than for accelerating the per-token attention kernel during active decoding. The method combines 3 stages: (1) feature decorrelation via a PCA basis computed from a calibration dataset and reused across requests; (2) adaptive, variable-precision quantization with bit allocation solved via dynamic programming (DP), including groupwise scaling/shift overhead; and (3) lossless entropy coding (DEFLATE via nvCOMP in the reference implementation) to exploit residual redundancy after quantization. The central empirical claim is that KV tensors contain large, exploitable redundancy across heads and layers, enabling approximately 20× compression versus a 16-bit baseline with negligible degradation across a broad set of accuracy and long-context benchmarks, with materially higher compression (≥40×) available at modest quality cost in some regimes. The system claim is that such compression materially improves the economics of multi-turn, prefix-reuse serving by extending effective KV cache capacity in GPU HBM and host tiers (DRAM/NVMe) and by reducing inter-node and GPU↔host bandwidth demands, thereby improving cache hit rates and reducing time-to-first-token (TTFT) relative to recomputation when caches would otherwise be evicted. KV CACHE AS THE DOMINANT STATE VARIABLE IN INFERENCE ECONOMICS KV cache growth is linear in context length and is multiplicative in layers and attention heads, making it an increasingly dominant constraint as (a) context lengths expand, (b) models add layers and maintain large hidden dimensions, and (c) production workloads shift toward iterative and tool-augmented interactions that repeatedly reuse long prefixes. The paper uses the canonical 16-bit KV cache size formula (4·l·h·d_head·t) bytes and reports 16-bit KV cache sizes per 1K tokens of context that are already operationally large: 128MiB for Llama 3.1 8B, 160MiB for Mistral NeMo 12B, and 320MiB for Llama 3.3 70B Instruct. In binary units, these figures imply per-token KV footprints of 128KiB/token (Llama 3.1 8B), 160KiB/token (Mistral NeMo 12B), and 320KiB/token (Llama 3.3 70B Instruct) at 16-bit. For a 10K-token prompt (10×1K in the paper’s binary convention), the 16-bit KV cache sizes scale to approximately 1.25GiB (Llama 3.1 8B), 1.56GiB (Mistral NeMo 12B), and 3.13GiB (Llama 3.3 70B Instruct). These magnitudes explain why stale caches create a throughput–latency dilemma: retaining them in HBM maximizes responsiveness on future turns but crowds out concurrent sessions; evicting them forces quadratic-cost prefill recomputation and increases TTFT; offloading them to host or storage introduces large transfer overhead and consumes DRAM/NVMe capacity. A key operational nuance emphasized is that modern serving stacks increasingly treat KV caches as a database, leveraging block paging and shared-prefix reuse. In the common disaggregated serving design (separate prefill and decode nodes), KV cache transfer becomes a dominant category of cross-node traffic. Under that design, any reduction in KV cache size directly increases effective fabric capacity and reduces tail latency attributable to congestion, while also enabling longer cache lifetimes in “hot” (HBM) and “warm” (CPU DRAM) tiers that raise cache hit rates and reduce recomputation frequency. The paper’s quantitative example illustrates the economic stakes: a 1,000-line code file tokenized at ~10 tokens/line yields ~10K tokens; for Llama 3.3 70B, an 8-bit KV cache for that context is ~1.6GiB. Reuse across subsequent turns or parallel chats around the same file is valuable, but HBM scarcity makes retaining many such caches infeasible without compression. TECHNICAL MECHANISM: WHY KV CACHES ARE COMPRESSIBLE AND HOW KVTC EXPLOITS IT The technical rationale begins with an empirical observation: keys (and, to a lesser extent, values) across different attention heads can be aligned into a shared latent space using orthogonal transformations (Procrustes alignment). This supports the hypothesis that head-specific projections introduce rotations of a common subspace rather than completely distinct information, implying that concatenating across heads and layers should reveal low-rank structure suitable for linear decorrelation and dimensionality reduction. The method operationalizes this using a PCA/SVD basis learned from calibration data rather than recomputing a decomposition per prompt. This design choice targets production viability: per-prompt SVD is computationally expensive and scales poorly with long prompts and frequent cache updates. kvtc is explicitly structured as an offline-calibrated, online-applied codec: Calibration (performed 1 time per model and compression setting for DP allocation) A calibration dataset is forwarded through the model to collect KV caches. Token positions are pooled, and a subset of positions is sampled. Keys and values are processed separately. Several implementation choices are highlighted as decisive for stability: Rotary positional embeddings are effectively removed prior to compression (“undo positional rotations”), because positional rotations degrade the apparent low-rank structure of keys. “Attention sink” tokens (the earliest tokens in the sequence) and a sliding window of most recent tokens are excluded from compression because they disproportionately affect attention patterns and are empirically more sensitive to reconstruction error. Cross-layer concatenation is used: keys (or values) from multiple layers and heads at the same token position are concatenated along the feature axis to form a higher-dimensional feature vector. PCA is computed over these concatenated vectors, improving robustness relative to per-layer or per-head PCA. The PCA basis is computed via SVD of centered calibration data, using randomized SVD for scalability with a target rank cutoff. The paper reports calibration regimes of 160K tokens for several models with a 10K PCA dimension cutoff (8K for Qwen variants with fewer KV heads), selected to fit within a single 80GB H100 memory envelope and complete within minutes. A critical economic detail is that the same PCA basis can be reused across multiple compression ratios; only the DP-derived precision assignment changes per compression target. Compression (applied between inference phases) Compression operates on stored KV cache tensors, not on weights, and does not modify attention computation. The KV cache is projected into the PCA basis, quantized, packed, and then entropy-coded. Compression is positioned as a background or between-phase operation (after decoding, or between prefill and decode), executed on GPU or CPU depending on where the cache currently resides. The design intent is that compression should not sit on the critical per-token decoding path; it is a storage and transport optimization. Decompression (performed prior to reuse) Decompression reverses the entropy coding and quantization and applies the inverse PCA projection. A practical latency optimization is proposed: inverse projection can be performed layer-by-layer using submatrices of the PCA basis, allowing generation to begin before the full cache is reconstructed, reducing TTFT. Quantization and bit allocation are the core differentiators versus simpler PCA truncation. PCA provides ordered components by variance; kvtc uses DP to allocate a global bit budget across PCA coordinates (and across groups of coordinates) to minimize reconstruction error in the decorrelated domain. Groups of subsequent PCA coordinates share 16-bit shift and scale factors (a microscaling-inspired design), and the DP algorithm jointly selects group size and precision type under a bit budget, including the overhead of per-group metadata. DP commonly assigns 0 bits to many trailing PCA components, which both increases compression and provides a mechanism to trim the PCA basis to the subset of components that actually carry payload, reducing compute and storage overhead of the projection matrices in deployment. Lossless entropy coding then exploits the structure induced by quantization. DEFLATE is used in the reference implementation, and the paper emphasizes that the incremental gain from the lossless stage is content-dependent but meaningful, with an average uplift of ~1.23× on top of quantization in the reported regime. An ablation in the appendices indicates that GPU-friendly variants (GDeflate) can achieve nearly identical compression ratios (≤0.1 difference in measured cases), implying that throughput-optimized lossless codecs can likely be substituted without sacrificing meaningful compression. EMPIRICAL RESULTS: ACCURACY, COMPRESSION, AND LATENCY General-purpose 8B–12B dense models The paper evaluates Llama 3.1 8B, MN-Minitron 8B, and Mistral NeMo 12B across math/knowledge (GSM8K, MMLU) and long-context tasks (Qasper, Lost in the Middle, RULER Variable Tracking) under a simulated multi-turn regime where compression/decompression is applied periodically, with a sliding window of recent tokens excluded. A consistent pattern appears: kvtc maintains near-vanilla performance through 16× compression settings, and remains competitive at 32×, with degradation becoming task- and model-dependent at 64×, particularly on long-context retrieval metrics when compression is pushed aggressively. Selected quantitative anchor points from the paper’s standard-error table (all values are reported with the paper’s evaluation setup and token-window exclusions): Llama 3.1 8B Vanilla: GSM8K 56.8, MMLU 60.5, Qasper 40.4, LITM 99.4, RULER-VT 99.8 kvtc16×: GSM8K 56.9, MMLU 60.1, Qasper 40.7, LITM 99.3, RULER-VT 99.1 kvtc32×: GSM8K 57.8, MMLU 60.6, Qasper 39.4, LITM 99.1, RULER-VT 98.9 kvtc64×: GSM8K 57.2, MMLU 60.7, Qasper 37.8, LITM 90.2, RULER-VT 95.9 These results indicate that, for this model, long-context sensitivity emerges at 64× with meaningful drops in LITM and RULER-VT, while math/knowledge scores remain stable, implying a differential sensitivity consistent with key-vector precision being more critical for retrieval-style behavior. Mistral NeMo 12B Vanilla: GSM8K 61.9, MMLU 64.5, Qasper 38.4, LITM 99.5, RULER-VT 99.8 kvtc16×: GSM8K 62.0, MMLU 64.4, Qasper 37.6, LITM 99.8, RULER-VT 99.5 kvtc32×: GSM8K 62.2, MMLU 63.8, Qasper 37.5, LITM 99.6, RULER-VT 98.7 kvtc64×: GSM8K 61.9, MMLU 61.4, Qasper 38.0, LITM 95.3, RULER-VT 98.0 Here, degradation at 64× is visible but materially smaller than the Llama 3.1 8B LITM drop, suggesting model-architecture or training-data differences can change the tolerance envelope for aggressive KV cache distortion. MN-Minitron 8B Vanilla: GSM8K 59.1, MMLU 64.3, Qasper 38.2, LITM 99.8, RULER-VT 99.4 kvtc16×: GSM8K 60.3, MMLU 64.1, Qasper 38.6, LITM 99.3, RULER-VT 98.8 kvtc32×: GSM8K 59.1, MMLU 63.7, Qasper 37.7, LITM 86.9, RULER-VT 96.0 kvtc64×: GSM8K 57.8, MMLU 62.1, Qasper 38.1, LITM 59.5, RULER-VT 93.4 This model shows markedly higher sensitivity on LITM at 32× and 64×, despite stable short-context metrics, reinforcing that “compression safety” is not monotonic in parameter count and that pruning/distillation choices can alter KV cache redundancy or robustness. Comparisons to baselines The paper compares kvtc to quantization baselines (KIVI, GEAR, FP8) and eviction baselines (H2O, TOVA), plus an SVD-based prefill-optimization method (xKV). Across the reported tasks: Low-bit quantization methods at modest compression (2-bit KV schemes) show earlier degradation in long-context behavior than kvtc at substantially higher compression settings. Eviction methods perform poorly as generic compressors for long-context tasks, consistent with their objective function (selective pruning) being misaligned with “lossless-ish storage for reuse.” xKV shows competitive results on some tasks but a consistent underperformance on Qasper relative to kvtc and vanilla in the provided tables, consistent with method-specific distortions introduced by its decomposition regime. Reasoning models and high-variance tasks For DeepSeek-R1-distilled Qwen 2.5 reasoning models, the paper evaluates AIME 2024/2025 and LiveCodeBench coding. Results are averaged over 8 runs with large variance, but a key inference is that kvtc at ~9×–21× compression achieves broadly similar AIME scores within variance bands, while coding performance remains stable at ~9× and degrades more visibly at ~18×–21× on the 7B model. An important nuance is that smaller reasoning models already have smaller KV footprints (reported ~29KiB/token for Qwen R1 1.5B versus 131KiB/token for Llama 3.1 8B), so the economic value of aggressive KV cache compression is proportionally higher for large models and long contexts than for small models with short contexts, unless the serving system’s bottleneck is dominated by cache transfer rather than HBM capacity. Multi-GPU inference and pipeline parallel For Llama 3.3 70B Instruct run pipeline-parallel across 4 GPUs (20 layers per GPU), the paper compresses KV cache chunks independently per GPU. On MATH-500, the reported accuracy declines from 75.6 (vanilla) to 74.4 at 10× and 72.6 at 20×, with standard errors near ~1.9. NIAH and LITM remain at 100.0 for all tested ratios in that table. The paper notes that joint compression across chunks could improve accuracy for some offload scenarios but is not required for feasibility, highlighting an engineering trade-off between deployment simplicity in distributed settings and optimal global compression. Latency and TTFT economics A critical system result is the measured compression/decompression latency on an H100 for a non-fused implementation. For Mistral NeMo 12B in bfloat16: BS=8, CTX=8K: compression 379ms, decompression 267ms; vanilla recompute TTFT 3098ms; kvtc decompression TTFT 380ms BS=2, CTX=16K: compression 194ms, decompression 143ms; vanilla recompute TTFT 1780ms; kvtc decompression TTFT 208ms These measurements imply that, when a cache would otherwise be recomputed, decompressing a stored compressed cache can reduce TTFT by ~8×–9× in these scenarios, even without kernel fusion. The decomposition of runtime shows PCA projection and entropy coding as the largest contributors, implying that GPU-optimized kernels and faster GPU-native lossless codecs could reduce overhead further. The fundamental economic conclusion is that, in multi-turn settings with long prefixes, compression-induced overhead is likely dominated by the avoided prefill compute and avoided transfer overhead for uncompressed caches. KEY DEPLOYMENT-SENSITIVE DESIGN CHOICES AND FAILURE MODES Several design choices appear to be “hard requirements” rather than optional optimizations: Sink tokens and sliding window exclusions The paper’s ablations show that compressing early “sink” tokens can catastrophically degrade accuracy at high compression ratios (example: Llama 3.1 8B at 64× collapses on multiple tasks when sink tokens are compressed). Similarly, compressing the most recent tokens hurts performance, motivating a sliding window (default 128 tokens) that remains uncompressed. This introduces a predictable engineering constraint: kvtc is not a uniform compression of the full cache; it is a policy-driven, token-position-dependent codec. Production integration therefore requires correct handling of token positions, attention sinks, and window management, and these policies must be aligned with attention-kernel behavior and model-specific sink dynamics. RoPE handling Removing positional rotations prior to compression is described as important for preserving low-rank structure. In deployment, this implies that the codec must be position-aware and must invert and reapply RoPE correctly. This is an additional source of complexity relative to pure per-token quantization and is sensitive to model variants and RoPE parameterizations. Calibration set representativeness The method’s quality hinges on the PCA basis generalizing from calibration data to production data. The paper demonstrates relative stability with 160K–200K calibration tokens and explores domain shifts (general web text vs math traces vs code). Results suggest that moderate domain mismatch is tolerated at 16×–64×, while extreme compression (e.g., 256× in ablations) becomes materially more sensitive to calibration choice. In production, this implies that operators targeting the “negligible degradation” regime should be able to calibrate with broadly representative corpora, while operators targeting ultra-high compression for specialized workloads should expect tighter coupling between calibration domain and achieved quality. PCA matrix storage overhead and operational footprint A non-trivial hidden cost is the need to store PCA projection matrices per model. The paper reports that, prior to DP trimming, PCA matrices stored at 16-bit can amount to a meaningful fraction of model parameter count (examples reported: ~2.4% for Llama 3.3 70B, ~8.7% for Llama 3.1 8B). This overhead is amortized across all cached sessions for a model but competes with HBM/DRAM budgets in multi-model serving. DP-driven trimming can reduce this overhead at higher compression ratios by removing zero-bit components, but the directionality is not guaranteed at low compression ratios if many components remain active. In distributed inference (pipeline parallel), per-chunk PCA can reduce matrix sizes, but may reduce cross-layer decorrelation benefits if fewer layers are concatenated. SYSTEM-LEVEL IMPLICATIONS FOR GENERATIVE AI INFRASTRUCTURE GPU AND HBM The principal infrastructure implication is that KV cache compression at storage time targets the dominant memory allocator stressor in stateful serving: the accumulation of idle or warm conversation state. For workloads with long reusable prefixes (code assistants, enterprise agents with large system prompts, repeated RAG scaffolds, document chat), the limiting resource frequently becomes HBM reserved for KV caches rather than compute. By compressing stale caches by ~20× (or more), the same HBM budget can retain a materially larger working set of cached prefixes, increasing cache hit rates and reducing recomputation. This effect is multiplicative with cache-aware routing and prefix sharing: more prefixes can remain resident (hot or warm) and can be routed to nodes that already hold them, improving both throughput and tail latency. However, kvtc as described does not reduce the active KV cache footprint during the actual attention computation for a currently decoding sequence, because the model operates on decompressed KV caches during decoding. Therefore, the method does not directly reduce HBM bandwidth consumed by attention kernels during steady-state decode, and does not directly address the “memory traffic per generated token” bottleneck that motivates online KV quantization and eviction strategies. The primary HBM benefit is increased effective capacity for caches between turns and reduced HBM pressure from storing many idle sessions, not reduced per-token decode bandwidth. Compression and decompression themselves consume GPU compute and memory bandwidth. The measured decompression TTFT of ~208ms–380ms in the provided benchmarks indicates that the overhead is real but can be materially smaller than recomputation of long prefixes. In an HBM-constrained serving environment, this overhead can be interpreted as a trade between (a) maintaining more caches warm and paying decompression on reuse versus (b) evicting caches and paying full prefill recomputation. The decision boundary will depend on distribution of inter-turn idle times, probability of reuse, and SLA sensitivity to TTFT. kvtc expands the feasible region where keeping caches is economically rational, especially for long prompts. CPU AND DRAM The method implies a stronger role for CPU DRAM as a warm KV cache tier. A ~20× compression ratio changes the practical scale of “warm state” that can be stored per server. Using the paper’s reported KV cache sizes, a 10K-token 16-bit KV cache for Llama 3.3 70B is ~3.13GiB; compressing by ~20× would reduce this to ~160MiB. At that size, storing hundreds to thousands of warm conversation states in DRAM becomes materially more feasible, increasing cache hit rates and reducing NVMe dependence. This can shift system design from “HBM-only hot caches with aggressive eviction” toward “HBM hot + DRAM warm with long retention,” which is structurally analogous to CPU page cache hierarchies in classical systems design. CPU compute implications depend on where compression is executed. The paper explicitly allows compression on CPU if the cache is already in storage, but the strongest bandwidth savings are achieved when compression happens before moving KV caches off the GPU. If an operator chooses GPU-side compression prior to PCIe/NVLink transfer, CPU compute overhead is modest (orchestrating and DP calibration offline). If an operator instead transfers uncompressed caches to CPU for compression, bandwidth savings are forfeited and CPU memory bandwidth becomes a bottleneck. Therefore, the most economically coherent deployment path is GPU-native compression/decompression with CPU DRAM used as the warm storage reservoir.

TheValueist

16,549 Aufrufe • vor 7 Monaten