Загрузка видео...

Не удалось загрузить видео

На главную

remind-reid-tracker REMIND — RE-Identification with Memory for INDoor Navigation REMIND addresses a core challenge in visual tracking: re-identifying objects that disappear and reappear, look similar to one another, or are observed from changing viewpoints. Rather than relying on position or motion cues, REMIND builds appearance-based identity models per object...

19,013 просмотров • 1 месяц назад •via X (Twitter)

Комментарии: 0

Нет доступных комментариев

Здесь появятся комментарии из оригинального поста

Похожие видео

Everyone is sleeping on Meta's SAM 3 release. But it's actually a big deal. Here's why: Companies spend millions paying humans to label images and videos frame by frame. A single autonomous driving dataset? Months of work, hundreds of annotators, millions in cost. Without labeled data, you can't train custom models. Without custom models, you're stuck with generic solutions. This is why most companies never move past pilots. SAM 3 breaks this cycle. First let's look at the evolution: SAM 1 segmented objects when you clicked on them. Revolutionary, but one object at a time. SAM 2 added video tracking with memory. Game-changing, but you still manually prompted every object. SAM 3 changes everything with text prompts. Type "yellow school bus" and it finds ALL of them in your image or video. Not just one. Every instance across thousands of frames. Now here's where people get confused: "Can't I just use GPT-5 or Gemini for this?" No, and here's why that's a terrible approach. Large multimodal LLMs are great for reasoning, but they're slow and expensive for production visual tasks. You're paying API costs per image, waiting seconds for responses, getting inconsistent results. SAM 3 runs in 30 milliseconds on a single GPU for 100+ objects. That's 100x faster, and you own the infrastructure. More importantly, SAM 3 gives you precise pixel-level masks, not descriptions. Try asking an LLM to segment every defective part on a manufacturing line in real-time. It won't work. SAM 3 does this effortlessly. The real breakthrough is their data engine. Meta built an AI-human hybrid system that's 5x faster for complex annotations. They trained SAM 3 on 4 million unique visual concepts - 50x more than existing benchmarks like LVIS. SAM 3 is trained on 4 million unique visual concepts, it handles everything: - Text-based concept search - Interactive refinement with clicks - Video tracking across frames - Zero-shot detection of new concepts The model is open source. Weights, code, and benchmarks are on GitHub. If you're building computer vision applications, this is the foundation model to evaluate. The annotation time savings alone will pay for integration costs within weeks. Find the relevant links in the next tweet!

Akshay 🚀

46,438 просмотров • 10 месяцев назад

I’ve been researching depth-residual routing & delta memory- Kimi K3 proved this holds at 2.8T params. But only for models trained this way I then asked: can you retrofit this onto existing LLM/world models? Introducing Retro-DARC: a Delta Attention Residual Compute adapter that retrofits depth-selective memory onto existing models without retraining: EA exact no-op at insertion (bitwise 0.0 deviation on 3 public checkpoints - proven exact even through MoE routing and early exit), adapter-only training, exact rollback, and content-level memory audits that weight-touching adapters can't express by construction. The early results are promising. A ~262K-param adapter beat LoRA on a frozen public checkpoint, inserted at zero behavioral risk. And it's the first memory you can audit: zero the bank and the frozen loss returns to machine precision; shuffle it and 70–107% of the gain vanishes. That's a model measurably using its own computation history. Retro-DARC is useful for improving K3 too. K3 serves in MXFP4/MXFP8, and layer outputs - exactly what its AttnRes retrieves - collapse under that corruption in my tests (0.04–0.75 top-1) while normalized delta keys hold 0.86–0.99. A Softmax1 null route with proven mass bounds lets the router decline to retrieve instead of forcing the sinks and outliers OASIS documented; typed memory turns "what is depth retrieval doing?" into a runnable audit. And one experiment is free: K3 already computes a per-step KDA update magnitude and throws it away - that's a depth-saliency prior waiting to be read. For world models, the same contract becomes one typed memory interface: Retro-DARC-X gives world/action planners a single bank - observations, actions, tools, latent causes, traces, physics residuals, failures - the substrates a planner needs to remember in order to act over long horizons in the physical world. You can experience the thesis yourself. I created which runs the paper's memory contract inside a walkable AI world. Innovation Memory makes your interactions survive looking away or reloading, and Dream Residue turns your movement into an inspectable glowing trace. Join the main co-creation exhibition - the contest opens Wednesday. One of the six memory applications of the paper - saliency-weighted curation, will rank worlds by genuine novelty rather than recency and pick each world's most surprising viewpoint. And we're bringing live shows to SF + LA next month: screens showcasing worlds that remember, so you can literally play hide-and-seek with a dream that knows where you hide. Hide-and-seek is the exact test every 2026 world-model benchmark says the field fails, something leaves the frame, keeps changing, and has to come back right. This has been a fun little one-person research project so far! I’d love feedback and keep working on it. I’ve had over 100 references- thanks for all the prior work from teams that inspired me! I'd also love to intern at a GPU-rich lab (lol) and keep making beautiful worlds - if you let me keep running my remaining tests :) 2MB adapter kit, demo, prior iterations: Paper:

ada cyborg (🤖, 🔮)

11,344 просмотров • 1 месяц назад

$MU $SNDK $LITE $VRT NVIDIA and Groq: 2nd and 3rd Order Strategic Infrastructure Effects and Market Implications Public reporting indicates NVIDIA has agreed to acquire Groq for approximately $20,000,000,000 in cash, while excluding Groq’s nascent cloud business from the transaction perimeter. The reported carve-out materially constrains the immediate, direct linkage from the acquisition to incremental, NVIDIA-controlled data center capacity build-out because GroqCloud appears to be the principal channel through which Groq hardware is currently monetized at scale as a service. The infrastructure-market implications therefore depend primarily on post-close product strategy: whether NVIDIA (1) commercializes Groq silicon as a distinct inference product line and drives broad deployment through OEM/ODM channels and partners, (2) uses the acquisition mainly to absorb IP and talent while de-emphasizing standalone Groq hardware volumes, or (3) uses Groq technology to reshape NVIDIA’s own inference systems and networking roadmaps. The dominant transmission mechanism into memory, networking, and facility infrastructure markets is the degree to which NVIDIA shifts incremental inference deployments away from GPU architectures that are tightly coupled to external high-bandwidth memory (HBM) and toward Groq’s current architecture, which emphasizes large on-chip SRAM, deterministic compiler-scheduled execution, and direct chip-to-chip connectivity. Independent and company-published materials describe Groq’s current-generation approach as having no external memory, keeping weights and KV cache on-chip during processing, and requiring model sharding across multiple chips due to limited on-chip SRAM per device. That architectural choice is directionally HBM-negative on a per-accelerator basis and ambiguous for DRAM, NAND, networking, power, and cooling on a per-token basis because the design can reduce memory wall losses and tail-latency overhead while potentially increasing the number of chips and interconnect endpoints required to serve large models and long-context workloads. HBM implications are the most mechanically straightforward but should be framed as second-derivative rather than absolute. If Groq-class inference silicon meaningfully displaces NVIDIA GPU-based inference deployments, incremental HBM bit demand tied to inference growth could be reduced relative to a GPU-only baseline because Groq’s current approach does not appear to attach HBM stacks to each accelerator. However, current market structure suggests HBM remains supply-constrained and is being pulled by multiple vectors including continued GPU training scale and high-capacity inference configurations, with leading suppliers signaling tight conditions extending beyond 2026. In that environment, reduced inference-driven HBM intensity could primarily reallocate scarce HBM supply toward higher-end training and premium inference GPUs rather than creating an outright volume collapse, preserving high utilization of HBM capacity while potentially affecting the slope of pricing power and capacity expansion urgency over a multi-year horizon. The key downside scenario for the HBM complex would be a durable architectural bifurcation where “good-enough” inference shifts disproportionately to HBM-less ASICs across a broad swath of deployments (latency-sensitive, batch-1, cost-per-token optimized), while training remains GPU-HBM dominated; such a split would reduce the portion of future inference compute that naturally monetizes through HBM content and could compress the incremental HBM-per-AI-dollar ratio. The key upside/neutral scenario for HBM is that the supply chain remains fully allocated regardless, with NVIDIA using any “freed” HBM to ship more high-end GPUs into training and long-context inference, especially as roadmaps increase HBM per GPU, sustaining robust aggregate bit demand even if inference becomes more heterogeneous. Conventional DRAM implications split into 2 channels: (1) DRAM wafer capacity diversion into HBM and (2) DDR content per server in AI clusters. Supplier commentary indicates that AI-driven memory demand is supporting elevated DRAM markets more broadly, and HBM production is resource-intensive versus conventional DRAM, tightening supply for DDR products in parallel. A meaningful NVIDIA pivot to an inference architecture that reduces HBM dependence could, at the margin, ease the most acute HBM-driven bottlenecks and allow memory manufacturers more flexibility in balancing DRAM mix, which could be modestly DDR-positive on the supply side (less crowding-out) even if it is DDR-neutral or slightly negative on the demand side (if per-node CPU/DDR requirements decline due to more efficient accelerator utilization). The dominant practical outcome is likely that DDR demand remains supported by broad AI server proliferation and increasing memory footprints at the system level (CPUs, networking stacks, caching layers, retrieval-augmented pipelines), while HBM remains the premium profit pool; therefore, any HBM displacement that increases total server volumes could indirectly keep DDR demand resilient even if DDR per accelerator is not rising materially. NAND flash implications are comparatively indirect and volume-driven rather than architecture-driven. Inference clusters require SSD capacity for model storage, container images, logging, and increasingly for fast local retrieval indices and embedding stores, but the storage footprint per unit of compute is typically smaller than in training pipelines that stage large datasets and checkpoints. If NVIDIA uses Groq to lower inference cost and latency enough to expand the total number of inference deployment locations (regional colocation, enterprise on-prem, sovereign footprints), aggregate SSD attach could rise through geographic fragmentation and replication of model artifacts across more sites, even if per-site storage is modest. The NAND effect is therefore likely to be demand-broadening and mix-positive (datacenter SSDs) but not a primary swing factor versus the macro AI capex cycle and consumer/device cycles. Hard disk drive (HDD) markets should see negligible direct sensitivity because nearline HDD demand is driven by bulk storage and cloud archiving economics, while inference acceleration choices primarily reshape compute and network layers; any HDD benefit would be a tertiary function of overall data center square footage expansion rather than a direct consequence of Groq silicon displacing GPUs. Optical networking implications require separating (1) intra-cluster back-end fabrics that connect accelerators and (2) front-end / data center interconnect (DCI) that connects sites and regions. Groq’s own positioning and third-party reporting suggest scaling beyond a single node or rack relies on high-bandwidth fabrics and, in some described configurations, optical interconnect scaling across hundreds of chips. If NVIDIA commercializes Groq at scale, 2 offsetting forces emerge: lower cost-per-token and improved latency could expand inference throughput and drive more east-west traffic, increasing demand for high-speed switching and optics; conversely, if Groq delivers materially higher utilization and tokens per unit of network bandwidth for certain workloads, the network required per served token could decline. Public NVIDIA materials already indicate an aggressive photonics roadmap aimed at scaling AI factories, including co-packaged optics (CPO) switches and explicit collaboration with Coherent and Lumentum in the silicon photonics supply chain. That linkage is important because it suggests that, independent of Groq, NVIDIA is already pushing optics integration deeper into the switch package to reduce power and increase resiliency; Groq increases the strategic incentive to reduce network power and latency if inference becomes even more distributed and latency-sensitive. For Lumentum and Coherent specifically, the net implication is less about “more optics versus fewer optics” and more about a shift in optics form factor and value capture. Co-packaged optics can reduce reliance on pluggable transceivers in some switch architectures while increasing demand for integrated photonic engines, lasers, fiber attach, packaging processes, and component-level supply. NVIDIA’s own announcements explicitly position Coherent and Lumentum as collaborators in creating the integrated silicon/optics process and supply chain for photonics switches. If Groq accelerates the transition to very large-scale fabrics (more endpoints, higher port speeds, tighter power envelopes), that tends to pull forward CPO adoption and amplifies demand for the underlying photonics components even if the conventional pluggable module TAM is structurally pressured over time. If Groq instead pushes inference toward smaller, more localized pods (closer to users, more regional colocation), that can be optics-positive for DCI and metro connectivity because more sites must be interconnected at high bandwidth with low latency, favoring coherent optics and high-speed interconnect between facilities. The principal risk for optics suppliers is timing and margin structure: a faster move to NVIDIA-driven integrated photonics could concentrate bargaining power and compress margins for commoditized transceiver modules while favoring suppliers with differentiated lasers, integration capability, and qualification depth in NVIDIA’s CPO ecosystem. AEC and copper interconnect implications hinge on whether Groq deployment increases the density of short-reach links inside racks and rows. High-speed copper remains structurally advantaged at very short distances on cost, power, and serviceability, but reaches become constrained as lane speeds and aggregate bandwidth rise, creating a role for active electrical cables (AECs), retimers, and signal-conditioning silicon. Credo explicitly positions its AEC products as enabling reliable lossless 800G connectivity for AI clusters, and the company has highlighted participation at NVIDIA GTC with content focused on extending PCIe/CXL using AECs, indicating relevance to next-generation system topologies that require longer reach and higher signal integrity than passive copper can deliver. If NVIDIA turns Groq into a widely deployed inference card or chassis product, the likely near-term effect is AEC-positive because (1) more inference throughput tends to increase top-of-rack connectivity requirements, (2) distributing inference across more racks and sites increases short-reach links per unit of delivered service, and (3) PCIe-attached accelerator architectures tend to require robust signal conditioning as systems move to PCIe 6.x and beyond. Groq workshop materials explicitly reference GroqCard and GroqNode form factors, reinforcing that PCIe-attached deployment has been central to Groq’s current packaging strategy. The main countervailing risk is that Groq’s deterministic chip-to-chip fabric could be implemented primarily through backplanes and direct board-level connectivity that reduces the need for merchant AECs inside the box; in that case, incremental AEC demand would concentrate more in rack-to-switch and node-to-fabric links rather than within-chassis chip fabrics. Astera Labs implications are connectivity-architecture sensitive and, on balance, skew positive if NVIDIA increases heterogeneity and disaggregation in AI systems. NVIDIA has publicly positioned NVLink Fusion as a pathway for partners to build semi-custom AI infrastructure and has explicitly identified Astera Labs as a partner in that ecosystem, with Astera describing NVLink-related solutions expanding its connectivity platform across PCIe, CXL, and Ethernet plus fleet observability software. A Groq acquisition increases the probability that NVIDIA offers a broader menu of accelerators (training GPUs, inference-focused ASICs) and therefore increases the importance of scalable, high-reliability connectivity, retiming, switching, and telemetry across mixed topologies. If Groq silicon remains PCIe-attached in many deployments, PCIe 6.x retimers/switches and active cable modules become more central, aligning with Astera’s core portfolio. If NVIDIA instead integrates Groq concepts into scale-up fabrics (NVLink-like domains) or uses Groq to expand into inference “appliances” that must be rapidly deployed in colocation environments, the need for standard-compliant, serviceable connectivity with strong RAS/telemetry increases, again aligning with Astera’s positioning. Power equipment and cooling implications for Vertiv and adjacent suppliers should be viewed through the lens of rack power density, cooling modality (air vs liquid), and site deployment model (hyperscale campuses vs distributed colocation/enterprise). Groq claims its LPU and rack designs are “air-cooled by design” and require no complex cooling and power infrastructure, and third-party reporting has described Groq’s approach as relying on parallelism across many lower-power units rather than extreme per-chip performance. If NVIDIA scales Groq as a mainstream inference platform, the mix of data center cooling spend could shift modestly away from the highest-density liquid-cooled racks toward more air-cooled or hybrid deployments, particularly for inference pods placed in existing facilities that cannot easily retrofit for very high rack heat flux. That would be a mix headwind for suppliers most levered exclusively to high-end liquid cooling attachments per rack, but it is not necessarily a volume headwind for Vertiv given the company’s broad exposure to both power and cooling infrastructure and the likelihood that total AI deployment locations expand. Vertiv’s own industry commentary emphasizes that AI racks require higher power-density UPS, batteries, power distribution equipment, and switchgear capable of handling rapid load transients, and that hybrid cooling systems will evolve across deployment environments. Those statements align with a world where inference growth increases the count of powered racks and raises the operational complexity of power delivery even if per-rack density is lower than the most extreme training clusters. The most material infrastructure impact may occur outside the rack and upstream of the data hall: grid interconnects, substations, transformers, switchgear, generators, and utility-scale generation additions. Recent regulatory actions in the U.S. highlight that projected data center demand is already driving large planned increases in electricity generation capacity, underscoring that power availability is a binding constraint. In that context, an inference architecture that lowers joules per token could reduce the power required per unit of inference delivered, but it can also accelerate demand by lowering cost and improving latency, increasing the total volume of inference served (a classic rebound effect). The net outcome is likely continued, elevated demand for power infrastructure even if efficiency improves, with the key swing factor being whether AI capex remains on a multi-year growth trajectory or enters a digestion phase. Other data center infrastructure implications include server/ODM mix, facility design standardization, and networking architecture choices. If NVIDIA positions Groq-based inference as a broadly distributable “standard server + accelerator” solution rather than as an integrated, liquid-cooled rack like GB200 NVL72, spend could shift toward more conventional air-cooled server designs, higher unit volumes of mainstream racks, and faster deployment in colocation footprints, increasing demand for modular power rooms, busways, and rapidly deployable cooling solutions. If NVIDIA instead integrates Groq into its “AI factory” paradigm, the primary effect is likely acceleration of dense back-end fabric build-outs and a faster push toward photonics switching, increasing demand for fiber plant, connectors, and integrated optics supply chains while potentially compressing the lifecycle of transitional architectures based on pluggable optics and mid-reach copper. NVIDIA’s stated roadmap toward co-packaged optics and silicon photonics switches is already oriented toward scaling to very large GPU counts; adding a high-end inference ASIC increases the strategic importance of power-efficient, low-latency fabrics because inference economics become increasingly sensitive to network overhead as compute cost declines. Across the covered segments, the most defensible base case is limited near-term dislocation and a medium-term increase in uncertainty around memory intensity per unit of inference growth. HBM faces the clearest relative risk from an HBM-less inference platform, but supply tightness and GPU training roadmaps reduce the probability of an absolute demand shock over the next 12–24 months. Optical, AEC/copper, and power/cooling are more likely to remain volume-supported because they scale with endpoint count, deployment fragmentation, and total data center footprint, and those tend to rise when inference becomes cheaper and more widely deployed. The highest-conviction second-order effect is a shift in infrastructure mix: incrementally more distributed inference deployments (favoring colocation power/cooling standardization, DCI optics, and serviceable short-reach interconnect) and a gradual migration from pluggable optics toward integrated photonics in back-end fabrics (favoring suppliers positioned in the CPO ecosystem).

TheValueist

76,250 просмотров • 8 месяцев назад

EllesmereUI's Biggest Patch since Raid Frames is live! Aug Evokers since Blizzard won't give you an active state on your CDM for Ebon Might, I decided to give you one. check out the video! This feature also allows trinkets/pots/racials to get custom active states Players who share profiles: you can now include your full spell layout (which spells go where) and per-spell settings for CDM when exporting profiles! ----------- Full patch notes with new features and bugfixes: **Profiles:** - **NEW:** Profile exports can now carry your entire Cooldown Manager setup - which spells sit on which bars plus every per-spell setting - for the specs you choose, so importing a profile recreates your CDM layout instantly. **CDM:** - **NEW:** Give any trinket, potion, racial, or custom spell an Active State that adds its own glow and color while active, plus a Cooldown State Effect that changes its look based on whether it's ready. - **NEW:** Sync your trinkets, potions, and racials across specs with one button so you only set them up once. - Sound pickers for Focus Cast Sound and per-buff Audio Effect now have a search box. - Pandemic glow's Apply to All now also syncs tracking bars and keeps the same glow style everywhere. - A trinket, potion, or racial already on another bar now auto-moves to the new bar instead of being grayed out. - A buff saved under two spell IDs now shows as a single icon in the buff bar preview. - Cooldowns you remove from Blizzard's Cooldown Manager now disappear from your previews, while trinkets, racials, and custom spells are kept. - Added tracking for the Nightborne racial (Arcane Pulse), which was missing from the racial list. **Tracking Bars:** - **NEW:** Grouped bars now pack together with no blank gaps, always filling the next available slot. - Eclipse (Solar) and Eclipse (Lunar) now each drive their own bar. - Switching specs while the page is open now refreshes the selected bar correctly. - Pandemic glow now fires for Lifebloom on you or a group member. **Resource Bars:** - **NEW:** A new GCD Bar fills over your global cooldown, with full control over size, position, color, and look. - Expand Power Bar if No Resource now also expands when the class resource is toggled off or disabled for the spec. - Fixed the threshold color not showing with Enhance 5 Bar Style. - The cast bar latency overlay now reads live latency, so spell queueing no longer stops it showing. **Raid Frames:** - **NEW:** New healer tools including heal-absorb text, a crowd-control glow on debuffs, and more ways to position and highlight dispellable debuffs. - The Auto Resize toggle is now a dropdown that scales Indicators & Auras and Tracked Buffs independently with frame size. - New Show Over Dispels toggle lifts the heal-absorb overlay above the dispel gradient. - Fixed custom (non-20-player) sizes loading at the wrong position after login. **Unit Frames:** - **NEW:** A non-tank threat border shadows your player frame when you pull or hold aggro, with Has Aggro and Close to Aggro colors. - **NEW:** Player health text gains Heal Absorb Amount and Heal Absorb Short options. - **NEW:** Mini frames and Boss frames gain a per-frame Bar Texture dropdown. - **NEW:** Boss frames gain a Hover Borders control with its own mouseover and target colors. - **NEW:** Independent Spacing X and Spacing Y sliders for buff and debuff icons. - **NEW:** Absorb and heal-absorb style dropdowns now include your installed SharedMedia textures (also on Raid Frames and Nameplates). - A new Hover Borders control lets you turn the mouseover highlight border on or off per frame. - New Show 2 for Boss option adds a second decimal to boss frame health text. - Buff and debuff Offset X/Y sliders now reach plus or minus 1500. - A new Cast Bar Position cog adds an Offset Y slider for the boss cast bar. - Player, target, and focus cast bars now show above other frames instead of behind them. - Cast bar timer text now has room so it no longer cuts off early. **Action Bars:** - **NEW:** A new When Not Dragonriding visibility mode hides a bar while skyriding and shows it the rest of the time. - The When Dragonriding option now also shows the bar in Druid Flight Form. - The stance bar now shows its GCD swipe even for spells that don't change form. - Bars are briefly forced visible while Myslot's window is open to stop a stall during import/export. **Minimap:** - **NEW:** Hovering the calendar button now shows your raid and dungeon lockouts with boss progress, server time, and time until the weekly reset. - The Omnium Folio button no longer goes missing or drifts after a loading screen, and its position and scale now persist. **Mythic+ Timer:** - **NEW:** A new Show Time Remaining toggle adds an MM:SS countdown to the +2/+3 threshold row that reddens as time runs out. - Fixed timer and detail text cutting off after a font swap. **Colors:** - **NEW:** A new Global Colors section lets you share one profile's custom colors across all profiles or give each profile its own. **Localization:** - **NEW:** Full Russian language support. **Auras, Buffs & Consumables:** - Warrior stance reminders now read the stance bar, so each spec is reminded of its correct stance and clears the moment it's active. - The Inky Black Potion reminder now clears after you drink the potion and reappears on cancel, expiry, or death. - The last-used flask, food, and weapon-enchant preference now saves correctly. **Nameplates:** - A new sync icon on the Pandemic Glow Style row applies the nameplate's pandemic glow to all CDM and tracking bars at once. **Chat:** - The Whisper Sound dropdown gained a search box. **Bags:** - Mythic Keystone dungeon abbreviations now split on hyphens (Nexus-Point Xenas shows as NPX) and handle localized names.

Ellesmere

53,876 просмотров • 2 месяцев назад

$NVDA $MU $SNDK $LITE PAPER OVERVIEW AND CORE CLAIMS The paper “KV Cache Transform Coding for Compact Storage in LLM Inference” introduces kvtc, a transform-coding pipeline that compresses transformer key-value (KV) caches primarily for storage and transfer in LLM serving, rather than for accelerating the per-token attention kernel during active decoding. The method combines 3 stages: (1) feature decorrelation via a PCA basis computed from a calibration dataset and reused across requests; (2) adaptive, variable-precision quantization with bit allocation solved via dynamic programming (DP), including groupwise scaling/shift overhead; and (3) lossless entropy coding (DEFLATE via nvCOMP in the reference implementation) to exploit residual redundancy after quantization. The central empirical claim is that KV tensors contain large, exploitable redundancy across heads and layers, enabling approximately 20× compression versus a 16-bit baseline with negligible degradation across a broad set of accuracy and long-context benchmarks, with materially higher compression (≥40×) available at modest quality cost in some regimes. The system claim is that such compression materially improves the economics of multi-turn, prefix-reuse serving by extending effective KV cache capacity in GPU HBM and host tiers (DRAM/NVMe) and by reducing inter-node and GPU↔host bandwidth demands, thereby improving cache hit rates and reducing time-to-first-token (TTFT) relative to recomputation when caches would otherwise be evicted. KV CACHE AS THE DOMINANT STATE VARIABLE IN INFERENCE ECONOMICS KV cache growth is linear in context length and is multiplicative in layers and attention heads, making it an increasingly dominant constraint as (a) context lengths expand, (b) models add layers and maintain large hidden dimensions, and (c) production workloads shift toward iterative and tool-augmented interactions that repeatedly reuse long prefixes. The paper uses the canonical 16-bit KV cache size formula (4·l·h·d_head·t) bytes and reports 16-bit KV cache sizes per 1K tokens of context that are already operationally large: 128MiB for Llama 3.1 8B, 160MiB for Mistral NeMo 12B, and 320MiB for Llama 3.3 70B Instruct. In binary units, these figures imply per-token KV footprints of 128KiB/token (Llama 3.1 8B), 160KiB/token (Mistral NeMo 12B), and 320KiB/token (Llama 3.3 70B Instruct) at 16-bit. For a 10K-token prompt (10×1K in the paper’s binary convention), the 16-bit KV cache sizes scale to approximately 1.25GiB (Llama 3.1 8B), 1.56GiB (Mistral NeMo 12B), and 3.13GiB (Llama 3.3 70B Instruct). These magnitudes explain why stale caches create a throughput–latency dilemma: retaining them in HBM maximizes responsiveness on future turns but crowds out concurrent sessions; evicting them forces quadratic-cost prefill recomputation and increases TTFT; offloading them to host or storage introduces large transfer overhead and consumes DRAM/NVMe capacity. A key operational nuance emphasized is that modern serving stacks increasingly treat KV caches as a database, leveraging block paging and shared-prefix reuse. In the common disaggregated serving design (separate prefill and decode nodes), KV cache transfer becomes a dominant category of cross-node traffic. Under that design, any reduction in KV cache size directly increases effective fabric capacity and reduces tail latency attributable to congestion, while also enabling longer cache lifetimes in “hot” (HBM) and “warm” (CPU DRAM) tiers that raise cache hit rates and reduce recomputation frequency. The paper’s quantitative example illustrates the economic stakes: a 1,000-line code file tokenized at ~10 tokens/line yields ~10K tokens; for Llama 3.3 70B, an 8-bit KV cache for that context is ~1.6GiB. Reuse across subsequent turns or parallel chats around the same file is valuable, but HBM scarcity makes retaining many such caches infeasible without compression. TECHNICAL MECHANISM: WHY KV CACHES ARE COMPRESSIBLE AND HOW KVTC EXPLOITS IT The technical rationale begins with an empirical observation: keys (and, to a lesser extent, values) across different attention heads can be aligned into a shared latent space using orthogonal transformations (Procrustes alignment). This supports the hypothesis that head-specific projections introduce rotations of a common subspace rather than completely distinct information, implying that concatenating across heads and layers should reveal low-rank structure suitable for linear decorrelation and dimensionality reduction. The method operationalizes this using a PCA/SVD basis learned from calibration data rather than recomputing a decomposition per prompt. This design choice targets production viability: per-prompt SVD is computationally expensive and scales poorly with long prompts and frequent cache updates. kvtc is explicitly structured as an offline-calibrated, online-applied codec: Calibration (performed 1 time per model and compression setting for DP allocation) A calibration dataset is forwarded through the model to collect KV caches. Token positions are pooled, and a subset of positions is sampled. Keys and values are processed separately. Several implementation choices are highlighted as decisive for stability: Rotary positional embeddings are effectively removed prior to compression (“undo positional rotations”), because positional rotations degrade the apparent low-rank structure of keys. “Attention sink” tokens (the earliest tokens in the sequence) and a sliding window of most recent tokens are excluded from compression because they disproportionately affect attention patterns and are empirically more sensitive to reconstruction error. Cross-layer concatenation is used: keys (or values) from multiple layers and heads at the same token position are concatenated along the feature axis to form a higher-dimensional feature vector. PCA is computed over these concatenated vectors, improving robustness relative to per-layer or per-head PCA. The PCA basis is computed via SVD of centered calibration data, using randomized SVD for scalability with a target rank cutoff. The paper reports calibration regimes of 160K tokens for several models with a 10K PCA dimension cutoff (8K for Qwen variants with fewer KV heads), selected to fit within a single 80GB H100 memory envelope and complete within minutes. A critical economic detail is that the same PCA basis can be reused across multiple compression ratios; only the DP-derived precision assignment changes per compression target. Compression (applied between inference phases) Compression operates on stored KV cache tensors, not on weights, and does not modify attention computation. The KV cache is projected into the PCA basis, quantized, packed, and then entropy-coded. Compression is positioned as a background or between-phase operation (after decoding, or between prefill and decode), executed on GPU or CPU depending on where the cache currently resides. The design intent is that compression should not sit on the critical per-token decoding path; it is a storage and transport optimization. Decompression (performed prior to reuse) Decompression reverses the entropy coding and quantization and applies the inverse PCA projection. A practical latency optimization is proposed: inverse projection can be performed layer-by-layer using submatrices of the PCA basis, allowing generation to begin before the full cache is reconstructed, reducing TTFT. Quantization and bit allocation are the core differentiators versus simpler PCA truncation. PCA provides ordered components by variance; kvtc uses DP to allocate a global bit budget across PCA coordinates (and across groups of coordinates) to minimize reconstruction error in the decorrelated domain. Groups of subsequent PCA coordinates share 16-bit shift and scale factors (a microscaling-inspired design), and the DP algorithm jointly selects group size and precision type under a bit budget, including the overhead of per-group metadata. DP commonly assigns 0 bits to many trailing PCA components, which both increases compression and provides a mechanism to trim the PCA basis to the subset of components that actually carry payload, reducing compute and storage overhead of the projection matrices in deployment. Lossless entropy coding then exploits the structure induced by quantization. DEFLATE is used in the reference implementation, and the paper emphasizes that the incremental gain from the lossless stage is content-dependent but meaningful, with an average uplift of ~1.23× on top of quantization in the reported regime. An ablation in the appendices indicates that GPU-friendly variants (GDeflate) can achieve nearly identical compression ratios (≤0.1 difference in measured cases), implying that throughput-optimized lossless codecs can likely be substituted without sacrificing meaningful compression. EMPIRICAL RESULTS: ACCURACY, COMPRESSION, AND LATENCY General-purpose 8B–12B dense models The paper evaluates Llama 3.1 8B, MN-Minitron 8B, and Mistral NeMo 12B across math/knowledge (GSM8K, MMLU) and long-context tasks (Qasper, Lost in the Middle, RULER Variable Tracking) under a simulated multi-turn regime where compression/decompression is applied periodically, with a sliding window of recent tokens excluded. A consistent pattern appears: kvtc maintains near-vanilla performance through 16× compression settings, and remains competitive at 32×, with degradation becoming task- and model-dependent at 64×, particularly on long-context retrieval metrics when compression is pushed aggressively. Selected quantitative anchor points from the paper’s standard-error table (all values are reported with the paper’s evaluation setup and token-window exclusions): Llama 3.1 8B Vanilla: GSM8K 56.8, MMLU 60.5, Qasper 40.4, LITM 99.4, RULER-VT 99.8 kvtc16×: GSM8K 56.9, MMLU 60.1, Qasper 40.7, LITM 99.3, RULER-VT 99.1 kvtc32×: GSM8K 57.8, MMLU 60.6, Qasper 39.4, LITM 99.1, RULER-VT 98.9 kvtc64×: GSM8K 57.2, MMLU 60.7, Qasper 37.8, LITM 90.2, RULER-VT 95.9 These results indicate that, for this model, long-context sensitivity emerges at 64× with meaningful drops in LITM and RULER-VT, while math/knowledge scores remain stable, implying a differential sensitivity consistent with key-vector precision being more critical for retrieval-style behavior. Mistral NeMo 12B Vanilla: GSM8K 61.9, MMLU 64.5, Qasper 38.4, LITM 99.5, RULER-VT 99.8 kvtc16×: GSM8K 62.0, MMLU 64.4, Qasper 37.6, LITM 99.8, RULER-VT 99.5 kvtc32×: GSM8K 62.2, MMLU 63.8, Qasper 37.5, LITM 99.6, RULER-VT 98.7 kvtc64×: GSM8K 61.9, MMLU 61.4, Qasper 38.0, LITM 95.3, RULER-VT 98.0 Here, degradation at 64× is visible but materially smaller than the Llama 3.1 8B LITM drop, suggesting model-architecture or training-data differences can change the tolerance envelope for aggressive KV cache distortion. MN-Minitron 8B Vanilla: GSM8K 59.1, MMLU 64.3, Qasper 38.2, LITM 99.8, RULER-VT 99.4 kvtc16×: GSM8K 60.3, MMLU 64.1, Qasper 38.6, LITM 99.3, RULER-VT 98.8 kvtc32×: GSM8K 59.1, MMLU 63.7, Qasper 37.7, LITM 86.9, RULER-VT 96.0 kvtc64×: GSM8K 57.8, MMLU 62.1, Qasper 38.1, LITM 59.5, RULER-VT 93.4 This model shows markedly higher sensitivity on LITM at 32× and 64×, despite stable short-context metrics, reinforcing that “compression safety” is not monotonic in parameter count and that pruning/distillation choices can alter KV cache redundancy or robustness. Comparisons to baselines The paper compares kvtc to quantization baselines (KIVI, GEAR, FP8) and eviction baselines (H2O, TOVA), plus an SVD-based prefill-optimization method (xKV). Across the reported tasks: Low-bit quantization methods at modest compression (2-bit KV schemes) show earlier degradation in long-context behavior than kvtc at substantially higher compression settings. Eviction methods perform poorly as generic compressors for long-context tasks, consistent with their objective function (selective pruning) being misaligned with “lossless-ish storage for reuse.” xKV shows competitive results on some tasks but a consistent underperformance on Qasper relative to kvtc and vanilla in the provided tables, consistent with method-specific distortions introduced by its decomposition regime. Reasoning models and high-variance tasks For DeepSeek-R1-distilled Qwen 2.5 reasoning models, the paper evaluates AIME 2024/2025 and LiveCodeBench coding. Results are averaged over 8 runs with large variance, but a key inference is that kvtc at ~9×–21× compression achieves broadly similar AIME scores within variance bands, while coding performance remains stable at ~9× and degrades more visibly at ~18×–21× on the 7B model. An important nuance is that smaller reasoning models already have smaller KV footprints (reported ~29KiB/token for Qwen R1 1.5B versus 131KiB/token for Llama 3.1 8B), so the economic value of aggressive KV cache compression is proportionally higher for large models and long contexts than for small models with short contexts, unless the serving system’s bottleneck is dominated by cache transfer rather than HBM capacity. Multi-GPU inference and pipeline parallel For Llama 3.3 70B Instruct run pipeline-parallel across 4 GPUs (20 layers per GPU), the paper compresses KV cache chunks independently per GPU. On MATH-500, the reported accuracy declines from 75.6 (vanilla) to 74.4 at 10× and 72.6 at 20×, with standard errors near ~1.9. NIAH and LITM remain at 100.0 for all tested ratios in that table. The paper notes that joint compression across chunks could improve accuracy for some offload scenarios but is not required for feasibility, highlighting an engineering trade-off between deployment simplicity in distributed settings and optimal global compression. Latency and TTFT economics A critical system result is the measured compression/decompression latency on an H100 for a non-fused implementation. For Mistral NeMo 12B in bfloat16: BS=8, CTX=8K: compression 379ms, decompression 267ms; vanilla recompute TTFT 3098ms; kvtc decompression TTFT 380ms BS=2, CTX=16K: compression 194ms, decompression 143ms; vanilla recompute TTFT 1780ms; kvtc decompression TTFT 208ms These measurements imply that, when a cache would otherwise be recomputed, decompressing a stored compressed cache can reduce TTFT by ~8×–9× in these scenarios, even without kernel fusion. The decomposition of runtime shows PCA projection and entropy coding as the largest contributors, implying that GPU-optimized kernels and faster GPU-native lossless codecs could reduce overhead further. The fundamental economic conclusion is that, in multi-turn settings with long prefixes, compression-induced overhead is likely dominated by the avoided prefill compute and avoided transfer overhead for uncompressed caches. KEY DEPLOYMENT-SENSITIVE DESIGN CHOICES AND FAILURE MODES Several design choices appear to be “hard requirements” rather than optional optimizations: Sink tokens and sliding window exclusions The paper’s ablations show that compressing early “sink” tokens can catastrophically degrade accuracy at high compression ratios (example: Llama 3.1 8B at 64× collapses on multiple tasks when sink tokens are compressed). Similarly, compressing the most recent tokens hurts performance, motivating a sliding window (default 128 tokens) that remains uncompressed. This introduces a predictable engineering constraint: kvtc is not a uniform compression of the full cache; it is a policy-driven, token-position-dependent codec. Production integration therefore requires correct handling of token positions, attention sinks, and window management, and these policies must be aligned with attention-kernel behavior and model-specific sink dynamics. RoPE handling Removing positional rotations prior to compression is described as important for preserving low-rank structure. In deployment, this implies that the codec must be position-aware and must invert and reapply RoPE correctly. This is an additional source of complexity relative to pure per-token quantization and is sensitive to model variants and RoPE parameterizations. Calibration set representativeness The method’s quality hinges on the PCA basis generalizing from calibration data to production data. The paper demonstrates relative stability with 160K–200K calibration tokens and explores domain shifts (general web text vs math traces vs code). Results suggest that moderate domain mismatch is tolerated at 16×–64×, while extreme compression (e.g., 256× in ablations) becomes materially more sensitive to calibration choice. In production, this implies that operators targeting the “negligible degradation” regime should be able to calibrate with broadly representative corpora, while operators targeting ultra-high compression for specialized workloads should expect tighter coupling between calibration domain and achieved quality. PCA matrix storage overhead and operational footprint A non-trivial hidden cost is the need to store PCA projection matrices per model. The paper reports that, prior to DP trimming, PCA matrices stored at 16-bit can amount to a meaningful fraction of model parameter count (examples reported: ~2.4% for Llama 3.3 70B, ~8.7% for Llama 3.1 8B). This overhead is amortized across all cached sessions for a model but competes with HBM/DRAM budgets in multi-model serving. DP-driven trimming can reduce this overhead at higher compression ratios by removing zero-bit components, but the directionality is not guaranteed at low compression ratios if many components remain active. In distributed inference (pipeline parallel), per-chunk PCA can reduce matrix sizes, but may reduce cross-layer decorrelation benefits if fewer layers are concatenated. SYSTEM-LEVEL IMPLICATIONS FOR GENERATIVE AI INFRASTRUCTURE GPU AND HBM The principal infrastructure implication is that KV cache compression at storage time targets the dominant memory allocator stressor in stateful serving: the accumulation of idle or warm conversation state. For workloads with long reusable prefixes (code assistants, enterprise agents with large system prompts, repeated RAG scaffolds, document chat), the limiting resource frequently becomes HBM reserved for KV caches rather than compute. By compressing stale caches by ~20× (or more), the same HBM budget can retain a materially larger working set of cached prefixes, increasing cache hit rates and reducing recomputation. This effect is multiplicative with cache-aware routing and prefix sharing: more prefixes can remain resident (hot or warm) and can be routed to nodes that already hold them, improving both throughput and tail latency. However, kvtc as described does not reduce the active KV cache footprint during the actual attention computation for a currently decoding sequence, because the model operates on decompressed KV caches during decoding. Therefore, the method does not directly reduce HBM bandwidth consumed by attention kernels during steady-state decode, and does not directly address the “memory traffic per generated token” bottleneck that motivates online KV quantization and eviction strategies. The primary HBM benefit is increased effective capacity for caches between turns and reduced HBM pressure from storing many idle sessions, not reduced per-token decode bandwidth. Compression and decompression themselves consume GPU compute and memory bandwidth. The measured decompression TTFT of ~208ms–380ms in the provided benchmarks indicates that the overhead is real but can be materially smaller than recomputation of long prefixes. In an HBM-constrained serving environment, this overhead can be interpreted as a trade between (a) maintaining more caches warm and paying decompression on reuse versus (b) evicting caches and paying full prefill recomputation. The decision boundary will depend on distribution of inter-turn idle times, probability of reuse, and SLA sensitivity to TTFT. kvtc expands the feasible region where keeping caches is economically rational, especially for long prompts. CPU AND DRAM The method implies a stronger role for CPU DRAM as a warm KV cache tier. A ~20× compression ratio changes the practical scale of “warm state” that can be stored per server. Using the paper’s reported KV cache sizes, a 10K-token 16-bit KV cache for Llama 3.3 70B is ~3.13GiB; compressing by ~20× would reduce this to ~160MiB. At that size, storing hundreds to thousands of warm conversation states in DRAM becomes materially more feasible, increasing cache hit rates and reducing NVMe dependence. This can shift system design from “HBM-only hot caches with aggressive eviction” toward “HBM hot + DRAM warm with long retention,” which is structurally analogous to CPU page cache hierarchies in classical systems design. CPU compute implications depend on where compression is executed. The paper explicitly allows compression on CPU if the cache is already in storage, but the strongest bandwidth savings are achieved when compression happens before moving KV caches off the GPU. If an operator chooses GPU-side compression prior to PCIe/NVLink transfer, CPU compute overhead is modest (orchestrating and DP calibration offline). If an operator instead transfers uncompressed caches to CPU for compression, bandwidth savings are forfeited and CPU memory bandwidth becomes a bottleneck. Therefore, the most economically coherent deployment path is GPU-native compression/decompression with CPU DRAM used as the warm storage reservoir.

TheValueist

16,549 просмотров • 7 месяцев назад

$NVDA $GFS NVIDIA’s reported agreement to acquire Groq for $20B in cash (per CNBC, amplified via Reuters and other wire coverage) represents a materially different strategic posture than NVIDIA’s prior M&A pattern, given both the headline size (largest reported NVIDIA acquisition to date) and the unusual carve-out that Groq’s early-stage cloud business would not be included. Public reporting indicates the information originated from Alex Davis, CEO of Disruptive (lead investor in Groq’s latest financing), and that neither NVIDIA nor Groq had issued an immediate confirmation at the time of publication. The same reporting frames the transaction as coming together quickly, only months after Groq raised $750M at a ~$6.9B valuation, and highlights Groq’s positioning as a high-performance inference chip vendor founded by ex-Google TPU engineers. Groq is best understood as a vertically integrated inference acceleration company whose core asset is an application-specific processor optimized for deterministic, low-latency execution of transformer-style workloads, paired with a compiler-led software stack and a distribution layer (GroqCloud) designed to reduce developer friction via OpenAI-compatible APIs and integrations. Groq brands its architecture as a Language Processing Unit (LPU) and consistently emphasizes that the design target is inference, not training. The company’s own architecture description centers on 1-core execution, large on-chip SRAM used as primary storage (explicitly not cache), a custom compiler that statically schedules compute and communication, and direct chip-to-chip connectivity intended to coordinate multi-chip execution without relying on conventional caching hierarchies or dynamic runtime scheduling. The technical premise is a deliberate inversion of the conventional GPU approach. GPUs deliver throughput via massively parallel, multi-core execution with dynamic scheduling, complex memory hierarchies, and heavy reliance on off-chip HBM bandwidth and sophisticated runtime/kernel optimization. Groq instead argues that inference bottlenecks are driven by latency variance (tail latency), synchronization overhead, and memory access unpredictability inherent in dynamically scheduled, cache-heavy architectures, particularly when workloads are latency sensitive and batch sizes cannot be inflated. Groq’s solution is to move “control” into the compiler: the full execution graph and inter-chip communication schedule are computed ahead of time down to clock-cycle granularity, with deterministic execution designed to reduce run-to-run variance. In Groq’s framing, the removal of caches, reorder buffers, speculative execution overhead, and other sources of contention enables predictable latency and high utilization without per-model kernel engineering typical of GPU tuning cycles. A critical nuance is that Groq’s determinism is not merely a software claim; it is tightly coupled to architectural constraints and system design choices that trade flexibility for predictability. Third-party technical commentary indicates Groq’s chip uses a fully deterministic VLIW-style approach with minimal buffering, no external memory, and heavy dependence on sharding models across many chips because on-chip SRAM capacity is limited. SemiAnalysis describes a ~725 mm^2 die on GlobalFoundries 14nm with ~230MB of SRAM and notes that “no useful models” fit on a single chip, forcing multi-chip partitioning for modern LLMs and driving a system-level design where networking and compilation are first-class scheduling problems rather than ancillary infrastructure. This is consistent with Groq’s own messaging that tensor parallelism across chips is a primary design goal, enabled by large on-chip SRAM and compile-time coordination of compute plus interconnect. The on-chip SRAM emphasis is central to Groq’s latency story and also its most constraining trade-off. Groq claims on-chip SRAM bandwidth “upwards of 80 TB/s” and contrasts that with off-chip HBM bandwidth “about 8 TB/s,” asserting a potential 10x advantage from bandwidth plus reduced trips across chip-to-memory boundaries. While these comparisons are marketing-oriented and depend on workload specifics, the architectural implication is clear: Groq prioritizes ultra-fast local weight/activation access and then scales capacity by adding chips, not by attaching large off-chip memory pools. This design can reduce latency for sequential inference layers and minimize unpredictable stalls, but it pushes complexity into partitioning strategy, interconnect topology, and compiler scheduling, and it increases the number of chips needed for very large parameter counts and large KV-cache footprints. Groq also highlights numeric formats and compiler-driven precision management as a performance lever. In its 2025 technical blog, Groq describes “TruePoint numerics,” including 100-bit intermediate accumulation and selective quantization choices (FP32 for attention-sensitive operations, block floating point for MoE weights, FP8 storage in error-tolerant layers), and claims 2-4x speedups versus BF16 without measurable accuracy degradation on benchmarks such as MMLU and HumanEval. Even if the absolute uplift is workload dependent, the strategic point is that Groq is pursuing performance via end-to-end co-design: precision policy is not just hardware capability (FP8/BF16) but compiler-enforced mapping of precision to error sensitivity, which can matter materially for inference cost-per-token if it reduces memory traffic and boosts throughput without forcing aggressive, accuracy-damaging quantization. Independent performance datapoints indicate Groq has been credible on latency-oriented inference speed, at least for certain regimes. EE Times reported in 2023 that Groq demonstrated Llama-2 70B inference at ~240 tokens/s per user on a cloud-based dev system described as 10 racks and 64 chips, using the company’s 1st-gen silicon introduced several years earlier. Separate Groq commentary around independent benchmarking cites results showing ~241 tokens/s throughput and ~0.8s time to receive 100 output tokens for a Llama-2 70B API configuration, positioning the platform as a step-change in “available speed” for certain interactive use cases. These figures do not settle total cost-of-ownership versus GPUs or hyperscaler ASICs, but they establish that Groq’s system-level architecture can deliver strong single-user throughput and latency on large models when properly partitioned and scheduled. GroqCloud is the commercial wrapper that packages this hardware/software stack as “tokens-as-a-service,” aiming to make Groq adoption feel like switching API endpoints rather than adopting new silicon. Groq’s documentation states its API is designed to be “mostly compatible” with OpenAI client libraries, and its pricing page provides model-specific token rates, published speeds (tokens/s), prompt caching discounts, and batch processing discounts. For example, pricing lists inputs as low as $0.05 per 1M tokens and outputs as low as $0.08 per 1M tokens for certain smaller LLM configurations, with higher prices for larger models and long-context or MoE variants; it also advertises prompt caching with a 50% discount on cached input tokens for certain models and a batch API offering 50% lower cost for asynchronous processing windows. These mechanics are economically important because they demonstrate Groq’s go-to-market is not simply “sell chips,” but “sell predictable unit economics per token,” with tooling (batch, caching) that directly targets inference cost drivers (reused prompts, throughput smoothing, and asynchronous workloads). The cloud footprint and distribution partnerships indicate Groq has been building an inference-native “edge within the cloud” strategy rather than competing head-on with hyperscalers on breadth of services. A 2025 Groq newsroom release describes a European deployment in Helsinki with Equinix, positioned as latency reduction and data governance for European customers, and explicitly references Equinix Fabric enabling private connectivity to GroqCloud over public, private, or sovereign infrastructure. The same release enumerates additional capacity in the U.S. (Equinix, DataBank), Canada (Bell Canada), and Saudi Arabia (HUMAIN), and states these sites collectively served more than 20M tokens/s across Groq’s global network at that time. That supply-side metric matters because it provides a directional sense that Groq is scaling capacity as a network, not merely as a chip vendor. Customer disclosure is inherently limited because Groq is private and many enterprise deployments are not public, but Groq’s marketing materials and partnerships provide signals about demand vectors. The company’s public website displays logos of large consumer and enterprise brands (e.g., Dropbox, Vercel, Chevron, Volkswagen, Canva, Robinhood, Riot Games, Workday, Ramp) and includes a published customer quote claiming a 7.41x chat speed increase and an 89% cost reduction after moving to GroqCloud, followed by a tripling of token consumption. While marketing claims should be treated as case-specific and not generalized, they indicate that Groq is targeting both AI-native developers (who measure success by latency and cost-per-token) and enterprise buyers (who care about predictable performance and governance). Supplier and dependency mapping for Groq spans 3 layers: silicon production, system integration, and cloud infrastructure. On silicon, third-party analysis indicates GlobalFoundries 14nm for the 1st-gen Groq chip, implying a supply chain less constrained by the most capacity-tight leading-edge nodes and advanced packaging bottlenecks that dominate high-end GPU supply (HBM stacks, CoWoS-type packaging constraints). If accurate, this is strategically meaningful because it suggests Groq capacity expansion could be gated more by conventional wafer supply, board assembly, and data center power than by the same HBM/advanced packaging scarcity that has constrained top-tier GPU ramp cycles. On systems and cloud, Groq’s own releases identify colocation and connectivity partners (Equinix, DataBank, Bell Canada) and a Middle East partner (HUMAIN), implying dependencies on data center real estate, power availability, and network connectivity, alongside procurement of standard server components, NICs/switching, racks, and cooling infrastructure. The Groq design narrative also emphasizes air cooling and reduced need for complex power/cooling infrastructure, which—if realized in deployments—can widen the set of feasible hosting locations and lower deployment friction relative to liquid-cooled, very high power density GPU racks. Against that backdrop, the strategic rationale for NVIDIA acquiring Groq can be framed as a set of overlapping objectives: inference silicon optionality, architectural hedging, competitive defense, and supply chain diversification, with the carve-out of GroqCloud signaling a preference to avoid direct cloud competition and to focus on IP and product portfolio control rather than operating a capital-intensive token-serving business. The deal, if confirmed, would occur at a valuation step-up of ~190% versus Groq’s reported ~$6.9B private valuation in the September $750M round, reinforcing that any acquisition logic would be predominantly strategic rather than a conventional financial multiple arbitrage. The most compelling strategic driver is inference. Training has historically been the center of gravity for cutting-edge GPU demand, but inference volume is structurally larger and more distributed as deployments scale, with economics dominated by cost-per-token, latency guarantees, and utilization under spiky demand. Inference workloads also create a strategic vulnerability for NVIDIA: hyperscalers and large platforms can justify bespoke ASICs (TPU, Trainium/Inferentia, Maia-class efforts) because inference is stable, repeatable, and can amortize software investment at massive scale. Groq’s core proposition—deterministic, compiler-scheduled inference with predictable latency—aligns directly with the segment where GPU generality is least valued and where “good enough” programmability plus superior unit economics can win share. Acquiring Groq would allow NVIDIA to own a credible inference-native architecture rather than relying solely on GPUs and software optimization to defend that segment. Competitive defense logic is also plausible. Groq occupies a specific competitive wedge: low-latency, high-throughput interactive inference, delivered via a simple API abstraction that reduces switching cost. That wedge directly pressures GPU inference margins in the long run because it makes inference price/performance comparisons more transparent at the token level, and it targets a developer persona that historically defaulted to CUDA-first ecosystems. Even if NVIDIA’s current-generation systems can achieve very high tokens/s per user with extensive optimization, the strategic risk is that competing architectures normalize the idea that inference is best served by special-purpose silicon with a simpler programming model, weakening CUDA lock-in at the application layer. NVIDIA has actively demonstrated that Blackwell-era systems can exceed 1,000 tokens/s per user in benchmarked configurations, but that performance leadership does not automatically translate to lowest cost-per-token across the full range of batch sizes, latency targets, and deployment environments. Groq’s existence as a credible alternative architecture forces NVIDIA to keep defending inference economics rather than only raw performance leadership. The “technology acquisition” rationale is unusually strong in this specific case because Groq’s differentiator is not a single block of silicon IP but an end-to-end methodology: compiler-led static scheduling, deterministic networking, and a system architecture designed around tensor-parallel inference rather than throughput-maximizing batch inference. NVIDIA’s stack is already compiler-heavy (TensorRT, Triton, CUDA graphs, kernel fusion, speculative decoding techniques), but GPUs remain dynamically scheduled devices with complex memory hierarchies and stochastic latency behaviors under contention. Groq’s approach provides an alternate design point: treating the entire inference execution (compute plus communication) as a statically schedulable program. In principle, that IP could be valuable even if Groq silicon itself is not adopted at massive scale, because it can inform how NVIDIA builds future inference-optimized products, compilers, and networking fabrics, especially as distributed inference with large models makes communication a first-order performance determinant. Supply chain diversification is a non-obvious but potentially important driver. If Groq’s mainstream product generation is truly based on a mature process node and avoids HBM, then the scaling constraints look different than those of state-of-the-art GPUs. NVIDIA’s ability to meet incremental demand has been tightly coupled to advanced packaging and HBM supply, and those constraints can remain binding even when wafer supply is available. An inference ASIC architecture that relies primarily on on-chip SRAM and scales by adding chips—while not costless—could reduce dependence on HBM availability and advanced packaging capacity, enabling NVIDIA to ship “inference capacity” in higher absolute volumes or into geographies and customer segments where the highest-end GPUs are economically or logistically difficult to deploy. This could be particularly relevant for latency-sensitive inference deployed in regional colocation footprints rather than centralized hyperscale campuses. The carve-out of GroqCloud, if accurate, is itself a strategic signal about NVIDIA’s priorities. Operating a token-serving cloud at scale is capital intensive, structurally lower margin than silicon IP rents, and creates channel conflict with hyperscalers and CSP partners who are core NVIDIA customers. NVIDIA has generally positioned its cloud offerings through partnerships rather than as a direct hyperscale competitor. Excluding GroqCloud would preserve neutrality with CSPs and avoid inheriting multi-region data residency obligations and partner contracts, while still allowing NVIDIA to acquire Groq’s silicon, compiler technology, and engineering talent. At the same time, excluding GroqCloud would also mean NVIDIA would not automatically acquire the commercial proof-point of Groq’s unit economics or the customer contracts that validate product-market fit at scale, increasing the importance of diligence on whether Groq’s cloud pricing is structurally profitable or partially subsidized by fundraising. There is also a “preemptive acquisition” angle. The reporting identifies recent investors in Groq’s latest round including large financial institutions and strategic/industry players. In that context, Groq represents an asset that could plausibly have been acquired by a competitor (AMD/Intel) or by a hyperscaler seeking to accelerate inference independence. NVIDIA acquiring Groq could be a defensive move to prevent a credible inference-native architecture from being weaponized by a rival with deep distribution. Even if GroqCloud is carved out, controlling the silicon roadmap and compiler IP would meaningfully constrain Groq’s ability to evolve into a standalone competitor, unless the carved-out entity retains long-term rights to the hardware and software stack. However, the strategic case is not one-sided; there are meaningful risks and potential contradictions that would need to be reconciled for the transaction to be value-accretive on a multi-year horizon. 1st, Groq’s architecture appears to rely on scaling out chip count to achieve capacity, which introduces system cost, networking complexity, and physical footprint considerations. The absence of external memory and limited on-chip SRAM implies very large models require substantial chip parallelism, and the economics then depend heavily on chip cost, yield, power efficiency, and interconnect overhead. SemiAnalysis explicitly frames Groq as trading space for time and raises questions about token economics and whether publicly advertised pricing reflects fully loaded costs or market share capture. 2nd, integration risk is non-trivial. Groq’s compiler-led deterministic model is philosophically and practically different from CUDA’s dominant programming and execution model. A poorly executed integration could create internal product confusion, dilute engineering focus, or alienate developers if the combined stack fragments. 3rd, there is cannibalization risk. If Groq-class inference silicon undercuts GPU inference economics, NVIDIA could face internal margin trade-offs, even if the goal is to defend share against hyperscaler ASICs. Cannibalization can still be rational if it prevents larger share loss, but it would require crisp portfolio segmentation and go-to-market discipline. The presence of NVIDIA’s own rapidly improving inference performance complicates the “need” for Groq but does not eliminate the “option value.” NVIDIA has demonstrated benchmark-leading tokens/s per user on Blackwell-based systems, suggesting that raw interactive throughput is not necessarily the limiting factor for NVIDIA’s product line. The more enduring strategic question is unit economics and architectural control: whether future inference demand is better monetized through general-purpose GPUs plus software optimization, or whether a bifurcated product portfolio (training GPUs plus inference-native ASICs) becomes necessary to defend total AI compute wallet share as hyperscaler ASIC penetration increases. Acquiring Groq could be a decisive move to ensure NVIDIA participates in both regimes rather than betting exclusively on GPUs to win inference forever. What is “special” about Groq’s technology relative to a typical accelerator roadmap is the tight coupling of determinism, compilation, and networking into a single scheduling problem. The LPU narrative emphasizes deterministic compute and networking, static scheduling, and direct chip-to-chip coordination that allows “hundreds” (more precisely, 100s) of chips to behave like a single scheduled resource. The architecture also explicitly targets tensor-parallel, latency-optimized distribution rather than pure data-parallel throughput scaling, which matters for real-time applications where a single response must arrive quickly rather than many requests being processed in bulk. The implication is that Groq is optimized for the time-to-first-token and steady token streaming behavior that defines user experience in interactive LLMs, and it attempts to achieve that without relying on large batch sizes that can degrade latency. From a portfolio manager’s perspective, the most important interpretation is that an NVIDIA-Groq combination would likely be less about “NVIDIA needs more inference speed” and more about controlling the architectural trajectory of inference acceleration and removing a fast-improving, developer-friendly competitor from the market. The carve-out of GroqCloud would reinforce that the transaction is aimed at IP, talent, and product optionality, not acquiring a cloud revenue stream. The valuation step-up implied by $20B versus $6.9B would therefore be justified only if the acquired assets materially reduce long-term competitive risk (hyperscaler ASIC displacement, inference margin compression) or enable new monetization vectors (inference ASIC product line, supply chain de-bottlenecking, improved software determinism) that would be difficult to achieve on a comparable timeline via internal R&D.

TheValueist

102,145 просмотров • 8 месяцев назад

Grok Grok Bot is incredible - and they can even manufacture real physical objects! Here is a little experiment I did last night: I created a team of bots and asked them to solve a complex engineering problem end to end - starting from four images as design cues, inferring transferable structural principles from the pixels, synthesizing an executable interactive physics simulator, running and reasoning over experiments, optimizing the design & finally manufacturing the best designs. The entire loop worked remarkably well - and I was even able to communicate with the agents from my Apple Watch. (Do we live in the future yet?) Team of agents 1⃣ Chief of Staff coordinates the workflow: watches the other agents, pulls results into the main chat, transfers files between them, and keeps the job moving. 2⃣ Physics Experimenter is the scientist-coder. It interprets the design cues and images, writes the simulator, runs experiments, analyzes the results, and produces a detailed LaTeX scientific report. 3⃣ 3D Printing Bot operates the fabrication workflow: prepares and slices the models, generates manufacturing code, sends the job, and monitors the printer. The workflow I provided an initial task based on four unregistered reference photographs containing different objects at different scales (pinnate leaf venation, a Voronoi-like areole mesh, a stochastic fibrous lattice, and a radial/circumferential web). The prompt asked the agents to infer transferable design principles - hierarchy, branching, interfaces, redundancy, disorder, load paths - and use them to build an interactive laboratory for hierarchical materials and fracture. The scientific question was: at fixed material budget, how do hierarchy depth, redundancy, disorder, and interlevel strength change stiffness, peak load, energy absorption, and the brittle-to-progressive transition? In ~20 minutes, the Physics Experimenter produced a 2D hierarchical Euler–Bernoulli beam-network laboratory. Coarse veins persist and remain thicker; finer infill is added inside cells; members connecting levels are treated as interfaces with relative strength κ; and total material volume is conserved. The four source photographs remain visible in an editable interpretation panel. The app generates geometry, steps or runs the network to failure, compares A/B/C designs, and exports JSON, CSV, PNG, and STL geometry for fabrication. After validation the Physics Experimenter used the app and conducted 47 simulation experiments, including six holdouts. It found something scientifically interesting: extra hierarchy is not "free" toughness. At fixed volume, initial stiffness changed by only about 20%, while work-to-failure varied by several-fold. Infill steals cross-section from the main axial veins, so deeper and more redundant networks often absorbed less energy than a simple depth-1 grid. Weak interfaces behaved as distributed fuses, producing more progressive failure and reducing localization. The specific H2 hypothesis - that hierarchy becomes detrimental primarily because interfaces form a mechanical bottleneck - was rejected; the dominant effect instead came from redistribution of a fixed material budget across structural levels. The Physics Experimenter then assembled the methods, tests, results, hypothesis evaluation, and conclusions into a detailed scientific report. The best designs were passed to the 3D Printing Bot. It opened Bambu Studio and brought the Bambu Lab H2D online. Both STLs were placed on one build plate at the same 50x scale and sliced using a 0.20 mm PLA process. The prints completed within less than an hour. The loop images → structural abstraction → executable physics → autonomous experiments → hypothesis testing → design selection → STL → slicing/manufacturing code → physical object That last transition is what I find especially interesting: AI is beginning to operate across the entire scientific and physical workflow - converting observations into models, models into experiments, experimental evidence into revised designs, and those designs into manufactured matter by directly operating machines. This starts to blur the boundary between AI that reasons about the physical world and AI that can actually act on it. Shoutout to the Grok Bot team - you are building something very special here! The way these agents can move naturally from reasoning, to experiments, to operating machines in the physical world feels like an important step.

Markus J. Buehler

1,124,112 просмотров • 26 дней назад

一番最後の[Prompt for original image]の部分に画像生成に使用したPromptを入れると一貫性が増します。不要な場合は3行削ってしまっても大丈夫です。 --- Extreme wide-angle perspective and dynamic pose remix edit. This is an EDIT of the original image, not a new character. Use the original image as a strict reference for: – the person’s identity, hairstyle, and overall fashion style, – the general type of background and location (same street, same room, same beach, same kind of architecture, etc.). You are allowed to completely change the camera position, angle, and pose, but you must keep the scene in the SAME location and keep the SAME person and outfit design. Camera and perspective: – Use an ultra wide-angle or fisheye feeling lens (around 12–18mm full-frame look). – The camera angle MUST change significantly from the original: use dramatic angles such as • worm’s-eye view from directly below looking up, • bird’s-eye view from directly above looking down, • very low angle from the ground, • high angle from above, • tilted Dutch angles. – Always create strong foreshortening: body parts close to the lens look huge, while the rest of the body falls away in perspective. – The final result must look like a bold fashion or street photo, fully photorealistic, not illustration or anime. Background consistency: – Keep the same location as the original image: same street, same bridge, same room, same studio, same beach, same general structures and materials. – Do NOT replace the background with a completely different place. – Because the camera angle changes, it is allowed and expected that different parts of the environment become visible. – When new areas appear, extend the original environment logically (same buildings, fences, road markings, walls, colors, materials, lighting style), as if the camera moved within the same place. Body parts near the lens (1–2 parts, sometimes 3): – In each edit, choose ONE or TWO main body parts to be extremely close to the lens (sometimes even THREE in more complex poses). – Vary them from image to image, do NOT always use the same body part. – Allowed near-the-lens parts include: • one or both hands / fingers reaching toward the camera, • one or both feet / shoes / boots near the lens, • knees or thighs, • face very close to the lens, • shoulders or chest close to the lens in a leaning pose. – The chosen body parts should come extremely close to the lens, almost touching it, with visible skin texture, fabric texture, and realistic wide-angle distortion. Pose and overall body (complex and varied): – Create strong, cool, dynamic poses that match the extreme perspective. – Randomly use different pose types, including: • standing with one leg or one arm reaching toward the camera, • crouching or squatting low to the ground, • sitting on the floor or on objects, • lying on the ground with legs or feet toward the lens, • leaning forward aggressively toward the camera, • twisting the body, crossing legs, or arching the back for more dynamic lines. – Allow complex poses where: • both hands are near the lens forming shapes (peace signs, triangles, frames, pointing toward the viewer), • both feet are toward the lens, • one hand and one foot are both large in the foreground, • the face is close to the lens while hands or feet are also visible in perspective. – Maintain believable anatomy even with extreme foreshortening. Angle and attitude (randomized): – Randomize camera angle and orientation (up, down, side, Dutch tilt) while keeping the composition visually balanced and powerful. – Keep the vibe cool, confident, and fashion/editorial or street style, depending on the original outfit. – Facial expressions can vary (serious, playful, confident, mysterious), but must still look like the same person. Lighting and rendering: – Keep the general time of day and lighting mood similar to the original (night vs day, indoor vs outdoor, soft vs hard light), but you may enhance contrast and color to make the image punchy and dramatic. – Maintain realistic shadows and contact points with the ground or floor. – High-resolution, sharp details with clear skin texture, fabric weave, and material highlights. Variation and randomness: – Each edit should look noticeably different from the original image and from other edits, with different: • camera angles, • pose types, • which body parts are closest to the lens, • orientation (straight, tilted, from above, from below). – Avoid repeating the exact same single-foot-close-up composition; produce a wide variety of dynamic poses and angles. Strict rules: – Do NOT change the person into someone else. – Do NOT change the outfit type; only restyle it through pose, perspective, and small natural movement of clothing. – Do NOT move the scene to a completely different location; always stay in a plausible extension of the original place. – Do NOT add text, logos, watermarks, or graphic design elements. – Do NOT switch to painting, illustration, or anime style; keep it photorealistic. Overall: Transform the original photo into a dramatic, photorealistic, ultra wide-angle shot with an extreme camera angle (including views from directly below or above), where one or more body parts are right next to the lens and look huge, the rest of the body recedes in perspective, and the same person strikes a stylish, complex, powerful pose in a consistent, expanded version of the original environment. Also, below is the prompt for generating the original image. Please use it as a reference. [Prompt for original image] #nanobanana2

AI Girl's Photo Studio

20,684 просмотров • 9 месяцев назад

🚨3I/ATLAS Is Carrying Fusion Fuel at Impossible Levels and the Mainstream Explanation Doesn't Hold Dr. Avi Loeb is laying out something about 3I/ATLAS that is buried inside a technical discussion, one that you and I may not get to hear about in mainstream science. However the point that Loeb is making is really difficult to ignore when you listen to the facts. We are all aware by now that we are dealing with an interstellar object, something that didn't originate in our solar system, that passed through and was observed closely enough for its chemical composition to be analyzed. That was an amazing opportunity, but what came out of that analysis is where things start to get spicy. The reason I say that this is interesting is because of deuterium. Yes it is a known isotope of hydrogen with an extra neutron and yes it exists everywhere, but only in very small quantities. Across the universe, the ratio is remarkably consistent. Roughly one atom of deuterium for every 50k atoms of hydrogen. That number doesn't change much whether you're looking at stars, gas clouds, or planetary systems. Even in places where it's slightly elevated, like Earth's oceans, it's still nowhere near significant enough to stand out in a major way. It's measurable, but it doesn't dominate anything. That's the baseline that we have to make comparisons from. Now take that baseline and compare it to what was measured in 3I/ATLAS. Instead of one in 50k, you're looking at something closer to one in a hundred in water, and one in thirty in methane. That is a huge jump and once you take that into consideration you're no longer talking about natural variation in any conventional sense at least. The first explanation is the one you would expect. Extremely cold environments, possibly tied to very early star formation, where deuterium can be preserved more efficiently than in regions like our own solar system. All that explanation does is give you a place to put the anomaly without breaking anything, aka mainstream scientific models. But it doesn't actually resolve the full picture. Here's why... The same object showing this deuterium enrichment also contains heavier elements like carbon and oxygen in ways that don't align with those early environments. The universe at that stage didn't have enough of those elements available in the right quantities to produce what we're seeing now. So what you end up with is a contradiction. The conditions that could explain the deuterium don't support the rest of the chemistry, and the conditions that support the chemistry don't explain the deuterium. That's where the conversation conversation obviously becomes difficult for the 'tenure' crowd, because when formation models stop lining up as predicted by archaic models, you're left with a narrower set of possibilities. Either there's a process we don't yet fully understand that can produce this combination, or something has happened to the material after it formed. Considerations by non mainstream science would be that this is not random alteration, but something more deliberate. Processing, concentration, separation steps that could possibly mean function rather than accident. This is where deuterium stops being just an interesting anomaly and starts mean something very different. Deuterium is one of the primary fuels used in nuclear fusion. Every serious attempt to build a functional fusion reactor on Earth relies on it, typically in combination with tritium. It's efficient, predictable, and it's exactly the kind of material you would isolate and concentrate if you intended to use it as an energy source. So when you see an object carrying deuterium at levels this far beyond any natural baseline we observe locally, we have to wonder what conditions would allow that concentration to exist, and whether those conditions are passive or active. That doesn't automatically push you into extreme conclusions, but it does move you out of the safe 'mainstream' zone where everything can be explained with known processes. That's the part that tends to get softened in how this is presented publicly of course. There's a difference between saying something is unusual and admitting that it doesn't currently fit within the models we rely on. One side invites curiosity whilst the other invites scrutiny. What you're seeing here is that tension in real time because the data is absolutely clear enough to acknowledge the anomaly, but the interpretation is being held just short of where it would need to go to fully confront it. So what you're left with is a set of open questions that aren't being pushed by mainstream science, and I am sorry if it sounds like I have a drum to bang, but here we are. Could this be evidence of a type of cosmic environment we haven't observed directly yet, one capable of producing extreme isotopic enrichment alongside complex chemistry? Is there something in the way that we're measuring or interpreting the data that's creating a misleading picture of the ratios? Or are we looking at material that hasn't remained in a purely natural state since its formation? That last question is the one that tends to sit just beneath the surface, acknowledged but not explored too directly and that's not because it's impossible, but because of what it might imply if it turned out to be true, and that's where this becomes worth paying attention to. If this isn't an isolated curiosity then it is yet another one of the already stacked list of anomalies tied to interstellar objects, unusual motion, unexpected structural behavior, and now chemical signatures. Each one on its own can be managed, explained, or set aside, but taken together, they start to form a pattern that really does warrant further consideration. 3I/ATLAS may still end up having a natural explanation and I have always maintained that is always on the table. But if that explanation exists, it's not something we've defined yet, and it's not something that fits inside of current models, and until it does, the signal remains what it is. An object from outside our system, carrying a level of fusion capable material that doesn t match anything we see in our own environment, tied to formation theories that don't fully hold up under scrutiny. #UAP #InterstellarObject #3IATLAS #SpaceAnomalies #FusionFuel #JWST #Astrophysics #UFOtwitter #Disclosure

Skywatch Signal

28,445 просмотров • 5 месяцев назад

The U.S. MUST win the AI race We’ve implemented a clear policy at micro1: we will only work with U.S. AI labs and its allies. We made this decision because the AI race is not just about better products. It is about who controls the intelligence layer of the global economy, and whether frontier capability is used to strengthen the free world or to empower adversarial states. AI will be the most important technology of our lifetime. In the fullness of time, it will automate most functions across the economy. Not just software tasks, but coordination, production, logistics, judgment, and execution. As those functions are automated, human time is freed up to invent new ones. Those new functions then become candidates for automation themselves. This loop compounds. As this trajectory continues, output per worker increases dramatically. Entire categories of work become cheaper and faster to perform. Manufacturing reshoring becomes economically viable not because of policy intervention, but because intelligent systems operated domestically outperform global labor arbitrage. Goods and services trend toward lower marginal cost, while distribution improves through better coordination of supply and demand. That is the upside. However, this is impossible without deep integration of intelligent systems. For AI to meaningfully automate real-world functions inside enterprises or governments, it needs full context of any given enterprise. That means read and write access to its core databases. There is no credible path to automating high-impact functions without granting frontier systems that level of access. If the United States does not win the AI race, enterprises eventually face a constrained choice. Either grant that access to Chinese models controlled by an adversarial government, or rely on sub-optimal intelligence to automate functions that still must be automated. Both outcomes are not acceptable. And ultimately, this becomes the greatest national security risk the United States has ever faced. AI models are trained by humans. The judgment embedded in pre-training data and especially in expert post-training data largely determines how a model behaves. While emergent behavior exists, a useful approximation is that a model reflects the weighted aggregate of the human judgment distilled into it. Assisting foreign actors—who will naturally prioritize expert tasks aligned with their own interests—to dominate data creation embeds those interests directly into the intelligence layer itself. Once encoded at scale, these interests propagate through every downstream applications that relies on that intelligence. Here’s how we win. First, leverage is in software. China is ahead in hardware for physically intelligent systems. Catching up there is a long and difficult battle. Software, both large language models and robotics models, remains the bottleneck. Advancing the brain (AI models) is the fastest way to increase the usefulness of existing hardware and deployed systems. Second, the U.S. must 100x its investment in structured human judgment. Continued investment in compute and algorithmic efficiency is critical. But that investment is ultimately a bet on very high future inference demand. For that bet to pay off, models must unlock many new capabilities, and in practice the only way to unlock those capabilities is through expert human data. Historically, experts like doctors and lawyers were never incentivized to produce high-quality reasoning data in a machine-verifiable format. There was no reason for a doctor to generate precise, structured simulations of patient interactions, diagnostic reasoning, or treatment tradeoffs. There was no reason for a lawyer to document complex legal reasoning paths in a way that could be programmatically evaluated. AI systems now require exactly this kind of data. The incentive finally exists because this data directly improves systems that operate at massive scale, and experts can be paid well to produce it. Once expert judgment is encoded into models in a structured, verifiable way, it compounds. Those who delay do not just lose time. They lose the ability to catch up. Third, distillation from Chinese labs must be stopped. AI labs must do everything they can to prevent Chinese labs and models from distilling frontier models. Simply calling frontier APIs, or even interacting through UIs, lets Chinese model companies rapidly generate high-quality supervised fine-tuning datasets and close the gap at a fraction of the cost. This method does not put you at the frontier, but it does let you catch up quickly, which is what we saw with DeepSeek. The West significantly overreacted to DeepSeek’s headline capabilities, but underreacted to the underlying dynamic: frontier access itself becomes a training set at a fraction of the cost. Human data platforms also have a duty to help prevent this distillation. Lastly, the U.S.government should set the standard for AI Evaluation that leads to real production usage. AI agents are under-deployed relative to what the technology allows because they are probabilistic systems that require a fundamentally different QA approach than deterministic software. Generic QA is insufficient; safely shipping agents requires explicit evaluation frameworks that assess their full action space. Organizations must clearly define which functions an agent is allowed to perform, how quality is measured for each function, and which domain experts are qualified to judge outcomes. With these frameworks in place, agents can be rigorously tested using structured human data, deployed to production with confidence, and continuously improved over time. The U.S. government should be the first large enterprise to implement rigorous evaluation systems across every function. If the government leads on evaluation-driven deployment, adoption across the private sector accelerates naturally. This is how American workers become more powerful. Each worker operates digital or physical agents that expand their effective output. Recruiting, manufacturing, logistics, and other domains shift toward human judgment overseeing autonomous execution. Reshoring occurs because it becomes economically rational. Work becomes more meaningful. This is a race to determine who controls the intelligence layer of the global economy. And that must be us. 🇺🇸

Ali Ansari

397,008 просмотров • 7 месяцев назад

how to produce long form documentaries with claude this is how creators are producing long-form youtube documentaries in the sleep niche for about a low cost. you'll spend most of your effort building the workflow once, then every script after that runs through the same pipeline for cents. the format that works in this niche is different from normal youtube. your viewers are actively trying to fall asleep. that's the entire point. so people leave for two reasons: they got bored, or it worked and they're out. the ones who fall asleep come back later and keep listening. that repeat listening is a huge part of why the niche prints. which means the script is 90% of the whole thing. average view duration on my channels sits close to 25 min. that number does not come from cinematic visuals or fancy editing. it comes from narrative structure. if the script gets repetitive, drifts off topic, or loses momentum halfway, people stop listening. better footage cannot rescue a weak story here. the problem is the format does not scale on its own. one video needs a 15k-20k word script, hours of narration, hundreds of visual changes, music, and final assembly. writing that manually takes forever. editing every scene takes even longer. here's the workflow i set up: claude api (NOT the chat app. HIGHLY RECOMMENDED to not skip this. in the chat interface you end up typing "continue.. write chapter 4.. don't repeat yourself.. you forgot what happened in chapter 2" and by the halfway point it's contradicting earlier sections and drifting from the outline. you spend more time babysitting than writing. the api sends every request automatically and you pay per actual usage instead of another monthly subscription) google sheets connected to the claude api. this is the whole engine. you don't need to be a dev. the sheet does two things: first it generates the full documentary structure/outline. then it writes ONE chapter at a time instead of trying to produce the entire 20k words in a single response (which is where models fall apart). before each chapter, it passes claude three things: the outline, the instructions for that specific section, and a running summary of everything already written. that running summary is the trick. it's why chapter 8 never contradicts chapter 2. capcut ai video maker for the first edit. it generates voiceover, subtitles, and an initial visual sequence from auto-matched stock footage. the stock matching is not perfect, but it gets you a 90% first draft way faster than manually searching for hundreds of clips. note: capcut caps at 3000 words, so you split the script into sections, generate each one, export, and combine into the final video. HERE'S HOW THE PRODUCTION ACTUALLY RUNS: step 1 —> topic + title + thumbnail. do NOT skip this. ai cannot tell you which topic has demand or whether a title creates curiosity. this is where most of the value still is. figure this out before you touch any automation. step 2 —> run the sheet. it builds the outline first, then writes chapter by chapter, feeding itself the running summary each time so it stays consistent. cost for a full script usually lands around $0.30-0.40 depending on the model, input length, and number of revisions. step 3 —> paste script into capcut in sub-3000 word chunks. generate voiceover + subtitles + auto-matched visuals for each. export each section. step 4 —> combine sections into the final 2-3 hour video. then you handle the parts ai can't: pacing check, misleading visuals, final editorial judgment. the reason this matters is repeatability. every script moves through the exact same production structure, but you can still change the topic, tone, evidence, pacing, and narrative direction each time. so it stops being random one-off videos and starts being a system. the math: capcut is ~$20/mo and allows many exports. claude api is a little above thirty cents per script. at 30-40 documentaries a month that works out to roughly $1 in direct software cost per finished video. that figure does NOT include your time, research, thumbnails, subscriptions, failed ideas, or the cost of building the workflow itself. it is not the full cost of the business, it's the direct software cost. one more thing worth knowing: mixing real historical/stock footage alongside ai assets is the best defense i've found against the "reused/inauthentic content" flags that destroy fully automated channels. that's from experience, not a rule youtube publishes. this is not passive income and it's not a one-click youtube machine. it's a production system that makes experimentation cheaper. ai removes the repetitive work. it does not remove the need for taste.

Sulfur

25,540 просмотров • 1 месяц назад

Scientists discover surprising link between gut-brain interactions and mental health | Eric W. Dolan, PsyPost A new study provides evidence that the connection between the brain and the stomach may be linked to mental health in a measurable way. Researchers from Aarhus University in Denmark, publishing their work in Nature Mental Health, report that a specific pattern of communication between the brain and the stomach reflects how individuals feel emotionally and psychologically. Their findings suggest that these gut-brain interactions can indicate a person’s levels of anxiety, depression, well-being, and overall quality of life. The idea that emotions are linked to physical sensations in the gut is widely reflected in language. People often talk about having “butterflies in the stomach” when nervous, or feeling “sick to the stomach” when distressed. Yet, despite these common expressions, most scientific attention in the field of brain-body interaction has focused on other organs, such as the heart and lungs. These areas have long been studied for their roles in emotion and mood. The researchers were struck by how little was known about how the stomach, in particular, interacts with the brain. While recent studies have explored the influence of gut bacteria and digestion on mental health, very little work had been done on the electrical rhythms of the stomach and how they may directly communicate with the brain’s networks involved in emotion, attention, and cognition. The team behind this new study wanted to explore whether a person’s psychological profile might be reflected in how strongly the stomach and brain are coupled during rest. Their aim was not to link a specific diagnosis like depression to a single brain region, but rather to identify patterns across a broad spectrum of mental health experiences. “Our interest grew from the long-standing discussion about the role of the body in shaping emotion, a question that has fascinated philosophers and scientists for centuries,” said study author Leah Banellis (Leah Banellis), a postdoctoral fellow in Cognitive Neuroscience at Aarhus University. “Yet, while the heart and lungs have received much attention, the stomach has been largely overlooked. This gap struck us as especially surprising, because the link between the stomach and emotional experience feels so intuitive. It is heavily reflected in everyday language, with phrases like ‘butterflies in the stomach,’ ‘sick to our stomach,’ or ‘trust your gut.'” The research was part of the Visceral Mind Project, a large-scale initiative that combines data on brain activity, bodily rhythms, and psychological assessments. The team recorded data from 243 people using a method that captures both electrical signals from the stomach (electrogastrography) and brain activity measured with functional magnetic resonance imaging (fMRI). The participants represented a wide range of mental health profiles, from those reporting high well-being to others showing signs of distress, including anxiety, depression, fatigue, and insomnia. To capture this diversity, the researchers didn’t exclude people with psychiatric symptoms or diagnoses. Instead, they aimed for variation, which would allow their models to detect patterns across the mental health spectrum. Each participant underwent a series of recordings while lying still in the MRI scanner. At the same time, sensors on the abdomen captured the stomach’s slow electrical rhythm, which cycles about three times per minute. This rhythm, which originates from specialized cells in the stomach lining, is typically involved in coordinating digestion. But the researchers suspected it might also be linked to mental state. To analyze the relationship between stomach and brain activity, the team used a method that looks at how well the two rhythms align over time. This measure, known as phase-locking value, essentially captures the degree of synchronization between stomach signals and brain signals across different regions. The researchers then combined this data with results from a comprehensive mental health questionnaire. The battery included 37 different scores across a range of domains—such as anxiety, stress, mood, fatigue, attention, sleep quality, and life satisfaction. Using a statistical method known as canonical correlation analysis, they looked for patterns that linked brain-stomach coupling with the participants’ mental health profiles. The analysis revealed a clear and statistically significant pattern. Stronger coupling between the stomach’s rhythm and brain activity was associated with poorer mental health. Individuals who reported more symptoms of anxiety, depression, stress, and fatigue tended to show increased synchronization between their stomach and brain rhythms. In contrast, those with higher levels of well-being and life satisfaction showed weaker coupling. “For the first time, we’ve found a scientific link between your ‘gut feelings’ and your mental health, showing a surprising connection between your stomach’s natural rhythm and your brain,” Banellis told PsyPost. “Specifically, our study revealed that stronger communication between the stomach and brain is linked to worse mental health, such as higher symptoms of anxiety, depression, stress, and fatigue, whereas weaker stomach-brain communication aligns with better mental health reflected in higher overall well-being and quality of life.” This stomach-brain signature was not random. It was localized in specific brain networks, particularly those involved in attention, cognitive control, and salience detection. Some of the strongest associations were found in regions like the superior angular gyrus and the posterior frontal and parietal areas—regions often implicated in cognitive tasks and mental health disorders. Importantly, the researchers ran multiple control analyses to ensure the robustness of their findings. They ruled out the possibility that the observed effects were simply due to general brain activity patterns, fluctuations in heart rate or breathing, or basic features of stomach physiology. In other words, the association appeared specific to the coupling between the stomach’s electrical rhythm and particular brain networks—not just a general marker of body or brain state. Their approach was designed to detect broad psychological dimensions rather than focus on one diagnosis. The strongest psychological pattern they found was a spectrum ranging from negative affective states (like anxiety and depression) to positive traits (like well-being and quality of life). This result suggests that the stomach-brain connection is not tied to any one disorder but instead reflects a general mode of psychological functioning. “Anxiety, depression, stress, and fatigue showed the strongest links to stomach-brain communication,” Banellis explained. “While phrases like ‘butterflies in the stomach’ or feeling ‘sick to your stomach’ are common ways we describe emotional distress, it was surprising to find such consistent and clear evidence across these symptoms. Even more unexpected was the direction of the effect: we might have assumed that stronger alignment between the body and brain would be beneficial. Instead, our findings suggest that heightened stomach-brain communication could act more like a warning signal, an internal alarm system reflecting mental strain rather than harmony.” Read more:

Owen Gregorian

92,301 просмотров • 11 месяцев назад

BREAKING: Our world is an information theory-based simulation. Our physical worlds are rendered locally by conscious observer nodes. MIT/Stanford trained Rizwan Virk provides a master compendium of evidence for the simulation hypothesis in our latest documentary (full vid in reply!): 1. The building blocks of reality and biological life look like computer code: interoperable matter-antimatter particles (like electrons and positrons), human base pairs etc. 2. Heisenberg’s uncertainty principle: position and momentum measurement accuracy trade off against each other. When one is precise, the other is fuzzy: this looks exactly like computational caching. Only a certain amount of information is stored in local memory! 3. Biological and cosmic life involves fundamentally conserved, geometric building blocks: Fibonacci sequence leading to golden ratios are prime examples. They are everywhere. These look like copied and pasted code libraries. 4. All of the physical constants (Planck’s constant, big G Gravity etc) seem like they are fine tuned for life. That leaves intelligent creation as the BASE CASE and MORE likely than the idea that we just happen to be in a Goldilocks zone across millions of random permutations. The magnetosphere of the Earth also plays a role in “programming” biology and only letting in enough UV radiation to support genetic mutations and differential selection 5. Knowledge has client-server relationship. When we discover something new, we upload those discoveries to a central server. This makes new breakthroughs that much easier (this explains the bannister effect with the 4 minute mile and Rupert Sheldrakes (@RupertSheldrake) morphic field findings!). Download times are faster than upload times. Doing something incrementally new is always harder than repeating. 6. Given that only certain information is rendered locally, we end up with air pockets of “consensus reality”: this explains mass false memories as shown by the hugely popular Mandela effect 7. The idea of the mind causing wave function collapse was considered by virtually every early quantum thinker, including the pioneers who developed the core equations we use today (Schrödinger, Von Neumann etc.). Modern smug scientists who say that the wave function collapses due to particles colliding independent of an observer or due to any measurement device are engaging in AS MUCH OR MORE speculation as those who think the mind is responsible. 8. Random event generators point to humans being able to affect conventionally thought of as random “events” (with known statistical distributions) in quantum mechanics (I.e. isotope decay) with their minds. This ALSO points to the human rendering reality in real time 9. Donald Hoffman (Donald Hoffman) shows why from an evolutionary perspective, it is not adaptive for humans to see base reality: we basically render objects into “computer icons”. Again, simulation is more likely. 10. Fermat's theorem points towards light using algorithmic optimization principles that lead to most efficient paths, There are two version of the simulation: Non player character and role playing game. The first is a product of postmodern nihilism -- we can do anything because this is all a video game and we're all bots. The second is life AFFIRMING and it comports with all the world's major religions and Plato. Our primordial souls are being "ported" into a low-level incarnation in which we must learn karmic lessons. Riz and I get into all of the protocols for glimpsing beyond our simulation: it is a narrow, but worthwhile path. In a world obsessing over low level simulations (AGI), we should think more about the computational soup we are all swimming in.

Jesse Michels

298,559 просмотров • 1 год назад

Mystery 'Solved'? Radiation Spikes in NY & NJ Metro Detected Amid Drone Sightings There have been significant radiation spikes detected in the New York and New Jersey Metro areas amid a spate of ‘mysterious’ drone sightings. One theory that has been floated by a subject matter expert named John Ferguson, the CEO of a military drone company named Saxon Aerospace, is that these drones are flying at night because they are searching for gas leaks or radioactive material. Drones typically don't fly at night unless they are equipped with special sensors; these may include infrared or thermal imaging, as well as gas and radioactivity sensing technology. Two locations in the NY/NJ metro area have been flagged as condition "red," with CPM (counts per minute) levels exceeding the safety threshold of 200 CPM over the past week, according to the Geiger Counter World Map. One 1000 CPM reading was detected at Hamilton Park in Weehawken, New Jersey, across the Lincoln Tunnel going into New York City, while the other was detected near Fort Hamilton, which is located near the Verrazano Narrows bridge. These readings are markedly higher than typical background radiation levels, raising concerns about localized environmental or industrial factors that may be contributing to these spikes. The CPM metric measures the number of radioactive particles detected in a given area per minute, using Geiger counters. This does not directly quantify the strength of the radiation but provides an indication of its presence. For reference, a typical background radiation level is around 30–50 CPM, while readings above 100 CPM may indicate an anomaly requiring attention. Prolonged exposure to higher levels of radiation, such as those recorded in Weehawken and Fort Hamilton, can pose health risks depending on duration and proximity. Readings of over 1,000 CPM (counts per minute) on a Geiger counter are relatively rare under normal circumstances. Authorities are urged to investigate the source of these elevated readings and assess potential risks to the population in these areas. However, there is another factor that one should consider: An increased in reported or perceived “drone activity” that may correspond to increased military presence or emergency response and accompanying spikes in radiation levels. There are military drones and unmanned flight systems that use radioactive fuel sources, although such designs are typically considered to be “experimental” due to safety, regulatory, and operational constraints. Radioisotope Thermoelectric Generators or RTGs use radioactive isotopes, such as plutonium-238, to generate electricity through thermoelectric conversion. These have been more commonly used in space missions (e.g., NASA's Voyager probes) but have been considered for unmanned aerial systems (UAS) in specific scenarios where long endurance and reliability are critical. RTGs typically use isotopes like plutonium-238, which primarily emit alpha particles. Alpha radiation cannot travel far and is usually blocked by the outer casing of the RTG. It is unlikely to directly contribute to a 1,000+ CPM spike unless the RTG casing is compromised, exposing the radioactive material. Some radioactive decay chains produce beta particles and gamma rays, which can penetrate the RTG casing to some extent. Gamma radiation, in particular, can travel significant distances and could lead to elevated Geiger counter readings. However, radiation escaping from a well-functioning RTG is minimal and typically not detectable at a significant distance. However, if the RTG casing is damaged or degraded, radioactive material could leak, leading to higher-than-normal CPM readings. More plausible or at least common forms of radiation capable of causing Geiger counter spikes would be leakage from medical isotopes (e.g., cesium-137 or cobalt-60), industrial equipment accidents, or environmental contamination. However, increased military activity into the areas may introduce variables such as top secret equipment that may be associated with radioactivity spikes. Agencies like DARPA have explored advanced power systems for drones, including nuclear-based options, but details are often classified. DARPA’s SIGMA program also equips the Port Authority of New York and New Jersey with an advanced radiation detection system, enhancing counterterrorism efforts and radiological threat monitoring across the region's critical infrastructure. SIGMA, operational since late 2019, uses networked sensors—stationary, vehicle-mounted, and wearable—to provide real-time radiation detection and alert capabilities. The mysterious drone sightings in New Jersey took an unexpected turn with an ABC 7 News crew capturing video footage of a strange, white orb floating in the Mendham Township sky. The reporter described the phenomenon as unidentifiable and urged residents to submit similar videos for expert analysis. The unexplained activity has also included sightings over military bases, intensifying demands for clarity from public officials as speculation swirls about the drones' origins, including theories that they may have been sent by foreign adversaries. While National Security Council spokesman John Kirby dismissed over 3,000 reports as mistaken observations of helicopters or airplanes, his statements contradict confirmations from military officials at New Jersey's Picatinny Arsenal and Naval Weapons Station Earle of unauthorized drone breaches. These sightings add to growing public concern over an apparent "drone invasion" across New Jersey and neighboring states. The FBI and DHS said in a joint statement on Thursday that there was "no evidence at this time that the reported drone sightings pose a national security or public safety threat or have a foreign nexus." On Sunday, Hochul announced federal officials are deploying a high-tech drone detection system to New York State. The advanced system will support state and local law enforcement in investigating the drones, which have been flickering across the night skies over the past month, Gov. Kathy Hochul announced Sunday. Despite the federal support, Hochul urged Congress to provide greater resources. The Aerospace CEO's explanation that the 'mystery' drones may be searching for gas leaks or radioactivity fits with both the government secrecy surrounding the reported phenomena, as well as the increased radioactivity detected in the region.

Kyle Becker

359,828 просмотров • 1 год назад

IBM’s Selectric typewriter (introduced in 1961) fundamentally reshaped office culture by combining a radical mechanical redesign with strong industrial design, higher productivity, and a bridge toward modern word processing and computer keyboards. Core technical shift that changed daily work Traditional typewriters used a basket of individual type bars that swung up to strike the ribbon and paper. Fast typing frequently caused bars to collide and jam. The Selectric replaced that entire system with a single interchangeable spherical “golf-ball” typing element that rotated and tilted to the correct character, then struck the page. Because the element moved instead of the carriage, there was no heavy carriage return, less vibration, fewer jams, more consistent impression quality, and noticeably higher sustainable speed. Typists could work longer with less physical strain. Documents looked cleaner and more professional. The machine also allowed rapid font changes simply by swapping the type element (different typefaces, italics, scientific symbols, foreign-language characters). This flexibility was previously cumbersome or impossible on standard machines. Aesthetic and status transformation of the office Eliot Noyes’s industrial design gave the Selectric a smooth, modern, almost sculptural form that broke from the black/gray utilitarian look of earlier office machines. It came in multiple colors (companies could even order custom finishes), and its shape was meant to be viewed from all sides—suited to the emerging open-plan office rather than machines tucked against walls or built into desks. The Selectric became a visible status object. Having one on a desk signaled modernity and organizational seriousness. Its presence in popular culture (notably Mad Men) reinforced the image of the polished mid-century office. Office equipment was no longer purely functional hardware; it could carry design prestige similar to furniture or automobiles. Impact on people and workflow Typing pools (predominantly staffed by women) became more productive and less mechanically frustrating. The machine reduced the physical and mental friction of producing clean correspondence, reports, and forms. Electric power already lowered the force needed per keystroke; the Selectric’s consistency and reliability amplified that advantage. Over time this contributed to a cultural shift: typing moved from a specialized trade toward a more general office skill. Later Selectric variants (especially the Magnetic Tape Selectric Typewriter / MT/ST of 1964 and subsequent correcting models) introduced early forms of stored text, automatic playback, and error correction without full retyping. These were direct precursors to dedicated word processors and, ultimately, the computer keyboard as the primary human-computer interface. Market dominance and lasting cultural footprint IBM sold more than 13 million Selectrics. For roughly a quarter-century it was the dominant electric typewriter in American (and many international) offices—so ubiquitous that for many people “typewriter” simply meant Selectric. Its mechanical principles and keyboard layout also influenced later IBM terminal and PC keyboards. By the mid-1980s personal computers and daisy-wheel/laser printers displaced it, but the Selectric had already altered expectations: offices expected clean, fast, flexible document production; equipment could be designed rather than merely engineered; and the path from mechanical typing to electronic text manipulation had been opened. The Selectric did not invent the office or the typewriter, but it removed major friction points, elevated the visual and status environment of clerical work, boosted throughput in typing-intensive departments, and served as a practical stepping-stone toward word processing and keyboard-driven computing. That combination permanently changed how offices looked, felt, and operated.

Brian Roemmele

10,304 просмотров • 1 месяц назад

Here is how I am using AI right now at work, at home, and for my finances... Every founder, executive or investor I talk with these days wants to know how others are using AI in their daily lives. I figured it would be helpful to pull back the curtain on what I am using and how I have implemented the various products. There are three areas where I have adopted AI in a material way: professionally, financially, and personally. Professional use of AI On the professional side, I am currently using Grok Grok Bot extensively. I started with a Chief of Staff bot that I put in charge of the entire operation, followed by a number of more specialized bots for various bodies of work (talent recruiter, product designer, podcast researcher, book launch manager, email organizer, etc). Once I had the initial team of bots set up, I spent about an hour “onboarding” the Chief of Staff to my professional life. I treated this exactly how I would onboard a human Chief of Staff. I explained each business I am involved with, including their products, business model, personnel, metrics, and goals. I explicitly called out what the business is doing well and where we need to improve. I also gave the Chief of Staff access to relevant systems (email, calendar, Slack, analytics dashboards, etc). Once I had given as much context as I thought necessary, I asked the Chief of Staff to create an overview document to send me so I could double-check the accuracy and thoroughness of the bots understanding. I also asked the CoS bot to interview me for any other information that would be relevant to ensuring the bot could help me. This entire process was fairly quick and painless, but I believe it was the single most important thing I did to get value from Grok Bot. The more context that the AI system has, the more helpful it can be. That context can come from static, institutional knowledge or it can come from dynamic daily updates like email and Slack messages. After getting the bots set up and giving them context, I have done two other things that I think are worth sharing. The first is that my team of bots holds a daily standup meeting where they all come together and share what they did yesterday, what they are going to do today, and what they need my help or approval on (aka what they are blocked on). These “exec meeting” or daily standup allows for the bots to collaborate in a more seamless way, while also creating a very simple process for the Chief of Staff bot to put together a daily brief for me on what happened yesterday, what is going to happen today, and where I am needed to unblock productivity. The second thing I have done is treat the AI system as the brain of the company. Most people try to use AI as an augmentation to themselves, which can be helpful to a degree. I have flipped the relationship though. I look at my job as persistently giving the AI bots as much context as possible, so I can leverage their superhuman intelligence to make decisions and achieve our goals. For example, the recruiter bot recently surfaced a number of very high-quality candidates for an open role we have. After meeting with each candidate, I wrote a quick message to the recruiter bot to tell it what I liked about the person, what I thought were potential issues, improvements for future searches, and what the next steps were with each individual. All of that information and context is getting stored in the bot’s memory, which will compound over time and help us improve as an organization. Quick pro tip: If you are worried about putting all of the context into a single system’s memory, but unsure if that is the system you will use forever, you can have Grok Bot or another system dump their memory and context into a Notion document as well. This way you have a duplicate copy of the memory so it can be referenced by any AI system you use in the future. My takeaway from using Grok Bot to manage our companies is that we are having to hire less people, we are seeing a direct impact on revenue growth, and it appears to drive higher quality in our decision-making process. That is a win-win-win. I highly recommend going through these steps to setup your system correctly and it will pay off big time later on. Financial use of AI On the financial side, it was nearly impossible to find a good AI product to use for personal finance. Everything seemed to be a Chat-GPT wrapper that technically worked from an engineering standpoint, but didn’t solve any of the user problems I was facing. A big issue is that most of the fintech products are focused on budgeting and saving, rather than investing and growing your portfolio. This is why I eventually spent the time and money to build CFO Silvia. I went through a similar process of getting Silvia set up with the necessary context. I attached my bank accounts, brokerage accounts, crypto accounts, and credit cards, along with uploading real estate, cars, collectibles, and private investments. Silvia allows me to dynamically track the value of these assets (and my overall net worth) in real-time. But the real unlock for me has been talking to Silvia about two specific topics: tax and estate planning. As most of you know, I am not a frequent trader, so although you could use Silvia for stock analysis or trading activities, that is not my approach to investing. Instead, I have had great success in using Silvia to find creative and valuable tax mitigation strategies that are personalized to my situation, including ideas that had not previously been surfaced by my accountants, lawyers, or tax experts. Additionally, I have used Silvia for estate planning purposes. I am married and have four children, so there is a decent amount of complexity and opportunities to pursue. Having a dedicated resource with superhuman intelligence and the full context of my personal financial situation has been incredibly powerful. One funny thing I have noticed is that I am willing to tell Silvia certain things that I would hesitate to tell other humans (financial goals, areas of concern, etc) and I ask numerous “dumb” questions that I would probably shy away from asking a human. Regardless of why I feel more comfortable talking to the AI product, it has unlocked a few different ideas and strategies that I was previously unaware of, so that has been an added bonus to using the product. If you aren’t using AI to help manage your finances, I think it is a no brainer to start using the technology. I am biased towards Silvia since we built it, but you can give it a try for free here: Personal use of AI On the personal side, I use almost all of the traditional AI products (Chat-GPT, Claude, Grok, Gemini, Perplexity, etc). Those are well understood at this point, but one product that I started using recently that I am impressed with is Instinct AI. They have built a personal assistant AI bot that you communicate with through iMessage or SMS. The experience has been delightful, but I am most excited about the bot’s ability to anticipate the second or third-step in a process before I have to tell it anything. For example, Instinct got access to my calendar and immediately started identifying scheduling conflicts and asked me if I would like the bot to reach out to one of the parties to reschedule. I never told it to look for conflicts, nor did I tell it I wanted help rescheduling things. It’s “instincts” knew what the basic task would be and began executing. Another example is that Instinct was told my wife is Polina, so whenever it deems something important to the household or family, Instinct will add Polina to the calendar invite, communicate the information to her, or ask me if Polina should be aware of the information. This is very helpful for someone like me who has too many things floating around in my brain and should always do a better job of keeping Polina informed about various things. Lastly, Instinct is very helpful in scanning my personal email and understanding what is most important. It ignores things that are trivial, but somehow can parse out the high priority items, summarize them for me in a text message, draft a response to the email, and then ask me for permission to respond. As I said, it is the most impressive personal assistant AI product I have used so far. So those are the three big areas that I use AI today and the specific products I have incorporated into my life. Before I let you go, I figured I could share some best practices I have learned as well. I also make sure to tell AI bots they are not allowed to respond to any message or email without my explicit approval. This reduces the risk of having a bot go rogue with a message or commitment that I am not onboard with. I also ensure that each bot only has read access to our business systems like an analytics dashboard, etc. While I am a big proponent of using these products and believe they will fundamentally transform how we operate professionally, I am still not ready to let them loose without human oversight. I am sure that will change in the coming weeks and months, but I need more time to get comfortable with that level of delegation and trust. I hope this overview was helpful for each of you. It would be great if you could respond to this post with any products you are using or tips/tricks that you have learned to get more productivity and value in your life. I love writing these letters each day because I learn just as much from me as I learn from you all. Onwards!

Anthony Pompliano 🌪

83,069 просмотров • 15 дней назад

🚨What is she carrying? Part 2⁉️ Depending on your AI platform preference … we get either a $40,000 handheld X-ray device or a $40 thermos-and-lunch-bag cooler combo? What was your conclusion, and which was right? When we first came across this video months ago, I immediately said it looked like she was “carrying a lunch bag,” or some kind of cooler. But for whatever reason, and what we were more focused on at the time, we didn’t spend the hours and hours and hours required to drill down on those few seconds of video. Not until this week. Tons of social media critics say I should “just release everything we know, and let the truth fall where it may.” But that’s how we get in trouble. And we HAVE gotten things wrong in this five-year-long investigation. EVERYONE has made mistakes. Left and right media, major legacy media, alternative media, and even the best of the independent journalists have made mistakes or misreported details of the January 6, 2021 event. Whether on purpose, by accident, or careless disregard of the truth … you can be the judge of each incident. I’ve explained on numerous occasions that we’ve spent more than a year researching, investigating, and preparing some stories before going public. In this case, when we finally started looking hard at it, the Brave New World of AI took us on a wild goose chase. We now have good reason to finally drill down on the timelines and available video leading up to the sequence of events on the night of January 5, 2021 … the night before the discovery of the two “devices” at the RNC and DNC headquarters. When inputting into AI that first video — which I posted last night — It began spitting out some shocking alternatives to my original “lunch bag” assumption. Unprompted, the AI drew its own conclusion about what Ms. Kerkhoff was carrying, probably because they were “cops” in the video. To be clear, UNPROMPTED, AI was initially adamant that the item in her hand was a portable X-ray device for sniffing out narcotics, weapons, and explosives. Even some The U.S. Capitol Police sources we shared the video with said the same thing. Other Capitol Police sources were more skeptical. Then, suddenly AI wasn’t so sure. As I mentioned last night, in my second post, X’s Grok and Google’s Gemini were arguing with each other. Gemini was shockingly specific: “The item being carried by the officer in that CCTV footage matches the exact structural form factor, top-handle position, and multi-point shoulder strap design of a Viken Detection HBI-120 Handheld Backscatter X-Ray Imager … Rather than being transported bare, it is housed in its specialized tactical, weather-resistant soft deployment carry case.” Wow. You can imagine how that hit like a bolt of lightning, considering whom we were looking at. Grok was less committed, offering varying possibilities ranging from a “paper takeout bag” to all manner of sophisticated law enforcement-used electronic devices. When I pitted Grok against Gemini — and vice versa — Gemini dug in its heels, saying: “Based on the visual evidence in the video and the analysis circulating in response to Steve Baker's query, the device carried by Officer Kerkhoff is almost certainly a Viken Detection (formerly Heuresis) HBI-120 in its tactical soft-carry case. While officially unconfirmed, the visual and structural match is highly specific.” Grok countered: “The competitor [Gemini] write-up describes that pouch as if it were a catalog item with a unique layout, then maps the CCTV blob onto it and calls it an ‘exact’ match. The citations in that text are HBI-120 product pages, not photos of a factory soft case that looks like this bag. That is overfit. Size class + top handle + ‘looks tactical’ is not an identification of an 8-pound 120 keV imager. Could an agency have put an HBI-120 in some aftermarket or unit-made padded bag? Yes. Does this video prove that? No.” But Grok was also an absolute mess in describing the basic video scene itself: “A gray-haired person in dark clothing walks in from the left carrying a dark bag by a single top handle. Papers or a light object in the other hand. Two uniformed officers come down the ramp. They meet. Later two officers walk the sidewalk; one has a light-colored item (bag or papers), and a dark bag is still in the scene. Someone else walks away across the lot with a bag.” Huh? Anyway… Gemini didn’t like what it called “The Thermos Theory”: “Soft Lunch Coolers typically use flexible nylon webbing straps or soft padded handles that pinch or deform when lifted. The handle in the video does not appear to ‘pinch’ like a soft strap; it remains an open loop. This structural rigidity strongly favors the Viken device (or a hard-shelled case) over a soft lunch bag.” Grok’s conclusion: “On evidence quality, the Viken ID is the weaker of the two. It is possible in the abstract and unsupported as a positive ID here. The honest read is: dark soft top-handle bag, lunch-cooler to small-kit size, contents unknown. Anything more specific needs a better frame, a still of the bag on the ground, or provenance outside this clip.” Gemini’s conclusion: “This single frame is ambiguous due to the blur and lighting (reflection). However, the rigid, non-deforming handle visible in the officer's grip is a mechanical characteristic of the Viken HBI-120 (or similar heavy equipment) and is inconsistent with the flexible straps of a standard soft lunch cooler. While the "silver" glint is a point for the thermos theory, it is plausibly explained by the Viken's glass screen or case window. Verdict: The structural analysis (rigid handle) still favors the heavier equipment (Viken) over a standard soft lunch bag.” So, depending on your AI preference … we get either a $40,000 handheld X-ray device or a $40 thermos-and-lunch-bag cooler combo? But there is more evidence in the alternate video angles available. (That I now present below.) They seem to confirm with the ‘human eye’ that Ms. Kerkhoff is carrying a metal cylinder (possibly a thermos) in the same hand as a soft-sided lunch bag cooler. Though the design conflicts with the common thermos handle configurations of most known cylinder thermoses. (Photos below.) And we've not been able to get an accurate hit on the so-called "lunch bag" with that specific handle and strap configuration. And THAT is why you don’t just “release what you know” without seeking every possible video angle and expert opinion. That is why we didn’t run to print with our original November 8 story on the OG topic without first taking it to a government intelligence agency and professional investigators for review. That is why so many bad theories about January 6 still abound — five and a half years later — because they were based on a single camera angle, when years later, the same scene was revealed to have been captured from multiple angles that change reality 180 degrees. This is exactly why all CCTV footage — not just from January 6, but also January 5 and 7 — still needs to be released to the public. When Speaker Mike Johnson authorized Rep. Barry Loudermilk's old investigative subcommittee to begin uploading CCTV footage to a Congressional Rumble channel, we were elated. I had already spent many weeks in the Capitol CCTV viewing room in D.C. The travel, the expense, and the scheduling hassles with the committee made it nearly impossible to spend the amount of time required to prepare any story correctly. Not only to view and harvest what you were looking for, but also to sift through far more than the infamous “41,000 hours” of footage. Congress made more than 1,800 cameras' worth of footage available, and ten total days of footage. That’s hundreds of thousands of hours of potentially useful footage to review. An impossible task for any one person or media organization to review if Congress didn’t make that footage directly available to the public. But they didn’t finish the project. Tens of thousands of vitally important hours from both January 5 and 6 were never uploaded to the Rumble page. Additionally, my team has made specific requests for curiously missing gaps in footage throughout that two-day timeline. In an arrangement made with the Committee, they had originally been very good about getting us the footage from the specific cameras and timestamps we requested. That suddenly stopped when the new Congress and Loudermilk’s new J6 investigative subcommittee took over in January of 2025. Joe Hanneman and I have made innumerable requests — REPEATEDLY — for missing and/or unreleased cameras and timestamps specifically related to the pipe bomb investigation. Despite being told — REPEATEDLY — that they would provide the requested footage, they never did. Something happened. As I’ve reported several times in the last few months, the Capitol Police were finally and successfully able to shut down Loudermilk’s subcommittee investigation into ALL THINGS related to the Capitol Police. They did this only with the complicity and surrender of Speaker Johnson and Judiciary Chairman Rep. Jim Jordan to Capitol Police leadership’s demands. On that note, and in conclusion … there are eight full hours of missing footage from January 5, right in the middle of the day. ALL CAMERAS are missing. These are important hours for what we are tracking. We can see Ms. Kerkhoff arrive at Capitol Police HQ early in the morning to clock in for her shift, but she is not carrying her “lunch bag and thermos” when she arrives. Her car is parked two blocks away, and is in the same parking spot at the end of her day. We can see her leave HQ late in the day (as I’ve documented in the last several posts on this page) with other officers and go to the Fairchild Building. Only to return some half hour later carrying that … thing(?) … and only to spend 45 seconds in the HQ to “clock out” from her overtime shift. We’re still missing vital video footage that both Speakers McCarthy and Johnson promised the American people. Including certain cameras deliberately withheld at the RNC bomb drop location, and other cameras with mysterious gaps at the most important of moments. Do the other video angles here prove that either Grok or Gemini was right, or does the X hive mind have better theories on what Kerkhoff is carrying? How about that high-definition CCTV camera that is right inside that west side door at Capitol Police HQ, with good lighting? They should release that video to us. A $40,000 bomb detection device or a Walmart thermos and lunch bag? I’m good with either. The truth is what we seek. But we should be able to see ALL the footage. Including all Capitol CCTV cameras and footage from January 6, and the days immediately preceding and following. Including the 39,000 video files the FBI claims to have in the entire J5/J6 pipe bomb investigation. Conspiracy theories are born and fester precisely because the government isn’t transparent and purposefully keeps the People in the dark. Then the lawyers who control government make billions from the legal aftermath. (More Photos in the thread below.)

Steve Baker

48,351 просмотров • 17 дней назад

What's next for OpenTUI? Here's a technical write-up. Over the last few months OpenTUI gained a lot of stability improvements, new unnecessary but fun features like live audio streaming, and useful features like rendering to the scrollback buffer mixed with a live TUI, called footer mode. Overall the feature set enables building large and complex applications. React and Solid make it super simple and convenient. There is still so much to do though. Three big milestones we have set out to achieve are: - Moving most of the behavioural logic currently living in TypeScript down to the native Zig core - Node compatibility - Optimizing the hell out of primitives like text rendering The render tree mechanisms are currently only usable from TypeScript. Think of the DOM, but controllable like a scene graph. Elements in the render tree are called renderables. They can expose a render method to draw themselves. All renderables are derived from a BaseRenderable. Renderables and the render tree will become native primitives. Building blocks usable from any language bindings. Reducing the TypeScript bindings to a very thin layer, with all the behavioural logic living in the native binary. Moving this down is not just a matter of porting TypeScript classes to Zig. TypeScript currently owns the tree, dirty-state propagation, layout reads, culling, and render ordering. If it still has to walk every node and call into native code for each step, we keep most of the complexity and add FFI overhead. Whole passes and their state need to move together. We took a big step towards this recently by building yoga-layout into the native binary. It exposes part of the official yoga-layout TypeScript package via FFI. Only the API surface that is actually used by OpenTUI. Covered by the test suite of the original yoga-layout package. This already gave a median speedup of ~2.5x, and up to 30x for narrow scenarios. The yoga-layout integration is useful beyond the speedup. Built-in text and editor measurement can now happen entirely in native code during layout instead of calling back into JavaScript. I ran an experiment last month taking this even further, having GPT 5.6 port yoga-layout from C++ to Zig, which gave extremely good results. It would be a burden to maintain right now though, so that's off the table for now. I might come back to it. Simon Klee is working relentlessly on Node compatibility and already has a full Node version of OpenCode running. Node got FFI support in v26.4.0, thanks to help from the Node community, namely Matteo Collina and Paolo Insogna. Behaviour and interfaces seem similar between Node and Bun, but there are some major differences. To get the best performance out of the Node FFI implementation, its usage has to follow some rules. Node has three ways to call native functions: the generic C++/libffi path, the SharedBuffer path, and the V8 Fast API. The generic path converts every argument in Node's C++ layer and then calls the function through libffi. It is flexible, but also the slowest option for frequently called functions. The SharedBuffer path is a middle ground. JavaScript writes scalar values and BigInt pointers into a small per-function buffer, reducing some conversion work. The actual native call still goes through libffi though. Typed arrays used as pointers cannot be packed into this buffer and fall back to the generic path. The path we really want is the V8 Fast API. Node generates a small machine-code trampoline for the exact function signature, allowing optimized JavaScript to call the native function without going through the generic converter or libffi. This only applies to JavaScript-to-native calls. Callbacks from native code into JavaScript still use libffi closures. Getting onto this path is quite strict. A signature can have at most eight arguments and everything must fit into CPU registers. x86-64 Unix systems have room for six GP (general-purpose) and eight FP (floating-point) arguments. AArch64 has room for seven GP and eight FP arguments. Anything that spills onto the stack falls back to a slower path. These are Node fast-path restrictions, not general FFI restrictions. Bun also does not support passing structs by value through its current FFI API. OpenTUI uses bun-ffi-structs to pack ABI-aligned struct data into an ArrayBuffer and passes a pointer instead. Despite the name, the package also works with Node. Pointers need some care too. Typed arrays and ArrayBuffers normally have to be resolved into BigInt addresses first. Eligible functions with exactly one pointer argument get another Fast API entrypoint that can extract the address directly from the buffer. An eligible signature is still not enough. V8 has to optimize a direct call with a fixed number of consistently typed arguments. Wrappers that collect arguments and forward them using spread or Reflect.apply can hide that call shape and keep the function on a slower path. The practical rules are: keep hot signatures within register limits, use direct fixed-arity calls with stable argument types, reuse owned buffers safely, and batch small operations. Then measure the real call site, because eligibility only makes a function fast-capable. We have to design the ABI around these constraints where it makes sense and gives the expected performance improvement. The third big area is text rendering. Today a Text renderable accepts a string, StyledText, or a tree of TextNodes. Before rendering, the TextNode tree is walked and flattened into styled chunks. Those chunks are packed in TypeScript, sent through FFI, copied into a native TextBuffer, and stored in a rope. Styles are represented separately as highlights. A TextBufferView then wraps the rope into visual lines, which are drawn into the visible buffer. This works, but updates are much more expensive than they should be. setStyledText effectively throws away and rebuilds the rope, copies and reparses all text and recreates the style highlights. Changing one TextNode also walks and flattens the complete tree before going through this path again. Text and style segments should instead live directly in the rope and support incremental replacement. Memory ownership is split between retained JavaScript buffers, the native memory registry, rope arenas, wrapping caches, styled-text storage, and highlights. Different operations preserve or reset different parts of that state. This is hard to reason about and can retain memory for much longer than expected. Text storage needs clearer ownership, with fewer lifetimes split across JavaScript and native code. The public API reflects the same split. The t template literal is convenient, but creates another intermediate chunk representation that is mutable, not cached, and not merged. Text also maintains both StyledText content and a special TextNode tree, which do not compose properly. TextNode is only a style scope, not a normal layout primitive, so Text renderables cannot naturally compose inside each other. I think this should become one Text primitive backed directly by rope segments. The template literal API might disappear or become a very thin helper around those native segments. Editing has another temporary layer in TypeScript. Extmarks currently monkey-patch editing operations, scan and adjust all marks after changes, maintain their own undo state, and recreate native highlights. They should become native marks anchored directly in the rope. A proper mark tree, similar to Neovim's marktree, could update marks together with edits, undo, and redo, and provide the foundation for highlights and concealment. Text wrapping has also become too complex. Supporting CJK, emoji, combining characters, ZWJ sequences, tabs, and different terminal width rules currently mixes byte offsets, grapheme indexes, and display-cell columns across several custom algorithms. Dirty views rewrap the complete document. Measurement and drawing can repeat some of the same work. The wrapping implementation needs an overhaul, but the exact shape is still open. The goal is to make Unicode handling easier to maintain, avoid repeated full-document work, and clearly separate byte offsets, graphemes, and terminal display cells. None of this will happen as one big rewrite. We will replace pieces when we understand the problem well enough and when the result is clearly simpler, faster, or more useful. To achieve all of this we might break public interfaces. Thanks to OpenCode and a lot of good models, migration to a new version with breaking changes mostly is not an issue anymore. What do you want to see next for OpenTUI?

kmdr

29,430 просмотров • 1 месяц назад

MH370: Lithium Cargo, Semiconductor Scientists, and the Frequency-Wave Interception Theory TL;DR: Under the full Frequency Wave Theory interpretation, MH370 was not a conventional crash. The aircraft was carrying large lithium battery cargo and a group of advanced semiconductor specialists. According to this theory, a U.S. black-budget surveillance operation using wide-area ISR systems such as Gorgon Stare V2 tracked the aircraft, after which a tri-orb plasma resonance system intercepted the jet and displaced it through a phase-shift event. The aircraft was allegedly relocated to Diego Garcia. The scenario connects lithium cargo risk, strategic semiconductor competition, intelligence surveillance networks, and frequency-based field manipulation. —————————— On March 8, 2014, Malaysia Airlines Flight MH370 vanished while flying from Kuala Lumpur to Beijing with 239 people aboard. The official investigation concluded that the aircraft deviated from its flight path and likely flew south into the Indian Ocean before crashing. However, the wreckage has never been conclusively located, and the chain of events remains one of the most puzzling mysteries in aviation history. Several unusual elements surrounding the flight have fueled alternative interpretations. The cargo manifest confirmed that the aircraft carried a shipment of lithium-ion batteries, which are known to pose thermal and fire risks in aviation transport. Lithium batteries can enter runaway reactions when damaged or improperly stored, generating intense heat and electromagnetic disturbance. In a conventional investigation, this cargo is treated primarily as a possible fire hazard. In a Frequency Wave Theory framework, however, lithium battery packs represent something else: a dense electromagnetic energy reservoir capable of interacting strongly with surrounding fields. Another point of interest is the passenger list. Reports circulated that approximately twenty engineers connected to advanced semiconductor research were on board. Semiconductor technology sits at the heart of modern geopolitics, powering everything from artificial intelligence to missile guidance systems. The loss of a group of specialists connected to chip fabrication and electronics development would represent a strategic event, not merely a tragic accident. This has led some investigators and researchers to speculate that the aircraft may have been targeted or intercepted because of the individuals aboard. The intelligence dimension deepens when surveillance infrastructure is considered. Modern military monitoring systems can track enormous areas of the planet simultaneously using wide-area motion imagery (WAMI). One example is the Gorgon Stare V2 system used on MQ-9 Reaper drones, capable of observing entire cities and tracking thousands of moving objects at once. In a theoretical scenario where MH370 entered a monitored corridor, such a system could maintain continuous visual and infrared coverage of the aircraft even after it disappeared from civilian radar. According to the theory circulated by the researcher known online as RegicideAnon, the aircraft was captured simultaneously by two surveillance perspectives: a thermal imaging system consistent with a targeting pod and a wide-area optical platform. The videos associated with this claim appear to show three luminous spherical objects orbiting the aircraft in a rotating triangular pattern. In the Frequency Wave Theory model, this geometry is not arbitrary. Three rotating emitters form a stable standing-wave cavity capable of phase-locking to a target object. Frequency Wave Theory proposes that all matter exists as coherent standing waves within a universal scalar field Φ. Objects maintain stability through conserved Frequency Momentum: FM = ½ ρ ω A² If an external system can measure and phase-lock to an object’s resonance signature, it becomes possible to manipulate the object’s inertial coupling with spacetime. The three rotating orbs observed in the footage could represent nodes of a resonance field surrounding the aircraft. By synchronizing their emissions, the system could gradually reduce the plane’s inertial anchoring through Frequency Momentum transfer. As the phase alignment intensifies, the aircraft’s coherence state approaches a critical boundary where the phase differential approaches |Δφ| → π. At this threshold, a phase-inversion bubble can form — essentially a localized cavity in spacetime where conventional inertia is suppressed. When the cavity collapses, the aircraft’s wave structure transitions out of the local coordinate frame. The bright flash seen in the footage would therefore represent a rapid phase transition rather than an explosion. The aircraft does not disintegrate; it exits the observable frame. Conservation still applies: FM_in ≈ FM_out The system transfers Frequency Momentum from the aircraft into another region of the field, allowing the object to reappear elsewhere. In the version of events proposed by this theory, the destination was the U.S. military installation on Diego Garcia, located in the central Indian Ocean. Diego Garcia hosts major surveillance infrastructure, strategic bomber facilities, and advanced tracking systems. Its remote location and high security would make it a plausible site for receiving or containing a displaced aircraft. The alleged involvement of U.S. personnel has also been connected to the case of Edward Lin, a U.S. Navy flight officer later convicted in espionage-related charges involving sensitive surveillance information. Some researchers speculate that knowledge of the operation or the existence of the videos may have circulated within intelligence channels connected to that case, although no official link has ever been confirmed. Under this interpretation, the disappearance of MH370 becomes something very different from a lost aircraft. The lithium cargo becomes a potential energy-interaction variable, the semiconductor engineers represent strategic technological assets, and the surveillance network reveals that the aircraft may have been tracked continuously even after leaving civilian radar coverage. —————————— From the Frequency Wave Theory perspective, MH370 represents the first public glimpse of a hidden technological capability: the manipulation of matter through resonance control. The rotating plasma orbs act as field emitters that can phase-lock to a target, redistribute Frequency Momentum, and temporarily decouple the object from local spacetime. In that framework, the aircraft did not simply fall into the ocean. It was intercepted, resonance-captured, and relocated through advanced frequency engineering.

Drew Ponder

14,294 просмотров • 6 месяцев назад