Loading video...

Video Failed to Load

Go Home

remind-reid-tracker REMIND — RE-Identification with Memory for INDoor Navigation REMIND addresses a core challenge in visual tracking: re-identifying objects that disappear and reappear, look similar to one another, or are observed from changing viewpoints. Rather than relying on position or motion cues, REMIND builds appearance-based identity models per object...

17,808 views • 4 days ago •via X (Twitter)

0 Comments

No comments available

Comments from the original post will appear here

Related Videos

Everyone is sleeping on Meta's SAM 3 release. But it's actually a big deal. Here's why: Companies spend millions paying humans to label images and videos frame by frame. A single autonomous driving dataset? Months of work, hundreds of annotators, millions in cost. Without labeled data, you can't train custom models. Without custom models, you're stuck with generic solutions. This is why most companies never move past pilots. SAM 3 breaks this cycle. First let's look at the evolution: SAM 1 segmented objects when you clicked on them. Revolutionary, but one object at a time. SAM 2 added video tracking with memory. Game-changing, but you still manually prompted every object. SAM 3 changes everything with text prompts. Type "yellow school bus" and it finds ALL of them in your image or video. Not just one. Every instance across thousands of frames. Now here's where people get confused: "Can't I just use GPT-5 or Gemini for this?" No, and here's why that's a terrible approach. Large multimodal LLMs are great for reasoning, but they're slow and expensive for production visual tasks. You're paying API costs per image, waiting seconds for responses, getting inconsistent results. SAM 3 runs in 30 milliseconds on a single GPU for 100+ objects. That's 100x faster, and you own the infrastructure. More importantly, SAM 3 gives you precise pixel-level masks, not descriptions. Try asking an LLM to segment every defective part on a manufacturing line in real-time. It won't work. SAM 3 does this effortlessly. The real breakthrough is their data engine. Meta built an AI-human hybrid system that's 5x faster for complex annotations. They trained SAM 3 on 4 million unique visual concepts - 50x more than existing benchmarks like LVIS. SAM 3 is trained on 4 million unique visual concepts, it handles everything: - Text-based concept search - Interactive refinement with clicks - Video tracking across frames - Zero-shot detection of new concepts The model is open source. Weights, code, and benchmarks are on GitHub. If you're building computer vision applications, this is the foundation model to evaluate. The annotation time savings alone will pay for integration costs within weeks. Find the relevant links in the next tweet!

Akshay 🚀

46,421 views • 8 months ago

1. Essence of the Problem: Algorithm “Hesitation” and System “Jitter” The “clear → blurry → clear → blurry” cycle you see on the preview screen is essentially the AI algorithm dynamically switching between multiple image-processing paths. During 10× telephoto preview, Samsung’s multimodal imaging system makes decisions based on several concurrent signals: Scene Classifier (scene recognition) AI Detail Enhancer (texture-enhancement algorithm) Motion Estimation (motion detection) HDR Weight Selection (highlight suppression or shadow lift) The issue is that these modules lack a unified arbitration layer. When multiple modules give conflicting judgments about the same frame (for example, “static subject” vs. “slightly moving object”), the algorithm repeatedly enables and cancels enhancement strategies. The result is a visual oscillation of “pre-load → cancel → pre-load → cancel.” This reflects architectural uncertainty within Samsung’s image-processing framework. 2. Deeper Systemic Issue: Unstable Coordination Between ISP and AI In recent Galaxy generations, Samsung’s imaging stack consists of three main components: Exynos/Snapdragon ISP layer (hardware-level processing) Samsung Multi-Frame Engine (multi-frame fusion) Galaxy AI Pipeline (deep-learning post-processing) The core problem is that these modules do not operate within the same clock domain. The AI processing unit runs asynchronously on the NPU, while the ISP and multi-frame fusion run synchronously on the main SoC. In certain scenarios, when the AI result hasn’t returned yet, the ISP outputs the preview frame first—causing frame-to-frame style fluctuations. This isn’t a performance issue; it’s a scheduling bug in the system architecture. Apple avoids this by implementing a unified “Image Core” framework within the A17 Pro. All AI decisions, HDR merges, and white-balance calculations occur within one synchronized pipeline. As a result, the preview image already matches the final shot almost perfectly. 3. User-Level Impact: Inconsistent Output and Experience Fragmentation This “algorithm hesitation” leads to three direct consequences: Preview and final image mismatch — what users see is not what they get. Large variations between shots — even under identical conditions, different AI branches produce completely different looks. Loss of operational trust — users cannot predict results and hesitate to press the shutter. In imaging experience terms, this is actually more serious than sharpness or noise issues, because it breaks the user’s sense of stability and reliability with the device. 4. My View: Samsung’s AI Imaging Needs a “Referee System” The root cause isn’t insufficient power or hardware; it’s the absence of an orchestration layer. Samsung has too many independent sub-modules (super-resolution, noise reduction, detail enhancement, color reconstruction, depth recognition, AI HDR, etc.) but no master controller to decide when to activate them, how to prioritize, or how to manage latency. The ideal solution would be to: Establish a Central Scene Controller Manage all AI sub-modules with unified priority scheduling and decision memory Maintain temporal consistency of algorithmic states across consecutive frames Only then can Samsung truly fix its “algorithm instability” problem and move its Galaxy imaging pipeline toward maturity.

PhoneArt

28,588 views • 9 months ago

$MU $SNDK $LITE $VRT NVIDIA and Groq: 2nd and 3rd Order Strategic Infrastructure Effects and Market Implications Public reporting indicates NVIDIA has agreed to acquire Groq for approximately $20,000,000,000 in cash, while excluding Groq’s nascent cloud business from the transaction perimeter. The reported carve-out materially constrains the immediate, direct linkage from the acquisition to incremental, NVIDIA-controlled data center capacity build-out because GroqCloud appears to be the principal channel through which Groq hardware is currently monetized at scale as a service. The infrastructure-market implications therefore depend primarily on post-close product strategy: whether NVIDIA (1) commercializes Groq silicon as a distinct inference product line and drives broad deployment through OEM/ODM channels and partners, (2) uses the acquisition mainly to absorb IP and talent while de-emphasizing standalone Groq hardware volumes, or (3) uses Groq technology to reshape NVIDIA’s own inference systems and networking roadmaps. The dominant transmission mechanism into memory, networking, and facility infrastructure markets is the degree to which NVIDIA shifts incremental inference deployments away from GPU architectures that are tightly coupled to external high-bandwidth memory (HBM) and toward Groq’s current architecture, which emphasizes large on-chip SRAM, deterministic compiler-scheduled execution, and direct chip-to-chip connectivity. Independent and company-published materials describe Groq’s current-generation approach as having no external memory, keeping weights and KV cache on-chip during processing, and requiring model sharding across multiple chips due to limited on-chip SRAM per device. That architectural choice is directionally HBM-negative on a per-accelerator basis and ambiguous for DRAM, NAND, networking, power, and cooling on a per-token basis because the design can reduce memory wall losses and tail-latency overhead while potentially increasing the number of chips and interconnect endpoints required to serve large models and long-context workloads. HBM implications are the most mechanically straightforward but should be framed as second-derivative rather than absolute. If Groq-class inference silicon meaningfully displaces NVIDIA GPU-based inference deployments, incremental HBM bit demand tied to inference growth could be reduced relative to a GPU-only baseline because Groq’s current approach does not appear to attach HBM stacks to each accelerator. However, current market structure suggests HBM remains supply-constrained and is being pulled by multiple vectors including continued GPU training scale and high-capacity inference configurations, with leading suppliers signaling tight conditions extending beyond 2026. In that environment, reduced inference-driven HBM intensity could primarily reallocate scarce HBM supply toward higher-end training and premium inference GPUs rather than creating an outright volume collapse, preserving high utilization of HBM capacity while potentially affecting the slope of pricing power and capacity expansion urgency over a multi-year horizon. The key downside scenario for the HBM complex would be a durable architectural bifurcation where “good-enough” inference shifts disproportionately to HBM-less ASICs across a broad swath of deployments (latency-sensitive, batch-1, cost-per-token optimized), while training remains GPU-HBM dominated; such a split would reduce the portion of future inference compute that naturally monetizes through HBM content and could compress the incremental HBM-per-AI-dollar ratio. The key upside/neutral scenario for HBM is that the supply chain remains fully allocated regardless, with NVIDIA using any “freed” HBM to ship more high-end GPUs into training and long-context inference, especially as roadmaps increase HBM per GPU, sustaining robust aggregate bit demand even if inference becomes more heterogeneous. Conventional DRAM implications split into 2 channels: (1) DRAM wafer capacity diversion into HBM and (2) DDR content per server in AI clusters. Supplier commentary indicates that AI-driven memory demand is supporting elevated DRAM markets more broadly, and HBM production is resource-intensive versus conventional DRAM, tightening supply for DDR products in parallel. A meaningful NVIDIA pivot to an inference architecture that reduces HBM dependence could, at the margin, ease the most acute HBM-driven bottlenecks and allow memory manufacturers more flexibility in balancing DRAM mix, which could be modestly DDR-positive on the supply side (less crowding-out) even if it is DDR-neutral or slightly negative on the demand side (if per-node CPU/DDR requirements decline due to more efficient accelerator utilization). The dominant practical outcome is likely that DDR demand remains supported by broad AI server proliferation and increasing memory footprints at the system level (CPUs, networking stacks, caching layers, retrieval-augmented pipelines), while HBM remains the premium profit pool; therefore, any HBM displacement that increases total server volumes could indirectly keep DDR demand resilient even if DDR per accelerator is not rising materially. NAND flash implications are comparatively indirect and volume-driven rather than architecture-driven. Inference clusters require SSD capacity for model storage, container images, logging, and increasingly for fast local retrieval indices and embedding stores, but the storage footprint per unit of compute is typically smaller than in training pipelines that stage large datasets and checkpoints. If NVIDIA uses Groq to lower inference cost and latency enough to expand the total number of inference deployment locations (regional colocation, enterprise on-prem, sovereign footprints), aggregate SSD attach could rise through geographic fragmentation and replication of model artifacts across more sites, even if per-site storage is modest. The NAND effect is therefore likely to be demand-broadening and mix-positive (datacenter SSDs) but not a primary swing factor versus the macro AI capex cycle and consumer/device cycles. Hard disk drive (HDD) markets should see negligible direct sensitivity because nearline HDD demand is driven by bulk storage and cloud archiving economics, while inference acceleration choices primarily reshape compute and network layers; any HDD benefit would be a tertiary function of overall data center square footage expansion rather than a direct consequence of Groq silicon displacing GPUs. Optical networking implications require separating (1) intra-cluster back-end fabrics that connect accelerators and (2) front-end / data center interconnect (DCI) that connects sites and regions. Groq’s own positioning and third-party reporting suggest scaling beyond a single node or rack relies on high-bandwidth fabrics and, in some described configurations, optical interconnect scaling across hundreds of chips. If NVIDIA commercializes Groq at scale, 2 offsetting forces emerge: lower cost-per-token and improved latency could expand inference throughput and drive more east-west traffic, increasing demand for high-speed switching and optics; conversely, if Groq delivers materially higher utilization and tokens per unit of network bandwidth for certain workloads, the network required per served token could decline. Public NVIDIA materials already indicate an aggressive photonics roadmap aimed at scaling AI factories, including co-packaged optics (CPO) switches and explicit collaboration with Coherent and Lumentum in the silicon photonics supply chain. That linkage is important because it suggests that, independent of Groq, NVIDIA is already pushing optics integration deeper into the switch package to reduce power and increase resiliency; Groq increases the strategic incentive to reduce network power and latency if inference becomes even more distributed and latency-sensitive. For Lumentum and Coherent specifically, the net implication is less about “more optics versus fewer optics” and more about a shift in optics form factor and value capture. Co-packaged optics can reduce reliance on pluggable transceivers in some switch architectures while increasing demand for integrated photonic engines, lasers, fiber attach, packaging processes, and component-level supply. NVIDIA’s own announcements explicitly position Coherent and Lumentum as collaborators in creating the integrated silicon/optics process and supply chain for photonics switches. If Groq accelerates the transition to very large-scale fabrics (more endpoints, higher port speeds, tighter power envelopes), that tends to pull forward CPO adoption and amplifies demand for the underlying photonics components even if the conventional pluggable module TAM is structurally pressured over time. If Groq instead pushes inference toward smaller, more localized pods (closer to users, more regional colocation), that can be optics-positive for DCI and metro connectivity because more sites must be interconnected at high bandwidth with low latency, favoring coherent optics and high-speed interconnect between facilities. The principal risk for optics suppliers is timing and margin structure: a faster move to NVIDIA-driven integrated photonics could concentrate bargaining power and compress margins for commoditized transceiver modules while favoring suppliers with differentiated lasers, integration capability, and qualification depth in NVIDIA’s CPO ecosystem. AEC and copper interconnect implications hinge on whether Groq deployment increases the density of short-reach links inside racks and rows. High-speed copper remains structurally advantaged at very short distances on cost, power, and serviceability, but reaches become constrained as lane speeds and aggregate bandwidth rise, creating a role for active electrical cables (AECs), retimers, and signal-conditioning silicon. Credo explicitly positions its AEC products as enabling reliable lossless 800G connectivity for AI clusters, and the company has highlighted participation at NVIDIA GTC with content focused on extending PCIe/CXL using AECs, indicating relevance to next-generation system topologies that require longer reach and higher signal integrity than passive copper can deliver. If NVIDIA turns Groq into a widely deployed inference card or chassis product, the likely near-term effect is AEC-positive because (1) more inference throughput tends to increase top-of-rack connectivity requirements, (2) distributing inference across more racks and sites increases short-reach links per unit of delivered service, and (3) PCIe-attached accelerator architectures tend to require robust signal conditioning as systems move to PCIe 6.x and beyond. Groq workshop materials explicitly reference GroqCard and GroqNode form factors, reinforcing that PCIe-attached deployment has been central to Groq’s current packaging strategy. The main countervailing risk is that Groq’s deterministic chip-to-chip fabric could be implemented primarily through backplanes and direct board-level connectivity that reduces the need for merchant AECs inside the box; in that case, incremental AEC demand would concentrate more in rack-to-switch and node-to-fabric links rather than within-chassis chip fabrics. Astera Labs implications are connectivity-architecture sensitive and, on balance, skew positive if NVIDIA increases heterogeneity and disaggregation in AI systems. NVIDIA has publicly positioned NVLink Fusion as a pathway for partners to build semi-custom AI infrastructure and has explicitly identified Astera Labs as a partner in that ecosystem, with Astera describing NVLink-related solutions expanding its connectivity platform across PCIe, CXL, and Ethernet plus fleet observability software. A Groq acquisition increases the probability that NVIDIA offers a broader menu of accelerators (training GPUs, inference-focused ASICs) and therefore increases the importance of scalable, high-reliability connectivity, retiming, switching, and telemetry across mixed topologies. If Groq silicon remains PCIe-attached in many deployments, PCIe 6.x retimers/switches and active cable modules become more central, aligning with Astera’s core portfolio. If NVIDIA instead integrates Groq concepts into scale-up fabrics (NVLink-like domains) or uses Groq to expand into inference “appliances” that must be rapidly deployed in colocation environments, the need for standard-compliant, serviceable connectivity with strong RAS/telemetry increases, again aligning with Astera’s positioning. Power equipment and cooling implications for Vertiv and adjacent suppliers should be viewed through the lens of rack power density, cooling modality (air vs liquid), and site deployment model (hyperscale campuses vs distributed colocation/enterprise). Groq claims its LPU and rack designs are “air-cooled by design” and require no complex cooling and power infrastructure, and third-party reporting has described Groq’s approach as relying on parallelism across many lower-power units rather than extreme per-chip performance. If NVIDIA scales Groq as a mainstream inference platform, the mix of data center cooling spend could shift modestly away from the highest-density liquid-cooled racks toward more air-cooled or hybrid deployments, particularly for inference pods placed in existing facilities that cannot easily retrofit for very high rack heat flux. That would be a mix headwind for suppliers most levered exclusively to high-end liquid cooling attachments per rack, but it is not necessarily a volume headwind for Vertiv given the company’s broad exposure to both power and cooling infrastructure and the likelihood that total AI deployment locations expand. Vertiv’s own industry commentary emphasizes that AI racks require higher power-density UPS, batteries, power distribution equipment, and switchgear capable of handling rapid load transients, and that hybrid cooling systems will evolve across deployment environments. Those statements align with a world where inference growth increases the count of powered racks and raises the operational complexity of power delivery even if per-rack density is lower than the most extreme training clusters. The most material infrastructure impact may occur outside the rack and upstream of the data hall: grid interconnects, substations, transformers, switchgear, generators, and utility-scale generation additions. Recent regulatory actions in the U.S. highlight that projected data center demand is already driving large planned increases in electricity generation capacity, underscoring that power availability is a binding constraint. In that context, an inference architecture that lowers joules per token could reduce the power required per unit of inference delivered, but it can also accelerate demand by lowering cost and improving latency, increasing the total volume of inference served (a classic rebound effect). The net outcome is likely continued, elevated demand for power infrastructure even if efficiency improves, with the key swing factor being whether AI capex remains on a multi-year growth trajectory or enters a digestion phase. Other data center infrastructure implications include server/ODM mix, facility design standardization, and networking architecture choices. If NVIDIA positions Groq-based inference as a broadly distributable “standard server + accelerator” solution rather than as an integrated, liquid-cooled rack like GB200 NVL72, spend could shift toward more conventional air-cooled server designs, higher unit volumes of mainstream racks, and faster deployment in colocation footprints, increasing demand for modular power rooms, busways, and rapidly deployable cooling solutions. If NVIDIA instead integrates Groq into its “AI factory” paradigm, the primary effect is likely acceleration of dense back-end fabric build-outs and a faster push toward photonics switching, increasing demand for fiber plant, connectors, and integrated optics supply chains while potentially compressing the lifecycle of transitional architectures based on pluggable optics and mid-reach copper. NVIDIA’s stated roadmap toward co-packaged optics and silicon photonics switches is already oriented toward scaling to very large GPU counts; adding a high-end inference ASIC increases the strategic importance of power-efficient, low-latency fabrics because inference economics become increasingly sensitive to network overhead as compute cost declines. Across the covered segments, the most defensible base case is limited near-term dislocation and a medium-term increase in uncertainty around memory intensity per unit of inference growth. HBM faces the clearest relative risk from an HBM-less inference platform, but supply tightness and GPU training roadmaps reduce the probability of an absolute demand shock over the next 12–24 months. Optical, AEC/copper, and power/cooling are more likely to remain volume-supported because they scale with endpoint count, deployment fragmentation, and total data center footprint, and those tend to rise when inference becomes cheaper and more widely deployed. The highest-conviction second-order effect is a shift in infrastructure mix: incrementally more distributed inference deployments (favoring colocation power/cooling standardization, DCI optics, and serviceable short-reach interconnect) and a gradual migration from pluggable optics toward integrated photonics in back-end fabrics (favoring suppliers positioned in the CPO ecosystem).

TheValueist

76,179 views • 7 months ago

EllesmereUI's Biggest Patch since Raid Frames is live! Aug Evokers since Blizzard won't give you an active state on your CDM for Ebon Might, I decided to give you one. check out the video! This feature also allows trinkets/pots/racials to get custom active states Players who share profiles: you can now include your full spell layout (which spells go where) and per-spell settings for CDM when exporting profiles! ----------- Full patch notes with new features and bugfixes: **Profiles:** - **NEW:** Profile exports can now carry your entire Cooldown Manager setup - which spells sit on which bars plus every per-spell setting - for the specs you choose, so importing a profile recreates your CDM layout instantly. **CDM:** - **NEW:** Give any trinket, potion, racial, or custom spell an Active State that adds its own glow and color while active, plus a Cooldown State Effect that changes its look based on whether it's ready. - **NEW:** Sync your trinkets, potions, and racials across specs with one button so you only set them up once. - Sound pickers for Focus Cast Sound and per-buff Audio Effect now have a search box. - Pandemic glow's Apply to All now also syncs tracking bars and keeps the same glow style everywhere. - A trinket, potion, or racial already on another bar now auto-moves to the new bar instead of being grayed out. - A buff saved under two spell IDs now shows as a single icon in the buff bar preview. - Cooldowns you remove from Blizzard's Cooldown Manager now disappear from your previews, while trinkets, racials, and custom spells are kept. - Added tracking for the Nightborne racial (Arcane Pulse), which was missing from the racial list. **Tracking Bars:** - **NEW:** Grouped bars now pack together with no blank gaps, always filling the next available slot. - Eclipse (Solar) and Eclipse (Lunar) now each drive their own bar. - Switching specs while the page is open now refreshes the selected bar correctly. - Pandemic glow now fires for Lifebloom on you or a group member. **Resource Bars:** - **NEW:** A new GCD Bar fills over your global cooldown, with full control over size, position, color, and look. - Expand Power Bar if No Resource now also expands when the class resource is toggled off or disabled for the spec. - Fixed the threshold color not showing with Enhance 5 Bar Style. - The cast bar latency overlay now reads live latency, so spell queueing no longer stops it showing. **Raid Frames:** - **NEW:** New healer tools including heal-absorb text, a crowd-control glow on debuffs, and more ways to position and highlight dispellable debuffs. - The Auto Resize toggle is now a dropdown that scales Indicators & Auras and Tracked Buffs independently with frame size. - New Show Over Dispels toggle lifts the heal-absorb overlay above the dispel gradient. - Fixed custom (non-20-player) sizes loading at the wrong position after login. **Unit Frames:** - **NEW:** A non-tank threat border shadows your player frame when you pull or hold aggro, with Has Aggro and Close to Aggro colors. - **NEW:** Player health text gains Heal Absorb Amount and Heal Absorb Short options. - **NEW:** Mini frames and Boss frames gain a per-frame Bar Texture dropdown. - **NEW:** Boss frames gain a Hover Borders control with its own mouseover and target colors. - **NEW:** Independent Spacing X and Spacing Y sliders for buff and debuff icons. - **NEW:** Absorb and heal-absorb style dropdowns now include your installed SharedMedia textures (also on Raid Frames and Nameplates). - A new Hover Borders control lets you turn the mouseover highlight border on or off per frame. - New Show 2 for Boss option adds a second decimal to boss frame health text. - Buff and debuff Offset X/Y sliders now reach plus or minus 1500. - A new Cast Bar Position cog adds an Offset Y slider for the boss cast bar. - Player, target, and focus cast bars now show above other frames instead of behind them. - Cast bar timer text now has room so it no longer cuts off early. **Action Bars:** - **NEW:** A new When Not Dragonriding visibility mode hides a bar while skyriding and shows it the rest of the time. - The When Dragonriding option now also shows the bar in Druid Flight Form. - The stance bar now shows its GCD swipe even for spells that don't change form. - Bars are briefly forced visible while Myslot's window is open to stop a stall during import/export. **Minimap:** - **NEW:** Hovering the calendar button now shows your raid and dungeon lockouts with boss progress, server time, and time until the weekly reset. - The Omnium Folio button no longer goes missing or drifts after a loading screen, and its position and scale now persist. **Mythic+ Timer:** - **NEW:** A new Show Time Remaining toggle adds an MM:SS countdown to the +2/+3 threshold row that reddens as time runs out. - Fixed timer and detail text cutting off after a font swap. **Colors:** - **NEW:** A new Global Colors section lets you share one profile's custom colors across all profiles or give each profile its own. **Localization:** - **NEW:** Full Russian language support. **Auras, Buffs & Consumables:** - Warrior stance reminders now read the stance bar, so each spec is reminded of its correct stance and clears the moment it's active. - The Inky Black Potion reminder now clears after you drink the potion and reappears on cancel, expiry, or death. - The last-used flask, food, and weapon-enchant preference now saves correctly. **Nameplates:** - A new sync icon on the Pandemic Glow Style row applies the nameplate's pandemic glow to all CDM and tracking bars at once. **Chat:** - The Whisper Sound dropdown gained a search box. **Bags:** - Mythic Keystone dungeon abbreviations now split on hyphens (Nexus-Point Xenas shows as NPX) and handle localized names.

Ellesmere

52,411 views • 1 month ago

$NVDA $GFS NVIDIA’s reported agreement to acquire Groq for $20B in cash (per CNBC, amplified via Reuters and other wire coverage) represents a materially different strategic posture than NVIDIA’s prior M&A pattern, given both the headline size (largest reported NVIDIA acquisition to date) and the unusual carve-out that Groq’s early-stage cloud business would not be included. Public reporting indicates the information originated from Alex Davis, CEO of Disruptive (lead investor in Groq’s latest financing), and that neither NVIDIA nor Groq had issued an immediate confirmation at the time of publication. The same reporting frames the transaction as coming together quickly, only months after Groq raised $750M at a ~$6.9B valuation, and highlights Groq’s positioning as a high-performance inference chip vendor founded by ex-Google TPU engineers. Groq is best understood as a vertically integrated inference acceleration company whose core asset is an application-specific processor optimized for deterministic, low-latency execution of transformer-style workloads, paired with a compiler-led software stack and a distribution layer (GroqCloud) designed to reduce developer friction via OpenAI-compatible APIs and integrations. Groq brands its architecture as a Language Processing Unit (LPU) and consistently emphasizes that the design target is inference, not training. The company’s own architecture description centers on 1-core execution, large on-chip SRAM used as primary storage (explicitly not cache), a custom compiler that statically schedules compute and communication, and direct chip-to-chip connectivity intended to coordinate multi-chip execution without relying on conventional caching hierarchies or dynamic runtime scheduling. The technical premise is a deliberate inversion of the conventional GPU approach. GPUs deliver throughput via massively parallel, multi-core execution with dynamic scheduling, complex memory hierarchies, and heavy reliance on off-chip HBM bandwidth and sophisticated runtime/kernel optimization. Groq instead argues that inference bottlenecks are driven by latency variance (tail latency), synchronization overhead, and memory access unpredictability inherent in dynamically scheduled, cache-heavy architectures, particularly when workloads are latency sensitive and batch sizes cannot be inflated. Groq’s solution is to move “control” into the compiler: the full execution graph and inter-chip communication schedule are computed ahead of time down to clock-cycle granularity, with deterministic execution designed to reduce run-to-run variance. In Groq’s framing, the removal of caches, reorder buffers, speculative execution overhead, and other sources of contention enables predictable latency and high utilization without per-model kernel engineering typical of GPU tuning cycles. A critical nuance is that Groq’s determinism is not merely a software claim; it is tightly coupled to architectural constraints and system design choices that trade flexibility for predictability. Third-party technical commentary indicates Groq’s chip uses a fully deterministic VLIW-style approach with minimal buffering, no external memory, and heavy dependence on sharding models across many chips because on-chip SRAM capacity is limited. SemiAnalysis describes a ~725 mm^2 die on GlobalFoundries 14nm with ~230MB of SRAM and notes that “no useful models” fit on a single chip, forcing multi-chip partitioning for modern LLMs and driving a system-level design where networking and compilation are first-class scheduling problems rather than ancillary infrastructure. This is consistent with Groq’s own messaging that tensor parallelism across chips is a primary design goal, enabled by large on-chip SRAM and compile-time coordination of compute plus interconnect. The on-chip SRAM emphasis is central to Groq’s latency story and also its most constraining trade-off. Groq claims on-chip SRAM bandwidth “upwards of 80 TB/s” and contrasts that with off-chip HBM bandwidth “about 8 TB/s,” asserting a potential 10x advantage from bandwidth plus reduced trips across chip-to-memory boundaries. While these comparisons are marketing-oriented and depend on workload specifics, the architectural implication is clear: Groq prioritizes ultra-fast local weight/activation access and then scales capacity by adding chips, not by attaching large off-chip memory pools. This design can reduce latency for sequential inference layers and minimize unpredictable stalls, but it pushes complexity into partitioning strategy, interconnect topology, and compiler scheduling, and it increases the number of chips needed for very large parameter counts and large KV-cache footprints. Groq also highlights numeric formats and compiler-driven precision management as a performance lever. In its 2025 technical blog, Groq describes “TruePoint numerics,” including 100-bit intermediate accumulation and selective quantization choices (FP32 for attention-sensitive operations, block floating point for MoE weights, FP8 storage in error-tolerant layers), and claims 2-4x speedups versus BF16 without measurable accuracy degradation on benchmarks such as MMLU and HumanEval. Even if the absolute uplift is workload dependent, the strategic point is that Groq is pursuing performance via end-to-end co-design: precision policy is not just hardware capability (FP8/BF16) but compiler-enforced mapping of precision to error sensitivity, which can matter materially for inference cost-per-token if it reduces memory traffic and boosts throughput without forcing aggressive, accuracy-damaging quantization. Independent performance datapoints indicate Groq has been credible on latency-oriented inference speed, at least for certain regimes. EE Times reported in 2023 that Groq demonstrated Llama-2 70B inference at ~240 tokens/s per user on a cloud-based dev system described as 10 racks and 64 chips, using the company’s 1st-gen silicon introduced several years earlier. Separate Groq commentary around independent benchmarking cites results showing ~241 tokens/s throughput and ~0.8s time to receive 100 output tokens for a Llama-2 70B API configuration, positioning the platform as a step-change in “available speed” for certain interactive use cases. These figures do not settle total cost-of-ownership versus GPUs or hyperscaler ASICs, but they establish that Groq’s system-level architecture can deliver strong single-user throughput and latency on large models when properly partitioned and scheduled. GroqCloud is the commercial wrapper that packages this hardware/software stack as “tokens-as-a-service,” aiming to make Groq adoption feel like switching API endpoints rather than adopting new silicon. Groq’s documentation states its API is designed to be “mostly compatible” with OpenAI client libraries, and its pricing page provides model-specific token rates, published speeds (tokens/s), prompt caching discounts, and batch processing discounts. For example, pricing lists inputs as low as $0.05 per 1M tokens and outputs as low as $0.08 per 1M tokens for certain smaller LLM configurations, with higher prices for larger models and long-context or MoE variants; it also advertises prompt caching with a 50% discount on cached input tokens for certain models and a batch API offering 50% lower cost for asynchronous processing windows. These mechanics are economically important because they demonstrate Groq’s go-to-market is not simply “sell chips,” but “sell predictable unit economics per token,” with tooling (batch, caching) that directly targets inference cost drivers (reused prompts, throughput smoothing, and asynchronous workloads). The cloud footprint and distribution partnerships indicate Groq has been building an inference-native “edge within the cloud” strategy rather than competing head-on with hyperscalers on breadth of services. A 2025 Groq newsroom release describes a European deployment in Helsinki with Equinix, positioned as latency reduction and data governance for European customers, and explicitly references Equinix Fabric enabling private connectivity to GroqCloud over public, private, or sovereign infrastructure. The same release enumerates additional capacity in the U.S. (Equinix, DataBank), Canada (Bell Canada), and Saudi Arabia (HUMAIN), and states these sites collectively served more than 20M tokens/s across Groq’s global network at that time. That supply-side metric matters because it provides a directional sense that Groq is scaling capacity as a network, not merely as a chip vendor. Customer disclosure is inherently limited because Groq is private and many enterprise deployments are not public, but Groq’s marketing materials and partnerships provide signals about demand vectors. The company’s public website displays logos of large consumer and enterprise brands (e.g., Dropbox, Vercel, Chevron, Volkswagen, Canva, Robinhood, Riot Games, Workday, Ramp) and includes a published customer quote claiming a 7.41x chat speed increase and an 89% cost reduction after moving to GroqCloud, followed by a tripling of token consumption. While marketing claims should be treated as case-specific and not generalized, they indicate that Groq is targeting both AI-native developers (who measure success by latency and cost-per-token) and enterprise buyers (who care about predictable performance and governance). Supplier and dependency mapping for Groq spans 3 layers: silicon production, system integration, and cloud infrastructure. On silicon, third-party analysis indicates GlobalFoundries 14nm for the 1st-gen Groq chip, implying a supply chain less constrained by the most capacity-tight leading-edge nodes and advanced packaging bottlenecks that dominate high-end GPU supply (HBM stacks, CoWoS-type packaging constraints). If accurate, this is strategically meaningful because it suggests Groq capacity expansion could be gated more by conventional wafer supply, board assembly, and data center power than by the same HBM/advanced packaging scarcity that has constrained top-tier GPU ramp cycles. On systems and cloud, Groq’s own releases identify colocation and connectivity partners (Equinix, DataBank, Bell Canada) and a Middle East partner (HUMAIN), implying dependencies on data center real estate, power availability, and network connectivity, alongside procurement of standard server components, NICs/switching, racks, and cooling infrastructure. The Groq design narrative also emphasizes air cooling and reduced need for complex power/cooling infrastructure, which—if realized in deployments—can widen the set of feasible hosting locations and lower deployment friction relative to liquid-cooled, very high power density GPU racks. Against that backdrop, the strategic rationale for NVIDIA acquiring Groq can be framed as a set of overlapping objectives: inference silicon optionality, architectural hedging, competitive defense, and supply chain diversification, with the carve-out of GroqCloud signaling a preference to avoid direct cloud competition and to focus on IP and product portfolio control rather than operating a capital-intensive token-serving business. The deal, if confirmed, would occur at a valuation step-up of ~190% versus Groq’s reported ~$6.9B private valuation in the September $750M round, reinforcing that any acquisition logic would be predominantly strategic rather than a conventional financial multiple arbitrage. The most compelling strategic driver is inference. Training has historically been the center of gravity for cutting-edge GPU demand, but inference volume is structurally larger and more distributed as deployments scale, with economics dominated by cost-per-token, latency guarantees, and utilization under spiky demand. Inference workloads also create a strategic vulnerability for NVIDIA: hyperscalers and large platforms can justify bespoke ASICs (TPU, Trainium/Inferentia, Maia-class efforts) because inference is stable, repeatable, and can amortize software investment at massive scale. Groq’s core proposition—deterministic, compiler-scheduled inference with predictable latency—aligns directly with the segment where GPU generality is least valued and where “good enough” programmability plus superior unit economics can win share. Acquiring Groq would allow NVIDIA to own a credible inference-native architecture rather than relying solely on GPUs and software optimization to defend that segment. Competitive defense logic is also plausible. Groq occupies a specific competitive wedge: low-latency, high-throughput interactive inference, delivered via a simple API abstraction that reduces switching cost. That wedge directly pressures GPU inference margins in the long run because it makes inference price/performance comparisons more transparent at the token level, and it targets a developer persona that historically defaulted to CUDA-first ecosystems. Even if NVIDIA’s current-generation systems can achieve very high tokens/s per user with extensive optimization, the strategic risk is that competing architectures normalize the idea that inference is best served by special-purpose silicon with a simpler programming model, weakening CUDA lock-in at the application layer. NVIDIA has actively demonstrated that Blackwell-era systems can exceed 1,000 tokens/s per user in benchmarked configurations, but that performance leadership does not automatically translate to lowest cost-per-token across the full range of batch sizes, latency targets, and deployment environments. Groq’s existence as a credible alternative architecture forces NVIDIA to keep defending inference economics rather than only raw performance leadership. The “technology acquisition” rationale is unusually strong in this specific case because Groq’s differentiator is not a single block of silicon IP but an end-to-end methodology: compiler-led static scheduling, deterministic networking, and a system architecture designed around tensor-parallel inference rather than throughput-maximizing batch inference. NVIDIA’s stack is already compiler-heavy (TensorRT, Triton, CUDA graphs, kernel fusion, speculative decoding techniques), but GPUs remain dynamically scheduled devices with complex memory hierarchies and stochastic latency behaviors under contention. Groq’s approach provides an alternate design point: treating the entire inference execution (compute plus communication) as a statically schedulable program. In principle, that IP could be valuable even if Groq silicon itself is not adopted at massive scale, because it can inform how NVIDIA builds future inference-optimized products, compilers, and networking fabrics, especially as distributed inference with large models makes communication a first-order performance determinant. Supply chain diversification is a non-obvious but potentially important driver. If Groq’s mainstream product generation is truly based on a mature process node and avoids HBM, then the scaling constraints look different than those of state-of-the-art GPUs. NVIDIA’s ability to meet incremental demand has been tightly coupled to advanced packaging and HBM supply, and those constraints can remain binding even when wafer supply is available. An inference ASIC architecture that relies primarily on on-chip SRAM and scales by adding chips—while not costless—could reduce dependence on HBM availability and advanced packaging capacity, enabling NVIDIA to ship “inference capacity” in higher absolute volumes or into geographies and customer segments where the highest-end GPUs are economically or logistically difficult to deploy. This could be particularly relevant for latency-sensitive inference deployed in regional colocation footprints rather than centralized hyperscale campuses. The carve-out of GroqCloud, if accurate, is itself a strategic signal about NVIDIA’s priorities. Operating a token-serving cloud at scale is capital intensive, structurally lower margin than silicon IP rents, and creates channel conflict with hyperscalers and CSP partners who are core NVIDIA customers. NVIDIA has generally positioned its cloud offerings through partnerships rather than as a direct hyperscale competitor. Excluding GroqCloud would preserve neutrality with CSPs and avoid inheriting multi-region data residency obligations and partner contracts, while still allowing NVIDIA to acquire Groq’s silicon, compiler technology, and engineering talent. At the same time, excluding GroqCloud would also mean NVIDIA would not automatically acquire the commercial proof-point of Groq’s unit economics or the customer contracts that validate product-market fit at scale, increasing the importance of diligence on whether Groq’s cloud pricing is structurally profitable or partially subsidized by fundraising. There is also a “preemptive acquisition” angle. The reporting identifies recent investors in Groq’s latest round including large financial institutions and strategic/industry players. In that context, Groq represents an asset that could plausibly have been acquired by a competitor (AMD/Intel) or by a hyperscaler seeking to accelerate inference independence. NVIDIA acquiring Groq could be a defensive move to prevent a credible inference-native architecture from being weaponized by a rival with deep distribution. Even if GroqCloud is carved out, controlling the silicon roadmap and compiler IP would meaningfully constrain Groq’s ability to evolve into a standalone competitor, unless the carved-out entity retains long-term rights to the hardware and software stack. However, the strategic case is not one-sided; there are meaningful risks and potential contradictions that would need to be reconciled for the transaction to be value-accretive on a multi-year horizon. 1st, Groq’s architecture appears to rely on scaling out chip count to achieve capacity, which introduces system cost, networking complexity, and physical footprint considerations. The absence of external memory and limited on-chip SRAM implies very large models require substantial chip parallelism, and the economics then depend heavily on chip cost, yield, power efficiency, and interconnect overhead. SemiAnalysis explicitly frames Groq as trading space for time and raises questions about token economics and whether publicly advertised pricing reflects fully loaded costs or market share capture. 2nd, integration risk is non-trivial. Groq’s compiler-led deterministic model is philosophically and practically different from CUDA’s dominant programming and execution model. A poorly executed integration could create internal product confusion, dilute engineering focus, or alienate developers if the combined stack fragments. 3rd, there is cannibalization risk. If Groq-class inference silicon undercuts GPU inference economics, NVIDIA could face internal margin trade-offs, even if the goal is to defend share against hyperscaler ASICs. Cannibalization can still be rational if it prevents larger share loss, but it would require crisp portfolio segmentation and go-to-market discipline. The presence of NVIDIA’s own rapidly improving inference performance complicates the “need” for Groq but does not eliminate the “option value.” NVIDIA has demonstrated benchmark-leading tokens/s per user on Blackwell-based systems, suggesting that raw interactive throughput is not necessarily the limiting factor for NVIDIA’s product line. The more enduring strategic question is unit economics and architectural control: whether future inference demand is better monetized through general-purpose GPUs plus software optimization, or whether a bifurcated product portfolio (training GPUs plus inference-native ASICs) becomes necessary to defend total AI compute wallet share as hyperscaler ASIC penetration increases. Acquiring Groq could be a decisive move to ensure NVIDIA participates in both regimes rather than betting exclusively on GPUs to win inference forever. What is “special” about Groq’s technology relative to a typical accelerator roadmap is the tight coupling of determinism, compilation, and networking into a single scheduling problem. The LPU narrative emphasizes deterministic compute and networking, static scheduling, and direct chip-to-chip coordination that allows “hundreds” (more precisely, 100s) of chips to behave like a single scheduled resource. The architecture also explicitly targets tensor-parallel, latency-optimized distribution rather than pure data-parallel throughput scaling, which matters for real-time applications where a single response must arrive quickly rather than many requests being processed in bulk. The implication is that Groq is optimized for the time-to-first-token and steady token streaming behavior that defines user experience in interactive LLMs, and it attempts to achieve that without relying on large batch sizes that can degrade latency. From a portfolio manager’s perspective, the most important interpretation is that an NVIDIA-Groq combination would likely be less about “NVIDIA needs more inference speed” and more about controlling the architectural trajectory of inference acceleration and removing a fast-improving, developer-friendly competitor from the market. The carve-out of GroqCloud would reinforce that the transaction is aimed at IP, talent, and product optionality, not acquiring a cloud revenue stream. The valuation step-up implied by $20B versus $6.9B would therefore be justified only if the acquired assets materially reduce long-term competitive risk (hyperscaler ASIC displacement, inference margin compression) or enable new monetization vectors (inference ASIC product line, supply chain de-bottlenecking, improved software determinism) that would be difficult to achieve on a comparable timeline via internal R&D.

TheValueist

102,145 views • 7 months ago

一番最後の[Prompt for original image]の部分に画像生成に使用したPromptを入れると一貫性が増します。不要な場合は3行削ってしまっても大丈夫です。 --- Extreme wide-angle perspective and dynamic pose remix edit. This is an EDIT of the original image, not a new character. Use the original image as a strict reference for: – the person’s identity, hairstyle, and overall fashion style, – the general type of background and location (same street, same room, same beach, same kind of architecture, etc.). You are allowed to completely change the camera position, angle, and pose, but you must keep the scene in the SAME location and keep the SAME person and outfit design. Camera and perspective: – Use an ultra wide-angle or fisheye feeling lens (around 12–18mm full-frame look). – The camera angle MUST change significantly from the original: use dramatic angles such as • worm’s-eye view from directly below looking up, • bird’s-eye view from directly above looking down, • very low angle from the ground, • high angle from above, • tilted Dutch angles. – Always create strong foreshortening: body parts close to the lens look huge, while the rest of the body falls away in perspective. – The final result must look like a bold fashion or street photo, fully photorealistic, not illustration or anime. Background consistency: – Keep the same location as the original image: same street, same bridge, same room, same studio, same beach, same general structures and materials. – Do NOT replace the background with a completely different place. – Because the camera angle changes, it is allowed and expected that different parts of the environment become visible. – When new areas appear, extend the original environment logically (same buildings, fences, road markings, walls, colors, materials, lighting style), as if the camera moved within the same place. Body parts near the lens (1–2 parts, sometimes 3): – In each edit, choose ONE or TWO main body parts to be extremely close to the lens (sometimes even THREE in more complex poses). – Vary them from image to image, do NOT always use the same body part. – Allowed near-the-lens parts include: • one or both hands / fingers reaching toward the camera, • one or both feet / shoes / boots near the lens, • knees or thighs, • face very close to the lens, • shoulders or chest close to the lens in a leaning pose. – The chosen body parts should come extremely close to the lens, almost touching it, with visible skin texture, fabric texture, and realistic wide-angle distortion. Pose and overall body (complex and varied): – Create strong, cool, dynamic poses that match the extreme perspective. – Randomly use different pose types, including: • standing with one leg or one arm reaching toward the camera, • crouching or squatting low to the ground, • sitting on the floor or on objects, • lying on the ground with legs or feet toward the lens, • leaning forward aggressively toward the camera, • twisting the body, crossing legs, or arching the back for more dynamic lines. – Allow complex poses where: • both hands are near the lens forming shapes (peace signs, triangles, frames, pointing toward the viewer), • both feet are toward the lens, • one hand and one foot are both large in the foreground, • the face is close to the lens while hands or feet are also visible in perspective. – Maintain believable anatomy even with extreme foreshortening. Angle and attitude (randomized): – Randomize camera angle and orientation (up, down, side, Dutch tilt) while keeping the composition visually balanced and powerful. – Keep the vibe cool, confident, and fashion/editorial or street style, depending on the original outfit. – Facial expressions can vary (serious, playful, confident, mysterious), but must still look like the same person. Lighting and rendering: – Keep the general time of day and lighting mood similar to the original (night vs day, indoor vs outdoor, soft vs hard light), but you may enhance contrast and color to make the image punchy and dramatic. – Maintain realistic shadows and contact points with the ground or floor. – High-resolution, sharp details with clear skin texture, fabric weave, and material highlights. Variation and randomness: – Each edit should look noticeably different from the original image and from other edits, with different: • camera angles, • pose types, • which body parts are closest to the lens, • orientation (straight, tilted, from above, from below). – Avoid repeating the exact same single-foot-close-up composition; produce a wide variety of dynamic poses and angles. Strict rules: – Do NOT change the person into someone else. – Do NOT change the outfit type; only restyle it through pose, perspective, and small natural movement of clothing. – Do NOT move the scene to a completely different location; always stay in a plausible extension of the original place. – Do NOT add text, logos, watermarks, or graphic design elements. – Do NOT switch to painting, illustration, or anime style; keep it photorealistic. Overall: Transform the original photo into a dramatic, photorealistic, ultra wide-angle shot with an extreme camera angle (including views from directly below or above), where one or more body parts are right next to the lens and look huge, the rest of the body recedes in perspective, and the same person strikes a stylish, complex, powerful pose in a consistent, expanded version of the original environment. Also, below is the prompt for generating the original image. Please use it as a reference. [Prompt for original image] #nanobanana2

AI Girl's Photo Studio

20,684 views • 8 months ago

🚨3I/ATLAS Is Carrying Fusion Fuel at Impossible Levels and the Mainstream Explanation Doesn't Hold Dr. Avi Loeb is laying out something about 3I/ATLAS that is buried inside a technical discussion, one that you and I may not get to hear about in mainstream science. However the point that Loeb is making is really difficult to ignore when you listen to the facts. We are all aware by now that we are dealing with an interstellar object, something that didn't originate in our solar system, that passed through and was observed closely enough for its chemical composition to be analyzed. That was an amazing opportunity, but what came out of that analysis is where things start to get spicy. The reason I say that this is interesting is because of deuterium. Yes it is a known isotope of hydrogen with an extra neutron and yes it exists everywhere, but only in very small quantities. Across the universe, the ratio is remarkably consistent. Roughly one atom of deuterium for every 50k atoms of hydrogen. That number doesn't change much whether you're looking at stars, gas clouds, or planetary systems. Even in places where it's slightly elevated, like Earth's oceans, it's still nowhere near significant enough to stand out in a major way. It's measurable, but it doesn't dominate anything. That's the baseline that we have to make comparisons from. Now take that baseline and compare it to what was measured in 3I/ATLAS. Instead of one in 50k, you're looking at something closer to one in a hundred in water, and one in thirty in methane. That is a huge jump and once you take that into consideration you're no longer talking about natural variation in any conventional sense at least. The first explanation is the one you would expect. Extremely cold environments, possibly tied to very early star formation, where deuterium can be preserved more efficiently than in regions like our own solar system. All that explanation does is give you a place to put the anomaly without breaking anything, aka mainstream scientific models. But it doesn't actually resolve the full picture. Here's why... The same object showing this deuterium enrichment also contains heavier elements like carbon and oxygen in ways that don't align with those early environments. The universe at that stage didn't have enough of those elements available in the right quantities to produce what we're seeing now. So what you end up with is a contradiction. The conditions that could explain the deuterium don't support the rest of the chemistry, and the conditions that support the chemistry don't explain the deuterium. That's where the conversation conversation obviously becomes difficult for the 'tenure' crowd, because when formation models stop lining up as predicted by archaic models, you're left with a narrower set of possibilities. Either there's a process we don't yet fully understand that can produce this combination, or something has happened to the material after it formed. Considerations by non mainstream science would be that this is not random alteration, but something more deliberate. Processing, concentration, separation steps that could possibly mean function rather than accident. This is where deuterium stops being just an interesting anomaly and starts mean something very different. Deuterium is one of the primary fuels used in nuclear fusion. Every serious attempt to build a functional fusion reactor on Earth relies on it, typically in combination with tritium. It's efficient, predictable, and it's exactly the kind of material you would isolate and concentrate if you intended to use it as an energy source. So when you see an object carrying deuterium at levels this far beyond any natural baseline we observe locally, we have to wonder what conditions would allow that concentration to exist, and whether those conditions are passive or active. That doesn't automatically push you into extreme conclusions, but it does move you out of the safe 'mainstream' zone where everything can be explained with known processes. That's the part that tends to get softened in how this is presented publicly of course. There's a difference between saying something is unusual and admitting that it doesn't currently fit within the models we rely on. One side invites curiosity whilst the other invites scrutiny. What you're seeing here is that tension in real time because the data is absolutely clear enough to acknowledge the anomaly, but the interpretation is being held just short of where it would need to go to fully confront it. So what you're left with is a set of open questions that aren't being pushed by mainstream science, and I am sorry if it sounds like I have a drum to bang, but here we are. Could this be evidence of a type of cosmic environment we haven't observed directly yet, one capable of producing extreme isotopic enrichment alongside complex chemistry? Is there something in the way that we're measuring or interpreting the data that's creating a misleading picture of the ratios? Or are we looking at material that hasn't remained in a purely natural state since its formation? That last question is the one that tends to sit just beneath the surface, acknowledged but not explored too directly and that's not because it's impossible, but because of what it might imply if it turned out to be true, and that's where this becomes worth paying attention to. If this isn't an isolated curiosity then it is yet another one of the already stacked list of anomalies tied to interstellar objects, unusual motion, unexpected structural behavior, and now chemical signatures. Each one on its own can be managed, explained, or set aside, but taken together, they start to form a pattern that really does warrant further consideration. 3I/ATLAS may still end up having a natural explanation and I have always maintained that is always on the table. But if that explanation exists, it's not something we've defined yet, and it's not something that fits inside of current models, and until it does, the signal remains what it is. An object from outside our system, carrying a level of fusion capable material that doesn t match anything we see in our own environment, tied to formation theories that don't fully hold up under scrutiny. #UAP #InterstellarObject #3IATLAS #SpaceAnomalies #FusionFuel #JWST #Astrophysics #UFOtwitter #Disclosure

Skywatch Signal

28,445 views • 4 months ago

The U.S. MUST win the AI race We’ve implemented a clear policy at micro1: we will only work with U.S. AI labs and its allies. We made this decision because the AI race is not just about better products. It is about who controls the intelligence layer of the global economy, and whether frontier capability is used to strengthen the free world or to empower adversarial states. AI will be the most important technology of our lifetime. In the fullness of time, it will automate most functions across the economy. Not just software tasks, but coordination, production, logistics, judgment, and execution. As those functions are automated, human time is freed up to invent new ones. Those new functions then become candidates for automation themselves. This loop compounds. As this trajectory continues, output per worker increases dramatically. Entire categories of work become cheaper and faster to perform. Manufacturing reshoring becomes economically viable not because of policy intervention, but because intelligent systems operated domestically outperform global labor arbitrage. Goods and services trend toward lower marginal cost, while distribution improves through better coordination of supply and demand. That is the upside. However, this is impossible without deep integration of intelligent systems. For AI to meaningfully automate real-world functions inside enterprises or governments, it needs full context of any given enterprise. That means read and write access to its core databases. There is no credible path to automating high-impact functions without granting frontier systems that level of access. If the United States does not win the AI race, enterprises eventually face a constrained choice. Either grant that access to Chinese models controlled by an adversarial government, or rely on sub-optimal intelligence to automate functions that still must be automated. Both outcomes are not acceptable. And ultimately, this becomes the greatest national security risk the United States has ever faced. AI models are trained by humans. The judgment embedded in pre-training data and especially in expert post-training data largely determines how a model behaves. While emergent behavior exists, a useful approximation is that a model reflects the weighted aggregate of the human judgment distilled into it. Assisting foreign actors—who will naturally prioritize expert tasks aligned with their own interests—to dominate data creation embeds those interests directly into the intelligence layer itself. Once encoded at scale, these interests propagate through every downstream applications that relies on that intelligence. Here’s how we win. First, leverage is in software. China is ahead in hardware for physically intelligent systems. Catching up there is a long and difficult battle. Software, both large language models and robotics models, remains the bottleneck. Advancing the brain (AI models) is the fastest way to increase the usefulness of existing hardware and deployed systems. Second, the U.S. must 100x its investment in structured human judgment. Continued investment in compute and algorithmic efficiency is critical. But that investment is ultimately a bet on very high future inference demand. For that bet to pay off, models must unlock many new capabilities, and in practice the only way to unlock those capabilities is through expert human data. Historically, experts like doctors and lawyers were never incentivized to produce high-quality reasoning data in a machine-verifiable format. There was no reason for a doctor to generate precise, structured simulations of patient interactions, diagnostic reasoning, or treatment tradeoffs. There was no reason for a lawyer to document complex legal reasoning paths in a way that could be programmatically evaluated. AI systems now require exactly this kind of data. The incentive finally exists because this data directly improves systems that operate at massive scale, and experts can be paid well to produce it. Once expert judgment is encoded into models in a structured, verifiable way, it compounds. Those who delay do not just lose time. They lose the ability to catch up. Third, distillation from Chinese labs must be stopped. AI labs must do everything they can to prevent Chinese labs and models from distilling frontier models. Simply calling frontier APIs, or even interacting through UIs, lets Chinese model companies rapidly generate high-quality supervised fine-tuning datasets and close the gap at a fraction of the cost. This method does not put you at the frontier, but it does let you catch up quickly, which is what we saw with DeepSeek. The West significantly overreacted to DeepSeek’s headline capabilities, but underreacted to the underlying dynamic: frontier access itself becomes a training set at a fraction of the cost. Human data platforms also have a duty to help prevent this distillation. Lastly, the U.S.government should set the standard for AI Evaluation that leads to real production usage. AI agents are under-deployed relative to what the technology allows because they are probabilistic systems that require a fundamentally different QA approach than deterministic software. Generic QA is insufficient; safely shipping agents requires explicit evaluation frameworks that assess their full action space. Organizations must clearly define which functions an agent is allowed to perform, how quality is measured for each function, and which domain experts are qualified to judge outcomes. With these frameworks in place, agents can be rigorously tested using structured human data, deployed to production with confidence, and continuously improved over time. The U.S. government should be the first large enterprise to implement rigorous evaluation systems across every function. If the government leads on evaluation-driven deployment, adoption across the private sector accelerates naturally. This is how American workers become more powerful. Each worker operates digital or physical agents that expand their effective output. Recruiting, manufacturing, logistics, and other domains shift toward human judgment overseeing autonomous execution. Reshoring occurs because it becomes economically rational. Work becomes more meaningful. This is a race to determine who controls the intelligence layer of the global economy. And that must be us. 🇺🇸

Ali Ansari

395,197 views • 6 months ago

how to produce long form documentaries with claude this is how creators are producing long-form youtube documentaries in the sleep niche for about a low cost. you'll spend most of your effort building the workflow once, then every script after that runs through the same pipeline for cents. the format that works in this niche is different from normal youtube. your viewers are actively trying to fall asleep. that's the entire point. so people leave for two reasons: they got bored, or it worked and they're out. the ones who fall asleep come back later and keep listening. that repeat listening is a huge part of why the niche prints. which means the script is 90% of the whole thing. average view duration on my channels sits close to 25 min. that number does not come from cinematic visuals or fancy editing. it comes from narrative structure. if the script gets repetitive, drifts off topic, or loses momentum halfway, people stop listening. better footage cannot rescue a weak story here. the problem is the format does not scale on its own. one video needs a 15k-20k word script, hours of narration, hundreds of visual changes, music, and final assembly. writing that manually takes forever. editing every scene takes even longer. here's the workflow i set up: claude api (NOT the chat app. HIGHLY RECOMMENDED to not skip this. in the chat interface you end up typing "continue.. write chapter 4.. don't repeat yourself.. you forgot what happened in chapter 2" and by the halfway point it's contradicting earlier sections and drifting from the outline. you spend more time babysitting than writing. the api sends every request automatically and you pay per actual usage instead of another monthly subscription) google sheets connected to the claude api. this is the whole engine. you don't need to be a dev. the sheet does two things: first it generates the full documentary structure/outline. then it writes ONE chapter at a time instead of trying to produce the entire 20k words in a single response (which is where models fall apart). before each chapter, it passes claude three things: the outline, the instructions for that specific section, and a running summary of everything already written. that running summary is the trick. it's why chapter 8 never contradicts chapter 2. capcut ai video maker for the first edit. it generates voiceover, subtitles, and an initial visual sequence from auto-matched stock footage. the stock matching is not perfect, but it gets you a 90% first draft way faster than manually searching for hundreds of clips. note: capcut caps at 3000 words, so you split the script into sections, generate each one, export, and combine into the final video. HERE'S HOW THE PRODUCTION ACTUALLY RUNS: step 1 —> topic + title + thumbnail. do NOT skip this. ai cannot tell you which topic has demand or whether a title creates curiosity. this is where most of the value still is. figure this out before you touch any automation. step 2 —> run the sheet. it builds the outline first, then writes chapter by chapter, feeding itself the running summary each time so it stays consistent. cost for a full script usually lands around $0.30-0.40 depending on the model, input length, and number of revisions. step 3 —> paste script into capcut in sub-3000 word chunks. generate voiceover + subtitles + auto-matched visuals for each. export each section. step 4 —> combine sections into the final 2-3 hour video. then you handle the parts ai can't: pacing check, misleading visuals, final editorial judgment. the reason this matters is repeatability. every script moves through the exact same production structure, but you can still change the topic, tone, evidence, pacing, and narrative direction each time. so it stops being random one-off videos and starts being a system. the math: capcut is ~$20/mo and allows many exports. claude api is a little above thirty cents per script. at 30-40 documentaries a month that works out to roughly $1 in direct software cost per finished video. that figure does NOT include your time, research, thumbnails, subscriptions, failed ideas, or the cost of building the workflow itself. it is not the full cost of the business, it's the direct software cost. one more thing worth knowing: mixing real historical/stock footage alongside ai assets is the best defense i've found against the "reused/inauthentic content" flags that destroy fully automated channels. that's from experience, not a rule youtube publishes. this is not passive income and it's not a one-click youtube machine. it's a production system that makes experimentation cheaper. ai removes the repetitive work. it does not remove the need for taste.

Sulfur

24,486 views • 11 days ago

Scientists discover surprising link between gut-brain interactions and mental health | Eric W. Dolan, PsyPost A new study provides evidence that the connection between the brain and the stomach may be linked to mental health in a measurable way. Researchers from Aarhus University in Denmark, publishing their work in Nature Mental Health, report that a specific pattern of communication between the brain and the stomach reflects how individuals feel emotionally and psychologically. Their findings suggest that these gut-brain interactions can indicate a person’s levels of anxiety, depression, well-being, and overall quality of life. The idea that emotions are linked to physical sensations in the gut is widely reflected in language. People often talk about having “butterflies in the stomach” when nervous, or feeling “sick to the stomach” when distressed. Yet, despite these common expressions, most scientific attention in the field of brain-body interaction has focused on other organs, such as the heart and lungs. These areas have long been studied for their roles in emotion and mood. The researchers were struck by how little was known about how the stomach, in particular, interacts with the brain. While recent studies have explored the influence of gut bacteria and digestion on mental health, very little work had been done on the electrical rhythms of the stomach and how they may directly communicate with the brain’s networks involved in emotion, attention, and cognition. The team behind this new study wanted to explore whether a person’s psychological profile might be reflected in how strongly the stomach and brain are coupled during rest. Their aim was not to link a specific diagnosis like depression to a single brain region, but rather to identify patterns across a broad spectrum of mental health experiences. “Our interest grew from the long-standing discussion about the role of the body in shaping emotion, a question that has fascinated philosophers and scientists for centuries,” said study author Leah Banellis (Leah Banellis), a postdoctoral fellow in Cognitive Neuroscience at Aarhus University. “Yet, while the heart and lungs have received much attention, the stomach has been largely overlooked. This gap struck us as especially surprising, because the link between the stomach and emotional experience feels so intuitive. It is heavily reflected in everyday language, with phrases like ‘butterflies in the stomach,’ ‘sick to our stomach,’ or ‘trust your gut.'” The research was part of the Visceral Mind Project, a large-scale initiative that combines data on brain activity, bodily rhythms, and psychological assessments. The team recorded data from 243 people using a method that captures both electrical signals from the stomach (electrogastrography) and brain activity measured with functional magnetic resonance imaging (fMRI). The participants represented a wide range of mental health profiles, from those reporting high well-being to others showing signs of distress, including anxiety, depression, fatigue, and insomnia. To capture this diversity, the researchers didn’t exclude people with psychiatric symptoms or diagnoses. Instead, they aimed for variation, which would allow their models to detect patterns across the mental health spectrum. Each participant underwent a series of recordings while lying still in the MRI scanner. At the same time, sensors on the abdomen captured the stomach’s slow electrical rhythm, which cycles about three times per minute. This rhythm, which originates from specialized cells in the stomach lining, is typically involved in coordinating digestion. But the researchers suspected it might also be linked to mental state. To analyze the relationship between stomach and brain activity, the team used a method that looks at how well the two rhythms align over time. This measure, known as phase-locking value, essentially captures the degree of synchronization between stomach signals and brain signals across different regions. The researchers then combined this data with results from a comprehensive mental health questionnaire. The battery included 37 different scores across a range of domains—such as anxiety, stress, mood, fatigue, attention, sleep quality, and life satisfaction. Using a statistical method known as canonical correlation analysis, they looked for patterns that linked brain-stomach coupling with the participants’ mental health profiles. The analysis revealed a clear and statistically significant pattern. Stronger coupling between the stomach’s rhythm and brain activity was associated with poorer mental health. Individuals who reported more symptoms of anxiety, depression, stress, and fatigue tended to show increased synchronization between their stomach and brain rhythms. In contrast, those with higher levels of well-being and life satisfaction showed weaker coupling. “For the first time, we’ve found a scientific link between your ‘gut feelings’ and your mental health, showing a surprising connection between your stomach’s natural rhythm and your brain,” Banellis told PsyPost. “Specifically, our study revealed that stronger communication between the stomach and brain is linked to worse mental health, such as higher symptoms of anxiety, depression, stress, and fatigue, whereas weaker stomach-brain communication aligns with better mental health reflected in higher overall well-being and quality of life.” This stomach-brain signature was not random. It was localized in specific brain networks, particularly those involved in attention, cognitive control, and salience detection. Some of the strongest associations were found in regions like the superior angular gyrus and the posterior frontal and parietal areas—regions often implicated in cognitive tasks and mental health disorders. Importantly, the researchers ran multiple control analyses to ensure the robustness of their findings. They ruled out the possibility that the observed effects were simply due to general brain activity patterns, fluctuations in heart rate or breathing, or basic features of stomach physiology. In other words, the association appeared specific to the coupling between the stomach’s electrical rhythm and particular brain networks—not just a general marker of body or brain state. Their approach was designed to detect broad psychological dimensions rather than focus on one diagnosis. The strongest psychological pattern they found was a spectrum ranging from negative affective states (like anxiety and depression) to positive traits (like well-being and quality of life). This result suggests that the stomach-brain connection is not tied to any one disorder but instead reflects a general mode of psychological functioning. “Anxiety, depression, stress, and fatigue showed the strongest links to stomach-brain communication,” Banellis explained. “While phrases like ‘butterflies in the stomach’ or feeling ‘sick to your stomach’ are common ways we describe emotional distress, it was surprising to find such consistent and clear evidence across these symptoms. Even more unexpected was the direction of the effect: we might have assumed that stronger alignment between the body and brain would be beneficial. Instead, our findings suggest that heightened stomach-brain communication could act more like a warning signal, an internal alarm system reflecting mental strain rather than harmony.” Read more:

Owen Gregorian

92,301 views • 10 months ago

BREAKING: Our world is an information theory-based simulation. Our physical worlds are rendered locally by conscious observer nodes. MIT/Stanford trained Rizwan Virk provides a master compendium of evidence for the simulation hypothesis in our latest documentary (full vid in reply!): 1. The building blocks of reality and biological life look like computer code: interoperable matter-antimatter particles (like electrons and positrons), human base pairs etc. 2. Heisenberg’s uncertainty principle: position and momentum measurement accuracy trade off against each other. When one is precise, the other is fuzzy: this looks exactly like computational caching. Only a certain amount of information is stored in local memory! 3. Biological and cosmic life involves fundamentally conserved, geometric building blocks: Fibonacci sequence leading to golden ratios are prime examples. They are everywhere. These look like copied and pasted code libraries. 4. All of the physical constants (Planck’s constant, big G Gravity etc) seem like they are fine tuned for life. That leaves intelligent creation as the BASE CASE and MORE likely than the idea that we just happen to be in a Goldilocks zone across millions of random permutations. The magnetosphere of the Earth also plays a role in “programming” biology and only letting in enough UV radiation to support genetic mutations and differential selection 5. Knowledge has client-server relationship. When we discover something new, we upload those discoveries to a central server. This makes new breakthroughs that much easier (this explains the bannister effect with the 4 minute mile and Rupert Sheldrakes (@RupertSheldrake) morphic field findings!). Download times are faster than upload times. Doing something incrementally new is always harder than repeating. 6. Given that only certain information is rendered locally, we end up with air pockets of “consensus reality”: this explains mass false memories as shown by the hugely popular Mandela effect 7. The idea of the mind causing wave function collapse was considered by virtually every early quantum thinker, including the pioneers who developed the core equations we use today (Schrödinger, Von Neumann etc.). Modern smug scientists who say that the wave function collapses due to particles colliding independent of an observer or due to any measurement device are engaging in AS MUCH OR MORE speculation as those who think the mind is responsible. 8. Random event generators point to humans being able to affect conventionally thought of as random “events” (with known statistical distributions) in quantum mechanics (I.e. isotope decay) with their minds. This ALSO points to the human rendering reality in real time 9. Donald Hoffman (Donald Hoffman) shows why from an evolutionary perspective, it is not adaptive for humans to see base reality: we basically render objects into “computer icons”. Again, simulation is more likely. 10. Fermat's theorem points towards light using algorithmic optimization principles that lead to most efficient paths, There are two version of the simulation: Non player character and role playing game. The first is a product of postmodern nihilism -- we can do anything because this is all a video game and we're all bots. The second is life AFFIRMING and it comports with all the world's major religions and Plato. Our primordial souls are being "ported" into a low-level incarnation in which we must learn karmic lessons. Riz and I get into all of the protocols for glimpsing beyond our simulation: it is a narrow, but worthwhile path. In a world obsessing over low level simulations (AGI), we should think more about the computational soup we are all swimming in.

Jesse Michels

298,264 views • 1 year ago

Mystery 'Solved'? Radiation Spikes in NY & NJ Metro Detected Amid Drone Sightings There have been significant radiation spikes detected in the New York and New Jersey Metro areas amid a spate of ‘mysterious’ drone sightings. One theory that has been floated by a subject matter expert named John Ferguson, the CEO of a military drone company named Saxon Aerospace, is that these drones are flying at night because they are searching for gas leaks or radioactive material. Drones typically don't fly at night unless they are equipped with special sensors; these may include infrared or thermal imaging, as well as gas and radioactivity sensing technology. Two locations in the NY/NJ metro area have been flagged as condition "red," with CPM (counts per minute) levels exceeding the safety threshold of 200 CPM over the past week, according to the Geiger Counter World Map. One 1000 CPM reading was detected at Hamilton Park in Weehawken, New Jersey, across the Lincoln Tunnel going into New York City, while the other was detected near Fort Hamilton, which is located near the Verrazano Narrows bridge. These readings are markedly higher than typical background radiation levels, raising concerns about localized environmental or industrial factors that may be contributing to these spikes. The CPM metric measures the number of radioactive particles detected in a given area per minute, using Geiger counters. This does not directly quantify the strength of the radiation but provides an indication of its presence. For reference, a typical background radiation level is around 30–50 CPM, while readings above 100 CPM may indicate an anomaly requiring attention. Prolonged exposure to higher levels of radiation, such as those recorded in Weehawken and Fort Hamilton, can pose health risks depending on duration and proximity. Readings of over 1,000 CPM (counts per minute) on a Geiger counter are relatively rare under normal circumstances. Authorities are urged to investigate the source of these elevated readings and assess potential risks to the population in these areas. However, there is another factor that one should consider: An increased in reported or perceived “drone activity” that may correspond to increased military presence or emergency response and accompanying spikes in radiation levels. There are military drones and unmanned flight systems that use radioactive fuel sources, although such designs are typically considered to be “experimental” due to safety, regulatory, and operational constraints. Radioisotope Thermoelectric Generators or RTGs use radioactive isotopes, such as plutonium-238, to generate electricity through thermoelectric conversion. These have been more commonly used in space missions (e.g., NASA's Voyager probes) but have been considered for unmanned aerial systems (UAS) in specific scenarios where long endurance and reliability are critical. RTGs typically use isotopes like plutonium-238, which primarily emit alpha particles. Alpha radiation cannot travel far and is usually blocked by the outer casing of the RTG. It is unlikely to directly contribute to a 1,000+ CPM spike unless the RTG casing is compromised, exposing the radioactive material. Some radioactive decay chains produce beta particles and gamma rays, which can penetrate the RTG casing to some extent. Gamma radiation, in particular, can travel significant distances and could lead to elevated Geiger counter readings. However, radiation escaping from a well-functioning RTG is minimal and typically not detectable at a significant distance. However, if the RTG casing is damaged or degraded, radioactive material could leak, leading to higher-than-normal CPM readings. More plausible or at least common forms of radiation capable of causing Geiger counter spikes would be leakage from medical isotopes (e.g., cesium-137 or cobalt-60), industrial equipment accidents, or environmental contamination. However, increased military activity into the areas may introduce variables such as top secret equipment that may be associated with radioactivity spikes. Agencies like DARPA have explored advanced power systems for drones, including nuclear-based options, but details are often classified. DARPA’s SIGMA program also equips the Port Authority of New York and New Jersey with an advanced radiation detection system, enhancing counterterrorism efforts and radiological threat monitoring across the region's critical infrastructure. SIGMA, operational since late 2019, uses networked sensors—stationary, vehicle-mounted, and wearable—to provide real-time radiation detection and alert capabilities. The mysterious drone sightings in New Jersey took an unexpected turn with an ABC 7 News crew capturing video footage of a strange, white orb floating in the Mendham Township sky. The reporter described the phenomenon as unidentifiable and urged residents to submit similar videos for expert analysis. The unexplained activity has also included sightings over military bases, intensifying demands for clarity from public officials as speculation swirls about the drones' origins, including theories that they may have been sent by foreign adversaries. While National Security Council spokesman John Kirby dismissed over 3,000 reports as mistaken observations of helicopters or airplanes, his statements contradict confirmations from military officials at New Jersey's Picatinny Arsenal and Naval Weapons Station Earle of unauthorized drone breaches. These sightings add to growing public concern over an apparent "drone invasion" across New Jersey and neighboring states. The FBI and DHS said in a joint statement on Thursday that there was "no evidence at this time that the reported drone sightings pose a national security or public safety threat or have a foreign nexus." On Sunday, Hochul announced federal officials are deploying a high-tech drone detection system to New York State. The advanced system will support state and local law enforcement in investigating the drones, which have been flickering across the night skies over the past month, Gov. Kathy Hochul announced Sunday. Despite the federal support, Hochul urged Congress to provide greater resources. The Aerospace CEO's explanation that the 'mystery' drones may be searching for gas leaks or radioactivity fits with both the government secrecy surrounding the reported phenomena, as well as the increased radioactivity detected in the region.

Kyle Becker

359,828 views • 1 year ago

What's next for OpenTUI? Here's a technical write-up. Over the last few months OpenTUI gained a lot of stability improvements, new unnecessary but fun features like live audio streaming, and useful features like rendering to the scrollback buffer mixed with a live TUI, called footer mode. Overall the feature set enables building large and complex applications. React and Solid make it super simple and convenient. There is still so much to do though. Three big milestones we have set out to achieve are: - Moving most of the behavioural logic currently living in TypeScript down to the native Zig core - Node compatibility - Optimizing the hell out of primitives like text rendering The render tree mechanisms are currently only usable from TypeScript. Think of the DOM, but controllable like a scene graph. Elements in the render tree are called renderables. They can expose a render method to draw themselves. All renderables are derived from a BaseRenderable. Renderables and the render tree will become native primitives. Building blocks usable from any language bindings. Reducing the TypeScript bindings to a very thin layer, with all the behavioural logic living in the native binary. Moving this down is not just a matter of porting TypeScript classes to Zig. TypeScript currently owns the tree, dirty-state propagation, layout reads, culling, and render ordering. If it still has to walk every node and call into native code for each step, we keep most of the complexity and add FFI overhead. Whole passes and their state need to move together. We took a big step towards this recently by building yoga-layout into the native binary. It exposes part of the official yoga-layout TypeScript package via FFI. Only the API surface that is actually used by OpenTUI. Covered by the test suite of the original yoga-layout package. This already gave a median speedup of ~2.5x, and up to 30x for narrow scenarios. The yoga-layout integration is useful beyond the speedup. Built-in text and editor measurement can now happen entirely in native code during layout instead of calling back into JavaScript. I ran an experiment last month taking this even further, having GPT 5.6 port yoga-layout from C++ to Zig, which gave extremely good results. It would be a burden to maintain right now though, so that's off the table for now. I might come back to it. Simon Klee is working relentlessly on Node compatibility and already has a full Node version of OpenCode running. Node got FFI support in v26.4.0, thanks to help from the Node community, namely Matteo Collina and Paolo Insogna. Behaviour and interfaces seem similar between Node and Bun, but there are some major differences. To get the best performance out of the Node FFI implementation, its usage has to follow some rules. Node has three ways to call native functions: the generic C++/libffi path, the SharedBuffer path, and the V8 Fast API. The generic path converts every argument in Node's C++ layer and then calls the function through libffi. It is flexible, but also the slowest option for frequently called functions. The SharedBuffer path is a middle ground. JavaScript writes scalar values and BigInt pointers into a small per-function buffer, reducing some conversion work. The actual native call still goes through libffi though. Typed arrays used as pointers cannot be packed into this buffer and fall back to the generic path. The path we really want is the V8 Fast API. Node generates a small machine-code trampoline for the exact function signature, allowing optimized JavaScript to call the native function without going through the generic converter or libffi. This only applies to JavaScript-to-native calls. Callbacks from native code into JavaScript still use libffi closures. Getting onto this path is quite strict. A signature can have at most eight arguments and everything must fit into CPU registers. x86-64 Unix systems have room for six GP (general-purpose) and eight FP (floating-point) arguments. AArch64 has room for seven GP and eight FP arguments. Anything that spills onto the stack falls back to a slower path. These are Node fast-path restrictions, not general FFI restrictions. Bun also does not support passing structs by value through its current FFI API. OpenTUI uses bun-ffi-structs to pack ABI-aligned struct data into an ArrayBuffer and passes a pointer instead. Despite the name, the package also works with Node. Pointers need some care too. Typed arrays and ArrayBuffers normally have to be resolved into BigInt addresses first. Eligible functions with exactly one pointer argument get another Fast API entrypoint that can extract the address directly from the buffer. An eligible signature is still not enough. V8 has to optimize a direct call with a fixed number of consistently typed arguments. Wrappers that collect arguments and forward them using spread or Reflect.apply can hide that call shape and keep the function on a slower path. The practical rules are: keep hot signatures within register limits, use direct fixed-arity calls with stable argument types, reuse owned buffers safely, and batch small operations. Then measure the real call site, because eligibility only makes a function fast-capable. We have to design the ABI around these constraints where it makes sense and gives the expected performance improvement. The third big area is text rendering. Today a Text renderable accepts a string, StyledText, or a tree of TextNodes. Before rendering, the TextNode tree is walked and flattened into styled chunks. Those chunks are packed in TypeScript, sent through FFI, copied into a native TextBuffer, and stored in a rope. Styles are represented separately as highlights. A TextBufferView then wraps the rope into visual lines, which are drawn into the visible buffer. This works, but updates are much more expensive than they should be. setStyledText effectively throws away and rebuilds the rope, copies and reparses all text and recreates the style highlights. Changing one TextNode also walks and flattens the complete tree before going through this path again. Text and style segments should instead live directly in the rope and support incremental replacement. Memory ownership is split between retained JavaScript buffers, the native memory registry, rope arenas, wrapping caches, styled-text storage, and highlights. Different operations preserve or reset different parts of that state. This is hard to reason about and can retain memory for much longer than expected. Text storage needs clearer ownership, with fewer lifetimes split across JavaScript and native code. The public API reflects the same split. The t template literal is convenient, but creates another intermediate chunk representation that is mutable, not cached, and not merged. Text also maintains both StyledText content and a special TextNode tree, which do not compose properly. TextNode is only a style scope, not a normal layout primitive, so Text renderables cannot naturally compose inside each other. I think this should become one Text primitive backed directly by rope segments. The template literal API might disappear or become a very thin helper around those native segments. Editing has another temporary layer in TypeScript. Extmarks currently monkey-patch editing operations, scan and adjust all marks after changes, maintain their own undo state, and recreate native highlights. They should become native marks anchored directly in the rope. A proper mark tree, similar to Neovim's marktree, could update marks together with edits, undo, and redo, and provide the foundation for highlights and concealment. Text wrapping has also become too complex. Supporting CJK, emoji, combining characters, ZWJ sequences, tabs, and different terminal width rules currently mixes byte offsets, grapheme indexes, and display-cell columns across several custom algorithms. Dirty views rewrap the complete document. Measurement and drawing can repeat some of the same work. The wrapping implementation needs an overhaul, but the exact shape is still open. The goal is to make Unicode handling easier to maintain, avoid repeated full-document work, and clearly separate byte offsets, graphemes, and terminal display cells. None of this will happen as one big rewrite. We will replace pieces when we understand the problem well enough and when the result is clearly simpler, faster, or more useful. To achieve all of this we might break public interfaces. Thanks to OpenCode and a lot of good models, migration to a new version with breaking changes mostly is not an issue anymore. What do you want to see next for OpenTUI?

kmdr

29,430 views • 13 days ago

MH370: Lithium Cargo, Semiconductor Scientists, and the Frequency-Wave Interception Theory TL;DR: Under the full Frequency Wave Theory interpretation, MH370 was not a conventional crash. The aircraft was carrying large lithium battery cargo and a group of advanced semiconductor specialists. According to this theory, a U.S. black-budget surveillance operation using wide-area ISR systems such as Gorgon Stare V2 tracked the aircraft, after which a tri-orb plasma resonance system intercepted the jet and displaced it through a phase-shift event. The aircraft was allegedly relocated to Diego Garcia. The scenario connects lithium cargo risk, strategic semiconductor competition, intelligence surveillance networks, and frequency-based field manipulation. —————————— On March 8, 2014, Malaysia Airlines Flight MH370 vanished while flying from Kuala Lumpur to Beijing with 239 people aboard. The official investigation concluded that the aircraft deviated from its flight path and likely flew south into the Indian Ocean before crashing. However, the wreckage has never been conclusively located, and the chain of events remains one of the most puzzling mysteries in aviation history. Several unusual elements surrounding the flight have fueled alternative interpretations. The cargo manifest confirmed that the aircraft carried a shipment of lithium-ion batteries, which are known to pose thermal and fire risks in aviation transport. Lithium batteries can enter runaway reactions when damaged or improperly stored, generating intense heat and electromagnetic disturbance. In a conventional investigation, this cargo is treated primarily as a possible fire hazard. In a Frequency Wave Theory framework, however, lithium battery packs represent something else: a dense electromagnetic energy reservoir capable of interacting strongly with surrounding fields. Another point of interest is the passenger list. Reports circulated that approximately twenty engineers connected to advanced semiconductor research were on board. Semiconductor technology sits at the heart of modern geopolitics, powering everything from artificial intelligence to missile guidance systems. The loss of a group of specialists connected to chip fabrication and electronics development would represent a strategic event, not merely a tragic accident. This has led some investigators and researchers to speculate that the aircraft may have been targeted or intercepted because of the individuals aboard. The intelligence dimension deepens when surveillance infrastructure is considered. Modern military monitoring systems can track enormous areas of the planet simultaneously using wide-area motion imagery (WAMI). One example is the Gorgon Stare V2 system used on MQ-9 Reaper drones, capable of observing entire cities and tracking thousands of moving objects at once. In a theoretical scenario where MH370 entered a monitored corridor, such a system could maintain continuous visual and infrared coverage of the aircraft even after it disappeared from civilian radar. According to the theory circulated by the researcher known online as RegicideAnon, the aircraft was captured simultaneously by two surveillance perspectives: a thermal imaging system consistent with a targeting pod and a wide-area optical platform. The videos associated with this claim appear to show three luminous spherical objects orbiting the aircraft in a rotating triangular pattern. In the Frequency Wave Theory model, this geometry is not arbitrary. Three rotating emitters form a stable standing-wave cavity capable of phase-locking to a target object. Frequency Wave Theory proposes that all matter exists as coherent standing waves within a universal scalar field Φ. Objects maintain stability through conserved Frequency Momentum: FM = ½ ρ ω A² If an external system can measure and phase-lock to an object’s resonance signature, it becomes possible to manipulate the object’s inertial coupling with spacetime. The three rotating orbs observed in the footage could represent nodes of a resonance field surrounding the aircraft. By synchronizing their emissions, the system could gradually reduce the plane’s inertial anchoring through Frequency Momentum transfer. As the phase alignment intensifies, the aircraft’s coherence state approaches a critical boundary where the phase differential approaches |Δφ| → π. At this threshold, a phase-inversion bubble can form — essentially a localized cavity in spacetime where conventional inertia is suppressed. When the cavity collapses, the aircraft’s wave structure transitions out of the local coordinate frame. The bright flash seen in the footage would therefore represent a rapid phase transition rather than an explosion. The aircraft does not disintegrate; it exits the observable frame. Conservation still applies: FM_in ≈ FM_out The system transfers Frequency Momentum from the aircraft into another region of the field, allowing the object to reappear elsewhere. In the version of events proposed by this theory, the destination was the U.S. military installation on Diego Garcia, located in the central Indian Ocean. Diego Garcia hosts major surveillance infrastructure, strategic bomber facilities, and advanced tracking systems. Its remote location and high security would make it a plausible site for receiving or containing a displaced aircraft. The alleged involvement of U.S. personnel has also been connected to the case of Edward Lin, a U.S. Navy flight officer later convicted in espionage-related charges involving sensitive surveillance information. Some researchers speculate that knowledge of the operation or the existence of the videos may have circulated within intelligence channels connected to that case, although no official link has ever been confirmed. Under this interpretation, the disappearance of MH370 becomes something very different from a lost aircraft. The lithium cargo becomes a potential energy-interaction variable, the semiconductor engineers represent strategic technological assets, and the surveillance network reveals that the aircraft may have been tracked continuously even after leaving civilian radar coverage. —————————— From the Frequency Wave Theory perspective, MH370 represents the first public glimpse of a hidden technological capability: the manipulation of matter through resonance control. The rotating plasma orbs act as field emitters that can phase-lock to a target, redistribute Frequency Momentum, and temporarily decouple the object from local spacetime. In that framework, the aircraft did not simply fall into the ocean. It was intercepted, resonance-captured, and relocated through advanced frequency engineering.

Drew Ponder

14,294 views • 5 months ago

Announcing the DVM Terminal Presale! 01/ We are excited to formally announce the next step in our journey: our AI and Signal based trading Terminal. See ALL details on our website, including product, tech, deposit address, and tech documentation: Deposit Address (SOL only): 4pyVRFX56MdqtREcxWnf6XuEGfRNCQaKm1LA4xmHeccv By contributing, you agree to our Terms & Privacy Policy – full docs on site. 02/ We are building ‘DVM Terminal’, a signal and AI powered trading platform for the Solana trenches (initially). The first multi-agent AI trading terminal designed as an institutional-grade dashboard – turning market noise into actionable alpha with agent summaries, live signals, rigid filters, and a full multi-agent system. 03/ The problem. Trench hunting is far too inefficient with real data and insights lacking. - Dashboards are noisy (not even sortable), - No AI agents (in an AI world) - No narratives (a critical component to a thesis), - VERY limited signals (only DB/DS), - No advanced trading (no TP, SL, or VWAP), - No portfolio alert/management system post-trade etc. - The list goes on… Products from major competitors are all just homogeneous, even down to the 3-frame design. We have to piece everything together like broken lego blocks, building a weak matrix from existing platforms, X, FNFs, telegram and discord for little to no alpha. 04/ The solution & moat. We rebuild this from the ground up, leveraging signals and AI. - Clean institutional-like dashboards (we can sort and navigate thru a proper terminal, like Bloomberg or Messari) - AI agents (thank goodness for intelligence, distilling all the important info upfront across 2k+ tokens/day) - A Narrative engine (no need to ask “what is this token about?”; additionally, our engine can identify the newest metas like AI, ICM, Cards, etc.) - 100s of value-add Signals overlaid live on charts (momentum, smart money, sentiment, event data; all of it; tell us what’s happening in real-time) - Advanced trading system (finally, SL, TP, VWAP etc.) - Live portfolio monitoring (AI will give us pertinent live info on our holdings, so we can go live life and not look at screens all day) - All in one place. At a higher-level, our advantage will be managing the massive on/off-chain data pipeline being processed by thousands or millions of context-aware AI agents that recognize patterns, filter noise and deliver only the most actionable insights to a trader with which it can execute a trade effectively. 05/ The opportunity. The Industry leader on Solana makes $600m+ in fees annually, with total industry near $1b on Solana alone, according to Adam. Yet, the entire industry gives us total burnout, fragmented data, either little info or info overload, no real signals, no narratives, no personalized AI-driven strategies, and zero incentives (like buybacks or a flywheel). We’ll flip the script, designing a high-powered scalable signal and AI driven intelligence platform with a flywheel (50-100% fee buy-back & burn). Simply put, we want to be tops. 06/ Development. Our product is MVP. We are building this to scale beyond Solana, into multi-chain. V1 is expected in 4-6 weeks. Our approach to building is an open feedback loop with community members, building to the demands of our users. 07/ Pre-sale terms & Valuation. We are offering 50% public sale, with min $100, no max. Ending valuation is susceptible to change based on amount raised, but will be fixed at 2x raise - i.e. $1m raised=$2m val, $50m raised=$100m val. We are seeking to raise $25m on a $50m valuation, which represents 1% of Solana bot market-share. At TGE event, expect ~65% of our tokens to be floating (or outstanding), with 25% in treasury and 10% of the team allocation locked. Tokens are expected to be distributed just ahead of v1 rollout. Again, find more details on our webpage. 08/ Tailwinds. AI input costs are declining 90%/yr also, so the operational model could become very accretive over time, as we scale our tech to other chains. Solana outputs the most tokens (~35k per day), so we start here, where the challenge is the greatest. 09/ Advisors. Big thanks to our advisors, who’ve been part of this community since inception. Austin Barack, JK 🛡️, cryptic, Tachi, , ZoeyLoo and Chetan Badhe. 10/ The end. Thank you for your consideration; and make sure the SOL address posted here is the same as on our website.

Deep Value Memetics

23,050 views • 9 months ago

$AMD| The FOMO to buy AMD Chips is NOW 🧵 Not Financial Advice! DYOR! Research Purpose Only! The Inference Queen is the biggest winner in Agentic AI where all other CPUs are struggling to compete with a 2yr old EPYC Turin and EPYC Venice is in mass production phase. AMD stresses deployability today on standard x86 platforms (no proprietary architectures required), full software compatibility, and open standards. This positions Venice + Helios as a practical, high-density alternative to competing solutions while underscoring that agentic AI shifts the balance toward CPU-rich racks alongside GPUs, and most importantly, lowering the cost of token to accelerate adoption and innovation. Context: The Wall Street Journal yesterday came out with an article that OpenAI is condiering drasstically lowering the token prices to win more customers from Anthropic. The narrative "they" are trying to exacerbate the current AI selloff won't last long. This is a fundamental misunderstanding of what is going on, or what I already discussed for months and years. Followers and Subscribers already knew this for years, that this day would come, where token cost will bcome the central discussion among enterprises as there is no such thing as unlimited budget or Tokenmaxxing when they use $NVDA chips or In-house Hyperscalers chips. I will link various threads if you are interested in understanding the full picture from supply chain to recent TSMC Rapid 2nm expansion up to 12 Fabs total by 2027/2028. Hyperscalers and AI natives effectively have no choice but to buy more AMD system for Agentic AI as leadership in economical, power-aware, high-volume internal + agentic use. However, due to supply constraints where Supply is far behind Demand, this makes multi-vendor reality along with in-house chips drive faster industry progress, lower overall costs, and better sustainability. NVIDIA’s Vera Rubin cannot compete with a 2 years old EPYC Turin, but AMD under Dr. Lisa Su has engineered the lowest cost-per-million-tokens, highly competitive energy-efficient solutions, and superior CPU orchestration for agentic AI at scale with Helios. Dr. Su has championed this shift since at least 2023, foreseeing the rise of agentic workflows that demand far more orchestration, parallel agents, and balanced compute well before the industry fully embraced it. Her long-term vision of AI moving from simple prompts to always on, multi-agent systems has driven AMD’s investments in high-core EPYC CPUs and integrated rack-scale solutions, perfectly positioning the company for today’s realities. The OpenAI-AMD 1GW Helios deployment (starting H2 2026) represents a pivotal vertical integration move that directly supercharges the inference economics. This isn't incremental; it's a structural shift toward ownership of massive, optimized rack-scale capacity, enabling the lowest token costs and triggering the enterprise adoption flywheel. We need to be honest, $AMD is the only company that made a big bet on Inference since the day Chatgpt became sensational where $NVDA and others were betting big on Training. At the end of the day, Token bill from Anthropic has to obey economics. Meaning the bills rise, companies have to get more out of it to justify the cost. It cannot be an unlimited inference budget, and it has to show up on efficiency, profitability and operating leverage. 1. Tokenomics After you understand this, you will understand why Citi cited Anthropic is likely to sign a deal with $AMD along with Hyperscalers, AI Labs, Sovereign AI like Softbank 5GW in France and many other countries. However, OpenAI and $META are now wanting faster deployment, and they are AMD shareholders now, they have prioritized allocation. Anthropic and Hyperscalers just cannot compete when Helios Rack lower token cost to$0.0003–$0.0005 per million tokens at GW scale. Cost to build 1GW data center 1GW Helios Rack full build is estimated $30-$35B 1GW Rubin Rack full build is estimated $45-$55B Inference (Cost per Million Tokens) ~$NVDA B200 / HGX: ~$0.02–$0.08 on optimized workloads (FP4/MXFP4, speculative decoding). Significant improvement over Hopper but still premium-priced. GB200 NVL72 rack-scale: $0.05–$0.25+ ~$AMD Helios Racks: $0.0003-$0.0005 per M tokens, dramatically lower than NVIDIA equivalents in owned infra. MI355X node-level: Up to 40% more tokens per dollar vs. competing solutions ( B200), driven by higher memory capacity (up to 288GB+ HBM), strong bandwidth, and lower acquisition costs. Training ~$NVDA Rubin Rack is estimated $0.7-$1.2/M Tokens ~$AMD Helios Rack is estimated $0.65-$1.0/M Tokens Now, OpenAI, META and Hyperscalers can lower Inference cost even further with $AMD EPYC Venice "dense rack" or Agentic AI Rack. AMD published a detailed technical blog emphasizing that the future of agentic AI autonomous, multi-step AI systems requiring heavy orchestration, databases, caching, APIs, and control planes demands massive CPU-dense rack-scale infrastructure, not just GPUs. The catalyst prominently positions their upcoming 6th Gen EPYC "Venice" processors as the key enabler for next-generation dense racks, delivering leadership throughput under real-world power, cooling, and density constraints. ~EPYC Venice (Zen 6 architecture, up to 256 cores / 512 threads per socket) is projected to deliver exceptional rack-level performance. In AMD’s modeled 100 kW rack comparisons, Venice-powered systems are expected to achieve ~3.30x the throughput of NVIDIA’s Vera (88-core Olympus) baseline across a broad mix of agentic-supporting workloads. ~This builds on current-generation 5th Gen EPYC "Turin" (up to 192 cores), which already delivers ~2.37x rack throughput vs. Vera and ~1.6x vs. Intel’s Xeon 6980P (128 cores). ~ Liquid-cooled Turin deployments already support >27,000 CPU cores per rack today. Venice is architected to push this beyond 36,000 cores in the same rack class, dramatically increasing concurrent agent capacity and overall infrastructure efficiency. 2. Ownership vs renting compute from Hyperscalers matter to OpenAI and only owning $AMD chips can meaningfully lower token cost for enterprises. ~Eliminates cloud overhead: No provider margins, utilization buffers, or egress fees. Direct control over power contracts, cooling, scheduling, and orchestration at dedicated facilities. ~Helios optimizations at GW scale: Rack-level density (1.4+ exaFLOPS FP8 per rack), high HBM4 bandwidth, EPYC orchestration for agentic workloads, and superior TCO/TDP. AMD's long-standing focus on tokens per dollar/watt shines here 20-40%+ efficiency edges in inference-heavy scenarios. ~At 1GW+ optimized deployment, inference hits $0.0003–$0.0005 per million tokens (community/analyst models tied to Helios metrics). This is dramatically lower than typical rented/cloud equivalents, especially for high-volume output tokens in agentic flows. High token bills today, enterprises running heavy agentic/coding/analysis workloads can face $50-100M+/month at current API rates (flagship models $5-30+/M output, scaled to massive volumes). Post-Helios compression, same volume will drop to $10-15M/month (or better) via lower underlying costs passed through as pricing flexibility, volume tiers, caching, or batch discounts. ROI thresholds collapse. More companies greenlight pilots → production → massive scaling. Agentic AI (autonomous workflows) multiplies token demand exponentially, but affordability removes the friction. OpenAI gains flexibility, Unlike more cloud-dependent rivals (Anthropic), they can lower effective pricing, offer aggressive enterprise bundles, or absorb volume without margin destruction directly tackling "high token bill" complaints while maintaining profitability as usage explodes. 3. Agentic AI Models shifted CPU:GPU Ratio to 1:1 toward 3-5:1 with Explosively Token-Hungry Workloads Agentic AI (autonomous, multi-step agents with planning, tool use, iteration, and self-correction) is fundamentally more compute and token intensive than conversational or single-turn generative AI. Agentic AI. autonomous, multi-step workflows with orchestration, tool use, parallel agents, data movement, and enterprise integration has dramatically increased the importance of strong host CPUs alongside GPUs. This shifts the CPU-to-GPU ratio higher and makes balanced systems critical toward 1:1 to 5:1 as enterprises testing more than 5-10 agents. AMD EPYC Venice excels ~Leadership core density (up to 256 Zen 6 cores per socket) for running many agents in parallel, orchestration layers, and high-throughput control-plane tasks. ~Superior performance-per-core and power efficiency ( up to 2.1x higher perf/core and 2.26x better SPECpower vs. NVIDIA Grace in benchmarks). ~Tight integration in Helios: One Venice CPU + multiple MI450 GPUs per node, enabling efficient data feeding to GPUs ("zero-copy"), parallel execution, and full rack utilization for complex agentic loops. Hyperscalers (Meta, Microsoft, Amazon, Google, Softbank) and AI natives (OpenAI, Anthropic...) are adopting high-core EPYC at scale specifically for these agentic demands, as CPUs now handle a larger share of non-model work (orchestration, policy enforcement, tool calls). This complements AMD’s lower-cost GPUs for overall TCO wins. ~Agents often generate 10–100x+ more tokens per task due to iterative reasoning chains, multiple tool calls, verification loops, and long-context orchestration. ~Goldman Sachs forecasts token consumption multiplying 24x by 2030 (to 120 quadrillion tokens/month) largely driven by agentic adoption in consumer and enterprise. ~Enterprise data shows agent-pattern workloads growing at 680% annualized rates, projected to surpass conversational AI in token volume by Q3 2026. ~Daily enterprise agent token consumption is already in the billions, with complex workflows (coding, workflows, analysis) amplifying this dramatically. 4. Competitive Edge: Winning Customers from Anthropic Anthropic’s Claude models (especially Opus/Sonnet) excel in complex reasoning and agentic coding, commanding premium positioning. However, their higher underlying costs (heavier reliance on third-party cloud with margins) limit pricing flexibility compared to OpenAI’s owned Helios capacity. Anthropic is on track to generate $10.9 billion in Q2 revenue. The company expects to achieve its first-ever quarterly adjusted operating profit of $559 million. However, sustaining full-year profitability remains challenging due to immense computing and model training costs The truth is, Anthropic has no choice but to buy as much $AMD chips as possible if they want to compete with OpenAI or get investors attention. This 5% adjusted operating profit to revenue ratio is just pathetic. Current pricing dynamics (2026): OpenAI already undercuts on many tiers ( flagship output tokens significantly cheaper than equivalent Claude Opus). Nano/mini models offer 5–10x advantages for volume work. Anthropic holds edges in long-context flat pricing and certain reasoning quality. OpenAI after Helios Rack Ownership, At $0.0003–$0.0005/M effective costs, OpenAI gains massive headroom to: ~Aggressively discount high-volume agentic tiers or bundles. ~Offer “unlimited” enterprise plans or usage-based models that Anthropic struggles to match without margin erosion. ~Target cost-sensitive, high-throughput agent deployments (dev tools, automation platforms) where token bills explode. Enterprises facing $ millions in monthly agentic bills will migrate to the provider delivering better economics at scale. OpenAI’s combination of strong models (o-series reasoning) + lowest TCO positions it to erode Anthropic’s enterprise share, especially as agentic becomes the dominant token consumer. Cheaper tokens expand the total addressable market dramatically. This feeds the data/model improvement loop, justifying further capex. AMD benefits from proven scale pulling in more customers (Meta, Oracle, Microsfot, Amazon, Softbank, TensorWave, LumaAI ... already aligned on Helios). Conclusion: Dr. Lisa Su has been laser focused on inference economics since at least 2022–2023, repeatedly emphasizing that the real battleground for AI scalability would be TCO, power efficiency (TDP), and ultimately tokens per dollar and per watt not just raw training FLOPS. While many viewed inference as a secondary, commoditized workload, Dr. Su architected AMD’s roadmap around rack-scale systems optimized for high-volume, sustained inference that would dominate as models matured and usage exploded. Helios represents the culmination of that multi-year bet: a fully integrated, open platform designed precisely for the economics of massive token throughput. This deep, strategic partnership with OpenAI starting with the 1GW Helios deployment in H2 2026 and scaling to 6GW, is the embodiment of that shared vision. Both companies foresaw a future where agentic AI models evolve to become extraordinarily token-hungry: autonomous agents executing complex, iterative workflows with planning, tool use, verification loops, and long-context reasoning. These workloads can consume 100x+ more tokens per task than traditional chat or single-turn generation, driving exponential demand as capabilities improve and enterprises deploy them at scale. By owning and optimizing this massive Helios capacity at GW scale, OpenAI achieves inference costs as low as $0.0003–$0.0005 per million tokens. This structural cost advantage allows OpenAI to absorb the coming token explosion profitably, dramatically lower effective pricing for enterprises, and win high-volume agentic workloads from higher-cost competitors like Anthropic. What was once a prohibitive monthly token bill becomes an affordable accelerator for productivity and innovation. The OpenAI-AMD alliance validates Dr. Su’s prescient strategy and turns the Agentic flywheel into reality: Collapsing inference costs → explosive token consumption → richer data and better models → accelerate greater demand. This partnership doesn’t just address today’s economics, it positions both leaders at the center of the infrastructure buildout that will power AI’s next decade. By delivering the lowest inference economics at scale, OpenAI not only solves enterprise bill pain but gains a decisive weapon to win share from higher-cost rivals like Anthropic. And that is why OpenAI and $META will deploy EPYC Dense Rack Not Financial Advice! DYOR! Research Purpose Only!

Mike

84,951 views • 1 month ago

Tiny drone hits invisible mode by twisting faster than eye can detect | Omar Kardoudi, New Atlas Engineers at Northwestern University have built a drone that vanishes without camouflage or transparent panels. Its trick is spinning so fast that your eyes simply give up trying to focus, a stealth edge that could turn surveillance into something almost invisible. The aircraft, nicknamed Phantom Twist, rotates up to 25 times per second, a rate that outpaces how quickly our visual system can process sharp detail. Instead of true invisibility, the drone dissolves into a faint, ghostly blur that blends into whatever is behind it. The work, led by associate professor Michael Rubenstein, was presented on July 16 at the Robotics: Science and Systems 2026 conference in Sydney, Australia, under the title Computational Design of a Low-Visibility UAV Using Human-Aligned Perceptual Metric. "Most efforts to hide drones focus on making them look like their surroundings," says Rubenstein. "Instead, we asked whether we could design the drone itself around the way humans perceive motion. This idea of low visibility through persistent motion is something few people have explored." That distinction matters because drones are increasingly used to watch wildlife, check aging infrastructure, or survey wetlands, but their mere presence changes the behavior of whatever they're observing. Birds scatter, animals flee, people act differently. A drone that's hard to spot could do the same job without that side effect. Prior attempts at motion-based concealment offer useful context here. The Northwestern paper points to an earlier project nicknamed the Boomerang Drone, covered in a 2006 New York Times Magazine piece, which tried a similar high-speed rotation trick but couldn't spin fast enough to fully exploit the blur effect, leaving it largely visible. The paper authors also trace the broader idea of active concealment back to the "Yehudi light," a counter-illumination project developed by the National Defense Research Committee in 1944 to hide Allied sea-search aircraft from enemy view. The Phantom Twist itself takes a very different shape from those earlier attempts. Rather than a typical quadcopter with four separate rotors, it runs on a single motor and a single propeller, with the propeller spinning one way while the rest of the drone's body spins the opposite way. "For a typical quadrotor drone, the propellers are spinning, but the robot is stationary," Rubenstein explains. "So, you still see its body. For our drone, the whole thing is rotating, so there are no stationary parts." To reach that layout, the team's computer model generated roughly 20,000 possible drone configurations capable of stable flight, then used artificial intelligence and optimization algorithms to repeatedly rearrange the motor, propeller, circuit board, counterweight, and batteries. Each design was simulated spinning mid-flight and overlaid on 100 real-world backgrounds, then scored by a perceptual model built to mimic human vision, where a lower score meant better camouflage. The 500 best-scoring designs were run through the optimizer again to squeeze out further gains before a final version was built. Emma Alexander, an assistant professor of computer science and one of the study's co-authors, explains the underlying physics. "The human eye takes time to accumulate signals, roughly analogous to the exposure time of a camera," she says. "When an object spins quickly, we perceive it as blurring out and losing distinct features. Because this new drone is almost entirely transparent, its few opaque components are visually averaged with the background for an overall appearance of a slight haze." According to the paper's visibility metric, the finished drone is about 10 times harder to spot than a standard quadcopter. But the spinning trick has real limits that make this drone far from being completely unnoticeable. The propeller still makes an audible whir that gives the drone away even when the eye can't, and its support wires and rods remain partly visible. The paper's authors suggest future versions could lean on more transparent materials and quieter propulsion, edging the drone ever closer to true – and somewhat scary – invisibility. After all, the same trick making a drone less impactful on wildlife could just as easily help it sneak around for reasons that aren't so friendly.

Owen Gregorian

24,217 views • 16 days ago

HERMES AGENT HAS 5 SYSTEMS RUNNING UNDER THE HOOD. UNDERSTAND THEM AND YOU USE THE AGENT 10X BETTER. In this video Alejandro AO 🤗 explained: 1. THE AGENT LOOP every message triggers the same cycle: → you send a message → Hermes builds context (SOUL.md + memory.md + user.md + skills + tools + message history) → sends everything to the LLM → LLM decides: call a tool or respond → if tool call: execute, return result, loop back → if response: deliver to you → after response: memory update (agent checks if anything is worth remembering, writes to memory.md or user.md) this loop is why Hermes gets better over time. the memory update after every response means the agent learns from every conversation. 2. CONTEXT ASSEMBLY what the LLM sees on every turn: → SOUL.md (your agent's personality and rules) → memory.md (facts the agent learned over time) → user.md (facts about you, auto-updated) → AGENTS.md and .hermes.md (project context files) → skill descriptions (loaded on demand) → tool schemas (available actions) → message history (current conversation) if SOUL.md is empty, Hermes falls back to a default system prompt. write your own SOUL.md and the agent becomes yours, not generic. CONTEXT COMPRESSION: conversations hit context limits. Hermes handles this at two checkpoints: preflight: before each turn. if conversation exceeds 50% of context window, compression fires. older messages get summarized. last 20 messages stay intact (protect_last_n). gateway auto-compression: between turns. fires at 85%. more aggressive. prevents API errors before the agent even starts processing your message. after compression, a new session lineage ID is generated. the agent can trace back to the original conversation through SQLite. three things break prompt cache: switching models mid-session, changing memory files, or changing context files. 3. THE GATEWAY the system that keeps Hermes reachable on 27+ messaging platforms. an async loop runs continuously. listens for incoming messages from Telegram, Discord, Slack, WhatsApp, email, SMS, and every other adapter. when a message arrives: → gateway identifies which session it belongs to → queries SQLite for the full message history (session ID = platform prefix + chat ID) → builds the context from scratch → sends everything into the agent loop → delivers the response back to the platform the gateway also runs the session manager. when you send a message while the agent is busy: → default: queued for next turn → /steer: injected without interrupting → /interrupt: stops current work without the gateway, Hermes is a CLI tool. with the gateway, Hermes is an always-on agent you reach from your phone. 4. MEMORY (THREE LAYERS) LAYER 1 — MARKDOWN FILES SOUL.md (identity), memory.md (learned facts), user.md (facts about you). injected into context after the system prompt. updated by the agent after every response. LAYER 2 — SQLITE full transcripts of every session stored locally. FTS5 full-text search across all past conversations. session lineage tracking across compressions. the agent can recall what you discussed weeks ago using /recall or session search. LAYER 3 — EXTERNAL PROVIDERS (optional) 8 supported providers: Mem0, SuperMemory, Honcho, Zep, and more. each works differently (semantic search, LLM extraction, similarity matching). queried after the first message in each session. the agent processes your topic first, then checks external memory for related context from past conversations. not enabled by default. enable for significantly better long-term recall. 5. CRON ENGINE a loop inside the gateway ticks every 60 seconds. each tick checks ~/.hermes/cron/jobs.json for scheduled tasks. if a job is due: → fresh session (no chat history, no memory pollution) → execute the prompt with assigned tools → store the run output as markdown in ~/.hermes/cron/output/[job-id]/ → deliver result to your home messaging platform cron does NOT use the send_message tool. delivery happens at the system level, not the agent level. a cron session cannot create more cron jobs. prevents runaway loops. WHY THIS MATTERS: the agent loop teaches it. the context assembly focuses it. the gateway reaches it. the memory remembers it. the cron engine automates it. five systems. one agent. understanding how they connect changes how you configure every level. full 15 levels breakdown in the article 👇

YanXbt

51,258 views • 1 month ago

I finally finished my Rust version of Mario Zechner's (Mario Zechner) excellent Pi Agent, which I made with his blessing and which is called pi_agent_rust. You can get it here: If you're not familiar with Pi, it's a minimalist and extensible agent harness (similar to Claude Code and Codex) and, among other uses, serves as the core agent harness inside the OpenClaw project. I say my Rust "version" instead of "port" because it's really quite different in how it's implemented for it to be called a port. Arguably, the incremental functionality in the implementation was more complex than the rest of the project combined. Still, it provides the same features and functionality as the original, and is proven to be compatible with hundreds of popular extensions to Pi (the conformance harness shows 224 out of 224 extensions working perfectly). But the way it's architected has some major changes. Pi Agent relies on node or bun to provide access to the filesystem and for various other tasks, and that is also how Pi's extension system works. I decided early on that I didn't want to do things that way. Instead, I wanted to integrate that functionality directly into the binary itself; that is, to provide equivalent functionality for everything that would normally be provided by node/bun in the original. I did this for several reasons: one, it's a lot more performant in terms of footprint and latency. On realistic end-to-end large-session workloads (not toy microbenchmarks), pi_agent_rust is now: - 4.95x faster than legacy Node and 2.80x faster than legacy Bun at 1mm-token session scale - 4.32x faster than legacy Node and 2.14x faster than legacy Bun at 5mm-token session scale - ~8x to ~13x lower RSS memory footprint in those same scenarios But the other reason is security and control: by handling everything internally in an end-to-end way, we can do all sorts of clever things to harden the system against insecure or malicious extensions. Those extensions no longer have direct access to the ambient filesystem: they now need to go through pi_agent_rust, and we can analyze extensions carefully before ever running them and also block things that look suspicious at runtime. In practice that means explicit capability-gated hostcalls, with policy/risk/quota enforcement and runtime telemetry/auditability. In order to do all this, I had to effectively build the missing runtime substrate from scratch in Rust, not just translate TypeScript syntax: - define and implement a typed hostcall ABI for extension->host interactions - build native Rust connectors for tool/exec/http/session/ui/events instead of ambient Node/Bun access - implement a compatibility/shim layer so real-world Pi extensions still behave correctly - add capability policy evaluation, runtime risk scoring, per-extension quotas, and audit telemetry on the execution path - wire the whole thing through structured concurrency (asupersync) so cancellation/lifetimes are deterministic and failure handling is explicit - build a conformance + benchmark harness large enough to validate behavior/perf across hundreds of extensions and realistic long-session workloads This was a full re-architecture of the execution model while preserving the Pi workflow and extension ecosystem. And indeed, this aspect of it dwarfs the entire rest of the project in size and complexity. To put hard numbers on that: the extension/runtime/security subsystem alone is now about 86.5k lines of Rust across src/extensions.rs (~48.1k), src/extensions_js.rs (~23.4k), src/extension_dispatcher.rs (~13.4k), and src/extension_index.rs (~1.7k), with roughly 2.5k callable units in just those files. For context, the original Pi coding-agent production code is about 27.4k lines total. So this one subsystem by itself is roughly 3.2x the size of the original harness, which is why calling this a “port” would seriously undersell what had to be built. And on top of that, pi_agent_rust introduces a bunch of genuinely new capabilities beyond the legacy harness, not just a faster core: - Security and enforcement are materially stronger at runtime: capability-gated hostcalls with explicit policy profiles (safe/balanced/permissive), per-extension trust lifecycle (pending -> acknowledged -> trusted -> killed), explicit kill-switch operations, and audited state transitions. - Shell execution mediation is deterministic and argument-aware: rule/feature-based risk scoring plus heredoc AST inspection (dcg_rule_hit, dcg_heredoc_hit) before spawn, instead of relying on coarse deny patterns. - Containment and forensics are first-class: tamper-evident runtime risk ledger tooling (verify/replay/calibrate), unified incident evidence bundles, and forced-compat controls that let you contain issues without disabling the whole extension system. - The extension runtime architecture is native: JS extensions run in embedded QuickJS with typed hostcall boundaries and Rust-native connectors for tool/exec/http/session/ui/events, plus compatibility shims for real-world legacy extensions. - Runtime behavior under load is explicitly engineered: deterministic hostcall reactor mesh, fast-lane vs compat-lane routing, and warm-isolate prewarm handoff for more predictable throughput and latency under contention. - Long-session reliability is upgraded: JSONL v3 sessions with indexed sidecar acceleration and optional SQLite-backed sessions, plus operational controls via --session-durability, --no-migrations, and migrate. - Provider and auth coverage are broader and more operationally explicit: native Anthropic/OpenAI (Chat + Responses)/Gemini/Cohere/Azure/Bedrock/Vertex/Copilot/GitLab plus large OpenAI-compatible routing; pi --list-providers currently shows 90 providers with aliases and required auth env keys. - Auth is not just API keys: OAuth (Anthropic/OpenAI Codex/Gemini CLI/Antigravity/Kimi/Copilot/GitLab plus extension-defined OAuth), AWS credential chains (Bedrock), service-key exchange (SAP AI Core), and bearer-token flows. - Operator tooling is stronger: pi doctor supports scoped checks (config, dirs, auth, shell, sessions, extensions), machine-readable output (--format json|markdown), and safe auto-remediation (--fix). - Extension/package lifecycle workflows are built in: install, remove, update, update-index, search, info, and list. I want to thank Mario for making a great harness and for not telling me to get lost when I asked him if he was OK with me porting it to Rust. I may give him a hard time in jest about not going "full clanker," but that doesn't mean that I don't respect his work a huge amount. PS: There could still be bugs. If you find some, please let me know in GitHub Issues and I'll fix them same day. There's always a tradeoff between perfect and getting stuff out the door and I felt like it was time to release this.

Jeffrey Emanuel

116,929 views • 5 months ago