正在加载视频...

视频加载失败

Simplify data collection with low cost,hand-held parallel jaw gripper #Pika AgileX Robotics Pika Gripper: ±1.5mm accuracy, dual-camera system Pika Sense: Lightweight (550g) Support ROS1/2 Perfect for researchers and developers driving innovation. #Robotics #EmbodiedAI #ROS

17,709 次观看 • 1 年前 •via X (Twitter)

2 条评论

MR MR 的头像
MR MR1 年前

cool

Pwnage 的头像
Pwnage2 年前

🚨PREORDER OPEN: StormBreaker Max CF The easy choice, the better choice. Get yours: Carbon Fiber Composite only 43g 5% larger in size Pwnage XERO Sensor, custom PixArt 3950 Adjustable Sensor Placement 8kHz polling rate, Wireless & Wired and much more!

相关视频

A gripper with two thumbs! 🐨 Handheld data collection is how a lot of manipulation datasets get built now. Someone picks up a gripper-shaped device, does the task by hand, and the recording becomes training data. Voila! 🤌🏼 The catch is that these devices are usually shaped to match whatever robot gripper already exists. The human ends up working around the robot's morphology, which costs both ergonomics and the quality of the demonstration being recorded. The RAI Institute team designed The Koala Gripper. It includes two devices with full parity: a capture device a person holds, and a robotic device that executes the learned policy. Anything you can do with one, the other reproduces. The shared feature set gets updated as the design iterates, so neither side quietly compromises the other. The morphology is where it gets interesting. Two independently controlled underactuated fingers are opposed by a single pivoting "dual thumb", which is where the koala name comes from. → Each finger is a 9-bar linkage with one actuated and one underactuated degree of freedom, with the spring sitting proximal to the phalanges to drive preshaping → Fingernails on the distal link for precise pinches, contoured pads for secure wrap grasps → The handheld trigger traces a natural fingertip arc instead of a straight line, lengthening the stroke and giving the operator better mechanical advantage → On the robot side, a frameless motor with an integrated high-pitch ballscrew keeps reflected inertia low, so the fingers stay backdrivable at an effective mass of tens of grams The dual thumb is what buys tool stabilising, singulating one object out of a pile, and tabletop pinches. Congrats to Zubin Kremer Guha and the RAI Institute team behind it 👏🏼 🔗 Project page: ~~ ♻️ Join the weekly robotics newsletter, and never miss any news →

Lukas Ziegler

19,233 次观看 • 13 天前

I’m thrilled to announce that we just released GraspGen, a multi-year project we have been cooking at NVIDIA Robotics 🚀 GraspGen: A Diffusion-Based Framework for 6-DOF Grasping Grasping is a foundational challenge in robotics 🤖 — whether for industrial picking or general-purpose humanoids. VLA + real data collection is all the rage now but is expensive and scales poorly for this task. For every new gripper and/or scene, you’ll have to recollect the dataset in this paradigm for the best perf. 💡Key Idea: Since grasping is such a well-defined task in simulation - why can’t we just scale synthetic data generation and train a generative model for grasping? By embracing modularity and standardized grasp formats, we can make this a turnkey technology that works zero-shot for multiple settings. GraspGen is a modular framework for diffusion-based 6-DOF grasp generation that scales across embodiment types, observability conditions, clutter, task complexity. Key Features: ✅ Multi-embodiment support: suction, parallel-jaw, and multi-fingered grippers ✅ Generalization to partial + complete 3D point clouds ✅ Generalization to single-objects + cluttered scenes ✅ Modular design uses other robotics modules and foundation models (SAM2, cuRobo, FoundationStereo, FoundationPose). This allows GraspGen to focus on only one thing - grasp generation ✅ Training recipe: grasp discriminator is trained with On-Generator data from the diffusion model - so that it learns to correct the mistakes (if any) of the diffusion generator ✅ Real-time performance (~20 Hz) before any GPU acceleration; low memory footprint 📊 Results: • SOTA on the FetchBench [Han et al. CoRL 2024] benchmark • Zero-shot sim-to-real transfer on unknown objects and cluttered scenes • Dataset of 53M simulated grasps across 8K objects from Objaverse 📄 arXiv: 🌐 Website: 💻 Code: A huge thank you to everyone involved in this journey — excited to see what the community builds on top of it! Joint work with Clemens Eppner , Balakumar Sundaralingam , Yu-Wei, Jun Yamada Wentao Yuan and other collaborators #robotics #diffusionmodels #physicalAI #simtoreal

Adithya Murali

24,309 次观看 • 1 年前

Once we started to work with large global retailers, we needed a better way to scale this process. Ideally, the staff at the store could do this themselves — rather than us flying our team across the world — and then we could lower the cost and timelines. So we built a self-serve version of our survey app, with a tutorial mode designed for beginners. Over time, we collected millions of data points, and so we were able to develop an algorithm which would auto-correct mistakes. In other words, if the surveyor accidentally placed their ground-truth location in the wrong place on the map, we could use our algorithms to detect it, and correct it. So now we have WiFi, and with and our efforts on producing a high quality survey, we have the best WiFi positioning available. With WiFi on its own, it’s achieving 3 meter accuracy. This is a great foundation to build on. WiFi + Motion data To refine this down to 1-meter accuracy, we realised that we could combine WiFi with the same technology behind self-driving cars and robotics: a motion system called SLAM (Simultaneous Localization and Mapping). SLAM uses the accelerometer, gyroscope and camera system to understand precise device motion. Imagine a car driving through a tunnel, using the motion since its last GPS ping to keep location accurate until it comes out the other side. On a phone, this technology is very reliable, and measures device motion with high precision. But SLAM is measuring motion within its own coordinate space, it’s not aligned with the real world. SLAM tracks the user’s relative motion, like “moved forward 2 meters, then turned left”, but does “forward” mean “north”, or some other direction? It’s not calibrated, so it could mean any location, any direction. We can’t rely on the compass to help us out with this, because phone compasses are notoriously incorrect — everyone knows the frustration of being sent the wrong way down a street. So our job was to align this motion data with the triangulation data we were receiving from WiFi. We designed an algorithm that could simulate every possibility, filter the unlikely scenarios, and hone in your location, using WiFi as an anchor. So WiFi gives us the initial blue dot, SLAM gives us motion, and as the user starts walking and we receive more data, our algorithms can refine location accuracy down to a consistent 1-meter accuracy. We’ve tested these algorithms in many locations, on hundreds of hours of ground-truth data:

Andrew Hart

91,047 次观看 • 1 年前

NEW ROBOT BENCHMARK: If your robot can do Origami, it can do anything! Called The Robotic Origami Challenge, it is a dexterous-manipulation competition and benchmark held at IROS 2026, organized by 13 co-organizers with the Nippon Origami Association as judge and task curator -> pretty cool to have them on board imho. The evaluation consists of single task: a traditional Japanese paper airplane, in exactly six folds, from a 15×15 cm sheet of ≥60 gsm paper, on a competition-supplied standardized rig (bimanual arms + Sharpa Hands), both remotely and on-site. Teams bring policies, not hardware. An "Origami Grand Master" declares pass/fail on crease accuracy, structural fidelity, symmetry and paper integrity. Among passes, faster folds rank higher, with a 10-minute-per-attempt ceiling and flight explicitly irrelevant to the score. -> I find it interesting how they chose to evaluate the task. Quality is a binary pass or fail, therefore speed becomes the only thing graded here. Speed is currently the bottleneck in dexterous manipulation though, so this choice makes sense. I wonder whether there could be finer ways to grade the qaulity of such a creative task though. When it comes to data, registered teams get 500+ teleoperation episodes (six camera streams, 65-D joint state/action, 10-fingertip 6-axis tactile), an NVIDIA Isaac Sim environment with thin-shell paper physics (plastic creasing + fold memory), digital twins of every partner hand, and a remote eval lab (upload a policy, queue an eval, get scored). Still, I think it is a great dexterity benchmark the field badly needs, it supplies the hardware, an outside human judges, and the pass criterion is externally defined -> all three degrees of freedom are checked! Neutral measurement layer, here we go! The task is engineered to be un-gameable and to isolate pure dexterity. A known figure, exactly six folds, judged on creases, with flight explicitly declared irrelevant (the latter makes sense to me). Therefore, this underlines the goal to focus on dexterity, not task-selection or other strategies. I really like origami as an ideal controlled dexterity task: deterministic goal, deformable medium, sequential, bimanual, precision-bound. I am just not quite satisfied again by the binary pass or fail, I think quality of execution could be finely graded! But again, I understand this is not the goal yet. Also interesting to see the Sharpa Hands as de facto standard for everyone. Total land-grab that anoints Sharpa as the reference dexterous hand, also featured in Google's Gemini Robotics 2. By providing the hardware, the benchmark measures software while quietly making "good on Sharpa" the definition of good, and Sharpa gets real world data and feedback as a bonus. That's smart, the data flywheel starts spinning. The provided dataset is the richest tactile-manipulation corpus I have seen yet: 10-fingertip 6-axis tactile, plus plastic creasing and fold memory. 500+ teleop episodes with six camera streams, 65-D joint state/action, and ten fingertip 6-axis tactile sensors. The force/tactile channel are parts of the the benchmark's data, this is the first time I see this. Credits where it's due: organizers include Yang Gao, Noriaki Hirose, Steve Xie, Chris Paxton, Jiafei Duan, Michael Cho - Rbt/Acc, Michael Yuan, Haoquan Fang, and others.

Léo

28,718 次观看 • 1 个月前

CHINA JUST SOLVED THE PROBLEM THAT'S BEEN BREAKING ROBOT AI FOR A DECADE. and the fix wasn't a smarter model. for years, every robot AI failure got the same diagnosis. the model isn't smart enough. so everyone scaled intelligence. bigger models. more parameters. better reasoning. AGIBOT asked a different question: what if the reasoning was never the problem? there's a gap that runs through every traditional robot AI system. reasoning on one side & motor commands on the other. the brain decides but the body executes something different, because thinking and moving were never actually connected. GO-2 fixes this by reasoning INSIDE the action space, not above it. before moving, it runs a complete mental simulation of every step - like a basketball player mentally tracing the arc of a shot before releasing the ball. watch the demo and you'll see exactly what this means. the robot works through a task queue autonomously. classify toiletries. upright the drink bottle. place headphones in the leather box. mid-execution, a new instruction drops: "my phone's missing. help me find it." it doesn't pause. doesn't reset. it processes the new task and keeps moving. that's not a scripted sequence. that's real-time instruction following on top of an active task queue. that one architectural change is where the numbers come from. > #1 on LIBERO across Spatial, Object, Goal, and Long tasks → 98.5% average success > 86.6% zero-shot accuracy in active disturbance environments > 47.4 on VLABench → best-in-class on objects and textures it's never seen before > 82.9% success trained on simulation only, tested on real hardware sim-to-real is the graveyard of robotics research. models trained in simulation collapse the moment they touch the real world. 82.9% means that graveyard just got a lot smaller. it holds because of how GO-2 trains. deliberately fed imperfect reasoning conditions, then trained to execute robustly anyway. not a researcher assumption. a design decision from a team that ships hardware and knows exactly what breaks. then there's the infrastructure layer. Genie Studio. fleet-wide data collection. cloud training. online post-training in live environments. 10x improvement in training efficiency. task startup reduced to minutes. 2-4x better success rates with 50%+ less data. the model gets smarter every time a robot fails in the field. this isn't a benchmark story. it's a compounding moat. dual CVPR 2026 + ACL 2026 acceptance. computer vision AND natural language processing. top conferences. simultaneously. that doesn't happen with incremental research. the US-China robotics race has been framed as a compute race. a model quality race. it was always an execution race. the robot that wins won't be the smartest one in the lab. it'll be the most reliable one on the floor. full breakdown: is execution reliability the real bottleneck, or are we still underestimating how far reasoning needs to go?

Shruti

18,622 次观看 • 5 个月前

A team tested Pi0, Pi0 Fast, Gr00t, and ACT on real robot arms in manufacturing tasks. (🔖 Bookmark this for later!) The task was precise: place thin rectangular frames from a messy stack into a holder. The team fine-tuned each model on 100 real trajectories and compared training time, inference speed, motion quality, and success rates. ⬇️ Here’s a breakdown of what they found Pi0 (Original) ✅ Strongest overall performance in precise pick-and-place ✅ High success rate even in edge cases ✅ Longest training time (~11 hours, ~$30 per run) ✅ Inference time of 80 ms causes short pauses between actions Despite delays, it handles complex scenarios well… solid for high-precision tasks, but slow to train. Gr00t ✅ Trains fast (~2 hours, ~$5 per run) ✅ Performs almost as well as Pi0 on large-object tasks ✅ Struggles with fine precision; random movement in some trials ✅ More training didn’t fix jitter or random offsets Best suited for tasks where exact precision isn’t critical. Not ready for manufacturing-grade accuracy without more tuning. Pi0 Fast ✅ Promised faster training, but results were underwhelming ✅ Training at 6 hours still showed low success rates ✅ Inference was slower than expected ✅ Not reliable for generalizing even slightly new tasks Currently too unstable for real-world deployment. Doesn’t live up to the “Fast” name yet. ACT (Baseline) ✅ 200MB model—lightweight, but limited ✅ Struggles with stacked objects or ambiguous scenes ✅ Success rates around 70% in best-case setups ✅ Can’t match newer models on precision or generalization Still a solid baseline, but clearly a generation behind in robustness. 🚨 Extra Notes All newer models share a common issue: •Inference takes longer than a frame (80 ms vs 33 ms), so robots “pause” between chunks. •This results in jittery movements, but not a dealbreaker unless tasks are time-sensitive. Language-conditioned tasks also fell short: after training on two labeled tasks, the model couldn’t generalize to a third unseen combination using only text prompts. ✅ The good news? These models adapt well to new robot arms with quick fine-tuning. ❌ The bad news? There’s still no plug-and-play solution for improving performance after deployment. Reinforcement learning or DAgger-style data collection during real-world operation may be the next big step, something many teams in robotics are actively working on.

Ilir Aliu

21,844 次观看 • 1 年前

Drew Bredvick compressed Vercel's sales team from 20 people to 2. And I think it's one of the best case studies in the history of AI and GTM. the problem: Sales development doesn't compound. Headcount does. Every additional SDR brings another salary, another ramp period, another personal definition of what "qualified" actually means. the solution: Drew built an AI agent that evaluates every inbound lead: researching the company, scoring intent, and routing only the credible opportunities to the sales team. Everything else is handled automatically. the result: Now two people, focused exclusively on edge cases and high-touch accounts handle the entire sales operation at the $10B company. Andddd the previous team wasn't let go. They were moved into "higher-value work" within the company. here's the play in six steps: 1. Shadow your best performer 2. Pull 90 days of historical data 3. Iterate until 95% agreement 4. Run in parallel with people 5. Get co-sign 6. Hand your top dogs the controller 1. shadow your best performer Sit next to your best SDR for a full day and document every decision: when they qualify, when they disqualify, every signal they check, every button they click. Drew found the real qualification criteria was not in process docs. Reps were checking LinkedIn profiles, scanning websites for tech stack indicators, and pattern-matching on how leads found Vercel. None of it was documented. 2. pull 90 days of historical data Export 90 days of contact form submissions with outcomes attached. Did they close? Ghost? Become a $500K whale? 3. iterate until 95% agreement Open any code editor with AI built in. Drop your CSV into a new project and start a conversation: "Look at this lead data. I'm going to give you a prompt to evaluate leads. Tell me if each one is qualified or not." Run this prompt against a batch. Compare the agent's calls to what actually happened—not what humans decided, but whether the lead converted. Find disagreements. Fix the prompt. Repeat. You're aiming for 95%+ agreement with historical outcomes. starter prompt: You are a lead qualification agent. For each lead, analyze the following signals and provide your reasoning BEFORE your decision: Company signals: website quality, tech stack, company stage, employee count Intent signals: how they found us, what they asked for, urgency indicators Fit signals: ICP match, use case alignment, budget indicators Structure your response as: REASONING: [Your analysis of each signal category] CONFIDENCE: [High/Medium/Low] DECISION: [Qualified/Not Qualified] NEXT ACTION: [Route to sales / Auto-respond / Request more info] Be conservative. You naturally want to qualify leads to make humans happy. Resist that urge. A false positive wastes sales time. A false negative just means we follow up later. 4. run in parallel with people Once the prompt works with historical data, prove it works live with your sales team: Here's what to track: - agreement rate: Agent vs. human decisions. - accuracy rate: Agent vs. actual outcomes. - processing time: Lead received → decision made. - confidence distribution: How often the agent is certain vs. uncertain - error log: When it got it wrong, and why. 5. get co-sign Drew started chatting with individual contributors. He got them to validate that the agent was making good calls. Then he partnered closely with the leader of the SDR team. He then let leaders of the sales team tweak qualification criteria, the leaders of the marketing team adjust scoring weights, and let leaders of the ops team define routing rules. b/c when leadership builds alongside you, they stop being gatekeepers and start being advocates. 6. hand your top dog the controller Flip the switch!! The agent processes every lead, makes a qualification decision, and even drafts the response. But a person reviews before anything goes out and has more time to check the genuinely f******* tough and weird cases. The system recommends; humans decide. The same loop should work In other parts of the org too: customer support triage, contract review, expense approvals, and even content moderation. Full playbook below w/ prompts and Drew's handholding. 👇

Alex Lieberman

81,801 次观看 • 7 个月前

What a ride! Made by using GPT Image 2 + Seedance 2.0 on Fish Creative Prompt reference_handling: "Image generation strictly for driver facial and wardrobe styling reference only — calm, composed features silver wristwatch on left wrist Image strictly for sports car styling and cabin reference only — low, wide Italian wedge-shaped body + bright yellow paint + strongly geometric body lines + hexagonal front intake + Y-shaped LED headlights + gloss black multi-spoke wheels + black leather cabin with orange stitching + left-hand-drive cabin (driver seat on left) + across this sequence the driver-side window (left side of car) is rolled down only for the cockpit reveal shot, all other windows remain as-is throughout. Image strictly for spire architectural geometry, Dubai downtown skyline, and warm hazy midday atmosphere reference only — tapered glass-and-steel spire that widens progressively toward the base + dense glass high-rise skyline below + wide multi-lane boulevard. Do not reproduce any specific camera angle, composition, or caption elements from the reference images" style: "REAL AERIAL + AUTOMOTIVE CINEMATOGRAPHY PLATE — not CGI rendering, not game-engine rendering, not an animated/illustrated look." visual_feel: "Strong overhead midday light + warm hazy atmosphere softening the horizon. Color strictly natural and true-to-life — not oversaturated, not faded, not washed out. Continuous soft haze and atmospheric layering from spire tip down to street level. Every camera move, whether aerial or ground-tracking, strictly gimbal-level smooth — absolutely no handheld feel, no shake, no roll or tilt, even during the FPV-paced dive segment or the accelerating side-pass segments. 16:9 frame + no stylized film-grain treatment, aiming for genuine cinematography texture across every shot" duration: "30 seconds (8-shot sequence)" aspect_ratio: "16:9" character_modeling: driver_suited_woman: base: " appearance and wardrobe strictly per Image generation reference. Present in the car throughout the sequence, but the face is strictly clearly visible only during the 0:20–0:21 cockpit reveal shot — in every other shot the face is strictly not shown or not resolvable, whether by camera position, angle, or framing" wardrobe: silver wristwatch on left wrist . complete, with no wrinkling, misalignment, or missing pieces throughout" presence: "In shots where the driver is not the subject (0:01–0:19, 0:22–0:30), the driver strictly remains seated in the left-side driving position, present but strictly not resolved facially due to camera side, distance, or angle. During the 0:20–0:21 cockpit reveal shot only, the face and posture are strictly fully clear — visibility achieved via a right-side cockpit camera position looking across the cabin, with natural light and open sightline entering through the already-lowered driver-side window (left side of car) forming an angled depth-of-view channel — strictly NOT via looking directly through a window immediately adjacent to the camera" sports_car_yellow: identity: "Low, wide Italian wedge-shaped supercar + bright yellow paint + strongly geometric body + hexagonal front intake + Y-shaped LED headlights + gloss black multi-spoke wheels + black leather cabin with orange stitching + left-hand-drive cabin, driver seat on left — appearance strictly per image2 reference. Strictly only this one car appears across all 8 shots + doors strictly closed throughout + strictly only the driver-side window (left side of car) is rolled down, and only for the 0:20–0:21 cockpit shot + all other windows strictly remain closed/unchanged throughout + left-hand-drive position strictly remains on the left side of the car in every shot — not mirrored, not flipped, regardless of which side the camera is on" physics: "In every shot showing the car in motion, tires strictly show real, visible load deformation on turns + suspension strictly compresses and rebounds continuously with road surface undulation + body strictly shows slight roll and pitch matching cornering or acceleration — strictly not a rigid-glide, zero-deformation model feel. The car strictly stays lane-centered along the boulevard's true path — strictly no crossing lines, no drifting, no hugging the curb" environment_spire_and_skyline: setting_lock: "Tapered glass-and-steel spire structure — body progressively widens toward the base + spire tip is the sequence's starting point + below is Dubai downtown's dense glass high-rise skyline and wide multi-lane boulevard + warm hazy midday light, architectural geometry, skyline, and atmosphere strictly per image3 reference. Shot 1 (0:01–0:12) strictly covers the spire exterior and the high-altitude-to-street transition. Shots 2–8 (0:12–0:30) strictly take place entirely at street level, on or beside the boulevard, with the skyline visible as background context only" cinematic_storyboard: shot_1_spire_descent_0_01_0_12: camera: "Continuous aerial dive — the same FPV-paced descent arc as the master establishing move: 0:01–0:02 camera approaches and briefly hovers directly above the spire tip, gimbal strictly steady, no roll or tilt. 0:02–0:12 camera descends along one continuous curved arc, vertical speed component smoothly decaying while horizontal speed component smoothly increasing, no perceptible docking point or speed jump. As the arc resolves near street level, the camera settles into a position that spotlights the yellow car on the right side of frame — car held in the right third of the composition as the shot closes. Lens strictly 35–50mm cine prime throughout, no zoom, no digital zoom, straight architectural lines keep true perspective." action: "Spire tip and city grid fill the frame at the open, then dissolve into recognizable streets and blocks as the dive continues. The yellow car appears as a small point mid-descent and grows continuously larger, coming to rest spotlighted on the right side of frame by 0:12, driving forward along the boulevard, wheels rotating forward, no reverse." lighting: "Strong overhead midday light at the spire tip with long shadows; as the descent continues, glass-facade and ground reflections shift continuously and smoothly with the changing angle — no abrupt lens-flare flicker." vfx: "Ground detail and color progressively sharpen through the descent, no sudden clarity jump. Ground shadows strictly limited to the car's own cast shadow — no operator or camera-rig shadow anywhere in frame." sfx: "High-altitude wind roar at the open, fading continuously into rising engine sound and city ambience as the car comes into view. No music, no voiceover, no captions." shot_2_side_pass_0_12_0_15: camera: "Hard cut to a static lateral profile position — camera holds a fixed side-view framing of the car, 35–50mm cine prime, gimbal-locked, no handheld sway. As the car accelerates, the camera lets it pull ahead and overtake past the camera's position, exiting frame screen-right." action: "Car holds briefly in profile, then accelerates hard — visible squat of the rear suspension under acceleration, tires gripping without slip, body pitching slightly rearward under load — before overtaking and leaving frame past the camera." lighting: "Even natural daylight, sun still overhead-midday, car's yellow paint reading true and saturated against the boulevard背景, no flat frontal wash." vfx: "Real suspension compression and rebound as the car surges forward. No motion blur artifacts beyond natural shutter response; no CG float." sfx: "Engine note rises sharply with the acceleration, a clean Doppler pass as the car overtakes the camera position; no music." shot_3_center_mirror_0_15_0_17: camera: "Hard cut to an interior point-of-view through the car's center rear-view mirror — camera framed as if looking through the mirror glass from just behind/above the driver's eyeline, mirror surface visibly framing the receding view." action: "Through the mirror, the boulevard and the Dubai skyline recede behind the car as it continues forward at speed; slight natural mirror-glass vignette at the frame edge." lighting: "Cabin interior in soft ambient light, mirror glass reflecting the bright exterior daylight and skyline without glare washing out the reflected image." vfx: "Mirror reflection stays optically clean and stable — no double image, no warping; road and skyline motion in the reflection reads as physically continuous with forward travel." sfx: "Muffled cabin-interior tone to engine and wind noise (heard as if from inside the car); no music, no dialogue." shot_4_front_view_0_18_0_19: camera: "Hard cut to a nose-on front view of the car — camera positioned directly ahead on the boulevard, framing the grille, headlights, and hood centered in frame, lens 35–50mm cine prime, static or minimal push, gimbal-steady." action: "Car approaches head-on at a steady, controlled speed, Y-shaped LED headlights and hexagonal intake clearly readable, wheels visibly rotating forward." lighting: "Overhead midday sun catches the hood and windshield with clean natural highlights, no artificial front-fill look." vfx: "Subtle heat-haze shimmer off the hot asphalt ahead of the car for realism; no CG gloss on the paint." sfx: "Engine sound growing louder as the car closes distance toward camera; no music." shot_5_cockpit_reveal_0_20_0_21: camera: "Hard cut to a right-side cockpit angle — camera positioned to the right-front of the car, sightline crossing through the windshield and, aided by the already-lowered driver-side (left) window, resolving the driver clearly inside the cabin. This is the sequence's only driver-reveal shot." action: "Driver's face and posture are fully visible — one hand resting lightly on the wheel, eyes on the road ahead, expression calm and composed, natural unstiff posture. Car maintains the same forward direction and steady speed with no lens or vehicle behavior change during the shot." lighting: "Even natural daylight, light falling cleanly across the yellow paint, windshield, and driver's face; windshield reflection kept light enough not to obscure visibility." vfx: "Windshield glass stays transparent and reflection-light, no glare occlusion of the driver." sfx: "Steady engine hum plus faint city ambience; strictly no dramatic sound swell or music entering at the reveal moment." shot_6_straight_road_rear_3_4_0_22_0_24: camera: "Hard cut to a rear-bumper 3/4 angle — camera positioned low and behind, off to one side, framing the car driving away down a straight stretch of boulevard, lens 35–50mm cine prime, gimbal-smooth tracking that holds pace with the car." action: "Car drives straight down the boulevard at a steady cruising speed, lane-centered, taillights and rear three-quarter bodywork clearly visible, wheels rotating forward, no drift or lane departure." lighting: "Overhead midday sun, road surface and rear bodywork evenly lit, skyline visible in soft haze in the background." vfx: "Light heat-shimmer off the straight road surface; ground shadow strictly limited to the car's own cast shadow." sfx: "Steady, sustained engine tone at cruising speed plus ambient city sound; no music." shot_7_side_pass_0_25_0_28: camera: "Hard cut back to a static lateral profile position, mirroring shot 2's setup — camera holds the side view of the car, gimbal-locked, no handheld sway." action: "Car holds briefly in profile again, then accelerates a second time and overtakes past the camera, exiting frame — same physical behavior as the first side-pass (visible suspension squat, tire grip, body pitch)." lighting: "Consistent overhead midday daylight, same natural exposure as shot 2 for continuity." vfx: "Real suspension compression and rebound under acceleration; clean natural motion, no CG float." sfx: "Engine note rising sharply into the pass, clean Doppler effect as the car overtakes camera; no music." shot_8_static_3_4_close_0_29_0_30: camera: "Hard cut to a static, locked-off 3/4 angle — camera fixed in position, no movement, gimbal-perfect stillness, framing a 3/4 view of the boulevard as the car enters and exits frame to close the sequence." action: "Car drives through the static frame at a steady speed and exits, completing the sequence; wheels rotating forward, no reverse, no lingering hold after exit." lighting: "Same consistent overhead midday daylight and natural color grade as the rest of the sequence, no shift in exposure for the closing shot." vfx: "Ground shadow strictly limited to the car's own cast shadow; no operator or rig shadow in frame." sfx: "Engine sound passing through and fading as the car exits frame; no music, no voiceover, no captions at any point in the closing shot." production_notes: multi_shot_cut_lock: "This sequence is strictly 8 distinct shots joined by hard cuts at the following points: 0:12, 0:15, 0:18, 0:20, 0:22, 0:25, 0:29 — strictly no smooth transitions, no cross-dissolves, no whip-pans between shots, no morphing between camera setups. Each shot is a clean cut to a new fixed or moving camera setup as specified; only shot 1 (0:01–0:12) is itself one continuous unbroken aerial move. Across every cut, the following must remain continuous: the car's identity and paint color, the boulevard geography and skyline, the direction of travel, the lighting direction and quality, and the absence of music/dialogue/captions." gimbal_stabilization_lock: "Every shot, aerial or ground-based, strictly holds gimbal-level smoothness — absolutely no handheld shake, no roll, no tilt, no high-frequency jitter, in any of the 8 shots, including the accelerating side-pass shots." optics_lock: "Every ground/tracking shot strictly uses a 35–50mm cinema prime feel (Sony Cine prime lens character), unchanged within each shot — strictly no zoom in, no zoom out, no digital zoom of any kind in any shot. Straight lines of buildings and roads strictly retain true perspective — strictly no wide-angle distortion, no fisheye curvature." forward_motion_lock: "In every shot, the car strictly drives forward — nose pointed in the direction of travel except where explicitly framed nose-on toward camera (shot 4), and never reversing. Wheels strictly rotate continuously in the direction of travel, reverse rotation strictly forbidden in any shot." vehicle_structural_identity_lock: "Strictly only this one car appears across all 8 shots — a second car of the same or different model is strictly forbidden. Doors strictly remain closed in every shot. Left-hand-drive layout strictly remains unchanged across all 8 shots — driver's seat strictly on the left side of the car in every shot, never mirrored or flipped regardless of camera side. The driver-side (left) window is strictly rolled down only during shot 5 (cockpit reveal, 0:20–0:21); in every other shot all windows strictly remain closed/unchanged." driver_appearance_window_lock: "The driver's face is strictly clearly visible only during shot 5 (0:20–0:21) — in shots 1–4 and 6–8 the face is strictly not resolved, whether due to distance, angle, motion, or framing. During shot 5, visibility is strictly achieved via the right-side cockpit camera angle through the windshield, aided by the already-lowered driver-side window — strictly NOT via a window directly facing the camera." no_operator_shadow_lock: "In every shot, the ground strictly shows only the car's own cast shadow — strictly no human-shaped shadow, camera-operator silhouette, photographer's figure, or rig shadow cast anywhere in any of the 8 shots. The aerial shot (shot 1) is strictly pure drone photography with no physical rig or ground-crew trace; the ground shots (2–8) are strictly framed with no visible operator, crew, or equipment in frame." audio_lock: "Sound throughout the sequence is strictly authentic diegetic sound only — wind, engine, and city ambience, shifting naturally shot to shot — strictly no background music or score of any kind, including any hidden musical layer, in any of the 8 shots + strictly no voiceover or dialogue of any kind + strictly no captions or on-screen text of any kind at any point." critical_constraint: "8 hard-cut shots across 30 seconds, cut points strictly at 0:12, 0:15, 0:18, 0:20, 0:22, 0:25, 0:29 — strictly no dissolves or blended transitions. Shot 1 is one continuous uncut aerial descent from the spire tip to street level, ending with the car spotlighted screen-right. Shots 2 and 7 are matching static side-profile setups where the car accelerates and passes the camera. Shot 5 is the sequence's only driver-face reveal, via the right-side cockpit angle through the windshield with the driver-side window down — strictly not visible in any other shot. The car is strictly the same single yellow LHD supercar throughout, doors closed except for the driver-side window during shot 5, always driving forward, wheels never reversing. All camera work is strictly gimbal-smooth with a 35–50mm cine-prime feel, no zoom, no distortion, no handheld shake in any shot. Ground shadows throughout strictly show only the car's own shadow, no operator or rig trace. Audio is strictly diegetic only — no music, no voiceover, no captions, throughout the entire 30 seconds." avoid: "dissolves, cross-fades, morphs, or whip-pans between shots — cuts must be hard cuts only, any shot other than shot 5 showing the driver's face clearly, the driver-side window rolled down in any shot other than shot 5, the passenger-side or any other window rolled down at any point, a second car appearing in any shot, car doors opening in any shot, the car reversing or wheels rotating backward in any shot, mirrored or flipped left-hand-drive layout, driver's seat appearing on the right side of the car, human-shaped shadow, camera-operator silhouette, or rig shadow in any of the 8 shots, camera shake, handheld feel, roll, tilt, or high-frequency jitter in any shot, zoom in, zoom out, digital zoom, wide-angle distortion, fisheye distortion, or curved building/road lines in any shot, smooth continuous single-take treatment of the whole 30 seconds (the sequence is strictly multi-shot with hard cuts, not one unbroken take beyond shot 1), background music, score, melody, hidden musical layer, voiceover, dialogue, narration, captions, or on-screen text at any point, CGI-rendered look, game-engine feel, plasticky car-paint gloss, static hovering with no sense of gravity" animation_style: "Shot 1 plays out as one continuous real-time aerial move; shots 2–8 are each a distinct, clean hard-cut setup, every shot internally in real time with no speed ramping — the sense of pace across the sequence comes from the editing rhythm of the cuts, not from slow motion or time manipulation within any single shot"

Aaliya

11,907 次观看 • 17 天前

$AMD| The FOMO to buy AMD Chips is NOW 🧵 Not Financial Advice! DYOR! Research Purpose Only! The Inference Queen is the biggest winner in Agentic AI where all other CPUs are struggling to compete with a 2yr old EPYC Turin and EPYC Venice is in mass production phase. AMD stresses deployability today on standard x86 platforms (no proprietary architectures required), full software compatibility, and open standards. This positions Venice + Helios as a practical, high-density alternative to competing solutions while underscoring that agentic AI shifts the balance toward CPU-rich racks alongside GPUs, and most importantly, lowering the cost of token to accelerate adoption and innovation. Context: The Wall Street Journal yesterday came out with an article that OpenAI is condiering drasstically lowering the token prices to win more customers from Anthropic. The narrative "they" are trying to exacerbate the current AI selloff won't last long. This is a fundamental misunderstanding of what is going on, or what I already discussed for months and years. Followers and Subscribers already knew this for years, that this day would come, where token cost will bcome the central discussion among enterprises as there is no such thing as unlimited budget or Tokenmaxxing when they use $NVDA chips or In-house Hyperscalers chips. I will link various threads if you are interested in understanding the full picture from supply chain to recent TSMC Rapid 2nm expansion up to 12 Fabs total by 2027/2028. Hyperscalers and AI natives effectively have no choice but to buy more AMD system for Agentic AI as leadership in economical, power-aware, high-volume internal + agentic use. However, due to supply constraints where Supply is far behind Demand, this makes multi-vendor reality along with in-house chips drive faster industry progress, lower overall costs, and better sustainability. NVIDIA’s Vera Rubin cannot compete with a 2 years old EPYC Turin, but AMD under Dr. Lisa Su has engineered the lowest cost-per-million-tokens, highly competitive energy-efficient solutions, and superior CPU orchestration for agentic AI at scale with Helios. Dr. Su has championed this shift since at least 2023, foreseeing the rise of agentic workflows that demand far more orchestration, parallel agents, and balanced compute well before the industry fully embraced it. Her long-term vision of AI moving from simple prompts to always on, multi-agent systems has driven AMD’s investments in high-core EPYC CPUs and integrated rack-scale solutions, perfectly positioning the company for today’s realities. The OpenAI-AMD 1GW Helios deployment (starting H2 2026) represents a pivotal vertical integration move that directly supercharges the inference economics. This isn't incremental; it's a structural shift toward ownership of massive, optimized rack-scale capacity, enabling the lowest token costs and triggering the enterprise adoption flywheel. We need to be honest, $AMD is the only company that made a big bet on Inference since the day Chatgpt became sensational where $NVDA and others were betting big on Training. At the end of the day, Token bill from Anthropic has to obey economics. Meaning the bills rise, companies have to get more out of it to justify the cost. It cannot be an unlimited inference budget, and it has to show up on efficiency, profitability and operating leverage. 1. Tokenomics After you understand this, you will understand why Citi cited Anthropic is likely to sign a deal with $AMD along with Hyperscalers, AI Labs, Sovereign AI like Softbank 5GW in France and many other countries. However, OpenAI and $META are now wanting faster deployment, and they are AMD shareholders now, they have prioritized allocation. Anthropic and Hyperscalers just cannot compete when Helios Rack lower token cost to$0.0003–$0.0005 per million tokens at GW scale. Cost to build 1GW data center 1GW Helios Rack full build is estimated $30-$35B 1GW Rubin Rack full build is estimated $45-$55B Inference (Cost per Million Tokens) ~$NVDA B200 / HGX: ~$0.02–$0.08 on optimized workloads (FP4/MXFP4, speculative decoding). Significant improvement over Hopper but still premium-priced. GB200 NVL72 rack-scale: $0.05–$0.25+ ~$AMD Helios Racks: $0.0003-$0.0005 per M tokens, dramatically lower than NVIDIA equivalents in owned infra. MI355X node-level: Up to 40% more tokens per dollar vs. competing solutions ( B200), driven by higher memory capacity (up to 288GB+ HBM), strong bandwidth, and lower acquisition costs. Training ~$NVDA Rubin Rack is estimated $0.7-$1.2/M Tokens ~$AMD Helios Rack is estimated $0.65-$1.0/M Tokens Now, OpenAI, META and Hyperscalers can lower Inference cost even further with $AMD EPYC Venice "dense rack" or Agentic AI Rack. AMD published a detailed technical blog emphasizing that the future of agentic AI autonomous, multi-step AI systems requiring heavy orchestration, databases, caching, APIs, and control planes demands massive CPU-dense rack-scale infrastructure, not just GPUs. The catalyst prominently positions their upcoming 6th Gen EPYC "Venice" processors as the key enabler for next-generation dense racks, delivering leadership throughput under real-world power, cooling, and density constraints. ~EPYC Venice (Zen 6 architecture, up to 256 cores / 512 threads per socket) is projected to deliver exceptional rack-level performance. In AMD’s modeled 100 kW rack comparisons, Venice-powered systems are expected to achieve ~3.30x the throughput of NVIDIA’s Vera (88-core Olympus) baseline across a broad mix of agentic-supporting workloads. ~This builds on current-generation 5th Gen EPYC "Turin" (up to 192 cores), which already delivers ~2.37x rack throughput vs. Vera and ~1.6x vs. Intel’s Xeon 6980P (128 cores). ~ Liquid-cooled Turin deployments already support >27,000 CPU cores per rack today. Venice is architected to push this beyond 36,000 cores in the same rack class, dramatically increasing concurrent agent capacity and overall infrastructure efficiency. 2. Ownership vs renting compute from Hyperscalers matter to OpenAI and only owning $AMD chips can meaningfully lower token cost for enterprises. ~Eliminates cloud overhead: No provider margins, utilization buffers, or egress fees. Direct control over power contracts, cooling, scheduling, and orchestration at dedicated facilities. ~Helios optimizations at GW scale: Rack-level density (1.4+ exaFLOPS FP8 per rack), high HBM4 bandwidth, EPYC orchestration for agentic workloads, and superior TCO/TDP. AMD's long-standing focus on tokens per dollar/watt shines here 20-40%+ efficiency edges in inference-heavy scenarios. ~At 1GW+ optimized deployment, inference hits $0.0003–$0.0005 per million tokens (community/analyst models tied to Helios metrics). This is dramatically lower than typical rented/cloud equivalents, especially for high-volume output tokens in agentic flows. High token bills today, enterprises running heavy agentic/coding/analysis workloads can face $50-100M+/month at current API rates (flagship models $5-30+/M output, scaled to massive volumes). Post-Helios compression, same volume will drop to $10-15M/month (or better) via lower underlying costs passed through as pricing flexibility, volume tiers, caching, or batch discounts. ROI thresholds collapse. More companies greenlight pilots → production → massive scaling. Agentic AI (autonomous workflows) multiplies token demand exponentially, but affordability removes the friction. OpenAI gains flexibility, Unlike more cloud-dependent rivals (Anthropic), they can lower effective pricing, offer aggressive enterprise bundles, or absorb volume without margin destruction directly tackling "high token bill" complaints while maintaining profitability as usage explodes. 3. Agentic AI Models shifted CPU:GPU Ratio to 1:1 toward 3-5:1 with Explosively Token-Hungry Workloads Agentic AI (autonomous, multi-step agents with planning, tool use, iteration, and self-correction) is fundamentally more compute and token intensive than conversational or single-turn generative AI. Agentic AI. autonomous, multi-step workflows with orchestration, tool use, parallel agents, data movement, and enterprise integration has dramatically increased the importance of strong host CPUs alongside GPUs. This shifts the CPU-to-GPU ratio higher and makes balanced systems critical toward 1:1 to 5:1 as enterprises testing more than 5-10 agents. AMD EPYC Venice excels ~Leadership core density (up to 256 Zen 6 cores per socket) for running many agents in parallel, orchestration layers, and high-throughput control-plane tasks. ~Superior performance-per-core and power efficiency ( up to 2.1x higher perf/core and 2.26x better SPECpower vs. NVIDIA Grace in benchmarks). ~Tight integration in Helios: One Venice CPU + multiple MI450 GPUs per node, enabling efficient data feeding to GPUs ("zero-copy"), parallel execution, and full rack utilization for complex agentic loops. Hyperscalers (Meta, Microsoft, Amazon, Google, Softbank) and AI natives (OpenAI, Anthropic...) are adopting high-core EPYC at scale specifically for these agentic demands, as CPUs now handle a larger share of non-model work (orchestration, policy enforcement, tool calls). This complements AMD’s lower-cost GPUs for overall TCO wins. ~Agents often generate 10–100x+ more tokens per task due to iterative reasoning chains, multiple tool calls, verification loops, and long-context orchestration. ~Goldman Sachs forecasts token consumption multiplying 24x by 2030 (to 120 quadrillion tokens/month) largely driven by agentic adoption in consumer and enterprise. ~Enterprise data shows agent-pattern workloads growing at 680% annualized rates, projected to surpass conversational AI in token volume by Q3 2026. ~Daily enterprise agent token consumption is already in the billions, with complex workflows (coding, workflows, analysis) amplifying this dramatically. 4. Competitive Edge: Winning Customers from Anthropic Anthropic’s Claude models (especially Opus/Sonnet) excel in complex reasoning and agentic coding, commanding premium positioning. However, their higher underlying costs (heavier reliance on third-party cloud with margins) limit pricing flexibility compared to OpenAI’s owned Helios capacity. Anthropic is on track to generate $10.9 billion in Q2 revenue. The company expects to achieve its first-ever quarterly adjusted operating profit of $559 million. However, sustaining full-year profitability remains challenging due to immense computing and model training costs The truth is, Anthropic has no choice but to buy as much $AMD chips as possible if they want to compete with OpenAI or get investors attention. This 5% adjusted operating profit to revenue ratio is just pathetic. Current pricing dynamics (2026): OpenAI already undercuts on many tiers ( flagship output tokens significantly cheaper than equivalent Claude Opus). Nano/mini models offer 5–10x advantages for volume work. Anthropic holds edges in long-context flat pricing and certain reasoning quality. OpenAI after Helios Rack Ownership, At $0.0003–$0.0005/M effective costs, OpenAI gains massive headroom to: ~Aggressively discount high-volume agentic tiers or bundles. ~Offer “unlimited” enterprise plans or usage-based models that Anthropic struggles to match without margin erosion. ~Target cost-sensitive, high-throughput agent deployments (dev tools, automation platforms) where token bills explode. Enterprises facing $ millions in monthly agentic bills will migrate to the provider delivering better economics at scale. OpenAI’s combination of strong models (o-series reasoning) + lowest TCO positions it to erode Anthropic’s enterprise share, especially as agentic becomes the dominant token consumer. Cheaper tokens expand the total addressable market dramatically. This feeds the data/model improvement loop, justifying further capex. AMD benefits from proven scale pulling in more customers (Meta, Oracle, Microsfot, Amazon, Softbank, TensorWave, LumaAI ... already aligned on Helios). Conclusion: Dr. Lisa Su has been laser focused on inference economics since at least 2022–2023, repeatedly emphasizing that the real battleground for AI scalability would be TCO, power efficiency (TDP), and ultimately tokens per dollar and per watt not just raw training FLOPS. While many viewed inference as a secondary, commoditized workload, Dr. Su architected AMD’s roadmap around rack-scale systems optimized for high-volume, sustained inference that would dominate as models matured and usage exploded. Helios represents the culmination of that multi-year bet: a fully integrated, open platform designed precisely for the economics of massive token throughput. This deep, strategic partnership with OpenAI starting with the 1GW Helios deployment in H2 2026 and scaling to 6GW, is the embodiment of that shared vision. Both companies foresaw a future where agentic AI models evolve to become extraordinarily token-hungry: autonomous agents executing complex, iterative workflows with planning, tool use, verification loops, and long-context reasoning. These workloads can consume 100x+ more tokens per task than traditional chat or single-turn generation, driving exponential demand as capabilities improve and enterprises deploy them at scale. By owning and optimizing this massive Helios capacity at GW scale, OpenAI achieves inference costs as low as $0.0003–$0.0005 per million tokens. This structural cost advantage allows OpenAI to absorb the coming token explosion profitably, dramatically lower effective pricing for enterprises, and win high-volume agentic workloads from higher-cost competitors like Anthropic. What was once a prohibitive monthly token bill becomes an affordable accelerator for productivity and innovation. The OpenAI-AMD alliance validates Dr. Su’s prescient strategy and turns the Agentic flywheel into reality: Collapsing inference costs → explosive token consumption → richer data and better models → accelerate greater demand. This partnership doesn’t just address today’s economics, it positions both leaders at the center of the infrastructure buildout that will power AI’s next decade. By delivering the lowest inference economics at scale, OpenAI not only solves enterprise bill pain but gains a decisive weapon to win share from higher-cost rivals like Anthropic. And that is why OpenAI and $META will deploy EPYC Dense Rack Not Financial Advice! DYOR! Research Purpose Only!

Mike

84,951 次观看 • 3 个月前

What if I told you ripple:native just moved closer to a financial universe doing $17.5 TRILLION in FX and interest-rate derivatives every single day? I’m not talking about some random prediction. I’m talking about BIS Working Paper No. 1374. This is going to be a long read, because the headline barely scratches the surface. Four of the five authors work at the Bank for International Settlements, and instead of only mentioning XRP Ledger in theory, the researchers actually built, tested and published an open-source XRPL-based prototype. That distinction matters. This is a research implementation, not a production BIS deployment. But the technical choice itself is what caught me. The researchers needed a public blockchain that could help prove official economic and financial data had not been altered. They chose XRP Ledger. And they explained why: low fees, fast finality, developer resources and existing research around its consensus system. This wasn’t somebody adding an XRP logo to a presentation. They built the gateway. They created XRPL transactions. They used institutional anchoring wallets. They put cryptographic proofs inside transaction memos. They linked publisher identities to XRPL addresses. They retrieved those transactions again during verification. Then they measured how the system performed. Median publication latency came in around 3–5 seconds. Verification took around 1–2 seconds. That is where my brain immediately went beyond the headline. Because what exactly were they trying to verify? The kind of information the entire financial system runs on. -Inflation. -GDP. -Interest rates. -Banking statistics. -Debt information. -Financial-stability data. -Regulatory reporting. Imagine a central bank publishes an inflation number. Today that number gets copied everywhere. -Websites. -News terminals. -Databases. -Screenshots. -AI models. -Trading systems. Once it spreads across the internet, how does another machine independently prove that the number it received is exactly what the institution originally published? That is the problem BIS researchers were attacking. Their model creates a cryptographic fingerprint of the official dataset. Individual statistical series can receive fingerprints too. Those hashes are combined through a Merkle tree. A final Merkle root gets anchored to XRPL. The underlying economic data do not need to be dumped onto the blockchain. XRPL simply keeps the proof. Think of it like this: The official institution publishes the document. XRPL holds the tamper-proof receipt. Someone changes even one part of the underlying file? The cryptographic fingerprint changes. Now a bank, regulator, investor, trading engine or AI agent can check: Is this the original data? Has it been changed? Did it really come from the institution claiming to publish it? And that second part is where this paper gets even more serious. The BIS prototype combines the data proof with a W3C Verifiable Credential for the publisher. The publisher’s cryptographic identity is connected to an XRPL address. The paper even uses the format: did:xrpl: So you are not only verifying the information. You are verifying who published it. Now picture a financial world where machines can check both automatically. A central bank publishes CPI. A model receives it. Before touching money, the software checks XRPL. Correct file. Correct publisher. No alteration. Then it acts. That sounds simple until you realize what financial markets actually do with official data. -Rates move. -Currencies move. -Bond prices move. -Derivatives reprice. -Collateral requirements change. -Loans reset. -Inflation-linked instruments adjust. -Portfolio risk changes. And this is where BIS Working Paper 1374 stops being a boring statistics paper for me. Because the authors themselves discuss putting verified information beside digital financial assets. They specifically mention: -CBDCs -stablecoins -tokenized deposits -derivatives. That one section changes the entire way I look at this. The vision is not simply: “Put a hash on a blockchain.” It becomes: verified economic information + digital money + tokenized assets + automated execution. Now remember what Ripple has been building around XRPL. -Multi-Purpose Tokens. -Credentials. -Permissioned Domains. -Permissioned DEX infrastructure. -Confidential Transfers. -Stablecoins. -Institutional lending. -Tokenized collateral. -FX. -Onchain credit. And Ripple has repeatedly positioned XRP across payments, liquidity and credit. Now put those pieces beside what the BIS researchers are exploring. An official institution needs an identity. XRPL can represent identity and credentials. A regulated participant needs permission to enter a market. XRPL is building permissioned infrastructure. A bond needs trustworthy economic information. The BIS prototype shows one way that information can be authenticated through XRPL. A financial asset needs a digital representation. XRPL is being built for tokenization. A transaction needs money. Stablecoins and tokenized deposits can provide the cash side. Then all those different assets need liquidity. That is where ripple:native becomes much more interesting to me. But before getting there, look at the scale surrounding BIS itself. The BIS does not process the world’s $9.6 trillion of daily FX transactions. It measures that market through its Triennial Central Bank Survey. That distinction matters. According to the numbers in the context here: global OTC FX turnover = $9.6 TRILLION every day. Then add: OTC interest-rate derivatives turnover = $7.9 TRILLION every day. Together: $17.5 TRILLION per day. Just the FX number annualized across roughly 250 trading days comes to around: $2.4 QUADRILLION per year. That is the financial universe BIS research sits over. -Currencies. -Banks. -Central banks. -FX swaps. -Rates. -Derivatives. -Cross-border capital. -Collateral. -Dollar funding. And researchers inside that institution just chose XRP Ledger for an actual technical prototype. That is why I keep telling people not to reduce this to transaction fees. Yes, the worked example uses an XRPL Payment transaction. Yes, the reference cost is only: 10 drops = 0.00001 XRP. Yes, transaction fees on XRPL are destroyed. So if this kind of anchoring eventually ran on mainnet, publishing data itself would consume XRP. But that is not the part that gets me excited. The fee is intentionally tiny. The much bigger question is: What happens when verified information starts triggering financial activity on the same broader infrastructure? The paper itself talks about: inflation-linked products perpetual futures tokenized financial instruments derivative settlement interest payments automated compliance and even: automated monetary-policy applications. Now we are talking about information causing money to move. Imagine an inflation-linked bond. The government publishes inflation. That release gets cryptographically anchored. The bond checks the proof. The CPI number is verified. The contract adjusts what is owed. Digital cash settles the payment. No one has to manually copy a number from a website into another system. No one has to blindly trust a third-party data feed. The financial instrument can verify the economic input itself. That is the idea I keep coming back to: self-verifying finance. And the researchers even discuss using the XRPL EVM-compatible sidechain for more advanced applications where data verification and programmable financial execution exist in the same broader ecosystem. They mention: access controls, permissioning, automated compliance, multisignature requirements, oracle integration, programmable validation. Now connect that with Ripple’s institutional roadmap. Credentials can prove who a participant is. Permissioned Domains can define who belongs inside a regulated environment. Tokenized assets can represent financial instruments. RLUSD can represent digital dollar liquidity. Lending can make those assets productive. XRP can provide native network resources and, where economically useful, liquidity between fragmented assets. That is a very different picture of XRPL than the one people were arguing about years ago. It is not simply: “Can XRP send a payment quickly?” The question becomes: Can XRPL sit underneath parts of a machine-readable financial system? And Working Paper 1374 just gave that question much more weight for me. There is another section that barely gets discussed. The architecture is not limited to one data publisher. The researchers designed a multi-publisher system. Different institutions can create their own Merkle roots. Those roots can be combined into one larger super-root. One XRPL transaction can anchor that shared proof. Yet each publisher remains independently accountable for its own data. Now imagine the participants. Central Bank A. Central Bank B. Regulator C. Statistical Office D. International Organization E. One public verification system. Different publishers. Independent cryptographic accountability. That begins to resemble infrastructure for cross-border public-sector data exchange. And the paper’s own conclusion talks about trustworthy exchange among: national statistical offices central banks international organizations. Then look at who already uses the statistical standard the paper builds around. SDMX is sponsored by institutions including: BIS European Central Bank Eurostat International Monetary Fund OECD United Nations World Bank Group International Labour Organization. That does not mean those institutions are adopting XRPL. But it tells you something important about the design philosophy. The researchers did not create a blockchain system that requires the existing financial world to throw everything away. They designed it to sit underneath an existing institutional standard. That matters a lot. Because the easiest technology to adopt is often the technology that does not force everyone to rebuild from zero. Existing systems can continue publishing. XRPL can provide the cryptographic proof underneath. Then comes BIS Open Tech. The paper says the open-source reference implementation is being released as a prototype through BIS Open Tech and the SDMX community. That means other institutions can inspect it. Reuse it. Modify it. Build on it. This is how technical ideas can spread inside serious institutions. Not through hype. Through code. Documentation. Standards. Reuse. That is the kind of adoption path I pay attention to. Then there is the AI angle. This is where the whole thesis becomes almost unfairly interesting. The authors explicitly discuss AI agents. An AI system receives economic information. Instead of blindly trusting what it scraped from somewhere, it can ask: Is this data authentic? It checks the XRPL proof. Valid? Continue. Invalid? Do nothing. Now compare that with what Ripple launched in June 2026: the XRPL AI Starter Kit, designed around autonomous agents making payments with XRP and RLUSD. Two completely separate directions suddenly sit beside each other. BIS research: AI verifies information through XRPL. Ripple ecosystem: AI moves value through XRPL. Now imagine both ideas eventually meeting. An agent receives official inflation data. It verifies the release cryptographically. It recalculates risk. It reprices a bond. It adjusts collateral. It changes an FX position. It executes a payment. It settles in RLUSD. It routes through XRP where XRP is the best available liquidity path. That is machine-native finance. And now go back to the scale. The BIS 2025 Triennial Survey says: $9.6T/day FX. The dollar appears on one side of 89% of FX trades. The euro is involved in 28.9%. The Japanese yen in 16.8%. FX swaps alone are around $4T every day. Then another $7.9T/day exists in OTC interest-rate derivatives turnover. Think about what happens if only part of those markets becomes tokenized. Digital USD deposits. Digital EUR deposits. Tokenized JPY. RLUSD. CBDCs. Tokenized Treasuries. Interest-rate derivatives. FX derivatives. Collateral. Money-market instruments. The first problem is getting the assets onchain. The second is verifying the information those assets depend on. The third is moving liquidity between all the different forms of value. This BIS paper attacks the second problem using XRPL. Ripple has spent years attacking the first and third. That is why the combination gets my attention. And you do not need XRPL to capture the whole market for the numbers to become enormous. For scale only: 0.1% of $9.6T daily FX turnover = $9.6B per day. 1% = $96B per day. Again, that is not a forecast. It shows what even tiny percentages mean when the underlying market is measured in trillions every day. And that is only FX. It does not include the additional $7.9T/day of interest-rate derivatives turnover BIS measures. This is where the XRP liquidity thesis changes from a crypto argument into a market-structure argument. Suppose the future has hundreds of tokenized currencies and financial products. Every possible pair cannot maintain perfect direct liquidity. USD token / EUR token. EUR token / JPY token. JPY token / RLUSD. RLUSD / Treasury token. Treasury token / derivative. Derivative / deposit token. The combinations explode. A common intermediate asset becomes useful whenever routing through it provides a better market. That is where XRP’s role becomes interesting. Not replacing the dollar. Not replacing the euro. Not replacing CBDCs. Not replacing bank deposits. Connecting liquidity between them when that route makes economic sense. Now imagine the system is automated. No trader needs to shout: “Use XRP.” Software looks at: price, spread, depth, settlement, availability. If the XRP path wins, the software uses XRP. That is the outcome I care about. Machine-selected liquidity. And if those transactions grow large enough, the XRP market itself has to change. Institutional market makers need inventory. Liquidity providers need inventory. Prime brokers need financing capacity. Order books need deeper capital. Large transactions need to clear without huge price impact. That is where the price thesis becomes different from retail speculation. If XRP ever helps support institutional flows inside markets measured in trillions per day, the relevant question is not: “How many retail holders bought today?” It becomes: How much dollar liquidity does the XRP market need to represent? That is an entirely different valuation conversation. There is one more thing I think people are missing. BIS Working Paper 1374 does not only talk about SDMX statistics. The researchers say the same architecture can extend to: XBRL regulatory filings FINREP COREP and other forms of structured official information. Now imagine banks submitting regulatory reports that receive immutable XRPL proofs. The bank cannot quietly change an old filing later. The regulator can verify the exact version. Auditors can verify it. Another authority can verify it. AI software can consume it. One system can prove both: who submitted the data and whether it changed. That gives XRPL a potential role far beyond payments. It starts touching the information layer of finance. And this is why the line “BIS used XRP Ledger” actually undersells the paper. What happened is more specific. Researchers inside BIS took a real institutional problem. They selected XRPL. They built a working implementation. They measured performance. They published the code direction. Then they explored how authenticated data could coexist with: CBDCs, stablecoins, tokenized deposits, derivatives, AI agents, automated financial instruments. That is what I am bullish on. Not a logo. Not a rumor. Not a screenshot. Technical work. And when I look at the direction Ripple is independently pushing XRPL, the overlap is hard for me to ignore. Trusted identities. Verified information. Regulated participants. Tokenized assets. Digital money. Automated execution. Credit. Collateral. FX. Liquidity. AI. Put together, the long-term architecture can look like this: Official institutions publish information. XRPL anchors the proof. Banks and regulators verify it. AI consumes it. Tokenized instruments use it. Stablecoins and tokenized deposits provide cash. Institutional markets execute trades. XRP supplies native network resources and can supply cross-asset liquidity where the route makes sense. That is not simply a faster payment network. That starts looking like part of a digital financial operating system. And then remember where this conversation is happening. Inside the research world of the institution that measures: $9.6 trillion of FX turnover every day plus $7.9 trillion of interest-rate derivatives turnover every day. A combined: $17.5 TRILLION DAILY. No, that is not XRPL volume. No, BIS does not process those trades. The significance is that BIS researchers just tested XRP Ledger while working inside the institutional world surrounding markets of that size. That is the fact. And now I’m asking the question that matters to me as an ripple:native holder: What happens if XRPL earns even a small role inside the tokenized version of that financial system? Because 0.1% of a trillion-dollar market is not small. And this market is not one trillion. It is trillions every single day. That is why Working Paper 1374 changed the scale of the conversation for me. For years, people asked whether XRP could become part of the future financial system. Now researchers inside the BIS have taken XRP Ledger, built institutional infrastructure on it, and explicitly discussed a future combining trusted information with digital money and programmable financial assets. We are still at the prototype stage. But for me, the direction is the real story. The next financial system will need trusted data, tokenized assets, automated execution and deep liquidity. XRPL is now showing up in all four conversations. And XRP sits natively underneath the network where those pieces can eventually meet. $17.5T a day. Now look at your ripple:native bag again. Enough?

X Finance Bull

68,256 次观看 • 9 天前

$NVDA $GFS NVIDIA’s reported agreement to acquire Groq for $20B in cash (per CNBC, amplified via Reuters and other wire coverage) represents a materially different strategic posture than NVIDIA’s prior M&A pattern, given both the headline size (largest reported NVIDIA acquisition to date) and the unusual carve-out that Groq’s early-stage cloud business would not be included. Public reporting indicates the information originated from Alex Davis, CEO of Disruptive (lead investor in Groq’s latest financing), and that neither NVIDIA nor Groq had issued an immediate confirmation at the time of publication. The same reporting frames the transaction as coming together quickly, only months after Groq raised $750M at a ~$6.9B valuation, and highlights Groq’s positioning as a high-performance inference chip vendor founded by ex-Google TPU engineers. Groq is best understood as a vertically integrated inference acceleration company whose core asset is an application-specific processor optimized for deterministic, low-latency execution of transformer-style workloads, paired with a compiler-led software stack and a distribution layer (GroqCloud) designed to reduce developer friction via OpenAI-compatible APIs and integrations. Groq brands its architecture as a Language Processing Unit (LPU) and consistently emphasizes that the design target is inference, not training. The company’s own architecture description centers on 1-core execution, large on-chip SRAM used as primary storage (explicitly not cache), a custom compiler that statically schedules compute and communication, and direct chip-to-chip connectivity intended to coordinate multi-chip execution without relying on conventional caching hierarchies or dynamic runtime scheduling. The technical premise is a deliberate inversion of the conventional GPU approach. GPUs deliver throughput via massively parallel, multi-core execution with dynamic scheduling, complex memory hierarchies, and heavy reliance on off-chip HBM bandwidth and sophisticated runtime/kernel optimization. Groq instead argues that inference bottlenecks are driven by latency variance (tail latency), synchronization overhead, and memory access unpredictability inherent in dynamically scheduled, cache-heavy architectures, particularly when workloads are latency sensitive and batch sizes cannot be inflated. Groq’s solution is to move “control” into the compiler: the full execution graph and inter-chip communication schedule are computed ahead of time down to clock-cycle granularity, with deterministic execution designed to reduce run-to-run variance. In Groq’s framing, the removal of caches, reorder buffers, speculative execution overhead, and other sources of contention enables predictable latency and high utilization without per-model kernel engineering typical of GPU tuning cycles. A critical nuance is that Groq’s determinism is not merely a software claim; it is tightly coupled to architectural constraints and system design choices that trade flexibility for predictability. Third-party technical commentary indicates Groq’s chip uses a fully deterministic VLIW-style approach with minimal buffering, no external memory, and heavy dependence on sharding models across many chips because on-chip SRAM capacity is limited. SemiAnalysis describes a ~725 mm^2 die on GlobalFoundries 14nm with ~230MB of SRAM and notes that “no useful models” fit on a single chip, forcing multi-chip partitioning for modern LLMs and driving a system-level design where networking and compilation are first-class scheduling problems rather than ancillary infrastructure. This is consistent with Groq’s own messaging that tensor parallelism across chips is a primary design goal, enabled by large on-chip SRAM and compile-time coordination of compute plus interconnect. The on-chip SRAM emphasis is central to Groq’s latency story and also its most constraining trade-off. Groq claims on-chip SRAM bandwidth “upwards of 80 TB/s” and contrasts that with off-chip HBM bandwidth “about 8 TB/s,” asserting a potential 10x advantage from bandwidth plus reduced trips across chip-to-memory boundaries. While these comparisons are marketing-oriented and depend on workload specifics, the architectural implication is clear: Groq prioritizes ultra-fast local weight/activation access and then scales capacity by adding chips, not by attaching large off-chip memory pools. This design can reduce latency for sequential inference layers and minimize unpredictable stalls, but it pushes complexity into partitioning strategy, interconnect topology, and compiler scheduling, and it increases the number of chips needed for very large parameter counts and large KV-cache footprints. Groq also highlights numeric formats and compiler-driven precision management as a performance lever. In its 2025 technical blog, Groq describes “TruePoint numerics,” including 100-bit intermediate accumulation and selective quantization choices (FP32 for attention-sensitive operations, block floating point for MoE weights, FP8 storage in error-tolerant layers), and claims 2-4x speedups versus BF16 without measurable accuracy degradation on benchmarks such as MMLU and HumanEval. Even if the absolute uplift is workload dependent, the strategic point is that Groq is pursuing performance via end-to-end co-design: precision policy is not just hardware capability (FP8/BF16) but compiler-enforced mapping of precision to error sensitivity, which can matter materially for inference cost-per-token if it reduces memory traffic and boosts throughput without forcing aggressive, accuracy-damaging quantization. Independent performance datapoints indicate Groq has been credible on latency-oriented inference speed, at least for certain regimes. EE Times reported in 2023 that Groq demonstrated Llama-2 70B inference at ~240 tokens/s per user on a cloud-based dev system described as 10 racks and 64 chips, using the company’s 1st-gen silicon introduced several years earlier. Separate Groq commentary around independent benchmarking cites results showing ~241 tokens/s throughput and ~0.8s time to receive 100 output tokens for a Llama-2 70B API configuration, positioning the platform as a step-change in “available speed” for certain interactive use cases. These figures do not settle total cost-of-ownership versus GPUs or hyperscaler ASICs, but they establish that Groq’s system-level architecture can deliver strong single-user throughput and latency on large models when properly partitioned and scheduled. GroqCloud is the commercial wrapper that packages this hardware/software stack as “tokens-as-a-service,” aiming to make Groq adoption feel like switching API endpoints rather than adopting new silicon. Groq’s documentation states its API is designed to be “mostly compatible” with OpenAI client libraries, and its pricing page provides model-specific token rates, published speeds (tokens/s), prompt caching discounts, and batch processing discounts. For example, pricing lists inputs as low as $0.05 per 1M tokens and outputs as low as $0.08 per 1M tokens for certain smaller LLM configurations, with higher prices for larger models and long-context or MoE variants; it also advertises prompt caching with a 50% discount on cached input tokens for certain models and a batch API offering 50% lower cost for asynchronous processing windows. These mechanics are economically important because they demonstrate Groq’s go-to-market is not simply “sell chips,” but “sell predictable unit economics per token,” with tooling (batch, caching) that directly targets inference cost drivers (reused prompts, throughput smoothing, and asynchronous workloads). The cloud footprint and distribution partnerships indicate Groq has been building an inference-native “edge within the cloud” strategy rather than competing head-on with hyperscalers on breadth of services. A 2025 Groq newsroom release describes a European deployment in Helsinki with Equinix, positioned as latency reduction and data governance for European customers, and explicitly references Equinix Fabric enabling private connectivity to GroqCloud over public, private, or sovereign infrastructure. The same release enumerates additional capacity in the U.S. (Equinix, DataBank), Canada (Bell Canada), and Saudi Arabia (HUMAIN), and states these sites collectively served more than 20M tokens/s across Groq’s global network at that time. That supply-side metric matters because it provides a directional sense that Groq is scaling capacity as a network, not merely as a chip vendor. Customer disclosure is inherently limited because Groq is private and many enterprise deployments are not public, but Groq’s marketing materials and partnerships provide signals about demand vectors. The company’s public website displays logos of large consumer and enterprise brands (e.g., Dropbox, Vercel, Chevron, Volkswagen, Canva, Robinhood, Riot Games, Workday, Ramp) and includes a published customer quote claiming a 7.41x chat speed increase and an 89% cost reduction after moving to GroqCloud, followed by a tripling of token consumption. While marketing claims should be treated as case-specific and not generalized, they indicate that Groq is targeting both AI-native developers (who measure success by latency and cost-per-token) and enterprise buyers (who care about predictable performance and governance). Supplier and dependency mapping for Groq spans 3 layers: silicon production, system integration, and cloud infrastructure. On silicon, third-party analysis indicates GlobalFoundries 14nm for the 1st-gen Groq chip, implying a supply chain less constrained by the most capacity-tight leading-edge nodes and advanced packaging bottlenecks that dominate high-end GPU supply (HBM stacks, CoWoS-type packaging constraints). If accurate, this is strategically meaningful because it suggests Groq capacity expansion could be gated more by conventional wafer supply, board assembly, and data center power than by the same HBM/advanced packaging scarcity that has constrained top-tier GPU ramp cycles. On systems and cloud, Groq’s own releases identify colocation and connectivity partners (Equinix, DataBank, Bell Canada) and a Middle East partner (HUMAIN), implying dependencies on data center real estate, power availability, and network connectivity, alongside procurement of standard server components, NICs/switching, racks, and cooling infrastructure. The Groq design narrative also emphasizes air cooling and reduced need for complex power/cooling infrastructure, which—if realized in deployments—can widen the set of feasible hosting locations and lower deployment friction relative to liquid-cooled, very high power density GPU racks. Against that backdrop, the strategic rationale for NVIDIA acquiring Groq can be framed as a set of overlapping objectives: inference silicon optionality, architectural hedging, competitive defense, and supply chain diversification, with the carve-out of GroqCloud signaling a preference to avoid direct cloud competition and to focus on IP and product portfolio control rather than operating a capital-intensive token-serving business. The deal, if confirmed, would occur at a valuation step-up of ~190% versus Groq’s reported ~$6.9B private valuation in the September $750M round, reinforcing that any acquisition logic would be predominantly strategic rather than a conventional financial multiple arbitrage. The most compelling strategic driver is inference. Training has historically been the center of gravity for cutting-edge GPU demand, but inference volume is structurally larger and more distributed as deployments scale, with economics dominated by cost-per-token, latency guarantees, and utilization under spiky demand. Inference workloads also create a strategic vulnerability for NVIDIA: hyperscalers and large platforms can justify bespoke ASICs (TPU, Trainium/Inferentia, Maia-class efforts) because inference is stable, repeatable, and can amortize software investment at massive scale. Groq’s core proposition—deterministic, compiler-scheduled inference with predictable latency—aligns directly with the segment where GPU generality is least valued and where “good enough” programmability plus superior unit economics can win share. Acquiring Groq would allow NVIDIA to own a credible inference-native architecture rather than relying solely on GPUs and software optimization to defend that segment. Competitive defense logic is also plausible. Groq occupies a specific competitive wedge: low-latency, high-throughput interactive inference, delivered via a simple API abstraction that reduces switching cost. That wedge directly pressures GPU inference margins in the long run because it makes inference price/performance comparisons more transparent at the token level, and it targets a developer persona that historically defaulted to CUDA-first ecosystems. Even if NVIDIA’s current-generation systems can achieve very high tokens/s per user with extensive optimization, the strategic risk is that competing architectures normalize the idea that inference is best served by special-purpose silicon with a simpler programming model, weakening CUDA lock-in at the application layer. NVIDIA has actively demonstrated that Blackwell-era systems can exceed 1,000 tokens/s per user in benchmarked configurations, but that performance leadership does not automatically translate to lowest cost-per-token across the full range of batch sizes, latency targets, and deployment environments. Groq’s existence as a credible alternative architecture forces NVIDIA to keep defending inference economics rather than only raw performance leadership. The “technology acquisition” rationale is unusually strong in this specific case because Groq’s differentiator is not a single block of silicon IP but an end-to-end methodology: compiler-led static scheduling, deterministic networking, and a system architecture designed around tensor-parallel inference rather than throughput-maximizing batch inference. NVIDIA’s stack is already compiler-heavy (TensorRT, Triton, CUDA graphs, kernel fusion, speculative decoding techniques), but GPUs remain dynamically scheduled devices with complex memory hierarchies and stochastic latency behaviors under contention. Groq’s approach provides an alternate design point: treating the entire inference execution (compute plus communication) as a statically schedulable program. In principle, that IP could be valuable even if Groq silicon itself is not adopted at massive scale, because it can inform how NVIDIA builds future inference-optimized products, compilers, and networking fabrics, especially as distributed inference with large models makes communication a first-order performance determinant. Supply chain diversification is a non-obvious but potentially important driver. If Groq’s mainstream product generation is truly based on a mature process node and avoids HBM, then the scaling constraints look different than those of state-of-the-art GPUs. NVIDIA’s ability to meet incremental demand has been tightly coupled to advanced packaging and HBM supply, and those constraints can remain binding even when wafer supply is available. An inference ASIC architecture that relies primarily on on-chip SRAM and scales by adding chips—while not costless—could reduce dependence on HBM availability and advanced packaging capacity, enabling NVIDIA to ship “inference capacity” in higher absolute volumes or into geographies and customer segments where the highest-end GPUs are economically or logistically difficult to deploy. This could be particularly relevant for latency-sensitive inference deployed in regional colocation footprints rather than centralized hyperscale campuses. The carve-out of GroqCloud, if accurate, is itself a strategic signal about NVIDIA’s priorities. Operating a token-serving cloud at scale is capital intensive, structurally lower margin than silicon IP rents, and creates channel conflict with hyperscalers and CSP partners who are core NVIDIA customers. NVIDIA has generally positioned its cloud offerings through partnerships rather than as a direct hyperscale competitor. Excluding GroqCloud would preserve neutrality with CSPs and avoid inheriting multi-region data residency obligations and partner contracts, while still allowing NVIDIA to acquire Groq’s silicon, compiler technology, and engineering talent. At the same time, excluding GroqCloud would also mean NVIDIA would not automatically acquire the commercial proof-point of Groq’s unit economics or the customer contracts that validate product-market fit at scale, increasing the importance of diligence on whether Groq’s cloud pricing is structurally profitable or partially subsidized by fundraising. There is also a “preemptive acquisition” angle. The reporting identifies recent investors in Groq’s latest round including large financial institutions and strategic/industry players. In that context, Groq represents an asset that could plausibly have been acquired by a competitor (AMD/Intel) or by a hyperscaler seeking to accelerate inference independence. NVIDIA acquiring Groq could be a defensive move to prevent a credible inference-native architecture from being weaponized by a rival with deep distribution. Even if GroqCloud is carved out, controlling the silicon roadmap and compiler IP would meaningfully constrain Groq’s ability to evolve into a standalone competitor, unless the carved-out entity retains long-term rights to the hardware and software stack. However, the strategic case is not one-sided; there are meaningful risks and potential contradictions that would need to be reconciled for the transaction to be value-accretive on a multi-year horizon. 1st, Groq’s architecture appears to rely on scaling out chip count to achieve capacity, which introduces system cost, networking complexity, and physical footprint considerations. The absence of external memory and limited on-chip SRAM implies very large models require substantial chip parallelism, and the economics then depend heavily on chip cost, yield, power efficiency, and interconnect overhead. SemiAnalysis explicitly frames Groq as trading space for time and raises questions about token economics and whether publicly advertised pricing reflects fully loaded costs or market share capture. 2nd, integration risk is non-trivial. Groq’s compiler-led deterministic model is philosophically and practically different from CUDA’s dominant programming and execution model. A poorly executed integration could create internal product confusion, dilute engineering focus, or alienate developers if the combined stack fragments. 3rd, there is cannibalization risk. If Groq-class inference silicon undercuts GPU inference economics, NVIDIA could face internal margin trade-offs, even if the goal is to defend share against hyperscaler ASICs. Cannibalization can still be rational if it prevents larger share loss, but it would require crisp portfolio segmentation and go-to-market discipline. The presence of NVIDIA’s own rapidly improving inference performance complicates the “need” for Groq but does not eliminate the “option value.” NVIDIA has demonstrated benchmark-leading tokens/s per user on Blackwell-based systems, suggesting that raw interactive throughput is not necessarily the limiting factor for NVIDIA’s product line. The more enduring strategic question is unit economics and architectural control: whether future inference demand is better monetized through general-purpose GPUs plus software optimization, or whether a bifurcated product portfolio (training GPUs plus inference-native ASICs) becomes necessary to defend total AI compute wallet share as hyperscaler ASIC penetration increases. Acquiring Groq could be a decisive move to ensure NVIDIA participates in both regimes rather than betting exclusively on GPUs to win inference forever. What is “special” about Groq’s technology relative to a typical accelerator roadmap is the tight coupling of determinism, compilation, and networking into a single scheduling problem. The LPU narrative emphasizes deterministic compute and networking, static scheduling, and direct chip-to-chip coordination that allows “hundreds” (more precisely, 100s) of chips to behave like a single scheduled resource. The architecture also explicitly targets tensor-parallel, latency-optimized distribution rather than pure data-parallel throughput scaling, which matters for real-time applications where a single response must arrive quickly rather than many requests being processed in bulk. The implication is that Groq is optimized for the time-to-first-token and steady token streaming behavior that defines user experience in interactive LLMs, and it attempts to achieve that without relying on large batch sizes that can degrade latency. From a portfolio manager’s perspective, the most important interpretation is that an NVIDIA-Groq combination would likely be less about “NVIDIA needs more inference speed” and more about controlling the architectural trajectory of inference acceleration and removing a fast-improving, developer-friendly competitor from the market. The carve-out of GroqCloud would reinforce that the transaction is aimed at IP, talent, and product optionality, not acquiring a cloud revenue stream. The valuation step-up implied by $20B versus $6.9B would therefore be justified only if the acquired assets materially reduce long-term competitive risk (hyperscaler ASIC displacement, inference margin compression) or enable new monetization vectors (inference ASIC product line, supply chain de-bottlenecking, improved software determinism) that would be difficult to achieve on a comparable timeline via internal R&D.

TheValueist

102,145 次观看 • 8 个月前