正在加载视频...

视频加载失败

Centralized AI inference runs on one provider. NEAR AI's confidential inference draws from multiple compute partners distributed across regions instead of a single data center. No KYC required. No visibility into what you're running. Intelligence owned by the user.

73,204 次观看 • 1 个月前 •via X (Twitter)

0 条评论

暂无评论

原始帖子的评论将显示在这里

相关视频

Perplexity CEO Aravind Srinivas on the biggest threat to the data center industry: It's not competition. It's not regulation. It's decentralisation. "The biggest threat to a data center is if the intelligence can be packed locally on a chip that's running on the device and then there's no need to inference all of it on like one centralized data center." He outlines how this could work in practice. Personalisation doesn't necessarily require on-device model training. Retrieval augmented generation, tool calls, and local data can already tailor AI to individual users. But the real unlock? Test time training. Aravind Srinivas describes a future where AI lives on your device, watches how you work and gradually automates your repetitive tasks. "Imagine we crack test time training where the AI watches tasks you repeatedly do on your local system, adapts to you over time and starts automating a lot of the things you do." The key insight: in this model, the intelligence belongs to you. It's your data, your device, your personalised AI brain. And if that future arrives, the economics of centralised infrastructure start to collapse. "That really disrupts the whole data center industry. It doesn't make sense to spend all this money, 500 billion, 5 trillion, whatever on building all the centralized data centers across the world that do a lot of the intelligence workloads for people." The companies spending trillions on centralised infrastructure may want to rethink where intelligence actually needs to live.

Big Brain AI

90,276 次观看 • 5 个月前

NEW: Premium Inference 101 The Economics & Infrastructure Behind Running Trillion Parameter Models Rodrigo Liang, CEO & Co-Founder of SambaNova "Inference has arrived. 70-80% of those racks are running inference." "[Inference services] are generating lots of revenue, but not enough margin. In order for them to sustain, they've gotta be more profitable." "With SambaNova, that min quantum is down to 1 rack. Where if you have other service providers, [with] say, a DeepSeek model, now 1.5 trillion parameters, to run that, the min for some of the other providers might be 10-20 racks." SambaNova builds full-stack inference infrastructure. 16 chips to a 10kW air-cooled rack that runs trillion parameter models, where a GPU rack pulls 130kW. They just demonstrated the fastest MiniMax M2.7 inference in the world, as benchmarked by Artificial Analysis. The demo paired one NVIDIA H200 rack for prefill with one SambaRack SN50 for decode. Disaggregated inference: GPUs load the context, RDUs generate the tokens. Now serving JPMorgan, SoftBank, Saudi Aramco & DOE national labs, just valued at $11B on a $1B Series F led by General Atlantic. We Cover: › Why inference will need orders of magnitude more chips than training ever did › The 10kW rack vs the 130kW rack, & why air cooling decides geography › Running a 1T parameter model in one rack at full precision, no quantization › The agent latency problem: 20 agents, 2 seconds each, 40 seconds gone › Revenue per rack, & why inference providers have revenue but no margin › JPMorgan, sovereignty, & the move back to on-prem Filmed at the RAISE Summit in Paris. Thank you to Brex, MongoDB & AssemblyAI for helping make this trip & content series happen. 𝐓𝐈𝐌𝐄𝐒𝐓𝐀𝐌𝐏𝐒 (00:00) Rodrigo Liang , Co-Founder & CEO at SambaNova Systems (00:59) SambaNova’s Series F: $1B raise at an $11 billion valuation (03:00) The Inference problem nobody saw coming (04:52) SambaNova's chip evolution (07:19) Running a trillion-parameter model on a single rack (11:00) Do $100 billion data centers actually make sense? (14:14) What "premium inference" really means (18:28) Speed is about to become AI's biggest price tag (20:43) Starlink, edge computing, & AI reaching every corner of the planet (24:27) Working alongside NVIDIA & rival chipmakers (27:49) How customers actually measure inference performance (32:07) The biggest bottlenecks in AI's global land grab (35:12) Justifying the billion-dollar AI valuations (37:53) Why SambaNova refuses to build its own cloud (41:03) The "AI sovereignty" debate (43:48) Data privacy fears are driving the return to on-prem AI (47:55) How to actually get ROI out of AI spend (51:16) The one question every business should be asking about AI (56:09) The mentors & lessons behind a 32-year career in chips (58:02) Unveiling SambaNova's newest chip, the SN50

Molly O’Shea

134,960 次观看 • 1 个月前

70,000 Phones, One AI Agent — The World's Largest Edge AI Fleet Runs on Hermes We turned 70,000 phones into a shared AI compute network. Any device owner contributes idle compute. Any developer taps distributed inference at a fraction of cloud cost. Not a concept. Not a whitepaper. 70K devices online today. The problem: orchestrating a shared network of heterogeneous edge devices — different chipsets, different memory, different thermal profiles, different owners — is a coordination nightmare no human team can handle manually. So we gave the network a brain: Nous Research Hermes Agent. Hermes connects to 16 MCP servers and runs 24/7: 🔬 Research Loop — Tracks every breakthrough in on-device inference: quantization (GPTQ/AWQ/GGUF), speculative decoding on mobile SoCs, federated learning protocols. Auto-imports papers into NotebookLM. 36 research topics, zero manual curation. 🌐 Network Intelligence — Monitors device availability, compute capacity, and workload distribution across the shared fleet. Surfaces bottlenecks before they cascade. 🧬 Tech Tree Optimizer — Maps the full optimization frontier: from KV-cache compression to on-device LoRA to peer-to-peer model sharding. Hermes autonomously identifies which research paths unlock the most network-wide throughput gains. The result: a self-improving shared compute network. Research compounds daily. The fleet gets smarter without human intervention. Cloud AI scales with money. We scale with people. #HermesHackathon Teknium 🪽 Delphi Digital Tommy

Oyster Republic 🦪📲🦞👓

20,721 次观看 • 5 个月前