Video wird geladen...

Video konnte nicht geladen werden

Zur Startseite

The Cost of Intelligence is Heading to Zero | Hyperspace P2P Distributed Cache We present to you our breakthrough cross-domain work across AI, distributed systems, cryptography, game theory to solve the primary structural inefficiency at the heart of AI infrastructure: most inference is redundant. Google has reported that only...

37,555 Aufrufe • vor 4 Monaten •via X (Twitter)

0 Kommentare

Keine Kommentare verfügbar

Kommentare vom Original-Post werden hier angezeigt

Ähnliche Videos

Micron is going to $4,000 and once you understand what inference actually is, the number stops sounding crazy (Save this). Dylan Patel just said that by 2030, OpenAI and Anthropic alone will need over 100 gigawatts of compute combined and by 2040, we may not even be measuring AI infrastructure in gigawatts anymore. We may be talking about terawatts. Every single one of those gigawatts needs memory to function. Without it, the compute is worthless. Most people heard that and thought about Nvidia but they should be thinking about Micron. Every AI model generating a response has two phases. The first is prefill, processing your prompt which is compute-heavy and the second is decode generating each word one token at a time and that phase is almost entirely memory-bound, not compute-bound. During decode, the GPU's processing units sit idle more than 95% of the time, waiting for data to arrive from memory. Google confirmed it in a research paper that decode-phase bottlenecks are dominated by memory bandwidth and capacity not raw compute. The GPU is not the bottleneck but the memory feeding the GPU is. This matters because inference is now where all the money lives. Training a model happens once, Inference happens billions of times a day every ChatGPT response, every Claude output, every agentic workflow running in the background and every one of those token streams is a billing event tied directly to memory performance. Adding more GPUs does not fix this because GPUs are already underutilized in inference because they are sitting idle waiting on memory. Adding more memory bandwidth and capacity is what directly reduces token cost, reduces latency, and allows the same cluster to serve dramatically more users simultaneously. Longer context windows compound the problem further, a model running a 1 million token context window requires dramatically more memory per session than a 10,000 token window, and every new model generation pushes context longer. The market treats memory as a downstream beneficiary of Nvidia orders. The correct framework is the opposite, Micron is the upstream constraint on how much value every Nvidia GPU can actually generate at inference scale. Micron guided Q4 to $50 billion in revenue, has HBM4 ramping at twice the pace of the prior generation, and CEO Sanjay Mehrotra has said supply will not catch demand before the end of 2027. At 8x forward earnings on $112 projected FY2027 EPS, Micron is the most undervalued infrastructure company in the entire AI stack. Inference is memory. Memory is Micron and the inference ramp has barely started. Milk Road Pro members are already up massively on this position and we're just getting started. If you want the full breakdown of what we're buying and why, come join us for just a dollar using the link below!

Milk Road AI

128,678 Aufrufe • vor 1 Monat

I pay Claude $20 a month. Most $TAO holders do too. There is a stack you can build in 15 minutes that fixes that completely. It runs on Bittensor. It costs $10. You do not write a single line of code. Here is how every AI chat product actually works under the hood. Three layers. Always three. The model. The brain. GPT, Claude, DeepSeek, Kimi, GLM. The inference layer. The GPU that runs the model when you hit send. The interface. The chat box you actually look at. ChatGPT and Claude bundle all three and hand you the result. You cannot change the model. You cannot change the inference. The interface is non-negotiable. Every prompt you type goes to a server run by a private company whose terms of service can quietly change next month. The anti-ChatGPT move is to pick each layer yourself. This is where $TAO comes in. Chutes is Subnet 64 on Bittensor. It is the inference layer. Open source models like DeepSeek, Kimi, GLM, and Llama get served by a global network of miner-operated GPUs. Validators score the output quality. The best inference wins the emissions. You hit send. A miner somewhere runs your prompt. You get the answer back. The TAO you hold is in part paying for the GPU you just used. The basic stack is one URL. chutes. ai/chat No account. No API key. No setup. Switch models mid-conversation. Web search built in. Image generation. File uploads. Free. The advanced stack is Chutes plus TypingMind. One-time license. No recurring fee. Plugins, agents, custom personas, a prompt library you build over months. Full model switching between Chutes, OpenAI, and Anthropic from the same window. Total cost: $10 a month to Chutes for inference. That $10 buys you $50 in actual usage. But here is the signal most people missed inside this story. Chutes ran a free tier until February. Then they killed it. Then they raised the minimum to $10 in May. Most people saw that as bad news. It is the opposite. Free things on the internet do not last. Real products do. Chutes is becoming a real product. A subnet that generates actual revenue from actual users paying actual money for actual AI inference. That is what $43 million in Q1 network revenue looks like at the individual subnet level. And there is one more thing ChatGPT and Claude cannot offer that Chutes already has. Trusted Execution Environments. Your prompt gets encrypted on your device, shipped to a confidential compute GPU, and the lock only breaks inside the chip. The miner running the model physically cannot read your prompt. ChatGPT cannot promise that. Claude cannot promise that. Bittensor already built it. You are holding a network where the subnets are generating real revenue, shipping real privacy infrastructure, and replacing $20 a month centralised subscriptions with $10 a month decentralised inference. The people who use the product always understand the investment better than the people who only watch the price.

2xnmore

27,088 Aufrufe • vor 2 Monaten

70,000 Phones, One AI Agent — The World's Largest Edge AI Fleet Runs on Hermes We turned 70,000 phones into a shared AI compute network. Any device owner contributes idle compute. Any developer taps distributed inference at a fraction of cloud cost. Not a concept. Not a whitepaper. 70K devices online today. The problem: orchestrating a shared network of heterogeneous edge devices — different chipsets, different memory, different thermal profiles, different owners — is a coordination nightmare no human team can handle manually. So we gave the network a brain: Nous Research Hermes Agent. Hermes connects to 16 MCP servers and runs 24/7: 🔬 Research Loop — Tracks every breakthrough in on-device inference: quantization (GPTQ/AWQ/GGUF), speculative decoding on mobile SoCs, federated learning protocols. Auto-imports papers into NotebookLM. 36 research topics, zero manual curation. 🌐 Network Intelligence — Monitors device availability, compute capacity, and workload distribution across the shared fleet. Surfaces bottlenecks before they cascade. 🧬 Tech Tree Optimizer — Maps the full optimization frontier: from KV-cache compression to on-device LoRA to peer-to-peer model sharding. Hermes autonomously identifies which research paths unlock the most network-wide throughput gains. The result: a self-improving shared compute network. Research compounds daily. The fleet gets smarter without human intervention. Cloud AI scales with money. We scale with people. #HermesHackathon Teknium 🪽 Delphi Digital Tommy

Oyster Republic 🦪📲🦞👓

20,721 Aufrufe • vor 5 Monaten

If intelligence is the log of compute… it starts with a lot of compute! And that’s why we’re scaling our GPU fleet faster than anyone else. Just last year, we added over 2 gigawatts of new capacity – roughly the output of 2 nuclear power plants. And today we’re going further, announcing the world's most powerful AI datacenter, located in southeastern Wisconsin. Fairwater is a seamless cluster of hundreds of thousands of NVIDIA GB200s, connected by enough fiber to circle the Earth 4.5 times. It will deliver 10x the performance of the world’s fastest supercomputer today, enabling AI training and inference workloads at a level never before seen. For AI training workloads, you need compute at exponential scale. That’s why we designed the datacenter, GPU fleet, and network together as one integrated system. This ensures a single job can run from day 1 at exponential scale across thousands of GPUs. Fairwater uses a liquid-cooled closed-loop system for cooling GPUs that requires zero water for operations after construction. And we’re matching all of the energy that is consumed with renewable sources. And of course, it is just one of several similar sites we’re lighting up across our 70+ regions. We have multiple identical Fairwater datacenters under construction in other locations across the US, in addition to our AI infrastructure already deployed in over 100 datacenters around the world, powering model training, test-time compute, RL tuning, and real-time inference at global scale. Too often during times like this, people go with the current and only later wonder, how did we get here? With Fairwater, we're charting a new path: doing the hard engineering work, bringing compute, network, and storage into one highly scaled cluster, and designing closed-loop energy systems to meet real-world computing needs. And partnering with local communities to ensure it's thoughtfully done in a way that is sustainable, creates new jobs, and expands opportunity. We are thrilled to see this take hold in Wisconsin, and we are just getting started.

Satya Nadella

2,022,693 Aufrufe • vor 10 Monaten

Chamath said AI is not like the internet. Every new user costs real money. And the infrastructure making it possible was built by everyone. His argument was the clearest case for government ownership of AI labs I have ever heard. And it had nothing to do with Bernie Sanders. Start with the internet comparison. Google and Facebook became the most profitable companies in human history because of one number. The marginal cost of adding a new user was effectively zero. One more search query cost Google nothing. One more Facebook profile cost Meta nothing. They could serve a billion people and the incremental cost of that billion person was rounding error. That is the money printer. Infinite scale at zero marginal cost. AI breaks that model completely. Every single user taxes a GPU. Every query costs electricity. Every response requires memory and compute. The marginal cost of AI is real, significant, and does not disappear at scale. You cannot print money the same way. Then Chamath made the point that landed hardest. The infrastructure these companies depend on, the power grid, the land, the data centers, the permitting, the national security apparatus that protects their chips from being stolen, none of that was built by Anthropic or OpenAI. It was built by the public. By taxpayers. By decades of government investment in the physical and legal foundation these companies are now running on. He compared it to the interstate highway system. If the federal government built the roads and two companies transported all the goods on them, a logical question at that point would be how much of that should I own? You are riding on my rails. His conclusion was direct. If he were running a sovereign wealth fund and had the negotiating leverage of the US government, he would own 75% of these companies when he was done. The internet had zero marginal cost. That is why the founders captured almost all of the value. AI has real marginal cost and runs on public infrastructure. That changes who has a claim on what gets built. WATCH THE FULL PODCAST ON The All-In Podcast

Ihtesham Ali

79,066 Aufrufe • vor 1 Monat

Mark Zuckerberg is explaining one of the most misunderstood dynamics in AI and it has direct investment implications (Save this). The concept he's describing is model distillation, and it's one of the most important techniques to emerge in AI over the past year. Here's how it works. You train a massive, enormously expensive model, in Meta's case, Llama 4 Behemoth, a 2 trillion parameter teacher model and then you use that model to teach a much smaller, cheaper model. The smaller model inherits roughly 90 to 95% of the intelligence of the giant while running at 10% of the cost and on a fraction of the compute. Meta already did this with the Llama 4 family and Behemoth serves as the teacher. Llama 4 Scout and Maverick, the publicly released open-source models were distilled from it. Scout runs on a single H100 GPU with a 10 million token context window and outperforms models that cost far more to operate. Maverick, at 17 billion active parameters, rivals DeepSeek V3 in coding at half the parameter count and beats GPT-4o on multimodal benchmarks. Both are completely free for commercial use. What Zuckerberg is pointing at is a structural shift in how AI gets deployed in the real world. Companies aren't taking a frontier model off the shelf and running it as-is but rather taking open-source models, fine-tuning them on their own proprietary data, distilling them into even smaller custom models tailored to their specific use case, and running them on infrastructure they control at a fraction of the cost of a closed frontier API. The investment implication of this is significant and runs in two directions. For Meta specifically, this is a strategic masterstroke. Every company that builds on Llama, fine-tunes it, distills it, or deploys it through their infrastructure is pulling into Meta's orbit while Meta builds the most powerful open teacher model. The ecosystem of companies using it grows and that ecosystem generates commercial activity across Meta's platforms and data services. Meta's AI research benefits from billions of real world deployment signals and it's a flywheel that closed model providers cannot replicate because their strategy requires charging per token, which is now a 65x cost disadvantage against the open-source alternative. For the broader market, distillation changes the economics of inference in a way that has barely been priced in. As intelligence becomes extractable into smaller and cheaper models, the absolute demand for compute doesn't decline but rather it explodes, because now the number of applications that are economically viable expands by orders of magnitude. Every task that was previously too expensive to automate at $3.25 per call becomes viable at $0.05 that means more total token usage, more total GPU utilization, and more demand for the infrastructure companies, the Nebiuses, the GE Vernovas, the Constellation Energies that supply the underlying compute and power.

Milk Road AI

27,908 Aufrufe • vor 1 Monat

Today we announced our new Fairwater datacenter in Atlanta, connected with our first Fairwater site in Wisconsin and our broader Azure footprint to create the world’s first AI superfactory. Fairwater exemplifies our vision for a fungible fleet: infra that can serve any workload, anywhere, on fit-for-purpose accelerators and network paths, with maximum performance and efficiency. AI workloads have evolved beyond large-scale pre-training. Today, they encompass fine-tuning, reinforcement learning (RL), synthetic data generation, evaluation pipelines, and more. Fairwater is built to support this full lifecycle: Max density: Fairwater’s two-story design and liquid cooling system lets us place racks in three dimensions and pack them with GPUs as densely as possible, minimizing cable runs and improving latency and effective bandwidth. Fleet: Each Fairwater DC can integrate hundreds of thousands of the latest NVIDIA GPUs into a single coherent cluster. This provides flexible infra that can support the full spectrum of workloads, and ensure no GPU is left unnecessarily idle. And that’s on top of the more than 100,000 GB300s coming online this quarter alone for inference across the rest of our fleet. For us, it’s all about turning every gigawatt into the maximum number of useful tokens. Not every GW is created equal! Planet-scale: Every Fairwater DC will connect through our continent-spanning AI WAN to prior generations of AI supercomputers, forming a truly fungible pool of compute. This enables developers to scale beyond the capacity of a single site and dynamically land workloads on the right infra for their needs. Together, these innovations let us bring together different generations of silicon and AI systems across DCs and geos into a single elastic system that scales seamlessly across training and inference workloads And this elastic AI capacity is all available alongside all the other cloud services (compute, storage, databases, app services) that AI agents and workloads need. This is what we mean when we talk about building a fungible fleet – a single, unified platform that pushes the limits of performance per watt and per dollar. Read more:

Satya Nadella

907,893 Aufrufe • vor 9 Monaten

The value of the work we're doing at Optimum is encapsulated quite well by the phrase "speed is money". In modern markets there are real economic advantages to latency reduction. This is nothing new. Wall Street firms have long been optimizing on latency, primarily through colocation and top of the line hardware. However, when it comes to decentralized systems, expensive hardware and geographic concentration are antithetical to their purpose. Therefore we should optimize decentralized network latency through software, which I'm thrilled about because it's exactly what I've spent the better part of the past 2 decades working on with Random Linear Network Coding. Now let’s talk about networking economics, the relationship between speed and money. First, it's important to note that users will only pay for low latency if it can be consistently guaranteed. Second, you can only make that latency guarantee for a certain number of users. This is a universal law of networking. We can model this relationship on a delay curve, shown below. The delay curve is determined by the utilization rate of the network, meaning how much traffic is flowing through the network divided by the network's throughput. As you approach a level of traffic equal to the available throughput, latency trends infinitely higher. On this delay curve we can impose some utility thresholds. These thresholds are the levels of latency which are important to different groups of users because of how that latency guarantee improves their economic outcomes. Finding the point on the curve where each threshold intersects will tell us what level of traffic we can guarantee that level of latency for. Essentially, there exists a finite supply of speed on a network and the highest utility users of that speed are willing to pay more for it. I like to think of this similarly to expedited shipping options on Amazon. This is why we say speed is money, and why we can create a Latency Marketplace. The only way to increase the supply of speed is to fundamentally increase network throughput. This is what we work on at Optimum by using Random Linear Network Coding. The same relationship between traffic and throughput still applies, but now the delay curve is shifted out further to the right. Now more traffic can be processed at the same latency, or the same traffic can be processed at a lower latency. More speed available to the network. More value unlocked for the network’s users. Crucially, that value is no longer only reserved for those who can afford to sit closest to the machine. Expanding the supply of speed widens who can reach each latency threshold, keeping the network's advantage decentralized rather than concentrated in the hands of a few. When nodes join Optimum and participate, they reap the benefits, but they also add to the capacity. Rather than vying against each other in a zero-sum game, nodes help themselves and others.

Muriel Medard

44,225 Aufrufe • vor 1 Monat

AI is the first technology in history where more customers makes you POORER. Every tech company in history got cheaper as it scaled. More users meant lower costs per user. That's the entire model. That's why Microsoft prints money. That's why Google prints money. That's why Meta prints money. Software has near-zero marginal cost. Build it once. Sell it a billion times. The 100 millionth user costs basically nothing to serve. This is the single most important rule in tech economics. But AI completely broke it. Every single query costs real compute. Every interaction burns real electricity. Every response depreciates real hardware. There is no "build once, sell forever." There is only "burn money every time someone asks a question." And the numbers prove it: OpenAI hit $20 billion in annualized revenue. Losses? $14 billion. For every dollar they earn, they spend $1.69 delivering it. Their losses TRIPLED as their revenue grew. Not because they're bad at business, but simply because the model itself is broken. Anthropic crossed $30 billion in annualized revenue. Still burning billions. Still not profitable. Still raising tens of billions just to keep the lights on. xAI is burning $1 billion every single month. Perplexity spent 164% of its revenue on compute costs from AWS, They literally spent more on running the AI than they made from selling it. This is not how technology is supposed to work. Google once estimated that adding AI to every search query would require 500,000 A100 servers. The cost of answering a single AI query is 10x MORE than a traditional search result. Traditional software: Serving 1 million users costs roughly the same as serving 100,000. The marginal cost is basically zero. AI: Serving 1 million users can cost 10 times what 100,000 costs. Every new user is a new expense. Every new query is a new dollar burned. This is reverse economics. The more successful you become, the faster you die. And nobody in the industry wants to talk about it because the entire narrative depends on you believing AI companies work like software companies. But they don't. They NEVER will. Software scales to infinity. AI scales to bankruptcy. HSBC ran the numbers on OpenAI specifically. Their conclusion: Even after every funding round, every investment, every deal, OpenAI still faces a $207 BILLION shortfall to reach profitability. The industry response has been to raise prices. ChatGPT went from free to $20 to $200 for the Pro plan. And it's still not enough because the cost of running these models grows FASTER than any price increase consumers will accept. Meanwhile 966 AI startups died in 2024. A 25.6% jump from the year before. AI startups burn cash twice as fast as non-AI tech companies. And the ones building on TOP of OpenAI and Anthropic are in even worse shape. Every wrapper app. Every "AI-powered" SaaS tool. Every startup whose entire product is someone else's model with a different skin on it. They're all margin-negative. Every single one. And these are the companies about to IPO. SpaceX, OpenAI, Anthropic, and Cerebras. $240 billion in combined raises planned for 2026. They're asking you to invest in an industry where the fundamental unit economics don't work. Where the MORE customers you get, the MORE money you lose. Where no company has figured out how to make the math positive. The dot-com bubble had the same pitch: "Revenue is growing. Profitability comes later." For most of them, later never came. The question isn't whether AI will change the world. It will. The question is whether it can do it without going broke first. And right now, every single number literally says no. How can they become profitable?

Ricardo

167,414 Aufrufe • vor 3 Monaten

"I used to be a bitcoiner. The transition to a new store of value only happens once every 3,000 years. That's the main prize -- just focus on that. But [security] is the criteria that ultimately convinced me to flip from Bitcoin to ETH." "I have a higher degree of certainty that Ethereum will be around longer [than Bitcoin]. The reason for that is because Bitcoin relies on proof-of-work, which is less efficient than proof-of-stake and doesn't scale with the value of the network. And as the block subsidy of Bitcoin halves every four years, it is increasingly becoming more and more reliant on transaction fees to fund the security budget paid to miners." "If you look at [Bitcoin's] security budget right now, about 0.6% of revenue to miners is transaction fees... The problem with that is if Bitcoin becomes 'digital gold', flips gold, and becomes a $30 trillion asset, but it only costs $10-20 billion to attack it, that's too asymmetric." "You want the security budget to scale with the market cap, similar to how countries spend a % of their GDP on defense. The more valuable something is, the more you need to spend to protect it." "Ethereum, with the Merge, migrated to proof-of-stake, which is fundamentally more secure because it's less reliant on transaction fees and it scales with the value of the network. If 1/3rd of ETH is staked and then you need 1/3rd of those ETH to censor the network, you're looking at roughly 10% of the total market cap as the cost to attack the network." "So if Ethereum flips Bitcoin and gold and becomes a $30 trillion asset, it'll cost ~$3 trillion to attack the Ethereum network versus Bitcoin at like $10 billion." "The other aspect here is that as AI hyperscalers invest more and more in AI, proof-of-work becomes increasingly vulnerable because the cost to attack the Bitcoin network is starting to look close to the quarterly CapEx these hyperscalers are spending on their data centers." Full interview on Bankless with Vivek Raman discussing the new Etherealize "Productive Money" report below.

Michael McGuiness

120,456 Aufrufe • vor 3 Monaten