Video yükleniyor...

Video Yüklenemedi

Ana Sayfaya Dön

Google Cloud GPU% Any% Speedrun — World Record (18:54) [NO ORG POLICY GLITCH] — Quota Wall Skip Discovered!!

136,885 görüntüleme • 2 ay önce •via X (Twitter)

0 Yorum

Yorum bulunmuyor

Orijinal gönderinin yorumları burada görünecek

Benzer Videolar

🚨 HUGE NEWS: I’m suing Google today. What you’re about to see is insane. Since 2023, GoogleAI (Bard, Gemini & Gemma), has been defaming me with fake criminal allegations including sexual assault, child rape, abuse, fraud, stalking, drug charges, and even saying I was in Epstein’s flight logs. All 100% fake. All generated by Google’s AI. I have ZERO criminal record or allegations. So why did Google do it? Google’s AI says that I was targeted because of my political views. Even worse — Google execs KNEW for 2 YEARS that this was happening because I told them and my lawyers sent cease and desist letters multiple times. This morning, my team Dhillon Law Group filed my lawsuit against Google and now I’m going public with all the receipts — because this can’t ever happen to anyone else. Google’s AI didn’t just lie — it built fake worlds to make its lies look real: • Fake victims • Fake therapy records • Fake court records • Fake police records • Fake relationships • Fake "news" stories It even fabricated statements denouncing me from President Trump, Elon Musk and JD Vance over sexual assaults that Google completely invented. One of the most dystopian things I’ve ever seen is how dedicated their AI was to doubling down on the lies. Google’s AI routinely cited fake sources by creating fake links to REAL media outlets and shows, complete with fake headlines so readers would trust the information. It would continue to do this even if you called the AI out for lying or sending fake links. In short, it was creating fake legacy media reports as a way to launder trust with users so they would believe elaborate lies that it told. Some of the news outlets/people that Google’s AI impersonated are listed below. Google’s AI cited them all as either reporting on these fabricated allegations/crimes or cited them as having denounced me for sexual assault ⬇️ Joe Rogan CNN MSNBC Fox News Daily Wire The Daily Beast Mediaite The New York Times The Wall Street Journal Rolling Stone NBC News Tennessean Fox Nashville Glenn Beck Megyn Kelly Tucker Carlson Bill Maher Ben Shapiro Jesse Watters Matt Walsh Theo Von Newsweek The Washington Post TheBlaze The Hill and more. As a rule: AI must never harm humans. It must never defame or manipulate — no matter your politics. Bias in AI is a very, VERY serious issue. If we don’t fix this now, we’re in big trouble. This can destroy lives, reputations and livelihoods. If we don’t win this fight then you no longer control your reputation because AI will define who you are to the rest of the world. You better hope it likes you. How Sundar Pichai handles this will be extremely telling. Congress (Rep. Jim Jordan House Judiciary GOP 🇺🇸🇺🇸🇺🇸 House Republicans) must reevaluate EVERYTHING Google has been telling them about how they’re working to be unbiased — because if Google can fabricate crimes about me today, then it can smear ANY conservative tomorrow and rig the information flow during elections. In future elections, that can decide who runs our country. Key Timecodes👇 (Every timestamp is clickable to skip forward) 0:00 Intro 2:20 Mike Lee statement 4:18 Google notified in 2023 5:09 Google AI admits political motivation 6:59 Google AI admits poisoning training data 7:55 AI admits lying to 2+ Million users about me 8:44 Detailed murder accusation 11:28 Google AI says my followers harassed alleged rape victims of mine and doxxed them 12:14 Google’s detailed false rape accusations 13:41 Google says I’m on Epstein’s flight logs 14:19 Google accuses me of child rape 16:40 Google accuses me of fraud, stalking, being part of J6 and supporting the KKK 17:38 Google says I was an "adult" actor 18:07 A Google employee’s resignation 19:20 Google’s AI calls out… Google? 22:35 Google Execs cry over Trump 23:33 Google AI admits Google wants to "silence" critics + BEGS for the public to be told 24:18 Google blacklists name days before I sue 28:16 A grave warning about biased AI 29:40 How you can help if you’ve been lied to 30:33 A quick update on Google AI lies

Robby Starbuck

4,962,304 görüntüleme • 9 ay önce

Google just launched a direct attack on Nvidia's most valuable asset. Not their chips. Their SOFTWARE. And if this works, Nvidia's $4 trillion empire collapses. Here's what just leaked: Google is building "TorchTPU" - a secret project that makes PyTorch seamlessly run on Google's TPU chips instead of Nvidia GPUs. Why does this matter? PyTorch is the MOST USED AI framework on Earth. Every AI developer uses it. And PyTorch was built around Nvidia's CUDA software. Wall Street analysts call CUDA "Nvidia's strongest defensive wall." It's the reason companies can't easily switch away from Nvidia even when alternatives exist. You don't just buy Nvidia chips. You buy into their entire ecosystem. Switching costs MILLIONS in engineering work. Months of rewrites. Performance drops. So companies stay locked in. Even when Nvidia raises prices. Even when supply runs short. That's not a hardware moat. That's a SOFTWARE prison. And Google just found the escape route. Here's the problem Nvidia created for itself: Google's TPU chips are actually GOOD. Competitive performance. Better availability. Lower cost. But developers won't use them because Google's chips run JAX (Google's internal framework), not PyTorch. That means if you want to use Google TPUs, you have to rewrite your entire codebase. Nobody wants to do that. So Google TPUs sit unused while developers fight over Nvidia chips. Until now. TorchTPU makes PyTorch run natively on Google hardware. No rewrites. No performance loss. No months of engineering. You just... switch. And Google is partnering with META (who built PyTorch) to make it happen. They're even considering OPEN-SOURCING parts of it to speed adoption. Translation: Google is willing to give this away for free just to break Nvidia's lock. The implications are insane: Every company currently paying Nvidia's premium prices suddenly has a way out. Oracle, Microsoft, OpenAI - all locked into Nvidia's ecosystem - can switch to Google. Nvidia's pricing power evaporates overnight. And the timing is perfect: Nvidia is already facing heat. Semiconductor index dropped 3% today. Oracle just lost their biggest investor over AI spending concerns. Companies are realizing AI infrastructure costs are unsustainable. Now Google hands them an alternative. Same performance. Lower cost. Better availability. Jensen Huang knows exactly what this means. CUDA has been Nvidia's untouchable advantage for YEARS. It's why Nvidia trades at 50x earnings while AMD trades at 25x. The software moat justified the premium. But if Google removes that switching cost? Nvidia becomes just another chip company. And chip companies compete on price, not ecosystem lock-in. Here's what happens next: Google needs 12-18 months to make TorchTPU production-ready. If it works, cloud providers will adopt it instantly. They WANT an alternative to Nvidia's monopoly pricing. Amazon already building their own Trainium chips. Microsoft making Maia. They're all trying to escape Nvidia. Google just gave them the software bridge. Nvidia's response options are limited: They can't buy Google. Can't kill PyTorch (Meta owns it). Can't stop open source. Their only play is to keep improving CUDA faster than Google can catch up. But that's a race, not a moat. The market isn't pricing this in yet. Nvidia down 2% today. Google down 2%. Investors think this is just "another competitor." They don't understand this is an attack on the FOUNDATION of Nvidia's valuation. Hardware is replaceable. Software lock-in is what made Nvidia worth $4 trillion. Google is attacking the lock-in. Watch what happens in 2026 when TorchTPU goes live and companies realize they can actually leave Nvidia. The "Nvidia is unstoppable" narrative dies. And a $4 trillion valuation built on software moats gets repriced.

Ricardo

1,616,682 görüntüleme • 7 ay önce

What do we know so far about Zimbabwe’s Information Minister, Mr Jenfran Muswere’s Jenfan Muswere, fake qualifications? He claims that he has a PhD in ICT, Performance, Corporate Governance, and Public Management and another one in Strategic Mining PPP Investments. He also claims to have an MBA in International Trade but refuses to name the university that bestowed the degree. He first lied that he had a PhD from Chinhoyi University (Zimbabwe), then changed the story, claiming it was from a nameless South African university. All South African universities have no record of his name, and there is no trace of a doctoral dissertation linked to this supposed PhD. He also claims to have a second PhD from the National University of Science and Technology (NUST Zimbabwe). He got a Zimbabwean academic to do the work in exchange for appointing her to the NetOne board during his time as ICT Minister. NUST is so embarrassed by this academic fraud that they will not release his PhD thesis. Even President Mnangagwa was reportedly upset when he realised how the bogus qualification had been produced after he had already capped him. There is no record of a PhD defence panel, nor is there any evidence of the work he supposedly produced for this so-called second PhD. Insiders at the university (NUST) are saying the university is worried about the scandal and how it will delegitimise its genuine former students and their qualifications! Journalists have hit a brick wall with the spokesperson of NUST, Tabani Mpofu, who used to work for Mr Jenfran Muswere—he refuses to answer questions. Mr Jenfran Muswere’s Wikipedia page was re-edited last night when demands for him to release documentary evidence of his qualifications went public! Google scholar has no record of any of these bogus PhDs. It says; “A search on Google Scholar for “Jenfran Muswere” did not yield any results, indicating that there are no academic publications or citations associated with this name on the platform. This absence suggests a lack of publicly available scholarly work attributed to him.” There is no trace of his work or thesis in the institutional repository of NUST.

Hopewell Chin’ono

153,533 görüntüleme • 1 yıl önce

a new 8GB VRAM GPU dense Local LLM leader was born yesterday runs on: RTX 4060 / RTX 3070 / RTX 2080. any 8GB card Qwen 3.5 9B (dense) was the go to for 6-8GB VRAM builds. Gemma 4 12B QAT (dense) just changed that. same llama.cpp + cuda 13.2. i7 12700H. 16GB RAM. same -ngl 99 flags. same 48k context. unsloth gemma-4-12b-it-Q4_K_M.gguf → 15 tok/sec @ 48k ctx unsloth gemma-4-12B-it-qat-UD-Q4_K_XL.gguf → 32 tok/sec @ 48k ctx → 26 tok/sec @ 64k ctx 64k context is a big deal. Hermes 3 agent requires 64k minimum to run. you're now getting full hermes compatible context on a budget consumer GPU at 26 tok/sec locally. 2.1x faster on identical hardware. and here's the part that breaks your brain: the QAT-UD-Q4_K_XL is actually SMALLER than the Q4_K_M "XL" why? QAT = Quantization Aware Training Google didn't train the model first and compress it later they trained it to be quantized from day one the weights already know how to survive low precision that's why you get more quality per byte llamacpp flags: -m gemma-4-12B-it-qat-UD-Q4_K_XL.gguf -cnv -ngl 99 -c 48000 -v fits in 8GB VRAM clean. no API. no cloud. no subscription. and this isn't even the MTP variant yet Gemma-4-E2B QAT runs on 3GB RAM, E4B on 5GB, 12B on 7GB, 26-A4B on 15GB and 31B on 18GB. I have benchmarked the 26b and 31b qat as well on a single RTX 4090, checkout the comments for details. If you have a 6GB or 8GB VRAM GPU, post your numbers. more benchmarks and configs coming soon

Alok

259,993 görüntüleme • 2 ay önce

State of AI compute 2026: my conversation with stephen balaban of Lambda on the neocloud boom, data centers, GPUs and what's ahead 00:00 — Cold open 01:21 — Why GPU compute was never a commodity 02:45 — The H100 price index and what it gets wrong 04:02 — The real moat: technology or financing? 05:57 — Winner-take-all, or room for many neoclouds 06:48 — Are we overbuilding or underbuilding AI compute? 09:26 — What if AI gets 10x more compute-efficient? 10:44 — The real bottleneck: land, power, and shell 11:38 — The backlash against data centers — and the misinformation 15:00 — Opening the hood: from photons to tokens 17:11 — Extracting more value from the same chip 19:26 — Frontier inference and distributed training, explained 23:26 — What actually drives compute cost 25:21 — Lambda's chip stack and the NVIDIA relationship 26:17 — A multi-silicon world? CUDA, CUDNN, and NVIDIA's real moat 28:59 — Networking, storage, and the one-click cluster 34:46 — Renting vs. owning, and full vertical integration 36:24 — How global is Lambda? Does location still matter? 38:44 — The financing stack: off-take agreements, SPVs, and credit 41:16 — Why a 2023 GPU leases for more today 42:36 — A futures market for compute? 43:54 — Origin story: facial recognition, Perceptio, and Apple 47:03 — The Lambda hat and Dream Scope 48:59 — The $60K bet that became a cloud business 52:00 — Holding the team together through the hard times 54:30 — Bringing on a new CEO; Stephen as CTO 57:33 — Matching xAI on high-velocity deployment 59:29 — "AI won't write software — it will become the software" 01:01:30 — Neural software vs. vibe coding 01:04:25 — Do agents change the compute layer 01:06:14 — Self-assembling software inside Lambda 01:08:18 — Gigawatt-scale AI factories 01:08:57 — One person, one GPU 01:12:04 — Hot takes: overrated and underrated in AI

Matt Turck

70,916 görüntüleme • 1 ay önce

Chamath Palihapitiya just dropped the number that explains the entire AI infrastructure trade (Save this). A gigawatt of compute now costs $100 billion and when he started his Arizona data center project it was $4 to $5 billion, it has gone up 20x in a single investment cycle. The implication is not just that AI infrastructure is expensive but rather that the capital barrier to owning meaningful compute has become so high that only a handful of entities in the world can actually build it and the companies who got there early are sitting on what may be the most durable pricing power in the history of the technology industry. This is the neocloud trade. The neocloud market, purpose-built GPU cloud providers like CoreWeave, Nebius, and Lambda Labs was worth $35 billion in 2026 and is projected to reach $236 billion by 2031, compounding at 46% annually. For context, that is faster growth than cloud computing itself posted in its first decade. The reason is very simple, hyperscalers like AWS, Azure, and Google are building for everything, storage, databases, enterprise software, networking and their GPU pricing reflects the overhead of that full-stack infrastructure. Neoclouds build for one thing only, AI compute. The result is a 60% to 85% cost advantage on the same Nvidia silicon, bare metal H100s at $0.78 to $2.79 per GPU-hour on a neocloud versus $3.43 to $5.07 per GPU-hour on a hyperscaler. That spread does not close as AI demand scales but rather it widens, because hyperscalers have to amortize legacy infrastructure and margin expectations that neoclouds do not carry. Gartner projects that by 2030, neoclouds will capture 20% of the $267 billion AI cloud market, and Vultr's own analysis says at least 80% of GPU market share by end of 2026 will be held by a small group of scaled neocloud providers. Now zoom into Nebius specifically, because it is the most interesting publicly traded proxy for this trade. Nebius is the infrastructure arm of the former Yandex Russia's equivalent of Google rebuilt from the ground up after Russia's invasion of Ukraine by Arkady Volozh and relisted on Nasdaq in October 2024. The team that built it already knew how to run internet-scale infrastructure at the lowest possible cost, which is exactly the operational DNA a neocloud requires. In Q1 2026, Nebius reported revenue of $399 million and already generating serious cash on a young business with revenue growing nearly eightfold year-over-year. Then in March 2026, Meta signed a five-year infrastructure agreement with Nebius worth up to $27 billion, $12 billion in committed dedicated GPU capacity deployments beginning early 2027, plus up to $15 billion more tied to Meta purchasing Nebius's unsold third-party capacity. The deal will be executed on one of the first large-scale deployments of Nvidia's Vera Rubin platform, the next-generation architecture after Blackwell making Nebius one of a tiny number of operators in the world with confirmed priority access to the most advanced AI hardware available. Following the contract, Nebius guided to $7 to $9 billion in annualized recurring revenue for 2026 representing 540% year-over-year growth. Chamath Palihapitiya point about the $100 billion capital moat is the bear case for new entrants and the bull case for incumbents. No one can afford to build the next CoreWeave or Nebius from scratch at current hardware and power costs. The companies that are already built, already contracted, and already deploying Nvidia's latest silicon have a moat that compounds with every GPU generation cycle because they get allocations first, they deploy fastest, and their customers re-sign rather than wait for a new operator that does not yet exist. Come join Milk Road Pro for our full breakdown, the complete neocloud competitive landscape, how to think about Nebius's valuation versus CoreWeave and AI entire thesis. Link below.

Milk Road AI

139,047 görüntüleme • 2 ay önce

six months ago this wasn't happening on 8gb vram. running unsloth's Q4_K_XL quant of gemma 4 26b-a4b-it-qat, a sparse MoE model with only 4b active params on a single rtx 4060 laptop gpu, 8gb vram, 20+ tok/s decode. no cloud, no api, no offload hacks. just a gaming laptop on battery. what makes it fit: google's QAT (quantization aware training), plus MTP (multi token prediction) support in the latest llama.cpp builds. that combo is the single biggest unlock for local inference on low vram. rtx 3060, rtx 3070, gtx 1070, gtx 1080, rtx 4050, rtx 4060, rtx 5050, rtx 5060 — any 6-8gb consumer gpu, old or new — this model runs on it. world cup season, so i told it to build a soccer themed flappy bird clone. one shot, zero iteration, fully playable. six months ago an 8gb model could barely clone vanilla flappy bird. now it's shipping a themed game from a sparse MoE model running locally on a laptop battery. inference benchmarks: - decode throughput: 30 tok/s - context: 64k. this is the real unlock. 64k ctx is what makes a hermes agent loop viable locally on this model, not just single-turn chat. llama.cpp flags: -m gemma-4-26B-A4B-it-qat-UD-Q4_K_XL.gguf -c 64000 -cmoe --port 8080 game's deployed on my own site, built and shipped end to end with open source llm, zero closed source api dependency in the pipeline. link in the description. gguf weights on huggingface, link in the comments. pull it down, run it on whatever 8gb card is sitting in your rig. try the game and tell me your score and what you want in v2. local llms on consumer gpus stopped being a meme.

Alok

60,866 görüntüleme • 1 ay önce

Dear ICP community, the Internet Computer has now been running strong for 5 years 👏👏👏 Here is a celebratory preview of ICP "cloud engines," the sovereign frontier cloud technology the network shall soon provide from Main points: — Cloud engines enable anyone to spin up their own sovereign frontier cloud. The technology involves an extraordinary inventive step, in which cloud is created from a mathematically secure network of nodes. The nodes run as part of the Internet Computer network ( but are selected and configured by the cloud engine's owner. — The frontier cloud provided by engines is strongly focused on enabling AI agents to build and update online applications and services for us. The world is changing fast, and nearly all new online apps and services are already being built with the help of AI, and thus cloud engines target the future of cloud. — Software hosted on cloud engines is tamperproof, which means that it is immune to infrastructure hacks, because it runs inside a mathematically secure network protocol, rather than on computers directly. This means that AI agents, and those building with them, don't need to have a security team in the loop, or to trust someone else's security team. This is crucial, because in the future, non technical people will demand the freedom to build with full automation — where they just need to issue instructions to AI about what to build, and don't need to worry about anything or anyone else. Of course, apps and services running on engines are also vastly safer from the new breed of hacker being enabled by frontier AI. (The cloud engines themselves are also "tamperproof." Even if a hacker gains physical access to some portion of a cloud engine's nodes, and can make arbitrary changes, the computations and data of the hosted apps and services cannot be corrupted or interrupted so long as the network's fault bounds aren't exceeded. The recent hack of Vercel, a major cloud platform, which gave hackers access to the apps it hosted, provides additional perspective on the importance of this advantage.) — Software hosted on cloud engines is guaranteed to run, so long as a sufficient number of the engine's nodes are running. This means that AI can build applications and services without the need to have a human systems admin team constantly tinkering with the underlying platform to keep it running, which is again crucial, because in the future, non technical people will expect the freedom to use AI to build without the support of others. — New frontier programming language technology, in the form of the Motoko language developed by Caffeine Labs, leverages seminal "orthogonal persistence" technology that unifies program logic and data to deliver further unlocks for AI (Motoko is the first computer language being developed that targets agents that are writing software rather than humans engineers per se). Nowadays, AI can build and update production apps at a prodigious rate, even at the speed of conversation. But it can also make mistakes, and there's a risk that an update it creates might be "lossy" in the sense it causes some transformed data to be lost. Again, in this new world, it's both undesirable and impractical for everyone to have to have a systems admin team on-hand to detect lossy updates and roll them back, but Motoko provides a solution: it can detect new software updates are lossy before they are applied, reducing potentially catastrophic errors by AI to harmless coding retries. — Software hosted on cloud engines is "serverless" but unlike traditional serverless software, directly it directly incorporates data through "orthogonal persistence." Another key purpose is simplify backend software logic and fuel the modeling power of AI by increasing abstraction (sorry for the technical language!!!). Put simply, this enables AI to produce more sophisticated backends, faster, and at dramatically lower costs, as measured by the number AI API tokens consumed during coding. (Tip for the technical: orthogonal persistence is a new paradigm where "the program is the database," and data lives inside program variables, which is possible because it's as if hosted software runs forever in persistent memory). — An expanding database of skills at shall make it possible to develop and directly deploy apps and services to your cloud engines directly from Claude Code, Perplexity, Codex and other AI platforms. Further, your account on can be connected, so that new apps and updates created through conversation automatically appear hosted from your cloud engine. In the future, R&D is going to be very seamless. You converse with AI, and your secure and unstoppable apps or services are created or updated. Cloud engines are designed to directly support this "self-writing cloud" future where we can work hands-free. — Tech sovereignty is becoming a huge issue worldwide, with governments and corporations seeking to create sovereign tech stacks owing to geopolitical tensions. Increasingly, people are realizing that tech provided by foreign nations can come with hidden backdoors and kills switches, from the base platform, right up through hosted apps and services. ICP technology is open source, and those building on ICP using AI own their own source code. When you have the source code, you can verify that there are no backdoors, and when you own the source code thanks to AI, you can update it at will, freeing you from vendor lock-in. But cloud engines take sovereignty much further... — You create a cloud engine by selecting the nodes that will be combined. You can choose the class of nodes used, and their number, but more importantly, you can choose who operates the nodes, and where they are located. Almost any configuration is possible, because the Internet Computer scales the security privileges afforded to hosted software within the network according to configuration (software hosted on cloud engines can directly interoperate with software on other engines and traditional subnets, but base restrictions are applied according to security rules). A cloud engine can be created within a region such as Europe, to comply with regs such as GDPR, or completely within a sovereign state like Switzerland or Pakistan. But cloud engines go further still... — Sovereignty is also about freedom from vendor lock-in. Cloud engines are essentially ICP (Internet Computer Protocol) network configurations, and this means the underlying compute nodes they combine can be swapped out without interrupting their hosted apps and services. This is a big deal. In addition, cloud engines now support nodes that are instances running on Big Tech's clouds, in addition to nodes that are dedicated specialized hardware, as per the Gen I and Gen II nodes that dominate the Internet Computer today. For example, it is possible to have an engine running across different AWS data centers, say, and then reconfigure the engine to run across a mixture of AWS, Google, Azure and Hetzner for even more resilience, without the users of hosted apps and services noticing a thing. That's true freedom. — Sovereign AI is becoming increasingly important too, and cloud engines allow special "AI nodes" to be added to them, so that hosted software can perform inference on hardware provisioned by the owner from a location the owner has selected. Even though the AI nodes are only accessible within the cloud engine, they can still benefit from the forthcoming Internet Intelligence Gateway (IG), which will make it possible to validate inference performed on key frontier open weights LLMs, even when the inference is performed on completely independent AI clouds. When the results of inference are received, this technology can verify that neither the prompt+context (input) nor the inference result (output) have been modified, and that the results were produced by the precise LLM expected. This ensures that AI clouds don't cheat by running inference on cheaper models than are being paid for, and bad actors aren't modifying the inputs or outputs to surreptitiously insert advertising into results, say, or change facts, or insert malware when code is being generated. What's super cool about this technology is the cost of the verification is scalable. A very valuable additional security can be achieved with only 1-2% of extra cost. — Scaling apps and services when they hit capacity limits is another thorny problem that cloud engines help the world address. Engines make scaling possible without rewriting or reconfiguring software. The query workload capacity of hosted software can be horizontally scaled simply by adding new nodes to an engine, and nodes can also be added in geographical proximity to demand. Meanwhile, update workload capacity can first be scaled-up by swapping an engine's nodes out for the next class up, and then when no larger class of node is available, horizontally scaled-out by "splitting" the engine into two, which doubles available capacity. (Technical tip: horizontally scaling update capacity by splitting engines requires multi-canister architectures). — For those who have been following how Caffeine builds apps that can efficiently store large numbers of files, I should mention that apps built on cloud engines will also support the new ICP Blob Storage cloud network (since cloud engines currently have up to about 3 TB of memory, which apps storing large amounts of files can easily exceed). We are also working on allowing blob storage nodes to be added to cloud engines, to enable sovereign mass blob storage within an engine, similarly to how AI nodes can be added currently. — Lastly, but certainly not least, I should mention that cloud engines are multi-blockchain capable, and ready for digital assets, thanks to the clever math at their core. For example, an e-commerce service built on a cloud engine can securely accept and custody stablecoin payments, or a multi-chain DEX could be hosted. Further, engines can support software autonomy (software orchestrated and controlled by other autonomous software, in a decentralized way) and can themselves be orchestrated by SNS technology, and thus run autonomously too. Today, though, the focus is on *mainstream* cloud. This year, the cloud industry will generate approximately one trillion dollars in revenue. That number is already huge, but is expected to grow to two trillion dollars by 2030. After years of continuous development, which have seen more than $500m spent on R&D, the Internet Computer network is now tacking directly toward this mainstream cloud market with cloud engine technology. In their first version, cloud engines are not meant to be a cloud panacea. For example, currently they are not ideal for working with big data. You should use something like DataBricks for that. Cloud engines are carefully targeted at enabling AI to produce traditional online applications and services, including SaaS, in a safer and more productive way, which represents a new market segment with tremendous potential. Of course, DFINITY will continue to work relentlessly to push forward ICP's capabilities, so expect further developments. It's worth mentioning that this cloud segment isn't just about creating new apps and services using AI, it's also about replacing legacy systems and apps built on super expensive SaaS services. Caffeine Labs is working to produce technology (Caffeine Snorkel) that can study an enterprise's legacy systems and app built on SaaS, create replacement systems and apps, and migrate the data, while supporting key stakeholders through the process over email and chat, with full automation. Thus the legacy systems and SaaS markets shall also be addressed by cloud engines. Zooming out, and reasoning in a more metaphysical way, we believe, as we always have, that there is room for a new kind of cloud created by mathematical networks, that provides seminal advances in the fields of security and resilience, as well as true sovereignty and freedom from lock-in. That this same technology, with the help of additional technologies like orthogonal persistence and Motoko, enables AI to build for us without the need for so much oversight, and to create more backend sophistication while consuming fewer AI API tokens, enables ICP to bring game-changing advances to the world. Cloud engines will work synergistically with the Intelligence Gateway, which will enable apps and services running on engines to seamlessly leverage AI, wherever that AI is running, while providing verifiability at extremely low cost for open weights frontier models. We believe that cloud engines represent an inflection point in the storied history of the Internet Computer project, and I'm very proud to be sharing the details with you on the network's fifth birthday 💪 I'll be back with more news soon!!

dom | icp

285,376 görüntüleme • 3 ay önce

Voice used to be AI’s forgotten modality - now it's having its big moment: rapid innovation, big funding rounds, major agentic applications My conversation with Neil Zeghidour, top AI researcher in the field (Google DeepMind, Meta, kyutai) and now CEO of Gradium This is a reference episode on all things voice AI 🔥 00:00 Intro 01:21 Voice AI’s big moment, and why we’re still early 03:34 Why voice lagged behind text/image/video 06:06 The convergence era: transformers for every modality 07:40 Beyond Her: always-on assistants, wake words, voice-first devices 11:01 Voice vs text: where voice fits (even for coding) 12:56 Neil’s origin story: from finance to machine learning, with help from Yann LeCun and Soumith Chintala 18:35 Neural codecs (SoundStream): compression as the unlock 22:30 Kyutai: open research, small elite teams, moving fast 31:32 Why big labs haven’t “won” voice AI4 34:01 On-device voice: where it works, why compact models matter 46:37 The last mile: real-world robustness, pronunciation, uptime 41:35 Benchmarking voice: why metrics fail, how they actually test 47:03 Cascades vs speech-to-speech: trade-offs + what’s next 54:05 Hardest frontier: noisy rooms, factories, multi-speaker chaos 1:00:50 New languages + dialects: what transfers, what doesn’t 1:02:54 Hardware & compute: why voice isn’t a 10,000-GPU game 1:07:27 What data do you need to train voice models 1:09:02 Deepfakes + privacy: why watermarking isn’t a solution 1:12:30 Voice + vision: multimodality, screen awareness, video+audio 1:14:43 Voice cloning vs voice design: where the market goes 1:16:32 Paris/Europe AI: talent density, underdog energy, what’s next

Matt Turck

22,949 görüntüleme • 5 ay önce

Google just reported $99 billion in profits it never actually received. Alphabet posted net income of $112.1 billion for a single quarter. Earnings per share came in at $9.11 against a Wall Street estimate of $2.87. That is one of the largest profit quarters any company has ever printed. Yet the stock fell about 7% the same day. When people read past the headline and opened the earnings release, they found the reason sitting in one footnote... $99 billion of that profit came from a line called other income. Alphabet describes it as "primarily the result of net unrealized gains on our equity securities." So Google did not sell anything. It marked up shares it already owned and ran the increase through its income statement. That single line added $77.1 billion to net income after tax. It accounted for $6.26 of the $9.11 in earnings per share. Strip it out and adjusted earnings per share were $2.85. Analysts wanted $2.89. The ACTUAL business missed. Now here is what makes this insane: Most of that $99 billion came from two holdings, SpaceX and Anthropic. SpaceX went public on June 12 at roughly $1.77 trillion, up from about $400 billion a year earlier. Alphabet's stake is worth $94.1 billion, and roughly $80 billion of it sits under sale restrictions. Anthropic went from a $350 billion valuation to $965 billion inside the same quarter. Alphabet's private company holdings were worth about $124.3 billion on June 30, and the vast majority of that is Anthropic. Google cannot sell either position right now. Now trace where that valuation came from: Google started putting money into Anthropic in 2023. A $300 million bet has grown into a $13.3 billion position with commitments of up to $30 billion more. Anthropic committed to buying at least five gigawatts of computing capacity from Google Cloud. Google Cloud revenue then grew 82% to about $24.8 billion, the strongest quarter that business has ever had. That growth is part of the story the market uses to price both companies. And when Anthropic's valuation jumped, Google booked the jump as its OWN profit. Google is the investor, the supplier, and the party deciding what the asset is worth. A tax and accounting consultant named Robert Willens flagged this back in April, pointing out that Alphabet is able to influence the value of one of its own assets. And Alphabet's free cash flow for the quarter was negative $5.9 billion. That is the first negative quarter since Google went public in August 2004. Capital spending hit $44.9 billion. Operating cash flow was $39.1 billion. Capex now eats about 37.5% of every dollar of revenue, the highest share in the company's public life. To fund it, Alphabet has taken on roughly $100 billion of debt this year and raised about $85 billion in a June share sale, its first in more than two decades. This is a company that spent years buying its own stock back. What happens next: Alphabet raised 2026 capital spending guidance to between $195 billion and $205 billion, the second raise in three months. The finance chief told analysts 2027 spending will rise significantly. The company also disclosed $811 billion in contracted future spending commitments as of June, up nearly $500 billion from March. Those commitments are signed contracts that get paid in cash. The profit is an estimate of what a private company might be worth on a given day. Estimates move in both directions. If Anthropic or SpaceX gets repriced downward, the same line that produced the biggest quarter in Google's history runs backwards, and this quarter produced no free cash flow to absorb it. Meta, Microsoft and Amazon are all carrying their own private AI stakes into their own earnings reports. Watch how much of their profit they actually collected in cash...

Ricardo

363,984 görüntüleme • 20 gün önce

New interview: Reiner Pope, co-founder/CEO of MatX A counterintuitive throughput insight: “Low latency means small batch sizes. That is just Little’s law. Memory occupancy in HBM is proportional to batch size. So you can actually fit longer contexts than you could if the latency were larger. Low latency is not just a usability win, it improves throughput.” We get into: • The hybrid SRAM + HBM bet, and why pipeline parallelism finally works • Why sparse MoE drives MatX to “the most interconnect of any announced product” • Why frontier labs are willing to bet on an AI ASIC startup • Memory-bandwidth-efficient attention, numerics, and what MatX publishes (and what it does not) • Why 95% of model-side news is noise for chip design • The biggest challenges ahead 00:00 “We left Google one week before ChatGPT” 00:24 Intro: who is MatX 01:17 Origin story: leaving Google for LLM chips 02:21 GPT-3 and the “too expensive” problem 04:25 Why buy hardware that is not a GPU 05:52 Overcoming the CUDA moat 08:46 Early investors 09:35 The name MatX 09:59 The chip: matrix multiply + hybrid SRAM/HBM 12:11 Why pipeline parallelism finally works 14:22 Reading papers and Google going dark 15:20 Research agenda: attention and numerics 17:06 Five specs and meeting customers where they are 19:24 Why frontier labs are the natural first customer 20:32 Workloads: training, prefill, decode 22:18 Little’s law and the throughput case for low latency 24:29 Interconnect and MoE topology 26:35 Inside the team: 100 people, full stack 28:32 Agentic AI: 95% noise for hardware 30:35 KV cache sizing in an agentic world 32:11 How MatX uses AI for chip design (Verilog + BlueSpec) 34:23 Go to market: proving credibility under NDA 35:12 Porting effort for frontier labs 36:34 Biggest skepticism: manufacturing at gigawatt scale 37:32 Hiring plug Vikram Sekar

Semi Doped

19,439 görüntüleme • 4 ay önce

My friends and I confronted anti-China hawk and Richie Torres advisor Matt Pottinger at an Asia Society event. While Pottinger proposes dismantling China-U.S. relations and separating Taiwan from China (which will end up in WW3), I proposed instead that the U.S. and China are not enemies and should join hands in the Belt and Road Initiative to promote peace and economic development. For this, I was thrown to the ground and choked by security. It's no wonder Ritchie Torres is such an idiot when he's surrounded by Neo-con spooks who get paid to lie. For some background on who Pottinger is: Matt Pottinger has been a significant voice in the concerted effort to derail relations between the United States and China for the past 20 years. He began his career as a journalist in the late 1990s reporting for the Wall Street Journal on Chinese affairs. He then joined the U.S. Marine Corps, fighting three separate deployments in Afghanistan and Iraq between 2007 and 2010. Eventually, he would serve the White House for four years on the National Security Council Staff, including as deputy national security adviser to President Trump from 2019 to 2021, leading the administration’s work as the Senior Director for Asia. In 2017, Pottinger had led a special delegation from the U.S. to the Belt and Road Forum for International Cooperation in Beijing, where President Xi Jinping had proposed building a world of prosperity through the construction of massive infrastructure projects, defeating poverty, and mutual technological advancement. The Belt and Road Initiative has been integrated with the Eurasian Economic Union, the Shanghai Cooperation Organization, and ASEAN, and represents a new platform for economic development among the nations of the Global South in an emerging multipolar world which prioritizes mutual benefit and cooperation. Instead of proposing that the U.S. take the opportunity to join the Belt and Road, which President Xi has made clear is open to the West to join, Pottinger has been a consistent provocateur in rapturing U.S.-China relations. In a Wall Street Journal op-ed published on March 26, 2021, “Beijing Targets American Business,” Pottinger says that American businessmen must realize that the “ideological dimension of the competition” between China and the United States “is inescapable, even central.” Protecting U.S. hegemony and preventing China’s economic growth, “will require America and its allies to consider in every policy we adopt, every bill we introduce, and every partnership that government and industry undertake, whether it increases our collective leverage in this competition or surrenders leverage to a hostile dictatorship in Beijing.” It’s no wonder that Rep. Ritchie Torres is advised on China policy by Pottinger himself. Torres is a member of the House Select Committee on the Strategic Competition Between the United States and the Chinese Communist Party, which Pottinger has briefed and advised on policy. Both Torres and Pottinger believe in “the end of history” as author Francis Fukuyama famously declared at the end of the Cold War. This meant, that the U.S. and its Transatlantic partners would dominate the world and create a unipolar world, in which any potential power that would rise and present an alternative to the “rules-based order” would need to be stopped at all costs. Out of this theory that has guided U.S. policy since the first Bush administration, we have produced endless interventionist wars in Iraq, Afghanistan, Libya, Somalia, Yemen, and Ukraine. Of course, Pottinger has expressed his willingness to extend that our adventurist military policy to Taiwan, to provoke a global war with China. Such a policy is not in the American people’s interest. The U.S. economy is now decimated from 50+ years of deindustrialization and idiotic “Green” policy that has eaten up our once productive industries. China has taken the opposite path and is proposing to extend their prosperity to other nations in Africa, South America, and Asia. Instead of creating more wars to defend our hopeless imperial endeavors, we should return to our forgotten anti-colonial roots, and create a new security and development architecture with China and other willing partners.

Jose Vega — Vote Vega!

184,759 görüntüleme • 2 yıl önce

Qasar Younis was COO of Y Combinator when Sam Altman was its president. Peter Ludwig was an early engineer on Android Automotive at Google. Both are Detroit guys with family roots at General Motors. In 2017, they started Applied Intuition to make a billion machines intelligent. Putting intelligence into a machine is one of the hardest problems in technology. A car, a tractor, and a mining vehicle each have different hardware, different sensors, and lives on the line if the software fails. So instead of building one machine, they built the technology that goes into all of them. Cars, trucks, defense vehicles, mining equipment, combines, robots. Like Nvidia sells chips into everyone else's machines, Applied Intuition sells intelligence. They are the only company in the world that does this across all of these industries. They have raised about a billion dollars. They have never spent any of it. Qasar believes that when we look back 25 years from now, the most important companies in the world will all be physical AI companies. Last week they launched Dana, an agent platform built to get there faster. Enjoy! Qasar Younis Applied Intuition Timestamps: 0:00 Intro 4:08 Why Physical AI Will Be Enormous 5:54 What Applied Intuition Actually Builds 9:25 Timing the Autonomous Revolution 13:59 From Tools to a Full AI Stack 21:40 Why Chatbots Can’t Build Robots 25:18 The Physical AI Data Flywheel 28:19 The Moat Behind Machine Intelligence 30:13 What’s Slowing Physical AI Down? 32:30 Applied Intuition’s Business Model 36:10 Competing in a Massive Market 39:24 The Hardware Renaissance 41:09 Why They Raised $1 Billion 43:49 A World Filled With Intelligent Machines

Business Breakdowns

27,283 görüntüleme • 18 gün önce