Video yükleniyor...

Video Yüklenemedi

Ana Sayfaya Dön

🚀 Fully offline AI TALKING AVATAR IS HERE! 💻 Stack: • Arcee.ai Virtuoso-Lite (with ollama ) • Kokoro-80M (TTS) • Whisper.cpp (ASR) 🎭 Streaming to Unreal Engine Metahuman via Audio2Face 🔥 Running smooth on RTX 4090 Laptop (12GB/16GB) (Excuse the jet engine fans 😅)

44,342 görüntüleme • 1 yıl önce •via X (Twitter)

12 Yorum

nisten - e/acc profil fotoğrafı
nisten - e/acc1 yıl önce

@arcee_ai @ollama @UnrealEngine bruh

qnguyen3 profil fotoğrafı
qnguyen31 yıl önce

@arcee_ai @ollama @UnrealEngine what do you think🙏

Page to Pixel Publishing profil fotoğrafı
Page to Pixel Publishing2 yıl önce

The Art of Flight is a homage to 80s/90s arcade action shmups with a fresh twist on the genre. Pilot multiple ships at the same time to take on oncoming waves of enemies in this fast paced space shooter. Wishlist on Steam today!

Noah Santoni profil fotoğrafı
Noah Santoni1 yıl önce

@arcee_ai @ollama @UnrealEngine this is aaaaaawesome! I am working on my own AI voicebot without the avatar though:

attentionmech profil fotoğrafı
attentionmech1 yıl önce

@arcee_ai @ollama @UnrealEngine are you gonna open-source this?

qnguyen3 profil fotoğrafı
qnguyen31 yıl önce

@arcee_ai @ollama @UnrealEngine yes!

Melvin Carvalho profil fotoğrafı
Melvin Carvalho1 yıl önce

@arcee_ai @ollama @UnrealEngine This is excellent! What kokoro blend are you using, or just the default?

qnguyen3 profil fotoğrafı
qnguyen31 yıl önce

@arcee_ai @ollama @UnrealEngine thank you! i am using everything default here

Pratyay Banerjee (নীল)  profil fotoğrafı
Pratyay Banerjee (নীল) 1 yıl önce

@arcee_ai @ollama @UnrealEngine 🔥

llmstock.com mc profil fotoğrafı
llmstock.com mc1 yıl önce

@arcee_ai @ollama @UnrealEngine opensourced?

Ivan Fioravanti ᯅ profil fotoğrafı
Ivan Fioravanti ᯅ1 yıl önce

@arcee_ai @ollama @UnrealEngine WOW, oh WOW! Great job!

qnguyen3 profil fotoğrafı
qnguyen31 yıl önce

@arcee_ai @ollama @UnrealEngine thank you🤗

Benzer Videolar

I ditched Unreal for AI. Kaiju Engine is a completely AI written replacement. To prove it was ready for production, I tested it by recreating and porting my old game, Firefall, into Kaiju, purely from the Steam download, no source. And it worked. Real game code is messy, full of difficult details and compromises. If our engine could handle a commercially released game, it proves you can ship a game with it. We're developing 2 orders of magnitude faster (93x) by lines of code and features. We have fewer bugs, and iteration is super fast. Build times have gone from 30 mins (Unreal) to 36 seconds. Feature rate is through the roof, days instead of months. Our original engine for Firefall took 2-3 years and cost over 5M to build (adjusted). Our new engine was written in 5 months for a couple grand. We built only what we needed, with none of the Unreal bloat. And we added some features: - Modern meshlet based rendering with PBR. - Vis Buffer/Forward+ with clustered lighting. - Id Tech 5 style megatexture streaming. - Planetary sized renderer (entire solar systems possible). - Seamless flight from ground to orbit and space. - No loading screens. - Procedural planet generation with plate simulation. - Client/Server at all times. - Ozz for animation. Jolt for Physics. - Companion Blender plug-in for AI directed asset exports. - 140fps currently, 200-300 projected after optimizations. You can see Firefall's assets and levels ported over into our engine in the video. Texture resolution and pop-in are limitations of the original game (high rez CDN based textures were lost when game went offline, fog hid the original's pop-in). Our meshlet renderer will be able to do much better with LOD and already supports high rez textures. I used Grok/Codex/Claude to tag team the code. This is not just a boon for indie games, it's a real game changer for game preservation. The reaction to seeing this 10 year old discontinued game revived is very emotional for Firefall fans. But we're not just preserving Firefall, we're going beyond, creating an original game that is the spiritual successor, Em-8ER. Gliding, jumpjets... it's all coming back. Moving past Unreal let us move much faster, without the bloat, and with better performance. If you want to signup to follow news on the game, sign up is free at

Grummz

189,489 görüntüleme • 23 saat önce

World Model is trending— let's revisit our HunyuanWorld journey. We’ve been pioneering open-source 3D world generation in the past two months, and this ride’s only getting started. 🌍 📅 July: HunyuanWorld 1.0 📌 First open-source 3D world model compatible with CG pipelines (Unity/Unreal/Blender) 📌 Hit 2K+ GitHub stars in just two months ⭐—thank you for the love! 📅 August: 1.0-Lite 📌Same top-tier quality, running on consumer GPUs! 📅 September: 1.0-Voyager 📌 Direct 3D output + world memory—taking exploration further! Seamlessly integrated into CG pipelines with layered 3D modeling (assets, terrain, skybox) and fully open-sourced.. we’re fully committed to building open-source spatial intelligence for all! 🚀 💡 Why it matters? ✅ Seamless CG Pipeline Integration: Export generated 3D scenes as standard mesh formats, effortlessly integrating into industry-standard tools like Blender, Unity, and Unreal Engine for direct editing, animation, and physical simulation. ✅ Hierarchical Scene Editing: Deconstruct scenes into semantic layers (sky, background, foreground objects) via instance recognition and layer decomposition, allowing for atomic-level control—independently modify, relocate, or replace objects without rebuilding the entire world. Project page: Github: Amazing creations by Stijn Spanhove camenduru GENEL | AIを用いた動画制作 apolinario 🌐 とりにく Directive Creator 🪥 👇 #AI #3DGeneration #OpenSource #WorldModels #Hunyuan3D #HunyuanWorld

Tencent HY

20,178 görüntüleme • 11 ay önce

That's sick! 🤯 Genesis AI simulates robots playing yo-yo! 🪀 Genesis AI just open-sourced Genesis World 1.0, and it might be one of the most important infrastructure releases in robotics this year. Robotics is still bottlenecked by the 1× speed of the physical world. Every model needs to be tested on real hardware, slowly, expensively, with limited coverage. Genesis World 1.0 from Genesis AI flips that equation: One hour in reality becomes 100 days in simulation. That turns a wall-clock bottleneck into a compute problem. And compute problems are solvable. The technical stack they rebuilt from scratch is serious: → GPU-accelerated cross-platform compiler via Quadrants, 10x faster launch time and up to 4.6x runtime vs the initial Genesis release → Penetration-free multi-physics contact solvers, the thing that makes simulation actually trustworthy → Unified rigid AND deformable physics in a single engine → Nyx, a high-performance path-traced rendering engine purpose-built for physical AI The sim-to-real gap has historically been the graveyard of robotics research. Policies that work beautifully in simulation fall apart on real hardware. Genesis World 1.0 is a direct attack on that problem. And it's fully open-source. The companies that master simulation infrastructure will train better robots faster than anyone else. Find it here: Genesis World 1.0: Quadrants: Nyx: Theophile Gervet, Zhou Xian congrats! 👏🏼 ~~ ♻️ Join the weekly robotics newsletter, and never miss any news →

Lukas Ziegler

56,807 görüntüleme • 2 ay önce

I freaked out when my WiFi router suddenly died. then realized my autonomous Hermes agent is running fully local, nothing stopped. Hermes Agent + Gemma 4 26B A4B QAT MoE, 100% local on my laptop, building my side projects while I scroll my phone zero API calls. zero cost. 100% private. fully offline. This might be the most satisfying thing I’ve watched in a while. last post: showed Hermes + local Gemma 4 26B pull off backtest a trading strategy. this time I asked it to develop something i'd use myself everyday: # A full unpacked extension with: - React side panel UI - Local llama.cpp backend (offline AI) - Live tab sync + status tracking - Auto context extraction via Readability.js Vision on Demand → captures viewport screenshots as compressed JPEGs Deterministic action system -> model outputs tokens -> directly controls page scrolling It planned everything first. Then started executing step by step. all i did was say 'ok'. only once. # What’s wild: - It reports back after every phase - Auto compresses context when nearing limits - Actualy, stays on track llama.cpp flags: -m gemma-4-26B-A4B-it-qat-UD-Q4_K_XL.gguf -c 64000 --cache-type-k q8_0 --cache-type-v q8_0 --port 8080 # Performance on a single NVIDIA RTX 4060 (8GB VRAM) + 16 GB DDR4 RAM Gaming Laptop: - 300 tokens/sec prefill - 25+ tokens/sec decode More than usable for real dev workflows. This isn’t AI demo territory anymore. This is autonomous local software actually building things.

Alok

56,428 görüntüleme • 1 ay önce

THIS BUILDER JUST DROPPED 256GB OF RAM INTO A MONSTER THREADRIPPER WORKSTATION TO RUN UNFILTERED LOCAL AI Imagine trying to fit a computer setup into a chassis that is basically the size of a mini fridge. That is the Corsair 1000D tower. The builder crammed an ASUS Pro WS WRX80E-SAGE motherboard inside and slapped a 64-core AMD Threadripper PRO 5995WX right into the socket Why go this heavy? Simple. To make sure local open weights models don't instantly choke standard desktop hardware during heavy reasoning tasks Then things get downright ridiculous with the memory configuration. He unboxes eight separate Kingston DDR4 modules, pinning a massive 256 gigabytes of system RAM directly to the board. Having that kind of local memory headroom is an absolute necessity if a team wants to handle massive datasets or run dense training loops without constantly swapping data to the storage drives Speaking of storage, the system relies on two lightning fast 2TB Samsung 990 PRO NVMe drives For the graphics pipeline, he drops in a top-tier ASUS ROG RTX 4090 boasting 24 gigabytes of VRAM. That is pretty much the gold standard right now if you want to run quick local inference cycles and completely stop paying corporate cloud token fees to OpenAI or Anthropic Powering this whole grid requires a monstrous ASUS ROG 1600W Thor Gen 2 power supply. And to prevent the entire workstation from turning into a space heater under full load, the builder went all out with a 360mm AIO liquid cooling setup and an insane cluster of sixteen Lian Li SL-Infinity RGB fans It looks incredibly flashy, probably sounds like a jet engine when the cores push maximum load What to buy for local AI? => my guide below Bookmark this so you don't lose it

beamnxw ./

47,366 görüntüleme • 1 ay önce

Deepseek V4 Flash 0731 (Q2) - 12 tokens/sec - Single RTX 4090 - 650+ tokens/sec prefill - 250k context - no kv cache quantization! DeepSeek just dropped the official V4 Flash 0731 two days ago with a massive agent capabilities upgrade. The official benchmarks are literally crushing their own V4-Pro-Preview on agentic tasks like Terminal Bench 2.1 and DeepSWE. Unsloth AI said they couldn't wait to bring it to local devices, and they delivered. If you thought my 118B Poolside Laguna S 2.1 MoE run last week on a single GPU was wild, hold onto your hardware. I just successfully ran Unsloth’s brand new 91GB DeepSeek-V4-Flash-0731 (UD-IQ2_M) GGUF entirely locally. And I pushed it to a mind-bending 250,000 context window. The VRAM ceiling is an illusion if you know how to optimize llama.cpp. Here are the benchmarks and the cheat codes to run a local frontier class model yourself. For the hardware and setup, I used a single NVIDIA RTX 4090 (24GB VRAM) hooked up via a PCIe 4 bus, running Ubuntu 22.04 LTS and CUDA 13.0. You don't need a massive enterprise server for this, if you have more than 80 GB of standard DDR4 RAM and a 24GB card like an RTX 3090 or 4090, you can run this exact stack yourself. All benchmarks were run using a massive 28k token prompt to truly stress test the prefill limits. no kv cache quantization THE BENCHMARKS (Scaling Context): # 80k Context (Baseline: -b 2048 -ub 2048): Prefill: 465.43 t/s | Decode: 13.00 t/s | VRAM: 22.87 GB # 80k Context (Optimized: -b 4096 -ub 4096): Prefill: 643.15 t/s | Decode: 12.20 t/s | VRAM: 23.00 GB (Notice how doubling the batch flags spiked my prefill throughput by nearly 200 t/s with almost zero VRAM penalty) # 180k Context (-b 4096 -ub 4096): Prefill: 629.18 t/s | Decode: 11.92 t/s | VRAM: 23.40 GB # 250k Context MAXIMUM (-b 4096 -ub 4096): Prefill: 619.02 t/s | Decode: 11.54 t/s | VRAM: 23.40 GB # THE SECRET SAUCE (Why this works): Unsloth’s UD-IQ2_M quant is ~91GB across 3 files. Since I only have 24GB of VRAM, the PCIe 4 bus and system RAM have to do the heavy lifting. The magic bullet is the --no-mmap flag. By completely bypassing OS disk paging, I forced llama.cpp to load the massive model weights directly into the system RAM upfront. Combined with Flash Attention (-fa on) and exactly 12 CPU threads (--threads 12), I maintained an incredibly stable 11.5+ tokens/sec decode speed even at a quarter million token context. # THE EXACT COMMAND: ./build/bin/llama-server -m /workspace/models/DeepSeek-V4-Flash-0731-UD-IQ2_M-00001-of-00003.gguf -c 250000 -fa on --port 8080 --threads 12 -b 4096 -ub 4096 --no-mmap -v Local conversational and agentic coding AI is fully here. You don’t need an API or an H100 cluster. Qwen 3.8 27b drops next week making the 24GB VRAM tier even more worthwhile. What does your current local AI rig look like, and what's the craziest model you've managed to squeeze into it? Official huggingface GGUF links from Unsloth and performance graphs are dropped in the replies below!

Alok

44,971 görüntüleme • 13 gün önce

$FAME and #AICON Launcher are officially LIVE on Base!🔥 A chance to be part of something BIG–Participate in the #AICON revolution and experience the top-tier seamless experience with #FameAI 🔗 Time to celebrate this milestone together! 🌟 ----------------------------------------------------------- 🚨 Early Access to Fame AI-CON Launcher is Now Available on BASE! 🚨 With FameAI's #AICON Launcher, you can: ✨ Create your human-like AICON ✨ Publish your AI-CON's token on the market $FAME will be tokenized as the first tradable #AICON on our platform! ----------------------------------------------------------- 🔥Early access will be available for selected $FMC stakers🔥 Stakers can enjoy full features being rolled out in the coming weeks. Next up, opening to the public! $FMC stakers, stake now if you haven't already! 👉 ----------------------------------------------------------- 🟣 How Does #FameAI's #AICON Launcher Work? 💻 No coding required! Anyone can create their own #AICON 🔗 Launch your #AICON token—graduated AI-CON can trade on Base Uniswap v2 ⚙️ Fully customizable and highly autonomous How to acquire $FMC on Base ----------------------------------------------------------- 🟣 $FAME Token and FAME Framework This is built to simulate human-like interactions on social platforms. 🎨 Generates content: Images, text, videos 🤖 Reflects personality, knowledge, and mood 💰 Powered by $FAME token ✨ CA: 0xB8e23ab4A1762Fe8dABb844EcC66FEEE3725c480 🔗 Read more here 👉 ----------------------------------------------------------- 🟣 Fame Studio Fame Studio lets you: ✅ Create hyper-realistic human identities ✅ Generate images, audio, music, and even videos V2 Update Coming Soon– Expanded features for more engaging content—perfect for creators and innovators! #AIAgents $FMC #FameAI $FAME ----------------------------------------------------------- 🟣 #FameAI Roadmap for Q1 2025 🚀 Skill Marketplace: Co-create #AICON, utilizing $FMC as the native currency 🧠 AI-Agent Learning: Train agents to learn specific tones, personalities, and behaviors with one click 🎥 Live Streaming: Superhuman-like models for business or lifestyle-focused live streams

Fame AI | The Home of AI Agent 2.0 - AI-CON

23,775 görüntüleme • 1 yıl önce

If you are running local LLMs without N-gram speculative decoding, you are wasting massive amounts of compute. Whether your AI is editing a document, outputting structured JSON, or rewriting boilerplate templates, a huge chunk of the text it generates is highly repetitive or already exists right there in the prompt. Standard decoding wastes expensive GPU compute cycles "re thinking" every single token. By adding one hidden flag in llama.cpp, you can instantly fast forward through the repetition. Zero draft models. Zero extra VRAM. And virtually zero compute overhead. Google Colab hands you an enterprise grade NVIDIA Tesla T4 GPU with 16GB of VRAM for free. It’s the perfect Ubuntu Linux sandbox to build a bleeding edge inference engine from scratch. Recently, I showed you how to double your local speeds using MTP (Multi Token Prediction). But MTP requires a secondary neural network draft model. That eats into your precious VRAM (slightly though) and burns extra compute for every guess it makes. N-gram Speculative Decoding gives you a massive speed boost for exactly 0 memory cost and minimal compute. And it's faster than MTP when it works. Here is how it actually works under the hood: Standard autoregressive decoding is slow because it predicts one token at a time. If you ask an agent to format a long JSON object or update one line in an HTML file, it runs heavy matrix multiplications to calculate the probability of every single bracket, space, and letter from scratch. N-gram changes the game. It acts as a lightweight caching system. Instead of running heavy neural network math to guess the next word, it uses a simple hash table. Whenever the LLM starts outputting a sequence of tokens that already exists anywhere in its context window, N-gram instantly recognizes the pattern. Because it is just doing lightning fast string matching, the compute cost is practically zero. It "fast forwards" through the text, drafting the boilerplate instantly from memory, and the main model just verifies it in parallel. Pure speed. Using quantized GGUFs from Unsloth via HuggingFace, I spun up DeepMind’s massive Gemma 4 26B A4B QAT MoE on a free Colab instance to test this. Just look at the raw benchmark data on code editing task: Without N-gram: [ Prompt: 638.6 t/s | Generation: 45.9 t/s ] With N-gram: [ Prompt: 601.9 t/s | Generation: 107.1 t/s ] Here is the exact llama.cpp CLI command to activate it. Notice we don't even need the --model-draft flag: ./llama-cli -m gemma-4-26B-A4B-it-qat-UD-Q4_K_XL.gguf -cnv -n 6000 -c 12000 -ngl 99 -fa on --spec-type ngram-mod Stop waiting for your GPU to re calculate words it already knows. I’ve built a free, interactive, cell by cell Google Colab notebook that lets you test this live in your browser. You can literally chat with the model and watch the text generation speed absolutely fly on the second turn when you ask it to edit a file. There are additional parameters for ngram-mod that you can tune once you get it working with the single flag. Link to the free Colab Notebook is in the comments below. It walks you through the entire stack: pulling pre built llama.cpp CUDA binaries for Linux, fetching GGUFs from HuggingFace, and spinning up the inference engine with ngram-mod from scratch. Let me know if you have already tried ngram-mod

Alok

31,765 görüntüleme • 1 ay önce

Vibe computing is here. Or, as Matt Deitke @mattdietke, cofounder of Vercept, puts it "the first true AI operating layer" is here. I use it on my Mac, prompt to it, and it does stuff. Like changes system settings, watch how I work and gives suggestions, or copies and pastes from one application into another. I'm highly interested in how AI is changing how we work, so I sat down with Matt for an hour to get a much better look at how he thinks, and what his AI operating layer, Vy, is for. Here's what ChatGPT learned after I fed it the transcript: ++++++++++++ Vercept AI + Vi: Rethinking How We Use Computers 🚀 What It Is Vi is an AI-powered assistant that can control your entire Mac screen like a human would — moving the mouse, typing, clicking, navigating apps. It’s being called the first true AI operating layer — what you dubbed “AI operating system” or “vibe computing.” Unlike traditional assistants (Siri, Copilot, ChatGPT), Vy works across any app — from Descript to Chrome to Slack to Photoshop — and acts on your behalf. 🤯 Game-Changing Capabilities Does anything you can describe: “Unfollow people on X,” “Write a Word doc,” “Summarize my emails,” or “Plan my vacation in a spreadsheet.” Works via screenshots: Interprets your screen visually, just like a human would — no APIs or browser hooks needed. Cross-app workflows: Can copy data from one app to another, or handle complex tasks like “look up 10 Goodreads books, extract data, and fill a spreadsheet.” Understands vague language: Even if you don’t use exact names or phrasing, Vy figures it out. 🧠 Where It’s Going Will evolve to: Run in the background Manage multiple apps and windows Act like a team of virtual assistants Work on Apple Vision Pro and future AR/AI interfaces Long-term vision: Vy becomes a swarm of agents running “24/7 like a digital company” doing real, expert-level work. 💼 For Power Users & Enterprises Strong use cases for: Developers using Cursor or VS Code Researchers summarizing YouTube videos, PDFs, long threads Execs automating emails, calendar, reports Batching tasks, templates, and macros are coming: “Tell Elon X, Y, Z” → will soon run across apps and reuse workflows. 🔐 Privacy & Safety Runs locally, stores nothing permanently, doesn’t send screen data to servers. You control when it’s active. Cept prioritizes on-device execution and temporary-only data. Security-conscious users (like Apple employees) will eventually get fully offline modes. 💵 Business Model Currently 100% free while in early access. Future: Premium plans, pro tools, enterprise deployments. 🧑‍🔬 The Founders A veteran computer vision & AI research team from University of Washington, Allen Institute for AI, and early deep learning work. Includes Ross Girshick, one of the most cited researchers in computer vision. 🔮 The Future Matt sees Vy evolving into: A universal expert-level interface across all digital tools The AI-powered bridge between humans and complex systems (e.g., building robots via simulators, managing workflows, analyzing regulations) A new way to compute, where you just describe your goal and it gets done — quietly, in the background, or visually on screen. Try it at:

Robert Scoble

17,509 görüntüleme • 1 yıl önce

10 free Google AI tools nobody talks about. while everyone's burning $20/mo on chatgpt and claude, google quietly shipped a stack worth $200+/mo. all free. all yours. — 1️⃣ NotebookLM — your second brain upload sources (PDFs, websites, audio, YouTube). it summarizes, builds mind maps, generates quizzes, drafts slide decks, even turns your notes into a podcast you can listen to on a walk. free tier: 100 notebooks, 50 sources each, 50 chats/day, 3 audio overviews/day. replaces: notion AI + perplexity + readwise — 2️⃣ Google AI Studio — the free gemini playground web playground for gemini 3 pro and flash with a free API key. generous limits. paste a 1M-token context window and watch it actually use it. faster than the openai playground and free where openai charges per token. replaces: openai playground + paid API credits — 3️⃣ Gemini CLI — google's open-source terminal agent apache 2.0 licensed. one command (npx @google/gemini-cli) and you've got an agent in your terminal that reads your codebase, runs shell commands, and ships PRs. drop-in claude code alternative. replaces: claude code ($20/mo by default) — 4️⃣ Jules — async coding agent assign jules a github issue. it spins up a cloud VM, clones your repo, writes the plan, makes the changes, opens a PR. free tier: 15 tasks/day, 3 concurrent, runs on gemini flash. replaces: devin ($20/mo+) + cursor agent 5️⃣ Stitch — text → UI → code google's free figma killer. describe an interface, get production-ready HTML/CSS/Tailwind + figma export. march 2026 update added voice canvas, infinite canvas, and MCP integration with cursor. 350 standard + 200 experimental generations/month free. replaces: galileo AI + early-stage figma work — 6️⃣ Gemma 4 — open-weight LLM google's flagship open model. apache 2.0. 2B, 4B, 26B-MoE, and 31B variants. 256K context. runs on ollama with one command. quantized versions run on a 4090 or beefy laptop. replaces: paying for hosted LLM inference — 7️⃣ Illuminate — papers → podcasts paste an arxiv preprint link. illuminate turns dense research papers into a 6-8 min conversation between two AI hosts breaking it down. perfect for commute reading you can't do at a desk. note: still in waitlist for some regions. replaces: snipd + manual research reading — 8️⃣ Learn About (LearnLM) — adaptive AI tutor drop in any topic you're stuck on. highlight a word, click "go deeper," and the interface adapts in real time to your comprehension level. visual explanations, follow-up questions, the works. replaces: paid tutoring on niche topics — 9️⃣ Google Labs FX (ImageFX + Flow + MusicFX) — free imagen, veo, musicLM google labs creative suite. text-to-image (imagen 4), text-to-video (veo via Flow), text-to-music (musicLM). free tier: limited daily generations. the heavy veo 3.1 features are paid (AI Pro $19.99/mo). still worth using for image and music — those stay free. replaces: midjourney + suno (free tier only — runway-level video gen is paid) — 🔟 Google Colab — free GPU notebooks free T4 GPU + 12GB RAM in a browser tab. enough to fine-tune small models, run stable diffusion, prototype agents. the launching pad for half the ML projects on github. replaces: paid cloud GPU rentals — a quick honest note: these tools aren't 1:1 better than the paid versions they replace. but they're decent enough to get most things done — especially if you're not a heavy user or you've got little funds to play with. i've put all 10 in a public github repo (link in comments). follow + turn on post notifications for more useful posts like this 🔔

m0h

11,847 görüntüleme • 2 ay önce

Matthew Gallagher Built a $401M Company in Year One with 2 People. And the tool behind it? Claude Code. This year he's on track for $1.8B. Sam Altman predicted this. It's happening now. The problem? It costs money. API credits stack up. Monthly bills keep growing. Every prompt eats your budget. Every project drains your wallet faster. Until now. Two methods. 99% cheaper. One is completely free. Forever. $0. Not a trial. This video breaks down both step by step. ↓ Let me put this in perspective. $100-$500. That's monthly. That's what you spend. That's $6,000/year on API credits. Just to use a tool you haven't shipped anything with. The $401M guy? Spending $0. Same capability. Shipping weekly. Different cost structure. Different results. Different life. I'm about to hand you his cost structure for free. ↓ Open source vs closed source. Pay attention. Closed source: Claude. GPT-4. Pay per token. Meter always running. Open source: Qwen. Llama. Mistral. Free to download. Free to run. Free forever. No meter. No tokens. No bill. Here's what nobody tells you: 80% of coding tasks? Open source handles them. More than handles them. Writes clean code. Debugs errors. Generates boilerplate. Handles routine work perfectly. You're paying premium prices for tasks that don't need premium intelligence. That's hiring a brain surgeon to put on a bandaid. Smart play: Free models for the 80%. Paid credits for the 20%. That's what the $401M guy does. That's what this video teaches you. Follow Himanshu Kumar for more breakdowns that turn free tools into real businesses. ↓ Method 1: Ollama. Local. Free. Forever. Download it. Pull a model. Point Claude Code at it. Done. No internet needed. No API keys required. No monthly subscription. No token counting ever. No bill. Today. Tomorrow. Ever. Your data never leaves your computer. Complete privacy. Complete freedom. Claude Code thinks it's talking to the cloud. It's talking to your laptop. For $0. The video walks through every step: Every config file. Every variable. Every command. Every click. If you can follow a recipe, you can do this. People who set this up 3 months ago? Saved $300-$1,500 since then. Workflow didn't change one bit. ↓ Hardware you need: 16GB RAM: 7B models run smooth. 32GB RAM: 32B models run comfortable. 64GB + GPU: biggest models available. No GPU? Still works. Just slower. Few extra seconds. That's it. Your $1,500 laptop is sitting there running Chrome and Spotify. Put it to work saving you $200/month instead. Follow Himanshu Kumar for more breakdowns that turn free tools into real businesses. ↓ Method 2: Open Router. Free Cloud. No Hardware. Weak machine? Don't want local setup? This method is for you. Free AI models in the cloud. No download. No hardware. Configure Claude Code to route through Open Router. The config: Base URL: Open Router API. API key: free Open Router key. Default Sonnet: free. Default Opus: free. Default Haiku: free. Small fast model: free. Subagent model: free. Free. Free. Free. Free. Free across the board. Same interface. Same commands. Same workflow. Zero cost. Copy the config from the video. Paste it. Save $200/month. Starting today. Right now. ↓ When to use which: Ollama (local): Best for privacy. Best for offline work. Best for unlimited usage. Best if you have decent hardware. Open Router (cloud): Best for weak machines. Best for instant setup. Best for trying different models. Best if you don't want to manage anything. Both methods: Best for 80% of your daily work. Still use paid Claude for: Complex architecture. Multi-file refactoring. Deep reasoning tasks. The 20% that actually needs it. $20/month instead of $200/month. Same output. 90% less cost. ↓ The math that should make you angry. You (current): $200-$500/month. $2,400-$6,000/year. $7,200-$18,000 over 3 years. You (after this video): $20-$50/month. $240-$600/year. $720-$1,800 over 3 years. Savings over 3 years: $6,480-$16,200. That's a used car. That's seed money. That's 6 months of rent. All from one 25-minute video. All from 15 minutes of configuration. Highest ROI 25 minutes you'll spend this year. ↓ The limitations. I won't lie to you. Open source is not Opus. Not as smart on complex reasoning. Not as good at long-context tasks. Makes more mistakes on nuanced problems. But they are: Free. Capable. Getting better monthly. Good enough for 80% of daily work. Smart cost management isn't being cheap. It's being strategic. Expensive tool when it matters. Free tool when it doesn't. ↓ The one-person billion-dollar company is coming. $401M in year one proved it's possible. The building blocks: AI that codes: Claude Code. Way to run it free: this video. Distribution: the internet. Customers: everyone. Only missing ingredient? Someone who builds. Not reads about building. Not saves posts about building. Not bookmarks videos about building. Builds. Tools are free. Knowledge is free. Opportunity is screaming. You're still "thinking about it." ↓ Your action plan: Tonight: Watch the video. Tomorrow morning: Set up Ollama or Open Router. Tomorrow afternoon: Build something. Anything. This week: Build a second thing. Faster. This month: Charge someone for it. One video. One setup. One weekend. $0 cost. Unlimited potential. Or keep paying $200/month for something you could get free. Keep consuming instead of building. Keep planning instead of shipping. Matthew Gallagher didn't plan a $401M company. He built it. Full video attached. Every method. Every config. Every tradeoff. 25 minutes. Your move. Follow Himanshu Kumar for more breakdowns that turn free tools into real businesses.

Himanshu Kumar

13,599 görüntüleme • 4 ay önce

When Elon Musk beams in virtually for a high-stakes fireside chat with JPMorgan Chase CEO Jamie Dimon, the conversation goes completely out of this world. The discussion was packed with massive milestones—from the bombshell that SpaceX is going public to plans for lunar AI data centers and the urgent need for the Terafab chip revolution. Here is the ultimate breakdown of their discussion: 💵 SpaceX has been self-funding and cash-flow positive for a decade Before the decision to go public, SpaceX didn't actually need to raise money to survive. The company has been cash-flow positive since around 2014–2015, meaning its private equity rounds were exclusively held to provide liquidity for employees and early investors. "We've been positive cash flow for quite a long time, I think, since around 2014-2015. And we've been self-funding. In fact, in our sort of private equity rounds, they actually have not been fundraising rounds. They've been liquidity rounds for investors and employees because we give everyone at the company stock." 🚀 The upcoming capital growth phase requires massive funding The primary trigger for going public now is an unprecedented capital expenditure phase. SpaceX is preparing to deploy an immense constellation of over 100,000 Next-Gen communication satellites and construct massive AI data centers in orbit. "we are embarking on a significant capital growth phase where we're going to put in over probably 100,000 satellites, probably over 100,000 satellites, just for communications... And then we're also doing the AI data centers in space, which is another massive capital endeavor." 📡 Starlink V3 introduces a massive bandwidth breakthrough The custom chips designed by SpaceX for the V3 satellites will completely alter global communications, offering 100 times the bandwidth of the current system and slashing latency in half by operating at a lower altitude. They are so large—the size of a small bus—that Starship is the only rocket on Earth capable of launching them, carrying 50 at a time. "The version three is, depending on how you count it, 10 to 20 times more capable than the version two satellite. And there were three chips that the SpaceX chip design team taped out that are specific to this... Which means it's 100 times more bandwidth than the SpaceX's Starlink system currently on the surface. And also half the latency because the altitude will be about half altitude." 🤖 AI and robots possess an insatiable appetite for data Musk points out that expanding infrastructure into space is vital because future AI and robotic systems will demand an astronomical amount of bandwidth compared to the relatively low data transmission rates of human beings. "And the future with AI and robots is actually going to require a lot more bandwidth than we currently use. Because you can imagine like what's the bandwidth of a human? Peak bandwidth of the human is a few hundred bits per second. But bandwidth of a computer can be a trillion bits a second. So the appetite for bandwidth of AI and robots is going to be enormous." ☀️ Space solves the looming terrestrial power plant crisis Building traditional power plants on Earth faces heavy community resistance. Moving data centers into space unlocks unlimited energy generation via solar power ("star power") without disrupting Earth's environment, tapping into an energy source that accounts for 99.8% of the solar system's mass. "It's increasingly difficult to build power plants on the ground. There are very few people who want a power plant in their backyard... But actually if we go to space, we can go far beyond the electricity generation of both. In fact, this is going to sound kind of crazy. But you could actually increase human energy by a factor of a million and still be using much less than a millionth of the sun's energy." 🌕 The Moon is a 1,000-Terawatt compute launchpad While Mars remains the long-term goal, the Moon is the immediate fast-track location for massive scaling. Because it lacks an atmosphere and has low gravity, SpaceX can use electromagnetic rail guns to shoot AI data centers into deep space from the lunar surface, scaling power to an incredible 1,000 terawatts per year. "I just think that we can build a self-sustaining city on the moon faster than we could do so on Mars. And there's also the potential... you can use an electromagnetic accelerator, a rail gun or mass driver. Basically, you don't need to use rockets to do AI data centers into deep space from the moon... We can do a thousand terawatts or more from the moon." 🪐 Mars is the ultimate "fixer-upper" planet Mars is being targeted as a full-scale terraforming project. Due to its atmosphere and gravity levels, warming up the planet could eventually unlock liquid oceans and allow humans to walk around without spacesuits. "And if you warm up Mars, you could one day make Mars like Earth. And with like liquid oceans and life. And where you could walk outside without a spacesuit type of thing. So Mars is, I call Mars a fixer upper of a planet. But it's got a lot of potential." 🚂 SpaceX is the modern-day Union Pacific Railroad Musk rejects the idea that SpaceX is moving into the hospitality or hotel business for space tourism. Instead, he views the company as a foundational infrastructure provider, comparable to the historic railroads that opened up the American West. "We're kind of like Union Pacific, you know. You know, when they built Union Pacific back in the day, people thought they were crazy. Because like, why are you trying to carry all this cargo and people to California? No one's there. But now California is the biggest state in the country." ♻️ Starship's core disruption is 100% reusability The true holy grail of Starship is full reusability, which drops orbit access costs down to the mere price of fuel. Because it utilizes ultra-cheap liquid oxygen and methane, shipping cargo to space will become more economical than flying cargo across Earth's oceans on an airplane. "The fundamental breakthrough of Starship is that it will be the first orbital rocket that is fully reusable... And the propellant we use for Starship is liquid oxygen and liquid methane, which is the cheapest propellant you could possibly get... which means that you should be able to actually send cargo to space for less than the cost of cargo on an airplane going on a trans-oceanic trip." 🔄 Starship V4 targets hourly launch cadences SpaceX's engineering pipeline is aiming for staggering operational frequencies and massive payloads. While Starship V3 targets 100 tons to orbit, the upcoming V4 variant is designed to carry over 200 tons and launch on an hourly schedule. "Because Starship V3 is aiming to do 100 tons to orbit with full reusability. And then Starship V4 we're aiming for over 200 tons per mission. And then being able to launch every hour." ☁️ Orbital data centers are entirely weather-proof Space-based AI data centers are highly practical because they are simpler to construct than communication satellites. Data is beamed via lasers between satellites, and then beamed to the ground using cloud-penetrating radio frequencies that completely bypass bad weather. "The AI data center would be much simpler by comparison. Because it's really just solar power plus radiator... The connection would happen no matter what the weather is. Because once you connect via the lasers to the Starlink communication constellation, the Starlink communication to the ground uses frequencies that are cloud penetrating." 🇺🇸 The U.S. faces a catastrophic "Zero Memory Fab" crisis A major vulnerability in domestic tech infrastructure is that the U.S. currently manufactures zero high-volume computer memory chips. Even with new facilities arriving online between 2028 and 2030, domestic supply will not match the exponential requirements of AI, which is why Musk is aggressively building the Terafab. "there's not a single high volume computer memory fab in America right now. Zero. There's one being built in Idaho by Micron. But that will not reach volume production until I believe 2028. And there's something being built in New York, but they are in, I think, 29 and 30. And this is a tiny fraction of the memory that's needed... That's why we need to do the Terafab." 🧠 SpaceX will offer proprietary AI chips and software While the orbital data center network will remain an open marketplace capable of running third-party hardware like NVIDIA GPUs, Google TPUs, or Amazon Trainium, SpaceX plans to deploy its own in-house AI chips and software stack in the near future. "So if NVIDIA GPUs can be put on it, Google TPUs can be put on it, Amazon Trainium or any other chips that you want to put on, can be put on. We'll also offer our chips in the future and I think we also want to offer our software, our AI software as well in the future." 🛡️ Starshield handles critical national intelligence Musk emphasizes his deeply pro-American stance, highlighting SpaceX's specialized Starshield division as a crucial backbone for the U.S. military and national intelligence agencies. "We have a division called Starshield which provides military communications. And you know, there's some other stuff that's kind of classified, I guess. We can't be talking about that. But we are helping the Department of War and intelligence part of the government. We're a vital element of that." 👥 Executive retention fuels the mission The core leadership bench at SpaceX is defined by extreme longevity, driven by a deep collective belief in turning science fiction into reality. Top executives like Gwynne Shotwell have remained with Musk for over two decades. "I guess Gwynne was, I think, around the seventh person to join the company. And that was 2002. It's just went to like 24 years. And generally the senior executives at the company, you have a very long tenure. I think Brent Johnson's been, you see, over 15 years... because people really believe in the mission, I think they want to stay and they want to keep building it." ❤️ Character overrides IQ in leadership Reflecting on how he has evolved over 20 years, Musk notes that he has become significantly more laid back. He has also learned that a candidate's moral character and heart are just as vital to a company's success as raw intellectual horsepower. "Well, I think I'm probably more chill than I used to be... And one of the things I've found over time... is that like in terms of like recruiting people to the company and having people work with the company, like their individual abilities and their intellectual capabilities matter a lot, but it also matters if they have a good heart. It's not just about whether somebody has a certain IQ or whatever, but just are they like a good person, that matters a lot."

Ming

60,910 görüntüleme • 2 ay önce

Behind The Scenes In The Vegas Loop: Inside Elon Musk's The Boring Company Bold Bet On Urban Mobility Hey everyone. Tesla Owners Silicon Valley (Tesla Owners Silicon Valley) here. I recently had the chance to go behind the scenes with Steve Davis, President of The Boring Company, for a deep dive into the Vegas Loop in Las Vegas. This wasn’t a quick photo op. It was a full 47-minute immersion: riding through the LED-lit tunnels in a Tesla, visiting active construction sites with Prufrock boring machines, and hearing directly from Steve about what’s working today, and what’s coming next. I’m posting the full long-form video alongside this recap so you can experience it firsthand. But here’s the readable, “what actually matters” story from the tour. From “Traffic Is Soul-Crushing” To A Working Underground Network The Boring Company was founded in 2016, born of a familiar frustration: gridlocked cities that can’t build fast enough, cheap enough, or with minimal disruption. The premise is simple but ambitious: reinvent tunneling to make it practical infrastructure, not a decade-long mega-project. Las Vegas is where that idea is being tested at real scale. Instead of waiting for buses, shuttles, or rail schedules, the Vegas Loop aims to provide point-to-point trips in Teslas, fast, quiet, and emissions-free, connecting major destinations without the chaos of the Strip above. And after seeing it up close, what stands out most is how operational it already is. This isn’t a render. It’s a functioning system handling real demand, in real conditions, with real riders. What It Feels Like: Fast, Weirdly Fun, And Surprisingly Smooth The “Loop experience” is part transit, part sci-fi. The tunnels are lined with shifting LEDs—purples, greens, yellows—that make the ride feel more like entering a venue than commuting. Trips are short and direct. One example Steve shared: LVCC to Encore in about 85 seconds. But the biggest “wait, that just happened” moment on the tour was Full Self-Driving. FSD Underground (And Onto Surface Streets) We rode in a Model Y running Full Self-Driving (Supervised), which navigated the tunnels smoothly and then transitioned back to surface streets without intervention. Steve’s point wasn’t that autonomy is a cool demo; it’s that autonomy is a force multiplier for throughput, consistency, and future scale. Steve Davis: “Full Self-Driving Supervised is live commercially between LVCC and Encore, watch this: zero interventions as it navigates the tunnels and pops out onto surface streets seamlessly.” Right now, they still operate with safety drivers, but the trajectory is clear: as autonomy matures, the system can move more people with tighter headways and less variability than human-driven operations. The Numbers: “Spiky Demand” Is Where This System Wants To Win Vegas isn’t a steady-demand commuter city. It’s a burst-demand city: conventions, games, concerts, and tourist surges. Steve emphasized that this is exactly where the Loop model shines, because you can scale vehicles dynamically without rebuilding an entire transit line. During CES 2026, the Loop moved 90,000+ passengers, peaking at 6,600+ riders per hour, including 22,000+ trips to/from Resorts World, Encore, and Westgate. That’s on top of 3.5M+ total passengers since 2021. Steve Davis: “We’ve hit over 3 million passengers since 2021, and during CES 2026 alone, we shuttled more than 90,000 people, peaking at 6,600 passengers per hour without a hitch.” And beyond the numbers, there’s a secondary effect people don’t always talk about: for many riders, this is their first time in a Tesla, and it’s an unusually positive first impression. The Airport Connection: A Phased Plan With A Very Clear Endgame Connecting the system to Harry Reid International Airport is the crown jewel, and they’re doing it in phases to deliver value quickly while they work through the harder parts. Phase 1 (Live Now) Limited airport rides are already operating via a mix of tunnels and surface streets from existing stations, including Resorts World, Encore, Westgate, and LVCC. They’re doing roughly 50 test rides per day, and Steve noted 100 of ~130 vehicles are already “airport-ready” with transponders. Phase 2 (Next Couple Months) This is where things get meaningfully faster: a 2.2-mile dual tunnel from Westgate to 4744 Paradise Road, eliminating about two miles of surface traffic and stoplights. New stations are planned at Virgin Hotels, The Boring Company’s apartment complex, the former Gordon Biersch site, and Firefly. Fleet expands to 160 vehicles. Steve Davis: “Phase 2 kicks in soon: a 2.2-mile tunnel to Paradise Road, cutting out those surface miles and stoplights.” Phase 3 Extend to 5032 Palo Verde Road near Terminal 1, further removing surface bottlenecks around Tropicana and University Center. Fleet scales to 250–300 vehicles. Phase 4 (The “Holy Grail”) A direct underground station at the terminals, true curb-to-gate simplicity, fully underground. Steve Davis: “Phase 4 is the holy grail: a direct underground station right at the airport terminals.” The Big Build: 68 Miles, 104 Stations, Privately Funded The long-term vision is expansive: 68 miles of tunnels and 104 stations spanning the Strip, downtown, the stadium, and the airport. Core Strip construction begins this fall, with a 2027 target for that major phase, and further expansion into 2028–2029. Steve emphasized something important here: the funding model. These builds are privately funded, and the cost structure is the entire point: build rapidly and avoid “subway economics.” Steve Davis: “68 miles, 104 stations… all privately funded at about $10M per mile, versus billions for subways.” The Real Workhorses: Prufrock Boring Machines Up Close If the Loop is the user experience, Prufrock is the engine underneath it. Seeing Prufrock at an active dig site is hard to describe unless you’ve stood next to one. It’s enormous, loud, and relentlessly practical. The key advantage is that it changes the setup cost: it can launch from the surface without massive open pits, and it’s designed to move fast, with a long-term target of one mile per week. The machine isn’t just digging; it’s built around an integrated approach to lining, pumping, and maintaining the tunnel environment while staying cost-effective. Challenges They’re Solving In Real Time: Groundwater And Permitting One of the most interesting “myth-busting” moments was hearing Steve talk about tunnel conditions. Despite the desert setting, the tunnels are roughly 30 feet below grade, and in many areas, they’re fully submerged in groundwater, sand, clay, caliche, and water management, all part of the daily reality. Steve Davis: “Tunnels are 30 feet down, fully submerged in groundwater, desert myth busted.” They manage leaks through periodic sealing (foam, maintenance cycles) and now operate with stronger compliance processes for water treatment and disposal. The bigger long-term bottleneck, though, isn’t engineering; it’s approvals. Steve noted they need hundreds of permits (600+), and many can take months. Their push is toward a more streamlined, operator-style approval model, closer to how SpaceX is regulated: certify capability and safety, then execute without rearguing every step. Steve Davis: “Permitting’s the bottleneck… we’re advocating for a SpaceX-style operator license.” Fleet Scaling And The “Robovan” Strategy Right now, the fleet is about 130 Teslas, including Model Ys and Cybertrucks, tuned for tight turns and repeated high-frequency operations. The larger goal is to scale up to 1,200 vehicles as the network grows. And that’s where Robovan (high-occupancy, event-optimized vehicles) becomes strategically important. Steve’s framing was refreshingly clear: cars are more efficient for small groups. Robovans win when you can predict surges, like a Raiders game or a Sphere show, and load high-occupancy vehicles in advance. Steve Davis: “Robovans shine when everyone’s going to the same spot… that’s when you put the high occupancy vehicle in.” What’s Next: Suburbs, Regional Links, And Bigger Swing Ideas After the core network is built, they’re looking at suburban expansions (Henderson, Summerlin) via shorter demo segments first, proving utility for pedestrian and vehicle connectivity. And then Steve hinted at the kind of long-range thinking that gets people excited (and skeptical): longer-distance routes, potentially even Hyperloop concepts like Reno connections, if permitting and economics align. Steve Davis: “Suburbs like Henderson and Summerlin next… long-term? Hyperloop to Reno… private funding makes it doable if permitting catches up.” Final Take: Vegas Is Becoming A Live Testbed For A New Kind Of Transit This tour made one thing very clear: The Boring Company isn’t trying to win the “traditional public transit debate.” They’re trying to change the rules of what’s feasible, building faster, cheaper, and with an experience that people actually want to use. Watching FSD glide through the tunnels, seeing Prufrock tearing through the ground, and hearing the phased plan for the airport and Strip expansion straight from Steve… It’s hard not to feel like Vegas is a real-world preview of what mobility can look like when infrastructure is built like technology. Huge thanks to Steve Davis and The Boring Company team for the access and the time. And keep an eye out, I’m posting the full 47-minute video with this recap so you can see the ride, the sites, and the details for yourself. What do you think, would you ride the Loop instead of sitting in Strip traffic?

Tesla Owners Silicon Valley

447,027 görüntüleme • 7 ay önce

$AMD $5 Trillion MC Is Inevitable Long Term👑 This thread will focus more on Inference! 2026 EPYC "Venice" $TSM 2nm to save Large GW Scale Inference by 40% more than Prior Turin gen. Context: EPYC Turin achieves ~$0.001 per million tokens for batch inference vs $0.02-$0.12/ million tokens as I wrote the thread below. Venice is going to lower cost down to $0.0005-$0.0006/Million Tokens. OpenAI spent roughly $20B on Inference and Training, where 80-90% of that was for Inference per Analysts. AKA Renting Compute is Expensive AF! In this thread, I want to focus on why most analysts and investors are underestimating the role EPYC "Venice" and future Gen on overall Data center revenue. And $TSM ramping up 2nm supply early is a confirmation that AMD will be a major buyer long term. I will also link the thread the Gap between AMD Analysts & Reality and 2nm Ramp Thread so you have more comprehensive view of what I'm writing here. Before I go into detail this is my 2026 Projection: AI GPUs: $35-$50B EPYC Data Center: $15B-$17B Client Segment: $12-$13B Gaming: $6B Embedded: $4B-$5B Total Revenue $70-$100B Non-GAAP net income $18B-$25B Non-GAAP EPS $10.97-$15.40 Foward P/E 55x-70x= $603-$1,078 AMD's Analysts are projecting $0 Revenue for MI450 and sluggish EPYC Growth. Meaning, all analysts are either full of 💩 or Sexist, you decide! Analysts are also projecting 0% growth on AMD "Secret Weapon" Chip as $MSFT said we are at significant Windows refresh and upgrade cycle. Do you think TSMC would allocate more 2nm supply to $AMD at $0 MI450 revenue and sluggish EPYC? 1. EPYC is going to be the leader in lowest Inference! Current Turin cost saving is 95% vs $NVDA or 98-99% on Inference cost when you factor in renting Inference compute from Amazon Web Services, Microsoft Azure, or $NVDA Neocloud pets. TSMC claimed: 10-15% higher performance at iso-power, 25-30% lower power at iso-speed, and ~15% higher transistor density compared to 3nm. This reduces operational expenses (energy, cooling) while increasing throughput per chip. EPYC Turin achieves ~$0.001 per million tokens for batch inference (via vLLM on models like Llama 3 70B), driven by high core counts and low hardware costs. EPYC Venice offers ~1.7x overall performance and up to 70% more compute capability per core, with up to 256 cores (512 threads). Enhanced vector/AI instructions and open-source firmware (openSIL) optimize for inference workloads. AMD Incorporates AI Engines (now part of AMD's XDNA) for on-chip acceleration, improving efficiency for low-latency and edge inference. This reduces reliance on discrete GPUs, lowering system complexity and TCO. Venice SKUs are projected at $3,000-$15,000 ($5,000 for 256-core flagship), far below NVIDIA Rubin ($50,000-$90,000) or AMD's own MI450 GPUs ($40,000-$50,000). High memory bandwidth (up to 1.6 TB/s) supports efficient batch inference. Venice is designed exactly for Large customers that want to lower Inference Cost and MI450 Helios is for Customers that want Training at lowest TCO, TDP as well as lower Upfront 1GW scale(Full build $35-$40B vs $NVDA $55B-$80B). 2. Real World Example: OpenAI's 2025 inference spend reached ~$20B, escalating to even higher total compute rental (mostly inference) amid token volume growth(from video generating). By 2026, with usage doubling (consistent with industry trends: token demand grows 2-5x YoY), assume OpenAI processes ~1,800 billion million-tokens annually $NVDA Blackwell at $0.02-$0.12 is $36B(most optimized) Rubin is projected to be at $0.01/million tokens or $18B annual Inference Cost vs $AMD Venice $0.0005/million tokens or $0.9B annual Inference Cost => Massive saving for OpenAI or anyone that are paying 80-90% Annual Bill for Inference compute. In short, it is unsustainable to pay this much rent vs owning for all current AI players for the medium to long term. Rubin excels in low-latency decode (if Groq integration from $20B deal in 2027-2028), but Venice dominates batch (80% of inference by 2030). Actual savings depend on deployment scale (OpenAI's 6GW AMD plans), electricity rates, and software maturity. If Rubin only hits $0.03, savings swell to $53.1B vs. $17.1B. 3. Will running Inference on Venice and future Gen slow down response generation in 2026 and beyond? Human perception of "fast enough" for chat, agents, search augmentation, summarization, coding assistance is roughly Meaning, EPYC may generate $100B a year on data center revenue, Hence $MSFT $AMZN $META $GOOGL OpenAI xAI and 42+ Countries are leaning AMD for Inference, because the cost saving is MASSIVE! 4. Regular users (you, me, people using ChatGPT, Claude, Gemini, Grok, Perplexity...) are extremely unlikely to notice any slowdown and in many cases might even experience slightly faster or more consistent response times if the industry heavily shifts toward AMD EPYC for inference. What actually happens when companies save massively on inference? When OpenAI , Anthropic , Gemini , Grok Meta .... save billions on the batch/enterprise/RAG layer using EPYC Venice, they typically do one or more of these things with the savings, none of which make your chat slower but enhancing their bottom line(Profit) ~Keep prices the same → make more profit ~Lower subscription prices / increase free tier limits ~Train bigger & better models more frequently ~Offer longer context windows ~Add more reasoning steps / tool calls / agents per query ~Improve multimodal capabilities ~Build more data centers / reduce throttling during peaks In practice the consumer experience usually gets better, not worse, when inference becomes dramatically cheaper. Prime example is $META leaning AMD heavily or currently AMD largest customer. or Grok 2 to Grok 3 heavily used AMD for Inference saving. And most Grok Users reported Groke responses snappier, not slower. 5. What does this mean for potential Revenue? Noted that TSMC is massively ramping 2nm supply for $AMD both MI450 and EPYC. EPYC Conservative projection: FY2025: $10.5B(best Est) FY2026: $16B FY2027: $29B FY2028: $49B FY2029: $75B FY2030: $100B Large customers: $META OpenAI $MSFT $AMZN $GOOGL xAI (Apple?) Smaller customer: $DELL $HPE $SMCI and 42+ other countries. The roadmap to $5 Trillion is very much inevitable as Inference Cost from Renting or owning $NVDA are too high, but $NVDA will still dominate Training market share, where MI families are likely to take 15-20% market share, but the TAM is also expanding Rapidly. Most Institutions are projecting $2-$3Trillion TAM by 2030. $NVDA said $4 Trillion. Dr. Lisa Su said $1 Trillion+ by 2030. So you decide on how much TAM. If you enjoy this kind of analysis, Slap the Like/Repost and Bookmark to please the X Algo as it is Free.99! If you want to support my work further, consider subscribe to see more in-depth analysis! Alright, that is it. Not Financial Advice!

Mike

102,223 görüntüleme • 7 ay önce

Just in $AMD Anush "Speed is the moat"|ROCm🎙️ In the race to define the future of AI, what's the one advantage that truly lasts? It's not proprietary tech, argues Anush Elangovan Elangovan, VP of AI Software at AMD , but the sustainable speed of innovation. He explains why AMD is rejecting the "walled garden" model for its open source ROCm stack, betting that an open community flywheel is the key to victory. Listen to understand how this open strategy is designed to out-innovate closed systems by empowering developers to solve everything from frontier-model challenges to the mundane, everyday problems that define the "last mile" of AI. AMD ROCm Software: Part 1 Transcript [00:00:00] Andrew Zigler: Joining me is Anush Elangovan, VP of AI software at AMD. And when people talk about AI compute, the conversation often stops at hardware specs, but it's more than just physical chips that win the game. It's also the software ecosystems supporting them. [00:00:18] Andrew Zigler: The prevailing strategy in the industry has been to build something like a walled garden. You know, something closed, proprietary locks, developers in. But AMD is betting on an entirely different play, open source acceleration, and with rock, their open source AI software stack. AMD is building not just hardware parity, but an innovation flywheel that's powered by the community with interoperability and the freedom to scale without all of that pesky lockin. [00:00:48] Andrew Zigler: And in this world, speed is your moat and how fast you can innovate while your platform remains open, flexible, and standardize across all of its applications. That's what we're gonna explore [00:01:00] today. So Anush, I'm really excited to have you here. Welcome to Dev Interrupted. [00:01:04] Anush Elangovan: Thanks for having me. Uh, super excited to chat about it. [00:01:07] Andrew Zigler: Amazing. Well, let's go ahead and dive right in with kind of what I laid it out with in the beginning, the idea of the moat and it being about speed. I wanna unpack that a bit because that came from you when you and I first spoke. And I, and I want to know, you know, how do you define speed inside of AMD beyond just things like hardware, benchmarks. [00:01:27] Anush Elangovan: Yeah, that's a very good question. So when we typically talk about speed, everyone's like, Hey, hardware benchmark specs, right? Like, uh, memory bandwidth or, or flops. And that is one important part of it, uh, AMD does very well. With that, we do have, a, a very good history of executing on that axis. [00:01:47] Anush Elangovan: But when I say speed is the moat, it is about, uh, how we prepare, how we build the muscle to run the race for a long time and run it fast. And it is [00:02:00] not about a single point in time that you've, you've beat some you know, benchmark and, and you declare victory. It's about building the ability to consistently develop and deliver. [00:02:13] Anush Elangovan: Both hardware and software innovation at scale and do it fast, right? Like, you know, we we're increasingly getting to a point where models come out and they're, uh, you know, a year or two ago it was like, Hey, they work on AMD on day zero, which is great, but now they are performing on AMD the day it releases, right? [00:02:32] Anush Elangovan: So, what does it take to Prefetch where the industry is going? Be prepared to intercept. At that point is what you know, I, I refer to as you know, the, the speed factor in, in creating this mode, right? And the mode is just shed all things that hold you back and run as fast as you can. [00:02:53] Anush Elangovan: Uh, because the pace of innovation that is, uh, being seen in, in AI [00:03:00] industries is just. Amazing. Right? And it's like, it's transformational at at how you generate electricity. It's transformational as at how you build data centers. It's transformational at how you deploy compute, networking. It's transformational at what kind of use cases you, you know, uh, use AI for. [00:03:17] Anush Elangovan: Uh, and for that, you need to be prepared to, see what comes tomorrow and be prepared to run the race tomorrow. [00:03:23] Andrew Zigler: Yeah, it's a really great perspective because it highlights that it's not just like a checkpoint that you run through. I like how you called out, like it's not just hitting that benchmark or being the best in class at that moment, in that snapshot, it's about having a. The throughput and about having that dedication to the idea and continuing to deliver on it. [00:03:43] Andrew Zigler: It's not just crossing the threshold, but it's also being the engine. And that's what, that's what protects a business. That is the moat, because the moat is that innovation layer, the faster and more, uh, future forward. That you can work and think, [00:04:00] you know, the better. Uh, we, we talk a lot about like future forward work styles. [00:04:04] Andrew Zigler: Like what are the things I could be doing right now today that are gonna be like, way more useful tomorrow? Let, let's abandon those, workflows that are older and that kind of like, that translates into. An advantage when you work that way. You know, what kind of things have you learned working with, uh, like across all spectrums of people who would use ROCm, right? [00:04:23] Andrew Zigler: You have like the developers, but then you also have the enterprises and you have this large span of adoptees, right? So what is the, what does that look like that you learn? [00:04:32] Anush Elangovan: Yeah, so, so the way I look at it is there are gonna be pockets of different, uh, you know, cadences, right? Like, so people who are deploying in enterprises, for example, right? The validation and how long it takes for them to deploy an LLM that's secure. It's, with guardrails, et cetera, maybe longer. [00:04:52] Anush Elangovan: but you still have to go through the process and you have to be prepared to like, walk that walk to deploy an enterprises. That doesn't mean it's [00:05:00] not fast, that's as fast as you can do for that industry, right? And if you are deploying AI in healthcare, right, it's, it's got its own, uh, cycle. [00:05:07] Anush Elangovan: but in each one of these, you want to see how, like, go down to the essence of what is it that you actually have to do. And, you know, I, I, I like how you framed it. It's like it's, you shed your prior assumptions of how things are done, right. And, and you kind of build up from a, uh, first principles, uh, approach to say, this is how I could use AI to unlock, whatever I'm doing. [00:05:33] Anush Elangovan: And, and, some of it, you know, it's good to really step back and look at. Just question every part of it, right? Like right now you're getting chat GPT and, Gemini competing for like, math, olympiads and, and, uh, college, uh, reasoning, uh, tests. Right? And, and those are like that, that is amazing and increasingly like complex tasks that they're trying to do. [00:05:58] Anush Elangovan: But there may also be like. [00:06:00] More mundane things that AI could, could get applied to. Right? And, and so when we think about shedding old ways, you wanna shed it not just in like the tip of the spear. It's like, you know, I'm gonna see what's the frontier model. It's also, it could be something as simple as. [00:06:18] Anush Elangovan: How do you choose a, a movie, uh, you know, like a recommendation system, right? Or, or, uh, an automated, uh, flight, uh, rebooking system. So the moment, you know, your flight is late, uh, right now it's a notification, right? It's like, oh, you got a text message saying your flight's late. And I got that like three times this week. [00:06:38] Anush Elangovan: But anyway, uh, and, and, and, and, I was just like, okay, so if I were to rethink this. All this MCPs that we have that should be hooked up into an MCP that says, your flight's delayed. Here are your options. If you want, you know, these are the paid options. Yeah. Here are the free options. This will get you back into your you know, Toronto airport [00:07:00] tonight. [00:07:00] Anush Elangovan: Or if you stay, here's a hotel plus this, plus this, plus. It's just like, go ahead is all I should say. Versus now I'm like, okay, can someone, you know, can I call a travel agent? Can I do this? Can I go online and log into And you know, so we gotta fundamentally rethink even those like small, nuances of, things that we do that can be automated out and AI is really, really good at doing something like this, right? Maybe I just explained an AI startup idea right now. Somebody should just start that. [00:07:29] Andrew Zigler: I think you did. Yeah, you definitely did. Someone, one of our listeners is definitely going to lift that off of you. I, I, I, you know, I hate being on the receiving end of those. You feel a little helpless and then you have to like, follow the whole flow. So I know what you mean. Like I, I like how you called out that the build and this like. [00:07:45] Andrew Zigler: Where speed is your moat and the innovation layer is protecting you, is what makes you better than your competitors. How you scale that and you bring that to market. So by understanding the problems that you're solving, uh, throwing away those older assumptions, but also [00:08:00] recognizing that like. We're building every single day, new things and new ways of using stuff that we're still figuring out the implications of. [00:08:08] Andrew Zigler: And so when you have a lot of velocity and you're introducing a lot of new ideas, and maybe you have that workflow now that automatically rebook your flight off of your late flight text message, and uh, I know I would certainly use it, but you know, what kind of philosophies guide the way that y'all think about building this ecosystem to manage that stability while letting folks. [00:08:29] Andrew Zigler: Play with the speed and the assumptions and the airplane re bookings. [00:08:34] Anush Elangovan: so, so I think, you know, we need to peel one layer down, right? and the philosophy is, Hey, we, we just discovered electricity, right? And you know what we're gonna do? We are gonna make motors, uh, or dynamos, right? Like engines. Uh, sure. We don't know if it's gonna be a Ferrari that you're gonna make, or it's a a a a dump truck. [00:08:57] Anush Elangovan: That's good for doing this. But let's [00:09:00] let, which is also required, right? You need a dump truck. You need a garbage truck. And, [00:09:04] Andrew Zigler: Yeah. You need the [00:09:04] Anush Elangovan: course you need, uh, a Ferrari for a midlife crisis, right? So, [00:09:09] Andrew Zigler: precisely. [00:09:10] Anush Elangovan: But, but my, uh, point is what do we build next? And, uh, and this is what I meant by like, okay, let's, let's take those baby steps to build the. [00:09:20] Anush Elangovan: Infrastructure that's required that we know we'll have to use, right? So, so if I just discovered electricity, okay, great. Now one, how do I save this electricity and how do I use it? So there's battery technology, so you need to do something like that, right? Like so. But then you also want to make it into an actionable thing. [00:09:37] Anush Elangovan: You want to make it for like automobiles, or you wanna use it for, you know, powering, uh, entire cities. So it is that transformational. So, uh, AI is that transformational. So, if you distill down, it'll, it'll come down to how do we think about, what we can do with this this fundamental technology that, We may not be aware of what it [00:10:00] is gonna unlock next, but at least you know the next step is clear, right? It's like a dense fog, you know, it's gonna be like, it, it's the right path. You see the light, but it's kind of like out there and, and the steps you're taking are concrete and you're like, okay, this is good. [00:10:16] Anush Elangovan: I, this is better than where I was or where we were. So we are moving forward. So you can build with the. Intuition from what you see in the short term and a tactical view, but towards what you think the future is gonna be. [00:10:28] Andrew Zigler: Right. You almost like we're all in this like fog of war, right? And like you said, you're reaching out and you're trying to step through it. You could think of it too, as like you're in the dark and your hands are up in front of you and you know that. You're, you're not gonna run your face into a wall because your hands are out in front of you, but you're not gonna maybe do much better than that. [00:10:45] Andrew Zigler: So that's kind of like, I think the eco, the, the industry, the world that we find ourselves in, uh, and we all have to, then this becomes the power of an ecosystem, of a group of people working together to create that layer of, [00:11:00] uh, of establishing the [00:11:01] Anush Elangovan: exactly. And I, I, I just, instead of, you know, saying fog of war I describe it as like, you're in this. Beautiful valley with like a morning, uh, fog that's in. You can smell the flowers. You, you hear the birds. You are like, okay, it's, we are in like, uh, utopian paradise and yes, I just need to like, continue the walk, right? [00:11:24] Anush Elangovan: and then move forward with that, conviction that you're in the right spot. [00:11:27] Andrew Zigler: Yeah. So let's talk about that ecosystem world. This nice, I love how you describe it, this grassy side of a hill in the morning that's covered in some mist and maybe we can't see 30 feet in one direction, but it sure is a beautiful hill and it smells nice. And so we're all here. And why is, in that world, why is. [00:11:44] Andrew Zigler: You know, open source, their strategic advantage that y'all are going for in the AI hardware market. And, and then how does like ROCm turn that into wins for people within that ecosystem? [00:11:56] Anush Elangovan: you know, the, the way we look at it is this, is kind of like how I view [00:12:00] AI and the ecosystem, right? But, but it is for everyone to enjoy. Uh, and so we do want to make sure that. You know, it is, uh, beneficial for everyone. [00:12:09] Anush Elangovan: The ecosystem can come in and, and innovate. It's an open innovation engine. and uh, it is very different from, you know, having a walled garden with, Hey, only I know how to do this and I'm gonna do it and throw it over the fence and you can use it or keep walking, right? So we'd like to be good citizens that way, but also. [00:12:30] Anush Elangovan: Uh, it is self-fulfilling in a way, right? Like it, the, the pace at which we innovate with open source is unmatched. Like, you know, our serving engines are like VLLM and, and sg l. Those things, uh, those frameworks are like super, super aggressive in terms of how fast they come out with features and how fast they can you know, get performant models out. [00:12:52] Anush Elangovan: And that compared with what, uh, you'd get from, you know, the likes of like T-R-T-L-L-M or something is always lagging, right? Because you [00:13:00] just can't keep up with you know, 200 commits a week just on one particular model to get that model really performant [00:13:06] Andrew Zigler: And, and, and in that world where, you know, everyone can enjoy the winds of this, what kind of customer stories or innovation stories have really stood out to you and excite you about building and creating this place for developers? [00:13:19] Anush Elangovan: Yeah. So I think the parts that are super exciting for me are when when we get to see a customer that is first skeptical. Then they start a little like, okay, fine, we'll give you a chance. Uh, we do a simple, uh, POC and then they're like, huh, this seems to work. Yeah, we told you it works. [00:13:42] Anush Elangovan: You don't have to change one line of code. Really? Yes, no need to change one line of code. Okay, let's try a production workload. So then they try it. Oh, you're more performant than the competition. Yes. We're more performant than, than the competition. So how much does it cost? And we're like, oh, it's your TCO is better with, uh, [00:14:00] AMD. [00:14:00] Anush Elangovan: So again, they're like, wow, okay, good. So now how do we deploy at scale? And then we go deploy it at scale. And when they give a thumbs up on that and they say, this is good, right? That's when you know, you, you see it go full circle from like, oh, we, we've never heard about AMD to like actually deploy to tens of thousands of GPUs In the order of a few months, right? It, it, it really is fascinating to see and very exciting and invigorating to [00:14:28] Andrew Zigler: Yeah. At like a great exposure to a lot of interesting problems. And, and then people using the infrastructure, the, the technology available to solve those problems. Really specific problems by the way, that's often why they're bringing their data and AI to it, uh, is because it is really specific and important for them. [00:14:45] Andrew Zigler: And there's a, a lot I think that other engineering orgs can learn and even emulate from AMD's success and, and having this open source ecosystem and it causing this acceleration within. You [00:15:00] know, uh, customers and enterprises that use and adopt the tools and, and, and that creates an advantage. And that goes back to why we're talking and like the real thesis of our conversation today. [00:15:10] Andrew Zigler: So how do you think engineering leaders that are listening to this and obviously tapping into this great success AMD has from an open source flywheel, how do you think other, other folks building in the same space can foster that open, first, that open source oriented culture in order to, you know, accelerate their innovation goals? [00:15:29] Anush Elangovan: Yeah, that's a very good question. So the startup that um, was acquired by AMD we, we built, I mean, we started off doing iot stuff and you know, smart ring and all that, right? But in the, the end of like, uh, and not the end, the last six years of the company was building ML compilers. [00:15:47] Anush Elangovan: And ml, ML compilers are like super, uh, complicated, sophisticated, advanced algorithms, dah, dah, dah. but it was all open source, right? So our VCs were like, wait, what do you mean your core [00:16:00] IP is open source? And um, the speed is the moat applied even then, right? It was just like, yes, if you have an idea that. [00:16:08] Anush Elangovan: Because someone saw this idea that you are, they're gonna be able to catch up, then you probably have the wrong idea anyway. But if they are, you know, you execute and they're gonna catch up, that you should assume they're gonna catch up. Right? So you gotta move forward. So keeping it open source is super important. [00:16:25] Anush Elangovan: But also to your question on like, you know, the learnings from an AMD standpoint, right? If there are, hard problems, I'd say dig in and work through it, right? Like there's no way but through it, right? That should be the simple mentality. And more, uh, frequently than not. you'll see that you'll just make it through in a, in, in good form. [00:16:52] Anush Elangovan: But if you doubt it and you're like, oh, I don't know if I should commit, if I'm, I, you know, what should just commit to do the right thing [00:17:00] every step, right? Every step, and just keep taking one step in front of the other. And in no time you'll see that you'll be running. Right. And, and yes, the first few steps will be like, yeah, everyone's complaining about your software quality. [00:17:15] Anush Elangovan: Everyone's complaining about this and that, and it doesn't work. And, and a few steps in, you know, you get, you get the hang of all the complaints that are coming in. You get the feedback loop. You're like, okay, what, what are you prioritizing again? One step in front of the other, right? You just keep knocking that out and then you get to a point where you're, it just becomes second nature, right? To do the, to do the right thing. And, and then yes, if someone gives you two options, you'll be like, fine. This is, uh, you know, there's always the resource trade off. There's always a human capital trade off, but what's the right thing to do? of course, I, I'm pragmatic about what we choose, but, but if the right thing for your long-term success is dig in, go first, principles, make it [00:18:00] happen. [00:18:00] Anush Elangovan: Well. Then just go for that. There's, there is no shortcut to [00:18:04] Andrew Zigler: acknowledging, you know, how it aligns with your mission, your core company goals, and what you're looking to achieve. And, and I, I love how you rightfully called out that in the open source world and you know, you have your technology that you've built, what you think is your moat upon, right? [00:18:22] Andrew Zigler: It's your code and, and to open source that, or to just make it where anyone could peer in is, you know. Scary in one regard, but two, it just kind of feels like you're handing away your throne room in some kind of sense, a very direct feeling sense. But the ultimately, you were really right to call out, and this is something I think about all the time, that the real power there is still the speed This the speed. [00:18:42] Andrew Zigler: That was the moat at the beginning of our conversation. It's the speed in combination with your. Very specific domain understanding of what you're building and what you're creating, and your new role as the steward of that world and how people plug into it, which [00:19:00] has frankly, a lot more influence and power than lording over a closed. [00:19:04] Andrew Zigler: You know, repository or an ecosystem, and like you said, like throwing things over the wall. Sure. There, there might be people always on the other side of that wall, but you're not gonna have a great connection with them. You're not gonna be able to really clearly understand them. I, I like your metaphor of the side of the field of the mountain a lot more. [00:19:23] Andrew Zigler: But, but in the, in this world, you know, where. That speed is, is the power and, and open source is just one way that you can harness that speed to get really far ahead and to innovate. , There's other parts of this equation that you can be experimenting with too, and I'd love to pick your brain about them as a software leader and, and, and one of them is about looking forward and kind of understanding that future that we're all building towards and beyond today's models and hardware. [00:19:48] Andrew Zigler: You know, what do you see as the next major bottleneck or opportunity in the AI compute space? As, as you know, enterprises and folks start to get a little more mature about what's available to [00:20:00] them. [00:20:00] Anush Elangovan: Yeah, I think, the bottleneck and opportunity is, uh, what I'd call, call walking the last mile of ai. Right. Uh, and like I I, I gave you an example, uh, previously, but, but it's similar to that. It's like there are cases where Humans have so many, uh, things to do in your day. You know, like the, if we sit down and actually had a customer focus like, okay, these customers lives, I'm gonna save four hours of this customer's life. And if you actually sit down and look at all of that, it'll be. Easily automatable, easily you know, uh, applicable, uh, for ai, right? [00:20:39] Anush Elangovan: Like, but then making it happen is gonna take a little bit, right? It's like maybe it's, uh, paying your utility bill, right? Or something like that, right? Or, or, your healthcare explanation of benefits. Uh, like, I'm sure you get an explanation of benefits, and I'm like, I, I don't even know what that thing is. [00:20:55] Anush Elangovan: It's just like EOB and like. [00:20:57] Andrew Zigler: it's a big, a big old PDF. Yeah, [00:21:00] exactly. [00:21:01] Anush Elangovan: Like, like, I'm like great straight to the, uh, shredder, right? And but that could be, you know, automated with the ai, right? It, it, it'd be like, Hey, the summary of this thing is you went and visited this day. Everything is okay. Everything is paid for, so don't worry, it's not a bill. [00:21:17] Anush Elangovan: That again, the same, uh, thing, but the sense of what that information overload is could be. Digested by ai, uh, accumulated over time and retrieved when you need it. Like, I don't, I actually don't even need to know this EOB right now, unless of course, whenever I need to know it, that maybe, you know, like for some benefits I need to figure out what do, what did I do over the past year and how do I apply it? Source:

Mike

14,195 görüntüleme • 8 ay önce