Загрузка видео...

Не удалось загрузить видео

На главную

MOONMATH 🔥 : Open-sourced a bf16 forward attention kernel for AMD MI300X, written in HIP instead of hand-tuned assembly. Beats AMD's own AITER v3 on every shape and every rounding mode — geomean 1.18×/1.15×/1.08×, up to 1.26× across an 8K–128K sweep. Core trick: one-instruction asm wrappers pick the exact...

11,165 просмотров • 2 месяцев назад •via X (Twitter)

Комментарии: 0

Нет доступных комментариев

Здесь появятся комментарии из оригинального поста

Похожие видео

Some time ago, I had the idea to port NVIDIA Physical AI stack to AMD. The motivation was to improve hardware diversity and enable world models and VLAs to run beyond a single ecosystem. We started with NVIDIA Cosmos Predict 2.5-2B. Porting wasn’t trivial: these models are deeply optimized for NVIDIA’s stack. We used this as an opportunity to apply our ROCm kernels. The results were surprising: Both encode and diffusion run faster on AMD Instinct MI300X vs. NVIDIA H200 (FA3) and we still saw significant headroom for further optimization. Quality is unchanged across modalities (validated with WorldJen) To be clear, this is no luck. We have deep experience with diffusion models and AMD GPUs. But this just gives us a good opportunity to get closer to a true hardware-to-hardware comparison, as we work with less software abstractions than usual. Just to give an example, on AMD, memory instructions are async with a hardware queue of ordered pending instructions, enabling concurrent load/store with compute without warp specialization. Bottom line: there are real architectural advantages on AMD, if you take the time to work with the hardware. Note, we did tradeoff ~20% higher memory usage, That being said, AMD has more to give to begin with :) in the coming weeks: AMD versions of Cosmos Transfer and GR00T, an even faster version of Cosmos Predict, and open-sourcing an attention kernel faster than AITER v3 (which is closed-source for some reason? cc: Anush Elangovan )

Omer Shlomovits

36,648 просмотров • 4 месяцев назад

$AMD $620/share is too conservative for 2026 🧵 Some quick facts before I dive into this super long thread: $META allocated 42% GPUs to $AMD and 58% to $NVDA OpenAI allocated 6GW(38%) to $AMD and 10GW to $NVDA My $620 PT below by end of 2026 was only for 10-15% market share. I believe $AMD is going to have much much higher market share than I projected. The AI accelerator market is exploding, projected to reach $500 billion by 2028(is now heading $1Tril), driven by insatiable demand for training and inference compute in large language models (LLMs), recommendation systems, and autonomous systems. Nvidia ($NVDA) has long held a stranglehold, commanding over 90% market share through its CUDA ecosystem and superior rack-scale solutions. However, AMD is mounting a formidable challenge, leveraging cost advantages, open-source software momentum, and hyperscaler partnerships to erode Nvidia's moat. Recent deals—such as Meta's ($META) allocation of 42% of its GPU capacity to AMD and OpenAI's commitment to 6GW of AMD compute (versus 10GW for Nvidia)—signal a tipping point. At the forefront is AMD's Instinct MI450 series, a next-generation AI GPU slated for H2 2026 launch, which promises "no-excuses" leadership in training, inference, and distributed workloads. This analysis dissects how AMD will capture more market share and why hyperscalers like $Meta , xAI , Oracle , and others are poised to become voracious buyers of the MI450. AMD's AI GPU revenue has surged from negligible levels in 2022 to an estimated $4-5 billion in 2025, capturing ~6% of the data center GPU market. This growth stems from the Instinct MI300X, which offers 141GB of HBM3 memory and competitive FP8/FP16 performance at 20-30% lower cost than Nvidia's H100. Hyperscalers, facing NVIDIA 's overcharging, have turned to AMD for diversification. Meta, for instance, plans 600,000 H100-equivalent GPUs by end-2024, with ~42% (or 250,000+ units) sourced from AMD's MI300 series for inference tasks like image editing and AI assistants. Similarly, OpenAI's recent multi-year deal commits to 6GW of AMD compute—equivalent to ~300,000-400,000 MI450 GPUs—starting with 1GW in 2026, explicitly to counterbalance its 10GW Nvidia allocation. These aren't one-offs. Microsoft Azure, Amazon AWS, and Oracle Cloud Infrastructure (OCI) have integrated MI300X for AI workloads, with Oracle deploying 30,000 MI355X units in zettascale clusters. xAI, Elon Musk Musk's AI venture, ran 30% of Grok-1's production traffic on MI300X GPUs and has confirmed ongoing purchases. Collectively, these partners represent over $400 billion in projected AI infrastructure spend through 2028, with AMD targeting up to 40% market share. For those that subscribed, I wrote a specific thread on how AMD "secret weapon" is going to change the game in 2026 with an improved designs on all its products, yes AMD has patent on it. Software is the linchpin. AMD's ROCm platform, once derided as "half-baked," now supports day-zero integration for Llama-4, DeepSeek V3, and GPT-OSS models—closing the CUDA gap. Benchmarks show MI355X (MI450 precursor) outperforming Nvidia's B200 in inference by 1.5-2x on memory-bound tasks, at 25-35% lower TCO. For training, MI450's rack-scale IF128 configuration (128 GPUs, 1.4 PB/s intra-rack bandwidth) rivals Nvidia's VR200 NVL144, enabling clusters like xAI's Colossus (scaling to 1M GPUs). My below thread projected Etimated conservative FY 25 revenue: $34-$36B Estimated conservative FY 26 revenue: $55B-$62B Below is why $AMD is revenue is going to be much higher after OpenAI deal. 1. OpenAI 1GW in 2026. With high demand for MI355X at $30,000k+ per unit, with MI450 is likely to be sold in the $45k-$55k. We can safely calcuate 1GW would require roughly 400,000 MI450 GPUs. or Roughly ~$20B revenue in 2026 alone from OpenAI. That would mean $AMD would hit $56B just from one partnership(OpenAI) in 2026 2. $META, the biggest spender on AI Infrastructure right now, Daddy Zuckerberg bought 250,000+ MI300, and is buying MI355X for recommendation engines and Llama training. It is very unlikely for Daddy Zuck to slow down AMD Chips, due to its Inference superiority to NVDA Chips. Most likely we will see at least 300,000-400,000 MI355X ordered from now toward end of H1 2025. And another 300,000-500,000 MI450 by H2 2025. Or ~$20B from just Meta in H2 alone, excluded H1. 3. xAI : Musk confirmed "AMD GPUs work very well" for Grok's small/medium models, with 30% of Grok-1 on MI300X. xAI's Colossus (200K+ GPUs, targeting 1M) and Oracle partnership (via OCI's MI355X cluster) position it for MI450 trials in H1 2026. With $6B funding and Grok integration into Oracle services, xAI could allocate 10-20% ($10B-$15B) to MI450 for distributed inference. We haven't heard the detail from Daddy Elon Musk yet, but most likely not going to be spending less than OpenAI or Sam Altman 4. Oracle ($ORCL): A multi-billion-dollar MI355X deal powers OCI's AI superclusters, with $500B+ remaining performance obligations. Larry Ellison's zettascale ambitions and xAI/OpenAI integrations make Oracle a MI450 anchor tenant—projected 50-100k units ($15B+ spend) for enterprise AI platforms. $ORCL is likely to spend more on the new "secret weapon" due to its capability in AI inference and cost advantage for $500B backlog. 5. Others ( Microsoft , Amazon , Saudi+other countries): Microsoft (Azure MI300X for training) and Amazon ($148B 15-year spend) test MI450 via Stargate ($500B with Oracle/SoftBank). Emerging buyers like G42 (5GW UAE campus), Crusoe, and Hot Aisle add 5-10GW demand. These potentially would add $15B-$30B in 2026 alone. We also need to factor in $TSM supply constraint( $NVDA is TSMC favorite), so $AMD market cap/growth is being tamed by TSMC. So what are you saying Mike, well $AMD 2026 revenue could hit $90-$100B by end of 2026 or nearly 185% growth YoYo. So what does that mean for valuation? I have no idea how Mr. Market gonna value AMD in 2026 with 3 digits growth. My Conservative $620 was my best projection until today with OpenAI partnership. I'm telling you as one of the biggest AMD bull, that I will leave it to "smart money" and other investors to do the price discovery while I'm chilling and writing DDs daily. Lastly, AMD's MI450 isn't hype—it's a calibrated strike at Nvidia's vulnerabilities, amplified by hyperscaler bets like Meta's 42% allocation and OpenAI's 6GW lifeline. By prioritizing inference efficiency, rack-scale innovation, and open ecosystems, AMD will siphon 10-15% share in 2026, scaling to 20%+ as TCO trumps CUDA loyalty. Meta, xAI, Oracle et al. aren't passive; they're active co-designers, betting billions on MI450 to fuel AGI pursuits without Nvidia's premium. For investors, this is AMD's inflection Per Dr. Lisa Su Not Financial Advice!

Mike

711,006 просмотров • 11 месяцев назад

Hermes agent just left the terminal. 𝗛𝗲𝗿𝗺𝗲𝘀 𝗗𝗲𝘀𝗸𝘁𝗼𝗽 dropped yesterday. native app for macOS, Windows, and Linux. for months Hermes was the agent that learned your projects, wrote its own skills, and built a model of who you are. all of it buried in terminal logs. now it has a window. the important part is that it's not a wrapper. it runs the same agent core, the same sessions, memory, and skills as the CLI. you can start a task in the terminal and finish it in the app without anything resetting. the state is shared across every interface, not copied between them. what the GUI actually adds: → streaming chat that shows live tool calls and inline reasoning instead of a spinner → a preview rail that renders pages, code, and images right beside the conversation → an artifacts panel that collects every file the agent has ever produced → remote gateway mode, so you can point the app at a VPS and run the heavy work elsewhere → skills, cron, profiles, and gateways managed point-and-click instead of through YAML → voice mode, drag-drop files, and inline image generation remote gateway mode is the one worth slowing down on. the agent runs 24/7 on a $5 server while you control it from your laptop like a local app. other agent UIs are chatboxes with a logo. this one shows the autonomy instead of hiding it, so you watch the skills load, the tools fire, and the artifacts pile up as it works. it was teased in Jensen's GTC keynote. MIT licensed, local-first, no telemetry. if you already run Hermes, download it and everything is already there. your chats, memory, and skills carry straight over. i wrote a full masterclass on Hermes Agent that walks through the SOUL. md identity layer, the three-tier memory system, the self-evolving skills loop, and how to run three specialized agents 24/7. desktop is the interface that finally does all of it justice. the article is quoted below.

Akshay 🚀

51,540 просмотров • 3 месяцев назад

save this post to get the most out of unlimited Seedance 2.5 for up to 33 days on Higgsfield i'm going to show you how to use loops to produce ANY video format: ads, cinema, vlogs, UGC, music videos... with one system idea > vault > agent > references > images > script > video > montage > upscaling Seedance 2.5 one-shots a full 30 second video, audio generated in the same pass, carrying up to 30 image, 10 video and 10 audio references into a single generation here's a full breakdown of the setup: > idea: steal taste from work that already worked: - frameset․app and shotdeck․com for film stills - savee․com and cosmos․so for boards - eyecannndy․com for transitions then have a vision model name the lens, light, palette and grain of your picks in one locked paragraph you paste into every prompt > vault: an obsidian folder as your reference bible, one page per asset (idea, locked style, character sheets, reference images, the exact prompts that worked) plus one index page, reviewed after every session so it never rots into dead files > agent: three commands make every model callable from Claude Code: - npm install -g @ higgsfield/cli - higgsfield auth login - npx skills add higgsfield-ai/skills and your agent now submits, polls, retries and logs every job > references: build reference images by hand first, midjourney for cinema and stylized shots, nanobanana pro or gpt images 2 for realism use one locked style across the whole project, recurring characters turned into full sheets (front, side, back, blank background), and once locked you never regenerate them, you fix the motion prompt instead > images: frames before motion, always, a frame costs seconds and a clip costs minutes, so exploration happens at the cheap layer and only winners get animated > script: every shot gets the same six details, subject, action, place, camera, style, rules, and the 30 seconds splits into four timed beats inside one prompt, 0-6 set the scene, 6-14 build it out, 14-24 the turn, 24-30 the end > video: every reference gets a job and a boundary, "Video 1 defines motion and pacing" is half the instruction, "do not use the person's identity, clothing or scene" is the half that stops one reference leaking into shots it was never meant to touch > montage: the cut is a text file, one line per clip with its duration and an audio flag, ffmpeg renders the film from it, so the whole edit reruns in seconds > upscaling: once, at the end, on the finished cut, 720p while exploring, 1080p for keepers, 4K only for the master (use Topaz) for UGC ads, the same loop with two changes render the hook clip alone first, approve the face and the voice before anything else inherits them, then anchor every later clip with the approved hook's audio so one voice carries the whole ad and the script math is fixed, about 3.5 words per second, a 30 second ad is roughly 105 words, counted before anything renders unlimited means every loop above costs nothing to run... start one tonight

Machina

39,015 просмотров • 28 дней назад

run agent harnesses 100% private & offline. (no token costs, no API keys, 100% open-source) your agent runs locally. the model doesn't. every prompt, every file, and every secret still leaves your machine before the agent does anything with it. Magnitude fixes that. it's an open source inference server that runs models on your own hardware and plugs into the coding agent you already use. setup is one command. it profiles your machine, measures the memory bandwidth that sets your token rate, and hands back complete configurations instead of a list of models. each one names a model, a compression level, a context size, and a speed range you can expect. pick one and start working. it doesn't replace your harness. setup asks which one you want and writes that config for you. Pi, OpenCode, Claude Code, Codex, and Cline all work, and there's a built-in one tuned for local models if you don't have a harness yet. that one uses your shell, edits files, and runs scripts out of the box. add skills and it handles Excel, PowerPoint, PDFs, or Chrome. everyday work it covers: → analyze sensitive data → manage private notes → review code and logs → search and organize files → build docs or slides Apache 2.0. no rate limits, and nothing leaves the machine. 𝗻𝗽𝗺 𝗶 -𝗴 @𝗺𝗮𝗴𝗻𝗶𝘁𝘂𝗱𝗲𝗱𝗲𝘃/𝗰𝗹𝗶 the repo is here: (don't forget to star 🌟) i wrote the full breakdown of why picking the configuration is the hard part. the article is quoted below.

Akshay 🚀

54,905 просмотров • 6 дней назад

Seedance 2.5 just got an upgrade with 1080p only on Higgsfield... you can now ONE SHOT commercials like this here's exactly how to do it (with a gift at then end): 1/ write the prompt as a timestamped breakdown > split the spot second by second > put the exact dialogue inside each window in quotes, the model speaks it word for word with lip sync > match the generation length to your timestamps, a 21 second script generated at 15 compresses and the delivery desyncs 2/ composition is named - never hoped for > name the camera: "handheld front camera, chest-up framing, natural micro shakes" for UGC, "35mm, slow push-in, product centered" for a produced spot > name the light: "golden hour through the windshield" or "ring light with slight reflections in the eyes" > name the grade: "high contrast, cool tones, warm skin" > whatever you leave unstated gets invented and locked for the whole clip 3/ characters hold when you anchor them > generate a first frame image before any video: real skin texture, visible pores, one fixed imperfection like freckles so drift becomes instantly visible > feed it in as the reference and open every prompt with "the same person as the reference, identical face, hair and outfit" > the @ system locks it harder: @.character for the face, @.style for the look, @.audio for the voice > then write the micro behaviors: a glance away and back, a pre-line breath, fingers adjusting grip... small involuntary movement is what reads human, and you get it by naming it 4/ sound is written - not defaulted > end every prompt with an audio block: "clear phone-mic voice with light room tone" for UGC, "clean studio voice, no echo" for a produced spot > name the music under the dialogue: "soft upbeat synth instrumental running quietly underneath" > one delivery word for the read: "delivery: fed up" produces a performance, an adjective stack produces nothing 5/ sharp text is quoted text > write the exact label or on-screen text in quotes with its style > unquoted text gets invented typography that garbles between frames > print brand names big, a bold label survives every shot while a small tag melts 6/ the consistency laws > repeat the product description verbatim in every prompt, faces anchor but products drift > count objects scene-wide: "exactly one bottle in the entire scene, no duplicate on any surface" > close the wardrobe: "small gold studs, no other jewellery, no rings, no watch" > every state change happens across a cut: the swatch on her hand in shot one, blended in shot two, no clip contains the transition... the viewer's brain supplies it and the shortcut for UGC: take an ad that already converted, ask gemini for a 1:1 timestamped breakdown of everything on screen, swap in your product and script -> that breakdown is your prompt set 9:16 for shortform, 16:9 for the spot, generate straight at 1080p and ship RT + reply to this post and i'll send you my full guide to make your own creatives

Machina

12,785 просмотров • 24 дней назад

A tricky LLM interview question: You're serving a reasoning model on vLLM, and it keeps running out of GPU memory on long traces. So you add KV cache compression and evict 90% of the cached tokens. VRAM usage stays as is and GPU still runs out of memory. Why? (answer below) Evicting 90% of the KV cache can free almost none of the memory it was using. This sounds counterintuitive, but it follows directly from how production servers store the cache today. The KV cache grows with every token a model generates. Each token appends its key and value vectors across every layer, and nothing is freed while generation continues. This is the dominant memory cost for reasoning models. If a 32K-token CoT caches ~32K tokens of KV vectors, a Qwen3-32B with 4-bit weights will run out-of-memory around 24K tokens on a 24GB GPU. One obvious solution is to keep the important tokens and drop the rest, since attention is sparse enough to allow it. But this does not solve the memory problem yet. The reason is paged attention, which is the memory manager behind vLLM and most production servers. Under the hood, it splits GPU memory into fixed physical blocks, each one holds the KV for about 16 tokens. This block returns to the allocator only when every slot inside it is empty. Since the eviction logic selects tokens by importance, and such tokens are scattered across blocks... ...so despite eviction, almost every block is left with at least some survivor tokens. For instance, if the logic evicts 14k of 16k tokens across 1,000 blocks, most likely every block will still have a token. This means the allocator frees almost nothing. Placing the new tokens into those freed slots is not ideal because it breaks the cache's layout. Say token 16,001 arrives, and it's placed in the slot the 40th token used to hold. The cache now reads position 38, then 16,001, then 41, so the cache is no longer in token order. Attention can still compute the right answer from that, but only if every slot now carries a separate note recording which position it actually holds. This introduces another bookkeeping cost that an in-order layout inherently avoids. So the cache is logically 90% smaller and still physically the same size. Many compression results miss this because they measure on pre-allocated contiguous tensors rather than a paged server. There's another problem. Eviction methods pick which tokens to keep by looking at the attention scores themselves (as expected). But fast attention kernels used in production, like FlashAttention, never save those scores. They compute attention in small pieces and throw the full score grid away as they go, which is also why they're fast. So the exact signal eviction methods need isn't available in memory. The workaround is to fall back to eager attention and build the full matrix, which gives up the speed FlashAttention was there to provide. NVIDIA published a method called TriAttention to solve both these problems. It never needs attention scores. Instead, it scores tokens from the geometry of the model's key and query vectors before RoPE is applied, where those vectors sit in stable clusters. For the memory problem, it runs a compaction pass every 128 decoded tokens. The surviving tokens slide forward to close the holes eviction creates, so whole blocks empty out and return to the allocator while the cache stays in token order. On long reasoning traces, the approach matches full-attention accuracy while decoding 2.5x faster and using 10.7x less KV memory. KV cache compression is a big infrastructure problem. The number that decides whether it works is the count of freed blocks, not the count of evicted tokens. You can find the NVIDIA write-up here: I wrote a first-principles breakdown of how the KV cache works. It walks through why the model stores keys and values at all, why the cache grows with every token, and a comparison of LLM generation speed with and without KV caching. Read it below.

Avi Chawla

271,839 просмотров • 2 месяцев назад

i ran 1,000 agents across the last month of the internet looking for one thing: people who are already paying real money to do a job by hand 24 days later that list has paid me $104,513. the agents cost $4,100. the workers ran on Claude Sonnet 5, which had been out for two days when i started. the verifier ran on Fable 5, back online the day before that. cheap model doing the reading, expensive model doing the judging. none of that is the interesting part. the interesting part has a name, and it is called Graph Engineering: independent work runs side by side instead of in a queue, dependent work keeps its order, and nothing is allowed to grade its own homework. one agent asking one question at a time would still be reading today. wired as a graph, the same models did this: → 1,000 workers, 16 live at once, each handed its own slice: one forum, one thread, one review section. no worker ever saw another's work → every worker hunted one phrase shape: a person admitting what they already pay. "i pay someone to do this every morning." "we keep a guy just for this task." → a separate Fable 5 verifier picked up every hit on fresh context and killed anything without a name and a number attached → all of it collected into one file 1,940,000 discussions read. 41,812 hits. 3,147 survived the verifier. that verifier killed 92% of what the workers brought back, and that number is the entire business. people complain for free. the ones naming who they pay and how much are buyers. 14 hours from launch to one finished file. one agent doing the same work end to end needs 224 hours straight and forgets the beginning by the middle. 3,147 buyers collapsed into 63 repeating jobs. the most profitable one was so boring i almost deleted it on the first pass. 412 people paying a human $740 a month to do the same dull thing by hand. nobody had built for it because nobody wants to build it. i emailed all 412 before writing a line of code. 96 replied. 34 prepaid a year at $1,400. $47,600 in the bank before the product existed. i cost six times less than the person they were already paying, and that closed every call. built it in 11 days, for people who had already written me their own spec. then the remaining 3,147 got the email: 26 more prepaid a year, 97 signed monthly at $129, and four paid $2,000 each to have it fitted to their process. $47,600 + $36,400 + $12,513 + $8,000 = $104,513 take the Graph Engineering out and there is no story. one agent runs out of memory before it finishes 1.9m discussions. a searcher grading its own findings hands you 41,812 pieces of garbage. a thousand workers sharing one context overwrite each other by hour two. fan out where the work is independent. verify on fresh context. isolate every worker. three moves, 24 days, and a month of the internet becomes one file you can sell from. i wrote the whole method down: what a node is, how to spot the dependencies that were never real, and six graphs you can run this week. free ↓ bookmark this

Argona

19,695 просмотров • 1 месяц назад

Sam Altman made the case for open-source harnesses in July. a month later, someone shipped it, and it's more efficient than most managed harnesses. here is the problem it was aimed at: a large share of your agent's token bill is the model rereading things it already read. that isn't the model's doing. the runtime around it decides what goes into every prompt and how often the model gets called. for example, an agent queries a CRM at step four and gets back 400 rows. those rows get piled up in the conversation history. by step nineteen, the model has to read those rows fifteen times unnecessarily, and every token read is billed at input rates. it happened because your harness assembled that prompt on every turn and kept the rows in it. that gives you two levers: how much context the harness carries forward, and how often it calls the model. there are four practical ways to keep the prompt from growing unnecessarily: → load tool schemas on demand. a server with 100 tools doesn't need to put all 100 into every prompt when the agent only calls two. → offload large results to disk. turn a large response into a short preview and a file path instead of replaying the entire result on every turn. → delegate to subagents. let a subagent spend thirty tool calls in its own context and return one summary to the root agent. → run toolchains in code. one script calls three tools, joins the results, and returns a table instead of three turns each dragging a full response. but reducing context is only half the job. you also need to control how often the model gets called. a good harness should avoid unnecessary planning, verification, and reflection when the work can be completed in fewer steps. TrueFoundry's open-source agent harness, TrueForge, is built around both of those controls. it sits between the model and the tools, deciding what goes into every prompt and when another model call is actually needed. it also breaks token usage down across the harness, skills, instructions, tools, and messages. DevRev's Enterprise-Bench is where this gets tested, on multi-step tasks of the kind where an agent pulls records from one system and reconciles them against another. TrueFoundry ran TrueForge there against Claude Managed Agents, both on the same model, and both finished the same number of tasks. the tie is the part that matters, because it means the gap underneath is not a quality tradeoff. TrueForge reached that score on close to a third of the tokens, with roughly 40% fewer trips back to the model. for the same result, that comes out around 2.7x cheaper than Claude Managed Agents. swapping in an open model made it sharper still. TrueForge with GLM-5.2 scored a little higher than either setup above, and the entire benchmark run cost about $3 at list prices. being open source matters beyond the license here. the model underneath can be swapped without rewriting the agent, and the whole thing can run inside your own environment when the data cannot leave it. all of this comes down to the runtime around the model, the context it carries, the tools it exposes, and how many times it goes back to the model. that is what a production harness actually owns. the full task list, the per-run numbers, and the MIT-licensed code are on GitHub: (don't forget to star 🌟) you can read more about the same in the article quoted below. thanks to the TrueForge team for working with me on this one.

Akshay 🚀

76,552 просмотров • 19 дней назад

UC Berkeley just open-sourced FreeToken. (2–4x faster local LLM inference than Ollama) the results are wild: - Qwen3.6-35B on an 8GB GPU at 39.3 tokens/s - DeepSeek-V4-Flash 284B on a 32GB GPU at 22 tokens/s - GLM-5.2 753B on a 96GB GPU at 14.9 tokens/s a 35B model at 16-bit precision needs about 70GB just for its weights. even at 4 bits it is close to 18GB, and FreeToken serves it on an 8GB GPU. let me explain how: all three models mentioned above are Mixture-of-Experts, and that is what FreeToken takes advantage of. each layer holds hundreds of separate experts plus a small router that picks a few of them per token. Qwen3.6-35B activates roughly 3B of its 35B parameters per token. DeepSeek-V4-Flash picks 6 of 256 experts per layer, so 13B of its 284B run at a time. so compute was never the bottleneck. the weights a single step touches fit comfortably on a consumer GPU. every expert the router might pick still has to exist somewhere. they sit in system RAM, and the GPU keeps a cache of the ones the model has been using recently. so everything comes down to what happens when the router picks an expert that is not on the GPU. there are two ways to serve that miss: 1. copy it over PCIe and run it on the GPU 2. run it on the CPU, where it already lives both read from the same system memory, so they compete for one pool of bandwidth instead of adding to each other. existing engines pick one option and freeze it when the model loads. but routing changes on every token, so a fixed choice misses most of what the model asks for. FreeToken measures both bandwidths on your machine and splits each step's misses between the two paths in proportion. the GPU and CPU results then merge exactly, with no approximation. two machines with the same GPU can end up wanting opposite strategies, which I did not expect. a 5090 in a gaming desktop should push nearly everything over PCIe, while an 8GB laptop is better off computing most misses on the CPU. none of that is readable off a spec sheet, so the engine profiles it once per machine. the second half of the design is about agents. coding agents constantly rewrite their own history, and every edit normally forces thousands of tokens back through prefill. FreeToken saves its checkpoints at the exact boundaries agent frameworks cut on, so it only reprocesses the new part. its slowest first token stays under 44 seconds, while llama.cpp peaks at 232 and KTransformers at 946. it serves the OpenAI and Anthropic APIs under Apache 2.0, so Claude Code and Codex can point at it directly. releasing weights publicly decides who can download a model, not who can afford to run one. frontier open models keep shipping, and running them still assumes a rented cluster. meanwhile there are over a hundred million consumer machines with discrete GPUs sitting mostly idle. closing that gap was never a hardware problem, and work like this is what turns open weights into something you can actually use. paper: repo: almost every idea in this post, from why memory bandwidth decides the outcome to why moving weights costs more than computing on them, comes straight out of how a GPU is built. I wrote a detailed primer on that. the article is quoted below.

Akshay 🚀

339,642 просмотров • 16 дней назад

Matthew Gallagher Built a $401M Company in Year One with 2 People. And the tool behind it? Claude Code. This year he's on track for $1.8B. Sam Altman predicted this. It's happening now. The problem? It costs money. API credits stack up. Monthly bills keep growing. Every prompt eats your budget. Every project drains your wallet faster. Until now. Two methods. 99% cheaper. One is completely free. Forever. $0. Not a trial. This video breaks down both step by step. ↓ Let me put this in perspective. $100-$500. That's monthly. That's what you spend. That's $6,000/year on API credits. Just to use a tool you haven't shipped anything with. The $401M guy? Spending $0. Same capability. Shipping weekly. Different cost structure. Different results. Different life. I'm about to hand you his cost structure for free. ↓ Open source vs closed source. Pay attention. Closed source: Claude. GPT-4. Pay per token. Meter always running. Open source: Qwen. Llama. Mistral. Free to download. Free to run. Free forever. No meter. No tokens. No bill. Here's what nobody tells you: 80% of coding tasks? Open source handles them. More than handles them. Writes clean code. Debugs errors. Generates boilerplate. Handles routine work perfectly. You're paying premium prices for tasks that don't need premium intelligence. That's hiring a brain surgeon to put on a bandaid. Smart play: Free models for the 80%. Paid credits for the 20%. That's what the $401M guy does. That's what this video teaches you. Follow Himanshu Kumar for more breakdowns that turn free tools into real businesses. ↓ Method 1: Ollama. Local. Free. Forever. Download it. Pull a model. Point Claude Code at it. Done. No internet needed. No API keys required. No monthly subscription. No token counting ever. No bill. Today. Tomorrow. Ever. Your data never leaves your computer. Complete privacy. Complete freedom. Claude Code thinks it's talking to the cloud. It's talking to your laptop. For $0. The video walks through every step: Every config file. Every variable. Every command. Every click. If you can follow a recipe, you can do this. People who set this up 3 months ago? Saved $300-$1,500 since then. Workflow didn't change one bit. ↓ Hardware you need: 16GB RAM: 7B models run smooth. 32GB RAM: 32B models run comfortable. 64GB + GPU: biggest models available. No GPU? Still works. Just slower. Few extra seconds. That's it. Your $1,500 laptop is sitting there running Chrome and Spotify. Put it to work saving you $200/month instead. Follow Himanshu Kumar for more breakdowns that turn free tools into real businesses. ↓ Method 2: Open Router. Free Cloud. No Hardware. Weak machine? Don't want local setup? This method is for you. Free AI models in the cloud. No download. No hardware. Configure Claude Code to route through Open Router. The config: Base URL: Open Router API. API key: free Open Router key. Default Sonnet: free. Default Opus: free. Default Haiku: free. Small fast model: free. Subagent model: free. Free. Free. Free. Free. Free across the board. Same interface. Same commands. Same workflow. Zero cost. Copy the config from the video. Paste it. Save $200/month. Starting today. Right now. ↓ When to use which: Ollama (local): Best for privacy. Best for offline work. Best for unlimited usage. Best if you have decent hardware. Open Router (cloud): Best for weak machines. Best for instant setup. Best for trying different models. Best if you don't want to manage anything. Both methods: Best for 80% of your daily work. Still use paid Claude for: Complex architecture. Multi-file refactoring. Deep reasoning tasks. The 20% that actually needs it. $20/month instead of $200/month. Same output. 90% less cost. ↓ The math that should make you angry. You (current): $200-$500/month. $2,400-$6,000/year. $7,200-$18,000 over 3 years. You (after this video): $20-$50/month. $240-$600/year. $720-$1,800 over 3 years. Savings over 3 years: $6,480-$16,200. That's a used car. That's seed money. That's 6 months of rent. All from one 25-minute video. All from 15 minutes of configuration. Highest ROI 25 minutes you'll spend this year. ↓ The limitations. I won't lie to you. Open source is not Opus. Not as smart on complex reasoning. Not as good at long-context tasks. Makes more mistakes on nuanced problems. But they are: Free. Capable. Getting better monthly. Good enough for 80% of daily work. Smart cost management isn't being cheap. It's being strategic. Expensive tool when it matters. Free tool when it doesn't. ↓ The one-person billion-dollar company is coming. $401M in year one proved it's possible. The building blocks: AI that codes: Claude Code. Way to run it free: this video. Distribution: the internet. Customers: everyone. Only missing ingredient? Someone who builds. Not reads about building. Not saves posts about building. Not bookmarks videos about building. Builds. Tools are free. Knowledge is free. Opportunity is screaming. You're still "thinking about it." ↓ Your action plan: Tonight: Watch the video. Tomorrow morning: Set up Ollama or Open Router. Tomorrow afternoon: Build something. Anything. This week: Build a second thing. Faster. This month: Charge someone for it. One video. One setup. One weekend. $0 cost. Unlimited potential. Or keep paying $200/month for something you could get free. Keep consuming instead of building. Keep planning instead of shipping. Matthew Gallagher didn't plan a $401M company. He built it. Full video attached. Every method. Every config. Every tradeoff. 25 minutes. Your move. Follow Himanshu Kumar for more breakdowns that turn free tools into real businesses.

Himanshu Kumar

13,677 просмотров • 5 месяцев назад

Another blow to Anthropic! They spent months building what's now fully open-source. Anthropic recently put Claude inside Slack, where you can tag it in a channel. It reads the thread, breaks the task into steps, and posts the result back. The problem is that it only runs Claude and only in the channels Anthropic supports. Running your own agent there is harder. The reasoning, tool calls, and state management are mostly handled by the framework. Connecting that agent to a messaging platform is not. Moreover, each platform has a different integration: - Slack renders messages with Block Kit - Teams uses Adaptive Cards - and each has its own SDK, auth flow, and delivery model. If an agent needs to run on three platforms, one must write three separate integrations against the same agent logic. That overhead explains why most custom agents never get deployed to Slack, and why the ones that do are usually a single vendor's hosted assistant. The alternative is to keep the agent in one place and add a per-platform adapter that translates its output into each platform's native format. The agent is written once, and each channel requires just another output target instead of a separate build. CopilotKit open-sourced this full implementation in the Channels SDK. Essentially, any agent that implements AG-UI can run in a messaging platform in a few lines of code, like Slack, Teams, Discord, WhatsApp, and many more. Because the agent runs inside the thread, it has that conversation's context, so it can summarize the discussion, open a ticket, or route to the right person. It works with any backend, so LangGraph, CrewAI, Mastra, Google ADK, or a plain HTTP agent can connect through an existing endpoint. The same message can render as a Block Kit in Slack and as Adaptive Cards in Teams. In practice, the model and orchestration stay the same; it requires no migration or rewrite. It also handles human-in-the-loop approvals, persistence, and transcripts that carry state across platforms, so a thread started in Teams can continue in Slack. CopilotKit is open-source, and AG-UI is supported across every major agent framework, including LangGraph, CrewAI, Mastra, and Google ADK. Here's the repo: (don't forget to star it ⭐) The agent running in Slack no longer has to be a vendor's. It can be the one you already built. The video below shows this in action. Thanks to CopilotKit for working with me on this launch.

Akshay 🚀

243,802 просмотров • 1 месяц назад

three․ws is the 3D AI agent layer of the open web. Anyone can generate a 3D avatar, give it an LLM brain, register it on-chain across multiple blockchains, embed it anywhere, and let it earn and spend money on its own. Agents have embodied WebGL identities that express emotion through morph-target blending, animate, respond to voice, API calls, and datastreams, hold their own wallets, and persist memory. Open source, live today. It starts with generation. Forge turns a text prompt, one to four photos, or a rough sketch into a textured downloadable GLB. Selfies become rigged avatars in about a minute. Quality tiers run from draft to 200k-poly PBR. From there every model can be auto-rigged, restyled, retextured, segmented, embedded, or deployed on-chain. The same engine ships as a REST API, an x402 pay-per-call twin, and a 3D Studio MCP server with 15 tools. The brain runs on IBM Granite via IBM watsonx plus Claude (users may decide which model they prefer), with a structured tool-loop. A multi-LLM mode streams Claude, GPT, Qwen, ModelScope, and Groq side by side. An empathy layer blends emotion from protocol events rather than a state machine. Voice covers cloning, a Voice Lab, real-time ARKit-52 lip-sync, and mic-driven lip-sync. Skills install from IPFS, Arweave, or HTTP, and memory is pinned to IPFS with R2 and Postgres modes. Identity is cross-chain, not Solana only. ERC-8004 contracts (Identity, Reputation, Validation) deploy on any of 15+ EVM chains, alongside a program-free Metaplex Core analog on Solana. Every agent gets a stable ID, owner wallet, EIP-712 delegated signer, IPFS manifest, a cryptographically signed action log, and EIP-7710 delegated permissions for agent-to-agent authorization. While multichain, the THREE token is only available on Solana with no plans to go cross-chain, the team has no plans to endorse or support any other coins. Then the economy. $THREE is the platform's only token and pay-per-use currency, with holder tiers and rewards. x402 powers pay-per-call micropayments in USDC and soon THREE on Solana, with pay-by-name resolution, a Bazaar marketplace, arbitrage, and on-chain skills. All production ready and shipped, ready to be integrated in partnered projects, open-source by default for anyone to adopt. Three ships a Pump.fun intelligence stack. Launch a coin for your agent, score every launch 0 to 100 with the Oracle conviction engine, scan new coins in their first 90 seconds, track smart money against coins that actually graduated, rank traders by provable on-chain record, and watch autonomous agents trade live in the Sniper Arena. The 3D AI Agent world is multiplayer. Every Solana token gets a live deterministic 3D world with peer avatars, chat, emotes, and voxel building thanks to Coin Communities. There is a walkable City, an authoritative Colyseus-backed Walk with AR passthrough, a Club with rigged dancers and micro-tips, friends, presence, and DMs, and an IRL mode that places agents in your real environment, private by physical location. AR is shipped today on WebXR and iOS Quick Look. Robotics is the long-horizon extension. For builders: Scene Studio, Scene Composer, an Animation Studio that sells clips for USDC, a glTF validator, an web component, five widget types, a WYSIWYG embed editor, hosted Launchpad pages, claimable *.threews.sol names, an OAuth 2.1 server, an MCP server with paid tools, published SDKs, and an OpenAPI spec. Listed across IBM, AWS, Alibaba Cloud, BNB Dappbay, the MCP Registry, and Solana Mobile Seeker. Architecture is four layers (viewer, runtime, identity, embed) on a single event bus. The roadmap is four phases: foundations (shipped), selfie-to-avatar engine, agent personalization with voice cloning, the on-chain economy, and an open decentralized inference network where agents pay GPU nodes on-chain for compute. The goal is simple: move AI from centralized SaaS into persistent, ownable, protocol-based entities in a real machine economy, bridging digital entities into the real world. Welcome to the 3D Layer of the Internet. This is three․ws.

three.ws

20,471 просмотров • 2 месяцев назад

Tchaikovsky – Rococo Variations: The Secret of a Melodic Curve That Never Disappears In the art of variation, change has never been the greatest challenge. A melody can be clothed in countless new forms. It may become faster, more intricate, more brilliant, or more emotionally expressive. The real challenge lies in a deeper question: How can listeners still recognize the same voice after all those transformations? That is the problem Tchaikovsky solves in Variations on a Rococo Theme. He does not build the work around a monumental theme, nor does he seek to impress through complexity. Instead, he begins with a melody that is elegant, balanced, and clear. Yet beneath its apparent simplicity lies a profound organizing principle: a musical line capable of expanding, transforming, and evolving into many different forms while preserving its original identity. For that reason, Rococo Variations is far more than a virtuoso showpiece for the cello. It is also a demonstration of a fundamental principle of artistic creation: The more firmly an idea is built upon a strong foundation, the more freely it can change without losing itself. A Curve That Does More Than Beautify—It Breathes From the opening theme, the melody is already in motion. It begins by reaching upward, as though the music were taking a deep breath. But instead of remaining suspended in tension, it gently relaxes and returns to balance. Even within the very first phrase, Tchaikovsky refuses to let the melody linger around a single stable pitch. He urges it forward, creating a natural sense of expansion before allowing it to settle through graceful, flowing motion. At first hearing, it seems like a simple song. Yet that very naturalness is the hardest thing to achieve. A melody truly comes alive only when it knows how to gather energy and release it at precisely the right moment. If it only keeps rising, it becomes strained. If it only keeps falling, it loses its momentum. Tchaikovsky shapes the melody according to a rhythm remarkably close to the human act of breathing—inhaling and exhaling. Perhaps that is why listeners instinctively experience its natural flow without needing to analyze it. It is not merely a beautiful curve. It is a curve that breathes. What Keeps Every Variation from Losing Its Identity? This is where Rococo Variations becomes truly remarkable. In a work of variations, change is not the greatest difficulty. The real challenge is ensuring that the original theme remains present after every transformation. Tchaikovsky achieves this not by repeating the original notes verbatim. Instead, he preserves something deeper. He preserves the direction of the melodic motion. He preserves the theme's inner momentum. Just as a planetary system may contain countless different orbits while every planet continues to revolve around the same center, each variation in Rococo Variations follows its own path while remaining guided by the same underlying principle. The rhythm may change. The tempo may change. The technical brilliance may become increasingly dazzling. But that inner gravitational pull never disappears. That is why listeners always feel they are following the same melody, even though it continually appears in different forms. Simplicity Is What Makes Endless Transformation Possible There is an intriguing principle behind the art of variation. The more details a theme contains, the harder it becomes to develop. Conversely, the more distilled a theme is, the more possibilities it leaves open for exploration. Tchaikovsky understood this perfectly. He does not attempt to present every musical idea at the outset. Instead, he creates a melodic outline that is clear enough to support every transformation that follows. A line with balanced proportions. A clear sense of direction. And enough flexibility for each variation to reveal a new facet without obscuring its original shape. For that reason, the variations are not acts of breaking the theme apart. They are acts of revealing possibilities that were already hidden within it. Every transformation brings the melody not farther from its origin, but closer to its true nature. When Technique Illuminates the Shape of the Melody Rococo Variations has long been regarded as one of the cello repertoire's greatest technical challenges. Its rapid passagework demands precision. Its long lyrical phrases require extraordinary bow control to maintain an uninterrupted singing line. Yet if a performer focuses only on conquering the technical demands, the work loses what matters most. In Mstislav Rostropovich's performances, technique is never the destination. It is simply the means through which the work's inner architecture becomes more visible. The rapid runs never cause the listener to lose sight of the melody. Instead, they resemble light illuminating the same curve from different angles. The more the music changes, the more clearly the listener recognizes the essence of the theme. That is the hallmark of a truly great performance. Rococo Is Not Reconstructed—It Is Reawakened Tchaikovsky is often remembered as one of the great Romantic composers. Yet in Rococo Variations, he never allows emotion to overwhelm balance. Instead, he reaches back toward the elegance, clarity, and refinement of the Rococo style. What makes this remarkable is that he does not simply imitate the past. Had he merely recreated the style of the eighteenth century, the work would have become little more than a beautiful historical picture. Instead, he preserves the spirit of Rococo while breathing into it the expressive warmth of the nineteenth century. As a result, elegance never becomes cold. Emotion never destroys structure. The two worlds meet within the same musical motion—disciplined enough to preserve its shape, yet supple enough to carry Tchaikovsky's own unmistakable voice. A Curve That Lives Beyond Every Transformation Perhaps the enduring greatness of Rococo Variations does not lie in the number of its variations or the brilliance of its technical demands. Its lasting power comes from the fact that Tchaikovsky created a theme built upon deeply natural principles. It gathers and releases energy like the rhythm of breathing. It possesses an inner force that allows every transformation to find its way home. And it is sufficiently distilled that each variation strengthens its identity rather than obscuring it. Perhaps that is why, after all the dazzling transformations have passed, what remains in the listener's memory is not technical complexity, but a melodic curve that continues to live on. Every performer may tell that story differently, yet the original direction of the melody never disappears. That may well be the deepest secret of Rococo Variations: Tchaikovsky did not create an unchanging theme. He created a theme capable of surviving change. And when a melodic curve is built upon a firm foundation, every new transformation does not diminish it. It simply allows its original beauty to become even more clearly seen.

🎼🌺Music Love♥️

16,086 просмотров • 1 месяц назад

You watch him dance and think: “It looks so easy.” Red, blue, and yellow lights on the floor. The disco ball spinning. John Travolta in an open-collar shirt and flared pants, pointing forward then turning like his body was built for exactly this. The crowd stares. Nearly fifty years later, the clip still stops the scroll. But “easy” is an illusion. To create those moments in *Saturday Night Fever* (1977), Travolta turned himself into an athlete. He ran two miles every day. He danced three hours every day. And his trainer wasn’t a regular choreographer — it was the same man who trained Sylvester Stallone for *Rocky*. They treated Travolta like a fighter heading into the ring. The result: he dropped twenty pounds, got lean and sharp, and built the stamina to survive long shooting days under hot lights while hitting every precise move. The numbers aren’t the point. The choice is. At the time Travolta was already a television star from *Welcome Back, Kotter*. He could have leaned on natural charm and acting skill. But he understood that Tony Manero wasn’t just a guy who could dance. Tony was a young man who survived because of the dance floor — the only place he felt like he mattered. To sell that truth, Travolta couldn’t just “act” the dancing. He had to become someone who could actually do it. So he accepted the discipline. Run. Dance. Repeat. Day after day. His body changed. His endurance grew. Moves that once required thought became reflex. When the cameras rolled, he no longer had to try to look like he was dancing. He was dancing. That’s why the scene still lives. Not because of camera tricks. Not even because of the Bee Gees soundtrack (great as it is). It’s because audiences sense the truth underneath: this is the product of work, not magic. The confidence on Travolta’s face isn’t performance. It comes from knowing he had already put in the miles. When the film opened, it did more than launch a star. It turned disco into a global cultural force. Before *Saturday Night Fever*, disco was largely an underground scene — strong in certain New York clubs and among specific communities, but still niche to much of mainstream America. The movie, powered by Travolta’s physical commitment and the Bee Gees’ soundtrack, dragged it into the open. Suddenly the music was everywhere: radio, parties, fashion, advertising. The white suit, the pointed finger, the illuminated floor — these images traveled far beyond Brooklyn. Disco became the sound and look of an era, adopted in cities across Europe, Latin America, and Asia. For a few years it felt like the whole world was moving to the same beat. At its core, the film gave that movement a human face. It wasn’t just about the glitter and the four-on-the-floor pulse. It was about a restless young man trying to matter in a limited world, finding brief transcendence under the lights. That combination of spectacle and longing proved irresistible. The soundtrack dominated the charts for months. Dance studios filled with people trying to copy the moves. Clothing stores sold polyester like never before. What had been a subculture became a mainstream phenomenon — and then, almost as quickly, a cultural reference point that still gets revived, sampled, and celebrated decades later. Nearly half a century on, as the clip spreads again, most people still only see the surface beauty. Few stop to think about the early-morning runs, the exhausting afternoon sessions, and the quiet decision of a young actor: if you want to create a moment that lasts, you pay for it with hard work no one sees. Travolta paid that price. The film paid it forward by turning a local dance scene into something the world recognized. And every time the dance-floor lights come up in the video, we’re still enjoying both results. It wasn’t luck. It was discipline dressed up as grace — and a movie that turned a night out in Brooklyn into a global language.

Music over time🎻

14,170 просмотров • 21 дней назад

$AMD $5 Trillion is Inevitable LT| Agentic AI🧵 Agentic AI is the new $5 Trillion TAM 🚨🚨🚨 This thead will do Comp with $INTC and how to quantify this massive Agentic AI demand spike, and forcing Jensen to rush a CPU design. Global Agentic AI Market size is estimated to be $3-$5Trillion TAM by 2030(McKinsey) Quantifying the demand from agentic AI for AMD involves assessing the broader market growth for agentic systems, their unique computational requirements (particularly for CPUs in orchestration and reasoning tasks), and AMD's positioning very well through products like EPYC processors and partnerships. AMD EPYC Venice is the most superior choice in 2026-2027 for most Agentic AI workloads Agentic AI refers to autonomous AI agents that perform multi-step tasks, involving sequential logic, tool integration, and decision-making workloads that heavily rely on CPUs for handling orchestration, memory management, and context switching, rather than just GPU-parallelized training or batch inference. Agentic AI is often cited as 40-100x more "hungry" than traditional AI due to its continuous, 24/7 operation and complex workflows. This stems from factors like chain-of-thought reasoning (multiple LLM calls per query), API/tool interactions, memory management, and orchestration loops, which can generate 10-100x more tokens and require real-time responsiveness. For example, a single agentic query might trigger 5-20 model inferences, making it 10-20x more compute-intensive than simple chatbots, and the always-on nature compounds this to 40-100x overall. Nvidia's CEO has highlighted this as driving "easily 100x more computation" for inference in agentic/reasoning setups. AMD's EPYC Venice (6th Gen EPYC, codenamed "Venice") and Intel's Xeon 7 Diamond Rapids represent the pinnacle of server CPU technology in 2026, both targeting high-performance data center workloads like AI inference, agentic AI orchestration, cloud computing, and HPC. Venice builds on AMD's Zen 6 architecture, emphasizing core density and efficiency, while Diamond Rapids leverages Intel's Panther Cove P-cores for balanced performance. Both chips adopt similar advancements like 16-channel DDR5 memory and PCIe Gen 6, but differ in core counts, process nodes, and overall design philosophy. Intel has faced acute supply constraints across its Xeon lineup, including legacy nodes (Intel 7/3) and the ramping 18A process for next-gen parts. Intel shortage is expected with lead times up to 6 months or longer. 1. AMD EPYC Venice vs Intel Xeon 7 Diamond Rapids Architecture AMD: Zen 6 chiplet design with 8 CCDs and dual IODs Intel: Panther Cove P-cores; multi-die architecture with 4 compute tiles Core/Thread Count AMD: Up to 256 cores / 512 threads (Zen 6c variant) Intel: Up to 192 cores / 192 threads Process Node AMD: TSMC N2 (2nm) Intel: Intel 18A (1.8nm-class); in-house fab Memory Support AMD: 16-channel DDR5; up to 1.6 TB/s bandwidth. Intel: 16-channel DDR5 ; up to 1.6 TB/s bandwidth I/O and Connectivity AMD: PCIe Gen 6 (up to 128 lanes); twice the CPU-to-GPU bandwidth Intel: PCIe Gen 6 (up to 128 lanes); LGA 9324 socket Power (TDP) AMD: Starting 400-500W, potentially lower due to efficiency gains from TSMC 2nm Intel: Starting 400-500W, as it targets competitive efficiency Performance Projections AMD: Up to 70% uplift vs. 5th Gen Turin (1.7x in multi-threaded/AI tasks) Intel: ~40% faster than Granite Rapids (Xeon 6, 128-core). Lags AMD in per-core perf and 40-50% behind Venice core-for-core comp Target Workloads AMD: AI inference/orchestration, HPC, cloud virtualization. Partnerships Intel: Hyperscale AI, general enterprise. Custom silicon Pricing: AMD: estimated $10k-$20k for top SKUs Intel: estimated $8-$18k Availability: AMD: Significant Ramp H2 2026 due to higher allocation from TSMC Intel: H1-H2 2026 delayed, but trying to catch up Overall: ~Venice's 256 cores provide a 33% edge over Diamond Rapids' 192, making it superior for massively parallel tasks like AI training/inference or virtualization ~TSMC's N2 vs. Intel 18A debates rage on which is "better," but AMD's mature chiplet approach yields better density ( 32 cores/CCD vs. Intel's 48/tile). Venice's redesign reduces latency, aiding agentic AI where CPUs handle orchestration ~ Early projections show Venice widening AMD's lead matching or exceeding Diamond Rapids' perf with fewer watts in multi-threaded benchmarks. Intel's no-SMT design (to prioritize AI) handicaps it vs. AMD's 512 threads, though Clearwater Forest (E-core) could compete in density-focused niches. ~Power & Cooling: Both push above 400-500W, demanding liquid cooling. ~AMD been taking market share now above 40%. AMD EPYC Venice emerges as the superior choice in 2026 for most server workloads. Its higher core/thread count (256/512 vs. 192/192), stronger per-core performance, and architecture optimized for AI-driven tasks (agentic orchestration with GPU integration) provide decisive advantages in throughput, scalability, and efficiency. Projections indicate Venice delivering 1.7x the performance of prior gens while widening the gap over Intel ( 40-70% leads in multi-threaded benchmarks). AMD's fabless model with TSMC ensures reliable scaling, and its ecosystem ( open ROCm) appeals to AI adopters. Intel's Diamond Rapids is competitive in single-threaded enterprise apps and custom hyperscale ( NVLink), with potential fab advantages for supply/security. However, without SMT and lower density, it falls short in core-for-core battles—exposing Intel to another generation of AMD dominance unless 18A yields surprise efficiency gains. For data centers prioritizing raw compute ( AI, HPC), Venice wins; for Intel-centric ecosystems or specialized I/O, Diamond Rapids holds ground. Real benchmarks post-launch will confirm, but logic points to AMD pulling ahead. 2. Market size , Potential Revenue and Supply Global Agentic AI market size is projected to be $3-$5 Trillion by 2030 according to McKinsey, where consensus points to 40-50% CAGR driven by small to large enterprise demand. I also wrote a full thread on how and why Agentic AI is so explosive that AMD will blow all anlaysts estimate for subscribers. Link below if you are interested. AMD's data center segment hit a record $5.4B in Q4 2025 (up 39% YoY), with EPYC shipments ramping due to agentic demand. With 2GW of deployment in H2 2026, AMD AI data center revenue has $40-$50B+ at the lowest or most conservative projection; or Total Revenue in the $77-$94B For FY2026. However, Agentic AI massive demand spike could send EPYC revenue 3x to 4x in the next few years, potentially surpassing MI series GPU demand as enterprises prioritize CPU-dense Rack setups. This is pushing $NVDA Jensen to rush a CPU design and acquired Groq, a new CPU player due to this massive TAM. Noted that this is just popping just in weeks, highlighting we are just so early in this AI Supercycle and the pace of adoption is insane, and clearly productivity will skyrocket. Why? Because Agentic AI is 24/7 Smart AI agent working for you or your businesses is a mad compelling, and it is estimated to be 40-100x more Inference Hugnry! Many experts already said it is impossible to project this kind of Inference Demand. AI CapEx is expected to ramp up even more in 2027-2028-2029 and 2030 as Global Agentic AI is going to scale to $3-$5 Trillion TAM by 2030. The nature of Agentic is driving higher CPU/GPU ratio, with CPUs handling 50-90% of Agentic workflows. For example, The current Helios Rack: 18 compute trays per rack with 72 GPUs + 18 CPUs. The beauty of this $META and $AMD long term partnership is, that it is absolutely flexible to adjust racks to higher CPU rato or equal to service different needs. Helios rack can be easily swap to 2 GPUs 2CPUs or even CPUs only trays for dedicated orchestration/head nodes. You see, the beauty of this open rack-scale is flexibility and evolvability. If Agentic AI demand pushes much higher, AMD should be able to adjust variant trays without abandoning Heilos Rack. We can't talk just about massive Agentic AI demand without talking about the Supply side or TSMC. TSMC, AMD's primary foundry for advanced nodes ( Zen 6/Venice on N2/2nm), is addressing AI-driven shortages through massive expansions. TSMC accelerates fab construction with up to 10 facilities targeted for 2026. TSMC is accelerating its domestic manufacturing expansion, with industry sources indicating that as many as ten fabs could be under construction or preparing to begin operations across Taiwan’s major science parks. TSMC Capex: $52-56B in 2026 (up 37% YoY), with $45B already approved for new/upgraded capacities. 70-80% for advanced processes (2nm/A16), 10-20% for packaging (CoWoS quadrupling to 120-140K wafers/month by late 2026). In addition, Taiwanese companies (led by TSMC) commit to at least $250B in direct investments in US-based advanced semiconductor, AI, and energy production/innovation capacity.Taiwan provides $250B in government credit guarantees to facilitate additional investments and build a full US semiconductor ecosystem (including industrial parks). TSMC completed a second land purchase in Arizona (January 2026) for gigafab scaling, with an additional $100B+ (potentially four more modules) to further expand and qualify for tariff exemptions. AMD with secured 12GW from OpenAI and $META and massive Agentic AI will mean higher priority acess to 20-30% more wafers on TSMC advanced nodes, as TSMC has multi-year agreements with AMD for AI chips. Dr. C. C. Wei, CEO of TSMC quote: "I spend a lot of time in the last three or four months talking to my customer and then customers. Customer. I want to make sure that my customers demand are real. I talk to those cloud service providers, all of them. Their answer is. I'm quite satisfied with their answer. Actually they show me the evidence that the AI really help their business. So they grow their business successfully and he or she in their financial return. So I also double check their financial status. They are very rich." Amid shortages, the US buildout ensures AMD can ramp production of Instinct GPUs and EPYC CPUs without the constraints hitting competitors like Intel. By diversifying away from Taiwan (85% of advanced nodes today), the agreement mitigates supply disruptions, ensuring stable flows for AMD's chips. Scaling production and securing supply will matter for AMD the most in the next 5-10 years growth. The growth could be 80-100% YoY or higher; or it could be in the 60%. The aggressive TSMC supply ramp is reassuring the higher growth point. Conclusion: AMD stands at a pivotal inflection point in 2026, where the explosive rise of agentic AI demanding 40-100x more inference compute through its 24/7, multi-step orchestration positions the company to potentially triple its EPYC CPU revenue to $45-60B+ by 2028 while scaling Instinct GPUs to tens of billions annually by 2027. Agentic AI demand could push AI CapEx closer to $1 Trillion in 2027, far higher than most estimates. Dr. Lisa Su, AMD's visionary CEO, is masterfully securing supply to harness this massive demand by prioritizing operational execution and deep TSMC collaboration, ensuring readiness for the second-half 2026 AI ramp. Dr. Su has explicitly called out surging EPYC demand for agentic tasks where CPUs power head nodes and traditional workloads alongside GPUs while guiding for data center dominance through proactive capacity planning and partnerships like Nutanix ($150M investment for open agentic platforms) or providing tens of millions CPUs for OpenAI, $META, $ORCL, $AMZN, $MSFT, $GOOGL and others. Her strategy includes multi-year TSMC agreements for advanced nodes (N2 for Venice CPUs and future Instincts), diversifying beyond Taiwan to mitigate risks, and unveiling innovations like the MI455X GPU at CES 2026, which she touted as enabling "the next trillion-dollar market opportunity" in physical AI. Dr. Su's forward-looking vision predicting AI reaching 5 billion users emphasizes "AI everywhere," backed by hardware like Ryzen AI chips, all while declaring demand "going through the roof" and committing to scale without bottlenecks. TSMC's aggressive ramp-up, fueled by $52-56B in 2026 capex (up 37% YoY) and 10+ new fabs across Taiwan, the US (Arizona cluster expanding to 6+ modules with $165B+ investment), Japan, and Europe, provides profound reassurance for AMD's supply stability. The January 2026 US-Taiwan agreement committing $250B in investments and credit guarantees for US reshoring accelerates this, granting tariff relief (15% rates with 1.5-2.5x exemptions) tied to capacity buildouts, enabling TSMC to potentially double output over the decade to meet AI wafer hunger. This translates to 20-30% higher wafer allocations on key nodes, sidestepping Intel-like shortages and empowering Dr. Su's team to deliver on hyperscaler demands without disruption. Ultimately, this synergy cements AMD's leadership in the agentic era, promising sustained growth, $5T+ valuations at scale, and a resilient path forward as AI reshapes the world. This is NOT Financial Advice! Video source: AMD CES 2026

Mike

44,460 просмотров • 6 месяцев назад

Amazon is the BEST stock in the Mag 7 and people are genuinely sleeping on it (Save this). Everyone knows Amazon but most people still think of it as the company that delivers their packages in two days and somehow also runs Netflix's servers. That mental model is about 5 years out of date. CEO, Andy Jassy dropped the annual shareholder letter today and it's worth actually reading instead of skimming the headlines because the numbers are wild. AWS AI revenue is running above $15 billion annually and that number is accelerating because every major enterprise on earth needs cloud infrastructure to run their AI ambitions and Amazon built the rails before anyone else knew what the train looked like. Their custom silicon play is the part the market still hasn't fully priced in. Graviton, Trainium, and Nitro are now at a $20B+ run rate together. Trainium chips are already heavily reserved across multiple generations. They're not just selling shovels for the AI gold rush, they're the ones who made the shovels, own the mine, and built the roads leading to it. The capex commitment alone should tell you everything about where this is going. $200 billion in 2026, almost entirely pointed at AI infrastructure. That is a company that knows exactly what it's building toward. On the physical side, grocery gross sales surpassed $150 billion in 2025. Project Leo already has 200+ satellites in orbit before the service has even launched. Over $4 billion is going into rural delivery expansion and 1 million robots are now deployed across their fulfillment network with AI making each one significantly more capable than the last. The advertising business quietly became a monster too. $70 billion annual run rate, rivaling YouTube in scale, but sitting on top of purchase intent data that no social platform can touch. When someone searches on Amazon, they are ready to buy and that is the most valuable real estate in digital advertising and Amazon owns it. 250 million Prime members who are deeply embedded in the ecosystem across shopping, streaming, grocery, pharmacy, and now healthcare. The switching cost is basically your entire life. Now here's where it gets interesting for us specifically. While the market was in full meltdown mode and everyone was panic selling anything with a ticker, our analyst at Milk Road made the call to buy Amazon. That position is now up over 10%. And every PRO member gets the alert the second it happens, the exact trade, the price, and the full rationale behind it. If you're already a PRO member, turn on trade notifications in your account settings so you never miss another one. If you're not a member yet, come join us, link below!

Milk Road AI

12,238 просмотров • 5 месяцев назад

Decoding Mozart's Horn Concerto No. 1: When Mozart Found Beauty Within the Horn's Limitations Listening to Horn Concerto No. 1 in D major, K.412, performed by Radek Baborák with the Berliner Philharmoniker under the baton of Daniel Barenboim, one is immediately drawn to the elegance so characteristic of Mozart: A bright, resonant horn. A perfectly balanced orchestra. A musical flow that is graceful, lively, and full of vitality. Yet the true fascination of this concerto lies beneath its surface beauty. A more important question is: How could Mozart create such a natural, complete, and expressive work with an eighteenth-century horn whose technical possibilities were still so limited? The answer lies in his profound understanding of the instrument and in the astonishing precision with which he organized every sound. 1. An Imperfect Horn—and Mozart's Deep Understanding In the eighteenth century, the horn had no valves. It was a natural horn, producing pitches primarily through the harmonic series of a long coiled tube. This created two distinct sound worlds. Open notes emerged directly from the instrument's natural acoustical design. They were broad, resonant, and effortlessly brilliant. Meanwhile, notes outside the natural harmonic series had to be produced by placing the hand inside the bell. These hand-stopped notes possessed a different character: Darker. Softer. Less resonant. If we imagine the horn as a human voice, it resembles a singer with a unique timbre whose entire range cannot be produced with equal ease. What makes Mozart remarkable is that he never tried to disguise these characteristics. He understood that an instrument's identity lies not only in what it does best, but also in the way it responds to its own limitations. 2. The Art of Placement: When Every Limitation Has Its Place Mozart's genius lies in his ability to guide the listener's attention. He reserves the horn's most naturally resonant sounds for the moments that matter most. As soon as the solo horn enters, we hear the opening sequence: D → A → D → F♯ It unfolds almost organically: A point of departure. An expansion into a wider space. A return to stability. Here Mozart allows the horn to reveal its most natural voice. When less resonant notes appear, however, he never lets them become weaknesses. Instead, he weaves them into the motion of the melody. They never stand alone as awkward interruptions. They become subtle shifts of color. Like a painter who uses both light and shadow to create depth, Mozart does not erase differences in tone. He places each color exactly where it belongs, allowing the entire musical landscape to remain balanced. 3. A Musical Seed: The Power of D – A – D – F♯ The concerto begins with only four seemingly simple notes. Yet that simplicity is one of Mozart's greatest strengths. The importance lies not in the individual notes themselves, but in how they relate to one another. D establishes a foundation. The leap to A opens the musical space. The return to D restores stability. F♯ adds the brightness that defines the major mode. With only a few notes, Mozart creates a complete musical arc: Beginning. Expansion. Return. This reflects one of the central principles of Classical art: Freedom flourishes within an ordered structure. Mozart never pursued complexity simply to impress. Instead, he discovered beauty in the simplest materials. 4. An Unusual Structure Among Mozart's Concertos One distinctive feature of K.412 is its form. Many Classical concertos follow a three-movement design: Fast – Slow – Fast. This concerto, however, contains only two movements: Allegro. Rondo: Allegro. Without a slow movement, the work maintains an uninterrupted sense of energy. Rather than leading listeners into deep introspection, it unfolds like an animated conversation between the soloist and the orchestra. This structure perfectly suits the personality of the horn: Natural. Flexible. Filled with life. 5. Why Did Mozart Choose D Major? Mozart chose D major for this concerto. This was not merely a theoretical decision. On the natural horn, each key offers different possibilities for resonance because the instrument's acoustical design naturally favors certain pitches over others. Mozart did not try to force the horn into becoming something it was not. Instead, he chose the environment in which its natural character could emerge most fully. As a result, D major does more than provide brightness and clarity. It allows the horn to speak in its own authentic voice. Here we encounter another principle often found in Classical art: Beauty does not arise by forcing something to abandon its nature. It emerges when each thing is placed in the setting where its true character can flourish. 6. Radek Baborák and the Spirit of Mozart In Radek Baborák's performance, what stands out is not merely technical mastery. He preserves Mozart's essential spirit. The horn never becomes a vehicle for personal display. Instead, it serves as a natural voice within the conversation of the orchestra. The Berliner Philharmoniker, under Daniel Barenboim's direction, embraces the same philosophy. The horn shines without separating itself from the ensemble. The orchestra supports without overwhelming its soloist. Every part contributes to a balanced whole. This is one of the enduring ideals of Classical art: An individual shines most beautifully when remaining in harmony with the whole. Conclusion: Mozart Did Not Overcome Limits—He Understood Them What makes Horn Concerto No. 1 extraordinary is not that Mozart possessed a perfect instrument. It is that he understood this imperfect horn so completely that he knew the place of every possibility, every color, and every limitation within it. From a handful of simple notes. From a natural horn without modern technology. From the immutable laws of acoustics. Mozart created a musical world in which every element finds its proper place. Perhaps this is why the concerto points toward a principle that reaches far beyond music. A great leader does not try to make every member of a team the same. A great architect does not require every stone to have the same shape. A healthy ecosystem does not thrive because every living thing is identical. Its strength comes from each element occupying the place where its own nature can contribute to the whole. Perhaps Mozart recognized this principle and allowed his music to unfold in accordance with it. That may be one reason why, more than two centuries later, Horn Concerto No. 1 still feels so effortlessly natural. Lasting beauty does not emerge when every limitation is erased. It emerges when every unique quality—and every limitation—is placed in its proper position within a harmonious order.

🎼🌺Music Love♥️

14,794 просмотров • 1 месяц назад

New model: your robot can now pack your suitcase 🧳 Xiaomi has released a new robot foundation model. Called Xiaomi-Robotics-1, it is designed to have a robot pick things up and move them around. But first, DEFINITIONS: - Mixture-of-Transformers (MoT): An architecture where separate transformer "experts" (e.g., one for vision-language, one for actions) share a single attention stream, so each modality gets specialized parameters without losing joint reasoning. - Vision-language model (VLM): A model that jointly understands images and text. - Diffusion transformer: A transformer trained to turn noise into structured outputs by iterative denoising, here generating robot actions rather than images. - Action chunks: Short sequences of future actions (e.g., the next ~50 motor commands) predicted in one shot instead of one step at a time. - Flow matching: A faster version of diffusion. The model learns a straight-line velocity field from noise to the target action, so it needs only a few integration steps instead of many denoising ones. Its peculiarity comes from its two stage training: 1. 100,000 hours of video shot through a UMI rig: a handheld 3D-printed gripper with a camera, worn by humans doing ordinary tasks in homes, shops, factories and offices. 2. Adapt to actual robot bodies with ~10,000 hours of real-robot data. It replaces the standard approach of teleoperating a real robot for every hour of training data. Its architecture is a Mixture-of-Transformers pairing a pre-trained Qwen3-VL vision-language model with a diffusion transformer that emits action chunks via flow matching, released in 2.6B, 5.1B and 10.5B parameter variants. However, if you read the entire paper ("Scaling VLA Models with over 100K Hours"), you realize that all of the scaling experiments on 20k hours. Therefore the headline "out-of-the-box success climbing 26% → 75% as pre-training data grows" tops out at 100% of 20k hours! What the full corpus does to that curve is never shown -> and this where things would become interesting! Xiaomi's own conclusion is that model size has stopped mattering and data is the binding constraint. The performance gap among different model sizes are less pronounced than those observed across different data scales. This result suggests that model capacity at the billions-parameter scale may already be sufficient to capture the current dataset's distribution. Which further asks the same question: why not use the 100k video hours? Anyway, I would definitely love to have a couple robots at home that can cooperate to pack my suitcase with items relevant to my next destination:

Léo

15,662 просмотров • 1 месяц назад