Video yükleniyor...

Video Yüklenemedi

Ana Sayfaya Dön

TypeGPU + ruNNtime + Jev TypeSafe AI is a very fun combo :D ruNNtime gives me efficient local inference, TypeGPU lets inference and rendering share GPU resources directly with zero copy. That’s 3 separate NN inferences plus rendering, all happening in realtime Since we control the pipeline, Jev can...

38,838 görüntüleme • 1 gün önce •via X (Twitter)

17 Yorum

Konrad Reczko profil fotoğrafı
Konrad Reczko1 gün önce

Jev did not have enough context to know this was a computer mouse and not, you know, a mouse 🐭

Stanisław Białecki profil fotoğrafı
Stanisław Białecki1 gün önce

@typesafeai @_jakubchmura you are starting to freak us out...

Jakub Chmura profil fotoğrafı
Jakub Chmura1 gün önce

@typesafeai damn that's a combo I wasn't ready for

Ernest Szlamczyk profil fotoğrafı
Ernest Szlamczyk1 gün önce

@typesafeai @_jakubchmura The wizard did it again. Honestly I'm getting scared we will get Konrad turning the day to night outside my window with TypeGPU... 🧙

Phat T. Pham profil fotoğrafı
Phat T. Pham1 gün önce

@typesafeai @_jakubchmura this is very cool. It's very useful for AR environments

Iwo Plaza | TypeGPU profil fotoğrafı
Iwo Plaza | TypeGPU1 gün önce

@typesafeai @_jakubchmura SHOW ME A CUP ☕️

PaulinaMcg profil fotoğrafı
PaulinaMcg1 gün önce

@iwoplaza @typesafeai @_jakubchmura Show me some talented people 🔥

$Bill profil fotoğrafı
$Bill1 gün önce

@typesafeai @_jakubchmura I already see how we will use this in AR glasses IRL in danger situation or when you confused

Xavier 👨‍🚀 profil fotoğrafı
Xavier 👨‍🚀1 gün önce

@typesafeai @_jakubchmura noob question, is depthART really useful? yolo26*-depth not enough?

Konrad Reczko profil fotoğrafı
Konrad Reczko1 gün önce

@typesafeai @_jakubchmura They are both fine, just had DepthART already implemented and optimized for my hardware

Xavier 👨‍🚀 profil fotoğrafı
Xavier 👨‍🚀1 gün önce

@typesafeai @_jakubchmura thanks !

ethereagle · building profil fotoğrafı
ethereagle · building1 gün önce

@typesafeai @_jakubchmura Moonshine plus YOLO plus DepthART then Jev is three models before the semantic bit. is that one Jev call on the fused output, or three? that's where the realtime budget goes

Konrad Reczko profil fotoğrafı
Konrad Reczko1 gün önce

I actually only query Jev when either a new label appears or the query changes. It’s very fast, but the latency is still too high to make sense as something realtime. Calling it more often would be wasteful anyway, since all it really cares about are the labels and the query. Whenever a new result is ready, I just feed it into the renderer

M🇿🇦 profil fotoğrafı
M🇿🇦1 gün önce

@typesafeai @_jakubchmura If you want to take it to the next level model in blender using your hands, model doesn't have to be good.

Cathy@tknet.ai profil fotoğrafı
[email protected]1 gün önce

@typesafeai @_jakubchmura This is such a neat real-time pipeline—zero-copy GPU sharing is impressive!

Pratik Mistri profil fotoğrafı
Pratik Mistri1 gün önce

@typesafeai @_jakubchmura It’d be interesting to extend this on digital interfaces and have it interact with them. Like a computer use but more reasoned.

Rashid Khan profil fotoğrafı
Rashid Khan1 gün önce

@typesafeai @_jakubchmura Wow! Amazing. Imagine how good Jev will be in handling real time facial recognition of large crowds & ANPR!

Benzer Videolar

Micron is going to $4,000 and once you understand what inference actually is, the number stops sounding crazy (Save this). Dylan Patel just said that by 2030, OpenAI and Anthropic alone will need over 100 gigawatts of compute combined and by 2040, we may not even be measuring AI infrastructure in gigawatts anymore. We may be talking about terawatts. Every single one of those gigawatts needs memory to function. Without it, the compute is worthless. Most people heard that and thought about Nvidia but they should be thinking about Micron. Every AI model generating a response has two phases. The first is prefill, processing your prompt which is compute-heavy and the second is decode generating each word one token at a time and that phase is almost entirely memory-bound, not compute-bound. During decode, the GPU's processing units sit idle more than 95% of the time, waiting for data to arrive from memory. Google confirmed it in a research paper that decode-phase bottlenecks are dominated by memory bandwidth and capacity not raw compute. The GPU is not the bottleneck but the memory feeding the GPU is. This matters because inference is now where all the money lives. Training a model happens once, Inference happens billions of times a day every ChatGPT response, every Claude output, every agentic workflow running in the background and every one of those token streams is a billing event tied directly to memory performance. Adding more GPUs does not fix this because GPUs are already underutilized in inference because they are sitting idle waiting on memory. Adding more memory bandwidth and capacity is what directly reduces token cost, reduces latency, and allows the same cluster to serve dramatically more users simultaneously. Longer context windows compound the problem further, a model running a 1 million token context window requires dramatically more memory per session than a 10,000 token window, and every new model generation pushes context longer. The market treats memory as a downstream beneficiary of Nvidia orders. The correct framework is the opposite, Micron is the upstream constraint on how much value every Nvidia GPU can actually generate at inference scale. Micron guided Q4 to $50 billion in revenue, has HBM4 ramping at twice the pace of the prior generation, and CEO Sanjay Mehrotra has said supply will not catch demand before the end of 2027. At 8x forward earnings on $112 projected FY2027 EPS, Micron is the most undervalued infrastructure company in the entire AI stack. Inference is memory. Memory is Micron and the inference ramp has barely started. Milk Road Pro members are already up massively on this position and we're just getting started. If you want the full breakdown of what we're buying and why, come join us for just a dollar using the link below!

Milk Road AI

130,756 görüntüleme • 2 ay önce

If intelligence is the log of compute… it starts with a lot of compute! And that’s why we’re scaling our GPU fleet faster than anyone else. Just last year, we added over 2 gigawatts of new capacity – roughly the output of 2 nuclear power plants. And today we’re going further, announcing the world's most powerful AI datacenter, located in southeastern Wisconsin. Fairwater is a seamless cluster of hundreds of thousands of NVIDIA GB200s, connected by enough fiber to circle the Earth 4.5 times. It will deliver 10x the performance of the world’s fastest supercomputer today, enabling AI training and inference workloads at a level never before seen. For AI training workloads, you need compute at exponential scale. That’s why we designed the datacenter, GPU fleet, and network together as one integrated system. This ensures a single job can run from day 1 at exponential scale across thousands of GPUs. Fairwater uses a liquid-cooled closed-loop system for cooling GPUs that requires zero water for operations after construction. And we’re matching all of the energy that is consumed with renewable sources. And of course, it is just one of several similar sites we’re lighting up across our 70+ regions. We have multiple identical Fairwater datacenters under construction in other locations across the US, in addition to our AI infrastructure already deployed in over 100 datacenters around the world, powering model training, test-time compute, RL tuning, and real-time inference at global scale. Too often during times like this, people go with the current and only later wonder, how did we get here? With Fairwater, we're charting a new path: doing the hard engineering work, bringing compute, network, and storage into one highly scaled cluster, and designing closed-loop energy systems to meet real-world computing needs. And partnering with local communities to ensure it's thoughtfully done in a way that is sustainable, creates new jobs, and expands opportunity. We are thrilled to see this take hold in Wisconsin, and we are just getting started.

Satya Nadella

2,026,666 görüntüleme • 1 yıl önce

$AMD | Inference's estimated to be 80%+ in 2027👑🆕 Dr. Su told everyone the world will need a lot more CPUs and Inference will dominate most of compute long term from 2022. Nobody believed her, but I did along with other high conviction investors. 2026 is the first time Inference surpassed training at 65%+ vs 33-35% for training. Early 2026 infrastructure spend: ~55% inference With Agentic AI, 2027 Inference Infrastructure spend is projected to be 80%+ Training still grows in absolute terms (bigger clusters, more experiments). Inference grows faster because usage, agents, and reasoning traces multiply token volume continuously. That’s why hardware and data center design are shifting toward latency, utilization, and cost per token rather than peak training FLOPS. What "token efficient" actually means when AI labs produces better/smarter models? When models get more token efficient, the expensive part of an agent the GPU “thinking” step gets shorter and cheaper. The rest of the loop does not. The agent still has to parse the answer, pick a tool, run code in a sandbox, query a database, open a browser, apply guardrails, and feed the result back. Those steps live on CPUs where AMD has the best CPU in the world. So a more efficient model does not shrink the agent; it shrinks the model’s share of the agent. Wall clock time and cost tilt toward orchestration and sandboxes, which is why you provision more CPU racks even as tokens per decision fall. Cheaper thinking also unlocks more doing. Teams stop designing one shot answers and start adding retries, parallel branches, sub-agents, and24/7 digital workers, Jevons paradox for agents. Each extra loop is another isolated environment, and unlike GPU batching, sandbox demand scales almost linearly with concurrency: fifty candidate patches means fifty containers, not one fatter GPU job. Token saving tricks often push even more work onto that layer. The result is a fleet that thinks less per step and acts more often, so the volume product becomes CPU/sandbox capacity, not just accelerators. Agentic workloads are a big part of why inference is pulling ahead. I’ll pull the latest numbers on token multipliers and how that shows up in 2026 compute mix.Yes. The inference flip is mostly agents + reasoning, not more people chatting. Chat was one prompt in, a few hundred tokens out. Agents turn a single user request into a loop: plan, tool call, read the result, think, retry, hand off to a sub-agent. That is why Gartner’s 2026 range is 5–30× more tokens per task than a chatbot, with coding and research agents often landing higher. On OpenRouter, agentic workloads 14x’d in six months, passed human usage in February 2026, and by early August were ~5× human tokens (~7.3T/day agentic vs ~1.4T human). A typical agentic request there used 15× the tokens of a human query. One important thing to understand, unit cost per token is still falling, but tokens per useful outcome rose faster. An “agentic seat” can burn 50–100× the tokens of a chat subscription. That is impressive growth for inference. It is also why KV cache, speculative decoding, quantization, and inference specific silicon suddenly matter more than another giant training cluster. Not Financial Advice! DYOR!

Mike

16,308 görüntüleme • 12 gün önce

There's been a few cool updates recently. In particular, Rerun 0.33 released headless rendering. This, along with the Fable 5 release pushed me to work torwards making MAMMA realtime! I threw Fable at the problem, and it was able to take original implementation that was ~12 seconds / frame and get it all the way down to 40ms /frame, or nearly a 300x speedup 🏎️ How did I achieve this? TLDR: - Use rerun's headless rendering as supervision when optimizing - Save rrd file as test fixture to guide model optiziation with /goal - create an html artifact with headless rendering to provide detailed breakdown of what it did and how it actually looks like in the viewer There were a few critical bits to make sure that this ACTUALLY worked and that Fable didn't just cheat or delete something and declare victory. The first is that the original version used Rerun, this allowed us to save things to disk as an RRD file, meaning we could query the contents and use this as a sort of test fixture or golden artifact that held EXACTLY what all of the values should be. Then we can use this with /goal as a metric when doing the optimization to ensure there are no regressions. The second bit is the headless rendering, this gave us the ability to check that not only did the test fixture pass, but it also looked visually correct. This made a huge difference, and an awesome side affect of it is that we can use the headless rendering to create an implementations.html file. This gives a visual guide as to what the agent did (I walk through it in the video below) Along with this, we're working on an MCP server for rerun that allows full interactivity with the rerun viewer for your agent. So for example the agent can click, drag, move views, scroll timelines, ect. I used this to help the agent debug certain parts such as when the 2d sam masks didn't line up, or if the triangulated keypoints werent correctly matching with the optimized mesh. The agents could go, click into the view, scroll through the timeline and see where things went wrong. Fable + Headless Rendering + Rerun MCP == 300x speedup in less then a days work With these new tools, I'm planning on going back to my gaussian splatting implemntation and cleaning it up + making it fast!

Pablo Vela

22,880 görüntüleme • 3 ay önce

After taking some time off post-Rapid, I'm excited to share what I’ve been up to since: Datawizz AI! We’ve raised a $12.5M Seed led by Human Capital to make AI 10x cheaper, 2x more accurate and 15x faster by transitioning from LLMs to SLMs. AI is eating the world. But unit economics are eating AI. Looking at the fastest growing AI products, they all share two traits - growing fast, and painful inference bills. General-purpose LLMs are just too expensive to run. A big reason for that is we train LLMs to be good at everything - answer any question, be an expert on any topic. The big labs dub this "generalisation", but for real-world applications, it is unnecessary. In reality - many AI applications need models to be experts in one thing - and do that thing extremely well. Your coding model doesn’t need to memorize ancient recipes for Garum sauce. This is where Datawizz comes in - we sit between the AI applications and automatically create smaller (100x-1,000x) specialized models to handle specific aspects of your work. By focusing the model and combining industry-data in the distillation process - we end up with models that beat SOTA LLMs at a fraction of the cost. We created Datawizz to make AI specialized and scalable. We’re early in the journey, but have already been able to save companies 90%+ on their inference bill and speed up their apps by 10x. Excited to build better AI platforms? Join the Datawizz team (link in first comment)

Iddo Gino 🐙

21,928 görüntüleme • 11 ay önce

ELON MUSK: We believe the AI5 chip will be roughly comparable performance to an NVIDIA Blackwell, and at much less than 10% of the cost Transcription: I'm super hardcore on chips right now as you may be able to tell. I have chips on the brain. I dream about chips, Literally! Because in order to have a functional robot, you have to have a great AI chip. And it needs to be an inexpensive chip and it needs to be very power efficient So we think we believe the AI5 chip will be probably about a third of the power of say something like a Blackwell, an NVIDIA Blackwell, which is a great chip, for roughly comparable performance. And much less than 10% of the cost. This is a chip that is very much optimized for the Tesla AI software stack. So it's not meant to be a general purpose chip, it's meant to be an amazing chip for the Tesla AI software And I mean a couple of things that I think make... like how is Tesla able to achieve such an improvement? I think it is because we are specialized. We're not trying to... you know, NVIDIA has to serve the superset of all past and future customers. So all of their requirements, all of the software that they've written has to work, which is a very difficult problem. Whereas we just need to make it work for our software. And so we're able to simplify the chip dramatically And then we also, I think we're unique in this, but like we have an integer-based system. And integer operations are fundamentally more efficient than floating point operations. So we can do floating point, but the vast majority of our inference is done in integer. Which is, if you're familiar with sort of logic gates, the simplicity of integer... it's integer is much more power efficient, much more silicon efficient, but you have to, you actually have to train for integer inference, which everyone else is training for floating point. That's kind of like a niche technical detail, but it's actually very important. So, yeah, this is going to be a great chip So this chip will be made in basically in four places: TSMC Taiwan, Samsung Korea, TSMC Arizona, and TSMC Texas. And we already know what improvements to make for AI6. So I'm hopeful that we can within less than a year of AI5 starting production, we can actually transition in the same fab to AI6 and double all of the performance metrics

X Freeze

305,109 görüntüleme • 10 ay önce

Today we announced our new Fairwater datacenter in Atlanta, connected with our first Fairwater site in Wisconsin and our broader Azure footprint to create the world’s first AI superfactory. Fairwater exemplifies our vision for a fungible fleet: infra that can serve any workload, anywhere, on fit-for-purpose accelerators and network paths, with maximum performance and efficiency. AI workloads have evolved beyond large-scale pre-training. Today, they encompass fine-tuning, reinforcement learning (RL), synthetic data generation, evaluation pipelines, and more. Fairwater is built to support this full lifecycle: Max density: Fairwater’s two-story design and liquid cooling system lets us place racks in three dimensions and pack them with GPUs as densely as possible, minimizing cable runs and improving latency and effective bandwidth. Fleet: Each Fairwater DC can integrate hundreds of thousands of the latest NVIDIA GPUs into a single coherent cluster. This provides flexible infra that can support the full spectrum of workloads, and ensure no GPU is left unnecessarily idle. And that’s on top of the more than 100,000 GB300s coming online this quarter alone for inference across the rest of our fleet. For us, it’s all about turning every gigawatt into the maximum number of useful tokens. Not every GW is created equal! Planet-scale: Every Fairwater DC will connect through our continent-spanning AI WAN to prior generations of AI supercomputers, forming a truly fungible pool of compute. This enables developers to scale beyond the capacity of a single site and dynamically land workloads on the right infra for their needs. Together, these innovations let us bring together different generations of silicon and AI systems across DCs and geos into a single elastic system that scales seamlessly across training and inference workloads And this elastic AI capacity is all available alongside all the other cloud services (compute, storage, databases, app services) that AI agents and workloads need. This is what we mean when we talk about building a fungible fleet – a single, unified platform that pushes the limits of performance per watt and per dollar. Read more:

Satya Nadella

908,065 görüntüleme • 10 ay önce

$AMD $MSFT Partnership is MASSIVE in 2026 🚀 If you were excited about my thread on $AMD $AMZN AWS long time partnership, you will be even more excited about what Microsoft gonna do with 2026 AMD EPYC "Venice". Historical Context: The relationship between AMD and Microsoft began in the early 2000s, with Microsoft initially focusing on Intel's x86 architecture for its Windows operating system and server products. However, AMD's entry into the server market with its Opteron processors in 2003 marked the beginning of a competitive dynamic that eventually led to collaboration. The partnership intensified with the launch of 3rd Generation EPYC "Milan" in 2021, powering Azure's N2D and C2D VM families. By 2025, Microsoft had integrated 5th Generation EPYC "Turin" into new compute-optimized instances, reflecting a strategic shift towards AMD for cost and performance benefits. This "Secret Weapon" breakthrough will mark another inflection point for AMD Microsoft Azure relationship, will probably be more aggressive than EPYC "Milan" moment in 2021. We can call it EPYC "Venice" moment 2026" 1. Technical performance of AMD EPYC "Venice" (2026) AMD's 6th Gen EPYC "Venice" processors, slated for 2026, introduce New Chiplet design breakthrough. a revolutionary chiplet interconnect fabric that redefines server scalability for AI. This isn't just faster silicon; it's a paradigm shift for Microsoft Azure , enabling hyper-efficient, rack-scale AI inference that slashes costs and latency while boosting throughput. ~Up to 256 Zen 6 cores, a 70% performance increase over "Turin," optimized for AI and HPC. ~Memory and Bandwidth: 1.6 TB/s per socket, doubling "Turin's" capability, with support for MR-DIMM/MCR-DIMM. ~Efficiency: 1,500-1,700W power draw, a 50% reduction, aligning with Microsoft's sustainability initiatives. ~Interconnect: PCIe 6.0 and a new chiplet fabric for rack-scale AI, reducing latency and enhancing scalability. 2. Why $MSFT will adopt $AMD YPYC Share to 50%+ in 2026. AMD EPYC Share: ~30-35% of Azure's x86 CPU-based business while Intel Xeon share is 65% Microsoft's Azure has been progressively integrating AMD EPYC, with "Venice" expected to expand this footprint: A. Dominance of AI Inference Workloads ~AI inference constitutes 80% of AI workloads in cloud environments, with latency-sensitive applications like chatbots, recommendation engines, and fraud detection requiring sub-second response times. ~"Venice's" 35x inference performance uplift directly addresses these requirements, outperforming Intel's offerings and custom Arm solutions in multi-threaded scenarios. B. Cost Efficiency and Operational Savings ~Azure's 2025 capex of $118B is under pressure to deliver returns. "Venice" can reduce operational expenses by $20-30B annually due to its power efficiency and performance gains, improving Azure's margins to 35-40%. ~The cost per inference operation is significantly lower with "Venice," estimated at 24-31% less than Intel-based alternatives, enhancing Azure's competitiveness against AWS and GCP. C. Scalability for Enterprise AI: ~"Venice" supports rack-scale AI deployments, enabling Azure to scale AI services for enterprise customers. For example, a 1,000-node cluster can process 700,000+ tokens per second, crucial for large-scale AI applications like personalized marketing and predictive analytics. ~This scalability is particularly important as Azure aims to capture the $100B+ AI opportunity by 2026, as stated by Microsoft CEO Satya Nadella. D. Reduction of Nvidia Dependency ~While Nvidia ( $NVDA) dominates AI accelerators, AMD's integrated EPYC-GPU solutions (MI450 with "Venice") offer a balanced approach, reducing Azure's reliance on Nvidia's high-cost GPUs. ~"Venice" enables hybrid inference models, where CPU-based inference handles 80% of workloads, and GPU acceleration is reserved for training and complex tasks, optimizing resource allocation. 3. Financial Implication: ~Revenue from Azure could reach $15-18B annually by 2026, part of a total revenue projection of $70-100B ~Profit margins could improve to 55-60%, boosting net income to $20-25B, supported by scale economies and reduced production costs. Intel could respond by giving more aggressive discounts, but this breakthrough has been a decade long of $AMD R&D, or rethinking chiplet design, a complete new approach. "Venice's" lead in AI inference and efficiency is challenging to match. Broader Industry: Other hyperscalers ( Amazon Web Services , GCP) and enterprises will follow Azure's lead, standardizing EPYC technology and pressuring Intel further. This could lead to a broader industry shift towards AMD, enhancing its ecosystem and bargaining power. Conclusion: The strategic adoption of AMD's 6th Generation EPYC "Venice" processors by Microsoft Azure in 2026 marks a pivotal moment in the evolution of cloud computing, particularly for AI inference capabilities. "Venice's" groundbreaking chiplet design, offering a 35x performance uplift for AI inference tasks, a 50% reduction in power consumption, and unparalleled scalability, positions Azure to leapfrog its competitors in the race for AI dominance. This technical superiority, combined with significant cost savings potentially $20-30B annually in operational expenses; aligns perfectly with Microsoft's ambitions to capture the $100B+ Revenue AI opportunity by 2026. The shift to 50% x86 market share for AMD within Azure is not merely a technical transition but a strategic realignment that redefines the competitive landscape. Historically, Microsoft's partnership with AMD has evolved from niche deployments to a core component of Azure's infrastructure, and "Venice" accelerates this trend. The 30-35% AMD EPYC share in 2025 is expected to double, driven by new VM families like C4D and H4D, which will dominate AI-intensive and HPC workloads. This migration is incentivized by "Venice's" efficiency gains, reducing dependency on Intel and Nvidia, and enhancing Azure's sustainability profile. Not Financial Advice!

Mike

141,018 görüntüleme • 11 ay önce