Загрузка видео...

Не удалось загрузить видео

На главную

NVIDIA & Microsoft Advance Development on RTX AI PCs.🤝 TensorRT has been reimagined for RTX. Available via Microsoft’s new inference stack, Windows ML, they streamline development and unlock peak performance. Explore #MSBuild & #COMPUTEX2025 news 👉

Комментарии: 3

Фото профиля DavetheWave
DavetheWave1 год назад

@nvidia Why are you counting generated frames as fps? That's false advertising.

Фото профиля prakash raja
prakash raja1 год назад

• NVIDIA and Microsoft's TensorRT and Windows ML integration boosts AI performance on RTX PCs by over 50%, simplifying deployment.

Фото профиля Lark Davis
Lark Davis1 год назад

2013: $BTC topped 367 days post-halving 2017: 526 days 2021: 548 days If the trend holds, next peak could hit in Oct 2025. Subscribe now and join 130K+ investors prepping for the next big cycle.

Похожие видео

$AMD $MSFT Partnership is MASSIVE in 2026 🚀 If you were excited about my thread on $AMD $AMZN AWS long time partnership, you will be even more excited about what Microsoft gonna do with 2026 AMD EPYC "Venice". Historical Context: The relationship between AMD and Microsoft began in the early 2000s, with Microsoft initially focusing on Intel's x86 architecture for its Windows operating system and server products. However, AMD's entry into the server market with its Opteron processors in 2003 marked the beginning of a competitive dynamic that eventually led to collaboration. The partnership intensified with the launch of 3rd Generation EPYC "Milan" in 2021, powering Azure's N2D and C2D VM families. By 2025, Microsoft had integrated 5th Generation EPYC "Turin" into new compute-optimized instances, reflecting a strategic shift towards AMD for cost and performance benefits. This "Secret Weapon" breakthrough will mark another inflection point for AMD Microsoft Azure relationship, will probably be more aggressive than EPYC "Milan" moment in 2021. We can call it EPYC "Venice" moment 2026" 1. Technical performance of AMD EPYC "Venice" (2026) AMD's 6th Gen EPYC "Venice" processors, slated for 2026, introduce New Chiplet design breakthrough. a revolutionary chiplet interconnect fabric that redefines server scalability for AI. This isn't just faster silicon; it's a paradigm shift for Microsoft Azure , enabling hyper-efficient, rack-scale AI inference that slashes costs and latency while boosting throughput. ~Up to 256 Zen 6 cores, a 70% performance increase over "Turin," optimized for AI and HPC. ~Memory and Bandwidth: 1.6 TB/s per socket, doubling "Turin's" capability, with support for MR-DIMM/MCR-DIMM. ~Efficiency: 1,500-1,700W power draw, a 50% reduction, aligning with Microsoft's sustainability initiatives. ~Interconnect: PCIe 6.0 and a new chiplet fabric for rack-scale AI, reducing latency and enhancing scalability. 2. Why $MSFT will adopt $AMD YPYC Share to 50%+ in 2026. AMD EPYC Share: ~30-35% of Azure's x86 CPU-based business while Intel Xeon share is 65% Microsoft's Azure has been progressively integrating AMD EPYC, with "Venice" expected to expand this footprint: A. Dominance of AI Inference Workloads ~AI inference constitutes 80% of AI workloads in cloud environments, with latency-sensitive applications like chatbots, recommendation engines, and fraud detection requiring sub-second response times. ~"Venice's" 35x inference performance uplift directly addresses these requirements, outperforming Intel's offerings and custom Arm solutions in multi-threaded scenarios. B. Cost Efficiency and Operational Savings ~Azure's 2025 capex of $118B is under pressure to deliver returns. "Venice" can reduce operational expenses by $20-30B annually due to its power efficiency and performance gains, improving Azure's margins to 35-40%. ~The cost per inference operation is significantly lower with "Venice," estimated at 24-31% less than Intel-based alternatives, enhancing Azure's competitiveness against AWS and GCP. C. Scalability for Enterprise AI: ~"Venice" supports rack-scale AI deployments, enabling Azure to scale AI services for enterprise customers. For example, a 1,000-node cluster can process 700,000+ tokens per second, crucial for large-scale AI applications like personalized marketing and predictive analytics. ~This scalability is particularly important as Azure aims to capture the $100B+ AI opportunity by 2026, as stated by Microsoft CEO Satya Nadella. D. Reduction of Nvidia Dependency ~While Nvidia ( $NVDA) dominates AI accelerators, AMD's integrated EPYC-GPU solutions (MI450 with "Venice") offer a balanced approach, reducing Azure's reliance on Nvidia's high-cost GPUs. ~"Venice" enables hybrid inference models, where CPU-based inference handles 80% of workloads, and GPU acceleration is reserved for training and complex tasks, optimizing resource allocation. 3. Financial Implication: ~Revenue from Azure could reach $15-18B annually by 2026, part of a total revenue projection of $70-100B ~Profit margins could improve to 55-60%, boosting net income to $20-25B, supported by scale economies and reduced production costs. Intel could respond by giving more aggressive discounts, but this breakthrough has been a decade long of $AMD R&D, or rethinking chiplet design, a complete new approach. "Venice's" lead in AI inference and efficiency is challenging to match. Broader Industry: Other hyperscalers ( Amazon Web Services , GCP) and enterprises will follow Azure's lead, standardizing EPYC technology and pressuring Intel further. This could lead to a broader industry shift towards AMD, enhancing its ecosystem and bargaining power. Conclusion: The strategic adoption of AMD's 6th Generation EPYC "Venice" processors by Microsoft Azure in 2026 marks a pivotal moment in the evolution of cloud computing, particularly for AI inference capabilities. "Venice's" groundbreaking chiplet design, offering a 35x performance uplift for AI inference tasks, a 50% reduction in power consumption, and unparalleled scalability, positions Azure to leapfrog its competitors in the race for AI dominance. This technical superiority, combined with significant cost savings potentially $20-30B annually in operational expenses; aligns perfectly with Microsoft's ambitions to capture the $100B+ Revenue AI opportunity by 2026. The shift to 50% x86 market share for AMD within Azure is not merely a technical transition but a strategic realignment that redefines the competitive landscape. Historically, Microsoft's partnership with AMD has evolved from niche deployments to a core component of Azure's infrastructure, and "Venice" accelerates this trend. The 30-35% AMD EPYC share in 2025 is expected to double, driven by new VM families like C4D and H4D, which will dominate AI-intensive and HPC workloads. This migration is incentivized by "Venice's" efficiency gains, reducing dependency on Intel and Nvidia, and enhancing Azure's sustainability profile. Not Financial Advice!

Mike

141,018 просмотров • 9 месяцев назад

Progress Development Updates There is one ethos at $SPECT, there isn’t a single day without development and product building. Pleasure to share what we achieved today: 1. There were some user reports about the SE watchlist cards resetting/disappearing after a page refresh or switching to another page. We've fixed this issue by implementing a new handler structure, and it has already been deployed for public use. 2. The Research Zone Lite version now has a Main Project Tweets Preview section. Previously, we displayed project tweets as a list. Now, you can check tweets as embedded blocks with integrated images and videos. This update has also been deployed for public use. (1st attached video) 3. Market Widgets - Mobile Version: The Sector Performance, Market News, X Trending, Partners in Focus, and Compare sections are now optimized for mobile and tablet devices. We're making final adjustments to the full-screen charts and will deploy this mobile version for public use tomorrow. (2nd attached video) 4. Background Development: While we’re adding new features to the Search Engine, there is ongoing background work on our other ecosystem products. The main focus remains on X Bubbles, Monarch AI Agent, and Spectre DexScan. These products will change the game in how we research and trade, as we aim to implement all the functionality DeFi users truly want to see and use daily. Our goal never changes. We build for and with our users to deliver the best tools and products for this space. #ai $SPECT

SPECTRE AI

11,772 просмотров • 1 год назад

🚀New Amazon Q Developer agent for software development is available to customers: This agent is based on a new agent architecture that has exciting results coming from the SWE-bench scores (on the full and verified benchmarks) representing AI models’ ability to resolve real-world coding problems. Interesting aspect of Q Agent is that with these newest updates, Q drove nearly 50% more successful coding tasks completed. What makes Q Dev Agent remarkable? The agent architecture is not just about using the best LLMs (which we do), but also giving the agent the ability to constantly explore multiple paths to find the best way to resolve a particular problem (and back tracking when it has reached dead end like a developer would do). Needless to say, we are just getting started on the developer agent and we are constantly pushing to advance our AI capabilities while maintaining quality, security, privacy, and reliability to keep Amazon Q Developer an innovative and trusted option available to our customers using agents for software development. We highlighted the results of our first SWE-bench submission of Amazon Q Developer back in June blog post; with these updates, our new agent resolves 51% more coding tasks than its previous iteration on the SWE-bench verified dataset, and 43% more on the full dataset. That’s the difference a few months make, and I can’t wait to share what our teams will deliver at re:Invent this December. Here's a quick demo showcasing our new Agent in action:

Swami Sivasubramanian

28,946 просмотров • 1 год назад

MEET THE NVIDIA KILLER: OpenAI bet $10 BILLION on this company that makes chips 20x faster than Nvidia's. If this plays out as expected, it’s over for Nvidia. Cerebras Systems just locked in 750 megawatts of computing power to OpenAI through 2028. For reference: that's equivalent to the annual power consumption of 600,000 US homes. The deal? Over $10 billion. Here's what nobody understands: Cerebras doesn't make normal chips. Nvidia sells you thousands of tiny chips that you connect together. Cerebras makes ONE chip. A single wafer-scale processor the size of a dinner plate. 900,000 AI cores. 4 trillion transistors. All on one piece of silicon. The result? When OpenAI tested it, Cerebras ran inference 20X FASTER than Nvidia GPUs. That's not incremental improvement. That's a different category of performance. But here's where the story gets wild: Four months ago, Cerebras was a struggling company. Their IPO filing revealed that 87% of their revenue came from ONE customer: G42, a UAE-based AI firm. The US government launched a national security review. G42 had ties to Huawei. Ties to China. The IPO collapsed. Investors panicked. Cerebras withdrew their filing in October 2025. Most startups would've been dead. Instead, Cerebras did the opposite. They raised $1.1 billion at an $8.1 billion valuation. Kicked G42 out of the cap table entirely. Got CFIUS clearance. Then landed the OpenAI deal. Now they're raising ANOTHER $1 billion at a $22 billion valuation. They more than DOUBLED their valuation in 4 months. From near-death to $22 billion. While getting rid of their biggest customer. Why OpenAI chose them: ChatGPT has 900 million weekly users. Sam Altman keeps saying they have a "severe shortage" of compute. They need SPEED, not just power. When you ask ChatGPT a question, there's a loop happening: You send request → model thinks → sends response back Nvidia chips are fast at training models. Cerebras chips are built specifically for inference. For real-time responses. For the exact bottleneck OpenAI is trying to solve. Sachin Katti from OpenAI said it best: "Cerebras adds a dedicated low-latency inference solution to our platform. That means faster responses, more natural interactions, and a stronger foundation to scale real-time AI to many more people." In other words: "We need this to scale ChatGPT." The competitive landscape just shifted: Nvidia announced a $100 billion deal with OpenAI in September. But it's still not finalized. Meanwhile, Cerebras closed their deal before Thanksgiving. And it's ALREADY being deployed. Here's the part that should terrify Nvidia: In December, Nvidia bought Groq for $20 billion. Groq makes fast inference chips. Just like Cerebras. So why would Nvidia spend $20 billion buying a competitor to something they supposedly already dominate? Because they know what's coming. Inference is the new battleground. And Cerebras is winning it. The IPO is coming Q2 2026. After this OpenAI deal, Cerebras now has: ✓ IBM contracts ✓ Department of Energy contracts ✓ OpenAI locked in for 3 years ✓ $22 billion valuation ✓ CFIUS clearance ✓ Zero customer concentration risk They went from 87% revenue dependency on one customer to the most diversified chip company outside Nvidia. In four months. The lesson? Smart money doesn't follow headlines. It follows where the AI leaders are actually spending. OpenAI didn't announce this deal for publicity. They need Cerebras hardware to scale ChatGPT. That's a $10 billion vote of confidence. While everyone's watching Nvidia stock, the real war is happening in inference. And the company with ONE giant chip just beat the company with thousands of tiny ones. What do you think happens when Cerebras IPOs?

Ricardo

28,088 просмотров • 6 месяцев назад

Micron is going to $4,000 and once you understand what inference actually is, the number stops sounding crazy (Save this). Dylan Patel just said that by 2030, OpenAI and Anthropic alone will need over 100 gigawatts of compute combined and by 2040, we may not even be measuring AI infrastructure in gigawatts anymore. We may be talking about terawatts. Every single one of those gigawatts needs memory to function. Without it, the compute is worthless. Most people heard that and thought about Nvidia but they should be thinking about Micron. Every AI model generating a response has two phases. The first is prefill, processing your prompt which is compute-heavy and the second is decode generating each word one token at a time and that phase is almost entirely memory-bound, not compute-bound. During decode, the GPU's processing units sit idle more than 95% of the time, waiting for data to arrive from memory. Google confirmed it in a research paper that decode-phase bottlenecks are dominated by memory bandwidth and capacity not raw compute. The GPU is not the bottleneck but the memory feeding the GPU is. This matters because inference is now where all the money lives. Training a model happens once, Inference happens billions of times a day every ChatGPT response, every Claude output, every agentic workflow running in the background and every one of those token streams is a billing event tied directly to memory performance. Adding more GPUs does not fix this because GPUs are already underutilized in inference because they are sitting idle waiting on memory. Adding more memory bandwidth and capacity is what directly reduces token cost, reduces latency, and allows the same cluster to serve dramatically more users simultaneously. Longer context windows compound the problem further, a model running a 1 million token context window requires dramatically more memory per session than a 10,000 token window, and every new model generation pushes context longer. The market treats memory as a downstream beneficiary of Nvidia orders. The correct framework is the opposite, Micron is the upstream constraint on how much value every Nvidia GPU can actually generate at inference scale. Micron guided Q4 to $50 billion in revenue, has HBM4 ramping at twice the pace of the prior generation, and CEO Sanjay Mehrotra has said supply will not catch demand before the end of 2027. At 8x forward earnings on $112 projected FY2027 EPS, Micron is the most undervalued infrastructure company in the entire AI stack. Inference is memory. Memory is Micron and the inference ramp has barely started. Milk Road Pro members are already up massively on this position and we're just getting started. If you want the full breakdown of what we're buying and why, come join us for just a dollar using the link below!

Milk Road AI

128,522 просмотров • 24 дней назад

Jensen is using Nebius to fight the hyperscalers and this is why they will be a $1T hyperscaler (Save this) According to a new Schedule 13G filing, Nvidia beneficially owns 22.25 million Class A shares of Nebius, made up of 1.19 million shares held directly and 21.07 million shares tied to pre-funded warrants acquired back in March 2026. That warrant stake traces back to a $2 billion deal Nvidia struck with Nebius on March where Nvidia bought pre-funded warrants for roughly 21 million shares at an exercise price of essentially zero, structured to work almost like an upfront equity check. Nvidia is currently restricted from exercising or selling any of those warrant-backed shares until September 11, 2026, so this stake has been locked up and largely out of the news cycle until the filing just brought it back into view. That deal came bundled with a much bigger strategic partnership. Alongside the investment, Nvidia and Nebius announced a plan to build out more than 5 gigawatts of Nvidia-powered AI cloud infrastructure by the end of 2030, giving Nebius early access to Nvidia's next-generation Rubin platform, Vera CPUs, and BlueField storage systems well ahead of most competitors. This stake fits a pattern Jensen Huang has been running for a while now. Huang reportedly hates a world where hyperscalers control all the compute, since Google TPUs and Amazon Trainium getting stronger is the one outcome that actually threatens Nvidia long-term. That's why Nvidia keeps putting money into neoclouds like Nebius and CoreWeave and backstopping their GPU clusters, effectively betting on a wide field of players rather than letting three or four hyperscalers dominate the entire compute layer. A GPU sold to Nebius costs Nvidia the same as a GPU sold to Google today, but five years out, every neocloud that survives and scales is one more customer that isn't building its own competing chip and one more reason inference keeps running on open, non-hyperscaler infrastructure instead of a closed ecosystem Nvidia doesn't control. That's the real bull case for Nebius becoming a trillion dollar hyperscaler in its own right. It already has $27 billion locked in from Meta, $17.3 billion from Microsoft, direct equity backing from Nvidia and priority access to Nvidia's next generation chip roadmap before most competitors get it, giving it the capital, the customer base, and the hardware edge all at once, exactly the combination Nvidia needs someone to have if it wants a real fifth hyperscaler standing up against Google, Amazon, Microsoft, and Meta. I remain extremely bullish on Nebius, follow me Melvin for more infrastructure plays and make sure to check out the link below for more!

Melvin

74,607 просмотров • 4 дней назад

Jensen Huang just looked at every tech giant building custom silicon. And laughed. Every major player is burning billions to escape the Nvidia tax. Google is building TPUs. Amazon is building Trainium. Meta is building MTIA. The logic makes sense on a spreadsheet. Design a chip perfectly tailored to your workload. Cut out the middleman. Own the stack. But the spreadsheet assumes the middleman is standing still. Huang: “Look at the number of ASICs that have been canceled… It’s not sensible, actually.” He is not questioning their engineering. He is questioning their math. Custom silicon takes years. Every design choice is a bet on a target that exists today. And Nvidia does not let today exist for long. Huang: “Because of our scale, our velocity, we’re the only company in the world that’s cranking it out every single year.” That is the real weapon. Not the chip. The clock. Nvidia stopped selling performance. They started selling time. By the time a custom ASIC tapes out, Nvidia has already shipped the next generation. The chip arrives obsolete. Not because it failed. Because it was built for a world that no longer exists. The graveyard of custom silicon is not filled with bad engineers. It is filled with slow ones. You cannot aim a three-year development cycle at a one-year moving target. Every company building custom silicon thinks they are building an escape route. They are building a time capsule.

Dustin

43,254 просмотров • 3 месяцев назад

This week, we have had a lot of discussions around artificial intelligence, inspired by the Global AI Summit in Kigali, Rwanda. Many African countries are doing great things to motivate young people to take advantage of AI because it represents the future in problem solving. Unfortunately, Zimbabwe’s ICT Minister Tatenda Mavetera and her permanent secretary did not attend, showing how these things are not taken seriously by our government. Zimbabwe’s richest man, Strive Masiyiwa, who has not been to Zimbabwe for decades, made it clear at the summit, which he co-chaired, that investment will not go where the environment is not conducive. Our government talks about anything topical without delivering anything meaningful—they are doing the same with AI. The majority of schools have no computers. Zimbabweans receive electricity for only four hours a day. As the Under-Secretary-General and Executive Secretary of the United Nations Economic Commission for Africa, Claver Gatete, explains, a country needs electricity for AI data centres to work. Yet only 600 million out of 1.5 billion people in Africa have access to electricity, not even energy. Yet energy is an integral part of AI development. Energy is an essential component for the successful development and implementation of AI in a country because AI systems require massive amounts of data to function effectively. This data must be stored and processed in data centres and servers, which depend on electricity to power both the hardware and the cooling systems. You cannot achieve this in a country that delivers only four hours of electricity a day to its citizens like Zimbabwe. AI applications require continuous and uninterrupted access to data and computing resources to deliver accurate and timely results. Electricity is also crucial for powering research institutions, universities, and Research and Development centres that drive AI advancement. Without reliable access to electricity, these institutions will struggle to conduct research, develop new algorithms, or train AI models. The tragedy of Zimbabwe is that the Vice-President of Google responsible for AI, Dr James Manyika, is Zimbabwean; one of the key presenters at the summit, Prof Arthur Mutambara, who has just released a book on AI, is Zimbabwean; Strive Masiyiwa, who has partnered with Nvidia to bring supercomputer technology to the continent by building Africa’s first artificial intelligence factory in South Africa with data centres in Kenya and Egypt, is Zimbabwean. Yet, none of them are working in or with Zimbabwe. Our political leaders have let us down on all these fronts, yet they keep yapping about AI when there is nothing on the ground! Instead of slogans and dancing at rallies, they should see how other countries are doing it. A country that doesn’t focus on technology for development will be a dusty village in 25 years, and its people will not be able to compete at all, rendering it just a dot on the global map. Too much political yada yada without anything delivered, and with people like Tatenda Mavetera, who forge qualifications, in the driving seat, Zimbabwe’s fortunes will continue to dwindle! Add to that the historic looting of public funds meant for building power plants to give us electricity, future generations will curse on our graves!

Hopewell Chin’ono

33,660 просмотров • 1 год назад