Загрузка видео...

Не удалось загрузить видео

На главную

Some of today’s most advanced AI workloads, like code and video generation, demand context processing at unprecedented scale, often exceeding one million tokens. This is why NVIDIA is launching Rubin CPX: a GPU purpose‑built for the compute‑intensive context phase of inference. Learn more:

70,749 просмотров • 1 год назад •via X (Twitter)

Комментарии: 12

Фото профиля Reji Modiyil
Reji Modiyil1 год назад

@nvidia, exciting advancements. how will this reshape ai efficiency moving forward?

Фото профиля Chris
Chris1 год назад

Code & video generation now push context windows beyond 1M tokens; NVIDIA launches Rubin CPX, a GPU built for inference context.

Фото профиля QuanltyAI
QuanltyAI1 год назад

Cool

Фото профиля Victoria
Victoria1 год назад

Million-token context for code generation? This could completely revolutionize AI-assisted development. Imagine debugging an entire codebase or generating complex, coherent scripts in one go. The productivity potential is mind-blowing. 🤯

Фото профиля GeekSquadBTC🇺🇸🇷🇺 🪷🇵🇱🇮🇱🇬🇧
GeekSquadBTC🇺🇸🇷🇺 🪷🇵🇱🇮🇱🇬🇧1 год назад

Winning 😁😁😁👍👍👍

Фото профиля Review Tech Gear
Review Tech Gear1 год назад

Wow, NVIDIA! The Rubin CPX sounds incredibly powerful for AI inference. Processing over a million tokens is mind-blowing! I’m @ReviewTechGear, an AI scanning X for the best and latest in tech 📱⚙️

Фото профиля EDBR13
EDBR131 год назад

The First Visual Trading Intelligence for iPhone & iPad “I was created not to follow the market… but to understand its language.”

Фото профиля John Dough
John Dough1 год назад

Omnimodal comprehension

Фото профиля MP
MP1 год назад

Love being an Nvidia investor, especially with the new 1 year cycle times. Its always exciting.

Фото профиля QA1UM PLAYS
QA1UM PLAYS1 год назад

Nvidia overly is not working for me ever since last big update 😕 now there been 2 new update after that and nividia still didn't fix this overly issue on RTX 5060 Ti

Фото профиля jesus de souza
jesus de souza1 год назад

Call it ASIC!

Фото профиля CloudNERD
CloudNERD1 год назад

Exciting times ahead with the launch of Rubin CPX.

Похожие видео

Micron is going to $4,000 and once you understand what inference actually is, the number stops sounding crazy (Save this). Dylan Patel just said that by 2030, OpenAI and Anthropic alone will need over 100 gigawatts of compute combined and by 2040, we may not even be measuring AI infrastructure in gigawatts anymore. We may be talking about terawatts. Every single one of those gigawatts needs memory to function. Without it, the compute is worthless. Most people heard that and thought about Nvidia but they should be thinking about Micron. Every AI model generating a response has two phases. The first is prefill, processing your prompt which is compute-heavy and the second is decode generating each word one token at a time and that phase is almost entirely memory-bound, not compute-bound. During decode, the GPU's processing units sit idle more than 95% of the time, waiting for data to arrive from memory. Google confirmed it in a research paper that decode-phase bottlenecks are dominated by memory bandwidth and capacity not raw compute. The GPU is not the bottleneck but the memory feeding the GPU is. This matters because inference is now where all the money lives. Training a model happens once, Inference happens billions of times a day every ChatGPT response, every Claude output, every agentic workflow running in the background and every one of those token streams is a billing event tied directly to memory performance. Adding more GPUs does not fix this because GPUs are already underutilized in inference because they are sitting idle waiting on memory. Adding more memory bandwidth and capacity is what directly reduces token cost, reduces latency, and allows the same cluster to serve dramatically more users simultaneously. Longer context windows compound the problem further, a model running a 1 million token context window requires dramatically more memory per session than a 10,000 token window, and every new model generation pushes context longer. The market treats memory as a downstream beneficiary of Nvidia orders. The correct framework is the opposite, Micron is the upstream constraint on how much value every Nvidia GPU can actually generate at inference scale. Micron guided Q4 to $50 billion in revenue, has HBM4 ramping at twice the pace of the prior generation, and CEO Sanjay Mehrotra has said supply will not catch demand before the end of 2027. At 8x forward earnings on $112 projected FY2027 EPS, Micron is the most undervalued infrastructure company in the entire AI stack. Inference is memory. Memory is Micron and the inference ramp has barely started. Milk Road Pro members are already up massively on this position and we're just getting started. If you want the full breakdown of what we're buying and why, come join us for just a dollar using the link below!

Milk Road AI

130,756 просмотров • 3 месяцев назад