Загрузка видео...

Не удалось загрузить видео

На главную

Introducing The AI CUDA Engineer: An agentic AI system that automates the production of highly optimized CUDA kernels. The AI CUDA Engineer can produce highly optimized CUDA kernels, reaching 10-100x speedup over common machine learning operations in PyTorch. Our system is also able to produce highly optimized CUDA kernels...

1,160,948 просмотров • 1 год назад •via X (Twitter)

Комментарии: 10

Фото профиля vittorio
vittorio1 год назад

japan bros are back again

Фото профиля Globant
Globant1 год назад

🚀 Agentic AI Systems are changing what AI can do by having the power to act independently. Unlike traditional AI, which needs constant human supervision, these systems can operate more autonomously. Learn how this shift is leading to smarter solutions that can transform industries.➡️ #TechTrends2025

Фото профиля ludwig
ludwig1 год назад

I’m going to sleep if I wake up to this having 1M+ views I will read the paper tomorrow morning else pls give me a vibe check chat

Фото профиля Bing Xu
Bing Xu1 год назад

I quickly take a look of their report on phone, there are a few misleading parts: 1. Torch C++ code is not CUDA kernel, it is calling CUDNN under hood. 2. The highlighted example Conv3D GroupNorm, conv code is not generated at all. The speedup doesn’t make sense if numerical is wrong. 3. It claims wmma can be faster than PyTorch (CUBLAS), is definitely wrong. Probably benchmark error.

Фото профиля main
main1 год назад

isn't there clearly something wrong with level_1->15_Matmul_for_lower_triangular_matrices? claimed 152.9x speedup for the kernel on the left over the code on the right. really?

Фото профиля Viraat
Viraat1 год назад

Hey - wondering if you all are only working with large enterprises right now. If not, we’d love to chat! This would be extremely useful to us - we’re building low-bit models to run efficiently on Jetsons. Generating optimized CUDA code for these would be a game-changer!

Фото профиля aizk ✡️
aizk ✡️1 год назад

A slow 14 seconds in AI developments

Фото профиля Dan Mac
Dan Mac1 год назад

guys seriously I can't take it anymore need to slow down

Фото профиля Kristof
Kristof1 год назад

Japan is back

Фото профиля Dr Futuro - e/acc
Dr Futuro - e/acc1 год назад

Wow, Japan is back! 🇯🇵

Похожие видео

Ben Thompson explains how LLMs greatly diminished Nvidia's CUDA moat even as they sent the stock to the moon "So the weird thing about large language models is they were obviously incredible for Nvidia. That's why their stock went to the moon." "They have been on and off the most valuable company in the world." "It was also very bad for Nvidia. And the reason it was bad for Nvidia is that the play with CUDA is to build a developer ecosystem on top of CUDA." "But CUDA only works on Nvidia GPUs. So you get CUDA for free. It's easier to use, and it's a tremendous investment. Nvidia almost went under trying to build CUDA at a time when no one understood what they were doing or why they were wasting money on it." "And that's why Jensen Huang will get bristly, particularly when people question their rent-seeking or profit, whatever. It's like, no, they earned their spot fair and square." "Absolutely. It shouldn't be forgotten. They have earned every dollar they've gotten through 25 years of taking massive risks." "It bottomed out in October 2022. I wrote an article like three weeks before ChatGPT came out, tracing their bottoming-out history and their search for what was next." "'Nvidia in the Valley.' So, go back to this GTC. So I wrote an article at the time called 'Nvidia Waves and Moats'." "And what was interesting about that GTC was, number one, it was very boring. All the cool stuff kind of got scrubbed out." "Now, Jensen Huang has brought that stuff back, so the last few GTCs he's more talking about other things. Now it comes across as, oh, you're still looking for something beyond the LLM." "Because the problem with the LLM is it shifts the developer platform far above where Nvidia sits. All the activity is happening on top of LLMs. And so no one who's writing an AI application today is using CUDA." "Now, some people are, if you're training your own model and you're doing some low-level things or non-LLM things." "But the vast majority of the energy and all the money and the ecosystem is far removed from CUDA." "They have no idea and don't need to know or care what chips their application is running on. They're just on the OpenAI API, or the Anthropic API, or using Bedrock on Amazon, and it's sitting on Trainium, and they're using a Chinese open-source model. It's totally abstracted away, and this is why LLMs were bad for Nvidia." "Now, again, all the money they made along the way is worth it, but their moat has been tremendously diminished." "CUDA is still a moat if you need to do stuff that requires CUDA. But the vast majority of stuff, in energy, doesn't require CUDA, like in a post-LLM world."

Fireside Alpha

12,980 просмотров • 1 месяц назад