正在加载视频...

视频加载失败

Introducing The AI CUDA Engineer: An agentic AI system that automates the production of highly optimized CUDA kernels. The AI CUDA Engineer can produce highly optimized CUDA kernels, reaching 10-100x speedup over common machine learning operations in PyTorch. Our system is also able to produce highly optimized CUDA kernels...

1,160,948 次观看 • 1 年前 •via X (Twitter)

10 条评论

vittorio 的头像
vittorio1 年前

japan bros are back again

Globant 的头像
Globant1 年前

🚀 Agentic AI Systems are changing what AI can do by having the power to act independently. Unlike traditional AI, which needs constant human supervision, these systems can operate more autonomously. Learn how this shift is leading to smarter solutions that can transform industries.➡️ #TechTrends2025

ludwig 的头像
ludwig1 年前

I’m going to sleep if I wake up to this having 1M+ views I will read the paper tomorrow morning else pls give me a vibe check chat

Bing Xu 的头像
Bing Xu1 年前

I quickly take a look of their report on phone, there are a few misleading parts: 1. Torch C++ code is not CUDA kernel, it is calling CUDNN under hood. 2. The highlighted example Conv3D GroupNorm, conv code is not generated at all. The speedup doesn’t make sense if numerical is wrong. 3. It claims wmma can be faster than PyTorch (CUBLAS), is definitely wrong. Probably benchmark error.

main 的头像
main1 年前

isn't there clearly something wrong with level_1->15_Matmul_for_lower_triangular_matrices? claimed 152.9x speedup for the kernel on the left over the code on the right. really?

Viraat 的头像
Viraat1 年前

Hey - wondering if you all are only working with large enterprises right now. If not, we’d love to chat! This would be extremely useful to us - we’re building low-bit models to run efficiently on Jetsons. Generating optimized CUDA code for these would be a game-changer!

aizk ✡️ 的头像
aizk ✡️1 年前

A slow 14 seconds in AI developments

Dan Mac 的头像
Dan Mac1 年前

guys seriously I can't take it anymore need to slow down

Kristof 的头像
Kristof1 年前

Japan is back

Dr Futuro - e/acc 的头像
Dr Futuro - e/acc1 年前

Wow, Japan is back! 🇯🇵

相关视频

Ben Thompson explains how LLMs greatly diminished Nvidia's CUDA moat even as they sent the stock to the moon "So the weird thing about large language models is they were obviously incredible for Nvidia. That's why their stock went to the moon." "They have been on and off the most valuable company in the world." "It was also very bad for Nvidia. And the reason it was bad for Nvidia is that the play with CUDA is to build a developer ecosystem on top of CUDA." "But CUDA only works on Nvidia GPUs. So you get CUDA for free. It's easier to use, and it's a tremendous investment. Nvidia almost went under trying to build CUDA at a time when no one understood what they were doing or why they were wasting money on it." "And that's why Jensen Huang will get bristly, particularly when people question their rent-seeking or profit, whatever. It's like, no, they earned their spot fair and square." "Absolutely. It shouldn't be forgotten. They have earned every dollar they've gotten through 25 years of taking massive risks." "It bottomed out in October 2022. I wrote an article like three weeks before ChatGPT came out, tracing their bottoming-out history and their search for what was next." "'Nvidia in the Valley.' So, go back to this GTC. So I wrote an article at the time called 'Nvidia Waves and Moats'." "And what was interesting about that GTC was, number one, it was very boring. All the cool stuff kind of got scrubbed out." "Now, Jensen Huang has brought that stuff back, so the last few GTCs he's more talking about other things. Now it comes across as, oh, you're still looking for something beyond the LLM." "Because the problem with the LLM is it shifts the developer platform far above where Nvidia sits. All the activity is happening on top of LLMs. And so no one who's writing an AI application today is using CUDA." "Now, some people are, if you're training your own model and you're doing some low-level things or non-LLM things." "But the vast majority of the energy and all the money and the ecosystem is far removed from CUDA." "They have no idea and don't need to know or care what chips their application is running on. They're just on the OpenAI API, or the Anthropic API, or using Bedrock on Amazon, and it's sitting on Trainium, and they're using a Chinese open-source model. It's totally abstracted away, and this is why LLMs were bad for Nvidia." "Now, again, all the money they made along the way is worth it, but their moat has been tremendously diminished." "CUDA is still a moat if you need to do stuff that requires CUDA. But the vast majority of stuff, in energy, doesn't require CUDA, like in a post-LLM world."

Fireside Alpha

12,980 次观看 • 1 个月前