Video wird geladen...

Video konnte nicht geladen werden

Zur Startseite

Introducing The AI CUDA Engineer: An agentic AI system that automates the production of highly optimized CUDA kernels. The AI CUDA Engineer can produce highly optimized CUDA kernels, reaching 10-100x speedup over common machine learning operations in PyTorch. Our system is also able to produce highly optimized CUDA kernels...

1,160,518 Aufrufe • vor 1 Jahr •via X (Twitter)

10 Kommentare

Profilbild von vittorio
vittoriovor 1 Jahr

japan bros are back again

Profilbild von Globant
Globantvor 1 Jahr

🚀 Agentic AI Systems are changing what AI can do by having the power to act independently. Unlike traditional AI, which needs constant human supervision, these systems can operate more autonomously. Learn how this shift is leading to smarter solutions that can transform industries.➡️ #TechTrends2025

Profilbild von ludwig
ludwigvor 1 Jahr

I’m going to sleep if I wake up to this having 1M+ views I will read the paper tomorrow morning else pls give me a vibe check chat

Profilbild von Bing Xu
Bing Xuvor 1 Jahr

I quickly take a look of their report on phone, there are a few misleading parts: 1. Torch C++ code is not CUDA kernel, it is calling CUDNN under hood. 2. The highlighted example Conv3D GroupNorm, conv code is not generated at all. The speedup doesn’t make sense if numerical is wrong. 3. It claims wmma can be faster than PyTorch (CUBLAS), is definitely wrong. Probably benchmark error.

Profilbild von main
mainvor 1 Jahr

isn't there clearly something wrong with level_1->15_Matmul_for_lower_triangular_matrices? claimed 152.9x speedup for the kernel on the left over the code on the right. really?

Profilbild von Viraat
Viraatvor 1 Jahr

Hey - wondering if you all are only working with large enterprises right now. If not, we’d love to chat! This would be extremely useful to us - we’re building low-bit models to run efficiently on Jetsons. Generating optimized CUDA code for these would be a game-changer!

Profilbild von aizk ✡️
aizk ✡️vor 1 Jahr

A slow 14 seconds in AI developments

Profilbild von Dan Mac
Dan Macvor 1 Jahr

guys seriously I can't take it anymore need to slow down

Profilbild von Kristof
Kristofvor 1 Jahr

Japan is back

Profilbild von Dr Futuro - e/acc
Dr Futuro - e/accvor 1 Jahr

Wow, Japan is back! 🇯🇵

Ähnliche Videos

Ben Thompson explains how LLMs greatly diminished Nvidia's CUDA moat even as they sent the stock to the moon "So the weird thing about large language models is they were obviously incredible for Nvidia. That's why their stock went to the moon." "They have been on and off the most valuable company in the world." "It was also very bad for Nvidia. And the reason it was bad for Nvidia is that the play with CUDA is to build a developer ecosystem on top of CUDA." "But CUDA only works on Nvidia GPUs. So you get CUDA for free. It's easier to use, and it's a tremendous investment. Nvidia almost went under trying to build CUDA at a time when no one understood what they were doing or why they were wasting money on it." "And that's why Jensen Huang will get bristly, particularly when people question their rent-seeking or profit, whatever. It's like, no, they earned their spot fair and square." "Absolutely. It shouldn't be forgotten. They have earned every dollar they've gotten through 25 years of taking massive risks." "It bottomed out in October 2022. I wrote an article like three weeks before ChatGPT came out, tracing their bottoming-out history and their search for what was next." "'Nvidia in the Valley.' So, go back to this GTC. So I wrote an article at the time called 'Nvidia Waves and Moats'." "And what was interesting about that GTC was, number one, it was very boring. All the cool stuff kind of got scrubbed out." "Now, Jensen Huang has brought that stuff back, so the last few GTCs he's more talking about other things. Now it comes across as, oh, you're still looking for something beyond the LLM." "Because the problem with the LLM is it shifts the developer platform far above where Nvidia sits. All the activity is happening on top of LLMs. And so no one who's writing an AI application today is using CUDA." "Now, some people are, if you're training your own model and you're doing some low-level things or non-LLM things." "But the vast majority of the energy and all the money and the ecosystem is far removed from CUDA." "They have no idea and don't need to know or care what chips their application is running on. They're just on the OpenAI API, or the Anthropic API, or using Bedrock on Amazon, and it's sitting on Trainium, and they're using a Chinese open-source model. It's totally abstracted away, and this is why LLMs were bad for Nvidia." "Now, again, all the money they made along the way is worth it, but their moat has been tremendously diminished." "CUDA is still a moat if you need to do stuff that requires CUDA. But the vast majority of stuff, in energy, doesn't require CUDA, like in a post-LLM world."

Fireside Alpha

12,980 Aufrufe • vor 6 Tagen