Загрузка видео...

Не удалось загрузить видео

На главную

Part 1: CUDA Graphs - Intro & Basics - how normal CPU ←→ GPU execution works - what a kernel launch is and where CPU launch overhead comes from - why launching many small kernels individually can become a bottleneck - what CUDA Graphs are and how they reduce...

29,135 просмотров • 2 месяцев назад •via X (Twitter)

Комментарии: 13

Фото профиля Manoj
Manoj2 месяцев назад

So if a beginner like me can follow u from day 1 Any prerequisites for this?

Фото профиля mohit
mohit2 месяцев назад

i would say you would still need some inference knowledge or context of if u get what I mean because I'm doing topics that I haven't deeply done or missing in upcoming days, so following my exact path as beginner won't work but if u happen to learn about cuda graphs or torch.compile, later ptx etc, you can easily watch and follow these topics

Фото профиля adi
adi2 месяцев назад

my goat

Фото профиля prasanna
prasanna2 месяцев назад

Bro one shottted ? Sick man keep it coming

Фото профиля mohit
mohit2 месяцев назад

nah lol, ~60 tries for 2 clips - its 2 seperate clips that i merged

Фото профиля Sisyphus
Sisyphus2 месяцев назад

Start a youtube

Фото профиля jbz
jbz2 месяцев назад

Please help me with DMing you.

Фото профиля mohit
mohit2 месяцев назад

sure

Фото профиля jbz
jbz2 месяцев назад

still can’t

Фото профиля mohit
mohit2 месяцев назад

dmed you

Фото профиля AI Apps API
AI Apps API2 месяцев назад

Great breakdown. Where this really bites is LLM inference decoding: you launch a flurry of tiny kernels per token and CPU launch overhead dominates at small batch sizes. Capturing the decode step as a graph and replaying it is one of the biggest easy wins for tokens per second.

Фото профиля Vishnu
Vishnu2 месяцев назад

🗿🗿

Фото профиля Vighnesh
Vighnesh2 месяцев назад

second follow

Похожие видео

Introducing The AI CUDA Engineer: An agentic AI system that automates the production of highly optimized CUDA kernels. The AI CUDA Engineer can produce highly optimized CUDA kernels, reaching 10-100x speedup over common machine learning operations in PyTorch. Our system is also able to produce highly optimized CUDA kernels that are much faster than existing CUDA kernels commonly used in production. We believe that fundamentally, AI systems can and should be as resource-efficient as the human brain, and that the best path to achieve this efficiency is to use AI to make AI more efficient! We are excited to publish our paper, The AI CUDA Engineer: Agentic CUDA Kernel Discovery, Optimization and Composition. We also release a dataset of over 17,000 verified CUDA kernels produced by The AI CUDA Engineer. Paper: Kernel Archive Webpage: HuggingFace Dataset: The AI CUDA Engineer utilizes evolutionary LLM-driven code optimization to autonomously improve the runtime of machine learning operations. Our system is not only able to convert PyTorch code into CUDA kernels, but through the use of evolution, it can also optimize the runtime performance of CUDA kernels, fuse multiple operations, and even discover novel solutions for writing efficient CUDA operations by learning from past innovations! We believe The AI CUDA Engineer opens a new era of AI-driven acceleration of AI and automated inference time optimization. We (Robert Lange, Aaditya Prasad 🇺🇸, sssss, Maxence Faldor, Yujin Tang, hardmaru) are excited to continue Sakana AI's mission of leveraging AI to improve AI.

Sakana AI

1,160,948 просмотров • 1 год назад

Build better RAG by letting a team of agents extract and connect your reference materials into a knowledge graph. Our new short course, “Agentic Knowledge Graph Construction,” taught by Neo4j Innovation Lead Andreas Kollegger, shows you how. Knowledge graphs are an important way to store information accurately but they are a lot of work to build manually. In this course you’ll learn how to build a team of agents that turn data– in this case product reviews and invoices from suppliers–into structured graphs of entities and relationships for RAG. Learn how agents can automatically handle the time-consuming work of building graphs — extracting entities and relationships (e.g., Product "contains" Assembly, Part "supplied_by" Supplier, Customer review "mentions" Product), deduplicating them, fact-checking them, and committing them to a graph database — so your retrieval system can find right information to generate accurate output. For example, you can use agents to help trace customer complaints directly to specific suppliers, manufacturing processes, and product hierarchies, thus turning fragmented information into queryable business intelligence. Skills you’ll gain: - Build, store, and access knowledge graphs using the Neo4j graph database - Build multi-agent systems using Google’s Agent Development Kit (ADK) - Set up a loop of agentic workflows to propose and refine a graph schema through fact-checking - Connect agent-generated graphs of unstructured and structured data into a unified knowledge graph This course gets into the practicum of why knowledge graphs give more accurate information retrieval than vector search alone, especially for high-stakes applications where precision matters more than fuzzy similarity matching. Sign up here:

Andrew Ng

168,153 просмотров • 1 год назад

Ben Thompson explains how LLMs greatly diminished Nvidia's CUDA moat even as they sent the stock to the moon "So the weird thing about large language models is they were obviously incredible for Nvidia. That's why their stock went to the moon." "They have been on and off the most valuable company in the world." "It was also very bad for Nvidia. And the reason it was bad for Nvidia is that the play with CUDA is to build a developer ecosystem on top of CUDA." "But CUDA only works on Nvidia GPUs. So you get CUDA for free. It's easier to use, and it's a tremendous investment. Nvidia almost went under trying to build CUDA at a time when no one understood what they were doing or why they were wasting money on it." "And that's why Jensen Huang will get bristly, particularly when people question their rent-seeking or profit, whatever. It's like, no, they earned their spot fair and square." "Absolutely. It shouldn't be forgotten. They have earned every dollar they've gotten through 25 years of taking massive risks." "It bottomed out in October 2022. I wrote an article like three weeks before ChatGPT came out, tracing their bottoming-out history and their search for what was next." "'Nvidia in the Valley.' So, go back to this GTC. So I wrote an article at the time called 'Nvidia Waves and Moats'." "And what was interesting about that GTC was, number one, it was very boring. All the cool stuff kind of got scrubbed out." "Now, Jensen Huang has brought that stuff back, so the last few GTCs he's more talking about other things. Now it comes across as, oh, you're still looking for something beyond the LLM." "Because the problem with the LLM is it shifts the developer platform far above where Nvidia sits. All the activity is happening on top of LLMs. And so no one who's writing an AI application today is using CUDA." "Now, some people are, if you're training your own model and you're doing some low-level things or non-LLM things." "But the vast majority of the energy and all the money and the ecosystem is far removed from CUDA." "They have no idea and don't need to know or care what chips their application is running on. They're just on the OpenAI API, or the Anthropic API, or using Bedrock on Amazon, and it's sitting on Trainium, and they're using a Chinese open-source model. It's totally abstracted away, and this is why LLMs were bad for Nvidia." "Now, again, all the money they made along the way is worth it, but their moat has been tremendously diminished." "CUDA is still a moat if you need to do stuff that requires CUDA. But the vast majority of stuff, in energy, doesn't require CUDA, like in a post-LLM world."

Fireside Alpha

12,980 просмотров • 1 месяц назад

Loops vs. Graphs, clearly explained! loops are great, but they have a ceiling: a loop makes one unit of work better. it cannot decide which units exist. so you end up with a very good agent running the wrong three steps, in the wrong order, one at a time. Graph engineering fixes this by moving the decision up a layer: what runs, what runs at the same time, and what never runs at all. you need both. here's how it works: a graph splits your system into two kinds of decision. ↳ inside a unit: the loop. produce, check, correct, repeat until green ↳ between units: the graph. split, fan out, merge, gate, send back Prompts → Context → Harness → Loops → Graphs you get parallel work, isolated contexts, and steps that stop running when nothing needs them. the trick is being selective about what becomes a node. only spend a model where judgment lives. merging, ranking, deduping and schema checks are edges, and edges are code. free, instant, and they cannot be argued out of a verdict. a graph where every edge is an agent pays rent on its own wiring. one thing to know before you scale it. a graph has two return paths, and almost everyone builds one. ↳ the correction edge is short. a gate rejects one unit back to the step that produced it, and it fixes the run you are in ↳ the learning edge is long. an accepted result goes back to the splitter as a constraint, and it fixes every run after skip the second and you get a graph that is fast and never gets smarter. next week it starts from the same place with the same blind spots. and a smaller one that eats whole nights: when a unit fails, return that unit, not the batch. send back four slices because one failed and you have just rewritten three correct ones. do it twice in a run and the run never converges. below i have quoted my full guide on graph engineering. it covers the three topologies, the verifier patterns, and where the gate should actually open. save this and read it below ↓

Hanako

73,867 просмотров • 1 месяц назад

Context vs. Graphs, clearly explained! context engineering is great, but it has a ceiling: you can make one window perfect. there is still only one of it. every technique on that layer is rationing the same scarce thing. compact, retrieve less, delete, defer. all of it is deciding what to drop. Graph engineering fixes this by moving the decision up a layer: not what goes in the window, but how many windows there are and what each one is for. you need both. here's how it works: ↳ inside a window: context engineering. what loads, in what order, what gets compacted ↳ between windows: the graph. how many lanes, what each one is allowed to see, what comes back Prompts → Context → Harness → Loops → Graphs each lane gets a clean window, nothing in one competes with anything in another, and your main thread stops filling up. the trick is being selective about what comes back. a subagent reads six thousand tokens of files and hands you a four hundred token summary. that ratio is the whole point. send back the raw material instead and you have moved the problem, not solved it. one thing to know before you scale it. not everything survives compaction equally, and almost nobody knows the table. ↳ the project-root rules file and auto memory are re-injected from disk. they come back intact ↳ path-scoped rules and nested rules files live in message history. they get summarized away and do not return until a matching file is read again so a rule that genuinely must persist cannot be path-scoped. move it to the root and pay the always-loaded cost, or accept that it is advisory in any long session. and the one that eats whole nights: shared context makes parallel agents converge. four auditors on one window produce one opinion with three echoes. you paid four times for it. below i have quoted my full guide on graph engineering. it covers the three topologies, the verifier patterns, and where the gate should actually open. save this and read it below ↓

Hanako

36,970 просмотров • 1 месяц назад