Video wird geladen...

Video konnte nicht geladen werden

Zur Startseite

Today PyTorch releases 2.9.1 GA which comes with fixes to cuDNN attention numerical issues. End users of cuDNN in PyTorch have been running into NaN & convergence issues with cuDNN as such cuDNN was disabled by default in PyTorch until 2.9.0. But unfortunately in 2.9.0, end users also reported...

24,842 Aufrufe • vor 10 Monaten •via X (Twitter)

0 Kommentare

Keine Kommentare verfügbar

Kommentare vom Original-Post werden hier angezeigt

Ähnliche Videos

I gave a talk at GPU MODE workshop last week on llm.c - the origin story of llm.c - being naked in the world without PyTorch and having to re-invent Array, Autograd, Device, Dtype, Compile, Distributed - how to port a PyTorch layer to 1) explicit PyTorch - and then to 2) write the backward pass - 3) port forward & backward pass to C - 4) string all the layers together - achieving one file of C with no dependencies that compiles and runs ~instantly, where all memory is pre-planned and allocated a single time, fully deterministic, portable code that can run on a potato or a von Neumann probe - how most of llm.c was built at 1am-7am in a water villa porch in Maldives and why this is the recommended way to develop software - convert all of it to run in CUDA on GPU in fp32 - port matmul to cuBLAS - port attention to cuDNN flash-attention - introduce bfloat16 mixed precision - introduce many more optimizations and features like kernel fusions, Packed128, stochastic rounding, full determinism - add multi-GPU training, NCCL, sharded optimizer - add multi-node with MPI or file system or socket - reproduce GPT-2 (1.6B) on one 8XH100 node in 24 hours for $672 in llm.c, achieving (at the time) 29% less memory, 19% faster training that PyTorch nightly, and much faster compile & run - how open source development attracts Avengers from the internet - port to training Llama 3 imminent (branch exists) - many other notable forks - last thought: how software abstractions like Python/PyTorch and everything else really exist only because humans are finite in knowledge, IQ and attention, and how with increasing AI capability LLMs may export custom binaries like llm.c for any application directly, tearing apart and refactoring all abstractions as needed. More links in reply

Andrej Karpathy

337,695 Aufrufe • vor 1 Jahr

New short course: Attention in Transformers: Concepts and Code in PyTorch. Last week we released a course on how LLM transformers work. This week, go deeper and learn about the technical ideas behind the attention mechanism, and see how to code it in PyTorch. This course is built with Joshua Starmer, Founder and CEO of StatQuest. The attention mechanism was a breakthrough that led to transformers, the architecture powering large language models like ChatGPT. Transformers, introduced in the 2017 paper: "Attention is All You Need" by Viswani and others, took off because of its highly scalable design. In this course, you’ll learn how the attention mechanism, a key element of transformer-based LLMs, works and implement it in PyTorch. You'll develop deep intuition about building reliable, functional, and scalable AI applications. What you will do: - Understand the evolution of the attention mechanism, a key breakthrough that led to transformers. - Learn the relationships between word embeddings, positional embeddings, and attention. - Learn about the Query, Key, and Value matrices, and how to produce and use them in attention. - Walk through the math required to calculate self-attention and masked self-attention to learn why and how they work. - Understand the difference between self-attention and masked self-attention and how one is used in the encoder to build context-aware embeddings and the other is used in the decoder for generative outputs. - Learn the details of the encoder-decoder architecture, cross-attention, and multi-head attention and how they are all incorporated into a transformer. - Use PyTorch to code a class that implements self-attention, masked self-attention, and multi-head attention. There're lots of exciting technical details in this course. Please sign up here:

Andrew Ng

132,400 Aufrufe • vor 1 Jahr

Google just launched a direct attack on Nvidia's most valuable asset. Not their chips. Their SOFTWARE. And if this works, Nvidia's $4 trillion empire collapses. Here's what just leaked: Google is building "TorchTPU" - a secret project that makes PyTorch seamlessly run on Google's TPU chips instead of Nvidia GPUs. Why does this matter? PyTorch is the MOST USED AI framework on Earth. Every AI developer uses it. And PyTorch was built around Nvidia's CUDA software. Wall Street analysts call CUDA "Nvidia's strongest defensive wall." It's the reason companies can't easily switch away from Nvidia even when alternatives exist. You don't just buy Nvidia chips. You buy into their entire ecosystem. Switching costs MILLIONS in engineering work. Months of rewrites. Performance drops. So companies stay locked in. Even when Nvidia raises prices. Even when supply runs short. That's not a hardware moat. That's a SOFTWARE prison. And Google just found the escape route. Here's the problem Nvidia created for itself: Google's TPU chips are actually GOOD. Competitive performance. Better availability. Lower cost. But developers won't use them because Google's chips run JAX (Google's internal framework), not PyTorch. That means if you want to use Google TPUs, you have to rewrite your entire codebase. Nobody wants to do that. So Google TPUs sit unused while developers fight over Nvidia chips. Until now. TorchTPU makes PyTorch run natively on Google hardware. No rewrites. No performance loss. No months of engineering. You just... switch. And Google is partnering with META (who built PyTorch) to make it happen. They're even considering OPEN-SOURCING parts of it to speed adoption. Translation: Google is willing to give this away for free just to break Nvidia's lock. The implications are insane: Every company currently paying Nvidia's premium prices suddenly has a way out. Oracle, Microsoft, OpenAI - all locked into Nvidia's ecosystem - can switch to Google. Nvidia's pricing power evaporates overnight. And the timing is perfect: Nvidia is already facing heat. Semiconductor index dropped 3% today. Oracle just lost their biggest investor over AI spending concerns. Companies are realizing AI infrastructure costs are unsustainable. Now Google hands them an alternative. Same performance. Lower cost. Better availability. Jensen Huang knows exactly what this means. CUDA has been Nvidia's untouchable advantage for YEARS. It's why Nvidia trades at 50x earnings while AMD trades at 25x. The software moat justified the premium. But if Google removes that switching cost? Nvidia becomes just another chip company. And chip companies compete on price, not ecosystem lock-in. Here's what happens next: Google needs 12-18 months to make TorchTPU production-ready. If it works, cloud providers will adopt it instantly. They WANT an alternative to Nvidia's monopoly pricing. Amazon already building their own Trainium chips. Microsoft making Maia. They're all trying to escape Nvidia. Google just gave them the software bridge. Nvidia's response options are limited: They can't buy Google. Can't kill PyTorch (Meta owns it). Can't stop open source. Their only play is to keep improving CUDA faster than Google can catch up. But that's a race, not a moat. The market isn't pricing this in yet. Nvidia down 2% today. Google down 2%. Investors think this is just "another competitor." They don't understand this is an attack on the FOUNDATION of Nvidia's valuation. Hardware is replaceable. Software lock-in is what made Nvidia worth $4 trillion. Google is attacking the lock-in. Watch what happens in 2026 when TorchTPU goes live and companies realize they can actually leave Nvidia. The "Nvidia is unstoppable" narrative dies. And a $4 trillion valuation built on software moats gets repriced.

Ricardo

1,617,348 Aufrufe • vor 8 Monaten

Q: How do you decide which customers to listen to? As Superhuman founder & CEO Rahul Vohra puts it: “In a world where you’re drowning in feedback—and most startups are drowning in feedback—you have to filter it down to only the stuff that’s going to increase the number of people who fall in love with your product.” Most startups will listen to all feedback from on-the-fence customers, but this isn’t targeted enough and will often lead to a muddled, incoherent product. As Rahul argues in the clip below, you need to identify the main benefit of your product—for Superhuman this was speed. And then focus on the feedback of on-the-fence users who also view this as the main benefit—there’s often something small holding them back. Users for whom your main benefit does not resonate (e.g. Superhuman users who value offline capabilities rather than speed), are unlikely to ever fall in love with your product. When Superhuman ran this analysis in 2015, they found that the main thing holding back users who viewed speed as the main benefit was their lack of a mobile app. Probing further, they found some less obvious and more interesting requests, such as integrations, attachment handling, calendering, unified inbox and read receipts. With a clear understanding of their main benefit and missing features, they were able to move this cohort of users from on-the-fence into the territory of enthusiastic advocates. As Rahul puts it in his Product Market Fit Engine article: “To increase your product/market fit score, spend half your time doubling down on what users already love and the other half on addressing what’s holding others back.” But make sure you’re focusing on users who love the main benefit of your product. Users who don’t are unlikely to ever fall in love with your product.

Michael McGuiness

89,882 Aufrufe • vor 3 Jahren