
PyTorch
@PyTorch • 510,343 subscribers
Tensors and neural networks in Python with strong hardware acceleration. PyTorch is an open source project at the Linux Foundation. #PyTorchFoundation
Videos

At PyTorch Conference North America 2026, hear directly from and connect with the people working on PyTorch, vLLM, DeepSpeed, ray, Helion, and Safetensors, alongside experts working across the AI stack. Keynote speaker Simon Mo (vLLM, Inferact) says vLLM’s goal is to become “the easiest to use and most efficient inference engine.” At #PyTorchCon NA 2025, he spoke about the difficulty and cost of using large language models and continuing to push inference efficiency forward. PyTorch conferences are the open source AI community’s town square, where what’s next gets decided. Register by September 4 to save on your conference pass:
PyTorch21,300 görüntüleme • 9 gün önce

PyTorch 2.13 brings FlexAttention to Apple Silicon, cuts peak memory by up to 4× for large-vocabulary models with nn.LinearCrossEntropyLoss, and updates distributed training, compilation, profiling, and on-device inference. Our live Q&A examined CUDA version support, CuTeDSL in TorchInductor, Python 3.15, ExecuTorch, torch.compile, and plans for PyTorch 2.14. Andrey Talman (Meta), albanD (Meta), and Piotr Bialecki (NVIDIA) joined moderator Chris Gottbrath (Gottbrath Technologies). This clip explains how FlexAttention compiles kernels specialized for the requested masking pattern, allowing them to skip unnecessary computation. It also covers speedups of up to approximately 12× over SDPA in tested sparse configurations and a deterministic backward path on CUDA that removes one source of nondeterminism in gradient computation. 🔗 Watch the full PyTorch 2.13 Release Live Q&A:
PyTorch18,591 görüntüleme • 1 ay önce

FlexAttention is a novel compiler-driven programming model that allows implementing the majority of attention variants in a few lines of idiomatic PyTorch code Boyuan Feng & Avik Chaudhuri show how many existing attention variants can be implemented via FlexAttention & that we achieve competitive performance compared to handwritten kernels. 📺 Watch the full video on YouTube:
PyTorch17,929 görüntüleme • 1 yıl önce
Daha fazla içerik yok.
