Loading video...

Video Failed to Load

Go Home

Scale Python AI workloads from a single notebook to thousands of CPUs/GPUs. In #AzureFriday, Scott Hanselman 🌮 talks with Omar Shorbaji about building, training, and serving AI apps on Anyscale on Azure, powered by Ray, and running on AKS (without managing Kubernetes). Watch:

16,550 views • 2 months ago •via X (Twitter)

0 Comments

No comments available

Comments from the original post will appear here

Related Videos

Today we announced our new Fairwater datacenter in Atlanta, connected with our first Fairwater site in Wisconsin and our broader Azure footprint to create the world’s first AI superfactory. Fairwater exemplifies our vision for a fungible fleet: infra that can serve any workload, anywhere, on fit-for-purpose accelerators and network paths, with maximum performance and efficiency. AI workloads have evolved beyond large-scale pre-training. Today, they encompass fine-tuning, reinforcement learning (RL), synthetic data generation, evaluation pipelines, and more. Fairwater is built to support this full lifecycle: Max density: Fairwater’s two-story design and liquid cooling system lets us place racks in three dimensions and pack them with GPUs as densely as possible, minimizing cable runs and improving latency and effective bandwidth. Fleet: Each Fairwater DC can integrate hundreds of thousands of the latest NVIDIA GPUs into a single coherent cluster. This provides flexible infra that can support the full spectrum of workloads, and ensure no GPU is left unnecessarily idle. And that’s on top of the more than 100,000 GB300s coming online this quarter alone for inference across the rest of our fleet. For us, it’s all about turning every gigawatt into the maximum number of useful tokens. Not every GW is created equal! Planet-scale: Every Fairwater DC will connect through our continent-spanning AI WAN to prior generations of AI supercomputers, forming a truly fungible pool of compute. This enables developers to scale beyond the capacity of a single site and dynamically land workloads on the right infra for their needs. Together, these innovations let us bring together different generations of silicon and AI systems across DCs and geos into a single elastic system that scales seamlessly across training and inference workloads And this elastic AI capacity is all available alongside all the other cloud services (compute, storage, databases, app services) that AI agents and workloads need. This is what we mean when we talk about building a fungible fleet – a single, unified platform that pushes the limits of performance per watt and per dollar. Read more:

Satya Nadella

908,065 views • 9 months ago

🌟Quilibrium’s AI Breakthrough: Encrypted Training on CPUs In her latest live stream ( - minute 14) Cassie unveiled a groundbreaking AI training method that allows models to be trained on encrypted data using CPUs while achieving performance comparable to Nvidia’s A100 GPU (blue line in the graph below). Traditionally, AI training requires expensive GPUs because matrix multiplications—the core of deep learning—are highly computational. Running these calculations on CPUs is painfully slow, often taking hours or days for even small models. The problem worsens when trying to train AI on encrypted data, as standard encryption methods add a massive computational burden. Quilibrium’s breakthrough removes this bottleneck. Instead of relying on traditional matrix multiplication, their method uses a completely different mathematical approach, allowing AI models to be trained securely and efficiently without exposing the raw data. Cassie didn’t reveal the exact technique, only hinting that it’s inspired by existing AI research and will be detailed in a future open-source AGPL-licensed paper. The key advantage? AI can now be trained at GPU speeds on standard CPUs, making privacy-preserving machine learning far more accessible. This innovation has major implications. It slashes AI infrastructure costs, allowing organizations to train powerful models without investing in expensive hardware. It also enables private AI training on personal or corporate data without revealing sensitive information, a game-changer for industries like healthcare and finance. If Quilibrium’s method delivers on its promise, it could reshape AI development, making privacy-first computing the new standard. $QUIL $wQUIL

Quilibrium Community

18,919 views • 1 year ago