
Alex Cheema
@alexocheema • 52,248 subscribers
building @exolabs | prev @UniOfOxford We're hiring: https://t.co/UlkApFndnH
Shorts
Videos

GLM-5.2 running on Mac Studio Clusters or RTX GPU rigs isn't how Local AI will get mass adopted. For mass adoption we need to focus on two axes: 1. Intelligence/Speed Frontier. This is being tracked on local dot ai. Aggressive co-design is needed across the entire stack to push up the frontier specifically for local hardware. Better kernels, inference tricks like spec decoding, quantizing models, model routing. 2. Usability / UX. To use Local AI today, you need to install CLIs, select a model, a quantization, a harness that works best with that model, and you need to configure your inference engine. It's not very accessible. More work needs to be done on making this seamless, so the user doesn't even know the AI is running locally.
Alex Cheema30,443 Aufrufe • vor 11 Tagen

2 MacBooks is all you need. Llama 3.1 405B running distributed across 2 MacBooks using exo intern home AI cluster
Alex Cheema1,288,858 Aufrufe • vor 2 Jahren

"Somebody got one of the small versions of Llama to run on Windows 98...We could've been talking to our computers in English for the last 30 years" - Marc Andreessen 🇺🇸 It was me! I got Llama running on a Pentium II machine with 128MB RAM running Windows 98. Details below.
Alex Cheema765,909 Aufrufe • vor 1 Jahr

NVIDIA sent us 2 DGX Sparks. For a while we wondered what we would do with them. The memory bandwidth is 273GB/s making it 3x slower than an M3 Ultra (819GB/s) for batch_size=1 inference. But it has 4x more FLOPS (100 TFLOPS compared to 26 TFLOPS). So we thought, what if we could combine the DGX Spark & M3 Ultra, and make use of both the massive compute on the DGX Spark and the massive memory-bandwidth on the M3 Ultra. We came up with a way to split inference across both devices and achieve a speedup of up to 4x for long prompts compared to the M3 Ultra on its own. Full details in the blog post linked below.
Alex Cheema281,225 Aufrufe • vor 9 Monaten

NVIDIA DGX Spark. World’s smallest AI supercomputer with 128GB memory
Alex Cheema312,131 Aufrufe • vor 1 Jahr

Running GLM-4.7-Flash on 4 x M4 Pro Mac Minis using EXO Labs. Uses tensor parallelism with RDMA over Thunderbolt & MLX backend (h/t Awni Hannun). Runs at 100 tok/sec. We're working on optimizing this at EXO Labs. Aiming to hit ~200 tok/sec on this setup soon.
Alex Cheema62,195 Aufrufe • vor 6 Monaten

Running GLM 4.7 Flash (8-bit) with Tensor Parallel / RDMA on 2 M4 Pro Mac Minis at 60 tok/sec. mlx-lm 0.30.5 features huge speedups for GLM 4.7 Flash for long context (h/t N8 Programs & Awni Hannun). M5 Pro (~28 Jan) will have ~4x faster prefill and ~1.3x faster decode.
Alex Cheema56,555 Aufrufe • vor 5 Monaten

"Mac Minis for example are a very good fit" - Andrej Karpathy Andrej Karpathy shouted out my work on EXO Labs in his keynote at Y Combinator AI SUS! Here's the breakdown: Right now most AI workloads run in the cloud where requests from different users are continuously batched together. These workloads are FLOPS-bound and favors hardware with the best unit economics of $ per FLOP, i.e. enterprise GPUs. The personal computing revolution will shift these workloads to personal devices with lower batch sizes (mostly batch_size=1). batch_size=1 inference is memory-bound, because all of the model parameters need to be loaded into the GPU every time a token is generated. Apple Silicon with its Unified Memory architecture has a lot of memory and memory bandwidth per $ compared to other hardware: - M4 Pro Mac Mini, 24GB @ 273GB/s, $58.33/GB, $5.13/GB/s - H100, 80GB @ 3350GB/s, $625/GB, $14.93/GB/s The unit economics of Apple Silicon are becoming more compelling with every release of the Mac. The future of AI inference looks more like open weights models (OpenAI open weights model soonTM) run at low batch_size on personal devices.
Alex Cheema103,566 Aufrufe • vor 1 Jahr

Latest Fireship video featured my run of DeepSeek R1 on M4 Mac Minis. Apple Silicon dominates in memory/bw unit economics, ideal for huge MoE models like R1 at batch_size=1 (the real-world use-case). SOTA AI may come from China, but it will run on American hardware.
Alex Cheema - e/acc79,775 Aufrufe • vor 1 Jahr
Keine weiteren Inhalte verfügbar