
Elliot Arledge
@elliotarledge • 30,601 subscribers
Shorts
Videos

timelapse #83 (22 hrs): - it was very easy to dive super deep into anything i needed to (this is what i focused on today because not all days are like this) - finding the grok code fast 1 + grok 4 for deep thinking and verification combo to be super useful in cursor. speed was solid - hard to imagine myself spending many more mental clock cycles in a 24 hr period - had to pull out qwen3-next’s gated deltanet + linear attention from bleeding edge hf transformers to begin implementing a multi-gpu fp8 trainer from scratch. this is so damn bleeding edge and i underestimated how much effort this has and will require lol - lots of diet coke and oats - shipped the template which the core chapters of my book will be built on: - all im missing now is flash attention 1/2 mastery (fa2 tmrw), intuition on making topk faster (for arbitrary row length), what i should and shouldnt teach in cutlass/cute, hopper/blackwell gemm kernel mastery (down to fp4) — shoutout to Pranjal for making this easier for me. his blog post is amazing - caught up w/ Mati Roy - im feeling great mentally but not so good physically as im writing this and about to pass out
Elliot Arledge2,255,154 views • 11 months ago

timelapse #165 (13 hrs) - stayed up til 3:30am getting the kimi k3 kernelbench article reviewed and polished (with help of the kimi team!) - private contract work through the day and night - optimizing a full game craftax cuda implementation on rtx pro 6000 - assembled a drone and flight tested it in 2 hrs (dead motor... so getting a replacement) - juggled other side projects
Elliot Arledge216,110 views • 1 month ago

timelapse #147 (15.5 hrs) - woke up to MiniMax (official) M3 launch including my kernelbench-hard (lowest score at below 30% which emphasizes the hardness) - did a space with the minimax and together ai folks - burned 1.1B tokens - got nanogpt at nvfp4 training stability to match bf16. this is a prereq for another problem im trying to solve - got my timelapse workflow nailed with a solid html page lol - loosing patience from anthropic rate limits
Elliot Arledge257,252 views • 3 months ago

timelapse #119 (15 hrs): - been trying out the 4am wake up and 8pm sleep schedule and its pretty good for getting ahead in the day. feels like less pressure is on me - upgraded my cpu and ram to ryzen 9 9950x3d + 96gb ddr5 - speedrunning contract work - planning out the entire software and hardware stack for what a space datacenter might look like and how its manufactured - built an MCP server called "thinkingcap" which ill be pushing out for you all today - getting used to opus 4.5 and gemini 3 pro still, although grok 4.1 has been very helpful to me
Elliot Arledge617,066 views • 9 months ago

Chang: "Explain this to me. Daniela runs day-to-day operations. All the leadership team reports to you." Dario: "Yes." Chang: "No one reports to you. That sounds like a pretty sweet job." Dario: "It's incredibly freeing. It lets me do all the things that I do much more easily than I would otherwise..." Chang: "And she does all the work. Is that what you're saying?" Dario: "If you had to go through the things I had to go through during the pandemic, the things I had to go through during the DoW... No"
Elliot Arledge149,262 views • 2 months ago

I wrote the full Craftax RL environment + PPO in one CUDA file. zero python in the loop. training from scratch on my RTX PRO 6000 > 0.6s: first wood, first crafting table, first dungeon entered > 1.7s: first zombie + skeleton kills, first bow found > 2.9s: first iron > 14s: first iron sword > 18s: first diamond > 71s: first diamond pickaxe > 16.7 min: first time reaching floor 3 of 9 (gnomish mines) > 75 min: 25,000,000,000 steps done up to 524,288 games running at once. env stepping alone hits 312M steps/s.
Elliot Arledge51,332 views • 1 month ago

timelapse #85 (27.5 hrs): - currently cant rely on any other coding models except grok code fast 1 + grok 4 fast (for complex reasoning grok 4 fast is 20 cents for 1M tokens) - wrote qwen3-next trainer entirely from scratch to make it more managable - each piece completely done by grok-code-fast-1 in cursor as it seems to handle this task pretty well without the grok 4 fast reasoning - take on smaller problems and complete them quickly (makes it easier with 400 toks/sec over the api) - got distributed fp8 qwen3-next trainer running at 0.8 seconds per step on 8xH100s (still need to finish checkpoint loading logic) - perfect timing as the fp8 version of qwen3-next drops as im writing this - ill be in LA in 2 days (will visit SF mid way through as well) - 12.5% margarita - steak dinner with family - gained intuition on FlashAttention in very long context settings - caught up w/ Kearm h/eng and Arnie Ramesh
Elliot Arledge283,820 views • 11 months ago

timelapse #162 (17 hrs) - Idk what it is, grok makes it super easy to get into flow state. No model does this for me - Went deep into spec decoding architecture learning with grok - Continued Minecraft reimplementation in C. The biggest bottleneck is samples to verify against (capturing my own gameplay for a few mins rn) as well as fly wheel speed for an agent to iterate. was bottlenecked by replay and compile, now bottlenecked by tokens. this is the hardest task ive given an llm to date. - Contract work - Getting tinygrad’s glm 5.2 at c=1 to be stable - more meetings
Elliot Arledge48,116 views • 1 month ago

This is my favorite clip of the new Elon pod. He opens up saying xAI struggles with memory usage/bandwidth and CUDA kernel optimization (matmul, attention, MoE, etc). If you are good kernel or performance engineering in general, you should apply. Steer the world in a better direction.
Elliot Arledge163,922 views • 7 months ago

timelapse #153 (11 hrs) - Woke up early and went to gym - Read Paul graham’s article and did some cold outreach for RL sim engineering (just wanna see where it goes) - got wireless keyboard and mouse. feels cleaner - Co-worked virtually with a friend - Trying out RL sims for battery materials optimization since the space seems pretty open for GPU optimization - Paused some ambitious projects for when fable comes back
Elliot Arledge57,927 views • 2 months ago

Timelapse #156 (36 hrs) - Worked with the tiny corp on getting GLM 5.2 running on 8xMI300X (sglang won here) - Launched KernelBench-Mega and updated Kernelbench-Hard with h100 and b200 sweeps - Took care of boring business stuff - Did some training sweeps for specialized technical vocab audio model - Some bugs with putting kernelbench-mega and hard on cloud instances so had to do some reruns. Learned a lot though - Setting up my own local rl infra and profiling concurrency 128 rollouts with vllm. Became clear to me that I need to serve in nvfp4, use MoE only for throughput, reap so training doesn’t OOM, dig into the vllm kernel graph itself to not underutilize my hardware from poor flashinfer/cutlass selections for my rtx pro 6000 sm120 architecture - Might do online distillation from glm 5.2 but for now taking it one step at a time - Slept for a bit then woke up and showered - Fixed an issue with SGLang tensor parallel deadlock on GLM 5.2 architecture with MTP enabled - GLM 5.2 inference is 2-3x faster than coding plans and running on amd boxes - Spent time with family - Hung out with some friends - Recorded some yoctogpt lectures with the revamped notebook (high taste btw) - Setting up dflash training for GLM 5.2
Elliot Arledge45,447 views • 2 months ago

timelapse #164 (15.5 hrs) - Burning midnight oil with the Kimi team benchmarking on kernels - Getting a Persistent Fused Mega Kernel for a Puffer Lib Environment (the full Craftax game) - Dipped for the night to have models run my overnight training runs on autopilot, including some kernel-based GRPO on qwen3.6 - private contract work - downloaded 1.1 TB of minecraft videos to train a foundation model on - optimized my minecraft foundation model training pipeline and kernels to train on 500k hrs of 128px at 10 fps in ~28 hrs on single gpu (5 hrs of learning per second on rtx pro 6000 blackwell) - doing it for the love of the game
Elliot Arledge28,210 views • 1 month ago

Timelapse #177 (14 hrs) - Picked up an RTX Pro 6000 Blackwell and wired in another 1000W PSU - running two of them now - DeepSeek v4 Flash running locally on both GPUs - Distilling DeepSeek v4 Flash down into Nanbeige (still cooking) - Kernel bench runs for DeepSeek v4 Flash and Qwen 3.8 Max - Got it into chat but couldn't reliably get it into the OMP harness yet - Private contract work - days like this are stuff I actually want to build. and when it's the high-intensity grind, i still have to show up and get it done
Elliot Arledge17,401 views • 29 days ago

timelapse #74 (11.5 hrs): - 95% done the most insane transformer training and inference chapter ever (competing w/ llm.c at this point) - talking with Luminal team - contract work - watching Minecraft videos while waiting for claude code and build scripts - starting learning multiple things at same time so I can parallelize chapter creation in my book based on what im feeling at a given moment - went a layer deeper into quantization: training challenges, group-wise vs block-wise vs tensor-wise vs channel-wise vs all the wises, input type vs compute type vs accumulate type vs epilogue, dealing w/ outliers
Elliot Arledge121,805 views • 11 months ago

timelapse #86 (15 hrs): - got my first OOM on 8xB200 node - defaulting back to grok-code-fast-1, the fastest reliable coding model with by far most intuitive instruction following, combined with grok 4 fast reasoning to plan before i let grok code work its magic - drank 2 large tim hortons iced capps, loaded myself w/ creatine, daily nootropics - tried out gpt-5-codex but it simply doesnt match the speed i require when i go deep into one thing at a time sequentially - got caught watching youtube videos in the middle, need to make sure i block any and all content that could get in my way - caught up on all book revisions so getting super ahead with other chapters - developed an overnight addiction to switching color themes in cursor - did some pair programming w/ Kearm h/eng using Tuple on free trial - applying for O-1
Elliot Arledge103,006 views • 11 months ago