LLMs require more GPU memory as they generate longer... responses. Can we make GPU memory constant without significantly sacrificing accuracy? IceCache is a new method for managing KV caches that leverages Dynamic Continuous Indexing (DCI) to efficiently group and retrieve tokens by semantics. Joint work w/ Yuzhen Mao, Qitong Wang and Martin Ester. For details, check out the links below.show more

Ke Li 🍁
21,163 次观看 • 5 个月前
🚨 YOUR GPU IS PROBABLY WASTING MORE THAN YOU... THINK. vLLM is built to squeeze far more useful work out of your GPU when serving LLMs. Running an LLM at scale isn’t just about having a powerful GPU. The real problem is how efficiently you use its memory and compute. That’s where vLLM comes in. → High-throughput LLM inference and serving → PagedAttention for smarter KV-cache memory management → Continuous batching to keep GPUs busy → Prefix caching + chunked prefill → OpenAI-compatible API out of the box → Supports a huge range of modern LLM architectures → Quantization support for running models more efficiently And the crazy part? You can start an OpenAI-compatible inference server with: `vllm serve ` So your application can talk to your own model almost like it’s talking to OpenAI. The bigger idea: Don’t just buy more GPUs. Make the GPUs you already have work harder. That’s why vLLM has become such a major project in LLM inference. 🔥 #vLLM #AI #LLM #Inference #GPU #MachineLearning #AIInfrastructure #OpenSource #AIAgents #Developersshow more

Vikas gupta
14,138 次观看 • 24 天前
Day 11/90 of Inference Engineering How does vLLM work... and how is it used in production? Before we discuss how vLLM works internally, it helps to understand what vLLM is. At a high level, vLLM is an inference engine that is designed to serve LLMs to thousands of concurrent users efficiently while managing scarce compute and memory. The goal for vLLM is to maximize throughput and minimize latency; optimizing for the best inference economics and experience for end users. With every request from the end user, it eventually ends up in the engine core, gets scheduled alongside other requests from other concurrent users, executes on the GPU, and updates the KV cache with the new key and value vectors, and streams the tokens back to the user. The Scheduler decides what requests should execute next while continuously batching requests together to maximize GPU utilization. Continuous batching is an inference optimization that allows new requests to join a running batch as other requests finish generating tokens. This helps with keeping the GPU utilization high instead of letting it sit idle waiting for an entire batch to complete generating. After the scheduler dispatches the selected batch to the Model Executor, the Model Executor prepares the tensors and metadata required for inference, retrieves each request’s block table from KV Cache Manager, launches the optimized transformer forward pass on the GPU, computes the logits, updates the KV cache with the new key and value vectors, and finally returns the results for sampling and streaming. The KV Cache Manager uses the PagedAttention memory layout to allocate fixed-size cache blocks on demand and maintains a Free Block Queue on the CPU that tracks which blocks in the GPU’s Paged KV Cache are currently free. When a request needs additional KV cache space, the KV Cache manager takes a free block from the queue and assigns it to that request, thus avoiding an expensive search through GPU memory for available cache blocks. All of these components form the core of vLLM’s inference engine. The Scheduler determines what requests are executed, the Model Executor determines how those requests are executed, the KV Cache Manager determines where each request’s KV cache lives using the PagedAttention Memory Layout. This architecture enables vLLM to serve thousands of concurrent requests with high throughput, low latency, and efficient GPU memory utilization. Heres a little animation that visualizes everything! - I've also completed the forward pass for my mnist.c project. I had a nice chat with shrey birmiwal, such a knowledgeable guy. Excited to learn more about vLLM and implement a tiny-vLLM one day.show more

max fu
70,797 次观看 • 2 个月前
1-bit Qwen3.8-27B just more than DOUBLED its MMLU-Pro score.... They didn’t change the decoder. 29.04% -> 61.54%, according to a new paper from ISTA-DASLab. Same lab behind GSQ-RCO, some of the best tiny Qwen3.8-27B quants in my book. Their method is called Disaggregated Quantization. You use different weights to read the prompt and generate the answer. A trained NVFP4 prefiller processes your input and builds the cache. Then the tiny GGUF decoder takes over. They train the prefiller specifically for that decoder, so it learns to produce representations the heavily compressed model can use. The extra checkpoint is 12.8 GiB, but you don’t need all of it in VRAM. Stream it from SSD layer by layer, reusing GPU memory. With a long enough prompt, loading can overlap computation. That’s a pretty fucking good reason to pay attention if your GPU is short on memory. I want to test this on my 4x3090 rig with NVMe. The released implementation targets Blackwell, so this needs adaptation. First I’ll check whether the quality gains survive on Ampere, then measure whether offloading actually helps. Paper:show more

Alexey Fateev
19,210 次观看 • 7 天前
Free NVIDIA GPU with 16 GB VRAM GPU for... Running Local LLMs! If you want to master local LLMs but you're waiting until you can afford a $1,500 GPU, you're honestly not going to make it. The open source AI ecosystem is moving way too fast for you to wait on your budget to catch up. Especially when you can build a bleeding edge inference engine from scratch right now, completely for free. You don't need a heavy local rig to start. Google is literally letting you use an enterprise grade NVIDIA Tesla T4 GPU for $0/hour. At standard cloud computing rates (~$0.20/hr), Google Colab’s 4 hour daily free tier hands you roughly $24 worth of data center tier GPU compute every single month. And most people just waste it. Let’s talk about the hardware you get access to for free. The NVIDIA Tesla T4 is an absolute workhorse: - Architecture: NVIDIA Turing (TU104) - VRAM: 16GB GDDR6 (320 GB/s bandwidth) - Compute: 320 Tensor Cores | 2560 CUDA Cores - Performance: 130 TOPS INT8 | 8.1 TFLOPS FP32 - Power: Sipping energy at a max 70W TDP This is the exact same hardware I used to run DeepMind's Gemma 4 26B A4B QAT MoE at a 250,000 context window without a single Out Of Memory (OOM) crash. If you have a web browser and 10 minutes, you have everything you need. I’ve put together a fully documented, cell by cell Google Colab notebook that teaches you exactly how to do this. Here is what the notebook actually teaches you: - How to provision an Ubuntu Linux environment with CUDA 13.0 and verify your driver stack. - How to pull the source code and compile the latest llama.cpp C++ binaries from scratch, specifically optimizing the build for your exact GPU using the -DCMAKE_CUDA_ARCHITECTURES=native flag. - How to directly download quantized local LLMs (GGUF format) straight from HuggingFace using the CLI. - How to manage 16GB VRAM limits, offload neural network layers to the GPU, and push massive context windows. Compile raw llama.cpp, ollama run a model, or spin up the LM Studio CLI. Pick whatever stack you are comfortable with. just start building. No hardware. No credit card. No excuses. Bookmark this post right now so you don't lose the tutorial. Even if you don't have time to run it today, you are going to want this workflow in your engineering toolkit. The link to the free Colab Notebook is in the comments below. Lemme know if you need more tutorials like this.show more

Alok
182,483 次观看 • 3 个月前
Introducing StreamingLLM. Imagine chatting with an AI assistant that... can contextually reference your conversations from weeks or months ago. Or summarizing reports that span thousands of pages. StreamingLLM makes this possible by enabling language models to smoothly handle endless texts without losing steam. Current LLMs are like students cramming for an exam - they can only memorize a limited context. StreamingLLM is the valedictorian with a photographic memory of everything you've ever discussed. It works by identifying and preserving the model's inherent "attention sinks" - initial tokens that anchored its reasoning. Combined with a rolling cache of recent tokens, StreamingLLM delivers up to 22x faster inference without any drop in accuracy. You know that irksome feeling when chatbots forget your earlier conversations? StreamingLLM abolishes that frustration. It remembers the touchdowns from your last game and your newborn's name without missing a beat. Monumental books, verbose contracts, drawn out debates - StreamingLLM takes them all in its stride. No shortcuts, no forgetfulness. It's like upgrading your assistant's RAM to handle heavier workloads flawlessly.show more

Carlos E. Perez
557,849 次观看 • 3 年前
The Whistle Tactic by Delhi Police You may have... seen recent protest videos at Jantar Mantar with the continuous sound of whistles. This is a new tactic adopted by Delhi Police against the protesters. Let's break down the psychology behind continuous whistling: 1. It disrupts communication The loud whistles make it difficult for protesters to hear speakers, volunteers and each other. 2. It breaks the rhythm of the protest Slogans, speeches and coordinated responses become difficult when the crowd cannot hear clearly. 3. It affects media interviews When a protester speaks to a journalist, continuous whistling can make the audio difficult to understand. The video is recorded, but the message can get lost. 4. It creates irritation and fatigue A continuous high-pitched sound can cause stress, frustration and mental fatigue. Some people may eventually leave simply because they cannot tolerate the noise. 5. It creates confusion When people cannot hear instructions, coordination becomes harder. A large, organized crowd can gradually become a collection of smaller, disconnected groups. 6. It creates psychological pressure Constant whistles reinforce the feeling that the police control the environment. This can make protesters feel intimidated even without direct physical intervention. Let us come up with ideas about how to counter this tactic of Delhi Police.show more

Kapil
49,414 次观看 • 20 小时前
🚨 Anthropic committed up to 1M TPU chips for... Claude. Openai is leasing TPUs for chatgpt inference. Here's How kernels work on TPUs (deep dive 2/6 by emilio andere) pallas is Google's answer to kernel writing. a python kernel SDK built on JAX. still very experimental (jax.experimental.pallas). on TPU it compiles through mosaic; on GPU it lowers to triton. if you know CUDA, the syntax will feel familiar but the execution model is completely different. in CUDA, grid=(4,4) launches 16 blocks running simultaneously across SMs. in pallas, those 16 iterations run one after another in lexicographic order. no threads. no warps. no blocks. no occupancy tuning. a TPU is a sequential machine with a very wide vector register — more like a CPU than a GPU. performance comes from width: a 128x128 systolic array doing matmul and an 8x128 SIMD vector unit doing everything else. maximum parallelism on chip: 2, one per TensorCore in megacore mode. three concepts replace CUDA's thread/block/grid hierarchy. Refs are mutable memory references. because execution is sequential, each iteration safely accumulates without atomics. in CUDA you'd need atomics or a separate reduction pass. the memory model is also very different from NVIDIA's. zero hardware caches. VMEM is 32-128 MiB of software-managed scratchpad — 500-1000x larger than GPU shared memory per SM. all data must be explicitly DMA'd from HBM to VMEM before any computation touches it. four levels: HBM → VMEM → VREGs → MXU/VPU, plus SMEM for scalar control data. every byte of data movement is your responsibility. this is like CUDA shared memory except it's 500x bigger and there's no cache fallback. pipelining is mandatory. without double-buffering HBM→VMEM transfers, the MXU just stalls waiting for data. this is the single most important optimization on TPU. and because grid execution is sequential and deterministic, consecutive iterations that need the same input block skip the redundant HBM transfer automatically, impossible on GPU where block execution order is undefined. the compilation pipeline is unlike anything in this series: python → jaxpr → stableHLO → XLA HLO (71+ optimization passes) → LLO (78+ passes) → 322-bit VLIW bundles. the compiler packs instructions for scalar, vector, matrix, and DMA units into a single 322-bit word. everything in that bundle executes in parallel, with no runtime scheduling.show more

wafer
33,655 次观看 • 2 个月前
A wrist force sensor fires at 100Hz. The policy... only ever sees it at 30Hz, downsampled to land on the same control step as the camera and the joint state. That's not a bug, it's the whole point, and it sits inside a bigger pattern in VLA research this year. Every major release has been Markovian at its core, mapping the current frame straight to the next action. The fix everyone reaches for is more vision: more history frames, longer image context. FM-VLA makes a clean case that the fix is the wrong channel for a whole class of tasks. Press a button three times and stop. A camera watching that has almost nothing to work with, the scene barely changes between press one and press three. Force doesn't have that ambiguity problem. Each press is a sharp, distinct spike in the wrench signal, whether or not the camera noticed anything at all. So FM-VLA doesn't add more frames. It compresses the wrench history into eight tokens with a VAE, pretrained purely on reconstructing force signals, frozen before it ever touches the policy, then hands those tokens to the action expert alongside a short window of joint state. That's the entire memory system. Averaged across three contact-rich tasks, FM-VLA hits 83.3 percent success against 33.3 percent for the strongest vision-memory baseline on the button-counting task specifically, where the ambiguity problem is worst, 72.2 percent for FM-VLA there. Strip out the short-state window and force-only performance drops well below the combined system, so force alone isn't the answer either. The two channels are doing different jobs. The field has defaulted to one memory channel for every kind of temporal problem. This is a clean data point that the channel should match the ambiguity you're actually trying to resolve, not just get bigger. Source: Paper: Credit to the teams at Tsinghua University, Microsoft Research, and Fudan University. #Robotics #PhysicalAI #RobotLearningshow more

Stephen James
11,658 次观看 • 2 个月前
A good technical LLM interview question: Your LLM chatbot... takes 12s before it generates the first token, and the users are complaining. So you move the model onto a GPU with 3x the computing power. The time to first token barely improves. Why did this happen? (answer below) Latency in an LLM app is a placement problem disguised as a model problem. If you profile the 12 seconds, the model's prefill itself may only account for around 1.5 seconds of it. So halving the prefill step saves just 750ms out of 12000, which is under 7%. The rest is spread across stages that never touch the GPU. The request first travels to whatever region the app runs in, and a cross-continent round trip could cost over a second before any code executes. Then the request handler starts. On a container-based serverless platform under load, this adds several seconds of cold start, paid before auth, rate limiting, or prompt assembly even begins. Retrieval adds its own hop, and the response streams back across the same distance. Optimizing a stage that was already fast cannot alter the latency that's majorly affected by other stages. Those other stages are slow for a structural reason. An LLM app runs two workloads that want opposite machines. - The request path is short, spiky, and needs to sit close to users - Inference is long-running, GPU-bound, and billed hourly, whether requests arrive or not. So the actual decision is not which model to run, but where each of these two workloads runs. There are three options, each with its own tradeoffs: > A dedicated GPU box removes inference cold starts, but it bills around the clock and lives in one location, so distant users wait out the round trip on every request > Container-based serverless scales to zero, but the request path pays a cold start, and most of these platforms have no GPU behind them. > Edge runtimes start in under a millisecond, because a WebAssembly module carries no OS or container image to boot. They handle the request path well and cannot hold a model. So the answer is not to pick one, but to split the app across two of them. The request path runs close to users, and inference runs on a dedicated GPU it calls into. That also explains the failed upgrade. More compute made a stage that was already fast faster, and left the 10.5 seconds around it untouched. To actually learn how it's done in practice, Akamai's GitHub has a reference implementation for each half. - vllm-on-lke serves Qwen2.5-7B-Instruct behind an OpenAI-compatible endpoint on one RTX 4000 Ada GPU in Linode Kubernetes Engine, with Terraform creating the cluster, both firewalls, and the GPU operator in one apply. - akamai-functions-llm-chatbot covers the front, where a WebAssembly API checks a KV cache and only calls the GPU-backed instance on a miss. Both are available on Akamai’s new Developer Hub, alongside their tutorials and code samples. It also links to Edge Case, their Discord, where four developer advocates architect and deploy a production app live every other Wednesday. If you create a new Akamai Cloud account, you can also get $300 in credits for joining. Join here: That said, this post treats generation as a single 1.5s block, but that block has its own structure, and knowing it well tells you whether a model is slow to start or slow to stream. I wrote a first-principles walkthrough of it, covering the prefill and decode split, KV caching, and where the time actually goes inside each one. Read it below. Thanks to Akamai Cloud for partnering today!show more

Avi Chawla
21,786 次观看 • 1 个月前
$IREN "we haven't disclosed the specific amount of GPUs"... 1. 🤮 reminds me of $NBIS 2. Setting a terrible precedent here for future deals 3. Making it purposely difficult, to not let analysts properly value your 2027 revenue 4. Increasing the polarized view on IREN by the market However: "approximately 60MW of air-cooled Blackwells" 1. You typically don't talk about gross capacity in a deployment like this 2. If it would be gross capacity, the GPU hour rate at IT level would be crazy high (at PUE 1.2, $680m / 50 = 13.6m/MW) 3. At 60MW IT load, and ~14kW draw at DGX server level, we can get to ~4,286 DGX systems with 8 GPUs per. 4. Based on this we can conclude that 60MW of IT load can run approximately 34k DGX B300. 5. 34k DGX B300 at $680m/yr, would represent a GPU hour price of $2.28 Now this is the problem with not disclosing your GPU quantity. You purposely make your business model look bad, because by approach, you get to a GPU hour price that would imply a payback period of 4 years, where only the last year of the contract is 100% margin. But of course, we can also take "the glass is half full" approach. IREN has ordered 50K B300s from Dell. They have 2 purchase orders for this, 1 between Dell Canada and IE CA Leasing Ltd for 4 phases, and 1 between Dell USA and IE US Hardware 1 Inc (amended from IE US Hardware 4 Inc on April 27, 2026). The order for Canada is divided in 4 phases, and are going to Mackenzie for 80MW of gross capacity, which happens to be 4 buildings of 20MW. The order for Childress is divided in 2 phases, and are going to DC35 and DC36, (as depicted in the earnings presentation) and those are 50MW gross. The purchase price of the order for Childress was $1.2B, and for Canada it was $2.3B If we go with 50,000 B300s for a total of $3.5B then $1.2 would represent 34.285% of the 50,000 GPUs, or 17,140 B300s rounded down. For this calculation I will consider that $IREN will deploy 17,140 GPUs in 50MW gross capacity in DC35 and DC36 of block 3 in Childress.. That would imply at 1.2 PUE, IREN can run 17,140 B300s in 41.67MW IT load. Now by that ratio, they can run 24,680 GPUs in 60MW IT load — a massive difference with 34k units through the Nvidia DGX reference calculation. If common sense is applied, you can still get to 2 completely different outcomes, that show a difference of more than 9k GPUs. The GPU hour rate at 24.68k GPUs would be $3.145 per B300, as MASSIVE difference from the earlier calculated $2.28. Sure, the DGX system may be a factor here. And I'm sure that the reality is somewhere in the middle. But I personally hate this as an investor, to be unable to calculate profitability on unit economic basis. After all, contracts are signed on a $/GPU hour basis. Why hide this from your investors? Not being able to calculate payback periods, unable to calculate ROIC. And most importantly, we cannot properly assess the $NVDA deal on a contract basis. I really hope the payback period of this contract is not 4 years. I want the glass to be half full, but by starting to censor the purchases, IREN is taking a step in the wrong direction. Not a fan of this.show more

Frans Bakker
148,167 次观看 • 4 个月前
India should host the biggest vibe coding conference the... world has seen 🚀 Everything about Vibe coding(talks, panels, AI product building, hackathon, including the biggest names in the world) come here for an ultimate showdown(think 10k+ people in a Vibecoding conference) 🔥 And I have a plan Why? Because the number of learners, founders and professionals who are building via vibe coding in India and powering the global platforms is testament to the fact that if it should happen anywhere, it is here For the last 4+ years, I have enabled 3000+ people to build and continue to do so in a new avatar(announcement soon) and I believe if the entire ecosystem is willing to come together we can create a real spectacle We have the skills, the access and the experience to pull this off. The only thing it needs is for everyone seeing this to reach out and join hands(partners, sponsors & more) to make this as big as possible I am so hyped about it that the website is set, the name is set(Vibecon) and we can make an incredible run to make it happen. Reply to this post if you think we should do this, if we have >500 responses we will get to work 🥳 Reach out if you have ideas to partner to make the biggest Vibecoding conference a reality ♥️show more

Prashant Sharma
11,030 次观看 • 1 年前
TWIN FLAME LOVE IS NOT A SPARK THAT FADES... WITH ME. It is not sustained by excitement, romance, or constant closeness. It is sustained by truth. This is why it does not burn out. Twin flames share a frequency, not just emotions. Even when they are apart, the connection continues to exist because it is not dependent on words, actions, or physical presence. It lives in the nervous system, in memory, in the way awareness shifts after the meeting. Once activated, it does not return to what it was before. This love does not consume itself the way ordinary passion can. It matures. It moves through phases of intensity, silence, confusion, distance, and clarity, yet the core recognition remains unchanged. What changes is the capacity of each person to hold it without fear. When separation happens, the love does not disappear. It reorganizes. It turns inward and begins to work on unresolved wounds, attachment patterns, and old survival responses. The connection stays alive because it is no longer ted by chasing or longing, but by integration. Twin flame love endures because it is not trying to prove itself. It does not require constant reassurance. It is quiet when needed, intense when allowed, and steady beneath all cycles. Even in moments of doubt, something deeper continues to recognize the other as familiar, safe, and true. This is why twin flame love does not burn out. It is not fueled by emotion alone. It is carried by awareness. And awareness, once awakened, does not extinguish. ~ Twinflame Infinity ✨🙌🏽💫show more

🧬Maxpein🧬
18,135 次观看 • 9 个月前
Twin flame love is not a spark that fades... with time. It is not sustained by excitement, romance, or constant closeness. It is sustained by truth. This is why it does not burn out. Twin flames share a frequency, not just emotions. Even when they are apart, the connection continues to exist because it is not dependent on words, actions, or physical presence. It lives in the nervous system, in memory, in the way awareness shifts after the meeting. Once activated, it does not return to what it was before. This love does not consume itself the way ordinary passion can. It matures. It moves through phases of intensity, silence, confusion, distance, and clarity, yet the core recognition remains unchanged. What changes is the capacity of each person to hold it without fear. When separation happens, the love does not disappear. It reorganizes. It turns inward and begins to work on unresolved wounds, attachment patterns, and old survival responses. The connection stays alive because it is no longer fed by chasing or longing, but by integration. Twin flame love endures because it is not trying to prove itself. It does not require constant reassurance. It is quiet when needed, intense when allowed, and steady beneath all cycles. Even in moments of doubt, something deeper continues to recognize the other as familiar, safe, and true. This is why twin flame love does not burn out. It is not fueled by emotion alone. It is carried by awareness. And awareness, once awakened, does not extinguish. ~ Twinflames.Infinity ✨🙌🏿💫show more

Cosmic Insights 🌻
14,295 次观看 • 7 个月前
Run Gemma 4 26B MoE on 8GB VRAM with... 250k context at 20+ tokens/sec If you own any 8GB VRAM graphics card, stop what you are doing. Local AI just had its absolute "Holy Shit" moment for budget hardware. Yesterday, I benchmarked Unsloth Gemma 4 12B Q4_K_XL on an 8GB card. The community went wild but immediately demanded more: "Can we run a 25B+ model on budget GPUs?" Today, I’m delivering exactly that. I am running a massive 26B parameter Mixture of Experts (MoE) model locally on a standard 8GB VRAM setup with 250k full native context!. If you own an RTX 3060, 3070, 4060, or any budget GPU with 8GB of VRAM, the local AI paradigm has completely changed. The performance metrics are astonishing: - 20 tokens/sec flat decode throughput. - Stable, flat decode speed even with massive prompts. - I threw a 60k token prompt at it, and it still clocked in at 20 TPS without dropping a single frame. # What about prefill? Yes, Time To First Token (TTFT) is slightly high when swallowing massive contexts. But with a solid 200 tokens/sec prefill speed, the wait is barely noticeable and highly usable. And this is running completely without Multi Token Prediction (MTP) active. How is this possible? It’s the magic of Google's new QAT (Quantization Aware Training) quants for Gemma 4. The model weight file (unsloth gemma-4-26B-A4B-it-qat-UD-Q4_K_XL.gguf) is only 13.2 GB, making it the ultimate local powerhouse. # The Test Setup: CPU: Intel Core i7 RAM: 16GB System RAM GPU: NVIDIA GeForce RTX 4060 Laptop GPU (8GB VRAM) # The Secret Sauce (The -cmoe Flag) To make this work properly on any 8GB card, you must use the -cmoe (CPU MoE) flag in llama.cpp. This flag isolates the heavy MoE expert weights directly to system memory (CPU/RAM) while letting your GPU focus strictly on the Attention layers and the KV Cache. It prevents VRAM spillage and holds the throughput rock solid. # The flags: -m "gemma-4-26B-A4B-it-qat-UD-Q4_K_XL.gguf" -cmoe -c 248000 -v Once running, just open the UI on localhost and toggle the new reasoning lightbulb icon in the text input box to watch the model perform multi step thinking. Are you still running smaller models, or are you ready to scale up your budget local setups? Let's discuss in the repliesshow more

Alok
295,215 次观看 • 4 个月前
Dad’s home! 🏠 Because it’s one of the most... frequent questions I receive, here’s your periodic reminder that your baby doesn’t forget you when you go to work. There’s a developmental milestone known as object permanence and sometimes parents overthink it. In a nutshell it’s the understanding that things continue to exist even when they’re out of sight - and it typically develops somewhere between 6-9 months. Does that mean you cease to exist (to your baby) when you go to work or the grocery store? It can be kind of a scary thought to new parents. The answer, even before the development of object permanence, is no. But it may mean that (even for your toddler) you’re misattributing adult forms of memory to your little one. Infants and toddlers both live very much in the moment and are still developing the type of robust narrative memory we possess as adults. So while you’re gone, it’s not that you cease to exist. It’s more a case of “out of sight, out of mind.” Will they recognize you when you return? Will they be excited to see you? I’ll let this beautiful video from beingritty on IG show you the answer.show more

Dan Wuori
30,443 次观看 • 1 年前
Model-Free Reinforcement Learning (MFRL) has been alluring, especially with... supercharged compute with physics on GPU. However, the methods use 0-th order gradients, and are often not the best optimizers. Can we do better than PPO in continuous control for robotics? Turns out yes! 🥳 tl;dr: Faster, better RL than PPO in continuous control 💪 The answer lies in using more information from the simulation. We are juicing the simulation on GPU as it is, why not use it for gradients as well? This has been a driving question in a series of our works. We first studied this problem in ICLR 2022 paper on Short Horizon Actor Critic Naive gradient based methods are stuck in local minima and have exploding/vanishing gradients. SHAC solved this problem truncated rollouts and model based value estimation, where the model is Differentiable Sim. This boosted sample efficiency and wall-clock time immensely especially in high dimensional systems such as humanoids Yet, given enough compute PPO often caught up. Our follow up paper on on Adaptive Horizon Actor Critic at ICML 2024 discovers the cause and provides a fix. However, we find that even when given ground-truth dynamics, not all gradients are useful due to sample error. 1st-Order Model-Based Reinforcement Learning methods employing differentiable simulation provide gradients with reduced variance but are susceptible to bias in scenarios involving stiff dynamics, such as physical contact. We find that back-propagating through contact and long trajectories drastically reduces gradient accuracy. Using this insight, we propose AHAC to dynamically adapt its roll-out horizon to avoid differentiating through stiff contact. AHAC is a first-order model-based RL algorithm that learns high-dimensional tasks in minutes (wall clock) and outperforms PPO by 40%, even in the limit of data provided to PPO. This work is led by Ignat Georgiev alongside Krishnan Srinivasan, Jie Xu, Eric Heiden and ample assistance from warp team at NVIDIA Robotics (Miles Macklin)show more

Animesh Garg
52,308 次观看 • 2 年前
What if you kept asking an LLM to "make... it better"? In some recent work at FAIR, we investigate how we can efficiently use RL to fine-tune LLMs to iteratively self-improve on their previous solutions at inference-time. Training for iterated self-improvement can be costly. The naive approach to training for K self-improvement steps leads to K times the number of rollout steps per episode. We introduce Exploratory Iteration (ExIt), an RL-based automatic curriculum method that bootstraps diverse training distributions of self-improvement tasks by upcycling the LLM's own responses at previous turns as the starting points for both self-improvement and *self-divergence.* In order to decide what task to train on next, the curriculum prioritizes sampling of partial turn histories that led to higher return variance in its GRPO group (a learnability score that comes for free). This automatic curriculum over the bootstrapped task space teaches the model how to perform iterated self-improvement while only ever training the model on single-step self-improvement tasks. We look at ExIt's impact in both single-turn (contest math problems) and multi-turn (BFCLv3 multi-turn tasks), as well as MLE-bench, where the LLM is run in a search scaffold to produce solutions to real Kaggle competitions. Across these eval settings, we find ExIt produces models with greater capacity for inference-time self-improvement compared to GRPO. Notably, ExIt models can self-improve on test tasks for many more steps than the typical solution depth encountered during training, including a 22% improvement in MLE-bench performance compared to GRPO.show more

Minqi Jiang
41,154 次观看 • 1 年前
I had the same thought so I've been playing... with it in nanochat. E.g. here's 8 agents (4 claude, 4 codex), with 1 GPU each running nanochat experiments (trying to delete logit softcap without regression). The TLDR is that it doesn't work and it's a mess... but it's still very pretty to look at :) I tried a few setups: 8 independent solo researchers, 1 chief scientist giving work to 8 junior researchers, etc. Each research program is a git branch, each scientist forks it into a feature branch, git worktrees for isolation, simple files for comms, skip Docker/VMs for simplicity atm (I find that instructions are enough to prevent interference). Research org runs in tmux window grids of interactive sessions (like Teams) so that it's pretty to look at, see their individual work, and "take over" if needed, i.e. no -p. But ok the reason it doesn't work so far is that the agents' ideas are just pretty bad out of the box, even at highest intelligence. They don't think carefully though experiment design, they run a bit non-sensical variations, they don't create strong baselines and ablate things properly, they don't carefully control for runtime or flops. (just as an example, an agent yesterday "discovered" that increasing the hidden size of the network improves the validation loss, which is a totally spurious result given that a bigger network will have a lower validation loss in the infinite data regime, but then it also trains for a lot longer, it's not clear why I had to come in to point that out). They are very good at implementing any given well-scoped and described idea but they don't creatively generate them. But the goal is that you are now programming an organization (e.g. a "research org") and its individual agents, so the "source code" is the collection of prompts, skills, tools, etc. and processes that make it up. E.g. a daily standup in the morning is now part of the "org code". And optimizing nanochat pretraining is just one of the many tasks (almost like an eval). Then - given an arbitrary task, how quickly does your research org generate progress on it?show more

Andrej Karpathy
1,657,913 次观看 • 7 个月前