Video wird geladen...

Video konnte nicht geladen werden

Zur Startseite

A RUSSIAN MATHEMATICIAN BUILT A SYSTEM WHERE MEMORY AND EVAL WORK AS ONE PIPELINE NOT TWO SEPARATE TOOLS Most setups run memory and evaluation as separate systems that never actually talk to each other. He wired them into one loop instead, memory stores every past output, eval scores each...

39,872 Aufrufe • vor 1 Monat •via X (Twitter)

0 Kommentare

Keine Kommentare verfügbar

Kommentare vom Original-Post werden hier angezeigt

Ähnliche Videos

a contractor in Shenzhen priced a ¥12,470,900 hospital contract, about $1.7m, in one afternoon and beat firms carrying forty people he explained how he did it: the bid consultancy he used to pay took three days and ¥46,000 for the same envelope. he did this one alone, off one screen, at 11.4% margin, uploaded before the 17:00 cutoff 214 pages of tender documents read, 68 binding clauses pulled out, 9,485 building parts loaded, 14 places found where a duct and a beam sit in the same cubic metre, deepest one 38mm, all of them fixed, 3,318 lines of quantities priced and the package encrypted and uploaded before the 17:00 cutoff this is Graph Engineering: the job gets cut into small nodes, one narrow task each, wired so that one node's output is the next node's input, and any node is allowed to stop the whole run. it turns a model that answers you into a machine that finishes the job: - give every node one job and one output. a node doing two things fails at both and you cannot tell which one broke - put the cheapest rejection first. his qualification node reads clause 7.4, foreign-owned firms barred, and ends the run four seconds in, before anything expensive touches the model - what moves between nodes is a file. the model travels as a model, the quantities as a table, the price as a number - build exactly one loop: the checker finds 14 collisions, the fixer drops the duct 550mm, the checker runs again, and nothing moves on until the count is zero - cap that loop, or a graph will grind on three impossible clashes until the deadline passes - keep one node whose only job is to say no, and give it authority over everything above it - log each node's output on its own, because when the price comes out wrong you need to know which node believed the wrong thing - run the expensive nodes last, always the catch is that a graph is an extremely confident machine: point it at an outdated rate book and it prices an entire hospital off it without a single node noticing, because no node is asked to doubt the input, only to process it so the nodes that earn their keep are the ones that reject, and almost nobody builds those first bookmark this, the full build with all nine nodes and what each one hands to the next is written out in the article ↓

Argona

38,189 Aufrufe • vor 2 Monaten

Researchers made KMeans 200x faster. And the new technique also beats approaches like cuML and FAISS. Flash-KMeans is an IO-aware implementation of exact KMeans that redesigns the algorithm around modern GPU bottlenecks. By attacking the memory bottlenecks directly, Flash-KMeans achieves: - 33x speedup over cuML - 200x speedup over FAISS This speedup comes from how it moves through GPU memory. Standard KMeans runs in two steps, and both are bottlenecked by reads and writes to GPU memory: 1) The first step matches every point to its nearest centroid. Standard KMeans computes the full point-to-centroid distance matrix, writes it out to GPU memory, then reads it back to find each nearest centroid. That write-then-read round trip is the bottleneck. Flash-KMeans combines the distance calculation with the nearest-centroid step, so the result is computed on-chip and the full matrix is never written out. 2) The second step recomputes each centroid by averaging the points assigned to it. Standard KMeans has thousands of threads writing into the same centroid slots at once, so they stall waiting for their turn. Flash-KMeans sorts points by cluster first, turning scattered writes into sequential reductions that read and write memory in one efficient pass. Using these two optimizations at the million-scale, Flash-KMeans completes a standard KMeans iteration in a few milliseconds. The video below depicts this in action. Several reasons why this is important: KMeans has always been an offline primitive. Something you run once to preprocess data and move on. These speedups make the approach viable in several runtime-critical systems. ↳ Vector indices like FAISS use KMeans to build search indices. Faster KMeans means you can re-index dynamically as data changes. ↳ LLM quantization methods need KMeans to find optimal weight codebooks, per layer, repeatedly. What takes hours could now take minutes. ↳ MoE models need fast token routing at inference time. Flash-KMeans makes it viable to run this inside the inference loop, not just in preprocessing. I have shared the paper in the replies. That said, memory is the real constraint Flash-KMeans solves, and the problem is not just limited to clustering. The vectors a RAG system stores after indexing create similar bottlenecks. I wrote a detailed walkthrough recently on cutting this vector memory by 32x with binary quantization, querying 36M+ vectors in a few milliseconds. Read it below.

Avi Chawla

90,063 Aufrufe • vor 3 Monaten

Day 12/90 of Inference Engineering What is chunked prefill within vLLM? In continuation of yesterday's post on the high level architecture of vLLM, I want to dive deeper into vLLM core engine starting with the mechanics of chunked prefill. In this post, I will closely follow the original blog on the anatomy of vLLM. To start, let's define chunked prefill. It's a runtime inference optimization technique that splits a long input request so that it doesn’t monopolize the whole GPU. Keep in mind this is all within the context of vLLM. And since vLLM is an inference engine that's meant to serve a model to multiple concurrent users, having a GPU that’s fully monopolized on a single user's request means other users' requests would be in queue waiting to be processed. It isn’t too good to have the whole GPU occupied on a single request when the GPU is meant to be shared! So the key idea behind chunked prefill is to break the long request into smaller chunks, so that each chunk along with other users' requests gets processed and written into the KV cache together. Suppose we split up the long request into chunks and each chunk has 8 tokens. Now each memory block can hold 4 tokens. Therefore, 8 tokens can fit into 2 blocks of memory. After the first forward pass, 2 blocks are occupied, and after the second forward pass, 4 blocks of memory are occupied and so forth. Each forward pass handles a small chunk of the long request so that there's room in the same pass to keep serving other users' requests. Here's a small animation that I made today to fully visualize the idea behind chunked prefill when learning this topic~

max fu

29,585 Aufrufe • vor 2 Monaten

I just launched a brand new channel with an INSANE strategy. I'm calling it the Launch Loop. I've been working on this channel for over a year, and took the actually massive risk to hold off on posting until I had all 5 launch videos ready. Each idea was carefully selected for this strategy, and each one was carefully written to execute it. I view this channel launch, and these videos, not as individual videos but as a season of a new show. The plan was this: come up with video ideas that fit together, and write the end of each video to smoothly blend into the next. I treated the conclusion of each video as if it was the intro of the next video, so that if people enjoyed what I made and wanted to see more, they could click the next video on the End Card This was inspired by work I've done in the past: a few years ago on my YouTube strategy channel, I posted a video about Ryan Trahan's thumbnail strategy which has gotten over 2M lifetime views. But at the end of that video, I carefully wrote a transition showing video my comprehensive breakdown of 21 thumbnail formats. The End Screen video CTR on that video is 29.4% 🤯. I'm convinced that some of those 2M views must have come from the success of that end screen. I've come to believe that it's really important to bring in an audience to your CHANNEL, not just one video. That was a mistake I made in another channel I worked on before! Getting virality on individual videos, but not keeping viewers for the long term or growing an actualy community. So I thought, what if I launched a channel where every video pointed to the next, in a loop? Maybe there will be a flywheel where success on one brings success on the next, and the next, and the next... But for this plan to work the videos I made have to actually be good 😅

Tyler Fleischman ⏸️

73,039 Aufrufe • vor 13 Tagen

A 24-year-old built two AI girls with Claude and now clears $21,800 a month from them. The build took 15 days. He trained separate LoRAs for both girls, locked their identity seeds, and kept small imperfections on purpose: a loose strand of hair, tiny skin marks, slightly uneven framing. Perfect symmetry gets flagged. Small inconsistencies make them look real. He posts 5 times a day across TikTok, Instagram and X. Morning routines, gym sessions, mirror videos, outfit changes, pool clips, and videos of the two girls together. The content is designed so they look like two real friends who actually live in the same world. The smartest part is that the accounts interact with each other. One girl comments on the other’s posts, appears in her videos, and references things they supposedly did together. Followers stop seeing them as two AI models and start following the relationship between the characters. Within three months they crossed 312,000 followers combined and started receiving hundreds of DMs every night. The private channel sits at $25 a month, while an AI memory agent keeps track of every conversation, previous message, favorite post, and personal detail each follower has shared. Replies come back in under 30 seconds. The agent checks the user's previous conversations before answering, so the response feels consistent with the personality of the girl they are talking to instead of sounding like another generic AI chatbot. By the end of month three, the two accounts were generating $13,900 from subscriptions and private chats, another $5,700 from brand deals, and around $2,200 from digital products. The brands came after the audience started growing: clothing companies, beauty products, fitness brands, and lifestyle products wanted access to the same audience that was already following the two characters every day. The Claude stack that locked them: 1Full identity, personality, lighting, camera style and body proportions locked into separate character systems. 2Separate LoRAs trained only on each girl's approved character frames. 3Apartment, bedroom, gym and outdoor locations generated once and reused to keep the world consistent. 4Every video built around natural movement, imperfect framing and small variations instead of polished AI-perfect shots. 5Memory agent connected to the conversations so both girls remember what followers previously said. 6Upscaling, face consistency and final post-processing before everything goes live. The first girl brings people into the account. The second gives them another character to follow, another story to watch, and another reason to come back. The content gets them interested. The relationship between the two characters keeps them watching. The memory agent turns that attention into recurring revenue.

genuenci

678,031 Aufrufe • vor 1 Monat

whoever leaked this has bigger balls than sense someone gave a fleet of Claude agents shared memory so they would stop contradicting each other, then measured both the bill and the output: the version that talked most made 2.4x the api calls of the version that won, and hallucinated 34% more than doing nothing at all, 0.658 against 0.492 i ran the same question past two of my own agents afterwards and got two different answers about which file owns the config. each one was individually right and the pair was wrong, which is the whole failure in one line this is Graph Engineering, the layer that decides which agents may talk to each other at all, and it installs into the agent you already pay for: - decide which agents may share state at all, because every edge you draw is a channel a mistake can travel down - measure divergence per PAIR instead of as a fleet average, across what they believe about place, time and task history - gate on that number and stop the pair above your threshold before it reasons, rather than repairing the output afterwards - let compressed summaries replace whole states: the verified protocol landed 0.463 against 0.658 for full broadcast - cut the sync frequency until it hurts, since the winning setup used 58% fewer calls than the one that broke it - never propagate a state nobody checked, because the contamination effect came in at d=1.18, a full standard deviation of extra lying - keep the shared layer small enough to diff, which is what a written standard does and a running conversation cannot - re-run the check after every model upgrade, because this was 8 scenarios on one model family at n=30 per condition - and learn where it does not bite: on plain software tasks every condition converged under 0.2 and the whole effect vanished turns out the ranking is the uncomfortable part: verified summaries 0.463, no synchronisation at all 0.492, full broadcast 0.658. the middle option is doing nothing, and it beat the thing everyone builds first the group agreeing is what it looks like when every agent copied the same mistake, which is why a fleet that hallucinates has a replication problem and keeps getting handed a smarter model instead so the question for your own setup: if you asked two of your agents the same thing right now, would they answer the same way bookmark this one. the layer underneath it, deciding which arrows between agents exist at all, is built step by step in the piece below ↓

Argona

724,665 Aufrufe • vor 2 Monaten

A wrist force sensor fires at 100Hz. The policy only ever sees it at 30Hz, downsampled to land on the same control step as the camera and the joint state. That's not a bug, it's the whole point, and it sits inside a bigger pattern in VLA research this year. Every major release has been Markovian at its core, mapping the current frame straight to the next action. The fix everyone reaches for is more vision: more history frames, longer image context. FM-VLA makes a clean case that the fix is the wrong channel for a whole class of tasks. Press a button three times and stop. A camera watching that has almost nothing to work with, the scene barely changes between press one and press three. Force doesn't have that ambiguity problem. Each press is a sharp, distinct spike in the wrench signal, whether or not the camera noticed anything at all. So FM-VLA doesn't add more frames. It compresses the wrench history into eight tokens with a VAE, pretrained purely on reconstructing force signals, frozen before it ever touches the policy, then hands those tokens to the action expert alongside a short window of joint state. That's the entire memory system. Averaged across three contact-rich tasks, FM-VLA hits 83.3 percent success against 33.3 percent for the strongest vision-memory baseline on the button-counting task specifically, where the ambiguity problem is worst, 72.2 percent for FM-VLA there. Strip out the short-state window and force-only performance drops well below the combined system, so force alone isn't the answer either. The two channels are doing different jobs. The field has defaulted to one memory channel for every kind of temporal problem. This is a clean data point that the channel should match the ambiguity you're actually trying to resolve, not just get bigger. Source: Paper: Credit to the teams at Tsinghua University, Microsoft Research, and Fudan University. #Robotics #PhysicalAI #RobotLearning

Stephen James

11,658 Aufrufe • vor 2 Monaten