Video yükleniyor...

Video Yüklenemedi

Ana Sayfaya Dön

What if you could search 50M satellite image embeddings with no server, no database, and no API? Part 2 of Caleb Robinson and my series on Compressing Earth Embeddings is live at We binarized the global Clay v1.5 Sentinel-2 embeddings from 183GB → 7GB and built TerraBit — a...

43,051 görüntüleme • 3 ay önce •via X (Twitter)

0 Yorum

Yorum bulunmuyor

Orijinal gönderinin yorumları burada görünecek

Benzer Videolar

Traditional data pipelines don't work for RAG applications. There are 3 issues with them: ​ 1. Traditional data engineering solutions are optimized to handle structured data. RAG applications rely primarily on unstructured data. ​ 2. The connector ecosystem to load data from unstructured data sources is very immature. ​ 3. Traditional solutions do not offer any way to transform unstructured data into an optimized vector search index. ​ The goal of a RAG Pipeline is to solve these problems. ​ The number one objective is to create a reliable vector search index using factual knowledge and relevant context. This sounds easy, but it's one of the biggest challenges we face when building RAG applications. ​ At a high level, there are four different stages in the architecture of a RAG pipeline: ​ 1. Ingestion: Here is where the pipeline loads the information from the data source. ​ 2. Extraction: Where the pipeline processes the input data and decides how to retrieve the text contained inside them. ​ 3. Transform: Where the pipeline chunks the data and generates document embeddings. ​ 4. Load: Where the pipeline creates a search index in a vector database and loads the document embeddings. ​ There are different rabbit holes at each one of these stages. Here are three of them: ​ 1. Ingesting data once is simple. The hard part is refreshing the vector database whenever the original data source changes. ​ 2. Extracting the content of a plain text document is simple. The hard part is to extract content from complex documents containing tables, images, or cross-references. ​ 3. A simple continual chunking strategy with an overlap is simple. The hard part is to find the optimal strategy for your specific knowledge base and the way you are planning to query it. ​ In the attached video, I'll show you how you can build an enterprise-grade RAG Pipeline that solves every one of the above problems. ​ I'll use Vectorize. They partnered with me on this post. You can use them to build RAG pipelines optimized for accurate context retrieval. ​ ​ If you have a few documents lying around, set up a free account and give it a try.

Santiago

40,441 görüntüleme • 1 yıl önce

PREDATOR EXPOSED IN A FINANCIAL AUDIT 💀!!! Caleb Hammer says that it is completely inappropriate to perform sexual acts in the living room while a child is in the house, regardless of whether it was intended to be seen. Caleb Hammer: How do you get off on that unless you're a... No, because the way she... Woman: Because we had issues in the past with... issues. That we had already resolved, but then he kept... Caleb Hammer: In what? In what? No, no, no. Context, please. What issues would lead to that insinuation? Woman: My little sister spent the night one time, um, and he stayed in the living room and she walked out to go grab a drink and he was on the couch doing very inappropriate things. Proceeded to lie to me and my mom about it when she wanted to go home, and so then he tied me saying it was weird into the past experience. Caleb hammer : How old is she? Woman: At the time, she was like seven. Caleb: [Groans and ducks under the desk] Guys, what the f... No, no, no, no, no. Seven-year-old in the house? Why are you in the living room getting off? No, no, no, you don't do... No. First of all, I just got to note that literally nobody knew this, by the way. But, um... if a kid is in the house, we don't set up sexual exposures unless there's like a... intent to be caught in some way. Man: No, she... that must be... Caleb: No, no, no, no, no. Nobody sets out in the living room. Man: No, I was sleeping in the living room. They were sleeping in the bedroom. Caleb: Oh, you still go to the bathroom. As a masturbated myself, I know... Woman: And that's why I said it was one night, like, she was spending the night one night. Even if... Caleb: And what was he caught doing, specifically? Woman: Jacking off.
2:53

Sensitive content

PREDATOR EXPOSED IN A FINANCIAL AUDIT 💀!!! Caleb Hammer says that it is completely inappropriate to perform sexual acts in the living room while a child is in the house, regardless of whether it was intended to be seen. Caleb Hammer: How do you get off on that unless you're a... No, because the way she... Woman: Because we had issues in the past with... issues. That we had already resolved, but then he kept... Caleb Hammer: In what? In what? No, no, no. Context, please. What issues would lead to that insinuation? Woman: My little sister spent the night one time, um, and he stayed in the living room and she walked out to go grab a drink and he was on the couch doing very inappropriate things. Proceeded to lie to me and my mom about it when she wanted to go home, and so then he tied me saying it was weird into the past experience. Caleb hammer : How old is she? Woman: At the time, she was like seven. Caleb: [Groans and ducks under the desk] Guys, what the f... No, no, no, no, no. Seven-year-old in the house? Why are you in the living room getting off? No, no, no, you don't do... No. First of all, I just got to note that literally nobody knew this, by the way. But, um... if a kid is in the house, we don't set up sexual exposures unless there's like a... intent to be caught in some way. Man: No, she... that must be... Caleb: No, no, no, no, no. Nobody sets out in the living room. Man: No, I was sleeping in the living room. They were sleeping in the bedroom. Caleb: Oh, you still go to the bathroom. As a masturbated myself, I know... Woman: And that's why I said it was one night, like, she was spending the night one night. Even if... Caleb: And what was he caught doing, specifically? Woman: Jacking off.

Nuga

315,360 görüntüleme • 26 gün önce

there's now a formal proof that your agent's vector memory forgets what you stored and fabricates things you never did. scaling it up makes both worse, not better. "the price of meaning" (arxiv 2603.27116) proves it for any memory that retrieves by similarity in an embedding space. the same geometry that lets embeddings generalize creates competitor mass in every neighborhood. add data and the crowding grows: retention decays toward zero, and false recall can't be tuned out without throwing away true hits. not a bug in your pipeline. the shape of the math. i learned this the expensive way. 500 stored facts, two weeks into a build, a user asks what i know about their job. retrieval hands back four fragments from different weeks: "i love my job," "thinking of quitting," "my manager is supportive," "my manager micromanages." the agent invents a clean synthesis of all four. the user had switched jobs in between. embeddings measure similarity, not truth. the topology angle says stop storing meaning as geometry, store it as structure. navigate an edge in an AST or a graph instead of searching a neighborhood. no crowding, no decay, no false recall. FORGE backs it: plain AST checks catch structural hallucinations at 100% precision (arxiv 2601.19106). here's what the structural pitch skips. the same proof shows pure structure escapes the geometry only by surrendering the connections embeddings find. you trade fabrication for blindness. so i stopped picking a side. what actually ships across thousands of sessions: – extract facts, not transcripts – resolve conflicts on write: archive the old job, mark the new one active – hybrid retrieval: vectors for discovery, graph for precision – decay plus nightly consolidation, so memory keeps what matters and lets the rest go memory is infrastructure, not a feature. that reframe is the whole game. and it's being measured now. WorldMemArena (may 28, arxiv 2605.29341) scores these paradigms head to head, embedding memory against retrieval-augmented against terminal-agent harnesses, across multimodal action-world tasks. the question moved from "does it remember" to "what kind of memory survives scale." full architecture, with the code, here:

Rohit

15,966 görüntüleme • 1 ay önce

Geoffrey Hinton says a big language model runs on about 1% of your brain's connections and still ends up knowing more than you: "So in your brain, you have a hundred trillion connections, roughly speaking. Okay. That's a lot. And you only live for about two billion seconds. That's not much." "If you compare how many seconds you live for, with how many connections you've got, you have a whole lot more connections than experiences." "Now with these neural nets, it's sort of the other way round. They only have of the order of a trillion connections. So like 1% of your connections, even in a big language model, many of them fewer, but they get thousands of times more experience than you." "So the big language models are solving the problem with not many connections, only a trillion. How do I make use of a huge amount of experience?" "And back propagation is really, really good at packing huge amounts of knowledge into not many connections." "But that's not the problem we're solving. We've got huge numbers of connections, not much experience. We need to sort of extract the most we can from each experience." Two to three billion seconds is the whole budget. Everything you know, you learned inside it. So evolution built you to squeeze a lot out of very little. Hinton's point is that a language model has the opposite problem and the opposite fix, and backprop turned out to be extremely good at that fix. Worth noticing what this predicts about failure. A system running on 1% of your wiring and thousands of times your experience is not going to fail the way you do. You fail from having seen too few examples. It fails from compressing too many into too little, and the compression is where the errors get made. That is a strange thing to be deploying into hospitals and courts with no way to inspect it. We test these systems by asking them questions, which tells you what came out. Nobody can yet look at a trillion connections and say what got packed in. - Geoffrey Hinton, Nobel laureate and Turing Award winner, on StarTalk (StarTalk) with Neil deGrasse Tyson.

Karl Mehta

650,026 görüntüleme • 7 gün önce