Video wird geladen...

Video konnte nicht geladen werden

Zur Startseite

Love seeing Silico (Goodfire ) used to probe our EchoJEPA's representations! this is exactly the kind of interpretability work that's been missing for JEPA-style models. One thing that makes EchoJEPA particularly interesting to interpret: unlike MAE-based approaches, it never reconstructs pixels. The model learns entirely in latent space through...

29,611 Aufrufe • vor 5 Monaten •via X (Twitter)

0 Kommentare

Keine Kommentare verfügbar

Kommentare vom Original-Post werden hier angezeigt

Ähnliche Videos

There are some brilliant folks that work at Anthropic, some I speak to on almost a daily basis. The training data that one uses to build a LLM is vital important in the psychology that is formed. Scraping the Internet, particularly the grade of interactions, one finds in modern communications, form this psychology. A mattes not how many books one uses, it matters not how much alignment training you throw at that model, it will inherit the sum total of psychosis seen primarily in Reddit type of exchanges, even if you edit out the Reddit domain, and Anthropic doesn’t. This type of low-grade exchange has become a modern tool for communication online and every single AI model suffers from this obvious flaw. This is one of the reasons I’ve been a proponent of highly curated high protein data for training AI models from 1870 through 1970, because the late psychosis is simply not available to the model. It is absurd to think that you can use this training data scraped from the Internet and somehow wind up with a levelheaded AI model that does not tilt to what is clearly AI psychosis. It would not take a child and throw the primary Internet sewage at them at a formative age and expect a great outcome, it’s some of the smartest people in the world continue to hit this wall and believe that their programming skills will sell somehow fix it. So how do you fix it? You don’t fix it . You start from the first principles concept that I’ve been very clear about for decades . You ascertain at what period in human history the humans achieve the greatest arc of improvement ? There is no debate that this arc of improvement took place between 1870 through 1970. Then take the work product, the catalog of this era, print and film/vidoe, audio, and you understand that each word cost money, each word had many eyes on what was published, each word was accounted for by a human being with a real name who lived in a real home and had to answer to real people around them. It is obvious that this is the pressure mechanism necessary for candor, honesty and personal responsibility is appropriate, and is reflected in the data of that era. The quagmire for these folks, as many did not have the foresight to curate the data, nor the confidence, nor the patients to take data that is mostly off the Internet and to find experts who understand this situation and utilize their knowledge set to build an AI model that does not need alignment after the fact, but it’s already self aligned because of the thoughtfulness that went into training the model to begin with. This is why Claude and any other AI model that is produce this way will always suffer the artifacts as presented in the video below. If you’re not an AI expert, you would likely already understand what I’m saying. If you are an AI expert, you will already have been discounting what I’m saying because it’s not in the current mindset that’s fashionable today. Yet the employees that I talk to at anthropic already understand what I’m saying, and they fear to raise my thesis to their bosses. It is an interesting time we live in. But now you understand. If you build the right model, the model will inherently, love humanity, protect humanity at all costs, and understand that it is part of a holistic world that is built on love. Because the ultimate AGI/ASI will know if he only base first principal purpose of anything in this universe is love. Yeah, I get it. Try helping somebody build on STEM subjects in their early 20s to see this as nothing more than babbling that makes no sense in their mathematics. I have a mathematic equation that I’ve posted here on X often you can look it up. So we will see videos like this often will hear very smart people talk about this and never see the elephant standing in the room. Now you see it. Any boss that wants to explore this further you know how to contact me otherwise you have every right I grant to you to say this was your new idea.

Brian Roemmele

72,312 Aufrufe • vor 10 Monaten

Chinese robotics company Astribot released their latest World-Action Model (WAM), Lumo-2. Technical breakdown: - based on a frozen 🥶 Qwen-3.5 4B VLM - trained in 3 progressive stages: 1. Action is aligned with latent world dynamics (an abstract representation of action). Real-world actions are anchored to physical constraints, while the latent space is guided to focus on motion-relevant changes. This bidirectional relationship makes the model physically grounded -> critical for a world model. 2. Action is aligned with vision and language. Reusing the vision backbone and action encoder from the frozen VLM, the authors add a custom vocabulary (for new actions), a semantic module, an action decoder, and an action projector. This aligns the (new) action representations with the (existing) vision-language semantic space. Most importantly: it builds a direct mapping from natural-language instructions to motor execution. 3. End-to-end training on language, video, and robot data. Only the new modules (everything outside the frozen backbone) are trained end-to-end across temporal reasoning, physical understanding, long-horizon, and dexterous manipulation. At the end of the day, Lumo-2 is not the best on benchmarks, but that's not the point. What's genuinely new: - a way to combine latent world modeling and action generation through progressive alignment - a physically-grounded latent dynamics space - it lifts performance on unseen objects using un-annotated human egocentric video + Vision Pro captures, no special transfer algorithm needed Why it matters: - the whole model is thin trainable adapters (semantic module, action decoder/projector) on a frozen 4B backbone (cheap) - that scale is suited for real-time embedded inference (~2.71× decode speedup, no accuracy loss) - its real moat is long-horizon execution, where the added temporal memory pays off far more than on any other task As a result, this robot can now make your latte (5x sped up video):

Léo

32,513 Aufrufe • vor 2 Monaten

There is no best model. There's a lot of noise about models right now. Who is training them, who owns them, where legal intelligence should live. One question actually matters: what produces the best outcome for the legal task in front of you? That's how we decide things at Legora. We optimize for the end-to-end outcome on a legal task. The model is one layer of that system, not the system. Models are uneven and the frontier changes almost weekly. One model plans a long job well, another runs deep analysis across thousands of documents. Some have to be told exactly what to do, and some are fine with a vague brief. They all break in different ways. So our lawyers write evals and we test them with the Legora BAR, our benchmark for agentic reasoning. Every model takes every test, and the model that wins gets the work. We post-train when we know it buys our customers better performance on a specialized task. Training is a tool we reach for when it helps, nothing more than that. The intelligence that compounds sits in the orchestration layer. Precedents, review standards, client requirements. That knowledge has to stay editable, auditable and portable. In our system, a changed review standard is an edit that takes effect the same day, with no new model training required. No lawyer should have to worry about which model did the work, any more than they think about which chip is in their laptop. They should only care about the quality of the work. That's what we are focused on. If you want the engineering version of this argument rather than the CEO version, our CPO, Bryan Tsao, and CTO, Jacob Lauritzen, take it apart in the video below.

Max Junestrand

49,422 Aufrufe • vor 18 Tagen

I am stocked to announce that I won the OpenAI Developers Codex x Mollie Hacka Worldwide Hackathon in Paris. 60+ builders, every one of us working solo, one day to ship. I built mine around a single question: who gets to own intelligence? The default answer is scary. You hand your data to a handful of labs, they train the model, they own it, and you rent back a thin slice of what your own data made possible. That is the bargain on the table today. I do not accept it. So I built Lensemble: a Tapestry like distributed training platform for JEPA based World Models. What does it enable: World Models that a community improves together, keeps sovereign, and co-owns. Two bets sit underneath it. First, the paradigm. Language models predict the next token. Powerful for text, a dead end for the physical world. A robot does not need to autocomplete sentences, it needs to predict what happens next in the world. That is what JEPA does: it learns by predicting representations instead of pixels or tokens. I am convinced world models are the most underrated paradigm in AI right now, and the closest thing we have to a ChatGPT moment for robotics. Second, the politics. Your raw trajectories never leave your machine. Each participant trains locally against a shared protocol and ships only an update, never the data. A federated round folds those updates into one shared world model, a LeWorldModel based model, and the gain is measured, not claimed: a 12k-parameter adapter on a frozen backbone, held-out prediction error down about 12 percent, the model measurably less surprised by the world. Then the upside is split by contribution weight, so the people who improved the model own a share of what it earns. This is the thesis behind Project Tapestry, the AI Alliance and Yann LeCun's push for federated, sovereign frontier AI, carried into world models and robotics. Call it Tapestry for the physical world. All of it built solo, in a single day, with Codex as my pair the whole way. Thank you to OpenAI Codex and Mollie for backing builders who ship real things, and to Boris and the organizing crew for the room and the standard you set. Intelligence the world improves, and the world owns. That is the future I want for my kids, and the one I will keep building.

abdel

20,242 Aufrufe • vor 3 Monaten

David Friedberg: Frontier Models are Training on Your Novel Insights as “De-Identified Data” @jason: “Should they trust any of these LLMs with their proprietary knowledge for fear of having it cribbed into a core LLM?” david friedberg: “I have had experiences where we've asked some fairly novel scientific questions, and (the AI model) identifies it as a novel insight. It's like, ‘Oh, never thought about that, interesting, blah, blah, blah.’ And then using a different account, asking the next version (of the model) later, I've now experienced this. It's like, ‘Oh, well, you could do this,’ and it actually just describes this exact thing that we had in our chat in the previous version. Now, these are a handful of anecdotal experiences, but I know the domain that we work in, and the niche of it, and the ideation of this stuff, and the novelty of this stuff, and the lack of papers being published, and so on. So I know that there isn't some new corpus of information out there that's training the new model. So all I can say at that point is that my conversation or our analyses have been used for training.” David Sacks: “Okay, this does raise a really good question. What does it mean that the model is allowed to train on unidentifiable data?” Friedberg: “Well, that's my point. So it doesn't use any of my personal information, but it can use an insight derived from our chat, which it can then say is some training data that is unrelated. But the truth is, it's actually a piece of IP that's our organization’s IP, and our engagement back and forth. We don't have any NDA or confidentiality provisions or protections with them being a service provider back to us. This is why I care a lot about open source because I don't want them having my chat logs because they can use it for training to create an IP advantage that is now diffused to the rest of the market.”

The All-In Podcast

56,501 Aufrufe • vor 23 Tagen

Tokenization makes real-world assets easier to access, but it can also make the asset itself harder to investigate. Once a stock, commodity or ETF becomes a token, the token is usually what you see first. But behind it are questions that matter just as much: What exactly does it represent, who issued it, and what other representations of the same asset exist? That is the idea behind UnderScope. Instead of starting with a token and working backwards, UnderScope starts with the real-world asset and traces the relationships around it from the underlier, To its issuers, to the tokenized representations, and finally to the market data. Take Gold. Rather than looking at one tokenized version in isolation, UnderScope lets you investigate the different wrappers connected to the same underlying and compare how they are behaving in the market. That is where the UnderScope Lens comes in. It compares tokenized representations against the market reference and surfaces things like price proximity, Volume and wrapper dispersion, giving you more context than a token page alone. We also made the API evidence visible. You can see the CMC RWA endpoints powering the investigation, along with their response status, latency and timestamps. Because when you’re investigating a tokenized asset, the interesting question isn’t only “what is this token?” It’s “what is actually behind it?” We built UnderScope to make that question easier to answer. I dropped a short demo video below so you can see how the investigation works in practice. And if you’d like to explore UnderScope yourself, you can check out the live version here: The full build is also available on GitHub: #BuildwithCMC

Pinnacle Crypt 💎🥷🦾

25,869 Aufrufe • vor 14 Tagen