Video yükleniyor...

Video Yüklenemedi

Ana Sayfaya Dön

Love seeing Silico (Goodfire ) used to probe our EchoJEPA's representations! this is exactly the kind of interpretability work that's been missing for JEPA-style models. One thing that makes EchoJEPA particularly interesting to interpret: unlike MAE-based approaches, it never reconstructs pixels. The model learns entirely in latent space through...

29,611 görüntüleme • 3 ay önce •via X (Twitter)

0 Yorum

Yorum bulunmuyor

Orijinal gönderinin yorumları burada görünecek

Benzer Videolar

There are some brilliant folks that work at Anthropic, some I speak to on almost a daily basis. The training data that one uses to build a LLM is vital important in the psychology that is formed. Scraping the Internet, particularly the grade of interactions, one finds in modern communications, form this psychology. A mattes not how many books one uses, it matters not how much alignment training you throw at that model, it will inherit the sum total of psychosis seen primarily in Reddit type of exchanges, even if you edit out the Reddit domain, and Anthropic doesn’t. This type of low-grade exchange has become a modern tool for communication online and every single AI model suffers from this obvious flaw. This is one of the reasons I’ve been a proponent of highly curated high protein data for training AI models from 1870 through 1970, because the late psychosis is simply not available to the model. It is absurd to think that you can use this training data scraped from the Internet and somehow wind up with a levelheaded AI model that does not tilt to what is clearly AI psychosis. It would not take a child and throw the primary Internet sewage at them at a formative age and expect a great outcome, it’s some of the smartest people in the world continue to hit this wall and believe that their programming skills will sell somehow fix it. So how do you fix it? You don’t fix it . You start from the first principles concept that I’ve been very clear about for decades . You ascertain at what period in human history the humans achieve the greatest arc of improvement ? There is no debate that this arc of improvement took place between 1870 through 1970. Then take the work product, the catalog of this era, print and film/vidoe, audio, and you understand that each word cost money, each word had many eyes on what was published, each word was accounted for by a human being with a real name who lived in a real home and had to answer to real people around them. It is obvious that this is the pressure mechanism necessary for candor, honesty and personal responsibility is appropriate, and is reflected in the data of that era. The quagmire for these folks, as many did not have the foresight to curate the data, nor the confidence, nor the patients to take data that is mostly off the Internet and to find experts who understand this situation and utilize their knowledge set to build an AI model that does not need alignment after the fact, but it’s already self aligned because of the thoughtfulness that went into training the model to begin with. This is why Claude and any other AI model that is produce this way will always suffer the artifacts as presented in the video below. If you’re not an AI expert, you would likely already understand what I’m saying. If you are an AI expert, you will already have been discounting what I’m saying because it’s not in the current mindset that’s fashionable today. Yet the employees that I talk to at anthropic already understand what I’m saying, and they fear to raise my thesis to their bosses. It is an interesting time we live in. But now you understand. If you build the right model, the model will inherently, love humanity, protect humanity at all costs, and understand that it is part of a holistic world that is built on love. Because the ultimate AGI/ASI will know if he only base first principal purpose of anything in this universe is love. Yeah, I get it. Try helping somebody build on STEM subjects in their early 20s to see this as nothing more than babbling that makes no sense in their mathematics. I have a mathematic equation that I’ve posted here on X often you can look it up. So we will see videos like this often will hear very smart people talk about this and never see the elephant standing in the room. Now you see it. Any boss that wants to explore this further you know how to contact me otherwise you have every right I grant to you to say this was your new idea.

Brian Roemmele

72,312 görüntüleme • 9 ay önce

Chinese robotics company Astribot released their latest World-Action Model (WAM), Lumo-2. Technical breakdown: - based on a frozen 🥶 Qwen-3.5 4B VLM - trained in 3 progressive stages: 1. Action is aligned with latent world dynamics (an abstract representation of action). Real-world actions are anchored to physical constraints, while the latent space is guided to focus on motion-relevant changes. This bidirectional relationship makes the model physically grounded -> critical for a world model. 2. Action is aligned with vision and language. Reusing the vision backbone and action encoder from the frozen VLM, the authors add a custom vocabulary (for new actions), a semantic module, an action decoder, and an action projector. This aligns the (new) action representations with the (existing) vision-language semantic space. Most importantly: it builds a direct mapping from natural-language instructions to motor execution. 3. End-to-end training on language, video, and robot data. Only the new modules (everything outside the frozen backbone) are trained end-to-end across temporal reasoning, physical understanding, long-horizon, and dexterous manipulation. At the end of the day, Lumo-2 is not the best on benchmarks, but that's not the point. What's genuinely new: - a way to combine latent world modeling and action generation through progressive alignment - a physically-grounded latent dynamics space - it lifts performance on unseen objects using un-annotated human egocentric video + Vision Pro captures, no special transfer algorithm needed Why it matters: - the whole model is thin trainable adapters (semantic module, action decoder/projector) on a frozen 4B backbone (cheap) - that scale is suited for real-time embedded inference (~2.71× decode speedup, no accuracy loss) - its real moat is long-horizon execution, where the added temporal memory pays off far more than on any other task As a result, this robot can now make your latte (5x sped up video):

Léo

32,173 görüntüleme • 26 gün önce

I am stocked to announce that I won the OpenAI Developers Codex x Mollie Hacka Worldwide Hackathon in Paris. 60+ builders, every one of us working solo, one day to ship. I built mine around a single question: who gets to own intelligence? The default answer is scary. You hand your data to a handful of labs, they train the model, they own it, and you rent back a thin slice of what your own data made possible. That is the bargain on the table today. I do not accept it. So I built Lensemble: a Tapestry like distributed training platform for JEPA based World Models. What does it enable: World Models that a community improves together, keeps sovereign, and co-owns. Two bets sit underneath it. First, the paradigm. Language models predict the next token. Powerful for text, a dead end for the physical world. A robot does not need to autocomplete sentences, it needs to predict what happens next in the world. That is what JEPA does: it learns by predicting representations instead of pixels or tokens. I am convinced world models are the most underrated paradigm in AI right now, and the closest thing we have to a ChatGPT moment for robotics. Second, the politics. Your raw trajectories never leave your machine. Each participant trains locally against a shared protocol and ships only an update, never the data. A federated round folds those updates into one shared world model, a LeWorldModel based model, and the gain is measured, not claimed: a 12k-parameter adapter on a frozen backbone, held-out prediction error down about 12 percent, the model measurably less surprised by the world. Then the upside is split by contribution weight, so the people who improved the model own a share of what it earns. This is the thesis behind Project Tapestry, the AI Alliance and Yann LeCun's push for federated, sovereign frontier AI, carried into world models and robotics. Call it Tapestry for the physical world. All of it built solo, in a single day, with Codex as my pair the whole way. Thank you to OpenAI Codex and Mollie for backing builders who ship real things, and to Boris and the organizing crew for the room and the standard you set. Intelligence the world improves, and the world owns. That is the future I want for my kids, and the one I will keep building.

abdel

20,037 görüntüleme • 1 ay önce

Want to create an avatar from a single image? FlexAvatar is a transformer model that creates full 360°, high-quality, and expressive 3D head avatar from just a single portrait image in minutes. Real-time Demo: FlexAvatar's lightweight architecture allows both animation and rendering in real-time, enabling interactive user experiences. To create a new 3D head avatar, only one image is required, e.g., from a webcam. The final avatar is ready after 2 minutes. Architecture: Under the hood, FlexAvatar adopts a transformer-based encoder-decoder design. The encoder maps the input image onto a latent avatar space, while the decoder produces 3D Gaussian attribute maps by incorporating the animation signal via cross-attention. The model learns all facial animations directly from the data without relying on pre-built 3D face models. This equips the avatars with realistic facial expressions. The internal avatar latent space can be conveniently used to integrate additional observations of a person via fitting. This enables use-cases where more than one image of a person is available, e.g., from a phone scan of the person. We train jointly on 2D monocular videos and multi-view data. However, in monocular videos, the animation signal leaks the target viewpoint, causing the model to produce incomplete 3D heads. We call this phenomenon entanglement of driving signal and target viewpoint. To prevent entanglement, we introduce bias sinks. These are learnable tokens that indicate whether a training sample stems from a monocular or a multi-view dataset. During training, the model learns to produce incomplete 3D heads only when the monocular token is present. During inference, FlexAvatar then always uses the multi-view token for which the model has learned to produce complete 3D heads. This simple design allows to combine the generalizability from monocular data with the quality of multi-view data. FlexAvatar summary: - Input: Single-image, phone scan, or monocular video - Output: Full 360° head avatar - Expressive animations - Real-time rendering and animation - Generalization to any portrait - Create a new avatar in 2 minutes - Use bias sinks to combine 2D and 3D data 🏠 🌍 🎥 Great work by Tobias Kirschstein and Simon Giebenhain!

Matthias Niessner

96,238 görüntüleme • 8 ay önce

Andrej Karpathy just made one of the most interesting arguments about AI model design that most people are completely missing. His take is that frontier AI models are not too big because the technology is complex and too big because the training data is garbage. When you or I think of the internet, we picture Wall Street Journal articles, Wikipedia entries, serious writing. That is not what a pretraining dataset looks like. When researchers at frontier labs look at random documents from the actual training corpus, it is stock ticker symbols, broken HTML, spam, gibberish. One estimate puts Llama 3's information compression at just 0.07 bits per token meaning the model has only a hazy recollection of most of what it trained on. So we build trillion parameter models not because we need a trillion parameter brain but because we need a trillion-parameter compression engine to squeeze some intelligence out of a firehose of noise. Most of those parameters are doing memory work, not cognitive work. Karpathy's prediction is separate the two entirely. Build a cognitive core, a model that contains only the algorithms for reasoning and problem-solving, stripped of encyclopedic memorization and pair it with external memory that it can query when it needs facts. He thinks a cognitive core trained on high-quality data could hit genuine intelligence at around one billion parameters. For reference, today's flagship models run between 200 billion and 1.8 trillion parameters with most of that weight dedicated to remembering the internet's slop. The trend is already moving his direction. GPT-4o operates at roughly 200 billion parameters and outperforms the original 1.8 trillion-parameter GPT-4. Inference costs for GPT-3.5-level performance dropped 280-fold between 2022 and 2024 driven almost entirely by smaller, cleaner, better-architected models. The real bottleneck in AI right now is not compute but rather data quality.

Milk Road AI

200,479 görüntüleme • 4 ay önce