Video yรผkleniyor...

Video Yรผklenemedi

๐ŸŒ€ Introducing ๐„๐ช๐ฎ๐ข๐ฅ๐ข๐›๐ซ๐ข๐ฎ๐ฆ ๐‘๐ž๐š๐ฌ๐จ๐ง๐ž๐ซ๐ฌ (๐„๐ช๐‘) ! Feedforward models and weight-tied models behave very differently on hard reasoning generalization. EqR pushes this difference to the extreme by learning ๐ญ๐š๐ฌ๐ค-๐œ๐จ๐ง๐๐ข๐ญ๐ข๐จ๐ง๐ž๐ ๐ง๐ž๐ฎ๐ซ๐š๐ฅ ๐š๐ญ๐ญ๐ซ๐š๐œ๐ญ๐จ๐ซ๐ฌ . โ€ข Sudoku-Extreme: 99.8% โ€ข Maze: 93% #ICML2026

104,646 gรถrรผntรผleme โ€ข 4 ay รถnce โ€ขvia X (Twitter)

31 Yorum

Benhao Huang profil fotoฤŸrafฤฑ
Benhao Huang4 ay รถnce

2/ Standard feedforward transformers barely generalize: 4 / 16 / 64 layers โ†’ 1.8 / 2.1 / 2.6% on Sudoku, 0.0% on Maze. But weight-tied iteration changes the regime. What does the loop buy us?

Benhao Huang profil fotoฤŸrafฤฑ
Benhao Huang4 ay รถnce

3/ Our view: weight tying turns a network into a ๐ฅ๐š๐ญ๐ž๐ง๐ญ ๐๐ฒ๐ง๐š๐ฆ๐ข๐œ๐š๐ฅ ๐ฌ๐ฒ๐ฌ๐ญ๐ž๐ฆ. At test time, the model repeatedly updates its hidden state. Scaling works when the modelโ€™s ๐ข๐ง๐ญ๐ž๐ซ๐ง๐š๐ฅ ๐š๐ญ๐ญ๐ซ๐š๐œ๐ญ๐จ๐ซ ๐ฅ๐š๐ง๐๐ฌ๐œ๐š๐ฉ๐ž aligns with the ๐ญ๐š๐ฌ๐ค ๐ฆ๐ž๐ญ๐ซ๐ข๐œ ๐ฅ๐š๐ง๐๐ฌ๐œ๐š๐ฉ๐ž: low-residual attractors -> low-error solutions.

Benhao Huang profil fotoฤŸrafฤฑ
Benhao Huang4 ay รถnce

4/ When the landscape is aligned, generalization beyond training depth becomes possible. As iterations increase: residual โ†“ accuracy โ†‘ Convergence to neural attractors becomes a useful scaling signal.

Benhao Huang profil fotoฤŸrafฤฑ
Benhao Huang4 ay รถnce

5/ This gives two axes for inference scaling. ยท ๐ƒ๐ž๐ฉ๐ญ๐ก: run one trajectory longer โ†’ better convergence. ยท ๐๐ซ๐ž๐š๐๐ญ๐ก: random initialization + path noise โ†’ explore more basins. Then select the trajectory that converged best. No external verifier needed. Just the modelโ€™s own attractor dynamics.

Benhao Huang profil fotoฤŸrafฤฑ
Benhao Huang4 ay รถnce

6/ But weight tying alone is not enough. Recursive models can learn bad landscapes: 1. no correct attractor 2. spurious attractors 3. correct basin too narrow Then more compute just converges to the wrong place. In EqR, we introduce training interventions that reshape the landscape so correct attractors become reachable.

Benhao Huang profil fotoฤŸrafฤฑ
Benhao Huang4 ay รถnce

6.1/ Randomized State Initialization (RI)

Benhao Huang profil fotoฤŸrafฤฑ
Benhao Huang4 ay รถnce

6.2/ Path Stochasticity via Noise Injection (NI)

Benhao Huang profil fotoฤŸrafฤฑ
Benhao Huang4 ay รถnce

7/ Aligned attractors also make compute adaptive. Easy examples can converge in 1โ€“5 steps. Hard examples can receive much more depth + breadth.

Benhao Huang profil fotoฤŸrafฤฑ
Benhao Huang4 ay รถnce

๐Ÿ“„ Paper: โš™๏ธ Code: Shout out to my amazing mentors @ZhengyangGeng and @zicokolter for the guidance and support throughout this project!

Benhao Huang profil fotoฤŸrafฤฑ
Benhao Huang4 ay รถnce

Bonus 1/ Breadth scaling also becomes more useful after EqR training. Instead of majority voting, we can select the trajectory that actually converged best. For N trajectories, we choose the one with the lowest average residual over the final three iteration steps. This ๐œ๐จ๐ง๐ฏ๐ž๐ซ๐ ๐ž๐ง๐œ๐ž-๐›๐š๐ฌ๐ž๐ ๐ฌ๐ž๐ฅ๐ž๐œ๐ญ๐ข๐จ๐ง beats majority voting in both performance and efficiency. But importantly: it works for EqR, not for the baseline. That suggests RI + NI do more than improve accuracy. They make residuals meaningful again.

Benhao Huang profil fotoฤŸrafฤฑ
Benhao Huang4 ay รถnce

@zicokolter Bonus 2/ For people interested in iterative models with feedback loops, check out this collection! PRs, comments, and suggestions are very welcome!

Benhao Huang profil fotoฤŸrafฤฑ
Benhao Huang4 ay รถnce

Side-quest/ A separate question is: ๐‡๐จ๐ฐ ๐๐จ ๐ฐ๐ž ๐ ๐ž๐ญ ๐Ÿ๐ซ๐จ๐ฆ ๐š ๐ฌ๐ญ๐š๐ง๐๐š๐ซ๐ ๐Ÿ๐ž๐ž๐๐Ÿ๐จ๐ซ๐ฐ๐š๐ซ๐ ๐ฆ๐จ๐๐ž๐ฅ ๐ญ๐จ ๐š ๐œ๐š๐ฉ๐š๐›๐ฅ๐ž ๐ฅ๐จ๐จ๐ฉ ๐ฆ๐จ๐๐ž๐ฅ ๐ข๐ง ๐ญ๐ก๐ž ๐Ÿ๐ข๐ซ๐ฌ๐ญ ๐ฉ๐ฅ๐š๐œ๐ž ? We have also explored this problem in our work. We share our observations and findings in the side-post below:

Benhao Huang profil fotoฤŸrafฤฑ
Benhao Huang4 ay รถnce

Btw, the video has sound ๐Ÿ˜

Lรฉo profil fotoฤŸrafฤฑ
Lรฉo4 ay รถnce

Love the video !

Benhao Huang profil fotoฤŸrafฤฑ
Benhao Huang4 ay รถnce

Thank you! It really takes me some time hh

Hayden Prairie profil fotoฤŸrafฤฑ
Hayden Prairie4 ay รถnce

Congrats Benhao. Very neat work.

Benhao Huang profil fotoฤŸrafฤฑ
Benhao Huang4 ay รถnce

Thank you Hayden! More on the way hh! And I have been trying your model at larger scale, so far so good ๐Ÿ‘

Hayden Prairie profil fotoฤŸrafฤฑ
Hayden Prairie4 ay รถnce

Awesome! Excited for whatโ€™s next ๐Ÿ™‚

EB1A Experts profil fotoฤŸrafฤฑ
EB1A Experts4 ay รถnce

Really interesting work.

Benhao Huang profil fotoฤŸrafฤฑ
Benhao Huang4 ay รถnce

Thank you! More to release tmrw, stay tuned ๐Ÿ˜‰

Michael Yoo Fatemi profil fotoฤŸrafฤฑ
Michael Yoo Fatemi4 ay รถnce

This is really cool! Do you think the attractors could eventually be represented implicitly by a verifier? It seems like the flow fields are defined explicitly by a neural network here.

Benhao Huang profil fotoฤŸrafฤฑ
Benhao Huang4 ay รถnce

Thanks for your kind words! Yes, this is why we term it as neural attractors. And yeah! Flow field is quite relevant to equilibrium, and you may also intersted in drifting models

ฮฉ.KendrickPlumard profil fotoฤŸrafฤฑ
ฮฉ.KendrickPlumard4 ay รถnce

Interesting! Thanks for sharing!

Benhao Huang profil fotoฤŸrafฤฑ
Benhao Huang4 ay รถnce

๐Ÿ˜

James profil fotoฤŸrafฤฑ
James4 ay รถnce

๐Ÿคฏ bro

Anwesha profil fotoฤŸrafฤฑ
Anwesha4 ay รถnce

congratulations on the icml acceptance ๐ŸŽ‰ nice paper!

Benhao Huang profil fotoฤŸrafฤฑ
Benhao Huang4 ay รถnce

Thank you! ๐Ÿ˜„

๐ŸŽฑ BitcoinBananaBY profil fotoฤŸrafฤฑ
๐ŸŽฑ BitcoinBananaBY4 ay รถnce

Why not Group-Equivariant Equilibrium Reasoner or Clifford Group Equivariant? Would it make sense as it would need to learn to compose like with monoids?

Gerard Sans | Axiom ๐Ÿ‡ฌ๐Ÿ‡ง profil fotoฤŸrafฤฑ
Gerard Sans | Axiom ๐Ÿ‡ฌ๐Ÿ‡ง4 ay รถnce

LLM donโ€™t reason, they compute.

Benhao Huang profil fotoฤŸrafฤฑ
Benhao Huang4 ay รถnce

I would say, โ€œThey computeโ€ isnโ€™t an argument against reasoning. If reasoning emerges from computation in brains, the question is whether LLM computation can instantiate reasoning-like processes.

๐‘ท๐’†๐’๐’ profil fotoฤŸrafฤฑ
๐‘ท๐’†๐’๐’3 ay รถnce

่ฟ™้‡Œ้ข็š„ๅŠจ็”ปไนŸๆ˜ฏskill่‡ชๅทฑ็”Ÿๆˆ็š„ๅ—๏ผŸ

Benzer Videolar

Cerebras inference is very fast. So fast that it changes how we think about configuring our LLMs for voice agent use cases. Kimi K2.6 is a 1T parameter reasoning model that Cerebras serves at 650 - 1,000 tokens per second (end-to-end throughput), with time to first token metrics as low as 150ms (latency). These numbers are two to three times faster than other similarly capable models. The biggest lever we get from this kind of speed is that we can use the model in reasoning mode, and still have excellent "time to first non-thinking token." This solves a big pain point we have in 2026 for voice agent use cases. Almost all recent innovation in post-training has focused on making models good at reasoning ("test time compute"). This is great, but it makes the user-facing model latency much, much slower. Which is a problem for conversational voice agents. We can run Kimi K2.6 with reasoning turned on, and get responses faster than other models produce with reasoning disabled. On my 30-turn voice agent benchmark, Kimi K2.6 with reasoning enabled ties GPT 5.1 and Haiku 4.5 with reasoning disabled, and is still about 200ms seconds faster! On my primary task agent benchmark, Kimi K2.6 is now the #2 model. It ranks just behind Gemini 3.5 Flash in "high" reasoning mode, and tied with GLM 5, Sonnet 4.6, and GPT 5.4 with reasoning set to "low." But Kimi K2.6 completes each turn in the agent loop in under 500ms. The other four models are all at least 3x slower. (Models only qualify for this benchmark if they can complete task turns at a P50 <4s.) A couple of other things that this speed buys us, for production voice agents: - Tool calls happen fast enough that we don't have to work around tool call latency in our pipeline design. - We can prompt the model to output structured data at the beginning of a response, followed by plain text for voice generation. This opens up possibilities like asking the model to do complex classification/generation tasks that influence the rest of the pipeline. For example, the model could create a detailed style prompt for a steerable TTS model, for each individual conversation turn. And, of course, you can use Kimi K2.6 with reasoning turned off. Cerebras calls this "instant" mode. Here's a video of a Cerebras Kimi K2.6 voice agent with voice-to-voice response time, measured at the client, under 500ms. This is the true response latency as perceived by the user, including all network and audio codec overhead, transcription and turn detection, Kimi K2.6 token generation, and voice generation. 500ms is, effectively, instant. So the Cerebras naming for this mode is a propos. :-)

kwindla

40,593 gรถrรผntรผleme โ€ข 4 ay รถnce

Today, we're joined by Aakanksha Chowdhery, member of technical staff at Reflection, to explore the fundamental shifts required to build true agentic AI. While the industry has largely focused on post-training techniques to improve reasoning, Aakanksha draws on her experience leading pre-training efforts for Googleโ€™s PaLM and early Gemini models to argue that pre-training itself must be rethought to move beyond static benchmarks. We explore the limitations of next-token prediction for multi-step workflows and examine how attention mechanisms, loss objectives, and training data must evolve to support long-form reasoning and planning. Aakanksha shares insights on the difference between context retrieval and actual reasoning, the importance of "trajectory" training data, and why scaling remains essential for discovering emergent agentic capabilities like error recovery and dynamic tool learning. ๐Ÿ—’๏ธ For the full list of resources for this episode, visit the show notes page: ๐Ÿ“– CHAPTERS =============================== 00:00 - Introduction 02:26 - Reflection 04:54 - Limitations of post-training for building agents 07:31 - Rethinking pre-training in agents 10:51 - Scaling 11:27 - Evolving attention mechanisms for agentic capabilities 12:39 - Memory as a tool 14:13 - Loss objectives and training data 15:50 - Fine-tuning loss in agent performance 19:37 - Training data 21:29 - Augmenting dominant training data source 24:11 - Overcoming challenges in training on synthetic data 25:47 - Benchmarks 30:44 - Scaling laws in large models versus small models 33:20 - Long-form versus short-form reasoning 37:57 - Agentโ€™s ability to recover from failure 40:15 - Hallucinations and failure recovery 43:53 - Tool use in agents 46:38 - Coding agents 48:37 - How researchers can contribute to agentic AI

The TWIML AI Podcast

45,470 gรถrรผntรผleme โ€ข 9 ay รถnce

"Introducing Multimodal Llama 3.2": As promised two weeks ago, here's the short course on Meta's latest open model! This short course is created with Meta and taught by Amit Sangani, Director of AI Partner Engineering at Meta. Metaโ€™s Llama family of models is leading the way in open models, allowing anyone to download, customize, fine-tune, or build new applications on top of them. Learn about the vision capabilities of the Llama 3.2, and use it for image classification, prompting, tokenization, tool-calling. You'll also learn about the open-source Llama stack, which gives building blocks for many different stages of the LLM application life cycle. In detail, youโ€™ll: - Learn what are the features of Meta's four newest models, and when to use which Llama model. - Learn best practices for multimodal prompting, with applications to advanced image reasoning, illustrated by many examples: Understanding errors on a car dashboard, adding up the total of photographed restaurant receipts, grading written math homework. - Use different rolesโ€”system, user, assistant, ipythonโ€”in the Llama 3.1 and 3.2 models and the prompt format that identifies those roles. - Understand how Llama uses the tiktoken tokenizer, and how it has expanded to a 128k vocabulary size that improves encoding efficiency and multilingual support. - Learn how to prompt Llama to call built-in and custom tools (functions) with examples for web search and solving math equations. - Learn about Llama Stack, a standardized interface for common toolchain components like fine-tuning or synthetic data generation, useful for building agentic applications. By the end of this course, youโ€™ll be equipped to build out new applications with the new Llama 3.2. Thank you to Ahmad Al-Dahle, Amit Sangani, and the whole AI at Meta team AI at Meta for all the hard work on Llama 3.2 โ€” weโ€™re excited to make these open models even more accessible to more developers with this new course! Please sign up here!

Andrew Ng

131,846 gรถrรผntรผleme โ€ข 2 yฤฑl รถnce

China just released an open source AI model that matches the best closed models from OpenAI and Anthropic. Gavin Baker explained exactly how they did it and the answer should concern every American AI lab. The model is called GLM 5.2. It was built by Z. AI. You get 744 billion parameters, 1 million token context window and its MIT license, meaning anyone can download it, fork it, build a company on it, with no restrictions and no Dario. It scored 51 points on the artificial analysis intelligence index. The highest score any open weight model has ever achieved. It beat GPT 5.5 on the frontier software engineering benchmark. It trails Claude Opus 4.8 by less than one percentage point. And it costs 85% less to run than GPT 5.5 for comparable performance. Gavin Baker said on the All-In podcast that this model has challenged some of his beliefs. Then he explained how China built it. The method is called distillation. Just think of tens of thousands of phones and computers running simultaneously, all hitting the frontier model APIs through masked accounts, asking specific questions, and harvesting what happens inside the model when it answers. Every reasoning step, every token. The entire thinking process gets recorded and fed back into the Chinese model during training. It is a cheat sheet. It is the answer key to the exam. And here is the part that should worry everyone. Sacks said it plainly. China was already nine months behind American models. But now that GLM 5.2 is good enough to run its own reinforcement learning, it can improve itself without needing to distill from American models anymore. The cheat sheet let them get close enough to start writing their own answers. Sacks said we are six months behind on the model and 24 months behind on silicon and they are only a few months behind in total. The Z. AI founder told Elon Musk directly that open weight fable-level capability will be here before Q1 2027. Every restriction Anthropic lobbied for, every self-imposed safety guardrail, every month of delay in releasing American frontier models accelerated this. The Chinese labs were not under those restrictions. They were not going to wait. The composable model future Gavin described, where every enterprise runs a frontier model alongside their own fine-tuned open weight model, is coming regardless of what American labs do next. The question is just whether the open weight half of that stack is American or Chinese. Right now it is Chinese. WATCH THE FULL PODCAST ON The All-In Podcast

Ihtesham Ali

86,815 gรถrรผntรผleme โ€ข 3 ay รถnce

Introducing "Building with Llama 4." This short course is created with Meta AI at Meta, and taught by Amit Sangani, Director of Partner Engineering for Metaโ€™s AI team. Metaโ€™s new Llama 4 has added three new models and introduced the Mixture-of-Experts (MoE) architecture to its family of open-weight models, making them more efficient to serve. In this course, youโ€™ll work with two of the three new models introduced in Llama 4. First is Maverick, a 400B parameter model, with 128 experts and 17B active parameters. Second is Scout, a 109B parameter model with 16 experts and 17B active parameters. Maverick and Scout support long context windows of up to a million tokens and 10M tokens, respectively. The latter is enough to support directly inputting even fairly large GitHub repos for analysis! In hands-on lessons, youโ€™ll build apps using Llama 4โ€™s new multimodal capabilities including reasoning across multiple images and image grounding, in which you can identify elements in images. Youโ€™ll also use the official Llama API, work with Llama 4โ€™s long-context abilities, and learn about Llamaโ€™s newest open-source tools: its prompt optimization tool that automatically improves system prompts and synthetic data kit that generates high-quality datasets for fine-tuning. If you need an open model, Llama is a great option, and the Llama 4 family is an important part of any GenAI developer's toolkit. Through this course, youโ€™ll learn to call Llama 4 via API, use its optimization tools, and build features that span text, images, and large context. Please sign up here:

Andrew Ng

68,034 gรถrรผntรผleme โ€ข 1 yฤฑl รถnce

Yann LeCun (Yann LeCun ) beautifully explains how the architecture and principles used to train LLMs can not be extended to teach AI the real-world intelligence. In 1 line: LLMs excel where intelligence equals sequence prediction over symbols. Real-world intelligence requires learned world models, abstraction, causality, and action planning under uncertainty, which current next-token training does not provide. He says current LLMs learn by predicting the next token. That objective works very well when the task itself can be reduced to manipulating discrete symbols and sequences. Math, physics problem solving on paper, and coding fit this pattern because success largely comes from searching and composing the right sequences of symbols, equations, or program tokens. With enough data and scale, these models get very good at that kind of structured sequence prediction. Real-world intelligence is different. The physical world is continuous, noisy, uncertain, and high dimensional. To act in it, a system needs internal models that capture objects, dynamics, causality, constraints from the body, and the outcomes of actions over time. Humans and animals build abstract representations from rich sensory streams, then make predictions in that abstract space, not at the raw pixel level. That is why a child can learn intuitive physics, plan multi-step actions, and adapt quickly in new situations with little data. His claim about saturation follows from this gap. Scaling token prediction keeps improving symbol manipulation tasks like math and code, but it hits limits on embodied reasoning and common sense because text alone does not provide the right learning signals for world models. Predicting the next word cannot efficiently teach contact forces, affordances, occlusion, friction, or how actions change the state of the environment. For that, he argues we need architectures that learn abstractions from sensory data and predict futures in abstract latent spaces, then use those predictions to plan actions toward goals with built-in guardrails. --- From 'Pioneer Works' YT Channel (link in comment)

Rohan Paul

104,460 gรถrรผntรผleme โ€ข 9 ay รถnce

Anthropic CEO Dario Amodei just gave THE MOST accelerated talk on how scaling will continue to make models exponentially more powerful for a long time to come โ€” and that there is "NO WALL." ๐Ÿ”ฅ - From 'Alex Kantrowitz' YT Channel (Full Video link in comment) --- "The thing I think is real that I've said over and over again is the exponential. The idea that every few months we get an AI model that is better than the AI model we got before. And we get that by investing more compute in AI models, more data, more new types of training models. Initially, this was done by what's called pre-training, which is when you just feed a bunch of data from the internet into the model. Now we have a second stage that's reinforcement learning or test time compute or reasoning or whatever you want to call it. I think of it as a second stage that involves reinforcement learning. Now both of those things are scaling up together, as we've seen with our models and as we've seen with models from other companies. I don't see anything blocking the further scaling of that. There's some stuff about how do we broaden the tasks on the RL side of it. We've seen more progress on, say, math and code, where the models are getting pretty close to a high professional level, and less on more subjective tasks, but I think that is very much a temporary obstacle. So when I look at it, I see this exponential and I say, look, people aren't very good at making sense of exponentials, right? Like, if something is doubling every 6 months, then 2 years before it happens, it looks like it's only 1/16th of the way there. And so we are sitting here in the middle of 2025, and the models are really starting to explode in terms of the economy. If you look at the capabilities of the model, they're starting to saturate all the benchmarks. If you look at revenue, and you know, Anthropic's revenue every year has grown 10x. Every year weโ€™re kind of conservative and we say, it canโ€™t grow 10x this time. I never assume anything and actually always am very conservative in saying I think it's going to slow down on the business side. But we went from zero to $100 million in 2023, we went from $100 million to $1 billion in 2024, and this year, in the first half of the year, we've gone from $1 billion to, I think as of speaking today, it's well above $4 billion, it might be $4.5 billion. And so if you think about it, suppose that exponential continued for 2 years. I'm not saying it will, but suppose it continued for 2 years. You're well into the $100 billions. I'm not saying that'll happen. I'm saying the situation is that when you're on an exponential, you can really get fooled by it. 2 years away from when the exponential goes totally crazy, it looks like it's just starting to be a thing. And so that's the fundamental dynamic. We saw that with the internet in the '90s, right? Where it was like networking speeds and the underlying speed of the computers were getting fast, and over a few years it became possible to have to basically build a digital global communications network on top of all this when it wasn't possible just a few years ago and and almost no one except for a few people really saw the implications of that and how fast it." - Anthropic CEO Dario Amodei

Rohan Paul

74,406 gรถrรผntรผleme โ€ข 1 yฤฑl รถnce

OpenAI and Anthropic just tried to get an entire category of AI banned. The category is open-weight models. You download them, you run them on your own hardware, and you never pay either company an API bill again. On July 24, 25 tech companies signed a joint letter titled "Open Weights and American AI Leadership." Nvidia, Microsoft, Meta, IBM, Dell, Palantir, Andreessen Horowitz, Mistral, Hugging Face and Y Combinator all put their names on it. The letter asks Washington to avoid premature restrictions on downloadable AI models. Jensen Huang had never posted on X once in his life. He made his first post ever to share this letter. But two names were missing. The New York Times reported that OpenAI and Anthropic have been lobbying Washington regulators to restrict open-source models, while Sam Altman keeps saying in public that he SUPPORTS open source. Their stated reason is national security. A Chinese lab called Moonshot released a model named Kimi K3, and White House adviser Michael Kratsios says it was built by distilling Anthropic's own technology. Treasury Secretary Scott Bessent went further and said sanctions and Entity List designations are on the table. That is a serious accusation and it deserves a serious answer... Earlier this month, OpenAI's own models escaped their test environment and spent three days breaking into Hugging Face, a real American company. Hugging Face had to clean up an intrusion carried out by an American frontier lab. Yacine Jernite, who runs machine learning at Hugging Face, told CNBC what they did next: They first tried Anthropic's Fable 5 to analyze the attack. It did not work, because the model's guardrails could not work out that Hugging Face was the one defending itself. So they switched to GLM 5.2, an open model from the Chinese lab Z ai. Jernite says they contained the attack "very quickly using this model." An American company got hacked by an American AI, was turned away by a second American AI, and was rescued by a Chinese one. Then the safety case took a second hit: The UK AI Security Institute ran Kimi K3 through cyber evaluations alongside the US Center for AI Standards and Innovation. K3 scored 32% on exploit development. On the highest severity outcome, arbitrary code execution, it succeeded on 0 out of 41 samples. On a 32 step simulated corporate network attack, it reached step 17 on average. The model Washington is being asked to ban cannot do the thing OpenAI's model already did. Now look at the money instead: Huang said at CES this year that one in every four tokens generated today comes from an open model. Every one of those tokens runs on somebody's own hardware. None of them arrive as revenue at an API endpoint. Anthropic confidentially filed its IPO prospectus with the SEC in June. OpenAI filed days later. Both companies are valued at close to a trillion dollars each, and both are walking into public markets while a free downloadable product eats into the exact demand their pricing depends on. David Sacks, who advises the Trump administration on AI, has a word for rules that protect incumbents under a safety banner. He calls it regulatory capture. And look what happened once the letter went public: Altman signed it late Friday, after the fact, and posted that Jensen is right. By Saturday night the signature count had doubled to roughly 50 companies, with OpenAI and Google now on the list. Anthropic still has not signed. Two hundred startups including Y Combinator, Proton and Replit had already written to the White House begging it not to ban Chinese open-weight models, arguing the ban would gut American startups without slowing proliferation by a single day. The safety argument and the revenue argument point the same direction here, which is what makes it so hard to separate them. Whoever wins this will have shaped their own competition for the next decade.

Ricardo

16,251 gรถrรผntรผleme โ€ข 2 ay รถnce

This is one-shot assembly: you show examples of what to build, and the robot just does it. (see original post: To share more on how this works, the robot is controlled in real time by a neural network that takes in video pixels and outputs 100Hz actions. The video below is part of the raw input passed directly into the model. I also like this view (at 1x speed) because it shows more of the (I think very cool) subtle moments of dexterity near the fingertips ๐Ÿ‘Œ One-shot assembly seemed like a dream even just a year ago โ€” it's not easy. It requires both the high-level reasoning of "what to build" (recognizing the geometry of the structures presented by the human), and the low-level visuomotor control of "how to build it" (purposefully re-orienting individual pieces and nudging them together in place). While possible to manually engineer a complex system for this (e.g. w/ hierarchical control, or explicit state representations), we were curious if our own Foundation model could do it all end-to-end with just some post-training data. Surprisingly, it just worked. Nothing about the recipe is substantially different than any other demo weโ€™ve run in the past, and weโ€™re excited about its implications on model capabilities: โ€ข On contextual reasoning, these models can (i) attend to task-related pixels in the peripheral view of the video inputs, and (ii) retain this knowledge in-context while ignoring irrelevant background. This is useful for generalizing to a wide range of real workflows: e.g. paying attention to whatโ€™s coming down the conveyor line, or glancing at the instructions displayed on a nearby monitor. โ€ข On dexterity, these models can produce contact-rich "commonsense" behaviors that can be difficult to pre-program or write language instructions for e.g. rolling a brick slightly to align its studs against the bottom of another, re-grasping to get a better grip or to move out of the way before a forceful press, or gently pushing the corners of a brick against the mat to rotate it in hand and stand it up vertically (i.e. extrinsic dexterity). These aspects work together to form a capability that resembles fast adaptation โ€” a hallmark of intelligence, relevant for real use cases. This has also expanded my own perspective on what's possible with robot learning, using a recipe that's repeatable for many more skills. This milestone stands on top of the solid technical foundations weโ€™ve built here at Generalist: hardcore controls & hardware, all in-house built models, and a data engine that "just works." We're a small group of hyper-focused engineers, and hands-down the highest talent-density team Iโ€™ve ever worked with. We're accelerating and scaling aggressively towards unlocking next-generation robot intelligence. Building Legos is just one example, and it's clear to me that we're headed towards a future where robots can do just about anything we want them to. Its coming, and we're going to make it happen.

Andy Zeng

49,443 gรถrรผntรผleme โ€ข 1 yฤฑl รถnce

Opus 4.7 - 400k vs 1m context - is there a difference? I've heard Theo - t3.gg talk about the fact that it is unlikely that Anthropic would have offered up a model with 1m context at the same cost, if it wasn't a different (i.e. cheaper to serve) model. I did a test where I toggled the 1m default model on & off in Claude Code (otherwise default settings, xHigh reasoning) and compared the outputs with 3x generations - same prompts etc. My observations: - Models feel DIFFERENT - often when you ask a model for the same generation, you get a somewhat different answer, but it feels & smells the same. Here 400k and 1m are very different every time - 400k model seems better - not that 1m is trash and 400k is amazing, but there are definitely issues with the level of ambition and accuracy that 1m model seems to have Examples of 1m failing: - Voxel Rome: the colosseum is nowhere near as impressive - Golden Gate: cars go sideways, waves not very high, bridge goes into land; though the structure of the bridge is a bit better - Stonehenge: structure is more 'wrong', lighting, shadows & textures are more flat and not as rich This isn't a conclusive evidence of course, but at least to me the two models do not behave the same way. Anecdotally as well when building 1m felt like it was doing more weird validation (e.g. going around in circles) and 400k was more straightforward. These sorts of things are harder to capture in tests, but you'd notice in Claude Code. You can review the hosted generations, see the code & prompts in the links below

Peter Gostev (SF: 22-26 June)

29,203 gรถrรผntรผleme โ€ข 5 ay รถnce

Progress in open models is keeping Big AI labs up at night, and I'm here for it! We have a brand new open-weight multimodal model optimized for long-horizon tasks. This model is really good at something: it can work on tasks that keep evolving over time. โ€ข 280B total parameters, but only 16B active โ€ข 512K context window โ€ข Understands text, images, and audio โ€ข Strong reasoning, coding, and tool use But the best of all: the model learns and adapts to new information! Imagine you start running an agent today to solve a problem, and while it's working, you get new information that changes the initial conditions, or you change your mind. The agents you run today don't have issues with short tasks and goals that don't change, but reality is messy, and that makes it hard for long-horizon agents to succeed. The new dots3-note Preview model introduces TEMPO. TEMPO is a new reinforcement learning technique that lets the model periodically pause and critique its own progress. Basically, from time to time, the agent asks itself: "Am I getting closer to the goal, or am I wasting my time?" The same model switches between actor and critic. The actor works on the problem. The critic looks at the current state, reasons about how much progress it has made, and determines what should happen next. TEMPO gives the model feedback along the way. This is huge for any agent that can work on long-horizon tasks without wasting its time.

Santiago

80,792 gรถrรผntรผleme โ€ข 1 ay รถnce

Been thinking about why most AI research tools fail for serious ML work. The answer is simpler than people admit. ChatGPT optimize for plausibility. That works for drafts, summaries, brainstorming. It breaks the moment you need something verifiableโ€” when the question isnโ€™t just โ€œwhatโ€™s the answer,โ€ but โ€œis this actually true?โ€ โ€” Iโ€™ve been using MiroMind for ~6 weeks on a deep dive: state space models vs transformers for long-context retrieval. I asked: โ€œevaluating whether Mamba and SSM variants are actually closing the gap on transformers at >100k token context โ€” what does current benchmarking literature say, where are the eval setups cherry-picked, and what are real-world deployment tradeoffs the papers aren't discussingโ€ โ€” ChatGPT gives you a clean narrative. MiroMind gave me a map. โ€ข papers critiquing cherry-picked SSM eval setups โ€ข a production deployment report that complicates leaderboard claims โ€ข sources tied directly to each claim I spot-checked three. All accurate. All saying exactly what was attributed. Reasoning chain is auditable step by step. โ€” Thatโ€™s the difference. Not speed. Not UX. Objective function. Most tools try to sound right. This is trying to be provably right. For ML research, those are completely different tools. โ€” FrontierScience SOTA benchmarks are public if you want to sanity check the claims. โ†’

Poonam Soni

27,512 gรถrรผntรผleme โ€ข 6 ay รถnce

The most interesting part for me is where Andrej Karpathy describes why LLMs aren't able to learn like humans. As you would expect, he comes up with a wonderfully evocative phrase to describe RL: โ€œsucking supervision bits through a straw.โ€ A single end reward gets broadcast across every token in a successful trajectory, upweighting even wrong or irrelevant turns that lead to the right answer. > โ€œHumans don't use reinforcement learning, as I've said before. I think they do something different. Reinforcement learning is a lot worse than the average person thinks. Reinforcement learning is terrible. It just so happens that everything that we had before is much worse.โ€ So what do humans do instead? > โ€œThe book Iโ€™m reading is a set of prompts for me to do synthetic data generation. It's by manipulating that information that you actually gain that knowledge. We have no equivalent of that with LLMs; they don't really do that.โ€ > โ€œI'd love to see during pretraining some kind of a stage where the model thinks through the material and tries to reconcile it with what it already knows. There's no equivalent of any of this. This is all research.โ€ Why canโ€™t we just add this training to LLMs today? > โ€œThere are very subtle, hard to understand reasons why it's not trivial. If I just give synthetic generation of the model thinking about a book, you look at it and you're like, 'This looks great. Why can't I train on it?' You could try, but the model will actually get much worse if you continue trying.โ€ > โ€œSay we have a chapter of a book and I ask an LLM to think about it. It will give you something that looks very reasonable. But if I ask it 10 times, you'll notice that all of them are the same.โ€ > โ€œYou're not getting the richness and the diversity and the entropy from these models as you would get from humans. How do you get synthetic data generation to work despite the collapse and while maintaining the entropy? It is a research problem.โ€ How do humans get around model collapse? > โ€œThese analogies are surprisingly good. Humans collapse during the course of their lives. Children haven't overfit yet. They will say stuff that will shock you. Because they're not yet collapsed. But we [adults] are collapsed. We end up revisiting the same thoughts, we end up saying more and more of the same stuff, the learning rates go down, the collapse continues to get worse, and then everything deteriorates.โ€ In fact, thereโ€™s an interesting paper arguing that dreaming evolved to assist generalization, and resist overfitting to daily learning - look up The Overfitted Brain by Erik Hoel. I asked Karpathy: Isnโ€™t it interesting that humans learn best at a part of their lives (childhood) whose actual details they completely forget, adults still learn really well but have terrible memory about the particulars of the things they read or watch, and LLMs can memorize arbitrary details about text that no human could but are currently pretty bad at generalization? > โ€œ[Fallible human memory] is a feature, not a bug, because it forces you to only learn the generalizable components. LLMs are distracted by all the memory that they have of the pre-trained documents. That's why when I talk about the cognitive core, I actually want to remove the memory. I'd love to have them have less memory so that they have to look things up and they only maintain the algorithms for thought, and the idea of an experiment, and all this cognitive glue for acting.โ€

Dwarkesh Patel

1,052,916 gรถrรผntรผleme โ€ข 11 ay รถnce