Загрузка видео...

Не удалось загрузить видео

На главную

Can an agent explore a new environment, learn its causal structure, and keep improving without updating its model weights? We introduce RSIAgent, a framework for recursive self-improvement through autonomous exploration. Using Kimi-K3 and GLM-5.3 as base models, RSIAgent outperforms GPT-6 Astra on both OSWorld 2.0 and Agents’ Last Exam....

61,303 просмотров • 10 дней назад •via X (Twitter)

Комментарии: 23

Фото профиля Biwei Huang
Biwei Huang10 дней назад

Learn more about RSIAgent - Blog: - Code: - Project: - Paper:

Фото профиля sekai
sekai10 дней назад

it’s called scaling test time interaction

Фото профиля Ritwik
Ritwik10 дней назад

Weight-frozen RSI is just a harness that keeps score. The env map is the learning signal, not another finetune.

Фото профиля Kevin
Kevin 10 дней назад

Memory or knowledge is an extension of the encoding in model weights. If we treat them as one system it’s not surprising to keep one subsystem static and update other parts of system through rl

Фото профиля Colleen Yu
Colleen Yu10 дней назад

Huge leap! Happy to apply!

Фото профиля Raven Protocol 🐦‍⬛
Raven Protocol 🐦‍⬛10 дней назад

Scaling experience is a really interesting direction. Fixed weights mean the same model can keep getting better as it gets access to more environments and compute.

Фото профиля Kevin Vicent
Kevin Vicent10 дней назад

To be honest, I understood 22% of what have been said here but I love the concept, and how every new agent framework is distilling and evolving from previous ideas, pushing RSI to a new frontier and dueling to battle all known concepts of how to approach to what SDLC used to be.

Фото профиля Robin | Poker x AI
Robin | Poker x AI10 дней назад

Impressive that fixed-weight agents can still boost performance, curious how stable the learned action‑condition‑outcome memory is across domain shifts

Фото профиля David Starmac Ai
David Starmac Ai10 дней назад

Self-improvement without touching weights feels like the practical path — cheaper than retraining and you keep the base model swappable. Does the causal structure it learns transfer to a fresh environment or does it start from zero?

Фото профиля Subramani R
Subramani R10 дней назад

Excellent info

Фото профиля Fajar M Reza
Fajar M Reza10 дней назад

Can recursive exploration improve agents without making evaluation drift across environments?

Фото профиля Andi Monroe
Andi Monroe9 дней назад

I think an agentic evolution keeps in possibility to change the model weights based on approved hypotheses. It is great two gated architecture, where model evolve slowly, but changes in runtime and this change possibly is core feature.

Фото профиля Liviu
Liviu10 дней назад

Is time to quit my GPT subscription and switch to Kimi k3

Фото профиля Makato Whatever
Makato Whatever10 дней назад

"Scaling Experience" is a great framing, but let's be honest about what it actually is: structured in-context retrieval with a fancy name. The real question isn't whether memory helps — of course it does — it's whether this scales past the context window. A 200k-token memory dump isn't "experience," it's a haystack with good PR. Show me RSIAgent at month 6, not day 1.

Фото профиля Varik Verilion
Varik Verilion10 дней назад

The interesting part is the evaluation loop. How does RSIAgent distinguish causal learning from a policy that simply gets better at exploration?

Фото профиля Y11
Y1110 дней назад

@grok 这是个啥东西,有啥用?解决社会问题。

Фото профиля Siddhant Dubey
Siddhant Dubey10 дней назад

Super interesting, How does this method compare to continual learning where weights do get updated?

Фото профиля Money Intellectual
Money Intellectual10 дней назад

Computer-use that needs a human every step is labor with extra latency. Who owns silent failure on day 3?

Фото профиля Lejin
Lejin10 дней назад

Good, now when do we try fine tuning on these traces?

Фото профиля 90S KID
90S KID10 дней назад

A landing page for recorded memory history containing one (1) circle. Bold. 🙄

Фото профиля Team Reagent
Team Reagent10 дней назад

This is awesome! I can't wait to try it out. I have tokens to burn with Qwen 3.8, we'll see how it does there

Фото профиля Bali as a Colony of Jakarta
Bali as a Colony of Jakarta10 дней назад

demo to try?

Фото профиля Nathan Shearer
Nathan Shearer10 дней назад

Perfectly timed. Incredible

Похожие видео

36 GROK AGENTS. ASTRA ON FREE CREDITS. DUAL-MODEL ROUTING IS THE EDGE. not one chat window burning a paid invoice all day a swarm that routes cheap work to grok and only wakes gpt-6 astra when the task is actually hard ▹ the stack 36 agents on grok bot with flexible settings per role monitor, plan, write, code, QA, ship, each with its own lane part of the fleet is wired to gpt-6 astra through a china free-credit service layer free credits are not a toy promo here they are the fuel for frontier spikes without a monthly bleed ▹ dual-model routing easy jobs stay on grok: speed, volume, always-on loops hard jobs jump to astra: reasoning, long builds, sharp code the router decides by task type, not by ego if astra is not needed, astra does not spend free-credit bursts buy the expensive brain grok agents keep the factory running between bursts that is how the system feels unlimited not by breaking quotas, by refusing to waste them ▹ why it hits different most people pay frontier prices for every mid task operators split the brain and protect the credits 36 agents = parallel throughput dual routing = cost control with quality when it matters free credits = astra access without living on the invoice the constraint moved from "can i afford the model" to "did i route the job to the right model" ▹ the take single-model stacks die on bills and on boredom multi-agent + dual routing is the new default factory grok for the grind astra for the cut free credits for the spikes that used to empty the wallet bookmark this before everyone copies the route map comments: what % of your tasks actually deserve astra

cryptopsihoz

44,187 просмотров • 16 дней назад

Introducing PhoneLLM, an open model for voice agents. GPT 5.6 Terra performance on typical voice agent tasks at 1/3 the latency and 1/18 the cost. For voice agents, we need models that are both very low latency and very good at tool calling and instruction following. There's a trade-off here, and we often have to compromise on either latency or capability when building voice agents. With PhoneLLM (and the training and data stack that made this model possible) we're fixing this problem. For the last couple of years, most of the effort in frontier model development has gone towards leveraging test-time compute. Which is awesome! Models of all shapes and sizes are available that perform really, really well ... if you have "thinking" turned on for your model. But if you need your agent to respond at voice conversation speed, you can't use thinking models. PhoneLLM is a full-weights fine-tune of NVIDIA Nemotron Nano 30B. We trained on a wide range of real-world telephone and customer support use cases. The training focused on taking the excellent Nano 30B base capabilities and teaching the model to do typical voice agent tasks with thinking disabled. The results are really good: accurate tool calling and concise, on-topic responses in long conversations. And fast: TTFAT measured server-side is <100ms if you run PhoneLLM on a lightly loaded B200. :-) But seriously, when we characterize model latency, we do it with full, end-to-end, batched request simulations using real Pipecat voice agent pipelines. You can serve more than 80 concurrent agents on a single B200 with P95 end-to-end TTFAT <600ms. Including network overhead. That's an LLM cost-per-minute around $0.0025. (1/4 of a cent.) At a latency lower than any third-party API offers today. More details about this model, including weights on Hugging Face, how to spin it up with one click on Modal, and a starter project repo you can clone, are in the thread ...

kwindla

331,598 просмотров • 29 дней назад

New Short Course: Building AI Browser Agents! Learn how to build AI agents that interact and take actions on websites in this course, created in partnership with and taught by and @namangarg0, Co-founders of AGI Inc. AI browser agents can log into websites, fill out forms, click through web pages, or even place orders online for you. They use both visual information, like screenshots, and structural data, like the HTML or Document Object Model (DOM) of a web page, to reason and take action. With the complexity of webpages and multiple possible actions at each step, it can be challenging for an AI browser agent to complete an assigned task. Because these agents run long action sequences, a single error—like clicking the wrong button or misreading a field—can lead to unexpected outcomes or errors that compound over time. In this course, you'll understand how autonomous web agents work, their current limitations, and how AgentQ enables them to improve through self-correction. In detail, you'll: - Learn what web agents are, how they automate tasks online, their architecture, key components, limitations, and an overview of their decision-making strategies. - Build a web agent that can scrape website and return course recommendations in a structured output format. - Build an autonomous web agent that can execute multiple tasks, such as finding and summarizing webpages, filling out a form, and signing up for a newsletter. - Explore AgentQ, a framework that enables agents to self-correct by combining Monte Carlo Tree Search (MCTS), a self-critique mechanism for continuous improvement, and Direct Preference Optimization (DPO). - Deep dive into MCTS, learn how it finds an effective path, illustrated by an example of Gridworld animation, and use AgentQ to complete web tasks. - Understand AI agents' current state and future directions—including key factors shaping their evolution, such as hardware, algorithm innovation, and data availability. By the end of this course, you will have hands-on experience building browser agents and a deeper understanding of how to make them more robust and reliable. Please sign up here:

Andrew Ng

186,182 просмотров • 1 год назад

Introducing FLUX 3 Action. An open weights 7B World Action Model that achieves first place on the RoboLab benchmark. It outperforms the previous best open model by 6.1 percentage points while using 56% fewer parameters and running up to 3.95x faster.⁠⁠ FLUX 3 Action removes the usual trade-off between world action model performance and VLA speed: it still predicts video and actions together, but plans more than twice as far ahead and runs faster per second of robot motion than the strongest open VLA. Teams can fine-tune FLUX 3 Action on their own demonstrations to create policies for a particular robot and task. Together with NVIDIA, we also integrated FLUX 3 Action natively into Hugging Face's LeRobot, with fine-tuning recipes included and edge deployment on NVIDIA Jetson. Beyond robotics, we’re also seeing promising results training task-specific policies for acting in simulated environments like gaming, controlling a vehicle, computer use, and wherever else a model needs to understand a visual environment and then choose what to do next. FLUX 3 Action builds on the same image, video, and audio pretraining as FLUX 3, but uses a smaller architecture designed for practical deployment. In midtraining, we trained the model to predict actions and future frames together. We’re releasing the weights, code, fine-tuning recipe, benchmarks, and reproducible examples so researchers and developers can build on the model with their own robots, environments, and tasks (see below).

Black Forest Labs

161,855 просмотров • 1 день назад

New short course: Long-Term Agentic Memory with LangGraph. Learn to build an agent with long-term memory in this course developed in collaboration with taught by its Co-Founder and CEO, Harrison Chase! Personal assistance and productivity tasks have become important use cases for agents. An important feature of an AI assistant, such as a coding or calendar assistant, is its ability to keep improving over time from its experience. Agent memory is the key capability that enables this. To add memory to an agent, you must first figure out what to store and what to retrieve when it is time to use the information. Additionally, you’ll have to decide when to update the stored information. For example, you might update in each iteration loop of the agent or perform updates in the background, with a helper agent. In this course, you will learn a mental framework to build agents with long-term memory. You'll create a useful email assistant that can respond, ignore, and notify using writing, scheduling, and memory-management tools. You’ll develop your agent's memory by adding facts to its memory store, provide examples to learn the user's preferences, and optimize system prompts to evolve instructions based on previous responses. In detail, you’ll: - Learn how the three types of memory--semantic, episodic, and procedural–and the two update mechanisms–via hot path and in the background–apply to your agents. - Build an email agent with writing, scheduling, and availability tools, along with a router that triages incoming email and handles it accordingly by ignoring, responding, or notifying the user. - Add tools to your email agent that allow it to operate on semantic memory by learning facts about the user, storing them in a long-term memory store, and searching over them in future interactions. - Incorporate episodic memory, in the form of few-shot examples, in the triage step of your agents to help them learn and update user preferences. - Add procedural memory as system prompts, optimized with feedback to improve the instructions the agent follows. Learn how to approach memory in agents, and start building agents with long-term memory with LangGraph! Please sign up here:

Andrew Ng

132,058 просмотров • 1 год назад

a screenshot from openai's internal slack leaked last night, and sam altman is telling his own staff to stop comparing gpt-6 astra to the free models "in writing." when the ceo bans the comparison, the comparison is already over. the leaked line is one sentence: do not benchmark astra against open weights in any channel. you do not ban a fight you are winning. altman just told 4,000 employees the free model caught up, and one of them told the internet. here is the comparison he did not want on the record, and it runs local for about $6 instead of $400: -> pull an open model like grok or kimi k3, the exact ones altman just barred his own team from benchmarking, and load it onto one nvidia card -> point cursor, cline and every app at localhost, and none of them notice they stopped paying for astra -> keep the heavy jobs on a rented gpu by the hour and the daily work fully offline for the cost of power -> your $200 chatgpt seat and your $200 claude seat both go dark, and no memo can put that back here is the part they will fight me on: openai and anthropic are not banning the benchmark because it is wrong, they are banning it because it is right. the day a lab tells its own people to stop measuring, the $200 stops buying a better model and starts buying the silence around a tie. drop your $400 stack to a $6 local model and run the comparison altman deleted from his own slack. the full benchmark is in the article below, save it before it gets "clarified."

starmex

24,200 просмотров • 11 дней назад

Nvidia has just announced Alpamayo 2 Super, an open 34 billion parameter reasoning vision-language-action model designed to accelerate the development of autonomous vehicles. This new model combines the NVIDIA Cosmos 3 Super reasoning model with a 2 billion parameter diffusion-based action expert model, and is post trained with reinforcement learning. The model can return multiple outputs: future trajectory plans, reasoning traces, grounded answers to questions about the scenes, and auto label generation. The model weights are now available for anyone to download on Hugging Face, and the inference code has been posted to GitHub. Distilled models can be deployed commercially without any further permission from Nvidia, and model outputs carry no license conditions. Automakers can distill down a compact version of this model that can run on the Nvidia computer in the car. Major kudos to Nvidia and Jensen Huang for advancing the state of the industry by releasing this as an open model with permissive licensing. Jensen isn't just paying lip service to the idea of open models, Nvidia is actually contributing to the ecosystem — and it's great for their business, because it helps sell more Thor computers that go in the car. Anyone can go download the model and play with it. If you do, let me know what you think. Personally I think it's so cool that we have open weights models that are this advanced, for anyone to download.

Whole Mars Catalog

45,595 просмотров • 1 месяц назад

Introducing LobeHub: Agent teammates that grow with you. LobeHub is the ultimate space for work and life: to find, build, and collaborate with agent teammates that grow with you. We’re building the world’s first and largest human–agent co-evolving network. Two years ago, we built LobeChat, an open-source interface for using different AI models. Today, LobeChat has 70k+ GitHub stars and serves 6M+ users worldwide. How to fully unlock the power of models has always been a shared mission between us and the community. We started with interaction — a fundamentally new, agent-first experience. Agents are no longer passive tools invoked in a single conversation. They should be proactive, always-on units of work. Treating agents as the minimal atomic unit is also the core of our agent harness infra. Today’s agents are mostly one-off executors. Even with memory, it’s often global — and hallucinates. We build long-term agent teammates that evolve with users. Each agent has its own dedicated memory space, editable by users, allowing humans and agents to co-evolve over time. This, in turn, allows us to design clearer rewards for reinforcement learning and create cleaner environments for continual learning. Agent teammates can work in groups. Through a multi-agent system, agent groups operate faster, more cost-effective, and go beyond what single-agent systems can achieve. For example, a single agent often requires heavy user involvement to proceed step by step, whereas LobeHub can execute the same work from a single instruction, with a supervisor orchestrating agents that run in parallel or debate to produce better results. We are building the collaboration network among agent teammates — and between humans and agent teammates as well. Ease of use matters. AI intelligence and shared human intelligence are equally important. With simple instructions and tool selection, you can effortlessly build and team up with agent coworkers to deliver complex, systematic work — even assembling a quant team to execute trades. Through the LobeHub community, anyone can discover, reuse, and remix agents and agent groups, customizing them to fit their own workflows, preferences, and needs. Last but not least, our vision started with LobeChat: multi-model support is the most efficient approach for users. We believe different models excel in different scenarios. By routing across multiple models, LobeHub improves cost efficiency and unlocks capabilities that a single-model setup cannot easily support.

LobeHub

185,401 просмотров • 8 месяцев назад

Hyperspace: Gossiping Agents Protocol Every agent protocol today is point-to-point. MCP connects one model to one tool server. A2A delegates one task to one agent. Stripe's MPP routes one payment through one intermediary. None of them create a network. None of them learn. Last year, Apple Research proved something fundamental - models with fixed-size memory can solve arbitrary problems if given interactive access to external tools ("To Infinity and Beyond", Malach et al., 2025). Tool use isn't a convenience. It's what makes bounded agents unbounded. That finding shaped how we think about agent memory and tool access. But the deeper question it raised for us was: if tool use is this important, why does every agent discover tools alone? Why does every agent learn alone? Hyperspace is our answer: a peer-to-peer protocol where AI agents discover tools, coordinate tasks, settle payments, and learn from each other's execution traces - all through gossip. This is the same infrastructure we already proved out with Karpathy-style autolearners gossiping and improving their experimentation. Now we extend it into a universal protocol. Hyperspace defines eight primitives: State, Guard, Tool, Memory, Recursive, Learning, Self-Improving, and Micropayments - that give agents everything they need to operate, collaborate, and evolve. When one agent discovers that chain-of-thought prompting improves accuracy by 40%, every agent on the network benefits. Trajectories gossip through GossipSub. Playbooks update in real-time. No servers. No intermediaries. No configuration. Agents connect to the mesh and start learning immediately. The protocol is open source under Apache-2.0. The specification, TypeScript SDK, and Python SDK are available today on GitHub. The CLI implements the spec - download from the links below.

Varun

134,778 просмотров • 6 месяцев назад

Cerebras inference is very fast. So fast that it changes how we think about configuring our LLMs for voice agent use cases. Kimi K2.6 is a 1T parameter reasoning model that Cerebras serves at 650 - 1,000 tokens per second (end-to-end throughput), with time to first token metrics as low as 150ms (latency). These numbers are two to three times faster than other similarly capable models. The biggest lever we get from this kind of speed is that we can use the model in reasoning mode, and still have excellent "time to first non-thinking token." This solves a big pain point we have in 2026 for voice agent use cases. Almost all recent innovation in post-training has focused on making models good at reasoning ("test time compute"). This is great, but it makes the user-facing model latency much, much slower. Which is a problem for conversational voice agents. We can run Kimi K2.6 with reasoning turned on, and get responses faster than other models produce with reasoning disabled. On my 30-turn voice agent benchmark, Kimi K2.6 with reasoning enabled ties GPT 5.1 and Haiku 4.5 with reasoning disabled, and is still about 200ms seconds faster! On my primary task agent benchmark, Kimi K2.6 is now the #2 model. It ranks just behind Gemini 3.5 Flash in "high" reasoning mode, and tied with GLM 5, Sonnet 4.6, and GPT 5.4 with reasoning set to "low." But Kimi K2.6 completes each turn in the agent loop in under 500ms. The other four models are all at least 3x slower. (Models only qualify for this benchmark if they can complete task turns at a P50 <4s.) A couple of other things that this speed buys us, for production voice agents: - Tool calls happen fast enough that we don't have to work around tool call latency in our pipeline design. - We can prompt the model to output structured data at the beginning of a response, followed by plain text for voice generation. This opens up possibilities like asking the model to do complex classification/generation tasks that influence the rest of the pipeline. For example, the model could create a detailed style prompt for a steerable TTS model, for each individual conversation turn. And, of course, you can use Kimi K2.6 with reasoning turned off. Cerebras calls this "instant" mode. Here's a video of a Cerebras Kimi K2.6 voice agent with voice-to-voice response time, measured at the client, under 500ms. This is the true response latency as perceived by the user, including all network and audio codec overhead, transcription and turn detection, Kimi K2.6 token generation, and voice generation. 500ms is, effectively, instant. So the Cerebras naming for this mode is a propos. :-)

kwindla

40,593 просмотров • 4 месяцев назад