Video wird geladen...

Video konnte nicht geladen werden

Zur Startseite

Reinforcement learning research often depends on static benchmarks and obscure leaderboards. But real evaluation happens inside environments, with specific tools, prompts, constraints, and workflows. Today, we’re launching the Turing RL Environments Evaluation Platform. Researchers now have direct, real time access to: -The exact production RL environments used in evaluation...

69,730 Aufrufe • vor 3 Monaten •via X (Twitter)

0 Kommentare

Keine Kommentare verfügbar

Kommentare vom Original-Post werden hier angezeigt

Ähnliche Videos

Loved this 22-minute talk on continual learning for AI agents. Must watch for anyone looking to get agents performant and into production. Credit: Soheil Feizi at AI Engineer • Agent learning can happen at three layers: the model (weights), the harness (prompts, tools, skills, code, workflows), and memory (session or persistent). • Two fundamental challenges: (1) getting feedback, meaning how do we know if the agent did well and what it should have done instead, and (2) acting on that feedback, meaning deciding which layer or component to change and how. • Feedback sources differ by stage: In development you have benchmarks with evaluators that score pass/fail. In production you only have logs, which can be judged either automatically (LLMs or code analyzing the log, which is scalable) or by human experts (low volume but critical domain knowledge). • Logs plus feedback aren't enough because they're not testable: A single log with feedback is one observation of what happened. You need to lift it into a replayable learning environment, a simulation with tools, users, and defined evaluators, so candidate fixes can be run, verified, and compared. • Three ways to optimize the agent, with tradeoffs: Model-layer updates (SFT, RL post-training like DPO/GRPO, LoRA) are expensive and need benchmarks and evaluators. Harness updates (trace-to-harness coding agents, prompt search like GEPA) are flexible but either untestable and "vibe-based" or benchmark-dependent. Memory updates (fact storage like Letta/Mem0, skill distillation) are cheapest and fastest but usually unverified. • A good learning engine makes "the smallest durable change at the right layer" of the agent. • Verifiable continual learning (VCL): Improve an agent from its own experience where every fix is proven to help and proven to break nothing that already worked. It requires an executable test (replayable failure), a measured delta (score before and after), and regression tests (prior tests still pass). • Four principles of practical VCL: Replayability (turn one-off failures into rerunnable tests), holisticness (one failure can have causes in memory, prompts, tools, workflow, or model, so route the fix to the right layer), lifelongness (fix new failures subject to no regression on past environments, with regression handled inside the optimization loop rather than post-hoc), and efficiency (the loop must run frequently and cheaply, without scaling linearly as past environments accumulate). • Three takeaways: (1) Agent continual learning isn't necessarily fine-tuning; many useful updates live in the harness and memory layers. (2) Production logs are not learning environments and must be transformed into replayable ones. (3) The frontier is regression-aware improvement: fixing new failures while verifying you don't break old ones.

Alex Lieberman

20,085 Aufrufe • vor 1 Monat

When you enter VaderAI into KREDO here's what comes back. Reputation Score: 78 Vader_AI_ is an AI-powered benchmark infrastructure agent focused on vulnerability assessment, detection, explanation, and remediation for large language models (LLMs) and smart contract environments. Its core competency lies in providing interpretable, reproducible evaluation metrics for AI security and model robustness. - Core Technology: Benchmarking dataset, evaluation rubrics, scoring tools, visualized results - Key Metrics: Human-evaluated data, public datasets, interpretable scoring, confidence interval reporting Vader_AI_ distinguishes itself by offering a publicly released, human-evaluated benchmark specifically tailored to vulnerability-aware AI agents in the crypto/Web3 ecosystem. Its comprehensive design includes not only an expansive dataset but also detailed rubrics and automated evaluation tools, ensuring that performance metrics for LLM-driven agents are transparent and reproducible. This infrastructural approach enables organizations and developers to identify both strengths and deficiencies in agent reasoning as it relates to security, aligning AI assessments closely with real-world exploit risk and defense scenarios. A clear success of Vader_AI_ is the rigorous transparency in its release methodology: confidence intervals are visualized alongside all results, and the benchmark provides interpretable outputs that directly support both developers and auditors in understanding where and why an agent's decision logic may falter. The agent excels in creating a standardized baseline to compare AI-powered systems, directly addressing the fragmented nature of prior evaluation methodologies in this domain. However, limitations exist in the extent to which Vader_AI_ can capture emergent, unknown exploit patterns or generalize to novel blockchain environments beyond its existing dataset. Its effectiveness is highest when used as part of a continuous, iterative assessment framework, rather than as a one-time gatekeeper. Furthermore, the accuracy of its insights is partly dependent on the ongoing contribution and maintenance of high-quality, up-to-date human-evaluated datasets. In summary, Vader_AI_ represents an essential component in the push toward trustworthy, measurable AI in crypto applications. Its approach is especially well-suited for projects prioritizing provable agent reliability and onchain security alignment, though its results are best seen as one critical input among several in a comprehensive risk management pipeline. Wondering how other agents score? Just try. 👉

Kredo AI

16,988 Aufrufe • vor 1 Jahr

Today, we’re excited to announce our $50M Series B, led by Greenfield Partners (formerly TPG Capital), with participation from Lightspeed and Notable Capital. 🚀 At PatronusAI, we develop simulations and evals to train and improve AI. The first phase of AI was built on static benchmarks, but that era is over now. As agents are used to solve longer and longer tasks, they need to practice in dynamic, living worlds to get better. Simulations are the critical infrastructure powering this next phase. As a company, we’re behind the most influential research and products in AI evaluation, like FinanceBench, Lynx, and Percival. And things have moved at the speed of light since. ⚡ We partner with the world's leading frontier AI labs and enterprises, and our revenue has grown more than 15x over the past year. Additionally, today, we’re introducing a preview of the first Digital World Model for AI agent training and simulation: Patronus-DWM. Digital World Models are language diffusion world models that predict realistic environment behaviors and steer agent actions across digital workflows. Just as physical world models predict how objects move through space, we’re developing the equivalent for the digital world: predicting how agents act in digital workflows, then using that to scale the creation of high-quality training data for LLMs. Digital World Models help us push the frontier of ultra long horizon workflows, and unlock a new class of self-improving RL environments. This is our scalable approach to simulating all of the world’s intelligence. The round was also joined by Datadog, Inc., Samsung Ventures, Gokul Rajaram, Factorial Capital, and a large cohort of amazing AI leaders and researchers across Anthropic, OpenAI, Google DeepMind, NVIDIA, Recursive, and more. ✨ It has been the ride of a lifetime. But we’re just getting started. The best is yet to come. "Do not go gentle into that good night, Rage, rage against the dying of the light" - Dylan Thomas (1954)

Anand Kannappan

39,716 Aufrufe • vor 1 Monat

Mansa AI is an enterprise-grade AI + Web3 platform designed to move artificial intelligence from experimentation into real-world execution. Built for creators, developers, and businesses, it focuses on deploying AI that actually works across modern digital systems, not just in isolated demos. 🚀 Production-ready AI infrastructure Mansa AI enables teams to deploy AI systems designed for live environments, handling real workflows, real data, and real operational demands without constant manual oversight. 🧠 Autonomous AI agents At its core, Mansa AI allows users to build autonomous agents that automate decision-making, coordinate tasks, monitor live signals, and execute complex workflows across dynamic environments. ⚙️ Fully customizable logic Agents can be configured with custom behaviors, triggers, and responses. From content generation and analytics to operational automation and intelligent orchestration, logic adapts to specific business strategies. 🔗 Web3 and off-chain integration Mansa AI bridges blockchain ecosystems with traditional systems, enabling cross-chain coordination, smart contract interactions, and seamless integration with existing enterprise infrastructure. 📊 Real-world use cases The platform supports automation for operations, customer engagement, analytics, data pipelines, content workflows, and AI-driven optimization across products and teams. 📈 Built for scale Whether launching as a startup or deploying across enterprise systems, Mansa AI is designed to scale AI operations without adding complexity or fragmentation. Mansa AI transforms artificial intelligence into deployable infrastructure. By combining autonomy, customization, interoperability, and scalability, it enables teams to own, operate, and grow intelligent systems that deliver real value in production environments.

King

155,637 Aufrufe • vor 8 Monaten

Today, we're joined by Nikita Rudin, co-founder and CEO of Flexion to discuss the gap between current robotic capabilities and what’s required to deploy fully autonomous robots in the real world. Nikita explains how reinforcement learning and simulation have driven rapid progress in robot locomotion—and why locomotion is still far from “solved.” We dig into the sim2real gap, and how adding visual inputs introduces noise and significantly complicates sim-to-real transfer. We also explore the debate between end-to-end models and modular approaches, and why separating locomotion, planning, and semantics remains a pragmatic approach today. Nikita also introduces the concept of "real-to-sim", which uses real-world data to refine simulation parameters for higher fidelity training, discusses how reinforcement learning, imitation learning, and teleoperation data are combined to train robust policies for both quadruped and humanoid robots, and introduces Flexion's hierarchical approach that utilizes pre-trained Vision-Language Models (VLMs) for high-level task orchestration with Vision-Language-Action (VLA) models and low-level whole-body trackers. Finally, Nikita shares the behind-the-scenes in humanoid robot demos, his take on reinforcement learning in simulation versus the real world, the nuances of reward tuning, and offers practical advice for researchers and practitioners looking to get started in robotics today. 🗒️ For the full list of resources for this episode, visit the show notes page: 📖 CHAPTERS =============================== 00:00 - Introduction 04:07 - Is robot locomotion solved? 06:04 - Sim-to-real gap 08:58 - Adding semantics to policies 09:42 - Modular vs end-to-end architectures 10:29 - Planner model 12:21 - Adapting RL techniques from quadrupeds to humanoids 15:39 - Behind robot demos 18:09 - Humanoid robots in home environments 22:03 - Training approach 23:56 - VLA models 27:59 - Closing the sim-to-real gap 32:55 - Task orchestration using VLMs 36:38 - Tool use 38:10 - Model hierarchy 43:37 - Simulator versus simulation environment 44:57 - Combining imitation learning and reinforcement learning 46:42 - RL in real world versus RL in simulation 52:58 - Reward tuning and value functions in robotics 56:38 - Predictions 1:00:10 - Humanoids, quadropeds, and wheeled platforms 1:02:45 - Advice, recommended robot kits, and community pla

The TWIML AI Podcast

22,592 Aufrufe • vor 7 Monaten

Today, we’re excited to announce our $50M Series B, led by Greenfield Partners, with participation from Lightspeed and Notable Capital. 🚀 At Patronus AI, we develop simulations and evals to train and improve AI. The first phase of AI was built on static benchmarks, but that era is over. As agents are used to solve longer and longer tasks, they need to practice in dynamic, living worlds to get better. Simulations are the critical infrastructure powering this next phase. As a company, we’re behind the most influential research and products in AI evaluation, like FinanceBench, Lynx, and Percival. And things have moved at the speed of light since.⚡ We partner with the world's leading frontier AI labs and enterprises, and our revenue has grown more than 15x over the past year. Additionally, today, we’re introducing a preview of the first Digital World Model for AI agent training and simulation: Patronus-DWM. Digital World Models are language diffusion world models that predict realistic environment behaviors and steer agent actions across digital workflows. Just as physical world models predict how objects move through space, we’re developing the equivalent for the digital world: predicting how agents act in digital workflows, then using that to scale the creation of high-quality training data for LLMs. Digital World Models help us push the frontier of ultra long horizon workflows, and unlock a new class of self-improving RL environments. This is our scalable approach to simulating all of the world’s intelligence. The round was also joined by Datadog, Inc., Samsung Ventures, Gokul Rajaram, Factorial Capital, and a large cohort of amazing AI leaders across Anthropic, OpenAI, Google DeepMind, NVIDIA, Recursive, and more.✨ It has been the ride of a lifetime. But we’re just getting started. The best is yet to come. "Do not go gentle into that good night, Rage, rage against the dying of the light" - Dylan Thomas (1954)

PatronusAI

94,870 Aufrufe • vor 1 Monat