PatronusAI's banner
PatronusAI's profile picture

PatronusAI

@PatronusAI2,310 subscribers

Simulation research and infrastructure for human-aligned AGI https://t.co/8X6bVgvCHd

Shorts

1/ 🔥🔥 Big news: We’re launching Percival, the first AI agent that can evaluate and fix other AI agents! 🤖 Percival is an evaluation agent that doesn’t just detect failures in agent traces — it can fix them. Percival outperformed SOTA LLMs by 2.9x on the TRAIL dataset, containing human annotated errors from GAIA and SWE-Bench. 🦾 Here’s what Percival can do for you: - Automatically suggest prompt fixes for your agent - Catch 20+ types of agent failures spanning tool use, planning and coordination, domain specific errors - Reduce manual debugging time from hours to < 1 minute

1/ 🔥🔥 Big news: We’re launching Percival, the first AI agent that can evaluate and fix other AI agents! 🤖 Percival is an evaluation agent that doesn’t just detect failures in agent traces — it can fix them. Percival outperformed SOTA LLMs by 2.9x on the TRAIL dataset, containing human annotated errors from GAIA and SWE-Bench. 🦾 Here’s what Percival can do for you: - Automatically suggest prompt fixes for your agent - Catch 20+ types of agent failures spanning tool use, planning and coordination, domain specific errors - Reduce manual debugging time from hours to < 1 minute

23,988 views

1/ Introducing Glider - the smallest model to beat GPT-4o-mini on eval tasks ⚡🚀 - Open source, open weights, open code - Explainable evaluations by nature - Trained on 183 criteria and 685 domains Try it out for free at 🔥

1/ Introducing Glider - the smallest model to beat GPT-4o-mini on eval tasks ⚡🚀 - Open source, open weights, open code - Explainable evaluations by nature - Trained on 183 criteria and 685 domains Try it out for free at 🔥

14,856 views

Videos

PatronusAI's profile picture

Today, we’re excited to announce our $50M Series B, led by Greenfield Partners, with participation from Lightspeed and Notable Capital. 🚀 At Patronus AI, we develop simulations and evals to train and improve AI. The first phase of AI was built on static benchmarks, but that era is over. As agents are used to solve longer and longer tasks, they need to practice in dynamic, living worlds to get better. Simulations are the critical infrastructure powering this next phase. As a company, we’re behind the most influential research and products in AI evaluation, like FinanceBench, Lynx, and Percival. And things have moved at the speed of light since.⚡ We partner with the world's leading frontier AI labs and enterprises, and our revenue has grown more than 15x over the past year. Additionally, today, we’re introducing a preview of the first Digital World Model for AI agent training and simulation: Patronus-DWM. Digital World Models are language diffusion world models that predict realistic environment behaviors and steer agent actions across digital workflows. Just as physical world models predict how objects move through space, we’re developing the equivalent for the digital world: predicting how agents act in digital workflows, then using that to scale the creation of high-quality training data for LLMs. Digital World Models help us push the frontier of ultra long horizon workflows, and unlock a new class of self-improving RL environments. This is our scalable approach to simulating all of the world’s intelligence. The round was also joined by Datadog, Inc., Samsung Ventures, Gokul Rajaram, Factorial Capital, and a large cohort of amazing AI leaders across Anthropic, OpenAI, Google DeepMind, NVIDIA, Recursive, and more.✨ It has been the ride of a lifetime. But we’re just getting started. The best is yet to come. "Do not go gentle into that good night, Rage, rage against the dying of the light" - Dylan Thomas (1954)

PatronusAI

94,014 views • 25 days ago

No more content to load