
Anand Kannappan
@anandnk24 • 1,760 subscribers
co-founder and ceo @PatronusAI (series B). prev AI @Meta evals, RL, data ⚡️
Shorts
Videos

Today, we’re releasing SpeedrunBench: the first benchmark that measures how fast agents can beat video games. We asked frontier models not to simply complete video games, but to speedrun them - across 10 titles like Mario Kart and Pokemon. On simpler games, agents get close to the human world record. However, on more complicated games like Pokemon Blue, the best models Kimi-K3 and Opus 5 are ~4x off the world record. Benchmark, paper, and demo below. PatronusAI
Anand Kannappan142,594 Aufrufe • vor 1 Monat

Today, we’re excited to announce our $50M Series B, led by Greenfield Partners (formerly TPG Capital), with participation from Lightspeed and Notable Capital. 🚀 At PatronusAI, we develop simulations and evals to train and improve AI. The first phase of AI was built on static benchmarks, but that era is over now. As agents are used to solve longer and longer tasks, they need to practice in dynamic, living worlds to get better. Simulations are the critical infrastructure powering this next phase. As a company, we’re behind the most influential research and products in AI evaluation, like FinanceBench, Lynx, and Percival. And things have moved at the speed of light since. ⚡ We partner with the world's leading frontier AI labs and enterprises, and our revenue has grown more than 15x over the past year. Additionally, today, we’re introducing a preview of the first Digital World Model for AI agent training and simulation: Patronus-DWM. Digital World Models are language diffusion world models that predict realistic environment behaviors and steer agent actions across digital workflows. Just as physical world models predict how objects move through space, we’re developing the equivalent for the digital world: predicting how agents act in digital workflows, then using that to scale the creation of high-quality training data for LLMs. Digital World Models help us push the frontier of ultra long horizon workflows, and unlock a new class of self-improving RL environments. This is our scalable approach to simulating all of the world’s intelligence. The round was also joined by Datadog, Inc., Samsung Ventures, Gokul Rajaram, Factorial Capital, and a large cohort of amazing AI leaders and researchers across Anthropic, OpenAI, Google DeepMind, NVIDIA, Recursive, and more. ✨ It has been the ride of a lifetime. But we’re just getting started. The best is yet to come. "Do not go gentle into that good night, Rage, rage against the dying of the light" - Dylan Thomas (1954)
Anand Kannappan45,854 Aufrufe • vor 3 Monaten
Keine weiteren Inhalte verfügbar