Loading video...
Video Failed to Load
How well do today’s frontier models handle long-horizon, multi-step web agent tasks, such as identifying the top 25 U.S. CS PhD programs with ML/AI faculty likely accepting students and compiling the results into a structured sheet? Check out our new work on Odysseys: Benchmarking Web Agents on Realistic Long... show more
22,518 views • 3 months ago •via X (Twitter)
0 Comments
No comments available
Comments from the original post will appear here

