Loading video...
Video Failed to Load
We are releasing AutoResearchExam, a benchmark on open-ended machine learning and engineering tasks. Our benchmark covers seven research areas including model training, data curation, AI safety and interpretability. In each task, we give agents 24 hours with a CPU or GPU machine to develop and improve their solutions through... show more
1,762,414 views • 2 days ago •via X (Twitter)
0 Comments
No comments available
Comments from the original post will appear here
