正在加载视频...
视频加载失败
🚀 Introducing Community Benchmarks on Kaggle! As AI evolves at an unprecedented pace, measuring intelligence requires more than a few AI research labs alone – it requires the imagination and collective expertise of the global community. That’s why we’re launching Community Benchmarks. It lets you build, run, and share... show more
54,677 次观看 • 8 个月前 •via X (Twitter)
8 条评论

With community benchmarks you get: ✨ Free access (within quota) to leading models 🧠 Multimodal, multi-step & tool-based tasks 📊 Compare model performance on a leaderboard 🛠 Powered by the kaggle-benchmarks SDK

this matters more than people realize. lab benchmarks test what researchers think matters. community benchmarks test what builders actually need. biggest gap i see: "can it follow complex multi-step instructions without hallucinating?" no benchmark for that. real builders run their own tests daily.

This is a strong step toward democratizing AI evaluation. Community driven benchmarks can surface real world tasks that academic suites miss and keep model progress grounded in practical impact. Looking forward to seeing how this evolves and what the community builds.

Community-driven benchmarks could surface edge cases that internal teams miss. Transparency in AI evaluation is overdue.

Is it possible to use a dataset with input and expected output while using this feature? If so, how to do that ?

Awesome.

@kaggle, the community benchmarks will be crucial as AI's pace of evolution accelerates, so it will be useful to understand how these will be.

Hey guys, help me explain to people what's wrong with LLMs if I'm right? I don't care about being wrong. I'm worried if I'm right, LLMs and Humans experience hallucinations in similar ways. Deviation in time from context and perspective. Help me find the plot for everyone else, please? Best Wishes, Aaron Bowen 🍀✨
