Video wird geladen...
Video konnte nicht geladen werden
Earlier today we released local development for Kaggle Benchmarks. 🚀 You can now write, validate and run AI evaluation tasks directly from your preferred dev environment — VSCode, Antigravity, Claude Code, and more. Go from idea to working eval using natural language with the write-kaggle-benchmarks skill.
16,790 Aufrufe • vor 4 Monaten •via X (Twitter)
9 Kommentare

Drop the skill into your agent to get started 👇

nice :)

oh i like this!

This is a big step for AI evaluation workflows. Bringing benchmark creation and validation directly into developers' existing environments removes friction and makes experimentation much faster. Natural language → working evals is exactly the kind of productivity unlock teams need. 🚀

Local eval dev matters because the expensive part is not trusting a leaderboard. It is rerunning the exact task after every prompt, model, or tool change. If the check lives next to the code, teams may actually use it.

This is important because beginners need a way to test AI outputs, not just trust them. Even simple evals change the workflow: write the expected behavior, let AI build, run the check, fix what fails. That habit makes vibe coding a lot less random.

本地跑评估终于不用来回切页面了

I happened to come across this update by chance, although I last participated on Kaggle about 7 years ago. Could you tell me what these Kaggle benchmarks are for, after all?

not sure 'idea to working eval' is the bottleneck tbh. when i tried writing evals for pennywise, setup took an hour but figuring out what 'correct' even means took days. does local dev actually help with the metric definition side or mostly just the scaffolding?
