Video wird geladen...
Video konnte nicht geladen werden
Introducing the Open Axis Benchmark, a living benchmark engine for honest evaluation of robot manipulation models, built together with OpenRoboto. Robotic models evolve faster every month, while most benchmarks stay frozen. Models overfit to fixed task sets, scores stop reflecting real generalization, and that distorted signal misleads and holds... show more
27,911 Aufrufe • vor 3 Tagen •via X (Twitter)
35 Kommentare

@openroboto

@openroboto finally a benchmark that actually keeps up with model progress

@openroboto That’s the way to do it! 🎯 The Axis to Jumper pipeline is looking strong. Keep riding that alpha wave! 🌊

@openroboto Fresh task sets make robot evaluation much more meaningful

@openroboto You're doing well

@openroboto Dope addition tbh.

@openroboto gAxis

@openroboto interesting

@openroboto gAxis

@openroboto time to explore

@openroboto bullish Axis 🙌🙌

@openroboto That's huge

@openroboto up up up

@openroboto Been grinding Axis robotics task and I would probably say that Libero tasks are challenging and good for learning the mastery in controlling the robot arms

@openroboto let’s explore

@openroboto this is amazing

@openroboto old benchmarks jjust train robots to memorize

@openroboto Good to know that robotics models evolve faster every month

@openroboto this is bullish! LFAxis!

@openroboto Axis on fire

@openroboto a living eval for robot hands just shipped. the scoreboard finally has a home

@openroboto Wow

@openroboto gAxis

@openroboto gAxis

@openroboto optimize for specific task parameters rather than generalized spatial reasoning

@openroboto LFG!!!

@openroboto Each task will have a reward pool paid in USD (so the token doesn't get inflated). Whoever performs the tasks better will earn more dollars. We need to make sure the reward pool still offers a profitable ratio so that people stay excited to join.

@openroboto impressive work, countless hours fueling progress

@openroboto Sobrang dami nating matututunan dito! Ang lakas talaga ni @openroboto!

@openroboto honestly fresh task sets are the only way to know it actually generalizes

@openroboto frozen task sets get gamed so fast tho

@openroboto fresh tasks make benchmarks way more meaningful

@openroboto axis bigger picture is here...

@openroboto gAxis

@openroboto Continuous evaluation could make robotic progress more measurable
