Загрузка видео...
Не удалось загрузить видео
Introducing Nav Arena, over 150 real and simulated environments to benchmark Agents on robots As a first experiment, we put Jev, Astra, Fable, Dimcode, and other models + harnesses to the test on 2000+ navigation tasks Fully open source! Code, data, and results below
16,456 просмотров • 9 дней назад •via X (Twitter)
Комментарии: 14

2/ Despite Jev inference at 2 times per second it uses limited context. GPT 5.6 has the highest acceleration throughout navigation runs. Dots mark run completion. Astra and Fable context growth is constant, but still massive. On avg 100k tokens for a single short nav task.

3/ Jev outperforms Fable on shorter paths, and wins dramatically on speed and cost Jev costs $0.081 per run versus Astra at $1.610 per run and Fable at $0.963 per run

4/ Paper & complete results:

@dimensionalos nice! open source yay

@dimensionalos This is really interesting robotics benchmark. Time to start digging into @dimensionalos

@dimensionalos 这个仿真环境里面做的是轨迹数据吗,我正在做一个轨迹预测的模型我认为这些世界模型的新方向。

@dimensionalos It’s possible to make a custom environment ?

@dimensionalos Reporting time and cost alongside navigation results makes this easier to use. Cost per completed route would also help show whether cheap failed runs change the comparison.

@dimensionalos Thanks for this 👍👌

@dimensionalos what’s the setup for deploying models/agents in the arena?

@dimensionalos very nice! but you need to post a longer/better video!

@dimensionalos twitter attention spans are max 5 seconds but yes we should as longer form

@dimensionalos well there's nothing there in the 10 secs. it's 'teaser'. the camera's angled where you can't see other environments so you wasted the 10 seconds you had, you would have done better to post 5 2-second sped up clips of 5 different environments since it's the point. but I love it

@dimensionalos you can view all 2000+ runs here in our nav data explorer:
