Video wird geladen...
Video konnte nicht geladen werden
Can Jev Nav? We put Jev, Astra, Fable, Dimcode, and other models + harnesses to the test on 2000+ navigation tasks in 133 real and simulated environments How far are LMs from zero-shotting complex real-time control tasks? Full dataset and code open source ⤵️
20,068 Aufrufe • vor 11 Tagen •via X (Twitter)
12 Kommentare

2/ Agents were provided with environments in JSON/string format, as Jev can only receive language as input. We use the below method to convert lidar and XML scenegraphs to WorldState, which is a string representation of the environment, objects, and obstacles.

3/ Despite Jev inference at 2 times per second it uses limited context. GPT 5.6 has the highest acceleration throughout navigation runs. Dots mark run completion. Astra and Fable context growth is constant, but still massive. On avg 100k tokens for a single short nav task.

4/ Jev outperforms Fable on shorter paths, and wins dramatically on speed and cost Jev costs $0.081 per run versus Astra at $1.610 per run and Fable at $0.963 per run

5/ Initial results ranked by SPL (a nav trajectory quality metric)

6/ Full paper, results, and dataset navigator here!

@dimensionalos Nav arena looks so cool.

@dimensionalos Cool stuff!

@dimensionalos

@dimensionalos Real vs simulated split is the useful part of this benchmark, most navigation evals stay purely sim and quietly assume the gap closes on its own. Curious how much the numbers diverge between the two environment types once you break it down that way.

@dimensionalos Is there more info on the WorldState concept? How do you determine what’s included at any given frame?

@dimensionalos TLDR we look at free space, close objects, obstacles, cartesian distance to target object. And for each keyed object we include distance in meters, orientation (bearing out of 360), and relative position in language (ahead, back left, etc.)

@dimensionalos How did you come up with this method? Is there some related work / papers etc?
