Loading video...

Video Failed to Load

Go Home

Can Jev Nav? We put Jev, Astra, Fable, Dimcode, and other models + harnesses to the test on 2000+ navigation tasks in 133 real and simulated environments How far are LMs from zero-shotting complex real-time control tasks? Full dataset and code open source ⤵️

20,068 views • 11 days ago •via X (Twitter)

12 Comments

stash's profile picture
stash11 days ago

2/ Agents were provided with environments in JSON/string format, as Jev can only receive language as input. We use the below method to convert lidar and XML scenegraphs to WorldState, which is a string representation of the environment, objects, and obstacles.

stash's profile picture
stash11 days ago

3/ Despite Jev inference at 2 times per second it uses limited context. GPT 5.6 has the highest acceleration throughout navigation runs. Dots mark run completion. Astra and Fable context growth is constant, but still massive. On avg 100k tokens for a single short nav task.

stash's profile picture
stash11 days ago

4/ Jev outperforms Fable on shorter paths, and wins dramatically on speed and cost Jev costs $0.081 per run versus Astra at $1.610 per run and Fable at $0.963 per run

stash's profile picture
stash11 days ago

5/ Initial results ranked by SPL (a nav trajectory quality metric)

stash's profile picture
stash11 days ago

6/ Full paper, results, and dataset navigator here!

Sitarama Chekuri's profile picture
Sitarama Chekuri11 days ago

@dimensionalos Nav arena looks so cool.

Dhruv Atreja's profile picture
Dhruv Atreja11 days ago

@dimensionalos Cool stuff!

asul's profile picture
asul11 days ago

@dimensionalos

Soroush Fadaeimanesh's profile picture
Soroush Fadaeimanesh11 days ago

@dimensionalos Real vs simulated split is the useful part of this benchmark, most navigation evals stay purely sim and quietly assume the gap closes on its own. Curious how much the numbers diverge between the two environment types once you break it down that way.

Johannes Tscharn's profile picture
Johannes Tscharn11 days ago

@dimensionalos Is there more info on the WorldState concept? How do you determine what’s included at any given frame?

stash's profile picture
stash11 days ago

@dimensionalos TLDR we look at free space, close objects, obstacles, cartesian distance to target object. And for each keyed object we include distance in meters, orientation (bearing out of 360), and relative position in language (ahead, back left, etc.)

Johannes Tscharn's profile picture
Johannes Tscharn11 days ago

@dimensionalos How did you come up with this method? Is there some related work / papers etc?

Related Videos

🔥 JUST IN: Open-source robotics dataset from 100% real-world scenarios! 🤯 Chinese robotics company AGIBOT just released AGIBOT WORLD 2026, an open-source dataset systematically covering key embodied AI research directions. Built entirely from real-world environments: commercial spaces, and homes. Collected using AGIBOT G2 robots in free-form collection mode, providing structured, accurately annotated, high-quality data. Digital twin technology creates 1:1 scale replicas in simulation matching the real environments. Both real-world and simulation data are open-sourced. The AGIBOT G2 platform collects multiple data types simultaneously: RGB(D) cameras, tactile sensors, force sensors, LiDAR, IMU, and full-body joint states. Whole-body control coordinates arms, waist, and hands for complex tasks. First-person teleoperation lets operators control the robot from its perspective. The tasks covered are fine-grained manipulation, ultra-long-horizon tasks, spatial navigation, dual-arm coordination, and multi-agent/human-robot collaboration. The dataset includes error-recovery trajectories with annotations. Most datasets only show successful demonstrations. AGIBOT includes failures and how the robot recovers, teaching models how to handle mistakes. After collection, data is tested through policy training and real-robot deployment to ensure quality. Then processed through industrial quality control with multiple screening and cleaning rounds. Making it open-source accelerates embodied AI research by giving researchers access to high-quality real-world robot data at scale. 🇨🇳 Learn more here: ~~ ♻️ Join the weekly robotics newsletter, and never miss any news →

Lukas Ziegler

40,583 views • 6 months ago

this is pure f*cking treasure Claude Code tip: once Opus 5.5 is your main model, stop leaving Fable 5.1 sitting idle and stop burning Opus tokens on tasks Sonnet 5.5 can swarm put it on call with /goal claude --auto-mode "/goal all tests pass and lint is clean" Opus 5.5 keeps writing the code in auto mode Sonnet 5.5 swarms the routine work at medium Fable 5.1 inspects the full transcript and only speaks up at the turn boundary: → not yet met: am I drifting from acceptance criteria? → met: do the tests, diffs, and lint hold? → impossible: am I chasing an unreachable state? Fable grades. Opus ships JEV engineering is the same move one layer down: the forks that need no thinker (which file, which tool, retry or halt) go to JEV in under 16ms, and the big models only see the ones that split — the full tree > Opus 5.5 on high runs the main session > worker edits and patches code on Sonnet 5.5 > explorer indexes AST maps and call sites on Sonnet 5.5 > verifier runs tests and linters on Sonnet 5.5 > all three subagents on medium effort > Fable 5.1 on call as the Stop hook evaluator > JEV routes micro-forks under 16ms paste the tree and this prompt into Claude Code ↓ "Reconfigure my Claude Code setup around this tree: 1. Set the main session to Opus 5.5 on high effort with auto mode enabled. 2. Put Fable 5.1 on call as a session-scoped evaluator attached to the Stop hook: - Read-only transcript evaluation with zero tool calls - Three structured outcomes: not yet met (with steering notes), met (auto-clear), and impossible (abort early) 3. Spawn Sonnet 5.5 subagents on medium effort for worker, explorer, and verifier roles. Defer goal checks while background tasks run. 4. Route fast deterministic forks (tool resolution, path selection, auto-retries) through JEV. 5. Add one rule to CLAUDE.md: run long tasks with: claude --auto-mode \"/goal \" Show me every configuration change as a diff first. No edits until I say go." ↳

mirku

40,973 views • 3 days ago

Jev has been exploding in popularity recently. If you already have access to the Jev API but aren't sure how to start experimenting with it, just copy this checklist: 1. jev-ultrafast Browser Use's fastest agent. Jev decides the next action and which element to click, and a language model is only called when text has to be typed. 2. typesafe-mario Jev plays Super Mario Bros. from structured emulator state, choosing every action from features pulled out of the game. 3. jev-plays-pokemon Reads Pokémon Red's game state as text, answers typed questions each turn, and lets plain code turn the answers into moves. 4. jev-drone A camera-only autonomous drone in MuJoCo with a Jev judgment model sitting in the control loop at 2.5 Hz. 5. robo-harness A real SO-101 robot arm workbench where Jev picks bounded joint steps from typed candidate actions under a spend budget. 6. fast-jev-compaction Claude Code plugin that replaces the compaction summary with Jev decisions, scoring every tool call for whether it is still needed. 7. jev-claude Routes Claude Code's own judgment calls through Jev: typed choices with probabilities at plan approval, on questions, and before risky commands. 8. is-malicious Supply-chain check before you run anything: Jev Noul checks over source and build files, returning the implicated files and lines. 9. sqlite-jev Jev inside SQL. Noul, Choice and Score judgments exposed as SQLite functions, with confidence on every row. 10. jevinci Paints images by having Jev predict every pixel's colour in parallel, with confidence deciding how wide each stroke is drawn. Copy these complete Jev blueprints - then read full Jev setup below ↓ ↓

Hanako

29,894 views • 15 days ago