How it works ยท Getting Started ยท Results
English | ็ฎไฝไธญๆ
DepthJev is a navigation agent for EmbodiedBench EB-Navigation. It converts each RGB frame into metric depth with Depth Anything 3, detects the target with OWLv2, and turns the geometry into short text facts. The Jev decision model reads only these facts, never the image, and picks the next action.
- ๐ Strong without pixels. 46.7% success on all 300 EB-Navigation episodes, above Claude-3.5-Sonnet (44.7%) and 29 points above text-only GPT-4o (17.4%) as reported by EmbodiedBench.
- ๐ Metric grounding. Free space in five sectors, a collision check for the next 0.25 m step and the target distance all come from monocular metric depth, so Jev reasons in metres, not pixels.
- โก Fast. 0.76 s per step on one H100: DA3 0.15 s, OWLv2 0.22 s, Jev 0.34 s.
| Agent | Input to the decision model | EB-Navigation success |
|---|---|---|
| GPT-4o | image + text | 57.7 |
| Gemini-2.0-flash | image + text | 48.7 |
| DepthJev (ours) | text facts only | 46.7 |
| Claude-3.5-Sonnet | image + text | 44.7 |
| InternVL2.5-78B | image + text | 30.7 |
| Qwen2-VL-72B | image + text | 21.2 |
| GPT-4o | text only | 17.4 |
Baselines are from Table 3 of the EmbodiedBench paper.
- ๐๏ธ Perceive. Depth Anything 3 estimates metric depth; OWLv2 finds the target.
- ๐ Describe. Geometry and action history become text facts about free space, target distance and movement constraints.
- ๐ฏ Act. Jev reads the facts, selects one of eight navigation actions, and the loop repeats with the next observation.
โI need a vessel to boil pasta for dinner. Can you navigate to that object and stay close?โ DepthJev resolves the target to Pot and reaches it in 9 steps. Boxes are OWLv2 detections, distances come from DA3.
What Jev reads at step 0
{
"task": {
"instruction": "I need a vessel to boil pasta for dinner. Can you navigate to that object and stay close?",
"target_object": "pot",
"success_rule": "the episode succeeds as soon as the robot stands within 1 m of the target object",
"step": 1,
"steps_left": 20
},
"target": {
"visible": true,
"sector": "center",
"distance": "over 2 m"
},
"free_distance_ahead_by_sector": {
"far_left": "1 to 2 m",
"left": "over 2 m",
"center": "over 2 m",
"right": "1 to 2 m",
"far_right": "under 0.5 m"
},
"move_check": {
"move_ahead": "clear",
"move_back": "unknown",
"move_left": "unknown",
"move_right": "unknown"
},
"move_ahead_check_saw_the_floor": false,
"camera_pitch": "level",
"recent_actions": []
}Decision: move_ahead ยท probability 1.00 ยท 1,491 input tokens ยท 0.37 s.
All commands run from the repository root.
1. Clone the repository and its dependencies
git clone https://github.com/ZJUCQR/DepthJev.git && cd DepthJev
git clone https://github.com/EmbodiedBench/EmbodiedBench.git repos/EmbodiedBench
git clone https://github.com/ByteDance-Seed/Depth-Anything-3.git repos/Depth-Anything-32. Server environment
conda create -y -p envs/depthjev python=3.11 pip
conda activate ./envs/depthjev
pip install -r requirements.txt
hf download depth-anything/DA3METRIC-LARGE --local-dir checkpoints/DA3METRIC-LARGE
hf download google/owlv2-base-patch16-ensemble --local-dir checkpoints/owlv2-base-patch16-ensemble3. Evaluation environment
conda create -y -p envs/depthjev-eval python=3.9.21 pip
conda activate ./envs/depthjev-eval
pip install -r requirements-eval.txt4. Headless rendering for AI2-THOR
sudo apt-get install -y --no-install-recommends xvfb x11-utils libgl1-mesa-dri libgl1-mesa-glx \
libglu1-mesa libxcursor1 libxrandr2 libxinerama1 libxi6 libxxf86vm15. Jev API key
cp .env.example .env # then fill in TYPESAFE_API_KEY6. Run
bash scripts/run.sh # smoke test, 3 base episodes
bash scripts/run.sh full --sets all --ratio 1 # all 300 episodes, about 2.2 h
bash scripts/run.sh full --sets all --ratio 1 --parallel # one server per subset, add --gpus 0 to share one card| Option | Default | Description |
|---|---|---|
RUN_NAME |
smoke |
name of the run |
--sets |
base |
comma-separated subsets, or all |
--ratio |
0.05 |
EmbodiedBench down_sample_ratio, 1 runs every task |
--gpus |
visible GPUs | CUDA devices for the servers |
--parallel |
off | one server + evaluator per subset |
Server and evaluator in separate terminals
DEPTHJEV_RUN_NAME=dev bash scripts/server.sh
SERVER_URL=http://127.0.0.1:23333/process bash scripts/eval.sh exp_name=dev eval_sets=[base] down_sample_ratio=0.05The server listens on 127.0.0.1 by default. For an evaluator on another machine, start it with bash scripts/server.sh --host 0.0.0.0 and point SERVER_URL at the server.
All settings live in config.json.
| Key | Default |
|---|---|
server_env, eval_env |
envs/depthjev, envs/depthjev-eval |
da3_dir, owlv2_dir |
checkpoints/... |
jev_model, jev_timeout |
jev-latest, 20 |
port, detection_threshold |
23333, 0.1 |
All 300 EB-Navigation episodes on one H100, run with bash scripts/run.sh full --sets all --ratio 1.
| Subset | Successes | Success rate | Steps / episode | Target type correct |
|---|---|---|---|---|
| base | 32/60 | 0.533 | 15.2 | 60/60 |
| common_sense | 31/60 | 0.517 | 15.4 | 58/60 |
| complex_instruction | 30/60 | 0.500 | 15.3 | 60/60 |
| visual_appearance | 22/60 | 0.367 | 16.9 | 25/60 |
| long_horizon | 25/60 | 0.417 | 17.3 | 60/60 |
| All | 140/300 | 0.467 | 16.0 | 263/300 |
โฑ๏ธ Server latency
Mean, p50 and p90 are for one run; the last column shows three runs sharing one GPU. All values are in seconds.
| Stage | Mean | p50 | p90 | Mean, 3 runs/GPU |
|---|---|---|---|---|
| DA3 depth | 0.15 | 0.13 | 0.23 | 0.27 |
| OWLv2 detection | 0.22 | 0.21 | 0.22 | 0.32 |
| Jev decision | 0.34 | 0.27 | 0.51 | 0.38 |
| Target resolution, once per episode | 0.82 | 0.61 | 1.63 | 1.42 |
| Whole step | 0.76 | 0.64 | 1.04 | 1.06 |
With AI2-THOR software rendering, a single run averages about 21 s per episode.
๐ Failure analysis: 160 unsuccessful episodes
| Pattern | Episodes |
|---|---|
| Target rarely detected or incorrectly identified | 51 |
| Sidestepping left and right without progress | 40 |
| Stuck against obstacles | 32 |
| Distance underestimated or false target detection | 27 |
| Other | 10 |
- Detection. The first group includes 17
visual_appearanceepisodes with an incorrect target type. Adding "trash can" and "bin" improved GarbageCan success from 0/5 to 5/5 on a development sample; low-score detections (0.1โ0.2) remain a common source of errors. - Action constraints. Attaching constraints directly to action options prevented repeated
look_downactions more effectively than a general rule. - Scene difficulty. Across the same 60 scenes and targets, 12 succeeded under all five instruction variants and 22 failed under all five.
long_horizonalso changes the starting orientation by 180ยฐ.
depthjev/
โโโ server.py # Flask /process + /health, one JSONL line per request
โโโ policy.py # one step: parse โ depth โ detect โ facts โ Jev โ EmbodiedBench JSON
โโโ prompt_parse.py # instruction and action history from the EmbodiedBench prompt
โโโ depth.py # DA3METRIC-LARGE โ metric depth
โโโ detect.py # OWLv2 target detection
โโโ geometry.py # back-projection, floor estimate, sector distances, step collision
โโโ facts.py # the state Jev reads, move checks, search status
โโโ jev_client.py # the three Jev Choices, rules, iTHOR types, detector aliases
โโโ actions.py # the eight EB-Navigation actions
โโโ config.py # config.json โ shell variables
โโโ eb_launch.py # starts EmbodiedBench in the evaluation environment
โโโ report.py # latency, success and per-episode tables
scripts/
โโโ run.sh # end-to-end evaluation, single or --parallel
โโโ server.sh # server only
โโโ eval.sh # evaluator only
- EmbodiedBench for EB-Navigation
- Depth Anything 3 for metric depth
- OWLv2 for open-vocabulary detection
- TypeSafe Jev for the decision model
- AI2-THOR for the simulator and its object types
DepthJev is released under the Apache-2.0 License.

