Skip to content

Latest commit

ย 

History

6 Commits

Folders and files

NameName
Last commit message
Last commit date
ย 
ย 
ย 
ย 
ย 
ย 
ย 
ย 
ย 
ย 
ย 
ย 
ย 
ย 
ย 
ย 
ย 
ย 
ย 
ย 
ย 
ย 
ย 
ย 
ย 
ย 
ย 
ย 
ย 
ย 

Repository files navigation

๐Ÿงญ DepthJev

Turning Depth into Text for Embodied Navigation

Project page lint Python 3.11 Apache-2.0 EB-Navigation success 46.7% 0.76 s per step

How it works ยท Getting Started ยท Results

English | ็ฎ€ไฝ“ไธญๆ–‡

DepthJev is a navigation agent for EmbodiedBench EB-Navigation. It converts each RGB frame into metric depth with Depth Anything 3, detects the target with OWLv2, and turns the geometry into short text facts. The Jev decision model reads only these facts, never the image, and picks the next action.

DepthJev pipeline: RGB observation, task and history โ†’ metric depth and target detection โ†’ text facts โ†’ Jev navigation action โ†’ next observation

โœจ Highlights

  • ๐Ÿ“ˆ Strong without pixels. 46.7% success on all 300 EB-Navigation episodes, above Claude-3.5-Sonnet (44.7%) and 29 points above text-only GPT-4o (17.4%) as reported by EmbodiedBench.
  • ๐Ÿ“ Metric grounding. Free space in five sectors, a collision check for the next 0.25 m step and the target distance all come from monocular metric depth, so Jev reasons in metres, not pixels.
  • โšก Fast. 0.76 s per step on one H100: DA3 0.15 s, OWLv2 0.22 s, Jev 0.34 s.

๐Ÿ† Where it stands

Agent Input to the decision model EB-Navigation success
GPT-4o image + text 57.7
Gemini-2.0-flash image + text 48.7
DepthJev (ours) text facts only 46.7
Claude-3.5-Sonnet image + text 44.7
InternVL2.5-78B image + text 30.7
Qwen2-VL-72B image + text 21.2
GPT-4o text only 17.4

Baselines are from Table 3 of the EmbodiedBench paper.

๐Ÿ—๏ธ How it works

  1. ๐Ÿ‘๏ธ Perceive. Depth Anything 3 estimates metric depth; OWLv2 finds the target.
  2. ๐Ÿ“ Describe. Geometry and action history become text facts about free space, target distance and movement constraints.
  3. ๐ŸŽฏ Act. Jev reads the facts, selects one of eight navigation actions, and the loop repeats with the next observation.

๐Ÿ Example

โ€œI need a vessel to boil pasta for dinner. Can you navigate to that object and stay close?โ€ DepthJev resolves the target to Pot and reaches it in 9 steps. Boxes are OWLv2 detections, distances come from DA3.

DepthJev reaching a pot in 9 steps: frames of steps 0, 3, 6, 7 and 8 with the detected pot and its distance

What Jev reads at step 0
{
  "task": {
    "instruction": "I need a vessel to boil pasta for dinner. Can you navigate to that object and stay close?",
    "target_object": "pot",
    "success_rule": "the episode succeeds as soon as the robot stands within 1 m of the target object",
    "step": 1,
    "steps_left": 20
  },
  "target": {
    "visible": true,
    "sector": "center",
    "distance": "over 2 m"
  },
  "free_distance_ahead_by_sector": {
    "far_left": "1 to 2 m",
    "left": "over 2 m",
    "center": "over 2 m",
    "right": "1 to 2 m",
    "far_right": "under 0.5 m"
  },
  "move_check": {
    "move_ahead": "clear",
    "move_back": "unknown",
    "move_left": "unknown",
    "move_right": "unknown"
  },
  "move_ahead_check_saw_the_floor": false,
  "camera_pitch": "level",
  "recent_actions": []
}

Decision: move_ahead ยท probability 1.00 ยท 1,491 input tokens ยท 0.37 s.

๐Ÿš€ Getting Started

All commands run from the repository root.

1. Clone the repository and its dependencies

git clone https://github.com/ZJUCQR/DepthJev.git && cd DepthJev
git clone https://github.com/EmbodiedBench/EmbodiedBench.git repos/EmbodiedBench
git clone https://github.com/ByteDance-Seed/Depth-Anything-3.git repos/Depth-Anything-3

2. Server environment

conda create -y -p envs/depthjev python=3.11 pip
conda activate ./envs/depthjev
pip install -r requirements.txt
hf download depth-anything/DA3METRIC-LARGE --local-dir checkpoints/DA3METRIC-LARGE
hf download google/owlv2-base-patch16-ensemble --local-dir checkpoints/owlv2-base-patch16-ensemble

3. Evaluation environment

conda create -y -p envs/depthjev-eval python=3.9.21 pip
conda activate ./envs/depthjev-eval
pip install -r requirements-eval.txt

4. Headless rendering for AI2-THOR

sudo apt-get install -y --no-install-recommends xvfb x11-utils libgl1-mesa-dri libgl1-mesa-glx \
    libglu1-mesa libxcursor1 libxrandr2 libxinerama1 libxi6 libxxf86vm1

5. Jev API key

cp .env.example .env    # then fill in TYPESAFE_API_KEY

6. Run

bash scripts/run.sh                                       # smoke test, 3 base episodes
bash scripts/run.sh full --sets all --ratio 1             # all 300 episodes, about 2.2 h
bash scripts/run.sh full --sets all --ratio 1 --parallel  # one server per subset, add --gpus 0 to share one card
Option Default Description
RUN_NAME smoke name of the run
--sets base comma-separated subsets, or all
--ratio 0.05 EmbodiedBench down_sample_ratio, 1 runs every task
--gpus visible GPUs CUDA devices for the servers
--parallel off one server + evaluator per subset
Server and evaluator in separate terminals
DEPTHJEV_RUN_NAME=dev bash scripts/server.sh
SERVER_URL=http://127.0.0.1:23333/process bash scripts/eval.sh exp_name=dev eval_sets=[base] down_sample_ratio=0.05

The server listens on 127.0.0.1 by default. For an evaluator on another machine, start it with bash scripts/server.sh --host 0.0.0.0 and point SERVER_URL at the server.

โš™๏ธ Configuration

All settings live in config.json.

Key Default
server_env, eval_env envs/depthjev, envs/depthjev-eval
da3_dir, owlv2_dir checkpoints/...
jev_model, jev_timeout jev-latest, 20
port, detection_threshold 23333, 0.1

๐Ÿ“Š Results

All 300 EB-Navigation episodes on one H100, run with bash scripts/run.sh full --sets all --ratio 1.

Subset Successes Success rate Steps / episode Target type correct
base 32/60 0.533 15.2 60/60
common_sense 31/60 0.517 15.4 58/60
complex_instruction 30/60 0.500 15.3 60/60
visual_appearance 22/60 0.367 16.9 25/60
long_horizon 25/60 0.417 17.3 60/60
All 140/300 0.467 16.0 263/300
โฑ๏ธ Server latency

Mean, p50 and p90 are for one run; the last column shows three runs sharing one GPU. All values are in seconds.

Stage Mean p50 p90 Mean, 3 runs/GPU
DA3 depth 0.15 0.13 0.23 0.27
OWLv2 detection 0.22 0.21 0.22 0.32
Jev decision 0.34 0.27 0.51 0.38
Target resolution, once per episode 0.82 0.61 1.63 1.42
Whole step 0.76 0.64 1.04 1.06

With AI2-THOR software rendering, a single run averages about 21 s per episode.

๐Ÿ” Failure analysis: 160 unsuccessful episodes
Pattern Episodes
Target rarely detected or incorrectly identified 51
Sidestepping left and right without progress 40
Stuck against obstacles 32
Distance underestimated or false target detection 27
Other 10
  • Detection. The first group includes 17 visual_appearance episodes with an incorrect target type. Adding "trash can" and "bin" improved GarbageCan success from 0/5 to 5/5 on a development sample; low-score detections (0.1โ€“0.2) remain a common source of errors.
  • Action constraints. Attaching constraints directly to action options prevented repeated look_down actions more effectively than a general rule.
  • Scene difficulty. Across the same 60 scenes and targets, 12 succeeded under all five instruction variants and 22 failed under all five. long_horizon also changes the starting orientation by 180ยฐ.

๐Ÿ“ Project Structure

depthjev/
โ”œโ”€โ”€ server.py         # Flask /process + /health, one JSONL line per request
โ”œโ”€โ”€ policy.py         # one step: parse โ†’ depth โ†’ detect โ†’ facts โ†’ Jev โ†’ EmbodiedBench JSON
โ”œโ”€โ”€ prompt_parse.py   # instruction and action history from the EmbodiedBench prompt
โ”œโ”€โ”€ depth.py          # DA3METRIC-LARGE โ†’ metric depth
โ”œโ”€โ”€ detect.py         # OWLv2 target detection
โ”œโ”€โ”€ geometry.py       # back-projection, floor estimate, sector distances, step collision
โ”œโ”€โ”€ facts.py          # the state Jev reads, move checks, search status
โ”œโ”€โ”€ jev_client.py     # the three Jev Choices, rules, iTHOR types, detector aliases
โ”œโ”€โ”€ actions.py        # the eight EB-Navigation actions
โ”œโ”€โ”€ config.py         # config.json โ†’ shell variables
โ”œโ”€โ”€ eb_launch.py      # starts EmbodiedBench in the evaluation environment
โ””โ”€โ”€ report.py         # latency, success and per-episode tables
scripts/
โ”œโ”€โ”€ run.sh            # end-to-end evaluation, single or --parallel
โ”œโ”€โ”€ server.sh         # server only
โ””โ”€โ”€ eval.sh           # evaluator only

๐Ÿ™ Acknowledgements

๐Ÿ“„ License

DepthJev is released under the Apache-2.0 License.

About

An embodied navigation agent that converts RGB observations into metric depth and text facts for Jev to choose actions in EmbodiedBench EB-Navigation.

Resources

Stars

1 star

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages