Skip to content

About

open-vocabulary 3D scene graph for semantic navigation

Resources

Stars

1 star

Watchers

0 watching

Forks

Latest commit

 

History

1 Commit

Folders and files

Repository files navigation

CampusGraph

An open-vocabulary, typed 3D scene graph for language-guided outdoor navigation.

Mobile robots can plan through geometric maps, but a geometric map does not identify the entities people name, and it does not record the physical relationships needed to interpret a request. Open-vocabulary vision widens what can be recognised, yet detections from individual images do not by themselves establish persistent entities, nor determine how those entities relate inside a reconstructed environment. This matters on a pedestrian campus, where a request may refer to an object, a ground surface, a building structure, an entrance, or a functional place.

CampusGraph grounds open-vocabulary camera observations in registered LiDAR geometry and consolidates them across views. Objects, surfaces, structures, entrances and functional places stay independently addressable, and the relations between them are derived from reconstructed geometry rather than from text similarity. A language model translates a natural-language request into logical constraints, which are then executed deterministically over the graph. Because every answer keeps its recovered spatial extent, a retrieved target can be handed directly to a global navigation roadmap as a goal.

A web interface for the scene graph is included: the reconstructed scene with its typed nodes and relations, natural-language queries answered over the graph, and routes planned to any returned node. See demo/README.md to run it, and the recorded walkthrough for a full session.

CampusGraph system interface

System pipeline

CampusGraph system pipeline

The pipeline runs left to right, from raw sequence to planned route. Open-vocabulary perception tags, grounds and segments each RGB frame and lifts every mask onto the registered LiDAR map; in parallel, a geometric partition voxelises that map and separates traversable ground from vertical structure. Layer assignment pools the two by multi-channel weighted agreement, routing each detection to the object, surface or structure stream. Object detections are then associated across views by spatial IoU, caption and SBERT similarity into persistent objects, and groups of those objects form functional places, while semantic projection and DBSCAN clustering consolidate map cells into surface and structure regions. The resulting layers are promoted to a single typed graph, over which a natural-language query is interpreted, executed deterministically, and bound to a navigation goal with a planned route.

What the system produces

Running the full pipeline over one sequence reconstructs a single graph with five node types and four relation types:

Layer node_type Nodes Built from
Objects instance 196 fused open-vocabulary object detections
Structure regions structure_region 53 vertical building geometry with semantic support
Surface regions surface_region 39 ground geometry with semantic support
Functional places zone 18 class-anchored clusters of nearby objects
Entrances entrance 15 facade openings on structure regions

Relations are NEAR, ON_SURFACE, BELONGS_TO and PART_OF, giving 756 unique typed relations over 321 nodes. Each node carries its world-frame points, a consolidated caption with its embedding, and a label elected from channel-attributed detection votes.

Repository layout

campusgraph/
├── configs/        Hydra configuration; one block per pipeline stage
├── datasets/       Oxford Spires loading, calibration, undistortion
├── perception/     detection, geometry, routing, fusion, background, topology
├── graph/          node construction, relations, retrieval, navigation binding
├── zones/          class-anchored functional-place clustering
└── pipelines/      the runnable entry points described below

Every stage reads its inputs from disk and writes its own directory under paths.results_root, so stages can be run individually or as a chain, and a stage refuses to overwrite an existing output directory.

Environment

The pipeline was developed and run on Imperial College London's RDS cluster under a conda environment holding PyTorch, GroundingDINO, Tokenize-Anything, Recognize-Anything, OpenCLIP and Sentence-Transformers.

conda activate opengraph

Machine-specific roots live in one place, paths: in campusgraph/configs/perception.yaml. Point dependencies_dir at the model weights, sequence_results_root at where results should be written, and dataset.basedir / dataset.calib_dir in campusgraph/configs/oxford_spires_dataset.yaml at the sequence. Any setting can be overridden per run on the command line.

Running the pipeline end to end

All commands are run from the repository root. Stages marked GPU load models; the rest are CPU-only and are safe on a login node.

1. Semantic perception

python -m campusgraph.pipelines.run_detections                # GPU

For every configured frame and camera: Tag2Text proposes an open-set prompt list, GroundingDINO grounds those prompts to boxes, TAP returns one mask, concept and caption per box, and OpenCLIP and SBERT encode the region and its caption. Mask pixels are matched to projected LiDAR returns and stored in world coordinates, which is what makes a 2D detection a 3D observation. Writes instance_detections/.

python -m campusgraph.pipelines.run_detection_evidence        # CPU

Extracts the grammatical head of each caption and materialises the LiDAR points behind person-like and other dynamic detections, so the geometry stage can exclude them. Writes detection_evidence/.

2. Geometric partition

python -m campusgraph.pipelines.run_geometry                  # CPU

Accumulates the registered LiDAR scans into a dense world-frame cloud with dynamic points removed, voxelises it, and segments ground from vertical structure using a robust height-plane fit taken from a radius around each voxel. The walked trajectory seeds a region-grow that confirms the traversable surface manifold. Writes geometry/.

3. Layer assignment

python -m campusgraph.pipelines.run_detection_routing         # CPU

Routes every detection to the object, surface or structure stream by combining the detector term, the caption head and the geometry under the detection. A detection whose channels disagree contributes only where each channel agrees, which is what later allows a label to be elected from attributed votes rather than asserted. Writes routing/.

4. Persistent regions and objects

python -m campusgraph.pipelines.run_background_builder        # CPU
python -m campusgraph.pipelines.run_instance_fusion           # CPU

The background builder projects routed semantic evidence onto the segmented geometry and clusters it with DBSCAN into surface and structure regions, detecting facade openings as entrances. Instance fusion associates object-stream detections across frames and cameras by aggregated spatial IoU, caption and SBERT similarity, requiring nonzero spatial overlap so that description similarity alone can never merge two physically separate objects. Writes background/ and object_fusion/.

5. Navigation topology

python -m campusgraph.pipelines.run_navgrid                   # CPU
python -m campusgraph.pipelines.run_gvd_graph                 # CPU
python -m campusgraph.pipelines.run_path_planner              # CPU

Projects the geometry voxels to a 2.5D navigation grid, extracts the generalised Voronoi diagram of the free space as a maximum-clearance skeleton, and contracts it into a routable graph. The planner then produces routes over that graph. Writes topology/.

6. Functional places and captions

python -m campusgraph.pipelines.run_zone_clustering           # CPU
python -m campusgraph.pipelines.run_caption_consolidation     # GPU
python -m campusgraph.pipelines.run_zone_captions             # GPU

Zone clustering groups nearby instances of anchor classes into functional places, such as a bike-parking area or a garden. Caption consolidation reduces the many captions collected on each node to one label-independent description and its embedding, falling back to a deterministic dominant caption whenever the evidence is thin. Writes zones/ and node_captions/.

7. Scene graph assembly

python -m campusgraph.pipelines.run_graph                     # CPU

Promotes every persistent entity to a typed node, elects each instance label from its accumulated channel votes, joins the consolidated captions, and derives the typed relations between nodes from their geometry. Writes graph/.

8. Retrieval and goal binding

python -m campusgraph.pipelines.run_graph_retrieval \
    retrieval.examples_file=<query-set.json> \
    retrieval.evidence.enabled=true

A language model is shown a catalog of what the graph actually contains and returns a structured plan: a target, typed relations to reference entities, and any ordering or aggregation. The plan is validated against a closed vocabulary and then executed deterministically over the graph, so the model decides what was asked and never what the answer is. Each retrieved node is bound to the navigation graph by its spatial extent and returned with a planned route. Writes retrieval_results/.

For a stage-by-stage account of what the pipeline does, see docs/pipeline.md. To evaluate retrieval on your own scene — writing a query set, annotating the graph, computing relevance and comparing against baselines — see docs/benchmarking.md.

Results

Evaluated on the Observatory Quarter sequence of the Oxford Spires dataset, against a purpose-built benchmark with scene-grounded answer sets annotated per query. Over 56 queries, structured graph retrieval reaches a mean average precision of 0.800, against 0.536 for the strongest similarity baseline. Flattening the graph shows that preserving node and relation types is what stops nearby but physically different entities from satisfying the same constraint. Treating retrieved targets as navigation goals, the top-ranked policy yields successful planned routes in 82.4% of 255 scored episodes. Routes are planned in the reconstructed environment and are not executed on a physical robot.

Dataset

Oxford Spires, sequence 2024-03-13-observatory-quarter-01, using the multi-camera rig, the undistorted LiDAR scans and the HBA trajectory.

Citation

If you use this work, please cite the accompanying MSc thesis:

A. H. Haidar, Campus Graph: Open-vocabulary 3D Scene Graph for Semantic Navigation, MSc thesis, Imperial College London.

About

open-vocabulary 3D scene graph for semantic navigation

Resources

Stars

1 star

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages