An open-vocabulary, typed 3D scene graph for language-guided outdoor navigation.
Mobile robots can plan through geometric maps, but a geometric map does not identify the entities people name, and it does not record the physical relationships needed to interpret a request. Open-vocabulary vision widens what can be recognised, yet detections from individual images do not by themselves establish persistent entities, nor determine how those entities relate inside a reconstructed environment. This matters on a pedestrian campus, where a request may refer to an object, a ground surface, a building structure, an entrance, or a functional place.
CampusGraph grounds open-vocabulary camera observations in registered LiDAR geometry and consolidates them across views. Objects, surfaces, structures, entrances and functional places stay independently addressable, and the relations between them are derived from reconstructed geometry rather than from text similarity. A language model translates a natural-language request into logical constraints, which are then executed deterministically over the graph. Because every answer keeps its recovered spatial extent, a retrieved target can be handed directly to a global navigation roadmap as a goal.
A web interface for the scene graph is included: the reconstructed scene with its typed nodes and relations, natural-language queries answered over the graph, and routes planned to any returned node. See demo/README.md to run it, and the recorded walkthrough for a full session.
The pipeline runs left to right, from raw sequence to planned route. Open-vocabulary perception tags, grounds and segments each RGB frame and lifts every mask onto the registered LiDAR map; in parallel, a geometric partition voxelises that map and separates traversable ground from vertical structure. Layer assignment pools the two by multi-channel weighted agreement, routing each detection to the object, surface or structure stream. Object detections are then associated across views by spatial IoU, caption and SBERT similarity into persistent objects, and groups of those objects form functional places, while semantic projection and DBSCAN clustering consolidate map cells into surface and structure regions. The resulting layers are promoted to a single typed graph, over which a natural-language query is interpreted, executed deterministically, and bound to a navigation goal with a planned route.
Running the full pipeline over one sequence reconstructs a single graph with five node types and four relation types:
| Layer | node_type |
Nodes | Built from |
|---|---|---|---|
| Objects | instance |
196 | fused open-vocabulary object detections |
| Structure regions | structure_region |
53 | vertical building geometry with semantic support |
| Surface regions | surface_region |
39 | ground geometry with semantic support |
| Functional places | zone |
18 | class-anchored clusters of nearby objects |
| Entrances | entrance |
15 | facade openings on structure regions |
Relations are NEAR, ON_SURFACE, BELONGS_TO and PART_OF, giving 756
unique typed relations over 321 nodes. Each node carries its world-frame
points, a consolidated caption with its embedding, and a label elected from
channel-attributed detection votes.
campusgraph/
├── configs/ Hydra configuration; one block per pipeline stage
├── datasets/ Oxford Spires loading, calibration, undistortion
├── perception/ detection, geometry, routing, fusion, background, topology
├── graph/ node construction, relations, retrieval, navigation binding
├── zones/ class-anchored functional-place clustering
└── pipelines/ the runnable entry points described below
Every stage reads its inputs from disk and writes its own directory under
paths.results_root, so stages can be run individually or as a chain, and a
stage refuses to overwrite an existing output directory.
The pipeline was developed and run on Imperial College London's RDS cluster under a conda environment holding PyTorch, GroundingDINO, Tokenize-Anything, Recognize-Anything, OpenCLIP and Sentence-Transformers.
conda activate opengraphMachine-specific roots live in one place, paths: in
campusgraph/configs/perception.yaml. Point dependencies_dir at the model
weights, sequence_results_root at where results should be written, and
dataset.basedir / dataset.calib_dir in
campusgraph/configs/oxford_spires_dataset.yaml at the sequence. Any setting
can be overridden per run on the command line.
All commands are run from the repository root. Stages marked GPU load models; the rest are CPU-only and are safe on a login node.
python -m campusgraph.pipelines.run_detections # GPUFor every configured frame and camera: Tag2Text proposes an open-set prompt
list, GroundingDINO grounds those prompts to boxes, TAP returns one mask,
concept and caption per box, and OpenCLIP and SBERT encode the region and its
caption. Mask pixels are matched to projected LiDAR returns and stored in world
coordinates, which is what makes a 2D detection a 3D observation.
Writes instance_detections/.
python -m campusgraph.pipelines.run_detection_evidence # CPUExtracts the grammatical head of each caption and materialises the LiDAR points
behind person-like and other dynamic detections, so the geometry stage can
exclude them. Writes detection_evidence/.
python -m campusgraph.pipelines.run_geometry # CPUAccumulates the registered LiDAR scans into a dense world-frame cloud with
dynamic points removed, voxelises it, and segments ground from vertical
structure using a robust height-plane fit taken from a radius around each
voxel. The walked trajectory seeds a region-grow that confirms the traversable
surface manifold. Writes geometry/.
python -m campusgraph.pipelines.run_detection_routing # CPURoutes every detection to the object, surface or structure stream by combining
the detector term, the caption head and the geometry under the detection. A
detection whose channels disagree contributes only where each channel agrees,
which is what later allows a label to be elected from attributed votes rather
than asserted. Writes routing/.
python -m campusgraph.pipelines.run_background_builder # CPU
python -m campusgraph.pipelines.run_instance_fusion # CPUThe background builder projects routed semantic evidence onto the segmented
geometry and clusters it with DBSCAN into surface and structure regions,
detecting facade openings as entrances. Instance fusion associates object-stream
detections across frames and cameras by aggregated spatial IoU, caption and
SBERT similarity, requiring nonzero spatial overlap so that description
similarity alone can never merge two physically separate objects. Writes
background/ and object_fusion/.
python -m campusgraph.pipelines.run_navgrid # CPU
python -m campusgraph.pipelines.run_gvd_graph # CPU
python -m campusgraph.pipelines.run_path_planner # CPUProjects the geometry voxels to a 2.5D navigation grid, extracts the
generalised Voronoi diagram of the free space as a maximum-clearance skeleton,
and contracts it into a routable graph. The planner then produces routes over
that graph. Writes topology/.
python -m campusgraph.pipelines.run_zone_clustering # CPU
python -m campusgraph.pipelines.run_caption_consolidation # GPU
python -m campusgraph.pipelines.run_zone_captions # GPUZone clustering groups nearby instances of anchor classes into functional
places, such as a bike-parking area or a garden. Caption consolidation reduces
the many captions collected on each node to one label-independent description
and its embedding, falling back to a deterministic dominant caption whenever
the evidence is thin. Writes zones/ and node_captions/.
python -m campusgraph.pipelines.run_graph # CPUPromotes every persistent entity to a typed node, elects each instance label
from its accumulated channel votes, joins the consolidated captions, and derives
the typed relations between nodes from their geometry. Writes graph/.
python -m campusgraph.pipelines.run_graph_retrieval \
retrieval.examples_file=<query-set.json> \
retrieval.evidence.enabled=trueA language model is shown a catalog of what the graph actually contains and
returns a structured plan: a target, typed relations to reference entities, and
any ordering or aggregation. The plan is validated against a closed vocabulary
and then executed deterministically over the graph, so the model decides what
was asked and never what the answer is. Each retrieved node is bound to the
navigation graph by its spatial extent and returned with a planned route.
Writes retrieval_results/.
For a stage-by-stage account of what the pipeline does, see docs/pipeline.md. To evaluate retrieval on your own scene — writing a query set, annotating the graph, computing relevance and comparing against baselines — see docs/benchmarking.md.
Evaluated on the Observatory Quarter sequence of the Oxford Spires dataset, against a purpose-built benchmark with scene-grounded answer sets annotated per query. Over 56 queries, structured graph retrieval reaches a mean average precision of 0.800, against 0.536 for the strongest similarity baseline. Flattening the graph shows that preserving node and relation types is what stops nearby but physically different entities from satisfying the same constraint. Treating retrieved targets as navigation goals, the top-ranked policy yields successful planned routes in 82.4% of 255 scored episodes. Routes are planned in the reconstructed environment and are not executed on a physical robot.
Oxford Spires,
sequence 2024-03-13-observatory-quarter-01, using the multi-camera rig, the
undistorted LiDAR scans and the HBA trajectory.
If you use this work, please cite the accompanying MSc thesis:
A. H. Haidar, Campus Graph: Open-vocabulary 3D Scene Graph for Semantic Navigation, MSc thesis, Imperial College London.

