Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension


Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
15 changes: 15 additions & 0 deletions README.md
Original file line number Diff line number Diff line change
Expand Up @@ -178,6 +178,20 @@ above gives SpatialRust a measured 1.07× 4K reuse lead, with maximum error zero
VGA and 1080p remain narrow OpenCV wins. See the
[acceleration receipt](notes/2026-07-15_exact_edt_acceleration.md).

For AI detection post-processing, the seeded Python NMS harness uses identical
float32 boxes, scores, and thresholds and requires exact kept-index parity
before publishing timings:

| NMS candidates | OpenCV `dnn.NMSBoxes` | SpatialRust `nms` | Result |
| ---: | ---: | ---: | ---: |
| 100 | 0.298 ms | 0.033 ms | **SpatialRust 8.95×** |
| 1,000 | 8.720 ms | 2.286 ms | **SpatialRust 3.82×** |
| 8,400 (YOLO-style) | 407.086 ms | 126.562 ms | **SpatialRust 3.22×** |

These Windows-host medians include each Python API call and returned indices;
see the [NMS harness](bench/opencv_nms_comparison/) and dated
[receipt](notes/2026-07-15_nms_opencv_acceleration.md).

#### Vision accuracy

The same deterministic RGB inputs passed all VGA, 1080p, and 4K gates:
Expand All @@ -204,6 +218,7 @@ On dense `H×W×3` XYZ (320×240, OpenCL off, local Windows laptop), `spatialrus
python bench\opencv_vision_comparison\run.py
python bench\opencv_vision_comparison\performance.py
python bench\opencv_rgbd_comparison\run.py
python bench\opencv_nms_comparison\performance.py
```

### Registration methods
Expand Down
2 changes: 2 additions & 0 deletions bench/opencv_comparison/README.md
Original file line number Diff line number Diff line change
Expand Up @@ -32,6 +32,7 @@ VGA, 1080p, and 4K profiles and the initial competitive workload set:
10. colored RGB-D to point cloud
11. AI preprocessing
12. RGB-D to voxel end-to-end
13. detection NMS post-processing

Exact matches use a JSON `null` PSNR (mathematically infinite) so reports remain
strict RFC-compatible JSON. Numerical comparisons retain max/mean/RMS and
Expand All @@ -49,6 +50,7 @@ then run both current suites:

```powershell
python bench\opencv_comparison\run.py
python bench\opencv_nms_comparison\performance.py
```

Reports are written under `target/opencv-comparison/`. Run one suite with
Expand Down
1 change: 1 addition & 0 deletions bench/opencv_comparison/manifest.json
Original file line number Diff line number Diff line change
Expand Up @@ -47,6 +47,7 @@
{ "id": "depth_to_xyz", "domain": "rgbd", "modes": ["allocate", "reuse"] },
{ "id": "rgbd_to_point_cloud", "domain": "spatial-e2e", "modes": ["allocate"] },
{ "id": "ai_preprocess", "domain": "dnn-adapter", "modes": ["allocate", "reuse"] },
{ "id": "nms", "domain": "dnn-adapter", "modes": ["postprocess"] },
{ "id": "rgbd_to_voxel", "domain": "spatial-e2e", "modes": ["allocate"] }
]
}
1 change: 1 addition & 0 deletions bench/opencv_comparison/test_report.py
Original file line number Diff line number Diff line change
Expand Up @@ -117,6 +117,7 @@ def test_manifest_reserves_representative_profiles_and_workloads(self) -> None:
self.assertGreaterEqual(len(workloads), 10)
self.assertIn("rgbd_to_voxel", workloads)
self.assertIn("ai_preprocess", workloads)
self.assertIn("nms", workloads)
self.assertIn("coefficient_of_variation", statistics)
self.assertIn("median_absolute_deviation", statistics)
self.assertIn("batch_size", statistics)
Expand Down
15 changes: 15 additions & 0 deletions bench/opencv_nms_comparison/README.md
Original file line number Diff line number Diff line change
@@ -0,0 +1,15 @@
# OpenCV NMS comparison

This harness compares SpatialRust `nms` with OpenCV `dnn.NMSBoxes` using the
same deterministic float32 boxes, scores, score threshold, and IoU threshold.
It covers small post-processing, 1,000-candidate, and YOLO-style 8,400-candidate
profiles. Returned indices must match exactly before timings are published.

```powershell
python bench/opencv_nms_comparison/performance.py `
--output target/opencv-nms-performance.json
```

The report follows `spatialrust.opencv-comparison.v1` and records raw samples,
dispersion, library versions, thread policy, and the host environment. Results
are machine-specific and must not be generalized beyond the named workload.
124 changes: 124 additions & 0 deletions bench/opencv_nms_comparison/performance.py
Original file line number Diff line number Diff line change
@@ -0,0 +1,124 @@
"""Reproducible Python NMS performance and parity comparison with OpenCV."""

from __future__ import annotations

import argparse
import sys
from pathlib import Path

import cv2
import numpy as np
import spatialrust as sr

sys.path.insert(0, str(Path(__file__).resolve().parents[1]))
from opencv_comparison.report import emit_report, environment, make_report, timed_pair


PROFILES = {
"small_100": (100, 50),
"medium_1000": (1_000, 20),
"yolo_8400": (8_400, 8),
}


def parse_args() -> argparse.Namespace:
parser = argparse.ArgumentParser()
parser.add_argument("--output", type=Path)
parser.add_argument("--profiles", default=",".join(PROFILES))
parser.add_argument("--warmup", type=int, default=3)
return parser.parse_args()


def main() -> None:
args = parse_args()
selected = [name.strip() for name in args.profiles.split(",") if name.strip()]
unknown = sorted(set(selected) - PROFILES.keys())
if unknown:
raise ValueError(f"unknown profiles: {', '.join(unknown)}")
if args.warmup < 0:
raise ValueError("warmup must be non-negative")
if hasattr(cv2, "ocl"):
cv2.ocl.setUseOpenCL(False)

rng = np.random.default_rng(115)
results: dict[str, object] = {}
for profile in selected:
count, repeats = PROFILES[profile]
centers = rng.uniform(0.0, 640.0, size=(count, 2)).astype(np.float32)
sizes = rng.uniform(5.0, 120.0, size=(count, 2)).astype(np.float32)
boxes_xyxy = np.empty((count, 4), dtype=np.float32)
boxes_xyxy[:, :2] = centers - sizes * 0.5
boxes_xyxy[:, 2:] = centers + sizes * 0.5
boxes_xywh = boxes_xyxy.copy()
boxes_xywh[:, 2:] -= boxes_xywh[:, :2]
scores = rng.random(count, dtype=np.float32)

def opencv_nms() -> np.ndarray:
return np.asarray(
cv2.dnn.NMSBoxes(boxes_xywh, scores, 0.25, 0.5)
).reshape(-1)

def spatialrust_nms() -> np.ndarray:
return sr.nms(boxes_xyxy, scores, 0.25, 0.5)

expected = opencv_nms().astype(np.int64, copy=False)
actual = spatialrust_nms()
exact = bool(np.array_equal(expected, actual))
if not exact:
raise AssertionError(f"{profile} NMS index mismatch")

_, _, opencv_timing, spatialrust_timing = timed_pair(
opencv_nms,
spatialrust_nms,
warmup=args.warmup,
repeats=repeats,
seed=117,
min_sample_time_ms=1.0,
)
opencv_ms = float(opencv_timing["median"])
spatialrust_ms = float(spatialrust_timing["median"])
results[profile] = {
"box_count": count,
"kept_count": int(actual.size),
"score_threshold": 0.25,
"iou_threshold": 0.5,
"indices_exact": exact,
"opencv": opencv_timing,
"spatialrust": spatialrust_timing,
"spatialrust_speedup": opencv_ms / spatialrust_ms,
"faster_implementation": (
"spatialrust" if spatialrust_ms < opencv_ms else "opencv"
),
}

receipt = environment(opencv_version=cv2.__version__, spatialrust_version=sr.__version__)
receipt["opencv_threads"] = cv2.getNumThreads()
receipt["opencv_opencl_enabled"] = bool(
hasattr(cv2, "ocl") and cv2.ocl.useOpenCL()
)
report = make_report(
suite="opencv-nms-performance",
kind="performance",
status="pass",
environment_receipt=receipt,
results={
"methodology": {
"timing_scope": "Python API call returning kept indices",
"paired_interleaved": True,
"input_seed": 115,
"random_order_seed": 117,
"minimum_sample_time_ms": 1.0,
"box_format": {
"opencv": "xywh float32 NumPy array",
"spatialrust": "xyxy float32 NumPy array",
},
"thread_policy": "library defaults; OpenCV thread count recorded",
},
"profiles": results,
},
)
emit_report(report, args.output)


if __name__ == "__main__":
main()
12 changes: 10 additions & 2 deletions crates/spatialrust-py/src/lib.rs
Original file line number Diff line number Diff line change
Expand Up @@ -3038,8 +3038,16 @@ fn nms<'py>(
native_boxes
.push(BoundingBox2::try_new(row[0], row[1], row[2], row[3]).map_err(to_py_err)?);
}
let scores: Vec<f32> = scores.as_array().iter().copied().collect();
let indices = nms_op(&native_boxes, &scores, score_threshold, iou_threshold)
let scores_view = scores.as_array();
let packed_scores;
let scores = match scores_view.as_slice() {
Some(scores) => scores,
None => {
packed_scores = scores_view.iter().copied().collect::<Vec<_>>();
packed_scores.as_slice()
}
};
let indices = nms_op(&native_boxes, scores, score_threshold, iou_threshold)
.map_err(to_py_err)?
.into_iter()
.map(|index| index as i64)
Expand Down
3 changes: 3 additions & 0 deletions crates/spatialrust-py/tests/test_bindings.py
Original file line number Diff line number Diff line change
Expand Up @@ -176,6 +176,9 @@ def test_detection_nms_and_soft_nms():
)
scores = np.array([0.9, 0.8, 0.7], dtype=np.float32)
np.testing.assert_array_equal(sr.nms(boxes, scores), [0, 2])
score_storage = np.empty(scores.size * 2, dtype=np.float32)
score_storage[::2] = scores
np.testing.assert_array_equal(sr.nms(boxes, score_storage[::2]), [0, 2])
indices, updated = sr.soft_nms(boxes, scores, method="linear")
assert indices[0] == 0
assert len(indices) == len(updated) == 3
Expand Down
5 changes: 5 additions & 0 deletions crates/spatialrust-vision/Cargo.toml
Original file line number Diff line number Diff line change
Expand Up @@ -73,6 +73,11 @@ name = "dense"
harness = false
required-features = ["dense"]

[[bench]]
name = "detection"
harness = false
required-features = ["detection"]

[[bench]]
name = "canny"
harness = false
Expand Down
55 changes: 55 additions & 0 deletions crates/spatialrust-vision/benches/detection.rs
Original file line number Diff line number Diff line change
@@ -0,0 +1,55 @@
use criterion::{black_box, criterion_group, criterion_main, BenchmarkId, Criterion, Throughput};
use spatialrust_vision::{nms, BoundingBox2};

fn benchmark_nms(c: &mut Criterion) {
let mut group = c.benchmark_group("nms_xyxy_f32");
group.sample_size(10);
for &count in &[100_usize, 1_000, 8_400] {
let (boxes, scores) = detections(count);
group.throughput(Throughput::Elements(count as u64));
group.bench_function(BenchmarkId::from_parameter(count), |b| {
b.iter(|| {
black_box(
nms(black_box(&boxes), black_box(&scores), black_box(0.25), black_box(0.5))
.unwrap(),
)
});
});
}
group.finish();
}

fn detections(count: usize) -> (Vec<BoundingBox2>, Vec<f32>) {
let mut state = 115_u64;
let mut boxes = Vec::with_capacity(count);
let mut scores = Vec::with_capacity(count);
for _ in 0..count {
let center_x = sample(&mut state) * 640.0;
let center_y = sample(&mut state) * 640.0;
let width = 5.0 + sample(&mut state) * 115.0;
let height = 5.0 + sample(&mut state) * 115.0;
boxes.push(
BoundingBox2::try_new(
center_x - width * 0.5,
center_y - height * 0.5,
center_x + width * 0.5,
center_y + height * 0.5,
)
.unwrap(),
);
scores.push(sample(&mut state));
}
(boxes, scores)
}

fn sample(state: &mut u64) -> f32 {
*state = state.wrapping_add(0x9E37_79B9_7F4A_7C15);
let mut value = *state;
value = (value ^ (value >> 30)).wrapping_mul(0xBF58_476D_1CE4_E5B9);
value = (value ^ (value >> 27)).wrapping_mul(0x94D0_49BB_1331_11EB);
value ^= value >> 31;
(value >> 40) as f32 / (1_u32 << 24) as f32
}

criterion_group!(benches, benchmark_nms);
criterion_main!(benches);
Loading
Loading