diff --git a/README.md b/README.md index bf6b7f1..80c9e4c 100644 --- a/README.md +++ b/README.md @@ -178,6 +178,20 @@ above gives SpatialRust a measured 1.07× 4K reuse lead, with maximum error zero VGA and 1080p remain narrow OpenCV wins. See the [acceleration receipt](notes/2026-07-15_exact_edt_acceleration.md). +For AI detection post-processing, the seeded Python NMS harness uses identical +float32 boxes, scores, and thresholds and requires exact kept-index parity +before publishing timings: + +| NMS candidates | OpenCV `dnn.NMSBoxes` | SpatialRust `nms` | Result | +| ---: | ---: | ---: | ---: | +| 100 | 0.298 ms | 0.033 ms | **SpatialRust 8.95×** | +| 1,000 | 8.720 ms | 2.286 ms | **SpatialRust 3.82×** | +| 8,400 (YOLO-style) | 407.086 ms | 126.562 ms | **SpatialRust 3.22×** | + +These Windows-host medians include each Python API call and returned indices; +see the [NMS harness](bench/opencv_nms_comparison/) and dated +[receipt](notes/2026-07-15_nms_opencv_acceleration.md). + #### Vision accuracy The same deterministic RGB inputs passed all VGA, 1080p, and 4K gates: @@ -204,6 +218,7 @@ On dense `H×W×3` XYZ (320×240, OpenCL off, local Windows laptop), `spatialrus python bench\opencv_vision_comparison\run.py python bench\opencv_vision_comparison\performance.py python bench\opencv_rgbd_comparison\run.py +python bench\opencv_nms_comparison\performance.py ``` ### Registration methods diff --git a/bench/opencv_comparison/README.md b/bench/opencv_comparison/README.md index cbc6a34..69c4e7a 100644 --- a/bench/opencv_comparison/README.md +++ b/bench/opencv_comparison/README.md @@ -32,6 +32,7 @@ VGA, 1080p, and 4K profiles and the initial competitive workload set: 10. colored RGB-D to point cloud 11. AI preprocessing 12. RGB-D to voxel end-to-end +13. detection NMS post-processing Exact matches use a JSON `null` PSNR (mathematically infinite) so reports remain strict RFC-compatible JSON. Numerical comparisons retain max/mean/RMS and @@ -49,6 +50,7 @@ then run both current suites: ```powershell python bench\opencv_comparison\run.py +python bench\opencv_nms_comparison\performance.py ``` Reports are written under `target/opencv-comparison/`. Run one suite with diff --git a/bench/opencv_comparison/manifest.json b/bench/opencv_comparison/manifest.json index d664787..c15dec3 100644 --- a/bench/opencv_comparison/manifest.json +++ b/bench/opencv_comparison/manifest.json @@ -47,6 +47,7 @@ { "id": "depth_to_xyz", "domain": "rgbd", "modes": ["allocate", "reuse"] }, { "id": "rgbd_to_point_cloud", "domain": "spatial-e2e", "modes": ["allocate"] }, { "id": "ai_preprocess", "domain": "dnn-adapter", "modes": ["allocate", "reuse"] }, + { "id": "nms", "domain": "dnn-adapter", "modes": ["postprocess"] }, { "id": "rgbd_to_voxel", "domain": "spatial-e2e", "modes": ["allocate"] } ] } diff --git a/bench/opencv_comparison/test_report.py b/bench/opencv_comparison/test_report.py index f7328cd..69f0e94 100644 --- a/bench/opencv_comparison/test_report.py +++ b/bench/opencv_comparison/test_report.py @@ -117,6 +117,7 @@ def test_manifest_reserves_representative_profiles_and_workloads(self) -> None: self.assertGreaterEqual(len(workloads), 10) self.assertIn("rgbd_to_voxel", workloads) self.assertIn("ai_preprocess", workloads) + self.assertIn("nms", workloads) self.assertIn("coefficient_of_variation", statistics) self.assertIn("median_absolute_deviation", statistics) self.assertIn("batch_size", statistics) diff --git a/bench/opencv_nms_comparison/README.md b/bench/opencv_nms_comparison/README.md new file mode 100644 index 0000000..52d3371 --- /dev/null +++ b/bench/opencv_nms_comparison/README.md @@ -0,0 +1,15 @@ +# OpenCV NMS comparison + +This harness compares SpatialRust `nms` with OpenCV `dnn.NMSBoxes` using the +same deterministic float32 boxes, scores, score threshold, and IoU threshold. +It covers small post-processing, 1,000-candidate, and YOLO-style 8,400-candidate +profiles. Returned indices must match exactly before timings are published. + +```powershell +python bench/opencv_nms_comparison/performance.py ` + --output target/opencv-nms-performance.json +``` + +The report follows `spatialrust.opencv-comparison.v1` and records raw samples, +dispersion, library versions, thread policy, and the host environment. Results +are machine-specific and must not be generalized beyond the named workload. diff --git a/bench/opencv_nms_comparison/performance.py b/bench/opencv_nms_comparison/performance.py new file mode 100644 index 0000000..e6d784b --- /dev/null +++ b/bench/opencv_nms_comparison/performance.py @@ -0,0 +1,124 @@ +"""Reproducible Python NMS performance and parity comparison with OpenCV.""" + +from __future__ import annotations + +import argparse +import sys +from pathlib import Path + +import cv2 +import numpy as np +import spatialrust as sr + +sys.path.insert(0, str(Path(__file__).resolve().parents[1])) +from opencv_comparison.report import emit_report, environment, make_report, timed_pair + + +PROFILES = { + "small_100": (100, 50), + "medium_1000": (1_000, 20), + "yolo_8400": (8_400, 8), +} + + +def parse_args() -> argparse.Namespace: + parser = argparse.ArgumentParser() + parser.add_argument("--output", type=Path) + parser.add_argument("--profiles", default=",".join(PROFILES)) + parser.add_argument("--warmup", type=int, default=3) + return parser.parse_args() + + +def main() -> None: + args = parse_args() + selected = [name.strip() for name in args.profiles.split(",") if name.strip()] + unknown = sorted(set(selected) - PROFILES.keys()) + if unknown: + raise ValueError(f"unknown profiles: {', '.join(unknown)}") + if args.warmup < 0: + raise ValueError("warmup must be non-negative") + if hasattr(cv2, "ocl"): + cv2.ocl.setUseOpenCL(False) + + rng = np.random.default_rng(115) + results: dict[str, object] = {} + for profile in selected: + count, repeats = PROFILES[profile] + centers = rng.uniform(0.0, 640.0, size=(count, 2)).astype(np.float32) + sizes = rng.uniform(5.0, 120.0, size=(count, 2)).astype(np.float32) + boxes_xyxy = np.empty((count, 4), dtype=np.float32) + boxes_xyxy[:, :2] = centers - sizes * 0.5 + boxes_xyxy[:, 2:] = centers + sizes * 0.5 + boxes_xywh = boxes_xyxy.copy() + boxes_xywh[:, 2:] -= boxes_xywh[:, :2] + scores = rng.random(count, dtype=np.float32) + + def opencv_nms() -> np.ndarray: + return np.asarray( + cv2.dnn.NMSBoxes(boxes_xywh, scores, 0.25, 0.5) + ).reshape(-1) + + def spatialrust_nms() -> np.ndarray: + return sr.nms(boxes_xyxy, scores, 0.25, 0.5) + + expected = opencv_nms().astype(np.int64, copy=False) + actual = spatialrust_nms() + exact = bool(np.array_equal(expected, actual)) + if not exact: + raise AssertionError(f"{profile} NMS index mismatch") + + _, _, opencv_timing, spatialrust_timing = timed_pair( + opencv_nms, + spatialrust_nms, + warmup=args.warmup, + repeats=repeats, + seed=117, + min_sample_time_ms=1.0, + ) + opencv_ms = float(opencv_timing["median"]) + spatialrust_ms = float(spatialrust_timing["median"]) + results[profile] = { + "box_count": count, + "kept_count": int(actual.size), + "score_threshold": 0.25, + "iou_threshold": 0.5, + "indices_exact": exact, + "opencv": opencv_timing, + "spatialrust": spatialrust_timing, + "spatialrust_speedup": opencv_ms / spatialrust_ms, + "faster_implementation": ( + "spatialrust" if spatialrust_ms < opencv_ms else "opencv" + ), + } + + receipt = environment(opencv_version=cv2.__version__, spatialrust_version=sr.__version__) + receipt["opencv_threads"] = cv2.getNumThreads() + receipt["opencv_opencl_enabled"] = bool( + hasattr(cv2, "ocl") and cv2.ocl.useOpenCL() + ) + report = make_report( + suite="opencv-nms-performance", + kind="performance", + status="pass", + environment_receipt=receipt, + results={ + "methodology": { + "timing_scope": "Python API call returning kept indices", + "paired_interleaved": True, + "input_seed": 115, + "random_order_seed": 117, + "minimum_sample_time_ms": 1.0, + "box_format": { + "opencv": "xywh float32 NumPy array", + "spatialrust": "xyxy float32 NumPy array", + }, + "thread_policy": "library defaults; OpenCV thread count recorded", + }, + "profiles": results, + }, + ) + emit_report(report, args.output) + + +if __name__ == "__main__": + main() diff --git a/crates/spatialrust-py/src/lib.rs b/crates/spatialrust-py/src/lib.rs index 80f2228..3a31cdf 100644 --- a/crates/spatialrust-py/src/lib.rs +++ b/crates/spatialrust-py/src/lib.rs @@ -3038,8 +3038,16 @@ fn nms<'py>( native_boxes .push(BoundingBox2::try_new(row[0], row[1], row[2], row[3]).map_err(to_py_err)?); } - let scores: Vec = scores.as_array().iter().copied().collect(); - let indices = nms_op(&native_boxes, &scores, score_threshold, iou_threshold) + let scores_view = scores.as_array(); + let packed_scores; + let scores = match scores_view.as_slice() { + Some(scores) => scores, + None => { + packed_scores = scores_view.iter().copied().collect::>(); + packed_scores.as_slice() + } + }; + let indices = nms_op(&native_boxes, scores, score_threshold, iou_threshold) .map_err(to_py_err)? .into_iter() .map(|index| index as i64) diff --git a/crates/spatialrust-py/tests/test_bindings.py b/crates/spatialrust-py/tests/test_bindings.py index 63b1bc0..4283eac 100644 --- a/crates/spatialrust-py/tests/test_bindings.py +++ b/crates/spatialrust-py/tests/test_bindings.py @@ -176,6 +176,9 @@ def test_detection_nms_and_soft_nms(): ) scores = np.array([0.9, 0.8, 0.7], dtype=np.float32) np.testing.assert_array_equal(sr.nms(boxes, scores), [0, 2]) + score_storage = np.empty(scores.size * 2, dtype=np.float32) + score_storage[::2] = scores + np.testing.assert_array_equal(sr.nms(boxes, score_storage[::2]), [0, 2]) indices, updated = sr.soft_nms(boxes, scores, method="linear") assert indices[0] == 0 assert len(indices) == len(updated) == 3 diff --git a/crates/spatialrust-vision/Cargo.toml b/crates/spatialrust-vision/Cargo.toml index bdcf7b2..36ab142 100644 --- a/crates/spatialrust-vision/Cargo.toml +++ b/crates/spatialrust-vision/Cargo.toml @@ -73,6 +73,11 @@ name = "dense" harness = false required-features = ["dense"] +[[bench]] +name = "detection" +harness = false +required-features = ["detection"] + [[bench]] name = "canny" harness = false diff --git a/crates/spatialrust-vision/benches/detection.rs b/crates/spatialrust-vision/benches/detection.rs new file mode 100644 index 0000000..154bee1 --- /dev/null +++ b/crates/spatialrust-vision/benches/detection.rs @@ -0,0 +1,55 @@ +use criterion::{black_box, criterion_group, criterion_main, BenchmarkId, Criterion, Throughput}; +use spatialrust_vision::{nms, BoundingBox2}; + +fn benchmark_nms(c: &mut Criterion) { + let mut group = c.benchmark_group("nms_xyxy_f32"); + group.sample_size(10); + for &count in &[100_usize, 1_000, 8_400] { + let (boxes, scores) = detections(count); + group.throughput(Throughput::Elements(count as u64)); + group.bench_function(BenchmarkId::from_parameter(count), |b| { + b.iter(|| { + black_box( + nms(black_box(&boxes), black_box(&scores), black_box(0.25), black_box(0.5)) + .unwrap(), + ) + }); + }); + } + group.finish(); +} + +fn detections(count: usize) -> (Vec, Vec) { + let mut state = 115_u64; + let mut boxes = Vec::with_capacity(count); + let mut scores = Vec::with_capacity(count); + for _ in 0..count { + let center_x = sample(&mut state) * 640.0; + let center_y = sample(&mut state) * 640.0; + let width = 5.0 + sample(&mut state) * 115.0; + let height = 5.0 + sample(&mut state) * 115.0; + boxes.push( + BoundingBox2::try_new( + center_x - width * 0.5, + center_y - height * 0.5, + center_x + width * 0.5, + center_y + height * 0.5, + ) + .unwrap(), + ); + scores.push(sample(&mut state)); + } + (boxes, scores) +} + +fn sample(state: &mut u64) -> f32 { + *state = state.wrapping_add(0x9E37_79B9_7F4A_7C15); + let mut value = *state; + value = (value ^ (value >> 30)).wrapping_mul(0xBF58_476D_1CE4_E5B9); + value = (value ^ (value >> 27)).wrapping_mul(0x94D0_49BB_1331_11EB); + value ^= value >> 31; + (value >> 40) as f32 / (1_u32 << 24) as f32 +} + +criterion_group!(benches, benchmark_nms); +criterion_main!(benches); diff --git a/crates/spatialrust-vision/src/detection.rs b/crates/spatialrust-vision/src/detection.rs index a9860b0..7beaa22 100644 --- a/crates/spatialrust-vision/src/detection.rs +++ b/crates/spatialrust-vision/src/detection.rs @@ -172,10 +172,13 @@ pub fn nms( .filter_map(|(index, &score)| (score >= score_threshold).then_some(index)) .collect(); sort_indices_by_score(&mut order, scores); + let areas: Vec = boxes.iter().copied().map(BoundingBox2::area).collect(); let mut keep: Vec = Vec::with_capacity(order.len()); 'candidate: for index in order { for &selected in &keep { - if boxes[index].iou(boxes[selected]) > iou_threshold { + if iou_with_areas(boxes[index], areas[index], boxes[selected], areas[selected]) + > iou_threshold + { continue 'candidate; } } @@ -211,11 +214,17 @@ pub fn batched_nms( .filter_map(|(index, &score)| (score >= score_threshold).then_some(index)) .collect(); sort_indices_by_score(&mut order, &scores); + let areas: Vec = detections.iter().map(|detection| detection.bbox.area()).collect(); let mut keep: Vec = Vec::with_capacity(order.len()); 'candidate: for index in order { for &selected in &keep { if detections[index].class_id == detections[selected].class_id - && detections[index].bbox.iou(detections[selected].bbox) > iou_threshold + && iou_with_areas( + detections[index].bbox, + areas[index], + detections[selected].bbox, + areas[selected], + ) > iou_threshold { continue 'candidate; } @@ -245,6 +254,7 @@ pub fn soft_nms( .enumerate() .map(|(index, score)| ScoredIndex { index, score }) .collect(); + let areas: Vec = boxes.iter().copied().map(BoundingBox2::area).collect(); let mut output = Vec::new(); while !candidates.is_empty() { candidates.sort_by(score_order); @@ -254,7 +264,12 @@ pub fn soft_nms( } output.push(selected); for candidate in &mut candidates { - let overlap = boxes[selected.index].iou(boxes[candidate.index]); + let overlap = iou_with_areas( + boxes[selected.index], + areas[selected.index], + boxes[candidate.index], + areas[candidate.index], + ); let weight = match method { SoftNmsMethod::Hard => { if overlap <= iou_threshold { @@ -279,6 +294,18 @@ pub fn soft_nms( Ok(output) } +fn iou_with_areas(left: BoundingBox2, left_area: f32, right: BoundingBox2, right_area: f32) -> f32 { + let intersection_width = (left.x_max.min(right.x_max) - left.x_min.max(right.x_min)).max(0.0); + let intersection_height = (left.y_max.min(right.y_max) - left.y_min.max(right.y_min)).max(0.0); + let intersection = intersection_width * intersection_height; + let union = left_area + right_area - intersection; + if union > 0.0 { + intersection / union + } else { + 0.0 + } +} + fn validate_nms_inputs( boxes: &[BoundingBox2], scores: &[f32], @@ -325,7 +352,9 @@ fn score_order(left: &ScoredIndex, right: &ScoredIndex) -> Ordering { #[cfg(test)] mod tests { - use super::{batched_nms, nms, soft_nms, BoundingBox2, Detection, SoftNmsMethod}; + use super::{ + batched_nms, iou_with_areas, nms, soft_nms, BoundingBox2, Detection, SoftNmsMethod, + }; fn bbox(x0: f32, y0: f32, x1: f32, y1: f32) -> BoundingBox2 { BoundingBox2::try_new(x0, y0, x1, y1).unwrap() @@ -345,6 +374,21 @@ mod tests { assert_eq!(nms(&boxes, &[0.9, 0.8, 0.7], 0.0, 0.5).unwrap(), vec![0, 2]); } + #[test] + fn cached_area_iou_matches_public_iou() { + let boxes = [ + BoundingBox2::try_new(0.0, 0.0, 20.0, 20.0).unwrap(), + BoundingBox2::try_new(2.0, 2.0, 18.0, 18.0).unwrap(), + BoundingBox2::try_new(30.0, 30.0, 45.0, 45.0).unwrap(), + BoundingBox2::try_new(5.0, 5.0, 5.0, 8.0).unwrap(), + ]; + for &left in &boxes { + for &right in &boxes { + assert_eq!(iou_with_areas(left, left.area(), right, right.area()), left.iou(right),); + } + } + } + #[test] fn batched_nms_keeps_overlapping_different_classes() { let detections = [ diff --git a/docs/ROADMAP.md b/docs/ROADMAP.md index 9cacee2..47d3c33 100644 --- a/docs/ROADMAP.md +++ b/docs/ROADMAP.md @@ -618,6 +618,7 @@ to one implicitly, and GPU receipts must retain named upload/readback stages. | 114D | Planned | Bound worker creation and temporary memory | thread-count and peak-memory receipt | | 114E | Complete | Exact EDT binary-row fast path, tiled transpose, and bounded pool dispatch | VGA/1080p/4K Criterion and OpenCV receipt | | 114F | Complete | Cache EDT parabola heights and balance column tasks for dense masks | exact OpenCV parity and 4K Python reuse win | +| 114G | Complete | Cache NMS box geometry and avoid packed Python score copies | exact OpenCV index parity and 100/1,000/8,400-candidate wins | ### Epic 115 delivery slices diff --git a/docs/site/algorithms.html b/docs/site/algorithms.html index ac52cc0..176870d 100644 --- a/docs/site/algorithms.html +++ b/docs/site/algorithms.html @@ -35,7 +35,8 @@

Algorithm catalog

MorphologyErode, dilate, open, close, gradient, top-hat, black-hat with Rect/Cross/Ellipse/Diamond/custom elementsspatialrust-vision · imgproc-morphologyCPU / GPU Image analysisFixed/Otsu/adaptive threshold, histogram, equalization, CLAHE, integral image, Cannyspatialrust-vision · imgproc-analysis, imgproc-cannyCPU Local featuresFAST, Harris, Shi–Tomasi, ORB, descriptor matching, grid selection, pyramidal Lucas–Kanade trackingspatialrust-vision · feature2dCPU - Dense visionExact Euclidean distance transform with reusable workspace/output, connected components, contours, polygon approximation, mask RLE, depth/confidence/flow/point maps, NMS/Soft-NMSspatialrust-vision · dense, detectionCPU parallel + Dense visionExact Euclidean distance transform with reusable workspace/output, connected components, contours, polygon approximation, mask RLE, depth/confidence/flow/point mapsspatialrust-vision · denseCPU parallel + Detection post-processingIoU/GIoU, greedy NMS, class-aware batched NMS, and hard/linear/Gaussian Soft-NMS. Seeded Python NMS is 3.22×–8.95× faster than OpenCV on the recorded 100/1,000/8,400-candidate workloads with exact kept indices. Harness and methodology.spatialrust-vision · detectionCPU Multiview geometryHomography, fundamental/essential matrices, RANSAC, triangulation, relative pose, PnP/PnP-RANSACspatialrust-vision · geometryCPU Stereo and odometryStereo rectification, block matching, disparity-to-depth/XYZ, monocular and RGB-D visual odometryspatialrust-vision · geometry, odometryCPU VideoDense block flow, adaptive background model, multi-object tracker, timestamped pull sourcesspatialrust-vision · videoCPU diff --git a/docs/site/vision2.html b/docs/site/vision2.html index e666b91..89da5ac 100644 --- a/docs/site/vision2.html +++ b/docs/site/vision2.html @@ -51,6 +51,7 @@

Latest measured outcome

Exact EDT at 4K

With caller-owned output and workspace reuse, SpatialRust measured 40.66 ms versus OpenCV 43.33 ms: a 1.07× lead on the recorded Windows host.

Accuracy preserved

The VGA, 1080p, and 4K canonical masks retain maximum absolute error 0.0 against OpenCV's precise L2 distance transform.

+

NMS post-processing

SpatialRust measured 3.22× faster for 8,400 YOLO-style candidates and up to 8.95× faster for 100 candidates, with kept indices exactly matching OpenCV.

This is a workload- and host-specific result. VGA and 1080p reuse remain narrow OpenCV wins; the repository receipt contains the reproducible methodology.

diff --git a/notes/2026-07-15_nms_opencv_acceleration.md b/notes/2026-07-15_nms_opencv_acceleration.md new file mode 100644 index 0000000..4b0f168 --- /dev/null +++ b/notes/2026-07-15_nms_opencv_acceleration.md @@ -0,0 +1,54 @@ +# NMS OpenCV acceleration receipt — 2026-07-15 + +## Outcome + +SpatialRust greedy NMS now precomputes each valid box area once and reuses it +for all overlap comparisons. The same cached geometry is used by class-aware +batched NMS and Soft-NMS. The Python binding borrows contiguous float32 scores +directly and only packs non-contiguous inputs. + +This preserves the public safe API, score-descending deterministic ordering, +half-open `xyxy` box semantics, generic non-contiguous NumPy fallback, and the +existing score/IoU validation contract. + +## Correctness contract + +The comparison uses OpenCV [`dnn.NMSBoxes`](https://docs.opencv.org/master/df/d57/namespacecv_1_1dnn.html) +as the oracle. Both implementations receive the same deterministic float32 +coordinates and scores (represented as corresponding `xyxy` and `xywh` boxes), score +threshold `0.25`, and IoU threshold `0.5`. + +All three profiles returned exactly the same ordered kept indices: + +| Profile | Candidates | Kept | Exact indices | +| --- | ---: | ---: | ---: | +| Small post-process | 100 | 71 | yes | +| Medium post-process | 1,000 | 654 | yes | +| YOLO-style output | 8,400 | 3,675 | yes | + +Rust coverage also checks cached-area IoU bit-for-bit against the public IoU +implementation for overlapping, disjoint, identical, and zero-area boxes. + +## Performance + +Host: Windows 11, Intel 6-core/12-thread CPU, CPython 3.12.10, OpenCV 4.10, +OpenCL disabled. Timings are randomized, interleaved Python API medians after +three warmups and include the returned index array. + +| Candidates | Repeats | OpenCV | SpatialRust | SpatialRust speedup | +| ---: | ---: | ---: | ---: | ---: | +| 100 | 50 | 0.2978 ms | 0.0333 ms | **8.95×** | +| 1,000 | 20 | 8.7205 ms | 2.2856 ms | **3.82×** | +| 8,400 | 8 | 407.0856 ms | 126.5618 ms | **3.22×** | + +Native Criterion medians for a separately seeded workload were approximately +17.1 microseconds, 1.17 milliseconds, and 50.3 milliseconds respectively. +These results are host- and workload-specific, not universal performance +claims. + +Reproduce the machine-readable report with: + +```powershell +python bench/opencv_nms_comparison/performance.py ` + --output target/opencv-nms-performance.json +```