diff --git a/README.md b/README.md index 052b349..a5b6c09 100644 --- a/README.md +++ b/README.md @@ -222,6 +222,26 @@ All profiles exactly matched OpenCV's kept-index order; updated float32 scores stayed within `1.79e-7`. See the [Soft-NMS harness](bench/opencv_soft_nms_comparison/) and dated [receipt](notes/2026-07-15_soft_nms_opencv_acceleration.md). +Connected-component labeling uses horizontal runs plus union-find instead of +per-pixel flood fill. Packed NumPy masks are borrowed without an input copy, +and all non-zero `uint8` values are foreground, matching OpenCV. Against +OpenCV 4.13's explicit row-major SAUF algorithm on structured masks: + +| Profile | Pattern | OpenCV SAUF | SpatialRust | Result | +| --- | --- | ---: | ---: | ---: | +| VGA | Segmentation blobs | 1.284 ms | 0.413 ms | **SpatialRust 3.11×** | +| VGA | Document lines | 1.271 ms | 0.352 ms | **SpatialRust 3.61×** | +| 1080p | Segmentation blobs | 6.763 ms | 2.815 ms | **SpatialRust 2.40×** | +| 1080p | Document lines | 6.649 ms | 2.407 ms | **SpatialRust 2.76×** | +| 4K | Segmentation blobs | 21.356 ms | 9.838 ms | **SpatialRust 2.17×** | +| 4K | Document lines | 21.075 ms | 8.606 ms | **SpatialRust 2.45×** | + +Labels, areas, and bounding boxes matched exactly on every canonical profile +and 320 additional seeded randomized 4/8-connectivity cases. The speed claim is +limited to the named structured masks; dense random noise still favors OpenCV. +See the [connected-components harness](bench/opencv_connected_components_comparison/) +and dated [receipt](notes/2026-07-15_connected_components_opencv_acceleration.md). + #### Vision accuracy The same deterministic RGB inputs passed all VGA, 1080p, and 4K gates: @@ -236,6 +256,7 @@ The same deterministic RGB inputs passed all VGA, 1080p, and 4K gates: | Morphology open 5×5 | Exact pixels (max error 0) | | Canny | Precision, recall, F1, and IoU all 1.0 | | Exact Euclidean distance transform | Exact values on canonical profiles; separate irregular-mask max float error `9.54e-7` | +| Connected components (SAUF ordering) | Exact labels, areas, and bounding boxes on structured profiles and 320 randomized cases | The broader correctness harness also checks filters, analysis, keypoints, matching, and geometry with documented tolerances (exact pixels where we claim diff --git a/bench/opencv_comparison/README.md b/bench/opencv_comparison/README.md index a1e5e68..70caa4a 100644 --- a/bench/opencv_comparison/README.md +++ b/bench/opencv_comparison/README.md @@ -33,6 +33,7 @@ VGA, 1080p, and 4K profiles and the initial competitive workload set: 11. AI preprocessing 12. RGB-D to voxel end-to-end 13. detection NMS, class-aware batched NMS, and Soft-NMS post-processing +14. connected components on structured segmentation and document masks Exact matches use a JSON `null` PSNR (mathematically infinite) so reports remain strict RFC-compatible JSON. Numerical comparisons retain max/mean/RMS and @@ -53,6 +54,7 @@ python bench\opencv_comparison\run.py python bench\opencv_nms_comparison\performance.py python bench\opencv_batched_nms_comparison\performance.py python bench\opencv_soft_nms_comparison\performance.py +python bench\opencv_connected_components_comparison\performance.py ``` Reports are written under `target/opencv-comparison/`. Run one suite with diff --git a/bench/opencv_comparison/manifest.json b/bench/opencv_comparison/manifest.json index 0d3c9ba..854d486 100644 --- a/bench/opencv_comparison/manifest.json +++ b/bench/opencv_comparison/manifest.json @@ -33,6 +33,7 @@ { "id": "canny", "domain": "imgproc", "modes": ["allocate"] }, { "id": "morphology_open", "domain": "imgproc", "modes": ["allocate"] }, { "id": "distance_transform_edt", "domain": "imgproc", "modes": ["allocate"] }, + { "id": "connected_components", "domain": "imgproc", "modes": ["structured-mask"] }, { "id": "orb", "domain": "feature2d", "modes": ["allocate"] }, { "id": "stereo_bm", "domain": "calib3d", "modes": ["allocate"] }, { "id": "pinhole_calibration", "domain": "calib3d", "modes": ["allocate"] }, diff --git a/bench/opencv_comparison/test_report.py b/bench/opencv_comparison/test_report.py index 01753c0..a60c532 100644 --- a/bench/opencv_comparison/test_report.py +++ b/bench/opencv_comparison/test_report.py @@ -120,6 +120,7 @@ def test_manifest_reserves_representative_profiles_and_workloads(self) -> None: self.assertIn("nms", workloads) self.assertIn("batched_nms", workloads) self.assertIn("soft_nms", workloads) + self.assertIn("connected_components", workloads) self.assertIn("coefficient_of_variation", statistics) self.assertIn("median_absolute_deviation", statistics) self.assertIn("batch_size", statistics) diff --git a/bench/opencv_connected_components_comparison/README.md b/bench/opencv_connected_components_comparison/README.md new file mode 100644 index 0000000..2309c6d --- /dev/null +++ b/bench/opencv_connected_components_comparison/README.md @@ -0,0 +1,17 @@ +# OpenCV connected-components comparison + +This harness compares SpatialRust connected-component labeling with OpenCV's +explicit SAUF implementation on structured segmentation-blob and document-line +masks. SAUF is selected because OpenCV documents it as the algorithm that +guarantees row-major label ordering, matching SpatialRust's public contract. + +```powershell +python bench/opencv_connected_components_comparison/performance.py ` + --output target/opencv-connected-components-performance.json +``` + +Labels, component areas, and half-open bounding boxes must match exactly before +timings are published. The report follows `spatialrust.opencv-comparison.v1` +and records raw samples, dispersion, library versions, OpenCV thread policy, +and the host environment. Calls are batched to at least 20 ms to stabilize +allocator-heavy short profiles. Results are scoped to the named structured masks. diff --git a/bench/opencv_connected_components_comparison/performance.py b/bench/opencv_connected_components_comparison/performance.py new file mode 100644 index 0000000..f1ab7d5 --- /dev/null +++ b/bench/opencv_connected_components_comparison/performance.py @@ -0,0 +1,194 @@ +"""Reproducible structured-mask connected-components comparison with OpenCV.""" + +from __future__ import annotations + +import argparse +import sys +from pathlib import Path + +import cv2 +import numpy as np +import spatialrust as sr + +sys.path.insert(0, str(Path(__file__).resolve().parents[1])) +from opencv_comparison.report import emit_report, environment, make_report, timed_pair + + +PROFILES = { + "vga": (640, 480, 40), + "1080p": (1920, 1080, 30), + "4k": (3840, 2160, 20), +} +PATTERNS = ("segmentation_blobs", "document_lines") + + +def parse_args() -> argparse.Namespace: + parser = argparse.ArgumentParser() + parser.add_argument("--output", type=Path) + parser.add_argument("--profiles", default=",".join(PROFILES)) + parser.add_argument("--patterns", default=",".join(PATTERNS)) + parser.add_argument("--warmup", type=int, default=8) + return parser.parse_args() + + +def make_mask(width: int, height: int, pattern: str) -> np.ndarray: + x = np.arange(width, dtype=np.int32)[None, :] + y = np.arange(height, dtype=np.int32)[:, None] + if pattern == "segmentation_blobs": + foreground = (x % 97 < 23) & (y % 83 < 19) + elif pattern == "document_lines": + foreground = (y % 32 < 3) & (x % 211 > 8) & (x % 211 < 190) + else: + raise ValueError(f"unknown pattern: {pattern}") + return foreground.astype(np.uint8) * np.uint8(255) + + +def assert_sauf_compatible( + mask: np.ndarray, connectivity: int, context: str +) -> tuple[int, np.ndarray, np.ndarray]: + count, expected_labels, expected_stats, _ = ( + cv2.connectedComponentsWithStatsWithAlgorithm( + mask, connectivity, cv2.CV_32S, cv2.CCL_SAUF + ) + ) + actual_labels, actual_stats = sr.connected_components_image( + mask, connectivity=connectivity + ) + labels_exact = bool(np.array_equal(actual_labels, expected_labels)) + actual_areas = np.asarray([value[1] for value in actual_stats], dtype=np.int64) + expected_areas = expected_stats[1:count, cv2.CC_STAT_AREA].astype(np.int64) + areas_exact = bool(np.array_equal(actual_areas, expected_areas)) + actual_boxes = np.asarray([value[2] for value in actual_stats], dtype=np.float64) + expected_boxes = expected_stats[1:count, :4].astype(np.float64) + expected_boxes[:, 2] += expected_boxes[:, 0] + expected_boxes[:, 3] += expected_boxes[:, 1] + boxes_exact = bool(np.array_equal(actual_boxes, expected_boxes)) + if not labels_exact or not areas_exact or not boxes_exact: + raise AssertionError( + f"{context} mismatch: labels={labels_exact}, " + f"areas={areas_exact}, boxes={boxes_exact}" + ) + return count, expected_labels, expected_stats + + +def validate_randomized_cases(cases_per_connectivity: int = 160) -> int: + rng = np.random.default_rng(20_260_715) + checked = 0 + for connectivity in (4, 8): + for case in range(cases_per_connectivity): + height = int(rng.integers(1, 90)) + width = int(rng.integers(1, 120)) + if case % 4 == 0: + density = float(rng.uniform(0.01, 0.8)) + mask = (rng.random((height, width)) < density).astype(np.uint8) * 255 + else: + mask = np.zeros((height, width), dtype=np.uint8) + for _ in range(int(rng.integers(1, 25))): + y0 = int(rng.integers(height)) + y1 = int(rng.integers(y0 + 1, height + 1)) + x0 = int(rng.integers(width)) + x1 = int(rng.integers(x0 + 1, width + 1)) + mask[y0:y1, x0:x1] = int(rng.integers(1, 256)) + assert_sauf_compatible(mask, connectivity, f"random/{connectivity}/{case}") + checked += 1 + return checked + + +def main() -> None: + args = parse_args() + profiles = [name.strip() for name in args.profiles.split(",") if name.strip()] + patterns = [name.strip() for name in args.patterns.split(",") if name.strip()] + unknown_profiles = sorted(set(profiles) - PROFILES.keys()) + unknown_patterns = sorted(set(patterns) - set(PATTERNS)) + if unknown_profiles: + raise ValueError(f"unknown profiles: {', '.join(unknown_profiles)}") + if unknown_patterns: + raise ValueError(f"unknown patterns: {', '.join(unknown_patterns)}") + if args.warmup < 0: + raise ValueError("warmup must be non-negative") + if not hasattr(cv2, "connectedComponentsWithStatsWithAlgorithm"): + raise RuntimeError("OpenCV build does not expose explicit CCL algorithms") + if hasattr(cv2, "ocl"): + cv2.ocl.setUseOpenCL(False) + + randomized_cases = validate_randomized_cases() + + results: dict[str, object] = {} + for profile in profiles: + width, height, repeats = PROFILES[profile] + profile_results: dict[str, object] = {} + for pattern in patterns: + mask = make_mask(width, height, pattern) + + def opencv_components() -> tuple[object, ...]: + return cv2.connectedComponentsWithStatsWithAlgorithm( + mask, 8, cv2.CV_32S, cv2.CCL_SAUF + ) + + def spatialrust_components() -> tuple[object, ...]: + return sr.connected_components_image(mask, connectivity=8) + + count, _, _ = assert_sauf_compatible(mask, 8, f"{profile}/{pattern}") + labels_exact = True + areas_exact = True + boxes_exact = True + + _, _, opencv_timing, spatialrust_timing = timed_pair( + opencv_components, + spatialrust_components, + warmup=args.warmup, + repeats=repeats, + seed=173, + min_sample_time_ms=20.0, + ) + opencv_ms = float(opencv_timing["median"]) + spatialrust_ms = float(spatialrust_timing["median"]) + profile_results[pattern] = { + "width": width, + "height": height, + "connectivity": 8, + "component_count": int(count - 1), + "labels_exact": labels_exact, + "areas_exact": areas_exact, + "bounding_boxes_exact": boxes_exact, + "opencv": opencv_timing, + "spatialrust": spatialrust_timing, + "spatialrust_speedup": opencv_ms / spatialrust_ms, + "faster_implementation": ( + "spatialrust" if spatialrust_ms < opencv_ms else "opencv" + ), + } + results[profile] = profile_results + + receipt = environment( + opencv_version=cv2.__version__, spatialrust_version=sr.__version__ + ) + receipt["opencv_threads"] = cv2.getNumThreads() + receipt["opencv_opencl_enabled"] = bool( + hasattr(cv2, "ocl") and cv2.ocl.useOpenCL() + ) + report = make_report( + suite="opencv-connected-components-performance", + kind="performance", + status="pass", + environment_receipt=receipt, + results={ + "methodology": { + "timing_scope": "Python API call returning labels and foreground statistics", + "paired_interleaved": True, + "random_order_seed": 173, + "minimum_sample_time_ms": 20.0, + "opencv_algorithm": "CCL_SAUF", + "mask_values": "packed uint8; 0 background, 255 foreground", + "randomized_correctness_seed": 20_260_715, + "randomized_correctness_cases": randomized_cases, + "thread_policy": "library defaults; OpenCV thread count recorded", + }, + "profiles": results, + }, + ) + emit_report(report, args.output) + + +if __name__ == "__main__": + main() diff --git a/crates/spatialrust-py/README.md b/crates/spatialrust-py/README.md index 734cdc8..cb89e9d 100644 --- a/crates/spatialrust-py/README.md +++ b/crates/spatialrust-py/README.md @@ -108,7 +108,7 @@ reloaded = sr.read("labeled.las") | `resize_image` / `letterbox_image` / `normalize_image_chw` | Model-ready RGB resize, padding, and float32 CHW packing | | `rgb_to_gray_image` / `rgb_to_hsv_image` / `remap_image` | CPU color conversion and coordinate-map resampling | | `nms` / `batched_nms` / `soft_nms` | Detection post-processing for `(N,4)` xyxy boxes, including class-aware suppression | -| `connected_components_image` / `find_mask_contours` | Binary-mask labeling and contour extraction | +| `connected_components_image` / `find_mask_contours` | Row-major 4/8-connected labeling (borrowed packed input; any non-zero byte is foreground) and contour extraction | | `encode_mask_rle` / `decode_mask_rle` | Row-major or COCO column-major binary-mask RLE | | `point_map_to_point_cloud` | Filter a dense `(H,W,3)` point map into a native point cloud | | `knn_graph(cloud, k)` / `radius_graph(cloud, radius)` | PyG-style `(2, E)` `edge_index` for GNNs | diff --git a/crates/spatialrust-py/src/lib.rs b/crates/spatialrust-py/src/lib.rs index f734efc..aa1a350 100644 --- a/crates/spatialrust-py/src/lib.rs +++ b/crates/spatialrust-py/src/lib.rs @@ -72,9 +72,10 @@ use spatialrust::transform::{ use spatialrust::vision::{ adaptive_threshold as adaptive_threshold_op, approximate_polygon as approximate_contour, batched_nms as batched_nms_op, bilateral_filter as bilateral_filter_op, canny as canny_op, - clahe as clahe_op, connected_components as label_components, decode_rle as decode_mask_runs, - detect_and_describe_orb as detect_and_describe_orb_op, detect_fast as detect_fast_op, - detect_harris as detect_harris_op, detect_shi_tomasi as detect_shi_tomasi_op, + clahe as clahe_op, connected_components_u8 as label_components_u8, + decode_rle as decode_mask_runs, detect_and_describe_orb as detect_and_describe_orb_op, + detect_fast as detect_fast_op, detect_harris as detect_harris_op, + detect_shi_tomasi as detect_shi_tomasi_op, distance_transform_edt_u8_into as distance_transform_edt_u8_into_op, distance_transform_edt_with_spacing as distance_transform_edt_op, encode_rle as encode_mask_runs, equalize_histogram as equalize_histogram_op, @@ -3159,15 +3160,23 @@ fn connected_components_image<'py>( mask: PyReadonlyArray2<'_, u8>, connectivity: u8, ) -> PyResult<(Bound<'py, PyArray2>, ComponentStats)> { - let image = gray_u8_image_from_numpy(mask)?; - let mask = - BinaryMask::try_new(image.width(), image.height(), image.into_vec()).map_err(to_py_err)?; + let view = mask.as_array(); + let shape = view.shape(); + let (height, width) = (shape[0], shape[1]); + let packed; + let pixels = match mask.as_slice() { + Ok(slice) => slice, + Err(_) => { + packed = view.iter().copied().collect::>(); + packed.as_slice() + } + }; let connectivity = match connectivity { 4 => Connectivity::Four, 8 => Connectivity::Eight, _ => return Err(PyValueError::new_err("connectivity must be 4 or 8")), }; - let result = label_components(&mask, connectivity).map_err(to_py_err)?; + let result = label_components_u8(width, height, pixels, connectivity).map_err(to_py_err)?; let stats = result .components .iter() @@ -3184,11 +3193,8 @@ fn connected_components_image<'py>( ) }) .collect(); - let labels = Array2::from_shape_vec( - (result.labels.height(), result.labels.width()), - result.labels.as_slice().to_vec(), - ) - .map_err(to_py_err)?; + let labels = Array2::from_shape_vec((height, width), result.labels.into_image().into_vec()) + .map_err(to_py_err)?; Ok((labels.into_pyarray_bound(py), stats)) } diff --git a/crates/spatialrust-py/tests/test_bindings.py b/crates/spatialrust-py/tests/test_bindings.py index 7604164..076b02c 100644 --- a/crates/spatialrust-py/tests/test_bindings.py +++ b/crates/spatialrust-py/tests/test_bindings.py @@ -206,6 +206,12 @@ def test_mask_components_contours_and_rle(): assert labels.shape == mask.shape assert labels.dtype == np.uint32 assert sorted(stat[1] for stat in stats) == [4, 4] + mask_255 = mask * np.uint8(255) + storage = np.zeros((5, 14), dtype=np.uint8) + storage[:, ::2] = mask_255 + view_labels, view_stats = sr.connected_components_image(storage[:, ::2], connectivity=4) + np.testing.assert_array_equal(view_labels, labels) + assert view_stats == stats contours = sr.find_mask_contours(mask) assert len(contours) == 2 diff --git a/crates/spatialrust-vision/benches/dense.rs b/crates/spatialrust-vision/benches/dense.rs index 41e6a0a..b36c607 100644 --- a/crates/spatialrust-vision/benches/dense.rs +++ b/crates/spatialrust-vision/benches/dense.rs @@ -1,6 +1,7 @@ use criterion::{black_box, criterion_group, criterion_main, BenchmarkId, Criterion, Throughput}; use spatialrust_vision::{ - distance_transform_edt, distance_transform_edt_into, BinaryMask, DistanceTransformWorkspace, + connected_components, distance_transform_edt, distance_transform_edt_into, BinaryMask, + Connectivity, DistanceTransformWorkspace, }; fn benchmark_exact_distance_transform(c: &mut Criterion) { @@ -61,5 +62,33 @@ fn benchmark_exact_distance_transform(c: &mut Criterion) { group.finish(); } -criterion_group!(benches, benchmark_exact_distance_transform); +fn benchmark_connected_components(c: &mut Criterion) { + let mut group = c.benchmark_group("connected_components_8_structured"); + group.sample_size(10); + for &(profile, width, height) in &[("vga", 640, 480), ("1080p", 1920, 1080), ("4k", 3840, 2160)] + { + for pattern in ["segmentation-blobs", "document-lines"] { + let data = (0..width * height) + .map(|index| { + let (x, y) = (index % width, index / width); + match pattern { + "segmentation-blobs" => u8::from(x % 97 < 23 && y % 83 < 19), + "document-lines" => u8::from(y % 32 < 3 && x % 211 > 8 && x % 211 < 190), + _ => unreachable!(), + } + }) + .collect(); + let mask = BinaryMask::try_new(width, height, data).unwrap(); + group.throughput(Throughput::Elements((width * height) as u64)); + group.bench_function(BenchmarkId::new(pattern, profile), |b| { + b.iter(|| { + black_box(connected_components(black_box(&mask), Connectivity::Eight).unwrap()) + }); + }); + } + } + group.finish(); +} + +criterion_group!(benches, benchmark_exact_distance_transform, benchmark_connected_components); criterion_main!(benches); diff --git a/crates/spatialrust-vision/src/dense.rs b/crates/spatialrust-vision/src/dense.rs index c00adc4..c7b8a9b 100644 --- a/crates/spatialrust-vision/src/dense.rs +++ b/crates/spatialrust-vision/src/dense.rs @@ -1,6 +1,6 @@ //! Dense mask, depth, flow, confidence, and point-map primitives. -use std::collections::{BTreeMap, VecDeque}; +use std::collections::BTreeMap; use rayon::prelude::*; use spatialrust_image::{ColorSpace, Image, ImageMetadata, ImageView}; @@ -801,6 +801,12 @@ impl LabelImage { pub fn image(&self) -> &Image { &self.image } + + /// Consumes the wrapper and returns its image. + #[must_use] + pub fn into_image(self) -> Image { + self.image + } } /// Per-component statistics. @@ -830,82 +836,191 @@ pub fn connected_components( mask: &BinaryMask, connectivity: Connectivity, ) -> VisionResult { - let width = mask.width(); - let height = mask.height(); + connected_components_u8(mask.width(), mask.height(), mask.image.as_slice(), connectivity) +} + +/// Labels connected non-zero pixels from packed `u8` mask storage. +/// +/// This borrowed-input variant avoids constructing an owned [`BinaryMask`] +/// when the caller already has packed mask storage. Like OpenCV, every +/// non-zero byte is foreground. Positive labels follow the first foreground +/// run of each component in row-major order. +pub fn connected_components_u8( + width: usize, + height: usize, + pixels: &[u8], + connectivity: Connectivity, +) -> VisionResult { + let expected = width + .checked_mul(height) + .ok_or_else(|| VisionError::InvalidDimensions("mask dimensions overflow".to_owned()))?; + if pixels.len() != expected { + return Err(VisionError::InvalidDimensions(format!( + "mask storage has {} pixels, expected {expected}", + pixels.len() + ))); + } let mut labels = vec![0_u32; width * height]; - let mut components = Vec::new(); - let mut queue = VecDeque::new(); - let mut next_label = 1_u32; - for seed_y in 0..height { - for seed_x in 0..width { - let seed = seed_y * width + seed_x; - if !mask.contains(seed_x, seed_y) || labels[seed] != 0 { - continue; + let mut runs = Vec::new(); + let mut previous_runs = 0..0; + + for y in 0..height { + let current_start = runs.len(); + let row = &pixels[y * width..(y + 1) * width]; + let mut x = 0; + let mut previous_cursor = previous_runs.start; + while x < width { + while x < width && row[x] == 0 { + x += 1; } - labels[seed] = next_label; - queue.push_back((seed_x, seed_y)); - let mut area = 0_usize; - let mut min_x = seed_x; - let mut min_y = seed_y; - let mut max_x = seed_x; - let mut max_y = seed_y; - let mut sum_x = 0.0_f64; - let mut sum_y = 0.0_f64; - while let Some((x, y)) = queue.pop_front() { - area += 1; - min_x = min_x.min(x); - min_y = min_y.min(y); - max_x = max_x.max(x); - max_y = max_y.max(y); - sum_x += x as f64 + 0.5; - sum_y += y as f64 + 0.5; - for (nx, ny) in neighbors(x, y, width, height, connectivity) { - let index = ny * width + nx; - if mask.contains(nx, ny) && labels[index] == 0 { - labels[index] = next_label; - queue.push_back((nx, ny)); - } - } + if x == width { + break; + } + let start = x; + while x < width && row[x] != 0 { + x += 1; + } + let end = x; + let run_index = runs.len(); + runs.push(ComponentRun { y, start, end, parent: run_index }); + + while previous_cursor < previous_runs.end + && run_precedes(runs[previous_cursor], start, connectivity) + { + previous_cursor += 1; } - components.push(ComponentStats { - label: next_label, - area, - bbox: BoundingBox2 { - x_min: min_x as f32, - y_min: min_y as f32, - x_max: (max_x + 1) as f32, - y_max: (max_y + 1) as f32, - }, - centroid: [sum_x / area as f64, sum_y / area as f64], - }); + let mut overlap = previous_cursor; + while overlap < previous_runs.end && !run_follows(runs[overlap], end, connectivity) { + union_component_runs(&mut runs, run_index, overlap); + overlap += 1; + } + } + previous_runs = current_start..runs.len(); + } + + let mut root_labels = vec![0_u32; runs.len()]; + let mut accumulators = Vec::::new(); + let mut next_label = 1_u32; + for run_index in 0..runs.len() { + let run = runs[run_index]; + let root = component_run_root(&mut runs, run_index); + let label = if root_labels[root] == 0 { + let label = next_label; + root_labels[root] = label; + accumulators.push(ComponentAccumulator::new(run)); next_label = next_label .checked_add(1) .ok_or_else(|| VisionError::InvalidDimensions("too many components".to_owned()))?; - } + label + } else { + root_labels[root] + }; + labels[run.y * width + run.start..run.y * width + run.end].fill(label); + accumulators[label as usize - 1].include(run); } + + let components = accumulators + .into_iter() + .enumerate() + .map(|(index, accumulator)| accumulator.finish(index as u32 + 1)) + .collect(); let metadata = ImageMetadata { color_space: ColorSpace::Label, ..Default::default() }; let image = Image::try_new_with_metadata(width, height, labels, metadata)?; Ok(ConnectedComponents { labels: LabelImage { image }, components }) } -fn neighbors( - x: usize, +#[derive(Clone, Copy)] +struct ComponentRun { y: usize, - width: usize, - height: usize, - connectivity: Connectivity, -) -> impl Iterator { - let offsets: &[(isize, isize)] = match connectivity { - Connectivity::Four => &[(0, -1), (-1, 0), (1, 0), (0, 1)], - Connectivity::Eight => { - &[(-1, -1), (0, -1), (1, -1), (-1, 0), (1, 0), (-1, 1), (0, 1), (1, 1)] + start: usize, + end: usize, + parent: usize, +} + +fn run_precedes(run: ComponentRun, start: usize, connectivity: Connectivity) -> bool { + match connectivity { + Connectivity::Four => run.end <= start, + Connectivity::Eight => run.end < start, + } +} + +fn run_follows(run: ComponentRun, end: usize, connectivity: Connectivity) -> bool { + match connectivity { + Connectivity::Four => run.start >= end, + Connectivity::Eight => run.start > end, + } +} + +fn component_run_root(runs: &mut [ComponentRun], mut index: usize) -> usize { + let mut root = index; + while runs[root].parent != root { + root = runs[root].parent; + } + while runs[index].parent != index { + let parent = runs[index].parent; + runs[index].parent = root; + index = parent; + } + root +} + +fn union_component_runs(runs: &mut [ComponentRun], left: usize, right: usize) { + let left_root = component_run_root(runs, left); + let right_root = component_run_root(runs, right); + if left_root != right_root { + let (root, child) = + if left_root < right_root { (left_root, right_root) } else { (right_root, left_root) }; + runs[child].parent = root; + } +} + +struct ComponentAccumulator { + area: usize, + min_x: usize, + min_y: usize, + max_x: usize, + max_y: usize, + sum_x: f64, + sum_y: f64, +} + +impl ComponentAccumulator { + fn new(run: ComponentRun) -> Self { + Self { + area: 0, + min_x: run.start, + min_y: run.y, + max_x: run.end, + max_y: run.y + 1, + sum_x: 0.0, + sum_y: 0.0, } - }; - offsets.iter().filter_map(move |&(dx, dy)| { - let nx = x.checked_add_signed(dx)?; - let ny = y.checked_add_signed(dy)?; - (nx < width && ny < height).then_some((nx, ny)) - }) + } + + fn include(&mut self, run: ComponentRun) { + let length = run.end - run.start; + self.area += length; + self.min_x = self.min_x.min(run.start); + self.min_y = self.min_y.min(run.y); + self.max_x = self.max_x.max(run.end); + self.max_y = self.max_y.max(run.y + 1); + self.sum_x += length as f64 * (run.start as f64 + run.end as f64) * 0.5; + self.sum_y += length as f64 * (run.y as f64 + 0.5); + } + + fn finish(self, label: u32) -> ComponentStats { + ComponentStats { + label, + area: self.area, + bbox: BoundingBox2 { + x_min: self.min_x as f32, + y_min: self.min_y as f32, + x_max: self.max_x as f32, + y_max: self.max_y as f32, + }, + centroid: [self.sum_x / self.area as f64, self.sum_y / self.area as f64], + } + } } /// One closed polygonal contour on pixel-grid corner coordinates. @@ -1304,12 +1419,13 @@ impl PointMap { #[cfg(test)] mod tests { use super::{ - approximate_polygon, connected_components, decode_rle, distance_transform_edt, - distance_transform_edt_u8_into, distance_transform_edt_with_spacing, encode_rle, - find_contours, BinaryMask, Connectivity, DepthMap, DistanceTransformWorkspace, FlowField, - PointMap, RleOrder, + approximate_polygon, connected_components, connected_components_u8, decode_rle, + distance_transform_edt, distance_transform_edt_u8_into, + distance_transform_edt_with_spacing, encode_rle, find_contours, BinaryMask, Connectivity, + DepthMap, DistanceTransformWorkspace, FlowField, PointMap, RleOrder, }; use spatialrust_image::Image; + use std::collections::VecDeque; #[test] fn threshold_and_components_find_two_regions() { @@ -1327,6 +1443,116 @@ mod tests { assert_eq!(result.components[0].centroid, [1.0, 1.0]); } + #[test] + fn run_length_components_match_pixel_flood_fill() { + let mut state = 0x9e37_79b9_u32; + for connectivity in [Connectivity::Four, Connectivity::Eight] { + for case in 0..96 { + let width = 1 + case % 37; + let height = 1 + (case * 11) % 29; + let cutoff = match case % 4 { + 0 => 40, + 1 => 250, + 2 => 600, + _ => 850, + }; + let data = (0..width * height) + .map(|_| { + state = state.wrapping_mul(1_664_525).wrapping_add(1_013_904_223); + u8::from(state % 1_000 < cutoff) + }) + .collect::>(); + let mask = BinaryMask::try_new(width, height, data.clone()).unwrap(); + let actual = connected_components(&mask, connectivity).unwrap(); + let expected = flood_fill_labels(&data, width, height, connectivity); + assert_eq!(actual.labels.as_slice(), expected, "case {case:?} {connectivity:?}"); + + let max_label = expected.iter().copied().max().unwrap_or(0) as usize; + assert_eq!(actual.components.len(), max_label); + for component in &actual.components { + let mut area = 0; + let mut min_x = width; + let mut min_y = height; + let mut max_x = 0; + let mut max_y = 0; + let mut sum_x = 0.0; + let mut sum_y = 0.0; + for (index, &label) in expected.iter().enumerate() { + if label == component.label { + let (x, y) = (index % width, index / width); + area += 1; + min_x = min_x.min(x); + min_y = min_y.min(y); + max_x = max_x.max(x + 1); + max_y = max_y.max(y + 1); + sum_x += x as f64 + 0.5; + sum_y += y as f64 + 0.5; + } + } + assert_eq!(component.area, area); + assert_eq!(component.bbox.x_min, min_x as f32); + assert_eq!(component.bbox.y_min, min_y as f32); + assert_eq!(component.bbox.x_max, max_x as f32); + assert_eq!(component.bbox.y_max, max_y as f32); + assert_eq!(component.centroid, [sum_x / area as f64, sum_y / area as f64]); + } + } + } + } + + #[test] + fn borrowed_u8_components_accept_all_nonzero_values() { + let pixels = [0, 2, 0, 255, 0, 7]; + let result = connected_components_u8(3, 2, &pixels, Connectivity::Eight).unwrap(); + assert_eq!(result.labels.as_slice(), [0, 1, 0, 1, 0, 1]); + assert_eq!(result.components[0].area, 3); + assert!(connected_components_u8(3, 2, &pixels[..5], Connectivity::Eight).is_err()); + } + + fn flood_fill_labels( + data: &[u8], + width: usize, + height: usize, + connectivity: Connectivity, + ) -> Vec { + let offsets: &[(isize, isize)] = match connectivity { + Connectivity::Four => &[(0, -1), (-1, 0), (1, 0), (0, 1)], + Connectivity::Eight => { + &[(-1, -1), (0, -1), (1, -1), (-1, 0), (1, 0), (-1, 1), (0, 1), (1, 1)] + } + }; + let mut labels = vec![0_u32; data.len()]; + let mut queue = VecDeque::new(); + let mut next_label = 1_u32; + for seed in 0..data.len() { + if data[seed] == 0 || labels[seed] != 0 { + continue; + } + labels[seed] = next_label; + queue.push_back((seed % width, seed / width)); + while let Some((x, y)) = queue.pop_front() { + for &(dx, dy) in offsets { + let Some(nx) = x.checked_add_signed(dx) else { + continue; + }; + let Some(ny) = y.checked_add_signed(dy) else { + continue; + }; + if nx >= width || ny >= height { + continue; + } + let index = ny * width + nx; + if data[index] != 0 && labels[index] == 0 { + labels[index] = next_label; + queue.push_back((nx, ny)); + } + } + } + next_label += 1; + } + labels + } + #[test] fn contours_trace_outer_and_hole_loops() { let mask = BinaryMask::try_new(3, 3, vec![1, 1, 1, 1, 0, 1, 1, 1, 1]).unwrap(); diff --git a/docs/ROADMAP.md b/docs/ROADMAP.md index 5ef0a45..3932ea3 100644 --- a/docs/ROADMAP.md +++ b/docs/ROADMAP.md @@ -621,6 +621,7 @@ to one implicitly, and GPU receipts must retain named upload/readback stages. | 114G | Complete | Cache NMS box geometry and avoid packed Python score copies | exact OpenCV index parity and 100/1,000/8,400-candidate wins | | 114H | Complete | Bucket class-aware NMS keeps and expose one-call Python batched NMS | exact OpenCV parity and 26.38×/97.25× wins | | 114I | Complete | One-pass active-set Soft-NMS selection, disjoint IoU exit, and borrowed Python scores | exact indices, bounded scores, and 3.42×–7.40× wins | +| 114J | Complete | Run-length union-find connected components with borrowed/non-zero Python masks | exact SAUF labels/stats and 2.17×–3.61× structured-mask wins | ### Epic 115 delivery slices diff --git a/docs/site/algorithms.html b/docs/site/algorithms.html index 1adc840..6950744 100644 --- a/docs/site/algorithms.html +++ b/docs/site/algorithms.html @@ -35,7 +35,7 @@

Algorithm catalog

MorphologyErode, dilate, open, close, gradient, top-hat, black-hat with Rect/Cross/Ellipse/Diamond/custom elementsspatialrust-vision · imgproc-morphologyCPU / GPU Image analysisFixed/Otsu/adaptive threshold, histogram, equalization, CLAHE, integral image, Cannyspatialrust-vision · imgproc-analysis, imgproc-cannyCPU Local featuresFAST, Harris, Shi–Tomasi, ORB, descriptor matching, grid selection, pyramidal Lucas–Kanade trackingspatialrust-vision · feature2dCPU - Dense visionExact Euclidean distance transform with reusable workspace/output, connected components, contours, polygon approximation, mask RLE, depth/confidence/flow/point mapsspatialrust-vision · denseCPU parallel + Dense visionExact Euclidean distance transform with reusable workspace/output; row-major run-length/union-find connected components; contours, polygon approximation, mask RLE, and dense spatial maps. Structured VGA/1080p/4K masks measured 2.17×–3.61× faster than OpenCV SAUF with exact labels, areas, and boxes across canonical plus 320 randomized cases. Harness.spatialrust-vision · denseCPU Detection post-processingIoU/GIoU, greedy NMS, class-aware batched NMS, and hard/linear/Gaussian Soft-NMS. Seeded Python NMS is 3.22×–8.95× faster than OpenCV; batched NMS is 26.38×–97.25× faster; linear/Gaussian Soft-NMS is 3.42×–7.40× faster. NMS indices are exact and Soft-NMS scores stay within 1.79e-7. NMS · batched · Soft-NMS.spatialrust-vision · detectionCPU Multiview geometryHomography, fundamental/essential matrices, RANSAC, triangulation, relative pose, PnP/PnP-RANSACspatialrust-vision · geometryCPU Stereo and odometryStereo rectification, block matching, disparity-to-depth/XYZ, monocular and RGB-D visual odometryspatialrust-vision · geometry, odometryCPU diff --git a/docs/site/vision2.html b/docs/site/vision2.html index b6ccfdd..5e0a0ab 100644 --- a/docs/site/vision2.html +++ b/docs/site/vision2.html @@ -52,8 +52,9 @@

Latest measured outcome

Exact EDT at 4K

With caller-owned output and workspace reuse, SpatialRust measured 40.66 ms versus OpenCV 43.33 ms: a 1.07× lead on the recorded Windows host.

Accuracy preserved

The VGA, 1080p, and 4K canonical masks retain maximum absolute error 0.0 against OpenCV's precise L2 distance transform.

NMS post-processing

SpatialRust measured 3.22×–8.95× faster for NMS, 26.38×–97.25× for batched NMS, and 3.42×–7.40× for linear/Gaussian Soft-NMS. Indices exactly match OpenCV; Soft-NMS scores stay within 1.79e-7.

+

Structured-mask labeling

Run-length union-find connected components measured 2.17×–3.61× faster than OpenCV SAUF at VGA, 1080p, and 4K. Labels, areas, and boxes are exact across canonical and 320 randomized cases.

-

This is a workload- and host-specific result. VGA and 1080p reuse remain narrow OpenCV wins; the repository receipt contains the reproducible methodology.

+

These are workload- and host-specific results. The connected-components speed claim covers structured segmentation/document masks; dense random noise is not claimed. The repository receipts contain the reproducible methodology.

diff --git a/notes/2026-07-15_connected_components_opencv_acceleration.md b/notes/2026-07-15_connected_components_opencv_acceleration.md new file mode 100644 index 0000000..045a455 --- /dev/null +++ b/notes/2026-07-15_connected_components_opencv_acceleration.md @@ -0,0 +1,102 @@ +# Connected-components OpenCV acceleration receipt — 2026-07-15 + +## Outcome + +SpatialRust's connected-component labeling now uses row runs and union-find +instead of queue-based per-pixel flood fill. On the recorded structured +segmentation and document masks it is **2.17×–3.61× faster** than OpenCV's +explicit SAUF implementation while returning exactly the same row-major labels, +foreground areas, and bounding boxes. + +Connected components are a foundational post-segmentation operation for object +counting, instance extraction, OCR/document regions, defect inspection, and +binary-mask cleanup. The optimized path remains pure safe Rust. + +## Algorithm choice + +OpenCV exposes SAUF, BBDT, and Spaghetti labeling. Its documentation states that +SAUF is the variant that forces row-major label ordering, so SAUF is the +compatibility oracle for SpatialRust's existing row-major contract. OpenCV's +default 8-connected path uses Spaghetti and does not promise the same ordering. + +The SpatialRust implementation scans each row into maximal foreground runs, +unions only overlapping runs in the preceding row, path-compresses equivalence +classes, then emits consecutive labels and statistics in first-run order. Long +structured regions therefore avoid queue traffic and repeated neighbor checks. +This follows the same union-find optimization direction described by Wu, Otoo, +and Suzuki while exploiting run structure rather than inspecting every +foreground pixel's neighborhood. + +References: + +- OpenCV connected-components API and ordering contract: +- OpenCV implementation dispatch: +- Wu, Otoo, Suzuki, *Optimizing Two-Pass Connected-Component Labeling Algorithms*: +- Bolelli et al., *Spaghetti Labeling*: + +## Reproduction environment + +- Windows 11 `10.0.26300`, AMD64 +- Intel Family 6 Model 158, 6 cores / 12 logical CPUs +- CPython 3.12.10 +- OpenCV 4.13.0, 12 reported threads, OpenCL disabled +- SpatialRust 1.0.0 release wheel +- 8-connectivity structured masks; `uint8` values 0/255 +- paired/interleaved order, eight warmups, calls batched to at least 20 ms +- 40 VGA, 30 1080p, and 20 4K samples per pattern + +Run: + +```powershell +python bench/opencv_connected_components_comparison/performance.py ` + --output target/opencv-connected-components-performance.json +``` + +## Python API medians + +| Profile | Pattern | OpenCV SAUF | SpatialRust | Speedup | +| --- | --- | ---: | ---: | ---: | +| VGA | Segmentation blobs | 1.284 ms | 0.413 ms | 3.11× | +| VGA | Document lines | 1.271 ms | 0.352 ms | 3.61× | +| 1080p | Segmentation blobs | 6.763 ms | 2.815 ms | 2.40× | +| 1080p | Document lines | 6.649 ms | 2.407 ms | 2.76× | +| 4K | Segmentation blobs | 21.356 ms | 9.838 ms | 2.17× | +| 4K | Document lines | 21.075 ms | 8.606 ms | 2.45× | + +The timing scope includes each Python API call, label image, and foreground +statistics. Packed NumPy input is borrowed; generated label storage is moved +into NumPy rather than copied. + +## Correctness gates + +- exact label image against OpenCV `CCL_SAUF` +- exact foreground component count, areas, and half-open bounding boxes +- all canonical profile/pattern pairs passed +- 320 seeded randomized rectangle/noise masks passed for both 4- and + 8-connectivity, including arbitrary non-zero foreground byte values +- 192 deterministic Rust cases match an independent pixel flood-fill reference, + including areas, boxes, and SpatialRust pixel-center centroids + +## Native Criterion medians + +| Profile | Segmentation blobs | Document lines | +| --- | ---: | ---: | +| VGA | 376 µs | 269 µs | +| 1080p | 2.319 ms | 1.703 ms | +| 4K | 8.682 ms | 7.806 ms | + +The corresponding median throughputs range from 816.7 million to 1.143 billion +pixels/s. Reproduce with: + +```powershell +cargo bench -p spatialrust-vision --bench dense --features dense -- ` + connected_components_8_structured +``` + +## Scope boundary + +The speed claim is intentionally limited to masks with useful horizontal run +structure, such as region/instance segmentation and document lines. Preliminary +random-noise tests improved substantially over the former flood fill but still +favored OpenCV for dense or highly fragmented masks; no general random-mask win +is claimed.