Performance attribution
Separate native kernel, allocation, Python conversion, upload, execution, and readback costs. Publish MPix/s, ns/pixel, memory, and thread policy.
Turn measured OpenCV gaps into faster native kernels, reusable memory, explicit GPU-resident chains, and fail-closed performance evidence—without weakening accuracy or ownership contracts.
Targets compare SpatialRust with its Epic 112 baseline on the same host. OpenCV remains an accuracy and workload reference; results never imply universal superiority.
flowchart TD
classDef done fill:#12362c,stroke:#57d3b0,color:#e8eef8;
classDef doing fill:#2b2510,stroke:#ffcf70,color:#e8eef8;
classDef todo fill:#172842,stroke:#28405f,color:#9fb0c7;
E112["112 · Performance attribution"]:::done
E113["113 · Reusable outputs and workspaces"]:::doing
E114["114 · Safe CPU dispatch"]:::doing
E115["115 · Resize and color"]:::done
E116["116 · Gaussian and Sobel"]:::done
E117["117 · Morphology"]:::done
E118["118 · Fused Canny"]:::done
E119["119 · Explicit GPU chain"]:::done
E120["120 · Vision 2 release gate"]:::done
E104["104 · GPU Image v2"]:::done
E112 --> E113 --> E114
E114 --> E115 --> E119
E114 --> E116 --> E119
E114 --> E117
E114 --> E118 --> E119
E104 --> E119
E119 --> E120
E112 --> E120
Node color follows the ROADMAP status: green is complete, amber is in progress, slate is planned. The canonical status table stays in ROADMAP.md.
Separate native kernel, allocation, Python conversion, upload, execution, and readback costs. Publish MPix/s, ns/pixel, memory, and thread policy.
Add caller-owned outputs and explicit scratch storage for Gaussian, Sobel, morphology, and Canny while preserving packed and strided behavior.
Use scalar small-image paths, packed specializations, bounded row/tile parallelism, and fully compatible generic fallbacks.
Reusable Q11 bilinear plans, exact packed RGB8 half-scale, packed nearest/area, target-dispatched Q14 RGB-to-gray, fused resize-to-gray, and fused resize-normalize-CHW are delivered.
Reusable separable intermediates, cached kernels, isolated borders, paired gradients, and a high-precision Q15 7×7 path are delivered.
Introduce rectangular sliding min/max, small-kernel paths, ping-pong scratch, and exact generic-shape fallback.
Avoid materializing public intermediates in the standard path while retaining an opt-in inspectable result.
Upload once, execute resize-to-AI stages on pooled textures, and read back only when the caller asks.
Enforce cross-platform accuracy, native/Python performance, memory, allocation, thread, and transfer budgets.
Bounded numerical error for resize, gray, and Gaussian; exact Sobel/morphology comparison; binary F1 and IoU for Canny.
CPU APIs remain on CPU. GPU upload/readback stages stay named, explicit, and auditable.
Vision 1 ownership and error behavior remain compatible; new workspace surfaces are additive.
Every Epic follows implement → test → commit → PR → merge with its own reproducible receipt.
The typed gate covers Linux, Windows, and macOS conformance; native/Python allocate and reuse latency; peak memory; steady-state allocations; worker policy; and exact GPU transfer ceilings. Generated receipt · Migration guide.
One explicit upload feeds resize, gray, blur, Sobel, morphology, and normalized CHW packing. On the recorded Intel UHD 630 Vulkan adapter, the synchronized resident loop measured 1.696 ms at VGA, 7.423 ms at 1080p, and 16.005 ms at 4K, with no readback before caller request.
On the recorded 12-logical-CPU Windows host, caller-output bilinear resize measured 93.18 MPix/s at VGA, 93.05 MPix/s at 1080p, and 85.45 MPix/s at 4K in default-thread mode. The native kernel was the largest measured component; explicit Python packing cost roughly half as much, while output allocation stayed below 0.025 ms.
With caller-owned output and workspace reuse, SpatialRust measured 40.66 ms versus OpenCV 43.33 ms: a 1.07× lead on the recorded Windows host.
The VGA, 1080p, and 4K canonical masks retain maximum absolute error 0.0 against OpenCV's precise L2 distance transform.
SpatialRust measured 3.22×–8.95× faster for NMS, 26.38×–97.25× for batched NMS, and 3.42×–7.40× for linear/Gaussian Soft-NMS. Indices exactly match OpenCV; Soft-NMS scores stay within 1.79e-7.
Run-length union-find connected components measured 2.17×–3.61× faster than OpenCV SAUF at VGA, 1080p, and 4K. Labels, areas, and boxes are exact across canonical and 320 randomized cases.
Exact 3×3 abs(Gx) + abs(Gy) avoids four OpenCV materialization stages. Allocated Python calls measured 1.86× faster at 1080p, 2.19× at 4K, and 2.42× at 8K across 300 randomized parity cases.
A weak-candidate hysteresis frontier avoids revisiting every initial strong edge. Caller-output sensor-noise workloads measured 2.59× faster than OpenCV at 1080p and 2.75× at 4K, while document-line wins remain 1.38× and 1.47×. VGA dense noise remains an OpenCV win.
Worker-owned halo bands now run the Q8 horizontal and vertical passes back-to-back. Allocated Python calls improve another 1.80× at 1080p and 1.70× at 4K with exact canonical OpenCV parity; standalone OpenCV still leads by 1.74× and 1.68× on the recorded host.
Precomputed Q11 coefficients and an exact row-parallel half-scale path reduce the old 26×–146× gap. VGA caller-output measured 0.120 ms versus OpenCV's 0.133 ms (1.10× faster); 1080p/4K/8K remain scoped optimization targets.
Q14 BT.601 coefficients, target-feature dispatch, and size-aware row blocks improve native reuse by 5.7×–10.6×. Allocated Python calls measured 1.03× faster than OpenCV at 1080p and 1.05× at 4K; 8K caller-output reuse measured 1.02× faster.
A single-pass Q11 bilinear + Q14 BT.601 path removes the intermediate RGB image. The allocated 1920×1080→960×540 pipeline measured 0.677 ms versus OpenCV's 0.755 ms (1.12× faster), with bit-exact SpatialRust unfused parity.
Bilinear resize, float normalization, and planar CHW packing now write model input directly. Allocated calls measured 2.21× faster than OpenCV blobFromImage at 1080p→640×640, 2.02× at 4K→640×640, and 2.33× at 4K→1280×720.
These are workload- and host-specific results. The component baseline uses an explicit strided-to-packed Python copy and reports input-pixel throughput; it is not an OpenCV comparison, and CPU-only upload/execution/readback stages are recorded as not applicable. The resize win covers packed RGB8 640×480→320×240 caller-owned output; OpenCV remains 1.49×–1.67× faster for allocated 1080p–8K and 1.85×–2.40× faster for reuse. RGB-to-gray wins cover allocated 1080p/4K and caller-owned 8K; OpenCV remains faster at VGA and for 1080p/4K reuse. The fused resize-to-gray win covers allocated 1080p→540p only; 4K allocation is effectively tied, and OpenCV leads 8K allocation plus all reuse profiles. The fused CHW claim compares scale=1/255, zero mean, unit std against OpenCV blobFromImage at the named model shapes; arbitrary mean/std retain exact SpatialRust unfused parity. The Sobel claim covers fused L1 magnitude allocation, not standalone paired gradients; reuse ties at 1080p and favors OpenCV at 4K/8K. Gaussian improvement compares against SpatialRust's prior generic engine and is not an OpenCV win. The connected-components claim covers structured segmentation/document masks; dense random noise is not claimed. Repository receipts contain the reproducible methodology.