Repository navigation
Matrix kernels zero decode - #27
Merged
Merged
Conversation
Matrix-dot kernels (dot ids 7-24), all through sparsify_and_dot(data, indices, indptr), the dot id saying how the arrays are read: - csc-run (start + length), csc-nosplit (histogram), csc-nosplit-moment, bsb-csr-nosplit, and the experimental set in bslz4_csc_variants.hpp / _csc_variants.py: dump-bin, u16 starts, delta and walk index streams, 2D tile, split moment (q at the column front), fixed-point u32/u16 weights with an int64 powder, FAZIT permutation (csc-permute, -runs). - c2py spec: types only (powder d/q/any width for the packed output; data f/I/H; indices 4-byte, H or a headed byte stream B); C prototype void*. u16 sparsify, very sparse frames 2.6x (4M) / 3.2x (16M) faster: - collect tier 6 "avx512cs" (default with AVX-512 VBMI2): compress values and indices into registers, advance by popcnt; 23-57 % faster when many pixels are selected. - plane extraction (default on): blocks > 24x compressed with <= 48 non-zero pixels are read straight from the OR of their bit-planes. - byte skip (default on, u16, AVX-512 VBMI + GFNI): blocks < 256 untranspose only the 8 low planes, straight into u16 (kcb's GFNI transpose). - zero-aware lz4 block decoder (bslz4_lz4zero.h, default on for blocks > 24x): zero runs and the zero tail are not written; nz_end tells the consumers where data ends. Fuzzed against LZ4_decompress_safe (~300M inputs, ASan+UBSan): identical output, same rejections, plus offset 0 rejected (lz4 #1631). - lz4 submodule: branch bslz4/v1.10.0 (local commit, not yet pushed), the offset 8..15 overlap-copy fix; also patches/lz4/*.patch. 16M dense lz4 ~4.5x faster. Benchmarks and tools: tools/bench_suite.py (real pyFAI geometries, 4M/16M, beam mid/corner), bench_perf.py (perf counters, per-stage profile), lz4_blocks.c / lz4_scan.py, membw*.c, compress_matrices.py. Findings in notes_dot_layouts.md. Tests: 576 passed. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
… division Measured on WAu5um (Eiger2 4M, frames 400-599 of eiger_0008 and 0012, mask as 65535 and as 0): - plane extraction: limit 48 -> 16, now off by default (it loses on real frames with photon noise, ~10-35 non-zero px per block); its bitmap is AND-ed with the caller's fixed mask, so masked pixels never count. - zero-aware lz4 decoder: used above 48x (was 24x; lost up to 6 % on real blocks with many short sequences). - set_mask_planes (off by default): the fixed mask packed into the bit-plane layout once per block per batch and AND-ed into the planes after lz4; blocks with nothing masked skip it. -1..-5 % on mask=65535 data at cut 0 (lets the byte skip apply), overhead on mask=0 data. - the mask-plane AND first divided per 64-byte chunk (~1 us/block): fixed; no division per pixel remains in the hot loops. - coo(): reciprocal multiply instead of np.divmod per pixel (exact for every pixel of the 4M/16M shapes), 1.6-1.7x faster. Only the caller's fixed mask is ever applied: 65535 in unmasked pixels (saturation, dynamic masking) is always collected; tests cover masked and saturated pixels at the maximum under every switch combination. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…nes collect Zero-aware lz4 decoder: after a copied match, zero_from is set past its last non-zero byte, so the long offset-2 zero matches that follow on real Eiger blocks are skipped instead of copied (~1.2 kB/block of libc memmove on WAu mask=0). Short literals, matches and zero fills are fixed-size inline copies that may write into caller-provided slack past the block. Fuzzed against LZ4_decompress_safe (~940M runs, real blocks as seeds). Byte skip: u16 blocks with empty high planes are transposed and collected in one pass (bslz4_lowplanes_collect_u16), for plain sparsify and the sparse dot route (precompacted list via bslz4_collect_nz_w). WAu mask=0 vs 45f830d: 0012 sparsify cut 2 -28 %, 0008 cut 2 -16 %; cut 0 and CSC -6..+0 %. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
- Fused u16 low-planes collect in two passes (record every 64-px group, then emit the groups with pixels by bit scan): the per-group branch mispredicted on photon noise at cut 0. - Sparse CSC dot: software prefetch of indptr/entries ahead of use, and staged entries for blocks averaging 1.5-4 entries per pixel. - lz4.c compiled with 64-byte function/loop alignment: its speed moved ~25 % with the size of the code linked before it. WAu mask=0 cut 0: sparsify -28..-34 %, CSC 1D/2D/rings -20..-31 %; synthetic 9 % frames: 1D bbox -15 %, 2D bbox -11 %. Cut 2 on very sparse frames +4-6 %. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
tools/compiler_suite.py builds six variants of the extension, runs the test suite against each and times them interleaved on real and synthetic frames, per host, into a shared directory. build_extension.py takes BSLZ4_EXTRA_CFLAGS / BSLZ4_EXTRA_LDFLAGS (for PGO); tools/zig-cc and tools/zig-c++ drive zig as a baseline-CPU C/C++ compiler (x86_64, ppc64le). Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…itecture Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…ne.json Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
… and p9 Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…I; fix counter overflow - Collect tier 7 "avx2cs": the selection packed with a 256-entry lane table (vpermd for indices, pshufb for u16 values), full-width stores, advance by popcount; a group with one pixel keeps the bit loop. Default after avx512cs. - Byte skip without VBMI + GFNI: the 8 low bit-planes go through kcb as one-byte elements, then are widened to u16 (AVX2 / SSE2). - bslz4_counters had 16 slots per stage but dot ids reach 24: counting dots 16-24 wrote past the array into the static data after it (it hit the new lane table). 32 slots, a _Static_assert against the dot table, and _COUNTER_SLOTS for the tests. EPYC 7543 (hpc6), vs ce9504c: dense sparsify -55 %, dense rings -53 %, 9 % frames -9..-21 %, 0.1 % frames -16..-22 %, WAu -7..-24 %. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…MI/GFNI When a u16 block's high planes are empty, kcb untransposes the low planes as bytes into scratch and bslz4_collect_u8_avx2cs packs the pixels > cut (plain sparsify) or the non-zero ones (sparse dot route) straight from them: 32 pixels per compare, the avx2cs lane table per 8, no u16 block. EPYC 7543 (hpc6), vs ce9504c: sparsify -24..-63 %, dots -17..-55 %; WAu0012 sparsify -40 %, CSC 1D -32 %. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…ut VBMI/GFNI bslz4_lowplanes_collect_avx2: kcb's AVX2 bit transpose for the 8 low planes, empty planes read from a zero buffer instead of zero filling, the 64-pixel groups and their selections recorded in pass 1, only the groups with pixels packed in pass 2. Replaces kcb + the u8 collect (BSLZ4_FUSED_U8 1) as the default. EPYC 7543 (hpc6), vs ce9504c: sparsify -29..-64 %, dots -18..-54 %; WAu0012 sparsify -49 %, CSC 1D -39 %, rings -46 %. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
The planes are ORed 32 bytes at a time into a bitmap of non-empty groups; pass A transposes and compares only those (recording, no branch per group), pass B packs only the groups with selected pixels. EPYC 7543 (hpc6), vs the two-pass version: 0.1 % frames -27..-40 %, WAu0012 -15..-35 %, WAu0008 -8..-19 %, 9 % frames -1..-4 %, dense -5..+2 %. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…lap scan A copied match of <= 64 bytes goes in 16-byte chunks (a 64-byte load of a source just written by smaller stores stalled on store forwarding) and the chunks' non-zero bytes give its trailing zeros by one bit scan; that byte loop was 26 % of the decoder on WAu blocks. An overlapping copy repeats with period off, so its zero scan stops after off bytes. SSE2 only (baseline x86-64); the scalar path stays for other architectures. Real WAu blocks: 1077 -> 781 cycles/block (0012), 1289 -> 919 (0008), output identical on all 138613; libFuzzer vs LZ4_decompress_safe 224M runs clean. hpc6: WAu0012 -5..-14 %, WAu0008 -2..-6 %. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Non-overlapping matches longer than 64 bytes are copied in 16-byte chunks, keeping the last chunk with a non-zero byte, instead of memcpy and a byte scan back over the (mostly zero) copy. Synthetic sparse blocks of 32-48x had ~15 such copies per block: 4924 -> 2091 cycles/block (stock lz4 2941), WAu unchanged (1154 -> 1109 on 0012). Output identical to the previous decoder on all real and band blocks; libFuzzer 115M runs clean. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
With chunked match copies it beats stock LZ4_decompress_safe on nearly all blocks, dense ones included (it never writes the empty high byte-planes). Per block, hpc5: dense 2-8x 2894 -> 1406 cycles, 9 % frames 16-24x 5481 -> 4821, 8-16x 6170 -> 6463. End to end, hpc5 vs 48x: WAu0008 -17..-36 %, 0.1 % frames -6..-13 %, dense sparsify -15 %, dense rings -17 %, 9 % frames -7..+1 %, WAu0012 -1..-3 %. To be confirmed on Zen 3/4/5. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
On Zen 5 (hpc8) the zero decoder for every block (6f9d43d) made 9 %- occupied frames 8-30 % slower: their 8-24x blocks are ~90-130 short sequences per 8 kB, where stock's loop is faster there (8-16x: 2408 vs 3142 cycles/block). Dense blocks (< 8x, literal heavy, empty high planes) and sparse ones (> 24x) stay with the zero decoder. hpc8 vs last week (ce9504c): 9 % frames -1..+4 %, WAu0012 -10..-25 %, WAu0008 -8..-23 %, 0.1 % frames -5..-20 %, dense -2..-12 %. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
As on the AVX2 path: the planes ORed 64 bytes (8 groups) at a time give a bitmap of non-empty groups; for blocks compressed more than 48x only those are transposed (bit-scan loop), otherwise the plain counted loop (a bit-scan loop over every group cost a few % on denser blocks). hpc8 (Zen 5, medians of 6, node heavily loaded): 0.1 % frames -25 %, WAu0012 -8 %, WAu0008 -4 %, 9 % frames within noise. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
On ppc64le/aarch64, u16 blocks with empty high planes are untransposed as bytes (kcb, 8 planes, no byte transpose) and collected from those bytes with a portable loop that skips all-zero 8-byte words; plain sparsify and the sparse dot route. POWER9 profile before: kcb's scalar 16-plane bit transpose + byte transpose ~80 % of sparsify. p9-10 (POWER9 3.8 GHz), ms/frame: 0.1 % frames 10.97 -> 5.57, WAu0012 11.70 -> 4.87, WAu0008 13.52 -> 6.05, 9 % frames 26.7 -> 14.7, dense 28.3 -> 12.9. Tests pass on ppc64le and x86_64. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Per 64-pixel group the 8 low planes' chunks are regrouped by two vec_perm rounds so each doubleword holds a byte of every plane, and vgbbd (vec_gb, an 8x8 bit transpose per doubleword) gives the pixel bytes; all-zero groups are skipped, pixels collected from the bytes. Byte order checked against a reference transpose on POWER9 (no reversal needed on LE). p9-10, ms/frame (before today -> portable byte skip -> VSX): WAu0012 11.70 -> 4.86 -> 2.21, 0.1 % frames 10.97 -> 5.56 -> 2.98, WAu0008 13.52 -> 6.02 -> 3.81, dense 28.3 -> 12.9 -> 8.4, 9 % frames 26.7 -> 14.7 -> 12.8; WAu0012 CSC 1D 12.1 -> 5.48 -> 2.78. Tests pass on ppc64le and x86_64. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Scattered selections (9 % frames) mispredicted the per-pixel branch; storing every pixel and advancing by its selection is -18 % there, but a serial chain on dense blocks where nearly every pixel passes (+89 %), so dense blocks keep the branch. p9-10: 9 % sparsify 12.8 -> 10.6 ms, 9 % CSC 1D 18.3 -> 15.8 ms; dense and WAu unchanged. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
POWER9's stock lz4 loses to the zero decoder in the 8-24x band as well (8-16x 3781 vs 4812 ns/block), so off x86 the cutoff is 1: p9-10 9 % sparsify 10.6 -> 9.3 ms, CSC 1D 15.8 -> 14.5 ms, others flat. x86 keeps the 8x/24x band. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
The zero decoder's SSE2 path (also taken under _M_X64) used __builtin_clz; the dot prefetch used __builtin_prefetch; the counter-slot check used _Static_assert. Now _BitScanReverse / _mm_prefetch under MSVC, and a negative-array typedef for the compile-time check. Built with gcc 13 (-Wall -Wextra clean apart from a pre-existing padded.hpp warning). Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
The submodule pointed at a local commit (our offset 8..15 fix) that exists on no remote, so a clone could not check it out. It now points at upstream v1.10.0 (ebb370ca, also lz4's release branch), and build_extension.py compiles a copy of lib/lz4.c with patches/lz4/*.patch applied by a strict pure-Python applier (no git/patch binary needed, e.g. on Windows; the patches are in the source digest). The patched copy is byte-identical to the previous submodule commit; 605 tests pass. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…t fixes - build_extension: apply patches with universal newlines so a core.autocrlf checkout still matches; .gitattributes pins *.patch to LF - MSVC: objects, import library and .exp go to the build dir instead of the cwd and src/ - test_dot: skip without pyFAI instead of failing collection - test_vs_hdf5plugin: quote-safe call of make_testcases.py Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…amplers - patches/kcb: PE/COFF has no ifunc; on _WIN32 kcb's resolvers fill a function pointer from a load-time constructor (was SSE2 only with BITSHUF_USE_IFUNC=0: 0.1 % frames 0.745 -> 0.599 ms on an i7-8700) - build_extension: patches for any submodule (lz4, kcb); gcc for Windows gets -Wa,-muse-unaligned-vector-move (GCC bug 54412: aligned AVX stack stores crashed the fused AVX2 collect); no -fPIC on Windows - MANIFEST.in: ship patches/ (an sdist built unpatched lz4 silently) - tools/sampler: the LD_PRELOAD / in-process SIGPROF samplers used on hpc6 and POWER9, a Windows sampler (winprof.py) and a one-case driver - notes: i7-8700 results; the gcc cross build is 2.0x MSVC Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
build-windows now runs on ubuntu-24.04 with Ubuntu's mingw-w64 gcc 13 (binutils 2.42); the MSVC build had none of the SIMD paths and was 2.1x slower on an i7-8700. test-windows checks the SIMD collect is active and runs the corrupt-chunk tests too. MSVC stays as build-windows-msvc: build + decode matrix only, so the sources keep compiling with it. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…etuptools
test-py27 (ubuntu 18.04, gcc 7) failed: target("gfni"/"avx512vbmi2"), the
VBMI2/GFNI intrinsics and their __builtin_cpu_supports names need gcc >= 8
or clang >= 7. BSLZ4_HAVE_VBMI_GFNI (bslz4_common.h) now guards the
low-planes VBMI/GFNI kernels and the avx512cs collect; older compilers get
stubs and capability probes that say no, so the AVX2 paths run.
BSLZ4_NO_VBMI_GFNI forces the fallback (605 tests pass either way on Zen 4;
all sources compile under gcc 7.5 in a container, the previous driver did
not: 12 errors, as in CI).
build-windows-msvc failed at `python setup.py build`: the runner's Python
3.12 has no setuptools.
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
No description provided.