Skip to content

Matrix kernels zero decode - #27

Merged
jonwright merged 35 commits into
mainfrom
matrix-kernels-zero-decode
Oct 5, 2026
Merged

jonwright merged 35 commits into
mainfrom
matrix-kernels-zero-decode

Conversation

@jonwright

Copy link
Copy Markdown
Owner

No description provided.

jonwright and others added 30 commits October 4, 2026 07:03
Matrix-dot kernels (dot ids 7-24), all through sparsify_and_dot(data,
indices, indptr), the dot id saying how the arrays are read:
- csc-run (start + length), csc-nosplit (histogram), csc-nosplit-moment,
  bsb-csr-nosplit, and the experimental set in bslz4_csc_variants.hpp /
  _csc_variants.py: dump-bin, u16 starts, delta and walk index streams,
  2D tile, split moment (q at the column front), fixed-point u32/u16
  weights with an int64 powder, FAZIT permutation (csc-permute, -runs).
- c2py spec: types only (powder d/q/any width for the packed output; data
  f/I/H; indices 4-byte, H or a headed byte stream B); C prototype void*.

u16 sparsify, very sparse frames 2.6x (4M) / 3.2x (16M) faster:
- collect tier 6 "avx512cs" (default with AVX-512 VBMI2): compress values
  and indices into registers, advance by popcnt; 23-57 % faster when many
  pixels are selected.
- plane extraction (default on): blocks > 24x compressed with <= 48
  non-zero pixels are read straight from the OR of their bit-planes.
- byte skip (default on, u16, AVX-512 VBMI + GFNI): blocks < 256 untranspose
  only the 8 low planes, straight into u16 (kcb's GFNI transpose).
- zero-aware lz4 block decoder (bslz4_lz4zero.h, default on for blocks
  > 24x): zero runs and the zero tail are not written; nz_end tells the
  consumers where data ends.  Fuzzed against LZ4_decompress_safe
  (~300M inputs, ASan+UBSan): identical output, same rejections, plus
  offset 0 rejected (lz4 #1631).
- lz4 submodule: branch bslz4/v1.10.0 (local commit, not yet pushed), the
  offset 8..15 overlap-copy fix; also patches/lz4/*.patch.  16M dense lz4
  ~4.5x faster.

Benchmarks and tools: tools/bench_suite.py (real pyFAI geometries, 4M/16M,
beam mid/corner), bench_perf.py (perf counters, per-stage profile),
lz4_blocks.c / lz4_scan.py, membw*.c, compress_matrices.py.  Findings in
notes_dot_layouts.md.  Tests: 576 passed.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
… division

Measured on WAu5um (Eiger2 4M, frames 400-599 of eiger_0008 and 0012, mask
as 65535 and as 0):
- plane extraction: limit 48 -> 16, now off by default (it loses on real
  frames with photon noise, ~10-35 non-zero px per block); its bitmap is
  AND-ed with the caller's fixed mask, so masked pixels never count.
- zero-aware lz4 decoder: used above 48x (was 24x; lost up to 6 % on real
  blocks with many short sequences).
- set_mask_planes (off by default): the fixed mask packed into the bit-plane
  layout once per block per batch and AND-ed into the planes after lz4;
  blocks with nothing masked skip it.  -1..-5 % on mask=65535 data at cut 0
  (lets the byte skip apply), overhead on mask=0 data.
- the mask-plane AND first divided per 64-byte chunk (~1 us/block): fixed;
  no division per pixel remains in the hot loops.
- coo(): reciprocal multiply instead of np.divmod per pixel (exact for every
  pixel of the 4M/16M shapes), 1.6-1.7x faster.

Only the caller's fixed mask is ever applied: 65535 in unmasked pixels
(saturation, dynamic masking) is always collected; tests cover masked and
saturated pixels at the maximum under every switch combination.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…nes collect

Zero-aware lz4 decoder: after a copied match, zero_from is set past its
last non-zero byte, so the long offset-2 zero matches that follow on real
Eiger blocks are skipped instead of copied (~1.2 kB/block of libc memmove
on WAu mask=0).  Short literals, matches and zero fills are fixed-size
inline copies that may write into caller-provided slack past the block.
Fuzzed against LZ4_decompress_safe (~940M runs, real blocks as seeds).

Byte skip: u16 blocks with empty high planes are transposed and collected
in one pass (bslz4_lowplanes_collect_u16), for plain sparsify and the
sparse dot route (precompacted list via bslz4_collect_nz_w).

WAu mask=0 vs 45f830d: 0012 sparsify cut 2 -28 %, 0008 cut 2 -16 %;
cut 0 and CSC -6..+0 %.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
- Fused u16 low-planes collect in two passes (record every 64-px group,
  then emit the groups with pixels by bit scan): the per-group branch
  mispredicted on photon noise at cut 0.
- Sparse CSC dot: software prefetch of indptr/entries ahead of use, and
  staged entries for blocks averaging 1.5-4 entries per pixel.
- lz4.c compiled with 64-byte function/loop alignment: its speed moved
  ~25 % with the size of the code linked before it.

WAu mask=0 cut 0: sparsify -28..-34 %, CSC 1D/2D/rings -20..-31 %;
synthetic 9 % frames: 1D bbox -15 %, 2D bbox -11 %.  Cut 2 on very sparse
frames +4-6 %.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
tools/compiler_suite.py builds six variants of the extension, runs the
test suite against each and times them interleaved on real and synthetic
frames, per host, into a shared directory.  build_extension.py takes
BSLZ4_EXTRA_CFLAGS / BSLZ4_EXTRA_LDFLAGS (for PGO); tools/zig-cc and
tools/zig-c++ drive zig as a baseline-CPU C/C++ compiler (x86_64, ppc64le).

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…itecture

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…ne.json

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
… and p9

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…I; fix counter overflow

- Collect tier 7 "avx2cs": the selection packed with a 256-entry lane
  table (vpermd for indices, pshufb for u16 values), full-width stores,
  advance by popcount; a group with one pixel keeps the bit loop.  Default
  after avx512cs.
- Byte skip without VBMI + GFNI: the 8 low bit-planes go through kcb as
  one-byte elements, then are widened to u16 (AVX2 / SSE2).
- bslz4_counters had 16 slots per stage but dot ids reach 24: counting
  dots 16-24 wrote past the array into the static data after it (it hit the
  new lane table).  32 slots, a _Static_assert against the dot table, and
  _COUNTER_SLOTS for the tests.

EPYC 7543 (hpc6), vs ce9504c: dense sparsify -55 %, dense rings -53 %,
9 % frames -9..-21 %, 0.1 % frames -16..-22 %, WAu -7..-24 %.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…MI/GFNI

When a u16 block's high planes are empty, kcb untransposes the low planes
as bytes into scratch and bslz4_collect_u8_avx2cs packs the pixels > cut
(plain sparsify) or the non-zero ones (sparse dot route) straight from
them: 32 pixels per compare, the avx2cs lane table per 8, no u16 block.

EPYC 7543 (hpc6), vs ce9504c: sparsify -24..-63 %, dots -17..-55 %;
WAu0012 sparsify -40 %, CSC 1D -32 %.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…ut VBMI/GFNI

bslz4_lowplanes_collect_avx2: kcb's AVX2 bit transpose for the 8 low
planes, empty planes read from a zero buffer instead of zero filling, the
64-pixel groups and their selections recorded in pass 1, only the groups
with pixels packed in pass 2.  Replaces kcb + the u8 collect
(BSLZ4_FUSED_U8 1) as the default.

EPYC 7543 (hpc6), vs ce9504c: sparsify -29..-64 %, dots -18..-54 %;
WAu0012 sparsify -49 %, CSC 1D -39 %, rings -46 %.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
The planes are ORed 32 bytes at a time into a bitmap of non-empty groups;
pass A transposes and compares only those (recording, no branch per
group), pass B packs only the groups with selected pixels.

EPYC 7543 (hpc6), vs the two-pass version: 0.1 % frames -27..-40 %, WAu0012
-15..-35 %, WAu0008 -8..-19 %, 9 % frames -1..-4 %, dense -5..+2 %.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…lap scan

A copied match of <= 64 bytes goes in 16-byte chunks (a 64-byte load of a
source just written by smaller stores stalled on store forwarding) and the
chunks' non-zero bytes give its trailing zeros by one bit scan; that byte
loop was 26 % of the decoder on WAu blocks.  An overlapping copy repeats
with period off, so its zero scan stops after off bytes.  SSE2 only
(baseline x86-64); the scalar path stays for other architectures.

Real WAu blocks: 1077 -> 781 cycles/block (0012), 1289 -> 919 (0008),
output identical on all 138613; libFuzzer vs LZ4_decompress_safe 224M runs
clean.  hpc6: WAu0012 -5..-14 %, WAu0008 -2..-6 %.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Non-overlapping matches longer than 64 bytes are copied in 16-byte chunks,
keeping the last chunk with a non-zero byte, instead of memcpy and a byte
scan back over the (mostly zero) copy.  Synthetic sparse blocks of 32-48x
had ~15 such copies per block: 4924 -> 2091 cycles/block (stock lz4
2941), WAu unchanged (1154 -> 1109 on 0012).  Output identical to the
previous decoder on all real and band blocks; libFuzzer 115M runs clean.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
With chunked match copies it beats stock LZ4_decompress_safe on nearly all
blocks, dense ones included (it never writes the empty high byte-planes).
Per block, hpc5: dense 2-8x 2894 -> 1406 cycles, 9 % frames 16-24x
5481 -> 4821, 8-16x 6170 -> 6463.  End to end, hpc5 vs 48x: WAu0008
-17..-36 %, 0.1 % frames -6..-13 %, dense sparsify -15 %, dense rings
-17 %, 9 % frames -7..+1 %, WAu0012 -1..-3 %.  To be confirmed on Zen 3/4/5.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
On Zen 5 (hpc8) the zero decoder for every block (6f9d43d) made 9 %-
occupied frames 8-30 % slower: their 8-24x blocks are ~90-130 short
sequences per 8 kB, where stock's loop is faster there (8-16x: 2408 vs
3142 cycles/block).  Dense blocks (< 8x, literal heavy, empty high planes)
and sparse ones (> 24x) stay with the zero decoder.

hpc8 vs last week (ce9504c): 9 % frames -1..+4 %, WAu0012 -10..-25 %,
WAu0008 -8..-23 %, 0.1 % frames -5..-20 %, dense -2..-12 %.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
As on the AVX2 path: the planes ORed 64 bytes (8 groups) at a time give a
bitmap of non-empty groups; for blocks compressed more than 48x only those
are transposed (bit-scan loop), otherwise the plain counted loop (a
bit-scan loop over every group cost a few % on denser blocks).

hpc8 (Zen 5, medians of 6, node heavily loaded): 0.1 % frames -25 %,
WAu0012 -8 %, WAu0008 -4 %, 9 % frames within noise.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
On ppc64le/aarch64, u16 blocks with empty high planes are untransposed
as bytes (kcb, 8 planes, no byte transpose) and collected from those bytes
with a portable loop that skips all-zero 8-byte words; plain sparsify and
the sparse dot route.  POWER9 profile before: kcb's scalar 16-plane bit
transpose + byte transpose ~80 % of sparsify.

p9-10 (POWER9 3.8 GHz), ms/frame: 0.1 % frames 10.97 -> 5.57, WAu0012
11.70 -> 4.87, WAu0008 13.52 -> 6.05, 9 % frames 26.7 -> 14.7, dense
28.3 -> 12.9.  Tests pass on ppc64le and x86_64.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Per 64-pixel group the 8 low planes' chunks are regrouped by two vec_perm
rounds so each doubleword holds a byte of every plane, and vgbbd (vec_gb,
an 8x8 bit transpose per doubleword) gives the pixel bytes; all-zero
groups are skipped, pixels collected from the bytes.  Byte order checked
against a reference transpose on POWER9 (no reversal needed on LE).

p9-10, ms/frame (before today -> portable byte skip -> VSX): WAu0012
11.70 -> 4.86 -> 2.21, 0.1 % frames 10.97 -> 5.56 -> 2.98, WAu0008
13.52 -> 6.02 -> 3.81, dense 28.3 -> 12.9 -> 8.4, 9 % frames 26.7 ->
14.7 -> 12.8; WAu0012 CSC 1D 12.1 -> 5.48 -> 2.78.  Tests pass on ppc64le
and x86_64.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Scattered selections (9 % frames) mispredicted the per-pixel branch;
storing every pixel and advancing by its selection is -18 % there, but a
serial chain on dense blocks where nearly every pixel passes (+89 %), so
dense blocks keep the branch.  p9-10: 9 % sparsify 12.8 -> 10.6 ms,
9 % CSC 1D 18.3 -> 15.8 ms; dense and WAu unchanged.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
POWER9's stock lz4 loses to the zero decoder in the 8-24x band as well
(8-16x 3781 vs 4812 ns/block), so off x86 the cutoff is 1: p9-10 9 %
sparsify 10.6 -> 9.3 ms, CSC 1D 15.8 -> 14.5 ms, others flat.  x86 keeps
the 8x/24x band.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
The zero decoder's SSE2 path (also taken under _M_X64) used __builtin_clz;
the dot prefetch used __builtin_prefetch; the counter-slot check used
_Static_assert.  Now _BitScanReverse / _mm_prefetch under MSVC, and a
negative-array typedef for the compile-time check.  Built with gcc 13
(-Wall -Wextra clean apart from a pre-existing padded.hpp warning).

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
jonwright and others added 5 commits October 5, 2026 11:05
The submodule pointed at a local commit (our offset 8..15 fix) that exists
on no remote, so a clone could not check it out.  It now points at upstream
v1.10.0 (ebb370ca, also lz4's release branch), and build_extension.py
compiles a copy of lib/lz4.c with patches/lz4/*.patch applied by a strict
pure-Python applier (no git/patch binary needed, e.g. on Windows; the
patches are in the source digest).  The patched copy is byte-identical to
the previous submodule commit; 605 tests pass.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…t fixes

- build_extension: apply patches with universal newlines so a core.autocrlf
  checkout still matches; .gitattributes pins *.patch to LF
- MSVC: objects, import library and .exp go to the build dir instead of the
  cwd and src/
- test_dot: skip without pyFAI instead of failing collection
- test_vs_hdf5plugin: quote-safe call of make_testcases.py

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…amplers

- patches/kcb: PE/COFF has no ifunc; on _WIN32 kcb's resolvers fill a
  function pointer from a load-time constructor (was SSE2 only with
  BITSHUF_USE_IFUNC=0: 0.1 % frames 0.745 -> 0.599 ms on an i7-8700)
- build_extension: patches for any submodule (lz4, kcb); gcc for Windows
  gets -Wa,-muse-unaligned-vector-move (GCC bug 54412: aligned AVX stack
  stores crashed the fused AVX2 collect); no -fPIC on Windows
- MANIFEST.in: ship patches/ (an sdist built unpatched lz4 silently)
- tools/sampler: the LD_PRELOAD / in-process SIGPROF samplers used on hpc6
  and POWER9, a Windows sampler (winprof.py) and a one-case driver
- notes: i7-8700 results; the gcc cross build is 2.0x MSVC

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
build-windows now runs on ubuntu-24.04 with Ubuntu's mingw-w64 gcc 13
(binutils 2.42); the MSVC build had none of the SIMD paths and was 2.1x
slower on an i7-8700.  test-windows checks the SIMD collect is active and
runs the corrupt-chunk tests too.  MSVC stays as build-windows-msvc: build +
decode matrix only, so the sources keep compiling with it.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…etuptools

test-py27 (ubuntu 18.04, gcc 7) failed: target("gfni"/"avx512vbmi2"), the
VBMI2/GFNI intrinsics and their __builtin_cpu_supports names need gcc >= 8
or clang >= 7.  BSLZ4_HAVE_VBMI_GFNI (bslz4_common.h) now guards the
low-planes VBMI/GFNI kernels and the avx512cs collect; older compilers get
stubs and capability probes that say no, so the AVX2 paths run.
BSLZ4_NO_VBMI_GFNI forces the fallback (605 tests pass either way on Zen 4;
all sources compile under gcc 7.5 in a container, the previous driver did
not: 12 errors, as in CI).

build-windows-msvc failed at `python setup.py build`: the runner's Python
3.12 has no setuptools.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
@jonwright
jonwright merged commit 057cbcc into main Oct 5, 2026
21 checks passed
@jonwright
jonwright deleted the matrix-kernels-zero-decode branch October 6, 2026 20:41
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant