Skip to content

Jonwright/corrupt chunk test - #24

Merged
jonwright merged 3 commits into
mainfrom
jonwright/corrupt-chunk-test
Sep 30, 2026
Merged

jonwright merged 3 commits into
mainfrom
jonwright/corrupt-chunk-test

Conversation

@jonwright

Copy link
Copy Markdown
Owner

With thanks to @jadball and the tape archive for the motivating corrupt hdf5 file example

jonwright and others added 3 commits September 30, 2026 18:08
The field files (scan1107 chunk 5172, scan2169 chunk 2669) each have one chunk
whose bitshuffle-LZ4 stream loses framing (a block length word holds garbage).
test_corrupt_chunk.py reproduces the classes of damage on a hand-built chunk
placed against a PROT_NONE guard page. Damaged length words raise cleanly;
truncated chunks segfault and are marked xfail until the decoder bounds-checks.

Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
The decoder took each block's length word from the data and never compared it
with the chunk size, so a corrupt or truncated chunk could read past its end
(SIGSEGV on truncation) or into a neighbouring chunk of a shared buffer.

Core (bslz4_core.hpp, both plain and CSC decoders):
- a chunk shorter than its 12 byte header, or with a negative length, is
  rejected before any header is read;
- each block length word must lie inside the chunk, and the block it announces
  must fit in what remains;
- the raw tail must start at or after the last compressed block.
All return the new ERR_CORRUPT_CHUNK (-107).

Base-buffer variant (bslz4_csc_multi_base_*): takes base.len and checks every
(offset, size) from HDF5 metadata against it, overflow-safe, before any pointer
is formed and before offsets is mutated. Returns ERR_BAD_CHUNK_BOUNDS (-108).
Spec also checks lengths.n == offsets.n (and compressed_lengths.n ==
compressed_ptrs.n). note_chunk marks a chunk > INT32_MAX as -1 (rejected)
instead of truncating its length.

Python: error codes now map to messages. A failed batch says at least one chunk
in it is bad, that its output must not be used, and to decode one at a time;
it does not try to identify the frame. npx_out is zeroed on failure so partial
counts cannot be mistaken for results.

Generated files: src/bslz4_to_sparse.cpp now begins with a note that it is
generated by tools/regenerate_source.py. The generator had drifted from the
committed .cpp (the Windows 'L'/'l' format-char checks from a2c678e were only
hand-edited into the output and would have been lost on regeneration); the
generator now carries them. Wrapper regenerated with c2py23 v0.5.8 (ee81c97,
current upstream head).

Tests: test_corrupt_chunk.py no longer xfails; adds base-buffer bounds cases
and a no-cross-talk case (first chunk declared short in a shared buffer). Run
on the Linux CI jobs. Decode speed unchanged (kcb 0.38 ms/frame before/after).

Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
fixture/test-linux/test-windows/test-macos only use test/ and an installed
wheel, so they no longer clone the four submodules (one less thing to fail).
Only the jobs that compile keep recursive checkout. lz4 was the one submodule
pointing at plain http://, so use https like the others.

Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
@jonwright
jonwright merged commit 32f60b7 into main Sep 30, 2026
40 checks passed
@jonwright
jonwright deleted the jonwright/corrupt-chunk-test branch September 30, 2026 17:49
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant