Skip to content

feat: add the zarr-indexing package (TensorStore-style index transforms, ndsel wire format) - #4196

Open
d-v-b wants to merge 10 commits into
zarr-developers:mainfrom
d-v-b:feat/zarr-indexing-package
Open

feat: add the zarr-indexing package (TensorStore-style index transforms, ndsel wire format)#4196
d-v-b wants to merge 10 commits into
zarr-developers:mainfrom
d-v-b:feat/zarr-indexing-package

Conversation

@d-v-b

@d-v-b d-v-b commented Jul 28, 2026

Copy link
Copy Markdown
Contributor

Summary

This PR defines zarr-indexing, a new sub-package in the repository for lazy chunked array indexing. This package currently contains an implementation of tensorstore-style explicit-coordinate indexing, which complies with the spec here: https://github.com/zarr-developers/ndsel.

If we merge this, I would move forward on defining zarr-python's current indexing datastructures in zarr-indexing, and then making zarr-python depend on zarr-indexing. At that point we can start releasing new zarr-indexing features and bumping the minimum zarr-indexing version in zarr-python.

For reviewers

This is a big AI-written PR. That might make it boring. But it also represents progress towards a major improvement in how zarr-python models array indexing, and it's safe -- it has no impact on zarr-python itself. If you are interested in lazy indexing, please have a look at the data structures in this new package.

Author attestation

  • I am a human, these are my changes, and I have reviewed and understood every change and can explain why each is correct.

TODO

  • Add unit tests and/or doctests in docstrings
  • Add docstrings and API docs for any new/modified user-facing classes and functions
  • New/modified features documented in docs/user-guide/*.md
  • Changes documented as a new file in changes/
  • GitHub Actions have all passed
  • Test coverage is 100% (Codecov passes)

This was written by claude. See claude's PR here: d-v-b#249

d-v-b added 7 commits July 14, 2026 13:25
* fix: byte-order handling for structured dtypes in the bytes codec

The bytes codec neither byte-swapped structured-dtype fields to its
configured endian on encode (numpy reports byteorder '|' for void
dtypes, so the top-level byteorder comparison never detected a
mismatch) nor honored its endian when decoding, silently corrupting
any structured data whose field byte order differed from the stored
one (e.g. virtual references to external big-endian data).

Encode now detects byte-order mismatches by comparing full dtypes via
newbyteorder, and decode reinterprets raw bytes in the stored byte
order before converting to the data type's declared byte order, so the
stored layout (codec state) and the in-memory layout (array data type)
are independent.

Closes zarr-developers#4141

Assisted-by: ClaudeCode:claude-fable-5

* test: fold structured byte-order cases into existing bytes codec tests

Extend test_endian's parametrization with structured dtypes and
test_bytes_codec_sync_roundtrip with endian/dtype parametrization plus
stored-layout and decoded-dtype assertions, instead of adding parallel
test functions for the same properties.

Assisted-by: ClaudeCode:claude-fable-5

* refactor: rename stored_dtype to view_dtype in BytesCodec decode

The variable is the dtype used to view the raw chunk bytes (byte order
from the codec's endian configuration), not a property of the stored
data or of the returned buffer, which always carries the array's
declared dtype.

Assisted-by: ClaudeCode:claude-fable-5

* docs: note that the decode-side byte-order conversion copies the chunk

Assisted-by: ClaudeCode:claude-fable-5
Standalone workspace package extracted from the lazy-indexing branch
(zarr-developers#3906): composable, lazy coordinate transforms
(IndexTransform / IndexDomain / output maps), dependency-aware chunk
resolution against a DimensionGridLike protocol, and an ndsel-conformant
JSON wire format validated against the vendored conformance corpus.

zarr itself does not depend on zarr-indexing yet — the runtime wiring
lands separately once 0.1.0 is published. The package is numpy-only;
its tests exercise chunk resolution against zarr's concrete ChunkGrid,
so they run from the workspace root (uv sync --all-packages).

Assisted-by: ClaudeCode:claude-fable-5
@github-actions github-actions Bot added the needs release notes Automatically applied to PRs which haven't added release notes label Jul 28, 2026
@d-v-b

d-v-b commented Jul 28, 2026

Copy link
Copy Markdown
Contributor Author

ping @jbms in case you have comments or suggestions

@d-v-b d-v-b changed the title feat: add the zarr-indexing package (TensorStore-style index transforms, ndsel wire format) - #249 feat: add the zarr-indexing package (TensorStore-style index transforms, ndsel wire format) Jul 28, 2026
@d-v-b
d-v-b marked this pull request as ready for review July 28, 2026 20:03
@codecov

codecov Bot commented Jul 28, 2026

Copy link
Copy Markdown

Codecov Report

✅ All modified and coverable lines are covered by tests.
✅ Project coverage is 93.90%. Comparing base (0c07bf5) to head (d1e931f).
⚠️ Report is 5 commits behind head on main.

Additional details and impacted files
@@            Coverage Diff             @@
##             main    #4196      +/-   ##
==========================================
+ Coverage   93.88%   93.90%   +0.02%     
==========================================
  Files          91       91              
  Lines       12622    12672      +50     
==========================================
+ Hits        11850    11900      +50     
  Misses        772      772              

see 6 files with indirect coverage changes

🚀 New features to boost your workflow:
  • ❄️ Test Analytics: Detect flaky tests, report on failures, and find test suite problems.

Candidate-chunk enumeration took the cartesian product of each correlated
ArrayMap's per-dimension distinct chunk ids and relied on intersect() to
filter untouched combinations. For a diagonal selection of P scattered
points that is P**2 intersect calls — quadratic in the number of selected
points, the same workload shape as zarr-developers#4174 (400 points:
~2.6s; 10k points: ~30min).

Group correlated maps jointly instead: broadcast their per-point chunk
ids, take the distinct rows (np.unique(axis=0), O(P log P)), and
enumerate exactly the touched combinations. Candidate slots now carry
chunk-coordinate tuples covering one or more output dimensions;
orthogonal/constant/slice dimensions keep their existing per-dimension
candidates. 400-point diagonal resolution drops from 2628ms to 14ms and
scales linearly.

Assisted-by: ClaudeCode:claude-fable-5
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

needs release notes Automatically applied to PRs which haven't added release notes

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant