Skip to content

DEV-1230: Fix B+tree node decoding that corrupts concurrent reads - #6

Merged
razvan-danit-tq merged 2 commits into
base-jena-6.2.0from
DEV-1230-bptree
Sep 23, 2026
Merged

razvan-danit-tq merged 2 commits into
base-jena-6.2.0from
DEV-1230-bptree

Conversation

@razvan-danit-tq

Copy link
Copy Markdown

Internal patched build of Apache Jena, published as 6.2.0-tq-2. Cumulative on top of 6.2.0-tq-1: it carries the journal write-back fix from #5 as well as this one.

BPTreeNodeMgr.formatBPTreeNode decoded a node with relative position(), limit(), slice() and rewind() calls on block.getByteBuffer(). That returns the block's single buffer instance, and both BlockMgrJournal.getRead and BlockMgrCache.getRead hand the same Block to every concurrent reader, so two readers decoding the same node tore each other's view and read a garbage child pointer. It surfaces as BlockException: BlockAccessBase: Bounds exception on node2id.dat, thrown from find() in a read transaction with no journal replay involved.

The fix uses absolute slice(index, length), so decoding reads the buffer without changing state other threads can see. Six lines in, twelve out, and no lock.

This is the second race described in DEV-1230, the one the write-back patch does not address. The approach is the one sketched in that ticket.

Validation

  • jena-tdb1 suite on this branch: 1005 tests, 0 failures.
  • Reproduction harness from DEV-1230, 10 runs of 300s per build on 4 pinned cores: 6.2.0-tq-1 failed 2 of 10, both BlockException on node2id.dat with no data loss; 6.2.0-tq-2 failed 0 of 10.
  • Independent earlier run at 60s on 12 cores: 5 of 60 versus 0 of 60.

Not urgent. Master is on 6.2.0-tq-1 and Richard is observing it. This is ready whenever it is wanted; tq-1 remains published and tagged, so moving between them is a one-line change to ver.jena.

🤖 Generated with Claude Code

razvan-danit-tq and others added 2 commits September 23, 2026 15:20
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
formatBPTreeNode decoded a node with relative position(), limit(), slice() and
rewind() calls on block.getByteBuffer(). That returns the block's single buffer
instance, and both BlockMgrJournal.getRead and BlockMgrCache.getRead hand the
same Block to every concurrent reader, so two readers decoding the same node
tore each other's view and read a garbage child pointer. It surfaces as
BlockException: BlockAccessBase: Bounds exception on node2id.dat, from find()
in a read transaction with no journal replay involved.

Use absolute slice(index, length) instead, so decoding reads the buffer without
changing any state other threads can see.

Reported in DEV-1230; the fix is the one sketched there.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
@razvan-danit-tq
razvan-danit-tq merged commit 1d0e179 into base-jena-6.2.0 Sep 23, 2026
2 checks passed
@razvan-danit-tq
razvan-danit-tq deleted the DEV-1230-bptree branch September 23, 2026 12:47
@razvan-danit-tq
razvan-danit-tq restored the DEV-1230-bptree branch September 23, 2026 12:49
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant