Skip to content

DEV-1347: Fix TDB1 query failures when a writer promotes the same block - #7

Merged
razvan-danit-tq merged 2 commits into
base-jena-6.2.0from
DEV-1230-bptree
Sep 23, 2026
Merged

razvan-danit-tq merged 2 commits into
base-jena-6.2.0from
DEV-1230-bptree

Conversation

@razvan-danit-tq

@razvan-danit-tq razvan-danit-tq commented Sep 23, 2026 •

Copy link
Copy Markdown

Internal patched build of Apache Jena, published as 6.2.0-tq-2. Cumulative on top of 6.2.0-tq-1: it carries the journal write-back fix from #5 as well as this one.

BPTreeNodeMgr.formatBPTreeNode decoded a node with relative position(), limit(), slice() and rewind() calls on block.getByteBuffer(). That returns the block's single buffer instance, and the same Block is shared across transactions: after a writer commits while readers are active, later transactions are stacked on its BlockMgrJournal, whose getRead hands out its writeBlocks entries directly (BlockMgrCache.getRead shares cached blocks the same way in direct file mode).

When the next writer takes such a block for writing, BlockMgrJournal._promote copies it via Block.replicate, which moves the shared buffer's position and limit without holding the block monitor. A reader decoding the same block at that moment slices the wrong byte range and reads a garbage child pointer. It surfaces as BlockException: BlockAccessBase: Bounds exception on node2id.dat, from find() in a read transaction with no journal replay involved.

The fix uses absolute slice(index, length), so decoding neither reads nor changes the shared buffer's position. Six lines in, twelve out, and no lock.

Tracked as DEV-1347. This is the second race described in DEV-1230, the one the write-back patch does not address; it has its own ticket because it is a separate Jena bug with its own symptoms. The approach is the one sketched in DEV-1230.

Validation

  • jena-tdb1 suite on this branch: 1005 tests, 0 failures.
  • Reproduction harness from DEV-1230, 10 runs of 300s per build on 4 pinned cores: 6.2.0-tq-1 failed 2 of 10, both BlockException on node2id.dat with no data loss; 6.2.0-tq-2 failed 0 of 10.
  • Independent earlier run at 60s on 12 cores: 5 of 60 versus 0 of 60.

Not urgent. Master is on 6.2.0-tq-1 and that build is being observed. This is ready whenever it is wanted; tq-1 remains published and tagged, so moving between them is a one-line change to ver.jena.

(Supersedes #6, which was merged prematurely by mistake and has been undone — base-jena-6.2.0 is back at the tq-1 state.)

🤖 Generated with Claude Code

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
@razvan-danit-tq razvan-danit-tq changed the title DEV-1230: Fix B+tree node decoding that corrupts concurrent reads DEV-1347: Fix B+tree node decoding that corrupts concurrent reads Sep 23, 2026

@cygri cygri left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

The commit message states that there is a race between two readers. Claude tells me this is wrong and a reader race cannot happen; it is a race between a reader and a writer. It proposes this commit message instead:

TDB1: decode B+tree nodes without mutating the shared block buffer

formatBPTreeNode decoded a node with relative position(), limit(), slice() and
rewind() calls on block.getByteBuffer(). That returns the block's single buffer
instance, and the same Block is shared across transactions: after a writer
commits while readers are active, later transactions are stacked on its
BlockMgrJournal, whose getRead hands out its writeBlocks entries directly
(BlockMgrCache.getRead shares cached blocks the same way in direct file mode).

When the next writer takes such a block for writing, BlockMgrJournal._promote
copies it via Block.replicate, which moves the shared buffer's position and
limit without holding the block monitor. A reader decoding the same block at
that moment slices the wrong byte range and reads a garbage child pointer. It
surfaces as BlockException: BlockAccessBase: Bounds exception on node2id.dat,
from find() in a read transaction with no journal replay involved.

Use absolute slice(index, length) instead, so decoding neither reads nor
changes the shared buffer's position.

Reported in DEV-1230; the fix is the one sketched there.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

A similar inaccuracy is in the PR description.

@razvan-danit-tq razvan-danit-tq changed the title DEV-1347: Fix B+tree node decoding that corrupts concurrent reads DEV-1347: Fix TDB1 query failures when a writer promotes the same block Sep 23, 2026
formatBPTreeNode decoded a node with relative position(), limit(), slice() and
rewind() calls on block.getByteBuffer(). That returns the block's single buffer
instance, and the same Block is shared across transactions: after a writer
commits while readers are active, later transactions are stacked on its
BlockMgrJournal, whose getRead hands out its writeBlocks entries directly
(BlockMgrCache.getRead shares cached blocks the same way in direct file mode).

When the next writer takes such a block for writing, BlockMgrJournal._promote
copies it via Block.replicate, which moves the shared buffer's position and
limit without holding the block monitor. A reader decoding the same block at
that moment slices the wrong byte range and reads a garbage child pointer. It
surfaces as BlockException: BlockAccessBase: Bounds exception on node2id.dat,
from find() in a read transaction with no journal replay involved.

Use absolute slice(index, length) instead, so decoding neither reads nor
changes the shared buffer's position.

Reported in DEV-1230 and split out as DEV-1347; the fix is the one sketched
there.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
@razvan-danit-tq
razvan-danit-tq merged commit 4540dd2 into base-jena-6.2.0 Sep 23, 2026
2 checks passed
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants