Skip to content

DEV-1230: Fix TDB1 write-back race that loses committed data - #5

Merged
razvan-danit-tq merged 2 commits into
base-jena-6.2.0from
DEV-1230
Sep 21, 2026
Merged

razvan-danit-tq merged 2 commits into
base-jena-6.2.0from
DEV-1230

Conversation

@razvan-danit-tq

@razvan-danit-tq razvan-danit-tq commented Sep 21, 2026 •

Copy link
Copy Markdown

Internal patched build of Apache Jena 6.2.0, published as 6.2.0-tq-1.

Two commits: the POM version bump across all modules, then Richard's fix cherry-picked unchanged from tdb1-txn-start-race (4634b5c). It applied to the 6.2.0 release tag without conflicts.

The fix records a transaction's start under the same monitor that created it, so a reader finishing in that window can no longer replay the journal underneath a starting transaction and cause committed quads to be lost.

Validation on this branch: jena-tdb1 suite 1005 tests, 0 failures, matching the baseline on the ticket. A deterministic regression test in TopBraid-Suite fails 20/20 against stock 6.2.0 (180 of 240 committed quads lost) and passes 20/20 against this build.

Not for upstream yet — per DEV-1230 we use it internally first.

Intending to merge this into base-jena-6.2.0 rather than close it and reference the tag, following the 5.6.0-tq-1 precedent. That way a later -tq-N branched off the base inherits the fix instead of silently losing it. The jena-6.2.0-tq-1 tag gets created either way.

🤖 Generated with Claude Code

razvan-danit-tq and others added 2 commits September 18, 2026 16:19
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
beginInternal() called begin$() (synchronized) and then noteTxnStart()
after the monitor had been released. A transaction finishing in that gap
(notifyCommit/notifyAbort are synchronized) observed
activeReaders == 0 && activeWriters == 0 and ran processDelayedReplayQueue,
which clears the queued transactions' BlockMgrJournal write sets and then
replays the journal into the base dataset. The new transaction, built on
the last queued transaction's view (determineBaseDataset), meanwhile read
blocks that were in neither place and saw stale, pre-commit content. For a
writer this surfaces as "Secondary index duplicate: GSPO->..." on a later
add, and once such a writer commits its stale blocks the loss is durable:
committed quads vanish from some indexes, or node table lookups fail.

Move noteTxnStart() under the same monitor as begin$() (beginSync$), and
make promoteExec$() synchronized, since READ_COMMITTED promotion reaches
it with the same gap. Nothing in the moved code can block: the only lock
taken is the base dataset's MRSW lock in READ mode, which no code path in
jena-tdb1 takes in WRITE mode.

Reproduced with a stress harness on unmodified jena-tdb1 6.2.0 (4 of 4
on-disk runs of 90-150 s corrupted committed data) and deterministically
with fault-injection hooks. With this change: jena-tdb1 test suite passes
(1005 tests), the deterministic scenario no longer loses data, and the
stress harness runs clean, including with promoted transactions.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_013KmWBMLwuUqWg8pQtHHe8x
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants