Conversation
…g them Two TransactWriteItems that name the same items in different orders could deadlock in PostgreSQL. The victim came back as HTTP 500, and each deadlock held its row locks for deadlock_timeout first. Run the ops in table and key order, so every write transaction takes its row locks in one order and two of them cannot deadlock on each other. The results keep their request positions. When PostgreSQL still aborts the transaction with deadlock_detected (40P01) or serialization_failure (40001), cancel it with a TransactionConflict reason on the item that hit the abort, the shape Amazon DynamoDB returns. An abort outside the per-item work names every item. To read the SQLSTATE, the transaction helpers now map database errors through db_error. Assisted-by: pi claude-opus-5-5
… known With the ops running in key order, a request whose earliest invalid op is known no longer runs and locks the rest. A storage test pins that the earliest invalid op in request order is still the one reported, and CI now runs the conflict storage tests against its PostgreSQL. Assisted-by: pi claude-opus-5-5
Bound the tolerated InternalServerError count to a fixed 5 per run instead of a share of the cancellations, so a server that cancels a lot cannot hide a 5xx rate far above the service's. Assisted-by: pi claude-opus-5-5
Add the contention difference to differences-from-dynamodb.md and the lock order rule for blocking-lock backends to the storage extension guide. Assisted-by: pi claude-opus-5-5
…ocks A storage test holds the next op's row from outside the transaction and expects the ValidationException within 5 s. The pytest docstring now says which test can deadlock, and the file ends with one newline. Assisted-by: pi claude-opus-5-5
Assisted-by: pi claude-opus-5-5
yesyayen
marked this pull request as ready for review
October 5, 2026 17:54
yesyayen
requested review from
LeeroyHannigan,
amrith,
c33howard,
jcshepherd,
pdf-amzn and
robinnsc
as code owners
October 5, 2026 17:54
This was referenced Oct 5, 2026
This branch has not been deployed
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
What
On PostgreSQL, two TransactWriteItems that touch the same items in different orders could deadlock. PostgreSQL aborted one of them after
deadlock_timeout(1 s by default), and ExtendDB returned it as HTTP 500Internal server error.transact_write_items_impl()incrates/storage-postgres/src/data/transactions.rsnow runs the ops sorted by table and key, so every write transaction locks its items in the same order and two of them cannot deadlock. Items that do not exist yet are covered too: the create-race insert waits at the same point in the order. Cancellation reasons, stream records, and GSI writes keep request order.40P01or40001(a deadlock with some other lock holder, or a stricterdefault_transaction_isolation), the request is canceled with aTransactionConflictreason on the item that hit it. The transaction helpers now map database errors throughdb_errorso the SQLSTATE survives.CancellationReason::transaction_conflict(), also used by the two existing create-race cancellations.differences-from-dynamodb.md(on PostgreSQL a conflicting transaction waits for the lock instead of being canceled) and a lock-order bullet in the storage extension guide. CI runs the new storage tests in the PostgreSQL storage-level step.SQLite runs one writer at a time and MongoDB retries write conflicts, so neither backend has this bug.
Why
Found with a contended TransactWriteItems load test on 0.1.12, 2 items per transaction over 100 hot keys: 1, 2, and 16 HTTP 500s at 500, 1000, and 2000 rps, and each deadlock stalled those keys for 1 s, so p99 went to 0.6 to 3.1 s.
Amazon DynamoDB (captured 2026-10-01) cancels the losing transaction with HTTP 400
TransactionCanceledException, messageTransaction cancelled, please refer cancellation reasons for specific reasons [TransactionConflict, None], with{"Code": "TransactionConflict", "Message": "Transaction is ongoing for the item"}on the contended item and{"Code": "None"}on the others. That tells the caller nothing was written and which item conflicted. A 500 says the outcome is unknown, and SDKs retry it as a server error (botocore's standard retry mode retries a 500 and does not retry the cancellation).Fixes: n/a, found by a load test, no issue filed
Result
cargo test --workspace: 1,230 passed.Testing done
tests/test_transact_write_conflict.py(new, dual-target): 4 clients run opposite-order and shared-plus-private transactions with SDK retries off. Every failure must be the cancellation above, a run may see at most 5 HTTP 500s (Amazon DynamoDB itself returns a fewInternalServerErrorunder this contention, about 1 in 1000 transactions), and the counters must match the commits.crates/storage-postgres/tests/twi_conflict.rs(new, needsEXTENDDB_TEST_PG_CONNECTION_STRING): a forced deadlock with an outside lock holder returns[TransactionConflict, None]at the right request position and writes nothing; 8 writers x 50 opposite-order transactions finish with no error; the earliest invalid op in request order is reported; an invalid first op does not wait on a later op's lock.Checklist
cargo test --workspace)cargo fmt --check)cargo clippy -- -W clippy::pedantic)Storagetrait, auth model, on-diskformat, or public CLI surface, an RFC has been accepted or is linked
below. Otherwise, an ADR captures the decision (link below).
ADR / RFC: n/a. Lock order inside one backend's write transaction; no wire, trait, auth, on-disk, or CLI change.
Breaking changes
None.
By submitting this pull request, I confirm that my contribution is made under
the terms of the Apache License 2.0 and I agree to the Developer Certificate of
Origin (DCO). See CONTRIBUTING.md for details.