Skip to content

fix(postgres): lock TransactWriteItems items in key order, cancel on deadlock - #385

Open
yesyayen wants to merge 6 commits into
ExtendDB:mainfrom
yesyayen:fix/f7-twi-deadlock
Open

yesyayen wants to merge 6 commits into
ExtendDB:mainfrom
yesyayen:fix/f7-twi-deadlock

Conversation

@yesyayen

@yesyayen yesyayen commented Oct 5, 2026

Copy link
Copy Markdown
Member

What

On PostgreSQL, two TransactWriteItems that touch the same items in different orders could deadlock. PostgreSQL aborted one of them after deadlock_timeout (1 s by default), and ExtendDB returned it as HTTP 500 Internal server error.

  • transact_write_items_impl() in crates/storage-postgres/src/data/transactions.rs now runs the ops sorted by table and key, so every write transaction locks its items in the same order and two of them cannot deadlock. Items that do not exist yet are covered too: the create-race insert waits at the same point in the order. Cancellation reasons, stream records, and GSI writes keep request order.
  • If PostgreSQL still aborts the transaction with SQLSTATE 40P01 or 40001 (a deadlock with some other lock holder, or a stricter default_transaction_isolation), the request is canceled with a TransactionConflict reason on the item that hit it. The transaction helpers now map database errors through db_error so the SQLSTATE survives.
  • The earliest invalid op in request order is still the one reported, and the loop stops once that answer is known.
  • New CancellationReason::transaction_conflict(), also used by the two existing create-race cancellations.
  • Docs: a row in differences-from-dynamodb.md (on PostgreSQL a conflicting transaction waits for the lock instead of being canceled) and a lock-order bullet in the storage extension guide. CI runs the new storage tests in the PostgreSQL storage-level step.

SQLite runs one writer at a time and MongoDB retries write conflicts, so neither backend has this bug.

Why

Found with a contended TransactWriteItems load test on 0.1.12, 2 items per transaction over 100 hot keys: 1, 2, and 16 HTTP 500s at 500, 1000, and 2000 rps, and each deadlock stalled those keys for 1 s, so p99 went to 0.6 to 3.1 s.

Amazon DynamoDB (captured 2026-10-01) cancels the losing transaction with HTTP 400 TransactionCanceledException, message Transaction cancelled, please refer cancellation reasons for specific reasons [TransactionConflict, None], with {"Code": "TransactionConflict", "Message": "Transaction is ongoing for the item"} on the contended item and {"Code": "None"} on the others. That tells the caller nothing was written and which item conflicted. A 500 says the outcome is unknown, and SDKs retry it as a server error (botocore's standard retry mode retries a 500 and does not retry the cancellation).

Fixes: n/a, found by a load test, no issue filed

Result

  • Conformance: the cancellation shape above was captured on Amazon DynamoDB on 2026-10-01, and the new dual-target tests pass there in 10 of 10 runs.
  • Tests: the new pytest passes on PostgreSQL, SQLite, and MongoDB, and fails on main with 10 to 23 HTTP 500s per run. 4 new PostgreSQL storage tests pass, and 2 of them fail on main. cargo test --workspace: 1,230 passed.
  • Performance: 8 clients, 50 transactions each, every transaction adds 1 to the same two items, half the clients name them (a, b) and half (b, a), SDK retries off, release build on one host over loopback:
main this PR
HTTP 500 316 of 400 0
p50 3.0 s 5.6 ms
p99 7.0 s 105 ms

Testing done

  • tests/test_transact_write_conflict.py (new, dual-target): 4 clients run opposite-order and shared-plus-private transactions with SDK retries off. Every failure must be the cancellation above, a run may see at most 5 HTTP 500s (Amazon DynamoDB itself returns a few InternalServerError under this contention, about 1 in 1000 transactions), and the counters must match the commits.
  • crates/storage-postgres/tests/twi_conflict.rs (new, needs EXTENDDB_TEST_PG_CONNECTION_STRING): a forced deadlock with an outside lock holder returns [TransactionConflict, None] at the right request position and writes nothing; 8 writers x 50 opposite-order transactions finish with no error; the earliest invalid op in request order is reported; an invalid first op does not wait on a later op's lock.
  • Transaction, condition, stream, GSI, batch, and item pytest suites against a server built from this branch: 298 passed, 1 xfailed.
cargo fmt --all -- --check
cargo clippy --all-targets -- -D warnings
cargo test --workspace          # 1,230 passed, with a live PostgreSQL

Checklist

  • I have read CONTRIBUTING.md
  • All tests pass (cargo test --workspace)
  • Code is formatted (cargo fmt --check)
  • Clippy is clean (cargo clippy -- -W clippy::pedantic)
  • I have added or updated tests for new functionality
  • I have updated documentation if behavior changed
  • Breaking changes are noted below (if any)
  • If this changes the wire protocol, Storage trait, auth model, on-disk
    format, or public CLI surface, an RFC has been accepted or is linked
    below. Otherwise, an ADR captures the decision (link below).

ADR / RFC: n/a. Lock order inside one backend's write transaction; no wire, trait, auth, on-disk, or CLI change.

Breaking changes

None.


By submitting this pull request, I confirm that my contribution is made under
the terms of the Apache License 2.0 and I agree to the Developer Certificate of
Origin (DCO). See CONTRIBUTING.md for details.

…g them

Two TransactWriteItems that name the same items in different orders could deadlock in PostgreSQL. The victim came back as HTTP 500, and each deadlock held its row locks for deadlock_timeout first.

Run the ops in table and key order, so every write transaction takes its row locks in one order and two of them cannot deadlock on each other. The results keep their request positions.

When PostgreSQL still aborts the transaction with deadlock_detected (40P01) or serialization_failure (40001), cancel it with a TransactionConflict reason on the item that hit the abort, the shape Amazon DynamoDB returns. An abort outside the per-item work names every item. To read the SQLSTATE, the transaction helpers now map database errors through db_error.

Assisted-by: pi claude-opus-5-5
… known

With the ops running in key order, a request whose earliest invalid op is known no longer runs and locks the rest. A storage test pins that the earliest invalid op in request order is still the one reported, and CI now runs the conflict storage tests against its PostgreSQL.

Assisted-by: pi claude-opus-5-5
Bound the tolerated InternalServerError count to a fixed 5 per run instead of a share of the cancellations, so a server that cancels a lot cannot hide a 5xx rate far above the service's.

Assisted-by: pi claude-opus-5-5
Add the contention difference to differences-from-dynamodb.md and the lock order rule for blocking-lock backends to the storage extension guide.

Assisted-by: pi claude-opus-5-5
…ocks

A storage test holds the next op's row from outside the transaction and expects the ValidationException within 5 s. The pytest docstring now says which test can deadlock, and the file ends with one newline.

Assisted-by: pi claude-opus-5-5

This branch has not been deployed

No deployments
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant