Skip to content

fix(postgres): order a PutItem that loses a create race after the winner - #388

Open
yesyayen wants to merge 8 commits into
ExtendDB:mainfrom
yesyayen:fix/put-create-race
Open

yesyayen wants to merge 8 commits into
ExtendDB:mainfrom
yesyayen:fix/put-create-race

Conversation

@yesyayen

@yesyayen yesyayen commented Oct 5, 2026

Copy link
Copy Markdown
Member

What

On PostgreSQL, a PutItem that lost the race to create a new item returned ConditionalCheckFailedException, even when it had no condition. It now orders itself after the winner, as UpdateItem already does.

  • put_item_impl() in crates/storage-postgres/src/data/put_item.rs: when the INSERT ... ON CONFLICT DO NOTHING of a missing item affects no row, the put re-reads the row FOR UPDATE, checks its condition against the winner, and overwrites it. ReturnValues ALL_OLD returns the winner, and the stream record and index maintenance use it as the old image. If the winner was deleted before the re-read, the insert is retried, up to 5 inserts, the same bound as UpdateItem.
  • This is the transactional path. A PutItem takes it when it has a condition, ReturnValues ALL_OLD, ReturnConsumedCapacity, a secondary or vector index, or a stream. BatchWriteItem puts take it on the same tables, and with ReturnConsumedCapacity INDEXES. TransactWriteItems already handled this race.
  • Docs: one bullet on missing items in section 6.3 of docs/design/02-high-level-design.md.
  • CI runs the new storage tests in the PostgreSQL storage-level step.

SQLite runs one writer at a time, so it does not have this bug. On MongoDB the losing put gets HTTP 500 instead, and a separate PR fixes that.

Why

Found while checking TransactGetItems isolation: 8 clients put the same new key at once. Measured on c178814:

Request shape (8 writers per round) Amazon DynamoDB PostgreSQL on main: ConditionalCheckFailed
ReturnValues ALL_OLD, hash table 200 of 200 succeed 44 of 200
ReturnValues ALL_OLD, hash-range table 200 of 200 succeed 32 of 200
condition attribute_not_exists(zz), which every racer passes, hash table 200 of 200 succeed 130 of 200
same condition with ALL_OLD, hash-range table 200 of 200 succeed 81 of 200
unconditional, table with a GSI 200 of 200 succeed 83 of 200
unconditional, table with a stream 200 of 200 succeed 83 of 200
BatchWriteItem with one PutRequest, table with a stream (10 rounds) 80 of 80 succeed 41 of 80
ReturnConsumedCapacity TOTAL (10 rounds) 80 of 80 succeed 21 of 80
condition attribute_not_exists(pk) (control) 25 of 200 succeed, 1 per round 25 of 200 succeed, 1 per round

On Amazon DynamoDB the ALL_OLD images of one round form a single chain, from no item to the final one, so the puts were applied one after another.

Fixes: n/a, found by a concurrency probe, no issue filed

Result

  • Conformance: the Amazon DynamoDB column was captured on 2026-10-05, and the new dual-target test passes there, 9 of 9.
  • Tests: test_put_item_create_race.py (9 tests) fails 8 of 9 on main. It passes 9 of 9 on this branch in every run, on SQLite, and on MongoDB with its fix. 7 of the 8 new PostgreSQL storage tests fail on main, and all 8 pass here. Two mutants, one that drops the old image for the index updates and one that drops it for the stream record, each fail the matching tests. cargo test --release --workspace: 1,230 passed.
  • Performance: not measured, the change only adds a re-read on the losing side of a create race.

Testing done

  • tests/test_put_item_create_race.py (new, dual-target): 8 clients put the same new key, 15 rounds per test. Shapes: ReturnValues ALL_OLD on hash and hash-range tables, ReturnConsumedCapacity, a condition every racer passes, a GSI table, an LSI table, a stream table, BatchWriteItem on the stream table, and the attribute_not_exists(pk) control. Every put must succeed unless its own condition fails, and the ALL_OLD images must form one chain. The GSI and LSI must hold only the final item's entry. For each key, the stream must hold one INSERT and then MODIFY records in sequence order, each with the record before it as its old image.
  • crates/storage-postgres/tests/put_create_race.rs (new, needs EXTENDDB_TEST_PG_CONNECTION_STRING): an outside transaction holds an uncommitted insert of the item, and the test commits it once the put waits on it. The put must overwrite the winner and return it as ALL_OLD on hash and hash-range tables, pass a condition that holds for the winner, and fail a condition that the winner breaks, returning the winner. Four more tests park the put's insert behind a trigger and delete the winner before the put re-reads it, on hash and hash-range tables: the put retries and creates the item, and after 5 lost inserts it gives up with nothing written.
  • Item, conditional, stream, GSI, batch, concurrency and transaction pytest suites against a server built from this branch: 287 passed, 1 xfailed, and the same 2 GSI propagation-delay failures as on main.
cargo fmt --all -- --check
cargo clippy --release --all-targets -- -D warnings
cargo test --release --workspace                    # 1,230 passed
cargo test --release -p extenddb-storage-postgres   # with a live PostgreSQL

Checklist

  • I have read CONTRIBUTING.md
  • All tests pass (cargo test --workspace)
  • Code is formatted (cargo fmt --check)
  • Clippy is clean (cargo clippy -- -W clippy::pedantic)
  • I have added or updated tests for new functionality
  • I have updated documentation if behavior changed
  • Breaking changes are noted below (if any)
  • If this changes the wire protocol, Storage trait, auth model, on-disk
    format, or public CLI surface, an RFC has been accepted or is linked
    below. Otherwise, an ADR captures the decision (link below).

ADR / RFC: n/a. One backend's create-race handling; no wire, trait, auth, on-disk, or CLI change.

Merge order: after the MongoDB fix for the same race, without which the new test gets HTTP 500s on MongoDB. The CI step conflicts by one line with #385: keep both --test flags.

Breaking changes

None.


By submitting this pull request, I confirm that my contribution is made under
the terms of the Apache License 2.0 and I agree to the Developer Certificate of
Origin (DCO). See CONTRIBUTING.md for details.

Racing puts to a new item must all succeed unless a put's own condition fails against the item before it, and their ReturnValues ALL_OLD images must form one chain, as on Amazon DynamoDB. The tests cover hash and range tables, a condition, a GSI table, and a stream table.

Assisted-by: pi claude-opus-5-5
A PutItem takes the transactional path when it has a condition, ReturnValues ALL_OLD, ReturnConsumedCapacity, an index, a vector index, or a stream. BatchWriteItem puts take the same path on the same tables, and with ReturnConsumedCapacity INDEXES. On that path a put inserts a missing item with ON CONFLICT DO NOTHING. When a concurrent writer created the item first, the put returned ConditionalCheckFailed, even with no condition. BatchWriteItem returned it too, which Amazon DynamoDB never does.

The put now re-reads the winner's row, checks its condition against it, and overwrites it, as UpdateItem already does. ALL_OLD, the index updates, and the stream record get the winner as the old image. If the winner was deleted before the re-read, the put retries the insert, up to 5 inserts. Storage tests hold the competing create open against a live PostgreSQL, and CI runs them.

Assisted-by: pi claude-opus-5-5
The GSI and LSI tests give each put its own index key and check that only the final item's entry stays in the index. The stream tests check every round in sequence order: each record's old image is the record before it. New tests race BatchWriteItem puts on the stream table, and PutItem calls with ReturnConsumedCapacity.

Assisted-by: pi claude-opus-5-5
Two storage tests park the put's insert behind a trigger while an outside transaction commits a winner and a second one locks it. The winner is then deleted before the put re-reads it. In one test the put retries its insert and creates the item. In the other it loses five inserts and returns an internal error, with nothing written. The create race tests now wait for a backend that the creator blocks, not for any lock waiter.

Assisted-by: pi claude-opus-5-5
The high-level design says that a missing item has no row to lock, and how a PutItem or UpdateItem that loses the insert orders itself after the winner. Two comments in the put code now match what the code does.

Assisted-by: pi claude-opus-5-5
The deleted-winner and give-up tests now run on a hash table and on a hash and range table, so the retry arm and the bound of the range branch have tests too. When a driver step fails, the driver opens the gate, and the test reports what the put returned instead of only that no backend waited.

Assisted-by: pi claude-opus-5-5
The module doc of the storage tests still said two.

Assisted-by: pi claude-opus-5-5
On MongoDB, a put with no condition on a table with no index and no stream returns HTTP 500 when it loses the create race, and the MongoDB write-race fix repairs it. That fix lands as its own PR. Until it is on main, the three affected tests are marked xfail when the MongoDB test runner runs them. The marker is not strict, so the tests also pass once the fix is in. Remove the marker then.

Assisted-by: pi claude-opus-5-5
@yesyayen
yesyayen marked this pull request as ready for review October 5, 2026 22:31

This branch has not been deployed

No deployments
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant