perf(db): index MetadataFiles/ExtraFiles by (AuthorId, BookId) - #267
Open
jordanfelle wants to merge 1 commit into
Open
jordanfelle wants to merge 1 commit into
jordanfelle wants to merge 1 commit into
Conversation
ExtraFileService<T> reads (GetFilesByBook) and deletes (DeleteForBook) its rows with WHERE "AuthorId" = @A AND "BookId" = @b, and DeleteForBook runs once per deleted book for every ExtraFile type via BookDeletedEvent. Neither MetadataFiles nor ExtraFiles had an index on those columns (only the primary key, plus BookFileId on MetadataFiles), so each call was a sequential scan. On a live database (MetadataFiles ~69.5k rows) pg_stat_statements showed that DELETE at ~5 ms/call against ~0.05-0.1 ms for every other statement on the delete path, and 38% of all database execution time over a 20 s window of an author refresh that was pruning books. EXPLAIN ANALYZE on a copy: 4.2 ms Seq Scan -> 0.019 ms Index Scan (index 1.3 MB). Migration 110 adds IX_MetadataFiles_AuthorId_BookId and IX_ExtraFiles_AuthorId_BookId, guarded like 105: skipped when the table/columns are absent or the index already exists. Tests (SQLite): both indexes exist in query column order; the per-book delete is planned as an index search; missing tables/columns and a pre-existing index do not throw. Two of them fail if the index is not created.
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Problem
ExtraFileService<T>reads (GetFilesByBook) and deletes (DeleteForBook) its rows withWHERE "AuthorId" = @a AND "BookId" = @b.DeleteForBookruns once per deleted book, for every extra-file type, viaBookDeletedEvent. NeitherMetadataFilesnorExtraFileshad an index on those columns (only the primary key, plusBookFileIdonMetadataFiles), so each call was a sequential scan.Evidence (live PostgreSQL,
pg_stat_statements)DELETE FROM "MetadataFiles" WHERE "AuthorId" = $1 AND "BookId" = $2: mean ~5 ms over 33k calls, against ~0.05-0.1 ms for every other statement on the delete path (Books,Editions,ExtraFiles, aliases).EXPLAIN ANALYZEon a copy of the table (69.5k rows): 4.2 ms Seq Scan -> 0.019 ms Index Scan. Index size 1.3 MB.Change
Migration 110 adds
IX_MetadataFiles_AuthorId_BookIdandIX_ExtraFiles_AuthorId_BookId. Guarded like migration 105: skipped when the table or columns are absent or the index already exists. No application code changes.Migration number: 110 is simply the next free number in the branch I run (108/109 are taken by other open PRs); renumber to whatever is next on merge.
Tests
ExtraFileAuthorBookIndexMigrationFixture(SQLite): both indexes exist in query column order; the per-book delete is planned as an index search rather than a scan; missing tables/columns and a pre-existing index don't throw. Two of the four fail if the index is not created. FullChaptarr.Core.Test: 3039 passed.Not measured
The 4.2 ms -> 0.019 ms figure is the statement on a copy; I have no end-to-end refresh timing yet.
History(500k rows) andDownloadHistory(275k) also haveAuthorId+BookIdand no index on the pair; I did not touch them because nothing on this path queries them that way.