Skip to content

perf(db): index MetadataFiles/ExtraFiles by (AuthorId, BookId) - #267

Open
jordanfelle wants to merge 1 commit into
Chaptarr:developfrom
jordanfelle:perf-extrafiles-author-book-index
Open

jordanfelle wants to merge 1 commit into
Chaptarr:developfrom
jordanfelle:perf-extrafiles-author-book-index

Conversation

@jordanfelle

Copy link
Copy Markdown

Problem

ExtraFileService<T> reads (GetFilesByBook) and deletes (DeleteForBook) its rows with WHERE "AuthorId" = @a AND "BookId" = @b. DeleteForBook runs once per deleted book, for every extra-file type, via BookDeletedEvent. Neither MetadataFiles nor ExtraFiles had an index on those columns (only the primary key, plus BookFileId on MetadataFiles), so each call was a sequential scan.

Evidence (live PostgreSQL, pg_stat_statements)

  • DELETE FROM "MetadataFiles" WHERE "AuthorId" = $1 AND "BookId" = $2: mean ~5 ms over 33k calls, against ~0.05-0.1 ms for every other statement on the delete path (Books, Editions, ExtraFiles, aliases).
  • In a 20 s window of an author refresh that was pruning books, it was 38% of all database execution time.
  • EXPLAIN ANALYZE on a copy of the table (69.5k rows): 4.2 ms Seq Scan -> 0.019 ms Index Scan. Index size 1.3 MB.

Change

Migration 110 adds IX_MetadataFiles_AuthorId_BookId and IX_ExtraFiles_AuthorId_BookId. Guarded like migration 105: skipped when the table or columns are absent or the index already exists. No application code changes.

Migration number: 110 is simply the next free number in the branch I run (108/109 are taken by other open PRs); renumber to whatever is next on merge.

Tests

ExtraFileAuthorBookIndexMigrationFixture (SQLite): both indexes exist in query column order; the per-book delete is planned as an index search rather than a scan; missing tables/columns and a pre-existing index don't throw. Two of the four fail if the index is not created. Full Chaptarr.Core.Test: 3039 passed.

Not measured

The 4.2 ms -> 0.019 ms figure is the statement on a copy; I have no end-to-end refresh timing yet. History (500k rows) and DownloadHistory (275k) also have AuthorId+BookId and no index on the pair; I did not touch them because nothing on this path queries them that way.

ExtraFileService<T> reads (GetFilesByBook) and deletes (DeleteForBook) its rows with
WHERE "AuthorId" = @A AND "BookId" = @b, and DeleteForBook runs once per deleted book for
every ExtraFile type via BookDeletedEvent. Neither MetadataFiles nor ExtraFiles had an index
on those columns (only the primary key, plus BookFileId on MetadataFiles), so each call was a
sequential scan.

On a live database (MetadataFiles ~69.5k rows) pg_stat_statements showed that DELETE at
~5 ms/call against ~0.05-0.1 ms for every other statement on the delete path, and 38% of all
database execution time over a 20 s window of an author refresh that was pruning books.
EXPLAIN ANALYZE on a copy: 4.2 ms Seq Scan -> 0.019 ms Index Scan (index 1.3 MB).

Migration 110 adds IX_MetadataFiles_AuthorId_BookId and IX_ExtraFiles_AuthorId_BookId,
guarded like 105: skipped when the table/columns are absent or the index already exists.

Tests (SQLite): both indexes exist in query column order; the per-book delete is planned as
an index search; missing tables/columns and a pre-existing index do not throw. Two of them
fail if the index is not created.
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant