Skip to content

perf(extras): batch the post-scan metadata-file cleanup and probe the right base path first - #273

Open
jordanfelle wants to merge 2 commits into
Chaptarr:developfrom
jordanfelle:perf-clean-metadata-files-batch
Open

jordanfelle wants to merge 2 commits into
Chaptarr:developfrom
jordanfelle:perf-clean-metadata-files-batch

Conversation

@jordanfelle

@jordanfelle jordanfelle commented Sep 29, 2026 •

Copy link
Copy Markdown

Addresses #271.

Problem

CleanExtraFileService.Clean runs after every author scan. For each MetadataFiles row it probed the author's base paths in a fixed audiobook, ebook, legacy order, and it deleted missing rows one at a time. DiskProviderBase.FileExists falls back to listing directories when the exact path misses, so:

  • an ebook sidecar always paid for a miss under the audiobook folder before its hit, and
  • every missing row cost its own DELETE.

Change

  • Probe the right base path first. Sidecars are written under ExtraFilePathHelper.GetPreferredBasePath(author, bookFile), so for rows that point at a book file that path is tried first. Rows with no book file keep the existing order, and every base path is still tried before a row counts as missing, so which rows get deleted is unchanged. Adds an ExtraFilePathHelper.GetAuthorBasePaths(author, preferredBasePath) overload and one IMediaFileService.GetFilesByAuthor query per author (optional constructor dependency, no effect when absent).
  • One batched delete (DeleteMany) for all missing rows instead of one Delete per row.
  • Skip rows with a blank RelativePath instead of throwing from Path.Combine.
  • Return early for an author with no metadata files.

What this does not change

FileExists itself, the case-insensitive fallback, and the "list each folder once" idea from #271. That would change matching semantics and belongs in its own change.

Tests

CleanExtraFileServiceFixture (9 tests): one batched delete with the exact ids; nothing deleted when all files exist; a linked ebook sidecar is found on the first probe and the audiobook folder is never touched; a file that exists only under the other base path is kept; default order for rows with no book file; blank paths ignored; no work for an author with no rows; a row whose linked book file no longer exists falls back to the default order; works without the media file service. The batching and preferred-base tests fail if either change is reverted (checked). Full Chaptarr.Core.Test: 3044 passed.

Not measured

I have not measured the wall-clock effect of this on a real library; the claim is fewer probes and fewer statements, not a speedup figure. The cost of a full-library scan on network storage is what #271 describes.

Note on cost

The added GetFilesByAuthor is one join (BookFile, Edition, Book, Author) per scan, and only when the author has metadata rows. For an author with many book files and very few sidecars it could cost more than the probes it saves; the existing lighter GetMappedFilePathEvidenceByAuthor returns no book-file id, so it cannot be used here.

… right base path first

CleanExtraFileService.Clean ran after every author scan and, for each MetadataFiles row, probed the
author's base paths in a fixed audiobook -> ebook -> legacy order and deleted missing rows one by
one. DiskProviderBase.FileExists falls back to listing directories when the exact path misses, so an
ebook sidecar always paid for a miss under the audiobook folder before its hit, and every missing row
cost its own DELETE.

- Probe the base path the row's book file lives under first. Sidecars are written under
  ExtraFilePathHelper.GetPreferredBasePath(author, bookFile), so that is where they are. Rows with no
  book file keep the existing order, and every base path is still tried before a row counts as missing.
  New ExtraFilePathHelper.GetAuthorBasePaths(author, preferredBasePath) overload; one
  IMediaFileService.GetFilesByAuthor query per author (optional constructor dependency, no effect
  when absent).
- Delete all missing rows with one DeleteMany instead of one Delete per row.
- Skip rows with a blank RelativePath instead of throwing from Path.Combine.
- Return early for an author with no metadata files.

Tests: one batched delete with the exact ids; nothing deleted when all files exist; a linked ebook
sidecar is found on the first probe and the audiobook folder is never touched; a file that exists only
under the other base path is kept; default order for rows without a book file; blank paths ignored;
no work for an author with no rows; works without the media file service. The batching and
preferred-base tests fail if either change is reverted.

Refs Chaptarr#271
… exists

Adversarial review of Chaptarr#273: a BookFileId that is not among the author's book files (deleted, or never mapped) must fall back to the default probe order and still keep a file that exists under another base path.
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant