Skip to content

Collapse duplicate search rows for one symbol page - #255

Merged
malfet merged 1 commit into
pytorch_sphinx_theme2from
search-collapse-duplicate-pages
Sep 2, 2026
Merged

malfet merged 1 commit into
pytorch_sphinx_theme2from
search-collapse-duplicate-pages

Conversation

@malfet

@malfet malfet commented Aug 28, 2026

Copy link
Copy Markdown
Collaborator

Follow-up to #252. That fixed the ranking half of pytorch/pytorch#4954; this is the "finds duplicates" half, which it left untouched.

On the docs preview from pytorch/pytorch#195042, searching linear returned eight rows all rendering as Linear pointing at eight different pages, and nn.linear returned four rows that were the same page reached three different ways.

Cause

performSearch runs a title pass, an object pass and a fulltext pass, each emitting its own row for a page. The built-in dedup keys on [docname, title, anchor, descr, filename] — which differs per pass, so nothing collapses:

 1. "Adam"              #adam              -> generated/torch.optim.Adam   (title pass)
 9. "torch.optim.Adam"  #torch.optim.Adam  -> generated/torch.optim.Adam   (object pass)
12. "Adam"              (no anchor)        -> generated/torch.optim.Adam   (fulltext pass)

Separately, rows are labelled with whichever heading matched — for autodoc symbol pages that's the bare leaf name, making sibling pages indistinguishable.

Fix

searchtools.js applies Scorer.score to every result before it dedupes, and hands score() the live result array. Rewriting the fields the dedup keys on therefore makes Sphinx's own dedup collapse the rows — no patching of Search.query:

var generated = /(?:^|\/)generated\/(.+)$/.exec(String(docname));
if (generated) { result[1] = generated[1]; result[2] = ""; result[3] = null; }

Relabelling to the page's own symbol fixes the display at the same time. Only generated/<dotted.path> pages are canonicalised — prose pages are left alone, since distinct sections of one page are useful deep links rather than duplicates.

Before / after

linear, before:  Linear, Linear, Linear, Linear, linear, Linear, Linear, Linear
linear, after:   torch.nn.Linear
                 torch.ao.nn.qat.Linear
                 torch.ao.nn.quantized.Linear
                 torch.nn.modules.linear.Linear
                 torch.ao.nn.qat.dynamic.Linear

Same-page duplicates in the top 10: nn.linear 4 → 0, conv2d 2 → 0, adam 1 → 0. All eight symbol queries keep their target at rank 1; cuda semantics, autograd mechanics and broadcasting semantics keep their existing top result.

Known limits

  • Ordering dependency. score() must run ahead of the dedup in Search.query. That holds in Sphinx 7.2.6, which docs builds pin. If a future Sphinx reorders it, rows duplicate again exactly as they do today — it degrades to current behaviour rather than breaking. Called out in a comment in the source.
  • Does not merge genuinely separate pages. torch.optim.Adam and torch.optim.adam.Adam_class are two pages documenting one class. No search change can merge those; that's an autodoc/autosummary duplication on the PyTorch side and deserves its own issue.

Alternative rejected

Switching to a document-oriented engine cannot produce same-page duplicates, since it indexes one entry per page. MiniSearch over a page-level corpus does return zero duplicates — but ranks the wanted page 3rd for adam and linear where this returns it 1st. That trades #252's ranking win for the dedup win instead of getting both.


Authored with the assistance of an AI coding agent.

Searching `linear` returned eight rows all rendering as "Linear", pointing
at eight different pages, and `nn.linear` returned four rows that were the
same page reached three different ways. This is the "finds duplicates" half
of pytorch/pytorch#4954, which the ranking fix in #252 did not address.

Two causes, one fix. performSearch runs a title pass, an object pass and a
fulltext pass, each emitting its own row for a page, and the built-in dedup
keys on [docname, title, anchor, descr, filename] -- which differs per pass,
so nothing collapses. Separately the rows are labelled with whichever
heading matched, and for autodoc symbol pages that is the bare leaf name, so
sibling pages are visually indistinguishable.

searchtools.js applies Scorer.score to every result before it dedupes, and
passes score() the live result array. Rewriting the fields the dedup keys on
therefore makes Sphinx's own dedup collapse the rows, with no patching of
Search.query. Relabelling to the page's own symbol fixes the display at the
same time: `torch.ao.nn.qat.Linear` instead of a third row saying "Linear".

Only `generated/<dotted.path>` pages are canonicalised. Prose pages are left
alone, since distinct sections of one page are useful deep links rather than
duplicates.

The one fragility is the ordering dependency: score() must run ahead of the
dedup in Search.query. That holds in Sphinx 7.2.6, which docs builds pin. If
a future Sphinx reorders it, rows duplicate again exactly as they do today,
so it degrades to current behaviour rather than breaking.

An alternative considered and rejected: switching to a document-oriented
engine, which cannot produce same-page duplicates because it indexes one
entry per page. MiniSearch over a page-level corpus does return zero
duplicates, but ranks the wanted page 3rd for `adam` and `linear` where this
returns it 1st, so it would trade the #252 ranking win for the dedup win
rather than getting both.

Test Plan:

Ran the real searchtools.js under node against the PyTorch docs preview
index from pytorch/pytorch#195042 (3472 docs, 8918 objects):

```
node rank2.js
node cmp.js
```

Same-page duplicates in the top 10: `nn.linear` 4 -> 0, `conv2d` 2 -> 0,
`adam` 1 -> 0. Every one of eight symbol queries keeps its target page at
rank 1, and `cuda semantics`, `autograd mechanics` and `broadcasting
semantics` keep their existing top result.

Authored with the assistance of an AI coding agent.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
@netlify

netlify Bot commented Aug 28, 2026 •

Copy link
Copy Markdown

✅ Deploy Preview for pytorchsphinxtheme ready!

Name Link
🔨 Latest commit d0f1679
🔍 Latest deploy log https://app.netlify.com/projects/pytorchsphinxtheme/deploys/6a9182c3755e580008943eee
😎 Deploy Preview https://deploy-preview-255--pytorchsphinxtheme.netlify.app
📱 Preview on mobile
Toggle QR Code...

QR Code

Use your smartphone camera to open QR code link.

To edit notification comments on pull requests, go to your Netlify project configuration.

@meta-cla meta-cla Bot added the cla signed label Aug 28, 2026
@malfet
malfet marked this pull request as ready for review September 2, 2026 13:46
@malfet
malfet merged commit 628e3f3 into pytorch_sphinx_theme2 Sep 2, 2026
7 checks passed
@malfet malfet mentioned this pull request Sep 2, 2026
malfet added a commit that referenced this pull request Sep 2, 2026
Picks up the duplicate search result collapsing from #255, which is the
other half of pytorch/pytorch#4954 from the ranking fix released in 0.4.12.

Test Plan:

```
python3 -c "import pytorch_sphinx_theme2 as t; print(t.__version__)"
grep -rn '0\.4\.13' setup.py pytorch_sphinx_theme2/__init__.py
```

All three version strings report 0.4.13.

Authored with the assistance of an AI coding agent.

Co-authored-by: Claude Opus 5 <noreply@anthropic.com>
pytorchmergebot pushed a commit to pytorch/pytorch that referenced this pull request Sep 9, 2026
Picks up **both** halves of the docs search fix for #4954, in a single bump.

### Ranking — theme [#252](pytorch/pytorch_sphinx_theme#252), released in 0.4.12

Sphinx's stock scorer ranked module stub pages above the API symbols they document: the Python domain gives modules search priority 0 (worth `+15`) against `+5` for classes and functions, and `performSearch` merges three passes onto one list without normalising their scales. Searching `nn.linear` returned twelve `*.modules.linear` module pages ahead of `torch.nn.Linear`, which sat at rank 29.

### Duplicates — theme [#255](pytorch/pytorch_sphinx_theme#255), released in 0.4.13

Those same three passes each emitted a row for one page, and the built-in dedup keyed on `[docname, title, anchor, descr, filename]` — which differs per pass — so a single symbol page was listed repeatedly. Rows were also labelled with whichever heading matched, so `linear` returned eight indistinguishable entries reading `Linear`.

### Note on the Docker rebuild

This touches `.ci/docker/`, so it forces a docs image rebuild. That's unavoidable rather than accidental — `docs/requirements.txt` is a symlink to `.ci/docker/requirements-docs.txt`, and the image is what the docs job actually installs from. Both theme fixes are taken in one bump so the rebuild happens once rather than twice.

### Test plan

CI builds the docs and publishes a preview. On that preview, searching `nn.linear` should return `torch.nn.Linear` first rather than the quantized module pages, and `linear` should return distinct qualified names rather than repeated `Linear` rows.

Verified before landing against the preview index from an earlier revision of this PR (3472 docs, 8918 objects), running the real `searchtools.js` under node.

Rank of the wanted page:

| query | before | after |
|---|---|---|
| `nn.linear` → `torch.nn.Linear` | 29 | **1** |
| `linear` → `torch.nn.Linear` | 3 | **1** |
| `conv2d` → `torch.nn.Conv2d` | 3 | **1** |
| `adam` → `torch.optim.Adam` | 2 | **1** |
| `softmax` → `torch.nn.Softmax` | 2 | **1** |

Same-page duplicates in the top 10: `nn.linear` 4 → **0**, `conv2d` 2 → **0**, `adam` 1 → **0**. Prose queries (`cuda semantics`, `autograd mechanics`, `broadcasting semantics`) keep their existing top result.

### Not fixed here

`torch.optim.Adam` and `torch.optim.adam.Adam_class` are two separate generated pages documenting the same class (likewise `AdamW`, `Adamax`, `NAdam`, `RAdam`). No search change can merge those — it's an autodoc/autosummary duplication that puts genuine duplicates in the index, and it deserves its own issue.

---

Authored with the assistance of an AI coding agent.
Pull Request resolved: #195042
Approved by: https://github.com/atalman

Co-authored-by: Claude Opus 5 <noreply@anthropic.com>
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant