Skip to content

docs(15.9): add the operating facts of the chunk job, embedding settings and engine-side fusion - #566

Merged
marevol merged 1 commit into
mainfrom
docs/hybrid-search-operations
Oct 4, 2026
Merged

marevol merged 1 commit into
mainfrom
docs/hybrid-search-operations

Conversation

@marevol

@marevol marevol commented Oct 4, 2026

Copy link
Copy Markdown
Contributor

Summary

Add to the 15.9 semantic search pages the operating facts they left out, and correct one statement they got wrong. Each statement was re-read in the Fess source at the current main (66b1fa64b) before it was written.

All seven languages (ja, en, de, es, fr, ko, zh-cn), development tree (15.9) only. The ja and en text is the source; de, es, fr, ko and zh-cn are translated from it, reusing each page's existing terms.

Changes

  • config/rank-fusion.rst
    • Hit counts: with engine-side fusion the total hit count and the facet counts cover the union of the keyword and semantic hits (each searcher contributes at most rank.fusion.pagination_depth results per shard, the semantic searcher at most content_chunker.search.knn.k). With Fess-side fusion the count is the main searcher's count plus the existing correction (at most window_size / 2), and the facets are the main searcher's, so the two modes report different counts for the same query.
    • Paging: a page whose start is rank.fusion.pagination_depth or more is refused; in the v2 search API this is HTTP 400 invalid_request.
    • rank.fusion.timeout does not apply to engine-side fusion, so an unresponsive embedding provider is bounded by its own timeout times its attempts (built-in opensearch provider: timeout x retry.max, 60 s x 3 by default, plus the waits between attempts).
  • config/search-semantic.rst
    • content_chunker.embedding.dimension: an unset value gives a 768-dimension mapping without a warning (the page said it warned; only an empty, non-numeric, non-positive or over-16000 value warns). The mapping's dimension cannot be changed afterwards. New section If the Index Was Created Without a Dimension: what happens (the runtime read has no fallback, so documents become fail) and how to recover.
    • content_chunk_status: skipped covers empty or whitespace-only bodies, documents with no chunk produced, documents over max_chunks_per_document, and every document when content_chunker.chunker.name names no chunker. skipped and fail are terminal (fail is selected again only with content_chunker.job.retry_failed=true). How to read the status distribution as the state of the chunk job, and how to clear skipped after fixing the cause.
    • content_chunker.embedding.opensearch.retry.max is the maximum number of attempts including the first (3 = three requests, two waits), not the number of retries.
    • content_chunker.search.knn.k is per shard; with Fess-side fusion the effective value has a floor of rank.fusion.window_size divided by the number of searchers (100 by default), so a smaller value has no effect; with engine-side fusion it is used as configured.
    • chunk_size is in characters while the model's limit is in tokens, and Fess does not truncate the text it embeds; a measured example for paraphrase-multilingual-MiniLM-L12-v2.
    • New section Changing to Another Model with the Same Dimension: it is accepted silently because only the dimension is checked, old vectors and new query vectors then live in different spaces, and how to switch.
  • api/api-search.rst: q is limited by api.param.max.length (default 1000); a longer value returns 400 invalid_request. This is not a search engine limit.

config/properties.rst is generated and is not touched.

Not documented

Behaviour of how the chunk job handles provider failures and what its summary line reports is deliberately left out of the new text.

Verification

  • python3 -m unittest discover -s tools -p 'test_*.py': 79 tests OK
  • python3 tools/update_eol.py --check, python3 tools/gen_properties_doc.py --check, python3 tools/check_headings.py on the 15.8 and 15.9 trees: pass
  • Sphinx build (default builder, minimal configuration) of each language's 15.9 config/ and api/ directories before and after the change: the set of warnings is identical, and no unrendered inline markup remains in the changed pages
  • docutils parse (doctitle_xform=False) of every changed file: no new messages, and the two new sections, the new notes and the new literal blocks are present in each language

…ngs and engine-side fusion

Add what the 15.9 semantic search pages left out or got wrong. Each statement
was checked against the Fess source.

rank-fusion:
- With engine-side fusion the total hit count and the facet counts cover the
  union of the keyword and semantic hits. With Fess-side fusion the count is
  the main searcher's count plus a correction of at most window_size / 2, and
  the facets are the main searcher's.
- A page whose start is at or beyond rank.fusion.pagination_depth is refused
  (HTTP 400 invalid_request in the v2 search API).
- An unresponsive embedding provider delays an engine-fused search by its
  timeout times its attempts, since rank.fusion.timeout does not apply.

search-semantic:
- content_chunker.embedding.dimension: an unset value gives a 768-dimension
  mapping without any warning (the page said it warned) that cannot be
  changed afterwards. New section on the failure that follows and how to
  recover.
- content_chunk_status: which documents become skipped, including every
  document when content_chunker.chunker.name names no chunker; skipped and
  fail are terminal; how to read the status distribution as the state of the
  chunk job.
- content_chunker.embedding.opensearch.retry.max is the number of attempts,
  not the number of retries.
- content_chunker.search.knn.k is per shard, and with Fess-side fusion it has
  a floor of window_size divided by the number of searchers.
- chunk_size is in characters while the model's limit is in tokens (measured
  example for paraphrase-multilingual-MiniLM-L12-v2).
- Changing to another model of the same dimension is accepted silently and
  mixes vector spaces; how to switch.

api-search:
- q is limited by api.param.max.length (default 1000); a longer value returns
  400 invalid_request.

All seven languages, 15.9 tree only.
@marevol marevol self-assigned this Oct 4, 2026
@marevol
marevol merged commit bed61f1 into main Oct 4, 2026
2 checks passed
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant