docs(15.9): add the operating facts of the chunk job, embedding settings and engine-side fusion - #566
Merged
Merged
Conversation
…ngs and engine-side fusion Add what the 15.9 semantic search pages left out or got wrong. Each statement was checked against the Fess source. rank-fusion: - With engine-side fusion the total hit count and the facet counts cover the union of the keyword and semantic hits. With Fess-side fusion the count is the main searcher's count plus a correction of at most window_size / 2, and the facets are the main searcher's. - A page whose start is at or beyond rank.fusion.pagination_depth is refused (HTTP 400 invalid_request in the v2 search API). - An unresponsive embedding provider delays an engine-fused search by its timeout times its attempts, since rank.fusion.timeout does not apply. search-semantic: - content_chunker.embedding.dimension: an unset value gives a 768-dimension mapping without any warning (the page said it warned) that cannot be changed afterwards. New section on the failure that follows and how to recover. - content_chunk_status: which documents become skipped, including every document when content_chunker.chunker.name names no chunker; skipped and fail are terminal; how to read the status distribution as the state of the chunk job. - content_chunker.embedding.opensearch.retry.max is the number of attempts, not the number of retries. - content_chunker.search.knn.k is per shard, and with Fess-side fusion it has a floor of window_size divided by the number of searchers. - chunk_size is in characters while the model's limit is in tokens (measured example for paraphrase-multilingual-MiniLM-L12-v2). - Changing to another model of the same dimension is accepted silently and mixes vector spaces; how to switch. api-search: - q is limited by api.param.max.length (default 1000); a longer value returns 400 invalid_request. All seven languages, 15.9 tree only.
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Summary
Add to the 15.9 semantic search pages the operating facts they left out, and correct one statement they got wrong. Each statement was re-read in the Fess source at the current
main(66b1fa64b) before it was written.All seven languages (ja, en, de, es, fr, ko, zh-cn), development tree (15.9) only. The ja and en text is the source; de, es, fr, ko and zh-cn are translated from it, reusing each page's existing terms.
Changes
config/rank-fusion.rstrank.fusion.pagination_depthresults per shard, the semantic searcher at mostcontent_chunker.search.knn.k). With Fess-side fusion the count is the main searcher's count plus the existing correction (at mostwindow_size / 2), and the facets are the main searcher's, so the two modes report different counts for the same query.startisrank.fusion.pagination_depthor more is refused; in the v2 search API this is HTTP 400invalid_request.rank.fusion.timeoutdoes not apply to engine-side fusion, so an unresponsive embedding provider is bounded by its own timeout times its attempts (built-inopensearchprovider:timeoutxretry.max, 60 s x 3 by default, plus the waits between attempts).config/search-semantic.rstcontent_chunker.embedding.dimension: an unset value gives a 768-dimension mapping without a warning (the page said it warned; only an empty, non-numeric, non-positive or over-16000 value warns). The mapping's dimension cannot be changed afterwards. New section If the Index Was Created Without a Dimension: what happens (the runtime read has no fallback, so documents becomefail) and how to recover.content_chunk_status:skippedcovers empty or whitespace-only bodies, documents with no chunk produced, documents overmax_chunks_per_document, and every document whencontent_chunker.chunker.namenames no chunker.skippedandfailare terminal (failis selected again only withcontent_chunker.job.retry_failed=true). How to read the status distribution as the state of the chunk job, and how to clearskippedafter fixing the cause.content_chunker.embedding.opensearch.retry.maxis the maximum number of attempts including the first (3= three requests, two waits), not the number of retries.content_chunker.search.knn.kis per shard; with Fess-side fusion the effective value has a floor ofrank.fusion.window_sizedivided by the number of searchers (100 by default), so a smaller value has no effect; with engine-side fusion it is used as configured.chunk_sizeis in characters while the model's limit is in tokens, and Fess does not truncate the text it embeds; a measured example forparaphrase-multilingual-MiniLM-L12-v2.api/api-search.rst:qis limited byapi.param.max.length(default 1000); a longer value returns 400invalid_request. This is not a search engine limit.config/properties.rstis generated and is not touched.Not documented
Behaviour of how the chunk job handles provider failures and what its summary line reports is deliberately left out of the new text.
Verification
python3 -m unittest discover -s tools -p 'test_*.py': 79 tests OKpython3 tools/update_eol.py --check,python3 tools/gen_properties_doc.py --check,python3 tools/check_headings.pyon the 15.8 and 15.9 trees: passconfig/andapi/directories before and after the change: the set of warnings is identical, and no unrendered inline markup remains in the changed pagesdoctitle_xform=False) of every changed file: no new messages, and the two new sections, the new notes and the new literal blocks are present in each language