diff --git a/de/15.9/admin/backup-guide.rst b/de/15.9/admin/backup-guide.rst index d119c6d4..41b824e4 100644 --- a/de/15.9/admin/backup-guide.rst +++ b/de/15.9/admin/backup-guide.rst @@ -45,6 +45,11 @@ doc.json doc.json enthält die Mapping-Informationen des Fess-Index. +chat_log.ndjson +::::::::::::::: + +chat_log.ndjson enthält die Nutzungsprotokolle des KI-Chats. + click_log.ndjson :::::::::::::::: diff --git a/de/15.9/admin/docreport-guide.rst b/de/15.9/admin/docreport-guide.rst new file mode 100644 index 00000000..90558c02 --- /dev/null +++ b/de/15.9/admin/docreport-guide.rst @@ -0,0 +1,71 @@ +=============== +Dokumentbericht +=============== + +Übersicht +========= + +Die Seite Dokumentbericht hilft beim Aufräumen gecrawlter Dateiserver und Ähnlichem. Sie listet +Dokumente mit gleichem Inhalt und Dokumente auf, die lange nicht geändert wurden; jede Liste lässt +sich als CSV herunterladen. + +Um die Seite zu öffnen, wählen Sie im linken Menü [Systeminformationen > Dokumentbericht]. Zum +Anzeigen ist die Rolle ``admin-docreport`` oder ``admin-docreport-view`` erforderlich. Die Seite +zeigt Berichte nur an und lädt sie herunter; sie ändert keine Dokumente. + +Beide Registerkarten lassen sich mit „URL-Präfix“ eingrenzen, zum Beispiel ``smb://server/share/``. + +Duplikate +========= + +Dokumente mit gleichem oder nahezu gleichem Inhalt werden gruppiert, die größte Gruppe zuerst. Die +Gruppen beruhen auf der beim Indexieren berechneten Inhaltssignatur (``content_minhash_bits``, +dieselbe, die doppelte Suchergebnisse zusammenfasst); eine Neuindexierung ist daher nicht nötig. +Dokumente, deren Inhalt keine Wörter enthält (etwa leere Dateien), werden ausgelassen. + +Die Seite zeigt bis zu ``docreport.duplicate.group.size`` (Standard: 100) Gruppen und je Gruppe bis zu +``docreport.duplicate.docs.size`` (Standard: 10) Dokumente. Mit [CSV herunterladen] erhalten Sie alle +Gruppen. Die CSV-Datei liest auch bei einem großen Index alle Gruppen, mit den Spalten +``group, groupSize, url, title, filename, contentLength, lastModified, owner, lastModifier, clickCount, docId``. + +.. note:: + + Bei einem Index, der die Inhaltssignatur nicht speichert (die Mappings ``cloud`` und ``aws``), ist + der Duplikatbericht nicht verfügbar. + +Inaktive Dokumente +================== + +Dokumente, deren letzte Änderung älter als die angegebene Anzahl von Tagen ist („Nicht geändert seit +(Tagen)“, Standard 365 aus ``docreport.dormant.days``), werden aufgelistet, die ältesten zuerst. +Dokumente ohne Änderungsdatum werden nicht aufgeführt. Mit „Nie aus den Suchergebnissen geöffnet“ +werden Dokumente ausgelassen, die aus Suchergebnissen angeklickt wurden. + +Die Seite zeigt die Anzahl der passenden Dokumente, ihre Gesamtgröße und eine seitenweise Liste. Das +Blättern endet bei ``indexer.max.result.window.size``; Dokumente darüber hinaus erhalten Sie mit +[CSV herunterladen]. + +Einstellungen +============= + +Die folgenden Einstellungen in ``fess_config.properties`` steuern den Bericht. + +.. list-table:: + :header-rows: 1 + :widths: 40 45 15 + + * - Eigenschaft + - Beschreibung + - Standard + * - ``docreport.duplicate.group.size`` + - Höchstzahl der Duplikatgruppen, die die Seite anzeigt + - ``100`` + * - ``docreport.duplicate.docs.size`` + - Höchstzahl der Dokumente, die die Seite je Gruppe auflistet + - ``10`` + * - ``docreport.duplicate.export.page.size`` + - Anzahl der Inhaltssignaturen, die beim CSV-Download je Anfrage gelesen werden + - ``10000`` + * - ``docreport.dormant.days`` + - Standardanzahl der Tage seit der letzten Änderung, ab der ein Dokument als inaktiv gilt + - ``365`` diff --git a/de/15.9/admin/general-guide.rst b/de/15.9/admin/general-guide.rst index a1a93677..868b6584 100644 --- a/de/15.9/admin/general-guide.rst +++ b/de/15.9/admin/general-guide.rst @@ -111,6 +111,13 @@ Letzte Änderung prüfen Aktivieren Sie dies für differenzielles Crawling. +Bei HTTP/HTTPS-URLs wird, wenn der ``Last-Modified``-Wert der HEAD-Antwort keine Entscheidung +erlaubt (kein Änderungsdatum im Index, kein ``Last-Modified`` in der HEAD-Antwort oder ein anderer +Status als 200/404), der GET als bedingter GET mit ``If-None-Match`` (dem indexierten ``ETag``) +und/oder ``If-Modified-Since`` gesendet. Eine mit ``304 Not Modified`` beantwortete Seite wird wie +eine unveränderte Seite behandelt und nicht erneut abgerufen. Um alles erneut abzurufen, etwa nach +einer Änderung der Crawl-Einstellungen, deaktivieren Sie diese Einstellung für einen Crawl. + Gleichzeitige Crawler-Konfiguration ::::::::::::::::::::::::::::::::::::: diff --git a/de/15.9/admin/index.rst b/de/15.9/admin/index.rst index 91419e8f..b69b8f21 100644 --- a/de/15.9/admin/index.rst +++ b/de/15.9/admin/index.rst @@ -50,6 +50,7 @@ Rollen bis zu Protokollen und Sicherungen. log-guide failureurl-guide searchlist-guide + docreport-guide backup-guide maintenance-guide esreq-guide diff --git a/de/15.9/admin/labeltype-guide.rst b/de/15.9/admin/labeltype-guide.rst index 01a2c634..be87afa6 100644 --- a/de/15.9/admin/labeltype-guide.rst +++ b/de/15.9/admin/labeltype-guide.rst @@ -85,11 +85,46 @@ Anzeigereihenfolge Geben Sie die Anzeigereihenfolge der Labels an. +Art +::: + +Geben Sie „Label“ oder „Tag“ an. Ein gewöhnliches Label ist „Label“. „Tag“ ist ein Tag, den Benutzer +auf der Suchseite hinzufügen (siehe „Tags“ unten). Ein bestehendes Label ohne Art wird als „Label“ +behandelt. + + Konfiguration löschen --------------------- Klicken Sie auf den Konfigurationsnamen auf der Übersichtsseite und dann auf die Schaltfläche „Löschen". Es wird ein Bestätigungsbildschirm angezeigt. Klicken Sie auf die Schaltfläche „Löschen", um die Konfiguration zu löschen. +Tags +---- + +Mit ``user.tag.enabled=true`` (Standard: ``false``) in ``fess_config.properties`` können +angemeldete Benutzer Suchergebnisse taggen. Im mitgelieferten Theme ``bootstrap`` werden Tags an den +Ergebnissen angezeigt, Benutzer können Tags hinzufügen und eigene entfernen, und eine Facette „Tags“ +grenzt die Ergebnisse ein. Zur API siehe :doc:`../api/api-tag`. + +Ein Tag wird als Label der Art „Tag“ gespeichert: Der Name ist der Tag-Name, der Wert der SHA-256 +des Namens, die eingeschlossenen Pfade sind die getaggten URLs (eine je Zeile, exakte +Übereinstimmung), und die Berechtigungen bestimmen, wer das Tag sehen kann. Ein Benutzer, der ein Tag +hinzufügt, wird zu dessen Berechtigungen hinzugefügt. + +- Ein Tag ist nur sichtbar, wenn die Berechtigungen seines Labels auf den Aufrufer zutreffen. + Administratoren können ein Tag auf dieser Seite bearbeiten, um es mit einer Rolle oder Gruppe zu + teilen, oder es löschen. +- Tags mit gleichem Namen werden zu einem Label zusammengeführt; Benutzer, die ein Tag gleichen + Namens hinzugefügt haben, sehen daher gegenseitig, wo ihre Tags gesetzt sind. +- Tags erscheinen weder in der Label-Listen-API (``/api/v2/labels``) noch in der Label-Auswahl der + Suchseite. +- Tags zählen zum Label-Limit (``page.labeltype.max.fetch.size``, Standard: 1000). Ist das Limit + erreicht, kann kein neues Tag erstellt werden. +- Ändert oder löscht ein Administrator ein Tag auf dieser Seite, behalten die indexierten Dokumente die + alten Werte, bis sie erneut gecrawlt werden oder der Job „Label Updater“ läuft. +- Ein Dokument kann bis zu ``user.tag.max.document.tags`` (Standard: 100) Tags haben, ein Tag-Name bis + zu ``user.tag.name.max.length`` (Standard: 50) Zeichen. + .. |image0| image:: ../../../resources/images/en/15.9/admin/labeltype-1.png .. |image1| image:: ../../../resources/images/en/15.9/admin/labeltype-2.png diff --git a/de/15.9/admin/mapping-guide.rst b/de/15.9/admin/mapping-guide.rst index b6efe5cf..69d59bc3 100644 --- a/de/15.9/admin/mapping-guide.rst +++ b/de/15.9/admin/mapping-guide.rst @@ -49,6 +49,33 @@ Upload Sie können im Mapping-Wörterbuchformat hochladen. +Mitgelieferte Mapping-Wörterbücher +================================== + +Das Standard-``mapping.txt`` wird bei der Analyse von Suchfeldern wie ``title`` und ``content`` +verwendet. Es vereinheitlicht Hiragana, kleine Kana und Halbbreiten-Katakana zu Vollbreiten-Katakana, +sodass りんご, リンゴ und リンゴ einander finden. In |Fess| 15.9 werden zusätzlich folgende +Schreibweisen vereinheitlicht: + +- ゐ und ゑ (zu イ und エ), kleine Kana wie ゎ, ゕ, ゖ, ヮ, ヵ, ヶ und ㇰ-ㇿ sowie ゝ und ゞ (zu ヽ und ヾ) +- ヴ, ヴャ, ヴュ und ヴョ, Hiragana ゔ, Halbbreiten-ヴ sowie ウ oder う mit einem kombinierenden + Stimmhaftigkeitszeichen (U+3099); zum Beispiel wird ラヴ zu ラブ und レヴュー zu レビユー + +Außerdem werden ‐ ‑ ‒ – — ― ⁻ ₋ − und -, die direkt nach Kana stehen, als Längungszeichen ー +behandelt (``prolonged_sound_mark_filter``), sodass サ―バ- und サ−バ‐ サーバー finden. Ein +ASCII-Bindestrich (``-``) und ein Strich nach Kanji, Buchstaben oder Ziffern (東京-大阪, +2026−10−02) werden nicht verändert. + +``ja/mapping.txt`` für Japanisch (die ``*_ja``-Felder) lässt Hiragana und kleine Kana unverändert, +weil die morphologische Analyse sie benötigt, und vereinheitlicht nur Schreibweisen wie ヴ. + +.. note:: + + Diese Einstellungen gelten für einen neu erstellten Dokumentindex. Ein bestehender Index behält + seine Analyse-Einstellungen und Wörterbücher, bis er neu indexiert wird. Um sie auf einen + bestehenden Index anzuwenden, indexieren Sie auf der Seite :doc:`maintenance-guide` mit + aktiviertem „Wörterbücher zurücksetzen“ neu. Das Zurücksetzen überschreibt Änderungen an + ``mapping.txt`` / ``ja/mapping.txt``, die in der Administrationsoberfläche vorgenommen wurden. .. |image0| image:: ../../../resources/images/en/15.9/admin/mapping-1.png .. |image1| image:: ../../../resources/images/en/15.9/admin/mapping-2.png diff --git a/de/15.9/admin/relatedquery-guide.rst b/de/15.9/admin/relatedquery-guide.rst index 79397f2f..a854779c 100644 --- a/de/15.9/admin/relatedquery-guide.rst +++ b/de/15.9/admin/relatedquery-guide.rst @@ -54,6 +54,76 @@ Konfiguration löschen Klicken Sie auf den Konfigurationsnamen auf der Übersichtsseite und dann auf die Schaltfläche „Löschen". Es wird ein Bestätigungsbildschirm angezeigt. Klicken Sie auf die Schaltfläche „Löschen", um die Konfiguration zu löschen. +Aus Suchprotokollen generieren +------------------------------ + +Klicken Sie auf der Listenseite auf [Aus Suchprotokollen generieren], um verwandte Abfragen aus den +letzten Suchprotokollen zu erstellen. Eine Suche, die dieselbe Benutzersitzung kurz nach einer +anderen Suche ausführt (ein Tippfehler gefolgt von seiner Korrektur oder ein allgemeiner Begriff +gefolgt von einem genaueren), gilt als Verfeinerung. Für häufig gesuchte Begriffe werden die +häufigsten Verfeinerungen zu den verwandten Abfragen des Begriffs. + +Verwandte Abfragen gelten für alle Benutzer und erweitern jede Suche nach ihrem Begriff; die +Generierung ist daher zurückhaltend: + +- Es werden nur Suchen verwendet, die ein Gast sehen kann. Ein Suchprotokoll wird nur gelesen, wenn + alle seine Rollen ``suggest.search.log.permissions`` erfüllen (dieselbe Einstellung wie bei Suggest). +- Suchbegriffe mit einem Feldfilter wie ``label:"x"``, Operatoren, Platzhaltern, ``sort:`` oder + einem führenden ``+`` / ``-`` werden nicht verwendet. +- Wörter, die unter [Vorschlagen > Schlechtes Wort] registriert sind, werden weder als Begriff noch als + verwandte Abfrage verwendet. +- Ein Begriff und jede seiner verwandten Abfragen müssen aus mindestens + ``related_query.generate.min.sessions`` Sitzungen stammen, und die Verfeinerungen müssen Treffer + haben. +- Einträge werden für jeden virtuellen Host getrennt erzeugt. Suchprotokolle ohne virtuellen Host + gelten als Standardhost. +- Begriffe, die bereits verwandte Abfragen haben, werden nicht verändert (das Ergebnis nennt die + Anzahl der übersprungenen), und es werden nicht mehr Einträge erzeugt, als der Cache der + verwandten Abfragen laden kann (``page.relatedquery.max.fetch.size``). + +Erzeugte verwandte Abfragen lassen sich wie manuell registrierte bearbeiten oder löschen. Sie +können nicht erzeugt werden, solange „Suchprotokoll“ oder „Benutzerprotokoll“ unter +[System > Allgemein] deaktiviert ist, und ein zweiter Lauf kann nicht starten, während einer läuft. + +Die folgenden Einstellungen in ``fess_config.properties`` steuern die Generierung. + +.. list-table:: + :header-rows: 1 + :widths: 45 40 15 + + * - Eigenschaft + - Beschreibung + - Standard + * - ``related_query.generate.days`` + - Anzahl der Tage an Suchprotokollen, die gelesen werden + - ``30`` + * - ``related_query.generate.term.size`` + - Höchstzahl an Begriffen je virtuellem Host + - ``100`` + * - ``related_query.generate.query.size`` + - Höchstzahl an verwandten Abfragen je Begriff + - ``5`` + * - ``related_query.generate.min.sessions`` + - Mindestzahl an Sitzungen, in denen ein Begriff und seine verwandte Abfrage vorkommen müssen + - ``3`` + * - ``related_query.generate.session.interval`` + - Zeitraum, in dem eine Suche als Verfeinerung gilt (Minuten) + - ``10`` + * - ``related_query.generate.seed.log.size`` + - Höchstzahl der je Begriff gelesenen Suchprotokolle + - ``1000`` + * - ``related_query.generate.seed.session.size`` + - Höchstzahl der je Begriff gelesenen Sitzungen + - ``200`` + * - ``related_query.generate.log.fetch.size`` + - Höchstzahl der je Begriff gelesenen Folgesuchen + - ``2000`` + * - ``related_query.generate.query.min.length`` + - Mindestlänge eines Begriffs und einer verwandten Abfrage (Zeichen) + - ``2`` + * - ``related_query.generate.query.max.length`` + - Höchstlänge eines Begriffs und einer verwandten Abfrage (Zeichen) + - ``50`` .. |image0| image:: ../../../resources/images/en/15.9/admin/relatedquery-1.png .. |image1| image:: ../../../resources/images/en/15.9/admin/relatedquery-2.png diff --git a/de/15.9/admin/searchlog-guide.rst b/de/15.9/admin/searchlog-guide.rst index 5dd1f826..ed4bc867 100644 --- a/de/15.9/admin/searchlog-guide.rst +++ b/de/15.9/admin/searchlog-guide.rst @@ -1,27 +1,105 @@ -=========== +============= Suchprotokoll -=========== +============= Übersicht ========= -Such-, Klick- und Favoriten-Ausführungsergebnisse werden aufgezeichnet, und Suchprotokolle können auf diesem Verwaltungsbildschirm überprüft werden. +Suchen, Klicks und Favoriten werden aufgezeichnet. Die Seite Suchprotokoll zeigt Analyseberichte, die +sie zusammenfassen, sowie eine Liste der einzelnen Protokolle. -Verwaltung -========== +Um die Seite zu öffnen, wählen Sie im linken Menü [Systeminformationen > Suchprotokoll]. Zuerst wird +die Registerkarte „Übersicht“ angezeigt. Zum Anzeigen ist die Rolle ``admin-searchlog`` oder +``admin-searchlog-view`` erforderlich; mit ``admin-searchlog-view`` lassen sich keine Protokolle +löschen. -Übersicht -========= +Analyseberichte +=============== + +Zeitraum und Filter +------------------- + +Wählen Sie oben auf jeder Registerkarte den Zeitraum aus „Heute“, „Gestern“, „Letzte 7 Tage“, +„Letzte 28 Tage“ und „Letzte 90 Tage“, oder geben Sie mit „Benutzerdefiniert“ ein Start- und +Enddatum an (höchstens 366 Tage). Mit „Mit vorherigem Zeitraum vergleichen“ werden die Werte mit dem +gleich langen Zeitraum davor verglichen. Außerdem lassen sich die Zugriffsart und die Zeilenzahl der +Tabellen wählen. Zeiträume und Diagrammintervalle folgen den Kalendertagen der Zeitzone des Servers. + +Registerkarten +-------------- + +- **Übersicht**: Suchanfragen, Benutzer, Quote ohne Treffer, Klickrate und durchschnittliche + Antwortzeit, jeweils mit einem kleinen Verlaufsdiagramm und der Veränderung gegenüber dem + vorherigen Zeitraum; ein Verlaufsdiagramm mit umschaltbarer Kennzahl (beim Vergleich wird der + vorherige Zeitraum gestrichelt dargestellt); sowie die häufigsten Suchbegriffe und die Suchbegriffe + ohne Treffer. +- **Suchbegriffe**: je Suchbegriff die Suchanfragen, Benutzer, durchschnittlichen Treffer, Klicks, + Klickrate und durchschnittliche Klickposition. Außerdem die Suchbegriffe ohne Treffer (mit dem + Zeitpunkt der letzten Suche) und die Suchbegriffe ohne Klicks, die Treffer hatten, deren Ergebnisse + aber nie angeklickt wurden. +- **Klicks**: die am häufigsten angeklickten URLs, die am häufigsten favorisierten URLs, die + Verteilung der Klickpositionen und der Anteil der Aufrufe ab Seite 2. +- **Leistung**: durchschnittliche Antwortzeit, Median (p50), p95 und p99, die Verteilung der + Antwortzeiten, die langsamsten Suchbegriffe und die Abfragezeit. +- **Zielgruppe**: neue und wiederkehrende Benutzer, Zugriffsarten, Suchanfragen nach Wochentag und + Stunde sowie die häufigsten User-Agents, Referrer, Sprachen und virtuellen Hosts. „Suchanfragen + nach Rolle und Gruppe“ zeigt für jede Rolle und Gruppe die Suchanfragen, Benutzer und die Quote ohne + Treffer. Eine Suche zählt für jede Rolle und Gruppe des ausführenden Benutzers, daher kann die Summe + der Zeilen über der Gesamtzahl liegen. Einzelne Benutzer werden nicht aufgeführt. +- **KI-Chat**: Anfragen, Benutzer, Tokens insgesamt, durchschnittliche Antwortzeit und Fehlerquote des + KI-Suchmodus (RAG-Chat) sowie die häufigsten Benutzer und die Anfragen nach Absicht und nach + Modell. Die Chat-Nutzung wird aufgezeichnet, solange ``rag.chat.log.enabled`` (Standard: ``true``) + aktiviert ist. Fragen und Antworten werden nicht aufgezeichnet. Token-Zahlen werden nur + aufgezeichnet, wenn das LLM-Plugin sie meldet. +- **Protokolle**: die Liste der einzelnen Protokolle; siehe „Protokollliste“ unten. + +.. note:: + + Klick-Kennzahlen je Suchbegriff und die Suchbegriffe ohne Klicks zählen nur Klicks, die + aufgezeichnet wurden, seit Suchbegriffe mit den Klicks gespeichert werden (ab |Fess| 15.9). Die + gesamten Klicks und die Klickrate enthalten auch ältere Klicks. Einige Werte, etwa die Zahl der + Benutzer, sind Näherungswerte. -In der Übersicht können Sie Such-, Klick- und Favoriten-Suchprotokolle überprüfen. -Um Details des Suchprotokolls anzuzeigen, klicken Sie auf das entsprechende Suchprotokoll. +Suchbegriffe ohne Treffer untersuchen +------------------------------------- + +Klicken Sie auf den Registerkarten „Übersicht“ und „Suchbegriffe“ auf einen Suchbegriff ohne Treffer, +um die Registerkarte „Protokolle“ mit den Suchprotokollen dieses Begriffs zu öffnen, gefiltert auf +„Nur ohne Treffer“. So sehen Sie, welche Suchen nichts gefunden haben, und können Dokumente, +Synonyme oder verwandte Abfragen ergänzen. + +CSV herunterladen +----------------- + +Jede Tabelle und jedes Diagramm der Analyseberichte hat einen CSV-Link, der die Werte für den +aktuellen Zeitraum, Vergleich, die Zugriffsart und die Größe herunterlädt. Die Filterleiste hat +außerdem einen Link auf eine CSV-Datei der Kennzahlen. Zahlen werden unverändert ausgegeben +(Anteile als 0 bis 1, Zeiten in Millisekunden). Ein verglichenes Diagramm erhält eine zusätzliche +Spalte ``_previous``. + +Protokollliste +============== + +Die Registerkarte „Protokolle“ listet Suchprotokolle, Klickprotokolle, Favoritenprotokolle und +Benutzerprotokolle auf. Sie lassen sich nach Protokollart, Abfrage-ID, Benutzer-ID, Zeitraum, +Zugriffsart und Suchbegriff filtern, Suchprotokolle zusätzlich nach Trefferzahl („Alle“, „Nur ohne +Treffer“, „Mindestens ein Treffer“). Um die Details eines Protokolls anzuzeigen, klicken Sie darauf. |image0| +Klicken Sie auf [CSV herunterladen], um die Protokolle, die dem aktuellen Filter entsprechen, ohne +Zeilenbegrenzung und neueste zuerst als CSV herunterzuladen. Die Kopfzeile enthält die Feldnamen und +hängt daher nicht von der Sprache der Oberfläche ab. + +CSV-Dateien, auch die der Analyseberichte, werden in der Kodierung von ``csv.file.encoding`` +geschrieben; eine UTF-8-Datei beginnt mit einer Byte-Order-Mark. Ein Wert, der mit ``=``, ``+``, +``-``, ``@``, einem Tabulator oder einem Wagenrücklauf beginnt, erhält ein vorangestelltes ``'``, +damit eine Tabellenkalkulation ihn nicht als Formel ausführt. + Details -======= +------- -Durch Klicken auf ein Suchprotokoll in der Übersicht werden die Details des entsprechenden Suchprotokolls angezeigt. +Klicken Sie in der Liste auf ein Protokoll, um seine Details anzuzeigen. |image1| diff --git a/de/15.9/api/admin/api-admin-backup.rst b/de/15.9/api/admin/api-admin-backup.rst index b7a807a2..8a284b46 100644 --- a/de/15.9/api/admin/api-admin-backup.rst +++ b/de/15.9/api/admin/api-admin-backup.rst @@ -79,12 +79,13 @@ Im Folgenden ein Beispiel bei den Standardeinstellungen (``index.backup.targets` { "id": "system.properties", "name": "system.properties" }, { "id": "fess.json", "name": "fess.json" }, { "id": "doc.json", "name": "doc.json" }, + { "id": "chat_log.ndjson", "name": "chat_log.ndjson" }, { "id": "click_log.ndjson", "name": "click_log.ndjson" }, { "id": "favorite_log.ndjson", "name": "favorite_log.ndjson" }, { "id": "search_log.ndjson", "name": "search_log.ndjson" }, { "id": "user_info.ndjson", "name": "user_info.ndjson" } ], - "total": 10 + "total": 11 } } @@ -112,7 +113,7 @@ Je nach Art von ``{id}`` wechselt der Antwortinhalt wie folgt. - Die Mapping-Definitionsdatei des Index selbst (``fess_indices/fess.json`` / ``fess_indices/fess/doc.json``), unverändert (``application/octet-stream``) * - ``*.bulk`` oder Indexname ohne Erweiterung - Durch Scrollen des gleichnamigen Index erzeugte Bulk-Daten (``application/octet-stream``). Der Name ohne ``.bulk`` wird als Indexname behandelt. - * - ``*.ndjson`` (``search_log`` / ``user_info`` / ``click_log`` / ``favorite_log``) + * - ``*.ndjson`` (``search_log`` / ``user_info`` / ``click_log`` / ``favorite_log`` / ``chat_log``) - NDJSON-Daten des entsprechenden Protokolls (``application/x-ndjson``) .. note:: diff --git a/de/15.9/api/api-export.rst b/de/15.9/api/api-export.rst new file mode 100644 index 00000000..ef72afee --- /dev/null +++ b/de/15.9/api/api-export.rst @@ -0,0 +1,98 @@ +======================= +Suchergebnis-Export-API +======================= + +Dieses Dokument beschreibt die v2-Export-API von |Fess|, mit der Suchergebnisse als CSV- oder +JSON-Datei heruntergeladen werden. Informationen zum gemeinsamen Antwort-Envelope und zum +Fehlermodell finden Sie unter :doc:`api-overview`. + +Die Basis-URL lautet ``http:///api/v2/`` (Beispiel für eine lokale Umgebung: ``http://localhost:8080/api/v2``). + +.. note:: + + Der Export ist standardmäßig deaktiviert. Um ihn zu nutzen, setzen Sie ``api.search.export=true`` + in ``fess_config.properties``. Ist er aktiviert, zeigt das mitgelieferte Theme ``bootstrap`` neben + der Trefferzahl ein Exportmenü (CSV / JSON). ``features.search_export`` von ``/api/v2/ui/config`` + meldet den Zustand. + +Suchergebnisse herunterladen +============================ + +Anfrage +------- + +================== ==================================================== +HTTP-Methode GET +Endpunkt ``/api/v2/documents/export`` +================== ==================================================== + +Gibt die zur Suche passenden Dokumente als Datei-Download zurück (``Content-Disposition: +attachment``, Dateiname ``search_results.csv`` oder ``search_results.json``). + +- Es gilt derselbe Rollenfilter wie bei ``/api/v2/search``. Bei ``login.required=true`` kann wie bei + ``/api/v2/search`` ein Zugriffstoken verwendet werden. +- Höchstens ``api.search.export.max.size`` (Standard: ``1000``) Dokumente werden exportiert. Die + Blätterparameter (``start``, ``num``) werden nicht verwendet. +- Exportiert werden die Felder aus ``api.search.export.fields`` (Standard: + ``title,url_link,last_modified,content_length,filetype``), die auch in API-Antworten erscheinen + dürfen. +- Anfragen sind auf ``api.search.export.rate.limit.per.minute`` (Standard: ``10``; ``0`` bedeutet + unbegrenzt) pro Minute begrenzt, gezählt je angemeldetem Benutzer bzw. je Client-IP bei Gästen. + Darüber antwortet der Endpunkt mit ``429`` und einem ``Retry-After``-Header. +- Ein Export wird nicht im Suchprotokoll aufgezeichnet. + +Anfrageparameter +---------------- + +Es können dieselben Suchbedingungen wie bei ``/api/v2/documents/all`` angegeben werden, etwa ``q``, +``ex_q``, ``fields.*``, ``sort`` und ``lang`` (siehe :doc:`api-search`). Zusätzlich: + +.. tabularcolumns:: |p{4cm}|p{11cm}| +.. list-table:: Anfrageparameter + + * - ``format`` + - Dateiformat: ``csv`` (Standard) oder ``json``. Jeder andere Wert ergibt ``invalid_request`` (400). + +Tabelle: Anfrageparameter + +Antwort +------- + +Die CSV-Datei hat eine Kopfzeile mit den Feldnamen und wird in der Kodierung von +``csv.file.encoding`` geschrieben (eine UTF-8-Datei beginnt mit einer Byte-Order-Mark). Ein Wert, der +mit ``=``, ``+``, ``-``, ``@``, einem Tabulator oder einem Wagenrücklauf beginnt, erhält ein +vorangestelltes ``'``, damit eine Tabellenkalkulation ihn nicht als Formel ausführt. Mehrwertige +Felder werden mit einem Leerzeichen verbunden. + +:: + + "title","url_link","last_modified","content_length","filetype" + "Example","https://example.com/","2025-01-01T00:00:00.000Z","1234","html" + +Die JSON-Datei hat die Form ``{"data":[{...},...]}``; mehrwertige Felder bleiben Arrays. + +Ein Fehler vor Beginn der Datei liefert den üblichen Fehler-Envelope. Ein Fehler danach lässt sich in +der Datei nicht melden; der Download endet vorzeitig mit einer abgeschnittenen CSV-Datei oder einem +nicht parsebaren JSON. + +Fehlerantwort +------------- + +.. tabularcolumns:: |p{4cm}|p{11cm}| +.. list-table:: Fehlerantwort + + * - Statuscode + - Beschreibung + * - 400 Bad Request + - Eine fehlerhafte Abfrage, ein anderes ``format`` als ``csv`` / ``json`` oder ein mit + ``api.search.export=false`` deaktivierter Export. + * - 401 Unauthorized + - Wenn eine Authentifizierung erforderlich ist (z. B. anonymer Aufrufer bei ``login.required=true``). + * - 405 Method Not Allowed + - Wenn die HTTP-Methode nicht zulässig ist. + * - 429 Too Many Requests + - Wenn das Anfragelimit pro Minute überschritten ist. + * - 500 Internal Server Error + - Wenn ein interner Serverfehler auftritt. + +Tabelle: Fehlerantwort diff --git a/de/15.9/api/api-search-history.rst b/de/15.9/api/api-search-history.rst new file mode 100644 index 00000000..d4bcb6d0 --- /dev/null +++ b/de/15.9/api/api-search-history.rst @@ -0,0 +1,117 @@ +=============== +Suchverlauf-API +=============== + +Dieses Dokument beschreibt die v2-Suchverlauf-API von |Fess|. +Informationen zum gemeinsamen Antwort-Envelope und zum Fehlermodell finden Sie unter :doc:`api-overview`. + +Die Basis-URL lautet ``http:///api/v2/`` (Beispiel für eine lokale Umgebung: ``http://localhost:8080/api/v2``). + +.. note:: + + Der Suchverlauf ist verfügbar, solange sowohl ``search.history.enabled`` (Standard: ``true``) als + auch das Suchprotokoll aktiviert sind. ``features.search_history`` von ``/api/v2/ui/config`` meldet + den Zustand. + +Letzte Suchen abrufen +===================== + +Anfrage +------- + +================== ==================================================== +HTTP-Methode GET +Endpunkt ``/api/v2/search-history`` +================== ==================================================== + +Gibt die letzten Suchen zurück, die der angemeldete Benutzer mit ``/api/v2/search`` auf dem aktuellen +virtuellen Host ausgeführt hat, die neueste zuerst. Ein Client kann eine davon mit den +zurückgegebenen Bedingungen erneut ausführen. + +- Nur Suchen auf der ersten Ergebnisseite werden aufgeführt. Suchen mit gleichen Bedingungen werden + zur neuesten zusammengefasst, Suchen ohne Suchbegriff ausgelassen. +- Höchstens ``search.history.size`` (Standard: ``10``) Suchen werden zurückgegeben. +- Suchprotokolle werden von einem minütlich laufenden Job geschrieben; eine Suche kann daher bis zu + etwa einer Minute brauchen, bis sie erscheint. +- Der Verlauf gehört zum angemeldeten Benutzer der Sitzung. Anonyme Aufrufer erhalten + ``auth_required`` (401); ein Zugriffstoken ersetzt keine Anmeldung. +- Ist der Suchverlauf deaktiviert, antwortet der Endpunkt mit ``invalid_request`` (400). + +Es gibt keine Anfrageparameter. + +Antwort +------- + +Bei Erfolg (200) werden die folgenden Felder direkt unter ``response`` des gemeinsamen Envelopes zurückgegeben. + +:: + + { + "response": { + "status": 0, + "record_count": 1, + "data": [ + { + "q": "fess", + "fields": { "label": ["docs"] }, + "sort": "last_modified.desc", + "requested_at": "2026-10-01T09:00:00Z", + "hit_count": 42 + } + ] + } + } + +.. tabularcolumns:: |p{4cm}|p{11cm}| +.. list-table:: Antwortfelder + + * - ``record_count`` + - Anzahl der Suchen in ``data`` (int). + * - ``data`` + - Die letzten Suchen, die neueste zuerst. Die Schlüssel der Bedingungen sind die + Anfrageparameternamen von ``/api/v2/search``; ein Schlüssel fehlt, wenn die Suche ihn nicht + verwendet hat. + * - ``data[].q`` + - Der Suchbegriff (str). + * - ``data[].fields`` + - Mit ``fields.`` angegebene Feldbedingungen, nach Feldnamen, jeweils mit ihren Werten. + * - ``data[].ex_q`` + - Zusätzliche Abfragen (Array von str). + * - ``data[].sort`` + - Sortierreihenfolge (str). + * - ``data[].lang`` + - Mit ``lang`` angeforderte Sprachen (Array von str). + * - ``data[].requested_at`` + - Zeitpunkt der Suche (UTC, ISO-8601). + * - ``data[].hit_count`` + - Trefferzahl der Suche (int64). + +Tabelle: Antwortfelder + +Fehlerantwort +------------- + +.. tabularcolumns:: |p{4cm}|p{11cm}| +.. list-table:: Fehlerantwort + + * - Statuscode + - Beschreibung + * - 400 Bad Request + - Wenn der Suchverlauf deaktiviert ist. + * - 401 Unauthorized + - Wenn der Aufrufer nicht angemeldet ist. + * - 405 Method Not Allowed + - Wenn die HTTP-Methode nicht zulässig ist. + * - 500 Internal Server Error + - Wenn ein interner Serverfehler auftritt. + +Tabelle: Fehlerantwort + +Im mitgelieferten Theme +======================= + +Im mitgelieferten Theme ``bootstrap`` sieht ein angemeldeter Benutzer die letzten Suchen in der +Vorschlagsliste, wenn er in das leere Suchfeld klickt oder darin die Pfeil-nach-unten-Taste drückt. +Ein ausgewählter Eintrag führt dieselbe Suche erneut aus, einschließlich Bedingungen wie Labels. +Suchprotokolle, die vor |Fess| 15.9 aufgezeichnet wurden, enthalten keine Bedingungen und werden nicht +aufgeführt. diff --git a/de/15.9/api/api-tag.rst b/de/15.9/api/api-tag.rst new file mode 100644 index 00000000..403a38f9 --- /dev/null +++ b/de/15.9/api/api-tag.rst @@ -0,0 +1,151 @@ +======== +Tags-API +======== + +Dieses Dokument beschreibt die v2-Tags-API von |Fess|, mit der Benutzer Dokumente taggen. +Informationen zum gemeinsamen Antwort-Envelope, zum Fehlermodell und zu CSRF finden Sie unter :doc:`api-overview`. + +Die Basis-URL lautet ``http:///api/v2/`` (Beispiel für eine lokale Umgebung: ``http://localhost:8080/api/v2``). + +.. note:: + + Tags sind standardmäßig deaktiviert. Um sie zu nutzen, setzen Sie ``user.tag.enabled=true`` in + ``fess_config.properties``. ``features.user_tag`` von ``/api/v2/ui/config`` meldet den Zustand. + +Ein Tag ist ein Label der Art „Tag“ (siehe :doc:`../admin/labeltype-guide`): Der Labelname ist der +Tag-Name, der Wert der SHA-256 des Namens in Hexadezimalform, die eingeschlossenen Pfade sind die +getaggten URLs, und die Berechtigungen bestimmen, wer das Tag sehen kann. Ein Tag ist nur sichtbar, +wenn sein Label für den Aufrufer sichtbar ist. + +Die Such-API (``/api/v2/search``) gibt zu jedem Treffer die für den Aufrufer sichtbaren Tags als +``tags`` zurück. ``fields.tag=`` grenzt die Ergebnisse auf Dokumente mit einem Tag ein, und +``facet.field=tag`` liefert eine Tag-Facette. Das Indexfeld ``tag`` selbst wird nicht zurückgegeben. + +Tags abrufen +============ + +Anfrage +------- + +================== ==================================================== +HTTP-Methode GET +Endpunkt ``/api/v2/documents/{docId}/tags`` +================== ==================================================== + +Gibt die für den Aufrufer sichtbaren Tags des Dokuments zurück. Kann der Aufrufer das Dokument nicht +durchsuchen, antwortet der Endpunkt mit ``not_found`` (404). + +Antwort +------- + +Bei Erfolg (200) werden die folgenden Felder direkt unter ``response`` des gemeinsamen Envelopes zurückgegeben. + +:: + + { + "response": { + "status": 0, + "doc_id": "a1b2c3d4e5f6", + "addable": true, + "tags": [ + { "value": "9f86d081884c7d65...", "name": "zu-pruefen", "mine": true } + ] + } + } + +.. tabularcolumns:: |p{4cm}|p{11cm}| +.. list-table:: Antwortfelder + + * - ``doc_id`` + - Dokument-ID (str). + * - ``addable`` + - ``true``, wenn der Aufrufer angemeldet ist und Tags hinzufügen kann (bool). + * - ``added`` + - Nur POST. ``false``, wenn der Aufrufer das Dokument bereits getaggt hatte (bool). + * - ``removed`` + - Nur DELETE (bool). + * - ``tags`` + - Die für den Aufrufer sichtbaren Tags. Jedes hat ``value`` (den Labelwert für ``fields.tag``), + ``name`` (den Tag-Namen) und ``mine`` (``true``, wenn der Aufrufer zu den Berechtigungen des + Tags gehört). + +Tabelle: Antwortfelder + +Ein Tag hinzufügen +================== + +Anfrage +------- + +================== ==================================================== +HTTP-Methode POST +Endpunkt ``/api/v2/documents/{docId}/tags`` +================== ==================================================== + +Taggt die URL des Dokuments für den angemeldeten Benutzer; ein Zugriffstoken ersetzt keine Anmeldung. +Als zustandsändernde Anfrage erfordert sie den Header ``X-Fess-CSRF-Token``. + +- Gibt es bereits ein Tag dieses Namens, wird die URL zu seinen eingeschlossenen Pfaden und der + Benutzer zu seinen Berechtigungen hinzugefügt. Andernfalls wird ein Tag erstellt, das nur der + Benutzer sehen kann. Tags gleichen Namens werden daher zu einem zusammengeführt, und Benutzer, die + ein Tag gleichen Namens hinzugefügt haben, sehen gegenseitig, wo ihre Tags gesetzt sind. +- Erneutes Taggen desselben Dokuments ergibt ``added: false``. +- Ein Dokument kann höchstens ``user.tag.max.document.tags`` (Standard: ``100``) Tags haben. + +Senden Sie ``Content-Type: application/json`` mit dem Tag-Namen in ``name``. + +:: + + { + "name": "zu-pruefen" + } + +Der Name wird NFKC-normalisiert, Leerzeichenfolgen werden zusammengefasst und er wird getrimmt. Er +muss 1 bis ``user.tag.name.max.length`` (Standard: ``50``) Zeichen lang sein; Namen mit einem +Steuer- oder Formatzeichen (etwa einem Nullbreitenzeichen oder einer Bidi-Überschreibung) werden +abgelehnt. + +Ein Tag entfernen +================= + +Anfrage +------- + +================== ==================================================== +HTTP-Methode DELETE +Endpunkt ``/api/v2/documents/{docId}/tags?value=`` +================== ==================================================== + +Entfernt den angemeldeten Benutzer aus den Berechtigungen des mit ``value`` angegebenen Tags. Bleibt +keine Berechtigung eines Benutzers, einer Gruppe oder einer Rolle übrig, wird das Tag gelöscht und aus +den Dokumenten entfernt. Gehört der Benutzer nicht zu den Berechtigungen des Tags, antwortet der +Endpunkt mit ``forbidden`` (403). Der Header ``X-Fess-CSRF-Token`` ist erforderlich. + +Fehlerantwort +============= + +.. tabularcolumns:: |p{4cm}|p{11cm}| +.. list-table:: Fehlerantwort + + * - Statuscode + - Beschreibung + * - 400 Bad Request + - Wenn die Anfrage ungültig ist (auch wenn Tags deaktiviert sind, der Tag-Name ungültig ist oder + ein Tag-Limit überschritten wird). + * - 401 Unauthorized + - POST oder DELETE ohne Anmeldung. + * - 403 Forbidden + - Ein fehlendes oder abgelaufenes CSRF-Token oder ein DELETE eines Tags, das der Benutzer nicht + hinzugefügt hat. + * - 404 Not Found + - Wenn das Dokument nicht gefunden wird oder der Aufrufer es nicht durchsuchen kann. + * - 405 Method Not Allowed + - Wenn die HTTP-Methode nicht zulässig ist. + * - 413 Payload Too Large + - Wenn der Anfrage-Body die Größenbegrenzung überschreitet. + * - 415 Unsupported Media Type + - Wenn der ``Content-Type`` nicht unterstützt wird. + * - 500 Internal Server Error + - Wenn ein interner Serverfehler auftritt. + +Tabelle: Fehlerantwort diff --git a/de/15.9/api/api-uiconfig.rst b/de/15.9/api/api-uiconfig.rst index bd127b85..9eb2067d 100644 --- a/de/15.9/api/api-uiconfig.rst +++ b/de/15.9/api/api-uiconfig.rst @@ -44,6 +44,9 @@ Bei Erfolg (HTTP 200, UiConfigResponse) wird eine Antwort im gemeinsamen Envelop }, "features": { "user_favorite": false, + "search_history": true, + "search_export": false, + "user_tag": false, "popular_word": true, "suggest_search_log": true, "suggest_documents": true, @@ -186,6 +189,15 @@ Alle Felder sind Pflichtfelder. * - ``user_favorite`` - boolean - Gibt an, ob die Benutzerfavoriten-Funktion aktiviert ist. + * - ``search_history`` + - boolean + - Ob der Suchverlauf (``GET /api/v2/search-history``) verfügbar ist (``true``, wenn ``search.history.enabled`` und das Suchprotokoll aktiviert sind). + * - ``search_export`` + - boolean + - Ob der Export von Suchergebnissen (``GET /api/v2/documents/export``) aktiviert ist (``api.search.export``). + * - ``user_tag`` + - boolean + - Ob Tags (``/api/v2/documents/{docId}/tags``) aktiviert sind (``user.tag.enabled``). * - ``popular_word`` - boolean - Gibt an, ob die Beliebte-Wörter-Funktion aktiviert ist. diff --git a/de/15.9/api/index.rst b/de/15.9/api/index.rst index ee9d6636..a5bb4fee 100644 --- a/de/15.9/api/index.rst +++ b/de/15.9/api/index.rst @@ -23,6 +23,7 @@ Authentifizierung. :caption: Such-API api-search + api-export api-label api-popularword api-suggest @@ -41,6 +42,8 @@ Authentifizierung. :caption: Benutzerfunktionen-API api-favorite + api-search-history + api-tag api-click api-cache diff --git a/de/15.9/config/crawler-ocr.rst b/de/15.9/config/crawler-ocr.rst index 68c920c1..32d446e6 100644 --- a/de/15.9/config/crawler-ocr.rst +++ b/de/15.9/config/crawler-ocr.rst @@ -167,6 +167,7 @@ Betriebshinweise - Bei neuen Web-Crawl-Konfigurationen schließt das standardmäßige Ausschlussmuster Bild-URLs (jpg, png, gif usw.) aus. Um Bilder auf einer Website zu crawlen, entfernen Sie diese aus „Vom Crawlen ausgeschlossene URL“. Beim Dateisystem-Crawl werden Bilder berücksichtigt. - Auch die Größenbeschränkungen des Crawlers gelten. Die Größenbeschränkung für die Indexierung je Dateityp (Standard: 10 MB) finden Sie unter :doc:`crawler-basic`. - Die OCR-Genauigkeit hängt von der Qualität des Scans ab. Handschrift wird im Allgemeinen nicht gut erkannt. +- Dateien, die vor dem Aktivieren von OCR indexiert wurden, werden nicht automatisch per OCR verarbeitet. Bei aktiviertem inkrementellem Crawling ("Letzte Änderung prüfen" in :doc:`../admin/general-guide`) wird eine Datei, deren Änderungszeit sich nicht geändert hat, beim erneuten Crawlen nicht noch einmal abgerufen. Um OCR auf diese Dateien anzuwenden, deaktivieren Sie "Letzte Änderung prüfen" für einen Crawl, oder löschen Sie die Dokumente aus dem Index und crawlen Sie erneut. Hinweise zum Upgrade ==================== diff --git a/de/15.9/config/rate-limiting.rst b/de/15.9/config/rate-limiting.rst index 0f10c69a..b9290f42 100644 --- a/de/15.9/config/rate-limiting.rst +++ b/de/15.9/config/rate-limiting.rst @@ -102,6 +102,39 @@ Bei Setzen auf ``true`` wird die robots.txt-Verarbeitung einschließlich Crawl-d # robots.txt ignorieren (Standard: false) crawler.ignore.robots.txt=false +Crawl-delay gilt je Origin und ist auf 60 Sekunden begrenzt. Es taktet URLs, nicht Anfragen; bei +einem inkrementellen Crawl werden HEAD und GET einer URL daher direkt nacheinander gesendet. Das +Intervall einer Crawl-Konfiguration wirkt weiterhin getrennt davon als Wartezeit vor der nächsten URL. + +robots.txt wird gemäß RFC 9309 ausgewertet. Zwischen ``Allow:`` und ``Disallow:`` gewinnt die längste +passende Regel, und auch die Start-URLs werden gegen robots.txt geprüft. Eine von robots.txt +verbotene URL wird in ``fess-crawler.log`` auf INFO protokolliert und nicht als Fehler-URL erfasst. + +Backoff nach 429/503-Antworten +------------------------------ + +Antwortet ein Server mit ``429 Too Many Requests`` oder ``503 Service Unavailable``, pausiert |Fess| +die Anfragen an diesen Origin und wiederholt die URL bis zu dreimal. Gewartet wird gemäß dem +``Retry-After``-Header, falls vorhanden, sonst exponentiell ab 10 Sekunden (höchstens 5 Minuten). +Schlägt auch der letzte Versuch fehl, wird ein WARN in ``fess-crawler.log`` geschrieben. + +Wenn robots.txt nicht abgerufen werden kann +------------------------------------------- + +Schlägt der Abruf von robots.txt mit 5xx, 429 oder einer Zeitüberschreitung fehl, werden die URLs +dieses Origins bis zum Ende des Backoffs wieder in die Warteschlange gestellt. Scheitern nach dem +ersten Fehler auch drei weitere Versuche, wird für den Rest des Crawls keine URL dieses Origins +gecrawlt, und ein einzelnes WARN wird in ``fess-crawler.log`` geschrieben. Bis 15.8 bedeutete eine +nicht abrufbare robots.txt „alles erlaubt“. Um dieses Verhalten wiederherzustellen, geben Sie in den +„Konfigurationsparametern“ der Web-Crawl-Konfiguration Folgendes an: + +:: + + client.robotsTxtAllowOnUnavailable=true + +Um die robots.txt-Verarbeitung ganz abzuschalten, verwenden Sie ``client.robotsTxtEnabled=false`` +(je Crawl-Konfiguration) oder ``crawler.ignore.robots.txt=true``. + Alle Rate-Limiting-Einstellungen ================================= diff --git a/de/15.9/dev/theme-development.rst b/de/15.9/dev/theme-development.rst index a5462b9a..0ee6af2d 100644 --- a/de/15.9/dev/theme-development.rst +++ b/de/15.9/dev/theme-development.rst @@ -171,6 +171,17 @@ Auslieferung und API von |Fess| selbst zulässt (Inline-Styles sind erlaubt, Inline-Skripte nicht). Schriften oder Skripte von einem externen CDN werden daher nicht geladen; liefern Sie sie im Theme mit. +- Die ``Content-Security-Policy`` des Einstiegs-HTML enthält + ``frame-ancestors 'none'``, und die Antwort trägt zusätzlich + ``X-Frame-Options: DENY``; die Seite wird daher nicht im Frame einer + anderen Seite angezeigt. Der Wert von ``frame-ancestors`` lässt sich mit + ``theme.index.frame.ancestors`` in ``fess_config.properties`` ändern + (Standard: ``'none'``). Ein leerer Wert lässt ``frame-ancestors`` weg + (``X-Frame-Options: DENY`` wird weiterhin gesendet). Unter + ``frame-ancestors 'none'`` lassen WebKit-basierte Browser wie Safari die + Frames leer, die ein Theme aus einer ``blob:``-URL anzeigt (PDF-Vorschau + oder zwischengespeicherte Kopie). Setzen Sie den Wert leer, um sie + anzuzeigen. - Die SPA des Themes ruft Daten wie Suchergebnisse und Chat über die ``/api/v2/*`` API ab. diff --git a/de/15.9/install/fess-setup.rst b/de/15.9/install/fess-setup.rst index 9eea44e2..5f75e37c 100644 --- a/de/15.9/install/fess-setup.rst +++ b/de/15.9/install/fess-setup.rst @@ -139,6 +139,20 @@ Damit werden die Versionsliste, die JARs und ihre Prüfsummen aus diesem einen M bezogen, etwa einem internen Mirror, statt aus den standardmäßigen Release- und Snapshot-Repositories und von GitHub. +``--repository`` akzeptiert auch eine URL, die mit ``file:///`` beginnt. Legen Sie auf einem Server +ohne Internetzugang eine Kopie des Maven-Repositorys ab und geben Sie dieses Verzeichnis an. Anzugeben +ist das Verzeichnis, das die Verzeichnisse der einzelnen Plugins enthält (``/maven-metadata.xml`` +usw.), also die Stelle, die ``https://maven.codelibs.org/release/org/codelibs/fess/`` des +Standard-Repositorys entspricht. Versionen werden wie bei einem HTTP-Repository aufgelöst und +Prüfsummen ebenso geprüft. + +:: + + $ bin/fess-setup install plugin fess-ds-csv --repository file:///opt/maven-repo/org/codelibs/fess/ + +Ein anderes Schema als ``http``, ``https`` und ``file`` oder eine URL ohne Schema wird mit einer +einzeiligen Fehlermeldung abgelehnt. + install plugin -------------- @@ -159,6 +173,14 @@ Snapshot-Builds seiner eigenen Versionslinie und bevorzugt diese. Beispiele finden Sie unter :doc:`../admin/plugin-guide`. +Installiert werden können nur die Plugin-Typen, die |Fess| lädt: Namen, die mit ``fess-ds-``, +``fess-ingest-``, ``fess-script-``, ``fess-webapp-``, ``fess-thumbnail-``, ``fess-crawler-``, +``fess-llm-``, ``fess-storage-`` oder ``fess-sso-`` beginnen, also dieselben, die ``list plugins`` +anzeigt. Jeder andere Name (``fess-theme-*`` oder eine Bibliothek wie ``fess`` oder ``fess-crawler``) +wird vor jedem Download mit einer einzeiligen Fehlermeldung und Exit-Code 2 abgelehnt. Wird einer von +mehreren Namen abgelehnt, wird keiner installiert. Ein statisches Theme installieren Sie mit +``install theme``. + list plugins ------------ @@ -222,6 +244,10 @@ Bringen Sie in einer solchen Umgebung die JARs der Plugins selbst mit: Plugin-Installationsseite in der Administrationsoberfläche hochladen; siehe :doc:`../admin/plugin-guide`. 3. Starten Sie |Fess|. Wenn Sie die JARs während des Betriebs abgelegt haben, starten Sie es neu. +Statt jedes JAR einzeln mitzubringen, können Sie auch den benötigten Teil des Maven-Repositorys (die +Verzeichnisse ``org/codelibs/fess//``) auf den Server kopieren und mit ``install plugin`` und +``--repository file:///...`` installieren. Auch dabei werden die Prüfsummen geprüft. + Themes verwalten ================ diff --git a/de/15.9/user/search-field.rst b/de/15.9/user/search-field.rst index d488243f..fc6c5110 100644 --- a/de/15.9/user/search-field.rst +++ b/de/15.9/user/search-field.rst @@ -68,6 +68,12 @@ Standardmäßig können die folgenden Felder für die Suche angegeben werden. * - favorite_count - Anzahl der Favorisierungen des Dokuments - Numerisch + * - owner + - Kontoname des Dateibesitzers + - Keyword + * - last_modifier + - Letzter Bearbeiter der Datei + - Keyword Tabelle: Liste der verfügbaren Felder @@ -82,6 +88,16 @@ Wenn kein Feld angegeben wird, erfolgt die Suche über title und content. Je nac .. note:: Je nach Crawling-Ziel werden manche Felder nicht mit einem Wert belegt. So wird anchor beispielsweise nur beim Web-Crawling registriert, und lang nur, wenn das HTML ein Sprachattribut enthält. Außerdem lassen sich Felder wie segment (eine Sitzungs-ID, die den jeweiligen Crawling-Lauf kennzeichnet) oder doc_id (eine vom System vergebene interne ID) angeben, diese werden jedoch bei der normalen Suche in der Regel nicht verwendet. +owner und last_modifier werden beim Crawlen von Dateiservern und Ähnlichem registriert. owner +enthält den Kontonamen des Dateibesitzers aus SMB-, Dateisystem- und FTP-Crawls (ein Wert wie +``DOMAIN\alice`` wird zu ``alice``). last_modifier enthält den letzten Bearbeiter, der aus einem +Office-Dokument o. Ä. extrahiert wird, ersatzweise den Besitzer. Für HTML aus einem Web-Crawl wird +kein owner registriert. Suchen Sie zum Beispiel mit ``owner:alice`` oder +``last_modifier:"Taro Yamada"``. Im mitgelieferten Theme lassen sich Besitzer und letzter Bearbeiter +in der erweiterten Suche angeben. Administratoren können die Felder mit +``crawler.document.file.owner.enabled`` und ``crawler.document.file.last.modifier.enabled`` +(beide standardmäßig ``true``) ein- und ausschalten. + Wenn HTML-Dateien als Suchziel verwendet werden, wird das title-Tag im title-Feld und die Zeichenkette unter dem body-Tag im content-Feld registriert. Verwendung diff --git a/en/15.9/admin/backup-guide.rst b/en/15.9/admin/backup-guide.rst index 8bd0f103..63a11dad 100644 --- a/en/15.9/admin/backup-guide.rst +++ b/en/15.9/admin/backup-guide.rst @@ -48,6 +48,11 @@ doc.json doc.json contains mapping information for the fess index. +chat_log.ndjson +::::::::::::::: + +chat_log.ndjson includes AI chat usage log information. + click_log.ndjson :::::::::::::::: diff --git a/en/15.9/admin/docreport-guide.rst b/en/15.9/admin/docreport-guide.rst new file mode 100644 index 00000000..6d17c774 --- /dev/null +++ b/en/15.9/admin/docreport-guide.rst @@ -0,0 +1,70 @@ +=============== +Document Report +=============== + +Overview +======== + +The Document Report page helps you clean up crawled file servers and the like. It lists documents +with the same content and documents that have not been modified for a long time, and each list can +be downloaded as CSV. + +To open the page, select [System Info > Document Report] in the left menu. Viewing requires the +``admin-docreport`` or ``admin-docreport-view`` role. The page only shows and downloads reports; it +does not change any document. + +Both tabs can be narrowed with "URL prefix", for example ``smb://server/share/``. + +Duplicates +========== + +Documents whose content is the same or nearly the same are grouped, largest group first. The groups +use the content signature computed at index time (``content_minhash_bits``, the same one that +collapses duplicate search results), so no reindex is needed. Documents whose content has no words +(such as empty files) are left out. + +The screen shows up to ``docreport.duplicate.group.size`` (default: 100) groups and up to +``docreport.duplicate.docs.size`` (default: 10) documents per group. Use [Download CSV] to get every +group. The CSV reads all groups, even on a large index, with the columns +``group, groupSize, url, title, filename, contentLength, lastModified, owner, lastModifier, clickCount, docId``. + +.. note:: + + On an index that does not keep the content signature (the ``cloud`` and ``aws`` mappings), the + duplicate report is not available. + +Dormant Documents +================= + +Documents whose last modification is older than the given number of days ("Not modified for +(days)", default 365 from ``docreport.dormant.days``) are listed, oldest first. Documents without a +last modification date are not listed. With "Never opened from search results", documents that have +been clicked from search results are left out. + +The screen shows the number of matching documents, their total size and a paged list. Paging stops +at ``indexer.max.result.window.size``; use [Download CSV] to get the documents beyond it. + +Settings +======== + +The following settings in ``fess_config.properties`` adjust the report. + +.. list-table:: + :header-rows: 1 + :widths: 40 45 15 + + * - Property + - Description + - Default + * - ``docreport.duplicate.group.size`` + - Maximum number of duplicate groups the screen shows + - ``100`` + * - ``docreport.duplicate.docs.size`` + - Maximum number of documents the screen lists for each group + - ``10`` + * - ``docreport.duplicate.export.page.size`` + - Number of content signatures read per request when downloading the CSV + - ``10000`` + * - ``docreport.dormant.days`` + - Default number of days since the last modification after which a document is dormant + - ``365`` diff --git a/en/15.9/admin/general-guide.rst b/en/15.9/admin/general-guide.rst index 84f46d4a..bd0baf48 100644 --- a/en/15.9/admin/general-guide.rst +++ b/en/15.9/admin/general-guide.rst @@ -111,6 +111,13 @@ Check Last Modified Enable to perform differential crawling. +For HTTP/HTTPS URLs, when the ``Last-Modified`` of the HEAD response cannot decide (no last +modified date is indexed, the HEAD response has no ``Last-Modified``, or the HEAD returns a status +other than 200/404), the GET is sent as a conditional GET with ``If-None-Match`` (the indexed +``ETag``) and/or ``If-Modified-Since``. A page answered with ``304 Not Modified`` is handled like an +unchanged page and is not fetched again. To fetch everything again, for example after changing crawl +settings, turn this setting off for a crawl. + Concurrent Crawler Config ::::::::::::::::::::::::: diff --git a/en/15.9/admin/index.rst b/en/15.9/admin/index.rst index 699fef20..2e98d5d7 100644 --- a/en/15.9/admin/index.rst +++ b/en/15.9/admin/index.rst @@ -50,6 +50,7 @@ logs, and backups. log-guide failureurl-guide searchlist-guide + docreport-guide backup-guide maintenance-guide esreq-guide diff --git a/en/15.9/admin/labeltype-guide.rst b/en/15.9/admin/labeltype-guide.rst index 02a7074a..611b473f 100644 --- a/en/15.9/admin/labeltype-guide.rst +++ b/en/15.9/admin/labeltype-guide.rst @@ -84,11 +84,43 @@ Display Order Specifies the display order of labels. +Kind +:::: + +Specify "Label" or "Tag". An ordinary label is "Label". "Tag" is a tag that users add from the +search screen (see "Tags" below). An existing label without a kind is treated as "Label". + + Deleting Configuration ---------------------- Click the configuration name on the list page, then click the Delete button to display a confirmation screen. Click the Delete button to remove the configuration. +Tags +---- + +With ``user.tag.enabled=true`` (default: ``false``) in ``fess_config.properties``, logged-in users +can tag search results. In the bundled ``bootstrap`` theme, tags are shown on the results, users can +add tags and remove their own, and a "Tags" facet narrows the results. For the API, see +:doc:`../api/api-tag`. + +A tag is stored as a label of the kind "Tag": the name is the tag name, the value is the SHA-256 of +the name, the included paths are the tagged URLs (one per line, exact match), and the permissions +decide who can see the tag. A user who adds a tag is added to its permissions. + +- A tag is visible only when the permissions of its label match the caller. Administrators can edit + a tag on this page to share it with a role or a group, or delete it. +- Tags with the same name are merged into one label, so users who added a tag of the same name can + see where each other's tags are. +- Tags are not included in the label list API (``/api/v2/labels``) or in the label choices of the + search screen. +- Tags count toward the label limit (``page.labeltype.max.fetch.size``, default: 1000). Once the + limit is reached, no new tag can be created. +- After an administrator changes or deletes a tag on this page, the indexed documents keep the old + values until they are crawled again or the "Label Updater" job runs. +- A document can have up to ``user.tag.max.document.tags`` (default: 100) tags, and a tag name can be + up to ``user.tag.name.max.length`` (default: 50) characters long. + .. |image0| image:: ../../../resources/images/en/15.9/admin/labeltype-1.png .. |image1| image:: ../../../resources/images/en/15.9/admin/labeltype-2.png \ No newline at end of file diff --git a/en/15.9/admin/mapping-guide.rst b/en/15.9/admin/mapping-guide.rst index e375c24d..6a9fb803 100644 --- a/en/15.9/admin/mapping-guide.rst +++ b/en/15.9/admin/mapping-guide.rst @@ -49,5 +49,31 @@ Upload You can upload in mapping dictionary format. +Bundled Mapping Dictionaries +============================ + +The default ``mapping.txt`` is used to analyze search fields such as ``title`` and ``content``. It +folds hiragana, small kana and half-width katakana into full-width katakana, so that りんご, リンゴ +and リンゴ match each other. In |Fess| 15.9, it also folds the following spellings: + +- ゐ and ゑ (to イ and エ), small kana such as ゎ, ゕ, ゖ, ヮ, ヵ, ヶ and ㇰ-ㇿ, and ゝ and ゞ (to ヽ and ヾ) +- ヴ, ヴャ, ヴュ and ヴョ, hiragana ゔ, half-width ヴ, and ウ or う followed by a combining voiced + sound mark (U+3099); for example, ラヴ becomes ラブ and レヴュー becomes レビユー + +In addition, ‐ ‑ ‒ – — ― ⁻ ₋ − and - written right after kana are treated as the long vowel mark +ー (``prolonged_sound_mark_filter``), so サ―バ- and サ−バ‐ match サーバー. An ASCII hyphen-minus +(``-``) and a dash after kanji, letters or digits (東京-大阪, 2026−10−02) are not changed. + +``ja/mapping.txt`` for Japanese (the ``*_ja`` fields) keeps hiragana and small kana as they are, +because the morphological analyzer needs them, and folds only spellings such as ヴ. + +.. note:: + + These settings apply to a newly created document index. An existing index keeps its analysis + settings and dictionaries until it is reindexed. To apply them to an existing index, reindex + with "Reset Dictionaries" enabled on the :doc:`maintenance-guide` page. Resetting the + dictionaries overwrites any edits to ``mapping.txt`` / ``ja/mapping.txt`` made in the + administration screen. + .. |image0| image:: ../../../resources/images/en/15.9/admin/mapping-1.png .. |image1| image:: ../../../resources/images/en/15.9/admin/mapping-2.png diff --git a/en/15.9/admin/relatedquery-guide.rst b/en/15.9/admin/relatedquery-guide.rst index ed2f2022..91d7b3bd 100644 --- a/en/15.9/admin/relatedquery-guide.rst +++ b/en/15.9/admin/relatedquery-guide.rst @@ -53,5 +53,74 @@ Deleting Configuration Click the configuration name on the list page, then click the Delete button to display a confirmation screen. Click the Delete button to remove the configuration. +Generating from Search Logs +--------------------------- + +Click the [Generate from Search Logs] button on the list page to create related queries from recent +search logs. A search that the same user session makes shortly after another search (a typo +followed by its correction, or a broad term followed by a more specific one) counts as a +refinement. For frequently searched terms, the most common refinements become the related queries +of the term. + +Related queries apply to everyone and broaden every search for their term, so the generation is +conservative: + +- Only searches a guest can see are used. A search log is read only when all of its roles pass + ``suggest.search.log.permissions`` (the same setting as suggest). +- Search terms that contain a field filter such as ``label:"x"``, operators, wildcards, ``sort:`` + or a leading ``+`` / ``-`` are not used. +- Words registered under [Suggest > Bad Word] are used neither as terms nor as related queries. +- A term and each of its related queries must come from at least + ``related_query.generate.min.sessions`` sessions, and the refinements must have hits. +- Entries are generated separately for each virtual host. Search logs without a virtual host are + treated as the default host. +- Terms that already have related queries are not changed (the result shows how many were + skipped), and no more entries are created than the related query cache can load + (``page.relatedquery.max.fetch.size``). + +Generated related queries can be edited or deleted like those registered by hand. They cannot be +generated while "Search Log" or "User Log" is disabled under [System > General], and a second run +cannot start while one is in progress. + +The following settings in ``fess_config.properties`` adjust the generation. + +.. list-table:: + :header-rows: 1 + :widths: 45 40 15 + + * - Property + - Description + - Default + * - ``related_query.generate.days`` + - Days of search logs to read + - ``30`` + * - ``related_query.generate.term.size`` + - Maximum number of terms per virtual host + - ``100`` + * - ``related_query.generate.query.size`` + - Maximum number of related queries per term + - ``5`` + * - ``related_query.generate.min.sessions`` + - Minimum number of sessions in which a term and its related query must appear + - ``3`` + * - ``related_query.generate.session.interval`` + - Interval within which a search counts as a refinement (minutes) + - ``10`` + * - ``related_query.generate.seed.log.size`` + - Maximum number of search logs read per term + - ``1000`` + * - ``related_query.generate.seed.session.size`` + - Maximum number of sessions read per term + - ``200`` + * - ``related_query.generate.log.fetch.size`` + - Maximum number of follow-up search logs read per term + - ``2000`` + * - ``related_query.generate.query.min.length`` + - Minimum length of a term and a related query (characters) + - ``2`` + * - ``related_query.generate.query.max.length`` + - Maximum length of a term and a related query (characters) + - ``50`` + .. |image0| image:: ../../../resources/images/en/15.9/admin/relatedquery-1.png .. |image1| image:: ../../../resources/images/en/15.9/admin/relatedquery-2.png \ No newline at end of file diff --git a/en/15.9/admin/searchlog-guide.rst b/en/15.9/admin/searchlog-guide.rst index 1ae3ac99..e1b366bf 100644 --- a/en/15.9/admin/searchlog-guide.rst +++ b/en/15.9/admin/searchlog-guide.rst @@ -5,24 +5,95 @@ Search Log Overview ======== -Search Log page manages logs for search, click and liked. +Searches, clicks and favorites are recorded. The Search Log page shows analytics reports that +aggregate them, and a list of the individual logs. -Management Operations -===================== +To open the page, select [System Info > Search Log] in the left menu. The "Overview" tab is shown +first. Viewing requires the ``admin-searchlog`` or ``admin-searchlog-view`` role; +``admin-searchlog-view`` cannot delete logs. -Display Search Log +Analytics Reports +================= + +Period and Filters ------------------ -Select System Info > Search Log in the left menu to display a list page of Search Log, as below. +At the top of each tab, choose the period to aggregate from "Today", "Yesterday", "Last 7 days", +"Last 28 days" and "Last 90 days", or specify a start and end date with "Custom" (up to 366 days). +With "Compare to previous period", the values are compared with the period of the same length just +before. You can also choose the access type and the number of table rows. Periods and chart buckets +follow the calendar days of the server's time zone. + +Tabs +---- + +- **Overview**: the searches, users, zero-hit rate, click-through rate and average response time, + each with a small trend chart and its change from the previous period; a trend chart with a + metric switcher (the previous period is drawn as dashed lines when compared); and the top queries + and zero-hit queries. +- **Queries**: per query, the searches, users, average hits, clicks, click-through rate and average + click position. It also shows the zero-hit queries (with the time they were last searched) and the + zero-click queries, which had hits but whose results were never clicked. +- **Clicks**: the most clicked URLs, the most favorited URLs, the click position distribution and the + rate of viewing page 2 and later. +- **Performance**: the average, median (p50), p95 and p99 response time, the response time + distribution, the slowest queries and the query time. +- **Audience**: new and returning users, access types, searches by weekday and hour, and the top + user agents, referers, languages and virtual hosts. "Searches by Role and Group" shows the + searches, users and zero-hit rate of each role and group. A search counts toward every role and + group of the user who ran it, so the rows can add up to more than the total. Individual users are + not listed. +- **AI Chat**: the requests, users, total tokens, average response time and error rate of the AI + search mode (RAG chat), with the top users and the requests by intent and by model. Chat usage is + recorded while ``rag.chat.log.enabled`` (default: ``true``) is on. Questions and answers are not + recorded. Token counts are recorded only when the LLM plugin reports them. +- **Logs**: the list of individual logs; see "Log List" below. + +.. note:: + + Per-query click metrics and the zero-click queries count only clicks recorded since search words + started to be recorded with clicks (|Fess| 15.9 and later). The overall clicks and click-through + rate include older clicks. Some values, such as the number of users, are approximate. + +Looking into Zero-Hit Queries +----------------------------- + +On the "Overview" and "Queries" tabs, click a zero-hit query to open the "Logs" tab with the search +logs of that query, filtered to "Zero hits only". You can see which searches found nothing and use +that to add documents, synonyms or related queries. + +Downloading CSV +--------------- + +Every table and chart of the analytics reports has a CSV link that downloads what it aggregates for +the current period, comparison, access type and size. The filter bar also has a link to a CSV of the +metrics. Numbers are written as they are (ratios as 0 to 1, times in milliseconds). A compared chart +gets an extra ``_previous`` column. + +Log List +======== + +The "Logs" tab lists search logs, click logs, favorite logs and user logs. You can filter them by +log type, query ID, user ID, time range, access type and search word, and search logs also by hit +count ("All", "Zero hits only", "One or more hits"). To see the details of a log, click it. |image0| +Click [Download CSV] to download the logs that match the current filter as CSV, newest first, with +no row limit. The header row holds the field names, so it does not depend on the UI language. + +CSV files, including those of the analytics reports, are written in the encoding of +``csv.file.encoding``; a UTF-8 file starts with a byte order mark. A value that starts with ``=``, +``+``, ``-``, ``@``, a tab or a carriage return gets a leading ``'``, so that a spreadsheet does not +run it as a formula. + Details ------- -To display details for a specific search log, click the row of search log data. +Click a log in the list to show its details. |image1| + .. |image0| image:: ../../../resources/images/en/15.9/admin/searchlog-1.png .. |image1| image:: ../../../resources/images/en/15.9/admin/searchlog-2.png diff --git a/en/15.9/api/admin/api-admin-backup.rst b/en/15.9/api/admin/api-admin-backup.rst index d37f9553..2d576ced 100644 --- a/en/15.9/api/admin/api-admin-backup.rst +++ b/en/15.9/api/admin/api-admin-backup.rst @@ -79,12 +79,13 @@ The following is an example under the default settings (when ``index.backup.targ { "id": "system.properties", "name": "system.properties" }, { "id": "fess.json", "name": "fess.json" }, { "id": "doc.json", "name": "doc.json" }, + { "id": "chat_log.ndjson", "name": "chat_log.ndjson" }, { "id": "click_log.ndjson", "name": "click_log.ndjson" }, { "id": "favorite_log.ndjson", "name": "favorite_log.ndjson" }, { "id": "search_log.ndjson", "name": "search_log.ndjson" }, { "id": "user_info.ndjson", "name": "user_info.ndjson" } ], - "total": 10 + "total": 11 } } @@ -112,7 +113,7 @@ Depending on the type of ``{id}``, the response content switches as follows. - The index mapping definition file itself (``fess_indices/fess.json`` / ``fess_indices/fess/doc.json``), returned as-is (``application/octet-stream``) * - ``*.bulk`` or an index name without an extension - Bulk data generated by scrolling the index with the same name as the target (``application/octet-stream``). The name with ``.bulk`` removed is treated as the index name. - * - ``*.ndjson`` (``search_log`` / ``user_info`` / ``click_log`` / ``favorite_log``) + * - ``*.ndjson`` (``search_log`` / ``user_info`` / ``click_log`` / ``favorite_log`` / ``chat_log``) - NDJSON data of the corresponding log (``application/x-ndjson``) .. note:: diff --git a/en/15.9/api/api-export.rst b/en/15.9/api/api-export.rst new file mode 100644 index 00000000..2ae82ce4 --- /dev/null +++ b/en/15.9/api/api-export.rst @@ -0,0 +1,94 @@ +======================== +Search Result Export API +======================== + +This document describes the v2 Export API of |Fess|, which downloads search results as a CSV or +JSON file. For the common response envelope and error model, see :doc:`api-overview`. + +The base URL is ``http:///api/v2/`` (local environment example: ``http://localhost:8080/api/v2``). + +.. note:: + + Export is disabled by default. To use it, set ``api.search.export=true`` in + ``fess_config.properties``. When it is enabled, the bundled ``bootstrap`` theme shows an export + menu (CSV / JSON) beside the result count. ``features.search_export`` of ``/api/v2/ui/config`` + reports the state. + +Downloading Search Results +========================== + +Request +------- + +================== ==================================================== +HTTP Method GET +Endpoint ``/api/v2/documents/export`` +================== ==================================================== + +Returns the documents matching the search as a file download (``Content-Disposition: attachment``, +file name ``search_results.csv`` or ``search_results.json``). + +- The same role filter as ``/api/v2/search`` applies. With ``login.required=true``, an access token + can be used the same way as with ``/api/v2/search``. +- At most ``api.search.export.max.size`` (default: ``1000``) documents are exported. The paging + parameters (``start``, ``num``) are not used. +- The exported fields are those in ``api.search.export.fields`` (default: + ``title,url_link,last_modified,content_length,filetype``) that may also appear in API responses. +- Requests are limited to ``api.search.export.rate.limit.per.minute`` (default: ``10``; ``0`` means + unlimited) per minute, counted per logged-in user, or per client IP for a guest. Over the limit the + endpoint answers ``429`` with a ``Retry-After`` header. +- An export is not recorded in the search log. + +Request Parameters +------------------ + +The same search condition parameters as ``/api/v2/documents/all`` can be given, such as ``q``, +``ex_q``, ``fields.*``, ``sort`` and ``lang`` (see :doc:`api-search`). In addition: + +.. tabularcolumns:: |p{4cm}|p{11cm}| +.. list-table:: Request Parameters + + * - ``format`` + - File format: ``csv`` (default) or ``json``. Any other value produces ``invalid_request`` (400). + +Table: Request Parameters + +Response +-------- + +The CSV file has a header row of field names and is written in the encoding of +``csv.file.encoding`` (a UTF-8 file starts with a byte order mark). A value that starts with ``=``, +``+``, ``-``, ``@``, a tab or a carriage return gets a leading ``'`` so that a spreadsheet does not +run it as a formula. A multi-valued field is joined with a space. + +:: + + "title","url_link","last_modified","content_length","filetype" + "Example","https://example.com/","2025-01-01T00:00:00.000Z","1234","html" + +The JSON file is ``{"data":[{...},...]}``, and multi-valued fields stay arrays. + +A failure before the file starts returns the usual error envelope. A failure after that cannot be +reported in the file, so the download ends early: a truncated CSV, or JSON that does not parse. + +Error Response +-------------- + +.. tabularcolumns:: |p{4cm}|p{11cm}| +.. list-table:: Error Response + + * - Status Code + - Description + * - 400 Bad Request + - A malformed query, a ``format`` other than ``csv`` / ``json``, or export disabled with + ``api.search.export=false``. + * - 401 Unauthorized + - When authentication is required (for example, an anonymous caller with ``login.required=true``). + * - 405 Method Not Allowed + - When the HTTP method is not allowed. + * - 429 Too Many Requests + - When the per-minute request limit is exceeded. + * - 500 Internal Server Error + - When an internal server error occurs. + +Table: Error Response diff --git a/en/15.9/api/api-search-history.rst b/en/15.9/api/api-search-history.rst new file mode 100644 index 00000000..56e71e11 --- /dev/null +++ b/en/15.9/api/api-search-history.rst @@ -0,0 +1,113 @@ +================== +Search History API +================== + +This document describes the v2 Search History API of |Fess|. +For the common response envelope and error model, see :doc:`api-overview`. + +The base URL is ``http:///api/v2/`` (local environment example: ``http://localhost:8080/api/v2``). + +.. note:: + + Search history is available while both ``search.history.enabled`` (default: ``true``) and the + search log are enabled. ``features.search_history`` of ``/api/v2/ui/config`` reports the state. + +Getting Recent Searches +======================= + +Request +------- + +================== ==================================================== +HTTP Method GET +Endpoint ``/api/v2/search-history`` +================== ==================================================== + +Returns the recent searches that the logged-in user made with ``/api/v2/search`` on the current +virtual host, newest first. A client can run one of them again with the returned conditions. + +- Only first-page searches are listed. Searches with the same conditions are merged into the newest + one, and searches without a query are left out. +- At most ``search.history.size`` (default: ``10``) searches are returned. +- Search logs are written by a job that runs every minute, so a search can take up to about a minute + to appear. +- The history is keyed on the logged-in user of the session. Anonymous callers receive + ``auth_required`` (401); an access token does not stand in for a login. +- When search history is disabled, the endpoint responds with ``invalid_request`` (400). + +There are no request parameters. + +Response +-------- + +On success (200), the following fields are returned directly under ``response`` of the common envelope. + +:: + + { + "response": { + "status": 0, + "record_count": 1, + "data": [ + { + "q": "fess", + "fields": { "label": ["docs"] }, + "sort": "last_modified.desc", + "requested_at": "2026-10-01T09:00:00Z", + "hit_count": 42 + } + ] + } + } + +.. tabularcolumns:: |p{4cm}|p{11cm}| +.. list-table:: Response Fields + + * - ``record_count`` + - Number of searches in ``data`` (int). + * - ``data`` + - Recent searches, newest first. The condition keys use the request parameter names of + ``/api/v2/search``; a key is omitted when the search did not use it. + * - ``data[].q`` + - The query (str). + * - ``data[].fields`` + - Field conditions given with ``fields.``, keyed by field name, each with its values. + * - ``data[].ex_q`` + - Extra queries (array of str). + * - ``data[].sort`` + - Sort order (str). + * - ``data[].lang`` + - Languages requested with ``lang`` (array of str). + * - ``data[].requested_at`` + - When the search was made (UTC, ISO-8601). + * - ``data[].hit_count`` + - Number of hits the search returned (int64). + +Table: Response Fields + +Error Response +-------------- + +.. tabularcolumns:: |p{4cm}|p{11cm}| +.. list-table:: Error Response + + * - Status Code + - Description + * - 400 Bad Request + - When search history is disabled. + * - 401 Unauthorized + - When the caller is not logged in. + * - 405 Method Not Allowed + - When the HTTP method is not allowed. + * - 500 Internal Server Error + - When an internal server error occurs. + +Table: Error Response + +In the Bundled Theme +==================== + +In the bundled ``bootstrap`` theme, a logged-in user sees the recent searches in the suggest +dropdown when they click the empty search box or press the down arrow key in it. Choosing an entry +runs the same search again, including conditions such as labels. Search logs recorded before +|Fess| 15.9 carry no conditions and are not listed. diff --git a/en/15.9/api/api-tag.rst b/en/15.9/api/api-tag.rst new file mode 100644 index 00000000..7103bb70 --- /dev/null +++ b/en/15.9/api/api-tag.rst @@ -0,0 +1,149 @@ +======== +Tags API +======== + +This document describes the v2 Tags API of |Fess|, which lets users tag documents. +For the common response envelope, error model, and CSRF, see :doc:`api-overview`. + +The base URL is ``http:///api/v2/`` (local environment example: ``http://localhost:8080/api/v2``). + +.. note:: + + Tags are disabled by default. To use them, set ``user.tag.enabled=true`` in + ``fess_config.properties``. ``features.user_tag`` of ``/api/v2/ui/config`` reports the state. + +A tag is a label of the kind "Tag" (see :doc:`../admin/labeltype-guide`): the label name is the tag +name, the value is the SHA-256 of the name in hex, the included paths are the tagged URLs, and the +permissions decide who can see the tag. A tag is visible only when its label is visible to the +caller. + +The search API (``/api/v2/search``) returns the tags of each hit that the caller can see as +``tags``. ``fields.tag=`` narrows the results to the documents with a tag, and +``facet.field=tag`` returns a tag facet. The ``tag`` index field itself is not returned. + +Getting Tags +============ + +Request +------- + +================== ==================================================== +HTTP Method GET +Endpoint ``/api/v2/documents/{docId}/tags`` +================== ==================================================== + +Returns the tags of the document that the caller can see. When the caller cannot search the +document, the endpoint responds with ``not_found`` (404). + +Response +-------- + +On success (200), the following fields are returned directly under ``response`` of the common envelope. + +:: + + { + "response": { + "status": 0, + "doc_id": "a1b2c3d4e5f6", + "addable": true, + "tags": [ + { "value": "9f86d081884c7d65...", "name": "to-review", "mine": true } + ] + } + } + +.. tabularcolumns:: |p{4cm}|p{11cm}| +.. list-table:: Response Fields + + * - ``doc_id`` + - Document ID (str). + * - ``addable`` + - ``true`` when the caller is logged in and can add tags (bool). + * - ``added`` + - POST only. ``false`` when the caller had already tagged the document (bool). + * - ``removed`` + - DELETE only (bool). + * - ``tags`` + - The tags that the caller can see. Each has ``value`` (the label value, used with + ``fields.tag``), ``name`` (the tag name) and ``mine`` (``true`` when the caller is in the + permissions of the tag). + +Table: Response Fields + +Adding a Tag +============ + +Request +------- + +================== ==================================================== +HTTP Method POST +Endpoint ``/api/v2/documents/{docId}/tags`` +================== ==================================================== + +Tags the document's URL for the logged-in user; an access token does not stand in for a login. As a +state-changing request, it requires the ``X-Fess-CSRF-Token`` header. + +- When a tag of the name exists, the URL is added to its included paths and the user to its + permissions. Otherwise, a tag that only the user can see is created. Tags with the same name + therefore merge into one, and users who added a tag of the same name can see where each other's + tags are. +- Tagging the same document again answers ``added: false``. +- A document can have at most ``user.tag.max.document.tags`` (default: ``100``) tags. + +Send ``Content-Type: application/json`` with the tag name in ``name``. + +:: + + { + "name": "to-review" + } + +The name is NFKC-normalized, runs of whitespace are collapsed and it is trimmed. It must be 1 to +``user.tag.name.max.length`` (default: ``50``) characters, and names with a control or format +character (such as a zero-width character or a bidi override) are refused. + +Removing a Tag +============== + +Request +------- + +================== ==================================================== +HTTP Method DELETE +Endpoint ``/api/v2/documents/{docId}/tags?value=`` +================== ==================================================== + +Removes the logged-in user from the permissions of the tag given by ``value``. When no permission of +a user, a group or a role is left, the tag is deleted and removed from the documents. When the user +is not in the permissions of the tag, the endpoint responds with ``forbidden`` (403). The +``X-Fess-CSRF-Token`` header is required. + +Error Response +============== + +.. tabularcolumns:: |p{4cm}|p{11cm}| +.. list-table:: Error Response + + * - Status Code + - Description + * - 400 Bad Request + - When the request is invalid (including when tags are disabled, the tag name is invalid or a + tag limit is exceeded). + * - 401 Unauthorized + - POST or DELETE without a login. + * - 403 Forbidden + - A missing or expired CSRF token, or a DELETE of a tag the user did not add. + * - 404 Not Found + - When the document is not found or the caller cannot search it. + * - 405 Method Not Allowed + - When the HTTP method is not allowed. + * - 413 Payload Too Large + - When the request body exceeds the size limit. + * - 415 Unsupported Media Type + - When the ``Content-Type`` is not supported. + * - 500 Internal Server Error + - When an internal server error occurs. + +Table: Error Response diff --git a/en/15.9/api/api-uiconfig.rst b/en/15.9/api/api-uiconfig.rst index 1caff994..39592fe4 100644 --- a/en/15.9/api/api-uiconfig.rst +++ b/en/15.9/api/api-uiconfig.rst @@ -44,6 +44,9 @@ On success (HTTP 200, UiConfigResponse), the following response is returned in t }, "features": { "user_favorite": false, + "search_history": true, + "search_export": false, + "user_tag": false, "popular_word": true, "suggest_search_log": true, "suggest_documents": true, @@ -186,6 +189,15 @@ All fields are required. * - ``user_favorite`` - boolean - Whether the user favorites feature is enabled. + * - ``search_history`` + - boolean + - Whether search history (``GET /api/v2/search-history``) is available (``true`` when both ``search.history.enabled`` and the search log are enabled). + * - ``search_export`` + - boolean + - Whether search result export (``GET /api/v2/documents/export``) is enabled (``api.search.export``). + * - ``user_tag`` + - boolean + - Whether tags (``/api/v2/documents/{docId}/tags``) are enabled (``user.tag.enabled``). * - ``popular_word`` - boolean - Whether the popular word feature is enabled. diff --git a/en/15.9/api/index.rst b/en/15.9/api/index.rst index 5fd967f1..c565ca18 100644 --- a/en/15.9/api/index.rst +++ b/en/15.9/api/index.rst @@ -22,6 +22,7 @@ the AI chat API, request and response formats, and authentication. :caption: Search APIs api-search + api-export api-label api-popularword api-suggest @@ -40,6 +41,8 @@ the AI chat API, request and response formats, and authentication. :caption: User Feature APIs api-favorite + api-search-history + api-tag api-click api-cache diff --git a/en/15.9/config/crawler-ocr.rst b/en/15.9/config/crawler-ocr.rst index e78a5c4a..ff182a3e 100644 --- a/en/15.9/config/crawler-ocr.rst +++ b/en/15.9/config/crawler-ocr.rst @@ -167,6 +167,7 @@ Operational Notes - For new web crawl configurations, the default excluded URL pattern excludes image URLs (jpg, png, gif, etc.). To crawl images on a web site, remove these from "Excluded URLs for Crawling". File crawls include images. - The crawler's size limits also apply. For the indexing size limit per file type (default: 10 MB), see :doc:`crawler-basic`. - OCR accuracy depends on the quality of the scan. Handwriting is generally not recognized well. +- Files that were indexed before OCR was enabled are not OCR'd as they are. With incremental crawling ("Check Last Modified" in :doc:`../admin/general-guide`) on, a file whose modification time has not changed is not fetched again by a re-crawl. To OCR those files, turn "Check Last Modified" off for one crawl, or delete the documents from the index and crawl again. Upgrade Notes ============= diff --git a/en/15.9/config/rate-limiting.rst b/en/15.9/config/rate-limiting.rst index 0328ebc4..0e78b178 100644 --- a/en/15.9/config/rate-limiting.rst +++ b/en/15.9/config/rate-limiting.rst @@ -102,6 +102,38 @@ Setting it to ``true`` disables robots.txt handling, including Crawl-delay. # Ignore robots.txt (default: false) crawler.ignore.robots.txt=false +Crawl-delay applies per origin and is capped at 60 seconds. It paces URLs, not requests, so in an +incremental crawl the HEAD and the GET of one URL are sent back to back. The interval of a crawl +configuration still works separately, as the wait before the next URL. + +robots.txt is interpreted according to RFC 9309. Between ``Allow:`` and ``Disallow:``, the longest +matching rule wins, and the start URLs are checked against robots.txt too. A URL that robots.txt +disallows is logged at INFO in ``fess-crawler.log`` and is not recorded as a failure URL. + +Backoff after 429/503 responses +------------------------------- + +When a server answers ``429 Too Many Requests`` or ``503 Service Unavailable``, |Fess| pauses +requests to that origin and retries the URL up to three times. The wait is the ``Retry-After`` +header when it is present, otherwise an exponential wait that starts at 10 seconds (up to 5 +minutes). When the last retry fails too, a WARN is written to ``fess-crawler.log``. + +When robots.txt cannot be fetched +--------------------------------- + +When fetching robots.txt fails with a 5xx, a 429 or a timeout, the URLs of that origin are put back +into the queue until the backoff ends. If three retries after the first failure also fail, no URL of +that origin is crawled for the rest of the crawl, and one WARN is written to ``fess-crawler.log``. +Up to 15.8, a robots.txt that could not be fetched meant "allow all". To restore that behavior, +set the following in the "Config Parameters" of the web crawl configuration: + +:: + + client.robotsTxtAllowOnUnavailable=true + +To turn robots.txt handling off entirely, use ``client.robotsTxtEnabled=false`` (per crawl +configuration) or ``crawler.ignore.robots.txt=true``. + All Rate Limiting Properties ============================ diff --git a/en/15.9/dev/theme-development.rst b/en/15.9/dev/theme-development.rst index 87061476..50d209eb 100644 --- a/en/15.9/dev/theme-development.rst +++ b/en/15.9/dev/theme-development.rst @@ -163,6 +163,16 @@ Serving and API itself (inline styles are allowed; inline scripts are not). Fonts or scripts from an external CDN are therefore not loaded; ship them in the theme. +- The ``Content-Security-Policy`` of the entry HTML contains + ``frame-ancestors 'none'``, and the response also carries + ``X-Frame-Options: DENY``, so the page is not shown in a frame of + another page. The ``frame-ancestors`` value can be changed with + ``theme.index.frame.ancestors`` in ``fess_config.properties`` + (default: ``'none'``). An empty value drops ``frame-ancestors`` + (``X-Frame-Options: DENY`` is still sent). Under + ``frame-ancestors 'none'``, WebKit-based browsers such as Safari leave + blank the frames that a theme shows from a ``blob:`` URL (a PDF + preview or a cached copy). Set the value to empty to show them. - The theme's SPA retrieves data such as search results and chat from the ``/api/v2/*`` API. diff --git a/en/15.9/install/fess-setup.rst b/en/15.9/install/fess-setup.rst index f677b711..450fd930 100644 --- a/en/15.9/install/fess-setup.rst +++ b/en/15.9/install/fess-setup.rst @@ -129,6 +129,19 @@ the System > Plugin page in the administration screen; see :doc:`../admin/plugin the version list, the jars and their checksums from that one Maven repository, such as an internal mirror, instead of the default release and snapshot repositories and GitHub. +``--repository`` also accepts a URL that starts with ``file:///``. On a server with no route to the +Internet, put a copy of the Maven repository on the server and name that directory. Name the directory +that contains the directory of each plugin (``/maven-metadata.xml`` and so on), the place that +corresponds to ``https://maven.codelibs.org/release/org/codelibs/fess/`` of the default repository. +Versions are resolved and checksums are verified the same way as with an HTTP repository. + +:: + + $ bin/fess-setup install plugin fess-ds-csv --repository file:///opt/maven-repo/org/codelibs/fess/ + +A scheme other than ``http``, ``https`` and ``file``, or a URL without a scheme, is refused with a +one-line error. + install plugin -------------- @@ -147,6 +160,13 @@ build of |Fess| also installs the snapshot builds of its own line, and prefers t See :doc:`../admin/plugin-guide` for examples. +Only the plugin types that |Fess| loads can be installed: names that start with ``fess-ds-``, +``fess-ingest-``, ``fess-script-``, ``fess-webapp-``, ``fess-thumbnail-``, ``fess-crawler-``, +``fess-llm-``, ``fess-storage-`` or ``fess-sso-``, the same ones ``list plugins`` shows. Any other name +(``fess-theme-*``, or a library such as ``fess`` or ``fess-crawler``) is refused with a one-line error +and exit code 2 before anything is downloaded. If one of several names is refused, none of them is +installed. Install a static theme with ``install theme``. + list plugins ------------ @@ -210,6 +230,10 @@ In such an environment, bring in the plugin jars: screen; see :doc:`../admin/plugin-guide`. 3. Start |Fess|. If you put the jars there while it was running, restart it. +Instead of bringing in each jar, you can also copy the needed part of the Maven repository (the +``org/codelibs/fess//`` directories) to the server and install with ``install plugin`` and +``--repository file:///...``. The checksums are verified in this case too. + Managing Themes =============== diff --git a/en/15.9/user/search-field.rst b/en/15.9/user/search-field.rst index 83afc6af..a6976b4f 100644 --- a/en/15.9/user/search-field.rst +++ b/en/15.9/user/search-field.rst @@ -66,6 +66,12 @@ By default, you can search using the following fields: * - favorite_count - Number of times the document has been added as a favorite - Numeric + * - owner + - Account name of the file owner + - Keyword + * - last_modifier + - Last modifier of the file + - Keyword Table: Available Field List @@ -80,6 +86,15 @@ If no field is specified, the search targets the title and content fields. Depen .. note:: Depending on the crawl target, some fields may not have values registered. For example, anchor is registered only during web crawling, and lang is registered only when the HTML has a language attribute. Fields such as segment (a session ID representing the crawl execution unit) and doc_id (an internal ID assigned by the system) can also be specified, but they are not used in normal searches. +owner and last_modifier are registered when crawling file servers and the like. owner holds the +account name of the file owner obtained by SMB, file system and FTP crawls (a value such as +``DOMAIN\alice`` becomes ``alice``). last_modifier holds the last author extracted from an Office +document or similar, and falls back to the owner. No owner is registered for HTML from a web crawl. +Search them as ``owner:alice`` or ``last_modifier:"Taro Yamada"``. In the bundled theme, the +advanced search can specify the owner and the last modifier. Administrators can turn the fields on +and off with ``crawler.document.file.owner.enabled`` and +``crawler.document.file.last.modifier.enabled`` (both ``true`` by default). + For HTML files, the title tag is stored in the title field, and the text under the body tag is stored in the content field. Usage diff --git a/es/15.9/admin/backup-guide.rst b/es/15.9/admin/backup-guide.rst index 90f5b107..839877c8 100644 --- a/es/15.9/admin/backup-guide.rst +++ b/es/15.9/admin/backup-guide.rst @@ -45,6 +45,11 @@ doc.json doc.json contiene información de mapeo del índice fess. +chat_log.ndjson +::::::::::::::: + +chat_log.ndjson contiene información de los registros de uso del chat con IA. + click_log.ndjson :::::::::::::::: diff --git a/es/15.9/admin/docreport-guide.rst b/es/15.9/admin/docreport-guide.rst new file mode 100644 index 00000000..4e771bc0 --- /dev/null +++ b/es/15.9/admin/docreport-guide.rst @@ -0,0 +1,72 @@ +===================== +Informe de documentos +===================== + +Descripción general +=================== + +La página Informe de documentos ayuda a ordenar servidores de archivos rastreados y similares. Muestra +los documentos con el mismo contenido y los que llevan mucho tiempo sin modificarse, y cada lista se +puede descargar como CSV. + +Para abrir la página, seleccione [Información del sistema > Informe de documentos] en el menú +izquierdo. Para verla se necesita el rol ``admin-docreport`` o ``admin-docreport-view``. La página +solo muestra y descarga informes; no modifica ningún documento. + +Ambas pestañas se pueden acotar con "Prefijo de URL", por ejemplo ``smb://server/share/``. + +Duplicados +========== + +Los documentos con el mismo contenido, o casi el mismo, se agrupan, empezando por el grupo más +grande. Los grupos usan la firma de contenido calculada al indexar (``content_minhash_bits``, la +misma que agrupa los resultados de búsqueda duplicados), por lo que no hace falta reindexar. Se +excluyen los documentos cuyo contenido no tiene palabras (como los archivos vacíos). + +La pantalla muestra hasta ``docreport.duplicate.group.size`` (predeterminado: 100) grupos y hasta +``docreport.duplicate.docs.size`` (predeterminado: 10) documentos por grupo. Use [Descargar CSV] para +obtener todos los grupos. El CSV lee todos los grupos, incluso en un índice grande, con las columnas +``group, groupSize, url, title, filename, contentLength, lastModified, owner, lastModifier, clickCount, docId``. + +.. note:: + + En un índice que no guarda la firma de contenido (las asignaciones ``cloud`` y ``aws``), el + informe de duplicados no está disponible. + +Documentos inactivos +==================== + +Se muestran, del más antiguo al más reciente, los documentos cuya última modificación es anterior al +número de días indicado ("Sin modificar durante (días)", 365 de forma predeterminada según +``docreport.dormant.days``). Los documentos sin fecha de última modificación no se muestran. Con +"Nunca abiertos desde los resultados de búsqueda" se excluyen los documentos en los que se ha hecho +clic desde los resultados de búsqueda. + +La pantalla muestra el número de documentos coincidentes, su tamaño total y una lista paginada. La +paginación se detiene en ``indexer.max.result.window.size``; use [Descargar CSV] para obtener los +documentos que quedan más allá. + +Configuración +============= + +Los siguientes ajustes de ``fess_config.properties`` ajustan el informe. + +.. list-table:: + :header-rows: 1 + :widths: 40 45 15 + + * - Propiedad + - Descripción + - Predeterminado + * - ``docreport.duplicate.group.size`` + - Número máximo de grupos de duplicados que muestra la pantalla + - ``100`` + * - ``docreport.duplicate.docs.size`` + - Número máximo de documentos que la pantalla muestra por grupo + - ``10`` + * - ``docreport.duplicate.export.page.size`` + - Número de firmas de contenido leídas por solicitud al descargar el CSV + - ``10000`` + * - ``docreport.dormant.days`` + - Número predeterminado de días desde la última modificación a partir del cual un documento está inactivo + - ``365`` diff --git a/es/15.9/admin/general-guide.rst b/es/15.9/admin/general-guide.rst index 19b1157a..6718f247 100644 --- a/es/15.9/admin/general-guide.rst +++ b/es/15.9/admin/general-guide.rst @@ -111,6 +111,13 @@ Comprobar fecha de última modificación Habilite esto para realizar rastreo diferencial. +Para las URL HTTP/HTTPS, cuando el ``Last-Modified`` de la respuesta HEAD no permite decidir (no hay +fecha de modificación en el índice, la respuesta HEAD no tiene ``Last-Modified`` o el HEAD devuelve un +estado distinto de 200/404), el GET se envía como GET condicional con ``If-None-Match`` (el ``ETag`` +indexado) y/o ``If-Modified-Since``. Una página que responde ``304 Not Modified`` se trata como una +página sin cambios y no se vuelve a obtener. Para obtener todo de nuevo, por ejemplo tras cambiar la +configuración de rastreo, desactive esta opción durante un rastreo. + Configuración de rastreadores simultáneos :::::::::::::::::::::::::::::::::::::::::: diff --git a/es/15.9/admin/index.rst b/es/15.9/admin/index.rst index b57b10de..a62d3bb1 100644 --- a/es/15.9/admin/index.rst +++ b/es/15.9/admin/index.rst @@ -50,6 +50,7 @@ roles, los registros y las copias de seguridad. log-guide failureurl-guide searchlist-guide + docreport-guide backup-guide maintenance-guide esreq-guide diff --git a/es/15.9/admin/labeltype-guide.rst b/es/15.9/admin/labeltype-guide.rst index f0cf62cb..9dcc23d4 100644 --- a/es/15.9/admin/labeltype-guide.rst +++ b/es/15.9/admin/labeltype-guide.rst @@ -85,11 +85,49 @@ Orden de clasificación Especifique el orden de clasificación de las etiquetas. +Tipo +:::: + +Indique "Etiqueta" o "Etiqueta de usuario". Una etiqueta normal es "Etiqueta". "Etiqueta de usuario" +es una etiqueta que los usuarios añaden desde la pantalla de búsqueda (consulte "Etiquetas de +usuario" más abajo). Una etiqueta existente sin tipo se trata como "Etiqueta". + + Eliminar configuración ---------------------- Haga clic en el nombre de la configuración en la página de lista y haga clic en el botón de eliminar para que aparezca una pantalla de confirmación. Al presionar el botón de eliminar, se eliminará la configuración. +Etiquetas de usuario +-------------------- + +Con ``user.tag.enabled=true`` (predeterminado: ``false``) en ``fess_config.properties``, los +usuarios que han iniciado sesión pueden etiquetar los resultados de búsqueda. En el tema incluido +``bootstrap``, las etiquetas se muestran en los resultados, los usuarios pueden añadir etiquetas y +quitar las suyas, y una faceta "Etiquetas" acota los resultados. Para la API, consulte +:doc:`../api/api-tag`. + +Una etiqueta de usuario se guarda como una etiqueta del tipo "Etiqueta de usuario": el nombre es el +nombre de la etiqueta, el valor es el SHA-256 del nombre, las rutas incluidas son las URL etiquetadas +(una por línea, coincidencia exacta) y los permisos deciden quién puede verla. El usuario que añade +una etiqueta se agrega a sus permisos. + +- Una etiqueta solo es visible cuando los permisos de su etiqueta coinciden con quien llama. Los + administradores pueden editarla en esta página para compartirla con un rol o un grupo, o + eliminarla. +- Las etiquetas con el mismo nombre se combinan en una sola, por lo que los usuarios que añadieron una + etiqueta con el mismo nombre ven dónde están las etiquetas de los demás. +- Las etiquetas de usuario no se incluyen en la API de lista de etiquetas (``/api/v2/labels``) ni en + las opciones de etiqueta de la pantalla de búsqueda. +- Cuentan para el límite de etiquetas (``page.labeltype.max.fetch.size``, predeterminado: 1000). Al + alcanzarlo, no se pueden crear nuevas. +- Después de que un administrador cambie o elimine una etiqueta de usuario en esta página, los + documentos indexados conservan los valores anteriores hasta que se vuelven a rastrear o se ejecuta + el trabajo "Label Updater". +- Un documento puede tener hasta ``user.tag.max.document.tags`` (predeterminado: 100) etiquetas de + usuario, y un nombre puede tener hasta ``user.tag.name.max.length`` (predeterminado: 50) + caracteres. + .. |image0| image:: ../../../resources/images/en/15.9/admin/labeltype-1.png .. |image1| image:: ../../../resources/images/en/15.9/admin/labeltype-2.png diff --git a/es/15.9/admin/mapping-guide.rst b/es/15.9/admin/mapping-guide.rst index 071b7b9b..66f6aa5c 100644 --- a/es/15.9/admin/mapping-guide.rst +++ b/es/15.9/admin/mapping-guide.rst @@ -49,6 +49,33 @@ Carga Puede cargar en el formato de diccionario de mapeo. +Diccionarios de mapeo incluidos +=============================== + +El ``mapping.txt`` predeterminado se usa al analizar campos de búsqueda como ``title`` y +``content``. Unifica hiragana, kana pequeños y katakana de ancho medio en katakana de ancho completo, +de modo que りんご, リンゴ y リンゴ coinciden entre sí. En |Fess| 15.9 también unifica las siguientes +grafías: + +- ゐ y ゑ (a イ y エ), kana pequeños como ゎ, ゕ, ゖ, ヮ, ヵ, ヶ y ㇰ-ㇿ, y ゝ y ゞ (a ヽ y ヾ) +- ヴ, ヴャ, ヴュ y ヴョ, el hiragana ゔ, ヴ de ancho medio, y ウ o う seguidos de una marca de sonoridad + combinable (U+3099); por ejemplo, ラヴ pasa a ラブ y レヴュー a レビユー + +Además, ‐ ‑ ‒ – — ― ⁻ ₋ − y - escritos justo después de un kana se tratan como la marca de vocal +larga ー (``prolonged_sound_mark_filter``), por lo que サ―バ- y サ−バ‐ coinciden con サーバー. El +guion ASCII (``-``) y los guiones después de kanji, letras o dígitos (東京-大阪, 2026−10−02) no se +modifican. + +``ja/mapping.txt`` para japonés (los campos ``*_ja``) mantiene el hiragana y los kana pequeños tal +cual, porque el analizador morfológico los necesita, y solo unifica grafías como ヴ. + +.. note:: + + Estos ajustes se aplican a un índice de documentos recién creado. Un índice existente conserva + sus ajustes de análisis y diccionarios hasta que se reindexa. Para aplicarlos a un índice + existente, reindexe con "Restablecer diccionarios" activado en la página :doc:`maintenance-guide`. + Restablecer los diccionarios sobrescribe las modificaciones de ``mapping.txt`` / + ``ja/mapping.txt`` hechas en la pantalla de administración. .. |image0| image:: ../../../resources/images/en/15.9/admin/mapping-1.png .. |image1| image:: ../../../resources/images/en/15.9/admin/mapping-2.png diff --git a/es/15.9/admin/relatedquery-guide.rst b/es/15.9/admin/relatedquery-guide.rst index 75747a2c..b8939b49 100644 --- a/es/15.9/admin/relatedquery-guide.rst +++ b/es/15.9/admin/relatedquery-guide.rst @@ -54,6 +54,76 @@ Eliminar configuración Haga clic en el nombre de la configuración en la página de lista y haga clic en el botón de eliminar para que aparezca una pantalla de confirmación. Al presionar el botón de eliminar, se eliminará la configuración. +Generar a partir de los registros de búsqueda +--------------------------------------------- + +Haga clic en el botón [Generar a partir de los registros de búsqueda] de la página de lista para +crear consultas relacionadas a partir de los registros de búsqueda recientes. Una búsqueda que la +misma sesión de usuario hace poco después de otra (una errata seguida de su corrección, o un +término amplio seguido de otro más específico) cuenta como una reformulación. Para los términos +buscados con frecuencia, las reformulaciones más habituales pasan a ser las consultas relacionadas +del término. + +Las consultas relacionadas se aplican a todos y amplían cada búsqueda de su término, por lo que la +generación es conservadora: + +- Solo se usan las búsquedas que puede ver un invitado. Un registro de búsqueda solo se lee cuando + todos sus roles cumplen ``suggest.search.log.permissions`` (el mismo ajuste que usa la sugerencia). +- No se usan términos de búsqueda que contengan un filtro de campo como ``label:"x"``, operadores, + comodines, ``sort:`` o un ``+`` / ``-`` inicial. +- Las palabras registradas en [Sugerir > Palabra no deseada] no se usan ni como términos ni como + consultas relacionadas. +- Un término y cada una de sus consultas relacionadas deben proceder de al menos + ``related_query.generate.min.sessions`` sesiones, y las reformulaciones deben tener resultados. +- Las entradas se generan por separado para cada host virtual. Los registros de búsqueda sin host + virtual se tratan como el host predeterminado. +- Los términos que ya tienen consultas relacionadas no se modifican (el resultado indica cuántos se + omitieron), y no se crean más entradas de las que la caché de consultas relacionadas puede cargar + (``page.relatedquery.max.fetch.size``). + +Las consultas relacionadas generadas se pueden editar o eliminar como las registradas a mano. No se +pueden generar mientras "Registro de búsqueda" o "Registro de usuario" estén desactivados en +[Sistema > General], y no puede empezar una segunda ejecución mientras otra está en curso. + +Los siguientes ajustes de ``fess_config.properties`` ajustan la generación. + +.. list-table:: + :header-rows: 1 + :widths: 45 40 15 + + * - Propiedad + - Descripción + - Predeterminado + * - ``related_query.generate.days`` + - Días de registros de búsqueda que se leen + - ``30`` + * - ``related_query.generate.term.size`` + - Número máximo de términos por host virtual + - ``100`` + * - ``related_query.generate.query.size`` + - Número máximo de consultas relacionadas por término + - ``5`` + * - ``related_query.generate.min.sessions`` + - Número mínimo de sesiones en las que deben aparecer un término y su consulta relacionada + - ``3`` + * - ``related_query.generate.session.interval`` + - Intervalo en el que una búsqueda cuenta como reformulación (minutos) + - ``10`` + * - ``related_query.generate.seed.log.size`` + - Número máximo de registros de búsqueda leídos por término + - ``1000`` + * - ``related_query.generate.seed.session.size`` + - Número máximo de sesiones leídas por término + - ``200`` + * - ``related_query.generate.log.fetch.size`` + - Número máximo de búsquedas posteriores leídas por término + - ``2000`` + * - ``related_query.generate.query.min.length`` + - Longitud mínima de un término y de una consulta relacionada (caracteres) + - ``2`` + * - ``related_query.generate.query.max.length`` + - Longitud máxima de un término y de una consulta relacionada (caracteres) + - ``50`` .. |image0| image:: ../../../resources/images/en/15.9/admin/relatedquery-1.png .. |image1| image:: ../../../resources/images/en/15.9/admin/relatedquery-2.png diff --git a/es/15.9/admin/searchlog-guide.rst b/es/15.9/admin/searchlog-guide.rst index c6c79d1c..f8102d5e 100644 --- a/es/15.9/admin/searchlog-guide.rst +++ b/es/15.9/admin/searchlog-guide.rst @@ -1,27 +1,103 @@ ==================== -Registro de Búsqueda +Registro de búsqueda ==================== Descripción general =================== -Los resultados de ejecución de búsqueda, clics y favoritos se registran, y los registros de búsqueda se pueden verificar en esta pantalla de administración. +Las búsquedas, los clics y los favoritos se registran. La página Registro de búsqueda muestra +informes de análisis que los agregan y una lista de los registros individuales. -Método de gestión -================== +Para abrir la página, seleccione [Información del sistema > Registro de búsqueda] en el menú +izquierdo. Primero se muestra la pestaña "Resumen". Para verla se necesita el rol +``admin-searchlog`` o ``admin-searchlog-view``; con ``admin-searchlog-view`` no se pueden eliminar +registros. + +Informes de análisis +==================== + +Período y filtros +----------------- + +En la parte superior de cada pestaña, elija el período que se agrega entre "Hoy", "Ayer", "Últimos +7 días", "Últimos 28 días" y "Últimos 90 días", o indique una fecha de inicio y de fin con +"Personalizado" (hasta 366 días). Con "Comparar con el período anterior", los valores se comparan con +el período de la misma duración inmediatamente anterior. También puede elegir el tipo de acceso y el +número de filas de las tablas. Los períodos y los intervalos de los gráficos siguen los días +naturales de la zona horaria del servidor. + +Pestañas +-------- + +- **Resumen**: las búsquedas, los usuarios, la tasa de cero resultados, la tasa de clics y el tiempo + de respuesta medio, cada uno con un pequeño gráfico de tendencia y su variación respecto al período + anterior; un gráfico de tendencia con selector de métrica (al comparar, el período anterior se + dibuja con líneas discontinuas); y las consultas más frecuentes y las consultas sin resultados. +- **Consultas**: por consulta, las búsquedas, los usuarios, los resultados medios, los clics, la tasa + de clics y la posición media de clic. También muestra las consultas sin resultados (con la fecha de + la última búsqueda) y las consultas sin clics, que tuvieron resultados que nunca se abrieron. +- **Clics**: las URL más pulsadas, las URL más añadidas a favoritos, la distribución de la posición de + clic y la proporción de visitas a la página 2 y siguientes. +- **Rendimiento**: el tiempo de respuesta medio, la mediana (p50), p95 y p99, la distribución del + tiempo de respuesta, las consultas más lentas y el tiempo de consulta. +- **Audiencia**: usuarios nuevos y recurrentes, tipos de acceso, búsquedas por día de la semana y hora, + y los agentes de usuario, referentes, idiomas y hosts virtuales más frecuentes. "Búsquedas por rol y + grupo" muestra las búsquedas, los usuarios y la tasa de cero resultados de cada rol y grupo. Una + búsqueda cuenta para todos los roles y grupos del usuario que la hizo, por lo que la suma de las + filas puede superar el total. No se muestran usuarios individuales. +- **Chat con IA**: las solicitudes, los usuarios, el total de tokens, el tiempo de respuesta medio y la + tasa de errores del modo de búsqueda con IA (chat RAG), con los usuarios más activos y las + solicitudes por intención y por modelo. El uso del chat se registra mientras + ``rag.chat.log.enabled`` (predeterminado: ``true``) está activado. Las preguntas y las respuestas no + se registran. El número de tokens solo se registra cuando el plugin de LLM lo informa. +- **Registros**: la lista de registros individuales; consulte "Lista de registros" más abajo. + +.. note:: -Lista -===== + Las métricas de clics por consulta y las consultas sin clics solo cuentan los clics registrados + desde que la palabra de búsqueda se guarda con los clics (|Fess| 15.9 y posteriores). El total de + clics y la tasa de clics incluyen los clics anteriores. Algunos valores, como el número de + usuarios, son aproximados. -En la lista puede verificar los registros de búsqueda de búsquedas, clics y favoritos. -Si desea verificar los detalles del registro de búsqueda, haga clic en el registro de búsqueda objetivo. +Revisar las consultas sin resultados +------------------------------------ + +En las pestañas "Resumen" y "Consultas", haga clic en una consulta sin resultados para abrir la +pestaña "Registros" con los registros de búsqueda de esa consulta, filtrados por "Solo sin +resultados". Así puede ver qué búsquedas no encontraron nada y añadir documentos, sinónimos o +consultas relacionadas. + +Descargar CSV +------------- + +Cada tabla y gráfico de los informes de análisis tiene un enlace CSV que descarga lo que agrega para +el período, la comparación, el tipo de acceso y el tamaño actuales. La barra de filtros también +tiene un enlace a un CSV de las métricas. Los números se escriben tal cual (proporciones de 0 a 1, +tiempos en milisegundos). Un gráfico comparado incluye una columna adicional ``_previous``. + +Lista de registros +================== + +La pestaña "Registros" muestra los registros de búsqueda, de clics, de favoritos y de usuario. Puede +filtrarlos por tipo de registro, ID de consulta, ID de usuario, intervalo de tiempo, tipo de acceso y +palabra de búsqueda, y los registros de búsqueda también por número de resultados ("Todos", "Solo +sin resultados", "Uno o más resultados"). Para ver los detalles de un registro, haga clic en él. |image0| +Haga clic en [Descargar CSV] para descargar como CSV los registros que coinciden con el filtro +actual, del más reciente al más antiguo y sin límite de filas. La fila de encabezado contiene los +nombres de los campos, por lo que no depende del idioma de la interfaz. + +Los archivos CSV, incluidos los de los informes de análisis, se escriben con la codificación de +``csv.file.encoding``; un archivo UTF-8 empieza con una marca de orden de bytes. Un valor que empieza +por ``=``, ``+``, ``-``, ``@``, un tabulador o un retorno de carro recibe un ``'`` inicial para que una +hoja de cálculo no lo ejecute como fórmula. + Detalles -======== +-------- -Al hacer clic en un registro de búsqueda desde la lista, se muestran los detalles del registro de búsqueda objetivo. +Haga clic en un registro de la lista para mostrar sus detalles. |image1| diff --git a/es/15.9/api/admin/api-admin-backup.rst b/es/15.9/api/admin/api-admin-backup.rst index 3702b4e7..1814260a 100644 --- a/es/15.9/api/admin/api-admin-backup.rst +++ b/es/15.9/api/admin/api-admin-backup.rst @@ -79,12 +79,13 @@ A continuación se muestra un ejemplo con la configuración predeterminada (cuan { "id": "system.properties", "name": "system.properties" }, { "id": "fess.json", "name": "fess.json" }, { "id": "doc.json", "name": "doc.json" }, + { "id": "chat_log.ndjson", "name": "chat_log.ndjson" }, { "id": "click_log.ndjson", "name": "click_log.ndjson" }, { "id": "favorite_log.ndjson", "name": "favorite_log.ndjson" }, { "id": "search_log.ndjson", "name": "search_log.ndjson" }, { "id": "user_info.ndjson", "name": "user_info.ndjson" } ], - "total": 10 + "total": 11 } } @@ -112,7 +113,7 @@ Según el tipo de ``{id}``, el contenido de la respuesta cambia de la siguiente - El propio archivo de definición de mapeo del índice (``fess_indices/fess.json`` / ``fess_indices/fess/doc.json``), sin modificar (``application/octet-stream``) * - ``*.bulk`` o nombre de índice sin extensión - Datos masivos generados al recorrer el índice con el mismo nombre que el objetivo (``application/octet-stream``). El nombre sin ``.bulk`` se trata como nombre del índice. - * - ``*.ndjson`` (``search_log`` / ``user_info`` / ``click_log`` / ``favorite_log``) + * - ``*.ndjson`` (``search_log`` / ``user_info`` / ``click_log`` / ``favorite_log`` / ``chat_log``) - Datos NDJSON del registro correspondiente (``application/x-ndjson``) .. note:: diff --git a/es/15.9/api/api-export.rst b/es/15.9/api/api-export.rst new file mode 100644 index 00000000..f34cccb6 --- /dev/null +++ b/es/15.9/api/api-export.rst @@ -0,0 +1,100 @@ +============================================ +API de exportación de resultados de búsqueda +============================================ + +Este documento describe la API de exportación v2 de |Fess|, que descarga los resultados de búsqueda +como un archivo CSV o JSON. Para el sobre de respuesta común y el modelo de errores, consulte +:doc:`api-overview`. + +La URL base es ``http:///api/v2/`` (ejemplo en entorno local: ``http://localhost:8080/api/v2``). + +.. note:: + + La exportación está deshabilitada de forma predeterminada. Para usarla, configure + ``api.search.export=true`` en ``fess_config.properties``. Cuando está habilitada, el tema incluido + ``bootstrap`` muestra un menú de exportación (CSV / JSON) junto al número de resultados. + ``features.search_export`` de ``/api/v2/ui/config`` indica el estado. + +Descargar los resultados de búsqueda +==================================== + +Solicitud +--------- + +================== ==================================================== +Método HTTP GET +Endpoint ``/api/v2/documents/export`` +================== ==================================================== + +Devuelve los documentos que coinciden con la búsqueda como una descarga de archivo +(``Content-Disposition: attachment``, con el nombre ``search_results.csv`` o ``search_results.json``). + +- Se aplica el mismo filtro de roles que en ``/api/v2/search``. Con ``login.required=true``, se puede + usar un token de acceso igual que con ``/api/v2/search``. +- Se exportan como máximo ``api.search.export.max.size`` (predeterminado: ``1000``) documentos. Los + parámetros de paginación (``start``, ``num``) no se usan. +- Se exportan los campos de ``api.search.export.fields`` (predeterminado: + ``title,url_link,last_modified,content_length,filetype``) que también pueden aparecer en las + respuestas de la API. +- Las solicitudes se limitan a ``api.search.export.rate.limit.per.minute`` (predeterminado: ``10``; + ``0`` significa sin límite) por minuto, contadas por usuario con sesión iniciada o, para un + invitado, por IP de cliente. Por encima del límite, el endpoint responde ``429`` con una cabecera + ``Retry-After``. +- Una exportación no se registra en el registro de búsqueda. + +Parámetros de solicitud +----------------------- + +Se pueden indicar los mismos parámetros de condición de búsqueda que en ``/api/v2/documents/all``, +como ``q``, ``ex_q``, ``fields.*``, ``sort`` y ``lang`` (consulte :doc:`api-search`). Además: + +.. tabularcolumns:: |p{4cm}|p{11cm}| +.. list-table:: Parámetros de solicitud + + * - ``format`` + - Formato del archivo: ``csv`` (predeterminado) o ``json``. Cualquier otro valor produce ``invalid_request`` (400). + +Tabla: Parámetros de solicitud + +Respuesta +--------- + +El archivo CSV tiene una fila de encabezado con los nombres de los campos y se escribe con la +codificación de ``csv.file.encoding`` (un archivo UTF-8 empieza con una marca de orden de bytes). Un +valor que empieza por ``=``, ``+``, ``-``, ``@``, un tabulador o un retorno de carro recibe un ``'`` +inicial para que una hoja de cálculo no lo ejecute como fórmula. Un campo con varios valores se une +con un espacio. + +:: + + "title","url_link","last_modified","content_length","filetype" + "Example","https://example.com/","2025-01-01T00:00:00.000Z","1234","html" + +El archivo JSON tiene la forma ``{"data":[{...},...]}`` y los campos con varios valores siguen siendo +arrays. + +Un fallo antes de que empiece el archivo devuelve el sobre de error habitual. Un fallo posterior no se +puede indicar en el archivo, por lo que la descarga termina antes de tiempo: un CSV truncado o un +JSON que no se puede analizar. + +Respuesta de error +------------------ + +.. tabularcolumns:: |p{4cm}|p{11cm}| +.. list-table:: Respuesta de error + + * - Código de estado + - Descripción + * - 400 Bad Request + - Una consulta mal formada, un ``format`` distinto de ``csv`` / ``json``, o la exportación + deshabilitada con ``api.search.export=false``. + * - 401 Unauthorized + - Cuando se requiere autenticación (por ejemplo, una llamada anónima con ``login.required=true``). + * - 405 Method Not Allowed + - Cuando el método HTTP no está permitido. + * - 429 Too Many Requests + - Cuando se supera el límite de solicitudes por minuto. + * - 500 Internal Server Error + - Cuando se produce un error interno del servidor. + +Tabla: Respuesta de error diff --git a/es/15.9/api/api-search-history.rst b/es/15.9/api/api-search-history.rst new file mode 100644 index 00000000..ef9bddf7 --- /dev/null +++ b/es/15.9/api/api-search-history.rst @@ -0,0 +1,116 @@ +============================ +API de historial de búsqueda +============================ + +Este documento describe la API de historial de búsqueda v2 de |Fess|. +Para el sobre de respuesta común y el modelo de errores, consulte :doc:`api-overview`. + +La URL base es ``http:///api/v2/`` (ejemplo en entorno local: ``http://localhost:8080/api/v2``). + +.. note:: + + El historial de búsqueda está disponible mientras ``search.history.enabled`` (predeterminado: + ``true``) y el registro de búsqueda estén habilitados. ``features.search_history`` de + ``/api/v2/ui/config`` indica el estado. + +Obtener las búsquedas recientes +=============================== + +Solicitud +--------- + +================== ==================================================== +Método HTTP GET +Endpoint ``/api/v2/search-history`` +================== ==================================================== + +Devuelve las búsquedas recientes que el usuario que ha iniciado sesión hizo con ``/api/v2/search`` en +el host virtual actual, de la más reciente a la más antigua. Un cliente puede volver a ejecutar una +de ellas con las condiciones devueltas. + +- Solo se incluyen las búsquedas de la primera página. Las búsquedas con las mismas condiciones se + combinan en la más reciente, y se omiten las búsquedas sin consulta. +- Se devuelven como máximo ``search.history.size`` (predeterminado: ``10``) búsquedas. +- Los registros de búsqueda los escribe un trabajo que se ejecuta cada minuto, por lo que una búsqueda + puede tardar hasta un minuto aproximadamente en aparecer. +- El historial pertenece al usuario de la sesión. Las llamadas anónimas reciben ``auth_required`` + (401); un token de acceso no sustituye al inicio de sesión. +- Cuando el historial de búsqueda está deshabilitado, el endpoint responde con ``invalid_request`` (400). + +No hay parámetros de solicitud. + +Respuesta +--------- + +Si tiene éxito (200), se devuelven los siguientes campos directamente bajo ``response`` del sobre común. + +:: + + { + "response": { + "status": 0, + "record_count": 1, + "data": [ + { + "q": "fess", + "fields": { "label": ["docs"] }, + "sort": "last_modified.desc", + "requested_at": "2026-10-01T09:00:00Z", + "hit_count": 42 + } + ] + } + } + +.. tabularcolumns:: |p{4cm}|p{11cm}| +.. list-table:: Campos de respuesta + + * - ``record_count`` + - Número de búsquedas en ``data`` (int). + * - ``data`` + - Búsquedas recientes, de la más reciente a la más antigua. Las claves de las condiciones son los + nombres de los parámetros de ``/api/v2/search``; una clave se omite cuando la búsqueda no la usó. + * - ``data[].q`` + - La consulta (str). + * - ``data[].fields`` + - Condiciones de campo indicadas con ``fields.``, por nombre de campo, con sus valores. + * - ``data[].ex_q`` + - Consultas adicionales (array de str). + * - ``data[].sort`` + - Orden (str). + * - ``data[].lang`` + - Idiomas solicitados con ``lang`` (array de str). + * - ``data[].requested_at`` + - Momento de la búsqueda (UTC, ISO-8601). + * - ``data[].hit_count`` + - Número de resultados de la búsqueda (int64). + +Tabla: Campos de respuesta + +Respuesta de error +------------------ + +.. tabularcolumns:: |p{4cm}|p{11cm}| +.. list-table:: Respuesta de error + + * - Código de estado + - Descripción + * - 400 Bad Request + - Cuando el historial de búsqueda está deshabilitado. + * - 401 Unauthorized + - Cuando quien llama no ha iniciado sesión. + * - 405 Method Not Allowed + - Cuando el método HTTP no está permitido. + * - 500 Internal Server Error + - Cuando se produce un error interno del servidor. + +Tabla: Respuesta de error + +En el tema incluido +=================== + +En el tema incluido ``bootstrap``, un usuario que ha iniciado sesión ve sus búsquedas recientes en la +lista de sugerencias al hacer clic en el cuadro de búsqueda vacío o pulsar la flecha hacia abajo en +él. Al elegir una entrada se vuelve a ejecutar la misma búsqueda, incluidas condiciones como las +etiquetas. Los registros de búsqueda anteriores a |Fess| 15.9 no guardan las condiciones y no se +muestran. diff --git a/es/15.9/api/api-tag.rst b/es/15.9/api/api-tag.rst new file mode 100644 index 00000000..7b24ab4b --- /dev/null +++ b/es/15.9/api/api-tag.rst @@ -0,0 +1,154 @@ +=========================== +API de etiquetas de usuario +=========================== + +Este documento describe la API de etiquetas de usuario v2 de |Fess|, que permite a los usuarios +etiquetar documentos. Para el sobre de respuesta común, el modelo de errores y CSRF, consulte +:doc:`api-overview`. + +La URL base es ``http:///api/v2/`` (ejemplo en entorno local: ``http://localhost:8080/api/v2``). + +.. note:: + + Las etiquetas de usuario están deshabilitadas de forma predeterminada. Para usarlas, configure + ``user.tag.enabled=true`` en ``fess_config.properties``. ``features.user_tag`` de + ``/api/v2/ui/config`` indica el estado. + +Una etiqueta de usuario es una etiqueta del tipo "Etiqueta de usuario" (consulte +:doc:`../admin/labeltype-guide`): el nombre de la etiqueta es el nombre, el valor es el SHA-256 del +nombre en hexadecimal, las rutas incluidas son las URL etiquetadas y los permisos deciden quién puede +verla. Solo es visible cuando su etiqueta es visible para quien llama. + +La API de búsqueda (``/api/v2/search``) devuelve en ``tags`` las etiquetas de usuario de cada +resultado que quien llama puede ver. ``fields.tag=`` acota los resultados a los documentos con +una etiqueta, y ``facet.field=tag`` devuelve una faceta de etiquetas. El campo de índice ``tag`` no se +devuelve. + +Obtener las etiquetas +===================== + +Solicitud +--------- + +================== ==================================================== +Método HTTP GET +Endpoint ``/api/v2/documents/{docId}/tags`` +================== ==================================================== + +Devuelve las etiquetas de usuario del documento que quien llama puede ver. Si quien llama no puede +buscar el documento, el endpoint responde con ``not_found`` (404). + +Respuesta +--------- + +Si tiene éxito (200), se devuelven los siguientes campos directamente bajo ``response`` del sobre común. + +:: + + { + "response": { + "status": 0, + "doc_id": "a1b2c3d4e5f6", + "addable": true, + "tags": [ + { "value": "9f86d081884c7d65...", "name": "revisar", "mine": true } + ] + } + } + +.. tabularcolumns:: |p{4cm}|p{11cm}| +.. list-table:: Campos de respuesta + + * - ``doc_id`` + - ID del documento (str). + * - ``addable`` + - ``true`` cuando quien llama ha iniciado sesión y puede añadir etiquetas (bool). + * - ``added`` + - Solo POST. ``false`` cuando quien llama ya había etiquetado el documento (bool). + * - ``removed`` + - Solo DELETE (bool). + * - ``tags`` + - Las etiquetas de usuario que quien llama puede ver. Cada una tiene ``value`` (el valor de la + etiqueta, usado con ``fields.tag``), ``name`` (el nombre) y ``mine`` (``true`` cuando quien + llama está en los permisos de la etiqueta). + +Tabla: Campos de respuesta + +Añadir una etiqueta +=================== + +Solicitud +--------- + +================== ==================================================== +Método HTTP POST +Endpoint ``/api/v2/documents/{docId}/tags`` +================== ==================================================== + +Etiqueta la URL del documento para el usuario que ha iniciado sesión; un token de acceso no sustituye +al inicio de sesión. Como solicitud que cambia el estado, requiere la cabecera ``X-Fess-CSRF-Token``. + +- Si ya existe una etiqueta con ese nombre, la URL se añade a sus rutas incluidas y el usuario a sus + permisos. Si no, se crea una etiqueta que solo ese usuario puede ver. Por eso las etiquetas con el + mismo nombre se combinan en una, y los usuarios que añadieron una etiqueta con el mismo nombre ven + dónde están las etiquetas de los demás. +- Volver a etiquetar el mismo documento responde ``added: false``. +- Un documento puede tener como máximo ``user.tag.max.document.tags`` (predeterminado: ``100``) + etiquetas de usuario. + +Envíe ``Content-Type: application/json`` con el nombre en ``name``. + +:: + + { + "name": "revisar" + } + +El nombre se normaliza con NFKC, los espacios consecutivos se reducen a uno y se recorta. Debe tener +entre 1 y ``user.tag.name.max.length`` (predeterminado: ``50``) caracteres, y se rechazan los nombres +con un carácter de control o de formato (como un carácter de ancho cero o una anulación +bidireccional). + +Quitar una etiqueta +=================== + +Solicitud +--------- + +================== ==================================================== +Método HTTP DELETE +Endpoint ``/api/v2/documents/{docId}/tags?value=`` +================== ==================================================== + +Quita al usuario que ha iniciado sesión de los permisos de la etiqueta indicada en ``value``. Cuando no +queda ningún permiso de usuario, grupo o rol, la etiqueta se elimina y se quita de los documentos. Si +el usuario no está en los permisos de la etiqueta, el endpoint responde con ``forbidden`` (403). Se +requiere la cabecera ``X-Fess-CSRF-Token``. + +Respuesta de error +================== + +.. tabularcolumns:: |p{4cm}|p{11cm}| +.. list-table:: Respuesta de error + + * - Código de estado + - Descripción + * - 400 Bad Request + - Cuando la solicitud no es válida (también cuando las etiquetas de usuario están deshabilitadas, + el nombre no es válido o se supera un límite). + * - 401 Unauthorized + - POST o DELETE sin iniciar sesión. + * - 403 Forbidden + - Un token CSRF ausente o caducado, o un DELETE de una etiqueta que el usuario no añadió. + * - 404 Not Found + - Cuando el documento no se encuentra o quien llama no puede buscarlo. + * - 405 Method Not Allowed + - Cuando el método HTTP no está permitido. + * - 413 Payload Too Large + - Cuando el cuerpo de la solicitud supera el límite de tamaño. + * - 415 Unsupported Media Type + - Cuando el ``Content-Type`` no es compatible. + * - 500 Internal Server Error + - Cuando se produce un error interno del servidor. + +Tabla: Respuesta de error diff --git a/es/15.9/api/api-uiconfig.rst b/es/15.9/api/api-uiconfig.rst index 8847d8c6..450009f3 100644 --- a/es/15.9/api/api-uiconfig.rst +++ b/es/15.9/api/api-uiconfig.rst @@ -44,6 +44,9 @@ En caso de éxito (HTTP 200, UiConfigResponse), se devuelve una respuesta con el }, "features": { "user_favorite": false, + "search_history": true, + "search_export": false, + "user_tag": false, "popular_word": true, "suggest_search_log": true, "suggest_documents": true, @@ -186,6 +189,15 @@ Todos los campos son obligatorios. * - ``user_favorite`` - boolean - Si la función de favoritos de usuario está habilitada. + * - ``search_history`` + - boolean + - Si el historial de búsqueda (``GET /api/v2/search-history``) está disponible (``true`` cuando ``search.history.enabled`` y el registro de búsqueda están habilitados). + * - ``search_export`` + - boolean + - Si la exportación de resultados de búsqueda (``GET /api/v2/documents/export``) está habilitada (``api.search.export``). + * - ``user_tag`` + - boolean + - Si las etiquetas de usuario (``/api/v2/documents/{docId}/tags``) están habilitadas (``user.tag.enabled``). * - ``popular_word`` - boolean - Si la función de palabras populares está habilitada. diff --git a/es/15.9/api/index.rst b/es/15.9/api/index.rst index 2f751d75..7193f684 100644 --- a/es/15.9/api/index.rst +++ b/es/15.9/api/index.rst @@ -23,6 +23,7 @@ petición y respuesta, y la autenticación. :caption: API de búsqueda api-search + api-export api-label api-popularword api-suggest @@ -41,6 +42,8 @@ petición y respuesta, y la autenticación. :caption: API de funciones de usuario api-favorite + api-search-history + api-tag api-click api-cache diff --git a/es/15.9/config/crawler-ocr.rst b/es/15.9/config/crawler-ocr.rst index 28e673c6..cc3424c3 100644 --- a/es/15.9/config/crawler-ocr.rst +++ b/es/15.9/config/crawler-ocr.rst @@ -167,6 +167,7 @@ Notas operativas - En las configuraciones de rastreo web nuevas, el patrón de URL excluidas predeterminado excluye las URL de imágenes (jpg, png, gif, etc.). Para rastrear imágenes de un sitio web, elimínelas de "URL excluidas del rastreo". El rastreo de archivos incluye las imágenes. - También se aplican los límites de tamaño del rastreador. Para el límite de tamaño de indexación por tipo de archivo (predeterminado: 10 MB), consulte :doc:`crawler-basic`. - La precisión del OCR depende de la calidad del escaneo. En general, la escritura a mano no se reconoce bien. +- Los archivos indexados antes de activar el OCR no se procesan con OCR por sí solos. Con el rastreo incremental activado ("Comprobar fecha de última modificación" en :doc:`../admin/general-guide`), un archivo cuya fecha de modificación no ha cambiado no se vuelve a obtener al rastrear de nuevo. Para aplicar el OCR a esos archivos, desactive "Comprobar fecha de última modificación" durante un rastreo, o elimine los documentos del índice y vuelva a rastrear. Notas de actualización ====================== diff --git a/es/15.9/config/rate-limiting.rst b/es/15.9/config/rate-limiting.rst index 080e3866..a603b70a 100644 --- a/es/15.9/config/rate-limiting.rst +++ b/es/15.9/config/rate-limiting.rst @@ -102,6 +102,39 @@ Al establecerlo en ``true``, se deshabilita el manejo de robots.txt, incluyendo # Ignorar robots.txt (predeterminado: false) crawler.ignore.robots.txt=false +Crawl-delay se aplica por origen y tiene un máximo de 60 segundos. Espacia URL, no solicitudes, por +lo que en un rastreo incremental el HEAD y el GET de una URL se envían uno tras otro. El intervalo de +una configuración de rastreo sigue actuando por separado, como espera antes de la siguiente URL. + +robots.txt se interpreta según RFC 9309. Entre ``Allow:`` y ``Disallow:`` prevalece la regla +coincidente más larga, y las URL iniciales también se comprueban con robots.txt. Una URL que +robots.txt no permite se registra en INFO en ``fess-crawler.log`` y no se registra como URL fallida. + +Espera tras respuestas 429/503 +------------------------------ + +Cuando un servidor responde ``429 Too Many Requests`` o ``503 Service Unavailable``, |Fess| detiene +temporalmente las solicitudes a ese origen y reintenta la URL hasta tres veces. La espera es el valor +de la cabecera ``Retry-After`` si existe; si no, una espera exponencial que empieza en 10 segundos +(hasta 5 minutos). Si el último reintento también falla, se escribe un WARN en ``fess-crawler.log``. + +Cuando no se puede obtener robots.txt +------------------------------------- + +Cuando la obtención de robots.txt falla con un 5xx, un 429 o un tiempo de espera agotado, las URL de +ese origen vuelven a la cola hasta que termina la espera. Si tras el primer fallo también fallan tres +reintentos, no se rastrea ninguna URL de ese origen durante el resto del rastreo y se escribe un único +WARN en ``fess-crawler.log``. Hasta la versión 15.8, un robots.txt que no se podía obtener significaba +"permitir todo". Para recuperar ese comportamiento, indique lo siguiente en los "Parámetros de +configuración" de la configuración de rastreo web: + +:: + + client.robotsTxtAllowOnUnavailable=true + +Para desactivar por completo el tratamiento de robots.txt, use ``client.robotsTxtEnabled=false`` +(por configuración de rastreo) o ``crawler.ignore.robots.txt=true``. + Todas las opciones de configuración de límite de tasa ===================================================== diff --git a/es/15.9/dev/theme-development.rst b/es/15.9/dev/theme-development.rst index f5b22363..b12bc49b 100644 --- a/es/15.9/dev/theme-development.rst +++ b/es/15.9/dev/theme-development.rst @@ -176,6 +176,17 @@ Distribución y API estilos en línea, pero no los scripts en línea). Por lo tanto, las fuentes o los scripts de una CDN externa no se cargan; inclúyalos en el tema. +- La ``Content-Security-Policy`` del HTML de entrada contiene + ``frame-ancestors 'none'`` y la respuesta incluye también + ``X-Frame-Options: DENY``, por lo que la página no se muestra dentro del + marco de otra página. El valor de ``frame-ancestors`` se puede cambiar con + ``theme.index.frame.ancestors`` en ``fess_config.properties`` + (predeterminado: ``'none'``). Un valor vacío omite ``frame-ancestors`` + (``X-Frame-Options: DENY`` se sigue enviando). Con + ``frame-ancestors 'none'``, los navegadores basados en WebKit, como Safari, + dejan en blanco los marcos que un tema muestra desde una URL ``blob:`` + (la vista previa de un PDF o la copia en caché). Deje el valor vacío para + mostrarlos. - La SPA del tema obtiene datos como los resultados de búsqueda y el chat a través de la API ``/api/v2/*``. diff --git a/es/15.9/install/fess-setup.rst b/es/15.9/install/fess-setup.rst index bc0cf2e9..48d11c56 100644 --- a/es/15.9/install/fess-setup.rst +++ b/es/15.9/install/fess-setup.rst @@ -134,6 +134,20 @@ opción, la lista de versiones, los jar y sus sumas de comprobación se obtienen repositorio Maven, como un espejo interno, en lugar de los repositorios predeterminados de versiones publicadas y de snapshots y de GitHub. +``--repository`` también acepta una URL que empiece por ``file:///``. En un servidor sin acceso a +Internet, coloque una copia del repositorio Maven en el servidor e indique ese directorio. Indique el +directorio que contiene el directorio de cada plugin (``/maven-metadata.xml``, etc.), es decir, +el lugar que corresponde a ``https://maven.codelibs.org/release/org/codelibs/fess/`` del repositorio +predeterminado. Las versiones se resuelven y las sumas de comprobación se verifican igual que con un +repositorio HTTP. + +:: + + $ bin/fess-setup install plugin fess-ds-csv --repository file:///opt/maven-repo/org/codelibs/fess/ + +Un esquema distinto de ``http``, ``https`` y ``file``, o una URL sin esquema, se rechaza con un error +de una línea. + install plugin -------------- @@ -154,6 +168,13 @@ línea, y les da preferencia. Consulte :doc:`../admin/plugin-guide` para ver ejemplos. +Solo se pueden instalar los tipos de plugin que |Fess| carga: nombres que empiezan por ``fess-ds-``, +``fess-ingest-``, ``fess-script-``, ``fess-webapp-``, ``fess-thumbnail-``, ``fess-crawler-``, +``fess-llm-``, ``fess-storage-`` o ``fess-sso-``, los mismos que muestra ``list plugins``. Cualquier +otro nombre (``fess-theme-*`` o una biblioteca como ``fess`` o ``fess-crawler``) se rechaza con un error +de una línea y el código de salida 2 antes de descargar nada. Si se rechaza uno de varios nombres, no +se instala ninguno. Instale un tema estático con ``install theme``. + list plugins ------------ @@ -217,6 +238,10 @@ En un entorno así, lleve los jar de los plugins e instálelos de esta forma: instalación de plugins de la pantalla de administración (consulte :doc:`../admin/plugin-guide`). 3. Inicie |Fess|. Si colocó los jar con |Fess| en ejecución, reinícielo. +En lugar de llevar cada jar, también puede copiar la parte necesaria del repositorio Maven (los +directorios ``org/codelibs/fess//``) al servidor e instalar con ``install plugin`` y +``--repository file:///...``. En este caso también se verifican las sumas de comprobación. + Gestión de Temas ================ diff --git a/es/15.9/user/search-field.rst b/es/15.9/user/search-field.rst index b6ed37fa..c8a8d53b 100644 --- a/es/15.9/user/search-field.rst +++ b/es/15.9/user/search-field.rst @@ -66,6 +66,12 @@ Por defecto, puede buscar especificando los siguientes campos: * - favorite_count - Número de veces que el documento se agregó a favoritos - Numérico + * - owner + - Nombre de cuenta del propietario del archivo + - Palabra clave + * - last_modifier + - Última persona que modificó el archivo + - Palabra clave Tabla: Lista de campos disponibles @@ -80,6 +86,16 @@ Si no se especifica ningún campo, la búsqueda se realiza sobre los campos titl .. note:: Según el objetivo del rastreo, hay campos en los que no se registra ningún valor. Por ejemplo, anchor solo se registra durante el rastreo web, y lang solo cuando el HTML tiene un atributo de idioma. Además, también se pueden especificar campos como segment (el ID de sesión que representa la unidad de ejecución del rastreo) o doc_id (el ID interno asignado por el sistema), aunque normalmente no se utilizan en las búsquedas habituales. +owner y last_modifier se registran al rastrear servidores de archivos y similares. owner contiene el +nombre de cuenta del propietario del archivo obtenido en rastreos SMB, de sistema de archivos y FTP +(un valor como ``DOMAIN\alice`` pasa a ser ``alice``). last_modifier contiene el último autor +extraído de un documento de Office o similar y, si no existe, el propietario. No se registra owner +para el HTML de un rastreo web. Busque, por ejemplo, con ``owner:alice`` o +``last_modifier:"Taro Yamada"``. En el tema incluido, la búsqueda avanzada permite indicar el +propietario y el último modificador. Los administradores pueden activar y desactivar estos campos +con ``crawler.document.file.owner.enabled`` y ``crawler.document.file.last.modifier.enabled`` +(ambos ``true`` de forma predeterminada). + Cuando los archivos HTML son el objetivo de búsqueda, la etiqueta title se registra en el campo title, y el texto debajo de la etiqueta body se registra en el campo content. Cómo utilizar diff --git a/fr/15.9/admin/backup-guide.rst b/fr/15.9/admin/backup-guide.rst index ab2cbaf4..7d9c9e5c 100644 --- a/fr/15.9/admin/backup-guide.rst +++ b/fr/15.9/admin/backup-guide.rst @@ -45,6 +45,11 @@ doc.json doc.json contient les informations de mapping de l'index fess. +chat_log.ndjson +::::::::::::::: + +chat_log.ndjson contient les informations du journal d'utilisation du chat IA. + click_log.ndjson :::::::::::::::: diff --git a/fr/15.9/admin/docreport-guide.rst b/fr/15.9/admin/docreport-guide.rst new file mode 100644 index 00000000..66bb0eaf --- /dev/null +++ b/fr/15.9/admin/docreport-guide.rst @@ -0,0 +1,74 @@ +==================== +Rapport de documents +==================== + +Présentation +============ + +La page Rapport de documents aide à faire le ménage dans les serveurs de fichiers crawlés et autres. +Elle liste les documents au contenu identique et ceux qui n'ont pas été modifiés depuis longtemps ; +chaque liste peut être téléchargée en CSV. + +Pour ouvrir la page, sélectionnez [Informations système > Rapport de documents] dans le menu de +gauche. La consultation nécessite le rôle ``admin-docreport`` ou ``admin-docreport-view``. La page ne +fait qu'afficher et télécharger des rapports ; elle ne modifie aucun document. + +Les deux onglets peuvent être restreints avec « Préfixe d’URL », par exemple ``smb://server/share/``. + +Doublons +======== + +Les documents au contenu identique ou presque identique sont regroupés, le groupe le plus grand en +premier. Les groupes reposent sur la signature de contenu calculée à l'indexation +(``content_minhash_bits``, la même qui regroupe les résultats de recherche en double) ; aucune +réindexation n'est donc nécessaire. Les documents dont le contenu ne contient aucun mot (comme les +fichiers vides) sont exclus. + +L'écran affiche au plus ``docreport.duplicate.group.size`` (par défaut : 100) groupes et au plus +``docreport.duplicate.docs.size`` (par défaut : 10) documents par groupe. Utilisez [Télécharger le +CSV] pour obtenir tous les groupes. Le CSV lit tous les groupes, même sur un grand index, avec les +colonnes +``group, groupSize, url, title, filename, contentLength, lastModified, owner, lastModifier, clickCount, docId``. + +.. note:: + + Sur un index qui ne conserve pas la signature de contenu (les mappings ``cloud`` et ``aws``), le + rapport de doublons n'est pas disponible. + +Documents inactifs +================== + +Les documents dont la dernière modification est antérieure au nombre de jours indiqué (« Non modifiés +depuis (jours) », 365 par défaut d'après ``docreport.dormant.days``) sont listés, du plus ancien au +plus récent. Les documents sans date de dernière modification ne sont pas listés. Avec « Jamais +ouverts depuis les résultats de recherche », les documents cliqués depuis les résultats de recherche +sont exclus. + +L'écran affiche le nombre de documents correspondants, leur taille totale et une liste paginée. La +pagination s'arrête à ``indexer.max.result.window.size`` ; utilisez [Télécharger le CSV] pour obtenir +les documents au-delà. + +Paramètres +========== + +Les réglages suivants de ``fess_config.properties`` ajustent le rapport. + +.. list-table:: + :header-rows: 1 + :widths: 40 45 15 + + * - Propriété + - Description + - Défaut + * - ``docreport.duplicate.group.size`` + - Nombre maximal de groupes de doublons affichés par l’écran + - ``100`` + * - ``docreport.duplicate.docs.size`` + - Nombre maximal de documents listés par groupe à l’écran + - ``10`` + * - ``docreport.duplicate.export.page.size`` + - Nombre de signatures de contenu lues par requête lors du téléchargement du CSV + - ``10000`` + * - ``docreport.dormant.days`` + - Nombre de jours par défaut depuis la dernière modification au-delà duquel un document est inactif + - ``365`` diff --git a/fr/15.9/admin/general-guide.rst b/fr/15.9/admin/general-guide.rst index 2d20e2c8..4a6a75e9 100644 --- a/fr/15.9/admin/general-guide.rst +++ b/fr/15.9/admin/general-guide.rst @@ -111,6 +111,13 @@ Vérifier la dernière modification Active le crawl différentiel. +Pour les URL HTTP/HTTPS, quand le ``Last-Modified`` de la réponse HEAD ne permet pas de décider +(aucune date de modification dans l'index, pas de ``Last-Modified`` dans la réponse HEAD, ou un statut +HEAD autre que 200/404), le GET est envoyé comme GET conditionnel avec ``If-None-Match`` (l'``ETag`` +indexé) et/ou ``If-Modified-Since``. Une page qui répond ``304 Not Modified`` est traitée comme une +page inchangée et n'est pas récupérée à nouveau. Pour tout récupérer à nouveau, par exemple après +avoir modifié les paramètres de crawl, désactivez ce réglage le temps d'un crawl. + Configuration du robot d'exploration simultané ::::::::::::::::::::::::::::::::::::::::::::::: diff --git a/fr/15.9/admin/index.rst b/fr/15.9/admin/index.rst index 861f362e..88da8cf0 100644 --- a/fr/15.9/admin/index.rst +++ b/fr/15.9/admin/index.rst @@ -50,6 +50,7 @@ et rôles, journaux et sauvegardes. log-guide failureurl-guide searchlist-guide + docreport-guide backup-guide maintenance-guide esreq-guide diff --git a/fr/15.9/admin/labeltype-guide.rst b/fr/15.9/admin/labeltype-guide.rst index b56d48d4..b9d2aae5 100644 --- a/fr/15.9/admin/labeltype-guide.rst +++ b/fr/15.9/admin/labeltype-guide.rst @@ -85,11 +85,47 @@ Ordre de tri Spécifie l'ordre d'affichage des étiquettes. +Type +:::: + +Indiquez « Étiquette » ou « Tag ». Une étiquette ordinaire est une « Étiquette ». Un « Tag » est un tag +que les utilisateurs ajoutent depuis l'écran de recherche (voir « Tags » ci-dessous). Une étiquette +existante sans type est traitée comme une « Étiquette ». + + Suppression de configuration ---------------------------- Cliquez sur le nom de la configuration dans la page de liste, puis cliquez sur le bouton Supprimer pour afficher l'écran de confirmation. Appuyer sur le bouton Supprimer supprimera la configuration. +Tags +---- + +Avec ``user.tag.enabled=true`` (par défaut : ``false``) dans ``fess_config.properties``, les +utilisateurs connectés peuvent taguer les résultats de recherche. Dans le thème fourni +``bootstrap``, les tags s'affichent sur les résultats, les utilisateurs peuvent ajouter des tags et +retirer les leurs, et une facette « Tags » restreint les résultats. Pour l'API, voir +:doc:`../api/api-tag`. + +Un tag est enregistré comme une étiquette du type « Tag » : le nom est le nom du tag, la valeur est le +SHA-256 du nom, les chemins inclus sont les URL taguées (une par ligne, correspondance exacte) et les +permissions décident qui peut voir le tag. L'utilisateur qui ajoute un tag est ajouté à ses +permissions. + +- Un tag n'est visible que si les permissions de son étiquette correspondent à l'appelant. Les + administrateurs peuvent modifier un tag sur cette page pour le partager avec un rôle ou un groupe, + ou le supprimer. +- Les tags de même nom sont fusionnés en une seule étiquette ; les utilisateurs qui ont ajouté un tag + de même nom voient donc où se trouvent les tags des autres. +- Les tags ne figurent ni dans l'API de liste des étiquettes (``/api/v2/labels``) ni dans les choix + d'étiquettes de l'écran de recherche. +- Les tags comptent dans la limite des étiquettes (``page.labeltype.max.fetch.size``, par défaut : + 1000). Une fois la limite atteinte, aucun nouveau tag ne peut être créé. +- Après qu'un administrateur a modifié ou supprimé un tag sur cette page, les documents indexés + conservent les anciennes valeurs jusqu'à un nouveau crawl ou l'exécution du job « Label Updater ». +- Un document peut avoir jusqu'à ``user.tag.max.document.tags`` (par défaut : 100) tags, et un nom + de tag peut compter jusqu'à ``user.tag.name.max.length`` (par défaut : 50) caractères. + .. |image0| image:: ../../../resources/images/en/15.9/admin/labeltype-1.png .. |image1| image:: ../../../resources/images/en/15.9/admin/labeltype-2.png diff --git a/fr/15.9/admin/mapping-guide.rst b/fr/15.9/admin/mapping-guide.rst index 26d3e742..8506fcf7 100644 --- a/fr/15.9/admin/mapping-guide.rst +++ b/fr/15.9/admin/mapping-guide.rst @@ -49,6 +49,33 @@ Téléversement Vous pouvez téléverser au format de dictionnaire de mapping. +Dictionnaires de mapping fournis +================================ + +Le ``mapping.txt`` par défaut sert à l'analyse des champs de recherche comme ``title`` et +``content``. Il ramène les hiragana, les petits kana et les katakana demi-chasse aux katakana pleine +chasse, de sorte que りんご, リンゴ et リンゴ se correspondent. Dans |Fess| 15.9, il unifie aussi les +graphies suivantes : + +- ゐ et ゑ (en イ et エ), les petits kana comme ゎ, ゕ, ゖ, ヮ, ヵ, ヶ et ㇰ-ㇿ, ainsi que ゝ et ゞ (en ヽ et ヾ) +- ヴ, ヴャ, ヴュ et ヴョ, le hiragana ゔ, ヴ demi-chasse, et ウ ou う suivi d'une marque de voisement + combinante (U+3099) ; par exemple, ラヴ devient ラブ et レヴュー devient レビユー + +De plus, ‐ ‑ ‒ – — ― ⁻ ₋ − et - écrits juste après un kana sont traités comme la marque de voyelle +longue ー (``prolonged_sound_mark_filter``) : サ―バ- et サ−バ‐ correspondent donc à サーバー. Le +trait d'union ASCII (``-``) et un tiret après un kanji, une lettre ou un chiffre (東京-大阪, +2026−10−02) ne sont pas modifiés. + +``ja/mapping.txt`` pour le japonais (les champs ``*_ja``) conserve les hiragana et les petits kana +tels quels, car l'analyseur morphologique en a besoin, et n'unifie que des graphies comme ヴ. + +.. note:: + + Ces réglages s'appliquent à un index de documents nouvellement créé. Un index existant conserve + ses réglages d'analyse et ses dictionnaires jusqu'à sa réindexation. Pour les appliquer à un + index existant, réindexez avec « Réinitialiser les dictionnaires » activé sur la page + :doc:`maintenance-guide`. La réinitialisation écrase les modifications de ``mapping.txt`` / + ``ja/mapping.txt`` faites dans l'écran d'administration. .. |image0| image:: ../../../resources/images/en/15.9/admin/mapping-1.png .. |image1| image:: ../../../resources/images/en/15.9/admin/mapping-2.png diff --git a/fr/15.9/admin/relatedquery-guide.rst b/fr/15.9/admin/relatedquery-guide.rst index 69af7c00..e9c2808d 100644 --- a/fr/15.9/admin/relatedquery-guide.rst +++ b/fr/15.9/admin/relatedquery-guide.rst @@ -54,6 +54,77 @@ Suppression de configuration Cliquez sur le nom de la configuration dans la page de liste, puis cliquez sur le bouton Supprimer pour afficher l'écran de confirmation. Appuyer sur le bouton Supprimer supprimera la configuration. +Générer à partir des journaux de recherche +------------------------------------------ + +Cliquez sur le bouton [Générer à partir des journaux de recherche] de la page de liste pour créer +des requêtes associées à partir des journaux de recherche récents. Une recherche que la même session +utilisateur lance peu après une autre (une faute de frappe suivie de sa correction, ou un terme +général suivi d'un terme plus précis) compte comme une reformulation. Pour les termes souvent +recherchés, les reformulations les plus fréquentes deviennent les requêtes associées du terme. + +Les requêtes associées s'appliquent à tous et élargissent chaque recherche de leur terme ; la +génération est donc prudente : + +- Seules les recherches visibles par un invité sont utilisées. Un journal de recherche n'est lu que + si tous ses rôles satisfont ``suggest.search.log.permissions`` (le même réglage que la suggestion). +- Les termes contenant un filtre de champ comme ``label:"x"``, des opérateurs, des jokers, + ``sort:`` ou un ``+`` / ``-`` initial ne sont pas utilisés. +- Les mots enregistrés dans [Suggérer > Mot incorrect] ne servent ni de termes ni de requêtes + associées. +- Un terme et chacune de ses requêtes associées doivent provenir d'au moins + ``related_query.generate.min.sessions`` sessions, et les reformulations doivent avoir des + résultats. +- Les entrées sont générées séparément pour chaque hôte virtuel. Les journaux sans hôte virtuel sont + traités comme l'hôte par défaut. +- Les termes qui ont déjà des requêtes associées ne sont pas modifiés (le résultat indique combien + ont été ignorés), et pas plus d'entrées ne sont créées que le cache des requêtes associées ne peut + en charger (``page.relatedquery.max.fetch.size``). + +Les requêtes associées générées se modifient ou se suppriment comme celles saisies à la main. Elles +ne peuvent pas être générées tant que « Journal de recherche » ou « Journal utilisateur » est +désactivé dans [Système > Général], et une seconde exécution ne peut pas démarrer pendant qu'une +autre est en cours. + +Les réglages suivants de ``fess_config.properties`` ajustent la génération. + +.. list-table:: + :header-rows: 1 + :widths: 45 40 15 + + * - Propriété + - Description + - Défaut + * - ``related_query.generate.days`` + - Nombre de jours de journaux de recherche lus + - ``30`` + * - ``related_query.generate.term.size`` + - Nombre maximal de termes par hôte virtuel + - ``100`` + * - ``related_query.generate.query.size`` + - Nombre maximal de requêtes associées par terme + - ``5`` + * - ``related_query.generate.min.sessions`` + - Nombre minimal de sessions où un terme et sa requête associée doivent apparaître + - ``3`` + * - ``related_query.generate.session.interval`` + - Délai dans lequel une recherche compte comme reformulation (minutes) + - ``10`` + * - ``related_query.generate.seed.log.size`` + - Nombre maximal de journaux de recherche lus par terme + - ``1000`` + * - ``related_query.generate.seed.session.size`` + - Nombre maximal de sessions lues par terme + - ``200`` + * - ``related_query.generate.log.fetch.size`` + - Nombre maximal de recherches suivantes lues par terme + - ``2000`` + * - ``related_query.generate.query.min.length`` + - Longueur minimale d’un terme et d’une requête associée (caractères) + - ``2`` + * - ``related_query.generate.query.max.length`` + - Longueur maximale d’un terme et d’une requête associée (caractères) + - ``50`` .. |image0| image:: ../../../resources/images/en/15.9/admin/relatedquery-1.png .. |image1| image:: ../../../resources/images/en/15.9/admin/relatedquery-2.png diff --git a/fr/15.9/admin/searchlog-guide.rst b/fr/15.9/admin/searchlog-guide.rst index 757787d3..67c3993e 100644 --- a/fr/15.9/admin/searchlog-guide.rst +++ b/fr/15.9/admin/searchlog-guide.rst @@ -5,23 +5,102 @@ Journal de recherche Présentation ============ -Les résultats d'exécution des recherches, clics et favoris sont enregistrés, et les journaux de recherche peuvent être vérifiés dans cet écran d'administration. +Les recherches, les clics et les favoris sont enregistrés. La page Journal de recherche affiche des +rapports d'analyse qui les agrègent, ainsi qu'une liste des journaux individuels. -Gestion -======= +Pour ouvrir la page, sélectionnez [Informations système > Journal de recherche] dans le menu de +gauche. L'onglet « Aperçu » s'affiche en premier. La consultation nécessite le rôle +``admin-searchlog`` ou ``admin-searchlog-view`` ; ``admin-searchlog-view`` ne permet pas de supprimer +des journaux. -Liste -===== +Rapports d'analyse +================== -Dans la liste, vous pouvez vérifier les journaux de recherche, de clics et de favoris. -Si vous souhaitez vérifier les détails d'un journal de recherche, cliquez sur le journal de recherche cible. +Période et filtres +------------------ + +En haut de chaque onglet, choisissez la période à agréger parmi « Aujourd’hui », « Hier », +« 7 derniers jours », « 28 derniers jours » et « 90 derniers jours », ou indiquez une date de début +et de fin avec « Personnalisée » (366 jours au plus). Avec « Comparer à la période précédente », les +valeurs sont comparées à la période de même durée qui précède. Vous pouvez aussi choisir le type +d'accès et le nombre de lignes des tableaux. Les périodes et les intervalles des graphiques suivent +les jours calendaires du fuseau horaire du serveur. + +Onglets +------- + +- **Aperçu** : les recherches, les utilisateurs, le taux de zéro résultat, le taux de clics et le + temps de réponse moyen, chacun avec un petit graphique de tendance et son évolution par rapport à la + période précédente ; un graphique de tendance avec choix de la métrique (la période précédente est + tracée en pointillés lors d'une comparaison) ; ainsi que les requêtes les plus fréquentes et les + requêtes sans résultat. +- **Requêtes** : par requête, les recherches, les utilisateurs, le nombre moyen de résultats, les + clics, le taux de clics et la position moyenne des clics. Il affiche aussi les requêtes sans + résultat (avec la date de la dernière recherche) et les requêtes sans clic, qui avaient des + résultats jamais ouverts. +- **Clics** : les URL les plus cliquées, les URL les plus ajoutées aux favoris, la distribution des + positions de clic et la proportion de consultations de la page 2 et suivantes. +- **Performances** : le temps de réponse moyen, médian (p50), p95 et p99, la distribution des temps + de réponse, les requêtes les plus lentes et le temps de requête. +- **Audience** : les utilisateurs nouveaux et récurrents, les types d'accès, les recherches par jour + de la semaine et par heure, ainsi que les principaux user agents, référents, langues et hôtes + virtuels. « Recherches par rôle et par groupe » affiche les recherches, les utilisateurs et le taux + de zéro résultat de chaque rôle et groupe. Une recherche compte pour chaque rôle et groupe de + l'utilisateur qui l'a lancée, de sorte que la somme des lignes peut dépasser le total. Aucun + utilisateur individuel n'est affiché. +- **Chat IA** : les requêtes, les utilisateurs, le total de tokens, le temps de réponse moyen et le + taux d'erreur du mode de recherche IA (chat RAG), avec les principaux utilisateurs et les requêtes + par intention et par modèle. L'utilisation du chat est enregistrée tant que + ``rag.chat.log.enabled`` (par défaut : ``true``) est activé. Les questions et les réponses ne sont + pas enregistrées. Le nombre de tokens n'est enregistré que si le plugin LLM le communique. +- **Journaux** : la liste des journaux individuels ; voir « Liste des journaux » ci-dessous. + +.. note:: + + Les métriques de clics par requête et les requêtes sans clic ne comptent que les clics enregistrés + depuis que le mot recherché est enregistré avec les clics (|Fess| 15.9 et ultérieur). Le total des + clics et le taux de clics incluent les clics plus anciens. Certaines valeurs, comme le nombre + d'utilisateurs, sont approximatives. + +Examiner les requêtes sans résultat +----------------------------------- + +Dans les onglets « Aperçu » et « Requêtes », cliquez sur une requête sans résultat pour ouvrir +l'onglet « Journaux » avec les journaux de recherche de cette requête, filtrés sur « Sans résultat +uniquement ». Vous voyez ainsi quelles recherches n'ont rien trouvé et pouvez ajouter des documents, +des synonymes ou des requêtes associées. + +Télécharger le CSV +------------------ + +Chaque tableau et chaque graphique des rapports d'analyse a un lien CSV qui télécharge ce qu'il agrège +pour la période, la comparaison, le type d'accès et la taille en cours. La barre de filtres a aussi +un lien vers un CSV des métriques. Les nombres sont écrits tels quels (proportions de 0 à 1, durées +en millisecondes). Un graphique comparé reçoit une colonne supplémentaire ``_previous``. + +Liste des journaux +================== + +L'onglet « Journaux » liste les journaux de recherche, de clics, de favoris et utilisateur. Vous +pouvez les filtrer par type de journal, ID de requête, ID utilisateur, plage horaire, type d'accès et +mot recherché, et les journaux de recherche aussi par nombre de résultats (« Tous », « Sans résultat +uniquement », « Un résultat ou plus »). Pour voir les détails d'un journal, cliquez dessus. |image0| +Cliquez sur [Télécharger le CSV] pour télécharger en CSV les journaux qui correspondent au filtre en +cours, du plus récent au plus ancien et sans limite de lignes. La ligne d'en-tête contient les noms +des champs ; elle ne dépend donc pas de la langue de l'interface. + +Les fichiers CSV, y compris ceux des rapports d'analyse, sont écrits dans l'encodage de +``csv.file.encoding`` ; un fichier UTF-8 commence par une marque d'ordre des octets. Une valeur qui +commence par ``=``, ``+``, ``-``, ``@``, une tabulation ou un retour chariot reçoit un ``'`` initial, +afin qu'un tableur ne l'exécute pas comme une formule. + Détails -======= +------- -En cliquant sur un journal de recherche dans la liste, les détails du journal de recherche cible s'affichent. +Cliquez sur un journal de la liste pour afficher ses détails. |image1| diff --git a/fr/15.9/api/admin/api-admin-backup.rst b/fr/15.9/api/admin/api-admin-backup.rst index 67db0b23..bf697143 100644 --- a/fr/15.9/api/admin/api-admin-backup.rst +++ b/fr/15.9/api/admin/api-admin-backup.rst @@ -79,12 +79,13 @@ Voici un exemple avec la configuration par défaut (valeurs par défaut de ``ind { "id": "system.properties", "name": "system.properties" }, { "id": "fess.json", "name": "fess.json" }, { "id": "doc.json", "name": "doc.json" }, + { "id": "chat_log.ndjson", "name": "chat_log.ndjson" }, { "id": "click_log.ndjson", "name": "click_log.ndjson" }, { "id": "favorite_log.ndjson", "name": "favorite_log.ndjson" }, { "id": "search_log.ndjson", "name": "search_log.ndjson" }, { "id": "user_info.ndjson", "name": "user_info.ndjson" } ], - "total": 10 + "total": 11 } } @@ -112,7 +113,7 @@ Selon le type de ``{id}``, le contenu de la réponse change comme suit. - Le fichier de définition de mapping de l'index lui-même (``fess_indices/fess.json`` / ``fess_indices/fess/doc.json``), tel quel (``application/octet-stream``) * - ``*.bulk`` ou nom d'index sans extension - Données en masse générées en parcourant (scroll) l'index du même nom que la cible (``application/octet-stream``). Le nom sans ``.bulk`` est traité comme le nom de l'index. - * - ``*.ndjson`` (``search_log`` / ``user_info`` / ``click_log`` / ``favorite_log``) + * - ``*.ndjson`` (``search_log`` / ``user_info`` / ``click_log`` / ``favorite_log`` / ``chat_log``) - Données NDJSON du journal correspondant (``application/x-ndjson``) .. note:: diff --git a/fr/15.9/api/api-export.rst b/fr/15.9/api/api-export.rst new file mode 100644 index 00000000..7b71a645 --- /dev/null +++ b/fr/15.9/api/api-export.rst @@ -0,0 +1,100 @@ +============================================ +API d'exportation des résultats de recherche +============================================ + +Ce document décrit l'API d'exportation v2 de |Fess|, qui télécharge les résultats de recherche sous +forme de fichier CSV ou JSON. Pour l'enveloppe de réponse commune et le modèle d'erreur, voir +:doc:`api-overview`. + +L'URL de base est ``http:///api/v2/`` (exemple en environnement local : ``http://localhost:8080/api/v2``). + +.. note:: + + L'exportation est désactivée par défaut. Pour l'utiliser, définissez ``api.search.export=true`` + dans ``fess_config.properties``. Lorsqu'elle est activée, le thème fourni ``bootstrap`` affiche un + menu d'exportation (CSV / JSON) à côté du nombre de résultats. ``features.search_export`` de + ``/api/v2/ui/config`` indique l'état. + +Télécharger les résultats de recherche +====================================== + +Requête +------- + +==================== ==================================================== +Méthode HTTP GET +Point de terminaison ``/api/v2/documents/export`` +==================== ==================================================== + +Renvoie les documents correspondant à la recherche sous forme de téléchargement de fichier +(``Content-Disposition: attachment``, nom ``search_results.csv`` ou ``search_results.json``). + +- Le même filtre de rôles que pour ``/api/v2/search`` s'applique. Avec ``login.required=true``, un + jeton d'accès s'utilise comme avec ``/api/v2/search``. +- Au plus ``api.search.export.max.size`` (par défaut : ``1000``) documents sont exportés. Les + paramètres de pagination (``start``, ``num``) ne sont pas utilisés. +- Les champs exportés sont ceux de ``api.search.export.fields`` (par défaut : + ``title,url_link,last_modified,content_length,filetype``) qui peuvent aussi figurer dans les + réponses de l'API. +- Les requêtes sont limitées à ``api.search.export.rate.limit.per.minute`` (par défaut : ``10`` ; + ``0`` signifie illimité) par minute, comptées par utilisateur connecté ou, pour un invité, par IP + cliente. Au-delà, l'endpoint répond ``429`` avec un en-tête ``Retry-After``. +- Une exportation n'est pas enregistrée dans le journal de recherche. + +Paramètres de requête +--------------------- + +Les mêmes paramètres de condition de recherche que pour ``/api/v2/documents/all`` peuvent être +indiqués, comme ``q``, ``ex_q``, ``fields.*``, ``sort`` et ``lang`` (voir :doc:`api-search`). En +plus : + +.. tabularcolumns:: |p{4cm}|p{11cm}| +.. list-table:: Paramètres de requête + + * - ``format`` + - Format du fichier : ``csv`` (par défaut) ou ``json``. Toute autre valeur produit ``invalid_request`` (400). + +Tableau : Paramètres de requête + +Réponse +------- + +Le fichier CSV a une ligne d'en-tête avec les noms des champs et est écrit dans l'encodage de +``csv.file.encoding`` (un fichier UTF-8 commence par une marque d'ordre des octets). Une valeur qui +commence par ``=``, ``+``, ``-``, ``@``, une tabulation ou un retour chariot reçoit un ``'`` initial +afin qu'un tableur ne l'exécute pas comme une formule. Un champ à plusieurs valeurs est joint par un +espace. + +:: + + "title","url_link","last_modified","content_length","filetype" + "Example","https://example.com/","2025-01-01T00:00:00.000Z","1234","html" + +Le fichier JSON a la forme ``{"data":[{...},...]}`` et les champs à plusieurs valeurs restent des +tableaux. + +Un échec avant le début du fichier renvoie l'enveloppe d'erreur habituelle. Un échec ultérieur ne +peut pas être signalé dans le fichier : le téléchargement s'arrête prématurément, avec un CSV tronqué +ou un JSON non analysable. + +Réponses d'erreur +----------------- + +.. tabularcolumns:: |p{4cm}|p{11cm}| +.. list-table:: Réponses d'erreur + + * - Code de statut + - Description + * - 400 Bad Request + - Une requête mal formée, un ``format`` autre que ``csv`` / ``json``, ou l'exportation + désactivée par ``api.search.export=false``. + * - 401 Unauthorized + - Lorsqu'une authentification est requise (par exemple, un appelant anonyme avec ``login.required=true``). + * - 405 Method Not Allowed + - Lorsque la méthode HTTP n'est pas autorisée. + * - 429 Too Many Requests + - Lorsque la limite de requêtes par minute est dépassée. + * - 500 Internal Server Error + - Lorsqu'une erreur interne du serveur se produit. + +Tableau : Réponses d'erreur diff --git a/fr/15.9/api/api-search-history.rst b/fr/15.9/api/api-search-history.rst new file mode 100644 index 00000000..c608d3ad --- /dev/null +++ b/fr/15.9/api/api-search-history.rst @@ -0,0 +1,116 @@ +================================ +API de l'historique de recherche +================================ + +Ce document décrit l'API de l'historique de recherche v2 de |Fess|. +Pour l'enveloppe de réponse commune et le modèle d'erreur, voir :doc:`api-overview`. + +L'URL de base est ``http:///api/v2/`` (exemple en environnement local : ``http://localhost:8080/api/v2``). + +.. note:: + + L'historique de recherche est disponible tant que ``search.history.enabled`` (par défaut : + ``true``) et le journal de recherche sont activés. ``features.search_history`` de + ``/api/v2/ui/config`` indique l'état. + +Obtenir les recherches récentes +=============================== + +Requête +------- + +==================== ==================================================== +Méthode HTTP GET +Point de terminaison ``/api/v2/search-history`` +==================== ==================================================== + +Renvoie les recherches récentes que l'utilisateur connecté a lancées avec ``/api/v2/search`` sur +l'hôte virtuel courant, de la plus récente à la plus ancienne. Un client peut relancer l'une d'elles +avec les conditions renvoyées. + +- Seules les recherches de la première page sont listées. Les recherches aux conditions identiques + sont fusionnées dans la plus récente, et les recherches sans requête sont omises. +- Au plus ``search.history.size`` (par défaut : ``10``) recherches sont renvoyées. +- Les journaux de recherche sont écrits par un job exécuté chaque minute ; une recherche peut donc + mettre environ une minute à apparaître. +- L'historique est lié à l'utilisateur connecté de la session. Les appelants anonymes reçoivent + ``auth_required`` (401) ; un jeton d'accès ne remplace pas une connexion. +- Si l'historique de recherche est désactivé, l'endpoint répond ``invalid_request`` (400). + +Il n'y a pas de paramètres de requête. + +Réponse +------- + +En cas de succès (200), les champs suivants sont renvoyés directement sous ``response`` de l'enveloppe commune. + +:: + + { + "response": { + "status": 0, + "record_count": 1, + "data": [ + { + "q": "fess", + "fields": { "label": ["docs"] }, + "sort": "last_modified.desc", + "requested_at": "2026-10-01T09:00:00Z", + "hit_count": 42 + } + ] + } + } + +.. tabularcolumns:: |p{4cm}|p{11cm}| +.. list-table:: Champs de réponse + + * - ``record_count`` + - Nombre de recherches dans ``data`` (int). + * - ``data`` + - Les recherches récentes, de la plus récente à la plus ancienne. Les clés des conditions sont + les noms des paramètres de ``/api/v2/search`` ; une clé est omise si la recherche ne l'a pas + utilisée. + * - ``data[].q`` + - La requête (str). + * - ``data[].fields`` + - Conditions de champ données avec ``fields.``, par nom de champ, avec leurs valeurs. + * - ``data[].ex_q`` + - Requêtes supplémentaires (tableau de str). + * - ``data[].sort`` + - Ordre de tri (str). + * - ``data[].lang`` + - Langues demandées avec ``lang`` (tableau de str). + * - ``data[].requested_at`` + - Date de la recherche (UTC, ISO-8601). + * - ``data[].hit_count`` + - Nombre de résultats de la recherche (int64). + +Tableau : Champs de réponse + +Réponses d'erreur +----------------- + +.. tabularcolumns:: |p{4cm}|p{11cm}| +.. list-table:: Réponses d'erreur + + * - Code de statut + - Description + * - 400 Bad Request + - Lorsque l'historique de recherche est désactivé. + * - 401 Unauthorized + - Lorsque l'appelant n'est pas connecté. + * - 405 Method Not Allowed + - Lorsque la méthode HTTP n'est pas autorisée. + * - 500 Internal Server Error + - Lorsqu'une erreur interne du serveur se produit. + +Tableau : Réponses d'erreur + +Dans le thème fourni +==================== + +Dans le thème fourni ``bootstrap``, un utilisateur connecté voit ses recherches récentes dans la liste +de suggestions lorsqu'il clique dans la zone de recherche vide ou y appuie sur la flèche vers le bas. +Choisir une entrée relance la même recherche, conditions comme les étiquettes comprises. Les journaux +de recherche enregistrés avant |Fess| 15.9 ne contiennent pas de conditions et ne sont pas listés. diff --git a/fr/15.9/api/api-tag.rst b/fr/15.9/api/api-tag.rst new file mode 100644 index 00000000..57e4e222 --- /dev/null +++ b/fr/15.9/api/api-tag.rst @@ -0,0 +1,151 @@ +============ +API des tags +============ + +Ce document décrit l'API des tags v2 de |Fess|, qui permet aux utilisateurs de taguer des documents. +Pour l'enveloppe de réponse commune, le modèle d'erreur et les jetons CSRF, voir :doc:`api-overview`. + +L'URL de base est ``http:///api/v2/`` (exemple en environnement local : ``http://localhost:8080/api/v2``). + +.. note:: + + Les tags sont désactivés par défaut. Pour les utiliser, définissez ``user.tag.enabled=true`` dans + ``fess_config.properties``. ``features.user_tag`` de ``/api/v2/ui/config`` indique l'état. + +Un tag est une étiquette du type « Tag » (voir :doc:`../admin/labeltype-guide`) : le nom de +l'étiquette est le nom du tag, la valeur est le SHA-256 du nom en hexadécimal, les chemins inclus sont +les URL taguées et les permissions décident qui peut voir le tag. Un tag n'est visible que si son +étiquette est visible pour l'appelant. + +L'API de recherche (``/api/v2/search``) renvoie dans ``tags`` les tags de chaque résultat que +l'appelant peut voir. ``fields.tag=`` restreint les résultats aux documents portant un tag, +et ``facet.field=tag`` renvoie une facette de tags. Le champ d'index ``tag`` lui-même n'est pas +renvoyé. + +Obtenir les tags +================ + +Requête +------- + +==================== ==================================================== +Méthode HTTP GET +Point de terminaison ``/api/v2/documents/{docId}/tags`` +==================== ==================================================== + +Renvoie les tags du document que l'appelant peut voir. Si l'appelant ne peut pas rechercher le +document, l'endpoint répond ``not_found`` (404). + +Réponse +------- + +En cas de succès (200), les champs suivants sont renvoyés directement sous ``response`` de l'enveloppe commune. + +:: + + { + "response": { + "status": 0, + "doc_id": "a1b2c3d4e5f6", + "addable": true, + "tags": [ + { "value": "9f86d081884c7d65...", "name": "a-relire", "mine": true } + ] + } + } + +.. tabularcolumns:: |p{4cm}|p{11cm}| +.. list-table:: Champs de réponse + + * - ``doc_id`` + - ID du document (str). + * - ``addable`` + - ``true`` lorsque l'appelant est connecté et peut ajouter des tags (bool). + * - ``added`` + - POST uniquement. ``false`` lorsque l'appelant avait déjà tagué le document (bool). + * - ``removed`` + - DELETE uniquement (bool). + * - ``tags`` + - Les tags que l'appelant peut voir. Chacun a ``value`` (la valeur de l'étiquette, utilisée avec + ``fields.tag``), ``name`` (le nom du tag) et ``mine`` (``true`` lorsque l'appelant figure dans + les permissions du tag). + +Tableau : Champs de réponse + +Ajouter un tag +============== + +Requête +------- + +==================== ==================================================== +Méthode HTTP POST +Point de terminaison ``/api/v2/documents/{docId}/tags`` +==================== ==================================================== + +Tague l'URL du document pour l'utilisateur connecté ; un jeton d'accès ne remplace pas une connexion. +Comme requête qui modifie l'état, elle exige l'en-tête ``X-Fess-CSRF-Token``. + +- Si un tag de ce nom existe, l'URL est ajoutée à ses chemins inclus et l'utilisateur à ses + permissions. Sinon, un tag visible par ce seul utilisateur est créé. Les tags de même nom + fusionnent donc en un seul, et les utilisateurs qui ont ajouté un tag de même nom voient où se + trouvent les tags des autres. +- Taguer à nouveau le même document renvoie ``added: false``. +- Un document peut avoir au plus ``user.tag.max.document.tags`` (par défaut : ``100``) tags. + +Envoyez ``Content-Type: application/json`` avec le nom du tag dans ``name``. + +:: + + { + "name": "a-relire" + } + +Le nom est normalisé en NFKC, les suites d'espaces sont réduites et il est rogné. Il doit compter de +1 à ``user.tag.name.max.length`` (par défaut : ``50``) caractères, et les noms contenant un caractère +de contrôle ou de format (comme un caractère de largeur nulle ou un forçage bidirectionnel) sont +refusés. + +Retirer un tag +============== + +Requête +------- + +==================== ==================================================== +Méthode HTTP DELETE +Point de terminaison ``/api/v2/documents/{docId}/tags?value=`` +==================== ==================================================== + +Retire l'utilisateur connecté des permissions du tag indiqué par ``value``. Lorsqu'il ne reste aucune +permission d'utilisateur, de groupe ou de rôle, le tag est supprimé et retiré des documents. Si +l'utilisateur ne figure pas dans les permissions du tag, l'endpoint répond ``forbidden`` (403). +L'en-tête ``X-Fess-CSRF-Token`` est requis. + +Réponses d'erreur +================= + +.. tabularcolumns:: |p{4cm}|p{11cm}| +.. list-table:: Réponses d'erreur + + * - Code de statut + - Description + * - 400 Bad Request + - Lorsque la requête est invalide (y compris lorsque les tags sont désactivés, que le nom est + invalide ou qu'une limite est dépassée). + * - 401 Unauthorized + - POST ou DELETE sans connexion. + * - 403 Forbidden + - Un jeton CSRF absent ou expiré, ou un DELETE d'un tag que l'utilisateur n'a pas ajouté. + * - 404 Not Found + - Lorsque le document est introuvable ou que l'appelant ne peut pas le rechercher. + * - 405 Method Not Allowed + - Lorsque la méthode HTTP n'est pas autorisée. + * - 413 Payload Too Large + - Lorsque le corps de la requête dépasse la taille maximale. + * - 415 Unsupported Media Type + - Lorsque le ``Content-Type`` n'est pas pris en charge. + * - 500 Internal Server Error + - Lorsqu'une erreur interne du serveur se produit. + +Tableau : Réponses d'erreur diff --git a/fr/15.9/api/api-uiconfig.rst b/fr/15.9/api/api-uiconfig.rst index b3d0b22a..b4bc8606 100644 --- a/fr/15.9/api/api-uiconfig.rst +++ b/fr/15.9/api/api-uiconfig.rst @@ -44,6 +44,9 @@ En cas de succès (HTTP 200, UiConfigResponse), une réponse au format d'envelop }, "features": { "user_favorite": false, + "search_history": true, + "search_export": false, + "user_tag": false, "popular_word": true, "suggest_search_log": true, "suggest_documents": true, @@ -186,6 +189,15 @@ Tous les champs sont obligatoires. * - ``user_favorite`` - boolean - Indique si la fonctionnalité de favoris utilisateur est activée. + * - ``search_history`` + - boolean + - Si l'historique de recherche (``GET /api/v2/search-history``) est disponible (``true`` lorsque ``search.history.enabled`` et le journal de recherche sont activés). + * - ``search_export`` + - boolean + - Si l'exportation des résultats de recherche (``GET /api/v2/documents/export``) est activée (``api.search.export``). + * - ``user_tag`` + - boolean + - Si les tags (``/api/v2/documents/{docId}/tags``) sont activés (``user.tag.enabled``). * - ``popular_word`` - boolean - Indique si la fonctionnalité de mots populaires est activée. diff --git a/fr/15.9/api/index.rst b/fr/15.9/api/index.rst index 575efa0d..6b69b9dd 100644 --- a/fr/15.9/api/index.rst +++ b/fr/15.9/api/index.rst @@ -23,6 +23,7 @@ et l'authentification. :caption: API de recherche api-search + api-export api-label api-popularword api-suggest @@ -41,6 +42,8 @@ et l'authentification. :caption: API des fonctions utilisateur api-favorite + api-search-history + api-tag api-click api-cache diff --git a/fr/15.9/config/crawler-ocr.rst b/fr/15.9/config/crawler-ocr.rst index ea74f5a8..11581b89 100644 --- a/fr/15.9/config/crawler-ocr.rst +++ b/fr/15.9/config/crawler-ocr.rst @@ -167,6 +167,7 @@ Remarques d'exploitation - Pour les nouvelles configurations de crawl Web, le motif d'URL exclues par défaut exclut les URL d'images (jpg, png, gif, etc.). Pour explorer les images d'un site Web, supprimez-les de « URL exclues du crawl ». Le crawl de fichiers inclut les images. - Les limites de taille du crawler s'appliquent également. Pour la limite de taille d'indexation par type de fichier (par défaut : 10 Mo), consultez :doc:`crawler-basic`. - La précision de l'OCR dépend de la qualité de la numérisation. L'écriture manuscrite n'est en général pas bien reconnue. +- Les fichiers indexés avant l'activation de l'OCR ne sont pas traités par l'OCR tels quels. Avec le crawl incrémental activé (« Vérifier la dernière modification » dans :doc:`../admin/general-guide`), un fichier dont la date de modification n'a pas changé n'est pas récupéré à nouveau lors d'un nouveau crawl. Pour appliquer l'OCR à ces fichiers, désactivez « Vérifier la dernière modification » le temps d'un crawl, ou supprimez les documents de l'index puis relancez le crawl. Remarques sur la mise à niveau ============================== diff --git a/fr/15.9/config/rate-limiting.rst b/fr/15.9/config/rate-limiting.rst index 762e5745..e35b3a5d 100644 --- a/fr/15.9/config/rate-limiting.rst +++ b/fr/15.9/config/rate-limiting.rst @@ -102,6 +102,42 @@ En le définissant sur ``true``, le traitement de robots.txt, y compris Crawl-de # Ignorer robots.txt (defaut : false) crawler.ignore.robots.txt=false +Crawl-delay s'applique par origine et est plafonné à 60 secondes. Il espace les URL, pas les +requêtes : lors d'un crawl incrémental, le HEAD et le GET d'une même URL partent donc l'un après +l'autre. L'intervalle d'une configuration de crawl agit toujours séparément, comme attente avant +l'URL suivante. + +robots.txt est interprété selon la RFC 9309. Entre ``Allow:`` et ``Disallow:``, la règle +correspondante la plus longue l'emporte, et les URL de départ sont elles aussi vérifiées. Une URL +interdite par robots.txt est journalisée au niveau INFO dans ``fess-crawler.log`` et n'est pas +enregistrée comme URL en échec. + +Attente après les réponses 429/503 +---------------------------------- + +Quand un serveur répond ``429 Too Many Requests`` ou ``503 Service Unavailable``, |Fess| suspend +les requêtes vers cette origine et réessaie l'URL jusqu'à trois fois. L'attente est la valeur de +l'en-tête ``Retry-After`` s'il est présent, sinon une attente exponentielle qui commence à +10 secondes (5 minutes au plus). Si la dernière tentative échoue aussi, un WARN est écrit dans +``fess-crawler.log``. + +Quand robots.txt ne peut pas être récupéré +------------------------------------------ + +Quand la récupération de robots.txt échoue avec un 5xx, un 429 ou un délai dépassé, les URL de cette +origine sont remises dans la file jusqu'à la fin de l'attente. Si trois nouvelles tentatives +échouent après le premier échec, aucune URL de cette origine n'est crawlée jusqu'à la fin du crawl, +et un seul WARN est écrit dans ``fess-crawler.log``. Jusqu'à la version 15.8, un robots.txt +impossible à récupérer valait « tout autoriser ». Pour retrouver ce comportement, indiquez ceci dans +les « Paramètres de configuration » de la configuration de crawl web : + +:: + + client.robotsTxtAllowOnUnavailable=true + +Pour désactiver entièrement le traitement de robots.txt, utilisez ``client.robotsTxtEnabled=false`` +(par configuration de crawl) ou ``crawler.ignore.robots.txt=true``. + Liste complète des propriétés de limitation de débit ====================== diff --git a/fr/15.9/dev/theme-development.rst b/fr/15.9/dev/theme-development.rst index 7a5149e6..7558fa34 100644 --- a/fr/15.9/dev/theme-development.rst +++ b/fr/15.9/dev/theme-development.rst @@ -168,6 +168,16 @@ Diffusion et API |Fess| lui-même (les styles en ligne sont autorisés ; les scripts en ligne ne le sont pas). Les polices ou scripts provenant d'un CDN externe ne sont donc pas chargés ; incluez-les dans le thème. +- La ``Content-Security-Policy`` du HTML d'entrée contient + ``frame-ancestors 'none'`` et la réponse porte aussi + ``X-Frame-Options: DENY`` ; la page n'est donc pas affichée dans un cadre + d'une autre page. La valeur de ``frame-ancestors`` se modifie avec + ``theme.index.frame.ancestors`` dans ``fess_config.properties`` + (par défaut : ``'none'``). Une valeur vide retire ``frame-ancestors`` + (``X-Frame-Options: DENY`` est toujours envoyé). Avec + ``frame-ancestors 'none'``, les navigateurs basés sur WebKit, comme Safari, + laissent vides les cadres qu'un thème affiche depuis une URL ``blob:`` + (aperçu d'un PDF ou copie en cache). Videz la valeur pour les afficher. - La SPA du thème récupère les données telles que les résultats de recherche ou le chat depuis l'API ``/api/v2/*``. diff --git a/fr/15.9/install/fess-setup.rst b/fr/15.9/install/fess-setup.rst index d50faafc..5be1f291 100644 --- a/fr/15.9/install/fess-setup.rst +++ b/fr/15.9/install/fess-setup.rst @@ -136,6 +136,20 @@ peuvent aussi être gérés depuis la page **Système > Plugin** de l'écran d'a option prend la liste des versions, les jars et leurs sommes de contrôle dans ce seul dépôt Maven, par exemple un miroir interne, au lieu des dépôts release et snapshot par défaut et de GitHub. +``--repository`` accepte aussi une URL qui commence par ``file:///``. Sur un serveur sans accès à +Internet, placez une copie du dépôt Maven sur le serveur et indiquez ce répertoire. Indiquez le +répertoire qui contient le répertoire de chaque plugin (``/maven-metadata.xml``, etc.), c'est-à-dire +l'emplacement qui correspond à ``https://maven.codelibs.org/release/org/codelibs/fess/`` du dépôt par +défaut. Les versions sont résolues et les sommes de contrôle vérifiées de la même façon qu'avec un dépôt +HTTP. + +:: + + $ bin/fess-setup install plugin fess-ds-csv --repository file:///opt/maven-repo/org/codelibs/fess/ + +Un schéma autre que ``http``, ``https`` et ``file``, ou une URL sans schéma, est refusé avec une erreur +d'une ligne. + install plugin -------------- @@ -155,6 +169,14 @@ de développement de |Fess| installe aussi les builds snapshot de sa propre lign Consultez :doc:`../admin/plugin-guide` pour des exemples. +Seuls les types de plugins que |Fess| charge peuvent être installés : les noms qui commencent par +``fess-ds-``, ``fess-ingest-``, ``fess-script-``, ``fess-webapp-``, ``fess-thumbnail-``, +``fess-crawler-``, ``fess-llm-``, ``fess-storage-`` ou ``fess-sso-``, c'est-à-dire ceux qu'affiche +``list plugins``. Tout autre nom (``fess-theme-*``, ou une bibliothèque comme ``fess`` ou +``fess-crawler``) est refusé avec une erreur d'une ligne et le code de sortie 2, avant tout +téléchargement. Si l'un de plusieurs noms est refusé, aucun n'est installé. Installez un thème +statique avec ``install theme``. + list plugins ------------ @@ -219,6 +241,10 @@ Dans un tel environnement, apportez les jars des plugins et installez-les ainsi 3. Démarrez |Fess|. Si vous avez placé les jars pendant que |Fess| était en cours d'exécution, redémarrez-le. +Au lieu d'apporter chaque jar, vous pouvez aussi copier la partie nécessaire du dépôt Maven (les +répertoires ``org/codelibs/fess//``) sur le serveur et installer avec ``install plugin`` et +``--repository file:///...``. Les sommes de contrôle sont vérifiées dans ce cas aussi. + Gestion des thèmes ================== diff --git a/fr/15.9/user/search-field.rst b/fr/15.9/user/search-field.rst index 6bf1a611..bc6ae88e 100644 --- a/fr/15.9/user/search-field.rst +++ b/fr/15.9/user/search-field.rst @@ -68,6 +68,12 @@ Par défaut, vous pouvez effectuer une recherche en spécifiant les champs suiva * - favorite_count - Nombre de fois où le document a été ajouté aux favoris - Numérique + * - owner + - Nom de compte du propriétaire du fichier + - Mot-clé + * - last_modifier + - Dernier modificateur du fichier + - Mot-clé Table : Liste des champs disponibles @@ -82,6 +88,16 @@ Si aucun champ n'est spécifié, la recherche porte sur les champs title et cont .. note:: Selon la cible du crawl, certains champs peuvent ne pas être renseignés. Par exemple, anchor n'est enregistré que lors d'un crawl Web, et lang uniquement lorsque le document HTML possède un attribut de langue. Par ailleurs, des champs tels que segment (identifiant de session représentant une exécution de crawl) ou doc_id (identifiant interne attribué par le système) peuvent également être spécifiés, mais ils ne sont pas utilisés dans le cadre d'une recherche normale. +owner et last_modifier sont enregistrés lors du crawl de serveurs de fichiers et autres. owner +contient le nom de compte du propriétaire du fichier obtenu par les crawls SMB, système de fichiers +et FTP (une valeur comme ``DOMAIN\alice`` devient ``alice``). last_modifier contient le dernier +auteur extrait d'un document Office ou similaire, à défaut le propriétaire. Aucun owner n'est +enregistré pour le HTML issu d'un crawl Web. Recherchez par exemple avec ``owner:alice`` ou +``last_modifier:"Taro Yamada"``. Dans le thème fourni, la recherche avancée permet d'indiquer le +propriétaire et le dernier modificateur. Les administrateurs peuvent activer ou désactiver ces +champs avec ``crawler.document.file.owner.enabled`` et +``crawler.document.file.last.modifier.enabled`` (``true`` par défaut tous les deux). + Lorsque des fichiers HTML sont ciblés par la recherche, le contenu de la balise title est enregistré dans le champ title, et le texte situé sous la balise body est enregistré dans le champ content. Utilisation diff --git a/ja/15.9/admin/backup-guide.rst b/ja/15.9/admin/backup-guide.rst index 8277caf3..ca61e9d8 100644 --- a/ja/15.9/admin/backup-guide.rst +++ b/ja/15.9/admin/backup-guide.rst @@ -45,6 +45,11 @@ doc.json doc.jsonはfessインデックスのマッピング情報を含みます。 +chat_log.ndjson +::::::::::::::: + +chat_log.ndjsonはAIチャットの利用ログの情報を含みます。 + click_log.ndjson :::::::::::::::: diff --git a/ja/15.9/admin/docreport-guide.rst b/ja/15.9/admin/docreport-guide.rst new file mode 100644 index 00000000..dc959cf6 --- /dev/null +++ b/ja/15.9/admin/docreport-guide.rst @@ -0,0 +1,55 @@ +==================== +ドキュメントレポート +==================== + +概要 +==== + +ドキュメントレポートは、クロールしたファイルサーバーなどの整理に役立つ管理画面です。内容が同じ文書と、長い間更新されていない文書を一覧表示し、それぞれ CSV でダウンロードできます。 + +画面を開くには、左メニューの [システム情報 > ドキュメントレポート] をクリックします。閲覧には ``admin-docreport`` または ``admin-docreport-view`` のロールが必要です。この画面は表示とダウンロードだけを行い、文書を変更しません。 + +どちらのタブも「URL の前方一致」(例: ``smb://server/share/`` )で対象を絞り込めます。 + +重複文書 +======== + +内容が同一またはほぼ同一の文書をグループにまとめ、文書数の多い順に表示します。グループ分けには、索引時に計算される内容の署名( ``content_minhash_bits`` 、検索結果の重複表示の折りたたみと同じもの)を使うため、再インデクシングは不要です。内容に語がない文書(空のファイルなど)は対象外です。 + +画面には最大 ``docreport.duplicate.group.size`` (デフォルト: 100)グループを表示し、各グループの文書は ``docreport.duplicate.docs.size`` (デフォルト: 10)件まで表示します。すべてのグループは [CSVをダウンロード] で取得できます。CSV は大きなインデックスでもすべてのグループを読み取り、 ``group, groupSize, url, title, filename, contentLength, lastModified, owner, lastModifier, clickCount, docId`` の列を出力します。 + +.. note:: + + 内容の署名を保持しないインデックス( ``cloud`` 、 ``aws`` 用のマッピング)では、重複文書レポートは利用できません。 + +休眠文書 +======== + +最終更新日時が指定した日数(「未更新の日数」、デフォルトは ``docreport.dormant.days`` の 365)より前の文書を、古い順に表示します。最終更新日時を持たない文書は表示しません。「検索結果から一度も開かれていない」を有効にすると、検索結果からクリックされたことのある文書を除きます。 + +画面には該当する文書の件数と合計サイズ、文書の一覧を表示します。一覧のページ送りには上限( ``indexer.max.result.window.size`` )があり、それを超える文書は [CSVをダウンロード] で取得します。 + +設定 +==== + +``fess_config.properties`` の次の設定で動作を調整できます。 + +.. list-table:: + :header-rows: 1 + :widths: 40 45 15 + + * - プロパティ + - 説明 + - デフォルト + * - ``docreport.duplicate.group.size`` + - 画面に表示する重複グループの最大数 + - ``100`` + * - ``docreport.duplicate.docs.size`` + - 画面で 1 つのグループに表示する文書の最大数 + - ``10`` + * - ``docreport.duplicate.export.page.size`` + - CSV のダウンロード時に 1 回に読み取る署名の数 + - ``10000`` + * - ``docreport.dormant.days`` + - 休眠文書とみなす、最終更新からの日数のデフォルト + - ``365`` diff --git a/ja/15.9/admin/general-guide.rst b/ja/15.9/admin/general-guide.rst index 52b35549..0b2a6f83 100644 --- a/ja/15.9/admin/general-guide.rst +++ b/ja/15.9/admin/general-guide.rst @@ -111,6 +111,8 @@ SSOタイプ 差分クロールを行う場合に有効にします。 +HTTP/HTTPS の URL では、HEAD の ``Last-Modified`` で判断できない場合(索引に最終更新日時がない、HEAD の応答に ``Last-Modified`` がない、HEAD が 200/404 以外を返した)に、GET に ``If-None-Match`` (索引に保存した ``ETag`` )や ``If-Modified-Since`` を付ける条件付き GET を行います。サーバーが ``304 Not Modified`` を返したページは、更新されていないページと同じく取得し直されません。設定を変えた後などにすべてを取得し直すには、この設定を無効にしてクロールします。 + 同時クローラー設定 :::::::::::::: diff --git a/ja/15.9/admin/index.rst b/ja/15.9/admin/index.rst index c011e646..a645b622 100644 --- a/ja/15.9/admin/index.rst +++ b/ja/15.9/admin/index.rst @@ -49,6 +49,7 @@ log-guide failureurl-guide searchlist-guide + docreport-guide backup-guide maintenance-guide esreq-guide diff --git a/ja/15.9/admin/labeltype-guide.rst b/ja/15.9/admin/labeltype-guide.rst index cbaf0317..19d74f67 100644 --- a/ja/15.9/admin/labeltype-guide.rst +++ b/ja/15.9/admin/labeltype-guide.rst @@ -85,11 +85,31 @@ ラベルの表示順を指定します。 +種類 +:::: + +「ラベル」または「タグ」を指定します。通常のラベルは「ラベル」です。「タグ」は、ユーザーが検索画面から追加するタグです(下の「タグ」を参照)。種類を指定していない既存のラベルは「ラベル」として扱われます。 + + 設定の削除 -------- 一覧ページの設定名をクリックし、削除ボタンをクリックすると確認画面が表示されます。 削除ボタンを押すと設定が削除されます。 +タグ +---- + +``fess_config.properties`` で ``user.tag.enabled=true`` (デフォルト: ``false`` )にすると、ログインしたユーザーが検索結果にタグを付けられるようになります。同梱の ``bootstrap`` テーマでは、結果にタグが表示され、タグの追加と自分が付けたタグの削除ができ、「タグ」のファセットで絞り込めます。API については :doc:`../api/api-tag` を参照してください。 + +タグは種類が「タグ」のラベルとして保存されます。名前がタグ名、値がタグ名の SHA-256、対象とするパスがタグを付けた URL(1 行に 1 つ、完全一致)、パーミッションがタグを見られるユーザーです。タグを付けたユーザーはパーミッションに追加されます。 + +- タグが見えるのは、そのラベルのパーミッションが呼び出し元に一致する場合だけです。管理者はこの画面でタグを編集して、ロールやグループと共有したり、削除したりできます。 +- 同じ名前のタグは 1 つのラベルにまとまります。そのため、同じ名前のタグを付けたユーザーどうしは、互いのタグの付与先を見られます。 +- タグはラベルの一覧 API( ``/api/v2/labels`` )や検索画面のラベルの選択肢には含まれません。 +- タグはラベルの件数の上限( ``page.labeltype.max.fetch.size`` 、デフォルト: 1000)に含まれます。上限に達すると新しいタグを作成できません。 +- 管理者がこの画面でタグを変更または削除した後、インデックスのドキュメントは、再クロールするか「Label Updater」ジョブを実行するまで以前の値のままです。 +- 1 つのドキュメントのタグの数は ``user.tag.max.document.tags`` (デフォルト: 100)、タグ名の長さは ``user.tag.name.max.length`` (デフォルト: 50)文字までです。 + .. |image0| image:: ../../../resources/images/ja/15.9/admin/labeltype-1.png .. |image1| image:: ../../../resources/images/ja/15.9/admin/labeltype-2.png diff --git a/ja/15.9/admin/mapping-guide.rst b/ja/15.9/admin/mapping-guide.rst index 672cf292..c057f5fe 100644 --- a/ja/15.9/admin/mapping-guide.rst +++ b/ja/15.9/admin/mapping-guide.rst @@ -49,6 +49,21 @@ マッピングの辞書形式でアップロードすることができます。 +同梱のマッピング辞書 +==================== + +既定の ``mapping.txt`` は、 ``title`` や ``content`` などの検索用フィールドの解析で使われ、ひらがな、小書きのかな、半角カナを全角のカタカナにそろえます。そのため「りんご」「リンゴ」「リンゴ」は互いに一致します。 |Fess| 15.9 では、次の表記もそろえるようになりました。 + +- ゐ・ゑ(イ・エ)、ゎ・ゕ・ゖ・ヮ・ヵ・ヶ・ㇰ〜ㇿ などの小書きのかな、ゝ・ゞ(ヽ・ヾ) +- ヴ、ヴャ・ヴュ・ヴョ、ひらがなの ゔ、半角の ヴ、ウ・う と結合用濁点(U+3099)の組み合わせ(例: ラヴ → ラブ、レヴュー → レビユー) + +また、かなの直後に書かれた ‐ ‑ ‒ – — ― ⁻ ₋ − - は、長音記号「ー」として扱われます( ``prolonged_sound_mark_filter`` )。このため「サ―バ-」や「サ−バ‐」は「サーバー」と一致します。半角のハイフン( ``-`` )や、漢字・英数字の後のダッシュ(東京-大阪、2026−10−02)は変換されません。 + +日本語用の ``ja/mapping.txt`` ( ``*_ja`` フィールド)は、形態素解析のためにひらがなと小書きのかなをそのまま残し、ヴ などの表記だけをそろえます。 + +.. note:: + + これらの設定は、新しく作成された文書インデックスに適用されます。既存のインデックスは、再インデクシングするまで以前の解析設定と辞書を使います。既存のインデックスに適用するには、 :doc:`maintenance-guide` で「辞書の初期化」を有効にして再インデクシングしてください。辞書の初期化は、管理画面で編集した ``mapping.txt`` / ``ja/mapping.txt`` の内容を上書きします。 .. |image0| image:: ../../../resources/images/ja/15.9/admin/mapping-1.png .. |image1| image:: ../../../resources/images/ja/15.9/admin/mapping-2.png diff --git a/ja/15.9/admin/relatedquery-guide.rst b/ja/15.9/admin/relatedquery-guide.rst index a3294135..2e371ff4 100644 --- a/ja/15.9/admin/relatedquery-guide.rst +++ b/ja/15.9/admin/relatedquery-guide.rst @@ -54,6 +54,61 @@ 一覧ページの設定名をクリックし、削除ボタンをクリックすると確認画面が表示されます。 削除ボタンを押すと設定が削除されます。 +検索ログから生成 +---------------- + +一覧ページの [検索ログから生成] ボタンをクリックすると、最近の検索ログから関連クエリーを作成します。同じ利用者のセッションで、ある検索の後に短い間隔で続けて行われた検索(誤字の後の正しい語や、広い語の後のより具体的な語など)を言い換えとみなし、よく検索される語について、最も多い言い換えをその語の関連クエリーにします。 + +関連クエリーは誰にでも適用され、その語の検索をすべて広げるため、生成は控えめに行われます。 + +- ゲストに見える検索だけを使います。検索ログのロールがすべて ``suggest.search.log.permissions`` (サジェストと同じ設定)を満たす場合だけ読み込みます。 +- ``label:"x"`` などのフィールド指定、演算子、ワイルドカード、 ``sort:`` 、先頭の ``+`` / ``-`` を含む検索語は使いません。 +- [サジェスト > 除外ワード] に登録した語は、語にも関連クエリーにも使いません。 +- 語とその関連クエリーは、それぞれ ``related_query.generate.min.sessions`` 以上のセッションから得られたものに限ります。言い換え後の検索はヒットがあるものに限ります。 +- 仮想ホストごとに作成します。仮想ホストのない検索ログは既定のホストとして扱います。 +- 既に関連クエリーがある語は変更しません(スキップした件数が結果に表示されます)。また、関連クエリーのキャッシュに読み込める件数( ``page.relatedquery.max.fetch.size`` )を超えて作成しません。 + +作成された関連クエリーは、手で登録したものと同じように編集、削除できます。[システム > 全般] の「検索ログ」または「ユーザーログ」が無効な場合は生成できません。生成の実行中に、もう一度実行することはできません。 + +生成の動作は ``fess_config.properties`` の次の設定で調整できます。 + +.. list-table:: + :header-rows: 1 + :widths: 45 40 15 + + * - プロパティ + - 説明 + - デフォルト + * - ``related_query.generate.days`` + - 読み込む検索ログの日数 + - ``30`` + * - ``related_query.generate.term.size`` + - 仮想ホストごとの語の最大数 + - ``100`` + * - ``related_query.generate.query.size`` + - 1 つの語に作る関連クエリーの最大数 + - ``5`` + * - ``related_query.generate.min.sessions`` + - 語と関連クエリーが現れる必要のある最小のセッション数 + - ``3`` + * - ``related_query.generate.session.interval`` + - 言い換えとみなす間隔(分) + - ``10`` + * - ``related_query.generate.seed.log.size`` + - 1 つの語について読み込む検索ログの最大数 + - ``1000`` + * - ``related_query.generate.seed.session.size`` + - 1 つの語について読み込むセッションの最大数 + - ``200`` + * - ``related_query.generate.log.fetch.size`` + - 1 つの語について読み込む後続の検索ログの最大数 + - ``2000`` + * - ``related_query.generate.query.min.length`` + - 語と関連クエリーの最小の長さ(文字数) + - ``2`` + * - ``related_query.generate.query.max.length`` + - 語と関連クエリーの最大の長さ(文字数) + - ``50`` .. |image0| image:: ../../../resources/images/ja/15.9/admin/relatedquery-1.png .. |image1| image:: ../../../resources/images/ja/15.9/admin/relatedquery-2.png diff --git a/ja/15.9/admin/searchlog-guide.rst b/ja/15.9/admin/searchlog-guide.rst index 92350ff9..80d1eafd 100644 --- a/ja/15.9/admin/searchlog-guide.rst +++ b/ja/15.9/admin/searchlog-guide.rst @@ -1,27 +1,62 @@ -====== +======== 検索ログ -====== +======== 概要 ==== -検索、クリック、お気に入りの実行結果は記録され、検索ログはこの管理画面で確認することができます。 +検索、クリック、お気に入りの実行結果は記録されます。検索ログの画面では、それらを集計した分析レポートと、個々のログの一覧を確認できます。 -管理方法 -====== +画面を開くには、左メニューの [システム情報 > 検索ログ] をクリックします。最初に「概要」タブが表示されます。閲覧には ``admin-searchlog`` または ``admin-searchlog-view`` のロールが必要です。 ``admin-searchlog-view`` ではログを削除できません。 -一覧 -==== +分析レポート +============ + +期間と絞り込み +-------------- + +各タブの上部で、集計する期間を「今日」「昨日」「過去7日間」「過去28日間」「過去90日間」から選ぶか、「カスタム」で開始日と終了日(最大 366 日)を指定します。「前の期間と比較」を有効にすると、同じ日数だけ前の期間の値と比較します。アクセス種別と、表の行数も指定できます。期間とグラフの区切りは、サーバーのタイムゾーンの日付に従います。 + +タブ +---- + +- **概要**: 検索数、利用者数、ゼロヒット率、クリック率、平均応答時間の指標を、推移の小さなグラフと前の期間からの変化とともに表示します。指標を切り替えられる推移のグラフ(比較時は前の期間を破線で表示)と、人気の検索語、ゼロヒット語も表示します。 +- **検索語**: 検索語ごとの検索数、利用者数、平均ヒット数、クリック数、クリック率、平均クリック順位を表示します。ゼロヒット語(最終検索日時付き)と、ヒットしたのに結果が一度もクリックされなかったゼロクリック語も表示します。 +- **クリック**: クリック数の多い URL、お気に入り数の多い URL、クリック順位の分布、2 ページ目以降の閲覧率を表示します。 +- **パフォーマンス**: 応答時間の平均と中央値 (p50)・p95・p99、応答時間の分布、応答の遅い検索語、クエリ時間を表示します。 +- **利用者・環境**: 新規利用者と既存利用者、アクセス種別、曜日・時間帯別の検索数、上位の User-Agent、リファラ、言語、仮想ホストを表示します。「ロール・グループ別の検索数」は、ロールやグループごとの検索数、利用者数、ゼロヒット率を表示します。検索は実行した利用者のすべてのロールとグループに計上されるため、合計が全体の検索数より多くなることがあります。利用者個人は表示されません。 +- **AIチャット**: AI 検索モード(RAG チャット)のリクエスト数、利用者数、合計トークン数、平均応答時間、エラー率と、上位ユーザー、意図別・モデル別のリクエストを表示します。チャットの利用ログは ``rag.chat.log.enabled`` (デフォルト: ``true`` )が有効な場合に記録されます。質問と回答の内容は記録されません。トークン数は LLM プラグインが報告した場合にのみ記録されます。 +- **ログ**: 個々のログの一覧です。下の「ログの一覧」を参照してください。 + +.. note:: + + 検索語ごとのクリックの指標とゼロクリック語は、検索語が記録されるようになった後( |Fess| 15.9 以降)のクリックだけを集計します。全体のクリック数とクリック率には、それ以前のクリックも含まれます。利用者数などの一部の値は概算です。 + +ゼロヒット語の確認 +------------------ -一覧では検索、クリック、お気に入りの検索ログを確認することができます。 -検索ログの詳細を確認したい場合は対象の検索ログをクリックします。 +「概要」と「検索語」のタブで、ゼロヒット語をクリックすると、「ログ」タブにヒット件数「0件のみ」で絞り込んだその語の検索ログが表示されます。どのような検索で結果が見つからなかったかを確認し、文書や同義語、関連クエリーの追加に役立てられます。 + +CSV のダウンロード +------------------ + +分析レポートの各表とグラフには CSV へのリンクがあり、表示中の期間、比較、アクセス種別、行数の条件で集計した内容をダウンロードできます。絞り込みの欄には、指標をまとめた CSV へのリンクもあります。数値は加工せずに出力されます(比率は 0〜1、時間はミリ秒)。比較時のグラフには ``<系列>_previous`` の列が加わります。 + +ログの一覧 +========== + +「ログ」タブでは、検索ログ、クリックログ、お気に入りログ、ユーザーログを確認できます。ログの種別、クエリーID、ユーザーID、時刻の範囲、アクセスタイプ、検索語で絞り込めます。検索ログでは、ヒット件数(「すべて」「0件のみ」「1件以上」)でも絞り込めます。ログの詳細を確認したい場合は、対象のログをクリックします。 |image0| +[CSVをダウンロード] ボタンをクリックすると、現在の条件に一致するログを、件数の上限なしに新しい順で CSV としてダウンロードできます。見出し行はフィールド名で、画面の言語によって変わりません。 + +CSV ファイルは、分析レポートのものも含めて ``csv.file.encoding`` の文字コードで出力され、UTF-8 の場合は BOM が付きます。 ``=`` 、 ``+`` 、 ``-`` 、 ``@`` 、タブ、復帰文字(CR)で始まる値には、表計算ソフトで数式として扱われないように先頭に ``'`` が付きます。 + 詳細 -==== +---- -一覧から検索ログをクリックすることで、対象の検索ログ詳細が表示されます。 +一覧からログをクリックすると、対象のログの詳細が表示されます。 |image1| diff --git a/ja/15.9/api/admin/api-admin-backup.rst b/ja/15.9/api/admin/api-admin-backup.rst index b656685a..f665c3be 100644 --- a/ja/15.9/api/admin/api-admin-backup.rst +++ b/ja/15.9/api/admin/api-admin-backup.rst @@ -79,12 +79,13 @@ Backup APIは、|Fess| のバックアップ対象データを参照・ダウン { "id": "system.properties", "name": "system.properties" }, { "id": "fess.json", "name": "fess.json" }, { "id": "doc.json", "name": "doc.json" }, + { "id": "chat_log.ndjson", "name": "chat_log.ndjson" }, { "id": "click_log.ndjson", "name": "click_log.ndjson" }, { "id": "favorite_log.ndjson", "name": "favorite_log.ndjson" }, { "id": "search_log.ndjson", "name": "search_log.ndjson" }, { "id": "user_info.ndjson", "name": "user_info.ndjson" } ], - "total": 10 + "total": 11 } } @@ -112,7 +113,7 @@ Backup APIは、|Fess| のバックアップ対象データを参照・ダウン - インデックスのマッピング定義ファイル(``fess_indices/fess.json`` / ``fess_indices/fess/doc.json``)そのもの(``application/octet-stream``) * - ``*.bulk`` または拡張子なしのインデックス名 - 対象名と同名のインデックスをスクロールして生成したバルクデータ(``application/octet-stream``)。\ ``.bulk`` を取り除いた名前をインデックス名として扱います。 - * - ``*.ndjson`` (``search_log`` / ``user_info`` / ``click_log`` / ``favorite_log``) + * - ``*.ndjson`` (``search_log`` / ``user_info`` / ``click_log`` / ``favorite_log`` / ``chat_log``) - 対応するログのNDJSONデータ(``application/x-ndjson``) .. note:: diff --git a/ja/15.9/api/api-export.rst b/ja/15.9/api/api-export.rst new file mode 100644 index 00000000..5421c373 --- /dev/null +++ b/ja/15.9/api/api-export.rst @@ -0,0 +1,78 @@ +======================= +検索結果エクスポートAPI +======================= + +このドキュメントでは、検索結果を CSV または JSON のファイルとしてダウンロードする |Fess| の v2 エクスポート API について説明します。共通のレスポンスエンベロープ・エラーモデルについては :doc:`api-overview` を参照してください。 + +ベースURLは ``http:///api/v2/`` です(ローカル環境の例: ``http://localhost:8080/api/v2`` )。 + +.. note:: + + エクスポートはデフォルトで無効です。利用するには ``fess_config.properties`` で ``api.search.export=true`` を設定してください。有効な場合、同梱の ``bootstrap`` テーマでは検索結果の件数の横にエクスポートのメニュー(CSV / JSON)が表示されます。 ``/api/v2/ui/config`` の ``features.search_export`` で状態を確認できます。 + +検索結果のダウンロード +====================== + +リクエスト +---------- + +================== ==================================================== +HTTPメソッド GET +エンドポイント ``/api/v2/documents/export`` +================== ==================================================== + +検索に一致するドキュメントを、ファイルのダウンロード( ``Content-Disposition: attachment`` 、ファイル名は ``search_results.csv`` または ``search_results.json`` )として返します。 + +- ``/api/v2/search`` と同じロールの絞り込みが適用されます。 ``login.required=true`` の場合も、 ``/api/v2/search`` と同じようにアクセストークンを使えます。 +- 出力する件数は最大 ``api.search.export.max.size`` (デフォルト: ``1000`` )件です。ページングのパラメーター( ``start`` 、 ``num`` )は使いません。 +- 出力するフィールドは ``api.search.export.fields`` (デフォルト: ``title,url_link,last_modified,content_length,filetype`` )のうち、API のレスポンスに含められるフィールドだけです。 +- リクエストは 1 分あたり ``api.search.export.rate.limit.per.minute`` (デフォルト: ``10`` 、 ``0`` で無制限)回に制限されます。ログインしているユーザーはユーザーごと、ゲストはクライアントの IP アドレスごとに数えます。超えると ``429`` と ``Retry-After`` ヘッダーが返ります。 +- エクスポートは検索ログに記録されません。 + +リクエストパラメーター +---------------------- + +``q`` 、 ``ex_q`` 、 ``fields.*`` 、 ``sort`` 、 ``lang`` など、 ``/api/v2/documents/all`` と同じ検索条件のパラメーターを指定できます( :doc:`api-search` を参照)。これに加えて次のパラメーターがあります。 + +.. tabularcolumns:: |p{4cm}|p{11cm}| +.. list-table:: リクエストパラメーター + + * - ``format`` + - ファイルの形式。 ``csv`` (デフォルト)または ``json`` 。それ以外は ``invalid_request`` (400) になります。 + +表: リクエストパラメーター + +レスポンス +---------- + +CSV は 1 行目がフィールド名の見出し行で、 ``csv.file.encoding`` の文字コードで出力されます(UTF-8 の場合は BOM が付きます)。 ``=`` 、 ``+`` 、 ``-`` 、 ``@`` 、タブ、復帰文字(CR)で始まる値には、表計算ソフトで数式として扱われないように先頭に ``'`` が付きます。複数の値を持つフィールドは空白でつなげます。 + +:: + + "title","url_link","last_modified","content_length","filetype" + "Example","https://example.com/","2025-01-01T00:00:00.000Z","1234","html" + +JSON は ``{"data":[{...},...]}`` の形式で、複数の値を持つフィールドは配列のままです。 + +ファイルの出力が始まる前に失敗した場合は、通常のエラーエンベロープが返ります。出力が始まった後に失敗した場合は、ファイルが途中で終わります(CSV は途中まで、JSON は解析できない内容になります)。 + +エラーレスポンス +---------------- + +.. tabularcolumns:: |p{4cm}|p{11cm}| +.. list-table:: エラーレスポンス + + * - ステータスコード + - 説明 + * - 400 Bad Request + - 不正なクエリ、 ``format`` が ``csv`` / ``json`` 以外、または ``api.search.export=false`` でエクスポートが無効な場合。 + * - 401 Unauthorized + - 認証が必要な場合( ``login.required=true`` で匿名の呼び出し元など)。 + * - 405 Method Not Allowed + - HTTP メソッドが許可されていない場合。 + * - 429 Too Many Requests + - 1 分あたりのリクエスト数の上限を超えた場合。 + * - 500 Internal Server Error + - サーバー内部エラーが発生した場合。 + +表: エラーレスポンス diff --git a/ja/15.9/api/api-search-history.rst b/ja/15.9/api/api-search-history.rst new file mode 100644 index 00000000..2e04229d --- /dev/null +++ b/ja/15.9/api/api-search-history.rst @@ -0,0 +1,103 @@ +=========== +検索履歴API +=========== + +このドキュメントでは、 |Fess| の v2 検索履歴 API について説明します。共通のレスポンスエンベロープ・エラーモデルについては :doc:`api-overview` を参照してください。 + +ベースURLは ``http:///api/v2/`` です(ローカル環境の例: ``http://localhost:8080/api/v2`` )。 + +.. note:: + + 検索履歴は、 ``search.history.enabled`` (デフォルト: ``true`` )と検索ログの記録の両方が有効な場合に利用できます。 ``/api/v2/ui/config`` の ``features.search_history`` で状態を確認できます。 + +最近の検索の取得 +================ + +リクエスト +---------- + +================== ==================================================== +HTTPメソッド GET +エンドポイント ``/api/v2/search-history`` +================== ==================================================== + +ログインしているユーザーが、現在の仮想ホストで ``/api/v2/search`` を使って行った最近の検索を、新しい順に返します。返った条件で同じ検索を実行し直せます。 + +- 1 ページ目の検索だけが対象です。同じ条件の検索は、最も新しい 1 件にまとめられます。検索語のない検索は含まれません。 +- 返す件数は最大 ``search.history.size`` (デフォルト: ``10`` )件です。 +- 検索ログは 1 分ごとのジョブで書き込まれるため、検索が履歴に現れるまで最大 1 分ほどかかります。 +- 履歴はセッションのログインユーザーに紐づきます。匿名の呼び出し元は ``auth_required`` (401) になり、アクセストークンはログインの代わりになりません。 +- 検索履歴の機能が無効な場合は ``invalid_request`` (400) になります。 + +リクエストパラメーターはありません。 + +レスポンス +---------- + +成功時(200)は、共通エンベロープの ``response`` 直下に以下のフィールドが返ります。 + +:: + + { + "response": { + "status": 0, + "record_count": 1, + "data": [ + { + "q": "fess", + "fields": { "label": ["docs"] }, + "sort": "last_modified.desc", + "requested_at": "2026-10-01T09:00:00Z", + "hit_count": 42 + } + ] + } + } + +.. tabularcolumns:: |p{4cm}|p{11cm}| +.. list-table:: レスポンスフィールド + + * - ``record_count`` + - ``data`` 内の検索の件数(int)。 + * - ``data`` + - 最近の検索の配列(新しい順)。条件のキーは ``/api/v2/search`` のリクエストパラメーター名で、その検索で使われなかったキーは省略されます。 + * - ``data[].q`` + - 検索語(str)。 + * - ``data[].fields`` + - ``fields.`` で指定したフィールドの条件(フィールド名をキーとする値の配列)。 + * - ``data[].ex_q`` + - 追加のクエリー(str の配列)。 + * - ``data[].sort`` + - ソート順(str)。 + * - ``data[].lang`` + - ``lang`` で指定した言語(str の配列)。 + * - ``data[].requested_at`` + - 検索した日時(UTC、ISO-8601)。 + * - ``data[].hit_count`` + - その検索のヒット件数(int64)。 + +表: レスポンスフィールド + +エラーレスポンス +---------------- + +.. tabularcolumns:: |p{4cm}|p{11cm}| +.. list-table:: エラーレスポンス + + * - ステータスコード + - 説明 + * - 400 Bad Request + - 検索履歴の機能が無効な場合。 + * - 401 Unauthorized + - ログインしていない場合。 + * - 405 Method Not Allowed + - HTTP メソッドが許可されていない場合。 + * - 500 Internal Server Error + - サーバー内部エラーが発生した場合。 + +表: エラーレスポンス + +同梱のテーマでの表示 +==================== + +同梱の ``bootstrap`` テーマでは、ログインしたユーザーが空の検索ボックスをクリックするか、検索ボックスで下矢印キーを押すと、サジェストのドロップダウンに最近の検索が表示されます。項目を選ぶと、ラベルなどの条件も含めて同じ検索を実行します。 |Fess| 15.9 より前に記録された検索ログには条件が保存されていないため、履歴には表示されません。 diff --git a/ja/15.9/api/api-tag.rst b/ja/15.9/api/api-tag.rst new file mode 100644 index 00000000..caab7d62 --- /dev/null +++ b/ja/15.9/api/api-tag.rst @@ -0,0 +1,129 @@ +======= +タグAPI +======= + +このドキュメントでは、ドキュメントにタグを付ける |Fess| の v2 タグ API について説明します。共通のレスポンスエンベロープ・エラーモデル・CSRF については :doc:`api-overview` を参照してください。 + +ベースURLは ``http:///api/v2/`` です(ローカル環境の例: ``http://localhost:8080/api/v2`` )。 + +.. note:: + + タグ機能はデフォルトで無効です。利用するには ``fess_config.properties`` で ``user.tag.enabled=true`` を設定してください。 ``/api/v2/ui/config`` の ``features.user_tag`` で状態を確認できます。 + +タグは、種類が「タグ」のラベルです( :doc:`../admin/labeltype-guide` を参照)。ラベルの名前がタグ名、値がタグ名の SHA-256(16 進数)、対象とするパスがタグを付けた URL、パーミッションがタグを見られるユーザーです。タグが見えるのは、そのラベルが呼び出し元に見える場合だけです。 + +検索 API( ``/api/v2/search`` )は、各ヒットに呼び出し元に見えるタグを ``tags`` として返します。 ``fields.tag=<値>`` でタグの付いたドキュメントに絞り込め、 ``facet.field=tag`` でタグのファセットを取得できます。インデックスの ``tag`` フィールドそのものはレスポンスに含まれません。 + +タグの取得 +========== + +リクエスト +---------- + +================== ==================================================== +HTTPメソッド GET +エンドポイント ``/api/v2/documents/{docId}/tags`` +================== ==================================================== + +指定したドキュメントのタグのうち、呼び出し元に見えるものを返します。呼び出し元がそのドキュメントを検索できない場合は ``not_found`` (404) になります。 + +レスポンス +---------- + +成功時(200)は、共通エンベロープの ``response`` 直下に以下のフィールドが返ります。 + +:: + + { + "response": { + "status": 0, + "doc_id": "a1b2c3d4e5f6", + "addable": true, + "tags": [ + { "value": "9f86d081884c7d65...", "name": "要確認", "mine": true } + ] + } + } + +.. tabularcolumns:: |p{4cm}|p{11cm}| +.. list-table:: レスポンスフィールド + + * - ``doc_id`` + - ドキュメント ID(str)。 + * - ``addable`` + - 呼び出し元がログインしている(タグを追加できる)場合 ``true`` (bool)。 + * - ``added`` + - POST のみ。呼び出し元が既にそのドキュメントにタグを付けていた場合は ``false`` (bool)。 + * - ``removed`` + - DELETE のみ(bool)。 + * - ``tags`` + - 呼び出し元に見えるタグの配列。各要素は ``value`` (ラベルの値。 ``fields.tag`` に指定する値)、 ``name`` (タグ名)、 ``mine`` (呼び出し元がタグのパーミッションに含まれる場合 ``true`` )。 + +表: レスポンスフィールド + +タグの追加 +========== + +リクエスト +---------- + +================== ==================================================== +HTTPメソッド POST +エンドポイント ``/api/v2/documents/{docId}/tags`` +================== ==================================================== + +ログインしているユーザーとして、ドキュメントの URL にタグを付けます。アクセストークンはログインの代わりになりません。状態を変更するリクエストのため、 ``X-Fess-CSRF-Token`` ヘッダーが必要です。 + +- 同じ名前のタグがあれば、その対象とするパスに URL を、パーミッションにユーザーを追加します。なければ、そのユーザーだけが見られるタグを作成します。そのため、同じ名前のタグは 1 つにまとまり、同じ名前のタグを付けた利用者どうしは互いのタグの付与先を見られます。 +- 同じドキュメントにもう一度付けると ``added`` は ``false`` になります。 +- 1 つのドキュメントに付けられるタグは最大 ``user.tag.max.document.tags`` (デフォルト: ``100`` )個です。 + +リクエストボディは ``Content-Type: application/json`` で、 ``name`` にタグ名を指定します。 + +:: + + { + "name": "要確認" + } + +タグ名は NFKC で正規化され、連続する空白は 1 つにまとめられ、前後の空白は取り除かれます。長さは 1 〜 ``user.tag.name.max.length`` (デフォルト: ``50`` )文字で、制御文字や書式文字(ゼロ幅文字や双方向の上書きなど)を含む名前は拒否されます。 + +タグの削除 +========== + +リクエスト +---------- + +================== ==================================================== +HTTPメソッド DELETE +エンドポイント ``/api/v2/documents/{docId}/tags?value=<タグの値>`` +================== ==================================================== + +ログインしているユーザーを、 ``value`` で指定したタグのパーミッションから外します。ユーザー、グループ、ロールのパーミッションが 1 つも残らなくなると、タグは削除され、ドキュメントからも取り除かれます。ユーザーがタグのパーミッションに含まれていない場合は ``forbidden`` (403) になります。 ``X-Fess-CSRF-Token`` ヘッダーが必要です。 + +エラーレスポンス +================ + +.. tabularcolumns:: |p{4cm}|p{11cm}| +.. list-table:: エラーレスポンス + + * - ステータスコード + - 説明 + * - 400 Bad Request + - リクエストが不正な場合(タグ機能が無効な場合、タグ名が不正な場合、タグの上限を超えた場合を含む)。 + * - 401 Unauthorized + - POST・DELETE でログインしていない場合。 + * - 403 Forbidden + - CSRF トークンの欠落・失効、または自分が追加していないタグを DELETE した場合。 + * - 404 Not Found + - ドキュメントが見つからない、または呼び出し元が検索できない場合。 + * - 405 Method Not Allowed + - HTTP メソッドが許可されていない場合。 + * - 413 Payload Too Large + - リクエストボディがサイズ上限を超えている場合。 + * - 415 Unsupported Media Type + - サポートされていない ``Content-Type`` の場合。 + * - 500 Internal Server Error + - サーバー内部エラーが発生した場合。 + +表: エラーレスポンス diff --git a/ja/15.9/api/api-uiconfig.rst b/ja/15.9/api/api-uiconfig.rst index 156dc7de..fdd7203d 100644 --- a/ja/15.9/api/api-uiconfig.rst +++ b/ja/15.9/api/api-uiconfig.rst @@ -44,6 +44,9 @@ SPA が必要とする初期設定を返します。 }, "features": { "user_favorite": false, + "search_history": true, + "search_export": false, + "user_tag": false, "popular_word": true, "suggest_search_log": true, "suggest_documents": true, @@ -186,6 +189,15 @@ features * - ``user_favorite`` - boolean - ユーザーお気に入り機能が有効かどうか。 + * - ``search_history`` + - boolean + - 検索履歴( ``GET /api/v2/search-history`` )が利用できるかどうか( ``search.history.enabled`` と検索ログの両方が有効な場合に ``true`` )。 + * - ``search_export`` + - boolean + - 検索結果のエクスポート( ``GET /api/v2/documents/export`` )が有効かどうか( ``api.search.export`` )。 + * - ``user_tag`` + - boolean + - タグ機能( ``/api/v2/documents/{docId}/tags`` )が有効かどうか( ``user.tag.enabled`` )。 * - ``popular_word`` - boolean - 人気ワード機能が有効かどうか。 diff --git a/ja/15.9/api/index.rst b/ja/15.9/api/index.rst index c2dc99b7..2a26a1dd 100644 --- a/ja/15.9/api/index.rst +++ b/ja/15.9/api/index.rst @@ -22,6 +22,7 @@ :caption: 検索API api-search + api-export api-label api-popularword api-suggest @@ -40,6 +41,8 @@ :caption: ユーザー機能API api-favorite + api-search-history + api-tag api-click api-cache diff --git a/ja/15.9/config/crawler-ocr.rst b/ja/15.9/config/crawler-ocr.rst index 07f26f5a..7d02e9fd 100644 --- a/ja/15.9/config/crawler-ocr.rst +++ b/ja/15.9/config/crawler-ocr.rst @@ -167,6 +167,7 @@ Docker 環境で OCR を使うには、|Fess| のイメージに Tesseract を - Web クロール設定では、新規作成時のデフォルトで画像の URL(jpg、png、gif など)がクロール対象から除外されます。Web サイトの画像をクロールするには、「クロール対象から除外するURL」からこれらを削除してください。ファイルクロールでは、画像もクロール対象になります。 - クローラーの取得サイズの上限も適用されます。ファイルの種類ごとのインデックスサイズの上限(デフォルトは 10MB)については :doc:`crawler-basic` を参照してください。 - OCR の精度は、スキャンした画像の品質に左右されます。手書きの文字は、一般に認識されにくくなります。 +- OCR を有効にする前に索引に登録済みのファイルには、そのままでは OCR が適用されません。差分クロール(:doc:`../admin/general-guide` の「最終更新日時の確認」)が有効な場合、ファイルの更新日時が変わっていなければ再クロールでも取得し直されないためです。既存のファイルに OCR を適用するには、「最終更新日時の確認」を一時的に無効にしてクロールするか、対象のドキュメントを索引から削除してからクロールし直してください。 アップグレード時の注意 ====================== diff --git a/ja/15.9/config/rate-limiting.rst b/ja/15.9/config/rate-limiting.rst index 55b6702b..e8eadfda 100644 --- a/ja/15.9/config/rate-limiting.rst +++ b/ja/15.9/config/rate-limiting.rst @@ -102,6 +102,26 @@ robots.txtの処理は ``app/WEB-INF/classes/fess_config.properties`` の # robots.txtを無視する(デフォルト: false) crawler.ignore.robots.txt=false +Crawl-delay は接続先(オリジン)ごとに適用され、上限は 60 秒です。待機は URL 単位で行われるため、差分クロールで 1 つの URL に送る HEAD と GET は続けて送信されます。クロール設定の「間隔」は、これとは別に次の URL までの待機時間として働きます。 + +robots.txt は RFC 9309 に従って解釈されます。 ``Allow:`` と ``Disallow:`` は最も長く一致した規則が優先され、開始 URL も robots.txt で確認されます。robots.txt で拒否された URL は ``fess-crawler.log`` に INFO で記録され、障害 URL には登録されません。 + +429/503 応答時のバックオフ +-------------------------- + +サーバーが ``429 Too Many Requests`` または ``503 Service Unavailable`` を返した場合、そのオリジンへのアクセスを一時的に止め、その URL を最大 3 回まで再試行します。待機時間は、 ``Retry-After`` ヘッダーがあればその値、なければ 10 秒から始まる指数関数的な時間(最大 5 分)です。最後の再試行にも失敗すると、 ``fess-crawler.log`` に WARN が出力されます。 + +robots.txt を取得できない場合 +----------------------------- + +robots.txt の取得が 5xx、429、タイムアウトなどで失敗した場合、そのオリジンの URL はバックオフが終わるまでキューに戻されます。最初の失敗の後、3 回の再試行にも失敗すると、そのクロールの間はそのオリジンのすべての URL がクロール対象外になり、 ``fess-crawler.log`` に WARN が 1 回出力されます。15.8 以前は、取得できなかった robots.txt はすべて許可として扱われていました。この動作に戻すには、ウェブクロール設定の「設定パラメーター」に次を指定します。 + +:: + + client.robotsTxtAllowOnUnavailable=true + +robots.txt の処理そのものを無効にするには、 ``client.robotsTxtEnabled=false`` (クロール設定ごと)または ``crawler.ignore.robots.txt=true`` を使います。 + レート制限の全設定項目 ====================== diff --git a/ja/15.9/dev/theme-development.rst b/ja/15.9/dev/theme-development.rst index 1142f229..7ba56537 100644 --- a/ja/15.9/dev/theme-development.rst +++ b/ja/15.9/dev/theme-development.rst @@ -154,6 +154,13 @@ ``Content-Security-Policy`` ヘッダーが付きます(インラインのスタイルは許可され、 インラインのスクリプトは許可されません)。そのため外部の CDN のフォントやスクリプトは 読み込まれません。テーマに同梱してください。 +- エントリー HTML の ``Content-Security-Policy`` には ``frame-ancestors 'none'`` が含まれ、 + ``X-Frame-Options: DENY`` ヘッダーも付くため、ページは他のページのフレームに表示されません。 + ``frame-ancestors`` の値は ``fess_config.properties`` の ``theme.index.frame.ancestors`` + (デフォルト: ``'none'`` )で変更できます。空にすると ``frame-ancestors`` を付けません + ( ``X-Frame-Options: DENY`` は付いたままです)。WebKit 系のブラウザー(Safari など)は、 + ``frame-ancestors 'none'`` のもとではテーマが ``blob:`` URL で表示するフレーム + (PDF のプレビューやキャッシュの表示)を空にします。これらを表示するには値を空にします。 パッケージング -------------- diff --git a/ja/15.9/install/fess-setup.rst b/ja/15.9/install/fess-setup.rst index 2f5eee6a..7add824c 100644 --- a/ja/15.9/install/fess-setup.rst +++ b/ja/15.9/install/fess-setup.rst @@ -94,6 +94,14 @@ Playwright クローラが必要とする Node.js を、 |Fess| のディレク ``install plugin`` 、 ``list plugins`` 、 ``upgrade plugins`` には ``--repository `` を指定できます。指定すると、バージョンの一覧、jar、チェックサムをすべてその 1 つの Maven リポジトリ(社内のミラーなど)から取得し、既定のリリースリポジトリ、スナップショットリポジトリ、GitHub は使用しません。 +``--repository`` には ``file:///`` で始まる URL も指定できます。インターネットに接続できないサーバーでは、Maven リポジトリのコピーをサーバー上に置き、そのディレクトリを指定します。指定するのは、各プラグインのディレクトリ( ``/maven-metadata.xml`` など)を含むディレクトリで、既定のリポジトリの ``https://maven.codelibs.org/release/org/codelibs/fess/`` に当たる場所です。バージョンの解決とチェックサムの検証は、HTTP のリポジトリと同じように行われます。 + +:: + + $ bin/fess-setup install plugin fess-ds-csv --repository file:///opt/maven-repo/org/codelibs/fess/ + +``http`` 、 ``https`` 、 ``file`` 以外のスキームや、スキームのない URL は 1 行のエラーで拒否されます。 + install plugin -------------- @@ -107,6 +115,8 @@ jar はプラグインの GitHub リリースから取得し、リリースに 使用例は :doc:`../admin/plugin-guide` を参照してください。 +導入できるのは、 |Fess| が読み込む種類のプラグイン(名前が ``fess-ds-`` 、 ``fess-ingest-`` 、 ``fess-script-`` 、 ``fess-webapp-`` 、 ``fess-thumbnail-`` 、 ``fess-crawler-`` 、 ``fess-llm-`` 、 ``fess-storage-`` 、 ``fess-sso-`` で始まるもの)だけです。 ``list plugins`` に表示されるものと同じです。それ以外の名前( ``fess-theme-*`` 、 ``fess`` 、 ``fess-crawler`` などのライブラリ)を指定すると、何もダウンロードせずに 1 行のエラーを表示し、終了コード 2 で終了します。複数の名前のうち 1 つでも該当すると、どれも導入しません。静的テーマは ``install theme`` で導入してください。 + list plugins ------------ @@ -156,6 +166,8 @@ remove plugin 2. 対象の環境で ZIP を展開し、jar を ``app/WEB-INF/plugin/`` に置きます。 |Fess| の起動後に、管理画面のプラグインのインストール画面の「ローカル」タブからアップロードすることもできます( :doc:`../admin/plugin-guide` を参照)。 3. |Fess| を起動します。起動中に jar を置いた場合は再起動してください。 +プラグインの jar を個別に持ち込む代わりに、Maven リポジトリの必要な部分( ``org/codelibs/fess//`` のディレクトリ)をコピーして持ち込み、 ``--repository file:///...`` を指定して ``install plugin`` で導入することもできます。この場合も、チェックサムが検証されます。 + テーマの管理 ============ diff --git a/ja/15.9/user/search-field.rst b/ja/15.9/user/search-field.rst index 72f9b2ca..8fe0edc1 100644 --- a/ja/15.9/user/search-field.rst +++ b/ja/15.9/user/search-field.rst @@ -68,6 +68,12 @@ * - favorite_count - ドキュメントがお気に入り登録された回数 - 数値 + * - owner + - ファイルの所有者のアカウント名 + - キーワード + * - last_modifier + - ファイルの最終更新者 + - キーワード 表: 利用可能なフィールド一覧 @@ -82,6 +88,8 @@ .. note:: クロール対象によっては値が登録されないフィールドもあります。たとえば anchor は Web クロール時のみ、lang は HTML に言語属性がある場合のみ登録されます。また、segment(クロール実行単位を表すセッション ID)や doc_id(システムが採番する内部 ID)などのフィールドも指定できますが、通常の検索では利用しません。 +owner と last_modifier は、ファイルサーバーなどのクロールで登録されます。owner には SMB、ファイルシステム、FTP のクロールで取得したファイルの所有者のアカウント名が入ります( ``DOMAIN\alice`` のような値は ``alice`` になります)。last_modifier には Office 文書などから抽出した最終更新者が入り、取得できない場合は所有者が入ります。Web クロールした HTML には owner は登録されません。たとえば ``owner:alice`` や ``last_modifier:"Taro Yamada"`` のように検索します。同梱のテーマでは、詳細検索で所有者と最終更新者を指定できます。登録するかどうかは、管理者が ``crawler.document.file.owner.enabled`` と ``crawler.document.file.last.modifier.enabled`` (どちらもデフォルトは ``true`` )で切り替えられます。 + HTML ファイルを検索対象としている場合、title タグが title フィールドに、body タグ以下の文字列が content フィールドに登録されます。 利用方法 diff --git a/ko/15.9/admin/backup-guide.rst b/ko/15.9/admin/backup-guide.rst index f60b5f20..d4748651 100644 --- a/ko/15.9/admin/backup-guide.rst +++ b/ko/15.9/admin/backup-guide.rst @@ -45,6 +45,11 @@ doc.json doc.json은 fess 인덱스의 매핑 정보를 포함합니다. +chat_log.ndjson +::::::::::::::: + +chat_log.ndjson은 AI 채팅 이용 로그 정보를 포함합니다. + click_log.ndjson :::::::::::::::: diff --git a/ko/15.9/admin/docreport-guide.rst b/ko/15.9/admin/docreport-guide.rst new file mode 100644 index 00000000..ed97dee6 --- /dev/null +++ b/ko/15.9/admin/docreport-guide.rst @@ -0,0 +1,58 @@ +=========== +문서 리포트 +=========== + +개요 +==== + +문서 리포트는 크롤링한 파일 서버 등을 정리하는 데 도움이 되는 관리 화면입니다. 내용이 같은 문서와 오랫동안 갱신되지 않은 문서를 목록으로 표시하며, 각각 CSV 로 다운로드할 수 있습니다. + +화면을 열려면 왼쪽 메뉴의 [시스템 정보 > 문서 리포트] 를 클릭합니다. 열람하려면 ``admin-docreport`` 또는 ``admin-docreport-view`` 역할이 필요합니다. 이 화면은 표시와 다운로드만 하며 문서를 변경하지 않습니다. + +두 탭 모두 「URL 접두사」(예: ``smb://server/share/`` )로 대상을 좁힐 수 있습니다. + +중복 문서 +========= + +내용이 같거나 거의 같은 문서를 그룹으로 묶어 문서 수가 많은 순으로 표시합니다. +그룹화에는 색인 시 계산되는 내용 서명( ``content_minhash_bits`` , 검색 결과의 중복 묶기와 같은 것)을 사용하므로 재인덱싱은 필요하지 않습니다. +내용에 단어가 없는 문서(빈 파일 등)는 대상에서 제외됩니다. + +화면에는 최대 ``docreport.duplicate.group.size`` (기본값: 100)개 그룹을 표시하고, 각 그룹의 문서는 ``docreport.duplicate.docs.size`` (기본값: 10)건까지 표시합니다. 모든 그룹은 [CSV 다운로드] 로 가져올 수 있습니다. CSV 는 큰 인덱스에서도 모든 그룹을 읽어 들이며 ``group, groupSize, url, title, filename, contentLength, lastModified, owner, lastModifier, clickCount, docId`` 열을 출력합니다. + +.. note:: + + 내용 서명을 보관하지 않는 인덱스( ``cloud`` , ``aws`` 용 매핑)에서는 중복 문서 리포트를 사용할 수 없습니다. + +휴면 문서 +========= + +최종 갱신 일시가 지정한 일수(「미갱신 일수」, 기본값은 ``docreport.dormant.days`` 의 365)보다 이전인 문서를 오래된 순으로 표시합니다. 최종 갱신 일시가 없는 문서는 표시하지 않습니다. +「검색 결과에서 한 번도 열리지 않음」을 활성화하면 검색 결과에서 클릭된 적이 있는 문서를 제외합니다. + +화면에는 해당 문서의 건수와 합계 크기, 문서 목록을 표시합니다. 목록의 페이지 이동에는 상한( ``indexer.max.result.window.size`` )이 있으며, 이를 넘는 문서는 [CSV 다운로드] 로 가져옵니다. + +설정 +==== + +``fess_config.properties`` 의 다음 설정으로 동작을 조정할 수 있습니다. + +.. list-table:: + :header-rows: 1 + :widths: 40 45 15 + + * - 속성 + - 설명 + - 기본값 + * - ``docreport.duplicate.group.size`` + - 화면에 표시하는 중복 그룹의 최대 수 + - ``100`` + * - ``docreport.duplicate.docs.size`` + - 화면에서 그룹 하나에 표시하는 문서의 최대 수 + - ``10`` + * - ``docreport.duplicate.export.page.size`` + - CSV 다운로드 시 한 번에 읽어 들이는 서명의 수 + - ``10000`` + * - ``docreport.dormant.days`` + - 휴면 문서로 간주하는, 최종 갱신 이후 일수의 기본값 + - ``365`` diff --git a/ko/15.9/admin/general-guide.rst b/ko/15.9/admin/general-guide.rst index be80decf..bbb1e5b8 100644 --- a/ko/15.9/admin/general-guide.rst +++ b/ko/15.9/admin/general-guide.rst @@ -111,6 +111,8 @@ SSO 유형 차분 크롤을 수행하는 경우 활성화합니다. +HTTP/HTTPS URL 에서는 HEAD 의 ``Last-Modified`` 로 판단할 수 없는 경우(색인에 최종 갱신 일시가 없거나, HEAD 응답에 ``Last-Modified`` 가 없거나, HEAD 가 200/404 이외를 반환한 경우), GET 에 ``If-None-Match`` (색인에 저장한 ``ETag`` )나 ``If-Modified-Since`` 를 붙인 조건부 GET 을 수행합니다. 서버가 ``304 Not Modified`` 를 반환한 페이지는 갱신되지 않은 페이지와 마찬가지로 다시 가져오지 않습니다. 설정을 변경한 후 등 모든 것을 다시 가져오려면 이 설정을 비활성화하고 크롤링하십시오. + 동시 크롤러 설정 :::::::::::::: diff --git a/ko/15.9/admin/index.rst b/ko/15.9/admin/index.rst index 3482e868..d68c9059 100644 --- a/ko/15.9/admin/index.rst +++ b/ko/15.9/admin/index.rst @@ -49,6 +49,7 @@ log-guide failureurl-guide searchlist-guide + docreport-guide backup-guide maintenance-guide esreq-guide diff --git a/ko/15.9/admin/labeltype-guide.rst b/ko/15.9/admin/labeltype-guide.rst index 8ab0a8cc..b5cbc82d 100644 --- a/ko/15.9/admin/labeltype-guide.rst +++ b/ko/15.9/admin/labeltype-guide.rst @@ -85,11 +85,31 @@ 라벨의 표시 순서를 지정합니다. +종류 +:::: + +「라벨」 또는 「태그」를 지정합니다. 일반 라벨은 「라벨」입니다. 「태그」는 사용자가 검색 화면에서 추가하는 태그입니다(아래의 「태그」를 참조). 종류를 지정하지 않은 기존 라벨은 「라벨」로 취급됩니다. + + 설정 삭제 -------- 목록 페이지의 설정 이름을 클릭하고 삭제 버튼을 클릭하면 확인 화면이 표시됩니다. 삭제 버튼을 누르면 설정이 삭제됩니다. +태그 +---- + +``fess_config.properties`` 에서 ``user.tag.enabled=true`` (기본값: ``false`` )로 설정하면 로그인한 사용자가 검색 결과에 태그를 붙일 수 있습니다. 기본 제공 ``bootstrap`` 테마에서는 결과에 태그가 표시되고, 태그 추가와 자신이 붙인 태그의 삭제가 가능하며, 「태그」 패싯으로 결과를 좁힐 수 있습니다. API 에 대해서는 :doc:`../api/api-tag` 를 참조하십시오. + +태그는 종류가 「태그」인 라벨로 저장됩니다. 이름이 태그 이름, 값이 태그 이름의 SHA-256, 대상 경로가 태그를 붙인 URL(한 줄에 하나, 정확히 일치), 권한이 태그를 볼 수 있는 사용자입니다. 태그를 붙인 사용자는 권한에 추가됩니다. + +- 태그는 그 라벨의 권한이 호출자와 일치하는 경우에만 보입니다. 관리자는 이 화면에서 태그를 편집하여 역할이나 그룹과 공유하거나 삭제할 수 있습니다. +- 같은 이름의 태그는 하나의 라벨로 합쳐집니다. 따라서 같은 이름의 태그를 붙인 사용자끼리는 서로의 태그가 붙은 곳을 볼 수 있습니다. +- 태그는 라벨 목록 API( ``/api/v2/labels`` )나 검색 화면의 라벨 선택지에 포함되지 않습니다. +- 태그는 라벨 건수 상한( ``page.labeltype.max.fetch.size`` , 기본값: 1000)에 포함됩니다. 상한에 도달하면 새 태그를 만들 수 없습니다. +- 관리자가 이 화면에서 태그를 변경 또는 삭제한 후, 인덱스의 문서는 다시 크롤링하거나 「Label Updater」 작업을 실행할 때까지 이전 값을 유지합니다. +- 문서 하나에 붙일 수 있는 태그 수는 ``user.tag.max.document.tags`` (기본값: 100), 태그 이름의 길이는 ``user.tag.name.max.length`` (기본값: 50)자까지입니다. + .. |image0| image:: ../../../resources/images/en/15.9/admin/labeltype-1.png .. |image1| image:: ../../../resources/images/en/15.9/admin/labeltype-2.png diff --git a/ko/15.9/admin/mapping-guide.rst b/ko/15.9/admin/mapping-guide.rst index df49a2fd..3ae8cd07 100644 --- a/ko/15.9/admin/mapping-guide.rst +++ b/ko/15.9/admin/mapping-guide.rst @@ -49,6 +49,21 @@ 매핑 사전 형식으로 업로드할 수 있습니다. +기본 제공 매핑 사전 +=================== + +기본 ``mapping.txt`` 는 ``title`` 이나 ``content`` 등 검색용 필드의 분석에 사용되며, 히라가나, 작은 가나, 반각 가타카나를 전각 가타카나로 통일합니다. 따라서 「りんご」「リンゴ」「リンゴ」는 서로 일치합니다. |Fess| 15.9 에서는 다음 표기도 통일하게 되었습니다. + +- ゐ・ゑ(イ・エ), ゎ・ゕ・ゖ・ヮ・ヵ・ヶ・ㇰ〜ㇿ 등의 작은 가나, ゝ・ゞ(ヽ・ヾ) +- ヴ, ヴャ・ヴュ・ヴョ, 히라가나 ゔ, 반각 ヴ, ウ・う 와 결합용 탁점(U+3099)의 조합(예: ラヴ → ラブ, レヴュー → レビユー) + +또한 가나 바로 뒤에 쓰인 ‐ ‑ ‒ – — ― ⁻ ₋ − - 는 장음 기호 「ー」로 취급됩니다( ``prolonged_sound_mark_filter`` ). 따라서 「サ―バ-」나 「サ−バ‐」는 「サーバー」와 일치합니다. 반각 하이픈( ``-`` )이나 한자・영숫자 뒤의 대시(東京-大阪, 2026−10−02)는 변환되지 않습니다. + +일본어용 ``ja/mapping.txt`` ( ``*_ja`` 필드)는 형태소 분석을 위해 히라가나와 작은 가나를 그대로 두고, ヴ 등의 표기만 통일합니다. + +.. note:: + + 이 설정은 새로 생성된 문서 인덱스에 적용됩니다. 기존 인덱스는 재인덱싱할 때까지 이전의 분석 설정과 사전을 사용합니다. 기존 인덱스에 적용하려면 :doc:`maintenance-guide` 에서 「사전 초기화」를 활성화하고 재인덱싱하십시오. 사전 초기화는 관리 화면에서 편집한 ``mapping.txt`` / ``ja/mapping.txt`` 의 내용을 덮어씁니다. .. |image0| image:: ../../../resources/images/en/15.9/admin/mapping-1.png .. |image1| image:: ../../../resources/images/en/15.9/admin/mapping-2.png diff --git a/ko/15.9/admin/relatedquery-guide.rst b/ko/15.9/admin/relatedquery-guide.rst index e771d845..9c62e4ad 100644 --- a/ko/15.9/admin/relatedquery-guide.rst +++ b/ko/15.9/admin/relatedquery-guide.rst @@ -54,6 +54,63 @@ 목록 페이지의 설정 이름을 클릭하고 삭제 버튼을 클릭하면 확인 화면이 표시됩니다. 삭제 버튼을 누르면 설정이 삭제됩니다. +검색 로그에서 생성 +------------------ + +목록 페이지의 [검색 로그에서 생성] 버튼을 클릭하면 최근 검색 로그에서 관련 쿼리를 만듭니다. +같은 사용자 세션에서 어떤 검색 직후 짧은 간격으로 이어서 수행된 검색(오타 다음의 올바른 단어나, 넓은 단어 다음의 보다 구체적인 단어 등)을 바꿔 말하기로 보고, 자주 검색되는 검색어에 대해 가장 많은 바꿔 말하기를 그 검색어의 관련 쿼리로 만듭니다. + +관련 쿼리는 모든 사용자에게 적용되고 그 검색어의 검색을 모두 넓히므로, 생성은 신중하게 이루어집니다. + +- 게스트가 볼 수 있는 검색만 사용합니다. 검색 로그의 역할이 모두 ``suggest.search.log.permissions`` (추천과 같은 설정)를 만족하는 경우에만 읽어 들입니다. +- ``label:"x"`` 등의 필드 지정, 연산자, 와일드카드, ``sort:`` , 앞에 붙은 ``+`` / ``-`` 를 포함한 검색어는 사용하지 않습니다. +- [추천 > 제외 단어] 에 등록한 단어는 검색어로도 관련 쿼리로도 사용하지 않습니다. +- 검색어와 그 관련 쿼리는 각각 ``related_query.generate.min.sessions`` 이상의 세션에서 얻은 것으로 한정합니다. 바꿔 말한 후의 검색은 히트가 있는 것으로 한정합니다. +- 가상 호스트별로 생성합니다. 가상 호스트가 없는 검색 로그는 기본 호스트로 취급합니다. +- 이미 관련 쿼리가 있는 검색어는 변경하지 않습니다(건너뛴 건수가 결과에 표시됩니다). 또한 관련 쿼리 캐시에 읽어 들일 수 있는 건수( ``page.relatedquery.max.fetch.size`` )를 넘어서 생성하지 않습니다. + +생성된 관련 쿼리는 직접 등록한 것과 마찬가지로 편집, 삭제할 수 있습니다. +[시스템 > 일반] 의 「검색 로그」 또는 「사용자 로그」가 비활성화되어 있으면 생성할 수 없습니다. 생성이 실행 중일 때는 다시 실행할 수 없습니다. + +생성 동작은 ``fess_config.properties`` 의 다음 설정으로 조정할 수 있습니다. + +.. list-table:: + :header-rows: 1 + :widths: 45 40 15 + + * - 속성 + - 설명 + - 기본값 + * - ``related_query.generate.days`` + - 읽어 들이는 검색 로그의 일수 + - ``30`` + * - ``related_query.generate.term.size`` + - 가상 호스트별 검색어의 최대 수 + - ``100`` + * - ``related_query.generate.query.size`` + - 검색어 하나에 만드는 관련 쿼리의 최대 수 + - ``5`` + * - ``related_query.generate.min.sessions`` + - 검색어와 관련 쿼리가 나타나야 하는 최소 세션 수 + - ``3`` + * - ``related_query.generate.session.interval`` + - 바꿔 말하기로 간주하는 간격(분) + - ``10`` + * - ``related_query.generate.seed.log.size`` + - 검색어 하나에 대해 읽어 들이는 검색 로그의 최대 수 + - ``1000`` + * - ``related_query.generate.seed.session.size`` + - 검색어 하나에 대해 읽어 들이는 세션의 최대 수 + - ``200`` + * - ``related_query.generate.log.fetch.size`` + - 검색어 하나에 대해 읽어 들이는 후속 검색 로그의 최대 수 + - ``2000`` + * - ``related_query.generate.query.min.length`` + - 검색어와 관련 쿼리의 최소 길이(문자 수) + - ``2`` + * - ``related_query.generate.query.max.length`` + - 검색어와 관련 쿼리의 최대 길이(문자 수) + - ``50`` .. |image0| image:: ../../../resources/images/en/15.9/admin/relatedquery-1.png .. |image1| image:: ../../../resources/images/en/15.9/admin/relatedquery-2.png diff --git a/ko/15.9/admin/searchlog-guide.rst b/ko/15.9/admin/searchlog-guide.rst index 53171cf4..bfc2e9dd 100644 --- a/ko/15.9/admin/searchlog-guide.rst +++ b/ko/15.9/admin/searchlog-guide.rst @@ -1,27 +1,68 @@ -====== +========= 검색 로그 -====== +========= 개요 ==== -검색, 클릭, 즐겨찾기의 실행 결과가 기록되며, 검색 로그는 이 관리 화면에서 확인할 수 있습니다. +검색, 클릭, 즐겨찾기의 실행 결과는 기록됩니다. 검색 로그 화면에서는 이를 집계한 분석 리포트와 개별 로그 목록을 확인할 수 있습니다. -관리 방법 -====== +화면을 열려면 왼쪽 메뉴의 [시스템 정보 > 검색 로그] 를 클릭합니다. 처음에는 「개요」 탭이 표시됩니다. +열람하려면 ``admin-searchlog`` 또는 ``admin-searchlog-view`` 역할이 필요합니다. ``admin-searchlog-view`` 로는 로그를 삭제할 수 없습니다. -목록 -==== +분석 리포트 +=========== + +기간과 필터 +----------- + +각 탭 상단에서 집계할 기간을 「오늘」「어제」「최근 7일」「최근 28일」「최근 90일」 중에서 고르거나, 「사용자 지정」으로 시작일과 종료일(최대 366일)을 지정합니다. +「이전 기간과 비교」를 활성화하면 같은 일수만큼 앞선 기간의 값과 비교합니다. 접근 유형과 표의 행 수도 지정할 수 있습니다. +기간과 그래프의 구간은 서버 시간대의 날짜를 따릅니다. + +탭 +-- + +- **개요**: 검색 수, 사용자 수, 제로 히트율, 클릭률, 평균 응답 시간 지표를 작은 추이 그래프 및 이전 기간 대비 변화와 함께 표시합니다. 지표를 전환할 수 있는 추이 그래프(비교 시 이전 기간을 점선으로 표시)와 인기 검색어, 제로 히트 검색어도 표시합니다. +- **검색어**: 검색어별 검색 수, 사용자 수, 평균 히트 수, 클릭 수, 클릭률, 평균 클릭 순위를 표시합니다. 제로 히트 검색어(최종 검색 일시 포함)와, 히트했지만 결과가 한 번도 클릭되지 않은 제로 클릭 검색어도 표시합니다. +- **클릭**: 클릭 수가 많은 URL, 즐겨찾기 수가 많은 URL, 클릭 순위 분포, 2페이지 이후 열람률을 표시합니다. +- **성능**: 응답 시간의 평균과 중앙값(p50)・p95・p99, 응답 시간 분포, 응답이 느린 검색어, 쿼리 시간을 표시합니다. +- **사용자·환경**: 신규 사용자와 기존 사용자, 접근 유형, 요일・시간대별 검색 수, 상위 User-Agent, 리퍼러, 언어, 가상 호스트를 표시합니다. 「역할·그룹별 검색 수」는 역할과 그룹별 검색 수, 사용자 수, 제로 히트율을 표시합니다. 검색은 실행한 사용자의 모든 역할과 그룹에 집계되므로 합계가 전체 검색 수보다 많아질 수 있습니다. 개별 사용자는 표시되지 않습니다. +- **AI 채팅**: AI 검색 모드(RAG 채팅)의 요청 수, 사용자 수, 총 토큰 수, 평균 응답 시간, 오류율과 상위 사용자, 의도별・모델별 요청을 표시합니다. 채팅 이용 로그는 ``rag.chat.log.enabled`` (기본값: ``true`` )가 활성화된 경우에 기록됩니다. 질문과 답변 내용은 기록되지 않습니다. 토큰 수는 LLM 플러그인이 보고한 경우에만 기록됩니다. +- **로그**: 개별 로그 목록입니다. 아래의 「로그 목록」을 참조하십시오. + +.. note:: + + 검색어별 클릭 지표와 제로 클릭 검색어는 클릭과 함께 검색어가 기록되게 된 이후( |Fess| 15.9 이후)의 클릭만 집계합니다. 전체 클릭 수와 클릭률에는 그 이전의 클릭도 포함됩니다. 사용자 수 등 일부 값은 근사치입니다. + +제로 히트 검색어 확인 +--------------------- -목록에서는 검색, 클릭, 즐겨찾기의 검색 로그를 확인할 수 있습니다. -검색 로그의 세부 정보를 확인하려면 대상 검색 로그를 클릭합니다. +「개요」와 「검색어」 탭에서 제로 히트 검색어를 클릭하면, 「로그」 탭에 히트 수 「0건만」으로 필터링한 그 검색어의 검색 로그가 표시됩니다. 어떤 검색에서 결과를 찾지 못했는지 확인하고 문서, 동의어, 관련 쿼리 추가에 활용할 수 있습니다. + +CSV 다운로드 +------------ + +분석 리포트의 각 표와 그래프에는 CSV 링크가 있어, 현재 기간, 비교, 접근 유형, 행 수 조건으로 집계한 내용을 다운로드할 수 있습니다. 필터 영역에는 지표를 정리한 CSV 링크도 있습니다. +숫자는 가공하지 않고 출력됩니다(비율은 0〜1, 시간은 밀리초). 비교 시의 그래프에는 ``_previous`` 열이 추가됩니다. + +로그 목록 +========= + +「로그」 탭에서는 검색 로그, 클릭 로그, 즐겨찾기 로그, 사용자 로그를 확인할 수 있습니다. +로그 유형, 쿼리 ID, 사용자 ID, 시간 범위, 접근 유형, 검색어로 필터링할 수 있으며, 검색 로그는 히트 수(「모두」「0건만」「1건 이상」)로도 필터링할 수 있습니다. +로그의 세부 정보를 확인하려면 대상 로그를 클릭합니다. |image0| +[CSV 다운로드] 버튼을 클릭하면 현재 조건에 일치하는 로그를 건수 제한 없이 최신순으로 CSV 로 다운로드할 수 있습니다. 머리글 행은 필드 이름이며 화면 언어에 따라 바뀌지 않습니다. + +CSV 파일은 분석 리포트의 것을 포함해 ``csv.file.encoding`` 의 문자 코드로 출력되며, UTF-8 의 경우 BOM 이 붙습니다. ``=`` , ``+`` , ``-`` , ``@`` , 탭, 캐리지 리턴으로 시작하는 값에는 스프레드시트에서 수식으로 처리되지 않도록 앞에 ``'`` 가 붙습니다. + 세부 정보 -==== +--------- -목록에서 검색 로그를 클릭하면 대상 검색 로그 세부 정보가 표시됩니다. +목록에서 로그를 클릭하면 대상 로그의 세부 정보가 표시됩니다. |image1| diff --git a/ko/15.9/api/admin/api-admin-backup.rst b/ko/15.9/api/admin/api-admin-backup.rst index 8ef3547d..999a1048 100644 --- a/ko/15.9/api/admin/api-admin-backup.rst +++ b/ko/15.9/api/admin/api-admin-backup.rst @@ -79,12 +79,13 @@ Backup API는 |Fess| 의 백업 대상 데이터를 참조 및 다운로드하 { "id": "system.properties", "name": "system.properties" }, { "id": "fess.json", "name": "fess.json" }, { "id": "doc.json", "name": "doc.json" }, + { "id": "chat_log.ndjson", "name": "chat_log.ndjson" }, { "id": "click_log.ndjson", "name": "click_log.ndjson" }, { "id": "favorite_log.ndjson", "name": "favorite_log.ndjson" }, { "id": "search_log.ndjson", "name": "search_log.ndjson" }, { "id": "user_info.ndjson", "name": "user_info.ndjson" } ], - "total": 10 + "total": 11 } } @@ -112,7 +113,7 @@ Backup API는 |Fess| 의 백업 대상 데이터를 참조 및 다운로드하 - 인덱스의 매핑 정의 파일 (``fess_indices/fess.json`` / ``fess_indices/fess/doc.json``) 그 자체 (``application/octet-stream``) * - ``*.bulk`` 또는 확장자 없는 인덱스 이름 - 대상 이름과 동일한 인덱스를 스크롤하여 생성한 벌크 데이터 (``application/octet-stream``). ``.bulk`` 를 제거한 이름을 인덱스 이름으로 처리합니다. - * - ``*.ndjson`` (``search_log`` / ``user_info`` / ``click_log`` / ``favorite_log``) + * - ``*.ndjson`` (``search_log`` / ``user_info`` / ``click_log`` / ``favorite_log`` / ``chat_log``) - 대응하는 로그의 NDJSON 데이터 (``application/x-ndjson``) .. note:: diff --git a/ko/15.9/api/api-export.rst b/ko/15.9/api/api-export.rst new file mode 100644 index 00000000..d9100b40 --- /dev/null +++ b/ko/15.9/api/api-export.rst @@ -0,0 +1,78 @@ +====================== +검색 결과 내보내기 API +====================== + +이 문서에서는 검색 결과를 CSV 또는 JSON 파일로 다운로드하는 |Fess| 의 v2 내보내기 API 에 대해 설명합니다. 공통 응답 엔벨로프·오류 모델에 대해서는 :doc:`api-overview` 를 참조하십시오. + +베이스 URL은 ``http:///api/v2/`` 입니다 (로컬 환경 예: ``http://localhost:8080/api/v2`` ). + +.. note:: + + 내보내기는 기본적으로 비활성화되어 있습니다. 사용하려면 ``fess_config.properties`` 에서 ``api.search.export=true`` 를 설정하십시오. 활성화하면 기본 제공 ``bootstrap`` 테마에서 검색 결과 건수 옆에 내보내기 메뉴(CSV / JSON)가 표시됩니다. ``/api/v2/ui/config`` 의 ``features.search_export`` 로 상태를 확인할 수 있습니다. + +검색 결과 다운로드 +================== + +요청 +---- + +================== ==================================================== +HTTP 메서드 GET +엔드포인트 ``/api/v2/documents/export`` +================== ==================================================== + +검색에 일치하는 문서를 파일 다운로드( ``Content-Disposition: attachment`` , 파일 이름은 ``search_results.csv`` 또는 ``search_results.json`` )로 반환합니다. + +- ``/api/v2/search`` 와 같은 역할 필터가 적용됩니다. ``login.required=true`` 인 경우에도 ``/api/v2/search`` 와 마찬가지로 액세스 토큰을 사용할 수 있습니다. +- 출력하는 건수는 최대 ``api.search.export.max.size`` (기본값: ``1000`` )건입니다. 페이징 파라미터( ``start`` , ``num`` )는 사용하지 않습니다. +- 출력하는 필드는 ``api.search.export.fields`` (기본값: ``title,url_link,last_modified,content_length,filetype`` ) 중 API 응답에 포함할 수 있는 필드뿐입니다. +- 요청은 1분당 ``api.search.export.rate.limit.per.minute`` (기본값: ``10`` , ``0`` 은 무제한)회로 제한됩니다. 로그인한 사용자는 사용자별, 게스트는 클라이언트 IP 주소별로 셉니다. 초과하면 ``429`` 와 ``Retry-After`` 헤더가 반환됩니다. +- 내보내기는 검색 로그에 기록되지 않습니다. + +요청 파라미터 +------------- + +``q`` , ``ex_q`` , ``fields.*`` , ``sort`` , ``lang`` 등 ``/api/v2/documents/all`` 과 같은 검색 조건 파라미터를 지정할 수 있습니다( :doc:`api-search` 참조). 이와 더불어 다음 파라미터가 있습니다. + +.. tabularcolumns:: |p{4cm}|p{11cm}| +.. list-table:: 요청 파라미터 + + * - ``format`` + - 파일 형식. ``csv`` (기본값) 또는 ``json`` . 그 밖의 값은 ``invalid_request`` (400)가 됩니다. + +표: 요청 파라미터 + +응답 +---- + +CSV 는 첫 행이 필드 이름의 머리글 행이며 ``csv.file.encoding`` 의 문자 코드로 출력됩니다(UTF-8 의 경우 BOM 이 붙습니다). ``=`` , ``+`` , ``-`` , ``@`` , 탭, 캐리지 리턴으로 시작하는 값에는 스프레드시트에서 수식으로 처리되지 않도록 앞에 ``'`` 가 붙습니다. 여러 값을 가진 필드는 공백으로 연결합니다. + +:: + + "title","url_link","last_modified","content_length","filetype" + "Example","https://example.com/","2025-01-01T00:00:00.000Z","1234","html" + +JSON 은 ``{"data":[{...},...]}`` 형식이며, 여러 값을 가진 필드는 배열 그대로입니다. + +파일 출력이 시작되기 전에 실패하면 일반 오류 엔벨로프가 반환됩니다. 출력이 시작된 후에 실패하면 파일이 도중에 끝납니다(CSV 는 도중까지, JSON 은 해석할 수 없는 내용이 됩니다). + +오류 응답 +--------- + +.. tabularcolumns:: |p{4cm}|p{11cm}| +.. list-table:: 오류 응답 + + * - 상태 코드 + - 설명 + * - 400 Bad Request + - 잘못된 쿼리, ``format`` 이 ``csv`` / ``json`` 이외, 또는 ``api.search.export=false`` 로 내보내기가 비활성화된 경우. + * - 401 Unauthorized + - 인증이 필요한 경우( ``login.required=true`` 에서 익명 호출자 등). + * - 405 Method Not Allowed + - HTTP 메서드가 허용되지 않는 경우. + * - 429 Too Many Requests + - 1분당 요청 수 상한을 초과한 경우. + * - 500 Internal Server Error + - 서버 내부 오류가 발생한 경우. + +표: 오류 응답 diff --git a/ko/15.9/api/api-search-history.rst b/ko/15.9/api/api-search-history.rst new file mode 100644 index 00000000..aa704557 --- /dev/null +++ b/ko/15.9/api/api-search-history.rst @@ -0,0 +1,104 @@ +============= +검색 기록 API +============= + +이 문서에서는 |Fess| 의 v2 검색 기록 API 에 대해 설명합니다. +공통 응답 엔벨로프·오류 모델에 대해서는 :doc:`api-overview` 를 참조하십시오. + +베이스 URL은 ``http:///api/v2/`` 입니다 (로컬 환경 예: ``http://localhost:8080/api/v2`` ). + +.. note:: + + 검색 기록은 ``search.history.enabled`` (기본값: ``true`` )와 검색 로그 기록이 모두 활성화된 경우에 사용할 수 있습니다. ``/api/v2/ui/config`` 의 ``features.search_history`` 로 상태를 확인할 수 있습니다. + +최근 검색 가져오기 +================== + +요청 +---- + +================== ==================================================== +HTTP 메서드 GET +엔드포인트 ``/api/v2/search-history`` +================== ==================================================== + +로그인한 사용자가 현재 가상 호스트에서 ``/api/v2/search`` 로 수행한 최근 검색을 최신순으로 반환합니다. 반환된 조건으로 같은 검색을 다시 실행할 수 있습니다. + +- 첫 페이지의 검색만 대상입니다. 같은 조건의 검색은 가장 최근의 1건으로 합쳐집니다. 검색어가 없는 검색은 포함되지 않습니다. +- 반환하는 건수는 최대 ``search.history.size`` (기본값: ``10`` )건입니다. +- 검색 로그는 1분마다 실행되는 작업으로 기록되므로, 검색이 기록에 나타나기까지 최대 1분 정도 걸립니다. +- 기록은 세션의 로그인 사용자에 연결됩니다. 익명 호출자는 ``auth_required`` (401)가 되며, 액세스 토큰은 로그인을 대신하지 않습니다. +- 검색 기록 기능이 비활성화된 경우 ``invalid_request`` (400)가 됩니다. + +요청 파라미터는 없습니다. + +응답 +---- + +성공 시(200)에는 공통 엔벨로프의 ``response`` 바로 아래에 다음 필드가 반환됩니다. + +:: + + { + "response": { + "status": 0, + "record_count": 1, + "data": [ + { + "q": "fess", + "fields": { "label": ["docs"] }, + "sort": "last_modified.desc", + "requested_at": "2026-10-01T09:00:00Z", + "hit_count": 42 + } + ] + } + } + +.. tabularcolumns:: |p{4cm}|p{11cm}| +.. list-table:: 응답 필드 + + * - ``record_count`` + - ``data`` 안의 검색 건수(int). + * - ``data`` + - 최근 검색의 배열(최신순). 조건의 키는 ``/api/v2/search`` 의 요청 파라미터 이름이며, 그 검색에서 사용하지 않은 키는 생략됩니다. + * - ``data[].q`` + - 검색어(str). + * - ``data[].fields`` + - ``fields.`` 으로 지정한 필드 조건(필드 이름을 키로 하는 값의 배열). + * - ``data[].ex_q`` + - 추가 쿼리(str 배열). + * - ``data[].sort`` + - 정렬 순서(str). + * - ``data[].lang`` + - ``lang`` 으로 지정한 언어(str 배열). + * - ``data[].requested_at`` + - 검색한 일시(UTC, ISO-8601). + * - ``data[].hit_count`` + - 그 검색의 히트 수(int64). + +표: 응답 필드 + +오류 응답 +--------- + +.. tabularcolumns:: |p{4cm}|p{11cm}| +.. list-table:: 오류 응답 + + * - 상태 코드 + - 설명 + * - 400 Bad Request + - 검색 기록 기능이 비활성화된 경우. + * - 401 Unauthorized + - 로그인하지 않은 경우. + * - 405 Method Not Allowed + - HTTP 메서드가 허용되지 않는 경우. + * - 500 Internal Server Error + - 서버 내부 오류가 발생한 경우. + +표: 오류 응답 + +기본 제공 테마에서의 표시 +========================= + +기본 제공 ``bootstrap`` 테마에서는 로그인한 사용자가 빈 검색 상자를 클릭하거나 검색 상자에서 아래쪽 화살표 키를 누르면 추천 드롭다운에 최근 검색이 표시됩니다. 항목을 선택하면 라벨 등의 조건을 포함해 같은 검색을 실행합니다. |Fess| 15.9 이전에 기록된 검색 로그에는 조건이 저장되어 있지 않으므로 기록에 표시되지 않습니다. diff --git a/ko/15.9/api/api-tag.rst b/ko/15.9/api/api-tag.rst new file mode 100644 index 00000000..89bd708e --- /dev/null +++ b/ko/15.9/api/api-tag.rst @@ -0,0 +1,130 @@ +======== +태그 API +======== + +이 문서에서는 문서에 태그를 붙이는 |Fess| 의 v2 태그 API 에 대해 설명합니다. +공통 응답 엔벨로프·오류 모델·CSRF 에 대해서는 :doc:`api-overview` 를 참조하십시오. + +베이스 URL은 ``http:///api/v2/`` 입니다 (로컬 환경 예: ``http://localhost:8080/api/v2`` ). + +.. note:: + + 태그 기능은 기본적으로 비활성화되어 있습니다. 사용하려면 ``fess_config.properties`` 에서 ``user.tag.enabled=true`` 를 설정하십시오. ``/api/v2/ui/config`` 의 ``features.user_tag`` 로 상태를 확인할 수 있습니다. + +태그는 종류가 「태그」인 라벨입니다( :doc:`../admin/labeltype-guide` 참조). 라벨 이름이 태그 이름, 값이 태그 이름의 SHA-256(16진수), 대상 경로가 태그를 붙인 URL, 권한이 태그를 볼 수 있는 사용자입니다. 태그는 그 라벨이 호출자에게 보이는 경우에만 보입니다. + +검색 API( ``/api/v2/search`` )는 각 히트에 호출자에게 보이는 태그를 ``tags`` 로 반환합니다. ``fields.tag=<값>`` 으로 태그가 붙은 문서로 좁힐 수 있고, ``facet.field=tag`` 로 태그 패싯을 가져올 수 있습니다. 인덱스의 ``tag`` 필드 자체는 응답에 포함되지 않습니다. + +태그 가져오기 +============= + +요청 +---- + +================== ==================================================== +HTTP 메서드 GET +엔드포인트 ``/api/v2/documents/{docId}/tags`` +================== ==================================================== + +지정한 문서의 태그 중 호출자에게 보이는 것을 반환합니다. 호출자가 그 문서를 검색할 수 없는 경우 ``not_found`` (404)가 됩니다. + +응답 +---- + +성공 시(200)에는 공통 엔벨로프의 ``response`` 바로 아래에 다음 필드가 반환됩니다. + +:: + + { + "response": { + "status": 0, + "doc_id": "a1b2c3d4e5f6", + "addable": true, + "tags": [ + { "value": "9f86d081884c7d65...", "name": "검토 필요", "mine": true } + ] + } + } + +.. tabularcolumns:: |p{4cm}|p{11cm}| +.. list-table:: 응답 필드 + + * - ``doc_id`` + - 문서 ID(str). + * - ``addable`` + - 호출자가 로그인하여 태그를 추가할 수 있는 경우 ``true`` (bool). + * - ``added`` + - POST 만. 호출자가 이미 그 문서에 태그를 붙인 경우 ``false`` (bool). + * - ``removed`` + - DELETE 만(bool). + * - ``tags`` + - 호출자에게 보이는 태그의 배열. 각 요소는 ``value`` (라벨 값. ``fields.tag`` 에 지정하는 값), ``name`` (태그 이름), ``mine`` (호출자가 태그 권한에 포함된 경우 ``true`` ). + +표: 응답 필드 + +태그 추가 +========= + +요청 +---- + +================== ==================================================== +HTTP 메서드 POST +엔드포인트 ``/api/v2/documents/{docId}/tags`` +================== ==================================================== + +로그인한 사용자로서 문서의 URL 에 태그를 붙입니다. 액세스 토큰은 로그인을 대신하지 않습니다. 상태를 변경하는 요청이므로 ``X-Fess-CSRF-Token`` 헤더가 필요합니다. + +- 같은 이름의 태그가 있으면 그 대상 경로에 URL 을, 권한에 사용자를 추가합니다. 없으면 그 사용자만 볼 수 있는 태그를 만듭니다. 따라서 같은 이름의 태그는 하나로 합쳐지며, 같은 이름의 태그를 붙인 사용자끼리는 서로의 태그가 붙은 곳을 볼 수 있습니다. +- 같은 문서에 다시 붙이면 ``added`` 는 ``false`` 가 됩니다. +- 문서 하나에 붙일 수 있는 태그는 최대 ``user.tag.max.document.tags`` (기본값: ``100`` )개입니다. + +요청 본문은 ``Content-Type: application/json`` 이며, ``name`` 에 태그 이름을 지정합니다. + +:: + + { + "name": "검토 필요" + } + +태그 이름은 NFKC 로 정규화되고, 연속된 공백은 하나로 합쳐지며, 앞뒤 공백은 제거됩니다. 길이는 1〜 ``user.tag.name.max.length`` (기본값: ``50`` )자이며, 제어 문자나 서식 문자(폭이 0인 문자나 양방향 재정의 등)를 포함한 이름은 거부됩니다. + +태그 삭제 +========= + +요청 +---- + +================== ==================================================== +HTTP 메서드 DELETE +엔드포인트 ``/api/v2/documents/{docId}/tags?value=<태그 값>`` +================== ==================================================== + +로그인한 사용자를 ``value`` 로 지정한 태그의 권한에서 제외합니다. 사용자, 그룹, 역할의 권한이 하나도 남지 않으면 태그는 삭제되고 문서에서도 제거됩니다. 사용자가 태그의 권한에 포함되어 있지 않으면 ``forbidden`` (403)이 됩니다. ``X-Fess-CSRF-Token`` 헤더가 필요합니다. + +오류 응답 +========= + +.. tabularcolumns:: |p{4cm}|p{11cm}| +.. list-table:: 오류 응답 + + * - 상태 코드 + - 설명 + * - 400 Bad Request + - 요청이 잘못된 경우(태그 기능이 비활성화된 경우, 태그 이름이 잘못된 경우, 태그 상한을 초과한 경우 포함). + * - 401 Unauthorized + - POST・DELETE 에서 로그인하지 않은 경우. + * - 403 Forbidden + - CSRF 토큰의 누락・만료, 또는 자신이 추가하지 않은 태그를 DELETE 한 경우. + * - 404 Not Found + - 문서를 찾을 수 없거나 호출자가 검색할 수 없는 경우. + * - 405 Method Not Allowed + - HTTP 메서드가 허용되지 않는 경우. + * - 413 Payload Too Large + - 요청 본문이 크기 상한을 초과한 경우. + * - 415 Unsupported Media Type + - 지원되지 않는 ``Content-Type`` 인 경우. + * - 500 Internal Server Error + - 서버 내부 오류가 발생한 경우. + +표: 오류 응답 diff --git a/ko/15.9/api/api-uiconfig.rst b/ko/15.9/api/api-uiconfig.rst index 06daf7f0..a215b0ea 100644 --- a/ko/15.9/api/api-uiconfig.rst +++ b/ko/15.9/api/api-uiconfig.rst @@ -44,6 +44,9 @@ SPA 가 필요로 하는 초기 설정을 반환합니다. }, "features": { "user_favorite": false, + "search_history": true, + "search_export": false, + "user_tag": false, "popular_word": true, "suggest_search_log": true, "suggest_documents": true, @@ -186,6 +189,15 @@ features * - ``user_favorite`` - boolean - 사용자 즐겨찾기 기능이 활성화되어 있는지 여부. + * - ``search_history`` + - boolean + - 검색 기록( ``GET /api/v2/search-history`` )을 사용할 수 있는지 여부( ``search.history.enabled`` 와 검색 로그가 모두 활성화된 경우 ``true`` ). + * - ``search_export`` + - boolean + - 검색 결과 내보내기( ``GET /api/v2/documents/export`` )가 활성화되어 있는지 여부( ``api.search.export`` ). + * - ``user_tag`` + - boolean + - 태그 기능( ``/api/v2/documents/{docId}/tags`` )이 활성화되어 있는지 여부( ``user.tag.enabled`` ). * - ``popular_word`` - boolean - 인기 검색어 기능이 활성화되어 있는지 여부. diff --git a/ko/15.9/api/index.rst b/ko/15.9/api/index.rst index 0bdd4a65..ce0bcb43 100644 --- a/ko/15.9/api/index.rst +++ b/ko/15.9/api/index.rst @@ -22,6 +22,7 @@ :caption: 검색 API api-search + api-export api-label api-popularword api-suggest @@ -40,6 +41,8 @@ :caption: 사용자 기능 API api-favorite + api-search-history + api-tag api-click api-cache diff --git a/ko/15.9/config/crawler-ocr.rst b/ko/15.9/config/crawler-ocr.rst index e3b015a4..296674a8 100644 --- a/ko/15.9/config/crawler-ocr.rst +++ b/ko/15.9/config/crawler-ocr.rst @@ -167,6 +167,7 @@ Docker 환경에서 OCR을 사용하려면 |Fess| 이미지에 Tesseract를 추 - 웹 크롤링 설정에서는 신규 생성 시 기본값으로 이미지 URL(jpg, png, gif 등)이 크롤링 대상에서 제외됩니다. 웹 사이트의 이미지를 크롤링하려면 「크롤링 대상에서 제외할 URL」에서 이를 삭제하십시오. 파일 크롤링에서는 이미지도 크롤링 대상이 됩니다. - 크롤러의 크기 제한도 적용됩니다. 파일 종류별 인덱스 크기 상한(기본값 10MB)에 대해서는 :doc:`crawler-basic`\ 을 참조하십시오. - OCR 정확도는 스캔한 이미지의 품질에 좌우됩니다. 손글씨는 일반적으로 인식되기 어렵습니다. +- OCR을 활성화하기 전에 색인된 파일에는 그대로는 OCR이 적용되지 않습니다. 증분 크롤링(:doc:`../admin/general-guide` 의 "최종 갱신 일시 확인")이 활성화되어 있으면 갱신 일시가 바뀌지 않은 파일은 다시 크롤링해도 새로 가져오지 않기 때문입니다. 기존 파일에 OCR을 적용하려면 "최종 갱신 일시 확인"을 일시적으로 비활성화하고 크롤링하거나, 해당 문서를 색인에서 삭제한 후 다시 크롤링하십시오. 업그레이드 시 주의 ================== diff --git a/ko/15.9/config/rate-limiting.rst b/ko/15.9/config/rate-limiting.rst index e330575e..5b3dc397 100644 --- a/ko/15.9/config/rate-limiting.rst +++ b/ko/15.9/config/rate-limiting.rst @@ -102,6 +102,32 @@ robots.txt 처리는 ``app/WEB-INF/classes/fess_config.properties`` 의 # robots.txt를 무시한다(기본값: false) crawler.ignore.robots.txt=false +Crawl-delay 는 접속 대상(오리진)별로 적용되며 상한은 60초입니다. 대기는 URL 단위로 이루어지므로, 증분 크롤링에서 하나의 URL 에 보내는 HEAD 와 GET 은 연이어 전송됩니다. +크롤링 설정의 「간격」은 이와 별도로 다음 URL 까지의 대기 시간으로 동작합니다. + +robots.txt 는 RFC 9309 에 따라 해석됩니다. ``Allow:`` 와 ``Disallow:`` 는 가장 길게 일치한 규칙이 우선하며, 시작 URL 도 robots.txt 로 확인합니다. robots.txt 로 거부된 URL 은 ``fess-crawler.log`` 에 INFO 로 기록되며 장애 URL 에는 등록되지 않습니다. + +429/503 응답 시 백오프 +---------------------- + +서버가 ``429 Too Many Requests`` 또는 ``503 Service Unavailable`` 을 반환하면 해당 오리진에 대한 접근을 일시적으로 멈추고 그 URL 을 최대 3회까지 재시도합니다. +대기 시간은 ``Retry-After`` 헤더가 있으면 그 값, 없으면 10초부터 시작하는 지수적 시간(최대 5분)입니다. +마지막 재시도도 실패하면 ``fess-crawler.log`` 에 WARN 이 출력됩니다. + +robots.txt 를 가져올 수 없는 경우 +--------------------------------- + +robots.txt 를 가져오는 것이 5xx, 429, 타임아웃 등으로 실패하면, 해당 오리진의 URL 은 백오프가 끝날 때까지 큐로 되돌려집니다. +첫 실패 후 3회의 재시도도 실패하면, 그 크롤링 동안 해당 오리진의 모든 URL 이 크롤링 대상에서 제외되고 ``fess-crawler.log`` 에 WARN 이 한 번 출력됩니다. +15.8 이전에는 가져올 수 없었던 robots.txt 는 모두 허용으로 취급되었습니다. +이 동작으로 되돌리려면 웹 크롤링 설정의 「설정 파라미터」에 다음을 지정합니다. + +:: + + client.robotsTxtAllowOnUnavailable=true + +robots.txt 처리 자체를 비활성화하려면 ``client.robotsTxtEnabled=false`` (크롤링 설정별) 또는 ``crawler.ignore.robots.txt=true`` 를 사용합니다. + 속도 제한 전체 설정 항목 ========================= diff --git a/ko/15.9/dev/theme-development.rst b/ko/15.9/dev/theme-development.rst index 76a0840b..15e61586 100644 --- a/ko/15.9/dev/theme-development.rst +++ b/ko/15.9/dev/theme-development.rst @@ -155,6 +155,14 @@ 허용하는 ``Content-Security-Policy`` 헤더와 함께 반환됩니다(인라인 스타일은 허용되지만 인라인 스크립트는 허용되지 않습니다). 따라서 외부 CDN 의 폰트나 스크립트는 로드되지 않으므로 테마에 포함하십시오. +- 엔트리 HTML 의 ``Content-Security-Policy`` 에는 ``frame-ancestors 'none'`` 이 + 포함되고 ``X-Frame-Options: DENY`` 헤더도 붙으므로, 페이지는 다른 페이지의 + 프레임에 표시되지 않습니다. ``frame-ancestors`` 의 값은 ``fess_config.properties`` + 의 ``theme.index.frame.ancestors`` (기본값: ``'none'`` )로 변경할 수 있습니다. + 값을 비우면 ``frame-ancestors`` 를 붙이지 않습니다( ``X-Frame-Options: DENY`` 는 + 그대로 붙습니다). WebKit 계열 브라우저(Safari 등)는 ``frame-ancestors 'none'`` + 에서 테마가 ``blob:`` URL 로 표시하는 프레임(PDF 미리보기나 캐시 표시)을 비워 둡니다. + 이를 표시하려면 값을 비우십시오. - 테마의 SPA 는 검색 결과나 채팅 등의 데이터를 ``/api/v2/*`` API 에서 가져옵니다. 패키징 diff --git a/ko/15.9/install/fess-setup.rst b/ko/15.9/install/fess-setup.rst index 492ebac2..4f846221 100644 --- a/ko/15.9/install/fess-setup.rst +++ b/ko/15.9/install/fess-setup.rst @@ -94,6 +94,14 @@ Playwright 크롤러가 필요로 하는 Node.js 를 |Fess| 디렉터리의 ``no ``install plugin`` , ``list plugins`` , ``upgrade plugins`` 에는 ``--repository `` 을 지정할 수 있습니다. 지정하면 버전 목록, jar, 체크섬을 모두 그 하나의 Maven 저장소(사내 미러 등)에서 가져오며, 기본 릴리스 저장소, 스냅숏 저장소, GitHub 은 사용하지 않습니다. +``--repository`` 에는 ``file:///`` 로 시작하는 URL 도 지정할 수 있습니다. 인터넷에 연결할 수 없는 서버에서는 Maven 저장소의 사본을 서버에 두고 그 디렉터리를 지정합니다. 지정하는 것은 각 플러그인의 디렉터리( ``/maven-metadata.xml`` 등)를 포함하는 디렉터리로, 기본 저장소의 ``https://maven.codelibs.org/release/org/codelibs/fess/`` 에 해당하는 위치입니다. 버전 확인과 체크섬 검증은 HTTP 저장소와 같은 방식으로 이루어집니다. + +:: + + $ bin/fess-setup install plugin fess-ds-csv --repository file:///opt/maven-repo/org/codelibs/fess/ + +``http`` , ``https`` , ``file`` 이외의 스킴이나 스킴이 없는 URL 은 한 줄의 오류로 거부됩니다. + install plugin -------------- @@ -107,6 +115,8 @@ jar 는 플러그인의 GitHub 릴리스에서 가져오며, 릴리스에 해당 사용 예는 :doc:`../admin/plugin-guide` 를 참조하십시오. +설치할 수 있는 것은 |Fess| 가 읽어 들이는 종류의 플러그인(이름이 ``fess-ds-`` , ``fess-ingest-`` , ``fess-script-`` , ``fess-webapp-`` , ``fess-thumbnail-`` , ``fess-crawler-`` , ``fess-llm-`` , ``fess-storage-`` , ``fess-sso-`` 로 시작하는 것)뿐입니다. ``list plugins`` 에 표시되는 것과 같습니다. 그 밖의 이름( ``fess-theme-*`` , ``fess`` , ``fess-crawler`` 등의 라이브러리)을 지정하면 아무것도 다운로드하지 않고 한 줄의 오류를 표시한 후 종료 코드 2로 종료합니다. 여러 이름 중 하나라도 해당하면 아무것도 설치하지 않습니다. 정적 테마는 ``install theme`` 으로 설치하십시오. + list plugins ------------ @@ -156,6 +166,8 @@ remove plugin 2. 대상 환경에서 ZIP 을 압축 해제하고, jar 를 ``app/WEB-INF/plugin/`` 에 둡니다. |Fess| 시작 후에 관리 화면 플러그인 설치 화면의 「로컬」 탭에서 업로드할 수도 있습니다( :doc:`../admin/plugin-guide` 참조). 3. |Fess| 를 시작합니다. 실행 중에 jar 를 둔 경우에는 재시작하십시오. +플러그인 jar 를 하나씩 가져오는 대신, Maven 저장소의 필요한 부분( ``org/codelibs/fess//`` 디렉터리)을 복사해 가져온 후 ``--repository file:///...`` 를 지정하여 ``install plugin`` 으로 설치할 수도 있습니다. 이 경우에도 체크섬이 검증됩니다. + 테마 관리 ========= diff --git a/ko/15.9/user/search-field.rst b/ko/15.9/user/search-field.rst index 1de2db29..ba9a2618 100644 --- a/ko/15.9/user/search-field.rst +++ b/ko/15.9/user/search-field.rst @@ -66,6 +66,12 @@ * - favorite_count - 문서가 즐겨찾기에 등록된 횟수 - 숫자 + * - owner + - 파일 소유자의 계정 이름 + - 키워드 + * - last_modifier + - 파일의 최종 수정자 + - 키워드 표: 사용 가능한 필드 목록 @@ -80,6 +86,8 @@ .. note:: 크롤링 대상에 따라 값이 등록되지 않는 필드도 있습니다. 예를 들어 anchor는 웹 크롤링 시에만, lang은 HTML에 언어 속성이 있는 경우에만 등록됩니다. 또한 segment(크롤링 실행 단위를 나타내는 세션 ID)나 doc_id(시스템이 부여하는 내부 ID) 등의 필드도 지정할 수 있지만, 일반적인 검색에서는 사용하지 않습니다. +owner 와 last_modifier 는 파일 서버 등을 크롤링할 때 등록됩니다. owner 에는 SMB, 파일 시스템, FTP 크롤링에서 가져온 파일 소유자의 계정 이름이 들어갑니다( ``DOMAIN\alice`` 와 같은 값은 ``alice`` 가 됩니다). last_modifier 에는 Office 문서 등에서 추출한 최종 수정자가 들어가며, 가져올 수 없으면 소유자가 들어갑니다. 웹 크롤링한 HTML 에는 owner 가 등록되지 않습니다. 예를 들어 ``owner:alice`` 나 ``last_modifier:"Taro Yamada"`` 와 같이 검색합니다. 기본 제공 테마에서는 고급 검색에서 소유자와 최종 수정자를 지정할 수 있습니다. 등록 여부는 관리자가 ``crawler.document.file.owner.enabled`` 와 ``crawler.document.file.last.modifier.enabled`` (둘 다 기본값은 ``true`` )로 전환할 수 있습니다. + HTML 파일을 검색 대상으로 하는 경우 title 태그가 title 필드에, body 태그 이하의 문자열이 content 필드에 등록됩니다. 사용 방법 diff --git a/zh-cn/15.9/admin/backup-guide.rst b/zh-cn/15.9/admin/backup-guide.rst index d5f95aa8..125c4d6a 100644 --- a/zh-cn/15.9/admin/backup-guide.rst +++ b/zh-cn/15.9/admin/backup-guide.rst @@ -45,6 +45,11 @@ doc.json doc.json包含fess索引的映射信息。 +chat_log.ndjson +::::::::::::::: + +chat_log.ndjson包含AI聊天使用日志信息。 + click_log.ndjson :::::::::::::::: diff --git a/zh-cn/15.9/admin/docreport-guide.rst b/zh-cn/15.9/admin/docreport-guide.rst new file mode 100644 index 00000000..64970817 --- /dev/null +++ b/zh-cn/15.9/admin/docreport-guide.rst @@ -0,0 +1,55 @@ +======== +文档报告 +======== + +概述 +==== + +文档报告是有助于整理已爬取的文件服务器等的管理页面。它会列出内容相同的文档和长期未更新的文档,并且各列表都可以下载为 CSV。 + +要打开该页面,请点击左侧菜单中的 [系统信息 > 文档报告]。查看需要 ``admin-docreport`` 或 ``admin-docreport-view`` 角色。该页面只进行显示和下载,不会修改任何文档。 + +两个选项卡都可以通过“URL 前缀”(例如 ``smb://server/share/`` )缩小范围。 + +重复文档 +======== + +将内容相同或几乎相同的文档分组,按文档数量从多到少显示。分组使用在编入索引时计算的内容签名( ``content_minhash_bits``\ ,与折叠重复搜索结果所用的相同),因此无需重新索引。内容中没有词语的文档(如空文件)不在对象范围内。 + +页面最多显示 ``docreport.duplicate.group.size`` (默认值:100)个组,每个组最多显示 ``docreport.duplicate.docs.size`` (默认值:10)个文档。可以通过 [下载 CSV] 获取所有组。即使索引很大,CSV 也会读取所有组,并输出 ``group, groupSize, url, title, filename, contentLength, lastModified, owner, lastModifier, clickCount, docId`` 列。 + +.. note:: + + 在不保存内容签名的索引( ``cloud``\ 、\ ``aws`` 用的映射)中,无法使用重复文档报告。 + +休眠文档 +======== + +按从旧到新的顺序显示最后修改时间早于指定天数(“未更新天数”,默认值为 ``docreport.dormant.days`` 的 365)的文档。没有最后修改时间的文档不会显示。启用“从未在搜索结果中被打开”后,会排除曾从搜索结果中被点击过的文档。 + +页面显示符合条件的文档数量、合计大小和文档列表。列表的翻页有上限( ``indexer.max.result.window.size`` ),超出该上限的文档请通过 [下载 CSV] 获取。 + +设置 +==== + +可以通过 ``fess_config.properties`` 中的以下设置调整其行为。 + +.. list-table:: + :header-rows: 1 + :widths: 40 45 15 + + * - 属性 + - 说明 + - 默认值 + * - ``docreport.duplicate.group.size`` + - 页面显示的重复组最大数量 + - ``100`` + * - ``docreport.duplicate.docs.size`` + - 页面中每个组列出的文档最大数量 + - ``10`` + * - ``docreport.duplicate.export.page.size`` + - 下载 CSV 时每次请求读取的内容签名数量 + - ``10000`` + * - ``docreport.dormant.days`` + - 视为休眠文档的、距最后修改的默认天数 + - ``365`` diff --git a/zh-cn/15.9/admin/general-guide.rst b/zh-cn/15.9/admin/general-guide.rst index 8756f0e9..0485b189 100644 --- a/zh-cn/15.9/admin/general-guide.rst +++ b/zh-cn/15.9/admin/general-guide.rst @@ -111,6 +111,8 @@ SSO类型 执行差分爬取时启用。 +对于 HTTP/HTTPS 的 URL,当无法通过 HEAD 响应的 ``Last-Modified`` 判断时(索引中没有最后修改时间、HEAD 响应中没有 ``Last-Modified``\ 、或 HEAD 返回 200/404 以外的状态),会在 GET 中附加 ``If-None-Match`` (索引中保存的 ``ETag`` )或 ``If-Modified-Since``\ ,进行条件 GET。服务器返回 ``304 Not Modified`` 的页面与未更新的页面一样,不会被重新获取。修改设置后等需要重新获取全部内容时,请禁用此设置后再爬取。 + 同时爬虫设置 :::::::::::::: diff --git a/zh-cn/15.9/admin/index.rst b/zh-cn/15.9/admin/index.rst index 628f0acb..20e15d39 100644 --- a/zh-cn/15.9/admin/index.rst +++ b/zh-cn/15.9/admin/index.rst @@ -49,6 +49,7 @@ log-guide failureurl-guide searchlist-guide + docreport-guide backup-guide maintenance-guide esreq-guide diff --git a/zh-cn/15.9/admin/labeltype-guide.rst b/zh-cn/15.9/admin/labeltype-guide.rst index b3665933..1e59f06d 100644 --- a/zh-cn/15.9/admin/labeltype-guide.rst +++ b/zh-cn/15.9/admin/labeltype-guide.rst @@ -85,11 +85,31 @@ 指定标签的排序顺序。 +类型 +:::: + +指定“标签”或“用户标签”。普通标签为“标签”。“用户标签”是用户在搜索页面添加的标签(请参阅下面的“用户标签”)。未指定类型的现有标签按“标签”处理。 + + 删除配置 -------- 点击列表页面中的配置名称,然后点击删除按钮,将显示确认画面。 点击删除按钮将删除配置。 +用户标签 +-------- + +在 ``fess_config.properties`` 中设置 ``user.tag.enabled=true`` (默认值: ``false`` )后,已登录的用户可以为搜索结果添加用户标签。在内置的 ``bootstrap`` 主题中,结果上会显示用户标签,用户可以添加用户标签并删除自己添加的用户标签,还可以通过“用户标签”分面缩小结果。关于 API,请参阅 :doc:`../api/api-tag`\ 。 + +用户标签以类型为“用户标签”的标签保存。名称为用户标签名,值为名称的 SHA-256,包含的路径为添加了用户标签的 URL(每行一个,完全匹配),权限决定谁可以查看该用户标签。添加用户标签的用户会被加入其权限中。 + +- 只有当标签的权限与调用者匹配时,用户标签才可见。管理员可以在此页面编辑用户标签,将其与角色或组共享,或将其删除。 +- 同名的用户标签会合并为一个标签。因此,添加了同名用户标签的用户之间可以看到彼此的用户标签添加在哪里。 +- 用户标签不包含在标签列表 API( ``/api/v2/labels`` )和搜索页面的标签选项中。 +- 用户标签计入标签数量上限( ``page.labeltype.max.fetch.size``\ ,默认值:1000)。达到上限后无法创建新的用户标签。 +- 管理员在此页面更改或删除用户标签后,索引中的文档在重新爬取或运行“Label Updater”作业之前仍保留以前的值。 +- 一个文档最多可以有 ``user.tag.max.document.tags`` (默认值:100)个用户标签,用户标签名最长为 ``user.tag.name.max.length`` (默认值:50)个字符。 + .. |image0| image:: ../../../resources/images/en/15.9/admin/labeltype-1.png .. |image1| image:: ../../../resources/images/en/15.9/admin/labeltype-2.png diff --git a/zh-cn/15.9/admin/mapping-guide.rst b/zh-cn/15.9/admin/mapping-guide.rst index 75020341..86de468e 100644 --- a/zh-cn/15.9/admin/mapping-guide.rst +++ b/zh-cn/15.9/admin/mapping-guide.rst @@ -49,6 +49,21 @@ 可以按映射词典格式上传。 +内置的映射词典 +============== + +默认的 ``mapping.txt`` 用于分析 ``title``\ 、\ ``content`` 等搜索字段,会将平假名、小写假名和半角片假名统一为全角片假名。因此“りんご”“リンゴ”“リンゴ”可以互相匹配。在 |Fess| 15.9 中,还会统一以下写法。 + +- ゐ・ゑ(イ・エ)、ゎ・ゕ・ゖ・ヮ・ヵ・ヶ・ㇰ〜ㇿ 等小写假名、ゝ・ゞ(ヽ・ヾ) +- ヴ、ヴャ・ヴュ・ヴョ、平假名 ゔ、半角 ヴ、ウ・う 与组合用浊音符(U+3099)的组合(例:ラヴ → ラブ、レヴュー → レビユー) + +此外,紧跟在假名之后的 ‐ ‑ ‒ – — ― ⁻ ₋ − - 会被视为长音符号“ー”( ``prolonged_sound_mark_filter`` )。因此“サ―バ-”和“サ−バ‐”可以匹配“サーバー”。半角连字符( ``-`` )以及汉字、字母、数字之后的破折号(東京-大阪、2026−10−02)不会被转换。 + +日语用的 ``ja/mapping.txt`` ( ``*_ja`` 字段)为了形态素分析,会保留平假名和小写假名,只统一 ヴ 等写法。 + +.. note:: + + 这些设置适用于新建的文档索引。现有索引在重新索引之前会继续使用以前的分析设置和词典。要应用到现有索引,请在 :doc:`maintenance-guide` 中启用“重置字典”后重新索引。重置字典会覆盖在管理界面中编辑过的 ``mapping.txt`` / ``ja/mapping.txt`` 的内容。 .. |image0| image:: ../../../resources/images/en/15.9/admin/mapping-1.png .. |image1| image:: ../../../resources/images/en/15.9/admin/mapping-2.png diff --git a/zh-cn/15.9/admin/relatedquery-guide.rst b/zh-cn/15.9/admin/relatedquery-guide.rst index 45361972..7a16a8ef 100644 --- a/zh-cn/15.9/admin/relatedquery-guide.rst +++ b/zh-cn/15.9/admin/relatedquery-guide.rst @@ -54,6 +54,61 @@ 点击列表页面中的配置名称,然后点击删除按钮,将显示确认画面。 点击删除按钮将删除配置。 +从搜索日志生成 +-------------- + +点击列表页面上的 [从搜索日志生成] 按钮,会根据最近的搜索日志创建相关查询。同一用户会话在某次搜索后短时间内接着进行的搜索(拼写错误之后的正确词,或宽泛词之后更具体的词等)被视为改写;对于经常被搜索的词,最常见的改写会成为该词的相关查询。 + +相关查询对所有人生效,并会扩展该词的所有搜索,因此生成会保守地进行。 + +- 只使用访客可见的搜索。只有当搜索日志的所有角色都满足 ``suggest.search.log.permissions`` (与建议相同的设置)时才会读取。 +- 不使用包含 ``label:"x"`` 等字段指定、运算符、通配符、 ``sort:`` 、开头的 ``+`` / ``-`` 的搜索词。 +- 在 [建议 > 不良词] 中登记的词既不作为搜索词,也不作为相关查询使用。 +- 搜索词及其相关查询各自必须来自至少 ``related_query.generate.min.sessions`` 个会话,并且改写后的搜索必须有命中。 +- 按虚拟主机分别生成。没有虚拟主机的搜索日志按默认主机处理。 +- 已有相关查询的词不会被更改(结果中会显示跳过的数量)。此外,生成的数量不会超过相关查询缓存可以加载的数量( ``page.relatedquery.max.fetch.size`` )。 + +生成的相关查询可以像手动登记的一样编辑和删除。当 [系统 > 通用] 中的“搜索日志”或“用户日志”被禁用时无法生成。生成正在执行时,不能再次执行。 + +可以通过 ``fess_config.properties`` 中的以下设置调整生成行为。 + +.. list-table:: + :header-rows: 1 + :widths: 45 40 15 + + * - 属性 + - 说明 + - 默认值 + * - ``related_query.generate.days`` + - 读取的搜索日志天数 + - ``30`` + * - ``related_query.generate.term.size`` + - 每个虚拟主机的搜索词最大数量 + - ``100`` + * - ``related_query.generate.query.size`` + - 每个搜索词生成的相关查询最大数量 + - ``5`` + * - ``related_query.generate.min.sessions`` + - 搜索词及其相关查询必须出现的最少会话数 + - ``3`` + * - ``related_query.generate.session.interval`` + - 视为改写的时间间隔(分钟) + - ``10`` + * - ``related_query.generate.seed.log.size`` + - 每个搜索词读取的搜索日志最大数量 + - ``1000`` + * - ``related_query.generate.seed.session.size`` + - 每个搜索词读取的会话最大数量 + - ``200`` + * - ``related_query.generate.log.fetch.size`` + - 每个搜索词读取的后续搜索日志最大数量 + - ``2000`` + * - ``related_query.generate.query.min.length`` + - 搜索词和相关查询的最小长度(字符数) + - ``2`` + * - ``related_query.generate.query.max.length`` + - 搜索词和相关查询的最大长度(字符数) + - ``50`` .. |image0| image:: ../../../resources/images/en/15.9/admin/relatedquery-1.png .. |image1| image:: ../../../resources/images/en/15.9/admin/relatedquery-2.png diff --git a/zh-cn/15.9/admin/searchlog-guide.rst b/zh-cn/15.9/admin/searchlog-guide.rst index fbb4c4dd..2004add0 100644 --- a/zh-cn/15.9/admin/searchlog-guide.rst +++ b/zh-cn/15.9/admin/searchlog-guide.rst @@ -1,27 +1,62 @@ -====== +======== 搜索日志 -====== +======== 概述 ==== -搜索、点击、收藏的执行结果会被记录,搜索日志可以在此管理界面中查看。 +搜索、点击、收藏的执行结果会被记录。在搜索日志页面中,可以查看对其进行汇总的分析报告,以及各条日志的列表。 -管理方法 -====== +要打开该页面,请点击左侧菜单中的 [系统信息 > 搜索日志]。首先显示“概览”选项卡。查看需要 ``admin-searchlog`` 或 ``admin-searchlog-view`` 角色。使用 ``admin-searchlog-view`` 无法删除日志。 -列表 -==== +分析报告 +======== + +期间和筛选 +---------- + +在各选项卡的顶部,从“今天”“昨天”“过去 7 天”“过去 28 天”“过去 90 天”中选择要汇总的期间,或者通过“自定义”指定开始日期和结束日期(最多 366 天)。启用“与上一时间段比较”后,会与之前相同天数的期间的值进行比较。还可以指定访问类型和表格的行数。期间和图表的分段按照服务器时区的日期划分。 + +选项卡 +------ + +- **概览**:显示搜索次数、用户数、零结果率、点击率、平均响应时间等指标,并附带小型趋势图和相对上一期间的变化。还显示可以切换指标的趋势图(比较时上一期间以虚线显示),以及热门查询词和零结果查询词。 +- **查询词**:显示每个查询词的搜索次数、用户数、平均命中数、点击数、点击率和平均点击位置。还显示零结果查询词(附最后搜索时间),以及有命中但结果从未被点击的零点击查询词。 +- **点击**:显示点击次数最多的 URL、收藏次数最多的 URL、点击位置分布,以及第 2 页及以后的浏览率。 +- **性能**:显示响应时间的平均值、中位数(p50)、p95、p99、响应时间分布、响应较慢的查询词和查询时间。 +- **用户与环境**:显示新用户和回访用户、访问类型、按星期和时段统计的搜索次数,以及排名靠前的 User-Agent、引用来源、语言和虚拟主机。“按角色和组统计的搜索次数”显示各角色和组的搜索次数、用户数和零结果率。搜索会计入执行该搜索的用户的所有角色和组,因此各行的合计可能超过总搜索次数。不会显示单个用户。 +- **AI 聊天**:显示 AI 搜索模式(RAG 聊天)的请求数、用户数、总令牌数、平均响应时间和错误率,以及排名靠前的用户、按意图和按模型统计的请求。聊天使用日志在 ``rag.chat.log.enabled`` (默认值: ``true`` )启用时记录。不会记录问题和回答的内容。令牌数仅在 LLM 插件报告时记录。 +- **日志**:各条日志的列表。请参阅下面的“日志列表”。 + +.. note:: + + 按查询词统计的点击指标和零点击查询词,只汇总开始与点击一起记录查询词之后( |Fess| 15.9 及以后)的点击。总点击数和点击率也包含之前的点击。用户数等部分数值是近似值。 + +查看零结果查询词 +---------------- -在列表中可以查看搜索、点击、收藏的搜索日志。 -如果要查看搜索日志的详细信息,请点击目标搜索日志。 +在“概览”和“查询词”选项卡中点击零结果查询词,会在“日志”选项卡中显示以命中数“仅零结果”筛选的该词的搜索日志。可以确认哪些搜索没有找到结果,并据此添加文档、同义词或相关查询。 + +下载 CSV +-------- + +分析报告中的每个表格和图表都有 CSV 链接,可以下载按当前期间、比较、访问类型和行数汇总的内容。筛选栏中还有汇总各项指标的 CSV 链接。数值按原样输出(比率为 0~1,时间为毫秒)。比较时的图表会增加 ``_previous`` 列。 + +日志列表 +======== + +在“日志”选项卡中,可以查看搜索日志、点击日志、收藏日志和用户日志。可以按日志类型、查询 ID、用户 ID、时间范围、访问类型和搜索词筛选,搜索日志还可以按命中数(“全部”“仅零结果”“至少一个结果”)筛选。要查看日志的详细信息,请点击目标日志。 |image0| +点击 [下载 CSV] 按钮,可以将符合当前条件的日志不限条数、按从新到旧的顺序下载为 CSV。标题行为字段名,不会随界面语言而变化。 + +CSV 文件(包括分析报告的 CSV)以 ``csv.file.encoding`` 的字符编码输出,UTF-8 时会附加 BOM。以 ``=``\ 、\ ``+``\ 、\ ``-``\ 、\ ``@``\ 、制表符或回车开头的值会在开头加上 ``'``\ ,以免在电子表格中被当作公式执行。 + 详情 -==== +---- -从列表中点击搜索日志后,将显示目标搜索日志的详细信息。 +从列表中点击日志后,将显示目标日志的详细信息。 |image1| diff --git a/zh-cn/15.9/api/admin/api-admin-backup.rst b/zh-cn/15.9/api/admin/api-admin-backup.rst index eee5bffa..97deb8a1 100644 --- a/zh-cn/15.9/api/admin/api-admin-backup.rst +++ b/zh-cn/15.9/api/admin/api-admin-backup.rst @@ -79,12 +79,13 @@ Backup API是用于参照和下载 |Fess| 备份对象数据的API。 { "id": "system.properties", "name": "system.properties" }, { "id": "fess.json", "name": "fess.json" }, { "id": "doc.json", "name": "doc.json" }, + { "id": "chat_log.ndjson", "name": "chat_log.ndjson" }, { "id": "click_log.ndjson", "name": "click_log.ndjson" }, { "id": "favorite_log.ndjson", "name": "favorite_log.ndjson" }, { "id": "search_log.ndjson", "name": "search_log.ndjson" }, { "id": "user_info.ndjson", "name": "user_info.ndjson" } ], - "total": 10 + "total": 11 } } @@ -112,7 +113,7 @@ Backup API是用于参照和下载 |Fess| 备份对象数据的API。 - 索引的映射定义文件(``fess_indices/fess.json`` / ``fess_indices/fess/doc.json``)本身,按原样返回(``application/octet-stream``) * - ``*.bulk`` 或不带扩展名的索引名 - 对与对象名同名的索引进行滚动(scroll)生成的批量数据(``application/octet-stream``)。去除 ``.bulk`` 后的名称作为索引名处理。 - * - ``*.ndjson`` (``search_log`` / ``user_info`` / ``click_log`` / ``favorite_log``) + * - ``*.ndjson`` (``search_log`` / ``user_info`` / ``click_log`` / ``favorite_log`` / ``chat_log``) - 对应日志的NDJSON数据(``application/x-ndjson``) .. note:: diff --git a/zh-cn/15.9/api/api-export.rst b/zh-cn/15.9/api/api-export.rst new file mode 100644 index 00000000..b52e9843 --- /dev/null +++ b/zh-cn/15.9/api/api-export.rst @@ -0,0 +1,78 @@ +================ +搜索结果导出 API +================ + +本文档介绍将搜索结果下载为 CSV 或 JSON 文件的 |Fess| v2 导出 API。公共响应信封和错误模型相关内容,请参阅 :doc:`api-overview`\ 。 + +基础 URL 为 ``http:///api/v2/``\ (本地环境示例:\ ``http://localhost:8080/api/v2``\ )。 + +.. note:: + + 导出默认禁用。要使用此功能,请在 ``fess_config.properties`` 中设置 ``api.search.export=true``\ 。启用后,内置的 ``bootstrap`` 主题会在搜索结果数量旁边显示导出菜单(CSV / JSON)。可以通过 ``/api/v2/ui/config`` 的 ``features.search_export`` 确认其状态。 + +下载搜索结果 +============ + +请求 +---- + +================== ==================================================== +HTTP 方法 GET +端点 ``/api/v2/documents/export`` +================== ==================================================== + +将与搜索匹配的文档作为文件下载( ``Content-Disposition: attachment``\ ,文件名为 ``search_results.csv`` 或 ``search_results.json`` )返回。 + +- 应用与 ``/api/v2/search`` 相同的角色筛选。即使在 ``login.required=true`` 时,也可以像 ``/api/v2/search`` 一样使用访问令牌。 +- 最多导出 ``api.search.export.max.size`` (默认值: ``1000`` )个文档。不使用分页参数( ``start``\ 、\ ``num`` )。 +- 导出的字段为 ``api.search.export.fields`` (默认值: ``title,url_link,last_modified,content_length,filetype`` )中可以包含在 API 响应中的字段。 +- 请求被限制为每分钟 ``api.search.export.rate.limit.per.minute`` (默认值: ``10``\ ,\ ``0`` 表示不限制)次。已登录用户按用户计数,访客按客户端 IP 地址计数。超出时返回 ``429`` 和 ``Retry-After`` 头。 +- 导出不会记录到搜索日志中。 + +请求参数 +-------- + +可以指定 ``q``\ 、\ ``ex_q``\ 、\ ``fields.*``\ 、\ ``sort``\ 、\ ``lang`` 等与 ``/api/v2/documents/all`` 相同的搜索条件参数(请参阅 :doc:`api-search` )。此外还有以下参数。 + +.. tabularcolumns:: |p{4cm}|p{11cm}| +.. list-table:: 请求参数 + + * - ``format`` + - 文件格式。 ``csv`` (默认)或 ``json``\ 。其他值返回 ``invalid_request`` (400)。 + +表: 请求参数 + +响应 +---- + +CSV 的第 1 行为字段名标题行,以 ``csv.file.encoding`` 的字符编码输出(UTF-8 时附加 BOM)。以 ``=``\ 、\ ``+``\ 、\ ``-``\ 、\ ``@``\ 、制表符或回车开头的值会在开头加上 ``'``\ ,以免在电子表格中被当作公式执行。具有多个值的字段以空格连接。 + +:: + + "title","url_link","last_modified","content_length","filetype" + "Example","https://example.com/","2025-01-01T00:00:00.000Z","1234","html" + +JSON 为 ``{"data":[{...},...]}`` 格式,具有多个值的字段保持为数组。 + +在文件开始输出之前失败时,返回通常的错误信封。输出开始后失败时,文件会中途结束(CSV 只输出到中途,JSON 会成为无法解析的内容)。 + +错误响应 +-------- + +.. tabularcolumns:: |p{4cm}|p{11cm}| +.. list-table:: 错误响应 + + * - 状态码 + - 说明 + * - 400 Bad Request + - 查询不正确、 ``format`` 不是 ``csv`` / ``json``\ ,或通过 ``api.search.export=false`` 禁用了导出时。 + * - 401 Unauthorized + - 需要认证时(例如 ``login.required=true`` 下的匿名调用者)。 + * - 405 Method Not Allowed + - HTTP 方法不被允许时。 + * - 429 Too Many Requests + - 超过每分钟请求数上限时。 + * - 500 Internal Server Error + - 发生服务器内部错误时。 + +表: 错误响应 diff --git a/zh-cn/15.9/api/api-search-history.rst b/zh-cn/15.9/api/api-search-history.rst new file mode 100644 index 00000000..40560b3f --- /dev/null +++ b/zh-cn/15.9/api/api-search-history.rst @@ -0,0 +1,103 @@ +============ +搜索历史 API +============ + +本文档介绍 |Fess| v2 搜索历史 API。公共响应信封和错误模型相关内容,请参阅 :doc:`api-overview`\ 。 + +基础 URL 为 ``http:///api/v2/``\ (本地环境示例:\ ``http://localhost:8080/api/v2``\ )。 + +.. note:: + + 当 ``search.history.enabled`` (默认值: ``true`` )和搜索日志记录都启用时,才能使用搜索历史。可以通过 ``/api/v2/ui/config`` 的 ``features.search_history`` 确认其状态。 + +获取最近的搜索 +============== + +请求 +---- + +================== ==================================================== +HTTP 方法 GET +端点 ``/api/v2/search-history`` +================== ==================================================== + +按从新到旧的顺序返回已登录用户在当前虚拟主机上通过 ``/api/v2/search`` 进行的最近搜索。可以使用返回的条件再次执行同样的搜索。 + +- 只包含第 1 页的搜索。条件相同的搜索会合并为最新的 1 条。不包含没有搜索词的搜索。 +- 最多返回 ``search.history.size`` (默认值: ``10`` )条。 +- 搜索日志由每分钟运行的作业写入,因此搜索最多需要大约 1 分钟才会出现在历史中。 +- 历史与会话中的登录用户关联。匿名调用者会得到 ``auth_required`` (401),访问令牌不能代替登录。 +- 搜索历史功能被禁用时,返回 ``invalid_request`` (400)。 + +没有请求参数。 + +响应 +---- + +成功时(200),在公共信封的 ``response`` 下直接返回以下字段。 + +:: + + { + "response": { + "status": 0, + "record_count": 1, + "data": [ + { + "q": "fess", + "fields": { "label": ["docs"] }, + "sort": "last_modified.desc", + "requested_at": "2026-10-01T09:00:00Z", + "hit_count": 42 + } + ] + } + } + +.. tabularcolumns:: |p{4cm}|p{11cm}| +.. list-table:: 响应字段 + + * - ``record_count`` + - ``data`` 中的搜索条数(int)。 + * - ``data`` + - 最近搜索的数组(从新到旧)。条件的键使用 ``/api/v2/search`` 的请求参数名,该搜索未使用的键会被省略。 + * - ``data[].q`` + - 搜索词(str)。 + * - ``data[].fields`` + - 通过 ``fields.`` 指定的字段条件(以字段名为键的值数组)。 + * - ``data[].ex_q`` + - 附加查询(str 数组)。 + * - ``data[].sort`` + - 排序顺序(str)。 + * - ``data[].lang`` + - 通过 ``lang`` 指定的语言(str 数组)。 + * - ``data[].requested_at`` + - 搜索的日期时间(UTC,ISO-8601)。 + * - ``data[].hit_count`` + - 该搜索的命中数(int64)。 + +表: 响应字段 + +错误响应 +-------- + +.. tabularcolumns:: |p{4cm}|p{11cm}| +.. list-table:: 错误响应 + + * - 状态码 + - 说明 + * - 400 Bad Request + - 搜索历史功能被禁用时。 + * - 401 Unauthorized + - 未登录时。 + * - 405 Method Not Allowed + - HTTP 方法不被允许时。 + * - 500 Internal Server Error + - 发生服务器内部错误时。 + +表: 错误响应 + +在内置主题中的显示 +================== + +在内置的 ``bootstrap`` 主题中,已登录的用户点击空的搜索框或在搜索框中按下方向键时,建议下拉列表中会显示最近的搜索。选择某一项后,会连同标签等条件一起再次执行同样的搜索。在 |Fess| 15.9 之前记录的搜索日志没有保存条件,因此不会显示在历史中。 diff --git a/zh-cn/15.9/api/api-tag.rst b/zh-cn/15.9/api/api-tag.rst new file mode 100644 index 00000000..4c449bb2 --- /dev/null +++ b/zh-cn/15.9/api/api-tag.rst @@ -0,0 +1,129 @@ +============ +用户标签 API +============ + +本文档介绍为文档添加用户标签的 |Fess| v2 用户标签 API。公共响应信封、错误模型及 CSRF 相关内容,请参阅 :doc:`api-overview`\ 。 + +基础 URL 为 ``http:///api/v2/``\ (本地环境示例:\ ``http://localhost:8080/api/v2``\ )。 + +.. note:: + + 用户标签功能默认禁用。要使用此功能,请在 ``fess_config.properties`` 中设置 ``user.tag.enabled=true``\ 。可以通过 ``/api/v2/ui/config`` 的 ``features.user_tag`` 确认其状态。 + +用户标签是类型为“用户标签”的标签(请参阅 :doc:`../admin/labeltype-guide` )。标签名称为用户标签名,值为名称的 SHA-256(十六进制),包含的路径为添加了用户标签的 URL,权限决定谁可以查看。只有当该标签对调用者可见时,用户标签才可见。 + +搜索 API( ``/api/v2/search`` )会将每个命中中调用者可见的用户标签作为 ``tags`` 返回。可以通过 ``fields.tag=<值>`` 缩小到带有该用户标签的文档,并通过 ``facet.field=tag`` 获取用户标签的分面。索引中的 ``tag`` 字段本身不会包含在响应中。 + +获取用户标签 +============ + +请求 +---- + +================== ==================================================== +HTTP 方法 GET +端点 ``/api/v2/documents/{docId}/tags`` +================== ==================================================== + +返回指定文档的用户标签中调用者可见的部分。调用者无法搜索该文档时,返回 ``not_found`` (404)。 + +响应 +---- + +成功时(200),在公共信封的 ``response`` 下直接返回以下字段。 + +:: + + { + "response": { + "status": 0, + "doc_id": "a1b2c3d4e5f6", + "addable": true, + "tags": [ + { "value": "9f86d081884c7d65...", "name": "待确认", "mine": true } + ] + } + } + +.. tabularcolumns:: |p{4cm}|p{11cm}| +.. list-table:: 响应字段 + + * - ``doc_id`` + - 文档 ID(str)。 + * - ``addable`` + - 调用者已登录(可以添加用户标签)时为 ``true`` (bool)。 + * - ``added`` + - 仅 POST。调用者已为该文档添加过该用户标签时为 ``false`` (bool)。 + * - ``removed`` + - 仅 DELETE(bool)。 + * - ``tags`` + - 调用者可见的用户标签数组。每个元素包含 ``value`` (标签值,用于 ``fields.tag`` )、 ``name`` (用户标签名)和 ``mine`` (调用者包含在该用户标签的权限中时为 ``true`` )。 + +表: 响应字段 + +添加用户标签 +============ + +请求 +---- + +================== ==================================================== +HTTP 方法 POST +端点 ``/api/v2/documents/{docId}/tags`` +================== ==================================================== + +以已登录用户的身份为文档的 URL 添加用户标签。访问令牌不能代替登录。由于是更改状态的请求,需要 ``X-Fess-CSRF-Token`` 头。 + +- 如果存在同名的用户标签,则将 URL 添加到其包含的路径,并将用户添加到其权限中。如果不存在,则创建一个只有该用户可见的用户标签。因此,同名的用户标签会合并为一个,添加了同名用户标签的用户之间可以看到彼此的用户标签添加在哪里。 +- 再次为同一文档添加时, ``added`` 为 ``false``\ 。 +- 一个文档最多可以有 ``user.tag.max.document.tags`` (默认值: ``100`` )个用户标签。 + +请求体使用 ``Content-Type: application/json``\ ,在 ``name`` 中指定用户标签名。 + +:: + + { + "name": "待确认" + } + +用户标签名会进行 NFKC 规范化,连续的空白会合并为一个,并去除首尾空白。长度为 1 至 ``user.tag.name.max.length`` (默认值: ``50`` )个字符,包含控制字符或格式字符(零宽字符、双向覆盖等)的名称会被拒绝。 + +删除用户标签 +============ + +请求 +---- + +================== ==================================================== +HTTP 方法 DELETE +端点 ``/api/v2/documents/{docId}/tags?value=<用户标签值>`` +================== ==================================================== + +将已登录用户从 ``value`` 指定的用户标签的权限中移除。当不再剩下任何用户、组或角色的权限时,该用户标签会被删除,并从文档中移除。用户不在该用户标签的权限中时,返回 ``forbidden`` (403)。需要 ``X-Fess-CSRF-Token`` 头。 + +错误响应 +======== + +.. tabularcolumns:: |p{4cm}|p{11cm}| +.. list-table:: 错误响应 + + * - 状态码 + - 说明 + * - 400 Bad Request + - 请求不正确时(包括用户标签功能被禁用、用户标签名不正确、超过用户标签上限的情况)。 + * - 401 Unauthorized + - 在 POST、DELETE 中未登录时。 + * - 403 Forbidden + - CSRF 令牌缺失或失效,或者 DELETE 了自己未添加的用户标签时。 + * - 404 Not Found + - 找不到文档,或调用者无法搜索该文档时。 + * - 405 Method Not Allowed + - HTTP 方法不被允许时。 + * - 413 Payload Too Large + - 请求体超过大小上限时。 + * - 415 Unsupported Media Type + - 不支持的 ``Content-Type`` 时。 + * - 500 Internal Server Error + - 发生服务器内部错误时。 + +表: 错误响应 diff --git a/zh-cn/15.9/api/api-uiconfig.rst b/zh-cn/15.9/api/api-uiconfig.rst index ff07f22f..67c39d17 100644 --- a/zh-cn/15.9/api/api-uiconfig.rst +++ b/zh-cn/15.9/api/api-uiconfig.rst @@ -44,6 +44,9 @@ HTTP 方法 GET }, "features": { "user_favorite": false, + "search_history": true, + "search_export": false, + "user_tag": false, "popular_word": true, "suggest_search_log": true, "suggest_documents": true, @@ -186,6 +189,15 @@ features * - ``user_favorite`` - boolean - 用户收藏功能是否启用。 + * - ``search_history`` + - boolean + - 搜索历史( ``GET /api/v2/search-history`` )是否可用( ``search.history.enabled`` 和搜索日志都启用时为 ``true`` )。 + * - ``search_export`` + - boolean + - 搜索结果导出( ``GET /api/v2/documents/export`` )是否启用( ``api.search.export`` )。 + * - ``user_tag`` + - boolean + - 用户标签功能( ``/api/v2/documents/{docId}/tags`` )是否启用( ``user.tag.enabled`` )。 * - ``popular_word`` - boolean - 热门词功能是否启用。 diff --git a/zh-cn/15.9/api/index.rst b/zh-cn/15.9/api/index.rst index f486f142..df23cb2b 100644 --- a/zh-cn/15.9/api/index.rst +++ b/zh-cn/15.9/api/index.rst @@ -22,6 +22,7 @@ AI 聊天 API、请求与响应格式,以及认证方式。 :caption: 搜索 API api-search + api-export api-label api-popularword api-suggest @@ -40,6 +41,8 @@ AI 聊天 API、请求与响应格式,以及认证方式。 :caption: 用户功能 API api-favorite + api-search-history + api-tag api-click api-cache diff --git a/zh-cn/15.9/config/crawler-ocr.rst b/zh-cn/15.9/config/crawler-ocr.rst index 0aa1a8bc..36424167 100644 --- a/zh-cn/15.9/config/crawler-ocr.rst +++ b/zh-cn/15.9/config/crawler-ocr.rst @@ -167,6 +167,7 @@ RHEL / Rocky Linux / AlmaLinux(请先启用 EPEL):: - 在 Web 爬取配置中,新建时的默认设置会将图像 URL(jpg、png、gif 等)从爬取对象中排除。要爬取网站上的图像,请从“从爬取对象中排除的URL”中删除这些项。文件爬取会将图像也作为爬取对象。 - 爬虫的大小限制同样适用。关于按文件类型设置的索引大小上限(默认值 10MB),请参阅 :doc:`crawler-basic`。 - OCR 的精度取决于扫描图像的质量。手写文字通常难以被识别。 +- 启用 OCR 之前已编入索引的文件不会自动进行 OCR。启用增量爬取( :doc:`../admin/general-guide` 中的“检查上次修改时间”)时,修改时间未变化的文件在重新爬取时不会被再次获取。要对这些文件应用 OCR,请临时禁用“检查上次修改时间”后爬取一次,或者先从索引中删除这些文档后重新爬取。 升级注意事项 ============ diff --git a/zh-cn/15.9/config/rate-limiting.rst b/zh-cn/15.9/config/rate-limiting.rst index bc55307b..cc2ffbb3 100644 --- a/zh-cn/15.9/config/rate-limiting.rst +++ b/zh-cn/15.9/config/rate-limiting.rst @@ -101,6 +101,26 @@ robots.txt的处理由 ``app/WEB-INF/classes/fess_config.properties`` 中的 ``c # 忽略robots.txt(默认: false) crawler.ignore.robots.txt=false +Crawl-delay 按访问目标(源)分别应用,上限为 60 秒。等待以 URL 为单位进行,因此在增量爬取中,对同一个 URL 发送的 HEAD 和 GET 会连续发送。爬取配置的“间隔”与此分开,作为到下一个 URL 之前的等待时间生效。 + +robots.txt 按照 RFC 9309 解释。 ``Allow:`` 和 ``Disallow:`` 以最长匹配的规则优先,起始 URL 也会根据 robots.txt 进行检查。被 robots.txt 禁止的 URL 会以 INFO 级别记录到 ``fess-crawler.log``\ ,不会登记为故障 URL。 + +429/503 响应时的退避 +-------------------- + +当服务器返回 ``429 Too Many Requests`` 或 ``503 Service Unavailable`` 时,会暂时停止对该源的访问,并最多重试该 URL 3 次。等待时间在有 ``Retry-After`` 头时为该值,否则为从 10 秒开始的指数增长时间(最长 5 分钟)。最后一次重试也失败时,会在 ``fess-crawler.log`` 中输出 WARN。 + +无法获取 robots.txt 时 +---------------------- + +当获取 robots.txt 因 5xx、429 或超时等失败时,该源的 URL 会被放回队列,直到退避结束。首次失败后再重试 3 次仍失败时,在本次爬取期间该源的所有 URL 都不会被爬取,并会在 ``fess-crawler.log`` 中输出一次 WARN。15.8 及更早版本会将无法获取的 robots.txt 一律视为允许。要恢复该行为,请在 Web 爬取配置的“配置参数”中指定以下内容。 + +:: + + client.robotsTxtAllowOnUnavailable=true + +要完全禁用 robots.txt 处理,请使用 ``client.robotsTxtEnabled=false`` (按爬取配置)或 ``crawler.ignore.robots.txt=true``\ 。 + 速率限制的全部配置项 ====================== diff --git a/zh-cn/15.9/dev/theme-development.rst b/zh-cn/15.9/dev/theme-development.rst index 718b0ca6..5e49bec7 100644 --- a/zh-cn/15.9/dev/theme-development.rst +++ b/zh-cn/15.9/dev/theme-development.rst @@ -148,6 +148,13 @@ - 入口 HTML 在返回时附带 ``Content-Security-Policy`` 头,仅允许来自 |Fess| 自身的脚本、样式、图片和连接(允许内联样式,不允许内联脚本)。 因此,来自外部 CDN 的字体或脚本不会被加载;请将它们包含在主题中。 +- 入口 HTML 的 ``Content-Security-Policy`` 中包含 ``frame-ancestors 'none'``\ , + 并且还带有 ``X-Frame-Options: DENY`` 头,因此页面不会显示在其他页面的框架中。 + ``frame-ancestors`` 的值可以通过 ``fess_config.properties`` 中的 + ``theme.index.frame.ancestors`` (默认值: ``'none'`` )更改。设为空值时不再附加 + ``frame-ancestors`` ( ``X-Frame-Options: DENY`` 仍会附加)。在 ``frame-ancestors 'none'`` + 下,基于 WebKit 的浏览器(如 Safari)会将主题通过 ``blob:`` URL 显示的框架 + (PDF 预览或缓存显示)显示为空白。要显示这些内容,请将该值设为空。 - 主题的 SPA 会通过 ``/api/v2/*`` API 获取搜索结果、聊天等数据。 打包 diff --git a/zh-cn/15.9/install/fess-setup.rst b/zh-cn/15.9/install/fess-setup.rst index 410c3f9b..b1e2694f 100644 --- a/zh-cn/15.9/install/fess-setup.rst +++ b/zh-cn/15.9/install/fess-setup.rst @@ -94,6 +94,14 @@ install nodejs ``install plugin``\ 、\ ``list plugins`` 和 ``upgrade plugins`` 可以指定 ``--repository ``\ 。指定后,版本列表、jar 及其校验和都从该单个 Maven 仓库(例如内部镜像)获取,而不使用默认的发布仓库、快照仓库和 GitHub。 +``--repository`` 也可以指定以 ``file:///`` 开头的 URL。在无法连接互联网的服务器上,可以将 Maven 仓库的副本放到服务器上,并指定该目录。应指定包含各插件目录( ``/maven-metadata.xml`` 等)的目录,即与默认仓库的 ``https://maven.codelibs.org/release/org/codelibs/fess/`` 相对应的位置。版本解析和校验和验证与 HTTP 仓库相同。 + +:: + + $ bin/fess-setup install plugin fess-ds-csv --repository file:///opt/maven-repo/org/codelibs/fess/ + +``http``\ 、\ ``https``\ 、\ ``file`` 以外的协议,或没有协议的 URL,会以一行错误信息被拒绝。 + install plugin -------------- @@ -107,6 +115,8 @@ jar 从插件的 GitHub 发布获取;发布中没有对应文件时,则从 M 使用示例请参阅 :doc:`../admin/plugin-guide`\ 。 +只能安装 |Fess| 会加载的插件类型,即名称以 ``fess-ds-``\ 、\ ``fess-ingest-``\ 、\ ``fess-script-``\ 、\ ``fess-webapp-``\ 、\ ``fess-thumbnail-``\ 、\ ``fess-crawler-``\ 、\ ``fess-llm-``\ 、\ ``fess-storage-`` 或 ``fess-sso-`` 开头的插件,与 ``list plugins`` 显示的相同。指定其他名称( ``fess-theme-*``\ ,或 ``fess``\ 、\ ``fess-crawler`` 等库)时,不会下载任何内容,而是显示一行错误并以退出码 2 结束。多个名称中只要有一个被拒绝,就不会安装任何插件。静态主题请使用 ``install theme`` 安装。 + list plugins ------------ @@ -156,6 +166,8 @@ remove plugin 2. 在目标环境中解压 ZIP,并将 jar 放入 ``app/WEB-INF/plugin/``\ 。也可以在 |Fess| 启动后,从管理界面插件安装页面的「本地」选项卡上传(参见 :doc:`../admin/plugin-guide` )。 3. 启动 |Fess|\ 。如果在运行中放入了 jar,请重启。 +也可以不逐个带入插件的 jar,而是复制 Maven 仓库中所需的部分( ``org/codelibs/fess//`` 目录)带入服务器,然后指定 ``--repository file:///...`` 使用 ``install plugin`` 安装。这种情况下同样会验证校验和。 + 管理主题 ======== diff --git a/zh-cn/15.9/user/search-field.rst b/zh-cn/15.9/user/search-field.rst index a0957810..6a0dc1e6 100644 --- a/zh-cn/15.9/user/search-field.rst +++ b/zh-cn/15.9/user/search-field.rst @@ -66,6 +66,12 @@ * - favorite_count - 文档被收藏的次数 - 数值 + * - owner + - 文件所有者的账户名 + - 关键字 + * - last_modifier + - 文件的最后修改者 + - 关键字 表: 可用字段列表 @@ -80,6 +86,8 @@ .. note:: 根据爬取对象的不同,有些字段可能不会被注册值。例如,anchor 仅在 Web 爬取时才会被注册,lang 仅在 HTML 具有语言属性时才会被注册。此外,还可以指定 segment(表示爬取执行单位的会话 ID)、doc_id(系统分配编号的内部 ID)等字段,但通常的搜索中不会使用它们。 +owner 和 last_modifier 会在爬取文件服务器等时注册。owner 中存放通过 SMB、文件系统、FTP 爬取获取的文件所有者的账户名( ``DOMAIN\alice`` 这样的值会变为 ``alice`` )。last_modifier 中存放从 Office 文档等中提取的最后修改者,无法获取时存放所有者。Web 爬取的 HTML 不会注册 owner。例如,可以像 ``owner:alice`` 或 ``last_modifier:"Taro Yamada"`` 这样搜索。在内置主题中,可以在高级搜索中指定所有者和最后修改者。管理员可以通过 ``crawler.document.file.owner.enabled`` 和 ``crawler.document.file.last.modifier.enabled`` (默认均为 ``true`` )切换是否注册。 + 当 HTML 文件作为搜索对象时,title 标签注册为 title 字段,body 标签以下的字符串注册为 content 字段。 使用方法