From 588e97da804d49c90ef1a04a3b26d5f837648c3a Mon Sep 17 00:00:00 2001 From: baraline Date: Thu, 1 Oct 2026 00:48:41 +0200 Subject: [PATCH 1/7] fix(content)!: spell literal text so python-markdown reads it as text from_transport handed markdownify's output on as Markdown, and markdownify escapes nothing it was not asked to: text a user typed into GLPI came back as syntax once any peer rendered it. Measured through this reader into python-markdown with the four extensions to_transport uses: \\serveur lost a backslash, __init__ and ______ became emphasis, a "-----------" line under text made a heading, "* point" and "> merci" lines made a list and a quote, "#4521" at a line start made a heading, "[1]: url" was consumed as a reference definition, and a | in a cell dropped the rest of the row. The converter is now a MarkdownConverter subclass. escape() escapes nothing itself: it stands each context-dependent character in for with a private-use code point, and each container python-markdown parses on its own -- paragraph, div, list item, quote, heading, cell, document -- decides every stand-in in it at once, with its whole Markdown in view. The rules replay python-markdown's own: its code-span pattern, its escape set (asked of the renderer, not copied), its bracket counting, its emphasis pairing including the NOT_STRONG exemption, its rule, setext, fence, list, quote and reference-definition line tests, and its e-mail autolink, which it matches only after code spans, escapes, links and images are stashed. A character is escaped only where python-markdown would misread it, with a backslash it removes again, or with a character reference where no backslash works (< & = ~): ordinary prose -- fichier_de_test_v2.xlsx, C:\Temp, R&D, a # or a - mid-sentence -- comes back byte for byte as before. Structure markdownify garbled on the way through is fixed where literal safety depended on it: continuation lines indented by four spaces, so nested lists nest and keep their numbers; content after a nested list starts a block of its own instead of joining the list's last item; a quote in an item gets the blank line python-markdown needs, and a quote opening an item keeps its later lines lazy; items python-markdown writes loose are written loose, so the second read equals the first; a list on a line with three bullets is spaced, because python-markdown cannot nest anything under such a line; a
 in an item or quote is an indented
code block, except where python-markdown reads none, where its lines are
kept as literal text. Script, style and title bodies are dropped instead
of leaking into the text, // keep their words without a
~~ python-markdown would display, images stay images in headings and
cells, and a source newline is a space, as HTML displays it. The HTML
standard's obsolete elements now make a body HTML.

The degraded path spells its stripped text through the same rules, so a
body too deep to convert is literal-safe too.

test_literal_text.py holds the property -- to_transport(from_transport(
html)) displays what html displays, compared with an HTML parser, and the
Markdown is a fixed point -- over the measured misreadings, one case per
construct, realistic bodies and a seeded fuzzer (8 seeds x 50 bodies in
the suite; 30,000 bodies across 600 seeds measured clean). What the
format cannot carry is an inventory of strict xfails, each asserted to
stabilise after one cycle. The nested-list round-trip loss is gone from
the corpus.

Co-Authored-By: Claude Opus 5.5 (1M context) 
---
 glpi_python_client/content/conversion.py      | 1664 ++++++++++++++++-
 .../content/tests/test_conversion.py          |   76 +-
 .../content/tests/test_literal_text.py        | 1470 +++++++++++++++
 .../testing/tests/test_content_roundtrip.py   |   17 +-
 4 files changed, 3164 insertions(+), 63 deletions(-)
 create mode 100644 glpi_python_client/content/tests/test_literal_text.py

diff --git a/glpi_python_client/content/conversion.py b/glpi_python_client/content/conversion.py
index 55a6a37..cdbe9c9 100644
--- a/glpi_python_client/content/conversion.py
+++ b/glpi_python_client/content/conversion.py
@@ -4,6 +4,29 @@
 package's canonical Markdown representation used by the rich content
 models.
 
+Literal text
+------------
+
+**Text in the HTML is literal, and the Markdown spells it so.** A user who
+types ``__init__`` into GLPI means eight characters; Markdown reads the
+same eight as bold ``init``. So :meth:`GlpiContentConverter.from_transport`
+escapes a character exactly where python-markdown -- with this module's
+extensions, which is what :meth:`GlpiContentConverter.to_transport` and any
+peer rendering the Markdown uses -- would otherwise read it as syntax, and
+nowhere else. It has to happen here: once the Markdown is written, nothing
+can tell a literal ``__`` from a bold one any more. The rule set is
+documented on :class:`_LiteralSafeConverter`; the property it is held to is
+that ``to_transport(from_transport(html))`` displays what ``html``
+displays, and that the Markdown is a fixed point of the round trip.
+
+Ordinary prose carries no escape at all -- ``fichier_de_test_v2.xlsx``,
+``C:\\Temp\\logs``, a ``#`` or a ``-`` mid-sentence, ``R&D`` -- because
+python-markdown already renders those literally. What is escaped is what it
+would not: ``\\\\serveur`` (it would lose a backslash), ``__init__``, a
+``#4521`` or a ``> merci`` at the start of a line, a ``|`` in a table cell,
+````, ``&``. None of this is visible to anyone reading the ticket
+in either ITSM: the escapes exist only in the Markdown between them.
+
 Neither *parse* recurses -- ``html.parser`` is an iterative scanner --
 but ``markdownify`` walks the finished tree recursively, at about two
 CPython frames per nesting level, so inbound conversion has a nesting
@@ -45,13 +68,25 @@
 from __future__ import annotations
 
 import re
+from collections.abc import Callable, Iterable
 from html import unescape
 from html.entities import html5 as _HTML5_REFERENCES
 from html.parser import HTMLParser
+from itertools import chain
+from typing import Any, NamedTuple
 
-from bs4 import ParserRejectedMarkup
+from bs4 import Comment, Doctype, ParserRejectedMarkup, Tag
+from bs4.element import PageElement
+from markdown import Markdown
 from markdown import markdown as markdown_to_html
-from markdownify import markdownify as html_to_markdown
+from markdown.blockprocessors import HRProcessor, ReferenceProcessor
+from markdown.inlinepatterns import (
+    AUTOLINK_RE,
+    AUTOMAIL_RE,
+    BACKTICK_RE,
+    NOT_STRONG_RE,
+)
+from markdownify import MarkdownConverter
 
 from glpi_python_client._errors import GlpiContentError
 
@@ -62,6 +97,12 @@
 #: ```` -- parses as an *unknown* tag, whose markup is dropped
 #: while its (usually empty) body is kept, so the token silently vanishes
 #: from the middle of a sentence.
+#:
+#: The second group is the HTML standard's *obsolete* elements -- its list
+#: of features that "must not be used by authors", which old editors and
+#: e-mail clients write all the same. Without them a body marked up only with
+#: ``URGENT`` or ``
`` failed the probe, took +#: the plain-text path and kept its tags as text. _HTML_ELEMENTS = frozenset( """ a abbr address area article aside audio b base bdi bdo blockquote body br @@ -73,6 +114,10 @@ section select slot small source span strong style sub summary sup table tbody td template textarea tfoot th thead time title tr track u ul var video wbr + + acronym applet basefont bgsound big blink center dir font frame frameset + isindex keygen listing marquee menuitem multicol nextid nobr noembed + noframes plaintext rb rtc spacer strike tt xmp """.split() ) @@ -137,9 +182,9 @@ #: the normal path treats it: markup dropped, body kept in place. _BLOCK_ELEMENTS = frozenset( """ - address article aside blockquote br col dd details dialog div dl dt - fieldset figcaption figure footer form h1 h2 h3 h4 h5 h6 header hgroup - hr li main menu nav ol p pre search section summary table tbody td + address article aside blockquote br center col dd details dialog dir div + dl dt fieldset figcaption figure footer form h1 h2 h3 h4 h5 h6 header + hgroup hr li main menu nav ol p pre search section summary table tbody td tfoot th thead tr ul """.split() ) @@ -336,10 +381,10 @@ def handle_data(self, data: str) -> None: """Keep character data, a raw-text element's body included. A ``", True, id="script-body"), - pytest.param("", True, id="style-body"), - pytest.param("", True, id="marked-section"), - pytest.param("", True, id="unterminated-marked-section"), - pytest.param("", True, id="processing-instruction"), + pytest.param("", False, False, id="resolved-comment"), + pytest.param("", False, False, id="doctype"), + pytest.param("", False, False, id="bogus-declaration"), + pytest.param("", False, True, id="script-body"), + pytest.param("", False, True, id="style-body"), + pytest.param("", True, True, id="marked-section"), + pytest.param("", True, True, id="unterminated-marked-section"), + pytest.param("", True, True, id="processing-instruction"), ], ) -def test_the_degraded_path_keeps_exactly_what_the_converter_keeps( - construct: str, kept: bool +def test_the_degraded_path_keeps_at_least_what_the_converter_keeps( + construct: str, converted_keeps: bool, degraded_keeps: bool ) -> None: """Parity, construct by construct, and not one of these was a guess. Each expectation here was read off the converting path rather than - reasoned about, and three came back the opposite way round from the - obvious answer -- a ``
code
", + "- code", + id="code-opening-an-item-after-a-script", + ), + pytest.param( + "
  • code
", + "- code", + id="code-opening-an-item-after-a-comment", + ), + pytest.param( + "
    • code
    ", + "- code", + id="code-opening-an-item-after-an-empty-list", + ), + pytest.param( + '
    • i
      code
      ' + "
    ", + "- ![i](https://example.org/i.png)\n\n code", + id="code-after-an-image", + ), + ], +) +def test_shapes_with_a_rule_of_their_own(html: str, expected: str) -> None: + """The converter's special cases, each on the shape that reaches it. + + What displays nothing -- a comment, a script, an empty list -- does not + count as content before a code block, so the code still opens its item. + """ + + assert read(html) == expected + + +def test_an_ordered_list_keeps_its_start_when_rendered() -> None: + """``sane_lists`` is what keeps ``3.`` from restarting the count at 1.""" + + assert '
      ' in render("intro\n\n3. trois\n4. quatre") + + +@pytest.mark.parametrize( + ("html", "expected"), + [ + pytest.param( + "
      • item
        #4521 code
      ", + "- item\n\n #4521 code", + id="in-a-list-item", + ), + pytest.param( + "
      #4521 C:\\Temp\n  indenté
      ", + "> #4521 C:\\Temp\n> indenté", + id="in-a-blockquote", + ), + pytest.param( + "
      ligne\n```\nfin
      ", + "````\nligne\n```\nfin\n````", + id="holding-a-fence-line", + ), + ], +) +def test_a_preformatted_block_stays_one(html: str, expected: str) -> None: + """A fence opens only at the start of a line, so nested code is indented. + + Inside a list item or a block quote, python-markdown never sees + ``` ``` ``` at the start of a line and reads the fence as text -- the + code's own ``#4521`` then became a heading. An indented code block is the + spelling it does read there. At the top level the fence is made longer + than any fence line the code holds. + """ + + assert read(html) == expected + assert_survives(html) + + +@pytest.mark.xfail( + strict=True, + reason=( + "python-markdown runs html.parser over its whole source to find raw " + "HTML, before an indented code block is recognised, and re-emits a " + "numeric reference written without its semicolon with one: the code " + "then shows 'ᆩ'. A fence is stashed before that pass, which is " + "why a top-level
       is unaffected; inside a list item or a quote no "
      +        "fence can open. Measured on 3.10.3."
      +    ),
      +)
      +def test_a_numeric_reference_in_nested_code_gains_a_semicolon() -> None:
      +    assert_survives("
      echo &#4521
      ") + + +def test_a_preformatted_block_opening_a_list_item_keeps_its_text() -> None: + """The one place python-markdown can start no code block at all. + + A list item's first line is its paragraph, so the code there degrades to + its lines, each escaped as the literal text it is -- the words survive, + the preformatting does not. + """ + + html = "
      • #4521 code\n  suite
      " + + markdown = read(html) + + assert markdown == "- \\#4521 code \n suite" + assert "#4521 code" in render(markdown) + assert read(render(markdown)) == markdown + + +@pytest.mark.parametrize( + ("html", "expected"), + [ + pytest.param( + "
      • item
        • sous-item
        #4521 code
      ", + "- item\n\n - sous-item\n\n \\#4521 code", + id="in-a-list-item", + ), + pytest.param( + "
      • item
      #4521 code
      ", + "> - item\n>\n> \\#4521 code", + id="in-a-quote", + ), + ], +) +def test_code_right_after_a_nested_list_keeps_its_text( + html: str, expected: str +) -> None: + """The other place python-markdown can start no code block. + + An indented code block right after a list in the same item or quote is + indented exactly as the list's last item's own content, and + python-markdown reads it as a paragraph of that item. The lines are + kept as literal text in the right place instead. + """ + + markdown = read(html) + + assert markdown == expected + assert "#4521 code" in render(markdown) + assert read(render(markdown)) == markdown + + +#: What the format loses, whatever the escaping does: each body with why. +#: +#: None of these is literal text misread. Each is structure python-markdown +#: has no spelling for, or spells as something else. +LOSSES = [ + pytest.param( + "
      • a
      • b
      ", + "python-markdown continues a list across a blank line: two lists are " + "read back as one list of loose items.", + id="two-adjacent-lists", + ), + pytest.param( + "
      a
      b
      ", + "python-markdown continues a quote across a blank line: two quotes are " + "read back as one quote of two paragraphs.", + id="two-adjacent-quotes", + ), + pytest.param( + "

      a

      b

      ", + "Markdown has no blank line inside a paragraph: two breaks in a row " + "are a paragraph break, and read back as one.", + id="two-breaks-in-a-row", + ), + pytest.param( + "

      ab

      ", + "'`a``b`' is one code span holding 'a``b' to python-markdown.", + id="two-adjacent-code-spans", + ), + pytest.param( + "

      a b c

      ", + "python-markdown pairs '*a **b** c*' as three emphasis runs, and b " + "loses its bold.", + id="strong-inside-emphasis", + ), + pytest.param( + "
      ab
      ", + "A Markdown table starts with its header row, so one without gains an " + "empty header.", + id="table-without-a-header", + ), + pytest.param( + "
      ab
      ", + "python-markdown renders a header-only table with one empty body row.", + id="table-without-a-body", + ), + pytest.param( + "
      a
      • x
      ", + "A table cell holds one line of inline Markdown: its list is written " + "as its text.", + id="list-in-a-cell", + ), + pytest.param( + "
      a
      x\ny
      ", + "A table cell holds one line of inline Markdown: its code block is " + "written as inline code.", + id="code-block-in-a-cell", + ), + pytest.param( + "
      • a
      ", + "An item's first block is its paragraph: the code is kept as text " + "(test_a_preformatted_block_opening_a_list_item_keeps_its_text).", + id="code-opening-a-list-item", + ), + pytest.param( + "
      • a
        • b
        c
      ", + "The code would be indented as the nested item's own content: it is " + "kept as text (test_code_right_after_a_nested_list_keeps_its_text).", + id="code-after-a-nested-list", + ), +] + + +@pytest.mark.parametrize(("html", "reason"), LOSSES) +def test_what_markdown_cannot_carry(html: str, reason: str) -> None: + """The inventory of losses, each asserted to still be one. + + ``xfail(strict=True)`` is the point: a loss that stops being one fails + here, so the inventory stays true. Struck and underlined text lose their + line too, which this comparison does not see: + ``test_struck_text_keeps_its_words_and_loses_its_line``. + """ + + with pytest.raises(AssertionError): + assert_survives(html) + pytest.xfail(reason) + + +@pytest.mark.parametrize(("html", "reason"), LOSSES) +def test_what_is_lost_is_lost_once(html: str, reason: str) -> None: + """Whatever a loss costs, it costs on the first cycle and never again.""" + + again = read(render(read(html))) + + assert read(render(again)) == again, reason + + +@pytest.mark.parametrize( + "html", + [ + pytest.param( + "

      Bonjour

      ", + id="style", + ), + pytest.param("

      Bonjour

      ", id="script"), + pytest.param( + "RE: Imprimante" + "" + "

      Bonjour

      ", + id="outlook-shaped", + ), + ], +) +def test_style_script_and_title_bodies_are_not_text(html: str) -> None: + """A browser displays none of these, so neither does the Markdown. + + ``strip=["script", "style"]`` used to be passed to ``markdownify``, and + ``strip`` skips an element's own converter -- ``convert_script`` and + ``convert_style`` return ``""`` -- so the bodies leaked into the text as + prose. ```` has no converter at all and leaked the same way. + """ + + assert read(html) == "Bonjour" + + +@pytest.mark.parametrize( + ("html", "expected"), + [ + pytest.param( + '<font color="red">URGENT</font> serveur HS', "URGENT serveur HS", id="font" + ), + pytest.param("<center>Titre</center> suite", "Titre\n\nsuite", id="center"), + pytest.param("<strike>ancien</strike> nouveau", "ancien nouveau", id="strike"), + pytest.param("<big>gros</big> texte", "gros texte", id="big"), + pytest.param("<tt>code</tt> texte", "code texte", id="tt"), + pytest.param("<nobr>sans coupure</nobr>", "sans coupure", id="nobr"), + ], +) +def test_a_body_marked_up_only_with_obsolete_elements_is_html( + html: str, expected: str +) -> None: + """Old editors still write these; without them the tags were kept as text.""" + + assert read(html) == expected + + +def test_every_obsolete_element_name_makes_a_body_html() -> None: + """The HTML standard's list of obsolete elements, all of them recognised.""" + + obsolete = ( + "acronym applet basefont bgsound big blink center dir font frame " + "frameset isindex keygen listing marquee menuitem multicol nextid nobr " + "noembed noframes plaintext rb rtc spacer strike tt xmp" + ).split() + + assert set(obsolete) <= conversion._HTML_ELEMENTS + + +@pytest.mark.parametrize( + "html", + [ + pytest.param("<p><s>ancien</s> nouveau</p>", id="s"), + pytest.param("<p><del>ancien</del> nouveau</p>", id="del"), + pytest.param("<p><strike>ancien</strike> nouveau</p>", id="strike"), + ], +) +def test_struck_text_keeps_its_words_and_loses_its_line(html: str) -> None: + """python-markdown has no strikethrough, so ``~~x~~`` would show literally. + + Keeping the words and losing the line is the honest loss: the text is + all there, and what is missing is recorded in the round-trip inventory. + """ + + assert read(html) == "ancien nouveau" + + +@pytest.mark.parametrize( + "html", ["<P>Bonjour</P>", "<BR>Bonjour", "<DIV><B>Bonjour</B></DIV>"] +) +def test_upper_case_tags_are_html(html: str) -> None: + """Element names are case-insensitive; the probe has to be too.""" + + assert "<" not in read(html) + + +@pytest.mark.parametrize( + ("value", "expected"), + [ + pytest.param(" texte ", "texte", id="plain-text"), + pytest.param(" <b>x</b> ", "**x**", id="html"), + pytest.param("<br>x<br>", "x", id="html-with-edge-breaks"), + ], +) +def test_the_result_is_stripped_on_both_paths(value: str, expected: str) -> None: + """Edge whitespace is never content, and a digest must not depend on it.""" + + assert read(value) == expected + + +@pytest.mark.parametrize( + ("html", "expected"), + [ + pytest.param( + '<h2>Titre <img src="https://example.org/i.png" alt="logo"></h2>', + "## Titre ![logo](https://example.org/i.png)", + id="heading", + ), + pytest.param( + "<table><tr><th>a</th></tr><tr><td>" + '<img src="https://example.org/i.png" alt="x"></td></tr></table>', + "| a |\n| --- |\n| ![x](https://example.org/i.png) |", + id="table-cell", + ), + ], +) +def test_an_image_in_a_heading_or_a_cell_stays_an_image( + html: str, expected: str +) -> None: + """``markdownify`` reduced these to their alt text; python-markdown needs not.""" + + assert read(html) == expected + assert_survives(html) + + +def test_a_newline_in_text_is_a_space() -> None: + """HTML displays a source newline as a space; ``nl2br`` would break the line. + + Outlook wraps its HTML source mid-sentence, so every e-mail created + ticket used to gain line breaks where the reader saw none. + """ + + html = "<p>Bonjour,\nle serveur\nest redémarré.</p>" + + assert read(html) == "Bonjour, le serveur est redémarré." + assert_survives(html) + + +@pytest.mark.parametrize( + ("html", "expected"), + [ + pytest.param("<p>a<br> b<br> c</p>", "a \nb \nc", id="space-after-break"), + pytest.param("<p>a <br>b</p>", "a \nb", id="space-before-break"), + pytest.param("<p><strong>a<br></strong>b</p>", "**a** \nb", id="edge-break"), + ], +) +def test_a_line_break_has_one_spelling(html: str, expected: str) -> None: + """Whitespace around ``<br>`` is not displayed, so it is not kept either. + + Without this the first read of ``a<br> b`` was ``a \\n b`` and the + second ``a \\nb``: the same body, read twice, disagreeing. + """ + + assert read(html) == expected + assert_survives(html) + + +# --------------------------------------------------------------------------- +# The degraded path spells text the same way +# --------------------------------------------------------------------------- + + +def test_a_body_too_deep_to_convert_is_still_literal_safe() -> None: + """The stripped text is Markdown too, and it is escaped like the rest.""" + + html = "<div>" * 600 + r"<p>__init__ \\serveur</p><p># pas un titre</p>" + + markdown = read(html) + + assert markdown == "\\_\\_init\\_\\_ \\\\\\serveur\n\n\\# pas un titre" + assert "__init__ \\\\serveur" in render(markdown) + + +@pytest.mark.parametrize("seed", range(4)) +def test_stripped_text_renders_as_itself(seed: int) -> None: + """Any text, line by line, renders as exactly that text. + + The degraded path's contract, over generated text dense with the + characters python-markdown treats as syntax. + """ + + rng = random.Random(seed) + for _ in range(60): + lines = [_text(rng, rng.randint(1, 6)) for _ in range(rng.randint(1, 4))] + text = "\n".join(line for line in lines if line) + if not text: + continue + + soup = BeautifulSoup(render(conversion._literal_markdown(text)), "html.parser") + for line_break in soup.find_all("br"): + line_break.replace_with("\u2029") + shown = soup.get_text().replace("\n", " ").replace("\u2029", "\n") + + assert [" ".join(line.split()) for line in shown.split("\n")] == [ + " ".join(line.split()) for line in text.split("\n") + ], text + + +# --------------------------------------------------------------------------- +# The property, over realistic bodies and over a fuzzer +# --------------------------------------------------------------------------- + +#: Ticket bodies shaped the way GLPI's editor and mail collector store them. +REALISTIC = [ + "<p>Bonjour,</p><p>Le PC du poste 12 ne démarre plus depuis ce matin.</p>", + "<p>Bonjour,<br>Le PC ne démarre plus.<br>Cordialement,<br>Jean</p>", + "<p>Merci de <strong>redémarrer</strong> le serveur <em>avant</em> 18h.</p>", + "<p>Étapes :</p><ul><li>ouvrir la session</li><li>lancer Outlook</li></ul>", + "<table><thead><tr><th>Poste</th><th>IP</th></tr></thead>" + "<tbody><tr><td>PC12</td><td>10.0.0.12</td></tr></tbody></table>", + "<h2>Contexte</h2><p>Migration du serveur.</p>", + '<p>Voir <a href="https://example.org/doc">https://example.org/doc</a></p>', + '<p>Voir <a title="doc" href="https://example.org/doc">la doc</a></p>', + '<p><a href="https://example.org/wiki/Test_(informatique)">wiki</a></p>', + "<p>prix 5*3 et note * importante</p>", + "<p>appuyer sur <Entrée> puis valider</p>", + "<p>Service R&D, bâtiment A & B</p>", + "<p>Montant : 12 000 €</p>", + '<p><span style="color: #e03e2d;">URGENT</span> <u>à traiter</u></p>', + '<p>Cordialement</p><p><img src="https://example.org/logo.png" alt="Logo"></p>', + "<p>Bonjour</p><blockquote><p>Message d'origine</p></blockquote>", + '<pre>Traceback (most recent call last):\n File "x.py", line 1\nError</pre>', + "<p>Lancer <code>ipconfig /all</code> puis envoyer.</p>", + "<p>Fichier mon_fichier_final.docx et __init__</p>", + "<p># pas un titre</p><p>- pas une liste</p>", + '<p>Contact <a href="mailto:support@example.org">support@example.org</a></p>', + "<div>Bonjour,</div><div><br></div><div>Le serveur est down.</div>", + "<p>* point un<br>* point deux</p>", + "<p><https://example.org/x></p>", + "<p>[INFO] tâche [1] terminée</p>", + "<p>[1]: https://example.org/note</p>", + "<p>m<sup>2</sup></p>", + r"<p>Chemin C:\Users\jdupont\Desktop</p>", + "<p>Merci 👍</p>", + "<p>Titre<br>=====</p>", + "<p>---</p><p>signature</p>", + "<p>1) un<br>2) deux</p>", + "<p><support@example.org></p>", + "<p><!-- note --></p>", + "<ul><li><p>point un</p><p>suite du point</p></li><li>point deux</li></ul>", + "<blockquote><ul><li>a</li><li>b</li></ul></blockquote>", + "<p>Nom : ______ Prénom : ______</p>", + "<p>~~pas barré~~</p>", + "<p>a | b | c</p>", + "<p>> pas une citation</p>", + "<p>+ un<br>+ deux</p>", + "<p>Cordialement,<br>Jean Dupont<br>--<br>Service IT</p>", + "<p>Merci<br>-----------<br>Jean Dupont</p>", + "<p>calcul 5 * 3 * 2 = 30</p>", + r"<p>voir \\srv\partage\__archive__\2026</p>", + "<p>Le 30/09, Jean a écrit :<br>> merci<br>> cordialement</p>", + r"<p>Accès au partage \\serveur\compta\2026 refusé</p>", + r"<p>Dossier C:\_temp\logs</p>", + "<p>voir la note [1]</p><p>[1]: https://example.org/note</p>", + "<ul><li>Réseau<ul><li>switch 3</li><li>borne wifi</li></ul></li></ul>", + "<p><b>Important</b> : voir <i>ci-dessous</i></p>", + '<p><font color="red">rouge</font> <span style="font-size:14px">texte</span></p>', + "<p>module __init__ et _x_</p>", + "<p>1) un<br>2) deux</p><pre>ligne 1\n ligne indentée</pre>", + "<p>Suite à la mise à jour, <strong>3 postes</strong> ne se connectent plus :" + "</p><ul><li>PC12 (salle 3)</li><li>PC14 – <em>poste d'accueil</em></li>" + '</ul><p>Voir <a href="https://example.org/kb/42">https://example.org/kb/42</a>' + ' et le <a href="https://example.org/kb/43" title="KB 43">KB 43</a>.</p>', +] + + +@pytest.mark.parametrize("html", REALISTIC) +def test_realistic_bodies_display_the_same_after_the_round_trip(html: str) -> None: + assert_survives(html) + + +#: Literal text snippets: every character python-markdown treats as syntax, +#: alone and in the shapes that trigger it, beside ordinary words so that +#: each lands at line starts, at word boundaries and inside words. +_SPECIAL = ( + "\\ \\\\ \\* \\_ \\[ \\# \\. \\` \\\\serveur * ** *** _ __ ___ ______ __init__ " + "_x_ ` `` ``` ~ ~~ ~~~ [ ] [x] [x](y) ![x](y) [1]: [1]:https://example.org/n " + "( ) (y) ! # ## #4521 > - -- --- + . 1. 2026. = == === | a|b --- < <b> </b> " + "<Enter> <!-- --> <https://example.org> <a@b.c> <3 <= & & < A " + "A ᆩ © © &D; { } : \" '" +).split(" ") +_WORDS = ( + "alpha beta snake_case C:\\Temp R&D x mot 5*3 a_b é fichier_de_test_v2.xlsx" +).split(" ") + + +#: Code content. A numeric reference with no semicolon is left out: inside an +#: indented code block python-markdown's raw-HTML pass adds the semicolon +#: (``test_a_numeric_reference_in_nested_code_gains_a_semicolon``). +_CODE_SPECIAL = [special for special in _SPECIAL if not special.startswith("&#")] + + +def _text(rng: random.Random, pieces: int, specials: list[str] = _SPECIAL) -> str: + out = [] + for _ in range(pieces): + out.append(rng.choice(specials) if rng.random() < 0.45 else rng.choice(_WORDS)) + out.append(rng.choice(["", " ", " ", " "])) + return "".join(out).strip() + + +def _code_text(rng: random.Random, pieces: int) -> str: + return escape(_text(rng, pieces, _CODE_SPECIAL), quote=False) + + +def _inline_html(rng: random.Random, formatted: bool = False, depth: int = 0) -> str: + """Inline HTML that Markdown can express. + + No emphasis inside emphasis, no link inside a link, and a word on either + side of every code element: python-markdown mis-pairs nested ``*`` runs + across constructs and merges two adjacent code spans, neither of which + involves literal text -- they are recorded in the round-trip inventory. + """ + + parts = [] + for _ in range(rng.randint(1, 4)): + kind = rng.random() + if kind < 0.55 or depth > 2: + parts.append(escape(_text(rng, rng.randint(1, 4)), quote=False)) + elif kind < 0.65 and not formatted: + parts.append(f" <strong>{_inline_html(rng, True, depth + 1)}</strong> ") + elif kind < 0.75 and not formatted: + parts.append(f" <em>{_inline_html(rng, True, depth + 1)}</em> ") + elif kind < 0.82 and depth == 0: + inner = _inline_html(rng, formatted, depth + 1) + parts.append( + f'<a href="https://example.org/{rng.randint(1, 9)}">{inner}</a>' + ) + elif kind < 0.88: + code = escape(rng.choice(_WORDS + _SPECIAL[:20]), quote=False) + parts.append(f" mot <code>{code}</code> mot ") + elif kind < 0.94: + alt = escape(_text(rng, rng.randint(0, 2)), quote=True) + parts.append( + f'<img src="https://example.org/{rng.randint(1, 9)}.png" alt="{alt}">' + ) + else: + parts.append(f"<span>{_inline_html(rng, formatted, depth + 1)}</span>") + parts.append(rng.choice(["", " ", " "])) + return "".join(parts).strip() or "mot" + + +def _paragraph(rng: random.Random) -> str: + return "<br>".join(_inline_html(rng) for _ in range(rng.randint(1, 3))) + + +def _list_html(rng: random.Random, depth: int) -> str: + """A list whose items hold text, and sometimes code, a list, and more text. + + The code goes before an item's nested list, never right after it: + python-markdown has no way to write a code block there + (``test_code_right_after_a_nested_list_keeps_its_text``). + """ + + tag = rng.choice(["ul", "ol"]) + items = [] + for _ in range(rng.randint(1, 3)): + shape = rng.random() + if depth < 3 and shape < 0.1: + # An item that opens with a list: one line holds both bullets. + items.append(f"<li>{_list_html(rng, depth + 1)}</li>") + continue + inner = f"<p>{_paragraph(rng)}</p>" if shape < 0.3 else _paragraph(rng) + if rng.random() < 0.1: + inner += f"<pre>{_code_text(rng, 3)}</pre>" + if depth < 3 and rng.random() < 0.3: + inner += _list_html(rng, depth + 1) + after = rng.random() + if after < 0.15: + inner += _inline_html(rng) + elif after < 0.25: + inner += f"<p>{_paragraph(rng)}</p>" + elif after < 0.3: + inner += f"<blockquote><p>{_paragraph(rng)}</p></blockquote>" + items.append(f"<li>{inner}</li>") + return f"<{tag}>{''.join(items)}</{tag}>" + + +def _block(rng: random.Random, depth: int = 0) -> str: + kind = rng.random() + if kind < 0.35 or depth > 1: + return f"<p>{_paragraph(rng)}</p>" + if kind < 0.45: + level = rng.randint(1, 3) + return f"<h{level}>{_inline_html(rng)}</h{level}>" + if kind < 0.60: + return _list_html(rng, depth + 1) + if kind < 0.70: + return f"<blockquote>{_block(rng, depth + 1)}</blockquote>" + if kind < 0.80: + head = "".join(f"<th>{_inline_html(rng)}</th>" for _ in range(2)) + rows = "".join( + "<tr>" + + "".join(f"<td>{_inline_html(rng)}</td>" for _ in range(2)) + + "</tr>" + for _ in range(rng.randint(1, 2)) + ) + return f"<table><tr>{head}</tr>{rows}</table>" + if kind < 0.87: + return f"<pre>{_code_text(rng, 4)}</pre>" + return f"<div>{_paragraph(rng)}</div>" + + +def _document(rng: random.Random) -> str: + """A body of one to four blocks. + + python-markdown merges two adjacent lists of one type, or two adjacent + block quotes, into one -- a limitation of the format -- so a paragraph + separates them. + """ + + blocks: list[str] = [] + for _ in range(rng.randint(1, 4)): + block = _block(rng) + mergeable = block[:4] in {"<ul>", "<ol>", "<blo"} + if blocks and mergeable and block[:4] == blocks[-1][:4]: + blocks.append("<p>mot</p>") + blocks.append(block) + return "".join(blocks) + + +@pytest.mark.parametrize("seed", range(8)) +def test_generated_bodies_display_the_same_after_the_round_trip(seed: int) -> None: + """The property over a seeded fuzzer, 50 bodies a seed.""" + + rng = random.Random(seed) + for _ in range(50): + assert_survives(_document(rng)) + + +# --------------------------------------------------------------------------- +# What the escaping rests on +# --------------------------------------------------------------------------- + + +def test_python_markdown_undoes_every_backslash_the_reader_writes() -> None: + """A backslash escape is only safe if python-markdown removes it again. + + Measured on 3.10.3 with the four extensions: ``\\``, backtick, ``*``, + ``_``, ``{``, ``}``, ``[``, ``]``, ``(``, ``)``, ``>``, ``#``, ``+``, + ``-``, ``.``, ``!`` and ``|``. ``=`` and ``~`` are not among them, which + is why those two are spelled as character references instead. + """ + + escapable = set(Markdown(extensions=conversion._MARKDOWN_EXTENSIONS).ESCAPED_CHARS) + backslashed = set(conversion._LITERAL) - set(conversion._SPELLED_OUT) | {"\\"} + + assert backslashed <= escapable + assert not {"=", "~", "<", "&"} & escapable + + +def test_a_body_holding_the_private_stand_ins_converts_like_any_other() -> None: + """Literal text is carried in private-use characters until it is spelled. + + A body that already holds one of them -- an icon font maps symbols + there -- must not be confused with them, so the converter picks a block + the body does not use. + """ + + first_block = [chr(0xF0000 + offset) for offset in range(32)] + html = "<p>" + "".join(first_block) + " __init__ et #4521</p>" + + markdown = read(html) + + assert "".join(first_block) in markdown + assert markdown.endswith("\\_\\_init\\_\\_ et #4521") diff --git a/glpi_python_client/testing/tests/test_content_roundtrip.py b/glpi_python_client/testing/tests/test_content_roundtrip.py index 5f685b1..fbe6801 100644 --- a/glpi_python_client/testing/tests/test_content_roundtrip.py +++ b/glpi_python_client/testing/tests/test_content_roundtrip.py @@ -171,6 +171,15 @@ def _lossy(reason: str) -> pytest.MarkDecorator: pytest.param("The snake_case name.", id="underscore"), pytest.param("5 * 3 = 15", id="asterisk"), pytest.param("# Title\n\n- alpha\n- beta\n\nClosing **note**.", id="mixed"), + pytest.param("- alpha\n - inner\n- beta", id="nested-list"), + pytest.param("1. one\n 1. inner\n2. two", id="nested-numbered-list"), + pytest.param("intro\n\n3. three\n4. four", id="numbered-from-three"), + # Literal text the reader escapes reads back escaped the same way, so a + # body that carries it is a fixed point too. + pytest.param(r"\#4521: module \_\_init\_\_.", id="escaped-literals"), + pytest.param(r"Share \\\server\share and C:\Temp.", id="backslashes"), + pytest.param(r"Press <Enter> and see \*x\*.", id="escaped-markup"), + pytest.param("| cmd |\n| --- |\n| ps aux \\| grep java |", id="pipe-in-a-cell"), pytest.param( "line one\nline two", id="soft-newline", @@ -180,14 +189,6 @@ def _lossy(reason: str) -> pytest.MarkDecorator: "equivalent and stable after one cycle; see issue #32." ), ), - pytest.param( - "- alpha\n - inner\n- beta", - id="nested-list", - marks=_lossy( - "markdownify indents nested items by 2 spaces; python-markdown " - "needs 4 to keep the nesting, so a second cycle flattens it." - ), - ), pytest.param( "```python\nx = 1\n```", id="fence-with-language", From 7ea833a884765a94d07ee27c2f246118c7c26355 Mon Sep 17 00:00:00 2001 From: baraline <antoine.guillaume45@gmail.com> Date: Thu, 1 Oct 2026 01:27:44 +0200 Subject: [PATCH 2/7] perf(content): keep every literal-text pass linear, and test every guard Four passes of the literal-text reader cost quadratic time on a body dense with one character, measured on 20,000 repetitions: unclosed brackets (settle_brackets counted forward from each [), words opening with an underscore (the pairing check compared every opener with every later run), backtick runs (_code_spans rebuilt the whole text after every span) and a long numbered list (each bullet counted its previous siblings, which markdownify's own converter did too -- 34 s before this change). They took 8 to 117 seconds each and take well under one now: - brackets are matched right to left against a stack of unclosed ], an escaped [ handing its ] back, which is python-markdown's own count; - underscore pairing is one pass remembering whether an opener was seen; - _code_spans searches on from each span's end, rebuilding the text only for the escaped-backslash run python-markdown consumes on its own; - bullets are numbered once per list, and a list's shown items are counted once. A complexity test holds each of the four under a ten-second budget. Mutation testing, in a disposable worktree, over 65 mutants of the reader's guards left eight alive before this commit. Six were test gaps, now closed: a backtick touching a generated code span, an escaped bracket in link text that python-markdown does not count, seven underscores mid-sentence, asterisks in an item and its nested item, a style block between a nested list and a code block, and a line break in a cell or a heading. Two were dead code, now gone: the nested list's blank line no longer checks for a following sibling -- the item strips it when nothing follows -- and from_transport no longer strips a result the document converter has already stripped, before settling it, which is where the strip matters. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> --- glpi_python_client/content/conversion.py | 137 +++++++++++------- .../content/tests/test_literal_text.py | 78 ++++++++++ 2 files changed, 166 insertions(+), 49 deletions(-) diff --git a/glpi_python_client/content/conversion.py b/glpi_python_client/content/conversion.py index cdbe9c9..8542881 100644 --- a/glpi_python_client/content/conversion.py +++ b/glpi_python_client/content/conversion.py @@ -1068,7 +1068,7 @@ def hide_escapes(self) -> None: continue position += 1 - def settle_brackets(self, start: int, end: int) -> None: + def settle_brackets(self, units: Iterable[tuple[int, int]]) -> None: """Escape every literal ``[`` that would open a link or an image. A ``[`` opens one when its balanced ``]`` is followed straight away @@ -1076,19 +1076,32 @@ def settle_brackets(self, start: int, end: int) -> None: escaping an inner ``[`` rebalances the brackets of an outer one; one pass is enough, since the brackets to a ``[``'s right are settled before it is. + + The pass keeps the ``]`` not yet closed by a ``[`` to their right on + a stack, so each ``[`` finds its partner on top of it: the first + ``]`` python-markdown's own count would stop at. An escaped ``[`` + hands its ``]`` back, to be closed by one further left. That keeps + the pass linear, where counting forward from each ``[`` cost a body + of unclosed brackets quadratic time. """ working = list(self.working()) - openers = [ - position - for position in range(start, end) - if self.view[position] == "[" and self.live(position) - ] - for opener in reversed(openers): - close = _balanced_close(working, opener + 1, end) - if close is not None and close + 1 < end and working[close + 1] == "(": - self.escape(opener) - working[opener] = _NEUTRAL + for start, end in units: + closers: list[int] = [] + for position in range(end - 1, start - 1, -1): + char = working[position] + if char == "]": + closers.append(position) + elif char == "[" and closers: + close = closers.pop() + if ( + self.live(position) + and close + 1 < end + and working[close + 1] == "(" + ): + self.escape(position) + working[position] = _NEUTRAL + closers.append(close) def settle_asterisks(self, units: Iterable[tuple[int, int]]) -> None: """Escape literal ``*`` wherever a unit holds two possible delimiters. @@ -1161,8 +1174,7 @@ def settle_inline(self, units: list[tuple[int, int]]) -> None: self.settle_bang() self.settle_backticks(units) self.hide_escapes() - for start, end in units: - self.settle_brackets(start, end) + self.settle_brackets(units) self.settle_automail() self.settle_asterisks(units) self.settle_underscores(units) @@ -1204,29 +1216,30 @@ def _underscores_pair( return True if len(runs) < 2: return False - for index, (begin, finish) in enumerate(runs): - if finish - begin >= 3: + if any(finish - begin >= 3 for begin, finish in runs): + return True + opened = False + for begin, finish in runs: + if opened and (finish == end or not _WORD.match(working[finish])): return True - if begin != start and _WORD.match(working[begin - 1]): - continue - for later_begin, later_finish in runs[index + 1 :]: - if later_finish - later_begin >= 3: - return True - if later_finish == end or not _WORD.match(working[later_finish]): - return True + if begin == start or not _WORD.match(working[begin - 1]): + opened = True return False def _code_spans(markdown: str) -> list[tuple[tuple[int, int], tuple[int, int]]]: """Return the code spans python-markdown finds, as delimiter spans. - Replays its backtick processor: each match is replaced before the next - search, and a run of escaped backslashes before a backtick is consumed - on its own, which leaves the backtick after it free to open. + Replays its backtick processor, which replaces each match with a + placeholder and searches on from just after it. After a code span that + is the same as searching on from the span's end: the one look-behind + there is ``(?<!\\\\)``, and the span ends in a backtick. A run of escaped + backslashes before a backtick is consumed on its own, though, and its + placeholder is what leaves the backtick after it free to open -- so + that run, and only that, is neutralised in the text searched. """ spans: list[tuple[tuple[int, int], tuple[int, int]]] = [] - working = list(markdown) text = markdown position = 0 while True: @@ -1241,13 +1254,11 @@ def _code_spans(markdown: str) -> list[tuple[tuple[int, int], tuple[int, int]]]: (found.end(3), found.end(3) + size), ) ) - replaced = range(found.start(), found.end()) + position = found.end() else: - replaced = range(found.start(1), found.end(1)) - for index in replaced: - working[index] = _NEUTRAL - text = "".join(working) - position = found.start() + start, end = found.span(1) + text = text[:start] + _NEUTRAL * (end - start) + text[end:] + position = end class _Link(NamedTuple): @@ -1720,14 +1731,23 @@ def _ends_in_a_list(node: PageElement) -> bool | None: return False -def _bullet(item: Tag) -> str: - """Return the marker the reader writes before a list item.""" +def _bullet(item: Tag, numbers: dict[int, int]) -> str: + """Return the marker the reader writes before a list item. + + ``numbers`` holds each ordered item's position among its list's items, + filled for a whole list the first time one of its items is asked for: + counting an item's previous siblings instead cost a long list quadratic + time. + """ parent = item.parent if parent is not None and parent.name == "ol": + if id(item) not in numbers: + for index, entry in enumerate(parent.find_all("li", recursive=False)): + numbers[id(entry)] = index start = str(parent.get("start") or "") first = int(start) if start.isdigit() else 1 - return f"{first + len(item.find_previous_siblings('li'))}. " + return f"{first + numbers[id(item)]}. " return "- " @@ -1773,14 +1793,14 @@ def _shown_items(element: Tag) -> int: return sum(_last_shown(item) is not None for item in items) -def _line_lead(element: Tag) -> tuple[str, ...]: +def _line_lead(element: Tag, numbers: dict[int, int]) -> tuple[str, ...]: """Return the bullets written on ``element``'s first line, its own included. ``- - -`` is a rule to python-markdown, so the first line of a block is - checked against its whole line. + checked against its whole line. ``numbers`` is :func:`_bullet`'s. """ - return tuple(_bullet(item) for item in _line_items(element)) + return tuple(_bullet(item, numbers) for item in _line_items(element)) def _code_block_fits(element: Tag) -> bool: @@ -1857,6 +1877,17 @@ def __init__(self, stand_ins: _StandIns, **options: Any) -> None: self.stand_ins = stand_ins #: The list items whose Markdown ends in a blank line, by ``id``. self._loose_items: set[int] = set() + #: Each ordered item's position in its list, by ``id`` (:func:`_bullet`). + self._numbers: dict[int, int] = {} + #: How many of a list's items display something, by the list's ``id``. + self._shown: dict[int, int] = {} + + def _shown_items(self, element: Tag) -> int: + """Return :func:`_shown_items` for a list, counted once per list.""" + + if id(element) not in self._shown: + self._shown[id(element)] = _shown_items(element) + return self._shown[id(element)] def _inherited(self, name: str) -> Callable[..., Any]: """Return ``markdownify``'s own converter ``name``. @@ -1911,20 +1942,28 @@ def _settle( ) def convert__document_(self, el: Tag, text: str, parent_tags: set[str]) -> str: + """Settle the document, stripped first: edge whitespace is never content. + + Stripped before it is settled, not after, because where the first + line starts decides what it would be read as: ``" # x"`` is not a + heading until the space goes. It is also why the Markdown is stripped + at all -- a digest of a body must not depend on its edges. + """ + text = self.stand_ins.settle_breaks(text).strip() return _settle_block(text, self.stand_ins, in_list=False, top_level=True) def convert_p(self, el: Tag, text: str, parent_tags: set[str]) -> str: if self._deferred(parent_tags): return str(self._inherited("convert_p")(el, text, parent_tags)) - lead = _line_lead(el) if "li" in parent_tags else () + lead = _line_lead(el, self._numbers) if "li" in parent_tags else () text = self._settle(text, parent_tags, lead=lead) return f"\n\n{text}\n\n" if text else "" def convert_div(self, el: Tag, text: str, parent_tags: set[str]) -> str: if self._deferred(parent_tags): return str(self._inherited("convert_div")(el, text, parent_tags)) - lead = _line_lead(el) if "li" in parent_tags else () + lead = _line_lead(el, self._numbers) if "li" in parent_tags else () text = self._settle(text, parent_tags, lead=lead) return f"\n\n{text}\n\n" if text else "" @@ -1960,7 +1999,7 @@ def convert_blockquote(self, el: Tag, text: str, parent_tags: set[str]) -> str: return "\n\n{}\n\n".format("\n".join(lines)) def convert_li(self, el: Tag, text: str, parent_tags: set[str]) -> str: - bullet = _bullet(el) + bullet = _bullet(el, self._numbers) spaced = False if self._deferred(parent_tags): text = (text or "").strip() @@ -1969,7 +2008,7 @@ def convert_li(self, el: Tag, text: str, parent_tags: set[str]) -> str: text = _LEADING_BLANK_LINES.sub("", text).rstrip(_EDGE) parent = el.parent deep = _deep_list(parent) - spaced = deep and parent is not None and _shown_items(parent) > 1 + spaced = deep and parent is not None and self._shown_items(parent) > 1 # The item after a loose one is loose too: python-markdown reads # it as the first item of a new block, and wraps its text. So is # an item that shares its first line with another's bullet: a @@ -1980,7 +2019,7 @@ def convert_li(self, el: Tag, text: str, parent_tags: set[str]) -> str: loose = deep or stacked or id(previous) in self._loose_items if not _opens_with_a_block(el): text = _loosened(text, loose) - lead = _line_lead(el) + lead = _line_lead(el, self._numbers) text = self._settle(text, parent_tags | {"li"}, lead=lead) if not text: return "\n" @@ -2126,19 +2165,18 @@ def convert_pre(self, el: Tag, text: str, parent_tags: set[str]) -> str: return f"\n\n{fence}\n{text}\n{fence}\n\n" def convert_list(self, el: Tag, text: str, parent_tags: set[str]) -> str: - """End a nested list with a blank line when its item goes on after it. + """End a nested list with a blank line, for whatever its item holds next. ``markdownify`` ends a list inside a list item with no line break at all, so the item's text after it joined the list's last item, and a - quote after it became that item's lazy continuation. + quote after it became that item's lazy continuation. When nothing + follows, the item strips the blank line with the rest of its edge. """ markdown = str(self._inherited("convert_list")(el, text, parent_tags)) if "li" not in parent_tags or self._deferred(parent_tags): return markdown - if any(_last_shown(sibling) is not None for sibling in el.next_siblings): - return markdown + "\n\n" - return markdown + return markdown + "\n\n" convert_ul = convert_list convert_ol = convert_list @@ -2299,7 +2337,8 @@ def from_transport(value: object) -> str: "Could not convert GLPI HTML content to Markdown " f"({type(exc).__name__}: {exc})." ) from exc - return str(markdown).strip() + # Already stripped: the document converter strips before it settles. + return markdown @staticmethod def to_transport(value: object) -> str: diff --git a/glpi_python_client/content/tests/test_literal_text.py b/glpi_python_client/content/tests/test_literal_text.py index 8675343..ddf9fee 100644 --- a/glpi_python_client/content/tests/test_literal_text.py +++ b/glpi_python_client/content/tests/test_literal_text.py @@ -32,6 +32,7 @@ from __future__ import annotations import random +import time from html import escape import pytest @@ -465,6 +466,11 @@ def test_the_measured_misreadings_read_back_as_their_text( ), # Emphasis delimiters. pytest.param("<p>_______</p>", r"\_" * 7, id="seven-underscores"), + pytest.param( + "<p>Code : _______ fin</p>", + "Code : " + r"\_" * 7 + " fin", + id="seven-underscores-mid-sentence", + ), pytest.param( "<p>a*b*c et 5*3 <strong>gras</strong></p>", r"a\*b\*c et 5\*3 **gras**", @@ -550,6 +556,11 @@ def test_the_measured_misreadings_read_back_as_their_text( r"[voir \[x\](y)](https://example.org/u)", id="link-syntax-in-link-text", ), + pytest.param( + '<p>[a <a href="https://example.org/u">[b</a> ](https://example.org/x)</p>', + r"\[a [\[b](https://example.org/u) ](https://example.org/x)", + id="escaped-bracket-in-link-text-is-not-counted", + ), pytest.param( '<p><img src="https://example.org/c.png" alt="capture [1].png"></p>', "![capture [1].png](https://example.org/c.png)", @@ -610,6 +621,16 @@ def test_the_measured_misreadings_read_back_as_their_text( r"a\`b\`c et `x` puis `", id="backtick-pair", ), + pytest.param( + "<p><code>x</code>`y</p>", + r"`x`\`y", + id="backtick-touching-a-code-span", + ), + pytest.param( + "<p>`<code>x</code> y</p>", + r"\``x` y", + id="backtick-opening-onto-a-code-span", + ), pytest.param( "<table><tr><th>a</th><th>b</th></tr>" "<tr><td>x`y</td><td>`z</td></tr></table>", @@ -664,6 +685,11 @@ def test_literal_text_is_escaped_where_python_markdown_would_read_it( "intro\n\n3. trois\n4. quatre", id="ordered-list-starting-at-three", ), + pytest.param( + "<ul><li>5*3<ul><li>2*4</li></ul></li></ul>", + "- 5*3\n - 2*4", + id="asterisks-an-item-and-its-nested-item-cannot-pair", + ), pytest.param( "<ul><li><p>point un</p><p>suite du point</p></li><li>point deux</li></ul>", "- point un\n\n suite du point\n\n- point deux", @@ -767,6 +793,19 @@ def test_blocks_inside_a_list_item_stay_in_it(html: str, expected: str) -> None: pytest.param( "<h2><blockquote>cité</blockquote></h2>", "## cité", id="quote-in-a-heading" ), + pytest.param( + "<table><tr><th>a</th></tr><tr><td>un<br>deux</td></tr></table>", + "| a |\n| --- |\n| un deux |", + id="break-in-a-cell", + ), + pytest.param( + "<h2>Titre<br>suite</h2>", "## Titre suite", id="break-in-a-heading" + ), + pytest.param( + '<p><img src="https://example.org/i.png" alt="ligne 1\n\n ligne 2"></p>', + "![ligne 1 ligne 2](https://example.org/i.png)", + id="alt-text-over-several-lines", + ), pytest.param( "<ul><li><script>x()</script><pre>code</pre></li></ul>", "- code", @@ -885,6 +924,12 @@ def test_a_preformatted_block_opening_a_list_item_keeps_its_text() -> None: "> - item\n>\n> \\#4521 code", id="in-a-quote", ), + pytest.param( + "<ul><li>item<ul><li>sous-item</li></ul><div><style>p{margin:0}</style>" + "</div><pre>#4521 code</pre></li></ul>", + "- item\n\n - sous-item\n\n \\#4521 code", + id="with-only-a-style-block-between", + ), ], ) def test_code_right_after_a_nested_list_keeps_its_text( @@ -962,6 +1007,16 @@ def test_code_right_after_a_nested_list_keeps_its_text( "written as inline code.", id="code-block-in-a-cell", ), + pytest.param( + "<table><tr><th>a</th></tr><tr><td>un<br>deux</td></tr></table>", + "A table cell holds one line: its line break is written as a space.", + id="break-in-a-cell", + ), + pytest.param( + "<h2>Titre<br>suite</h2>", + "A heading is one line: its line break is written as a space.", + id="break-in-a-heading", + ), pytest.param( "<ul><li><pre>a</pre></li></ul>", "An item's first block is its paragraph: the code is kept as text " @@ -1432,6 +1487,29 @@ def test_generated_bodies_display_the_same_after_the_round_trip(seed: int) -> No assert_survives(_document(rng)) +@pytest.mark.parametrize( + "html", + [ + pytest.param("<p>" + "[a " * 20_000 + "</p>", id="unclosed-brackets"), + pytest.param("<p>" + "_a " * 20_000 + "</p>", id="underscores-opening-words"), + pytest.param("<p>" + "``` " * 20_000 + "</p>", id="backtick-runs"), + pytest.param("<ol>" + "<li>x</li>" * 20_000 + "</ol>", id="long-numbered-list"), + ], +) +def test_a_body_dense_with_syntax_converts_in_linear_time(html: str) -> None: + """Every pass is linear, so a pathological body costs what its size does. + + Measured before the passes were made linear: 20,000 of these took from + 8 to 117 seconds each, where they take well under one now. The budget + is generous on purpose -- this guards the complexity, not the speed. + """ + + started = time.perf_counter() + read(html) + + assert time.perf_counter() - started < 10 + + # --------------------------------------------------------------------------- # What the escaping rests on # --------------------------------------------------------------------------- From 4fc3bed91a6f1d6fe95334640f96599cbb08ee23 Mon Sep 17 00:00:00 2001 From: baraline <antoine.guillaume45@gmail.com> Date: Thu, 1 Oct 2026 01:28:11 +0200 Subject: [PATCH 3/7] chore(release): 0.6.0 Literal-safe Markdown on the read path (588e97d, 7ea833a), and the markdownify floor it needs: >=1.2, since the converter subclasses MarkdownConverter and relies on the 1.x hooks -- escape(text, parent_tags), convert_*(el, text, parent_tags), convert__document_ -- which 0.13 does not have. 1.2.2 and 1.2.3 are the versions measured. The changelog records the fixes, the two behaviour changes a caller can see (a value holding one real HTML element is read as HTML throughout, write models included; Markdown read from GLPI changes once for any body the old reader let through as syntax) and the inventory of what Markdown cannot carry. The user guide and the API reference say what the escaping is and is not, with an example checked against the code, and the two skills that describe .content now say not to strip the backslashes. A correction to 7ea833a's message: nine of the 65 mutants survived before it, not eight, and seven were test gaps -- the seventh is alt text over several lines, which a blank line in it split out of its paragraph once the collapse was removed. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> --- CHANGELOG.md | 119 ++++++++++++++++++ docs/api_reference.rst | 8 ++ docs/user_guide.rst | 35 ++++++ glpi_python_client/__init__.py | 2 +- pyproject.toml | 4 +- skills/glpi-client-setup/SKILL.md | 2 +- skills/glpi-document-workflow/SKILL.md | 2 +- skills/glpi-knowledge-base/SKILL.md | 2 +- skills/glpi-plugin-fields/SKILL.md | 2 +- skills/glpi-reporting-and-context/SKILL.md | 2 +- skills/glpi-team-members/SKILL.md | 2 +- skills/glpi-ticket-timeline/SKILL.md | 3 +- skills/glpi-ticket-workflow/SKILL.md | 3 +- .../glpi-user-location-provisioning/SKILL.md | 2 +- 14 files changed, 176 insertions(+), 12 deletions(-) diff --git a/CHANGELOG.md b/CHANGELOG.md index b1dc065..6837c53 100644 --- a/CHANGELOG.md +++ b/CHANGELOG.md @@ -4,6 +4,125 @@ All notable changes to this project are documented in this file. The format is based on [Keep a Changelog](https://keepachangelog.com/en/1.1.0/). +## 0.6.0 — 2026-10-01 + +### Fixed + +- **Literal text came back as Markdown syntax.** `from_transport` handed + `markdownify`'s output on as Markdown, and `markdownify` escapes nothing + it is not asked to, so text a user typed into GLPI changed meaning as + soon as anything rendered it. Measured through the reader into + python-markdown with the four extensions `to_transport` uses: + `\\serveur\compta` lost a backslash, `__init__` and `______` became + emphasis, a `-----------` line under text made the text a heading, + `* point` and `> merci` lines made a list and a quote, `#4521` at the + start of a line made a heading, `[1]: https://...` was consumed as a + reference definition, and a `|` in a table cell dropped the rest of the + row. + + The Markdown now spells literal text so python-markdown reads it as + text. A character is escaped **exactly where python-markdown would + otherwise read it as syntax, and nowhere else**: ordinary prose -- + `fichier_de_test_v2.xlsx`, `C:\Temp\logs`, `R&D`, a `#` or a `-` + mid-sentence, `5 * 3` -- comes back byte for byte as before. The escape + is a backslash wherever python-markdown removes one (`\_`, `\#`, `\\`, + `\[`, `\|`, ...), and a character reference where no backslash works: + `<` for a `<` that would open a tag or an e-mail autolink, `&` + for an `&` that would start a reference, `=` and `~` for a + setext `=` underline or a `~~~` fence. A pasted URL link stays + `<https://...>`. + + The rules replay python-markdown's own: its code-span pattern, its + escape set (asked of the renderer, not copied), its bracket counting, + its emphasis pairing, the line tests for headings, rules, setext + underlines, fences, lists, quotes, table separators and reference + definitions, and its e-mail autolink, which it matches only after code + spans, links and images are set aside. They are held to two properties, + checked with an HTML parser over the measured cases, over realistic + bodies and over a seeded fuzzer: `to_transport(from_transport(html))` + **displays what `html` displays**, and the Markdown is **a fixed + point**, reading back as itself. 30,000 fuzzed bodies were measured + clean; 400 run in the suite. + +- **Nested lists flattened on the first write.** `markdownify` indents a + nested item by its bullet's width, two or three spaces, where + python-markdown nests at four, so a nested list lost a level per cycle + and a nested ordered list restarted its count. Continuation lines are + indented by four spaces, and numbering, `start` included, survives. + Around that, the blocks inside a list item now stay in it: text or a + quote after a nested list starts a block of its own instead of running + into the list's last item, a quote in an item gets the blank line + python-markdown needs, and a quote opening an item keeps its later + lines. An item python-markdown writes as loose is written loose, so the + second read equals the first; a list whose first line carries three + bullets (`- - - a`) is spaced out, since python-markdown cannot nest + anything under such a line tightly. + +- **`<script>`, `<style>` and `<title>` bodies leaked into the text.** + `strip=["script", "style"]` skips an element's own converter, which is + the one that drops its body, so Outlook's CSS arrived as prose. They are + dropped now, on the converting path, as a browser drops them. The + degraded path still keeps them, which its promise -- never less than + the converting path -- allows. + +- **Bodies marked up only with obsolete elements were read as text.** + `<font>`, `<center>`, `<big>`, `<tt>`, `<nobr>` and the rest of the HTML + standard's obsolete list now make a body HTML; `<center>` and `<dir>` + are blocks. + +- **`<s>`, `<del>` and `<strike>` became a literal `~~`.** python-markdown + has no strikethrough, so the markers showed. The words are kept and the + line is dropped -- recorded as a loss below. + +- `<pre>` inside a list item or a quote is an indented code block, since a + fence opens only at the start of a line and was read there as text; a + top-level fence is made longer than any fence line the code holds. An + image stays an image in a heading or a cell. A newline in the HTML + source is a space, as a browser shows it, rather than a line break + from `nl2br`. Whitespace around a `<br>` is dropped, so one body reads + one way. + +- The degraded rendering of a body too deep to convert is spelled through + the same rules, so it is literal-safe too. + +### Changed + +- **Mixed Markdown and HTML in one value is read as HTML.** `from_transport` + is also the validator on write models, and a value carrying a single real + HTML element takes the HTML path throughout: `**bold** <b>x</b>` keeps + its asterisks as text now, where it used to render them. Write Markdown + or HTML, not both. + +- **`markdownify>=1.2`** (was `>=0.13`). The converter is a + `MarkdownConverter` subclass, and it relies on the 1.x converter hooks -- + `escape(text, parent_tags)`, `convert_*(el, text, parent_tags)`, + `convert__document_` -- which 0.13 does not have. 1.2.2 and 1.2.3 are + the versions measured; 1.0 and 1.1 were not, so they are not claimed. + +- Markdown read from GLPI changes for any body whose text the old reader + let through as syntax, which is the point, and for nested lists. A + caller who stored the old Markdown or a digest of it sees a one-time + difference on the next read. + +### Known limitations + +Structure Markdown cannot carry, each asserted to still be a loss and to +stabilise after one cycle (`test_what_markdown_cannot_carry`): +strikethrough and underline (words kept, line lost); two adjacent lists, +or two adjacent quotes, read back as one; two `<br>` in a row are a +paragraph break; adjacent code spans merge; strong inside emphasis loses +its bold; a table without a header row gains an empty one, and one +without a body gains an empty row; a table cell or a heading holds one +line, so a list, a code block or a line break in one is flattened; a +`<pre>` that opens a list item, or directly follows a list inside the same +item or quote, is kept as text, because python-markdown reads no code +block there. Inside a list item or a quote, a numeric reference with no +semicolon in code gains one (`test_a_numeric_reference_in_nested_code_gains_a_semicolon`). + +Reading costs more: about 1.4x the old reader on realistic bodies and up +to 4x on bodies dense with syntax characters, every pass linear in the +body's size. + ## 0.5.0 — 2026-09-08 ### Fixed diff --git a/docs/api_reference.rst b/docs/api_reference.rst index 0c62a1b..e45db87 100644 --- a/docs/api_reference.rst +++ b/docs/api_reference.rst @@ -113,6 +113,14 @@ conversion happens, which buys two things: listing records costs nothing per body, and a body that cannot be converted no longer stops the rest of its page being read. +The Markdown spells a body's text as literal text: a character is escaped +exactly where python-markdown, with the package's four extensions, would +read it as syntax -- ``\_\_init\_\_``, ``\\\serveur``, ``\#4521`` at the +start of a line, ``<Entrée>`` -- and nowhere else, so rendering it +displays what GLPI displayed and reading that back gives the same Markdown. +A value holding one real HTML element is read as HTML throughout, write +models included. See :ref:`content-conversion`. + Very deeply nested HTML is the case worth knowing about. ``markdownify`` walks the document recursively and runs out of stack at around 494 levels of nesting. The converter does not try to predict diff --git a/docs/user_guide.rst b/docs/user_guide.rst index 9fb1b13..a90b45f 100644 --- a/docs/user_guide.rst +++ b/docs/user_guide.rst @@ -1552,6 +1552,41 @@ The whole page is built in one pass, so a single unconvertible record used to make its page-mates unreadable too. The failure is now scoped to the record whose body you actually read. +Text in a body is literal, and ``.content`` spells it so: rendering the +Markdown -- as the package does on the way back to GLPI, or with any +python-markdown using the same four extensions (``nl2br``, ``sane_lists``, +``fenced_code``, ``tables``) -- displays what GLPI displayed. Markdown has +one spelling for ``__init__`` typed by a user and for bold ``init``, so a +character is escaped exactly where python-markdown would otherwise read it +as syntax: + +.. code-block:: python + + from glpi_python_client.content import GlpiContentConverter + + GlpiContentConverter.from_transport( + "<p>Voir __init__ et \\\\serveur\\partage</p><p># pas un titre</p>" + ) + # 'Voir \\_\\_init\\_\\_ et \\\\\\serveur\\partage\n\n\\# pas un titre' + +Ordinary prose carries no escape at all -- ``fichier_de_test_v2.xlsx``, +``C:\Temp``, ``R&D``, a ``#`` or a ``-`` mid-sentence come back exactly as +typed -- and the Markdown is a fixed point: rendering it and reading it +back gives the same Markdown again. A backslash is used wherever +python-markdown removes one; ``<``, ``&``, ``=`` and ``~`` are spelled as +character references (``<``) where they would be read, since no +backslash escapes them. A pasted URL link stays ``<https://...>``. + +Two things change for a caller. A value holding a single real HTML element +is read as HTML throughout, on write models too, so ``**bold** <b>x</b>`` +keeps its asterisks as text: write either Markdown or HTML, not both. And +a few structures have no Markdown spelling, which the round-trip inventory +in ``test_literal_text.py`` records: struck-through and underlined text +keep their words but lose the line, adjacent lists or quotes merge, a line +break in a table cell or a heading becomes a space, and a ``<pre>`` that +opens a list item, or that directly follows a list inside the same item, +keeps its lines as text rather than as code. + .. note:: Deeply nested HTML is the case worth knowing about. The HTML-to-Markdown diff --git a/glpi_python_client/__init__.py b/glpi_python_client/__init__.py index 508acad..ffdb55c 100644 --- a/glpi_python_client/__init__.py +++ b/glpi_python_client/__init__.py @@ -110,7 +110,7 @@ date_window, ) -__version__ = "0.5.0" +__version__ = "0.6.0" __all__ = [ "AsyncGlpiClient", diff --git a/pyproject.toml b/pyproject.toml index b43c7a3..f945341 100644 --- a/pyproject.toml +++ b/pyproject.toml @@ -35,7 +35,7 @@ exclude = [ [project] name = "glpi-python-client" -version = "0.5.0" +version = "0.6.0" description = "A typed Python client for GLPI ITSM APIs." readme = "README.md" requires-python = ">=3.10" @@ -62,7 +62,7 @@ dependencies = [ "httpx>=0.28", "lxml>=4.9", "markdown>=3.6", - "markdownify>=0.13", + "markdownify>=1.2", "pydantic>=2.8", # Never imported by this package, and deliberately so. httpcore decides # whether it is running under asyncio or trio by probing for sniffio on diff --git a/skills/glpi-client-setup/SKILL.md b/skills/glpi-client-setup/SKILL.md index 0e2598f..5919a69 100644 --- a/skills/glpi-client-setup/SKILL.md +++ b/skills/glpi-client-setup/SKILL.md @@ -5,7 +5,7 @@ license: MIT compatibility: "Requires Python 3.10+, glpi-python-client, network access to a GLPI v2 API, and valid GLPI credentials." metadata: package: glpi-python-client - version: "0.5.0" + version: "0.6.0" --- # GLPI Client Setup diff --git a/skills/glpi-document-workflow/SKILL.md b/skills/glpi-document-workflow/SKILL.md index 4fe0272..3fe75c8 100644 --- a/skills/glpi-document-workflow/SKILL.md +++ b/skills/glpi-document-workflow/SKILL.md @@ -5,7 +5,7 @@ license: MIT compatibility: "Requires Python 3.10+, glpi-python-client, network access to the GLPI v2 API, and v1 credentials configured on the client for binary uploads." metadata: package: glpi-python-client - version: "0.5.0" + version: "0.6.0" --- # GLPI Document Workflow diff --git a/skills/glpi-knowledge-base/SKILL.md b/skills/glpi-knowledge-base/SKILL.md index 146d747..c2cbe2d 100644 --- a/skills/glpi-knowledge-base/SKILL.md +++ b/skills/glpi-knowledge-base/SKILL.md @@ -5,7 +5,7 @@ license: MIT compatibility: "Requires Python 3.10+, glpi-python-client, network access to the GLPI v2 API, and — for category writes only — a legacy v1 session (v1_base_url + v1_user_token)." metadata: package: glpi-python-client - version: "0.5.0" + version: "0.6.0" --- # GLPI Knowledge Base diff --git a/skills/glpi-plugin-fields/SKILL.md b/skills/glpi-plugin-fields/SKILL.md index e979583..f90548e 100644 --- a/skills/glpi-plugin-fields/SKILL.md +++ b/skills/glpi-plugin-fields/SKILL.md @@ -5,7 +5,7 @@ license: MIT compatibility: "Requires Python 3.10+, glpi-python-client, the GLPI Fields plugin installed server-side, and a legacy v1 session (v1_base_url + v1_user_token) — every method in this family goes over the v1 API." metadata: package: glpi-python-client - version: "0.5.0" + version: "0.6.0" --- # GLPI Plugin Fields diff --git a/skills/glpi-reporting-and-context/SKILL.md b/skills/glpi-reporting-and-context/SKILL.md index 427d595..70b9808 100644 --- a/skills/glpi-reporting-and-context/SKILL.md +++ b/skills/glpi-reporting-and-context/SKILL.md @@ -5,7 +5,7 @@ license: MIT compatibility: "Requires Python 3.10+, glpi-python-client, network access to the GLPI v2 API, and credentials allowed to read tickets, tasks, users, entities, and timeline records." metadata: package: glpi-python-client - version: "0.5.0" + version: "0.6.0" --- # GLPI Reporting And Context diff --git a/skills/glpi-team-members/SKILL.md b/skills/glpi-team-members/SKILL.md index 7f1a070..42ce77c 100644 --- a/skills/glpi-team-members/SKILL.md +++ b/skills/glpi-team-members/SKILL.md @@ -5,7 +5,7 @@ license: MIT compatibility: "Requires Python 3.10+, glpi-python-client, network access to the GLPI v2 API, and credentials allowed to manage ticket teams." metadata: package: glpi-python-client - version: "0.5.0" + version: "0.6.0" --- # GLPI Team Members diff --git a/skills/glpi-ticket-timeline/SKILL.md b/skills/glpi-ticket-timeline/SKILL.md index 7e70892..f779e79 100644 --- a/skills/glpi-ticket-timeline/SKILL.md +++ b/skills/glpi-ticket-timeline/SKILL.md @@ -5,7 +5,7 @@ license: MIT compatibility: "Requires Python 3.10+, glpi-python-client, and network access to the GLPI v2 API." metadata: package: glpi-python-client - version: "0.5.0" + version: "0.6.0" --- # GLPI Ticket Timeline @@ -120,6 +120,7 @@ await client.update_ticket_timeline_document( - `create_*` methods return new identifiers as plain `int`. `update_*` and `delete_*`/`unlink_*` return `None`. - Three enums carry this family's value vocabularies, all exported from `glpi_python_client` and all subclasses of `GlpiEnum` (itself an `IntEnum`, so a member serialises as its number and compares equal to one): `GlpiTaskState` on `PostTicketTask.state`/`PatchTicketTask.state` (`INFORMATION = 0`, `TODO = 1`, `DONE = 2` -- note `INFORMATION` is `0`, so `if task.state:` is false for it; test against `None`), `GlpiSolutionStatus` on `PostSolution.status`/`PatchSolution.status` (`NONE = 1`, `WAITING = 2`, `ACCEPTED = 3`, `REFUSED = 4`), and `GlpiTimelinePosition` on `timeline_position` (`INVALID = -1`, `NONE = 0`, `LEFT = 1`, `RIGHT = 2`, `LEFT_BIG = 3`, `RIGHT_BIG = 4`), which the followup, task and document models carry -- the solution models do not have the field at all. - Timeline `content` fields are Markdown on the Python side, not HTML. `PostFollowup`/`PostTicketTask`/`PostSolution` render Markdown to GLPI's HTML on serialisation. On `GetFollowup`/`GetTicketTask`/`GetSolution` the field is `content_html` -- the server's HTML verbatim -- and `content` is a cached property that converts it on **first read**, not on validation (`GetFollowup(content="<p>Hello <strong>world</strong></p>").content == "Hello **world**"` still holds; the wire spelling `content` is accepted as a validation alias). So `record.content` is always Markdown, but `list_ticket_followups` converts nothing until you read a body, and a body that cannot be converted no longer breaks the whole list. Authoring raw HTML on a write model is not an error but is round-tripped through the Markdown converter and can be reshaped; write Markdown. +- **`.content` spells literal text so it stays text.** A character is escaped exactly where python-markdown (with `nl2br`, `sane_lists`, `fenced_code`, `tables`) would read it as syntax: a user's `__init__` reads back as `\_\_init\_\_`, `\\serveur` as `\\\serveur`, `#4521` at a line start as `\#4521`, `<Entrée>` as `<Entrée>`. Ordinary prose -- `fichier_de_test_v2.xlsx`, `C:\Temp`, `R&D` -- carries no escape. Render the Markdown to display it; do not strip the backslashes, and do not mix Markdown and HTML in one value you write: one real HTML element makes the whole value HTML, so its `**bold**` is kept as literal asterisks. - HTML too deeply nested to walk is **stripped to text instead of converted**: `markdownify` recurses per level and dies around 494 from a shallow stack. The conversion is attempted rather than the depth predicted, so the real limit is whatever stack is left at the call site. Every character the normal rendering would have produced still appears, but structure does not: link targets, image alt text and code fencing are gone. Nothing raises. Any other conversion failure raises `GlpiContentError` -- from the `.content` read, not from the `list_*` call. Note that `GlpiTicketContext.to_markdown()` is usually the first thing to read every body, so it is where such an error surfaces. - Extra server fields (e.g. plugin keys) flow into `record.extra_payload` rather than raising. - `delete_ticket_*` and `unlink_ticket_timeline_document` accept a keyword-only `force` parameter; pass `force=True` to permanently delete. \ No newline at end of file diff --git a/skills/glpi-ticket-workflow/SKILL.md b/skills/glpi-ticket-workflow/SKILL.md index bb986b5..29502ba 100644 --- a/skills/glpi-ticket-workflow/SKILL.md +++ b/skills/glpi-ticket-workflow/SKILL.md @@ -5,7 +5,7 @@ license: MIT compatibility: "Requires Python 3.10+, glpi-python-client, network access to the GLPI v2 API, and credentials accepted by GlpiClient." metadata: package: glpi-python-client - version: "0.5.0" + version: "0.6.0" --- # GLPI Ticket Workflow @@ -74,6 +74,7 @@ ticket = PostTicket( - **`search_tickets` and every other `search_*` raise `GlpiStatusError` on a 4xx.** This changed: they used to check the response status only when the caller passed a `failure_message`, which none of the seven `search_*` helpers does, so a GLPI error body was coerced to `[]` and a malformed RSQL filter, a 403, a missing route and a genuinely empty result set were indistinguishable. `_resource_list` now checks the status on every call, so **an empty list means the server said the result set is empty**. The iterators inherit that: a 4xx raises instead of making the first page short and ending the walk silently. Note the *other* fail-open path is unchanged and still bites -- GLPI v2 ignores a filter field it does not recognise and answers 200 with the whole unfiltered table, so a filter that returns rows is still not proof it was applied. - `create_ticket` returns the new ticket ID. `update_ticket` and `delete_ticket` return `None`. - **`GetTicket` has no `content` field; it has `content_html` and a `content` property.** `content_html` holds the server's HTML verbatim (and accepts the wire spelling `content` on construction); `content` is a cached property that converts it to Markdown on **first read**, not on validation. Reading `ticket.content` is unchanged and still gives Markdown, so read-side code needs no edit — but `search_tickets` now converts nothing until a body is read, and a body that cannot be converted no longer breaks the rest of the page. Two things do change if you were relying on them: `"content" not in GetTicket.model_fields`, and `GetTicket(...).model_dump()` emits `content_html` with HTML where it used to emit `content` with Markdown (`by_alias=True` gives you the GLPI key). `PostTicket`/`PatchTicket` are untouched: plain `content` field, converted eagerly. +- **`ticket.content` spells literal text so it stays text.** A character is escaped exactly where python-markdown (with `nl2br`, `sane_lists`, `fenced_code`, `tables`) would read it as syntax — a user's `__init__` reads back as `\_\_init\_\_`, `#4521` at a line start as `\#4521`, `<Entrée>` as `<Entrée>` — and nowhere else, so `fichier_de_test_v2.xlsx` or `R&D` come back as typed. Render the Markdown to display it; do not strip the backslashes. One real HTML element in a value you write makes the whole value HTML, so write Markdown or HTML, not both. - HTML too deeply nested to walk is **stripped to text rather than converted** — `markdownify` recurses about twice per nesting level and dies around 494 from a shallow stack. The converter attempts the conversion and answers the `RecursionError` rather than predicting the depth, so the real limit is whatever stack is left at the call site, and the same body can convert from one and degrade from a deeper one. It degrades rather than truncating — every character the normal rendering would have produced still appears — but structure does not survive: link targets, image alt text, code fencing and `<pre>` indentation are gone. Anything else that goes wrong converting a body raises `GlpiContentError`, from the attribute read rather than from `get_ticket`. Treat a fetched ticket as immutable afterwards: the conversion is cached, so assigning to `content_html` (or `model_copy(update={"content_html": ...})`) leaves stale Markdown on `.content` with nothing in `repr`, `==` or `model_dump` to show it. - The GLPI server is the authoritative validator. Extra keys returned by the server flow into `ticket.extra_payload` rather than raising. Caller-provided `extra_payload` keys win on conflicts. - **`status` is writable, on `PatchTicket` only.** `client.update_ticket(tid, PatchTicket(status=GlpiTicketStatus.PENDING))` moves the ticket. This corrects an earlier claim in this skill that the field was read-only: GLPI's own contract publishes `Ticket.status.id` as `readOnly: true`, and that is simply wrong — a live GLPI 11 instance honours the `PATCH`. The same contract omits the `Major` level from `priority`, so treat it as a hint, not an authority. Two things it does *not* cover: `POST` **ignores** `status` (201, then the ticket reads back as `New`), which is why the field is declared on `PatchTicket` and not on `PostTicket`; and `status_id` is silently dropped, so use `status`. Posting a solution still moves the ticket to `SOLVED` on its own and remains the better call when there is a resolution to record — a bare status change leaves no trace of why. diff --git a/skills/glpi-user-location-provisioning/SKILL.md b/skills/glpi-user-location-provisioning/SKILL.md index caab841..d8f0c6a 100644 --- a/skills/glpi-user-location-provisioning/SKILL.md +++ b/skills/glpi-user-location-provisioning/SKILL.md @@ -5,7 +5,7 @@ license: MIT compatibility: "Requires Python 3.10+, glpi-python-client, network access to the GLPI v2 API, and credentials allowed to read or write users, locations, and entities." metadata: package: glpi-python-client - version: "0.5.0" + version: "0.6.0" --- # GLPI User, Location, And Entity Provisioning From 7c5e04ad3feae3fd3f503fa67d5924684a48426c Mon Sep 17 00:00:00 2001 From: baraline <antoine.guillaume45@gmail.com> Date: Thu, 1 Oct 2026 02:00:40 +0200 Subject: [PATCH 4/7] test(content): keep the stack-can-hold case inside CPython 3.10's margin The deepest case was 300 levels. CPython 3.10, which CI runs, spends about three frames per level where 3.12 spends two, so its cliff is about 328 levels from a shallow stack -- measured by easyvista_python_client's port of this converter, 2026-09-30 -- and 300 left 14 levels of margin under pytest. 250 still proves the point, being past the old fixed bound of 200. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> --- glpi_python_client/content/tests/test_conversion.py | 11 ++++++++--- 1 file changed, 8 insertions(+), 3 deletions(-) diff --git a/glpi_python_client/content/tests/test_conversion.py b/glpi_python_client/content/tests/test_conversion.py index 7c19969..ba64dfe 100644 --- a/glpi_python_client/content/tests/test_conversion.py +++ b/glpi_python_client/content/tests/test_conversion.py @@ -374,7 +374,7 @@ def test_an_unterminated_raw_text_element_reads_the_same_on_both_paths() -> None assert "keep" in GlpiContentConverter.from_transport("<p>keep</p><script>SECRET") -@pytest.mark.parametrize("depth", [1, 100, 200, 250, 300]) +@pytest.mark.parametrize("depth", [1, 100, 200, 250]) def test_a_document_the_stack_can_hold_is_converted_in_full(depth: int) -> None: """Everything that fits must convert, and structure has to survive. @@ -382,8 +382,13 @@ def test_a_document_the_stack_can_hold_is_converted_in_full(depth: int) -> None: predicted the depth and degraded past a fixed 200, which flattened every body between 200 and the real cliff of about 494 -- ordinary quoted mail threads among them -- to text, with no error to notice - and no way for a caller to ask for better. The 300 and 400 cases here - are the ones that used to come back as prose. + and no way for a caller to ask for better. The 250 case is one that + used to come back as prose. It is also as deep as this goes, so that + it holds on every interpreter CI runs: CPython 3.10 spends about three + frames per level where 3.12 spends two, its cliff is about 328 levels + from a shallow stack (measured by easyvista_python_client's port, + 2026-09-30), and a 300-level case left 14 levels of margin under + pytest there. """ html = "<div>" * depth + "<strong>offline</strong>" + "</div>" * depth From 917f0303252dfa00d044a43c23d0cfc3695833a1 Mon Sep 17 00:00:00 2001 From: baraline <antoine.guillaume45@gmail.com> Date: Thu, 1 Oct 2026 17:25:25 +0200 Subject: [PATCH 5/7] update parsing package and method --- CHANGELOG.md | 188 +- docs/api_reference.rst | 27 +- docs/user_guide.rst | 49 +- glpi_python_client/content/conversion.py | 2671 +++-------------- glpi_python_client/content/tests/display.py | 220 ++ .../content/tests/test_conversion.py | 1109 +------ .../content/tests/test_literal_text.py | 1548 ---------- .../content/tests/test_round_trip.py | 350 +++ .../models/api_schema/_content.py | 19 +- .../models/api_schema/tests/test_content.py | 17 +- .../models/custom_schema/_ticket_context.py | 18 +- .../tests/test_ticket_context.py | 39 + .../testing/tests/test_content_roundtrip.py | 120 +- pyproject.toml | 11 +- skills/glpi-asset-workflow/SKILL.md | 2 +- skills/glpi-contract-workflow/SKILL.md | 2 +- skills/glpi-ticket-timeline/SKILL.md | 4 +- skills/glpi-ticket-workflow/SKILL.md | 4 +- 18 files changed, 1375 insertions(+), 5023 deletions(-) create mode 100644 glpi_python_client/content/tests/display.py delete mode 100644 glpi_python_client/content/tests/test_literal_text.py create mode 100644 glpi_python_client/content/tests/test_round_trip.py diff --git a/CHANGELOG.md b/CHANGELOG.md index 6837c53..8b7abce 100644 --- a/CHANGELOG.md +++ b/CHANGELOG.md @@ -6,122 +6,90 @@ The format is based on [Keep a Changelog](https://keepachangelog.com/en/1.1.0/). ## 0.6.0 — 2026-10-01 -### Fixed - -- **Literal text came back as Markdown syntax.** `from_transport` handed - `markdownify`'s output on as Markdown, and `markdownify` escapes nothing - it is not asked to, so text a user typed into GLPI changed meaning as - soon as anything rendered it. Measured through the reader into - python-markdown with the four extensions `to_transport` uses: - `\\serveur\compta` lost a backslash, `__init__` and `______` became - emphasis, a `-----------` line under text made the text a heading, - `* point` and `> merci` lines made a list and a quote, `#4521` at the - start of a line made a heading, `[1]: https://...` was consumed as a - reference definition, and a `|` in a table cell dropped the rest of the - row. - - The Markdown now spells literal text so python-markdown reads it as - text. A character is escaped **exactly where python-markdown would - otherwise read it as syntax, and nowhere else**: ordinary prose -- - `fichier_de_test_v2.xlsx`, `C:\Temp\logs`, `R&D`, a `#` or a `-` - mid-sentence, `5 * 3` -- comes back byte for byte as before. The escape - is a backslash wherever python-markdown removes one (`\_`, `\#`, `\\`, - `\[`, `\|`, ...), and a character reference where no backslash works: - `<` for a `<` that would open a tag or an e-mail autolink, `&` - for an `&` that would start a reference, `=` and `~` for a - setext `=` underline or a `~~~` fence. A pasted URL link stays - `<https://...>`. - - The rules replay python-markdown's own: its code-span pattern, its - escape set (asked of the renderer, not copied), its bracket counting, - its emphasis pairing, the line tests for headings, rules, setext - underlines, fences, lists, quotes, table separators and reference - definitions, and its e-mail autolink, which it matches only after code - spans, links and images are set aside. They are held to two properties, - checked with an HTML parser over the measured cases, over realistic - bodies and over a seeded fuzzer: `to_transport(from_transport(html))` - **displays what `html` displays**, and the Markdown is **a fixed - point**, reading back as itself. 30,000 fuzzed bodies were measured - clean; 400 run in the suite. - -- **Nested lists flattened on the first write.** `markdownify` indents a - nested item by its bullet's width, two or three spaces, where - python-markdown nests at four, so a nested list lost a level per cycle - and a nested ordered list restarted its count. Continuation lines are - indented by four spaces, and numbering, `start` included, survives. - Around that, the blocks inside a list item now stay in it: text or a - quote after a nested list starts a block of its own instead of running - into the list's last item, a quote in an item gets the blank line - python-markdown needs, and a quote opening an item keeps its later - lines. An item python-markdown writes as loose is written loose, so the - second read equals the first; a list whose first line carries three - bullets (`- - - a`) is spaced out, since python-markdown cannot nest - anything under such a line tightly. - -- **`<script>`, `<style>` and `<title>` bodies leaked into the text.** - `strip=["script", "style"]` skips an element's own converter, which is - the one that drops its body, so Outlook's CSS arrived as prose. They are - dropped now, on the converting path, as a browser drops them. The - degraded path still keeps them, which its promise -- never less than - the converting path -- allows. - -- **Bodies marked up only with obsolete elements were read as text.** - `<font>`, `<center>`, `<big>`, `<tt>`, `<nobr>` and the rest of the HTML - standard's obsolete list now make a body HTML; `<center>` and `<dir>` - are blocks. - -- **`<s>`, `<del>` and `<strike>` became a literal `~~`.** python-markdown - has no strikethrough, so the markers showed. The words are kept and the - line is dropped -- recorded as a loss below. - -- `<pre>` inside a list item or a quote is an indented code block, since a - fence opens only at the start of a line and was read there as text; a - top-level fence is made longer than any fence line the code holds. An - image stays an image in a heading or a cell. A newline in the HTML - source is a space, as a browser shows it, rather than a line break - from `nl2br`. Whitespace around a `<br>` is dropped, so one body reads - one way. - -- The degraded rendering of a body too deep to convert is spelled through - the same rules, so it is literal-safe too. - -### Changed +### Changed (breaking) -- **Mixed Markdown and HTML in one value is read as HTML.** `from_transport` - is also the validator on write models, and a value carrying a single real - HTML element takes the HTML path throughout: `**bold** <b>x</b>` keeps - its asterisks as text now, where it used to render them. Write Markdown - or HTML, not both. +- **Content conversion is rebuilt on three libraries: markdownify, + mdformat and cmark-gfm.** `from_transport` reads GLPI's HTML with + `markdownify`, and `mdformat` re-renders that Markdown from its syntax + tree, so it keeps only the escapes CommonMark needs. `to_transport` + renders through `cmark-gfm`, the GitHub reference implementation, in + place of python-markdown. A thin layer of glue sits on top. Measured on + 346 real bodies sampled from a GLPI 11 instance: + - 322 display the same after a round trip, against 299 before; + - 344 read back as the same Markdown, against 240. + + Of 205 realistic caller-written Markdown documents, all 205 survive + Markdown → HTML → Markdown with the same display. +- **Markdown is rendered as CommonMark with GFM tables.** A newline is a + line break, as `nl2br` made it before. Raw HTML passes through. + Differences you may see in Markdown you write: + - a list or a table written straight after a line now starts a list or a + table, where python-markdown wanted a blank line first; + - lists nested by two or three spaces nest; + - `1)` starts a numbered list; + - `#Important`, with no space, is text rather than a heading; + - `*a **b** c*` keeps its bold; + - a backslash ending a line is a line break, so write `C:\Temp\` at the + end of a line as `` `C:\Temp\` ``; + - `<word>` is read as an HTML tag, so put a placeholder such as `<login>` + in backticks. +- **`.content` is spelled as canonical CommonMark.** A line break reads + back as `\` and a newline, a nested list is indented by its bullet's + width, and a table comes back unpadded. Text that would otherwise read as + syntax is escaped: `__init__` reads `\_\_init\_\_`, and a `* point` line + reads `\* point`. Stored digests of `.content` change once. +- **A plain-text body is literal text on the read path.** A value with no + HTML element used to come back verbatim and be rendered as Markdown. It is + now read as GLPI displays it, its lines as lines. + `GlpiContentConverter.from_transport` takes `plain_text_is_markdown`. + `True` is what the write models' validator passes: caller-authored + Markdown passes verbatim unless it starts with an HTML tag, so Markdown + carrying an inline `<br>` or `<kbd>` stays Markdown. +- **Dependencies.** + - Added: `cmarkgfm>=2025.10` (compiled wheels for CPython 3.10–3.14 on + Linux, macOS and Windows), `mdformat>=0.7.22,<0.8`, + `mdformat-tables>=1.0` and `markdown-it-py>=3.0`. + - Dropped: `markdown`. python-markdown 3.11 had broken the previous + reader. + - Raised: `beautifulsoup4>=4.15`, which fixed the parser defect that + dropped the text after a `<br />` in a body that also held a bare + `<br>`. That removes the workaround. -- **`markdownify>=1.2`** (was `>=0.13`). The converter is a - `MarkdownConverter` subclass, and it relies on the 1.x converter hooks -- - `escape(text, parent_tags)`, `convert_*(el, text, parent_tags)`, - `convert__document_` -- which 0.13 does not have. 1.2.2 and 1.2.3 are - the versions measured; 1.0 and 1.1 were not, so they are not claimed. +### Fixed -- Markdown read from GLPI changes for any body whose text the old reader - let through as syntax, which is the point, and for nested lists. A - caller who stored the old Markdown or a digest of it sees a one-time - difference on the next read. +- **Literal text came back as Markdown syntax.** These now read back as the + text a user typed: + - `\serveur\compta`, which had lost a backslash; + - `__init__` and `______`, which had become bold; + - a `-----` line under text, which had made a heading; + - `* point` and `> merci` lines, which had become a list and a quote; + - `[1]: https://...`, which had been consumed as a reference definition; + - a `|` in a table cell, which had dropped the rest of the row. +- **A table nested in a table cell lost all its text.** That is the usual + layout of an e-mail signature. The inner table is now written as its + cells' text, its line breaks kept as `<br>`. +- **Nested lists flattened on the first write.** Nested items, the text + after a nested list, and numbering, `start` included, survive. +- `<script>`, `<style>` and `<title>` bodies no longer leak into the text. +- A blank line inside a paragraph (`<br><br>`, or Outlook's + `<br> <br>`) is kept. +- A label in bold right before a figure, as in `<b>Total:</b>12`, keeps its + bold. +- A code fence keeps its language, and a `|` inside code in a table cell no + longer splits the row. +- A long body converts in linear time, and a body nested deeper than the + stack is read as its text instead of raising. ### Known limitations -Structure Markdown cannot carry, each asserted to still be a loss and to -stabilise after one cycle (`test_what_markdown_cannot_carry`): -strikethrough and underline (words kept, line lost); two adjacent lists, -or two adjacent quotes, read back as one; two `<br>` in a row are a -paragraph break; adjacent code spans merge; strong inside emphasis loses -its bold; a table without a header row gains an empty one, and one -without a body gains an empty row; a table cell or a heading holds one -line, so a list, a code block or a line break in one is flattened; a -`<pre>` that opens a list item, or directly follows a list inside the same -item or quote, is kept as text, because python-markdown reads no code -block there. Inside a list item or a quote, a numeric reference with no -semicolon in code gains one (`test_a_numeric_reference_in_nested_code_gains_a_semicolon`). - -Reading costs more: about 1.4x the old reader on realistic bodies and up -to 4x on bodies dense with syntax characters, every pass linear in the -body's size. +- A Markdown table needs a header row and the same number of cells in every + row. A header-less HTML table gains an empty header row, and a row that + spans the table gains empty cells. +- Struck-through text is kept as raw `<s>`, since CommonMark has no + strikethrough. +- Literal text that looks like syntax is sometimes escaped where + CommonMark would not need it, for example `5\*3` or `x \<= y`. It + displays as typed. ## 0.5.0 — 2026-09-08 diff --git a/docs/api_reference.rst b/docs/api_reference.rst index 4daabc8..0d259cd 100644 --- a/docs/api_reference.rst +++ b/docs/api_reference.rst @@ -113,26 +113,19 @@ conversion happens, which buys two things: listing records costs nothing per body, and a body that cannot be converted no longer stops the rest of its page being read. -The Markdown spells a body's text as literal text: a character is escaped -exactly where python-markdown, with the package's four extensions, would -read it as syntax -- ``\_\_init\_\_``, ``\\\serveur``, ``\#4521`` at the -start of a line, ``<Entrée>`` -- and nowhere else, so rendering it -displays what GLPI displayed and reading that back gives the same Markdown. -A value holding one real HTML element is read as HTML throughout, write -models included. See :ref:`content-conversion`. +The Markdown is CommonMark with GFM tables, rendered by cmark-gfm. It +spells a body's text as literal text -- ``\_\_init\_\_``, ``\\\serveur``, +``\# pas un titre`` -- so rendering it displays what GLPI displayed and +reading that back gives the same Markdown. A write model keeps the caller's +Markdown verbatim unless it starts with an HTML tag. See +:ref:`content-conversion`. Very deeply nested HTML is the case worth knowing about. ``markdownify`` walks the document recursively and runs out of stack at -around 494 levels of nesting. The converter does not try to predict -that: it attempts the conversion and, if the walk does not fit, strips -tags instead. It degrades, it never truncates, and it does not raise: -every character the normal rendering would have produced still appears. -What is lost is structure rather than words — link targets and image alt -text, code fencing and ``<pre>`` indentation, `` `` alignment. -Because the budget is the stack left when the conversion starts, the -same body can convert from one call site and degrade from a deeper one. -Anything else that goes wrong in either direction raises -:class:`GlpiContentError`. +a few hundred levels of nesting. The converter attempts the conversion +and, if the walk does not fit, reads the body's text instead, a line per +block: the words survive, the structure does not. Anything else that goes +wrong in either direction raises :class:`GlpiContentError`. Because the conversion is cached on first read, a read model should be treated as immutable afterwards: assigning to ``content_html``, or diff --git a/docs/user_guide.rst b/docs/user_guide.rst index e2cb70d..585df9c 100644 --- a/docs/user_guide.rst +++ b/docs/user_guide.rst @@ -1758,40 +1758,35 @@ The whole page is built in one pass, so a single unconvertible record used to make its page-mates unreadable too. The failure is now scoped to the record whose body you actually read. -Text in a body is literal, and ``.content`` spells it so: rendering the -Markdown -- as the package does on the way back to GLPI, or with any -python-markdown using the same four extensions (``nl2br``, ``sane_lists``, -``fenced_code``, ``tables``) -- displays what GLPI displayed. Markdown has -one spelling for ``__init__`` typed by a user and for bold ``init``, so a -character is escaped exactly where python-markdown would otherwise read it -as syntax: +The Markdown is CommonMark with GFM tables. Rendering it -- as the package +does on the way back to GLPI, with cmark-gfm -- displays what GLPI +displayed, and reading that rendering back gives the same Markdown. Text in +a body is literal: Markdown has one spelling for ``__init__`` typed by a +user and for bold ``init``, so text that would read as syntax is escaped: .. code-block:: python from glpi_python_client.content import GlpiContentConverter GlpiContentConverter.from_transport( - "<p>Voir __init__ et \\\\serveur\\partage</p><p># pas un titre</p>" + r"<p>Voir __init__ et \\serveur\compta</p><p># pas un titre</p>" ) - # 'Voir \\_\\_init\\_\\_ et \\\\\\serveur\\partage\n\n\\# pas un titre' - -Ordinary prose carries no escape at all -- ``fichier_de_test_v2.xlsx``, -``C:\Temp``, ``R&D``, a ``#`` or a ``-`` mid-sentence come back exactly as -typed -- and the Markdown is a fixed point: rendering it and reading it -back gives the same Markdown again. A backslash is used wherever -python-markdown removes one; ``<``, ``&``, ``=`` and ``~`` are spelled as -character references (``<``) where they would be read, since no -backslash escapes them. A pasted URL link stays ``<https://...>``. - -Two things change for a caller. A value holding a single real HTML element -is read as HTML throughout, on write models too, so ``**bold** <b>x</b>`` -keeps its asterisks as text: write either Markdown or HTML, not both. And -a few structures have no Markdown spelling, which the round-trip inventory -in ``test_literal_text.py`` records: struck-through and underlined text -keep their words but lose the line, adjacent lists or quotes merge, a line -break in a table cell or a heading becomes a space, and a ``<pre>`` that -opens a list item, or that directly follows a list inside the same item, -keeps its lines as text rather than as code. + # Voir \_\_init\_\_ et \\\serveur\compta + # + # \# pas un titre + +Ordinary prose stays as typed: ``fichier_de_test_v2.xlsx``, ``C:\Temp``, +``R&D``, a ``#`` mid-sentence. The spelling is canonical: a line break reads +back as ``\`` and a newline, a nested list is indented by its bullet's width, +and a table comes back unpadded. A body with no HTML element is plain text +and is read as GLPI shows it, its lines as lines. + +Writing, your Markdown is rendered by cmark-gfm. A newline is a line break, +GFM tables work, and raw HTML passes through, so put a placeholder such as +``<login>`` in backticks. A write model keeps your Markdown verbatim unless +it starts with an HTML tag, in which case it is read as HTML. A Markdown +table needs a header row, so a header-less HTML table reads back with an +empty one, and struck-through text stays as raw ``<s>``. .. note:: diff --git a/glpi_python_client/content/conversion.py b/glpi_python_client/content/conversion.py index 8542881..0bc1ff0 100644 --- a/glpi_python_client/content/conversion.py +++ b/glpi_python_client/content/conversion.py @@ -1,108 +1,54 @@ -"""Content conversion helpers for GLPI payloads. - -This module translates between GLPI's HTML transport format and the -package's canonical Markdown representation used by the rich content -models. - -Literal text ------------- - -**Text in the HTML is literal, and the Markdown spells it so.** A user who -types ``__init__`` into GLPI means eight characters; Markdown reads the -same eight as bold ``init``. So :meth:`GlpiContentConverter.from_transport` -escapes a character exactly where python-markdown -- with this module's -extensions, which is what :meth:`GlpiContentConverter.to_transport` and any -peer rendering the Markdown uses -- would otherwise read it as syntax, and -nowhere else. It has to happen here: once the Markdown is written, nothing -can tell a literal ``__`` from a bold one any more. The rule set is -documented on :class:`_LiteralSafeConverter`; the property it is held to is -that ``to_transport(from_transport(html))`` displays what ``html`` -displays, and that the Markdown is a fixed point of the round trip. - -Ordinary prose carries no escape at all -- ``fichier_de_test_v2.xlsx``, -``C:\\Temp\\logs``, a ``#`` or a ``-`` mid-sentence, ``R&D`` -- because -python-markdown already renders those literally. What is escaped is what it -would not: ``\\\\serveur`` (it would lose a backslash), ``__init__``, a -``#4521`` or a ``> merci`` at the start of a line, a ``|`` in a table cell, -``<Enter>``, ``&``. None of this is visible to anyone reading the ticket -in either ITSM: the escapes exist only in the Markdown between them. - -Neither *parse* recurses -- ``html.parser`` is an iterative scanner -- -but ``markdownify`` walks the finished tree recursively, at about two -CPython frames per nesting level, so inbound conversion has a nesting -ceiling. Measured from a shallow stack against the default 1000-frame -limit, the deepest document that converts is 494 levels: the same 494 -for ``<div>``, ``<p>``, ``<blockquote>`` and ``<table><tr><td>``, and -495 for ``<ul><li>``, which is what identifies the cost as per-level. - -**The ceiling is discovered rather than predicted.** Inbound conversion -is attempted, and a ``RecursionError`` is caught and answered by -stripping the document to its text instead -- see -:meth:`GlpiContentConverter.from_transport`. Estimating the depth up -front and degrading past a fixed bound was the previous design, and it -was wrong in both directions: it degraded bodies that would have -converted, because the bound had to assume the worst about the caller's -remaining stack, and three rounds of review found seven ways for the -estimate to come in *under* the real tree, each of which put a document -through ``markdownify`` and into the ``RecursionError`` the bound -existed to prevent. Trying the conversion cannot be wrong about whether -the conversion fits. - -Anything else that goes wrong in either direction surfaces as -:class:`~glpi_python_client.GlpiContentError`, so no parser fault -escapes the package's exception taxonomy. - -**``sys.setrecursionlimit`` is deliberately not called, here or anywhere -in the package.** It is process-global state that belongs to the -application, not to a library an application imported; and raising the -limit past what the C stack can hold turns a catchable -``RecursionError`` into a hard interpreter crash -- on Windows, an -access violation with no traceback. It moves the cliff and makes falling -off it worse. Degrading the one body that does not fit is the answer -that does not. Running the walk in a thread with a larger stack was -considered and rejected for the same reason: the recursion limit is a -counter rather than a measurement of the stack, so a deeper thread still -needs the global limit raised to use it. +"""Content conversion between GLPI's HTML and the package's Markdown. + +GLPI stores rich text as HTML; the package's surface is Markdown. + +* **Read** -- :meth:`GlpiContentConverter.from_transport`. ``markdownify`` + writes the HTML as Markdown with every character that could be syntax + escaped, then ``mdformat`` re-renders that Markdown from its syntax tree, + which keeps only the escapes CommonMark needs. Text is literal: + ``__init__`` typed into GLPI reads back as ``\\_\\_init\\_\\_``. +* **Write** -- :meth:`GlpiContentConverter.to_transport`. ``cmark-gfm`` + renders CommonMark with GFM tables; a newline is a line break and raw + HTML passes through. + +Rendering the Markdown a read gives displays what GLPI displayed, and +reading that back gives the same Markdown. The glue below covers what the +three libraries leave out: plain text, line breaks a browser does not +show, bold and italic CommonMark would not close, link targets, and an +mdformat set up without its nesting cap and with its quadratic lookups made +linear. A body nested too deeply for the stack is read as its text +(:func:`_text_of`); anything else that fails raises +:class:`~glpi_python_client.GlpiContentError`. """ from __future__ import annotations import re -from collections.abc import Callable, Iterable -from html import unescape -from html.entities import html5 as _HTML5_REFERENCES -from html.parser import HTMLParser -from itertools import chain -from typing import Any, NamedTuple - -from bs4 import Comment, Doctype, ParserRejectedMarkup, Tag -from bs4.element import PageElement -from markdown import Markdown -from markdown import markdown as markdown_to_html -from markdown.blockprocessors import HRProcessor, ReferenceProcessor -from markdown.inlinepatterns import ( - AUTOLINK_RE, - AUTOMAIL_RE, - BACKTICK_RE, - NOT_STRONG_RE, -) +import string +import unicodedata +from collections.abc import Callable, Iterator, Mapping, MutableMapping, Sequence +from functools import cached_property +from html import escape, unescape +from typing import Any, cast + +import cmarkgfm +import mdformat_tables +from bs4 import BeautifulSoup, ParserRejectedMarkup, Tag +from bs4.element import PageElement, PreformattedString +from markdown_it import MarkdownIt +from markdown_it.token import Token from markdownify import MarkdownConverter +from mdformat.renderer import ( + DEFAULT_RENDERERS, + MDRenderer, + RenderContext, + RenderTreeNode, +) from glpi_python_client._errors import GlpiContentError -#: Element names that make a ``<...>`` sequence markup rather than text. -#: -#: The HTML5 element set, which is what the parser behind ``markdownify`` -#: will actually recognise. Anything outside it -- ``<Enter>``, ``<T>``, -#: ``</dev/null>`` -- parses as an *unknown* tag, whose markup is dropped -#: while its (usually empty) body is kept, so the token silently vanishes -#: from the middle of a sentence. -#: -#: The second group is the HTML standard's *obsolete* elements -- its list -#: of features that "must not be used by authors", which old editors and -#: e-mail clients write all the same. Without them a body marked up only with -#: ``<font color="red">URGENT</font>`` or ``<center>`` failed the probe, took -#: the plain-text path and kept its tags as text. +#: Element names that make a ``<...>`` sequence markup rather than text: the +#: HTML5 elements, then the obsolete ones old editors and mail clients write. _HTML_ELEMENTS = frozenset( """ a abbr address area article aside audio b base bdi bdo blockquote body br @@ -121,127 +67,76 @@ """.split() ) -#: python-markdown extensions applied when rendering outbound content. -#: -#: ``fenced_code`` and ``tables`` are here because without them the two -#: constructs do not survive at all. A fence rendered without -#: ``fenced_code`` becomes inline ``<code>``, which the GLPI web UI shows -#: as one run-on line and which a later read writes back as inline code -- -#: so a pasted log degrades a little more on every edit. A table without -#: ``tables`` renders as literal pipe characters. -#: -#: A language tag is still lost: ``markdownify`` drops the -#: ``class="language-python"`` that ``fenced_code`` emits, so ```` ```python ```` -#: comes back as a bare fence. That is a limitation of the pair of -#: libraries, not something an extension list can fix. -_MARKDOWN_EXTENSIONS = ["nl2br", "sane_lists", "fenced_code", "tables"] - -#: One candidate tag: ``<`` or ``</`` immediately followed by a name. -#: -#: The ``<`` must abut the name, matching what an HTML parser accepts. That -#: is what keeps ``2 < 3 > 1`` and ``x <= y`` text: a space after ``<`` -#: means no tag, so arithmetic never reaches the HTML path in the first -#: place. +#: ``<`` or ``</`` right before a name: one candidate tag. _CANDIDATE_TAG = re.compile(r"</?([a-zA-Z][a-zA-Z0-9]*)\b[^<>]*>") -#: Elements that cannot contain anything, so nothing nests below them. -#: -#: ``html.parser`` -- the parser ``markdownify`` builds its tree with -- -#: closes these itself, so ``<br>`` a thousand times over is a thousand -#: siblings, not a thousand levels. Measured: ``"<br>" * 5000`` parses one -#: level deep and converts fine, while ``"<div>" * 5000`` parses 5000 deep -#: and raises. -#: -#: The HTML5 void set, plus the ten legacy names ``bs4``'s HTML-parser tree -#: builder also treats as empty. **The invariant is that this stays a -#: subset of what the parser treats as empty**, and the direction matters: -#: a name missing from here is counted as nesting when it does not, which -#: costs an unnecessary degradation, while a name wrongly *in* here hides -#: real nesting, which is a ``RecursionError``. Listing the legacy names -#: only makes the count exact on old markup. Copied rather than imported -- -#: it lives in ``bs4.builder`` as -#: ``HTMLTreeBuilder.DEFAULT_EMPTY_ELEMENT_TAGS``, which ``bs4``'s own -#: documentation marks ``:meta private:`` -- and the subset invariant is -#: asserted against a real parse in the unit tests, so a future ``bs4`` -#: cannot quietly break it. The two sets currently hold the same 24 names. -_VOID_ELEMENTS = frozenset( +#: Elements a browser lays out as blocks: a line ends at their edges. +_BLOCKS = frozenset( """ - area base basefont bgsound br col command embed frame hr image img - input isindex keygen link menuitem meta nextid param source spacer - track wbr + address article aside blockquote body caption center dd details dir div + dl dt fieldset figcaption figure footer form h1 h2 h3 h4 h5 h6 header hr + html li main menu nav ol p pre section summary table tbody td tfoot th + thead tr ul """.split() ) -#: Elements whose boundary becomes a line break when tags are stripped. -#: -#: Used only by :func:`_strip_tags`. Removing a block element outright runs -#: its neighbours together -- ``<p>a</p><p>b</p>`` becomes ``ab`` -- while -#: putting a separator at *every* tag breaks words apart, turning -#: ``<b>off</b>line`` into ``off line``. Splitting on the block/inline line -#: keeps both readable. An unrecognised name counts as inline, matching how -#: the normal path treats it: markup dropped, body kept in place. -_BLOCK_ELEMENTS = frozenset( - """ - address article aside blockquote br center col dd details dialog dir div - dl dt fieldset figcaption figure footer form h1 h2 h3 h4 h5 h6 header - hgroup hr li main menu nav ol p pre search section summary table tbody td - tfoot th thead tr ul - """.split() -) +#: Anything between angle brackets, for the parser-rejected fallback. +_ANY_TAG = re.compile(r"<[^<>]*>?") +#: Elements whose content a browser does not display. +_HIDDEN = ["head", "script", "style", "template", "title"] -#: One character reference, with or without its terminating semicolon. -#: -#: Resolved by :func:`_resolve_references` rather than by ``html.unescape`` -#: over the whole string. ``unescape`` implements the HTML5 rule of -#: consuming the longest *known* name it can find, so a semicolon-less -#: reference that is a prefix of a longer unknown word gets split -- -#: measured, ``http://x/?a=1©right=2`` becomes -#: ``http://x/?a=1©right=2``, and a URL pasted into a ticket is exactly -#: where ``©right=`` and ``¬anentity=`` occur. The parser behind the -#: converting path leaves those whole, so this does too. -_CHARACTER_REFERENCE = re.compile( - r"&(?:\#[0-9]+;?|\#[xX][0-9a-fA-F]+;?|[A-Za-z][A-Za-z0-9]*;?)" -) +#: The elements whose ``*`` markers are checked, and the attributes that +#: carry the characters displayed on either side of one (:func:`_note_sides`). +_EMPHASIS = frozenset({"b", "strong", "em", "i"}) +_BEFORE, _AFTER = "data-glpi-before", "data-glpi-after" +#: Text split into leading line breaks and spaces, content, trailing ones. +_EDGES = re.compile(r"((?:\\\n|\s)*)(.*?)((?:\\\n|\s)*)", re.DOTALL) -#: The ``/`` of a self-closing tag, with any space around it. -_VOID_SELF_CLOSE = re.compile(r"\s*/\s*>$") +#: A URL that is its own CommonMark autolink and that markdown-it leaves as it +#: is: printable ASCII it does not percent-encode. +_AUTOLINK = re.compile( + r"[A-Za-z][A-Za-z0-9+.-]{1,31}:[A-Za-z0-9;/?:@&=+$,\-_.!~*'()#%]*" +) +_BACKTICK_RUN = re.compile("`{3,}") -def _resolve_references(text: str) -> str: - """Resolve character references the way the real parser would. - A numeric reference always resolves. A named one resolves only when - the whole name is known, which is the difference from - ``html.unescape``: that consumes the longest known *prefix*, so - ``©right=2`` loses its ``©`` and leaves ``right=2`` behind. - See :data:`_CHARACTER_REFERENCE`. +def _language(pre: Tag) -> str | None: + """Return the language cmark-gfm wrote on a fence as ``class="language-x"``.""" + + for tag in (pre.find("code"), pre): + if isinstance(tag, Tag): + for name in tag.get_attribute_list("class"): + if name and name.startswith("language-"): + return str(name[len("language-") :]) + return None - Where the two rules differ this one keeps the reference literal, which - is the safe direction for :func:`_strip_tags`, whose promise is that no - text goes missing. - """ - def resolve(match: re.Match[str]) -> str: - token = match.group(0) - if token[1] == "#": - return unescape(token) - return unescape(token) if token[1:] in _HTML5_REFERENCES else token +#: ``markdownify`` options; the converter's overrides decide the rest. +_MARKDOWNIFY_OPTIONS: dict[str, Any] = { + "code_language_callback": _language, + "bullets": "-", + "escape_misc": True, + "heading_style": "atx", + "newline_style": "backslash", + "wrap": True, # a newline in HTML text is a space... + "wrap_width": None, # ...and no line is wrapped +} - return _CHARACTER_REFERENCE.sub(resolve, text) +#: cmark-gfm's options: a newline is a line break, raw HTML passes through. +_RENDER_OPTIONS = ( + cmarkgfm.cmark.Options.CMARK_OPT_HARDBREAKS + | cmarkgfm.cmark.Options.CMARK_OPT_UNSAFE +) def _looks_like_html(content: str) -> bool: - """Return whether ``content`` carries at least one real HTML element. - - Deciding on the element *name* rather than on the presence of angle - brackets is what separates markup from prose that merely contains - ``<`` and ``>``. It cannot separate them perfectly: ``a<b>c`` is - genuinely ambiguous, because ``b`` is both a real element and a - plausible variable, and no probe reading the text alone can resolve - that. It resolves every case where the name is not an element at all, - which is where the silent deletions came from. + """Return whether ``content`` holds at least one real HTML element. + + The element *name* decides, so ``use the <Enter> key`` and + ``if x<y then z>0`` are text. ``a<b>c`` is markup: ``b`` is an element. """ return any( @@ -250,2129 +145,531 @@ def _looks_like_html(content: str) -> bool: ) -class _ParserScan(HTMLParser): - """Walk a document with the parser that will convert it, not one like it. - - Both remaining questions this module asks about raw HTML -- which - self-closing void tags to rewrite, and what the text is when the tree - will not fit the stack -- are questions about ``html.parser``'s - dispatch. This subclass asks ``html.parser`` instead of describing - it. - - It replaces a regular expression that reproduced that dispatch by - imitation, and the imitation kept being wrong in ways that showed up - only after they shipped. A comment closes on ``--\\s*>`` - and not only on ``-->``, ``</ script>`` ends raw text, ``<![IGNORE[`` - opens a marked section, ``</ div foo>`` is a bogus comment rather than - an end tag, and ``<a href=/>`` leaves an element *open* because the - unquoted value swallows the ``/``. Two more were cost rather than - correctness: a run of whitespace inside a failing tag made the - attribute pattern backtrack as ``(a+)*``, and a 39-byte body took - 20.8 seconds. - - Those seven were found as depth under-counts, back when this module - predicted the nesting depth instead of attempting the conversion. - Predicting it is gone, but the pattern's other two readers were wrong - the same way and are still here: a derailed scan meant - :func:`_canonicalise_void_elements` never saw the ``<br />`` it - exists to rewrite, so - ``"<p>one<br>two</p><script>x</ script><p>three<br />TAIL</p>"`` lost - ``TAIL`` outright, and :func:`_strip_tags` inherited both the - misreadings and the backtracking. - - Reading the parser's own event stream cannot be wrong about the - parser, so none of those remain judgement calls. It is also not a new - dependency nor a new risk: ``markdownify`` builds its tree with - ``bs4``, and ``bs4`` builds it with this same ``html.parser``, so - every pathology the parser has was already in the pipeline. Measured - on the shapes that made the pattern backtrack, the ``markdownify`` - call costs what this scan costs, to within a few per cent. - - ``convert_charrefs`` is ``False`` because that is what ``bs4`` passes - (in ``bs4.builder._htmlparser``), and the difference shows: with it - on, ``html.unescape`` consumes the longest *known* name, so - ``©right=2`` in a pasted URL loses its ``©``. Off, each - reference arrives as its own event and :func:`_resolve_references` - applies the whole-name rule the converting path applies. - - ``collect_text`` separates the two callers: the void rewrite needs - only the spans, and accumulating the pieces of a 150 KB body for it - would be waste. - - Parameters - ---------- - content : str - The document to walk. Kept so spans can be sliced back out of it: - the parser reports what it found and where, and the source is the - only place the exact original spelling still exists. - collect_text : bool, optional - Whether to accumulate :attr:`pieces` for :func:`_strip_tags`. - """ - - def __init__(self, content: str, *, collect_text: bool = False) -> None: - super().__init__(convert_charrefs=False) - self._content = content - self._collect_text = collect_text - offsets = [0] - for line in content.splitlines(keepends=True): - offsets.append(offsets[-1] + len(line)) - self._line_offsets = offsets - #: Source spans of ``<void ... />`` tags, for - #: :func:`_canonicalise_void_elements`. - self.void_spans: list[tuple[int, int]] = [] - #: The document's text, in order, when ``collect_text`` is set. - self.pieces: list[str] = [] - #: Set when ``html.parser`` gave up on the document. - self.rejected = False - - def _at(self) -> int: - """Return the absolute offset of the construct being handled. - - ``goahead`` calls ``updatepos`` up to the start of each construct - before dispatching it, so ``getpos`` addresses the construct - itself. It reports a line and a column, and every span sliced - here needs an index, which is what the line table built in - ``__init__`` converts between. - """ - - lineno, offset = self.getpos() - return self._line_offsets[lineno - 1] + offset - - def _text(self, piece: str) -> None: - if self._collect_text: - self.pieces.append(piece) - - def _boundary(self, tag: str) -> None: - """Record a block element's edge as a line break. - - See :data:`_BLOCK_ELEMENTS` for why the block/inline line is the - one that matters here. - """ - - if self._collect_text and tag in _BLOCK_ELEMENTS: - self.pieces.append("\n") - - def handle_starttag(self, tag: str, attrs: object) -> None: - self._boundary(tag) - - def handle_startendtag(self, tag: str, attrs: object) -> None: - """Count a ``<foo/>`` as a leaf, and note a void one to rewrite. - - This is the event :func:`_canonicalise_void_elements` needs, and - the one a pattern cannot identify reliably. ``html.parser`` - reaches it only when the stripped remainder of the tag is exactly - ``/>``, so ``<br />`` arrives here while ``<br / >`` is an - ordinary start tag. Deciding it on the event means the workaround - fires on exactly the tags that trigger the ``bs4`` defect, and on - no others. - """ - - if tag in _VOID_ELEMENTS: - token = self.get_starttag_text() - start = self._at() - if token is not None and self._content.startswith(token, start): - self.void_spans.append((start, start + len(token))) - self._boundary(tag) - - def handle_endtag(self, tag: str) -> None: - self._boundary(tag) - - def handle_data(self, data: str) -> None: - """Keep character data, a raw-text element's body included. - - A ``<script>`` or ``<style>`` body arrives here because the parser - is in CDATA mode. The converting path drops such a body -- a - browser displays none of it -- and this keeps it anyway, which is - the safe direction for a fallback whose promise is that it says no - less than the conversion: see :func:`_strip_tags`. - """ - - self._text(data) - - def handle_entityref(self, name: str) -> None: - self._text(self._reference_at()) - - def handle_charref(self, name: str) -> None: - self._text(self._reference_at()) - - def _reference_at(self) -> str: - """Return a character reference exactly as it was written. - - The event carries the name but not whether a semicolon closed it, - and that is what decides whether the reference resolves -- so the - source is re-read rather than the token rebuilt from the name. - See :data:`_CHARACTER_REFERENCE`. - """ - - match = _CHARACTER_REFERENCE.match(self._content, self._at()) - return match.group(0) if match is not None else "" - - def handle_pi(self, data: str) -> None: - """Keep a processing instruction's body, which the converter prints. - - Measured, not assumed, and the delimiters are the detail that - matters: ``bs4`` files the instruction as a string node, so - ``<p>a</p><?php SECRET ?><p>b</p>`` converts to - ``"a\\n\\nphp SECRET ?\\n\\nb"`` -- the body, without its ``<?`` - and ``>``. Keeping the body alone therefore matches the - converting path exactly, where keeping the whole construct used - to leave punctuation in an output that promises none. - """ - - self._text(data) - - def keep_remainder(self) -> None: - """Hand back the text of the region the parser stopped on. - - Called only when it raised, so that the degraded path still - carries every word after the construct it could not read. - """ - - self._text(self._content[self._at() :]) - - def unknown_decl(self, data: str) -> None: - """Keep the body of a ``CDATA`` section or a marked section. - - Counter-intuitive, and measured rather than assumed: ``bs4`` files - both as a string node, and ``markdownify`` prints a string node - that is neither a comment nor a doctype. So ``<![IGNORE[x]]>`` - contributes ``IGNORE[x`` to the converted output, and dropping it - here would make the same document say less on the degraded path - than on the converting one. - """ - - self._text(data[6:] if data.startswith("CDATA[") else data) - - -def _scan(content: str, *, collect_text: bool = False) -> _ParserScan: - """Run one :class:`_ParserScan` over ``content`` and hand it back. - - The parser gives up on two constructs -- an unknown marked-section - keyword such as ``<![FOO[``, and a ``[`` where a declaration cannot - hold one -- by raising ``AssertionError`` from ``_markupbase``. That - is not a case to guess around, because ``bs4`` catches the same - ``AssertionError`` and re-raises it as ``ParserRejectedMarkup``: a - document that stops this scan is a document ``markdownify`` cannot - convert either. So the partial scan is kept and the remainder of the - source is handed to :attr:`_ParserScan.pieces` as text, which is what - lets :meth:`GlpiContentConverter.from_transport` answer that - rejection with the body's words instead of an exception. - - Parameters - ---------- - content : str - The document to walk. - collect_text : bool, optional - Whether the scan should accumulate the document's text. - - Returns - ------- - _ParserScan - The finished scan, whether or not the parser ran out of document. - """ - - scan = _ParserScan(content, collect_text=collect_text) - try: - scan.feed(content) - scan.close() - except AssertionError: - scan.rejected = True - scan.keep_remainder() - return scan - - -def _strip_tags(content: str) -> str: - """Reduce HTML to its text without building a tree, keeping every word. - - The fallback for a document ``markdownify`` could not walk -- deeper - than the caller's remaining stack, or refused by the parser outright. - One pass of :class:`_ParserScan`, then whitespace tidying. No tree, - no recursion and no ceiling of its own, which is what qualifies it as - the fallback: it answers for input of any shape and any depth, so - there is always something to give the caller. - - It **degrades and never truncates.** The property, stated as - something checkable: after collapsing whitespace, every character the - converting path would have produced also appears here, in order. A - superset, not an equality -- so no body says less because of the path - it took, which is the only guarantee worth making about a fallback. - - Establishing that meant measuring what the converting path really - keeps, construct by construct, rather than assuming. Two answers were - counter-intuitive and each was a silent deletion here before it was - checked: a ``CDATA`` body is kept, and so is the inside of any - ``<!``/``<?`` construct the parser could not resolve, which it hands - back as character data. A ``<script>``/``<style>``/``<title>`` body - goes the other way: the converting path drops it, as a browser does, - and this keeps it, which the superset promise allows. - - The text is plain, not Markdown: :meth:`GlpiContentConverter.from_transport` - spells it through :func:`_literal_markdown` before handing it back, so - it is escaped exactly as the converting path escapes literal text. - - What it does **not** reproduce, none of which loses a character of - prose: - - * Markup that only the converter can express: a link becomes its text - without the target, an image contributes nothing, and a fenced - block loses its fence -- so ``<pre>`` indentation is normalised - away with the rest. A pasted log comes back as its own lines of - text, not as a code block. - * Whitespace is normalised harder. Runs of spaces collapse, and - `` `` counts as whitespace, so `` ``-padded column - alignment does not survive. - * Character references are resolved even inside a region the parser - handed back as raw data, so a broken comment's ``&`` comes back - as ``&``. In the other direction, a handful of semicolon-less - references stay literal here that the converter resolves -- see - :func:`_resolve_references`, which errs that way on purpose. - * Whitespace falls differently at a markup boundary, in both - directions: the converter joins ``a<b>c`` as ``a**c**`` where this - joins it as ``ac``, and this breaks a line at a block edge the - converter runs together. Which is why the property is about the - order of the characters of prose and not about where the spaces - land. - - Parameters - ---------- - content : str - Raw HTML. - - Returns - ------- - str - The document's text, block boundaries preserved as line breaks - and character references resolved. - """ - - if ">" not in content: - # No ``>`` means no markup to skip, so the whole document is the - # text the parser would flush on ``close()``. Answering it here - # also keeps the scan away from the one shape that costs - # ``html.parser`` more than linear time: with no ``>`` to finish a - # tag, ``close()`` advances one character at a time and rescans - # the tail, and 32 KB of an unfinished tag takes 13 seconds. - text = _resolve_references(content) - else: - text = _resolve_references("".join(_scan(content, collect_text=True).pieces)) - text = re.sub(r"[^\S\n]*\n[^\S\n]*", "\n", text) - text = re.sub(r"[^\S\n]{2,}", " ", text) - text = re.sub(r"\n{3,}", "\n\n", text) - return text.strip() - - -def _canonicalise_void_elements(content: str) -> str: - """Rewrite ``<br />`` as ``<br>``, so the text after it is not lost. - - A workaround for a ``beautifulsoup4`` defect (measured on 4.14.3), - reachable from ordinary editor output and silent when it fires. - - ``bs4``'s ``html.parser`` builder auto-closes a bare ``<br>`` and - records the name in ``already_closed_empty_element``, a list keyed by - name alone, so a later ``</br>`` can be ignored as redundant. If no - ``</br>`` ever arrives the entry simply stays there. The next - ``<br />`` -- which reaches the builder as ``handle_startendtag`` -- - opens a real element and then closes it itself, and *that* close - finds the stale entry, treats the element as already closed, and - leaves it open. Every following sibling becomes a child of the - ``<br>``. - - ``get_text`` still walks those children, which is why the tree looks - intact, but ``markdownify``'s ``convert_br`` ignores an element's - children and returns a line break. The text is gone: - - ``"<p>line1<br>line2</p><p>para2<br />line4</p>"`` converted to - ``"line1 \\nline2\\n\\npara2"`` -- and note the two spellings are in - different paragraphs, because a name once recorded poisons the rest - of the document. ``<img>`` and ``<hr>`` lose text the same way; they - are the other two converters that discard children. One bare ``<br>`` - anywhere before one ``<br />`` is the whole precondition, and GLPI - bodies are edited by more than one client. - - Rewriting to the bare spelling removes the ``handle_startendtag`` - path for void elements, which is where the asymmetry lives; both - spellings already build the same node, so nothing else about the - output moves. Only the names in :data:`_VOID_ELEMENTS` are touched, - and only where the parser really reports ``handle_startendtag``: a - self-closed ``<div/>`` is left alone, and cannot be affected anyway, - since only a void name is ever recorded. - - Which tags those are is the parser's answer rather than this - module's, and that is not cosmetic. Deciding it by pattern meant - inheriting every way the pattern could be derailed, and a derailed - scan reinstates the very defect this works around: measured, - ``"<p>one<br>two</p><script>x</ script><p>three<br />TAIL</p>"`` - lost ``TAIL`` outright, because ``</ script>`` ends raw text for the - parser but not for the pattern, so the ``<br />`` after it was never - seen and never rewritten. ``<img>`` and ``<hr>`` lost their tails the - same way. - - Parameters - ---------- - content : str - Raw HTML. - - Returns - ------- - str - The same HTML with self-closing void tags written bare. - """ - - if "/>" not in content: - return content - spans = _scan(content).void_spans - if not spans: - return content - pieces: list[str] = [] - cursor = 0 - for start, end in spans: - pieces.append(content[cursor:start]) - pieces.append(_VOID_SELF_CLOSE.sub(">", content[start:end])) - cursor = end - pieces.append(content[cursor:]) - return "".join(pieces) - - -# --------------------------------------------------------------------------- -# Literal text -# --------------------------------------------------------------------------- - -#: The characters whose reading as Markdown depends on what surrounds them. -#: -#: Every one of them is literal in some positions and syntax in others -- -#: ``#`` is a heading only at the start of a line, ``_`` is emphasis only -#: when a partner closes it -- so a text node cannot decide them alone. Each -#: is carried as a private stand-in (:class:`_StandIns`) until the container -#: it lands in is assembled, and decided there. -_LITERAL = "\\<&*_`[]!#>-+.=|~" - -#: How a literal character is spelled when it would be misread, where that -#: is not a backslash in front of it. -#: -#: ``<`` and ``&`` have no backslash escape python-markdown honours -- ``\<`` -#: renders both characters and the ``<`` still opens a tag -- and neither do -#: ``=`` and ``~``. A character reference is the spelling it passes through -#: as the character. -_SPELLED_OUT = { - "\\": "\\\\", - "<": "<", - "&": "&", - "=": "=", - "~": "~", -} - -#: The characters python-markdown removes a backslash from, with these -#: extensions. Asked of python-markdown rather than written down, because a -#: backslash in front of anything else is displayed. -_ESCAPABLE = frozenset(Markdown(extensions=_MARKDOWN_EXTENSIONS).ESCAPED_CHARS) - -# python-markdown's own patterns, compiled the way its inline processors -# compile theirs (``re.DOTALL``: a code span or an e-mail autolink can run -# across a line break). Taken from the renderer rather than copied, so the -# reader cannot disagree with the version doing the rendering. -_CODE_SPAN = re.compile(BACKTICK_RE, re.DOTALL) -_STANDALONE = re.compile(NOT_STRONG_RE, re.DOTALL) -_AUTOMAIL = re.compile(AUTOMAIL_RE, re.DOTALL) -_AUTOLINK = re.compile(AUTOLINK_RE) -_RULE = re.compile(HRProcessor.RE) -_REFERENCE_DEFINITION = ReferenceProcessor.RE - -#: What follows an ``&`` that the renderer displays as a character. -#: -#: A named reference needs its semicolon. A numeric one does not: -#: python-markdown runs ``html.parser`` over its whole source to find raw -#: HTML, and that re-emits ``ᆩ`` as ``ᆩ`` -- measured, 3.10.3. -_REFERENCE_TAIL = re.compile(r"\#[0-9]|\#[xX][0-9a-fA-F]|[0-9A-Za-z]+;") - -#: What makes a ``<`` open a tag, a comment, a declaration or an instruction. -_TAG_START = re.compile(r"[A-Za-z/!?]") - -_LINE_BREAK = re.compile(" \n") -_WORD = re.compile(r"\w") -_SETEXT_UNDERLINE = re.compile(r"(?:=+|-+) *") -_ORDERED_MARKER = re.compile(r"\d+(?=\.[ ])") -_LEADING_BLANK_LINES = re.compile(r"\A(?:[ \t]*\n)+") -_WHITESPACE = re.compile(r"\s+") -_EDGE = " \t\r\n" - -#: A character no analysis below treats as syntax: not a word character, -#: not whitespace, not punctuation python-markdown reads. Stands in for -#: anything already decided -- an escaped character, a code span. -_NEUTRAL = "\x02" - -#: Elements ``markdownify`` leaves inline that a browser shows as blocks. -_EXTRA_BLOCKS = frozenset({"center", "dir", "menu"}) - - -class _StandIns: - """Private characters carrying literal text until its container spells it. - - ``markdownify`` calls ``escape`` on one text node at a time, and a text - node cannot see what it will end up beside: whether its ``#`` starts a - line, whether its ``*`` has a partner in the next node, whether its - ``<`` is followed by a letter from a ``<span>``. So escaping replaces - each character in :data:`_LITERAL` with a stand-in, and the converter - for the enclosing block -- a paragraph, a list item, a cell, the - document -- decides every stand-in in it at once, with the whole - assembled Markdown of that block in view (:class:`_Spelling`). - - The stand-ins are a run of consecutive supplementary private-use code - points the source does not already contain, so a body holding - private-use characters of its own -- icon fonts map symbols there -- - cannot be confused with them. One more code point marks a hard line - break, whose surrounding spaces are settled the same way, and one more - starts a line no list item indents (:attr:`lazy`). - - Parameters - ---------- - source : str - The text about to be converted, which the stand-ins must not occur - in. - """ - - def __init__(self, source: str) -> None: - size = len(_LITERAL) + 2 - for base in range(0xF0000, 0xFFFFE - size, size): - characters = [chr(base + offset) for offset in range(size)] - if not any(character in source for character in characters): - break - else: # pragma: no cover - needs a source using all of plane 15 - raise GlpiContentError( - "Could not convert GLPI HTML content to Markdown: it uses " - "every private-use code point the converter could work with." - ) - literal = characters[: len(_LITERAL)] - self.hard_break = characters[-2] - #: Starts a line that stays at the start of the line: a quote that - #: opens a list item is read by python-markdown only if its later - #: lines are its lazy continuation, since the item's first block is - #: never detabbed and ``>`` counts only three spaces in at most. - self.lazy = characters[-1] - self.characters = frozenset(literal) - self._shadow = str.maketrans(dict(zip(_LITERAL, literal, strict=True))) - self._restore = str.maketrans(dict(zip(literal, _LITERAL, strict=True))) - self._any = re.compile("[" + "".join(literal) + "]") - self._break = re.compile("[ ]*" + self.hard_break + "\n?[ ]*") - - def shadow(self, text: str) -> str: - """Return ``text`` with every character in :data:`_LITERAL` stood in for.""" - - return text.translate(self._shadow) - - def restore(self, text: str) -> str: - """Return ``text`` with every stand-in back as the character it is.""" - - return text.translate(self._restore) - - def carried(self, text: str) -> bool: - """Return whether ``text`` still holds an undecided stand-in.""" - - return self._any.search(text) is not None - - def positions(self, text: str) -> set[int]: - """Return where ``text`` holds a stand-in.""" - - return {found.start() for found in self._any.finditer(text)} - - def replace(self, text: str, spell: Callable[[int, str], str]) -> str: - """Replace each stand-in with ``spell(position, character)``.""" - - return self._any.sub( - lambda found: spell(found.start(), self.restore(found.group(0))), text - ) - - def settle_breaks(self, text: str) -> str: - """Spell every hard line break ``" \\n"``, dropping the spaces around it. - - A browser does not display whitespace at either side of a ``<br>``, - so it is not content, and keeping it made the same body read two - ways: ``a<br> b`` read as ``a \\n b`` the first time and as - ``a \\nb`` once python-markdown had rendered it. - """ - - if self.hard_break not in text: - return text - return self._break.sub(" \n", text) - - -def _spelled(character: str) -> str: - """Return the escaped spelling of one literal character.""" - - return _SPELLED_OUT.get(character) or "\\" + character - - -class _Spelling: - """One container's Markdown, and which of its literal characters to escape. - - Built over the container's assembled text, where a stand-in is literal - text and anything else is markup ``markdownify`` generated or text an - inner container already decided. Each ``settle_*`` method asks one of - python-markdown's questions of the text as it will be rendered, and - marks the literal characters that would be read as syntax; - :meth:`render` spells them. - - The order the methods run in is python-markdown's own inline order -- - code spans before escapes before links before emphasis -- because each - of those consumes text the next one would otherwise see. - - Parameters - ---------- - text : str - The container's text, stand-ins included. - stand_ins : _StandIns - The stand-ins in use. - """ - - def __init__(self, text: str, stand_ins: _StandIns) -> None: - self.text = text - self.stand_ins = stand_ins - self.view = stand_ins.restore(text) - self.literal = stand_ins.positions(text) - self.escaped: set[int] = set() - self.hidden: set[int] = set() - - # -- bookkeeping --------------------------------------------------------- - - def escape(self, position: int) -> None: - """Escape the character at ``position``, if it is literal.""" - - if position in self.literal: - self.escaped.add(position) - - def live(self, position: int) -> bool: - """Return whether ``position`` is literal and not yet escaped.""" - - return position in self.literal and position not in self.escaped - - def render(self) -> str: - """Return the container's Markdown, every literal character spelled.""" +def _plain_text_html(text: str) -> str: + """Return the HTML that displays ``text`` as GLPI does: literally, line by line.""" - escaped = self.escaped + lines = (escape(line, quote=False) for line in text.splitlines()) + return "<p>" + "<br>".join(lines) + "</p>" - def spell(position: int, character: str) -> str: - return _spelled(character) if position in escaped else character - return self.stand_ins.replace(self.text, spell) +def _soup(html: str) -> BeautifulSoup: + """Parse ``html`` and drop what a browser does not display.""" - def markdown(self, start: int, end: int) -> tuple[str, list[int]]: - """Return one span's Markdown as it stands, and each character's position.""" + soup = BeautifulSoup(html, "html.parser") + for hidden in soup.find_all(_HIDDEN): + hidden.extract() + return soup - pieces: list[str] = [] - origin: list[int] = [] - cursor = start - for position in sorted(p for p in self.escaped if start <= p < end): - pieces.append(self.view[cursor:position]) - origin.extend(range(cursor, position)) - form = _spelled(self.view[position]) - pieces.append(form) - origin.extend([position] * len(form)) - cursor = position + 1 - pieces.append(self.view[cursor:end]) - origin.extend(range(cursor, end)) - return "".join(pieces), origin - def working(self, *, line_breaks: bool = False) -> str: - """Return the view with everything already decided made neutral. - - With ``line_breaks``, a hard break is neutral too: python-markdown - replaces ``" \\n"`` with a placeholder before it looks for emphasis, - so the characters either side of one are not next to whitespace. - """ - - chars = list(self.view) - for position in self.hidden | self.escaped: - chars[position] = _NEUTRAL - working = "".join(chars) - if line_breaks: - working = _LINE_BREAK.sub(_NEUTRAL * 3, working) - return working - - # -- what depends only on the next few characters ---------------------- - - def settle_local(self) -> None: - """Decide ``\\``, ``<`` and ``&`` by the characters right after them. - - * ``\\`` before a character python-markdown would treat as escaped - (:data:`_ESCAPABLE`) -- ``\\\\serveur``, ``C:\\_temp``; - * ``<`` before a letter, ``/``, ``!`` or ``?`` -- anything - python-markdown would pass through as raw HTML; - * ``&`` opening a character reference python-markdown would pass - through for the browser to decode (:data:`_REFERENCE_TAIL`). - - A ``<`` opening an e-mail autolink is decided once the inline - constructs it could run through are known (:meth:`settle_automail`). - """ - - view = self.view - for position in self.literal: - char = view[position] - following = view[position + 1 : position + 2] - if char == "\\": - if following and following in _ESCAPABLE: - self.escape(position) - elif char == "<": - if following and _TAG_START.match(following): - self.escape(position) - elif char == "&" and _REFERENCE_TAIL.match(view, position + 1): - self.escape(position) - - def settle_automail(self) -> None: - """Escape every literal ``<`` that would open an e-mail autolink. - - python-markdown looks for ``<address@domain>`` only once code spans, - escapes, links, images and URL autolinks have each become a - placeholder, so an address can run through any of them: - ``<3![a b](x.png)@c>`` is one to it. The pattern is matched against - the text with each of those collapsed to one character, and right to - left, because a match may also run through a ``<`` decided to its - right: once that one is ``<``, nothing stops it. - - A link's text is read the same way on its own, since python-markdown - parses it after the link is found; an image's text is never parsed. - """ - - view = self.view - opening = sorted( - (p for p in self.literal if view[p] == "<" and p not in self.escaped), - reverse=True, - ) - if not opening or "@" not in view: - return - working = list(self.working()) - contexts: dict[tuple[int, int], _Collapsed] = {} - for position in opening: - start, end = 0, len(view) - while True: - if (start, end) not in contexts: - contexts[start, end] = self._collapsed(working, start, end) - context = contexts[start, end] - link = context.link_at(position) - if link is None: - at = context.index[position - start] - if _automail_at(context.text, context.tail, at): - self.escape(position) - context.tail[at] = "&" - break - if link.label is None or not link.label[0] <= position < link.label[1]: - break - start, end = link.label - - def _collapsed(self, working: list[str], start: int, end: int) -> _Collapsed: - """Return ``view[start:end]`` with every placeholder collapsed.""" - - view = self.view - links = _generated_links(view, working, self.literal, start, end) - masked = self.hidden | self.escaped - for link in links: - masked = masked | set(range(*link.span)) - pieces: list[str] = [] - index: list[int] = [] - for position in range(start, end): - if position not in masked: - pieces.append(view[position]) - elif position == start or position - 1 not in masked: - pieces.append(_NEUTRAL) - index.append(len(pieces) - 1) - return _Collapsed("".join(pieces), index, links) - - # -- python-markdown's inline patterns, in its own order --------------- - - def settle_bang(self) -> None: - """Escape a literal ``!`` right before a generated link's ``[``. - - Otherwise ``Attention!`` followed by a link renders an image. - """ +#: The parts of a table other than its cells. +_TABLE_PARTS = frozenset( + {"table", "thead", "tbody", "tfoot", "tr", "caption", "colgroup", "col"} +) - view = self.view - for position in self.literal: - if ( - view[position] == "!" - and view[position + 1 : position + 2] == "[" - and position + 1 not in self.literal - ): - self.escape(position) - - def settle_backticks(self, units: Iterable[tuple[int, int]]) -> None: - """Escape every literal backtick run python-markdown would pair. - - Found by running python-markdown's own code-span pattern over each - unit's Markdown as it stands, escaping the literal runs it paired, - and running it again until it pairs none: escaping an opener can hand - its closer to the next run. A literal backtick touching a generated - one is escaped first, since it would join the generated run and - change the length the closer is matched on. What is left paired is - the generated code spans, which the later passes must not look - inside. - """ - view = self.view - if "`" not in view: - return - for position in self.literal: - if view[position] == "`": - for neighbour in (position - 1, position + 1): - if ( - 0 <= neighbour < len(view) - and view[neighbour] == "`" - and neighbour not in self.literal - ): - self.escape(position) - for start, end in units: - if "`" in view[start:end]: - self._settle_code_spans(start, end) - - def _settle_code_spans(self, start: int, end: int) -> None: - while True: - markdown, origin = self.markdown(start, end) - offending: set[int] = set() - generated: list[tuple[int, int]] = [] - for opener, closer in _code_spans(markdown): - delimiters = [origin[i] for i in range(*opener)] - delimiters += [origin[i] for i in range(*closer)] - live = [position for position in delimiters if self.live(position)] - if live: - offending.update(live) +def _flatten_nested_tables(root: Tag) -> None: + """Write a table that sits inside a cell as its cells' text. + + A GFM cell holds one line, so a nested table -- the usual layout of an + e-mail signature -- written as a table split the outer row and lost + every word. Inside a cell, a table's parts become ``span`` and its cells + ``glpi-cell``, which :class:`_Converter` writes as spaced inline text. + Renaming in one walk keeps it linear however deeply tables nest. + """ + + cells = 0 + counted: list[bool] = [] + for node, entering in _walk(root): + if not isinstance(node, Tag): + continue + if not entering: + if counted.pop(): + cells -= 1 + continue + is_cell = node.name in ("td", "th") + if cells and (is_cell or node.name in _TABLE_PARTS): + node.name = "glpi-cell" if is_cell else "span" + counted.append(False) + continue + counted.append(is_cell) + cells += is_cell + + +def _walk(root: Tag) -> Iterator[tuple[PageElement, bool]]: + """Yield the nodes under ``root`` in document order, each tag in and out. + + A tag comes with ``True`` on the way in and ``False`` on the way out. + Comments and declarations are skipped. The walk keeps its own stack, so + it costs nothing in recursion however deep the document is. + """ + + stack = [(node, True) for node in reversed(root.contents)] + while stack: + node, entering = stack.pop() + if isinstance(node, PreformattedString): + continue + yield node, entering + if entering and isinstance(node, Tag): + stack.append((node, False)) + stack.extend((child, True) for child in reversed(node.contents)) + + +def _drop_trailing_breaks(root: Tag) -> None: + """Turn every ``<br>`` its block shows nothing after into a ``<wbr>``. + + A browser shows no line for such a break, and CommonMark shows a + backslash break that ends a block as a backslash. ``<wbr>`` displays + nothing either, and renaming costs nothing where removing a node from a + long run of siblings costs the length of the run. + """ + + pending: list[Tag] = [] + for node, entering in _walk(root): + if not isinstance(node, Tag): + if str(node).strip(): + pending.clear() + elif node.name in _BLOCKS: + for line_break in pending: + line_break.name = "wbr" + pending.clear() + elif entering and node.name == "br": + pending.append(node) + elif entering and node.name == "img": + pending.clear() + for line_break in pending: + line_break.name = "wbr" + + +def _note_sides(root: Tag) -> None: + """Note on each bold or italic element the characters displayed on either side. + + One pass: ``last`` is the last character shown so far on the current + line, and an element that has closed waits for the next one. A line edge + counts as a space, as it does for CommonMark. + """ + + last = " " + waiting: list[Tag] = [] + for node, entering in _walk(root): + if isinstance(node, Tag): + if node.name in _EMPHASIS: + if entering: + node[_BEFORE] = last else: - generated.append((origin[opener[0]], origin[closer[1] - 1] + 1)) - if not offending: - for first, last in generated: - self.hidden.update(range(first, last)) - return - for position in offending: - self.escape(position) - - def hide_escapes(self) -> None: - """Hide what a backslash already escaped in generated or decided text. - - Pairs are consumed left to right, as python-markdown's escape - pattern consumes them, so ``\\\\`` hides both backslashes and escapes - nothing after it. - """ - - view = self.view - position = 0 - while position < len(view): - if ( - view[position] == "\\" - and position not in self.hidden - and position not in self.literal - ): - following = position + 1 - if ( - following < len(view) - and view[following] in _ESCAPABLE - and following not in self.literal - ): - self.hidden.update((position, following)) - position += 2 + waiting.append(node) continue - position += 1 - - def settle_brackets(self, units: Iterable[tuple[int, int]]) -> None: - """Escape every literal ``[`` that would open a link or an image. - - A ``[`` opens one when its balanced ``]`` is followed straight away - by ``(``, which is python-markdown's test. Right to left, because - escaping an inner ``[`` rebalances the brackets of an outer one; one - pass is enough, since the brackets to a ``[``'s right are settled - before it is. - - The pass keeps the ``]`` not yet closed by a ``[`` to their right on - a stack, so each ``[`` finds its partner on top of it: the first - ``]`` python-markdown's own count would stop at. An escaped ``[`` - hands its ``]`` back, to be closed by one further left. That keeps - the pass linear, where counting forward from each ``[`` cost a body - of unclosed brackets quadratic time. - """ - - working = list(self.working()) - for start, end in units: - closers: list[int] = [] - for position in range(end - 1, start - 1, -1): - char = working[position] - if char == "]": - closers.append(position) - elif char == "[" and closers: - close = closers.pop() - if ( - self.live(position) - and close + 1 < end - and working[close + 1] == "(" - ): - self.escape(position) - working[position] = _NEUTRAL - closers.append(close) - - def settle_asterisks(self, units: Iterable[tuple[int, int]]) -> None: - """Escape literal ``*`` wherever a unit holds two possible delimiters. - - python-markdown pairs asterisks with no regard for word boundaries, - so any two in one block -- literal or generated -- can pair. The one - exemption is its own: a run of one to three that is literal through - and through and stands alone between whitespace is text to it - (``5 * 3``). A run that touches a generated delimiter is not alone. - """ - - working = self.working(line_breaks=True) - for start, end in units: - unit = working[start:end] - exempt: set[int] = set() - for found in _STANDALONE.finditer(unit): - run = range(start + found.start(3), start + found.end(3)) - if found.group(3).startswith("*") and all( - position in self.literal for position in run - ): - exempt.update(run) - delimiters = [ - start + offset - for offset, char in enumerate(unit) - if char == "*" and start + offset not in exempt - ] - if len(delimiters) >= 2: - for position in delimiters: - self.escape(position) - - def settle_underscores(self, units: Iterable[tuple[int, int]]) -> None: - """Escape literal ``_`` runs that python-markdown could pair. - - Underscores are only emphasis at a word boundary, which is what - keeps ``fichier_de_test_v2.xlsx`` literal, except in a run of three - or more, which may open emphasis anywhere and take a mid-word run as - its closer; a single run of seven pairs with itself. When a unit - holds a pairing, every run that could take part in one is escaped. - """ - - working = self.working(line_breaks=True) - for start, end in units: - unit = working[start:end] - exempt: set[int] = set() - for found in _STANDALONE.finditer(unit): - if found.group(3).startswith("_"): - exempt.update(range(start + found.start(3), start + found.end(3))) - runs = [ - (start + found.start(), start + found.end()) - for found in re.finditer("_+", unit) - if start + found.start() not in exempt - and any( - self.live(position) - for position in range(start + found.start(), start + found.end()) - ) - ] - if not _underscores_pair(runs, working, start, end): + if node.name not in _BLOCKS and node.name != "br": continue - triple = any(finish - begin >= 3 for begin, finish in runs) - for begin, finish in runs: - left = begin == start or not _WORD.match(working[begin - 1]) - right = finish == end or not _WORD.match(working[finish]) - if left or right or triple: - for position in range(begin, finish): - self.escape(position) - - def settle_inline(self, units: list[tuple[int, int]]) -> None: - """Run every inline question, in python-markdown's order.""" - - self.settle_bang() - self.settle_backticks(units) - self.hide_escapes() - self.settle_brackets(units) - self.settle_automail() - self.settle_asterisks(units) - self.settle_underscores(units) - - -#: The longest ``<...>`` checked for an e-mail autolink. An address is at -#: most 254 characters; the pattern has no bound of its own, and checking -#: every ``<`` to the end of a long body would cost that body quadratically. -_AUTOMAIL_SPAN = 320 - - -def _automail_at(view: str, tail: list[str], position: int) -> bool: - """Return whether an e-mail autolink opens at ``position``. - - ``tail`` is the view with every ``<`` already decided to the right - spelled as a character that is not one. The pattern cannot cross a - space, and ends at the first ``>``. - """ - - end = view.find(">", position, position + _AUTOMAIL_SPAN) - if end < 0 or " " in view[position:end]: - return False - return _AUTOMAIL.match("".join(tail[position : end + 1])) is not None - - -def _underscores_pair( - runs: list[tuple[int, int]], working: str, start: int, end: int -) -> bool: - """Return whether python-markdown could pair any of ``runs``. - - ``EM_STRONG2`` and ``STRONG_EM2`` take their inner text from anywhere, - the run itself included, so a run of seven is a match on its own and a - run of three or more pairs with any other run. The ``SMART`` patterns - need an opener at a left word boundary and a later closer at a right - one. - """ - - if any(finish - begin >= 7 for begin, finish in runs): - return True - if len(runs) < 2: - return False - if any(finish - begin >= 3 for begin, finish in runs): - return True - opened = False - for begin, finish in runs: - if opened and (finish == end or not _WORD.match(working[finish])): - return True - if begin == start or not _WORD.match(working[begin - 1]): - opened = True - return False - - -def _code_spans(markdown: str) -> list[tuple[tuple[int, int], tuple[int, int]]]: - """Return the code spans python-markdown finds, as delimiter spans. - - Replays its backtick processor, which replaces each match with a - placeholder and searches on from just after it. After a code span that - is the same as searching on from the span's end: the one look-behind - there is ``(?<!\\\\)``, and the span ends in a backtick. A run of escaped - backslashes before a backtick is consumed on its own, though, and its - placeholder is what leaves the backtick after it free to open -- so - that run, and only that, is neutralised in the text searched. - """ - - spans: list[tuple[tuple[int, int], tuple[int, int]]] = [] - text = markdown - position = 0 - while True: - found = _CODE_SPAN.search(text, position) - if found is None: - return spans - if found.group(3): - size = len(found.group(2)) - spans.append( - ( - (found.start(2), found.end(2)), - (found.end(3), found.end(3) + size), - ) - ) - position = found.end() + first = last = " " else: - start, end = found.span(1) - text = text[:start] + _NEUTRAL * (end - start) + text[end:] - position = end - - -class _Link(NamedTuple): - """A link, an image or a URL autolink the reader wrote, as view spans.""" - - span: tuple[int, int] - #: The link's text, which python-markdown parses on its own; ``None`` - #: for an image or an autolink, whose text it never parses. - label: tuple[int, int] | None - - -class _Collapsed: - """A span of a view with every construct python-markdown stashes collapsed. - - Parameters - ---------- - text : str - The span, each stashed construct one character. - index : list of int - Where each position of the span went in ``text``. - links : list of _Link - The links, images and autolinks collapsed. - """ - - def __init__(self, text: str, index: list[int], links: list[_Link]) -> None: - self.text = text - self.index = index - self.links = links - #: ``text`` with every ``<`` decided so far spelled as a character - #: that opens nothing. - self.tail = list(text) - - def link_at(self, position: int) -> _Link | None: - """Return the link or image ``position`` is inside, if any.""" - - for link in self.links: - if link.span[0] <= position < link.span[1]: - return link - return None - - -def _generated_links( - view: str, working: list[str], literal: set[int], start: int, end: int -) -> list[_Link]: - """Return the links, images and URL autolinks the reader wrote in a span. - - A generated ``[`` is one that is not literal, and it opens a link when - its balanced ``]`` is followed by ``(``; the link ends at the ``)`` - that balances that one. A generated ``<`` opens an autolink when the - text up to the next ``>`` is one. - """ - - links: list[_Link] = [] - position = start - while position < end: - char = working[position] - if char == "[" and position not in literal: - close = _balanced_close(working, position + 1, end) - if close is not None and view[close + 1 : close + 2] == "(": - finish = _balanced_paren(view, close + 2, end) - if finish is not None: - image = position > start and view[position - 1] == "!" - links.append( - _Link( - (position - image, finish + 1), - None if image else (position + 1, close), - ) - ) - position = finish + 1 - continue - elif char == "<" and position not in literal: - finish = view.find(">", position + 1, end) - if finish > 0 and _AUTOLINK.fullmatch(view, position, finish + 1): - links.append(_Link((position, finish + 1), None)) - position = finish + 1 + text = str(node) + if not text: continue - position += 1 - return links - - -def _balanced_paren(view: str, start: int, end: int) -> int | None: - """Return where the ``(`` before ``start`` closes, before ``end``.""" - - depth = 1 - for position in range(start, end): - if view[position] == "(": - depth += 1 - elif view[position] == ")": - depth -= 1 - if depth == 0: - return position - return None - - -def _balanced_close(working: list[str], start: int, end: int) -> int | None: - """Return where the ``[`` before ``start`` closes, as python-markdown counts.""" + first, last = text[0], text[-1] + for element in waiting: + element[_AFTER] = first + waiting.clear() + for element in waiting: + element[_AFTER] = " " - depth = 1 - for position in range(start, end): - char = working[position] - if char == "]": - depth -= 1 - if depth == 0: - return position - elif char == "[": - depth += 1 - return None +def _punctuation(char: str) -> bool: + """Return whether CommonMark counts ``char`` as punctuation.""" -def _blocks(view: str) -> list[list[tuple[int, int]]]: - """Return the text's blocks -- runs of non-blank lines -- as line spans.""" - - blocks: list[list[tuple[int, int]]] = [] - current: list[tuple[int, int]] = [] - start = 0 - for line in view.split("\n"): - end = start + len(line) - if line.strip(" \t"): - current.append((start, end)) - elif current: - blocks.append(current) - current = [] - start = end + 1 - if current: - blocks.append(current) - return blocks - - -def _units( - view: str, blocks: list[list[tuple[int, int]]], *, soft_breaks: bool -) -> list[tuple[int, int]]: - """Return the spans python-markdown parses inline text in, one at a time. - - A paragraph is one: its lines are joined by hard breaks, ``" \\n"``. - Every other line break the reader writes starts something python-markdown - parses on its own -- a nested list, a quote -- even with no blank line - before it, so a unit ends there. Pairing is only possible inside a unit, - so this is what keeps a list item's ``*`` from being escaped for a ``*`` - in the list nested under it. Stripped text is the exception: its plain - line breaks continue a paragraph, which ``soft_breaks`` says. - """ + return char in string.punctuation or unicodedata.category(char).startswith("P") - units: list[tuple[int, int]] = [] - for block in blocks: - start, end = block[0] - for line_start, line_end in block[1:]: - if soft_breaks or view[end - 2 : end] == " ": - end = line_end - continue - units.append((start, end)) - start, end = line_start, line_end - units.append((start, end)) - return units - - -def _settle_block( - text: str, - stand_ins: _StandIns, - *, - in_list: bool, - top_level: bool, - lead: tuple[str, ...] = (), - soft_breaks: bool = False, -) -> str: - """Spell the literal text of one block container. - - ``text`` is the container's content before its own prefix is added -- - a list item's bullet, a quote's ``>`` -- so every line starts where - python-markdown will read it from. Each line is checked for the block - syntax python-markdown would find there, then the whole container for - the inline syntax (:meth:`_Spelling.settle_inline`). The line rules: - - * ``#`` at the very start of a line: an ATX heading, on any line; - * ``>`` after up to three spaces: a block quote, on any line; - * ``-``, ``+`` or ``*`` followed by a space, and digits followed by - ``.`` and a space, after up to three spaces: a list item -- on a - block's first line, or on any line inside a list item, where a - continuation line starts a nested list; - * a line of ``=`` or ``-`` alone as a block's second line: a setext - underline, spelled ``=`` or ``\\-``; - * three or more ``-``, ``*`` or ``_``, spaces allowed between: a rule, - on any line -- counting the bullet a list item is about to get; - * ``[label]: destination`` on a line: a reference definition, which - python-markdown consumes and which would turn every ``[label]`` in - the document into a link; - * three backticks or tildes at the start of a top-level line: a fence; - * a line of only ``|``, ``:``, ``-`` and spaces: a table's separator. - - Parameters - ---------- - text : str - The container's text. - stand_ins : _StandIns - The stand-ins in use. - in_list : bool - Whether the container is inside a list item. - top_level : bool - Whether its lines start at column 0 of the document, where a fence - can open. - lead : tuple of str, optional - The bullets that will precede the first line on its line, outermost - first (:func:`_line_lead`). - soft_breaks : bool, optional - Whether a plain line break continues a paragraph (see :func:`_units`). - - Returns - ------- - str - The same text with every stand-in spelled. - """ - if not stand_ins.carried(text): - return text - spelling = _Spelling(text, stand_ins) - spelling.settle_local() - view = spelling.view - blocks = _blocks(view) - for block_number, block in enumerate(blocks): - block_start, block_end = block[0][0], block[-1][1] - for number, (start, end) in enumerate(block): - line = view[start:end] - indent = len(line) - len(line.lstrip(" ")) - if line.startswith("#"): - spelling.escape(start) - if number == 1 and _SETEXT_UNDERLINE.fullmatch(line): - spelling.escape(start) - # python-markdown reads each item's content from just after its - # own bullet, so the first line is a rule if any tail of the - # bullets before it makes one: ``1. - --`` holds ``- --``. - heads = ["".join(lead[i:]) for i in range(len(lead))] - bulleted = heads if block_number == 0 and number == 0 else [] - if any(_RULE.match(head + line) for head in ["", *bulleted]): - for position in range(start, end): - if view[position] in "-*_" and position in spelling.literal: - spelling.escape(position) - if view[position] == "-": - break - if top_level and line.startswith(("```", "~~~")): - if line.startswith("~"): - spelling.escape(start) - else: - position = start - while position < end and view[position] == "`": - spelling.escape(position) - position += 1 - if "|" in line and set(line) <= set("|:- "): - for position in range(start, end): - if view[position] == "|": - spelling.escape(position) - if indent > 3 or indent == len(line): - continue - content = start + indent - first = view[content] - if first == ">": - spelling.escape(content) - if in_list or number == 0: - if first in "-+*" and view[content + 1 : content + 2] == " ": - spelling.escape(content) - ordered = _ORDERED_MARKER.match(view, content, end) - if ordered: - spelling.escape(ordered.end()) - if first == "[" and _REFERENCE_DEFINITION.match( - view[block_start:block_end], start - block_start - ): - spelling.escape(content) - spelling.settle_inline(_units(view, blocks, soft_breaks=soft_breaks)) - return spelling.render() - - -def _settle_inline( - text: str, stand_ins: _StandIns, *, cell: bool = False, heading: bool = False -) -> str: - """Spell the literal text of a table cell or a heading. - - Neither holds block syntax, so only the inline questions are asked, - plus one each: - - * in a cell, every ``|`` -- the table splits the row on it -- and every - backtick, because the table pairs backtick runs across the whole row - to decide which ``|`` it splits on; - * in a heading, the run of ``#`` it ends with, which python-markdown - strips as a closing sequence, and a final ``\\``, after which it - cannot match the heading line at all. - """ +def _flanks(inner: str, before: str, after: str) -> bool: + """Return whether CommonMark reads ``*`` runs around ``inner`` as emphasis. - if not stand_ins.carried(text): - return text - spelling = _Spelling(text, stand_ins) - spelling.settle_local() - view = spelling.view - if cell: - for position in spelling.literal: - if view[position] in "|`": - spelling.escape(position) - if heading: - position = len(view) - 1 - while position >= 0 and view[position] == "#": - spelling.escape(position) - position -= 1 - if view.endswith("\\"): - spelling.escape(len(view) - 1) - spelling.settle_inline([(0, len(view))]) - return spelling.render() - - -def _settle_label(text: str, stand_ins: _StandIns) -> str: - """Spell the brackets in a link's text or an image's alt text. - - python-markdown finds a link's text by counting brackets, so literal - brackets inside one are safe when they balance -- ``rapport - [final].pdf`` -- and end the text early when they do not. They are - left alone when balanced, and all escaped otherwise, or when they would - form an image or a link of their own inside the text. The label's other - characters are left to the block it is in. + ``inner`` neither starts nor ends with whitespace. A run opening onto + punctuation needs whitespace or punctuation in front of it, and a run + closing after punctuation needs the same behind it. """ - if not stand_ins.carried(text): - return text - probe = _Spelling(text, stand_ins) - view = probe.view - brackets = sorted(position for position in probe.literal if view[position] in "[]") - if not brackets: - return text - # Only generated code spans can hide a bracket: a literal backtick is - # either escaped later or pairs with nothing. - for position in probe.literal: - if view[position] in "`\\": - probe.escaped.add(position) - probe.settle_backticks([(0, len(view))]) - probe.hide_escapes() - working = list(probe.working()) - depth = 1 - broken = False - for char in working: - if char == "[": - depth += 1 - elif char == "]": - depth -= 1 - if depth == 0: - broken = True - break - broken = broken or depth != 1 - if not broken: - for opener in (position for position in brackets if view[position] == "["): - close = _balanced_close(working, opener + 1, len(working)) - if close is not None and working[close + 1 : close + 2] == ["("]: - broken = True - break - pieces = list(text) - for position in brackets: - pieces[position] = _spelled(view[position]) if broken else view[position] - return "".join(pieces) - - -def _loosened(item: str, loose: bool = False) -> str: - """Put a blank line after a list item's own text, when it is loose anyway. - - python-markdown makes an item *loose* -- its text wrapped in a paragraph - -- as soon as anything inside it is separated by a blank line: a code - block, a second paragraph, a nested item of either. An item following a - loose one is loose too. Reading that output back gives a blank line - between the item's text and the nested list after it, where the first - read had one line break; writing the blank line from the start makes the - first read the one every later read gives. - - The item's own text ends at its first line break that is not a hard - break, since that is the only other newline the reader writes there. - - Parameters - ---------- - item : str - The item's Markdown, before its bullet and indentation. - loose : bool, optional - Whether the item is loose whatever it holds: it follows a loose one. - """ + opens = not _punctuation(inner[0]) or before.isspace() or _punctuation(before) + closes = not _punctuation(inner[-1]) or after.isspace() or _punctuation(after) + return opens and closes - if not loose and "\n\n" not in item: - return item - position = 0 - while True: - position = item.find("\n", position) - if position < 0: - return item - if item[position - 2 : position] != " ": - break - position += 1 - if item[position + 1 : position + 2] == "\n": - return item - return item[:position] + "\n" + item[position:] +def _destination(url: str) -> str: + """Spell ``url`` as a link destination that reads back as ``url``.""" -#: The elements ``markdownify`` writes as a list. -_LIST_ELEMENTS = frozenset({"ul", "ol", "dir", "menu"}) + url = re.sub(r"[\t\n\r]", "", url) + return "<" + re.sub(r"[\\<>]", r"\\\g<0>", url) + ">" -#: The elements whose text is not displayed, and that the converter drops. -_UNSHOWN_ELEMENTS = frozenset({"script", "style", "title"}) -#: The elements displayed with no text of their own. -_SHOWN_EMPTY = frozenset({"img", "hr"}) +def _title(title: str) -> str: + """Spell a link or image title, with the space before it.""" -#: The blocks whose second line follows their first with no blank line. -_LINED_BLOCKS = _LIST_ELEMENTS | {"blockquote", "table"} + title = " ".join(title.split()) + return ' "' + re.sub(r'[\\"]', r"\\\g<0>", title) + '"' if title else "" -def _shows(node: PageElement, within: PageElement) -> bool: - """Return whether ``node`` is text, an image or a rule displayed in ``within``. +class _Converter(MarkdownConverter): + """``markdownify``, with what CommonMark and the reader's glue need on top. - ``markdownify`` skips comments and doctypes, and whitespace between - blocks; a ``<br>`` on its own starts no block, so it does not count. + The stub ``markdownify`` ships declares the constructor and ``convert`` + alone, so its converters are reached through :meth:`_inherited`. """ - if isinstance(node, Tag): - if node.name not in _SHOWN_EMPTY: - return False - elif isinstance(node, (Comment, Doctype)) or not str(node).strip(): - return False - for parent in node.parents: - if parent is within: - break - if parent.name in _UNSHOWN_ELEMENTS: - return False - return True - - -def _last_shown(node: PageElement) -> PageElement | None: - """Return the last thing ``node`` displays, or ``None`` if it displays nothing.""" - - if isinstance(node, Tag) and node.name in _UNSHOWN_ELEMENTS: - return None - last = node - while isinstance(last, Tag) and last.contents: - last = last.contents[-1] - for candidate in chain([last], last.previous_elements): - if _shows(candidate, node): - return candidate - if candidate is node: - break - return None - - -def _opens_with_a_block(item: Tag) -> bool: - """Return whether a list item's first line is a list's, a quote's or a table's. - - The item then has no text of its own to set apart with a blank line: - the line after its first is that block's, and a blank line there would - split the block in two. - """ - - for candidate in item.descendants: - if _shows(candidate, item): - for element in candidate.parents: - if element is item: - return False - if element.name in _LINED_BLOCKS: - return True - return False - - -def _ends_in_a_list(node: PageElement) -> bool | None: - """Return whether what ``node`` displays last is in a list, if it shows anything.""" - - shown = _last_shown(node) - if shown is None: - return None - for element in chain([shown], shown.parents): - if isinstance(element, Tag) and element.name in _LIST_ELEMENTS: - return True - if element is node: - break - return False - - -def _bullet(item: Tag, numbers: dict[int, int]) -> str: - """Return the marker the reader writes before a list item. - - ``numbers`` holds each ordered item's position among its list's items, - filled for a whole list the first time one of its items is asked for: - counting an item's previous siblings instead cost a long list quadratic - time. - """ + def _inherited(self, name: str) -> Any: + return getattr(super(), name) - parent = item.parent - if parent is not None and parent.name == "ol": - if id(item) not in numbers: - for index, entry in enumerate(parent.find_all("li", recursive=False)): - numbers[id(entry)] = index - start = str(parent.get("start") or "") - first = int(start) if start.isdigit() else 1 - return f"{first + numbers[id(item)]}. " - return "- " + def _markup( + self, el: Tag, text: str, parent_tags: set[str], markers: str, tag: str + ) -> str: + """Wrap ``text`` in ``markers``, or in raw ``<tag>`` where they would not close. + Line breaks and spaces at the edges move outside, where they cannot + stop the markers closing. + """ -def _line_items(element: Tag) -> list[Tag]: - """Return the list items whose bullets are written on ``element``'s first line. + edges = _EDGES.fullmatch(text) + if "_noformat" in parent_tags or edges is None or not edges[2]: + return text + before, inner, after = edges.groups() + prev = before[-1:] or str(el.get(_BEFORE) or " ") + succ = after[:1] or str(el.get(_AFTER) or " ") + if markers and _flanks(inner, prev, succ): + return f"{before}{markers}{inner}{markers}{after}" + return f"{before}<{tag}>{inner}</{tag}>{after}" - ``element`` itself if it is an item, and every item it opens, through - every block that opens the next, outermost first: ``<li><ul><li>-`` is - written ``- - -``. A quote ends the line's items, since python-markdown - parses a quote's content on its own. - """ + def convert_b(self, el: Tag, text: str, parent_tags: set[str]) -> str: + return self._markup(el, text, parent_tags, "**", "strong") - items = [element] if element.name == "li" else [] - node = element - while node.parent is not None: - if any(_last_shown(sibling) is not None for sibling in node.previous_siblings): - break - node = node.parent - if node.name == "blockquote": - break - if node.name == "li": - items.insert(0, node) - return items - - -def _deep_list(element: Tag | None) -> bool: - """Return whether a list's first line carries three bullets or more. - - python-markdown's first pass over a block takes every line after - ``- - - a`` that is indented eight spaces or more as that line's lazy - continuation, so nothing of such a list can follow its first line - directly: its items' own blocks, and its other items, each start a block - of their own, which makes every item of the list loose. - """ + convert_strong = convert_b - return element is not None and len(_line_items(element)) >= 2 + def convert_em(self, el: Tag, text: str, parent_tags: set[str]) -> str: + return self._markup(el, text, parent_tags, "*", "em") + convert_i = convert_em -def _shown_items(element: Tag) -> int: - """Return how many of a list's items display something.""" + def convert_s(self, el: Tag, text: str, parent_tags: set[str]) -> str: + """Struck text: CommonMark has no strikethrough, so always raw ``<s>``.""" - items = element.find_all("li", recursive=False) - return sum(_last_shown(item) is not None for item in items) + return self._markup(el, text, parent_tags, "", "s") + convert_del = convert_s + convert_strike = convert_s -def _line_lead(element: Tag, numbers: dict[int, int]) -> tuple[str, ...]: - """Return the bullets written on ``element``'s first line, its own included. + def convert_glpi_cell(self, el: Tag, text: str, parent_tags: set[str]) -> str: + """A cell of a table nested in a cell: its text, set apart by spaces.""" - ``- - -`` is a rule to python-markdown, so the first line of a block is - checked against its whole line. ``numbers`` is :func:`_bullet`'s. - """ + return f" {text.strip()} " - return tuple(_bullet(item, numbers) for item in _line_items(element)) + def convert_br(self, el: Tag, text: str, parent_tags: set[str]) -> str: + if "pre" in parent_tags: + return "\n" + if "td" in parent_tags or "th" in parent_tags: + return "<br>" # a cell is one line of Markdown; cmark-gfm passes the tag + return str(self._inherited("convert_br")(el, text, parent_tags)) + + def convert_code(self, el: Tag, text: str, parent_tags: set[str]) -> str: + code = str(self._inherited("convert_code")(el, text, parent_tags)) + if "td" in parent_tags or "th" in parent_tags: + code = code.replace( + "|", "\\|" + ) # GFM splits a row on "|" before it reads code + return code + def convert_list(self, el: Tag, text: str, parent_tags: set[str]) -> str: + """A list; a nested one ends in a blank line. -def _code_block_fits(element: Tag) -> bool: - """Return whether python-markdown reads an indented code block where ``element`` is. + Otherwise the item's text after it joins the nested list's last item. + """ - Inside a list item or a quote a ``<pre>`` can only be an indented code - block, and there are two places where python-markdown reads none: as the - item's first block, which is always its paragraph, and right after a - list in the same item or quote, since the code's indentation is then - exactly that of the list's last item, and its lines become a paragraph - of that item. - """ + markdown = str(self._inherited("convert_list")(el, text, parent_tags)) + return markdown + "\n\n" if "li" in parent_tags else markdown - node: Tag = element - while node.parent is not None: - for sibling in node.previous_siblings: - in_a_list = _ends_in_a_list(sibling) - if in_a_list is not None: - return not in_a_list - node = node.parent - if node.name == "li": - return False - if node.name == "blockquote": - return True - return True - - -class _LiteralSafeConverter(MarkdownConverter): - """``markdownify``'s converter, spelling literal text so it stays text. - - ``escape`` is replaced, and escapes nothing itself: it stands each - character of :data:`_LITERAL` in for (:class:`_StandIns`). The - converters for the containers python-markdown parses on their own -- a - paragraph, a ``<div>``, a list item, a quote, a heading, a table cell, - the document -- then spell every stand-in in their text at once, before - they add their own prefix, with the whole of it in view: - :func:`_settle_block`, :func:`_settle_inline` and :func:`_settle_label` - list the rules. A container inside a heading, a cell or a link leaves its - stand-ins to the one it sits in, which is the unit python-markdown will - parse. - - It also changes what ``markdownify`` produces wherever the result lost - or garbled content that literal-safe escaping would otherwise have kept: - - * a list item's continuation lines are indented by four spaces, which is - what python-markdown nests at, rather than by the bullet's width; - * whatever follows a nested list in its item -- text, a quote -- starts - after a blank line, where ``markdownify`` ran it into the list's last - item; a quote anywhere in an item after its text does too, and a - quote that opens an item keeps its later lines at the start of the - line, the only place python-markdown reads them (:attr:`_StandIns.lazy`); - * an item whose first line it shares with another's bullet is loose, - and so is every item of a list whose first line carries three bullets - or more (:func:`_deep_list`), because python-markdown's first pass - over the block cannot nest anything under such a line; - * an item written loose by python-markdown is written loose here too, so - the second read is the first one (:func:`_loosened`); - * ``<script>``, ``<style>`` and ``<title>`` bodies are dropped, as a - browser drops them; - * ``<s>``, ``<del>`` and ``<strike>`` keep their words without the - ``~~`` markers python-markdown would display; - * an image stays an image inside a heading or a cell, and its alt text - is one line; - * a ``<pre>`` inside a list item or a quote becomes an indented code - block, since a fence only opens at the start of a line, except where - python-markdown reads no code block at all (:func:`_code_block_fits`), - where its lines are kept as literal text; a top-level fence is made - longer than any fence line the code holds; - * a newline in text is a space, as HTML displays it. - """ + convert_ul = convert_list + convert_ol = convert_list - def __init__(self, stand_ins: _StandIns, **options: Any) -> None: - super().__init__(**options) - self.stand_ins = stand_ins - #: The list items whose Markdown ends in a blank line, by ``id``. - self._loose_items: set[int] = set() - #: Each ordered item's position in its list, by ``id`` (:func:`_bullet`). - self._numbers: dict[int, int] = {} - #: How many of a list's items display something, by the list's ``id``. - self._shown: dict[int, int] = {} + def convert_pre(self, el: Tag, text: str, parent_tags: set[str]) -> str: + """A fenced block, its fence longer than any backtick run in the code.""" - def _shown_items(self, element: Tag) -> int: - """Return :func:`_shown_items` for a list, counted once per list.""" + markdown = str(self._inherited("convert_pre")(el, text, parent_tags)) + longest = max((len(run) for run in _BACKTICK_RUN.findall(text)), default=0) + if longest < 3 or markdown.count("```") < 2: + return markdown + fence = "`" * (longest + 1) + head, _, rest = markdown.partition("```") + body, _, tail = rest.rpartition("```") + return f"{head}{fence}{body}{fence}{tail}" - if id(element) not in self._shown: - self._shown[id(element)] = _shown_items(element) - return self._shown[id(element)] + def convert_a(self, el: Tag, text: str, parent_tags: set[str]) -> str: + """A link; ``<url>`` when its text is its URL and reads back unchanged.""" - def _inherited(self, name: str) -> Callable[..., Any]: - """Return ``markdownify``'s own converter ``name``. + href = str(el.get("href") or "") + edges = _EDGES.fullmatch(text) + if "_noformat" in parent_tags or not href or edges is None or not edges[2]: + return text + before, inner, after = edges.groups() + title = str(el.get("title") or "") + if not title and el.get_text() == href and _AUTOLINK.fullmatch(href): + return f"{before}<{href}>{after}" + return f"{before}[{inner}]({_destination(href)}{_title(title)}){after}" - Its type stub declares the constructor and ``convert`` alone, so the - converters this class extends are reached by name. - """ + def convert_img(self, el: Tag, text: str, parent_tags: set[str]) -> str: + """An image, wherever it is: a cell or a heading holds one too.""" - method: Callable[..., Any] = getattr(super(), name) - return method + alt = " ".join(str(el.get("alt") or "").split()) + alt = str(self._inherited("escape")(alt, parent_tags)) + src = _destination(str(el.get("src") or "")) + return f"![{alt}]({src}{_title(str(el.get('title') or ''))})" - # -- text ----------------------------------------------------------------- - def escape(self, text: str, parent_tags: set[str]) -> str: - return self.stand_ins.shadow(text) if text else "" +class _Node(RenderTreeNode): + """mdformat's syntax-tree node, answering in constant time what it asks often. - def process_text(self, el: object, parent_tags: set[str] | None = None) -> str: - """Drop the whitespace next to an element displayed as a block. + mdformat asks every text node for its next sibling, every list for its + previous one, and every list item whether its list is tight. The base + class answers each by scanning, which made a long paragraph or a long + list quadratic. + """ - ``markdownify`` already does for the elements it knows are blocks; - :data:`_EXTRA_BLOCKS` are the ones it does not. - """ + _index = -1 - text = str(self._inherited("process_text")(el, parent_tags=parent_tags)) - before = getattr(el, "previous_sibling", None) - after = getattr(el, "next_sibling", None) - if isinstance(before, Tag) and before.name in _EXTRA_BLOCKS: - text = text.lstrip(_EDGE) - if isinstance(after, Tag) and after.name in _EXTRA_BLOCKS: - text = text.rstrip(_EDGE) - return text + def _position(self) -> int: + if self._index < 0: + for index, sibling in enumerate(self.siblings): + sibling._index = index + return self._index - # -- containers that settle their own literal text ------------------------ + @property + def next_sibling(self) -> _Node | None: + siblings = self.siblings + index = self._position() + 1 + return siblings[index] if index < len(siblings) else None - @staticmethod - def _deferred(parent_tags: set[str]) -> bool: - """Return whether an enclosing heading, cell or link settles instead.""" + @property + def previous_sibling(self) -> _Node | None: + index = self._position() - 1 + return self.siblings[index] if index >= 0 else None - return "_inline" in parent_tags or "a" in parent_tags + @cached_property + def tight(self) -> bool: + """Whether this list is tight: no paragraph in it is shown as one.""" - def _settle( - self, text: str, parent_tags: set[str], lead: tuple[str, ...] = () - ) -> str: - text = self.stand_ins.settle_breaks(text) - text = _LEADING_BLANK_LINES.sub("", text).rstrip(_EDGE) - return _settle_block( - text, - self.stand_ins, - in_list="li" in parent_tags, - top_level="li" not in parent_tags and "blockquote" not in parent_tags, - lead=lead, + return all( + grandchild.hidden + for child in self.children + for grandchild in child.children + if grandchild.type == "paragraph" ) - def convert__document_(self, el: Tag, text: str, parent_tags: set[str]) -> str: - """Settle the document, stripped first: edge whitespace is never content. - - Stripped before it is settled, not after, because where the first - line starts decides what it would be read as: ``" # x"`` is not a - heading until the space goes. It is also why the Markdown is stripped - at all -- a digest of a body must not depend on its edges. - """ - text = self.stand_ins.settle_breaks(text).strip() - return _settle_block(text, self.stand_ins, in_list=False, top_level=True) - - def convert_p(self, el: Tag, text: str, parent_tags: set[str]) -> str: - if self._deferred(parent_tags): - return str(self._inherited("convert_p")(el, text, parent_tags)) - lead = _line_lead(el, self._numbers) if "li" in parent_tags else () - text = self._settle(text, parent_tags, lead=lead) - return f"\n\n{text}\n\n" if text else "" - - def convert_div(self, el: Tag, text: str, parent_tags: set[str]) -> str: - if self._deferred(parent_tags): - return str(self._inherited("convert_div")(el, text, parent_tags)) - lead = _line_lead(el, self._numbers) if "li" in parent_tags else () - text = self._settle(text, parent_tags, lead=lead) - return f"\n\n{text}\n\n" if text else "" - - convert_article = convert_div - convert_center = convert_div - convert_dl = convert_div - convert_section = convert_div - # ``markdownify`` writes a definition as ": text" under its term, one - # line break apart: python-markdown, with no definition-list extension, - # shows the colon and runs the two into one paragraph. Two paragraphs - # keep both texts and nothing else. - convert_dd = convert_div - convert_dt = convert_div - - def convert_blockquote(self, el: Tag, text: str, parent_tags: set[str]) -> str: - if self._deferred(parent_tags): - return str(self._inherited("convert_blockquote")(el, text, parent_tags)) - text = self._settle(text or "", parent_tags | {"blockquote"}) - if not text: - return "\n" - lines = [f"> {line}" if line else ">" for line in text.split("\n")] - if "li" not in parent_tags: - return "\n{}\n\n".format("\n".join(lines)) - if _line_items(el): - # The quote opens a list item: python-markdown never detabs an - # item's first block, so the quote's later lines have to be its - # lazy continuation, at the start of the line. - lines[1:] = [self.stand_ins.lazy + line for line in lines[1:]] - return "\n\n{}\n\n".format("\n".join(lines)) - # Anywhere else in an item a quote opens only after a blank line: the - # item's later lines are indented, and ``>`` counts three spaces in - # at most. - return "\n\n{}\n\n".format("\n".join(lines)) - - def convert_li(self, el: Tag, text: str, parent_tags: set[str]) -> str: - bullet = _bullet(el, self._numbers) - spaced = False - if self._deferred(parent_tags): - text = (text or "").strip() - else: - text = self.stand_ins.settle_breaks(text or "") - text = _LEADING_BLANK_LINES.sub("", text).rstrip(_EDGE) - parent = el.parent - deep = _deep_list(parent) - spaced = deep and parent is not None and self._shown_items(parent) > 1 - # The item after a loose one is loose too: python-markdown reads - # it as the first item of a new block, and wraps its text. So is - # an item that shares its first line with another's bullet: a - # list after its text is eight spaces in, too deep to follow - # that line directly (see _deep_list). - previous = el.find_previous_sibling("li") - stacked = len(_line_items(el)) > 1 - loose = deep or stacked or id(previous) in self._loose_items - if not _opens_with_a_block(el): - text = _loosened(text, loose) - lead = _line_lead(el, self._numbers) - text = self._settle(text, parent_tags | {"li"}, lead=lead) - if not text: - return "\n" - blocks = "\n\n" in text or spaced - if blocks: - self._loose_items.add(id(el)) - first_line, *rest = text.split("\n") - lazy = self.stand_ins.lazy - indented = [ - line if not line or line.startswith(lazy) else f" {line}" - for line in rest - ] - item = "\n".join([bullet + first_line, *indented]) + "\n" - # An item of several blocks needs a blank line after it, or the next - # item's line is read as the last block's lazy continuation. - return item + "\n" if blocks else item - - def convert_hN(self, n: int, el: Tag, text: str, parent_tags: set[str]) -> str: - if self._deferred(parent_tags): - return str(self._inherited("convert_hN")(n, el, text, parent_tags)) - text = _WHITESPACE.sub(" ", text.strip()) - text = _settle_inline(text, self.stand_ins, heading=True) - return "\n\n{} {}\n\n".format("#" * max(1, min(6, n)), text) - - def convert_td(self, el: Tag, text: str, parent_tags: set[str]) -> str: - span = str(el.get("colspan") or "") - colspan = max(1, min(1000, int(span))) if span.isdigit() else 1 - text = _settle_inline( - text.strip().replace("\n", " "), self.stand_ins, cell=True - ) - return " " + text + " |" * colspan +def _list_item(node: RenderTreeNode, context: RenderContext) -> str: + """Render a list item as mdformat does, asking its list's tightness once.""" - convert_th = convert_td + separator = "\n" if cast(_Node, node.parent).tight else "\n\n" + text = separator.join( + filter(None, (child.render(context) for child in node.children)) + ) + return text if text.strip() else "" - # -- markup ---------------------------------------------------------------- - def _edges(self, text: str) -> tuple[str, str, str]: - """Split ``text`` into what goes before markup, inside it, and after. +#: ``2k`` backslashes, which mdformat writes for ``k`` literal ones, before a +#: character no backslash escapes. ``2k - 1`` read as the same ``k`` there. +_NEEDLESS_DOUBLING = re.compile(r"(?<!\\)((?:\\\\)+)(?=[^!-/:-@\[-`{-~\\\s])") - ``markdownify``'s ``chomp`` moves spaces outside the markup so it - never opens or closes on whitespace; hard breaks at the edge have to - move out the same way, or ``**a \\n**`` is not emphasis. - """ - edge = _EDGE + self.stand_ins.hard_break - content = text.strip(edge) - head = text[: len(text) - len(text.lstrip(edge))] - tail = text[len(text.rstrip(edge)) :] if content else "" - return self._edge(head), content, self._edge(tail) +def _text(node: RenderTreeNode, context: RenderContext) -> str: + """Render text as mdformat does, less one backslash where it escapes nothing. - def _edge(self, side: str) -> str: - breaks = side.count(self.stand_ins.hard_break) - if breaks: - return (self.stand_ins.hard_break + "\n") * breaks - return " " if side else "" + ``C:\\Temp`` stays ``C:\\Temp`` instead of becoming ``C:\\\\Temp``. + """ - def _emphasis(self, markup: str, text: str, parent_tags: set[str]) -> str: - if "_noformat" in parent_tags: - return text - before, content, after = self._edges(text) - if not content: - return before - return before + markup + content + markup + after + text = DEFAULT_RENDERERS["text"](node, context) + return _NEEDLESS_DOUBLING.sub(lambda found: found.group(1)[:-1], text) - def convert_b(self, el: Tag, text: str, parent_tags: set[str]) -> str: - return self._emphasis("**", text, parent_tags) - convert_strong = convert_b +class _Lists: + """An mdformat extension: linear list items, short text and rules.""" - def convert_em(self, el: Tag, text: str, parent_tags: set[str]) -> str: - return self._emphasis("*", text, parent_tags) + RENDERERS: Mapping[str, Callable[[RenderTreeNode, RenderContext], str]] = { + "list_item": _list_item, + "text": _text, + "hr": lambda node, context: "---", + } - convert_i = convert_em - - def convert_s(self, el: Tag, text: str, parent_tags: set[str]) -> str: - """Keep struck text's words: python-markdown has no strikethrough.""" - return text +class _Renderer(MDRenderer): + """mdformat's renderer, over :class:`_Node`.""" - convert_del = convert_s - convert_strike = convert_s + def render( + self, + tokens: Sequence[Token], + options: Mapping[str, Any], + env: MutableMapping[Any, Any], + *, + finalize: bool = True, + ) -> str: + return self.render_tree(_Node(tokens), options, env, finalize=finalize) - def convert_title(self, el: Tag, text: str, parent_tags: set[str]) -> str: - """Drop a ``<title>``, which a browser shows only in its tab.""" - return "" +class _Parser(MarkdownIt): + """markdown-it keeping every link: it reformats Markdown, never renders it.""" - def convert_br(self, el: Tag, text: str, parent_tags: set[str]) -> str: - if "_inline" in parent_tags: - return " " - if "pre" in parent_tags: - return "\n" - if "_noformat" in parent_tags: - return " " - return self.stand_ins.hard_break + "\n" + def validateLink(self, url: str) -> bool: + return True - def convert_a(self, el: Tag, text: str, parent_tags: set[str]) -> str: - if "_noformat" in parent_tags: - return text - before, text, after = self._edges(text) - if not text: - return before - href = el.get("href") - title = el.get("title") - if not href: - return before + text + after - href = str(href) - pasted = self.stand_ins.restore(text) == href - if pasted and not title and _AUTOLINK.fullmatch(f"<{href}>"): - # A pasted URL: the one autolink python-markdown renders as a link. - return f"{before}<{href}>{after}" - text = _settle_label(text, self.stand_ins) - titled = ' "{}"'.format(str(title).replace('"', r"\"")) if title else "" - return f"{before}[{text}]({href}{titled}){after}" - def convert_img(self, el: Tag, text: str, parent_tags: set[str]) -> str: - alt = _WHITESPACE.sub(" ", str(el.get("alt") or "")).strip() - alt = _settle_label(self.stand_ins.shadow(alt), self.stand_ins) - src = str(el.get("src") or "") - title = _WHITESPACE.sub(" ", str(el.get("title") or "")).strip() - titled = ' "{}"'.format(title.replace('"', r"\"")) if title else "" - return f"![{alt}]({src}{titled})" +def _formatter() -> MarkdownIt: + """Return mdformat as ``mdformat.text`` builds it, less what this module cannot use. - def convert_pre(self, el: Tag, text: str, parent_tags: set[str]) -> str: - if not text: - return "" - text = re.sub(r"[ \n]*$", "", re.sub(r"^[ \n]*\n", "", text)) - nested = "li" in parent_tags or "blockquote" in parent_tags - if nested and not _code_block_fits(el): - # No code block can be written here (see _code_block_fits), so - # the lines are kept as the literal text they are. - lines = [line.strip(" ") for line in text.split("\n")] - joined = (self.stand_ins.hard_break + "\n").join(filter(None, lines)) - return self.stand_ins.shadow(joined) - if nested: - lines = text.split("\n") - code = "\n".join(f" {line}" if line else "" for line in lines) - return f"\n\n{code}\n\n" - fence = "```" - lines = [line.rstrip(" ") for line in text.split("\n")] - while fence in lines: - fence += "`" - return f"\n\n{fence}\n{text}\n{fence}\n\n" + ``mdformat.text`` keeps markdown-it's nesting cap of 20, beyond which the + parser drops the rest of the block without a word: a list ten deep lost + its text. The cap is lifted, so a document too deep raises + ``RecursionError`` instead, which :meth:`GlpiContentConverter.from_transport` + answers. Link validation is lifted too, or a ``file:`` link read back as + escaped text. + """ - def convert_list(self, el: Tag, text: str, parent_tags: set[str]) -> str: - """End a nested list with a blank line, for whatever its item holds next. + parser = _Parser( + "commonmark", + {"maxNesting": 1_000_000}, + renderer_cls=cast(Any, _Renderer), + ) + parser.options["mdformat"] = { + "wrap": "keep", + "number": True, + "compact_tables": True, + } + parser.options["store_labels"] = True + parser.options["parser_extension"] = [mdformat_tables, _Lists] + parser.options["codeformatters"] = {} + mdformat_tables.update_mdit(parser) + return parser + + +# The shared objects keep no state between calls: markdownify fills a +# per-tag cache of its own methods, the same on every thread, and markdown-it +# and cmark-gfm parse into fresh state. What they set up once -- markdown-it's +# rule chains, cmark-gfm's extension registry, which has no lock -- is set up +# here, under the import lock, before any thread can race for it. +_CONVERTER = _Converter(**_MARKDOWNIFY_OPTIONS) +_FORMATTER = _formatter() +_FORMATTER.render("*warm* [up](x)\n\n- a\n\n| a |\n| - |") +cmarkgfm.cmark.core_extensions_ensure_registered() - ``markdownify`` ends a list inside a list item with no line break at - all, so the item's text after it joined the list's last item, and a - quote after it became that item's lazy continuation. When nothing - follows, the item strips the blank line with the rest of its edge. - """ - markdown = str(self._inherited("convert_list")(el, text, parent_tags)) - if "li" not in parent_tags or self._deferred(parent_tags): - return markdown - return markdown + "\n\n" +def html_to_markdown(html: str) -> str: + """Convert HTML to Markdown that cmark-gfm renders as the same display. - convert_ul = convert_list - convert_ol = convert_list - convert_dir = convert_list - convert_menu = convert_list + Needs ``beautifulsoup4`` 4.15 or later: before it, a ``<br />`` in a body + that also held a bare ``<br>`` swallowed the text after it. + Raises + ------ + RecursionError + ``markdownify`` or mdformat recursed deeper than the stack left. + """ -#: ``markdownify`` options; the converter's overrides decide the rest. -#: -#: ``wrap`` is what turns a newline inside text into a space, and -#: ``wrap_width=None`` is what stops it wrapping lines. -_MARKDOWNIFY_OPTIONS: dict[str, Any] = {"wrap": True, "wrap_width": None} + soup = _soup(html) + _flatten_nested_tables(soup) + _drop_trailing_breaks(soup) + _note_sides(soup) + return str(_FORMATTER.render(_CONVERTER.convert_soup(soup))).strip() -def html_to_markdown(html: str) -> str: - """Convert HTML to literal-safe Markdown. - - Parameters - ---------- - html : str - The HTML to convert. - - Returns - ------- - str - The Markdown, every literal character spelled so python-markdown - renders it as that character. - """ +def markdown_to_html(markdown: str) -> str: + """Render Markdown as HTML: CommonMark with GFM tables, through cmark-gfm.""" - stand_ins = _StandIns(html) - markdown = str( - _LiteralSafeConverter(stand_ins, **_MARKDOWNIFY_OPTIONS).convert(html) + html: str = cmarkgfm.markdown_to_html_with_extensions( + markdown, options=_RENDER_OPTIONS, extensions=["table"] ) - return markdown.replace(stand_ins.lazy, "") + return html.strip() -def _literal_markdown(text: str) -> str: - """Spell plain text as Markdown that renders as that text. - - The degraded path's final step: :func:`_strip_tags` returns text, and - :meth:`GlpiContentConverter.from_transport` returns Markdown, so the - text is escaped by the same rules as any text node -- every line a line - start, one container per blank-line-separated block. - """ +def _text_of(html: str) -> str: + """Return the text ``html`` displays, a line per block, without recursing.""" - stand_ins = _StandIns(text) - return _settle_block( - stand_ins.shadow(text), - stand_ins, - in_list=False, - top_level=True, - soft_breaks=True, - ) + pieces: list[str] = [] + for node, _ in _walk(_soup(html)): + if not isinstance(node, Tag): + pieces.append(str(node)) + elif node.name in _BLOCKS or node.name == "br": + pieces.append("\n") + lines = (" ".join(line.split()) for line in "".join(pieces).split("\n")) + return "\n".join(line for line in lines if line) class GlpiContentConverter: - """Convert content between GLPI HTML payloads and canonical Markdown. - - The converter keeps the translation rules in one place so ticket, followup, - task, and solution parsing all share the same content normalization. - """ + """Convert content between GLPI HTML payloads and canonical Markdown.""" @staticmethod - def from_transport(value: object) -> str: - """Convert one GLPI transport value into canonical Markdown. - - Empty input stays empty, plain text is preserved, and HTML content is - converted through ``markdownify`` into Markdown that python-markdown - renders back as the same text: the HTML's text is literal, so every - character that would otherwise read as syntax is escaped, and nothing - else is (see the module docstring and :class:`_LiteralSafeConverter`). - - The HTML path is taken only when :func:`_looks_like_html` finds a real - element. Both directions of that decision matter, because this method - is also wired as the inbound validator for caller-authored content: - text sent down the HTML path loses whatever the parser does not - recognise, and Markdown sent down it comes back escaped -- a value - carrying one real element is read as HTML throughout, so its - ``**bold**`` is the eight characters it spells there. - - There are therefore three outcomes, not two. Real HTML that fits - the stack is converted. Real HTML that does not is stripped to its - text instead: ``markdownify`` recurses about two frames per - nesting level, so the conversion is *attempted* and its - ``RecursionError`` answered, rather than the depth predicted and a - bound applied. The caller gets a readable body either way; - **this method degrades, it does not truncate, and it does not - raise for depth.** - - Attempting it is what makes the answer exact. The budget is not - 1000 frames, it is whatever is left of the stack when the - conversion starts, and that belongs to the caller -- an - application reading ``.content`` from inside a request handler, a - template render or a recursive walk has less of it than a script - does. No bound computed in advance can know that number, so the - previous design guessed low, 200 against a measured cliff of 494, - and flattened bodies that would have converted. It also had to - estimate the depth of the tree, and three rounds of review found - seven ways for that estimate to come in *under* the real one -- - each of which sent a document to ``markdownify`` and into the - ``RecursionError`` the bound existed to prevent. - - One consequence is the price of that exactness and worth naming: - the outcome now depends on the caller's remaining stack, so the - same body can convert from one call site and degrade from a - deeper one. Nothing is lost either way -- the degraded rendering - keeps every character of prose -- but a caller comparing two - renderings of one body should know which knob moved it. - - Self-closing void tags are written bare before conversion, which - works around a ``beautifulsoup4`` defect that silently dropped - everything after the second spelling of ``<br>`` in a body that - used both -- see :func:`_canonicalise_void_elements`. - - A document ``html.parser`` refuses outright takes the degraded - path as well, rather than the exception it used to. ``<![FOO[`` - is the reachable case: an unknown marked-section keyword, which - ``_markupbase`` raises ``AssertionError`` for and ``bs4`` - re-raises as ``ParserRejectedMarkup``. A caller who can read - their text is better off than one holding an error, and there is - nothing else to be done with such a body, so it is stripped too. - The stripped text is escaped like any other literal text - (:func:`_literal_markdown`), so it renders as itself. + def from_transport(value: object, *, plain_text_is_markdown: bool = False) -> str: + """Convert one GLPI transport value into Markdown. + + Parameters + ---------- + value : object + HTML, or plain text: a value with no HTML element in it. + plain_text_is_markdown : bool, optional + How a value that is not an HTML document is read. ``False``, the + read path, reads plain text as GLPI displays it -- literal + characters, one line per line -- so ``__init__`` comes back + escaped. ``True`` is the write models' validator: the value is + the caller's own Markdown and passes verbatim unless it starts + with a real HTML tag, so Markdown carrying an inline ``<br>`` or + ``<kbd>`` is still Markdown. + + Returns + ------- + str + The Markdown, stripped; empty for an empty value. Raises ------ GlpiContentError - The parser failed for some reason other than depth. The original - exception is attached as ``__cause__``. Nothing is expected to - reach this -- it is here so a parser fault cannot escape - ``except GlpiError`` the way a bare ``RecursionError`` used to. + The value could not be converted. A body nested too deeply for + the stack left is read as its text instead, so this is a + backstop. """ - content = str(value or "") - if not content.strip(): + content = str(value or "").strip() + if not content: return "" + if plain_text_is_markdown and not ( + content.startswith("<") and _looks_like_html(content) + ): + return content if not _looks_like_html(content): - return content.strip() + content = _plain_text_html(content) try: - markdown = html_to_markdown(_canonicalise_void_elements(content)) - except (RecursionError, ParserRejectedMarkup): - # The tree is deeper than the stack left, or the parser will - # not build it at all. Both are answered with the text. try: - return _literal_markdown(_strip_tags(content)) - except RecursionError as exc: - # Reachable only from a caller already within a few frames - # of the limit, where stripping cannot run either. Named - # rather than allowed to escape as a bare builtin, which is - # what this taxonomy exists for. - raise GlpiContentError( - "Could not convert GLPI HTML content to Markdown: the " - "caller's stack left too little room even to strip its " - "tags." - ) from exc + return html_to_markdown(content) + except RecursionError: + return html_to_markdown(_plain_text_html(_text_of(content))) + except ParserRejectedMarkup: + # html.parser gives up on a few malformed declarations; the + # body's words are still worth more than an exception. + text = unescape(_ANY_TAG.sub(" ", content)) + return html_to_markdown(_plain_text_html(" ".join(text.split()))) except Exception as exc: raise GlpiContentError( "Could not convert GLPI HTML content to Markdown " f"({type(exc).__name__}: {exc})." ) from exc - # Already stripped: the document converter strips before it settles. - return markdown @staticmethod def to_transport(value: object) -> str: - """Convert one canonical Markdown value into GLPI HTML. - - Empty Markdown stays empty, while non-empty content is rendered through - the configured Markdown extensions used by the package. - - There is no depth ceiling on this direction and no degraded path. - Inbound content is whatever GLPI happens to hold, so it has to be - survivable; outbound content is what the caller just wrote, so a - failure is worth reporting rather than papering over. ``markdown`` - recurses on nested constructs too -- measured, a list indented 495 - levels raises -- so the failure is caught and named. + """Convert one Markdown value into GLPI HTML. Raises ------ GlpiContentError - The Markdown could not be rendered. The original exception is - attached as ``__cause__``. + The Markdown could not be rendered. """ markdown = str(value or "") if not markdown.strip(): return "" try: - html = markdown_to_html( - markdown, - extensions=_MARKDOWN_EXTENSIONS, - output_format="html5", - ) + return markdown_to_html(markdown) except Exception as exc: raise GlpiContentError( "Could not render Markdown content as GLPI HTML " f"({type(exc).__name__}: {exc})." ) from exc - return str(html).strip() diff --git a/glpi_python_client/content/tests/display.py b/glpi_python_client/content/tests/display.py new file mode 100644 index 0000000..ef7f2c6 --- /dev/null +++ b/glpi_python_client/content/tests/display.py @@ -0,0 +1,220 @@ +"""What a browser displays of a body, as comparable data. + +The content tests compare displays rather than HTML: the converters +legitimately respell markup -- ``<b>`` becomes ``<strong>``, a ``<div>`` a +``<p>`` -- and none of that is visible. What is kept is what a reader sees: +the blocks, the words in each (whitespace collapsed, as HTML collapses it), +and for every character whether it is bold, italic, code or a link. + +A table's rows are its own: a table nested in a cell is read as words in that +cell, as a browser shows it inside the cell. A table with no row displays +nothing. +""" + +from __future__ import annotations + +from bs4 import ( + BeautifulSoup, + Comment, + Declaration, + Doctype, + NavigableString, + ProcessingInstruction, + Tag, +) + +_HIDDEN = {"script", "style", "title", "head", "template"} +_PARAGRAPHS = {"p", "div", "section", "article", "center", "body", "html", "main"} +_HEADINGS = {f"h{level}" for level in range(1, 7)} +_BLOCKS = _PARAGRAPHS | _HEADINGS | {"ul", "ol", "li", "blockquote", "pre", "table"} +_FORMATS = { + "b": "strong", + "strong": "strong", + "i": "em", + "em": "em", + "code": "code", + "kbd": "code", + "samp": "code", +} +_SKIPPED = (Comment, Doctype, Declaration, ProcessingInstruction) + +Format = tuple[bool, bool, bool, str | None] + +#: No formatting, and not in a link. +_PLAIN: Format = (False, False, False, None) + + +class _Words: + """The words of one paragraph, each character with its formatting.""" + + def __init__(self) -> None: + self.words: list[tuple[object, ...]] = [] + self.current: list[object] = [] + + def char(self, char: str, fmt: Format) -> None: + if char.isspace(): + self.boundary() + else: + self.current.append((char, fmt)) + + def token(self, token: object) -> None: + self.current.append(token) + + def boundary(self) -> None: + if self.current: + self.words.append(tuple(self.current)) + self.current = [] + + def line_break(self) -> None: + self.boundary() + self.words.append(("BR",)) + + def finish(self) -> tuple[object, ...]: + """The paragraph's words, less the line breaks at its edges.""" + + self.boundary() + words = list(self.words) + while words and words[0] == ("BR",): + words.pop(0) + while words and words[-1] == ("BR",): + words.pop() + return tuple(words) + + +def _with_format(fmt: Format, name: str) -> Format: + strong, em, code, href = fmt + kind = _FORMATS.get(name) + return ( + strong or kind == "strong", + em or kind == "em", + code or kind == "code", + href, + ) + + +def _image(tag: Tag, fmt: Format) -> object: + alt = " ".join(str(tag.get("alt") or "").split()) + return (("IMG", str(tag.get("src") or ""), alt), fmt) + + +def _inline(node: Tag, words: _Words, fmt: Format) -> None: + for child in node.children: + if isinstance(child, _SKIPPED): + continue + if isinstance(child, NavigableString): + for char in str(child): + words.char(char, fmt) + elif isinstance(child, Tag) and child.name not in _HIDDEN: + if child.name == "br": + words.line_break() + elif child.name == "img": + words.token(_image(child, fmt)) + elif child.name == "a" and child.get("href"): + _inline(child, words, (fmt[0], fmt[1], fmt[2], str(child.get("href")))) + else: + _inline(child, words, _with_format(fmt, child.name)) + + +def _has_block(node: Tag) -> bool: + return any( + isinstance(child, Tag) and child.name in _BLOCKS for child in node.descendants + ) + + +def _list(node: Tag, fmt: Format) -> tuple[object, ...]: + """A list: each non-empty item with the number it displays.""" + + ordered = node.name == "ol" + start = str(node.get("start") or "1") + first = int(start) if start.isdigit() else 1 + items: list[object] = [] + for position, item in enumerate(node.find_all("li", recursive=False)): + content = _display_blocks(item, fmt) + if content: # Markdown cannot spell an empty list item + items.append((first + position if ordered else None, content)) + return (node.name, tuple(items)) if items else () + + +def _table(node: Tag, fmt: Format) -> tuple[object, ...]: + rows = [] + for row in node.find_all("tr"): + if row.find_parent("table") is not node: + continue # a nested table's row shows inside its cell + cells = [] + for cell in row.find_all(["td", "th"], recursive=False): + inner = _Words() + _inline(cell, inner, fmt) + cells.append((cell.name == "th", inner.finish())) + rows.append(tuple(cells)) + return ("table", tuple(rows)) if rows else () + + +def _display_blocks(node: Tag, fmt: Format = _PLAIN) -> tuple[object, ...]: + out: list[object] = [] + words = _Words() + + def flush() -> None: + nonlocal words + paragraph = words.finish() + if paragraph: + out.append(("p", paragraph)) + words = _Words() + + for child in node.children: + if isinstance(child, _SKIPPED): + continue + if isinstance(child, NavigableString): + for char in str(child): + words.char(char, fmt) + continue + if not isinstance(child, Tag) or child.name in _HIDDEN: + continue + name = child.name + if name in _PARAGRAPHS: + flush() + out.extend(_display_blocks(child, fmt)) + elif name in _HEADINGS: + flush() + heading = _Words() + _inline(child, heading, fmt) + out.append((name, heading.finish())) + elif name in {"ul", "ol"}: + flush() + listed = _list(child, fmt) + if listed: + out.append(listed) + elif name == "blockquote": + flush() + quoted = _display_blocks(child, fmt) + if quoted: + out.append(("quote", quoted)) + elif name == "pre": + flush() + out.append(("pre", child.get_text().strip("\n"))) + elif name == "hr": + flush() + out.append(("hr",)) + elif name == "table": + flush() + table = _table(child, fmt) + if table: + out.append(table) + elif name == "br": + words.line_break() + elif name == "img": + words.token(_image(child, fmt)) + elif _has_block(child): + flush() + out.extend(_display_blocks(child, _with_format(fmt, name))) + elif name == "a" and child.get("href"): + _inline(child, words, (fmt[0], fmt[1], fmt[2], str(child.get("href")))) + else: + _inline(child, words, _with_format(fmt, name)) + flush() + return tuple(out) + + +def displayed(html: str) -> tuple[object, ...]: + """Return what a browser displays of ``html``, as comparable data.""" + + return _display_blocks(BeautifulSoup(html, "html.parser")) diff --git a/glpi_python_client/content/tests/test_conversion.py b/glpi_python_client/content/tests/test_conversion.py index ba64dfe..3a1beb5 100644 --- a/glpi_python_client/content/tests/test_conversion.py +++ b/glpi_python_client/content/tests/test_conversion.py @@ -1,151 +1,35 @@ -from __future__ import annotations +"""The converter's contract: its two directions, its errors, and what it survives. + +How faithfully realistic content round-trips is :mod:`.test_round_trip`'s +subject; this module pins the interface around it. +""" -import re -from html.parser import HTMLParser +from __future__ import annotations import pytest -from bs4 import BeautifulSoup from glpi_python_client import GlpiContentError, GlpiError from glpi_python_client.content import conversion -from glpi_python_client.content.conversion import ( - GlpiContentConverter, - _strip_tags, -) - -#: Everything that is not a letter or a digit. -_NOT_PROSE = re.compile(r"[^0-9A-Za-z]+") - -#: A Markdown link or image destination. -#: -#: The converting path renders a target that the fallback documents as -#: dropped -- "a link becomes its text without the target, an image -#: contributes nothing" -- so a URL inside ``]( )`` is not prose either, -#: and comparing it would assert a difference the module declares. -_DESTINATION = re.compile(r"\]\([^)]*\)") - -#: A link, and the target only the *converting* path renders. -#: -#: There is no depth number to ask any more -- the converter attempts the -#: walk and answers the ``RecursionError`` -- so a test that needs to know -#: which path ran has to read the output. A link is the cheapest tell: -#: ``markdownify`` writes ``[probe](u)`` and :func:`_strip_tags` writes -#: ``probe``. -PROBE_LINK = '<a href="u">probe</a>' -PROBE_TARGET = "](u)" - - -def _prose(text: str) -> str: - """Reduce a rendering to its letters and digits, in order. - - Whitespace falls differently at a markup boundary on the two paths -- - the converter joins ``a<b>c`` as ``a**c**`` where the degraded path - joins it as ``ac``, and the degraded path breaks a line at a block - edge the converter runs together -- and the converter adds - punctuation of its own: table pipes, fence backticks, list bullets, - link and image brackets. None of that is prose, and none of it is - what the fallback promises to reproduce. - """ - - return _NOT_PROSE.sub("", _DESTINATION.sub("]", text)) - - -def _displayed(markdown: str) -> str: - """Return the text a reader is shown of the converting path's Markdown. - - That path escapes literal text -- ``<!--``, ``\\_`` -- so its - Markdown is not the text it says: rendering it, as every reader does, - is what gives the text back to compare. The degraded path hands back - plain text, which :func:`_strip_tags` returns as it is. - """ - - rendered = GlpiContentConverter.to_transport(markdown) - return BeautifulSoup(rendered, "html.parser").get_text() - - -def _is_subsequence(needle: str, haystack: str) -> bool: - remaining = iter(haystack) - return all(character in remaining for character in needle) - - -def _parser_rejects(html: str) -> bool: - """Return whether ``html.parser`` gives up on this document. - - ``_markupbase`` raises ``AssertionError`` for an unknown - marked-section keyword, and which keywords count has changed across - CPython patch releases -- so whether a given document is rejected is - a question to ask the running interpreter rather than to assume. - """ +from glpi_python_client.content.conversion import GlpiContentConverter +from glpi_python_client.content.tests.display import displayed - parser = HTMLParser(convert_charrefs=False) - try: - parser.feed(html) - parser.close() - except AssertionError: - return True - return False +read = GlpiContentConverter.from_transport +render = GlpiContentConverter.to_transport -def assert_the_degraded_path_says_no_less(shallow: str) -> None: - """Assert the module's one promise about the fallback, on this parser. - - ``html.parser``'s reading of a *malformed* construct is not stable - across CPython patch releases. Measured on the same three documents, - 3.12.3, 3.12.11 and 3.12.14 disagree about an unterminated - ``<script>``, about a comment with no ``-->``, and about an end tag - carrying a quoted ``>``: each build emits a different set of events. - Writing down the literal output of one of them made this suite assert - that interpreter's quirks rather than the module's contract, and it - duly went red on a patch bump while the module itself was fine. - - So the expectation is computed from the converting path rather than - written down. Both paths read the same parser, so they move together, - and the promise was never equality anyway -- it is inclusion: **a - body must not say less because of the path it took.** - """ - - # The padding is closed *before* the construct rather than wrapped - # around it. Wrapping changes what the construct means: an - # unterminated ``<!weird`` runs to the next ``>``, which inside a - # wrapper is the ``>`` of a ``</div>``, so the same text is a bogus - # comment there and character data at end of input. The point of the - # padding is only to be deeper than the converter can walk. - # - # The probe link is how the test knows which path ran, now that there - # is no depth number to consult: only the converting path renders a - # target, so its absence is the degradation. - # Both renderings are taken from the *same* document, and the - # fallback is called directly rather than provoked with a body deep - # enough to exhaust the stack. Provoking it costs a 600-level tree - # and a walk that runs until it raises, which under coverage - # instrumentation took this suite from 69 seconds to 333; and it - # tests the routing, which one test can do once, rather than the - # property, which is what every shape here is for. - # - # The probe link carries the document onto the HTML path. Without it - # a fragment whose only tag name is not an element -- ``<scripty>`` -- - # is plain text rather than markup, and the two sides would not be - # renderings of the same thing. - document = PROBE_LINK + shallow - converted = GlpiContentConverter.from_transport(document) - degraded = _strip_tags(document) - - assert PROBE_TARGET in converted, "the body must reach the converting path" - - assert _is_subsequence(_prose(_displayed(converted)), _prose(degraded)), ( - f"the degraded path said less than the converting one\n" - f" converted: {converted!r}\n degraded: {degraded!r}" +def test_content_is_markdown_in_python_and_html_for_glpi() -> None: + assert read("<p>The printer is <strong>offline</strong>.</p>") == ( + "The printer is **offline**." ) - - -def test_content_converter_uses_markdown_in_python_and_html_for_glpi() -> None: - markdown = GlpiContentConverter.from_transport( - "<p>Hello <strong>world</strong></p>" + assert render("The printer is **offline**.") == ( + "<p>The printer is <strong>offline</strong>.</p>" ) - html = GlpiContentConverter.to_transport("Hello **world**") - assert markdown == "Hello **world**" - assert html == "<p>Hello <strong>world</strong></p>" + +@pytest.mark.parametrize("value", [None, "", " ", "\n\t"]) +def test_an_empty_value_stays_empty(value: object) -> None: + assert read(value) == "" + assert render(value) == "" @pytest.mark.parametrize( @@ -154,969 +38,88 @@ def test_content_converter_uses_markdown_in_python_and_html_for_glpi() -> None: "use the <Enter> key", "cmd </dev/null > out", "if x<y then z>0", - "temp<max and p>min", - "a </close> b", - "<!-- a bare comment -->", "generic<T> in the signature", ], ) -def test_from_transport_preserves_text_whose_tags_are_not_html(text: str) -> None: - """Angle brackets around a non-element name are text, not markup. - - ``<Enter>`` parses as an unknown tag, and an unknown tag's markup is - dropped while its (empty) body is kept -- so the word disappears from the - middle of a sentence with nothing to show it was ever there. - """ - - assert GlpiContentConverter.from_transport(text) == text - - -def test_from_transport_still_converts_real_html() -> None: - """Tightening the probe must not stop genuine HTML being normalised.""" - - html = "<p>The printer is <strong>offline</strong>.</p>" - - assert GlpiContentConverter.from_transport(html) == "The printer is **offline**." - - -def test_from_transport_leaves_caller_markdown_untouched() -> None: - """Markdown authored by a caller survives the inbound normaliser. - - ``from_transport`` is wired as a Pydantic ``BeforeValidator``, so it also - runs on outbound content. Anything that sends caller Markdown down the - HTML path escapes it, and the ticket reaches GLPI showing literal - asterisks. - """ - - markdown = "The printer is **offline** and 5 * 3 = 15." - - assert GlpiContentConverter.from_transport(markdown) == markdown +def test_angle_brackets_that_are_not_html_stay_text(text: str) -> None: + """The element name decides what is markup: ``<Enter>`` is text.""" + shown = displayed(render(read(text))) -def test_fenced_code_block_survives_the_round_trip() -> None: - """A fence stays a fence. Pasted logs are the common case for this.""" + assert shown == displayed("<p>" + text.replace("<", "<") + "</p>") - markdown = "```\nblock\n```" - assert ( - GlpiContentConverter.from_transport(GlpiContentConverter.to_transport(markdown)) - == markdown - ) - - -def test_fenced_code_block_renders_as_a_pre_block() -> None: - """Outbound, a fence becomes ``<pre><code>`` rather than inline code. - - Inline ``<code>`` is what collapsed a multi-line log into one line in the - GLPI web UI, and what a read-modify-write then wrote back as inline code. - """ - - assert GlpiContentConverter.to_transport("```\nblock\n```") == ( - "<pre><code>block\n</code></pre>" - ) - - -def test_table_survives_the_round_trip() -> None: - """A Markdown table stays a table instead of degrading to text.""" - - rendered = GlpiContentConverter.from_transport( - GlpiContentConverter.to_transport("| a | b |\n| - | - |\n| 1 | 2 |") - ) +def test_a_newline_renders_as_a_line_break() -> None: + """GLPI's editor shows each line the caller typed, as nl2br did before.""" - assert rendered == "| a | b |\n| --- | --- |\n| 1 | 2 |" + assert render("ligne un\nligne deux") == "<p>ligne un<br />\nligne deux</p>" -@pytest.mark.parametrize( - ("html", "expected"), - [ - ("<p>snake_case name</p>", "snake_case name"), - ("<p>5 * 3 = 15</p>", "5 * 3 = 15"), - ], -) -def test_incoming_text_is_not_backslash_escaped(html: str, expected: str) -> None: - r"""Underscores and asterisks in prose stay readable. - - Escaping them turns ``snake_case`` into ``snake\_case`` on every read, - and the backslash accumulates across read-modify-write cycles. - """ +def test_a_body_using_both_spellings_of_a_line_break_keeps_its_text() -> None: + """beautifulsoup4 before 4.15 dropped the text after ``<br />`` in such a body.""" - assert GlpiContentConverter.from_transport(html) == expected + markdown = read("<p>line1<br>line2</p><p>para2<br />line4</p>") - -# --------------------------------------------------------------------------- -# Nesting depth -# --------------------------------------------------------------------------- -# -# ``markdownify`` walks the parsed tree recursively, so a deeply nested body -# used to exhaust the interpreter's stack -- and because the converter was -# wired as a Pydantic ``BeforeValidator``, the ``RecursionError`` surfaced -# from inside ``model_validate``, i.e. from inside ``get_ticket``. These -# tests pin the three things that fixed it: the depth is measured without -# recursing, the ceiling is enforced, and past it the body degrades to text -# rather than raising or being cut short. + assert "line4" in markdown @pytest.mark.parametrize( "html", [ - pytest.param("<div>" * 500 + "text" + "</div>" * 500, id="balanced"), - pytest.param("<p>" * 5000 + "text", id="unclosed"), - pytest.param("<ul><li>" * 500 + "text" + "</li></ul>" * 500, id="lists"), - pytest.param("<foo>" * 5000 + "<p>text</p>", id="unknown-elements"), + pytest.param("<div>" * 3000 + "deep" + "</div>" * 3000, id="balanced"), + pytest.param("<p>" * 5000 + "deep", id="unclosed"), + pytest.param("<ul><li>" * 2000 + "deep" + "</li></ul>" * 2000, id="lists"), pytest.param( - "<table><tr><td>" * 400 + "text" + "</td></tr></table>" * 400, + "<table><tr><td>" * 1500 + "deep" + "</td></tr></table>" * 1500, id="tables", ), - pytest.param("<div>" * 100_000 + "text", id="absurd"), ], ) -def test_deep_html_degrades_instead_of_raising(html: str) -> None: - """Past the ceiling the caller gets a usable body, not an exception. - - 494 levels was enough to exhaust the default 1000-frame limit from a - shallow stack. Every shape here is past it, including the unclosed and - unknown-element ones -- both of which the parser nests just as deeply - as the balanced case. - """ - - assert GlpiContentConverter.from_transport(html) == "text" - - -def test_the_degraded_path_keeps_every_word() -> None: - """It degrades; it does not truncate. +def test_a_body_too_deep_for_the_stack_degrades_to_its_text(html: str) -> None: + """markdownify recurses per level; past the stack the words are still read.""" - A body too deep to convert is still the only copy of what someone - wrote, so the fallback's contract is that all of the text comes back. - """ - - lines = [f"line {index}" for index in range(400)] - html = "".join(f"<div><p>{line}</p>" for line in lines) + "</div>" * 400 - - stripped = GlpiContentConverter.from_transport(html) - - assert all(line in stripped for line in lines) - assert "<" not in stripped - - -def test_the_degraded_path_resolves_entities_and_block_boundaries() -> None: - """Blocks become line breaks, inline tags vanish, references resolve. - - Dropping every tag outright would run ``<p>a</p><p>b</p>`` together as - ``ab``; separating at every tag would break ``<b>off</b>line`` into two - words. Only the block boundary gets a separator. - """ - - html = "<div>" * 600 + "<b>off</b>line & <p>next</p>" + "</div>" * 600 - - assert GlpiContentConverter.from_transport(html) == "offline &\nnext" + assert "deep" in read(html) @pytest.mark.parametrize( - ("construct", "converted_keeps", "degraded_keeps"), - [ - pytest.param("<!-- SECRET -->", False, False, id="resolved-comment"), - pytest.param("<!DOCTYPE SECRET>", False, False, id="doctype"), - pytest.param("<!SECRET>", False, False, id="bogus-declaration"), - pytest.param("<script>SECRET</script>", False, True, id="script-body"), - pytest.param("<style>SECRET</style>", False, True, id="style-body"), - pytest.param("<![CDATA[SECRET]]>", True, True, id="marked-section"), - pytest.param("<![CDATA[SECRET>", True, True, id="unterminated-marked-section"), - pytest.param("<?php SECRET ?>", True, True, id="processing-instruction"), - ], + "html", ["<p>a</p><![FOO[x]]><p>b</p>", "<p>a</p><![ x<p>b</p>"] ) -def test_the_degraded_path_keeps_at_least_what_the_converter_keeps( - construct: str, converted_keeps: bool, degraded_keeps: bool -) -> None: - """Parity, construct by construct, and not one of these was a guess. - - Each expectation here was read off the converting path rather than - reasoned about, and two came back the opposite way round from the - obvious answer -- a ``CDATA`` body is *kept*, and so is the inside of - any construct the parser could not resolve. Each of those was a - silent deletion in the degraded path until it was measured. - - A ``<script>`` or ``<style>`` body is where the two paths part, in the - one direction allowed. The converting path used to keep it -- - ``markdownify``'s ``strip=`` removed the element's markup and still - walked its children -- and now drops it, as a browser does; the - degraded path still keeps it. - - The bar is that a body must not say less because of the path it took, - so a divergence the other way would be a bug even if the text it lost - were JavaScript. - """ - - shallow = f"<p>a</p>{construct}<p>b</p>" - deep = "<div>" * 600 + shallow + "</div>" * 600 - - converted = GlpiContentConverter.from_transport(shallow) - degraded = GlpiContentConverter.from_transport(deep) - - # What was measured is asserted only where the running parser still - # agrees with the measurement -- the reading of a malformed construct - # moves between CPython patch releases, and it is the superset below, - # not the snapshot, that this module promises. - if ("SECRET" in converted) is converted_keeps: - assert ("SECRET" in degraded) is degraded_keeps - assert not ("SECRET" in converted and "SECRET" not in degraded) - - -def test_an_unterminated_raw_text_element_reads_the_same_on_both_paths() -> None: - """Whether an unclosed ``<script>`` body survives is the parser's call. +def test_markup_the_parser_rejects_degrades_to_its_text(html: str) -> None: + assert read(html) == "a b" - It made this test its own snapshot: 3.12.3 discards the body on - ``close()`` and 3.12.14 flushes it as character data, so the literal - that was correct on one was wrong on the other. What has to hold on - either is that the two paths agree. - """ - assert_the_degraded_path_says_no_less("<p>keep</p><script>SECRET") - assert "keep" in GlpiContentConverter.from_transport("<p>keep</p><script>SECRET") - - -@pytest.mark.parametrize("depth", [1, 100, 200, 250]) -def test_a_document_the_stack_can_hold_is_converted_in_full(depth: int) -> None: - """Everything that fits must convert, and structure has to survive. - - This is what attempting the conversion bought. The previous design - predicted the depth and degraded past a fixed 200, which flattened - every body between 200 and the real cliff of about 494 -- ordinary - quoted mail threads among them -- to text, with no error to notice - and no way for a caller to ask for better. The 250 case is one that - used to come back as prose. It is also as deep as this goes, so that - it holds on every interpreter CI runs: CPython 3.10 spends about three - frames per level where 3.12 spends two, its cliff is about 328 levels - from a shallow stack (measured by easyvista_python_client's port, - 2026-09-30), and a 300-level case left 14 levels of margin under - pytest there. - """ - - html = "<div>" * depth + "<strong>offline</strong>" + "</div>" * depth - - assert GlpiContentConverter.from_transport(html) == "**offline**" - - -def test_void_elements_do_not_spend_the_depth_budget() -> None: - """5000 ``<br>`` is one level, so this must take the converting path. - - The surviving ``**`` proves it: the degraded path strips markup, so - emphasis would be gone if the void tags had been counted. - """ - - html = "<p>" + "<br>" * 5000 + "<strong>offline</strong></p>" - - assert "**offline**" in GlpiContentConverter.from_transport(html) - - -# --------------------------------------------------------------------------- -# Failure taxonomy -# --------------------------------------------------------------------------- - - -def test_a_parser_fault_surfaces_as_a_glpi_error( +def test_an_inbound_fault_surfaces_as_a_glpi_error( monkeypatch: pytest.MonkeyPatch, ) -> None: - """No parser fault escapes ``except GlpiError``. - - Nothing ordinary reaches this, so the fault is injected. It matters - anyway: a caller who wrote ``except GlpiError`` around ``get_ticket`` - would otherwise watch a bare parser exception sail straight through - it. + def failing(html: str) -> str: + raise ValueError("converter fault") - A ``RecursionError`` is deliberately *not* the fault used here. It is - no longer a failure at all -- it is how the converter learns that the - document does not fit, and it is answered with the body's text; see - the test below. - """ - - def _boom(*args: object, **kwargs: object) -> str: - raise ValueError("the parser fell over") - - monkeypatch.setattr(conversion, "html_to_markdown", _boom) + monkeypatch.setattr(conversion, "html_to_markdown", failing) with pytest.raises(GlpiContentError) as caught: - GlpiContentConverter.from_transport("<p>offline</p>") + read("<p>x</p>") assert isinstance(caught.value, GlpiError) assert isinstance(caught.value.__cause__, ValueError) -def test_a_recursion_error_degrades_rather_than_raising( +def test_an_outbound_fault_surfaces_as_a_glpi_error( monkeypatch: pytest.MonkeyPatch, ) -> None: - """Running out of stack is answered, not reported. - - Injected rather than provoked, because the depth needed to provoke it - depends on the stack the test runner has already spent -- which is the - very reason the depth is no longer predicted. What is pinned is the - contract: the caller gets their words, not an exception. - """ + def failing(markdown: str) -> str: + raise RuntimeError("renderer fault") - def _boom(*args: object, **kwargs: object) -> str: - raise RecursionError("maximum recursion depth exceeded") - - monkeypatch.setattr(conversion, "html_to_markdown", _boom) - - assert ( - GlpiContentConverter.from_transport("<p>Le serveur ne repond plus.</p>") - == "Le serveur ne repond plus." - ) - - -def test_an_outbound_render_fault_surfaces_as_a_glpi_error( - monkeypatch: pytest.MonkeyPatch, -) -> None: - """The outbound direction is wrapped too; ``markdown`` recurses as well.""" - - def _boom(*args: object, **kwargs: object) -> str: - raise RecursionError("maximum recursion depth exceeded") - - monkeypatch.setattr(conversion, "markdown_to_html", _boom) + monkeypatch.setattr(conversion, "markdown_to_html", failing) with pytest.raises(GlpiContentError) as caught: - GlpiContentConverter.to_transport("offline") - - assert isinstance(caught.value, GlpiError) - assert isinstance(caught.value.__cause__, RecursionError) - - -def test_deeply_nested_markdown_does_not_raise_a_bare_recursion_error() -> None: - """The real outbound cliff, unmocked. - - ``markdown`` breaks between 495 and 500 levels of list indentation - (measured). Unlike the inbound direction this is not degraded -- - outbound content is what the caller just wrote, so a failure is worth - reporting -- but it has to be reported as a library error. - """ - - markdown = "\n".join(" " * level + "- x" for level in range(500)) - - with pytest.raises(GlpiContentError): - GlpiContentConverter.to_transport(markdown) - - -# --------------------------------------------------------------------------- -# The scan against the tree the parser really builds -# --------------------------------------------------------------------------- -# -# The three cases below are the ones a plain open/close counter gets wrong, -# and the first is not a corner case: a stray ``</p>`` or ``</span>`` is -# what a Word or Outlook paste leaves in a GLPI body. Each was measured -# under-counting -- the one direction that turns into a crash -- before the -# scan learned to pop by name. - - -def _parser_depth(html: str) -> int: - """Return the deepest element ``html.parser`` actually builds. - - The ground truth the scan is checked against, walked iteratively so - that measuring a pathological document does not hit the very limit - under test. - """ - - soup = BeautifulSoup(html, "html.parser") - deepest = 0 - pending = [(child, 1) for child in soup.children if getattr(child, "name", None)] - while pending: - node, depth = pending.pop() - deepest = max(deepest, depth) - pending.extend( - (child, depth + 1) - for child in node.children - if getattr(child, "name", None) - ) - return deepest - - -@pytest.mark.parametrize( - "html", - [ - pytest.param("<div></p>" * 600 + "kept", id="stray-close-p"), - pytest.param("<div></span>" * 800 + "kept", id="stray-close-span"), - pytest.param("<li></tr>" * 700 + "kept", id="stray-close-tr"), - ], -) -def test_a_stray_closing_tag_does_not_hide_real_nesting(html: str) -> None: - """The regression test for the under-count that reached ``markdownify``. - - A closing tag with no matching open element pops nothing in ``bs4``, so - these documents nest as deeply as their opening tags say. Counting the - close as a level down measured them at 1, they went to the recursive - converter, and it raised. - """ - - assert GlpiContentConverter.from_transport(html) == "kept" - - -def test_the_degraded_path_keeps_a_cdata_body() -> None: - """A ``CDATA`` section's body is text, and text is what survives. - - The declaration pattern that strips ``<!DOCTYPE ...>`` reaches the - first ``>``, and a ``CDATA`` section has none until its end, so it used - to take the body with it -- a silent deletion in the one path whose - whole promise is that nothing is deleted. - """ - - html = "<div>" * 600 + "<p>a<![CDATA[secret words]]>b</p>" + "</div>" * 600 - - assert GlpiContentConverter.from_transport(html) == "asecret wordsb" - - -def test_the_degraded_path_keeps_the_text_of_a_broken_comment() -> None: - """Parity with the normal path, even where the parser gave up. - - ``html.parser`` cannot resolve a comment with no ``-->``, so it hands - the region back as character data -- meaning ``markdownify`` would have - kept it. The degraded path is only trustworthy if which path a body - took never changes what it says, so it keeps it too. - """ - - assert_the_degraded_path_says_no_less("<p>keep</p><!--oops but keep this") - - -def test_the_degraded_path_drops_a_comment_the_parser_understood() -> None: - """A resolved comment is text on neither path, so it goes. - - The mirror of the test above, and the reason the two cannot share one - rule: telling them apart is the whole job of the ``-->``. - """ - - html = "<div>" * 600 + "<p>keep</p><!-- drop this -->" + "</div>" * 600 - - assert GlpiContentConverter.from_transport(html) == "keep" - - -def test_no_name_in_the_void_set_actually_nests() -> None: - """The one direction of the void set that would be a crash. - - A name listed as void that the parser really nests hides real depth, - and the document then reaches the recursive converter. The set is a - hand-copy of a private ``bs4`` table, so the invariant is asserted - against a real parse rather than against that table: for every name - claimed void, 300 of them must build one level, not 300. - """ - - understated = [ - name - for name in sorted(conversion._VOID_ELEMENTS) - if _parser_depth(f"<{name}>" * 300 + "x") > 1 - ] - - assert understated == [] - - -@pytest.mark.parametrize( - "html", - [ - pytest.param('<div title="</div>">x</div>', id="close-tag-in-attribute"), - pytest.param('<div title="<div>">x</div>', id="open-tag-in-attribute"), - pytest.param('<div title="a>b">x</div>', id="gt-in-attribute"), - pytest.param("<div data-x='a>b'><p>x</p></div>", id="single-quoted"), - pytest.param('<p title=">">one</p><p>two</p>', id="attribute-is-just-gt"), - pytest.param('<div title="</div>">' * 5 + "x", id="nested-and-quoted"), - ], -) -def test_a_tag_inside_a_quoted_attribute_is_not_read_as_markup(html: str) -> None: - """An attribute value may legally contain ``<`` and ``>``. - - Reading a quoted ``</div>`` as a real close tag lost the text after - it, and used to under-count the nesting without bound as well, back - when the nesting was predicted. What has to hold either way is that - the shape costs the body nothing: it converts, and if it is too deep - to convert it still says the same thing. - """ - - assert_the_degraded_path_says_no_less(html) - - -def test_a_quoted_close_tag_at_depth_still_degrades() -> None: - """The same shape scaled past the ceiling: degrades, does not raise.""" - - html = '<div title="</div>">' * 600 + "keep" - - assert GlpiContentConverter.from_transport(html) == "keep" - - -@pytest.mark.parametrize( - ("raw", "expected"), - [ - # A URL in a ticket body is exactly where a semicolon-less - # reference that is a prefix of a longer word shows up. - ("http://x/?a=1©right=2", "http://x/?a=1©right=2"), - ("http://x/?a=1¬anentity=2", "http://x/?a=1¬anentity=2"), - # Terminated references still resolve, named and numeric. - ("a & b < c", "a & b < c"), - ("AB", "AB"), - # ... and so does a semicolon-less reference whose whole name is - # known, because that is what the parser does. - ("© 2026", "\u00a9 2026"), - ("a&b", "a&b"), - ], -) -def test_the_degraded_path_resolves_references_like_the_parser( - raw: str, expected: str -) -> None: - """``html.unescape`` alone corrupts URLs; the parser's rule does not. - - ``unescape`` implements HTML5's longest-known-*prefix* rule, so - ``©right=2`` comes back as ``(c)right=2`` -- a query parameter - silently rewritten. The parser behind the converting path resolves a - semicolon-less reference only when the entire name is known, so it - leaves that URL alone, and the degraded path has to agree or a body - changes meaning according to how deeply it nests. - """ - - html = "<div>" * 600 + f"<p>{raw}</p>" + "</div>" * 600 - - assert GlpiContentConverter.from_transport(html) == expected - - -@pytest.mark.parametrize( - "prefix", - [ - pytest.param("<style=>", id="malformed-style-name"), - pytest.param("<script=>", id="malformed-script-name"), - pytest.param("<script/>", id="self-closed-script"), - pytest.param("<style />", id="self-closed-style"), - pytest.param("<p title=don't>", id="apostrophe-in-bare-value"), - pytest.param("<p alt=P<0.05>", id="lt-in-bare-value"), - ], -) -def test_a_malformed_tag_does_not_swallow_the_body_after_it(prefix: str) -> None: - """One malformed tag must not take the rest of the document with it. - - Read as raw text, ``"<style=>"`` and ``"<script/>"`` swallowed - everything after them. That showed up first as an unbounded depth - under-count -- ``"<style=>" + "<div>" * 600`` measured **1** level - against a real 601 and reached ``markdownify`` -- but the deletion was - always the real damage, and it is what this pins now that the depth is - no longer predicted. - """ - - html = prefix + "<div>" * 600 + "the printer is offline" - - assert GlpiContentConverter.from_transport(html) == "the printer is offline" - - -def test_the_degraded_path_does_not_emit_a_tag_it_could_not_read() -> None: - """A tag the scan mis-read used to be printed at the reader. - - Measured: the body below degraded to ``"<p title=don't>Le serveur ne - repond plus."`` -- the opening tag verbatim in text a person reads, - from a path whose whole promise is text. An apostrophe is ordinary in - French, so this needs no malice to reach a ticket. - """ - - html = "<div>" * 600 + "<p title=don't>Le serveur ne repond plus.</p>" - - degraded = GlpiContentConverter.from_transport(html) - - assert degraded == "Le serveur ne repond plus." - assert "<" not in degraded - - -def test_an_unterminated_declaration_at_end_of_input_is_kept() -> None: - """``close()`` flushes an incomplete declaration as text, so this does. - - A declaration is text on neither path only when it is *closed*: a - ``"<!weird"`` that never completes is flushed as character data when - the parser closes, and the converting path prints it. Dropping it lost - the tail of the body. - """ - - assert_the_degraded_path_says_no_less("<p>keep this</p><!weird") - assert "keep this" in GlpiContentConverter.from_transport("<p>keep this</p><!weird") - - -@pytest.mark.parametrize( - ("html", "expected"), - [ - pytest.param( - "<p>one<br>two<br />three</p>", - "one \ntwo \nthree", - id="br-both-spellings", - ), - pytest.param( - "<p>Bonjour,<br>Le serveur ne repond plus.<br />Merci de regarder.</p>", - "Bonjour, \nLe serveur ne repond plus. \nMerci de regarder.", - id="realistic-body", - ), - pytest.param( - "<p>line1<br>line2</p><p>para2<br />line4</p>", - "line1 \nline2\n\npara2 \nline4", - id="across-paragraphs", - ), - pytest.param( - "<p>a<br>b<br />c<br>d<br />e</p>", - "a \nb \nc \nd \ne", - id="alternating", - ), - pytest.param( - "<p>one<img>two<img />three</p>", - "one![]()two![]()three", - id="img", - ), - pytest.param( - "<p>one<hr>two<hr />three</p>", - "one\n\n---\n\ntwo\n\n---\n\nthree", - id="hr", - ), - ], -) -def test_a_body_using_both_spellings_of_a_void_tag_keeps_its_text( - html: str, expected: str -) -> None: - """Text after the second spelling of ``<br>`` used to be dropped. - - A ``beautifulsoup4`` defect, silent when it fires and reachable from - ordinary editor output: a bare ``<br>`` leaves its name in - ``already_closed_empty_element`` for a ``</br>`` that never comes, and - the next ``<br />`` closes itself against that stale entry and stays - open. Every later sibling becomes its child, and ``convert_br`` - discards an element's children. - - Note the paragraph case: the two spellings need not be near each - other, because a name once recorded poisons the rest of the document. - ``<img>`` and ``<hr>`` are the other two converters that drop - children, so they lose text the same way. - """ - - assert GlpiContentConverter.from_transport(html) == expected - - -@pytest.mark.parametrize( - "html", - [ - pytest.param("<div/>x", id="self-closed-non-void"), - pytest.param("<custom />x", id="self-closed-unknown"), - pytest.param("<p>a<br>b</p>", id="already-bare"), - pytest.param("<p>2 /> 3</p>", id="slash-gt-in-text"), - pytest.param('<div title="<br />">x</div>', id="in-an-attribute"), - pytest.param("<script>var s = '<br />';</script>x", id="in-a-script-body"), - pytest.param("<!-- <br /> -->x", id="in-a-comment"), - pytest.param("<p>a<br / >b</p>", id="slash-not-abutting-gt"), - ], -) -def test_the_void_rewrite_leaves_everything_else_alone(html: str) -> None: - """The rewrite is confined to void tags in real tag position. - - ``<div/>`` is left as it is -- rewriting it would change what the - document means, and it cannot be affected anyway, since only a void - name is ever recorded as already closed. The last case is the one - worth pinning: ``<br / >`` reaches the parser as an ordinary start - tag, because its ``/`` does not abut the ``>``, so it never takes the - path that loses text and needs no rewriting. - """ - - assert conversion._canonicalise_void_elements(html) == html - - -def test_the_void_rewrite_changes_nothing_for_one_spelling_alone() -> None: - """A body that picks a spelling and keeps it converts exactly as before. - - The rewrite exists to remove an asymmetry between two spellings of the - same node, so it must be invisible to every body that does not mix - them. Measured over 4000 fuzzed documents of each spelling: not one - output moved. - """ - - bare = "<p>a<br>b<img><hr>c</p>" - slashed = "<p>a<br />b<img /><hr />c</p>" - - assert GlpiContentConverter.from_transport(bare) == "a \nb![]()\n\n---\n\nc" - assert GlpiContentConverter.from_transport(slashed) == "a \nb![]()\n\n---\n\nc" - - -def test_both_paths_agree_on_a_body_using_both_spellings() -> None: - """The degraded path already kept this text; now the converting one does. - - This body was the one place where the fallback said *more* than the - conversion it stands in for, which is the wrong way round for a - fallback and was how the defect was noticed at all. - """ - - shallow = "<p>one<br>two<br />three</p>" - deep = "<div>" * 600 + shallow + "</div>" * 600 - - converted = GlpiContentConverter.from_transport(shallow) - degraded = GlpiContentConverter.from_transport(deep) - - for word in ("one", "two", "three"): - assert word in converted - assert word in degraded - - -@pytest.mark.parametrize( - "fragment", - [ - pytest.param('</x a="><div>">', id="end-tag-with-a-quoted-attribute"), - pytest.param('<div ="<p>', id="name-less-equals-quote"), - pytest.param('<p title="><span>">', id="quoted-gt-then-tag"), - ], -) -def test_a_misread_tag_end_does_not_swallow_the_body_after_it(fragment: str) -> None: - """Scaled past the cliff: degrades quietly, and keeps its words. - - Reading an end tag with start-tag rules consumed everything up to the - next quote, which deleted prose at any depth and under-counted the - nesting 1:1 with the repetition back when the nesting was predicted. - """ - - html = fragment * 600 + "the printer is offline" - - assert "the printer is offline" in GlpiContentConverter.from_transport(html) - - -def test_a_misread_tag_end_does_not_delete_prose() -> None: - """The other half of the same defect, and it needs no depth at all. - - Reading an end tag with attribute rules consumed everything between - the opening quote and its partner, so a degraded body said less than - the converting one -- the divergence the parity test forbids. - """ - - body = '<p>Bonjour</p title="> Le serveur ne repond plus. SECRET ">fin' - - assert_the_degraded_path_says_no_less(body) - assert "Bonjour" in GlpiContentConverter.from_transport(body) - - -def test_a_document_with_no_closing_bracket_is_answered_without_scanning() -> None: - """No ``>`` means no element, and saying so keeps a bad shape cheap. - - ``html.parser`` cannot finish a tag that never closes, so ``close()`` - flushes it one character at a time and rescans the tail at each step: - measured, 32 KB of ``'<div a="'`` costs it 13 seconds. Both readers - answer that shape directly instead. The text still survives, because - the parser flushes an unfinished tag as data when it closes. - """ - - html = '<div a="' * 4000 - - assert GlpiContentConverter.from_transport(html).startswith('<div a="') - assert "the printer" in GlpiContentConverter.from_transport(html + "the printer") - assert _strip_tags(html).startswith('<div a="') - - -def test_a_document_the_parser_rejects_degrades_instead_of_raising() -> None: - """An unknown marked-section keyword stops the parser, not the body. - - ``_markupbase.parse_marked_section`` raises ``AssertionError`` for a - keyword it does not know, and ``bs4`` catches that same - ``AssertionError`` and re-raises it as ``ParserRejectedMarkup``. So a - document carrying ``<![FOO[`` is one the converting path cannot - convert either, and the old behaviour was to let it try and fail: - the caller got a :class:`GlpiContentError` and none of their text. - - Reporting a depth past the ceiling instead sends it down the degraded - path, where it yields its words. Degrading beats raising when the - alternative is a body nobody can read. - """ - - html = "<p>Le serveur ne repond plus. SECRET</p><![FOO[x]]>" - - assert "SECRET" in GlpiContentConverter.from_transport(html) - if _parser_rejects(html): - # The converting path cannot run at all, so the answer is the text, - # spelled as the Markdown that renders as it. - assert GlpiContentConverter.from_transport(html) == ( - conversion._literal_markdown(_strip_tags(html)) - ) - - -def test_the_text_after_a_construct_the_parser_rejects_is_still_kept() -> None: - """The give-up point is not the end of the body. - - The scan stops where the parser stopped, so everything past that - construct would go missing unless it is handed back explicitly -- - and a body is far more likely to carry the marked section in the - middle than at the end. - """ - - html = "<p>avant</p><![FOO[x]]><p>apres SECRET</p>" - - assert "avant" in _strip_tags(html) - assert "SECRET" in _strip_tags(html) - assert "SECRET" in GlpiContentConverter.from_transport(html) - assert "avant" in GlpiContentConverter.from_transport(html) - - -def test_stripping_a_document_with_no_closing_bracket_keeps_all_of_it() -> None: - """The degraded path needs the same guard the depth scan needs. - - ``html.parser`` cannot complete a tag that never closes, so - ``close()`` flushes it one character at a time and rescans the tail - at each step: measured, 32 KB of ``'<div a="'`` costs it 13 seconds. - The whole document is that unfinished tag's text, which is the answer - the guard returns directly. - """ - - html = '<div a="' * 4000 + "le serveur ne repond plus" - - assert _strip_tags(html).endswith("le serveur ne repond plus") - assert _strip_tags(html).startswith('<div a="') - - -@pytest.mark.parametrize( - "html", - [ - # The repetition counts here are deliberately modest. They were - # large when this corpus guarded a depth *prediction*, where the - # error grew with the repetition; the shapes are what matter now, - # and each document has to be one the converting path can still - # walk on every supported interpreter for the comparison to mean - # anything. - pytest.param("<div></p>" * 60 + "kept", id="stray-close-p"), - pytest.param("<div></span>" * 60 + "kept", id="stray-close-span"), - pytest.param("<p></b>" * 60 + "kept", id="stray-close-b"), - pytest.param("<b><i>x</b></i>", id="interleaved"), - pytest.param("<div>" * 100 + "<br>" + "</div>" * 100, id="void-leaf"), - pytest.param("<div>" * 100 + "<img/>" + "</div>" * 100, id="self-closed-leaf"), - pytest.param("<br>" * 5000, id="void-only"), - pytest.param("<div><b>x</b></br></div>", id="close-of-a-void"), - pytest.param("<p>The printer is <strong>offline</strong>.</p>", id="realistic"), - pytest.param("<div><!-- <div><div> --><p>x</p></div>", id="tags-in-a-comment"), - pytest.param( - "<div><script>var s='<div><div>'</script>x</div>", id="tags-in-js" - ), - pytest.param("<table><tr><td>" * 30 + "x", id="tables"), - pytest.param("<blockquote>" * 100 + "x", id="blockquotes"), - pytest.param("<p>a</p>" * 100, id="siblings"), - ], -) -def test_both_paths_agree_about_what_is_markup(html: str) -> None: - """The two renderings of one body must not disagree about its markup. - - This corpus was built against a flat scan that predicted the nesting - depth, and it caught the scan reading markup differently from the - parser -- a stray close popping an element the parser keeps, a void - element counted as a parent, a tag inside a comment or a script body - counted at all. The prediction is gone; the corpus is not, because - the same disagreements would now show up as the fallback deleting or - inventing text relative to the converting path. - """ - - assert_the_degraded_path_says_no_less(html) - - -@pytest.mark.parametrize( - "html", - [ - # A comment with no ``-->`` is a *bogus comment*: the parser gives up - # at the first ``>``, so the ``</custom>`` inside it is text and the - # ``<br>`` lands inside ``<custom>``. Read that ``</custom>`` as a - # real close and the count comes back one level short -- which is - # how a document that needed degrading reached the converter. - pytest.param("<custom><!--oops</custom><br>", id="bogus-comment-eats-a-close"), - # ... and recovery ends at that ``>``. It does not swallow the rest - # of the document, so these really are two levels. - pytest.param("<!--oops><div><div>", id="bogus-comment-ends-at-its-close"), - # A processing instruction ends at the first ``>`` too, and here - # that lands inside what looks like a comment -- so the second - # ``<div>`` is a real element. Stripping comments globally before - # scanning gets this wrong in both directions at once. - pytest.param("<?php x<!-- <div><div> -->", id="pi-overlapping-a-comment"), - pytest.param("<div><!-- <div><div> --></div>", id="terminated-comment"), - pytest.param("<!DOCTYPE html><div><p>x</p></div>", id="doctype"), - pytest.param("<div><![CDATA[a<div>b]]><p>x</p></div>", id="marked-section"), - pytest.param("<div><script>a<div><div></script><p>x</p></div>", id="raw-text"), - pytest.param("<div><script>a<div>", id="unclosed-raw-text"), - ], -) -def test_every_markup_construct_is_read_the_way_the_parser_reads_it(html: str) -> None: - """Construct by construct, and the order they are tried in matters. - - The parser reads left to right and these constructs overlap: in - ``<?php x<!-- <div><div> -->`` the processing instruction ends at the - first ``>``, which lands inside what looks like a comment, so the - ``<div>`` after it is a real element. Handling any of them out of - order gets that document wrong in both directions at once. - """ - - assert_the_degraded_path_says_no_less(html) - - -@pytest.mark.parametrize( - "html", - [ - pytest.param("<p title=don't>x</p>", id="apostrophe-in-bare-value"), - pytest.param('<p title=say"hi>x</p>', id="quote-in-bare-value"), - pytest.param("<p alt=P<0.05>x</p>", id="lt-in-bare-value"), - pytest.param("<p title=Etape 1>x</p>", id="space-in-bare-value"), - pytest.param("<style=>x", id="malformed-style-name"), - pytest.param("<script=>x", id="malformed-script-name"), - pytest.param("<div=x>y", id="malformed-name"), - pytest.param('<div"a">y', id="quote-in-name"), - pytest.param("<scripty>x</scripty>", id="raw-name-is-a-prefix"), - pytest.param("<script/>x", id="self-closed-script"), - pytest.param("<style />x", id="self-closed-style"), - pytest.param('<li y=">mot<script data-x="</div>">tail', id="lt-after-value"), - pytest.param("<div a=1 <p>text", id="tag-inside-a-tag"), - pytest.param("<div class=a<b>text", id="lt-in-unquoted-value"), - ], -) -def test_a_malformed_tag_is_read_the_way_the_parser_reads_it(html: str) -> None: - """Malformed markup is where every rule taken from the spec was wrong. - - Each shape here was read one way by the HTML5 grammar and another by - ``tagfind_tolerant`` and ``locatestarttagend_tolerant``, which are - what ``markdownify`` actually builds its tree with. The module reads - the parser's events now, so the corpus is a guard against a future - change reintroducing a rule from the wrong place. - """ - - assert_the_degraded_path_says_no_less(html) - - -@pytest.mark.parametrize( - ("html", "reason"), - [ - pytest.param( - '</x a="><div>">', - "an end tag skips nothing, so the <div> after it is real", - id="end-tag-with-a-quoted-attribute", - ), - pytest.param( - '<div ="<p><p>">', - "a name-less = starts an attribute NAME, not a quoted value", - id="name-less-equals-quote", - ), - pytest.param( - '<div a ="x>y">', - "whitespace before = still leaves a quoted value", - id="space-before-equals", - ), - pytest.param( - "</ p>x", - "the strict end-tag pattern allows space after </", - id="space-in-end-tag", - ), - pytest.param("</p/>x", "a trailing slash on an end tag", id="slash-in-end-tag"), - pytest.param( - '<li y=">mot<script data-x="</div>">tail', - "< inside a tag", - id="lt-after-a-value", - ), - ], -) -def test_a_tag_end_is_read_the_way_the_parser_reads_it(html: str, reason: str) -> None: - """``parse_endtag`` falls back to ``rawdata.find(">")`` and skips nothing. - - CPython's own comment concedes the consequence -- "this is not 100% - correct, since we might have things like ``</tag attr=">">``" -- so - an end tag read with attribute rules consumes prose the parser keeps. - """ + render("**x**") - assert_the_degraded_path_says_no_less(html) + assert isinstance(caught.value.__cause__, RuntimeError) -def test_the_fallback_does_not_itself_recurse() -> None: - """Stripping 100k levels must not need 100k frames. +def test_deeply_nested_markdown_renders() -> None: + """cmark-gfm does not recurse in Python, so depth costs nothing outbound.""" - The fallback exists because the converting path ran out of stack, so - it cannot want a stack of its own. ``html.parser`` is an iterative - scanner and this walks its events into a list, which is what makes it - an answer for input of any depth rather than a second thing to guard. - """ + markdown = "".join(" " * (2 * level) + "- x\n" for level in range(600)) - assert _strip_tags("<div>" * 100_000 + "le serveur") == "le serveur" + assert render(markdown).count("<li>") == 600 diff --git a/glpi_python_client/content/tests/test_literal_text.py b/glpi_python_client/content/tests/test_literal_text.py deleted file mode 100644 index ddf9fee..0000000 --- a/glpi_python_client/content/tests/test_literal_text.py +++ /dev/null @@ -1,1548 +0,0 @@ -"""Literal text reads back as the text it was, not as Markdown syntax. - -``from_transport`` turns GLPI's HTML into Markdown, and ``to_transport`` -- or -any python-markdown with the same four extensions, which is what a peer -system renders the Markdown with -- turns it back. Text in the HTML is -literal: a ``__init__`` a user typed is eight characters, not bold ``init``. -Markdown has one spelling for both, so the converter has to spell the literal -one so that python-markdown cannot mistake it, and it has to do it in the -HTML-to-Markdown step: that is the last point where literal text and markup -can still be told apart. - -The property every test here comes back to, checked with an HTML parser -rather than by eye: - -* ``to_transport(from_transport(html))`` **displays what ``html`` displays** - -- the same words, the same formatting on each character, the same block - structure -- for every shape Markdown can express; and -* the Markdown is **a fixed point**: reading back what it renders gives the - same Markdown again. - -It is asserted over hand-picked regressions, over realistic ticket bodies, -and over a seeded fuzzer that puts every character python-markdown treats as -syntax at every position -- line start, word boundary, inside a word, in -cells, list items, headings and link text. - -Escaping is also **minimal**: a character is escaped only where -python-markdown would otherwise read it as syntax, so ordinary prose -- a -file name with underscores, a Windows path, a mid-sentence ``#``, a hyphen, -``R&D`` -- comes back exactly as it did before escaping existed. -""" - -from __future__ import annotations - -import random -import time -from html import escape - -import pytest -from bs4 import ( - BeautifulSoup, - Comment, - Declaration, - Doctype, - NavigableString, - ProcessingInstruction, - Tag, -) -from markdown import Markdown - -from glpi_python_client.content import conversion -from glpi_python_client.content.conversion import GlpiContentConverter - -read = GlpiContentConverter.from_transport -render = GlpiContentConverter.to_transport - -# --------------------------------------------------------------------------- -# What a browser displays of a body -# --------------------------------------------------------------------------- -# -# The comparison has to be on the display rather than on the HTML: the two -# converters legitimately respell markup -- ``<b>`` becomes ``<strong>``, a -# ``<div>`` becomes a ``<p>``, a list item's lone paragraph loses its -# ``<p>`` -- and none of that is visible. What is kept is what a reader sees: -# the blocks, the words in each (whitespace collapsed, as HTML collapses it), -# and for every character whether it is bold, italic, code or a link. - -_HIDDEN = {"script", "style", "title", "head", "template"} -_PARAGRAPHS = {"p", "div", "section", "article", "center", "body", "html", "main"} -_HEADINGS = {f"h{level}" for level in range(1, 7)} -_BLOCKS = _PARAGRAPHS | _HEADINGS | {"ul", "ol", "li", "blockquote", "pre", "table"} -_FORMATS = { - "b": "strong", - "strong": "strong", - "i": "em", - "em": "em", - "code": "code", - "kbd": "code", - "samp": "code", -} -_SKIPPED = (Comment, Doctype, Declaration, ProcessingInstruction) - -#: No formatting, and not in a link. -_PLAIN: tuple[bool, bool, bool, str | None] = (False, False, False, None) - - -class _Words: - """The words of one paragraph, each character with its formatting.""" - - def __init__(self) -> None: - self.words: list[tuple[object, ...]] = [] - self.current: list[object] = [] - - def char(self, char: str, fmt: tuple[bool, bool, bool, str | None]) -> None: - if char.isspace(): - self.boundary() - else: - self.current.append((char, fmt)) - - def token(self, token: object) -> None: - self.current.append(token) - - def boundary(self) -> None: - if self.current: - self.words.append(tuple(self.current)) - self.current = [] - - def line_break(self) -> None: - self.boundary() - self.words.append(("BR",)) - - def finish(self) -> tuple[object, ...]: - """The paragraph's words, less the line breaks at its edges.""" - - self.boundary() - words = list(self.words) - while words and words[0] == ("BR",): - words.pop(0) - while words and words[-1] == ("BR",): - words.pop() - return tuple(words) - - -def _with_format( - fmt: tuple[bool, bool, bool, str | None], name: str -) -> tuple[bool, bool, bool, str | None]: - strong, em, code, href = fmt - kind = _FORMATS.get(name) - return ( - strong or kind == "strong", - em or kind == "em", - code or kind == "code", - href, - ) - - -def _image(tag: Tag, fmt: tuple[bool, bool, bool, str | None]) -> object: - """An image as displayed: its source, and its alt text as it is read.""" - - alt = " ".join(str(tag.get("alt") or "").split()) - return (("IMG", str(tag.get("src") or ""), alt), fmt) - - -def _inline(node: Tag, words: _Words, fmt: tuple[bool, bool, bool, str | None]) -> None: - for child in node.children: - if isinstance(child, _SKIPPED): - continue - if isinstance(child, NavigableString): - for char in str(child): - words.char(char, fmt) - elif isinstance(child, Tag) and child.name not in _HIDDEN: - if child.name == "br": - words.line_break() - elif child.name == "img": - words.token(_image(child, fmt)) - elif child.name == "a" and child.get("href"): - link = (fmt[0], fmt[1], fmt[2], str(child.get("href"))) - _inline(child, words, link) - else: - _inline(child, words, _with_format(fmt, child.name)) - - -def _has_block(node: Tag) -> bool: - return any( - isinstance(child, Tag) and child.name in _BLOCKS for child in node.descendants - ) - - -def _list(node: Tag, fmt: tuple[bool, bool, bool, str | None]) -> tuple[object, ...]: - """A list: each non-empty item with the number it displays.""" - - ordered = node.name == "ol" - start = str(node.get("start") or "1") - first = int(start) if start.isdigit() else 1 - items: list[object] = [] - for position, item in enumerate(node.find_all("li", recursive=False)): - content = _display_blocks(item, fmt) - if content: # Markdown cannot spell an empty list item - items.append((first + position if ordered else None, content)) - return (node.name, tuple(items)) if items else () - - -def _display_blocks( - node: Tag, fmt: tuple[bool, bool, bool, str | None] = _PLAIN -) -> tuple[object, ...]: - out: list[object] = [] - words = _Words() - - def flush() -> None: - nonlocal words - paragraph = words.finish() - if paragraph: - out.append(("p", paragraph)) - words = _Words() - - for child in node.children: - if isinstance(child, _SKIPPED): - continue - if isinstance(child, NavigableString): - for char in str(child): - words.char(char, fmt) - continue - if not isinstance(child, Tag) or child.name in _HIDDEN: - continue - name = child.name - if name in _PARAGRAPHS: - flush() - out.extend(_display_blocks(child, fmt)) - elif name in _HEADINGS: - flush() - heading = _Words() - _inline(child, heading, fmt) - out.append((name, heading.finish())) - elif name in {"ul", "ol"}: - flush() - listed = _list(child, fmt) - if listed: - out.append(listed) - elif name == "blockquote": - flush() - out.append(("quote", _display_blocks(child, fmt))) - elif name == "pre": - flush() - out.append(("pre", child.get_text().strip("\n"))) - elif name == "hr": - flush() - out.append(("hr",)) - elif name == "table": - flush() - rows = [] - for row in child.find_all("tr"): - cells = [] - for cell in row.find_all(["td", "th"], recursive=False): - inner = _Words() - _inline(cell, inner, fmt) - cells.append((cell.name == "th", inner.finish())) - rows.append(tuple(cells)) - out.append(("table", tuple(rows))) - elif name == "br": - words.line_break() - elif name == "img": - words.token(_image(child, fmt)) - elif _has_block(child): - flush() - out.extend(_display_blocks(child, _with_format(fmt, name))) - elif name == "a" and child.get("href"): - _inline(child, words, (fmt[0], fmt[1], fmt[2], str(child.get("href")))) - else: - _inline(child, words, _with_format(fmt, name)) - flush() - return tuple(out) - - -def displayed(html: str) -> tuple[object, ...]: - """Return what a browser displays of ``html``, as comparable data.""" - - return _display_blocks(BeautifulSoup(html, "html.parser")) - - -def assert_survives(html: str) -> str: - """Assert the round-trip property for one body, and return its Markdown.""" - - markdown = read(html) - rendered = render(markdown) - - assert displayed(rendered) == displayed(html), ( - f"the Markdown does not display what the HTML did\n" - f" html: {html!r}\n markdown: {markdown!r}\n rendered: {rendered!r}" - ) - assert read(rendered) == markdown, ( - f"the Markdown is not a fixed point\n markdown: {markdown!r}\n" - f" again: {read(rendered)!r}" - ) - return markdown - - -# --------------------------------------------------------------------------- -# Ordinary prose is left alone -# --------------------------------------------------------------------------- - - -@pytest.mark.parametrize( - ("html", "expected"), - [ - pytest.param( - "<p>Voir fichier_de_test_v2.xlsx et mon_fichier_final.docx</p>", - "Voir fichier_de_test_v2.xlsx et mon_fichier_final.docx", - id="file-names", - ), - pytest.param( - r"<p>Chemin C:\Temp\logs et C:\Users\Admin\Documents</p>", - r"Chemin C:\Temp\logs et C:\Users\Admin\Documents", - id="windows-paths", - ), - pytest.param( - "<p>Le ticket # 3 et le #4521 sont liés, C# aussi</p>", - "Le ticket # 3 et le #4521 sont liés, C# aussi", - id="mid-sentence-hash", - ), - pytest.param( - "<p>Porte-monnaie - un tiret - et -- deux</p>", - "Porte-monnaie - un tiret - et -- deux", - id="dashes-and-hyphens", - ), - pytest.param( - "<p>Service R&D, bâtiment A & B</p>", - "Service R&D, bâtiment A & B", - id="ampersands", - ), - pytest.param( - "<p>5 * 3 = 15 et prix 5*3 et note * importante</p>", - "5 * 3 = 15 et prix 5*3 et note * importante", - id="lone-asterisks", - ), - pytest.param( - "<p>[INFO] tâche [1] terminée (voir note)</p>", - "[INFO] tâche [1] terminée (voir note)", - id="brackets", - ), - pytest.param( - "<p>si a < b et x <= y alors 2 < 3</p>", - "si a < b et x <= y alors 2 < 3", - id="less-than", - ), - pytest.param( - "<p>2 + 2 = 4, +33 6 12 34 56 78, 1) un, 3.14</p>", - "2 + 2 = 4, +33 6 12 34 56 78, 1) un, 3.14", - id="numbers", - ), - pytest.param( - "<p>Cordialement,<br>Jean Dupont<br>--<br>Service IT</p>", - "Cordialement, \nJean Dupont \n-- \nService IT", - id="signature-dashes-on-a-third-line", - ), - pytest.param( - "<p>Voir https://example.org/doc?a=1&b=2 ou support@example.org</p>", - "Voir https://example.org/doc?a=1&b=2 ou support@example.org", - id="bare-url-and-address", - ), - pytest.param( - "<p>Pourquoi ? Parce que ! 100 % a/b a=b ~5 minutes</p>", - "Pourquoi ? Parce que ! 100 % a/b a=b ~5 minutes", - id="punctuation", - ), - pytest.param( - "<p>l`imprimante et la variable user_id et _temp</p>", - "l`imprimante et la variable user_id et _temp", - id="unpaired-backtick-and-underscore", - ), - pytest.param( - "<p>ps aux | grep java</p>", - "ps aux | grep java", - id="pipe-outside-a-table", - ), - ], -) -def test_ordinary_prose_carries_no_escape(html: str, expected: str) -> None: - """Minimal means none of these grows a backslash or a reference. - - Each of these characters *can* be Markdown syntax, and none of them is - here, so python-markdown already renders every one of them literally. - Escaping them anyway would be harmless to the rendering and a nuisance to - everyone who reads the Markdown -- and a change for every existing body. - """ - - assert read(html) == expected - assert_survives(html) - - -# --------------------------------------------------------------------------- -# The measured misreadings, one by one -# --------------------------------------------------------------------------- - - -@pytest.mark.parametrize( - ("html", "expected"), - [ - # The reviewers' measurements through GLPI's reader into - # python-markdown, before this change. - pytest.param( - r"<p>Accès au partage \\serveur\compta\2026 refusé</p>", - r"Accès au partage \\\serveur\compta\2026 refusé", - id="unc-path-lost-a-backslash", - ), - pytest.param( - r"<p>Dossier C:\_temp\logs</p>", - r"Dossier C:\\_temp\logs", - id="backslash-underscore-lost-the-backslash", - ), - pytest.param( - "<p>Fichier mon_fichier_final.docx et __init__</p>", - r"Fichier mon_fichier_final.docx et \_\_init\_\_", - id="dunder-became-bold", - ), - pytest.param( - "<p>Nom : ______ Prénom : ______</p>", - r"Nom : \_\_\_\_\_\_ Prénom : \_\_\_\_\_\_", - id="form-blanks-became-emphasis", - ), - pytest.param( - "<p>Merci<br>-----------<br>Jean Dupont</p>", - "Merci \n\\----------- \nJean Dupont", - id="dash-line-made-a-heading", - ), - pytest.param( - "<p>* point un<br>* point deux</p>", - "\\* point un \n* point deux", - id="star-lines-made-a-list", - ), - pytest.param( - "<p>Le 30/09, Jean a écrit :<br>> merci<br>> cordialement</p>", - "Le 30/09, Jean a écrit : \n\\> merci \n\\> cordialement", - id="quoted-reply-made-a-blockquote", - ), - pytest.param( - "<p>voir la note [1]</p><p>[1]: https://example.org/note</p>", - "voir la note [1]\n\n\\[1]: https://example.org/note", - id="footnote-line-was-consumed", - ), - pytest.param( - "<p># pas un titre</p>", r"\# pas un titre", id="hash-made-a-heading" - ), - pytest.param( - "<p>#4521 est un doublon</p>", - r"\#4521 est un doublon", - id="ticket-number-made-a-heading", - ), - pytest.param( - "<p>2026. Une annee</p>", r"2026\. Une annee", id="year-made-a-list" - ), - pytest.param( - "<table><tr><th>Commande</th></tr>" - "<tr><td>ps aux | grep java</td></tr></table>", - "| Commande |\n| --- |\n| ps aux \\| grep java |", - id="pipe-in-a-cell-dropped-the-rest", - ), - ], -) -def test_the_measured_misreadings_read_back_as_their_text( - html: str, expected: str -) -> None: - """Every shape the reviewers measured, now spelled so it survives.""" - - assert read(html) == expected - assert_survives(html) - - -@pytest.mark.parametrize( - ("html", "expected"), - [ - # Backslashes: escaped exactly when the next character is one - # python-markdown would take as escaped. - pytest.param( - r"<p>D:\logs\.cache et \\srv\share\[archive] et HKLM\SOFTWARE\#1</p>", - r"D:\logs\\.cache et \\\srv\share\\[archive] et HKLM\SOFTWARE\\#1", - id="backslash-before-punctuation", - ), - pytest.param( - r"<p>fin de ligne \<br>suite</p>", - "fin de ligne \\ \nsuite", - id="backslash-before-a-break", - ), - pytest.param( - r"<h2>Chemin C:\</h2>", - r"## Chemin C:\\", - id="backslash-ending-a-heading", - ), - # Emphasis delimiters. - pytest.param("<p>_______</p>", r"\_" * 7, id="seven-underscores"), - pytest.param( - "<p>Code : _______ fin</p>", - "Code : " + r"\_" * 7 + " fin", - id="seven-underscores-mid-sentence", - ), - pytest.param( - "<p>a*b*c et 5*3 <strong>gras</strong></p>", - r"a\*b\*c et 5\*3 **gras**", - id="asterisks-beside-real-emphasis", - ), - pytest.param( - "<p><strong>x</strong>* suite</p>", - r"**x**\* suite", - id="asterisk-touching-a-delimiter", - ), - # Block syntax at the start of a line. - pytest.param("<p>+ un<br>+ deux</p>", "\\+ un \n+ deux", id="plus-list"), - pytest.param("<p>- pas une liste</p>", r"\- pas une liste", id="dash-list"), - pytest.param( - "<p>Bonjour<br>#4521 doublon</p>", - "Bonjour \n\\#4521 doublon", - id="hash-after-a-break", - ), - pytest.param("<p>Titre<br>=====</p>", "Titre \n=====", id="setext-equals"), - pytest.param( - "<p>---</p><p>signature</p>", - "\\---\n\nsignature", - id="rule-of-dashes", - ), - pytest.param("<p>***</p>", r"\*\*\*", id="rule-of-asterisks"), - pytest.param("<ul><li>--</li></ul>", r"- \--", id="bullet-completes-a-rule"), - pytest.param("<ul><li>___</li></ul>", r"- \_\_\_", id="rule-in-a-list-item"), - pytest.param( - "<ul><li><ul><li>-</li></ul></li></ul>", - r"- - \-", - id="two-bullets-complete-a-rule", - ), - pytest.param( - "<ol><li><ul><li>--</li></ul></li></ol>", - r"1. - \--", - id="rule-inside-a-numbered-item", - ), - pytest.param( - "<ul><li><p>--</p><p>suite</p></li></ul>", - "- \\--\n\n suite", - id="rule-in-an-item-paragraph", - ), - pytest.param("<h1>C#</h1>", r"# C\#", id="heading-trailing-hash"), - pytest.param("<h2>Titre ##</h2>", r"## Titre \#\#", id="heading-closing-run"), - pytest.param( - "<p>```<br>code<br>```</p>", - "\\`\\`\\` \ncode \n\\`\\`\\`", - id="backtick-fence", - ), - pytest.param( - "<p>~~~<br>code<br>~~~</p>", - "~~~ \ncode \n~~~", - id="tilde-fence", - ), - pytest.param( - "<p>a | b<br>--- | ---</p>", - "a | b \n--- \\| ---", - id="table-separator", - ), - # Link and image syntax typed as text. - pytest.param( - "<p>[x](https://example.org/y) et ![x](https://example.org/y.png)</p>", - r"\[x](https://example.org/y) et !\[x](https://example.org/y.png)", - id="link-and-image-syntax", - ), - pytest.param( - '<p>Attention!<a href="https://example.org/u">voir</a></p>', - r"Attention\![voir](https://example.org/u)", - id="bang-before-a-link", - ), - pytest.param( - '<p><a href="https://example.org/u">rapport [final].pdf</a></p>', - "[rapport [final].pdf](https://example.org/u)", - id="balanced-brackets-in-link-text", - ), - pytest.param( - '<p><a href="https://example.org/u">a]b</a></p>', - r"[a\]b](https://example.org/u)", - id="unbalanced-bracket-in-link-text", - ), - pytest.param( - '<p><a href="https://example.org/u">voir [x](y)</a></p>', - r"[voir \[x\](y)](https://example.org/u)", - id="link-syntax-in-link-text", - ), - pytest.param( - '<p>[a <a href="https://example.org/u">[b</a> ](https://example.org/x)</p>', - r"\[a [\[b](https://example.org/u) ](https://example.org/x)", - id="escaped-bracket-in-link-text-is-not-counted", - ), - pytest.param( - '<p><img src="https://example.org/c.png" alt="capture [1].png"></p>', - "![capture [1].png](https://example.org/c.png)", - id="balanced-brackets-in-alt-text", - ), - pytest.param( - '<p><img src="https://example.org/c.png" alt="a]b"></p>', - r"![a\]b](https://example.org/c.png)", - id="unbalanced-bracket-in-alt-text", - ), - # Raw HTML and references typed as text. - pytest.param( - "<p>appuyer sur <Entrée> puis valider</p>", - "appuyer sur <Entrée> puis valider", - id="angle-bracketed-word", - ), - pytest.param( - "<p>if x<y then z>0</p>", - "if x<y then z>0", - id="comparison-that-looks-like-a-tag", - ), - pytest.param( - "<p><https://example.org/x> et <support@example.org></p>", - "<https://example.org/x> et <support@example.org>", - id="autolink-syntax", - ), - pytest.param( - '<p><3<img src="https://example.org/i.png" alt="a b">@c></p>', - "<3![a b](https://example.org/i.png)@c>", - id="address-running-through-an-image", - ), - pytest.param( - '<p><a href="https://example.org/u"><<a@c></a></p>', - "[<<a@c>](https://example.org/u)", - id="address-in-link-text", - ), - pytest.param( - '<p><3<a href="https://example.org">https://example.org</a>@c></p>', - "<3<https://example.org>@c>", - id="address-running-through-an-autolink", - ), - pytest.param( - '<p><a <a href="https://example.org/T_(x)">wiki</a>b@c></p>', - "<a [wiki](https://example.org/T_(x))b@c>", - id="link-target-holding-parentheses", - ), - pytest.param( - "<p><!-- note --></p>", "<!-- note -->", id="comment-syntax" - ), - pytest.param( - "<p>&amp; &lt; &#65; &#4521 &copy; &copy</p>", - "&amp; &lt; &#65; &#4521 &copy; ©", - id="character-references", - ), - # Code spans typed as text. - pytest.param( - "<p>a`b`c et <code>x</code> puis `</p>", - r"a\`b\`c et `x` puis `", - id="backtick-pair", - ), - pytest.param( - "<p><code>x</code>`y</p>", - r"`x`\`y", - id="backtick-touching-a-code-span", - ), - pytest.param( - "<p>`<code>x</code> y</p>", - r"\``x` y", - id="backtick-opening-onto-a-code-span", - ), - pytest.param( - "<table><tr><th>a</th><th>b</th></tr>" - "<tr><td>x`y</td><td>`z</td></tr></table>", - "| a | b |\n| --- | --- |\n| x\\`y | \\`z |", - id="backticks-pair-across-cells", - ), - ], -) -def test_literal_text_is_escaped_where_python_markdown_would_read_it( - html: str, expected: str -) -> None: - """One case per construct, each with the escape it needs and no other.""" - - assert read(html) == expected - assert_survives(html) - - -# --------------------------------------------------------------------------- -# Structure that used to be lost on the way through -# --------------------------------------------------------------------------- - - -@pytest.mark.parametrize( - ("html", "expected"), - [ - pytest.param( - "<ul><li>Réseau<ul><li>switch 3</li><li>borne wifi</li></ul></li>" - "<li>Imprimante</li></ul>", - "- Réseau\n - switch 3\n - borne wifi\n- Imprimante", - id="ul-in-ul", - ), - pytest.param( - "<ol><li>Arreter</li><li>Sauvegarder<ol><li>la base</li>" - "<li>les fichiers</li></ol></li><li>Redemarrer</li></ol>", - "1. Arreter\n2. Sauvegarder\n 1. la base\n 2. les fichiers\n" - "3. Redemarrer", - id="ol-in-ol-keeps-its-numbering", - ), - pytest.param( - "<ol><li>un<ul><li>a</li></ul></li><li>deux</li></ol>", - "1. un\n - a\n2. deux", - id="ul-in-ol", - ), - pytest.param( - "<ul><li>a<ul><li>b<ul><li>c<ul><li>d</li></ul></li></ul></li></ul>" - "</li></ul>", - "- a\n - b\n - c\n - d", - id="four-levels", - ), - pytest.param( - '<p>intro</p><ol start="3"><li>trois</li><li>quatre</li></ol>', - "intro\n\n3. trois\n4. quatre", - id="ordered-list-starting-at-three", - ), - pytest.param( - "<ul><li>5*3<ul><li>2*4</li></ul></li></ul>", - "- 5*3\n - 2*4", - id="asterisks-an-item-and-its-nested-item-cannot-pair", - ), - pytest.param( - "<ul><li><p>point un</p><p>suite du point</p></li><li>point deux</li></ul>", - "- point un\n\n suite du point\n\n- point deux", - id="item-with-two-paragraphs", - ), - ], -) -def test_nested_lists_nest_and_keep_their_numbers(html: str, expected: str) -> None: - """python-markdown nests only at four spaces; the reader indents by four. - - ``markdownify`` indented a continuation by its bullet's width -- two for - ``- ``, three for ``1. `` -- so the first write flattened a nested list - and renumbered a nested ordered one, one level per pass. - """ - - assert read(html) == expected - assert_survives(html) - - -@pytest.mark.parametrize( - ("html", "expected"), - [ - pytest.param( - "<ul><li>Réseau<ul><li>switch 3</li></ul>à vérifier</li></ul>", - "- Réseau\n\n - switch 3\n\n à vérifier", - id="text-after-a-nested-list", - ), - pytest.param( - "<ul><li>Réseau<ul><li>switch 3</li></ul><blockquote>cité</blockquote>" - "</li></ul>", - "- Réseau\n\n - switch 3\n\n > cité", - id="quote-after-a-nested-list", - ), - pytest.param( - "<ul><li>Réponse :<blockquote>cité</blockquote></li></ul>", - "- Réponse :\n\n > cité", - id="quote-after-an-items-text", - ), - pytest.param( - "<ul><li><blockquote>cité<br>suite</blockquote></li></ul>", - "- > cité \n> suite", - id="quote-opening-an-item", - ), - pytest.param( - "<ul><li><ul><li>un</li><li>deux</li></ul></li></ul>", - "- - un\n - deux", - id="item-opening-with-a-list", - ), - pytest.param( - "<ul><li><ul><li>un<ul><li>a</li></ul></li></ul></li></ul>", - "- - un\n\n - a", - id="list-under-an-item-sharing-its-line", - ), - pytest.param( - "<ul><li><ul><li><ul><li>un</li><li>deux</li></ul></li></ul></li>" - "<li>trois</li></ul>", - "- - - un\n\n - deux\n\n- trois", - id="three-bullets-on-one-line", - ), - ], -) -def test_blocks_inside_a_list_item_stay_in_it(html: str, expected: str) -> None: - """python-markdown only nests a block where its first pass lets it. - - That pass never detabs an item's first block, and reads ``>`` only three - spaces in at most, so what follows a list item's text or its nested list - needs a blank line before it -- ``markdownify`` gave none after a nested - list, and the item's text ran into the list's last item -- and a quote - opening an item needs its later lines unindented. On a line already - carrying two bullets, anything eight spaces in is taken as that line's - continuation, so what follows it starts a block of its own. - """ - - assert read(html) == expected - assert_survives(html) - - -@pytest.mark.parametrize( - ("html", "expected"), - [ - pytest.param( - '<p><a name="_MailEndCompose">Bonjour</a> Jean</p>', - "Bonjour Jean", - id="anchor-without-a-target", - ), - pytest.param( - '<p><a href="https://example.org/u"></a>texte</p>', "texte", id="empty-link" - ), - pytest.param("<p><strong></strong>texte</p>", "texte", id="empty-emphasis"), - pytest.param( - '<p><a href="https://example.org/u">un<br>deux</a></p>', - "[un \ndeux](https://example.org/u)", - id="break-in-link-text", - ), - pytest.param("<p>a<center>b</center>c</p>", "a\n\nb\n\nc", id="center-block"), - pytest.param( - "<table><tr><th><h3>titre</h3></th></tr><tr><td>x</td></tr></table>", - "| titre |\n| --- |\n| x |", - id="heading-in-a-cell", - ), - pytest.param( - "<h2><blockquote>cité</blockquote></h2>", "## cité", id="quote-in-a-heading" - ), - pytest.param( - "<table><tr><th>a</th></tr><tr><td>un<br>deux</td></tr></table>", - "| a |\n| --- |\n| un deux |", - id="break-in-a-cell", - ), - pytest.param( - "<h2>Titre<br>suite</h2>", "## Titre suite", id="break-in-a-heading" - ), - pytest.param( - '<p><img src="https://example.org/i.png" alt="ligne 1\n\n ligne 2"></p>', - "![ligne 1 ligne 2](https://example.org/i.png)", - id="alt-text-over-several-lines", - ), - pytest.param( - "<ul><li><script>x()</script><pre>code</pre></li></ul>", - "- code", - id="code-opening-an-item-after-a-script", - ), - pytest.param( - "<ul><li><!-- note --><pre>code</pre></li></ul>", - "- code", - id="code-opening-an-item-after-a-comment", - ), - pytest.param( - "<ul><li><ul></ul><pre>code</pre></li></ul>", - "- code", - id="code-opening-an-item-after-an-empty-list", - ), - pytest.param( - '<ul><li><img src="https://example.org/i.png" alt="i"><pre>code</pre>' - "</li></ul>", - "- ![i](https://example.org/i.png)\n\n code", - id="code-after-an-image", - ), - ], -) -def test_shapes_with_a_rule_of_their_own(html: str, expected: str) -> None: - """The converter's special cases, each on the shape that reaches it. - - What displays nothing -- a comment, a script, an empty list -- does not - count as content before a code block, so the code still opens its item. - """ - - assert read(html) == expected - - -def test_an_ordered_list_keeps_its_start_when_rendered() -> None: - """``sane_lists`` is what keeps ``3.`` from restarting the count at 1.""" - - assert '<ol start="3">' in render("intro\n\n3. trois\n4. quatre") - - -@pytest.mark.parametrize( - ("html", "expected"), - [ - pytest.param( - "<ul><li>item<pre>#4521 code</pre></li></ul>", - "- item\n\n #4521 code", - id="in-a-list-item", - ), - pytest.param( - "<blockquote><pre>#4521 C:\\Temp\n indenté</pre></blockquote>", - "> #4521 C:\\Temp\n> indenté", - id="in-a-blockquote", - ), - pytest.param( - "<pre>ligne\n```\nfin</pre>", - "````\nligne\n```\nfin\n````", - id="holding-a-fence-line", - ), - ], -) -def test_a_preformatted_block_stays_one(html: str, expected: str) -> None: - """A fence opens only at the start of a line, so nested code is indented. - - Inside a list item or a block quote, python-markdown never sees - ``` ``` ``` at the start of a line and reads the fence as text -- the - code's own ``#4521`` then became a heading. An indented code block is the - spelling it does read there. At the top level the fence is made longer - than any fence line the code holds. - """ - - assert read(html) == expected - assert_survives(html) - - -@pytest.mark.xfail( - strict=True, - reason=( - "python-markdown runs html.parser over its whole source to find raw " - "HTML, before an indented code block is recognised, and re-emits a " - "numeric reference written without its semicolon with one: the code " - "then shows 'ᆩ'. A fence is stashed before that pass, which is " - "why a top-level <pre> is unaffected; inside a list item or a quote no " - "fence can open. Measured on 3.10.3." - ), -) -def test_a_numeric_reference_in_nested_code_gains_a_semicolon() -> None: - assert_survives("<blockquote><pre>echo &#4521</pre></blockquote>") - - -def test_a_preformatted_block_opening_a_list_item_keeps_its_text() -> None: - """The one place python-markdown can start no code block at all. - - A list item's first line is its paragraph, so the code there degrades to - its lines, each escaped as the literal text it is -- the words survive, - the preformatting does not. - """ - - html = "<ul><li><pre>#4521 code\n suite</pre></li></ul>" - - markdown = read(html) - - assert markdown == "- \\#4521 code \n suite" - assert "#4521 code" in render(markdown) - assert read(render(markdown)) == markdown - - -@pytest.mark.parametrize( - ("html", "expected"), - [ - pytest.param( - "<ul><li>item<ul><li>sous-item</li></ul><pre>#4521 code</pre></li></ul>", - "- item\n\n - sous-item\n\n \\#4521 code", - id="in-a-list-item", - ), - pytest.param( - "<blockquote><ul><li>item</li></ul><pre>#4521 code</pre></blockquote>", - "> - item\n>\n> \\#4521 code", - id="in-a-quote", - ), - pytest.param( - "<ul><li>item<ul><li>sous-item</li></ul><div><style>p{margin:0}</style>" - "</div><pre>#4521 code</pre></li></ul>", - "- item\n\n - sous-item\n\n \\#4521 code", - id="with-only-a-style-block-between", - ), - ], -) -def test_code_right_after_a_nested_list_keeps_its_text( - html: str, expected: str -) -> None: - """The other place python-markdown can start no code block. - - An indented code block right after a list in the same item or quote is - indented exactly as the list's last item's own content, and - python-markdown reads it as a paragraph of that item. The lines are - kept as literal text in the right place instead. - """ - - markdown = read(html) - - assert markdown == expected - assert "#4521 code" in render(markdown) - assert read(render(markdown)) == markdown - - -#: What the format loses, whatever the escaping does: each body with why. -#: -#: None of these is literal text misread. Each is structure python-markdown -#: has no spelling for, or spells as something else. -LOSSES = [ - pytest.param( - "<ul><li>a</li></ul><ul><li>b</li></ul>", - "python-markdown continues a list across a blank line: two lists are " - "read back as one list of loose items.", - id="two-adjacent-lists", - ), - pytest.param( - "<blockquote>a</blockquote><blockquote>b</blockquote>", - "python-markdown continues a quote across a blank line: two quotes are " - "read back as one quote of two paragraphs.", - id="two-adjacent-quotes", - ), - pytest.param( - "<p>a<br><br>b</p>", - "Markdown has no blank line inside a paragraph: two breaks in a row " - "are a paragraph break, and read back as one.", - id="two-breaks-in-a-row", - ), - pytest.param( - "<p><code>a</code><code>b</code></p>", - "'`a``b`' is one code span holding 'a``b' to python-markdown.", - id="two-adjacent-code-spans", - ), - pytest.param( - "<p><em>a <strong>b</strong> c</em></p>", - "python-markdown pairs '*a **b** c*' as three emphasis runs, and b " - "loses its bold.", - id="strong-inside-emphasis", - ), - pytest.param( - "<table><tr><td>a</td><td>b</td></tr></table>", - "A Markdown table starts with its header row, so one without gains an " - "empty header.", - id="table-without-a-header", - ), - pytest.param( - "<table><tr><th>a</th><th>b</th></tr></table>", - "python-markdown renders a header-only table with one empty body row.", - id="table-without-a-body", - ), - pytest.param( - "<table><tr><th>a</th></tr><tr><td><ul><li>x</li></ul></td></tr></table>", - "A table cell holds one line of inline Markdown: its list is written " - "as its text.", - id="list-in-a-cell", - ), - pytest.param( - "<table><tr><th>a</th></tr><tr><td><pre>x\ny</pre></td></tr></table>", - "A table cell holds one line of inline Markdown: its code block is " - "written as inline code.", - id="code-block-in-a-cell", - ), - pytest.param( - "<table><tr><th>a</th></tr><tr><td>un<br>deux</td></tr></table>", - "A table cell holds one line: its line break is written as a space.", - id="break-in-a-cell", - ), - pytest.param( - "<h2>Titre<br>suite</h2>", - "A heading is one line: its line break is written as a space.", - id="break-in-a-heading", - ), - pytest.param( - "<ul><li><pre>a</pre></li></ul>", - "An item's first block is its paragraph: the code is kept as text " - "(test_a_preformatted_block_opening_a_list_item_keeps_its_text).", - id="code-opening-a-list-item", - ), - pytest.param( - "<ul><li>a<ul><li>b</li></ul><pre>c</pre></li></ul>", - "The code would be indented as the nested item's own content: it is " - "kept as text (test_code_right_after_a_nested_list_keeps_its_text).", - id="code-after-a-nested-list", - ), -] - - -@pytest.mark.parametrize(("html", "reason"), LOSSES) -def test_what_markdown_cannot_carry(html: str, reason: str) -> None: - """The inventory of losses, each asserted to still be one. - - ``xfail(strict=True)`` is the point: a loss that stops being one fails - here, so the inventory stays true. Struck and underlined text lose their - line too, which this comparison does not see: - ``test_struck_text_keeps_its_words_and_loses_its_line``. - """ - - with pytest.raises(AssertionError): - assert_survives(html) - pytest.xfail(reason) - - -@pytest.mark.parametrize(("html", "reason"), LOSSES) -def test_what_is_lost_is_lost_once(html: str, reason: str) -> None: - """Whatever a loss costs, it costs on the first cycle and never again.""" - - again = read(render(read(html))) - - assert read(render(again)) == again, reason - - -@pytest.mark.parametrize( - "html", - [ - pytest.param( - "<style>p.MsoNormal{margin:0cm;font-size:11pt}</style><p>Bonjour</p>", - id="style", - ), - pytest.param("<p>Bonjour</p><script>track()</script>", id="script"), - pytest.param( - "<html><head><title>RE: Imprimante" - "" - "

      Bonjour

      ", - id="outlook-shaped", - ), - ], -) -def test_style_script_and_title_bodies_are_not_text(html: str) -> None: - """A browser displays none of these, so neither does the Markdown. - - ``strip=["script", "style"]`` used to be passed to ``markdownify``, and - ``strip`` skips an element's own converter -- ``convert_script`` and - ``convert_style`` return ``""`` -- so the bodies leaked into the text as - prose. ```` has no converter at all and leaked the same way. - """ - - assert read(html) == "Bonjour" - - -@pytest.mark.parametrize( - ("html", "expected"), - [ - pytest.param( - '<font color="red">URGENT</font> serveur HS', "URGENT serveur HS", id="font" - ), - pytest.param("<center>Titre</center> suite", "Titre\n\nsuite", id="center"), - pytest.param("<strike>ancien</strike> nouveau", "ancien nouveau", id="strike"), - pytest.param("<big>gros</big> texte", "gros texte", id="big"), - pytest.param("<tt>code</tt> texte", "code texte", id="tt"), - pytest.param("<nobr>sans coupure</nobr>", "sans coupure", id="nobr"), - ], -) -def test_a_body_marked_up_only_with_obsolete_elements_is_html( - html: str, expected: str -) -> None: - """Old editors still write these; without them the tags were kept as text.""" - - assert read(html) == expected - - -def test_every_obsolete_element_name_makes_a_body_html() -> None: - """The HTML standard's list of obsolete elements, all of them recognised.""" - - obsolete = ( - "acronym applet basefont bgsound big blink center dir font frame " - "frameset isindex keygen listing marquee menuitem multicol nextid nobr " - "noembed noframes plaintext rb rtc spacer strike tt xmp" - ).split() - - assert set(obsolete) <= conversion._HTML_ELEMENTS - - -@pytest.mark.parametrize( - "html", - [ - pytest.param("<p><s>ancien</s> nouveau</p>", id="s"), - pytest.param("<p><del>ancien</del> nouveau</p>", id="del"), - pytest.param("<p><strike>ancien</strike> nouveau</p>", id="strike"), - ], -) -def test_struck_text_keeps_its_words_and_loses_its_line(html: str) -> None: - """python-markdown has no strikethrough, so ``~~x~~`` would show literally. - - Keeping the words and losing the line is the honest loss: the text is - all there, and what is missing is recorded in the round-trip inventory. - """ - - assert read(html) == "ancien nouveau" - - -@pytest.mark.parametrize( - "html", ["<P>Bonjour</P>", "<BR>Bonjour", "<DIV><B>Bonjour</B></DIV>"] -) -def test_upper_case_tags_are_html(html: str) -> None: - """Element names are case-insensitive; the probe has to be too.""" - - assert "<" not in read(html) - - -@pytest.mark.parametrize( - ("value", "expected"), - [ - pytest.param(" texte ", "texte", id="plain-text"), - pytest.param(" <b>x</b> ", "**x**", id="html"), - pytest.param("<br>x<br>", "x", id="html-with-edge-breaks"), - ], -) -def test_the_result_is_stripped_on_both_paths(value: str, expected: str) -> None: - """Edge whitespace is never content, and a digest must not depend on it.""" - - assert read(value) == expected - - -@pytest.mark.parametrize( - ("html", "expected"), - [ - pytest.param( - '<h2>Titre <img src="https://example.org/i.png" alt="logo"></h2>', - "## Titre ![logo](https://example.org/i.png)", - id="heading", - ), - pytest.param( - "<table><tr><th>a</th></tr><tr><td>" - '<img src="https://example.org/i.png" alt="x"></td></tr></table>', - "| a |\n| --- |\n| ![x](https://example.org/i.png) |", - id="table-cell", - ), - ], -) -def test_an_image_in_a_heading_or_a_cell_stays_an_image( - html: str, expected: str -) -> None: - """``markdownify`` reduced these to their alt text; python-markdown needs not.""" - - assert read(html) == expected - assert_survives(html) - - -def test_a_newline_in_text_is_a_space() -> None: - """HTML displays a source newline as a space; ``nl2br`` would break the line. - - Outlook wraps its HTML source mid-sentence, so every e-mail created - ticket used to gain line breaks where the reader saw none. - """ - - html = "<p>Bonjour,\nle serveur\nest redémarré.</p>" - - assert read(html) == "Bonjour, le serveur est redémarré." - assert_survives(html) - - -@pytest.mark.parametrize( - ("html", "expected"), - [ - pytest.param("<p>a<br> b<br> c</p>", "a \nb \nc", id="space-after-break"), - pytest.param("<p>a <br>b</p>", "a \nb", id="space-before-break"), - pytest.param("<p><strong>a<br></strong>b</p>", "**a** \nb", id="edge-break"), - ], -) -def test_a_line_break_has_one_spelling(html: str, expected: str) -> None: - """Whitespace around ``<br>`` is not displayed, so it is not kept either. - - Without this the first read of ``a<br> b`` was ``a \\n b`` and the - second ``a \\nb``: the same body, read twice, disagreeing. - """ - - assert read(html) == expected - assert_survives(html) - - -# --------------------------------------------------------------------------- -# The degraded path spells text the same way -# --------------------------------------------------------------------------- - - -def test_a_body_too_deep_to_convert_is_still_literal_safe() -> None: - """The stripped text is Markdown too, and it is escaped like the rest.""" - - html = "<div>" * 600 + r"<p>__init__ \\serveur</p><p># pas un titre</p>" - - markdown = read(html) - - assert markdown == "\\_\\_init\\_\\_ \\\\\\serveur\n\n\\# pas un titre" - assert "__init__ \\\\serveur" in render(markdown) - - -@pytest.mark.parametrize("seed", range(4)) -def test_stripped_text_renders_as_itself(seed: int) -> None: - """Any text, line by line, renders as exactly that text. - - The degraded path's contract, over generated text dense with the - characters python-markdown treats as syntax. - """ - - rng = random.Random(seed) - for _ in range(60): - lines = [_text(rng, rng.randint(1, 6)) for _ in range(rng.randint(1, 4))] - text = "\n".join(line for line in lines if line) - if not text: - continue - - soup = BeautifulSoup(render(conversion._literal_markdown(text)), "html.parser") - for line_break in soup.find_all("br"): - line_break.replace_with("\u2029") - shown = soup.get_text().replace("\n", " ").replace("\u2029", "\n") - - assert [" ".join(line.split()) for line in shown.split("\n")] == [ - " ".join(line.split()) for line in text.split("\n") - ], text - - -# --------------------------------------------------------------------------- -# The property, over realistic bodies and over a fuzzer -# --------------------------------------------------------------------------- - -#: Ticket bodies shaped the way GLPI's editor and mail collector store them. -REALISTIC = [ - "<p>Bonjour,</p><p>Le PC du poste 12 ne démarre plus depuis ce matin.</p>", - "<p>Bonjour,<br>Le PC ne démarre plus.<br>Cordialement,<br>Jean</p>", - "<p>Merci de <strong>redémarrer</strong> le serveur <em>avant</em> 18h.</p>", - "<p>Étapes :</p><ul><li>ouvrir la session</li><li>lancer Outlook</li></ul>", - "<table><thead><tr><th>Poste</th><th>IP</th></tr></thead>" - "<tbody><tr><td>PC12</td><td>10.0.0.12</td></tr></tbody></table>", - "<h2>Contexte</h2><p>Migration du serveur.</p>", - '<p>Voir <a href="https://example.org/doc">https://example.org/doc</a></p>', - '<p>Voir <a title="doc" href="https://example.org/doc">la doc</a></p>', - '<p><a href="https://example.org/wiki/Test_(informatique)">wiki</a></p>', - "<p>prix 5*3 et note * importante</p>", - "<p>appuyer sur <Entrée> puis valider</p>", - "<p>Service R&D, bâtiment A & B</p>", - "<p>Montant : 12 000 €</p>", - '<p><span style="color: #e03e2d;">URGENT</span> <u>à traiter</u></p>', - '<p>Cordialement</p><p><img src="https://example.org/logo.png" alt="Logo"></p>', - "<p>Bonjour</p><blockquote><p>Message d'origine</p></blockquote>", - '<pre>Traceback (most recent call last):\n File "x.py", line 1\nError</pre>', - "<p>Lancer <code>ipconfig /all</code> puis envoyer.</p>", - "<p>Fichier mon_fichier_final.docx et __init__</p>", - "<p># pas un titre</p><p>- pas une liste</p>", - '<p>Contact <a href="mailto:support@example.org">support@example.org</a></p>', - "<div>Bonjour,</div><div><br></div><div>Le serveur est down.</div>", - "<p>* point un<br>* point deux</p>", - "<p><https://example.org/x></p>", - "<p>[INFO] tâche [1] terminée</p>", - "<p>[1]: https://example.org/note</p>", - "<p>m<sup>2</sup></p>", - r"<p>Chemin C:\Users\jdupont\Desktop</p>", - "<p>Merci 👍</p>", - "<p>Titre<br>=====</p>", - "<p>---</p><p>signature</p>", - "<p>1) un<br>2) deux</p>", - "<p><support@example.org></p>", - "<p><!-- note --></p>", - "<ul><li><p>point un</p><p>suite du point</p></li><li>point deux</li></ul>", - "<blockquote><ul><li>a</li><li>b</li></ul></blockquote>", - "<p>Nom : ______ Prénom : ______</p>", - "<p>~~pas barré~~</p>", - "<p>a | b | c</p>", - "<p>> pas une citation</p>", - "<p>+ un<br>+ deux</p>", - "<p>Cordialement,<br>Jean Dupont<br>--<br>Service IT</p>", - "<p>Merci<br>-----------<br>Jean Dupont</p>", - "<p>calcul 5 * 3 * 2 = 30</p>", - r"<p>voir \\srv\partage\__archive__\2026</p>", - "<p>Le 30/09, Jean a écrit :<br>> merci<br>> cordialement</p>", - r"<p>Accès au partage \\serveur\compta\2026 refusé</p>", - r"<p>Dossier C:\_temp\logs</p>", - "<p>voir la note [1]</p><p>[1]: https://example.org/note</p>", - "<ul><li>Réseau<ul><li>switch 3</li><li>borne wifi</li></ul></li></ul>", - "<p><b>Important</b> : voir <i>ci-dessous</i></p>", - '<p><font color="red">rouge</font> <span style="font-size:14px">texte</span></p>', - "<p>module __init__ et _x_</p>", - "<p>1) un<br>2) deux</p><pre>ligne 1\n ligne indentée</pre>", - "<p>Suite à la mise à jour, <strong>3 postes</strong> ne se connectent plus :" - "</p><ul><li>PC12 (salle 3)</li><li>PC14 – <em>poste d'accueil</em></li>" - '</ul><p>Voir <a href="https://example.org/kb/42">https://example.org/kb/42</a>' - ' et le <a href="https://example.org/kb/43" title="KB 43">KB 43</a>.</p>', -] - - -@pytest.mark.parametrize("html", REALISTIC) -def test_realistic_bodies_display_the_same_after_the_round_trip(html: str) -> None: - assert_survives(html) - - -#: Literal text snippets: every character python-markdown treats as syntax, -#: alone and in the shapes that trigger it, beside ordinary words so that -#: each lands at line starts, at word boundaries and inside words. -_SPECIAL = ( - "\\ \\\\ \\* \\_ \\[ \\# \\. \\` \\\\serveur * ** *** _ __ ___ ______ __init__ " - "_x_ ` `` ``` ~ ~~ ~~~ [ ] [x] [x](y) ![x](y) [1]: [1]:https://example.org/n " - "( ) (y) ! # ## #4521 > - -- --- + . 1. 2026. = == === | a|b --- < <b> </b> " - "<Enter> <!-- --> <https://example.org> <a@b.c> <3 <= & & < A " - "A ᆩ © © &D; { } : \" '" -).split(" ") -_WORDS = ( - "alpha beta snake_case C:\\Temp R&D x mot 5*3 a_b é fichier_de_test_v2.xlsx" -).split(" ") - - -#: Code content. A numeric reference with no semicolon is left out: inside an -#: indented code block python-markdown's raw-HTML pass adds the semicolon -#: (``test_a_numeric_reference_in_nested_code_gains_a_semicolon``). -_CODE_SPECIAL = [special for special in _SPECIAL if not special.startswith("&#")] - - -def _text(rng: random.Random, pieces: int, specials: list[str] = _SPECIAL) -> str: - out = [] - for _ in range(pieces): - out.append(rng.choice(specials) if rng.random() < 0.45 else rng.choice(_WORDS)) - out.append(rng.choice(["", " ", " ", " "])) - return "".join(out).strip() - - -def _code_text(rng: random.Random, pieces: int) -> str: - return escape(_text(rng, pieces, _CODE_SPECIAL), quote=False) - - -def _inline_html(rng: random.Random, formatted: bool = False, depth: int = 0) -> str: - """Inline HTML that Markdown can express. - - No emphasis inside emphasis, no link inside a link, and a word on either - side of every code element: python-markdown mis-pairs nested ``*`` runs - across constructs and merges two adjacent code spans, neither of which - involves literal text -- they are recorded in the round-trip inventory. - """ - - parts = [] - for _ in range(rng.randint(1, 4)): - kind = rng.random() - if kind < 0.55 or depth > 2: - parts.append(escape(_text(rng, rng.randint(1, 4)), quote=False)) - elif kind < 0.65 and not formatted: - parts.append(f" <strong>{_inline_html(rng, True, depth + 1)}</strong> ") - elif kind < 0.75 and not formatted: - parts.append(f" <em>{_inline_html(rng, True, depth + 1)}</em> ") - elif kind < 0.82 and depth == 0: - inner = _inline_html(rng, formatted, depth + 1) - parts.append( - f'<a href="https://example.org/{rng.randint(1, 9)}">{inner}</a>' - ) - elif kind < 0.88: - code = escape(rng.choice(_WORDS + _SPECIAL[:20]), quote=False) - parts.append(f" mot <code>{code}</code> mot ") - elif kind < 0.94: - alt = escape(_text(rng, rng.randint(0, 2)), quote=True) - parts.append( - f'<img src="https://example.org/{rng.randint(1, 9)}.png" alt="{alt}">' - ) - else: - parts.append(f"<span>{_inline_html(rng, formatted, depth + 1)}</span>") - parts.append(rng.choice(["", " ", " "])) - return "".join(parts).strip() or "mot" - - -def _paragraph(rng: random.Random) -> str: - return "<br>".join(_inline_html(rng) for _ in range(rng.randint(1, 3))) - - -def _list_html(rng: random.Random, depth: int) -> str: - """A list whose items hold text, and sometimes code, a list, and more text. - - The code goes before an item's nested list, never right after it: - python-markdown has no way to write a code block there - (``test_code_right_after_a_nested_list_keeps_its_text``). - """ - - tag = rng.choice(["ul", "ol"]) - items = [] - for _ in range(rng.randint(1, 3)): - shape = rng.random() - if depth < 3 and shape < 0.1: - # An item that opens with a list: one line holds both bullets. - items.append(f"<li>{_list_html(rng, depth + 1)}</li>") - continue - inner = f"<p>{_paragraph(rng)}</p>" if shape < 0.3 else _paragraph(rng) - if rng.random() < 0.1: - inner += f"<pre>{_code_text(rng, 3)}</pre>" - if depth < 3 and rng.random() < 0.3: - inner += _list_html(rng, depth + 1) - after = rng.random() - if after < 0.15: - inner += _inline_html(rng) - elif after < 0.25: - inner += f"<p>{_paragraph(rng)}</p>" - elif after < 0.3: - inner += f"<blockquote><p>{_paragraph(rng)}</p></blockquote>" - items.append(f"<li>{inner}</li>") - return f"<{tag}>{''.join(items)}</{tag}>" - - -def _block(rng: random.Random, depth: int = 0) -> str: - kind = rng.random() - if kind < 0.35 or depth > 1: - return f"<p>{_paragraph(rng)}</p>" - if kind < 0.45: - level = rng.randint(1, 3) - return f"<h{level}>{_inline_html(rng)}</h{level}>" - if kind < 0.60: - return _list_html(rng, depth + 1) - if kind < 0.70: - return f"<blockquote>{_block(rng, depth + 1)}</blockquote>" - if kind < 0.80: - head = "".join(f"<th>{_inline_html(rng)}</th>" for _ in range(2)) - rows = "".join( - "<tr>" - + "".join(f"<td>{_inline_html(rng)}</td>" for _ in range(2)) - + "</tr>" - for _ in range(rng.randint(1, 2)) - ) - return f"<table><tr>{head}</tr>{rows}</table>" - if kind < 0.87: - return f"<pre>{_code_text(rng, 4)}</pre>" - return f"<div>{_paragraph(rng)}</div>" - - -def _document(rng: random.Random) -> str: - """A body of one to four blocks. - - python-markdown merges two adjacent lists of one type, or two adjacent - block quotes, into one -- a limitation of the format -- so a paragraph - separates them. - """ - - blocks: list[str] = [] - for _ in range(rng.randint(1, 4)): - block = _block(rng) - mergeable = block[:4] in {"<ul>", "<ol>", "<blo"} - if blocks and mergeable and block[:4] == blocks[-1][:4]: - blocks.append("<p>mot</p>") - blocks.append(block) - return "".join(blocks) - - -@pytest.mark.parametrize("seed", range(8)) -def test_generated_bodies_display_the_same_after_the_round_trip(seed: int) -> None: - """The property over a seeded fuzzer, 50 bodies a seed.""" - - rng = random.Random(seed) - for _ in range(50): - assert_survives(_document(rng)) - - -@pytest.mark.parametrize( - "html", - [ - pytest.param("<p>" + "[a " * 20_000 + "</p>", id="unclosed-brackets"), - pytest.param("<p>" + "_a " * 20_000 + "</p>", id="underscores-opening-words"), - pytest.param("<p>" + "``` " * 20_000 + "</p>", id="backtick-runs"), - pytest.param("<ol>" + "<li>x</li>" * 20_000 + "</ol>", id="long-numbered-list"), - ], -) -def test_a_body_dense_with_syntax_converts_in_linear_time(html: str) -> None: - """Every pass is linear, so a pathological body costs what its size does. - - Measured before the passes were made linear: 20,000 of these took from - 8 to 117 seconds each, where they take well under one now. The budget - is generous on purpose -- this guards the complexity, not the speed. - """ - - started = time.perf_counter() - read(html) - - assert time.perf_counter() - started < 10 - - -# --------------------------------------------------------------------------- -# What the escaping rests on -# --------------------------------------------------------------------------- - - -def test_python_markdown_undoes_every_backslash_the_reader_writes() -> None: - """A backslash escape is only safe if python-markdown removes it again. - - Measured on 3.10.3 with the four extensions: ``\\``, backtick, ``*``, - ``_``, ``{``, ``}``, ``[``, ``]``, ``(``, ``)``, ``>``, ``#``, ``+``, - ``-``, ``.``, ``!`` and ``|``. ``=`` and ``~`` are not among them, which - is why those two are spelled as character references instead. - """ - - escapable = set(Markdown(extensions=conversion._MARKDOWN_EXTENSIONS).ESCAPED_CHARS) - backslashed = set(conversion._LITERAL) - set(conversion._SPELLED_OUT) | {"\\"} - - assert backslashed <= escapable - assert not {"=", "~", "<", "&"} & escapable - - -def test_a_body_holding_the_private_stand_ins_converts_like_any_other() -> None: - """Literal text is carried in private-use characters until it is spelled. - - A body that already holds one of them -- an icon font maps symbols - there -- must not be confused with them, so the converter picks a block - the body does not use. - """ - - first_block = [chr(0xF0000 + offset) for offset in range(32)] - html = "<p>" + "".join(first_block) + " __init__ et #4521</p>" - - markdown = read(html) - - assert "".join(first_block) in markdown - assert markdown.endswith("\\_\\_init\\_\\_ et #4521") diff --git a/glpi_python_client/content/tests/test_round_trip.py b/glpi_python_client/content/tests/test_round_trip.py new file mode 100644 index 0000000..9e573d0 --- /dev/null +++ b/glpi_python_client/content/tests/test_round_trip.py @@ -0,0 +1,350 @@ +"""Realistic content survives the trip between GLPI's HTML and Markdown, both ways. + +The contract, checked with an HTML parser (:mod:`.display`) rather than by eye: + +* **HTML first.** ``to_transport(from_transport(html))`` displays what + ``html`` displays -- the same blocks, words, and bold/italic/code/link on + each character -- and the Markdown is a fixed point: reading back what it + renders gives the same Markdown. +* **Markdown first.** Markdown a caller writes renders to HTML that reads + back as Markdown displaying the same, and that Markdown is a fixed point. + +The bodies are the shapes GLPI's editor and the mail collector store, and +the Markdown is what an integrator writes. Adversarial input -- syntax +characters packed into every position -- is out of scope: the converter is a +thin layer over markdownify, mdformat and cmark-gfm, and it is held to +realistic content. +""" + +from __future__ import annotations + +import time + +import pytest + +from glpi_python_client.content.conversion import GlpiContentConverter +from glpi_python_client.content.tests.display import displayed + +read = GlpiContentConverter.from_transport +render = GlpiContentConverter.to_transport + + +def assert_survives(html: str) -> str: + """Assert the HTML-first property for one body, and return its Markdown.""" + + markdown = read(html) + rendered = render(markdown) + + assert displayed(rendered) == displayed(html), ( + f"the Markdown does not display what the HTML did\n" + f" html: {html!r}\n markdown: {markdown!r}\n rendered: {rendered!r}" + ) + assert read(rendered) == markdown, ( + f"the Markdown is not a fixed point\n markdown: {markdown!r}\n" + f" again: {read(rendered)!r}" + ) + return markdown + + +# --------------------------------------------------------------------------- +# HTML first +# --------------------------------------------------------------------------- + +#: Ticket bodies shaped the way GLPI's editor and mail collector store them. +REALISTIC = [ + "<p>Bonjour,</p><p>Le PC du poste 12 ne démarre plus depuis ce matin.</p>", + "<p>Bonjour,<br>Le PC ne démarre plus.<br>Cordialement,<br>Jean</p>", + "<p>Merci de <strong>redémarrer</strong> le serveur <em>avant</em> 18h.</p>", + "<p>Étapes :</p><ul><li>ouvrir la session</li><li>lancer Outlook</li></ul>", + "<table><thead><tr><th>Poste</th><th>IP</th></tr></thead>" + "<tbody><tr><td>PC12</td><td>10.0.0.12</td></tr></tbody></table>", + "<h2>Contexte</h2><p>Migration du serveur.</p>", + '<p>Voir <a href="https://example.org/doc">https://example.org/doc</a></p>', + '<p>Voir <a title="doc" href="https://example.org/doc">la doc</a></p>', + '<p><a href="https://example.org/wiki/Test_(informatique)">wiki</a></p>', + "<p>prix 5*3 et note * importante</p>", + "<p>appuyer sur <Entrée> puis valider</p>", + "<p>Service R&D, bâtiment A & B</p>", + "<p>Montant : 12 000 €</p>", + '<p><span style="color: #e03e2d;">URGENT</span> <u>à traiter</u></p>', + '<p>Cordialement</p><p><img src="https://example.org/logo.png" alt="Logo"></p>', + "<p>Bonjour</p><blockquote><p>Message d'origine</p></blockquote>", + '<pre>Traceback (most recent call last):\n File "x.py", line 1\nError</pre>', + "<p>Lancer <code>ipconfig /all</code> puis envoyer.</p>", + "<p>Fichier mon_fichier_final.docx et __init__</p>", + "<p># pas un titre</p><p>- pas une liste</p>", + '<p>Contact <a href="mailto:support@example.org">support@example.org</a></p>', + "<div>Bonjour,</div><div><br></div><div>Le serveur est down.</div>", + "<p>* point un<br>* point deux</p>", + "<p>[INFO] tâche [1] terminée</p>", + "<p>[1]: https://example.org/note</p>", + r"<p>Chemin C:\Users\jdupont\Desktop</p>", + r"<p>Accès au partage \\serveur\compta\2026 refusé</p>", + "<p>Merci 👍</p>", + "<p>Titre<br>=====</p>", + "<p>---</p><p>signature</p>", + "<p>1) un<br>2) deux</p>", + "<ul><li><p>point un</p><p>suite du point</p></li><li>point deux</li></ul>", + "<blockquote><ul><li>a</li><li>b</li></ul></blockquote>", + "<p>Nom : ______ Prénom : ______</p>", + "<p>a | b | c</p>", + "<p>> pas une citation</p>", + "<p>Cordialement,<br>Jean Dupont<br>--<br>Service IT</p>", + "<p>Merci<br>-----------<br>Jean Dupont</p>", + "<p>Le 30/09, Jean a écrit :<br>> merci<br>> cordialement</p>", + "<ul><li>Réseau<ul><li>switch 3</li><li>borne wifi</li></ul></li></ul>", + "<ol><li>Arrêter</li><li>Sauvegarder<ol><li>la base</li><li>les fichiers</li>" + "</ol></li><li>Redémarrer</li></ol>", + '<ol start="3"><li>trois</li><li>quatre</li></ol>', + "<ul><li>Réseau<ul><li>switch 3</li></ul>à vérifier</li></ul>", + "<p><b>Important</b> : voir <i>ci-dessous</i></p>", + "<table><tr><th>Commande</th></tr><tr><td>ps aux | grep java</td></tr></table>", + "<table><tr><th>Commande</th></tr><tr><td><code>ps aux | grep java</code></td></tr>" + "</table>", + "<pre>ligne\n```\nfin</pre>", + "<p>Suite à la mise à jour, <strong>3 postes</strong> ne se connectent plus :" + "</p><ul><li>PC12 (salle 3)</li><li>PC14 – <em>poste d'accueil</em></li>" + '</ul><p>Voir <a href="https://example.org/kb/42">https://example.org/kb/42</a>' + ' et le <a href="https://example.org/kb/43" title="KB 43">KB 43</a>.</p>', +] + + +@pytest.mark.parametrize("html", REALISTIC) +def test_a_realistic_body_displays_the_same_after_the_round_trip(html: str) -> None: + assert_survives(html) + + +#: What e-mail clients add: Outlook's blank paragraphs and line breaks made of +#: a non-breaking space, a blank line inside a paragraph, a label in bold +#: right before a figure, and Gmail's quoted reply. +E_MAIL_SHAPES = [ + pytest.param( + '<p class="MsoNormal">Bonjour,<o:p></o:p></p>' + '<p class="MsoNormal"><o:p> </o:p></p>' + '<p class="MsoNormal">Le serveur répond.<o:p></o:p></p>', + id="outlook-blank-paragraph", + ), + pytest.param("<p>Bonjour<br> <br>Texte</p>", id="nbsp-blank-line"), + pytest.param("<p>Bonjour,<br><br>Texte</p>", id="blank-line-in-a-paragraph"), + pytest.param("<p><b>Total:</b>12 postes</p>", id="bold-label-before-a-figure"), + pytest.param( + '<div dir="ltr">Merci</div><div class="gmail_quote"><div>Le lun. a écrit :' + "</div><blockquote>Le serveur est down.</blockquote></div>", + id="gmail-quote", + ), + pytest.param("<p>Ligne<br></p><p>Suite</p>", id="break-ending-a-paragraph"), + pytest.param( + "<p>Source wrapped\nmid-sentence\nby Outlook.</p>", id="newlines-in-source" + ), +] + + +@pytest.mark.parametrize("html", E_MAIL_SHAPES) +def test_an_e_mail_shape_displays_the_same_after_the_round_trip(html: str) -> None: + assert_survives(html) + + +@pytest.mark.parametrize( + ("html", "expected"), + [ + pytest.param( + "<p>Voir fichier_de_test_v2.xlsx et mon_fichier_final.docx</p>", + "Voir fichier_de_test_v2.xlsx et mon_fichier_final.docx", + id="file-names", + ), + pytest.param( + r"<p>Chemin C:\Temp\logs et C:\Users\Admin\Documents</p>", + r"Chemin C:\Temp\logs et C:\Users\Admin\Documents", + id="windows-paths", + ), + pytest.param( + "<p>Le ticket # 3 et le #4521 sont liés, C# aussi</p>", + "Le ticket # 3 et le #4521 sont liés, C# aussi", + id="hashes", + ), + pytest.param( + "<p>Service R&D, bâtiment A & B</p>", + "Service R&D, bâtiment A & B", + id="ampersands", + ), + pytest.param( + "<p>[INFO] tâche [1] terminée (voir note)</p>", + "[INFO] tâche [1] terminée (voir note)", + id="brackets", + ), + pytest.param( + "<p>Voir https://example.org/doc?a=1&b=2 ou support@example.org</p>", + "Voir https://example.org/doc?a=1&b=2 ou support@example.org", + id="bare-url-and-address", + ), + pytest.param( + "<p>Fichier mon_fichier_final.docx et __init__</p>", + r"Fichier mon_fichier_final.docx et \_\_init\_\_", + id="dunder", + ), + ], +) +def test_ordinary_prose_reads_back_as_typed(html: str, expected: str) -> None: + """Prose carries no escape, and literal syntax is escaped only to stay text.""" + + assert assert_survives(html) == expected + + +def test_a_table_nested_in_a_cell_keeps_its_text() -> None: + """A signature laid out as a table inside a table keeps every word.""" + + html = ( + "<table><tr><th>Signature</th></tr><tr><td><table><tr><td>Jean Dupont</td>" + "<td>Service IT</td></tr></table></td></tr></table>" + ) + + markdown = read(html) + + assert "Jean Dupont" in render(markdown) + assert "Service IT" in render(markdown) + + +@pytest.mark.parametrize( + "html", + [ + pytest.param( + "<style>p.MsoNormal{margin:0cm}</style><p>Bonjour</p>", id="style" + ), + pytest.param("<p>Bonjour</p><script>track()</script>", id="script"), + pytest.param( + "<html><head><title>RE: Imprimante" + "

      Bonjour

      ", + id="title", + ), + ], +) +def test_what_a_browser_does_not_display_is_dropped(html: str) -> None: + assert read(html) == "Bonjour" + + +def test_a_fence_keeps_its_language() -> None: + markdown = "```powershell\nGet-Service\n```" + + assert read(render(markdown)) == markdown + + +# --------------------------------------------------------------------------- +# Plain text and the write path +# --------------------------------------------------------------------------- + + +@pytest.mark.parametrize( + "text", + [ + "__init__ et _x_", + "# pas un titre", + "* point un\n* point deux", + r"\\serveur\partage", + "if x0", + "Bonjour,\n\nMerci.", + ], +) +def test_a_plain_text_body_reads_as_the_text_it_is(text: str) -> None: + """A body with no HTML element is literal text, its lines lines.""" + + markdown = read(text) + shown = displayed(render(markdown)) + expected = displayed( + "

      " + text.replace("<", "<").replace("\n", "
      ") + "

      " + ) + + assert shown == expected + assert read(render(markdown)) == markdown + + +@pytest.mark.parametrize( + "markdown", + [ + "Run **passwd**, then check `logs`.", + "| a | b
      c |\n| --- | --- |", + "Press Ctrl + C.", + "line one\nline two", + ], +) +def test_the_write_path_keeps_caller_markdown_verbatim(markdown: str) -> None: + """The write models' validator never rewrites the caller's Markdown.""" + + assert read(markdown, plain_text_is_markdown=True) == markdown + + +def test_the_write_path_still_converts_an_html_document() -> None: + assert read("

      A bold move

      ", plain_text_is_markdown=True) == ( + "A **bold** move" + ) + + +# --------------------------------------------------------------------------- +# Markdown first +# --------------------------------------------------------------------------- + +#: What an integrator writes into followups, tasks, solutions and articles. +CALLER_MARKDOWN = [ + "The printer is **offline**.", + "Line one\nline two", + "# Procédure\n\n1. Arrêter le service\n2. Vider le cache\n3. Redémarrer", + "- Réseau\n - switch 3\n - borne wifi\n- Imprimante", + "- Réseau\n - switch 3\n- Imprimante", + "1. Sauvegarder\n 1. la base\n 2. les fichiers\n2. Redémarrer", + 'Voir [la procédure](https://example.org/kb/42 "KB 42").', + "Lien direct : ", + "![capture](https://example.org/c.png)", + "Lancer `ipconfig /all` puis envoyer le résultat.", + "```powershell\nGet-Service | Where-Object Status -eq Running\n```", + "> Le 30/09, Jean a écrit :\n> merci", + "| Poste | IP |\n| :--- | ---: |\n| PC12 | 10.0.0.12 |", + "| Commande | Effet |\n| --- | --- |\n| `ps aux \\| grep java` | processus |", + "Fichier `mon_fichier_final.docx` et chemin C:\\Temp\\logs.", + "Contact : support@example.org, R&D, 5 * 3 = 15.", + "Résolu ✅ — merci 👍", + "**Cause :** disque plein.\n\n**Solution :** purge des journaux.", + "Avant :\n\n---\n\nAprès.", + "Étapes :\n- ouvrir la session\n- lancer Outlook", +] + + +@pytest.mark.parametrize("markdown", CALLER_MARKDOWN) +def test_caller_markdown_displays_the_same_after_the_round_trip(markdown: str) -> None: + html = render(markdown) + back = read(html) + + assert displayed(render(back)) == displayed(html) + assert read(render(back)) == back + + +# --------------------------------------------------------------------------- +# Cost and depth +# --------------------------------------------------------------------------- + + +@pytest.mark.parametrize( + "html", + [ + pytest.param("

      " + "ligne
      " * 20_000 + "

      ", id="line-breaks"), + pytest.param("
        " + "
      • x
      • " * 20_000 + "
      ", id="list-items"), + pytest.param("
        " + "
      • " * 20_000 + "
      ", id="empty-items"), + pytest.param("

      " + "gras mot " * 20_000 + "

      ", id="emphasis"), + ], +) +def test_a_long_body_converts_in_linear_time(html: str) -> None: + """The budget is generous on purpose: this guards the complexity, not the speed.""" + + started = time.perf_counter() + read(html) + + assert time.perf_counter() - started < 20 + + +def test_a_body_too_deep_to_convert_keeps_its_text() -> None: + """markdownify recurses per nesting level; past the stack, the text is kept.""" + + html = "
      " * 3000 + "

      __init__ au fond

      " + "
      " * 3000 + + markdown = read(html) + + assert "init" in markdown + assert "au fond" in render(markdown) diff --git a/glpi_python_client/models/api_schema/_content.py b/glpi_python_client/models/api_schema/_content.py index 13fd521..50d0a5c 100644 --- a/glpi_python_client/models/api_schema/_content.py +++ b/glpi_python_client/models/api_schema/_content.py @@ -44,19 +44,18 @@ Pydantic resolves them, and would otherwise divert the wire's ``content`` into ``extra_payload``. -Plain-text content is preserved verbatim on the inbound path and rendered -as HTML paragraphs on the outbound path, matching the converter's default -behaviour. "Plain text" means text carrying no recognised HTML element: -``use the key`` and ``if x0`` are text, because ``Enter`` -and ``y`` are not elements, while ``ac`` is treated as markup because -``b`` is. ``None`` values are passed through unchanged so optional fields -and ``exclude_none`` semantics keep working. +On the read path a plain-text body -- one carrying no recognised HTML +element, so ``use the key`` and ``if x0`` are text while +``ac`` is markup -- is literal text, read as GLPI displays it. +``None`` values are passed through unchanged so optional fields and +``exclude_none`` semantics keep working. Note that the inbound converter also runs on **outbound** content: the ``BeforeValidator`` below fires when a caller constructs a ``Post*`` model, so caller-authored Markdown passes through it before the serializer renders -it. That is why the plain-text path has to stay verbatim -- routing Markdown -through the HTML normaliser escapes it, and GLPI receives literal asterisks. +it. It runs with ``plain_text_is_markdown=True``, which keeps the caller's +Markdown verbatim unless it starts with an HTML tag -- reading it as literal +text would escape it, and GLPI would receive literal asterisks. One sharp edge comes with the read side, from ``functools.cached_property``: assigning to ``content_html`` after ``content`` has been read leaves the @@ -110,7 +109,7 @@ def _from_transport(value: object) -> str | None: if value is None: return None - return GlpiContentConverter.from_transport(value) + return GlpiContentConverter.from_transport(value, plain_text_is_markdown=True) def markdown_view(raw: str | None) -> str | None: diff --git a/glpi_python_client/models/api_schema/tests/test_content.py b/glpi_python_client/models/api_schema/tests/test_content.py index 00cc8c1..52892e9 100644 --- a/glpi_python_client/models/api_schema/tests/test_content.py +++ b/glpi_python_client/models/api_schema/tests/test_content.py @@ -19,6 +19,7 @@ from glpi_python_client import GlpiContentError, GlpiError, GlpiValidationError from glpi_python_client._sync._testing import TransportRecorder, make_client from glpi_python_client._sync.clients.commons._payloads import model_to_payload +from glpi_python_client.content import conversion from glpi_python_client.models.api_schema._content import ( _from_transport, _to_transport, @@ -79,7 +80,9 @@ def test_a_read_model_reads_none_when_glpi_sent_no_body() -> None: assert ticket.content is None -def test_a_content_fault_on_the_write_path_stays_in_the_taxonomy() -> None: +def test_a_content_fault_on_the_write_path_stays_in_the_taxonomy( + monkeypatch: pytest.MonkeyPatch, +) -> None: """pydantic-core destroys a serializer's exception; this puts it back. Outbound conversion runs in a ``PlainSerializer``, and everything it @@ -88,19 +91,25 @@ def test_a_content_fault_on_the_write_path_stays_in_the_taxonomy() -> None: ``__context__`` both ``None``. Every ``create_*``/``update_*`` carrying a body was affected, so ``except GlpiError`` did not fire on the one path the package's error contract is most explicit about. + + cmark-gfm renders without recursing, so no Markdown makes it fail; the + fault is injected. """ - deep_markdown = "".join(" " * (4 * level) + "- x\n" for level in range(600)) + def failing_renderer(markdown: str) -> str: + raise RuntimeError("renderer fault") + + monkeypatch.setattr(conversion, "markdown_to_html", failing_renderer) client = make_client() TransportRecorder().install(client) with pytest.raises(GlpiContentError) as caught: - client.create_ticket(PostTicket(name="round trip", content=deep_markdown)) + client.create_ticket(PostTicket(name="round trip", content="**x**")) assert isinstance(caught.value, GlpiError) assert not isinstance(caught.value, ValueError) # the original fault, which pydantic-core had discarded - assert isinstance(caught.value.__cause__, RecursionError) + assert isinstance(caught.value.__cause__, RuntimeError) # and the serializer wrapper, kept where a reader would look for it assert isinstance(caught.value.__context__, PydanticSerializationError) diff --git a/glpi_python_client/models/custom_schema/_ticket_context.py b/glpi_python_client/models/custom_schema/_ticket_context.py index d645f4f..08c1bfa 100644 --- a/glpi_python_client/models/custom_schema/_ticket_context.py +++ b/glpi_python_client/models/custom_schema/_ticket_context.py @@ -15,6 +15,7 @@ from pydantic import Field +from glpi_python_client.content.conversion import GlpiContentConverter from glpi_python_client.models._base import GlpiModel from glpi_python_client.models.api_schema._common import IdNameRef from glpi_python_client.models.api_schema.assistance._ticket import GetTicket @@ -32,6 +33,17 @@ _MAX_DATETIME = datetime.max.replace(tzinfo=timezone.utc) +def _literal(text: str) -> str: + """Spell a name or a file name as one line of Markdown that shows it as typed. + + User data is literal text, as a body's text is: ``__init__`` must not turn + bold, nor a file name starting ``1.`` start a list. The converter reads + plain text that way already. + """ + + return GlpiContentConverter.from_transport(" ".join(text.split())) + + def _ref_label(ref: IdNameRef | None) -> str | None: """Return the human-readable label of one ``IdNameRef`` reference. @@ -83,7 +95,7 @@ def _subtitle_line(*parts: tuple[str, object | None]) -> str | None: for label, value in parts: rendered_value = _render_value(value) if rendered_value: - rendered_parts.append(f"{label}: {rendered_value}") + rendered_parts.append(f"{label}: {_literal(rendered_value)}") if not rendered_parts: return None return f"> {' | '.join(rendered_parts)}" @@ -255,7 +267,7 @@ def to_markdown( lines: list[str] = [] ticket = self.ticket - ticket_label = ticket.name or "(unnamed ticket)" + ticket_label = _literal(ticket.name) if ticket.name else "(unnamed ticket)" if ticket.id is not None: lines.append(f"# Ticket #{ticket.id} \u2014 {ticket_label}") else: @@ -364,7 +376,7 @@ def to_markdown( else "document" ) ) - lines.append(f"- {label}") + lines.append(f"- {_literal(label)}") return "\n".join(lines).rstrip() diff --git a/glpi_python_client/models/custom_schema/tests/test_ticket_context.py b/glpi_python_client/models/custom_schema/tests/test_ticket_context.py index a9ad98a..95dc23f 100644 --- a/glpi_python_client/models/custom_schema/tests/test_ticket_context.py +++ b/glpi_python_client/models/custom_schema/tests/test_ticket_context.py @@ -454,3 +454,42 @@ def test_to_markdown_orders_events_across_mixed_datetime_awareness() -> None: rendered = context.to_markdown() assert rendered.index("naive first") < rendered.index("aware second") + + +def test_to_markdown_shows_names_and_file_names_as_typed() -> None: + """A ticket name, a user name or a file name is literal text, like a body. + + ``__init__`` in a name used to turn bold and a file name's ``*final*`` + italic, and a file name starting ``1.`` became a numbered list. + """ + + from bs4 import BeautifulSoup + + from glpi_python_client.content import GlpiContentConverter + + context = GlpiTicketContext.model_validate( + { + "ticket": { + "id": 7, + "name": "__init__ échoue", + "content": "

      corps

      ", + "user_recipient": {"id": 3, "name": "*admin* [ext]"}, + }, + "documents": [ + {"id": 1, "filename": "rapport_*final*_v2.pdf"}, + {"id": 2, "filename": "1. lisez-moi.txt"}, + ], + } + ) + + html = GlpiContentConverter.to_transport(context.to_markdown()) + soup = BeautifulSoup(html, "html.parser") + + h1 = soup.find("h1") + assert h1 is not None + assert h1.get_text() == "Ticket #7 \u2014 __init__ échoue" + quote = soup.find("blockquote") + assert quote is not None + assert "Requester: *admin* [ext]" in quote.get_text() + items = [item.get_text() for item in soup.find_all("li")] + assert items == ["rapport_*final*_v2.pdf", "1. lisez-moi.txt"] diff --git a/glpi_python_client/testing/tests/test_content_roundtrip.py b/glpi_python_client/testing/tests/test_content_roundtrip.py index fbe6801..a2ce94b 100644 --- a/glpi_python_client/testing/tests/test_content_roundtrip.py +++ b/glpi_python_client/testing/tests/test_content_roundtrip.py @@ -16,6 +16,7 @@ from glpi_python_client._sync.clients.commons._payloads import model_to_payload from glpi_python_client.content import conversion +from glpi_python_client.content.tests.display import displayed from glpi_python_client.models._base import GlpiModel from glpi_python_client.models.api_schema.assistance import ( GetTicket, @@ -46,9 +47,9 @@ def _count_conversions(monkeypatch: pytest.MonkeyPatch) -> list[str]: seen: list[str] = [] real = conversion.GlpiContentConverter.from_transport - def _record(value: object) -> str: + def _record(value: object, **options: bool) -> str: seen.append(str(value)) - return real(value) + return real(value, **options) monkeypatch.setattr( conversion.GlpiContentConverter, "from_transport", staticmethod(_record) @@ -135,16 +136,15 @@ def test_outgoing_empty_string_renders_empty() -> None: # Round-trip corpus # --------------------------------------------------------------------------- # -# ``from_transport(to_transport(m)) == m`` is the property the content layer -# would like to hold. It does not hold universally, and cannot: the two -# libraries either side of the wire disagree about a handful of constructs, -# and no option on either fixes them. +# Markdown written by a caller goes to GLPI as HTML and reads back as +# Markdown. It reads back in the converter's canonical spelling -- a line +# break as a backslash, a nested list indented by its bullet's width, a table +# unpadded -- which displays the same and reads back as itself. Each entry +# names that spelling when it differs from the caller's, and the test asserts +# both: the canonical Markdown, and the same display as what was sent. # -# So the corpus is an inventory rather than a property test. Every case is -# listed, the lossy ones carry ``xfail(strict=True)``, and that strictness is -# the point -- fixing one of them turns its xfail into an XPASS and fails the -# suite, forcing the inventory to be updated rather than quietly drifting out -# of date. A regression in a passing case fails immediately. +# A loss carries ``xfail(strict=True)``, so fixing it fails the suite until +# the inventory is updated. def _lossy(reason: str) -> pytest.MarkDecorator: @@ -154,69 +154,77 @@ def _lossy(reason: str) -> pytest.MarkDecorator: ROUND_TRIP_CORPUS = [ - pytest.param("The printer is offline.", id="plain"), - pytest.param("The printer is **offline**.", id="bold"), - pytest.param("This is *emphasis*.", id="italic"), - pytest.param("Run `systemctl restart` now.", id="inline-code"), - pytest.param("# Title\n\nBody text.", id="heading"), - pytest.param("## Section\n\nBody text.", id="subheading"), - pytest.param("First para.\n\nSecond para.", id="paragraphs"), - pytest.param("line one \nline two", id="hard-break"), - pytest.param("- alpha\n- beta\n- gamma", id="bullets"), - pytest.param("1. one\n2. two", id="numbered"), - pytest.param("> quoted text", id="blockquote"), - pytest.param("See [the doc](https://example.test/doc).", id="link"), - pytest.param("```\nx = 1\n```", id="fence"), - pytest.param("| a | b |\n| --- | --- |\n| 1 | 2 |", id="table"), - pytest.param("The snake_case name.", id="underscore"), - pytest.param("5 * 3 = 15", id="asterisk"), - pytest.param("# Title\n\n- alpha\n- beta\n\nClosing **note**.", id="mixed"), - pytest.param("- alpha\n - inner\n- beta", id="nested-list"), - pytest.param("1. one\n 1. inner\n2. two", id="nested-numbered-list"), - pytest.param("intro\n\n3. three\n4. four", id="numbered-from-three"), - # Literal text the reader escapes reads back escaped the same way, so a - # body that carries it is a fixed point too. - pytest.param(r"\#4521: module \_\_init\_\_.", id="escaped-literals"), - pytest.param(r"Share \\\server\share and C:\Temp.", id="backslashes"), - pytest.param(r"Press <Enter> and see \*x\*.", id="escaped-markup"), - pytest.param("| cmd |\n| --- |\n| ps aux \\| grep java |", id="pipe-in-a-cell"), + pytest.param("The printer is offline.", None, id="plain"), + pytest.param("The printer is **offline**.", None, id="bold"), + pytest.param("This is *emphasis*.", None, id="italic"), + pytest.param("Run `systemctl restart` now.", None, id="inline-code"), + pytest.param("# Title\n\nBody text.", None, id="heading"), + pytest.param("## Section\n\nBody text.", None, id="subheading"), + pytest.param("First para.\n\nSecond para.", None, id="paragraphs"), + pytest.param("line one \nline two", "line one\\\nline two", id="hard-break"), + pytest.param("line one\nline two", "line one\\\nline two", id="soft-newline"), + pytest.param("- alpha\n- beta\n- gamma", None, id="bullets"), + pytest.param("1. one\n2. two", None, id="numbered"), + pytest.param("> quoted text", None, id="blockquote"), + pytest.param("See [the doc](https://example.test/doc).", None, id="link"), + pytest.param("```\nx = 1\n```", None, id="fence"), + pytest.param("```python\nx = 1\n```", None, id="fence-with-language"), pytest.param( - "line one\nline two", - id="soft-newline", - marks=_lossy( - "nl2br renders a lone newline as
      , which markdownify reads " - "back as a hard break (two trailing spaces). Semantically " - "equivalent and stable after one cycle; see issue #32." - ), + "| a | b |\n| --- | --- |\n| 1 | 2 |", + "| a | b |\n| -- | -- |\n| 1 | 2 |", + id="table", ), + pytest.param("The snake_case name.", None, id="underscore"), + pytest.param("5 * 3 = 15", None, id="asterisk"), + pytest.param("# Title\n\n- alpha\n- beta\n\nClosing **note**.", None, id="mixed"), pytest.param( - "```python\nx = 1\n```", - id="fence-with-language", - marks=_lossy( - "fenced_code emits class='language-python' and markdownify drops " - "the class, so the language tag cannot survive." - ), + "- alpha\n - inner\n- beta", "- alpha\n - inner\n- beta", id="nested-list" + ), + pytest.param( + "1. one\n 1. inner\n2. two", + "1. one\n 1. inner\n2. two", + id="nested-numbered-list", + ), + pytest.param("intro\n\n3. three\n4. four", None, id="numbered-from-three"), + pytest.param( + r"\#4521: module \_\_init\_\_.", + r"#4521: module \_\_init\_\_.", + id="escaped-literals", + ), + pytest.param(r"Share \\\server\share and C:\Temp.", None, id="backslashes"), + pytest.param( + r"Press <Enter> and see \*x\*.", + r"Press \ and see \*x\*.", + id="escaped-markup", + ), + pytest.param( + "| cmd |\n| --- |\n| ps aux \\| grep java |", + "| cmd |\n| -- |\n| ps aux \\| grep java |", + id="pipe-in-a-cell", ), pytest.param( "use the key", + None, id="angle-bracket-text", marks=_lossy( - "to_transport does not escape raw markup, so the text reaches " - "GLPI as a live unknown tag -- which the web UI drops too. " - "Escaping it is a separate change to the outbound direction." + "to_transport passes raw markup through, so the text reaches GLPI " + "as a live unknown tag, which the web UI drops too. Write it in " + "backticks." ), ), ] -@pytest.mark.parametrize("markdown", ROUND_TRIP_CORPUS) -def test_round_trip_corpus(markdown: str) -> None: +@pytest.mark.parametrize(("markdown", "canonical"), ROUND_TRIP_CORPUS) +def test_round_trip_corpus(markdown: str, canonical: str | None) -> None: """Markdown survives a full write-then-read cycle through GLPI's HTML.""" outgoing = model_to_payload(PostTicket(name="Round trip", content=markdown)) incoming = GetTicket.model_validate({"name": "Round trip", **outgoing}) - assert incoming.content == markdown + assert incoming.content == (canonical or markdown) + rendered = conversion.GlpiContentConverter.to_transport(incoming.content) + assert displayed(rendered) == displayed(outgoing["content"]) # --------------------------------------------------------------------------- diff --git a/pyproject.toml b/pyproject.toml index f09d51d..eb89aec 100644 --- a/pyproject.toml +++ b/pyproject.toml @@ -58,11 +58,18 @@ classifiers = [ "Typing :: Typed", ] dependencies = [ - "beautifulsoup4>=4.12", + # 4.15 fixed a parser defect that dropped the text after a "
      " in a + # body that also held a bare "
      ". + "beautifulsoup4>=4.15", + # Renders Markdown to HTML: CommonMark with GFM tables. + "cmarkgfm>=2025.10", "httpx>=0.28", "lxml>=4.9", - "markdown>=3.6", + "markdown-it-py>=3.0", "markdownify>=1.2", + # The reader extends mdformat's renderer, an API that may move at 0.8. + "mdformat>=0.7.22,<0.8", + "mdformat-tables>=1.0", "pydantic>=2.8", # Never imported by this package, and deliberately so. httpcore decides # whether it is running under asyncio or trio by probing for sniffio on diff --git a/skills/glpi-asset-workflow/SKILL.md b/skills/glpi-asset-workflow/SKILL.md index a88b819..d830a20 100644 --- a/skills/glpi-asset-workflow/SKILL.md +++ b/skills/glpi-asset-workflow/SKILL.md @@ -5,7 +5,7 @@ license: MIT compatibility: "Requires Python 3.10+, glpi-python-client, network access to the GLPI v2 API, and credentials allowed to read or write assets." metadata: package: glpi-python-client - version: "0.5.0" + version: "0.6.0" --- # GLPI Asset Workflow diff --git a/skills/glpi-contract-workflow/SKILL.md b/skills/glpi-contract-workflow/SKILL.md index 6f0ff9c..0a55c56 100644 --- a/skills/glpi-contract-workflow/SKILL.md +++ b/skills/glpi-contract-workflow/SKILL.md @@ -5,7 +5,7 @@ license: MIT compatibility: "Requires Python 3.10+, glpi-python-client, network access to the GLPI v2 API, and credentials allowed to read or write contracts." metadata: package: glpi-python-client - version: "0.5.0" + version: "0.6.0" --- # GLPI Contract Workflow diff --git a/skills/glpi-ticket-timeline/SKILL.md b/skills/glpi-ticket-timeline/SKILL.md index f779e79..cd18a59 100644 --- a/skills/glpi-ticket-timeline/SKILL.md +++ b/skills/glpi-ticket-timeline/SKILL.md @@ -120,7 +120,7 @@ await client.update_ticket_timeline_document( - `create_*` methods return new identifiers as plain `int`. `update_*` and `delete_*`/`unlink_*` return `None`. - Three enums carry this family's value vocabularies, all exported from `glpi_python_client` and all subclasses of `GlpiEnum` (itself an `IntEnum`, so a member serialises as its number and compares equal to one): `GlpiTaskState` on `PostTicketTask.state`/`PatchTicketTask.state` (`INFORMATION = 0`, `TODO = 1`, `DONE = 2` -- note `INFORMATION` is `0`, so `if task.state:` is false for it; test against `None`), `GlpiSolutionStatus` on `PostSolution.status`/`PatchSolution.status` (`NONE = 1`, `WAITING = 2`, `ACCEPTED = 3`, `REFUSED = 4`), and `GlpiTimelinePosition` on `timeline_position` (`INVALID = -1`, `NONE = 0`, `LEFT = 1`, `RIGHT = 2`, `LEFT_BIG = 3`, `RIGHT_BIG = 4`), which the followup, task and document models carry -- the solution models do not have the field at all. - Timeline `content` fields are Markdown on the Python side, not HTML. `PostFollowup`/`PostTicketTask`/`PostSolution` render Markdown to GLPI's HTML on serialisation. On `GetFollowup`/`GetTicketTask`/`GetSolution` the field is `content_html` -- the server's HTML verbatim -- and `content` is a cached property that converts it on **first read**, not on validation (`GetFollowup(content="

      Hello world

      ").content == "Hello **world**"` still holds; the wire spelling `content` is accepted as a validation alias). So `record.content` is always Markdown, but `list_ticket_followups` converts nothing until you read a body, and a body that cannot be converted no longer breaks the whole list. Authoring raw HTML on a write model is not an error but is round-tripped through the Markdown converter and can be reshaped; write Markdown. -- **`.content` spells literal text so it stays text.** A character is escaped exactly where python-markdown (with `nl2br`, `sane_lists`, `fenced_code`, `tables`) would read it as syntax: a user's `__init__` reads back as `\_\_init\_\_`, `\\serveur` as `\\\serveur`, `#4521` at a line start as `\#4521`, `` as `<Entrée>`. Ordinary prose -- `fichier_de_test_v2.xlsx`, `C:\Temp`, `R&D` -- carries no escape. Render the Markdown to display it; do not strip the backslashes, and do not mix Markdown and HTML in one value you write: one real HTML element makes the whole value HTML, so its `**bold**` is kept as literal asterisks. -- HTML too deeply nested to walk is **stripped to text instead of converted**: `markdownify` recurses per level and dies around 494 from a shallow stack. The conversion is attempted rather than the depth predicted, so the real limit is whatever stack is left at the call site. Every character the normal rendering would have produced still appears, but structure does not: link targets, image alt text and code fencing are gone. Nothing raises. Any other conversion failure raises `GlpiContentError` -- from the `.content` read, not from the `list_*` call. Note that `GlpiTicketContext.to_markdown()` is usually the first thing to read every body, so it is where such an error surfaces. +- **`.content` is CommonMark that spells literal text so it stays text.** Rendering it (the package uses cmark-gfm, with GFM tables and a newline as a line break) displays what GLPI displayed: a user's `__init__` reads back as `\_\_init\_\_`, `\\serveur` as `\\\serveur`, `# titre` at a line start as `\# titre`. Ordinary prose -- `fichier_de_test_v2.xlsx`, `C:\Temp`, `R&D` -- comes back as typed. A line break reads back as `\` and a newline. Render the Markdown to display it; do not strip the backslashes. Markdown you write is kept verbatim unless it starts with an HTML tag; put a placeholder such as `` in backticks, since raw HTML passes through to GLPI. +- HTML too deeply nested to walk is **stripped to text instead of converted**: `markdownify` recurses per level and runs out of stack a few hundred levels deep. The conversion is attempted rather than the depth predicted, so the real limit is whatever stack is left at the call site. Every character the normal rendering would have produced still appears, but structure does not: link targets, image alt text and code fencing are gone. Nothing raises. Any other conversion failure raises `GlpiContentError` -- from the `.content` read, not from the `list_*` call. Note that `GlpiTicketContext.to_markdown()` is usually the first thing to read every body, so it is where such an error surfaces. - Extra server fields (e.g. plugin keys) flow into `record.extra_payload` rather than raising. - `delete_ticket_*` and `unlink_ticket_timeline_document` accept a keyword-only `force` parameter; pass `force=True` to permanently delete. \ No newline at end of file diff --git a/skills/glpi-ticket-workflow/SKILL.md b/skills/glpi-ticket-workflow/SKILL.md index 29502ba..cfb3944 100644 --- a/skills/glpi-ticket-workflow/SKILL.md +++ b/skills/glpi-ticket-workflow/SKILL.md @@ -74,8 +74,8 @@ ticket = PostTicket( - **`search_tickets` and every other `search_*` raise `GlpiStatusError` on a 4xx.** This changed: they used to check the response status only when the caller passed a `failure_message`, which none of the seven `search_*` helpers does, so a GLPI error body was coerced to `[]` and a malformed RSQL filter, a 403, a missing route and a genuinely empty result set were indistinguishable. `_resource_list` now checks the status on every call, so **an empty list means the server said the result set is empty**. The iterators inherit that: a 4xx raises instead of making the first page short and ending the walk silently. Note the *other* fail-open path is unchanged and still bites -- GLPI v2 ignores a filter field it does not recognise and answers 200 with the whole unfiltered table, so a filter that returns rows is still not proof it was applied. - `create_ticket` returns the new ticket ID. `update_ticket` and `delete_ticket` return `None`. - **`GetTicket` has no `content` field; it has `content_html` and a `content` property.** `content_html` holds the server's HTML verbatim (and accepts the wire spelling `content` on construction); `content` is a cached property that converts it to Markdown on **first read**, not on validation. Reading `ticket.content` is unchanged and still gives Markdown, so read-side code needs no edit — but `search_tickets` now converts nothing until a body is read, and a body that cannot be converted no longer breaks the rest of the page. Two things do change if you were relying on them: `"content" not in GetTicket.model_fields`, and `GetTicket(...).model_dump()` emits `content_html` with HTML where it used to emit `content` with Markdown (`by_alias=True` gives you the GLPI key). `PostTicket`/`PatchTicket` are untouched: plain `content` field, converted eagerly. -- **`ticket.content` spells literal text so it stays text.** A character is escaped exactly where python-markdown (with `nl2br`, `sane_lists`, `fenced_code`, `tables`) would read it as syntax — a user's `__init__` reads back as `\_\_init\_\_`, `#4521` at a line start as `\#4521`, `` as `<Entrée>` — and nowhere else, so `fichier_de_test_v2.xlsx` or `R&D` come back as typed. Render the Markdown to display it; do not strip the backslashes. One real HTML element in a value you write makes the whole value HTML, so write Markdown or HTML, not both. -- HTML too deeply nested to walk is **stripped to text rather than converted** — `markdownify` recurses about twice per nesting level and dies around 494 from a shallow stack. The converter attempts the conversion and answers the `RecursionError` rather than predicting the depth, so the real limit is whatever stack is left at the call site, and the same body can convert from one and degrade from a deeper one. It degrades rather than truncating — every character the normal rendering would have produced still appears — but structure does not survive: link targets, image alt text, code fencing and `
      ` indentation are gone. Anything else that goes wrong converting a body raises `GlpiContentError`, from the attribute read rather than from `get_ticket`. Treat a fetched ticket as immutable afterwards: the conversion is cached, so assigning to `content_html` (or `model_copy(update={"content_html": ...})`) leaves stale Markdown on `.content` with nothing in `repr`, `==` or `model_dump` to show it.
      +- **`ticket.content` is CommonMark that spells literal text so it stays text.** Rendering it (the package uses cmark-gfm, with GFM tables and a newline as a line break) displays what GLPI displayed: a user's `__init__` reads back as `\_\_init\_\_`, `# titre` at a line start as `\# titre`, while `fichier_de_test_v2.xlsx` or `R&D` come back as typed. A plain-text body is read as the text it is. Render the Markdown to display it; do not strip the backslashes. `PostTicket`/`PatchTicket` keep your Markdown verbatim unless it starts with an HTML tag; put a placeholder such as `` in backticks, since raw HTML passes through to GLPI.
      +- HTML too deeply nested to walk is **stripped to text rather than converted** — `markdownify` recurses about twice per nesting level and runs out of stack a few hundred levels deep. The converter attempts the conversion and answers the `RecursionError` rather than predicting the depth, so the real limit is whatever stack is left at the call site, and the same body can convert from one and degrade from a deeper one. It degrades rather than truncating — every character the normal rendering would have produced still appears — but structure does not survive: link targets, image alt text, code fencing and `
      ` indentation are gone. Anything else that goes wrong converting a body raises `GlpiContentError`, from the attribute read rather than from `get_ticket`. Treat a fetched ticket as immutable afterwards: the conversion is cached, so assigning to `content_html` (or `model_copy(update={"content_html": ...})`) leaves stale Markdown on `.content` with nothing in `repr`, `==` or `model_dump` to show it.
       - The GLPI server is the authoritative validator. Extra keys returned by the server flow into `ticket.extra_payload` rather than raising. Caller-provided `extra_payload` keys win on conflicts.
       - **`status` is writable, on `PatchTicket` only.** `client.update_ticket(tid, PatchTicket(status=GlpiTicketStatus.PENDING))` moves the ticket. This corrects an earlier claim in this skill that the field was read-only: GLPI's own contract publishes `Ticket.status.id` as `readOnly: true`, and that is simply wrong — a live GLPI 11 instance honours the `PATCH`. The same contract omits the `Major` level from `priority`, so treat it as a hint, not an authority. Two things it does *not* cover: `POST` **ignores** `status` (201, then the ticket reads back as `New`), which is why the field is declared on `PatchTicket` and not on `PostTicket`; and `status_id` is silently dropped, so use `status`. Posting a solution still moves the ticket to `SOLVED` on its own and remains the better call when there is a resolution to record — a bare status change leaves no trace of why.
       - **Key on `status.id`, never on `status.name`.** The id set is hardcoded in GLPI's core — there is no `TicketStatus` itemtype (the v1 API answers 400 for it, while `ITILCategory` and `RequestType` return instance rows), and `listSearchOptions` reports `status` as `{"table": "glpi_tickets", "datatype": "specific"}`, i.e. a column of the tickets table rather than a foreign key into a configurable one. An administrator cannot add or rename a status. The **labels** are another matter: they are translations, so the same id reads back as `"Nouveau"` on a French instance and `"New"` on an English one. The ids `7`, `8`, `9`, `11`–`14` exist but belong to the sibling ITIL types, and they are literal integers in GLPI's source: `CommonITILObject` declares `INCOMING = 1` through `OBSERVED = 8` plus `APPROVAL = 10`, and `Change` adds `EVALUATION = 9`, `TEST = 11`, `QUALIFICATION = 12`, `REFUSED = 13`, `CANCELED = 14`. Each type's `getAllStatusArray()` is a hardcoded PHP array mapping those constants to *translated* labels — no database, no config table. The set is therefore coupled to the GLPI *version*, not to the instance.
      
      From ae9b4c4d2958b38f72f427e394f0156fa5f4022e Mon Sep 17 00:00:00 2001
      From: baraline 
      Date: Thu, 1 Oct 2026 20:22:41 +0200
      Subject: [PATCH 6/7] Fix conversion testing
      
      ---
       .../content/tests/test_conversion.py          | 35 ++++++++++++++++---
       1 file changed, 31 insertions(+), 4 deletions(-)
      
      diff --git a/glpi_python_client/content/tests/test_conversion.py b/glpi_python_client/content/tests/test_conversion.py
      index 3a1beb5..85cc58e 100644
      --- a/glpi_python_client/content/tests/test_conversion.py
      +++ b/glpi_python_client/content/tests/test_conversion.py
      @@ -7,6 +7,7 @@
       from __future__ import annotations
       
       import pytest
      +from bs4 import BeautifulSoup, ParserRejectedMarkup
       
       from glpi_python_client import GlpiContentError, GlpiError
       from glpi_python_client.content import conversion
      @@ -81,10 +82,36 @@ def test_a_body_too_deep_for_the_stack_degrades_to_its_text(html: str) -> None:
           assert "deep" in read(html)
       
       
      -@pytest.mark.parametrize(
      -    "html", ["

      a

      b

      ", "

      a

      b

      "] -) -def test_markup_the_parser_rejects_degrades_to_its_text(html: str) -> None: +_MALFORMED_DECLARATIONS = ["

      a

      b

      ", "

      a

      b

      "] + + +@pytest.mark.parametrize("html", _MALFORMED_DECLARATIONS) +def test_a_malformed_declaration_still_reads_both_sides(html: str) -> None: + """Older CPython patch releases reject these; newer ones parse them. + + 3.12.11 raises from ``html.parser``, 3.12.14 and 3.13.14 read a + comment, so only what both outcomes share is pinned here. + """ + + markdown = read(html) + assert markdown.index("a") < markdown.rindex("b") + + +@pytest.mark.parametrize("html", _MALFORMED_DECLARATIONS) +def test_markup_the_parser_rejects_degrades_to_its_text( + monkeypatch: pytest.MonkeyPatch, html: str +) -> None: + """Rejection is forced, so the fallback runs whatever ``html.parser`` does.""" + + parse = conversion._soup + + def rejecting(markup: str) -> BeautifulSoup: + if " Date: Fri, 2 Oct 2026 12:04:58 +0200 Subject: [PATCH 7/7] Drop 3.10, UTC tz --- .github/workflows/ci.yml | 2 +- .github/workflows/release.yml | 7 ++--- CHANGELOG.md | 4 ++- CONTRIBUTING.md | 2 +- README.md | 2 +- docs/conf.py | 6 +--- docs/installation.rst | 2 +- docs/publishing_rtd.rst | 2 +- glpi_python_client/_async/auth/_v1_session.py | 6 ++-- glpi_python_client/_async/auth/auth.py | 6 ++-- .../_async/auth/tests/test_auth.py | 12 ++++---- .../_async/clients/_base_client.py | 8 +----- .../api/management/tests/test_document.py | 4 +-- glpi_python_client/_async/clients/client.py | 7 +---- .../clients/commons/tests/test_payloads.py | 6 ++-- .../_async/tests/test_concurrency.py | 4 +-- glpi_python_client/_sync/auth/_v1_session.py | 6 ++-- glpi_python_client/_sync/auth/auth.py | 6 ++-- .../_sync/auth/tests/test_auth.py | 12 ++++---- .../_sync/clients/_base_client.py | 8 +----- .../api/management/tests/test_document.py | 4 +-- glpi_python_client/_sync/clients/client.py | 7 +---- .../clients/commons/tests/test_payloads.py | 6 ++-- .../models/custom_schema/_ticket_context.py | 6 ++-- .../tests/test_ticket_context.py | 28 +++++++++---------- glpi_python_client/models/tests/test_base.py | 8 +++--- .../testing/tests/test_method_invocation.py | 4 +-- .../testing/tests/test_packaging.py | 7 +---- .../testing/tests/test_version_agreement.py | 7 +---- glpi_python_client/tests/test_rsql.py | 10 +++---- pyproject.toml | 10 ++----- skills/glpi-asset-workflow/SKILL.md | 2 +- skills/glpi-client-setup/SKILL.md | 2 +- skills/glpi-contract-workflow/SKILL.md | 2 +- skills/glpi-document-workflow/SKILL.md | 2 +- skills/glpi-knowledge-base/SKILL.md | 2 +- skills/glpi-plugin-fields/SKILL.md | 2 +- skills/glpi-reporting-and-context/SKILL.md | 2 +- skills/glpi-team-members/SKILL.md | 2 +- skills/glpi-ticket-timeline/SKILL.md | 2 +- skills/glpi-ticket-workflow/SKILL.md | 2 +- .../glpi-user-location-provisioning/SKILL.md | 2 +- 42 files changed, 95 insertions(+), 136 deletions(-) diff --git a/.github/workflows/ci.yml b/.github/workflows/ci.yml index d6c2828..87d61d8 100644 --- a/.github/workflows/ci.yml +++ b/.github/workflows/ci.yml @@ -21,7 +21,7 @@ jobs: strategy: fail-fast: false matrix: - python-version: ["3.10", "3.11", "3.12", "3.13", "3.14"] + python-version: ["3.11", "3.12", "3.13", "3.14"] steps: - name: Check out repository diff --git a/.github/workflows/release.yml b/.github/workflows/release.yml index c49ddf3..fcd63a3 100644 --- a/.github/workflows/release.yml +++ b/.github/workflows/release.yml @@ -20,7 +20,7 @@ jobs: strategy: fail-fast: false matrix: - python-version: ["3.10", "3.11", "3.12", "3.13", "3.14"] + python-version: ["3.11", "3.12", "3.13", "3.14"] steps: - name: Check out repository @@ -100,10 +100,7 @@ jobs: normalized_tag="${release_tag#v}" pyproject_version="$(python - <<'PY' from pathlib import Path - try: - from tomllib import loads as toml_loads - except ModuleNotFoundError: - from tomli import loads as toml_loads + from tomllib import loads as toml_loads pyproject = toml_loads(Path('pyproject.toml').read_text(encoding='utf-8')) print(pyproject['project']['version']) diff --git a/CHANGELOG.md b/CHANGELOG.md index 8b7abce..1d77c27 100644 --- a/CHANGELOG.md +++ b/CHANGELOG.md @@ -8,6 +8,8 @@ The format is based on [Keep a Changelog](https://keepachangelog.com/en/1.1.0/). ### Changed (breaking) +- **Python 3.10 is no longer supported; 3.11 is the minimum.** The + `typing-extensions` and `tomli` backports it needed are dropped. - **Content conversion is rebuilt on three libraries: markdownify, mdformat and cmark-gfm.** `from_transport` reads GLPI's HTML with `markdownify`, and `mdformat` re-renders that Markdown from its syntax @@ -46,7 +48,7 @@ The format is based on [Keep a Changelog](https://keepachangelog.com/en/1.1.0/). Markdown passes verbatim unless it starts with an HTML tag, so Markdown carrying an inline `
      ` or `` stays Markdown. - **Dependencies.** - - Added: `cmarkgfm>=2025.10` (compiled wheels for CPython 3.10–3.14 on + - Added: `cmarkgfm>=2025.10` (compiled wheels for CPython 3.11–3.14 on Linux, macOS and Windows), `mdformat>=0.7.22,<0.8`, `mdformat-tables>=1.0` and `markdown-it-py>=3.0`. - Dropped: `markdown`. python-markdown 3.11 had broken the previous diff --git a/CONTRIBUTING.md b/CONTRIBUTING.md index 54192c6..abacc49 100644 --- a/CONTRIBUTING.md +++ b/CONTRIBUTING.md @@ -38,7 +38,7 @@ python -m sphinx -W --keep-going -b html docs docs/_build/html ## GitHub Actions -- `.github/workflows/ci.yml` runs tests for Python 3.10 through 3.14 on pull +- `.github/workflows/ci.yml` runs tests for Python 3.11 through 3.14 on pull requests and pushes to `main`. - The same workflow runs `ruff`, `mypy`, and a warning-free Sphinx build on Python 3.12. diff --git a/README.md b/README.md index d4a2d41..435e233 100644 --- a/README.md +++ b/README.md @@ -3,7 +3,7 @@ [![CI](https://github.com/baraline/glpi_python_client/actions/workflows/ci.yml/badge.svg?branch=main)](https://github.com/baraline/glpi_python_client/actions/workflows/ci.yml) [![Coverage](https://codecov.io/gh/baraline/glpi_python_client/branch/main/graph/badge.svg)](https://codecov.io/gh/baraline/glpi_python_client) [![License](https://img.shields.io/github/license/baraline/glpi_python_client)](LICENSE) -[![Python](https://img.shields.io/badge/python-3.10%2B-blue)](https://github.com/baraline/glpi_python_client) +[![Python](https://img.shields.io/badge/python-3.11%2B-blue)](https://github.com/baraline/glpi_python_client) [![Docs](https://readthedocs.org/projects/glpi-python-client/badge/?version=latest)](https://glpi-python-client.readthedocs.io/en/latest/) `glpi-python-client` is a typed Python client for the GLPI REST API. diff --git a/docs/conf.py b/docs/conf.py index 9fbb13c..562033a 100644 --- a/docs/conf.py +++ b/docs/conf.py @@ -5,11 +5,7 @@ from datetime import date from importlib.metadata import PackageNotFoundError, version from pathlib import Path - -try: - from tomllib import loads as toml_loads -except ModuleNotFoundError: - from tomli import loads as toml_loads +from tomllib import loads as toml_loads def _read_project_version() -> str: diff --git a/docs/installation.rst b/docs/installation.rst index bb1acea..c10cd41 100644 --- a/docs/installation.rst +++ b/docs/installation.rst @@ -4,7 +4,7 @@ Installation Requirements ------------ -``glpi-python-client`` supports Python 3.10 and newer. Runtime dependencies are installed +``glpi-python-client`` supports Python 3.11 and newer. Runtime dependencies are installed from the package metadata and include ``httpx``, ``tenacity``, ``beautifulsoup4``, ``lxml``, and ``pydantic``. diff --git a/docs/publishing_rtd.rst b/docs/publishing_rtd.rst index 91301ae..953a374 100644 --- a/docs/publishing_rtd.rst +++ b/docs/publishing_rtd.rst @@ -92,7 +92,7 @@ GitHub Actions and Read the Docs The repository ships with two GitHub Actions workflows: * ``.github/workflows/ci.yml`` runs on pull requests and pushes to ``main``. - It executes ``pytest`` on Python 3.10 through 3.14, then runs ``ruff``, + It executes ``pytest`` on Python 3.11 through 3.14, then runs ``ruff``, ``mypy``, and the Sphinx build on Python 3.12. * ``.github/workflows/release.yml`` runs on published GitHub releases. It repeats the quality checks, builds the source and wheel distributions, diff --git a/glpi_python_client/_async/auth/_v1_session.py b/glpi_python_client/_async/auth/_v1_session.py index a80f788..0c2304d 100644 --- a/glpi_python_client/_async/auth/_v1_session.py +++ b/glpi_python_client/_async/auth/_v1_session.py @@ -37,7 +37,7 @@ import json import logging -from datetime import datetime, timedelta, timezone +from datetime import UTC, datetime, timedelta from typing import Any, cast import httpx @@ -170,7 +170,7 @@ async def _init_session(self) -> None: raise GlpiProtocolError("GLPI v1 initSession returned no session_token") self._session_token = str(token) - self._session_started_at = datetime.now(tz=timezone.utc) + self._session_started_at = datetime.now(tz=UTC) logger.info("GLPI v1 session initialised.") async def _ensure_session(self) -> None: @@ -196,7 +196,7 @@ def _is_session_stale(self) -> bool: if self._session_started_at is None: return True - return datetime.now(tz=timezone.utc) >= ( + return datetime.now(tz=UTC) >= ( self._session_started_at + self._session_refresh_interval ) diff --git a/glpi_python_client/_async/auth/auth.py b/glpi_python_client/_async/auth/auth.py index 9386863..64d7c32 100644 --- a/glpi_python_client/_async/auth/auth.py +++ b/glpi_python_client/_async/auth/auth.py @@ -8,7 +8,7 @@ from __future__ import annotations import logging -from datetime import datetime, timedelta, timezone +from datetime import UTC, datetime, timedelta import httpx from tenacity import retry, retry_if_exception_type, stop_after_attempt, wait_fixed @@ -241,7 +241,7 @@ def _store_token_data( if refresh_token: self.refresh_token = refresh_token expires_in = int(str(token_data.get("expires_in") or 3600)) - now = datetime.now(tz=timezone.utc) + now = datetime.now(tz=UTC) self.token_updated_at = now self.token_expires_at = now + timedelta(seconds=expires_in) logger.info("GLPI OAuth token %s successfully.", label) @@ -399,7 +399,7 @@ async def ensure_token(self) -> None: await self._acquire_token() return - now = datetime.now(tz=timezone.utc) + now = datetime.now(tz=UTC) token_expired = ( self.token_expires_at is not None and now >= self.token_expires_at ) diff --git a/glpi_python_client/_async/auth/tests/test_auth.py b/glpi_python_client/_async/auth/tests/test_auth.py index a19f820..b7050e2 100644 --- a/glpi_python_client/_async/auth/tests/test_auth.py +++ b/glpi_python_client/_async/auth/tests/test_auth.py @@ -1,6 +1,6 @@ from __future__ import annotations -from datetime import datetime, timedelta, timezone +from datetime import UTC, datetime, timedelta from typing import cast import httpx @@ -125,8 +125,8 @@ async def test_token_manager_refreshes_when_configured_interval_elapses() -> Non ) auth.access_token = "old-token" auth.refresh_token = "refresh-token" - auth.token_updated_at = datetime.now(tz=timezone.utc) - timedelta(seconds=61) - auth.token_expires_at = datetime.now(tz=timezone.utc) + timedelta(hours=1) + auth.token_updated_at = datetime.now(tz=UTC) - timedelta(seconds=61) + auth.token_expires_at = datetime.now(tz=UTC) + timedelta(hours=1) await auth.ensure_token() @@ -166,7 +166,7 @@ def test_token_manager_logout_clears_cached_tokens() -> None: ) auth.access_token = "access-token" auth.refresh_token = "refresh-token" - auth.token_updated_at = datetime.now(tz=timezone.utc) + auth.token_updated_at = datetime.now(tz=UTC) auth.token_expires_at = auth.token_updated_at + timedelta(hours=1) auth.logout() @@ -282,8 +282,8 @@ def _make_refresh_ready_manager( ) manager.access_token = "stale-token" manager.refresh_token = "refresh-token" - manager.token_updated_at = datetime.now(tz=timezone.utc) - timedelta(hours=2) - manager.token_expires_at = datetime.now(tz=timezone.utc) - timedelta(seconds=1) + manager.token_updated_at = datetime.now(tz=UTC) - timedelta(hours=2) + manager.token_expires_at = datetime.now(tz=UTC) - timedelta(seconds=1) return manager diff --git a/glpi_python_client/_async/clients/_base_client.py b/glpi_python_client/_async/clients/_base_client.py index cf12737..8591f32 100644 --- a/glpi_python_client/_async/clients/_base_client.py +++ b/glpi_python_client/_async/clients/_base_client.py @@ -13,13 +13,7 @@ import logging import os -import sys -from typing import TYPE_CHECKING - -if sys.version_info >= (3, 11): - from typing import Self -else: # pragma: no cover - fallback for Python 3.10 - from typing_extensions import Self +from typing import TYPE_CHECKING, Self from glpi_python_client._async._concurrency import Lock from glpi_python_client._async.clients.commons._config import ( diff --git a/glpi_python_client/_async/clients/api/management/tests/test_document.py b/glpi_python_client/_async/clients/api/management/tests/test_document.py index 1e7f1a1..507d846 100644 --- a/glpi_python_client/_async/clients/api/management/tests/test_document.py +++ b/glpi_python_client/_async/clients/api/management/tests/test_document.py @@ -294,10 +294,10 @@ def _pin_token(client: Any) -> None: helper runs through it, so the token has to be supplied here instead. """ - from datetime import datetime, timedelta, timezone + from datetime import UTC, datetime, timedelta client._auth.access_token = "stub-token" - client._auth.token_expires_at = datetime.now(tz=timezone.utc) + timedelta(days=1) + client._auth.token_expires_at = datetime.now(tz=UTC) + timedelta(days=1) async def test_stream_document_content_yields_chunks(client: Any) -> None: diff --git a/glpi_python_client/_async/clients/client.py b/glpi_python_client/_async/clients/client.py index a605d51..15f4c3e 100644 --- a/glpi_python_client/_async/clients/client.py +++ b/glpi_python_client/_async/clients/client.py @@ -14,13 +14,8 @@ from __future__ import annotations import logging -import sys from types import TracebackType - -if sys.version_info >= (3, 11): - from typing import Self -else: # pragma: no cover - fallback for Python 3.10 - from typing_extensions import Self +from typing import Self from glpi_python_client._async.clients._base_client import _BaseGlpiClient from glpi_python_client._async.clients.api import ( diff --git a/glpi_python_client/_async/clients/commons/tests/test_payloads.py b/glpi_python_client/_async/clients/commons/tests/test_payloads.py index cbb6251..6b6611a 100644 --- a/glpi_python_client/_async/clients/commons/tests/test_payloads.py +++ b/glpi_python_client/_async/clients/commons/tests/test_payloads.py @@ -3,7 +3,7 @@ from __future__ import annotations import json -from datetime import datetime, timedelta, timezone +from datetime import UTC, datetime, timedelta, timezone from glpi_python_client._async.clients.commons._payloads import ( model_from_payload, @@ -101,7 +101,7 @@ def test_model_to_payload_rewrites_an_aware_datetime_onto_the_server_clock() -> offset has to be spent on the conversion instead of written out. """ - aware = datetime(2024, 1, 1, 12, 0, tzinfo=timezone.utc) + aware = datetime(2024, 1, 1, 12, 0, tzinfo=UTC) body = model_to_payload( PostTicketTask(planned_begin=aware), @@ -114,7 +114,7 @@ def test_model_to_payload_rewrites_an_aware_datetime_onto_the_server_clock() -> def test_model_to_payload_leaves_an_aware_datetime_alone_without_a_timezone() -> None: """Outside the client there is no server clock to convert onto.""" - aware = datetime(2024, 1, 1, 12, 0, tzinfo=timezone.utc) + aware = datetime(2024, 1, 1, 12, 0, tzinfo=UTC) body = model_to_payload(PostTicketTask(planned_begin=aware)) diff --git a/glpi_python_client/_async/tests/test_concurrency.py b/glpi_python_client/_async/tests/test_concurrency.py index 4f4d134..46b6062 100644 --- a/glpi_python_client/_async/tests/test_concurrency.py +++ b/glpi_python_client/_async/tests/test_concurrency.py @@ -15,7 +15,7 @@ from __future__ import annotations import asyncio -from datetime import datetime, timedelta, timezone +from datetime import UTC, datetime, timedelta from typing import Any import httpx @@ -53,7 +53,7 @@ async def _request(method: str, url: str, **kwargs: Any) -> _Response: client._session.request = _request # type: ignore[method-assign,assignment] client._auth.access_token = "stub-token" - client._auth.token_expires_at = datetime.now(tz=timezone.utc) + timedelta(days=365) + client._auth.token_expires_at = datetime.now(tz=UTC) + timedelta(days=365) return calls diff --git a/glpi_python_client/_sync/auth/_v1_session.py b/glpi_python_client/_sync/auth/_v1_session.py index 8f4f950..222afcd 100644 --- a/glpi_python_client/_sync/auth/_v1_session.py +++ b/glpi_python_client/_sync/auth/_v1_session.py @@ -37,7 +37,7 @@ import json import logging -from datetime import datetime, timedelta, timezone +from datetime import UTC, datetime, timedelta from typing import Any, cast import httpx @@ -170,7 +170,7 @@ def _init_session(self) -> None: raise GlpiProtocolError("GLPI v1 initSession returned no session_token") self._session_token = str(token) - self._session_started_at = datetime.now(tz=timezone.utc) + self._session_started_at = datetime.now(tz=UTC) logger.info("GLPI v1 session initialised.") def _ensure_session(self) -> None: @@ -196,7 +196,7 @@ def _is_session_stale(self) -> bool: if self._session_started_at is None: return True - return datetime.now(tz=timezone.utc) >= ( + return datetime.now(tz=UTC) >= ( self._session_started_at + self._session_refresh_interval ) diff --git a/glpi_python_client/_sync/auth/auth.py b/glpi_python_client/_sync/auth/auth.py index 4409650..a4ba60f 100644 --- a/glpi_python_client/_sync/auth/auth.py +++ b/glpi_python_client/_sync/auth/auth.py @@ -8,7 +8,7 @@ from __future__ import annotations import logging -from datetime import datetime, timedelta, timezone +from datetime import UTC, datetime, timedelta import httpx from tenacity import retry, retry_if_exception_type, stop_after_attempt, wait_fixed @@ -241,7 +241,7 @@ def _store_token_data( if refresh_token: self.refresh_token = refresh_token expires_in = int(str(token_data.get("expires_in") or 3600)) - now = datetime.now(tz=timezone.utc) + now = datetime.now(tz=UTC) self.token_updated_at = now self.token_expires_at = now + timedelta(seconds=expires_in) logger.info("GLPI OAuth token %s successfully.", label) @@ -399,7 +399,7 @@ def ensure_token(self) -> None: self._acquire_token() return - now = datetime.now(tz=timezone.utc) + now = datetime.now(tz=UTC) token_expired = ( self.token_expires_at is not None and now >= self.token_expires_at ) diff --git a/glpi_python_client/_sync/auth/tests/test_auth.py b/glpi_python_client/_sync/auth/tests/test_auth.py index f3c73f8..221297a 100644 --- a/glpi_python_client/_sync/auth/tests/test_auth.py +++ b/glpi_python_client/_sync/auth/tests/test_auth.py @@ -1,6 +1,6 @@ from __future__ import annotations -from datetime import datetime, timedelta, timezone +from datetime import UTC, datetime, timedelta from typing import cast import httpx @@ -125,8 +125,8 @@ def test_token_manager_refreshes_when_configured_interval_elapses() -> None: ) auth.access_token = "old-token" auth.refresh_token = "refresh-token" - auth.token_updated_at = datetime.now(tz=timezone.utc) - timedelta(seconds=61) - auth.token_expires_at = datetime.now(tz=timezone.utc) + timedelta(hours=1) + auth.token_updated_at = datetime.now(tz=UTC) - timedelta(seconds=61) + auth.token_expires_at = datetime.now(tz=UTC) + timedelta(hours=1) auth.ensure_token() @@ -166,7 +166,7 @@ def test_token_manager_logout_clears_cached_tokens() -> None: ) auth.access_token = "access-token" auth.refresh_token = "refresh-token" - auth.token_updated_at = datetime.now(tz=timezone.utc) + auth.token_updated_at = datetime.now(tz=UTC) auth.token_expires_at = auth.token_updated_at + timedelta(hours=1) auth.logout() @@ -282,8 +282,8 @@ def _make_refresh_ready_manager( ) manager.access_token = "stale-token" manager.refresh_token = "refresh-token" - manager.token_updated_at = datetime.now(tz=timezone.utc) - timedelta(hours=2) - manager.token_expires_at = datetime.now(tz=timezone.utc) - timedelta(seconds=1) + manager.token_updated_at = datetime.now(tz=UTC) - timedelta(hours=2) + manager.token_expires_at = datetime.now(tz=UTC) - timedelta(seconds=1) return manager diff --git a/glpi_python_client/_sync/clients/_base_client.py b/glpi_python_client/_sync/clients/_base_client.py index a30b8c2..9a097a4 100644 --- a/glpi_python_client/_sync/clients/_base_client.py +++ b/glpi_python_client/_sync/clients/_base_client.py @@ -13,13 +13,7 @@ import logging import os -import sys -from typing import TYPE_CHECKING - -if sys.version_info >= (3, 11): - from typing import Self -else: # pragma: no cover - fallback for Python 3.10 - from typing_extensions import Self +from typing import TYPE_CHECKING, Self from glpi_python_client._sync._concurrency import Lock from glpi_python_client._sync.clients.commons._config import ( diff --git a/glpi_python_client/_sync/clients/api/management/tests/test_document.py b/glpi_python_client/_sync/clients/api/management/tests/test_document.py index 87fb963..c6a6a38 100644 --- a/glpi_python_client/_sync/clients/api/management/tests/test_document.py +++ b/glpi_python_client/_sync/clients/api/management/tests/test_document.py @@ -294,10 +294,10 @@ def _pin_token(client: Any) -> None: helper runs through it, so the token has to be supplied here instead. """ - from datetime import datetime, timedelta, timezone + from datetime import UTC, datetime, timedelta client._auth.access_token = "stub-token" - client._auth.token_expires_at = datetime.now(tz=timezone.utc) + timedelta(days=1) + client._auth.token_expires_at = datetime.now(tz=UTC) + timedelta(days=1) def test_stream_document_content_yields_chunks(client: Any) -> None: diff --git a/glpi_python_client/_sync/clients/client.py b/glpi_python_client/_sync/clients/client.py index 7a2bcdc..95fc0ef 100644 --- a/glpi_python_client/_sync/clients/client.py +++ b/glpi_python_client/_sync/clients/client.py @@ -14,13 +14,8 @@ from __future__ import annotations import logging -import sys from types import TracebackType - -if sys.version_info >= (3, 11): - from typing import Self -else: # pragma: no cover - fallback for Python 3.10 - from typing_extensions import Self +from typing import Self from glpi_python_client._sync.clients._base_client import _BaseGlpiClient from glpi_python_client._sync.clients.api import ( diff --git a/glpi_python_client/_sync/clients/commons/tests/test_payloads.py b/glpi_python_client/_sync/clients/commons/tests/test_payloads.py index 3bcf86f..6520921 100644 --- a/glpi_python_client/_sync/clients/commons/tests/test_payloads.py +++ b/glpi_python_client/_sync/clients/commons/tests/test_payloads.py @@ -3,7 +3,7 @@ from __future__ import annotations import json -from datetime import datetime, timedelta, timezone +from datetime import UTC, datetime, timedelta, timezone from glpi_python_client._sync.clients.commons._payloads import ( model_from_payload, @@ -101,7 +101,7 @@ def test_model_to_payload_rewrites_an_aware_datetime_onto_the_server_clock() -> offset has to be spent on the conversion instead of written out. """ - aware = datetime(2024, 1, 1, 12, 0, tzinfo=timezone.utc) + aware = datetime(2024, 1, 1, 12, 0, tzinfo=UTC) body = model_to_payload( PostTicketTask(planned_begin=aware), @@ -114,7 +114,7 @@ def test_model_to_payload_rewrites_an_aware_datetime_onto_the_server_clock() -> def test_model_to_payload_leaves_an_aware_datetime_alone_without_a_timezone() -> None: """Outside the client there is no server clock to convert onto.""" - aware = datetime(2024, 1, 1, 12, 0, tzinfo=timezone.utc) + aware = datetime(2024, 1, 1, 12, 0, tzinfo=UTC) body = model_to_payload(PostTicketTask(planned_begin=aware)) diff --git a/glpi_python_client/models/custom_schema/_ticket_context.py b/glpi_python_client/models/custom_schema/_ticket_context.py index 08c1bfa..e5034e4 100644 --- a/glpi_python_client/models/custom_schema/_ticket_context.py +++ b/glpi_python_client/models/custom_schema/_ticket_context.py @@ -9,7 +9,7 @@ from __future__ import annotations from dataclasses import dataclass, field -from datetime import datetime, timezone +from datetime import UTC, datetime from enum import Enum from typing import Any @@ -30,7 +30,7 @@ ) from glpi_python_client.models.api_schema.management._document import GetDocument -_MAX_DATETIME = datetime.max.replace(tzinfo=timezone.utc) +_MAX_DATETIME = datetime.max.replace(tzinfo=UTC) def _literal(text: str) -> str: @@ -192,7 +192,7 @@ def _event_sort_key(event: Any) -> datetime: if created is None: return _MAX_DATETIME if created.tzinfo is None: - return created.replace(tzinfo=timezone.utc) + return created.replace(tzinfo=UTC) return created diff --git a/glpi_python_client/models/custom_schema/tests/test_ticket_context.py b/glpi_python_client/models/custom_schema/tests/test_ticket_context.py index 95dc23f..288ea48 100644 --- a/glpi_python_client/models/custom_schema/tests/test_ticket_context.py +++ b/glpi_python_client/models/custom_schema/tests/test_ticket_context.py @@ -2,7 +2,7 @@ from __future__ import annotations -from datetime import datetime, timezone +from datetime import UTC, datetime import pytest from pydantic import ValidationError @@ -75,8 +75,8 @@ def test_to_markdown_renders_ticket_subtitle_metadata() -> None: "name": "Printer broken", "user_recipient": {"id": 7, "name": "Alice"}, "user_editor": {"id": 8, "name": "Bob"}, - "date_creation": datetime(2024, 1, 1, 9, 30, tzinfo=timezone.utc), - "date_mod": datetime(2024, 1, 2, 11, 45, tzinfo=timezone.utc), + "date_creation": datetime(2024, 1, 1, 9, 30, tzinfo=UTC), + "date_mod": datetime(2024, 1, 2, 11, 45, tzinfo=UTC), } } ) @@ -99,12 +99,12 @@ def test_to_markdown_orders_events_by_creation_when_no_position() -> None: { "id": 2, "content": "second note", - "date_creation": datetime(2024, 1, 2, tzinfo=timezone.utc), + "date_creation": datetime(2024, 1, 2, tzinfo=UTC), }, { "id": 1, "content": "first note", - "date_creation": datetime(2024, 1, 1, tzinfo=timezone.utc), + "date_creation": datetime(2024, 1, 1, tzinfo=UTC), }, ], } @@ -123,7 +123,7 @@ def test_to_markdown_ignores_timeline_position_for_ordering() -> None: { "id": 1, "content": "no position late", - "date_creation": datetime(2024, 1, 5, tzinfo=timezone.utc), + "date_creation": datetime(2024, 1, 5, tzinfo=UTC), }, ], "tasks": [ @@ -131,7 +131,7 @@ def test_to_markdown_ignores_timeline_position_for_ordering() -> None: "id": 2, "content": "left positioned", "timeline_position": 1, - "date_creation": datetime(2024, 1, 10, tzinfo=timezone.utc), + "date_creation": datetime(2024, 1, 10, tzinfo=UTC), }, ], } @@ -197,8 +197,8 @@ def test_to_markdown_renders_event_creator_editor_and_timestamps() -> None: "content": "note", "user": {"id": 7, "name": "Alice"}, "user_editor": {"id": 8, "name": "Bob"}, - "date_creation": datetime(2024, 1, 2, 10, 0, tzinfo=timezone.utc), - "date_mod": datetime(2024, 1, 2, 10, 5, tzinfo=timezone.utc), + "date_creation": datetime(2024, 1, 2, 10, 0, tzinfo=UTC), + "date_mod": datetime(2024, 1, 2, 10, 5, tzinfo=UTC), } ], } @@ -226,8 +226,8 @@ def test_to_markdown_renders_event_creator_editor_and_timestamps() -> None: "status": {"id": 2, "name": "Open"}, "user_recipient": {"id": 3, "name": "Alice"}, "user_editor": {"id": 4, "name": "Bob"}, - "date_creation": datetime(2024, 1, 1, tzinfo=timezone.utc), - "date_mod": datetime(2024, 1, 2, tzinfo=timezone.utc), + "date_creation": datetime(2024, 1, 1, tzinfo=UTC), + "date_mod": datetime(2024, 1, 2, tzinfo=UTC), }, "followups": [{"id": 10, "content": "followup body"}], "tasks": [{"id": 20, "content": "task body", "duration": 600}], @@ -368,7 +368,7 @@ def test_options_hide_event_dates() -> None: { "id": 5, "content": "note", - "date_creation": datetime(2024, 3, 1, tzinfo=timezone.utc), + "date_creation": datetime(2024, 3, 1, tzinfo=UTC), } ], } @@ -415,7 +415,7 @@ def test_to_markdown_sorts_aware_events_when_one_lacks_a_creation_date() -> None { "id": 1, "content": "dated note", - "date_creation": datetime(2024, 1, 1, tzinfo=timezone.utc), + "date_creation": datetime(2024, 1, 1, tzinfo=UTC), }, ], } @@ -440,7 +440,7 @@ def test_to_markdown_orders_events_across_mixed_datetime_awareness() -> None: { "id": 2, "content": "aware second", - "date_creation": datetime(2024, 1, 2, tzinfo=timezone.utc), + "date_creation": datetime(2024, 1, 2, tzinfo=UTC), }, { "id": 1, diff --git a/glpi_python_client/models/tests/test_base.py b/glpi_python_client/models/tests/test_base.py index 8e0fbe1..2990880 100644 --- a/glpi_python_client/models/tests/test_base.py +++ b/glpi_python_client/models/tests/test_base.py @@ -7,7 +7,7 @@ from __future__ import annotations -from datetime import datetime, timedelta, timezone +from datetime import UTC, datetime, timedelta, timezone from typing import Annotated import pytest @@ -145,7 +145,7 @@ def test_aware_datetime_is_converted_to_the_server_clock_and_stripped() -> None: the conversion rather than written out. """ - stamped = _Stamped(id=1, date=datetime(2024, 1, 1, 12, 0, tzinfo=timezone.utc)) + stamped = _Stamped(id=1, date=datetime(2024, 1, 1, 12, 0, tzinfo=UTC)) dumped = stamped.model_dump(mode="json", context={"server_timezone": _PARIS_WINTER}) @@ -169,7 +169,7 @@ def test_serialisation_without_a_context_leaves_the_offset_alone() -> None: one would be the same silent shift the conversion exists to prevent. """ - stamped = _Stamped(id=1, date=datetime(2024, 1, 1, 12, 0, tzinfo=timezone.utc)) + stamped = _Stamped(id=1, date=datetime(2024, 1, 1, 12, 0, tzinfo=UTC)) assert stamped.model_dump(mode="json")["date"] == "2024-01-01T12:00:00Z" @@ -179,7 +179,7 @@ def test_the_serialisation_timezone_reaches_a_nested_model() -> None: nested = _Nested( id=1, - inner=_Stamped(id=2, date=datetime(2024, 1, 1, 12, 0, tzinfo=timezone.utc)), + inner=_Stamped(id=2, date=datetime(2024, 1, 1, 12, 0, tzinfo=UTC)), ) dumped = nested.model_dump(mode="json", context={"server_timezone": _PARIS_WINTER}) diff --git a/glpi_python_client/testing/tests/test_method_invocation.py b/glpi_python_client/testing/tests/test_method_invocation.py index 5aa2f07..9111401 100644 --- a/glpi_python_client/testing/tests/test_method_invocation.py +++ b/glpi_python_client/testing/tests/test_method_invocation.py @@ -25,7 +25,7 @@ import inspect from collections.abc import AsyncIterator, Iterator from contextlib import asynccontextmanager, contextmanager -from datetime import datetime, timedelta, timezone +from datetime import UTC, datetime, timedelta from typing import Any, ClassVar, get_type_hints import pytest @@ -212,7 +212,7 @@ async def _astream( # Pretend a valid, non-expiring token is already held so no OAuth round # trip happens and the call log contains only endpoint traffic. client._auth.access_token = "stub-token" - client._auth.token_expires_at = datetime.now(tz=timezone.utc) + timedelta(days=365) + client._auth.token_expires_at = datetime.now(tz=UTC) + timedelta(days=365) # Several features (plugin fields, KB category writes, document upload, # actor statistics) run on the legacy v1 session rather than the v2 # transport. Stub it into the same log so they are exercised too. diff --git a/glpi_python_client/testing/tests/test_packaging.py b/glpi_python_client/testing/tests/test_packaging.py index c6e7563..348e82f 100644 --- a/glpi_python_client/testing/tests/test_packaging.py +++ b/glpi_python_client/testing/tests/test_packaging.py @@ -13,12 +13,7 @@ from __future__ import annotations import pathlib -import sys - -if sys.version_info >= (3, 11): - import tomllib -else: # pragma: no cover - exercised on 3.10 only - import tomli as tomllib +import tomllib _REPO_ROOT = pathlib.Path(__file__).resolve().parents[3] diff --git a/glpi_python_client/testing/tests/test_version_agreement.py b/glpi_python_client/testing/tests/test_version_agreement.py index 4e2e55b..93b8a8e 100644 --- a/glpi_python_client/testing/tests/test_version_agreement.py +++ b/glpi_python_client/testing/tests/test_version_agreement.py @@ -17,12 +17,7 @@ import pathlib import re -import sys - -if sys.version_info >= (3, 11): - import tomllib -else: # pragma: no cover - exercised on 3.10 only - import tomli as tomllib +import tomllib import glpi_python_client diff --git a/glpi_python_client/tests/test_rsql.py b/glpi_python_client/tests/test_rsql.py index a8692a9..a3342e2 100644 --- a/glpi_python_client/tests/test_rsql.py +++ b/glpi_python_client/tests/test_rsql.py @@ -2,7 +2,7 @@ from __future__ import annotations -from datetime import date, datetime, timezone +from datetime import UTC, date, datetime from zoneinfo import ZoneInfo import pytest @@ -98,7 +98,7 @@ def test_changed_since_converts_an_aware_datetime_into_the_server_zone() -> None must be the server's rendering of that same instant. """ - aware = datetime(2026, 8, 12, 7, 33, tzinfo=timezone.utc) + aware = datetime(2026, 8, 12, 7, 33, tzinfo=UTC) assert changed_since(aware, tz=ZoneInfo("Europe/Paris")) == ( "date_mod=ge=2026-08-12 09:33:00" @@ -113,8 +113,8 @@ def test_changed_since_follows_dst_in_the_server_zone() -> None: """ paris = ZoneInfo("Europe/Paris") - winter = datetime(2026, 1, 15, 12, 0, tzinfo=timezone.utc) - summer = datetime(2026, 7, 15, 12, 0, tzinfo=timezone.utc) + winter = datetime(2026, 1, 15, 12, 0, tzinfo=UTC) + summer = datetime(2026, 7, 15, 12, 0, tzinfo=UTC) assert changed_since(winter, tz=paris) == "date_mod=ge=2026-01-15 13:00:00" assert changed_since(summer, tz=paris) == "date_mod=ge=2026-07-15 14:00:00" @@ -129,7 +129,7 @@ def test_changed_since_rejects_an_aware_datetime_without_a_zone() -> None: over-reads east of UTC and skips modifications west of it. """ - aware = datetime(2026, 1, 1, 13, 45, 30, tzinfo=timezone.utc) + aware = datetime(2026, 1, 1, 13, 45, 30, tzinfo=UTC) with pytest.raises(GlpiValidationError): changed_since(aware) diff --git a/pyproject.toml b/pyproject.toml index eb89aec..62f048e 100644 --- a/pyproject.toml +++ b/pyproject.toml @@ -38,7 +38,7 @@ name = "glpi-python-client" version = "0.6.0" description = "A typed Python client for GLPI ITSM APIs." readme = "README.md" -requires-python = ">=3.10" +requires-python = ">=3.11" license = { text = "MIT" } authors = [{ name = "glpi-python-client contributors" }] keywords = ["glpi", "itsm", "api", "client"] @@ -48,7 +48,6 @@ classifiers = [ "License :: OSI Approved :: MIT License", "Operating System :: OS Independent", "Programming Language :: Python :: 3", - "Programming Language :: Python :: 3.10", "Programming Language :: Python :: 3.11", "Programming Language :: Python :: 3.12", "Programming Language :: Python :: 3.13", @@ -89,7 +88,6 @@ dependencies = [ # CI and raises ZoneInfoNotFoundError on a developer machine. "tzdata>=2024.1; platform_system == 'Windows'", "tenacity>=8.2", - "typing-extensions>=4.7; python_version < '3.11'", ] [project.optional-dependencies] @@ -97,7 +95,6 @@ docs = [ "numpydoc>=1.8", "sphinx>=7.2,<8.2", "sphinx-rtd-theme>=2.0", - "tomli>=2.0; python_version < '3.11'", ] dev = [ "build>=1.2", @@ -110,7 +107,6 @@ dev = [ "ruff>=0.6", "sphinx>=7.2,<8.2", "sphinx-rtd-theme>=2.0", - "tomli>=2.0; python_version < '3.11'", "twine>=5.1", "unasync>=0.6", "vulture>=2.11", @@ -185,7 +181,7 @@ exclude_lines = [ [tool.ruff] line-length = 88 -target-version = "py310" +target-version = "py311" # The sync client tree is generated by unasync_build.py from _async/ and # must stay byte-identical to what regeneration produces -- that identity # is what the CI gate checks. Formatting it would fight the generator, and @@ -197,7 +193,7 @@ extend-exclude = ["glpi_python_client/_sync"] select = ["B", "E", "F", "I", "RUF", "UP"] [tool.mypy] -python_version = "3.10" +python_version = "3.11" packages = ["glpi_python_client"] strict = true warn_unreachable = true diff --git a/skills/glpi-asset-workflow/SKILL.md b/skills/glpi-asset-workflow/SKILL.md index d830a20..e7b32f9 100644 --- a/skills/glpi-asset-workflow/SKILL.md +++ b/skills/glpi-asset-workflow/SKILL.md @@ -2,7 +2,7 @@ name: glpi-asset-workflow description: "Search, fetch, create, update, and delete GLPI computers, and read or write the contracts covering them, with the synchronous glpi_python_client.GlpiClient or the asynchronous AsyncGlpiClient, and the GetComputer/PostComputer/PatchComputer/DeleteComputer and GetContractItem/PostContractItem models. Use for GLPI asset inventory, computer records, asset serial numbers, asset locations, or finding which contracts cover a machine." license: MIT -compatibility: "Requires Python 3.10+, glpi-python-client, network access to the GLPI v2 API, and credentials allowed to read or write assets." +compatibility: "Requires Python 3.11+, glpi-python-client, network access to the GLPI v2 API, and credentials allowed to read or write assets." metadata: package: glpi-python-client version: "0.6.0" diff --git a/skills/glpi-client-setup/SKILL.md b/skills/glpi-client-setup/SKILL.md index 5919a69..367709a 100644 --- a/skills/glpi-client-setup/SKILL.md +++ b/skills/glpi-client-setup/SKILL.md @@ -2,7 +2,7 @@ name: glpi-client-setup description: "Create and configure the synchronous glpi_python_client.GlpiClient or the asynchronous glpi_python_client.AsyncGlpiClient, including from_env, OAuth credential pairs, entity/profile headers, SSL settings, and the optional legacy v1 session (v1_base_url / v1_user_token) that backs document uploads, the Fields plugin helpers, KB category writes and actor-based statistics. Use before calling GLPI APIs, when configuring the v1 session for any of those features, or when the user asks how to connect to GLPI with glpi_python_client." license: MIT -compatibility: "Requires Python 3.10+, glpi-python-client, network access to a GLPI v2 API, and valid GLPI credentials." +compatibility: "Requires Python 3.11+, glpi-python-client, network access to a GLPI v2 API, and valid GLPI credentials." metadata: package: glpi-python-client version: "0.6.0" diff --git a/skills/glpi-contract-workflow/SKILL.md b/skills/glpi-contract-workflow/SKILL.md index 0a55c56..d0498a3 100644 --- a/skills/glpi-contract-workflow/SKILL.md +++ b/skills/glpi-contract-workflow/SKILL.md @@ -2,7 +2,7 @@ name: glpi-contract-workflow description: "Search, fetch, create, update, and delete GLPI contracts, their cost lines, and the contract-type dropdown, with the synchronous glpi_python_client.GlpiClient or the asynchronous AsyncGlpiClient, and the GetContract/PostContract/PatchContract/DeleteContract, GetContractCost/PostContractCost, and GetContractType/PostContractType models. Use for GLPI contract coverage, maintenance agreements, contract cost/budget lines, contract renewal type, or the contract-type dropdown." license: MIT -compatibility: "Requires Python 3.10+, glpi-python-client, network access to the GLPI v2 API, and credentials allowed to read or write contracts." +compatibility: "Requires Python 3.11+, glpi-python-client, network access to the GLPI v2 API, and credentials allowed to read or write contracts." metadata: package: glpi-python-client version: "0.6.0" diff --git a/skills/glpi-document-workflow/SKILL.md b/skills/glpi-document-workflow/SKILL.md index 3fe75c8..3c119f3 100644 --- a/skills/glpi-document-workflow/SKILL.md +++ b/skills/glpi-document-workflow/SKILL.md @@ -2,7 +2,7 @@ name: glpi-document-workflow description: "Manage GLPI document metadata, upload binary content via the legacy v1 fallback, download document binaries, and link documents to a ticket timeline with the synchronous glpi_python_client.GlpiClient or the asynchronous AsyncGlpiClient, and the GetDocument/PostDocument/PatchDocument/DeleteDocument models. Use for ticket attachments, document binary content, document metadata, or saving downloaded files." license: MIT -compatibility: "Requires Python 3.10+, glpi-python-client, network access to the GLPI v2 API, and v1 credentials configured on the client for binary uploads." +compatibility: "Requires Python 3.11+, glpi-python-client, network access to the GLPI v2 API, and v1 credentials configured on the client for binary uploads." metadata: package: glpi-python-client version: "0.6.0" diff --git a/skills/glpi-knowledge-base/SKILL.md b/skills/glpi-knowledge-base/SKILL.md index c2cbe2d..10875e1 100644 --- a/skills/glpi-knowledge-base/SKILL.md +++ b/skills/glpi-knowledge-base/SKILL.md @@ -2,7 +2,7 @@ name: glpi-knowledge-base description: "Search, read, create, update, and delete GLPI knowledge base articles, categories, comments, and revisions with the synchronous glpi_python_client.GlpiClient or the asynchronous AsyncGlpiClient, and the GetKBArticle/PostKBArticle/GetKBCategory/GetKBArticleComment/GetKBArticleRevision models. Use for GLPI knowledge base content, FAQ articles, article categories, article comments, article revision history, or assigning categories to a KB article." license: MIT -compatibility: "Requires Python 3.10+, glpi-python-client, network access to the GLPI v2 API, and — for category writes only — a legacy v1 session (v1_base_url + v1_user_token)." +compatibility: "Requires Python 3.11+, glpi-python-client, network access to the GLPI v2 API, and — for category writes only — a legacy v1 session (v1_base_url + v1_user_token)." metadata: package: glpi-python-client version: "0.6.0" diff --git a/skills/glpi-plugin-fields/SKILL.md b/skills/glpi-plugin-fields/SKILL.md index f90548e..ce8a0ec 100644 --- a/skills/glpi-plugin-fields/SKILL.md +++ b/skills/glpi-plugin-fields/SKILL.md @@ -2,7 +2,7 @@ name: glpi-plugin-fields description: "Discover and read/write GLPI Fields-plugin custom fields with the synchronous glpi_python_client.GlpiClient or the asynchronous AsyncGlpiClient — list_plugin_fields_containers, list_plugin_fields_fields, list_item_plugin_field_rows, create_item_plugin_field_row, update_item_plugin_field_row, and the Ticket-only get_ticket_custom_fields/set_ticket_custom_fields. Use for GLPI custom fields, the Fields plugin, per-instance extra ticket attributes, or reading a ticket's custom-field values." license: MIT -compatibility: "Requires Python 3.10+, glpi-python-client, the GLPI Fields plugin installed server-side, and a legacy v1 session (v1_base_url + v1_user_token) — every method in this family goes over the v1 API." +compatibility: "Requires Python 3.11+, glpi-python-client, the GLPI Fields plugin installed server-side, and a legacy v1 session (v1_base_url + v1_user_token) — every method in this family goes over the v1 API." metadata: package: glpi-python-client version: "0.6.0" diff --git a/skills/glpi-reporting-and-context/SKILL.md b/skills/glpi-reporting-and-context/SKILL.md index 70b9808..774fbb2 100644 --- a/skills/glpi-reporting-and-context/SKILL.md +++ b/skills/glpi-reporting-and-context/SKILL.md @@ -2,7 +2,7 @@ name: glpi-reporting-and-context description: "Aggregate GLPI ticket and task statistics and load grouped ticket contexts with the synchronous glpi_python_client.GlpiClient or the asynchronous AsyncGlpiClient. Use for operational reporting, ticket counts grouped by entity/status/priority/type, task duration totals grouped by user/entity/ticket, per-user activity reports, batch-streamed pagination of search results, or one-call ticket context retrieval bundling tickets with timeline records." license: MIT -compatibility: "Requires Python 3.10+, glpi-python-client, network access to the GLPI v2 API, and credentials allowed to read tickets, tasks, users, entities, and timeline records." +compatibility: "Requires Python 3.11+, glpi-python-client, network access to the GLPI v2 API, and credentials allowed to read tickets, tasks, users, entities, and timeline records." metadata: package: glpi-python-client version: "0.6.0" diff --git a/skills/glpi-team-members/SKILL.md b/skills/glpi-team-members/SKILL.md index 42ce77c..d6b811e 100644 --- a/skills/glpi-team-members/SKILL.md +++ b/skills/glpi-team-members/SKILL.md @@ -2,7 +2,7 @@ name: glpi-team-members description: "List, add, and remove GLPI ticket team members with the synchronous glpi_python_client.GlpiClient or the asynchronous AsyncGlpiClient, and the GetTeamMember/PostTeamMember models. Use when assigning users or groups to tickets, inspecting ticket teams, or removing GLPI ticket participants." license: MIT -compatibility: "Requires Python 3.10+, glpi-python-client, network access to the GLPI v2 API, and credentials allowed to manage ticket teams." +compatibility: "Requires Python 3.11+, glpi-python-client, network access to the GLPI v2 API, and credentials allowed to manage ticket teams." metadata: package: glpi-python-client version: "0.6.0" diff --git a/skills/glpi-ticket-timeline/SKILL.md b/skills/glpi-ticket-timeline/SKILL.md index cd18a59..02be123 100644 --- a/skills/glpi-ticket-timeline/SKILL.md +++ b/skills/glpi-ticket-timeline/SKILL.md @@ -2,7 +2,7 @@ name: glpi-ticket-timeline description: "Read GLPI ticket timeline records and create or update followups, tasks, solutions, and timeline document links with the synchronous glpi_python_client.GlpiClient or the asynchronous AsyncGlpiClient. Use when handling ticket notes, followups, tasks, solutions, or attached documents on a GLPI ticket timeline." license: MIT -compatibility: "Requires Python 3.10+, glpi-python-client, and network access to the GLPI v2 API." +compatibility: "Requires Python 3.11+, glpi-python-client, and network access to the GLPI v2 API." metadata: package: glpi-python-client version: "0.6.0" diff --git a/skills/glpi-ticket-workflow/SKILL.md b/skills/glpi-ticket-workflow/SKILL.md index cfb3944..b8ef0b8 100644 --- a/skills/glpi-ticket-workflow/SKILL.md +++ b/skills/glpi-ticket-workflow/SKILL.md @@ -2,7 +2,7 @@ name: glpi-ticket-workflow description: "Search, fetch, create, update, and delete GLPI tickets with the synchronous glpi_python_client.GlpiClient or the asynchronous AsyncGlpiClient, and the GetTicket/PostTicket/PatchTicket/DeleteTicket models. Use for GLPI ticket records, ticket filters, fields, pagination, status, priority, category, location, or instance-specific extra_payload values." license: MIT -compatibility: "Requires Python 3.10+, glpi-python-client, network access to the GLPI v2 API, and credentials accepted by GlpiClient." +compatibility: "Requires Python 3.11+, glpi-python-client, network access to the GLPI v2 API, and credentials accepted by GlpiClient." metadata: package: glpi-python-client version: "0.6.0" diff --git a/skills/glpi-user-location-provisioning/SKILL.md b/skills/glpi-user-location-provisioning/SKILL.md index d8f0c6a..b99557c 100644 --- a/skills/glpi-user-location-provisioning/SKILL.md +++ b/skills/glpi-user-location-provisioning/SKILL.md @@ -2,7 +2,7 @@ name: glpi-user-location-provisioning description: "Search GLPI users, locations, and entities, or create, update, and delete users and locations and entities with the synchronous glpi_python_client.GlpiClient or the asynchronous AsyncGlpiClient, and the matching Get/Post/Patch/Delete models. Use for user lookup, entity lookup, location lookup, user provisioning, location creation, GLPI entity defaults, or RSQL filters." license: MIT -compatibility: "Requires Python 3.10+, glpi-python-client, network access to the GLPI v2 API, and credentials allowed to read or write users, locations, and entities." +compatibility: "Requires Python 3.11+, glpi-python-client, network access to the GLPI v2 API, and credentials allowed to read or write users, locations, and entities." metadata: package: glpi-python-client version: "0.6.0"