diff --git a/CHANGELOG.md b/CHANGELOG.md index 1d77c27..1ea91ed 100644 --- a/CHANGELOG.md +++ b/CHANGELOG.md @@ -4,6 +4,134 @@ All notable changes to this project are documented in this file. The format is based on [Keep a Changelog](https://keepachangelog.com/en/1.1.0/). +## 0.6.1 — 2026-10-02 + +The reader now gives the same Markdown as `easyvista-python-client` 0.4.1, +whose converter is this package's 0.6.0 converter plus fifteen fixes and a +correction to one of them. This release ports those fixes and the +correction, with their tests. The two conversion modules now differ only in +names, error messages, docstrings and that package's optional-dependency +import guard. The converter's writing (`to_transport`) is unchanged; a +write model now sends a lone surrogate as U+FFFD. + +### Fixed + +- **Image alt text lost its escapes**: `a_b_c.png` read back as `abc.png`. + Escapes and character references in alt text are kept. +- **A line opening with `~~~` became a code fence** that swallowed the rest + of the body. It reads as `\~~~`. +- **A `!` right before a link turned the link into an image** + (`Attention!` then a link). It reads as `\!`. +- **An image whose alt text starts with `^` stopped being an image** once + rendered, since cmark-gfm never opens one on `![^`. The `^` is escaped. +- **`<` became an autolink** behind an escaped `<`. The second + `<` is escaped too. +- **A line break inside inline code showed as a literal backslash.** + ``, `` and `` holding a `
` become one span per line; + a `|` in `` or `` in a table cell no longer splits the row. +- **A table inside a heading or a link showed its pipes as text**, and + inside a link did not read back as itself. It is written as its cells' + text, as a table inside a cell already was. +- **A block inside such a flattened table broke its holder's line.** It + stays on that line. +- **Bold ending in punctuation at a flattened cell's edge came back as raw + ``**, which read back as `**`. It stays Markdown bold. +- **Two `
` blocks ran together.** `
` is a block, as `
` + is, except inside `
`.
+- **Underline, highlight and inserted text lost their formatting.** ``,
+  `` and `` are kept as raw tags, as `` already was. Round a
+  block (a table, a list, a heading, a quote, a code block, a rule or
+  paragraphs) the tag is dropped and the blocks read as 0.6.0 read them.
+  Markdown has no inline tag round blocks: `easyvista-python-client` 0.4.0,
+  which kept it, showed a table as pipe text and a list or a heading as its
+  Markdown source, and round a `
` left a fence open to the end of the
+  body. Inside a table cell or a heading, whose blocks are one line, the tag
+  stays.
+- **A long ordered list was quadratic to read**: 5,000 items took 3.1 s,
+  and take 0.6 s now (CPython 3.12.3). An `
    ` such as `²` or `½` + counts from 1, as a browser counts a start holding no digit, instead of + failing the body. +- **A run of spaces or line breaks at the edge of bold, italic or a link was + quadratic to read.** It is linear. +- **A `colspan` or `start` markdownify cannot read as a number raised + `GlpiContentError`** from `.content` (`colspan="²"`, or 5,000 digits). The + body degrades to its text instead, every word kept. Any other + `ValueError` raised while converting takes the same path. +- **A lone surrogate in a write model's Markdown failed the whole write** + with `GlpiContentError`: cmark-gfm renders UTF-8, which cannot encode + one. Each lone surrogate is written as U+FFFD instead, as CommonMark + replaces a NUL, and the rest of the body keeps its Markdown. Calling + `GlpiContentConverter.to_transport` directly still raises, as + `easyvista-python-client`'s converter does. +- **A run of unfinished tags at the end of a body was quadratic to read** on + CPython before 3.11.14, 3.12.12 and 3.13.6 (CVE-2025-6069). Every `<` + after the body's last `>` is read as text, so `x ` document + read with its structure went from about 490 levels to about 326 (CPython + 3.12.3 and 3.13.14, default recursion limit). `
    ` stays at + about 194. The words are kept either way. +- **Dependency bounds** are now those of `easyvista-python-client[content]` + 0.4.0, so the two packages resolve to the same libraries side by side: + `markdown-it-py>=3.0,<4` (the alt-text fix relies on 3.x's `text_special` + tokens, and `mdformat` 0.7.22 requires `<4` anyway), + `markdownify>=1.2.3,<1.3` and `mdformat-tables>=1.0,<1.1`. + `mdformat>=0.7.22,<0.8` is unchanged. The reader overrides private + surfaces of `markdownify` and `mdformat`; `pyproject.toml` gives the + reason for each bound. + +### Documentation + +- The user guide's rich-text section now says what each element becomes, + that neither direction sanitises, what survives a round trip, and the + known holes, all reproduced on 2026-10-02. Its claim that a read always + displays the same and reads back as itself is now stated as the aim it + is, as in the API reference and the module docstring. +- The deep-nesting note gives the measured depths, and that a read started + within about 30 frames of the recursion limit can still raise. +- `GlpiContentError`'s docstring named python-markdown as the writer; it is + cmark-gfm. +- The `glpi-ticket-timeline`, `glpi-ticket-workflow` and + `glpi-knowledge-base` skills describe raw ``, ``, `` and + ``, the write-model rule exactly, and the new depth. +- Corrected after review: a write model keeps the caller's Markdown + stripped at both ends, not verbatim; the `glpi-knowledge-base` and + `glpi-plugin-fields` skills no longer say a deep body never raises; any + `ValueError` degrades to text, not only a `colspan` or `start`; a + `start` counts as a browser counts it only when it holds no digit; and + the CVE-2025-6069 note says that everything after a body's last `>`, + an unterminated comment or a cut-short closing tag included, is shown. + +### Tests + +- `content/tests/test_fixes.py` pins each of the fifteen fixes; run against + 0.6.0, 36 of the 54 tests first ported fail on CPython 3.12.3 and 37 on + 3.13.14, the rest being controls and pins. Section 11 also pins the + correction to fix 11 with 42 tests, 40 of which fail on + `easyvista-python-client` 0.4.0's converter. Reverting any one fix alone fails at least + one test of the content suite on both. +- `content/tests/test_properties.py` holds seeded generators to three + properties: a body displays what its HTML displayed, is a fixed point, + and keeps its words in order. Against 0.6.0 all 16 tests fail. +- `content/tests/test_cost.py` replaces the 20-second budgets on long bodies + with growth tests, `time(4n) / time(n) < 8`, which a quadratic pass fails + at a few thousand items. +- Four guards that no test caught are pinned, each by a test that fails + without it: a second `<` before a space keeps one escape, a header cell + escapes a `|` in `` or `` and splits a line break in inline + code, a table inside an `` without `href` stays a table, and an ordered + item numbered 10 or more indents its content by its bullet's width. A + count pins that the walk finding blocks inside ``, `` and `` + checks each tag about once. +- The display oracle treats a flattened cell's edge as a word boundary, as a + browser does, and reads `start` with `isdecimal`. + ## 0.6.0 — 2026-10-01 ### Changed (breaking) diff --git a/docs/api_reference.rst b/docs/api_reference.rst index 0d259cd..2716eb5 100644 --- a/docs/api_reference.rst +++ b/docs/api_reference.rst @@ -115,17 +115,23 @@ its page being read. The Markdown is CommonMark with GFM tables, rendered by cmark-gfm. It spells a body's text as literal text -- ``\_\_init\_\_``, ``\\\serveur``, -``\# pas un titre`` -- so rendering it displays what GLPI displayed and -reading that back gives the same Markdown. A write model keeps the caller's -Markdown verbatim unless it starts with an HTML tag. See -:ref:`content-conversion`. +``\# pas un titre`` -- so that rendering it displays what GLPI displayed and +reading that back gives the same Markdown. That is the aim; the user guide +lists the shapes where it falls short. Underline, highlight and struck text +stay as raw ````, ````, ```` and ````. Neither direction +sanitises: raw HTML and ``javascript:`` links pass through. A write model +keeps the caller's Markdown as written, stripped at both ends, unless it +starts with ``<`` and holds an HTML element, and sends a lone surrogate as +U+FFFD. See :ref:`content-conversion`. Very deeply nested HTML is the case worth knowing about. ``markdownify`` walks the document recursively and runs out of stack at a few hundred levels of nesting. The converter attempts the conversion and, if the walk does not fit, reads the body's text instead, a line per -block: the words survive, the structure does not. Anything else that goes -wrong in either direction raises :class:`GlpiContentError`. +block: the words survive, the structure does not. Any ``ValueError`` from +the conversion degrades the same way: markdownify raises one for a +``colspan`` or ``start`` it cannot read as a number. Anything else that +goes wrong in either direction raises :class:`GlpiContentError`. Because the conversion is cached on first read, a read model should be treated as immutable afterwards: assigning to ``content_html``, or diff --git a/docs/development.md b/docs/development.md index c8e5cfe..bae786b 100644 --- a/docs/development.md +++ b/docs/development.md @@ -86,13 +86,17 @@ python -m pytest articles (`content` and `description`) and article revisions. It is wired into the models by `models/api_schema/_content.py`, eagerly on the write models and through a cached property on the read ones. - Inbound HTML too deep for `markdownify` to walk is tag-stripped + Inbound HTML too deep for `markdownify` to walk, or on which it + raises `ValueError` (a `colspan` or `start` it cannot read as a + number, among others), is tag-stripped rather than parsed. The depth is not predicted: the conversion is attempted and the `RecursionError` answered, because the budget is the caller's remaining stack and no bound computed in advance can know it. - The module docstring carries the derivation and the rejected + The 0.5.0 changelog entry carries the derivation and the rejected alternatives (`sys.setrecursionlimit`, and a thread with a larger - stack, which needs the same global). The prohibition is enforced by + stack, which needs the same global). The reader is kept identical, but + for names, messages, docstrings and an import guard, to + `easyvista-python-client`'s; change the two together. The prohibition is enforced by `testing/tests/test_raise_site_audit.py`, not just written down. - `glpi_python_client.testing` exposes `make_client` and `make_async_client` factories that produce in-memory clients with no diff --git a/docs/installation.rst b/docs/installation.rst index c10cd41..660e109 100644 --- a/docs/installation.rst +++ b/docs/installation.rst @@ -6,7 +6,12 @@ Requirements ``glpi-python-client`` supports Python 3.11 and newer. Runtime dependencies are installed from the package metadata and include ``httpx``, ``tenacity``, -``beautifulsoup4``, ``lxml``, and ``pydantic``. +``beautifulsoup4``, ``lxml``, and ``pydantic``, plus the libraries the +rich-text converter is built on: ``markdownify``, ``mdformat``, +``mdformat-tables`` and ``markdown-it-py`` to read, and ``cmarkgfm`` to +write. The reader overrides private surfaces of ``markdownify`` and +``mdformat``, so those four carry upper bounds; ``pyproject.toml`` gives the +reason for each. Install from PyPI ----------------- diff --git a/docs/user_guide.rst b/docs/user_guide.rst index 585df9c..9498d7b 100644 --- a/docs/user_guide.rst +++ b/docs/user_guide.rst @@ -1758,11 +1758,13 @@ The whole page is built in one pass, so a single unconvertible record used to make its page-mates unreadable too. The failure is now scoped to the record whose body you actually read. -The Markdown is CommonMark with GFM tables. Rendering it -- as the package -does on the way back to GLPI, with cmark-gfm -- displays what GLPI -displayed, and reading that rendering back gives the same Markdown. Text in -a body is literal: Markdown has one spelling for ``__init__`` typed by a -user and for bold ``init``, so text that would read as syntax is escaped: +The Markdown is CommonMark with GFM tables. The aim is that rendering it -- +as the package does on the way back to GLPI, with cmark-gfm -- displays what +GLPI displayed, and that reading that rendering back gives the same +Markdown. It is an aim, not a guarantee for every body: `What survives a +round trip`_ lists where it falls short. Text in a body is literal: Markdown +has one spelling for ``__init__`` typed by a user and for bold ``init``, so +text that would read as syntax is escaped: .. code-block:: python @@ -1781,28 +1783,106 @@ back as ``\`` and a newline, a nested list is indented by its bullet's width, and a table comes back unpadded. A body with no HTML element is plain text and is read as GLPI shows it, its lines as lines. +More of what GLPI displays, and the Markdown it reads as: + +.. code-block:: text + + GLPI displays Markdown + ---------------------------- ---------------------------- + __init__ \_\_init\_\_ + *important* \*important\* + # 4521 (at a line start) \# 4521 + 1. pas une liste 1\. pas une liste + - pas une liste \- pas une liste + > pas une citation \> pas une citation + [lien](x) \[lien\](x) + \ + ~~~ (at a line start) \~~~ + Attention! then a link Attention\![lien](https://example.org) + +Text GLPI displays as markup is read back as that text, so ``<b>`` +reads as ``\``, never as a live ````. What each kind of element +becomes: + +* **Bold and italic** become ``**`` and ``*``, or raw ```` and + ```` where CommonMark would not close the markers, for example bold + that ends in punctuation directly followed by a letter: + ``prix:(10)euros`` reads as ``prix:(10)euros``. +* **Underline, highlight and inserted text** (````, ````, + ````) stay as those raw tags, and **struck text** (````, + ````, ````) as a raw ````. CommonMark has no spelling for + any of them, and writing passes the tags through, so the formatting + survives. The cost is raw HTML in the Markdown. Markdown has no inline tag + round blocks, so a ````, ```` or ```` holding a table, a + list, a heading, a quote, a ``
    ``, a rule or paragraphs is dropped and
    +  its blocks kept; inside a table cell or a heading, which hold one line, it
    +  stays.
    +* **Links** become ``[text](https://... "title")``, and a link whose text is
    +  its own URL, a pasted link, becomes the autolink ````. No
    +  link target is filtered, ``javascript:`` included (see `It is not a
    +  sanitiser`_). An **image** becomes ``![alt](src "title")``, its alt text
    +  escaped like any other text.
    +* **Lists** nest and keep their numbers, including an ``
      ``. A + ``start`` that is not a decimal number counts from 1, as a browser counts + one holding no digit, such as ``²``; a browser reads ``" 3"``, ``"+3"`` + or ``"3abc"`` as 3. **Tables** become GFM tables, and a ``
      `` inside a cell + stays a raw ``
      ``, since a GFM cell is one line. A table inside a cell, + a heading or a link is written as its cells' text, spaced apart. +* **Preformatted blocks** become fences that keep the ``language-`` class + cmark-gfm writes, with a fence longer than any run of backticks in the + code. **Inline code** holding a line break becomes one code span per line. +* ``
      `` and ``
      `` are blocks. ````, ``
      code
      ", + "
      • code
      ", + "
        • code
        ", + '
        • i
          code
        • ' + "
        ", + ], + "struck": [ + "

        ancien nouveau

        ", + "

        ancien nouveau

        ", + "

        ancien nouveau

        ", + ], + "hidden": [ + "

        Bonjour

        ", + "RE: Imprimante

        Bonjour

        " + "", + ], + "misreading": [ + "

        # pas un titre

        ", + "

        #4521 est un doublon

        ", + "

        2026. Une annee

        ", + ], + "upper-case": [ + "

        Bonjour

        ", + "
        Bonjour", + "
        Bonjour
        ", + ], +} + + +@pytest.mark.parametrize( + "html", + [ + pytest.param(html, id=f"{group}-{index}") + for group, bodies in EARLIER_REGRESSIONS.items() + for index, html in enumerate(bodies) + ], +) +def test_an_earlier_regression_body_survives(html: str) -> None: + assert_survives(html) + + @pytest.mark.parametrize( ("html", "expected"), [ @@ -191,7 +378,10 @@ def test_ordinary_prose_reads_back_as_typed(html: str, expected: str) -> None: def test_a_table_nested_in_a_cell_keeps_its_text() -> None: - """A signature laid out as a table inside a table keeps every word.""" + """A table inside a table keeps every word. + + E-mail signatures and notification templates are often laid out so. + """ html = ( "
        Signature
        " @@ -228,6 +418,65 @@ def test_a_fence_keeps_its_language() -> None: assert read(render(markdown)) == markdown +# --------------------------------------------------------------------------- +# A notification template: a layout table around a table of label/value cells +# --------------------------------------------------------------------------- +# +# From easyvista-python-client 0.4.0. E-mail templates are commonly laid out +# as an outer one-cell table around an inner one-row table: each label in +# bold over its value, spacer cells between them, and here one label in a +# cell of its own beside its value's cell. That shape exercises the +# flattening of a table nested in a cell, fix 9 of test_fixes.py (bold that +# ends on punctuation at a flattened cell's edge still closes -- 0.6.0 lost +# the fixed point there) and
        in a cell, all at once. Every word below +# is invented. + +_LABELS = ["Zorvane", "Plimet", "Quastor", "Brindel"] +_VALUES = ["trelm 4471", "oskar-vint", "maludi pref", "kobra 12/09"] + + +def _notification_template() -> str: + cells = [] + for label, value in zip(_LABELS, _VALUES, strict=True): + cells.append("") # a spacer cell + cells.append( + f"" + ) + cells.append("") + inner = "
        Jean Dupont
        {label}


        {value}

        Velusk :sanquo
        " + "".join(cells) + "
        " + footer = ( + 'Grovak le suivi ' + 'Trelune' + ) + return f"
        {inner}
        {footer}
        " + + +def test_a_notification_template_keeps_every_word_in_order() -> None: + """Every word, in order, each label still bold and the link still a link. + + The layout does not survive: GFM has no nested table, so the inner row + becomes one line of text in the outer cell, and a GFM table always has a + header row, so the outer table gains an empty one. Both are limits of + the format. Apart from that header row, the display is the same. + """ + + html = _notification_template() + + markdown = read(html) + rendered = render(markdown) + + assert text_words(rendered) == text_words(html) + name, rows = cast(tuple[str, tuple[object, ...]], displayed(html)[0]) + empty_header = ((True, ()),) # the header row GFM requires: one empty cell + assert displayed(rendered) == ((name, (empty_header, *rows)),) + for label in [*_LABELS, "Velusk :"]: + assert f"**{label}**" in markdown + assert "](https://example.org/suivi?ref=PX-7)" in markdown + assert "![Trelune](https://example.org/marque.png)" in markdown + assert "glpi-cell" not in markdown and "data-glpi-" not in markdown + assert read(rendered) == markdown + + # --------------------------------------------------------------------------- # Plain text and the write path # --------------------------------------------------------------------------- @@ -242,21 +491,38 @@ def test_a_fence_keeps_its_language() -> None: r"\\serveur\partage", "if x0", "Bonjour,\n\nMerci.", + pytest.param("Velk ondra\r\nPrastin", id="crlf"), ], ) def test_a_plain_text_body_reads_as_the_text_it_is(text: str) -> None: - """A body with no HTML element is literal text, its lines lines.""" + """A body with no HTML element is literal text, its lines lines. + + The ``crlf`` case: a CR LF separates lines as LF does. An HTML reading + of the same characters would show a space instead. + """ markdown = read(text) shown = displayed(render(markdown)) - expected = displayed( - "

        " + text.replace("<", "<").replace("\n", "
        ") + "

        " - ) + lines = text.replace("<", "<").splitlines() + expected = displayed("

        " + "
        ".join(lines) + "

        ") assert shown == expected assert read(render(markdown)) == markdown +def test_an_entity_in_a_plain_text_body_is_read_literally() -> None: + """With no element, `` `` is six characters of text, not a space. + + Pinned because it is the consequence of reading a body with no HTML + element as text, and the one a reviewer would most likely expect the + other way. + """ + + markdown = read("Trasvel ok") + + assert displayed(render(markdown)) == displayed("

        Trasvel&nbsp;ok

        ") + + @pytest.mark.parametrize( "markdown", [ @@ -278,6 +544,38 @@ def test_the_write_path_still_converts_an_html_document() -> None: ) +@pytest.mark.parametrize( + ("markdown", "expected"), + [ + pytest.param( + " puis
        suite **gras**", + "puis\\\nsuite \\*\\*gras\\*\\*", + id="autolink-first", + ), + pytest.param( + " puis Ctrl **gras**", + "puis `Ctrl` \\*\\*gras\\*\\*", + id="angle-text-first", + ), + ], +) +def test_markdown_opening_with_angle_brackets_and_holding_html_is_read_as_html( + markdown: str, expected: str +) -> None: + """A limit, pinned so that the documentation stays true. + + ``plain_text_is_markdown=True`` passes a value through unless it starts + with ``<`` and holds a real HTML element *anywhere*. So Markdown that + opens with an autolink or with angle-bracketed text, and carries inline + HTML further on, is read as HTML: the autolink, an unknown element to + the HTML parser, is dropped, and the Markdown syntax is escaped as text. + The write models' validator passes ``plain_text_is_markdown=True``, so + ``PostTicket(content=...)`` reads such a value the same way. + """ + + assert read(markdown, plain_text_is_markdown=True) == expected + + # --------------------------------------------------------------------------- # Markdown first # --------------------------------------------------------------------------- @@ -314,37 +612,3 @@ def test_caller_markdown_displays_the_same_after_the_round_trip(markdown: str) - assert displayed(render(back)) == displayed(html) assert read(render(back)) == back - - -# --------------------------------------------------------------------------- -# Cost and depth -# --------------------------------------------------------------------------- - - -@pytest.mark.parametrize( - "html", - [ - pytest.param("

        " + "ligne
        " * 20_000 + "

        ", id="line-breaks"), - pytest.param("
          " + "
        • x
        • " * 20_000 + "
        ", id="list-items"), - pytest.param("
          " + "
        • " * 20_000 + "
        ", id="empty-items"), - pytest.param("

        " + "gras mot " * 20_000 + "

        ", id="emphasis"), - ], -) -def test_a_long_body_converts_in_linear_time(html: str) -> None: - """The budget is generous on purpose: this guards the complexity, not the speed.""" - - started = time.perf_counter() - read(html) - - assert time.perf_counter() - started < 20 - - -def test_a_body_too_deep_to_convert_keeps_its_text() -> None: - """markdownify recurses per nesting level; past the stack, the text is kept.""" - - html = "
        " * 3000 + "

        __init__ au fond

        " + "
        " * 3000 - - markdown = read(html) - - assert "init" in markdown - assert "au fond" in render(markdown) diff --git a/glpi_python_client/models/api_schema/_content.py b/glpi_python_client/models/api_schema/_content.py index 50d0a5c..e1c3f46 100644 --- a/glpi_python_client/models/api_schema/_content.py +++ b/glpi_python_client/models/api_schema/_content.py @@ -54,7 +54,8 @@ ``BeforeValidator`` below fires when a caller constructs a ``Post*`` model, so caller-authored Markdown passes through it before the serializer renders it. It runs with ``plain_text_is_markdown=True``, which keeps the caller's -Markdown verbatim unless it starts with an HTML tag -- reading it as literal +Markdown as written, stripped at both ends, unless it starts with ``<`` and +holds an HTML element -- reading it as literal text would escape it, and GLPI would receive literal asterisks. One sharp edge comes with the read side, from ``functools.cached_property``: @@ -83,6 +84,7 @@ from __future__ import annotations +import re from collections.abc import Iterator from contextlib import contextmanager from contextvars import ContextVar @@ -159,6 +161,9 @@ def markdown_view(raw: str | None) -> str | None: "glpi_last_content_fault", default=None ) +#: A UTF-16 surrogate standing alone, which UTF-8 cannot encode. +_LONE_SURROGATE = re.compile("[\ud800-\udfff]") + def _to_transport(value: str | None) -> str | None: """Render an outbound Markdown content value as the HTML GLPI expects. @@ -167,6 +172,12 @@ def _to_transport(value: str | None) -> str | None: drop unset fields from request bodies. Empty Markdown is rendered as an empty string to stay consistent with the inbound converter behaviour. + A lone surrogate is written as U+FFFD, the way CommonMark writes a NUL: + cmark-gfm renders UTF-8, which cannot carry one, so the whole write + failed on a single character. Python strings can hold one -- decoding + with ``surrogateescape``, or ``json.loads`` of an unpaired ``\\ud800`` + escape -- and the rest of the body keeps its Markdown. + A failure is recorded in :data:`_LAST_CONTENT_FAULT` before it is raised, because raising it is not enough on its own: see :func:`restoring_content_faults`. @@ -175,7 +186,7 @@ def _to_transport(value: str | None) -> str | None: if value is None: return None try: - return GlpiContentConverter.to_transport(value) + return GlpiContentConverter.to_transport(_LONE_SURROGATE.sub("\ufffd", value)) except GlpiContentError as exc: _LAST_CONTENT_FAULT.set(exc) raise diff --git a/glpi_python_client/models/api_schema/tests/test_content.py b/glpi_python_client/models/api_schema/tests/test_content.py index 52892e9..7a79a5e 100644 --- a/glpi_python_client/models/api_schema/tests/test_content.py +++ b/glpi_python_client/models/api_schema/tests/test_content.py @@ -114,6 +114,46 @@ def failing_renderer(markdown: str) -> str: assert isinstance(caught.value.__context__, PydanticSerializationError) +@pytest.mark.parametrize( + ("markdown", "html"), + [ + pytest.param("a\ud800b", "

        a\ufffdb

        ", id="high"), + pytest.param( + "**gras** a\udc00", "

        gras a\ufffd

        ", id="low" + ), + pytest.param( # built with chr(): \ud83d\ude00 in a literal is one character + chr(0xD83D) + chr(0xDE00), "

        \ufffd\ufffd

        ", id="a-pair-kept-apart" + ), + ], +) +def test_a_lone_surrogate_is_written_as_a_replacement_character( + markdown: str, html: str +) -> None: + """UTF-8 cannot encode a lone UTF-16 surrogate, so cmark-gfm could not + render a body holding one and the whole write failed with + ``GlpiContentError``. Each surrogate is written as U+FFFD instead, as + CommonMark replaces a NUL, and the rest of the body keeps its Markdown; + a pair left as two code points is two of them. The write model's own + value is left as the caller gave it.""" + + model = PostTicket(name="surrogate", content=markdown) + + assert model.content == markdown + assert model_to_payload(model)["content"] == html + + +def test_a_lone_surrogate_reaches_glpi_as_a_replacement_character() -> None: + """The same through the client: the write is sent, not refused.""" + + client = make_client() + recorder = TransportRecorder() + recorder.install(client) + + client.create_ticket(PostTicket(name="surrogate", content="a\ud800b")) + + assert recorder.calls[-1]["json"]["content"] == "

        a\ufffdb

        " + + def test_a_serialisation_fault_that_is_not_content_is_not_mislabelled() -> None: """Only a stashed content fault becomes a ``GlpiContentError``. diff --git a/pyproject.toml b/pyproject.toml index 62f048e..b92f8f7 100644 --- a/pyproject.toml +++ b/pyproject.toml @@ -35,7 +35,7 @@ exclude = [ [project] name = "glpi-python-client" -version = "0.6.0" +version = "0.6.1" description = "A typed Python client for GLPI ITSM APIs." readme = "README.md" requires-python = ">=3.11" @@ -64,11 +64,27 @@ dependencies = [ "cmarkgfm>=2025.10", "httpx>=0.28", "lxml>=4.9", - "markdown-it-py>=3.0", - "markdownify>=1.2", - # The reader extends mdformat's renderer, an API that may move at 0.8. + # The four bounds below are easyvista-python-client[content] 0.4.0's, + # whose reader is this one, so the two packages resolve to the same + # libraries when installed side by side. The reader overrides private + # surfaces of markdownify and mdformat, so raise a bound only with the + # content tests re-run against the new release. + # + # Keeping an image's alt-text escapes relies on markdown-it-py 3's + # text_special tokens. mdformat 0.7.22 itself requires markdown-it-py<4, + # and easyvista-python-client measured 4.x ending a ragged table early. + "markdown-it-py>=3.0,<4", + # The reader overrides markdownify's process_tag, escape and convert_* + # methods and relies on its _inline and _noformat pseudo-tags. 1.2.3 is + # easyvista-python-client's floor; this package's content tests pass on + # 1.2.2 and 1.2.3. The cap is a precaution against a minor release that + # moves those surfaces. + "markdownify>=1.2.3,<1.3", + # The reader extends mdformat's renderer, an API that may move at 0.8, + # and mdformat-tables 1.0 requires mdformat<0.8 too. "mdformat>=0.7.22,<0.8", - "mdformat-tables>=1.0", + # GFM tables for mdformat. The cap is a precaution, as for markdownify. + "mdformat-tables>=1.0,<1.1", "pydantic>=2.8", # Never imported by this package, and deliberately so. httpcore decides # whether it is running under asyncio or trio by probing for sniffio on diff --git a/skills/glpi-asset-workflow/SKILL.md b/skills/glpi-asset-workflow/SKILL.md index e7b32f9..1ab8218 100644 --- a/skills/glpi-asset-workflow/SKILL.md +++ b/skills/glpi-asset-workflow/SKILL.md @@ -5,7 +5,7 @@ license: MIT compatibility: "Requires Python 3.11+, glpi-python-client, network access to the GLPI v2 API, and credentials allowed to read or write assets." metadata: package: glpi-python-client - version: "0.6.0" + version: "0.6.1" --- # GLPI Asset Workflow diff --git a/skills/glpi-client-setup/SKILL.md b/skills/glpi-client-setup/SKILL.md index 367709a..43e4ae6 100644 --- a/skills/glpi-client-setup/SKILL.md +++ b/skills/glpi-client-setup/SKILL.md @@ -5,7 +5,7 @@ license: MIT compatibility: "Requires Python 3.11+, glpi-python-client, network access to a GLPI v2 API, and valid GLPI credentials." metadata: package: glpi-python-client - version: "0.6.0" + version: "0.6.1" --- # GLPI Client Setup diff --git a/skills/glpi-contract-workflow/SKILL.md b/skills/glpi-contract-workflow/SKILL.md index d0498a3..8c8381e 100644 --- a/skills/glpi-contract-workflow/SKILL.md +++ b/skills/glpi-contract-workflow/SKILL.md @@ -5,7 +5,7 @@ license: MIT compatibility: "Requires Python 3.11+, glpi-python-client, network access to the GLPI v2 API, and credentials allowed to read or write contracts." metadata: package: glpi-python-client - version: "0.6.0" + version: "0.6.1" --- # GLPI Contract Workflow diff --git a/skills/glpi-document-workflow/SKILL.md b/skills/glpi-document-workflow/SKILL.md index 3c119f3..8d22db8 100644 --- a/skills/glpi-document-workflow/SKILL.md +++ b/skills/glpi-document-workflow/SKILL.md @@ -5,7 +5,7 @@ license: MIT compatibility: "Requires Python 3.11+, glpi-python-client, network access to the GLPI v2 API, and v1 credentials configured on the client for binary uploads." metadata: package: glpi-python-client - version: "0.6.0" + version: "0.6.1" --- # GLPI Document Workflow diff --git a/skills/glpi-knowledge-base/SKILL.md b/skills/glpi-knowledge-base/SKILL.md index 10875e1..57196a5 100644 --- a/skills/glpi-knowledge-base/SKILL.md +++ b/skills/glpi-knowledge-base/SKILL.md @@ -5,7 +5,7 @@ license: MIT compatibility: "Requires Python 3.11+, glpi-python-client, network access to the GLPI v2 API, and — for category writes only — a legacy v1 session (v1_base_url + v1_user_token)." metadata: package: glpi-python-client - version: "0.6.0" + version: "0.6.1" --- # GLPI Knowledge Base @@ -165,7 +165,7 @@ await client.delete_kb_article(42, force=True) - KB write models use `IdNameRef` for every foreign key -- `categories[]`, `entity`, `user`, `parent` -- not `IdRef`. Passing `IdRef(id=4)` raises a pydantic `ValidationError`. (`GetKBArticleComment.parent` is the one KB field genuinely typed `IdRef`, and it is read-only.) - Article `content`/`description` and revision `content` are Markdown on the Python side and HTML on the wire; the conversion is automatic, so never author HTML. Comment `comment` is a plain `str` with no conversion at all -- the inconsistency is real, not an omission here. - On the **read** models the conversion is lazy. `GetKBArticle` stores `content_html`/`description_html` (the server's HTML verbatim, also accepting the wire spellings `content`/`description` on construction) and exposes `content`/`description` as cached properties that convert on first read; `GetKBArticleRevision` is the same with `content_html`. Reading `.content` is unchanged, but `search_kb_articles` and `list_kb_article_revisions` now convert nothing until a body is read -- which matters most here, because article search returns whole bodies and revision history returns every past one. Write models (`PostKBArticle`/`PatchKBArticle`) still convert eagerly and keep the plain `content`/`description` field names. -- HTML too deeply nested for the converter to walk is **stripped to text instead of converted** -- every character the normal rendering would have produced, none of the Markdown structure, and no exception. The conversion is attempted rather than the depth predicted, so the limit is whatever stack is left at the call site (about 494 levels from a shallow one). Any other conversion failure raises `GlpiContentError`, from the `.content` read rather than from the search call. +- HTML too deeply nested for the converter to walk is **stripped to text instead of converted** -- every character the normal rendering would have produced, none of the Markdown structure. Nothing raises unless the read starts within about 30 frames of the recursion limit. The conversion is attempted rather than the depth predicted, so the limit is whatever stack is left at the call site (about 326 `
        ` levels from a shallow one, measured on CPython 3.12 and 3.13). Any `ValueError` from the conversion degrades the same way: markdownify raises one for a `colspan` or `start` it cannot read as a number. Any other conversion failure raises `GlpiContentError`, from the `.content` read rather than from the search call. - `force` on `delete_kb_article`, `delete_kb_category` and `delete_kb_article_comment` is keyword-only and is serialised into the JSON request **body** via the matching `Delete*` model, not sent as a query parameter. `force=True` deletes permanently; omitting it or passing `False` moves the record to the GLPI trash. - Revisions are read-only: there is no create/update/delete helper and no Post/Patch/Delete revision model. A revision appears as a side effect of updating an article. `get_kb_article_revision(article_id, revision)` takes the revision **number** (`GetKBArticleRevision.revision`), not the row `id` -- the two differ on the model. - `GetKBArticle.revisions` and `.translations` hold two *different* private ref classes (leading underscore, not exported from the package root). Read their attributes; never import them. The two field sets are not interchangeable: a `revisions` entry has `.id`, `.revision`, `.language`, `.date`, while a `translations` entry has `.id`, `.language`, `.name`. `revisions[0].name` raises `AttributeError` -- the models allow unknown keys from the server, but that does not synthesise an attribute that was never sent. diff --git a/skills/glpi-plugin-fields/SKILL.md b/skills/glpi-plugin-fields/SKILL.md index ce8a0ec..39009a8 100644 --- a/skills/glpi-plugin-fields/SKILL.md +++ b/skills/glpi-plugin-fields/SKILL.md @@ -5,7 +5,7 @@ license: MIT compatibility: "Requires Python 3.11+, glpi-python-client, the GLPI Fields plugin installed server-side, and a legacy v1 session (v1_base_url + v1_user_token) — every method in this family goes over the v1 API." metadata: package: glpi-python-client - version: "0.6.0" + version: "0.6.1" --- # GLPI Plugin Fields @@ -191,7 +191,7 @@ async def upsert_plugin_field_row( - A container that has never had a value saved for that ticket is **silently absent** from the `get_ticket_custom_fields` result -- you do not get `{"container": {}}`. Use `result.get(name, {})`, never `result[name]`. An empty overall dict is **ambiguous**: `get_ticket_custom_fields` builds its result by skipping every container with no persisted row, so `{}` comes back both when the instance declares no Ticket containers at all and when it declares several but this ticket has saved nothing in any of them. The return value cannot tell you which -- call `list_plugin_fields_containers(itemtype="Ticket")` if you need to know. - Both discovery listings fetch **one fixed page, `range=0-999`**, with no pagination and no server-side filtering. The `itemtype` and `container_id` narrowing happens client-side *after* that cap, so an instance with more than 1000 containers or field declarations silently loses the tail -- possibly including the container you are looking for. - Neither listing filters on `is_active`, so **disabled containers and disabled or read-only fields come back from discovery looking exactly like live ones**. `set_ticket_custom_fields` accepts a field whose `is_readonly` is `True`, because its guard only checks that the name is declared. Check `container.is_active`, `field.is_active` and `field.is_readonly` yourself. -- Values are transmitted **verbatim -- there is no HTML/Markdown conversion on this path**, unlike ticket and KB article content. A `richtext` field takes raw HTML (`"

        test

        "`). Convert yourself if you want Markdown: `GlpiContentConverter` is not exported from the package root, import it from `glpi_python_client.content`. Calling it directly puts you on the same two rules the models get: HTML too deeply nested for the converter to walk comes back **tag-stripped rather than converted** (every character the normal rendering would have produced, none of the structure, no exception), and any other failure raises `GlpiContentError`. +- Values are transmitted **verbatim -- there is no HTML/Markdown conversion on this path**, unlike ticket and KB article content. A `richtext` field takes raw HTML (`"

        test

        "`). Convert yourself if you want Markdown: `GlpiContentConverter` is not exported from the package root, import it from `glpi_python_client.content`. Calling it directly puts you on the same two rules the models get: HTML too deeply nested for the converter to walk comes back **tag-stripped rather than converted** (every character the normal rendering would have produced, none of the structure; nothing raises unless the call starts within about 30 frames of the recursion limit), and so does a body on which the conversion raises `ValueError`, such as a `colspan` markdownify cannot read as a number. Any other failure raises `GlpiContentError`, and `to_transport` raises on a lone surrogate, which the write models send as U+FFFD. - The two convenience helpers are **Ticket-only** -- the itemtype is hardcoded. There is no `get_item_custom_fields` and no `set_item_custom_fields`. For Computer, Problem, Change and the rest, drive the generic row helpers. - Parameter-passing style is inconsistent across the family, and only one half of it is a rule. `list_item_plugin_field_rows(itemtype, items_id, container_name)` declares ordinary `POSITIONAL_OR_KEYWORD` parameters, so both `("Ticket", 1234, "extrainfo")` and `(itemtype="Ticket", items_id=1234, container_name="extrainfo")` are legal -- the examples above pass them positionally by choice, not by requirement. `create_item_plugin_field_row` and `update_item_plugin_field_row` are genuinely **keyword-only** (`*` in the signature): calling either writer positionally is a `TypeError`. - Error taxonomy on the write path: an unknown container name or an unknown field name raises `GlpiValidationError`; a container the server returned without an `id`, or a create whose v1 reply carries no numeric row id, raises `GlpiProtocolError`. Both inherit `ValueError`. A missing v1 session raises a plain `RuntimeError`, and a non-success v1 status raises `GlpiStatusError` (with `.status_code`, `.url` and `.response_text`). diff --git a/skills/glpi-reporting-and-context/SKILL.md b/skills/glpi-reporting-and-context/SKILL.md index 774fbb2..0210ee0 100644 --- a/skills/glpi-reporting-and-context/SKILL.md +++ b/skills/glpi-reporting-and-context/SKILL.md @@ -5,7 +5,7 @@ license: MIT compatibility: "Requires Python 3.11+, glpi-python-client, network access to the GLPI v2 API, and credentials allowed to read tickets, tasks, users, entities, and timeline records." metadata: package: glpi-python-client - version: "0.6.0" + version: "0.6.1" --- # GLPI Reporting And Context diff --git a/skills/glpi-team-members/SKILL.md b/skills/glpi-team-members/SKILL.md index d6b811e..5018630 100644 --- a/skills/glpi-team-members/SKILL.md +++ b/skills/glpi-team-members/SKILL.md @@ -5,7 +5,7 @@ license: MIT compatibility: "Requires Python 3.11+, glpi-python-client, network access to the GLPI v2 API, and credentials allowed to manage ticket teams." metadata: package: glpi-python-client - version: "0.6.0" + version: "0.6.1" --- # GLPI Team Members diff --git a/skills/glpi-ticket-timeline/SKILL.md b/skills/glpi-ticket-timeline/SKILL.md index 02be123..5a8b0f5 100644 --- a/skills/glpi-ticket-timeline/SKILL.md +++ b/skills/glpi-ticket-timeline/SKILL.md @@ -5,7 +5,7 @@ license: MIT compatibility: "Requires Python 3.11+, glpi-python-client, and network access to the GLPI v2 API." metadata: package: glpi-python-client - version: "0.6.0" + version: "0.6.1" --- # GLPI Ticket Timeline @@ -120,7 +120,7 @@ await client.update_ticket_timeline_document( - `create_*` methods return new identifiers as plain `int`. `update_*` and `delete_*`/`unlink_*` return `None`. - Three enums carry this family's value vocabularies, all exported from `glpi_python_client` and all subclasses of `GlpiEnum` (itself an `IntEnum`, so a member serialises as its number and compares equal to one): `GlpiTaskState` on `PostTicketTask.state`/`PatchTicketTask.state` (`INFORMATION = 0`, `TODO = 1`, `DONE = 2` -- note `INFORMATION` is `0`, so `if task.state:` is false for it; test against `None`), `GlpiSolutionStatus` on `PostSolution.status`/`PatchSolution.status` (`NONE = 1`, `WAITING = 2`, `ACCEPTED = 3`, `REFUSED = 4`), and `GlpiTimelinePosition` on `timeline_position` (`INVALID = -1`, `NONE = 0`, `LEFT = 1`, `RIGHT = 2`, `LEFT_BIG = 3`, `RIGHT_BIG = 4`), which the followup, task and document models carry -- the solution models do not have the field at all. - Timeline `content` fields are Markdown on the Python side, not HTML. `PostFollowup`/`PostTicketTask`/`PostSolution` render Markdown to GLPI's HTML on serialisation. On `GetFollowup`/`GetTicketTask`/`GetSolution` the field is `content_html` -- the server's HTML verbatim -- and `content` is a cached property that converts it on **first read**, not on validation (`GetFollowup(content="

        Hello world

        ").content == "Hello **world**"` still holds; the wire spelling `content` is accepted as a validation alias). So `record.content` is always Markdown, but `list_ticket_followups` converts nothing until you read a body, and a body that cannot be converted no longer breaks the whole list. Authoring raw HTML on a write model is not an error but is round-tripped through the Markdown converter and can be reshaped; write Markdown. -- **`.content` is CommonMark that spells literal text so it stays text.** Rendering it (the package uses cmark-gfm, with GFM tables and a newline as a line break) displays what GLPI displayed: a user's `__init__` reads back as `\_\_init\_\_`, `\\serveur` as `\\\serveur`, `# titre` at a line start as `\# titre`. Ordinary prose -- `fichier_de_test_v2.xlsx`, `C:\Temp`, `R&D` -- comes back as typed. A line break reads back as `\` and a newline. Render the Markdown to display it; do not strip the backslashes. Markdown you write is kept verbatim unless it starts with an HTML tag; put a placeholder such as `` in backticks, since raw HTML passes through to GLPI. -- HTML too deeply nested to walk is **stripped to text instead of converted**: `markdownify` recurses per level and runs out of stack a few hundred levels deep. The conversion is attempted rather than the depth predicted, so the real limit is whatever stack is left at the call site. Every character the normal rendering would have produced still appears, but structure does not: link targets, image alt text and code fencing are gone. Nothing raises. Any other conversion failure raises `GlpiContentError` -- from the `.content` read, not from the `list_*` call. Note that `GlpiTicketContext.to_markdown()` is usually the first thing to read every body, so it is where such an error surfaces. +- **`.content` is CommonMark that spells literal text so it stays text.** Rendering it (the package uses cmark-gfm, with GFM tables and a newline as a line break) displays what GLPI displayed: a user's `__init__` reads back as `\_\_init\_\_`, `\\serveur` as `\\\serveur`, `# titre` at a line start as `\# titre`. Ordinary prose -- `fichier_de_test_v2.xlsx`, `C:\Temp`, `R&D` -- comes back as typed. A line break reads back as `\` and a newline. Render the Markdown to display it; do not strip the backslashes. Underline, highlight and strike have no Markdown spelling and stay as raw ``, ``, `` and ``; a ``, `` or `` wrapped round a block (a table, a list, a paragraph) is dropped and the block kept, since Markdown has no inline tag round blocks. Markdown you write is kept as written, stripped at both ends, unless it starts with `<` and holds an HTML element; put a placeholder such as `` in backticks, since raw HTML passes through to GLPI. The converter is not a sanitiser: `