From 10b08ee4f14e769c7a65d23aed6e1a0cab0fc710 Mon Sep 17 00:00:00 2001 From: baraline Date: Wed, 30 Sep 2026 18:16:24 +0200 Subject: [PATCH 01/11] feat(content): port glpi_python_client's Markdown <-> memo HTML converter Adds easyvista_python_client.content.EasyvistaContentConverter, with two static methods: from_transport reads an EasyVista memo (rich-text HTML or plain text) as canonical Markdown, and to_transport renders Markdown as HTML. It is a port of glpi_python_client's content/conversion.py at 0d43528, helper for helper: markdownify inbound, python-markdown outbound, the same options and extensions, the RecursionError / ParserRejectedMarkup fallback that degrades a too-deep body to its text instead of raising, the parser-driven scan behind it, and the
rewrite. The module docstring says the two files should move together. Why here: until now this package could only strip HTML to plain text, so a downstream GLPI-to-EasyVista sync wrote its own converter. On 2026-09-30 a GLPI description holding a pasted URL and a titled link reached EasyVista with neither link clickable (tier 4, one instance: EasyVista stored exactly the HTML it was sent). That converter escaped the pivot's < a second time, fused a link title into the href, cut a URL at its first ")" and read two lone asterisks as emphasis. Markdown to ITSM-format conversion belongs in the connector library, and the GLPI library's converter already handled every one of those cases. Each is pinned here as a regression, with the host replaced by example.org. EasyvistaContentError joins the core taxonomy under EasyvistaError, raised for any converter fault other than depth, and is exported at top level so it can be caught without the extra. The one code divergence: the three libraries are an optional extra, easyvista-python-client[content], rather than hard dependencies, so the core stays httpx + pydantic + tenacity. Importing the subpackage without them raises an ImportError naming the pip command, and a fresh-interpreter test pins that `import easyvista_python_client` loads none of them. The three are repeated in dev (CI runs the tests) and docs (autodoc imports the module); a test fails if the copies drift. Re-measured rather than copied, 2026-09-30: the inbound depth cliff is 493 levels on CPython 3.12-3.14 but 328 on 3.10, which spends about three frames per level; and the beautifulsoup4
/
defect no longer reproduces on 4.15.0. The rewrite stays, since the extra accepts 4.12, and a new test checks it is still applied, which no output test can notice on 4.15. Co-Authored-By: Claude Opus 5.5 (1M context) --- docs/api_reference.rst | 2 + easyvista_python_client/__init__.py | 2 + easyvista_python_client/content/__init__.py | 15 + easyvista_python_client/content/conversion.py | 836 ++++++++++ .../content/tests/__init__.py | 0 .../content/tests/test_conversion.py | 1457 +++++++++++++++++ easyvista_python_client/exceptions.py | 24 + .../testing/test_public_api.py | 97 ++ .../tests/test_exceptions.py | 14 + pyproject.toml | 20 + 10 files changed, 2467 insertions(+) create mode 100644 easyvista_python_client/content/__init__.py create mode 100644 easyvista_python_client/content/conversion.py create mode 100644 easyvista_python_client/content/tests/__init__.py create mode 100644 easyvista_python_client/content/tests/test_conversion.py diff --git a/docs/api_reference.rst b/docs/api_reference.rst index f756070..d331ac0 100644 --- a/docs/api_reference.rst +++ b/docs/api_reference.rst @@ -154,6 +154,8 @@ Exceptions .. autoexception:: easyvista_python_client.exceptions.EasyvistaConnectionError +.. autoexception:: easyvista_python_client.exceptions.EasyvistaContentError + Resource engine --------------- diff --git a/easyvista_python_client/__init__.py b/easyvista_python_client/__init__.py index a573444..0e4235f 100644 --- a/easyvista_python_client/__init__.py +++ b/easyvista_python_client/__init__.py @@ -16,6 +16,7 @@ from .exceptions import ( EasyvistaAuthError, EasyvistaConnectionError, + EasyvistaContentError, EasyvistaError, EasyvistaNotFound, EasyvistaRateLimitError, @@ -63,6 +64,7 @@ "EasyvistaClient", "EasyvistaConfig", "EasyvistaConnectionError", + "EasyvistaContentError", "EasyvistaError", "EasyvistaNotFound", "EasyvistaRateLimitError", diff --git a/easyvista_python_client/content/__init__.py b/easyvista_python_client/content/__init__.py new file mode 100644 index 0000000..ea0c022 --- /dev/null +++ b/easyvista_python_client/content/__init__.py @@ -0,0 +1,15 @@ +"""Markdown <-> EasyVista memo HTML conversion: the optional ``content`` extra. + +Install it with ``pip install "easyvista-python-client[content]"``. Importing +this subpackage without the extra raises :class:`ImportError` naming that +command. Nothing else in the package imports it, so ``import +easyvista_python_client`` needs none of the extra's dependencies, and +:class:`~easyvista_python_client.EasyvistaContentError` -- the one error the +converter raises -- lives in the core package, catchable either way. +""" + +from __future__ import annotations + +from easyvista_python_client.content.conversion import EasyvistaContentConverter + +__all__ = ["EasyvistaContentConverter"] diff --git a/easyvista_python_client/content/conversion.py b/easyvista_python_client/content/conversion.py new file mode 100644 index 0000000..1dc9a09 --- /dev/null +++ b/easyvista_python_client/content/conversion.py @@ -0,0 +1,836 @@ +"""Content conversion between EasyVista memo HTML and canonical Markdown. + +**This module is a port.** It is ``glpi_python_client/content/conversion.py`` +from `glpi_python_client `_ +at commit ``0d43528``, and **the two should move together**: a defect fixed in +one is almost certainly present in the other, because the hard part of both is +the behaviour of the same three libraries -- ``beautifulsoup4``'s +``html.parser`` tree builder and ``markdownify`` inbound, ``python-markdown`` +outbound -- and not anything either ITSM does. The code is that file's, helper +for helper, with the converter and its error renamed for this package +(``GlpiContentConverter`` is :class:`EasyvistaContentConverter`, +``GlpiContentError`` is :class:`~easyvista_python_client.EasyvistaContentError`) +and the error messages naming EasyVista. The one other change carries a comment +starting "Diverges from glpi_python_client", with the reason. Prose that stated +something true only of GLPI now says whose measurement it was, and the figures +that depend on the interpreter or on a library version were measured again +here. + +What it converts +---------------- + +EasyVista keeps rich text in *memo* fields -- a ticket's ``COMMENT`` and +``DESCRIPTION``, an action's ``DESCRIPTION`` -- which the client reads with +``resolve_memo``. A memo holds the HTML it was sent: measured 2026-09-30 on +one instance, a memo written through the API was stored byte for byte, and +the web UI rendered its ``

`` elements as paragraphs. That is tier 4 -- one +instance, one date, and it may not generalise; no vendor documentation of the +memo format is recorded in ``docs/vendor-api-reference.md``. Nothing forces a +memo to be HTML either, since a caller can write plain text, so +:meth:`EasyvistaContentConverter.from_transport` takes both: HTML is converted +and plain text comes back as it was. + +Neither direction sanitises, exactly as in glpi_python_client. Raw HTML +inside the Markdown passes through :meth:`~EasyvistaContentConverter.to_transport` +untouched, a ``javascript:`` link target is rendered as a live ``href``, and +inbound, text that a memo *displays* as markup -- ``<script>`` -- comes +back as a raw ``", True, id="script-body"), + pytest.param("", True, id="style-body"), + pytest.param("", True, id="marked-section"), + pytest.param("", True, id="unterminated-marked-section"), + pytest.param("", True, id="processing-instruction"), + ], +) +def test_the_degraded_path_keeps_exactly_what_the_converter_keeps( + construct: str, kept: bool +) -> None: + """Parity, construct by construct, and not one of these was a guess. + + Each expectation here was read off the converting path rather than + reasoned about, and three came back the opposite way round from the + obvious answer -- a ``x", id="in-a-script-body"), + pytest.param("x", id="in-a-comment"), + pytest.param("

a
b

", id="slash-not-abutting-gt"), + ], +) +def test_the_void_rewrite_leaves_everything_else_alone(html: str) -> None: + """The rewrite is confined to void tags in real tag position. + + ``
`` is left as it is -- rewriting it would change what the + document means, and it cannot be affected anyway, since only a void + name is ever recorded as already closed. The last case is the one + worth pinning: ``
`` reaches the parser as an ordinary start + tag, because its ``/`` does not abut the ``>``, so it never takes the + path that loses text and needs no rewriting. + """ + + assert conversion._canonicalise_void_elements(html) == html + + +def test_the_void_rewrite_rewrites_a_self_closed_void_tag() -> None: + """The positive case, pinned on the rewrite itself. + + On ``beautifulsoup4`` 4.15.0 the defect no longer reproduces, so the + conversion tests above pass whether or not the rewrite ran; only this + one notices if it stops running. + """ + + rewritten = conversion._canonicalise_void_elements( + "

a
b
cd

" + ) + + assert rewritten == "

a
b
cd

" + + +def test_markdownify_is_handed_the_rewritten_document( + monkeypatch: pytest.MonkeyPatch, +) -> None: + """The rewrite is wired in, and not only written. + + The test above pins the rewrite; this pins that ``from_transport`` + applies it. On ``beautifulsoup4`` 4.15.0 dropping the call changes no + output, so an output test cannot notice it -- what ``markdownify`` is + given can. + """ + + seen: list[str] = [] + + def _record(html: str, **options: object) -> str: + seen.append(html) + return "recorded" + + monkeypatch.setattr(conversion, "html_to_markdown", _record) + + EasyvistaContentConverter.from_transport("

a
b
c

") + + assert seen == ["

a
b
c

"] + + +def test_the_void_rewrite_changes_nothing_for_one_spelling_alone() -> None: + """A body that picks a spelling and keeps it converts exactly as before. + + The rewrite exists to remove an asymmetry between two spellings of the + same node, so it must be invisible to every body that does not mix + them. glpi_python_client measured that over 4000 fuzzed documents of + each spelling: not one output moved. + """ + + bare = "

a
b


c

" + slashed = "

a
b


c

" + expected = "a \nb![]()\n\n---\n\nc" + + assert EasyvistaContentConverter.from_transport(bare) == expected + assert EasyvistaContentConverter.from_transport(slashed) == expected + + +def test_both_paths_agree_on_a_body_using_both_spellings() -> None: + """The degraded path keeps this text, and so must the converting one. + + In glpi_python_client this body was the one place where the fallback + said *more* than the conversion it stands in for, which is the wrong + way round for a fallback and was how the defect was noticed at all. + """ + + shallow = "

one
two
three

" + deep = "
" * 600 + shallow + "
" * 600 + + converted = EasyvistaContentConverter.from_transport(shallow) + degraded = EasyvistaContentConverter.from_transport(deep) + + for word in ("one", "two", "three"): + assert word in converted + assert word in degraded + + +@pytest.mark.parametrize( + "fragment", + [ + pytest.param('', id="end-tag-with-a-quoted-attribute"), + pytest.param('
', id="name-less-equals-quote"), + pytest.param('

', id="quoted-gt-then-tag"), + ], +) +def test_a_misread_tag_end_does_not_swallow_the_body_after_it(fragment: str) -> None: + """Scaled past the cliff: degrades quietly, and keeps its words. + + Reading an end tag with start-tag rules consumed everything up to the + next quote, which deleted prose at any depth and under-counted the + nesting 1:1 with the repetition back when the nesting was predicted. + """ + + html = fragment * 600 + "the printer is offline" + + assert "the printer is offline" in EasyvistaContentConverter.from_transport(html) + + +def test_a_misread_tag_end_does_not_delete_prose() -> None: + """The other half of the same defect, and it needs no depth at all. + + Reading an end tag with attribute rules consumed everything between + the opening quote and its partner, so a degraded body said less than + the converting one -- the divergence the parity test forbids. + """ + + body = '

Bonjour

fin' + + assert_the_degraded_path_says_no_less(body) + assert "Bonjour" in EasyvistaContentConverter.from_transport(body) + + +def test_a_document_with_no_closing_bracket_is_answered_without_scanning() -> None: + """No ``>`` means no element, and saying so keeps a bad shape cheap. + + ``html.parser`` cannot finish a tag that never closes, so ``close()`` + flushes it one character at a time and rescans the tail at each step: + measured in glpi_python_client, 32 KB of ``'
None: + """The give-up point is not the end of the body. + + The scan stops where the parser stopped, so everything past that + construct would go missing unless it is handed back explicitly -- + and a body is far more likely to carry the marked section in the + middle than at the end. + """ + + html = "

avant

apres SECRET

" + + assert "avant" in _strip_tags(html) + assert "SECRET" in _strip_tags(html) + assert "SECRET" in EasyvistaContentConverter.from_transport(html) + assert "avant" in EasyvistaContentConverter.from_transport(html) + + +def test_stripping_a_document_with_no_closing_bracket_keeps_all_of_it() -> None: + """The degraded path needs the same guard the converting path has. + + ``html.parser`` cannot complete a tag that never closes, so + ``close()`` flushes it one character at a time and rescans the tail + at each step. The whole document is that unfinished tag's text, which + is the answer the guard returns directly. + """ + + html = '

" * 60 + "kept", id="stray-close-p"), + pytest.param("
" * 60 + "kept", id="stray-close-span"), + pytest.param("

" * 60 + "kept", id="stray-close-b"), + pytest.param("x", id="interleaved"), + pytest.param("

" * 100 + "
" + "
" * 100, id="void-leaf"), + pytest.param("
" * 100 + "" + "
" * 100, id="self-closed-leaf"), + pytest.param("
" * 5000, id="void-only"), + pytest.param("
x
", id="close-of-a-void"), + pytest.param("

The printer is offline.

", id="realistic"), + pytest.param("

x

", id="tags-in-a-comment"), + pytest.param( + "
x
", id="tags-in-js" + ), + pytest.param("") # a spacer cell + cells.append( + f"" + ) + cells.append("") + inner = "
" * 30 + "x", id="tables"), + pytest.param("
" * 100 + "x", id="blockquotes"), + pytest.param("

a

" * 100, id="siblings"), + ], +) +def test_both_paths_agree_about_what_is_markup(html: str) -> None: + """The two renderings of one body must not disagree about its markup. + + This corpus was built in glpi_python_client against a flat scan that + predicted the nesting depth, and it caught the scan reading markup + differently from the parser -- a stray close popping an element the + parser keeps, a void element counted as a parent, a tag inside a + comment or a script body counted at all. The prediction is gone; the + corpus is not, because the same disagreements would show up as the + fallback deleting or inventing text relative to the converting path. + """ + + assert_the_degraded_path_says_no_less(html) + + +@pytest.mark.parametrize( + "html", + [ + # A comment with no ``-->`` is a *bogus comment*: the parser gives up + # at the first ``>``, so the ```` inside it is text and the + # ``
`` lands inside ````. Read that ```` as a + # real close and the count comes back one level short -- which is + # how a document that needed degrading reached the converter. + pytest.param("", id="pi-overlapping-a-comment"), + pytest.param("
", id="terminated-comment"), + pytest.param("

x

", id="doctype"), + pytest.param("
b]]>

x

", id="marked-section"), + pytest.param("

x

", id="raw-text"), + pytest.param("
") == ( + "" + ) + + +def test_a_javascript_link_target_is_rendered_live() -> None: + assert EasyvistaContentConverter.to_transport("[x](javascript:alert(1))") == ( + '

x

' + ) + + +def test_an_executable_scheme_in_angle_brackets_is_not_made_a_link() -> None: + """python-markdown autolinks ``http``, ``https``, ``ftp`` and ``ftps`` only. + + Anything else in angle brackets passes through as raw markup, which a + browser reads as an unknown element: inert, and invisible. + """ + + assert EasyvistaContentConverter.to_transport("") == ( + "

" + ) + + +def test_text_a_memo_displays_as_markup_comes_back_as_markup() -> None: + """Read and written back, escaped markup becomes live markup. + + ``markdownify`` resolves ``<`` and does not escape the ``<`` it + produces, so a memo *showing* the text ``" + assert EasyvistaContentConverter.to_transport(markdown) == ( + "" + ) + + +# --------------------------------------------------------------------------- +# Round trips +# --------------------------------------------------------------------------- +# +# ``from_transport(to_transport(m)) == m`` is the property the converter would +# like to hold. It does not hold universally, and cannot: the two libraries +# either side of the wire disagree about a handful of constructs, and no +# option on either fixes them. So, as in glpi_python_client, the corpus is an +# inventory rather than a property test. Every case is listed, the lossy ones +# carry ``xfail(strict=True)``, and that strictness is the point -- fixing one +# turns its xfail into an XPASS and fails the suite, which forces the +# inventory to be updated rather than quietly drifting out of date. +# +# The weaker property is the one a sync depends on: whatever one cycle does, +# a second changes nothing more. That is checked over the same corpus, with +# its own, shorter list of exceptions. + +#: Markdown a caller writes, by case name. +ROUND_TRIP_CORPUS = { + "plain": "The printer is offline.", + "bold": "The printer is **offline**.", + "italic": "This is *emphasis*.", + "inline-code": "Run `systemctl restart` now.", + "heading": "# Title\n\nBody text.", + "subheading": "## Section\n\nBody text.", + "paragraphs": "First para.\n\nSecond para.", + "hard-break": "line one \nline two", + "bullets": "- alpha\n- beta\n- gamma", + "numbered": "1. one\n2. two", + "blockquote": "> quoted text", + "link": "See [the doc](https://example.org/doc).", + "fence": "```\nx = 1\n```", + "table": "| a | b |\n| --- | --- |\n| 1 | 2 |", + "underscore": "The snake_case name.", + "asterisk": "5 * 3 = 15", + "mixed": "# Title\n\n- alpha\n- beta\n\nClosing **note**.", + "autolink": f"<{PASTED_URL}>", + "titled-link": f'[url]({PASTED_URL} "url")', + "parenthesised-url": "[wiki](https://example.org/wiki/Test_(informatique))", + "lone-asterisks": "prix 5*3 et note * importante", + "accents": "Le serveur ne répond plus, merci de vérifier.", + "query-string": "[doc](https://example.org/doc?a=1&b=2)", + "raw-less-than": "si a < b alors a", + "inner-nbsp": "Texte\xa0suite", + "soft-newline": "line one\nline two", + "nested-list": "- alpha\n - inner\n- beta", + "fence-with-language": "```python\nx = 1\n```", + "angle-bracket-text": "use the key", + "escaped-less-than": "si a < b alors a", + "email-autolink": "", + "synced-description": ( + f'Test !\xa0 \n \n<{PASTED_URL}> \n \n[url]({PASTED_URL} "url")' + ), + "defect-readback": f'Test !\n\n<{PASTED_URL}>\n\n[url]({PASTED_URL} "url")', +} + +#: Cases one write-then-read cycle does not reproduce exactly, and why. +#: +#: Each reason was read off the converter, 2026-09-30, python-markdown 3.10.3 +#: and markdownify 1.2.3; the first four were recorded by glpi_python_client. +LOSSY = { + "soft-newline": ( + "nl2br renders a lone newline as
, which markdownify reads back as " + "a hard break (two trailing spaces). Semantically equivalent, and " + "stable after one cycle." + ), + "nested-list": ( + "markdownify indents nested items by 2 spaces; python-markdown needs " + "4 to keep the nesting, so a second cycle flattens it." + ), + "fence-with-language": ( + "fenced_code emits class='language-python' and markdownify drops the " + "class, so the language tag cannot survive." + ), + "angle-bracket-text": ( + "to_transport does not escape raw markup, so the text reaches EasyVista " + "as a live unknown tag, and the inbound direction drops an unknown " + "tag's markup. The word is gone after one cycle and the doubled space " + "it leaves after two. How EasyVista's UI renders such a tag was not " + "measured." + ), + "escaped-less-than": ( + "markdownify resolves < and does not escape the < it produces, so " + "the reference reads back as a raw <. Both spellings render the same " + "HTML, so this is a change of spelling only." + ), + "email-autolink": ( + "python-markdown renders as an entity-obfuscated mailto: " + "anchor, whose text is not its href, so markdownify writes it as an " + "inline [user@host](mailto:user@host) link. Same link." + ), + "synced-description": ( + "the trailing no-break space and the hard breaks that end each " + "paragraph are whitespace markdownify drops at a paragraph's edge. The " + "text and both links survive." + ), + "defect-readback": ( + "the text <URL> reads back as , an autolink: what a memo " + "displayed as a literal URL in angle brackets becomes a working link " + "after one cycle." + ), +} + +#: Cases a second cycle still changes, and why. +UNSETTLED = { + "nested-list": LOSSY["nested-list"], + "angle-bracket-text": LOSSY["angle-bracket-text"], +} + + +def _inventory(known: dict[str, str]) -> list[object]: + """The corpus as parameters, each case in ``known`` a strict xfail.""" + + return [ + pytest.param( + markdown, + id=name, + marks=[pytest.mark.xfail(strict=True, reason=known[name])] + if name in known + else [], + ) + for name, markdown in ROUND_TRIP_CORPUS.items() + ] + + +def test_the_inventories_name_only_cases_in_the_corpus() -> None: + """A misspelt name would silently un-mark a lossy case, so check them. + + And every unsettled case is lossy: a case one cycle reproduces exactly + is, by that token, already settled. + """ + + assert set(LOSSY) <= set(ROUND_TRIP_CORPUS) + assert set(UNSETTLED) <= set(LOSSY) + + +@pytest.mark.parametrize("markdown", _inventory(LOSSY)) +def test_round_trip_corpus(markdown: str) -> None: + """Markdown survives one write-then-read cycle through memo HTML.""" + + html = EasyvistaContentConverter.to_transport(markdown) + + assert EasyvistaContentConverter.from_transport(html) == markdown + + +@pytest.mark.parametrize("markdown", _inventory(UNSETTLED)) +def test_one_cycle_reaches_a_fixed_point(markdown: str) -> None: + """Whatever one cycle changes, a second cycle changes nothing more. + + This is what keeps a two-way sync from rewriting a memo on every pass: + after the first write, reading back what was written gives exactly the + Markdown that was written. + """ + + once = EasyvistaContentConverter.from_transport( + EasyvistaContentConverter.to_transport(markdown) + ) + twice = EasyvistaContentConverter.from_transport( + EasyvistaContentConverter.to_transport(once) + ) + + assert twice == once + + +@pytest.mark.parametrize( + "memo", + [ + pytest.param( + "

Bonjour,

Le serveur ne répond plus.
" + "Merci de regarder.

", + id="paragraphs-entity-and-br", + ), + pytest.param( + '
Texte suite
', + id="styled-span-and-nbsp", + ), + pytest.param("
  • un
  • deux
", id="list"), + pytest.param( + "

ligne 1
ligne 2

para 2
ligne 4

", + id="both-br-spellings", + ), + pytest.param( + "
ab
12
", + id="table", + ), + pytest.param("
ligne 1\n  ligne 2
", id="preformatted"), + pytest.param("

Titre

corps

", id="heading"), + pytest.param( + "Bonjour,\r\nle serveur ne répond plus.", + id="plain-text-crlf", + marks=pytest.mark.xfail( + strict=True, + reason=( + "a plain-text memo is returned as it is, lone newline " + "included; the first write renders that newline as
, " + "which reads back as a hard break. Stable after that." + ), + ), + ), + pytest.param( + "Appuyer sur puis valider", + id="plain-text-angle-brackets", + marks=pytest.mark.xfail( + strict=True, + reason=( + "the plain-text memo is returned as it is, and the first " + "write sends as a live unknown tag, which reads " + "back as nothing: the word is lost." + ), + ), + ), + pytest.param( + "
  • a
    • b
  • c
", + id="nested-list", + marks=pytest.mark.xfail( + strict=True, + reason=( + "a nested list reads back with its items indented by 2 " + "spaces, which python-markdown does not nest, so the first " + "write flattens it." + ), + ), + ), + pytest.param( + "

Test !

" + f"

&lt;{PASTED_URL}>

" + f'

url

', + id="defect-memo", + marks=pytest.mark.xfail( + strict=True, + reason=( + "the literal <URL> text reads back as an autolink after " + "the first write, as in the defect-readback round trip." + ), + ), + ), + ], +) +def test_a_memo_read_and_written_back_reads_the_same(memo: str) -> None: + """Reading a memo, writing that Markdown back and reading it again. + + The memo-side twin of the fixed point: if this holds, the first sync of + an already-populated memo is the last one to change it. The exceptions + are the memos whose first write is itself lossy. + """ + + markdown = EasyvistaContentConverter.from_transport(memo) + written = EasyvistaContentConverter.to_transport(markdown) + + assert EasyvistaContentConverter.from_transport(written) == markdown + + # --------------------------------------------------------------------------- # Nesting depth # --------------------------------------------------------------------------- From ccd1a41dbdd46fa145dd978518928040deb34f7a Mon Sep 17 00:00:00 2001 From: baraline Date: Wed, 30 Sep 2026 18:25:32 +0200 Subject: [PATCH 03/11] docs(content): document the converter, the content extra, and the memo format Adds docs/content.rst to the user guide: what a memo holds, what each direction does, what survives a round trip, where the converter comes from, and -- in its own section -- that it is not a sanitiser, with the four things a caller relaying Markdown it did not write has to neutralise. The API reference gains the converter's autodoc entry under a new "Rich-text content" section; installation.rst and the README name the extra. docs/vendor-api-reference.md records the only evidence the converter's premise rests on, as tier 4: on 2026-09-30, on one instance, a ticket's COMMENT memo written through the API with HTML was stored byte for byte and the web UI rendered its paragraphs. No vendor documentation of the memo format is recorded, and the new open item O-MEMOFORMAT lists what is still unmeasured: what the UI's own editor writes, whether the UI treats the newlines between blocks as HTML whitespace, and what it shows for raw markup. The ticket-workflow and ticket-actions skills each gain a gotcha: a memo stores what it is sent, nothing renders Markdown for you, and the optional extra converts both ways. Prose only -- the skills contract admits imports from the package root alone, and the converter deliberately lives in a subpackage the root does not import. Co-Authored-By: Claude Opus 5.5 (1M context) --- README.md | 7 + docs/api_reference.rst | 9 ++ docs/content.rst | 155 ++++++++++++++++++++++ docs/index.rst | 1 + docs/installation.rst | 10 +- docs/vendor-api-reference.md | 17 +++ skills/easyvista-ticket-actions/SKILL.md | 11 ++ skills/easyvista-ticket-workflow/SKILL.md | 12 ++ 8 files changed, 221 insertions(+), 1 deletion(-) create mode 100644 docs/content.rst diff --git a/README.md b/README.md index 08ab850..873c0be 100644 --- a/README.md +++ b/README.md @@ -29,6 +29,13 @@ Build it locally with `pip install -e ".[docs]"` then pip install easyvista-python-client ``` +To read and write memo text as Markdown, add the optional `content` extra, which +brings a Markdown <-> HTML converter, `easyvista_python_client.content`: + +```bash +pip install "easyvista-python-client[content]" +``` + ## Usage (sync) ```python diff --git a/docs/api_reference.rst b/docs/api_reference.rst index d331ac0..12d6183 100644 --- a/docs/api_reference.rst +++ b/docs/api_reference.rst @@ -137,6 +137,15 @@ custom field (per EasyVista); official ``E_``-columns like ``E_MAIL`` stay offic because they are declared model fields. Resolve a link's text with ``client.resolve_memo(href)``. +Rich-text content +----------------- + +The optional ``content`` extra, ``pip install "easyvista-python-client[content]"``: +Markdown to and from the HTML a memo holds. See :doc:`content` for what each +direction does, what survives a round trip, and why it is not a sanitiser. + +.. autoclass:: easyvista_python_client.content.EasyvistaContentConverter + Exceptions ---------- diff --git a/docs/content.rst b/docs/content.rst new file mode 100644 index 0000000..4b23f95 --- /dev/null +++ b/docs/content.rst @@ -0,0 +1,155 @@ +.. _content-conversion: + +Rich-text content +================= + +EasyVista keeps a ticket's and an action's text in *memo* fields, and a memo +holds the HTML it was sent. The optional ``content`` extra converts between that +HTML and Markdown in both directions, so a caller can read memos as Markdown and +write Markdown to them without handling HTML itself. + +.. code-block:: bash + + pip install "easyvista-python-client[content]" + +The extra adds ``beautifulsoup4``, ``markdown`` and ``markdownify``. It is +optional so that the core package keeps its three runtime dependencies: nothing +outside ``easyvista_python_client.content`` imports them, and importing that +subpackage without them raises an :class:`ImportError` naming the command above. +:class:`~easyvista_python_client.exceptions.EasyvistaContentError`, the one error +the converter raises, is part of the core package, so it can be caught either +way. + +What a memo holds +----------------- + +A ticket's body lives in its ``COMMENT`` or its ``DESCRIPTION`` memo, depending +on the deployment, and an action's in ``DESCRIPTION``, or in ``COMMENT`` when +``DESCRIPTION`` is empty (see :doc:`the user guide `). The client +reads any of them with +:meth:`~easyvista_python_client.EasyvistaClient.resolve_memo`. + +What a memo contains is whatever was written to it. **Tier 4** -- measured +2026-09-30 on one instance, which may not generalise: a ticket memo written +through the API with HTML was stored byte for byte, and the web UI rendered its +``

`` elements as paragraphs. No vendor documentation of the memo format is +recorded in ``docs/vendor-api-reference.md``. Nothing forces a memo to be HTML +either -- a caller can write plain text -- so the reading direction accepts +both. + +Nothing in the client converts memos for you. The read models keep a memo +exactly as the API returned it, and +:meth:`~easyvista_python_client.context.TicketContext.to_markdown` still reduces +memos to plain text with the core package's dependency-free reducer. Convert +where you want Markdown: + +.. code-block:: python + + from easyvista_python_client import EasyvistaClient, RequestUpdate + from easyvista_python_client.content import EasyvistaContentConverter + + with EasyvistaClient(config) as client: + # Read a memo as Markdown. `resolve_memo` may return None, which reads as "". + memo = client.resolve_memo(f"requests/{rfc_number}/comment") + markdown = EasyvistaContentConverter.from_transport(memo) + + # Write Markdown to it: render first, because a memo stores what it is sent. + html = EasyvistaContentConverter.to_transport("The printer is **offline**.") + client.update_ticket(rfc_number, RequestUpdate(description=html)) + +Reading: ``from_transport`` +--------------------------- + +:meth:`~easyvista_python_client.content.EasyvistaContentConverter.from_transport` +returns ``""`` for an empty memo and a memo with no real HTML element unchanged, +so plain text and Markdown pass through. The test is the element *name*, not the +presence of angle brackets: ``use the key`` and ``if x0`` are +text, because neither ``Enter`` nor ``y`` is an HTML element. + +Real HTML goes through ``markdownify`` with ATX headings, ``-`` bullets, +```` + goes out as a live ``" + assert markdown == "<script>alert(1)</script>" assert EasyvistaContentConverter.to_transport(markdown) == ( - "" + "

<script>alert(1)</script>

" ) @@ -610,10 +631,6 @@ def test_text_a_memo_displays_as_markup_comes_back_as_markup() -> None: "a hard break (two trailing spaces). Semantically equivalent, and " "stable after one cycle." ), - "nested-list": ( - "markdownify indents nested items by 2 spaces; python-markdown needs " - "4 to keep the nesting, so a second cycle flattens it." - ), "fence-with-language": ( "fenced_code emits class='language-python' and markdownify drops the " "class, so the language tag cannot survive." @@ -640,16 +657,10 @@ def test_text_a_memo_displays_as_markup_comes_back_as_markup() -> None: "paragraph are whitespace markdownify drops at a paragraph's edge. The " "text and both links survive." ), - "defect-readback": ( - "the text <URL> reads back as , an autolink: what a memo " - "displayed as a literal URL in angle brackets becomes a working link " - "after one cycle." - ), } #: Cases a second cycle still changes, and why. UNSETTLED = { - "nested-list": LOSSY["nested-list"], "angle-bracket-text": LOSSY["angle-bracket-text"], } @@ -756,29 +767,13 @@ def test_one_cycle_reaches_a_fixed_point(markdown: str) -> None: ), ), pytest.param( - "
  • a
    • b
  • c
", - id="nested-list", - marks=pytest.mark.xfail( - strict=True, - reason=( - "a nested list reads back with its items indented by 2 " - "spaces, which python-markdown does not nest, so the first " - "write flattens it." - ), - ), + "
  • a
    • b
  • c
", id="nested-list" ), pytest.param( "

Test !

" f"

&lt;{PASTED_URL}>

" f'

url

', id="defect-memo", - marks=pytest.mark.xfail( - strict=True, - reason=( - "the literal <URL> text reads back as an autolink after " - "the first write, as in the defect-readback round trip." - ), - ), ), ], ) @@ -865,34 +860,38 @@ def test_the_degraded_path_resolves_entities_and_block_boundaries() -> None: @pytest.mark.parametrize( - ("construct", "kept"), + ("construct", "converted_keeps", "degraded_keeps"), [ - pytest.param("", False, id="resolved-comment"), - pytest.param("", False, id="doctype"), - pytest.param("", False, id="bogus-declaration"), - pytest.param("", True, id="script-body"), - pytest.param("", True, id="style-body"), - pytest.param("", True, id="marked-section"), - pytest.param("", True, id="unterminated-marked-section"), - pytest.param("", True, id="processing-instruction"), + pytest.param("", False, False, id="resolved-comment"), + pytest.param("", False, False, id="doctype"), + pytest.param("", False, False, id="bogus-declaration"), + pytest.param("", False, True, id="script-body"), + pytest.param("", False, True, id="style-body"), + pytest.param("", True, True, id="marked-section"), + pytest.param("", True, True, id="unterminated-marked-section"), + pytest.param("", True, True, id="processing-instruction"), ], ) -def test_the_degraded_path_keeps_exactly_what_the_converter_keeps( - construct: str, kept: bool +def test_the_degraded_path_keeps_at_least_what_the_converter_keeps( + construct: str, converted_keeps: bool, degraded_keeps: bool ) -> None: """Parity, construct by construct, and not one of these was a guess. Each expectation here was read off the converting path rather than - reasoned about, and three came back the opposite way round from the - obvious answer -- a ``
code
", + "- code", + id="code-opening-an-item-after-a-script", + ), + pytest.param( + "
  • code
", + "- code", + id="code-opening-an-item-after-a-comment", + ), + pytest.param( + "
    • code
    ", + "- code", + id="code-opening-an-item-after-an-empty-list", + ), + pytest.param( + '
    • i
      code
      ' + "
    ", + "- ![i](https://example.org/i.png)\n\n code", + id="code-after-an-image", + ), + ], +) +def test_shapes_with_a_rule_of_their_own(html: str, expected: str) -> None: + """The converter's special cases, each on the shape that reaches it. + + What displays nothing -- a comment, a script, an empty list -- does not + count as content before a code block, so the code still opens its item. + """ + + assert read(html) == expected + + +def test_an_ordered_list_keeps_its_start_when_rendered() -> None: + """``sane_lists`` is what keeps ``3.`` from restarting the count at 1.""" + + assert '
      ' in render("intro\n\n3. trois\n4. quatre") + + +@pytest.mark.parametrize( + ("html", "expected"), + [ + pytest.param( + "
      • item
        #4521 code
      ", + "- item\n\n #4521 code", + id="in-a-list-item", + ), + pytest.param( + "
      #4521 C:\\Temp\n  indenté
      ", + "> #4521 C:\\Temp\n> indenté", + id="in-a-blockquote", + ), + pytest.param( + "
      ligne\n```\nfin
      ", + "````\nligne\n```\nfin\n````", + id="holding-a-fence-line", + ), + ], +) +def test_a_preformatted_block_stays_one(html: str, expected: str) -> None: + """A fence opens only at the start of a line, so nested code is indented. + + Inside a list item or a block quote, python-markdown never sees + ``` ``` ``` at the start of a line and reads the fence as text -- the + code's own ``#4521`` then became a heading. An indented code block is the + spelling it does read there. At the top level the fence is made longer + than any fence line the code holds. + """ + + assert read(html) == expected + assert_survives(html) + + +@pytest.mark.xfail( + strict=True, + reason=( + "python-markdown runs html.parser over its whole source to find raw " + "HTML, before an indented code block is recognised, and re-emits a " + "numeric reference written without its semicolon with one: the code " + "then shows 'ᆩ'. A fence is stashed before that pass, which is " + "why a top-level
       is unaffected; inside a list item or a quote no "
      +        "fence can open. Measured on 3.10.3."
      +    ),
      +)
      +def test_a_numeric_reference_in_nested_code_gains_a_semicolon() -> None:
      +    assert_survives("
      echo &#4521
      ") + + +def test_a_preformatted_block_opening_a_list_item_keeps_its_text() -> None: + """The one place python-markdown can start no code block at all. + + A list item's first line is its paragraph, so the code there degrades to + its lines, each escaped as the literal text it is -- the words survive, + the preformatting does not. + """ + + html = "
      • #4521 code\n  suite
      " + + markdown = read(html) + + assert markdown == "- \\#4521 code \n suite" + assert "#4521 code" in render(markdown) + assert read(render(markdown)) == markdown + + +@pytest.mark.parametrize( + ("html", "expected"), + [ + pytest.param( + "
      • item
        • sous-item
        #4521 code
      ", + "- item\n\n - sous-item\n\n \\#4521 code", + id="in-a-list-item", + ), + pytest.param( + "
      • item
      #4521 code
      ", + "> - item\n>\n> \\#4521 code", + id="in-a-quote", + ), + pytest.param( + "
      • item
        • sous-item
        " + "
        #4521 code
      ", + "- item\n\n - sous-item\n\n \\#4521 code", + id="with-only-a-style-block-between", + ), + ], +) +def test_code_right_after_a_nested_list_keeps_its_text( + html: str, expected: str +) -> None: + """The other place python-markdown can start no code block. + + An indented code block right after a list in the same item or quote is + indented exactly as the list's last item's own content, and + python-markdown reads it as a paragraph of that item. The lines are + kept as literal text in the right place instead. + """ + + markdown = read(html) + + assert markdown == expected + assert "#4521 code" in render(markdown) + assert read(render(markdown)) == markdown + + +#: What the format loses, whatever the escaping does: each body with why. +#: +#: None of these is literal text misread. Each is structure python-markdown +#: has no spelling for, or spells as something else. +LOSSES = [ + pytest.param( + "
      • a
      • b
      ", + "python-markdown continues a list across a blank line: two lists are " + "read back as one list of loose items.", + id="two-adjacent-lists", + ), + pytest.param( + "
      a
      b
      ", + "python-markdown continues a quote across a blank line: two quotes are " + "read back as one quote of two paragraphs.", + id="two-adjacent-quotes", + ), + pytest.param( + "

      a

      b

      ", + "Markdown has no blank line inside a paragraph: two breaks in a row " + "are a paragraph break, and read back as one.", + id="two-breaks-in-a-row", + ), + pytest.param( + "

      ab

      ", + "'`a``b`' is one code span holding 'a``b' to python-markdown.", + id="two-adjacent-code-spans", + ), + pytest.param( + "

      a b c

      ", + "python-markdown pairs '*a **b** c*' as three emphasis runs, and b " + "loses its bold.", + id="strong-inside-emphasis", + ), + pytest.param( + "
      ab
      ", + "A Markdown table starts with its header row, so one without gains an " + "empty header.", + id="table-without-a-header", + ), + pytest.param( + "
      ab
      ", + "python-markdown renders a header-only table with one empty body row.", + id="table-without-a-body", + ), + pytest.param( + "
      a
      • x
      ", + "A table cell holds one line of inline Markdown: its list is written " + "as its text.", + id="list-in-a-cell", + ), + pytest.param( + "
      a
      x\ny
      ", + "A table cell holds one line of inline Markdown: its code block is " + "written as inline code.", + id="code-block-in-a-cell", + ), + pytest.param( + "
      a
      un
      deux
      ", + "A table cell holds one line: its line break is written as a space.", + id="break-in-a-cell", + ), + pytest.param( + "

      Titre
      suite

      ", + "A heading is one line: its line break is written as a space.", + id="break-in-a-heading", + ), + pytest.param( + "
      • a
      ", + "An item's first block is its paragraph: the code is kept as text " + "(test_a_preformatted_block_opening_a_list_item_keeps_its_text).", + id="code-opening-a-list-item", + ), + pytest.param( + "
      • a
        • b
        c
      ", + "The code would be indented as the nested item's own content: it is " + "kept as text (test_code_right_after_a_nested_list_keeps_its_text).", + id="code-after-a-nested-list", + ), +] + + +@pytest.mark.parametrize(("html", "reason"), LOSSES) +def test_what_markdown_cannot_carry(html: str, reason: str) -> None: + """The inventory of losses, each asserted to still be one. + + ``xfail(strict=True)`` is the point: a loss that stops being one fails + here, so the inventory stays true. Struck and underlined text lose their + line too, which this comparison does not see: + ``test_struck_text_keeps_its_words_and_loses_its_line``. + """ + + with pytest.raises(AssertionError): + assert_survives(html) + pytest.xfail(reason) + + +@pytest.mark.parametrize(("html", "reason"), LOSSES) +def test_what_is_lost_is_lost_once(html: str, reason: str) -> None: + """Whatever a loss costs, it costs on the first cycle and never again.""" + + again = read(render(read(html))) + + assert read(render(again)) == again, reason + + +@pytest.mark.parametrize( + "html", + [ + pytest.param( + "

      Bonjour

      ", + id="style", + ), + pytest.param("

      Bonjour

      ", id="script"), + pytest.param( + "RE: Imprimante" + "" + "

      Bonjour

      ", + id="outlook-shaped", + ), + ], +) +def test_style_script_and_title_bodies_are_not_text(html: str) -> None: + """A browser displays none of these, so neither does the Markdown. + + ``strip=["script", "style"]`` used to be passed to ``markdownify``, and + ``strip`` skips an element's own converter -- ``convert_script`` and + ``convert_style`` return ``""`` -- so the bodies leaked into the text as + prose. ```` has no converter at all and leaked the same way. + """ + + assert read(html) == "Bonjour" + + +@pytest.mark.parametrize( + ("html", "expected"), + [ + pytest.param( + '<font color="red">URGENT</font> serveur HS', "URGENT serveur HS", id="font" + ), + pytest.param("<center>Titre</center> suite", "Titre\n\nsuite", id="center"), + pytest.param("<strike>ancien</strike> nouveau", "ancien nouveau", id="strike"), + pytest.param("<big>gros</big> texte", "gros texte", id="big"), + pytest.param("<tt>code</tt> texte", "code texte", id="tt"), + pytest.param("<nobr>sans coupure</nobr>", "sans coupure", id="nobr"), + ], +) +def test_a_body_marked_up_only_with_obsolete_elements_is_html( + html: str, expected: str +) -> None: + """Old editors still write these; without them the tags were kept as text.""" + + assert read(html) == expected + + +def test_every_obsolete_element_name_makes_a_body_html() -> None: + """The HTML standard's list of obsolete elements, all of them recognised.""" + + obsolete = ( + "acronym applet basefont bgsound big blink center dir font frame " + "frameset isindex keygen listing marquee menuitem multicol nextid nobr " + "noembed noframes plaintext rb rtc spacer strike tt xmp" + ).split() + + assert set(obsolete) <= conversion._HTML_ELEMENTS + + +@pytest.mark.parametrize( + "html", + [ + pytest.param("<p><s>ancien</s> nouveau</p>", id="s"), + pytest.param("<p><del>ancien</del> nouveau</p>", id="del"), + pytest.param("<p><strike>ancien</strike> nouveau</p>", id="strike"), + ], +) +def test_struck_text_keeps_its_words_and_loses_its_line(html: str) -> None: + """python-markdown has no strikethrough, so ``~~x~~`` would show literally. + + Keeping the words and losing the line is the honest loss: the text is + all there, and what is missing is recorded in the round-trip inventory. + """ + + assert read(html) == "ancien nouveau" + + +@pytest.mark.parametrize( + "html", ["<P>Bonjour</P>", "<BR>Bonjour", "<DIV><B>Bonjour</B></DIV>"] +) +def test_upper_case_tags_are_html(html: str) -> None: + """Element names are case-insensitive; the probe has to be too.""" + + assert "<" not in read(html) + + +@pytest.mark.parametrize( + ("value", "expected"), + [ + pytest.param(" texte ", "texte", id="plain-text"), + pytest.param(" <b>x</b> ", "**x**", id="html"), + pytest.param("<br>x<br>", "x", id="html-with-edge-breaks"), + ], +) +def test_the_result_is_stripped_on_both_paths(value: str, expected: str) -> None: + """Edge whitespace is never content, and a digest must not depend on it.""" + + assert read(value) == expected + + +@pytest.mark.parametrize( + ("html", "expected"), + [ + pytest.param( + '<h2>Titre <img src="https://example.org/i.png" alt="logo"></h2>', + "## Titre ![logo](https://example.org/i.png)", + id="heading", + ), + pytest.param( + "<table><tr><th>a</th></tr><tr><td>" + '<img src="https://example.org/i.png" alt="x"></td></tr></table>', + "| a |\n| --- |\n| ![x](https://example.org/i.png) |", + id="table-cell", + ), + ], +) +def test_an_image_in_a_heading_or_a_cell_stays_an_image( + html: str, expected: str +) -> None: + """``markdownify`` reduced these to their alt text; python-markdown needs not.""" + + assert read(html) == expected + assert_survives(html) + + +def test_a_newline_in_text_is_a_space() -> None: + """HTML displays a source newline as a space; ``nl2br`` would break the line. + + Outlook wraps its HTML source mid-sentence, so every body that came + from an e-mail used to gain line breaks where the reader saw none. + """ + + html = "<p>Bonjour,\nle serveur\nest redémarré.</p>" + + assert read(html) == "Bonjour, le serveur est redémarré." + assert_survives(html) + + +@pytest.mark.parametrize( + ("html", "expected"), + [ + pytest.param("<p>a<br> b<br> c</p>", "a \nb \nc", id="space-after-break"), + pytest.param("<p>a <br>b</p>", "a \nb", id="space-before-break"), + pytest.param("<p><strong>a<br></strong>b</p>", "**a** \nb", id="edge-break"), + ], +) +def test_a_line_break_has_one_spelling(html: str, expected: str) -> None: + """Whitespace around ``<br>`` is not displayed, so it is not kept either. + + Without this the first read of ``a<br> b`` was ``a \\n b`` and the + second ``a \\nb``: the same body, read twice, disagreeing. + """ + + assert read(html) == expected + assert_survives(html) + + +# --------------------------------------------------------------------------- +# The degraded path spells text the same way +# --------------------------------------------------------------------------- + + +def test_a_body_too_deep_to_convert_is_still_literal_safe() -> None: + """The stripped text is Markdown too, and it is escaped like the rest.""" + + html = "<div>" * 600 + r"<p>__init__ \\serveur</p><p># pas un titre</p>" + + markdown = read(html) + + assert markdown == "\\_\\_init\\_\\_ \\\\\\serveur\n\n\\# pas un titre" + assert "__init__ \\\\serveur" in render(markdown) + + +@pytest.mark.parametrize("seed", range(4)) +def test_stripped_text_renders_as_itself(seed: int) -> None: + """Any text, line by line, renders as exactly that text. + + The degraded path's contract, over generated text dense with the + characters python-markdown treats as syntax. + """ + + rng = random.Random(seed) + for _ in range(60): + lines = [_text(rng, rng.randint(1, 6)) for _ in range(rng.randint(1, 4))] + text = "\n".join(line for line in lines if line) + if not text: + continue + + soup = BeautifulSoup(render(conversion._literal_markdown(text)), "html.parser") + for line_break in soup.find_all("br"): + line_break.replace_with("\u2029") + shown = soup.get_text().replace("\n", " ").replace("\u2029", "\n") + + assert [" ".join(line.split()) for line in shown.split("\n")] == [ + " ".join(line.split()) for line in text.split("\n") + ], text + + +# --------------------------------------------------------------------------- +# The property, over realistic bodies and over a fuzzer +# --------------------------------------------------------------------------- + +#: Bodies shaped the way an HTML editor and a mail collector store them -- +#: glpi_python_client's corpus, which EasyVista memos share: nobody working +#: on this package has measured EasyVista's own editor output. +REALISTIC = [ + "<p>Bonjour,</p><p>Le PC du poste 12 ne démarre plus depuis ce matin.</p>", + "<p>Bonjour,<br>Le PC ne démarre plus.<br>Cordialement,<br>Jean</p>", + "<p>Merci de <strong>redémarrer</strong> le serveur <em>avant</em> 18h.</p>", + "<p>Étapes :</p><ul><li>ouvrir la session</li><li>lancer Outlook</li></ul>", + "<table><thead><tr><th>Poste</th><th>IP</th></tr></thead>" + "<tbody><tr><td>PC12</td><td>10.0.0.12</td></tr></tbody></table>", + "<h2>Contexte</h2><p>Migration du serveur.</p>", + '<p>Voir <a href="https://example.org/doc">https://example.org/doc</a></p>', + '<p>Voir <a title="doc" href="https://example.org/doc">la doc</a></p>', + '<p><a href="https://example.org/wiki/Test_(informatique)">wiki</a></p>', + "<p>prix 5*3 et note * importante</p>", + "<p>appuyer sur <Entrée> puis valider</p>", + "<p>Service R&D, bâtiment A & B</p>", + "<p>Montant : 12 000 €</p>", + '<p><span style="color: #e03e2d;">URGENT</span> <u>à traiter</u></p>', + '<p>Cordialement</p><p><img src="https://example.org/logo.png" alt="Logo"></p>', + "<p>Bonjour</p><blockquote><p>Message d'origine</p></blockquote>", + '<pre>Traceback (most recent call last):\n File "x.py", line 1\nError</pre>', + "<p>Lancer <code>ipconfig /all</code> puis envoyer.</p>", + "<p>Fichier mon_fichier_final.docx et __init__</p>", + "<p># pas un titre</p><p>- pas une liste</p>", + '<p>Contact <a href="mailto:support@example.org">support@example.org</a></p>', + "<div>Bonjour,</div><div><br></div><div>Le serveur est down.</div>", + "<p>* point un<br>* point deux</p>", + "<p><https://example.org/x></p>", + "<p>[INFO] tâche [1] terminée</p>", + "<p>[1]: https://example.org/note</p>", + "<p>m<sup>2</sup></p>", + r"<p>Chemin C:\Users\jdupont\Desktop</p>", + "<p>Merci 👍</p>", + "<p>Titre<br>=====</p>", + "<p>---</p><p>signature</p>", + "<p>1) un<br>2) deux</p>", + "<p><support@example.org></p>", + "<p><!-- note --></p>", + "<ul><li><p>point un</p><p>suite du point</p></li><li>point deux</li></ul>", + "<blockquote><ul><li>a</li><li>b</li></ul></blockquote>", + "<p>Nom : ______ Prénom : ______</p>", + "<p>~~pas barré~~</p>", + "<p>a | b | c</p>", + "<p>> pas une citation</p>", + "<p>+ un<br>+ deux</p>", + "<p>Cordialement,<br>Jean Dupont<br>--<br>Service IT</p>", + "<p>Merci<br>-----------<br>Jean Dupont</p>", + "<p>calcul 5 * 3 * 2 = 30</p>", + r"<p>voir \\srv\partage\__archive__\2026</p>", + "<p>Le 30/09, Jean a écrit :<br>> merci<br>> cordialement</p>", + r"<p>Accès au partage \\serveur\compta\2026 refusé</p>", + r"<p>Dossier C:\_temp\logs</p>", + "<p>voir la note [1]</p><p>[1]: https://example.org/note</p>", + "<ul><li>Réseau<ul><li>switch 3</li><li>borne wifi</li></ul></li></ul>", + "<p><b>Important</b> : voir <i>ci-dessous</i></p>", + '<p><font color="red">rouge</font> <span style="font-size:14px">texte</span></p>', + "<p>module __init__ et _x_</p>", + "<p>1) un<br>2) deux</p><pre>ligne 1\n ligne indentée</pre>", + "<p>Suite à la mise à jour, <strong>3 postes</strong> ne se connectent plus :" + "</p><ul><li>PC12 (salle 3)</li><li>PC14 – <em>poste d'accueil</em></li>" + '</ul><p>Voir <a href="https://example.org/kb/42">https://example.org/kb/42</a>' + ' et le <a href="https://example.org/kb/43" title="KB 43">KB 43</a>.</p>', +] + + +@pytest.mark.parametrize("html", REALISTIC) +def test_realistic_bodies_display_the_same_after_the_round_trip(html: str) -> None: + assert_survives(html) + + +#: Literal text snippets: every character python-markdown treats as syntax, +#: alone and in the shapes that trigger it, beside ordinary words so that +#: each lands at line starts, at word boundaries and inside words. +_SPECIAL = ( + "\\ \\\\ \\* \\_ \\[ \\# \\. \\` \\\\serveur * ** *** _ __ ___ ______ __init__ " + "_x_ ` `` ``` ~ ~~ ~~~ [ ] [x] [x](y) ![x](y) [1]: [1]:https://example.org/n " + "( ) (y) ! # ## #4521 > - -- --- + . 1. 2026. = == === | a|b --- < <b> </b> " + "<Enter> <!-- --> <https://example.org> <a@b.c> <3 <= & & < A " + "A ᆩ © © &D; { } : \" '" +).split(" ") +_WORDS = ( + "alpha beta snake_case C:\\Temp R&D x mot 5*3 a_b é fichier_de_test_v2.xlsx" +).split(" ") + + +#: Code content. A numeric reference with no semicolon is left out: inside an +#: indented code block python-markdown's raw-HTML pass adds the semicolon +#: (``test_a_numeric_reference_in_nested_code_gains_a_semicolon``). +_CODE_SPECIAL = [special for special in _SPECIAL if not special.startswith("&#")] + + +def _text(rng: random.Random, pieces: int, specials: list[str] = _SPECIAL) -> str: + out = [] + for _ in range(pieces): + out.append(rng.choice(specials) if rng.random() < 0.45 else rng.choice(_WORDS)) + out.append(rng.choice(["", " ", " ", " "])) + return "".join(out).strip() + + +def _code_text(rng: random.Random, pieces: int) -> str: + return escape(_text(rng, pieces, _CODE_SPECIAL), quote=False) + + +def _inline_html(rng: random.Random, formatted: bool = False, depth: int = 0) -> str: + """Inline HTML that Markdown can express. + + No emphasis inside emphasis, no link inside a link, and a word on either + side of every code element: python-markdown mis-pairs nested ``*`` runs + across constructs and merges two adjacent code spans, neither of which + involves literal text -- they are recorded in the round-trip inventory. + """ + + parts = [] + for _ in range(rng.randint(1, 4)): + kind = rng.random() + if kind < 0.55 or depth > 2: + parts.append(escape(_text(rng, rng.randint(1, 4)), quote=False)) + elif kind < 0.65 and not formatted: + parts.append(f" <strong>{_inline_html(rng, True, depth + 1)}</strong> ") + elif kind < 0.75 and not formatted: + parts.append(f" <em>{_inline_html(rng, True, depth + 1)}</em> ") + elif kind < 0.82 and depth == 0: + inner = _inline_html(rng, formatted, depth + 1) + parts.append( + f'<a href="https://example.org/{rng.randint(1, 9)}">{inner}</a>' + ) + elif kind < 0.88: + code = escape(rng.choice(_WORDS + _SPECIAL[:20]), quote=False) + parts.append(f" mot <code>{code}</code> mot ") + elif kind < 0.94: + alt = escape(_text(rng, rng.randint(0, 2)), quote=True) + parts.append( + f'<img src="https://example.org/{rng.randint(1, 9)}.png" alt="{alt}">' + ) + else: + parts.append(f"<span>{_inline_html(rng, formatted, depth + 1)}</span>") + parts.append(rng.choice(["", " ", " "])) + return "".join(parts).strip() or "mot" + + +def _paragraph(rng: random.Random) -> str: + return "<br>".join(_inline_html(rng) for _ in range(rng.randint(1, 3))) + + +def _list_html(rng: random.Random, depth: int) -> str: + """A list whose items hold text, and sometimes code, a list, and more text. + + The code goes before an item's nested list, never right after it: + python-markdown has no way to write a code block there + (``test_code_right_after_a_nested_list_keeps_its_text``). + """ + + tag = rng.choice(["ul", "ol"]) + items = [] + for _ in range(rng.randint(1, 3)): + shape = rng.random() + if depth < 3 and shape < 0.1: + # An item that opens with a list: one line holds both bullets. + items.append(f"<li>{_list_html(rng, depth + 1)}</li>") + continue + inner = f"<p>{_paragraph(rng)}</p>" if shape < 0.3 else _paragraph(rng) + if rng.random() < 0.1: + inner += f"<pre>{_code_text(rng, 3)}</pre>" + if depth < 3 and rng.random() < 0.3: + inner += _list_html(rng, depth + 1) + after = rng.random() + if after < 0.15: + inner += _inline_html(rng) + elif after < 0.25: + inner += f"<p>{_paragraph(rng)}</p>" + elif after < 0.3: + inner += f"<blockquote><p>{_paragraph(rng)}</p></blockquote>" + items.append(f"<li>{inner}</li>") + return f"<{tag}>{''.join(items)}</{tag}>" + + +def _block(rng: random.Random, depth: int = 0) -> str: + kind = rng.random() + if kind < 0.35 or depth > 1: + return f"<p>{_paragraph(rng)}</p>" + if kind < 0.45: + level = rng.randint(1, 3) + return f"<h{level}>{_inline_html(rng)}</h{level}>" + if kind < 0.60: + return _list_html(rng, depth + 1) + if kind < 0.70: + return f"<blockquote>{_block(rng, depth + 1)}</blockquote>" + if kind < 0.80: + head = "".join(f"<th>{_inline_html(rng)}</th>" for _ in range(2)) + rows = "".join( + "<tr>" + + "".join(f"<td>{_inline_html(rng)}</td>" for _ in range(2)) + + "</tr>" + for _ in range(rng.randint(1, 2)) + ) + return f"<table><tr>{head}</tr>{rows}</table>" + if kind < 0.87: + return f"<pre>{_code_text(rng, 4)}</pre>" + return f"<div>{_paragraph(rng)}</div>" + + +def _document(rng: random.Random) -> str: + """A body of one to four blocks. + + python-markdown merges two adjacent lists of one type, or two adjacent + block quotes, into one -- a limitation of the format -- so a paragraph + separates them. + """ + + blocks: list[str] = [] + for _ in range(rng.randint(1, 4)): + block = _block(rng) + mergeable = block[:4] in {"<ul>", "<ol>", "<blo"} + if blocks and mergeable and block[:4] == blocks[-1][:4]: + blocks.append("<p>mot</p>") + blocks.append(block) + return "".join(blocks) + + +@pytest.mark.parametrize("seed", range(8)) +def test_generated_bodies_display_the_same_after_the_round_trip(seed: int) -> None: + """The property over a seeded fuzzer, 50 bodies a seed.""" + + rng = random.Random(seed) + for _ in range(50): + assert_survives(_document(rng)) + + +@pytest.mark.parametrize( + "html", + [ + pytest.param("<p>" + "[a " * 20_000 + "</p>", id="unclosed-brackets"), + pytest.param("<p>" + "_a " * 20_000 + "</p>", id="underscores-opening-words"), + pytest.param("<p>" + "``` " * 20_000 + "</p>", id="backtick-runs"), + pytest.param("<ol>" + "<li>x</li>" * 20_000 + "</ol>", id="long-numbered-list"), + ], +) +def test_a_body_dense_with_syntax_converts_in_linear_time(html: str) -> None: + """Every pass is linear, so a pathological body costs what its size does. + + Measured before the passes were made linear: 20,000 of these took from + 8 to 117 seconds each, where they take well under one now. The budget + is generous on purpose -- this guards the complexity, not the speed. + """ + + started = time.perf_counter() + read(html) + + assert time.perf_counter() - started < 10 + + +# --------------------------------------------------------------------------- +# What the escaping rests on +# --------------------------------------------------------------------------- + + +def test_python_markdown_undoes_every_backslash_the_reader_writes() -> None: + """A backslash escape is only safe if python-markdown removes it again. + + Measured on 3.10.3 with the four extensions: ``\\``, backtick, ``*``, + ``_``, ``{``, ``}``, ``[``, ``]``, ``(``, ``)``, ``>``, ``#``, ``+``, + ``-``, ``.``, ``!`` and ``|``. ``=`` and ``~`` are not among them, which + is why those two are spelled as character references instead. + """ + + escapable = set(Markdown(extensions=conversion._MARKDOWN_EXTENSIONS).ESCAPED_CHARS) + backslashed = set(conversion._LITERAL) - set(conversion._SPELLED_OUT) | {"\\"} + + assert backslashed <= escapable + assert not {"=", "~", "<", "&"} & escapable + + +def test_a_body_holding_the_private_stand_ins_converts_like_any_other() -> None: + """Literal text is carried in private-use characters until it is spelled. + + A body that already holds one of them -- an icon font maps symbols + there -- must not be confused with them, so the converter picks a block + the body does not use. + """ + + first_block = [chr(0xF0000 + offset) for offset in range(32)] + html = "<p>" + "".join(first_block) + " __init__ et #4521</p>" + + markdown = read(html) + + assert "".join(first_block) in markdown + assert markdown.endswith("\\_\\_init\\_\\_ et #4521") diff --git a/pyproject.toml b/pyproject.toml index d51abbc..8824176 100644 --- a/pyproject.toml +++ b/pyproject.toml @@ -91,13 +91,13 @@ Source = "https://github.com/baraline/easyvista_python_client" content = [ "beautifulsoup4>=4.12", "markdown>=3.6", - "markdownify>=0.13", + "markdownify>=1.2", ] dev = [ "beautifulsoup4>=4.12", "markdown>=3.6", - "markdownify>=0.13", + "markdownify>=1.2", "build>=1.2", "mypy>=1.11", "numpydoc>=1.8", @@ -138,7 +138,7 @@ dev = [ docs = [ "beautifulsoup4>=4.12", "markdown>=3.6", - "markdownify>=0.13", + "markdownify>=1.2", "numpydoc>=1.8", "sphinx>=7.2,<8.2", "sphinx-rtd-theme>=2.0", From 9e99a69fcba594a651a9335421f7db1f9e71e0ef Mon Sep 17 00:00:00 2001 From: baraline <antoine.guillaume45@gmail.com> Date: Thu, 1 Oct 2026 01:55:46 +0200 Subject: [PATCH 07/11] docs(content): say what the literal-safe reader escapes, and cite O-MEMOFORMAT docs/content.rst describes the reader as it is now: text in a memo is literal and the Markdown spells it so, with the example checked against the code; ordinary prose carries no escape; nested lists nest; a memo holding one real HTML element is read as HTML throughout. "It is not a sanitiser" keeps the three outbound cases and states the one inbound change -- text a memo displays as markup reads back as text, where it used to come back live. "What survives a round trip" drops the nested list from the losses, adds the memo-side property and the inventory of what Markdown cannot carry, and the port reference moves to 4fc3bed. The 0.4.0 changelog entry, which has not shipped, is corrected to match: markdownify>=1.2 and why, the port commit, what the reader escapes, and six lossy round-trip shapes rather than eight. The two skills that document the content extra say not to strip the backslashes and not to mix Markdown and HTML in one memo. models/request.py no longer calls the memo format "unverified (O4)": it is whatever its writer sent, measured on one instance, and what is still unknown is open item O-MEMOFORMAT in docs/vendor-api-reference.md. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> --- CHANGELOG.md | 44 ++++++++--- docs/content.rst | 89 ++++++++++++++++------- easyvista_python_client/models/request.py | 6 +- skills/easyvista-ticket-actions/SKILL.md | 4 +- skills/easyvista-ticket-workflow/SKILL.md | 7 +- 5 files changed, 109 insertions(+), 41 deletions(-) diff --git a/CHANGELOG.md b/CHANGELOG.md index 07b051f..eee6df9 100644 --- a/CHANGELOG.md +++ b/CHANGELOG.md @@ -27,15 +27,32 @@ public surface, not for a break. - `easyvista_python_client.content.EasyvistaContentConverter`, behind the new optional extra `easyvista-python-client[content]` (`beautifulsoup4>=4.12`, - `markdown>=3.6`, `markdownify>=0.13`). Two static methods: + `markdown>=3.6`, `markdownify>=1.2`). Two static methods: `from_transport(value)` reads a memo -- rich-text HTML or plain text -- as canonical Markdown, and `to_transport(value)` renders Markdown as the HTML a memo is written with. - It is a port of `glpi_python_client`'s `content/conversion.py` at `0d43528`, - with the same options, extensions and edge-case handling, so with the same - library versions the two produce the same Markdown from the same HTML. **The - two should move together.** The only code divergence is the optional extra. + It is a port of `glpi_python_client`'s `content/conversion.py` at `4fc3bed`, + the literal-safe converter, with the same rules, extensions and edge-case + handling, so with the same library versions the two produce the same + Markdown from the same HTML. **The two should move together.** The only code + divergence is the optional extra. + + **Text in a memo is literal, and the Markdown spells it so**: a character is + escaped exactly where python-markdown, with the four extensions + `to_transport` uses, would otherwise read it as syntax -- a typed `__init__` + reads as `\_\_init\_\_`, `\\serveur` as `\\\serveur`, `#4521` at a line start + as `\#4521`, `<Entrée>` as `<Entrée>` -- and nowhere else, so ordinary + prose such as `fichier_de_test_v2.xlsx` or `R&D` comes back as typed. + Rendering the Markdown displays what the memo displayed, and reading that + back gives the same Markdown: both are tested with an HTML parser over + realistic memos and a seeded fuzzer. Nested lists nest and keep their + numbers, `<script>`, `<style>` and `<title>` bodies are dropped, and the + obsolete elements (`<font>`, `<center>`, ...) make a memo HTML. The + `markdownify` floor is 1.2 because the converter subclasses + `MarkdownConverter` and relies on its 1.x hooks, which 0.13 does not have. + A memo holding one real HTML element is read as HTML throughout, so + Markdown syntax beside it is kept as literal characters. Why it is here: until now this package could only strip HTML to plain text, so a downstream GLPI-to-EasyVista sync wrote its own converter. Measured @@ -51,9 +68,9 @@ public surface, not for a break. `RecursionError` answered with the memo's words, and the cliff is measured at 493 levels on CPython 3.12-3.14 but 328 on 3.10. And **it is not a sanitiser**: raw HTML and `javascript:` link targets in the Markdown go out - live, and text a memo displays as markup (`<script>`) reads back as - raw markup. A caller relaying Markdown it did not write must neutralise - both. + live. A caller relaying Markdown it did not write must neutralise both. Text + a memo *displays* as markup (`<script>`) reads back escaped, as the + text it is. Importing `easyvista_python_client.content` without the extra raises an `ImportError` naming `pip install "easyvista-python-client[content]"`. @@ -80,13 +97,16 @@ public surface, not for a break. ### Notes -- The round-trip inventory is a test, not a promise: eight Markdown shapes do +- The round-trip inventory is a test, not a promise: six Markdown shapes do not survive one write-then-read cycle exactly, each a strict xfail with the measured reason -- among them a lone newline, which comes back as a hard - break, and a nested list, which a second cycle flattens. Past the first - cycle a second changes nothing more, except for that nested list and for + break. Past the first cycle a second changes nothing more, except for angle-bracket text such as `use the <Enter> key`, which is sent as a live - unknown tag and lost. + unknown tag and lost. What Markdown cannot carry at all when a memo is read + -- strikethrough, adjacent lists merging, two `<br>` in a row, a line break + in a table cell, a `<pre>` opening a list item -- is a second inventory, + `test_what_markdown_cannot_carry`, each case asserted to stabilise after + one cycle. - The `beautifulsoup4` defect glpi_python_client works around -- text after a `<br />` dropped in a body that also used `<br>` -- no longer reproduces on 4.15.0 (measured 2026-09-30, CPython 3.10 and 3.12-3.14). The workaround diff --git a/docs/content.rst b/docs/content.rst index 3ee77ac..04ad7f0 100644 --- a/docs/content.rst +++ b/docs/content.rst @@ -70,13 +70,36 @@ The test is the element *name*, not the presence of angle brackets: ``use the <Enter> key`` and ``if x<y then z>0`` are text, because neither ``Enter`` nor ``y`` is an HTML element. -Real HTML goes through ``markdownify`` with ATX headings, ``-`` bullets, -``<script>`` and ``<style>`` markup stripped, and underscores and asterisks in -prose left unescaped, so ``snake_case`` does not grow a backslash on every read. -An anchor whose text is its own URL -- a pasted link -- reads as the autolink -``<https://...>``; a ``title`` attribute reads as a link title, +Real HTML goes through ``markdownify``, and **the text in it is literal**: the +Markdown spells it so that rendering it -- with ``to_transport``, or any +python-markdown using the same four extensions -- displays what the memo +displayed. Markdown has one spelling for a ``__init__`` a user typed and for +bold ``init``, so a character is escaped exactly where python-markdown would +otherwise read it as syntax, and nowhere else: + +.. code-block:: python + + EasyvistaContentConverter.from_transport( + "<p>Voir __init__ et \\\\serveur\\partage</p><p># pas un titre</p>" + ) + # 'Voir \\_\\_init\\_\\_ et \\\\\\serveur\\partage\n\n\\# pas un titre' + +Ordinary prose carries no escape at all -- ``fichier_de_test_v2.xlsx``, +``C:\Temp``, ``R&D``, ``snake_case``, a ``#`` or a ``-`` mid-sentence come back +exactly as typed -- and the Markdown is a fixed point: rendering it and reading +it back gives the same Markdown again. A backslash is used wherever +python-markdown removes one; ``<``, ``&``, ``=`` and ``~`` are spelled as +character references (``<``) where they would be read, since no backslash +escapes them. Nested lists nest, at four spaces, and keep their numbers; +``<script>``, ``<style>`` and ``<title>`` bodies are dropped, as a browser drops +them. An anchor whose text is its own URL -- a pasted link -- reads as the +autolink ``<https://...>``; a ``title`` attribute reads as a link title, ``[text](https://... "title")``. +A memo holding a single real HTML element is read as HTML throughout, so +Markdown syntax in the same memo -- ``**bold** <b>x</b>`` -- is read as the +literal characters it is: write a memo as HTML or as Markdown, not both. + **Deep nesting degrades, it does not raise.** ``markdownify`` walks the document recursively, so a deeply nested memo can exhaust the interpreter's stack: from a shallow stack, the deepest ``<div>`` document that converts is 493 levels on @@ -105,53 +128,69 @@ reported -- and a list nested around 500 levels deep raises It is not a sanitiser --------------------- -Neither direction neutralises anything, by design, exactly as in -``glpi_python_client``: the Markdown is the caller's own. So: +Spelling literal text as text is not sanitising, and the writing direction +neutralises nothing, by design, exactly as in ``glpi_python_client``: the +Markdown is the caller's own. So: * raw HTML in the Markdown is rendered verbatim -- ``<script>alert(1)</script>`` goes out as a live ``<script>``; * a ``javascript:`` link target is rendered as a live ``href``; * ``<javascript:alert(1)>`` is not made a link (python-markdown autolinks ``http``, ``https``, ``ftp`` and ``ftps`` only) but passes through as raw - markup; -* text a memo *displays* as markup, ``<script>``, reads back as a raw - ``<script>`` -- which the writing direction would then emit live. + markup. + +What the reading direction does is keep text a memo *displays* as text: +``<script>`` reads back as ``<script>``, and is written back as the +same ``<script>``. It used to read back as a raw ``<script>``, which the +writing direction then emitted live. A caller relaying Markdown it did not write -- a sync between two ITSMs, for one --- must neutralise raw HTML and executable link schemes before calling +-- must still neutralise raw HTML and executable link schemes before calling ``to_transport``. What survives a round trip -------------------------- Markdown written and read back is the same Markdown for paragraphs, emphasis, -headings, lists, block quotes, fences, tables, links -- titled ones and URLs -containing parentheses included -- autolinks, underscores, lone asterisks, -accents and query strings. The exceptions, each pinned by a test: +headings, lists -- nested ones included -- block quotes, fences, tables, links +-- titled ones and URLs containing parentheses included -- autolinks, +underscores, lone asterisks, accents, query strings and escaped literal text. +The exceptions, each pinned by a test: * a lone newline comes back as a hard break (two trailing spaces); -* a nested list comes back indented by two spaces, which python-markdown does - not nest, so a second cycle flattens it; * a fence's language tag is dropped; * text in angle brackets that is not a URL, ``use the <Enter> key``, is sent as a live unknown tag and does not come back; -* ``<`` comes back as a raw ``<``, which renders the same; +* a ``<`` that opens no tag, ``a < b``, comes back as a raw ``<``, which + renders the same; * an e-mail autolink comes back as an inline ``mailto:`` link; * a no-break space or hard break at the end of a paragraph is dropped. -Past the first cycle, a second changes nothing more, except for the nested list -and the angle-bracket text above. That is the property a two-way sync relies on: -once a text has made one trip, writing what was read back and reading it again -gives exactly the same Markdown. +Past the first cycle, a second changes nothing more, except for the +angle-bracket text above. That is the property a two-way sync relies on: once +a text has made one trip, writing what was read back and reading it again gives +exactly the same Markdown. + +The other direction -- a memo read, written back and read again -- is held to +more: what the Markdown displays is what the memo displayed, compared with an +HTML parser over realistic memos and a seeded fuzzer, and the Markdown is a +fixed point from the first read. A few structures have no Markdown spelling, +and the inventory in ``test_literal_text.py`` records each: struck-through and +underlined text keep their words and lose the line, adjacent lists or quotes +merge, two ``<br>`` in a row become a paragraph break, adjacent code spans +merge, strong inside emphasis loses its bold, a table without a header row +gains an empty one, a table cell or a heading holds one line, and a ``<pre>`` +that opens a list item or directly follows a list inside the same item keeps +its lines as text rather than as code. Where it comes from ------------------- The converter is a port of ``glpi_python_client``'s -``content/conversion.py`` at commit ``0d43528``, with the same options, the same -extensions and the same edge-case handling, and **the two should move -together**: the hard part of both is the behaviour of the same three libraries, -not anything either ITSM does. Names and error messages aside, the only +``content/conversion.py`` at commit ``4fc3bed``, the literal-safe converter, +with the same rules, the same extensions and the same edge-case handling, and +**the two should move together**: the hard part of both is the behaviour of the +same three libraries, not anything either ITSM does. Names and error messages aside, the only difference in code is that the three libraries are an optional extra here rather than dependencies. One measurement differs from the one recorded there: the ``beautifulsoup4`` defect that dropped diff --git a/easyvista_python_client/models/request.py b/easyvista_python_client/models/request.py index 3a3f02b..05edf4b 100644 --- a/easyvista_python_client/models/request.py +++ b/easyvista_python_client/models/request.py @@ -55,8 +55,10 @@ class Request(EasyvistaModel): title: str | None = Field(default=None, alias="TITLE") # The list view returns DESCRIPTION inline (a string); the single-ticket GET # expands it into an HREF reference object (``{"HREF": ".../description"}``). - # Accept either so both read paths validate. Whether the resolved text is - # HTML or plain text is still unverified (spec open item O4). + # Accept either so both read paths validate. The resolved text is whatever + # its writer sent, HTML or plain text: measured 2026-09-30 on one instance + # (tier 4, may not generalise), and what is still unknown is open item + # O-MEMOFORMAT in docs/vendor-api-reference.md. description: str | dict[str, Any] | None = Field(default=None, alias="DESCRIPTION") external_reference: str | None = Field(default=None, alias="EXTERNAL_REFERENCE") diff --git a/skills/easyvista-ticket-actions/SKILL.md b/skills/easyvista-ticket-actions/SKILL.md index cb9934d..e3ee712 100644 --- a/skills/easyvista-ticket-actions/SKILL.md +++ b/skills/easyvista-ticket-actions/SKILL.md @@ -529,7 +529,9 @@ with EasyvistaClient.from_env() as client: `EasyvistaContentConverter.to_transport(markdown)` for the `description` you write, `EasyvistaContentConverter.from_transport(memo)` on what `resolve_memo` returns. Import it from the `easyvista_python_client.content` - subpackage. It sanitises nothing: raw HTML and `javascript:` link targets go + subpackage. Reading spells a note's text as literal text (`__init__` comes + back as `\_\_init\_\_`), so render the Markdown rather than stripping its + backslashes. It sanitises nothing: raw HTML and `javascript:` link targets go out live, so neutralise both in Markdown you did not write — a comment sync relaying another ITSM's text is exactly that case. - **`create_action` resolves an implicit parent** and needs exactly **one** open diff --git a/skills/easyvista-ticket-workflow/SKILL.md b/skills/easyvista-ticket-workflow/SKILL.md index 623f9ac..8b92b8e 100644 --- a/skills/easyvista-ticket-workflow/SKILL.md +++ b/skills/easyvista-ticket-workflow/SKILL.md @@ -231,7 +231,12 @@ with EasyvistaClient.from_env() as client: `EasyvistaContentConverter.to_transport(markdown)` before the write, `EasyvistaContentConverter.from_transport(memo)` on what `resolve_memo` returns. It lives in the `easyvista_python_client.content` subpackage, not - the package root, so it is not imported unless you ask for it. It sanitises + the package root, so it is not imported unless you ask for it. Reading + spells a memo's text as literal text -- a typed `__init__` comes back as + `\_\_init\_\_`, `#4521` at a line start as `\#4521`, `<Entrée>` as + `<Entrée>` -- so render the Markdown to display it rather than + stripping the backslashes, and write a memo as HTML or as Markdown, not + both: one real HTML element makes the whole value HTML. It sanitises nothing: raw HTML and `javascript:` link targets in the Markdown go out live, so neutralise both in Markdown you did not write. - A `description` passed to **`PostRequest`** at create time was not readable From 2bfe220ae5ed39f3c05c1418325e4fc52699160d Mon Sep 17 00:00:00 2001 From: baraline <antoine.guillaume45@gmail.com> Date: Fri, 2 Oct 2026 15:33:22 +0200 Subject: [PATCH 08/11] build!: drop Python 3.10 CPython 3.10 reaches end of life in October 2026 (PEP 619), and glpi_python_client, whose converter the content extra ports, dropped it in 524304a. requires-python is now >=3.11, the 3.10 classifier is gone, and the CI and release matrices run 3.11-3.14. typing-extensions and tomli, needed only on 3.10, go with it, and ruff and mypy target 3.11. datetime.UTC replaces timezone.utc (the same object); the timestamp parser accepts and refuses exactly the values it did. On 3.10, pip keeps resolving 0.3.0, which has no content extra; README and the installation page say so. Three release.yml comments were wrong before this change and are corrected: the release job gates no coverage (ci.yml's coverage job does), every tag is unprefixed, and "same coverage as 3.10" is history. The two model test files carry ruff 0.16.0's formatting, so the pre-commit hook does not rewrite them at commit time. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> --- .github/workflows/ci.yml | 12 +++---- .github/workflows/release.yml | 27 ++++++++-------- CONTRIBUTING.md | 4 +++ README.md | 5 ++- docs/conf.py | 6 +--- docs/development.rst | 4 +++ docs/installation.rst | 6 +++- docs/publishing.rst | 2 +- easyvista_python_client/models/common.py | 13 +++++--- .../models/tests/test_common.py | 11 ++++--- .../models/tests/test_request.py | 22 ++++--------- .../testing/test_unasync_codegen.py | 10 ++---- .../tests/test_reporting.py | 13 ++++---- .../tests/test_timestamps.py | 10 +++--- easyvista_python_client/timestamps.py | 32 +++++++++++-------- integration_tests/test_live_change_window.py | 6 ++-- pyproject.toml | 10 ++---- scripts/run_coverage_gate.py | 6 +++- skills/easyvista-asset-workflow/SKILL.md | 2 +- skills/easyvista-client-setup/SKILL.md | 2 +- skills/easyvista-directory/SKILL.md | 2 +- skills/easyvista-document-workflow/SKILL.md | 2 +- skills/easyvista-instance-discovery/SKILL.md | 2 +- .../easyvista-reporting-and-context/SKILL.md | 2 +- skills/easyvista-search-syntax/SKILL.md | 2 +- skills/easyvista-ticket-actions/SKILL.md | 2 +- skills/easyvista-ticket-workflow/SKILL.md | 2 +- 27 files changed, 114 insertions(+), 103 deletions(-) diff --git a/.github/workflows/ci.yml b/.github/workflows/ci.yml index 15ca8b3..8fa972a 100644 --- a/.github/workflows/ci.yml +++ b/.github/workflows/ci.yml @@ -11,7 +11,7 @@ jobs: matrix: # Keep in step with the `Programming Language :: Python` classifiers in # pyproject.toml and with the same matrix in release.yml. - python-version: ["3.10", "3.11", "3.12", "3.13", "3.14"] + python-version: ["3.11", "3.12", "3.13", "3.14"] steps: - uses: actions/checkout@v4 - uses: actions/setup-python@v5 @@ -47,13 +47,13 @@ jobs: # including a one-file local run -- measure the whole package and fail. Now # that addopts is clean, THIS job is where the floor is actually asserted; # delete it and fail_under becomes a number nothing reads. `needs: test` - # keeps the failure legible: a broken test reports as a broken test on five - # Python versions, not as an under-coverage number here. + # keeps the failure legible: a broken test reports as a broken test on every + # Python version in the matrix, not as an under-coverage number here. # # One version, not the matrix: coverage of this package does not vary by - # interpreter (there is no version-gated code), so five runs would produce - # five identical percentages, five uploads, and a Codecov report whose - # totals depend on which one landed last. + # interpreter (there is no version-gated code), so one run per matrix entry + # would produce identical percentages, one upload each, and a Codecov + # report whose totals depend on which one landed last. needs: test runs-on: ubuntu-latest steps: diff --git a/.github/workflows/release.yml b/.github/workflows/release.yml index e4963eb..685221e 100644 --- a/.github/workflows/release.yml +++ b/.github/workflows/release.yml @@ -26,18 +26,19 @@ jobs: fail-fast: false matrix: # Mirrors ci.yml and the `Programming Language :: Python` classifiers in - # pyproject.toml. requires-python is >=3.10; the upper end is bounded by + # pyproject.toml. requires-python is >=3.11; the upper end is bounded by # what the suite is verified on, not by anything in the code. The # generated-_sync/ gate is the reason a new interpreter is not just # added on faith: it is a byte-equality check whose output comes from # tokenize-rt re-tokenizing the source, so a tokenizer change can fail # it on one version alone. The suite and that gate were both run on - # 3.13.14 and 3.14.6 before either was listed here (600 passed, same - # coverage as 3.10, _sync/ regenerated identical under both). + # 3.13.14 and 3.14.6 before either was listed here (600 passed, the + # same coverage as 3.10 then gave, _sync/ regenerated identical under + # both). # These entries are minor versions, not pinned patches: setup-python # resolves a bare "3.14" to the newest STABLE 3.14.x and never to a # pre-release, which needs allow-prereleases. - python-version: ["3.10", "3.11", "3.12", "3.13", "3.14"] + python-version: ["3.11", "3.12", "3.13", "3.14"] steps: - name: Check out repository @@ -59,8 +60,9 @@ jobs: # than left to skip on absent credentials, so a leaked secret could not # make a release run go live. --strict-markers/--strict-config must be # passed here, not in addopts, where pytest 9 silently ignores them. - # Coverage runs implicitly: --cov lives in pyproject's addopts and - # [tool.coverage.report] fail_under = 95 gates it. + # This job does not gate coverage: pyproject has no addopts, so no + # --cov runs here. In CI the 95% floor is enforced only by ci.yml's + # `coverage` job, and locally by the pre-push hook. - name: Run tests run: python -m pytest -m "not integration" --strict-markers --strict-config @@ -135,10 +137,10 @@ jobs: # easyvista_python_client.__version__ -- and the git tag is a third. PyPI # takes whatever pyproject says, so a tag that disagrees publishes a # release nobody can find by version, and a __version__ that disagrees - # misreports at runtime. Both are unfixable after upload. The repo's only - # existing tag, 0.1.0, is UNPREFIXED; v-prefixing starts at v0.2.0. The - # leading v is therefore stripped before comparing, so both forms - # validate. + # misreports at runtime. Both are unfixable after upload. Every tag so + # far (0.1.0, 0.2.0, 0.3.0) is unprefixed, as CHANGELOG.md says tags + # are. A leading v is still stripped before comparing, as a tolerance, + # so a v-prefixed tag validates too. - name: Validate release tag matches package version if: github.event_name == 'release' shell: bash @@ -149,10 +151,7 @@ jobs: normalized_tag="${release_tag#v}" pyproject_version="$(python - <<'PY' from pathlib import Path - try: - from tomllib import loads as toml_loads - except ModuleNotFoundError: - from tomli import loads as toml_loads + from tomllib import loads as toml_loads pyproject = toml_loads(Path('pyproject.toml').read_text(encoding='utf-8')) print(pyproject['project']['version']) diff --git a/CONTRIBUTING.md b/CONTRIBUTING.md index a3f47a2..e9c6ab9 100644 --- a/CONTRIBUTING.md +++ b/CONTRIBUTING.md @@ -8,6 +8,10 @@ this page is the short version. ## Development Setup +The project needs Python 3.11 or newer. `requires-python` is `>=3.11` since +0.4.0, so a `.venv` created with 3.10 refuses the editable install below: +recreate it with a newer interpreter. + ```bash python -m venv .venv .venv\Scripts\activate diff --git a/README.md b/README.md index 873c0be..a07acc8 100644 --- a/README.md +++ b/README.md @@ -3,7 +3,7 @@ [![CI](https://github.com/baraline/easyvista_python_client/actions/workflows/ci.yml/badge.svg?branch=main)](https://github.com/baraline/easyvista_python_client/actions/workflows/ci.yml) [![Coverage](https://codecov.io/gh/baraline/easyvista_python_client/branch/main/graph/badge.svg)](https://codecov.io/gh/baraline/easyvista_python_client) [![License](https://img.shields.io/github/license/baraline/easyvista_python_client)](https://github.com/baraline/easyvista_python_client/blob/main/LICENSE) -[![Python](https://img.shields.io/badge/python-3.10%2B-blue)](https://github.com/baraline/easyvista_python_client) +[![Python](https://img.shields.io/badge/python-3.11%2B-blue)](https://github.com/baraline/easyvista_python_client) [![Docs](https://readthedocs.org/projects/easyvista-python-client/badge/?version=latest)](https://easyvista-python-client.readthedocs.io/en/latest/) @@ -29,6 +29,9 @@ Build it locally with `pip install -e ".[docs]"` then pip install easyvista-python-client ``` +Python 3.11 or newer. 0.4.0 dropped 3.10; on 3.10, pip installs 0.3.0, which +has no `content` extra: the converter below needs 3.11 or newer. + To read and write memo text as Markdown, add the optional `content` extra, which brings a Markdown <-> HTML converter, `easyvista_python_client.content`: diff --git a/docs/conf.py b/docs/conf.py index d483258..eb92ed9 100644 --- a/docs/conf.py +++ b/docs/conf.py @@ -6,11 +6,7 @@ from importlib.metadata import PackageNotFoundError from importlib.metadata import version as _installed_version from pathlib import Path - -try: - from tomllib import loads as toml_loads -except ModuleNotFoundError: # Python < 3.11 - from tomli import loads as toml_loads +from tomllib import loads as toml_loads def _read_project_version() -> str: diff --git a/docs/development.rst b/docs/development.rst index adbf45a..5fbd893 100644 --- a/docs/development.rst +++ b/docs/development.rst @@ -4,6 +4,10 @@ Development Setup ----- +Use Python 3.11 or newer. ``requires-python`` is ``>=3.11`` since 0.4.0, so an +environment created with 3.10 refuses the editable install: recreate it with a +newer interpreter. + .. code-block:: bash pip install -e ".[dev]" diff --git a/docs/installation.rst b/docs/installation.rst index c6a3736..ef32653 100644 --- a/docs/installation.rst +++ b/docs/installation.rst @@ -4,7 +4,11 @@ Installation Requirements ------------ -* Python 3.10 or newer. +* Python 3.11 or newer. Release 0.4.0 dropped 3.10, which reaches end of life + in October 2026 (PEP 619). On 3.10, ``pip`` installs 0.3.0, the last release that + supports it. 0.3.0 has no ``content`` extra: asked for + ``easyvista-python-client[content]`` on 3.10, ``pip`` warns and installs + 0.3.0 without the converter, which needs 3.11 or newer. * Runtime dependencies (installed automatically): ``httpx``, ``pydantic>=2``, ``tenacity``. From PyPI diff --git a/docs/publishing.rst b/docs/publishing.rst index e9ec6e6..ab7d7db 100644 --- a/docs/publishing.rst +++ b/docs/publishing.rst @@ -45,7 +45,7 @@ Cutting a release ``v`` before comparing, so a prefixed tag would still *build* -- which is exactly why this drifted unnoticed.) -The workflow then runs the test matrix (3.10--3.14) and the quality gates -- Ruff, mypy, +The workflow then runs the test matrix (3.11--3.14) and the quality gates -- Ruff, mypy, the generated-``_sync``-tree check, the hand-written-twin lint and a warnings-as-errors Sphinx build -- validates the tag against the package version, builds the wheel and the sdist, runs ``twine check``, uploads to PyPI, and finally triggers a Read the Docs build diff --git a/easyvista_python_client/models/common.py b/easyvista_python_client/models/common.py index f14894a..78ecbb3 100644 --- a/easyvista_python_client/models/common.py +++ b/easyvista_python_client/models/common.py @@ -4,7 +4,7 @@ import re from collections.abc import Sequence -from datetime import datetime, timezone +from datetime import UTC, datetime from decimal import Decimal from typing import Annotated, Any @@ -140,7 +140,7 @@ def _parse_with_context_formats(value: Any, info: ValidationInfo) -> datetime | parsed = datetime.strptime(value.strip(), pattern) except (TypeError, ValueError): continue - return parsed if parsed.tzinfo else parsed.replace(tzinfo=timezone.utc) + return parsed if parsed.tzinfo else parsed.replace(tzinfo=UTC) return None @@ -204,10 +204,13 @@ def _empty_str_to_none_datetime(value: Any, info: ValidationInfo) -> Any: EasyVista returns ISO 8601 with an explicit UTC offset and millisecond precision (``2026-08-17T15:40:41.610+02:00``), and ``""`` for an unset date — -verified live 2026-08-17. Python 3.10's ``fromisoformat`` rejects the 3-digit -fraction outright, which is why this goes through +verified live 2026-08-17. This goes through :func:`~easyvista_python_client.parse_ev_datetime` rather than letting pydantic -parse the string itself. A naive ``datetime`` passed in directly (not just a +parse the string itself because pydantic's parser is far more permissive than +that format and turns an ISO-basic or epoch-shaped value into a plausible but +wrong instant (see :func:`_empty_str_to_none_datetime`); the original reason, +Python 3.10's ``fromisoformat`` rejecting the 3-digit fraction, went with the +3.10 floor. A naive ``datetime`` passed in directly (not just a wire string) is normalized to aware UTC the same way, so the ``| None`` aside, this type's value is always timezone-aware, never naive. """ diff --git a/easyvista_python_client/models/tests/test_common.py b/easyvista_python_client/models/tests/test_common.py index 413da6d..c3754a0 100644 --- a/easyvista_python_client/models/tests/test_common.py +++ b/easyvista_python_client/models/tests/test_common.py @@ -1,4 +1,4 @@ -from datetime import datetime, timedelta, timezone +from datetime import UTC, datetime, timedelta, timezone import pydantic import pytest @@ -155,7 +155,7 @@ def test_a_naive_datetime_input_comes_back_aware(): aware -- OptionalDateTime promises "An aware `datetime | None`" for every accepted input, not only for strings.""" got = _Probe.model_validate({"when": datetime(2026, 1, 1, 9, 0, 0)}).when - assert got == datetime(2026, 1, 1, 9, 0, 0, tzinfo=timezone.utc) + assert got == datetime(2026, 1, 1, 9, 0, 0, tzinfo=UTC) # --- the caller's own timestamp formats, opt-in and empty by default --------- @@ -192,7 +192,7 @@ def test_a_named_format_is_accepted_and_stamped_utc(): {"when": "17/08/2026 15:40:00"}, context={"datetime_input_formats": ["%d/%m/%Y %H:%M:%S"]}, ).when - assert got == datetime(2026, 8, 17, 15, 40, 0, tzinfo=timezone.utc) + assert got == datetime(2026, 8, 17, 15, 40, 0, tzinfo=UTC) def test_a_context_format_never_shadows_the_native_iso_form(): @@ -254,8 +254,9 @@ def test_extra_payload_serializes_verbatim_without_prefix() -> None: def test_extra_payload_overrides_a_declared_field() -> None: """A caller reaching past the model wins; losing silently would be worse.""" - payload = PostRequest(catalog_code="X", title="declared", - extra_payload={"title": "override"}) + payload = PostRequest( + catalog_code="X", title="declared", extra_payload={"title": "override"} + ) assert payload.to_api()["title"] == "override" diff --git a/easyvista_python_client/models/tests/test_request.py b/easyvista_python_client/models/tests/test_request.py index 38fffef..ecebadc 100644 --- a/easyvista_python_client/models/tests/test_request.py +++ b/easyvista_python_client/models/tests/test_request.py @@ -1,4 +1,4 @@ -from datetime import datetime, timedelta, timezone +from datetime import UTC, datetime, timedelta import pydantic import pytest @@ -135,8 +135,8 @@ def test_request_declares_title_and_core_scalars(): assert req.owner_id == 14 assert req.external_reference == "REF-1" # No offset in the fixture -> parse_ev_datetime treats it as UTC. - assert req.submit_date_ut == datetime(2026, 1, 1, 9, 0, 0, tzinfo=timezone.utc) - assert req.last_update == datetime(2026, 1, 2, 10, 30, 0, tzinfo=timezone.utc) + assert req.submit_date_ut == datetime(2026, 1, 1, 9, 0, 0, tzinfo=UTC) + assert req.last_update == datetime(2026, 1, 2, 10, 30, 0, tzinfo=UTC) def test_request_coerces_empty_string_numerics_to_none(): @@ -220,15 +220,9 @@ def test_request_declares_the_official_time_fields(): } ) # No offset in the fixtures -> parse_ev_datetime treats them as UTC. - assert ticket.creation_date_ut == datetime( - 2026, 7, 28, 9, 0, 0, tzinfo=timezone.utc - ) - assert ticket.max_resolution_date_ut == datetime( - 2026, 7, 30, 9, 0, 0, tzinfo=timezone.utc - ) - assert ticket.expected_date_ut == datetime( - 2026, 7, 29, 9, 0, 0, tzinfo=timezone.utc - ) + assert ticket.creation_date_ut == datetime(2026, 7, 28, 9, 0, 0, tzinfo=UTC) + assert ticket.max_resolution_date_ut == datetime(2026, 7, 30, 9, 0, 0, tzinfo=UTC) + assert ticket.expected_date_ut == datetime(2026, 7, 29, 9, 0, 0, tzinfo=UTC) assert ticket.end_date_ut is None # "" sentinel assert ticket.sla_id == 4 assert ticket.time_used_to_solve_request == "3600" @@ -431,9 +425,7 @@ def test_department_id_and_recipient_id_accept_the_documented_string() -> None: longer rewritten: the string reaches the wire as written, and an int still reaches it as an int. """ - body = PostRequest( - catalog_code="X", department_id="9", recipient_id="42" - ).to_api() + body = PostRequest(catalog_code="X", department_id="9", recipient_id="42").to_api() assert body["department_id"] == "9" assert body["recipient_id"] == "42" ints = PostRequest(catalog_code="X", department_id=9, recipient_id=42).to_api() diff --git a/easyvista_python_client/testing/test_unasync_codegen.py b/easyvista_python_client/testing/test_unasync_codegen.py index b35157e..0fe3c41 100644 --- a/easyvista_python_client/testing/test_unasync_codegen.py +++ b/easyvista_python_client/testing/test_unasync_codegen.py @@ -32,6 +32,7 @@ import ast import importlib.util import pathlib +import tomllib import pytest @@ -124,8 +125,8 @@ def _rewritable_names(node: ast.AST, keys: set[str]) -> list[str]: (``ast.MatchAs.name``), ``case [*AsyncRetrying]:`` (``ast.MatchStar.name``) and ``case {**aclose}:`` (``ast.MatchMapping.rest``) each bind a plain name the same way, and none of them is an ``ast.Name`` either. ``match`` - is valid on this project's 3.10 floor, so it is a live construct even - though nothing in the tree uses it today. + is valid on every Python this project supports (it arrived in 3.10), so + it is a live construct even though nothing in the tree uses it today. The exemption for a deliberate rename is scoped to the *specific* role it is legitimate in -- a class name for ``AsyncEasyvistaClient``, a @@ -262,11 +263,6 @@ def test_coverage_omits_the_generated_modules_and_nothing_else(build): ratio. A *hand-written* module wrongly listed in ``omit`` is measured by nothing, which is the defect this test was written for. """ - try: - import tomllib - except ModuleNotFoundError: # Python 3.10 - import tomli as tomllib - config = tomllib.loads((_REPO_ROOT / "pyproject.toml").read_text(encoding="utf-8")) omit = set(config["tool"]["coverage"]["run"]["omit"]) diff --git a/easyvista_python_client/tests/test_reporting.py b/easyvista_python_client/tests/test_reporting.py index 0b05674..f364913 100644 --- a/easyvista_python_client/tests/test_reporting.py +++ b/easyvista_python_client/tests/test_reporting.py @@ -1,4 +1,4 @@ -from datetime import datetime, timezone +from datetime import UTC, datetime import pytest @@ -12,7 +12,8 @@ def test_parse_offset_with_3_digit_milliseconds(): - # EasyVista's CREATION_DATE_UT format; 3.10's fromisoformat rejects 3-digit ms. + # EasyVista's CREATION_DATE_UT format, 3-digit ms (which 3.10's + # fromisoformat rejected, back when 3.10 was supported). dt = _parse_iso_datetime("2025-11-28T11:35:22.900+01:00") assert dt is not None assert dt.year == 2025 and dt.month == 11 and dt.day == 28 @@ -21,7 +22,7 @@ def test_parse_offset_with_3_digit_milliseconds(): def test_parse_trailing_z_is_utc(): dt = _parse_iso_datetime("2025-01-02T03:04:05Z") - assert dt == datetime(2025, 1, 2, 3, 4, 5, tzinfo=timezone.utc) + assert dt == datetime(2025, 1, 2, 3, 4, 5, tzinfo=UTC) def test_parse_no_fraction(): @@ -31,13 +32,13 @@ def test_parse_no_fraction(): def test_parse_naive_string_becomes_utc(): dt = _parse_iso_datetime("2025-06-15T08:00:00") - assert dt == datetime(2025, 6, 15, 8, 0, 0, tzinfo=timezone.utc) + assert dt == datetime(2025, 6, 15, 8, 0, 0, tzinfo=UTC) def test_parse_datetime_passthrough_makes_naive_utc(): naive = datetime(2025, 6, 15, 8, 0, 0) - assert _parse_iso_datetime(naive) == naive.replace(tzinfo=timezone.utc) - aware = datetime(2025, 6, 15, 8, 0, 0, tzinfo=timezone.utc) + assert _parse_iso_datetime(naive) == naive.replace(tzinfo=UTC) + aware = datetime(2025, 6, 15, 8, 0, 0, tzinfo=UTC) assert _parse_iso_datetime(aware) == aware diff --git a/easyvista_python_client/tests/test_timestamps.py b/easyvista_python_client/tests/test_timestamps.py index af22e09..f98c943 100644 --- a/easyvista_python_client/tests/test_timestamps.py +++ b/easyvista_python_client/tests/test_timestamps.py @@ -65,13 +65,15 @@ def test_format_refuses_a_naive_datetime(): def test_the_iso_basic_form_is_refused_on_every_python(basic): """Separator-less ISO input must return ``None`` regardless of interpreter. - This is a portability guard, not a formatting preference. ``fromisoformat`` - accepts the ISO "basic" form from Python 3.11 and rejects it on 3.10, and - this package supports 3.10 through 3.14 -- so before the explicit refusal - the *same wire value* parsed to an instant on four of the five supported + This began as a portability guard, not a formatting preference. + ``fromisoformat`` accepts the ISO "basic" form from Python 3.11 and rejects + it on 3.10, and when this package still supported 3.10 through 3.14 the + *same wire value* parsed to an instant on four of the five supported versions and raised on the fifth. CI caught it exactly that way: the 3.10 job was green while 3.11 and 3.12 failed ``test_a_numeric_shaped_value_raises_instead_of_becoming_an_epoch_instant``. + With 3.11 as the floor every supported interpreter accepts these, so the + explicit refusal is now the only thing between them and a parsed instant. EasyVista's timestamps always carry separators, so none of these is one of its values on any interpreter, and accepting them would let a genuine diff --git a/easyvista_python_client/timestamps.py b/easyvista_python_client/timestamps.py index fc5be40..ad259d6 100644 --- a/easyvista_python_client/timestamps.py +++ b/easyvista_python_client/timestamps.py @@ -33,7 +33,7 @@ from __future__ import annotations import re -from datetime import datetime, timezone +from datetime import UTC, datetime from typing import Any _FRACTION_RE = re.compile(r"\.(\d+)") @@ -48,18 +48,24 @@ def parse_ev_datetime(value: Any) -> datetime | None: """Parse an EasyVista timestamp to a timezone-aware ``datetime``, or ``None``. Accepts a ``datetime`` (returned as-is; a naive one is treated as UTC) or an - ISO-8601 string. Normalizes for Python 3.10's stricter ``fromisoformat``: - maps a trailing ``Z`` to ``+00:00`` and pads/truncates fractional seconds to - 6 digits — EasyVista sends 3, which 3.10 rejects outright. Unparseable input - returns ``None`` rather than raising, so a single malformed column never - fails a whole record. + ISO-8601 string. Before ``fromisoformat`` it maps a trailing ``Z`` or ``z`` + to ``+00:00`` and pads/truncates fractional seconds to 6 digits. Both rules + date from the Python 3.10 floor, whose ``fromisoformat`` rejected + EasyVista's 3-digit fraction outright. From 3.11 ``fromisoformat`` takes a + ``Z`` and any fraction length itself but still refuses a lowercase ``z`` + (measured 2026-10-02 on 3.11.13 and 3.14.6), so the rules are kept and the + set of accepted values did not move when 3.10 was dropped. Unparseable + input returns ``None`` rather than raising, so a single malformed column + never fails a whole record. **A value must start with an extended ISO date** (``YYYY-MM-DD``) or it is - refused, on every interpreter. From 3.11 ``fromisoformat`` also accepts the - ISO *basic* forms — ``"20260817"``, ``"20260817T154041.610"``, week dates - like ``"2026W331"`` — which 3.10 rejects, so without this rule the same wire - value parsed to an instant on four of the five supported Pythons and raised - on the fifth. CI found it precisely that way: 3.10 green, 3.11 and 3.12 red. + refused, on every interpreter. ``fromisoformat`` on every supported Python + (3.11+) also accepts the ISO *basic* forms — ``"20260817"``, + ``"20260817T154041.610"``, week dates like ``"2026W331"`` — so without this + rule each would parse to a plausible instant. The rule predates the 3.11 + floor: 3.10 rejected those forms, so the same wire value then parsed to an + instant on four of the five supported Pythons and raised on the fifth, and + CI found it precisely that way: 3.10 green, 3.11 and 3.12 red. The rule is stated positively because the reject-list version of it was wrong: "digits only" catches ``"20260817"`` and misses both a basic @@ -71,7 +77,7 @@ def parse_ev_datetime(value: Any) -> datetime | None: which is tried after this returns ``None``. """ if isinstance(value, datetime): - return value if value.tzinfo else value.replace(tzinfo=timezone.utc) + return value if value.tzinfo else value.replace(tzinfo=UTC) if not isinstance(value, str) or not value.strip(): return None text = value.strip() @@ -92,7 +98,7 @@ def parse_ev_datetime(value: Any) -> datetime | None: parsed = datetime.fromisoformat(text) except ValueError: return None - return parsed if parsed.tzinfo else parsed.replace(tzinfo=timezone.utc) + return parsed if parsed.tzinfo else parsed.replace(tzinfo=UTC) def format_ev_datetime(value: datetime) -> str: diff --git a/integration_tests/test_live_change_window.py b/integration_tests/test_live_change_window.py index 9b3634a..4a91c4f 100644 --- a/integration_tests/test_live_change_window.py +++ b/integration_tests/test_live_change_window.py @@ -16,7 +16,7 @@ from __future__ import annotations import uuid -from datetime import timedelta, timezone +from datetime import UTC, timedelta from itertools import pairwise import pytest @@ -299,8 +299,8 @@ def test_only_some_timestamp_renderings_are_accepted_as_an_interval_bound( later = parse_ev_datetime(late) assert moment is not None, "split_instants did not yield a parseable literal" assert later is not None, "split_instants did not yield a parseable literal" - as_utc = moment.astimezone(timezone.utc) - later_utc = later.astimezone(timezone.utc) + as_utc = moment.astimezone(UTC) + later_utc = later.astimezone(UTC) # For the date-only rendering the second bound is the day AFTER the late # instant, not its own day: on an instance whose sampled stamps all fall on # one day the two dates would otherwise be equal and the differential empty. diff --git a/pyproject.toml b/pyproject.toml index 8824176..4e0e9b9 100644 --- a/pyproject.toml +++ b/pyproject.toml @@ -47,7 +47,7 @@ name = "easyvista-python-client" version = "0.4.0" description = "Typed Python client for the EasyVista Service Manager REST API" readme = "README.md" -requires-python = ">=3.10" +requires-python = ">=3.11" license = "MIT" license-files = ["LICENSE"] authors = [{ name = "easyvista-python-client contributors" }] @@ -57,7 +57,6 @@ classifiers = [ "Intended Audience :: Developers", "Operating System :: OS Independent", "Programming Language :: Python :: 3", - "Programming Language :: Python :: 3.10", "Programming Language :: Python :: 3.11", "Programming Language :: Python :: 3.12", "Programming Language :: Python :: 3.13", @@ -70,7 +69,6 @@ dependencies = [ "httpx>=0.27", "pydantic>=2.8", "tenacity>=8.2", - "typing-extensions>=4.7; python_version < '3.11'", ] [project.urls] @@ -110,7 +108,6 @@ dev = [ "pre-commit>=4.0", "sphinx>=7.2,<8.2", "sphinx-rtd-theme>=2.0", - "tomli>=2.0; python_version < '3.11'", # >=7.0 for Metadata-Version 2.5 support: hatchling >=1.32 stamps 2.5, and # twine <=6.2 monkeypatches packaging's valid-metadata list to end at 2.4, so # `twine check` fails a perfectly good wheel and sdist (measured). @@ -142,7 +139,6 @@ docs = [ "numpydoc>=1.8", "sphinx>=7.2,<8.2", "sphinx-rtd-theme>=2.0", - "tomli>=2.0; python_version < '3.11'", ] [tool.hatch.build.targets.wheel] @@ -178,7 +174,7 @@ include = [ [tool.ruff] line-length = 88 -target-version = "py310" +target-version = "py311" # The sync client tree is generated by unasync_build.py from _async/ and must # stay byte-identical to what regeneration produces -- that identity is what # the CI gate checks. The conflict is narrow but permanent: stripping the @@ -193,7 +189,7 @@ extend-exclude = ["easyvista_python_client/_sync"] select = ["B", "E", "F", "I", "RUF", "UP"] [tool.mypy] -python_version = "3.10" +python_version = "3.11" strict = true ignore_missing_imports = true exclude = ["tests/", "integration_tests/", "testing/"] diff --git a/scripts/run_coverage_gate.py b/scripts/run_coverage_gate.py index c6468e9..13b23e5 100644 --- a/scripts/run_coverage_gate.py +++ b/scripts/run_coverage_gate.py @@ -116,7 +116,11 @@ def main() -> int: "Could not find an interpreter with the project's dev dependencies " "installed, so the coverage gate did not run. Tried:\n " + "\n ".join(str(path) for path in tried) - + "\n\nCreate the environment CONTRIBUTING.md describes:\n" + + "\n\nAn interpreter older than requires-python in pyproject.toml " + "(3.11) fails this probe too, even with every dependency installed: " + "the package itself no longer imports there.\n\n" + "Create the environment CONTRIBUTING.md describes, on Python 3.11 or " + "newer:\n" " python -m venv .venv\n" ' .venv\\Scripts\\python.exe -m pip install -e ".[dev]"', file=sys.stderr, diff --git a/skills/easyvista-asset-workflow/SKILL.md b/skills/easyvista-asset-workflow/SKILL.md index 4f77c52..41a8753 100644 --- a/skills/easyvista-asset-workflow/SKILL.md +++ b/skills/easyvista-asset-workflow/SKILL.md @@ -2,7 +2,7 @@ name: easyvista-asset-workflow description: "Create, fetch, search and iterate EasyVista assets with easyvista_python_client — create_asset, get_asset, search_assets and iter_assets with PostAsset and Asset. Use for equipment, hardware or CI records: registering a new asset, looking one up by tag, or listing a department's assets." license: MIT -compatibility: "Requires Python 3.10+, easyvista-python-client, network access to an EasyVista Service Manager REST API, and a profile authorized for the assets resource." +compatibility: "Requires Python 3.11+, easyvista-python-client, network access to an EasyVista Service Manager REST API, and a profile authorized for the assets resource." metadata: package: easyvista-python-client version: "0.4.0" diff --git a/skills/easyvista-client-setup/SKILL.md b/skills/easyvista-client-setup/SKILL.md index d6efa9d..c66cea8 100644 --- a/skills/easyvista-client-setup/SKILL.md +++ b/skills/easyvista-client-setup/SKILL.md @@ -2,7 +2,7 @@ name: easyvista-client-setup description: "Create and configure the synchronous easyvista_python_client.EasyvistaClient or the asynchronous AsyncEasyvistaClient — server/account/api_version, Bearer token or HTTP Basic credentials, EasyvistaConfig.from_env, timeouts, retries, TLS verification, default page size, and the EasyvistaError hierarchy. Use before calling any EasyVista API, or when the user asks how to connect to EasyVista with easyvista_python_client." license: MIT -compatibility: "Requires Python 3.10+, easyvista-python-client, network access to an EasyVista Service Manager REST API, and valid EasyVista credentials." +compatibility: "Requires Python 3.11+, easyvista-python-client, network access to an EasyVista Service Manager REST API, and valid EasyVista credentials." metadata: package: easyvista-python-client version: "0.4.0" diff --git a/skills/easyvista-directory/SKILL.md b/skills/easyvista-directory/SKILL.md index 8207d24..ac21e27 100644 --- a/skills/easyvista-directory/SKILL.md +++ b/skills/easyvista-directory/SKILL.md @@ -2,7 +2,7 @@ name: easyvista-directory description: "Look up and provision EasyVista departments and employees with easyvista_python_client — get_department, search_departments, iter_departments, find_departments, get_department_comment, create_department, update_department and the matching employee methods, plus Reference and FieldClassification for reading instance-specific columns. Use to resolve a department by name or code, list a department's people, read a directory memo, or create/update directory records." license: MIT -compatibility: "Requires Python 3.10+, easyvista-python-client, network access to an EasyVista Service Manager REST API, and a profile authorized for the departments and employees resources (writes are additionally profile-gated)." +compatibility: "Requires Python 3.11+, easyvista-python-client, network access to an EasyVista Service Manager REST API, and a profile authorized for the departments and employees resources (writes are additionally profile-gated)." metadata: package: easyvista-python-client version: "0.4.0" diff --git a/skills/easyvista-document-workflow/SKILL.md b/skills/easyvista-document-workflow/SKILL.md index 420a43b..85a4ecd 100644 --- a/skills/easyvista-document-workflow/SKILL.md +++ b/skills/easyvista-document-workflow/SKILL.md @@ -2,7 +2,7 @@ name: easyvista-document-workflow description: "Attach, list, download, stream and delete files on an EasyVista ticket with easyvista_python_client — add_document, list_documents, download_document, stream_document and delete_document with the Document model. Use for ticket attachments, uploading evidence or logs to a request, fetching an attachment's bytes whole or chunk by chunk without buffering a large file, or removing one." license: MIT -compatibility: "Requires Python 3.10+, easyvista-python-client, network access to an EasyVista Service Manager REST API, and a profile authorized for the documents sub-resource." +compatibility: "Requires Python 3.11+, easyvista-python-client, network access to an EasyVista Service Manager REST API, and a profile authorized for the documents sub-resource." metadata: package: easyvista-python-client version: "0.4.0" diff --git a/skills/easyvista-instance-discovery/SKILL.md b/skills/easyvista-instance-discovery/SKILL.md index 812ef1c..19e7c55 100644 --- a/skills/easyvista-instance-discovery/SKILL.md +++ b/skills/easyvista-instance-discovery/SKILL.md @@ -2,7 +2,7 @@ name: easyvista-instance-discovery description: "Discover what one EasyVista deployment actually exposes with easyvista_python_client — get_api_spec reads the instance's own OpenAPI, list_reference_table reads any list route into column-free records, discover resolves one reference name to the ids/labels/codes/GUIDs in use, and describe_instance profiles the lot into an InstanceProfile. Use before hardcoding any id, when a ticket create is rejected for an unknown catalog, urgency, impact or group, when you need a STATUS_GUID for set_status or close_ticket, or when you need to know which routes a deployment declares at all." license: MIT -compatibility: "Requires Python 3.10+, easyvista-python-client, and network access to an EasyVista Service Manager REST API. Every call here is a GET; nothing is created, updated or deleted." +compatibility: "Requires Python 3.11+, easyvista-python-client, and network access to an EasyVista Service Manager REST API. Every call here is a GET; nothing is created, updated or deleted." metadata: package: easyvista-python-client version: "0.4.0" diff --git a/skills/easyvista-reporting-and-context/SKILL.md b/skills/easyvista-reporting-and-context/SKILL.md index 3317705..bf1d543 100644 --- a/skills/easyvista-reporting-and-context/SKILL.md +++ b/skills/easyvista-reporting-and-context/SKILL.md @@ -2,7 +2,7 @@ name: easyvista-reporting-and-context description: "Aggregate EasyVista tickets into counts and per-dimension breakdowns, and assemble one-call context bundles, with easyvista_python_client — count_tickets, ticket_statistics, aggregate_tickets, TicketStatistics, get_ticket_context, TicketContext.to_markdown and get_department_context. Use for ticket dashboards, per-status or per-department counts, and for exporting a ticket or a department as an LLM-ready document." license: MIT -compatibility: "Requires Python 3.10+, easyvista-python-client, and network access to an EasyVista Service Manager REST API." +compatibility: "Requires Python 3.11+, easyvista-python-client, and network access to an EasyVista Service Manager REST API." metadata: package: easyvista-python-client version: "0.4.0" diff --git a/skills/easyvista-search-syntax/SKILL.md b/skills/easyvista-search-syntax/SKILL.md index 464900c..a880b19 100644 --- a/skills/easyvista-search-syntax/SKILL.md +++ b/skills/easyvista-search-syntax/SKILL.md @@ -2,7 +2,7 @@ name: easyvista-search-syntax description: "Write correct EasyVista server-side search expressions for search_tickets, iter_tickets, count_tickets, search_assets, search_departments and search_employees using ev_equals_filter, ev_in_filter, ev_contains_filter, ev_starts_with_filter, ev_since_filter, ev_between_filter, escape_ev_value and is_safe_ev_value. Use whenever building a search= argument, filtering EasyVista records, filtering by a date/time window, or debugging a filter that returned everything or nothing — EasyVista silently ignores conditions it cannot honour and returns the whole table." license: MIT -compatibility: "Requires Python 3.10+, easyvista-python-client, and network access to an EasyVista Service Manager REST API." +compatibility: "Requires Python 3.11+, easyvista-python-client, and network access to an EasyVista Service Manager REST API." metadata: package: easyvista-python-client version: "0.4.0" diff --git a/skills/easyvista-ticket-actions/SKILL.md b/skills/easyvista-ticket-actions/SKILL.md index e3ee712..18d16e9 100644 --- a/skills/easyvista-ticket-actions/SKILL.md +++ b/skills/easyvista-ticket-actions/SKILL.md @@ -2,7 +2,7 @@ name: easyvista-ticket-actions description: "Read and write the action log on an EasyVista ticket with easyvista_python_client — create_task and PostTask (the one call that posts a COMMENT: a task is an action born already ended, so its text shows in the history), plus create_action, end_action (an action is born OPEN and its text does not show until ended), list_actions, iter_actions, get_action and update_action with PostAction, Action and ActionUpdate. Covers why there is no private-comment flag and that visibility is the action TYPE instead, how to recover a created action's id, how to page a whole log past the one-page cap, and how to resolve an action's note text, which the list endpoint does not return. Use for ticket comments, followups, work notes, internal or private comments, progress entries or any per-ticket action history." license: MIT -compatibility: "Requires Python 3.10+, easyvista-python-client, network access to an EasyVista Service Manager REST API, and a profile authorized for the actions sub-resource." +compatibility: "Requires Python 3.11+, easyvista-python-client, network access to an EasyVista Service Manager REST API, and a profile authorized for the actions sub-resource." metadata: package: easyvista-python-client version: "0.4.0" diff --git a/skills/easyvista-ticket-workflow/SKILL.md b/skills/easyvista-ticket-workflow/SKILL.md index 8b92b8e..065b0e1 100644 --- a/skills/easyvista-ticket-workflow/SKILL.md +++ b/skills/easyvista-ticket-workflow/SKILL.md @@ -2,7 +2,7 @@ name: easyvista-ticket-workflow description: "Create, read, search, paginate, update and close EasyVista tickets (requests) with easyvista_python_client — PostRequest, Request, RequestUpdate, create_ticket, create_tickets, get_ticket, search_tickets, iter_tickets, count_tickets, update_ticket and close_ticket. Use for any ticket/incident/request operation, including discovering the instance-specific catalog codes and ids a create needs." license: MIT -compatibility: "Requires Python 3.10+, easyvista-python-client, network access to an EasyVista Service Manager REST API, and a profile authorized for the requests resource." +compatibility: "Requires Python 3.11+, easyvista-python-client, network access to an EasyVista Service Manager REST API, and a profile authorized for the requests resource." metadata: package: easyvista-python-client version: "0.4.0" From e856ed1007de693fdf01376480efd03a923fec5d Mon Sep 17 00:00:00 2001 From: baraline <antoine.guillaume45@gmail.com> Date: Fri, 2 Oct 2026 15:33:58 +0200 Subject: [PATCH 09/11] feat(content): rebuild the converter on markdownify, mdformat and cmark-gfm The content extra followed glpi_python_client's python-markdown design (4fc3bed). glpi_python_client replaced that design in 917f030: markdownify plus mdformat inbound, cmark-gfm outbound. This package now carries that converter, plus 15 fixes measured on this package's synthetic corpus and on a 367-memo preproduction sample: alt-text escapes, ~~~ and ! before a link, a ^ alt, a second <, <br> in inline code, tables in headings and links, <center>, raw <u>/<mark>/<ins>, linear list numbering and edge splitting, unreadable colspan/start, and the CVE-2025-6069 tail. The extra now installs beautifulsoup4>=4.15, cmarkgfm>=2025.10.22, markdown-it-py>=3.0,<4, markdownify>=1.2.3,<1.3, mdformat>=0.7.22,<0.8 and mdformat-tables>=1.0,<1.1; markdown is gone. Tests: GLPI's display oracle, round-trip and contract tests are ported, plus: - one regression test per fix (GLPI 917f030 fails 36 of 54); - seeded property families, including underline/highlight and blocks in flattened cells (GLPI fails 39-100% of each); - growth tests asserting time(4n)/time(n) < 8: a linear pass reads 2.4-6.8, a reverted fix 11-37; - the "not a sanitiser" contract, pinned exactly; - the 107 HTML bodies of this package's earlier regression tests, held to display plus fixed point. The pre-commit mypy hook installs the extra's four typed packages with pyproject's bounds, and a test keeps them in step; cmarkgfm and mdformat-tables add no types. Nothing checked that requires-python, the classifiers and the two CI matrices agree, so dropping 3.10 had to chase four lists by hand: test_the_supported_pythons_agree_everywhere_they_are_written now binds them. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> --- .pre-commit-config.yaml | 20 +- easyvista_python_client/content/__init__.py | 7 +- easyvista_python_client/content/conversion.py | 2859 ++++------------- .../content/tests/display.py | 325 ++ .../content/tests/test_conversion.py | 1782 +--------- .../content/tests/test_cost.py | 205 ++ .../content/tests/test_fixes.py | 558 ++++ .../content/tests/test_literal_text.py | 1553 --------- .../content/tests/test_properties.py | 786 +++++ .../content/tests/test_round_trip.py | 630 ++++ easyvista_python_client/exceptions.py | 10 +- .../testing/test_public_api.py | 149 +- pyproject.toml | 40 +- 13 files changed, 3434 insertions(+), 5490 deletions(-) create mode 100644 easyvista_python_client/content/tests/display.py create mode 100644 easyvista_python_client/content/tests/test_cost.py create mode 100644 easyvista_python_client/content/tests/test_fixes.py delete mode 100644 easyvista_python_client/content/tests/test_literal_text.py create mode 100644 easyvista_python_client/content/tests/test_properties.py create mode 100644 easyvista_python_client/content/tests/test_round_trip.py diff --git a/.pre-commit-config.yaml b/.pre-commit-config.yaml index d4ed1f3..cfdf9a8 100644 --- a/.pre-commit-config.yaml +++ b/.pre-commit-config.yaml @@ -15,7 +15,25 @@ repos: rev: v2.3.0 hooks: - id: mypy - additional_dependencies: ["pydantic>=2.8", "httpx", "tenacity"] + # The runtime dependencies, then the TYPED packages of the `content` + # extra. The hook's venv holds only this list, so without those four + # every converter base class (MarkdownConverter, MDRenderer, + # RenderTreeNode, MarkdownIt) resolves to Any and strict mode's + # disallow_subclassing_any fails content/conversion.py even though + # `mypy easyvista_python_client` in CI is clean. The extra's other two, + # cmarkgfm and mdformat-tables, ship no type information, so they are + # Any either way and are left out; cmarkgfm would also make every + # contributor without a wheel for their platform compile it before + # any commit. Each line here must equal the extra's requirement in + # pyproject; testing/test_public_api.py fails if one drifts. + additional_dependencies: + - "pydantic>=2.8" + - "httpx" + - "tenacity" + - "beautifulsoup4>=4.15" + - "markdown-it-py>=3.0,<4" + - "markdownify>=1.2.3,<1.3" + - "mdformat>=0.7.22,<0.8" # Must mirror [tool.mypy] exclude in pyproject.toml. pre-commit passes # filenames explicitly, and mypy ignores its own `exclude` for files named # on the command line -- so this regex is the ONLY thing keeping the hook diff --git a/easyvista_python_client/content/__init__.py b/easyvista_python_client/content/__init__.py index ea0c022..fd04f8c 100644 --- a/easyvista_python_client/content/__init__.py +++ b/easyvista_python_client/content/__init__.py @@ -4,8 +4,11 @@ this subpackage without the extra raises :class:`ImportError` naming that command. Nothing else in the package imports it, so ``import easyvista_python_client`` needs none of the extra's dependencies, and -:class:`~easyvista_python_client.EasyvistaContentError` -- the one error the -converter raises -- lives in the core package, catchable either way. +:class:`~easyvista_python_client.EasyvistaContentError` -- the error the +converter raises -- lives in the core package, catchable either way. The +one exception: called from within a few frames of the recursion limit, +the converter can let a bare :class:`RecursionError` through, having no +stack left to report it otherwise (see ``docs/content.rst``). """ from __future__ import annotations diff --git a/easyvista_python_client/content/conversion.py b/easyvista_python_client/content/conversion.py index 0258174..029782d 100644 --- a/easyvista_python_client/content/conversion.py +++ b/easyvista_python_client/content/conversion.py @@ -1,160 +1,100 @@ -"""Content conversion between EasyVista memo HTML and canonical Markdown. - -**This module is a port.** It is ``glpi_python_client/content/conversion.py`` -from `glpi_python_client <https://github.com/baraline/glpi_python_client>`_ -at commit ``4fc3bed``, the literal-safe converter, and **the two should move -together**: a defect fixed in one is almost certainly present in the other, -because the hard part of both is the behaviour of the same three libraries -- -``beautifulsoup4``'s ``html.parser`` tree builder and ``markdownify`` inbound, -``python-markdown`` outbound -- and not anything either ITSM does. The code is -that file's, helper for helper, with the converter and its error renamed for -this package (``GlpiContentConverter`` is :class:`EasyvistaContentConverter`, -``GlpiContentError`` is :class:`~easyvista_python_client.EasyvistaContentError`) -and the error messages naming EasyVista. The one other change carries a comment -starting "Diverges from glpi_python_client", with the reason. Prose that stated -something true only of GLPI now says whose measurement it was, and the figures -that depend on the interpreter or on a library version were measured again -here. - -What it converts ----------------- +"""Content conversion between EasyVista memo HTML and the package's Markdown. EasyVista keeps rich text in *memo* fields -- a ticket's ``COMMENT`` and -``DESCRIPTION``, an action's ``DESCRIPTION`` -- which the client reads with -``resolve_memo``. A memo holds the HTML it was sent: measured 2026-09-30 on -one instance, a memo written through the API was stored byte for byte, and -the web UI rendered its ``<p>`` elements as paragraphs. That is tier 4 -- one -instance, one date, and it may not generalise; no vendor documentation of the -memo format is recorded in ``docs/vendor-api-reference.md``. Nothing forces a -memo to be HTML either, since a caller can write plain text, so -:meth:`EasyvistaContentConverter.from_transport` takes both: HTML is converted -and plain text comes back as it was. - -Literal text ------------- - -**Text in the HTML is literal, and the Markdown spells it so.** A user who -types ``__init__`` into a memo means eight characters; Markdown reads the -same eight as bold ``init``. So :meth:`EasyvistaContentConverter.from_transport` -escapes a character exactly where python-markdown -- with this module's -extensions, which is what :meth:`EasyvistaContentConverter.to_transport` and -any peer rendering the Markdown uses -- would otherwise read it as syntax, and -nowhere else. It has to happen here: once the Markdown is written, nothing -can tell a literal ``__`` from a bold one any more. The rule set is -documented on :class:`_LiteralSafeConverter`; the property it is held to is -that ``to_transport(from_transport(html))`` displays what ``html`` -displays, and that the Markdown is a fixed point of the round trip. - -Ordinary prose carries no escape at all -- ``fichier_de_test_v2.xlsx``, -``C:\\Temp\\logs``, a ``#`` or a ``-`` mid-sentence, ``R&D`` -- because -python-markdown already renders those literally. What is escaped is what it -would not: ``\\\\serveur`` (it would lose a backslash), ``__init__``, a -``#4521`` or a ``> merci`` at the start of a line, a ``|`` in a table cell, -``<Enter>``, ``&``. None of this is visible to anyone reading the memo -in either ITSM: the escapes exist only in the Markdown between them. - -It is not a sanitiser, which is a different job. Inbound, text a memo -*displays* as markup -- ``<script>`` -- now comes back escaped, as the -literal text it is, where it used to come back as a raw ``<script>``; but -Markdown a caller writes passes through -:meth:`~EasyvistaContentConverter.to_transport` untouched, raw HTML -included, and a ``javascript:`` link target is rendered as a live ``href``. -Markdown a caller did not write needs those neutralised before it is -rendered, and that is the caller's job. - -Nesting depth -------------- - -Neither *parse* recurses -- ``html.parser`` is an iterative scanner -- but -``markdownify`` walks the finished tree recursively, so inbound conversion has -a nesting ceiling. Measured 2026-09-30 from a shallow stack against the -default 1000-frame limit, with ``markdownify`` 1.2.3 and ``beautifulsoup4`` -4.15.0: the deepest ``<div>`` document that converts is 493 levels on CPython -3.12, 3.13 and 3.14, and 328 on CPython 3.10, where each level costs about -three frames instead of two. (glpi_python_client recorded 494 for ``<div>``, -``<p>``, ``<blockquote>`` and ``<table><tr><td>``.) The per-level cost is -what differs between interpreters, which is one more reason the number is not -something to build on. - -**The ceiling is discovered rather than predicted.** Inbound conversion is -attempted, and a ``RecursionError`` is caught and answered by stripping the -document to its text instead -- see -:meth:`EasyvistaContentConverter.from_transport`. glpi_python_client's -earlier design estimated the depth up front and degraded past a fixed bound, -and it was wrong in both directions: it degraded bodies that would have -converted, because the bound had to assume the worst about the caller's -remaining stack, and three rounds of review found seven ways for the estimate -to come in *under* the real tree, each of which put a document through -``markdownify`` and into the ``RecursionError`` the bound existed to prevent. -Trying the conversion cannot be wrong about whether the conversion fits. - -Anything else that goes wrong in either direction surfaces as -:class:`~easyvista_python_client.EasyvistaContentError`, so no parser fault -escapes the package's exception taxonomy. - -**``sys.setrecursionlimit`` is deliberately not called, here or anywhere in -the package.** It is process-global state that belongs to the application, -not to a library an application imported; and raising the limit past what the -C stack can hold turns a catchable ``RecursionError`` into a hard interpreter -crash -- on Windows, an access violation with no traceback. It moves the -cliff and makes falling off it worse. Degrading the one body that does not fit -is the answer that does not. Running the walk in a thread with a larger stack -was considered and rejected for the same reason: the recursion limit is a -counter rather than a measurement of the stack, so a deeper thread still needs -the global limit raised to use it. +``DESCRIPTION``, an action's ``DESCRIPTION`` and ``COMMENT`` -- and the +package's surface is Markdown. A memo holds what it was sent: measured +2026-09-30 on one instance, a memo written through the API was stored byte +for byte (tier 4: one instance, one date, and it may not generalise). So a +memo is HTML or plain text, and both are read here. + +* **Read** -- :meth:`EasyvistaContentConverter.from_transport`. ``markdownify`` + writes the HTML as Markdown with every character that could be syntax + escaped, then ``mdformat`` re-renders that Markdown from its syntax tree, + which keeps only the escapes CommonMark needs. Text is literal: + ``__init__`` typed into a memo reads back as ``\\_\\_init\\_\\_``. +* **Write** -- :meth:`EasyvistaContentConverter.to_transport`. ``cmark-gfm`` + renders CommonMark with GFM tables; a newline is a line break and raw + HTML passes through. + +The aim is that rendering the Markdown a read gives displays what the memo +displayed, and that reading that back gives the same Markdown. On the 367 +memos of the one sample measured (tier 4, one instance, may not +generalise), every word was kept and every memo was a fixed point, and the +65 that displayed differently did so in the ways ``docs/content.rst`` +lists. Some synthetic shapes still lose words or the fixed point; that +page lists them too. The glue below covers what the three libraries leave +out: plain text, line breaks a browser does not show, bold and italic +CommonMark would not close, link targets, and an mdformat set up without +its nesting cap and with its quadratic lookups made linear. A body nested +too deeply for the stack is read as its text (:func:`_text_of`); anything +else that fails raises +:class:`~easyvista_python_client.EasyvistaContentError`. + +A memo with no HTML element in it is read as literal lines, a line break +per line. How EasyVista's web UI displays such a memo is unverified: the +vendor's form editor documents a MEMO object (text) beside a TEXT AREA +object that accepts HTML [doc:https://docs.easyvista.com/docs/form], and +its comment-log page configures a custom field gathering a request's +comments as a "Text area" +[doc:https://docs.easyvista.com/docs/service-manager-comment-log-creation]. +That leans towards the UI showing memo text as HTML, but neither page says +which object the built-in description or the action history is. + +**This module is a port** of ``glpi_python_client/content/conversion.py`` at +commit ``917f030``, plus fifteen fixes measured on this package's corpus +and still to be proposed to it. The two should move together: the hard part +of both is the behaviour of the same libraries, not anything either ITSM +does. The code is that file's, helper for helper, with +``GlpiContentConverter`` renamed :class:`EasyvistaContentConverter`, +``GlpiContentError`` renamed +:class:`~easyvista_python_client.EasyvistaContentError`, the converter's +private markers renamed from ``glpi`` to ``ev`` and the messages naming +EasyVista. """ from __future__ import annotations import re -from collections.abc import Callable, Iterable -from html import unescape -from html.entities import html5 as _HTML5_REFERENCES -from html.parser import HTMLParser -from itertools import chain -from typing import Any, NamedTuple +import string +import unicodedata +from collections.abc import Callable, Iterator, Mapping, MutableMapping, Sequence +from functools import cached_property +from html import escape, unescape +from typing import Any, cast from easyvista_python_client.exceptions import EasyvistaContentError -# Diverges from glpi_python_client: there the three libraries are hard +# Diverges from glpi_python_client: there these libraries are hard # dependencies. Here they are the optional ``content`` extra, so that # ``import easyvista_python_client`` stays as light as it was before this # module existed; nothing outside this subpackage imports it. A missing one is -# answered with the command that installs all three, rather than with a bare -# ``No module named 'markdownify'`` that names one package and not the extra. +# answered with the command that installs them all, rather than with a bare +# ``No module named 'mdformat'`` that names one package and not the extra. try: - from bs4 import Comment, Doctype, ParserRejectedMarkup, Tag - from bs4.element import PageElement - from markdown import Markdown - from markdown import markdown as markdown_to_html - from markdown.blockprocessors import HRProcessor, ReferenceProcessor - from markdown.inlinepatterns import ( - AUTOLINK_RE, - AUTOMAIL_RE, - BACKTICK_RE, - NOT_STRONG_RE, - ) + import cmarkgfm + import mdformat_tables + from bs4 import BeautifulSoup, ParserRejectedMarkup, Tag + from bs4.element import PageElement, PreformattedString + from markdown_it import MarkdownIt + from markdown_it.token import Token from markdownify import MarkdownConverter + from mdformat.renderer import ( + DEFAULT_RENDERERS, + MDRenderer, + RenderContext, + RenderTreeNode, + ) + from mdformat.renderer.typing import Postprocess except ImportError as exc: raise ImportError( "easyvista_python_client.content needs the optional 'content' extra " - "(beautifulsoup4, markdown and markdownify). Install it with: " + "(beautifulsoup4, cmarkgfm, markdown-it-py, markdownify, mdformat and " + "mdformat-tables). Install it with: " 'pip install "easyvista-python-client[content]"' ) from exc -#: Element names that make a ``<...>`` sequence markup rather than text. -#: -#: The HTML5 element set, which is what the parser behind ``markdownify`` -#: will actually recognise. Anything outside it -- ``<Enter>``, ``<T>``, -#: ``</dev/null>`` -- parses as an *unknown* tag, whose markup is dropped -#: while its (usually empty) body is kept, so the token silently vanishes -#: from the middle of a sentence. -#: -#: The second group is the HTML standard's *obsolete* elements -- its list -#: of features that "must not be used by authors", which old editors and -#: e-mail clients write all the same. Without them a body marked up only with -#: ``<font color="red">URGENT</font>`` or ``<center>`` failed the probe, took -#: the plain-text path and kept its tags as text. +#: Element names that make a ``<...>`` sequence markup rather than text: the +#: HTML5 elements, then the obsolete ones old editors and mail clients write. _HTML_ELEMENTS = frozenset( """ a abbr address area article aside audio b base bdi bdo blockquote body br @@ -173,130 +113,81 @@ """.split() ) -#: python-markdown extensions applied when rendering outbound content. -#: -#: ``fenced_code`` and ``tables`` are here because without them the two -#: constructs do not survive at all. A fence rendered without -#: ``fenced_code`` becomes inline ``<code>``, which any HTML renderer that -#: does not style it as preformatted shows as one run-on line -- -#: glpi_python_client saw exactly that in GLPI's web UI -- and which a later -#: read writes back as inline code, so a pasted log degrades a little more on -#: every edit. A table without ``tables`` renders as literal pipe characters. -#: -#: A language tag is still lost: ``markdownify`` drops the -#: ``class="language-python"`` that ``fenced_code`` emits, so ```` ```python ```` -#: comes back as a bare fence. That is a limitation of the pair of -#: libraries, not something an extension list can fix. -_MARKDOWN_EXTENSIONS = ["nl2br", "sane_lists", "fenced_code", "tables"] - -#: One candidate tag: ``<`` or ``</`` immediately followed by a name. -#: -#: The ``<`` must abut the name, matching what an HTML parser accepts. That -#: is what keeps ``2 < 3 > 1`` and ``x <= y`` text: a space after ``<`` -#: means no tag, so arithmetic never reaches the HTML path in the first -#: place. +#: ``<`` or ``</`` right before a name: one candidate tag. _CANDIDATE_TAG = re.compile(r"</?([a-zA-Z][a-zA-Z0-9]*)\b[^<>]*>") -#: Elements that cannot contain anything, so nothing nests below them. -#: -#: ``html.parser`` -- the parser ``markdownify`` builds its tree with -- -#: closes these itself, so ``<br>`` a thousand times over is a thousand -#: siblings, not a thousand levels. Measured: ``"<br>" * 5000`` parses one -#: level deep and converts fine, while ``"<div>" * 5000`` parses 5000 deep -#: and raises. -#: -#: The HTML5 void set, plus the ten legacy names ``bs4``'s HTML-parser tree -#: builder also treats as empty. **The invariant is that this stays a -#: subset of what the parser treats as empty**, and the direction matters: -#: a name missing from here is counted as nesting when it does not, which -#: costs an unnecessary degradation, while a name wrongly *in* here hides -#: real nesting, which is a ``RecursionError``. Listing the legacy names -#: only makes the count exact on old markup. Copied rather than imported -- -#: it lives in ``bs4.builder`` as -#: ``HTMLTreeBuilder.DEFAULT_EMPTY_ELEMENT_TAGS``, which ``bs4``'s own -#: documentation marks ``:meta private:`` -- and the subset invariant is -#: asserted against a real parse in the unit tests, so a future ``bs4`` -#: cannot quietly break it. The two sets held the same 24 names when -#: glpi_python_client copied them, and still do on ``beautifulsoup4`` 4.15.0 -#: (checked 2026-09-30). -_VOID_ELEMENTS = frozenset( +#: Elements a browser lays out as blocks: a line ends at their edges. +_BLOCKS = frozenset( """ - area base basefont bgsound br col command embed frame hr image img - input isindex keygen link menuitem meta nextid param source spacer - track wbr + address article aside blockquote body caption center dd details dir div + dl dt fieldset figcaption figure footer form h1 h2 h3 h4 h5 h6 header hr + html li main menu nav ol p pre section summary table tbody td tfoot th + thead tr ul """.split() ) -#: Elements whose boundary becomes a line break when tags are stripped. -#: -#: Used only by :func:`_strip_tags`. Removing a block element outright runs -#: its neighbours together -- ``<p>a</p><p>b</p>`` becomes ``ab`` -- while -#: putting a separator at *every* tag breaks words apart, turning -#: ``<b>off</b>line`` into ``off line``. Splitting on the block/inline line -#: keeps both readable. An unrecognised name counts as inline, matching how -#: the normal path treats it: markup dropped, body kept in place. -_BLOCK_ELEMENTS = frozenset( - """ - address article aside blockquote br center col dd details dialog dir div - dl dt fieldset figcaption figure footer form h1 h2 h3 h4 h5 h6 header - hgroup hr li main menu nav ol p pre search section summary table tbody td - tfoot th thead tr ul - """.split() -) +#: Anything between angle brackets, for the parser-rejected fallback. +_ANY_TAG = re.compile(r"<[^<>]*>?") + +#: Elements whose content a browser does not display. +_HIDDEN = ["head", "script", "style", "template", "title"] +#: The elements whose ``*`` markers are checked, and the attributes that +#: carry the characters displayed on either side of one (:func:`_note_sides`). +_EMPHASIS = frozenset({"b", "strong", "em", "i"}) +_BEFORE, _AFTER = "data-ev-before", "data-ev-after" +_NUMBER = "data-ev-number" # an ordered item's number (_Converter.convert_li) -#: One character reference, with or without its terminating semicolon. -#: -#: Resolved by :func:`_resolve_references` rather than by ``html.unescape`` -#: over the whole string. ``unescape`` implements the HTML5 rule of -#: consuming the longest *known* name it can find, so a semicolon-less -#: reference that is a prefix of a longer unknown word gets split -- -#: measured, ``http://x/?a=1©right=2`` becomes -#: ``http://x/?a=1©right=2``, and a URL pasted into a memo is exactly -#: where ``©right=`` and ``¬anentity=`` occur. The parser behind the -#: converting path leaves those whole, so this does too. -_CHARACTER_REFERENCE = re.compile( - r"&(?:\#[0-9]+;?|\#[xX][0-9a-fA-F]+;?|[A-Za-z][A-Za-z0-9]*;?)" +#: Text split into leading line breaks and spaces, content, trailing ones. +#: The content ends on its last character that is neither, found greedily: +#: a lazy ``.*?`` rescanned the run after it at every step, quadratic. +_EDGES = re.compile( + r"((?:\\\n|\s)*)((?:.*(?:[^\\\s]|\\(?!\n)))?)((?:\\\n|\s)*)", re.DOTALL ) +#: A URL that is its own CommonMark autolink and that markdown-it leaves as it +#: is: printable ASCII it does not percent-encode. +_AUTOLINK = re.compile( + r"[A-Za-z][A-Za-z0-9+.-]{1,31}:[A-Za-z0-9;/?:@&=+$,\-_.!~*'()#%]*" +) -#: The ``/`` of a self-closing tag, with any space around it. -_VOID_SELF_CLOSE = re.compile(r"\s*/\s*>$") +_BACKTICK_RUN = re.compile("`{3,}") -def _resolve_references(text: str) -> str: - """Resolve character references the way the real parser would. +def _language(pre: Tag) -> str | None: + """Return the language cmark-gfm wrote on a fence as ``class="language-x"``.""" - A numeric reference always resolves. A named one resolves only when - the whole name is known, which is the difference from - ``html.unescape``: that consumes the longest known *prefix*, so - ``©right=2`` loses its ``©`` and leaves ``right=2`` behind. - See :data:`_CHARACTER_REFERENCE`. + for tag in (pre.find("code"), pre): + if isinstance(tag, Tag): + for name in tag.get_attribute_list("class"): + if name and name.startswith("language-"): + return str(name[len("language-") :]) + return None - Where the two rules differ this one keeps the reference literal, which - is the safe direction for :func:`_strip_tags`, whose promise is that no - text goes missing. - """ - def resolve(match: re.Match[str]) -> str: - token = match.group(0) - if token[1] == "#": - return unescape(token) - return unescape(token) if token[1:] in _HTML5_REFERENCES else token +#: ``markdownify`` options; the converter's overrides decide the rest. +_MARKDOWNIFY_OPTIONS: dict[str, Any] = { + "code_language_callback": _language, + "bullets": "-", + "escape_misc": True, + "heading_style": "atx", + "newline_style": "backslash", + "wrap": True, # a newline in HTML text is a space... + "wrap_width": None, # ...and no line is wrapped +} - return _CHARACTER_REFERENCE.sub(resolve, text) +#: cmark-gfm's options: a newline is a line break, raw HTML passes through. +_RENDER_OPTIONS = ( + cmarkgfm.cmark.Options.CMARK_OPT_HARDBREAKS + | cmarkgfm.cmark.Options.CMARK_OPT_UNSAFE +) def _looks_like_html(content: str) -> bool: - """Return whether ``content`` carries at least one real HTML element. - - Deciding on the element *name* rather than on the presence of angle - brackets is what separates markup from prose that merely contains - ``<`` and ``>``. It cannot separate them perfectly: ``a<b>c`` is - genuinely ambiguous, because ``b`` is both a real element and a - plausible variable, and no probe reading the text alone can resolve - that. It resolves every case where the name is not an element at all, - which is where the silent deletions came from. + """Return whether ``content`` holds at least one real HTML element. + + The element *name* decides, so ``use the <Enter> key`` and + ``if x<y then z>0`` are text. ``a<b>c`` is markup: ``b`` is an element. """ return any( @@ -305,2184 +196,628 @@ def _looks_like_html(content: str) -> bool: ) -class _ParserScan(HTMLParser): - """Walk a document with the parser that will convert it, not one like it. - - Both remaining questions this module asks about raw HTML -- which - self-closing void tags to rewrite, and what the text is when the tree - will not fit the stack -- are questions about ``html.parser``'s - dispatch. This subclass asks ``html.parser`` instead of describing - it. - - In glpi_python_client it replaced a regular expression that reproduced - that dispatch by imitation, and the imitation kept being wrong in ways - that showed up only after they shipped. A comment closes on - ``--\\s*>`` and not only on ``-->``, ``</ script>`` ends raw text, - ``<![IGNORE[`` opens a marked section, ``</ div foo>`` is a bogus - comment rather than an end tag, and ``<a href=/>`` leaves an element - *open* because the unquoted value swallows the ``/``. Two more were - cost rather than correctness: a run of whitespace inside a failing tag - made the attribute pattern backtrack as ``(a+)*``, and a 39-byte body - took 20.8 seconds. - - Those seven were found as depth under-counts, back when that module - predicted the nesting depth instead of attempting the conversion. The - pattern's other two readers were wrong the same way: a derailed scan - meant :func:`_canonicalise_void_elements` never saw the ``<br />`` it - exists to rewrite, so - ``"<p>one<br>two</p><script>x</ script><p>three<br />TAIL</p>"`` lost - ``TAIL`` outright, and :func:`_strip_tags` inherited both the - misreadings and the backtracking. - - Reading the parser's own event stream cannot be wrong about the - parser, so none of those remain judgement calls. It is also not a new - dependency nor a new risk: ``markdownify`` builds its tree with - ``bs4``, and ``bs4`` builds it with this same ``html.parser``, so - every pathology the parser has was already in the pipeline. Measured - by glpi_python_client on the shapes that made the pattern backtrack, - the ``markdownify`` call costs what this scan costs, to within a few - per cent. - - ``convert_charrefs`` is ``False`` because that is what ``bs4`` passes - (in ``bs4.builder._htmlparser``), and the difference shows: with it - on, ``html.unescape`` consumes the longest *known* name, so - ``©right=2`` in a pasted URL loses its ``©``. Off, each - reference arrives as its own event and :func:`_resolve_references` - applies the whole-name rule the converting path applies. - - ``collect_text`` separates the two callers: the void rewrite needs - only the spans, and accumulating the pieces of a 150 KB body for it - would be waste. - - Parameters - ---------- - content : str - The document to walk. Kept so spans can be sliced back out of it: - the parser reports what it found and where, and the source is the - only place the exact original spelling still exists. - collect_text : bool, optional - Whether to accumulate :attr:`pieces` for :func:`_strip_tags`. - """ - - def __init__(self, content: str, *, collect_text: bool = False) -> None: - super().__init__(convert_charrefs=False) - self._content = content - self._collect_text = collect_text - offsets = [0] - for line in content.splitlines(keepends=True): - offsets.append(offsets[-1] + len(line)) - self._line_offsets = offsets - #: Source spans of ``<void ... />`` tags, for - #: :func:`_canonicalise_void_elements`. - self.void_spans: list[tuple[int, int]] = [] - #: The document's text, in order, when ``collect_text`` is set. - self.pieces: list[str] = [] - #: Set when ``html.parser`` gave up on the document. - self.rejected = False - - def _at(self) -> int: - """Return the absolute offset of the construct being handled. - - ``goahead`` calls ``updatepos`` up to the start of each construct - before dispatching it, so ``getpos`` addresses the construct - itself. It reports a line and a column, and every span sliced - here needs an index, which is what the line table built in - ``__init__`` converts between. - """ - - lineno, offset = self.getpos() - return self._line_offsets[lineno - 1] + offset - - def _text(self, piece: str) -> None: - if self._collect_text: - self.pieces.append(piece) - - def _boundary(self, tag: str) -> None: - """Record a block element's edge as a line break. - - See :data:`_BLOCK_ELEMENTS` for why the block/inline line is the - one that matters here. - """ - - if self._collect_text and tag in _BLOCK_ELEMENTS: - self.pieces.append("\n") - - def handle_starttag(self, tag: str, attrs: object) -> None: - self._boundary(tag) - - def handle_startendtag(self, tag: str, attrs: object) -> None: - """Count a ``<foo/>`` as a leaf, and note a void one to rewrite. - - This is the event :func:`_canonicalise_void_elements` needs, and - the one a pattern cannot identify reliably. ``html.parser`` - reaches it only when the stripped remainder of the tag is exactly - ``/>``, so ``<br />`` arrives here while ``<br / >`` is an - ordinary start tag. Deciding it on the event means the workaround - fires on exactly the tags that trigger the ``bs4`` defect, and on - no others. - """ - - if tag in _VOID_ELEMENTS: - token = self.get_starttag_text() - start = self._at() - if token is not None and self._content.startswith(token, start): - self.void_spans.append((start, start + len(token))) - self._boundary(tag) - - def handle_endtag(self, tag: str) -> None: - self._boundary(tag) - - def handle_data(self, data: str) -> None: - """Keep character data, a raw-text element's body included. - - A ``<script>`` or ``<style>`` body arrives here because the parser - is in CDATA mode. The converting path drops such a body -- a - browser displays none of it -- and this keeps it anyway, which is - the safe direction for a fallback whose promise is that it says no - less than the conversion: see :func:`_strip_tags`. - """ - - self._text(data) - - def handle_entityref(self, name: str) -> None: - self._text(self._reference_at()) +def _plain_text_html(text: str) -> str: + """Return the HTML that displays ``text`` literally, line by line.""" - def handle_charref(self, name: str) -> None: - self._text(self._reference_at()) + lines = (escape(line, quote=False) for line in text.splitlines()) + return "<p>" + "<br>".join(lines) + "</p>" - def _reference_at(self) -> str: - """Return a character reference exactly as it was written. - The event carries the name but not whether a semicolon closed it, - and that is what decides whether the reference resolves -- so the - source is re-read rather than the token rebuilt from the name. - See :data:`_CHARACTER_REFERENCE`. - """ +def _soup(html: str) -> BeautifulSoup: + """Parse ``html`` and drop what a browser does not display.""" - match = _CHARACTER_REFERENCE.match(self._content, self._at()) - return match.group(0) if match is not None else "" + soup = BeautifulSoup(html, "html.parser") + for hidden in soup.find_all(_HIDDEN): + hidden.extract() + return soup - def handle_pi(self, data: str) -> None: - """Keep a processing instruction's body, which the converter prints. - Measured, not assumed, and the delimiters are the detail that - matters: ``bs4`` files the instruction as a string node, so - ``<p>a</p><?php SECRET ?><p>b</p>`` converts to - ``"a\\n\\nphp SECRET ?\\n\\nb"`` -- the body, without its ``<?`` - and ``>``. Keeping the body alone therefore matches the - converting path exactly, where keeping the whole construct used - to leave punctuation in an output that promises none. - """ +#: The parts of a table other than its cells. +_TABLE_PARTS = frozenset( + {"table", "thead", "tbody", "tfoot", "tr", "caption", "colgroup", "col"} +) - self._text(data) +#: What Markdown writes on one line: a table inside one cannot be a table. +_ONE_LINE = frozenset({"td", "th", "a", "h1", "h2", "h3", "h4", "h5", "h6"}) + + +def _flatten_nested_tables(root: Tag) -> None: + """Write a table that sits inside a cell, a heading or a link as its cells' text. + + A GFM cell holds one line, so a nested table -- a layout e-mail + signatures often use -- written as a table split the outer row and lost + every word; in a heading or a link its pipes became text. Inside one, a + table's parts become ``span`` and its cells ``ev-cell``, which + :class:`_Converter` writes as spaced inline text. Renaming in one walk + keeps it linear however deeply tables nest. + """ + + cells = 0 + counted: list[bool] = [] + for node, entering in _walk(root): + if not isinstance(node, Tag): + continue + if not entering: + if counted.pop(): + cells -= 1 + continue + is_cell = node.name in ("td", "th") + if cells and (is_cell or node.name in _TABLE_PARTS): + node.name = "ev-cell" if is_cell else "span" + counted.append(False) + continue + holds = node.name in _ONE_LINE and (node.name != "a" or bool(node.get("href"))) + counted.append(holds) + cells += holds + + +def _walk(root: Tag) -> Iterator[tuple[PageElement, bool]]: + """Yield the nodes under ``root`` in document order, each tag in and out. + + A tag comes with ``True`` on the way in and ``False`` on the way out. + Comments and declarations are skipped. The walk keeps its own stack, so + it costs nothing in recursion however deep the document is. + """ + + stack = [(node, True) for node in reversed(root.contents)] + while stack: + node, entering = stack.pop() + if isinstance(node, PreformattedString): + continue + yield node, entering + if entering and isinstance(node, Tag): + stack.append((node, False)) + stack.extend((child, True) for child in reversed(node.contents)) + + +def _drop_trailing_breaks(root: Tag) -> None: + """Turn every ``<br>`` its block shows nothing after into a ``<wbr>``. + + A browser shows no line for such a break, and CommonMark shows a + backslash break that ends a block as a backslash. ``<wbr>`` displays + nothing either, and renaming costs nothing where removing a node from a + long run of siblings costs the length of the run. + """ + + pending: list[Tag] = [] + for node, entering in _walk(root): + if not isinstance(node, Tag): + if str(node).strip(): + pending.clear() + elif node.name in _BLOCKS: + for line_break in pending: + line_break.name = "wbr" + pending.clear() + elif entering and node.name == "br": + pending.append(node) + elif entering and node.name == "img": + pending.clear() + for line_break in pending: + line_break.name = "wbr" + + +def _note_sides(root: Tag) -> None: + """Note on each bold or italic element the characters displayed on either side. + + One pass: ``last`` is the last character shown so far on the current + line, and an element that has closed waits for the next one. A line edge + counts as a space, as it does for CommonMark, and so does the edge of a + flattened cell, which :class:`_Converter` spaces. + """ + + last = " " + waiting: list[Tag] = [] + for node, entering in _walk(root): + if isinstance(node, Tag): + if node.name in _EMPHASIS: + if entering: + node[_BEFORE] = last + else: + waiting.append(node) + continue + if node.name not in _BLOCKS and node.name not in ("br", "ev-cell"): + continue + first = last = " " + else: + text = str(node) + if not text: + continue + first, last = text[0], text[-1] + for element in waiting: + element[_AFTER] = first + waiting.clear() + for element in waiting: + element[_AFTER] = " " - def keep_remainder(self) -> None: - """Hand back the text of the region the parser stopped on. - Called only when it raised, so that the degraded path still - carries every word after the construct it could not read. - """ +def _punctuation(char: str) -> bool: + """Return whether CommonMark counts ``char`` as punctuation.""" - self._text(self._content[self._at() :]) + return char in string.punctuation or unicodedata.category(char).startswith("P") - def unknown_decl(self, data: str) -> None: - """Keep the body of a ``CDATA`` section or a marked section. - Counter-intuitive, and measured rather than assumed: ``bs4`` files - both as a string node, and ``markdownify`` prints a string node - that is neither a comment nor a doctype. So ``<![IGNORE[x]]>`` - contributes ``IGNORE[x`` to the converted output, and dropping it - here would make the same document say less on the degraded path - than on the converting one. - """ +def _flanks(inner: str, before: str, after: str) -> bool: + """Return whether CommonMark reads ``*`` runs around ``inner`` as emphasis. - self._text(data[6:] if data.startswith("CDATA[") else data) - - -def _scan(content: str, *, collect_text: bool = False) -> _ParserScan: - """Run one :class:`_ParserScan` over ``content`` and hand it back. - - The parser gives up on two constructs -- an unknown marked-section - keyword such as ``<![FOO[``, and a ``[`` where a declaration cannot - hold one -- by raising ``AssertionError`` from ``_markupbase``. That - is not a case to guess around, because ``bs4`` catches the same - ``AssertionError`` and re-raises it as ``ParserRejectedMarkup``: a - document that stops this scan is a document ``markdownify`` cannot - convert either. So the partial scan is kept and the remainder of the - source is handed to :attr:`_ParserScan.pieces` as text, which is what - lets :meth:`EasyvistaContentConverter.from_transport` answer that - rejection with the body's words instead of an exception. - - Parameters - ---------- - content : str - The document to walk. - collect_text : bool, optional - Whether the scan should accumulate the document's text. - - Returns - ------- - _ParserScan - The finished scan, whether or not the parser ran out of document. + ``inner`` neither starts nor ends with whitespace. A run opening onto + punctuation needs whitespace or punctuation in front of it, and a run + closing after punctuation needs the same behind it. """ - scan = _ParserScan(content, collect_text=collect_text) - try: - scan.feed(content) - scan.close() - except AssertionError: - scan.rejected = True - scan.keep_remainder() - return scan - - -def _strip_tags(content: str) -> str: - """Reduce HTML to its text without building a tree, keeping every word. - - The fallback for a document ``markdownify`` could not walk -- deeper - than the caller's remaining stack, or refused by the parser outright. - One pass of :class:`_ParserScan`, then whitespace tidying. No tree, - no recursion and no ceiling of its own, which is what qualifies it as - the fallback: it answers for input of any shape and any depth, so - there is always something to give the caller. - - It **degrades and never truncates.** The property, stated as - something checkable: after collapsing whitespace, every character the - converting path would have produced also appears here, in order. A - superset, not an equality -- so no body says less because of the path - it took, which is the only guarantee worth making about a fallback. - - Establishing that meant measuring what the converting path really - keeps, construct by construct, rather than assuming. Two answers were - counter-intuitive and each was a silent deletion in glpi_python_client - before it was checked: a ``CDATA`` body is kept, and so is the inside of - any ``<!``/``<?`` construct the parser could not resolve, which it hands - back as character data. A ``<script>``/``<style>``/``<title>`` body - goes the other way: the converting path drops it, as a browser does, - and this keeps it, which the superset promise allows. - - The text is plain, not Markdown: - :meth:`EasyvistaContentConverter.from_transport` spells it through - :func:`_literal_markdown` before handing it back, so it is escaped - exactly as the converting path escapes literal text. - - What it does **not** reproduce, none of which loses a character of - prose: - - * Markup that only the converter can express: a link becomes its text - without the target, an image contributes nothing, and a fenced - block loses its fence -- so ``<pre>`` indentation is normalised - away with the rest. A pasted log comes back as its own lines of - text, not as a code block. - * Whitespace is normalised harder. Runs of spaces collapse, and - `` `` counts as whitespace, so `` ``-padded column - alignment does not survive. - * Character references are resolved even inside a region the parser - handed back as raw data, so a broken comment's ``&`` comes back - as ``&``. In the other direction, a handful of semicolon-less - references stay literal here that the converter resolves -- see - :func:`_resolve_references`, which errs that way on purpose. - * Whitespace falls differently at a markup boundary, in both - directions: the converter joins ``a<b>c`` as ``a**c**`` where this - joins it as ``ac``, and this breaks a line at a block edge the - converter runs together. Which is why the property is about the - order of the characters of prose and not about where the spaces - land. - - Parameters - ---------- - content : str - Raw HTML. - - Returns - ------- - str - The document's text, block boundaries preserved as line breaks - and character references resolved. - """ + opens = not _punctuation(inner[0]) or before.isspace() or _punctuation(before) + closes = not _punctuation(inner[-1]) or after.isspace() or _punctuation(after) + return opens and closes - if ">" not in content: - # No ``>`` means no markup to skip, so the whole document is the - # text the parser would flush on ``close()``. Answering it here - # also keeps the scan away from the one shape that costs - # ``html.parser`` more than linear time: with no ``>`` to finish a - # tag, ``close()`` advances one character at a time and rescans - # the tail, and 32 KB of an unfinished tag takes 13 seconds - # (measured by glpi_python_client). - text = _resolve_references(content) - else: - text = _resolve_references("".join(_scan(content, collect_text=True).pieces)) - text = re.sub(r"[^\S\n]*\n[^\S\n]*", "\n", text) - text = re.sub(r"[^\S\n]{2,}", " ", text) - text = re.sub(r"\n{3,}", "\n\n", text) - return text.strip() - - -def _canonicalise_void_elements(content: str) -> str: - """Rewrite ``<br />`` as ``<br>``, so the text after it is not lost. - - A workaround for a ``beautifulsoup4`` defect, reachable from ordinary - editor output and silent when it fires. glpi_python_client measured it - on 4.14.3. **It no longer reproduces on 4.15.0** (measured 2026-09-30 on - CPython 3.10, 3.12, 3.13 and 3.14, every shape below); the workaround - stays because the ``content`` extra accepts ``beautifulsoup4`` 4.12 and - later, and on current releases it changes no output. - - ``bs4``'s ``html.parser`` builder auto-closes a bare ``<br>`` and - records the name in ``already_closed_empty_element``, a list keyed by - name alone, so a later ``</br>`` can be ignored as redundant. If no - ``</br>`` ever arrives the entry simply stays there. The next - ``<br />`` -- which reaches the builder as ``handle_startendtag`` -- - opens a real element and then closes it itself, and *that* close - finds the stale entry, treats the element as already closed, and - leaves it open. Every following sibling becomes a child of the - ``<br>``. - - ``get_text`` still walks those children, which is why the tree looks - intact, but ``markdownify``'s ``convert_br`` ignores an element's - children and returns a line break. The text is gone: - ``"<p>line1<br>line2</p><p>para2<br />line4</p>"`` converted to - ``"line1 \\nline2\\n\\npara2"`` on 4.14.3 -- and note the two - spellings are in different paragraphs, because a name once recorded - poisons the rest of the document. ``<img>`` and ``<hr>`` lose text the - same way; they are the other two converters that discard children. One - bare ``<br>`` anywhere before one ``<br />`` is the whole precondition, - and a memo is written by more than one client -- the web UI and every - API caller. - - Rewriting to the bare spelling removes the ``handle_startendtag`` - path for void elements, which is where the asymmetry lives; both - spellings already build the same node, so nothing else about the - output moves. Only the names in :data:`_VOID_ELEMENTS` are touched, - and only where the parser really reports ``handle_startendtag``: a - self-closed ``<div/>`` is left alone, and cannot be affected anyway, - since only a void name is ever recorded. - - Which tags those are is the parser's answer rather than this - module's, and that is not cosmetic. Deciding it by pattern meant - inheriting every way the pattern could be derailed, and a derailed - scan reinstates the very defect this works around: measured by - glpi_python_client, - ``"<p>one<br>two</p><script>x</ script><p>three<br />TAIL</p>"`` - lost ``TAIL`` outright, because ``</ script>`` ends raw text for the - parser but not for the pattern, so the ``<br />`` after it was never - seen and never rewritten. ``<img>`` and ``<hr>`` lost their tails the - same way. - - Parameters - ---------- - content : str - Raw HTML. - - Returns - ------- - str - The same HTML with self-closing void tags written bare. - """ - if "/>" not in content: - return content - spans = _scan(content).void_spans - if not spans: - return content - pieces: list[str] = [] - cursor = 0 - for start, end in spans: - pieces.append(content[cursor:start]) - pieces.append(_VOID_SELF_CLOSE.sub(">", content[start:end])) - cursor = end - pieces.append(content[cursor:]) - return "".join(pieces) - - -# --------------------------------------------------------------------------- -# Literal text -# --------------------------------------------------------------------------- - -#: The characters whose reading as Markdown depends on what surrounds them. -#: -#: Every one of them is literal in some positions and syntax in others -- -#: ``#`` is a heading only at the start of a line, ``_`` is emphasis only -#: when a partner closes it -- so a text node cannot decide them alone. Each -#: is carried as a private stand-in (:class:`_StandIns`) until the container -#: it lands in is assembled, and decided there. -_LITERAL = "\\<&*_`[]!#>-+.=|~" - -#: How a literal character is spelled when it would be misread, where that -#: is not a backslash in front of it. -#: -#: ``<`` and ``&`` have no backslash escape python-markdown honours -- ``\<`` -#: renders both characters and the ``<`` still opens a tag -- and neither do -#: ``=`` and ``~``. A character reference is the spelling it passes through -#: as the character. -_SPELLED_OUT = { - "\\": "\\\\", - "<": "<", - "&": "&", - "=": "=", - "~": "~", -} +def _destination(url: str) -> str: + """Spell ``url`` as a link destination that reads back as ``url``.""" -#: The characters python-markdown removes a backslash from, with these -#: extensions. Asked of python-markdown rather than written down, because a -#: backslash in front of anything else is displayed. -_ESCAPABLE = frozenset(Markdown(extensions=_MARKDOWN_EXTENSIONS).ESCAPED_CHARS) - -# python-markdown's own patterns, compiled the way its inline processors -# compile theirs (``re.DOTALL``: a code span or an e-mail autolink can run -# across a line break). Taken from the renderer rather than copied, so the -# reader cannot disagree with the version doing the rendering. -_CODE_SPAN = re.compile(BACKTICK_RE, re.DOTALL) -_STANDALONE = re.compile(NOT_STRONG_RE, re.DOTALL) -_AUTOMAIL = re.compile(AUTOMAIL_RE, re.DOTALL) -_AUTOLINK = re.compile(AUTOLINK_RE) -_RULE = re.compile(HRProcessor.RE) -_REFERENCE_DEFINITION = ReferenceProcessor.RE - -#: What follows an ``&`` that the renderer displays as a character. -#: -#: A named reference needs its semicolon. A numeric one does not: -#: python-markdown runs ``html.parser`` over its whole source to find raw -#: HTML, and that re-emits ``ᆩ`` as ``ᆩ`` -- measured, 3.10.3. -_REFERENCE_TAIL = re.compile(r"\#[0-9]|\#[xX][0-9a-fA-F]|[0-9A-Za-z]+;") - -#: What makes a ``<`` open a tag, a comment, a declaration or an instruction. -_TAG_START = re.compile(r"[A-Za-z/!?]") - -_LINE_BREAK = re.compile(" \n") -_WORD = re.compile(r"\w") -_SETEXT_UNDERLINE = re.compile(r"(?:=+|-+) *") -_ORDERED_MARKER = re.compile(r"\d+(?=\.[ ])") -_LEADING_BLANK_LINES = re.compile(r"\A(?:[ \t]*\n)+") -_WHITESPACE = re.compile(r"\s+") -_EDGE = " \t\r\n" - -#: A character no analysis below treats as syntax: not a word character, -#: not whitespace, not punctuation python-markdown reads. Stands in for -#: anything already decided -- an escaped character, a code span. -_NEUTRAL = "\x02" - -#: Elements ``markdownify`` leaves inline that a browser shows as blocks. -_EXTRA_BLOCKS = frozenset({"center", "dir", "menu"}) - - -class _StandIns: - """Private characters carrying literal text until its container spells it. - - ``markdownify`` calls ``escape`` on one text node at a time, and a text - node cannot see what it will end up beside: whether its ``#`` starts a - line, whether its ``*`` has a partner in the next node, whether its - ``<`` is followed by a letter from a ``<span>``. So escaping replaces - each character in :data:`_LITERAL` with a stand-in, and the converter - for the enclosing block -- a paragraph, a list item, a cell, the - document -- decides every stand-in in it at once, with the whole - assembled Markdown of that block in view (:class:`_Spelling`). - - The stand-ins are a run of consecutive supplementary private-use code - points the source does not already contain, so a body holding - private-use characters of its own -- icon fonts map symbols there -- - cannot be confused with them. One more code point marks a hard line - break, whose surrounding spaces are settled the same way, and one more - starts a line no list item indents (:attr:`lazy`). - - Parameters - ---------- - source : str - The text about to be converted, which the stand-ins must not occur - in. - """ + url = re.sub(r"[\t\n\r]", "", url) + return "<" + re.sub(r"[\\<>]", r"\\\g<0>", url) + ">" - def __init__(self, source: str) -> None: - size = len(_LITERAL) + 2 - for base in range(0xF0000, 0xFFFFE - size, size): - characters = [chr(base + offset) for offset in range(size)] - if not any(character in source for character in characters): - break - else: # pragma: no cover - needs a source using all of plane 15 - raise EasyvistaContentError( - "Could not convert EasyVista HTML content to Markdown: it uses " - "every private-use code point the converter could work with." - ) - literal = characters[: len(_LITERAL)] - self.hard_break = characters[-2] - #: Starts a line that stays at the start of the line: a quote that - #: opens a list item is read by python-markdown only if its later - #: lines are its lazy continuation, since the item's first block is - #: never detabbed and ``>`` counts only three spaces in at most. - self.lazy = characters[-1] - self.characters = frozenset(literal) - self._shadow = str.maketrans(dict(zip(_LITERAL, literal, strict=True))) - self._restore = str.maketrans(dict(zip(literal, _LITERAL, strict=True))) - self._any = re.compile("[" + "".join(literal) + "]") - self._break = re.compile("[ ]*" + self.hard_break + "\n?[ ]*") - def shadow(self, text: str) -> str: - """Return ``text`` with every character in :data:`_LITERAL` stood in for.""" +def _title(title: str) -> str: + """Spell a link or image title, with the space before it.""" - return text.translate(self._shadow) + title = " ".join(title.split()) + return ' "' + re.sub(r'[\\"]', r"\\\g<0>", title) + '"' if title else "" - def restore(self, text: str) -> str: - """Return ``text`` with every stand-in back as the character it is.""" - return text.translate(self._restore) +class _Converter(MarkdownConverter): + """``markdownify``, with what CommonMark and the reader's glue need on top. - def carried(self, text: str) -> bool: - """Return whether ``text`` still holds an undecided stand-in.""" + The stub ``markdownify`` ships declares the constructor and ``convert`` + alone, so its converters are reached through :meth:`_inherited`. + """ - return self._any.search(text) is not None + def _inherited(self, name: str) -> Any: + return getattr(super(), name) - def positions(self, text: str) -> set[int]: - """Return where ``text`` holds a stand-in.""" + def escape(self, text: str, parent_tags: set[str]) -> str: + """Escape a text's last ``!`` as well: before a link it makes an image. - return {found.start() for found in self._any.finditer(text)} + A ``!`` meets a ``[`` only there, markdownify escaping a ``[`` in text; + mdformat keeps that escape before a link and drops it anywhere else. + """ - def replace(self, text: str, spell: Callable[[int, str], str]) -> str: - """Replace each stand-in with ``spell(position, character)``.""" + text = str(self._inherited("escape")(text, parent_tags)) + return text[:-1] + "\\!" if text.endswith("!") else text - return self._any.sub( - lambda found: spell(found.start(), self.restore(found.group(0))), text - ) + def process_tag(self, node: Any, parent_tags: Any = None) -> str: + if node.name == "ev-cell": # one line, as a cell is: its blocks are inline + parent_tags = {*(parent_tags or ()), "_inline"} + return str(self._inherited("process_tag")(node, parent_tags)) - def settle_breaks(self, text: str) -> str: - """Spell every hard line break ``" \\n"``, dropping the spaces around it. + def _markup( + self, el: Tag, text: str, parent_tags: set[str], markers: str, tag: str + ) -> str: + """Wrap ``text`` in ``markers``, or in raw ``<tag>`` where they would not close. - A browser does not display whitespace at either side of a ``<br>``, - so it is not content, and keeping it made the same body read two - ways: ``a<br> b`` read as ``a \\n b`` the first time and as - ``a \\nb`` once python-markdown had rendered it. + Line breaks and spaces at the edges move outside, where they cannot + stop the markers closing. """ - if self.hard_break not in text: + edges = _EDGES.fullmatch(text) + if "_noformat" in parent_tags or edges is None or not edges[2]: return text - return self._break.sub(" \n", text) - - -def _spelled(character: str) -> str: - """Return the escaped spelling of one literal character.""" - - return _SPELLED_OUT.get(character) or "\\" + character - - -class _Spelling: - """One container's Markdown, and which of its literal characters to escape. - - Built over the container's assembled text, where a stand-in is literal - text and anything else is markup ``markdownify`` generated or text an - inner container already decided. Each ``settle_*`` method asks one of - python-markdown's questions of the text as it will be rendered, and - marks the literal characters that would be read as syntax; - :meth:`render` spells them. - - The order the methods run in is python-markdown's own inline order -- - code spans before escapes before links before emphasis -- because each - of those consumes text the next one would otherwise see. - - Parameters - ---------- - text : str - The container's text, stand-ins included. - stand_ins : _StandIns - The stand-ins in use. - """ - - def __init__(self, text: str, stand_ins: _StandIns) -> None: - self.text = text - self.stand_ins = stand_ins - self.view = stand_ins.restore(text) - self.literal = stand_ins.positions(text) - self.escaped: set[int] = set() - self.hidden: set[int] = set() + before, inner, after = edges.groups() + prev = before[-1:] or str(el.get(_BEFORE) or " ") + succ = after[:1] or str(el.get(_AFTER) or " ") + if markers and _flanks(inner, prev, succ): + return f"{before}{markers}{inner}{markers}{after}" + return f"{before}<{tag}>{inner}</{tag}>{after}" - # -- bookkeeping --------------------------------------------------------- - - def escape(self, position: int) -> None: - """Escape the character at ``position``, if it is literal.""" - - if position in self.literal: - self.escaped.add(position) - - def live(self, position: int) -> bool: - """Return whether ``position`` is literal and not yet escaped.""" - - return position in self.literal and position not in self.escaped - - def render(self) -> str: - """Return the container's Markdown, every literal character spelled.""" + def convert_b(self, el: Tag, text: str, parent_tags: set[str]) -> str: + return self._markup(el, text, parent_tags, "**", "strong") - escaped = self.escaped + convert_strong = convert_b - def spell(position: int, character: str) -> str: - return _spelled(character) if position in escaped else character + def convert_em(self, el: Tag, text: str, parent_tags: set[str]) -> str: + return self._markup(el, text, parent_tags, "*", "em") - return self.stand_ins.replace(self.text, spell) + convert_i = convert_em - def markdown(self, start: int, end: int) -> tuple[str, list[int]]: - """Return one span's Markdown as it stands, and each character's position.""" + def convert_s(self, el: Tag, text: str, parent_tags: set[str]) -> str: + """Struck text: CommonMark has no strikethrough, so always raw ``<s>``.""" - pieces: list[str] = [] - origin: list[int] = [] - cursor = start - for position in sorted(p for p in self.escaped if start <= p < end): - pieces.append(self.view[cursor:position]) - origin.extend(range(cursor, position)) - form = _spelled(self.view[position]) - pieces.append(form) - origin.extend([position] * len(form)) - cursor = position + 1 - pieces.append(self.view[cursor:end]) - origin.extend(range(cursor, end)) - return "".join(pieces), origin + return self._markup(el, text, parent_tags, "", "s") - def working(self, *, line_breaks: bool = False) -> str: - """Return the view with everything already decided made neutral. + convert_del = convert_s + convert_strike = convert_s - With ``line_breaks``, a hard break is neutral too: python-markdown - replaces ``" \\n"`` with a placeholder before it looks for emphasis, - so the characters either side of one are not next to whitespace. - """ + def convert_u(self, el: Tag, text: str, parent_tags: set[str]) -> str: + """Underlined or highlighted text: CommonMark has neither, so raw HTML.""" - chars = list(self.view) - for position in self.hidden | self.escaped: - chars[position] = _NEUTRAL - working = "".join(chars) - if line_breaks: - working = _LINE_BREAK.sub(_NEUTRAL * 3, working) - return working + return self._markup(el, text, parent_tags, "", el.name) - # -- what depends only on the next few characters ---------------------- + convert_mark = convert_ins = convert_u - def settle_local(self) -> None: - """Decide ``\\``, ``<`` and ``&`` by the characters right after them. + def convert_center(self, el: Tag, text: str, parent_tags: set[str]) -> str: + """``<center>``, obsolete, is a block, as ``<div>`` is.""" - * ``\\`` before a character python-markdown would treat as escaped - (:data:`_ESCAPABLE`) -- ``\\\\serveur``, ``C:\\_temp``; - * ``<`` before a letter, ``/``, ``!`` or ``?`` -- anything - python-markdown would pass through as raw HTML; - * ``&`` opening a character reference python-markdown would pass - through for the browser to decode (:data:`_REFERENCE_TAIL`). + if "pre" in parent_tags: + return text + return str(self._inherited("convert_div")(el, text, parent_tags)) - A ``<`` opening an e-mail autolink is decided once the inline - constructs it could run through are known (:meth:`settle_automail`). - """ + def convert_ev_cell(self, el: Tag, text: str, parent_tags: set[str]) -> str: + """A cell of a table nested in a cell: its text, set apart by spaces.""" - view = self.view - for position in self.literal: - char = view[position] - following = view[position + 1 : position + 2] - if char == "\\": - if following and following in _ESCAPABLE: - self.escape(position) - elif char == "<": - if following and _TAG_START.match(following): - self.escape(position) - elif char == "&" and _REFERENCE_TAIL.match(view, position + 1): - self.escape(position) - - def settle_automail(self) -> None: - """Escape every literal ``<`` that would open an e-mail autolink. - - python-markdown looks for ``<address@domain>`` only once code spans, - escapes, links, images and URL autolinks have each become a - placeholder, so an address can run through any of them: - ``<3![a b](x.png)@c>`` is one to it. The pattern is matched against - the text with each of those collapsed to one character, and right to - left, because a match may also run through a ``<`` decided to its - right: once that one is ``<``, nothing stops it. - - A link's text is read the same way on its own, since python-markdown - parses it after the link is found; an image's text is never parsed. - """ + return f" {text.strip()} " - view = self.view - opening = sorted( - (p for p in self.literal if view[p] == "<" and p not in self.escaped), - reverse=True, + def convert_br(self, el: Tag, text: str, parent_tags: set[str]) -> str: + if "_noformat" in parent_tags: + return "\n" # in <pre>, or in inline code, which convert_code splits + if "td" in parent_tags or "th" in parent_tags: + return "<br>" # a cell is one line of Markdown; cmark-gfm passes the tag + return str(self._inherited("convert_br")(el, text, parent_tags)) + + def convert_code(self, el: Tag, text: str, parent_tags: set[str]) -> str: + """Inline code, a span per line: a code span cannot hold a line break.""" + + cell = "td" in parent_tags or "th" in parent_tags + lines = [text] if "_noformat" in parent_tags else text.split("\n") + join = "<br>" if cell else " " if "_inline" in parent_tags else "\\\n" + code = join.join( + str(self._inherited("convert_code")(el, line, parent_tags)) + for line in lines ) - if not opening or "@" not in view: - return - working = list(self.working()) - contexts: dict[tuple[int, int], _Collapsed] = {} - for position in opening: - start, end = 0, len(view) - while True: - if (start, end) not in contexts: - contexts[start, end] = self._collapsed(working, start, end) - context = contexts[start, end] - link = context.link_at(position) - if link is None: - at = context.index[position - start] - if _automail_at(context.text, context.tail, at): - self.escape(position) - context.tail[at] = "&" - break - if link.label is None or not link.label[0] <= position < link.label[1]: - break - start, end = link.label - - def _collapsed(self, working: list[str], start: int, end: int) -> _Collapsed: - """Return ``view[start:end]`` with every placeholder collapsed.""" - - view = self.view - links = _generated_links(view, working, self.literal, start, end) - masked = self.hidden | self.escaped - for link in links: - masked = masked | set(range(*link.span)) - pieces: list[str] = [] - index: list[int] = [] - for position in range(start, end): - if position not in masked: - pieces.append(view[position]) - elif position == start or position - 1 not in masked: - pieces.append(_NEUTRAL) - index.append(len(pieces) - 1) - return _Collapsed("".join(pieces), index, links) - - # -- python-markdown's inline patterns, in its own order --------------- - - def settle_bang(self) -> None: - """Escape a literal ``!`` right before a generated link's ``[``. - - Otherwise ``Attention!`` followed by a link renders an image. - """ - - view = self.view - for position in self.literal: - if ( - view[position] == "!" - and view[position + 1 : position + 2] == "[" - and position + 1 not in self.literal - ): - self.escape(position) - - def settle_backticks(self, units: Iterable[tuple[int, int]]) -> None: - """Escape every literal backtick run python-markdown would pair. - - Found by running python-markdown's own code-span pattern over each - unit's Markdown as it stands, escaping the literal runs it paired, - and running it again until it pairs none: escaping an opener can hand - its closer to the next run. A literal backtick touching a generated - one is escaped first, since it would join the generated run and - change the length the closer is matched on. What is left paired is - the generated code spans, which the later passes must not look - inside. - """ - - view = self.view - if "`" not in view: - return - for position in self.literal: - if view[position] == "`": - for neighbour in (position - 1, position + 1): - if ( - 0 <= neighbour < len(view) - and view[neighbour] == "`" - and neighbour not in self.literal - ): - self.escape(position) - for start, end in units: - if "`" in view[start:end]: - self._settle_code_spans(start, end) - - def _settle_code_spans(self, start: int, end: int) -> None: - while True: - markdown, origin = self.markdown(start, end) - offending: set[int] = set() - generated: list[tuple[int, int]] = [] - for opener, closer in _code_spans(markdown): - delimiters = [origin[i] for i in range(*opener)] - delimiters += [origin[i] for i in range(*closer)] - live = [position for position in delimiters if self.live(position)] - if live: - offending.update(live) - else: - generated.append((origin[opener[0]], origin[closer[1] - 1] + 1)) - if not offending: - for first, last in generated: - self.hidden.update(range(first, last)) - return - for position in offending: - self.escape(position) - - def hide_escapes(self) -> None: - """Hide what a backslash already escaped in generated or decided text. - - Pairs are consumed left to right, as python-markdown's escape - pattern consumes them, so ``\\\\`` hides both backslashes and escapes - nothing after it. - """ + if cell: + code = code.replace( + "|", "\\|" + ) # GFM splits a row on "|" before it reads code + return code - view = self.view - position = 0 - while position < len(view): - if ( - view[position] == "\\" - and position not in self.hidden - and position not in self.literal - ): - following = position + 1 - if ( - following < len(view) - and view[following] in _ESCAPABLE - and following not in self.literal - ): - self.hidden.update((position, following)) - position += 2 - continue - position += 1 - - def settle_brackets(self, units: Iterable[tuple[int, int]]) -> None: - """Escape every literal ``[`` that would open a link or an image. - - A ``[`` opens one when its balanced ``]`` is followed straight away - by ``(``, which is python-markdown's test. Right to left, because - escaping an inner ``[`` rebalances the brackets of an outer one; one - pass is enough, since the brackets to a ``[``'s right are settled - before it is. - - The pass keeps the ``]`` not yet closed by a ``[`` to their right on - a stack, so each ``[`` finds its partner on top of it: the first - ``]`` python-markdown's own count would stop at. An escaped ``[`` - hands its ``]`` back, to be closed by one further left. That keeps - the pass linear, where counting forward from each ``[`` cost a body - of unclosed brackets quadratic time. - """ + convert_kbd = convert_samp = convert_code - working = list(self.working()) - for start, end in units: - closers: list[int] = [] - for position in range(end - 1, start - 1, -1): - char = working[position] - if char == "]": - closers.append(position) - elif char == "[" and closers: - close = closers.pop() - if ( - self.live(position) - and close + 1 < end - and working[close + 1] == "(" - ): - self.escape(position) - working[position] = _NEUTRAL - closers.append(close) - - def settle_asterisks(self, units: Iterable[tuple[int, int]]) -> None: - """Escape literal ``*`` wherever a unit holds two possible delimiters. - - python-markdown pairs asterisks with no regard for word boundaries, - so any two in one block -- literal or generated -- can pair. The one - exemption is its own: a run of one to three that is literal through - and through and stands alone between whitespace is text to it - (``5 * 3``). A run that touches a generated delimiter is not alone. - """ + def convert_list(self, el: Tag, text: str, parent_tags: set[str]) -> str: + """A list; a nested one ends in a blank line. - working = self.working(line_breaks=True) - for start, end in units: - unit = working[start:end] - exempt: set[int] = set() - for found in _STANDALONE.finditer(unit): - run = range(start + found.start(3), start + found.end(3)) - if found.group(3).startswith("*") and all( - position in self.literal for position in run - ): - exempt.update(run) - delimiters = [ - start + offset - for offset, char in enumerate(unit) - if char == "*" and start + offset not in exempt - ] - if len(delimiters) >= 2: - for position in delimiters: - self.escape(position) - - def settle_underscores(self, units: Iterable[tuple[int, int]]) -> None: - """Escape literal ``_`` runs that python-markdown could pair. - - Underscores are only emphasis at a word boundary, which is what - keeps ``fichier_de_test_v2.xlsx`` literal, except in a run of three - or more, which may open emphasis anywhere and take a mid-word run as - its closer; a single run of seven pairs with itself. When a unit - holds a pairing, every run that could take part in one is escaped. + Otherwise the item's text after it joins the nested list's last item. """ - working = self.working(line_breaks=True) - for start, end in units: - unit = working[start:end] - exempt: set[int] = set() - for found in _STANDALONE.finditer(unit): - if found.group(3).startswith("_"): - exempt.update(range(start + found.start(3), start + found.end(3))) - runs = [ - (start + found.start(), start + found.end()) - for found in re.finditer("_+", unit) - if start + found.start() not in exempt - and any( - self.live(position) - for position in range(start + found.start(), start + found.end()) - ) - ] - if not _underscores_pair(runs, working, start, end): - continue - triple = any(finish - begin >= 3 for begin, finish in runs) - for begin, finish in runs: - left = begin == start or not _WORD.match(working[begin - 1]) - right = finish == end or not _WORD.match(working[finish]) - if left or right or triple: - for position in range(begin, finish): - self.escape(position) - - def settle_inline(self, units: list[tuple[int, int]]) -> None: - """Run every inline question, in python-markdown's order.""" - - self.settle_bang() - self.settle_backticks(units) - self.hide_escapes() - self.settle_brackets(units) - self.settle_automail() - self.settle_asterisks(units) - self.settle_underscores(units) - - -#: The longest ``<...>`` checked for an e-mail autolink. An address is at -#: most 254 characters; the pattern has no bound of its own, and checking -#: every ``<`` to the end of a long body would cost that body quadratically. -_AUTOMAIL_SPAN = 320 - - -def _automail_at(view: str, tail: list[str], position: int) -> bool: - """Return whether an e-mail autolink opens at ``position``. - - ``tail`` is the view with every ``<`` already decided to the right - spelled as a character that is not one. The pattern cannot cross a - space, and ends at the first ``>``. - """ - - end = view.find(">", position, position + _AUTOMAIL_SPAN) - if end < 0 or " " in view[position:end]: - return False - return _AUTOMAIL.match("".join(tail[position : end + 1])) is not None - - -def _underscores_pair( - runs: list[tuple[int, int]], working: str, start: int, end: int -) -> bool: - """Return whether python-markdown could pair any of ``runs``. - - ``EM_STRONG2`` and ``STRONG_EM2`` take their inner text from anywhere, - the run itself included, so a run of seven is a match on its own and a - run of three or more pairs with any other run. The ``SMART`` patterns - need an opener at a left word boundary and a later closer at a right - one. - """ - - if any(finish - begin >= 7 for begin, finish in runs): - return True - if len(runs) < 2: - return False - if any(finish - begin >= 3 for begin, finish in runs): - return True - opened = False - for begin, finish in runs: - if opened and (finish == end or not _WORD.match(working[finish])): - return True - if begin == start or not _WORD.match(working[begin - 1]): - opened = True - return False - - -def _code_spans(markdown: str) -> list[tuple[tuple[int, int], tuple[int, int]]]: - """Return the code spans python-markdown finds, as delimiter spans. - - Replays its backtick processor, which replaces each match with a - placeholder and searches on from just after it. After a code span that - is the same as searching on from the span's end: the one look-behind - there is ``(?<!\\\\)``, and the span ends in a backtick. A run of escaped - backslashes before a backtick is consumed on its own, though, and its - placeholder is what leaves the backtick after it free to open -- so - that run, and only that, is neutralised in the text searched. - """ + markdown = str(self._inherited("convert_list")(el, text, parent_tags)) + return markdown + "\n\n" if "li" in parent_tags else markdown - spans: list[tuple[tuple[int, int], tuple[int, int]]] = [] - text = markdown - position = 0 - while True: - found = _CODE_SPAN.search(text, position) - if found is None: - return spans - if found.group(3): - size = len(found.group(2)) - spans.append( - ( - (found.start(2), found.end(2)), - (found.end(3), found.end(3) + size), - ) - ) - position = found.end() - else: - start, end = found.span(1) - text = text[:start] + _NEUTRAL * (end - start) + text[end:] - position = end + convert_ul = convert_list + convert_ol = convert_list + def convert_li(self, el: Tag, text: str, parent_tags: set[str]) -> str: + """An item; an ordered one numbered one past the item before it. -class _Link(NamedTuple): - """A link, an image or a URL autolink the reader wrote, as view spans.""" + markdownify counts every item before each one, quadratic in a long + list, so each item keeps its number for the next. ``isdecimal``, + where markdownify's ``isnumeric`` let ``int("²")`` raise. + """ - span: tuple[int, int] - #: The link's text, which python-markdown parses on its own; ``None`` - #: for an image or an autolink, whose text it never parses. - label: tuple[int, int] | None + if el.parent is None or el.parent.name != "ol": + return str(self._inherited("convert_li")(el, text, parent_tags)) + before = el.find_previous_sibling("li") + start = str(el.parent.get("start") or "") + first = int(start) if start.isdecimal() else 1 + number = int(str(before[_NUMBER])) + 1 if isinstance(before, Tag) else first + el[_NUMBER] = str(number) + if not text.strip(): + return "\n" + bullet = f"{number}. " + body = re.sub("^(?=.)", " " * len(bullet), text.strip(), flags=re.M) + return bullet + body[len(bullet) :] + "\n" + def convert_pre(self, el: Tag, text: str, parent_tags: set[str]) -> str: + """A fenced block, its fence longer than any backtick run in the code.""" -class _Collapsed: - """A span of a view with every construct python-markdown stashes collapsed. + markdown = str(self._inherited("convert_pre")(el, text, parent_tags)) + longest = max((len(run) for run in _BACKTICK_RUN.findall(text)), default=0) + if longest < 3 or markdown.count("```") < 2: + return markdown + fence = "`" * (longest + 1) + head, _, rest = markdown.partition("```") + body, _, tail = rest.rpartition("```") + return f"{head}{fence}{body}{fence}{tail}" - Parameters - ---------- - text : str - The span, each stashed construct one character. - index : list of int - Where each position of the span went in ``text``. - links : list of _Link - The links, images and autolinks collapsed. - """ + def convert_a(self, el: Tag, text: str, parent_tags: set[str]) -> str: + """A link; ``<url>`` when its text is its URL and reads back unchanged.""" - def __init__(self, text: str, index: list[int], links: list[_Link]) -> None: - self.text = text - self.index = index - self.links = links - #: ``text`` with every ``<`` decided so far spelled as a character - #: that opens nothing. - self.tail = list(text) + href = str(el.get("href") or "") + edges = _EDGES.fullmatch(text) + if "_noformat" in parent_tags or not href or edges is None or not edges[2]: + return text + before, inner, after = edges.groups() + title = str(el.get("title") or "") + if not title and el.get_text() == href and _AUTOLINK.fullmatch(href): + return f"{before}<{href}>{after}" + return f"{before}[{inner}]({_destination(href)}{_title(title)}){after}" - def link_at(self, position: int) -> _Link | None: - """Return the link or image ``position`` is inside, if any.""" + def convert_img(self, el: Tag, text: str, parent_tags: set[str]) -> str: + """An image, wherever it is: a cell or a heading holds one too.""" - for link in self.links: - if link.span[0] <= position < link.span[1]: - return link - return None + alt = " ".join(str(el.get("alt") or "").split()) + alt = str(self._inherited("escape")(alt, parent_tags)) + if alt.startswith("^") and "_noformat" not in parent_tags: + alt = "\\" + alt # cmark-gfm reads "![^" as "!" and a link + src = _destination(str(el.get("src") or "")) + return f"![{alt}]({src}{_title(str(el.get('title') or ''))})" -def _generated_links( - view: str, working: list[str], literal: set[int], start: int, end: int -) -> list[_Link]: - """Return the links, images and URL autolinks the reader wrote in a span. +class _Node(RenderTreeNode): + """mdformat's syntax-tree node, answering in constant time what it asks often. - A generated ``[`` is one that is not literal, and it opens a link when - its balanced ``]`` is followed by ``(``; the link ends at the ``)`` - that balances that one. A generated ``<`` opens an autolink when the - text up to the next ``>`` is one. + mdformat asks every text node for its next sibling, every list for its + previous one, and every list item whether its list is tight. The base + class answers each by scanning, which made a long paragraph or a long + list quadratic. """ - links: list[_Link] = [] - position = start - while position < end: - char = working[position] - if char == "[" and position not in literal: - close = _balanced_close(working, position + 1, end) - if close is not None and view[close + 1 : close + 2] == "(": - finish = _balanced_paren(view, close + 2, end) - if finish is not None: - image = position > start and view[position - 1] == "!" - links.append( - _Link( - (position - image, finish + 1), - None if image else (position + 1, close), - ) - ) - position = finish + 1 - continue - elif char == "<" and position not in literal: - finish = view.find(">", position + 1, end) - if finish > 0 and _AUTOLINK.fullmatch(view, position, finish + 1): - links.append(_Link((position, finish + 1), None)) - position = finish + 1 - continue - position += 1 - return links - - -def _balanced_paren(view: str, start: int, end: int) -> int | None: - """Return where the ``(`` before ``start`` closes, before ``end``.""" - - depth = 1 - for position in range(start, end): - if view[position] == "(": - depth += 1 - elif view[position] == ")": - depth -= 1 - if depth == 0: - return position - return None - - -def _balanced_close(working: list[str], start: int, end: int) -> int | None: - """Return where the ``[`` before ``start`` closes, as python-markdown counts.""" - - depth = 1 - for position in range(start, end): - char = working[position] - if char == "]": - depth -= 1 - if depth == 0: - return position - elif char == "[": - depth += 1 - return None - + _index = -1 -def _blocks(view: str) -> list[list[tuple[int, int]]]: - """Return the text's blocks -- runs of non-blank lines -- as line spans.""" - - blocks: list[list[tuple[int, int]]] = [] - current: list[tuple[int, int]] = [] - start = 0 - for line in view.split("\n"): - end = start + len(line) - if line.strip(" \t"): - current.append((start, end)) - elif current: - blocks.append(current) - current = [] - start = end + 1 - if current: - blocks.append(current) - return blocks - - -def _units( - view: str, blocks: list[list[tuple[int, int]]], *, soft_breaks: bool -) -> list[tuple[int, int]]: - """Return the spans python-markdown parses inline text in, one at a time. - - A paragraph is one: its lines are joined by hard breaks, ``" \\n"``. - Every other line break the reader writes starts something python-markdown - parses on its own -- a nested list, a quote -- even with no blank line - before it, so a unit ends there. Pairing is only possible inside a unit, - so this is what keeps a list item's ``*`` from being escaped for a ``*`` - in the list nested under it. Stripped text is the exception: its plain - line breaks continue a paragraph, which ``soft_breaks`` says. - """ + def _position(self) -> int: + if self._index < 0: + for index, sibling in enumerate(self.siblings): + sibling._index = index + return self._index - units: list[tuple[int, int]] = [] - for block in blocks: - start, end = block[0] - for line_start, line_end in block[1:]: - if soft_breaks or view[end - 2 : end] == " ": - end = line_end - continue - units.append((start, end)) - start, end = line_start, line_end - units.append((start, end)) - return units - - -def _settle_block( - text: str, - stand_ins: _StandIns, - *, - in_list: bool, - top_level: bool, - lead: tuple[str, ...] = (), - soft_breaks: bool = False, -) -> str: - """Spell the literal text of one block container. - - ``text`` is the container's content before its own prefix is added -- - a list item's bullet, a quote's ``>`` -- so every line starts where - python-markdown will read it from. Each line is checked for the block - syntax python-markdown would find there, then the whole container for - the inline syntax (:meth:`_Spelling.settle_inline`). The line rules: - - * ``#`` at the very start of a line: an ATX heading, on any line; - * ``>`` after up to three spaces: a block quote, on any line; - * ``-``, ``+`` or ``*`` followed by a space, and digits followed by - ``.`` and a space, after up to three spaces: a list item -- on a - block's first line, or on any line inside a list item, where a - continuation line starts a nested list; - * a line of ``=`` or ``-`` alone as a block's second line: a setext - underline, spelled ``=`` or ``\\-``; - * three or more ``-``, ``*`` or ``_``, spaces allowed between: a rule, - on any line -- counting the bullet a list item is about to get; - * ``[label]: destination`` on a line: a reference definition, which - python-markdown consumes and which would turn every ``[label]`` in - the document into a link; - * three backticks or tildes at the start of a top-level line: a fence; - * a line of only ``|``, ``:``, ``-`` and spaces: a table's separator. - - Parameters - ---------- - text : str - The container's text. - stand_ins : _StandIns - The stand-ins in use. - in_list : bool - Whether the container is inside a list item. - top_level : bool - Whether its lines start at column 0 of the document, where a fence - can open. - lead : tuple of str, optional - The bullets that will precede the first line on its line, outermost - first (:func:`_line_lead`). - soft_breaks : bool, optional - Whether a plain line break continues a paragraph (see :func:`_units`). - - Returns - ------- - str - The same text with every stand-in spelled. - """ + @property + def next_sibling(self) -> _Node | None: + siblings = self.siblings + index = self._position() + 1 + return siblings[index] if index < len(siblings) else None - if not stand_ins.carried(text): - return text - spelling = _Spelling(text, stand_ins) - spelling.settle_local() - view = spelling.view - blocks = _blocks(view) - for block_number, block in enumerate(blocks): - block_start, block_end = block[0][0], block[-1][1] - for number, (start, end) in enumerate(block): - line = view[start:end] - indent = len(line) - len(line.lstrip(" ")) - if line.startswith("#"): - spelling.escape(start) - if number == 1 and _SETEXT_UNDERLINE.fullmatch(line): - spelling.escape(start) - # python-markdown reads each item's content from just after its - # own bullet, so the first line is a rule if any tail of the - # bullets before it makes one: ``1. - --`` holds ``- --``. - heads = ["".join(lead[i:]) for i in range(len(lead))] - bulleted = heads if block_number == 0 and number == 0 else [] - if any(_RULE.match(head + line) for head in ["", *bulleted]): - for position in range(start, end): - if view[position] in "-*_" and position in spelling.literal: - spelling.escape(position) - if view[position] == "-": - break - if top_level and line.startswith(("```", "~~~")): - if line.startswith("~"): - spelling.escape(start) - else: - position = start - while position < end and view[position] == "`": - spelling.escape(position) - position += 1 - if "|" in line and set(line) <= set("|:- "): - for position in range(start, end): - if view[position] == "|": - spelling.escape(position) - if indent > 3 or indent == len(line): - continue - content = start + indent - first = view[content] - if first == ">": - spelling.escape(content) - if in_list or number == 0: - if first in "-+*" and view[content + 1 : content + 2] == " ": - spelling.escape(content) - ordered = _ORDERED_MARKER.match(view, content, end) - if ordered: - spelling.escape(ordered.end()) - if first == "[" and _REFERENCE_DEFINITION.match( - view[block_start:block_end], start - block_start - ): - spelling.escape(content) - spelling.settle_inline(_units(view, blocks, soft_breaks=soft_breaks)) - return spelling.render() - - -def _settle_inline( - text: str, stand_ins: _StandIns, *, cell: bool = False, heading: bool = False -) -> str: - """Spell the literal text of a table cell or a heading. - - Neither holds block syntax, so only the inline questions are asked, - plus one each: - - * in a cell, every ``|`` -- the table splits the row on it -- and every - backtick, because the table pairs backtick runs across the whole row - to decide which ``|`` it splits on; - * in a heading, the run of ``#`` it ends with, which python-markdown - strips as a closing sequence, and a final ``\\``, after which it - cannot match the heading line at all. - """ + @property + def previous_sibling(self) -> _Node | None: + index = self._position() - 1 + return self.siblings[index] if index >= 0 else None - if not stand_ins.carried(text): - return text - spelling = _Spelling(text, stand_ins) - spelling.settle_local() - view = spelling.view - if cell: - for position in spelling.literal: - if view[position] in "|`": - spelling.escape(position) - if heading: - position = len(view) - 1 - while position >= 0 and view[position] == "#": - spelling.escape(position) - position -= 1 - if view.endswith("\\"): - spelling.escape(len(view) - 1) - spelling.settle_inline([(0, len(view))]) - return spelling.render() - - -def _settle_label(text: str, stand_ins: _StandIns) -> str: - """Spell the brackets in a link's text or an image's alt text. - - python-markdown finds a link's text by counting brackets, so literal - brackets inside one are safe when they balance -- ``rapport - [final].pdf`` -- and end the text early when they do not. They are - left alone when balanced, and all escaped otherwise, or when they would - form an image or a link of their own inside the text. The label's other - characters are left to the block it is in. - """ + @cached_property + def tight(self) -> bool: + """Whether this list is tight: no paragraph in it is shown as one.""" - if not stand_ins.carried(text): - return text - probe = _Spelling(text, stand_ins) - view = probe.view - brackets = sorted(position for position in probe.literal if view[position] in "[]") - if not brackets: - return text - # Only generated code spans can hide a bracket: a literal backtick is - # either escaped later or pairs with nothing. - for position in probe.literal: - if view[position] in "`\\": - probe.escaped.add(position) - probe.settle_backticks([(0, len(view))]) - probe.hide_escapes() - working = list(probe.working()) - depth = 1 - broken = False - for char in working: - if char == "[": - depth += 1 - elif char == "]": - depth -= 1 - if depth == 0: - broken = True - break - broken = broken or depth != 1 - if not broken: - for opener in (position for position in brackets if view[position] == "["): - close = _balanced_close(working, opener + 1, len(working)) - if close is not None and working[close + 1 : close + 2] == ["("]: - broken = True - break - pieces = list(text) - for position in brackets: - pieces[position] = _spelled(view[position]) if broken else view[position] - return "".join(pieces) - - -def _loosened(item: str, loose: bool = False) -> str: - """Put a blank line after a list item's own text, when it is loose anyway. - - python-markdown makes an item *loose* -- its text wrapped in a paragraph - -- as soon as anything inside it is separated by a blank line: a code - block, a second paragraph, a nested item of either. An item following a - loose one is loose too. Reading that output back gives a blank line - between the item's text and the nested list after it, where the first - read had one line break; writing the blank line from the start makes the - first read the one every later read gives. - - The item's own text ends at its first line break that is not a hard - break, since that is the only other newline the reader writes there. - - Parameters - ---------- - item : str - The item's Markdown, before its bullet and indentation. - loose : bool, optional - Whether the item is loose whatever it holds: it follows a loose one. - """ + return all( + grandchild.hidden + for child in self.children + for grandchild in child.children + if grandchild.type == "paragraph" + ) - if not loose and "\n\n" not in item: - return item - position = 0 - while True: - position = item.find("\n", position) - if position < 0: - return item - if item[position - 2 : position] != " ": - break - position += 1 - if item[position + 1 : position + 2] == "\n": - return item - return item[:position] + "\n" + item[position:] +def _list_item(node: RenderTreeNode, context: RenderContext) -> str: + """Render a list item as mdformat does, asking its list's tightness once.""" -#: The elements ``markdownify`` writes as a list. -_LIST_ELEMENTS = frozenset({"ul", "ol", "dir", "menu"}) + separator = "\n" if cast(_Node, node.parent).tight else "\n\n" + text = separator.join( + filter(None, (child.render(context) for child in node.children)) + ) + return text if text.strip() else "" -#: The elements whose text is not displayed, and that the converter drops. -_UNSHOWN_ELEMENTS = frozenset({"script", "style", "title"}) -#: The elements displayed with no text of their own. -_SHOWN_EMPTY = frozenset({"img", "hr"}) +#: ``2k`` backslashes, which mdformat writes for ``k`` literal ones, before a +#: character no backslash escapes. ``2k - 1`` read as the same ``k`` there. +_NEEDLESS_DOUBLING = re.compile(r"(?<!\\)((?:\\\\)+)(?=[^!-/:-@\[-`{-~\\\s])") -#: The blocks whose second line follows their first with no blank line. -_LINED_BLOCKS = _LIST_ELEMENTS | {"blockquote", "table"} +#: The ``<`` mdformat leaves bare after one it escaped: its pattern eats the +#: next character, so ``<<a@b.c>`` became ``\<<a@b.c>``, an autolink. +_SECOND_LESS_THAN = re.compile(r"(?<=\\<)<(?=[^ ]|$)") -def _shows(node: PageElement, within: PageElement) -> bool: - """Return whether ``node`` is text, an image or a rule displayed in ``within``. +def _text(node: RenderTreeNode, context: RenderContext) -> str: + """Render text as mdformat does, less one backslash where it escapes nothing. - ``markdownify`` skips comments and doctypes, and whitespace between - blocks; a ``<br>`` on its own starts no block, so it does not count. + ``C:\\Temp`` stays ``C:\\Temp`` instead of becoming ``C:\\\\Temp``. """ - if isinstance(node, Tag): - if node.name not in _SHOWN_EMPTY: - return False - elif isinstance(node, (Comment, Doctype)) or not str(node).strip(): - return False - for parent in node.parents: - if parent is within: - break - if parent.name in _UNSHOWN_ELEMENTS: - return False - return True - - -def _last_shown(node: PageElement) -> PageElement | None: - """Return the last thing ``node`` displays, or ``None`` if it displays nothing.""" - - if isinstance(node, Tag) and node.name in _UNSHOWN_ELEMENTS: - return None - last = node - while isinstance(last, Tag) and last.contents: - last = last.contents[-1] - for candidate in chain([last], last.previous_elements): - if _shows(candidate, node): - return candidate - if candidate is node: - break - return None - - -def _opens_with_a_block(item: Tag) -> bool: - """Return whether a list item's first line is a list's, a quote's or a table's. + text = DEFAULT_RENDERERS["text"](node, context) + text = _SECOND_LESS_THAN.sub(r"\\<", text) + return _NEEDLESS_DOUBLING.sub(lambda found: found.group(1)[:-1], text) - The item then has no text of its own to set apart with a blank line: - the line after its first is that block's, and a blank line there would - split the block in two. - """ - for candidate in item.descendants: - if _shows(candidate, item): - for element in candidate.parents: - if element is item: - return False - if element.name in _LINED_BLOCKS: - return True - return False - - -def _ends_in_a_list(node: PageElement) -> bool | None: - """Return whether what ``node`` displays last is in a list, if it shows anything.""" - - shown = _last_shown(node) - if shown is None: - return None - for element in chain([shown], shown.parents): - if isinstance(element, Tag) and element.name in _LIST_ELEMENTS: - return True - if element is node: - break - return False - - -def _bullet(item: Tag, numbers: dict[int, int]) -> str: - """Return the marker the reader writes before a list item. - - ``numbers`` holds each ordered item's position among its list's items, - filled for a whole list the first time one of its items is asked for: - counting an item's previous siblings instead cost a long list quadratic - time. - """ +#: A paragraph line opening with ``~~~``, which CommonMark reads as a fence. +_TILDE_FENCE = re.compile("^~~~", re.MULTILINE) - parent = item.parent - if parent is not None and parent.name == "ol": - if id(item) not in numbers: - for index, entry in enumerate(parent.find_all("li", recursive=False)): - numbers[id(entry)] = index - start = str(parent.get("start") or "") - first = int(start) if start.isdigit() else 1 - return f"{first + numbers[id(item)]}. " - return "- " +class _Lists: + """An mdformat extension: linear list items, short text and rules.""" -def _line_items(element: Tag) -> list[Tag]: - """Return the list items whose bullets are written on ``element``'s first line. + RENDERERS: Mapping[str, Callable[[RenderTreeNode, RenderContext], str]] = { + "list_item": _list_item, + "text": _text, + "hr": lambda node, context: "---", + } + #: What mdformat drops: an escape or entity in an image's alt text, which + #: markdown-it-py 3 leaves a ``text_special`` token mdformat renders as + #: nothing, and the escape on a paragraph line opening with ``~~~``. + POSTPROCESSORS: Mapping[str, Postprocess] = { + "text_special": lambda text, node, context: node.markup, + "paragraph": lambda text, node, context: _TILDE_FENCE.sub(r"\\~~~", text), + } - ``element`` itself if it is an item, and every item it opens, through - every block that opens the next, outermost first: ``<li><ul><li>-`` is - written ``- - -``. A quote ends the line's items, since python-markdown - parses a quote's content on its own. - """ - items = [element] if element.name == "li" else [] - node = element - while node.parent is not None: - if any(_last_shown(sibling) is not None for sibling in node.previous_siblings): - break - node = node.parent - if node.name == "blockquote": - break - if node.name == "li": - items.insert(0, node) - return items - - -def _deep_list(element: Tag | None) -> bool: - """Return whether a list's first line carries three bullets or more. - - python-markdown's first pass over a block takes every line after - ``- - - a`` that is indented eight spaces or more as that line's lazy - continuation, so nothing of such a list can follow its first line - directly: its items' own blocks, and its other items, each start a block - of their own, which makes every item of the list loose. - """ +class _Renderer(MDRenderer): + """mdformat's renderer, over :class:`_Node`.""" - return element is not None and len(_line_items(element)) >= 2 + def render( + self, + tokens: Sequence[Token], + options: Mapping[str, Any], + env: MutableMapping[Any, Any], + *, + finalize: bool = True, + ) -> str: + return self.render_tree(_Node(tokens), options, env, finalize=finalize) -def _shown_items(element: Tag) -> int: - """Return how many of a list's items display something.""" +class _Parser(MarkdownIt): + """markdown-it keeping every link: it reformats Markdown, never renders it.""" - items = element.find_all("li", recursive=False) - return sum(_last_shown(item) is not None for item in items) + def validateLink(self, url: str) -> bool: + return True -def _line_lead(element: Tag, numbers: dict[int, int]) -> tuple[str, ...]: - """Return the bullets written on ``element``'s first line, its own included. +def _formatter() -> MarkdownIt: + """Return mdformat as ``mdformat.text`` builds it, less what this module cannot use. - ``- - -`` is a rule to python-markdown, so the first line of a block is - checked against its whole line. ``numbers`` is :func:`_bullet`'s. + ``mdformat.text`` keeps markdown-it's nesting cap of 20, beyond which the + parser drops the rest of the block without a word: a list ten deep lost + its text. The cap is lifted, so a document too deep raises + ``RecursionError`` instead, which :meth:`EasyvistaContentConverter.from_transport` + answers. Link validation is lifted too, or a ``file:`` link read back as + escaped text. """ - return tuple(_bullet(item, numbers) for item in _line_items(element)) + parser = _Parser( + "commonmark", + {"maxNesting": 1_000_000}, + renderer_cls=cast(Any, _Renderer), + ) + parser.options["mdformat"] = { + "wrap": "keep", + "number": True, + "compact_tables": True, + } + parser.options["store_labels"] = True + parser.options["parser_extension"] = [mdformat_tables, _Lists] + parser.options["codeformatters"] = {} + mdformat_tables.update_mdit(parser) + return parser + + +# The shared objects keep no state between calls: markdownify fills a +# per-tag cache of its own methods, the same on every thread, and markdown-it +# and cmark-gfm parse into fresh state. What they set up once -- markdown-it's +# rule chains, cmark-gfm's extension registry, which has no lock -- is set up +# here, under the import lock, before any thread can race for it. +_CONVERTER = _Converter(**_MARKDOWNIFY_OPTIONS) +_FORMATTER = _formatter() +_FORMATTER.render("*warm* [up](x)\n\n- a\n\n| a |\n| - |") +cmarkgfm.cmark.core_extensions_ensure_registered() -def _code_block_fits(element: Tag) -> bool: - """Return whether python-markdown reads an indented code block where ``element`` is. +def html_to_markdown(html: str) -> str: + """Convert HTML to Markdown that cmark-gfm renders as the same display. - Inside a list item or a quote a ``<pre>`` can only be an indented code - block, and there are two places where python-markdown reads none: as the - item's first block, which is always its paragraph, and right after a - list in the same item or quote, since the code's indentation is then - exactly that of the list's last item, and its lines become a paragraph - of that item. - """ + Needs ``beautifulsoup4`` 4.15 or later: before it, a ``<br />`` in a body + that also held a bare ``<br>`` swallowed the text after it. - node: Tag = element - while node.parent is not None: - for sibling in node.previous_siblings: - in_a_list = _ends_in_a_list(sibling) - if in_a_list is not None: - return not in_a_list - node = node.parent - if node.name == "li": - return False - if node.name == "blockquote": - return True - return True - - -class _LiteralSafeConverter(MarkdownConverter): - """``markdownify``'s converter, spelling literal text so it stays text. - - ``escape`` is replaced, and escapes nothing itself: it stands each - character of :data:`_LITERAL` in for (:class:`_StandIns`). The - converters for the containers python-markdown parses on their own -- a - paragraph, a ``<div>``, a list item, a quote, a heading, a table cell, - the document -- then spell every stand-in in their text at once, before - they add their own prefix, with the whole of it in view: - :func:`_settle_block`, :func:`_settle_inline` and :func:`_settle_label` - list the rules. A container inside a heading, a cell or a link leaves its - stand-ins to the one it sits in, which is the unit python-markdown will - parse. - - It also changes what ``markdownify`` produces wherever the result lost - or garbled content that literal-safe escaping would otherwise have kept: - - * a list item's continuation lines are indented by four spaces, which is - what python-markdown nests at, rather than by the bullet's width; - * whatever follows a nested list in its item -- text, a quote -- starts - after a blank line, where ``markdownify`` ran it into the list's last - item; a quote anywhere in an item after its text does too, and a - quote that opens an item keeps its later lines at the start of the - line, the only place python-markdown reads them (:attr:`_StandIns.lazy`); - * an item whose first line it shares with another's bullet is loose, - and so is every item of a list whose first line carries three bullets - or more (:func:`_deep_list`), because python-markdown's first pass - over the block cannot nest anything under such a line; - * an item written loose by python-markdown is written loose here too, so - the second read is the first one (:func:`_loosened`); - * ``<script>``, ``<style>`` and ``<title>`` bodies are dropped, as a - browser drops them; - * ``<s>``, ``<del>`` and ``<strike>`` keep their words without the - ``~~`` markers python-markdown would display; - * an image stays an image inside a heading or a cell, and its alt text - is one line; - * a ``<pre>`` inside a list item or a quote becomes an indented code - block, since a fence only opens at the start of a line, except where - python-markdown reads no code block at all (:func:`_code_block_fits`), - where its lines are kept as literal text; a top-level fence is made - longer than any fence line the code holds; - * a newline in text is a space, as HTML displays it. + Raises + ------ + RecursionError + ``markdownify`` or mdformat recursed deeper than the stack left. """ - def __init__(self, stand_ins: _StandIns, **options: Any) -> None: - super().__init__(**options) - self.stand_ins = stand_ins - #: The list items whose Markdown ends in a blank line, by ``id``. - self._loose_items: set[int] = set() - #: Each ordered item's position in its list, by ``id`` (:func:`_bullet`). - self._numbers: dict[int, int] = {} - #: How many of a list's items display something, by the list's ``id``. - self._shown: dict[int, int] = {} - - def _shown_items(self, element: Tag) -> int: - """Return :func:`_shown_items` for a list, counted once per list.""" - - if id(element) not in self._shown: - self._shown[id(element)] = _shown_items(element) - return self._shown[id(element)] - - def _inherited(self, name: str) -> Callable[..., Any]: - """Return ``markdownify``'s own converter ``name``. - - Its type stub declares the constructor and ``convert`` alone, so the - converters this class extends are reached by name. - """ - - method: Callable[..., Any] = getattr(super(), name) - return method - - # -- text ----------------------------------------------------------------- - - def escape(self, text: str, parent_tags: set[str]) -> str: - return self.stand_ins.shadow(text) if text else "" - - def process_text(self, el: object, parent_tags: set[str] | None = None) -> str: - """Drop the whitespace next to an element displayed as a block. - - ``markdownify`` already does for the elements it knows are blocks; - :data:`_EXTRA_BLOCKS` are the ones it does not. - """ - - text = str(self._inherited("process_text")(el, parent_tags=parent_tags)) - before = getattr(el, "previous_sibling", None) - after = getattr(el, "next_sibling", None) - if isinstance(before, Tag) and before.name in _EXTRA_BLOCKS: - text = text.lstrip(_EDGE) - if isinstance(after, Tag) and after.name in _EXTRA_BLOCKS: - text = text.rstrip(_EDGE) - return text - - # -- containers that settle their own literal text ------------------------ - - @staticmethod - def _deferred(parent_tags: set[str]) -> bool: - """Return whether an enclosing heading, cell or link settles instead.""" - - return "_inline" in parent_tags or "a" in parent_tags - - def _settle( - self, text: str, parent_tags: set[str], lead: tuple[str, ...] = () - ) -> str: - text = self.stand_ins.settle_breaks(text) - text = _LEADING_BLANK_LINES.sub("", text).rstrip(_EDGE) - return _settle_block( - text, - self.stand_ins, - in_list="li" in parent_tags, - top_level="li" not in parent_tags and "blockquote" not in parent_tags, - lead=lead, - ) - - def convert__document_(self, el: Tag, text: str, parent_tags: set[str]) -> str: - """Settle the document, stripped first: edge whitespace is never content. - - Stripped before it is settled, not after, because where the first - line starts decides what it would be read as: ``" # x"`` is not a - heading until the space goes. It is also why the Markdown is stripped - at all -- a digest of a body must not depend on its edges. - """ - - text = self.stand_ins.settle_breaks(text).strip() - return _settle_block(text, self.stand_ins, in_list=False, top_level=True) - - def convert_p(self, el: Tag, text: str, parent_tags: set[str]) -> str: - if self._deferred(parent_tags): - return str(self._inherited("convert_p")(el, text, parent_tags)) - lead = _line_lead(el, self._numbers) if "li" in parent_tags else () - text = self._settle(text, parent_tags, lead=lead) - return f"\n\n{text}\n\n" if text else "" - - def convert_div(self, el: Tag, text: str, parent_tags: set[str]) -> str: - if self._deferred(parent_tags): - return str(self._inherited("convert_div")(el, text, parent_tags)) - lead = _line_lead(el, self._numbers) if "li" in parent_tags else () - text = self._settle(text, parent_tags, lead=lead) - return f"\n\n{text}\n\n" if text else "" - - convert_article = convert_div - convert_center = convert_div - convert_dl = convert_div - convert_section = convert_div - # ``markdownify`` writes a definition as ": text" under its term, one - # line break apart: python-markdown, with no definition-list extension, - # shows the colon and runs the two into one paragraph. Two paragraphs - # keep both texts and nothing else. - convert_dd = convert_div - convert_dt = convert_div - - def convert_blockquote(self, el: Tag, text: str, parent_tags: set[str]) -> str: - if self._deferred(parent_tags): - return str(self._inherited("convert_blockquote")(el, text, parent_tags)) - text = self._settle(text or "", parent_tags | {"blockquote"}) - if not text: - return "\n" - lines = [f"> {line}" if line else ">" for line in text.split("\n")] - if "li" not in parent_tags: - return "\n{}\n\n".format("\n".join(lines)) - if _line_items(el): - # The quote opens a list item: python-markdown never detabs an - # item's first block, so the quote's later lines have to be its - # lazy continuation, at the start of the line. - lines[1:] = [self.stand_ins.lazy + line for line in lines[1:]] - return "\n\n{}\n\n".format("\n".join(lines)) - # Anywhere else in an item a quote opens only after a blank line: the - # item's later lines are indented, and ``>`` counts three spaces in - # at most. - return "\n\n{}\n\n".format("\n".join(lines)) - - def convert_li(self, el: Tag, text: str, parent_tags: set[str]) -> str: - bullet = _bullet(el, self._numbers) - spaced = False - if self._deferred(parent_tags): - text = (text or "").strip() - else: - text = self.stand_ins.settle_breaks(text or "") - text = _LEADING_BLANK_LINES.sub("", text).rstrip(_EDGE) - parent = el.parent - deep = _deep_list(parent) - spaced = deep and parent is not None and self._shown_items(parent) > 1 - # The item after a loose one is loose too: python-markdown reads - # it as the first item of a new block, and wraps its text. So is - # an item that shares its first line with another's bullet: a - # list after its text is eight spaces in, too deep to follow - # that line directly (see _deep_list). - previous = el.find_previous_sibling("li") - stacked = len(_line_items(el)) > 1 - loose = deep or stacked or id(previous) in self._loose_items - if not _opens_with_a_block(el): - text = _loosened(text, loose) - lead = _line_lead(el, self._numbers) - text = self._settle(text, parent_tags | {"li"}, lead=lead) - if not text: - return "\n" - blocks = "\n\n" in text or spaced - if blocks: - self._loose_items.add(id(el)) - first_line, *rest = text.split("\n") - lazy = self.stand_ins.lazy - indented = [ - line if not line or line.startswith(lazy) else f" {line}" - for line in rest - ] - item = "\n".join([bullet + first_line, *indented]) + "\n" - # An item of several blocks needs a blank line after it, or the next - # item's line is read as the last block's lazy continuation. - return item + "\n" if blocks else item - - def convert_hN(self, n: int, el: Tag, text: str, parent_tags: set[str]) -> str: - if self._deferred(parent_tags): - return str(self._inherited("convert_hN")(n, el, text, parent_tags)) - text = _WHITESPACE.sub(" ", text.strip()) - text = _settle_inline(text, self.stand_ins, heading=True) - return "\n\n{} {}\n\n".format("#" * max(1, min(6, n)), text) - - def convert_td(self, el: Tag, text: str, parent_tags: set[str]) -> str: - span = str(el.get("colspan") or "") - colspan = max(1, min(1000, int(span))) if span.isdigit() else 1 - text = _settle_inline( - text.strip().replace("\n", " "), self.stand_ins, cell=True - ) - return " " + text + " |" * colspan - - convert_th = convert_td - - # -- markup ---------------------------------------------------------------- - - def _edges(self, text: str) -> tuple[str, str, str]: - """Split ``text`` into what goes before markup, inside it, and after. + soup = _soup(html) + _flatten_nested_tables(soup) + _drop_trailing_breaks(soup) + _note_sides(soup) + return str(_FORMATTER.render(_CONVERTER.convert_soup(soup))).strip() - ``markdownify``'s ``chomp`` moves spaces outside the markup so it - never opens or closes on whitespace; hard breaks at the edge have to - move out the same way, or ``**a \\n**`` is not emphasis. - """ - - edge = _EDGE + self.stand_ins.hard_break - content = text.strip(edge) - head = text[: len(text) - len(text.lstrip(edge))] - tail = text[len(text.rstrip(edge)) :] if content else "" - return self._edge(head), content, self._edge(tail) - - def _edge(self, side: str) -> str: - breaks = side.count(self.stand_ins.hard_break) - if breaks: - return (self.stand_ins.hard_break + "\n") * breaks - return " " if side else "" - - def _emphasis(self, markup: str, text: str, parent_tags: set[str]) -> str: - if "_noformat" in parent_tags: - return text - before, content, after = self._edges(text) - if not content: - return before - return before + markup + content + markup + after - def convert_b(self, el: Tag, text: str, parent_tags: set[str]) -> str: - return self._emphasis("**", text, parent_tags) +def markdown_to_html(markdown: str) -> str: + """Render Markdown as HTML: CommonMark with GFM tables, through cmark-gfm.""" - convert_strong = convert_b - - def convert_em(self, el: Tag, text: str, parent_tags: set[str]) -> str: - return self._emphasis("*", text, parent_tags) - - convert_i = convert_em - - def convert_s(self, el: Tag, text: str, parent_tags: set[str]) -> str: - """Keep struck text's words: python-markdown has no strikethrough.""" - - return text - - convert_del = convert_s - convert_strike = convert_s - - def convert_title(self, el: Tag, text: str, parent_tags: set[str]) -> str: - """Drop a ``<title>``, which a browser shows only in its tab.""" - - return "" - - def convert_br(self, el: Tag, text: str, parent_tags: set[str]) -> str: - if "_inline" in parent_tags: - return " " - if "pre" in parent_tags: - return "\n" - if "_noformat" in parent_tags: - return " " - return self.stand_ins.hard_break + "\n" - - def convert_a(self, el: Tag, text: str, parent_tags: set[str]) -> str: - if "_noformat" in parent_tags: - return text - before, text, after = self._edges(text) - if not text: - return before - href = el.get("href") - title = el.get("title") - if not href: - return before + text + after - href = str(href) - pasted = self.stand_ins.restore(text) == href - if pasted and not title and _AUTOLINK.fullmatch(f"<{href}>"): - # A pasted URL: the one autolink python-markdown renders as a link. - return f"{before}<{href}>{after}" - text = _settle_label(text, self.stand_ins) - titled = ' "{}"'.format(str(title).replace('"', r"\"")) if title else "" - return f"{before}[{text}]({href}{titled}){after}" - - def convert_img(self, el: Tag, text: str, parent_tags: set[str]) -> str: - alt = _WHITESPACE.sub(" ", str(el.get("alt") or "")).strip() - alt = _settle_label(self.stand_ins.shadow(alt), self.stand_ins) - src = str(el.get("src") or "") - title = _WHITESPACE.sub(" ", str(el.get("title") or "")).strip() - titled = ' "{}"'.format(title.replace('"', r"\"")) if title else "" - return f"![{alt}]({src}{titled})" - - def convert_pre(self, el: Tag, text: str, parent_tags: set[str]) -> str: - if not text: - return "" - text = re.sub(r"[ \n]*$", "", re.sub(r"^[ \n]*\n", "", text)) - nested = "li" in parent_tags or "blockquote" in parent_tags - if nested and not _code_block_fits(el): - # No code block can be written here (see _code_block_fits), so - # the lines are kept as the literal text they are. - lines = [line.strip(" ") for line in text.split("\n")] - joined = (self.stand_ins.hard_break + "\n").join(filter(None, lines)) - return self.stand_ins.shadow(joined) - if nested: - lines = text.split("\n") - code = "\n".join(f" {line}" if line else "" for line in lines) - return f"\n\n{code}\n\n" - fence = "```" - lines = [line.rstrip(" ") for line in text.split("\n")] - while fence in lines: - fence += "`" - return f"\n\n{fence}\n{text}\n{fence}\n\n" - - def convert_list(self, el: Tag, text: str, parent_tags: set[str]) -> str: - """End a nested list with a blank line, for whatever its item holds next. - - ``markdownify`` ends a list inside a list item with no line break at - all, so the item's text after it joined the list's last item, and a - quote after it became that item's lazy continuation. When nothing - follows, the item strips the blank line with the rest of its edge. - """ - - markdown = str(self._inherited("convert_list")(el, text, parent_tags)) - if "li" not in parent_tags or self._deferred(parent_tags): - return markdown - return markdown + "\n\n" - - convert_ul = convert_list - convert_ol = convert_list - convert_dir = convert_list - convert_menu = convert_list - - -#: ``markdownify`` options; the converter's overrides decide the rest. -#: -#: ``wrap`` is what turns a newline inside text into a space, and -#: ``wrap_width=None`` is what stops it wrapping lines. -_MARKDOWNIFY_OPTIONS: dict[str, Any] = {"wrap": True, "wrap_width": None} - - -def html_to_markdown(html: str) -> str: - """Convert HTML to literal-safe Markdown. - - Parameters - ---------- - html : str - The HTML to convert. - - Returns - ------- - str - The Markdown, every literal character spelled so python-markdown - renders it as that character. - """ - - stand_ins = _StandIns(html) - markdown = str( - _LiteralSafeConverter(stand_ins, **_MARKDOWNIFY_OPTIONS).convert(html) + html: str = cmarkgfm.markdown_to_html_with_extensions( + markdown, options=_RENDER_OPTIONS, extensions=["table"] ) - return markdown.replace(stand_ins.lazy, "") + return html.strip() -def _literal_markdown(text: str) -> str: - """Spell plain text as Markdown that renders as that text. - - The degraded path's final step: :func:`_strip_tags` returns text, and - :meth:`EasyvistaContentConverter.from_transport` returns Markdown, so the - text is escaped by the same rules as any text node -- every line a line - start, one container per blank-line-separated block. - """ +def _text_of(html: str) -> str: + """Return the text ``html`` displays, a line per block, without recursing.""" - stand_ins = _StandIns(text) - return _settle_block( - stand_ins.shadow(text), - stand_ins, - in_list=False, - top_level=True, - soft_breaks=True, - ) + pieces: list[str] = [] + for node, _ in _walk(_soup(html)): + if not isinstance(node, Tag): + pieces.append(str(node)) + elif node.name in _BLOCKS or node.name == "br": + pieces.append("\n") + lines = (" ".join(line.split()) for line in "".join(pieces).split("\n")) + return "\n".join(line for line in lines if line) class EasyvistaContentConverter: - """Convert content between EasyVista memo HTML and canonical Markdown. - - Two static methods, one per direction: - :meth:`from_transport` reads a memo -- rich-text HTML or plain text -- - as Markdown, and :meth:`to_transport` renders Markdown as the HTML a - memo is written with. - - Nothing in the client calls this. The read models keep a memo exactly - as the API returned it, and ``TicketContext.to_markdown`` still reduces - memos with the package's dependency-free text reducer, so a caller - converts where it wants Markdown. That is also why the rules live in one - place: every caller converting a memo gets the same Markdown for it. - - Semantics are ``glpi_python_client``'s ``GlpiContentConverter`` at - ``4fc3bed``; see the module docstring for what the port changed. - """ + """Convert content between EasyVista memo HTML and canonical Markdown.""" @staticmethod - def from_transport(value: object) -> str: - """Convert one EasyVista memo value into canonical Markdown. - - Empty input stays empty, plain text is preserved, and HTML content - is converted through ``markdownify`` into Markdown that - python-markdown renders back as the same text: the HTML's text is - literal, so every character that would otherwise read as syntax is - escaped, and nothing else is (see the module docstring and - :class:`_LiteralSafeConverter`). - - The HTML path is taken only when :func:`_looks_like_html` finds a - real element, and both directions of that decision matter, because - a memo is not always HTML: a memo holds what it was sent (tier 4, - see the module docstring), so text a caller wrote through the API - without rendering it is plain text or Markdown. Text sent down the - HTML path loses whatever the parser does not recognise, and Markdown - sent down it comes back escaped -- a value carrying one real element - is read as HTML throughout, so its ``**bold**`` is the eight - characters it spells there. - - There are therefore three outcomes, not two. Real HTML that fits - the stack is converted. Real HTML that does not is stripped to its - text instead: ``markdownify`` recurses two to three frames per - nesting level depending on the interpreter, so the conversion is - *attempted* and its ``RecursionError`` answered, rather than the - depth predicted and a bound applied. The caller gets a readable body - either way; **this method degrades, it does not truncate, and it - does not raise for depth.** - - Attempting it is what makes the answer exact. The budget is not - 1000 frames, it is whatever is left of the stack when the - conversion starts, and that belongs to the caller -- an application - converting a memo from inside a request handler, a template render - or a recursive walk has less of it than a script does. No bound - computed in advance can know that number: glpi_python_client's - earlier design guessed low, 200 against a measured cliff of 494, - and flattened bodies that would have converted. - - One consequence is the price of that exactness and worth naming: - the outcome depends on the caller's remaining stack, so the same - body can convert from one call site and degrade from a deeper one. - Nothing is lost either way -- the degraded rendering keeps every - character of prose -- but a caller comparing two renderings of one - body should know which knob moved it. - - Self-closing void tags are written bare before conversion, which - works around a ``beautifulsoup4`` defect that silently dropped - everything after the second spelling of ``<br>`` in a body that - used both -- see :func:`_canonicalise_void_elements`. - - A document ``html.parser`` refuses outright takes the degraded - path as well. ``<![FOO[`` is the reachable case: an unknown - marked-section keyword, which ``_markupbase`` raises - ``AssertionError`` for and ``bs4`` re-raises as - ``ParserRejectedMarkup``. A caller who can read their text is - better off than one holding an error, and there is nothing else to - be done with such a body, so it is stripped too. The stripped text - is escaped like any other literal text (:func:`_literal_markdown`), - so it renders as itself. - - **It is not a sanitiser.** Text a memo displays as markup -- - ``<script>`` -- comes back escaped, as the text it is; raw HTML - in Markdown a caller writes is another matter, and - :meth:`to_transport` renders it as it is. See the module docstring. + def from_transport(value: object, *, plain_text_is_markdown: bool = False) -> str: + """Convert one EasyVista memo value into Markdown. Parameters ---------- value : object - The memo as ``resolve_memo`` returned it: HTML, plain text or - ``None``. Anything else is converted with ``str`` first. + The memo: HTML, or plain text -- a value with no HTML element + in it. Anything else is converted with ``str`` first. + plain_text_is_markdown : bool, optional + How a value that is not an HTML document is read. ``False``, the + read path, reads plain text as literal characters, one line per + line -- so ``__init__`` comes back escaped. ``True`` is for a + value that is the caller's own Markdown: it passes through, + stripped, unless it starts with ``<`` and holds a real HTML + element anywhere. So Markdown carrying an inline ``<br>`` or + ``<kbd>`` is still Markdown, but Markdown that opens with an + autolink or other angle-bracketed text and carries inline HTML + further on is read as HTML, and loses that autolink. Returns ------- str - The memo as Markdown, stripped of leading and trailing - whitespace; ``""`` for an empty memo. + The Markdown, stripped; empty for an empty value. Raises ------ EasyvistaContentError - The parser failed for some reason other than depth, or the - caller's stack was too short even to strip the tags. The - original exception is attached as ``__cause__``. Nothing is - expected to reach this -- it is here so a parser fault cannot - escape ``except EasyvistaError`` as a bare builtin. + The value could not be converted. A body nested too deeply for + the stack left is read as its text instead, so this is a + backstop. + RecursionError + Only when called from within a few frames of the recursion + limit, where no stack is left even to report the failure as + ``EasyvistaContentError`` (``docs/content.rst`` gives the + measured depths). """ - content = str(value or "") - if not content.strip(): + content = str(value or "").strip() + if not content: return "" + if plain_text_is_markdown and not ( + content.startswith("<") and _looks_like_html(content) + ): + return content if not _looks_like_html(content): - return content.strip() + content = _plain_text_html(content) + # No "<" after the last ">" can finish a tag; CPython's html.parser before + # 3.11.14/3.12.12/3.13.6 rescans to the end for each (quadratic). + head, end, tail = content.rpartition(">") + content = head + end + tail.replace("<", "<") try: - markdown = html_to_markdown(_canonicalise_void_elements(content)) - except (RecursionError, ParserRejectedMarkup): - # The tree is deeper than the stack left, or the parser will - # not build it at all. Both are answered with the text. try: - return _literal_markdown(_strip_tags(content)) - except RecursionError as exc: - # Reachable only from a caller already within a few frames - # of the limit, where stripping cannot run either. Named - # rather than allowed to escape as a bare builtin, which is - # what this taxonomy exists for. - raise EasyvistaContentError( - "Could not convert EasyVista HTML content to Markdown: " - "the caller's stack left too little room even to strip " - "its tags." - ) from exc + return html_to_markdown(content) + except (RecursionError, ValueError): # ValueError: markdownify's + # int() of a colspan or start such as "²" or 5,000 digits + return html_to_markdown(_plain_text_html(_text_of(content))) + except ParserRejectedMarkup: + # html.parser gives up on a few malformed declarations; the + # body's words are still worth more than an exception. + text = unescape(_ANY_TAG.sub(" ", content)) + return html_to_markdown(_plain_text_html(" ".join(text.split()))) except Exception as exc: raise EasyvistaContentError( - "Could not convert EasyVista HTML content to Markdown " + "Could not convert EasyVista memo HTML to Markdown " f"({type(exc).__name__}: {exc})." ) from exc - # Already stripped: the document converter strips before it settles. - return markdown @staticmethod def to_transport(value: object) -> str: - """Convert one canonical Markdown value into EasyVista memo HTML. - - Empty Markdown stays empty, while non-empty content is rendered - through python-markdown with the ``nl2br``, ``sane_lists``, - ``fenced_code`` and ``tables`` extensions and ``html5`` output, so a - lone newline becomes ``<br>``, a fence a ``<pre><code>`` block and a - table a ``<table>``. - - There is no depth ceiling on this direction and no degraded path. - Inbound content is whatever EasyVista happens to hold, so it has to - be survivable; outbound content is what the caller just wrote, so a - failure is worth reporting rather than papering over. ``markdown`` - recurses on nested constructs too -- measured, a list indented 500 - levels raises -- so the failure is caught and named. - - **Nothing is sanitised**, as in glpi_python_client: raw HTML in the - Markdown is passed through verbatim, and a ``javascript:`` link - target becomes a live ``href``. Neutralise both before calling this - on Markdown you did not write. - - Parameters - ---------- - value : object - Markdown, or ``None``. Anything else is converted with ``str`` - first. - - Returns - ------- - str - The rendered HTML, stripped of leading and trailing whitespace; - ``""`` for empty Markdown, never ``<p></p>``. + """Convert one Markdown value into EasyVista memo HTML. Raises ------ EasyvistaContentError - The Markdown could not be rendered. The original exception is - attached as ``__cause__``. + The Markdown could not be rendered. """ markdown = str(value or "") if not markdown.strip(): return "" try: - html = markdown_to_html( - markdown, - extensions=_MARKDOWN_EXTENSIONS, - output_format="html5", - ) + return markdown_to_html(markdown) except Exception as exc: raise EasyvistaContentError( - "Could not render Markdown content as EasyVista HTML " + "Could not render Markdown as EasyVista memo HTML " f"({type(exc).__name__}: {exc})." ) from exc - return str(html).strip() diff --git a/easyvista_python_client/content/tests/display.py b/easyvista_python_client/content/tests/display.py new file mode 100644 index 0000000..6fb4490 --- /dev/null +++ b/easyvista_python_client/content/tests/display.py @@ -0,0 +1,325 @@ +"""What a browser displays of a body, as comparable data. + +The content tests compare displays rather than HTML: the converters +legitimately respell markup -- ``<b>`` becomes ``<strong>``, a ``<div>`` a +``<p>`` -- and none of that is visible. What is kept is what a reader sees: +the blocks, the words in each (whitespace collapsed, as HTML collapses it), +and for every character whether it is bold, italic, code or a link. + +A table's rows are its own: a table nested in a cell is read as words in that +cell, as a browser shows it inside the cell. A table with no row displays +nothing. + +A port of ``glpi_python_client``'s ``content/tests/display.py`` at commit +``524304a``. It replaces the oracle this package's tests used before, which +counted a nested table's rows as the outer table's too: a nested-table +regression read as a match there. :func:`one_line`, :func:`text_words` and +:func:`loose` are new here. Two rules differ from GLPI's, each marked where +it is: + +* inside one line -- a cell, a heading, a link -- the edge of a block or of + a nested table's cell is a word boundary, as a browser shows it (a new + line, a new box) and as the converter writes it (a space); and +* an ordered list's ``start`` is read with ``isdecimal``, as the converter + reads it, where ``isdigit`` made the oracle itself raise on ``"²"``. + +What the oracle cannot see: underline and highlight (``<u>``, ``<mark>``, +``<ins>``), struck text, link and image titles, and line breaks inside a +heading. +""" + +from __future__ import annotations + +import urllib.parse + +from bs4 import BeautifulSoup +from bs4.element import ( + Comment, + Declaration, + Doctype, + NavigableString, + ProcessingInstruction, + Tag, +) + +_HIDDEN = {"script", "style", "title", "head", "template"} +_PARAGRAPHS = {"p", "div", "section", "article", "center", "body", "html", "main"} +_HEADINGS = {f"h{level}" for level in range(1, 7)} +_BLOCKS = _PARAGRAPHS | _HEADINGS | {"ul", "ol", "li", "blockquote", "pre", "table"} +_FORMATS = { + "b": "strong", + "strong": "strong", + "i": "em", + "em": "em", + "code": "code", + "kbd": "code", + "samp": "code", +} +_SKIPPED = (Comment, Doctype, Declaration, ProcessingInstruction) + +#: Differs from GLPI: inside one line -- a cell, a heading, a link -- the +#: edge of a block or of a nested table's cell separates words. A browser +#: starts a new line or draws a new box there, so ``<td>a</td><td>b</td>`` +#: or ``<p>a</p><p>b</p>`` inside a cell never displays ``ab``; GLPI's +#: oracle joined them, which failed the converter for writing ``a b``. +_WORD_EDGES = _BLOCKS | {"td", "th", "tr", "caption"} + +Format = tuple[bool, bool, bool, str | None] + +#: No formatting, and not in a link. +_PLAIN: Format = (False, False, False, None) + + +class _Words: + """The words of one paragraph, each character with its formatting.""" + + def __init__(self) -> None: + self.words: list[tuple[object, ...]] = [] + self.current: list[object] = [] + + def char(self, char: str, fmt: Format) -> None: + if char.isspace(): + self.boundary() + else: + self.current.append((char, fmt)) + + def token(self, token: object) -> None: + self.current.append(token) + + def boundary(self) -> None: + if self.current: + self.words.append(tuple(self.current)) + self.current = [] + + def line_break(self) -> None: + self.boundary() + self.words.append(("BR",)) + + def finish(self) -> tuple[object, ...]: + """The paragraph's words, less the line breaks at its edges.""" + + self.boundary() + words = list(self.words) + while words and words[0] == ("BR",): + words.pop(0) + while words and words[-1] == ("BR",): + words.pop() + return tuple(words) + + +def _with_format(fmt: Format, name: str) -> Format: + strong, em, code, href = fmt + kind = _FORMATS.get(name) + return ( + strong or kind == "strong", + em or kind == "em", + code or kind == "code", + href, + ) + + +def _image(tag: Tag, fmt: Format) -> object: + alt = " ".join(str(tag.get("alt") or "").split()) + return (("IMG", str(tag.get("src") or ""), alt), fmt) + + +def _inline(node: Tag, words: _Words, fmt: Format) -> None: + for child in node.children: + if isinstance(child, _SKIPPED): + continue + if isinstance(child, NavigableString): + for char in str(child): + words.char(char, fmt) + elif isinstance(child, Tag) and child.name not in _HIDDEN: + if child.name == "br": + words.line_break() + elif child.name == "img": + words.token(_image(child, fmt)) + elif child.name == "a" and child.get("href"): + _inline(child, words, (fmt[0], fmt[1], fmt[2], str(child.get("href")))) + elif child.name in _WORD_EDGES: + words.boundary() + _inline(child, words, _with_format(fmt, child.name)) + words.boundary() + else: + _inline(child, words, _with_format(fmt, child.name)) + + +def _has_block(node: Tag) -> bool: + return any( + isinstance(child, Tag) and child.name in _BLOCKS for child in node.descendants + ) + + +def _list(node: Tag, fmt: Format) -> tuple[object, ...]: + """A list: each non-empty item with the number it displays.""" + + ordered = node.name == "ol" + start = str(node.get("start") or "1") + # Differs from GLPI: isdecimal, as the converter reads it. isdigit + # accepts "²", and int("²") raises. + first = int(start) if start.isdecimal() else 1 + items: list[object] = [] + for position, item in enumerate(node.find_all("li", recursive=False)): + content = _display_blocks(item, fmt) + if content: # Markdown cannot spell an empty list item + items.append((first + position if ordered else None, content)) + return (node.name, tuple(items)) if items else () + + +def _table(node: Tag, fmt: Format) -> tuple[object, ...]: + rows = [] + for row in node.find_all("tr"): + if row.find_parent("table") is not node: + continue # a nested table's row shows inside its cell + cells = [] + for cell in row.find_all(["td", "th"], recursive=False): + inner = _Words() + _inline(cell, inner, fmt) + cells.append((cell.name == "th", inner.finish())) + rows.append(tuple(cells)) + return ("table", tuple(rows)) if rows else () + + +def _display_blocks(node: Tag, fmt: Format = _PLAIN) -> tuple[object, ...]: + out: list[object] = [] + words = _Words() + + def flush() -> None: + nonlocal words + paragraph = words.finish() + if paragraph: + out.append(("p", paragraph)) + words = _Words() + + for child in node.children: + if isinstance(child, _SKIPPED): + continue + if isinstance(child, NavigableString): + for char in str(child): + words.char(char, fmt) + continue + if not isinstance(child, Tag) or child.name in _HIDDEN: + continue + name = child.name + if name in _PARAGRAPHS: + flush() + out.extend(_display_blocks(child, fmt)) + elif name in _HEADINGS: + flush() + heading = _Words() + _inline(child, heading, fmt) + out.append((name, heading.finish())) + elif name in {"ul", "ol"}: + flush() + listed = _list(child, fmt) + if listed: + out.append(listed) + elif name == "blockquote": + flush() + quoted = _display_blocks(child, fmt) + if quoted: + out.append(("quote", quoted)) + elif name == "pre": + flush() + out.append(("pre", child.get_text().strip("\n"))) + elif name == "hr": + flush() + out.append(("hr",)) + elif name == "table": + flush() + table = _table(child, fmt) + if table: + out.append(table) + elif name == "br": + words.line_break() + elif name == "img": + words.token(_image(child, fmt)) + elif _has_block(child): + flush() + out.extend(_display_blocks(child, _with_format(fmt, name))) + elif name == "a" and child.get("href"): + _inline(child, words, (fmt[0], fmt[1], fmt[2], str(child.get("href")))) + else: + _inline(child, words, _with_format(fmt, name)) + flush() + return tuple(out) + + +def displayed(html: str) -> tuple[object, ...]: + """Return what a browser displays of ``html``, as comparable data.""" + + return _display_blocks(BeautifulSoup(html, "html.parser")) + + +def one_line(html: str) -> tuple[object, ...]: + """Return ``html`` read as one paragraph of inline content. + + New here, for the one shape whose display the converter changes on + purpose: a table inside a link, which a browser draws as a table and + the converter writes as the link's text (Markdown has no table inside a + link). What it owes there is this: every word, in order, each with its + formatting and its link. + """ + + words = _Words() + _inline(BeautifulSoup(html, "html.parser"), words, _PLAIN) + paragraph = words.finish() + return (("p", paragraph),) if paragraph else () + + +#: Where a browser starts a new word whatever the text says: a block's +#: edges, a line break, a cell and an image. +_BREAKING = frozenset( + """ + address article aside blockquote body br caption center dd details dir div + dl dt fieldset figcaption figure footer form h1 h2 h3 h4 h5 h6 header hr + html img li main menu nav ol p pre section summary table tbody td tfoot th + thead tr ul + """.split() +) + + +def text_words(html: str) -> list[str]: + """Return the words ``html`` displays, in order, without their formatting. + + New here: a blunt check that does not depend on :func:`displayed`. A + word ends only where a browser breaks the text -- see ``_BREAKING`` -- + never at an inline tag's edge, so ``foo<b>bar</b>`` is one word, as it + displays. The walk keeps its own stack, so a deep body costs no + recursion. + """ + + soup = BeautifulSoup(html, "html.parser") + pieces: list[str] = [] + stack: list[object] = [soup] + while stack: + node = stack.pop() + if isinstance(node, Tag): + if node.name in _HIDDEN: + continue + edge = " " if node.name in _BREAKING else "" + pieces.append(edge) + stack.append(edge) # the closing edge, taken after the children + stack.extend(reversed(node.contents)) + elif isinstance(node, _SKIPPED): + continue + else: # a text node, or the str closing edge pushed above + pieces.append(str(node)) + return "".join(pieces).split() + + +def loose(display: object) -> object: + """Return ``display`` with every string percent-decoded. + + cmark-gfm percent-encodes a link or image target it renders -- a space + becomes ``%20``, a backslash ``%5C`` -- and a browser follows either + spelling to the same place. Comparing loosely lets a test hold the + converter to the display without holding it to one URL spelling. + """ + + if isinstance(display, str): + return urllib.parse.unquote(display) + if isinstance(display, tuple): + return tuple(loose(value) for value in display) + return display diff --git a/easyvista_python_client/content/tests/test_conversion.py b/easyvista_python_client/content/tests/test_conversion.py index 315b7cc..7205bdb 100644 --- a/easyvista_python_client/content/tests/test_conversion.py +++ b/easyvista_python_client/content/tests/test_conversion.py @@ -1,176 +1,48 @@ -"""Unit tests for :mod:`easyvista_python_client.content.conversion`. - -Most of this module is a port of ``glpi_python_client``'s -``content/tests/test_conversion.py`` at commit ``0d43528``, brought up to -``4fc3bed`` with the converter: the tests of the converter that module ports. -The two converters drive the same three libraries (``beautifulsoup4``'s -``html.parser`` tree builder, ``markdownify`` and ``python-markdown``), so every -edge case found there is an edge case here, and the tests move with the code. -Where a ported docstring cited something only true of GLPI, it now says whose -measurement it was. The literal-text property -- what a memo displays, the -Markdown displays -- is ``test_literal_text.py``, ported whole. - -Three sections are new: the link and literal-text regressions that motivated -the port (a GLPI description synced to EasyVista on 2026-09-30 arrived with -neither of its links clickable), a handful of memo shapes, and the import guard -for the optional ``content`` extra. - -Every URL here is under ``example.org``. +"""The converter's contract: its two directions, its errors, and what it survives. + +How faithfully realistic content round-trips is :mod:`.test_round_trip`'s +subject; this module pins the interface around it. + +A port of ``glpi_python_client``'s ``content/tests/test_conversion.py`` at +commit ``524304a``, with ``GlpiContentConverter`` renamed +:class:`EasyvistaContentConverter` and ``GlpiContentError`` +:class:`~easyvista_python_client.EasyvistaContentError`. New here: two +tests that a ``ValueError`` is answered by the text fallback before it is +reported (fix 14 of :mod:`.test_fixes`), and fails only if that fails too; +and the section pinning that the converter is not a sanitiser, which +``docs/content.rst`` documents and this package's earlier tests pinned. The +import guard for the optional ``content`` extra is tested in +``easyvista_python_client/testing/test_public_api.py``, in a fresh +interpreter per dependency, which an in-process test here cannot match. """ from __future__ import annotations -import importlib -import re -import sys -from html.parser import HTMLParser - import pytest -from bs4 import BeautifulSoup +from bs4 import BeautifulSoup, ParserRejectedMarkup from easyvista_python_client import EasyvistaContentError, EasyvistaError -from easyvista_python_client.content import EasyvistaContentConverter, conversion -from easyvista_python_client.content.conversion import _strip_tags - -#: Everything that is not a letter or a digit. -_NOT_PROSE = re.compile(r"[^0-9A-Za-z]+") - -#: A Markdown link or image destination. -#: -#: The converting path renders a target that the fallback documents as -#: dropped -- "a link becomes its text without the target, an image -#: contributes nothing" -- so a URL inside ``]( )`` is not prose either, -#: and comparing it would assert a difference the module declares. -_DESTINATION = re.compile(r"\]\([^)]*\)") - -#: A link, and the target only the *converting* path renders. -#: -#: There is no depth number to ask -- the converter attempts the walk and -#: answers the ``RecursionError`` -- so a test that needs to know which path -#: ran has to read the output. A link is the cheapest tell: ``markdownify`` -#: writes ``[probe](u)`` and :func:`_strip_tags` writes ``probe``. -PROBE_LINK = '<a href="u">probe</a>' -PROBE_TARGET = "](u)" - - -def _prose(text: str) -> str: - """Reduce a rendering to its letters and digits, in order. - - Whitespace falls differently at a markup boundary on the two paths -- - the converter joins ``a<b>c`` as ``a**c**`` where the degraded path - joins it as ``ac``, and the degraded path breaks a line at a block - edge the converter runs together -- and the converter adds - punctuation of its own: table pipes, fence backticks, list bullets, - link and image brackets. None of that is prose, and none of it is - what the fallback promises to reproduce. - """ - - return _NOT_PROSE.sub("", _DESTINATION.sub("]", text)) - - -def _displayed(markdown: str) -> str: - """Return the text a reader is shown of the converting path's Markdown. - - That path escapes literal text -- ``<!--``, ``\\_`` -- so its - Markdown is not the text it says: rendering it, as every reader does, - is what gives the text back to compare. The degraded path hands back - plain text, which :func:`_strip_tags` returns as it is. - """ - - rendered = EasyvistaContentConverter.to_transport(markdown) - return BeautifulSoup(rendered, "html.parser").get_text() +from easyvista_python_client.content import conversion +from easyvista_python_client.content.conversion import EasyvistaContentConverter +from easyvista_python_client.content.tests.display import displayed +read = EasyvistaContentConverter.from_transport +render = EasyvistaContentConverter.to_transport -def _is_subsequence(needle: str, haystack: str) -> bool: - remaining = iter(haystack) - return all(character in remaining for character in needle) - -def _parser_rejects(html: str) -> bool: - """Return whether ``html.parser`` gives up on this document. - - ``_markupbase`` raises ``AssertionError`` for an unknown - marked-section keyword, and which keywords count has changed across - CPython patch releases -- so whether a given document is rejected is - a question to ask the running interpreter rather than to assume. - """ - - parser = HTMLParser(convert_charrefs=False) - try: - parser.feed(html) - parser.close() - except AssertionError: - return True - return False - - -def assert_the_degraded_path_says_no_less(shallow: str) -> None: - """Assert the module's one promise about the fallback, on this parser. - - ``html.parser``'s reading of a *malformed* construct is not stable - across CPython patch releases. glpi_python_client measured it on the - same three documents: 3.12.3, 3.12.11 and 3.12.14 disagree about an - unterminated ``<script>``, about a comment with no ``-->``, and about an - end tag carrying a quoted ``>``, each build emitting a different set of - events. Writing down the literal output of one of them made that suite - assert the interpreter's quirks rather than the module's contract, and - it went red on a patch bump while the module itself was fine. - - So the expectation is computed from the converting path rather than - written down. Both paths read the same parser, so they move together, - and the promise was never equality anyway -- it is inclusion: **a - body must not say less because of the path it took.** - """ - - # The padding is closed *before* the construct rather than wrapped - # around it. Wrapping changes what the construct means: an - # unterminated ``<!weird`` runs to the next ``>``, which inside a - # wrapper is the ``>`` of a ``</div>``, so the same text is a bogus - # comment there and character data at end of input. The point of the - # padding is only to be deeper than the converter can walk. - # - # The probe link is how the test knows which path ran: only the - # converting path renders a target, so its absence is the degradation. - # Both renderings are taken from the *same* document, and the - # fallback is called directly rather than provoked with a body deep - # enough to exhaust the stack. Provoking it costs a 600-level tree - # and a walk that runs until it raises, which under coverage - # instrumentation took glpi_python_client's suite from 69 seconds to - # 333; and it tests the routing, which one test can do once, rather - # than the property, which is what every shape here is for. - # - # The probe link carries the document onto the HTML path. Without it - # a fragment whose only tag name is not an element -- ``<scripty>`` -- - # is plain text rather than markup, and the two sides would not be - # renderings of the same thing. - document = PROBE_LINK + shallow - converted = EasyvistaContentConverter.from_transport(document) - degraded = _strip_tags(document) - - assert PROBE_TARGET in converted, "the body must reach the converting path" - - assert _is_subsequence(_prose(_displayed(converted)), _prose(degraded)), ( - f"the degraded path said less than the converting one\n" - f" converted: {converted!r}\n degraded: {degraded!r}" +def test_content_is_markdown_in_python_and_html_for_easyvista() -> None: + assert read("<p>The printer is <strong>offline</strong>.</p>") == ( + "The printer is **offline**." ) - - -def test_content_converter_uses_markdown_in_python_and_html_for_easyvista() -> None: - markdown = EasyvistaContentConverter.from_transport( - "<p>Hello <strong>world</strong></p>" + assert render("The printer is **offline**.") == ( + "<p>The printer is <strong>offline</strong>.</p>" ) - html = EasyvistaContentConverter.to_transport("Hello **world**") - - assert markdown == "Hello **world**" - assert html == "<p>Hello <strong>world</strong></p>" -@pytest.mark.parametrize("value", [None, "", " ", "\n\t "]) -def test_empty_input_stays_empty_in_both_directions(value: object) -> None: - """An empty memo reads as ``""`` and renders as ``""``, never ``<p></p>``.""" - - assert EasyvistaContentConverter.from_transport(value) == "" - assert EasyvistaContentConverter.to_transport(value) == "" +@pytest.mark.parametrize("value", [None, "", " ", "\n\t"]) +def test_an_empty_value_stays_empty(value: object) -> None: + assert read(value) == "" + assert render(value) == "" @pytest.mark.parametrize( @@ -179,1588 +51,210 @@ def test_empty_input_stays_empty_in_both_directions(value: object) -> None: "use the <Enter> key", "cmd </dev/null > out", "if x<y then z>0", - "temp<max and p>min", - "a </close> b", - "<!-- a bare comment -->", "generic<T> in the signature", ], ) -def test_from_transport_preserves_text_whose_tags_are_not_html(text: str) -> None: - """Angle brackets around a non-element name are text, not markup. - - ``<Enter>`` parses as an unknown tag, and an unknown tag's markup is - dropped while its (empty) body is kept -- so the word disappears from the - middle of a sentence with nothing to show it was ever there. - """ - - assert EasyvistaContentConverter.from_transport(text) == text - - -def test_from_transport_still_converts_real_html() -> None: - """Tightening the probe must not stop genuine HTML being normalised.""" - - html = "<p>The printer is <strong>offline</strong>.</p>" - - assert EasyvistaContentConverter.from_transport(html) == ( - "The printer is **offline**." - ) - - -def test_from_transport_leaves_caller_markdown_untouched() -> None: - """A memo holding Markdown or plain text survives the inbound normaliser. - - Not every memo is HTML. A memo holds what it was sent (measured - 2026-09-30 on one instance; may not generalise), so one written through - the API by a caller that did not render its text holds Markdown or plain - text. Anything that sends that down the HTML path escapes it, and the - caller gets back literal backslashes. - """ - - markdown = "The printer is **offline** and 5 * 3 = 15." - - assert EasyvistaContentConverter.from_transport(markdown) == markdown - - -def test_fenced_code_block_survives_the_round_trip() -> None: - """A fence stays a fence. Pasted logs are the common case for this.""" - - markdown = "```\nblock\n```" - - assert ( - EasyvistaContentConverter.from_transport( - EasyvistaContentConverter.to_transport(markdown) - ) - == markdown - ) - - -def test_fenced_code_block_renders_as_a_pre_block() -> None: - """Outbound, a fence becomes ``<pre><code>`` rather than inline code. - - Inline ``<code>`` collapses a multi-line log onto one line in any HTML - renderer that does not style it as preformatted -- glpi_python_client - saw exactly that in GLPI's web UI -- and a read-modify-write then writes - it back as inline code. - """ - - assert EasyvistaContentConverter.to_transport("```\nblock\n```") == ( - "<pre><code>block\n</code></pre>" - ) - - -def test_a_lone_newline_renders_as_a_bare_br() -> None: - """``nl2br`` makes a lone newline a break, spelled the html5 way. - - ``output_format="html5"`` is what writes ``<br>`` rather than - ``<br />``: the bare spelling, which is also the one the inbound void - rewrite normalises to, so what this direction writes is what the other - reads without rewriting. - """ - - assert EasyvistaContentConverter.to_transport("line one\nline two") == ( - "<p>line one<br>\nline two</p>" - ) - - -def test_table_survives_the_round_trip() -> None: - """A Markdown table stays a table instead of degrading to text.""" - - rendered = EasyvistaContentConverter.from_transport( - EasyvistaContentConverter.to_transport("| a | b |\n| - | - |\n| 1 | 2 |") - ) - - assert rendered == "| a | b |\n| --- | --- |\n| 1 | 2 |" - - -@pytest.mark.parametrize( - ("html", "expected"), - [ - ("<p>snake_case name</p>", "snake_case name"), - ("<p>5 * 3 = 15</p>", "5 * 3 = 15"), - ], -) -def test_incoming_text_is_not_backslash_escaped(html: str, expected: str) -> None: - r"""Underscores and asterisks in prose stay readable. - - Escaping them turns ``snake_case`` into ``snake\_case`` on every read, - and the backslash accumulates across read-modify-write cycles. - """ - - assert EasyvistaContentConverter.from_transport(html) == expected - - -# --------------------------------------------------------------------------- -# Links and literal text: the defects this port was made to fix -# --------------------------------------------------------------------------- -# -# Measured 2026-09-30 (tier 4, one instance, may not generalise): a GLPI -# description holding a pasted URL and a titled link was synced to EasyVista -# by a downstream converter written for the purpose, and neither link was -# clickable in EasyVista afterwards. EasyVista stored exactly the HTML it was -# sent, so the damage was all in what was sent: the pasted URL arrived as the -# literal text ``<https://...>`` and the titled link as an ``href`` with -# the title inside it. The same converter also turned two lone asterisks into -# emphasis, doubled the escaping of ``<``, and cut a URL at its first -# ``)``. These tests pin every one of those inputs through this converter, -# with the host replaced by ``example.org``. - -PASTED_URL = "https://example.org/front/ticket.form.php" - - -def test_an_anchor_whose_text_is_its_url_reads_as_an_autolink() -> None: - """A pasted URL is an anchor whose text is its own ``href``. - - ``markdownify`` writes that as a CommonMark autolink, ``<URL>``, which is - the spelling a downstream sanitiser must not mistake for markup. - """ - - html = f'<p><a href="{PASTED_URL}">{PASTED_URL}</a></p>' - - assert EasyvistaContentConverter.from_transport(html) == f"<{PASTED_URL}>" - - -def test_an_autolink_renders_as_a_live_anchor() -> None: - """``<URL>`` goes back out as a clickable anchor, not as literal text.""" - - assert EasyvistaContentConverter.to_transport(f"<{PASTED_URL}>") == ( - f'<p><a href="{PASTED_URL}">{PASTED_URL}</a></p>' - ) - - -@pytest.mark.parametrize( - "markdown", - [ - pytest.param("si a < b alors a", id="escaped"), - pytest.param("si a < b alors a", id="raw"), - ], -) -def test_a_literal_less_than_renders_escaped_exactly_once(markdown: str) -> None: - """``<`` is sent as ``<``, never as ``&lt;``. - - The downstream converter this port replaces HTML-escaped the pivot's - ``<`` a second time, so EasyVista displayed the five characters - ``<`` instead of ``<``. python-markdown passes a character reference - through and escapes a bare ``<`` that starts no tag, so both spellings - land on the same HTML. - """ - - assert EasyvistaContentConverter.to_transport(markdown) == ( - "<p>si a < b alors a</p>" - ) - - -def test_a_literal_less_than_reads_back_as_the_character() -> None: - """Inbound, a ``<`` that opens nothing comes back as a raw ``<``. - - Worth pinning because it is the opposite of the spelling a caller may - use for its own pivot. The reader escapes a ``<`` only where - python-markdown would read one -- before a letter, ``/``, ``!`` or ``?``, - or opening an e-mail autolink -- so ``a < b`` is left as it is, and a - ``<script>`` a memo displays as text comes back as ``<script>`` - (:func:`test_text_a_memo_displays_as_markup_comes_back_as_text`). - """ - - assert EasyvistaContentConverter.from_transport("<p>si a < b alors a</p>") == ( - "si a < b alors a" - ) - - -def test_a_link_title_becomes_a_title_attribute() -> None: - """The title is not part of the target, so it cannot end up in ``href``.""" - - html = EasyvistaContentConverter.to_transport(f'[url]({PASTED_URL} "url")') - - assert html == f'<p><a href="{PASTED_URL}" title="url">url</a></p>' - - -def test_a_title_attribute_reads_back_as_a_link_title() -> None: - html = f'<p><a href="{PASTED_URL}" title="url">url</a></p>' - - assert EasyvistaContentConverter.from_transport(html) == ( - f'[url]({PASTED_URL} "url")' - ) - - -def test_parentheses_inside_a_url_stay_in_the_href() -> None: - """A balanced ``(...)`` belongs to the URL; the link is not cut at ``)``.""" - - target = "https://example.org/wiki/Test_(informatique)" - - assert EasyvistaContentConverter.to_transport(f"[wiki]({target})") == ( - f'<p><a href="{target}">wiki</a></p>' - ) - assert EasyvistaContentConverter.from_transport(f'<a href="{target}">wiki</a>') == ( - f"[wiki]({target})" - ) - - -def test_lone_asterisks_stay_literal() -> None: - """``5*3`` and a free-standing ``*`` are not an emphasis pair. - - python-markdown treats an asterisk with whitespace on both sides as - literal before it looks for emphasis, so the two never pair up. - """ - - text = "prix 5*3 et note * importante" - - assert EasyvistaContentConverter.to_transport(text) == f"<p>{text}</p>" - assert EasyvistaContentConverter.from_transport(f"<p>{text}</p>") == text - - -def test_the_synced_description_renders_both_of_its_links() -> None: - """The exact description of 2026-09-30, host replaced, end to end outbound. - - This is the Markdown ``glpi_python_client`` produced for that GLPI - description -- a no-break space, hard breaks and all. Rendered here, - both links come out as working anchors and the title lands in - ``title``, where the downstream converter produced literal text and a - broken ``href``. - """ - - markdown = f'Test !\xa0 \n \n<{PASTED_URL}> \n \n[url]({PASTED_URL} "url")' - - assert EasyvistaContentConverter.to_transport(markdown) == ( - "<p>Test !\xa0 </p>\n" - f'<p><a href="{PASTED_URL}">{PASTED_URL}</a> </p>\n' - f'<p><a href="{PASTED_URL}" title="url">url</a></p>' - ) - - -def test_the_memo_the_defect_stored_reads_back_as_text_and_a_titled_link() -> None: - """What an already-synced memo reads as, and what writing it back sends. - - The HTML below is the ``COMMENT`` memo EasyVista stored on 2026-09-30, - host replaced (tier 4: that memo, on one instance). It *displays* the - characters ``<`` in front of the URL, and reading it through this - converter keeps exactly that text: the ``&`` would start a reference, so - it is spelled ``&``, and writing the Markdown back sends the very - paragraph the memo holds. The broken ``href`` -- the title fused into - the target -- reads back as a link and a title, so the next write sends a - working titled link. The first link stays text: what EasyVista displayed - was literal text, and literal text is what reads back. - """ - - stored = ( - "<p>Test !</p>" - f"<p>&lt;{PASTED_URL}></p>" - f'<p><a href="{PASTED_URL} "url"">url</a></p>' - ) - - markdown = EasyvistaContentConverter.from_transport(stored) - - assert markdown == (f'Test !\n\n&lt;{PASTED_URL}>\n\n[url]({PASTED_URL} "url")') - assert EasyvistaContentConverter.to_transport(markdown) == ( - "<p>Test !</p>\n" - f"<p>&lt;{PASTED_URL}></p>\n" - f'<p><a href="{PASTED_URL}" title="url">url</a></p>' - ) - - -# --------------------------------------------------------------------------- -# Memo shapes -# --------------------------------------------------------------------------- - - -@pytest.mark.parametrize( - ("memo", "expected"), - [ - # Plain text, as a caller writing through the API without rendering - # stores it. Nothing here is an element, so it is returned as it is, - # line endings included. - pytest.param( - "Bonjour,\r\nle serveur ne répond plus.", - "Bonjour,\r\nle serveur ne répond plus.", - id="plain-text-crlf", - ), - pytest.param( - "Appuyer sur <Entrée> puis valider", - "Appuyer sur <Entrée> puis valider", - id="plain-text-angle-brackets", - ), - # Rich text in the shapes an HTML editor commonly produces. These are - # not transcriptions of EasyVista's own editor output, which nobody - # working on this package has measured. - pytest.param( - "<p>Bonjour,</p><p>Le serveur ne répond plus.<br />" - "Merci de regarder.</p>", - "Bonjour,\n\nLe serveur ne répond plus. \nMerci de regarder.", - id="paragraphs-entity-and-br", - ), - pytest.param( - '<div><span style="font-family: Arial">Texte</span> suite</div>', - "Texte\xa0suite", - id="styled-span-and-nbsp", - ), - pytest.param( - "<ul><li>un</li><li>deux</li></ul>", - "- un\n- deux", - id="list", - ), - pytest.param( - "<p>ligne 1<br>ligne 2</p><p>para 2<br />ligne 4</p>", - "ligne 1 \nligne 2\n\npara 2 \nligne 4", - id="both-br-spellings", - ), - ], -) -def test_memo_shapes_read_as_markdown(memo: str, expected: str) -> None: - assert EasyvistaContentConverter.from_transport(memo) == expected - - -# --------------------------------------------------------------------------- -# What this converter does not do: sanitise -# --------------------------------------------------------------------------- -# -# Pinned so that nobody mistakes it for a sanitiser. glpi_python_client's -# converter behaves the same way, by design: the Markdown is the caller's own. -# A caller relaying Markdown it did not write -- a sync between two ITSMs, for -# one -- has to neutralise raw HTML and executable link schemes itself before -# rendering, and the first three are what it has to cover. The fourth is the -# one thing the inbound direction does now: text a memo displays is literal, -# and comes back spelled as text. - - -def test_raw_html_in_markdown_is_rendered_verbatim() -> None: - assert EasyvistaContentConverter.to_transport("<script>alert(1)</script>") == ( - "<script>alert(1)</script>" - ) - - -def test_a_javascript_link_target_is_rendered_live() -> None: - assert EasyvistaContentConverter.to_transport("[x](javascript:alert(1))") == ( - '<p><a href="javascript:alert(1)">x</a></p>' - ) - - -def test_an_executable_scheme_in_angle_brackets_is_not_made_a_link() -> None: - """python-markdown autolinks ``http``, ``https``, ``ftp`` and ``ftps`` only. - - Anything else in angle brackets passes through as raw markup, which a - browser reads as an unknown element: inert, and invisible. - """ - - assert EasyvistaContentConverter.to_transport("<javascript:alert(1)>") == ( - "<p><javascript:alert(1)></p>" - ) - - -def test_text_a_memo_displays_as_markup_comes_back_as_text() -> None: - """Read and written back, text that looks like markup stays text. - - This used to be the other way round: ``markdownify`` resolved ``<`` - and did not escape the ``<`` it produced, so a memo *showing* the text - ``<script>...`` read as Markdown holding a raw ``<script>`` element, which - the outbound direction then passed through. A ``<`` that python-markdown - would read as a tag is spelled ``<`` now, so the memo is written back - showing what it showed. - """ - - markdown = EasyvistaContentConverter.from_transport( - "<p><script>alert(1)</script></p>" - ) - - assert markdown == "<script>alert(1)</script>" - assert EasyvistaContentConverter.to_transport(markdown) == ( - "<p><script>alert(1)</script></p>" - ) - - -# --------------------------------------------------------------------------- -# Round trips -# --------------------------------------------------------------------------- -# -# ``from_transport(to_transport(m)) == m`` is the property the converter would -# like to hold. It does not hold universally, and cannot: the two libraries -# either side of the wire disagree about a handful of constructs, and no -# option on either fixes them. So, as in glpi_python_client, the corpus is an -# inventory rather than a property test. Every case is listed, the lossy ones -# carry ``xfail(strict=True)``, and that strictness is the point -- fixing one -# turns its xfail into an XPASS and fails the suite, which forces the -# inventory to be updated rather than quietly drifting out of date. -# -# The weaker property is the one a sync depends on: whatever one cycle does, -# a second changes nothing more. That is checked over the same corpus, with -# its own, shorter list of exceptions. - -#: Markdown a caller writes, by case name. -ROUND_TRIP_CORPUS = { - "plain": "The printer is offline.", - "bold": "The printer is **offline**.", - "italic": "This is *emphasis*.", - "inline-code": "Run `systemctl restart` now.", - "heading": "# Title\n\nBody text.", - "subheading": "## Section\n\nBody text.", - "paragraphs": "First para.\n\nSecond para.", - "hard-break": "line one \nline two", - "bullets": "- alpha\n- beta\n- gamma", - "numbered": "1. one\n2. two", - "blockquote": "> quoted text", - "link": "See [the doc](https://example.org/doc).", - "fence": "```\nx = 1\n```", - "table": "| a | b |\n| --- | --- |\n| 1 | 2 |", - "underscore": "The snake_case name.", - "asterisk": "5 * 3 = 15", - "mixed": "# Title\n\n- alpha\n- beta\n\nClosing **note**.", - "autolink": f"<{PASTED_URL}>", - "titled-link": f'[url]({PASTED_URL} "url")', - "parenthesised-url": "[wiki](https://example.org/wiki/Test_(informatique))", - "lone-asterisks": "prix 5*3 et note * importante", - "accents": "Le serveur ne répond plus, merci de vérifier.", - "query-string": "[doc](https://example.org/doc?a=1&b=2)", - "raw-less-than": "si a < b alors a", - "inner-nbsp": "Texte\xa0suite", - "soft-newline": "line one\nline two", - "nested-list": "- alpha\n - inner\n- beta", - "fence-with-language": "```python\nx = 1\n```", - "angle-bracket-text": "use the <Enter> key", - "escaped-less-than": "si a < b alors a", - "email-autolink": "<someone@example.org>", - "synced-description": ( - f'Test !\xa0 \n \n<{PASTED_URL}> \n \n[url]({PASTED_URL} "url")' - ), - "defect-readback": f'Test !\n\n<{PASTED_URL}>\n\n[url]({PASTED_URL} "url")', -} - -#: Cases one write-then-read cycle does not reproduce exactly, and why. -#: -#: Each reason was read off the converter, 2026-09-30, python-markdown 3.10.3 -#: and markdownify 1.2.3; the first four were recorded by glpi_python_client. -LOSSY = { - "soft-newline": ( - "nl2br renders a lone newline as <br>, which markdownify reads back as " - "a hard break (two trailing spaces). Semantically equivalent, and " - "stable after one cycle." - ), - "fence-with-language": ( - "fenced_code emits class='language-python' and markdownify drops the " - "class, so the language tag cannot survive." - ), - "angle-bracket-text": ( - "to_transport does not escape raw markup, so the text reaches EasyVista " - "as a live unknown tag, and the inbound direction drops an unknown " - "tag's markup. The word is gone after one cycle and the doubled space " - "it leaves after two. How EasyVista's UI renders such a tag was not " - "measured." - ), - "escaped-less-than": ( - "markdownify resolves < and does not escape the < it produces, so " - "the reference reads back as a raw <. Both spellings render the same " - "HTML, so this is a change of spelling only." - ), - "email-autolink": ( - "python-markdown renders <user@host> as an entity-obfuscated mailto: " - "anchor, whose text is not its href, so markdownify writes it as an " - "inline [user@host](mailto:user@host) link. Same link." - ), - "synced-description": ( - "the trailing no-break space and the hard breaks that end each " - "paragraph are whitespace markdownify drops at a paragraph's edge. The " - "text and both links survive." - ), -} - -#: Cases a second cycle still changes, and why. -UNSETTLED = { - "angle-bracket-text": LOSSY["angle-bracket-text"], -} - - -def _inventory(known: dict[str, str]) -> list[object]: - """The corpus as parameters, each case in ``known`` a strict xfail.""" - - return [ - pytest.param( - markdown, - id=name, - marks=[pytest.mark.xfail(strict=True, reason=known[name])] - if name in known - else [], - ) - for name, markdown in ROUND_TRIP_CORPUS.items() - ] - - -def test_the_inventories_name_only_cases_in_the_corpus() -> None: - """A misspelt name would silently un-mark a lossy case, so check them. +def test_angle_brackets_that_are_not_html_stay_text(text: str) -> None: + """The element name decides what is markup: ``<Enter>`` is text.""" - And every unsettled case is lossy: a case one cycle reproduces exactly - is, by that token, already settled. - """ - - assert set(LOSSY) <= set(ROUND_TRIP_CORPUS) - assert set(UNSETTLED) <= set(LOSSY) - - -@pytest.mark.parametrize("markdown", _inventory(LOSSY)) -def test_round_trip_corpus(markdown: str) -> None: - """Markdown survives one write-then-read cycle through memo HTML.""" - - html = EasyvistaContentConverter.to_transport(markdown) - - assert EasyvistaContentConverter.from_transport(html) == markdown - - -@pytest.mark.parametrize("markdown", _inventory(UNSETTLED)) -def test_one_cycle_reaches_a_fixed_point(markdown: str) -> None: - """Whatever one cycle changes, a second cycle changes nothing more. + shown = displayed(render(read(text))) - This is what keeps a two-way sync from rewriting a memo on every pass: - after the first write, reading back what was written gives exactly the - Markdown that was written. - """ - - once = EasyvistaContentConverter.from_transport( - EasyvistaContentConverter.to_transport(markdown) - ) - twice = EasyvistaContentConverter.from_transport( - EasyvistaContentConverter.to_transport(once) - ) + assert shown == displayed("<p>" + text.replace("<", "<") + "</p>") - assert twice == once +def test_a_newline_renders_as_a_line_break() -> None: + """A newline the caller typed is a line break: cmark-gfm's hard-break option.""" -@pytest.mark.parametrize( - "memo", - [ - pytest.param( - "<p>Bonjour,</p><p>Le serveur ne répond plus.<br />" - "Merci de regarder.</p>", - id="paragraphs-entity-and-br", - ), - pytest.param( - '<div><span style="font-family: Arial">Texte</span> suite</div>', - id="styled-span-and-nbsp", - ), - pytest.param("<ul><li>un</li><li>deux</li></ul>", id="list"), - pytest.param( - "<p>ligne 1<br>ligne 2</p><p>para 2<br />ligne 4</p>", - id="both-br-spellings", - ), - pytest.param( - "<table><tr><th>a</th><th>b</th></tr><tr><td>1</td><td>2</td></tr></table>", - id="table", - ), - pytest.param("<pre>ligne 1\n ligne 2</pre>", id="preformatted"), - pytest.param("<h2>Titre</h2><p>corps</p>", id="heading"), - pytest.param( - "Bonjour,\r\nle serveur ne répond plus.", - id="plain-text-crlf", - marks=pytest.mark.xfail( - strict=True, - reason=( - "a plain-text memo is returned as it is, lone newline " - "included; the first write renders that newline as <br>, " - "which reads back as a hard break. Stable after that." - ), - ), - ), - pytest.param( - "Appuyer sur <Entrée> puis valider", - id="plain-text-angle-brackets", - marks=pytest.mark.xfail( - strict=True, - reason=( - "the plain-text memo is returned as it is, and the first " - "write sends <Entrée> as a live unknown tag, which reads " - "back as nothing: the word is lost." - ), - ), - ), - pytest.param( - "<ul><li>a<ul><li>b</li></ul></li><li>c</li></ul>", id="nested-list" - ), - pytest.param( - "<p>Test !</p>" - f"<p>&lt;{PASTED_URL}></p>" - f'<p><a href="{PASTED_URL} "url"">url</a></p>', - id="defect-memo", - ), - ], -) -def test_a_memo_read_and_written_back_reads_the_same(memo: str) -> None: - """Reading a memo, writing that Markdown back and reading it again. - - The memo-side twin of the fixed point: if this holds, the first sync of - an already-populated memo is the last one to change it. The exceptions - are the memos whose first write is itself lossy. - """ + assert render("ligne un\nligne deux") == "<p>ligne un<br />\nligne deux</p>" - markdown = EasyvistaContentConverter.from_transport(memo) - written = EasyvistaContentConverter.to_transport(markdown) - assert EasyvistaContentConverter.from_transport(written) == markdown +def test_a_body_using_both_spellings_of_a_line_break_keeps_its_text() -> None: + """beautifulsoup4 before 4.15 dropped the text after ``<br />`` in such a body.""" + markdown = read("<p>line1<br>line2</p><p>para2<br />line4</p>") -# --------------------------------------------------------------------------- -# Nesting depth -# --------------------------------------------------------------------------- -# -# ``markdownify`` walks the parsed tree recursively, so a deeply nested body -# exhausts the interpreter's stack. In glpi_python_client, where this -# converter comes from, the ``RecursionError`` used to surface from inside -# ``model_validate`` -- i.e. from inside ``get_ticket``. These tests pin the -# contract that replaced that: the conversion is attempted, and past the -# stack the body degrades to text rather than raising or being cut short. + assert "line4" in markdown @pytest.mark.parametrize( "html", [ - pytest.param("<div>" * 500 + "text" + "</div>" * 500, id="balanced"), - pytest.param("<p>" * 5000 + "text", id="unclosed"), - pytest.param("<ul><li>" * 500 + "text" + "</li></ul>" * 500, id="lists"), - pytest.param("<foo>" * 5000 + "<p>text</p>", id="unknown-elements"), + pytest.param("<div>" * 3000 + "deep" + "</div>" * 3000, id="balanced"), + pytest.param("<p>" * 5000 + "deep", id="unclosed"), + pytest.param("<ul><li>" * 2000 + "deep" + "</li></ul>" * 2000, id="lists"), pytest.param( - "<table><tr><td>" * 400 + "text" + "</td></tr></table>" * 400, + "<table><tr><td>" * 1500 + "deep" + "</td></tr></table>" * 1500, id="tables", ), - pytest.param("<div>" * 100_000 + "text", id="absurd"), ], ) -def test_deep_html_degrades_instead_of_raising(html: str) -> None: - """Past the ceiling the caller gets a usable body, not an exception. - - Measured from a shallow stack against the default 1000-frame limit, - 493 levels of ``<div>`` is the deepest that converts on CPython 3.12 to - 3.14 and 328 on 3.10. Every shape here is past both, including the - unclosed and unknown-element ones -- which the parser nests just as - deeply as the balanced case. - """ - - assert EasyvistaContentConverter.from_transport(html) == "text" +def test_a_body_too_deep_for_the_stack_degrades_to_its_text(html: str) -> None: + """markdownify recurses per level; past the stack the words are still read.""" + assert "deep" in read(html) -def test_the_degraded_path_keeps_every_word() -> None: - """It degrades; it does not truncate. - A body too deep to convert is still the only copy of what someone - wrote, so the fallback's contract is that all of the text comes back. - """ - - lines = [f"line {index}" for index in range(400)] - html = "".join(f"<div><p>{line}</p>" for line in lines) + "</div>" * 400 - - stripped = EasyvistaContentConverter.from_transport(html) - - assert all(line in stripped for line in lines) - assert "<" not in stripped +_MALFORMED_DECLARATIONS = ["<p>a</p><![FOO[x]]><p>b</p>", "<p>a</p><![ x<p>b</p>"] -def test_the_degraded_path_resolves_entities_and_block_boundaries() -> None: - """Blocks become line breaks, inline tags vanish, references resolve. +@pytest.mark.parametrize("html", _MALFORMED_DECLARATIONS) +def test_a_malformed_declaration_still_reads_both_sides(html: str) -> None: + """Older CPython patch releases reject these; newer ones parse them. - Dropping every tag outright would run ``<p>a</p><p>b</p>`` together as - ``ab``; separating at every tag would break ``<b>off</b>line`` into two - words. Only the block boundary gets a separator. + Measured in glpi_python_client: 3.12.11 raises from ``html.parser``, + 3.12.14 and 3.13.14 read a comment, so only what both outcomes share is + pinned here. """ - html = "<div>" * 600 + "<b>off</b>line & <p>next</p>" + "</div>" * 600 + markdown = read(html) + assert markdown.index("a") < markdown.rindex("b") - assert EasyvistaContentConverter.from_transport(html) == "offline &\nnext" - -@pytest.mark.parametrize( - ("construct", "converted_keeps", "degraded_keeps"), - [ - pytest.param("<!-- SECRET -->", False, False, id="resolved-comment"), - pytest.param("<!DOCTYPE SECRET>", False, False, id="doctype"), - pytest.param("<!SECRET>", False, False, id="bogus-declaration"), - pytest.param("<script>SECRET</script>", False, True, id="script-body"), - pytest.param("<style>SECRET</style>", False, True, id="style-body"), - pytest.param("<![CDATA[SECRET]]>", True, True, id="marked-section"), - pytest.param("<![CDATA[SECRET>", True, True, id="unterminated-marked-section"), - pytest.param("<?php SECRET ?>", True, True, id="processing-instruction"), - ], -) -def test_the_degraded_path_keeps_at_least_what_the_converter_keeps( - construct: str, converted_keeps: bool, degraded_keeps: bool +@pytest.mark.parametrize("html", _MALFORMED_DECLARATIONS) +def test_markup_the_parser_rejects_degrades_to_its_text( + monkeypatch: pytest.MonkeyPatch, html: str ) -> None: - """Parity, construct by construct, and not one of these was a guess. - - Each expectation here was read off the converting path rather than - reasoned about, and two came back the opposite way round from the - obvious answer -- a ``CDATA`` body is *kept*, and so is the inside of - any construct the parser could not resolve. Each of those was a silent - deletion in glpi_python_client's degraded path until it was measured. - - A ``<script>`` or ``<style>`` body is where the two paths part, in the - one direction allowed. The converting path used to keep it -- - ``markdownify``'s ``strip=`` removed the element's markup and still - walked its children -- and now drops it, as a browser does; the - degraded path still keeps it. - - The bar is that a body must not say less because of the path it took, - so a divergence the other way would be a bug even if the text it lost - were JavaScript. - """ - - shallow = f"<p>a</p>{construct}<p>b</p>" - deep = "<div>" * 600 + shallow + "</div>" * 600 - - converted = EasyvistaContentConverter.from_transport(shallow) - degraded = EasyvistaContentConverter.from_transport(deep) - - # What was measured is asserted only where the running parser still - # agrees with the measurement -- the reading of a malformed construct - # moves between CPython patch releases, and it is the superset below, - # not the snapshot, that this module promises. - if ("SECRET" in converted) is converted_keeps: - assert ("SECRET" in degraded) is degraded_keeps - assert not ("SECRET" in converted and "SECRET" not in degraded) - - -def test_an_unterminated_raw_text_element_reads_the_same_on_both_paths() -> None: - """Whether an unclosed ``<script>`` body survives is the parser's call. - - It made glpi_python_client's version of this test its own snapshot: - 3.12.3 discards the body on ``close()`` and 3.12.14 flushes it as - character data, so the literal that was correct on one was wrong on the - other. What has to hold on either is that the two paths agree. - """ - - assert_the_degraded_path_says_no_less("<p>keep</p><script>SECRET") - assert "keep" in EasyvistaContentConverter.from_transport( - "<p>keep</p><script>SECRET" - ) - - -@pytest.mark.parametrize("depth", [1, 100, 200, 250]) -def test_a_document_the_stack_can_hold_is_converted_in_full(depth: int) -> None: - """Everything that fits must convert, and structure has to survive. - - This is what attempting the conversion buys. glpi_python_client's - earlier design predicted the depth and degraded past a fixed 200, - which flattened every body between 200 and the real cliff to text, with - no error to notice and no way for a caller to ask for better; the 250 - case is one that came back as prose. It is also as deep as this goes, - so that it holds on every supported interpreter with room to spare: - CPython 3.10 spends about three frames per level where 3.12 and later - spend two, so its cliff is 328 levels from a shallow stack (measured - 2026-09-30), pytest's own frames come off that, and a 300-level case - left 14 levels of margin. - """ - - html = "<div>" * depth + "<strong>offline</strong>" + "</div>" * depth + """Rejection is forced, so the fallback runs whatever ``html.parser`` does.""" - assert EasyvistaContentConverter.from_transport(html) == "**offline**" + parse = conversion._soup + def rejecting(markup: str) -> BeautifulSoup: + if "<![" in markup: + raise ParserRejectedMarkup(AssertionError("marked section")) + return parse(markup) -def test_void_elements_do_not_spend_the_depth_budget() -> None: - """5000 ``<br>`` is one level, so this must take the converting path. - - The surviving ``**`` proves it: the degraded path strips markup, so - emphasis would be gone if the void tags had nested. - """ - - html = "<p>" + "<br>" * 5000 + "<strong>offline</strong></p>" - - assert "**offline**" in EasyvistaContentConverter.from_transport(html) + monkeypatch.setattr(conversion, "_soup", rejecting) + assert read(html) == "a b" -# --------------------------------------------------------------------------- -# Failure taxonomy -# --------------------------------------------------------------------------- - -def test_a_parser_fault_surfaces_as_an_easyvista_error( +def test_an_inbound_fault_surfaces_as_an_easyvista_error( monkeypatch: pytest.MonkeyPatch, ) -> None: - """No parser fault escapes ``except EasyvistaError``. - - Nothing ordinary reaches this, so the fault is injected. It matters - anyway: a caller who wrapped a sync loop in ``except EasyvistaError`` - would otherwise watch a bare parser exception sail straight through it. + """Differs from GLPI: the fault is a ``RuntimeError``. - A ``RecursionError`` is deliberately *not* the fault used here. It is - not a failure at all -- it is how the converter learns that the - document does not fit, and it is answered with the body's text; see - the test below. + GLPI's test raised ``ValueError``, which since fix 14 goes to the text + fallback first (:func:`test_a_value_error_is_retried_as_text_first`); + a ``RuntimeError`` reaches the error path directly. """ - def _boom(*args: object, **kwargs: object) -> str: - raise ValueError("the parser fell over") + def failing(html: str) -> str: + raise RuntimeError("converter fault") - monkeypatch.setattr(conversion, "html_to_markdown", _boom) + monkeypatch.setattr(conversion, "html_to_markdown", failing) with pytest.raises(EasyvistaContentError) as caught: - EasyvistaContentConverter.from_transport("<p>offline</p>") + read("<p>x</p>") assert isinstance(caught.value, EasyvistaError) - assert isinstance(caught.value.__cause__, ValueError) + assert isinstance(caught.value.__cause__, RuntimeError) + assert "Could not convert EasyVista memo HTML to Markdown" in str(caught.value) -def test_a_recursion_error_degrades_rather_than_raising( +def test_a_value_error_is_retried_as_text_first( monkeypatch: pytest.MonkeyPatch, ) -> None: - """Running out of stack is answered, not reported. + """Fix 14: a ``ValueError`` is answered like a ``RecursionError``, by the text. - Injected rather than provoked, because the depth needed to provoke it - depends on the stack the test runner has already spent -- which is the - very reason the depth is not predicted. What is pinned is the contract: - the caller gets their words, not an exception. + markdownify calls ``int()`` on a ``colspan`` or a ``start``, which raises + on ``"²"`` or on more digits than CPython converts. The body is read again + as its text; only if that fails too is the error reported. """ - def _boom(*args: object, **kwargs: object) -> str: - raise RecursionError("maximum recursion depth exceeded") + calls: list[str] = [] + convert = conversion.html_to_markdown - monkeypatch.setattr(conversion, "html_to_markdown", _boom) + def failing_once(html: str) -> str: + calls.append(html) + if len(calls) == 1: + raise ValueError("invalid literal for int()") + return convert(html) - assert ( - EasyvistaContentConverter.from_transport("<p>Le serveur ne repond plus.</p>") - == "Le serveur ne repond plus." - ) + monkeypatch.setattr(conversion, "html_to_markdown", failing_once) + assert read("<p>un <b>deux</b></p>") == "un deux" + assert len(calls) == 2 -def test_a_stack_too_short_even_to_strip_raises_an_easyvista_error( + +def test_a_value_error_on_the_text_path_too_is_an_easyvista_error( monkeypatch: pytest.MonkeyPatch, ) -> None: - """The one ``RecursionError`` that is reported, and it is reported named. - - Stripping needs a few frames of its own, so a caller already at the - limit cannot have its body degraded either. That is the only case the - inbound direction raises for depth, and it raises inside the taxonomy. - """ - - def _boom(*args: object, **kwargs: object) -> str: - raise RecursionError("maximum recursion depth exceeded") + def failing(html: str) -> str: + raise ValueError("converter fault") - monkeypatch.setattr(conversion, "html_to_markdown", _boom) - monkeypatch.setattr(conversion, "_strip_tags", _boom) + monkeypatch.setattr(conversion, "html_to_markdown", failing) with pytest.raises(EasyvistaContentError) as caught: - EasyvistaContentConverter.from_transport("<p>offline</p>") + read("<p>x</p>") - assert isinstance(caught.value.__cause__, RecursionError) + assert isinstance(caught.value.__cause__, ValueError) -def test_an_outbound_render_fault_surfaces_as_an_easyvista_error( +def test_an_outbound_fault_surfaces_as_an_easyvista_error( monkeypatch: pytest.MonkeyPatch, ) -> None: - """The outbound direction is wrapped too; ``markdown`` recurses as well.""" - - def _boom(*args: object, **kwargs: object) -> str: - raise RecursionError("maximum recursion depth exceeded") + def failing(markdown: str) -> str: + raise RuntimeError("renderer fault") - monkeypatch.setattr(conversion, "markdown_to_html", _boom) + monkeypatch.setattr(conversion, "markdown_to_html", failing) with pytest.raises(EasyvistaContentError) as caught: - EasyvistaContentConverter.to_transport("offline") + render("**x**") - assert isinstance(caught.value, EasyvistaError) - assert isinstance(caught.value.__cause__, RecursionError) - - -def test_deeply_nested_markdown_does_not_raise_a_bare_recursion_error() -> None: - """The real outbound cliff, unmocked. - - ``markdown`` breaks between 495 and 500 levels of list indentation - (measured by glpi_python_client, and again here on 2026-09-30 with - python-markdown 3.10.3). Unlike the inbound direction this is not - degraded -- outbound content is what the caller just wrote, so a - failure is worth reporting -- but it has to be reported as a library - error. - """ - - markdown = "\n".join(" " * level + "- x" for level in range(500)) + assert isinstance(caught.value.__cause__, RuntimeError) + assert "Could not render Markdown as EasyVista memo HTML" in str(caught.value) - with pytest.raises(EasyvistaContentError): - EasyvistaContentConverter.to_transport(markdown) +def test_deeply_nested_markdown_renders() -> None: + """cmark-gfm does not recurse in Python, so depth costs nothing outbound.""" -@pytest.mark.parametrize("dependency", ["bs4", "markdown", "markdownify"]) -def test_importing_without_the_extra_names_the_install_command( - monkeypatch: pytest.MonkeyPatch, dependency: str -) -> None: - """The converter's three dependencies are an extra; say which one. + markdown = "".join(" " * (2 * level) + "- x\n" for level in range(600)) - ``sys.modules[name] = None`` is how the import system spells "this - module is not installed" for the length of the test: the next import - of it raises ``ImportError``. The subpackage is dropped from - ``sys.modules`` so that it really is imported again, and monkeypatch - puts every entry back afterwards. - """ - - monkeypatch.setitem(sys.modules, dependency, None) - monkeypatch.delitem(sys.modules, "easyvista_python_client.content") - monkeypatch.delitem(sys.modules, "easyvista_python_client.content.conversion") - - with pytest.raises(ImportError) as caught: - importlib.import_module("easyvista_python_client.content") - - assert 'pip install "easyvista-python-client[content]"' in str(caught.value) - assert isinstance(caught.value.__cause__, ImportError) + assert render(markdown).count("<li>") == 600 # --------------------------------------------------------------------------- -# The scan against the tree the parser really builds +# It is not a sanitiser # --------------------------------------------------------------------------- # -# The three cases below are the ones a plain open/close counter gets wrong, -# and the first is not a corner case: a stray ``</p>`` or ``</span>`` is -# what a Word or Outlook paste leaves in a rich-text body. Each was measured -# under-counting -- the one direction that turns into a crash -- in -# glpi_python_client, before its scan learned to pop by name. - - -def _parser_depth(html: str) -> int: - """Return the deepest element ``html.parser`` actually builds. - - The ground truth the scan is checked against, walked iteratively so - that measuring a pathological document does not hit the very limit - under test. - """ - - soup = BeautifulSoup(html, "html.parser") - deepest = 0 - pending = [(child, 1) for child in soup.children if getattr(child, "name", None)] - while pending: - node, depth = pending.pop() - deepest = max(deepest, depth) - pending.extend( - (child, depth + 1) - for child in node.children - if getattr(child, "name", None) - ) - return deepest - - -@pytest.mark.parametrize( - "html", - [ - pytest.param("<div></p>" * 600 + "kept", id="stray-close-p"), - pytest.param("<div></span>" * 800 + "kept", id="stray-close-span"), - pytest.param("<li></tr>" * 700 + "kept", id="stray-close-tr"), - ], -) -def test_a_stray_closing_tag_does_not_hide_real_nesting(html: str) -> None: - """The regression test for the under-count that reached ``markdownify``. - - A closing tag with no matching open element pops nothing in ``bs4``, so - these documents nest as deeply as their opening tags say. Counting the - close as a level down measured them at 1, they went to the recursive - converter, and it raised. - """ - - assert EasyvistaContentConverter.from_transport(html) == "kept" - - -def test_the_degraded_path_keeps_a_cdata_body() -> None: - """A ``CDATA`` section's body is text, and text is what survives. - - The declaration pattern that strips ``<!DOCTYPE ...>`` reaches the - first ``>``, and a ``CDATA`` section has none until its end, so it used - to take the body with it -- a silent deletion in the one path whose - whole promise is that nothing is deleted. - """ - - html = "<div>" * 600 + "<p>a<![CDATA[secret words]]>b</p>" + "</div>" * 600 - - assert EasyvistaContentConverter.from_transport(html) == "asecret wordsb" - - -def test_the_degraded_path_keeps_the_text_of_a_broken_comment() -> None: - """Parity with the normal path, even where the parser gave up. - - ``html.parser`` cannot resolve a comment with no ``-->``, so it hands - the region back as character data -- meaning ``markdownify`` would have - kept it. The degraded path is only trustworthy if which path a body - took never changes what it says, so it keeps it too. - """ - - assert_the_degraded_path_says_no_less("<p>keep</p><!--oops but keep this") - - -def test_the_degraded_path_drops_a_comment_the_parser_understood() -> None: - """A resolved comment is text on neither path, so it goes. - - The mirror of the test above, and the reason the two cannot share one - rule: telling them apart is the whole job of the ``-->``. - """ - - html = "<div>" * 600 + "<p>keep</p><!-- drop this -->" + "</div>" * 600 - - assert EasyvistaContentConverter.from_transport(html) == "keep" - - -def test_no_name_in_the_void_set_actually_nests() -> None: - """The one direction of the void set that would be a crash. - - A name listed as void that the parser really nests hides real depth, - and the document then reaches the recursive converter. The set is a - hand-copy of a private ``bs4`` table, so the invariant is asserted - against a real parse rather than against that table: for every name - claimed void, 300 of them must build one level, not 300. - """ - - understated = [ - name - for name in sorted(conversion._VOID_ELEMENTS) - if _parser_depth(f"<{name}>" * 300 + "x") > 1 - ] - - assert understated == [] - - -@pytest.mark.parametrize( - "html", - [ - pytest.param('<div title="</div>">x</div>', id="close-tag-in-attribute"), - pytest.param('<div title="<div>">x</div>', id="open-tag-in-attribute"), - pytest.param('<div title="a>b">x</div>', id="gt-in-attribute"), - pytest.param("<div data-x='a>b'><p>x</p></div>", id="single-quoted"), - pytest.param('<p title=">">one</p><p>two</p>', id="attribute-is-just-gt"), - pytest.param('<div title="</div>">' * 5 + "x", id="nested-and-quoted"), - ], -) -def test_a_tag_inside_a_quoted_attribute_is_not_read_as_markup(html: str) -> None: - """An attribute value may legally contain ``<`` and ``>``. - - Reading a quoted ``</div>`` as a real close tag lost the text after - it, and used to under-count the nesting without bound as well, back - when the nesting was predicted. What has to hold either way is that - the shape costs the body nothing: it converts, and if it is too deep - to convert it still says the same thing. - """ - - assert_the_degraded_path_says_no_less(html) - - -def test_a_quoted_close_tag_at_depth_still_degrades() -> None: - """The same shape scaled past the ceiling: degrades, does not raise.""" - - html = '<div title="</div>">' * 600 + "keep" - - assert EasyvistaContentConverter.from_transport(html) == "keep" - - -@pytest.mark.parametrize( - ("raw", "expected"), - [ - # A URL in a memo is exactly where a semicolon-less reference that - # is a prefix of a longer word shows up. - ("http://x/?a=1©right=2", "http://x/?a=1©right=2"), - ("http://x/?a=1¬anentity=2", "http://x/?a=1¬anentity=2"), - # Terminated references still resolve, named and numeric. - ("a & b < c", "a & b < c"), - ("AB", "AB"), - # ... and so does a semicolon-less reference whose whole name is - # known, because that is what the parser does. - ("© 2026", "© 2026"), - ("a&b", "a&b"), - ], -) -def test_the_degraded_path_resolves_references_like_the_parser( - raw: str, expected: str -) -> None: - """``html.unescape`` alone corrupts URLs; the parser's rule does not. - - ``unescape`` implements HTML5's longest-known-*prefix* rule, so - ``©right=2`` comes back as ``(c)right=2`` -- a query parameter - silently rewritten. The parser behind the converting path resolves a - semicolon-less reference only when the entire name is known, so it - leaves that URL alone, and the degraded path has to agree or a body - changes meaning according to how deeply it nests. - """ - - html = "<div>" * 600 + f"<p>{raw}</p>" + "</div>" * 600 - - assert EasyvistaContentConverter.from_transport(html) == expected - - -@pytest.mark.parametrize( - "prefix", - [ - pytest.param("<style=>", id="malformed-style-name"), - pytest.param("<script=>", id="malformed-script-name"), - pytest.param("<script/>", id="self-closed-script"), - pytest.param("<style />", id="self-closed-style"), - pytest.param("<p title=don't>", id="apostrophe-in-bare-value"), - pytest.param("<p alt=P<0.05>", id="lt-in-bare-value"), - ], -) -def test_a_malformed_tag_does_not_swallow_the_body_after_it(prefix: str) -> None: - """One malformed tag must not take the rest of the document with it. - - Read as raw text, ``"<style=>"`` and ``"<script/>"`` swallowed - everything after them. In glpi_python_client that showed up first as - an unbounded depth under-count -- ``"<style=>" + "<div>" * 600`` - measured **1** level against a real 601 and reached ``markdownify`` -- - but the deletion was always the real damage, and it is what this pins. - """ - - html = prefix + "<div>" * 600 + "the printer is offline" - - assert EasyvistaContentConverter.from_transport(html) == "the printer is offline" - - -def test_the_degraded_path_does_not_emit_a_tag_it_could_not_read() -> None: - """A tag the scan mis-read used to be printed at the reader. - - Measured in glpi_python_client: the body below degraded to ``"<p - title=don't>Le serveur ne repond plus."`` -- the opening tag verbatim - in text a person reads, from a path whose whole promise is text. An - apostrophe is ordinary in French, so this needs no malice to reach a - memo. - """ - - html = "<div>" * 600 + "<p title=don't>Le serveur ne repond plus.</p>" - - degraded = EasyvistaContentConverter.from_transport(html) - - assert degraded == "Le serveur ne repond plus." - assert "<" not in degraded - - -def test_an_unterminated_declaration_at_end_of_input_is_kept() -> None: - """``close()`` flushes an incomplete declaration as text, so this does. - - A declaration is text on neither path only when it is *closed*: a - ``"<!weird"`` that never completes is flushed as character data when - the parser closes, and the converting path prints it. Dropping it lost - the tail of the body. - """ - - assert_the_degraded_path_says_no_less("<p>keep this</p><!weird") - assert "keep this" in EasyvistaContentConverter.from_transport( - "<p>keep this</p><!weird" - ) +# What docs/content.rst ("It is not a sanitiser") and the 0.4.0 entry of +# CHANGELOG.md say, pinned exactly: a cmark-gfm option or release that +# changed any of it would make the documentation wrong in one direction or +# the other. The writing direction neutralises nothing; the reading +# direction keeps a memo's link targets, and keeps text a memo displays as +# text. @pytest.mark.parametrize( - ("html", "expected"), + ("markdown", "html"), [ pytest.param( - "<p>one<br>two<br />three</p>", - "one \ntwo \nthree", - id="br-both-spellings", - ), - pytest.param( - "<p>Bonjour,<br>Le serveur ne repond plus.<br />Merci de regarder.</p>", - "Bonjour, \nLe serveur ne repond plus. \nMerci de regarder.", - id="realistic-body", - ), - pytest.param( - "<p>line1<br>line2</p><p>para2<br />line4</p>", - "line1 \nline2\n\npara2 \nline4", - id="across-paragraphs", - ), - pytest.param( - "<p>a<br>b<br />c<br>d<br />e</p>", - "a \nb \nc \nd \ne", - id="alternating", + "<script>alert(1)</script>", "<script>alert(1)</script>", id="raw-html" ), pytest.param( - "<p>one<img>two<img />three</p>", - "one![]()two![]()three", - id="img", + "[x](javascript:alert(1))", + '<p><a href="javascript:alert(1)">x</a></p>', + id="link-target", ), - pytest.param( - "<p>one<hr>two<hr />three</p>", - "one\n\n---\n\ntwo\n\n---\n\nthree", - id="hr", + pytest.param( # python-markdown, before 0.4.0, left this as raw markup + "<javascript:alert(1)>", + '<p><a href="javascript:alert(1)">javascript:alert(1)</a></p>', + id="autolink", ), ], ) -def test_a_body_using_both_spellings_of_a_void_tag_keeps_its_text( - html: str, expected: str -) -> None: - """Text after the second spelling of ``<br>`` used to be dropped. - - A ``beautifulsoup4`` defect, silent when it fires and reachable from - ordinary editor output: a bare ``<br>`` leaves its name in - ``already_closed_empty_element`` for a ``</br>`` that never comes, and - the next ``<br />`` closes itself against that stale entry and stays - open. Every later sibling becomes its child, and ``convert_br`` - discards an element's children. - - Note the paragraph case: the two spellings need not be near each - other, because a name once recorded poisons the rest of the document. - ``<img>`` and ``<hr>`` are the other two converters that drop - children, so they lose text the same way. - - glpi_python_client measured the defect on ``beautifulsoup4`` 4.14.3. - On 4.15.0 it no longer reproduces (measured 2026-09-30, CPython 3.10, - 3.12, 3.13 and 3.14), so on current releases this passes with or - without the workaround; the extra still accepts 4.12, and the workaround - is what makes it pass there. - """ - - assert EasyvistaContentConverter.from_transport(html) == expected - - -@pytest.mark.parametrize( - "html", - [ - pytest.param("<div/>x", id="self-closed-non-void"), - pytest.param("<custom />x", id="self-closed-unknown"), - pytest.param("<p>a<br>b</p>", id="already-bare"), - pytest.param("<p>2 /> 3</p>", id="slash-gt-in-text"), - pytest.param('<div title="<br />">x</div>', id="in-an-attribute"), - pytest.param("<script>var s = '<br />';</script>x", id="in-a-script-body"), - pytest.param("<!-- <br /> -->x", id="in-a-comment"), - pytest.param("<p>a<br / >b</p>", id="slash-not-abutting-gt"), - ], -) -def test_the_void_rewrite_leaves_everything_else_alone(html: str) -> None: - """The rewrite is confined to void tags in real tag position. - - ``<div/>`` is left as it is -- rewriting it would change what the - document means, and it cannot be affected anyway, since only a void - name is ever recorded as already closed. The last case is the one - worth pinning: ``<br / >`` reaches the parser as an ordinary start - tag, because its ``/`` does not abut the ``>``, so it never takes the - path that loses text and needs no rewriting. - """ - - assert conversion._canonicalise_void_elements(html) == html - - -def test_the_void_rewrite_rewrites_a_self_closed_void_tag() -> None: - """The positive case, pinned on the rewrite itself. - - On ``beautifulsoup4`` 4.15.0 the defect no longer reproduces, so the - conversion tests above pass whether or not the rewrite ran; only this - one notices if it stops running. - """ - - rewritten = conversion._canonicalise_void_elements( - "<p>a<br>b<br />c<img src=x />d</p>" - ) - - assert rewritten == "<p>a<br>b<br>c<img src=x>d</p>" - - -def test_markdownify_is_handed_the_rewritten_document( - monkeypatch: pytest.MonkeyPatch, -) -> None: - """The rewrite is wired in, and not only written. - - The test above pins the rewrite; this pins that ``from_transport`` - applies it. On ``beautifulsoup4`` 4.15.0 dropping the call changes no - output, so an output test cannot notice it -- what ``markdownify`` is - given can. - """ - - seen: list[str] = [] - - def _record(html: str, **options: object) -> str: - seen.append(html) - return "recorded" - - monkeypatch.setattr(conversion, "html_to_markdown", _record) - - EasyvistaContentConverter.from_transport("<p>a<br>b<br />c</p>") - - assert seen == ["<p>a<br>b<br>c</p>"] - - -def test_the_void_rewrite_changes_nothing_for_one_spelling_alone() -> None: - """A body that picks a spelling and keeps it converts exactly as before. - - The rewrite exists to remove an asymmetry between two spellings of the - same node, so it must be invisible to every body that does not mix - them. glpi_python_client measured that over 4000 fuzzed documents of - each spelling: not one output moved. - """ - - bare = "<p>a<br>b<img><hr>c</p>" - slashed = "<p>a<br />b<img /><hr />c</p>" - expected = "a \nb![]()\n\n---\n\nc" - - assert EasyvistaContentConverter.from_transport(bare) == expected - assert EasyvistaContentConverter.from_transport(slashed) == expected - - -def test_both_paths_agree_on_a_body_using_both_spellings() -> None: - """The degraded path keeps this text, and so must the converting one. - - In glpi_python_client this body was the one place where the fallback - said *more* than the conversion it stands in for, which is the wrong - way round for a fallback and was how the defect was noticed at all. - """ - - shallow = "<p>one<br>two<br />three</p>" - deep = "<div>" * 600 + shallow + "</div>" * 600 - - converted = EasyvistaContentConverter.from_transport(shallow) - degraded = EasyvistaContentConverter.from_transport(deep) - - for word in ("one", "two", "three"): - assert word in converted - assert word in degraded - - -@pytest.mark.parametrize( - "fragment", - [ - pytest.param('</x a="><div>">', id="end-tag-with-a-quoted-attribute"), - pytest.param('<div ="<p>', id="name-less-equals-quote"), - pytest.param('<p title="><span>">', id="quoted-gt-then-tag"), - ], -) -def test_a_misread_tag_end_does_not_swallow_the_body_after_it(fragment: str) -> None: - """Scaled past the cliff: degrades quietly, and keeps its words. - - Reading an end tag with start-tag rules consumed everything up to the - next quote, which deleted prose at any depth and under-counted the - nesting 1:1 with the repetition back when the nesting was predicted. - """ - - html = fragment * 600 + "the printer is offline" - - assert "the printer is offline" in EasyvistaContentConverter.from_transport(html) - +def test_the_write_path_renders_what_it_is_given_live(markdown: str, html: str) -> None: + assert render(markdown) == html -def test_a_misread_tag_end_does_not_delete_prose() -> None: - """The other half of the same defect, and it needs no depth at all. - Reading an end tag with attribute rules consumed everything between - the opening quote and its partner, so a degraded body said less than - the converting one -- the divergence the parity test forbids. - """ - - body = '<p>Bonjour</p title="> Le serveur ne repond plus. SECRET ">fin' - - assert_the_degraded_path_says_no_less(body) - assert "Bonjour" in EasyvistaContentConverter.from_transport(body) - - -def test_a_document_with_no_closing_bracket_is_answered_without_scanning() -> None: - """No ``>`` means no element, and saying so keeps a bad shape cheap. - - ``html.parser`` cannot finish a tag that never closes, so ``close()`` - flushes it one character at a time and rescans the tail at each step: - measured in glpi_python_client, 32 KB of ``'<div a="'`` costs it 13 - seconds. Both readers answer that shape directly instead. The text - still survives, because the parser flushes an unfinished tag as data - when it closes. - """ - - html = '<div a="' * 4000 - - assert EasyvistaContentConverter.from_transport(html).startswith('<div a="') - assert "the printer" in EasyvistaContentConverter.from_transport( - html + "the printer" +def test_the_read_path_keeps_a_memos_javascript_link() -> None: + assert read('<p><a href="javascript:alert(1)">x</a></p>') == ( + "[x](<javascript:alert(1)>)" ) - assert _strip_tags(html).startswith('<div a="') - - -def test_a_document_the_parser_rejects_degrades_instead_of_raising() -> None: - """An unknown marked-section keyword stops the parser, not the body. - - ``_markupbase.parse_marked_section`` raises ``AssertionError`` for a - keyword it does not know, and ``bs4`` catches that same - ``AssertionError`` and re-raises it as ``ParserRejectedMarkup``. So a - document carrying ``<![FOO[`` is one the converting path cannot - convert either. Letting it try and fail would hand the caller an - :class:`EasyvistaContentError` and none of their text; sending it down - the degraded path yields its words instead. Degrading beats raising - when the alternative is a body nobody can read. - """ - - html = "<p>Le serveur ne repond plus. SECRET</p><![FOO[x]]>" - - assert "SECRET" in EasyvistaContentConverter.from_transport(html) - if _parser_rejects(html): - # The converting path cannot run at all, so the answer is the text, - # spelled as the Markdown that renders as it. - assert EasyvistaContentConverter.from_transport(html) == ( - conversion._literal_markdown(_strip_tags(html)) - ) -def test_the_text_after_a_construct_the_parser_rejects_is_still_kept() -> None: - """The give-up point is not the end of the body. - - The scan stops where the parser stopped, so everything past that - construct would go missing unless it is handed back explicitly -- - and a body is far more likely to carry the marked section in the - middle than at the end. - """ - - html = "<p>avant</p><![FOO[x]]><p>apres SECRET</p>" - - assert "avant" in _strip_tags(html) - assert "SECRET" in _strip_tags(html) - assert "SECRET" in EasyvistaContentConverter.from_transport(html) - assert "avant" in EasyvistaContentConverter.from_transport(html) - - -def test_stripping_a_document_with_no_closing_bracket_keeps_all_of_it() -> None: - """The degraded path needs the same guard the converting path has. - - ``html.parser`` cannot complete a tag that never closes, so - ``close()`` flushes it one character at a time and rescans the tail - at each step. The whole document is that unfinished tag's text, which - is the answer the guard returns directly. - """ - - html = '<div a="' * 4000 + "le serveur ne repond plus" - - assert _strip_tags(html).endswith("le serveur ne repond plus") - assert _strip_tags(html).startswith('<div a="') - - -@pytest.mark.parametrize( - "html", - [ - # The repetition counts here are deliberately modest. They were - # large when this corpus guarded a depth *prediction*, where the - # error grew with the repetition; the shapes are what matter now, - # and each document has to be one the converting path can still - # walk on every supported interpreter for the comparison to mean - # anything. - pytest.param("<div></p>" * 60 + "kept", id="stray-close-p"), - pytest.param("<div></span>" * 60 + "kept", id="stray-close-span"), - pytest.param("<p></b>" * 60 + "kept", id="stray-close-b"), - pytest.param("<b><i>x</b></i>", id="interleaved"), - pytest.param("<div>" * 100 + "<br>" + "</div>" * 100, id="void-leaf"), - pytest.param("<div>" * 100 + "<img/>" + "</div>" * 100, id="self-closed-leaf"), - pytest.param("<br>" * 5000, id="void-only"), - pytest.param("<div><b>x</b></br></div>", id="close-of-a-void"), - pytest.param("<p>The printer is <strong>offline</strong>.</p>", id="realistic"), - pytest.param("<div><!-- <div><div> --><p>x</p></div>", id="tags-in-a-comment"), - pytest.param( - "<div><script>var s='<div><div>'</script>x</div>", id="tags-in-js" - ), - pytest.param("<table><tr><td>" * 30 + "x", id="tables"), - pytest.param("<blockquote>" * 100 + "x", id="blockquotes"), - pytest.param("<p>a</p>" * 100, id="siblings"), - ], -) -def test_both_paths_agree_about_what_is_markup(html: str) -> None: - """The two renderings of one body must not disagree about its markup. - - This corpus was built in glpi_python_client against a flat scan that - predicted the nesting depth, and it caught the scan reading markup - differently from the parser -- a stray close popping an element the - parser keeps, a void element counted as a parent, a tag inside a - comment or a script body counted at all. The prediction is gone; the - corpus is not, because the same disagreements would show up as the - fallback deleting or inventing text relative to the converting path. - """ - - assert_the_degraded_path_says_no_less(html) - - -@pytest.mark.parametrize( - "html", - [ - # A comment with no ``-->`` is a *bogus comment*: the parser gives up - # at the first ``>``, so the ``</custom>`` inside it is text and the - # ``<br>`` lands inside ``<custom>``. Read that ``</custom>`` as a - # real close and the count comes back one level short -- which is - # how a document that needed degrading reached the converter. - pytest.param("<custom><!--oops</custom><br>", id="bogus-comment-eats-a-close"), - # ... and recovery ends at that ``>``. It does not swallow the rest - # of the document, so these really are two levels. - pytest.param("<!--oops><div><div>", id="bogus-comment-ends-at-its-close"), - # A processing instruction ends at the first ``>`` too, and here - # that lands inside what looks like a comment -- so the second - # ``<div>`` is a real element. Stripping comments globally before - # scanning gets this wrong in both directions at once. - pytest.param("<?php x<!-- <div><div> -->", id="pi-overlapping-a-comment"), - pytest.param("<div><!-- <div><div> --></div>", id="terminated-comment"), - pytest.param("<!DOCTYPE html><div><p>x</p></div>", id="doctype"), - pytest.param("<div><![CDATA[a<div>b]]><p>x</p></div>", id="marked-section"), - pytest.param("<div><script>a<div><div></script><p>x</p></div>", id="raw-text"), - pytest.param("<div><script>a<div>", id="unclosed-raw-text"), - ], -) -def test_every_markup_construct_is_read_the_way_the_parser_reads_it(html: str) -> None: - """Construct by construct, and the order they are tried in matters. - - The parser reads left to right and these constructs overlap: in - ``<?php x<!-- <div><div> -->`` the processing instruction ends at the - first ``>``, which lands inside what looks like a comment, so the - ``<div>`` after it is a real element. Handling any of them out of - order gets that document wrong in both directions at once. - """ - - assert_the_degraded_path_says_no_less(html) - - -@pytest.mark.parametrize( - "html", - [ - pytest.param("<p title=don't>x</p>", id="apostrophe-in-bare-value"), - pytest.param('<p title=say"hi>x</p>', id="quote-in-bare-value"), - pytest.param("<p alt=P<0.05>x</p>", id="lt-in-bare-value"), - pytest.param("<p title=Etape 1>x</p>", id="space-in-bare-value"), - pytest.param("<style=>x", id="malformed-style-name"), - pytest.param("<script=>x", id="malformed-script-name"), - pytest.param("<div=x>y", id="malformed-name"), - pytest.param('<div"a">y', id="quote-in-name"), - pytest.param("<scripty>x</scripty>", id="raw-name-is-a-prefix"), - pytest.param("<script/>x", id="self-closed-script"), - pytest.param("<style />x", id="self-closed-style"), - pytest.param('<li y=">mot<script data-x="</div>">tail', id="lt-after-value"), - pytest.param("<div a=1 <p>text", id="tag-inside-a-tag"), - pytest.param("<div class=a<b>text", id="lt-in-unquoted-value"), - ], -) -def test_a_malformed_tag_is_read_the_way_the_parser_reads_it(html: str) -> None: - """Malformed markup is where every rule taken from the spec was wrong. - - Each shape here was read one way by the HTML5 grammar and another by - ``tagfind_tolerant`` and ``locatestarttagend_tolerant``, which are - what ``markdownify`` actually builds its tree with. The module reads - the parser's events, so the corpus is a guard against a future change - reintroducing a rule from the wrong place. - """ - - assert_the_degraded_path_says_no_less(html) - - -@pytest.mark.parametrize( - ("html", "reason"), - [ - pytest.param( - '</x a="><div>">', - "an end tag skips nothing, so the <div> after it is real", - id="end-tag-with-a-quoted-attribute", - ), - pytest.param( - '<div ="<p><p>">', - "a name-less = starts an attribute NAME, not a quoted value", - id="name-less-equals-quote", - ), - pytest.param( - '<div a ="x>y">', - "whitespace before = still leaves a quoted value", - id="space-before-equals", - ), - pytest.param( - "</ p>x", - "the strict end-tag pattern allows space after </", - id="space-in-end-tag", - ), - pytest.param("</p/>x", "a trailing slash on an end tag", id="slash-in-end-tag"), - pytest.param( - '<li y=">mot<script data-x="</div>">tail', - "< inside a tag", - id="lt-after-a-value", - ), - ], -) -def test_a_tag_end_is_read_the_way_the_parser_reads_it(html: str, reason: str) -> None: - """``parse_endtag`` falls back to ``rawdata.find(">")`` and skips nothing. - - CPython's own comment concedes the consequence -- "this is not 100% - correct, since we might have things like ``</tag attr=">">``" -- so - an end tag read with attribute rules consumes prose the parser keeps. - """ - - assert_the_degraded_path_says_no_less(html) - - -def test_the_fallback_does_not_itself_recurse() -> None: - """Stripping 100k levels must not need 100k frames. - - The fallback exists because the converting path ran out of stack, so - it cannot want a stack of its own. ``html.parser`` is an iterative - scanner and this walks its events into a list, which is what makes it - an answer for input of any depth rather than a second thing to guard. - """ +def test_markup_a_memo_displays_as_text_stays_text_both_ways() -> None: + markdown = read("<p><script></p>") - assert _strip_tags("<div>" * 100_000 + "le serveur") == "le serveur" + assert markdown == "\\<script>" + assert render(markdown) == "<p><script></p>" diff --git a/easyvista_python_client/content/tests/test_cost.py b/easyvista_python_client/content/tests/test_cost.py new file mode 100644 index 0000000..436abf9 --- /dev/null +++ b/easyvista_python_client/content/tests/test_cost.py @@ -0,0 +1,205 @@ +"""What a long or a deep body costs: growth, not seconds. + +A converter that is quadratic somewhere converts an ordinary memo in +milliseconds and a pathological one in minutes, so the cost tests measure +how the time grows: converting a body four times as long should take about +four times as long. A quadratic pass takes about sixteen times as long. The +tests assert ``time(4n) / time(n) < 8``, halfway between the two on a log +scale. + +An absolute budget would be the wrong instrument. ``glpi_python_client``'s +tests at commit ``524304a`` give each long body 20 seconds. That is +generous enough for a shared CI runner, so it catches a quadratic pass only +once a body is long enough to cost 20 seconds, and the tests then spend +most of their time converting. A ratio needs far shorter bodies -- a few +thousand items where GLPI's have 20,000 -- and the runner's speed cancels +out. + +The shapes are GLPI's four long bodies, ported in this form; three bodies +dense with syntax from this package's earlier tests; and one per fix that +removed a quadratic. Measured 2026-10-02 on CPython 3.12.11, timed the way +:func:`growth` times, at these sizes: this converter grew by 2.4 to 6.8 +over 70 readings, three per shape on an idle machine and four per shape +with four copies running at once. GLPI 917f030 grew by 11.0 on the ordered +list, 14.2 on bold holding line breaks and 37.2 on the unfinished-tag tail; +reverting fix 12, 13 or 15 alone grew by 12.5, 16.2 and 32.2 on its shape +(one reading each, idle). The earlier form, +``time(2n) / time(n) < 3``, let a linear pass read up to 2.9 and a +reverted fix 12 read 3.3, and missed GLPI's bold shape in one run of three +(a review, 2026-10-02). + +A ratio between 8 and 10 is measured once more and the second reading +decides, so that one burst of load cannot fail a linear pass; a reading of +10 or more fails at once. +""" + +from __future__ import annotations + +import gc +import math +import time +from collections.abc import Callable + +import pytest + +from easyvista_python_client.content.conversion import EasyvistaContentConverter + +read = EasyvistaContentConverter.from_transport +render = EasyvistaContentConverter.to_transport + +#: ``time(4n) / time(n)`` at or above this fails: linear reads about 4, +#: quadratic about 16. +_LIMIT = 8 + +#: The shortest reading worth trusting: a body that converts faster is read +#: again until the clock has run this long, and the time is the average. +_SHORTEST = 0.02 + + +def _timed(html: str) -> float: + """Seconds to read ``html`` once, with the cyclic collector out of the way. + + A collection starting mid-conversion is charged to whichever body + happens to be converting, so the heap is collected before the clock + starts and the collector stays off until it stops. A body read in a few + milliseconds is read again until :data:`_SHORTEST` has passed, so that + the scheduler's tick does not decide the ratio. + """ + + gc.collect() + gc.disable() + try: + reads = 0 + started = time.perf_counter() + while True: + read(html) + reads += 1 + took = time.perf_counter() - started + if took >= _SHORTEST: + return took / reads + finally: + gc.enable() + + +def growth(make: Callable[[int], str], n: int, rounds: int = 3) -> float: + """Return ``time(make(4n)) / time(make(n))``, each the best of ``rounds``. + + The two sizes alternate, so a burst of load on a shared machine slows + both rather than one; the best of several runs drops the runs it hit. + """ + + small, large = make(n), make(4 * n) + best_small = best_large = math.inf + for _ in range(rounds): + best_small = min(best_small, _timed(small)) + best_large = min(best_large, _timed(large)) + return best_large / best_small + + +def assert_linear(make: Callable[[int], str], n: int) -> None: + ratio = growth(make, n) + if _LIMIT <= ratio < 1.25 * _LIMIT: # near the line: see the module docstring + ratio = growth(make, n) + assert ratio < _LIMIT, ( + f"converting a body four times as long took {ratio:.1f} times as long" + ) + + +@pytest.mark.parametrize( + ("make", "n"), + [ + pytest.param( + lambda n: "<p>" + "ligne<br>" * n + "</p>", 1500, id="line-breaks" + ), + pytest.param( + lambda n: "<ul>" + "<li>x</li>" * n + "</ul>", 1000, id="list-items" + ), + pytest.param( + lambda n: "<ul>" + "<li></li>" * n + "</ul>", 2000, id="empty-items" + ), + pytest.param( + lambda n: "<p>" + "<b>gras</b> mot " * n + "</p>", 750, id="emphasis" + ), + ], +) +def test_a_long_body_converts_in_linear_time( + make: Callable[[int], str], n: int +) -> None: + """GLPI's long bodies: each pass mdformat makes is linear.""" + + assert_linear(make, n) + + +@pytest.mark.parametrize( + ("make", "n"), + [ + pytest.param( + lambda n: "<p>" + "[a " * n + "</p>", 2500, id="unclosed-brackets" + ), + pytest.param( + lambda n: "<p>" + "_a " * n + "</p>", 2500, id="underscores-opening-words" + ), + pytest.param(lambda n: "<p>" + "``` " * n + "</p>", 2000, id="backtick-runs"), + ], +) +def test_a_body_dense_with_syntax_converts_in_linear_time( + make: Callable[[int], str], n: int +) -> None: + """One paragraph packed with what CommonMark scans for, each escaped. + + From this package's earlier tests, where 20,000 of each took from 8 to + 117 seconds before the passes of the converter of the day were made + linear. This converter is not quite linear on them either: the ratio + reads about 4.3 at these sizes and about 7 at eight times them + (measured 2026-10-02, CPython 3.12.11), a superlinear term the lab + traced to markdown-it-py 3.0 joining the text of one long escape-dense + line. So the sizes are kept small: the test catches a pass that turns + quadratic, not that term. + """ + + assert_linear(make, n) + + +def test_a_long_ordered_list_converts_in_linear_time() -> None: + """Fix 12: each item takes its number from the item before it. + + markdownify counted every sibling before each item, which was quadratic. + On 5,000 items GLPI 917f030 took 2.3 seconds against 0.4 with the fix + (measured 2026-10-02, CPython 3.12.11, idle machine). The list is + written one item per line, as an editor writes it: each newline is one + more sibling to count, which makes the quadratic easier to see. + """ + + assert_linear(lambda n: "<ol>\n" + "<li>x</li>\n" * n + "</ol>", 1250) + + +def test_bold_holding_a_long_run_of_line_breaks_converts_in_linear_time() -> None: + """Fix 13: the regex that moves an element's edge breaks outside it is greedy. + + The lazy one rescanned the run after it at every step. + """ + + assert_linear(lambda n: "<p><b>a" + "<br>\n" * n + "b</b></p>", 2000) + + +def test_a_tail_of_unfinished_tags_converts_in_linear_time() -> None: + """Fix 15: no ``<`` after the last ``>`` reaches the parser as a ``<``. + + CPython's ``html.parser`` before 3.11.14, 3.12.12 and 3.13.6 rescanned + to the end of the input for each one (CVE-2025-6069). On a patched + interpreter the parser is linear anyway, so there this test cannot see + the guard go; on an unpatched one it fails within seconds without it. + """ + + assert_linear(lambda n: "<p>r</p>" + "x <a " * n, 500) + + +def test_a_body_too_deep_to_convert_keeps_its_text() -> None: + """markdownify recurses per nesting level; past the stack, the text is kept.""" + + html = "<div>" * 3000 + "<p>__init__ au fond</p>" + "</div>" * 3000 + + markdown = read(html) + + assert "init" in markdown + assert "au fond" in render(markdown) diff --git a/easyvista_python_client/content/tests/test_fixes.py b/easyvista_python_client/content/tests/test_fixes.py new file mode 100644 index 0000000..63e9ba8 --- /dev/null +++ b/easyvista_python_client/content/tests/test_fixes.py @@ -0,0 +1,558 @@ +"""One regression test per fix this package's converter carries over GLPI's. + +The converter is ``glpi_python_client``'s at commit ``917f030`` plus fifteen +fixes, each verified on 2026-10-02 against synthetic generators and against +this package's earlier test corpus. Each section below pins one, numbered +as in the record of that verification: + +1. an image's alt keeps its escapes and entities; +2. ``~~~`` opening a line stays text, not a code fence; +3. a ``!`` ending a text right before a link does not make it an image; +4. an alt opening with ``^`` stays an image; +5. a second ``<`` after an escaped one does not open an autolink; +6. a line break in inline code (``code``, ``kbd``, ``samp``) is kept in a + paragraph and a cell, and is a space in a heading, which is one line; + and a ``|`` in one of them in a cell stays in the cell; +7. a table inside a heading or a link is written as its cells' text, as + one inside a cell already was; +8. a block in such a table stays on its holder's line; +9. bold at a flattened cell's edge still closes; +10. ``<center>`` is a block, except inside ``<pre>``; +11. ``<u>``, ``<mark>`` and ``<ins>`` stay raw HTML; +12. an ordered list is numbered in linear time, and a ``start`` that is + not a decimal number counts from 1; +13. the regex splitting a text's edges is greedy, so linear; +14. a number markdownify cannot read sends the body to the text fallback; +15. an unfinished tag at the very end is read as text (the CVE-2025-6069 + guard). + +The cost side of 12, 13 and 15 is tested in :mod:`.test_cost`. Run against +GLPI 917f030 on 2026-10-02 (CPython 3.12.11), 36 of these 54 tests failed. +Of the 18 that passed there, 17 are controls, guards on a correction to a +fix, pins of output a fix leaves as it was, or behaviour a cost fix had to +keep, and each says which. The other is test 15's ``attribute`` case, which +passed only because 3.12.11 predates CPython's own fix; that test says +why. Every word is invented and every URL is under ``example.org``. +""" + +from __future__ import annotations + +import sys + +import pytest +from bs4 import BeautifulSoup + +from easyvista_python_client.content import conversion +from easyvista_python_client.content.conversion import EasyvistaContentConverter +from easyvista_python_client.content.tests.display import ( + displayed, + one_line, + text_words, +) +from easyvista_python_client.content.tests.test_round_trip import assert_survives + +read = EasyvistaContentConverter.from_transport +render = EasyvistaContentConverter.to_transport + + +def _alts(html: str) -> list[str]: + """The alt of every image, as the HTML parser reads it.""" + + soup = BeautifulSoup(html, "html.parser") + return [str(image.get("alt")) for image in soup.find_all("img")] + + +# --------------------------------------------------------------------------- +# 1. An image's alt keeps its escapes and entities +# --------------------------------------------------------------------------- + + +@pytest.mark.parametrize( + "alt", + [ + pytest.param("a_b C:\\Temp *x* [y] & `z`", id="syntax"), + pytest.param("R&amp;D &copy;", id="literal-entities"), + pytest.param("\\_ et \\*", id="backslashes"), + pytest.param("x < y", id="less-than"), + ], +) +def test_1_an_image_alt_keeps_its_escapes_and_entities(alt: str) -> None: + """markdown-it-py 3 leaves an escape or an entity in alt text as a + ``text_special`` token, which mdformat rendered as nothing.""" + + html = f'<p><img src="https://example.org/i.png" alt="{alt}"></p>' + + markdown = assert_survives(html) + + assert _alts(render(markdown)) == _alts(html) + + +# --------------------------------------------------------------------------- +# 2. '~~~' opening a line stays text +# --------------------------------------------------------------------------- + + +@pytest.mark.parametrize( + "html", + [ + pytest.param("<p>~~~</p>", id="paragraph"), + pytest.param("<p>intro<br>~~~<br>suite</p>", id="after-a-break"), + pytest.param("<ul><li>a<br>~~~~ b</li></ul>", id="in-a-list-item"), + pytest.param("<blockquote><p>a<br>~~~ x</p></blockquote>", id="quoted"), + ], +) +def test_2_a_tilde_run_opening_a_line_stays_text(html: str) -> None: + """CommonMark reads ``~~~`` at a line start as a fence, which swallowed + the rest of the body; mdformat has no guard for it.""" + + markdown = assert_survives(html) + + assert "\\~~~" in markdown + + +def test_2_a_tilde_run_inside_a_line_is_not_escaped() -> None: + """The control: only a run that opens a line is a fence, so only it is escaped.""" + + assert assert_survives("<p>a ~~~ b</p>") == "a ~~~ b" + + +# --------------------------------------------------------------------------- +# 3. A trailing '!' before a link +# --------------------------------------------------------------------------- + + +@pytest.mark.parametrize( + "html", + [ + pytest.param( + '<p>Attention!<a href="https://example.org/u">voir</a></p>', id="link" + ), + pytest.param( + '<p>a!<a href="https://example.org/u">' + '<img src="https://example.org/i.png" alt="x"></a></p>', + id="linked-image", + ), + pytest.param( # the control: "!<url>" was never an image + '<p>a!<a href="https://example.org/u">https://example.org/u</a></p>', + id="autolink", + ), + ], +) +def test_3_a_bang_before_a_link_does_not_make_an_image(html: str) -> None: + """``!`` then ``[text](url)`` is an image in Markdown.""" + + assert_survives(html) + + +def test_3_only_a_trailing_bang_is_escaped() -> None: + """A ``!`` meets a ``[`` only at a text's end, so no other is escaped. + + A pin on the output, not a guard on the fix's narrow form. mdformat + re-renders from the syntax tree and drops an escape that changes + nothing, so escaping every ``!`` gives this same Markdown: a review on + 2026-10-02 found no body, among 3,414 from this suite's generators, on + which the two forms differ. The narrow form was kept as the smaller + change. + """ + + assert assert_survives("<p>Hi! there! ok</p>") == "Hi! there! ok" + + +# --------------------------------------------------------------------------- +# 4. An alt opening with '^' +# --------------------------------------------------------------------------- + + +def test_4_an_alt_opening_with_a_caret_stays_an_image() -> None: + """cmark-gfm never opens an image on ``![^``, the footnote syntax.""" + + html = '<p><img src="https://example.org/i.png" alt="^x"> fin</p>' + + markdown = assert_survives(html) + + assert markdown.startswith("![\\^x]") + + +def test_4_in_code_the_caret_is_not_escaped() -> None: + """Inside code nothing is syntax, so a backslash there would show. + + This guards the fix's correction, not the fix. + """ + + markdown = read( + '<p><code>x<img src="https://example.org/i.png" alt="^x"></code></p>' + ) + + assert "\\^" not in markdown + + +# --------------------------------------------------------------------------- +# 5. A second '<' does not open an autolink +# --------------------------------------------------------------------------- + + +@pytest.mark.parametrize( + "html", + [ + pytest.param("<p><<a@b.c> et <<b> c</p>", id="address"), + pytest.param("<p><<https://example.org></p>", id="url"), + ], +) +def test_5_a_second_less_than_stays_text(html: str) -> None: + """mdformat's escape pattern ate the character after an escaped ``<``, + so ``<<a@b.c>`` became ``\\<<a@b.c>``: an autolink.""" + + markdown = assert_survives(html) + + assert "\\<\\<" in markdown + + +# --------------------------------------------------------------------------- +# 6. A line break in inline code +# --------------------------------------------------------------------------- + + +@pytest.mark.parametrize( + ("html", "expected"), + [ + pytest.param( + "<p>Sortie : <code>ligne un<br>ligne deux</code> fin</p>", + "Sortie : `ligne un`\\\n`ligne deux` fin", + id="paragraph", + ), + pytest.param( + "<p>x <code>ligne un<br> ligne deux</code> y</p>", + "x `ligne un`\\\n`ligne deux` y", + id="space-after-the-break", + ), + pytest.param( + "<table><tr><th>H</th></tr><tr><td>a <code>un<br>deux</code> b</td></tr>" + "</table>", + "| H |\n| -- |\n| a `un`<br>`deux` b |", + id="cell", + ), + ], +) +def test_6_a_line_break_in_inline_code_splits_the_span( + html: str, expected: str +) -> None: + """A code span cannot hold a line break, so the span ends and resumes.""" + + assert assert_survives(html) == expected + + +def test_6_a_line_break_in_inline_code_in_a_heading_is_a_space() -> None: + """A heading is one line in Markdown: the break becomes a space.""" + + markdown = read("<h2>a <code>un<br>deux</code> b</h2>") + + assert markdown == "## a `un` `deux` b" + assert displayed(render(markdown)) == displayed( + "<h2>a <code>un</code> <code>deux</code> b</h2>" + ) + + +@pytest.mark.parametrize("tag", ["code", "kbd", "samp"]) +def test_6_a_pipe_in_inline_code_in_a_cell_stays_in_the_cell(tag: str) -> None: + """GFM splits a row on ``|`` before it reads code. ``code`` is the control: + GLPI 917f030 escaped it there already, but not in ``kbd`` or ``samp``.""" + + html = f"<table><tr><th>H</th></tr><tr><td><{tag}>a|b</{tag}> fin</td></tr></table>" + + markdown = assert_survives(html) + + assert "`a\\|b` fin" in markdown + + +# --------------------------------------------------------------------------- +# 7. A table inside a heading or a link +# --------------------------------------------------------------------------- + + +def test_7_a_table_inside_a_heading_is_its_cells_text() -> None: + markdown = assert_survives( + "<h2>Titre <table><tr><td>a</td><td>b</td></tr></table></h2>" + ) + + assert markdown == "## Titre a b" + + +def test_7_a_table_inside_a_link_is_the_links_text() -> None: + """A browser draws the table inside the link; Markdown cannot, so every + word is kept, in order, still linked.""" + + html = ( + '<p>avant <a href="https://example.org/u"><table><tr><td>un</td><td>deux</td>' + "</tr></table></a> apres</p>" + ) + + markdown = read(html) + + assert markdown == "avant [un deux](https://example.org/u) apres" + assert displayed(render(markdown)) == one_line(html) + assert read(render(markdown)) == markdown + + +def test_7_a_table_inside_a_cell_is_its_cells_text() -> None: + """The control: GLPI 917f030 already flattened a table nested in a cell.""" + + markdown = assert_survives( + "<table><tr><th>A</th></tr><tr><td><table><tr><td>un</td><td>deux</td></tr>" + "</table></td></tr></table>" + ) + + assert markdown == "| A |\n| -- |\n| un deux |" + + +# --------------------------------------------------------------------------- +# 8. A block in a flattened table stays on its holder's line +# --------------------------------------------------------------------------- + + +def test_8_blocks_in_a_table_inside_a_link_stay_on_the_links_line() -> None: + html = ( + '<p>avant <a href="https://example.org/u"><table><tr><td><p>un</p><p>deux</p>' + "</td><td>z</td></tr></table></a> apres</p>" + ) + + markdown = read(html) + + assert markdown == "avant [un deux z](https://example.org/u) apres" + assert displayed(render(markdown)) == one_line(html) + + +# --------------------------------------------------------------------------- +# 9. Bold at a flattened cell's edge still closes +# --------------------------------------------------------------------------- + + +@pytest.mark.parametrize( + "cells", + [ + pytest.param("<td>x</td><td><b>gras.</b></td><td>y</td>", id="ends-on-a-dot"), + pytest.param("<td><b>(gras)</b></td><td>y</td>", id="parenthesised"), + ], +) +def test_9_bold_at_a_flattened_cells_edge_is_markdown_bold(cells: str) -> None: + """The converter spaces a flattened cell, so its edge counts as a space. + GLPI 917f030 counted the next cell's text and wrote raw ``<strong>``, + which read back as ``**`` and was not a fixed point.""" + + markdown = assert_survives( + f"<table><tr><th>A</th></tr><tr><td><table><tr>{cells}</tr></table></td></tr>" + "</table>" + ) + + assert "<strong>" not in markdown + + +# --------------------------------------------------------------------------- +# 10. <center> is a block, except inside <pre> +# --------------------------------------------------------------------------- + + +@pytest.mark.parametrize( + ("html", "expected"), + [ + pytest.param("<center>Titre</center> suite", "Titre\n\nsuite", id="alone"), + pytest.param("<p>a<center>b</center>c</p>", "a\n\nb\n\nc", id="in-a-line"), + ], +) +def test_10_center_is_a_block(html: str, expected: str) -> None: + assert assert_survives(html) == expected + + +def test_10_center_inside_pre_is_left_alone() -> None: + """The correction to the fix, not the fix: as a block there it split the code.""" + + markdown = assert_survives("<pre>un\n<center>deux</center>\ntrois</pre>") + + assert markdown == "```\nun\ndeux\ntrois\n```" + + +# --------------------------------------------------------------------------- +# 11. <u>, <mark> and <ins> stay raw HTML +# --------------------------------------------------------------------------- + + +@pytest.mark.parametrize("tag", ["u", "mark", "ins"]) +def test_11_underline_and_highlight_stay_raw_html(tag: str) -> None: + """CommonMark has neither; GLPI 917f030 dropped the formatting.""" + + markdown = assert_survives(f"<p>a <{tag}>trilo</{tag}> fin</p>") + + assert markdown == f"a <{tag}>trilo</{tag}> fin" + assert f"<{tag}>trilo</{tag}>" in render(markdown) + + +def test_11_the_edges_move_outside_the_tag() -> None: + assert read("<p>a<u> trilo </u>b</p>") == "a <u>trilo</u> b" + + +def test_11_underline_inside_a_link() -> None: + markdown = assert_survives('<p><a href="https://example.org/u"><u>lien</u></a></p>') + + assert markdown == "[<u>lien</u>](https://example.org/u)" + + +# --------------------------------------------------------------------------- +# 12. Ordered-list numbering +# --------------------------------------------------------------------------- + + +@pytest.mark.parametrize( + ("html", "expected"), + [ + pytest.param( + '<ol start="3"><li>a</li><li>b</li><li>c<ol><li>d</li></ol></li></ol>', + "3. a\n4. b\n5. c\n 1. d", + id="start-and-a-nested-list", + ), + pytest.param( + '<ol start="0"><li>a</li><li>b</li></ol>', "0. a\n1. b", id="start-zero" + ), + pytest.param('<ol start="½"><li>a</li><li>b</li></ol>', "1. a\n2. b", id="½"), + pytest.param('<ol start="²"><li>a</li><li>b</li></ol>', "1. a\n2. b", id="²"), + ], +) +def test_12_an_ordered_item_is_numbered_one_past_the_item_before( + html: str, expected: str +) -> None: + """Each item keeps its number for the next, where markdownify counted + every item before each one. ``isdecimal``, where markdownify's + ``isnumeric`` let ``int("½")`` raise -- a browser counts such a list + from 1, and so does the converter now. + + With a decimal ``start``, GLPI 917f030 numbered the same: what the fix + changed there is the cost (:mod:`.test_cost`), and the first two cases + pin that the rewrite numbers as markdownify did. + """ + + assert assert_survives(html) == expected + + +def test_12_an_empty_ordered_item_still_counts_for_the_next() -> None: + """An empty item takes a number and writes nothing. + + Markdown has no empty ordered item, so the item cannot survive: the + browser shows ``3. a`` and ``5. c``, and the Markdown, renumbered by + mdformat, shows ``3. a`` and ``4. c``. The words and the fixed point + are kept. GLPI 917f030 wrote the same; this pins that the rewrite still + does, and no other test reaches its branch for an empty ordered item. + """ + + html = '<ol start="3"><li>a</li><li></li><li>c</li></ol>' + + markdown = read(html) + + assert markdown == "3. a\n4. c" + assert text_words(render(markdown)) == text_words(html) + assert read(render(markdown)) == markdown + + +# --------------------------------------------------------------------------- +# 13. The edge regex is greedy +# --------------------------------------------------------------------------- + + +@pytest.mark.parametrize( + ("text", "groups"), + [ + pytest.param(" a b ", (" ", "a b", " "), id="spaces"), + pytest.param("\\\n a\\\n", ("\\\n ", "a", "\\\n"), id="hard-breaks"), + pytest.param("a\\b", ("", "a\\b", ""), id="inner-backslash"), + pytest.param("x\\", ("", "x\\", ""), id="trailing-backslash"), + pytest.param(" \n ", (" \n ", "", ""), id="nothing-inside"), + pytest.param("", ("", "", ""), id="empty"), + ], +) +def test_13_the_edge_regex_splits_leading_content_and_trailing( + text: str, groups: tuple[str, str, str] +) -> None: + """Leading breaks and spaces, the content, trailing ones. The greedy + pattern finds the content's last character from the end, where the + lazy one rescanned the run after it at every step. The lazy one split + the same way, only slower: these pin that the rewrite still does, and + :mod:`.test_cost` pins the cost.""" + + edges = conversion._EDGES.fullmatch(text) + + assert edges is not None + assert edges.groups() == groups + + +# --------------------------------------------------------------------------- +# 14. A number markdownify cannot read +# --------------------------------------------------------------------------- + + +@pytest.mark.parametrize( + "html", + [ + pytest.param( + '<table><tr><td colspan="²">trelm</td><td>vosk</td></tr></table>', + id="colspan-²", + ), + pytest.param( + '<ol start="' + "7" * 5000 + '"><li>trelm</li><li>vosk</li></ol>', + id="start-of-5000-digits", + ), + ], +) +def test_14_a_number_markdownify_cannot_read_falls_back_to_text(html: str) -> None: + """markdownify's ``int()`` raised ``ValueError`` -- on ``"²"``, or on more + digits than CPython converts -- and the whole memo failed. Now the body + is read as its text: every word, one line per row or item. + + The digit limit is pinned to CPython's default of 4,300: the + interpreter takes another from ``PYTHONINTMAXSTRDIGITS``, and ``0`` + lifts it, under which 5,000 digits are a number and the case tests + nothing. + """ + + limit = sys.get_int_max_str_digits() + sys.set_int_max_str_digits(4300) + try: + markdown = read(html) + finally: + sys.set_int_max_str_digits(limit) + + assert text_words(render(markdown)) == ["trelm", "vosk"] + assert read(render(markdown)) == markdown + + +# --------------------------------------------------------------------------- +# 15. An unfinished tag at the very end +# --------------------------------------------------------------------------- + + +@pytest.mark.parametrize( + ("html", "expected"), + [ + pytest.param("<p>r</p>x <a b", "r\n\nx \\<a b", id="attribute"), + pytest.param("<p>r</p>x <<a", "r\n\nx \\<\\<a", id="doubled"), + pytest.param( # the control: a bare '<' was text on every release + "<p>r</p>fin <", "r\n\nfin \\<", id="bare" + ), + ], +) +def test_15_an_unfinished_tag_at_the_end_reads_as_text( + html: str, expected: str +) -> None: + """No ``<`` after the last ``>`` can finish a tag, so each is text. + + CPython's ``html.parser`` before 3.11.14, 3.12.12 and 3.13.6 rescanned + to the end for each such ``<`` -- quadratic, CVE-2025-6069 -- and the + releases that fixed it drop the unfinished tag instead: on 3.14.6, GLPI + 917f030 read ``x <a b`` as ``x``, where 3.12.11 kept it (measured + 2026-10-02). With the guard every release reads it the same way, as + text. + + So what these cases detect depends on the interpreter. ``attribute`` + fails without the guard only on a patched release; on an unpatched one, + 3.12.11 included, the parser keeps the text by itself, and only the + cost test in :mod:`.test_cost` notices the guard is gone. ``doubled`` + also depends on fix 5, and fails on every release without that one. + """ + + assert read(html) == expected diff --git a/easyvista_python_client/content/tests/test_literal_text.py b/easyvista_python_client/content/tests/test_literal_text.py deleted file mode 100644 index 329d26a..0000000 --- a/easyvista_python_client/content/tests/test_literal_text.py +++ /dev/null @@ -1,1553 +0,0 @@ -"""Literal text reads back as the text it was, not as Markdown syntax. - -A port of ``glpi_python_client``'s ``content/tests/test_literal_text.py`` at -commit ``4fc3bed``, with the converter renamed: the two converters are the -same code, so the property and its corpus move with it. - -``from_transport`` turns a memo's HTML into Markdown, and ``to_transport`` -- -or any python-markdown with the same four extensions, which is what a peer -system renders the Markdown with -- turns it back. Text in the HTML is -literal: a ``__init__`` a user typed is eight characters, not bold ``init``. -Markdown has one spelling for both, so the converter has to spell the literal -one so that python-markdown cannot mistake it, and it has to do it in the -HTML-to-Markdown step: that is the last point where literal text and markup -can still be told apart. - -The property every test here comes back to, checked with an HTML parser -rather than by eye: - -* ``to_transport(from_transport(html))`` **displays what ``html`` displays** - -- the same words, the same formatting on each character, the same block - structure -- for every shape Markdown can express; and -* the Markdown is **a fixed point**: reading back what it renders gives the - same Markdown again. - -It is asserted over hand-picked regressions, over realistic bodies, -and over a seeded fuzzer that puts every character python-markdown treats as -syntax at every position -- line start, word boundary, inside a word, in -cells, list items, headings and link text. - -Escaping is also **minimal**: a character is escaped only where -python-markdown would otherwise read it as syntax, so ordinary prose -- a -file name with underscores, a Windows path, a mid-sentence ``#``, a hyphen, -``R&D`` -- comes back exactly as it did before escaping existed. -""" - -from __future__ import annotations - -import random -import time -from html import escape - -import pytest -from bs4 import ( - BeautifulSoup, - Comment, - Declaration, - Doctype, - NavigableString, - ProcessingInstruction, - Tag, -) -from markdown import Markdown - -from easyvista_python_client.content import EasyvistaContentConverter, conversion - -read = EasyvistaContentConverter.from_transport -render = EasyvistaContentConverter.to_transport - -# --------------------------------------------------------------------------- -# What a browser displays of a body -# --------------------------------------------------------------------------- -# -# The comparison has to be on the display rather than on the HTML: the two -# converters legitimately respell markup -- ``<b>`` becomes ``<strong>``, a -# ``<div>`` becomes a ``<p>``, a list item's lone paragraph loses its -# ``<p>`` -- and none of that is visible. What is kept is what a reader sees: -# the blocks, the words in each (whitespace collapsed, as HTML collapses it), -# and for every character whether it is bold, italic, code or a link. - -_HIDDEN = {"script", "style", "title", "head", "template"} -_PARAGRAPHS = {"p", "div", "section", "article", "center", "body", "html", "main"} -_HEADINGS = {f"h{level}" for level in range(1, 7)} -_BLOCKS = _PARAGRAPHS | _HEADINGS | {"ul", "ol", "li", "blockquote", "pre", "table"} -_FORMATS = { - "b": "strong", - "strong": "strong", - "i": "em", - "em": "em", - "code": "code", - "kbd": "code", - "samp": "code", -} -_SKIPPED = (Comment, Doctype, Declaration, ProcessingInstruction) - -#: No formatting, and not in a link. -_PLAIN: tuple[bool, bool, bool, str | None] = (False, False, False, None) - - -class _Words: - """The words of one paragraph, each character with its formatting.""" - - def __init__(self) -> None: - self.words: list[tuple[object, ...]] = [] - self.current: list[object] = [] - - def char(self, char: str, fmt: tuple[bool, bool, bool, str | None]) -> None: - if char.isspace(): - self.boundary() - else: - self.current.append((char, fmt)) - - def token(self, token: object) -> None: - self.current.append(token) - - def boundary(self) -> None: - if self.current: - self.words.append(tuple(self.current)) - self.current = [] - - def line_break(self) -> None: - self.boundary() - self.words.append(("BR",)) - - def finish(self) -> tuple[object, ...]: - """The paragraph's words, less the line breaks at its edges.""" - - self.boundary() - words = list(self.words) - while words and words[0] == ("BR",): - words.pop(0) - while words and words[-1] == ("BR",): - words.pop() - return tuple(words) - - -def _with_format( - fmt: tuple[bool, bool, bool, str | None], name: str -) -> tuple[bool, bool, bool, str | None]: - strong, em, code, href = fmt - kind = _FORMATS.get(name) - return ( - strong or kind == "strong", - em or kind == "em", - code or kind == "code", - href, - ) - - -def _image(tag: Tag, fmt: tuple[bool, bool, bool, str | None]) -> object: - """An image as displayed: its source, and its alt text as it is read.""" - - alt = " ".join(str(tag.get("alt") or "").split()) - return (("IMG", str(tag.get("src") or ""), alt), fmt) - - -def _inline(node: Tag, words: _Words, fmt: tuple[bool, bool, bool, str | None]) -> None: - for child in node.children: - if isinstance(child, _SKIPPED): - continue - if isinstance(child, NavigableString): - for char in str(child): - words.char(char, fmt) - elif isinstance(child, Tag) and child.name not in _HIDDEN: - if child.name == "br": - words.line_break() - elif child.name == "img": - words.token(_image(child, fmt)) - elif child.name == "a" and child.get("href"): - link = (fmt[0], fmt[1], fmt[2], str(child.get("href"))) - _inline(child, words, link) - else: - _inline(child, words, _with_format(fmt, child.name)) - - -def _has_block(node: Tag) -> bool: - return any( - isinstance(child, Tag) and child.name in _BLOCKS for child in node.descendants - ) - - -def _list(node: Tag, fmt: tuple[bool, bool, bool, str | None]) -> tuple[object, ...]: - """A list: each non-empty item with the number it displays.""" - - ordered = node.name == "ol" - start = str(node.get("start") or "1") - first = int(start) if start.isdigit() else 1 - items: list[object] = [] - for position, item in enumerate(node.find_all("li", recursive=False)): - content = _display_blocks(item, fmt) - if content: # Markdown cannot spell an empty list item - items.append((first + position if ordered else None, content)) - return (node.name, tuple(items)) if items else () - - -def _display_blocks( - node: Tag, fmt: tuple[bool, bool, bool, str | None] = _PLAIN -) -> tuple[object, ...]: - out: list[object] = [] - words = _Words() - - def flush() -> None: - nonlocal words - paragraph = words.finish() - if paragraph: - out.append(("p", paragraph)) - words = _Words() - - for child in node.children: - if isinstance(child, _SKIPPED): - continue - if isinstance(child, NavigableString): - for char in str(child): - words.char(char, fmt) - continue - if not isinstance(child, Tag) or child.name in _HIDDEN: - continue - name = child.name - if name in _PARAGRAPHS: - flush() - out.extend(_display_blocks(child, fmt)) - elif name in _HEADINGS: - flush() - heading = _Words() - _inline(child, heading, fmt) - out.append((name, heading.finish())) - elif name in {"ul", "ol"}: - flush() - listed = _list(child, fmt) - if listed: - out.append(listed) - elif name == "blockquote": - flush() - out.append(("quote", _display_blocks(child, fmt))) - elif name == "pre": - flush() - out.append(("pre", child.get_text().strip("\n"))) - elif name == "hr": - flush() - out.append(("hr",)) - elif name == "table": - flush() - rows = [] - for row in child.find_all("tr"): - cells = [] - for cell in row.find_all(["td", "th"], recursive=False): - inner = _Words() - _inline(cell, inner, fmt) - cells.append((cell.name == "th", inner.finish())) - rows.append(tuple(cells)) - out.append(("table", tuple(rows))) - elif name == "br": - words.line_break() - elif name == "img": - words.token(_image(child, fmt)) - elif _has_block(child): - flush() - out.extend(_display_blocks(child, _with_format(fmt, name))) - elif name == "a" and child.get("href"): - _inline(child, words, (fmt[0], fmt[1], fmt[2], str(child.get("href")))) - else: - _inline(child, words, _with_format(fmt, name)) - flush() - return tuple(out) - - -def displayed(html: str) -> tuple[object, ...]: - """Return what a browser displays of ``html``, as comparable data.""" - - return _display_blocks(BeautifulSoup(html, "html.parser")) - - -def assert_survives(html: str) -> str: - """Assert the round-trip property for one body, and return its Markdown.""" - - markdown = read(html) - rendered = render(markdown) - - assert displayed(rendered) == displayed(html), ( - f"the Markdown does not display what the HTML did\n" - f" html: {html!r}\n markdown: {markdown!r}\n rendered: {rendered!r}" - ) - assert read(rendered) == markdown, ( - f"the Markdown is not a fixed point\n markdown: {markdown!r}\n" - f" again: {read(rendered)!r}" - ) - return markdown - - -# --------------------------------------------------------------------------- -# Ordinary prose is left alone -# --------------------------------------------------------------------------- - - -@pytest.mark.parametrize( - ("html", "expected"), - [ - pytest.param( - "<p>Voir fichier_de_test_v2.xlsx et mon_fichier_final.docx</p>", - "Voir fichier_de_test_v2.xlsx et mon_fichier_final.docx", - id="file-names", - ), - pytest.param( - r"<p>Chemin C:\Temp\logs et C:\Users\Admin\Documents</p>", - r"Chemin C:\Temp\logs et C:\Users\Admin\Documents", - id="windows-paths", - ), - pytest.param( - "<p>Le ticket # 3 et le #4521 sont liés, C# aussi</p>", - "Le ticket # 3 et le #4521 sont liés, C# aussi", - id="mid-sentence-hash", - ), - pytest.param( - "<p>Porte-monnaie - un tiret - et -- deux</p>", - "Porte-monnaie - un tiret - et -- deux", - id="dashes-and-hyphens", - ), - pytest.param( - "<p>Service R&D, bâtiment A & B</p>", - "Service R&D, bâtiment A & B", - id="ampersands", - ), - pytest.param( - "<p>5 * 3 = 15 et prix 5*3 et note * importante</p>", - "5 * 3 = 15 et prix 5*3 et note * importante", - id="lone-asterisks", - ), - pytest.param( - "<p>[INFO] tâche [1] terminée (voir note)</p>", - "[INFO] tâche [1] terminée (voir note)", - id="brackets", - ), - pytest.param( - "<p>si a < b et x <= y alors 2 < 3</p>", - "si a < b et x <= y alors 2 < 3", - id="less-than", - ), - pytest.param( - "<p>2 + 2 = 4, +33 6 12 34 56 78, 1) un, 3.14</p>", - "2 + 2 = 4, +33 6 12 34 56 78, 1) un, 3.14", - id="numbers", - ), - pytest.param( - "<p>Cordialement,<br>Jean Dupont<br>--<br>Service IT</p>", - "Cordialement, \nJean Dupont \n-- \nService IT", - id="signature-dashes-on-a-third-line", - ), - pytest.param( - "<p>Voir https://example.org/doc?a=1&b=2 ou support@example.org</p>", - "Voir https://example.org/doc?a=1&b=2 ou support@example.org", - id="bare-url-and-address", - ), - pytest.param( - "<p>Pourquoi ? Parce que ! 100 % a/b a=b ~5 minutes</p>", - "Pourquoi ? Parce que ! 100 % a/b a=b ~5 minutes", - id="punctuation", - ), - pytest.param( - "<p>l`imprimante et la variable user_id et _temp</p>", - "l`imprimante et la variable user_id et _temp", - id="unpaired-backtick-and-underscore", - ), - pytest.param( - "<p>ps aux | grep java</p>", - "ps aux | grep java", - id="pipe-outside-a-table", - ), - ], -) -def test_ordinary_prose_carries_no_escape(html: str, expected: str) -> None: - """Minimal means none of these grows a backslash or a reference. - - Each of these characters *can* be Markdown syntax, and none of them is - here, so python-markdown already renders every one of them literally. - Escaping them anyway would be harmless to the rendering and a nuisance to - everyone who reads the Markdown -- and a change for every existing body. - """ - - assert read(html) == expected - assert_survives(html) - - -# --------------------------------------------------------------------------- -# The measured misreadings, one by one -# --------------------------------------------------------------------------- - - -@pytest.mark.parametrize( - ("html", "expected"), - [ - # glpi_python_client's reviewers' measurements through its reader into - # python-markdown, before this change. - pytest.param( - r"<p>Accès au partage \\serveur\compta\2026 refusé</p>", - r"Accès au partage \\\serveur\compta\2026 refusé", - id="unc-path-lost-a-backslash", - ), - pytest.param( - r"<p>Dossier C:\_temp\logs</p>", - r"Dossier C:\\_temp\logs", - id="backslash-underscore-lost-the-backslash", - ), - pytest.param( - "<p>Fichier mon_fichier_final.docx et __init__</p>", - r"Fichier mon_fichier_final.docx et \_\_init\_\_", - id="dunder-became-bold", - ), - pytest.param( - "<p>Nom : ______ Prénom : ______</p>", - r"Nom : \_\_\_\_\_\_ Prénom : \_\_\_\_\_\_", - id="form-blanks-became-emphasis", - ), - pytest.param( - "<p>Merci<br>-----------<br>Jean Dupont</p>", - "Merci \n\\----------- \nJean Dupont", - id="dash-line-made-a-heading", - ), - pytest.param( - "<p>* point un<br>* point deux</p>", - "\\* point un \n* point deux", - id="star-lines-made-a-list", - ), - pytest.param( - "<p>Le 30/09, Jean a écrit :<br>> merci<br>> cordialement</p>", - "Le 30/09, Jean a écrit : \n\\> merci \n\\> cordialement", - id="quoted-reply-made-a-blockquote", - ), - pytest.param( - "<p>voir la note [1]</p><p>[1]: https://example.org/note</p>", - "voir la note [1]\n\n\\[1]: https://example.org/note", - id="footnote-line-was-consumed", - ), - pytest.param( - "<p># pas un titre</p>", r"\# pas un titre", id="hash-made-a-heading" - ), - pytest.param( - "<p>#4521 est un doublon</p>", - r"\#4521 est un doublon", - id="ticket-number-made-a-heading", - ), - pytest.param( - "<p>2026. Une annee</p>", r"2026\. Une annee", id="year-made-a-list" - ), - pytest.param( - "<table><tr><th>Commande</th></tr>" - "<tr><td>ps aux | grep java</td></tr></table>", - "| Commande |\n| --- |\n| ps aux \\| grep java |", - id="pipe-in-a-cell-dropped-the-rest", - ), - ], -) -def test_the_measured_misreadings_read_back_as_their_text( - html: str, expected: str -) -> None: - """Every shape the reviewers measured, now spelled so it survives.""" - - assert read(html) == expected - assert_survives(html) - - -@pytest.mark.parametrize( - ("html", "expected"), - [ - # Backslashes: escaped exactly when the next character is one - # python-markdown would take as escaped. - pytest.param( - r"<p>D:\logs\.cache et \\srv\share\[archive] et HKLM\SOFTWARE\#1</p>", - r"D:\logs\\.cache et \\\srv\share\\[archive] et HKLM\SOFTWARE\\#1", - id="backslash-before-punctuation", - ), - pytest.param( - r"<p>fin de ligne \<br>suite</p>", - "fin de ligne \\ \nsuite", - id="backslash-before-a-break", - ), - pytest.param( - r"<h2>Chemin C:\</h2>", - r"## Chemin C:\\", - id="backslash-ending-a-heading", - ), - # Emphasis delimiters. - pytest.param("<p>_______</p>", r"\_" * 7, id="seven-underscores"), - pytest.param( - "<p>Code : _______ fin</p>", - "Code : " + r"\_" * 7 + " fin", - id="seven-underscores-mid-sentence", - ), - pytest.param( - "<p>a*b*c et 5*3 <strong>gras</strong></p>", - r"a\*b\*c et 5\*3 **gras**", - id="asterisks-beside-real-emphasis", - ), - pytest.param( - "<p><strong>x</strong>* suite</p>", - r"**x**\* suite", - id="asterisk-touching-a-delimiter", - ), - # Block syntax at the start of a line. - pytest.param("<p>+ un<br>+ deux</p>", "\\+ un \n+ deux", id="plus-list"), - pytest.param("<p>- pas une liste</p>", r"\- pas une liste", id="dash-list"), - pytest.param( - "<p>Bonjour<br>#4521 doublon</p>", - "Bonjour \n\\#4521 doublon", - id="hash-after-a-break", - ), - pytest.param("<p>Titre<br>=====</p>", "Titre \n=====", id="setext-equals"), - pytest.param( - "<p>---</p><p>signature</p>", - "\\---\n\nsignature", - id="rule-of-dashes", - ), - pytest.param("<p>***</p>", r"\*\*\*", id="rule-of-asterisks"), - pytest.param("<ul><li>--</li></ul>", r"- \--", id="bullet-completes-a-rule"), - pytest.param("<ul><li>___</li></ul>", r"- \_\_\_", id="rule-in-a-list-item"), - pytest.param( - "<ul><li><ul><li>-</li></ul></li></ul>", - r"- - \-", - id="two-bullets-complete-a-rule", - ), - pytest.param( - "<ol><li><ul><li>--</li></ul></li></ol>", - r"1. - \--", - id="rule-inside-a-numbered-item", - ), - pytest.param( - "<ul><li><p>--</p><p>suite</p></li></ul>", - "- \\--\n\n suite", - id="rule-in-an-item-paragraph", - ), - pytest.param("<h1>C#</h1>", r"# C\#", id="heading-trailing-hash"), - pytest.param("<h2>Titre ##</h2>", r"## Titre \#\#", id="heading-closing-run"), - pytest.param( - "<p>```<br>code<br>```</p>", - "\\`\\`\\` \ncode \n\\`\\`\\`", - id="backtick-fence", - ), - pytest.param( - "<p>~~~<br>code<br>~~~</p>", - "~~~ \ncode \n~~~", - id="tilde-fence", - ), - pytest.param( - "<p>a | b<br>--- | ---</p>", - "a | b \n--- \\| ---", - id="table-separator", - ), - # Link and image syntax typed as text. - pytest.param( - "<p>[x](https://example.org/y) et ![x](https://example.org/y.png)</p>", - r"\[x](https://example.org/y) et !\[x](https://example.org/y.png)", - id="link-and-image-syntax", - ), - pytest.param( - '<p>Attention!<a href="https://example.org/u">voir</a></p>', - r"Attention\![voir](https://example.org/u)", - id="bang-before-a-link", - ), - pytest.param( - '<p><a href="https://example.org/u">rapport [final].pdf</a></p>', - "[rapport [final].pdf](https://example.org/u)", - id="balanced-brackets-in-link-text", - ), - pytest.param( - '<p><a href="https://example.org/u">a]b</a></p>', - r"[a\]b](https://example.org/u)", - id="unbalanced-bracket-in-link-text", - ), - pytest.param( - '<p><a href="https://example.org/u">voir [x](y)</a></p>', - r"[voir \[x\](y)](https://example.org/u)", - id="link-syntax-in-link-text", - ), - pytest.param( - '<p>[a <a href="https://example.org/u">[b</a> ](https://example.org/x)</p>', - r"\[a [\[b](https://example.org/u) ](https://example.org/x)", - id="escaped-bracket-in-link-text-is-not-counted", - ), - pytest.param( - '<p><img src="https://example.org/c.png" alt="capture [1].png"></p>', - "![capture [1].png](https://example.org/c.png)", - id="balanced-brackets-in-alt-text", - ), - pytest.param( - '<p><img src="https://example.org/c.png" alt="a]b"></p>', - r"![a\]b](https://example.org/c.png)", - id="unbalanced-bracket-in-alt-text", - ), - # Raw HTML and references typed as text. - pytest.param( - "<p>appuyer sur <Entrée> puis valider</p>", - "appuyer sur <Entrée> puis valider", - id="angle-bracketed-word", - ), - pytest.param( - "<p>if x<y then z>0</p>", - "if x<y then z>0", - id="comparison-that-looks-like-a-tag", - ), - pytest.param( - "<p><https://example.org/x> et <support@example.org></p>", - "<https://example.org/x> et <support@example.org>", - id="autolink-syntax", - ), - pytest.param( - '<p><3<img src="https://example.org/i.png" alt="a b">@c></p>', - "<3![a b](https://example.org/i.png)@c>", - id="address-running-through-an-image", - ), - pytest.param( - '<p><a href="https://example.org/u"><<a@c></a></p>', - "[<<a@c>](https://example.org/u)", - id="address-in-link-text", - ), - pytest.param( - '<p><3<a href="https://example.org">https://example.org</a>@c></p>', - "<3<https://example.org>@c>", - id="address-running-through-an-autolink", - ), - pytest.param( - '<p><a <a href="https://example.org/T_(x)">wiki</a>b@c></p>', - "<a [wiki](https://example.org/T_(x))b@c>", - id="link-target-holding-parentheses", - ), - pytest.param( - "<p><!-- note --></p>", "<!-- note -->", id="comment-syntax" - ), - pytest.param( - "<p>&amp; &lt; &#65; &#4521 &copy; &copy</p>", - "&amp; &lt; &#65; &#4521 &copy; ©", - id="character-references", - ), - # Code spans typed as text. - pytest.param( - "<p>a`b`c et <code>x</code> puis `</p>", - r"a\`b\`c et `x` puis `", - id="backtick-pair", - ), - pytest.param( - "<p><code>x</code>`y</p>", - r"`x`\`y", - id="backtick-touching-a-code-span", - ), - pytest.param( - "<p>`<code>x</code> y</p>", - r"\``x` y", - id="backtick-opening-onto-a-code-span", - ), - pytest.param( - "<table><tr><th>a</th><th>b</th></tr>" - "<tr><td>x`y</td><td>`z</td></tr></table>", - "| a | b |\n| --- | --- |\n| x\\`y | \\`z |", - id="backticks-pair-across-cells", - ), - ], -) -def test_literal_text_is_escaped_where_python_markdown_would_read_it( - html: str, expected: str -) -> None: - """One case per construct, each with the escape it needs and no other.""" - - assert read(html) == expected - assert_survives(html) - - -# --------------------------------------------------------------------------- -# Structure that used to be lost on the way through -# --------------------------------------------------------------------------- - - -@pytest.mark.parametrize( - ("html", "expected"), - [ - pytest.param( - "<ul><li>Réseau<ul><li>switch 3</li><li>borne wifi</li></ul></li>" - "<li>Imprimante</li></ul>", - "- Réseau\n - switch 3\n - borne wifi\n- Imprimante", - id="ul-in-ul", - ), - pytest.param( - "<ol><li>Arreter</li><li>Sauvegarder<ol><li>la base</li>" - "<li>les fichiers</li></ol></li><li>Redemarrer</li></ol>", - "1. Arreter\n2. Sauvegarder\n 1. la base\n 2. les fichiers\n" - "3. Redemarrer", - id="ol-in-ol-keeps-its-numbering", - ), - pytest.param( - "<ol><li>un<ul><li>a</li></ul></li><li>deux</li></ol>", - "1. un\n - a\n2. deux", - id="ul-in-ol", - ), - pytest.param( - "<ul><li>a<ul><li>b<ul><li>c<ul><li>d</li></ul></li></ul></li></ul>" - "</li></ul>", - "- a\n - b\n - c\n - d", - id="four-levels", - ), - pytest.param( - '<p>intro</p><ol start="3"><li>trois</li><li>quatre</li></ol>', - "intro\n\n3. trois\n4. quatre", - id="ordered-list-starting-at-three", - ), - pytest.param( - "<ul><li>5*3<ul><li>2*4</li></ul></li></ul>", - "- 5*3\n - 2*4", - id="asterisks-an-item-and-its-nested-item-cannot-pair", - ), - pytest.param( - "<ul><li><p>point un</p><p>suite du point</p></li><li>point deux</li></ul>", - "- point un\n\n suite du point\n\n- point deux", - id="item-with-two-paragraphs", - ), - ], -) -def test_nested_lists_nest_and_keep_their_numbers(html: str, expected: str) -> None: - """python-markdown nests only at four spaces; the reader indents by four. - - ``markdownify`` indented a continuation by its bullet's width -- two for - ``- ``, three for ``1. `` -- so the first write flattened a nested list - and renumbered a nested ordered one, one level per pass. - """ - - assert read(html) == expected - assert_survives(html) - - -@pytest.mark.parametrize( - ("html", "expected"), - [ - pytest.param( - "<ul><li>Réseau<ul><li>switch 3</li></ul>à vérifier</li></ul>", - "- Réseau\n\n - switch 3\n\n à vérifier", - id="text-after-a-nested-list", - ), - pytest.param( - "<ul><li>Réseau<ul><li>switch 3</li></ul><blockquote>cité</blockquote>" - "</li></ul>", - "- Réseau\n\n - switch 3\n\n > cité", - id="quote-after-a-nested-list", - ), - pytest.param( - "<ul><li>Réponse :<blockquote>cité</blockquote></li></ul>", - "- Réponse :\n\n > cité", - id="quote-after-an-items-text", - ), - pytest.param( - "<ul><li><blockquote>cité<br>suite</blockquote></li></ul>", - "- > cité \n> suite", - id="quote-opening-an-item", - ), - pytest.param( - "<ul><li><ul><li>un</li><li>deux</li></ul></li></ul>", - "- - un\n - deux", - id="item-opening-with-a-list", - ), - pytest.param( - "<ul><li><ul><li>un<ul><li>a</li></ul></li></ul></li></ul>", - "- - un\n\n - a", - id="list-under-an-item-sharing-its-line", - ), - pytest.param( - "<ul><li><ul><li><ul><li>un</li><li>deux</li></ul></li></ul></li>" - "<li>trois</li></ul>", - "- - - un\n\n - deux\n\n- trois", - id="three-bullets-on-one-line", - ), - ], -) -def test_blocks_inside_a_list_item_stay_in_it(html: str, expected: str) -> None: - """python-markdown only nests a block where its first pass lets it. - - That pass never detabs an item's first block, and reads ``>`` only three - spaces in at most, so what follows a list item's text or its nested list - needs a blank line before it -- ``markdownify`` gave none after a nested - list, and the item's text ran into the list's last item -- and a quote - opening an item needs its later lines unindented. On a line already - carrying two bullets, anything eight spaces in is taken as that line's - continuation, so what follows it starts a block of its own. - """ - - assert read(html) == expected - assert_survives(html) - - -@pytest.mark.parametrize( - ("html", "expected"), - [ - pytest.param( - '<p><a name="_MailEndCompose">Bonjour</a> Jean</p>', - "Bonjour Jean", - id="anchor-without-a-target", - ), - pytest.param( - '<p><a href="https://example.org/u"></a>texte</p>', "texte", id="empty-link" - ), - pytest.param("<p><strong></strong>texte</p>", "texte", id="empty-emphasis"), - pytest.param( - '<p><a href="https://example.org/u">un<br>deux</a></p>', - "[un \ndeux](https://example.org/u)", - id="break-in-link-text", - ), - pytest.param("<p>a<center>b</center>c</p>", "a\n\nb\n\nc", id="center-block"), - pytest.param( - "<table><tr><th><h3>titre</h3></th></tr><tr><td>x</td></tr></table>", - "| titre |\n| --- |\n| x |", - id="heading-in-a-cell", - ), - pytest.param( - "<h2><blockquote>cité</blockquote></h2>", "## cité", id="quote-in-a-heading" - ), - pytest.param( - "<table><tr><th>a</th></tr><tr><td>un<br>deux</td></tr></table>", - "| a |\n| --- |\n| un deux |", - id="break-in-a-cell", - ), - pytest.param( - "<h2>Titre<br>suite</h2>", "## Titre suite", id="break-in-a-heading" - ), - pytest.param( - '<p><img src="https://example.org/i.png" alt="ligne 1\n\n ligne 2"></p>', - "![ligne 1 ligne 2](https://example.org/i.png)", - id="alt-text-over-several-lines", - ), - pytest.param( - "<ul><li><script>x()</script><pre>code</pre></li></ul>", - "- code", - id="code-opening-an-item-after-a-script", - ), - pytest.param( - "<ul><li><!-- note --><pre>code</pre></li></ul>", - "- code", - id="code-opening-an-item-after-a-comment", - ), - pytest.param( - "<ul><li><ul></ul><pre>code</pre></li></ul>", - "- code", - id="code-opening-an-item-after-an-empty-list", - ), - pytest.param( - '<ul><li><img src="https://example.org/i.png" alt="i"><pre>code</pre>' - "</li></ul>", - "- ![i](https://example.org/i.png)\n\n code", - id="code-after-an-image", - ), - ], -) -def test_shapes_with_a_rule_of_their_own(html: str, expected: str) -> None: - """The converter's special cases, each on the shape that reaches it. - - What displays nothing -- a comment, a script, an empty list -- does not - count as content before a code block, so the code still opens its item. - """ - - assert read(html) == expected - - -def test_an_ordered_list_keeps_its_start_when_rendered() -> None: - """``sane_lists`` is what keeps ``3.`` from restarting the count at 1.""" - - assert '<ol start="3">' in render("intro\n\n3. trois\n4. quatre") - - -@pytest.mark.parametrize( - ("html", "expected"), - [ - pytest.param( - "<ul><li>item<pre>#4521 code</pre></li></ul>", - "- item\n\n #4521 code", - id="in-a-list-item", - ), - pytest.param( - "<blockquote><pre>#4521 C:\\Temp\n indenté</pre></blockquote>", - "> #4521 C:\\Temp\n> indenté", - id="in-a-blockquote", - ), - pytest.param( - "<pre>ligne\n```\nfin</pre>", - "````\nligne\n```\nfin\n````", - id="holding-a-fence-line", - ), - ], -) -def test_a_preformatted_block_stays_one(html: str, expected: str) -> None: - """A fence opens only at the start of a line, so nested code is indented. - - Inside a list item or a block quote, python-markdown never sees - ``` ``` ``` at the start of a line and reads the fence as text -- the - code's own ``#4521`` then became a heading. An indented code block is the - spelling it does read there. At the top level the fence is made longer - than any fence line the code holds. - """ - - assert read(html) == expected - assert_survives(html) - - -@pytest.mark.xfail( - strict=True, - reason=( - "python-markdown runs html.parser over its whole source to find raw " - "HTML, before an indented code block is recognised, and re-emits a " - "numeric reference written without its semicolon with one: the code " - "then shows 'ᆩ'. A fence is stashed before that pass, which is " - "why a top-level <pre> is unaffected; inside a list item or a quote no " - "fence can open. Measured on 3.10.3." - ), -) -def test_a_numeric_reference_in_nested_code_gains_a_semicolon() -> None: - assert_survives("<blockquote><pre>echo &#4521</pre></blockquote>") - - -def test_a_preformatted_block_opening_a_list_item_keeps_its_text() -> None: - """The one place python-markdown can start no code block at all. - - A list item's first line is its paragraph, so the code there degrades to - its lines, each escaped as the literal text it is -- the words survive, - the preformatting does not. - """ - - html = "<ul><li><pre>#4521 code\n suite</pre></li></ul>" - - markdown = read(html) - - assert markdown == "- \\#4521 code \n suite" - assert "#4521 code" in render(markdown) - assert read(render(markdown)) == markdown - - -@pytest.mark.parametrize( - ("html", "expected"), - [ - pytest.param( - "<ul><li>item<ul><li>sous-item</li></ul><pre>#4521 code</pre></li></ul>", - "- item\n\n - sous-item\n\n \\#4521 code", - id="in-a-list-item", - ), - pytest.param( - "<blockquote><ul><li>item</li></ul><pre>#4521 code</pre></blockquote>", - "> - item\n>\n> \\#4521 code", - id="in-a-quote", - ), - pytest.param( - "<ul><li>item<ul><li>sous-item</li></ul><div><style>p{margin:0}</style>" - "</div><pre>#4521 code</pre></li></ul>", - "- item\n\n - sous-item\n\n \\#4521 code", - id="with-only-a-style-block-between", - ), - ], -) -def test_code_right_after_a_nested_list_keeps_its_text( - html: str, expected: str -) -> None: - """The other place python-markdown can start no code block. - - An indented code block right after a list in the same item or quote is - indented exactly as the list's last item's own content, and - python-markdown reads it as a paragraph of that item. The lines are - kept as literal text in the right place instead. - """ - - markdown = read(html) - - assert markdown == expected - assert "#4521 code" in render(markdown) - assert read(render(markdown)) == markdown - - -#: What the format loses, whatever the escaping does: each body with why. -#: -#: None of these is literal text misread. Each is structure python-markdown -#: has no spelling for, or spells as something else. -LOSSES = [ - pytest.param( - "<ul><li>a</li></ul><ul><li>b</li></ul>", - "python-markdown continues a list across a blank line: two lists are " - "read back as one list of loose items.", - id="two-adjacent-lists", - ), - pytest.param( - "<blockquote>a</blockquote><blockquote>b</blockquote>", - "python-markdown continues a quote across a blank line: two quotes are " - "read back as one quote of two paragraphs.", - id="two-adjacent-quotes", - ), - pytest.param( - "<p>a<br><br>b</p>", - "Markdown has no blank line inside a paragraph: two breaks in a row " - "are a paragraph break, and read back as one.", - id="two-breaks-in-a-row", - ), - pytest.param( - "<p><code>a</code><code>b</code></p>", - "'`a``b`' is one code span holding 'a``b' to python-markdown.", - id="two-adjacent-code-spans", - ), - pytest.param( - "<p><em>a <strong>b</strong> c</em></p>", - "python-markdown pairs '*a **b** c*' as three emphasis runs, and b " - "loses its bold.", - id="strong-inside-emphasis", - ), - pytest.param( - "<table><tr><td>a</td><td>b</td></tr></table>", - "A Markdown table starts with its header row, so one without gains an " - "empty header.", - id="table-without-a-header", - ), - pytest.param( - "<table><tr><th>a</th><th>b</th></tr></table>", - "python-markdown renders a header-only table with one empty body row.", - id="table-without-a-body", - ), - pytest.param( - "<table><tr><th>a</th></tr><tr><td><ul><li>x</li></ul></td></tr></table>", - "A table cell holds one line of inline Markdown: its list is written " - "as its text.", - id="list-in-a-cell", - ), - pytest.param( - "<table><tr><th>a</th></tr><tr><td><pre>x\ny</pre></td></tr></table>", - "A table cell holds one line of inline Markdown: its code block is " - "written as inline code.", - id="code-block-in-a-cell", - ), - pytest.param( - "<table><tr><th>a</th></tr><tr><td>un<br>deux</td></tr></table>", - "A table cell holds one line: its line break is written as a space.", - id="break-in-a-cell", - ), - pytest.param( - "<h2>Titre<br>suite</h2>", - "A heading is one line: its line break is written as a space.", - id="break-in-a-heading", - ), - pytest.param( - "<ul><li><pre>a</pre></li></ul>", - "An item's first block is its paragraph: the code is kept as text " - "(test_a_preformatted_block_opening_a_list_item_keeps_its_text).", - id="code-opening-a-list-item", - ), - pytest.param( - "<ul><li>a<ul><li>b</li></ul><pre>c</pre></li></ul>", - "The code would be indented as the nested item's own content: it is " - "kept as text (test_code_right_after_a_nested_list_keeps_its_text).", - id="code-after-a-nested-list", - ), -] - - -@pytest.mark.parametrize(("html", "reason"), LOSSES) -def test_what_markdown_cannot_carry(html: str, reason: str) -> None: - """The inventory of losses, each asserted to still be one. - - ``xfail(strict=True)`` is the point: a loss that stops being one fails - here, so the inventory stays true. Struck and underlined text lose their - line too, which this comparison does not see: - ``test_struck_text_keeps_its_words_and_loses_its_line``. - """ - - with pytest.raises(AssertionError): - assert_survives(html) - pytest.xfail(reason) - - -@pytest.mark.parametrize(("html", "reason"), LOSSES) -def test_what_is_lost_is_lost_once(html: str, reason: str) -> None: - """Whatever a loss costs, it costs on the first cycle and never again.""" - - again = read(render(read(html))) - - assert read(render(again)) == again, reason - - -@pytest.mark.parametrize( - "html", - [ - pytest.param( - "<style>p.MsoNormal{margin:0cm;font-size:11pt}</style><p>Bonjour</p>", - id="style", - ), - pytest.param("<p>Bonjour</p><script>track()</script>", id="script"), - pytest.param( - "<html><head><title>RE: Imprimante" - "" - "

      Bonjour

      ", - id="outlook-shaped", - ), - ], -) -def test_style_script_and_title_bodies_are_not_text(html: str) -> None: - """A browser displays none of these, so neither does the Markdown. - - ``strip=["script", "style"]`` used to be passed to ``markdownify``, and - ``strip`` skips an element's own converter -- ``convert_script`` and - ``convert_style`` return ``""`` -- so the bodies leaked into the text as - prose. ```` has no converter at all and leaked the same way. - """ - - assert read(html) == "Bonjour" - - -@pytest.mark.parametrize( - ("html", "expected"), - [ - pytest.param( - '<font color="red">URGENT</font> serveur HS', "URGENT serveur HS", id="font" - ), - pytest.param("<center>Titre</center> suite", "Titre\n\nsuite", id="center"), - pytest.param("<strike>ancien</strike> nouveau", "ancien nouveau", id="strike"), - pytest.param("<big>gros</big> texte", "gros texte", id="big"), - pytest.param("<tt>code</tt> texte", "code texte", id="tt"), - pytest.param("<nobr>sans coupure</nobr>", "sans coupure", id="nobr"), - ], -) -def test_a_body_marked_up_only_with_obsolete_elements_is_html( - html: str, expected: str -) -> None: - """Old editors still write these; without them the tags were kept as text.""" - - assert read(html) == expected - - -def test_every_obsolete_element_name_makes_a_body_html() -> None: - """The HTML standard's list of obsolete elements, all of them recognised.""" - - obsolete = ( - "acronym applet basefont bgsound big blink center dir font frame " - "frameset isindex keygen listing marquee menuitem multicol nextid nobr " - "noembed noframes plaintext rb rtc spacer strike tt xmp" - ).split() - - assert set(obsolete) <= conversion._HTML_ELEMENTS - - -@pytest.mark.parametrize( - "html", - [ - pytest.param("<p><s>ancien</s> nouveau</p>", id="s"), - pytest.param("<p><del>ancien</del> nouveau</p>", id="del"), - pytest.param("<p><strike>ancien</strike> nouveau</p>", id="strike"), - ], -) -def test_struck_text_keeps_its_words_and_loses_its_line(html: str) -> None: - """python-markdown has no strikethrough, so ``~~x~~`` would show literally. - - Keeping the words and losing the line is the honest loss: the text is - all there, and what is missing is recorded in the round-trip inventory. - """ - - assert read(html) == "ancien nouveau" - - -@pytest.mark.parametrize( - "html", ["<P>Bonjour</P>", "<BR>Bonjour", "<DIV><B>Bonjour</B></DIV>"] -) -def test_upper_case_tags_are_html(html: str) -> None: - """Element names are case-insensitive; the probe has to be too.""" - - assert "<" not in read(html) - - -@pytest.mark.parametrize( - ("value", "expected"), - [ - pytest.param(" texte ", "texte", id="plain-text"), - pytest.param(" <b>x</b> ", "**x**", id="html"), - pytest.param("<br>x<br>", "x", id="html-with-edge-breaks"), - ], -) -def test_the_result_is_stripped_on_both_paths(value: str, expected: str) -> None: - """Edge whitespace is never content, and a digest must not depend on it.""" - - assert read(value) == expected - - -@pytest.mark.parametrize( - ("html", "expected"), - [ - pytest.param( - '<h2>Titre <img src="https://example.org/i.png" alt="logo"></h2>', - "## Titre ![logo](https://example.org/i.png)", - id="heading", - ), - pytest.param( - "<table><tr><th>a</th></tr><tr><td>" - '<img src="https://example.org/i.png" alt="x"></td></tr></table>', - "| a |\n| --- |\n| ![x](https://example.org/i.png) |", - id="table-cell", - ), - ], -) -def test_an_image_in_a_heading_or_a_cell_stays_an_image( - html: str, expected: str -) -> None: - """``markdownify`` reduced these to their alt text; python-markdown needs not.""" - - assert read(html) == expected - assert_survives(html) - - -def test_a_newline_in_text_is_a_space() -> None: - """HTML displays a source newline as a space; ``nl2br`` would break the line. - - Outlook wraps its HTML source mid-sentence, so every body that came - from an e-mail used to gain line breaks where the reader saw none. - """ - - html = "<p>Bonjour,\nle serveur\nest redémarré.</p>" - - assert read(html) == "Bonjour, le serveur est redémarré." - assert_survives(html) - - -@pytest.mark.parametrize( - ("html", "expected"), - [ - pytest.param("<p>a<br> b<br> c</p>", "a \nb \nc", id="space-after-break"), - pytest.param("<p>a <br>b</p>", "a \nb", id="space-before-break"), - pytest.param("<p><strong>a<br></strong>b</p>", "**a** \nb", id="edge-break"), - ], -) -def test_a_line_break_has_one_spelling(html: str, expected: str) -> None: - """Whitespace around ``<br>`` is not displayed, so it is not kept either. - - Without this the first read of ``a<br> b`` was ``a \\n b`` and the - second ``a \\nb``: the same body, read twice, disagreeing. - """ - - assert read(html) == expected - assert_survives(html) - - -# --------------------------------------------------------------------------- -# The degraded path spells text the same way -# --------------------------------------------------------------------------- - - -def test_a_body_too_deep_to_convert_is_still_literal_safe() -> None: - """The stripped text is Markdown too, and it is escaped like the rest.""" - - html = "<div>" * 600 + r"<p>__init__ \\serveur</p><p># pas un titre</p>" - - markdown = read(html) - - assert markdown == "\\_\\_init\\_\\_ \\\\\\serveur\n\n\\# pas un titre" - assert "__init__ \\\\serveur" in render(markdown) - - -@pytest.mark.parametrize("seed", range(4)) -def test_stripped_text_renders_as_itself(seed: int) -> None: - """Any text, line by line, renders as exactly that text. - - The degraded path's contract, over generated text dense with the - characters python-markdown treats as syntax. - """ - - rng = random.Random(seed) - for _ in range(60): - lines = [_text(rng, rng.randint(1, 6)) for _ in range(rng.randint(1, 4))] - text = "\n".join(line for line in lines if line) - if not text: - continue - - soup = BeautifulSoup(render(conversion._literal_markdown(text)), "html.parser") - for line_break in soup.find_all("br"): - line_break.replace_with("\u2029") - shown = soup.get_text().replace("\n", " ").replace("\u2029", "\n") - - assert [" ".join(line.split()) for line in shown.split("\n")] == [ - " ".join(line.split()) for line in text.split("\n") - ], text - - -# --------------------------------------------------------------------------- -# The property, over realistic bodies and over a fuzzer -# --------------------------------------------------------------------------- - -#: Bodies shaped the way an HTML editor and a mail collector store them -- -#: glpi_python_client's corpus, which EasyVista memos share: nobody working -#: on this package has measured EasyVista's own editor output. -REALISTIC = [ - "<p>Bonjour,</p><p>Le PC du poste 12 ne démarre plus depuis ce matin.</p>", - "<p>Bonjour,<br>Le PC ne démarre plus.<br>Cordialement,<br>Jean</p>", - "<p>Merci de <strong>redémarrer</strong> le serveur <em>avant</em> 18h.</p>", - "<p>Étapes :</p><ul><li>ouvrir la session</li><li>lancer Outlook</li></ul>", - "<table><thead><tr><th>Poste</th><th>IP</th></tr></thead>" - "<tbody><tr><td>PC12</td><td>10.0.0.12</td></tr></tbody></table>", - "<h2>Contexte</h2><p>Migration du serveur.</p>", - '<p>Voir <a href="https://example.org/doc">https://example.org/doc</a></p>', - '<p>Voir <a title="doc" href="https://example.org/doc">la doc</a></p>', - '<p><a href="https://example.org/wiki/Test_(informatique)">wiki</a></p>', - "<p>prix 5*3 et note * importante</p>", - "<p>appuyer sur <Entrée> puis valider</p>", - "<p>Service R&D, bâtiment A & B</p>", - "<p>Montant : 12 000 €</p>", - '<p><span style="color: #e03e2d;">URGENT</span> <u>à traiter</u></p>', - '<p>Cordialement</p><p><img src="https://example.org/logo.png" alt="Logo"></p>', - "<p>Bonjour</p><blockquote><p>Message d'origine</p></blockquote>", - '<pre>Traceback (most recent call last):\n File "x.py", line 1\nError</pre>', - "<p>Lancer <code>ipconfig /all</code> puis envoyer.</p>", - "<p>Fichier mon_fichier_final.docx et __init__</p>", - "<p># pas un titre</p><p>- pas une liste</p>", - '<p>Contact <a href="mailto:support@example.org">support@example.org</a></p>', - "<div>Bonjour,</div><div><br></div><div>Le serveur est down.</div>", - "<p>* point un<br>* point deux</p>", - "<p><https://example.org/x></p>", - "<p>[INFO] tâche [1] terminée</p>", - "<p>[1]: https://example.org/note</p>", - "<p>m<sup>2</sup></p>", - r"<p>Chemin C:\Users\jdupont\Desktop</p>", - "<p>Merci 👍</p>", - "<p>Titre<br>=====</p>", - "<p>---</p><p>signature</p>", - "<p>1) un<br>2) deux</p>", - "<p><support@example.org></p>", - "<p><!-- note --></p>", - "<ul><li><p>point un</p><p>suite du point</p></li><li>point deux</li></ul>", - "<blockquote><ul><li>a</li><li>b</li></ul></blockquote>", - "<p>Nom : ______ Prénom : ______</p>", - "<p>~~pas barré~~</p>", - "<p>a | b | c</p>", - "<p>> pas une citation</p>", - "<p>+ un<br>+ deux</p>", - "<p>Cordialement,<br>Jean Dupont<br>--<br>Service IT</p>", - "<p>Merci<br>-----------<br>Jean Dupont</p>", - "<p>calcul 5 * 3 * 2 = 30</p>", - r"<p>voir \\srv\partage\__archive__\2026</p>", - "<p>Le 30/09, Jean a écrit :<br>> merci<br>> cordialement</p>", - r"<p>Accès au partage \\serveur\compta\2026 refusé</p>", - r"<p>Dossier C:\_temp\logs</p>", - "<p>voir la note [1]</p><p>[1]: https://example.org/note</p>", - "<ul><li>Réseau<ul><li>switch 3</li><li>borne wifi</li></ul></li></ul>", - "<p><b>Important</b> : voir <i>ci-dessous</i></p>", - '<p><font color="red">rouge</font> <span style="font-size:14px">texte</span></p>', - "<p>module __init__ et _x_</p>", - "<p>1) un<br>2) deux</p><pre>ligne 1\n ligne indentée</pre>", - "<p>Suite à la mise à jour, <strong>3 postes</strong> ne se connectent plus :" - "</p><ul><li>PC12 (salle 3)</li><li>PC14 – <em>poste d'accueil</em></li>" - '</ul><p>Voir <a href="https://example.org/kb/42">https://example.org/kb/42</a>' - ' et le <a href="https://example.org/kb/43" title="KB 43">KB 43</a>.</p>', -] - - -@pytest.mark.parametrize("html", REALISTIC) -def test_realistic_bodies_display_the_same_after_the_round_trip(html: str) -> None: - assert_survives(html) - - -#: Literal text snippets: every character python-markdown treats as syntax, -#: alone and in the shapes that trigger it, beside ordinary words so that -#: each lands at line starts, at word boundaries and inside words. -_SPECIAL = ( - "\\ \\\\ \\* \\_ \\[ \\# \\. \\` \\\\serveur * ** *** _ __ ___ ______ __init__ " - "_x_ ` `` ``` ~ ~~ ~~~ [ ] [x] [x](y) ![x](y) [1]: [1]:https://example.org/n " - "( ) (y) ! # ## #4521 > - -- --- + . 1. 2026. = == === | a|b --- < <b> </b> " - "<Enter> <!-- --> <https://example.org> <a@b.c> <3 <= & & < A " - "A ᆩ © © &D; { } : \" '" -).split(" ") -_WORDS = ( - "alpha beta snake_case C:\\Temp R&D x mot 5*3 a_b é fichier_de_test_v2.xlsx" -).split(" ") - - -#: Code content. A numeric reference with no semicolon is left out: inside an -#: indented code block python-markdown's raw-HTML pass adds the semicolon -#: (``test_a_numeric_reference_in_nested_code_gains_a_semicolon``). -_CODE_SPECIAL = [special for special in _SPECIAL if not special.startswith("&#")] - - -def _text(rng: random.Random, pieces: int, specials: list[str] = _SPECIAL) -> str: - out = [] - for _ in range(pieces): - out.append(rng.choice(specials) if rng.random() < 0.45 else rng.choice(_WORDS)) - out.append(rng.choice(["", " ", " ", " "])) - return "".join(out).strip() - - -def _code_text(rng: random.Random, pieces: int) -> str: - return escape(_text(rng, pieces, _CODE_SPECIAL), quote=False) - - -def _inline_html(rng: random.Random, formatted: bool = False, depth: int = 0) -> str: - """Inline HTML that Markdown can express. - - No emphasis inside emphasis, no link inside a link, and a word on either - side of every code element: python-markdown mis-pairs nested ``*`` runs - across constructs and merges two adjacent code spans, neither of which - involves literal text -- they are recorded in the round-trip inventory. - """ - - parts = [] - for _ in range(rng.randint(1, 4)): - kind = rng.random() - if kind < 0.55 or depth > 2: - parts.append(escape(_text(rng, rng.randint(1, 4)), quote=False)) - elif kind < 0.65 and not formatted: - parts.append(f" <strong>{_inline_html(rng, True, depth + 1)}</strong> ") - elif kind < 0.75 and not formatted: - parts.append(f" <em>{_inline_html(rng, True, depth + 1)}</em> ") - elif kind < 0.82 and depth == 0: - inner = _inline_html(rng, formatted, depth + 1) - parts.append( - f'<a href="https://example.org/{rng.randint(1, 9)}">{inner}</a>' - ) - elif kind < 0.88: - code = escape(rng.choice(_WORDS + _SPECIAL[:20]), quote=False) - parts.append(f" mot <code>{code}</code> mot ") - elif kind < 0.94: - alt = escape(_text(rng, rng.randint(0, 2)), quote=True) - parts.append( - f'<img src="https://example.org/{rng.randint(1, 9)}.png" alt="{alt}">' - ) - else: - parts.append(f"<span>{_inline_html(rng, formatted, depth + 1)}</span>") - parts.append(rng.choice(["", " ", " "])) - return "".join(parts).strip() or "mot" - - -def _paragraph(rng: random.Random) -> str: - return "<br>".join(_inline_html(rng) for _ in range(rng.randint(1, 3))) - - -def _list_html(rng: random.Random, depth: int) -> str: - """A list whose items hold text, and sometimes code, a list, and more text. - - The code goes before an item's nested list, never right after it: - python-markdown has no way to write a code block there - (``test_code_right_after_a_nested_list_keeps_its_text``). - """ - - tag = rng.choice(["ul", "ol"]) - items = [] - for _ in range(rng.randint(1, 3)): - shape = rng.random() - if depth < 3 and shape < 0.1: - # An item that opens with a list: one line holds both bullets. - items.append(f"<li>{_list_html(rng, depth + 1)}</li>") - continue - inner = f"<p>{_paragraph(rng)}</p>" if shape < 0.3 else _paragraph(rng) - if rng.random() < 0.1: - inner += f"<pre>{_code_text(rng, 3)}</pre>" - if depth < 3 and rng.random() < 0.3: - inner += _list_html(rng, depth + 1) - after = rng.random() - if after < 0.15: - inner += _inline_html(rng) - elif after < 0.25: - inner += f"<p>{_paragraph(rng)}</p>" - elif after < 0.3: - inner += f"<blockquote><p>{_paragraph(rng)}</p></blockquote>" - items.append(f"<li>{inner}</li>") - return f"<{tag}>{''.join(items)}</{tag}>" - - -def _block(rng: random.Random, depth: int = 0) -> str: - kind = rng.random() - if kind < 0.35 or depth > 1: - return f"<p>{_paragraph(rng)}</p>" - if kind < 0.45: - level = rng.randint(1, 3) - return f"<h{level}>{_inline_html(rng)}</h{level}>" - if kind < 0.60: - return _list_html(rng, depth + 1) - if kind < 0.70: - return f"<blockquote>{_block(rng, depth + 1)}</blockquote>" - if kind < 0.80: - head = "".join(f"<th>{_inline_html(rng)}</th>" for _ in range(2)) - rows = "".join( - "<tr>" - + "".join(f"<td>{_inline_html(rng)}</td>" for _ in range(2)) - + "</tr>" - for _ in range(rng.randint(1, 2)) - ) - return f"<table><tr>{head}</tr>{rows}</table>" - if kind < 0.87: - return f"<pre>{_code_text(rng, 4)}</pre>" - return f"<div>{_paragraph(rng)}</div>" - - -def _document(rng: random.Random) -> str: - """A body of one to four blocks. - - python-markdown merges two adjacent lists of one type, or two adjacent - block quotes, into one -- a limitation of the format -- so a paragraph - separates them. - """ - - blocks: list[str] = [] - for _ in range(rng.randint(1, 4)): - block = _block(rng) - mergeable = block[:4] in {"<ul>", "<ol>", "<blo"} - if blocks and mergeable and block[:4] == blocks[-1][:4]: - blocks.append("<p>mot</p>") - blocks.append(block) - return "".join(blocks) - - -@pytest.mark.parametrize("seed", range(8)) -def test_generated_bodies_display_the_same_after_the_round_trip(seed: int) -> None: - """The property over a seeded fuzzer, 50 bodies a seed.""" - - rng = random.Random(seed) - for _ in range(50): - assert_survives(_document(rng)) - - -@pytest.mark.parametrize( - "html", - [ - pytest.param("<p>" + "[a " * 20_000 + "</p>", id="unclosed-brackets"), - pytest.param("<p>" + "_a " * 20_000 + "</p>", id="underscores-opening-words"), - pytest.param("<p>" + "``` " * 20_000 + "</p>", id="backtick-runs"), - pytest.param("<ol>" + "<li>x</li>" * 20_000 + "</ol>", id="long-numbered-list"), - ], -) -def test_a_body_dense_with_syntax_converts_in_linear_time(html: str) -> None: - """Every pass is linear, so a pathological body costs what its size does. - - Measured before the passes were made linear: 20,000 of these took from - 8 to 117 seconds each, where they take well under one now. The budget - is generous on purpose -- this guards the complexity, not the speed. - """ - - started = time.perf_counter() - read(html) - - assert time.perf_counter() - started < 10 - - -# --------------------------------------------------------------------------- -# What the escaping rests on -# --------------------------------------------------------------------------- - - -def test_python_markdown_undoes_every_backslash_the_reader_writes() -> None: - """A backslash escape is only safe if python-markdown removes it again. - - Measured on 3.10.3 with the four extensions: ``\\``, backtick, ``*``, - ``_``, ``{``, ``}``, ``[``, ``]``, ``(``, ``)``, ``>``, ``#``, ``+``, - ``-``, ``.``, ``!`` and ``|``. ``=`` and ``~`` are not among them, which - is why those two are spelled as character references instead. - """ - - escapable = set(Markdown(extensions=conversion._MARKDOWN_EXTENSIONS).ESCAPED_CHARS) - backslashed = set(conversion._LITERAL) - set(conversion._SPELLED_OUT) | {"\\"} - - assert backslashed <= escapable - assert not {"=", "~", "<", "&"} & escapable - - -def test_a_body_holding_the_private_stand_ins_converts_like_any_other() -> None: - """Literal text is carried in private-use characters until it is spelled. - - A body that already holds one of them -- an icon font maps symbols - there -- must not be confused with them, so the converter picks a block - the body does not use. - """ - - first_block = [chr(0xF0000 + offset) for offset in range(32)] - html = "<p>" + "".join(first_block) + " __init__ et #4521</p>" - - markdown = read(html) - - assert "".join(first_block) in markdown - assert markdown.endswith("\\_\\_init\\_\\_ et #4521") diff --git a/easyvista_python_client/content/tests/test_properties.py b/easyvista_python_client/content/tests/test_properties.py new file mode 100644 index 0000000..5d35691 --- /dev/null +++ b/easyvista_python_client/content/tests/test_properties.py @@ -0,0 +1,786 @@ +"""Seeded property tests: generated bodies keep their display, words and Markdown. + +Each family puts the characters and shapes one of the converter's fixes is +about at every position the generator can reach. For every body, the +converter owes three things: + +* **display** -- ``to_transport(from_transport(html))`` displays what + ``html`` displays, by :func:`.display.displayed`, compared loosely on URLs + (cmark-gfm percent-encodes a target it renders, and a browser follows + either spelling to the same place); +* **fixed point** -- reading back what the Markdown renders gives the same + Markdown; and +* **words** -- the words displayed, in order, are the same + (:func:`.display.text_words`, which does not depend on the oracle). + +The generators are the synthetic ones the converter's fixes were verified +with on 2026-10-02, ported here: the image-alt, tilde and ``!`` attack +families, inline code holding line breaks, tables nested in a cell, a +heading or a link (their cells sometimes holding a ``<p>`` or a ``<div>``), +underline and highlight in every holder, and this package's earlier +document generator. Their words are invented or generic; every URL is +under ``example.org``. Each section names the fixes it exercises, numbered +as in :mod:`.test_fixes`. At the sizes below, GLPI 917f030 failed 39 to +100 percent of each test's bodies, and this converter none (measured +2026-10-02, CPython 3.12.11). + +What Markdown cannot spell is either expected as the converter writes it or +left out of the generators, each marked where it applies: + +* expected: a line break inside a heading -- an ATX heading is one line, so + the converter writes a space; +* expected: a table inside a link -- Markdown has no table there, so the + converter writes the cells as the link's text + (:func:`.display.one_line`); +* expected: underline, highlight and inserted text, which the display + oracle cannot see, kept as raw tags around the same words + (:func:`_raw_spans`); +* left out: a definition list, a block inside bold, bold inside bold, and + an image title holding a backslash -- the converter keeps their words + but not their display or, for the title, its backslash. + +``docs/content.rst`` lists the shapes no generator here reaches that the +converter is known to get wrong, such as a table whose rows have more +cells than its first. + +The sizes are a third to a tenth of the 500 to 5,400 bodies a family the +fixes were verified with, so that the whole content suite runs in about 30 +seconds (measured 2026-10-02 on CPython 3.12.11 and 3.14.6 on a developer +machine). +""" + +from __future__ import annotations + +import html as html_module +import random +import re +import string +from collections.abc import Callable + +import pytest +from bs4 import BeautifulSoup + +from easyvista_python_client.content.conversion import EasyvistaContentConverter +from easyvista_python_client.content.tests.display import ( + displayed, + loose, + one_line, + text_words, +) + +read = EasyvistaContentConverter.from_transport +render = EasyvistaContentConverter.to_transport + +# --------------------------------------------------------------------------- +# The property +# --------------------------------------------------------------------------- + +_HEADING = re.compile(r"<h([1-6])>.*?</h\1>", re.DOTALL) + + +def _heading_breaks_as_spaces(html: str) -> str: + """``html`` with every ``<br>`` inside a heading displayed as a space. + + An ATX heading is one line, so Markdown has no line break inside one; + the converter writes a space, and this is the display it owes. + """ + + return _HEADING.sub(lambda heading: heading[0].replace("<br>", " "), html) + + +def _problems(html: str, expected: Callable[[str], object] = displayed) -> list[str]: + """Return what the converter got wrong about ``html``: empty when nothing.""" + + markdown = read(html) + rendered = render(markdown) + problems = [] + if loose(displayed(rendered)) != loose(expected(_heading_breaks_as_spaces(html))): + problems.append("display") + if read(rendered) != markdown: + problems.append("fixed point") + if text_words(rendered) != text_words(html): + problems.append("words") + if "ev-cell" in markdown or "data-ev-" in markdown: + problems.append("a private marker leaked") + return problems + + +def assert_property( + bodies: list[str], expected: Callable[[str], object] = displayed +) -> None: + failures = [] + for html in bodies: + problems = _problems(html, expected) + if problems: + failures.append(f"{', '.join(problems)}: {html!r} -> {read(html)!r}") + assert not failures, ( + f"{len(failures)} of {len(bodies)} bodies failed; the first ones:\n " + + "\n ".join(failures[:5]) + ) + + +# --------------------------------------------------------------------------- +# Image alt text, '~~~' at a line start, '!' before a link (fixes 1 to 5) +# --------------------------------------------------------------------------- + +_PUNCTUATION = list(string.punctuation) +_ATOMS = ["a", "b", "x1", "C:", "Temp", "mot", "été", "—"] +_SYNTAX = [ + "~~~", "!!", "![", "](", "^", "\\", "\\\\", "a_b", "*x*", "[x]", "(y)", "<x>", + "&", "©", " ", "`x`", "``", "#", "1.", "-", "+", "|", "\\_", "\\*", + "\\[", "^^", "!^", +] # fmt: skip + + +def _piece(rng: random.Random) -> str: + draw = rng.random() + if draw < 0.35: + return rng.choice(_PUNCTUATION) + if draw < 0.65: + return rng.choice(_ATOMS) + return rng.choice(_SYNTAX) + + +def _raw_text(rng: random.Random, low: int = 1, high: int = 6) -> str: + separator = rng.choice(["", " ", ""]) + return separator.join(_piece(rng) for _ in range(rng.randint(low, high))) + + +def _text(rng: random.Random, low: int = 1, high: int = 6) -> str: + return html_module.escape(_raw_text(rng, low, high), quote=False) + + +def _image(rng: random.Random, alt: str | None = None) -> str: + """An image. Never with a title: mdformat writes a title raw, so a + backslash in one is lost -- a known limit, out of these families.""" + + alt = _raw_text(rng) if alt is None else alt + alt = html_module.escape(alt, quote=True) + return f'<img src="https://example.org/i{rng.randint(0, 9)}.png" alt="{alt}">' + + +def _link(rng: random.Random, inner: str | None = None) -> str: + inner = _text(rng) if inner is None else inner + return f'<a href="https://example.org/p{rng.randint(0, 9)}">{inner}</a>' + + +_CONTEXTS = [ + "<p>{x}</p>", + "<p>mot {x} mot</p>", + "<p>mot{x}mot</p>", + "<ul><li>{x}</li></ul>", + "<ol><li>a<ul><li>b<ol><li>{x}</li></ol></li></ul></li></ol>", + "<blockquote><p>{x}</p></blockquote>", + "<blockquote><ul><li>{x}</li></ul></blockquote>", + "<ul><li><blockquote><p>{x}</p></blockquote></li></ul>", + "<h2>{x}</h2>", + "<table><tr><th>h</th></tr><tr><td>{x}</td></tr></table>", + "<p><b>{x}</b></p>", + "<p><em>{x}</em></p>", + "<div>{x}</div>", + "<p>a<br>{x}</p>", + "<p>{x}<br>z</p>", +] + + +def alt_bodies(rng: random.Random, random_bodies: int) -> list[str]: + """Every punctuation character alone and in shapes, as alt and link text, + then random mixes of images and links in every context.""" + + bodies = [] + for char in [*_PUNCTUATION, "~~~", "^x", "^", "!", "![", "\\", "\\\\x", "&", "a b"]: + for shape in ("{c}", "{c}y", "y{c}", "y {c} z", "{c}{c}"): + alt = shape.format(c=char) + bodies.append(f"<p>{_image(rng, alt)}</p>") + bodies.append(f"<p>{_link(rng, html_module.escape(alt, quote=False))}</p>") + bodies.append(f"<p>{_link(rng, _image(rng, alt))}</p>") + for _ in range(random_bodies): + draw = rng.random() + if draw < 0.3: + inline = _image(rng) + elif draw < 0.5: + inline = _link(rng) + elif draw < 0.7: + inline = _link(rng, _text(rng, 0, 2) + _image(rng) + _text(rng, 0, 2)) + elif draw < 0.85: + inline = ( + _text(rng, 0, 2) + _image(rng) + _text(rng, 0, 3) + _link(rng) + + _text(rng, 0, 2) + ) # fmt: skip + else: + inline = _image(rng) + _image(rng) + _link(rng) + _image(rng) + bodies.append(rng.choice(_CONTEXTS).format(x=inline)) + return bodies + + +_TILDES = [ + "~~~", "~~~~", "~~~~~~", "~~~ info", "~~~x", "~~~ ~~~", "~~~`", "~~~ a`b", + "~~~", "~~~", + "~ ~ ~", "~~", "```", "\\~~~", "a~~~", # controls: no fence to guard +] # fmt: skip +_TILDE_PREFIXES = [ + "", " ", " ", "\t", "<span></span>", "<b></b>", "<span> </span>", + "​", "    ", "\n", +] # fmt: skip +#: Where a run can open a line. A definition list is left out: markdownify +#: writes it in a syntax CommonMark does not have. +_TILDE_PLACES = [ + "<p>{t} rest</p>", + "<p>{t}</p>", + "<p>intro<br>{t}<br>suite</p>", + "<p>intro<br>{t}</p>", + '<p><a href="https://example.org/u">intro<br>{t}</a></p>', + "<p><strong>intro<br>{t}</strong> z</p>", + "<p><em>intro<br>{t}</em></p>", + "<p><s>intro<br>{t}</s></p>", + "<p><span>intro<br>{t}</span></p>", + "<ul><li>{t}</li></ul>", + "<ul><li>a<br>{t}</li></ul>", + '<ol start="3"><li>{t}</li><li>b</li></ol>', + "<ul><li>a<ul><li>b<ol><li>{t}</li></ol></li></ul></li></ul>", + "<ul><li>a<ul><li>b<ol><li>c<br>{t}</li></ol></li></ul></li></ul>", + "<ul><li><p>a</p><p>{t}</p></li></ul>", + "<blockquote>{t}</blockquote>", + "<blockquote><p>a<br>{t}</p></blockquote>", + "<blockquote><blockquote><p>a<br>{t}</p></blockquote></blockquote>", + "<ul><li><blockquote><p>{t}</p></blockquote></li></ul>", + "<blockquote><ul><li>x<br>{t}</li></ul></blockquote>", + "<table><tr><th>h</th></tr><tr><td>{t}</td></tr></table>", + "<table><tr><th>h</th></tr><tr><td>a<br>{t}</td></tr></table>", + "<h3>{t}</h3>", + "<h3>a<br>{t}</h3>", + '<p><img src="https://example.org/i.png" alt="x"><br>{t}</p>', + "<p>{t}<br>code<br>{t}</p>", + "<p>{t}</p><p>text</p><p>{t}</p>", + "<div>{t}</div>", + "<div>a</div>{t}", + "{t}<br><b>x</b>", + "<pre>{t}\ncode\n{t}</pre>", + "<p><code>{t}</code></p>", + "<p>x<br><code>{t}</code></p>", + '<p>a<br><a href="https://example.org/u">{t}</a></p>', + "<p>a<br>!{t}</p>", + '<p>a<br><img src="https://example.org/i.png" alt="{t}"></p>', +] + + +def tilde_bodies(prefixes: int) -> list[str]: + """Every run at every place, behind ``prefixes`` of the prefixes in turn. + + The full product is 5,180 bodies; rotating the prefixes keeps every + run at every place and spreads the prefixes over them. + """ + + bodies = [] + for i, place in enumerate(_TILDE_PLACES): + for j, run in enumerate(_TILDES): + for k in range(prefixes): + prefix = _TILDE_PREFIXES[(i + j + k) % len(_TILDE_PREFIXES)] + bodies.append(place.replace("{t}", prefix + run)) + return bodies + + +#: A '!' right before every inline thing a '[' can open. +_BANGS = [ + '<p>a!<a href="https://example.org/u">voir</a></p>', + '<p>a!<a href="https://example.org/u">https://example.org/u</a></p>', + '<p>a!<a href="mailto:x@example.org">x@example.org</a></p>', + '<p>a!<a href="mailto:x@example.org">mailto:x@example.org</a></p>', + '<p>a!<img src="i.png" alt="x"></p>', + '<p>a!!<img src="i.png" alt="x"></p>', + '<p>a!!<a href="u">v</a></p>', + '<p>!<a href="u">v</a></p>', + '<p>a!<a href="u"><img src="i.png" alt="x"></a></p>', + '<p>a!<a href="u">t<img src="i.png" alt="x"></a></p>', + '<p>a!<br><a href="u">voir</a></p>', + '<p>a!\n<a href="u">voir</a></p>', + '<p>a! <a href="u">voir</a></p>', + '<p><b>a!</b><a href="u">v</a></p>', + '<p>x<b>a!</b><a href="u">v</a></p>', + '<p><em>a!</em><a href="u">v</a></p>', + '<p><code>a!</code><a href="u">v</a></p>', + '<p><a href="u">a!</a><a href="v">b</a></p>', + '<p>a!<b><a href="u">v</a></b></p>', + '<p>a!<span><a href="u">v</a></span></p>', + '<p>a!<s><a href="u">v</a></s></p>', + '<p>a!<a href="u">v</a></p>', + '<p>a\\!<a href="u">v</a></p>', + '<p>a\\\\!<a href="u">v</a></p>', + "<p>a![x](y)</p>", + "<p>a![x](y)</p>", + "<p>a!<a>[x]</a>(y)</p>", + '<p>a!<a href="u">^x</a></p>', + '<p>a!<a href="u" title="t">v</a></p>', + '<p>a<br>!<a href="u">v</a></p>', + '<p>a!<a href="u"></a><a href="v">x</a></p>', + '<p>a!<img src="i.png" alt="^x"></p>', + '<p><img src="i.png" alt="x!"><a href="u">v</a></p>', + '<p>a!<a href="u"><b>v</b></a></p>', + '<p>a!<a href="u"><code>v</code></a></p>', + "<p>!!!</p>", + "<p>Hello! World!</p>", + "<p>a!b!c</p>", + "<p><b>!</b></p>", + "<p>a<b>!x</b>b</p>", + '<p>a<s>!</s><a href="u">v</a></p>', + '<p>a!<br>!<a href="u">v</a></p>', + "<p>!<code>x</code></p>", + '<p>a !<a href="https://example.org/u">https://example.org/u</a></p>', + '<p>a!<a href="u">v</a>!<a href="u">w</a>!</p>', + '<p>a!<a href="https://example.org/u">https://example.org/u</a>!</p>', +] +_BANG_CONTEXTS = [ + "<ul><li>{x}</li></ul>", + "<blockquote>{x}</blockquote>", + "<table><tr><th>h</th></tr><tr><td>{x}</td></tr></table>", + "<h2>{x}</h2>", + "<ol><li><ul><li>{x}</li></ul></li></ol>", + "<div><b>{x}</b></div>", +] + + +def bang_bodies(rng: random.Random, random_bodies: int) -> list[str]: + """Each case alone and in every context, then random ``!``-then-inline runs. + + Bold inside bold is left out of the ``<b>`` context: CommonMark spells + both with ``**``, and ``****`` is not bold. + """ + + bodies = [] + for case in _BANGS: + bodies.append(case) + inner = case[3:-4] + for context in _BANG_CONTEXTS: + if not ("<b>" in context and "<b>" in inner): + bodies.append(context.format(x=inner)) + for _ in range(random_bodies): + before = _text(rng, 0, 3) + rng.choice(["!", "!!", "!", "\\!", "! "]) + after = rng.choice( + [ + _link(rng), + _image(rng), + _link(rng, _image(rng)), + "<br>" + _link(rng), + '<a href="https://example.org/q">https://example.org/q</a>', + "<b>" + _link(rng) + "</b>", + "<code>c</code>" + _link(rng), + ] + ) + context = rng.choice(_CONTEXTS) + if "<b>" in context and "<b>" in after: + context = "<p>{x}</p>" # bold inside bold: left out, as above + bodies.append(context.format(x=before + after + _text(rng, 0, 2))) + return bodies + + +@pytest.mark.parametrize("seed", [1, 2]) +def test_image_alt_and_link_text_survive(seed: int) -> None: + rng = random.Random(seed) + bodies = alt_bodies(rng, 150) + if seed != 1: # the enumerated part is the same for every seed + bodies = bodies[-150:] + assert_property(bodies) + + +def test_a_tilde_run_opening_a_line_stays_text() -> None: + assert_property(tilde_bodies(prefixes=1)) + + +@pytest.mark.parametrize("seed", [1, 2]) +def test_a_bang_before_a_link_or_an_image_stays_text(seed: int) -> None: + rng = random.Random(seed) + bodies = bang_bodies(rng, 150) + if seed != 1: + bodies = bodies[-150:] + assert_property(bodies) + + +# --------------------------------------------------------------------------- +# Inline code holding line breaks (fix 6) +# --------------------------------------------------------------------------- + +_CODE_WORDS = [ + "alpha", "beta", "C:\\Temp", "a_b", "x|y", "5*3", "R&D", "[x]", "`tick`", + "--", "#4521", "<b>", +] # fmt: skip + + +def _code_text(rng: random.Random) -> str: + out = [] + for _ in range(rng.randint(1, 5)): + out.append(rng.choice(_CODE_WORDS)) + out.append(rng.choice([" ", " ", "<br>", "<br><br>", ""])) + if rng.random() < 0.2: + out.insert(0, "<br>") + return "".join(out).strip() or "x" + + +def code_break_body(rng: random.Random) -> str: + """``<code>``, ``<kbd>`` or ``<samp>`` holding line breaks, in every holder.""" + + tag = rng.choice(["code", "code", "kbd", "samp"]) + code = f"<{tag}>{_code_text(rng)}</{tag}>" + before = rng.choice(["", "mot ", "a<br>"]) + after = rng.choice(["", " fin", "<br>b", "fin"]) + inline = f"{before}{code}{after}" + shape = rng.choice(["p", "li", "quote", "h2", "a", "td", "p-strong"]) + if shape == "p": + return f"<p>{inline}</p>" + if shape == "p-strong": + return f"<p>x <strong>{inline}</strong> y</p>" + if shape == "li": + return f"<ul><li>{inline}</li><li>deux</li></ul>" + if shape == "quote": + return f"<blockquote><p>{inline}</p></blockquote>" + if shape == "h2": + return f"<h2>{inline}</h2>" + if shape == "a": + return f'<p><a href="https://example.org/{rng.randint(1, 9)}">{inline}</a></p>' + return ( + "<table><tr><th>H</th><th>I</th></tr>" + f"<tr><td>{inline}</td><td>z</td></tr></table>" + ) + + +@pytest.mark.parametrize("seed", [20261001, 2]) +def test_a_line_break_in_inline_code_survives(seed: int) -> None: + rng = random.Random(seed) + assert_property([code_break_body(rng) for _ in range(300)]) + + +# --------------------------------------------------------------------------- +# Generic inline content, and tables nested in a cell, a heading or a link +# (fixes 7 to 9; 8 is a block in a flattened cell) +# --------------------------------------------------------------------------- + +#: Literal text snippets: Markdown syntax alone and in the shapes that +#: trigger it, beside ordinary words, so that each lands at line starts, at +#: word boundaries and inside words. +_SPECIAL = ( + "\\ \\\\ \\* \\_ \\[ \\# \\. \\` \\\\serveur * ** *** _ __ ___ ______ __init__ " + "_x_ ` `` ``` ~ ~~ ~~~ [ ] [x] [x](y) ![x](y) [1]: [1]:https://example.org/n " + "( ) (y) ! # ## #4521 > - -- --- + . 1. 2026. = == === | a|b --- < <b> </b> " + "<Enter> <!-- --> <https://example.org> <a@b.c> <3 <= & & < A " + "A ᆩ © © &D; { } : \" '" +).split(" ") +_WORDS = ( + "alpha beta snake_case C:\\Temp R&D x mot 5*3 a_b é fichier_de_test_v2.xlsx" +).split(" ") + +#: Code content: a numeric reference with no semicolon is left out, as it +#: was from the generator this one is ported from. +_CODE_SPECIAL = [special for special in _SPECIAL if not special.startswith("&#")] + + +def _prose(rng: random.Random, pieces: int, specials: list[str] = _SPECIAL) -> str: + out = [] + for _ in range(pieces): + out.append(rng.choice(specials) if rng.random() < 0.45 else rng.choice(_WORDS)) + out.append(rng.choice(["", " ", " ", " "])) + return "".join(out).strip() + + +def _inline_html(rng: random.Random, formatted: bool = False, depth: int = 0) -> str: + """Inline HTML Markdown can express. + + No emphasis inside emphasis -- ``****`` is not bold -- no link inside a + link, and a word on either side of every code element, which keeps two + code spans from touching. + """ + + parts = [] + for _ in range(rng.randint(1, 4)): + kind = rng.random() + if kind < 0.55 or depth > 2: + parts.append( + html_module.escape(_prose(rng, rng.randint(1, 4)), quote=False) + ) + elif kind < 0.65 and not formatted: + parts.append(f" <strong>{_inline_html(rng, True, depth + 1)}</strong> ") + elif kind < 0.75 and not formatted: + parts.append(f" <em>{_inline_html(rng, True, depth + 1)}</em> ") + elif kind < 0.82 and depth == 0: + inner = _inline_html(rng, formatted, depth + 1) + parts.append( + f'<a href="https://example.org/{rng.randint(1, 9)}">{inner}</a>' + ) + elif kind < 0.88: + code = html_module.escape(rng.choice(_WORDS + _SPECIAL[:20]), quote=False) + parts.append(f" mot <code>{code}</code> mot ") + elif kind < 0.94: + alt = html_module.escape(_prose(rng, rng.randint(0, 2)), quote=True) + parts.append( + f'<img src="https://example.org/{rng.randint(1, 9)}.png" alt="{alt}">' + ) + else: + parts.append(f"<span>{_inline_html(rng, formatted, depth + 1)}</span>") + parts.append(rng.choice(["", " ", " "])) + return "".join(parts).strip() or "mot" + + +def _nested_table(rng: random.Random, depth: int) -> str: + """A table whose cells hold inline content, a block of it, or a table. + + A block -- a ``<p>`` or a ``<div>`` -- in a flattened cell is what fix 8 + keeps on its holder's line. + """ + + rows = [] + for _ in range(rng.randint(1, 2)): + cells = [] + for _ in range(rng.randint(1, 3)): + if depth < 3 and rng.random() < 0.35: + inner = _nested_table(rng, depth + 1) + content = rng.choice([inner, "avant " + inner, inner + " apres"]) + elif rng.random() < 0.3: + block = rng.choice(["p", "div"]) + content = "".join( + f"<{block}>{_inline_html(rng)}</{block}>" + for _ in range(rng.randint(1, 2)) + ) + else: + content = _inline_html(rng) + tag = "th" if rng.random() < 0.2 else "td" + cells.append(f"<{tag}>{content}</{tag}>") + rows.append("<tr>" + "".join(cells) + "</tr>") + caption = "<caption>Legende</caption>" if rng.random() < 0.15 else "" + return f"<table>{caption}{''.join(rows)}</table>" + + +def nested_table_body(rng: random.Random, where: str) -> str: + """Tables nested up to three deep, inside a cell, a heading or a link.""" + + inner = _nested_table(rng, 2) + if where == "cell": + return ( + "<table><tr><th>A</th><th>B</th></tr>" + f"<tr><td>x</td><td>{inner}</td></tr></table>" + ) + if where == "heading": + return f"<h3>Titre {inner}</h3>" + # No link inside a link. A space where each tag was, or removing a link + # around emphasis could leave it touching the next: "****" is not bold. + inner = re.sub(r"</?a\b[^>]*>", " ", inner) + return f'<p><a href="https://example.org/t">{inner}</a></p>' + + +@pytest.mark.parametrize("where", ["cell", "heading"]) +def test_a_nested_table_reads_as_its_cells_words(where: str) -> None: + """A browser shows the nested table inside its cell or heading; the + converter writes its cells' words there, in order, spaced.""" + + rng = random.Random(20261001) + assert_property([nested_table_body(rng, where) for _ in range(80)]) + + +def test_a_table_inside_a_link_reads_as_the_links_text() -> None: + """A browser draws the table inside the link; Markdown has no table there, + so the converter writes its cells' words as the link's text, in order, + each still linked and formatted -- which is what is checked.""" + + rng = random.Random(20261001) + assert_property([nested_table_body(rng, "link") for _ in range(80)], one_line) + + +# --------------------------------------------------------------------------- +# Underline, highlight and inserted text (fix 11) +# --------------------------------------------------------------------------- + +_RAW_TAGS = ("u", "mark", "ins") +_MARKED_WORDS = [ + "alpha", "beta", "a_b", "*x*", "C:\\Temp", "R&D", "[y]", "!", "~~~", + "<<z", +] # fmt: skip + + +def _marked_inline( + rng: random.Random, inside: frozenset[str] = frozenset(), depth: int = 0 +) -> str: + """Inline content with ``<u>``, ``<mark>`` and ``<ins>`` at any depth. + + The three nest in each other and in themselves. Left out, as in + :func:`_inline_html`: bold or italic inside either, a link inside a + link, and anything but a word inside code; bold, italic and code keep a + space on each side. Let in, with those spaces dropped too, bold and + italic failed 129 of 2,000 bodies (measured 2026-10-02): they are what + the generator this one is ported from tripped on. + """ + + parts = [] + for _ in range(rng.randint(1, 4)): + draw = rng.random() + if draw < 0.4 or depth > 2: + parts.append(rng.choice(_MARKED_WORDS)) + elif draw < 0.6: + tag = rng.choice(_RAW_TAGS) + pad = rng.choice(["", " "]), rng.choice(["", " "]) + body = _marked_inline(rng, inside, depth + 1) + parts.append(f"<{tag}>{pad[0]}{body}{pad[1]}</{tag}>") + elif draw < 0.68 and not inside & {"b", "em"}: + tag = rng.choice(["b", "em"]) + body = _marked_inline(rng, inside | {"b", "em"}, depth + 1) + parts.append(f" <{tag}>{body}</{tag}> ") + elif draw < 0.72 and "s" not in inside: + parts.append(f"<s>{_marked_inline(rng, inside | {'s'}, depth + 1)}</s>") + elif draw < 0.78: + parts.append(f" mot <code>{rng.choice(_MARKED_WORDS)}</code> mot ") + elif draw < 0.86 and "a" not in inside: + body = _marked_inline(rng, inside | {"a"}, depth + 1) + parts.append( + f'<a href="https://example.org/{rng.randint(1, 9)}">{body}</a>' + ) + elif draw < 0.93: + parts.append("<br>") + else: + alt = rng.choice(_MARKED_WORDS) + parts.append(f'<img src="https://example.org/i.png" alt="{alt}">') + return rng.choice([" ", "", " "]).join(parts) + + +def marked_body(rng: random.Random) -> str: + """One to three blocks of :func:`_marked_inline`, in every kind of holder.""" + + def block() -> str: + draw = rng.random() + if draw < 0.15: + level = rng.randint(1, 3) + return f"<h{level}>{_marked_inline(rng)}</h{level}>" + if draw < 0.3: + return ( + f"<ul><li>{_marked_inline(rng)}</li><li>{_marked_inline(rng)}</li></ul>" + ) + if draw < 0.4: + return ( + "<table><tr><th>h</th><th>i</th></tr>" + f"<tr><td>{_marked_inline(rng)}</td><td>{_marked_inline(rng)}</td></tr>" + "</table>" + ) + if draw < 0.5: + return f"<blockquote><p>{_marked_inline(rng)}</p></blockquote>" + return f"<p>{_marked_inline(rng)}</p>" + + blocks: list[str] = [] + for _ in range(rng.randint(1, 3)): + new = block() + if blocks and new[:4] in {"<ul>", "<blo"} and new[:4] == blocks[-1][:4]: + blocks.append("<p>mot</p>") # two of these in a row merge in Markdown + blocks.append(new) + return "".join(blocks) + + +def _raw_spans(html: str) -> list[tuple[str, str]]: + """Each ``<u>``, ``<mark>`` and ``<ins>`` with text, and its text, in order. + + The display oracle cannot see these tags, so this is what holds the + converter to keeping them. The text is the tag's words + (:func:`.display.text_words`): the converter moves a tag's edge spaces + and line breaks outside it. + """ + + soup = BeautifulSoup(html, "html.parser") + spans = [] + for tag in soup.find_all(_RAW_TAGS): + words = text_words(str(tag)) + if words: + spans.append((tag.name, " ".join(words))) + return spans + + +@pytest.mark.parametrize("seed", [20261002, 2]) +def test_underline_and_highlight_keep_their_text_and_their_tags(seed: int) -> None: + rng = random.Random(seed) + bodies = [marked_body(rng) for _ in range(150)] + + assert_property(bodies) + lost = [ + html for html in bodies if _raw_spans(render(read(html))) != _raw_spans(html) + ] + assert not lost, f"{len(lost)} bodies lost a raw tag; the first: {lost[0]!r}" + + +# --------------------------------------------------------------------------- +# Whole documents +# --------------------------------------------------------------------------- + + +def _paragraph(rng: random.Random) -> str: + return "<br>".join(_inline_html(rng) for _ in range(rng.randint(1, 3))) + + +def _list_html(rng: random.Random, depth: int) -> str: + """A list whose items hold text, and sometimes code, a list, and more text.""" + + tag = rng.choice(["ul", "ol"]) + items = [] + for _ in range(rng.randint(1, 3)): + shape = rng.random() + if depth < 3 and shape < 0.1: + items.append(f"<li>{_list_html(rng, depth + 1)}</li>") + continue + inner = f"<p>{_paragraph(rng)}</p>" if shape < 0.3 else _paragraph(rng) + if rng.random() < 0.1: + code = html_module.escape(_prose(rng, 3, _CODE_SPECIAL), quote=False) + inner += f"<pre>{code}</pre>" + if depth < 3 and rng.random() < 0.3: + inner += _list_html(rng, depth + 1) + after = rng.random() + if after < 0.15: + inner += _inline_html(rng) + elif after < 0.25: + inner += f"<p>{_paragraph(rng)}</p>" + elif after < 0.3: + inner += f"<blockquote><p>{_paragraph(rng)}</p></blockquote>" + items.append(f"<li>{inner}</li>") + return f"<{tag}>{''.join(items)}</{tag}>" + + +def _block(rng: random.Random, depth: int = 0) -> str: + kind = rng.random() + if kind < 0.35 or depth > 1: + return f"<p>{_paragraph(rng)}</p>" + if kind < 0.45: + level = rng.randint(1, 3) + return f"<h{level}>{_inline_html(rng)}</h{level}>" + if kind < 0.60: + return _list_html(rng, depth + 1) + if kind < 0.70: + return f"<blockquote>{_block(rng, depth + 1)}</blockquote>" + if kind < 0.80: + head = "".join(f"<th>{_inline_html(rng)}</th>" for _ in range(2)) + rows = "".join( + "<tr>" + + "".join(f"<td>{_inline_html(rng)}</td>" for _ in range(2)) + + "</tr>" + for _ in range(rng.randint(1, 2)) + ) + return f"<table><tr>{head}</tr>{rows}</table>" + if kind < 0.87: + code = html_module.escape(_prose(rng, 4, _CODE_SPECIAL), quote=False) + return f"<pre>{code}</pre>" + return f"<div>{_paragraph(rng)}</div>" + + +def document(rng: random.Random) -> str: + """A body of one to four blocks. + + Markdown merges two adjacent lists of one type, or two adjacent block + quotes, into one -- a limit of the format -- so a paragraph separates + them. + """ + + blocks: list[str] = [] + for _ in range(rng.randint(1, 4)): + block = _block(rng) + mergeable = block[:4] in {"<ul>", "<ol>", "<blo"} + if blocks and mergeable and block[:4] == blocks[-1][:4]: + blocks.append("<p>mot</p>") + blocks.append(block) + return "".join(blocks) + + +@pytest.mark.parametrize("seed", range(4)) +def test_a_generated_document_survives(seed: int) -> None: + """The earlier generator of this package, 40 bodies a seed.""" + + rng = random.Random(seed) + assert_property([document(rng) for _ in range(40)]) diff --git a/easyvista_python_client/content/tests/test_round_trip.py b/easyvista_python_client/content/tests/test_round_trip.py new file mode 100644 index 0000000..c8d3dde --- /dev/null +++ b/easyvista_python_client/content/tests/test_round_trip.py @@ -0,0 +1,630 @@ +"""Realistic content survives the trip between memo HTML and Markdown, both ways. + +The contract, checked with an HTML parser (:mod:`.display`) rather than by eye: + +* **HTML first.** ``to_transport(from_transport(html))`` displays what + ``html`` displays -- the same blocks, words, and bold/italic/code/link on + each character -- and the Markdown is a fixed point: reading back what it + renders gives the same Markdown. +* **Markdown first.** Markdown a caller writes renders to HTML that reads + back as Markdown displaying the same, and that Markdown is a fixed point. + +A port of ``glpi_python_client``'s ``content/tests/test_round_trip.py`` at +commit ``524304a``, with the converter renamed. Its bodies are that +package's synthetic corpus -- the shapes a rich-text editor and a mail +collector produce, in invented or generic French -- and they are kept as +they are: the converter is the same code, so its corpus is too. What +EasyVista itself stores is not claimed here. New here: a body shaped like +an e-mail notification template, two plain-text cases (a CR LF, an +entity), the note that how EasyVista's own UI displays a plain-text memo +is unverified, a pin on what ``plain_text_is_markdown=True`` reads as HTML, +and the HTML of this package's earlier regression bodies. GLPI's long-body +tests moved, reshaped as growth tests, to :mod:`.test_cost`. + +Adversarial input -- syntax characters packed into every position -- is +:mod:`.test_properties`' subject; this module holds the converter to +realistic content. Every URL is under ``example.org``. +""" + +from __future__ import annotations + +from typing import cast + +import pytest + +from easyvista_python_client.content.conversion import EasyvistaContentConverter +from easyvista_python_client.content.tests.display import displayed, text_words + +read = EasyvistaContentConverter.from_transport +render = EasyvistaContentConverter.to_transport + + +def assert_survives(html: str) -> str: + """Assert the HTML-first property for one body, and return its Markdown.""" + + markdown = read(html) + rendered = render(markdown) + + assert displayed(rendered) == displayed(html), ( + f"the Markdown does not display what the HTML did\n" + f" html: {html!r}\n markdown: {markdown!r}\n rendered: {rendered!r}" + ) + assert read(rendered) == markdown, ( + f"the Markdown is not a fixed point\n markdown: {markdown!r}\n" + f" again: {read(rendered)!r}" + ) + return markdown + + +# --------------------------------------------------------------------------- +# HTML first +# --------------------------------------------------------------------------- + +#: Bodies shaped the way a rich-text editor and a mail collector store them. +REALISTIC = [ + "<p>Bonjour,</p><p>Le PC du poste 12 ne démarre plus depuis ce matin.</p>", + "<p>Bonjour,<br>Le PC ne démarre plus.<br>Cordialement,<br>Jean</p>", + "<p>Merci de <strong>redémarrer</strong> le serveur <em>avant</em> 18h.</p>", + "<p>Étapes :</p><ul><li>ouvrir la session</li><li>lancer Outlook</li></ul>", + "<table><thead><tr><th>Poste</th><th>IP</th></tr></thead>" + "<tbody><tr><td>PC12</td><td>10.0.0.12</td></tr></tbody></table>", + "<h2>Contexte</h2><p>Migration du serveur.</p>", + '<p>Voir <a href="https://example.org/doc">https://example.org/doc</a></p>', + '<p>Voir <a title="doc" href="https://example.org/doc">la doc</a></p>', + '<p><a href="https://example.org/wiki/Test_(informatique)">wiki</a></p>', + "<p>prix 5*3 et note * importante</p>", + "<p>appuyer sur <Entrée> puis valider</p>", + "<p>Service R&D, bâtiment A & B</p>", + "<p>Montant : 12 000 €</p>", + # The display oracle cannot see underline. Since fix 11 (test_fixes.py) + # the <u> is kept as raw HTML in the Markdown, where GLPI 917f030 + # dropped it; this row checks the words and the fixed point. + '<p><span style="color: #e03e2d;">URGENT</span> <u>à traiter</u></p>', + '<p>Cordialement</p><p><img src="https://example.org/logo.png" alt="Logo"></p>', + "<p>Bonjour</p><blockquote><p>Message d'origine</p></blockquote>", + '<pre>Traceback (most recent call last):\n File "x.py", line 1\nError</pre>', + "<p>Lancer <code>ipconfig /all</code> puis envoyer.</p>", + "<p>Fichier mon_fichier_final.docx et __init__</p>", + "<p># pas un titre</p><p>- pas une liste</p>", + '<p>Contact <a href="mailto:support@example.org">support@example.org</a></p>', + "<div>Bonjour,</div><div><br></div><div>Le serveur est down.</div>", + "<p>* point un<br>* point deux</p>", + "<p>[INFO] tâche [1] terminée</p>", + "<p>[1]: https://example.org/note</p>", + r"<p>Chemin C:\Users\jdupont\Desktop</p>", + r"<p>Accès au partage \\serveur\compta\2026 refusé</p>", + "<p>Merci 👍</p>", + "<p>Titre<br>=====</p>", + "<p>---</p><p>signature</p>", + "<p>1) un<br>2) deux</p>", + "<ul><li><p>point un</p><p>suite du point</p></li><li>point deux</li></ul>", + "<blockquote><ul><li>a</li><li>b</li></ul></blockquote>", + "<p>Nom : ______ Prénom : ______</p>", + "<p>a | b | c</p>", + "<p>> pas une citation</p>", + "<p>Cordialement,<br>Jean Dupont<br>--<br>Service IT</p>", + "<p>Merci<br>-----------<br>Jean Dupont</p>", + "<p>Le 30/09, Jean a écrit :<br>> merci<br>> cordialement</p>", + "<ul><li>Réseau<ul><li>switch 3</li><li>borne wifi</li></ul></li></ul>", + "<ol><li>Arrêter</li><li>Sauvegarder<ol><li>la base</li><li>les fichiers</li>" + "</ol></li><li>Redémarrer</li></ol>", + '<ol start="3"><li>trois</li><li>quatre</li></ol>', + "<ul><li>Réseau<ul><li>switch 3</li></ul>à vérifier</li></ul>", + "<p><b>Important</b> : voir <i>ci-dessous</i></p>", + "<table><tr><th>Commande</th></tr><tr><td>ps aux | grep java</td></tr></table>", + "<table><tr><th>Commande</th></tr><tr><td><code>ps aux | grep java</code></td></tr>" + "</table>", + "<pre>ligne\n```\nfin</pre>", + "<p>Suite à la mise à jour, <strong>3 postes</strong> ne se connectent plus :" + "</p><ul><li>PC12 (salle 3)</li><li>PC14 – <em>poste d'accueil</em></li>" + '</ul><p>Voir <a href="https://example.org/kb/42">https://example.org/kb/42</a>' + ' et le <a href="https://example.org/kb/43" title="KB 43">KB 43</a>.</p>', +] + + +@pytest.mark.parametrize("html", REALISTIC) +def test_a_realistic_body_displays_the_same_after_the_round_trip(html: str) -> None: + assert_survives(html) + + +#: What e-mail clients add: Outlook's blank paragraphs and line breaks made of +#: a non-breaking space, a blank line inside a paragraph, a label in bold +#: right before a figure, and Gmail's quoted reply. +E_MAIL_SHAPES = [ + pytest.param( + '<p class="MsoNormal">Bonjour,<o:p></o:p></p>' + '<p class="MsoNormal"><o:p> </o:p></p>' + '<p class="MsoNormal">Le serveur répond.<o:p></o:p></p>', + id="outlook-blank-paragraph", + ), + pytest.param("<p>Bonjour<br> <br>Texte</p>", id="nbsp-blank-line"), + pytest.param("<p>Bonjour,<br><br>Texte</p>", id="blank-line-in-a-paragraph"), + pytest.param("<p><b>Total:</b>12 postes</p>", id="bold-label-before-a-figure"), + pytest.param( + '<div dir="ltr">Merci</div><div class="gmail_quote"><div>Le lun. a écrit :' + "</div><blockquote>Le serveur est down.</blockquote></div>", + id="gmail-quote", + ), + pytest.param("<p>Ligne<br></p><p>Suite</p>", id="break-ending-a-paragraph"), + pytest.param( + "<p>Source wrapped\nmid-sentence\nby Outlook.</p>", id="newlines-in-source" + ), +] + + +@pytest.mark.parametrize("html", E_MAIL_SHAPES) +def test_an_e_mail_shape_displays_the_same_after_the_round_trip(html: str) -> None: + assert_survives(html) + + +# --------------------------------------------------------------------------- +# Bodies from this package's earlier tests +# --------------------------------------------------------------------------- +# +# The HTML of the regression bodies this package's tests held before 0.4.0, +# when its converter was python-markdown's dialect: a body does not depend on +# which converter reads it, and the fixes this converter carries were +# measured against these. Only the HTML is kept. The Markdown those tests +# expected was python-markdown's, so each body is held to the HTML-first +# property instead -- the same display, and a fixed point. Grouped by the +# test each came from; every body is synthetic, and every URL is under +# example.org. The long bodies dense with syntax are growth tests in +# test_cost.py. + +#: The earlier tests' bodies, by the test they came from. +EARLIER_REGRESSIONS: dict[str, list[str]] = { + "obsolete": [ + '<font color="red">URGENT</font> serveur HS', + "<strike>ancien</strike> nouveau", + "<big>gros</big> texte", + "<tt>code</tt> texte", + "<nobr>sans coupure</nobr>", + ], + "line-break": [ + "<p>a<br> b<br> c</p>", + "<p>a <br>b</p>", + "<p><strong>a<br></strong>b</p>", + ], + "pre": [ + "<ul><li>item<pre>#4521 code</pre></li></ul>", + "<blockquote><pre>#4521 C:\\Temp\n indenté</pre></blockquote>", + ], + "image": [ + '<h2>Titre <img src="https://example.org/i.png" alt="logo"></h2>', + "<table><tr><th>a</th></tr><tr><td>" + '<img src="https://example.org/i.png" alt="x"></td></tr></table>', + ], + "block-in-item": [ + "<ul><li>Réseau<ul><li>switch 3</li></ul><blockquote>cité</blockquote></li>" + "</ul>", + "<ul><li>Réponse :<blockquote>cité</blockquote></li></ul>", + "<ul><li><blockquote>cité<br>suite</blockquote></li></ul>", + "<ul><li><ul><li>un</li><li>deux</li></ul></li></ul>", + "<ul><li><ul><li>un<ul><li>a</li></ul></li></ul></li></ul>", + "<ul><li><ul><li><ul><li>un</li><li>deux</li></ul></li></ul></li><li>trois" + "</li></ul>", + ], + "code-after-list": [ + "<ul><li>item<ul><li>sous-item</li></ul><pre>#4521 code</pre></li></ul>", + "<blockquote><ul><li>item</li></ul><pre>#4521 code</pre></blockquote>", + "<ul><li>item<ul><li>sous-item</li></ul><div><style>p{margin:0}</style>" + "</div><pre>#4521 code</pre></li></ul>", + ], + "literal-syntax": [ + "<p>D:\\logs\\.cache et \\\\srv\\share\\[archive] et HKLM\\SOFTWARE\\#1</p>", + "<p>fin de ligne \\<br>suite</p>", + "<h2>Chemin C:\\</h2>", + "<p>_______</p>", + "<p>Code : _______ fin</p>", + "<p>a*b*c et 5*3 <strong>gras</strong></p>", + "<p><strong>x</strong>* suite</p>", + "<p>+ un<br>+ deux</p>", + "<p>- pas une liste</p>", + "<p>Bonjour<br>#4521 doublon</p>", + "<p>***</p>", + "<ul><li>--</li></ul>", + "<ul><li>___</li></ul>", + "<ul><li><ul><li>-</li></ul></li></ul>", + "<ol><li><ul><li>--</li></ul></li></ol>", + "<ul><li><p>--</p><p>suite</p></li></ul>", + "<h1>C#</h1>", + "<h2>Titre ##</h2>", + "<p>```<br>code<br>```</p>", + "<p>~~~<br>code<br>~~~</p>", + "<p>a | b<br>--- | ---</p>", + "<p>[x](https://example.org/y) et ![x](https://example.org/y.png)</p>", + '<p><a href="https://example.org/u">rapport [final].pdf</a></p>', + '<p><a href="https://example.org/u">a]b</a></p>', + '<p><a href="https://example.org/u">voir [x](y)</a></p>', + '<p>[a <a href="https://example.org/u">[b</a> ](https://example.org/x)</p>', + '<p><img src="https://example.org/c.png" alt="capture [1].png"></p>', + '<p><img src="https://example.org/c.png" alt="a]b"></p>', + "<p>if x<y then z>0</p>", + "<p><https://example.org/x> et <support@example.org></p>", + '<p><3<img src="https://example.org/i.png" alt="a b">@c></p>', + '<p><a href="https://example.org/u"><<a@c></a></p>', + '<p><3<a href="https://example.org">https://example.org</a>@c></p>', + '<p><a <a href="https://example.org/T_(x)">wiki</a>b@c></p>', + "<p><!-- note --></p>", + "<p>&amp; &lt; &#65; &#4521 &copy; &copy</p>", + "<p>a`b`c et <code>x</code> puis `</p>", + "<p><code>x</code>`y</p>", + "<p>`<code>x</code> y</p>", + "<table><tr><th>a</th><th>b</th></tr><tr><td>x`y</td><td>`z</td></tr></table>", + ], + "nested-list": [ + "<ul><li>Réseau<ul><li>switch 3</li><li>borne wifi</li></ul></li>" + "<li>Imprimante</li></ul>", + "<ol><li>Arreter</li><li>Sauvegarder<ol><li>la base</li><li>les fichiers" + "</li></ol></li><li>Redemarrer</li></ol>", + "<ol><li>un<ul><li>a</li></ul></li><li>deux</li></ol>", + "<ul><li>a<ul><li>b<ul><li>c<ul><li>d</li></ul></li></ul></li></ul></li></ul>", + '<p>intro</p><ol start="3"><li>trois</li><li>quatre</li></ol>', + "<ul><li>5*3<ul><li>2*4</li></ul></li></ul>", + ], + "prose": [ + "<p>Porte-monnaie - un tiret - et -- deux</p>", + "<p>5 * 3 = 15 et prix 5*3 et note * importante</p>", + "<p>si a < b et x <= y alors 2 < 3</p>", + "<p>2 + 2 = 4, +33 6 12 34 56 78, 1) un, 3.14</p>", + "<p>Pourquoi ? Parce que ! 100 % a/b a=b ~5 minutes</p>", + "<p>l`imprimante et la variable user_id et _temp</p>", + "<p>ps aux | grep java</p>", + ], + "realistic": [ + "<p><https://example.org/x></p>", + "<p>m<sup>2</sup></p>", + "<p><support@example.org></p>", + "<p>~~pas barré~~</p>", + "<p>calcul 5 * 3 * 2 = 30</p>", + "<p>voir \\\\srv\\partage\\__archive__\\2026</p>", + "<p>Dossier C:\\_temp\\logs</p>", + "<p>voir la note [1]</p><p>[1]: https://example.org/note</p>", + '<p><font color="red">rouge</font> <span style="font-size:14px">texte</span>' + "</p>", + "<p>module __init__ et _x_</p>", + "<p>1) un<br>2) deux</p><pre>ligne 1\n ligne indentée</pre>", + ], + "own-rule": [ + '<p><a name="_MailEndCompose">Bonjour</a> Jean</p>', + '<p><a href="https://example.org/u"></a>texte</p>', + "<p><strong></strong>texte</p>", + '<p><a href="https://example.org/u">un<br>deux</a></p>', + "<table><tr><th><h3>titre</h3></th></tr><tr><td>x</td></tr></table>", + "<h2><blockquote>cité</blockquote></h2>", + '<p><img src="https://example.org/i.png" alt="ligne 1\n\n ligne 2"></p>', + "<ul><li><script>x()</script><pre>code</pre></li></ul>", + "<ul><li><!-- note --><pre>code</pre></li></ul>", + "<ul><li><ul></ul><pre>code</pre></li></ul>", + '<ul><li><img src="https://example.org/i.png" alt="i"><pre>code</pre></li>' + "</ul>", + ], + "struck": [ + "<p><s>ancien</s> nouveau</p>", + "<p><del>ancien</del> nouveau</p>", + "<p><strike>ancien</strike> nouveau</p>", + ], + "hidden": [ + "<style>p.MsoNormal{margin:0cm;font-size:11pt}</style><p>Bonjour</p>", + "<html><head><title>RE: Imprimante

      Bonjour

      " + "", + ], + "misreading": [ + "

      # pas un titre

      ", + "

      #4521 est un doublon

      ", + "

      2026. Une annee

      ", + ], + "upper-case": [ + "

      Bonjour

      ", + "
      Bonjour", + "
      Bonjour
      ", + ], +} + + +@pytest.mark.parametrize( + "html", + [ + pytest.param(html, id=f"{group}-{index}") + for group, bodies in EARLIER_REGRESSIONS.items() + for index, html in enumerate(bodies) + ], +) +def test_an_earlier_regression_body_survives(html: str) -> None: + assert_survives(html) + + +@pytest.mark.parametrize( + ("html", "expected"), + [ + pytest.param( + "

      Voir fichier_de_test_v2.xlsx et mon_fichier_final.docx

      ", + "Voir fichier_de_test_v2.xlsx et mon_fichier_final.docx", + id="file-names", + ), + pytest.param( + r"

      Chemin C:\Temp\logs et C:\Users\Admin\Documents

      ", + r"Chemin C:\Temp\logs et C:\Users\Admin\Documents", + id="windows-paths", + ), + pytest.param( + "

      Le ticket # 3 et le #4521 sont liés, C# aussi

      ", + "Le ticket # 3 et le #4521 sont liés, C# aussi", + id="hashes", + ), + pytest.param( + "

      Service R&D, bâtiment A & B

      ", + "Service R&D, bâtiment A & B", + id="ampersands", + ), + pytest.param( + "

      [INFO] tâche [1] terminée (voir note)

      ", + "[INFO] tâche [1] terminée (voir note)", + id="brackets", + ), + pytest.param( + "

      Voir https://example.org/doc?a=1&b=2 ou support@example.org

      ", + "Voir https://example.org/doc?a=1&b=2 ou support@example.org", + id="bare-url-and-address", + ), + pytest.param( + "

      Fichier mon_fichier_final.docx et __init__

      ", + r"Fichier mon_fichier_final.docx et \_\_init\_\_", + id="dunder", + ), + ], +) +def test_ordinary_prose_reads_back_as_typed(html: str, expected: str) -> None: + """Prose carries no escape, and literal syntax is escaped only to stay text.""" + + assert assert_survives(html) == expected + + +def test_a_table_nested_in_a_cell_keeps_its_text() -> None: + """A table inside a table keeps every word. + + E-mail signatures and notification templates are often laid out so. + """ + + html = ( + "
      Signature
      " + "
      Jean DupontService IT
      " + ) + + markdown = read(html) + + assert "Jean Dupont" in render(markdown) + assert "Service IT" in render(markdown) + + +@pytest.mark.parametrize( + "html", + [ + pytest.param( + "

      Bonjour

      ", id="style" + ), + pytest.param("

      Bonjour

      ", id="script"), + pytest.param( + "RE: Imprimante" + "

      Bonjour

      ", + id="title", + ), + ], +) +def test_what_a_browser_does_not_display_is_dropped(html: str) -> None: + assert read(html) == "Bonjour" + + +def test_a_fence_keeps_its_language() -> None: + markdown = "```powershell\nGet-Service\n```" + + assert read(render(markdown)) == markdown + + +# --------------------------------------------------------------------------- +# A notification template: a layout table around a table of label/value cells +# --------------------------------------------------------------------------- +# +# New here. E-mail templates are commonly laid out as an outer one-cell +# table around an inner one-row table: each label in bold over its value, +# spacer cells between them, and here one label in a cell of its own beside +# its value's cell. That shape exercises the flattening of a table nested in +# a cell, fix 9 of test_fixes.py (bold that ends on punctuation at a +# flattened cell's edge still closes -- GLPI 917f030 lost the fixed point +# there, measured 2026-10-02) and
      in a cell, all at once. Every word +# below is invented. + +_LABELS = ["Zorvane", "Plimet", "Quastor", "Brindel"] +_VALUES = ["trelm 4471", "oskar-vint", "maludi pref", "kobra 12/09"] + + +def _notification_template() -> str: + cells = [] + for label, value in zip(_LABELS, _VALUES, strict=True): + cells.append("

    {label}


    {value}

    Velusk :sanquo
    " + "".join(cells) + "
    " + footer = ( + 'Grovak le suivi ' + 'Trelune' + ) + return f"
    {inner}
    {footer}
    " + + +def test_a_notification_template_keeps_every_word_in_order() -> None: + """Every word, in order, each label still bold and the link still a link. + + The layout does not survive: GFM has no nested table, so the inner row + becomes one line of text in the outer cell, and a GFM table always has a + header row, so the outer table gains an empty one. Both are limits of + the format. Apart from that header row, the display is the same. + """ + + html = _notification_template() + + markdown = read(html) + rendered = render(markdown) + + assert text_words(rendered) == text_words(html) + name, rows = cast(tuple[str, tuple[object, ...]], displayed(html)[0]) + empty_header = ((True, ()),) # the header row GFM requires: one empty cell + assert displayed(rendered) == ((name, (empty_header, *rows)),) + for label in [*_LABELS, "Velusk :"]: + assert f"**{label}**" in markdown + assert "](https://example.org/suivi?ref=PX-7)" in markdown + assert "![Trelune](https://example.org/marque.png)" in markdown + assert "ev-cell" not in markdown and "data-ev-" not in markdown + assert read(rendered) == markdown + + +# --------------------------------------------------------------------------- +# Plain text and the write path +# --------------------------------------------------------------------------- +# +# A memo with no HTML element in it is read as literal lines, a line break +# per line. How EasyVista's own UI displays such a memo is UNVERIFIED -- the +# vendor's form editor documents a MEMO object (text) beside a TEXT AREA +# object that accepts HTML [doc:https://docs.easyvista.com/docs/form], and +# its comment-log page configures a custom field gathering a request's +# comments as a "Text area" +# [doc:https://docs.easyvista.com/docs/service-manager-comment-log-creation, +# read 2026-10-02]. That leans towards the UI showing memo text as HTML, +# not as the literal lines read here, but neither page says which object +# the built-in description or the action history is. These tests pin what +# the converter does, not what EasyVista shows. + + +@pytest.mark.parametrize( + "text", + [ + "__init__ et _x_", + "# pas un titre", + "* point un\n* point deux", + r"\\serveur\partage", + "if x0", + "Bonjour,\n\nMerci.", + pytest.param("Velk ondra\r\nPrastin", id="crlf"), + ], +) +def test_a_plain_text_body_reads_as_the_text_it_is(text: str) -> None: + """A body with no HTML element is literal text, its lines lines. + + The ``crlf`` case is new: a CR LF separates lines as LF does. An HTML + reading of the same characters would show a space instead. + """ + + markdown = read(text) + shown = displayed(render(markdown)) + lines = text.replace("<", "<").splitlines() + expected = displayed("

    " + "
    ".join(lines) + "

    ") + + assert shown == expected + assert read(render(markdown)) == markdown + + +def test_an_entity_in_a_plain_text_body_is_read_literally() -> None: + """New here: with no element, `` `` is six characters of text, not a space. + + Pinned because it is the consequence of reading a tag-less memo as text, + and the one a reviewer would most likely expect the other way; which + reading matches EasyVista's UI is unverified (see the section comment). + """ + + markdown = read("Trasvel ok") + + assert displayed(render(markdown)) == displayed("

    Trasvel&nbsp;ok

    ") + + +@pytest.mark.parametrize( + "markdown", + [ + "Run **passwd**, then check `logs`.", + "| a | b
    c |\n| --- | --- |", + "Press Ctrl + C.", + "line one\nline two", + ], +) +def test_the_write_path_keeps_caller_markdown_verbatim(markdown: str) -> None: + """``plain_text_is_markdown=True`` never rewrites the caller's Markdown.""" + + assert read(markdown, plain_text_is_markdown=True) == markdown + + +def test_the_write_path_still_converts_an_html_document() -> None: + assert read("

    A bold move

    ", plain_text_is_markdown=True) == ( + "A **bold** move" + ) + + +@pytest.mark.parametrize( + ("markdown", "expected"), + [ + pytest.param( + " puis
    suite **gras**", + "puis\\\nsuite \\*\\*gras\\*\\*", + id="autolink-first", + ), + pytest.param( + " puis Ctrl **gras**", + "puis `Ctrl` \\*\\*gras\\*\\*", + id="angle-text-first", + ), + ], +) +def test_markdown_opening_with_angle_brackets_and_holding_html_is_read_as_html( + markdown: str, expected: str +) -> None: + """A limit, pinned so that the documentation stays true. + + ``plain_text_is_markdown=True`` passes a value through unless it starts + with ``<`` and holds a real HTML element *anywhere*. So Markdown that + opens with an autolink or with angle-bracketed text, and carries inline + HTML further on, is read as HTML: the autolink, an unknown element to + the HTML parser, is dropped, and the Markdown syntax is escaped as text. + Inherited from ``glpi_python_client`` 917f030, whose converter makes the + same check. + """ + + assert read(markdown, plain_text_is_markdown=True) == expected + + +# --------------------------------------------------------------------------- +# Markdown first +# --------------------------------------------------------------------------- + +#: What an integrator writes into a ticket, an action or a comment. +CALLER_MARKDOWN = [ + "The printer is **offline**.", + "Line one\nline two", + "# Procédure\n\n1. Arrêter le service\n2. Vider le cache\n3. Redémarrer", + "- Réseau\n - switch 3\n - borne wifi\n- Imprimante", + "- Réseau\n - switch 3\n- Imprimante", + "1. Sauvegarder\n 1. la base\n 2. les fichiers\n2. Redémarrer", + 'Voir [la procédure](https://example.org/kb/42 "KB 42").', + "Lien direct : ", + "![capture](https://example.org/c.png)", + "Lancer `ipconfig /all` puis envoyer le résultat.", + "```powershell\nGet-Service | Where-Object Status -eq Running\n```", + "> Le 30/09, Jean a écrit :\n> merci", + "| Poste | IP |\n| :--- | ---: |\n| PC12 | 10.0.0.12 |", + "| Commande | Effet |\n| --- | --- |\n| `ps aux \\| grep java` | processus |", + "Fichier `mon_fichier_final.docx` et chemin C:\\Temp\\logs.", + "Contact : support@example.org, R&D, 5 * 3 = 15.", + "Résolu ✅ — merci 👍", + "**Cause :** disque plein.\n\n**Solution :** purge des journaux.", + "Avant :\n\n---\n\nAprès.", + "Étapes :\n- ouvrir la session\n- lancer Outlook", +] + + +@pytest.mark.parametrize("markdown", CALLER_MARKDOWN) +def test_caller_markdown_displays_the_same_after_the_round_trip(markdown: str) -> None: + html = render(markdown) + back = read(html) + + assert displayed(render(back)) == displayed(html) + assert read(render(back)) == back diff --git a/easyvista_python_client/exceptions.py b/easyvista_python_client/exceptions.py index e8f9996..48eebee 100644 --- a/easyvista_python_client/exceptions.py +++ b/easyvista_python_client/exceptions.py @@ -66,11 +66,11 @@ class is part of the core package, so ``except EasyvistaContentError`` works whether or not the extra is installed. It exists so that no failure of the content layer escapes the package's - taxonomy. The conversion runs third-party parsers (``markdownify`` - inbound, ``markdown`` outbound), and a parser fault would otherwise reach - the caller as a bare builtin that ``except EasyvistaError`` does not - catch. HTML nested too deeply to convert is not an error at all: it is - answered with the memo's text instead. + taxonomy. The conversion runs third-party parsers (``markdownify`` and + ``mdformat`` inbound, ``cmark-gfm`` outbound), and a parser fault would + otherwise reach the caller as a bare builtin that ``except EasyvistaError`` + does not catch. HTML nested too deeply to convert is not an error at all: + it is answered with the memo's text instead. No request is involved, so the converter raises it with a message alone and ``status_code``, ``ev_code``, ``ev_message`` and ``body`` stay diff --git a/easyvista_python_client/testing/test_public_api.py b/easyvista_python_client/testing/test_public_api.py index 74f4212..619b224 100644 --- a/easyvista_python_client/testing/test_public_api.py +++ b/easyvista_python_client/testing/test_public_api.py @@ -1,13 +1,44 @@ +import re import subprocess import sys +import tomllib from pathlib import Path +import pytest + import easyvista_python_client _REPO_ROOT = Path(__file__).resolve().parents[2] -#: The import names of the optional ``content`` extra's three distributions. -_CONTENT_MODULES = ("bs4", "markdown", "markdownify") +#: Each distribution in the optional ``content`` extra, mapped to the module it +#: installs. The two names differ for three of the six, so neither can be +#: derived from the other; ``test_the_content_import_names_match_the_extra`` +#: keeps the keys equal to what pyproject.toml actually declares. +_CONTENT_IMPORT_NAMES = { + "beautifulsoup4": "bs4", + "cmarkgfm": "cmarkgfm", + "markdown-it-py": "markdown_it", + "markdownify": "markdownify", + "mdformat": "mdformat", + "mdformat-tables": "mdformat_tables", +} + +#: The import names of the optional ``content`` extra's distributions. +_CONTENT_MODULES = tuple(sorted(_CONTENT_IMPORT_NAMES.values())) + + +def _optional_dependencies() -> dict[str, list[str]]: + """``[project.optional-dependencies]`` as pyproject.toml declares it.""" + config = tomllib.loads((_REPO_ROOT / "pyproject.toml").read_text(encoding="utf-8")) + extras: dict[str, list[str]] = config["project"]["optional-dependencies"] + return extras + + +def _distribution_name(requirement: str) -> str: + """The normalized project name a PEP 508 requirement string names.""" + match = re.match(r"[A-Za-z0-9._-]+", requirement.strip()) + assert match, f"not a requirement: {requirement!r}" + return re.sub(r"[-_.]+", "-", match.group(0)).lower() def _run_python(code: str) -> subprocess.CompletedProcess[str]: @@ -134,7 +165,7 @@ def test_importing_the_package_does_not_import_the_content_extra(): def test_the_package_works_without_the_content_extra(): - """With the three uninstalled, only the ``content`` subpackage refuses. + """With the extra uninstalled, only the ``content`` subpackage refuses. ``sys.modules[name] = None`` makes the next import of ``name`` raise ``ImportError``, which is what an environment without the extra does. @@ -166,18 +197,112 @@ def test_dev_and_docs_install_the_content_extra(): Without it in ``dev`` the converter's tests cannot import on CI, and without it in ``docs`` autodoc cannot import the module it documents. The - three requirements are written out in each extra, and this is what keeps - the copies in step. + requirements are written out in each extra, bounds included, and this is + what keeps the copies in step. """ - try: - import tomllib - except ModuleNotFoundError: # Python 3.10 - import tomli as tomllib - - config = tomllib.loads((_REPO_ROOT / "pyproject.toml").read_text(encoding="utf-8")) - extras = config["project"]["optional-dependencies"] + extras = _optional_dependencies() content = set(extras["content"]) assert content, "the content extra declares nothing" assert content <= set(extras["dev"]), sorted(content - set(extras["dev"])) assert content <= set(extras["docs"]), sorted(content - set(extras["docs"])) + + +def test_the_pre_commit_mypy_hook_pins_the_content_extra_as_pyproject_does(): + """The mypy hook's copies of the extra's requirements equal pyproject's. + + The hook builds its own venv from ``additional_dependencies``, so the + extra's typed packages are written out there a second time, and nothing + else keeps that copy in step. Read with a regex rather than a YAML parser, + which the test environment does not otherwise need. The hook lists only + the typed packages, so a package of the extra may be absent from it, but + one that is listed must carry the extra's exact bounds. + """ + config = (_REPO_ROOT / ".pre-commit-config.yaml").read_text(encoding="utf-8") + hook = re.search(r"- id: mypy\n(.*?)\n\s*exclude:", config, re.DOTALL) + assert hook, "no mypy hook found in .pre-commit-config.yaml" + listed = re.findall(r'^\s*-\s*"([^"]+)"\s*$', hook.group(1), re.MULTILINE) + extra = {_distribution_name(r): r for r in _optional_dependencies()["content"]} + + copies = { + _distribution_name(r): r for r in listed if _distribution_name(r) in extra + } + + assert copies, "the mypy hook lists none of the content extra's packages" + drifted = {name: (r, extra[name]) for name, r in copies.items() if r != extra[name]} + assert not drifted, f"hook requirement != pyproject's: {drifted}" + + +def test_the_supported_pythons_agree_everywhere_they_are_written(): + """requires-python, the classifiers and both CI matrices name one range. + + Each lists the supported interpreters by hand, and a floor raised in one + place and not the others is what this package's drop of 3.10 had to + chase through all four. + """ + config = tomllib.loads((_REPO_ROOT / "pyproject.toml").read_text(encoding="utf-8")) + project = config["project"] + prefix = "Programming Language :: Python :: 3." + classified = sorted( + int(c.removeprefix(prefix)) + for c in project["classifiers"] + if c.startswith(prefix) and c.removeprefix(prefix).isdecimal() + ) + floor = re.fullmatch(r">=3\.(\d+)", project["requires-python"]) + assert floor, project["requires-python"] + assert classified, "no Python 3.x classifier" + assert classified[0] == int(floor.group(1)) + assert classified == list(range(classified[0], classified[-1] + 1)) + + for workflow in ("ci.yml", "release.yml"): + text = (_REPO_ROOT / ".github" / "workflows" / workflow).read_text( + encoding="utf-8" + ) + matrices = re.findall(r"python-version:\s*\[([^\]]*)\]", text) + assert len(matrices) == 1, f"{workflow}: {len(matrices)} matrices" + versions = re.findall(r'"3\.(\d+)"', matrices[0]) + assert sorted(int(v) for v in versions) == classified, workflow + + +def test_the_content_import_names_match_the_extra(): + """``_CONTENT_MODULES`` names exactly the modules the extra installs. + + The two import-isolation tests above are only as good as that list: a + distribution added to the extra but missing here is never checked, and one + dropped from the extra but left here is checked for nothing -- which is how + ``markdown`` outlived the python-markdown converter it was listed for. So + the list is bound to pyproject.toml by distribution name. + """ + declared = {_distribution_name(r) for r in _optional_dependencies()["content"]} + + assert declared == set(_CONTENT_IMPORT_NAMES), ( + f"in pyproject only: {sorted(declared - set(_CONTENT_IMPORT_NAMES))}; " + f"listed here only: {sorted(set(_CONTENT_IMPORT_NAMES) - declared)}" + ) + + +@pytest.mark.parametrize("module", _CONTENT_MODULES) +def test_each_content_dependency_alone_triggers_the_install_hint(module): + """Any ONE missing distribution makes the subpackage name the extra. + + A partial install is the realistic failure -- a pinned environment that + predates a dependency the extra gained -- and it must fail at import with + the install command, not later with a bare ``ModuleNotFoundError`` from + inside a conversion. A fresh child per module rather than one child + blocking and unblocking in turn: un-caching a package does not un-cache its + submodules, so a re-import in the same process can fail for reasons that + have nothing to do with the guard. + """ + result = _run_python( + "import sys\n" + f"sys.modules[{module!r}] = None\n" + "try:\n" + " import easyvista_python_client.content\n" + "except ImportError as exc:\n" + " print(exc)\n" + "else:\n" + " raise SystemExit('the content subpackage imported without it')\n" + ) + + assert result.returncode == 0, result.stderr + assert 'pip install "easyvista-python-client[content]"' in result.stdout diff --git a/pyproject.toml b/pyproject.toml index 4e0e9b9..fdb23b0 100644 --- a/pyproject.toml +++ b/pyproject.toml @@ -81,21 +81,36 @@ Source = "https://github.com/baraline/easyvista_python_client" # The Markdown <-> memo HTML converter, `easyvista_python_client.content`. # Optional so the core stays httpx + pydantic + tenacity: nothing outside that # subpackage imports these, and importing it without them raises an -# ImportError naming `pip install "easyvista-python-client[content]"`. The same -# three floors as glpi_python_client, whose converter this one ports. They are +# ImportError naming `pip install "easyvista-python-client[content]"`. They are # repeated in `dev` (CI installs `.[dev]` and runs the converter's tests) and # in `docs` (Read the Docs installs `.[docs]`, and autodoc imports the module); # testing/test_public_api.py fails if a copy drifts. +# +# The upper bounds are deliberate: the reader extends markdownify's converters +# and mdformat's renderer, private surfaces that may move. markdown-it-py <4 and +# mdformat <0.8 were measured against the next major (2026-10-02): mdformat-tables +# 1.0.0 itself requires mdformat<0.8, which holds markdown-it-py below 4; the +# alt-text fix needs markdown-it-py 3's `text_special` tokens (2.x measured +# worse), and markdown-it-py 4 ends a ragged table early. markdownify <1.3 and +# mdformat-tables <1.1 are precautionary caps on releases that did not exist +# on 2026-10-02. Raise a bound only with the converter's tests re-run against +# the new version. content = [ - "beautifulsoup4>=4.12", - "markdown>=3.6", - "markdownify>=1.2", + "beautifulsoup4>=4.15", + "cmarkgfm>=2025.10.22", + "markdown-it-py>=3.0,<4", + "markdownify>=1.2.3,<1.3", + "mdformat>=0.7.22,<0.8", + "mdformat-tables>=1.0,<1.1", ] dev = [ - "beautifulsoup4>=4.12", - "markdown>=3.6", - "markdownify>=1.2", + "beautifulsoup4>=4.15", + "cmarkgfm>=2025.10.22", + "markdown-it-py>=3.0,<4", + "markdownify>=1.2.3,<1.3", + "mdformat>=0.7.22,<0.8", + "mdformat-tables>=1.0,<1.1", "build>=1.2", "mypy>=1.11", "numpydoc>=1.8", @@ -133,9 +148,12 @@ dev = [ ] docs = [ - "beautifulsoup4>=4.12", - "markdown>=3.6", - "markdownify>=1.2", + "beautifulsoup4>=4.15", + "cmarkgfm>=2025.10.22", + "markdown-it-py>=3.0,<4", + "markdownify>=1.2.3,<1.3", + "mdformat>=0.7.22,<0.8", + "mdformat-tables>=1.0,<1.1", "numpydoc>=1.8", "sphinx>=7.2,<8.2", "sphinx-rtd-theme>=2.0", From 4292f1657a521c7a6337b4eecd6334ad0b888d08 Mon Sep 17 00:00:00 2001 From: baraline Date: Fri, 2 Oct 2026 15:33:58 +0200 Subject: [PATCH 10/11] docs(content): document the rebuilt converter as measured, holes included docs/content.rst describes the CommonMark-with-GFM-tables dialect, what the reader escapes, and the dependency bounds. Two bounds were measured against the next major; two are precautionary caps on releases that do not exist yet. The round trip is stated as an aim with its measurement: on 367 memos read from one preproduction instance on 2026-10-01 and measured 2026-10-02 (tier 4, may not generalise), no word changed and all were fixed points; 65 displayed differently in listed, word-keeping ways. A "Known holes" list names the synthetic shapes that do lose words or display: ragged rows, a | in a cell's code or link, caption/colgroup without tbody, an image title holding ", a list or
    in a cell, and more. Also stated plainly: - plain_text_is_markdown=True reads Markdown that opens with < and holds HTML as HTML; - the converter is not a sanitiser; - a caller near the recursion limit can still get a bare RecursionError. Tier 1: the vendor's comment-log page (read 2026-10-02) sets a custom comment field to "Text area". That leans towards HTML display, against the converter's literal-lines reading of a tag-less memo, but says nothing of the built-in memos. O-MEMOFORMAT records it. The CHANGELOG's 0.4.0 entry matches, and says the 15 fixes are still to be proposed to glpi_python_client. README and the installation page list the extra's six packages. Co-Authored-By: Claude Opus 5.5 --- CHANGELOG.md | 304 +++++++++++---- README.md | 3 +- docs/content.rst | 452 ++++++++++++++++------ docs/installation.rst | 12 +- docs/vendor-api-reference.md | 53 ++- skills/easyvista-ticket-actions/SKILL.md | 18 +- skills/easyvista-ticket-workflow/SKILL.md | 22 +- 7 files changed, 647 insertions(+), 217 deletions(-) diff --git a/CHANGELOG.md b/CHANGELOG.md index eee6df9..b206869 100644 --- a/CHANGELOG.md +++ b/CHANGELOG.md @@ -15,102 +15,250 @@ is the error. Tags carry no `v` prefix. ## [Unreleased] -## [0.4.0] - 2026-09-30 +## [0.4.0] - 2026-10-02 -Adds Markdown <-> memo HTML conversion, as an optional extra. **Nothing moves -for a caller who does not install it**: the core package imports none of the -extra's dependencies, no existing call or model changes, and the one new name -at the package root is an exception class. The minor bump is for the new -public surface, not for a break. +Adds Markdown <-> memo HTML conversion as an optional extra, and drops Python +3.10. A `0.4.0` section was first prepared on 2026-09-30 around a different +converter. It was never tagged or uploaded, and this section replaces it; +`### Changed since the 2026-09-30 preparation` says what moved, for anyone who +built against that branch. + +**Upgrading.** One breaking change: **Python 3.10 is no longer supported** +(see `### Removed`). On 3.10, pip keeps resolving `0.3.0`, the last release +that installs there. On 3.11 and later, **nothing moves for a caller who does +not install the extra**: the core package imports none of its dependencies, no +existing call or model changes, and the one new name at the package root is an +exception class. ### Added - `easyvista_python_client.content.EasyvistaContentConverter`, behind the new - optional extra `easyvista-python-client[content]` (`beautifulsoup4>=4.12`, - `markdown>=3.6`, `markdownify>=1.2`). Two static methods: - `from_transport(value)` reads a memo -- rich-text HTML or plain text -- as - canonical Markdown, and `to_transport(value)` renders Markdown as the HTML a - memo is written with. - - It is a port of `glpi_python_client`'s `content/conversion.py` at `4fc3bed`, - the literal-safe converter, with the same rules, extensions and edge-case - handling, so with the same library versions the two produce the same - Markdown from the same HTML. **The two should move together.** The only code - divergence is the optional extra. - - **Text in a memo is literal, and the Markdown spells it so**: a character is - escaped exactly where python-markdown, with the four extensions - `to_transport` uses, would otherwise read it as syntax -- a typed `__init__` - reads as `\_\_init\_\_`, `\\serveur` as `\\\serveur`, `#4521` at a line start - as `\#4521`, `` as `<Entrée>` -- and nowhere else, so ordinary - prose such as `fichier_de_test_v2.xlsx` or `R&D` comes back as typed. - Rendering the Markdown displays what the memo displayed, and reading that - back gives the same Markdown: both are tested with an HTML parser over - realistic memos and a seeded fuzzer. Nested lists nest and keep their - numbers, ``` - goes out as a live ```` goes out as a live ``