Skip to content

fix(content)!: spell literal text so python-markdown reads it as text (0.6.0) - #39

Merged
baraline merged 8 commits into
mainfrom
fix/literal-safe-markdown
Oct 2, 2026
Merged

baraline merged 8 commits into
mainfrom
fix/literal-safe-markdown

Conversation

@baraline

@baraline baraline commented Oct 1, 2026 •

Copy link
Copy Markdown
Owner

Why

from_transport passed markdownify's output on as Markdown with escaping switched off. Literal text in a GLPI body therefore changed meaning as soon as python-markdown rendered it, which is what to_transport and any consumer does.

Measured through the reader into python-markdown, with the four extensions to_transport uses:

  • \\serveur\compta lost a backslash, and __init__ became bold;
  • a ----- line under text made a heading;
  • * point and > merci lines became a list and a quote;
  • #4521 at the start of a line became a heading;
  • [1]: https://… was consumed as a reference definition;
  • a | inside a table cell dropped the rest of the row;
  • nested lists flattened.

What changes

  • The reader (GlpiContentConverter.from_transport, and so every read model's .content) escapes a character only where python-markdown would read it as syntax:
    • with a backslash where python-markdown removes one;
    • with a character reference (<, &, =, ~) where no backslash escape exists.
  • What stays untouched:
    • ordinary prose comes back byte for byte, for example fichier_de_test_v2.xlsx, C:\Temp, R&D or a # in mid-sentence;
    • pasted URLs stay <https://…>;
    • code spans and code blocks are not escaped.
  • Structure:
    • nested lists indent by four spaces and keep their numbering;
    • blocks inside list items stay inside them;
    • <script>, <style> and <title> bodies are dropped, where they used to leak as prose;
    • obsolete elements (font, center, strike, …) are recognised;
    • <s> and <del> keep their words without a literal ~~.
  • Dependencies: markdownify>=1.2, because the subclass needs the 1.x hooks. Version 0.6.0.

Breaking for consumers

  • .content now carries backslash escapes and character references wherever literal text would otherwise be misread. Render it before displaying or indexing it.
  • The Markdown read from GLPI changes once for most e-mail-created bodies: new escapes, dropped CSS, and source newlines read as spaces. Stored digests of .content therefore change once.
  • A write-model value that holds one real HTML element is read as HTML throughout. Its newlines become spaces, and any Markdown syntax beside the element stays literal.

Verification

  • pytest -m "not integration" with coverage: 1569 passed, 17 xfailed, 97.6 % (the gate is 95 %).
  • ruff, mypy (strict), unasync_build.py --check and sphinx -W pass. Sphinx was built offline, with intersphinx disabled locally. The wheel was checked for test files.
  • Property tests use an HTML-parser oracle and check that to_transport(from_transport(html)) displays what html displays. 50,000 fuzzed bodies gave 0 failures.
  • An independent review re-ran the gates and tried about 3 million generated bodies. Its open findings are listed below.

Still to do before this leaves draft

These come from that review.

  • A table nested inside a table cell loses all its text, which is common in Outlook signatures. The bug predates this branch (0.5.0 does the same). The fix is to flatten an inner table to its cells' text.
  • A body with no HTML element is still returned verbatim, so its literal text is still misread. Spell it literally on the read path only; the write-model validator must stay verbatim.
  • Edge whitespace such as &nbsp; defeats the line rules in three places: the start of the first block, the start of a list item, and the end of the last block.
  • Two list shapes take quadratic time, for example thousands of empty items. The CHANGELOG's "every pass is linear" needs that fix to be true.
  • The setext-underline rule is copied from python-markdown 3.10.3, and 3.10.2 behaves differently. Derive the rule from the installed version, or raise the floor.
  • Five guards have no test that fails without them:
    • fences in separate paragraphs;
    • SRV___PROD___01;
    • leading whitespace before top-level syntax;
    • a titled pasted URL;
    • a _ b _ c.
  • Spell link destinations, link titles and the ticket-context transcript literally too.
  • Narrow the documented fixed-point guarantee: some e-mail shapes settle after one extra cycle, with no visible change.
  • Head the CHANGELOG section "Changed (breaking)".

The EasyVista port of this converter: baraline/easyvista_python_client#6

🤖 Generated with Claude Code

baraline and others added 4 commits October 1, 2026 00:48
from_transport handed markdownify's output on as Markdown, and markdownify
escapes nothing it was not asked to: text a user typed into GLPI came back
as syntax once any peer rendered it. Measured through this reader into
python-markdown with the four extensions to_transport uses: \\serveur lost
a backslash, __init__ and ______ became emphasis, a "-----------" line
under text made a heading, "* point" and "> merci" lines made a list and
a quote, "#4521" at a line start made a heading, "[1]: url" was consumed
as a reference definition, and a | in a cell dropped the rest of the row.

The converter is now a MarkdownConverter subclass. escape() escapes
nothing itself: it stands each context-dependent character in for with a
private-use code point, and each container python-markdown parses on its
own -- paragraph, div, list item, quote, heading, cell, document --
decides every stand-in in it at once, with its whole Markdown in view.
The rules replay python-markdown's own: its code-span pattern, its
escape set (asked of the renderer, not copied), its bracket counting,
its emphasis pairing including the NOT_STRONG exemption, its rule,
setext, fence, list, quote and reference-definition line tests, and its
e-mail autolink, which it matches only after code spans, escapes, links
and images are stashed. A character is escaped only where python-markdown
would misread it, with a backslash it removes again, or with a character
reference where no backslash works (< & = ~): ordinary prose --
fichier_de_test_v2.xlsx, C:\Temp, R&D, a # or a - mid-sentence -- comes
back byte for byte as before.

Structure markdownify garbled on the way through is fixed where literal
safety depended on it: continuation lines indented by four spaces, so
nested lists nest and keep their numbers; content after a nested list
starts a block of its own instead of joining the list's last item; a
quote in an item gets the blank line python-markdown needs, and a quote
opening an item keeps its later lines lazy; items python-markdown writes
loose are written loose, so the second read equals the first; a list on
a line with three bullets is spaced, because python-markdown cannot nest
anything under such a line; a <pre> in an item or quote is an indented
code block, except where python-markdown reads none, where its lines are
kept as literal text. Script, style and title bodies are dropped instead
of leaking into the text, <s>/<del>/<strike> keep their words without a
~~ python-markdown would display, images stay images in headings and
cells, and a source newline is a space, as HTML displays it. The HTML
standard's obsolete elements now make a body HTML.

The degraded path spells its stripped text through the same rules, so a
body too deep to convert is literal-safe too.

test_literal_text.py holds the property -- to_transport(from_transport(
html)) displays what html displays, compared with an HTML parser, and the
Markdown is a fixed point -- over the measured misreadings, one case per
construct, realistic bodies and a seeded fuzzer (8 seeds x 50 bodies in
the suite; 30,000 bodies across 600 seeds measured clean). What the
format cannot carry is an inventory of strict xfails, each asserted to
stabilise after one cycle. The nested-list round-trip loss is gone from
the corpus.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Four passes of the literal-text reader cost quadratic time on a body
dense with one character, measured on 20,000 repetitions: unclosed
brackets (settle_brackets counted forward from each [), words opening
with an underscore (the pairing check compared every opener with every
later run), backtick runs (_code_spans rebuilt the whole text after
every span) and a long numbered list (each bullet counted its previous
siblings, which markdownify's own converter did too -- 34 s before this
change). They took 8 to 117 seconds each and take well under one now:

- brackets are matched right to left against a stack of unclosed ],
  an escaped [ handing its ] back, which is python-markdown's own count;
- underscore pairing is one pass remembering whether an opener was seen;
- _code_spans searches on from each span's end, rebuilding the text only
  for the escaped-backslash run python-markdown consumes on its own;
- bullets are numbered once per list, and a list's shown items are
  counted once.

A complexity test holds each of the four under a ten-second budget.

Mutation testing, in a disposable worktree, over 65 mutants of the
reader's guards left eight alive before this commit. Six were test gaps,
now closed: a backtick touching a generated code span, an escaped
bracket in link text that python-markdown does not count, seven
underscores mid-sentence, asterisks in an item and its nested item,
a style block between a nested list and a code block, and a line break
in a cell or a heading. Two were dead code, now gone: the nested list's
blank line no longer checks for a following sibling -- the item strips
it when nothing follows -- and from_transport no longer strips a result
the document converter has already stripped, before settling it, which
is where the strip matters.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Literal-safe Markdown on the read path (588e97d, 7ea833a), and the
markdownify floor it needs: >=1.2, since the converter subclasses
MarkdownConverter and relies on the 1.x hooks -- escape(text,
parent_tags), convert_*(el, text, parent_tags), convert__document_ --
which 0.13 does not have. 1.2.2 and 1.2.3 are the versions measured.

The changelog records the fixes, the two behaviour changes a caller can
see (a value holding one real HTML element is read as HTML throughout,
write models included; Markdown read from GLPI changes once for any body
the old reader let through as syntax) and the inventory of what Markdown
cannot carry. The user guide and the API reference say what the escaping
is and is not, with an example checked against the code, and the two
skills that describe .content now say not to strip the backslashes.

A correction to 7ea833a's message: nine of the 65 mutants survived
before it, not eight, and seven were test gaps -- the seventh is alt
text over several lines, which a blank line in it split out of its
paragraph once the collapse was removed.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
The deepest case was 300 levels. CPython 3.10, which CI runs, spends about
three frames per level where 3.12 spends two, so its cliff is about 328
levels from a shallow stack -- measured by easyvista_python_client's port
of this converter, 2026-09-30 -- and 300 left 14 levels of margin under
pytest. 250 still proves the point, being past the old fixed bound of 200.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
@codecov-commenter

Copy link
Copy Markdown

⚠️ Please install the 'codecov app svg image' to ensure uploads and comments are reliably processed by Codecov.

Codecov Report

❌ Patch coverage is 97.12230% with 8 lines in your changes missing coverage. Please review.
✅ Project coverage is 97.64%. Comparing base (f93411d) to head (524304a).

Files with missing lines Patch % Lines
glpi_python_client/content/conversion.py 96.91% 8 Missing ⚠️
❗ Your organization needs to install the Codecov GitHub app to enable full functionality.
Additional details and impacted files
@@            Coverage Diff             @@
##             main      #39      +/-   ##
==========================================
+ Coverage   97.59%   97.64%   +0.04%     
==========================================
  Files          90       90              
  Lines        3163     3313     +150     
==========================================
+ Hits         3087     3235     +148     
- Misses         76       78       +2     

☔ View full report in Codecov by Harness.
📢 Have feedback on the report? Share it here.

🚀 New features to boost your workflow:
  • ❄️ Test Analytics: Detect flaky tests, report on failures, and find test suite problems.

@baraline
baraline merged commit 8cb46fc into main Oct 2, 2026
7 checks passed
@baraline
baraline deleted the fix/literal-safe-markdown branch October 2, 2026 10:28
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants