Skip to content

fix(content): port easyvista-python-client's reader fixes, and keep blocks inside u/mark/ins (0.6.1) - #40

Merged
baraline merged 16 commits into
mainfrom
fix/port-easyvista-reader-fixes
Oct 3, 2026
Merged

baraline merged 16 commits into
mainfrom
fix/port-easyvista-reader-fixes

Conversation

@baraline

@baraline baraline commented Oct 2, 2026 •

Copy link
Copy Markdown
Owner

Why

easyvista-python-client carries this package's 0.6.0 converter plus fifteen reader fixes, so the same HTML could read as different Markdown through the two packages. Some of those differences change what a body shows:

  • image alt text lost its escapes;
  • a "!" before a link turned it into an image;
  • , and were dropped;
  • a line opening with ~~~ became a code fence;
  • some shapes read in quadratic time.

A colspan or start the converter could not read as a number failed the whole body.

A review of the port found one more problem, which this branch also fixes. Fix 11 kept , and raw even around blocks, so a table inside one read as pipe text, a list or a heading read as Markdown source, and

...
left a fence open over the rest of the body. Separately, a lone surrogate in a write model's Markdown failed the whole create or update.

What

  • Reader (content/conversion.py): the fifteen fixes, plus a correction to fix 11. A , or holding a block is dropped and its blocks kept, as 0.6.0 read them; inside a table cell or a heading it stays. The module now differs from easyvista-python-client 0.4.1's only in names, error messages, docstrings and that package's optional-dependency import guard. The converter's writing is unchanged.
  • Write models: a lone surrogate is sent as U+FFFD, as CommonMark replaces a NUL; the rest of the body keeps its Markdown. GlpiContentConverter.to_transport called directly still raises.
  • Tests:
    • one regression section per fix;
    • the fix-11 correction, for every block kind, the stray fence, nested tags and the one-line controls;
    • one test each for four guards no test caught (a second "<" before a space, a header cell, an anchor without href, an item numbered 10 or more);
    • a count of the block walk's checks;
    • growth tests replacing the 20-second budgets;
    • seeded property tests;
    • the surrogate write path.
  • Dependency bounds aligned with easyvista-python-client[content].
  • Docs:
    • the user guide's rich-text section: what each element becomes, what survives a round trip, the known holes;
    • corrected claims: start values, any ValueError degrading to text, the write-model rule (stripped, not verbatim), the skills' depth notes, and the text shown after a body's last ">".

Verification

  • Every fix's tests fail on the converter before it, and reverting any one guard fails at least one test (mutation runs in disposable copies). The fix-11 correction's tests fail 40 of 42 on easyvista-python-client 0.4.0's converter.
  • Parity:
    • The two conversion modules' ASTs are identical once names, messages, docstrings and the import guard are set aside.
    • Differential: 22,623 inputs (both packages' recorded test bodies plus a seeded fuzz corpus) through both converters with the same libraries, 6 operations each: 135,738 comparisons, 0 mismatches.
    • Against the previous head, only reads of bodies holding , or change. None that displayed the same and was a fixed point stopped being either, and 200 that were wrong became right.
  • Gates:
    • pytest -m "not integration" with coverage: 1761 passed, 97.87%;
    • ruff check, ruff format --check, mypy strict, unasync_build.py --check;
    • sphinx -E -W, built offline.
  • Every documented behaviour was re-checked on CPython 3.12 and 3.13.

Known follow-ups (from the review; all Minor)

  • , and round blocks remain a documented hole, unchanged since 0.6.0.
  • The converter reads start with isdecimal; a browser reads " 3", "+3" or "3abc" as 3. This is documented.
  • A lone surrogate in a non-content field still fails in the HTTP library's JSON encoding.
  • The cmarkgfm floor differs from easyvista-python-client's (writer only).
  • An inline wrapper (<span>, <font>, and now a dropped <u>/<mark>/<ins>) round a list, followed by inline text, reads that text into the list's last item. This is markdownify's general behaviour, the same on every tree; the fix belongs in the list converter's next-sibling handling.
  • A data-glpi-block / data-ev-block attribute already present in stored HTML is trusted by the block walk. Only hand-made or sanitiser-kept HTML carries one; clear it on read.
  • No test pins that the block walk runs after nested tables are flattened (formatting-only effect on a rare shape).
  • The documented exceptions to the second-cycle fixed point are still too narrow: two adjacent lists separated by anything that reads as nothing (a comment, an empty <div>) also change on the second cycle.
  • The start-value docs omit a negative start, a full-width digit, and a start of 10 or more digits, which reads as a paragraph.

Twin PR: baraline/easyvista_python_client#8

🤖 Generated with Claude Code

baraline and others added 16 commits October 2, 2026 17:01
easyvista-python-client 0.4.0 carries this package's converter plus 15
reader fixes. Porting those fixes starts with the oracle their tests are
written against, kept identical but for the package's names:

- inside one line (a cell, a heading, a link), the edge of a block or of
  a nested table's cell is a word boundary, as a browser shows it and as
  the converter writes it (a space); the 0.6.0 oracle joined the words;
- an ordered list's start is read with isdecimal, as the converter will
  read it, where isdigit made the oracle itself raise on "²";
- one_line, text_words and loose are new: the display a table inside a
  link owes, the displayed words without formatting, and a comparison
  loose on percent-encoded URLs.

The existing content tests pass unchanged against it.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Ports fixes 1 to 5 of easyvista-python-client 0.4.0's reader, which is
this converter plus 15 fixes:

1. An escape or character reference in an image's alt text is kept.
   markdown-it-py 3 leaves it a text_special token, which mdformat
   rendered as nothing: alt "a_b_c.png" read back as "abc.png".
2. A paragraph line opening with ~~~ is escaped; CommonMark read it as a
   code fence, which swallowed the rest of the body.
3. A "!" ending a text right before a link is escaped; "Attention!" then
   a link read back as an image.
4. An alt text opening with "^" is escaped; cmark-gfm never opens an image
   on "![^".
5. A "<" mdformat left bare after an escaped one is escaped; "<<a@b.c>"
   read back as "\<" and an autolink.

Tests first: test_fixes.py sections 1 to 5, ported with the package's
names; 13 of their 17 tests failed before the change (the other 4 are
controls), all pass after it.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Ports fixes 6 to 11 of easyvista-python-client 0.4.0's reader:

6. A <br> inside <code>, <kbd> or <samp> splits the span: a line break
   in a paragraph, a <br> in a cell, a space in a heading. It read back
   as a literal backslash inside one span. A "|" in <kbd> or <samp> in a
   cell is escaped, as one in <code> already was.
7. A table inside a heading or a link is written as its cells' text, as
   one inside a cell already was; its pipes showed as text, and inside a
   link the reading was not a fixed point.
8. A block inside such a flattened table stays on its holder's line.
9. A flattened cell's edge counts as a space when deciding whether "**"
   closes, so bold ending on punctuation there stays Markdown bold.
10. <center> is a block, as <div> is, except inside <pre>; two <center>
    blocks ran together.
11. <u>, <mark> and <ins> are kept as raw tags; the formatting was dropped.

Tests first: test_fixes.py sections 6 to 11; 18 of their 21 tests failed
before the change (the other 3 are controls or guard a correction to a
fix), all pass after it.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…de bad numbers

Ports fixes 12 to 15 of easyvista-python-client 0.4.0's reader:

12. Ordered items are numbered in linear time: each item takes its number
    from the item before it, where markdownify counted every earlier
    sibling. start is read with isdecimal, so start="²" or "½" counts from
    1 as a browser does instead of failing the body.
13. The pattern that splits emphasis and links into edges and content is
    greedy, so linear on runs of spaces or <br>; the lazy one was
    quadratic.
14. A colspan or start markdownify cannot read as a number ("²", or more
    digits than CPython converts) degrades the body to its text, as a
    body too deep for the stack does, instead of raising GlpiContentError.
15. Every "<" after a body's last ">" is read as text. CPython's
    html.parser before 3.11.14, 3.12.12 and 3.13.6 is quadratic on a run
    of unfinished tags there (CVE-2025-6069); patched releases dropped
    the unfinished tag, so "x <a b" now reads the same on every release.

Tests first. test_fixes.py is now complete: on 0.6.0, 36 of its 54
tests failed (CPython 3.12.3); all pass. test_cost.py replaces the four
20-second long-body budgets of test_round_trip.py with growth tests,
time(4n)/time(n) < 8: on 3.12.3 this converter reads 3.0 to 7.5 and
0.6.0 read 10.3 (ordered list), 14.2 (bold holding line breaks) and 34.4
(unfinished-tag tail). test_conversion.py gains the ValueError retry
tests; its inbound-fault test now injects a RuntimeError, since a
ValueError takes the text fallback first.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…d-trip tests

- test_properties.py: the seeded generators that package verified the 15
  reader fixes with, each body held to three things: it displays what the
  HTML displayed, it is a fixed point, and it keeps its words in order.
  Families: image alt and link text, "~~~" at a line start, "!" before a
  link or image, inline code holding line breaks, tables nested in a
  cell, a heading or a link, raw <u>/<mark>/<ins> in every holder, and
  whole generated documents. On 0.6.0 all 16 tests failed, 39 to 100
  percent of each test's bodies (CPython 3.12.3); none fails now.
- test_round_trip.py: that package's earlier regression bodies (HTML
  only, held to display plus fixed point), an e-mail notification
  template laid out as nested tables, a CR LF and an entity in plain
  text, and a pin on what the write models' validator reads as HTML.
- test_conversion.py: the converter is not a sanitiser, pinned exactly in
  both directions, as the user guide will say.

Every body is synthetic and every URL is under example.org. Reverting any
one of the 15 fixes alone, in a disposable copy, fails at least one test
of this suite (CPython 3.12.3; fix 15's functional case needs a patched
interpreter and was confirmed on 3.13.14). The content suite runs in
about 21 seconds.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…libraries

The ported reader fixes rely on private surfaces of markdownify and
mdformat, and the two packages' readers should resolve to the same
libraries when installed side by side. The bounds are now those of
easyvista-python-client[content] 0.4.0, each explained in pyproject.toml:

- markdown-it-py>=3.0,<4: the alt-text fix relies on 3.x's text_special
  tokens; 4.x ends a ragged table early.
- markdownify>=1.2.3,<1.3: the overrides use process_tag, escape and the
  _inline/_noformat pseudo-tags; 1.2.3 is the release the fixes were
  measured with, and the cap is a precaution.
- mdformat>=0.7.22,<0.8: unchanged; mdformat-tables 1.0 requires it too.
- mdformat-tables>=1.0,<1.1: the cap is a precaution.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
25a014a's comments stated what this package cannot check from its own
repository: that markdownify 1.2.3 is the release the fixes were measured
with, and that markdownify 1.3 and mdformat-tables 1.1 did not exist on
2026-10-02. They now say what can be checked:

- mdformat 0.7.22 itself requires markdown-it-py<4 (its metadata), and
  easyvista-python-client measured markdown-it-py 4 ending a ragged table
  early;
- 1.2.3 is easyvista-python-client's markdownify floor, and this
  package's content tests pass on 1.2.2 (CPython 3.12.3) and 1.2.3
  (3.13.14);
- the markdownify and mdformat-tables caps are precautions.

The bounds themselves are unchanged.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
96847df's docstring figures were read while other runs shared the
machine. Re-measured 2026-10-02 with growth() itself, three readings per
shape, nothing else running:

- this converter: 3.8 to 5.4 on CPython 3.12.3, 2.9 to 4.5 on 3.13.14
  (the limit is 8);
- 0.6.0 on 3.12.3: 10.9 to 11.3 (ordered list), 13.0 to 17.1 (bold
  holding line breaks), 28.5 to 31.1 (unfinished-tag tail);
- reverting fix 12, 13 or 15 alone: 10.4 to 12.0, 16.8 to 17.4 and 31.9
  to 33.0 on its shape;
- 5,000 ordered items, best of three: 3.1 s on 0.6.0, 0.6 s now.

All within the ranges the docstring gave; only the figures change.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
The user guide, the API reference, the module docstring and three skills
now describe the reader as 0.6.1 has it. Every behaviour stated was
reproduced on this branch on 2026-10-02.

- User guide, rich-text section: the escape table, what each element
  becomes (raw <u>/<mark>/<ins>/<s> included), that neither direction
  sanitises, what survives a round trip, the structures Markdown cannot
  spell, the known holes, and the CVE-2025-6069 guard. The statement
  that a read always displays the same and reads back as itself, which
  0.6.0 made unqualified, is now the aim it is, there, in the API
  reference and in the module docstring.
- The deep-nesting note gives measured depths: about 326 <div> levels
  from a shallow stack (0.6.0: about 490; the reader spends one more
  frame per level) and 194 <blockquote>, on CPython 3.12.3 and 3.13.14.
  It no longer says a read never raises: started within about 30 frames
  of the recursion limit, it raises GlpiContentError, then a bare
  RecursionError (measured: from about 975 and 994 frames).
- The write-model rule is stated exactly: Markdown passes through unless
  it starts with "<" and holds an HTML element anywhere, so an opening
  autolink followed by inline HTML is lost.
- from_transport documents that RecursionError; GlpiContentError's
  docstring named python-markdown as the writer, which is cmark-gfm.
- The installation page names the converter's libraries.
  development.md pointed at a module docstring that no longer holds the
  depth derivation; the 0.5.0 changelog entry does.
- Skills glpi-ticket-timeline, glpi-ticket-workflow and
  glpi-knowledge-base: raw tags, the write-model rule, the sanitiser
  caveat and the new depth.

Found while checking the guide: the claim, which easyvista-python-client's
docs also make, that a second write-and-read cycle of Markdown changes
nothing more failed on 23 of 4,500 generated bodies -- adjacent lists
with different bullets, and a fence whose info string holds a character
reference. The guide lists both.

The conversion module still differs from easyvista-python-client 0.4.0's
only in names, messages, docstrings and its import guard (AST comparison).

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
The reader of easyvista-python-client 0.4.0, ported: its fifteen fixes,
their regression, property and growth tests, its display oracle and its
dependency bounds. The two conversion modules now differ only in names,
error messages, docstrings and that package's optional-dependency import
guard, and gave the same Markdown and the same HTML on 20,443 inputs.

The changelog lists one entry per fix, effect first, and what changes
for a caller: .content changes once for the shapes fixed, nested HTML
degrades to text at about 326 <div> levels instead of about 490, and
the converter's libraries carry easyvista-python-client's bounds.

Version 0.6.1 in pyproject.toml, __version__ and all eleven skills.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…ocks

Fix 11 kept these tags as raw HTML around whatever they held. Markdown
has no inline tag round blocks, so a table inside one read as pipe text,
a list, a heading, a quote or a rule as its Markdown source, paragraphs
were not a fixed point, and <u><pre>code</pre></u> left a fence open
that turned the rest of the body into code.

A walk marks each such tag that holds a block, innermost first, stopping
at one already marked, so each is marked once however deep they nest.
convert_u drops a marked tag and keeps its blocks, as 0.6.0 read them.
Inside a table cell or a heading, whose blocks are written on one line,
the tag stays. The tag is kept around inline content as before.

Tests: 42 in test_fixes section 11, one per block kind and tag, the
stray fence, nested tags and the one-line controls. Against the
unfixed converter 40 fail; the 2 that pass are the controls. The
display oracle now counts an <hr> inside an inline element as a rule.

Docs: the changelog entry, the user guide's element list and a known
hole for <b>, <em> and <s> round blocks (unchanged since 0.6.0), and two
skills.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
A review removed each sub-guard of the ported fixes in turn and found
four whose removal no test noticed. Each now has a test that fails
without it, run in disposable copies against the whole content suite:

- fix 5: before a space a second "<" keeps mdformat's single escape
  ("<< b" reads as "\<< b"); spelling only, but digests depend on it;
- fix 6: a header cell (th) is a cell, so a "|" in <kbd> or <samp>
  stays escaped and a line break in inline code splits the span;
- fix 7: an <a> without href is no link, so a table in it stays a table;
- fix 12: an ordered item numbered 10 or more indents its content by its
  bullet's width (four for "10. "), where a fixed three let a nested
  list or a second paragraph leave the item.

test_cost also counts the checks of the walk that finds blocks inside
<u>, <mark> and <ins>: the stop at a tag already marked changes no
output, only cost, and a growth ratio at affordable sizes could not tell
the two apart, so the count does (149 checks for 50 levels, against
2,550 without the stop).

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
cmark-gfm renders UTF-8, which cannot encode a lone UTF-16 surrogate, so
a write model whose Markdown held one failed the whole create or update
with GlpiContentError. Python strings can hold one: decoding with
surrogateescape, or json.loads of an unpaired escape. easyvista-python-
client's converter raises the same way, and the caller relaying to it
answers with U+FFFD; here the conversion runs inside the write model's
serialiser, out of the caller's reach, so the model does it.

The serialiser replaces each surrogate code point with U+FFFD before
rendering, as CommonMark replaces a NUL, so the rest of the body keeps
its Markdown. The shared conversion module is unchanged:
GlpiContentConverter.to_transport called directly still raises.

Tests: the payload of a model holding a high, a low, and a pair kept as
two code points, and a create through the client. All four failed with
GlpiContentError before the change.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
- start: a start that is not a decimal number counts from 1, as a
  browser counts one holding no digit; a browser reads " 3", "+3" or
  "3abc" as 3 (the HTML standard's integer rules), which the converter
  counts from 1. The display oracle reads start as the converter does,
  and now says so.
- ValueError: any ValueError raised while converting degrades the body
  to text, not only an unreadable colspan or start. The module and
  from_transport docstrings, the user guide, the API reference,
  development.md and three skills say so.
- The knowledge-base and plugin-fields skills no longer say a deep body
  never raises: a read started within about 30 frames of the recursion
  limit can, as the timeline skill and the user guide already said.
- Write models keep the caller's Markdown stripped at both ends, not
  verbatim; stripping unindents a body that opens with an indented code
  block, so the guide says to open such a body with a fence.
- CVE-2025-6069 guard: whatever follows a body's last ">" reads as text,
  an unterminated comment and a closing tag cut short inside a link
  included, where a patched CPython drops some of it. Test 15 pins both.

Each stated behaviour was re-checked on CPython 3.12.3 and 3.13.14.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
The correction to fix 11 lands in both packages, as glpi-python-client
0.6.1 and easyvista-python-client 0.4.1, so the statements that the
other package's converter "is this one" now name 0.4.1; provenance
("ported from 0.4.0") stays as it was. The 0.6.1 changelog says the
reader matches 0.4.1, that the converter's writing is unchanged while a
write model sends a lone surrogate as U+FFFD, and how many of
test_fixes's tests fail on which converter.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
The replacement character went into the source and its tests as the
literal character, which reads as an encoding error in a diff. It is now
the escape �; the value is the same, and so is every test. A test
comment that meant the escaped pair 😀 had become the emoji
itself; it now says what it meant.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
@codecov-commenter

Copy link
Copy Markdown

⚠️ Please install the 'codecov app svg image' to ensure uploads and comments are reliably processed by Codecov.

Codecov Report

✅ All modified and coverable lines are covered by tests.
✅ Project coverage is 97.86%. Comparing base (8cb46fc) to head (655d153).
❗ Your organization needs to install the Codecov GitHub app to enable full functionality.

Additional details and impacted files
@@            Coverage Diff             @@
##             main      #40      +/-   ##
==========================================
+ Coverage   97.64%   97.86%   +0.22%     
==========================================
  Files          90       90              
  Lines        3313     3377      +64     
==========================================
+ Hits         3235     3305      +70     
+ Misses         78       72       -6     

☔ View full report in Codecov by Harness.
📢 Have feedback on the report? Share it here.

🚀 New features to boost your workflow:
  • ❄️ Test Analytics: Detect flaky tests, report on failures, and find test suite problems.

@baraline
baraline marked this pull request as ready for review October 3, 2026 07:56
@baraline
baraline merged commit a9e60dd into main Oct 3, 2026
7 checks passed
@baraline
baraline deleted the fix/port-easyvista-reader-fixes branch October 3, 2026 07:59
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants