fix(content): port easyvista-python-client's reader fixes, and keep blocks inside u/mark/ins (0.6.1) - #40
Merged
Conversation
easyvista-python-client 0.4.0 carries this package's converter plus 15 reader fixes. Porting those fixes starts with the oracle their tests are written against, kept identical but for the package's names: - inside one line (a cell, a heading, a link), the edge of a block or of a nested table's cell is a word boundary, as a browser shows it and as the converter writes it (a space); the 0.6.0 oracle joined the words; - an ordered list's start is read with isdecimal, as the converter will read it, where isdigit made the oracle itself raise on "²"; - one_line, text_words and loose are new: the display a table inside a link owes, the displayed words without formatting, and a comparison loose on percent-encoded URLs. The existing content tests pass unchanged against it. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Ports fixes 1 to 5 of easyvista-python-client 0.4.0's reader, which is this converter plus 15 fixes: 1. An escape or character reference in an image's alt text is kept. markdown-it-py 3 leaves it a text_special token, which mdformat rendered as nothing: alt "a_b_c.png" read back as "abc.png". 2. A paragraph line opening with ~~~ is escaped; CommonMark read it as a code fence, which swallowed the rest of the body. 3. A "!" ending a text right before a link is escaped; "Attention!" then a link read back as an image. 4. An alt text opening with "^" is escaped; cmark-gfm never opens an image on "![^". 5. A "<" mdformat left bare after an escaped one is escaped; "<<a@b.c>" read back as "\<" and an autolink. Tests first: test_fixes.py sections 1 to 5, ported with the package's names; 13 of their 17 tests failed before the change (the other 4 are controls), all pass after it. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Ports fixes 6 to 11 of easyvista-python-client 0.4.0's reader:
6. A <br> inside <code>, <kbd> or <samp> splits the span: a line break
in a paragraph, a <br> in a cell, a space in a heading. It read back
as a literal backslash inside one span. A "|" in <kbd> or <samp> in a
cell is escaped, as one in <code> already was.
7. A table inside a heading or a link is written as its cells' text, as
one inside a cell already was; its pipes showed as text, and inside a
link the reading was not a fixed point.
8. A block inside such a flattened table stays on its holder's line.
9. A flattened cell's edge counts as a space when deciding whether "**"
closes, so bold ending on punctuation there stays Markdown bold.
10. <center> is a block, as <div> is, except inside <pre>; two <center>
blocks ran together.
11. <u>, <mark> and <ins> are kept as raw tags; the formatting was dropped.
Tests first: test_fixes.py sections 6 to 11; 18 of their 21 tests failed
before the change (the other 3 are controls or guard a correction to a
fix), all pass after it.
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…de bad numbers
Ports fixes 12 to 15 of easyvista-python-client 0.4.0's reader:
12. Ordered items are numbered in linear time: each item takes its number
from the item before it, where markdownify counted every earlier
sibling. start is read with isdecimal, so start="²" or "½" counts from
1 as a browser does instead of failing the body.
13. The pattern that splits emphasis and links into edges and content is
greedy, so linear on runs of spaces or <br>; the lazy one was
quadratic.
14. A colspan or start markdownify cannot read as a number ("²", or more
digits than CPython converts) degrades the body to its text, as a
body too deep for the stack does, instead of raising GlpiContentError.
15. Every "<" after a body's last ">" is read as text. CPython's
html.parser before 3.11.14, 3.12.12 and 3.13.6 is quadratic on a run
of unfinished tags there (CVE-2025-6069); patched releases dropped
the unfinished tag, so "x <a b" now reads the same on every release.
Tests first. test_fixes.py is now complete: on 0.6.0, 36 of its 54
tests failed (CPython 3.12.3); all pass. test_cost.py replaces the four
20-second long-body budgets of test_round_trip.py with growth tests,
time(4n)/time(n) < 8: on 3.12.3 this converter reads 3.0 to 7.5 and
0.6.0 read 10.3 (ordered list), 14.2 (bold holding line breaks) and 34.4
(unfinished-tag tail). test_conversion.py gains the ValueError retry
tests; its inbound-fault test now injects a RuntimeError, since a
ValueError takes the text fallback first.
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…d-trip tests - test_properties.py: the seeded generators that package verified the 15 reader fixes with, each body held to three things: it displays what the HTML displayed, it is a fixed point, and it keeps its words in order. Families: image alt and link text, "~~~" at a line start, "!" before a link or image, inline code holding line breaks, tables nested in a cell, a heading or a link, raw <u>/<mark>/<ins> in every holder, and whole generated documents. On 0.6.0 all 16 tests failed, 39 to 100 percent of each test's bodies (CPython 3.12.3); none fails now. - test_round_trip.py: that package's earlier regression bodies (HTML only, held to display plus fixed point), an e-mail notification template laid out as nested tables, a CR LF and an entity in plain text, and a pin on what the write models' validator reads as HTML. - test_conversion.py: the converter is not a sanitiser, pinned exactly in both directions, as the user guide will say. Every body is synthetic and every URL is under example.org. Reverting any one of the 15 fixes alone, in a disposable copy, fails at least one test of this suite (CPython 3.12.3; fix 15's functional case needs a patched interpreter and was confirmed on 3.13.14). The content suite runs in about 21 seconds. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…libraries The ported reader fixes rely on private surfaces of markdownify and mdformat, and the two packages' readers should resolve to the same libraries when installed side by side. The bounds are now those of easyvista-python-client[content] 0.4.0, each explained in pyproject.toml: - markdown-it-py>=3.0,<4: the alt-text fix relies on 3.x's text_special tokens; 4.x ends a ragged table early. - markdownify>=1.2.3,<1.3: the overrides use process_tag, escape and the _inline/_noformat pseudo-tags; 1.2.3 is the release the fixes were measured with, and the cap is a precaution. - mdformat>=0.7.22,<0.8: unchanged; mdformat-tables 1.0 requires it too. - mdformat-tables>=1.0,<1.1: the cap is a precaution. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
25a014a's comments stated what this package cannot check from its own repository: that markdownify 1.2.3 is the release the fixes were measured with, and that markdownify 1.3 and mdformat-tables 1.1 did not exist on 2026-10-02. They now say what can be checked: - mdformat 0.7.22 itself requires markdown-it-py<4 (its metadata), and easyvista-python-client measured markdown-it-py 4 ending a ragged table early; - 1.2.3 is easyvista-python-client's markdownify floor, and this package's content tests pass on 1.2.2 (CPython 3.12.3) and 1.2.3 (3.13.14); - the markdownify and mdformat-tables caps are precautions. The bounds themselves are unchanged. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
96847df's docstring figures were read while other runs shared the machine. Re-measured 2026-10-02 with growth() itself, three readings per shape, nothing else running: - this converter: 3.8 to 5.4 on CPython 3.12.3, 2.9 to 4.5 on 3.13.14 (the limit is 8); - 0.6.0 on 3.12.3: 10.9 to 11.3 (ordered list), 13.0 to 17.1 (bold holding line breaks), 28.5 to 31.1 (unfinished-tag tail); - reverting fix 12, 13 or 15 alone: 10.4 to 12.0, 16.8 to 17.4 and 31.9 to 33.0 on its shape; - 5,000 ordered items, best of three: 3.1 s on 0.6.0, 0.6 s now. All within the ranges the docstring gave; only the figures change. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
The user guide, the API reference, the module docstring and three skills now describe the reader as 0.6.1 has it. Every behaviour stated was reproduced on this branch on 2026-10-02. - User guide, rich-text section: the escape table, what each element becomes (raw <u>/<mark>/<ins>/<s> included), that neither direction sanitises, what survives a round trip, the structures Markdown cannot spell, the known holes, and the CVE-2025-6069 guard. The statement that a read always displays the same and reads back as itself, which 0.6.0 made unqualified, is now the aim it is, there, in the API reference and in the module docstring. - The deep-nesting note gives measured depths: about 326 <div> levels from a shallow stack (0.6.0: about 490; the reader spends one more frame per level) and 194 <blockquote>, on CPython 3.12.3 and 3.13.14. It no longer says a read never raises: started within about 30 frames of the recursion limit, it raises GlpiContentError, then a bare RecursionError (measured: from about 975 and 994 frames). - The write-model rule is stated exactly: Markdown passes through unless it starts with "<" and holds an HTML element anywhere, so an opening autolink followed by inline HTML is lost. - from_transport documents that RecursionError; GlpiContentError's docstring named python-markdown as the writer, which is cmark-gfm. - The installation page names the converter's libraries. development.md pointed at a module docstring that no longer holds the depth derivation; the 0.5.0 changelog entry does. - Skills glpi-ticket-timeline, glpi-ticket-workflow and glpi-knowledge-base: raw tags, the write-model rule, the sanitiser caveat and the new depth. Found while checking the guide: the claim, which easyvista-python-client's docs also make, that a second write-and-read cycle of Markdown changes nothing more failed on 23 of 4,500 generated bodies -- adjacent lists with different bullets, and a fence whose info string holds a character reference. The guide lists both. The conversion module still differs from easyvista-python-client 0.4.0's only in names, messages, docstrings and its import guard (AST comparison). Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
The reader of easyvista-python-client 0.4.0, ported: its fifteen fixes, their regression, property and growth tests, its display oracle and its dependency bounds. The two conversion modules now differ only in names, error messages, docstrings and that package's optional-dependency import guard, and gave the same Markdown and the same HTML on 20,443 inputs. The changelog lists one entry per fix, effect first, and what changes for a caller: .content changes once for the shapes fixed, nested HTML degrades to text at about 326 <div> levels instead of about 490, and the converter's libraries carry easyvista-python-client's bounds. Version 0.6.1 in pyproject.toml, __version__ and all eleven skills. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…ocks Fix 11 kept these tags as raw HTML around whatever they held. Markdown has no inline tag round blocks, so a table inside one read as pipe text, a list, a heading, a quote or a rule as its Markdown source, paragraphs were not a fixed point, and <u><pre>code</pre></u> left a fence open that turned the rest of the body into code. A walk marks each such tag that holds a block, innermost first, stopping at one already marked, so each is marked once however deep they nest. convert_u drops a marked tag and keeps its blocks, as 0.6.0 read them. Inside a table cell or a heading, whose blocks are written on one line, the tag stays. The tag is kept around inline content as before. Tests: 42 in test_fixes section 11, one per block kind and tag, the stray fence, nested tags and the one-line controls. Against the unfixed converter 40 fail; the 2 that pass are the controls. The display oracle now counts an <hr> inside an inline element as a rule. Docs: the changelog entry, the user guide's element list and a known hole for <b>, <em> and <s> round blocks (unchanged since 0.6.0), and two skills. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
A review removed each sub-guard of the ported fixes in turn and found
four whose removal no test noticed. Each now has a test that fails
without it, run in disposable copies against the whole content suite:
- fix 5: before a space a second "<" keeps mdformat's single escape
("<< b" reads as "\<< b"); spelling only, but digests depend on it;
- fix 6: a header cell (th) is a cell, so a "|" in <kbd> or <samp>
stays escaped and a line break in inline code splits the span;
- fix 7: an <a> without href is no link, so a table in it stays a table;
- fix 12: an ordered item numbered 10 or more indents its content by its
bullet's width (four for "10. "), where a fixed three let a nested
list or a second paragraph leave the item.
test_cost also counts the checks of the walk that finds blocks inside
<u>, <mark> and <ins>: the stop at a tag already marked changes no
output, only cost, and a growth ratio at affordable sizes could not tell
the two apart, so the count does (149 checks for 50 levels, against
2,550 without the stop).
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
cmark-gfm renders UTF-8, which cannot encode a lone UTF-16 surrogate, so a write model whose Markdown held one failed the whole create or update with GlpiContentError. Python strings can hold one: decoding with surrogateescape, or json.loads of an unpaired escape. easyvista-python- client's converter raises the same way, and the caller relaying to it answers with U+FFFD; here the conversion runs inside the write model's serialiser, out of the caller's reach, so the model does it. The serialiser replaces each surrogate code point with U+FFFD before rendering, as CommonMark replaces a NUL, so the rest of the body keeps its Markdown. The shared conversion module is unchanged: GlpiContentConverter.to_transport called directly still raises. Tests: the payload of a model holding a high, a low, and a pair kept as two code points, and a create through the client. All four failed with GlpiContentError before the change. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
- start: a start that is not a decimal number counts from 1, as a browser counts one holding no digit; a browser reads " 3", "+3" or "3abc" as 3 (the HTML standard's integer rules), which the converter counts from 1. The display oracle reads start as the converter does, and now says so. - ValueError: any ValueError raised while converting degrades the body to text, not only an unreadable colspan or start. The module and from_transport docstrings, the user guide, the API reference, development.md and three skills say so. - The knowledge-base and plugin-fields skills no longer say a deep body never raises: a read started within about 30 frames of the recursion limit can, as the timeline skill and the user guide already said. - Write models keep the caller's Markdown stripped at both ends, not verbatim; stripping unindents a body that opens with an indented code block, so the guide says to open such a body with a fence. - CVE-2025-6069 guard: whatever follows a body's last ">" reads as text, an unterminated comment and a closing tag cut short inside a link included, where a patched CPython drops some of it. Test 15 pins both. Each stated behaviour was re-checked on CPython 3.12.3 and 3.13.14. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
The correction to fix 11 lands in both packages, as glpi-python-client
0.6.1 and easyvista-python-client 0.4.1, so the statements that the
other package's converter "is this one" now name 0.4.1; provenance
("ported from 0.4.0") stays as it was. The 0.6.1 changelog says the
reader matches 0.4.1, that the converter's writing is unchanged while a
write model sends a lone surrogate as U+FFFD, and how many of
test_fixes's tests fail on which converter.
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
The replacement character went into the source and its tests as the literal character, which reads as an encoding error in a diff. It is now the escape �; the value is the same, and so is every test. A test comment that meant the escaped pair 😀 had become the emoji itself; it now says what it meant. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
|
Codecov Report✅ All modified and coverable lines are covered by tests. Additional details and impacted files@@ Coverage Diff @@
## main #40 +/- ##
==========================================
+ Coverage 97.64% 97.86% +0.22%
==========================================
Files 90 90
Lines 3313 3377 +64
==========================================
+ Hits 3235 3305 +70
+ Misses 78 72 -6 ☔ View full report in Codecov by Harness. 🚀 New features to boost your workflow:
|
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Why
easyvista-python-client carries this package's 0.6.0 converter plus fifteen reader fixes, so the same HTML could read as different Markdown through the two packages. Some of those differences change what a body shows:
A colspan or start the converter could not read as a number failed the whole body.
A review of the port found one more problem, which this branch also fixes. Fix 11 kept , and raw even around blocks, so a table inside one read as pipe text, a list or a heading read as Markdown source, and
left a fence open over the rest of the body. Separately, a lone surrogate in a write model's Markdown failed the whole create or update.What
Verification
Known follow-ups (from the review; all Minor)
round blocks remain a documented hole, unchanged since 0.6.0.<span>,<font>, and now a dropped<u>/<mark>/<ins>) round a list, followed by inline text, reads that text into the list's last item. This is markdownify's general behaviour, the same on every tree; the fix belongs in the list converter's next-sibling handling.data-glpi-block/data-ev-blockattribute already present in stored HTML is trusted by the block walk. Only hand-made or sanitiser-kept HTML carries one; clear it on read.<div>) also change on the second cycle.Twin PR: baraline/easyvista_python_client#8
🤖 Generated with Claude Code