Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension


Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
56 changes: 55 additions & 1 deletion CHANGELOG.md
Original file line number Diff line number Diff line change
Expand Up @@ -15,6 +15,59 @@ is the error. Tags carry no `v` prefix.

## [Unreleased]

## [0.4.1] - 2026-10-02

A patch release of the content reader. A `<u>`, `<mark>` or `<ins>` round a
block no longer reads as broken Markdown, four guards of 0.4.0's fixes gain
the tests they lacked, and `docs/content.rst` corrects what it said about
`start` values, `ValueError`, the CVE-2025-6069 tail and a second round
trip. Writing is unchanged, and so is everything outside
`easyvista_python_client.content`. `glpi_python_client`'s port of this
converter makes the same correction.

### Fixed

- **A block inside `<u>`, `<mark>` or `<ins>` read as broken Markdown**, since
0.4.0 kept these tags raw (fix 11). Markdown has no inline tag round blocks:
a table inside one read as pipe text, a list, a heading, a quote or a rule
as its Markdown source, paragraphs were not a fixed point, and
`<u><pre>code</pre></u>` left a fence open that showed the rest of the memo
as code. Round a block the tag is now dropped and the blocks kept, as
`glpi_python_client`'s reader at `917f030` read them. Inside a table cell
or a heading, whose blocks are one line, the tag stays, and around inline
content nothing changes. The Markdown read from such a memo changes once.

### Documentation

- `docs/content.rst`, corrected after a review of the converter:
- a `start` that is not a decimal number counts from 1 as a browser
counts one holding no digit; a browser reads `" 3"`, `"+3"` or
`"3abc"` as 3, which the converter counts from 1;
- any `ValueError` raised while converting degrades the memo to text, not
only an unreadable `colspan` or `start` (the module and
`from_transport` docstrings say so too);
- the CVE-2025-6069 note says that everything after a memo's last `>`
reads as text, an unterminated comment and a closing tag cut short
inside a link included, where a patched CPython drops some of it;
- a second write-and-read cycle changes nothing more except for two
shapes, adjacent lists with different bullets and a fence whose info
string holds a character reference, where 0.4.0's page and its
changelog said it changes nothing;
- `<u>`, `<mark>` and `<ins>` round a block, and a known hole for
`<b>`, `<em>` and `<s>` round blocks.
- The `easyvista-ticket-actions` skill says a `<u>` round a block is
dropped.

### Notes

- Four guards of the 0.4.0 fixes that no test caught are pinned, each by a
test that fails without it: a second `<` before a space keeps one escape
(fix 5), a header cell escapes a `|` in `<kbd>` or `<samp>` and splits a
line break in inline code (fix 6), a table inside an `<a>` without `href`
stays a table (fix 7), and an ordered item numbered 10 or more indents its
content by its bullet's width (fix 12). A count pins that the walk finding
blocks inside `<u>`, `<mark>` and `<ins>` checks each tag about once.

## [0.4.0] - 2026-10-02

Adds Markdown <-> memo HTML conversion as an optional extra, drops Python
Expand Down Expand Up @@ -1592,7 +1645,8 @@ Initial public release.
status/error code, with non-retryable validation errors (HTTP 590, code 2013).
- `py.typed` marker — the package ships inline type information.

[Unreleased]: https://github.com/baraline/easyvista_python_client/compare/0.4.0...HEAD
[Unreleased]: https://github.com/baraline/easyvista_python_client/compare/0.4.1...HEAD
[0.4.1]: https://github.com/baraline/easyvista_python_client/compare/0.4.0...0.4.1
[0.4.0]: https://github.com/baraline/easyvista_python_client/compare/0.3.0...0.4.0
[0.3.0]: https://github.com/baraline/easyvista_python_client/compare/0.2.0...0.3.0
[0.2.0]: https://github.com/baraline/easyvista_python_client/compare/0.1.0...0.2.0
Expand Down
64 changes: 43 additions & 21 deletions docs/content.rst
Original file line number Diff line number Diff line change
Expand Up @@ -141,15 +141,21 @@ What each kind of element becomes:
stay as those raw tags, and **struck text** (``<s>``, ``<del>``,
``<strike>``) as a raw ``<s>``. CommonMark has no spelling for any of them,
and ``to_transport`` passes the tags through, so the formatting survives.
The cost is raw HTML in the Markdown.
The cost is raw HTML in the Markdown. Markdown has no inline tag round
blocks, so a ``<u>``, ``<mark>`` or ``<ins>`` holding a table, a list, a
heading, a quote, a ``<pre>``, a rule or paragraphs is dropped and its
blocks kept; inside a table cell or a heading, which hold one line, it
stays.
* **Links** become ``[text](https://... "title")``, and a link whose text is
its own URL, a pasted link, becomes the autolink ``<https://...>``. No link
target is filtered, ``javascript:`` included (see `It is not a sanitiser`_).
* **Lists** nest and keep their numbers, including an ``<ol start>``.
**Tables** become GFM tables, and a ``<br>`` inside a cell stays a raw
``<br>``, since a GFM cell is one line. **Preformatted blocks** become fences
that keep the ``language-`` class ``cmark-gfm`` writes, with a fence longer
than any run of backticks in the code.
* **Lists** nest and keep their numbers, including an ``<ol start>``. A
``start`` that is not a decimal number counts from 1, as a browser counts
one holding no digit, such as ``²``; a browser reads ``" 3"``, ``"+3"`` or
``"3abc"`` as 3. **Tables** become GFM tables, and a ``<br>`` inside a cell
stays a raw ``<br>``, since a GFM cell is one line. **Preformatted blocks**
become fences that keep the ``language-`` class ``cmark-gfm`` writes, with a
fence longer than any run of backticks in the code.
* ``<head>``, ``<script>``, ``<style>``, ``<template>`` and ``<title>`` are
dropped, as a browser does not display them. Styling such as ``<font>``
colours or ``<span style>`` keeps its text and loses the style.
Expand Down Expand Up @@ -180,10 +186,12 @@ and 3.14.6 (measured 2026-10-02, default recursion limit). The converter does
not predict that. It attempts the conversion and, if the walk does not fit,
reads the memo as its text instead, a line per block. Every word the
conversion would have produced is still there, in order. What is lost is
structure: link targets, image alt text, emphasis and code fencing. A
``colspan`` or ``start`` attribute ``markdownify`` cannot read as a number,
such as ``"²"``, takes the same path, and so does a document ``html.parser``
refuses outright, such as one carrying an unknown ``<![FOO[`` marked section.
structure: link targets, image alt text, emphasis and code fencing. Any
``ValueError`` from the conversion takes the same path -- ``markdownify``
raises one for a ``colspan`` or ``start`` attribute it cannot read as a
number, such as a ``colspan`` of ``"²"`` -- and so does a document
``html.parser`` refuses outright, such as one carrying an unknown
``<![FOO[`` marked section.

Because the budget is whatever stack is left when the call starts, the same
memo can convert from one call site and degrade from a deeper one. A caller
Expand Down Expand Up @@ -275,9 +283,13 @@ A few things do not come back:
* a line holding only ``*`` is an empty list item in CommonMark, and reads
back as nothing.

A second cycle changes nothing more. That is the property a two-way sync
relies on: once a text has made one trip, writing what was read back and
reading it again gives exactly the same Markdown.
After that first cycle, a further one changes nothing more, with two
exceptions (reproduced 2026-10-02): two adjacent lists with different
bullets read as one loose list, then as one tight list; and a fence whose
info string holds a character reference, such as ``&amp;amp;``, loses one
level of it on each cycle. A two-way sync relies on the rest: once a text
has made one trip, writing what was read back and reading it again gives
the same Markdown.

**A memo read, written back and read again.** This direction is held to more.
The aim is that what the Markdown displays is what the memo displayed, and
Expand Down Expand Up @@ -332,7 +344,12 @@ in the preproduction sample:
* a table whose ``<td>`` and ``<tr>`` are never closed folds into one cell,
keeping its words;
* a definition list (``<dl>``) reads as a ``term`` line and a
``: definition`` line, so its display gains the colon.
``: definition`` line, so its display gains the colon;
* bold, italic or struck text (``<b>``, ``<em>``, ``<s>`` and their
synonyms) wrapped round blocks, other than a single paragraph, shows its
markers or its tag as text, and a list, a table or a heading inside it as
Markdown source; round a ``<pre>`` it leaves a fence open, so the rest of
the memo shows as code.

Some shapes are not fixed points at the first read, and settle after one more
cycle. None broke a fixed point in the preproduction sample. Adjacent lists
Expand Down Expand Up @@ -390,9 +407,12 @@ fixed in those releases according to their changelogs). Many unfinished tags
after a memo's last ``>`` is the shape that reaches the converter, and a memo
is outside data. So ``from_transport`` spells every ``<`` after the last ``>``
as ``&lt;`` before parsing, since no ``<`` there can finish a tag. As a side
effect, an unfinished tag at the very end, ``x <a``, reads as the text
``x \<a`` on every interpreter, whereas a patched CPython would drop it. Prefer
a patched interpreter anyway: the guard covers the converter's input, and the
effect, whatever follows the memo's last ``>`` reads as text on every
interpreter, where a patched CPython drops some of it: an unfinished tag at
the very end, ``x <a``, reads as ``x \<a``; an unterminated comment,
``<!-- note``, as ``\<!-- note``; and a closing tag cut short inside a link,
``<a href="...">lien</a``, leaves ``\</a`` in the link's text. Prefer a
patched interpreter anyway: the guard covers the converter's input, and the
CPython fix covers the parser itself.

Where it comes from
Expand All @@ -414,12 +434,14 @@ proposed to ``glpi_python_client``. The fixes cover:
* ``<br>`` inside inline code;
* tables inside headings and links, and the spacing of flattened cells;
* ``<center>``;
* underline and highlight, as raw tags;
* underline and highlight, as raw tags, except round a block (corrected in
0.4.1);
* numbering a long ordered list in linear time;
* a quadratic pattern on runs of spaces;
* unreadable ``colspan`` and ``start`` values;
* the CVE-2025-6069 tail.

``CHANGELOG.md`` lists them under 0.4.0. Apart from those fixes, the names and
the error messages, the only difference from ``glpi_python_client`` is that the
libraries are an optional extra here, not dependencies.
``CHANGELOG.md`` lists them under 0.4.0, and the correction under 0.4.1.
Apart from those fixes, the names and the error messages, the only
difference from ``glpi_python_client`` is that the libraries are an optional
extra here, not dependencies.
2 changes: 1 addition & 1 deletion easyvista_python_client/_version.py
Original file line number Diff line number Diff line change
Expand Up @@ -5,4 +5,4 @@
live only in ``__init__``.
"""

__version__ = "0.4.0"
__version__ = "0.4.1"
54 changes: 48 additions & 6 deletions easyvista_python_client/content/conversion.py
Original file line number Diff line number Diff line change
Expand Up @@ -26,9 +26,10 @@
out: plain text, line breaks a browser does not show, bold and italic
CommonMark would not close, link targets, and an mdformat set up without
its nesting cap and with its quadratic lookups made linear. A body nested
too deeply for the stack is read as its text (:func:`_text_of`); anything
else that fails raises
:class:`~easyvista_python_client.EasyvistaContentError`.
too deeply for the stack, or on which the conversion raises ``ValueError``
-- markdownify does for a ``colspan`` or ``start`` it cannot read as a
number -- is read as its text (:func:`_text_of`); anything else that fails
raises :class:`~easyvista_python_client.EasyvistaContentError`.

A memo with no HTML element in it is read as literal lines, a line break
per line. How EasyVista's web UI displays such a memo is unverified: the
Expand Down Expand Up @@ -138,6 +139,11 @@
_BEFORE, _AFTER = "data-ev-before", "data-ev-after"
_NUMBER = "data-ev-number" # an ordered item's number (_Converter.convert_li)

#: The elements kept as raw tags (:meth:`_Converter.convert_u`), and the
#: attribute marking one that holds a block (:func:`_note_blocks`).
_RAW_INLINE = frozenset({"u", "mark", "ins"})
_HOLDS_BLOCK = "data-ev-block"

#: Text split into leading line breaks and spaces, content, trailing ones.
#: The content ends on its last character that is neither, found greedily:
#: a lazy ``.*?`` rescanned the run after it at every step, quadratic.
Expand Down Expand Up @@ -330,6 +336,33 @@ def _note_sides(root: Tag) -> None:
element[_AFTER] = " "


def _note_blocks(root: Tag) -> None:
"""Mark each ``<u>``, ``<mark>`` or ``<ins>`` that holds a block.

Markdown has no inline tag around blocks: such a tag wrapped round them
showed a table, a list or a heading as its Markdown source, and round a
``<pre>`` left a fence open to the end of the body. One walk: a block
marks the open ones from the innermost out and stops at one already
marked, whose holders were marked with it, so each is marked once
however deeply they nest.
"""

holders: list[Tag] = []
for node, entering in _walk(root):
if not isinstance(node, Tag):
continue
if node.name in _RAW_INLINE:
if entering:
holders.append(node)
else:
holders.pop()
elif entering and node.name in _BLOCKS:
for holder in reversed(holders):
if holder.has_attr(_HOLDS_BLOCK):
break
holder[_HOLDS_BLOCK] = ""


def _punctuation(char: str) -> bool:
"""Return whether CommonMark counts ``char`` as punctuation."""

Expand Down Expand Up @@ -426,8 +459,14 @@ def convert_s(self, el: Tag, text: str, parent_tags: set[str]) -> str:
convert_strike = convert_s

def convert_u(self, el: Tag, text: str, parent_tags: set[str]) -> str:
"""Underlined or highlighted text: CommonMark has neither, so raw HTML."""
"""Underlined or highlighted text: CommonMark has neither, so raw HTML.

Round a block the tag is dropped and the blocks kept, except on the
one line of a cell or a heading, where the blocks are inline text.
"""

if el.has_attr(_HOLDS_BLOCK) and "_inline" not in parent_tags:
return text
return self._markup(el, text, parent_tags, "", el.name)

convert_mark = convert_ins = convert_u
Expand Down Expand Up @@ -705,6 +744,7 @@ def html_to_markdown(html: str) -> str:
_flatten_nested_tables(soup)
_drop_trailing_breaks(soup)
_note_sides(soup)
_note_blocks(soup)
return str(_FORMATTER.render(_CONVERTER.convert_soup(soup))).strip()


Expand Down Expand Up @@ -762,8 +802,10 @@ def from_transport(value: object, *, plain_text_is_markdown: bool = False) -> st
------
EasyvistaContentError
The value could not be converted. A body nested too deeply for
the stack left is read as its text instead, so this is a
backstop.
the stack left, or on which the conversion raises
``ValueError`` (a ``colspan`` or ``start`` markdownify cannot
read as a number, among others), is read as its text instead,
so this is a backstop.
RecursionError
Only when called from within a few frames of the recursion
limit, where no stack is left even to report the failure as
Expand Down
13 changes: 11 additions & 2 deletions easyvista_python_client/content/tests/display.py
Original file line number Diff line number Diff line change
Expand Up @@ -21,7 +21,14 @@
a nested table's cell is a word boundary, as a browser shows it (a new
line, a new box) and as the converter writes it (a space); and
* an ordered list's ``start`` is read with ``isdecimal``, as the converter
reads it, where ``isdigit`` made the oracle itself raise on ``"²"``.
reads it, where ``isdigit`` made the oracle itself raise on ``"²"``. So
the oracle cannot see where a browser numbers otherwise: from 3 for
``" 3"``, ``"+3"`` or ``"3abc"``, and from 1 for a full-width 3
(U+FF13).

Since 0.4.1, with ``glpi_python_client`` 0.6.1, a rule (``<hr>``) inside an
inline element counts as a block, where the oracle read it as nothing: a
browser draws ``<u><hr></u>`` as a rule.

What the oracle cannot see: underline and highlight (``<u>``, ``<mark>``,
``<ins>``), struck text, link and image titles, and line breaks inside a
Expand Down Expand Up @@ -146,8 +153,10 @@ def _inline(node: Tag, words: _Words, fmt: Format) -> None:


def _has_block(node: Tag) -> bool:
# Since 0.4.1: a rule is a block too, so <u><hr></u> draws its rule.
return any(
isinstance(child, Tag) and child.name in _BLOCKS for child in node.descendants
isinstance(child, Tag) and (child.name in _BLOCKS or child.name == "hr")
for child in node.descendants
)


Expand Down
41 changes: 41 additions & 0 deletions easyvista_python_client/content/tests/test_cost.py
Original file line number Diff line number Diff line change
Expand Up @@ -41,7 +41,9 @@
from collections.abc import Callable

import pytest
from bs4 import Tag

from easyvista_python_client.content import conversion
from easyvista_python_client.content.conversion import EasyvistaContentConverter

read = EasyvistaContentConverter.from_transport
Expand Down Expand Up @@ -194,6 +196,45 @@ def test_a_tail_of_unfinished_tags_converts_in_linear_time() -> None:
assert_linear(lambda n: "<p>r</p>" + "x <a " * n, 500)


def test_the_walk_round_blocks_checks_each_tag_about_once(
monkeypatch: pytest.MonkeyPatch,
) -> None:
"""The walk marking a ``<u>``, ``<mark>`` or ``<ins>`` that holds a block
stops at the first tag already marked, so a block checks only the tags
marked since the last one.

Without the stop, every block checked every tag open around it: blocks
times depth. The output is the same either way, and a body nested past
markdownify's stack, which still pays for the walk before it falls back
to text, is where the cost would show. A ratio at sizes a test can
afford does not tell the two apart reliably, so this counts the checks:
about one per tag here, against one per tag per block without the stop.
The floor fails a rewrite that checks some other way, rather than
letting it pass on a count of nothing.
"""

checks = 0
has_attr, get = Tag.has_attr, Tag.get

def counting_has_attr(self: Tag, key: str) -> bool:
nonlocal checks
checks += key == conversion._HOLDS_BLOCK
return has_attr(self, key)

def counting_get(self: Tag, key: str, default: object = None) -> object:
nonlocal checks
checks += key == conversion._HOLDS_BLOCK
return get(self, key, default)

monkeypatch.setattr(Tag, "has_attr", counting_has_attr)
monkeypatch.setattr(Tag, "get", counting_get)
depth = 50

read("<u>" * depth + "<p>x</p>" * depth + "</u>" * depth)

assert depth <= checks <= 4 * depth


def test_a_body_too_deep_to_convert_keeps_its_text() -> None:
"""markdownify recurses per nesting level; past the stack, the text is kept."""

Expand Down
Loading
Loading