Skip to content

Give each engine archive its own feed, and make the feeds validate - #59

Open
yusuf-gundogdu wants to merge 2 commits into
mainfrom
blog-engine-rss
Open

yusuf-gundogdu wants to merge 2 commits into
mainfrom
blog-engine-rss

Conversation

@yusuf-gundogdu

Copy link
Copy Markdown
Member

The blog groups its posts by engine, but it publishes one feed for all of them, so anyone who wants a single engine has to filter us from the outside. Planet for the MySQL Community does exactly that today, and it asked us for a better input.

Two commits, in this order, so the change to the feed people already subscribe to can be read and reverted on its own.

What was wrong

The only thing a third party can filter on is the title. We are aggregated at planet.oursqlcommunity.org through siftrss with an include/title//mysql/i rule, because that is all an outside service can see. Eight published posts are about MySQL and the rule reaches seven. The one it drops is "Nine servers, one connection type, different answers", whose title names the behaviour rather than the engine, which is how we title most posts. The maintainer predicted this when we submitted and asked whether we could expose a category feed instead.

We are paying for that workaround in public. Our row in their blogroll points its RSS icon at siftrss.com rather than at libredb.org, and carries their filtering icon beside our name. The maintainer also had to create the siftrss link himself and now maintains it.

Our feed does not validate. It writes the author's name into <author>, which RSS 2.0 defines as an email address. The W3C validator answers Invalid email address once per item and refuses the feed, 104 times on /rss.xml as it stands.

One deleted line could have broken every feed silently. The namespaces are declared on <rss>; without them the documents are not well-formed XML and every reader rejects them. I checked: deleting that line left the whole suite green.

What changed

Each engine archive serves the feed for its own posts at /blog/engine/<engine>/rss.xml, built from the post set the archive page lists, with the engine as the first category on every item so a category filter works without touching the title. The archive page links it in the body and declares it in the head ahead of the site feed, which stays advertised so the whole blog is still reachable from there.

The author's name moved to <dc:creator> on every feed and each channel now carries atom:link rel="self".

The grouping and its threshold already existed three times: the archive route, the sibling links on an archive, and the engine index on /blog. The feed would have been a fourth, so all four now read engineArchives() in lib/posts and the threshold is one constant.

before after
MySQL posts an aggregator can reach 7 8
feeds published 1 18
feeds the W3C validator accepts 0 18
copies of the engine grouping and its threshold 3 1

Tests

Six new assertions, each checked to fail without the thing it guards rather than assumed to:

removed from the source what went red
the xmlns declaration feeds parse as XML
dc:creator, back to <author> that, plus the escaping test
atom:link rel="self" feeds parse as XML
pairing between page route and feed route pairs one feed with one archive
dist/blog/engine/mysql/rss.xml keeps serving the feed an aggregator subscribed to

The fixture behind the archive tests counted every post file on disk while the routes count published ones, so saving a draft turned the suite red with nothing wrong in the site. I reproduced that, then filtered the fixture on status, which also settles nine older tests that shared it.

Follow-up outside this repo

Once this is live their planet.ini entry can point at https://libredb.org/blog/engine/mysql/rss.xml and drop the siftrss link, the filtering comments and the _FI_ title prefix that draws the icon. The guids are identical between the two feeds, so switching will not re-flood their front page with posts they have already carried. I am answering the submission thread with the address; their own contributing history says the maintainer makes that edit himself, so I am not sending a pull request there unless he asks.

yusuf-gundogdu added 2 commits September 26, 2026 01:59
The site feed wrote the author's name into <author>, which RSS 2.0 defines as an email address. The W3C feed validator refuses the feed over it, once per item, so https://libredb.org/rss.xml does not validate today. The name moves to <dc:creator>, which is what readers and planet aggregators already read, and the channel gains the atom:link rel="self" the validator asks for.

Both live behind namespace declarations on <rss>, so removing either declaration leaves every feed on the site not well-formed XML and every reader rejecting it. Nothing in the suite noticed that. dist-smoke now parses each built feed, checks that every prefix it uses is declared, and pins the escaping of a creator name carrying an ampersand, which is the one we ship.
The blog groups posts by engine, but it publishes one feed for all of them, so an aggregator that wants a single engine has to filter us from the outside. The only thing an outside service can filter on is the title, and our titles name the behaviour rather than the engine, so it drops posts. Planet for the MySQL Community reads us through such a filter today and misses one.

Each engine archive now serves the feed for its own posts at /blog/engine/<engine>/rss.xml, built from the post set the archive page lists, with the engine as the first category on every item. The archive page links it in the body and declares it in the head ahead of the site feed, which stays advertised so the whole blog is still reachable from there.

The grouping and its threshold already existed three times: the archive route, the sibling links on an archive, and the engine index on /blog. The feed would have been a fourth, so all four now read engineArchives() in lib/posts and the threshold is one constant.

The fixture behind the archive tests counted every post file on disk while the routes count published ones, so saving a draft turned the suite red with nothing wrong in the site. It filters on status now, which also settles nine older tests that shared it.

This branch has not been deployed

No deployments
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant