Skip to content

SEO: one URL per page, and give the blog its engine axis - #52

Merged
cevheri merged 9 commits into
mainfrom
seo/one-url-per-page-and-engine-archives
Sep 20, 2026
Merged

cevheri merged 9 commits into
mainfrom
seo/one-url-per-page-and-engine-archives

Conversation

@mfatihdayan

Copy link
Copy Markdown
Collaborator

A crawl of all 122 pages turned up three structural findings and two budget overruns. Four commits, split so the content diff can be read separately from the code.

What was wrong

Every internal link spent a redirect. trailingSlash: 'ignore' with build.format: 'directory' means the host serves /features/ and answers /features with a 301 — but links, canonicals and the RSS feed were all written without the slash, while the sitemap used it. Three signals, three answers. 2,949 links and 121 canonicals affected; a canonical naming a redirect is not a canonical.

Posts related to nothing. "Keep reading" took the two newest posts and showed them to all 104. A reader on a Postgres page got the same two links as a reader on a Redis page, and never the other eight Postgres posts. The blog has always grouped by engine — the grouping just lived in the filename where nothing could reach it.

Titles and descriptions overran what a listing renders. 96 of 146 titles past 60 characters — but — LibreDB Studio is 17 of them, charged to every page, and removing it alone dropped the count to 11. 23 descriptions past 160, because the same string is the lede printed on the page and is written to open a post rather than to end at a limit.

What changed

before after
internal links costing a 301 2,949 0
canonicals naming a redirect 121 0
distinct related-post recommendations 3 86
BlogPosting.image / BreadcrumbList 0 104
rendered breadcrumb / prev-next 0 104
engine archives 0 17
in-body engine cross-links 0 157
sitemap lastmod 0 121
titles over 60 chars 96 13
descriptions over 160 chars 23 0

pagePath() in src/lib/site.ts is the single place that decides a URL's shape, and it leaves assets alone. src/lib/posts.ts reads the engine from the slug, sourcing the id list from engines.ts rather than keeping a second copy, matching longest-first because sqlserver, sqlite and libsql all contain sql.

Judgement calls worth reviewing

  • Only the engine chip navigates. The declared tags stay plain text: "Engineering" is on all 104 posts and "Databases" on all but four, so a link on either leads to a page that duplicates /blog instead of narrowing it.
  • The brand suffix is fitted, not fixed — full name where it fits, — LibreDB where that fits, nothing where neither does. 75 pages still carry it, none at the cost of their own words. The 13 still over 60 are editorial headlines that were long before any suffix; seoTitle is the per-post escape hatch and leaves the H1 alone.
  • Descriptions are trimmed for the tag only, at a sentence boundary. The lede on the page stays whole. seoDescription overrides.
  • lastmod only where a date is known — the posts' front matter, and each archive's newest post. Marketing pages stay undated: a lastmod refreshed on every deploy is the signal Google learns to ignore.
  • Archives need two posts. One post is the post, not an archive of it. 17 of 18 engines qualify.
  • The cross-linker skips libredb. In this blog "LibreDB" almost always means LibreDB Studio, not the embedded database, so the archive would be the wrong target. It also caps at three per post and never touches code, headings, tables, existing links or block quotations — a link injected into a quoted sentence is an edit to someone's words.
  • No aggregateRating. It would put stars in the SERP, but Google requires real reviews on the site and there are none.

Content commit

content: touches 101 posts with two mechanical passes — 176 body links given the slash the host serves, and 157 first mentions of another engine turned into links. No prose rewritten. The rules are in the commit message and asserted in tests/engine-archives.test.ts, so a rerun that gets greedy fails rather than ships.

Verification

bun run gate passes. Tests 166 → 212; the four new suites assert the built output rather than the helpers, because the helpers were never the part that broke.

Three findings from a crawl of all 122 pages, all of them structural.

**One URL per page.** `trailingSlash: 'ignore'` with `build.format:
'directory'` meant the host served `/features/` and answered `/features`
with a 301 — but every link, every canonical and the RSS feed were
written without the slash. So all 2,949 internal links spent a redirect,
and 121 pages carried a canonical naming the redirect rather than the
page it was stamped on. The sitemap, meanwhile, listed the slashed form:
three signals, three answers.

`pagePath()` in src/lib/site.ts is now the single place that decides,
and it leaves assets alone — `/og/default.png/` is not a file.

**Posts that relate to nothing.** "Keep reading" took the two newest
posts and showed them to all 104, so a reader on a Postgres page was
offered the same two links as a reader on a Redis page and never the
other eight Postgres posts. The blog has always grouped by engine — nine
PostgreSQL, eight MySQL — but the grouping lived in the filename where
nothing could reach it. src/lib/posts.ts reads it from there, sourcing
the id list from engines.ts rather than keeping a second copy, and
matching longest-first because sqlserver, sqlite and libsql all contain
`sql`.

That gives each post same-engine recommendations, sequential
prev/next within its engine, and a clickable chip. The declared tags
stay plain text on purpose: "Engineering" is on all 104 posts and
"Databases" on all but four, so a link on either leads to a page that
duplicates /blog instead of narrowing it.

**Archives.** /blog/engine/<id>/ for the seventeen engines with two or
more posts. One post is the post, not an archive of it.

Also: BlogPosting gains the image it needed for Article rich results
(og:image was already computed one line away), BreadcrumbList lands on
every post and archive, mainEntityOfPage names the served URL, and the
sitemap carries lastmod where a date is actually known — the posts' own
front matter, and each archive's newest post. Marketing pages stay
undated; a lastmod refreshed on every deploy is the signal Google learns
to ignore.

`updatedAt` is optional and uses the preprocessor pattern, not a bare
`.optional()` — tests/content.test.ts explains why.

Tests 166 → 198. The three new suites assert the built output, because
the helpers were never the part that broke.

Not done, deliberately: aggregateRating. It would put stars in the SERP,
but Google requires it to come from real reviews on the site, and there
are none.
Two mechanical passes over the posts, no prose rewritten.

**Slashes.** 176 body links were written `/features`, which the host
answers with a 301 to `/features/`. Same fix as the templates got.

**Cross-links.** 226 places named another engine in plain prose with no
link — 59 posts mention PostgreSQL, 41 SQLite, 38 DuckDB. 157 of those
first mentions now point at that engine's archive, which is where a
reader who just read the comparison wants to go.

The rules the pass applied, in case it is ever run again:

- at most three per post — a body full of chips reads as stuffing and
  dilutes every link on the page
- never a post's own engine
- first mention only, never every occurrence
- nothing inside code fences, inline code, headings, tables, existing
  links, or block quotations. A link injected into someone else's quoted
  sentence is an edit to their words.
- `libredb` excluded. In this blog "LibreDB" almost always means LibreDB
  Studio, not the embedded database of the same name, so the archive
  would be the wrong target.

tests/engine-archives.test.ts asserts each of those against the built
HTML, so a rerun that gets greedy fails rather than ships.
Google renders about 60 characters of a title. 96 of 146 pages ran past
it — but the writing was rarely the reason. ` — LibreDB Studio` is 17
characters charged to every page, and removing it alone drops the count
to 11. The brand was crowding out the words that tell the two pages
apart.

So the suffix is now fitted rather than fixed: the full name where it
fits, `— LibreDB` where that fits, nothing where neither does. 75 pages
still show the brand, none of them at the cost of their own words, and
the overrun falls 96 → 13. The thirteen are editorial headlines that are
long before any suffix is added; `seoTitle` lets a page title itself for
a listing without touching its H1.

Descriptions had the same shape of problem from a different cause: the
string is both the meta description and the lede printed on the page, so
it is written to open a post, not to end at 160 characters. 23 ran over,
up to 264 — losing the closing clause, which is usually the part that
earns the click. The trim now happens in buildSeo, for the tag only, and
stops at a sentence rather than mid-word. The lede stays whole.
`seoDescription` overrides it where a post wants a different listing.

Also: /security, /compare and /docker-compose declare TechArticle and
named no author. Same Organization the posts use.

tests/serp-budget.test.ts asserts the rules and the built pages both.
104 posts in one flat list, and the only way in was to scroll it. The
per-engine archives exist now, so the listing names them: a row of
engines with their post counts, above the cards.

Plain links, not a client-side filter. The grouping is then crawlable,
each subset has an address someone can send, and it works with
JavaScript off — which a filter would not, for no gain the filter offers
over seventeen links.
Comment thread tests/serp-budget.test.ts Fixed
mfatihdayan and others added 5 commits September 20, 2026 13:25
… renders

`building-universal-database-provider-typescript` arrived from somewhere
self-contained and kept the furniture it needed there. On this site that
furniture is duplicated:

- `> **Author:** Cevheri & The LibreDB Studio Engineering Team` sat two
  lines under the byline the template already prints from front matter,
  so the page named the author twice in a row. `Topic:` restated the
  tags; `Target Audience:` is an artefact of the original format.
- A hand-written `## Table of Contents` sat beside the generated one in
  the sidebar — and, being an `<h2>`, appeared *inside* that sidebar as
  its own first entry. The page offered a table of contents of its table
  of contents.

Both removed, along with the `---` rule that closed the header block:
with the block gone it would have been left stranded against the front
matter, where the parser stops treating what follows as body. That is
not hypothetical — it swallowed the Introduction section on the first
attempt at this edit.

Only this post is affected; the other 103 carry neither pattern.
Verified after the change that all thirteen headings still render and
the sidebar leads with Introduction again.
…lace

Two follow-ups on this branch.

serp-budget's rendered() chained six .replace() calls, so each ran over what
the previous produced: '&amp;lt;' decoded to '&lt;' and then to '<'. The count
it exists for came out right either way, but CodeQL reports it as
js/double-escaping at high severity and the required check stays red. Matching
once and consuming the match is the only ordering that cannot double-decode.

The sitemap serializer re-derived a post's engine with a bare startsWith(),
which is the second copy of a mapping this file's own comment warns against,
and it drops the longest-match rule that stops 'sqlite' claiming a 'sqlserver-'
post. It now calls postEngine() like everything else does.

Sitemap output unchanged where it was already right: engine/sqlite 2026-07-20,
engine/sqlserver 2026-06-17, engine/libsql 2026-05-06, engine/libredb
2026-05-28. Gate green, 212 tests.
--radius-md is not in the vocabulary: the scale is xs/s/m/l/xl/2xl. An
undefined custom property with no fallback makes the declaration invalid at
computed-value time, so border-radius resolved to 0 and the card rendered with
square corners while every card around it is rounded.

.post__seqlink is a bordered surface card in a grid next to PostCard, so it
takes --radius-l (12px), the token the design system documents as the card
radius, and now matches them.
/helper landed while this branch was open, so it was written against
trailingSlash: 'ignore' and never saw these tests.

Its link to /get-started spent a 301, and its TechArticle node had no author,
which is the one Google treats as required on an article. Both are what the
tests on this branch exist to catch; nothing else on the page changes.

Author is the organisation, matching what every blog post already emits.
@cevheri
cevheri merged commit bdd8226 into main Sep 20, 2026
5 checks passed
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

3 participants