Project guide for Claude Code working in this repository. Tells you what this site is, how the AEO platform is wired together, where the dependencies live, and the conventions to follow.
stackql.io is the marketing and documentation site for StackQL, built on Docusaurus 3.10. Three audiences:
- Humans - default site nav. This is a docs-only site: the docs tree is the site root (
/installing-stackql,/providers,/getting-started/*,/command-line-usage/*,/mcp/*,/language-spec/*,/developers/*,/quick-starts/*) and docs/index.md is the homepage. Blog at/blog/<section>/*(three sections, see "The blog" below). A few top-level React pages remain (/features,/privacy, redirect stubs). - AI agents and answer engines - a parallel content surface at
/ai/*(canonical definitions, comparisons, how-tos, FAQs, troubleshooting, etc.) reachable by deep link or viallms.txt, but not linked from the human nav. - LLM crawlers -
llms.txtandllms-full.txtat site root; raw markdown twin (/foo.md) for every doc and blog page.
The two surfaces serve the same URLs to all visitors (no UA-based cloaking). The /ai/* pages just don't appear in the human-facing nav.
- Docusaurus 3.10.1, classic preset
- React 18 (pinned - several deps require it)
- MUI 5/6 (already in deps - Material UI and emotion)
- Node 22 toolchain, yarn is the package manager (yarn.lock is the only lockfile - package-lock.json is gitignored)
- Deployed via Netlify, build =
npm run build, publish =build/ - Search via Algolia DocSearch (
ALGOLIA_APP_ID,ALGOLIA_API_KEY,ALGOLIA_INDEX_NAMEenv vars required for prod builds)
The classic preset's docs instance has routeBasePath: '/' (Docusaurus docs-only mode). docs/index.md carries slug: / and is the homepage; there is no src/pages/index.js. The marketing homepage was collapsed into the docs site in August 2025, and until October 2026 the root was a meta-refresh stub to /docs, which is why Google indexed /docs as the home and showed no sitelinks.
- Inbound
/docs/*links are Netlify 301s to the same path without the prefix (/docs->/,/docs.md->/index.md,/docs/*->/:splat). Those rules sit below the/docs/query-library/*proxy and the specific legacy/docs/...rules in netlify.toml. Netlify applies the first matching rule, so keep that order. - The query library stays proxied at
/docs/query-library/*: it is a separate site whosebaseUrlis a cross-repo contract (see "The query library" below). It is the only thing left under/docs. - Internal links use root paths (
/installing-stackql, never/docs/installing-stackql).onBrokenLinks: 'throw'catches relative mistakes; absolutestackql.io/docs/...links need a grep. - Root-level routes that are not docs:
/blog/*,/ai/*,/features,/privacy,/search, the provider routes (/providers/<slug>,/registry/*) and the stubs/install,/stackqldocs,/downloads,/stackql-deploy,/contact-us. A new doc whose slug collides with one of these would clash with it, so check before adding root-level docs. - Getting Started and Command Line Usage are sidebar categories with generated index pages (
/getting-started,/command-line-usage), so every major section has a landing page Google can surface as a sitelink: installation, providers, getting started, command line usage, MCP.
Two custom plugins published to npm and one secondary content-docs instance handle the AEO pipeline. Both plugins are in sibling repos and developed alongside this site.
Source: ../docusaurus-plugin-structured-data/ (sibling to this repo)
Published: npm registry, current version pinned in package.json
Lifecycle: postBuild + allContentLoaded
Emits JSON-LD <script type="application/ld+json"> blocks into the <head> of every emitted HTML page. The shape:
WebPage+BreadcrumbList+WebSite+Organizationon every pageArticle+ImageObject+Personon blog posts (/blog/<section>/<slug>). Breadcrumbs follow the blog instance base path (Home > Blog > Section > Post) andArticle.articleSectionis['Blog', '<Section>']- needs plugin >= 1.6.0TechArticleon every doc page of the default andaidocs instances (techArticleDocsInstances: ['default', 'ai'], plugin >= 1.6.0), except the instance roots/and/ai, which stayWebPage.techArticleRoutePrefixes: ['/ai/']is kept so older plugin versions still mark/ai/*.FAQPage,HowTo,SoftwareApplicationopt-in via frontmatter (faq:,howTo:,softwareApplication:)SpeakableSpecificationon every WebPage with default selectors- Connected
@graphviamainEntitylinking (TechArticle.mainEntity -> FAQPage when both present)
Config lives at themeConfig.structuredData in docusaurus.config.js. Key settings:
techArticleDocsInstancesandtechArticleRoutePrefixes- which pages get TechArticle (see above)excludedRoutes: []- nothing is excluded;/providersused to be a custom React grid and is a doc page todayauthors:- blog author identity graphorganization:- StackQL Studios identity, contact, address, and a square 512px logo (Google wants a logo mark, not a cover image)breadcrumbLabelMap:- friendly names for breadcrumb segments. Keys are single segments ('quick-starts') or full paths ('/blog/providers'), and a full path wins. Ancestor segments only become linked crumbs when a page exists at that path; otherwise they fold into the leaf name.
If a page has no <meta name="description">, the plugin falls back to siteConfig.tagline. The homepage explicitly sets a description meta in docusaurus.config.js themeConfig.metadata to avoid the fallback on the most-cited page.
Source: ../docusaurus-plugin-aeo/ (sibling to this repo)
Published: npm registry, current version pinned in package.json
Lifecycle: postBuild + allContentLoaded + theme component injection
Four features:
.mdcompanion files - for every doc and blog page, emits a sibling.mdat the same path (e.g./docs/foo->/docs/foo.md). Mirrors the raw MDX source. Only emits when there is a source markdown file (React pages and auto-generated index routes are correctly skipped - they have no source to mirror).llms.txt+llms-full.txt- at site root.llms.txtis the corpus index (Markdown bullet list with title and description per page).llms-full.txtis the concatenated body of every.mdcompanion, separated by\n\n---\n\n.- "Ask AI" dropdown - swizzled into the breadcrumb row of every doc page and the header of every blog post. MUI outlined Button + Menu. Three providers: Claude, ChatGPT, Perplexity (Gemini was removed - it does not accept URL-encoded prompts). Brand icons are hand-rolled inline SVGs in
src/theme/AskAiButton/brand-icons/to avoid React-version conflicts with icon libraries. /ai/*helpers - exported from@stackql/docusaurus-plugin-aeo/helpers. Used to integrate the/ai/*content surface with the structured-data plugin.
Config lives at the plugin options object in the plugins array in docusaurus.config.js:
llmsTxt.instanceSections- section titles + ordering for thellms.txtindex (AI Reference -> Documentation -> one "Blog -" section per blog instance, generated from blogSections)askAi.providerOrder- dropdown ordering (defaults to claude, chatgpt, perplexity)askAi.promptTemplate- the prefilled prompt sent to the AI surface. Default is self-contained ("Read {pageUrl}.md and help me understand it. Summarize the key points, then ask me one clarifying question to dig deeper."). The user can edit it before submitting.
A second @docusaurus/plugin-content-docs instance with id: 'ai', routeBasePath: '/ai', path: 'ai-content', and sidebarPath: false. Configured in docusaurus.config.js.
Directory structure under ai-content/:
ai-content/
├── index.md # /ai landing
├── canonical-definitions/ # "What is X?" pages
├── comparisons/ # "StackQL vs Y" pages
├── how-tos/ # task guides
├── concepts/ # design rationale + best practices
├── faqs/ # topic-grouped Q&A
├── architecture/ # internals
├── troubleshooting/ # error -> resolution
├── industry-positioning/ # where StackQL fits
├── tutorials/ # end-to-end walkthroughs
└── providers/ # auto-generated per-provider reference (placeholder)
Each section has an index.md landing. Individual reference pages go directly inside each section dir.
sidebarPath: false is what keeps these out of the human nav. They're reachable via direct URL, the sitemap, and llms.txt.
The preset blog is disabled (blog: false). Three @docusaurus/plugin-content-blog instances are generated from the blogSections array near the top of docusaurus.config.js (GitHub issue #288). A tag-based split was ruled out because the blog sidebar is built per instance and cannot be filtered by tag.
| Instance id | Content dir | Routes | Nav label |
|---|---|---|---|
product |
blog/product/ |
/blog/product/* |
Product Announcements |
providers |
blog/providers/ |
/blog/providers/* |
Provider Announcements |
tutorials |
blog/tutorials/ |
/blog/tutorials/* |
Tutorials |
- Each instance has its own list page, sidebar (
blogSidebarCount: 'ALL'), tags, pagination and feeds (/blog/<id>/rss.xml,atom.xml,feed.json). blog/authors.yml is shared by all three viaauthorsMapPath: '../authors.yml'. - New posts go straight into the section directory. The directory is the type - there is no marker tag. Slugs are set in front matter as before and must be unique across all three sections (the redirect generator checks this).
/blogis a landing page built by the local plugin plugins/blog-landing/index.js. It reads the three instances inallContentLoaded(the only hook that sees other plugins' content) and adds a route rendering src/components/BlogLanding/index.jsx with the newest five posts per section.- Every pre-split post URL (
/blog/<slug>and its.mdcompanion) has a 301 to its new home in netlify.toml. The per-post block between theBEGIN/END generated blog redirectsmarkers is owned by scripts/generate-blog-redirects.js, which derives rules from front matter slugs. Rerun it only if a pre-split post's slug or section changes; posts written after the split never had an old URL and need no rule. Old/blog/tags/*,/blog/page/*and/blog/archivego to/blog; the old feed URLs go to the product announcements feeds. - The navbar "More" dropdown and the footer list the three sections plus Quick Starts; both are derived from
blogSections. The two announcement entries carry a bullhorn (navLabel) in the header only. The/bloglanding page is deliberately not linked from the header or footer; it is reachable by URL and from the sitemap. - "Tutorials" in the nav means the blog section. The docs walkthroughs formerly at
/docs/tutorials/*are "Quick Starts" at/quick-starts/*(directorydocs/quick-starts/, sidebar category in sidebars.js). Old URLs are 301'd in netlify.toml, including/tutorials->/blog/tutorialsand/cookbooks->/quick-starts(both were meta-refresh React stubs, now deleted). - Sitemap
ignorePatternscover/blog/*/tags/**and/blog/*/page/**. - The
breadcrumbLabelMapentries for the section ids are generated fromblogSections, so JSON-LD breadcrumbs read "Product Announcements" rather than "product". - The shared nav used by the provider microsites lives in
../docusaurus-config(vendored by those sites at build time). Its Blog/Tutorials links must be kept in step with the main site nav.
netlify.toml sets MIME types and cache headers for the AEO files. Critical rules:
*.md->text/markdown; charset=utf-8(top-level and nested)/llms.txtand/llms-full.txt->text/plain; charset=utf-8- All three get
X-Robots-Tag: index, followandmax-age=300cache
Without these rules Netlify serves .md as application/octet-stream (browsers download instead of display) and crawlers may skip it.
Redirect rules live in the same file and their order matters (first match wins): the query library proxy, then specific legacy /docs/... rules, then the /docs/* -> /:splat catch-all, then the remaining hand-written rules, then the generated per-post blog block.
The StackQL query library lives in its own repo/site (query-library.stackql.io)
and is served under the canonical path https://stackql.io/docs/query-library/*
via a Netlify 200 proxy rewrite in netlify.toml. This repo
holds no library content, build tooling or CI - just the rewrite rule and a
navbar link to /docs/query-library/ (an href, not to, so it is a full
page load rather than a client-side route). Response headers for the proxied
path (content types, CORS, cache) come from the origin site, not this repo's
netlify.toml. The path 404s under yarn serve/yarn start - expected,
Netlify-only behaviour. Do not add files under static/docs/query-library/:
Netlify serves matching static files in preference to redirect rules, so any
file there would shadow the proxied site.
src/configs/providers.json is the single source of truth for everything provider-related on this site: config, not code. It is an array of categories, each with providers of { name, href, icon, invertOnDark?, featured?, shortName?, registryAliases? }. The code that reads it is src/lib/providers.js, which validates it, derives each provider's slug and path and exposes PROVIDER_CATEGORIES, FEATURED_PROVIDERS, providerRoutes() and registryRoutes(). Nothing else in the repo holds provider lists; the former src/configs/providers-data.json, providers.ts and the unused ProviderCards component were removed. The catalog drives:
- the tiles and table of contents on docs/providers.md, which imports from
src/lib/providers. Tiles link to/providers/<slug>, not straight to the microsite. - the navbar "Providers" dropdown in docusaurus.config.js: entries with
featured: true, labelled byshortNameorname, in catalog order - two families of redirect routes registered by the local plugin plugins/provider-redirects/index.js, each rendering src/components/ProviderRedirect/index.jsx, a Docusaurus head redirect (meta refresh plus canonical) to
https://<slug>-provider.stackql.io/:/providers/<slug>is explicit: exactly one route per catalog entry, no exceptions. Internal use (tiles, navbar)./registry/<name>is the inbound surface for external links: one route per catalog entry plus each entry'sregistryAliases, so a provider family exposes one canonical inbound link (/registry/databricks-> the Databricks Account microsite)./providersitself is the catalog doc page (docs/providers.md) now that docs live at the root, so the plugin does not register it; the bare/registryredirects to it.
The slug is derived from href, which must be exactly https://<slug>-provider.stackql.io/; the module throws at config load on a missing field, a malformed href or a duplicate slug or alias. To add a provider, add one entry to the JSON and nothing else. Do not create pages under src/pages/providers or src/pages/registry - those directories were removed and a file there would clash with the generated routes. The retired /providers/databricks URL is a Netlify 301 to /registry/databricks.
The provider microsites' own nav comes from ../docusaurus-config, which keeps its own featured list (PROVIDER_SLUGS). Update it by hand when featured changes here.
src/components/Gist/index.jsx - local replacement for the unmaintained react-gist package (was blocking React 18 upgrade). Drop-in compatible: same <Gist id="..." /> API. Used by two blog posts.
src/theme/DocItem/Footer/index.js - copy of the theme-classic doc footer with one addition: a doc with hide_last_update: true in its front matter drops the "Last updated" row while keeping tags and the edit link. Only the homepage uses it, so search results do not show a modification date on a landing page. showLastUpdateTime stays on globally. Keep this file in step with theme-classic when Docusaurus is upgraded.
src/theme/DocCard/index.js - copy of the theme-classic doc card with one addition: a sidebar item's customProps can replace the default emoji with iconComponent (a React node), icon (an image path under static/, with invertOnDark to invert it in dark mode) or emoji. Link items and category items both honour it. The tiles on docs/providers.md and the four Quick Starts provider categories in sidebars.js use icon; the Quick Starts entries look the icon up in the provider catalog by slug, so the cards follow src/configs/providers.json.
Every doc has its own description:; the old boilerplate ("Query and Deploy Cloud Infrastructure and Resources using SQL") was shared by 89 pages and is gone. Google uses the description as the snippet under a sitelink and treats repetition as a reason to withhold sitelinks, so a new doc needs a one-sentence description written from its content, not a copied one. The homepage title is deliberately descriptive ("SQL for cloud infrastructure, SaaS APIs and AI agents") with sidebar_label: Welcome to StackQL keeping the sidebar entry short.
The structured-data plugin reads several frontmatter fields directly. Use these on /ai/* pages especially.
---
title: What is StackQL?
description: One-sentence definition - this becomes the JSON-LD WebPage.description, the llms.txt entry's description, and the OG description.
keywords: [stackql, sql, cloud, api]
proficiencyLevel: Beginner # Beginner | Intermediate | Expert - sets TechArticle.proficiencyLevel
dependencies: stackql >= 0.6 # optional string - sets TechArticle.dependencies
faq:
- question: Is StackQL a database?
answer: No. StackQL is a query runtime...
- question: Does StackQL replace Terraform?
answer: Not directly. Terraform's primary job is...
---When faq: is present, the plugin emits a FAQPage JSON-LD node and links it to the TechArticle via mainEntity. The .md companion preserves the frontmatter verbatim, so an LLM ingesting the markdown gets the FAQ pairs as content too.
howTo: and softwareApplication: work the same way. See the structured-data plugin README for the full shapes.
- Build with the AEO env vars set:
ALGOLIA_APP_ID=dummy ALGOLIA_API_KEY=dummy ALGOLIA_INDEX_NAME=dummy npm run build(local builds only - production gets real values). - Check that the build emits expected
.mdcompanions:find build -name "*.md" -type f | wc -lshould be ~300. - Check that
build/llms.txtandbuild/llms-full.txtexist and are non-empty. - For
/ai/*pages, verify the JSON-LD by inspecting the rendered HTML forTechArticle+FAQPagetypes (script we wrote in earlier sessions can be reproduced if needed).
- Pick the right directory by answer intent (definition, comparison, how-to, etc.).
- Write the page in the same voice as ai-content/canonical-definitions/what-is-stackql.md - direct, declarative, technical-encyclopedia tone (think Wikipedia, not a vendor blog).
- Always include
description:in frontmatter - this drives JSON-LD, llms.txt, and OG metadata. - Use
faq:frontmatter rather than the<script type="application/json" data-aeo-faq>MDX pattern. Both work; frontmatter is cleaner. - Cross-link to related
/ai/*pages even if they don't exist yet - mark them as "not yet written" so Docusaurus'sonBrokenLinks: 'throw'doesn't fail the build. As pages are written, convert the plain-text references to real links.
- The structured-data plugin doesn't care if you change content - it re-runs on every build. But if you add
faq:,howTo:,proficiencyLevel:, ordependencies:to frontmatter, those will surface in the JSON-LD automatically.
The two AEO plugins are developed in sibling repos. To bump:
yarn add @stackql/docusaurus-plugin-structured-data@<version>
yarn add @stackql/docusaurus-plugin-aeo@<version>
If the build breaks after a bump, the plugin's CHANGELOG.md is the first place to look. Both plugins maintain detailed changelogs noting breaking changes and config migrations.
npm run start does NOT run postBuild, so:
- No
.mdcompanions are emitted - No
llms.txtis emitted - JSON-LD from the structured-data plugin is NOT injected (it runs in postBuild)
But the Ask AI button (theme component) DOES render in dev. To verify the full AEO pipeline locally, run npm run build && npm run serve instead.
After plugin changes, dev cache must be cleared: rm -rf .docusaurus && npm run start. Hard refresh the browser too (Ctrl+F5).
react-gist@1.2.4 declares react: <=17 as a peer dep. Replaced with the local src/components/Gist/index.jsx component to drop the conflict. Do not reintroduce react-gist.
When the live site changes, AI fetch tools (Claude.ai's web_fetch, ChatGPT's browse, Perplexity) may serve cached responses for some hours. If a recently-deployed .md URL appears as 404 in an LLM response, check the URL directly with curl first - if curl returns 200, the bot is using a stale cache. Wait a few hours and retest.
/install, /blog, /blog/<section>, /stackql-deploy, /stackqldocs, /downloads, generated category indexes
These routes are React pages, the blog landing plugin route, blog list pages, or generated category index pages (/getting-started, /command-line-usage, /quick-starts/*) with no source markdown. The AEO plugin correctly does not emit .md companions for them. They appear in the human nav but not in llms.txt or anywhere requiring a .md twin. This is by design - do not "fix" by trying to force .md emission. /install, /stackqldocs and /downloads are meta-refresh stubs kept for inbound links; the navbar and footer link to the canonical pages (/, /installing-stackql, /providers) instead, so crawlers see real site structure. Keep it that way when adding chrome links. /tutorials and /cookbooks are Netlify 301s, so they 404 under yarn serve.
The Ask AI button is hidden below 997px viewport width (the Docusaurus mobile breakpoint) to keep the breadcrumb row uncluttered. To verify the button on a desktop test, the browser window must be wider than 997px.
# Local dev server (no AEO postBuild, no JSON-LD)
npm run start
# Full production build (AEO + JSON-LD + .md companions)
ALGOLIA_APP_ID=dummy ALGOLIA_API_KEY=dummy ALGOLIA_INDEX_NAME=dummy npm run build
# Serve the production build locally to verify
npm run serve
# Clear dev cache after plugin changes
rm -rf .docusaurus build
# Inspect emitted JSON-LD on a page
grep -A1 'application/ld+json' build/command-line-usage/exec.html | head -20
grep -A1 'application/ld+json' build/blog/product/stackql-mcp-server-now-available.html | head -20
# Count .md companions
find build -name "*.md" -type f | wc -l
# Check llms.txt structure
grep -E '^## ' build/llms.txt
# Regenerate the per-post blog redirect block in netlify.toml
node scripts/generate-blog-redirects.js../docusaurus-plugin-structured-data- JSON-LD emission plugin source. Bug fixes for stackql.io-specific issues land here first, then ship to npm.../docusaurus-plugin-aeo- AEO plugin source. Same pattern.../../stackql-registry- StackQL provider registry. Houses the StackqlDeployDropdown component whose styling the Ask AI button was modeled on.