Skip to content

Add bin/pipeline.sh: sync, compact, publish the dashboard, from cron - #25

Merged
edsu merged 2 commits into
mainfrom
pipeline-script
Aug 24, 2026
Merged

edsu merged 2 commits into
mainfrom
pipeline-script

Conversation

@edsu

@edsu edsu commented Aug 24, 2026 •

Copy link
Copy Markdown
Collaborator

Adds bin/pipeline.sh, the scheduled pipeline: sync-all, compact, then podlake-web's refresh. The push that refresh makes is what deploys the live dashboard — via podlake-web's own Actions workflow — so nothing here touches Pages directly.

Running in cron on the production host for a few weeks now; it has been working well.

Why one script rather than three cron entries

The ordering is load-bearing. sync-all applies each POD resource in its own transaction, so a reader never sees a torn write — but it can see a lake where one institution's latest dump has landed and another's has not. Every headline figure on the dashboard compares institutions to each other, so publishing from a part-way-synced lake produces numbers that are wrong in a way that looks entirely plausible. Hence set -e and a failed sync publishing nothing: stale beats half-synced. Separate cron entries can't express that either, since a sync with months of catch-up runs for hours.

Notes on the things that bite under cron

  • Defaults need no configuration. Everything derives from the script's own location, assuming the sibling layout podlake-web already requires (it depends on podlake by path). PODLAKE_DIR, POD_ROOT, WEB_DIR, LOG_DIR, LOCK_FILE, CATALOG, PUBLISH_BRANCH, DEPLOY_KEY override the pieces.
  • It never updates either checkout. Deploying code is a decision, not a side effect of the schedule. It does check that podlake-web is current first, and if that checkout is behind upstream it syncs the lake anyway, skips the publish, and exits nonzero. The check runs up front because the failure it replaces is a quiet one: refresh pushes HEAD, so a stale checkout has its push rejected after the hour-long extract, and the next run finds a clean tree, sees no change against its own unpushed commit, and exits 0 — success reported while the site silently stops updating.
  • Failure notification. cron mails based on whether the job wrote something, not on exit status, and every step's output goes to the log. So the script keeps real stderr on fd 8 and writes a verdict plus the log tail there on any nonzero exit: a good run is silent, a bad one is actionable mail.
  • flock keeps a long catch-up sync from piling up on the next window; its absence is checked explicitly, since a bare failure would look exactly like "another run holds the lock", i.e. like a successful no-op.
  • PATH gets ~/.local/bin for uv — the classic reason a job that works by hand fails on a schedule.

PUBLISH_BRANCH set to anything but main stages the figures on a branch for review instead of updating the live site.

README

A new section covers scheduling it, plus setting up the GitHub credential: cron has no SSH agent, so it must be a passphrase-less per-repo deploy key with write access, with ssh -T git@github.com run once to seed known_hosts. DEPLOY_KEY pins it via GIT_SSH_COMMAND with IdentitiesOnly=yes — without that, ssh offers every key it can find, GitHub accepts the first that authenticates, and a deploy key for a different repo authenticates fine, after which the push fails as "repository not found".

🤖 Generated with Claude Code

edsu and others added 2 commits August 20, 2026 16:20
The three steps have to happen in order and nothing expressed that. sync-all
applies each resource in its own transaction, so a reader never sees a torn
write — but it can see a lake where one institution's latest dump has landed and
another's has not, and every headline figure on podlake-web compares
institutions to each other. Separate cron entries can't express the ordering
either, since a sync with months of catch-up runs for hours. So: one script,
set -e, and a failed sync publishing nothing. Stale beats half-synced.

Paths derive from the script's own location, so the sibling checkouts podlake-web
already requires (it depends on podlake by path) need no configuration.

It never updates either checkout. Deploying code is a decision, not a side effect
of the schedule. But it does check that podlake-web is current first, because the
failure that replaces is a quiet one: refresh pushes HEAD, so a stale checkout
has its push rejected *after* the hour-long extract, and the run after that finds
a clean tree, sees no change against its own unpushed commit, and exits 0 —
reporting success while the site silently stops updating. Being behind blocks
only the publish; syncing the lake is useful work either way.

cron mails on output, not on exit status, and every step here is redirected to a
log — so a failure would otherwise be entirely silent. Keep the real stderr on
fd 8 and report there on any nonzero exit, with the log tail, so the mail is
actionable without logging into the host. The two deliberate exit 0 paths (lock
held, nothing to publish) stay quiet.

DEPLOY_KEY sets GIT_SSH_COMMAND with IdentitiesOnly=yes. Without it ssh offers
every key it can find, GitHub accepts the first that authenticates, and a deploy
key for a different repo authenticates fine — after which the push fails as
"repository not found", which reads like a permissions problem and isn't.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
It sat inside "Building the lake" as a subsection, between the build cycle and
the query examples — which put a hundred lines of cron, fd juggling and deploy
keys in front of every reader who only wanted to know how to query the lake.
Scheduling is a maintainer's last step, so it reads better after both halves it
composes.

Now a top-level section (its own subsections promoted a level to match), and the
opening sentence names the cycle it refers to now that it's no longer adjacent.
@edsu
edsu merged commit 92b38f3 into main Aug 24, 2026
1 check passed
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant