Repository navigation
Add bin/pipeline.sh: sync, compact, publish the dashboard, from cron - #25
Merged
Merged
Conversation
The three steps have to happen in order and nothing expressed that. sync-all applies each resource in its own transaction, so a reader never sees a torn write — but it can see a lake where one institution's latest dump has landed and another's has not, and every headline figure on podlake-web compares institutions to each other. Separate cron entries can't express the ordering either, since a sync with months of catch-up runs for hours. So: one script, set -e, and a failed sync publishing nothing. Stale beats half-synced. Paths derive from the script's own location, so the sibling checkouts podlake-web already requires (it depends on podlake by path) need no configuration. It never updates either checkout. Deploying code is a decision, not a side effect of the schedule. But it does check that podlake-web is current first, because the failure that replaces is a quiet one: refresh pushes HEAD, so a stale checkout has its push rejected *after* the hour-long extract, and the run after that finds a clean tree, sees no change against its own unpushed commit, and exits 0 — reporting success while the site silently stops updating. Being behind blocks only the publish; syncing the lake is useful work either way. cron mails on output, not on exit status, and every step here is redirected to a log — so a failure would otherwise be entirely silent. Keep the real stderr on fd 8 and report there on any nonzero exit, with the log tail, so the mail is actionable without logging into the host. The two deliberate exit 0 paths (lock held, nothing to publish) stay quiet. DEPLOY_KEY sets GIT_SSH_COMMAND with IdentitiesOnly=yes. Without it ssh offers every key it can find, GitHub accepts the first that authenticates, and a deploy key for a different repo authenticates fine — after which the push fails as "repository not found", which reads like a permissions problem and isn't. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
It sat inside "Building the lake" as a subsection, between the build cycle and the query examples — which put a hundred lines of cron, fd juggling and deploy keys in front of every reader who only wanted to know how to query the lake. Scheduling is a maintainer's last step, so it reads better after both halves it composes. Now a top-level section (its own subsections promoted a level to match), and the opening sentence names the cycle it refers to now that it's no longer adjacent.
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Adds
bin/pipeline.sh, the scheduled pipeline:sync-all,compact, then podlake-web'srefresh. The push thatrefreshmakes is what deploys the live dashboard — via podlake-web's own Actions workflow — so nothing here touches Pages directly.Running in cron on the production host for a few weeks now; it has been working well.
Why one script rather than three cron entries
The ordering is load-bearing.
sync-allapplies each POD resource in its own transaction, so a reader never sees a torn write — but it can see a lake where one institution's latest dump has landed and another's has not. Every headline figure on the dashboard compares institutions to each other, so publishing from a part-way-synced lake produces numbers that are wrong in a way that looks entirely plausible. Henceset -eand a failed sync publishing nothing: stale beats half-synced. Separate cron entries can't express that either, since a sync with months of catch-up runs for hours.Notes on the things that bite under cron
PODLAKE_DIR,POD_ROOT,WEB_DIR,LOG_DIR,LOCK_FILE,CATALOG,PUBLISH_BRANCH,DEPLOY_KEYoverride the pieces.refreshpushesHEAD, so a stale checkout has its push rejected after the hour-long extract, and the next run finds a clean tree, sees no change against its own unpushed commit, and exits 0 — success reported while the site silently stops updating.flockkeeps a long catch-up sync from piling up on the next window; its absence is checked explicitly, since a bare failure would look exactly like "another run holds the lock", i.e. like a successful no-op.PATHgets~/.local/binforuv— the classic reason a job that works by hand fails on a schedule.PUBLISH_BRANCHset to anything butmainstages the figures on a branch for review instead of updating the live site.README
A new section covers scheduling it, plus setting up the GitHub credential: cron has no SSH agent, so it must be a passphrase-less per-repo deploy key with write access, with
ssh -T git@github.comrun once to seedknown_hosts.DEPLOY_KEYpins it viaGIT_SSH_COMMANDwithIdentitiesOnly=yes— without that, ssh offers every key it can find, GitHub accepts the first that authenticates, and a deploy key for a different repo authenticates fine, after which the push fails as "repository not found".🤖 Generated with Claude Code