ci: move all CI workflows to the dedicated ExtendDB runner pool - #379
Conversation
ubuntu-latest is an Amazon-wide pool of 1,000 concurrent jobs shared first-come first-served across every org; our integration jobs sat queued 10-90 minutes behind other teams' load on 2026-09-29/30, which is what broke the merge queue (its 30-minute check timeout expired before jobs even started). All seven jobs now run on extenddb_ubuntu-2404_4-core: scoped to this repo, 4 vCPU / 16 GB (matching ubuntu-latest), capped at 20 concurrent, pinned to Ubuntu 24.04 ahead of ubuntu-latest's Oct-Nov migration to 26.04. Also: timeout-minutes: 30 on every job (none was set; on a capped pool a hung job holds a slot for hours), and a merge-queue-safe concurrency block that supersedes an in-flight run of the same PR on a new push without touching merge-group runs.
Independent audit finding: merge-group runs reuse the same ref (gh-readonly-queue/main/pr-N-<sha> recurs across re-queues of one PR, observed on three runs with different SHAs, two overlapping), and GitHub concurrency replaces a PENDING run in a group even with cancel-in-progress false. Keying non-PR runs by ref therefore let one merge-group run cancel another and evict its queue entry — the exact failure the block claimed to prevent. Non-PR runs now get a unique group (run_id); PR runs keep the per-PR supersede. Also corrects the 'seven slots' comment (six upstream jobs run concurrently).
|
Pushed a fix from an independent audit of this PR: the concurrency group keyed non-PR runs by |
Same three changes #379 made to integration.yml, applied to integration-mongodb, test, clippy, fmt, and licenses: runs-on extenddb_ubuntu-2404_4-core (pinned Ubuntu 24.04), a per-job timeout (30 min for the MongoDB suite whose longest job is 18 min; 15 min for the sub-5-minute workflows; licenses keeps its existing 20), and the corrected concurrency block — per-PR supersede on pull_request, run_id-unique groups for everything else so merge-group runs can never cancel one another.
Second independent audit: four integration jobs reached 18-20 min on recent main runs (run-integration-sqlite 20:03), i.e. above 60% of the 30-minute cap. A timeout exists to catch HUNG jobs; at 30 min it would also clip the slow tail of healthy ones on a bad runner day and turn a passing merge-group round into a blocked merge. 40 min is 2x the worst observed duration while still bounding a hang.
|
Second independent audit (real-defects-only pass), two findings: Fixed here: the 30-minute timeout on the integration suites was too tight — four jobs hit 18–20 min on recent Not fixable in this PR, flagged for the admin discussion: the repo currently has no required status checks configured at all — the active Everything else cleared with runtime evidence: concurrency expression evaluates correctly for every event type (and a PR push did supersede the prior run on this branch — run 36939668898 replaced 36938931696, while merge-group/push runs get unique groups); job and matrix names identical to |
jcshepherd
left a comment
There was a problem hiding this comment.
I see a few automated suggestions from CodeQL but I think we can review and address those at another time, if we wish. Rest looks good.
What
Three changes across all six CI workflows (
integration,integration-mongodb,test,clippy,fmt,licenses— 16 jobs):runs-on: extenddb_ubuntu-2404_4-coreon every job — the dedicated runner group provisioned for this repo (4 vCPU / 16 GB, matchingubuntu-latest; capped at 20 concurrent). Zeroubuntu-latestremains in CI workflows. Pinned to Ubuntu 24.04 deliberately:ubuntu-latestmigrates to 26.04 between Oct 19 and Nov 19, and we compile Rust on it — we move to 26 when we choose.timeout-minutes, sized from the last several runs on main: 30 for the two integration suites (longest observed jobs 19:08 and 18 min), 15 for the sub-5-minute workflows (test 3 min, clippy/fmt ~1 min);licenseskeeps its existing 20. None was set before; on a capped pool a hung job holds a slot for hours.concurrencyblock: a new push supersedes the in-flight run of the same PR; everything else is keyed byrun_id. Therun_idpart comes from an independent audit of an earlier revision: merge-group runs reuse a ref across re-queues of one PR (observed on three real runs, two overlapping), and GitHub replaces a pending run in a group regardless ofcancel-in-progress— so a ref-keyed group could cancel a merge-group run and evict its queue entry.Why
ubuntu-latesthere is an Amazon-wide pool: 1,000 concurrent jobs, first-come first-served across every org, no reservations. On Sept 29–30 our jobs waited 10–90 minutes to start behind other teams' load — a 13-secondfmtjob waited 15 minutes for a runner. That, not any failing check, broke the merge queue: its 30-minute status-check timeout expired before jobs had begun, five evictions in a row. The integration jobs run in parallel, so queue wait was the whole problem.Confirmed and provisioned by the runner team (~742 min/month usage, under $10/month).
Testing done
yaml.safe_loadon all six; 16runs-onswapped, a timeout on every job, one concurrency block per file.pgvector,postgres:16, MongoDB) succeeded, not merely started — same Ubuntu 24.04 image release as GitHub-hosted, no missing preinstalled tooling.Checklist
ADR / RFC: n/a — CI infrastructure only.
Breaking changes
None.
By submitting this pull request, I confirm that my contribution is made under the terms of the Apache License 2.0 and I agree to the Developer Certificate of Origin (DCO). See CONTRIBUTING.md for details.