Skip to content

ci: move all CI workflows to the dedicated ExtendDB runner pool - #379

Merged
yesyayen merged 5 commits into
mainfrom
ci/dedicated-runner-pool
Oct 2, 2026
Merged

yesyayen merged 5 commits into
mainfrom
ci/dedicated-runner-pool

Conversation

@robinnsc

@robinnsc robinnsc commented Oct 1, 2026 •

Copy link
Copy Markdown
Collaborator

What

Three changes across all six CI workflows (integration, integration-mongodb, test, clippy, fmt, licenses — 16 jobs):

  1. runs-on: extenddb_ubuntu-2404_4-core on every job — the dedicated runner group provisioned for this repo (4 vCPU / 16 GB, matching ubuntu-latest; capped at 20 concurrent). Zero ubuntu-latest remains in CI workflows. Pinned to Ubuntu 24.04 deliberately: ubuntu-latest migrates to 26.04 between Oct 19 and Nov 19, and we compile Rust on it — we move to 26 when we choose.
  2. Per-job timeout-minutes, sized from the last several runs on main: 30 for the two integration suites (longest observed jobs 19:08 and 18 min), 15 for the sub-5-minute workflows (test 3 min, clippy/fmt ~1 min); licenses keeps its existing 20. None was set before; on a capped pool a hung job holds a slot for hours.
  3. A merge-queue-safe concurrency block: a new push supersedes the in-flight run of the same PR; everything else is keyed by run_id. The run_id part comes from an independent audit of an earlier revision: merge-group runs reuse a ref across re-queues of one PR (observed on three real runs, two overlapping), and GitHub replaces a pending run in a group regardless of cancel-in-progress — so a ref-keyed group could cancel a merge-group run and evict its queue entry.

Why

ubuntu-latest here is an Amazon-wide pool: 1,000 concurrent jobs, first-come first-served across every org, no reservations. On Sept 29–30 our jobs waited 10–90 minutes to start behind other teams' load — a 13-second fmt job waited 15 minutes for a runner. That, not any failing check, broke the merge queue: its 30-minute status-check timeout expired before jobs had begun, five evictions in a row. The integration jobs run in parallel, so queue wait was the whole problem.

Confirmed and provisioned by the runner team (~742 min/month usage, under $10/month).

Testing done

  • yaml.safe_load on all six; 16 runs-on swapped, a timeout on every job, one concurrency block per file.
  • Live on this branch: all six workflows ran on the new pool. Every job was picked up in 24–27 seconds (vs. 10–90 minutes on Sept 29–30). All service-container jobs (pgvector, postgres:16, MongoDB) succeeded, not merely started — same Ubuntu 24.04 image release as GitHub-hosted, no missing preinstalled tooling.
  • Independent adversarial audit of the integration.yml revision: found the concurrency defect above (fixed), confirmed the aggregator job fails closed on cancelled upstreams, confirmed timeout sizing against actual job durations.

Checklist

  • Tests / fmt / clippy — not applicable, CI only
  • Documentation — inline comments explain the pool, the pin, and the concurrency rationale
  • Breaking changes noted below

ADR / RFC: n/a — CI infrastructure only.

Breaking changes

None.


By submitting this pull request, I confirm that my contribution is made under the terms of the Apache License 2.0 and I agree to the Developer Certificate of Origin (DCO). See CONTRIBUTING.md for details.

ubuntu-latest is an Amazon-wide pool of 1,000 concurrent jobs shared
first-come first-served across every org; our integration jobs sat
queued 10-90 minutes behind other teams' load on 2026-09-29/30,
which is what broke the merge queue (its 30-minute check timeout
expired before jobs even started). All seven jobs now run on
extenddb_ubuntu-2404_4-core: scoped to this repo, 4 vCPU / 16 GB
(matching ubuntu-latest), capped at 20 concurrent, pinned to Ubuntu
24.04 ahead of ubuntu-latest's Oct-Nov migration to 26.04.

Also: timeout-minutes: 30 on every job (none was set; on a capped
pool a hung job holds a slot for hours), and a merge-queue-safe
concurrency block that supersedes an in-flight run of the same PR on
a new push without touching merge-group runs.
Independent audit finding: merge-group runs reuse the same ref
(gh-readonly-queue/main/pr-N-<sha> recurs across re-queues of one PR,
observed on three runs with different SHAs, two overlapping), and
GitHub concurrency replaces a PENDING run in a group even with
cancel-in-progress false. Keying non-PR runs by ref therefore let one
merge-group run cancel another and evict its queue entry — the exact
failure the block claimed to prevent. Non-PR runs now get a unique
group (run_id); PR runs keep the per-PR supersede. Also corrects the
'seven slots' comment (six upstream jobs run concurrently).
@robinnsc

robinnsc commented Oct 1, 2026

Copy link
Copy Markdown
Collaborator Author

Pushed a fix from an independent audit of this PR: the concurrency group keyed non-PR runs by github.ref, but merge-group runs reuse the same ref across re-queues of one PR (observed on three real runs, two overlapping), and GitHub replaces a pending run in a group regardless of cancel-in-progress — so the block could cancel a merge-group run, the exact thing it claimed to prevent. Non-PR runs now key on run_id (always unique); PR runs keep the per-PR supersede. Also corrected the body's "~2× the longest job" claim: longest observed is 19:08 (run-integration-sqlite), so 30 min is ~1.6×, still comfortable. Everything else in the audit checked out: all seven jobs on the new label, all service-container jobs succeeded on the dedicated runners (same Ubuntu 24.04 image release as GitHub-hosted), no missing preinstalled tooling, aggregator fails closed on cancellation.

Same three changes #379 made to integration.yml, applied to
integration-mongodb, test, clippy, fmt, and licenses: runs-on
extenddb_ubuntu-2404_4-core (pinned Ubuntu 24.04), a per-job
timeout (30 min for the MongoDB suite whose longest job is 18 min;
15 min for the sub-5-minute workflows; licenses keeps its existing
20), and the corrected concurrency block — per-PR supersede on
pull_request, run_id-unique groups for everything else so merge-group
runs can never cancel one another.
@robinnsc robinnsc changed the title ci(integration): move to the dedicated ExtendDB runner pool ci: move all CI workflows to the dedicated ExtendDB runner pool Oct 1, 2026
Second independent audit: four integration jobs reached 18-20 min on
recent main runs (run-integration-sqlite 20:03), i.e. above 60% of
the 30-minute cap. A timeout exists to catch HUNG jobs; at 30 min it
would also clip the slow tail of healthy ones on a bad runner day and
turn a passing merge-group round into a blocked merge. 40 min is 2x
the worst observed duration while still bounding a hang.
@robinnsc

robinnsc commented Oct 1, 2026

Copy link
Copy Markdown
Collaborator Author

Second independent audit (real-defects-only pass), two findings:

Fixed here: the 30-minute timeout on the integration suites was too tight — four jobs hit 18–20 min on recent main runs (run-integration-sqlite 20:03), so a slow-runner day would have turned passing jobs into spurious failures that block merges. Raised to 40 min (2× the worst observed; still bounds a genuine hang). The fast workflows' 15-min caps are fine (jobs run 1–3 min).

Not fixable in this PR, flagged for the admin discussion: the repo currently has no required status checks configured at all — the active main ruleset carries only deletion / non-fast-forward / pull-request rules, and classic branch protection is absent. So the integration and mongodb-integration aggregator jobs don't actually gate anything today; a merge with a red integration round would go through. The merge-queue requirement that blocked Monday's admin merge therefore lives in an org/enterprise ruleset above the repo. Both belong in the same admin conversation already planned: require the two aggregator checks, and raise the queue's status-check timeout.

Everything else cleared with runtime evidence: concurrency expression evaluates correctly for every event type (and a PR push did supersede the prior run on this branch — run 36939668898 replaced 36938931696, while merge-group/push runs get unique groups); job and matrix names identical to main so no required-check rename risk; runner label resolves to a GitHub-hosted larger runner in group extenddb (not an unmatched self-hosted label); all service containers healthy, 511 MongoDB tests passed on the pool; no permissions/secrets/env changes in the diff; aggregators use if: always() and fail closed on cancelled upstreams (confirmed on a real cancelled run).

@jcshepherd jcshepherd left a comment

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

I see a few automated suggestions from CodeQL but I think we can review and address those at another time, if we wish. Rest looks good.

@yesyayen
yesyayen added this pull request to the merge queue Oct 2, 2026
Merged via the queue into main with commit 67a619d Oct 2, 2026
22 checks passed
@yesyayen
yesyayen deleted the ci/dedicated-runner-pool branch October 2, 2026 18:08
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

3 participants