feat: record job queue/run time history from an incremental run feed - #19
Merged
Merged
Conversation
Replace per-workflow latest-run polling with a per-repo run feed: list runs created since a watermark, fetch jobs once per newly completed run, and record github_job_queue_seconds, github_job_run_seconds and github_workflow_runs_total. The github_repo_ci_last_run_* snapshot is kept, seeded by a startup bootstrap. Measured on drumandbytes: 106 API calls per org refresh, down from 236.
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Replaces the per-workflow latest-run polling with an incremental run feed per repo, so Prometheus can hold job run/queue time history (avg / p95 per repo, workflow, job, runner and runner pool).
What changes
/actions/runs?created=>=…, paginated) and fetches/runs/{id}/jobsonce for each newly completed run.github_job_queue_secondshistogram:created_at→started_atgithub_job_run_secondshistogram:started_at→completed_atgithub_workflow_runs_total{repo,workflow,conclusion}counterrepo, workflow, job, runner, runner_label, conclusion.runner_labelis the job'sruns-onlabels minusself-hosted(oracle-x64/oracle-arm64). GitHub-hosted runners are collapsed torunner="github-hosted". Buckets run from 5s to 60m.github_repo_ci_last_run_*keep their meaning. They're updated from the feed and filled at startup by one bootstrap pass withLatestRunForWorkflow. They no longer disappear while a workflow's newest run is in progress.status=completedfilter: with it, a slow run created before a faster one would be skipped for good. An in-progress run holds the watermark back for at most 24h.repo/windowvariables and a CI History row with:API calls
Measured on drumandbytes (26 repos, 156 active workflows, ~1000 runs/week) by counting real requests against the org:
Startup bootstrap is a one-off 262. With the runner poll, total usage is ~28% of the 5,000/h budget, down from ~59%.
Tests
These are the first Go unit tests in the repo:
internal/github: pagination for runs and jobs, and the encoding of thecreated>=filter, against an httptest server.internal/orgstats, using a fake client:Build's per-repo error swallowing and its error propagation