feat: poll GitHub on a timer instead of on scrape - #21
Merged
Merged
Conversation
Each cache now refreshes itself every TTL in the background, so a scrape always reads data at most one interval old instead of getting the stale value and triggering a refresh for the next scrape. Get never blocks, and Describe lists its descriptors instead of collecting, so /healthz answers immediately instead of after the first org poll. RUNNER_CACHE_TTL now defaults to 15s (was 30s).
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
The runner busy panel lagged: when the cache expired, a scrape got the stale value and only kicked off a refresh, so new state appeared one scrape later still. Short jobs often never showed as busy.
Changes
fetch.Fetchernow polls on its own timer (the domain's TTL), independent of scrapes. A scrape always reads data at most one interval old, andGetnever blocks. It returns the latest result, an error until the first poll finishes, or an error past max-stale.RUNNER_CACHE_TTLnow defaults to 15s (was 30s). A runner picking up a job shows within ~15s plus one scrape.Collector.Describelists its descriptors explicitly instead ofDescribeByCollect. Registering no longer waits for the first GitHub poll, so/healthzanswers immediately; locally it took 1s, previously ~85s.github_org_up/github_runners_upare 0 until their first poll finishes (~80s for the org poll on drumandbytes).GithubOrgStatsStalewaits 30m, so this doesn't fire alerts.No setup needed: GitHub webhooks were considered and rejected as overkill. They need a public endpoint, a webhook secret and an org webhook, and polling would still be needed as a backstop.
API budget
Runner polling goes from ~120 to ~240 calls/h. The org poll is unchanged (~1,280/h), so the total is ~1,520/h, about 30% of 5,000/h. Polling now also runs when nothing scrapes, which costs the same as being scraped continuously.
Tests
internal/fetchtests:Getcallsgo test -race ./...passes./healthzreturned 200 after 1s, runner metrics were present after 1s, and the org metrics appeared after ~80s.