Skip to content

feat: poll GitHub on a timer instead of on scrape - #21

Merged
JustMaris merged 1 commit into
mainfrom
feat/background-refresh
Oct 4, 2026
Merged

JustMaris merged 1 commit into
mainfrom
feat/background-refresh

Conversation

@JustMaris

Copy link
Copy Markdown
Member

The runner busy panel lagged: when the cache expired, a scrape got the stale value and only kicked off a refresh, so new state appeared one scrape later still. Short jobs often never showed as busy.

Changes

  • fetch.Fetcher now polls on its own timer (the domain's TTL), independent of scrapes. A scrape always reads data at most one interval old, and Get never blocks. It returns the latest result, an error until the first poll finishes, or an error past max-stale.
  • RUNNER_CACHE_TTL now defaults to 15s (was 30s). A runner picking up a job shows within ~15s plus one scrape.
  • Collector.Describe lists its descriptors explicitly instead of DescribeByCollect. Registering no longer waits for the first GitHub poll, so /healthz answers immediately; locally it took 1s, previously ~85s.
  • github_org_up / github_runners_up are 0 until their first poll finishes (~80s for the org poll on drumandbytes). GithubOrgStatsStale waits 30m, so this doesn't fire alerts.

No setup needed: GitHub webhooks were considered and rejected as overkill. They need a public endpoint, a webhook secret and an org webhook, and polling would still be needed as a backstop.

API budget

Runner polling goes from ~120 to ~240 calls/h. The org poll is unchanged (~1,280/h), so the total is ~1,520/h, about 30% of 5,000/h. Polling now also runs when nothing scrapes, which costs the same as being scraped continuously.

Tests

  • New internal/fetch tests:
    • error before the first poll and after a failed first poll
    • a stale value is served after a failed refresh
    • nothing is served past max-stale
    • the loop refreshes with no Get calls
  • go test -race ./... passes.
  • Local run against drumandbytes: /healthz returned 200 after 1s, runner metrics were present after 1s, and the org metrics appeared after ~80s.

Each cache now refreshes itself every TTL in the background, so a scrape
always reads data at most one interval old instead of getting the stale
value and triggering a refresh for the next scrape. Get never blocks, and
Describe lists its descriptors instead of collecting, so /healthz answers
immediately instead of after the first org poll.

RUNNER_CACHE_TTL now defaults to 15s (was 30s).
@JustMaris
JustMaris merged commit ae644a6 into main Oct 4, 2026
8 checks passed
@JustMaris
JustMaris deleted the feat/background-refresh branch October 4, 2026 16:09
@dnb-robot dnb-robot Bot mentioned this pull request Oct 4, 2026
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant