Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
10 changes: 6 additions & 4 deletions CLAUDE.md
Original file line number Diff line number Diff line change
Expand Up @@ -19,7 +19,9 @@ depending on which TTL "wins".
- `internal/github` — REST client for every GitHub endpoint this exporter
calls (`orgs/{org}/actions/runners`, `orgs/{org}/repos`,
`repos/{owner}/{repo}/actions/workflows`, `.../actions/workflows/{id}/runs`,
`.../pulls`, `.../dependabot/alerts`, `/rate_limit`). `OpenPRCount` reads
`.../pulls`, `.../dependabot/alerts`). The rate-limit budget comes from
the `X-RateLimit-*` headers of those responses, not `GET /rate_limit`,
which reports used=0 for our tokens. `OpenPRCount` reads
the page count off the `Link` response header instead of paginating —
deliberate: it's one request regardless of how many open PRs a repo has,
and stays on the *core* rate limit rather than the separately-throttled
Expand All @@ -34,8 +36,8 @@ depending on which TTL "wins".
the client's raw responses into each domain's `Summary`. `orgstats.Build`
swallows per-repo call failures deliberately (one repo the token can't
see, or with Dependabot disabled, shouldn't blank out every other
repo's data) — but propagates a failure of `Repos`/`RateLimit`
themselves, since nothing else can proceed without those.
repo's data) — but propagates a failure of `Repos` itself, since nothing else can
proceed without it.
- `internal/collector` — adapts both `Summary` types into Prometheus
metrics for `/metrics`, on one shared `prometheus.Collector`.
- `internal/config` — env var parsing (see README's Configuration table).
Expand All @@ -52,7 +54,7 @@ The watermark is pinned by the oldest in-progress run (capped at 24h) —
don't add `status=completed` to the runs query, or a slow run created
before a faster one gets skipped forever. Dedupe is by (run ID, attempt)
and job ID, kept 48h. A restart starts from "now"; nothing is replayed.
Measured on drumandbytes (26 repos, 156 active workflows): ~106 calls per
Measured on drumandbytes (26 repos, 156 active workflows): ~105 calls per
refresh vs ~236 with the old per-workflow polling.

## Build / test / run
Expand Down
16 changes: 14 additions & 2 deletions README.md
Original file line number Diff line number Diff line change
Expand Up @@ -43,8 +43,9 @@ scrape.
| `github_runner_busy` | `runner`, `os` | 1 if the runner is currently executing a job, 0 if idle |
| `github_org_up` | | 1 if the last org/repo stats poll succeeded, 0 if a stale cache is being served or the first poll is still running (about a minute after startup at ~25 repos) |
| `github_org_repos_total` | `visibility` | Number of non-archived repos, by `public`/`private` |
| `github_rate_limit_remaining` | | Remaining core API rate-limit budget |
| `github_rate_limit_limit` | | Total core API rate-limit budget |
| `github_rate_limit_remaining` | | Remaining core API budget in the scarcest active window: the one that runs out first |
| `github_rate_limit_limit` | | Total core API budget of that window |
| `github_rate_limit_window_remaining` | `reset` | Remaining budget per active core window, labelled with its reset time. GitHub counts per region, so one token can have several windows at once, and which one a call counts against depends on the endpoint. A window's series disappears once it resets |
| `github_repo_open_prs` | `repo` | Number of open pull requests |
| `github_repo_ci_last_run_conclusion` | `repo`, `workflow`, `url`, `conclusion` | Always 1 - an "info" metric. `conclusion` is GitHub's own string verbatim (`success`, `failure`, `cancelled`, `skipped`, `neutral`, `timed_out`, `action_required`, `stale`), not collapsed to pass/fail here - what counts as "actually broken" is a dashboard-level call. `url` links to the run on github.com. Absent if the workflow has never run |
| `github_repo_ci_last_run_timestamp_seconds` | `repo`, `workflow` | Unix timestamp of that workflow's latest completed run |
Expand Down Expand Up @@ -97,6 +98,17 @@ histogram_quantile(0.95, sum by (le, runner_label) (rate(github_job_queue_second
sum by (conclusion) (increase(github_workflow_runs_total[1d]))
```

## Rate limit

The budget comes from the `X-RateLimit-*` headers of the exporter's own API
responses, which GitHub documents as authoritative. `GET /rate_limit` isn't
used: it can disagree with the headers, and for our tokens it reported
`used: 0` throughout. Because GitHub serves requests from several regions, a
token can be counted in two `core` windows at once, with different reset
times and counts. On drumandbytes, runners, workflows and Dependabot alerts
landed in one window and everything else in another. The exporter tracks every
window it sees and alerts on the lowest.

## Configuration

| Env var | Default | Required |
Expand Down
14 changes: 11 additions & 3 deletions dashboards/github-actions-runner-exporter.json
Original file line number Diff line number Diff line change
Expand Up @@ -8,7 +8,7 @@
"prometheus"
],
"schemaVersion": 39,
"version": 2,
"version": 3,
"editable": true,
"time": {
"from": "now-24h",
Expand Down Expand Up @@ -952,12 +952,20 @@
"targets": [
{
"expr": "max(github_rate_limit_remaining)",
"legendFormat": "remaining",
"legendFormat": "remaining (scarcest window)",
"range": true,
"instant": false,
"refId": "A"
},
{
"expr": "max by (reset) (github_rate_limit_window_remaining)",
"legendFormat": "window resetting {{reset}}",
"range": true,
"instant": false,
"refId": "B"
}
]
],
"description": "remaining: the scarcest core window, which is what runs out first. GitHub counts per region, so a token can have several windows at once; each shows as its own line until it resets."
},
{
"id": 18,
Expand Down
33 changes: 27 additions & 6 deletions internal/collector/collector.go
Original file line number Diff line number Diff line change
Expand Up @@ -8,6 +8,7 @@ import (
"github.com/prometheus/client_golang/prometheus"

"github.com/drumandbytes/github-actions-runner-exporter/internal/fetch"
"github.com/drumandbytes/github-actions-runner-exporter/internal/github"
"github.com/drumandbytes/github-actions-runner-exporter/internal/orgstats"
"github.com/drumandbytes/github-actions-runner-exporter/internal/runners"
)
Expand All @@ -17,6 +18,7 @@ const namespace = "github"
type Collector struct {
runnerFetcher *fetch.Fetcher[runners.Summary]
orgFetcher *fetch.Fetcher[orgstats.Summary]
rateLimitFn func() map[time.Time]github.RateLimit

runnerUp *prometheus.Desc
runnerBusy *prometheus.Desc
Expand All @@ -25,6 +27,7 @@ type Collector struct {
reposTotal *prometheus.Desc
rateLimit *prometheus.Desc
rateLimitCap *prometheus.Desc
rateWindow *prometheus.Desc

repoOpenPRs *prometheus.Desc
repoCILastRunConclusion *prometheus.Desc
Expand All @@ -34,7 +37,7 @@ type Collector struct {
}

// New builds the collector. The intervals only feed HELP text; polling is the Fetchers' job.
func New(runnerFetcher *fetch.Fetcher[runners.Summary], orgFetcher *fetch.Fetcher[orgstats.Summary], runnerCacheTTL, orgCacheTTL time.Duration) *Collector {
func New(runnerFetcher *fetch.Fetcher[runners.Summary], orgFetcher *fetch.Fetcher[orgstats.Summary], rateLimits func() map[time.Time]github.RateLimit, runnerCacheTTL, orgCacheTTL time.Duration) *Collector {
runnerNote := fmt.Sprintf(" Polled every %s.", runnerCacheTTL)
orgNote := fmt.Sprintf(" Polled every %s - not real-time by design, see internal/orgstats.", orgCacheTTL)
desc := func(subsystem, name, help string, labels []string) *prometheus.Desc {
Expand All @@ -44,6 +47,7 @@ func New(runnerFetcher *fetch.Fetcher[runners.Summary], orgFetcher *fetch.Fetche
return &Collector{
runnerFetcher: runnerFetcher,
orgFetcher: orgFetcher,
rateLimitFn: rateLimits,

runnersUp: desc("runners", "up",
"Whether the last scrape of the runners API succeeded (1) or a stale cache is being served (0)."+runnerNote, nil),
Expand All @@ -57,9 +61,11 @@ func New(runnerFetcher *fetch.Fetcher[runners.Summary], orgFetcher *fetch.Fetche
reposTotal: desc("org", "repos_total",
"Number of non-archived repos in the org."+orgNote, []string{"visibility"}),
rateLimit: desc("rate_limit", "remaining",
"Remaining core API rate-limit budget."+orgNote, nil),
"Remaining core API rate-limit budget, as of the latest GitHub response.", nil),
rateLimitCap: desc("rate_limit", "limit",
"Total core API rate-limit budget."+orgNote, nil),
"Total core API rate-limit budget.", nil),
rateWindow: desc("rate_limit", "window_remaining",
"Remaining budget per core rate-limit window. GitHub counts per region, so a token can have several windows at once; github_rate_limit_remaining is the lowest.", []string{"reset"}),

repoOpenPRs: desc("repo", "open_prs",
"Number of open pull requests."+orgNote, []string{"repo"}),
Expand All @@ -82,7 +88,7 @@ func New(runnerFetcher *fetch.Fetcher[runners.Summary], orgFetcher *fetch.Fetche
func (c *Collector) Describe(ch chan<- *prometheus.Desc) {
for _, d := range []*prometheus.Desc{
c.runnersUp, c.runnerUp, c.runnerBusy,
c.orgUp, c.reposTotal, c.rateLimit, c.rateLimitCap,
c.orgUp, c.reposTotal, c.rateLimit, c.rateLimitCap, c.rateWindow,
c.repoOpenPRs, c.repoCILastRunConclusion, c.repoCILastRunAt, c.repoCIDuration, c.repoDependabot,
} {
ch <- d
Expand All @@ -92,6 +98,23 @@ func (c *Collector) Describe(ch chan<- *prometheus.Desc) {
func (c *Collector) Collect(ch chan<- prometheus.Metric) {
c.collectRunners(ch)
c.collectOrgStats(ch)
c.collectRateLimits(ch)
}

// collectRateLimits reports every active core window plus the scarcest one,
// which is what runs out first. Absent until the first GitHub response.
func (c *Collector) collectRateLimits(ch chan<- prometheus.Metric) {
var low github.RateLimit
for reset, rl := range c.rateLimitFn() {
ch <- prometheus.MustNewConstMetric(c.rateWindow, prometheus.GaugeValue, float64(rl.Remaining), reset.Format(time.RFC3339))
if low.Limit == 0 || rl.Remaining < low.Remaining {
low = rl
}
}
if low.Limit > 0 {
ch <- prometheus.MustNewConstMetric(c.rateLimit, prometheus.GaugeValue, float64(low.Remaining))
ch <- prometheus.MustNewConstMetric(c.rateLimitCap, prometheus.GaugeValue, float64(low.Limit))
}
}

func (c *Collector) collectRunners(ch chan<- prometheus.Metric) {
Expand Down Expand Up @@ -122,8 +145,6 @@ func (c *Collector) collectOrgStats(ch chan<- prometheus.Metric) {
return
}
ch <- prometheus.MustNewConstMetric(c.orgUp, prometheus.GaugeValue, 1)
ch <- prometheus.MustNewConstMetric(c.rateLimit, prometheus.GaugeValue, float64(s.RateLimit.Remaining))
ch <- prometheus.MustNewConstMetric(c.rateLimitCap, prometheus.GaugeValue, float64(s.RateLimit.Limit))

visibilityCounts := map[string]int{}
for _, r := range s.Repos {
Expand Down
52 changes: 46 additions & 6 deletions internal/github/client.go
Original file line number Diff line number Diff line change
Expand Up @@ -9,21 +9,29 @@ import (
"net/url"
"regexp"
"strconv"
"sync"
"time"
)

const apiBase = "https://api.github.com"

type Client struct {
baseURL string // apiBase; tests point it at httptest
now func() time.Time
org string
token string
httpClient *http.Client

mu sync.Mutex // both pollers share the client
// core budgets by reset time: GitHub counts per region, so one token has
// more than one window at once, depending on which region serves an endpoint
budgets map[int64]RateLimit
}

func NewClient(org, token string, timeout time.Duration) *Client {
return &Client{
baseURL: apiBase,
now: time.Now,
org: org,
token: token,
httpClient: &http.Client{Timeout: timeout},
Expand All @@ -43,9 +51,34 @@ func (c *Client) request(ctx context.Context, url string) (*http.Response, error
if err != nil {
return nil, fmt.Errorf("requesting %s: %w", url, err)
}
c.recordRateLimit(resp.Header)
return resp, nil
}

// recordRateLimit keeps the core budget from a response's headers. GET
// /rate_limit can't be trusted for this: it reports used=0 for our tokens
// while these headers show the real count.
func (c *Client) recordRateLimit(h http.Header) {
if r := h.Get("X-RateLimit-Resource"); r != "" && r != "core" {
return
}
limit, lErr := strconv.Atoi(h.Get("X-RateLimit-Limit"))
remaining, rErr := strconv.Atoi(h.Get("X-RateLimit-Remaining"))
reset, sErr := strconv.ParseInt(h.Get("X-RateLimit-Reset"), 10, 64)
if lErr != nil || rErr != nil || sErr != nil {
return
}
c.mu.Lock()
defer c.mu.Unlock()
if c.budgets == nil {
c.budgets = map[int64]RateLimit{}
}
// within one window, the lowest count seen is the latest
if b, ok := c.budgets[reset]; !ok || remaining < b.Remaining {
c.budgets[reset] = RateLimit{Limit: limit, Remaining: remaining}
}
}

func (c *Client) get(ctx context.Context, url string, out interface{}) error {
_, err := c.getPage(ctx, url, out)
return err
Expand Down Expand Up @@ -192,11 +225,18 @@ func (c *Client) DependabotAlerts(ctx context.Context, repo string) ([]Dependabo
return out, nil
}

// RateLimit returns the core rate-limit budget this client draws from.
func (c *Client) RateLimit(ctx context.Context) (RateLimit, error) {
var out rateLimitResponse
if err := c.get(ctx, c.baseURL+"/rate_limit", &out); err != nil {
return RateLimit{}, err
// RateLimits returns every core window that hasn't reset yet, by reset time.
func (c *Client) RateLimits() map[time.Time]RateLimit {
c.mu.Lock()
defer c.mu.Unlock()
now := c.now().Unix()
out := map[time.Time]RateLimit{}
for reset, b := range c.budgets {
if reset <= now {
delete(c.budgets, reset)
continue
}
out[time.Unix(reset, 0).UTC()] = b
}
return out.Resources.Core, nil
return out
}
47 changes: 47 additions & 0 deletions internal/github/client_test.go
Original file line number Diff line number Diff line change
Expand Up @@ -3,6 +3,7 @@ package github
import (
"context"
"fmt"
"maps"
"net/http"
"net/http/httptest"
"testing"
Expand Down Expand Up @@ -61,3 +62,49 @@ func TestRunJobsPaginates(t *testing.T) {
t.Fatalf("jobs = %+v", jobs)
}
}

func TestRateLimitFromHeaders(t *testing.T) {
now := time.Unix(1000, 0)
// path -> remaining, reset; two windows at once, like GitHub does
budgets := map[string][2]string{
"/orgs/org/repos": {"3790", "2000"},
"/repos/org/repo/dependabot/alerts": {"4900", "3000"},
"/search": {"1", "2000"},
}
srv := httptest.NewServer(http.HandlerFunc(func(w http.ResponseWriter, r *http.Request) {
b := budgets[r.URL.Path]
w.Header().Set("X-RateLimit-Limit", "5000")
w.Header().Set("X-RateLimit-Remaining", b[0])
w.Header().Set("X-RateLimit-Reset", b[1])
w.Header().Set("X-RateLimit-Resource", "core")
if r.URL.Path == "/search" {
w.Header().Set("X-RateLimit-Resource", "search")
}
_, _ = w.Write([]byte(`[]`))
}))
defer srv.Close()

c := NewClient("org", "token", time.Second)
c.baseURL = srv.URL
c.now = func() time.Time { return now }
if got := c.RateLimits(); len(got) != 0 {
t.Fatalf("before any request: %+v", got)
}
ctx := context.Background()
var out []any
_, _ = c.Repos(ctx)
_, _ = c.DependabotAlerts(ctx, "repo")
_ = c.get(ctx, srv.URL+"/search", &out) // another resource's budget
want := map[time.Time]RateLimit{
time.Unix(2000, 0).UTC(): {Limit: 5000, Remaining: 3790},
time.Unix(3000, 0).UTC(): {Limit: 5000, Remaining: 4900},
}
if got := c.RateLimits(); !maps.Equal(got, want) {
t.Fatalf("RateLimits = %+v, want both core windows and not search", got)
}

now = time.Unix(2500, 0) // the 3790 window has reset
if got := c.RateLimits(); len(got) != 1 || got[time.Unix(3000, 0).UTC()].Remaining != 4900 {
t.Fatalf("after reset: %+v, want only the later window", got)
}
}
12 changes: 3 additions & 9 deletions internal/github/types.go
Original file line number Diff line number Diff line change
Expand Up @@ -85,14 +85,8 @@ type DependabotAlert struct {
} `json:"security_advisory"`
}

// RateLimit is the "core" resource from GET /rate_limit.
// RateLimit is the "core" API budget, from X-RateLimit-* response headers.
type RateLimit struct {
Limit int `json:"limit"`
Remaining int `json:"remaining"`
}

type rateLimitResponse struct {
Resources struct {
Core RateLimit `json:"core"`
} `json:"resources"`
Limit int
Remaining int
}
5 changes: 1 addition & 4 deletions internal/orgstats/feed_test.go
Original file line number Diff line number Diff line change
Expand Up @@ -28,9 +28,6 @@ type fakeClient struct {
func (c *fakeClient) Repos(context.Context) ([]github.Repo, error) {
return []github.Repo{{Name: "repo"}, {Name: "old", Archived: true}}, c.reposErr
}
func (c *fakeClient) RateLimit(context.Context) (github.RateLimit, error) {
return github.RateLimit{Limit: 5000, Remaining: 4000}, nil
}
func (c *fakeClient) OpenPRCount(context.Context, string) (int, error) {
return 0, errors.New("forbidden")
}
Expand Down Expand Up @@ -245,7 +242,7 @@ func TestBuild(t *testing.T) {
t.Fatal(err)
}
// archived repo dropped, disabled workflow dropped, per-repo errors swallowed
if len(s.Repos) != 1 || len(s.Repos[0].Workflows) != 1 || !s.Repos[0].Workflows[0].HasRun || s.RateLimit.Remaining != 4000 {
if len(s.Repos) != 1 || len(s.Repos[0].Workflows) != 1 || !s.Repos[0].Workflows[0].HasRun {
t.Fatalf("summary = %+v", s)
}

Expand Down
Loading
Loading