Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension


Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
7 changes: 7 additions & 0 deletions .changeset/fluffy-hounds-fail.md
Original file line number Diff line number Diff line change
@@ -0,0 +1,7 @@
---
"ci-grafana-alert-test": major
"go-conditional-tests": minor
---

feat: first release of Grafana Alert check, fix: update distribution of
go-conditional-tests
144 changes: 144 additions & 0 deletions actions/ci-grafana-alert-test/README.md
Original file line number Diff line number Diff line change
@@ -0,0 +1,144 @@
# ci-grafana-alert-test

A CD quality gate for Grafana alerts. It bookends a release with two calls to
this action — `record` before the deploy or tests, `check` after the work is
done — and answers one question: _was any watched alert in a bad state at any
point during the release window?_

It wraps the
[`grafana-alertcheck`](https://github.com/smartcontractkit/chainlink-testing-framework/tree/main/grafana-alertcheck)
CLI from `chainlink-testing-framework`. The action downloads a prebuilt release
binary rather than compiling from source — the version is pinned in one place in
the action and bumped by editing that single line when a new release ships.

## Usage

```yaml
- uses: smartcontractkit/.github/actions/ci-grafana-alert-test@ci-grafana-alert-test/v1
with:
mode: record
grafana-url: ${{ vars.GRAFANA_URL }}
grafana-token: ${{ secrets.GRAFANA_TOKEN }}
alerts: |
My Service Latency
My Service Error Rate

- id: deploy
run: ./deploy.sh # emits deployed_at=<RFC3339> when the rollout is stable

- id: work
run: ./verify.sh # emits finished_at=<RFC3339> when done (tests, traffic, whatever)

- uses: smartcontractkit/.github/actions/ci-grafana-alert-test@ci-grafana-alert-test/v1
with:
mode: check
grafana-url: ${{ vars.GRAFANA_URL }}
grafana-token: ${{ secrets.GRAFANA_TOKEN }}
from: ${{ steps.deploy.outputs.deployed_at }}
to: ${{ steps.work.outputs.finished_at }} # ...or `duration: 10m` when there is no done event — never both
```

`record` and `check` must run in the **same job, on the same runner** — nothing
is passed between jobs or between run attempts. `from` must come from the deploy
step's own completion output, never from a wrapper step around it; a single step
must not serve its own completion as `from`, or the window between landing and
finishing is never observed at all.

## What it checks, and what it does not

- The gate checks the **state and health** of an alert. It does **not** check
whether a notification was ever delivered. **A silenced alert that fires still
fails the gate.**
- The gate needs Grafana 13.x.
- **A fix that stops emitting a metric does not look like a recovery.** An
instance that vanishes while bad stays a failure — a missing series is a
discontinuity, not evidence of health.
- **A paused rule fails the gate by default** (`allow-paused: 'false'`). If
someone else paused an alert you're watching, your release fails on it — the
alternative is silently watching fewer alerts than you asked for.

## Retries

**A retry is a new deploy, not a replay.** There is no cheap re-check: each
`check` run classifies its own freshly recorded window, and on failure the
evidence log is **uploaded, never downloaded** — so a rerun cannot replay old
evidence to pass. A second attempt legitimately relabeling the same commit
`newly_bad` on attempt 1 and `persistently_bad` on attempt 2 is correct, not a
bug — the exit code is the same, the label is more accurate.

## Deploy and test in separate jobs

`record` and `check` would ideally live in one job on one runner, because
`check` finds the recorded log by convention on the local filesystem. If your
deploy and your verification/test work run in **different jobs**, you have to
choose where the gate lives, and that choice trades off against coverage:

- **Record/check in the deploy job only** — the window observes the deploy and
whatever falls inside its `duration`. Whether it also covers your test job's
activity depends entirely on how long the gate runs versus when (and how long)
the test job runs; there is no automated way to know for sure, so any alert
that fires under test traffic could fall just outside the window.
- **Record/check in the test job only** — there is a **blind window** between
the deployment becoming ready and the test job's recorder starting. Alerts
that fire in that span are never seen.

There is no way to shrink that blind window by leaning on one job alone — the
recorder cannot see back in time, it can only watch from the moment it starts.
The way to get **zero gap** is to run the gate in **both** jobs: each job
records its own window, and as long as the deploy job's window end overlaps (or
touches) the test job's window start, the two observations together cover the
whole span with no uncovered interval. The price is two overlapping windows to
classify and, when the same alert fires across the boundary, two violations to
reconcile — but that is strictly better than a silent gap.

## Timing

`record` blocks for a short time — until it has observed every non-paused
watched alert at least once — before it detaches and returns. This is
intentional: it closes the blind interval between the deploy landing and the
gate actually watching it.

A gate with a 10-minute window (`to − from`) holds the runner for
**approximately 10 minutes plus grace and drain time**, printed at the start of
the `check` step. There is no early exit — the gate observes the full window
even after it already knows the answer, because early-exiting is exactly what
would reopen the coverage gap this whole tool exists to close. Make sure the
surrounding job's timeout accounts for this.

## Failure behaviour

- `fail-on-violation: 'false'` suppresses a **violation** (exit 1) only. A
**could-not-check** result (exit 2 — auth failure, coverage gap, an
unobservable rule, a schedule that doesn't fit, and so on) always fails the
job: an inability to answer is never a pass.
- On any non-zero `check` exit, the JSONL evidence log is uploaded as
`grafana-alert-gate-log-${{ github.run_id }}-${{ github.run_attempt }}` for
diagnosis after the runner is gone.
- `to` and `duration` are mutually exclusive on `mode: check` — give exactly
one. There is deliberately no default for either; a 10-minute gate is a choice
you make explicitly, not one this action makes for you.

## Inputs

See [action.yml](action.yml) for the full, authoritative list with defaults. The
ones worth calling out:

| Input | Notes |
| -------------------------- | -------------------------------------------------------------------------------------------------------------------------------------------- |
| `mode` | `record` or `check` |
| `alerts` | One alert name per line. `record` only — `check` reads the set from the recorded log, and giving both is an error |
| `from` / `to` / `duration` | `check` only. `from` is when the deploy landed; `to` is when the work ended; `duration` replaces `to` when there is no distinct "done" event |
| `fail-on-violation` | Default `true`. Stops exit 1 only, never exit 2 |

## Outputs

`record` sets `log-path` and `pidfile` for transparency; `check` finds them by
convention, so you never need to wire them through yourself. `check` sets
`passed`, `violation-count`, `violations` (JSON), and `outcomes` (JSON, one
`{alert, outcome}` entry per resolved rule).

## Runner requirements

Linux (`ubuntu-*`) and macOS runners, `amd64` or `arm64`. The release provides
no other platform binaries. The action runs on the `node24` runtime, and the
window arithmetic is done in TypeScript, so it works on both OSes.
104 changes: 104 additions & 0 deletions actions/ci-grafana-alert-test/action.yml
Original file line number Diff line number Diff line change
@@ -0,0 +1,104 @@
name: ci-grafana-alert-test
description:
"CD quality gate: bookend a release with a live-recorded observation of
Grafana alerts (`record`) and fail the job if any watched alert was bad during
the window (`check`). Wraps the grafana-alertcheck CLI from
chainlink-testing-framework, downloaded as a prebuilt binary from the released
`grafana-alertcheck/v0.1.0` version."

inputs:
mode:
description:
"'record' (start the recorder) or 'check' (classify the window and exit)"
required: true
grafana-url:
description: "Base URL of the Grafana instance"
required: true
grafana-token:
description: "Grafana API token (secret)"
required: true
alerts:
description:
"One alert rule name per line. Required with mode: record. Refused with
mode: check — the recorded log already carries its own alert set."
required: false
states:
description:
"Comma-separated bad states to classify against. mode: check only.
Default: firing"
required: false
from:
description:
"RFC3339 timestamp, explicit offset, of when the deploy landed. Required
with mode: check. Must come from the deploy step's own completion output,
never a wrapper step."
required: false
to:
description:
"RFC3339 timestamp of the end of the window. Mutually exclusive with
duration; give exactly one. mode: check only."
required: false
duration:
description:
"Duration expressed as one or more integer h/m/s segments (e.g. 10m,
1h30m) added to `from` to compute `to` when there is no distinct 'done'
event. Mutually exclusive with `to`; give exactly one. mode: check only."
required: false
preexisting:
description:
"fail-unless-recovered (default) | fail | ignore — how to judge an
instance already bad at `from`. mode: check only."
required: false
min-observed:
description:
"Minimum number of rules that must be observed. Default: every resolved
rule. mode: check only."
required: false
allow-paused:
description:
"Do not count a rule paused before the window against min-observed. mode:
check only."
required: false
default: "false"
nodata-is-unobservable:
description:
"Treat a sustained health=nodata as unobservable rather than a note. mode:
check only."
required: false
default: "false"
poll-interval:
description:
"Override every watched rule's poll cadence (Go duration, e.g. 30s).
Default: automatic, per rule. mode: record only."
required: false
concurrency:
description: "Maximum concurrent requests to Grafana"
required: false
folder:
description: "Default folder to scope an unqualified alert name to"
required: false
fail-on-violation:
description:
"Set to 'false' to stop a violation (exit 1) from failing the job. A
could-not-check result (exit 2) always fails the job regardless."
required: false
default: "true"

outputs:
log-path:
description: "Path of the JSONL evidence log. mode: record only."
pidfile:
description: "Path of the recorder's pidfile. mode: record only."
passed:
description: "'true' if the check exited 0. mode: check only."
violation-count:
description: "Number of violations. mode: check only."
violations:
description: "JSON array of violations. mode: check only."
outcomes:
description:
"JSON array of {alert, outcome} per resolved rule. mode: check only."

runs:
using: node24
main: "dist/index.js"
Loading
Loading