Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
3 changes: 3 additions & 0 deletions grafana-alertcheck/.changeset/v0.1.7.md
Original file line number Diff line number Diff line change
@@ -0,0 +1,3 @@
- `watch` now stamps `ready_at` when its first-observation pass completes, and `check` refuses a `from` before it; single-step `check` declares that startup interval as a blind spot and classifies from the pass completion. A window can no longer open inside the startup pass and surface as a spurious `heartbeat_gap` mid-run.
- The detached recorder and the live poller continue the first-observation pass's schedule instead of drawing fresh phases, and a new startup-handoff budget simulates the poller's first cycles and refuses, before the window, any schedule whose first polls would exceed a rule's `maxGap`.
- Budget errors now name rules by title as well as UID (`"Title" (uid)`) and state the minimum `--concurrency` that would fit.
13 changes: 11 additions & 2 deletions grafana-alertcheck/docs/advanced.md
Original file line number Diff line number Diff line change
Expand Up @@ -18,13 +18,22 @@ The scheduler staggers each rule's initial next-due time across its cadence, and

## The check budget

The gate records one observation of every rule up front and checks the schedule against those **measured** latencies (payload sizes varied ~230× across existing rules, so a fixed estimate would be meaningless). It errors at start — before waiting — if any of three conditions hold:
The gate records one observation of every rule up front and checks the schedule against those **measured** latencies (payload sizes varied ~230× across existing rules, so a fixed estimate would be meaningless). It errors at start — before waiting — if any of four conditions hold:

- **Utilization** — total request rate exceeds `--concurrency`.
- **Per-rule** — one rule's request can't fit its own cadence.
- **Burst bound** — the slowest request exceeds the fleet's tightest cadence, which can open a mid-run gap.
- **Startup handoff** — draining the first-observation pass's backlog at `--concurrency` would leave some rule unpolled past its own `maxGap`. A rule the pass observed early is seeded overdue, and a tight rule observed late can queue behind every rule due before it. The gate simulates the poller's first cycles from the recorded observation times and measured latencies — each wake takes every rule due at that instant, polls the batch at `--concurrency`, and wakes again when it ends — and refuses if any rule's first poll would land past its `maxGap`. Steady-state utilization cannot see this — a long pass at low concurrency is exactly the case it passes.

The error names the three levers only: raise `--concurrency`, raise `--poll-interval`, or watch fewer alerts. It never prescribes a single interval.
The error names only the levers that can fix it: the minimum `--concurrency` when the schedule is concurrency-bound, and `--poll-interval` or a smaller alert set for single-request shapes concurrency cannot shorten. It never prescribes a single interval.

## The startup pass and `ready_at`

Before detaching, `watch` observes every non-paused rule once, sequentially bounded by `--concurrency`. With many alerts and a low concurrency that pass takes real time (120 rules at ~230 ms each and the default concurrency of 1 is ~27 s). The header's `ready_at` stamps the moment the pass completed, and `check` refuses a `from` before it: a window opening inside the pass names observations that do not exist yet, and the earliest rules have no next poll until the detached recorder starts. This is a startup validation, checked from the immutable header before the wait, so a `from` emitted before `watch` returns fails immediately with a named reason instead of surfacing as a heartbeat gap mid-window. Emit `from` only after `watch` returns.

The detached recorder then continues the schedule the first observations were on (each rule's next poll is one cadence after its last recorded observation) rather than drawing fresh phases, so the handoff adds no extra up-to-one-cadence delay to the rules the pass observed first. Continuing the schedule is necessary but not sufficient: the startup-handoff budget above proves the backlog can actually be drained before any rule's `maxGap`, and refuses the run at startup when it cannot.

Single-step `check` runs the same pass itself. It cannot watch before it started, so a `from` inside the pass is a declared blind interval: the run warns, classifies from the pass completion, and the live poller continues the pass's schedule. It never classifies a window that opens before every rule has been observed.

## Why the gate never queries state history

Expand Down
2 changes: 1 addition & 1 deletion grafana-alertcheck/docs/architecture.md
Original file line number Diff line number Diff line change
Expand Up @@ -55,7 +55,7 @@ This, plus the declared supported range (Grafana >= 13.0.0, < 14.0.0), is how a

`watch` detaches a background recorder so observation survives the step boundary:

1. Parent resolves the alert set (names or labels), writes the header, observes every non-paused rule once, checks the budget.
1. Parent resolves the alert set (names or labels), observes every non-paused rule once, checks the budget, then writes the header — whose `ReadyAt` stamps the pass completion — and the observations. `ReadyAt` is what `check` uses to refuse a `from` that falls inside the pass.
2. Parent re-execs itself as the child (`--daemon-child`) under a new session/process group, stdout/stderr to the daemon log.
3. Child re-reads the header, reopens the log `O_APPEND`, takes the exclusive `flock`, and writes one readiness byte on `--ready-fd`.
4. Parent writes the pidfile **after** the readiness report, then returns.
Expand Down
2 changes: 1 addition & 1 deletion grafana-alertcheck/docs/index.md
Original file line number Diff line number Diff line change
Expand Up @@ -60,7 +60,7 @@ Skip the recorder and observe the window inline, from inside `check` itself:
grafana-alertcheck check --alerts alerts.txt --to "$finished_at"
```

In single-step mode the window starts at `check`'s first observation; if you give no `--from`, the interval before that first observation is declared as a blind spot with a warning (not an error).
In single-step mode the window starts at `check`'s first-observation pass completion; if you give no `--from`, or a `--from` inside the pass, the interval before that point is declared as a blind spot with a warning (not an error).

## Exit codes

Expand Down
4 changes: 2 additions & 2 deletions grafana-alertcheck/docs/reference/cli.md
Original file line number Diff line number Diff line change
Expand Up @@ -43,7 +43,7 @@ grafana-alertcheck watch --out <file> [--pidfile F] [--daemon-log F] \
| `--concurrency` | `1` | Max concurrent requests to Grafana |
| `--until` | run until signalled | Optional hard stop |

`watch` writes the header, observes every non-paused rule once, checks the budget, then detaches a background recorder and returns. Recording is **unfiltered** — there is no `--states` here, so the same log can be re-classified later under different `--states` without re-recording.
`watch` observes every non-paused rule once, checks the budget, writes the header (with `ready_at` stamped once the observation pass completes) and those observations, then detaches a background recorder and returns. Recording is **unfiltered** — there is no `--states` here, so the same log can be re-classified later under different `--states` without re-recording.

## `stop` — reap the recorder

Expand Down Expand Up @@ -92,7 +92,7 @@ grafana-alertcheck check [--in <file>] [--pidfile F] --from RFC3339 --to RFC3339

By default `check` **exits early** on a failure that cannot become a pass: a post-`from` bad onset, or an inability (a heartbeat gap, a sustained `health=error`, a stale evaluation, an in-window pause, an absent rule). This is a latency optimization, not a weaker gate — it never exits `0` early. The one observable difference is that an early exit can report `1` where a full run would have discovered an inability later and reported `2`. `--no-fail-fast` always waits for `to + transitionGrace` and the full coverage proof; the `Result` then carries no `terminated_early` marker. With early exit the JSON result includes `terminated_early` naming the rule, kind, reason and time.

`--from` and `--to` are RFC3339 with an explicit offset and must come from your work — `from` from the deploy step, `to` from the step that finishes. In recorder mode an absent `--from` is a hard error; in single-step mode it falls back (with a warning) to the start of the step.
`--from` and `--to` are RFC3339 with an explicit offset and must come from your work — `from` from the deploy step, `to` from the step that finishes. In recorder mode an absent `--from` is a hard error, and a `from` before the recording's first-observation pass is refused; in single-step mode an absent `from`, or one inside `check`'s own first-observation pass, is a declared blind interval — the window is classified from the pass completion, with a warning.

## Naming alerts

Expand Down
2 changes: 2 additions & 0 deletions grafana-alertcheck/docs/reference/log-format.md
Original file line number Diff line number Diff line change
Expand Up @@ -31,6 +31,7 @@ The header must be line 1, appear once, and carry `schema_version` `1` (any othe
"url": "https://grafana.example.com",
"grafana_version": "13.1.0",
"started_at": "2026-09-07T10:00:00Z",
"ready_at": "2026-09-07T10:00:27Z",
"rules": [
{
"uid": "rule0000001",
Expand All @@ -49,6 +50,7 @@ The header must be line 1, appear once, and carry `schema_version` `1` (any othe
```

- `url` and `rules` are the log's identity — `check` validates them against the current environment and a fresh ruler read.
- `started_at` is when the recording opened; `ready_at` is when the first-observation pass completed and every watched, non-paused rule had been observed once. The pass is sequential, so `check` refuses a `from` before `ready_at` (a window opening inside the pass would rest on observations that do not exist). `ready_at` is absent on logs written before the field existed; `check` then falls back to `started_at`.
- `is_paused` records the pause state at record start (the moment `paused` means).
- `poll_every_seconds` is the cadence the recording **actually used** (after any `--poll-interval` override). `check` derives `maxGap` from it, never from `interval_seconds`.
- `for_seconds`, `interval_seconds`, `no_data_state`, `exec_err_state` are forensic only — `check` re-resolves definitions and never reads them back.
Expand Down
62 changes: 49 additions & 13 deletions grafana-alertcheck/internal/gate/check.go
Original file line number Diff line number Diff line change
Expand Up @@ -54,9 +54,9 @@ type Config struct {
NoFailFast bool

// From is the moment the deploy finished and To is the end of the work.
// They are different moments and both come from the work. In recorder mode
// an absent From is a hard error; in single-step mode it falls back to the
// start of this step, with a blind-interval warning.
// In recorder mode an absent From is a hard error; in single-step mode it
// falls back to the start of this step, and a From before the
// first-observation pass completes is a declared blind interval.
From, To time.Time

// Log is the path of a recording made by watch; "" selects single-step
Expand Down Expand Up @@ -165,7 +165,7 @@ func (cfg Config) validate() error {
return errors.New("check: no `from` in recorder mode: the deploy step must emit a completion timestamp")
case from.IsZero():
// Single-step only. The caller sees the resulting blind interval named
// exactly, once the first observation has fixed its end.
// exactly, once the first-observation pass has fixed its end.
from = now
}

Expand Down Expand Up @@ -245,6 +245,14 @@ func check(ctx context.Context, cfg Config, src Source) (Result, error) {
return Result{}, fmt.Errorf("check: `from` %s is before recording started at %s",
from.Format(time.RFC3339), earlyHdr.StartedAt.Format(time.RFC3339))
}
// The pass is sequential, so a `from` inside it names a window
// whose earliest rules were never watched. Knowable from the
// immutable header, so it fails before the wait.
if from.Truncate(time.Second).Before(earlyHdr.ReadyAt.Truncate(time.Second)) {
return Result{}, fmt.Errorf(
"check: `from` %s is inside the recorder's initial observation pass, which completed at %s; emit `from` after `watch` returns (watch observes every watched rule before returning)",
from.Format(time.RFC3339), earlyHdr.ReadyAt.Format(time.RFC3339))
Comment thread
Tofel marked this conversation as resolved.
}
}
} else {
resolved, notes, err = resolveAlertSet(allDefs, cfg.namedAlerts(), cfg.IncludeLabels, cfg.ExcludeLabels, cfg.Folder)
Expand Down Expand Up @@ -308,35 +316,56 @@ func check(ctx context.Context, cfg Config, src Source) (Result, error) {
// proves that interval, not this timestamp.
startedAt := cfg.Clock.Now()
active := activeRules(resolved)
activeRT := activeTimingsOf(active, rt)
var measured map[string]time.Duration
initial, measured, err = firstObservations(ctx, src, active, reducer, cfg.Concurrency, cfg.Notes)
if err != nil {
return Result{}, err
}
if err := CheckBudget(activeTimingsOf(active, rt), measured, cfg.Concurrency); err != nil {
// The pass is over; the live poller will start from this instant, so
// both budget checks judge the schedule it will actually run.
readyAt := cfg.Clock.Now()
if err := CheckBudget(activeRT, measured, cfg.Concurrency); err != nil {
return Result{}, err
}
// The clamp below is exact, so the window cannot open before readyAt.
if err := CheckStartupHandoff(activeRT, measured, initial, readyAt, readyAt, cfg.Concurrency); err != nil {
return Result{}, err
}

// Single-step synthesis — how the pure layer stays unconditional. The
// shell builds the Header and later stamps the sentinel itself, so the
// sentinel and from-bounds coverage checks run exactly as they do over
// a recording and no mode flag ever reaches proveCoverage or decide.
//
// ReadyAt is the pass completion: single-step cannot watch before it
// started, so a `from` inside the pass is declared blind and the window
// is classified from ReadyAt.
header = Header{
SchemaVersion: LogSchemaVersion,
URL: cfg.URL,
GrafanaVersion: version,
StartedAt: startedAt,
ReadyAt: readyAt,
Rules: loggedRules(resolved, rt),
}
if from.Before(startedAt) {
if from.Before(readyAt) {
// The declared blind interval: in single-step mode this is a
// warning and a pass, and ONLY here. Recorder mode keeps the
// from-bounds coverage check strict, because there the recorder
// was supposed to be watching and the gap means it was not.
fmt.Fprintf(cfg.Notes, "warning: cannot see [%s, %s) — %s before the first observation; the window is classified from %s\n",
from.Format(time.RFC3339), startedAt.Format(time.RFC3339),
startedAt.Sub(from).Round(time.Second), startedAt.Format(time.RFC3339))
from = startedAt
//
// A clamp past `to` leaves no requested window to classify (and an
// inverted window can prove nothing), so fail closed instead.
if !readyAt.Before(cfg.To) {
return Result{}, fmt.Errorf(
"check: the first-observation pass completed at %s, at or after `to` %s: no window remains to classify",
readyAt.Format(time.RFC3339), cfg.To.Format(time.RFC3339))
}
fmt.Fprintf(cfg.Notes, "warning: cannot see [%s, %s) — %s before the first observation pass completed; the window is classified from %s\n",
from.Format(time.RFC3339), readyAt.Format(time.RFC3339),
readyAt.Sub(from).Round(time.Second), readyAt.Format(time.RFC3339))
from = readyAt
}
}

Expand Down Expand Up @@ -369,7 +398,7 @@ func check(ctx context.Context, cfg Config, src Source) (Result, error) {
guard terminalCheck
)
if !logHasHdr {
poller = newLivePoller(src, reducer, activeRules(resolved), rt, cfg.Concurrency, cfg.Clock.Now())
poller = newLivePoller(src, reducer, activeRules(resolved), rt, cfg.Concurrency, cfg.Clock.Now(), initial)
}
failFast := !cfg.NoFailFast
if failFast {
Expand Down Expand Up @@ -539,19 +568,26 @@ type livePoller struct {
concurrency int
}

// newLivePoller builds single-step mode's collection engine. seed, when
// present, is the measurement pass's observations: the poller continues their
// schedule instead of drawing fresh phases. nil falls back to the fresh stagger.
func newLivePoller(src Source, reducer *Reducer, active []Definition, rt map[string]RuleTimings,
concurrency int, now time.Time) *livePoller {
concurrency int, now time.Time, seed []Poll) *livePoller {

titles := make(map[string]string, len(active))
cadence := make(map[string]time.Duration, len(active))
for _, d := range active {
titles[d.UID] = d.Title
cadence[d.UID] = rt[d.UID].pollEvery
}
sched := NewScheduler(cadence, now)
if len(seed) > 0 {
sched = NewSchedulerFromPolls(cadence, seed, now)
}
return &livePoller{
src: src,
reducer: reducer,
sched: NewScheduler(cadence, now),
sched: sched,
titles: titles,
concurrency: concurrency,
}
Expand Down
Loading
Loading