You signed in with another tab or window. Reload to refresh your session.You signed out in another tab or window. Reload to refresh your session.You switched accounts on another tab or window. Reload to refresh your session.Dismiss alert
{{ message }}
Repository navigation
Commit fd38475
Browse filesBrowse the repository at this point in the historyBrowse files
-`watch` now stamps `ready_at` when its first-observation pass completes, and `check` refuses a `from` before it; single-step `check` declares that startup interval as a blind spot and classifies from the pass completion. A window can no longer open inside the startup pass and surface as a spurious `heartbeat_gap` mid-run.
2
+
- The detached recorder and the live poller continue the first-observation pass's schedule instead of drawing fresh phases, and a new startup-handoff budget simulates the poller's first cycles and refuses, before the window, any schedule whose first polls would exceed a rule's `maxGap`.
3
+
- Budget errors now name rules by title as well as UID (`"Title" (uid)`) and state the minimum `--concurrency` that would fit.
Copy file name to clipboardExpand all lines: grafana-alertcheck/docs/advanced.md
+11-2Lines changed: 11 additions & 2 deletions
Display the source diff
Display the rich diff
Original file line number
Diff line number
Diff line change
@@ -18,13 +18,22 @@ The scheduler staggers each rule's initial next-due time across its cadence, and
18
18
19
19
## The check budget
20
20
21
-
The gate records one observation of every rule up front and checks the schedule against those **measured** latencies (payload sizes varied ~230× across existing rules, so a fixed estimate would be meaningless). It errors at start — before waiting — if any of three conditions hold:
21
+
The gate records one observation of every rule up front and checks the schedule against those **measured** latencies (payload sizes varied ~230× across existing rules, so a fixed estimate would be meaningless). It errors at start — before waiting — if any of four conditions hold:
22
22
23
23
-**Utilization** — total request rate exceeds `--concurrency`.
24
24
-**Per-rule** — one rule's request can't fit its own cadence.
25
25
-**Burst bound** — the slowest request exceeds the fleet's tightest cadence, which can open a mid-run gap.
26
+
-**Startup handoff** — draining the first-observation pass's backlog at `--concurrency` would leave some rule unpolled past its own `maxGap`. A rule the pass observed early is seeded overdue, and a tight rule observed late can queue behind every rule due before it. The gate simulates the poller's first cycles from the recorded observation times and measured latencies — each wake takes every rule due at that instant, polls the batch at `--concurrency`, and wakes again when it ends — and refuses if any rule's first poll would land past its `maxGap`. Steady-state utilization cannot see this — a long pass at low concurrency is exactly the case it passes.
26
27
27
-
The error names the three levers only: raise `--concurrency`, raise `--poll-interval`, or watch fewer alerts. It never prescribes a single interval.
28
+
The error names only the levers that can fix it: the minimum `--concurrency` when the schedule is concurrency-bound, and `--poll-interval` or a smaller alert set for single-request shapes concurrency cannot shorten. It never prescribes a single interval.
29
+
30
+
## The startup pass and `ready_at`
31
+
32
+
Before detaching, `watch` observes every non-paused rule once, sequentially bounded by `--concurrency`. With many alerts and a low concurrency that pass takes real time (120 rules at ~230 ms each and the default concurrency of 1 is ~27 s). The header's `ready_at` stamps the moment the pass completed, and `check` refuses a `from` before it: a window opening inside the pass names observations that do not exist yet, and the earliest rules have no next poll until the detached recorder starts. This is a startup validation, checked from the immutable header before the wait, so a `from` emitted before `watch` returns fails immediately with a named reason instead of surfacing as a heartbeat gap mid-window. Emit `from` only after `watch` returns.
33
+
34
+
The detached recorder then continues the schedule the first observations were on (each rule's next poll is one cadence after its last recorded observation) rather than drawing fresh phases, so the handoff adds no extra up-to-one-cadence delay to the rules the pass observed first. Continuing the schedule is necessary but not sufficient: the startup-handoff budget above proves the backlog can actually be drained before any rule's `maxGap`, and refuses the run at startup when it cannot.
35
+
36
+
Single-step `check` runs the same pass itself. It cannot watch before it started, so a `from` inside the pass is a declared blind interval: the run warns, classifies from the pass completion, and the live poller continues the pass's schedule. It never classifies a window that opens before every rule has been observed.
Copy file name to clipboardExpand all lines: grafana-alertcheck/docs/architecture.md
+1-1Lines changed: 1 addition & 1 deletion
Display the source diff
Display the rich diff
Original file line number
Diff line number
Diff line change
@@ -55,7 +55,7 @@ This, plus the declared supported range (Grafana >= 13.0.0, < 14.0.0), is how a
55
55
56
56
`watch` detaches a background recorder so observation survives the step boundary:
57
57
58
-
1. Parent resolves the alert set (names or labels), writes the header, observes every non-paused rule once, checks the budget.
58
+
1. Parent resolves the alert set (names or labels), observes every non-paused rule once, checks the budget, then writes the header — whose `ReadyAt` stamps the pass completion — and the observations. `ReadyAt` is what `check` uses to refuse a `from` that falls inside the pass.
59
59
2. Parent re-execs itself as the child (`--daemon-child`) under a new session/process group, stdout/stderr to the daemon log.
60
60
3. Child re-reads the header, reopens the log `O_APPEND`, takes the exclusive `flock`, and writes one readiness byte on `--ready-fd`.
61
61
4. Parent writes the pidfile **after** the readiness report, then returns.
In single-step mode the window starts at `check`'s firstobservation; if you give no `--from`, the interval before that first observation is declared as a blind spot with a warning (not an error).
63
+
In single-step mode the window starts at `check`'s first-observation pass completion; if you give no `--from`, or a `--from` inside the pass, the interval before that point is declared as a blind spot with a warning (not an error).
|`--concurrency`|`1`| Max concurrent requests to Grafana |
44
44
|`--until`| run until signalled | Optional hard stop |
45
45
46
-
`watch`writes the header, observes every non-paused rule once, checks the budget, then detaches a background recorder and returns. Recording is **unfiltered** — there is no `--states` here, so the same log can be re-classified later under different `--states` without re-recording.
46
+
`watch` observes every non-paused rule once, checks the budget, writes the header (with `ready_at` stamped once the observation pass completes) and those observations, then detaches a background recorder and returns. Recording is **unfiltered** — there is no `--states` here, so the same log can be re-classified later under different `--states` without re-recording.
By default `check` **exits early** on a failure that cannot become a pass: a post-`from` bad onset, or an inability (a heartbeat gap, a sustained `health=error`, a stale evaluation, an in-window pause, an absent rule). This is a latency optimization, not a weaker gate — it never exits `0` early. The one observable difference is that an early exit can report `1` where a full run would have discovered an inability later and reported `2`. `--no-fail-fast` always waits for `to + transitionGrace` and the full coverage proof; the `Result` then carries no `terminated_early` marker. With early exit the JSON result includes `terminated_early` naming the rule, kind, reason and time.
94
94
95
-
`--from` and `--to` are RFC3339 with an explicit offset and must come from your work — `from` from the deploy step, `to` from the step that finishes. In recorder mode an absent `--from` is a hard error; in single-step mode it falls back (with a warning) to the start of the step.
95
+
`--from` and `--to` are RFC3339 with an explicit offset and must come from your work — `from` from the deploy step, `to` from the step that finishes. In recorder mode an absent `--from` is a hard error, and a `from` before the recording's first-observation pass is refused;in single-step mode an absent `from`, or one inside `check`'s own first-observation pass, is a declared blind interval — the window is classified from the pass completion, with a warning.
Copy file name to clipboardExpand all lines: grafana-alertcheck/docs/reference/log-format.md
+2Lines changed: 2 additions & 0 deletions
Display the source diff
Display the rich diff
Original file line number
Diff line number
Diff line change
@@ -31,6 +31,7 @@ The header must be line 1, appear once, and carry `schema_version` `1` (any othe
31
31
"url": "https://grafana.example.com",
32
32
"grafana_version": "13.1.0",
33
33
"started_at": "2026-09-07T10:00:00Z",
34
+
"ready_at": "2026-09-07T10:00:27Z",
34
35
"rules": [
35
36
{
36
37
"uid": "rule0000001",
@@ -49,6 +50,7 @@ The header must be line 1, appear once, and carry `schema_version` `1` (any othe
49
50
```
50
51
51
52
-`url` and `rules` are the log's identity — `check` validates them against the current environment and a fresh ruler read.
53
+
-`started_at` is when the recording opened; `ready_at` is when the first-observation pass completed and every watched, non-paused rule had been observed once. The pass is sequential, so `check` refuses a `from` before `ready_at` (a window opening inside the pass would rest on observations that do not exist). `ready_at` is absent on logs written before the field existed; `check` then falls back to `started_at`.
52
54
-`is_paused` records the pause state at record start (the moment `paused` means).
53
55
-`poll_every_seconds` is the cadence the recording **actually used** (after any `--poll-interval` override). `check` derives `maxGap` from it, never from `interval_seconds`.
54
56
-`for_seconds`, `interval_seconds`, `no_data_state`, `exec_err_state` are forensic only — `check` re-resolves definitions and never reads them back.
"check: `from` %s is inside the recorder's initial observation pass, which completed at %s; emit `from` after `watch` returns (watch observes every watched rule before returning)",
0 commit comments