diff --git a/.github/workflows/ci.yaml b/.github/workflows/ci.yaml index 8eb8db9c..37b73231 100644 --- a/.github/workflows/ci.yaml +++ b/.github/workflows/ci.yaml @@ -1,236 +1,30 @@ name: ci -# GATE COVERAGE. scripts/qa/ holds seventeen scripts outside the -# test_*.sh suite; two of those (ac-commands.sh, helpers.sh) are not gates — -# see below. This block is the one place that names, for each of the -# remaining fifteen gates AS COUNTED HERE AND NOW (this figure moves as gate -# scripts are added or retired — it is not the sixteen -# docs/spec/review-strategy.md's own count names, since that count predates -# copy-verify.sh, render-verify.sh, and later additions and retirements), -# whether CI runs it and why not when it doesn't — docs/spec/review-strategy.md -# §10 (G2) is the gap this closes. -# scripts/qa/gate-coverage-check.sh checks this block against the directory -# mechanically, so this listing cannot drift silently. -# -# RUNS HERE, by name, in `repo-gates`. On `pull_request` AND on every -# `push` (git-diff scripts in base-ref mode; the base is `base.sha` for a -# PR and `github.event.before` for a push, falling back to the root commit -# when that is the all-zero SHA — see the `repo-gates` job comment below -# for the one gap that remains, which is that this job reports rather than -# blocks): -# secret-scan.sh, self-hygiene.sh -# -# RUNS ELSEWHERE IN THIS FILE, already: -# genericity.sh — invoked by `./scripts/qa.sh` (`qa` job) on -# every `pull_request` and every `push` to -# `main`/a tag (this file's own `on:` block). -# copy-verify.sh — same: invoked by `./scripts/qa.sh`. Both are -# render-verify.sh whole-repo, diff-independent checks (no -# DOCKET_* env needed — verified empty-env), -# the same shape as genericity.sh, not tied -# to a specific workflow step or run/step -# identity. -# THIS BLOCK IS AUTHORITATIVE over the -# earlier acceptance criterion that listed -# these two as excluded. That AC was written from the -# premise that both need run context; the -# premise was false — neither reads a -# DOCKET_* variable, and both exit 0 under -# `env -i` — so they are RUN, which is -# strictly more coverage than the AC asked -# for. The other five it named are still -# excluded, below. -# render-verify.sh runs here with ONE HALF -# VACUOUS: its coverage check reads -# working-tree/staged/untracked state, which -# a `pull_request` checkout does not have, -# so that half reports and cannot fail. Its -# ANSI-escape half is whole-repo and does -# bite here. See the script's own comment at -# the `changed=` assignment. That is the -# same criterion the EXCLUDED block below -# uses; render-verify still earns its place -# here because its second half is real, -# whereas those four are vacuous entire. -# gate-baseref-regression.sh — also invoked by `./scripts/qa.sh`; pins the -# gate-coverage-check.sh base-ref mode of secret-scan.sh/ -# self-hygiene.sh and this block itself. -# build.sh — its check (`go1.26.6 build ./...`) is what -# the `test` job's `go build ./...` step -# already verifies, given the same Go -# toolchain: `actions/setup-go`'s -# `go-version-file: go.mod` resolves the -# `toolchain go1.26.6` line there, which the -# `test` job's own `go version` step below -# makes self-checking rather than assumed. -# tests.sh — same equivalence, for `go test ./...`. -# -# EXCLUDED — read only working-tree/staged/untracked state, with no -# base-ref mode (the gap secret-scan.sh/self-hygiene.sh had before this fix -# round) to read a committed PR diff instead. CI has no staged or unstaged -# state, so wiring these as-is would pass every PR vacuously rather than -# fail closed. Adding base-ref mode is the same three-line pattern applied -# to secret-scan.sh/self-hygiene.sh; deferred to a follow-up rather than -# done in this round — the missing base-ref mode is the actual blocker, -# not run/step identity (each already runs with an empty DOCKET_* env): -# doc-validate.sh, citation-check.sh, tdd-preflight.sh, -# reserved-name-check.sh -# -# EXCLUDED — for other stated reasons: -# sdet-abuse.sh — needs the `go1.26.6` binary by that exact name on -# PATH (the script invokes it directly, not `go test`); -# `actions/setup-go` does not add a version-suffixed -# alias. It also cannot verify its cases match a run's -# threat model without step identity — see the script's -# own header. Either gap is enough; deferred, not new -# CI infrastructure this round. -# vuln-scan.sh — needs `govulncheck` installed; deferred as a follow-up -# so this change stays wiring, not new CI infrastructure. -# ac-commands.sh — a reporter for the `verify` step's context bundle; -# ALWAYS EXITS 0 by design (see the script's own header), -# so it has no pass/fail verdict for CI to act on. Not a -# gate. -# helpers.sh — shared functions `qa.sh`'s own test files source; not -# an executable gate. - on: + # pull_request covers every commit on a branch with an open PR. push is + # scoped to main/tags only, so that branch's commits aren't built twice — + # once for pull_request, once for a push the "**" pattern used to also + # match. pull_request: push: - # ALL branches, not just `main`. This repository's work lands as - # direct commits to feature branches, so a `main`-only push trigger meant - # the gates in `repo-gates` never ran on the path the code actually takes. branches: - - "**" + - main tags: - "*" jobs: - test: - runs-on: ubuntu-latest - steps: - - uses: actions/checkout@v7 - - - uses: actions/setup-go@v6 - with: - go-version-file: go.mod - cache: true - - # Asserts, not just prints: the GATE COVERAGE block claims this job - # covers build.sh/tests.sh's version-pinned invocations because setup-go - # resolves go.mod's toolchain line — a bare `go version` step cannot - # fail, so it was evidence of nothing. `go build ./...` below already - # fails if go is missing entirely; this is the one property that step - # alone does not check. - # - # The expected version is READ FROM go.mod, never hard-coded here. - # go.mod carries two version lines — `go 1.26.0` and - # `toolchain go1.26.6` — and a literal here pinned only the second. A - # toolchain bump then had to be made in two files or CI failed every - # PR naming a version go.mod no longer contains. Deriving it keeps the - # property (setup-go really did resolve the toolchain line) while - # making go.mod the single place the version lives. - # - # ACCEPTING EITHER LINE is deliberate. Which one setup-go resolves is - # its behaviour, not ours; the property worth asserting is that the Go - # on PATH is one of the versions go.mod actually names, so a bump to - # either line stays green and an unrelated toolchain still fails. - - run: | - go_directive=$(sed -n 's/^go \([0-9.]*\)$/go\1/p' go.mod) - toolchain=$(sed -n 's/^toolchain \(go[0-9.]*\)$/\1/p' go.mod) - if [ -z "$go_directive" ] && [ -z "$toolchain" ]; then - echo "could not read a version from go.mod; refusing to pass" >&2 - exit 1 - fi - actual=$(go version) - for want in $toolchain $go_directive; do - case "$actual" in - *"$want "*|*"$want") echo "toolchain ok: $want"; exit 0 ;; - esac - done - echo "toolchain mismatch: go.mod names ${toolchain:-none}/${go_directive:-none}, got:" >&2 - echo "$actual" >&2 - exit 1 - - - run: go build ./... - - # internal/workflow/routing_sweep_test.go sweeps the workflow corpus - # config.Config.InstanceConfigDirs resolves, which on a bare CI - # checkout is empty — there is no shared store at $HOME/.docket and - # this repo ships no `.docket/config` of its own. Pointing the sweep - # at the hermetic fixture under internal/workflow/testdata gives it a - # corpus without unioning into anyone's real shared store (see that - # directory's README.md for why it must stay override-only). - - run: go test ./... - env: - DOCKET_SWEEP_WORKFLOWS_DIR: ${{ github.workspace }}/internal/workflow/testdata/ci-routing-corpus - - qa: - runs-on: ubuntu-latest - steps: - - uses: actions/checkout@v7 - - - uses: actions/setup-go@v6 - with: - go-version-file: go.mod - cache: true - - # jq and perl are preinstalled on ubuntu-latest; qa.sh hard-fails without them. - - run: ./scripts/qa.sh - - # repo-gates. The scripts here read a COMMITTED diff, not - # working-tree state, so they need the PR's base commit fetched and passed - # explicitly — see the base-ref mode each script's header now documents. - # - # KNOWN LIMIT: `BASE_SHA` below is `github.event.pull_request.base.sha`, the - # base branch tip at event time — not necessarily an ancestor of the - # checked-out `pull_request` ref's own history. If the base branch advances - # after the PR's last push, the scanned range can widen to base-branch - # commits the PR did not author (INFERRED — no local seam to exercise a - # real Actions `pull_request` event to confirm this against a live run). - # - # RUNS ON PUSH AS WELL AS PULL_REQUEST. - # - # `pull_request`-only was a gate that never fired on this repository's - # actual regime. Work here lands as direct commits to a feature branch - # (`git log --merges` is empty), and `ci-nightly.yaml` tags `main`'s HEAD - # daily and publishes through the push-on-tags path. Both paths excluded - # `repo-gates`, so the credential gate ran on nothing that actually - # happens here — literally satisfied, practically absent. - # - # The base ref differs per event, which is the whole reason this was - # deferred once: - # pull_request — `base.sha`, the base branch tip at event time. - # push — `github.event.before`, the branch's previous tip. - # - # `before` IS THE ALL-ZERO SHA ON A BRANCH'S FIRST PUSH (and on a tag - # push), which is not a readable ref. That case is handled explicitly - # below rather than left to fail confusingly: it falls back to the - # repository's first commit, so the scan covers the entire branch. For a - # brand-new branch that is the correct interpretation — every commit on it - # is new — and it fails toward scanning MORE, never less. - # - # STILL UNGATED, stated plainly: nothing here can make - # this a REQUIRED check. Branch protection is not version-controlled in - # this repo, so a push whose workflow run fails still lands. This job - # reports; it does not block. Making it blocking is a repository-settings - # change, outside the tree. - repo-gates: + # check runs first: secret-scan and self-hygiene are diff-only checks + # (no build, no test binary), done in seconds, so a leaked credential fails + # the run before build-shell/build spend minutes producing a binary from + # code that's already going to be rejected, and a CLI surface change is + # named in the log for the reviewer. + check: runs-on: ubuntu-latest steps: - uses: actions/checkout@v7 with: fetch-depth: 0 - # Resolved in its own step so the fallback is testable as shell rather - # than buried in a `${{ }}` ternary, and so both gates below read one - # value that has already been validated. - # - # Named via `env:` rather than interpolated straight into `run:`. - # GitHub substitutes `${{ }}` into the step's text before the shell - # parses it; both values are GitHub-computed rather than contributor - # text, so there is no injection here today, but the indirection is the - # shape that stays safe if a future edit swaps in something - # contributor-controlled. - id: base env: PR_BASE: ${{ github.event.pull_request.base.sha }} @@ -261,7 +55,7 @@ jobs: # valid diff would, never less, so it cannot narrow what gets swept. if ! git cat-file -e "$base^{commit}" 2>/dev/null; then widened=$(git rev-list --max-parents=0 HEAD | tail -1) - echo "repo-gates: base ref '$base' is unreadable (likely a force-push over it); widening to the root commit $widened instead of refusing" >&2 + echo "check: base ref '$base' is unreadable (likely a force-push over it); widening to the root commit $widened instead of refusing" >&2 base="$widened" fi @@ -273,38 +67,27 @@ jobs: - name: self-hygiene run: ./scripts/qa/self-hygiene.sh "${{ steps.base.outputs.sha }}" - build-shell: - runs-on: ${{ matrix.runner }} - strategy: - matrix: - runner: - - macos-latest - - macos-latest-large - - ubuntu-latest - - ubuntu-latest-arm64 + # The vorpal build compiles cmd/docket only; it runs no tests, so the Go + # suite needs its own job. + go-test: + needs: + - check + runs-on: ubuntu-latest steps: - uses: actions/checkout@v7 - - uses: ALT-F4-LLC/setup-vorpal-action@0.2.0 + - uses: actions/setup-go@v6 with: - registry-backend: s3 - registry-backend-s3-bucket: altf4llc-vorpal-registry - version: 0.4.0 - env: - AWS_ACCESS_KEY_ID: ${{ secrets.AWS_ACCESS_KEY_ID }} - AWS_DEFAULT_REGION: ${{ vars.AWS_DEFAULT_REGION }} - AWS_SECRET_ACCESS_KEY: ${{ secrets.AWS_SECRET_ACCESS_KEY }} - - - run: vorpal build 'docket-shell' + go-version-file: go.mod + cache: true - - uses: actions/upload-artifact@v7 - with: - name: vorpal-lock-${{ runner.arch }}-${{ runner.os }} - path: Vorpal.lock + - run: go test ./... + env: + DOCKET_SWEEP_WORKFLOWS_DIR: ${{ github.workspace }}/internal/workflow/testdata/ci-routing-corpus build: needs: - - build-shell + - check runs-on: ${{ matrix.runner }} strategy: matrix: @@ -328,6 +111,8 @@ jobs: AWS_DEFAULT_REGION: ${{ vars.AWS_DEFAULT_REGION }} AWS_SECRET_ACCESS_KEY: ${{ secrets.AWS_SECRET_ACCESS_KEY }} + - run: vorpal build '${{ matrix.artifact }}-shell' + - run: | ARTIFACT_PATH=$(vorpal build --path '${{ matrix.artifact }}')/bin/docket echo "ARTIFACT_PATH=$ARTIFACT_PATH" >> $GITHUB_ENV @@ -346,10 +131,29 @@ jobs: name: docket-dist-${{ env.ARCH }}-${{ env.OS }} path: ${{ env.ARCHIVE_NAME }} + test: + needs: + - build + runs-on: ubuntu-latest + steps: + - uses: actions/checkout@v7 + + - uses: actions/download-artifact@v8 + with: + name: docket-dist-x86_64-linux + path: dist + + - run: | + tar -xzf dist/docket-x86_64-linux.tar.gz -C dist + chmod +x dist/docket + echo "DOCKET_BIN=$(pwd)/dist/docket" >> "$GITHUB_ENV" + + - run: ./scripts/qa.sh "$DOCKET_BIN" + release: if: ${{ github.event_name == 'push' && contains(github.ref, 'refs/tags/') }} needs: - - build + - test permissions: attestations: write contents: write diff --git a/Makefile b/Makefile index 847abd06..8d30e842 100644 --- a/Makefile +++ b/Makefile @@ -1,4 +1,22 @@ -.PHONY: build test lint vet install clean demo +# Developer targets and the gate targets `docket trust list` invokes. +# +# The gate targets exist so a trust entry can name `make ` instead of an +# absolute path into one worktree's checkout. Trust entries are bound to the +# repository, not the worktree, and the engine spawns each gate with cwd set to +# the step's worktree — an absolute path pins every worktree's gate to whichever +# checkout happened to be current when the entry was approved, and goes dead the +# day that worktree is removed. `make ` resolves against the caller's cwd, +# so it is correct from every worktree and survives their coming and going. +# +# Each gate target's NAME MATCHES ITS TRUST ENTRY NAME exactly, so a workflow's +# `gates = ["build", "tests"]` reads the same as the Makefile and the same as +# `docket trust list`. That is why the compile gate owns the name `build` and +# the binary build is `make bin`. + +.PHONY: bin test lint vet install clean demo \ + build tests qa-test self-hygiene doc-validate citation-check secret-scan \ + vuln-scan sdet-abuse tdd-preflight reserved-name-check \ + render-verify copy-verify ac-commands VERSION ?= $(shell git describe --tags --always --dirty 2>/dev/null || echo "dev") COMMIT ?= $(shell git rev-parse --short HEAD 2>/dev/null || echo "none") @@ -6,7 +24,15 @@ BUILD_DATE ?= $(shell date -u '+%Y-%m-%dT%H:%M:%SZ') LDFLAGS := -X github.com/ALT-F4-LLC/docket/internal/cli.version=$(VERSION) -X github.com/ALT-F4-LLC/docket/internal/cli.commit=$(COMMIT) -X github.com/ALT-F4-LLC/docket/internal/cli.buildDate=$(BUILD_DATE) -build: +QA := scripts/qa + +# --------------------------------------------------------------------------- +# Developer targets +# --------------------------------------------------------------------------- + +# bin — produce ./bin/docket. Formerly `build`; renamed so the compile gate can +# own that name (see the header). +bin: CGO_ENABLED=0 go build -ldflags "$(LDFLAGS)" -o ./bin/docket ./cmd/docket test: @@ -24,6 +50,55 @@ install: clean: rm -rf ./bin/ -demo: build +demo: bin @command -v vhs >/dev/null 2>&1 || { echo "vhs is required: brew install vhs"; exit 1; } vhs scripts/demo.tape + +# --------------------------------------------------------------------------- +# Gate targets — one per `docket trust list` entry +# +# Each delegates to the script that already carries the gate's logic and its +# failure explanation; this file adds a stable, worktree-independent name and +# nothing else. A gate's behavior is changed in its script, not here. +# --------------------------------------------------------------------------- + +build: + @bash $(QA)/build.sh + +tests: + @bash $(QA)/tests.sh + +qa-test: tests + +self-hygiene: + @bash $(QA)/self-hygiene.sh + +doc-validate: + @bash $(QA)/doc-validate.sh + +citation-check: + @bash $(QA)/citation-check.sh + +secret-scan: + @bash $(QA)/secret-scan.sh + +vuln-scan: + @bash $(QA)/vuln-scan.sh + +sdet-abuse: + @bash $(QA)/sdet-abuse.sh + +tdd-preflight: + @bash $(QA)/tdd-preflight.sh + +reserved-name-check: + @bash $(QA)/reserved-name-check.sh + +render-verify: + @bash $(QA)/render-verify.sh + +copy-verify: + @bash $(QA)/copy-verify.sh + +ac-commands: + @bash $(QA)/ac-commands.sh diff --git a/README.md b/README.md index a37ffe40..a5af7910 100644 --- a/README.md +++ b/README.md @@ -40,26 +40,7 @@ DOCKET_INSTALL_DIR=/usr/local/bin curl -fsSL https://raw.githubusercontent.com/A ### Agent Skill Setup -If you're working with Docket through an AI coding agent, also install the companion skill so the agent has the full CLI workflow and command reference available without re-deriving usage from `--help` output. Copy this repo's [`skills/docket/SKILL.md`](skills/docket/SKILL.md) into your agent's skill directory: - -| Harness | Project-scoped | User-scoped | -|---------|-----------------|-------------| -| Claude Code | `.claude/skills/docket/SKILL.md` | `~/.claude/skills/docket/SKILL.md` | -| Codex | `.agents/skills/docket/SKILL.md` | `~/.agents/skills/docket/SKILL.md` | -| Opencode | `.opencode/skills/docket/SKILL.md` | `~/.config/opencode/skills/docket/SKILL.md` | - -```bash -# Claude Code (project-scoped) -mkdir -p .claude/skills/docket && cp skills/docket/SKILL.md .claude/skills/docket/SKILL.md - -# Codex (project-scoped) -mkdir -p .agents/skills/docket && cp skills/docket/SKILL.md .agents/skills/docket/SKILL.md - -# Opencode (project-scoped) -mkdir -p .opencode/skills/docket && cp skills/docket/SKILL.md .opencode/skills/docket/SKILL.md -``` - -See [Drop-in Skill](#drop-in-skill) for what the skill teaches. +If you're working with Docket through an AI coding agent, also install the companion skill so the agent has the full CLI workflow and command reference available without re-deriving usage from `--help` output. The skill is maintained in [ALT-F4-LLC/dotfiles.vorpal](https://github.com/ALT-F4-LLC/dotfiles.vorpal) at `src/user/claude_code/skills/docket/`; copy that directory into your agent's skill directory (e.g. `.claude/skills/docket/` project-scoped, or `~/.claude/skills/docket/` user-scoped for Claude Code). ### From Source @@ -67,7 +48,7 @@ Requires Go 1.26.0+ (toolchain go1.26.5). ```bash # Build the binary to ./bin/docket -make build +make bin ./bin/docket --help # Install to $GOPATH/bin @@ -142,7 +123,7 @@ Any agent that can run shell commands works with Docket. Point it at `docket nex ### Drop-in Skill -[`skills/docket/SKILL.md`](skills/docket/SKILL.md) is a thorough reference teaching the full Docket CLI workflow and command/flag reference in one file, so your agent doesn't need to re-derive usage from `--help` output. See [Agent Skill Setup](#agent-skill-setup) in the Installation section above for where to drop it in for Claude Code, Codex, and Opencode. +A companion skill teaching the full Docket CLI workflow and command/flag reference is maintained in [ALT-F4-LLC/dotfiles.vorpal](https://github.com/ALT-F4-LLC/dotfiles.vorpal) at `src/user/claude_code/skills/docket/`, so your agent doesn't need to re-derive usage from `--help` output. See [Agent Skill Setup](#agent-skill-setup) in the Installation section above.
Verbose JSON examples @@ -237,6 +218,8 @@ docket plan --json --root DKT-3 docket issue list --json -s todo -s in-progress -p high ``` +Listings — `issue list`, `next`, `plan`, and `board` — emit summary rows under `--json`: every issue field except `description`, plus `description_bytes`. Read one issue's full description with `issue show`, or pass `--with-body` to include every description in the listing. +
@@ -306,9 +289,16 @@ docket issue list --json -s todo -s in-progress -p high | Command | Description | |---------|-------------| | `docket next` | Show work-ready issues (unblocked, sorted by priority) | -| `docket plan` | Compute a phased execution plan from the dependency graph (filterable by `--status`/`-s`, `--label`/`-l`, `--priority`/`-p`, `--type`/`-T`, `--assignee`/`-a`) | +| `docket plan` | Compute a phased execution plan from the dependency graph (filterable by `--status`/`-s`, `--label`/`-l`, `--priority`/`-p`, `--size`, `--type`/`-T`, `--assignee`/`-a`) | | `docket board` | Kanban board view in the terminal | +### Report Commands + +| Command | Description | +|---------|-------------| +| `docket run report RUN-N` | Roll up one run's budget, steps, gates, actions, artifacts, findings, and metadata | +| `docket report executors` | Group fix-loop routings, override-pass rulings, reaps, and cluster corroboration per executor hint and voter name across runs (`--since RUN-N` or a date, `--all-projects`); read-only, never consulted by `next` | + ### Top-Level Commands | Command | Description | @@ -375,7 +365,9 @@ one. ### Consolidating legacy per-repo stores Move an old repo-local database into the shared store with export/import — -colliding ids are remapped automatically, nothing is dropped: +ids colliding with another project's rows are remapped automatically, nothing +is dropped. Merging a project's own export back with `--merge` skips the rows +it already holds: ```bash DOCKET_PATH=/path/to/repo/.docket docket export -f repo.json @@ -462,7 +454,8 @@ scripts/ ```bash git clone https://github.com/ALT-F4-LLC/docket.git cd docket -make build # Build to ./bin/docket +make bin # Build to ./bin/docket +make build # Run the compile gate (scripts/qa/build.sh) make test # Run unit tests make lint # Run staticcheck + go vet make clean # Remove build artifacts diff --git a/cmd/docket/main.go b/cmd/docket/main.go index 8b9dcc56..900d68e6 100644 --- a/cmd/docket/main.go +++ b/cmd/docket/main.go @@ -1,3 +1,4 @@ +// Command docket is the CLI entry point for Docket. package main import ( diff --git a/cmd/vorpal/main.go b/cmd/vorpal/main.go index ee3859a9..add3572d 100644 --- a/cmd/vorpal/main.go +++ b/cmd/vorpal/main.go @@ -11,14 +11,7 @@ import ( func main() { ctx := config.GetContext() - ctxTarget := ctx.GetTarget() - - systems := []string{ - "aarch64-darwin", - "aarch64-linux", - "x86_64-darwin", - "x86_64-linux", - } + ctxTarget := ctx.GetTargetStr() ffmpeg, err := ctx.FetchArtifactAlias("ffmpeg:8.0.1") if err != nil { @@ -81,7 +74,7 @@ func main() { } _, err = artifact. - NewDevelopmentEnvironment("docket-shell", systems). + NewDevelopmentEnvironment("docket-shell", config.SYSTEMS). WithArtifacts([]*string{ ffmpeg, gobin, @@ -104,7 +97,7 @@ func main() { log.Fatalf("error building project environment: %v", err) } - _, err = language.NewGo("docket", systems). + _, err = language.NewGo("docket", config.SYSTEMS). WithBuildDirectory("cmd/docket"). WithIncludes([]string{ "cmd/docket", diff --git a/docs/design/engine-core.md b/docs/design/engine-core.md index 3507c77c..cd783ab3 100644 --- a/docs/design/engine-core.md +++ b/docs/design/engine-core.md @@ -289,7 +289,15 @@ Budget and control inputs are core-owned (engine-spec §11): per-step `expected_ in workflow definitions, the per-run cap (`docket run start --budget`, `docket config` default), attempt caps (step fields, config defaults), and concurrency + lease TTLs per executor class (`[limits]`). The instance's `policy.toml` -carries only dispatcher-side spawn routing — the engine never reads it. +carries spawn routing, and the engine reads it. Activation pins its bytes by content +hash like any other config file. When a run pins one, the engine reads that file +whenever it renders a step row that needs routing, such as a `next` row, and +refuses with a CONFLICT if its bytes no longer match the hash recorded at +activation. From it the engine resolves each executor +row's `{model, effort, variant}` and each vote row's seat assignments. A variant that +`[sizes]` maps from the issue's size label replaces the `[executors]` starting variant +(internal/engine/policy_resolve.go `sizeVariant`). A run that pins no `policy.toml` +leaves those fields absent. Usage is recorded per step (executors attach what they observe; the reference harness back-fills from its dispatch journal, source recorded), but the cap does not rest on @@ -352,7 +360,7 @@ consequences, mechanically: so no torn state exists. - **Dispatcher session dies** ⇒ nothing happens. Leases lapse; the run sits `active` with a ready set. Any new dispatcher session picks up exactly where things stood — - the reference harness injects `docket run status --active` at session start so + the reference harness injects `docket run status` at session start so resumption is automatic (upstream 03 §7). There is no handoff narrative, no crash-recovery doctrine. - **Docket contention** ⇒ SQLite WAL + busy timeout + single-transaction claims diff --git a/docs/design/engine-spec.md b/docs/design/engine-spec.md index 6651f30f..a0893095 100644 --- a/docs/design/engine-spec.md +++ b/docs/design/engine-spec.md @@ -44,7 +44,11 @@ docket guard spawn|record|stop|gate --step NAME [--input …] docket workflow init [--template NAME] # scaffold instance config from shipped # optional templates (zero-authoring start) -docket run start --request-file … ; activate; pause|resume|abandon; status; report +docket run start --request-file … ; activate; conduct; pause|resume|abandon; status; report +docket report executors [--since RUN-N|DATE] [--all-projects] + # cross-run ledger per executor hint and + # voter name; read-only, operator-facing, + # never consulted by next (DKT-2453) docket run budget RUN-N --set N # raise/lower a live cap (CAS, # event-logged — DKT-29, stage 7) docket next --run RUN-N --json # step-level ready set (engine-core §5) @@ -52,6 +56,11 @@ docket dispatch open|close|verify|abandon --run RUN-N # batch manifest (TTL'd, one open per run, # CAS); next refuses while open or # discrepancies exist; abandon unsticks +docket run note add|list RUN-N # a standing statement recorded once + # against the run, rendered as + # `== RUN NOTE N` in EVERY packet the + # run renders from then on and carried + # by `context.notes` (DKT-1079) docket step claim STEP-N # atomic; mints a capability token and # returns token + the step CONTEXT bundle docket step heartbeat|complete|fail # token via DOCKET_TOKEN env or stdin @@ -78,7 +87,8 @@ surface differs: **Workflows.** Registered TOML in the generic grammar (§11.1): `[match]` predicates over kind/labels; steps with `executor` (opaque hint), `after`, `inputs`, `fanout`, -`gates`, `threshold`, `on_fail` routing (closed vocabulary), loops with specified +`gates`, `threshold`, `on_fail` routing (a closed vocabulary, or — on an +`executor` step — the name of a `type="vote"` triage panel), loops with specified re-expansion rules (step identity per loop entry, threshold re-application, gate re-runs — the DSL's loop semantics are part of this spec, not implementation discretion), `max_attempts`, `expected_cost`, `metadata` passthrough. Validation runs @@ -144,6 +154,38 @@ override-pass, note recorded); `approve|reject` belongs to `type=human` gate ste only, and a human gate's reject routing may not itself be `waiting-human` (register-time VALIDATION_ERROR). +**Park class.** Every park records `park_class` beside `park_reason`: the reason +is prose for the person the park escalates to, the class is the closed enum a +conductor or panel routes on. It is assigned inside the routing transaction from +that decision's own facts — a gate row's verdict word, a tally, an +`EnterLoop` outcome — and never by reading the reason text, so rewording a +message cannot move a row between classes. `SetStepRoutingTx` writes both halves +in one statement, under the same condition, and REFUSES a park that names no +class; a resolution, which writes a different status, overwrites neither. +A blank class means the row parked before the column existed. + +| value | produced by | +|---|---| +| `gate-failed` | a completion gate ran and did not pass — the routing stage's failed-verdict branch | +| `gate-unmatched` | a gate named no trust entry, so it never ran: the same branch when a recorded row's verdict is `unmatched`, and the gate-resolution path that parks on an unmatched entry outright | +| `gate-skipped` | a gate measured nothing because the judged commit could not be bound — the routing stage's unmeasured branch, decided before the failed-verdict one | +| `threshold-routed` | a declared `threshold` evaluated to a park, at the routing stage and at a vote step's approved-with-concerns evaluation | +| `loop-bound` | a `fix-loop` entry refused because the ordinal would exceed `max_fix_loops` plus grants — every `applyFixLoop` caller, from `EnterLoop`'s refusal rather than its wording | +| `vote-rejected` | a vote step tallied to a rejection | +| `held-rejected` | an operator rejected a held cluster; the consequence lands on the ROUTING step, never the materialized one, which ends `done` either way | +| `gap-only` | a completion whose only recorded artifacts were gaps, decided before any gate verdict | +| `unchanged-handback` | a loop round handed back the commit its previous round recorded, so the review chain would re-read one tree | +| `pass-floor` | a `pass` would have exited with declared `pass_floor` work still standing | +| `action-failed` | an action step's computation could not run — bad params, an unorderable value, an unmatched command name, or a non-zero exit | +| `attempts-exhausted` | the step spent its `max_attempts` budget | +| `join-missed` | a join completed below `min_siblings` | +| `triage-undecided` | a triage panel (DKT-1901) closed without a usable verdict — declined, no quorum, or retired without a tally — falling back to the same `waiting-human` a step with no panel would have reached | +| `held` | the run's conductor parked one READY step for the operator through `step hold` (`HoldStep`, DKT-3287), with its own stated reason; `step resolve` then retries, skips, or fails it | + +`paused` is not among them: it is a RUN status, never written to `steps.status`, +so a paused run removes its steps from the scheduler through their run rather +than by parking them, and no routing transaction produces it. + **Scheduling.** `next` computes readiness (engine-core §5): dependencies, predecessors, scope non-overlap, run active, concurrency headroom per executor-hint class (a generic knob; the reference instance's config sets its write class to 1 — @@ -177,14 +219,32 @@ or a pinned instance template file. Packet *layout* is thereby core mechanics wh instance's harness hands the packet to an LLM as its prompt; core neither knows nor cares.) +**Run notes.** `docket run note add RUN-N --text "…"` (or `--file F`, `-` for stdin) +records a standing statement against a run — once — and every packet the run +renders from then on carries it verbatim as a `== RUN NOTE N` section directly +after `== REQUEST`, for every step of every issue; `step context` carries it as +`context.notes`, so a contract can name it. It is the one steering channel that is +run-wide: `step resolve -m` reaches the step it rules on (and the round it +authorizes), which is right for a ruling about one step's work and useless for a +fact about the whole run — a required gate known to fail on clean HEAD, the issue +tracking it, and the disposition already given, which a dispatcher learns before +the first dispatch and which every worker otherwise rediscovers and re-files +(DKT-1079). Issue comments remain an audit surface and never render. Notes are +append-only (a changed ruling is a second note, rendered after the first), capped +at 16 KiB because each rides every packet, refused on a terminal run, and each is +recorded as a `run-note-added` event carrying its text, attributed human. +`docket run note list RUN-N` reads them back in render order. + **Guards.** `docket guard spawn|record|stop|gate` are deterministic allow/deny predicates over engine state for harness enforcement points (exit 0 allow / exit 2 deny with reason): `spawn` — proposed rows byte-match the open dispatch and no unacknowledged write reaps; `record` — no unreconciled dispatch; `stop` — no pending -work outside `waiting-human`; `gate --step NAME` — an approved `type=human` step of -that name exists for the active run (the reference instance's commit hook shims -`guard gate --step commit-gate`). Any harness's hook mechanism wires these as -one-liners; the logic lives here. (`step heartbeat` serves the heartbeat hook — an +work outside `waiting-human`; `gate --step NAME [--run RUN-N]` — an approved +`type=human` or `type=vote` step of that name exists in the named run, or with no +`--run` in any active run of the project (the reference instance's commit hook shims +`guard gate --step commit-gate`; the unscoped form lets one run's approval answer +for every caller in the project, so a hook that knows its run should pass it). Any +harness's hook mechanism wires these as one-liners; the logic lives here. (`step heartbeat` serves the heartbeat hook — an engine verb, though not a guard predicate.) **Payloads and thresholds.** `schema register` accepts JSON Schema plus an @@ -192,9 +252,9 @@ engine verb, though not a guard predicate.) (`threshold = { "fix-loop" = "any(severity >= high)" }` works because *the user's schema* declared the order — core never knows what a severity is). Aggregations beyond comparison are **action steps**. One is builtin and generic: `action = "aggregate"` -with `params = { field, method = median|max|min, hold_spread, output }` computes over -any ordered-enum payload field — median, spread-hold, and a recorded demotion trail -work for severities, priorities, or tiers alike. Cluster membership arrives in the +with `params = { field, method = median|max|min, hold_spread, output, route_at, source_field }` +computes over any ordered-enum payload field — median, spread-hold, and a recorded +demotion trail work for severities, priorities, or tiers alike. Cluster membership arrives in the payload itself: each element is one cluster, whose `field` value is either a scalar (a one-member cluster — the identity case) or an array of the cluster's member values; the builtin's input payload is the concatenated payloads of the step's @@ -205,7 +265,23 @@ that cluster's payload index (`-held@k#i`), together gating the routing st *(amended 2026-08-07, DKT-15: one step for the whole hold made approve/reject binary over the set, so an operator who wanted two clusters escalated and two accepted could not say so)*. The routing step waits for every one of them and routes once: per -`on_fail` if any was rejected, otherwise through the threshold. The +`on_fail` if any was rejected, otherwise through the threshold. An optional +`route_at = ""` names a routing floor in the field's declared order: only +clusters whose reduced value's position is at or above it are emitted to the output +payload — what the threshold evaluates and downstream `inputs` read — while the rest +are recorded, fully reduced, on the aggregate's own `action_results` row and never +enter the loop; a held cluster is never routed below the floor (the operator's +decision, not the untrusted computed value, decides it), an unknown `route_at` value +is a register-time refusal naming it, and with `route_at` absent the output is +byte-for-byte what it always was *(amended 2026-08-23, DKT-593)*. An optional +`source_field = ""` names a property of each input element holding an +array of opaque source labels (a workflow's own step-ref convention, say); +core neither validates its contents nor reduces by it — G3's verbatim +carry-through already moves whatever key it names into the output payload +unchanged — the param exists solely so the run report can group a round's +clusters by that key without core hardcoding one corpus's field name, the same +reason `route_at` takes its floor as a value rather than assuming a field +*(amended 2026-09-13, DKT-2462)*. The aggregate's output payload — per-cluster value, members, held flag, `demoted_from`, `operator_resolved` — validates against the shipped `aggregate@1` schema, and `operator_resolved` is set per cluster, on the approved ones only. The @@ -278,6 +354,21 @@ anything — under a trust model fit for an OSS tool: - Trust entries default to **full-argv hashes**; prefix entries are explicit opt-in (`trust add --prefix`, with an over-authorization warning). Tokens pass via env/stdin, never argv; claim markers are 0600 in a per-user runtime dir. +- **Conductor capability (DKT-2465).** The operator verbs — `step approve`, `step hold` (DKT-3287), + `step reject`, `step resolve`, `step reap`, `run pause`, `run resume`, `run + abandon` — are not token-free. Each requires the run's *conductor capability*: a + 256-bit token minted at the run's first activation (returned once, hash-only + storage on `runs.conductor_token_hash`), re-minted by `run conduct RUN-N`, and + presented via `DOCKET_TOKEN`/stdin like a lease token. "Repository access is the + authority" stopped holding once a harness gave every executor the operator's + checkout; the capability is what an executor is never handed, so the documented + path refuses it the way `step record` refuses an unclaimed worker. A run + activated before the capability existed stays open until it is conducted (the + guard-with-no-engine posture: nothing is required where nothing was minted). The + mechanism is tamper-evident, not tamper-proof: `run conduct` is deliberately + open to any caller, retires the standing token, and records who took the seat + (`conductor-seated`, with actor and cwd); a harness that keys its callers keeps + executors off that one verb. - **Conversational trust (zero-touch posture):** in this solution the session proposes, the human approves in-chat, and the session runs `trust add --yes` — the harness's own command-permission prompt is the human-confirmation backstop. The @@ -411,6 +502,79 @@ registered. `--restore` reverses a retirement. There is deliberately **no delete verb**: old versions stay registered, which is what keeps lineage readable. +A registered schema version may likewise be retired with +`docket schema deprecate @` *(amended 2026-09-24 — DKT-2792)*. +The contract is the workflow one applied to the other registry: the row stays, +`schema show` at an explicit `@version` still renders it, `--body` still emits +the registered bytes, and a run that pinned it still validates its payloads +against it. What retirement stops is **new references**: `workflow register`, +`workflow lint`, and activation's auto-registration refuse a step whose +`payload` names a retired version, naming the schema and the restore remedy. +`schema list` hides retired versions unless `--deprecated` is passed, a bare +`schema show NAME` resolves the highest version still in service, and +`registry audit` reports an orphaned schema whose every version is retired as +`retired: true`. Two refusals the workflow verb lacks: a schema that a workflow +version **still in service** names as a payload is refused (`CONFLICT`, listing +the referencing versions, no override — retire or re-version those workflows +first), and the builtin schema shipped in the binary cannot be retired, since it +is visible to every project through its flag. `--restore` reverses a +retirement. + +A registry is **per project** (`UNIQUE(project_id, name, version)`), and +`workflow register`, `workflow deprecate`, `schema register`, and +`schema deprecate` therefore +accept `--project ` and `--all-projects` *(amended 2026-08-24 — DKT-615; +`schema deprecate` added 2026-09-24 — DKT-2792, reporting `in-use` as its own +per-project outcome)*. +Without either flag each verb writes to the project the working directory +resolves to, unchanged, and emits the row it always emitted. With either flag it +emits a **per-project report** instead — one outcome per target +(`registered` / `unchanged` / `deprecated` / `restored` / `already-binding` / +`already-deprecated` / `conflict` / `not-registered` / `invalid`) — and each +project's own idempotency and conflict rules are decided *there*: a CONFLICT in +one project neither cancels nor hides another's registration, and the sweep +runs to the end. The process exits with the failures' shared code when they +agree and GENERAL_ERROR when they do not, having already written the report. +`workflow register`'s **environment validation runs per target**, because +`vote_rule` and `payload` references resolve against the registry of the project +being written to — the same bytes can be valid in one project and name a schema +that does not exist in the next, and storing them there anyway would defer a +guaranteed activation failure. Register schemas store-wide first. + +The two registry-**reading** verbs that plan those writes, `workflow list` and +`workflow lint`, accept `--project ` alone *(amended 2026-09-24 — +DKT-2793)*, taking the same ref forms (prefix, name, identity path, row id) and +failing on an unknown ref with the same error. `workflow list --project B` +lists B's registry, with `--deprecated`, `--orphans`, and `--name` applying +there; `workflow lint --project B` resolves `vote_rule` and `payload` +references in B and reports `new` / `unchanged` / CONFLICT against B's rows — +the verdict `workflow register --project B` would reach. `workflow lint` also +takes `--all-projects`, writing the per-project report `workflow register +--all-projects` writes with each project's verdict as its outcome (`new` / +`unchanged` / `conflict` / `invalid`); `workflow list` does not, since a +store-wide registry reading is `registry audit`'s. A project whose checkout is +missing from this machine can therefore be read from any other, as it could +already be written. + +`[match]` also accepts `domain_paths = [..]`, a list of path globs naming the paths +this pipeline's domain occupies *(amended 2026-09-03 — DKT-1182)*. **It binds nothing.** +It is not evaluated by the match predicate and no scope agreement can make a workflow +bind an issue its labels do not select; routing stays keyed on `kind` and labels. Its +sole consumer is an activation **lint**: an issue whose declared scope +(`issues.scope_globs`) lies *entirely* inside some other bindable workflow's +`domain_paths`, while lacking only the labels that workflow's `[match]` requires, is +named in the activation report (`binding_warnings` in the JSON envelope, a stderr line +in human mode) alongside both workflows and the missing labels. This is the +**exactly-one-WRONG-match** case: exactly-one-match refuses zero and refuses several, +but a mis-labelled issue matches its wrong workflow exactly once, so the count has +nothing to catch — HRN-1118 was scoped entirely to a TUI test file, carried `qa` and +not `ui`, bound the label-less baseline and silently lost the UI pipeline's judge +fanout and render gates. The lint **warns and never refuses** (a cross-domain binding +may be deliberate), stays quiet unless *every* scope glob is inside the domain, and +stays quiet where no labelling could have bound the other workflow anyway — a `kind` +it does not list, or an `unless_labels` entry that fires. A workflow that declares no +`domain_paths` is linted against nothing. + `[limits]` — optional map of executor *class* → `{ max = N, lease_ttl = "45m" }` (bare int = shorthand for `max`). When a run pins multiple workflows, the most restrictive limit per class wins; unset values fall back to `docket config` defaults. Classes also @@ -432,19 +596,29 @@ is exactly `"write" = { max = 1, lease_ttl = "45m", max_step_duration = "2h" }`. | `emits` | artifact-kind string, required on executor steps | binds the step to its recorded artifact kind (`inputs` resolution; an instance contract may mirror it for its worker's benefit — the workflow is authoritative) | | `payload` | `schema@ver`, optional | payload validated at `complete`; threshold fields check against it at register time; required on `action = "aggregate"` steps *(amended 2026-08-03, DKT-25)* | | `voters`, `vote_rule` | [executor hints], proposal-config name | required on `type="vote"` steps — who casts, which existing Docket threshold config tallies | +| `reviews` | step name, optional; unset by default; only on `type="vote"` steps (V44) | names the step whose produced work this vote step reviews — the declared link `recuse = "executor"` compares a caster against. Register-time validated like `voters`/`vote_rule`: a name that is not a step of this workflow, or the vote step's own name, is a `VALIDATION_ERROR` (V44). A vote step with no `reviews` set has no declared producer, so `recuse = "executor"` is a no-op there *(added 2026-09-23, DKT-2468)* | +| `roster` | `"open"` \| `"strict"`, default `"open"`; only on `type="vote"` steps (V43) | who a cast may be attributed to: `open` counts the voter list and lets a cast under any name fill a seat (every pre-existing ballot); `strict` refuses `docket vote cast` from a `--voter` not among the step's pinned `voters` as a `VALIDATION_ERROR`, writing no vote. Declared here and pinned at activation rather than in engine config, because `config set` has no per-caller identity and a constrained seat could rewrite a config key before casting *(added 2026-09-23, DKT-2562, DKT-2511)* | +| `weighting` | `"declared"` \| `"equal"`, default `"declared"`; only on `type="vote"` steps (V43) | what one cast is worth in the tally: `declared` prices it at the caster's own confidence × domain relevance (the existing arithmetic); `equal` counts every cast at 1.0 × 1.0 — the vote row still records the declared values, they price nothing. Same home as `roster`, for the same reason *(added 2026-09-23, DKT-2562, DKT-2512)* | +| `recuse` | `"none"` \| `"executor"`, default `"none"`; only on `type="vote"` steps (V43) | who is excluded: `executor` refuses a cast whose `--voter` equals the executor hint of the step named by `reviews` — its scalar `executor`, or any sibling name of its `fanout` — as a `VALIDATION_ERROR` naming both steps, so the party whose work is on the ballot cannot judge it. With `reviews` unset there is nothing to compare against and the switch does nothing *(added 2026-09-23, DKT-2562, DKT-2525)* | | `after` | [step names], **required** except the first step and `loop = true` steps (whose ordering comes from loop entry, §11.3) | intra-workflow predecessors; `[]` = root (implicit topology was a footgun) | -| `inputs` | [`"."` \| `".*"` \| `"issue.body"` \| `"issue.diff"`] | artifacts inlined into the context bundle, in order. `issue.diff` = the engine-computed VCS diff for the issue's scope, snapshotted and fingerprinted when its producing step completed (git in v1 — the one declared VCS coupling, §7) | +| `after_fired` | [step names], optional; every entry must also appear in `after` | predecessors this step runs ONLY IF THEY FIRED: when every instance of a named step ends `skipped` — an interposed gate its threshold routed elsewhere (§11.2), a false `when`, an `on_fail = "skip"` routing, an operator's `--as skip` — this step is terminalized `skipped` in the same transaction, and the skip cascades through every step declaring `after_fired` on it in turn. Additive beside `after`, whose meaning is unchanged: `skipped` still releases an `after` join (J1), and a skipped `after_fired` step contributes no input (J3). The corpus case is a `drain-highs` executor that runs only on rounds `security-vote` actually decided *(added 2026-09-02, DKT-1085)* | +| `inputs` | [`"."` \| `".*"` \| `".vote-record"` \| `"issue.body"` \| `"issue.diff"` \| `"issue.linked.."`] | artifacts inlined into the context bundle, in order. `issue.diff` = the engine-computed VCS diff for the issue's scope, snapshotted and fingerprinted when its producing step completed (git in v1 — the one declared VCS coupling, §7). `.vote-record` = the named `type="vote"` step's recorded proposal — tally outcome, weighted score, and every cast with its rationale — engine-served from the existing vote machinery; the named step must be a vote step, and the `vote-record` kind is reserved from `emits` *(amended 2026-08-22, DKT-545)*. `issue.linked..` = a CROSS-ISSUE input: the latest recorded artifact of `` held by each issue this issue is linked to by `` (a relation type or its inverse form — `depends_on`, `dependency_of`, `blocks`, `blocked_by`, `relates_to`, `duplicates`, `duplicate_of`), resolved and pinned by artifact id at activation inside the fat transaction; activation fails loudly when the relation is missing or no linked issue holds the kind, so the binding is enforced rather than an issue-body citation. V11's produced-kind table deliberately does not apply — the producer is another issue's run — and the `issue.linked` name is reserved from step names as `issue.latest` is *(amended 2026-08-22, DKT-547)* | | `gates` | [trusted gate names \| `{name, source="fence:", pre=bool}`] | `pre = true` gates run at claim with results included in the context bundle (measure-then-judge steps); the rest run in order inside `complete` (§2, §4) | | `params` | opaque KV table | arguments to `action` steps (e.g. the builtin `aggregate`) | | `min_siblings` | int, default = all | fanout join quorum (§2 Fanout joins); the default is the plain join — quorum semantics (the `on_fail` routing at join) apply only when declared below the sibling count *(clarified 2026-08-03)* | -| `threshold` | table: routing → predicate (11.2) | routing computed over the step's recorded payloads | -| `on_fail` | `"fix-loop"` \| `"waiting-human"` \| `"skip"` \| `"abandon-issue"`; default `"waiting-human"` | routing for gate failure / attempts exhausted; `type="human"` steps must declare it explicitly and `"waiting-human"` is invalid there — reject routes per `on_fail` (§2's reject-routing rule; amended 2026-08-03) | +| `threshold` | table: routing → predicate (11.2) | routing computed over the step's recorded payloads; on `type="vote"` steps, over the tally's cast set after an APPROVED tally (11.2) *(amended 2026-08-22, DKT-545)* | +| `pass_floor` | `{ field, at }`, optional; requires `payload`, and `at` must be a value of `field`'s declared order (V37/V37a) | exit bar on a `pass` routing: when the routing resolves to `pass` but the step's recorded payload holds an element whose `field` value sits at or above `at`'s position — and the element is neither `held` nor `operator_resolved` — the step parks `waiting-human` instead of exiting, naming `--as override-pass` and `--as fix-round` as the ways out. Both values are opaque tokens compared only by position, `route_at`'s discipline; declared nowhere, nothing changes *(amended 2026-08-26, DKT-870: RUN-58's reconcile routed `pass` with all 16 clusters open, six at the order's high position — "converged" in the ledger meaning "dispositioned")* | +| `on_fail` | `"fix-loop"` \| `"waiting-human"` \| `"skip"` \| `"abandon-issue"` \| the name of a `type="vote"` step of this workflow (V40); default `"waiting-human"` | routing for gate failure / attempts exhausted; `type="human"` steps must declare it explicitly and `"waiting-human"` is invalid there — reject routes per `on_fail` (§2's reject-routing rule; amended 2026-08-03). A STEP NAME routes the failure to that TRIAGE PANEL instead of to an operator: the routing step must be an `executor` step, the named step must be `type="vote"`, and its `after` must include the routing step (V40). Only an executor may route this way because only its gate-failure path writes the suspension — every other failure source (a human gate's reject, a rejected tally, a quorum miss, a failed action) would terminalize the step `done` while naming the panel, releasing its successors as though it had passed. The failed step is then SUSPENDED — it keeps the non-terminal `gated` status, so nothing downstream is released and, unlike a `waiting-human` park, neither the issue (R2b) nor the run is held while the panel sits. The panel's proposal opens carrying the failure evidence: the failed step's instance and its `gate_results` rows with each gate's verdict, exit, and captured output. Its verdict is applied by `on_fail_routes` below *(added 2026-09-15, DKT-1901: a lane's first executor step has no review upstream, so `fix-loop` has nothing to route to and `waiting-human` was the only routing left — 190 of 328 parks on 2026-09-07 were an implement step failing its gates, 164 resolved `override-pass`)* | +| `on_fail_routes` | table: `"approved"` \| `"rejected"` → `"retry"` \| `"fix-round"` \| `"abandon-issue"` \| `"waiting-human"`; only on `type="vote"` steps (V40a), required on a panel some `on_fail` names (V40c) | what a triage panel's tally does to the step that routed its failure here, applying the same transitions `docket step resolve --as` performs for an operator: `retry` resets the retry budget, releases the lease and returns the step to `pending`; `fix-round` grants an authorized loop round and supersedes the step (refused at register unless a `loop = true` body serves the routing step, V40b — a first-lane executor usually has none); `abandon-issue` routes the step `failed-routed`; `waiting-human` is the panel declining to decide, which converts the suspension into the ordinary park with the panel's rationale beside the gate rows. LIMITS: the keys are the vote's whole verdict vocabulary — a tally either reaches its threshold or does not — so ONE panel expresses at most two of the four routings; a workflow needing others declares another panel. A tally that reaches NO verdict the mapping is keyed on (a quorum miss, a proposal closed by `docket vote close`, an operator's manual commit) applies no mapped routing. THE PANEL'S OWN `on_fail` — which V13a already requires it to declare — then disposes of the suspended step, and that is the human backstop: `waiting-human` (the default) parks it for an operator, `abandon-issue` routes it `failed-routed` with the issue cascade, `fix-loop` buys an authorized round; `skip` is refused on a panel some step routes to (V40c), because routing a failed executor away is not a triage outcome. The panel's own row is then `skipped`. The suspension may not outlive the panel — the operator's resolution verbs act on a park, so a step left suspended behind a closed panel would have nothing able to resolve it — which is why the closure always disposes of it. On a tally that DID reach a verdict the panel routes `pass`, since a rejection is the answer it was convened to give rather than its own failure. Every mapped routing is followed by the same issue-and-run reconcile `docket step resolve --as` performs, so a panel's `abandon-issue` runs the same cascade an operator's does. A PANEL ANSWERS ONCE PER ORDINAL: its proposal is keyed `(run, issue, instance)`, so a step that failed, was ruled `retry`, and failed again at the same ordinal parks `waiting-human` naming the spent ruling rather than suspending for a panel that can no longer answer. A step whose routing did NOT select the panel — the ordinary passing lane — terminalizes it `skipped`, exactly as an unrouted threshold target is *(added 2026-09-15, DKT-1901)* | +| `on_exhausted` | `"waiting-human"` \| `"abandon-issue"` \| the name of a `type="vote"` step of this workflow \| the name of an executor step of this workflow; default `"waiting-human"`; only on a step that routes `fix-loop` under a positive `max_fix_loops` (V41) | routing for FIX-LOOP EXHAUSTION — a `fix-loop` routing whose next ordinal would exceed `max_fix_loops` plus `fix-round` grants (11.3 (1)). `waiting-human` parks for an operator, byte for byte what every exhaustion did before the key existed; `abandon-issue` stops the run's work on the issue; a `type="vote"` step name opens that panel's proposal, which is how an author declares a loop-extension vote under a hard ceiling; an executor step name runs that step, so a file-and-stop default can run machine-side. A named step is an interposed target exactly as `on_fail` reads one. It is the exhaustion's routing, not a second bound: the `max_fix_loops` arithmetic is unchanged, and this says only what happens once it refuses. The refused entry also writes the loop history (11.3 (1)) on the triggering step's row *(added 2026-09-24, DKT-1902: 36 exhaustions in the shared store on 2026-09-07 were answered by an operator 34 different ways, work a workflow could not declare while the park was unconditional)* | | `loop` | bool, default false | marks loop-body steps (11.3) | | `after_loop` | step name | re-entry target after a loop body completes | +| `serves` | [step names], only on `loop = true` steps | scopes the body to the named steps' `fix-loop` routings — its loop CLUSTER (11.3); omitted = serves every trigger. Entries must name steps that can route `fix-loop`, and every step that can must be served by at least one body *(amended 2026-08-22, DKT-544)* | | `max_attempts` | int, default engine config | per-instance retry budget | -| `max_fix_loops` | int, default engine config | loop-entry budget per issue | +| `max_fix_loops` | int, default engine config | loop-entry budget per issue — ONE counter over EVERY `fix-loop` routing source (threshold, `on_fail`, rejected vote/human gate, quorum miss), read off whichever non-cluster step declares it. Each admitted entry post-increments the counter to its own 1-indexed ordinal; an entry whose new count exceeds the bound is refused with the counter restored, so `= N` admits exactly N entries and routes the N+1th per the triggering step's `on_exhausted` (default `waiting-human`). Only a `fix-round` grant (one per resolution, effective bound = declared + grants) admits more *(amended 2026-08-23, DKT-587)*. On a `serves`-scoped loop body it is instead that CLUSTER's round budget, checked independently under the issue-level ceiling — it never raises or lowers it *(amended 2026-08-22, DKT-544)* | +| `max_stalled_rounds` | int ≥ 0, default 0 (never fires); only on a step that can route `fix-loop` and records an artifact (V38) | non-convergence tolerance over THIS step's routed volume: a `fix-loop` entry after that many consecutive measured rounds in which the element count of the step's recorded payload never fell below the smallest count any earlier round recorded is refused in the non-convergence park's exact shape — counter restored, nothing instantiated, `waiting-human` naming `--as fix-round` as the way out, an authorized entry waived. "No improvement" means no new strict minimum, so volumes oscillating around a floor still park while a genuinely shrinking set never does *(amended 2026-08-26, DKT-870: RUN-51 held 8-12 clusters flat across TEN rounds and RUN-50 7-10 across six, both ended only by operator action — the plateau was the corpus's own non-convergence signal and nothing in the engine read it)* | | `expected_cost` | number ≥ 0, default 0 | budget-floor contribution per claim (§2) | -| `when` | predicate over issue `kind`/`labels` | step is `skipped` when false | +| `when` | predicate over issue `kind`/`labels` — clauses ` <==\|!=\|contains> ` or `labels contains-any (a, b, c)` / `labels contains_any [a, b, c]`, joined by `and` throughout or by `or` throughout | step is `skipped` when false. `or` holds when at least one clause does; a predicate MIXING `and` and `or` is a VALIDATION_ERROR (V22), because the grammar has no parentheses and therefore no reading of `a and b or c` to prefer — the mixed case is expressed as two steps, which is what the disjunction removed the need for in the common case *(amended 2026-08-22, DKT-548)*. `labels contains-any (…)` is the step-level spelling of the `labels_any` [match] clause and holds when the list intersects the issue's labels — a CLAUSE, not a connective, so "kind X and any of these labels" is one homogeneous-`and` predicate rather than a mix V22 would refuse. The list needs at least one element and its values carry no whitespace *(amended 2026-08-22, DKT-550)*. The operator is spelled `contains-any` or `contains_any` and its list is delimited by `(…)` or `[…]`; all four combinations are the same clause, and the delimiters must pair — `[a, b)` is a VALIDATION_ERROR. Both spellings were admitted rather than one because `contains-any (…)` is what registered definitions carry and `contains_any [a, b]` is how a list is written everywhere else in a workflow TOML, so refusing either would make an author's first correct guess an error *(amended 2026-09-01, DKT-1000)* | | `metadata` | opaque KV table | recorded on the step; delivered in the context bundle | ### 11.2 Threshold predicates @@ -469,20 +643,98 @@ whose registered schema declares `ordered_enum` (§2). Fields and literals are validated against the registered schema at `workflow register` time. Example (standard-change): `threshold = { "fix-loop" = "any(severity >= high)" }`. +**Reserved `diff.*` fields** *(added 2026-09-24, DKT-2063/DKT-2518/DKT-2551)*: three +field names address the engine's own measurement of the change a tree-holding +executor or fanout step recorded, not its payload — `diff.lines` (added plus removed content +lines), `diff.files` (files touched), and `diff.empty` (whether the recorded body +holds a change). They are evaluated at record time over the step's in-scope +`issue.diff` round record, the object a review reads; the aggregation is not +applied (one number per step), and the counts compare numerically without an +`ordered_enum` because the engine measured them and knows their order. A step +that recorded nothing evaluates as `diff.empty == true`, `diff.lines == 0`, +`diff.files == 0` — decided, not unknown. No schema declares them, so V21a–V21c +skip exactly these three names (a lookalike such as `diff.bogus` is still an +undeclared field); register time instead refuses a `diff.*` predicate on a step +that does not hold the tree (V45), a non-integer literal on `diff.lines` or +`diff.files` (V46), and an ordered operator or non-boolean literal on +`diff.empty` (V47). Measurement rules and rationale: payloads-thresholds TDD +§5.1. Example (a change track): `threshold = { "review" = "any(diff.lines > 20)", +"review-lite" = "all(diff.lines <= 20)" }`. + +**Vote-step thresholds** *(amended 2026-08-22, DKT-545)*: on a `type="vote"` step, +`threshold` is evaluated over the proposal's recorded **casts** — one element per +cast, addressable fields `vote` / `verdict` (aliases for the cast's verdict) and +`voter` — and only after an **APPROVED** tally. A rejected tally routes per +`on_fail`, exactly as before; a manually committed proposal (an operator setting +the final outcome by hand) skips the threshold. The routing vocabulary is +restricted to `"fix-loop"` / `"waiting-human"` / `"pass"` — step-name +interposition is not available on vote steps — and operators to equality, because +casts have no registered schema and ordered comparisons are defined only over +`ordered_enum` fields (all register-time rules: V36). First match routes, no +match ⇒ `"pass"`, and a step declaring no threshold behaves exactly as it always +did. Example (an investigation read-gate): +`threshold = { "fix-loop" = "count>=2(vote == approve-with-concerns)" }` sends an +approved-but-concerned tally into the same revise loop a rejection enters, +instead of the concerns evaporating; the loop body reads what the panel said +through `inputs = [".vote-record"]` (§11.1). + +**Who may cast, and at what weight** *(added 2026-09-23, DKT-2541)*: a +`type="vote"` step's `roster`, `weighting` and `recuse` fields (§11.1) are its +authorization switches, pinned at activation with `voters` and read by `docket +vote cast` from the pinned definition — never from engine config, which has no +per-caller identity and which the constrained party could rewrite before +casting. `reviews` is the declared link recusal needs: it names the step whose +produced work the panel judges, is register-time validated (V44: an unknown +step name, or the vote step's own, is a `VALIDATION_ERROR`), and `recuse = +"executor"` (default `none`) reads it to refuse a cast from that step's +executor hint — its scalar `executor` or any sibling name of its `fanout`. A +vote step with no `reviews` set makes `recuse = "executor"` a no-op, since there +is no declared producer to compare against. Every switch defaults to the +behavior every earlier ballot had, and a proposal no vote step opened (an +operator's own, or a reap acknowledgment's) enforces none of them. The +`vote.rule..roster` and `.weighting` config keys that once carried the +first two are retired and refused by name (DKT-2764). + +**Evidence on findings** *(amended 2026-09-14, DKT-2451)*: an entry of a cast's +structured findings may cite `artifact:ARTIFACT-N` (an artifact the run holds) +or `gate:` (a gate result the run recorded). `vote cast` resolves every +reference against the run the proposal was opened for before the cast records +and refuses an unresolvable one by name; core checks that a reference resolves +and never reads what it points at. `run report` lists every recorded finding +with its evidence and marks an entry that cited nothing `unsupported`, so an +asserted finding and a reproduced one are distinguishable in the record. An +entry with no evidence keeps its bare-string wire form, so nothing that reads +`findings_json` or the vote-record packet changes shape until a cast cites +something. + ### 11.3 Loop semantics (normative) Step instances are identified `name@k#i` — `k` = loop ordinal (0 at initial expansion), `#i` = fanout index (absent when not fanned out). When a routing resolves -to `"fix-loop"`: (1) the issue's loop counter increments; exceeding `max_fix_loops` -routes `waiting-human` instead — loops are bounded by construction. A round is also +to `"fix-loop"`: (1) the issue's loop counter increments; an entry exceeding +`max_fix_loops` (plus `fix-round` grants) routes per the triggering step's +`on_exhausted` instead: `waiting-human` when undeclared (the default), +`abandon-issue`, the name of a `type="vote"` step, or the name of an executor +step — only `waiting-human` parks, and loops are bounded by construction. A round is also refused as NON-CONVERGENT when the round below it left the issue's scope byte-identical to the round below that (`issue.diff` fingerprints), because the next -round would read the same tree and reach the same verdict; that refusal takes the same -shape as the bound's (nothing superseded, nothing instantiated, counter restored, -`waiting-human` naming `step resolve --as fix-round` as the way out) and is waived by -that resolution. It never fires on a diff that records no content — an empty or +round would read the same tree and reach the same verdict; that refusal always routes +`waiting-human` regardless of `on_exhausted` (nothing superseded, nothing instantiated, +counter restored, `waiting-human` naming `step resolve --as fix-round` as the way out) +and is waived by that resolution. It never fires on a diff that records no content — an empty or unresolvable-base measurement is evidence the tree was not measured, not evidence that -it did not move. (2) Not-yet-claimed +it did not move. A refused entry writes THREE LOOP-HISTORY FACTS on the triggering +step's row, in the same transaction as the routing decision, so no reader sees an +exhaustion whose history has not landed: the rounds that actually ran against the +cap, the instance whose verdict opened the loop (the triggering step's own +`name@k#i`), and the routing verdict the latest round ended on (the refused +`fix-loop`). They are readable as `loop_rounds_run`, `loop_trigger_step`, and +`loop_latest_verdict` on the step row `docket step show STEP-N --json=v2` emits +(11.4), `omitempty` so every row outside an exhaustion serializes exactly as before. +They are persisted rather than reconstructed because the event log they would +otherwise come from is prunable, and a router reading the park needs all three: +was the cap reached, what started the loop, and was the last round still finding +problems *(added 2026-09-24, DKT-2522/DKT-2546)*. (2) Not-yet-claimed instances downstream of `after_loop` (e.g. a pending `verify@0`) transition to the terminal status **`superseded`** (event-logged); claimed/running instances finish, but routing from a superseded lineage is inert. (3) Steps marked `loop = true` — excluded @@ -498,6 +750,24 @@ the highest existing ordinal ≤ the consumer's (mirroring input binding); re-instantiation never spans steps outside the `after_loop` chain. *(Clarified 2026-08-03, S3 stage review.)* +**Cluster scoping** *(amended 2026-08-22, DKT-544)*: a `loop = true` body may declare +`serves = [step names]`, scoping it to the named steps' `fix-loop` routings. The +TRIGGERING step — the one whose routing resolved to `fix-loop` — selects its +cluster: clauses (2)–(4) then apply to the serving bodies and to the downstream +chains of THOSE bodies' `after_loop` roots only (a body or `after_loop` declarer +without `serves` serves every trigger, so a workflow declaring no `serves` anywhere +has exactly one cluster and the original behavior, unchanged). The loop counter, +its ordinal sequence, the `max_fix_loops` ceiling read off non-cluster steps, the +non-convergence refusal, and `fix-round` grants all stay issue-level across every +cluster. A `max_fix_loops` declared on a `serves`-scoped body additionally bounds +that cluster's own rounds — counted as the distinct ordinals holding its scoped +bodies' instances (bodies serving several triggers are counted wherever they ran) — +and its refusal takes the ceiling's exact shape, waived once by the same +`fix-round` resolution. The `loop-entered` event data names the trigger alongside +the ordinal. Register-time rules: `serves` is valid only on `loop = true` steps, +every entry must name a step of the workflow that can route `fix-loop` (V35), and +every step that can route `fix-loop` must be served by at least one body (V17c). + Engine-enforced numbers live core-side, never in opaque pins: per-class lease TTLs and concurrency (`[limits]` / `docket config`), attempt caps (step fields / config defaults), the per-run budget cap (`docket run start --budget N`, config default), and @@ -511,7 +781,8 @@ next row { step, instance, issue, run, executor, class, attempt, claim response { step, token, lease_expires_ms, context } context { step: , issue: {id, title, body_snapshot, kind, labels, scope}, inputs: [{artifact, kind, producer_step, body, payload?}], - pins: [{path, sha256}], loop_entry, metadata, pre_gates? } + pins: [{path, sha256}], loop_entry, metadata, pre_gates?, + notes?: [{id, text, recorded_at_ms}] } # notes: DKT-1079 dispatch { dispatch, run, opened_seq, rows: […] } # verify = byte-equality on rows complete args --artifact-file F [--payload-file F] [--usage '{"unit":n,…}'] [--metadata '{…}'] (token via DOCKET_TOKEN env or stdin — §4) @@ -521,6 +792,11 @@ action result { step, action, argv, exit, duration_ms, output, truncated, verdict, builtin, reason? } # argv/exit NULL for builtin|unmatched # (added 2026-08-03, DKT-24) event { seq, at_ms, kind, run?, step?, step_id?, data } # step_id: DKT-34 +step row { , status, routing?, park_reason?, park_class?, + loop_rounds_run?, loop_trigger_step?, loop_latest_verdict? } + # `step show --json=v2`; the three + # loop_* keys are present only on a + # fix-loop-exhausted step (11.3 (1)) ``` `instance` is the rendered `name@k#i` identity (§11.3) carried alongside the `STEP-N` @@ -530,6 +806,11 @@ DKT-15)*. `pre_gates?` is an array of §11.4-shaped gate results for the step's a verdict is `unmatched` or timed out, null on ordinary pass/fail *(added 2026-08-03, DKT-19 / DKT-20)*. +`notes?` is the run's recorded notes (`docket run note add`, DKT-1079) in insertion +order, present only when the run has at least one, so every bundle of a run without +notes is byte-identical to before the member existed. `run note list RUN-N` and +`run note add`'s answer carry the same `{id, text, recorded_at_ms}` shape. + `docket step context STEP-N` re-emits `context` read-only (no token required; local inspection). `--meta` on it reports per-section byte counts — the closure-size record -(engine-core §8). +(engine-core §8) — `notes_bytes` among them, since a note rides every packet. diff --git a/docs/tdd/attempt-numbering.md b/docs/tdd/attempt-numbering.md index 1e26c8c9..6ecca3d2 100644 --- a/docs/tdd/attempt-numbering.md +++ b/docs/tdd/attempt-numbering.md @@ -6,7 +6,9 @@ docs/tdd/claims-leases.md §2, schema v6) across every surface that renders it, after RUN-5 (2026-08-08) found the rendering inconsistent enough to mislead two independent consumers. Amended for DKT-490 (schema v23): the retry-reset claim v16 obsoleted is corrected, and the failure-vs-reap marker §"What to -check" called for now exists (`failed_attempts` / `reaped_claims`). +check" called for now exists (`failed_attempts` / `reaped_claims`). Amended +again for DKT-1279 (schema v27): the breakdown's own gap — it counts, it +does not say which ending is the LAST one — gets `prior_attempt_end`. ## The one counter @@ -93,6 +95,20 @@ from the same number: back-fills nothing, so zero on pre-v23 claims means "no recorded breakdown"). An escalation policy that means "escalate after N *failures*" reads `failed_attempts`, and never `attempt`. +3. The breakdown counters answer "how many of each has this step EVER had", + which goes ambiguous the moment a step's history mixes both — one failure + and one reap, in either order, both read `failed_attempts: 1, + reaped_claims: 1`, and neither counter says which ending THIS re-offer + follows. `prior_attempt_end` (schema v27, DKT-1279) answers that directly: + `"failed"` or `"reaped"`, the end reason of the MOST RECENT claim to leave + the step, overwritten every time one ends (never a tally). RUN-80 + DISPATCH-400 is the motivating incident — ten leases reaped after a + session was killed mid-wave, the steps re-dispatched at `attempt` + incremented, and an `on_failure` escalation policy read that as "failed + once" and routed all ten a tier up, because a reap is a liveness event and + nothing on the row said so. Rides the same rows as `failed_attempts`/ + `reaped_claims` (`next`, `dispatch open`, `step show`, `step list`), + `omitempty`, absent for a step that has never had a claim end either way. ## Regression coverage @@ -108,3 +124,9 @@ at two moments, not two counters. `step fail` counts the reverse in both its branches, a lazy reap at `claim` counts like the scheduler's, a recorded park and its `resolve --as retry` touch neither, and `attempt` moves at claims alone throughout. + +`internal/engine/prior_attempt_end_test.go` pins `prior_attempt_end` +(DKT-1279): a lease-expiry reap's own `next` call renders the re-offered row +carrying `"reaped"` — the same-transaction snapshot reflection, not a second +read — an explicit `step fail` renders `"failed"`, a forced `step reap` +renders `"reaped"` too, and a step never claimed omits the field entirely. diff --git a/docs/tdd/claims-leases.md b/docs/tdd/claims-leases.md index 2dc1bc94..e6b5d299 100644 --- a/docs/tdd/claims-leases.md +++ b/docs/tdd/claims-leases.md @@ -256,6 +256,13 @@ Mapped to this stage's verbs. Every row is proven by a test (§7): | R7 | non-holder closes a live-leased issue | close | `AUTH_ERROR` | 5 | | R8 | claim against an expired lease | claim | **succeeds**, `attempt++` | 0 | +**Step claims diverge at R5.** An issue claim against a live lease returns +`CONFLICT` for every caller, including the lease's own owner. A step claim whose +`--owner` equals the live lease's recorded owner instead succeeds with a +re-minted token that replaces the prior one. That path trusts `--owner`, a +caller-supplied and publicly shown string, as the holder identity; see +engine-spine.md §6.9. + **R2 and R3 are both AUTH_ERROR, deliberately.** "This issue is unclaimed" and "your token is wrong" are the same answer to the caller — you do not hold this lease — and distinguishing them leaks whether a lease exists to a caller holding @@ -275,6 +282,28 @@ same guarantee §9.3 states for `complete`. refusal never writes — including never bumping `attempt` and never touching the CAS `version`. Proven by asserting the version is unchanged after each refusal. +### 4.0 The conductor capability's rows (DKT-2465, v29) + +The seven operator verbs — `step approve`, `step reject`, `step resolve`, +`step reap`, `run pause`, `run resume`, `run abandon` — reuse the matrix's +shape against the RUN's conductor capability (reliability-delta §2, the v29 +amendment; `internal/engine/conductor.go`). Same transport, same channels, +same codes: + +| # | Situation | Verb | Code | Exit | +|---|---|---|---|---| +| R9 | run bound, no token supplied | the seven | `VALIDATION_ERROR` | 3 | +| R10 | run bound, wrong token | the seven | `AUTH_ERROR` | 5 | +| R11 | run unbound (activated before v29, never conducted) | the seven | **allowed**, as before | 0 | +| R12 | `run conduct` on a terminal run | conduct | `CONFLICT` | 4 | + +R9 and R10 are checked after the step or run is found and before its status +is inspected or anything is written, so a caller without the capability learns +only that the target exists. R11 is the dormancy guarantee at the run level: +the check binds the moment a capability exists, and every run this binary +activates is bound at birth. There is no STALE_LEASE row: the capability has +no TTL, and it ends only with the run or with a `run conduct` that retires it. + ### 4.1 Unleased issues stay unleased An issue with `owner IS NULL` is not "claimed by nobody who must be checked" — diff --git a/docs/tdd/completion-metadata.md b/docs/tdd/completion-metadata.md index babd66ea..c41fec76 100644 --- a/docs/tdd/completion-metadata.md +++ b/docs/tdd/completion-metadata.md @@ -11,6 +11,12 @@ metadata"). Tracker unit: **DKT-68**. Precedent: the M2a bind-to-highest patch Spec of record is engine-spec.md; deviations become DKT amendment issues per docs/design/amendments.md, never silent changes. +**Amended 2026-08-23 (DKT-592): §1.7 adds the claim-side write.** A bag that +lands only at completion is a bag that failed and crashed steps never record, +so the dispatcher's half of a requested/resolved pair now lands AT CLAIM and +only the worker-reported half stays at completion. §1.1–§1.5 are unchanged for +the completion path; §2.D's "both are worker-reported" is corrected there. + **The defect is not a missing feature; it is a missing assignment.** Every other part of the path already exists and is correct: the flag parses (`internal/cli/step.go:310`), the option field is declared and documented as @@ -197,6 +203,78 @@ would lose exactly the diagnostic an operator wants. But that is DKT-69's call to make with its own acceptance criteria, and nothing here forecloses it: the merge function is indifferent, and the decision is one call site. +### 1.7 The claim-side write (DKT-592) — amendment, 2026-08-23 + +**The completion-time write is correct and incomplete.** A bag that lands only +at `complete` is a bag that a step which never completes never records — and a +step that failed or crashed is exactly the step whose dispatch facts an +operator most wants. Two runs measured the hole: 76 of 119 steps carried the +routing keys, then 55 of 83. The rollup this note exists to feed went blind at +precisely the rows that motivate reading it. + +The rule, and it is a **split by who knows the fact, not by which verb is +convenient**: + +- A fact the **dispatcher already knows when it hands the step out** is + recorded **at claim**. §2.D calls the motivating routing keys "both + *worker*-reported"; that was wrong about the requested half. What was asked + for is settled before the work starts — it is the dispatcher's own decision — + and holding it until the work comes back makes recording it conditional on + the work succeeding. +- A fact **only the returning worker knows** stays at **completion**. The + resolved half is genuinely worker-reported: nothing before the work runs can + say what actually served it. It does **not** move to claim, and §1.1–§1.5 are + unchanged for it. + +`ClaimOptions.Metadata` (`--metadata` on `step claim`) is the channel. +Mechanically it is the pieces this note already built, with no new ones: + +- **The same merge.** `mergeMetadata` over the step's stored bag, + last-write-wins per top-level key, shallow. A definition-side bag survives; a + re-claim overlays the dead attempt's; **the completion bag later overlays the + claim's**, which is what makes a normally-completing step carry both pairs. + The pure function §1.6 promised now has three callers and still knows nothing + about which verb called it. +- **The same writer**, `db.SetStepMetadataTx`, in the claim's **transaction A** + — with the CAS that awarded the claim, after the status write and before the + `step-claimed` event, mirroring §1.3's ordering rationale: the row reaches + its final shape for the transition before the transition is recorded. A crash + between them rolls back both. +- **The same cap and the same shape check**, `MetadataMaxBytes`, measured on + the raw input, **before `conn.Begin()`**. That placement carries more weight + here than on `complete` or `fail`: transaction A performs the lazy reap and + the CAS, so validation drifting inside it would put a malformed bag one + change away from consuming an attempt. Only the remedy text differs — a claim + has no artifact or payload channel of its own, so the message names its + completion's. + +**Survival is a property of what the other writers do not touch**, and it is +worth stating as such rather than assuming it: `ReapStepTx`, +`MarkStepAttemptFailedTx`, `MarkStepClaimReapedTx` and the status writers all +leave `metadata` alone. So the claim's bag survives a failure, a reap, an +abandonment, and the return to `pending` — with no cleanup path to audit. + +The claim's **own context bundle** reports the merged bag: the snapshot is +updated in place before `AssembleContext` runs, so a worker reading +`context.metadata` sees what it was dispatched with rather than the row's state +before its own claim. + +**No schema change**, again (§1.4). `currentSchemaVersion` is untouched. + +**One spec deviation, unamended and named here rather than silently applied.** +engine-spec.md §11.4 lists a `complete args` line and no `claim args` line, so +`step claim --metadata` is surface the spec does not yet describe. Per +docs/design/amendments.md the spec is not edited from here: this owes a DKT +amendment issue proposing §11.4 gain + +``` +claim args --owner NAME [--ttl D] [--metadata '{…}'] +``` + +alongside the existing `complete args`. The engine behavior is additive and +every existing claimant is byte-identical without the flag, which is why the +implementation does not wait on the amendment — but the amendment is owed. + ## 2. Alternatives considered **A. Separate `completion_metadata` column.** Rejected. It requires a v11 @@ -228,6 +306,13 @@ appears, `--metadata` bags can carry their own provenance keys opaquely, which is precisely what the KV bag is for. Filing it as a deferred question rather than building it is the §2 discipline. +**Corrected by §1.7 (DKT-592):** "both are *worker*-reported" was wrong about +the requested half — what was asked for is the DISPATCHER's own decision, +settled before the work starts. The rejection of a provenance column stands +(the bag still carries its own provenance keys opaquely, and no schema change +was needed), but the sentence that justified it no longer describes the write +path: the requested half now lands at claim, the resolved half at completion. + **E. Refuse when the worker overwrites a definition key.** Rejected. It sounds protective and is actually core having an opinion about which keys matter. The reference instance's `model_resolved` overwriting a declared `model_requested` @@ -340,7 +425,41 @@ A new `scripts/qa/` section following the existing helper conventions 5. Genericity: `grep` the section's own fixtures for the banned words → zero hits, since the gate scans tests too. -### 4.5 Gates that must stay green +### 4.5 Go — the claim-side write (§1.7) + +In `internal/engine/metadata_test.go`, beside the completion and fail suites, +with the same neutral-key discipline: `tier_requested` / `desk_resolved` mirror +the SHAPE of a routing pair — one fact known at dispatch, one known only on +return — without naming any instance's vocabulary. + +- **The regression test:** claim with `--metadata`, assert `steps.metadata` + carries it with **no completion anywhere in the test**. +- **The failure case, which is the whole point:** claim with a bag, `fail` with + no bag of its own → the claim's keys are still on the row. +- **The crash case, which is not the failure case:** nobody calls `fail` at + all. The lease lapses, the next claim reaps it, and the bag survives the reap. +- **Both pairs on a normal completion:** claim's two keys plus completion's two + keys, all four present — the completion merge must not clobber the claim's. +- Definition-side bag present → merged, not clobbered; dispatcher's value wins + on a shared key. +- No `--metadata` → definition bag byte-identical, no row_version bump. +- The claim's own context bundle reports the merged bag. +- A claimed-then-FAILED step's keys appear in the R7 rollup — the read surface + the defect was observed in, with no read-side change. +- Refusals (invalid JSON, array, scalar, number, null) → `VALIDATION_ERROR`, + `row_version` unmoved, **status still `pending` and attempt still 0**, and + the step still claimable. The attempt assertion is the claim-specific half: + this transaction reaps and CASes, so "the refusal wrote nothing" has to mean + "it cost no attempt" too. +- The cap's own test: inclusive boundary in both directions, both numbers in + the message plus this verb's remedy, measured on raw input. +- **Source-position check**, `TestClaimValidatesMetadataBeforeTheTransactionOpens`: + the size and shape calls appear before `conn.Begin()` in + `claimStepWithGates`'s AST — the same argument + `TestFailValidatesMetadataBeforeTheTransactionOpens` makes, because a + rollback hides the difference from any runtime assertion. + +### 4.6 Gates that must stay green Full `scripts/qa.sh` plus `scripts/qa/genericity.sh`. `go build ./...` and `go test ./...`. No migration test changes, because there is no migration — diff --git a/docs/tdd/engine-spine.md b/docs/tdd/engine-spine.md index 11582c68..ec3f8eb5 100644 --- a/docs/tdd/engine-spine.md +++ b/docs/tdd/engine-spine.md @@ -260,7 +260,15 @@ bare-int shorthand (`"write" = 1` ⇒ `{max: 1}`), per §11.1. ## 4.3 Register-time validation table Every row is a `VALIDATION_ERROR` (exit 3) naming the workflow, the step, and the -offending field. Every row is a test case (§4.6). The table is the phase's contract: +offending field. Every row is a test case (§4.6). The table is the phase's contract, +and its rows match the `RuleIDs` slice in `internal/workflow/validate.go` one for one. +V12, the old `on_fail` closed-vocabulary rule, is gone: a value outside that +vocabulary now names a triage panel, and V40 resolves it. + +Most rows are decided by `Validate`, a pure function of the definition's bytes. +V26 runs in `ValidateVoteRules`, and V21a–V21d, V25a, V28a, V29, V30, and V37a run +in `ValidateSchemas`, because they consult the project's registered vote rules and +schemas. V21d records a case that is *not* an error. | # | Rule | Spec line | |---|---|---| @@ -275,22 +283,57 @@ offending field. Every row is a test case (§4.6). The table is the phase's cont | V9 | every `after` entry names a step in this workflow | §11.1 `after`: "intra-workflow predecessors" | | V10 | `after = []` is legal and means root; a **missing** `after` on a non-exempt step is the error | §11.1: "`[]` = root (implicit topology was a footgun)" | | V11 | `inputs` entries match `.` \| `.*` \| `issue.body` \| `issue.diff`; the named step exists; a `.` names a kind that step actually produces (§4.3.1) | §11.1 `inputs` | -| V12 | `on_fail` ∈ {`fix-loop`, `waiting-human`, `skip`, `abandon-issue`} | §11.1 `on_fail`; engine-core §4 "closed vocabulary" | +| V11a | a step may not produce an artifact of the reserved kind `gate-results`; `.gate-results` is the engine-served input form for recorded gate results, so such an artifact could never be addressed | §11.1 `inputs`; §4.3.1 | +| V11b | a step may not produce an artifact of the reserved kind `vote-record`; `.vote-record` is the engine-served input form for a vote step's recorded proposal | §11.1 `inputs` (`.vote-record`); §4.3.1 | | V13 | **a `type="human"` step's reject routing may not be `waiting-human`** — evaluated against the **effective** routing (declared `on_fail`, else the §11.1 default), so the default cannot smuggle the deadlock back in (§4.3.2) | §2: "a human gate's reject routing may not itself be `waiting-human` (register-time VALIDATION_ERROR)"; §11.1 `on_fail` (amended 2026-08-03) | -| V13a | **`on_fail` is required, explicitly, on `type="human"` steps** — the corollary of V13 over the effective value; the error names the step and the three legal values | §11.1 `on_fail`: "`type="human"` steps must declare it explicitly and `"waiting-human"` is invalid there" (amended 2026-08-03) | +| V13a | **`on_fail` is required, explicitly, on `type="human"` and `type="vote"` steps** — the corollary of V13 over the effective value; the error names the step and the legal values for its type (`waiting-human` is legal on a vote step, where it escalates to an operator) | §11.1 `on_fail`: "`type="human"` steps must declare it explicitly and `"waiting-human"` is invalid there" (amended 2026-08-03) | | V14 | `voters` and `vote_rule` required on `type="vote"`, forbidden elsewhere | §11.1 `voters`, `vote_rule` | | V15 | `fanout` non-empty when present | §11.1 `fanout` | | V16 | `min_siblings` ≥ 1 and ≤ `len(fanout)`; only on fanout steps | §11.1 `min_siblings`; §2 fanout joins | | V17 | `after_loop` names an existing step; only meaningful with `loop = true` in the workflow | §11.3 | | V17b | a step that can route `fix-loop` (via `on_fail` or `threshold`) requires a `loop = true` step in the workflow (DKT-196) | §11.3 | +| V17c | when every `loop = true` body declares `serves`, a step that can route `fix-loop` (via `on_fail`, `threshold`, or a triage panel's `fix-round`) must appear in at least one body's `serves`; a body without `serves` serves every trigger | §11.3 "every step that can route `fix-loop` must be served by at least one body" | | V18 | `loop = true` steps have no `after` (their ordering comes from loop entry) | §11.1 `after`; §11.3 (3) | -| V19 | `max_attempts` ≥ 1; `max_fix_loops` ≥ 0; `expected_cost` ≥ 0 | §11.1 | +| V19 | `max_attempts` ≥ 1; `max_fix_loops` ≥ 0; `max_stalled_rounds` ≥ 0; `expected_cost` ≥ 0 | §11.1 | | V20 | `threshold` keys ∈ {`fix-loop`, `waiting-human`, `pass`} ∪ step names in this workflow | §11.2 | | V21 | `threshold` predicate parses as `agg(field op literal)`, `agg ∈ {any, all, count>=n}`, `op ∈ {==, !=, >=, >, <=, <}` | §11.2 | -| V22 | `when` parses as a predicate over `kind`/`labels` only | §11.1 `when`; engine-core §4 "conditions (predicates over issue kind/labels only)" | +| V21a | each `threshold` predicate's field is a top-level property of the item schema the step's `payload` names; the reserved `diff.*` fields are exempt (V45–V47) | payloads-thresholds §4.9.1 | +| V21b | each predicate's literal is a member of the field's `enum` when it declares one, and otherwise parses as its declared `type` (`number`, `integer`, `boolean`) | payloads-thresholds §4.9.1 | +| V21c | an ordered operator (`>=`, `>`, `<=`, `<`) requires the field to declare `ordered_enum` | payloads-thresholds §4.9.1 | +| V21d | a step with a `threshold` and no `payload` gets V21's grammar check only: no field check, and T3 at runtime. Not an error; no code path emits it | payloads-thresholds §4.9.1 | +| V22 | `when` parses as a predicate over `kind`/`labels` only, its clauses joined by `and` throughout or `or` throughout — a mix of the two is refused (DKT-548). A clause is ` <==\|!=\|contains> ` or the set form `labels contains-any (a, b, c)`, whose list must be non-empty, comma-separated, and free of whitespace inside its values; `contains-any` is `labels`-only (DKT-550). The set operator is equivalently spelled `contains_any` and its list equivalently delimited `[a, b, c]`, with the delimiters required to pair (DKT-1000) | §11.1 `when`; engine-core §4 "conditions (predicates over issue kind/labels only)" | | V23 | `class` defaults to the `executor` value when unset | §11.1 `class`: "default = executor value" | | V24 | `[limits]` values: `max` ≥ 1, `lease_ttl`/`max_step_duration` parse as durations | §11.1 `[limits]` | | V25 | `payload` matches `name@version` shape (**shape only** at S3 — §6.14) | §11.1 `payload` | +| V25a | `payload` names a registered `name@version` that is not retired | payloads-thresholds §4.9.1 | +| V26 | a `type="vote"` step's `vote_rule` names a registered threshold configuration; the error lists the registered rules and, when other projects configure the rule, gives the `--global` remedy | gates-trust §8.2 | +| V27 | a step `name` may not end in the reserved `-held` suffix, and `action` may not name a reserved builtin this build does not implement | payloads-thresholds §7.1 | +| V28 | `action = "aggregate"` requires `params.field`, `params.method` ∈ {`median`, `max`, `min`}, and `params.output`; `hold_spread` is an integer ≥ 0, and `route_at` and `source_field` are non-empty strings, when present; no other `params` keys | payloads-thresholds §7.1 | +| V28a | an `aggregate` step's `params.route_at` names a value in the order its schema declares for `params.field` | payloads-thresholds §7.1 | +| V29 | an `aggregate` step declares `payload`, and that schema declares `params.field` as `ordered_enum` | payloads-thresholds §7.1 | +| V30 | an `aggregate` step's schema accepts a probe output document built from the schema's own declared values | payloads-thresholds §7.1 | +| V31 | `action = "aggregate"` requires a non-empty `inputs` | §2 `aggregate`: its input is the step's declared `inputs` artifacts | +| V32 | each `packet` entry is non-empty, relative (no leading `/` or `~`), and has no `..` segment, with backslashes read as `/`; shape only, since a file's existence is checked at activation | packet-composition §1.6 | +| V33 | a `packet` entry with the `{executor}` token requires the step to declare `executor` or `fanout` | packet-composition §1.6 | +| V34 | a step may not be named `issue.latest` or `issue.linked`, or start with `issue.latest.` or `issue.linked.`; those are engine-served input forms | §11.1 `inputs` | +| V35 | `serves` is valid only on `loop = true` steps; each entry is non-empty and names a step of this workflow that can route `fix-loop` | §11.3 "`serves` is valid only on `loop = true` steps" | +| V36 | on a `type="vote"` step, each `threshold` routing is `fix-loop`, `waiting-human`, or `pass` (no step names), the operator is `==` or `!=`, and the field is `vote`, `verdict`, or `voter` | §11.2 | +| V37 | `pass_floor` declares both `field` and `at`, and the step declares `payload` | §11.1 `pass_floor` | +| V37a | `pass_floor.field` is declared as `ordered_enum` in the step's schema, and `pass_floor.at` is a value of that order | §11.1 `pass_floor`; payloads-thresholds §4.9.1 | +| V38 | a positive `max_stalled_rounds` requires a step that can route `fix-loop` and records an artifact (not a `type` step) | §11.1 `max_stalled_rounds` | +| V39 | every `after_fired` entry names a step in this workflow | §11.1 `after_fired` | +| V39a | every `after_fired` entry also appears in the step's `after` | §11.1 `after_fired`: "every entry must also appear in `after`" | +| V40 | an `on_fail` outside the closed vocabulary names a triage panel: the routing step must be an executor, and the name must resolve to a `type="vote"` step whose `after` includes the routing step | §11.1 `on_fail` | +| V40a | `on_fail_routes` is valid only on `type="vote"` steps, and its keys are `approved` or `rejected` | §11.1 `on_fail_routes` | +| V40b | each `on_fail_routes` value is `retry`, `fix-round`, `abandon-issue`, or `waiting-human`; `fix-round` requires a `loop = true` body serving every step that routes to the panel | §11.1 `on_fail_routes` | +| V40c | a vote step named by another step's `on_fail` must declare `on_fail_routes`, and its own `on_fail` may not be `skip` | §11.1 `on_fail_routes` | +| V41 | `on_exhausted` is valid only on a step that can route `fix-loop` and declares a positive `max_fix_loops`; a value other than `waiting-human` or `abandon-issue` names a `type="vote"` or executor step whose `after` includes this step | §11.1 `on_exhausted` | +| V42 | every `[match].sizes_any` entry is a valid issue size: `trivial`, `small`, `bounded`, `needs-design`, or `unknown`. A definition-level rule, so the error names the field and no step | §11.1 `[match]`; reliability-delta "AMENDMENT — the span extends to v34" | +| V43 | `roster`, `weighting`, and `recuse` are valid only on `type="vote"` steps, with values in {`open`, `strict`}, {`declared`, `equal`}, and {`none`, `executor`} | §11.1 `roster`, `weighting`, `recuse` | +| V44 | `reviews` is valid only on `type="vote"` steps and names another step of this workflow, not the vote step itself | §11.1 `reviews` | +| V45 | a `threshold` predicate over a reserved `diff.*` field requires a tree-holding step: an executor or fanout step that does not declare `holds_tree = false` | §11.2; payloads-thresholds §5.1 | +| V46 | a predicate over `diff.lines` or `diff.files` compares against an integer literal, under any operator | §11.2; payloads-thresholds §5.1 | +| V47 | a predicate over `diff.empty` uses `==` or `!=` and compares against a boolean literal | §11.2; payloads-thresholds §5.1 | **V5 and V8 are the two the spec calls out by name**, and they are the two that make a workflow's topology unambiguous. V13/V13a are called out because they are the @@ -370,8 +413,17 @@ verbs): | `docket step approve STEP-N [--note …]` | step → `done`, `step-approved` event; downstream `after` successors become ready by the ordinary §6.3 predicate. An approve records **no artifact** (§4.3.1: human steps produce none) | | `docket step reject STEP-N [--note …]` | step → routed per the step's **effective `on_fail`** — which V13/V13a guarantee is one of `fix-loop`, `skip`, `abandon-issue`, never `waiting-human`; `step-rejected` then `step-routed` events, in the one routing transaction of §6.8 stage N+1 | -Both refuse a non-`human` step with `VALIDATION_ERROR` (§6.9 R10) and both are -token-free (§6.10): a human gate is resolved by an operator, who never claimed it. +Both refuse a non-`human` step with `VALIDATION_ERROR` (§6.9 R10) and neither +takes a lease token (§6.10): a human gate is resolved by an operator, who never +claimed it. Both — with `step resolve`, `step reap`, and the run lifecycle +verbs — require the RUN'S CONDUCTOR CAPABILITY instead (DKT-2465, +reliability-delta §2's v29 amendment): the token the run's first activation +returned once, or `run conduct` re-minted, presented via `DOCKET_TOKEN`/stdin; +none is the R1 `VALIDATION_ERROR`, a wrong one the R3 `AUTH_ERROR`, and a run +activated before the capability existed asks for none. The event records who +(DKT-2450): `step-approved`, `step-rejected`, `step-resolved`, and a forced +`lease-reaped` carry `actor` and `cwd` beside the note, exactly as trust events +do — see runs-dispatch §8.7. ### 4.3.3 Register-time DAG lints @@ -407,7 +459,7 @@ one classification feeds the lint, expansion, and the engine's readiness latch | Verb | Flags | Effect | |---|---|---| | `docket workflow register ` | `--json[=v2]` | parse + validate + lint; insert `name@version`; idempotent on identical bytes; `CONFLICT` on differing bytes at an existing `name@version` | -| `docket workflow list` | `--name`, `--limit`, `--json[=v2]` | registered workflows; a `Collection` (reliability-delta §4.1) so v2 renders `{items,total,truncated}` | +| `docket workflow list` | `--name`, `--limit`, `--orphans`, `--json[=v2]` | registered workflows; a `Collection` (reliability-delta §4.1) so v2 renders `{items,total,truncated}`; `--orphans` narrows to registrations whose NAME no file in any instance-config root declares any more (DKT-609), stamping each row with an `origin` verdict and refusing outright when there is no root to scan | | `docket workflow show [@]` | `--source`, `--json[=v2]` | the parsed definition; `@version` omitted ⇒ highest registered; `--source` emits the stored TOML verbatim | | `docket workflow init` | `--template NAME`, `--dir PATH`, `--force` | writes template files into `.docket/config/` (default), refusing to overwrite without `--force` | @@ -444,7 +496,7 @@ lifecycle. It creates the directory tree if absent and never overwrites without | Situation | Code | Exit | |---|---|---| -| grammar/validation/lint failure (V1–V25 incl. V13a, L1–L4) | `VALIDATION_ERROR` | 3 | +| grammar/validation/lint failure (every rule in `RuleIDs`, §4.3; L1–L4) | `VALIDATION_ERROR` | 3 | | file unreadable / not found | `NOT_FOUND` | 2 | | re-register differing bytes at existing `name@version` | `CONFLICT` | 4 | | `workflow show` on an unregistered name/version | `NOT_FOUND` | 2 | @@ -465,10 +517,9 @@ in the style of `internal/planner/plan_test.go`: `TestValidationTableIsComplete`, which asserts one test case per documented rule ID. (A validation table that drifts from its tests is a table that documents behavior the code does not have.) **The rule IDs are the authority, not a count**: - the test enumerates `V1…V25` *including* `V13a`, i.e. **26 rules across 25 numbered - IDs**, and asserts set equality between the documented IDs and the test cases' - IDs — never `len(cases) == 25`, which is exactly the assertion that breaks when a - rule is split. V13 and V13a are separate cases: V13 feeds a human step with an + the test reads every rule in `RuleIDs`, including split rules such as `V13a`, and + asserts set equality between those IDs and the test cases' IDs — never a fixed + `len(cases)`, which is exactly the assertion that breaks when a rule is split. V13 and V13a are separate cases: V13 feeds a human step with an explicit `on_fail = "waiting-human"`, V13a feeds one that declares no `on_fail`, and each asserts its own distinct message (§4.3.2). - **The lint table L1–L4**, including L2's two exceptions (a threshold-interposed @@ -612,6 +663,60 @@ mid-run edit immunity; freezing the scheduler would ignore a correction that exi precisely to prevent a collision. Both are stated so neither is "fixed" into the other later. +**No AUTOMATIC path refreshes the snapshot** (DKT-741). It is written once, at +activation stage 4, and nothing rewrites it for the life of the run — not a claim, not +a fix-round, not a re-instantiation, and not an `issue edit`. So the consumers that +read it — the packet's `context.issue.scope` (§6.6) and the recorded `issue.diff` +scope (§6.7.1 D1) — cannot drift apart from each other or from what the run was +activated on. Re-snapshotting at claim time instead would break exactly that: two +steps of one run would render two different declared scopes and record their diffs +over two different path sets, and a packet would stop being reproducible from the +ledger. `docket issue edit --scope` therefore reaches the live column and nothing +else, which is correct and is also a trap: + +| what the operator wants | what `issue edit --scope` does | +|---|---| +| stop a collision the scheduler is about to allow | works, immediately — R4 reads live | +| widen an authorized scope so a live step's packet says so | **does nothing**; the packet renders the frozen snapshot | + +The second row has **two** dispositions, and which one is right depends on whether the +run's premise changed or one declaration was corrected. + +**Where the premise changed**, take the issue out of the run and re-plan it — +`docket run abandon RUN-N --issue DKT-M --reason "scope widened"`, then plan it into a +new run, whose activation snapshots the widened scope afresh. It is expensive on +purpose: a mid-run scope widen can invalidate the premise every step of that issue +already executed under, and re-planning is what re-establishes it. + +**Where the premise is intact**, `docket run refresh-scope RUN-N --issue DKT-M +--reason R` copies `issues.scope_globs` into that one run-issue's snapshot and +rewrites nothing else in it (DKT-869, `RefreshIssueScopeInRun`). DKT-741 had ruled out +any refresh verb; RUN-52 (VPL-434) then charged twice for that ruling on an intact +premise — the panel rejected work as out of scope, the operator agreed and widened it, +the already-minted `fix@2` step still rendered the old scope, and the issue was +abandoned mid-loop. The freeze keeps its default and gains an explicit exception whose +four properties are what keep it from being a hole in §9 item 5: + +1. **It carries no scope of its own.** There is no `--scope` on it; `issue create|edit + --scope` stays the sole writer of the column it copies, so the refresh cannot make + real a scope that was not declared through the one gate widening has always had. A + refresh with no widen behind it is **refused** (CONFLICT), not silently no-op'd. +2. **No step straddles it.** It refuses while any of the issue's steps is `claimed`, + `running`, or `gated`, and while a dispatch is open — the repin quiescence rule + (DKT-408) applied to the other frozen premise. `pending` and `waiting-human` are the + refreshable states. +3. **It rewrites no history.** Terminal steps keep their artifacts and the scope their + diffs were computed over; only the remaining steps' renders move. +4. **The discontinuity is in the ledger.** One `issue-scope-refreshed` event (actor + `human`) carries the old scope, the new scope, the instances reached, and the + operator's reason — so two steps of one run declaring two different scopes is a + dated, attributable fact rather than drift a reader must infer. + +`issue edit --scope` **warns**, naming the run, the frozen scope, the count of live +steps, and **both** verbs, whenever the edit changes the scope of an issue that still +has non-terminal steps in a non-terminal run (`ScopeEditFrozenForActiveRuns`). It +reports rather than refuses, because the write is real for the scheduling half. + ## 5.2 Run status and the minimal subset engine-core §1.1: `planning → active ⇄ waiting-human → done | abandoned`. All five @@ -624,7 +729,7 @@ statuses exist at S3 because the step lifecycle routes into `waiting-human` and | `docket run activate RUN-N` | the fat transaction (§5.3) | | `docket run pause\|resume RUN-N [--reason R]` | `active ⇄ waiting-human`; pause blocks new claims, honors in-flight completes (engine-core §3.4) | | `docket run abandon RUN-N --reason R` | terminal; revokes live leases | -| `docket run status [RUN-N] [--active]` | read-only; effective status computed at read | +| `docket run status [RUN-N] [--all]` | read-only; effective status computed at read; the list hides done and abandoned runs unless `--all` | `--budget N` on `run start` is **accepted and stored** but enforces nothing until S6. Accepting it now means the S6 upgrade adds enforcement, not a flag — and a flag @@ -715,13 +820,15 @@ walks a map without sorting — which is the only realistic way this property br | Rule | Behavior | |---|---| | RA1 | re-activating an `active` run lints again and expands **only** issues with `expanded_at_ms IS NULL` whose predecessors are now satisfied | -| RA2 | the pin set is **inherited, never recomputed** — a re-activation after a workflow re-register or a pinned file edit uses the original `pins` rows, unchanged | +| RA2 | existing pins are **inherited, never recomputed** — a re-activation after a workflow re-register or a pinned file edit uses the original `pins` rows, unchanged. Re-activation can **add** a pin for a ref the set never held: a `policy.toml` the config now carries is pinned, and rows still waiting then resolve against it (runs-dispatch §9.4, F15) | | RA3 | new issues added to the run since activation are bound and snapshotted at re-activation, and pinned against the **already-pinned** workflow version if one exists for that name | | RA4 | refused while a dispatch is open — **vacuously true at S3** (dispatches are S6); the check is written as a seam that queries a not-yet-existing table via a helper returning `false`, so S6 adds a query, not a call site | | RA5 | re-activating a `done` or `abandoned` run is `CONFLICT` (exit 4) | RA2 is the reproducibility guarantee. If re-activation re-pinned, an in-flight run would silently adopt an edited workflow — precisely what engine-core §4 forbids. +Adding a pin the set never held does not break that guarantee for anything already +pinned, but it does change what waiting rows resolve against, so it is not a no-op. ## 5.5 Error taxonomy usage (phase 2) @@ -792,6 +899,14 @@ including `unless_labels` beating `labels_any`; absent clauses matching anything **Go unit tests** (`internal/engine/activate_test.go`): - exactly-one-match: zero matches and two matches each `VALIDATION_ERROR`, each naming the issue **and** every candidate workflow (asserted by substring). +- **orphan annotation** (DKT-609, `internal/engine/dkt609_test.go`): each named + candidate whose NAME no file in any instance-config root declares any more is + marked `(no source on disk — orphaned registration, deprecation candidate)`, + and the refusal carries the remedy. It DECORATES the candidate set and never + changes it — an orphaned registration still binds, because a registration is + a row and not a file. With no root to scan the verdict is `unchecked` and the + message is byte-identical to the pre-DKT-609 one: "nothing was checked" must + never render as "nothing is orphaned". - **bind-to-highest** (§11.1 as amended 2026-08-05, DKT-40): the candidate set is the **highest registered version of each name**, so exactly-one-match applies across NAMES. `TestBindingUsesHighestVersionOfEachName` is DKT-8's M2a wedge as @@ -953,6 +1068,7 @@ engine-core §5 and §1.3, as a conjunction. A step is ready iff **all** hold: |---|---|---| | R1 | the run is `active` | §1.3; §2 scheduling | | R2 | the issue's `depends_on` predecessors are satisfied | §1.3 "its issue's dependencies are satisfied" | +| R2b | no step of the issue is parked `waiting-human` — a park holds its own issue only; the run stays `active` while any other issue has unfinished steps, and rolls up to `waiting-human` once no unparked work remains (`reconcileRun`, reconcile.go) | RUN-90: eleven single-issue parks each stopped all 45 issues | | R3 | its intra-workflow predecessors (`after`) are **done** — and for a fanned-out predecessor, **joined** (§7.4) | §1.3; §2 fanout joins | | R4 | its scope conflicts with no `claimed`/`running` step (glob intersection) | §1.3; §5 mutual exclusion | | R5 | per-class concurrency headroom exists | §2 "concurrency headroom per executor-hint class" | @@ -964,6 +1080,16 @@ issue's; age is the step's `created_at_ms`, tie-broken by `id` so the order is t and reproducible. `sortIssues`'s existing `priorityRank` is reused for the priority half — same ranking, no second definition to drift. +**R3 carries two interposition clauses** for §11.2's interposed gate. First, a step +named as a `threshold` step-name target is latched until a routing predecessor's +recorded routing names it; a routing that resolves elsewhere terminalizes it +`skipped` (DKT-38). Second, a step's ordinary downstream is held by any of its +`after` predecessors' threshold targets that is itself a `type=vote` or `type=human` +gate and is still open (not yet terminal) — those are decisions the downstream's own +validity depends on. An open interposed executor target (a step with no `type`) does +**not** hold the downstream: it is ordinary work the routing chose to add, and no +downstream reads its output. + **`--limit`** applies after ordering, with the v2 truncation contract (reliability-delta §4.2/§5): `readyTotal` before slicing, `truncated` computed, negative limit `VALIDATION_ERROR` under v2. `next --run` is a *new* verb surface, so @@ -1085,8 +1211,10 @@ in `internal/db/leases.go` are **generalized over a table name**, not copied: a `ClaimIssue`/`HeartbeatIssue`/`ReleaseIssue` become thin wrappers over the shared implementation. `authorizeHolder`'s three-way refusal (ErrNotHolder / ErrLeaseExpired / ErrLeaseHeld) is reused unchanged — which is why the S3 refusal matrix (§6.9) is -the S2 matrix with step verbs substituted, and why it cannot drift between the two -entities. +the S2 matrix with step verbs substituted, and why the rows those helpers decide +cannot drift between the two entities. R5 is the one step-level divergence: a step +claim whose `--owner` matches a live lease's recorded owner re-mints the token +instead of returning `CONFLICT` (§6.9). Issue claims have no re-mint path. The claim response returns token **and** context in one response — "one atomic mediation: an unclaimed executor has nothing, a claimed one has everything" @@ -1116,6 +1244,21 @@ same property at the CLI level. `claim --render` returns the assembled packet instead, atomically (§2). +**A read-back replays what the claim recorded** (DKT-1054). Source 4 is resolved +over run state, and run state keeps moving after a step is handed out: the +fixture's `fix@1` binds `reconcile@0` at claim, then `review@1`, `synthesize@1`, +and `reconcile@1` complete at fix@1's own ordinal, and a live re-resolution binds +`reconcile@1` — an artifact produced by reviewing fix@1's diff. The claim writes +the bindings it handed over to `step_inputs` (§6.1) in its own transaction, and +`step context`, `step render`, and `step show`'s target ref read a claimed step +(`attempt > 0`, not back at `pending`) over exactly that set, through the same +resolver, so the read-back is the claim-time bundle however far the run has moved. +A re-claim (a reaped lease, `resolve --as retry`) records its own bindings in place +of the last attempt's. A step not yet handed out — every `action`, `human`, and +`vote` step, and an executor step still `pending` — reads live, since the claim +that will hand it out is what a read of it previews; `step context --live` asks +that question of any step. + ## 6.7 Input resolution §2, verbatim: "Downstream `inputs` resolve over siblings that RECORDED their work @@ -1174,6 +1317,27 @@ This is what makes the fixture's `fix@1 → review@1` cycle correct: `review@1` `issue.diff` to the artifact `fix@1` produced, not the one `implement@0` produced, because ordinal 1 beats ordinal 0 under D3 — without any rule specific to loops. +**The round record.** At a loop re-entry the artifact also carries a small JSON +payload: the hand-back `head`, the declared `worktree`, and `round_base` — the +commit the packet's appended round-delta section diffs from, "this round's work +alone". `round_base` derives from **the head of the newest `issue.diff` a done +step CONSUMED without recording one of its own** — a step that read the tree +rather than wrote it, which is a review — and not from the newest recorded head. +The distinction matters only when a fix round goes unjudged, and then it decides +whether anyone ever reads it: basing the delta on the previous fix round's own +commit puts that round INSIDE the base, so the next panel is handed just what the +latest round moved and its delta clause scopes it to a change no judge has seen. +RUN-14/HRN-27 lost a whole +1258/-550 round that way, with the following round's +74 lines rendered as the entire object under judgment. Reaching back to the last +head a review actually judged keeps every unreviewed fix round inside the next +reviewed delta. With no such consumer yet — a fix round minted before any review +recorded — the base falls back to the newest recorded head, so the round still +renders a delta. + +Reviewers are identified by that consumed-but-produced-nothing pair rather than by +a class or step name, for §6.5's genericity reason: core attaches no meaning to +class names, so a filter keyed on one would be instance policy living in core. + ## 6.8 The complete saga §2, verbatim, is the specification: @@ -1281,8 +1445,9 @@ per `on_fail` when attempts are exhausted, per the status machine. ## 6.9 Refusal matrix (§9 item 3, at step level) -The S2 matrix with step verbs substituted — same helper, same codes, so it cannot -drift: +The S2 matrix with step verbs substituted, using the same helpers and codes. R5 +diverges from S2: a step claim by the live lease's own owner re-mints the token, +where an issue claim always returns `CONFLICT`. | # | Situation | Verb | Code | Exit | |---|---|---|---|---| @@ -1290,7 +1455,7 @@ drift: | R2 | token supplied, step unclaimed | heartbeat/complete/fail | `AUTH_ERROR` | 5 | | R3 | token supplied, wrong value | heartbeat/complete/fail | `AUTH_ERROR` | 5 | | R4 | correct token, lease expired | heartbeat/complete/fail | `STALE_LEASE` | 6 | -| R5 | claim against a live lease | claim | `CONFLICT` | 4 | +| R5 | claim against a live lease | claim | `CONFLICT` for any other `--owner`; **succeeds** with a re-minted token when `--owner` equals the recorded owner | 4; 0 | | R6 | N concurrent claims on one ready step | claim | 1 × exit 0, N−1 × `CONFLICT` | 4 | | R7 | claim against an expired lease | claim | **succeeds**, `attempt++` | 0 | | R8 | claim a step that is not ready (R1–R7 of §6.3 unmet) | claim | `CONFLICT` naming the unmet condition | 4 | @@ -1299,6 +1464,22 @@ drift: | R11 | `resolve` on a step not in `waiting-human` | resolve | `VALIDATION_ERROR` | 3 | | R12 | artifact exceeding 1MiB | complete | `VALIDATION_ERROR` | 3 | +**R5 trusts `--owner` as the holder identity.** A claim against a live lease whose +`--owner` equals the lease's recorded owner exits 0 with `re_minted`. The engine +mints a fresh token that replaces the prior one, so the prior token stops +authorizing. `attempt`, accrual and expiry are unchanged. The re-mint recovers a +token whose claim response never reached its holder. A claim with any other owner +returns `CONFLICT` (exit 4), as at S2. While the claim's pre-gate results are +still unrecorded, a same-owner claim also returns `CONFLICT`. + +The owner match is the only check. `--owner` is caller-supplied text, `step show` +prints it, and dispatcher conventions such as `wave:STEP-N:attempt` make it +predictable, so it is not a secret. Any caller that presents the owner string +receives a working token and voids the holder's. A dispatcher must therefore +give each concurrent claimant a distinct owner. Claimants that share one become +a single holder: they get re-minted tokens instead of R6's `CONFLICT`, and only +the last token issued can record. + **R8 is new relative to S2** and it matters: `claim` must enforce readiness itself, not trust that the caller ran `next`. A dispatcher racing a scope conflict would otherwise claim a step `next` would never have offered. The refusal names which @@ -1316,8 +1497,8 @@ recording a duplicate artifact. | `docket step heartbeat STEP-N` | yes | extends the lease; does not touch `attempt` | | `docket step complete STEP-N --artifact-file F [--payload-file F] [--usage '{…}'] [--metadata '{…}']` | yes (stage 0–1) | the saga | | `docket step fail STEP-N [--note …] [--metadata '{…}']` | yes | records the failure; routes per `on_fail` when attempts are exhausted (the counter is bumped by CLAIMS, not by `fail` — E-8) | -| `docket step approve\|reject STEP-N [--note …]` | no | `type=human` gate steps **only** (§2) | -| `docket step resolve STEP-N --as retry\|skip\|abandon-issue\|override-pass [--note …]` | no | `waiting-human` resolutions (§2); `retry` **resets attempts** | +| `docket step approve\|reject STEP-N [--note …]` | conductor (DKT-2465) | `type=human` gate steps **only** (§2); no lease token, but the RUN's conductor capability — see §4.3.2 | +| `docket step resolve STEP-N --as retry\|skip\|abandon-issue\|override-pass [--note …]` | conductor (DKT-2465) | `waiting-human` resolutions (§2); `retry` **resets attempts** | | `docket step show STEP-N` | no | read-only; effective status | | `docket step list (--run RUN-N \| --issue ISSUE-N)` | no | read-only; steps with id, run, instance, issue, kind, effective status, attempt, expected_cost (DKT-54: step ids are a store-wide sequence, so nothing else enumerates a run). `--issue` lists one issue's steps across every run holding one, since a re-activation mints a fresh round under a new run (DKT-244) | | `docket step context STEP-N [--meta]` | no | re-emits `context` read-only (§11.4) | @@ -1422,7 +1603,7 @@ engine state: **exit 0 allow / exit 2 deny with reason** (§2). | Guard | Allows when | |---|---| | `docket guard stop` | no pending work outside `waiting-human` — i.e. no step in `pending`/`ready`/`claimed`/`running`/`gated` for any active run | -| `docket guard gate --step NAME` | an **approved** `type=human` step of that name exists for the active run | +| `docket guard gate --step NAME [--run RUN-N]` | an **approved** `type=human` or `type=vote` step of that name exists — in the named run, or with no `--run` in **any** active run of the project | Exit 2 collides numerically with `NOT_FOUND`'s exit 2, and that is **intentional and specified**: §2 defines the guard contract as "exit 0/2 + reason", independent of @@ -1432,6 +1613,21 @@ into the JSON envelope's `error` under `--json`. This is recorded here because a reviewer will otherwise read it as a taxonomy violation; it is the spec's own contract, quoted. +**`gate`'s scope is the caller's choice, and the two forms ask different +questions.** `--run RUN-N` asks about ONE run's gate: an approval is a decision +about one run's change, so another run's approval says nothing about it. That is +the form for a caller that knows its run — an executor's brief carries its step id, +and `step show` resolves the run from it. The named run is honored regardless of +project, a run that does not exist is `NOT_FOUND` rather than a verdict, and a run +that has ended denies. Without `--run` the guard answers over **every** active run +of the project and the first approved gate of that name allows. That reading is +cross-run by construction: one run's approval opens the gate for every caller in +the project until that run finishes, a second run's undecided gate included. It +exists for callers with no run context — an operator session's commit hook cannot +know which run a git write belongs to — and a hook that has a run should scope. +Both gate kinds answer in both forms: a tallied vote approval is a decision as much +as a human approval is. + `guard spawn|record` are **not** here — §10 assigns them to stage 6. ## 6.13 The action seam (S5 boundary), specified diff --git a/docs/tdd/events-follow.md b/docs/tdd/events-follow.md index 2fe2c01d..d95cb226 100644 --- a/docs/tdd/events-follow.md +++ b/docs/tdd/events-follow.md @@ -114,7 +114,7 @@ migration at all, and `TestSchemaVersionIsUnchangedAtS7` asserts | # | Clause | |---|---| -| D1 | `events prune` is a verb nobody is obliged to run. Docket still deletes nothing on its own — there is no automatic retention sweep, no prune inside `next`, no compaction at `run done`. The retention *config key* exists and defaults to **0, meaning "retain everything"**, which is the posture operations.md §2 already documents | +| D1 | `events prune` is a verb nobody is obliged to run. Docket still deletes nothing on its own — there is no automatic retention sweep, no prune inside `next`, no compaction at `run done`. The retention *config key* exists and defaults to **0, which imposes no retention window**: prune at 0 is bounded only by `--before` or `--before-run` and the live-run refusal. Docket deletes nothing it was not asked to delete, the posture operations.md §2 documents | | D2 | `--follow` is a flag on a read verb. Without it, `events list` is byte-identical to S6's, which `TestEventsListIsUnchangedByFollow` asserts by diffing the same call before and after | | D3 | `run budget --set` is a new sub-verb. `run start --budget` is untouched, and a run nobody re-caps has the same `budget` column value it always had | | D4 | The dormancy sweep runs against **engine-s6** and must show ZERO diffs on every existing verb's output — the standing check each stage has carried | @@ -197,7 +197,7 @@ docket events prune --before-run RUN-N # everything belonging to that run | # | Clause | |---|---| -| P8 | An event belonging to a run whose status is not `done` or `abandoned` is **never deleted**. `model.RunStatus.Terminal()` is the predicate — the same one `run status --active` and re-activation already use, so "terminal" has one definition | +| P8 | An event belonging to a run whose status is not `done` or `abandoned` is **never deleted**. `model.RunStatus.Terminal()` is the predicate — the same one the default `run status` list and re-activation already use, so "terminal" has one definition | | P9 | The refusal is a **CONFLICT (exit 4) naming the runs**, not a silent skip. A prune that quietly retained half its range would leave an operator believing space was reclaimed and a consumer believing a boundary moved | | P10 | It is evaluated **inside the delete's transaction** (F3), so the set refused and the set deleted are computed over one snapshot | | P11 | Events with **no run** — trust grants — are prunable by `--before`, because there is no run whose liveness could forbid it. They are the one class `--before-run` can never reach, and the help says so | @@ -219,7 +219,7 @@ refusal is protecting arithmetic, not sentiment. | # | Clause | |---|---| | P12 | The boundary is a **config key: `events.retain`** — a duration. Events younger than it are never pruned, whatever `--before` says | -| P13 | Default **`0`, meaning retain everything**, which makes prune a verb that refuses everything until an operator states a policy. That is the dormant posture D1 requires and the one operations.md §2 documents | +| P13 | Default **`0`, which imposes no retention window**, so prune at 0 is bounded only by `--before` or `--before-run` and the live-run refusal, not by age. Nothing is deleted until an operator runs prune with a target. That is the dormant posture D1 requires and the one operations.md §2 documents | | P14 | A `--before` that would cross the boundary is **clamped and reported**, not silently truncated: the answer names how many rows the boundary held back | | P15 | `--before-run` on a terminal run is **not** clamped by the boundary. A run that is done and whose artifacts an operator is discarding wholesale is the case §3's boundary is not about | diff --git a/docs/tdd/gates-trust.md b/docs/tdd/gates-trust.md index fea02d4b..ad64ac8e 100644 --- a/docs/tdd/gates-trust.md +++ b/docs/tdd/gates-trust.md @@ -1,6 +1,6 @@ # TDD: gates, the execution trust model, and the exec runner (stage 4) -Status: draft, revised per security review — 2026-08-03; §3.6 amended per DKT-81 — 2026-08-08 +Status: draft, revised per security review — 2026-08-03; §3.6 amended per DKT-81 — 2026-08-08; §3.6 removal order amended per DKT-2198 — 2026-09-15; §7.4 lock scope narrowed to the shared checkout and §7.6.2 PG5 claim-time pre-gate budget added — 2026-09-25; §7.6.2 PG6 detached pre-gate run added — 2026-10-08; §7.6.2 PG6 readiness hold while a detached run is in flight — 2026-10-09 (docs/tdd/gates-trust-review.md, verdict SOUND WITH FIXES — F1–F5 folded in; see that file's response table for the per-finding mapping) · implements docs/design/engine-spec.md **§4 (whole)** @@ -49,7 +49,7 @@ Out of scope, explicitly, each with the stage that owns it: | Budget enforcement and the floor; `run report`; `dispatch open\|close\|verify\|abandon`; `guard spawn\|record`; `events list --since` | S6 | §10 stage 6 | | The **write-class reap-acknowledgment half of §9 item 10** ("a reaped write-class step cannot gain a successor until the reap is acknowledged") | S6 | it is dispatch mechanics — `guard spawn` surfaces the `reaped` event (§2), and `guard spawn` is stage 6's. DKT-4's own AC text says so: "gate re-run half only". **This stage proves the gate half in full** (§9.2) | | `events --follow`, `events prune` | S7 | §10 stage 7 | -| Worktree-isolated parallel writes (upstream D9) | never core | engine-core §5 names it "an optional instance optimization"; the tree mutex (§7.4) is the always-available baseline | +| Worktree-isolated parallel writes (upstream D9) | never core | engine-core §5 names it "an optional instance optimization"; the tree mutex (§7.4) is the always-available baseline for the shared checkout, and an isolated worktree's gates run beside it unlocked (§7.4 L2) | ### 1.1 Genericity check (CLAUDE.md PR bar, docs/design/genericity.md) @@ -166,7 +166,7 @@ an operator who runs `docket` in it. | T10 | **Trust-entry over-authorization via `--prefix`.** A prefix entry for `make` authorizes `make anything`, including a target that a repo defines to do something hostile. | Prefix entries are **explicit opt-in only** (§3.3): full-argv hashes are the default, `--prefix` is a separate flag that **prints an over-authorization warning naming what it authorizes**, the entry records `prefix = true`, and **a prefix entry matches only when the entry opted in** — a full-argv entry never matches by prefix (§7.2 M3) | §9.1 `TestPrefixMatchingRequiresOptIn` + the warning asserted by substring | | T11 | **Gate result forgery / silent stubbing.** A run shows green gates that never ran. | Every recorded result carries its **real `argv`, `exit`, `duration_ms`, and captured `output`** (§6.1). The S3 `stub: true` field is **absent on every result this stage produces** (§6.2), and the migration marks S3-era trail results so the two are distinguishable forever. An `unmatched` gate records a distinct **`verdict = "unmatched"`**, never `pass` | §9.2 `TestRealResultsCarryNoStubField`; QA asserts `stub` is absent post-migration on new results and present on migrated S3 rows | | T12 | **Double execution of a non-idempotent gate on resume** (§9 item 10). A crash between `gate-started` and the result record; resume re-runs a gate that committed, deployed, or charged something. | §2's at-least-once rule, implemented exactly: a started-but-unrecorded gate re-runs **only if its trust entry is flagged `re-runnable`**, else the step **parks `waiting-human`** (§7.5). The flag is per-**entry**, i.e. the operator's declaration about their own command — core never infers idempotence | §9.2 `TestCrashAtGateBoundaryNeverDoubleRunsNonRerunnable` — every saga boundary, both flag values | -| T13 | **Concurrent tree mutation.** Two gates that touch the working tree (`tree = true`) run concurrently from parallel read-step completions and race a build. | An **engine-held per-repo mutex** (§7.4), which is **not** the database transaction — gates run outside transactions by construction (§6 of engine-spec: "No subprocess ever executes inside a transaction"). The mechanism is pinned in §7.4: an OS-level advisory lock on a lockfile in the repo's `.docket/` directory | §9.2 `TestTreeGatesSerialize` — two concurrent `tree=true` gates, each recording entry/exit timestamps, asserted non-overlapping; and a crash-holding-the-lock case | +| T13 | **Concurrent tree mutation.** Two gates that touch the working tree (`tree = true`) run concurrently from parallel read-step completions and race a build. | An **engine-held mutex on the shared checkout** (§7.4), which is **not** the database transaction — gates run outside transactions by construction (§6 of engine-spec: "No subprocess ever executes inside a transaction"). The mechanism is pinned in §7.4: an OS-level advisory lock on a lockfile in the repo's `.docket/` directory, taken only by a gate running in the shared checkout; a gate in a step's own worktree or a scratch reconstruction has its tree to itself and takes none (L2) | §9.2 `TestTreeGatesSerialize` — two concurrent `tree=true` gates, each recording entry/exit timestamps, asserted non-overlapping; a crash-holding-the-lock case; and `TestIsolatedWorktreeGatesDoNotSerialize` — eight tree gates on eight worktrees run concurrently with no lock wait recorded | | T14 | **Pre-gate as a bypass.** Pre-gates run at claim (§11.1), earlier in the lifecycle and on a different code path than the saga's gates — an implementation could reasonably "simplify" them into a trusted path. | **Same trust model, no exceptions** (§7.6). Pre-gates resolve through the identical matcher, the identical runner, the identical env allowlist and timeout, and an unmatched pre-gate is reported `unmatched` and does **not** execute. The only differences are *when* they run and *where the result goes* | §9.2 `TestPreGatesUseTheSameTrustPath` — an untrusted pre-gate command does not execute and the claim still succeeds with the result recorded `unmatched` | | T15 | **Path resolution hijack.** `argv[0]` is `make`; a repo ships `./make`, or `PATH` contains `.`, so a repo-controlled binary runs. | Resolution is `exec.LookPath` against the **allowlisted `PATH`** (§5.3), which is inherited from the operator's environment and never modified by docket; **the current directory is never prepended** and a relative `argv[0]` containing no separator is **not** resolved against cwd. A trust entry may name an absolute path, which resolves to itself. §5.2 records the residual: an operator whose own `PATH` contains `.` is already exposed everywhere, and docket does not repair that | §9.1 `TestArgv0IsNotResolvedAgainstTheWorkingDirectory` — a `./make` planted in the repo root is not executed | | T16 | **Fenced command harvested from an issue body an attacker can write.** Not a clone — a *live* repo where an attacker can file an issue. | The gate still needs a **matching trust entry** (T1's mechanism), so filing an issue grants nothing. §2 (engine-spec) adds the operator-facing half: activation "surfaces what activation will bind — including every harvested fenced command, verbatim". §7.7 makes that concrete at this stage: `run activate` **prints every harvested command and its trust-match status** (matched / unmatched), so an operator sees `unmatched` commands before the run, not after | §9.2 `TestActivationReportsFenceTrustStatus`; QA asserts the verbatim print | @@ -561,6 +561,33 @@ It is deliberately NOT the pre-existing `gate_results.stub`, which marks a row migrated from an S3 `gate_trail` — that is a fact about which era produced the row, and one column carrying both would answer neither question. +### AMENDMENT (DKT-607) — a stub records its reason + +A stub entry may carry `stub_reason`, set by `docket trust add --stub +--stub-reason ""`. It records the DECISION behind the +placeholder: why no real check exists yet and which issue tracks replacing it +(e.g. `"no scanner selected yet; removal tracked by DKT-607"`). + +**The problem it solves.** DKT-265 made hollow green visible; it did not make it +EXPLAINED. Two tribunal seats on DKT-V196 independently rediscovered the same +corpus stubs (`secret-scan`, `sdet-abuse`) because the decision that they remain +stubs lived only in tribunal transcripts. The project's stub-gate policy +requires every stub to have a removal-tracking issue; this field is where that +reference becomes discoverable from the surfaces an operator actually reads. + +**Where it surfaces.** The activation gate preflight prints it under the stub's +own line (and carries it as `stub_reason` in the JSON row); a stub with NO +recorded reason gets a remedy line naming `--stub-reason`. `trust list` renders +it inside the `stub(no-real-check: …)` marker, and it rides §3.6's event beside +`stub`. + +**Constraints.** It only makes sense alongside `stub = true`: a reason on a +non-stub entry is refused at parse and at add, the closed direction. It is +OPTIONAL on a stub — every pre-DKT-607 stub entry has none and keeps loading +with an empty reason. Changing or erasing it on a re-add is a `CONFLICT`, since +the reason is the documented decision and a silent rewrite would swap one +decision for another under a re-approval. + ## 3.6 Trust changes are event-logged (T9) `trust add` and `trust rm` write a **`trust-added` / `trust-removed` event** into @@ -597,12 +624,32 @@ T9's residual auditable rather than invisible; and `events` gains two kinds, which extends the closed set (engine-spine §7.6) — §6.4 lists every event kind this stage adds, so §9 item 2's closed-set check keeps passing. -Recording is **mandatory-or-fail** inside a repo (DKT-81): the event is written -before the store, and a recording failure fails the verb with the store -untouched, so the ledger and the allowlist cannot silently diverge. An +Recording is **mandatory-or-fail** inside a repo (DKT-81). The event and the +store are two writes, and one invariant fixes their order: whichever write +fails, the surviving partial failure must never let the record claim **less** +authority than the store grants. The invariant points opposite ways for +granting and revoking: + +- **`trust add`** writes the event before the store. A recording failure fails + the verb with the store untouched; a store failure after the event leaves a + recorded grant that never landed, which over-reports authority. +- **`trust rm`** publishes the store before it records the event. A publish + failure removes nothing and records nothing; a recording failure after a + successful publish fails the verb with the entry **already removed**, and the + verb says so. The record then still shows the entry as trusted, which + over-reports authority. Recording first would leave the reverse: a + revocation on record for an entry that still authorizes execution. + +Neither verb lets the ledger and the allowlist diverge silently. An **idempotent re-add emits no event** — an event proves a change, never mere repetition. +**AMENDMENT (DKT-2198, 2026-09-15): removal publishes before it records.** The +DKT-81 decision wrote the event before the store for both verbs. That order +still governs `trust add`; for `trust rm` it could leave a recorded revocation +over an entry the store still grants, so the removal verb takes the +store-first order above. + `trust list` and `trust rm` outside a repo work and write no event; the trust store is user-level and does not require a repo to manage. `add` and `rm` outside a repo say on stderr (and in the JSON warnings) that nothing was @@ -897,6 +944,71 @@ the next environment variable anyone invents. | `CI` | `1` | the near-universal convention for "non-interactive"; it makes tools skip prompts and progress spinners without docket having to know each tool | | `DOCKET_GATE` | the gate name | so a check can behave differently under docket if its author wants; opaque to core | | `DOCKET_REPO` | the repo root | the same value as `Dir`, for tools that need it in an env | +| `DOCKET_GATE_BASE` | the step's base commit sha, **worktree-recorded completion gates only** | so a range-shaped check can scan exactly the step's committed change — `DOCKET_GATE_BASE..HEAD` of the tree it runs in — see below *(added 2026-09-01, DKT-992)* | +| `DOCKET_STEP` | the step's reference, `STEP-N` | so a gate can ask the engine for its **own inputs** — the identity `docket step context` and `docket step artifacts` take — instead of re-deriving which step it is from `DOCKET_ISSUE` plus an instance-name convention; see below *(added 2026-09-03, DKT-1186)* | +| `GOLANGCI_LINT_CACHE`, `STATICCHECK_CACHE` | a scratch directory deleted with the tree, **gates measuring a reconstruction only** | both tools cache issues by package content while storing the absolute path each was found at, and re-open that path to find the `//nolint` that suppresses it. A reconstruction outlives neither, so its entries must not either — see §7.6's DKT-1166 amendment *(added 2026-09-03, DKT-1166)* | + +**`DOCKET_GATE_BASE` — the step's committed range** *(DKT-992)*. Executors +commit **before** `step record`, so at gate time a worktree-recorded step's +tree is clean: a working-tree-only scan measures zero lines however large the +change (RUN-66's secret-scan passed 8/8 write steps that way), and a gate +guessing `git diff HEAD~1` is wrong for every multi-commit step. The engine +already knows the step's base — the worktree's **fork point**, the same +resolution the diff stage's `runDiffBase` applies — so completion gates of a +`--worktree`-recorded step export it: + +- **Worktree-recorded step**: `DOCKET_GATE_BASE` names the commit the worktree + was created from. `git diff $DOCKET_GATE_BASE..HEAD` in the gate's own cwd + (the worktree, per DKT-9) is exactly the step's committed change — the same + range the recorded `issue.diff` describes. +- **Non-worktree step**: the variable is **unset** — that is the documented + pick between the two admissible encodings (unset, or equal to `HEAD`). The + shared checkout has no fork point, the run's pinned commit is not this + step's base (sibling work lands between them, DKT-42's over-attribution), + and a live `HEAD` read is a value docket cannot vouch for as a range + endpoint. Absence — never an invented sha — is the encoding, the same + convention as `DOCKET_SCOPE`. +- The variable is also unset when the fork point cannot be resolved, and on + the pre-claim path (a pre-gate measures the tree under review, not a + recorded completion; after integration sweeps a worktree, no honest base + survives to export). +- **Fail closed on absence**: a range-shaped gate that finds the variable + absent while the tree is clean has nothing it can honestly scan, and should + fail rather than pass having measured nothing — "we couldn't check, so + carry on" is what makes a control decorative (N3). + +**`DOCKET_STEP` — the gate's own identity** *(DKT-1186)*. A gate frequently +needs an **artifact an earlier step of the same issue produced** — a threat +model feeding an abuse-case check, a synthesis feeding a verifier. Nothing in +the child environment named the step, so the only route was to rebuild the +gate's identity from outside the engine: `docket step list --issue +$DOCKET_ISSUE`, pick the row by a **hardcoded instance-name convention**, parse +that listing's JSON shape, then `docket step artifacts` on the result. That is +three couplings to things the engine is free to change — instance naming, +listing order, wire shape — and each one breaks silently when it moves. It was +observed in the wild (RUN-80's activation gate, `agentic-services` commit +`897c0a7`) and flagged as a precedent not to set. + +- **Value**: the step's rendered reference, `STEP-N` — precisely the argument + `docket step context STEP-N` and `docket step artifacts STEP-N` take. The + bundle's `inputs` **are** the artifacts the step was handed, so + `docket step context $DOCKET_STEP` answers "what were my inputs" in one verb + against a stable identity, and `docket step artifact ARTIFACT-N` fetches a + body from there. +- **Set on both paths**: completion gates (the saga) and pre-gates (the + pre-claim path) alike. The pre-claim path is where it matters most — a + pre-gate runs *before* the claim hands the bundle over, so asking the engine + by reference is its only route to the chain's artifacts. +- **No new authority.** Both verbs are read-only and take no token, and the + reference is an identifier the child could already reconstruct by hand. This + makes the lookup cheap and correct rather than conventional and fragile; it + does not widen what a gate may do. `DOCKET_PATH` and `DOCKET_TOKEN` remain + excluded, unchanged. +- **Unset, never `STEP-0`**: a gate spawned with no step in hand (a bare runner, + a future caller) sees the variable **absent**. A well-formed id that resolves + to nothing fails inside the gate's own tooling with a misleading message, + where absence is a condition the gate can test — the same encoding + `DOCKET_SCOPE` and `DOCKET_GATE_BASE` use. **Excluded, by name, in addition to being absent from the allowlist:** @@ -1179,11 +1291,11 @@ repo, acquired for the duration of a `tree = true` gate's execution. | # | Clause | |---|---| | L1 | The lock is `flock(2)` (`syscall.Flock`, `LOCK_EX`) on a lockfile created `0600` at `/.docket/tree.lock`. It is **advisory and process-scoped** — held by the docket process running the gate, released when that process exits **by any means, including SIGKILL**, because the kernel releases flocks on fd close. This is the property a database row or a lockfile-with-a-pid cannot match: a crashed engine leaves no stale lock to clear | -| L2 | It is **per repo**, keyed by the repo root — the same identity as §3.4's, so a second checkout of the same project does not serialize against the first | +| L2 | It guards **the shared checkout**, keyed by the project — the same identity as §3.4's. **Only a gate whose working directory is the shared checkout takes it.** A gate running in a step's own `--worktree`, or in a pre-gate's scratch reconstruction (§7.6), takes **no lock**: that tree has one step behind it, and keying every tree's gate to one per-project lock serialized a whole wave's records and claim-time pre-gates behind each other until the bound expired, recording correct work as unmeasured and holding claims past the executor's tool timeout. The accepted trade: two gates on the **same** isolated worktree are not serialized against each other either. Engine actions (docs/tdd/payloads-thresholds.md §6.2, A7; `action_exec.go`) always run in the shared checkout and always take it | | L3 | It is acquired **immediately before the spawn** and released **immediately after the process exits**, outside every transaction, and it is never held across a database write | -| L4 | Acquisition **blocks**, with the gate's own timeout as its bound: a gate waiting on the mutex longer than its timeout records `verdict='fail'` with `reason` naming the wait. Blocking rather than failing fast is correct — the whole purpose is to make the second gate wait for the first | +| L4 | Acquisition **blocks**, with the gate's own timeout as its bound: a gate waiting on the mutex longer than its timeout records `verdict='skipped'` with `reason` naming the wait — nothing ran, so nothing about the tree was measured, and the execution verdict stays `fail` so routing remains fail-closed. Blocking rather than failing fast is correct — the whole purpose is to make the second gate wait for the first | | L5 | An in-process `sync.Mutex` is held **in addition**, because `flock` semantics between two fds in the *same* process are not exclusion; two goroutines in one engine must serialize on the Go mutex, and two engine processes on the flock. Both, or the single-process case silently races | -| L6 | Gates without `tree = true` take **no lock** and run in parallel freely — engine-core §5's "read-only fan-outs parallelize freely (the proven win)" | +| L6 | Gates without `tree = true` take **no lock** and run in parallel freely — engine-core §5's "read-only fan-outs parallelize freely (the proven win)". So do `tree = true` gates on an isolated tree (L2): parallel writes in distinct worktrees are the other proven win, and the lock must not take it back | | L7 | **The lockfile is opened `O_NOFOLLOW`, and an existing non-regular file is a refusal.** `.docket/tree.lock` sits inside the repository, so it is repo-shippable content: a hostile repo can commit it as a **symlink** (to `~/.ssh/config`, to a device node, to a FIFO), and an ordinary open would follow it. The open is `O_RDWR\|O_CREAT\|O_NOFOLLOW` with mode `0600`; if the path exists and `Lstat` reports anything other than a **regular file**, docket refuses with a `VALIDATION_ERROR` naming the path and what it found — the **identical language and disposition as §3.2's I1**, which established this discipline for the trust file. The `tree = true` gate does not run; it records `verdict='fail'` with the refusal as its `reason`, because the serialization it requires cannot be provided. | L7's blast radius is small on its own — the worst case is an `flock` taken on an @@ -1275,6 +1387,14 @@ step's inputs rather than of any gate: | sha reachable, tree gone | **Reconstruct** it — `git worktree add --detach` into a throwaway checkout, measure, release. Sweeping a checkout does not delete the object | | neither | `skipped`, spawning nothing, with a reason naming the sha or the swept path | +**The lock timeout is the same case (DKT-91).** A `tree = true` gate whose +working-tree mutex does not come free within its bound never spawns either, so +it records `skipped` with a reason naming the wait — no exit code, no duration, +no output — and parks as unmeasured rather than routing per `on_fail`. The +cause differs from a swept worktree; the fact does not. `fail` there spent the +token a genuinely failing build spends, and a fix loop entered on it would ask +a worker to fix a tree the engine never opened. + Reconstruction is what makes mode 2 a fixed bug rather than a documented park: parking every swept-worktree verify is honest and useless. It is NOT a best-effort fallback — if the sha cannot be checked out the gate skips, exactly @@ -1305,6 +1425,46 @@ PG4 is unchanged and still applies: a PRE-gate that could not bind its tree is data for the step's worker, not a park. Parking on it would be the engine judging a step by an input it handed the step itself. +### AMENDMENT (DKT-1166) — a throwaway tree gets throwaway linter caches + +DKT-254 gave a pre-gate the right tree. It did not give it a cache that dies +with that tree, and a class of tool needs exactly that. + +**The mechanism.** golangci-lint and staticcheck cache each reported issue +keyed by **package content**, storing the **absolute path** the issue was found +at, and re-open that path afterwards to look for the `//nolint` (or +`//lint:ignore`) comment that would suppress it. A reconstruction is deleted +within the minute; the content hash is not. So one reconstruction's entries are +replayed in the next, the suppression lookup re-opens a file that is gone, and +an already-suppressed issue is re-emitted as live. + +**Observed**: harness RUN-64/STEP-2939 recorded `ac-commands: fail, exit 2` over +a clean tree. Build and tests exited 0; `make lint` reported one forbidigo issue +at `../docket-pregate-4091742512/…/timelinecompare_test.go` — a directory an +earlier reconstruction had already removed — with golangci-lint warning it could +not read that file, while the source carried `//nolint:forbidigo` on the line +above and the same sha linted in place reported `0 issues.` + +**The cwd was never the problem.** DKT-254 already binds the reconstruction and +`gate_exec.go` already spawns in it, which is why the reported path was +*relative to* the current reconstruction. The carrier is the cache, and it +poisons in both directions: the operator's own persistent cache also receives +entries naming a `docket-pregate-*` path docket is about to delete. + +**The rule.** A gate that measures a tree docket will delete gets its +path-carrying result caches inside a scratch root docket deletes with that tree +(`GOLANGCI_LINT_CACHE`, `STATICCHECK_CACHE` — §5.3). A gate over a tree that +stays on disk keeps its shared caches: re-analysis is a real cost, and it is +only worth paying where the tree is genuinely throwaway. The Go **build** cache +is deliberately untouched — what makes an entry dangerous here is a stored +source path the tool re-opens to decide suppression, and relocating `GOCACHE` +would rebuild the standard library on every reconstruction for no such benefit. + +A cache root that cannot be created **fails the reconstruction**, which records +`skipped`, on the same reasoning as the rest of this section: a measurement that +can report a suppressed issue as live is measuring the wrong thing, and that is +the defect, while measuring nothing is a gap. + ### 7.6.1 Ordering and the claim restructure `ClaimStep` today is **one transaction** (internal/engine/claim.go:73–202): reap, @@ -1345,8 +1505,9 @@ Transaction A sets `expires_ms` when it wins the CAS. Phase 2 then runs **subprocesses** — each bounded by the gate's own timeout, which defaults to **5m** (§5.4 X5) and can be a per-entry override, and a step may declare several pre-gates that run one at a time. The wall time between the CAS and the response -is therefore unbounded in principle and minutes in practice, and **all of it is -deducted from a lease the caller has not yet received**. A pre-gate-heavy step +was therefore unbounded in principle and minutes in practice (it is now capped +by §7.6.2 PG5's 60s budget, which is still longer than a short TTL), and **all +of it is deducted from a lease the caller has not yet received**. A pre-gate-heavy step hands its worker a mostly-spent lease; with a short configured TTL it can hand over an **already-expired** one, so the worker's first `step complete` fails on a lease it never had a chance to use, the step is reaped, and the pre-gates run @@ -1378,13 +1539,26 @@ when the caller gets it. | # | Clause | |---|---| -| PG1 | A pre-gate resolves through the **identical** matcher (§7.2), the identical env allowlist (§5.3), the identical timeout and process-group kill (§5.4), the identical capture (§5.5), and the identical tree mutex when `tree = true` (§7.4). There is no pre-gate-specific path anywhere in `internal/exec` or `internal/trust`. | +| PG1 | A pre-gate resolves through the **identical** matcher (§7.2), the identical env allowlist (§5.3), the identical timeout and process-group kill (§5.4), the identical capture (§5.5), and the identical tree mutex when `tree = true` **and** it runs in the shared checkout (§7.4 L2) — a pre-gate measuring a step's worktree or a scratch reconstruction takes none, exactly as a completion gate there takes none. There is no pre-gate-specific path anywhere in `internal/exec` or `internal/trust`. | | PG2 | An **unmatched pre-gate does not execute**, records `verdict='unmatched'` with `pre = 1`, and **the claim still succeeds** — the result rides in the bundle as `unmatched`, and the step's worker sees that its measurement did not run. A pre-gate is a measurement whose *result* the step consumes; refusing the claim would make an untrusted command able to block work, which is a denial-of-service an issue author should not have. | | PG3 | A **failing** pre-gate (non-zero exit) likewise does not refuse the claim: the result is `fail` and rides in the bundle. §11.1 calls these "measure-then-judge steps" — the judging is the step's job, which is exactly why the failure is data rather than a refusal. | | PG4 | Pre-gate results are **excluded from the saga's gate verdict** (`gateVerdict` filters `pre = 1`). They are inputs to the step, not judgments of it. S3 already excludes `pre` gates from `completionGates` (saga.go:311–320); this is the read-side counterpart. | +| PG5 | **The pre-gate phase is bounded as a whole** by a **60s budget** (`claimPreGateBudget`, pregate.go), taken from the start of phase 2 so reconstruction, any lock wait, and every gate's execution spend from one purse. Each gate's timeout is clamped to what remains — a flaky entry's re-runs included, through `exec.Spec.Deadline` — and a gate whose turn comes with nothing left records `verdict='skipped'` with `reason` naming the budget. A gate the clamp cuts off records its timeout as `fail`, with `reason` naming the budget beside the entry's own timeout so the row does not read as a changed entry. **The claim still succeeds** (PG2, PG3): the budget decides *when* `docket step claim` returns, never *whether*. The bound exists because an executor runs the claim under a **120s tool timeout**, and a claim that outlives it is backgrounded with its token unread and its lease left to a forced reap; sixty seconds leaves the rest of that window for transactions A and B, context assembly, and the packet render under load. | +| PG6 | **A pre-gate the budget cannot hold is measured ahead of the claim and served to it** (`pregate_detached.go`). The budget stays: PG5 decides when the claim returns, and a gate longer than it — ac-commands, the full test suite — could therefore never hand the claim a complete result; every claim cut it off and recorded the cut, and a verify step whose evidence *is* that gate's exit had nothing to judge. So the engine starts the run **detached**, at the earliest point the step's target sha exists: the commit that records a tree-holding step's `issue.diff` round record (its `head`, `appendRoundDelta`), reached at the saga's close (`ResumeSaga`), and the two commits that move that record afterwards — a resolution's `--worktree` re-pin and `step annotate --integrated-sha`. Each resolves every **pending** pre-gated step of the run that consumes `issue.diff` to its target exactly as the claim will (`preGateTarget`'s resolution, so the key written is the key looked up) and launches `docket step pregate STEP-N --target SHA` in **its own session, with every standard stream on the null device, and the handle released**: the launching verb never waits, and the harness killing that verb's process group at its tool timeout does not reach the child. The child **reconstructs the sha** (§7.6, the same scratch tree, sidecar flock, and dispatch sweep) even when a worktree still stands — the row is keyed by the sha, so the sha is what it measures, and integration sweeps that worktree mid-run — runs each pre-gate through the **same body the claim runs** (`measurePreGate`: PG1 holds, with `Deadline` zero so the entry's own timeout is the bound), and records the rows with `gate_results.target_sha` set (schema v37). **At claim** (`runPreGates`): a **complete** result — a process ran and exited — for **this step and this exact target** is served into `pre_gates` and the packet's gate-results and **nothing spawns for that gate**; any other state — no result, a run still in flight, a result for an earlier target, a run that recorded only `skipped`/`unmatched` — leaves the gate on PG5's path, unchanged. **The claim never waits** on a detached run; **the scheduler holds the step instead** (`CondPreGatePending`, ready.go): a pending pre-gated step whose (step, resolved target) in-flight lock is held right now is not ready — `next` and `dispatch open` do not offer it, the staged closure does not stage it, `step show`/`step list` report the condition as its `blocked_reason`, and a claim that arrives anyway is refused naming it rather than served a budget-cut row — until the lock lifts: the child's release once its rows are recorded, or the kernel's when the child dies. The hold is probed once per scheduler snapshot (`loadDetachedPreGateHolds`), non-blocking, for the exact target the claim would resolve, creating no lockfile and keeping none it acquires, and probed again for a step the lazy reap returns to pending inside that snapshot (`refreshDetachedPreGateHold`, from the claim's own reap and the shared one `next` and the dispatch verbs run), so a lapsed lease is not a way past it. It sits after every clause of the readiness conjunction but the budget (R7): a verify-ac step still behind its predecessors reports those and stays stageable while the lock is held, and a held step reports the hold rather than `no budget headroom`, so the claim's `--cost-multiplier` override — which admits a budget refusal on the real cost, trusting that nothing else failed — cannot claim through it. **Only a flock another holder refuses is a hold**: a lockfile that is missing, that cannot be opened, or whose flock fails for any reason other than a holder reads as not held — for the scheduler and for the launcher's "in flight" probe alike — so a filesystem where flock fails costs one child that exits at its own lock per completion, never a step held for the rest of the run; and a lock nobody took — a run that never launched, a lock directory that does not resolve — holds nothing, leaving readiness and the claim exactly as they were. A live run is bounded by its entries' own timeouts, so the hold is too. **A held step is not a dispatch discrepancy**: `dispatch verify` (and the reconcile's verify stage) reads a stored manifest row whose step is held — the staged verify-ac row of an implement → verify-ac lane, once the writer's completion launched the child — as `matched`, not `genuinely-missing`; the manifest promised a claim the wave makes once the lock lifts, and the hold changes nothing about the offer. The lock directory is fixed from the config the CLI already resolved (`PrimeDetachedPreGateLockDir`), so no scheduler snapshot resolves it under its transaction's write lock. **Crash safety**: rows land only after a gate finished, one transaction per gate, so a child killed mid-run leaves no row and nothing reads as a pass; its scratch tree is reclaimed by the next dispatch open or close through the sidecar flock the kernel released; no tree mutex is involved, because a sha-keyed target is never the shared checkout (§7.4 L2). **One run per (step, target)** is an exclusive non-blocking flock on `pregate-STEP-N-.lock` beside `tree.lock`, taken by the child for its lifetime and probed by the launcher — liveness is the flock, never a pid, as everywhere in this document. **At claim**, a served gate carries exactly **one** row: the latest complete row for (step, target); a stale `unmatched` beside a later pass is not served with it, and a relaunched child skips every gate that already holds a complete row. **The child declines to record** — inside the recording transaction, not only before the gate — when the step is no longer pending and unclaimed or no longer resolves the sha it measured: a result nobody asks about any more lands nowhere. A result keyed to the step's current target that lands **after** the claim ran its budgeted path is still recorded, at the next ordinal: `step show`, `step gates`, and a re-minted claim's replay (last ordinal per gate, skipping rows keyed to a target other than the claim's own) read it; the packet already rendered does not. | PG2 and PG3 together mean a pre-gate never blocks — a deliberate asymmetry with -completion gates, and the reason is in §11.1's own parenthetical. +completion gates, and the reason is in §11.1's own parenthetical. PG5 keeps +that asymmetry: a slow pre-gate is cut short and reported, not turned into a +refusal that a slow command could use to block work. PG6 keeps the budget and +moves the measurement: the same gate, the same trust path, the same tree, +measured where the budget does not reach and served to the claim whole. Its +hold is a readiness narrowing, not a claim wait: `docket step claim` still +returns inside PG5's budget, and a verify-ac step is simply not offered while +the detached run that will serve its claim is in flight — the moment that run +records, the step is ready and its claim carries the complete result. +Nothing about the gate-results row a packet carries changes shape under PG6; +a served row's `reason` names the target and says the measurement ran ahead of +the claim, and that is the whole of the difference a reader sees. ### 7.6.3 Where the results land in the bundle @@ -1502,6 +1676,8 @@ registry (SKILL.md's engine-configuration table), not a new table: |---|---|---| | `vote.rule..threshold` | float in (0,1] | the approval threshold this rule tallies at | | `vote.rule..criticality` | `low\|medium\|high\|critical` | the proposal's criticality | +| `vote.rule..sealed` | bool, default `false` | whether proposals opened under this rule withhold their casts from the read verbs until the tally closes them (DKT-2447, below) | +| `vote.rule..hold_on_dissent` | bool, default `false` | whether an approved tally that carries at least one `reject` cast parks the vote step for the operator instead of passing | `` is an opaque string, exactly as `lease.ttl.`'s class is. A rule "exists" iff `vote.rule..threshold` is set. This reuses the config @@ -1512,6 +1688,84 @@ machinery, its `VALIDATION_ERROR`-on-unknown-key behavior, and its §11.1 puts the voter list on the step. A rule is about *how strictly to tally*; the step is about *who casts*. +**Roster and weighting are declared on the vote step, not in config +(DKT-2562, DKT-2764).** Who may cast and what a cast is worth are +authorization decisions, and a config key carries none: `docket config set` +has no per-caller identity, so a seat constrained by a rule could rewrite the +rule before casting. The two switches therefore live in the vote step's +`[[step]]` table — `roster = "open"` (the default) or `"strict"`, and +`weighting = "declared"` (the default) or `"equal"` — and activation pins them +with the workflow beside `voters`, where a live ballot cannot have them +changed underneath it. They are **not engine-config keys**: the +`vote.rule..roster` and `.weighting` keys that once registered them are +retired, and `config set` refuses each by name pointing at the field that +replaced it. Enforcement applies on **vote steps only**: `docket vote cast` +resolves the proposal back to the pinned step and reads the switches there, +so a proposal no vote step opened (an operator's own, a reap acknowledgment's, +a materialized held cluster's minted panel) enforces neither. + +- `roster = "strict"` refuses a cast whose `--voter` is not one of the step's + pinned `voters` as a `VALIDATION_ERROR` naming the voter and the roster, and + writes no vote (DKT-2511). `open` is today's behavior: the list is counted + and a cast under any name fills a seat. +- `weighting = "equal"` tallies every cast at 1.0 × 1.0 while each vote row + keeps the confidence and domain relevance the seat declared (DKT-2512); + `declared` is `db.CastVote`'s existing arithmetic, reached through the same + function with one flag. + +**Why this reverses the opaque-voter decision, and only for steps that opt +in.** `internal/engine/vote.go`'s proposal construction records that "the +voter hints themselves are OPAQUE: core never interprets one" — it counts +them for `required_voters` and nothing more, and that stays true of every +ballot by default. A step declaring `roster = "strict"` opts into the one +comparison the cast path then makes: the name on the cast must be one of +those hints. Core still never dispatches to a hint and never reads meaning +into one; it compares two strings the author wrote. Making that comparison +the default would have turned every registered ballot strict at once, and a +threshold nobody chose is not a threshold — so the switch is opt-in, per +step, and pinned. The recusal seam (`reviews`, `recuse = "executor"`) is the +same shape and lives in engine-spec §11.1. + +**Sealed ballots (DKT-2447).** Every proposal used to render every recorded +cast — verdict, confidence, relevance, weight, findings, summary — while it +was still open, so a seat that read the proposal after a sibling had cast saw +the sibling's verdict and reasoning before casting its own: the public-board +channel for anchoring and collusion. `vote.rule..sealed = true` closes +that channel for the proposals a rule opens. The flag is resolved beside the +threshold in `resolveVoteRule` and **stored on the proposal at open** (v28's +`proposals.sealed`, reliability-delta §2), so an edit to the rule cannot +change what a live ballot renders under the seats mid-vote; `vote create +--sealed` is the same flag for a conversational proposal. + +While a sealed proposal is `open`, `vote show`, `vote result` and `gate +status` render only **who has cast and how many casts are in**: the human +views list voter names under a `sealed` note, the JSON `votes` array carries +`{voter_name, created_at}` entries and no verdict, confidence, relevance, +weight, findings or summary keys, and `gate status` reports each seat's +`cast` without its `verdict`. `vote list` already rendered only the count. +The moment the status leaves `open` — a tally, a `vote commit`, a `vote +close` — everything renders, and the `vote-record` context artifact a +downstream step consumes is composed only after the vote step routes, so it +never carries an open ballot's casts. + +Sealing is a **norm-level shield, not a security boundary.** It changes what +the read verbs *render* and nothing else: `db.CastVote`, its weighted score, +its quorum and the one-cast-per-voter constraint never read the flag, and the +vote rows stay exactly as readable as before through `docket export`, +`ListAllVotes`, and direct store access with `sqlite3`. A seat that wants to +read a sibling's cast can; the shield removes the default path that handed +it to every seat that merely looked at the proposal it was asked to judge. + +**Holding on dissent.** A rule's tally is a weighted mean, so two approve +casts outweigh one reject and a dissenting seat's verdict leaves no trace in +routing once the score clears the threshold. `vote.rule..hold_on_dissent = true` +closes that gap: when such a rule's proposal tallies approved but at least +one cast's verdict is `reject`, the vote step parks `waiting-human` instead +of passing, with a reason naming the dissenting seat's voter name (each one, +sorted, when more than one seat dissented) — except a panel that is deciding +the tally's own question, which keeps the routing that decision already gave +it. + ## 8.4 What is NOT in scope here - **No new vote verb.** Casting, showing, listing, and committing use the diff --git a/docs/tdd/packet-composition.md b/docs/tdd/packet-composition.md index 7e1be110..6ea3934d 100644 --- a/docs/tdd/packet-composition.md +++ b/docs/tdd/packet-composition.md @@ -311,7 +311,51 @@ would inflate the closure for no gain. The default template gains a section rendering each file's body between delimiters carrying its path and hash. `== PINNED` **stays** — it is still the honest list of what the run pinned, and now the files that were inlined are -also legible as content rather than only as pointers. +also legible as content rather than only as pointers. For a pin the section +still lists only by path and hash, `docket pin show RUN-N PATH` prints the +pinned bytes at that path, so a step can read text the packet names but does +not inline. + +#### 1.4.1 `issue.files` — an issue's attachments reach the same section (DKT-44) + +`issue.files` is an `inputs` entry, declared beside the other engine-produced +forms: + +| form | resolves to | +|---|---| +| `issue.body` | the issue's activation-frozen body snapshot | +| `issue.diff` | the computed VCS diff recorded for the issue | +| `issue.files` | **the bytes of every path the issue attaches** | +| `issue.latest.` | the issue's latest recorded artifact of one kind | +| `issue.linked..` | an artifact recorded under a linked issue | + +A step declaring it receives each attached path as its own + +``` +== FILE + +``` + +section — the same section §1.4's declared `packet` entries render into, +appended after them, because the contract and fragments are what a worker reads +before the material the contract applies to. + +It is the one input form whose resolution **reads the filesystem**. The others +answer from run state, which is why they are snapshot-pinned; an attachment is a +path, and a path's contents live where the project keeps them. The read is +against the **run's recorded exec root** — the project checkout the issue's +paths are relative to — never the invoking process's cwd, because the claim that +needs the bytes typically runs from a linked worktree that does not have them. +That is the defect the form closes: the attachments that forced it were +untracked, so they existed in the shared checkout and nowhere else, and an +isolated executor told its inputs arrive in the packet had no sanctioned way to +reach them. + +An attached path the engine cannot read **refuses** with a `VALIDATION_ERROR` +naming the path, rather than rendering a packet that silently omits a declared +input. Since `step claim --render` renders as a pre-claim preflight (§1.2's +refuse-rather-than-drift discipline, applied on the claim path), that refusal +costs no lease. ### 1.5 Closure size: counted where the spec says to count it @@ -329,6 +373,16 @@ pinned and hashed them. **This is the one place the fix reaches beyond render.** It is not scope creep; it is the difference between the caps meaning something and meaning nothing. +**`issue.files` attachments are the exception: their bytes are not counted**, +in `ContextSize` or against the §11.1 caps. The caps run when activation expands +the steps, against byte counts activation has just pinned. The form does not +pin attachments: §1.4.1 reads them live from the run's exec root when the +packet renders, because the attachments that forced the form were untracked. At +expansion there is nothing to measure, so attachments are deliberately +uncapped rather than estimated. The cost: a step declaring `issue.files` can +render a packet larger than its recorded closure size, and nothing refuses it. +Each attachment's `== FILE` header carries its path and hash, not its size. + ### 1.6 Registration and validation Two new validation rules. **The numbers are V32 and V33**: V29 and V30 are diff --git a/docs/tdd/payloads-thresholds.md b/docs/tdd/payloads-thresholds.md index 7ba6d3cb..ba22b6fd 100644 --- a/docs/tdd/payloads-thresholds.md +++ b/docs/tdd/payloads-thresholds.md @@ -415,13 +415,14 @@ U3 exists because it is not hypothetical: group 1 ships `schemas`, group 2 ships `action_results`, and the operator's own tracker is migrated by whichever binary happens to be built between them. -## 4.5 `docket schema register|list|show` +## 4.5 `docket schema register|list|show|deprecate` | Verb | Flags | Effect | |---|---|---| | `docket schema register ` | `--json[=v2]` | read; parse as JSON; validate as a schema document (the library compiles it — a schema that does not compile is refused here, not at first use); derive the ordered index (§4.3); insert per §4.4's three outcomes | -| `docket schema list` | `--json[=v2]` | registered schemas: `name`, `version`, `sha256`, `ordered_fields`, `builtin`, `created_at_ms`. A `Collection` envelope under v2, per the `workflow list` precedent | -| `docket schema show ` | `--json[=v2]`, `--body` | the row; `--body` emits the registered **bytes verbatim** (what a run validates against, not a re-serialization) | +| `docket schema list` | `--json[=v2]`, `--deprecated` | registered schemas: `name`, `version`, `sha256`, `ordered_fields`, `builtin`, `created_at_ms`. A `Collection` envelope under v2, per the `workflow list` precedent. Retired versions are hidden unless `--deprecated` is passed; v2 items carry `deprecated_at_ms` when set *(amended 2026-09-24 — DKT-2792)* | +| `docket schema show ` | `--json[=v2]`, `--body` | the row; `--body` emits the registered **bytes verbatim** (what a run validates against, not a re-serialization). A bare name resolves the highest version still in service; an explicit `@version` resolves a retired one | +| `docket schema deprecate ` | `--json[=v2]`, `--restore`, `--project`, `--all-projects` | sets `deprecated_at_ms` and never deletes; refuses (CONFLICT, no override) a version a workflow still in service names as `payload`, naming the referencers, and refuses the builtin. `workflow register`, `workflow lint`, and auto-registration refuse a new `payload` reference to a retired version; runs that pinned it are untouched *(amended 2026-09-24 — DKT-2792; engine-spec §11.1)* | **`name@version` is an argument, not two flags** — §1's surface line is `docket schema register name@v schema.json` and it is followed exactly. The @@ -703,6 +704,99 @@ mechanical: package-level variable — the same purity discipline as §4.9.2, and what lets the table tests run without a database. +### 5.1 The reserved `diff.*` family *(added 2026-09-24, DKT-2063, DKT-2518, DKT-2548–DKT-2551)* + +A write step's threshold can route on the SIZE of the change the step recorded, +which no payload predicate can reach: `implement` emits a markdown +change-summary with no payload at all, and the facts a change track wants — +how big was the change, was there one — exist only in the ledger. Three +reserved field names address them. They are the engine's own measurement, never +a payload field, and no schema declares them. + +| Field | Value | Measured as | +|---|---|---| +| `diff.lines` | integer | added PLUS removed content lines | +| `diff.files` | integer | files the diff touches (`diff --git` headers) | +| `diff.empty` | boolean | whether the recorded body holds a change | + +**Evaluated at record time, over the step's recorded `issue.diff` round +record.** The measurement is taken in the completion transaction of a +tree-holding executor step (a fanout sibling included), over the in-scope cumulative diff that completion +recorded (the object a review is sized against; the round-delta and +out-of-scope trailers are not counted, since neither is this issue's own +change). Routing then applies the ordinary interposed-target semantics: the +matched routing's step is released and the unrouted siblings skip in the same +transaction, exactly as a payload predicate would route. The aggregation is NOT +applied — `diff.lines` is one number for the step, not a column over a set, so +`any(diff.lines > 20)` and `all(diff.lines > 20)` are the same assertion. +`diff.lines` and `diff.files` compare numerically without an `ordered_enum`: +T3 exists because core does not know whether `high` outranks `medium`; it does +know that 21 > 20, so there is no order to guess and nothing to park on. + +Two measurement rules keep the facts honest about the object review reads: + +- `--- `/`+++ ` lines are file headers only between a block's `diff --git` + line and its first `@@` hunk header; inside a hunk they are content and + count toward `diff.lines` (DKT-2550). +- `diff.empty` agrees with the ledger's own record-or-drop test + (`diffRecordsNoChange`) over the same in-scope portion, so a rename-, mode-, + or binary-only diff the ledger keeps as a real change reads + `diff.empty == false` even though it carries no `+`/`-` line (DKT-2549). +- When an `--as retry` re-execution computes an empty diff and the DKT-259 + guard drops the re-record because the issue already holds a non-empty + `issue.diff`, the facts are measured from THAT latest recorded non-empty + body — what `issue.diff` resolves to and a review would read — not from the + empty body the retry computed (DKT-2548). That record's round record stands + too: downstream packets keep its `head` and `worktree` as the target, since + the retry's own record would name a tree at its base with no head to bind; + moving the target on purpose is the annotate-integration and + `--worktree` re-pin verbs' job. A FIRST empty diff, with no + record for the issue yet, still records: neither the DKT-259 guard nor the + byte-identical guard fires without an earlier record to protect or match. + +**The absent-record rule.** A tree-holding step whose completion recorded no +change evaluates as `diff.empty == true`, `diff.lines == 0`, and +`diff.files == 0` — a decided answer, not an unknown field. `any(diff.empty == +false)` is therefore DECIDED false for such a step and the review it keys is +skipped (cascading through `after_fired`), rather than falling through to the +no-such-field path. On a step that holds no tree there is no measurement, and +`diff.lines` is an ordinary undeclared field evaluated exactly as any other. + +**Register-time lint** (DKT-2518), decided in `Validate` on bytes alone, because +a step with no `payload` would otherwise take the V21d skip and reach the +engine with a predicate that can never mean anything: + +1. **V45 — placement.** A `diff.*` predicate on a step that does not hold the + tree — an `action` or `type` step, or an executor or fanout step declaring + `holds_tree = false` — is refused, naming the step and the tree-holding + requirement. The rule keys on the engine's evaluation condition (a + measurement exists exactly for a tree-holding executor row), not on a + `class` value: on any other step the facts are nil and the predicate would + silently never match. A fanout step is admitted because it expands to + executor rows, and each tree-holding sibling measures and routes on the + change it recorded *(amended 2026-09-24: the first cut refused fanout + outright, stricter than the engine)*. +2. **V46 — count literal.** A non-numeric literal under an ordered operator + (`<`, `<=`, `>`, `>=`) on `diff.lines` or `diff.files` is refused under its + own id, naming the literal and the operator: the counts are ordered + numerically, so the literal must be an integer. The same holds under `==` + and `!=` (`any(diff.files == none)`), since the engine parses a count + literal as an integer before it looks at the operator. +3. **V47 — boolean.** `diff.empty` under an ordered operator (it has no order) + or compared against a literal that is not a boolean is refused, naming the + field and the operator or literal. + +The engine still refuses every one of these shapes at record time +(`evaluateDiff`); the lint moves each refusal from a parked run to the author's +terminal, so a registered definition never reaches one. **V21a does +not apply** to the three names (DKT-2551): on a payload-declaring step the +cross-validation skips exactly `diff.lines`, `diff.files`, and `diff.empty` by +name — a lookalike such as `diff.bogus` is an ordinary undeclared field V21a +refuses — because the alternative told an author to add `diff.lines` to a +payload schema, which the design forbids. The names are the workflow package's +(`workflow.DiffFields`), read by the engine, so validator and evaluator share +one spelling as they do for `VoteCastFields`. + ## 6. The real `ActionRunner` `internal/engine/action.go`'s seam is honored: `NewEngine` swaps @@ -832,7 +926,7 @@ requirement rather than drift: |---|---|---| | M-a | **`ActionResult` gains `Held []int` and `Results []ActionResultRow`** | §2 requires the held-cluster outcome and §6.3 requires per-attempt records; a seam that could only return a payload could express neither | | M-b | **The routing stage gains a `held` branch** and the saga a `held` stage (§7.7) | §2's "materializes a `type=human` step … gating the routing step" is a *deferral* of routing, and a routing stage that always routes cannot defer | -| M-c | **`DecideStep` gains a materialized-step branch** (§7.7) | `approve`/`reject` on a materialized step resumes another step's saga, which the S3 path had no reason to do | +| M-c | **`DecideStepWith` gains a materialized-step branch** (§7.7) | `approve`/`reject` on a materialized step resumes another step's saga, which the S3 path had no reason to do | | M-d | **Stage 0 gains schema validation** (§4.8) and `EvaluateThreshold` gains a resolver parameter (§5) | both are this stage's subject | Recorded as a **note on DKT-5** (§11 A7), not as a spec amendment: no @@ -863,6 +957,8 @@ unchanged. | `method` | `median` \| `max` \| `min` | yes | the reduction | | `hold_spread` | integer ≥ 0, default 0 | no | hold when spread **≥** this; `0` never holds | | `output` | string | yes (already V11's, §4.3.1) | the produced artifact kind | +| `route_at` | string, a value of `field`'s declared order | no | routing floor *(amended 2026-08-23, DKT-593 — see the amendment below)*: a cluster whose reduced value's position is **≥** its position is emitted to the output payload; the rest go to the record. Absent ⇒ every cluster emits, byte-for-byte the pre-`route_at` output | +| `source_field` | string, a property name | no | *(amended 2026-09-13, DKT-2462 — see the amendment below)*: names a property of each input element holding an array of opaque source labels. Core neither validates nor reduces by it; the run report groups clusters by it. Absent ⇒ no such grouping in the report | New register-time rules, each a `VALIDATION_ERROR` naming workflow, step, and param, each a test case: @@ -870,7 +966,8 @@ param, each a test case: | # | Rule | Argument | |---|---|---| | V27 | a step's `name` may not end in **`-held`**, and `action` may not name a builtin other than `aggregate` | the first reserves the materialized identity (§7.7) so a definition cannot collide with one; the second turns "my trusted `aggregate` command never runs" into a register-time sentence | -| V28 | `action = "aggregate"` requires `field`, `method ∈ {median,max,min}`, `output`; `hold_spread` an integer ≥ 0 if present; **no other keys** | the discipline every V-rule follows. A typo'd `method = "medain"` is otherwise discovered hours into a run, on a step whose inputs are already spent | +| V28 | `action = "aggregate"` requires `field`, `method ∈ {median,max,min}`, `output`; `hold_spread` an integer ≥ 0 if present; `route_at` a non-empty string if present *(DKT-593)*; `source_field` a non-empty string if present *(DKT-2462)*; **no other keys** | the discipline every V-rule follows. A typo'd `method = "medain"` is otherwise discovered hours into a run, on a step whose inputs are already spent | +| V28a | `route_at`, when declared, must name a value of `params.field`'s declared order *(DKT-593)* | the floor is a **position** in that order, and a value with no position has no floor to name — G4's discipline asked at register time. Schema-aware, so it lives in `ValidateSchemas` beside V29 rather than in the pure-bytes V28 | | V29 | an `aggregate` step must declare `payload = name@version`, and that schema must declare `params.field` as **`ordered_enum`** | median, max, and min are all defined **only** over an order. An aggregate without a declared order is a step that can never compute, and §11.2's own restriction is the same restriction | | V30 | the declared schema must **accept an aggregate-shaped document**: a synthetic probe built from the schema's own declared enum values (§7.6) is validated against it at register time | the output must satisfy *both* the instance schema and `aggregate@1` (§7.6). An instance schema with `"additionalProperties": false` makes that conjunction unsatisfiable — and the failure would otherwise land at the end of a review fan-out, hours in. The probe is deterministic and invents nothing: every value in it comes from the schema being checked, and it carries the aggregate output keys (`members`, `held`) plus one carried-through extra key so the conjunction it tests is the real one (review F3) | @@ -881,6 +978,57 @@ magic (nothing in the grammar says an action reads its predecessor's payload), it is unstable under `inputs` edits, and it hides the one declaration an author most needs to see. +### AMENDMENT (DKT-593) — `route_at`: a routing floor + +RUN-43 spent 71.6% of its output tokens on fix rounds that closed roughly as +many clusters as they opened, because every reconciled cluster — whatever its +reduced value — entered the loop, and `contracts/fix.md` rightly forbids the +FIXER from filtering ("the reconciled set is the work"). The missing piece was +never the fixer's to add: which clusters are worth a round is a ROUTING +question, and routing is core's. + +`route_at = ""` names a floor in `field`'s declared order. Like every +value the builtin touches it is an **opaque token compared only by position** +(§4.3 I3) — a severity floor for a severity order, a ripeness floor for a +ripeness order, and core cannot tell the difference. + +| # | Clause | +|---|---| +| R1 | A cluster whose **reduced value's** position is **≥** the floor's position is emitted to the output payload — the wire the threshold evaluates and every downstream `inputs` reader consumes. The comparison is over the reduced value, exactly as the threshold's would be: a cluster whose members reach `high` but whose median is `low` is a `low` cluster to both | +| R2 | A cluster below the floor goes **to the record**: fully reduced — value, members, demotion trail — into the aggregate's own `action_results` row (`output` column, the audit channel every attempt already writes), and counted in the artifact body. It never enters the artifact payload, so the threshold and the loop never see it. Routed, not erased: the input artifacts remain immutable and addressable, and the reduction's trail is attributed in the row | +| R3 | A **held cluster is never routed below the floor.** Its computed value is exactly the value the hold refuses to trust — the members disagree, and the operator resolving it may accept a different value (`--value`) — so routing it out by that value would spend the decision the hold exists to ask. It is emitted, it gates the step (§7.7), and the floor's opinion waits for the operator's | +| R4 | `Held` indices address the **emitted** payload — the payload the artifact records and `-held@k#i` resolves against (H2a) — which coincides with input positions whenever nothing was routed below | +| R5 | **Absent ⇒ byte-for-byte today's output.** No key, no floor, nothing recorded, G2's identity property untouched — the same no-cliff discipline `hold_spread = 0` follows | +| R6 | Every cluster below the floor is legal: the emitted payload is then the **empty array**, a valid `aggregate@1` document over which any threshold predicate finds nothing — the round is simply not entered | +| R7 | Validation is split as V28/V28a: presence and type are pure bytes (V28); membership in the declared order needs the schema and lives in `ValidateSchemas` (V28a). At run time — a definition can arrive through a restored database — the same two refusals fire in `ParseAggregateParams` and `Aggregate` respectively, the second naming the value per G4 | + +This is **core routing, not fixer filtering**: the set the fixer receives has +already been routed, so `contracts/fix.md`'s prohibition stands unchanged — +there is nothing left in the reconciled set to filter. + +### AMENDMENT (DKT-2462) — `source_field`: reporting attribution without a domain vocabulary + +DKT-2452 gave the run report a per-`aggregate`-step count of unique versus +corroborated clusters (§7.6's `members`). It could not attribute a cluster to +the **executor** that produced a member: `members` is core's own reduced +record, with no per-member origin, and core must not resolve one by opening a +judge worker's own artifact payload — that is payload interpretation +(genericity, §1.1.1), and a finding's own `id` is unique only within its +producing worker's payload, so two workers can emit the same id and an id +alone cannot disambiguate. + +`source_field = ""` names a property of each input element holding an +array of opaque source labels — a workflow's own convention for a +member's producing step ref, say. Core reads the NAME to know which other key +to read; it never assumes the name `member_sources` or any other domain word. +G3 already carries that array through to the output element verbatim, so +`Aggregate` itself needs no change for the value to survive reduction — the +param exists solely so `docket run report` can group a round's clusters by +that key and resolve each value to the step instance (and its declared +executor hint) that produced it, without core hardcoding one corpus's field +name. Absent `source_field`, the report's grouping section is omitted +entirely, matching `route_at`'s absent-parameter convention. + ## 7.2 The input shape: what a "cluster" is §2 promises the output is "per-cluster value, members, held flag", and @@ -897,7 +1045,7 @@ is the *judged* half, performed by a synthesizing worker; reconciliation is the |---|---| | G1 | `field` holding an **array** ⇒ the members are its elements, in payload order (order does not affect any result; §7.3 sorts by position) | | G2 | `field` holding a **scalar** ⇒ a one-member cluster. The median/max/min of one value is that value, the spread is 0, nothing is held, and nothing is demoted | -| G3 | Every **other key of the element is carried through verbatim** into the output element. Core does not read them (genericity), and dropping them would strip an instance's own identifiers from the very payload a downstream step consumes | +| G3 | Every **other key of the element is carried through verbatim** into the output element, **except the core-owned names** (`members`, `held`, `demoted_from`, `operator_resolved`, `operator_note`, `operator_set_from`, `operator_set_mirrors`), which are dropped so that a key of that name in the output is always core's own write. Core does not read the rest (genericity), and dropping them would strip an instance's own identifiers from the very payload a downstream step consumes. **Aggregation never writes an author's key. Hold resolution is the one exception** (amended by DKT-42 and DKT-1548): when an operator resolves a held cluster with a corrected value, core writes that value onto the threshold field, retains the computed value it replaced in `operator_set_from`, carries the correction onto the other fields the step's own threshold compares, and names each field it rewrote in `operator_set_mirrors` — so every key core wrote on the author's behalf is recorded beside the decision that caused it | | G4 | A member value **not present** in the declared order (§4.3 I4) is a step failure with a reason naming the value and the field. **Not** sorted, not ignored | | G5 | An **empty** members array is a step failure naming the element index — an empty cluster has no median, and inventing one (null? the lowest?) is exactly the kind of guess this stage exists to refuse | @@ -1078,7 +1226,7 @@ clause below is this TDD's, and each is a test. | H12b | **Ordering inside the approve transaction** *(DKT-15)*. The decision is written to the cluster step **before** the payload is resolved, because resolution now reads each cluster step's own status to decide which elements to mark — with the old order the cluster being approved reads as undecided and nothing is marked. Under the previous whole-hold shape the order was immaterial, since approval was inferred from the caller rather than read back. Both remain in one transaction, so §7.7.3's atomicity is unchanged | | H13 | **Artifacts stay immutable** (engine-core §1.4). Approval records a **new** artifact rather than annotating the old one, which is the same rule a step re-run follows. The held payload remains addressable forever: what the engine computed and what the operator accepted are two records, not one overwritten one | | H14 | The materialized step ends **`done` on both** approve and reject — it recorded a decision, which is what a gate does. The *consequence* lands on the routing step. This is also why V13's rule (a human gate may not route rejects to `waiting-human`) is not violated when the routing step's `on_fail` **is** `waiting-human`: the park is on a **different** step, resolvable by `step resolve`, not a step parking on its own decision | -| H15 | Both verbs are **token-free**, per §2 — a human gate is resolved by an operator who never claimed it | +| H15 | Neither verb takes a **lease token**, per §2 — a human gate is resolved by an operator who never claimed it. Both require the run's **conductor capability** (DKT-2465, reliability-delta §2 v29), checked before the held step's status is inspected | | H16 | `approve`/`reject` on a materialized step whose routing step is **not** in `held` (a double approve, a resumed race) is `CONFLICT` naming both steps. The saga advance is CAS-guarded on `saga_stage = 'held'`, so the loser writes nothing | ### 7.7.4 §11.3 interplay @@ -1302,7 +1450,7 @@ starts from a shape rather than a blank page. | Non-goal | Why not now | What a future amendment would look like | |---|---|---| | **Dotted/nested threshold field paths** | §11.2's grammar is a bare field token; indexing nested orders would build a map no predicate can query (§4.2 O5) | extend `predicateShape` to `a.b`, and the ordered index to JSON-Pointer keys, in one change | -| **`step approve --set =`** for a held cluster | engine-core §6's "an accepted severity" needs a **typed** channel; parsing it from a note is core reading instance meaning out of prose (§7.8) | `--set field=value`, validated against the pinned schema's declared order, recorded as the resolved value with the computed one retained | +| **`step approve --set =`** for a held cluster | engine-core §6's "an accepted severity" needs a **typed** channel; parsing it from a note is core reading instance meaning out of prose (§7.8) | **Shipped in a different shape.** DKT-42 delivered `step approve --value ` — the value is the held threshold field's, validated against the pinned schema's declared order, recorded as the resolved value with the computed one retained in `operator_set_from`. DKT-1548 added the mirror correction: the corrected value is carried onto the other fields the step's threshold compares, each named in `operator_set_mirrors` (G3). A general `--set =` for an arbitrary field remains unbuilt | | **More builtins** (`mean`, `mode`, `percentile`) | §2 names exactly one builtin and says the rest are user-trusted commands. Each addition is core acquiring an opinion about what is worth computing | a named builtin plus its params in §2, with the same register-time validation table | | **A schema `deprecate`/`rm` verb** | schemas are immutable and pinned; deleting one breaks a completed run's ability to explain itself | a `retired_at_ms` column and a `list --all`, never a delete | | **Cross-field predicates** (`any(a >= b)`) | §11.2's right-hand side is a literal, and comparing two fields requires deciding whether their orders are commensurable — a question only the schema author can answer | an explicit `comparable_with` annotation, refused by default | diff --git a/docs/tdd/reliability-delta.md b/docs/tdd/reliability-delta.md index 38369a95..90eff36a 100644 --- a/docs/tdd/reliability-delta.md +++ b/docs/tdd/reliability-delta.md @@ -592,6 +592,537 @@ the counters are authoritative only for claims that ended after v23. The wire fields are `omitempty`, so every row with no counted outcome serializes byte-identically to v22's rendering. +### AMENDMENT — the span extends to v24 (DKT-546, 2026-08-22) + +**What changed.** v24 adds ONE table, `gate_override_grants`: one operator +ruling that a gate's failure signature — gate name + exit code + reason +classification, and since v30 the content fingerprint — is environmental for +the remainder of ONE run. A grant is +minted by `step resolve --as override-pass --batch` (one row per failed +completion gate of the parked step, in the resolution's own transaction), and +spent by the routing stage: a later step of the same run whose EVERY failing +gate matches a grant routes the same generic `pass` the operator's own +override-pass records, instead of parking. `covered_steps` counts the spends, +bumped in the routing transaction; both edges are event-logged +(`gate-override-granted` / `step-batch-overridden`, attributed human / +threshold), so the feed walks from every auto-pass back to the person. The +grant dies with its run — the `run_id` FK is the whole scope rule, and a new +run re-asks. A cover blocked by an interposed threshold target (DKT-470's +shape) parks as before, with the block named. + +**What it fixes.** Refit mining of 25 runs / 34 issues found the dominant +operator toil is environmental gate parks — build failed 30/46 recorded +verdicts, tests 26/45, self-hygiene 18/34 — virtually every one +operator-overridden as a sandbox artifact, not a code defect, and each park +resolved individually: in RUN-42 the operator's own resolution was +"override-pass each as it parks", the same ruling re-made per step. No +mechanism let one ruling cover subsequent identical failures in the same run. + +**Why the ratified arithmetic is untouched.** Like v11–v23, v24 is an +amendment, not a stage: one additive table, `CREATE TABLE IF NOT EXISTS` +throughout so the migration is idempotent and re-runnable, and a rewind guard +that probes the TABLE (the v7/v8 form, since v24 adds no column). It +BACK-FILLS NOTHING — no operator granted a batch override before the verb for +granting one existed — and it is dormant: a run that never records a grant +reads byte-identically to v23 on every verb. + +### AMENDMENT — the span extends to v25 (DKT-742, 2026-08-25) + +**What changed.** v25 adds ONE table, `stale_target_waivers`: one operator +ruling that a specific stale-target warning — one (step instance, target sha) +pair — has been adjudicated for the remainder of ONE run. A waiver is minted +by `dispatch waive-target` (one row per named step instance, all for one +target sha, in one transaction, each row event-logged as +`stale-target-waived`, attributed human), and consulted READ-ONLY by the +DKT-193/424/451 advisory judge at `dispatch open`/`verify` and at held +resolutions: a would-be warning whose (instance, target) pair matches a +waiver — the sha compared as a case-insensitive prefix of at least 7 hex +characters, because the advisory renders it at 12 — is dropped from the +answer. There is deliberately no "spent" counter and no per-application event: +the advisory is recomputed by `dispatch verify`, which writes nothing by +contract, and a suppressed warning changes no step's state. The waiver dies +with its run — the `run_id` FK is the whole scope rule, and a new run +re-warns. + +**What it fixes.** The stale-target advisory had no memory: RUN-52 fired the +IDENTICAL adjudicated warning four times across DISPATCH-295/297/301 (the +shared HEAD moves at every integration, so the pair recurs under a different +rendered reason each time), each firing costing an investigation and the +first an operator gate, until the operator issued a standing waiver that +lived only in session memory — where the engine could not see it. A different +target sha on the same row, or the same sha on an unnamed row, is a different +question and still warns, which is what keeps a new divergence from riding an +old ruling. + +**Why the ratified arithmetic is untouched.** Like v11–v24, v25 is an +amendment, not a stage: one additive table, `CREATE TABLE IF NOT EXISTS` +throughout so the migration is idempotent and re-runnable, and a rewind guard +that probes the TABLE (the v24 form, since v25 adds no column). It BACK-FILLS +NOTHING — no operator waived a stale-target warning before the verb for +waiving one existed — and it is dormant: a run that never records a waiver +reads byte-identically to v24 on every verb. + +### AMENDMENT — the span extends to v26 (DKT-1079, 2026-09-02) + +**What changed.** v26 adds ONE table, `run_notes`: a standing statement +whoever drives a run records ONCE against it, and every packet the run +renders from then on carries — for every step of every issue, verbatim, as a +`== RUN NOTE N` section directly after `== REQUEST`, and in `step context` as +`context.notes`. A note is minted by `run note add` (one row, event-logged as +`run-note-added` carrying the text, attributed human) and read by context +assembly as its sixth source, in the claim's own transaction beside the other +five. It is append-only — no edit, no delete — because a packet is the record +of what a worker was told, and two renders of one step that disagreed about +that with nothing in the ledger between them would be exactly the drift the +snapshot discipline exists to prevent; a changed ruling is a second note. The +note dies with its run — the `run_id` FK is the whole scope rule. + +**What it fixes.** A packet's every writable source was step-scoped. RUN-70's +conductor gate-probed before dispatch, found `tests` failing on clean HEAD, +got the operator's disposition ("file issue, override-pass"), filed DKT-1075 +— and had nowhere to put any of it that a packet reads: issue comments are an +audit surface (§6.6), the body froze at activation, and no step had a routing +record for `step resolve -m` to reach. The executor stashed, re-ran the +suite, re-derived the failure, and filed DKT-1076, a duplicate the conductor +then spent three more calls closing. + +**Why the ratified arithmetic is untouched.** Like v11–v25, v26 is an +amendment, not a stage: one additive table, `CREATE TABLE IF NOT EXISTS` +throughout so the migration is idempotent and re-runnable, and a rewind guard +that probes the TABLE (the v24/v25 form, since v26 adds no column). It +BACK-FILLS NOTHING — no dispatcher recorded a note before the verb for +recording one existed — and it is dormant: a run that never records a note +reads byte-identically to v25 on every verb, every bundle, and every packet +(`notes` is `omitempty`, and the template's section renders only over a +non-empty list). + +### AMENDMENT — the span extends to v27 (DKT-1279, 2026-09-03) + +**What changed.** v27 adds ONE column on `steps`, `last_claim_end TEXT NOT +NULL DEFAULT ''`: the END REASON of the MOST RECENT claim to leave this step, +`'failed'` or `'reaped'` (`db.ClaimEndFailed` / `db.ClaimEndReaped`), +overwritten by whichever of `MarkStepAttemptFailedTx` / `MarkStepClaimReapedTx` +ran last — the same two write sites v23's counters already use, now also +stamping this column. It rides the wire as `prior_attempt_end` on +`model.StepRow` (`next --run`, `dispatch open`, `step show`, `claim`'s +`context.step`) and on `StepListEntry` (`step list`), beside +`failed_attempts`/`reaped_claims`, `omitempty` on both. + +**What it fixes.** v23's counters answer "how many of each has this step EVER +had", and that is the wrong question the moment a step's history mixes both: +RUN-80 DISPATCH-400 was killed mid-wave by a session usage limit, the engine +reaped ten leases, the steps re-dispatched at `attempt` incremented, and +wave.js/policy's `on_failure` escalation read that as "failed once" and +routed all ten a tier up to opus/xhigh — a reap is a liveness event, not a +quality verdict, and no field on the row named which ending THIS re-offer +followed. A router does not want a tally; it wants the answer to "was the +attempt I am about to hop past a measured failure or a silence", and the +existing breakdown cannot give that answer once a step has failed once and +been reaped once, in either order — both readings are consistent with +`failed_attempts=1, reaped_claims=1`. + +**Why the ratified arithmetic is untouched.** Like v11–v26, v27 is an +amendment, not a stage: one additive column with a default, a +`hasColumn`-probed `ALTER` so the migration is idempotent and re-runnable, +and a rewind guard that probes the COLUMN (the v21–v23 form, since v27 adds no +table). It BACK-FILLS NOTHING, for v23's own reason: the event log holds +`step-failed` and `lease-reaped` rows a value could be derived from, but +events are prunable and a back-fill would assert more than the store can +promise. An empty string on a pre-v27 claim means "no recorded ending", the +same never-captured honesty v23's zero counters use, and the wire fields are +`omitempty`, so every row with no claim yet ended serializes byte-identically +to v26's rendering. The in-memory snapshot `reapOneTx` reflects mid-transaction +now carries the same field alongside the counter it already bumps, so W3's +same-call offer (the reap and its own re-offer, one answer) carries the +ending it just recorded rather than requiring a second read. + +### AMENDMENT — the span extends to v28 (DKT-2447, 2026-09-13) + +**What changed.** v28 adds ONE column on `proposals`, `sealed INTEGER NOT NULL +DEFAULT 0`: the RENDERING RULE the proposal was opened under. A proposal +opened under a vote rule with `vote.rule..sealed = true` (resolved in +`resolveVoteRule`, stored by `OpenVoteProposal`), or by `vote create --sealed`, +carries `1`. While such a proposal is still `open`, `vote show`, `vote result` +and `gate status` withhold every cast's verdict, confidence, relevance, +weight, findings and summary and render only who has cast and how many casts +are in; once the status leaves `open`, everything renders. It rides the wire +as `sealed` on the proposal object (`vote show`, `vote result`, `vote list`, +`vote create`, export/import), and the withheld casts ride as +`{voter_name, created_at}` entries in the same `votes` array. + +**What it fixes.** Every open proposal rendered every recorded cast in full, +so a seat that read the proposal after a sibling had cast saw the sibling's +verdict and reasoning before casting its own — the public-board channel for +anchoring and collusion that Anthropic's "Patterns and problems in emerging +multiagent systems" names (DOC-130 candidate 1). Sealing is OPT-IN per rule, +so every pre-existing rule and every pre-existing workflow renders exactly as +before. + +**Why the ratified arithmetic is untouched.** Like v11–v27, v28 is an +amendment, not a stage: one additive column with a default, a +`hasColumn`-probed `ALTER` so the migration is idempotent and re-runnable, +and a rewind guard that probes the COLUMN (the v27 form, since v28 adds no +table). It BACK-FILLS NOTHING: no proposal that opened before the rule +existed was opened sealed, and `0` says exactly that. The flag is stored on +the proposal rather than re-resolved from the rule at read time so an edit +to the rule cannot change what a live ballot renders under the seats +mid-vote. Sealing is a NORM-LEVEL SHIELD, not a security boundary: the tally +(`db.CastVote`) and the one-cast-per-voter constraint never read the column, +and the vote rows stay readable through `export`, `ListAllVotes` and direct +store access — gates-trust §8.3 records the same caveat beside the key. + +### AMENDMENT — the span extends to v29 (DKT-2465, 2026-09-14) + +**What changed.** v29 adds ONE column on `runs`, `conductor_token_hash TEXT` +(nullable, no default): the SHA-256 of the run's CONDUCTOR CAPABILITY. A run's +first activation mints a 256-bit token with `model.MintToken`, stores the hash +and returns the token exactly once (`conductor_token` in the activate +envelope, its own line in human mode); `docket run conduct RUN-N` re-mints it, +retiring any standing one, and records a `conductor-seated` event carrying +`actor`, `cwd` and `rotated`. The seven operator verbs — `step approve`, +`step reject`, `step resolve`, `step reap`, `run pause`, `run resume`, +`run abandon` (with or without `--issue`) — require the token on a bound run, +read from `DOCKET_TOKEN` or stdin exactly as a lease token is: none supplied +is the R1 `VALIDATION_ERROR`, a wrong one the R3 `AUTH_ERROR`, both checked +after the step or run is found and before its status is inspected or anything +is written. + +**What it fixes.** The seven verbs were token-free by design — "the authority +is repository access" — and under a harness every executor a wave spawns +shares the operator's checkout, filesystem and environment, so repository +access resolved to "any executor": an executor could approve the very +`commit-gate` its own `git commit` is guarded on, reject a sibling's gate, +resolve a parked step, or park or end the run it was working in, and the +engine had no field on which to tell a conductor's ruling from an executor's +(DKT-2450's `actor`/`cwd` are self-reported audit fields, not authorization). +The only guard was a harness hook covering one verb for one harness's +callers. + +**Why the ratified arithmetic is untouched.** Like v11–v28, v29 is an +amendment, not a stage: one additive nullable column, a `hasColumn`-probed +`ALTER` so the migration is idempotent and re-runnable, and a rewind guard +that probes the COLUMN (the v27/v28 form, since v29 adds no table). It +BACK-FILLS NOTHING, and it cannot: a capability is returned once to whoever +minted it, and a migration has nobody to return one to. `NULL` therefore +means UNBOUND, and an unbound run — every run activated before v29 and not +since conducted — stays open to the seven verbs exactly as before: the check +binds the moment a capability exists, the same posture as a guard reached +with no engine (allow, not deny). Every run this binary activates is bound at +birth. The mechanism is TAMPER-EVIDENT rather than tamper-proof, and that is +recorded beside it: `run conduct` is necessarily token-free (nothing +authenticates a caller, and a run whose conductor died must not be +un-pausable forever), so a caller that takes the seat retires the standing +token — the displaced conductor's next ruling refuses `AUTH_ERROR` and the +`conductor-seated` event names the taker. A harness that keys its callers +keeps executors off that one verb; the engine keeps them off the other seven. + +### AMENDMENT — the span extends to v30 (DKT-1796, 2026-09-15) + +**What changed.** v30 adds ONE column to each of two tables, +`gate_results.fingerprint` and `gate_override_grants.fingerprint` +(`TEXT NOT NULL DEFAULT ''`): the CONTENT half of a gate failure's signature. +**The batch override grant signature is now (gate, exit, reason, +fingerprint)** — `grantMatches` compares all four, and a grant carrying an +EMPTY fingerprint matches nothing at all. Empty means pre-v30 and nothing +else — a row recorded at v30 or later always carries a value, since a gate +that printed nothing hashes the empty capture — so that refusal reaches +grants minted before the column existed and leaves an `unmatched` park, whose +process never ran and whose capture is empty by nature, coverable exactly as +v24 intended. Every recorded +`gate_results` row carries a fingerprint, computed at record time in +`recordGateRows` (the single write path) and exposed as `fingerprint` by +`docket step gates STEP-N --json`. A `--batch` grant COPIES the fingerprint off +the parked step's own failing row rather than recomputing it, so the ruling +binds to the content the operator read. Both ledger edges name the signature: +`gate-override-granted` carries `# fp=<12 hex>` and +`step-batch-overridden` carries ` fp=<12 hex per grant>`, the +fingerprint appended after a space so the id list each payload already carried +is byte-identical to what a reader splitting on `#` or `,` saw before. + +**The fingerprint covers the stored capture, not the whole stream.** +`internal/exec` keeps only the first `CaptureCap` bytes of a gate's output and +sets the row's `truncated` flag when it dropped the rest, so a truncated row's +fingerprint names its head alone. Two failures that differ only past the cap +share a fingerprint. `step resolve --as override-pass --batch` therefore +refuses the WHOLE batch with a `VALIDATION_ERROR` naming the gate when ANY +failing completion row on the step is truncated. The refusal fires before the +resolution's transaction opens, so it mints no grant for any gate in the batch +and records no resolution: the step stays parked, and the operator resolves it +without `--batch`. That ruling decides only this step. + +**The normalization rules**, applied in this order before the SHA-256, and +nothing else: + +1. ANSI CSI and OSC escape sequences are removed. +2. Carriage returns are dropped and trailing spaces and tabs are stripped from + each line, so CRLF and a progress redraw hash as their plain form. +3. RFC3339-ish timestamps and bare `HH:MM:SS(.fff)` clock times become + `