diff --git a/CHANGELOG.md b/CHANGELOG.md index 812794622..05d8b8474 100644 --- a/CHANGELOG.md +++ b/CHANGELOG.md @@ -9,7 +9,7 @@ - **The problem.** A pull request that added a Cursor plugin's `mcp.json`, removed a `beforeShellExecution` guard from `.cursor/hooks.json`, gave a dotfiles package's `claude/.claude/settings.json` `Bash(*)`, or moved a marketplace plugin's pinned `sha` printed `No static host-grant changes detected`, as a docs-only change does. Re-running a 23-PR public corpus after #812 found 11 of 23 pull requests were such coverage gaps: 0 of the 9 comparable zero-row results named the changed relevant file, and 4 of them named a file the pull request did not touch while omitting the one it did. - **What is named.** `diff`, `verify` and the manifest-free PR comment list, under `What this run established`, each path in the comparison's own changed-file set that a bounded, documented candidate rule recognises and no reader of this entry read: `mcp.json` in a plugin directory, a plugin manifest's `mcpServers`, a Codex, Cursor or Copilot manifest's `hooks` and the hook files it names, a manifest or marketplace that does not parse, `.cursor/hooks.json`, host settings below the repository root, and an external marketplace plugin source — `plugins/demo/mcp.json (cursor): added, not read by this entry: MCP configuration in a plugin directory; no row, and loading is not established`. An external source names what it now points at, redacted, and is never fetched. The block's first line says the list includes them. Ordinary documentation, an unrelated `*.json` and an unchanged candidate name nothing. - **What it is not.** Never a row, a widening, a `check` violation or a claim that a host loads the file. Nothing is fetched or run, only plugin manifests and marketplaces are read, and at most 32 candidates are examined; the rest, and any whose rule needed a file that was not read or did not parse, are counted as not examined, on a line that names both causes. The rules are listed in `docs/host-boundary-support.md` under *Changed inputs named but not read*. - - **JSON.** A `changed_not_read` coverage item with its `candidate` rule, ranked right after the blocking limits and inside the existing cap; `read_sources_only` is `false` while one is named, and `unread_candidates` / `unread_candidates_not_examined` say whether the change set was examined. Verifier `0.20` → `0.21`, capability diff `0.3` → `0.4`, runtime contract 40 → 41; host-grants does not move for it (#819, below, moves it to `0.7` in the same contract), and `minimum_control_contract_version` stays `21`. A `0.20` verifier reads with the search not recorded. + - **JSON.** A `changed_not_read` coverage item with its `candidate` rule, ranked right after the blocking limits and inside the existing cap; `read_sources_only` is `false` while one is named, and `unread_candidates` / `unread_candidates_not_examined` say whether the change set was examined. Verifier `0.20` → `0.21`, capability diff `0.3` → `0.4`, runtime contract 40 → 41; host-grants does not move for it (#819 and #823, below, move it to `0.7` in the same contract), and `minimum_control_contract_version` stays `21`. A `0.20` verifier reads with the search not recorded. - **One route moves, on `verify` and `verify --preview` alike.** A manifest-free `verify` whose only host-relevant change is such an input, or a changed candidate it counts as not examined, now publishes the host comparison (advisory, exit `0`) instead of the setup route, which said nothing about the change. `verify --preview` moves the same way: its next action is now `discover` (`audit --host`) with the comparison published, where it was `initialize` (`init --write`) with none. That includes an agent-related workspace, as it already did when the change edited a host file this entry reads. The pilot ledger's source-tree column was re-measured for contract 41. Rows, digests, baselines, `audit --host`, `check` and the benchmark replays are unchanged. - A hook row now names what changed in the hook, and an MCP row names a change to the server's launch arguments. Before, `diff`, `verify`, the manifest-free PR comment and `check` printed `PostToolUse → PostToolUse` whether the edit was to the hook's matcher, its command or its timeout, and an MCP server whose version pin moved from `example-mcp-server@1.2.3` to `@latest` read `docs: no difference in the command name npx, env key names or header key names; the change is in a detail this output does not show, such as the command's path or arguments`: the grants carried none of it, and only `config_sha256` saw the edit. On five of 23 public pull requests measured on 2026-09-15, the hook rows showed only event names. (#819, slice 2 of #795; direction is #820, an unpinned-launch note #825) - **The rows a reviewer reads:** `PostToolUse: matcher Edit → Edit|Write|Bash`, `PostToolUse: command changed (lint.sh sha256:d075f5f4772e → curl sha256:a510416cbecc)`, `PostToolUse: timeout 10 → 600` and `docs: package example-mcp-server@1.2.3 → example-mcp-server@latest`, in `diff`, `verify` text, the PR comment and `check` text, and in `review.changes[].change` in `diff --json` and `verifier.json`; any other launch argument edit reads `launch arguments changed (sha256:… → sha256:…)`, a digest printed as its first twelve hex digits. A timeout written as text prints quoted, so it never reads as a number or a boolean: `timeout 5 → "5"`, `timeout true → "true"`. With several handlers under one event the entry names which one (`handler 2 timeout 5 → 50`), and an added or removed handler is listed as such. The same published handlers in another order read `the published handlers in a different order; a detail this output does not show may also differ, such as …`, never that they are the same handlers. An added or removed hook names its handlers, `SessionEnd (command cleanup.sh sha256:18d2c7ec39bc)`, and an added MCP server its package. When none of the published fields differ, the entry says the change is in a detail it does not show — another hook setting such as `async`, or a redacted or shortened matcher or timeout; for an MCP server, the command's path or another setting such as `cwd` — instead of repeating the same values. The row's direction, severity, `why` and loading basis are unchanged: plugin-selected (#714) and Codex hooks read as before, and no entry claims a direction (#820), runtime loading or what a command does. @@ -20,6 +20,8 @@ - **Unchanged:** grant equality and every inventory digest leave the new members out, so a change is a row exactly when it was one before, through `config_sha256`; a value the digest's own input redacts (after `--token`, `--api-key` or `--password`, a `--password=…` value, an `X-Api-Key:` header value, a URL's path) moves no published digest, so a change confined to it is no row, as on 1.1.0; every row value, the row count, `check`'s boundary result and the control envelope's `capability_rows` publish what they did; verifier `0.21` and capability diff `0.4` do not move for it; the host-config and cold-start benchmark replays reproduce their run-of-record scores. The digest's credential-assignment rule gained a lookahead that removes its quadratic time on a long run of name characters and matches exactly what it matched, so every `config_sha256` is unchanged. Re-running `diff --json` on the 80 vendored benchmark cases with the prepared `1.1.0` commit and this tree on 2026-09-23 gave byte-identical rows on all 80; 23 entries on 22 cases changed, and each of the 7 changed hook or MCP entries (six repositories; one is vendored in both benchmarks) that read `PreToolUse → PreToolUse` or `no difference in the command name …` now names the field that changed, such as `mcp-outline: package mcp-outline==1.10.0 → mcp-outline==1.10.1` or `PreToolUse: handler 2 timeout 30 → 120`. - **Compatibility:** a `0.6` baseline stays comparable with no new row or reason, and `audit --host --save-baseline` may now replace it; an older one is still refused, as before. Validators pinned to the `0.6` schemas reject a `0.7` inventory, baseline or drift payload; the `0.6` files stay published. See the [migration note](STABILITY.md#hook-mcp-detail-fields-819). +- Read how a coding agent is launched inside a CI workflow. A documented agent action's permission inputs (`anthropics/claude-code-action`, `anthropics/claude-code-base-action`, `openai/codex-action`), the permission flags of a `run:` that is one plain `claude -p` or `codex exec` command, and each `actions/checkout` step's `with.ref` are listed on the workflow grant and compared as text, never executed. Changing plain `claude_args` from `--allowedTools Read` to `--permission-mode bypassPermissions --allowedTools Bash`, adding a `claude -p --permission-mode acceptEdits Summarize` step, or checking out `${{ github.event.pull_request.head.sha }}` in a `pull_request_target` job each gave no row and now gives one naming `job/step` and both values in `diff`, `check`, manifest-free `verify` and the PR comment. Only a documented rule a job's launches gain widens — bypassed permission checks (a flag, or JSON settings whose `defaultMode` is `bypassPermissions`), a bypassed or `danger-full-access` sandbox (`permission-profile: :danger-full-access` included), `safety-strategy: unsafe`, or a user gate opened to `*` — and every other edit is `changed`; a workflow row that runs an agent now names the untrusted-input trigger, write scopes, secrets and pull request checkout beside each agent step, so a move to `issue_comment` with `pull-requests: write` says which agent step it reaches. Shell is not parsed: a `run:` is read only when it is one line of plain words (no quote, expansion, operator, redirection, comment or continuation) run by `bash` or `sh` (a `shell:` template only when it runs the script alone, never `bash -c '…' {0}`), and `claude_args` / `codex-args` only when they are a plain list of words with no `--settings` or `--mcp-config` flag. Any other `run:` that mentions `claude` or `codex` is listed in `unread_agent_runs` and named as a non-blocking limit in `audit --host`; it publishes none of its text, is never compared and gives no row, and the workflow's line in `What this run established` names what its grant does not read instead of env values and `apiKeyHelper`. Any other argument input, the #823 reproduction's quoted `--allowedTools "Read"` included, is `unread_arguments`: compared by a digest, so editing it is a `changed` row, and read for no rule. The rules each launch meets are read from the declared text and published as `widening_rules`, so redaction never hides one; one a launch already met in a job it left (a renamed job, a moved step) is moved, not gained, though never while that job keeps any unread step of that agent or a launch of it with an unread input the rule is read from, nor while any other job but the receiving one holds more of them than before, so renaming a job while quoting its launch, as another job adds that launch plainly, is a widening; one the job's unread step or unread input may already have met is named, not claimed; and an unread step that remains takes no gain from another launch. A launch that becomes one this audit does not read (`npx`, quoting, `codex` options before `exec`) is worded as no longer declaring a launch this audit reads, never as no longer starting an agent. A JSON object in a `settings` or `mcp_config` input publishes its shape and none of its free text: key names, with `env` and `headers` values and `apiKeyHelper` redacted and every other string a `` digest, except the strings a host reader publishes (a permission rule, a documented setting's value, an MCP server's command name and URL host), so an MCP server's arguments and a hook's command are compared but never published, and a value there that is neither a JSON object nor a plain file path publishes only a digest; a codex `--config` override publishes its key and, except for the sandbox, permission profile, approval policy and model, only a digest or `` for its value; a URL elsewhere publishes its scheme and host; other credential-shaped text in a setting, prose such as "never print bearer tokens" included, is published redacted, compared as published and named as a non-blocking limit, while a redacted checkout ref refuses as a redacted step reference does. Every published value goes through the #802 label redaction. Host-grants inventory, baseline and drift move to `0.7`, because `0.6` shipped in 1.1.0, and the unreleased runtime contract 41 (#821) is extended in place; a `0.4`–`0.6` baseline holding a workflow grant is incomparable (`baseline_workflow_agent_launches_unavailable`), and the verifier `0.21` and capability diff `0.4` #821 minted do not move here. See the `Migration Note: Unreleased` entry in [`STABILITY.md`](STABILITY.md#workflow-agent-launches-contract-v41-823). (#823) + - A plugin directory that cannot be compared no longer hides the host changes outside it. (#808) - **The problem.** A pull request that broke `plugins/demo/.claude-plugin/plugin.json` and also dropped a `deny` rule from `.claude/settings.json` printed `Cannot compare against main: head_inventory_incomplete` and no row on `diff`, `verify` and the manifest-free PR comment, where the published `1.0.0` showed the removed denial. The plugin-reference limit #714 introduced refused the whole comparison, including files that plugin cannot reach. - **What changes.** When every blocking limit that refused a comparison is a plugin-reference limit bounded by its plugin directory — a reference is followed only inside it, so that is all it can hide — and nothing outside the directory depends on it, the comparison is `partial`: the directory is left uncompared on both sides and named, and the rows outside it are published. `diff` opens with `Partial comparison against main (…) -> working tree: head_inventory_incomplete` and `Not compared: plugins/demo, a plugin directory this entry could not read completely, …` before any row; `verify` and the PR comment open with `Host capability comparison partial: …` and the same line. A partial result with no row says it is not a no-change answer and never prints `No static host-grant changes detected`. diff --git a/STABILITY.md b/STABILITY.md index 041edec56..8a39402f9 100644 --- a/STABILITY.md +++ b/STABILITY.md @@ -18,6 +18,39 @@ workspace too. `minimum_control_contract_version` stays `21`. See [the migration note](#unread-changed-inputs-821). +Also in unreleased runtime contract v41, extended in place: the host inventory +reads how a coding agent is launched inside a workflow job (#823). +Host-grants `0.6` shipped in 1.1.0, so host-grants inventory, baseline and +drift schemas move to `0.7`: a workflow grant adds `agent_launches[]` — a +documented agent action's permission inputs, or the permission flags of a +`run:` that is one line of plain words running `claude -p` or `codex exec`, +compared as text and never executed — `unread_agent_runs[]`, each other +`run:` that mentions an agent CLI, and `checkout_refs[]`, each +`actions/checkout` step's `with.ref`. Shell is not parsed: only a plain list +of words is read, as a `run:` command or as `claude_args` / `codex-args`, and +any other form is a named non-blocking limit that publishes none of its text — +an unread `run:` is never compared and gives no row, and an unread argument +input is compared by a digest and read for no rule. Only a documented rule a +job's launches gain widens (`workflow_agent_widened_`); every +other edit is a `changed` row naming `job/step`, and a workflow row that runs +an agent ends with the job facts beside each agent step. The rules each launch +meets are read from its declared text and published as `widening_rules`; a +rule a launch already met in a job it left is moved, not gained, and one the +job's unread step or unread input may already have met is named, not claimed. +A JSON object in a `settings` or `mcp_config` input publishes its shape and +none of its free text: key names, with each string a `` digest +except those a host reader publishes (a permission rule, a documented +setting's value, an MCP server's command name and URL host), so an MCP +server's arguments and a hook's command are compared but never published; a +URL publishes its scheme and host. A setting holding credential-shaped text, +prose included, is published redacted, compared as published and named as a +non-blocking limit, while a checkout ref holding it refuses as a redacted step +reference does. A `0.4`–`0.6` baseline holding a workflow grant is +incomparable (`baseline_workflow_agent_launches_unavailable`); one without a +workflow stays comparable. It moves neither #821's verifier `0.21` nor its +capability diff `0.4`, and `minimum_control_contract_version` stays +`21`. See [the migration note](#workflow-agent-launches-contract-v41-823). + Also in unreleased runtime contract v41: a hook row names what changed in the hook, and an MCP row a change to the server's launch arguments (#819). Host-grants inventory, baseline and drift schemas move to `0.7`: a hook grant @@ -32,7 +65,7 @@ PostToolUse`, a command edit `command changed` with both digests, and a version pin moving to `@latest` is a `package` difference. The members display what `config_sha256` already binds, so grant equality and the inventory digests leave them out: they move no row value, row count, verifier or capability-diff -schema, a `0.6` baseline stays comparable with no new row or reason, and +schema, a `0.6` baseline without workflow grants stays comparable with no new row or reason, and `audit --host --save-baseline` may replace it. A saved baseline holds none of the members. `minimum_control_contract_version` stays `21`. See [the migration note](#hook-mcp-detail-fields-819). @@ -46,7 +79,7 @@ move from `SHIP-HOST-BOUNDARY-CONFIG-PARSE-FAILED` to `SHIP-HOST-BOUNDARY-PERMISSION-WILDCARD-ALLOW`, and `defaultMode: dontAsk` from the wildcard check to `SHIP-HOST-BOUNDARY-PERMISSION-ALLOW-EXPANDED` at `medium`. Setting rows name the setting and value, and `enabledMcpjsonServers` -entries become grants. No schema, member or check id moves. See +entries become grants. See [the migration note](#claude-setting-ratings-827). Also unreleased, and moving no version of its own: `check` and `verify` route @@ -319,6 +352,52 @@ is added. --- + + +## Migration Note: Unreleased — workflow agent launches (host-grants `0.7`, contract v41, #823) + +Host-grants `0.6` shipped in 1.1.0, so this mints host-grants inventory, baseline and drift `0.7` rather than extending it in place, and extends in place the unreleased runtime contract `41` that #821 minted ([its note](#unread-changed-inputs-821)). The `0.6` schema files stay published and unchanged. A workflow grant adds three members, each present only when a step declares one; in a `0.7` grant their absence means the steps were read and declare none: + +```json +{ + "agent_launches": [ + { + "job": "review", "step": "steps[1]", "agent": "anthropics/claude-code-action", + "form": "read", "unresolved_reason": null, + "settings": [ + {"name": "claude_args", "value": "--permission-mode bypassPermissions --allowedTools Bash", "unresolved_reason": null} + ], + "widening_rules": [{"rule": "bypass_permissions", "setting": "claude_args"}], + "job_secrets": ["CLAUDE_CODE_OAUTH_TOKEN"] + } + ], + "unread_agent_runs": [ + {"job": "review", "step": "steps[2]", "agent": "claude"} + ], + "checkout_refs": [ + {"job": "review", "step": "steps[0]", "ref": "${{ github.event.pull_request.head.sha }}", "unresolved_reason": null} + ] +} +``` + +- **Shell is not parsed.** A value is read only as a plain list of words, which every parser involved splits at its blanks and nowhere else; every other form is named as a non-blocking limit, publishes none of its text and is never guessed at. +- **What is read.** A step whose `uses:` is `anthropics/claude-code-action`, `anthropics/claude-code-base-action` (also published as `anthropics/claude-code-action/base-action`) or `openai/codex-action` (at any ref, in any letter case) is an agent launch listing the documented inputs it sets. A `run:` is an agent launch only when it is one line of plain words (letters, digits and `_ . / : = , % + -`, separated by spaces or tabs), run by `bash`, `sh` or no declared `shell:` (a template only as `bash` or `sh` running the script alone, option words it still runs the script under then `{0}` last, so never one such as `bash -c '…' {0}` that may run a command of its own, or one with `-s`, `-n` or `--rcfile`), whose program, after any `NAME=value` assignments (skipped, never published), has the file name `claude` and passes `-p`/`--print`, or is `codex` followed by `exec` (`codex e`); it lists its documented permission flags under their primary spelling, and a `--settings` or `--mcp-config` flag leaves it unread. Any other `run:` that mentions `claude` or `codex` as a word of its own is one `unread_agent_runs` entry per agent CLI it mentions, holding only `job`, `step` and `agent`. Each `actions/checkout` step lists its `with.ref`, `null` for the default. The support page has [the input and flag tables](docs/host-boundary-support.md#known-unread-surfaces). Values are text: no action is fetched, no command run, no expression evaluated. `job_secrets` names the secrets the launch's job references and the workflow's `env` passes; it is context for the row and is not compared. `job`, `step`, values and secret names are published labels (#802). +- **Argument inputs.** `claude_args` and `codex-args` are read only when they are a plain list of words (the characters above and parentheses, separated by blanks or newlines, with no `--settings` or `--mcp-config` flag), and published as those words one space apart, so reformatting one is quiet. Any other value — a quote, a `${{ }}` expression, `$`, a backtick, a backslash, a `#` comment, a shell separator or redirection, a glob, JSON, a `--settings` or `--mcp-config` flag, any other character — is `unresolved_reason: unread_arguments`: its `value` is ``, a short digest, so an edit is a `changed` row showing the digest, and no documented widening rule is read from it. So a quoted `claude_args` such as the #823 reproduction's `--allowedTools "Read"` is compared only by its digest. +- **What is compared.** Each job's multiset of launches (`job`, `agent`, `form`, `unresolved_reason`, `settings` by `name`, `value` and `unresolved_reason`, and `widening_rules`) and of checkout refs (`job`, `ref`, `unresolved_reason`), never the step label, so a rename or reorder is quiet. An added, removed or changed entry is one `changed` row on the workflow naming `job/step` and the value on each side. Of a CLI launch only the documented permission flags and their values are compared, so `--model` and any undocumented flag are not; a flag `claude` reads as variadic (`--allowedTools`, `--disallowedTools`, `--add-dir`) takes every following word up to the next word starting with `-`, as the CLI reads it, so a prompt written after one is compared, and published, as one of its values; any other prompt is not. `unread_agent_runs` is never compared: adding, removing or editing an unread step gives no row. +- **What is withheld.** A JSON object in a `settings` or `mcp_config` input publishes, as canonical JSON, its key names, numbers, booleans and `null`, with each string replaced by ``: a short digest of what the host readers digest for that string, so editing it is still a `changed` row while none of its text is published. `env` and `headers` values, `apiKeyHelper` and every secret-named value are `` and not digested, as the host readers redact them, so rotating one is quiet. The strings a host reader publishes are kept: a `permissions.allow`, `ask` or `deny` rule, and the value of a documented Claude Code setting (`defaultMode`, the switches, `enabledMcpjsonServers` entries), as the settings reader publishes them; and an MCP server's command name and its URL's scheme and host, as the MCP reader publishes them, each followed by the digest when it drops something the digest reads (a command's arguments, a URL's query). So `{"mcpServers":{"remote":{"command":"npx","args":["mcp-remote","https://…","--header","Authorization: Bearer …"]}}}` publishes `{"mcpServers":{"remote":{"args":["","","",""],"command":"npx"}}}`, and a hook publishes its event names and no command, as `.mcp.json` and `.claude/settings.json` publish none of them. JSON passed through `claude_args`, `codex-args` or a `run:` is never read, so never published. A codex `--config` override in a plain list of words (`-c`, `--config=`, `-c`, `-c=`) publishes its key, and its value as `` under `env`, `headers` or a secret-named key, as written for `sandbox_mode`, `default_permissions`, `approval_policy` and `model`, and as a `` digest under any other key, such as an MCP server's `command` or `url`. The word after a secret-named word such as `--token` or `password` is ``, as the host readers redact it among an MCP server's arguments, and the value is then published redacted (`redacted`, below). Other argument text — a prompt word, a flag's value — is published as written through the #802 label redaction, except that a URL in it publishes its scheme, host and port, with `` for any path, as an MCP server's URL does (#723), so a change only to such a URL's path is not reported; a `${{ }}` expression is one word while this is decided, so one in a URL's userinfo is withheld with it. Text in a `settings` or `mcp_config` input that starts like JSON and does not parse is withheld whole (`unparsed_json`); one that is neither a JSON object nor a plain file path (path characters, and a `${{ }}` expression only as a plain context reference) publishes only a `` digest. +- **Direction.** Only a documented rule a job's launches gain widens: bypassed permission checks (`--dangerously-skip-permissions` or `--permission-mode bypassPermissions` in a plain `claude_args` or `run:`, the last of a repeated `--permission-mode` counting, as the action and the CLI keep it, or Claude Code settings written as JSON in the `settings` input whose `defaultMode` is `bypassPermissions`, read as the settings reader reads it; one rule), `--dangerously-bypass-approvals-and-sandbox`, a `danger-full-access` sandbox (`sandbox: danger-full-access`, `--sandbox danger-full-access` in any spelling clap reads, `-sdanger-full-access` and `-s=danger-full-access` included, or `permission-profile: :danger-full-access`, Codex's built-in full-access profile; and in a `codex exec` step that passes no `--sandbox`, which takes precedence, a `--config` override — `-c`, `--config=`, `-c` or `-c=` — of `sandbox_mode` to `danger-full-access` or of `default_permissions` to `:danger-full-access`, the last override of a key counting and `default_permissions` outranking `sandbox_mode`, as the CLI resolves them; never such an override in `codex-args`, after which the action appends its own `--sandbox` or `default_permissions` selection), `safety-strategy: unsafe`, or a `*` entry in `allowed_bots` (any bot), `allowed_non_write_users` or `allow-users` (any user). A `settings` path names a file this audit does not read, and meets none. The rules each launch meets are decided from its declared text when the workflow is read, before anything is withheld, and published as `widening_rules` (`rule`, and the `setting` it was read from), so redaction never hides one. A rule is read only from text this audit reads exactly: an argument input holding a `${{ }}` expression is not read at all, a user gate's entries that hold none are read, and a `sandbox`, `permission-profile`, `safety-strategy` or `settings` value holding one meets none; such a setting is published with `holds_expression: true` (omitted otherwise), and a row that changes it says the text the expression reaches is not read. Three gains are named in the `why` and not claimed: where, in the job, a step that may launch that agent in a form this audit does not read (an unread `run:`, or an action whose `with:` is not a mapping) is gone and a launch this audit reads is added, as for a job whose permissions were not explicit, because the added launch may be that step rewritten; where the job's launch held, before, a `${{ }}` expression or an argument input that was not a plain list of words in an input the rule is read from, which may already have met it; and where the rule moved between jobs, because the launch that met it in another job left that job (the job no longer exists, or the same launch now runs elsewhere while the receiving job's own launches of that agent still run there or in the job it left; a job that remains may still run its launch in a form this audit does not read, so a launch that only becomes a named unread step has not left it, nor has one whose job still has any named unread step of that agent, wherever it stands and whether or not it was there before, since an unread step carries no text that tells which launch it is, or a launch of it holding an expression or an unread argument input the rule is read from; nor while any job but the receiving one has more such steps and launches of that agent than it had before, a job new at the head holding any, because the launch may be one of them and the job it left may be that job under a new name, so renaming a job while quoting its launch or running it through `npx`, as another job adds that plain launch, is claimed, and so are two jobs renamed at once while one holds an unread step, in the safe direction), as a step reference moved between jobs adds no scope. An unread step that remains takes no gain from another launch. Otherwise the grant earns `workflow_agent_widened_changed` (or `_added` for a new workflow), the row is `widened` with `expands: true`, its `why` names the rule and step, and the Stop hook announces it. Every other edit is `changed` with `expands: false`, including a tool rule such as `--allowedTools Bash` (rating its reach is #824's), `acceptEdits`, a new plugin, an unread argument input and a head-ref checkout. `access` and `risk` still describe the token and triggers alone. +- **The note on a workflow row.** Whatever the row is about, when its workflow runs an agent its `why` ends with the job facts beside each agent step: an untrusted-input trigger (`issue_comment`, `issues`, `pull_request_target`, `workflow_run`), the job's write scopes, the job's secrets, and a checkout of pull request code in the job. It moves no direction and is not a verdict. An unread step is not an agent step here. A removed workflow gets none. +- **Unread and unreadable values are a named limit, not a blocking one.** An unread `run:` step, an argument input that is not a plain list of words (`unread_arguments`), an agent action whose `with:` is not a mapping (`form: unresolved`, `inputs_not_a_mapping`), a setting that is not a string (`not_a_string`) or holds text that starts like JSON and does not parse (`unparsed_json`), a setting holding credential-shaped text (`redacted`), and a ref that is not a string each record a non-blocking `unsupported` coverage issue naming its `job/step`, as an unread secret value does (#693): GitHub coverage stays `complete`. An unread `run:` step is never compared, so it gives no row whatever is edited, and it never says the step starts or does not start an agent. For the others, adding, removing or re-forming the entry, or its gaining a rule, is still a row; only an edit inside it that gains no rule is not reported (an unread argument input's edit is a `changed` row by its digest). `diff`, `verify` and `check` carry no limit for any of them: a change that only adds an unread step prints `No static host-grant changes detected`, and `audit --host` names the step. The workflow's coverage line then ends `so no row (text this entry does not read, such as a step's env or an unread agent step, is not compared; audit --host names each unread agent step)` instead of the redacted-values note another file's line carries. +- **Credential-shaped text.** Other text the #802 label redaction rewrites — a token shape, a credential assignment, a bearer or header value, a URL's userinfo, and prose such as "never print bearer tokens" in a system prompt — is published redacted with `unresolved_reason: redacted`. In a setting it is compared as published, beside the rules read from its declared text, and named by the non-blocking limit above, so a permission change or a rule gained beside it is still a row and only an edit inside what is redacted is not reported. A checkout ref names the code a job runs, so a redacted one refuses as a redacted step reference does (#767): a changed workflow's comparison is refused, and an unchanged one is named in `unchanged_limits`. +- **What is not read.** An action outside the table, a composite action (#701), a script the step runs, an agent CLI reached through a variable or a function, and a step's `env:` and `if:`. The support page lists them under Known unread surfaces. A launch this audit read that becomes one of them, or an unread `run:`, is a row saying the step no longer declares an agent launch this audit reads, never that it no longer starts an agent. The one exception is a rule that moved between jobs (Direction, above): when another job adds the same launch while this job keeps no named unread step of that agent and no launch of it whose input the rule is read from this audit did not read, and no job but the receiving one holds more of them than it did, the row says the launch moved there, as a step reference moved between jobs does, so a launch that became a script or an action outside the table in the same change reads as moved; while the job keeps such a step, or any job but the receiving one gains one, it never does. + +**Compatibility.** +- **A committed `0.4`, `0.5` or `0.6` baseline holding a workflow grant** is loaded but incomparable: it never read agent launches or checkout refs, so its silence is not evidence that none changed. `audit --host --drift` reports `comparison_status: incomparable` with `baseline_workflow_agent_launches_unavailable` among `incomparable_reasons` (beside the #771 and #693 reasons for a `0.4`/`0.5` one), `has_drift: null` and `next_action: null`, and exits `20` under `--fail-on-drift`; `preflight` raises a `high`, `actor: human` `host_grant_drift` signal naming it. To migrate, follow [the #771 steps](#workflow-step-action-references-contract-v40-771) from a checkout of the reviewed default branch, keeping the old file as `host-grants.v0.6.json`: review `audit --host`, move the baseline aside, `audit --host --save-baseline`, and confirm drift is comparable with `has_drift: false`. +- **A `0.4`–`0.6` baseline with no workflow grant** stays comparable for drift. `audit --host --save-baseline` may replace a `0.6` baseline with no workflow grant, preserving #819’s display-only compatibility. A `0.6` baseline holding a workflow grant, and every older baseline, is refused with `unsupported_baseline_schema`; move it aside after review and re-save. +- **Git-backed `diff`, `check` and manifest-free `verify`** read both refs with the current reader and need no migration. Their rows keep their shape, and this change moves neither verifier `0.21` nor capability diff `0.4`, which are #821's. What changes is values: a workflow row can now be `widened` for an agent launch, its `before`/`after` cells list changed launches and checkout refs, and the `why` of every workflow row whose workflow runs an agent gains the note, so a consumer that matches `why` text exactly sees new text. No check id is added or removed, and `check` decides as before. +- **Validators pinned to the `0.6` schemas** reject a `0.7` inventory, baseline or drift payload. +- **`minimum_control_contract_version`** stays `21`. + ## Migration Note: Unreleased — the changed inputs a host comparison does not read (verifier `0.21`, capability diff `0.4`, contract v41, #821) @@ -367,7 +446,7 @@ command, verdict, reader, row or control state is added. **One route moves, on `verify` and `verify --preview` alike.** `verify` without a `shipgate.yaml` returned to the setup route (`Shipgate config not found`, exit `2`) whenever neither side of the comparison held a host artifact, and that route says nothing about the change. A comparison that read no artifact but names a changed input this entry does not read, or counts one or more changed candidate inputs as not examined (`unread_candidates_not_examined` above `0`, the one place that change is mentioned), is now published instead, on the existing manifest-free host route: advisory, exit `0`, `control.state` `agent_action_required` with the `audit --host` next action that route already names. `verify --preview` runs the same comparison and moves the same way: where its next action was `initialize` (`init --write`) with `host_comparison: null`, it is now `discover` (`audit --host`) with the comparison published and the host route's headline; `control.state` stays `agent_action_required` and the exit stays `0`. That includes an agent-related workspace, such as one whose change also adds a tool: a published host comparison takes the preview route whenever one exists, exactly as it already did when the change edits a host file this entry reads, such as the root `.claude/settings.json`. A comparison that reads no artifact, names nothing and counts nothing as not examined still takes the setup route on `verify` and `initialize` on `verify --preview`, as before; so does one whose changed files could not be listed (`unread_candidates: not_examined`), which says nothing about whether a candidate changed. -**What does not change.** `comparison_status`, `incomparable_reasons`, `rows` and every row value, `review`, `unchanged_limits`, every other coverage item, the inventory digests, saved host-grants baselines and drift payloads (#821 moves no host-grants schema; #819, below, moves it to `0.7`), `audit --host`, `check`'s decision, rows and text, the control envelope's `capability_rows`, and every control state, permission and next action on a comparison that reads a host artifact. The host-config and cold-start benchmark replays reproduce their run-of-record scores. `minimum_control_contract_version` stays `21`. +**What does not change.** `comparison_status`, `incomparable_reasons`, `rows` and every row value, `review`, `unchanged_limits`, every other coverage item, the inventory digests, saved host-grants baselines and drift payloads (this change moves no host-grants schema; the unreleased host-grants `0.7` is #823's, [its note](#workflow-agent-launches-contract-v41-823)), `audit --host`, `check`'s decision, rows and text, the control envelope's `capability_rows`, and every control state, permission and next action on a comparison that reads a host artifact. The host-config and cold-start benchmark replays reproduce their run-of-record scores. `minimum_control_contract_version` stays `21`. **Compatibility.** `coverage` and its items are closed objects, so a reader validating against the published [`docs/verifier-schema.v0.20.json`](docs/verifier-schema.v0.20.json) rejects a `0.21` artifact's new members; that schema stays frozen. The current reader reads a `0.20` artifact as `0.21` with `unread_candidates: null`, which is what that build knew, and refuses one that claims a `changed_not_read` item, a `candidate`, `read_sources_only: false` or either `unread_candidates` member. A `diff --json` consumer sees `capability_diff_schema_version: "0.4"`. A consumer switching on `coverage.items[].status` should treat an unknown status as a change it must read, not as no change. @@ -400,12 +479,12 @@ baselines** below): - **What an MCP server publishes.** `package`: the first argument that is a package specification of a strict shape — npm `name@version` or `@scope/name@version`, with a version of two or three numeric parts (optionally with `^` or `~`, a leading `v`, a prerelease or a build) or one of the dist-tags `latest`, `next`, `beta`, `alpha`, `canary`, `rc`, `stable`, `experimental`, `nightly`, `insiders`, `dev` and `preview`; PyPI `name==version`, with a version of two or more numeric parts and extras allowed; or an OCI image reference with a registry or namespace path and a tag or `sha256` digest — that neither the digest's input redaction nor the published-label redaction rewrites, that is at most 200 characters, and that follows no flag but a package runner's own (`-y`, `--yes`, `--package`, `--from`, `--spec`, `-i`, `--interactive`, `--rm`, `--init`, `-q`, `--quiet`; not `uvx --with`, whose value is an extra requirement beside the server); `null` when none is. `args_sha256`: the SHA-256 of the declared `args` as `config_sha256`'s input holds them, the package replaced by a marker and its position digested beside them, so an edit to the package alone moves only `package`, and the package and the digest together determine the arguments even when one of them is a literal marker; `args` that is not a list is digested as declared. Both are `null` when no `args` is declared. `endpoint` is unchanged, still the command's name. - **Shape.** Only the documented hooks shape is read: a list of matcher groups, each an object with a `hooks` list of objects whose `command`, when present, is a string. Anything else publishes `handlers: null`, and its row reads `PostToolUse: matcher, command and timeout not shown: the declaration is not a list of matcher groups whose hooks are objects`. When only one side is outside the shape, the row names that side and lists the other side's handlers as an added hook's are: a change that brings a declaration into the shape reads `PostToolUse: base matcher, command and timeout not shown (the declaration is not a list of matcher groups whose hooks are objects); head (matcher Edit; command a.sh sha256:…)`, and one that takes it out of the shape names `head` and lists `base`. A plugin-selected hook (#714) and a Codex `.codex/hooks.json` hook publish the handlers their file declares and keep their loading basis: `access`, `risk`, the row's `why` and the expansion signal are unchanged. - **Display only.** Every new member is a function of the configuration as `config_sha256`'s input holds it, so it can move only when that digest does. Grant equality and every inventory digest (a baseline's `inventory_sha256`, the drift and comparison digests) leave them out, so a change is a row exactly when it was one before. The digests bind what the digest's input binds: rotating a positional token, a header value's words after its scheme or the value after a flag that input does not name (`--secret-key`) is still a row, which reads `command changed` or `launch arguments changed`. A value the digest's own input already redacts, such as the value after `--token`, `--api-key` or `--password`, a `--password=…` value, an `X-Api-Key:` header value or a URL's path, moves no digest, so a change confined to it is no row, as before. -- **Saved baselines.** `audit --host --save-baseline` writes each grant as comparisons read it, without `handlers`, `omitted_handlers`, `package` or `args_sha256`, in either scope. A baseline is committed ("Commit it"), and a matcher, executable name, digest or package read from `~/.claude/settings.json`, `~/.cursor/mcp.json` or managed settings under `--scope local-static`, or from a git-ignored `.claude/settings.local.json` in either scope, would otherwise carry facts about files that were never in the repository into it. A saved `0.7` baseline's grants are therefore exactly the grants a `0.6` baseline holds, and the `0.7` baseline schema, which forbids the members, differs from `0.6` only in its version. Nothing is lost: no comparison, row or digest reads them from a baseline, `inventory_sha256` is the same with or without them, and `audit --host --drift` against a saved baseline compares as it would have. In that drift payload a changed hook or MCP server's `baseline` side has none of the members; its `current` side has them. A comparison between two commits (`diff`, `check`, manifest-free `verify`) reads both sides fresh, saves nothing, and renders both sides' detail. +- **Saved baselines.** `audit --host --save-baseline` writes each grant as comparisons read it, without `handlers`, `omitted_handlers`, `package` or `args_sha256`, in either scope. A baseline is committed ("Commit it"), and a matcher, executable name, digest or package read from `~/.claude/settings.json`, `~/.cursor/mcp.json` or managed settings under `--scope local-static`, or from a git-ignored `.claude/settings.local.json` in either scope, would otherwise carry facts about files that were never in the repository into it. A saved `0.7` baseline retains the `0.6` hook and MCP grant shapes while also retaining #823’s workflow-launch evidence. Its schema forbids the display-only members. Nothing is lost: no comparison, row or digest reads them from a baseline, `inventory_sha256` is the same with or without them, and `audit --host --drift` against a saved baseline compares as it would have. In that drift payload a changed hook or MCP server's `baseline` side has none of the members; its `current` side has them. A comparison between two commits (`diff`, `check`, manifest-free `verify`) reads both sides fresh, saves nothing, and renders both sides' detail. - **The rows.** A changed hook names each differing field with its before and after: `PostToolUse: matcher Edit → Edit|Write|Bash`, `PostToolUse: command changed (lint.sh sha256:d075f5f4772e → curl sha256:a510416cbecc)` (a digest printed as its first twelve hex digits), `PostToolUse: timeout 10 → 600`, a timeout written as text quoted so it never reads as a number or a boolean (`timeout 5 → "5"`, `timeout true → "true"`); with several handlers, which one (`PreToolUse: handler 2 timeout 5 → 50`); a handler only one side declares as `+handler (…)` or `-handler (…)`, since nothing establishes which handler another replaced. The same published handlers in another order read `the published handlers in a different order; a detail this output does not show may also differ, such as another hook setting or a redacted or shortened matcher or timeout`: equal published handlers never establish equal handlers. An added or removed hook names its handlers, `SessionEnd (command cleanup.sh sha256:18d2c7ec39bc)`. A changed MCP server adds `package example-mcp-server@1.2.3 → example-mcp-server@latest` or `launch arguments changed (sha256:… → sha256:…)` beside its other published facts, and an added one names its package, `docs (command name npx; package example-mcp-server@2.0.0)`. When no published field differs, a hook reads `no difference in the matcher, command or timeout; the change is in a detail this output does not show, such as another hook setting or a redacted or shortened matcher or timeout` (with more than sixteen handlers, `… or timeout of the first 16 handlers; … such as a handler past the first 16, …`), and a command server `no difference in the command name npx, launch arguments, env key names or header key names; the change is in a detail this output does not show, such as the command's path or another setting`. No entry names a direction (#820). The entry is printed by `diff`, `verify` text, the PR comment and `check` text, and published as `review.changes[].change` in `diff --json` and `verifier.json`. Every row value and the row count are unchanged; `check`'s boundary result and the control envelope's `capability_rows` carry rows alone, as before. - **The PR comment.** Its 6,000-character bound cuts at the first line that does not fit, so long entries could hide the rows after them, the coverage block, the change count, the review question, the reproduction and the advisory. The lines `1.1.0` printed get their room first: the coverage block is given the room the other lines leave with every entry in its shortest form (below), which is at least the room `1.1.0`'s lines left it, and the agent instruction block is chosen on those lines too, so the entries get only what is left. The first of these that fits is printed: every entry whole; every longer entry cut to the widest length of at least 60 characters at which the comment fits, ending in `…`; entries in their shortest form, longest first, down to every one that has one; and that without the line below. An entry's shortest form is, for a field-level difference, the difference cut after the name it opens with (`PreToolUse: …`, `docs: …`), and for an added or removed grant its row's own `before → after` (`(absent) → PreToolUse`); it is printed only where it is the shorter. A permission rule's entry and a joined change have none: they are never shortened. No entry in its shortest form is longer than the one `1.1.0` printed for the same row — a hook's read `PreToolUse → PreToolUse`, an MCP server's its name and at least one difference, an added hook its row — so wherever `1.1.0`'s own lines fit, every one of them is kept, entries aside. One line after the rows, not one per entry, reads ``Some entries are shortened here to fit; `verifier.json` holds each entry whole.``; it is left out only where it does not fit beside every other line. When not even the last fits, the bound cuts the rest, as it cut `1.1.0`'s, and the line naming what was cut is as long as `1.1.0`'s. `verifier.json` and every other route keep every entry whole. A comment written without a readiness report points to `verifier.json` when it omits detail (``- … more human summary detail omitted; see `verifier.json`.``), since that route writes no `report.md`. **Compatibility.** -- **A committed `0.6` baseline** stays comparable. Drift reads its grants without the new members and reports what contract v40 reported, with no new row, expansion signal or incomparable reason. `audit --host --save-baseline` may now replace it and reports `status: updated`, with no move-aside step. A baseline older than `0.6` is still refused with `unsupported_baseline_schema`, as the [#771 note](#workflow-step-action-references-contract-v40-771) describes. +- **A committed `0.6` baseline without workflow grants** stays comparable. #823 requires review and replacement of older workflow baselines, which never read agent launches. Drift reads its grants without the new members and reports what contract v40 reported, with no new row, expansion signal or incomparable reason. `audit --host --save-baseline` may now replace it and reports `status: updated`, with no move-aside step. A baseline older than `0.6` is still refused with `unsupported_baseline_schema`, as the [#771 note](#workflow-step-action-references-contract-v40-771) describes. - **Git-backed `diff`, `check` and manifest-free `verify`** read both sides with the current reader and need no migration. - **Validators pinned to the `0.6` schemas** reject a `0.7` inventory, baseline or drift payload. The `0.6` schema files stay published. - **Verifier `0.21`, capability diff `0.4` (both moved by #821 in the same contract), `shipgate.agent_boundary_result/v3` and `minimum_control_contract_version` `21`** do not move for it. @@ -415,11 +494,12 @@ baselines** below): ## Migration Note: Unreleased — one rating per Claude Code setting (#827) This change moves no version of its own: no schema, member, check id or -`minimum_control_contract_version` moves for it, and every field it changes is -one host-grants `0.6` already carried as shipped in 1.1.0. The capability diff -`0.4`, verifier `0.21` and runtime contract `41` of the unreleased tree are -#821's ([migration note](#unread-changed-inputs-821)), and host-grants `0.7` -is #819's ([migration note](#hook-mcp-detail-fields-819)), not this change's. What moves +`minimum_control_contract_version` moves. Of the unreleased tree's versions, +host-grants `0.7` is #823's +([migration note](#workflow-agent-launches-contract-v41-823)), capability +diff `0.4` and verifier `0.21` are #821's +([migration note](#unread-changed-inputs-821)), and runtime contract `41` +carries both; none is this change's. What moves is the value of existing fields for the Claude Code settings the host inventory publishes as `permission_mode` grants. One table, `core/host_settings.py`, now rates each value, and the grant's `access` and `risk`, a row's `severity`, diff --git a/docs/INDEX.md b/docs/INDEX.md index e07c57b15..0c8019b7c 100644 --- a/docs/INDEX.md +++ b/docs/INDEX.md @@ -118,13 +118,13 @@ repository [`README.md`](../README.md) is the landing page that routes to both. - [`org-evidence-bundle-schema.v2.json`](org-evidence-bundle-schema.v2.json) — JSON Schema for `agents-shipgate org bundle`; compact CI/ledger ingestion artifact over verifier/report/attestation/org/host-grant evidence, not a release verdict - [`registry-schema.v0.4.json`](registry-schema.v0.4.json) — JSON Schema for `agents-shipgate registry query --json`, `registry summary --json`, `registry verify --json`, and `registry report --bypass --json` - [`registry-schema.v0.3.json`](registry-schema.v0.3.json) — frozen v0.3 registry reference -- [`host-grants-inventory-schema.v0.7.json`](host-grants-inventory-schema.v0.7.json) — current typed, redacted, scope-aware host inventory; records the in-tree links a read followed, each workflow step's action reference, and each hook's matcher, executable name, command digest and timeout and each MCP server's package and argument digest +- [`host-grants-inventory-schema.v0.7.json`](host-grants-inventory-schema.v0.7.json) — current typed, redacted, scope-aware host inventory; records the in-tree links a read followed, each workflow step's action reference, and each agent launch and checkout ref a workflow declares, plus hook handlers and MCP package/argument digests - [`host-grants-inventory-schema.v0.6.json`](host-grants-inventory-schema.v0.6.json) — frozen v0.6 reference - [`host-grants-inventory-schema.v0.5.json`](host-grants-inventory-schema.v0.5.json) — frozen v0.5 reference - [`host-grants-inventory-schema.v0.4.json`](host-grants-inventory-schema.v0.4.json) — frozen v0.4 reference - [`host-grants-inventory-schema.v0.2.json`](host-grants-inventory-schema.v0.2.json) — frozen prior reference; no inferred structural comparison - [`host-grants-baseline-schema.v0.7.json`](host-grants-baseline-schema.v0.7.json) — current acknowledged host-grant baseline -- [`host-grants-baseline-schema.v0.6.json`](host-grants-baseline-schema.v0.6.json) — frozen v0.6 reference; still compared by drift, and may be replaced by `--save-baseline` +- [`host-grants-baseline-schema.v0.6.json`](host-grants-baseline-schema.v0.6.json) — frozen v0.6 reference; compared by drift only when it holds no workflow grant - [`host-grants-baseline-schema.v0.5.json`](host-grants-baseline-schema.v0.5.json) — frozen v0.5 reference; compared by drift only when it holds no workflow grant - [`host-grants-baseline-schema.v0.4.json`](host-grants-baseline-schema.v0.4.json) — frozen v0.4 reference; compared by drift only when it holds no workflow grant - [`host-grants-baseline-schema.v0.2.json`](host-grants-baseline-schema.v0.2.json) — frozen prior reference; no inferred structural comparison diff --git a/docs/agent-contract-current.md b/docs/agent-contract-current.md index 5a4963740..00c6e059d 100644 --- a/docs/agent-contract-current.md +++ b/docs/agent-contract-current.md @@ -44,6 +44,53 @@ directory, still refuses its comparison. A `0.20` verifier claiming a partial comparison or a `scope` is refused. See [the migration note](../STABILITY.md#partial-host-comparison-808). +Runtime contract v41, extended in place, also reads how a coding agent is +launched inside a workflow job (#823). Host-grants `0.6` shipped in 1.1.0, so +host-grants inventory, baseline and drift schemas move to `0.7`, and a workflow +grant adds `agent_launches[]`, `unread_agent_runs[]` and `checkout_refs[]`, +each omitted when empty. An agent launch is a step whose `uses:` is a +documented agent action (`anthropics/claude-code-action`, +`anthropics/claude-code-base-action`, `openai/codex-action`) with the +permission inputs it declares, or a `run:` that is one line of plain words +running `claude -p` / `codex exec` under `bash` or `sh`, with its documented +permission flags; its `job`, `step`, `agent`, `form` (`read`, or `unresolved` +with `inputs_not_a_mapping`), `settings[]` (`name`, `value`, +`unresolved_reason`, `holds_expression`), `widening_rules[]` (`rule`, +`setting`) and `job_secrets[]`. Shell is not parsed: any other `run:` that +mentions `claude` or `codex` is an `unread_agent_runs[]` entry (`job`, `step`, +`agent`), a named non-blocking limit that publishes none of its text, is never +compared and gives no row; and `claude_args` / `codex-args` are read only as a +plain list of words, any other value being `unread_arguments`, compared by a +digest and read for no rule. A checkout ref is each `actions/checkout` step's +`with.ref`, `null` for the default. Values are compared as text and never +executed. A JSON object in a `settings` or `mcp_config` input publishes its +shape and none of its free text — key names, with each string a +`` digest except those a host reader publishes (a permission rule, +a documented setting's value, an MCP server's command name and URL host) — so +an MCP server's arguments and a hook's command are compared but never +published; a URL publishes its scheme and host. Only a documented rule a job's +launches gain — bypassed permission checks (a flag, or JSON settings whose +`defaultMode` is `bypassPermissions`), a bypassed or `danger-full-access` +sandbox (`permission-profile: :danger-full-access` included, and in a +`codex exec` step without `--sandbox` a `--config` override of `sandbox_mode` +or `default_permissions` that selects it), `safety-strategy: unsafe`, or a +user gate opened to `*` — raises `workflow_agent_widened_` and +makes the row `widened`. A rule is read only from text this audit reads +exactly, and one a launch already met in a job it left, one where an unread +step of the job became a read launch, or one where the job's launch before +held an expression or an unread argument input the rule is read from, is named +and not claimed. Every other edit is `changed`, and a workflow row that runs an +agent ends its `why` with the job facts beside each agent step. An action whose +`with:` is not a mapping is `unresolved` and a named non-blocking limit; a +setting holding credential-shaped text, prose included, is published redacted, +compared as published and a named non-blocking limit, and a checkout ref +holding it is a blocking limit, as a redacted step reference is. A `0.4`–`0.6` +baseline holding a workflow grant is incomparable +(`baseline_workflow_agent_launches_unavailable`); one without a workflow stays +comparable. It moves neither #821's verifier `0.21` nor its capability diff +`0.4`, and `minimum_control_contract_version` stays `21`. See +[the migration note](../STABILITY.md#workflow-agent-launches-contract-v41-823). + The same unreleased runtime contract v41 also names what changed in a hook and in an MCP server's launch arguments (#819). Host-grants inventory, baseline and drift schemas move to `0.7`: a hook grant adds `handlers[]` (each handler's @@ -57,7 +104,7 @@ row names the changed field, `PostToolUse: matcher Edit → Edit|Write|Bash` or a `package` difference, in the text and in `review.changes[].change`. The members display what `config_sha256` already binds, so grant equality and the inventory digests leave them out: they move no row value, row count, verifier -or capability-diff schema, a `0.6` baseline stays comparable with no new row +or capability-diff schema, a `0.6` baseline without workflow grants stays comparable with no new row or reason, and `minimum_control_contract_version` stays `21`. A saved baseline holds none of the members. See [the migration note](../STABILITY.md#hook-mcp-detail-fields-819). diff --git a/docs/design-partner-pilot-results.md b/docs/design-partner-pilot-results.md index d308757cd..ec7f614bf 100644 --- a/docs/design-partner-pilot-results.md +++ b/docs/design-partner-pilot-results.md @@ -93,22 +93,27 @@ the `disposition` field, and it publishes no `review` or `coverage` block; `1.1.0`'s `review.summary` counts 6 changes from 6 rows, 4 widening. On this fixture `1.1.0` changes what a run says about the rows, not which rows it finds. -#821 then moved this tree's runtime contract to 41, and the source-tree column -was rerun on 2026-09-22 through `./shipgate` on the fixture rebuilt from the -description below, beside the `v1.1.0` release commit (`e3c6cb0c`, runtime -contract 40) run the same way. The two returned identical cells except the -runtime contract, 40 against 41: host-grant inventory schema 0.6, `check` -blocking with four violations and visible coverage, the host-only `init` -handoff with no manifest or workflow written, manifest-free `verify` exiting 0 -with the same six advisory rows, drift naming all four expansion signals, and -`diff` against the fixture base exiting 0 with `comparison_status: comparable` -and the same six rows, four of them widening. The `diff` text is identical -apart from the fixture's commit ids. `diff --json` and `verifier.json` differ -only in their schema versions (capability diff 0.3 against 0.4, verifier 0.20 -against 0.21) and in the members #821 adds to the coverage block: each item's -`candidate` is `null`, and `unread_candidates` is `examined` with none left -unexamined. This fixture changes only the two files the entry reads, so #821 -names nothing on it and `read_sources_only` stays `true`. +#821 then moved this tree's runtime contract to 41, and #823, extending it in +place, the host-grant inventory schema to 0.7. With both, the source-tree +column was rerun on 2026-09-23 through this tree's engine on the fixture +rebuilt from the description below, beside the `v1.1.0` release commit +(`e3c6cb0c`, runtime contract 40) run the same way. The two returned identical +cells except the runtime contract, 40 against 41, and the host-grant inventory +schema, 0.6 against 0.7: `check` blocking with four violations and visible +coverage, its `agent-boundary-json` identical apart from the workspace path, +the host-only `init` handoff with no manifest or workflow written, +manifest-free `verify` exiting 0 with the same six advisory rows, drift naming +all four expansion signals, and `diff` against the fixture base exiting 0 with +`comparison_status: comparable` and the same six rows, four of them widening. +The `diff` text is identical apart from the fixture's commit ids, and the host +inventory apart from its schema version. `diff --json` and `verifier.json` +differ only in their schema versions (capability diff 0.3 against 0.4, +verifier 0.20 against 0.21) and in the members #821 adds to the coverage +block: each item's `candidate` is `null`, and `unread_candidates` is +`examined` with none left unexamined. This fixture changes only the two files +the entry reads and has no workflow, so #821 names nothing on it and +`read_sources_only` stays `true`, and #823, which reads agent launches and +checkout refs in workflows, adds nothing to its grants. #819 then published hook and MCP launch detail and moved the host-grant inventory schema to 0.7 within the same contract. With both in the tree, the @@ -145,7 +150,7 @@ expansion signals. It shipped as a qualified release. | `diff` against the fixture base (Git-backed Route H) | exit 0, `comparable`, 6 rows, 4 widening | not measured | exit 0, `comparable`, the same 6 rows | | Qualification | **none** — advisory channel, no qualification claim | **none** — no adjudicated corpus, nothing signed | not a distributed build | -Every 2026-09-22 rerun reproduced four boundary violations (`block` / +Every rerun on 2026-09-22 and 2026-09-23 reproduced four boundary violations (`block` / `critical`) and visible coverage. `init --write --ci` writes no manifest or workflow and routes the host-only fixture to the read-only audit route. Manifest-free `verify` now succeeds as an advisory comparison: six rows name three added permissions, two removed narrower rules, diff --git a/docs/distribution-surfaces.md b/docs/distribution-surfaces.md index 06ed28881..d4459b271 100644 --- a/docs/distribution-surfaces.md +++ b/docs/distribution-surfaces.md @@ -74,7 +74,7 @@ and this document are checked against each other by | `human_review_request` | `docs/human-review-request.md` | `release_decision_vocabulary` | `test_surface_enumerations_match_the_engine_vocabulary` | One complete-evidence documentation-quality class only; no authority or decision ingestion. | | `human_review_decision` | `docs/human-review-decision.md` | `release_decision_vocabulary` | `test_surface_enumerations_match_the_engine_vocabulary` | Host-neutral read-only evaluator; no GitHub acquisition, persistence or operation authority. | | `github_action` | `action.yml`, `scripts/github_action_outputs.py` | `merge_verdict_vocabulary` | `test_action_input_enumerates_engine_merge_verdicts`, `test_action_output_script_shares_the_engine_merge_verdicts` | The paired `shipgate_wheel`/`shipgate_wheel_sha256` inputs install a caller-supplied local wheel instead of a published version, so that route names no channel and claims no `executable_pin`; it is refused unless both halves are given, and it installs `--no-deps`. `tests/test_action_engine_install.py` proves the refusals. Every `python` the Action starts in the workspace runs with `-P` or as a script path, so a pull request's `pip/` or `agents_shipgate/` package cannot stand in for pip or the engine; the same file executes the install and merge-verdict steps against such a checkout. The `v1.0.0` tag predates that fix; the published `v1.1.0` carries it. | -| `capability_diff` | `src/agents_shipgate/cli/diff.py`, `src/agents_shipgate/core/capability_diff_rows.py`, `src/agents_shipgate/core/host_comparison.py`, `src/agents_shipgate/report/host_comparison.py`, `src/agents_shipgate/core/unread_inputs.py`, `src/agents_shipgate/cli/verify/changed_inputs.py` | — | — | Answers no question the engine answers: it emits no verdict, no release decision and no pin. Every field is read from the drift payload the engine already produces — `risk` is the engine's severity and `expansion_signals` is the engine's word on widening — so there is no second implementation to drift. A `permission_mode` or `sandbox` row names the setting and its value as the file spells it (`enableAllProjectMcpServers: true`, `defaultMode: dontAsk`), recovered from the grant's published value and digest, and a Claude Code setting's `why` is the basis the engine's one setting table (`core/host_settings.py`) records for the value; that table also rates the grant and `check`'s violation, so a row's severity and the violation's risk give one answer (#827, `tests/test_prompt_disabling_settings.py`). `verify`/PR and `check` reuse the host comparator (#684, `tests/test_manifest_free_pr_rows.py`), and the source name each named reusable-workflow secret refers to, also non-widening, with a redacting name or target refused rather than compared, and an unreadable value neither compared nor named on this surface — only the host inventory and `audit --host` name its `job/destination`, as on `1.0.0` (#693, `tests/test_reusable_workflow_secret_mappings.py`); check retains permission-rule argument redaction and its existing local-policy control. Missing comparison evidence never supplies empty comparable rows. Host route only; workflow rows compare effective writes and reusable secret recipients (#685, `tests/test_workflow_capability_diff.py`) and each job's remote step action references, as a non-widening change (#771, `tests/test_workflow_step_action_references.py`); every job id, step label, trigger and scope name those rows print is the label the engine published once where it built the grant, redacted, never re-derived here; `check`'s workflow evidence is derived from the raw declarations, which it still compares, and redacts job and scope names by the same rule; two distinct job ids or triggers in one workflow, or scope names in one `permissions` mapping, that publish alike are refused rather than compared, so while such a workflow exists `check` refuses on every run even when it is unchanged (#802, `tests/test_workflow_label_redaction.py`); artifact-only edits remain separate evidence. Tool-source subjects are #655. Where a partial or experimental surface is byte-identical on both sides, `diff` and `verify` compare the rest and name it in `unchanged_limits`; `check` keeps refusing, because its boundary result cannot carry a limit yet (#721). A surface the reader reaches through an in-tree link it reads through qualifies only when that link, a link with the same text at each link on the way, and the file it lands on, the same blob at the same path, are both unchanged, read from the base's Git tree entries against a commit's or, without following any link, the working tree's; any change to either is treated as before. The same proof decides which shared plugin-reference limits `check` leaves out, so behind such a link `check` compares, and publishes the rows it finds, exactly as for a limit at its own path, and a comparison `partial` only because of such a limit is `comparable` with it in `unchanged_limits` (#822, `tests/test_linked_unchanged_limits.py`). A hook row's `why` states the grant's loading basis, read from its published `source`, `access` and `risk` by the engine's `hook_loading_basis`; only a hook the host loads for this project earns an expansion signal — one a settings layer declares, or one a plugin selects that the repository's project settings enable from an in-repository marketplace — so a declared-only hook, or one a plugin selects without that enablement, is a row and never an expansion, and a removal names no basis (#714). `check` compares without a plugin-reference limit both sides share on an untouched source, which it cannot name and, untouched, does not route; a limit only one side carries makes its comparison incomparable. Those rows are not what routes a change: `check`, and the boundary check a manifest-backed `verify` runs, route a changed hook declaration of a plugin the project settings enable through the existing protected-surface rule, from the plugin hook reader's selection on both compared sides, and count a changed hook file such a plugin selects that the reader does not open as incomplete input; the rows beside either are unchanged (#809, `tests/test_enabled_plugin_hook_routing.py`). A partial clone that never fetched the base's objects is refused as `objects_missing`, exit `2`, never compared and never fetched; the refusal ends with the remediation sentence `verify` reports for the same reason, produced by the same function (#817, `tests/test_capability_diff_partial_clone.py`). The text of `diff`, `verify`, the PR comment and `check` reads the rows through one function, `review_changes`, and adds no row and changes no row value in any JSON projection (#795, `tests/test_host_diff_review_changes.py`): a permission rule is named with its disposition; an MCP server with the command name (never its path) or redacted URL, its package and argument digest (#819) and the env and header key names its grant already publishes, a URL printing only in the engine's sanitized scheme-and-host form and otherwise as `url not shown`, or, when none of those differ, a sentence naming what was compared and that the change is in a detail not shown, such as the command's path or another setting; a hook with each handler field that changed — its group's matcher, its command as its executable's name and digest, its timeout — before and after, a handler only one side declares, or the published handlers in a different order with a detail not shown that may also differ, and past the handler bound the same kind of sentence naming a handler past it, all read from the handlers its host-grants `0.7` grant publishes, which hold no command or argument text, and never re-derived here, and a declaration outside the documented hooks shape named as not shown rather than guessed (#819, `tests/test_hook_mcp_detail_fields.py`); those hook and MCP members display what `config_sha256` already binds, so grant equality and every inventory digest leave them out, a saved baseline holds none of them, and no row, row value, reason, digest or control answer moves; the PR comment gives the lines 1.1.0 printed their room first, the coverage block included, and prints an entry whole when the whole comment fits, otherwise cut to the widest length of at least 60 characters at which it does, or else in its shortest form (a difference cut after its name, an added or removed grant as its row), never longer than the entry 1.1.0 printed, with one line naming `verifier.json`, so no long entry hides a row, the coverage block, the change count, the review question, the reproduction or the advisory that 1.1.0 kept (#819 review, cycles 4 and 6); an allow rule the permission lattice decided another replaced (`widened` or `narrowed`), or the exact rule text that moved between dispositions in one host and source (`moved`), is one entry, never on the routes that redact rule arguments; and `diff` counts entries `from N rows` when one joins rows. Comparable results with entries end with one review question, naming the row count when an entry joins rows, and every result whose comparison names a base commit and a commit or working-tree head — a zero-row result and a refusal included (#812 follow-up, `tests/test_host_comparison_coverage.py`) — ends with the compared commits, the tool version and an `agents-shipgate diff --base ` reproduction, labelled `Inputs:` rather than `Compared:` where the comparison was refused, since that run compared nothing — and a refused comparison publishes no `review` object at all, so those two lines are the only place that run states its provenance, built from the `base_commit` it publishes beside the refusal; `check` and a provided diff print the question alone, and no result without a change asks a question. Every one of those facts is published beside the rows, so a machine consumer reads what a human reads (#795 slice 2, same test file): a row adds `disposition`, the `allow`/`ask`/`deny` list a permission rule is declared under and `null` for any other kind, on every route that publishes rows; and `review` in `diff --json` (capability diff `0.3`) and `host_comparison.review` in `verifier.json` (verifier `0.20`) — one object for one comparison — carry the presented changes, each naming the `row_indexes` it stands for, the `direction` the text uses (`widened`, `narrowed` and `moved` included, which no single row can carry), its cells, its `why` and one `expands`, plus a `summary` of `{rows, changes, widenings}` equal to `diff`'s summary line, the review question and the reproduction command. The block is refused unless its changes stand for every published row exactly once, its counters match and no joined change's two sides read alike, so the routes that redact rule arguments publish their rows alone and never a pair that reads `X → X`; `check`'s boundary result carries rows, with their dispositions, and no block. It is presentation, not a second opinion: it is the one `review_changes` projection the text prints, so the rows, their values, their count and every control answer are what they were. A comparison read back from JSON prints the changes it published, and one whose rows a caller sliced falls back to those rows. Each comparison also says what it established (#812, `tests/test_host_comparison_coverage.py`): `coverage` in `diff --json` (capability diff `0.3`) and `host_comparison.coverage` in `verifier.json` (verifier `0.20`) are the same object, printed as `What this run established` by `diff`, `verify` text and the PR comment. It is read off the grant changes, artifact changes, observed sources and blocking issues the comparator already computed: a file's rows, counting a source inside it (`#profiles.`, `#plugins.`); a file with no row and no artifact change called unchanged (`compared`, `0` rows) only when Git proves its blob identical, as the check `unchanged_limits` uses does, asked privately in one bounded batch and never published, because the artifact digest redacts `env` values and `apiKeyHelper`; a file that changed with no compared grant moving (`changed_without_grant_change`), whose artifact differs only in its digest or whose content Git shows differs while its artifact did not (never a difference a checkout line-ending conversion or a converting attribute explains, and no filter is run), never a plugin manifest or marketplace, a retargeted link or project settings while a hook's loading basis moved, worded as no compared grant changing and never as which fields changed; any other changed file with no row (`changed_without_rows`); a file Git neither proves identical nor shows differs — a provided diff, a link read, a redacted path, a working-tree file a checkout wrote with `CRLF` that Git reports unchanged — as `unchanged_not_proven`, never no change and never a change (#812 review cycle 3); the side that published a source, worded `published by` rather than `read in` for a plugin manifest or marketplace, which is published only while it declares hooks; and on a refused comparison each blocking source and its kind. Outside the bounded candidate rules below, a file no inventory observed is never an item and its absence is no claim, which the block states where it is read — one line under the heading and `read_sources_only` in the JSON — so a true list cannot be taken for the account of the change (#812 follow-up); a source already in `unchanged_limits` is not repeated; the list is capped at ten with `omitted_items`, ordered so what no row shows precedes a file's rows and, among blocking limits, by kind (`unreadable`, `parse_failed`, `unresolved_precedence`, then `unsupported`, `dynamic_source_excluded`, `remote_source_excluded`) — order, not severity, and a ranking of kinds rather than of items, since `unsupported` carries both a file this entry merely does not accept and one whose own text would not parse, so an item behind the count may still be one to repair; total down to every field an item is keyed by, the source name and then its side, limit and status — and counted in text as items not listed, ranked below those listed, and the PR comment lists only what fits in the room its entries, review question, reproduction, advisory, next action and evidence leave, at most 2000 characters, so the block never pushes out a line the comment prints without it (a row list that fills the comment by itself still truncates it, as on `1.0.0`); an instruction file's line carries no redacted-values note; sources are the inventory's redacted paths; `null` means not recorded, which is how a `0.19` verifier reads. It moves no row, reason, digest, baseline, control state or next action, and `check`'s boundary result and text carry none, so neither `check` nor a provided diff asks Git anything for it. The same list names the changed inputs this entry does not read (#821, `tests/test_unread_changed_inputs.py`): capability diff `0.4` and verifier `0.21` add a `changed_not_read` item, with the `candidate` rule that named it, for each path in the comparison's own changed-file set — the committed change, or the working tree's tracked and untracked changes — that a bounded, documented rule set recognises as plausibly agent configuration (`mcp.json` in a plugin directory, a plugin manifest's `mcpServers`, a Codex, Cursor or Copilot manifest's `hooks` and the hook files it names, a manifest or marketplace that does not parse, `.cursor/hooks.json`, host settings below the repository root, a marketplace entry's external `source`) and that no inventory published; a member is named whatever read its file, because no reader reads it. It is named from the path and, for a manifest or marketplace member, its text: nothing is fetched, run or read as a grant, so it is never a row, a widening, a `check` violation or a loading claim, and an external source is described redacted and never fetched. It ranks right after the blocking limits, inside the same cap; `read_sources_only` is `false` while one is named, and the first line says so instead; `unread_candidates` and `unread_candidates_not_examined` say whether the change set was examined and how many candidates were not — past the bound of 32, or because a file the rule needed was not read or did not parse, one count the text names both causes of. A manifest-free `verify` whose only host-relevant change is such an input, or a changed candidate it counts as not examined, publishes the comparison instead of the setup route, and `verify --preview` then names `audit --host` instead of `init --write`, in an agent-related workspace too; a `0.20` verifier reads with the search not recorded. A comparison refused only by plugin-reference limits, each bounded by its plugin directory, that no compared source depends on, is `partial` instead (#808, `tests/test_partial_host_comparison.py`); any other blocking limit it carries must be one both sides share on an unchanged source, named in `unchanged_limits` as on a comparable result. Capability diff `0.4` and verifier `0.21` publish `comparison_status: partial` with the refusal's `incomparable_reasons`, the rows, review and unchanged limits established outside those directories, and each directory (the outermost, where one holds another) as the reserved `coverage.items[].scope` on the `blocking_limit` items it bounds, and never call a changed project settings file without a row `changed_without_grant_change`, since the hooks whose loading basis it decides are not all compared; `diff`, `verify` text and the PR comment lead with `Partial comparison against …` or `Host capability comparison partial: …` and `Not compared: , …` before any entry, and a partial result with no entry is never printed as no change. Independence is read off the reader's reference graph, never off directory names: any other limit that is not unchanged, a reference leaving its plugin, a plugin at the root or holding project settings, a marketplace elsewhere declaring inline hooks for it, or a directory that does not publish as itself refuses as before. It answers no engine question and moves no control: a partial comparison is not comparable, `verify`'s control and route are the refusal's, the control envelope projects it as `incomparable` with no rows, and `check`, whose boundary result cannot name a directory, refuses its comparison and decides exactly as before. A `0.20` verifier claiming a partial comparison or a scope is refused. | +| `capability_diff` | `src/agents_shipgate/cli/diff.py`, `src/agents_shipgate/core/capability_diff_rows.py`, `src/agents_shipgate/core/host_comparison.py`, `src/agents_shipgate/report/host_comparison.py`, `src/agents_shipgate/core/unread_inputs.py`, `src/agents_shipgate/cli/verify/changed_inputs.py` | — | — | Answers no question the engine answers: it emits no verdict, no release decision and no pin. Every field is read from the drift payload the engine already produces — `risk` is the engine's severity and `expansion_signals` is the engine's word on widening — so there is no second implementation to drift. A `permission_mode` or `sandbox` row names the setting and its value as the file spells it (`enableAllProjectMcpServers: true`, `defaultMode: dontAsk`), recovered from the grant's published value and digest, and a Claude Code setting's `why` is the basis the engine's one setting table (`core/host_settings.py`) records for the value; that table also rates the grant and `check`'s violation, so a row's severity and the violation's risk give one answer (#827, `tests/test_prompt_disabling_settings.py`). `verify`/PR and `check` reuse the host comparator (#684, `tests/test_manifest_free_pr_rows.py`), and the source name each named reusable-workflow secret refers to, also non-widening, with a redacting name or target refused rather than compared, and an unreadable value neither compared nor named on this surface — only the host inventory and `audit --host` name its `job/destination`, as on `1.0.0` (#693, `tests/test_reusable_workflow_secret_mappings.py`); check retains argument redaction and its existing local-policy control. Missing comparison evidence never supplies empty comparable rows. Host route only; workflow rows compare effective writes and reusable secret recipients (#685, `tests/test_workflow_capability_diff.py`) and each job's remote step action references, as a non-widening change (#771, `tests/test_workflow_step_action_references.py`); and each job's agent launches — a documented agent action's permission inputs, the permission flags of a `run:` that is one plain `claude -p` / `codex exec` command — and checkout refs, compared as text, as a change unless the job gains a documented widening rule, which the engine names in `expansion_signals` (`workflow_agent_widened_*`) — a rule read only from text the engine reads exactly (no shell is parsed; an argument input that is not a plain list of words is compared by a digest and read for no rule), and a gain the engine does not claim (a rule moved in from a job the launch left, one an unread step of the job rewritten as a read launch may already have met, or one the job's launch held before in an expression or an unread argument input) named in the `why` from the same engine function, never counted — with a note on a workflow row naming the untrusted-input trigger, write scopes, secrets and pull request checkout beside each agent step, read off the grant the engine published and moving no direction; an unread `run:` agent step (never compared, so never a row), an unreadable value or a setting published redacted is named only by the host inventory and `audit --host`, as for an unread secret value, an unread argument input or an unresolved launch is named there and in the `why` of a row reporting its launch, and a checkout ref holding credential-shaped text is refused as a redacting step reference is (#823, `tests/test_workflow_agent_launches.py`); every job id, step label, trigger and scope name those rows print is the label the engine published once where it built the grant, redacted, never re-derived here; `check`'s workflow evidence is derived from the raw declarations, which it still compares, and redacts job and scope names by the same rule; two distinct job ids or triggers in one workflow, or scope names in one `permissions` mapping, that publish alike are refused rather than compared, so while such a workflow exists `check` refuses on every run even when it is unchanged (#802, `tests/test_workflow_label_redaction.py`); artifact-only edits remain separate evidence. Tool-source subjects are #655. Where a partial or experimental surface is byte-identical on both sides, `diff` and `verify` compare the rest and name it in `unchanged_limits`; `check` keeps refusing, because its boundary result cannot carry a limit yet (#721). A surface the reader reaches through an in-tree link it reads through qualifies only when that link, a link with the same text at each link on the way, and the file it lands on, the same blob at the same path, are both unchanged, read from the base's Git tree entries against a commit's or, without following any link, the working tree's; any change to either is treated as before. The same proof decides which shared plugin-reference limits `check` leaves out, so behind such a link `check` compares, and publishes the rows it finds, exactly as for a limit at its own path, and a comparison `partial` only because of such a limit is `comparable` with it in `unchanged_limits` (#822, `tests/test_linked_unchanged_limits.py`). A hook row's `why` states the grant's loading basis, read from its published `source`, `access` and `risk` by the engine's `hook_loading_basis`; only a hook the host loads for this project earns an expansion signal — one a settings layer declares, or one a plugin selects that the repository's project settings enable from an in-repository marketplace — so a declared-only hook, or one a plugin selects without that enablement, is a row and never an expansion, and a removal names no basis (#714). `check` compares without a plugin-reference limit both sides share on an untouched source, which it cannot name and, untouched, does not route; a limit only one side carries makes its comparison incomparable. Those rows are not what routes a change: `check`, and the boundary check a manifest-backed `verify` runs, route a changed hook declaration of a plugin the project settings enable through the existing protected-surface rule, from the plugin hook reader's selection on both compared sides, and count a changed hook file such a plugin selects that the reader does not open as incomplete input; the rows beside either are unchanged (#809, `tests/test_enabled_plugin_hook_routing.py`). A partial clone that never fetched the base's objects is refused as `objects_missing`, exit `2`, never compared and never fetched; the refusal ends with the remediation sentence `verify` reports for the same reason, produced by the same function (#817, `tests/test_capability_diff_partial_clone.py`). The text of `diff`, `verify`, the PR comment and `check` reads the rows through one function, `review_changes`, and adds no row and changes no row value in any JSON projection (#795, `tests/test_host_diff_review_changes.py`): a permission rule is named with its disposition; an MCP server with the command name (never its path) or redacted URL, its package and argument digest (#819) and the env and header key names its grant already publishes, a URL printing only in the engine's sanitized scheme-and-host form and otherwise as `url not shown`, or, when none of those differ, a sentence naming what was compared and that the change is in a detail not shown, such as the command's path or another setting; a hook with each handler field that changed — its group's matcher, its command as its executable's name and digest, its timeout — before and after, a handler only one side declares, or the published handlers in a different order with a detail not shown that may also differ, and past the handler bound the same kind of sentence naming a handler past it, all read from the handlers its host-grants `0.7` grant publishes, which hold no command or argument text, and never re-derived here, and a declaration outside the documented hooks shape named as not shown rather than guessed (#819, `tests/test_hook_mcp_detail_fields.py`); those hook and MCP members display what `config_sha256` already binds, so grant equality and every inventory digest leave them out, a saved baseline holds none of them, and no row, row value, reason, digest or control answer moves; the PR comment gives the lines 1.1.0 printed their room first, the coverage block included, and prints an entry whole when the whole comment fits, otherwise cut to the widest length of at least 60 characters at which it does, or else in its shortest form (a difference cut after its name, an added or removed grant as its row), never longer than the entry 1.1.0 printed, with one line naming `verifier.json`, so no long entry hides a row, the coverage block, the change count, the review question, the reproduction or the advisory that 1.1.0 kept (#819 review, cycles 4 and 6); an allow rule the permission lattice decided another replaced (`widened` or `narrowed`), or the exact rule text that moved between dispositions in one host and source (`moved`), is one entry, never on the routes that redact rule arguments; and `diff` counts entries `from N rows` when one joins rows. Comparable results with entries end with one review question, naming the row count when an entry joins rows, and every result whose comparison names a base commit and a commit or working-tree head — a zero-row result and a refusal included (#812 follow-up, `tests/test_host_comparison_coverage.py`) — ends with the compared commits, the tool version and an `agents-shipgate diff --base ` reproduction, labelled `Inputs:` rather than `Compared:` where the comparison was refused, since that run compared nothing — and a refused comparison publishes no `review` object at all, so those two lines are the only place that run states its provenance, built from the `base_commit` it publishes beside the refusal; `check` and a provided diff print the question alone, and no result without a change asks a question. Every one of those facts is published beside the rows, so a machine consumer reads what a human reads (#795 slice 2, same test file): a row adds `disposition`, the `allow`/`ask`/`deny` list a permission rule is declared under and `null` for any other kind, on every route that publishes rows; and `review` in `diff --json` (capability diff `0.3`) and `host_comparison.review` in `verifier.json` (verifier `0.20`) — one object for one comparison — carry the presented changes, each naming the `row_indexes` it stands for, the `direction` the text uses (`widened`, `narrowed` and `moved` included, which no single row can carry), its cells, its `why` and one `expands`, plus a `summary` of `{rows, changes, widenings}` equal to `diff`'s summary line, the review question and the reproduction command. The block is refused unless its changes stand for every published row exactly once, its counters match and no joined change's two sides read alike, so the routes that redact rule arguments publish their rows alone and never a pair that reads `X → X`; `check`'s boundary result carries rows, with their dispositions, and no block. It is presentation, not a second opinion: it is the one `review_changes` projection the text prints, so the rows, their values, their count and every control answer are what they were. A comparison read back from JSON prints the changes it published, and one whose rows a caller sliced falls back to those rows. Each comparison also says what it established (#812, `tests/test_host_comparison_coverage.py`): `coverage` in `diff --json` (capability diff `0.3`) and `host_comparison.coverage` in `verifier.json` (verifier `0.20`) are the same object, printed as `What this run established` by `diff`, `verify` text and the PR comment. It is read off the grant changes, artifact changes, observed sources and blocking issues the comparator already computed: a file's rows, counting a source inside it (`#profiles.`, `#plugins.`); a file with no row and no artifact change called unchanged (`compared`, `0` rows) only when Git proves its blob identical, as the check `unchanged_limits` uses does, asked privately in one bounded batch and never published, because the artifact digest redacts `env` values and `apiKeyHelper`; a file that changed with no compared grant moving (`changed_without_grant_change`), whose artifact differs only in its digest or whose content Git shows differs while its artifact did not (never a difference a checkout line-ending conversion or a converting attribute explains, and no filter is run), never a plugin manifest or marketplace, a retargeted link or project settings while a hook's loading basis moved, worded as no compared grant changing and never as which fields changed; any other changed file with no row (`changed_without_rows`); a file Git neither proves identical nor shows differs — a provided diff, a link read, a redacted path, a working-tree file a checkout wrote with `CRLF` that Git reports unchanged — as `unchanged_not_proven`, never no change and never a change (#812 review cycle 3); the side that published a source, worded `published by` rather than `read in` for a plugin manifest or marketplace, which is published only while it declares hooks; and on a refused comparison each blocking source and its kind. Outside the bounded candidate rules below, a file no inventory observed is never an item and its absence is no claim, which the block states where it is read — one line under the heading and `read_sources_only` in the JSON — so a true list cannot be taken for the account of the change (#812 follow-up); a source already in `unchanged_limits` is not repeated; the list is capped at ten with `omitted_items`, ordered so what no row shows precedes a file's rows and, among blocking limits, by kind (`unreadable`, `parse_failed`, `unresolved_precedence`, then `unsupported`, `dynamic_source_excluded`, `remote_source_excluded`) — order, not severity, and a ranking of kinds rather than of items, since `unsupported` carries both a file this entry merely does not accept and one whose own text would not parse, so an item behind the count may still be one to repair; total down to every field an item is keyed by, the source name and then its side, limit and status — and counted in text as items not listed, ranked below those listed, and the PR comment lists only what fits in the room its entries, review question, reproduction, advisory, next action and evidence leave, at most 2000 characters, so the block never pushes out a line the comment prints without it (a row list that fills the comment by itself still truncates it, as on `1.0.0`); an instruction file's line carries no redacted-values note; sources are the inventory's redacted paths; `null` means not recorded, which is how a `0.19` verifier reads. It moves no row, reason, digest, baseline, control state or next action, and `check`'s boundary result and text carry none, so neither `check` nor a provided diff asks Git anything for it. The same list names the changed inputs this entry does not read (#821, `tests/test_unread_changed_inputs.py`): capability diff `0.4` and verifier `0.21` add a `changed_not_read` item, with the `candidate` rule that named it, for each path in the comparison's own changed-file set — the committed change, or the working tree's tracked and untracked changes — that a bounded, documented rule set recognises as plausibly agent configuration (`mcp.json` in a plugin directory, a plugin manifest's `mcpServers`, a Codex, Cursor or Copilot manifest's `hooks` and the hook files it names, a manifest or marketplace that does not parse, `.cursor/hooks.json`, host settings below the repository root, a marketplace entry's external `source`) and that no inventory published; a member is named whatever read its file, because no reader reads it. It is named from the path and, for a manifest or marketplace member, its text: nothing is fetched, run or read as a grant, so it is never a row, a widening, a `check` violation or a loading claim, and an external source is described redacted and never fetched. It ranks right after the blocking limits, inside the same cap; `read_sources_only` is `false` while one is named, and the first line says so instead; `unread_candidates` and `unread_candidates_not_examined` say whether the change set was examined and how many candidates were not — past the bound of 32, or because a file the rule needed was not read or did not parse, one count the text names both causes of. A manifest-free `verify` whose only host-relevant change is such an input, or a changed candidate it counts as not examined, publishes the comparison instead of the setup route, and `verify --preview` then names `audit --host` instead of `init --write`, in an agent-related workspace too; a `0.20` verifier reads with the search not recorded. A comparison refused only by plugin-reference limits, each bounded by its plugin directory, that no compared source depends on, is `partial` instead (#808, `tests/test_partial_host_comparison.py`); any other blocking limit it carries must be one both sides share on an unchanged source, named in `unchanged_limits` as on a comparable result. Capability diff `0.4` and verifier `0.21` publish `comparison_status: partial` with the refusal's `incomparable_reasons`, the rows, review and unchanged limits established outside those directories, and each directory (the outermost, where one holds another) as the reserved `coverage.items[].scope` on the `blocking_limit` items it bounds, and never call a changed project settings file without a row `changed_without_grant_change`, since the hooks whose loading basis it decides are not all compared; `diff`, `verify` text and the PR comment lead with `Partial comparison against …` or `Host capability comparison partial: …` and `Not compared: , …` before any entry, and a partial result with no entry is never printed as no change. Independence is read off the reader's reference graph, never off directory names: any other limit that is not unchanged, a reference leaving its plugin, a plugin at the root or holding project settings, a marketplace elsewhere declaring inline hooks for it, or a directory that does not publish as itself refuses as before. It answers no engine question and moves no control: a partial comparison is not comparable, `verify`'s control and route are the refusal's, the control envelope projects it as `incomparable` with no rows, and `check`, whose boundary result cannot name a directory, refuses its comparison and decides exactly as before. A `0.20` verifier claiming a partial comparison or a scope is refused. | | `zero_install_detector` | `tools/shipgate-detect.py` | `agent_project_verdict` | `test_detector_verdict_matches_cli` | Emits no `diagnostics[]` and no `next_actions[]`; evidence strings and framework scores are simplified. See the script's own "Intentional simplifications". | | `emitted_ci_workflow` | `src/agents_shipgate/cli/discovery/ci_workflow.py` | `executable_pin` | `tests/test_adopter_pins_resolve.py::test_the_emitted_workflow_pins_the_release_and_not_the_source_tree`, `tests/test_release_source.py::test_candidate_workflow_uses_immutable_source_before_and_after_publication` | Ordinary/source/preview builds use the published fallback; a stamped candidate pins its verified Action SHA and package version. Before publication its smoke substitutes the exact local wheel inputs. Provenance asserts no qualification. | | `prompts` | `prompts/` | `contract_floor`, `executable_pin`, `placeholder_ownership`, `release_decision_vocabulary` | `test_executable_pin_resolves_in_a_published_channel`, `test_surface_enumerations_match_the_engine_vocabulary`, `test_surface_routes_human_owned_placeholders_to_a_human`, `tests/test_adopter_pins_resolve.py::test_every_pin_init_writes_into_an_adopter_repo_names_the_published_release`, `tests/test_adopter_pins_resolve.py::test_the_shipped_floor_is_decided_against_the_release_the_prompts_pin` | — | diff --git a/docs/host-boundary-support.md b/docs/host-boundary-support.md index 1cddf97bb..44bbb2183 100644 --- a/docs/host-boundary-support.md +++ b/docs/host-boundary-support.md @@ -19,7 +19,7 @@ and `audit --host`. | Claude Code | first-class | `.claude/settings.json`, `.claude/settings.local.json`, `.mcp.json`, `CLAUDE.md`, Claude skills | permission modes/rules, sandbox/network, additional paths, MCP restrictions, plugins and their marketplaces (`extraKnownMarketplaces`), hooks | | Cursor | first-class | `.cursor/cli.json`, `.cursor/mcp.json`, `.cursor/rules/**` | Shell/Read/Write rules, MCP declarations, instruction trust roots | | VS Code MCP | first-class | `.vscode/mcp.json` | MCP servers; `sandbox` and per-server `sandboxEnabled`; `${input:…}` references by name, never value; `envFile` recorded as a limit; other top-level keys partial | -| Shared/GitHub | first-class | `AGENTS.md`, Shipgate policies/state, skills, `.github/workflows/*` | instruction/gate weakening, workflow permissions and triggers, remote step action references, named secret sources passed to reusable workflows | +| Shared/GitHub | first-class | `AGENTS.md`, Shipgate policies/state, skills, `.github/workflows/*` | instruction/gate weakening, workflow permissions and triggers, remote step action references, named secret sources passed to reusable workflows, agent launches (documented agent action inputs, `run:` steps that are one plain `claude -p` / `codex exec` command, `run:` steps that mention an agent CLI in shell this audit does not parse, named as a limit) and `actions/checkout` refs | A registered adapter reports `complete`, `not_applicable`, `partial`, or `experimental` coverage. A relevant malformed, unreadable, binary, oversized, @@ -69,6 +69,28 @@ the changed inputs the candidate rules at the end of this section name (#821): them while that subagent runs; no adapter reads the file. A skill's `hooks` frontmatter is type-checked with the skill's instructions, never read as a hook grant, so its events get no hook row (#714). +- **An agent launched any way the workflow reader below does not read** + (#823): an action outside its table, even one that takes `claude_args`; a + composite action (#701); a script the step runs; a `run:` that is not one + line of plain words running `claude -p` or `codex exec` under `bash` or `sh` + (more than one line or command, quoting, an expansion, a redirection, a + comment, a here-doc, `npx`, `timeout`, `sudo`, `bash -c`, `codex` with an + option before `exec`); an agent CLI reached through a variable or a + function; and a step's `env:` and `if:`. + A `run:` among these that mentions `claude` or `codex` as a word of its own + is named as a non-blocking limit (below); editing, adding or removing it + gives no row. The rest give no row and name no limit. A launch this audit + read that becomes one of these is a row saying the step no longer declares + an agent launch this audit reads, and that it may still start one this way; + it never says the step no longer starts an agent. The one exception is a + rule that moved between jobs (below): when another job adds the same launch + while this job keeps no named unread step of that agent and no launch of it + whose input the rule is read from this audit did not read, and no job but + the one adding it holds more of them than it did, the row says the launch + moved there, as a step reference moved between jobs does, so a launch that + became a script or an action outside the table in the same change reads as + moved; while the job keeps such a step, or any job but the receiving one + gains one, it never does. Review changes to those files and fields as you would a change to the workflow, hook or server entry that holds them. @@ -122,6 +144,291 @@ reusable `uses:` target, containing credential-shaped text is published redacted and refuses the same way a step reference does, so two values that redact alike never compare as unchanged. +How a coding agent is launched inside a job is read (#823). Every value is +compared as the text it declares, less what the host readers withhold (below); +no action is fetched, no command is run and no expression is evaluated. Shell +is not parsed: a value is read only in a form every parser involved reads the +same way, a plain list of words, and every other form is named as a limit and +never guessed at. Four things are listed on the workflow grant, each naming its +`job/step` (the step's `id`, else its `name`, else `steps[N]`): + +- **A documented agent action** — a step whose `uses:` is one of these + `owner/repo` references, at any ref and in any letter case — with the inputs + it declares from this table. Other inputs, such as `prompt` or an API key, + are not listed. + + | Action | Inputs compared as text | Documented widening | + |---|---|---| + | `anthropics/claude-code-action` | `additional_permissions`, `allowed_bots`, `allowed_non_write_users`, `claude_args`, `plugin_marketplaces`, `plugins`, `settings`, and the earlier `allowed_tools`, `disallowed_tools`, `mcp_config` | a plain `claude_args` gains `--dangerously-skip-permissions` or `--permission-mode bypassPermissions` (the last `--permission-mode` counting); `settings`, written as JSON, gains `defaultMode: bypassPermissions` (under `permissions`, else at the top, as the settings reader reads `.claude/settings.json`); `allowed_bots` (any bot) or `allowed_non_write_users` (any user) gains a `*` entry | + | `anthropics/claude-code-base-action`, also published as `anthropics/claude-code-action/base-action` | `claude_args`, `plugin_marketplaces`, `plugins`, `settings`, `allowed_tools`, `disallowed_tools`, `mcp_config` | `claude_args` and `settings` as above | + | `openai/codex-action` | `allow-bot-users`, `allow-bots`, `allow-users`, `codex-args`, `permission-profile`, `safety-strategy`, `sandbox` | `sandbox` becomes `danger-full-access`; `permission-profile` becomes `:danger-full-access`, Codex's reserved name for its built-in full-access profile; `safety-strategy` becomes `unsafe`; a plain `codex-args` gains `--dangerously-bypass-approvals-and-sandbox` (`--yolo`) or `--sandbox danger-full-access` (`-s`, attached or not); `allow-users` gains a `*` entry. A sandbox `--config` override in `codex-args` meets none: after `codex-args` the action appends its own `--sandbox`, or its own `default_permissions` override for a `permission-profile`, which takes precedence | + + `claude_args` and `codex-args` are read only when they are a **plain list + of words**: words made of letters, digits and `_ . / : = , % + - ( )`, + separated by blanks or newlines, with no `--settings` or `--mcp-config` flag + in any spelling. The Claude actions (`base-action/src/parse-sdk-options.ts`, + shell-quote with `()|&;<>` made literal) and `openai/codex-action` + (string-argv) both split such text at its blanks and nowhere else, so it is + published as those words, one space apart — `claude_args: |` on several + lines reads as it would on one, and reformatting it is quiet — and, as the + Claude actions read it, a word starting with `--` is always a flag, never + another flag's value. Any other + value — holding a quote, a `${{ }}` expression, `$`, a backtick, a + backslash, a `#` comment, `;`, `&`, `|`, `<`, `>`, a glob, JSON, a + `--settings` or `--mcp-config` flag, or any other character — is **not + read** (`unresolved_reason: unread_arguments`): none of its text is + published, its `value` is ``, a short digest, so an edit to it + is a `changed` row whose cell shows the digest, no documented widening rule + is read from it, and it is a non-blocking limit (below). So the + reproduction's `--allowedTools "Read"` → + `--permission-mode bypassPermissions --allowedTools "Bash(*)"` is a + `changed` row saying the input is not read, while the same change written + `--allowedTools Read` → `--permission-mode bypassPermissions --allowedTools Bash` + is `widened`. + +- **An agent CLI launched by a `run:`** — only when the whole `run:` is **one + line of plain words**: letters, digits and `_ . / : = , % + -`, separated by + spaces or tabs, so it holds no quote, `$`, backtick, backslash, `#`, `;`, + `&`, `|`, `<`, `>`, parenthesis, brace, glob, `~`, `!`, `@` or second line, + and every POSIX shell runs it as exactly those words. It must run under + `bash`, `sh` or no declared `shell:` (the step's, else its job's or its + workflow's `defaults.run.shell`); a template is read only as `bash` or `sh` + running the script alone — option words it still runs the script under + (`set` flags, `-l`, `-i`, `-r`, `-o`/`-O` with an option name other than + `noexec`, `--noprofile`, `--norc`, `--posix`, `--login`, `--restricted`, + `--noediting`, `--verbose`), then `{0}` last, as in + `bash --noprofile --norc -eo pipefail {0}` — so one such as + `bash -c '…' {0}`, which may run a command of its own, or one with `-s`, + `-n` or `--rcfile`, is not. After any `NAME=value` assignments, which + are skipped and never published, the program's file name must be `claude` + with `-p`/`--print` among its arguments, or `codex` followed by `exec` + (`codex e`): `claude -p …`, `./node_modules/.bin/claude -p …` and + `CI=1 codex exec …` are read. Its documented permission flags are listed + under their primary spelling, and every other word — the prompt, `--model`, + an undocumented flag — is not compared. For `claude`: `--permission-mode`, + `--dangerously-skip-permissions`, `--allow-dangerously-skip-permissions`, + `--allowedTools`/`--allowed-tools`, `--disallowedTools`/`--disallowed-tools`, + `--add-dir` and `--permission-prompt-tool`; gaining + `--dangerously-skip-permissions` or `--permission-mode bypassPermissions` + (one rule, so moving between the spellings is not a widening; of a repeated + `--permission-mode`, the last counts, as the CLI keeps it) widens, and a + `--settings` or `--mcp-config` flag makes the step unread. For `codex exec`: + `--sandbox`/`-s`, `--dangerously-bypass-approvals-and-sandbox`/`--yolo`, + `--approve-for-me`/`--not-so-yolo`, `--dangerously-bypass-hook-trust`, + `--add-dir`, `--config`/`-c` and `--profile`/`-p`, a short flag's value + read attached as clap reads it (`-sdanger-full-access`, `-s=…`, + `-c`, `-c=`); gaining + `--dangerously-bypass-approvals-and-sandbox` or the full-access sandbox + widens. The full-access sandbox is `--sandbox danger-full-access` or, when + the command passes no `--sandbox`, which takes precedence, a `--config` + override in any of its four spellings that sets `sandbox_mode` to + `danger-full-access` (the setting `--sandbox` sets) or `default_permissions` + to `:danger-full-access` (the built-in full-access profile, which the + action's `permission-profile` input passes the CLI the same way). The last + override of a key counts, and a `default_permissions` override outranks a + `sandbox_mode` one. A key under another table, such as + `profiles..sandbox_mode`, and a `--profile`, which names a + configuration this audit does not read, meet none. A flag the CLI reads as + variadic (`--allowedTools`, `--add-dir`, …) takes every following word up + to the next word starting with `-`, as the CLI reads it, so a prompt + written after it is compared as one of its values; the row shows it. + +- **An unread agent step** — any other `run:` that mentions `claude` or + `codex` as a word of its own: more than one line or command, quoting, an + expansion or a `${{ }}` expression, a redirection or a here-doc, a comment, + a line continuation, another program such as `npx`, `timeout`, `sudo` or + `echo`, a subcommand that is not a headless launch (`claude mcp add`, + `codex login`), `codex` with an option before `exec`, a script named after + an agent, or a declared `shell:` other than `bash` or `sh` running the + script alone. It is listed in + `unread_agent_runs` once for each agent CLI it mentions, with no other + field, and is a non-blocking limit (below). Nothing more: none of its text + is published, it is never compared, so adding, removing or editing it gives + no row, and it never says that the step starts, or does not start, an agent. + `npm ci && claude -p --dangerously-skip-permissions "Review"`, a + `# Don't run this on forks` comment above a launch, a PR comment drafted in + a here-doc that mentions `claude -p` and `gh pr comment --body "$(claude -p …)"` + are all unread steps. + +- **Each `actions/checkout` step's `with.ref`**, or the default when it + declares none or an empty one. Adding `ref: + ${{ github.event.pull_request.head.sha }}` is a `changed` row naming the + step on both sides. + +Each job's multiset of launches and of checkout refs is compared, so renaming +or reordering steps is quiet. An added, removed or changed launch or ref is a +`changed` row on the workflow naming `job/step` and the value on each side. +Direction is claimed only by the documented rules above. The rules each launch +meets are decided when the workflow is read, from the declared text before +anything is withheld for publication, and published on the launch as +`widening_rules` (each rule and the setting it was read from), so redaction +never hides one. + +Only text this audit reads exactly meets a rule. GitHub substitutes a `${{ }}` +expression into an input before the action reads it, and the substituted text +may be anything, so an argument input holding one is not read at all; a user +gate's entries that hold no expression are read, since the substituted text +may add entries but cannot remove a literal one, so `"${{ vars.USERS }}, *"` +opens the gate; and a `sandbox`, `permission-profile`, `safety-strategy` or +`settings` value holding one meets none. Such a setting is published with +`holds_expression: true`, and a row that changes it says the text the +expression reaches is not read for a rule, rather than that none was gained. + +When a job's launches gain a rule, the workflow earns +`workflow_agent_widened_`, the row is `widened` and its `why` +names the rule and step. Three gains are named in the `why` and not claimed: + +- where, in the same job, a step that may launch that agent in a form this + audit does not read (an unread agent step, or an action whose `with:` is + not a mapping) is gone and a launch this audit reads is added, because the + added launch may be that step rewritten, which may already have met the + rule — so rewriting `npm ci && claude -p --dangerously-skip-permissions "Review"` + as two plain steps is not a widening. An unread step that remains takes no + gain from another launch: beside it, a new plain + `claude -p --dangerously-skip-permissions Review` step is a widening; +- where the job's launch of that agent held, before, text this audit did not + read for the rule in an input the rule is read from (`claude_args` or + `settings` for bypassed permission checks, `sandbox`, `permission-profile` + or `codex-args` for a full-access sandbox, the gate for a `*` entry): a + `${{ }}` expression, or an argument input that was not a plain list of + words, because it may already have met it — so unquoting + `--dangerously-skip-permissions --append-system-prompt "Review"` is not a + widening; +- where the rule moved between jobs: another job met it before and the launch + that met it left that job — the job no longer exists, as when it is renamed, + or the same launch now runs in this job while each launch of that agent this + job had still runs here or in that job, as when an agent step moves or two + jobs swap launches — as a step reference moved between jobs adds no scope. + A second job gaining a rule a first job keeps, or a different launch + gaining it while the first job still exists, is claimed: that job may still + run its launch in a form this audit does not read (`npx`, quoting), so a + launch that only becomes a named unread step has not left it, and a launch + edited in place into the one that job had gains the rule. A launch has not + left a job that still has any named unread step of that agent — wherever it + stands and whether or not it was there before, because an unread step + carries no text that tells which launch it is — or a launch of it holding + an expression or an unread argument input the rule is read from. So quoting + the prompt of `claude -p --dangerously-skip-permissions Review` in one job, + or merging that job's `npm i -g @anthropic-ai/claude-code` step into it, + while another job adds that plain step is a widening, and so is moving that + step to another job while the job it left keeps a `claude mcp add` step. + Nor has it left while any job but the one gaining the rule has more such + steps and launches of that agent than it had before — a job new at the + head holding any — because the launch may be one of them, and the job it + left may be that job under a new name. So renaming the job while quoting + that launch or running it through `npx`, or removing the job while another + job gains it quoted, as a third job adds the plain step, is a widening; + renaming a job with the install step it keeps beside the launch, or beside + another job that keeps its unread step, is a move. Two jobs renamed at + once, one of them holding an unread step, cannot be told apart from those, + so they claim the gain, in the safe direction. + +Any other edit — `--allowedTools Read` to `--allowedTools Bash`, +`acceptEdits`, a new plugin, an argument input this audit does not read, a +head-ref checkout — is `changed`; rating a tool rule's reach is a job for +issue #824. `access` and `risk` still describe the token and triggers alone. + +A workflow row whose workflow runs an agent ends its `why` with the job facts +beside each agent step, whatever else the row is about: an untrusted-input +trigger (`issue_comment`, `issues`, `pull_request_target`, `workflow_run`), the +job's write scopes, the secrets the job references (`${{ secrets.NAME }}` in the +job, or in the workflow's `env`), and a checkout in the job of pull request +code (`github.event.pull_request.head.sha`, `.head.ref` or `.merge_commit_sha`, +`github.head_ref`, `github.event.workflow_run.head_sha` or `.head_branch`, or +`refs/pull//head` and `/merge`). It is a note, not a verdict: it moves no +direction, and `if:` conditions and the default checkout of a `pull_request` +event are not read into it. An unread agent step is not an agent step here. A +removed workflow gets no note. + +A structured input publishes its shape and none of its free text, so a setting +never publishes what `.claude/settings.json` and `.mcp.json` would withhold +(#823 review). A JSON object — a `settings` or `mcp_config` value — publishes, +in canonical JSON: + +- its key names, numbers, booleans and `null`, so reordering keys compares + as unchanged and adding a key is a change; +- `` for `env` and `headers` values (and codex's `http_headers` and + `env_http_headers`), `apiKeyHelper`, every other secret-named value and the + word after a secret-named argument such as `--token`, as the host readers + redact them, so rotating one compares as unchanged; +- each other string as ``, a short digest of what the host + readers digest for it: the sanitized text, and a URL's query. Editing it is + a `changed` row, and none of its text is published. So an MCP server's + `args` — an `mcp-remote --header "Authorization: Bearer …"` included — and a + hook's command, matcher and type publish only digests; +- except the strings a host reader publishes: a `permissions.allow`, `ask` or + `deny` rule, and the value of a documented Claude Code setting + (`defaultMode`, the switches in [the ratings table](#claude-code-setting-ratings), + `enabledMcpjsonServers` entries), as the settings reader publishes them; and, + under `mcpServers` (or `mcp_servers`), a server's command name and its URL's + scheme and host, as the MCP reader publishes them. Each is followed by the + digest when it drops something the digest reads: `"command":"npx "` + for `npx -y some-server`, a URL's digest for its query. A URL's path is + neither published nor compared, as an MCP server's is not (#723). + +A `settings` or `mcp_config` value that is not a JSON object is published as +written only when it is a plain file path: path characters, and any `${{ }}` +expression in it a plain context reference. Any other value, such as a comment +line before the JSON, publishes only ``, a digest, so an edit to it +is still a `changed` row and none of its text is published. + +JSON passed through `claude_args`, `codex-args` or a `run:` is never read, so +it is never published: the argument input or step is not read at all. A codex +`--config` override in a plain list of words — `-c key=value`, +`--config=key=value`, `-ckey=value` or `-c=key=value` — publishes its key; +its value is `` under `env`, `headers` or a secret-named key, so +rotating it is quiet, published as written for `sandbox_mode`, +`default_permissions`, `approval_policy` and `model`, and ``, a +digest, under any other key, such as an MCP server's `command` or `url` or a +`shell_environment_policy` value. The word after a secret-named word such +as `--token` or `password` is ``, as the host readers redact it +among an MCP server's arguments, and the value is then credential-shaped +(below). Other argument text — a prompt word, a +flag's value — is published as written through the label redaction below, +except that a URL in it publishes its scheme, host and port, with +`` for any path, as an MCP server's URL does (#723); the rest +of the setting is compared, so a change only to such a URL's path — which +repository a `plugin_marketplaces` URL names, for one — is not reported, and +a zero-row result says redacted values are not compared. A `${{ }}` +expression is one word while this is decided, so one inside a URL's userinfo +is withheld with it. + +Any other text the #802 label redaction rewrites is credential-shaped: a token +shape, a credential assignment such as `token=…`, a bearer or header value, a +URL's userinfo — and ordinary prose too, such as "never print bearer tokens" in +a system prompt. In a setting, the value is published redacted with +`unresolved_reason: redacted` and compared as published, beside the rules read +from its declared text, so a permission change or a rule gained beside it is +still a row; only an edit inside what is redacted that gains no rule is not +reported, and that is named as a non-blocking limit below. A checkout ref names +the code a job runs, as a step reference does, so a redacted ref refuses the +same way a step reference does (#767), because two refs that redact alike +cannot be compared apart: GitHub coverage is `partial`, a changed workflow's +comparison is refused, and an unchanged one is named in `unchanged_limits`. + +An unread agent step; an argument input that is not a plain list of words +(`unread_arguments`); an action whose `with:` is not a mapping; a setting that +is not a string (`not_a_string`), that holds text starting like JSON that does +not parse (`unparsed_json`, whose values cannot be told from its keys), or +that is published redacted (`redacted`); and a checkout ref that is not a +string or whose `with:` is not a mapping record a **non-blocking** +`unsupported` coverage issue naming the `job/step`, printed under +`audit --host` → Coverage issues; none but a redacted setting publishes any of +the value's text. GitHub coverage stays complete, so `check`, baselines and +every other row are unaffected. An unread agent step is never compared, so +adding, removing or editing it gives no row; for the others, adding, removing +or re-forming the entry, or its gaining a documented rule, is still a row, and +only an edit inside it that gains no rule is not reported (an unread argument +input's edit is a `changed` row by its digest). `diff`, `verify` and `check` +carry no limit for any of them, as for an unread secret value (#693): a +`diff` of a change that only adds an unread agent step says +`No static host-grant changes detected`, and `audit --host` is where the step +is named. The workflow's coverage line says so: `… compared; changed, but no +grant this entry compares changed, so no row (text this entry does not read, +such as a step's env or an unread agent step, is not compared; audit --host +names each unread agent step)`, where another file's line names redacted +values such as env values and `apiKeyHelper`. + A workflow's labels are published redacted (#802). A job id, a step's `id` or `name`, a trigger and a permission scope name go through the same redaction as step text, and any userinfo after `scheme://` inside one is replaced, so a diff --git a/docs/host-grants-baseline-schema.v0.7.json b/docs/host-grants-baseline-schema.v0.7.json index 839fc5aa5..82115c889 100644 --- a/docs/host-grants-baseline-schema.v0.7.json +++ b/docs/host-grants-baseline-schema.v0.7.json @@ -239,7 +239,6 @@ }, "HostGrantsBaselineV7": { "additionalProperties": false, - "description": "A saved ``0.7`` baseline: the grants a ``0.6`` baseline holds, under the ``0.7`` version.\n\nA saved baseline holds no hook ``handlers`` and no MCP ``package`` or\n``args_sha256`` (#819): it is committed, and those members, read from a\nuser, managed or git-ignored file, would carry facts about files that were\nnever in the repository into it. No comparison, row or digest reads a\nsaved copy of them, so its ``inventory`` is the ``0.6`` snapshot, which\nforbids them.", "properties": { "host_grants_schema_version": { "const": "0.7", @@ -248,7 +247,7 @@ "type": "string" }, "inventory": { - "$ref": "#/$defs/HostGrantsNormalizedSnapshotV6" + "$ref": "#/$defs/HostGrantsNormalizedSnapshotV7" }, "inventory_sha256": { "title": "Inventory Sha256", @@ -271,7 +270,7 @@ "title": "HostGrantsBaselineV7", "type": "object" }, - "HostGrantsNormalizedSnapshotV6": { + "HostGrantsNormalizedSnapshotV7": { "additionalProperties": false, "properties": { "artifacts": { @@ -295,7 +294,7 @@ "profile": "#/$defs/HostProfileGrantV2", "requirement": "#/$defs/HostRequirementGrantV2", "sandbox": "#/$defs/HostSandboxGrantV2", - "workflow": "#/$defs/HostWorkflowGrantV6" + "workflow": "#/$defs/HostWorkflowGrantV7" }, "propertyName": "kind" }, @@ -328,7 +327,7 @@ "$ref": "#/$defs/HostRequirementGrantV2" }, { - "$ref": "#/$defs/HostWorkflowGrantV6" + "$ref": "#/$defs/HostWorkflowGrantV7" }, { "$ref": "#/$defs/HostInstructionGrantV2" @@ -357,7 +356,7 @@ "required": [ "scope" ], - "title": "HostGrantsNormalizedSnapshotV6", + "title": "HostGrantsNormalizedSnapshotV7", "type": "object" }, "HostHookGrantV2": { @@ -1283,9 +1282,211 @@ "title": "HostSandboxGrantV2", "type": "object" }, - "HostWorkflowGrantV6": { + "HostWorkflowAgentLaunchV7": { "additionalProperties": false, - "description": "A GitHub workflow's token permissions, triggers, reusable calls and step references.\n\nEvery job id, trigger and permission scope name is a published label\n(#802): credential-shaped text in it is redacted, one way in every field \u2014\n``permission_contexts``, ``reusable_calls``, ``step_actions``, ``triggers``,\nand the job and scope names in the ``write_scopes`` and\n``effective_write_scopes`` entries \u2014 and ``config_sha256`` is computed over\nthose labels. A single redacted label still compares. When two distinct\njob ids or triggers in the workflow, or two scope names in one\n``permissions`` mapping that a job's permissions are read from, publish\nalike, the inventory records a blocking coverage issue instead of comparing\nthem as one. A top-level mapping no job inherits is read only into\n``write_scopes``, which is neither compared nor digested.", + "description": "A step that launches a known coding agent, read as text and never run (#823).\n\n``agent`` is a documented action reference's ``owner/repo`` (the step's\n``uses:`` at any ref; the Claude base action also as the ``base-action``\ndirectory of ``anthropics/claude-code-action``), or a known agent CLI a\n``run:`` launches when the whole ``run:`` is one line of plain words\n(letters, digits and ``_ . / : = , % + -``, separated by spaces or tabs),\nrun by ``bash``, ``sh`` or the runner's default shell, whose program,\nafter any ``NAME=value`` assignments, has the file name ``claude`` and\npasses ``-p``/``--print``, or ``codex`` followed by ``exec`` (``e``).\n``form: read`` lists the documented permission inputs or flags the step\ndeclares in ``settings``, and the documented widening rules they meet in\n``widening_rules``, omitted when none. ``form: unresolved`` is an agent\naction whose ``with:`` is not a mapping (``inputs_not_a_mapping``), with\nno settings, and records a non-blocking coverage issue. Any other\n``run:`` that mentions an agent CLI is not a launch: it is listed in\n``unread_agent_runs``. ``job_secrets`` names the secrets the step's job\nreferences (``${{ secrets.NAME }}``) and the workflow-level ``env``\npasses: context for the row that names this step, never compared.\n``job`` and ``step`` are published labels (#802).", + "properties": { + "agent": { + "enum": [ + "anthropics/claude-code-action", + "anthropics/claude-code-base-action", + "anthropics/claude-code-action/base-action", + "openai/codex-action", + "claude", + "codex" + ], + "title": "Agent", + "type": "string" + }, + "form": { + "enum": [ + "read", + "unresolved" + ], + "title": "Form", + "type": "string" + }, + "job": { + "title": "Job", + "type": "string" + }, + "job_secrets": { + "items": { + "type": "string" + }, + "title": "Job Secrets", + "type": "array" + }, + "settings": { + "items": { + "$ref": "#/$defs/HostWorkflowAgentSettingV7" + }, + "title": "Settings", + "type": "array" + }, + "step": { + "title": "Step", + "type": "string" + }, + "unresolved_reason": { + "anyOf": [ + { + "const": "inputs_not_a_mapping", + "type": "string" + }, + { + "type": "null" + } + ], + "default": null, + "title": "Unresolved Reason" + }, + "widening_rules": { + "items": { + "$ref": "#/$defs/HostWorkflowAgentRuleV7" + }, + "title": "Widening Rules", + "type": "array" + } + }, + "required": [ + "job", + "step", + "agent", + "form" + ], + "title": "HostWorkflowAgentLaunchV7", + "type": "object" + }, + "HostWorkflowAgentRuleV7": { + "additionalProperties": false, + "description": "One documented widening rule an agent launch meets, and the setting it was read from (#823).\n\nDecided when the workflow is read, from the declared text, before any of\nit is withheld for publication, so redaction never hides a rule. Only\ntext this reader reads exactly meets one: ``claude_args`` or\n``codex-args`` only when it is a plain list of words (never when it holds\na ``${{ }}`` expression), the entries of a user gate that hold no\nexpression, and a mode or ``settings`` input that holds none. Claude Code\nsettings written as JSON in the ``settings`` input meet\n``bypass_permissions`` when their ``defaultMode`` is\n``bypassPermissions``, read as the settings reader reads it; a path to a\nsettings file is not read. ``setting`` is the input (``claude_args``,\n``allowed_bots``, ``sandbox``, ``permission-profile``, \u2026) or the CLI\nflag's primary spelling. One rule compares as one whatever setting meets\nit, except ``open_gate``, which is one rule per gate input.", + "properties": { + "rule": { + "enum": [ + "bypass_permissions", + "bypass_approvals_and_sandbox", + "danger_full_access", + "unsafe_safety_strategy", + "open_gate" + ], + "title": "Rule", + "type": "string" + }, + "setting": { + "title": "Setting", + "type": "string" + } + }, + "required": [ + "rule", + "setting" + ], + "title": "HostWorkflowAgentRuleV7", + "type": "object" + }, + "HostWorkflowAgentSettingV7": { + "additionalProperties": false, + "description": "One permission input or flag an agent launch declares, compared as text (#823).\n\n``name`` is the documented input (``claude_args``, ``sandbox``, \u2026) or the\nflag's primary spelling (``--allowedTools`` for ``--allowed-tools`` too).\n``value`` is the declared text, stripped, as it may be published; a flag\nthat takes no value has ``null``.\n\n``claude_args`` and ``codex-args`` are read only when they are a plain\nlist of words: letters, digits and ``_ . / : = , % + - ( )``, separated by\nblanks or newlines, with no ``--settings`` or ``--mcp-config`` flag. Every\nparser involved splits such text the same way, so it is published as\nthose words, one space apart. Any other value \u2014 holding a quote, a\n``${{ }}`` expression, ``$``, a backtick, a comment, a shell operator,\nJSON or another character \u2014 is ``unread_arguments``: ``value`` is\n````, a short digest, so an edit to it is still a change\nwhile none of its text is published; no documented widening rule is read\nfrom it; and it records a non-blocking coverage issue naming its\n``job/step`` (#823 review cycle 4). A codex ``--config`` override keeps\nits key; its value is ```` under ``env``, ``headers`` or a\nsecret-named key, as the host readers redact such values, published as\nwritten for ``sandbox_mode``, ``default_permissions``,\n``approval_policy`` and ``model``, and ```` otherwise.\n\nEvery other input is one value. A JSON object (a ``settings`` or\n``mcp_config`` value) publishes its shape and none of its free text: key\nnames, numbers, booleans and ``null``, with each string replaced by\n````, a short digest of what the host readers digest for it,\nso an edit to it is still a change. ``env`` and ``headers`` values,\n``apiKeyHelper`` and every secret-named value are ````, as the\nhost readers redact them. The strings a host reader publishes are kept:\na ``permissions.allow``/``ask``/``deny`` rule and a documented Claude\nCode setting's value such as ``defaultMode``, and an MCP server's command\nname and its URL's scheme and host, each followed by the digest when it\ndrops something the digest reads (a command's arguments, a URL's query).\nSo an MCP server's arguments and a hook's command publish nothing, as\n`.mcp.json` and `.claude/settings.json` do not (#823 review). A\n``settings`` or ``mcp_config`` value that neither starts like a JSON\nobject nor is a plain file path (path characters, and a ``${{ }}``\nexpression only as a plain context reference) is ````, a\ndigest and none of its text (#823 review cycle 5). A URL in\nother text publishes its scheme and host with ```` for its\npath and query (#723). Other text \u2014 a prompt, a flag's value \u2014 is\npublished through the workflow label redaction (#802). A value it\nrewrites is credential-shaped \u2014 a token, but also prose such as \"never\nprint bearer tokens\" \u2014 and is published redacted with\n``unresolved_reason: redacted``: it is compared as published, beside the\nrules read from its declared text, and records a non-blocking coverage\nissue naming its ``job/step``, because an edit inside what is redacted is\nnot reported. A value that is not a string (``not_a_string``), or one\nholding text that starts like JSON and does not parse (``unparsed_json``),\nis ``null`` and records a non-blocking coverage issue naming its\n``job/step``: it is neither published nor compared.\n\n``holds_expression`` is ``true`` when an input other than an argument\ninput holds a ``${{ }}`` expression, which GitHub substitutes before the\naction reads the input, and is omitted otherwise. A documented widening\nrule is then read only from the entries of a user gate that hold none,\nand from no mode or settings input, and a rule the launch gains in the\nsame job afterwards is not claimed, because the substituted text may\nalready have met it.", + "properties": { + "holds_expression": { + "default": false, + "title": "Holds Expression", + "type": "boolean" + }, + "name": { + "title": "Name", + "type": "string" + }, + "unresolved_reason": { + "anyOf": [ + { + "enum": [ + "not_a_string", + "redacted", + "unparsed_json", + "unread_arguments" + ], + "type": "string" + }, + { + "type": "null" + } + ], + "default": null, + "title": "Unresolved Reason" + }, + "value": { + "anyOf": [ + { + "type": "string" + }, + { + "type": "null" + } + ], + "title": "Value" + } + }, + "required": [ + "name", + "value" + ], + "title": "HostWorkflowAgentSettingV7", + "type": "object" + }, + "HostWorkflowCheckoutRefV7": { + "additionalProperties": false, + "description": "One ``actions/checkout`` step and the ``with.ref`` it declares, as text (#823).\n\n``ref`` is ``null`` when the step declares none, or an empty one: the\ncheckout's default for the triggering event. A ref the label redaction\nrewrites is published redacted with ``unresolved_reason: redacted`` and\nmakes the workflow a blocking limit, as a redacted step reference does\n(#767): a ref names the code the job runs, as a step reference does. A\nvalue that is not a string, or ``with:`` that is not a mapping,\nis ``null`` with ``unresolved_reason`` and records a non-blocking coverage\nissue. The ref is never resolved or fetched.", + "properties": { + "job": { + "title": "Job", + "type": "string" + }, + "ref": { + "anyOf": [ + { + "type": "string" + }, + { + "type": "null" + } + ], + "title": "Ref" + }, + "step": { + "title": "Step", + "type": "string" + }, + "unresolved_reason": { + "anyOf": [ + { + "enum": [ + "not_a_string", + "redacted", + "inputs_not_a_mapping" + ], + "type": "string" + }, + { + "type": "null" + } + ], + "default": null, + "title": "Unresolved Reason" + } + }, + "required": [ + "job", + "step", + "ref" + ], + "title": "HostWorkflowCheckoutRefV7", + "type": "object" + }, + "HostWorkflowGrantV7": { + "additionalProperties": false, + "description": "A v0.6 workflow grant plus the agent launches, unread agent steps and checkout refs its steps declare.\n\nEach list is present only when a step declares one. In a v0.7 grant an\nabsent list means the steps were read and declare none; the schema\nversion, not the key, separates that from a legacy grant that never read\nthem. ``unread_agent_runs`` is a named limit and is never compared.\n``access`` and ``risk`` still describe the workflow's token and triggers\nalone.", "properties": { "access": { "enum": [ @@ -1300,6 +1501,20 @@ "title": "Access", "type": "string" }, + "agent_launches": { + "items": { + "$ref": "#/$defs/HostWorkflowAgentLaunchV7" + }, + "title": "Agent Launches", + "type": "array" + }, + "checkout_refs": { + "items": { + "$ref": "#/$defs/HostWorkflowCheckoutRefV7" + }, + "title": "Checkout Refs", + "type": "array" + }, "config_sha256": { "title": "Config Sha256", "type": "string" @@ -1389,6 +1604,13 @@ "title": "Triggers", "type": "array" }, + "unread_agent_runs": { + "items": { + "$ref": "#/$defs/HostWorkflowUnreadAgentRunV7" + }, + "title": "Unread Agent Runs", + "type": "array" + }, "write_all": { "default": false, "title": "Write All", @@ -1414,7 +1636,7 @@ "effective_write_scopes", "reusable_calls" ], - "title": "HostWorkflowGrantV6", + "title": "HostWorkflowGrantV7", "type": "object" }, "HostWorkflowPermissionsV4": { @@ -1515,6 +1737,35 @@ "title": "HostWorkflowStepActionV6", "type": "object" }, + "HostWorkflowUnreadAgentRunV7": { + "additionalProperties": false, + "description": "A ``run:`` step that mentions a known agent CLI and is not read as an agent launch (#823 review cycle 4).\n\nAny ``run:`` holding ``claude`` or ``codex`` as a word of its own that is\nnot an agent launch this reader reads \u2014 more than one line or command, a\nquote, an expansion, a redirection, a comment, a continuation, a\n``${{ }}`` expression, another program such as ``npx`` or ``timeout``, a\nsubcommand that is not a headless launch, or a declared ``shell:`` other\nthan ``bash`` or ``sh`` run on the script alone (so ``bash -c '\u2026' {0}``\ntoo) \u2014 once for each agent CLI it mentions. It is a\nnamed, non-blocking limit and nothing more: none of the step's text is\npublished, it is never compared, so adding, removing or editing it gives\nno row, and it never says that the step starts, or does not start, an\nagent. ``job`` and ``step`` are published labels (#802).", + "properties": { + "agent": { + "enum": [ + "claude", + "codex" + ], + "title": "Agent", + "type": "string" + }, + "job": { + "title": "Job", + "type": "string" + }, + "step": { + "title": "Step", + "type": "string" + } + }, + "required": [ + "job", + "step", + "agent" + ], + "title": "HostWorkflowUnreadAgentRunV7", + "type": "object" + }, "InstructionStructureEvidence": { "additionalProperties": false, "properties": { diff --git a/docs/host-grants-inventory-schema.v0.7.json b/docs/host-grants-inventory-schema.v0.7.json index 4ea55331a..2b0f960ac 100644 --- a/docs/host-grants-inventory-schema.v0.7.json +++ b/docs/host-grants-inventory-schema.v0.7.json @@ -268,7 +268,7 @@ "profile": "#/$defs/HostProfileGrantV2", "requirement": "#/$defs/HostRequirementGrantV2", "sandbox": "#/$defs/HostSandboxGrantV2", - "workflow": "#/$defs/HostWorkflowGrantV6" + "workflow": "#/$defs/HostWorkflowGrantV7" }, "propertyName": "kind" }, @@ -301,7 +301,7 @@ "$ref": "#/$defs/HostRequirementGrantV2" }, { - "$ref": "#/$defs/HostWorkflowGrantV6" + "$ref": "#/$defs/HostWorkflowGrantV7" }, { "$ref": "#/$defs/HostInstructionGrantV2" @@ -1459,9 +1459,211 @@ "title": "HostSandboxGrantV2", "type": "object" }, - "HostWorkflowGrantV6": { + "HostWorkflowAgentLaunchV7": { "additionalProperties": false, - "description": "A GitHub workflow's token permissions, triggers, reusable calls and step references.\n\nEvery job id, trigger and permission scope name is a published label\n(#802): credential-shaped text in it is redacted, one way in every field \u2014\n``permission_contexts``, ``reusable_calls``, ``step_actions``, ``triggers``,\nand the job and scope names in the ``write_scopes`` and\n``effective_write_scopes`` entries \u2014 and ``config_sha256`` is computed over\nthose labels. A single redacted label still compares. When two distinct\njob ids or triggers in the workflow, or two scope names in one\n``permissions`` mapping that a job's permissions are read from, publish\nalike, the inventory records a blocking coverage issue instead of comparing\nthem as one. A top-level mapping no job inherits is read only into\n``write_scopes``, which is neither compared nor digested.", + "description": "A step that launches a known coding agent, read as text and never run (#823).\n\n``agent`` is a documented action reference's ``owner/repo`` (the step's\n``uses:`` at any ref; the Claude base action also as the ``base-action``\ndirectory of ``anthropics/claude-code-action``), or a known agent CLI a\n``run:`` launches when the whole ``run:`` is one line of plain words\n(letters, digits and ``_ . / : = , % + -``, separated by spaces or tabs),\nrun by ``bash``, ``sh`` or the runner's default shell, whose program,\nafter any ``NAME=value`` assignments, has the file name ``claude`` and\npasses ``-p``/``--print``, or ``codex`` followed by ``exec`` (``e``).\n``form: read`` lists the documented permission inputs or flags the step\ndeclares in ``settings``, and the documented widening rules they meet in\n``widening_rules``, omitted when none. ``form: unresolved`` is an agent\naction whose ``with:`` is not a mapping (``inputs_not_a_mapping``), with\nno settings, and records a non-blocking coverage issue. Any other\n``run:`` that mentions an agent CLI is not a launch: it is listed in\n``unread_agent_runs``. ``job_secrets`` names the secrets the step's job\nreferences (``${{ secrets.NAME }}``) and the workflow-level ``env``\npasses: context for the row that names this step, never compared.\n``job`` and ``step`` are published labels (#802).", + "properties": { + "agent": { + "enum": [ + "anthropics/claude-code-action", + "anthropics/claude-code-base-action", + "anthropics/claude-code-action/base-action", + "openai/codex-action", + "claude", + "codex" + ], + "title": "Agent", + "type": "string" + }, + "form": { + "enum": [ + "read", + "unresolved" + ], + "title": "Form", + "type": "string" + }, + "job": { + "title": "Job", + "type": "string" + }, + "job_secrets": { + "items": { + "type": "string" + }, + "title": "Job Secrets", + "type": "array" + }, + "settings": { + "items": { + "$ref": "#/$defs/HostWorkflowAgentSettingV7" + }, + "title": "Settings", + "type": "array" + }, + "step": { + "title": "Step", + "type": "string" + }, + "unresolved_reason": { + "anyOf": [ + { + "const": "inputs_not_a_mapping", + "type": "string" + }, + { + "type": "null" + } + ], + "default": null, + "title": "Unresolved Reason" + }, + "widening_rules": { + "items": { + "$ref": "#/$defs/HostWorkflowAgentRuleV7" + }, + "title": "Widening Rules", + "type": "array" + } + }, + "required": [ + "job", + "step", + "agent", + "form" + ], + "title": "HostWorkflowAgentLaunchV7", + "type": "object" + }, + "HostWorkflowAgentRuleV7": { + "additionalProperties": false, + "description": "One documented widening rule an agent launch meets, and the setting it was read from (#823).\n\nDecided when the workflow is read, from the declared text, before any of\nit is withheld for publication, so redaction never hides a rule. Only\ntext this reader reads exactly meets one: ``claude_args`` or\n``codex-args`` only when it is a plain list of words (never when it holds\na ``${{ }}`` expression), the entries of a user gate that hold no\nexpression, and a mode or ``settings`` input that holds none. Claude Code\nsettings written as JSON in the ``settings`` input meet\n``bypass_permissions`` when their ``defaultMode`` is\n``bypassPermissions``, read as the settings reader reads it; a path to a\nsettings file is not read. ``setting`` is the input (``claude_args``,\n``allowed_bots``, ``sandbox``, ``permission-profile``, \u2026) or the CLI\nflag's primary spelling. One rule compares as one whatever setting meets\nit, except ``open_gate``, which is one rule per gate input.", + "properties": { + "rule": { + "enum": [ + "bypass_permissions", + "bypass_approvals_and_sandbox", + "danger_full_access", + "unsafe_safety_strategy", + "open_gate" + ], + "title": "Rule", + "type": "string" + }, + "setting": { + "title": "Setting", + "type": "string" + } + }, + "required": [ + "rule", + "setting" + ], + "title": "HostWorkflowAgentRuleV7", + "type": "object" + }, + "HostWorkflowAgentSettingV7": { + "additionalProperties": false, + "description": "One permission input or flag an agent launch declares, compared as text (#823).\n\n``name`` is the documented input (``claude_args``, ``sandbox``, \u2026) or the\nflag's primary spelling (``--allowedTools`` for ``--allowed-tools`` too).\n``value`` is the declared text, stripped, as it may be published; a flag\nthat takes no value has ``null``.\n\n``claude_args`` and ``codex-args`` are read only when they are a plain\nlist of words: letters, digits and ``_ . / : = , % + - ( )``, separated by\nblanks or newlines, with no ``--settings`` or ``--mcp-config`` flag. Every\nparser involved splits such text the same way, so it is published as\nthose words, one space apart. Any other value \u2014 holding a quote, a\n``${{ }}`` expression, ``$``, a backtick, a comment, a shell operator,\nJSON or another character \u2014 is ``unread_arguments``: ``value`` is\n````, a short digest, so an edit to it is still a change\nwhile none of its text is published; no documented widening rule is read\nfrom it; and it records a non-blocking coverage issue naming its\n``job/step`` (#823 review cycle 4). A codex ``--config`` override keeps\nits key; its value is ```` under ``env``, ``headers`` or a\nsecret-named key, as the host readers redact such values, published as\nwritten for ``sandbox_mode``, ``default_permissions``,\n``approval_policy`` and ``model``, and ```` otherwise.\n\nEvery other input is one value. A JSON object (a ``settings`` or\n``mcp_config`` value) publishes its shape and none of its free text: key\nnames, numbers, booleans and ``null``, with each string replaced by\n````, a short digest of what the host readers digest for it,\nso an edit to it is still a change. ``env`` and ``headers`` values,\n``apiKeyHelper`` and every secret-named value are ````, as the\nhost readers redact them. The strings a host reader publishes are kept:\na ``permissions.allow``/``ask``/``deny`` rule and a documented Claude\nCode setting's value such as ``defaultMode``, and an MCP server's command\nname and its URL's scheme and host, each followed by the digest when it\ndrops something the digest reads (a command's arguments, a URL's query).\nSo an MCP server's arguments and a hook's command publish nothing, as\n`.mcp.json` and `.claude/settings.json` do not (#823 review). A\n``settings`` or ``mcp_config`` value that neither starts like a JSON\nobject nor is a plain file path (path characters, and a ``${{ }}``\nexpression only as a plain context reference) is ````, a\ndigest and none of its text (#823 review cycle 5). A URL in\nother text publishes its scheme and host with ```` for its\npath and query (#723). Other text \u2014 a prompt, a flag's value \u2014 is\npublished through the workflow label redaction (#802). A value it\nrewrites is credential-shaped \u2014 a token, but also prose such as \"never\nprint bearer tokens\" \u2014 and is published redacted with\n``unresolved_reason: redacted``: it is compared as published, beside the\nrules read from its declared text, and records a non-blocking coverage\nissue naming its ``job/step``, because an edit inside what is redacted is\nnot reported. A value that is not a string (``not_a_string``), or one\nholding text that starts like JSON and does not parse (``unparsed_json``),\nis ``null`` and records a non-blocking coverage issue naming its\n``job/step``: it is neither published nor compared.\n\n``holds_expression`` is ``true`` when an input other than an argument\ninput holds a ``${{ }}`` expression, which GitHub substitutes before the\naction reads the input, and is omitted otherwise. A documented widening\nrule is then read only from the entries of a user gate that hold none,\nand from no mode or settings input, and a rule the launch gains in the\nsame job afterwards is not claimed, because the substituted text may\nalready have met it.", + "properties": { + "holds_expression": { + "default": false, + "title": "Holds Expression", + "type": "boolean" + }, + "name": { + "title": "Name", + "type": "string" + }, + "unresolved_reason": { + "anyOf": [ + { + "enum": [ + "not_a_string", + "redacted", + "unparsed_json", + "unread_arguments" + ], + "type": "string" + }, + { + "type": "null" + } + ], + "default": null, + "title": "Unresolved Reason" + }, + "value": { + "anyOf": [ + { + "type": "string" + }, + { + "type": "null" + } + ], + "title": "Value" + } + }, + "required": [ + "name", + "value" + ], + "title": "HostWorkflowAgentSettingV7", + "type": "object" + }, + "HostWorkflowCheckoutRefV7": { + "additionalProperties": false, + "description": "One ``actions/checkout`` step and the ``with.ref`` it declares, as text (#823).\n\n``ref`` is ``null`` when the step declares none, or an empty one: the\ncheckout's default for the triggering event. A ref the label redaction\nrewrites is published redacted with ``unresolved_reason: redacted`` and\nmakes the workflow a blocking limit, as a redacted step reference does\n(#767): a ref names the code the job runs, as a step reference does. A\nvalue that is not a string, or ``with:`` that is not a mapping,\nis ``null`` with ``unresolved_reason`` and records a non-blocking coverage\nissue. The ref is never resolved or fetched.", + "properties": { + "job": { + "title": "Job", + "type": "string" + }, + "ref": { + "anyOf": [ + { + "type": "string" + }, + { + "type": "null" + } + ], + "title": "Ref" + }, + "step": { + "title": "Step", + "type": "string" + }, + "unresolved_reason": { + "anyOf": [ + { + "enum": [ + "not_a_string", + "redacted", + "inputs_not_a_mapping" + ], + "type": "string" + }, + { + "type": "null" + } + ], + "default": null, + "title": "Unresolved Reason" + } + }, + "required": [ + "job", + "step", + "ref" + ], + "title": "HostWorkflowCheckoutRefV7", + "type": "object" + }, + "HostWorkflowGrantV7": { + "additionalProperties": false, + "description": "A v0.6 workflow grant plus the agent launches, unread agent steps and checkout refs its steps declare.\n\nEach list is present only when a step declares one. In a v0.7 grant an\nabsent list means the steps were read and declare none; the schema\nversion, not the key, separates that from a legacy grant that never read\nthem. ``unread_agent_runs`` is a named limit and is never compared.\n``access`` and ``risk`` still describe the workflow's token and triggers\nalone.", "properties": { "access": { "enum": [ @@ -1476,6 +1678,20 @@ "title": "Access", "type": "string" }, + "agent_launches": { + "items": { + "$ref": "#/$defs/HostWorkflowAgentLaunchV7" + }, + "title": "Agent Launches", + "type": "array" + }, + "checkout_refs": { + "items": { + "$ref": "#/$defs/HostWorkflowCheckoutRefV7" + }, + "title": "Checkout Refs", + "type": "array" + }, "config_sha256": { "title": "Config Sha256", "type": "string" @@ -1565,6 +1781,13 @@ "title": "Triggers", "type": "array" }, + "unread_agent_runs": { + "items": { + "$ref": "#/$defs/HostWorkflowUnreadAgentRunV7" + }, + "title": "Unread Agent Runs", + "type": "array" + }, "write_all": { "default": false, "title": "Write All", @@ -1590,7 +1813,7 @@ "effective_write_scopes", "reusable_calls" ], - "title": "HostWorkflowGrantV6", + "title": "HostWorkflowGrantV7", "type": "object" }, "HostWorkflowPermissionsV4": { @@ -1691,6 +1914,35 @@ "title": "HostWorkflowStepActionV6", "type": "object" }, + "HostWorkflowUnreadAgentRunV7": { + "additionalProperties": false, + "description": "A ``run:`` step that mentions a known agent CLI and is not read as an agent launch (#823 review cycle 4).\n\nAny ``run:`` holding ``claude`` or ``codex`` as a word of its own that is\nnot an agent launch this reader reads \u2014 more than one line or command, a\nquote, an expansion, a redirection, a comment, a continuation, a\n``${{ }}`` expression, another program such as ``npx`` or ``timeout``, a\nsubcommand that is not a headless launch, or a declared ``shell:`` other\nthan ``bash`` or ``sh`` run on the script alone (so ``bash -c '\u2026' {0}``\ntoo) \u2014 once for each agent CLI it mentions. It is a\nnamed, non-blocking limit and nothing more: none of the step's text is\npublished, it is never compared, so adding, removing or editing it gives\nno row, and it never says that the step starts, or does not start, an\nagent. ``job`` and ``step`` are published labels (#802).", + "properties": { + "agent": { + "enum": [ + "claude", + "codex" + ], + "title": "Agent", + "type": "string" + }, + "job": { + "title": "Job", + "type": "string" + }, + "step": { + "title": "Step", + "type": "string" + } + }, + "required": [ + "job", + "step", + "agent" + ], + "title": "HostWorkflowUnreadAgentRunV7", + "type": "object" + }, "InstructionStructureEvidence": { "additionalProperties": false, "properties": { diff --git a/docs/integrations.md b/docs/integrations.md index edc5f1ab3..e3eebc811 100644 --- a/docs/integrations.md +++ b/docs/integrations.md @@ -224,7 +224,19 @@ Without a configured manifest, when every changed file is host configuration — stays quiet when no row widens what the agent can do. A workflow step moved to a different action reference, such as a pinned SHA to `@main`, is a non-widening row, so the hook stays quiet about it; `diff` and the PR comment -still show it. It names each widening +still show it. The same holds for an agent launch in a workflow whose settings +change without gaining a documented widening rule, for a checkout's ref, and +for an edit to an argument input the audit does not read, which is compared by +a digest, named in the row and named as a limit in `audit --host`. A `run:` +that mentions an agent CLI and that the audit does not read is different: it +gives no row in `diff` or the PR comment, whatever is edited, and is named only +as a limit in `audit --host`. An agent launch that gains a rule, such as a plain `claude_args` +gaining `--dangerously-skip-permissions` on any of its lines, widens, and the +hook announces it (#823), unless the launch may be a step of that job the audit +did not read, rewritten, or the job's launch held before a `${{ }}` expression +or an unread argument input the rule is read from, or the rule moved in from a +job the launch left, as a renamed job's does. +It names each widening row once, and repeats the announcement only when the change or its rows change. A missing base ref, an incomparable inventory or unparsed output is never quiet. When host configuration changes beside other files, the Stop hook diff --git a/llms-full.txt b/llms-full.txt index 2059269df..21c3d9c65 100644 --- a/llms-full.txt +++ b/llms-full.txt @@ -1618,6 +1618,53 @@ directory, still refuses its comparison. A `0.20` verifier claiming a partial comparison or a `scope` is refused. See [the migration note](../STABILITY.md#partial-host-comparison-808). +Runtime contract v41, extended in place, also reads how a coding agent is +launched inside a workflow job (#823). Host-grants `0.6` shipped in 1.1.0, so +host-grants inventory, baseline and drift schemas move to `0.7`, and a workflow +grant adds `agent_launches[]`, `unread_agent_runs[]` and `checkout_refs[]`, +each omitted when empty. An agent launch is a step whose `uses:` is a +documented agent action (`anthropics/claude-code-action`, +`anthropics/claude-code-base-action`, `openai/codex-action`) with the +permission inputs it declares, or a `run:` that is one line of plain words +running `claude -p` / `codex exec` under `bash` or `sh`, with its documented +permission flags; its `job`, `step`, `agent`, `form` (`read`, or `unresolved` +with `inputs_not_a_mapping`), `settings[]` (`name`, `value`, +`unresolved_reason`, `holds_expression`), `widening_rules[]` (`rule`, +`setting`) and `job_secrets[]`. Shell is not parsed: any other `run:` that +mentions `claude` or `codex` is an `unread_agent_runs[]` entry (`job`, `step`, +`agent`), a named non-blocking limit that publishes none of its text, is never +compared and gives no row; and `claude_args` / `codex-args` are read only as a +plain list of words, any other value being `unread_arguments`, compared by a +digest and read for no rule. A checkout ref is each `actions/checkout` step's +`with.ref`, `null` for the default. Values are compared as text and never +executed. A JSON object in a `settings` or `mcp_config` input publishes its +shape and none of its free text — key names, with each string a +`` digest except those a host reader publishes (a permission rule, +a documented setting's value, an MCP server's command name and URL host) — so +an MCP server's arguments and a hook's command are compared but never +published; a URL publishes its scheme and host. Only a documented rule a job's +launches gain — bypassed permission checks (a flag, or JSON settings whose +`defaultMode` is `bypassPermissions`), a bypassed or `danger-full-access` +sandbox (`permission-profile: :danger-full-access` included, and in a +`codex exec` step without `--sandbox` a `--config` override of `sandbox_mode` +or `default_permissions` that selects it), `safety-strategy: unsafe`, or a +user gate opened to `*` — raises `workflow_agent_widened_` and +makes the row `widened`. A rule is read only from text this audit reads +exactly, and one a launch already met in a job it left, one where an unread +step of the job became a read launch, or one where the job's launch before +held an expression or an unread argument input the rule is read from, is named +and not claimed. Every other edit is `changed`, and a workflow row that runs an +agent ends its `why` with the job facts beside each agent step. An action whose +`with:` is not a mapping is `unresolved` and a named non-blocking limit; a +setting holding credential-shaped text, prose included, is published redacted, +compared as published and a named non-blocking limit, and a checkout ref +holding it is a blocking limit, as a redacted step reference is. A `0.4`–`0.6` +baseline holding a workflow grant is incomparable +(`baseline_workflow_agent_launches_unavailable`); one without a workflow stays +comparable. It moves neither #821's verifier `0.21` nor its capability diff +`0.4`, and `minimum_control_contract_version` stays `21`. See +[the migration note](../STABILITY.md#workflow-agent-launches-contract-v41-823). + The same unreleased runtime contract v41 also names what changed in a hook and in an MCP server's launch arguments (#819). Host-grants inventory, baseline and drift schemas move to `0.7`: a hook grant adds `handlers[]` (each handler's @@ -1631,7 +1678,7 @@ row names the changed field, `PostToolUse: matcher Edit → Edit|Write|Bash` or a `package` difference, in the text and in `review.changes[].change`. The members display what `config_sha256` already binds, so grant equality and the inventory digests leave them out: they move no row value, row count, verifier -or capability-diff schema, a `0.6` baseline stays comparable with no new row +or capability-diff schema, a `0.6` baseline without workflow grants stays comparable with no new row or reason, and `minimum_control_contract_version` stays `21`. A saved baseline holds none of the members. See [the migration note](../STABILITY.md#hook-mcp-detail-fields-819). diff --git a/src/agents_shipgate/cli/host_audit.py b/src/agents_shipgate/cli/host_audit.py index 8ed74c647..62c437833 100644 --- a/src/agents_shipgate/cli/host_audit.py +++ b/src/agents_shipgate/cli/host_audit.py @@ -642,10 +642,15 @@ def _refuse_invalid_baseline_overwrite( next_action=INCOMPARABLE_BASELINE_REVIEW, command=None, ) from exc - # A v0.6 baseline compares exactly as its v0.7 reading does, so it may be - # replaced; every older one is still refused (#819). + # Display-only hook/MCP fields preserve v0.6 compatibility (#819), but + # workflow grants did not read agent launches until v0.7 (#823). + unread_workflow = baseline.get("host_grants_schema_version") == "0.6" and any( + grant.get("kind") == "workflow" + for grant in (baseline.get("inventory") or {}).get("grants", []) + ) if ( baseline.get("host_grants_schema_version") not in OVERWRITABLE_BASELINE_SCHEMA_VERSIONS + or unread_workflow or baseline.get("_load_error") ): reason = str(baseline.get("_load_error") or "unsupported_baseline_schema") diff --git a/src/agents_shipgate/core/capability_diff_rows.py b/src/agents_shipgate/core/capability_diff_rows.py index 165817674..bd18e8f3d 100644 --- a/src/agents_shipgate/core/capability_diff_rows.py +++ b/src/agents_shipgate/core/capability_diff_rows.py @@ -30,12 +30,19 @@ from agents_shipgate.core.host_grants import ( _PLAIN_TOKEN_RE, + AGENT_RULE_INPUTS, DETAIL_NOT_SHOWN, + UNTRUSTED_INPUT_TRIGGERS, + agent_launch_key, + agent_rule_gains, + agent_rule_text, + checkout_ref_key, hook_loading_basis, host_grant_expansion_signals, permission_rule_replacements, published_setting_value, published_workflow_label, + pull_request_code_ref, secret_mapping_key, step_action_key, ) @@ -241,6 +248,304 @@ def _moved_between_jobs( moved.append((item, match)) return moved, [*changed, *arriving] + +def _job_entry_changes( + before: dict[str, Any] | None, + after: dict[str, Any] | None, + field: str, + key: Any, +) -> tuple[list[dict[str, Any]], list[dict[str, Any]]]: + """The entries of ``field`` only one side of a changed workflow declares (#823). + + Compared by ``key``, the way the comparator decides whether the grant + changed, so a renamed or reordered step appears on neither side. + """ + + if not _is_workflow_pair(before, after): + return [], [] + + def only_in(side: list[dict[str, Any]], other: list[dict[str, Any]]) -> list[dict[str, Any]]: + surplus = Counter(key(item) for item in side) - Counter(key(item) for item in other) + picked: list[dict[str, Any]] = [] + for item in side: + if surplus[key(item)] > 0: + surplus[key(item)] -= 1 + picked.append(item) + return picked + + old, new = (before or {}).get(field, []), (after or {}).get(field, []) + return only_in(old, new), only_in(new, old) + + +def _agent_label(agent: str) -> str: + """How a reviewer recognises the agent a step launches.""" + + return {"claude": "claude -p", "codex": "codex exec"}.get(agent, agent) + + +def _agent_launch_value(item: dict[str, Any]) -> str: + """One agent launch as a cell shows it: where, which agent, and its declared settings.""" + + where = f"{item['job']}/{item['step']}" + label = _agent_label(str(item["agent"])) + reason = item.get("unresolved_reason") + if reason: + return f"{where}: runs {label} (unresolved: {str(reason).replace('_', ' ')})" + cli = item["agent"] in {"claude", "codex"} + parts = [] + for setting in item.get("settings") or []: + unread = setting.get("unresolved_reason") + name = str(setting["name"]) + if unread == "unread_arguments": + # Compared by its digest alone, so the cell shows the digest (#823 review cycle 4). + parts.append(f"{name} (not read; digest {setting['value']})") + # A redacted value is compared as published, so the cell shows it (#823 review F2). + elif unread and not (unread == "redacted" and setting.get("value") is not None): + parts.append(f"{name} (unresolved: {str(unread).replace('_', ' ')})") + elif setting.get("value") is None: + parts.append(name) + else: + parts.append(f"{name} {setting['value']}" if cli else f"{name}: {setting['value']}") + if not parts: + return f"{where}: runs {label} with no permission {'flags' if cli else 'inputs'}" + return f"{where}: runs {label} with " + "; ".join(parts) + + +def _checkout_ref_value(item: dict[str, Any]) -> str: + where = f"{item['job']}/{item['step']}" + reason = item.get("unresolved_reason") + if reason: + return f"{where}: checkout ref (unresolved: {str(reason).replace('_', ' ')})" + if item.get("ref") is None: + return f"{where}: checkout of the default ref" + return f"{where}: checkout of ref {item['ref']}" + + +def _agent_launch_reasons( + before: dict[str, Any] | None, + after: dict[str, Any] | None, + gone: list[dict[str, Any]], + new: list[dict[str, Any]], + gone_checkouts: list[dict[str, Any]], + new_checkouts: list[dict[str, Any]], +) -> list[str]: + """What changed in how an agent is launched, and which of it widens (#823). + + Only a documented rule the engine claims — ``agent_rule_gains(...).claimed``, + the rule behind ``workflow_agent_widened_*`` — is called a widening, and + the sentence says which rule and where. A rule gained where a step of the + job this audit does not read is gone, or whose launch held, in an + input the rule is read from, a ``${{ }}`` expression or an argument list + this audit does not read, or that moved in from another job, is named and + not called a widening, as the engine claims no expansion for it. Every other + agent-launch or checkout edit is a change: its settings are compared as + published text, and nothing here ranks one value against another. + """ + + def where(item: dict[str, Any]) -> str: + return f"{item['job']}/{item['step']}" + + gains = agent_rule_gains(before, after) + reasons: list[str] = [] + # Launches a sentence about a rule already names. + named: set[str] = set() + for _job, rule, detail, entry in gains.claimed: + named.add(where(entry)) + reasons.append(f"an agent launch now {agent_rule_text(rule, detail)} ({where(entry)})") + for (_job, rule, detail, entry), source in gains.unread_before: + named.add(where(entry)) + reasons.append( + f"an agent launch now {agent_rule_text(rule, detail)} ({where(entry)}), which is not counted as a " + f"widening: a step in this job that may launch the agent in a form this audit does not read " + f"is gone ({where(source)}), and this launch may be that step rewritten in a form this audit " + "reads, which may already have done the same" + ) + for (_job, rule, detail, entry), setting, how in gains.setting_before: + named.add(where(entry)) + held = ( + "held a `${{ }}` expression, whose substituted text this audit does not read" + if how == "expression" + else "was not a plain list of words this audit reads, so no rule was read from it" + ) + reasons.append( + f"an agent launch now {agent_rule_text(rule, detail)} ({where(entry)}), which is not counted as a " + f"widening: before, this job's {setting} {held}, and it may already have done the same" + ) + for (_job, rule, detail, entry), source in gains.moved: + named.update({where(entry), where(source)}) + reasons.append( + f"an agent launch that {agent_rule_text(rule, detail)} moved between jobs ({where(source)} → " + f"{where(entry)}), which is not counted as a widening: the launch already met that " + "rule in the job it left, and it now runs with the receiving job's token permissions" + ) + # The step label only words the sentence; what changed was decided by the + # comparator's key, which never reads it. + old = {where(item) for item in gone} + now = {where(item) for item in new} + groups: dict[str, list[str]] = {} + for item in (*new, *gone): + label = where(item) + if label in named: + continue + verb = "changed" if label in old and label in now else ("added" if label in now else "removed") + groups.setdefault(verb, []) + if label not in groups[verb]: + groups[verb].append(label) + phrases = [ + f"{wording} ({', '.join(groups[verb])})" + for verb, wording in ( + ("changed", "an agent launch's declared settings changed"), + ("added", "a step now launches an agent"), + # What was established is that the step declares no launch this + # audit reads, not that it starts no agent (#823 review). + ("removed", "a step no longer declares an agent launch this audit reads"), + ) + if verb in groups + ] + if phrases: + reasons.append( + f"{_joined_words(phrases)}; agent launch settings are compared as declared text, " + "and a change that gains no documented widening rule is not counted as a widening" + ) + if "removed" in groups and after is not None: + reasons.append( + "a step that no longer declares one may still start an agent in a way this audit " + "does not read, such as an action outside its table, a script, or a `run:` this " + "audit does not read as a launch (more than one command, quoting, an expansion, `npx`, " + "`codex` options before `exec`), so this row does not say that it no longer starts one" + ) + # A rule is read only from literal text a `${{ }}` expression cannot + # reach, so the row says where that leaves text unread (#823 review). + expressions = list(dict.fromkeys( + f"{setting['name']} at {where(item)}" + for item in new if item.get("form") == "read" + for setting in item.get("settings") or [] + if setting.get("holds_expression") and setting["name"] in AGENT_RULE_INPUTS + )) + if expressions: + reasons.append( + "an agent launch setting holds a `${{ }}` expression (" + ", ".join(expressions) + "), " + "which GitHub substitutes before the action reads it; documented widening rules are " + "read only from the literal text the expression cannot reach, so this row does not " + "say whether the text it reaches meets one" + ) + # An argument input that is not a plain list of words is compared by its + # digest and read for no rule (#823 review cycle 4). + unread_arguments = list(dict.fromkeys( + f"{setting['name']} at {where(item)}" + for item in new + for setting in item.get("settings") or [] + if setting.get("unresolved_reason") == "unread_arguments" + )) + if unread_arguments: + reasons.append( + "an agent launch's argument input is not a plain list of words this audit reads (" + + ", ".join(unread_arguments) + "); none of its text is published and it is compared by " + "a digest only, so this row does not say whether it meets a documented widening rule" + ) + unread = list(dict.fromkeys(where(item) for item in new if item.get("form") != "read")) + if unread: + reasons.append( + f"a step launches an agent in a form this audit does not read ({', '.join(unread)}); " + "its settings are not compared, so this row does not say what that agent may do" + ) + # A checkout step on one side only, such as one in an added job, is + # worded as added or removed, not as a changed ref (#823 review cycle 3). + old_checkouts = {where(item) for item in gone_checkouts} + new_checkouts_at = {where(item) for item in new_checkouts} + checkout_groups: dict[str, list[str]] = {} + for item in (*new_checkouts, *gone_checkouts): + label = where(item) + verb = ( + "changed" if label in old_checkouts and label in new_checkouts_at + else ("added" if label in new_checkouts_at else "removed") + ) + checkout_groups.setdefault(verb, []) + if label not in checkout_groups[verb]: + checkout_groups[verb].append(label) + checkout_phrases = [ + f"{wording} ({', '.join(checkout_groups[verb])})" + for verb, wording in ( + ("changed", "a checkout's declared ref changed"), + ("added", "a step now declares a checkout"), + ("removed", "a step no longer declares a checkout"), + ) + if verb in checkout_groups + ] + if checkout_phrases: + reasons.append( + f"{_joined_words(checkout_phrases)}; a ref names which commit's code the job runs " + "and adds no scope" + ) + return reasons + + +#: How many agent steps one note names before counting the rest. +_NOTE_STEP_LIMIT = 5 + + +def _agent_composition_note(grant: dict[str, Any]) -> str | None: + """The job facts beside each agent step, as a note on the workflow's row (#823). + + Named, never scored: an untrusted-input trigger, the job's write scopes, + the secrets the job references and a checkout of pull request code in the + job. It is not a verdict and moves no direction; it says where on the + workflow an agent already runs with those facts. + """ + + launches = grant.get("agent_launches") or [] + if not launches: + return None + triggers = [name for name in grant.get("triggers", []) if name in UNTRUSTED_INPUT_TRIGGERS] + contexts = {context["job"]: context for context in grant.get("permission_contexts", [])} + by_job: dict[str, list[dict[str, Any]]] = {} + for item in launches: + by_job.setdefault(str(item["job"]), []).append(item) + notes: list[str] = [] + named = 0 + for job, items in by_job.items(): + shown = items[: max(0, _NOTE_STEP_LIMIT - named)] + if not shown: + break + named += len(shown) + steps = ", ".join( + f"{item['job']}/{item['step']} ({_agent_label(str(item['agent']))})" for item in shown + ) + facts: list[str] = [] + if triggers: + noun = "trigger" if len(triggers) == 1 else "triggers" + facts.append(f"the untrusted-input {noun} {_joined_words(triggers)}") + context = contexts.get(job) or {} + writes = [scope for scope, level in (context.get("permissions") or {}).items() if level == "write"] + if "*" in writes: + facts.append("write-all token permissions") + elif writes: + noun = "scope" if len(writes) == 1 else "scopes" + facts.append(f"the write {noun} {_joined_words(writes)}") + secrets = sorted({name for item in items for name in item.get("job_secrets", [])}) + if secrets: + noun = "secret" if len(secrets) == 1 else "secrets" + facts.append(f"the {noun} {_joined_words(secrets)}") + pull_request_code = [ + f"{checkout['job']}/{checkout['step']}" + for checkout in grant.get("checkout_refs", []) + if checkout["job"] == job and pull_request_code_ref(checkout.get("ref")) + ] + if pull_request_code: + facts.append(f"a checkout of pull request code ({', '.join(pull_request_code)})") + note = f"an agent runs at {steps}" + if facts: + note += " beside " + _joined_words(facts) + notes.append(note) + rest = len(launches) - named + if rest: + notes.append(f"{rest} more agent step(s) run in this workflow") + return "; ".join(notes) + + +def _joined_words(items: list[str]) -> str: + return items[0] if len(items) == 1 else f"{', '.join(items[:-1])} and {items[-1]}" + #: Direction is deliberately coarse here. Presence is certain: a grant is #: in one side and not the other. *Width* is not — deciding that #: `Bash(npm *)` -> `Bash(npm test:*)` narrows needs the pattern lattice in @@ -258,14 +563,16 @@ def _grant_value( redact_permission_arguments: bool = False, step_actions: list[dict[str, Any]] | None = None, secret_mappings: list[SecretMapping] | None = None, + agent_launches: list[dict[str, Any]] | None = None, + checkout_refs: list[dict[str, Any]] | None = None, ) -> str: """What a reader recognises this grant by. A workflow has no single name — its authority *is* the combination of access and triggers, so both sides render that combination or the row - reads "workflow -> workflow" and says nothing. ``step_actions`` and - ``secret_mappings`` are the step references and named secrets this side - alone declares; unchanged ones are not repeated. + reads "workflow -> workflow" and says nothing. ``step_actions``, + ``secret_mappings``, ``agent_launches`` and ``checkout_refs`` are the + entries this side alone declares; unchanged ones are not repeated. """ if not grant: @@ -302,6 +609,8 @@ def _grant_value( parts.append(f"{call['job']}: {forwarding}{call['uses']}") parts.extend(_secret_mapping_value(item) for item in secret_mappings or []) parts.extend(_step_action_value(item) for item in step_actions or []) + parts.extend(_checkout_ref_value(item) for item in checkout_refs or []) + parts.extend(_agent_launch_value(item) for item in agent_launches or []) return ", ".join(part for part in parts if part) or kind if kind in _SETTING_KINDS and grant.get("setting"): # A setting row read `True` or `dontAsk` alone, which names no setting @@ -334,11 +643,14 @@ def _why( new_steps: list[dict[str, Any]] | None = None, gone_secrets: list[SecretMapping] | None = None, new_secrets: list[SecretMapping] | None = None, + agent_reasons: list[str] | None = None, ) -> str: """Why a reviewer should care, in the reviewer's terms. Stated as what the grant *permits*, never as a prediction about what the agent will do with it — the engine reads configuration, not behaviour. + A workflow that launches an agent ends with the job facts beside each + agent step (#823), read off ``grant``, the side the row describes. """ kind = str(grant.get("kind") or "") @@ -394,7 +706,11 @@ def _why( + "); it names different code to run with that job's existing " "token permissions and adds no scope" ) - return "; ".join(reasons) or "changes the workflow's own authority" + reasons.extend(agent_reasons or []) + why = "; ".join(reasons) or "changes the workflow's own authority" + # A removed workflow runs nothing any more, so it gets no note. + note = None if direction == REMOVED else _agent_composition_note(grant) + return f"{why}; {note}" if note else why if kind == "hook": # The basis, stated in the row, because the row is what a reviewer # reads: a parsed hook file is not proof a host loads it (#714). @@ -1072,10 +1388,24 @@ def capability_diff_rows( direction = WIDENED gone_steps, new_steps = _step_action_changes(before_grant, after_grant) gone_secrets, new_secrets = _secret_mapping_changes(before_grant, after_grant) + gone_agents, new_agents = _job_entry_changes( + before_grant, after_grant, "agent_launches", agent_launch_key + ) + gone_checkouts, new_checkouts = _job_entry_changes( + before_grant, after_grant, "checkout_refs", checkout_ref_key + ) + agent_reasons = ( + _agent_launch_reasons( + before_grant, after_grant, gone_agents, new_agents, gone_checkouts, new_checkouts + ) + if grant.get("kind") == "workflow" + else [] + ) why = _why( grant, direction, gone_steps=gone_steps, new_steps=new_steps, gone_secrets=gone_secrets, new_secrets=new_secrets, + agent_reasons=agent_reasons, ) row = CapabilityDiffRow( subject=_subject(grant), @@ -1084,12 +1414,16 @@ def capability_diff_rows( redact_permission_arguments=redact_permission_arguments, step_actions=gone_steps, secret_mappings=gone_secrets, + agent_launches=gone_agents, + checkout_refs=gone_checkouts, ), after=_grant_value( after_grant, redact_permission_arguments=redact_permission_arguments, step_actions=new_steps, secret_mappings=new_secrets, + agent_launches=new_agents, + checkout_refs=new_checkouts, ), direction=direction, why=why, diff --git a/src/agents_shipgate/core/host_grants.py b/src/agents_shipgate/core/host_grants.py index 2e345d61d..08e16d72a 100644 --- a/src/agents_shipgate/core/host_grants.py +++ b/src/agents_shipgate/core/host_grants.py @@ -54,6 +54,7 @@ ) from agents_shipgate.core.host_settings import ( CLAUDE_LIST_SETTINGS, + CLAUDE_SCALAR_SETTINGS, claude_setting_values, rate_claude_setting, ) @@ -1873,6 +1874,1466 @@ def step_action_key(entry: dict[str, Any]) -> tuple[str, str, str, str]: ) +# --- agent launches (#823) ---------------------------------------------------------- +# +# How a coding agent is launched inside a job: a known agent action's permission +# inputs, a literal agent CLI command in a `run:` step, and the ref each +# `actions/checkout` step declares. Every value is compared as declared text; no +# action is fetched, no command is run and no expression is evaluated. +# +# Shell is not parsed (#823 review cycle 4). A `run:` is read only when it is one +# line of plain words — no quote, expansion, operator, redirection, comment or +# continuation, so every POSIX shell splits it at its blanks and nowhere else — +# whose program is a known agent CLI; `claude_args` and `codex-args` are read +# only when they are such a list of words, which each action's own splitter reads +# the same way. Every other form is never guessed at: a `run:` that mentions an +# agent CLI is named as an unread step, and an argument input as an unread +# setting, each a non-blocking limit that publishes none of its text. + + +@dataclass(frozen=True) +class _AgentAction: + """What one documented agent action's inputs mean to this reader (#823). + + ``inputs`` are compared as text. ``gates`` are comma-separated user lists + where a ``*`` entry opens the gate to every user (to every bot, for + ``allowed_bots``). ``args`` names the input + that carries agent CLI arguments, read, when it is a plain list of words, + for the documented widening rules. ``modes`` are ``(input, value, rule)``: an + input whose value is itself a documented widening. ``settings`` names the + input that holds Claude Code settings, JSON or a path; written as JSON, its + ``defaultMode`` is read as the settings reader reads it (#823 review). + """ + + family: Literal["claude", "codex"] + inputs: tuple[str, ...] + gates: tuple[str, ...] = () + args: str | None = None + modes: tuple[tuple[str, str, str], ...] = () + settings: str | None = None + + +_CLAUDE_BASE_ACTION = _AgentAction( + family="claude", + inputs=( + "allowed_tools", "claude_args", "disallowed_tools", "mcp_config", + "plugin_marketplaces", "plugins", "settings", + ), + args="claude_args", + settings="settings", +) + +#: The documented agent actions, by ``owner/repo`` (matched case-insensitively, +#: at any ref). The input names are those the actions' own `action.yml` +#: declare; `allowed_tools`, `disallowed_tools` and `mcp_config` are the +#: Claude actions' earlier inputs, still read when a workflow sets them. The +#: base action is published both as its own repository and as the +#: `base-action` directory of `anthropics/claude-code-action`. +_AGENT_ACTIONS: dict[str, _AgentAction] = { + "anthropics/claude-code-action": _AgentAction( + family="claude", + inputs=( + "additional_permissions", "allowed_bots", "allowed_non_write_users", + "allowed_tools", "claude_args", "disallowed_tools", "mcp_config", + "plugin_marketplaces", "plugins", "settings", + ), + gates=("allowed_bots", "allowed_non_write_users"), + args="claude_args", + settings="settings", + ), + "anthropics/claude-code-base-action": _CLAUDE_BASE_ACTION, + "anthropics/claude-code-action/base-action": _CLAUDE_BASE_ACTION, + "openai/codex-action": _AgentAction( + family="codex", + inputs=( + "allow-bot-users", "allow-bots", "allow-users", "codex-args", + "permission-profile", "safety-strategy", "sandbox", + ), + gates=("allow-users",), + args="codex-args", + # `:danger-full-access` is Codex's reserved name for its built-in + # full-access permission profile, the profile form of the + # `danger-full-access` sandbox (#823 review). + modes=( + ("sandbox", "danger-full-access", "danger_full_access"), + ("permission-profile", ":danger-full-access", "danger_full_access"), + ("safety-strategy", "unsafe", "unsafe_safety_strategy"), + ), + ), +} + +#: Flag spelling -> (primary spelling, arity): ``0`` takes no value, ``1`` one, +#: ``None`` every following word up to the next one starting with ``-``, as the +#: CLI's variadic options read them. From Claude Code's CLI reference. +#: ``--settings`` and ``--mcp-config`` are not listed: a value passing one is +#: not read at all (:data:`_UNREAD_ARGUMENT_FLAGS`). +_CLAUDE_FLAGS: dict[str, tuple[str, int | None]] = { + "--permission-mode": ("--permission-mode", 1), + "--dangerously-skip-permissions": ("--dangerously-skip-permissions", 0), + "--allow-dangerously-skip-permissions": ("--allow-dangerously-skip-permissions", 0), + "--allowedTools": ("--allowedTools", None), + "--allowed-tools": ("--allowedTools", None), + "--disallowedTools": ("--disallowedTools", None), + "--disallowed-tools": ("--disallowedTools", None), + "--add-dir": ("--add-dir", None), + "--permission-prompt-tool": ("--permission-prompt-tool", 1), +} + +#: The same for ``codex exec``, from the Codex CLI's shared option definitions. +_CODEX_FLAGS: dict[str, tuple[str, int | None]] = { + "--sandbox": ("--sandbox", 1), + "-s": ("--sandbox", 1), + "--dangerously-bypass-approvals-and-sandbox": ("--dangerously-bypass-approvals-and-sandbox", 0), + "--yolo": ("--dangerously-bypass-approvals-and-sandbox", 0), + "--approve-for-me": ("--approve-for-me", 0), + "--not-so-yolo": ("--approve-for-me", 0), + "--dangerously-bypass-hook-trust": ("--dangerously-bypass-hook-trust", 0), + "--add-dir": ("--add-dir", 1), + "--config": ("--config", 1), + "-c": ("--config", 1), + "--profile": ("--profile", 1), + "-p": ("--profile", 1), +} + +_AGENT_FLAG_TABLES = {"claude": _CLAUDE_FLAGS, "codex": _CODEX_FLAGS} + +#: The agent action inputs a documented widening rule is read from. +AGENT_RULE_INPUTS: frozenset[str] = frozenset( + name + for spec in _AGENT_ACTIONS.values() + for name in ( + *spec.gates, + *(mode for mode, _value, _rule in spec.modes), + *(item for item in (spec.args, spec.settings) if item), + ) +) + +#: The widening each documented rule names, as a row's ``why`` says it. +AGENT_WIDENING_RULES: dict[str, str] = { + "bypass_permissions": "skips permission checks (bypassPermissions)", + "bypass_approvals_and_sandbox": "bypasses approvals and the sandbox", + "danger_full_access": "runs without a sandbox (danger-full-access)", + "unsafe_safety_strategy": "runs without privilege restrictions (safety-strategy: unsafe)", + "open_gate": "accepts runs triggered by any user", +} + +#: A user gate whose ``*`` entry admits any bot rather than any user. +_BOT_GATES = frozenset({"allowed_bots"}) + + +def agent_rule_text(rule: str, detail: str) -> str: + """How a row's ``why`` names one documented rule: ``detail`` is the gate an ``open_gate`` opened.""" + + if rule == "open_gate" and detail in _BOT_GATES: + return f"accepts runs triggered by any bot ({detail}: *)" + return AGENT_WIDENING_RULES[rule] + (f" ({detail}: *)" if detail else "") + + +#: A literal checkout ref that names pull request code, not the base branch's: +#: the documented pull request and workflow-run head expressions, and +#: ``refs/pull//head`` or ``/merge``. +_PULL_REQUEST_CODE_REF_RE = re.compile( + r"\$\{\{\s*(?:github\.event\.pull_request\.(?:head\.(?:sha|ref)|merge_commit_sha)" + r"|github\.head_ref|github\.event\.workflow_run\.head_(?:sha|branch))\s*\}\}" + r"|refs/pull/(?:\d+|\$\{\{[^}]*\}\})/(?:head|merge)" +) + +#: Triggers that run with the base repository's context on input from people +#: who need not hold write access, as the support page lists them. +UNTRUSTED_INPUT_TRIGGERS = ("issue_comment", "issues", "pull_request_target", "workflow_run") + +_ASSIGNMENT_RE = re.compile(r"[A-Za-z_][A-Za-z0-9_]*=") +_SECRET_REFERENCE_RE = re.compile(r"\bsecrets\.([A-Za-z_][A-Za-z0-9_]*)") +#: One ``${{ … }}`` expression's body and its closing ``}}``, or an unterminated +#: ``${{`` to the end of the text, which names no secret. A lazy match that +#: must find ``}}`` scanned to the end from every ``${{``, so text holding many +#: unterminated ones took quadratic time (#823 review cycle 7). +_EXPRESSION_RE = re.compile(r"\$\{\{(.*?)(\}\}|\Z)", re.S) + + +def pull_request_code_ref(ref: str | None) -> bool: + """Whether a published checkout ref names pull request code (#823).""" + + return ref is not None and _PULL_REQUEST_CODE_REF_RE.fullmatch(ref.strip()) is not None + + +# --- the only forms read: plain lists of words (#823 review cycle 4) -------------------- +# +# Four review cycles each found a shell form a parser mis-read, so none is +# parsed. A word made only of these characters means the same to a POSIX shell, +# to shell-quote (which the Claude actions split `claude_args` with, after making +# `()|&;<>` literal) and to string-argv (which `openai/codex-action` splits +# `codex-args` with): no quote, `$`, backtick, backslash, `#`, `~`, glob, +# brace, bracket, `!`, `@`, `;`, `&`, `|`, `<`, `>` or non-ASCII character is +# among them, so the words are exactly the text split at its blanks. + +#: A word of a plain `run:` command. +_PLAIN_RUN_WORD_RE = re.compile(r"[A-Za-z0-9_./:=,%+-]+") +#: A word of a plain argument input: parentheses too, which neither action's +#: splitter reads as anything but a character (`Bash(git:status)`), while a +#: shell would. +_PLAIN_ARGUMENT_WORD_RE = re.compile(r"[A-Za-z0-9_./:=,%+()-]+") +#: Flags whose value the host readers withhold or read from a file: a value +#: passing one, in any spelling, is not read (#823 review cycle 4). +_UNREAD_ARGUMENT_FLAGS = frozenset({"--settings", "--mcp-config"}) +#: A known agent CLI's name as a word of its own: a `run:` that holds one +#: mentions that agent. +_AGENT_CLI_MENTION_RE = re.compile(r"(? list[str] | None: + """``words``, when each is made of plain characters and none is an unread flag, else ``None``.""" + + words = [word for word in words if word] + if not words or not all(pattern.fullmatch(word) for word in words): + return None + if any(word.split("=", 1)[0] in _UNREAD_ARGUMENT_FLAGS for word in words): + return None + return words + + +def _argument_words(text: str) -> list[str] | None: + """An agent action's argument input as its words, when it is a plain list of them. + + Blanks and newlines separate words, as both actions' splitters read + them. ``None`` — the input is not read — for anything else: a quote, a + ``${{ }}`` expression, ``$``, a backtick, a comment, a shell separator, + JSON, a ``--settings`` or ``--mcp-config`` flag, or any other character + outside the plain set. + """ + + return _plain_words(re.split(r"[ \t\r\n]+", text), _PLAIN_ARGUMENT_WORD_RE) + + +def _run_words(run: str) -> list[str] | None: + """A ``run:`` as the words of its one command, when it is one line of plain words. + + Only spaces and tabs separate words, as the shell reads them; a second + line, and every character outside the plain set, is ``None``. + """ + + text = run.strip(" \t\n") + if "\n" in text: + return None + return _plain_words(re.split(r"[ \t]+", text), _PLAIN_RUN_WORD_RE) + + +def _run_command(run: str) -> tuple[str, list[str]] | None: + """The agent CLI a plain one-command ``run:`` launches headless, and its arguments. + + After any ``NAME=value`` assignments, the program's file name must be + ``claude`` with ``-p``/``--print`` among its arguments, or ``codex`` with + ``exec`` (or its alias ``e``) as its first argument. Any other command — + ``claude mcp add``, ``codex login``, ``npx …``, ``timeout 600 claude`` — + launches none this reader reads. + """ + + words = _run_words(run) + if words is None: + return None + index = 0 + while index < len(words) and _ASSIGNMENT_RE.match(words[index]): + index += 1 + if index >= len(words): + return None + program, arguments = words[index].rsplit("/", 1)[-1], words[index + 1:] + options = arguments[:arguments.index("--")] if "--" in arguments else arguments + if program == "claude" and any(word in {"-p", "--print"} for word in options): + return "claude", arguments + if program == "codex" and arguments[:1] in (["exec"], ["e"]): + return "codex", arguments[1:] + return None + + +def _declared_shell(step: dict[Any, Any], job: dict[Any, Any], workflow: dict[Any, Any]) -> Any: + """The ``shell:`` a step runs its ``run:`` with: its own, else its job's, else its workflow's default. + + ``None`` when none is declared, which is the runner's default shell. + """ + + if step.get("shell") is not None: + return step["shell"] + for container in (job, workflow): + defaults = container.get("defaults") + run = defaults.get("run") if isinstance(defaults, dict) else None + if isinstance(run, dict) and run.get("shell") is not None: + return run["shell"] + return None + + +#: An option word of a declared ``bash``/``sh`` template after which the shell +#: still runs the script: a cluster of ``set`` flags and ``-l``, ``-i`` or +#: ``-r``, or one of these long options. Never ``-c``, which runs a command +#: string of its own, ``-s``, which reads commands from standard input, ``-n``, +#: which runs none, ``--rcfile`` or anything else. A cluster ending in ``o`` +#: or ``O`` takes an option name next. +_SHELL_OPTION_RE = re.compile( + r"[-+](?:[abefhiklmprtuvxBCEHPT]+[oO]?|[oO])" + r"|--(?:noprofile|norc|posix|login|restricted|noediting|verbose)" +) +_SHELL_OPTION_NAME_RE = re.compile(r"[a-z_]+") +#: The option name that runs no command. +_SHELL_NO_EXEC = "noexec" + + +def _read_shell(shell: Any) -> bool: + """Whether a ``run:`` under this declared shell is read. + + None declared, ``bash`` or ``sh`` (by any path), or a template that runs + one of them on the script alone: option words it still runs the script + under, then ``{0}`` last, as in ``bash --noprofile --norc -eo pipefail + {0}``. Any other template, such as ``bash -c '…' {0}``, may run a command + of its own or none, so the step is not read (#823 review cycle 7). + """ + + if shell is None: + return True + if not isinstance(shell, str): + return False + words = shell.split() + if not words or words[0].rsplit("/", 1)[-1] not in _READ_SHELLS: + return False + if len(words) == 1: + return True + if words[-1] != "{0}": + return False + index = 1 + while index < len(words) - 1: + word = words[index] + if not _SHELL_OPTION_RE.fullmatch(word): + return False + index += 1 + if not word.startswith("--") and word[-1] in "oO": + name = words[index] if index < len(words) - 1 else "" + if not _SHELL_OPTION_NAME_RE.fullmatch(name) or name == _SHELL_NO_EXEC: + return False + index += 1 + return True + + +def _read_flags( + words: list[str] | tuple[str, ...], table: dict[str, tuple[str, int | None]] +) -> list[tuple[str, int | None, list[str] | None]]: + """The documented permission flags among ``words``: primary spelling, arity and value words. + + A flag is read by name wherever it is a word of its own; every other word + — the prompt, ``--model`` and any undocumented flag — is not compared. A + flag that takes no value, or is given none, has ``None``. A long flag's + value may be attached as ``--name=value``, and a short one's as + ``-s`` or ``-s=``, which clap reads as ``-s `` + (#823 review cycle 3). + """ + + # A standalone terminator makes every following word positional prompt + # text, even when it spells a permission flag. + flags: list[tuple[str, int | None, list[str] | None]] = [] + index = 0 + while index < len(words): + word = words[index] + if word == "--": + break + if word.startswith("--"): + name, equals, attached = word.partition("=") + elif len(word) > 2 and table.get(word[:2], ("", 0))[1] == 1: + name, equals, attached = word[:2], "=", word[2:].removeprefix("=") + else: + name, equals, attached = word, "", "" + spec = table.get(name) + index += 1 + if spec is None: + continue + primary, arity = spec + if arity == 0: + flags.append((primary, arity, None)) + continue + values = [attached] if equals else [] + if arity == 1 and not equals and index < len(words): + values.append(words[index]) + index += 1 + elif arity is None: + while index < len(words) and not words[index].startswith("-"): + values.append(words[index]) + index += 1 + flags.append((primary, arity, values or None)) + return flags + + +#: One ``${{ … }}`` expression, or an unterminated ``${{`` to the end of the text. +_EXPRESSION_SPAN_RE = re.compile(r"\$\{\{.*?(?:\}\}|\Z)", re.S) +#: What an expression reads as while a user gate's entries are split: a +#: private-use character, which is neither a comma nor a blank. +_EXPRESSION_MARK = "" + + +def holds_expression(text: str) -> bool: + """Whether declared text holds a ``${{ }}`` expression GitHub substitutes before the action reads it.""" + + return "${{" in text + + +# --- what an agent setting publishes (#823, #802) -------------------------------------- + +#: A codex ``--config`` override under one of these keys carries values the +#: codex host reader never publishes, as ``.mcp.json`` ``env``/``headers`` do not. +_CONFIG_WITHHELD_KEYS = frozenset({"env", "headers", "http_headers", "env_http_headers"}) + +#: Keys whose value maps MCP server names to servers, anywhere in a structured +#: value: Claude's ``mcpServers`` and codex's ``mcp_servers``. VS Code's +#: ``servers`` counts only at the top, where `.vscode/mcp.json` puts it. +_SERVER_MAP_KEYS = frozenset({"mcpServers", "mcp_servers"}) +_TOP_LEVEL_SERVER_MAP_KEYS = _SERVER_MAP_KEYS | {"servers"} +#: Where, from the top of a Claude Code settings value, the settings reader +#: publishes a string as it is: a ``permissions.allow``/``ask``/``deny`` rule, +#: and the value of a documented setting, under ``permissions`` or at the top. +_PUBLISHED_STRING_PATHS: frozenset[tuple[str, ...]] = frozenset({ + *(("permissions", name) for name in ("allow", "ask", "deny", *CLAUDE_SCALAR_SETTINGS)), + *((name,) for name in (*CLAUDE_SCALAR_SETTINGS, *CLAUDE_LIST_SETTINGS)), +}) + + +def _withheld_string(value: str, label: str | None = None) -> str: + """One string of a structured value as it is published (#823 review C2-F1). + + ````: a short digest of what the host readers digest for it — + the sanitized text, and a URL's query as an MCP server's is (#723) — so + editing it is still a change while none of its text is published. Where a + host reader publishes a label for the string (``label``: an MCP server's + command name or URL, a permission rule, a documented setting's value), the + label is published, followed by the digest only when the label drops + something the digest reads, such as a command's arguments or a URL's query. + """ + + withheld = f"" + if label is None: + return withheld + if label == _sanitize_sensitive_string(value) and not _url_capability_parts(value): + return label + return f"{label} {withheld}" + + +def _server_command(command: str) -> str: + """An MCP server's ``command`` as the MCP reader names it: its first word's file name.""" + + first = command.strip().split(maxsplit=1)[0] if command.strip() else "" + return _withheld_string(command, _sanitize_sensitive_string(Path(first).name or first) or None) + + +def _json_shape(value: Any, path: tuple[str, ...] = ()) -> Any: + """A structured value's shape: what a host reader would publish of it, and no free text. + + Key names, numbers, booleans and ``null`` are kept. Secret-bearing values + are ```` exactly as the host readers redact them before any + digest: ``env`` and ``headers`` (and codex's ``http_headers`` and + ``env_http_headers``) keep their keys, a secret-named key's value + (``apiKeyHelper``) and the word after a secret-named argument are + replaced, and ``policyHelper`` is excluded. Every other string is + :func:`_withheld_string`, keeping only what a host reader publishes: a + ``permissions.allow``/``ask``/``deny`` rule and a documented Claude Code + setting's value (``defaultMode``, ``enabledMcpjsonServers`` entries) at the + top of the value, as the settings reader publishes them; and an MCP + server's command name and URL scheme and host, as the MCP reader publishes + them. So an MCP server's arguments and a hook's command publish nothing + of their text (#823 review C2-F1). ``path`` is the keys above ``value``. + """ + + parent = path[-1] if path else None + if isinstance(value, dict): + in_server = len(path) >= 2 and ( + path[-2] in _SERVER_MAP_KEYS or (len(path) == 2 and path[0] in _TOP_LEVEL_SERVER_MAP_KEYS) + ) + shaped: dict[str, Any] = {} + for key, inner in value.items(): + key_text = str(key) + if key_text == "policyHelper": + shaped[key_text] = "" + elif _is_secret_key(key_text) or parent in _CREDENTIAL_CONTAINER_KEYS: + shaped[key_text] = "" + elif key_text in _CONFIG_WITHHELD_KEYS and isinstance(inner, dict): + shaped[key_text] = {str(name): "" for name in inner} + elif in_server and key_text == "command" and isinstance(inner, str): + shaped[key_text] = _server_command(inner) + elif in_server and key_text in {"url", "serverUrl"} and isinstance(inner, str): + shaped[key_text] = _withheld_string(inner, _sanitize_url(inner)) + else: + shaped[key_text] = _json_shape(inner, (*path, key_text)) + return shaped + if isinstance(value, list): + items: list[Any] = [] + redact_next = False + for item in value: + if redact_next: + items.append("") + redact_next = False + continue + if isinstance(item, str) and not _SECRET_ARG_RE.fullmatch(item): + redact_next = item.lower().lstrip("-").replace("-", "_") in _SECRET_KEY_MARKERS + items.append(_json_shape(item, path)) + return items + if isinstance(value, str): + published = path in _PUBLISHED_STRING_PATHS + return _withheld_string(value, _sanitize_sensitive_string(value) if published else None) + return value + + +def _withheld_json(value: Any) -> str | None: + """A structured value as it may be published: its :func:`_json_shape`, as canonical JSON. + + Canonical, so reformatting or reordering keys changes nothing. + """ + + try: + return json.dumps( + _json_shape(value), sort_keys=True, separators=(",", ":"), + ensure_ascii=False, default=str, + ) + except (RecursionError, TypeError, ValueError): + return None + + +def _withheld_word(word: str) -> str | None: + """One word as it may be published: a JSON object by :func:`_withheld_json`, else as written. + + ``None`` for a word that starts like a JSON object and does not parse: what + it holds cannot be told apart, so none of it may be published. + """ + + if not word.lstrip().startswith("{"): + return word + try: + loaded = json.loads(word) + except (ValueError, RecursionError): + return None + return _withheld_json(loaded) + + +#: The action inputs that hold JSON or a file path (#823 review). +_JSON_OR_PATH_INPUTS = frozenset({"mcp_config", "settings"}) +#: A file path as it may be published: plain path characters, and any +#: ``${{ }}`` expression in it a plain context reference. +_PLAIN_PATH_RE = re.compile(r"(?:[A-Za-z0-9_./~@+,:=%-]|\$\{\{[ A-Za-z0-9_.\-\[\]*]*\}\})*") + + +def _withheld_json_or_path(text: str) -> str | None: + """A ``settings`` or ``mcp_config`` value as it may be published. + + A JSON object by :func:`_withheld_word`, and a plain file path as + written. Any other text is neither, so the action rejects it, but it may + still hold what a host reader withholds, such as an ``env`` value after a + comment line: it publishes only ````, a digest, so an edit to + it is still a change (#823 review cycle 5). + """ + + if text.lstrip().startswith("{") or _PLAIN_PATH_RE.fullmatch(text): + return _withheld_word(text) + return _withheld_string(text) + + +#: The codex ``--config`` keys whose value is published as written: the +#: settings a documented widening rule reads (``sandbox_mode``, +#: ``default_permissions``), the approval policy and the model. Any other +#: key's value — an MCP server's command or URL, a shell environment value, a +#: provider — publishes only its digest (#823 review cycle 4). +_PUBLISHED_CONFIG_KEYS = frozenset({"sandbox_mode", "default_permissions", "approval_policy", "model"}) + + +def _withheld_config(text: str) -> str: + """A codex ``--config key=value`` override as it may be published. + + The key is published. Under ``env``, ``headers`` or a secret-named key the + value is ````, as the host readers redact such values, so + rotating it is quiet; under one of :data:`_PUBLISHED_CONFIG_KEYS` it is + published as written; under any other key it is ````, a + digest, so editing it is still a change while none of its text is + published. Only a plain word reaches here, so the value is never a TOML + table or array. + """ + + key, equals, value = text.partition("=") + if not equals: + return text + segments = [segment.strip() for segment in key.split(".")] + if any(segment in _CONFIG_WITHHELD_KEYS or _is_secret_key(segment) for segment in segments): + return f"{key}=" + if key.strip() in _PUBLISHED_CONFIG_KEYS: + return text + return f"{key}={_withheld_string(value)}" + + +def _withheld_words(words: list[str] | tuple[str, ...], *, family: str) -> tuple[list[str], bool]: + """Each plain argument word as it may be published, and whether one was redacted as credential-shaped. + + A codex ``--config`` override is withheld however it is attached to its + flag: a separate word after ``-c`` or ``--config``, ``--config=…``, and + ``-c`` or ``-c=``, which clap reads as ``-c ``. The + word after a secret-named word such as ``--token`` or ``password`` is + ````, as the host readers redact it in an MCP server's + arguments, and the value is then credential-shaped: an edit to that word + is not reported, so the caller names it as a limit. Every other word is + published as written, through the label redaction. + """ + + shown: list[str] = [] + config = False + secret = False + redacted = False + for word in words: + if secret: + item = "" + redacted = True + elif config: + item = _withheld_config(word) + elif family == "codex" and word.startswith("--config="): + item = "--config=" + _withheld_config(word.removeprefix("--config=")) + elif family == "codex" and word.startswith("-c") and len(word) > 2: + prefix = "-c=" if word.startswith("-c=") else "-c" + item = prefix + _withheld_config(word.removeprefix(prefix)) + else: + item = word + shown.append(item) + config = family == "codex" and word in {"-c", "--config"} + secret = not secret and word.lower().lstrip("-").replace("-", "_") in _SECRET_KEY_MARKERS + return shown, redacted + + +def _url_withheld(url: str) -> str: + """A URL with its path and query withheld as the host sanitizer withholds them. + + One holding userinfo is kept as written, so the label redaction still + finds the credential in it. + """ + + try: + netloc = urlsplit(url).netloc + except ValueError: + return url + return url if "@" in netloc else _sanitize_url(url) + + +#: Private-use characters that stand for one expression each while a value is redacted. +_EXPRESSION_SLOTS = range(0xE000, 0xF900) +_EXPRESSION_SLOT_RE = re.compile("[\ue000-\uf8ff]") + + +def _published_value(text: str) -> tuple[str, bool]: + """``text`` as it may be published, and whether credential-shaped text had to be redacted. + + A URL's path and query are withheld the way the host sanitizer withholds + them from an MCP server URL (#723), and the rest of the text is published + and compared: a URL path is not a credential. Anything else the #802 label + redaction rewrites — a token shape, a credential assignment, a bearer or + header value, a URL's userinfo — is credential-shaped text, and the value + is published redacted. The caller decides what that refuses. + + Each ``${{ }}`` expression is read as one word while this is decided, so an + expression inside a URL's userinfo is withheld with the userinfo rather + than splitting the URL (``https://x:${{ secrets.T }}@host/…`` publishes + ``https://host/``), and every expression left in the value is + published through the same label redaction. + """ + + expressions = _EXPRESSION_SPAN_RE.findall(text) + if len(expressions) > len(_EXPRESSION_SLOTS) or _EXPRESSION_SLOT_RE.search(text): + # No free character to stand for each expression: read the text as written. + expressions = [] + def url_withheld(match: re.Match[str]) -> str: + # A backslash the URL ends at is kept, so a JSON escape after it, as + # in a published permission rule, stays an escape (#823 review cycle 3). + url = match.group(0) + bare = url.rstrip("\\") + return _url_withheld(bare) + url[len(bare):] + + def urls_withheld(value: str) -> str: + return _URL_RE.sub(url_withheld, value) + + slots = iter(_EXPRESSION_SLOTS) + marked = _EXPRESSION_SPAN_RE.sub(lambda _match: chr(next(slots)), text) if expressions else text + withheld = urls_withheld(marked) + shown = published_workflow_label(withheld) + kept = [urls_withheld(expression) for expression in expressions] + labels = [published_workflow_label(expression) for expression in kept] + redacted = shown != withheld or labels != kept + + if expressions: + shown = _EXPRESSION_SLOT_RE.sub( + lambda match: labels[ord(match.group()) - _EXPRESSION_SLOTS.start], shown + ) + return shown, redacted + + +def _setting_text(value: Any) -> str | None: + """A ``with:`` value as the text GitHub passes, or ``None`` when it is not a scalar.""" + + if isinstance(value, bool): + return "true" if value else "false" + if value is None: + return "" + if isinstance(value, (int, float)): + return str(value) + if isinstance(value, str): + return value.strip() + return None + + +def _published_text(name: str, text: str | None, *, redacted: bool = False) -> dict[str, Any]: + """A setting's withheld text as it may be published; ``None`` text is ``unparsed_json``. + + ``redacted`` says credential-shaped text was already withheld from it. + """ + + if text is None: + return {"name": name, "value": None, "unresolved_reason": "unparsed_json"} + shown, rewritten = _published_value(text) + return {"name": name, "value": shown, "unresolved_reason": "redacted" if redacted or rewritten else None} + + +def _published_setting(name: str, value: Any, *, arguments: str | None = None) -> dict[str, Any]: + """One action input as it may be published: its text, or why it is not. + + ``arguments`` names the family whose action splits this input + (``claude_args``, ``codex-args``). Such an input is read only when it is a + plain list of words (:func:`_argument_words`), published as those words; + otherwise it is ``unread_arguments`` and publishes only ````, + a digest, so an edit to it is still a change while none of its text is + published (#823 review cycle 4). Every other input is one value, a JSON + object published by :func:`_withheld_json`; a ``settings`` or + ``mcp_config`` value that is neither JSON nor a plain path publishes only + a digest (:func:`_withheld_json_or_path`). + """ + + text = _setting_text(value) + if text is None: + return {"name": name, "value": None, "unresolved_reason": "not_a_string"} + if arguments is not None: + words = _argument_words(text) + if words is None: + return {"name": name, "value": _withheld_string(text), "unresolved_reason": "unread_arguments"} + shown, redacted = _withheld_words(words, family=arguments) + return _published_text(name, " ".join(shown), redacted=redacted) + setting = _published_text( + name, _withheld_json_or_path(text) if name in _JSON_OR_PATH_INPUTS else _withheld_word(text) + ) + # Read off the declared text: redaction may rewrite the expression away. + return {**setting, "holds_expression": True} if holds_expression(text) else setting + + +def _published_flag(family: str, name: str, arity: int | None, values: list[str] | None) -> dict[str, Any]: + """One CLI flag of a plain ``run:`` as it may be published: its value words, a ``--config`` override withheld.""" + + if values is None: + return {"name": name, "value": None, "unresolved_reason": None} + if family == "codex" and name == "--config": + shown, redacted = [_withheld_config(value) for value in values], False + else: + shown, redacted = _withheld_words(values, family=family) + return _published_text(name, shown[0] if arity == 1 else " ".join(shown), redacted=redacted) + + +def _setting_key(setting: dict[str, Any]) -> tuple[str, str, str]: + return ( + str(setting["name"]), + str(setting.get("unresolved_reason") or ""), + "" if setting.get("value") is None else str(setting["value"]), + ) + + +def _job_secrets(job: dict[Any, Any], workflow_env: Any) -> list[str]: + """The secret names a job references in ``${{ }}``, and those the workflow ``env`` passes. + + Context for the row that names an agent step in the job, never compared. + Each name is a published label (#802). + """ + + names: set[str] = set() + # Walked once per container, without recursion: a YAML alias can make a + # mapping hold itself, or fan one out many times over (#823 review cycle 5). + seen: set[int] = set() + pending: list[Any] = [job, workflow_env] + while pending: + value = pending.pop() + if isinstance(value, str): + for expression, closed in _EXPRESSION_RE.findall(value): + if closed: + names.update(_SECRET_REFERENCE_RE.findall(expression)) + elif isinstance(value, (dict, list)) and id(value) not in seen: + seen.add(id(value)) + pending.extend(value.values() if isinstance(value, dict) else value) + return sorted({published_workflow_label(name) for name in names}) + + +def _action_identity(uses: Any) -> str | None: + """A step's ``owner/repo`` in lower case when ``uses:`` is ``owner/repo@ref``.""" + + if not isinstance(uses, str): + return None + text = uses.strip() + if not _REMOTE_STEP_ACTION_RE.fullmatch(text): + return None + return text.rpartition("@")[0].casefold() + + +def _action_launch(job: str, step_label: str, agent: str, step: dict[Any, Any]) -> dict[str, Any]: + """A documented agent action's launch: its documented inputs as published, and the rules they meet. + + The rules are read from each input's declared text before any of it is + withheld for publication, so what redaction hides never hides a rule. + """ + + spec = _AGENT_ACTIONS[agent] + entry: dict[str, Any] = { + "job": job, "step": step_label, "agent": agent, "form": "read", + "unresolved_reason": None, "settings": [], + } + inputs = step.get("with") + if inputs is None: + return entry + if not isinstance(inputs, dict): + return {**entry, "form": "unresolved", "unresolved_reason": "inputs_not_a_mapping"} + wanted = {name.casefold(): name for name in spec.inputs} + declared = [ + (wanted[str(key).casefold()], value) + for key, value in inputs.items() + if str(key).casefold() in wanted + ] + settings = [ + _published_setting(name, value, arguments=spec.family if name == spec.args else None) + for name, value in declared + ] + return _with_rules( + {**entry, "settings": sorted(settings, key=_setting_key)}, _action_rules(spec, declared) + ) + + +def _run_launch(job: str, step_label: str, agent: str, arguments: list[str]) -> dict[str, Any]: + """A plain ``run:`` agent CLI command's launch: its documented permission flags, and the rules they meet.""" + + flags = _read_flags(arguments, _AGENT_FLAG_TABLES[agent]) + settings = [_published_flag(agent, name, arity, values) for name, arity, values in flags] + return _with_rules( + { + "job": job, "step": step_label, "agent": agent, "form": "read", + "unresolved_reason": None, "settings": sorted(settings, key=_setting_key), + }, + _flag_rules(agent, flags), + ) + + +def _step_agent_launches( + job: str, step: dict[Any, Any], index: int, shell: Any +) -> tuple[list[dict[str, Any]], list[dict[str, Any]]]: + """The agent launches one step declares, and the unread ``run:`` agent steps it is. + + A known agent action is a launch. A ``run:`` is a launch only when it is + one line of plain words (:func:`_run_command`) whose program is a known + agent CLI, run by ``bash`` or ``sh`` or no declared ``shell:`` (``shell``). + Any other ``run:`` that mentions ``claude`` or ``codex`` as a word of its + own is one unread step per agent CLI it mentions: a non-blocking limit + that publishes none of its text, is never compared and so gives no row, + and never says that the step starts or does not start an agent + (#823 review cycle 4). + """ + + if "uses" in step: + agent = _action_identity(step["uses"]) + if agent is not None and agent in _AGENT_ACTIONS: + return [_action_launch(job, _step_label(step, index), agent, step)], [] + return [], [] + run = step.get("run") + if not isinstance(run, str): + return [], [] + mentioned = sorted(set(_AGENT_CLI_MENTION_RE.findall(run))) + if not mentioned: + return [], [] + label = _step_label(step, index) + command = _run_command(run) if _read_shell(shell) else None + if command is None: + return [], [{"job": job, "step": label, "agent": agent} for agent in mentioned] + agent, arguments = command + return [_run_launch(job, label, agent, arguments)], [] + + +def _checkout_ref(job: str, step: dict[Any, Any], index: int) -> dict[str, Any] | None: + """An ``actions/checkout`` step and the ``with.ref`` it declares, else ``None``.""" + + if _action_identity(step.get("uses")) != "actions/checkout": + return None + entry: dict[str, Any] = { + "job": job, "step": _step_label(step, index), "ref": None, "unresolved_reason": None, + } + inputs = step.get("with") + if inputs is None: + return entry + if not isinstance(inputs, dict): + return {**entry, "unresolved_reason": "inputs_not_a_mapping"} + if "ref" not in inputs: + return entry + text = _setting_text(inputs["ref"]) + if text is None: + return {**entry, "unresolved_reason": "not_a_string"} + shown, redacted = _published_value(text) + # A redacted ref is published redacted, as a step reference is, and makes + # the workflow a blocking limit (#767): two refs may redact alike. + return {**entry, "ref": shown or None, "unresolved_reason": "redacted" if redacted else None} + + +def agent_launch_key(entry: dict[str, Any]) -> tuple[Any, ...]: + """What an agent launch is compared by: its job, agent, form, settings and rules, not its step. + + The widening rules are part of it: they are read from the declared text, + so a rule gained where redaction withholds the text is still a change. + """ + + return ( + str(entry["job"]), + str(entry["agent"]), + str(entry["form"]), + str(entry.get("unresolved_reason") or ""), + tuple(sorted(_setting_key(setting) for setting in entry.get("settings", []))), + tuple(sorted( + (str(item["rule"]), str(item["setting"])) for item in entry.get("widening_rules", []) + )), + ) + + +def checkout_ref_key(entry: dict[str, Any]) -> tuple[str, str, str]: + """What a checkout ref is compared by: its job and the declared ref, not its step.""" + + return ( + str(entry["job"]), + str(entry.get("unresolved_reason") or ""), + "" if entry.get("ref") is None else str(entry["ref"]), + ) + + +def _settings_bypass_permissions(text: str) -> bool: + """Whether Claude Code settings written as JSON set ``defaultMode: bypassPermissions`` (#823 review). + + Read as the settings reader reads `.claude/settings.json` + (:func:`claude_setting_values`): under ``permissions``, else at the top. + A path, or text that does not parse, meets nothing: no file is read. + """ + + if not text.lstrip().startswith("{"): + return False + try: + loaded = json.loads(text) + except (ValueError, RecursionError): + return False + return any( + item.setting == "defaultMode" and item.value == "bypassPermissions" + for item in claude_setting_values(loaded) + ) + + +def _claude_action_rules(words: list[str]) -> set[str]: + """The widening rules the plain words of a Claude action's ``claude_args`` meet. + + Read as ``parse-sdk-options.ts`` reads them: a word starting with ``--`` + is always a flag and never another flag's value, so + ``--dangerously-skip-permissions`` counts wherever it stands, and + ``--permission-mode`` takes the next word unless that starts with ``--``, + the last one counting, as it and the CLI keep the last value of a + repeated option (#823 review cycle 5). + """ + + rules: set[str] = set() + mode: str | None = None + for index, word in enumerate(words): + name, equals, attached = word.partition("=") + following = words[index + 1] if index + 1 < len(words) else "" + value = attached if equals else ("" if following.startswith("--") else following) + if name == "--dangerously-skip-permissions": + rules.add("bypass_permissions") + elif name == "--permission-mode": + mode = value + if mode == "bypassPermissions": + rules.add("bypass_permissions") + return rules + + +def _codex_override(text: str) -> tuple[str, Any] | None: + """A codex ``--config`` override's key and value, as the CLI's ``parse_overrides`` reads them. + + The key is the text before the first ``=``, trimmed; the value is parsed + as TOML, and text that does not parse is a string with its surrounding + quotes trimmed, so ``sandbox_mode=danger-full-access`` and + ``sandbox_mode="danger-full-access"`` set one value. + """ + + key, equals, value = text.partition("=") + if not equals or not key.strip(): + return None + raw = value.strip() + try: + loaded = tomllib.loads(f"value = {raw}").get("value") + except (ValueError, RecursionError): + # ``TOMLDecodeError`` is a ``ValueError``, and so is the error + # ``int()`` raises for an integer longer than Python's digit limit + # (#823 review cycle 6). + loaded = raw.strip("\"'") + return key.strip(), loaded + + +def _codex_config_full_access(flags: list[tuple[str, int | None, list[str] | None]]) -> bool: + """Whether ``codex exec``'s ``--config`` overrides select its full-access sandbox (#823 review cycle 3). + + ``sandbox_mode = "danger-full-access"`` is the setting ``--sandbox`` sets, + and ``default_permissions = ":danger-full-access"`` selects the built-in + full-access profile. As the CLI resolves them, the last override of a key + counts, a ``default_permissions`` override selects the permission + profile over a ``sandbox_mode`` one, and a ``--sandbox`` flag takes + precedence over both, so this reads none while one is passed. + """ + + if any(name == "--sandbox" for name, _arity, _values in flags): + return False + profile: Any = None + mode: Any = None + for name, _arity, values in flags: + if name != "--config": + continue + for value in values or []: + override = _codex_override(value) + if override is None: + continue + key, loaded = override + if key == "default_permissions": + profile = loaded + elif key == "sandbox_mode": + mode = loaded + if profile is not None: + return profile == ":danger-full-access" + return mode == "danger-full-access" + + +def _flag_rules( + family: str, + flags: list[tuple[str, int | None, list[str] | None]], + *, + config_selects_sandbox: bool = True, +) -> set[tuple[str, str]]: + """The widening rules a CLI's documented flags meet, each with the flag that met it. + + ``config_selects_sandbox`` is false for the words `openai/codex-action` + passes on from ``codex-args``: after them the action appends its own + ``--sandbox ``, or for a ``permission-profile`` its own + ``default_permissions`` override (``runCodexExec.ts``), and either takes + precedence over a sandbox ``--config`` override written before it, so + one there selects nothing. + """ + + rules: set[tuple[str, str]] = set() + if family == "codex" and config_selects_sandbox and _codex_config_full_access(flags): + rules.add(("danger_full_access", "--config")) + # The CLI keeps the last value of a repeated `--permission-mode` (#823 review cycle 5). + modes = [values[0] if values else None for name, _arity, values in flags if name == "--permission-mode"] + if family == "claude" and modes and modes[-1] == "bypassPermissions": + rules.add(("bypass_permissions", "--permission-mode")) + for name, _arity, values in flags: + value = values[0] if values else None + if family == "claude" and name == "--dangerously-skip-permissions": + rules.add(("bypass_permissions", name)) + if family == "codex" and name == "--dangerously-bypass-approvals-and-sandbox": + rules.add(("bypass_approvals_and_sandbox", name)) + if family == "codex" and name == "--sandbox" and value == "danger-full-access": + rules.add(("danger_full_access", name)) + return rules + + +def _action_rules(spec: _AgentAction, declared: list[tuple[str, Any]]) -> set[tuple[str, str]]: + """The documented widening rules an agent action's declared inputs meet, read from the raw text. + + Read here, before anything is withheld for publication, so redaction never + hides a rule. Only from text this reader reads exactly: an argument input + only when it is a plain list of words (:func:`_argument_words`), which a + ``${{ }}`` expression never is; a user gate's entries that hold no + expression, since GitHub substitutes one before the action reads the + input and the substituted text may add entries but cannot remove a + literal one; and a mode or settings input only when it holds none. Each + rule names the input it was read from. + """ + + rules: set[tuple[str, str]] = set() + for name, value in declared: + text = _setting_text(value) + if text is None: + continue + if name in spec.gates: + entries = _EXPRESSION_SPAN_RE.sub(_EXPRESSION_MARK, text).split(",") + if "*" in {entry.strip() for entry in entries}: + rules.add(("open_gate", name)) + for mode, widening, rule in spec.modes: + if name == mode and text == widening: + rules.add((rule, name)) + if name == spec.settings and not holds_expression(text) and _settings_bypass_permissions(text): + rules.add(("bypass_permissions", name)) + if name == spec.args: + words = _argument_words(text) + if words is None: + continue + if spec.family == "claude": + found = _claude_action_rules(words) + else: + found = { + rule for rule, _flag in _flag_rules( + "codex", _read_flags(words, _CODEX_FLAGS), config_selects_sandbox=False + ) + } + rules.update((rule, name) for rule in found) + return rules + + +def _with_rules(entry: dict[str, Any], rules: set[tuple[str, str]]) -> dict[str, Any]: + """``entry`` with the rules it meets, omitted when none, as the schema omits them.""" + + if not rules: + return entry + return {**entry, "widening_rules": [{"rule": rule, "setting": setting} for rule, setting in sorted(rules)]} + + +def agent_widening_rules(entry: dict[str, Any]) -> set[tuple[str, str]]: + """The documented widening rules one published agent launch meets (#823). + + Each is ``(rule, detail)``: ``detail`` names the input for an opened gate + (``allowed_non_write_users``) and is empty otherwise, so one rule spelled + two ways, or moved between the CLI and an action, is one rule. Read from + ``widening_rules``, which the reader decided from the declared text before + any of it was withheld; an unresolved launch meets none. + """ + + if entry.get("form") != "read": + return set() + return { + (str(item["rule"]), str(item["setting"]) if item["rule"] == "open_gate" else "") + for item in entry.get("widening_rules", []) + } + + +def agent_family(agent: str) -> str: + """``claude`` or ``codex``: which agent's rules a launch is read by.""" + + return _AGENT_ACTIONS[agent].family if agent in _AGENT_ACTIONS else agent + + +AgentWidening = tuple[str, str, str, dict[str, Any]] +#: A rule one job's launches of one agent family meet: ``(job, family, rule, detail)``. +_RuleKey = tuple[str, str, str, str] + +#: The agent action inputs each documented rule is read from; a user gate's +#: rule is read from the gate it names. +_RULE_SETTINGS: dict[str, frozenset[str]] = { + "bypass_permissions": frozenset({"claude_args", "settings"}), + "bypass_approvals_and_sandbox": frozenset({"codex-args"}), + "danger_full_access": frozenset({"sandbox", "permission-profile", "codex-args"}), + "unsafe_safety_strategy": frozenset({"safety-strategy"}), +} + + +def _unread_setting(entry: dict[str, Any], rule: str, detail: str) -> tuple[str, str] | None: + """The input of ``entry`` whose text this reader did not read for ``rule``, and why. + + ``expression``: it holds a ``${{ }}`` expression, whose substituted text + may meet the rule. ``unread``: it is an argument input that is not a + plain list of words (``unread_arguments``), so no rule was read from it. + """ + + names = frozenset({detail}) if rule == "open_gate" else _RULE_SETTINGS.get(rule, frozenset()) + for setting in entry.get("settings", []): + if setting["name"] not in names: + continue + if setting.get("unresolved_reason") == "unread_arguments": + return str(setting["name"]), "unread" + if setting.get("holds_expression"): + return str(setting["name"]), "expression" + return None + + +@dataclass(frozen=True) +class AgentRuleGains: + """The documented rules a workflow's agent launches meet at ``after`` and not at ``before`` (#823). + + Each widening is ``(job, rule, detail, entry)`` for the first launch at + ``after`` in that job that meets it. Only ``claimed`` is a widening: + + - ``unread_before``: the job has fewer steps at ``after`` than at + ``before`` that may launch that agent in a form this reader does not + read — a ``run:`` that mentions the agent CLI and is not one plain + command (``unread_agent_runs``), or an agent action whose ``with:`` is + not a mapping — and more launches of it this reader reads. The launch + that meets the rule may be one of those steps, rewritten in a form this + reader reads, and it may have met the rule already, as a job whose + permissions were not explicit may already have held a write scope + (``unknown_before``). Each names a step that is gone. Such a step that + remains is a limit and nothing more, and one that goes while no read + launch is added takes no gain from a launch that was read before + (#823 review cycle 4). + - ``setting_before``: at ``before``, the job's launch of that agent held, + in an input the rule is read from, text this reader did not read for a + rule — a ``${{ }}`` expression (``expression``), whose substituted text + may already have met it, or an argument input that is not a plain list + of words (``unread``). Each names that input and which. + - ``moved``: the rule left another job whose launch that met it left that + job — the job no longer exists, or the same launch now runs here while + each launch of that agent this job had still runs here or in that job — + as when a job is renamed, an agent step moves to another job or two + jobs swap launches (#823 review). A job that remains may still run its + launch in a form this reader does not read, so a launch that only stops + being read has not left it, and one edited in place into the launch + that job had gains the rule (#823 review cycle 3). The launch has not + left while the job it met the rule in has any unread step of that + agent, wherever it stands, or a read launch of that agent whose input + the rule is read from it did not read (#823 review cycles 5 and 6); + nor while any other job but this one has more such steps and launches + than it had before, as a job new at the head holding one has, because + the launch may be one of them: the job it met the rule in may be that + job under a new name (#823 review cycle 7). So two jobs renamed at + once, one of them holding an unread step, claim the gain, in the safe + direction. Each names the launch it left, the way a step reference + moved between jobs adds no scope (#771). + """ + + claimed: list[AgentWidening] + unread_before: list[tuple[AgentWidening, dict[str, Any]]] + setting_before: list[tuple[AgentWidening, str, str]] + moved: list[tuple[AgentWidening, dict[str, Any]]] + + +def agent_rule_gains(before: dict[str, Any] | None, after: dict[str, Any] | None) -> AgentRuleGains: + """Which documented rules the workflow's agent launches gain, and which of them are claimed. + + Keyed by job, agent family and rule, so moving a launch between steps or + spellings (``--dangerously-skip-permissions`` and + ``--permission-mode bypassPermissions`` are one rule) gains nothing. + """ + + def met(grant: dict[str, Any] | None) -> dict[_RuleKey, list[dict[str, Any]]]: + found: dict[_RuleKey, list[dict[str, Any]]] = {} + for entry in (grant or {}).get("agent_launches", []): + for rule, detail in sorted(agent_widening_rules(entry)): + key = (str(entry["job"]), agent_family(str(entry["agent"])), rule, detail) + found.setdefault(key, []).append(entry) + return found + + def by_job(entries: list[dict[str, Any]]) -> dict[tuple[str, str], list[dict[str, Any]]]: + found: dict[tuple[str, str], list[dict[str, Any]]] = {} + for entry in entries: + found.setdefault((str(entry["job"]), agent_family(str(entry["agent"]))), []).append(entry) + return found + + def launches(grant: dict[str, Any] | None) -> dict[tuple[str, str], list[dict[str, Any]]]: + return by_job((grant or {}).get("agent_launches", [])) + + def unread_steps(grant: dict[str, Any] | None) -> dict[tuple[str, str], list[dict[str, Any]]]: + # The steps that may launch an agent in a form this reader does not read. + return by_job([ + *(entry for entry in (grant or {}).get("agent_launches", []) if entry.get("form") != "read"), + *(grant or {}).get("unread_agent_runs", []), + ]) + + old, new = met(before), met(after) + launched_before, launched_after = launches(before), launches(after) + unread_before, unread_after = unread_steps(before), unread_steps(after) + lost = [key for key in old if key not in new] + gained = [key for key in new if key not in old] + + def compared(entries: list[dict[str, Any]]) -> set[tuple[Any, ...]]: + # A launch as it is compared, less its job. + return {agent_launch_key(entry)[1:] for entry in entries} + + def unread_forms(grant_unread: dict[tuple[str, str], list[dict[str, Any]]], + grant_launched: dict[tuple[str, str], list[dict[str, Any]]], + job: str, family: str, rule: str, detail: str) -> int: + # How many steps of the job may run a launch of that agent that this + # reader cannot tell meets the rule: its unread steps of that agent, + # and its read launches of it holding text in an input the rule is + # read from that this reader did not read. + return len(grant_unread.get((job, family), [])) + sum( + 1 for entry in grant_launched.get((job, family), []) + if entry.get("form") == "read" and _unread_setting(entry, rule, detail) + ) + + def may_still_meet(source: _RuleKey, target: _RuleKey) -> bool: + # The launch that met the rule may still run, in a form this reader + # does not read, somewhere other than the job gaining the rule. Such a + # launch has not left for that job. An unread step carries no text to + # tell which launch it is, so: + # - in the losing job, any one counts, wherever it stands and whether + # or not it was there before (#823 review cycles 5 and 6); + # - in any other job, one counts when that job has more of them than + # it had before, as a job new at the head holding one has — so a + # renamed losing job cannot hide the launch it kept, quoted, under + # its new name (#823 review cycle 7). + job, family, rule, detail = source + if unread_forms(unread_after, launched_after, job, family, rule, detail): + return True + others = {name for name, agent in (*unread_after, *launched_after) if agent == family} - {job, target[0]} + return any( + unread_forms(unread_after, launched_after, other, family, rule, detail) + > unread_forms(unread_before, launched_before, other, family, rule, detail) + for other in others + ) + + def same_launch(source: _RuleKey, target: _RuleKey) -> bool: + # The launch that met the rule in the losing job now runs here, the + # losing job does not still run it in a form this reader does not + # read, and none of this job's own launches of that agent changed in + # place: each still runs here or now runs in the losing job, as when + # two jobs swap. A job whose launch was edited into the one the + # losing job had, while the losing job still exists, gained it (#823 + # review cycle 3). + if not compared(old[source]) & compared(new[target]): + return False + if may_still_meet(source, target): + return False + family = target[1] + remaining = compared([ + entry for job in (target[0], source[0]) for entry in launched_after.get((job, family), []) + ]) + return compared(launched_before.get((target[0], family), [])) <= remaining + + jobs_after = {str(context["job"]) for context in (after or {}).get("permission_contexts", [])} + + def job_left(source: _RuleKey, target: _RuleKey) -> bool: + # The losing job no longer exists, as when it is renamed, and no + # other job may run its launch unread. A job that remains may still + # run its launch in a form this reader does not read, such as `npx`, + # so its launch is not taken to have left (#823 review cycle 3); a + # job renamed while it runs its launch unread is a new job holding + # an unread step (#823 review cycle 7). + return source[0] not in jobs_after and not may_still_meet(source, target) + + # A rule moved when the launch that met it left the losing job. The same + # launch arriving is matched first, so which of two gaining jobs a rule + # moved to does not depend on the order the jobs are declared in (#823 review). + sources: dict[_RuleKey, _RuleKey] = {} + for left in (same_launch, job_left): + for key in gained: + if key in sources: + continue + source = next( + (other for other in lost if other[1:] == key[1:] and left(other, key)), None + ) + if source is not None: + lost.remove(source) + sources[key] = source + + def read(entries: list[dict[str, Any]]) -> int: + return sum(1 for entry in entries if entry.get("form") == "read") + + gains = AgentRuleGains(claimed=[], unread_before=[], setting_before=[], moved=[]) + for key in gained: + entries = new[key] + job, family, rule, detail = key + widening: AgentWidening = (job, rule, detail, entries[0]) + if key in sources: + gains.moved.append((widening, old[sources[key]][0])) + continue + before_launches = launched_before.get((job, family), []) + was_unread = unread_before.get((job, family), []) + still_unread = unread_after.get((job, family), []) + if len(was_unread) > len(still_unread) and read(launched_after.get((job, family), [])) > read( + before_launches + ): + remaining = {str(entry["step"]) for entry in still_unread} + gone = next((entry for entry in was_unread if str(entry["step"]) not in remaining), was_unread[0]) + gains.unread_before.append((widening, gone)) + continue + setting = next( + (found for entry in before_launches if (found := _unread_setting(entry, rule, detail))), None + ) + if setting is not None: + gains.setting_before.append((widening, *setting)) + continue + gains.claimed.append(widening) + return gains + + +def gained_agent_widenings(before: dict[str, Any] | None, after: dict[str, Any] | None) -> list[AgentWidening]: + """Documented widening rules the workflow's agent launches gain and claim (``AgentRuleGains.claimed``).""" + + return agent_rule_gains(before, after).claimed + + +def uncompared_agent_launch_texts(grant: dict[str, Any]) -> list[str]: + """One message per agent launch setting, unread agent step or checkout ref a workflow does not compare (#823). + + Not blocking, like an unread secret value (#693). A launch's job, agent, + form and widening rules are still compared, so adding, removing or + re-forming one, or gaining a documented rule, is a row; only an edit + inside what is named here is not reported. A setting holding + credential-shaped text (#823 review) is compared by its redacted text and + its rules, so only an edit inside what is redacted is not reported. An + argument input that is not a plain list of words is compared by a digest + and read for no rule. An unread ``run:`` agent step is compared not at + all: it gives no row, whatever is edited (#823 review cycle 4). A + redacted checkout ref is not named here: :func:`_uncompared_workflow_text` + makes it a blocking limit, as a redacted step reference is (#767). + """ + + texts: list[str] = [] + for entry in grant.get("agent_launches", []): + where = f"{entry['job']}/{entry['step']} ({entry['agent']})" + if entry.get("unresolved_reason") == "inputs_not_a_mapping": + texts.append( + f"the agent launch at {where} is a step whose `with:` is not a mapping; its " + "settings are neither published nor compared, so an edit to them is not reported" + ) + for setting in entry.get("settings", []): + unread = setting.get("unresolved_reason") + if unread == "redacted": + texts.append( + f"the {setting['name']} value of the agent launch at {where} contains " + "credential-shaped text; it is published redacted and compared as published, " + "so an edit inside what is redacted that gains no documented widening rule " + "is not reported" + ) + continue + if unread == "unread_arguments": + texts.append( + f"the {setting['name']} value of the agent launch at {where} is not a plain list " + "of words this audit reads: it holds a quote, a `${{ }}` expression, `$`, a " + "backtick, a comment, a shell operator, JSON, `--settings` or `--mcp-config`, or " + "another character outside the plain set; none of its text is published and no " + "documented widening rule is read from it, and it is compared only by a digest, " + "so an edit to it is a change, never a widening" + ) + continue + what = { + "not_a_string": "is not a string", + "unparsed_json": ( + "holds text that starts like JSON and does not parse, so the values it may " + "hold cannot be told apart from its key names" + ), + }.get(str(unread)) + if what: + texts.append( + f"the {setting['name']} value of the agent launch at {where} {what}; it is " + "neither published nor compared, so an edit to it that gains no documented " + "widening rule is not reported" + ) + for entry in grant.get("unread_agent_runs", []): + texts.append( + f"the `run:` at {entry['job']}/{entry['step']} mentions {entry['agent']} and is not read " + "as an agent launch: only a single-line command of plain words run by bash or sh, whose " + "program is `claude -p` or `codex exec`, is read, so this step may start an agent that is " + "neither published nor compared, and adding, removing or editing it gives no row" + ) + for entry in grant.get("checkout_refs", []): + unread = entry.get("unresolved_reason") + what = { + "not_a_string": "a ref that is not a string", + "inputs_not_a_mapping": "a `with:` that is not a mapping", + }.get(str(unread)) + if what: + texts.append( + f"the checkout at {entry['job']}/{entry['step']} declares {what}; it is " + "neither published nor compared, so an edit to it is not reported" + ) + # Two steps whose labels publish alike name one limit once. + return list(dict.fromkeys(texts)) + + #: A whole ``secrets:`` value of this form names its source (#693). Only the #: property form is read; ``secrets['NAME']`` and every other expression stay #: unresolved rather than guessed. The ``secrets`` context name is matched as @@ -1996,12 +3457,21 @@ def _uncompared_workflow_text( ) -> str | None: """Why part of a workflow grant is published but cannot be compared, or ``None``. - One rule for every compared workflow text (#767, #693). A redacted step - reference, reusable target, or secret name could publish the same text as - a different one, so comparing the display would read a change as equal. + One rule for every compared workflow reference (#767, #693). A redacted + step reference, reusable target, secret name or checkout ref (#823) could + publish the same text as a different one, so comparing the display would + read a change to the code a job runs, or to where a secret goes, as equal. That is a blocking limit: a changed workflow refuses, and an unchanged one is named (#721). + A redacted agent launch setting is not counted here (#823 review). Its + documented widening rules are read from the declared text before it is + redacted, so comparing its published text and its rules loses no + direction, and ordinary prose such as "never print bearer tokens" in a + system prompt is as credential-shaped to the label redaction as a token + is. :func:`uncompared_agent_launch_texts` names it, not blocking, as a + redacted label is compared by what it publishes (#802). + A job id, trigger or permission scope name is compared by its published label (#802). One redacted label is still a distinct label, so it refuses nothing; ``collided`` names the kinds of which two distinct raw labels in @@ -2020,6 +3490,10 @@ def _uncompared_workflow_text( ("a reusable workflow secret name", any( entry["unresolved_reason"] == "redacted" for entry in mappings )), + # A checkout ref names the code a job runs, as a step reference does (#823). + ("a checkout ref", any( + item.get("unresolved_reason") == "redacted" for item in grant.get("checkout_refs", []) + )), ) if present ] merged = [kind for kind in _LABEL_KINDS if kind in collided] @@ -2090,10 +3564,12 @@ def _workflow_grant( Each label is published by :func:`published_workflow_label` before it is used anywhere, so the job in ``permission_contexts``, ``reusable_calls``, - ``step_actions``, the ``write_scopes`` and ``effective_write_scopes`` - prefixes, every row built from them, and ``config_sha256`` all hold the - same label, and none holds the raw text. ``collided`` receives the kinds - of label of which two distinct raw values publish alike. + ``step_actions``, ``agent_launches``, ``unread_agent_runs``, + ``checkout_refs``, the ``write_scopes`` and ``effective_write_scopes`` + prefixes, every row built + from them, and ``config_sha256`` all hold the same label, and none holds + the raw text. ``collided`` receives the kinds of label of which two + distinct raw values publish alike. """ if not isinstance(data, dict): @@ -2110,6 +3586,9 @@ def _workflow_grant( permission_contexts: list[dict[str, Any]] = [] reusable_calls: list[dict[str, Any]] = [] step_actions: list[dict[str, Any]] = [] + agent_launches: list[dict[str, Any]] = [] + unread_agent_runs: list[dict[str, Any]] = [] + checkout_refs: list[dict[str, Any]] = [] def collect(perms: Any, where: str) -> None: if perms == "write-all": @@ -2148,6 +3627,7 @@ def collect(perms: Any, where: str) -> None: if not isinstance(steps, list): step_actions.append(_unreadable_step(label, "steps", "steps_not_a_list")) steps = [] + job_launches: list[dict[str, Any]] = [] for index, step in enumerate(steps): if not isinstance(step, dict): step_actions.append( @@ -2157,6 +3637,21 @@ def collect(perms: Any, where: str) -> None: action = _step_action(label, step, index) if action is not None: step_actions.append(action) + launches, unread = _step_agent_launches( + label, step, index, _declared_shell(step, job, data) + ) + job_launches.extend(launches) + unread_agent_runs.extend(unread) + checkout = _checkout_ref(label, step, index) + if checkout is not None: + checkout_refs.append(checkout) + if job_launches: + # Context for the note on the row, never compared (#823). + secrets = _job_secrets(job, data.get("env")) + agent_launches.extend( + {**launch, "job_secrets": secrets} if secrets else launch + for launch in job_launches + ) pull_target = "pull_request_target" in triggers write_all = any(entry.endswith(": write-all") for entry in effective_write_scopes) unknown = not permission_contexts or any( @@ -2177,6 +3672,14 @@ def collect(perms: Any, where: str) -> None: # Omitted when empty, as the schema omits it, so a workflow whose steps # declare no listed reference keeps its earlier fingerprint. projection["step_actions"] = step_actions + # Omitted when empty for the same reason (#823). + if agent_launches: + projection["agent_launches"] = agent_launches + if unread_agent_runs: + # Named as a limit, never compared (#823 review cycle 4). + projection["unread_agent_runs"] = unread_agent_runs + if checkout_refs: + projection["checkout_refs"] = checkout_refs return { **_grant_base( host="github", scope="repository", source=source, kind="workflow", @@ -2335,7 +3838,12 @@ def _collect_file( _inventory_issue( kind="unsupported", host=host, source=source, message=text, blocking=False, ) - for text in uncompared_secret_mapping_texts(grant) + for text in ( + *uncompared_secret_mapping_texts(grant), + # An agent launch or checkout ref this reader does not + # compare narrows that entry, not the workflow (#823). + *uncompared_agent_launch_texts(grant), + ) ) elif host == "codex" and kind == "requirements": grants.extend(_codex_requirement_grants(data, scope=scope, source=source)) @@ -4130,15 +5638,29 @@ def _same_workflow_grant(before: dict | None, after: dict | None) -> bool: # Step references compare as each job's multiset of declared references: # a reordered, renamed or re-id'd step that declares the same reference # changes no modeled fact, and no execution-order dependency is evaluated. - ignored = {"write_scopes", "config_sha256", "step_actions"} + # Agent launches and checkout refs compare the same way (#823): as each + # job's multiset of declared facts, never by step label, and a launch's + # `job_secrets` is context for the row, never compared. + # An unread `run:` agent step is a named limit and is never compared, so + # adding, removing or editing one gives no row (#823 review cycle 4). + ignored = { + "write_scopes", "config_sha256", "step_actions", "agent_launches", "checkout_refs", + "unread_agent_runs", + } def comparable(grant: dict[str, Any]) -> dict[str, Any]: # Both sides are v0.4+ shapes here, and a drift payload refuses a - # v0.4/v0.5 baseline holding a workflow, so a missing key is "none". + # v0.4-v0.6 baseline holding a workflow, so a missing key is "none". projection = {key: value for key, value in grant.items() if key not in ignored} projection["step_actions"] = sorted( step_action_key(item) for item in grant.get("step_actions", []) ) + projection["agent_launches"] = sorted( + agent_launch_key(item) for item in grant.get("agent_launches", []) + ) + projection["checkout_refs"] = sorted( + checkout_ref_key(item) for item in grant.get("checkout_refs", []) + ) # Named secrets compare as each call's set of facts, whatever order a # saved snapshot lists them in (#693). projection["reusable_calls"] = [ @@ -4303,6 +5825,10 @@ def inherited_calls(grant): } if inherited_calls(after) - inherited_calls(previous): signals.append(f"workflow_secrets_inherited_{prefix}: {after['source']}") + # Only a documented rule gained by a job's agent launches widens + # (#823); every other agent-launch or checkout edit is a change. + if gained_agent_widenings(before, after): + signals.append(f"workflow_agent_widened_{prefix}: {after['source']}") return sorted(set(signals)) @@ -4460,20 +5986,14 @@ def _incomparable_payload( #: v0.4 inventory refused such a link as unreadable, and an incomplete #: inventory can never be saved, so a v0.4 baseline holds no artifact that v0.5 #: would describe differently. Accepting it keeps every saved baseline usable. -#: v0.6 adds workflow step references (#771); the rule below narrows which -#: v0.4/v0.5 baselines that acceptance still covers. v0.7 adds only the hook -#: and MCP detail no comparison reads (#819), so a v0.6 baseline compares as -#: it did. +#: v0.6 adds workflow step references (#771) and v0.7 agent launches and +#: checkout refs (#823); the rules below narrow which older baselines that +#: acceptance still covers. _COMPARABLE_BASELINE_SCHEMA_VERSIONS = frozenset( {"0.4", "0.5", "0.6", HOST_GRANTS_BASELINE_SCHEMA_VERSION} ) -#: Baseline versions ``audit --host --save-baseline`` may replace (#819). Every -#: grant a v0.6 baseline holds compares exactly as its v0.7 reading does: v0.7 -#: adds only the display members :data:`DISPLAY_ONLY_GRANT_FIELDS` names, which -#: no comparison and no digest reads, so refusing to replace one would make -#: every v0.6 baseline a move-aside step for no change in what is compared. -#: Older baselines stay refused, as they were (#771). +#: Baselines without workflow grants retain the v0.6 replacement route (#819). OVERWRITABLE_BASELINE_SCHEMA_VERSIONS = frozenset({"0.6", HOST_GRANTS_BASELINE_SCHEMA_VERSION}) #: Baseline versions whose workflow grants never read step action references @@ -4484,6 +6004,10 @@ def _incomparable_payload( #: claims were compared. _STEP_ACTIONS_UNREAD_BASELINE_SCHEMA_VERSIONS = frozenset({"0.4", "0.5"}) +#: Baseline versions whose workflow grants never read agent launches or +#: checkout refs (#823), by the same rule: silence is not evidence of none. +_AGENT_LAUNCHES_UNREAD_BASELINE_SCHEMA_VERSIONS = frozenset({"0.4", "0.5", "0.6"}) + def build_host_drift_payload( *, baseline: dict[str, Any], inventory: dict[str, Any], baseline_file: str @@ -4495,14 +6019,19 @@ def build_host_drift_payload( reasons.append( str(baseline.get("_load_error") or "unsupported_baseline_schema") ) - elif baseline.get("host_grants_schema_version") in _STEP_ACTIONS_UNREAD_BASELINE_SCHEMA_VERSIONS: + elif baseline.get("host_grants_schema_version") in _AGENT_LAUNCHES_UNREAD_BASELINE_SCHEMA_VERSIONS: + version = baseline.get("host_grants_schema_version") workflows = [ grant for grant in (baseline.get("inventory") or {}).get("grants", []) if grant.get("kind") == "workflow" ] - if workflows: + if workflows and version in _STEP_ACTIONS_UNREAD_BASELINE_SCHEMA_VERSIONS: reasons.append("baseline_workflow_step_actions_unavailable") - if any(grant.get("reusable_calls") for grant in workflows): + if workflows: + reasons.append("baseline_workflow_agent_launches_unavailable") + if version in _STEP_ACTIONS_UNREAD_BASELINE_SCHEMA_VERSIONS and any( + grant.get("reusable_calls") for grant in workflows + ): # Such a call's missing ``secret_mappings`` is not evidence that it # passed no named secret: those snapshots never read them (#693). reasons.append("baseline_reusable_workflow_secret_mappings_unavailable") diff --git a/src/agents_shipgate/report/host_comparison.py b/src/agents_shipgate/report/host_comparison.py index 0512b1f2d..9a7451708 100644 --- a/src/agents_shipgate/report/host_comparison.py +++ b/src/agents_shipgate/report/host_comparison.py @@ -281,19 +281,32 @@ def _side_text(item: HostComparisonCoverageItem) -> str: #: comparison compares them (#812). _REDACTED_VALUES = "(redacted values such as env values and apiKeyHelper are not compared)" +#: The same note for a workflow, which holds no `apiKeyHelper`: what its +#: grant does not compare, and where the unread agent steps are named (#823 +#: review cycle 7). `diff`, `verify` and `check` carry no limit for such a +#: step, so without this a change that only adds one reads as covered. +_WORKFLOW_UNREAD_TEXT = ( + "(text this entry does not read, such as a step's env or an unread agent step, " + "is not compared; audit --host names each unread agent step)" +) + def _redacted_values_note(source: str) -> str: - """The note on redacted values, only for a file that can hold them (#812 review cycle 5). + """The note on what is not compared, only for a file that can hold it (#812 review cycle 5). Not for a file the inventory reads as instructions (`AGENTS.md`, a `CLAUDE.md` link to it, a skill or a rule): it holds no `env` value or `apiKeyHelper`, and a docs-only change would print the note on every such - file it could not prove unchanged. The kind is the one the inventory gives - the file's path when it reads it; a path redacted past recognition keeps - the note. + file it could not prove unchanged. A workflow's note names what its grant + does not read instead (#823). The kind is the one the inventory gives the + file's path when it reads it; a path redacted past recognition keeps the + note. """ - return "" if _source_kind(source) == "instructions" else f" {_REDACTED_VALUES}" + kind = _source_kind(source) + if kind == "instructions": + return "" + return f" {_WORKFLOW_UNREAD_TEXT if kind == 'workflow' else _REDACTED_VALUES}" #: What each candidate rule names (#821), as the line says it. None of these diff --git a/src/agents_shipgate/schemas/contract.py b/src/agents_shipgate/schemas/contract.py index 82d61706b..4f9500136 100644 --- a/src/agents_shipgate/schemas/contract.py +++ b/src/agents_shipgate/schemas/contract.py @@ -214,20 +214,34 @@ # ``host_comparison.coverage`` and ``shipgate diff --json`` (capability diff # 0.3) the same block. It is evidence beside the rows and moves no state, # permission, route or row. A 0.19 verifier reads with coverage not recorded. -# v41 names the changed inputs a host comparison does not read (#821): verifier -# 0.21 and ``shipgate diff --json`` (capability diff 0.4) add a -# ``changed_not_read`` coverage item with its ``candidate`` rule, a -# ``read_sources_only`` that is ``false`` while one is named, and whether the -# comparison's changed files were examined. A name is never a row, a widening -# or a ``check`` violation. The one route it moves, on ``verify`` and -# ``verify --preview`` alike: a manifest-free comparison whose only -# host-relevant change is such an input, or a changed candidate it counts as -# not examined, is now published, on the host route's -# ``audit --host`` next action, instead of the setup route (``verify``) or the -# ``initialize`` next action (``verify --preview``). A 0.20 verifier reads -# with the search not recorded. ``MINIMUM_CONTROL_CONTRACT_VERSION`` stays at -# 21. -# v41 also keeps what a comparison established outside a plugin directory it +# v41, unreleased, carries three changes. It names the changed inputs a host +# comparison does not read (#821): verifier 0.21 and ``shipgate diff --json`` +# (capability diff 0.4) add a ``changed_not_read`` coverage item with its +# ``candidate`` rule, a ``read_sources_only`` that is ``false`` while one is +# named, and whether the comparison's changed files were examined. A name is +# never a row, a widening or a ``check`` violation. The one route it moves, on +# ``verify`` and ``verify --preview`` alike: a manifest-free comparison whose +# only host-relevant change is such an input, or a changed candidate it counts +# as not examined, is now published, on the host route's ``audit --host`` next +# action, instead of the setup route (``verify``) or the ``initialize`` next +# action (``verify --preview``). A 0.20 verifier reads with the search not +# recorded. And it reads how a coding agent is launched inside a workflow job +# (#823). Host-grants 0.6 shipped in 1.1.0, so this mints host-grants +# inventory, baseline and drift 0.7: a workflow grant adds ``agent_launches[]`` +# (a documented agent action's permission inputs, or the permission flags of a +# ``run:`` that is one plain ``claude -p`` / ``codex exec`` command, compared as +# text), ``unread_agent_runs[]`` (any other ``run:`` that mentions an agent CLI: +# a named limit, never compared) and ``checkout_refs[]`` (each +# ``actions/checkout`` step's ``with.ref``), each omitted when empty. Shell is +# not parsed; an argument input that is not a plain list of words is compared by +# a digest and read for no rule. Only a documented rule gained by a job's agent +# launches widens; every other edit is a ``changed`` row, and a workflow row +# that runs an agent names the job facts beside it. A 0.4-0.6 baseline holding a +# workflow grant is incomparable +# (``baseline_workflow_agent_launches_unavailable``); one without a workflow +# stays comparable. #823 moves neither verifier 0.21 nor capability diff 0.4: +# its rows keep their shape. ``MINIMUM_CONTROL_CONTRACT_VERSION`` stays at 21. +# And v41 keeps what a comparison established outside a plugin directory it # could not compare (#808), extended in place because v41 is unreleased: where # every blocking limit is a plugin-reference limit that its plugin directory # bounds, and no compared source depends on that directory, verifier 0.21 and diff --git a/src/agents_shipgate/schemas/host_grants.py b/src/agents_shipgate/schemas/host_grants.py index d99884739..b76e3ad38 100644 --- a/src/agents_shipgate/schemas/host_grants.py +++ b/src/agents_shipgate/schemas/host_grants.py @@ -701,6 +701,222 @@ class HostMcpServerGrantV7(HostMcpServerGrantV2): args_sha256: str | None = Field(pattern=r"^[0-9a-f]{64}$") +# v0.7 reads how a coding agent is launched inside a job (#823). The v0.6 +# workflow grant above stays frozen: a v0.4-v0.6 snapshot never read agent +# launches or checkout refs, so its silence cannot assert that none changed. +class HostWorkflowAgentSettingV7(BaseModel): + """One permission input or flag an agent launch declares, compared as text (#823). + + ``name`` is the documented input (``claude_args``, ``sandbox``, …) or the + flag's primary spelling (``--allowedTools`` for ``--allowed-tools`` too). + ``value`` is the declared text, stripped, as it may be published; a flag + that takes no value has ``null``. + + ``claude_args`` and ``codex-args`` are read only when they are a plain + list of words: letters, digits and ``_ . / : = , % + - ( )``, separated by + blanks or newlines, with no ``--settings`` or ``--mcp-config`` flag. Every + parser involved splits such text the same way, so it is published as + those words, one space apart. Any other value — holding a quote, a + ``${{ }}`` expression, ``$``, a backtick, a comment, a shell operator, + JSON or another character — is ``unread_arguments``: ``value`` is + ````, a short digest, so an edit to it is still a change + while none of its text is published; no documented widening rule is read + from it; and it records a non-blocking coverage issue naming its + ``job/step`` (#823 review cycle 4). A codex ``--config`` override keeps + its key; its value is ```` under ``env``, ``headers`` or a + secret-named key, as the host readers redact such values, published as + written for ``sandbox_mode``, ``default_permissions``, + ``approval_policy`` and ``model``, and ```` otherwise. + + Every other input is one value. A JSON object (a ``settings`` or + ``mcp_config`` value) publishes its shape and none of its free text: key + names, numbers, booleans and ``null``, with each string replaced by + ````, a short digest of what the host readers digest for it, + so an edit to it is still a change. ``env`` and ``headers`` values, + ``apiKeyHelper`` and every secret-named value are ````, as the + host readers redact them. The strings a host reader publishes are kept: + a ``permissions.allow``/``ask``/``deny`` rule and a documented Claude + Code setting's value such as ``defaultMode``, and an MCP server's command + name and its URL's scheme and host, each followed by the digest when it + drops something the digest reads (a command's arguments, a URL's query). + So an MCP server's arguments and a hook's command publish nothing, as + `.mcp.json` and `.claude/settings.json` do not (#823 review). A + ``settings`` or ``mcp_config`` value that neither starts like a JSON + object nor is a plain file path (path characters, and a ``${{ }}`` + expression only as a plain context reference) is ````, a + digest and none of its text (#823 review cycle 5). A URL in + other text publishes its scheme and host with ```` for its + path and query (#723). Other text — a prompt, a flag's value — is + published through the workflow label redaction (#802). A value it + rewrites is credential-shaped — a token, but also prose such as "never + print bearer tokens" — and is published redacted with + ``unresolved_reason: redacted``: it is compared as published, beside the + rules read from its declared text, and records a non-blocking coverage + issue naming its ``job/step``, because an edit inside what is redacted is + not reported. A value that is not a string (``not_a_string``), or one + holding text that starts like JSON and does not parse (``unparsed_json``), + is ``null`` and records a non-blocking coverage issue naming its + ``job/step``: it is neither published nor compared. + + ``holds_expression`` is ``true`` when an input other than an argument + input holds a ``${{ }}`` expression, which GitHub substitutes before the + action reads the input, and is omitted otherwise. A documented widening + rule is then read only from the entries of a user gate that hold none, + and from no mode or settings input, and a rule the launch gains in the + same job afterwards is not claimed, because the substituted text may + already have met it. + """ + + model_config = ConfigDict(extra="forbid") + + name: str + value: str | None + unresolved_reason: Literal[ + "not_a_string", "redacted", "unparsed_json", "unread_arguments", + ] | None = None + holds_expression: bool = Field(default=False, exclude_if=lambda value: not value) + + +class HostWorkflowAgentRuleV7(BaseModel): + """One documented widening rule an agent launch meets, and the setting it was read from (#823). + + Decided when the workflow is read, from the declared text, before any of + it is withheld for publication, so redaction never hides a rule. Only + text this reader reads exactly meets one: ``claude_args`` or + ``codex-args`` only when it is a plain list of words (never when it holds + a ``${{ }}`` expression), the entries of a user gate that hold no + expression, and a mode or ``settings`` input that holds none. Claude Code + settings written as JSON in the ``settings`` input meet + ``bypass_permissions`` when their ``defaultMode`` is + ``bypassPermissions``, read as the settings reader reads it; a path to a + settings file is not read. ``setting`` is the input (``claude_args``, + ``allowed_bots``, ``sandbox``, ``permission-profile``, …) or the CLI + flag's primary spelling. One rule compares as one whatever setting meets + it, except ``open_gate``, which is one rule per gate input. + """ + + model_config = ConfigDict(extra="forbid") + + rule: Literal[ + "bypass_permissions", + "bypass_approvals_and_sandbox", + "danger_full_access", + "unsafe_safety_strategy", + "open_gate", + ] + setting: str + + +class HostWorkflowAgentLaunchV7(BaseModel): + """A step that launches a known coding agent, read as text and never run (#823). + + ``agent`` is a documented action reference's ``owner/repo`` (the step's + ``uses:`` at any ref; the Claude base action also as the ``base-action`` + directory of ``anthropics/claude-code-action``), or a known agent CLI a + ``run:`` launches when the whole ``run:`` is one line of plain words + (letters, digits and ``_ . / : = , % + -``, separated by spaces or tabs), + run by ``bash``, ``sh`` or the runner's default shell, whose program, + after any ``NAME=value`` assignments, has the file name ``claude`` and + passes ``-p``/``--print``, or ``codex`` followed by ``exec`` (``e``). + ``form: read`` lists the documented permission inputs or flags the step + declares in ``settings``, and the documented widening rules they meet in + ``widening_rules``, omitted when none. ``form: unresolved`` is an agent + action whose ``with:`` is not a mapping (``inputs_not_a_mapping``), with + no settings, and records a non-blocking coverage issue. Any other + ``run:`` that mentions an agent CLI is not a launch: it is listed in + ``unread_agent_runs``. ``job_secrets`` names the secrets the step's job + references (``${{ secrets.NAME }}``) and the workflow-level ``env`` + passes: context for the row that names this step, never compared. + ``job`` and ``step`` are published labels (#802). + """ + + model_config = ConfigDict(extra="forbid") + + job: str + step: str + agent: Literal[ + "anthropics/claude-code-action", + "anthropics/claude-code-base-action", + "anthropics/claude-code-action/base-action", + "openai/codex-action", + "claude", + "codex", + ] + form: Literal["read", "unresolved"] + unresolved_reason: Literal["inputs_not_a_mapping"] | None = None + settings: list[HostWorkflowAgentSettingV7] = Field(default_factory=list) + widening_rules: list[HostWorkflowAgentRuleV7] = Field( + default_factory=list, exclude_if=lambda value: not value, + ) + job_secrets: list[str] = Field(default_factory=list, exclude_if=lambda value: not value) + + +class HostWorkflowUnreadAgentRunV7(BaseModel): + """A ``run:`` step that mentions a known agent CLI and is not read as an agent launch (#823 review cycle 4). + + Any ``run:`` holding ``claude`` or ``codex`` as a word of its own that is + not an agent launch this reader reads — more than one line or command, a + quote, an expansion, a redirection, a comment, a continuation, a + ``${{ }}`` expression, another program such as ``npx`` or ``timeout``, a + subcommand that is not a headless launch, or a declared ``shell:`` other + than ``bash`` or ``sh`` run on the script alone (so ``bash -c '…' {0}`` + too) — once for each agent CLI it mentions. It is a + named, non-blocking limit and nothing more: none of the step's text is + published, it is never compared, so adding, removing or editing it gives + no row, and it never says that the step starts, or does not start, an + agent. ``job`` and ``step`` are published labels (#802). + """ + + model_config = ConfigDict(extra="forbid") + + job: str + step: str + agent: Literal["claude", "codex"] + + +class HostWorkflowCheckoutRefV7(BaseModel): + """One ``actions/checkout`` step and the ``with.ref`` it declares, as text (#823). + + ``ref`` is ``null`` when the step declares none, or an empty one: the + checkout's default for the triggering event. A ref the label redaction + rewrites is published redacted with ``unresolved_reason: redacted`` and + makes the workflow a blocking limit, as a redacted step reference does + (#767): a ref names the code the job runs, as a step reference does. A + value that is not a string, or ``with:`` that is not a mapping, + is ``null`` with ``unresolved_reason`` and records a non-blocking coverage + issue. The ref is never resolved or fetched. + """ + + model_config = ConfigDict(extra="forbid") + + job: str + step: str + ref: str | None + unresolved_reason: Literal["not_a_string", "redacted", "inputs_not_a_mapping"] | None = None + + +class HostWorkflowGrantV7(HostWorkflowGrantV6): + """A v0.6 workflow grant plus the agent launches, unread agent steps and checkout refs its steps declare. + + Each list is present only when a step declares one. In a v0.7 grant an + absent list means the steps were read and declare none; the schema + version, not the key, separates that from a legacy grant that never read + them. ``unread_agent_runs`` is a named limit and is never compared. + ``access`` and ``risk`` still describe the workflow's token and triggers + alone. + """ + + agent_launches: list[HostWorkflowAgentLaunchV7] = Field( + default_factory=list, exclude_if=lambda value: not value, + ) + unread_agent_runs: list[HostWorkflowUnreadAgentRunV7] = Field( + default_factory=list, exclude_if=lambda value: not value, + ) + checkout_refs: list[HostWorkflowCheckoutRefV7] = Field( + default_factory=list, exclude_if=lambda value: not value, + ) + + HostGrantV7 = Annotated[ HostMcpServerGrantV7 | HostPermissionRuleGrantV2 @@ -711,7 +927,23 @@ class HostMcpServerGrantV7(HostMcpServerGrantV2): | HostPluginGrantV2 | HostProfileGrantV2 | HostRequirementGrantV2 - | HostWorkflowGrantV6 + | HostWorkflowGrantV7 + | HostInstructionGrantV2, + Field(discriminator="kind"), +] + + +HostBaselineGrantV7 = Annotated[ + HostMcpServerGrantV2 + | HostPermissionRuleGrantV2 + | HostPermissionModeGrantV2 + | HostHookGrantV2 + | HostSandboxGrantV2 + | HostAdditionalPathGrantV2 + | HostPluginGrantV2 + | HostProfileGrantV2 + | HostRequirementGrantV2 + | HostWorkflowGrantV7 | HostInstructionGrantV2, Field(discriminator="kind"), ] @@ -722,18 +954,14 @@ class HostGrantsInventoryV7(HostGrantsInventoryV6): grants: list[HostGrantV7] = Field(default_factory=list) -class HostGrantsBaselineV7(HostGrantsBaselineV6): - """A saved ``0.7`` baseline: the grants a ``0.6`` baseline holds, under the ``0.7`` version. - - A saved baseline holds no hook ``handlers`` and no MCP ``package`` or - ``args_sha256`` (#819): it is committed, and those members, read from a - user, managed or git-ignored file, would carry facts about files that were - never in the repository into it. No comparison, row or digest reads a - saved copy of them, so its ``inventory`` is the ``0.6`` snapshot, which - forbids them. - """ +class HostGrantsNormalizedSnapshotV7(HostGrantsNormalizedSnapshotV6): + # Saved baselines keep workflow evidence but omit display-only hook/MCP fields. + grants: list[HostBaselineGrantV7] = Field(default_factory=list) + +class HostGrantsBaselineV7(HostGrantsBaselineV6): host_grants_schema_version: Literal["0.7"] = "0.7" + inventory: HostGrantsNormalizedSnapshotV7 class HostGrantsDriftV7(HostGrantsDriftV6): diff --git a/tests/test_distribution_surface_parity.py b/tests/test_distribution_surface_parity.py index 79b6b8ec7..86e9dc8c5 100644 --- a/tests/test_distribution_surface_parity.py +++ b/tests/test_distribution_surface_parity.py @@ -226,13 +226,28 @@ def paths(self) -> list[Path]: # reserved coverage `scope`, and is read as incomparable by every # control route, so it adds no claim; every route to the same object, # and every refusal it must keep, is held by - # `tests/test_partial_host_comparison.py`. Naming an unchanged limit + # `tests/test_partial_host_comparison.py`. + # Naming an unchanged limit # the reader reached through an in-tree link (#822) rests on a Git # identity fact about the link and the file it lands on, the proof # every unchanged limit already rested on, so it restates no engine # answer and adds no claim; `tests/test_linked_unchanged_limits.py` # holds diff, verify, the PR comment and `check` to it, including the # shared plugin-reference limits `check` leaves out by the same proof. + # Its agent-launch cells and note (#823) restate no answer: direction + # comes from the engine's + # `workflow_agent_widened_*` expansion signal, itself read off the + # `widening_rules` the engine published on each launch; the `why`'s + # moved, unread-before and setting-before sentences read the same + # `agent_rule_gains` the signal is computed from, and the note reads the + # triggers, write scopes, secrets and checkout refs the engine already + # published on the grant, so it adds no claim; + # `tests/test_workflow_agent_launches.py` holds diff, verify, the PR + # comment, check and the control envelope to the same row. A + # workflow's coverage line naming what its grant does not read, in + # place of the redacted-values note, is fixed text chosen by the + # source's kind the inventory gives its path, and decides nothing, so + # it adds no claim either. # Its hook and MCP-argument # text (#819) renders the handlers, package and argument digest the # engine published on the grant, which hold no command or argument diff --git a/tests/test_reusable_workflow_secret_mappings.py b/tests/test_reusable_workflow_secret_mappings.py index 142803ef6..5ecfeef0b 100644 --- a/tests/test_reusable_workflow_secret_mappings.py +++ b/tests/test_reusable_workflow_secret_mappings.py @@ -786,6 +786,7 @@ def test_a_legacy_baseline_holding_a_reusable_call_does_not_assert_no_mappings(t assert drift["comparison_status"] == "incomparable" assert drift["incomparable_reasons"] == [ "baseline_reusable_workflow_secret_mappings_unavailable", + "baseline_workflow_agent_launches_unavailable", "baseline_workflow_step_actions_unavailable", ] assert drift["has_drift"] is None and drift["changes"] == [] diff --git a/tests/test_workflow_agent_launches.py b/tests/test_workflow_agent_launches.py new file mode 100644 index 000000000..700f41493 --- /dev/null +++ b/tests/test_workflow_agent_launches.py @@ -0,0 +1,2416 @@ +"""#823: how a coding agent is launched inside a workflow job is read as text. + +The workflow grant already read triggers, token permissions, reusable calls and +step `uses:` references (#771), and nothing that says how an agent is started. +It now lists each documented agent action's permission inputs, the permission +flags of a `run:` that is one plain `claude -p` / `codex exec` command, and each +`actions/checkout` step's `with.ref`. Nothing is executed, fetched or +evaluated. Only a documented rule a job's launches gain widens; every other +edit is `changed`; and a workflow row that runs an agent names the job facts +beside it. + +Shell is not parsed (#823 review cycle 4). A `run:` is read only as one line of +plain words whose program is a known agent CLI, and `claude_args` / +`codex-args` only as a plain list of words; every other form is a named, +non-blocking limit that publishes none of its text, and an unread `run:` step +is never compared, so it gives no row. +""" + +from __future__ import annotations + +import hashlib +import json +import subprocess +import time +from pathlib import Path + +import pytest +import yaml +from typer.testing import CliRunner + +from agents_shipgate.cli.main import app +from agents_shipgate.core.capability_diff_rows import capability_diff_rows +from agents_shipgate.core.host_grants import ( + _argument_words, + _run_command, + _uncompared_workflow_text, + _workflow_grant, + diff_host_grants, + host_grant_expansion_signals, + uncompared_agent_launch_texts, +) + +ROOT = Path(__file__).resolve().parents[1] +SOURCE = ".github/workflows/agent.yml" +CLAUDE = "anthropics/claude-code-action@v1" +HEAD_SHA = "${{ github.event.pull_request.head.sha }}" + + +def _workflow(*steps, trigger="pull_request", permissions=None, jobs=None, env=None, defaults=None): + data = { + "on": trigger, + "permissions": permissions if permissions is not None else {"contents": "read", "pull-requests": "read"}, + "jobs": jobs if jobs is not None else {"review": {"runs-on": "ubuntu-latest", "steps": list(steps)}}, + } + if env is not None: + data["env"] = env + if defaults is not None: + data["defaults"] = defaults + return data + + +def _agent(claude_args="--allowedTools Read", **extra): + return {"uses": CLAUDE, "with": {"claude_args": claude_args, **extra}} + + +def _reproduction(trigger="pull_request", pr="read", claude_args="--allowedTools Read", run="echo done", ref=None): + """The workflow of #823's reproduction, as its `wf` shell function writes it, with plain `claude_args`.""" + + checkout = {"uses": "actions/checkout@v4", **({"with": {"ref": ref}} if ref else {})} + return _workflow( + checkout, _agent(claude_args), {"run": run}, + trigger=trigger, permissions={"contents": "read", "pull-requests": pr}, + ) + + +def _grant(value): + return _workflow_grant(value, source=SOURCE) + + +def _changes(before, after): + return diff_host_grants({"grants": [_grant(before)]}, {"grants": [_grant(after)]}) + + +def _rows(before, after): + changes = _changes(before, after) + payload = {"changes": changes, "expansion_signals": host_grant_expansion_signals(changes)} + return capability_diff_rows(payload) + + +def _launches(value): + return _grant(value).get("agent_launches", []) + + +def _unread(value): + return _grant(value).get("unread_agent_runs", []) + + +def _digest(text): + """What a withheld string publishes: a digest of what the host readers digest for it.""" + + from agents_shipgate.core.host_grants import redacted_config_sha256 + + return f"" + + +# --- the four cases of the reproduction -------------------------------------------- + + +def test_args_gaining_bypass_permissions_is_one_widened_row_naming_job_step_and_both_values(): + row, = _rows( + _reproduction(), + _reproduction(claude_args="--permission-mode bypassPermissions --allowedTools Bash(git:status)"), + ) + + assert row.subject == f"github {SOURCE}" + assert (row.direction, row.expands) == ("widened", True) + assert "review/steps[1]: runs anthropics/claude-code-action with claude_args: --allowedTools Read" in row.before + assert ( + "review/steps[1]: runs anthropics/claude-code-action with claude_args: " + "--permission-mode bypassPermissions --allowedTools Bash(git:status)" + ) in row.after + assert "an agent launch now skips permission checks (bypassPermissions) (review/steps[1])" in row.why + + +def test_the_quoted_args_of_the_reproduction_are_a_changed_row_that_publishes_only_digests(): + """#823 review cycle 4 scope: quoted `claude_args` is not read, only compared by a digest.""" + + before, after = '--allowedTools "Read"', '--permission-mode bypassPermissions --allowedTools "Bash(*)"' + assert host_grant_expansion_signals(_changes(_reproduction(claude_args=before), _reproduction(claude_args=after))) == [] + row, = _rows(_reproduction(claude_args=before), _reproduction(claude_args=after)) + + assert (row.direction, row.expands) == ("changed", False) + assert f"claude_args (not read; digest {_digest(before)})" in row.before + assert f"claude_args (not read; digest {_digest(after)})" in row.after + assert "an agent launch's declared settings changed (review/steps[1])" in row.why + assert ( + "an agent launch's argument input is not a plain list of words this audit reads (claude_args at " + "review/steps[1]); none of its text is published and it is compared by a digest only, so this row " + "does not say whether it meets a documented widening rule" + ) in row.why + for text in ("bypassPermissions", "Bash(*)", '"Read"'): + assert text not in row.before + row.after + limit, = uncompared_agent_launch_texts(_grant(_reproduction(claude_args=after))) + assert limit.startswith( + "the claude_args value of the agent launch at review/steps[1] (anthropics/claude-code-action) is not " + "a plain list of words this audit reads" + ) + + +def test_a_literal_claude_run_step_is_one_changed_row_with_its_permission_flags(): + row, = _rows( + _reproduction(), + _reproduction(run="claude -p --permission-mode acceptEdits --allowedTools Edit Summarize"), + ) + + assert (row.direction, row.expands) == ("changed", False) + assert "review/steps[2]" not in row.before + # A variadic flag reads every following word up to the next flag, as the + # CLI reads it, so the trailing prompt word is part of --allowedTools. + assert ( + "review/steps[2]: runs claude -p with --allowedTools Edit Summarize; --permission-mode acceptEdits" + ) in row.after + assert "a step now launches an agent (review/steps[2])" in row.why + assert "not counted as a widening" in row.why + + +def test_a_head_ref_checkout_is_one_changed_row_naming_the_default_and_the_new_ref(): + row, = _rows( + _reproduction(trigger="pull_request_target"), + _reproduction(trigger="pull_request_target", ref=HEAD_SHA), + ) + + assert (row.direction, row.expands) == ("changed", False) + assert "review/steps[0]: checkout of the default ref" in row.before + assert f"review/steps[0]: checkout of ref {HEAD_SHA}" in row.after + assert "a checkout's declared ref changed (review/steps[0])" in row.why + assert ( + "an agent runs at review/steps[1] (anthropics/claude-code-action) beside the " + "untrusted-input trigger pull_request_target and a checkout of pull request code (review/steps[0])" + ) in row.why + + +def test_a_checkout_step_on_one_side_only_is_worded_as_added_or_removed(): + """#823 review cycle 3 (P3): a default checkout in an added job is not a changed ref.""" + + checkout = {"uses": "actions/checkout@v4"} + before = _jobs(test=[checkout, {"run": "make test"}]) + after = _jobs(test=[checkout, {"run": "make test"}], lint=[checkout, {"run": "make lint"}]) + + row, = _rows(before, after) + assert "lint/steps[0]: checkout of the default ref" in row.after + assert "a step now declares a checkout (lint/steps[0]); a ref names which commit's code" in row.why + assert "declared ref changed" not in row.why + + removed, = _rows(after, before) + assert "a step no longer declares a checkout (lint/steps[0])" in removed.why + assert "declared ref changed" not in removed.why + + +def test_an_untrusted_trigger_with_a_write_scope_names_the_agent_step_it_now_reaches(): + row, = _rows(_reproduction(), _reproduction(trigger="issue_comment", pr="write")) + + assert (row.direction, row.expands) == ("widened", True) + assert row.why.startswith("grants write permissions to workflow jobs") + assert ( + "an agent runs at review/steps[1] (anthropics/claude-code-action) beside the " + "untrusted-input trigger issue_comment and the write scope pull-requests" + ) in row.why + + +# --- what is read ------------------------------------------------------------------ + + +def test_only_the_documented_inputs_of_a_known_action_are_listed(): + launch, = _launches(_workflow({ + "uses": "Anthropics/Claude-Code-Action@0123456789abcdef0123456789abcdef01234567", + "with": { + "prompt": "Review this", + "anthropic_api_key": "${{ secrets.ANTHROPIC_API_KEY }}", + "claude_args": "--max-turns 5", + "Allowed_Non_Write_Users": "octocat", + "use_sticky_comment": True, + }, + })) + + assert launch["agent"] == "anthropics/claude-code-action" + assert (launch["job"], launch["step"], launch["form"]) == ("review", "steps[0]", "read") + assert launch["settings"] == [ + {"name": "allowed_non_write_users", "value": "octocat", "unresolved_reason": None}, + {"name": "claude_args", "value": "--max-turns 5", "unresolved_reason": None}, + ] + assert launch["job_secrets"] == ["ANTHROPIC_API_KEY"] + + +def test_the_codex_action_inputs_are_listed(): + launch, = _launches(_workflow({ + "uses": "openai/codex-action@v1", + "with": {"sandbox": "workspace-write", "safety-strategy": "drop-sudo", "prompt": "x", "allow-bots": False}, + })) + + assert launch["agent"] == "openai/codex-action" + assert launch["settings"] == [ + {"name": "allow-bots", "value": "false", "unresolved_reason": None}, + {"name": "safety-strategy", "value": "drop-sudo", "unresolved_reason": None}, + {"name": "sandbox", "value": "workspace-write", "unresolved_reason": None}, + ] + + +def test_an_action_step_with_no_inputs_is_still_an_agent_launch(): + launch, = _launches(_workflow({"uses": "anthropics/claude-code-base-action@beta"})) + assert (launch["agent"], launch["form"], launch["settings"]) == ("anthropics/claude-code-base-action", "read", []) + + +def test_a_literal_claude_command_publishes_its_permission_flags_under_their_primary_spelling(): + launch, = _launches(_workflow({ + "name": "Review", + "run": ( + "claude --print --allowed-tools=Read --disallowedTools Bash " + "--model sonnet --dangerously-skip-permissions --add-dir ../docs secret-prompt-text" + ), + })) + + assert (launch["agent"], launch["step"], launch["form"]) == ("claude", "Review", "read") + assert launch["settings"] == [ + {"name": "--add-dir", "value": "../docs secret-prompt-text", "unresolved_reason": None}, + {"name": "--allowedTools", "value": "Read", "unresolved_reason": None}, + {"name": "--dangerously-skip-permissions", "value": None, "unresolved_reason": None}, + {"name": "--disallowedTools", "value": "Bash", "unresolved_reason": None}, + ] + assert "sonnet" not in json.dumps(launch) + + +def test_a_literal_codex_exec_command_publishes_its_permission_flags(): + launch, = _launches(_workflow({"run": "codex e -s danger-full-access --yolo -c model=o3 fix-it"})) + + assert (launch["agent"], launch["form"]) == ("codex", "read") + assert launch["settings"] == [ + {"name": "--config", "value": "model=o3", "unresolved_reason": None}, + {"name": "--dangerously-bypass-approvals-and-sandbox", "value": None, "unresolved_reason": None}, + {"name": "--sandbox", "value": "danger-full-access", "unresolved_reason": None}, + ] + + +def test_literal_assignments_before_the_command_are_skipped_and_never_published(): + launch, = _launches(_workflow({"run": "CI=true ANTHROPIC_API_KEY=sk-canary claude -p --permission-mode plan go"})) + assert launch["form"] == "read" + assert launch["settings"] == [{"name": "--permission-mode", "value": "plan", "unresolved_reason": None}] + assert "sk-canary" not in json.dumps(_grant(_workflow({"run": "CI=true ANTHROPIC_API_KEY=sk-canary claude -p go"}))) + + +def test_an_agent_cli_named_by_its_path_is_read_by_its_file_name(): + for run in ("./node_modules/.bin/claude -p --dangerously-skip-permissions go", "/usr/local/bin/codex exec --yolo go"): + launch, = _launches(_workflow({"run": run})) + assert launch["form"] == "read" and "widening_rules" in launch + assert run.split()[0] not in json.dumps(launch) + + +def test_the_base_action_directory_of_the_claude_action_is_read_as_the_base_action(): + launch, = _launches(_workflow({ + "uses": "anthropics/claude-code-action/base-action@v1", + "with": {"claude_args": "--dangerously-skip-permissions", "allowed_bots": "*"}, + })) + + assert (launch["agent"], launch["form"]) == ("anthropics/claude-code-action/base-action", "read") + # `allowed_bots` is not a base-action input, so it is neither listed nor a rule. + assert launch["settings"] == [ + {"name": "claude_args", "value": "--dangerously-skip-permissions", "unresolved_reason": None}, + ] + assert launch["widening_rules"] == [{"rule": "bypass_permissions", "setting": "claude_args"}] + + +def test_inputs_that_are_not_a_mapping_are_unresolved(): + launch, = _launches(_workflow({"uses": CLAUDE, "with": ["claude_args"]})) + assert (launch["form"], launch["unresolved_reason"], launch["settings"]) == ( + "unresolved", "inputs_not_a_mapping", [], + ) + limit, = uncompared_agent_launch_texts(_grant(_workflow({"uses": CLAUDE, "with": ["claude_args"]}))) + assert "a step whose `with:` is not a mapping" in limit + + +def test_every_checkout_step_records_its_declared_ref(): + grant = _grant(_workflow( + {"uses": "actions/checkout@v4"}, + {"id": "head", "uses": "actions/checkout@v4", "with": {"ref": HEAD_SHA, "fetch-depth": 0}}, + {"uses": "actions/checkout@v4", "with": {"ref": ""}}, + {"uses": "actions/checkout@v4", "with": {"ref": ["main"]}}, + {"uses": "actions/setup-node@v4", "with": {"ref": "main"}}, + )) + + assert grant["checkout_refs"] == [ + {"job": "review", "step": "steps[0]", "ref": None, "unresolved_reason": None}, + {"job": "review", "step": "head", "ref": HEAD_SHA, "unresolved_reason": None}, + {"job": "review", "step": "steps[2]", "ref": None, "unresolved_reason": None}, + {"job": "review", "step": "steps[3]", "ref": None, "unresolved_reason": "not_a_string"}, + ] + + +@pytest.mark.parametrize( + ("ref", "pull_request_code"), + [ + (HEAD_SHA, True), + ("${{github.event.pull_request.head.ref}}", True), + ("${{ github.event.pull_request.merge_commit_sha }}", True), + ("${{ github.head_ref }}", True), + ("${{ github.event.workflow_run.head_sha }}", True), + ("refs/pull/${{ github.event.issue.number }}/head", True), + ("refs/pull/123/merge", True), + (None, False), + ("main", False), + ("${{ github.sha }}", False), + ("${{ github.event.pull_request.base.sha }}", False), + ("${{ steps.pr.outputs.sha }}", False), + ], +) +def test_pull_request_code_is_the_documented_head_refs_only(ref, pull_request_code): + from agents_shipgate.core.host_grants import pull_request_code_ref + + assert pull_request_code_ref(ref) is pull_request_code + + +def test_a_workflow_without_agents_or_checkouts_keeps_its_v0_6_shape(): + grant = _grant(_workflow({"run": "make test"}, {"uses": "actions/setup-python@v5"})) + assert "agent_launches" not in grant and "checkout_refs" not in grant and "unread_agent_runs" not in grant + + +def test_job_secrets_name_what_the_agent_job_and_the_workflow_env_reference(): + workflows = _workflow( + jobs={ + "review": {"env": {"GH": "${{ secrets.REVIEW_TOKEN }}"}, "steps": [ + {"run": "echo ${{ secrets.DEPLOY_KEY }}"}, + _agent(anthropic_api_key="${{ secrets.ANTHROPIC_API_KEY }}"), + ]}, + "other": {"steps": [{"run": "echo ${{ secrets.OTHER_JOB_ONLY }}"}]}, + }, + env={"SHARED": "${{ secrets.WORKFLOW_ENV }}"}, + ) + launch, = _launches(workflows) + assert launch["job_secrets"] == ["ANTHROPIC_API_KEY", "DEPLOY_KEY", "REVIEW_TOKEN", "WORKFLOW_ENV"] + + +def _job_env_workflow(env_lines: list[str]) -> str: + return "\n".join([ + "on: pull_request", + "permissions: {contents: read}", + "jobs:", + " review:", + " runs-on: ubuntu-latest", + " env:", + *(f" {line}" for line in env_lines), + " steps:", + f" - run: claude -p {BYPASS} Review", + ]) + "\n" + + +def test_job_secrets_read_a_yaml_alias_that_holds_itself_or_fans_out_once(): + """#823 review cycle 5 (P3): a self-referential alias raised RecursionError, a fan-out took minutes.""" + + holds_itself = yaml.safe_load(_job_env_workflow(["&env", "A: ${{ secrets.TOKEN }}", "B: *env"])) + launch, = _launches(holds_itself) + assert launch["job_secrets"] == ["TOKEN"] + + # Ten references at each of eight levels: 10**8 leaves if each is walked. + levels = ["l0: &l0 '${{ secrets.DEEP }}'"] + [ + f"l{level}: &l{level} [{', '.join([f'*l{level - 1}'] * 10)}]" for level in range(1, 9) + ] + started = time.monotonic() + launch, = _launches(yaml.safe_load(_job_env_workflow(levels))) + assert launch["job_secrets"] == ["DEEP"] + assert time.monotonic() - started < 5 + + +def test_job_secrets_read_unterminated_expressions_in_linear_time(): + """#823 review cycle 7 (C7-F2): every unterminated `${{` scanned to the end, 92 s at 300 KB. + + An unterminated expression names no secret, as before; a closed one + after it still does. + """ + + unterminated = "${{ secrets.NEVER " * 20_000 # about 360 KB, no closing `}}` + job = {"env": {"X": unterminated, "Y": "${{ secrets.CLOSED }}"}, "steps": [{"run": "claude -p Review"}]} + started = time.monotonic() + launch, = _launches(_workflow(jobs={"review": job})) + assert launch["job_secrets"] == ["CLOSED"] + # The same text in an agent action's settings input, which took as long. + row, = _rows(_workflow(_agent()), _workflow(_agent(settings="${{" * 100_000))) + assert (row.direction, row.expands) == ("changed", False) + assert time.monotonic() - started < 5 + # The body of an expression that closes after an unterminated one is still read. + launch, = _launches(_workflow(jobs={"review": { + "env": {"X": "${{ vars.A ${{ secrets.INNER }}"}, "steps": [{"run": "claude -p Review"}], + }})) + assert launch["job_secrets"] == ["INNER"] + + +# --- the only forms read: plain lists of words (#823 review cycle 4) -------------------- +# +# Four review cycles each found a shell form the `run:` reader mis-read, so no +# shell is parsed. A `run:` is read only as one line of plain words — letters, +# digits and `_ . / : = , % + -` — whose program is a known agent CLI; +# `claude_args` and `codex-args` only as such words (parentheses too) across +# blanks and newlines, with no `--settings` or `--mcp-config` flag. + + +@pytest.mark.parametrize( + ("value", "words"), + [ + ("--max-turns 5\n--allowedTools Read", ["--max-turns", "5", "--allowedTools", "Read"]), + (" --allowedTools\tBash(git:status),Read ", ["--allowedTools", "Bash(git:status),Read"]), + ("--permission-mode=bypassPermissions --add-dir ../docs", ["--permission-mode=bypassPermissions", "--add-dir", "../docs"]), + ("-c model=o3 -csandbox_mode=read-only --json", ["-c", "model=o3", "-csandbox_mode=read-only", "--json"]), + ], + ids=["lines", "blanks-and-parentheses", "attached-values", "codex-config"], +) +def test_a_plain_argument_input_is_its_words(value, words): + assert _argument_words(value) == words + + +@pytest.mark.parametrize( + "value", + [ + '--allowedTools "Read"', + "--allowedTools 'Read'", + "--model ${{ vars.CLAUDE_MODEL }}", + "--append-system-prompt $PROMPT", + "--append-system-prompt `cat prompt.md`", + "--append-system-prompt $(cat prompt.md)", + "# reviewer: alice\n--max-turns 5", + "--max-turns 5 # --dangerously-skip-permissions", + "--max-turns 5 notes#x", + "a\\ b", + "--allowedTools Read;Edit", + "--allowedTools Read|Edit", + "--allowedTools Read&Edit", + "--x out.md\n" + "# We'll post the result below\ngh pr comment \"$PR\" --body-file out.md", + "# it's gated\nif [ -n \"$X\" ]; then claude -p --dangerously-skip-permissions go; fi", + "# we don't pipe secrets\ngit diff | claude -p 'Review'", + "npm ci && claude -p --dangerously-skip-permissions go # it's fine", + "# can't\nset -e; codex exec --yolo 'review'", + "# To reproduce locally: npm ci; claude -p \"review this change\"\nnpm test", + "cat < review.md", + "claude -p go 2>&1", + "claude -p \\\n --dangerously-skip-permissions go", + "claude -p $CLAUDE_FLAGS go", + "claude -p \"$(cat prompt.md)\"", + "claude -p 'go'", + 'claude -p "Fix ${{ github.event.issue.title }}"', + 'gh pr comment "$PR" --body "$(claude -p --dangerously-skip-permissions \'go\')"', + "REVIEW=`claude -p --dangerously-skip-permissions go`", + "cat <<'EOF'\nReproduce locally with `claude -p \"review this change\"`.\nEOF", + "cat < comment.md <<'EOF'\nReproduce locally with `claude -p \"review this change\"`.\nEOF\n", + "npm install -g @anthropic-ai/claude-code", + ): + step = {"run": unread} + before = _workflow(step, permissions={"contents": "read", "pull-requests": "write"}) + after = _workflow( + step, {"run": "claude -p --dangerously-skip-permissions Review"}, + permissions={"contents": "read", "pull-requests": "write"}, + ) + assert host_grant_expansion_signals(_changes(before, after)) == [f"workflow_agent_widened_changed: {SOURCE}"] + row, = _rows(before, after) + assert (row.direction, row.expands) == ("widened", True) + assert "an agent launch now skips permission checks (bypassPermissions) (review/steps[1])" in row.why + assert "not counted as a widening" not in row.why.split("; an agent runs at")[0] + assert "review/steps[0]" not in row.why + row.before + row.after + + +def test_an_unread_step_rewritten_as_a_read_launch_does_not_claim_the_rule_it_may_already_have_met(): + before = _workflow({"run": 'npm ci && claude -p --dangerously-skip-permissions "Review"'}) + after = _workflow({"run": "npm ci"}, {"run": "claude -p --dangerously-skip-permissions Review"}) + + assert host_grant_expansion_signals(_changes(before, after)) == [] + row, = _rows(before, after) + assert (row.direction, row.expands) == ("changed", False) + assert ( + "an agent launch now skips permission checks (bypassPermissions) (review/steps[1]), which is not " + "counted as a widening: a step in this job that may launch the agent in a form this audit does not " + "read is gone (review/steps[0]), and this launch may be that step rewritten in a form this audit " + "reads, which may already have done the same" + ) in row.why + + +def test_an_unread_step_that_goes_while_a_read_launch_gains_a_rule_still_widens(): + """The read launch was read on both sides, so the gain is its own.""" + + before = _workflow({"run": "npm ci && claude -p go"}, _agent()) + after = _workflow(_agent("--dangerously-skip-permissions")) + + assert host_grant_expansion_signals(_changes(before, after)) == [f"workflow_agent_widened_changed: {SOURCE}"] + + +@pytest.mark.parametrize( + ("before", "after"), + [ + ("claude -p --allowedTools Read Review", + "npx @anthropic-ai/claude-code -p --dangerously-skip-permissions Review"), + ("claude -p --allowedTools Read Review", 'claude -p --dangerously-skip-permissions "Review"'), + ("codex exec -s workspace-write review", "codex -c sandbox_mode=danger-full-access exec review"), + ], + ids=["npx", "quoted", "codex-root-options"], +) +def test_a_read_launch_that_becomes_a_form_this_audit_does_not_read_is_not_called_gone(before, after): + """#823 review: what was established is that no launch this audit reads is declared.""" + + row, = _rows(_workflow({"run": before}), _workflow({"run": after})) + assert (row.direction, row.expands) == ("changed", False) + assert "a step no longer declares an agent launch this audit reads (review/steps[0])" in row.why + assert ( + "a step that no longer declares one may still start an agent in a way this audit does not read, " + "such as an action outside its table, a script, or a `run:` this audit does not read as a launch " + "(more than one command, quoting, an expansion, `npx`, `codex` options before `exec`), so this row " + "does not say that it no longer starts one" + ) in row.why + assert "no longer launches an agent" not in row.why + assert "dangerously" not in row.after and "danger-full-access" not in row.after + + +# --- an agent action's argument input (#823 review cycles 1 and 4) ------------------------ +# +# `claude_args` and `codex-args` are not shell text: each action splits its own +# input. A plain list of words is split alike by all of them, so only that is +# read; a widening rule is read from those words. + + +@pytest.mark.parametrize( + ("before", "after"), + [ + ("--max-turns 5\n--allowedTools Read", "--max-turns 5\n--dangerously-skip-permissions"), + ("--allowedTools Bash(git:status)", "--allowedTools Bash(git:status) --dangerously-skip-permissions"), + ("--max-turns 5", "--max-turns 5\n--permission-mode\nbypassPermissions"), + ("--max-turns 5", "--permission-mode=bypassPermissions"), + # A word starting with `--` is always a flag to the action, never a value. + ("--allowedTools Read", "--model --dangerously-skip-permissions"), + ], + ids=["several-lines", "parentheses", "mode-over-lines", "mode-attached", "never-a-value"], +) +def test_plain_claude_args_gaining_a_bypass_is_a_widening(before, after): + changes = _changes(_workflow(_agent(before)), _workflow(_agent(after))) + assert host_grant_expansion_signals(changes) == [f"workflow_agent_widened_changed: {SOURCE}"] + row, = _rows(_workflow(_agent(before)), _workflow(_agent(after))) + + assert (row.direction, row.expands) == ("widened", True) + assert "an agent launch now skips permission checks (bypassPermissions) (review/steps[0])" in row.why + assert "not counted as a widening" not in row.why + + +@pytest.mark.parametrize( + "after", + [ + "# --dangerously-skip-permissions\n--max-turns 5", + "--max-turns 5 # --dangerously-skip-permissions", + '--dangerously-skip-permissions --append-system-prompt "Review"', + "--dangerously-skip-permissions --model ${{ vars.CLAUDE_MODEL }}", + "--dangerously-skip-permissions --settings ./ci/settings.json", + "--dangerously-skip-permissions --mcp-config '{\"mcpServers\":{}}'", + "--dangerously-skip-permissions --allowedTools Bash(*)", + ], + ids=["comment-line", "inline-comment", "quoted-prompt", "expression", "settings", "mcp-config", "glob"], +) +def test_argument_input_this_audit_does_not_read_meets_no_rule_and_is_a_changed_row(after): + before = _workflow(_agent("--max-turns 5")) + changed = _workflow(_agent(after)) + launch, = _launches(changed) + + assert launch["settings"] == [{"name": "claude_args", "value": _digest(after), "unresolved_reason": "unread_arguments"}] + assert "widening_rules" not in launch + assert host_grant_expansion_signals(_changes(before, changed)) == [] + row, = _rows(before, changed) + assert (row.direction, row.expands) == ("changed", False) + assert "dangerously" not in row.after and "an agent launch now" not in row.why + # Editing it is still a change, compared by its digest. + edited, = _rows(changed, _workflow(_agent(after + "\n--max-turns 9"))) + assert (edited.direction, edited.expands) == ("changed", False) + assert f"claude_args (not read; digest {_digest(after)})" in edited.before + + +@pytest.mark.parametrize( + ("codex_args", "rule"), + [ + ("--json\n--dangerously-bypass-approvals-and-sandbox", "bypasses approvals and the sandbox"), + ("--full-auto --yolo", "bypasses approvals and the sandbox"), + ("--json\n--sandbox=danger-full-access", "runs without a sandbox (danger-full-access)"), + ("-s danger-full-access", "runs without a sandbox (danger-full-access)"), + ("-sdanger-full-access", "runs without a sandbox (danger-full-access)"), + ], + ids=["several-lines", "yolo", "attached-sandbox", "short-sandbox", "attached-short-sandbox"], +) +def test_plain_codex_args_gaining_a_rule_is_a_widening(codex_args, rule): + before = _workflow({"uses": "openai/codex-action@v1", "with": {"codex-args": "--json"}}) + after = _workflow({"uses": "openai/codex-action@v1", "with": {"codex-args": codex_args}}) + + row, = _rows(before, after) + assert (row.direction, row.expands) == ("widened", True) + assert rule in row.why + + +@pytest.mark.parametrize( + "codex_args", + [ + '["--json", "--yolo"]', + "-s 'danger-full-access'", + "--yolo ${{ vars.EXTRA }}", + # The action appends its own --sandbox, or its own default_permissions + # override for a permission-profile, after codex-args, and either + # takes precedence over a sandbox --config override written before it. + "-c sandbox_mode=danger-full-access", + "--config=default_permissions=:danger-full-access", + ], + ids=["json-array", "quoted-sandbox", "expression", "sandbox-mode-override", "profile-override"], +) +def test_codex_args_this_audit_does_not_read_or_that_selects_nothing_is_changed(codex_args): + before = _workflow({"uses": "openai/codex-action@v1", "with": {"codex-args": "--json"}}) + after = _workflow({"uses": "openai/codex-action@v1", "with": {"codex-args": codex_args}}) + + assert host_grant_expansion_signals(_changes(before, after)) == [] + row, = _rows(before, after) + assert (row.direction, row.expands) == ("changed", False) + + +def test_the_multi_line_widening_reaches_diff_and_the_review_summary(tmp_path): + repo = _repo(tmp_path, {SOURCE: _yaml(_workflow(_agent("--max-turns 5\n--allowedTools Read")))}) + _git(repo, "checkout", "-qb", "change") + _write(repo, {SOURCE: _yaml(_workflow(_agent("--max-turns 5\n--dangerously-skip-permissions")))}) + _git(repo, "commit", "-qam", "bypass on its own line") + + payload = _diff(repo) + row, = payload["rows"] + assert (row["direction"], row["expands"]) == ("widened", True) + assert payload["review"]["summary"]["widenings"] == 1 + + +# --- what is compared -------------------------------------------------------------- + + +@pytest.mark.parametrize( + "after", + [ + # the prompt and an undocumented flag are not compared + _workflow({"run": "claude -p a-different-prompt --model opus --allowedTools Read"}), + # renamed, respelled and reordered flags + _workflow({"name": "Renamed", "run": "claude --allowed-tools=Read --print review"}), + ], + ids=["prompt-and-model", "rename-and-reorder"], +) +def test_an_edit_outside_the_compared_flags_is_quiet(after): + before = _workflow({"run": "claude -p review --allowedTools Read"}) + assert _rows(before, after) == [] + + +def test_a_word_after_a_variadic_flag_is_read_as_its_value_as_the_cli_reads_it(): + before = _workflow({"run": "claude -p review --allowedTools Read"}) + after = _workflow({"run": "claude -p --allowedTools Read review"}) + + row, = _rows(before, after) + assert "--allowedTools Read review" in row.after and "--allowedTools Read," in row.before + "," + + +def test_plain_claude_args_are_compared_by_their_words(): + """Reformatting the list is quiet; a new word is a change.""" + + assert _rows(_workflow(_agent("--max-turns 5 --allowedTools Read")), _workflow(_agent("--max-turns 5\n --allowedTools Read"))) == [] + row, = _rows(_workflow(_agent("--allowedTools Read")), _workflow(_agent("--allowedTools Read,Edit"))) + assert (row.direction, row.expands) == ("changed", False) + assert "claude_args: --allowedTools Read,Edit" in row.after + + +def test_an_action_ref_bump_is_a_step_reference_change_only(): + row, = _rows(_workflow(_agent()), _workflow({**_agent(), "uses": "anthropics/claude-code-action@v2"})) + + assert "action reference changed (review/steps[0])" in row.why + assert "agent launch" not in row.why.split("; an agent runs at")[0] + assert "claude_args" not in row.before + row.after + + +def test_renaming_or_moving_an_agent_step_within_its_job_is_quiet(): + before = _workflow({"run": "make"}, _agent()) + after = _workflow({**_agent(), "name": "Claude review"}, {"run": "make"}) + assert _rows(before, after) == [] + + +# --- direction --------------------------------------------------------------------- + + +@pytest.mark.parametrize( + ("before", "after", "rule"), + [ + (_workflow(_agent()), _workflow(_agent("--dangerously-skip-permissions")), "skips permission checks"), + (_workflow({"run": "claude -p x"}), _workflow({"run": "claude -p --permission-mode=bypassPermissions x"}), + "skips permission checks"), + (_workflow(_agent(allowed_non_write_users="octocat")), _workflow(_agent(allowed_non_write_users="octocat, *")), + "accepts runs triggered by any user (allowed_non_write_users: *)"), + # The bot gate admits any bot, not any user (#823 review cycle 2). + (_workflow(_agent()), _workflow(_agent(allowed_bots="dependabot,*")), + "accepts runs triggered by any bot (allowed_bots: *)"), + (_workflow({"uses": "openai/codex-action@v1", "with": {"sandbox": "read-only"}}), + _workflow({"uses": "openai/codex-action@v1", "with": {"sandbox": "danger-full-access"}}), + "runs without a sandbox (danger-full-access)"), + (_workflow({"uses": "openai/codex-action@v1"}), + _workflow({"uses": "openai/codex-action@v1", "with": {"safety-strategy": "unsafe"}}), + "runs without privilege restrictions"), + (_workflow({"uses": "openai/codex-action@v1"}), + _workflow({"uses": "openai/codex-action@v1", "with": {"codex-args": "--yolo"}}), + "bypasses approvals and the sandbox"), + (_workflow({"uses": "openai/codex-action@v1", "with": {"allow-users": "a"}}), + _workflow({"uses": "openai/codex-action@v1", "with": {"allow-users": "*"}}), + "accepts runs triggered by any user (allow-users: *)"), + (_workflow({"run": "codex exec x"}), _workflow({"run": "codex exec --sandbox danger-full-access x"}), + "runs without a sandbox"), + # A gate entry an expression cannot remove still opens it (#823 review F4). + (_workflow(_agent()), _workflow(_agent(allowed_non_write_users="${{ vars.USERS }}, *")), + "accepts runs triggered by any user (allowed_non_write_users: *)"), + # #823 review C2-F2: the documented bypasses written through inputs the reader lists. + (_workflow(_agent(settings=json.dumps({"permissions": {"defaultMode": "default"}}))), + _workflow(_agent(settings=json.dumps({"permissions": {"defaultMode": "bypassPermissions"}}))), + "skips permission checks (bypassPermissions)"), + (_workflow(_agent()), _workflow(_agent(settings=json.dumps({"defaultMode": "bypassPermissions"}))), + "skips permission checks (bypassPermissions)"), + (_workflow({"uses": "openai/codex-action@v1", "with": {"permission-profile": ":workspace"}}), + _workflow({"uses": "openai/codex-action@v1", "with": {"permission-profile": ":danger-full-access"}}), + "runs without a sandbox (danger-full-access)"), + # the last of a repeated `--permission-mode` counts (#823 review cycle 5) + (_workflow(_agent("--permission-mode bypassPermissions --permission-mode default")), + _workflow(_agent("--permission-mode default --permission-mode bypassPermissions")), + "skips permission checks (bypassPermissions)"), + (_workflow({"run": "claude -p --permission-mode bypassPermissions --permission-mode default x"}), + _workflow({"run": "claude -p --permission-mode default --permission-mode=bypassPermissions x"}), + "skips permission checks (bypassPermissions)"), + ], + ids=["skip-flag", "mode-flag", "gate", "bots", "codex-sandbox", "codex-unsafe", "codex-args", "codex-users", + "codex-cli", "gate-entry-beside-an-expression", "settings-default-mode", + "settings-top-level-default-mode", "codex-permission-profile", "last-mode-args", "last-mode-cli"], +) +def test_a_documented_rule_gained_is_a_widening(before, after, rule): + changes = _changes(before, after) + assert host_grant_expansion_signals(changes) == [f"workflow_agent_widened_changed: {SOURCE}"] + row, = _rows(before, after) + assert (row.direction, row.expands) == ("widened", True) + assert rule in row.why + + +@pytest.mark.parametrize( + ("before", "after"), + [ + # one rule, two spellings + (_workflow(_agent("--dangerously-skip-permissions")), _workflow(_agent("--permission-mode bypassPermissions"))), + # the same rule moved from the CLI to the action + (_workflow({"run": "claude -p --dangerously-skip-permissions x"}), + _workflow(_agent("--dangerously-skip-permissions"))), + # narrowed + (_workflow(_agent("--dangerously-skip-permissions")), _workflow(_agent("--allowedTools Read"))), + (_workflow(_agent(allowed_non_write_users="*")), _workflow(_agent(allowed_non_write_users="octocat"))), + # text an expression can reach is never read for a rule + (_workflow(_agent()), _workflow(_agent("${{ inputs.extra }} --dangerously-skip-permissions"))), + (_workflow(_agent()), _workflow(_agent("--dangerously-skip-permissions ${{ inputs.extra }}"))), + (_workflow(_agent()), _workflow(_agent(allowed_non_write_users="*${{ vars.USERS }}"))), + (_workflow({"uses": "openai/codex-action@v1"}), + _workflow({"uses": "openai/codex-action@v1", "with": {"sandbox": "${{ vars.SANDBOX }}"}})), + # widened by a tool rule, which is #824's to rate + (_workflow(_agent("--allowedTools Read")), _workflow(_agent("--allowedTools Bash"))), + (_workflow({"run": "claude -p --permission-mode default x"}), + _workflow({"run": "claude -p --permission-mode acceptEdits x"})), + # one rule, written as a flag and as the settings it passes + (_workflow(_agent("--dangerously-skip-permissions")), + _workflow(_agent(settings=json.dumps({"permissions": {"defaultMode": "bypassPermissions"}})))), + # settings holding an expression, or naming a file, meet no rule + (_workflow(_agent()), + _workflow(_agent(settings='{"permissions":{"defaultMode":"${{ vars.MODE }}"}}'))), + (_workflow(_agent()), _workflow(_agent(settings=".github/claude-settings.json"))), + (_workflow(_agent(settings=json.dumps({"permissions": {"defaultMode": "default"}}))), + _workflow(_agent(settings=json.dumps({"permissions": {"defaultMode": "acceptEdits"}})))), + (_workflow({"uses": "openai/codex-action@v1", "with": {"permission-profile": ":read-only"}}), + _workflow({"uses": "openai/codex-action@v1", "with": {"permission-profile": ":workspace"}})), + # a `--permission-mode` a later one replaces meets no rule (#823 review cycle 5) + (_workflow(_agent("--permission-mode default")), + _workflow(_agent("--permission-mode bypassPermissions --permission-mode default"))), + (_workflow({"run": "claude -p --permission-mode default x"}), + _workflow({"run": "claude -p --permission-mode=bypassPermissions --permission-mode default x"})), + ], + ids=["respelled", "moved-to-action", "narrowed", "gate-closed", "after-an-expression", "before-an-expression", + "gate-entry-holding-an-expression", "mode-expression", "tool-rule", "accept-edits", "flag-to-settings", + "settings-expression", "settings-path", "settings-accept-edits", "codex-workspace-profile", + "replaced-mode-args", "replaced-mode-cli"], +) +def test_any_other_edit_is_changed(before, after): + assert host_grant_expansion_signals(_changes(before, after)) == [] + row, = _rows(before, after) + assert (row.direction, row.expands) == ("changed", False) + + +# --- codex exec's full-access sandbox, however the CLI reads it (#823 review cycle 3) --- + +CODEX_WORKSPACE = _workflow({"run": "codex exec -s workspace-write review"}) + + +@pytest.mark.parametrize( + "run", + [ + "codex exec -s danger-full-access review", + "codex exec --sandbox=danger-full-access review", + # clap reads a short option's attached value, with or without `=`. + "codex exec -sdanger-full-access review", + "codex exec -s=danger-full-access review", + # `sandbox_mode` is the setting --sandbox sets, in each way -c is written. + "codex exec -c sandbox_mode=danger-full-access review", + "codex exec --config=sandbox_mode=danger-full-access review", + "codex exec -csandbox_mode=danger-full-access review", + "codex exec -c=sandbox_mode=danger-full-access review", + # The built-in full-access profile, which is what the action's + # `permission-profile: :danger-full-access` passes the CLI. + "codex exec -c default_permissions=:danger-full-access review", + "codex exec --config default_permissions=:danger-full-access review", + # The last override of a key counts; a profile override outranks a sandbox one. + "codex exec -c sandbox_mode=read-only -c sandbox_mode=danger-full-access review", + "codex exec -c sandbox_mode=read-only -c default_permissions=:danger-full-access review", + ], + ids=["short", "long-attached", "short-attached", "short-equals", "config", "config-attached", + "c-attached", "c-equals", "profile", "profile-long", "last-override", "profile-over-sandbox-mode"], +) +def test_codex_exec_full_access_widens_in_every_spelling_the_cli_reads(run): + after = _workflow({"run": run}) + + assert host_grant_expansion_signals(_changes(CODEX_WORKSPACE, after)) == [ + f"workflow_agent_widened_changed: {SOURCE}" + ] + row, = _rows(CODEX_WORKSPACE, after) + assert (row.direction, row.expands) == ("widened", True) + assert "an agent launch now runs without a sandbox (danger-full-access) (review/steps[0])" in row.why + assert "no permission flags" not in row.after + + +@pytest.mark.parametrize( + "run", + [ + # --sandbox takes precedence over a --config override. + "codex exec -s workspace-write -c sandbox_mode=danger-full-access review", + "codex exec -sworkspace-write -c default_permissions=:danger-full-access review", + # The last override counts, and a profile override outranks sandbox_mode. + "codex exec -c sandbox_mode=danger-full-access -c sandbox_mode=read-only review", + "codex exec -c sandbox_mode=danger-full-access -c default_permissions=:workspace review", + # A key under another table, or another value, is not the setting. + "codex exec -c profiles.ci.sandbox_mode=danger-full-access review", + "codex exec -c default_permissions=danger-full-access review", + "codex exec -sread-only review", + ], + ids=["sandbox-flag-wins", "attached-sandbox-flag-wins", "last-override", "profile-over-sandbox-mode", + "profile-scoped-key", "custom-profile-name", "read-only"], +) +def test_a_codex_exec_sandbox_the_cli_does_not_select_is_changed(run): + after = _workflow({"run": run}) + + assert host_grant_expansion_signals(_changes(CODEX_WORKSPACE, after)) == [] + row, = _rows(CODEX_WORKSPACE, after) + assert (row.direction, row.expands) == ("changed", False) + + +@pytest.mark.parametrize( + "run", + [ + "codex exec -c sandbox_mode=" + "1" * 5000 + " Review", + "codex exec -c default_permissions=" + "1" * 4400 + " Review", + "codex e --config=sandbox_mode=" + "1" * 4301 + " Review", + ], + ids=["sandbox-mode", "default-permissions", "attached-config"], +) +def test_a_codex_config_integer_past_the_digit_limit_is_read_as_text_and_selects_nothing(run): + """#823 review cycle 6 (C6-F2): `tomllib` raises a plain `ValueError` for such an integer.""" + + before, after = _workflow({"run": "echo hi"}), _workflow({"run": run}) + + launch, = _launches(after) + assert not launch.get("widening_rules") + assert host_grant_expansion_signals(_changes(before, after)) == [] + row, = _rows(before, after) + assert (row.direction, row.expands) == ("changed", False) + + +def test_attached_short_values_publish_under_the_primary_spelling(): + launch, = _launches(_workflow({"run": "codex exec -sdanger-full-access -c=model=o3 -pci review"})) + + assert launch["settings"] == [ + {"name": "--config", "value": "model=o3", "unresolved_reason": None}, + {"name": "--profile", "value": "ci", "unresolved_reason": None}, + {"name": "--sandbox", "value": "danger-full-access", "unresolved_reason": None}, + ] + assert launch["widening_rules"] == [{"rule": "danger_full_access", "setting": "--sandbox"}] + # One setting, two spellings: respelling it is quiet. + assert _rows( + _workflow({"run": "codex exec -s danger-full-access review"}), + _workflow({"run": "codex exec -sdanger-full-access review"}), + ) == [] + # One rule, two spellings: moving between the flag and the override is not a widening. + before = _workflow({"run": "codex exec -s danger-full-access review"}) + after = _workflow({"run": "codex exec -c sandbox_mode=danger-full-access review"}) + assert host_grant_expansion_signals(_changes(before, after)) == [] + row, = _rows(before, after) + assert (row.direction, row.expands) == ("changed", False) + + +def _jobs(**steps): + return _workflow(jobs={job: {"runs-on": "ubuntu-latest", "steps": list(items)} for job, items in steps.items()}) + + +def _named(args, name="agent"): + return {"name": name, **_agent(args)} + + +BYPASS = "--dangerously-skip-permissions" + + +def test_renaming_a_job_that_launches_a_bypassing_agent_is_not_a_widening(): + """#823 review F3: a rule the launch already met in the job it left is moved, not gained.""" + + before, after = _jobs(review=[_named(BYPASS)]), _jobs(**{"code-review": [_named(BYPASS)]}) + + assert host_grant_expansion_signals(_changes(before, after)) == [] + row, = _rows(before, after) + assert (row.direction, row.expands) == ("changed", False) + assert ( + "an agent launch that skips permission checks (bypassPermissions) moved between jobs " + "(review/agent → code-review/agent), which is not counted as a widening" + ) in row.why + assert "an agent launch now" not in row.why + + +@pytest.mark.parametrize( + ("before", "after"), + [ + # the agent step moved to another job, which launched no agent before + (_jobs(lint=[_named(BYPASS)], review=[{"run": "make"}]), + _jobs(lint=[{"run": "make"}], review=[_named(BYPASS)])), + # renamed and edited in the same change + (_jobs(review=[_named(BYPASS)]), _jobs(**{"code-review": [_named(f"{BYPASS} --max-turns 5")]})), + # the same launch now runs in the other job, and the other job's in this one + (_jobs(lint=[_named(BYPASS)], review=[_named("--allowedTools Read")]), + _jobs(lint=[_named("--allowedTools Read")], review=[_named(BYPASS)])), + # a gate the launch opened, moved with it + (_jobs(review=[{"name": "agent", **_agent(allowed_bots="*")}]), + _jobs(triage=[{"name": "agent", **_agent(allowed_bots="*")}])), + # the job it left keeps a step that does not mention the agent + (_jobs(lint=[{"name": "agent", "run": f"claude -p {BYPASS} Review"}, {"run": "make"}], + review=[{"run": "make"}]), + _jobs(lint=[{"run": "make"}], + review=[{"name": "agent", "run": f"claude -p {BYPASS} Review"}])), + # #823 review cycle 7: a job renamed with the install step it had + # beside the launch; the unread step is in the job gaining the rule + (_jobs(review=[{"run": "npm i -g @anthropic-ai/claude-code"}, {"name": "agent", "run": f"claude -p {BYPASS} Review"}]), + _jobs(**{"code-review": [{"run": "npm i -g @anthropic-ai/claude-code"}, + {"name": "agent", "run": f"claude -p {BYPASS} Review"}]})), + # a job renamed beside another job that keeps the unread step it had + (_jobs(review=[_named(BYPASS)], setup=[{"run": "claude mcp add x"}]), + _jobs(**{"code-review": [_named(BYPASS)]}, setup=[{"run": "claude mcp add x"}])), + ], + ids=["step-moved", "renamed-and-edited", "swapped", "gate-moved", "moved-beside-a-step-that-is-no-launch", + "renamed-with-its-install-step", "renamed-beside-an-unread-step-another-job-keeps"], +) +def test_a_launch_that_left_one_job_for_another_moves_its_rules(before, after): + assert host_grant_expansion_signals(_changes(before, after)) == [] + row, = _rows(before, after) + assert (row.direction, row.expands) == ("changed", False) + assert "moved between jobs" in row.why + + +@pytest.mark.parametrize( + ("before", "after"), + [ + # a second job now bypasses, beside the one that still does + (_jobs(lint=[_named(BYPASS)], review=[_named("--allowedTools Read")]), + _jobs(lint=[_named(BYPASS)], review=[_named(BYPASS)])), + # the job that met the rule still launches the agent, and a different launch meets it elsewhere + (_jobs(lint=[_named(BYPASS)], review=[_named("--allowedTools Read")]), + _jobs(lint=[_named("--allowedTools Read")], review=[_named(f"{BYPASS} --max-turns 5")])), + # the job that met it remains, running its launch in a form this audit + # does not read (#823 review cycle 3), so that launch has not left it + (_jobs(lint=[{"name": "agent", "run": f"claude -p {BYPASS} Review"}], review=[_named("--allowedTools Read")]), + _jobs(lint=[{"name": "agent", "run": f"npx @anthropic-ai/claude-code -p {BYPASS} Review"}], + review=[_named(BYPASS)])), + # the same, with the other job's launch edited in place into the one this job had + (_jobs(lint=[{"name": "agent", "run": f"claude -p {BYPASS} Review"}], + review=[{"name": "agent", "run": "claude -p --allowedTools Read Review"}]), + _jobs(lint=[{"name": "agent", "run": f"npx @anthropic-ai/claude-code -p {BYPASS} Review"}], + review=[{"name": "agent", "run": f"claude -p {BYPASS} Review"}])), + # a step moved and edited while the job it left remains + (_jobs(lint=[_named(BYPASS)], review=[{"run": "make"}]), + _jobs(lint=[{"run": "make"}], review=[_named(f"{BYPASS} --max-turns 5")])), + # #823 review cycle 5 (M5): the job it met it in still runs it, now + # quoted and so unread, while the other job adds the same plain launch + (_jobs(lint=[{"name": "agent", "run": f"claude -p {BYPASS} Review"}], review=[{"run": "echo hi"}]), + _jobs(lint=[{"name": "agent", "run": f'claude -p {BYPASS} "Review"'}], + review=[{"name": "agent", "run": f"claude -p {BYPASS} Review"}])), + # (M6) the same, the job it met it in now running it through `npx` + (_jobs(lint=[{"name": "agent", "run": f"claude -p {BYPASS} Review"}], review=[{"run": "echo hi"}]), + _jobs(lint=[{"name": "agent", "run": f"npx @anthropic-ai/claude-code -p {BYPASS} Review"}], + review=[{"name": "agent", "run": f"claude -p {BYPASS} Review"}])), + # the same, the unread step now standing at another step label + (_jobs(lint=[{"run": f"claude -p {BYPASS} Review"}], review=[{"run": "echo hi"}]), + _jobs(lint=[{"run": "echo hi"}, {"run": f'claude -p {BYPASS} "Review"'}], + review=[{"name": "agent", "run": f"claude -p {BYPASS} Review"}])), + # #823 review cycle 6 (V1): the install step merged into the launch, + # so the job it met it in has as many unread steps as before, at a + # step the launch did not hold + (_jobs(lint=[{"run": "npm i -g @anthropic-ai/claude-code"}, {"run": f"claude -p {BYPASS} Review"}], + review=[{"run": "echo hi"}]), + _jobs(lint=[{"run": f"npm i -g @anthropic-ai/claude-code && claude -p {BYPASS} Review"}], + review=[{"name": "agent", "run": f"claude -p {BYPASS} Review"}])), + # (V2) an unread step removed while the launch becomes unread at another step + (_jobs(lint=[{"run": f"claude -p {BYPASS} Review"}, {"run": 'echo "claude"'}], review=[{"run": "echo hi"}]), + _jobs(lint=[{"run": "echo hi"}, {"run": f'claude -p {BYPASS} "Review"'}], + review=[{"name": "agent", "run": f"claude -p {BYPASS} Review"}])), + # the job it left keeps an unread step it already had, at another + # step: that step carries no text to tell it is not the launch (#823 + # review cycle 6 flips the cycle 5 guard, in the safe direction) + (_jobs(lint=[{"name": "agent", "run": f"claude -p {BYPASS} Review"}, {"name": "mcp", "run": "claude mcp add x"}], + review=[{"run": "make"}]), + _jobs(lint=[{"name": "mcp", "run": "claude mcp add x"}], + review=[{"name": "agent", "run": f"claude -p {BYPASS} Review"}])), + # the job it met it in still runs the action, with an argument input + # it no longer reads, or a settings input holding an expression + (_jobs(lint=[_named(BYPASS)], review=[{"run": "make"}]), + _jobs(lint=[_named(f'{BYPASS} --append-system-prompt "Review"')], review=[_named(BYPASS)])), + (_jobs(lint=[{"name": "agent", **_agent(settings='{"permissions":{"defaultMode":"bypassPermissions"}}')}], + review=[{"run": "make"}]), + _jobs(lint=[{"name": "agent", **_agent(settings='{"permissions":{"defaultMode":"${{ vars.MODE }}"}}')}], + review=[{"name": "agent", **_agent(settings='{"permissions":{"defaultMode":"bypassPermissions"}}')}])), + # #823 review cycle 7 (R6): the job it met it in is renamed and quotes + # the prompt, so under its new name it still runs the launch, unread + (_jobs(lint=[{"name": "agent", "run": f"claude -p {BYPASS} Review"}], review=[{"run": "echo hi"}]), + _jobs(**{"lint-renamed": [{"name": "agent", "run": f'claude -p {BYPASS} "Review"'}]}, + review=[{"name": "agent", "run": f"claude -p {BYPASS} Review"}])), + # (R7) renamed and run through `npx`, the other job spelling the bypass as a mode + (_jobs(lint=[{"name": "agent", "run": f"claude -p {BYPASS} Review"}], review=[{"run": "echo hi"}]), + _jobs(**{"lint-renamed": [{"name": "agent", "run": f"npx @anthropic-ai/claude-code -p {BYPASS} Review"}]}, + review=[{"name": "agent", "run": "claude -p --permission-mode bypassPermissions Review"}])), + # (R3) the job it met it in is removed, and an existing job gains the quoted launch + (_jobs(lint=[{"name": "agent", "run": f"claude -p {BYPASS} Review"}], review=[{"run": "echo hi"}], + docs=[{"run": "make"}]), + _jobs(review=[{"name": "agent", "run": f"claude -p {BYPASS} Review"}], + docs=[{"run": "make"}, {"name": "agent", "run": f'claude -p {BYPASS} "Review"'}])), + # the same while the job it left remains + (_jobs(lint=[{"name": "agent", "run": f"claude -p {BYPASS} Review"}, {"run": "make"}], + review=[{"run": "echo hi"}], docs=[{"run": "make"}]), + _jobs(lint=[{"run": "make"}], review=[{"name": "agent", "run": f"claude -p {BYPASS} Review"}], + docs=[{"run": "make"}, {"name": "agent", "run": f'claude -p {BYPASS} "Review"'}])), + # renamed into an action whose argument input this audit does not read + (_jobs(lint=[_named(BYPASS)], review=[{"run": "echo hi"}]), + _jobs(**{"lint-renamed": [_named(f'{BYPASS} --append-system-prompt "Review"')]}, + review=[_named(BYPASS)])), + # Two jobs renamed at once, one holding an unread step: which new job + # is which is not told, so the gain is claimed (the safe direction). + (_jobs(lint=[_named(BYPASS)], setup=[{"run": "npm i -g @anthropic-ai/claude-code"}]), + _jobs(review=[_named(BYPASS)], prepare=[{"run": "npm i -g @anthropic-ai/claude-code"}])), + ], + ids=["second-job", "narrowed-there-widened-here", "unread-in-the-job-it-met-it", "edited-into-the-same-launch", + "moved-and-edited", "quoted-in-the-job-it-met-it", "npx-in-the-job-it-met-it", + "unread-at-another-step-in-the-job-it-met-it", "install-merged-into-the-launch", + "unread-step-removed-while-the-launch-becomes-unread", "beside-an-unread-step-that-stays", + "arguments-unread-in-the-job-it-met-it", + "setting-expression-in-the-job-it-met-it", + "renamed-and-quoted", "renamed-and-npx", "removed-while-another-job-gains-it-quoted", + "left-for-another-job-quoted", "renamed-into-unread-arguments", "two-jobs-renamed"], +) +def test_a_rule_another_job_gains_while_no_launch_left_is_a_widening(before, after): + assert host_grant_expansion_signals(_changes(before, after)) == [f"workflow_agent_widened_changed: {SOURCE}"] + row, = _rows(before, after) + assert (row.direction, row.expands) == ("widened", True) + assert "an agent launch now skips permission checks (bypassPermissions) (review/agent)" in row.why + assert "the launch already met that rule in the job it left" not in row.why + + +@pytest.mark.parametrize("renamed_first", [True, False], ids=["renamed-declared-first", "renamed-declared-last"]) +def test_a_renamed_job_takes_the_move_whatever_order_the_jobs_are_declared_in(renamed_first): + """#823 review cycle 2 (P3): the rule moved to the job running the same launch, not the first one declared.""" + + renamed = ("code-review", [_named(BYPASS)]) + added = ("triage", [_named(f"{BYPASS} --max-turns 5")]) + before = _jobs(review=[_named(BYPASS)]) + after = _jobs(**dict([renamed, added] if renamed_first else [added, renamed])) + + row, = _rows(before, after) + assert (row.direction, row.expands) == ("widened", True) + assert "moved between jobs (review/agent → code-review/agent)" in row.why + assert "an agent launch now skips permission checks (bypassPermissions) (triage/agent)" in row.why + + +def test_a_setting_holding_an_expression_is_marked_and_the_row_says_what_it_leaves_unread(): + """#823 review F4: the row no longer reads as though no documented rule was gained.""" + + before = _workflow(_agent(allowed_non_write_users="${{ vars.USERS }}")) + after = _workflow(_agent(allowed_non_write_users="${{ vars.OTHER_USERS }}")) + + launch, = _launches(after) + gate = next(item for item in launch["settings"] if item["name"] == "allowed_non_write_users") + assert gate["holds_expression"] is True + assert "widening_rules" not in launch + row, = _rows(before, after) + assert (row.direction, row.expands) == ("changed", False) + assert ( + "an agent launch setting holds a `${{ }}` expression (allowed_non_write_users at review/steps[0]), " + "which GitHub substitutes before the action reads it; documented widening rules are read only from " + "the literal text the expression cannot reach, so this row does not say whether the text it reaches " + "meets one" + ) in row.why + # A setting without one says nothing of the kind, and the key is omitted; + # an argument input holding one is not read at all. + plain, = _launches(_workflow(_agent())) + assert "holds_expression" not in plain["settings"][0] + args, = _launches(_workflow(_agent("--model ${{ vars.M }}"))) + assert args["settings"] == [ + {"name": "claude_args", "value": _digest("--model ${{ vars.M }}"), "unresolved_reason": "unread_arguments"}, + ] + + +@pytest.mark.parametrize( + ("before", "after", "rule", "held"), + [ + (_workflow(_agent(allowed_non_write_users="${{ vars.EXTRA_USERS }}")), + _workflow(_agent(allowed_non_write_users="${{ vars.EXTRA_USERS }}, *")), + "accepts runs triggered by any user (allowed_non_write_users: *)", + "allowed_non_write_users held a `${{ }}` expression, whose substituted text this audit does not read"), + # the settings input is read for the rule only when it holds no expression (#823 review C2-F2) + (_workflow(_agent(settings='{"permissions":{"defaultMode":"${{ vars.MODE }}"}}')), + _workflow(_agent(settings='{"permissions":{"defaultMode":"bypassPermissions"}}')), + "skips permission checks (bypassPermissions)", + "settings held a `${{ }}` expression, whose substituted text this audit does not read"), + # an argument input this audit did not read may have met it (#823 review cycle 4) + (_workflow(_agent("--allowedTools Read --model ${{ vars.CLAUDE_MODEL }}")), + _workflow(_agent("--dangerously-skip-permissions --model opus")), + "skips permission checks (bypassPermissions)", + "claude_args was not a plain list of words this audit reads, so no rule was read from it"), + (_workflow(_agent('--dangerously-skip-permissions --append-system-prompt "Review"')), + _workflow(_agent("--dangerously-skip-permissions")), + "skips permission checks (bypassPermissions)", + "claude_args was not a plain list of words this audit reads, so no rule was read from it"), + ], + ids=["gate", "settings", "claude-args-expression", "claude-args-quoted"], +) +def test_a_rule_gained_where_the_setting_was_not_read_before_is_named_and_not_claimed(before, after, rule, held): + assert host_grant_expansion_signals(_changes(before, after)) == [] + row, = _rows(before, after) + assert (row.direction, row.expands) == ("changed", False) + assert ( + f"an agent launch now {rule} (review/steps[0]), which is not counted as a widening: before, this " + f"job's {held}, and it may already have done the same" + ) in row.why + + +def test_a_new_workflow_that_bypasses_permissions_is_an_added_widening(): + after = _grant(_workflow(_agent("--dangerously-skip-permissions"))) + changes = diff_host_grants({"grants": []}, {"grants": [after]}) + + assert host_grant_expansion_signals(changes) == [f"workflow_agent_widened_added: {SOURCE}"] + row, = capability_diff_rows({"changes": changes, "expansion_signals": host_grant_expansion_signals(changes)}) + assert (row.direction, row.expands) == ("added", True) + assert "an agent launch now skips permission checks (bypassPermissions) (review/steps[0])" in row.why + + +def test_access_and_risk_still_describe_the_token_and_triggers_alone(): + plain, bypass = (_grant(_workflow(_agent(args))) for args in ("", "--dangerously-skip-permissions")) + assert (plain["access"], plain["risk"]) == (bypass["access"], bypass["risk"]) + + +# --- negative controls ------------------------------------------------------------- + + +def test_a_non_agent_actions_inputs_are_not_read(): + before = _workflow({"uses": "someone/ai-review@v1", "with": {"claude_args": "--allowedTools Read"}}) + after = _workflow({"uses": "someone/ai-review@v1", "with": {"claude_args": "--dangerously-skip-permissions"}}) + + assert _launches(after) == [] + assert _rows(before, after) == [] + + +def test_a_run_that_does_not_mention_an_agent_cli_is_nothing(): + for run in ("npm test", "echo claudette", "CLAUDE_MODEL=opus make", "echo my_codex_notes"): + grant = _grant(_workflow({"run": run})) + assert "agent_launches" not in grant and "unread_agent_runs" not in grant + assert uncompared_agent_launch_texts(grant) == [] + + +def test_a_pull_request_workflow_with_a_default_checkout_claims_no_pull_request_code(): + row, = _rows(_reproduction(), _reproduction(pr="write")) + + assert "checkout" not in row.why + assert "untrusted-input" not in row.why + assert row.why.endswith( + "an agent runs at review/steps[1] (anthropics/claude-code-action) beside the write scope pull-requests" + ) + + +def test_a_composite_action_is_neither_a_launch_nor_a_limit(): + step = {"uses": "./.github/actions/claude-review", "with": {"claude_args": "--dangerously-skip-permissions"}} + grant = _grant(_workflow(step)) + assert "agent_launches" not in grant and "unread_agent_runs" not in grant + assert _rows(_workflow({"run": "echo"}), _workflow(step)) == [] + + +def test_a_removed_workflow_gets_no_note(): + before = _grant(_reproduction(trigger="issue_comment", pr="write")) + changes = diff_host_grants({"grants": [before]}, {"grants": []}) + row, = capability_diff_rows({"changes": changes, "expansion_signals": []}) + assert row.direction == "removed" + assert "an agent runs" not in row.why + + +# --- redaction (#802) -------------------------------------------------------------- + + +SETTINGS_JSON = json.dumps({ + "env": {"DB_PASSWORD": "hunter2-canary", "INTERNAL_KEY": "canary-9f8e7d"}, + "apiKeyHelper": "echo canary-helper-value", + "permissions": {"allow": ["Bash(npm test)"]}, +}) +MCP_JSON = json.dumps({"mcpServers": {"db": { + "command": "db-mcp", "env": {"DB_API_TOKEN": "canary-tok-123"}, "headers": {"X-API-Key": "canary-hdr-456"}, +}}}) +#: What the host readers publish for them: key names, every such value withheld. +SETTINGS_PUBLISHED = ( + '{"apiKeyHelper":"","env":{"DB_PASSWORD":"","INTERNAL_KEY":""},' + '"permissions":{"allow":["Bash(npm test)"]}}' +) +MCP_PUBLISHED = ( + '{"mcpServers":{"db":{"command":"db-mcp","env":{"DB_API_TOKEN":""},' + '"headers":{"X-API-Key":""}}}}' +) +JSON_CANARIES = ("hunter2-canary", "canary-9f8e7d", "canary-helper-value", "canary-tok-123", "canary-hdr-456") + +#: #823 review C2-F1: text the host readers never publish, in no secret-named +#: key: an `mcp-remote` bearer header among a server's `args`, and a hook's command. +REMOTE_MCP_JSON = json.dumps({"mcpServers": {"remote": {"command": "npx", "args": [ + "mcp-remote", "https://mcp.example.com/sse", "--header", "Authorization: Bearer tokCANARY0123456789abcdef", +]}}}, separators=(",", ":")) +HOOK_JSON = json.dumps({"hooks": {"Stop": [{"hooks": [{ + "type": "command", "command": 'curl -H "X-Auth-Token: hookCANARY77" https://hooks.example.com/notify', +}]}]}}, separators=(",", ":")) +REMOTE_CODEX_CONFIG = ( + 'mcp_servers.remote={command="npx", args=["mcp-remote", "https://mcp.example.com/sse", ' + '"--header", "Authorization: Bearer tokCANARY-codex-0123456789"]}' +) +SHAPE_CANARIES = ( + "tokCANARY0123456789abcdef", "tokCANARY-codex-0123456789", "hookCANARY77", "hooks.example.com", + "mcp.example.com", "mcp-remote", "X-Auth-Token", +) + + +def test_a_json_value_publishes_its_shape_and_none_of_its_free_text(): + """#823 review C2-F1: every string a host reader does not publish is withheld, and still compared. + + The same server in `.mcp.json` publishes `remote (command name npx)`, and + the same hook in `.claude/settings.json` publishes `Stop`; neither + publishes an argument or a command. The same JSON passed through + `claude_args` or a `run:` is not read at all (#823 review cycle 4). + """ + + args = f"--allowedTools Read --mcp-config '{REMOTE_MCP_JSON}'" + grant = _grant(_workflow( + _agent(args, settings=HOOK_JSON, mcp_config=REMOTE_MCP_JSON), + {"run": f"codex exec -c '{REMOTE_CODEX_CONFIG}' 'go'"}, + )) + action, = grant["agent_launches"] + server_args = ["mcp-remote", "https://mcp.example.com/sse", "--header", "Authorization: Bearer tokCANARY0123456789abcdef"] + server = json.dumps( + {"mcpServers": {"remote": {"args": [_digest(arg) for arg in server_args], "command": "npx"}}}, + separators=(",", ":"), + ) + hook = ( + '{"hooks":{"Stop":[{"hooks":[{"command":"' + + _digest('curl -H "X-Auth-Token: hookCANARY77" https://hooks.example.com/notify') + + '","type":"' + _digest("command") + '"}]}]}}' + ) + assert {item["name"]: item["value"] for item in action["settings"]} == { + "claude_args": _digest(args), + "mcp_config": server, + "settings": hook, + } + assert grant["unread_agent_runs"] == [{"job": "review", "step": "steps[1]", "agent": "codex"}] + text = json.dumps(grant) + for canary in SHAPE_CANARIES: + assert canary not in text + limits = uncompared_agent_launch_texts(grant) + assert [limit.split(";")[0] for limit in limits] == [ + "the claude_args value of the agent launch at review/steps[0] (anthropics/claude-code-action) is not a " + "plain list of words this audit reads: it holds a quote, a `${{ }}` expression, `$`, a backtick, a " + "comment, a shell operator, JSON, `--settings` or `--mcp-config`, or another character outside the " + "plain set", + "the `run:` at review/steps[1] mentions codex and is not read as an agent launch: only a " + "single-line command of plain words run by bash or sh, whose program is `claude -p` or `codex exec`, " + "is read, so this step may start an agent that is neither published nor compared, and adding, " + "removing or editing it gives no row", + ] + + # A withheld string is still compared: a new argument or command is a row. + edited = HOOK_JSON.replace("curl -H", "wget --header") + row, = _rows(_workflow(_agent(settings=HOOK_JSON)), _workflow(_agent(settings=edited))) + assert (row.direction, row.expands) == ("changed", False) + assert row.before != row.after and "wget" not in row.after + # A documented setting's value, a permission rule and an enabled server are + # published as the settings reader publishes them; other strings are not. + launch, = _launches(_workflow(_agent(settings=json.dumps({ + "permissions": {"defaultMode": "acceptEdits", "allow": ["Bash(npm test)"], "additionalDirectories": ["../x"]}, + "enabledMcpjsonServers": ["github"], "model": "claude-opus", + })))) + setting = next(item for item in launch["settings"] if item["name"] == "settings") + assert setting["value"] == ( + '{"enabledMcpjsonServers":["github"],"model":"' + _digest("claude-opus") + '",' + '"permissions":{"additionalDirectories":["' + _digest("../x") + '"],"allow":["Bash(npm test)"],' + '"defaultMode":"acceptEdits"}}' + ) + + +def test_an_mcp_server_url_publishes_its_scheme_and_host_and_compares_its_query_as_the_mcp_reader_does(): + def config(url): + return _workflow(_agent(mcp_config=json.dumps({"mcpServers": {"db": {"url": url}}}))) + + launch, = _launches(config("https://mcp.example.com/v1/sse")) + assert { + "name": "mcp_config", "value": '{"mcpServers":{"db":{"url":"https://mcp.example.com/"}}}', + "unresolved_reason": None, + } in launch["settings"] + # A query decides which tools a server exposes, so it is compared (#723); a path is not. + row, = _rows(config("https://mcp.example.com/sse?read_only=true"), config("https://mcp.example.com/sse")) + assert "read_only" not in row.before + row.after + assert _rows(config("https://mcp.example.com/a"), config("https://mcp.example.com/b")) == [] + + +def test_a_json_value_publishes_only_what_the_host_readers_publish(): + grant = _grant(_workflow( + _agent(f"--mcp-config '{MCP_JSON}' --allowedTools Read", settings=SETTINGS_JSON, mcp_config=MCP_JSON), + {"run": f"claude -p --mcp-config '{MCP_JSON}' --settings '{SETTINGS_JSON}' 'go'"}, + )) + action, = grant["agent_launches"] + + assert {item["name"]: item["value"] for item in action["settings"]} == { + "claude_args": _digest(f"--mcp-config '{MCP_JSON}' --allowedTools Read"), + "mcp_config": MCP_PUBLISHED, + "settings": SETTINGS_PUBLISHED, + } + assert grant["unread_agent_runs"] == [{"job": "review", "step": "steps[1]", "agent": "claude"}] + text = json.dumps(grant) + for canary in JSON_CANARIES: + assert canary not in text + assert hashlib.sha256(canary.encode()).hexdigest() not in text + assert _uncompared_workflow_text(grant) is None + + +@pytest.mark.parametrize( + "value", + [ + f"--allowedTools Read --settings='{SETTINGS_JSON}'", + f"--allowedTools Read --mcp-config='{MCP_JSON}'", + "--allowedTools Read --settings=./ci/settings.json", + "--allowedTools Read --mcp-config .mcp.json", + ], + ids=["settings-attached-json", "mcp-config-attached-json", "settings-path", "mcp-config-path"], +) +def test_a_settings_or_mcp_config_flag_leaves_claude_args_unread(value): + """#823 review cycle 4 scope: such a value is never published, however it is attached.""" + + launch, = _launches(_workflow(_agent(value))) + assert launch["settings"] == [{"name": "claude_args", "value": _digest(value), "unresolved_reason": "unread_arguments"}] + for canary in JSON_CANARIES: + assert canary not in json.dumps(launch) + + +def test_the_word_after_a_secret_named_argument_is_withheld(): + """As the host readers withhold it among an MCP server's arguments.""" + + grant = _grant(_workflow( + _agent("--allowedTools Read --token canary-arg-1 --max-turns 5"), + {"run": "claude -p --allowedTools Read password canary-arg-2 go"}, + )) + action, run = grant["agent_launches"] + assert action["settings"] == [{ + "name": "claude_args", "value": "--allowedTools Read --token --max-turns 5", + "unresolved_reason": "redacted", + }] + assert run["settings"] == [ + {"name": "--allowedTools", "value": "Read password go", "unresolved_reason": "redacted"}, + ] + assert "canary-arg" not in json.dumps(grant) + # An edit inside what is redacted is not reported, so it is named as a limit. + assert [limit.split(";")[0] for limit in uncompared_agent_launch_texts(grant)] == [ + "the claude_args value of the agent launch at review/steps[0] (anthropics/claude-code-action) contains " + "credential-shaped text", + "the --allowedTools value of the agent launch at review/steps[1] (claude) contains credential-shaped text", + ] + + +def test_a_withheld_json_value_compares_as_the_host_readers_compare_it(): + def settings(env): + return _workflow(_agent(settings=json.dumps({"env": env}))) + + # An env value is not compared, as in `.claude/settings.json`; an added key is. + assert _rows(settings({"DB": "one"}), settings({"DB": "two"})) == [] + row, = _rows(settings({"DB": "one"}), settings({"DB": "one", "EXTRA": "three"})) + assert '"EXTRA":""' in row.after + assert "three" not in row.after + + +@pytest.mark.parametrize("name", ["settings", "mcp_config"]) +def test_a_json_or_path_input_that_is_neither_publishes_only_a_digest(name): + """#823 review cycle 5 (P3): text before the JSON, such as a comment line, published the env value it holds.""" + + value = '// ci\n{"env": {"DB_PASSWORD_PLAIN": "canary-env-value"}}' + launch, = _launches(_workflow(_agent(**{name: value}))) + setting = next(item for item in launch["settings"] if item["name"] == name) + assert setting == {"name": name, "value": _digest(value), "unresolved_reason": None} + assert "canary-env-value" not in json.dumps(launch) + # Still compared: an edit to it is a row, and one showing neither text. + row, = _rows(_workflow(_agent(**{name: value})), _workflow(_agent(**{name: value.replace("canary", "other")}))) + assert (row.direction, row.expands) == ("changed", False) + assert "canary" not in row.before + row.after and "other-env" not in row.after + # A plain path, an expression naming one included, is published as written. + for path in (".github/claude-settings.json", "${{ github.workspace }}/ci/settings.json"): + launch, = _launches(_workflow(_agent(**{name: path}))) + assert next(item for item in launch["settings"] if item["name"] == name)["value"] == path + + +def test_a_codex_config_override_publishes_its_key_and_withholds_its_value(): + """#823 review cycle 4: only a rule-bearing key's value, the approval policy and the model are published.""" + + run = ( + "codex exec -c mcp_servers.db.env.TOKEN=canary-cfg -c mcp_servers.gh.command=/opt/canary-bin/gh " + "-c shell_environment_policy.set.LEVEL=canary-env -c model=o3 go" + ) + action_args = "-cmcp_servers.db.env.REGION=canary-short -c=mcp_servers.x.url=https://canary.example/p --yolo" + grant = _grant(_workflow( + {"run": run}, {"uses": "openai/codex-action@v1", "with": {"codex-args": action_args}}, + )) + cli, action = grant["agent_launches"] + + assert cli["settings"] == [ + {"name": "--config", "value": "mcp_servers.db.env.TOKEN=", "unresolved_reason": None}, + {"name": "--config", "value": f"mcp_servers.gh.command={_digest('/opt/canary-bin/gh')}", + "unresolved_reason": None}, + {"name": "--config", "value": "model=o3", "unresolved_reason": None}, + {"name": "--config", "value": f"shell_environment_policy.set.LEVEL={_digest('canary-env')}", + "unresolved_reason": None}, + ] + assert action["settings"] == [{ + "name": "codex-args", + "value": ( + "-cmcp_servers.db.env.REGION= " + f"-c=mcp_servers.x.url={_digest('https://canary.example/p')} --yolo" + ), + "unresolved_reason": None, + }] + assert action["widening_rules"] == [{"rule": "bypass_approvals_and_sandbox", "setting": "codex-args"}] + assert "canary" not in json.dumps(grant) + # Rotating a redacted value is quiet; editing a withheld one is a change. + assert _rows(_workflow({"run": run}), _workflow({"run": run.replace("canary-cfg", "rotated")})) == [] + row, = _rows(_workflow({"run": run}), _workflow({"run": run.replace("canary-env", "edited")})) + assert row.direction == "changed" + + +def test_text_that_starts_like_json_and_does_not_parse_is_withheld_and_named(): + value = '{"env": {"T": "canary-unparsed"}' + grant = _grant(_workflow(_agent(settings=value))) + launch, = grant["agent_launches"] + + setting = next(item for item in launch["settings"] if item["name"] == "settings") + assert setting == {"name": "settings", "value": None, "unresolved_reason": "unparsed_json"} + assert "canary-unparsed" not in json.dumps(grant) + limit, = uncompared_agent_launch_texts(grant) + assert limit.startswith("the settings value of the agent launch at review/steps[0] (anthropics/claude-code-action)") + assert "holds text that starts like JSON and does not parse" in limit + assert _uncompared_workflow_text(grant) is None + + +def test_a_url_path_is_withheld_while_the_rest_of_the_setting_and_a_rule_beside_it_are_read(): + before = _workflow(_agent("--append-system-prompt Follow https://example.com/style-guide --allowedTools Read")) + after = _workflow(_agent("--append-system-prompt Follow https://example.com/style-guide --dangerously-skip-permissions")) + + row, = _rows(before, after) + assert (row.direction, row.expands) == ("widened", True) + assert ( + "claude_args: --append-system-prompt Follow https://example.com/ " + "--dangerously-skip-permissions" + ) in row.after + assert "style-guide" not in row.before + row.after + assert uncompared_agent_launch_texts(_grant(after)) == [] + assert _uncompared_workflow_text(_grant(after)) is None + + +def test_a_marketplace_url_compares_by_scheme_and_host_as_an_mcp_server_url_does(): + def marketplace(url): + return _workflow(_agent(plugin_marketplaces=url)) + + launch, = _launches(marketplace("https://github.com/anthropics/claude-code.git")) + assert { + "name": "plugin_marketplaces", "value": "https://github.com/", "unresolved_reason": None, + } in launch["settings"] + row, = _rows(marketplace("https://github.com/org/a.git"), marketplace("https://gitlab.example.com/org/a.git")) + assert "plugin_marketplaces: https://gitlab.example.com/" in row.after + # The path is withheld as an MCP server URL's is (#723), so a change only + # there is not reported; the support page says so. + assert _rows(marketplace("https://github.com/org/a.git"), marketplace("https://github.com/org/b.git")) == [] + + +def test_credential_shaped_text_in_a_setting_is_published_redacted_and_named_while_a_ref_blocks(): + """#823 review F2: a redacted setting compares by its published text and rules, not blocking. + + A checkout ref names the code a job runs, as a step reference does, so a + redacted one still refuses (#767). + """ + + grant = _grant(_workflow( + _agent("--append-system-prompt use token=ARGCANARY --dangerously-skip-permissions", + plugin_marketplaces="https://robot:PWCANARY@github.com/org/repo.git"), + {"uses": "actions/checkout@v4", "with": {"ref": "token=REFCANARY"}}, + {"run": "claude -p --permission-prompt-tool ghp_" + "A" * 36 + " review token=SECRETCANARY"}, + )) + action, cli = grant["agent_launches"] + + assert action["settings"] == [ + {"name": "claude_args", + "value": "--append-system-prompt use token= --dangerously-skip-permissions", + "unresolved_reason": "redacted"}, + {"name": "plugin_marketplaces", "value": "https://github.com/", "unresolved_reason": "redacted"}, + ] + # The rule is read from the declared text, so redaction does not hide it. + assert action["widening_rules"] == [{"rule": "bypass_permissions", "setting": "claude_args"}] + assert cli["settings"] == [ + {"name": "--permission-prompt-tool", "value": "[REDACTED:github_token]", "unresolved_reason": "redacted"}, + ] + assert grant["checkout_refs"] == [ + {"job": "review", "step": "steps[1]", "ref": "token=", "unresolved_reason": "redacted"}, + ] + text = json.dumps(grant) + for canary in ("ARGCANARY", "PWCANARY", "REFCANARY", "SECRETCANARY", "ghp_"): + assert canary not in text + assert uncompared_agent_launch_texts(grant) == [ + f"the {name} value of the agent launch at {where} contains credential-shaped text; it is published " + "redacted and compared as published, so an edit inside what is redacted that gains no documented " + "widening rule is not reported" + for name, where in ( + ("claude_args", "review/steps[0] (anthropics/claude-code-action)"), + ("plugin_marketplaces", "review/steps[0] (anthropics/claude-code-action)"), + ("--permission-prompt-tool", "review/steps[2] (claude)"), + ) + ] + assert _uncompared_workflow_text(grant) == ( + "a checkout ref contains credential-shaped text; it is published redacted and cannot be compared" + ) + + +#: Prose the #802 label redaction rewrites, as security-review prompts write it: +#: ``claude_args``, and a ``run:`` whose prompt is a variadic flag's value. +PROSE = [ + ("--append-system-prompt Never print bearer tokens in review comments --allowedTools Read", + "claude -p --allowedTools Read Never print bearer tokens in review comments"), + ("--append-system-prompt Flag Authorization: headers logged in plain text --allowedTools Read", + "claude -p --allowedTools Read Flag Authorization: headers logged in plain text"), + ("--append-system-prompt check the secret=... assignment --allowedTools Read", + "claude -p --allowedTools Read check the secret=... assignment"), +] + + +@pytest.mark.parametrize(("prose", "run"), PROSE, ids=["bearer", "authorization", "assignment"]) +def test_prose_the_label_redaction_rewrites_is_a_named_limit_that_refuses_nothing(prose, run): + """#823 review F2: such prose used to make the whole workflow a blocking limit.""" + + for step in (_agent(prose), {"run": run}): + grant = _grant(_workflow(step)) + launch, = grant["agent_launches"] + assert [setting["unresolved_reason"] for setting in launch["settings"]] == ["redacted"] + assert _uncompared_workflow_text(grant) is None + limit, = uncompared_agent_launch_texts(grant) + assert "contains credential-shaped text" in limit + + # A permission change beside it keeps its row, and a rule gained beside it widens. + row, = _rows(_workflow(_agent(prose)), _workflow(_agent(prose), permissions={"pull-requests": "write"})) + assert row.direction == "widened" and "grants write permissions to workflow jobs" in row.why + row, = _rows(_workflow(_agent(prose)), _workflow(_agent(f"{prose} --dangerously-skip-permissions"))) + assert (row.direction, row.expands) == ("widened", True) + # Each cell shows the redacted text it is compared by, so the two sides differ. + assert "" in row.before and row.after.endswith("--dangerously-skip-permissions") + + +def test_an_expression_in_a_url_is_read_as_one_word_so_the_url_is_withheld_whole(): + """#823 review (P3): the expression's spaces used to split the URL, publishing its path.""" + + launch, = _launches(_workflow(_agent( + plugin_marketplaces="https://x-access-token:${{ secrets.MARKET_TOKEN }}@github.com/acme/market.git", + ))) + setting = next(item for item in launch["settings"] if item["name"] == "plugin_marketplaces") + assert setting == { + "name": "plugin_marketplaces", "value": "https://github.com/", + "unresolved_reason": "redacted", "holds_expression": True, + } + # An expression outside a URL is published as written. + ref, = _grant(_reproduction(ref=HEAD_SHA))["checkout_refs"] + assert ref["ref"] == HEAD_SHA + + +def test_token_shaped_job_and_step_labels_are_redacted_in_every_entry(): + job = "ghp_" + "B" * 36 + grant = _grant(_workflow(jobs={job: {"steps": [ + {"name": "Pull docker://ci:hunter2@gcr.io/x", "uses": "actions/checkout@v4"}, + {"name": "Run ghp_" + "C" * 36, "run": "claude -p go"}, + {"name": "Unread ghp_" + "E" * 36, "run": "npm ci && claude -p go"}, + ]}})) + + text = json.dumps({key: grant[key] for key in ("agent_launches", "checkout_refs", "unread_agent_runs")}) + assert "ghp_" not in text and "hunter2" not in text + assert grant["checkout_refs"][0]["step"] == "Pull docker://@gcr.io/x" + assert grant["agent_launches"][0]["job"] == "[REDACTED:github_token]" + assert grant["unread_agent_runs"][0]["job"] == "[REDACTED:github_token]" + assert "ghp_" not in " ".join(uncompared_agent_launch_texts(grant)) + + +# --- saved baselines --------------------------------------------------------------- + + +def _git(repo: Path, *args: str) -> str: + return subprocess.run( + ["git", "-C", str(repo), *args], check=True, capture_output=True, text=True + ).stdout.strip() + + +def _repo(tmp_path: Path, files: dict[str, str]) -> Path: + repo = tmp_path / "repo" + repo.mkdir() + _git(repo, "init", "-q", "-b", "main") + _git(repo, "config", "user.name", "Fixture") + _git(repo, "config", "user.email", "fixture@example.invalid") + _write(repo, {".gitignore": "agents-shipgate-reports/\n", **files}) + _git(repo, "add", ".") + _git(repo, "commit", "-qm", "base") + return repo + + +def _write(repo: Path, files: dict[str, str]) -> None: + for name, text in files.items(): + path = repo / name + path.parent.mkdir(parents=True, exist_ok=True) + path.write_text(text, encoding="utf-8") + + +def _yaml(value) -> str: + return yaml.safe_dump(value, sort_keys=False) + + +def _v06_baseline(workspace: Path): + from agents_shipgate.cli.host_audit import host_audit_inventory + from agents_shipgate.core.host_grants import build_host_grants_baseline, host_grants_sha256 + + current = host_audit_inventory(workspace) + legacy = build_host_grants_baseline(current) + legacy["host_grants_schema_version"] = "0.6" + for grant in legacy["inventory"]["grants"]: + grant.pop("agent_launches", None) + grant.pop("unread_agent_runs", None) + grant.pop("checkout_refs", None) + legacy["inventory_sha256"] = host_grants_sha256(legacy["inventory"]) + return current, legacy + + +def test_a_v0_6_baseline_holding_a_workflow_does_not_assert_no_agent_launches(tmp_path): + from agents_shipgate.core.host_grants import build_host_drift_payload, load_host_grants_baseline + + _write(tmp_path, {SOURCE: _yaml(_reproduction())}) + current, legacy = _v06_baseline(tmp_path) + path = tmp_path / "baseline.json" + path.write_text(json.dumps(legacy)) + + drift = build_host_drift_payload( + baseline=load_host_grants_baseline(path), inventory=current, baseline_file=str(path) + ) + assert drift["comparison_status"] == "incomparable" + assert drift["incomparable_reasons"] == ["baseline_workflow_agent_launches_unavailable"] + assert drift["has_drift"] is None and drift["changes"] == [] + + +def test_a_v0_6_baseline_without_a_workflow_stays_comparable_and_can_be_replaced(tmp_path): + from agents_shipgate.core.host_grants import build_host_drift_payload + + _write(tmp_path, {".claude/settings.json": json.dumps({"permissions": {"allow": ["Read(**)"]}})}) + current, legacy = _v06_baseline(tmp_path) + drift = build_host_drift_payload(baseline=legacy, inventory=current, baseline_file="b.json") + assert (drift["comparison_status"], drift["has_drift"]) == ("comparable", False) + + path = tmp_path / ".agents-shipgate" / "host-grants.json" + path.parent.mkdir() + original = json.dumps(legacy, indent=2, sort_keys=True) + "\n" + path.write_text(original) + audit = ["audit", "--host", "--workspace", str(tmp_path), "--baseline-file", str(path)] + resaved = CliRunner().invoke(app, [*audit, "--save-baseline"]) + assert resaved.exit_code == 0, resaved.output + assert json.loads(path.read_text())["host_grants_schema_version"] == "0.7" + + +def test_the_documented_migration_from_a_v0_6_baseline_holding_a_workflow(tmp_path): + from tests.test_preflight import _workspace + + root = _workspace(tmp_path) + _write(root, {SOURCE: _yaml(_reproduction())}) + _, legacy = _v06_baseline(root) + path = root / ".agents-shipgate" / "host-grants.json" + path.parent.mkdir(parents=True, exist_ok=True) + original = json.dumps(legacy, indent=2, sort_keys=True) + "\n" + path.write_text(original) + audit = ["audit", "--host", "--workspace", str(root), "--baseline-file", str(path)] + + payload = json.loads(CliRunner().invoke(app, [*audit, "--drift", "--json"]).stdout) + assert payload["comparison_status"] == "incomparable" + assert payload["incomparable_reasons"] == ["baseline_workflow_agent_launches_unavailable"] + assert payload["has_drift"] is None and payload["next_action"] is None + assert CliRunner().invoke(app, [*audit, "--drift", "--fail-on-drift", "--json"]).exit_code == 20 + + preflight = CliRunner().invoke(app, ["preflight", "--workspace", str(root), "--json"]) + assert preflight.exit_code == 0, preflight.output + signal, = [item for item in json.loads(preflight.stdout)["signals"] if item["kind"] == "host_grant_drift"] + assert (signal["severity"], signal["actor"]) == ("high", "human") + assert "baseline_workflow_agent_launches_unavailable" in json.dumps(signal) + + refused = CliRunner().invoke(app, [*audit, "--save-baseline"]) + assert refused.exit_code == 2 + assert "unsupported_baseline_schema" in refused.output + (refused.stderr or "") + assert path.read_text() == original + + path.rename(path.with_name("host-grants.v0.6.json")) + resaved = CliRunner().invoke(app, [*audit, "--save-baseline"]) + assert resaved.exit_code == 0, resaved.output + after = json.loads(CliRunner().invoke(app, [*audit, "--drift", "--json"]).stdout) + assert (after["comparison_status"], after["has_drift"]) == ("comparable", False) + assert path.with_name("host-grants.v0.6.json").read_text() == original + + +def test_a_current_baseline_compares_agent_launches_and_validates_against_the_schemas(tmp_path): + from jsonschema import Draft202012Validator + + from agents_shipgate.cli.host_audit import host_audit_inventory + from agents_shipgate.core.host_grants import ( + build_host_drift_payload, + build_host_grants_baseline, + ) + + path = tmp_path / SOURCE + path.parent.mkdir(parents=True) + path.write_text(_yaml(_reproduction(run="npm ci && claude -p 'x'", ref=HEAD_SHA))) + inventory = host_audit_inventory(tmp_path) + baseline = build_host_grants_baseline(inventory) + assert baseline["host_grants_schema_version"] == "0.7" + workflow, = [grant for grant in baseline["inventory"]["grants"] if grant["kind"] == "workflow"] + assert workflow["unread_agent_runs"] == [{"job": "review", "step": "steps[2]", "agent": "claude"}] + for name, payload in (("inventory", inventory), ("baseline", baseline)): + schema = json.loads((ROOT / f"docs/host-grants-{name}-schema.v0.7.json").read_text()) + Draft202012Validator(schema).validate(payload) + + path.write_text(_yaml(_reproduction(claude_args="--dangerously-skip-permissions", ref=HEAD_SHA))) + drift = build_host_drift_payload(baseline=baseline, inventory=host_audit_inventory(tmp_path), baseline_file="b.json") + assert (drift["comparison_status"], drift["has_drift"]) == ("comparable", True) + assert drift["expansion_signals"] == [f"workflow_agent_widened_changed: {SOURCE}"] + schema = json.loads((ROOT / "docs/host-grants-drift-schema.v0.7.json").read_text()) + Draft202012Validator(schema).validate(drift) + + +def test_an_unread_run_is_a_non_blocking_limit_that_leaves_coverage_complete(tmp_path): + from agents_shipgate.cli.host_audit import host_audit_inventory + + _write(tmp_path, {SOURCE: _yaml(_workflow({"run": "npm ci && claude -p 'go'"}))}) + inventory = host_audit_inventory(tmp_path) + + github, = [item for item in inventory["host_coverage"] if item["host"] == "github"] + assert github["status"] == "complete" + issue, = [item for item in inventory["issues"] if item["host"] == "github"] + assert (issue["kind"], issue["blocking"]) == ("unsupported", False) + assert issue["message"].startswith( + "the `run:` at review/steps[0] mentions claude and is not read as an agent launch" + ) + + +@pytest.mark.parametrize("run", UNREAD_RUNS[:8], ids=[f"review-cycle-4-{index}" for index in range(8)]) +def test_the_run_forms_of_review_cycle_4_give_no_row_and_are_a_named_coverage_issue(tmp_path, run): + """#823 review cycle 4 (a) end to end: named in `audit --host`, never a row, none of the text published.""" + + from agents_shipgate.cli.host_audit import host_audit_inventory + + permissions = {"contents": "write", "pull-requests": "write"} + base = _workflow({"uses": "actions/checkout@v4"}, trigger="issue_comment", permissions=permissions) + head = _workflow({"uses": "actions/checkout@v4"}, {"run": run}, trigger="issue_comment", permissions=permissions) + repo = _repo(tmp_path, {SOURCE: _yaml(base)}) + _git(repo, "checkout", "-qb", "change") + _write(repo, {SOURCE: _yaml(head)}) + _git(repo, "commit", "-qam", "an agent step") + + assert _diff(repo)["rows"] == [] + inventory = host_audit_inventory(repo) + workflow, = [grant for grant in inventory["grants"] if grant.get("kind") == "workflow"] + assert "agent_launches" not in workflow + issues = [item for item in inventory["issues"] if item["host"] == "github"] + assert issues and all(not item["blocking"] for item in issues) + assert all("review/steps[1] mentions" in item["message"] for item in issues) + audit = CliRunner().invoke(app, ["audit", "--host", "--workspace", str(repo)]) + assert "review/steps[1] mentions" in audit.output + joined = json.dumps(inventory) + audit.output + for text in ("dangerously", "yolo", "Review this PR", "review this change"): + assert text not in joined + + +def test_a_change_that_only_adds_an_unread_step_says_what_the_workflow_does_not_compare(tmp_path): + """#823 review cycle 7 (carried P3): the coverage line named env values and apiKeyHelper.""" + + base = _workflow({"uses": "actions/checkout@v4"}) + head = _workflow({"uses": "actions/checkout@v4"}, {"run": f"claude -p {BYPASS} 'Review'"}) + repo = _repo(tmp_path, {SOURCE: _yaml(base)}) + _git(repo, "checkout", "-qb", "change") + _write(repo, {SOURCE: _yaml(head)}) + _git(repo, "commit", "-qam", "an unread agent step") + + result = CliRunner().invoke(app, ["diff", "--workspace", str(repo), "--base", "main"]) + assert result.exit_code == 0, result.output + assert "No static host-grant changes detected." in result.output + assert ( + f"{SOURCE} (github): compared; changed, but no grant this entry compares changed, so no row " + "(text this entry does not read, such as a step's env or an unread agent step, is not " + "compared; audit --host names each unread agent step)" + ) in result.output + assert "apiKeyHelper" not in result.output + assert "dangerously" not in result.output + + +@pytest.mark.parametrize( + "still_there", + [f'claude -p {BYPASS} "Review"', f"npx @anthropic-ai/claude-code -p {BYPASS} Review"], + ids=["quoted", "npx"], +) +def test_a_launch_the_job_still_runs_unread_has_not_moved_to_the_job_that_adds_it(tmp_path, still_there): + """#823 review cycle 5 (C5-F1, M5 and M6) end to end: `diff` widens, `audit --host` names the unread step.""" + + from agents_shipgate.cli.host_audit import host_audit_inventory + + permissions = {"contents": "write"} + plain = f"claude -p {BYPASS} Review" + base = _workflow(jobs={"a": {"steps": [{"run": plain}]}, "b": {"steps": [{"run": "echo hi"}]}}, + permissions=permissions) + head = _workflow(jobs={"a": {"steps": [{"run": still_there}]}, "b": {"steps": [{"run": plain}]}}, + permissions=permissions) + repo = _repo(tmp_path, {SOURCE: _yaml(base)}) + _git(repo, "checkout", "-qb", "change") + _write(repo, {SOURCE: _yaml(head)}) + _git(repo, "commit", "-qam", "an agent step in b") + + row, = _diff(repo)["rows"] + assert (row["direction"], row["expands"]) == ("widened", True) + assert "an agent launch now skips permission checks (bypassPermissions) (b/steps[0])" in row["why"] + assert "a step no longer declares an agent launch this audit reads (a/steps[0])" in row["why"] + assert "moved between jobs" not in row["why"] + workflow, = [grant for grant in host_audit_inventory(repo)["grants"] if grant.get("kind") == "workflow"] + assert workflow["unread_agent_runs"] == [{"job": "a", "step": "steps[0]", "agent": "claude"}] + + +def test_a_launch_merged_into_an_unread_step_it_did_not_hold_has_not_moved(tmp_path): + """#823 review cycle 6 (C6-F1, V1) end to end: the job keeps as many unread steps, at another step.""" + + from agents_shipgate.cli.host_audit import host_audit_inventory + + permissions = {"contents": "write"} + plain = f"claude -p {BYPASS} Review" + install = "npm i -g @anthropic-ai/claude-code" + base = _workflow(jobs={"a": {"steps": [{"run": install}, {"run": plain}]}, "b": {"steps": [{"run": "echo hi"}]}}, + permissions=permissions) + head = _workflow(jobs={"a": {"steps": [{"run": f"{install} && {plain}"}]}, "b": {"steps": [{"run": plain}]}}, + permissions=permissions) + repo = _repo(tmp_path, {SOURCE: _yaml(base)}) + _git(repo, "checkout", "-qb", "change") + _write(repo, {SOURCE: _yaml(head)}) + _git(repo, "commit", "-qam", "merge the install into the launch, and launch in b") + + payload = _diff(repo) + row, = payload["rows"] + assert (row["direction"], row["expands"]) == ("widened", True) + assert "an agent launch now skips permission checks (bypassPermissions) (b/steps[0])" in row["why"] + assert "a step no longer declares an agent launch this audit reads (a/steps[1])" in row["why"] + assert "may still start an agent in a way this audit does not read" in row["why"] + assert "moved between jobs" not in row["why"] + workflow, = [grant for grant in host_audit_inventory(repo)["grants"] if grant.get("kind") == "workflow"] + assert workflow["unread_agent_runs"] == [{"job": "a", "step": "steps[0]", "agent": "claude"}] + + +def test_a_codex_config_integer_past_the_digit_limit_crashes_no_route(tmp_path): + """#823 review cycle 6 (C6-F2): `diff`, `audit --host` and `check` exit 0, and `verify` is no internal error.""" + + repo = _repo(tmp_path, {SOURCE: _yaml(_workflow({"run": "echo hi"}))}) + _git(repo, "checkout", "-qb", "change") + _write(repo, {SOURCE: _yaml(_workflow({"run": "codex exec -c sandbox_mode=" + "1" * 5000 + " Review"}))}) + _git(repo, "commit", "-qam", "a codex step") + + for args in ( + ["diff", "--workspace", str(repo), "--base", "main", "--json"], + ["audit", "--host", "--workspace", str(repo), "--json"], + ["check", "--workspace", str(repo), "--base", "main", "--head", "HEAD", "--format", "agent-boundary-json"], + ["verify", "--workspace", str(repo), "--base", "main", "--head", "HEAD", "--format", "text"], + ): + result = CliRunner().invoke(app, args) + assert result.exit_code == 0, (args[0], result.output, result.exception) + row, = _diff(repo)["rows"] + assert (row["direction"], row["expands"]) == ("changed", False) + + +# --- the same row on every route --------------------------------------------------- + + +@pytest.fixture +def pr(tmp_path): + repo = _repo(tmp_path, {SOURCE: _yaml(_reproduction())}) + _git(repo, "checkout", "-qb", "change") + _write(repo, {SOURCE: _yaml(_reproduction(claude_args="--permission-mode bypassPermissions --allowedTools Bash"))}) + _git(repo, "add", ".") + _git(repo, "commit", "-qm", "bypass permissions") + return repo + + +def _assert_the_row(row: dict) -> None: + assert row["subject"] == f"github {SOURCE}" + assert "review/steps[1]: runs anthropics/claude-code-action with claude_args: --allowedTools Read" in row["before"] + assert "--permission-mode bypassPermissions" in row["after"] + assert (row["direction"], row["expands"]) == ("widened", True) + assert "skips permission checks" in row["why"] + + +def _diff(repo: Path, *args: str) -> dict: + result = CliRunner().invoke(app, ["diff", "--workspace", str(repo), "--base", "main", *args, "--json"]) + assert result.exit_code == 0, result.output + return json.loads(result.output) + + +def test_diff_names_the_widening_in_json_and_text(pr): + payload = _diff(pr) + assert payload["comparison_status"] == "comparable" + row, = payload["rows"] + _assert_the_row(row) + + text = CliRunner().invoke(app, ["diff", "--workspace", str(pr), "--base", "main"]) + assert text.exit_code == 0, text.output + assert "⚠" in text.output and "review/steps[1]" in text.output + assert "1 widening what the agent may do" in text.output + + +def test_manifest_free_verify_and_its_pr_comment_name_the_widening(pr): + args = ["verify", "--workspace", str(pr), "--base", "main", "--head", "HEAD", "--format", "text"] + result = CliRunner().invoke(app, args) + assert result.exit_code == 0, result.output + assert "bypassPermissions" in result.output + + comment = (pr / "agents-shipgate-reports/pr-comment.md").read_text() + assert "bypassPermissions" in comment and "review/steps[1]" in comment + verifier = json.loads((pr / "agents-shipgate-reports/verifier.json").read_text()) + row, = verifier["host_comparison"]["rows"] + _assert_the_row(row) + + +def test_check_and_the_control_envelope_carry_the_same_row(pr): + args = ["check", "--workspace", str(pr), "--base", "main", "--head", "HEAD"] + machine = CliRunner().invoke(app, [*args, "--format", "agent-boundary-json"]) + assert machine.exit_code == 0, machine.output + row, = json.loads(machine.output)["rows"] + _assert_the_row(row) + + control = CliRunner().invoke(app, [*args, "--format", "agent-control-json"]) + assert control.exit_code == 0, control.output + envelope_row, = json.loads(control.output)["capability_rows"]["rows"] + _assert_the_row(envelope_row) + + +def test_the_stop_hook_announces_the_widening(pr, tmp_path): + from tests.test_install_hooks import _host_diff_workspace, _run_hook + + payload = _diff(pr) + hooked = tmp_path / "hooked" + hooked.mkdir() + _host_diff_workspace(hooked) + result = _run_hook(hooked, "verify", {}, diff_payload=json.dumps(payload)) + + assert result.returncode == 0, result.stderr + message = json.loads(result.stdout)["systemMessage"] + assert "These rows widen what the agent can do" in message + assert SOURCE in message + + +def _published_outputs(repo: Path) -> str: + """Every route's output for the change on ``repo``: diff, audit, check, verify and its reports.""" + + outputs = [] + for args in ( + ["diff", "--workspace", str(repo), "--base", "main", "--json"], + ["diff", "--workspace", str(repo), "--base", "main"], + ["audit", "--host", "--workspace", str(repo), "--json"], + ["audit", "--host", "--workspace", str(repo)], + ["check", "--workspace", str(repo), "--base", "main", "--head", "HEAD", "--format", "agent-boundary-json"], + ["check", "--workspace", str(repo), "--base", "main", "--head", "HEAD", "--format", "agent-control-json"], + ["verify", "--workspace", str(repo), "--base", "main", "--head", "HEAD", "--format", "text"], + ): + result = CliRunner().invoke(app, args) + outputs.append(result.output) + outputs.append((repo / "agents-shipgate-reports/pr-comment.md").read_text()) + outputs.append((repo / "agents-shipgate-reports/verifier.json").read_text()) + return "\n".join(outputs) + + +def _assert_absent(joined: str, canaries) -> None: + for canary in canaries: + assert canary not in joined, canary + digest = hashlib.sha256(canary.encode()).hexdigest() + for prefix in (digest, digest[:24], digest[:12]): + assert prefix not in joined, canary + + +def test_no_canary_reaches_any_published_output(tmp_path): + """The #802 sweep for agent launches, JSON-shaped canaries included (#823 review F3, cycle 4).""" + + canary = "sk-ant-api03-" + "Z" * 40 + job = "ghp_" + "D" * 36 + mcp = json.dumps({"mcpServers": {"db": { + "command": "db-mcp", "env": {"DB_API_TOKEN": "canary-tok-123"}, + "headers": {"X-API-Key": "canary-hdr-456", "Authorization": f"Bearer {canary}"}, + }}}) + codex_args = ( + "-c mcp_servers.db.env.TOKEN=canary-cfg-789 -cmcp_servers.db.env.REGION=canary-short-c " + "-c=mcp_servers.db.env.ZONE=canary-eq-c -c mcp_servers.db.command=/opt/canary-cmd-c --full-auto" + ) + # The `=` spellings of #823 review F1, beside the separate-word ones. + equals = json.dumps({"env": {"DB_PASSWORD": "hunter2-eqcanary"}, "apiKeyHelper": "echo helper-eqcanary"}) + equals_mcp = json.dumps({"mcpServers": {"db": { + "command": "db-mcp", "env": {"DB_API_TOKEN": "tok-eqcanary"}, "headers": {"X-API-Key": "hdr-eqcanary"}, + }}}) + base = _workflow(jobs={job: {"steps": [_agent()]}}) + head = _workflow(jobs={job: {"steps": [ + {"name": "Pull docker://ci:" + "p4ssCANARY" + "@gcr.io/x", "uses": "actions/checkout@v4"}, + _agent( + "--allowedTools Read --dangerously-skip-permissions", + settings=SETTINGS_JSON, mcp_config=mcp, + plugin_marketplaces="https://github.com/canary-org/canary-repo.git", + ), + # #823 review cycle 4: `claude_args` or a `run:` holding JSON, `--settings` + # or `--mcp-config` is not read, and publishes only a digest or nothing. + _agent(f"--mcp-config '{mcp}' --dangerously-skip-permissions"), + _agent(f"--allowedTools Read --settings='{equals}' --mcp-config='{equals_mcp}'"), + {"run": f"ANTHROPIC_API_KEY={canary} claude -p --allowedTools Read --mcp-config '{mcp}' 'go'"}, + {"run": f"ANTHROPIC_API_KEY={canary} claude -p --allowedTools Read go"}, + {"uses": "openai/codex-action@v1", "with": {"codex-args": codex_args}}, + # #823 review C2-F1: an MCP server's arguments and a hook's command, + # which the host readers never publish, in every spelling. + _agent("--allowedTools Read", mcp_config=REMOTE_MCP_JSON, settings=HOOK_JSON), + _agent(f"--allowedTools Read\n--mcp-config '{REMOTE_MCP_JSON}'"), + {"run": f"claude -p --settings '{HOOK_JSON}' --mcp-config '{REMOTE_MCP_JSON}' 'go'"}, + {"run": f"codex exec -c '{REMOTE_CODEX_CONFIG}' 'go'"}, + ]}}) + repo = _repo(tmp_path, {SOURCE: _yaml(base)}) + _git(repo, "checkout", "-qb", "change") + _write(repo, {SOURCE: _yaml(head)}) + _git(repo, "commit", "-qam", "canaries") + + payload = _diff(repo) + assert payload["comparison_status"] == "comparable" + row, = payload["rows"] + assert (row["direction"], row["expands"]) == ("widened", True) + assert '"headers":{"Authorization":"","X-API-Key":""}' in row["after"] + assert f"settings: {SETTINGS_PUBLISHED}" in row["after"] + joined = _published_outputs(repo) + + _assert_absent(joined, ( + canary, "p4ssCANARY", job, "canary-org", "canary-repo", "canary-cfg-789", *JSON_CANARIES, + "hunter2-eqcanary", "helper-eqcanary", "tok-eqcanary", "hdr-eqcanary", "canary-short-c", "canary-eq-c", + "canary-cmd-c", *SHAPE_CANARIES, + )) + assert "runs claude -p with --allowedTools Read" in joined + assert "mcp_servers.db.env.TOKEN=" in joined + assert '"command":"npx"' in joined and '"Stop":[{"hooks":[{"command":" --dangerously-skip-permissions" in row["after"] + audit = json.loads(CliRunner().invoke(app, ["audit", "--host", "--workspace", str(repo), "--json"]).stdout) + issue, = [item for item in audit["issues"] if item["host"] == "github"] + assert (issue["kind"], issue["blocking"]) == ("unsupported", False) + github, = [item for item in audit["host_coverage"] if item["host"] == "github"] + assert github["status"] == "complete" + _assert_absent(_published_outputs(repo), ("ARGCANARY",)) + + +def _prose_repo(tmp_path: Path, head: dict, extra: dict[str, str] | None = None) -> Path: + base = _workflow(_agent(PROSE[0][0])) + repo = _repo(tmp_path, {SOURCE: _yaml(base)}) + _git(repo, "checkout", "-qb", "change") + _write(repo, {SOURCE: _yaml(head), **(extra or {})}) + _git(repo, "add", ".") + _git(repo, "commit", "-qm", "change") + return repo + + +def test_a_permission_change_beside_redacted_prose_keeps_its_row_on_every_route(tmp_path): + """#823 review F2 (a): the prose used to refuse the comparison and hide this row.""" + + repo = _prose_repo(tmp_path, _workflow(_agent(PROSE[0][0]), permissions={"contents": "read", "pull-requests": "write"})) + + payload = _diff(repo) + assert payload["comparison_status"] == "comparable" + row, = payload["rows"] + assert row["direction"] == "widened" and "grants write permissions to workflow jobs" in row["why"] + result = CliRunner().invoke(app, ["verify", "--workspace", str(repo), "--base", "main", "--head", "HEAD"]) + assert result.exit_code == 0, result.output + verifier = json.loads((repo / "agents-shipgate-reports/verifier.json").read_text()) + assert verifier["host_comparison"]["comparison_status"] == "comparable" + verified, = verifier["host_comparison"]["rows"] + assert verified["direction"] == "widened" + assert "Host capability comparison unavailable" not in (repo / "agents-shipgate-reports/pr-comment.md").read_text() + + +def test_an_unchanged_workflow_holding_redacted_prose_leaves_check_comparable(tmp_path): + """#823 review F2 (b): an unchanged workflow used to make every `check` incomparable (#721).""" + + mcp = json.dumps({"mcpServers": {"docs": {"command": "docs-mcp"}}}) + repo = _prose_repo(tmp_path, _workflow(_agent(PROSE[0][0])), extra={".mcp.json": mcp}) + + args = ["check", "--workspace", str(repo), "--base", "main", "--head", "HEAD", "--format", "agent-boundary-json"] + result = CliRunner().invoke(app, args) + assert result.exit_code == 0, result.output + boundary = json.loads(result.output) + assert boundary["comparison_status"] == "comparable", boundary + row, = boundary["rows"] + assert row["direction"] == "added" and ".mcp.json" in row["subject"] + + +@pytest.mark.parametrize(("command", "flag"), [ + ("codex exec", "--yolo"), + ("codex exec", "--dangerously-bypass-approvals-and-sandbox"), + ("claude -p", "--dangerously-skip-permissions"), + ("claude --print", "--dangerously-skip-permissions"), +]) +def test_removing_end_of_options_activates_bypass_and_widens(command, flag): + before = _workflow({"run": f"{command} -- {flag}"}) + after = _workflow({"run": f"{command} {flag}"}) + assert _launches(before)[0]["settings"] == [] + row, = _rows(before, after) + assert (row.direction, row.expands) == ("widened", True) + + +@pytest.mark.parametrize("command", ["codex exec", "claude -p"]) +def test_end_of_options_prompt_is_not_published_as_a_permission_setting(command): + workflow = _workflow({"run": f"{command} -- --add-dir private-prompt-canary"}) + launch, = _launches(workflow) + assert launch["settings"] == [] + assert "private-prompt-canary" not in json.dumps(_grant(workflow)) + + +@pytest.mark.parametrize("flag", ["-p", "--print"]) +def test_claude_print_flag_after_end_of_options_is_not_headless(flag): + workflow = _workflow({"run": f"claude -- {flag}"}) + assert _launches(workflow) == [] + assert len(_unread(workflow)) == 1 + + +@pytest.mark.parametrize("command", [ + "codex exec --yolo", "claude -p --dangerously-skip-permissions", +]) +def test_end_of_options_preserves_bypass_before_it(command): + before = _workflow({"run": command}) + after = _workflow({"run": f"{command} -- Review"}) + assert _launches(after)[0]["widening_rules"] + assert _rows(before, after) == [] diff --git a/tests/test_workflow_step_action_references.py b/tests/test_workflow_step_action_references.py index 51107dd1d..61eaf2827 100644 --- a/tests/test_workflow_step_action_references.py +++ b/tests/test_workflow_step_action_references.py @@ -394,6 +394,9 @@ def _legacy_baseline(tmp_path: Path, version: str): for grant in legacy["inventory"]["grants"]: if grant["kind"] == "workflow": del grant["step_actions"] + # Nor did they read agent launches or checkout refs (#823). + grant.pop("agent_launches", None) + grant.pop("checkout_refs", None) legacy["inventory_sha256"] = host_grants_sha256(legacy["inventory"]) return current, legacy @@ -415,7 +418,10 @@ def test_a_legacy_baseline_holding_a_workflow_does_not_assert_no_step_references drift = build_host_drift_payload(baseline=loaded, inventory=current, baseline_file=str(baseline_path)) assert drift["comparison_status"] == "incomparable" - assert drift["incomparable_reasons"] == ["baseline_workflow_step_actions_unavailable"] + assert drift["incomparable_reasons"] == [ + "baseline_workflow_agent_launches_unavailable", + "baseline_workflow_step_actions_unavailable", + ] assert drift["has_drift"] is None and drift["changes"] == [] assert baseline_path.read_text() == original @@ -915,6 +921,8 @@ def test_the_documented_migration_from_a_legacy_baseline_holding_a_workflow(tmp_ for grant in legacy["inventory"]["grants"]: if grant["kind"] == "workflow": grant.pop("step_actions", None) + grant.pop("agent_launches", None) + grant.pop("checkout_refs", None) legacy["inventory_sha256"] = host_grants_sha256(legacy["inventory"]) path = root / ".agents-shipgate" / "host-grants.json" path.parent.mkdir(parents=True, exist_ok=True) @@ -925,7 +933,10 @@ def test_the_documented_migration_from_a_legacy_baseline_holding_a_workflow(tmp_ drift = CliRunner().invoke(app, [*audit, "--drift", "--json"]) payload = json.loads(drift.stdout) assert payload["comparison_status"] == "incomparable" - assert payload["incomparable_reasons"] == ["baseline_workflow_step_actions_unavailable"] + assert payload["incomparable_reasons"] == [ + "baseline_workflow_agent_launches_unavailable", + "baseline_workflow_step_actions_unavailable", + ] assert payload["has_drift"] is None and payload["next_action"] is None assert CliRunner().invoke(app, [*audit, "--drift", "--fail-on-drift", "--json"]).exit_code == 20