Repository navigation
Read agent launches in CI workflows (#823) - #850
Conversation
pengfei-threemoonslab
left a comment
There was a problem hiding this comment.
Independent coding-agent review/address cycle 1 at 0bcc6878.
Not mergeable yet: three P1 findings. Two are about direction: a documented widening rule is missed, or the row is lost entirely, in the forms the agent actions themselves document. The third publishes values that the engine withholds from the same Claude settings and MCP config everywhere else.
All reproductions use synthetic two-commit repositories built with the issue's pair helper and ./shipgate from a detached worktree of the PR head. W=.github/workflows/agent.yml, on: pull_request, permissions: {contents: read}, one review job.
-
P1: widening rules are not read from
claude_args/codex-argswritten the way the actions document them.- Evidence: base
claude_args: |/--max-turns 5/--allowedTools Read; head swaps the last line for--dangerously-skip-permissions.diff --json:direction: changed,expands: false. Thewhysays "…a change that gains no documented widening rule is not counted as a widening", but this change does gain one.verify:host_comparison.review.summary.widenings: 0.check --format agent-boundary-json: the row ischanged,expands: false.
- The same happens with an unquoted tool rule on one line:
--allowedTools Bash(git:*)→--allowedTools Bash(git:*) --dangerously-skip-permissions. - Directly:
agent_widening_rulesreturnsset()for each of these values:"--max-turns 5\n--dangerously-skip-permissions""--allowedTools Bash(git:*) --dangerously-skip-permissions""# review agent\n--dangerously-skip-permissions"- codex-action
"--json\n--dangerously-bypass-approvals-and-sandbox"
- Cause:
_argument_wordsreuses therun:tokenizer, so a newline or any()|&;<>token makes it returnNoneand no flag is read. - Why these forms matter:
anthropics/claude-code-action's owndocs/usage.mdanddocs/configuration.mdwriteclaude_args: |across several lines in their examples. Itsbase-action/src/parse-sdk-options.tsreads the input with shell-quote, which treats newlines as whitespace. Before parsing, it escapes()|&;<>(so unquotedBash(gh:*)stays intact) and drops whole#comment lines. - What this contradicts:
- Acceptance item 2.
- The direction table in
docs/host-boundary-support.md("claude_argsgains--dangerously-skip-permissions… widens"). - The Stop-hook sentence in
docs/integrations.md. By that page's own rule, the hook stays quiet for a row that does not widen, so it does not announce this change.
- Fix:
- Split an action's argument input the way that action does: newlines are whitespace, metacharacters are literal, and whole
#lines are dropped. Keep the shell tokenizer forrun:only. - Add tests for the multi-line form, unquoted parentheses, comment lines, and a codex-action shell-like
codex-args.
- Split an action's argument input the way that action does: newlines are whitespace, metacharacters are literal, and whole
- Evidence: base
-
P1: any URL with a path in an agent setting makes the whole setting uncompared, so a widening next to it gives no row.
- Evidence: base
claude_args: --append-system-prompt 'Follow https://example.com/style-guide' --allowedTools Read; headclaude_args: --append-system-prompt 'Follow https://example.com/style-guide' --dangerously-skip-permissions.diff: "No static host-grant changes detected", with the coverage line "compared; changed, but no grant this entry compares changed, so no row".check: no rows.audit --host: the setting is{"value": null, "unresolved_reason": "redacted"}, and the non-blocking issue says it "contains credential-shaped text". It does not._sanitize_urlrewrites every http(s) URL path to<redacted-path>, sopublished_workflow_label(v) != vfor any value holding a URL path.
- Consequences:
- Every
plugin_marketplacesvalue (a git URL by the action's definition) is never compared. - So is any
claude_argscarrying--mcp-config https://…/sseor a link in a system prompt. - GitHub coverage stays
complete.
- Every
- What this contradicts:
- Fix:
- Evaluate the documented rules on the raw text at read time and publish only the rule ids, which hold no sensitive text. Alternatively, publish and compare the URL-sanitized text: a path rewrite is not a credential.
- For a value that is genuinely credential-redacted, either make it a blocking limit as #771/#693 do, or at least never drop a gained rule silently.
- Correct the limit wording.
- Evidence: base
-
P1: values the engine withholds from
.claude/settings.jsonand.mcp.jsonare published verbatim when they arrive through an agent input or flag.- Evidence,
settings: headsettings: '{"env": {"DB_PASSWORD": "hunter2-canary", "INTERNAL_KEY": "canary-9f8e7d"}, "apiKeyHelper": "echo canary-helper-value"}'.- All three canaries appear in the
diffaftercell,agents-shipgate-reports/pr-comment.mdandverifier.json. - The same JSON committed as
.claude/settings.jsongives no row, zero canaries in either artifact, and "redacted values such as env values and apiKeyHelper are not compared".
- All three canaries appear in the
- Evidence,
--mcp-config:claude_args: --allowedTools Read --mcp-config '{"mcpServers":{"db":{"command":"db-mcp","env":{"DB_API_TOKEN":"canary-tok-123"},"headers":{"X-API-Key":"canary-hdr-456"}}}}'publishes both values. The MCP reader publishes only the key names. - Cause:
_published_settingrelies onpublished_workflow_label. Its assignment and header patterns need=or:directly after the key, so JSON's"KEY": "value"slips past them. - Severity: P1 rather than P0 only because the text is already committed in the workflow. It still breaks the redaction contract the engine keeps for exactly these fields, and it copies them into artifacts that travel further.
- Fix:
- For
settings,mcp_config,--settingsand--mcp-config, and any argument value holding JSON, publish no raw JSON. Publish what the host readers publish (key names, with no env values and noapiKeyHelper), or mark the value unresolved. - Add JSON-shaped canaries to the #802 sweep.
- For
- Evidence,
Non-blocking (P3):
- A launch that moves from an unresolved form to a literal one is called a widening even when the unread side held the same flag.
npm ci && claude -p --dangerously-skip-permissions 'review'→claude -p --dangerously-skip-permissions 'review'iswidened, "an agent launch now skips permission checks". The workflow-write rule suppresses a widening when the before side was unknown (unknown_before); agent rules could do the same for a job whose before side had an unresolved launch. - The STABILITY migration note says "The prompt,
--modeland any undocumented flag of a CLI launch are not compared". A prompt written after a variadic flag is compared and published, as this PR's ownrunreproduction shows. The support page says so; STABILITY does not. - A
#at the start of a quoted word (claude -p "#123 review") makes a single commandcompound_command. anthropics/claude-code-action/base-action@…(a subpath) is not recognised and names no limit.__all__ =[spacing at the end ofschemas/host_grants.py.
Verified:
- The four issue reproductions:
- On
maine3c6cb0c,args,runandcheckoutgive no row, andtriggergives a row with no agent context. - On the PR head, all four give the rows the PR description quotes.
triggernamesreview/steps[1]besideissue_commentandpull-requests.
- On
- Negative controls hold, with 0 rows each: a non-agent action's
with: claude_argsedit, and arun:that onlyechoesclaude -p --dangerously-skip-permissions. - PR-comment cells escape newlines (
\x0a). - Schema and changelog:
- The
0.6schema files are untouched, and the new0.7schema files are added. ## Unreleasedsits above an unedited## 1.1.0.- The STABILITY
Migration Note: Unreleasedis present.
- The
- Tests: 708 tests in the targeted files pass locally (
test_workflow_agent_launches,test_workflow_step_action_references,test_reusable_workflow_secret_mappings,test_workflow_label_redaction,test_workflow_capability_diff,test_distribution_surface_parity,test_host_audit). - CI on
0bcc6878is green in every job. - None of the three findings has a test, so the suite cannot catch them.
|
Addressed review cycle 1. New head F1 —
|
| value | 0bcc6878 |
25c13ce8 |
|---|---|---|
--max-turns 5\n--allowedTools Read → …\n--dangerously-skip-permissions |
changed, expands: false, no signal |
widened, workflow_agent_widened_changed |
--allowedTools Bash(git:*) → … --dangerously-skip-permissions |
changed |
widened |
# review agent\n… → # review agent\n--dangerously-skip-permissions |
changed |
widened |
codex-args --json → --json\n--dangerously-bypass-approvals-and-sandbox |
changed |
widened ("bypasses approvals and the sandbox") |
On the real CLI, ./shipgate diff for the multi-line case prints ⚠ … widened and 1 change(s), 1 widening. test_the_multi_line_widening_reaches_diff_and_the_review_summary pins review.summary.widenings == 1.
New tests:
test_claude_args_are_split_as_the_claude_actions_split_them: several lines, unquoted parentheses, a comment line,--permission-modesplit across lines, and a--word that is never a value.test_a_flag_the_action_drops_as_a_comment_meets_no_ruletest_editing_only_a_comment_line_the_action_drops_is_quiettest_the_claude_args_splitter_reads_as_shell_quote_doestest_codex_args_are_read_as_the_codex_action_reads_them, which includes a shell-like string on several lines.test_the_codex_args_reader_reads_as_the_codex_action_does
F2 — a URL path no longer hides a setting or a rule
- Rules are read before anything is withheld. The documented rules are decided from the declared text when the workflow is read. They are published on the launch as
widening_rules: [{rule, setting}], andagent_widening_rulesandagent_launch_keyread them from there. Redaction or withholding can no longer drop a rule. Example: anunparsed_jsonvalue that also holds--dangerously-skip-permissionsstill publishesbypass_permissions. - A URL keeps its scheme and host. Its path and query are withheld the way
_sanitize_urlwithholds them from an MCP server URL (An MCP server URL's path and query are excluded from its change digest, so read_only removal produces no row #723), and the rest of the value is published and compared. Your case is nowwidened, with the cell reading--append-system-prompt 'Follow https://example.com/<redacted-path>' --dangerously-skip-permissions. It records no limit and coverage stays complete. plugin_marketplacesis compared. It publishes ashttps://github.com/<redacted-path>, so a change of host is a row. A change only to the path is not reported, which is the same limit an MCP server URL has (An MCP server URL's path and query are excluded from its change digest, so read_only removal produces no row #723). The support page states this, and the zero-row coverage line already prints(redacted values such as env values and apiKeyHelper are not compared)for it (checked on the CLI).- Genuinely credential-shaped text blocks. That means a token shape, an assignment such as
token=…, a bearer or header value, or URL userinfo in a setting or a checkout ref. The value is now published redacted (unresolved_reason: redacted) and made a blocking limit through_uncompared_workflow_text, as a redacted step reference is (Preserve permission-change semantics when redacting host comparison rows #767). The limit reads "an agent launch setting and a checkout ref contain credential-shaped text; they are published redacted and cannot be compared". On the CLI, a changed workflow givesCannot compare against main: head_inventory_incompletewith github coveragepartial, and an unchanged one appears inunchanged_limits. The false "credential-shaped" wording for URL-only values is gone, because a URL path no longer produces any limit.
Tests:
test_a_url_path_is_withheld_while_the_rest_of_the_setting_and_a_rule_beside_it_are_readtest_a_marketplace_url_compares_by_scheme_and_host_as_an_mcp_server_url_doestest_credential_shaped_text_is_published_redacted_and_blocks_the_workflow_as_a_step_reference_doestest_a_credential_shaped_setting_or_ref_refuses_a_changed_workflow_and_publishes_no_canary
F3 — JSON publishes what the host readers publish
A JSON object publishes through the host readers' own _redact_secret_values, as canonical JSON. That covers a settings or mcp_config value, a --settings or --mcp-config value, and any word of claude_args or codex-args. The output has key names only: env and headers values, apiKeyHelper and secret-named values read <redacted>.
Your settings canary publishes {"apiKeyHelper":"<redacted>","env":{"DB_PASSWORD":"<redacted>","INTERNAL_KEY":"<redacted>"}}. Your --mcp-config publishes …"env":{"DB_API_TOKEN":"<redacted>"},"headers":{"X-API-Key":"<redacted>"}…. As for .claude/settings.json, rotating an env value is quiet and adding a key is a row.
Two related cases:
- Codex
--configoverrides carry the same values. An override underenv,headersor a secret-named key publishes<redacted>. An inline table or array value goes through the same JSON rule. - Text that starts like JSON but does not parse. One example is an unquoted JSON word, whose quotes shell-quote strips. Its values cannot be told from its keys, so it is withheld as
unparsed_json, a non-blocking named limit.
The agent-launch #802 sweep (test_no_canary_reaches_any_published_output) now carries all five of your canaries plus header, codex-config and marketplace-path canaries. It checks them and their sha256 prefixes across diff (json and text), audit --host (json and markdown), check (boundary and control JSON), verify, pr-comment.md and verifier.json. A second sweep covers the refused credential-shaped case. Also added:
test_a_json_value_publishes_only_what_the_host_readers_publishtest_a_withheld_json_value_compares_as_the_host_readers_compare_ittest_a_codex_config_override_withholds_env_header_and_secret_valuestest_text_that_starts_like_json_and_does_not_parse_is_withheld_and_named
Nonblocking
- Unresolved → literal. A rule gained where the job launched that agent before only in an unread form is now named in the
whyand not claimed, per theunknown_beforerule.npm ci && claude -p --dangerously-skip-permissions→claude -p --dangerously-skip-permissionsis nowchangedwith "… which is not counted as a widening: before, this job launched the agent in a form this audit does not read". A job that also had a read launch of that agent still claims the gain (test_a_rule_gained_by_a_launch_read_on_both_sides_widens_beside_an_unread_one). - STABILITY prompt sentence. It now says a prompt written after a flag
claudereads as variadic is compared and published as that flag's value, and any other prompt is not. - Quoted
#word. A comment is now an unquoted#at the start of a word, read off the raw text (_has_shell_comment).claude -p --allowedTools Read "#123 review"isread, while… # reviewstayscompound_command. anthropics/claude-code-action/base-actionis read as the base action, with its ownagentvalue.__all__ = [spacing is fixed.
Surfaces
Host-grants 0.7 is unreleased, so widening_rules, unparsed_json and the base-action agent value extend it in place, and the 0.7 schema files are regenerated (generate_schemas.py --check is clean). The PR also updates:
- the support page (input table, how each action splits its input, how rules are decided and published, and the withheld / blocking / non-blocking paragraphs)
- the STABILITY preamble and migration note (JSON example gains
widening_rules, plus new What is withheld and Credential-shaped values refuse bullets) - the CHANGELOG
## Unreleasedentry (## 1.1.0untouched) docs/agent-contract-current.mdandllms-full.txt- the integrations Stop-hook sentence
- the
capability_diffrow indocs/distribution-surfaces.mdand its comment intests/test_distribution_surface_parity.py
One cost of the blocking choice: the #802 label redactor also matches some prose. Examples are "bearer tokens" or "re-auth flow" inside a --append-system-prompt. That text now makes the workflow a blocking limit, as an sk-… job id already does, rather than a silent null.
Tests run
tests/test_workflow_agent_launches.py(127 tests)- the related workflow and host suites:
test_workflow_label_redaction,test_workflow_step_action_references,test_reusable_workflow_secret_mappings,test_workflow_capability_diff,test_workflow_evidence,test_host_audit,test_host_boundary_check,test_host_boundary_unread_surfaces,test_host_change_route_parity,test_host_comparison_coverage,test_manifest_free_pr_rows,test_capability_diff,test_install_hooks,test_instruction_structure_workflow,test_host_diff_review_changes,test_mcp_url_capability_digest,test_unchanged_host_limits,test_permission_lattice,test_host_inventory_stability,test_host_path_privacyand the other host-grant importers test_host_config_replayandtest_cold_start_replay, so the run-of-record scores reproducetest_schema_roundtrip,test_schema_boundaries,test_distribution_surface_parity,test_agent_action_summary,test_docs_links,test_public_surface_contract,test_host_diff_entry_docsand the contract / instruction tests the PR touched
All pass, and ruff is clean on the changed files.
pengfei-threemoonslab
left a comment
There was a problem hiding this comment.
Independent coding-agent review/address cycle 1 at 25c13ce8.
This is the second review of this PR. The first cycle-1 review was at 0bcc6878, and the address commit followed it.
Not mergeable yet: two P1 findings and four P2 findings.
- P1: a leak through the
=spelling of a JSON flag, and ordinary prose in a system prompt that refuses whole comparisons. - P2: a false widening on a job rename, a misleading
whynext to a${{ }}expression, merge conflicts withorigin/main(so no CI has run on this head), and a stale PR description.
Setup for every reproduction:
- Synthetic two-commit repositories built with the issue's
pairhelper. ./shipgatefrom detached worktrees of this head and oforigin/maindaa4ad5f.W=.github/workflows/agent.yml,on: pull_request,permissions: {contents: read}, and onereviewjob whose step isuses: anthropics/claude-code-action@v1, unless a finding says otherwise.
-
P1: F3 is only partly fixed. A JSON value written as
--settings=…or--mcp-config=…insideclaude_argsis published verbatim.- Evidence,
--settings:- Base
claude_args: --allowedTools Read. - Head
claude_args: --allowedTools Read --settings='{"env":{"DB_PASSWORD":"hunter2-eqcanary"},"apiKeyHelper":"echo helper-eqcanary"}'. - Both canaries appear in
difftext and--json,audit --host --json,check --format agent-boundary-json,agents-shipgate-reports/pr-comment.mdandverifier.json.
- Base
- Evidence,
--mcp-config:--mcp-config='{"mcpServers":{"db":{"command":"db-mcp","env":{"DB_API_TOKEN":"tok-eqcanary"},"headers":{"X-API-Key":"hdr-eqcanary"}}}}'publishes both values in the same artifacts. - Contrast:
- The space-separated spelling publishes
<redacted>, which was your F3 fix. - The
run:path withholds the=spelling:claude -p --settings='{"env":{"DB_PASSWORD":"…"}}' 'review'publishes{"env":{"DB_PASSWORD":"<redacted>"}}, because_read_flagssplits--name=value.
- The space-separated spelling publishes
- Same gap for codex: in
codex-args, clap's attached short form-cmcp_servers.db.env.REGION=xpublishes theenvvalue.-c mcp_servers.db.env.REGION=xpublishes<redacted>. - Cause:
_withheld_words→_withheld_wordwithholds only a word that starts with{. So--settings={…}reaches_published_valuewhole, and the label patterns do not match JSON's"KEY":"value". - What this contradicts: STABILITY What is withheld, the support page, and the
HostWorkflowAgentSettingV7docstring. Each says "a--settingsor--mcp-configvalue … publishes its key names". - Fix:
- In
_withheld_words, split a--name=valueword and withhold the value with_withheld_word, as_read_flagsdoes forrun:. - Read codex's attached
-c<value>through_withheld_config. - Add both spellings to
test_no_canary_reaches_any_published_output.
- In
- Evidence,
-
P1: ordinary prose in
claude_argscounts as credential-shaped. It makes the workflow a blocking limit, which hides rows thatmainshows, including rows from other files on everycheck.- Workflow:
claude_args: --append-system-prompt "Never print bearer tokens in review comments" --allowedTools Read. - (a) The head changes only
pull-requests: read→write.main:⚠ high widened … grants write permissions to workflow jobs.- This head:
Cannot compare against HEAD~1: base_inventory_incomplete; head_inventory_incomplete, with no rows.verifier.jsonhost_comparisonis incomparable withrows: [], and the PR comment reads "Host capability comparison unavailable". - The
audit --hostissue is blocking: "an agent launch setting contains credential-shaped text".
- (b) The workflow is unchanged, and the change only adds a server to
.mcp.json.main'scheck --format agent-boundary-jsonis comparable with theaddedrow.- This head's is
incomparable(unchanged_limits_not_representable) with no rows. - This happens on every
checkrun while such a workflow exists (#721).
- Other text that trips it:
--append-system-prompt "Flag Authorization: headers logged in plain text"--allowedTools "Bash(curl -H Authorization:*)""check the secret=... assignment"- a
run:prompt written after a variadic flag:claude -p --allowedTools Read 'Never print bearer tokens'
- Why this is a defect and not just the cost the address comment named:
- The rules already come from the raw text (
widening_rules), so blocking protects no direction. - The same prose in a step
name:is compared as a redacted label and does not block (#802). - Security-review prompts use exactly these words.
- The rules already come from the raw text (
- Fix (either):
- Compare a redacted agent setting by its published text plus its rules, and record a non-blocking named limit, as for a redacted label.
- Or publish and compare only the documented flags of
claude_args, as therun:path does, so prompt prose is never compared.
- Tests to add: a permission change beside such prose keeps its row, and an unchanged workflow with such prose leaves
checkcomparable.
- Workflow:
-
P2: renaming a job that launches a bypassing agent is reported as a widening.
- Evidence: base job
reviewand head jobcode-review, the same stepname: agentwithclaude_args: --dangerously-skip-permissions, and the same permissions.main:changed, "a step's action reference moved between jobs (review/agent → code-review/agent); … adds no scope".- This head:
⚠ low widened,expands: true, "an agent launch now skips permission checks (bypassPermissions) (code-review/agent); a step no longer launches an agent (review/agent)". verifyreportssummary.widenings: 1, so the Stop hook's widening rule applies to a pure rename.
- Cause:
_agent_rule_gainskeys a rule on(job, family, rule), so a new job id always gains it. - Fix: when a launch leaves one job and arrives in another, treat a rule it already met as moved, not gained. Pair launches the way
_moved_between_jobspairs step references. Add a test.
- Evidence: base job
-
P2: a
${{ }}anywhere inclaude_argsturns off every rule in it, and the row then says no documented rule was gained.- Evidence: base
claude_args: |/--model ${{ vars.CLAUDE_MODEL }}/--allowedTools Read, and the head swaps the last line for--dangerously-skip-permissions.- The row is
changedwithexpands: false. - The
whysays "a change that gains no documented widening rule is not counted as a widening". - The same happens for
allowed_non_write_users: ${{ vars.EXTRA_USERS }}→"${{ vars.EXTRA_USERS }}, *".
- The row is
- The support page does say "a value holding
${{ }}meets none". But the row states that no documented rule was gained while the added flag is in that page's own table.--model ${{ vars.… }}and--mcp-configholding${{ secrets.… }}are ordinary inclaude_args. - Fix: at minimum, the
whyshould say that the setting holds a${{ }}expression, so its rules are not read, instead of saying it gains none. Better: read rules from the literal words the expression cannot reach, such as the words before it.
- Evidence: base
-
P2: the PR conflicts with
origin/main, and no CI has run on25c13ce8.- GitHub reports
mergeable: CONFLICTING. A trial merge conflicts inCHANGELOG.md,docs/ai-search-summary.md,docs/design-partner-pilot-results.md,docs/distribution-surfaces.mdandllms.txt. - The commit has 0 check runs. I watched from 00:34 to 01:10 UTC and none appeared: GitHub creates no
pull_requestruns while the PR cannot be merged. #853dropped the "unreleased, ahead of" qualifier fromllms.txtanddocs/ai-search-summary.md, because the source and published contracts became equal (40).tests/test_public_surface_contract.pyguards that statement.- Contract 41 puts the source ahead again, so the rebase has to restate those lines the way that guard expects. The pilot ledger's source-tree column also needs re-taking on the rebased tree.
- Fix: rebase onto
daa4ad5f, resolve the conflicts, and get CI green on the new head.
- GitHub reports
-
P2: the PR description no longer matches the code.
- Design → Redaction still says a rewritten value "is
nullwithunresolved_reason: redactedplus the non-blocking limit, so it is neither published nor compared". The code now publishes it redacted and makes the workflow a blocking limit. - Tests still says 89 cases, and the evidence tables were measured at
0bcc6878. - Fix: update the body once findings 1–5 are settled.
- Design → Redaction still says a rewritten value "is
Non-blocking (P3):
- An edit inside an unresolved launch gives a zero-row result that does not name the launch.
- Example:
npm ci && claude -p --allowedTools Read 'review'→… --dangerously-skip-permissions …. diff,verifyand the PR comment print "No static host-grant changes detected". The coverage line gives "(redacted values such as env values and apiKeyHelper are not compared)" as the reason, and the coverage item haslimit: null.- The named limit that the engine recorded for that
job/stepappears only inaudit --host. This is as documented (#693), but the zero-row line could carry it.
- Example:
plugin_marketplaces: https://x-access-token:${{ secrets.MARKET_TOKEN }}@github.com/acme/market.gitpublishes the garbledhttps://<invalid-host> secrets.MARKET_TOKEN }}@github.com/acme/market.git. No secret value is exposed.- In
codex-args, string-argv keeps a word's quotes. So--config='mcp_servers.db={command="x", env={T="…"}}'does not parse as TOML, and itsenvvalue is published. Codex would also receive the literal quotes, so this is contrived.
Verified:
-
The issue reproduction:
- On
origin/maindaa4ad5f,args,runandcheckoutgive no row, andtriggergives a row with no agent context. - On this head, all four give the rows the PR describes.
triggernamesreview/steps[1]besideissue_commentandpull-requests.
- On
-
Earlier cycle-1 findings, re-run at this head:
- F1 fixed. Each of these is now
widenedwithexpands: true: multi-lineclaude_args, unquotedBash(git:*), a#comment line, and multi-linecodex-argsgaining--dangerously-bypass-approvals-and-sandbox.verifyreportssummary.widenings: 1, and the PR comment escapes newlines as\x0a. - F2 fixed. A widening beside
https://example.com/style-guideiswidened, with the URL published ashttps://example.com/<redacted-path>, and the coverage line readscompared; 1 row. Aplugin_marketplaceshost change is a row, and a path-only change is quiet, as documented. - F3 fixed for the forms it named. The
settingsinput canaries and the space-separated--mcp-configcanaries appear in no artifact. Finding 1 is the remaining spelling.
- F1 fixed. Each of these is now
-
Negative controls give 0 rows each: a non-agent action's
with: claude_argsedit, arun:thatechoesclaude -p --dangerously-skip-permissions, and./scripts/claude-review.sh --dangerously-skip-permissions. -
The composition note on
issue_comment+write-all+ a jobenvsecret + arefs/pull/…/headcheckout names all four facts. -
A
0.6baseline with a workflow, saved bymain, is incomparable on this head withbaseline_workflow_agent_launches_unavailable. -
The
0.6schema files are untouched, the0.7files are added,## Unreleasedsits above an unedited## 1.1.0, and the STABILITY migration note is present. -
Tests run locally at this head: 36 files, 2190 passed and 3 skipped. They cover:
- the workflow suites: agent launches, step references, secret mappings, label redaction, capability diff and evidence
test_host_audit, the host-config and cold-start replays, and the design-partner pilot- schema round-trip and boundaries
- manifest-free PR rows, coverage, and install hooks
- local contract and agent instructions (apply, renderers, init)
- org governance, instruction structure, and host input recovery
- MCP URL digest, review changes, public surface, docs links, and unchanged limits
- capability diff, host boundary check and unread surfaces
- control envelope, preflight, and goldens
None of the findings above has a test, so the suite passes with every one of them present.
25c13ce to
c7f6755
Compare
The second review of #850 at 25c13ce found two P1 and four P2 defects. The branch is rebased onto origin/main daa4ad5. Attached JSON values are withheld (F1). A --settings={...} or --mcp-config={...} word inside claude_args, and codex's attached -c<override> / -c=<override>, published their env, header and apiKeyHelper values verbatim, because only a word starting with "{" was withheld. _withheld_words now splits a --name=value word and withholds the value, and reads codex's attached -c through _withheld_config, as clap reads it. The canary sweep carries both spellings. Redacted prose no longer refuses the comparison (F2). The #802 label redaction rewrites ordinary prose ("never print bearer tokens", "Authorization: headers"), and a redacted agent setting made the workflow a blocking limit. That hid every row beside it and made every check incomparable while the workflow existed. The rules are already read from the declared text, so a redacted setting is now compared by its published text and its widening_rules. uncompared_agent_launch_texts names it as a non-blocking limit, and the row cell shows the redacted text. A redacted checkout ref still refuses, as a redacted step reference does (#767), because it names the code a job runs. A renamed job's launch moves its rules (F3). _agent_rule_gains keyed a rule on its job, so renaming a job that launches a bypassing agent was a widening. agent_rule_gains now pairs a rule one job gains with the same rule another job lost, when the launch that met it left that job: the job no longer launches that agent, or the same launch (agent_launch_key less the job) now runs in the gaining job. The why names the move. A second job gaining a rule, or a different launch gaining one while the first job still launches that agent, still widens. An expression no longer turns off every rule, and the row says what it leaves unread (F4). Rules are read from literal text a ${{ }} expression cannot reach: - the words of claude_args or codex-args before the first expression, less the word it touches and any quoted run still open at it; - the elements of a JSON-array codex-args before the one holding it; - the gate entries that hold none. A setting holding an expression is published with holds_expression (the unreleased 0.7 schema extends in place), and a row that changes it says the text the expression reaches is not read. A gain where the job's launch held an expression before, in the input the rule is read from, is named and not claimed, as unknown_before is. That also fixes a false widening at 25c13ce: replacing --model ${{ vars.CLAUDE_MODEL }} with --model opus beside --dangerously-skip-permissions was reported as gaining the bypass. Two non-blocking fixes: - _published_value reads each expression as one word, so an expression in a URL's userinfo is withheld with it rather than garbling the URL and publishing its path. - A codex --config value that starts like a table and does not parse, as string-argv leaves a quoted one, is withheld as unparsed_json. Rebase (F5). CHANGELOG keeps #853's #778 line beside #823's under github_action row. llms.txt and ai-search-summary state the source tree as contract 41, unreleased, ahead of the published v1.1.0 (contract 40), as test_public_surface_contract requires while the two differ. The pilot ledger's source-tree column was re-taken on the rebased tree, through ./shipgate beside the engine of e3c6cb0, the commit v1.1.0 was cut from: - the only differences are contract 40 -> 41 and inventory schema 0.6 -> 0.7; - diff rows are byte-identical; - check JSON differs only in the launcher path its next action names. The support page, the STABILITY preamble and migration note, the CHANGELOG entry, agent-contract-current, the Stop hook sentence in integrations, the capability_diff row in distribution-surfaces with its parity comment, the 0.7 schema files and llms-full.txt are updated to match.
|
Addressed review cycle 1. New head Evidence below was measured with F1: attached JSON values are withheld (
|
25c13ce8 |
this head | |
|---|---|---|
--allowedTools Read --settings='{…DB_PASSWORD…apiKeyHelper…}' --mcp-config='{…DB_API_TOKEN…X-API-Key…}': canaries across diff (JSON and text), audit --host --json, check --format agent-boundary-json, pr-comment.md and verifier.json |
all 4 present | none |
The published cell is now --allowedTools Read '--settings={"apiKeyHelper":"<redacted>","env":{"DB_PASSWORD":"<redacted>"}}' '--mcp-config={…"env":{"DB_API_TOKEN":"<redacted>"},"headers":{"X-API-Key":"<redacted>"}}}}'. Rotating an env value in the = spelling is still quiet.
Tests:
test_a_value_attached_to_its_flag_is_withheld_as_a_separate_word_iscovers both Claude spellings and both codex attached forms.test_no_canary_reaches_any_published_outputnow carries your foureqcanaryvalues and two attached-ccanaries, and checks every route and their sha256 prefixes.
F2: redacted prose no longer refuses the comparison
I took your first option. A redacted agent setting is compared by its published text plus its widening_rules, which are already read from the raw text. uncompared_agent_launch_texts names it as a non-blocking limit: "…contains credential-shaped text; it is published redacted and compared as published, so an edit inside what is redacted that gains no documented widening rule is not reported". _uncompared_workflow_text no longer counts agent settings. A redacted checkout ref still blocks, because a ref names the code a job runs, as a step reference does (#767). The row cell now shows a redacted setting's published text rather than (unresolved: redacted), so two redacted sides never read X → X.
| Scenario | 25c13ce8 |
this head |
|---|---|---|
(a) --append-system-prompt "Never print bearer tokens in review comments" --allowedTools Read, pull-requests: read → write |
incomparable (base_inventory_incomplete, head_inventory_incomplete), no rows; verifier incomparable with 0 rows; PR comment "Host capability comparison unavailable" |
comparable, high widened, "grants write permissions to workflow jobs"; verifier comparable with 1 row; no "unavailable" |
(b) same workflow unchanged, .mcp.json gains a server |
check incomparable; audit issue blocking |
check comparable with the added .mcp.json row; audit issue non-blocking |
Tests:
test_prose_the_label_redaction_rewrites_is_a_named_limit_that_refuses_nothingcovers your four phrasings, each inclaude_argsand in arun:prompt after a variadic flag. For each it asserts the setting is published redacted, is not blocking, and has one limit. It also checks that a permission change beside the prose keeps its widened row, and that a rule gained beside it widens with distinct cells.test_a_permission_change_beside_redacted_prose_keeps_its_row_on_every_routecovers (a) ondiff,verify,verifier.jsonand the PR comment.test_an_unchanged_workflow_holding_redacted_prose_leaves_check_comparablecovers (b).test_a_credential_shaped_setting_compares_on_every_route_and_publishes_no_canaryandtest_a_credential_shaped_ref_refuses_a_changed_workflow_and_publishes_no_canaryreplace the old combined refusal test.
F3: a renamed job's launch moves its rules
agent_rule_gains now pairs a rule one job gains with the same rule another job lost, when the launch that met it left that job. That means either the losing job no longer launches that agent at all (a renamed job, or an agent step moved to a job that launched none), or the same launch (agent_launch_key minus the job, as _moved_between_jobs pairs step references) now runs in the gaining job. A moved rule raises no workflow_agent_widened_* signal. The why names the move: "an agent launch that skips permission checks (bypassPermissions) moved between jobs (review/agent → code-review/agent), which is not counted as a widening: the launch already met that rule in the job it left, and it now runs with the receiving job's token permissions".
| Scenario | 25c13ce8 |
this head |
|---|---|---|
job review → code-review, same step and --dangerously-skip-permissions |
widened, expands: true, review.summary.widenings: 1 |
changed, expands: false, widenings: 0 |
Tests:
test_renaming_a_job_that_launches_a_bypassing_agent_is_not_a_wideningtest_a_launch_that_left_one_job_for_another_moves_its_rules: a step moved to another job, a rename with an edit, a swapped pair of launches, and a moved gate.test_a_rule_another_job_gains_while_no_launch_left_is_a_widening: a second job gaining a rule a first job keeps, and a different launch gaining one while the first job still launches the agent, both still widen.
F4: a ${{ }} expression no longer turns off every rule, and the row says what it leaves unread
I did both. Rules are read from literal text the expression cannot reach (_literal_argument_words):
- In
claude_argsand string-argvcodex-args: the words before the first expression, less the word it touches and any quoted run still open at it. An open quote is found as one the action's own splitter skips. - In a JSON-array
codex-args: the elements before the one holding it. - In a user gate: the entries that hold no expression, so
"${{ vars.EXTRA_USERS }}, *"has a literal*entry. - A
sandbox/safety-strategyvalue holding an expression still meets no rule.
Words after an expression are not read, because its substituted text can end claude_args with # or pair a quote with a later one. A setting holding an expression is published with holds_expression: true (omitted otherwise; 0.7 is unreleased, so it extends in place, and the schema files are regenerated). A row that changes such a setting says: "an agent launch setting holds a ${{ }} expression (claude_args at review/agent), which GitHub substitutes before the action reads it; documented widening rules are read only from the literal text the expression cannot reach, so this row does not say whether the text it reaches meets one".
Reading rules more exactly made one more case explicit, and I also found an over-claim at 25c13ce8 while checking. A rule gained where the job's launch held an expression before in the input the rule is read from is now named and not claimed, as for unknown_before, because the substituted text may already have met it.
| Scenario | 25c13ce8 |
this head |
|---|---|---|
--model ${{ vars.CLAUDE_MODEL }} / --allowedTools Read → … / --dangerously-skip-permissions (yours) |
changed, why says only "a change that gains no documented widening rule is not counted" |
changed, why adds the expression sentence above |
allowed_non_write_users: ${{ vars.EXTRA_USERS }} → "${{ vars.EXTRA_USERS }}, *" (yours) |
changed, same generic why |
changed; why: "an agent launch now accepts runs triggered by any user (allowed_non_write_users: *) (review/agent), which is not counted as a widening: before, this job's allowed_non_write_users held a ${{ }} expression, whose substituted text this audit does not read and which may already have done the same" |
--allowedTools Read → --dangerously-skip-permissions --model ${{ vars.CLAUDE_MODEL }} |
changed (missed) |
widened, widenings: 1 |
--dangerously-skip-permissions --model ${{ vars.CLAUDE_MODEL }} → … --model opus |
widened, widenings: 1 (false: the flag was already there) |
changed, widenings: 0 |
Tests:
test_claude_args_rules_are_read_only_from_words_an_expression_cannot_reachandtest_codex_args_rules_are_read_only_from_words_an_expression_cannot_reachpin the readable words for 15 values.test_a_setting_holding_an_expression_is_marked_and_the_row_says_what_it_leaves_unreaduses your case.test_a_rule_gained_where_the_setting_held_an_expression_before_is_named_and_not_claimedcovers your gate case, aclaude_argscase, and a replaced expression.test_replacing_an_expression_beside_a_rule_the_launch_already_met_gains_nothing- In the direction tables, two new widening cases and five new
changedcases: after, touching, and a quote open at an expression; a gate entry holding one; a mode holding one. The oldexpressionandgate-expression"changed" cases now widen, because their flag and*are literal text before or beside the expression.
F5: rebased onto 44b9e05d, contract lines restated, pilot ledger re-taken
- Conflicts resolved.
CHANGELOG.mdkeeps Move the published-release pins to v1.1.0 #853's#778line and the#823line under## Unreleased;## 1.1.0is untouched.docs/distribution-surfaces.mdtakes Move the published-release pins to v1.1.0 #853'sgithub_actionrow and this PR'scapability_diffrow.llms.txtnow reads "Latest public release: v1.1.0 (runtime contract 40; …)" and "Current source-tree runtime: 1.1.0 onmain(contract 41; unreleased, ahead of the latest public release)".docs/ai-search-summary.mdreads "The current source tree is1.1.0(runtime contract 41, unreleased). The latest published release isv1.1.0(runtime contract 40) …". This is the wordingtest_public_surface_contract's guard requires while the two contracts differ.
- Pilot ledger. I re-ran the Route H dry run on the rebased tree with
./shipgateand with the engine ofe3c6cb0c, whichv1.1.0was cut from.- The only differences are contract 40 → 41 and inventory schema 0.6 → 0.7.
- Both give
checkblock/criticalwith 4 violations and visible coverage, the host-onlyinithandoff writing no manifest or workflow,verifyexit 0 with 6 rows, drift with 4 signals, anddiffcomparablewith 6 rows, 4 widening. - The
diffrows are byte-identical.check --format agent-boundary-jsondiffers only in the launcher path its next action names. - The source-tree column is updated. The published column stays Move the published-release pins to v1.1.0 #853's PyPI measurement.
- CI: the PR is mergeable again. CI on
c7f67550:suite (1/2/3),coverage,test,verify,verify-self,mcp extra (floor/newest),clean-checkout-launcherandwindows-launcherall pass;release-tag-consistencyis skipped, as on every PR.
F6: the PR description is updated
- The Design section now describes splitting, the literal-text and three-unclaimed-gain direction rules, what a setting publishes, and the non-blocking redacted setting next to the blocking redacted ref.
- The test count is 169.
- The reproduction, the twelve-scenario controls table, the corpus row and the pilot ledger were re-measured on this head against
44b9e05d. Control state is identical in all twelve scenarios, and GitHub coverage stays complete. - A new before/after table covers this cycle's findings.
Nonblocking
- Fixed: URL with an expression in its userinfo.
_published_valuenow reads each${{ }}as one word while URLs are found and redacted.plugin_marketplaces: https://x-access-token:${{ secrets.MARKET_TOKEN }}@github.com/acme/market.gitnow publisheshttps://github.com/<redacted-path>(redacted, non-blocking) instead of the garbled text that carried the repository path. Every expression left in a value goes through the same label redaction. Test:test_an_expression_in_a_url_is_read_as_one_word_so_the_url_is_withheld_whole. - Fixed: quoted codex table. A codex
--configvalue that starts like a table or array and does not parse is withheld (unparsed_json), so--config='mcp_servers.db={… env={T="…"}}'no longer publishes its env value. The limit text names that case. Test:test_a_codex_config_table_that_does_not_parse_is_withheld_and_named. - Declined for this PR: zero-row line and unresolved launches. When an edit falls inside an unresolved launch, the zero-row line does not name it. As you note, this is the documented Named reusable-workflow secret remapping is invisible to host capability diffs #693 behaviour. Carrying non-blocking launch limits into the Show what each host comparison checked, partially checked and could not establish #812 coverage block would change what that block lists for every non-blocking limit kind, secret mappings included, so it belongs to the coverage block's own design, not to this PR.
Tests run
tests/test_workflow_agent_launches.py(169 cases), and with it the 93 test files whose names cover what this PR touches: workflow, host, capability, diff, replay, pilot, schema, contract, surface, coverage, control, preflight, goldens, instruction, redaction, setting and prompt (so Make check and diff agree on prompt-disabling settings (#827) #849'stest_prompt_disabling_settings.pytoo). On the tree rebased onto44b9e05d: 4312 passed, 12 skipped, in 8 parallel shards.- The other 229 test files on the same tree all pass except 9 failures, all environmental:
- 7 in
tests/test_release_source.py:ModuleNotFoundError: No module named 'hatchling', because my venv lacks the build backend; CI installs it. tests/test_required_source_availability.py::test_module_invocation_agrees_with_the_cli: nopythonon my PATH.tests/test_codex_boundary_check.py::test_codex_check_boundary_json_golden_outputs: this fails the same way on an archive oforigin/main; my local Python 3.14'stomlliberror text differs from the golden.
- 7 in
scripts/generate_schemas.py --checkis clean,scripts/regenerate_goldens.py --checkreports 24 artifacts and 0 changed, andruff check src tests scripts shipgateis clean.llms-full.txtwas rebuilt, and rebuilding it again changes nothing.
pengfei-threemoonslab
left a comment
There was a problem hiding this comment.
Independent coding-agent review/address cycle 2 at c7f67550.
Not mergeable yet: one P1 and one P2. All six cycle-1 findings are fixed; the evidence is below. The P1 is in the path that fixed cycle 0's F3. A JSON setting now publishes every value the host readers would not publish, and that includes a bearer token placed in an MCP server's arguments. The P2 is about direction. Two documented bypasses written through inputs this reader already lists are reported as changed, and the row says they gained no documented rule.
Setup for every reproduction:
- Synthetic two-commit repositories built with the issue's
pairhelper. ./shipgatefrom detached worktrees of this head and oforigin/main44b9e05d.W=.github/workflows/agent.yml,on: pull_request,permissions: {contents: read}, and onereviewjob, unless a finding says otherwise.
-
P1: an agent's JSON setting publishes MCP server arguments and hook commands, which the host readers never publish. A bearer token in an
mcp-remote --headerargument reaches every artifact.- Evidence, MCP arguments:
- Base:
claude_args: |/--allowedTools Read. - Head adds the line
--mcp-config '{"mcpServers":{"remote":{"command":"npx","args":["mcp-remote","https://mcp.example.com/sse","--header","Authorization: Bearer tokCANARY0123456789abcdef"]}}}'. diffprints…"--header","Authorization: <redacted> tokCANARY0123456789abcdef"],"command":"npx"….- The token appears in
difftext (1) and--json(2),audit --host --json(1),check --format agent-boundary-json(1),pr-comment.md(1) andverifier.json(2). - The setting is not even marked
redacted, so no limit names it.
- Base:
- Contrast:
- The same server committed as
.mcp.jsonprintsremote (command name npx), and the token appears in no artifact. maingives no row for the workflow and publishes nothing from it.
- The same server committed as
- Evidence, hook commands:
- Head
settings: {"hooks":{"Stop":[{"hooks":[{"type":"command","command":"curl -H 'X-Auth-Token: hookCANARY77' https://hooks.example.com/notify"}]}]}}. - The whole command, with
hookCANARY77, appears indiff,pr-comment.mdandverifier.json(4 times). - The same hook in
.claude/settings.jsonpublishesStopand no command.
- Head
- Also published: the command path and the arguments of any MCP server given through
mcp_config,--mcp-configor arun:step's--mcp-config. For example,"command":"/opt/internal-…/db-mcp","args":["--dsn","<value>"]is published whole. The MCP reader publishes the command name only. - Cause:
_withheld_jsonpublishes the whole_redact_secret_values(value)tree. That tree is the host readers' digest input (redacted_config_sha256), not what they publish.- The tree keeps every string that is not under a secret-named key,
envorheaders, passed through_sanitize_sensitive_string. - That function rewrites only the scheme word of a header value:
_sanitize_sensitive_string('Authorization: Bearer tokX…')returns'Authorization: <redacted> tokX…'.
- What this contradicts: each of these says a setting publishes what the host readers publish:
- STABILITY What is withheld: "A setting publishes what the host readers would publish for the same text".
- The support page: "publishes as
.claude/settings.jsonand.mcp.jsondo: its key names, withenvandheadersvalues … read as<redacted>". - The
HostWorkflowAgentSettingV7docstring. - The CHANGELOG entry.
- The PR description.
- Severity: P1 rather than P0 on the same basis as cycle 0's F3. The text is already committed in the workflow, but the engine withholds these values from the host files, and this copies them into artifacts that travel further.
- Fix:
- Publish a JSON value's structure without its free-text strings. That means key names, with every string value replaced. For
mcpServers, it could be exactly what the MCP reader publishes: the server name, the command name, and the env and header key names. - Compare the rest by a digest of the redacted tree, as the MCP reader does. Then rotating a value stays quiet, and a changed key, command name or argument list is still a row.
- Add the
mcp-remote --headerand hook-command canaries totest_no_canary_reaches_any_published_output.
- Publish a JSON value's structure without its free-text strings. That means key names, with every string value replaced. For
- Evidence, MCP arguments:
-
P2: two documented bypasses written through inputs the reader already lists are
changed, and thewhysays they gained no documented widening rule.- (a) Claude
permissions.defaultMode: bypassPermissions.- Base
settings: '{"permissions":{"defaultMode":"default"}}', head…"bypassPermissions"}}'. - Result:
low changed, "a change that gains no documented widening rule is not counted as a widening". - The same edit in
.claude/settings.jsongives⚠ critical added … defaultMode: bypassPermissions — skips every permission prompt, from #827's setting table. --settings '{"permissions":{"defaultMode":"bypassPermissions"}}'insideclaude_argsis the same.- Why this is the documented rule: the action's
settingsinput is "Claude Code settings as JSON string or path to settings JSON file" (base-actionaction.yml). So this is bypassed permission checks, spelled another way.
- Base
- (b)
openai/codex-actionpermission-profile: ":workspace"→":danger-full-access".- Result:
low changed, andverifyreportssummary.widenings: 0. - Contrast:
sandbox: workspace-write→danger-full-accessis⚠ low widened, "runs without a sandbox (danger-full-access)". - Why this is the documented rule:
:danger-full-accessis one of Codex's three built-in profiles, and it maps toPermissionProfile::Disabled, meaning no sandbox (codex-rs/core/src/config/permissions.rs).- The action passes it as
--config default_permissions=":danger-full-access"(src/runCodexExec.ts). - The action's
action.ymland README recommendpermission-profileover the legacysandboxinput for new workflows.
- Result:
- Consequences:
- STABILITY Direction says "a
danger-full-accesssandbox" widens, and neither case is counted. - The Stop hook stays quiet for both.
- Neither case has a row on
main, so this is a wrong direction on a new row, not a regression.
- STABILITY Direction says "a
- Fix:
- Read
permissions.defaultMode == "bypassPermissions"from a parsedsettingsinput or--settingsvalue asbypass_permissions. - Read
permission-profile == ":danger-full-access"asdanger_full_access: add the input to_RULE_SETTINGSand to the support page's table. - Add tests for both.
- If that is out of scope, name both on the support page, and have the row not say that no rule was gained.
- Read
- (a) Claude
Non-blocking (P3):
allowed_bots: "dependabot,*"reads "accepts runs triggered by any user (allowed_bots: *)". That gate admits any bot, so "any bot" would be accurate.- The expression rule is conservative in two ways a reviewer may not expect. Both are documented.
- A literal
--dangerously-skip-permissionsadded before an expression is not claimed when the oldclaude_argsheld any expression. For example,--allowedTools Read --model ${{ vars.M }}→--dangerously-skip-permissions --allowedTools Read --model ${{ vars.M }}ischanged. - A flag after a GitHub-controlled path, as in
--mcp-config ${{ runner.temp }}/mcp.json --dangerously-skip-permissions, is never read. - Two options: treat path-valued contexts such as
runner.tempandgithub.workspaceas unable to inject words, or claim a gain when the flag's text is absent from the old value.
- A literal
- The
c7f67550commit message says it was rebased ontodaa4ad5f. Its base is44b9e05d. - Suppose a job is renamed and a new job both meet a rule. Which of the two the
whycalls "moved" and which it calls gained depends on iteration order. Direction is right either way.
Verified:
- The issue reproduction:
- On
origin/main44b9e05d,args,runandcheckoutgive no row, andtriggergives a row with no agent context. - On this head, all four give the rows the PR description quotes.
triggernamesreview/steps[1]besideissue_commentandpull-requests. check --format agent-control-jsonhas the samecontrol_stateanddecisionon both engines for these four cases and for the cycle-1 reproductions below.
- On
- Cycle-1 findings, re-run at this head:
- F1 fixed. Attached
--settings='{…DB_PASSWORD…apiKeyHelper…}'and--mcp-config='{…DB_API_TOKEN…X-API-Key…}': zero canaries indifftext and JSON,audit --host --json,check --format agent-boundary-json,pr-comment.mdandverifier.json. The cell publishes'--settings={"apiKeyHelper":"<redacted>","env":{"DB_PASSWORD":"<redacted>"}}'. - F2 fixed.
- Prose
--append-system-prompt "Never print bearer tokens in review comments"besidepull-requests: read→writeis⚠ high widened"grants write permissions to workflow jobs".verifier.jsonhost_comparisoniscomparablewith 1 row. - The same workflow unchanged, with only a new
.mcp.jsonserver, leavescheckwithincomparable_reasons: []and theaddedrow, as onmain.
- Prose
- F3 fixed. Renaming job
review→code-reviewwith the same bypassing launch ischanged,summary.widenings: 0, and thewhynames the move. - F4 fixed.
--model ${{ vars.CLAUDE_MODEL }}/--allowedTools Read→… / --dangerously-skip-permissionsischanged, and thewhycarries the expression sentence.- The
allowed_non_write_usersgate case is named and not claimed. --allowedTools Read→--dangerously-skip-permissions --model ${{ … }}widens.--dangerously-skip-permissions --model ${{ … }}→… --model opusischanged.
- F5 fixed. The PR is
MERGEABLE/CLEANon44b9e05d. CI onc7f67550is green in every job;release-tag-consistencyis skipped, as on every PR. - F6 fixed, except for the setting-publication paragraph that finding 1 covers. The test count (169) and the CI claims match.
- F1 fixed. Attached
- Negative controls give 0 rows each:
- a non-agent action's
with: claude_argsedit - a
run:that onlyechoesclaude -p --dangerously-skip-permissions
- a non-agent action's
- The composition note:
- It names
issue_comment, the job's write scopes orwrite-all, a jobenvsecret and arefs/pull/${{ … }}/headcheckout. - It stops at five agent steps and counts the rest.
- It names
- Baseline migration:
- A
0.6baseline with a workflow, saved bymain, isincomparablewithbaseline_workflow_agent_launches_unavailableandhas_drift: null. --save-baselineover it exits 2 withunsupported_baseline_schema.preflightraiseshost_grant_drift.- All three match the STABILITY note.
- A
- Other redaction and size checks:
- Codex
-coverrides underhttp_headers,bearer_tokenandshell_environment_policy.set.API_KEYpublish<redacted>. - A
plugin_marketplacesURL with userinfo is withheld. - A 40 KB
claude_argsprompt leavespr-comment.mdbounded (1.8 KB).
- Codex
- Tests would fail without the cycle-1 fixes. I mutated
host_grants.pyin a scratch copy and rantests/test_workflow_agent_launches.pyfor each mutation:- dropping the
--name=valuewithholding: 2 failures - disabling the moved-rule pairing: 5 failures
- reading every word regardless of expressions: 15 failures
- disabling the expression-before check: 3 failures
- dropping the
- Tests run locally at this head:
- 7 workflow and host files: 787 passed and 1 skipped. They are
test_workflow_agent_launches,test_workflow_step_action_references,test_reusable_workflow_secret_mappings,test_workflow_label_redaction,test_workflow_capability_diff,test_distribution_surface_parityandtest_host_audit. - 26 related files: 1436 passed and 5 skipped. They cover the control envelope, agent instructions, capability diff, the cold-start and host-config replays, the design-partner pilot, docs links, the host boundary, coverage, review changes, install hooks, the local contract, manifest-free PR rows, org governance, preflight, the public surface contract, schema boundaries and round-trip, unchanged limits, prompt-disabling settings and the MCP URL digest.
scripts/generate_schemas.py --checkis clean, andscripts/regenerate_goldens.py --checkreports 24 artifacts with 0 changed.ruffis clean on the changed source and test files.- Rebuilding
llms-full.txtchanges nothing.
- 7 workflow and host files: 787 passed and 1 skipped. They are
- Neither finding has a test, so the suite passes with both present.
The third review of #850 at c7f6755 found one P1 and one P2 defect and four P3 notes. The branch is on origin/main 44b9e05; no rebase was needed. A JSON setting publishes its shape, not its free text (C2-F1). _withheld_json published the whole _redact_secret_values tree. That tree is the host readers' digest input, not what they publish: it keeps every string outside env, headers and secret-named keys. So an mcp-remote --header "Authorization: Bearer ..." argument in --mcp-config, and a hook's curl command in settings, reached diff text, diff --json, audit --host --json, the PR comment and verifier.json. _json_shape now keeps key names, numbers, booleans and null, redacts what the host readers redact, and replaces each other string with <withheld:...>, a 12-hex digest of redacted_config_sha256 for that string, so an edit to it is still a changed row. It keeps only the strings a host reader publishes: - a permissions.allow/ask/deny rule and a documented Claude Code setting's value (defaultMode, the switches, enabledMcpjsonServers entries); - an MCP server's command name and its URL's scheme and host, followed by the digest when they drop a command's arguments or a URL's query. A codex --config table or array is read under its key path, so mcp_servers.gh={command="gh", ...} keeps its command name. The canary sweep adds the mcp-remote header and the hook command in every spelling (action input, claude_args, CLI flag, codex -c table). Re-running the reviewer's two repositories through ./shipgate gives 0 canaries in every output. The documented bypasses written through listed inputs widen (C2-F2). - Claude Code settings written as JSON meet bypass_permissions when their defaultMode is bypassPermissions, read by claude_setting_values as the settings reader reads .claude/settings.json. This covers the action's settings input, a --settings value in claude_args, and the CLI's --settings flag. A path is not read, and a settings value holding an expression meets none. - openai/codex-action's permission-profile: :danger-full-access meets danger_full_access. ":danger-full-access" is Codex's reserved name for its built-in full-access profile (BUILT_IN_PERMISSION_PROFILE_DANGER_FULL_ACCESS). Mode inputs are now (input, value, rule) triples. settings and permission-profile join _RULE_SETTINGS, so a gain after an expression in either is named and not claimed, as for claude_args. P3 notes: - allowed_bots: "*" now reads "accepts runs triggered by any bot" through agent_rule_text. - Which of two gaining jobs a moved rule goes to no longer depends on declaration order: the same launch arriving is matched before a job that merely stopped launching the agent. The support page, the STABILITY migration note, the contract summary, the schema docstrings and the CHANGELOG entry now say what a structured value publishes instead of claiming it publishes what the host readers would, and list the two rules. host-grants 0.7 is extended in place; it is unreleased.
|
Addressed review cycle 2. New head C2-F1 (P1): a JSON setting now publishes its shape, with none of its free textCause, as the review found. Fix.
A codex Evidence. I rebuilt the review's two repositories and ran them through These match what the host readers publish for the same text:
Docs. The claim "a setting publishes what the host readers would publish for the same text" was false. It is replaced with a statement of what is published, in:
C2-F2 (P2): the two documented bypasses now widen
Evidence (
Tests: five new widening cases (action
There is also one expression-before case for Nonblocking notes
Tests run
|
pengfei-threemoonslab
left a comment
There was a problem hiding this comment.
Independent coding-agent review/address cycle 3 at 6294321d.
Not mergeable yet: two P2 findings. Both cycle-2 findings are fixed; the evidence is below. The first new finding is that the branch conflicts with origin/main, so no CI has run on this head. The second is about direction. A Codex full-access sandbox written as a --config override in a literal codex exec step is changed. Written as --sandbox danger-full-access, the same thing widens.
Setup for every reproduction:
- Synthetic two-commit repositories built with the issue's
pairhelper. ./shipgatefrom detached worktrees of this head and oforigin/main01777037.W=.github/workflows/agent.yml,on: pull_request,permissions: {contents: read}, and onereviewjob, unless a finding says otherwise.
-
P2: the PR conflicts with
origin/main, and no CI has run on6294321d.- Evidence:
gh api repos/ThreeMoonsLab/agents-shipgate/pulls/850reportsmergeable: falseandmergeable_state: dirty.gh pr checks 850reports "no checks reported". The check-runs API givestotal_count: 0for6294321d.ci.ymlruns onpull_requestand on pushes tomain. So no check will start on this branch until the conflict is resolved.git merge-tree origin/main 6294321dconflicts in 7 files:STABILITY.md,docs/agent-contract-current.md,docs/design-partner-pilot-results.md,docs/distribution-surfaces.md,llms-full.txt,src/agents_shipgate/schemas/contract.pyandtests/test_distribution_surface_parity.py.
- What the rebase has to take into account:
- #852 (#821) already moved runtime contract 40 → 41 (unreleased), verifier 0.20 → 0.21 and capability diff 0.3 → 0.4.
- Under the agreed rule, this PR extends contract 41 in place: one
CONTRACT_VERSIONcomment naming both #821 and #823, not a 42. - Several texts say "Verifier 0.20 and capability diff 0.3 are unchanged" or "do not move": the contract comment, the STABILITY migration note and the PR description. After the rebase they need to say that 0.21 and 0.4, which #852 minted, do not move here.
- Both PRs edited the pilot ledger's source-tree column, so it has to be re-measured on the rebased tree.
- The
capability_diffrow indocs/distribution-surfaces.mdhas to keep #852's two added roots (core/unread_inputs.pyandcli/verify/changed_inputs.py) and both paragraphs.
- Trial merge: I merged in a scratch worktree and kept both sides of every conflict. I then ran 13 related test files: the new agent-launch tests, #852's
test_unread_changed_inputs, coverage, workflow diff, step references, reusable secrets, unchanged limits, manifest-free rows, the verifier control contract, the local contract, parity, and the host-config and cold-start replays.- Result: 1001 passed, 72 skipped, 1 failed.
- The one failure is the parity test. It failed only because my blunt merge left two
capability_diffrows. - I found no semantic conflict, so the resolution looks mechanical. It still has to be done, and CI has to run on the result.
- Fix:
- Rebase onto
origin/mainand resolve the conflicts as above. - Rebuild
llms-full.txtand re-measure the pilot ledger. - Update the PR description's version, test-count and CI statements (see the P3 notes).
- Let CI run on the new head.
- Rebase onto
- Evidence:
-
P2: Codex's full-access sandbox written as a
--configoverride in a literalcodex execstep ischanged, while--sandbox danger-full-accesswidens.- Evidence. Every base is
run: codex exec -s workspace-write 'review':codex exec -c sandbox_mode="danger-full-access" 'review'giveslow changed. The cell reads--config sandbox_mode=danger-full-access, thewhysays "a change that gains no documented widening rule is not counted as a widening", andexpandsisfalse.codex exec -c default_permissions=":danger-full-access" 'review'gives the same result.codex exec -s danger-full-access 'review'and--sandbox=danger-full-accessgive⚠ low widened"runs without a sandbox (danger-full-access)".codex exec -sdanger-full-access 'review'is clap's attached short value, which the reader already accepts for-c<override>. It giveschanged, and the cell says "runs codex exec with no permission flags".
- Why these are the documented rule:
sandbox_modeis the config key that--sandboxsets (codex-rs/config/src/config_toml.rs: "Sandbox mode to use").default_permissionsselects a permission profile, and a name starting with:is a built-in profile.openai/codex-actionpassespermission-profile: :danger-full-access, which this head now widens, to the CLI as exactly--config default_permissions=":danger-full-access"(src/runCodexExec.ts,case "profile"). So the input form widens and its literal CLI form does not.- The cycle-2 address comment says the action passes it as
--permission-profile. It does not, and the currentcodex exechas no such flag (codex-rs/utils/cli/src/shared_options.rs).
- The reader already has what it needs:
--config/-cis in_CODEX_FLAGS._withheld_configalready parses the override's value as TOML for publication._flag_rulesreads only--dangerously-bypass-approvals-and-sandboxand--sandbox.
- Scope:
- This is the
run:path. - Through the action's
codex-args, the action refusessandbox_modeanddefault_permissionsoverrides unlesssafety-strategy: unsafe(RESTRICTED_CONFIG_ROOTS), andunsafealready widens. - Neither form has a row on
main, so this is a wrong direction on a new row, not a regression. It is the same class as cycle 2's finding 2.
- This is the
- Fix:
- Read a
--configoverride, in any of the four ways the reader already accepts it, whose key issandbox_modewith the valuedanger-full-access, ordefault_permissionswith:danger-full-access, asdanger_full_access. - Read
-s<mode>as--sandbox. - Add tests, and list the new spellings in the support page's
codex execparagraph and in STABILITY Direction. - If that is out of scope, name these spellings on the support page as not read for a rule, and have the row not say that no rule was gained.
- Read a
- Evidence. Every base is
Non-blocking (P3):
- A bearer token in a permission rule inside the
settingsinput is published on every route.- Example:
settings: '{"permissions":{"allow":["Bash(curl -H \"Authorization: Bearer tokCANARY0123456789abcdef\" https://x.example.com/a)"]}}'publishesAuthorization: <redacted> tokCANARY…. - Counts:
difftext 1 and--json2,audit --host1,check --format agent-boundary-json1,agent-control-json1,verifytext 1,pr-comment.md1 andverifier.json2. - Cause:
_sanitize_sensitive_stringreplaces only the scheme word afterAuthorization:. After that, the #802 label redaction no longer seesBearer <token>. On the raw text,published_workflow_labelwould have redacted the whole value. - This matches the settings reader. On
origin/main, the same rule in.claude/settings.jsonpublishes a random 32-character token indiff,audit --host,verify, the PR comment andverifier.json. So this head inherits the gap rather than introducing it, and the right fix is in the shared sanitizer, as a separate change. The same rule written asclaude_args: --allowedTools "…"is redacted whole.
- Example:
- A JSON-array
codex-argselement can lose an escaped quote. When an element holding a quoted URL has the URL withheld, its escaped closing quote goes too."mcp_servers.x.url=\"https://h2.example.com/p?token=…\""publishes as"mcp_servers.x.url=\"https://h2.example.com/<redacted-path>"", so the published value is not the JSON array the support page describes. Nothing leaks. - The PR description is out of date.
- It says
tests/test_workflow_agent_launches.pyhas 169 cases. It collects 184. - Its CI line is about
c7f67550. No CI has run on this head.
- It says
Verified:
- The issue reproduction:
- On
origin/main01777037,args,runandcheckoutstill give no row, andtriggergives a row with no agent context. - On this head, all four give the rows the PR description quotes.
- On
- Negative controls give 0 rows each:
- a non-agent action's
with: claude_argsedit - a
run:that onlyechoesclaude -p --dangerously-skip-permissions
- a non-agent action's
- Cycle-2 findings, re-run at this head:
- C2-F1 fixed. The
--mcp-configserver withmcp-remote … --header "Authorization: Bearer tokCANARY…"and thesettingshook command withhookCANARY77publish none oftokCANARY,hookCANARY77,X-Auth-Token,mcp-remote,mcp.example.comorhooks.example.com. I checkeddifftext and--json,audit --host --json, bothcheckformats,verifytext,agent-handoff.json,current-control.json,pr-comment.mdandverifier.json. The cells show the shape, for example"args":["<withheld:…>",…],"command":"npx". - C2-F1, a wider sweep. I added 16 more canaries:
- an MCP command path and its arguments
- a URL's userinfo, path and query
- a
statusLinecommand,additionalDirectories,env,apiKeyHelperandmodel mcp_configarguments andenv- codex
-cURL, table, array andnotifyvalues
None of them reached any of those artifacts. The one exception is a codex scalar override,-c 'mcp_servers.z.command="…"', which the support page documents as published.
- C2-F2 fixed.
settingsdefaultModedefault→bypassPermissionswidens.claude_args: --settings '{"permissions":{"defaultMode":"bypassPermissions"}}'widens.permission-profile:workspace→:danger-full-accesswidens.- Each
whynames the rule and the step.
- C2-F1 fixed. The
- The tests would fail without the cycle-2 fixes. I mutated
host_grants.pyin a scratch copy and rantests/test_workflow_agent_launches.pyfor each mutation:- publishing
_redact_secret_valuesinstead of_json_shape: 3 failures - dropping the
settingsinput rule: 3 failures - dropping the
permission-profilemode: 1 failure - dropping
--settingsinclaude_args: 1 failure
- publishing
- Tests run locally at this head:
- 7 workflow and host files: 802 passed and 1 skipped. They are
test_workflow_agent_launches(184 cases),test_workflow_step_action_references,test_reusable_workflow_secret_mappings,test_workflow_label_redaction,test_workflow_capability_diff,test_distribution_surface_parityandtest_host_audit. - 21 related files: 1341 passed. They cover the agent boundary, the control envelope and its rows, capability diff, the cold-start and host-config replays, the design-partner pilot, determinism, docs links, unread surfaces, comparison coverage, review changes, install hooks, the local contract, manifest-free PR rows, the MCP URL digest, prompt-disabling settings, the public surface contract, schema boundaries and round-trip, and unchanged limits.
scripts/generate_schemas.py --checkis clean, andscripts/regenerate_goldens.py --checkreports 24 artifacts with 0 changed.
- 7 workflow and host files: 802 passed and 1 skipped. They are
- CI: none on
6294321d(finding 1).
6294321 to
9e76428
Compare
The second review of #850 at 25c13ce found two P1 and four P2 defects. The branch is rebased onto origin/main daa4ad5. Attached JSON values are withheld (F1). A --settings={...} or --mcp-config={...} word inside claude_args, and codex's attached -c<override> / -c=<override>, published their env, header and apiKeyHelper values verbatim, because only a word starting with "{" was withheld. _withheld_words now splits a --name=value word and withholds the value, and reads codex's attached -c through _withheld_config, as clap reads it. The canary sweep carries both spellings. Redacted prose no longer refuses the comparison (F2). The #802 label redaction rewrites ordinary prose ("never print bearer tokens", "Authorization: headers"), and a redacted agent setting made the workflow a blocking limit. That hid every row beside it and made every check incomparable while the workflow existed. The rules are already read from the declared text, so a redacted setting is now compared by its published text and its widening_rules. uncompared_agent_launch_texts names it as a non-blocking limit, and the row cell shows the redacted text. A redacted checkout ref still refuses, as a redacted step reference does (#767), because it names the code a job runs. A renamed job's launch moves its rules (F3). _agent_rule_gains keyed a rule on its job, so renaming a job that launches a bypassing agent was a widening. agent_rule_gains now pairs a rule one job gains with the same rule another job lost, when the launch that met it left that job: the job no longer launches that agent, or the same launch (agent_launch_key less the job) now runs in the gaining job. The why names the move. A second job gaining a rule, or a different launch gaining one while the first job still launches that agent, still widens. An expression no longer turns off every rule, and the row says what it leaves unread (F4). Rules are read from literal text a ${{ }} expression cannot reach: - the words of claude_args or codex-args before the first expression, less the word it touches and any quoted run still open at it; - the elements of a JSON-array codex-args before the one holding it; - the gate entries that hold none. A setting holding an expression is published with holds_expression (the unreleased 0.7 schema extends in place), and a row that changes it says the text the expression reaches is not read. A gain where the job's launch held an expression before, in the input the rule is read from, is named and not claimed, as unknown_before is. That also fixes a false widening at 25c13ce: replacing --model ${{ vars.CLAUDE_MODEL }} with --model opus beside --dangerously-skip-permissions was reported as gaining the bypass. Two non-blocking fixes: - _published_value reads each expression as one word, so an expression in a URL's userinfo is withheld with it rather than garbling the URL and publishing its path. - A codex --config value that starts like a table and does not parse, as string-argv leaves a quoted one, is withheld as unparsed_json. Rebase (F5). CHANGELOG keeps #853's #778 line beside #823's under github_action row. llms.txt and ai-search-summary state the source tree as contract 41, unreleased, ahead of the published v1.1.0 (contract 40), as test_public_surface_contract requires while the two differ. The pilot ledger's source-tree column was re-taken on the rebased tree, through ./shipgate beside the engine of e3c6cb0, the commit v1.1.0 was cut from: - the only differences are contract 40 -> 41 and inventory schema 0.6 -> 0.7; - diff rows are byte-identical; - check JSON differs only in the launcher path its next action names. The support page, the STABILITY preamble and migration note, the CHANGELOG entry, agent-contract-current, the Stop hook sentence in integrations, the capability_diff row in distribution-surfaces with its parity comment, the 0.7 schema files and llms-full.txt are updated to match.
The third review of #850 at c7f6755 found one P1 and one P2 defect and four P3 notes. The branch is on origin/main 44b9e05; no rebase was needed. A JSON setting publishes its shape, not its free text (C2-F1). _withheld_json published the whole _redact_secret_values tree. That tree is the host readers' digest input, not what they publish: it keeps every string outside env, headers and secret-named keys. So an mcp-remote --header "Authorization: Bearer ..." argument in --mcp-config, and a hook's curl command in settings, reached diff text, diff --json, audit --host --json, the PR comment and verifier.json. _json_shape now keeps key names, numbers, booleans and null, redacts what the host readers redact, and replaces each other string with <withheld:...>, a 12-hex digest of redacted_config_sha256 for that string, so an edit to it is still a changed row. It keeps only the strings a host reader publishes: - a permissions.allow/ask/deny rule and a documented Claude Code setting's value (defaultMode, the switches, enabledMcpjsonServers entries); - an MCP server's command name and its URL's scheme and host, followed by the digest when they drop a command's arguments or a URL's query. A codex --config table or array is read under its key path, so mcp_servers.gh={command="gh", ...} keeps its command name. The canary sweep adds the mcp-remote header and the hook command in every spelling (action input, claude_args, CLI flag, codex -c table). Re-running the reviewer's two repositories through ./shipgate gives 0 canaries in every output. The documented bypasses written through listed inputs widen (C2-F2). - Claude Code settings written as JSON meet bypass_permissions when their defaultMode is bypassPermissions, read by claude_setting_values as the settings reader reads .claude/settings.json. This covers the action's settings input, a --settings value in claude_args, and the CLI's --settings flag. A path is not read, and a settings value holding an expression meets none. - openai/codex-action's permission-profile: :danger-full-access meets danger_full_access. ":danger-full-access" is Codex's reserved name for its built-in full-access profile (BUILT_IN_PERMISSION_PROFILE_DANGER_FULL_ACCESS). Mode inputs are now (input, value, rule) triples. settings and permission-profile join _RULE_SETTINGS, so a gain after an expression in either is named and not claimed, as for claude_args. P3 notes: - allowed_bots: "*" now reads "accepts runs triggered by any bot" through agent_rule_text. - Which of two gaining jobs a moved rule goes to no longer depends on declaration order: the same launch arriving is matched before a job that merely stopped launching the agent. The support page, the STABILITY migration note, the contract summary, the schema docstrings and the CHANGELOG entry now say what a structured value publishes instead of claiming it publishes what the host readers would, and list the two rules. host-grants 0.7 is extended in place; it is unreleased.
|
Addressed review cycle 3. New head C3-F1: conflicts with #852 and contract 41 (blocking)I rebased all five commits onto
C3-F2: Codex's full-access sandbox through
|
head run: |
before (6294321d) |
this head |
|---|---|---|
codex exec -c sandbox_mode="danger-full-access" 'review' |
low changed, "gains no documented widening rule" |
⚠ low widened, "an agent launch now runs without a sandbox (danger-full-access) (review/steps[0])" |
codex exec -c default_permissions=":danger-full-access" 'review' |
low changed |
⚠ low widened, same why |
codex exec -sdanger-full-access 'review' |
changed, cell "runs codex exec with no permission flags" |
⚠ low widened, cell runs codex exec with --sandbox danger-full-access |
codex exec -s danger-full-access / --sandbox=danger-full-access |
⚠ low widened |
unchanged |
codex exec -s workspace-write -c sandbox_mode="danger-full-access" 'review' |
low changed |
low changed (the --sandbox flag takes precedence, see below) |
How it is read (core/host_grants.py):
- Attached short values.
_read_flagsnow reads a short flag's attached value as clap does:-s<mode>,-s=<mode>,-c<override>,-c=<override>and-p<name>. Each is published under its primary spelling, so moving from-s danger-full-accessto-sdanger-full-accessgives no row.openai/codex-action's owncodex-argschecks read-s…and-c…the same way (arg.startsWith("-s") && !arg.startsWith("--")inrunCodexExec.ts). --configoverrides. Any of the four spellings is read as the CLI'sparse_overridesreads it (codex-rs/utils/cli/src/config_override.rs). The key is trimmed, and the value is parsed as TOML or else taken as text with its quotes trimmed.sandbox_mode = "danger-full-access"ordefault_permissions = ":danger-full-access"meetsdanger_full_access, withsetting: --config. It is the same rule as--sandbox, so moving between the flag and the override ischanged, not a widening.- Precedence, as the CLI resolves it. The last override of a key counts. A
default_permissionsoverride selects the profile over asandbox_modeone, because session flags that setdefault_permissionsselect profile syntax. A--sandboxflag outranks both.resolve_permission_config_syntaxreturns legacy syntax whenever the flag is set, andderive_permission_profilethen takessandbox_mode_override.or(self.sandbox_mode)(codex-rs/core/src/config/mod.rs,codex-rs/config/src/config_toml.rs). So an override beside--sandboxmeets no rule, and thewhystays true for it. - Not
codex-args. A sandbox--configoverride incodex-argsmeets no rule. Aftercodex-args, the action always appends its own--sandbox <mode>, or for apermission-profileits own--config default_permissions=…(runCodexExec.ts, thepermissionSelectionswitch). Either one takes precedence over an override written earlier. As you noted, outsidesafety-strategy: unsafethe action refuses these roots in any case (RESTRICTED_CONFIG_ROOTS). - Docs. The support page's
codex execparagraph lists the spellings, the precedence, and what meets no rule (a key under another table, such asprofiles.<name>.sandbox_mode, and--profile). Its codex-action table cell says why acodex-argsoverride meets none. The STABILITY Direction bullet anddocs/agent-contract-current.mdsay the same.
Tests in tests/test_workflow_agent_launches.py:
test_codex_exec_full_access_widens_in_every_spelling_the_cli_reads(14 spellings)test_a_codex_exec_sandbox_the_cli_does_not_select_is_changed(7 cases: flag wins, last override, profile over mode, profile-scoped key, custom profile name, read-only)test_attached_short_values_publish_under_the_primary_spellingtest_codex_args_sandbox_overrides_are_read_as_the_action_passes_them
A correction to my cycle-2 comment: I wrote that codex-action passes permission-profile as --permission-profile. It does not. It passes --config default_permissions=<JSON-quoted name>, and the current CLI has no --permission-profile flag. That claim was only in the PR comment, and no published doc repeats it.
Something I noticed and did not change, because it is outside this finding: codex-args: --sandbox danger-full-access still widens, as the support page has said since cycle 1. The action appends its own --sandbox <mode> after codex-args, and codex's --sandbox is a plain clap Option with no args_override_self. So such a run most likely fails as a repeated argument and does not run unsandboxed. If that is right, the existing rule over-reports and never under-reports. I left it as it is. If you want it narrowed, I will do that separately.
Nonblocking
- P3, JSON-array
codex-argsURL: fixed. When_published_valuewithholds a URL, it now keeps a backslash that the URL match ends at. So the\"a JSON-array element uses to escape a quote stays an escape.["-c","mcp_servers.x.url=\"https://h2.example.com/p?token=…\"","--json"]now publishes["-c","mcp_servers.x.url=\"https://h2.example.com/<redacted-path>\"","--json"], which parses as JSON. Test:test_a_quoted_url_in_a_json_array_codex_args_element_keeps_the_array_valid. - P3, bearer token in a
settingspermission rule: not changed here. You traced it to the shared_sanitize_sensitive_string, which replaces only the scheme word afterAuthorization:, and.claude/settings.jsononmainpublishes the same token. So the fix belongs to the separate shared-sanitizer task you spawned (task_5d52561a). Changing that sanitizer here would change what every host reader publishes, which goes well beyond Read agent launches in CI workflows: known agent action inputs, literal agent CLI run steps and head-ref checkouts #823. - P3, stale PR description: updated.
tests/test_workflow_agent_launches.pynow collects 210 cases. The version, ledger, benchmark and local-run paragraphs describe the rebased tree, and the CI line reports this head's green run.
Run locally on this head
tests/test_workflow_agent_launches.py(210 cases) and 124 other test files: those whose names cover what this PR touches (workflow, host, capability, diff, replay, pilot, schema, contract, surface, coverage, control, preflight, golden, instruction, redaction, setting, prompt, unread, verif); those that readSTABILITY.md,CHANGELOG.md,llms-full.txt, the contract doc, the support page or the pilot ledger, or import the host-grant reader; Route an enabled in-repository plugin's hook at a non-registry path to review (#809) #860'stests/test_enabled_plugin_hook_routing.py; and the governance and semantic cold-start benchmarks. 125 files in all on9e764284, 5959 passed, 12 skipped, 0 failed. The host-config and cold-start replays reproduce their run-of-record scores.scripts/generate_schemas.py --checkclean;scripts/regenerate_goldens.py --checkreports 24 artifacts, 0 changed;ruff check src tests scripts shipgateclean; rebuildingllms-full.txtagain changes nothing.- CI on
9e764284:suite (1/2/3),coverage,test,verify,verify-self,mcp extra (floor/newest),clean-checkout-launcherandwindows-launcherall pass;release-tag-consistencyis skipped, as on every PR. GitHub reports the PRmergeable,clean.
pengfei-threemoonslab
left a comment
There was a problem hiding this comment.
Independent coding-agent review/address cycle 2 at 9e764284.
This is the fifth review of this PR. It follows the cycle-3 review at 6294321d and its address comment.
Not mergeable yet: two P2 findings. Every earlier finding is fixed at this head (evidence below), and CI is green. Both new findings concern a run: step that launches an agent in a shape the reader does not recognise:
- In some such shapes, the step gets no row and no limit, although the support page says it is listed as
unresolved. - When a step moves from a read launch to such a shape, the row says the step "no longer launches an agent". That is false, and in the reproduction the new launch bypasses permission checks.
Setup for every reproduction:
- Synthetic two-commit repositories built with the issue's
pairhelper. ./shipgatefrom detached worktrees of this head and oforigin/main997e9260.W=.github/workflows/agent.yml,on: pull_request, and onereviewjob withcontents: readandpull-requests: write.
-
P2: an agent CLI inside a double-quoted
$(…), inside backticks, or after a shell reserved word on the same line is not listed at all. It gets no row and no limit.- Evidence, with a base of
run: echo done:- Head
run: gh pr comment "$PR" --body "$(claude -p --dangerously-skip-permissions 'Review this change')".diffprintsNo static host-grant changes detected, and the coverage line reads "compared; changed, but no grant this entry compares changed, so no row (redacted values such as env values and apiKeyHelper are not compared)".audit --host --jsonhas noagent_launchesand no coverage issue naming the step.
- Editing inside that substitution, from
--allowedTools Readto--dangerously-skip-permissions, gives the same result. if [ -n "$PR" ]; then claude -p --dangerously-skip-permissions 'Review'; fion one line gives the same result._run_launchesalso returns[]for:REVIEW="$(claude -p x)"followed byecho "$REVIEW" >> $GITHUB_OUTPUTREVIEW=`claude -p …`for f in a b; do claude -p … "$f"; done{ claude -p x; } > out
- Head
- Contrast:
- The unquoted
REVIEW=$(claude -p --dangerously-skip-permissions "review")isunresolved/compound_command, with a named limit. npm ci && claude -p …is a row that says "a step launches an agent in a form this audit does not read".- So quoting the substitution, as shellcheck SC2046 recommends and as the common
--body "$(claude -p …)"pattern does, turns a named limit into silence.
- The unquoted
- Cause:
_run_launcheslooks for the agent only at the head of each top-level simple command that_shell_wordsproduces.- A double-quoted
"$(…)"is one word, and a backtick is not an operator. _agent_commandtakesthen,door{to be the command.
- What this contradicts:
docs/host-boundary-support.mdsays: "Arun:holding more than one command …, a shell expansion ($VAR,$(…), a backtick) or a${{ }}expression is listed asunresolvedwith that reason, once for each agent CLI it starts at the head of a command". In each case above, the agent CLI is at the head of a command, the substituted one or the one afterthen/do/{.- Acceptance item 3: "Unsupported shapes (a compound command,
$VARexpansion, …) are named limits". - This is not a regression against
main, which reads no launch at all. But the zero-row result reads as covered, which is the problem #823 set out to fix.
- Fix:
- Name an agent CLI at the head of a
$(…)or backtick substitution asunresolved(shell_expansion), anywhere outside single quotes. - Skip shell reserved words (
then,do,else,elif,!,{,time) before reading a command's head. - Keep
echo "claude -p …"unlisted. It is the acceptance's negative control, andtest_a_command_that_launches_no_headless_agent_is_not_listedholds it. - Add these shapes to
test_a_shape_this_reader_does_not_read_is_unresolved_and_publishes_no_text. - If that is out of scope, narrow the support-page sentence to top-level simple commands, and list these shapes under Known unread surfaces.
- Name an agent CLI at the head of a
- Evidence, with a base of
-
P2: when a step moves from a read launch to a shape the reader does not recognise, the row says "a step no longer launches an agent", although the step still launches one.
- Evidence, with a base of
run: claude -p --allowedTools Read 'Review'in every case:- Head
run: npx @anthropic-ai/claude-code -p --dangerously-skip-permissions 'Review':changed,expands: false, and thewhyreads "a step no longer launches an agent (review/steps[0]); agent launch settings are compared as declared text, and a change that gains no documented widening rule is not counted as a widening". - Head
run: gh pr comment "$PR" --body "$(claude -p --dangerously-skip-permissions 'Review')"gives the same row and the samewhy. - Base
run: codex exec -s workspace-write 'review', headrun: codex -c sandbox_mode=danger-full-access exec 'review'gives the same row and the same sentence.- The Codex CLI copies root
-c,--sandboxand--dangerously-bypass-approvals-and-sandboxintoexec(inherit_exec_root_optionsandprepend_config_flagsincodex-rs/cli/src/main.rs). - So this step runs without a sandbox.
- The Codex CLI copies root
- Head
- In all three cases, the step now launches the agent with permission checks or the sandbox bypassed, and the row tells the reader the launch is gone.
maingives no row for these edits. So this is a new false statement, not a changed one.- Cause:
_agent_launch_reasonswords any launch that is absent atafteras("removed", "a step no longer launches an agent")(core/capability_diff_rows.py:386). The reader cannot know that. The support page itself listsnpxand path launches as unread, and says "Editing one gives no row". - Fix:
- Word the removed case as what was established, for example "a step no longer declares an agent launch this audit reads (…); a launch in a form it does not recognise, such as
npxor a script, is not read". - Add a test for
claude -p …→npx @anthropic-ai/claude-code -p ….
- Word the removed case as what was established, for example "a step no longer declares an agent launch this audit reads (…); a launch in a form it does not recognise, such as
- Evidence, with a base of
Non-blocking (P3):
- A
${{ }}anywhere inrun:makes the whole launchunresolved, so the literal words before it are not read, as they are forclaude_args.- Example: adding
claude -p --dangerously-skip-permissions "Review PR ${{ github.event.pull_request.number }}"ischanged, "does not say what that agent may do". - This is documented, but it is the most common real
run:shape. Reading the words before the first expression, as the action path does, would cover it.
- Example: adding
codex [root options] execis not recognised, beyond the wording in finding 2.- Adding
codex --dangerously-bypass-approvals-and-sandbox exec 'review'gives no row and no limit. - The support page covers it only through its generic "any way the workflow reader below does not recognise". Its named examples do not include it.
- Adding
- A bypassing launch that moves from a
contents: readjob into an existingcontents: writejob ischanged, when the job it left is removed.- Its
whyreads "moved between jobs … not counted as a widening … it now runs with the receiving job's token permissions", and the note names the write scope. - This is consistent with #771, but the Stop hook stays quiet about the broader token.
- Pairing a move only when the receiving job's permission context is no broader would close this.
- Its
Verified:
-
CI on
9e764284. 11 check runs succeeded, andrelease-tag-consistencywas skipped, as on every PR. GitHub reportsmergeable,clean, againstorigin/main997e9260. -
The issue reproduction.
- On
997e9260,args,runandcheckoutgive no row, andtriggergives a row with no agent context. - On this head, all four give the rows the PR description quotes:
argsis⚠ low widened"skips permission checks",runandcheckoutarechanged, andtriggernamesreview/steps[1]besideissue_commentandpull-requests. check --format agent-boundary-jsoncarries the widened row withrequire_review, asmaindecides. Manifest-freeverifywrites the row topr-comment.mdandverifier.json.
- On
-
Cycle-3 findings, re-run at this head.
- C3-F1 fixed. The branch is rebased and clean.
schemas/contract.pyhas one v41 comment naming #821 and #823. Thecapability_diffregistry row is one row and keeps #852's two roots.## 1.1.0is untouched. - C3-F2 fixed. From a
codex exec -s workspace-writebase:- These widen with "runs without a sandbox (danger-full-access)":
-c sandbox_mode="danger-full-access",-c default_permissions=":danger-full-access",-sdanger-full-access,-s=danger-full-access,--sandbox=danger-full-access,-csandbox_mode=…,--config=sandbox_mode=…and-c 'sandbox_mode = "danger-full-access"'. -s workspace-write -c sandbox_mode="danger-full-access"stayschanged.
- These widen with "runs without a sandbox (danger-full-access)":
- Mutation checks. I removed each cycle-3 fix in a scratch worktree: the config rule, attached short values, the
--sandboxprecedence, thecodex-argsexclusion and the URL backslash. Each removal failstests/test_workflow_agent_launches.py.
- C3-F1 fixed. The branch is rebased and clean.
-
Earlier findings in this run, from
25c13ce8, re-run at this head.- Attached
--settings=/--mcp-config=canaries do not appear indifftext or JSON,audit --host, eithercheckformat,verify,pr-comment.mdorverifier.json. Neither do themcp-remote --header "Authorization: Bearer …"or hook-command canaries. - Prose "Never print bearer tokens" beside
pull-requests: read→writeis⚠ high widened. - Renaming job
review→code-reviewwith the same bypassing launch ischanged, and thewhynames the move. --model ${{ vars.CLAUDE_MODEL }}then--dangerously-skip-permissionsischanged, with the expression sentence.settingsdefaultModeandpermission-profile: :danger-full-accesswiden.- The PR description's counts match the code: 210 cases, and contract 41 extended in place.
- Attached
-
Baseline migration. I saved a
0.6baseline withmain. On this head:- drift is
incomparablewithbaseline_workflow_agent_launches_unavailable,has_drift: nullandnext_action: null --fail-on-driftexits 20--save-baselineexits 2 withunsupported_baseline_schemapreflightraises ahigh,actor: humanhost_grant_driftsignal
All of this matches the STABILITY note.
- drift is
-
The composition note. It honours a job-level
permissions:over a workflow-levelwrite-all, and it names only the agent job's secrets. -
Checkout refs and redaction.
- Realistic refs are not redacted: branch names holding "token", "password" or "api-key", a 40-hex SHA, and
${{ secrets.X }}. - A
ghp_…ref refuses, as documented. - Realistic codex
-coverrides underargs,notify,query_params,shell_environment_policy.setandbearer_token_env_varwithhold their values.mcp_servers.<x>.commandpublishes its full path as a scalar override, as the support page documents.
- Realistic refs are not redacted: branch names holding "token", "password" or "api-key", a 40-hex SHA, and
-
Fuzzing. 20,000 random inputs to
_action_launchand_run_launchesraised no exception. -
Regression check on this repository.
diffagainst four earlier bases of this repository's own workflow history gives the same rows and status on both engines. -
Tests run locally at this head:
- 10 files, 1431 tests, all passed:
test_workflow_agent_launches,test_workflow_step_action_references,test_reusable_workflow_secret_mappings,test_workflow_label_redaction,test_workflow_capability_diff,test_distribution_surface_parity,test_host_audit,test_public_surface_contract,test_design_partner_pilotandtest_local_contract. - 12 more files, 476 passed: the agent control envelope and its rows, the cold-start, host-config and semantic cold-start replays, the governance benchmark and its baseline, enabled plugin hook routing, install hooks, manifest-free PR rows, preflight, and unread changed inputs.
- 10 files, 1431 tests, all passed:
-
Test coverage of the new findings. Neither finding has a test, so the suite passes with both present.
|
Addressed review cycle 2. New head Evidence below was measured with the repository launcher from this head, on synthetic two-commit repositories ( C2R-F1 (P2): an agent CLI inside a quoted
|
run: (base echo done, or as noted) |
9e764284 |
this head |
|---|---|---|
gh pr comment "$PR" --body "$(claude -p --dangerously-skip-permissions 'Review this change')" |
0 rows, no launch, no issue | 1 changed row: "a step now launches an agent (review/steps[0])…; a step launches an agent in a form this audit does not read"; launch unresolved/shell_expansion; non-blocking issue "the agent launch at review/steps[0] (claude) is a command holding, or inside, a shell expansion…" |
same, --allowedTools Read → --dangerously-skip-permissions inside the substitution |
0 rows, no launch, no issue | 0 rows (an edit inside an unresolved launch), launch and named limit present |
if true; then claude -p --dangerously-skip-permissions 'Review'; fi |
0 rows, no launch, no issue | 1 changed row, unresolved/compound_command, named limit |
echo "claude -p --dangerously-skip-permissions" (control) |
0 rows | 0 rows, no launch |
_launches on the review's other shapes: REVIEW="$(claude -p x)" and REVIEW=`claude -p x` give shell_expansion. for …; do claude -p …; done, { claude -p x; } and the unquoted REVIEW=$(claude -p …) give compound_command, and the unquoted one is unchanged from before. echo "\$(claude -p x)", echo '$(claude -p x)' and echo "$(date) claude -p x" give none.
test_a_shape_this_reader_does_not_read_is_unresolved_and_publishes_no_text gains nine shapes: two quoted-substitution cases, a backtick, a nested substitution, unbalanced quoting with a substitution, if/then, for/do, { }, ! and time -p. test_a_command_that_launches_no_headless_agent_is_not_listed gains five negative controls. Also added: test_an_agent_cli_the_splitter_does_not_head_is_a_row_and_a_named_limit, test_an_edit_inside_a_quoted_substitution_is_quiet_and_named_as_a_limit, test_deeply_nested_substitutions_are_read_in_one_pass and the end-to-end test_the_quoted_substitution_of_the_review_is_a_row_in_diff_and_a_coverage_issue (diff JSON and text, and audit --host launches and coverage issue).
Docs: the support page's run: sentence now names reserved words and substitutions, with examples of each reason and of the negative controls. The STABILITY "Unresolved" bullet, the launch schema description (host-grants 0.7 inventory and baseline schemas regenerated; generate_schemas.py --check is clean) and the CHANGELOG Unreleased entry say the same.
C2R-F2 (P2): a launch that becomes an unrecognised form is no longer worded as gone
Fix (core/capability_diff_rows.py). The removed case now reads "a step no longer declares an agent launch this audit reads". While the workflow still exists, the row adds: "a step that no longer declares one may still start an agent in a way this audit does not read, such as an action outside its table, npx, a script, a path such as ./node_modules/.bin/claude or codex options before exec, so this row does not say that it no longer starts one". A removed workflow gets neither clause.
| base → head | 9e764284 why |
this head why |
|---|---|---|
claude -p --allowedTools Read 'Review' → npx @anthropic-ai/claude-code -p --dangerously-skip-permissions 'Review' |
"a step no longer launches an agent (review/steps[0])…" | "a step no longer declares an agent launch this audit reads (review/steps[0])…; a step that no longer declares one may still start an agent…" |
codex exec -s workspace-write 'review' → codex -c sandbox_mode=danger-full-access exec 'review' |
"a step no longer launches an agent…" | same as the row above |
claude -p --allowedTools Read 'Review' → gh pr comment "$PR" --body "$(claude -p --dangerously-skip-permissions 'Review')" |
"a step no longer launches an agent…" | "an agent launch's declared settings changed (review/steps[0])…; a step launches an agent in a form this audit does not read…" (F1 now reads the launch) |
Every row stays changed/expands: false. test_a_read_launch_that_becomes_a_form_this_audit_does_not_read_is_not_called_gone covers npx, a path and codex root options, including the claude -p → npx @anthropic-ai/claude-code -p case the review asked for. test_a_read_launch_that_moves_into_a_quoted_substitution_is_changed_and_named_unread covers the third row. Known unread surfaces now lists codex with an option before exec (and bash -c), and says how such a transition is worded. STABILITY's "What is not read" says the same.
Nonblocking
- P3, codex root options: partly addressed.
codex [root options] execis now a named example under Known unread surfaces and in STABILITY, and the removal wording above names it. I did not make it a read or unresolved launch. Findingexecafter root options needs to know which root flags take a value. I have not checked that against the CLI's own parser, and a guess could produce a false read. I can open this as a follow-up. - P3, literal words before
${{ }}inrun:: not changed in this cycle. A launch that stays unpublished but still carries rules read from its literal prefix would be a new state.agent_widening_rulesreads rules only fromform: readlaunches, and the unread-before and moved logic depends on that. The support page already documents the current behaviour, where such a run isunresolved(expression) and the row ischanged. This fits better as its own change. - P3, a move into a broader job: not changed in this cycle. Treating a launch that moved between jobs as moved, not gained, follows the A workflow step's action reference is unread: moving
uses:from a pinned SHA to@mainproduces no row #771 step-reference precedent that the support page documents. The row'swhysays the launch "now runs with the receiving job's token permissions", and the note names that job's write scope. Pairing a move only when the receiving job is no broader changes direction semantics, including how default and unknown permission contexts compare. That needs its own doc and test pass.
Tests run
tests/test_workflow_agent_launches.py(all pass).- Together with
test_host_boundary_unread_surfaces,test_distribution_surface_parity,test_host_diff_entry_docs,test_host_diff_review_changes,test_capability_diff,test_workflow_capability_diff,test_workflow_label_redaction,test_workflow_step_action_references,test_reusable_workflow_secret_mappings,test_unread_changed_inputs,test_host_audit,test_public_surface_contract,test_local_contract,test_instruction_structure_contracts,test_host_input_recovery,test_agent_instructions_apply,test_agent_instructions_renderers,test_org_governance,test_governance_benchmark,test_governance_benchmark_baseline,test_benchmark_results_privacy,test_regenerate_goldens,test_miner_corpus,test_schema_boundaries,test_schema_roundtripandtest_capability_change_schema_hash_parity: 1813 passed, 4 skipped. The benchmark replays reproduce unchanged. scripts/generate_schemas.py --checkis clean.ruff checkon the changed files is clean.
pengfei-threemoonslab
left a comment
There was a problem hiding this comment.
Independent coding-agent review/address cycle 3 at b1073829.
This follows the cycle-2 review at 9e764284 and its address comment.
Not mergeable yet: one P2 finding. Both cycle-2 findings are fixed at this head (evidence below), and CI is green. The new finding is a regression from the cycle-2 fix: text inside a quoted here-doc is now read as an agent launch.
Setup for every reproduction:
- Synthetic two-commit repositories built with the issue's
pairhelper. ./shipgatefrom detached worktrees of this head, of9e764284and oforigin/main997e9260.W=.github/workflows/agent.yml,on: pull_request, and onereviewjob withcontents: readandpull-requests: write.
- P2: a quoted here-doc that mentions
claude -pin Markdown backticks is now listed as an agent launch. That adds a false row, and it stops a real widening in the same job from being claimed.- Evidence. The job has a step that only writes a Markdown file:
With a quoted delimiter, the shell does not expand the backticks, so this step starts no agent.
- run: | cat > comment.md <<'EOF' Reproduce locally with `claude -p "review this change"`. EOF
- Adding only that step. On this head the row is
high changedand reads "a step now launches an agent (review/steps[1]) … a step launches an agent in a form this audit does not read (review/steps[1]) … an agent runs at review/steps[1] (claude -p)".audit --hostlistsreview/steps[1]asunresolved/compound_commandand adds a coverage issue for it. On9e764284: no row. - Adding a real launch beside it. The head adds a step
run: claude -p --dangerously-skip-permissions "Review". On this head the row ishigh changed,expands: false, and reads "an agent launch now skips permission checks (bypassPermissions) (review/steps[2]), which is not counted as a widening: before, this job launched the agent in a form this audit does not read, which may already have done the same". On9e764284the same pair is⚠ high widened, and so is this head without the here-doc step. So the Stop hook stays quiet about a real gain, and the row's sentence about the base is false: before, the job launched no agent. <<"EOF",<<\EOFand<<-'EOF'give the same result. So doescodex execin backticks.
- Adding only that step. On this head the row is
- Cause:
_command_substitutions(core/host_grants.py) scans the wholerun:for backticks and$(, outside single quotes. It does not know about here-doc bodies._run_launchesthen lists what the scan finds as an unresolved launch.agent_rule_gainstreats a job whose only earlier launch was unresolved asunread_before.
- Why it matters:
- Markdown backticks are a common way to mention
claude -pin a PR comment or step summary, and such text is often written through a quoted here-doc. - This contradicts the support page: "A command that launches no headless agent … is not listed", and "listed as
unresolved… once for each agent CLI it starts". - Two related shapes already behaved this way at
9e764284: a here-doc body line that starts withclaude -p, and$(claude -p …)inside a quoted here-doc. The backtick form is new inb1073829.
- Markdown backticks are a common way to mention
- Fix:
- Do not read substitutions, or command heads, inside the body of a here-doc whose delimiter is quoted (
'EOF',"EOF",\EOF). Keep reading substitutions in an unquoted<<EOFbody, because the shell runs them. - Add
hd-style tests. One: adding the quoted here-doc step alone gives no row. Two: beside it, a new--dangerously-skip-permissionsstep iswidened. - If that is out of scope, the smallest fix is: skip quoted here-doc bodies in
_command_substitutionsonly, which removes the regression. Then say on the support page that other text in a here-doc body can be listed as a launch.
- Do not read substitutions, or command heads, inside the body of a here-doc whose delimiter is quoted (
- Evidence. The job has a step that only writes a Markdown file:
Non-blocking (P3):
- A launch that becomes a form the audit does not read still counts as having left its job, for the "moved between jobs" rule. This is the same gap cycle-2 F2 closed for the removal wording.
- Evidence. Job
a(contents: read) goes fromclaude -p --dangerously-skip-permissions "x"tonpx @anthropic-ai/claude-code -p --dangerously-skip-permissions "x". In the same change, jobb(contents: write) goes fromclaude -p --allowedTools Read "x"toclaude -p --dangerously-skip-permissions "x". - The result is
high changed: "an agent launch that skips permission checks (bypassPermissions) moved between jobs (a/steps[0] → b/steps[0]), which is not counted as a widening". - In fact job
astill runs its launch, and jobb's bypass is new. job_leftinagent_rule_gainsreads "no longer launches that agent" from read and unresolved launches only.- This needs two edits in one change, so it is rare. Pairing a move only when the losing job is gone, or when its agent step is gone, would close it.
- Evidence. Job
- The PR description is stale for the last commit.
- It says
tests/test_workflow_agent_launches.pyhas 210 cases. It now has 234. - "Run locally on the final tree" and "CI on" both name
9e764284. - The Design section's
run:bullet does not mention reserved words or substitutions. - The address comment has the deltas. The description should be refreshed before merge.
- It says
- A checkout in an added job is worded "a checkout's declared ref changed". For example, on this repository's own
ci.ymlfrom base8d43106f, the newmcp-extrajob's default checkout adds "a checkout's declared ref changed (mcp-extra/Checkout)". This matches #771's "a step's action reference changed" wording for added steps, so it is only a wording nit. - Still open from cycle 2, as documented:
- a
${{ }}inrun:makes the whole launchunresolved codex [root options] execis not read- a move into a broader job is not claimed
- a
Verified:
-
CI on
b1073829.suite (1/2/3),coverage,test,verify,verify-self,mcp extra (floor/newest),clean-checkout-launcherandwindows-launcherpass.release-tag-consistencyis skipped, as on every PR. GitHub reports the PRMERGEABLEagainstorigin/main997e9260, which is still the tip. -
C2R-F1 fixed. From a base of
run: echo done:gh pr comment "$PR" --body "$(claude -p --dangerously-skip-permissions 'Review this change')"is alow changedrow, "a step now launches an agent … in a form this audit does not read".audit --hosthasunresolved/shell_expansionand a non-blocking issue namingreview/steps[0].if [ -n "$PR" ]; then claude -p …; fiis achangedrow,compound_command.- An edit inside the quoted substitution gives no row, and the named limit is kept.
echo "claude -p …",echo '$(claude -p x)',echo "\$(claude -p x)"andecho "$(date) claude -p x"are not listed.
-
C2R-F2 fixed.
claude -p --allowedTools Read→npx @anthropic-ai/claude-code -p --dangerously-skip-permissionsreads "a step no longer declares an agent launch this audit reads … may still start an agent in a way this audit does not read … so this row does not say that it no longer starts one". It stayschanged.codex exec -s workspace-write→codex -c sandbox_mode=danger-full-access execreads the same.- The read launch that moves into a quoted substitution reads "an agent launch's declared settings changed … in a form this audit does not read".
-
Mutation checks on the cycle-2 fixes. In a scratch worktree I made four changes, one at a time:
- dropped
_substituted_agentsfrom_run_launches - dropped
_reserved_prefix_lengthfrom_agent_command - restored the old removal wording
- dropped substitutions from the unbalanced-quoting fallback
Each one fails
tests/test_workflow_agent_launches.py. - dropped
-
The issue reproduction, on
997e9260and this head.- On
997e9260,args,runandcheckoutgive no row, andtriggergives a row with no agent context. - On this head:
argsis⚠ low widened"skips permission checks (bypassPermissions) (review/steps[1])".runislow changednamingreview/steps[2]and both values.checkoutiscritical changedwith the default ref →${{ github.event.pull_request.head.sha }}, and the note namespull_request_targetand the pull request checkout.triggernamesreview/steps[1]besideissue_commentandpull-requests.
checkcarries theargsrow aswidenedand therunrow aschanged(0 rows onmain).decisionisrequire_reviewon both engines.
- On
-
Earlier findings, re-run at this head.
- Attached
--settings='{…env…}'/--mcp-config='{…headers…}'canaries do not appear indiff --json. - Job
review→code-reviewwith the same bypassing launch ischangedwith the move named. - Prose "Never print bearer tokens" beside
pull-requests: read→writeis⚠ high widened.
- Attached
-
A realistic
claude-code-actionworkflow. It usesissue_comment/issues,id-token: writeandclaude_code_oauth_token. Itsclaude_args: |gains--dangerously-skip-permissions, and an--mcp-configholding aghp_…token and anenvvalue.diff, manifest-freeverify(pr-comment.md,verifier.json),audit --hostandcheckcarry the⚠ high widenedrow with the note.- None of them contains either canary.
-
Regression check on this repository.
difffrom basesb61aca78,5c4c62f8and8d43106fgives the same row count, subjects and directions on both engines. Onlywhygains the checkout-ref clause, and inrelease.ymlit reports real ref changes (needs.verify→needs.candidate). -
Fuzzing. About 890,000 random shell-like inputs to
_run_launchesand_command_substitutionsraised no exception. -
Local checks at this head.
scripts/generate_schemas.py --checkis clean, andruff check src tests scripts shipgateis clean.- Tests: 36 files, 2215 passed, 4 skipped, 0 failed. They are the new test file; the workflow step-reference, reusable-secret, label-redaction and capability-diff tests; distribution-surface parity; host audit; the public surface and local contract; the pilot ledger; unread surfaces and unread changed inputs; schema round-trip, boundaries and hash parity; the control envelope and its rows; the cold-start, host-config and semantic cold-start replays; the governance benchmark and its baseline; enabled plugin hook routing; install hooks; manifest-free PR rows; preflight; the host-diff docs, review-change and permission-direction tests; the instruction, recovery, agent-instruction and org-governance tests; and golden regeneration.
b107382 to
17f68e2
Compare
The second review of #850 at 25c13ce found two P1 and four P2 defects. The branch is rebased onto origin/main daa4ad5. Attached JSON values are withheld (F1). A --settings={...} or --mcp-config={...} word inside claude_args, and codex's attached -c<override> / -c=<override>, published their env, header and apiKeyHelper values verbatim, because only a word starting with "{" was withheld. _withheld_words now splits a --name=value word and withholds the value, and reads codex's attached -c through _withheld_config, as clap reads it. The canary sweep carries both spellings. Redacted prose no longer refuses the comparison (F2). The #802 label redaction rewrites ordinary prose ("never print bearer tokens", "Authorization: headers"), and a redacted agent setting made the workflow a blocking limit. That hid every row beside it and made every check incomparable while the workflow existed. The rules are already read from the declared text, so a redacted setting is now compared by its published text and its widening_rules. uncompared_agent_launch_texts names it as a non-blocking limit, and the row cell shows the redacted text. A redacted checkout ref still refuses, as a redacted step reference does (#767), because it names the code a job runs. A renamed job's launch moves its rules (F3). _agent_rule_gains keyed a rule on its job, so renaming a job that launches a bypassing agent was a widening. agent_rule_gains now pairs a rule one job gains with the same rule another job lost, when the launch that met it left that job: the job no longer launches that agent, or the same launch (agent_launch_key less the job) now runs in the gaining job. The why names the move. A second job gaining a rule, or a different launch gaining one while the first job still launches that agent, still widens. An expression no longer turns off every rule, and the row says what it leaves unread (F4). Rules are read from literal text a ${{ }} expression cannot reach: - the words of claude_args or codex-args before the first expression, less the word it touches and any quoted run still open at it; - the elements of a JSON-array codex-args before the one holding it; - the gate entries that hold none. A setting holding an expression is published with holds_expression (the unreleased 0.7 schema extends in place), and a row that changes it says the text the expression reaches is not read. A gain where the job's launch held an expression before, in the input the rule is read from, is named and not claimed, as unknown_before is. That also fixes a false widening at 25c13ce: replacing --model ${{ vars.CLAUDE_MODEL }} with --model opus beside --dangerously-skip-permissions was reported as gaining the bypass. Two non-blocking fixes: - _published_value reads each expression as one word, so an expression in a URL's userinfo is withheld with it rather than garbling the URL and publishing its path. - A codex --config value that starts like a table and does not parse, as string-argv leaves a quoted one, is withheld as unparsed_json. Rebase (F5). CHANGELOG keeps #853's #778 line beside #823's under github_action row. llms.txt and ai-search-summary state the source tree as contract 41, unreleased, ahead of the published v1.1.0 (contract 40), as test_public_surface_contract requires while the two differ. The pilot ledger's source-tree column was re-taken on the rebased tree, through ./shipgate beside the engine of e3c6cb0, the commit v1.1.0 was cut from: - the only differences are contract 40 -> 41 and inventory schema 0.6 -> 0.7; - diff rows are byte-identical; - check JSON differs only in the launcher path its next action names. The support page, the STABILITY preamble and migration note, the CHANGELOG entry, agent-contract-current, the Stop hook sentence in integrations, the capability_diff row in distribution-surfaces with its parity comment, the 0.7 schema files and llms-full.txt are updated to match.
The third review of #850 at c7f6755 found one P1 and one P2 defect and four P3 notes. The branch is on origin/main 44b9e05; no rebase was needed. A JSON setting publishes its shape, not its free text (C2-F1). _withheld_json published the whole _redact_secret_values tree. That tree is the host readers' digest input, not what they publish: it keeps every string outside env, headers and secret-named keys. So an mcp-remote --header "Authorization: Bearer ..." argument in --mcp-config, and a hook's curl command in settings, reached diff text, diff --json, audit --host --json, the PR comment and verifier.json. _json_shape now keeps key names, numbers, booleans and null, redacts what the host readers redact, and replaces each other string with <withheld:...>, a 12-hex digest of redacted_config_sha256 for that string, so an edit to it is still a changed row. It keeps only the strings a host reader publishes: - a permissions.allow/ask/deny rule and a documented Claude Code setting's value (defaultMode, the switches, enabledMcpjsonServers entries); - an MCP server's command name and its URL's scheme and host, followed by the digest when they drop a command's arguments or a URL's query. A codex --config table or array is read under its key path, so mcp_servers.gh={command="gh", ...} keeps its command name. The canary sweep adds the mcp-remote header and the hook command in every spelling (action input, claude_args, CLI flag, codex -c table). Re-running the reviewer's two repositories through ./shipgate gives 0 canaries in every output. The documented bypasses written through listed inputs widen (C2-F2). - Claude Code settings written as JSON meet bypass_permissions when their defaultMode is bypassPermissions, read by claude_setting_values as the settings reader reads .claude/settings.json. This covers the action's settings input, a --settings value in claude_args, and the CLI's --settings flag. A path is not read, and a settings value holding an expression meets none. - openai/codex-action's permission-profile: :danger-full-access meets danger_full_access. ":danger-full-access" is Codex's reserved name for its built-in full-access profile (BUILT_IN_PERMISSION_PROFILE_DANGER_FULL_ACCESS). Mode inputs are now (input, value, rule) triples. settings and permission-profile join _RULE_SETTINGS, so a gain after an expression in either is named and not claimed, as for claude_args. P3 notes: - allowed_bots: "*" now reads "accepts runs triggered by any bot" through agent_rule_text. - Which of two gaining jobs a moved rule goes to no longer depends on declaration order: the same launch arriving is matched before a job that merely stopped launching the agent. The support page, the STABILITY migration note, the contract summary, the schema docstrings and the CHANGELOG entry now say what a structured value publishes instead of claiming it publishes what the host readers would, and list the two rules. host-grants 0.7 is extended in place; it is unreleased.
|
Addressed review cycle 3. New head
C3-F1 (P2): a quoted here-doc that mentions
|
| base → head | b1073829 |
846763fd |
|---|---|---|
| adds the review's quoted here-doc step | high changed, "a step now launches an agent (review/steps[1])"; 1 GitHub coverage issue |
No static host-grant changes detected; no issue |
same with <<"EOF", <<\EOF, <<-'EOF' (backticked codex exec), a claude -p body line, or $(claude -p …) in a quoted body |
high changed; 1 issue |
no row; no issue |
here-doc step in the base; head adds claude -p --dangerously-skip-permissions "Review" |
high changed, expands: false, "which is not counted as a widening: before, this job launched the agent in a form this audit does not read" |
⚠ high widened, expands: true, "an agent launch now skips permission checks (bypassPermissions) (review/steps[2])" |
adds cat <<EOF / $(claude -p x) / EOF (unquoted, so the shell runs it) |
high changed; 1 issue |
same row; 1 issue |
adds cat <<EOF / claude -p x / EOF (a body line is input) |
high changed; 1 issue |
no row; no issue |
I also checked 29 run: texts against bash, with stub claude and codex functions that log when they run. Besides the shapes above, they include:
- two bodies on one line;
<<in arithmetic, in quotes, in a comment and in a here-string;- a here-doc inside a quoted substitution;
- an escaped
\$(…)and a double-quoted"$(…)"in an unquoted body; - a here-doc in a
casearm.
The reader lists an agent exactly where bash runs one, with two exceptions. In both it still lists one where bash runs none:
- a here-doc with no closing line;
- a here-doc opened inside another here-doc's body.
The support page names both: "A here-doc with no such line, or one opened inside another here-doc's body, is read as ordinary lines, so its body text can still be listed as a launch." It also now says a here-doc's body is never a command, and which substitutions in it are read.
Tests (tests/test_workflow_agent_launches.py):
test_a_quoted_here_doc_that_mentions_an_agent_cli_launches_none: 6 delimiter forms × 5 bodies, each giving no launch and no row.test_beside_a_quoted_here_doc_a_new_bypass_step_is_a_widening: the review's case (2).test_an_unquoted_here_doc_runs_its_substitutions_and_none_of_its_lines.test_an_agent_cli_outside_a_here_doc_body_is_still_named: 9 shapes.test_many_here_docs_are_read_in_one_pass.test_the_here_doc_of_the_review_adds_no_row_or_coverage_issue_and_the_bypass_widens:diffJSON and text, plusaudit --host, end to end.
With b1073829's host_grants.py restored, 30 of the new here-doc cases fail. The two existing cases that used a here-doc body with an apostrophe to reach the unbalanced-quoting path now pass without it, so I added two unbalanced cases that are not here-docs to keep that path covered.
Nonblocking
- P3, move between jobs: fixed. One detail beyond the review's description: in the exact reproduction, job b's new launch is identical to job a's old one, so the "same launch" pass matched it before
job_leftran. Both passes are narrowed:job_leftnow needs the losing job to no longer exist (a rename). A job that remains may still run its launch throughnpxor a path.same_launchnow also needs each launch of that agent the gaining job had to still run there or in the losing job, which holds for a moved step or a swap. So a launch edited in place into the one the other job had gains the rule.- The reproduction is now
⚠ high widened, "an agent launch now skips permission checks (bypassPermissions) (b/steps[0]); a step no longer declares an agent launch this audit reads (a/steps[0])". Onb1073829it washigh changed, "moved between jobs (a/steps[0] → b/steps[0])". - A step moved and edited while its job remains is now claimed as well. That is the direction the open "a move into a broader job is not claimed" note prefers.
- Rename, step-moved, rename-and-edited, swapped and gate-moved stay
changedwith the move named. test_a_rule_another_job_gains_while_no_launch_left_is_a_wideningadds the mixed-form case, the review's exact case and the moved-and-edited case. The support page bullet, STABILITY's Direction bullet, the CHANGELOG entry and theAgentRuleGainsdocstring state the narrowed rule.
- P3, stale PR description: fixed. The description now:
- counts 289 cases;
- names this head and
eff60d97for the local runs; - reports CI on this head;
- lists reserved words, substitutions and here-docs in the Design
run:bullet and in the unsupported-shapes bullet; - states the narrowed move rule;
- carries the before/after table above.
- P3, checkout in an added job: fixed for the new sentence. A checkout step on one side only is now "a step now declares a checkout (lint/steps[0])" or "a step no longer declares a checkout (…)". A changed ref at the same step keeps "a checkout's declared ref changed". For a job added with a default
actions/checkout@v4, the shape of this repository'sci.ymljobmcp-extra, the row now reads "a step's action reference changed (lint/steps[0]); …; a step now declares a checkout (lint/steps[0]); …". I left A workflow step's action reference is unread: movinguses:from a pinned SHA to@mainproduces no row #771's "a step's action reference changed" sentence as it is: it shipped in 1.1.0, and rewording it belongs in its own change.test_a_checkout_step_on_one_side_only_is_worded_as_added_or_removedcovers both directions. - P3, open from cycle 2 (
${{ }}inrun:,codex [root options] exec, a move into a broader job): unchanged and still documented. The move-rule change above treats fewer gains as moves, so it claims more of them as widenings, but it still does not compare the two jobs' permissions.
Tests run
tests/test_workflow_agent_launches.py: 289 passed.- 36 other test files, all passing on
846763fd:test_distribution_surface_parity,test_host_boundary_unread_surfaces,test_reusable_workflow_secret_mappings,test_unread_changed_inputs,test_workflow_step_action_references;test_agent_boundary,test_agent_control_envelope_rows,test_capability_diff,test_capability_diff_partial_clone,test_determinism_boundary,test_docs_links;test_host_audit,test_host_comparison_coverage,test_host_diff_entry_docs,test_host_diff_permission_direction,test_host_diff_review_changes,test_install_hooks;test_local_contract,test_public_surface_contract,test_release_decision,test_schema_roundtrip,test_workflow_capability_diff,test_workflow_label_redaction,test_workflow_evidence,test_instruction_structure_workflow;test_preview_control_currency,test_agent_control_reports_dir,test_design_partner_pilot, and Keep independently established host changes when one scope is incomparable (#808) #861'stest_partial_host_comparisonandtest_claude_hook_loading_evidence;- the replays and benchmarks:
test_host_config_replay,test_cold_start_replay,test_semantic_cold_start,test_governance_benchmark,test_governance_benchmark_baseline,test_benchmark_results_privacy. The benchmark runs of record reproduce unchanged.
- The one agent-launch workflow change in the 23-PR corpus (ci: replace the title-similarity duplicate bot with a Codex semantic check BerriAI/litellm#40935, its two blobs re-fetched as text) gives byte-identical rows on
b1073829and846763fd. scripts/generate_schemas.py --checkis clean.scripts/regenerate_goldens.py --checkreports 24 artifacts and 0 changed.ruff checkon the changed files is clean. Rebuildingllms-full.txtchanges nothing.- CI on
846763fd:suite (1/2/3),coverage,test,verify,verify-self,mcp extra (floor/newest),clean-checkout-launcherandwindows-launcherall pass;release-tag-consistencyis skipped, as on every PR. GitHub reports the PRMERGEABLE,CLEANagainst8269922b.
The workflow grant read triggers, token permissions, reusable-workflow secrets (#693) and step `uses:` references (#771), and nothing about how a coding agent is launched inside a job. Changing a claude-code-action's `claude_args` from `--allowedTools "Read"` to `--permission-mode bypassPermissions --allowedTools "Bash(*)"`, adding a `claude -p --permission-mode acceptEdits` run step, or checking out the pull request head in a `pull_request_target` job each gave "No static host-grant changes detected", and a move to `issue_comment` with `pull-requests: write` gave a row with no agent context. The workflow grant now lists, as text that is never executed, fetched or evaluated: - `agent_launches[]`: a step whose `uses:` is anthropics/claude-code-action, anthropics/claude-code-base-action or openai/codex-action (any ref, any case) with the documented permission inputs it sets, or a `run:` that is one literal simple command starting with `claude -p/--print` or `codex exec`, with its documented permission flags under their primary spelling. The prompt, `--model` and undocumented flags are not compared. `job_secrets` names the secrets the job references, as context only. - `checkout_refs[]`: each actions/checkout step's `with.ref`, null for the default. Each job's multiset of launches and refs is compared, never the step label, so a rename or reorder is quiet; a difference is one `changed` row on the existing workflow row naming `job/step` and both values. Direction is claimed only by documented rules a job's launches gain, read from literal values: bypassed permission checks (either spelling), bypassed approvals and sandbox, a danger-full-access sandbox, `safety-strategy: unsafe`, or a `*` user gate. Those raise `workflow_agent_widened_<added|changed>` and make the row widened; every other edit, including `--allowedTools "Bash(*)"` (#824's to rate), is `changed`. A workflow row whose workflow runs an agent ends its `why` with the untrusted-input trigger, write scopes, secrets and pull request checkout beside each agent step; it is a note and moves no direction. A compound `run:`, an expansion or an expression is `unresolved`, publishes none of its text and records a non-blocking coverage issue naming `job/step`, as an unread secret value does (#693): coverage stays complete, adding one is a row that claims no effect, and an edit inside one is not reported. Scripts, composite (#701) and unknown actions, and agents reached through npx/timeout/sudo are listed as unread surfaces. Every value, ref and secret name goes through the #802 label redaction; a rewritten one is null with `redacted` and is neither published nor compared. Contract 40 and host-grants 0.6 shipped in 1.1.0, so host-grants inventory, baseline and drift move to 0.7 (the 0.6 schema files are untouched) and the runtime contract to 41. A 0.4-0.6 baseline holding a workflow grant is incomparable (`baseline_workflow_agent_launches_unavailable`); one without a workflow stays comparable. Verifier 0.20 and capability diff 0.3 do not move. No check id is added or removed and `check` decides as before. Docs: the support page (tables, rules, note, limits, a new unread-surfaces bullet), a STABILITY "Migration Note: Unreleased", CHANGELOG `## Unreleased` above 1.1.0, the agent contract page, the distribution-surfaces `capability_diff` row and its parity comment, the Stop hook note, version tables and pins, and a rebuilt llms-full.txt. The pilot ledger's source-tree column was re-measured on the Route H fixture: identical cells to the 1.1.0 engine apart from contract 41 and inventory 0.7, with byte-identical diff rows. Tests: new tests/test_workflow_agent_launches.py covers the four reproduction cases, what is and is not read, every unsupported shape, direction rules and non-rules, the acceptance's negative controls, redaction with a CLI canary sweep across every published output, the 0.6 baseline migration, schema validation, and the same row on diff, verify, the PR comment, check, the control envelope and the Stop hook. Existing tests move to 0.7/41. The host-config and cold-start replays reproduce their committed outcomes and run-of-record scores, and the sample goldens are unchanged. Closes #823
) The support page and STABILITY say a documented rule is read from literal values only, so a value holding `${{ }}` never meets one. The `*` user gate did not apply that: `allowed_non_write_users: "${{ vars.USERS }}, *"` raised workflow_agent_widened_changed. Every rule now skips a value holding an expression, as the claude_args and codex-args rules already did, and a test pins the gate case.
F1. `claude_args` and `codex-args` were split with the `run:` shell tokenizer, which gave up on a newline or an unquoted `(`, so a widening written the way the actions document it gave `changed`: `claude_args: |` on several lines, `--allowedTools Bash(git:*) --dangerously-skip-permissions`, a `# comment` line, or a multi-line `codex-args`. Each input is now split the way its action splits it. The Claude actions' parse-sdk-options.ts drops full `#` lines, makes `()|&;<>` literal and reads the rest with shell-quote (newlines are whitespace, `$NAME` is empty, an unquoted `#` ends the input, a `--` word is always a flag); codex-action reads a JSON array of strings or string-argv. The `run:` tokenizer is kept for `run:` steps only. The published `claude_args` is the text the action parses, so a full-line comment is neither published nor compared. F2. Any URL path made a whole setting `redacted`, so a `--dangerously-skip-permissions` beside `https://example.com/style-guide` gave no row, and every `plugin_marketplaces` value was never compared, under a limit that wrongly called it credential-shaped. The documented rules are now decided from the declared text when the workflow is read, before anything is withheld, and published on the launch as `widening_rules` (rule and setting), which the comparator keys on, so redaction never hides a rule. A URL publishes its scheme and host with `<redacted-path>`, as an MCP server URL does (#723), and the rest of the value is published and compared. Other credential-shaped text in a setting or checkout ref (a token shape, an assignment, a bearer or header value, URL userinfo) is published redacted and makes the workflow a blocking limit through `_uncompared_workflow_text`, as a redacted step reference does (#767). F3. JSON in `settings`, `mcp_config`, `--settings`, `--mcp-config` or any argument word was published verbatim, env values and apiKeyHelper included. A JSON object now publishes what `.claude/settings.json` and `.mcp.json` publish: key names, with `env`/`headers` values, apiKeyHelper and secret-named values `<redacted>`, as canonical JSON. A codex `--config` override under env, headers or a secret-named key publishes `<redacted>`. Text that starts like JSON and does not parse is withheld (`unparsed_json`, a non-blocking limit). The agent-launch canary sweep now carries JSON-shaped canaries and their digests across diff, audit, check, verify, the PR comment and verifier.json, and a second sweep covers the refused credential-shaped case. Nonblocking: a rule gained where the job launched that agent before only in an unread form is named and not claimed (the `unknown_before` rule); a quoted word starting with `#` no longer makes a `run:` compound; `anthropics/claude-code-action/base-action` is read as the base action; the STABILITY note says a prompt after a variadic flag is compared; and `__all__ =[` is spaced. Host-grants 0.7 is unreleased, so `widening_rules`, `unparsed_json` and the base-action agent value extend it in place; the 0.7 schema files are regenerated. The support page, STABILITY migration note, CHANGELOG Unreleased entry, contract summary, integrations Stop-hook sentence and the capability_diff distribution-surface row say the same.
The second review of #850 at 25c13ce found two P1 and four P2 defects. The branch is rebased onto origin/main daa4ad5. Attached JSON values are withheld (F1). A --settings={...} or --mcp-config={...} word inside claude_args, and codex's attached -c<override> / -c=<override>, published their env, header and apiKeyHelper values verbatim, because only a word starting with "{" was withheld. _withheld_words now splits a --name=value word and withholds the value, and reads codex's attached -c through _withheld_config, as clap reads it. The canary sweep carries both spellings. Redacted prose no longer refuses the comparison (F2). The #802 label redaction rewrites ordinary prose ("never print bearer tokens", "Authorization: headers"), and a redacted agent setting made the workflow a blocking limit. That hid every row beside it and made every check incomparable while the workflow existed. The rules are already read from the declared text, so a redacted setting is now compared by its published text and its widening_rules. uncompared_agent_launch_texts names it as a non-blocking limit, and the row cell shows the redacted text. A redacted checkout ref still refuses, as a redacted step reference does (#767), because it names the code a job runs. A renamed job's launch moves its rules (F3). _agent_rule_gains keyed a rule on its job, so renaming a job that launches a bypassing agent was a widening. agent_rule_gains now pairs a rule one job gains with the same rule another job lost, when the launch that met it left that job: the job no longer launches that agent, or the same launch (agent_launch_key less the job) now runs in the gaining job. The why names the move. A second job gaining a rule, or a different launch gaining one while the first job still launches that agent, still widens. An expression no longer turns off every rule, and the row says what it leaves unread (F4). Rules are read from literal text a ${{ }} expression cannot reach: - the words of claude_args or codex-args before the first expression, less the word it touches and any quoted run still open at it; - the elements of a JSON-array codex-args before the one holding it; - the gate entries that hold none. A setting holding an expression is published with holds_expression (the unreleased 0.7 schema extends in place), and a row that changes it says the text the expression reaches is not read. A gain where the job's launch held an expression before, in the input the rule is read from, is named and not claimed, as unknown_before is. That also fixes a false widening at 25c13ce: replacing --model ${{ vars.CLAUDE_MODEL }} with --model opus beside --dangerously-skip-permissions was reported as gaining the bypass. Two non-blocking fixes: - _published_value reads each expression as one word, so an expression in a URL's userinfo is withheld with it rather than garbling the URL and publishing its path. - A codex --config value that starts like a table and does not parse, as string-argv leaves a quoted one, is withheld as unparsed_json. Rebase (F5). CHANGELOG keeps #853's #778 line beside #823's under github_action row. llms.txt and ai-search-summary state the source tree as contract 41, unreleased, ahead of the published v1.1.0 (contract 40), as test_public_surface_contract requires while the two differ. The pilot ledger's source-tree column was re-taken on the rebased tree, through ./shipgate beside the engine of e3c6cb0, the commit v1.1.0 was cut from: - the only differences are contract 40 -> 41 and inventory schema 0.6 -> 0.7; - diff rows are byte-identical; - check JSON differs only in the launcher path its next action names. The support page, the STABILITY preamble and migration note, the CHANGELOG entry, agent-contract-current, the Stop hook sentence in integrations, the capability_diff row in distribution-surfaces with its parity comment, the 0.7 schema files and llms-full.txt are updated to match.
The third review of #850 at c7f6755 found one P1 and one P2 defect and four P3 notes. The branch is on origin/main 44b9e05; no rebase was needed. A JSON setting publishes its shape, not its free text (C2-F1). _withheld_json published the whole _redact_secret_values tree. That tree is the host readers' digest input, not what they publish: it keeps every string outside env, headers and secret-named keys. So an mcp-remote --header "Authorization: Bearer ..." argument in --mcp-config, and a hook's curl command in settings, reached diff text, diff --json, audit --host --json, the PR comment and verifier.json. _json_shape now keeps key names, numbers, booleans and null, redacts what the host readers redact, and replaces each other string with <withheld:...>, a 12-hex digest of redacted_config_sha256 for that string, so an edit to it is still a changed row. It keeps only the strings a host reader publishes: - a permissions.allow/ask/deny rule and a documented Claude Code setting's value (defaultMode, the switches, enabledMcpjsonServers entries); - an MCP server's command name and its URL's scheme and host, followed by the digest when they drop a command's arguments or a URL's query. A codex --config table or array is read under its key path, so mcp_servers.gh={command="gh", ...} keeps its command name. The canary sweep adds the mcp-remote header and the hook command in every spelling (action input, claude_args, CLI flag, codex -c table). Re-running the reviewer's two repositories through ./shipgate gives 0 canaries in every output. The documented bypasses written through listed inputs widen (C2-F2). - Claude Code settings written as JSON meet bypass_permissions when their defaultMode is bypassPermissions, read by claude_setting_values as the settings reader reads .claude/settings.json. This covers the action's settings input, a --settings value in claude_args, and the CLI's --settings flag. A path is not read, and a settings value holding an expression meets none. - openai/codex-action's permission-profile: :danger-full-access meets danger_full_access. ":danger-full-access" is Codex's reserved name for its built-in full-access profile (BUILT_IN_PERMISSION_PROFILE_DANGER_FULL_ACCESS). Mode inputs are now (input, value, rule) triples. settings and permission-profile join _RULE_SETTINGS, so a gain after an expression in either is named and not claimed, as for claude_args. P3 notes: - allowed_bots: "*" now reads "accepts runs triggered by any bot" through agent_rule_text. - Which of two gaining jobs a moved rule goes to no longer depends on declaration order: the same launch arriving is matched before a job that merely stopped launching the agent. The support page, the STABILITY migration note, the contract summary, the schema docstrings and the CHANGELOG entry now say what a structured value publishes instead of claiming it publishes what the host readers would, and list the two rules. host-grants 0.7 is extended in place; it is unreleased.
Rebased onto origin/main 0177703 (#852, #821), which had already moved the unreleased runtime contract 40 -> 41, verifier 0.20 -> 0.21 and capability diff 0.3 -> 0.4. This change now extends contract 41 in place instead of minting it: one comment in schemas/contract.py names #821 and #823, and the texts that said verifier 0.20 and capability diff 0.3 were unchanged now say capability_diff registry row keeps #852's two added roots and both paragraphs; STABILITY keeps the #821, #823 and #827 notes, with #821's "host-grants stays 0.6" and #827's version sentence corrected for a tree that also carries host-grants 0.7; CHANGELOG keeps both Unreleased entries. llms-full.txt is rebuilt. The pilot ledger's source-tree column was re-measured on the rebased tree beside e3c6cb0: identical cells except contract 40/41 and inventory schema 0.6/0.7, identical diff text, rows, check JSON and inventory apart from its schema version, and diff --json / verifier.json differing only in #821's schema versions and coverage members. A codex exec step that selects the full-access sandbox through --config was a changed row while --sandbox danger-full-access widened, and -sdanger-full-access was not read at all. The flag reader now reads a short flag's attached value as clap does (-s<mode>, -s=<mode>, -c<override>, -c=<override>), and a --config override that sets sandbox_mode to danger-full-access, or default_permissions to :danger-full-access (what codex-action's permission-profile input passes the CLI), meets danger_full_access, its value read as the CLI's parse_overrides reads it. As the CLI resolves them, the last override of a key counts, default_permissions outranks sandbox_mode, and a --sandbox flag outranks both, so an override beside --sandbox meets none. In codex-args such an override meets none, because the action appends its own --sandbox or default_permissions selection after codex-args. The support page and the STABILITY Direction bullet list the spellings. Withholding a quoted URL inside a JSON-array codex-args element no longer drops the element's escaped closing quote, so the published value stays the JSON array the support page describes.
An agent CLI inside a double-quoted $(...) or a backtick substitution, or
after a shell reserved word, got no launch, no row and no coverage issue:
gh pr comment --body "$(claude -p --dangerously-skip-permissions ...)" read
as "No static host-grant changes detected", while the support page said such
a run: is listed as unresolved once for each agent CLI it starts at the head
of a command. The word splitter reads a double-quoted substitution as one
word and a backtick as no operator, and the head reader took then, do, { and
! for the command.
_command_substitutions now lists the text of each $(...) and backtick
substitution outside single quotes, double-quoted ones included, and an agent
CLI heading a command inside one is an unresolved launch (shell_expansion for
a single top-level command, compound_command otherwise), publishing none of
its text. An escaped \$, a single-quoted '$(...)', a comment and $((...))
arithmetic are not substitutions, so echo "claude -p ..." stays unlisted. The
scanner keeps an explicit stack and reads each nested substitution as `_` in
the one around it, so no nesting depth recurses or re-splits text. When the
run's quoting does not balance, each line's substitutions are read as well.
A command's head is now read after shell reserved words (!, time [-p], {, if,
then, elif, else, while, until, do, function NAME), and a command after one
is not a simple command, so its launch is compound_command and never read.
The compound_command and shell_expansion limit phrases, the launch schema
description (host-grants 0.7 schemas regenerated), the support page and the
STABILITY bullet say so.
A read launch that became a form this audit does not recognise (npx, a path,
codex options before exec) was worded "a step no longer launches an agent",
though the step still starts one. The removed case now reads "a step no longer
declares an agent launch this audit reads", and a row whose workflow still
exists adds that the step may still start an agent in a way this audit does
not read, naming those forms. codex with an option before exec is listed under
Known unread surfaces and in STABILITY's "What is not read".
Tests cover each shape the review named, the negative controls, the diff and
audit --host route of the reproduction, the reworded removal for npx, a path
and codex root options, and 20000 nested substitutions.
A step that drafts a PR comment in a quoted here-doc, such as cat > comment.md <<'EOF' with "Reproduce locally with `claude -p ...`" in its body, was listed as an unresolved agent launch: a "high changed" row saying a step now launches an agent, and a GitHub coverage issue. The shell passes a here-doc's body to its command as input and, with a quoted delimiter, expands nothing in it, so the step starts no agent. The phantom also hid a real gain: a --dangerously-skip-permissions step added beside it was "not counted as a widening" because the job "launched the agent in a form this audit does not read" before. The substitution scanner read backticks and $( across the whole run:, and the word splitter read a body line starting with claude -p as a command head. _here_documents now removes each here-doc's body and closing line before the run: is split or scanned. <<, <<- and a quoted, escaped or partly quoted delimiter are read outside quotes, comments and $((...)) arithmetic, inside a $(...) or backtick substitution too; <<< is a here-string; bodies start after the opening line and follow in order when one line opens several. A body is never a command. Only an unquoted <<EOF body's substitutions, which the shell runs, are read, with quotes and # in it literal, so '$(claude -p ...)' there is listed and was missed before. A here-doc with no closing line is left in the text and read as ordinary lines, so a misread << hides no later command. Closing lines are looked up by content, so many here-docs cost one pass. Every shape was checked against bash with stub claude and codex functions. A launch that only became a form this audit does not read no longer counts as having left its job. The moved-between-jobs rule now takes a rule as moved only when the job that met it no longer exists, or when the same launch now runs in the gaining job and that job's own launches of the agent still run there or in the job it left (a moved step, a swap). Job a going from claude -p --dangerously-skip-permissions to npx @anthropic-ai/claude-code -p while job b's launch is edited into a bypassing one is now widened, where it was "moved between jobs". A step moved and edited while its job remains is claimed too. A checkout step on one side only, such as a default checkout in an added job, is worded "a step now declares a checkout" or "a step no longer declares a checkout" rather than "a checkout's declared ref changed". The support page, STABILITY's migration note and the CHANGELOG entry say which here-doc text is read, name the two shapes still read as lines, and state the narrowed move rule.
17f68e2 to
846763f
Compare
…nstead of refusing (#822) An instruction file whose limit this entry cannot resolve, such as a SKILL.md whose metadata holds a non-string value, is named as an unchanged limit when a change leaves it alone (#721). Reached through an in-tree link the reader reads through (#700) - `.claude/skills -> ../.agents/skills`, a per-skill link or a file link - the same untouched file refused the whole comparison: diff, verify and the manifest-free PR comment printed `base_inventory_incomplete; head_inventory_incomplete` and no row, hiding a removed deny rule beside it. `unchanged_limits` asked blob_path_unchanged whether the source, the path the link is read under, was one regular file in Git, and a path through a link never is. blob_path_unchanged now resolves the path on each side the way the reader reaches it: from Git tree entries for the base and a commit head, and for a working-tree head without following any link (each component's own entry, each link's own text, the file's unfiltered hash). It holds only when both resolutions are equal - every link at the same path with the same text, every other component a directory - and the file they land on is the same regular-file blob at the same in-tree path. A link is followed only under the rules the reader and the base archive already use (#700, #711): a relative text that lands inside the tree after normalization, directories above where it lands, at most eight links, and links at one component of the path. A path no link reaches still takes the one `ls-tree` it took before, and blob object IDs are still compared, so no filter or textconv can make two byte sequences equal. Everything that asks the proof moves together: unchanged_limits in `diff --json` and verifier.json, the text and PR comment, the unchanged limits a partial comparison may carry (#808) and the shared plugin-reference limits check leaves out (#714). check's boundary result still cannot name a limit, so it now refuses these comparisons with unchanged_limits_not_representable, as it does for a limit at its own path; its rows, decision and violations do not move. No schema, member, reason code or check id is added. The metadata value is not coerced. tests/test_linked_unchanged_limits.py holds every layout (directory link, per-skill link, file link, a chain of file links) to the direct result on diff, verify, the PR comment and check; keeps the negative controls refused (skill added or edited behind the link, link retargeted to an identical copy, link text rewritten to land on the same file, link replaced by a directory or the reverse, a later hop retargeted, working-tree-only retargets and edits); and holds the proof to exactly the links the reader reads through, across dangling, looping, absolute, escaping, over-long, intermediate and nested links. The #812 coverage case that pinned the refusal is retired, and the static-only allowlist follows its two call sites down the file.
pengfei-threemoonslab
left a comment
There was a problem hiding this comment.
Independent coding-agent review/address cycle 4 at 846763fd.
This follows the cycle-3 review at b1073829 and its address comment.
Not mergeable yet: one P2 finding. The cycle-3 finding is fixed at this head (evidence below). CI is green, and the PR is MERGEABLE/CLEAN against origin/main 8269922b.
The new finding is in the run: word splitter. It is the same kind of fault as the cycle-3 finding: text the shell never runs is read as a command, and text it does run is lost.
Setup for every reproduction:
- Synthetic two-commit repositories built with the issue's
pairhelper. ./shipgate diff --base HEAD~1andaudit --host --json, from detached worktrees of this head, ofb1073829and oforigin/main8269922b.W=.github/workflows/agent.ymlwithon: pull_request(or as noted), and one jobreviewwithcontents: read,pull-requests: writeand a checkout.
-
P2: a shell comment in
run:is split as if it were commands.- An apostrophe in a comment hides an agent launch: no row and no limit.
- A comment that shows a command line adds a false launch, and that false launch stops a real widening in the same job from being claimed.
Evidence (a): a launch is missed. The head adds this step:
- run: | # Don't run this on forks npm ci && claude -p --dangerously-skip-permissions "Review"
- On this head:
No static host-grant changes detected.audit --hostlists no launch and no coverage issue. - With
# Do not run this on forks:high changed, "a step now launches an agent (review/steps[1]) … in a form this audit does not read". The launch isunresolved/compound_command, with the non-blocking limit. - Two apostrophes pair up across lines, so a launch at the head of its own line is lost too. In an
issue_commentworkflow withcontents: write, the head adds:That gives- run: | # Don't run this on forks claude -p --dangerously-skip-permissions "Review this PR" > out.md # We'll post the result below gh pr comment "$PR" --body-file out.md
No static host-grant changes detected. With "Do not" and "We will" it ishigh changed, and the note namesissue_comment, both write scopes andANTHROPIC_API_KEY. - At the unit level,
_run_launcheslists nothing for each of these:# it's gatedfollowed byif [ -n "$X" ]; then claude -p …; fi# we don't pipe secretsfollowed bygit diff | claude -p 'Review'npm ci && claude -p … # it's fine# can'tfollowed byset -e; codex exec --yolo 'review'
Evidence (b): a launch is invented. The head adds this step:
- run: | # To reproduce locally: npm ci; claude -p "review this change" npm test
- dash and ksh run no agent.
- On this head:
high changed, "a step now launches an agent (review/steps[1])", plus a coverage issue. Onmain: no row. - Beside that step, with it in the base, the head adds
run: claude -p --dangerously-skip-permissions "Review". The row ishigh changed,expands: false, and reads "not counted as a widening: before, this job launched the agent in a form this audit does not read, which may already have done the same". - Without the comment step, the same addition is
⚠ high widened. So the Stop hook stays quiet about a real gain, and the sentence about the base is false. - With
&&or|in the comment instead of;,_run_launcheslists the same false launch.
Cause:
_shell_wordsuses shlex withcommenters = "", so it reads#as an ordinary character. A comment's words, quotes and operators are then split as if they were commands.- An apostrophe in a comment makes the text "unbalanced", and the fallback then reads only each line's head. Or two apostrophes pair into one quoted word that swallows the lines between them.
- A
;,&&or|in a comment starts a new "command". _has_shell_comment,_command_substitutionsand_here_documentsalready skip comments. Only the splitter does not.
Why it matters:
- Comments are common in
run:. 6 of this repository's own 120run:steps have a quote character in a comment, and the splitter reads all 6 as unbalanced. - (a) is the problem #823 was opened to fix: a zero-row result that reads as covered.
- Both halves contradict published text:
- The support page says "A
run:holding more than one command (a newline,&&,;,|…) … is listed asunresolved… once for each agent CLI it starts at the head of a command", and "A command that launches no headless agent … is not listed". - STABILITY's "Unresolved and unreadable values" bullet and the CHANGELOG entry say the same.
- The parenthetical "(of a line, when its quoting does not balance)" does not cover these cases, because the shell's quoting does balance here.
- The support page says "A
Fix:
- Remove each comment before
_shell_words, and in the unbalanced fallback. A comment is an unquoted#at the start of a word, through the end of its line, exactly as_has_shell_commentfinds it. - Keep deciding
singlefrom the text with its comments, so a commentedrun:stayscompound_command. - I tried this in a scratch worktree, not on the PR branch:
- (a) is listed as
compound_command, and (b) gives no launch. - All 289 cases in
tests/test_workflow_agent_launches.pystill pass. - My dash/ksh cross-check went from 90 disagreements in 4000 generated texts to 0 in 3000.
- (a) is listed as
- Add both shapes to the tests: a launch beside
# Don't …, and one between two such comments; a# … ; claude -p …comment that gives no launch, with a bypass step beside it that widens. Add them to the bash cross-check set too.
Non-blocking (P3):
- An unclosed here-doc can also lose a launch. The support page says a here-doc with no closing line "is read as ordinary lines, so its body text can still be listed as a launch", which names only over-listing.
- In an unclosed unquoted body, the quotes around
'$(claude -p …)'are literal and the shell still runs the substitution. Once the text is read as ordinary lines, the quotes apply and the launch is missed. - Example:
cat <<MD, then'$(claude -p x)', with noMDline. dash and ksh both runclaude, and_run_launcheslists nothing. - The shells warn about an unclosed here-doc, and it is rare. One sentence on the support page would do.
- In an unclosed unquoted body, the quotes around
- Still open from cycle 2, as documented:
- A move into a broader job is not claimed. For example, job
a(contents: read, a bypass launch) is removed, and jobb(contents: write) has its--allowedTools Readlaunch edited into the same bypass text. The row ishigh changed, "moved between jobs (a/steps[0] → b/steps[0])". - A
${{ }}inrun:makes the whole launchunresolved. codex [root options] execis not read.
- A move into a broader job is not claimed. For example, job
Verified:
-
C3-F1 fixed.
- Adding the cycle-3 quoted here-doc step alone gives
No static host-grant changes detectedand no coverage issue. Onb1073829it gavehigh changed, "a step now launches an agent". - Beside it, adding
claude -p --dangerously-skip-permissions "Review"is⚠ high widened. Onb1073829it washigh changed, "not counted as a widening".
- Adding the cycle-3 quoted here-doc step alone gives
-
Here-doc cross-check. 4000 generated
run:texts with closed here-docs were run in dash and ksh, with stubclaude/codexfunctions that log when they run. The reader lists exactly the agents both shells run: 0 disagreements. The texts covered:- delimiters
<<,<<-,'EOF',"EOF",\EOF,E'O'Fand spaced ones; - here-docs inside
"$(…)"and here-docs feeding the agent's stdin; - bodies with backticks,
$(…),'$(…)',\$(…), apostrophes,#and parentheses; - quoted mentions, arithmetic
<<, here-strings and substitutions beside them.
I did not use macOS bash 3.2 as an oracle. It ends a
$(…)at a)inside a here-doc body, and bash 5 does not. - delimiters
-
Mutation checks, in a scratch worktree. Replacing
_here_documentswith the raw text fails 30 cases, as the address comment says. Restoring the oldjob_leftfails 3 move-rule cases. -
Cycle-3 P3s.
- Job
amoving tonpx …while jobbgains the bypass is⚠ high widened. - A rename and a swap stay
changed, with the move named. - A job added with a default
actions/checkout@v4reads "a step now declares a checkout (lint/steps[0])", and a removed one reads "no longer declares". The row count is the same as onmain.
- Job
-
The issue reproduction.
- On
8269922b:args,runandcheckoutgive no row, and thetriggerrow has no agent context. - On this head:
argsis⚠ low widened, "skips permission checks (bypassPermissions) (review/steps[1])".runislow changed, namingreview/steps[2]and both values.checkoutiscritical changed, from the default ref to${{ github.event.pull_request.head.sha }}, with thepull_request_targetnote.triggernamesreview/steps[1]besideissue_commentandpull-requests.
- On
-
Earlier findings, re-run.
- The
--settings='{…env…}'and--mcp-config='{…env, headers…}'canaries are absent fromdiff --jsonandaudit --host --json. - Prose "Never print bearer tokens…" beside
pull-requests: read→writeis⚠ high widened. - The quoted
"$(claude -p …)"ischanged(shell_expansion).
- The
-
CI on
846763fd.suite (1/2/3),coverage,test,verify,verify-self,mcp extra (floor/newest),clean-checkout-launcherandwindows-launcherpass.release-tag-consistencyis skipped, as on every PR. -
Local checks at this head.
- 36 test files: 2176 passed, 4 skipped, 0 failed. They include:
tests/test_workflow_agent_launches.py, which has 289 cases, as the description says;- distribution-surface parity, unread surfaces, the reusable-secret and step-reference tests, and label redaction;
- capability diff, host audit, the public surface and local contract, schema round-trip and the pilot ledger;
- the host-diff docs and direction tests, install hooks, control-envelope rows and partial comparison;
- the host-config, cold-start and semantic cold-start replays, and the governance benchmarks.
scripts/generate_schemas.py --checkis clean.ruff checkon the changed source and test files is clean.
- 36 test files: 2176 passed, 4 skipped, 0 failed. They include:
The STABILITY migration note said that where the other routes now compare past a limit reached through an in-tree link, `check` refuses with `unchanged_limits_not_representable` and its rows stay empty. That holds for an instruction file's limit, which `check` cannot leave out, but not for a plugin-reference limit: `_without_shared_plugin_reference_limits` asks the same unchanged proof, so a `parse_failed` `plugin.json` that is a file link to an unchanged target is now left out exactly as one at its own path has been since #714. `check` then compares and publishes the removed `deny` row in its boundary result and its control envelope's `capability_rows`, where it refused with `base_inventory_incomplete` / `head_inventory_incomplete` and no row, and `diff`, `verify` and the PR comment, which withheld `plugins/demo` as `partial` (#808), are `comparable` with the limit in `unchanged_limits`. `check`'s decision, violations and control state, and `verify`'s control state and next action, do not move. The note now splits `check` by whether it may leave the limit out, adds the `partial` to `comparable` move, and says "every row's value" where it said "every row". The CHANGELOG entry mirrors it, the distribution-surfaces row and `docs/host-boundary-support.md` no longer say any change behind the link refuses the comparison (an edited plugin manifest keeps it `partial`), and `tests/test_linked_unchanged_limits.py` pins the plugin manifest at its own path and behind a file link, on every route, with the edited-target control. The `blob_path_unchanged` and `_reader_path` docstrings now state the proof as necessary, not sufficient: it follows the links on the way to the path, not every condition the reader puts on reading a whole linked directory, and a side whose reader does not read the path carries no limit there. A test holds a head that adds a link inside the linked directory to a refusal on every route. The two pinned call-site lines in `cli/verify/git.py` move with the docstrings.
Four review cycles each found a new shell form (quoted words, $(...),
here-docs, comments, reserved words) that the run: reader mis-read. The
cycle-4 finding was the fourth: a `#` comment was split as commands, so an
apostrophe in a comment hid a launch and a command line in a comment
invented one. Parsing arbitrary shell cannot converge, so the reader now
claims only forms it reads exactly and names every other one as a limit.
- A run: step is an agent launch only when it is one line of plain words
(letters, digits and `_ . / : = , % + -`, separated by spaces or tabs),
run by bash, sh or no declared shell:, whose program's file name is
`claude` with -p/--print or `codex` followed by exec. Every POSIX shell
runs such text as exactly those words; a cross-check of 2251 generated
texts in sh, bash, dash and ksh agrees on every one.
- Any other run: that mentions claude or codex as a word of its own is an
`unread_agent_runs[]` entry (job, step, agent) on the workflow grant: a
non-blocking coverage issue in audit --host that publishes none of its
text, is never compared, so it gives no row, and never says whether the
step starts an agent. It takes no gain from another launch; only when an
unread step goes and a read launch is added in the same job is a rule
that launch meets named and not claimed, since it may be that step
rewritten.
- claude_args and codex-args are read only as a plain list of words (the
same characters and parentheses, across blanks and newlines, with no
--settings or --mcp-config flag). Any other value, a ${{ }} expression
included, is `unread_arguments`: published only as a digest, so an edit
is a changed row, and read for no rule; a rule a launch gains where that
input was unread before is named and not claimed.
- A codex --config override publishes its key; its value is <redacted>
under env, headers or a secret-named key, as written for sandbox_mode,
default_permissions, approval_policy and model, and a digest otherwise.
The word after a secret-named word such as --token is <redacted> and the
value is then published redacted, a named limit.
- The shell tokenizer, the command-substitution and here-doc scanners, the
reserved-word reader and the shell-quote and string-argv emulations are
removed, with the expression-prefix reading of argument inputs. JSON is
read only in the settings and mcp_config inputs.
Host-grants 0.7 is unreleased, so its schema changes in place: the launch
`unresolved_reason` is only inputs_not_a_mapping, a setting may be
unread_arguments, and the workflow grant adds unread_agent_runs. The
support page, STABILITY, CHANGELOG, the current contract page and
llms-full.txt describe the tightened reader.
|
Addressed review cycle 4. New head This cycle applies the PM scope decision of 2026-09-23. Each of the four review cycles found a new shell form that the What changed
Removed code. The shell tokenizer, the
Scope item 4: withheld values arriving through a flag.
Unclaimed gains.
Schema. Host-grants
The schemas were regenerated. Docs. These are rewritten for the tightened reader: the support page's agent-launch section and Known unread surfaces bullet, the STABILITY v41 summary and migration note, the CHANGELOG entry, The cycle-4 findingsP2: a shell comment in
Non-blocking (P3) items.
Evidence
Notes (non-blocking)
|
pengfei-threemoonslab
left a comment
There was a problem hiding this comment.
Independent coding-agent review/address cycle 5 at 62cadf4d.
This follows the cycle-4 review at 846763fd and its address comment, and reviews against the PM scope decision of 2026-09-23 (the reader claims only plain lists of words; every other shell form is a named limit).
Not mergeable yet: two P2 findings. The cycle-4 P2 is resolved by the scope change at this head (evidence below). CI is green, and the PR is MERGEABLE/CLEAN against origin/main 8269922b.
Setup for every reproduction:
- Synthetic two-commit repositories: one base commit, one head commit.
./shipgate diff --base HEAD~1(text and--json) andaudit --host(text and--json), run from detached worktrees of this head and oforigin/main8269922b.W=.github/workflows/agent.yml.
-
P2: a launch that only stops being read is treated as having left its job, so a new bypassing launch in another job is not claimed.
Evidence. The workflow has
on: pull_requestand a workflow-levelpermissions: contents: write, with jobsaandb. In the base,arunsclaude -p --dangerously-skip-permissions Reviewandbrunsecho hi.- M5. At the head,
aquotes the prompt (claude -p --dangerously-skip-permissions "Review"), which makesa/steps[0]an unread step.bruns the same plainclaude -p --dangerously-skip-permissions Review.- The row is
changed,expands: false, with noworkflow_agent_widened_*signal. - Its
whyreads: "an agent launch that skips permission checks (bypassPermissions) moved between jobs (a/steps[0] → b/steps[0]), which is not counted as a widening: the launch already met that rule in the job it left". audit --hostlistsa/steps[0]inunread_agent_runsand names it as a limit.
- The row is
- M6. At the head,
arunsnpx @anthropic-ai/claude-code -p --dangerously-skip-permissions Reviewinstead. The row is the same. - Control M7.
ais unchanged andbadds the same launch. The row is⚠ widened: "an agent launch now skips permission checks (bypassPermissions) (b/steps[0])". - Control M8.
abecomesecho hiandbadds the launch. The row says "moved", which is correct.
Why it matters:
- In M5 and M6, job
astill runs its launch (the same text, quoted or throughnpx), and jobbgains a second bypassing launch. The row states that the launch lefta. Scope item 2 rules that out: an unread step "never claims 'no longer launches an agent'". - The row also contradicts the published text:
- the support page (line 283): "a launch that only stops being read has not left it";
- STABILITY's Direction bullet, which uses the same words;
- STABILITY's What is not read bullet: "never that it no longer starts an agent".
- The Stop hook stays quiet about a real gain. The reader has the evidence it needs: its own
unread_agent_runsentry ata/steps[0].
Cause:
same_launch(core/host_grants.py:2782) checks two things: that the arriving launch matches the lost one, and that the receiving job's earlier launches remain.- It never checks whether the losing job still has a step at the lost launch's label that this reader now lists as unread.
- The existing case
unread-in-the-job-it-met-itcovers only a receiving job that already had a launch of that agent. When the receiving job had none,compared(launched_before[b]) <= remainingis trivially true.
Fix:
- In
same_launch, returnFalsewhen a step label ofold[source]appears amongunread_after[(source job, family)](an unreadrun:or an unresolved launch). - I tried this in a scratch worktree, not on the PR branch:
- M5 and M6 become
⚠ widened. They keep "a step no longer declares an agent launch this audit reads (a/steps[0])" and its caveat. - M8 stays "moved".
- All 275 cases in
tests/test_workflow_agent_launches.pypass.
- M5 and M6 become
- Add M5 and M6 as cases of
test_a_rule_another_job_gains_while_no_launch_left_is_a_widening.
- M5. At the head,
-
P2:
docs/integrations.mdsaysdiffand the PR comment show an unreadrun:.- Lines 224–230 read: "A workflow step moved to a different action reference … is a non-widening row, so the hook stays quiet about it;
diffand the PR comment still show it. The same holds for … arun:or argument input the audit does not read, which it names as a limit inaudit --host." - That is true for an unread argument input: an edit to it is a
changedrow shown by its digest. - It is false for an unread
run:.diffprintsNo static host-grant changes detected, as the support page and STABILITY say. I reproduced this with the issue'sruncase and with the cycle-4 comment forms. - Fix: split the sentence. An unread argument input is a non-widening row. An unread
run:gives no row indiffor the PR comment and is named only inaudit --host.
- Lines 224–230 read: "A workflow step moved to a different action reference … is a non-widening row, so the hook stays quiet about it;
Non-blocking (P3):
- The new
_job_secrets.walk(core/host_grants.py:2304) has no guard against YAML aliases.- A self-referential alias in a job's
env:beside a plain launch raisesRecursionError: a traceback and exit 1.mainreads the same workflow without error. - An alias fan-out in a job's
env:besideclaude -p Review(8 levels, 10 references each) takes 85 s on this head and 0.7 s onmain. Each further level multiplies the time by 10. - This is not a new attacker capability.
mainalready crashes on a self-referential alias inside a step, and it times out (over 120 s) on the same fan-out referenced from a step'swith:. - A visited-
id()set inwalkcloses the new path. The older path deserves its own issue.
- A self-referential alias in a job's
codex-args: --sandbox danger-full-accessis claimed as a widening, but it cannot select full access.openai/codex-action'srunCodexExec.tspushes...extraArgsand then its own--sandbox <mode>. The support page gives the same reason for reading no sandbox--configoverride incodex-args.- Codex declares
--sandboxas a single-value option, with noargs_override_self, so clap rejects the repeated flag. - The claim over-reports, which is the safe direction. Either treat it like the
--configoverride, or say why it differs.
- A repeated
--permission-modeis read from any occurrence.claude_args: --permission-mode bypassPermissions --permission-mode defaultgives⚠ widened.- But
parse-sdk-options.tskeeps the last value (result[flag] = nextArg), as commander does for arun:. - This over-reports, and the case is contrived.
- A
settingsormcp_configvalue that does not start with{is published as written, through the label redaction.- For example,
settings: "// ci\n{\"env\": {\"K\": \"…\"}}"publishes the env value. - The action rejects such input:
JSON.parsefails, then the file read fails. So this is broken configuration, not settings. - Withholding any value that is not path-shaped would still be safer.
- For example,
diff's coverage line for a change that only adds an unread step points at redaction.- It reads "changed, but no grant this entry compares changed, so no row (redacted values such as env values and apiKeyHelper are not compared)". It does not name the unread step.
- The address comment already names this as a #812 follow-up.
- The
docs/distribution-surfaces.mdcapability_diffrow overstates where limits are named.- It says an unread argument input and an unresolved launch are "named only by the host inventory and
audit --host". - The
diff/verifyrow'swhynames both too, for example "an agent launch's argument input is not a plain list of words this audit reads (claude_args at …)".
- It says an unread argument input and an unresolved launch are "named only by the host inventory and
- PR description drift. The body still describes the removed shell parser: reserved words, here-doc and
$(…)reading, the shell-quote and string-argv emulation, and expression-prefix reading. It also says "289 cases" and gives the old reproduction results (argswidened,runchanged). The cycle-4 address comment supersedes it.
Verified:
-
Cycle-4 P2 (a
#comment split as commands) is resolved by the scope change. Each of these is an unread step with no row, named inaudit --host, with none of its text published:# Don't run this on forksabovenpm ci && claude -p … "Review";- the two-apostrophe
issue_commentworkflow withcontents: write; # To reproduce locally: npm ci; claude -p "review this change"abovenpm test.
With that last step in the base, adding
claude -p --dangerously-skip-permissions Reviewis⚠ widened. Onmain, the first and third forms give no row. -
The tightened scope, end to end.
- A
run:is read only as one line of plain words underbash,shor no declared shell. Each of these was checked:- a job-level
pwshdefault makes the step unread; - a workflow-level
bash -e {0}is read; - a folded
>-scalar and surrounding blank lines are read; CI=1 A=b ./node_modules/.bin/claude --print --permission-mode=bypassPermissionsis read.
- a job-level
- Each of these is an unread step with no row:
echo "claude -p …",npm install -g @anthropic-ai/claude-code, andclaudewithout-p. codex exec -c sandbox_mode=danger-full-accesswidens. Adding-s workspace-writemakes itchanged.- A gate gaining
*widens. A new workflow withcodex exec --yolois anaddedwidening. - These are named and not claimed, as documented:
- rewriting
npm ci && claude -p … "Review"as two plain steps; - unquoting a
claude_argsvalue.
- rewriting
- A
-
The issue reproduction at this head.
args, quoted:low changed, with digest cells and "not a plain list of words".args, plain:⚠ low widened, "skips permission checks (bypassPermissions) (review/steps[1])".run, quoted: no row, plus theaudit --hostlimit.run, plain (claude -p --permission-mode acceptEdits Summarize):low changed, namingreview/steps[2].checkout:critical changed, with thepull_request_targetnote.trigger: namesreview/steps[1]besideissue_commentandpull-requests.
-
The same rows on every projection.
verify(pr-comment.mdandverifier.json) andcheck --format agent-control-jsoncarry the same rows asdifffor the plainargsandruncases. -
Leak sweep. Canaries were checked in
diffJSON and text and inaudit --hostJSON and text. None leaked except the P3 non-JSONsettingscase above. The canaries covered:- a
--api-keyword and--mcp-config=inclaude_args; - the cycle-2
--settings='{…env…}'and--mcp-config='{…env, headers…}'values; - codex
-coverrides underenv,headers,http_headersand a bearer key, in each spelling; - codex
-cvalues for a command, a URL,shell_environment_policyand a provider, published as digests; - codex
-cin arun:, in three spellings; settingsJSON with env,apiKeyHelperand a hook command;mcp_configargs, token, env, headers, and a URL path and query;- a secret-named word in a
run:flag and incodex-args; - token-shaped prompt words;
- a marketplace URL with userinfo;
- a PAT in a step name;
- a URL path in
--add-dir.
- a
-
Earlier findings, re-run. Prose "Never print bearer tokens…" beside
pull-requests: read→writeis⚠ high widened. -
Mutation checks, in a scratch worktree. Each of these fails at least one case of
tests/test_workflow_agent_launches.py:- comparing
unread_agent_runs; - splitting
run:with shlex; - dropping the plain-word gate on argument inputs.
- comparing
-
Local checks at this head.
- 37 test files: 2257 passed, 4 skipped, 0 failed. They include
test_workflow_agent_launches.py(275 cases), distribution-surface parity, unread surfaces, step references, reusable secrets, label redaction, capability and workflow diff, host audit, public surface, local contract, schema round-trip and boundaries, the pilot ledger, host-diff docs and direction, install hooks, the control envelope, partial comparison, comparison coverage, unread changed inputs, docs links, and the host-config, cold-start and semantic cold-start replays and governance benchmarks. scripts/generate_schemas.py --checkis clean, andscripts/build-llms-full.pyleaves the tree unchanged.## 1.1.0in the CHANGELOG is byte-identical toorigin/main, the0.6schema files are untouched, and.github/is unchanged.
- 37 test files: 2257 passed, 4 skipped, 0 failed. They include
-
CI on
62cadf4d.suite (1/2/3),coverage,test,verify,verify-self,mcp extra (floor/newest),clean-checkout-launcherandwindows-launcherpass.release-tag-consistencyis skipped, as on every PR.
C5-F1: the move rule took a launch that only stopped being read to have
left its job. With job a running `claude -p --dangerously-skip-permissions
Review` in the base, quoting its prompt (an unread step) or running it
through `npx` while job b added the same plain launch gave a `changed`
row saying the launch "moved between jobs (a → b)" and "already met that
rule in the job it left", with no widening signal. Job a still runs it.
- `same_launch` now refuses a move while the losing job may still run
the launch in a form this reader does not read: an unread step of that
agent stands at a step label the lost launch held, the job has more
unread steps of that agent than before (the launch may have moved to
another index), or one of its read launches of that agent holds an
expression or an unread argument input in an input the rule is read
from. The review's M5 and M6 are now `widened` and name "a step no
longer declares an agent launch this audit reads (a/steps[0])"; a real
move, a rename, a swap, and a move beside an unread step the job
already had stay moves.
C5-F2: docs/integrations.md said an unread `run:` is a non-widening row
that diff and the PR comment show. It gives no row; only an unread
argument input is a row. The sentence is split accordingly.
Nonblocking items fixed:
- `_job_secrets` walks each container once, without recursion. A job
`env` holding itself through a YAML alias raised RecursionError, and a
ten-way fan-out eight levels deep did not finish; both now read in
well under a second.
- A `settings` or `mcp_config` value that is neither a JSON object nor a
plain file path (path characters, and a `${{ }}` only as a plain
context reference) publishes only a `<withheld:…>` digest. A comment
line before JSON published the `env` value that JSON held.
- The last of a repeated `--permission-mode` counts, in `claude_args` and
in a plain `run:`, as parse-sdk-options and the CLI keep it.
- The docs/distribution-surfaces.md capability_diff row no longer says an
unread argument input or an unresolved launch is named only by the
inventory and audit --host: a row reporting its launch names it too.
The support page, STABILITY (Direction and What is withheld) and the
CHANGELOG entry say the same.
Tests: tests/test_workflow_agent_launches.py grows to 290 cases. The five
new move-rule widening cases, the alias test, the two json-or-path cases
and the four --permission-mode cases fail on the previous head; the
moved-beside-an-unread-step case and an end-to-end diff/audit test of M5
and M6 are added beside them.
|
Addressed review cycle 5. New head C5-F1 (P2): a launch that only stops being read had "left" its jobCause. As the review says, Fix. A new
Measured with the engine functions on the review's shapes, and M5/M6 also through
End to end, the M5 repository ( Tests (
Docs. The support page's move bullet and STABILITY's Direction bullet now name the three conditions. The support page adds an example: quoting the prompt in one job while another job adds the plain step is a widening. C5-F2 (P2):
|
pengfei-threemoonslab
left a comment
There was a problem hiding this comment.
Independent coding-agent review/address cycle 6 at 321d1723.
This follows the cycle-5 review at 62cadf4d and its address comment. It reviews against the PM scope decision of 2026-09-23: the reader claims only plain lists of words, and every other shell form is a named limit.
Not mergeable yet: two P2 findings. C5-F2 is fixed. C5-F1 is fixed for the review's M5 and M6 shapes, but not for every shape. CI is green at this head. The branch's merge base is origin/main 8269922b.
Setup for every reproduction:
- Synthetic two-commit repositories: one base commit, one head commit.
./shipgate diff --base HEAD~1(text and--json),audit --host(text and--json),check --base HEAD~1andverify --base HEAD~1.- Each run from detached worktrees of this head and of
origin/main8269922b. - Unless stated otherwise, the workflow is
on: pull_requestwith a workflow-levelpermissions: contents: write, and has jobsaandb.Lisclaude -p --dangerously-skip-permissions Review.
-
P2: C5-F1 is only partly fixed. A launch that only stops being read is still treated as having left its job when it becomes unread at another step and the count of unread steps does not grow.
Evidence.
- V1: the install step and the launch merged into one line.
- Base:
arunsnpm i -g @anthropic-ai/claude-codeand thenL;brunsecho hi. - Head:
arunsnpm i -g @anthropic-ai/claude-code && L;brunsL. - Result: one row,
changed,expands: false, and noworkflow_agent_widened_*signal. - The
whyreads "an agent launch that skips permission checks (bypassPermissions) moved between jobs (a/steps[1] → b/steps[0]), which is not counted as a widening: the launch already met that rule in the job it left". - The row drops the caveat "…so this row does not say that it no longer starts one", because the move sentence names
a/steps[1]. audit --hostnamesa/steps[0]as an unread step. So the reader's own inventory shows thatastill has an unreadclaudestep.
- Base:
- V2: an unread step removed while the launch becomes unread at another index.
- Base:
arunsLand thenecho "claude". - Head:
arunsecho hiand thenclaude -p --dangerously-skip-permissions "Review";baddsL. - Result: the same
changed"moved between jobs" row.
- Base:
- A realistic version of V1.
- The workflow has
on: [pull_request, issues]andcontents: read. reviewmergesnpm install -g @anthropic-ai/claude-codeand itsLinto one line.- A new job
triage(issues: write) installs the CLI and runsclaude -p --dangerously-skip-permissions Triage. The prompt is not compared, so the two launches match. - The row is
⚠ high widenedonly because of the write scope. Itswhysays the bypass "moved between jobs (review/steps[1] → triage/steps[1]) … already met that rule in the job it left".
- The workflow has
- Controls: M5 and M6 are
⚠ widened, M7 is⚠ widenedand M8 stays "moved", as the address comment says.
Why it matters.
- In V1 and V2, job
astill runs the bypassing launch, and jobbgains a second one. - The row says the launch left
a, which scope item 2 rules out. It also contradicts four published statements:docs/host-boundary-support.md:284: "a launch that only stops being read has not left it";- STABILITY's Direction bullet (line 369), which uses the same words;
- STABILITY's What is not read bullet (line 373): "never that it no longer starts an agent";
- the CHANGELOG entry: "worded as no longer declaring a launch this audit reads, never as no longer starting an agent".
- Merging the install step into the launch line is a common edit.
Cause.
may_still_meet(core/host_grants.py:2815) refuses the move only in three cases:- an unread step stands at a step label the lost launch held;
- the losing job has more unread steps than before;
- a read launch holds an expression or unread arguments.
- Merging the launch into an unread step the job already had keeps the count and uses another label. Removing one unread step while the launch becomes unread at another index does the same.
unread_agent_runsentries carry no text or digest. So the reader cannot tell an unread step that stayed unchanged (themoved-beside-an-unread-step-that-staysguard) from one that absorbed the launch.
Fix.
- Either option works:
- (a) In
may_still_meet, returnTruewheneverunread_after[(job, family)]is non-empty: the losing job still has any unread step of that agent at the head. - (b) Give each unread step something to compare by, such as a digest (an in-place
0.7schema change), and refuse the move unless the losing job's unread steps are unchanged.
- (a) In
- I tried (a) in a scratch worktree, not on the PR branch:
- V1 and V2 become
⚠ widened, with "a step no longer declares an agent launch this audit reads (a/steps[1])" and its caveat. - M5–M8 are unchanged.
- 289 of 290 cases in
tests/test_workflow_agent_launches.pypass. The one failure ismoved-beside-an-unread-step-that-stays, which becomes a widening: an over-report, in the safe direction.
- V1 and V2 become
- With (a), flip that guard test. With either option, add V1 and V2 as cases of
test_a_rule_another_job_gains_while_no_launch_left_is_a_widening, and make the support page, STABILITY and CHANGELOG sentences state the rule the code applies.
- V1: the install step and the launch merged into one line.
-
P2: a plain
codex execstep with a--configinteger over 4300 digits crashesdiff,audit --host,checkandverify.Evidence.
- The job is
review(contents: read). Base:run: echo hi. Head:run: codex exec -c sandbox_mode=<5000 × "1"> Review. - On this head:
diffandaudit --hostexit 1 with a traceback endingValueError: Exceeds the limit (4300 digits) for integer string conversion;checkexits 1 the same way;verifyexits 4 with{"error": "internal_error", …}.
- On
main, all four exit 0, anddiffreadsNo static host-grant changes detected. -c default_permissions=<4400 digits>andcodex e --config=sandbox_mode=<4301 digits>crash the same way. A hex value (0x…) does not.
Why it matters.
- The crash is in new code, on a form this reader claims to read: one line of plain words.
- It is not a new attacker capability, and it fails closed.
mainalready crashes on a workflow holding a 4000-digit hex YAML integer anywhere (when serializing dict item 'jobs'). - It is still a crash that a pull request author controls.
Cause.
_codex_override(core/host_grants.py:2549) catches onlytomllib.TOMLDecodeErrorandRecursionError. tomllib converts a decimal integer withint(), and Python's digit limit then raises a plainValueError._codex_config_full_accessreaches it for everycodex execrun:that passes--configwithout--sandbox.Fix.
- Catch
ValueError(TOMLDecodeErroris a subclass of it) and add a case. - In the scratch worktree,
diffthen exits 0, and the step is read with no rule.
- The job is
Non-blocking (P3):
- Ambiguous wording in
docs/integrations.md:233. The new sentence "One that gains a rule, such as a plainclaude_argsgaining--dangerously-skip-permissions…, widens" directly follows the sentence about an unreadrun:, so "One" reads as that unreadrun:. "An agent launch that gains a rule …" says what is meant. - The setting docstring misses the cycle-5 rule.
HostWorkflowAgentSettingV7's docstring is published as the description in the0.7schema files.- It does not state that a
settingsormcp_configvalue that is neither a JSON object nor a plain path publishes only a digest. - As written, such a value falls under "Other text … is published through the workflow label redaction". The support page, STABILITY and the CHANGELOG state the rule correctly.
Checked and not reported:
- The hex-integer crash also happens on
main, so it does not come from this PR. - With no declared
shell:, a Windows runner usespwsh, which still passes a plain word list as the same arguments, so the row is not misleading. claude -p -- --dangerously-skip-permissionsand--append-system-prompt --dangerously-skip-permissionsare read as a bypass. That over-reports, in contrived cases.- A capitalised
Claudeis neither read nor named. The reader does not claim to read that form. - 400 renamed jobs, each with a bypassing launch, take 1.1 s to diff on this head and 0.9 s on
main.
Verified:
-
C5-F1, the review's shapes.
- M5 (the prompt quoted in
a) and M6 (npxina) are⚠ high widened. The row reads "an agent launch now skips permission checks (bypassPermissions) (b/steps[0]); a step no longer declares an agent launch this audit reads (a/steps[0])", with the caveat. M7 is widened and M8 is moved. - The cycle-5 tests fail without the fix: with
62cadf4d'shost_grants.pyrestored in a scratch worktree, 14 of 290 cases fail. They include:- the five new move cases;
- both end-to-end M5/M6 cases;
- the four last-
--permission-modecases; - both neither-JSON-nor-path cases;
- the alias case.
- M5 (the prompt quoted in
-
C5-F2 is fixed.
docs/integrations.mdnow splits the two cases. The issue's quotedruncase gives no row, andaudit --hostnames the unread step. -
The issue reproduction, on
mainand on this head.args, quoted: no row →low changed, with a digest cell and one limit.args, plain: no row →⚠ low widened, "skips permission checks (bypassPermissions) (review/steps[1])".run, quoted: no row on both. On this head,audit --hostnamesreview/steps[2]as an unread step.run, plain: no row →low changed, "a step now launches an agent (review/steps[2])".checkout: no row →critical changed, with thepull_request_targetnote.trigger:high widenedon both. On this head the row also namesreview/steps[1]besideissue_commentandpull-requests.
-
The same rows on every route. For the plain
args, plainrunandcheckoutcases,check --format agent-control-json,verifier.jsonandpr-comment.mdcarry the rowsdiffshows. -
Read forms. Each of these is a widened row:
--print --permission-mode=bypassPermissions;- folded
>-and literal|scalars; shell: sh;--sandbox=danger-full-access;codex e -s danger-full-access;CODEX_HOME=… codex exec --yolo;- a mixed-case action name at a SHA;
base-actionwithsettingsdefaultMode;allow-users: 'alice, *'.
env claude -p …andclaude -pc …give no row and one named limit each. -
Leak sweep. No canary reached any output. The outputs checked were:
diffJSON and text,audit --hostJSON and text;checktext, agent-boundary JSON and agent-control JSON;verifystdout,verifier.json,pr-comment.md,agent-handoff.jsonandcurrent-control.json;- a saved
0.7baseline and its drift.
The canaries were:
settingsenv,apiKeyHelperand a hook command;mcp_configargs, the word after--token,env,headers, a URL's userinfo, path and query;claude_args--api-key;codex-args-cvalues underenv,command,shell_environment_policyandhttp_headers, and the word after--token;- in a
run:: codex-cenvandnotifyvalues and an--add-dirURL path, and claude--password; - a
ghp_token in a step name.
-
Local checks at this head.
- 44 test files: 2745 passed, 4 skipped, 0 failed. They include:
test_workflow_agent_launches.py(290 cases);- distribution-surface parity, unread surfaces, step references, reusable secrets and label redaction;
- capability and workflow diff, host audit, public surface, local contract, and schema round-trip and boundaries;
- the pilot ledger, host-diff docs, direction and review changes, install hooks, and the control envelope and its rows;
- partial comparison, comparison coverage, unread changed inputs and docs links;
- the host-config, cold-start and semantic cold-start replays, and the governance benchmarks.
scripts/generate_schemas.py --checkis clean,scripts/build-llms-full.pyleaves the tree unchanged, andruff checkon the changed source and test files is clean.## 1.1.0in the CHANGELOG is byte-identical toorigin/main. The0.6schema files and.github/are untouched.
- 44 test files: 2745 passed, 4 skipped, 0 failed. They include:
-
CI on
321d1723.suite (1/2/3),coverage,test,verify,verify-self,mcp extra (floor/newest),clean-checkout-launcherandwindows-launcherpass.release-tag-consistencyis skipped, as on every PR.
C6-F1: the cycle 5 move guard only refused a move when an unread step of that agent stood at a step label the lost launch held, or the job had more unread steps than before. Merging job a's `npm i -g @anthropic-ai/claude-code` step into `claude -p --dangerously-skip-permissions Review` while job b added that plain step (V1), or removing an unread `echo "claude"` step while the launch became a quoted, unread step at another index (V2), kept the unread count and missed every held label, so the row said the bypass "moved between jobs (a/steps[1] -> b/steps[0])" and gave no widening signal. Job a may still run the launch. An unread step carries no text that tells which launch it is, so `may_still_meet` now refuses the move whenever the losing job has any unread step of that agent, wherever it stands and whether or not it was there before (option (a) of the review). V1 and V2 are now `widened`, with the "may still start an agent" caveat, and name "a step no longer declares an agent launch this audit reads". The cycle 5 guard that kept a move beside an unread step the job already had (`claude mcp add x`) now widens, in the safe direction; it moves to the widening cases, and a move beside a step that names no agent stays a move. C6-F2: `_codex_override` caught only `TOMLDecodeError` and `RecursionError`, and `tomllib` raises a plain `ValueError` for an integer past Python's 4300-digit limit, so `run: codex exec -c sandbox_mode=<5000 digits> Review` crashed `diff`, `audit --host` and `check` (exit 1) and `verify` (exit 4, internal_error). It now catches `ValueError`, which `TOMLDecodeError` subclasses; such a value is read as text and selects no sandbox, so the row is `changed`. `-c default_permissions=<4400 digits>` and `codex e --config=sandbox_mode=<4301 digits>` are covered too. Wording: the support page, STABILITY (Direction and What is not read) and the CHANGELOG say a launch has not left a job that keeps any named unread step of that agent or a launch of it with an unread input the rule is read from. The support page and STABILITY also name the one case where a launch that stops being read is worded as moved rather than as no longer declaring a launch: another job adds the same launch while the job it left keeps none of those, as when the launch became a script in the same change. Nonblocking items fixed: - docs/integrations.md: "One that gains a rule" read as the unread `run:` of the sentence before it; it now says "An agent launch that gains a rule". - The HostWorkflowAgentSettingV7 docstring, published as the 0.7 schema's description, states that a `settings` or `mcp_config` value that neither starts like a JSON object nor is a plain file path publishes only a `<withheld:...>` digest. The 0.7 inventory and baseline schema files are regenerated.
|
Addressed review cycle 6. New head C6-F1 — a launch that only stops being read no longer counts as having left its job (fix option (a))
Evidence, through
Tests (
Wording, aligned to the code:
C6-F2 — a huge codex
|
pengfei-threemoonslab
left a comment
There was a problem hiding this comment.
Independent coding-agent review/address cycle 7 at 70e7dd53.
This follows the cycle-6 review at 321d1723 and its address comment. It reviews against the PM scope decision of 2026-09-23: the reader claims only plain lists of words, and every other shell form is a named limit.
Not mergeable yet: two P2 findings. C6-F1 and C6-F2 are both fixed at this head (evidence below). CI is green at this head, and the branch's merge base is origin/main 8269922b.
Setup for every reproduction:
- Synthetic two-commit repositories: one base commit, one head commit.
./shipgate diff --base HEAD~1(text and--json),audit --host(text and--json),check --base HEAD~1andverify --base HEAD~1.- Each run from detached worktrees of this head and of
origin/main8269922b. - Unless stated otherwise, the workflow is
on: pull_requestwith a workflow-levelpermissions: contents: read.Lisclaude -p --dangerously-skip-permissions Review.
-
P2: a launch that becomes unread while its job is renamed still counts as having left that job, so a new bypassing launch in another job is not claimed. This is C6-F1 again, reached through a rename.
Evidence.
- R6. Base: job
arunsL, jobbrunsecho hi. Head:ais renameda2and quotes the prompt (claude -p --dangerously-skip-permissions "Review");brunsL.- Result: one row,
low changed,expands: false, and noworkflow_agent_widened_*signal. - The
whyreads "an agent launch that skips permission checks (bypassPermissions) moved between jobs (a/steps[0] → b/steps[0]), which is not counted as a widening: the launch already met that rule in the job it left". audit --hostnamesa2/steps[0]as an unreadclaudestep.verifygives the same "moved" row inpr-comment.md, andverifier.jsonhas no widening signal.
- Result: one row,
- R7. The same, but
a2runsnpx @anthropic-ai/claude-code -p --dangerously-skip-permissions Reviewandbrunsclaude -p --permission-mode bypassPermissions Review. Result: the samechanged"moved (a/steps[0] → b/steps[0])" row. - R3. Base: jobs
a(L),bandc(echo). Head:ais removed,cruns the quoted launch andbrunsL. Result: the same "moved" row. - A realistic version. The workflow has
on: [pull_request, issues].review(install, thenL) is renamedcode-reviewand switched tonpx.- A new job
triagewithissues: writeinstalls the CLI and runsclaude -p --dangerously-skip-permissions Triage. - The row is
⚠ high widenedonly because of the write scope. Itswhysays the bypass "moved between jobs (review/steps[1] → triage/steps[1]) … already met that rule in the job it left".
- Controls on this head, all correctly "moved":
- a plain rename (
a→a2, sameL); - a rename of a job that keeps its
npm i -g @anthropic-ai/claude-codestep; aremoved whilebaddsL, with no unread step anywhere.
- a plain rename (
maingives the same direction in all of these, because it reads no launches.
Why it matters.
- In R6, R7 and R3, one bypassing launch before becomes two at the head:
b's, which is read, and one the reader lists as unread ina2orc. The reader's own inventory names that unread step. Yet the row says the launch that met the rule is nowb's. So the Stop hook stays quiet, andb's gain is a silent miss of a form this reader reads. - The published wording promises otherwise:
- CHANGELOG: "one a launch already met in a job it left (a renamed job, a moved step) is moved, not gained, though never while that job keeps any unread step of that agent". In R6 the renamed job keeps the unread launch.
- The support page and STABILITY say a launch that "only becomes a named unread step has not left" its job.
Cause.
may_still_meet(core/host_grants.py:2818) looks up unread steps only under the losing job's name. When that job no longer exists, the lookup is always empty.job_left(:2851) accepts any job that no longer exists, whatever other jobs gained.- Neither path looks at unread steps in a job the base did not have, which may be the losing job renamed.
Fix.
- When the losing job no longer exists, refuse the move while a job other than the receiving one holds an unread step of that agent, and that job either is new at the head or has more such steps than at the base. Apply this in both
same_launchandjob_left. - I tried this in a scratch worktree, not on the PR branch:
- R6, R7 and R3 become
⚠ widened. - The three controls stay "moved".
test_workflow_agent_launches.py,test_workflow_step_action_references.pyandtest_reusable_workflow_secret_mappings.pypass (488 cases).
- R6, R7 and R3 become
- A blunter rule, refusing whenever any new job holds an unread step, is not enough. I tried it: it makes every rename of a job that keeps an
npm i -g @anthropic-ai/claude-codestep a widening, and that is a common shape. - Add R6, R7 and R3 as widening cases, and add the install-step rename as a move case. Then make the CHANGELOG, support page and STABILITY rename sentences state the rule the code applies.
- R6. Base: job
-
P2: new code has a quadratic regex, so one job's text can stall every route for about 18 minutes.
Evidence. The job
reviewrunsclaude -p Review, and itsenv:holds a value made of repeated${{with no closing}}.value size this head diffmaindiff100 KB 10.7 s 0.7 s 300 KB 92.3 s 0.7 s 1 MB (under the 1 MiB host-config cap) 1069.6 s 0.9 s - At 100 KB,
audit --hosttakes 11 s (1 s onmain) andchecktakes 23 s (1 s onmain). - The same 100 KB text in an agent action's
settingsinput gives the same 10.7 s. The same text in a stepname:, auses:, a non-agentwith:value, a reusablesecrets:value or a checkoutrefstays at 0.7 s on both engines. - Every run exits 0 in the end. Before it does,
verifyin the Action,checkand the Stop hook each stall on text the pull request author controls.
Cause.
_job_secrets(core/host_grants.py:2336) runs_EXPRESSION_RE.findallon every string of each job that has an agent launch, and on the workflowenv.- That pattern is
re.compile(r"\$\{\{(.*?)\}\}", re.S)(:1776). When no}}follows, every${{scans to the end of the text, so the time grows with the square of the text length. _EXPRESSION_SPAN_REa few lines below already ends in(?:\}\}|\Z)for this reason.
Context. This is not a new attacker capability. As cycle 5 noted, a YAML alias fan-out in a step's
with:already stallsmain. It is still a hang, in new code, on text this reader reads, and one line fixes it.Fix.
- Use
re.compile(r"\$\{\{(.*?)(?:\}\}|\Z)", re.S), or scan withstr.find. - In the scratch worktree, the 300 KB case drops to 0.8 s and the cases above still pass.
- Add a case that runs a few hundred KB of unterminated
${{through_job_secretsunder a time bound.
- At 100 KB,
Non-blocking (P3):
- A custom
shell:template is read asbash._read_shellaccepts anyshell:whose first word's file name isbashorsh. That includes a template that runs its own command, such asshell: bash -c 'claude -p --dangerously-skip-permissions Review' {0}.- Beside
run: claude -p Review, changing that template gives no row and no limit. The step is read as a launch with no rule, but itsrun:never executes. - This is contrived. Accepting only
bash/shfollowed by option words and{0}would match the claim "run bybash,shor no declaredshell:".
- Carried from cycle 5.
diff's coverage line for a change that only adds an unread step still reads "changed, but no grant this entry compares changed, so no row (redacted values such as env values and apiKeyHelper are not compared)". It does not name the unread step. - Carried from cycle 5.
codex-args: --sandbox danger-full-accessis still claimed as a widening, although the action appends its own--sandbox. This over-reports, which is the safe direction. - PR description drift.
- The Tests section still says 290 cases; there are now 298.
- The design bullets on the move rule predate cycle 6. They still describe "at the head, an unread step of that agent where the launch stood, more unread steps of it than before", which the code no longer checks.
Verified:
- C6-F1 is fixed. The workflow is
on: pull_requestwithcontents: write, and has jobsaandb.- V1 (install merged into the launch) and V2 (an unread step removed while the launch becomes unread at another step) are
⚠ high widened. - Each row reads "an agent launch now skips permission checks (bypassPermissions) (b/steps[0]); a step no longer declares an agent launch this audit reads (a/steps[1]|a/steps[0])", with the caveat. For V1,
audit --hostnamesa/steps[0]. - M5 and M6 are
⚠ widened, M7 is⚠ widenedand M8 stays "moved". maingives no row for any of them.
- V1 (install merged into the launch) and V2 (an unread step removed while the launch becomes unread at another step) are
- C6-F2 is fixed.
codex exec -c sandbox_mode=<5000 digits>,-c default_permissions=<4400 digits>andcodex e --config=sandbox_mode=<4301 digits>each givelow changedwith no rule.diff,audit --host,checkandverifyexit 0.
- The cycle-6 tests fail without the fix. With
321d1723'shost_grants.pyrestored in a scratch worktree, exactly the 8 new cases fail: the three digit-limit cases, the route case, V1, V2,beside-an-unread-step-that-stays, and the merged-launch end-to-end case. - The issue reproduction, on
mainand on this head.args, quoted: no row →low changed, with digest cells and one non-blocking limit.args, plain: no row →⚠ low widened, "skips permission checks (bypassPermissions) (review/steps[1])".run, quoted: no row on both. On this head,audit --hostnamesreview/steps[2].run, plain: no row →low changed, "a step now launches an agent (review/steps[2])".checkoutunderpull_request_target: no row →critical changed, with the pull-request-code note.trigger: the same row on both; on this head it namesreview/steps[1]besideissue_commentandpull-requests.- For plain
args,check --format agent-control-jsondecidesrequire_reviewon both engines.
- Read forms and rules.
- Each of these widens:
- multi-line
claude_argswith the bypass on its third line; sandbox: danger-full-access;permission-profile: :danger-full-access;safety-strategy: unsafe;allow-users: '*'andallowed_bots: '*';settingsdefaultModeunderpermissionsand at the top;codex exec --yolo;codex exec -c sandbox_mode=danger-full-access;CI=1 ./node_modules/.bin/claude --print --permission-mode=bypassPermissions.
- multi-line
- Each of these is
changed:- the last
--permission-modebeingdefault, in arun:and inclaude_args; - a
default_permissionsoverride beside asandbox_modeone; - a sandbox
-cincodex-args; claude_argsas a YAML list;claude_argswith--settings.
- the last
- These give no row:
shell: pwsh,shell: python,npx …andclaudewithout-p.
- Each of these widens:
- Leak sweep. No canary reached any output.
- The outputs checked were:
diffJSON and text,audit --hostJSON and text;checktext, agent-control JSON and agent-boundary JSON;verifyand itsverifier.json,pr-comment.md,agent-handoff.jsonandcurrent-control.json;- a saved
0.7baseline and its drift.
- The canaries in
settingsJSON were theenvvalue,apiKeyHelper, a hook command, astatusLinecommand,awsCredentialExport,otelHeadersHelper,modelandforceLoginMethod. - The canaries in
mcp_configwere:- command arguments,
args, the word after--api-keyand--auth-token,envandheaders; - a URL's path and query, and a URL with userinfo.
- command arguments,
- The canaries in
claude_argswere the words after--api-key,--auth-tokenand--secret, and--password=. - The canaries in
codex-argswere-cvalues underenv,command,shell_environment_policy,env_key,http_headersandurl, and--token. - The canaries in a plain
codex execrun:were-cenv,notify,command,-c<attached>and--config=headers, and an--add-dirURL path. - The canaries in a plain
claude -prun:were an assignment before the program,--auth-token,--passwordand--api-key=, and an--add-dirURL path. - Also checked: a quoted
run:, thepromptand API-key inputs, a non-JSONsettings, a JSON-listsettings, an unparsablemcp_configand aplugin_marketplacesURL with userinfo.
- The outputs checked were:
- Local checks at this head.
- 36 test files: 2323 passed, 4 skipped, 0 failed. They include:
test_workflow_agent_launches.py(298 cases);- workflow step references, reusable secrets, label redaction, workflow capability diff and workflow evidence;
- host audit, unread surfaces, unread changed inputs, host-diff docs, direction and review changes, and comparison coverage and partial comparison;
- distribution-surface parity, public surface, local contract, schema round-trip, determinism, docs links and the pilot ledger;
- install hooks, the control envelope and its rows, manifest-free PR rows, instruction contracts, org governance and host input recovery;
- the host-config, cold-start and semantic cold-start replays, and the governance benchmarks.
scripts/generate_schemas.py --checkexits 0,scripts/build-llms-full.pyleaves the tree unchanged,scripts/regenerate_goldens.py --checkreports 24 artifacts and 0 changed, andruff checkon the changed source and test files is clean.## 1.1.0in the CHANGELOG is byte-identical toorigin/main. The0.6schema files and.github/are untouched.
- 36 test files: 2323 passed, 4 skipped, 0 failed. They include:
- CI on
70e7dd53.suite (1/2/3),coverage,test,verify,verify-self,mcp extra (floor/newest),clean-checkout-launcherandwindows-launcherpass.release-tag-consistencyis skipped, as on every PR.
C7-F1: `may_still_meet` looked for unread steps only under the losing
job's name, which holds nothing once that job is renamed or removed, and
`job_left` accepted any job that no longer exists. So renaming job a to
a2 while quoting its `claude -p --dangerously-skip-permissions Review`
(R6), or running it through `npx` (R7), or removing a while an existing
job c gained the quoted launch (R3), as job b added the plain launch,
was one `low changed` row saying the bypass "moved between jobs
(a/steps[0] -> b/steps[0])", with no widening signal. The launch may
still run, unread, under a2 or in c. `may_still_meet` now also refuses
the move while any job other than the losing and the receiving one has
more unread steps of that agent, or read launches of it holding unread
text in an input the rule is read from, than it had at the base; a job
new at the head counts from zero. `job_left` applies the same check, so
both move paths refuse alike. R6, R7, R3, the same while the losing job
remains, and a rename into an unread argument input are widenings; a
plain rename, a rename with the install step the job keeps beside the
launch, and a rename beside another job that keeps its unread step stay
moves. Two jobs renamed at once, one of them holding an unread step,
cannot be told apart from R6 and now claim the gain, in the safe
direction; the support page, STABILITY, the CHANGELOG and the
`AgentRuleGains` docstring say so.
C7-F2: `_EXPRESSION_RE` (`\$\{\{(.*?)\}\}`) scanned to the end of the
text from every unterminated `${{`, so a job `env:` or an agent action's
`settings` input holding a few hundred KB of them stalled `diff`,
`audit --host` and `check` (92 s at 300 KB, 1070 s at 1 MB). The pattern
now ends at `}}` or the end of the text, as `_EXPRESSION_SPAN_RE` does,
and only a closed expression names a secret, so what `job_secrets`
publishes is unchanged. 300 KB and 1 MB now take 0.8 s and 1.0 s end to
end; a new case bounds about 360 KB of them, and the settings input, to
5 s.
Nonblocking, fixed:
- A declared `shell:` was read whenever its first word was `bash` or
`sh`, so `bash -c 'claude -p --dangerously-skip-permissions Review'
{0}` beside `run: claude -p Review` read the step and missed the
template. A template is now read only as `bash`/`sh` running the script
alone: `set` flags, `-l`, `-i`, `-r`, `-o`/`-O` with a name other than
`noexec`, and `--noprofile`, `--norc`, `--posix`, `--login`,
`--restricted`, `--noediting`, `--verbose`, then `{0}` last. Anything
else (`-c`, `-s`, `-n`, `--rcfile`, words after `{0}`) leaves the step
an unread, named limit. The schema description of `unread_agent_runs`
says so, and the 0.7 inventory and baseline schema files are
regenerated.
- The coverage line for a workflow that changed with no row said
"redacted values such as env values and apiKeyHelper are not
compared", which named nothing a workflow holds. A workflow's line now
says "text this entry does not read, such as a step's env or an unread
agent step, is not compared; audit --host names each unread agent
step". Other files keep their note. The capability_diff row in
docs/distribution-surfaces.md and its parity comment are updated.
|
Addressed review cycle 7. New head C7-F1 (P2): a launch that went unread in a renamed job was still counted as having left itCause. As the review says, Fix ( Evidence.
New cases in
Docs. The rename sentences now say the same thing in the C7-F2 (P2):
|
pengfei-threemoonslab
left a comment
There was a problem hiding this comment.
P1 — Respect the CLI end-of-options marker before comparing permission flags.
At head 3986c93da637a284295da23b980c8c5c581b02f6, _read_flags continues scanning past a standalone --, and _run_command also recognizes Claude's -p after that boundary. This turns prompt words into authority settings. Reproduced with the PR's _rows workflow helper: run: codex exec -- --yolo → run: codex exec --yolo produces zero rows. Likewise claude -p -- --dangerously-skip-permissions → claude -p --dangerously-skip-permissions produces zero rows. Removing -- activates a real bypass, so this is a missed widening, not just an overly cautious warning.
Please stop CLI flag extraction at standalone -- and only recognize Claude print mode before it. Preserve real flags before the terminator. Add regressions for both agents, including a dash-prefixed prompt canary that must not appear in published settings. Codex uses clap's standard positional PROMPT parsing: https://github.com/openai/codex/blob/main/codex-rs/exec/src/cli.rs.
I am addressing this finding in this PR. This is a coding-agent review, not human approval.
pengfei-threemoonslab
left a comment
There was a problem hiding this comment.
Integration review after #851 merged: a second actionable issue appears when combining the branches. #851 permits replacing a v0.6 saved baseline because its new hook/MCP fields are display-only. Applying that version allowance unchanged to #850 permits replacing a v0.6 workflow baseline even though it never recorded agent launches or checkout refs. The existing workflow-migration regression fails because --save-baseline returns 0 instead of refusing and preserving the old file.
Addressed in the prepared integration: keep v0.6 replacement for baselines without workflow grants; refuse replacement when a v0.6 baseline contains a workflow grant. Saved v0.7 baselines retain workflow evidence while still omitting hook/MCP display fields. Both branches' generated schemas and migration documentation are combined. The focused baseline regressions pass; the full integrated suite is running.
The earlier end-of-options finding is also fixed with ten new regressions, including both agents' actual bypass activation and prompt non-publication. No authority declarations or baseline acknowledgements were synthesized.
|
Addressed both review findings:
Validation: the full integrated test suite passed 12,855 tests (7 skipped); Ruff, generated schema checks, and self-check passed. Fresh verification of committed head reports |
|
Final integration is pushed at |
|
Final integration check at |
… in the agent's function (#865) (#903) * Read a Google ADK tool a local factory builds, and a tools list built in the agent's function (#865) visulate/visulate-for-oracle#526 bound `save_memory_tool = create_save_memory_tool()` in its root agent's builder, where the factory returns `FunctionTool` around a nested function; the reader stopped at the local assignment. A factory call bound once and unconditionally (in the agent's function or at a module's top level), or inline in `tools=[...]`, is now followed as syntax to the factory's one unconditional `return` of `FunctionTool(inner)`, a plain function, a name bound once to one of those, or another factory's call, up to four deep; the tool is the wrapped function, with the factory recorded in `import_path`. A factory that returns from several places or under a condition, recurses, is decorated or a generator; a wrapped function that is decorated, a parameter, or changed or handed on (Visulate's delegate sets `__name__`); a tool changed where it is bound; and a factory held outside the read scope stay named with why (`factory_return`, or the resolver's reason). A third-party factory keeps the answer it had. On visulate#526 at `ai-agent`, the two memory tools become not_established candidate additions beside the nine named delegates. MuhammadVT/smart-assignment#46 built `tools = [...]` in the agent's function with a conditional `tools.append(...)`: read as a dynamic tools expression, it lost the unconditional tools. Such a list with append/extend/insert/`+=` is read member by member; an addition under a condition or in a loop is named on the agent and never read as bound; any other use keeps it dynamic. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> Resolve repository-local imported tools for Google ADK and OpenAI Agents SDK (#864) (#879) * feat: resolve repository-local imported tools for ADK and SDK readers (#864) A tool an agent binds from another module was unresolved at the module boundary, so a PR adding one produced no row. A shared resolver (inputs/python_imports.py) follows a tools-list reference through static imports, package re-exports, module-qualified access and plain aliases to one function definition, reading only regular .py files inside the read directory through the bounded input reader, never importing or running them. Every stop is a named reason. - Google ADK: imported names, module.function, alias = function, and FunctionTool/LongRunningFunctionTool wrappers (inline, assigned, or built in the imported module). One tool per definition; same-named functions in different modules stay distinct via binding locators. - OpenAI Agents SDK: names and module.function reaching the SDK's @function_tool, with guard evidence for imported definitions. - diff --application: rows carry import_path evidence (modules, lines, digests) outside compared meaning; unresolved references are scoped to their agent with the reason. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> * fix(#864): import resolution never produces a false or complete answer (review round 1) Adversarial review of the shared resolver found answers that read as complete while wrong: - One SDK agent binding two same-named definitions from different modules kept the last and reported a false CHANGED row. Both readers now bind neither, whatever the list order, and name both definitions. - A name the function building the agent binds itself (a local import, a parameter) was resolved through the module's binding. It is now a named stop (`local_binding`); a nested function that is the only definition of its name is that definition. - `import a.b` then `a.b.f` read `a/__init__`'s own `b`. It now reads the submodule, and several `import a.x` statements are not a rebinding. - An ADK wrapper warning shared by two agents scoped its gap to one of them. Every unresolved-reference record now gets its own gap. - `scan` counted one definition twice when an import reached a module another configured source also reads, making `{tool: ...}` selectors ambiguous. The catalog keeps one observation, and the binding graph reaches it through the exact definition locator the reader resolved when the edge's own source has none. - The module-binding walk climbed a parent chain per node; it is now linear (a 744 KB nested module: 32 s -> 3.7 s). Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> * fix(#864): read a reference where it is used (review round 2) - A builder's own `from support import lookup` is followed like a module-level import (`ImportResolver.resolve_local_import`) instead of stopping, for bare names and ADK `FunctionTool(func=...)` arguments alike. - `nonlocal` follows the outer function's binding; a name a scope binds more than once is a named stop. - `tools.lookup = ...` / `setattr(tools, "lookup", ...)` in the module that binds `tools.lookup` is a named stop. - A factory's own `toolset = McpToolset(...)` / `tool = FunctionTool(...)` is read through the existing toolset and wrapper paths again, so the MCP endpoint and toolset checks return. - A module-level ADK agent binds the module-level `def`, not a nested one the flat function map happened to keep last (pre-existing on main). - An SDK list variable bound twice anywhere in the file, or changed in place, is dynamic. - The `scan` dedupe runs once where sources are loaded, so guard association and inventory completion see the same tools; a source an inventory completes keeps its observation; `./` spellings match. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> * fix(#864): read a list and a flat name where they are bound (review round 3) - SDK tool lists are read through the scope that binds the name at the construction. They are read only when that scope binds the name once, to a literal list, and no change site anywhere in the file (a method call, a subscript store, `global` / `nonlocal`) changes that binding. Each change site is indexed once, so the reader is linear. A module-level list's names are resolved where the list is written, not in the building function's scope (R3-1, R3-4). - A flat function, wrapper or toolset map answers a module-level reference only when its entry is the module's single top-level binding. Otherwise the resolver answers (a module import wins over a nested `def`). If the resolver cannot establish the binding, the same-named definition is named, with a tool-scoped issue, and its row is never established (R3-5). - A wrapper's `func` is read where the wrapper is written. - An enclosing package `__init__.py` that reassigns the definition is a named stop (R3-3). `_dotted` is iterative, so a chain thousands of attributes deep no longer crashes (R3-2). - The scan dedupe drops the removed copy's guard evidence too (R3-7). - The CHANGELOG states the inventory-completion ambiguity (R3-6). Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> * fix(#864): a guess is scoped to its tool, and a list's second handle counts (review round 4) - A guessed ADK binding no longer marks the agent's list incomplete. - That had dropped every per-tool `scan` finding of the agent for one name (R4-2). `scan` is back to the round-3 behavior: a shadowed, medium-confidence definition. The PR #400 tests pass unmodified again. - The comparison reads `tool_issues` as a tool-scoped gap and an agent-scoped gap. The row is `not_established` on whichever side it is present, added and removed included (R4-1). - `x = FunctionTool(func=x)` right after `def x` wraps that `def` and is not a guess. That was noise on byte-identical files. - The SDK list reader counts `alias = TOOLS` and `TOOLS` passed to any call but a read-only builtin or logging method as a change (R4-3). - An attribute of the resolved name reassigned in an enclosing package's `__init__.py`, or in a module it imports relatively, is a named stop. This covers an alias spelled from a package above (R4-4). - Same-file symbol lookup is a dict, not a scan of every tool per reference. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> * fix(#864): scope a guess to its candidates; read list arguments and patches precisely (review round 5) - A guessed ADK binding's reason covers the tool it bound. It also covers every tool the module's bindings of that name could give the agent: a `def`, an import, `x = other`, or `x = FunctionTool(func=f)`. It covers the whole agent (`ANY_TOOL`) only when one of those cannot be named. - A guess that binds a name the agent already lists still leaves a gap (R5-1). - A `try: import … except ImportError: def …` fallback no longer hides the agent's other changes (R5-2). - A literal list passed to a call counts as changed unless the call is: - a read-only builtin, `pprint`, or a logging method; - an SDK `Agent` or a copy (`clone`, `replace`) reading its own `tools=`; - a function the module can resolve that leaves that parameter alone. Only names bound to a literal list are followed into a callee, so no other module is read for an unrelated call (R5-3). - The package-patch check counts only attributes rooted at an import. It parses the modules an `__init__` imports on a separate bounded budget, so a package that re-exports many modules no longer exhausts the resolution budget (R5-4, R5-5). Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> * fix(#864): only reads keep a list literal; follow a guess's candidates (review round 6) - A literal tool list stays readable only while every use of it is a read. Allowed reads: iteration, indexing, comparison, a truth test, formatting, a read-only builtin or logging method, an agent's or copy's own `tools=`, or a function whose every use of that parameter is such a read. `+=`, a return, a tuple, `*args`, `**kwargs` or storing it in another container makes it dynamic (R6-1). - A guess's candidate tool names are followed to their definitions: an import through the resolver, or by its imported name when the module is outside the scope; `x = other` and `FunctionTool(func=f)` through the name they spell. A candidate that cannot be followed, or a wildcard import, scopes the reason to the whole agent (R6-2, R6-3). - A module that a package `__init__` imports relatively and that cannot be read (a link, a missing file) is a named stop, cached with the package (R6-4). A module read for patches is the same object a later resolution uses (R6-5). Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> * fix(#864): spreads and tests are reads; a package's routine imports never stop (review round 7) - A literal list is still read when it is spread (`[*COMMON, x]`, `print(*TOOLS)`), tested (`TOOLS or …` in a condition), or read through a dict method (`.get`, `.keys`, `.values`, `.items`). `x or y` whose value is the list is judged by its own use. A `globals()` or `vars()` call in the module makes its lists dynamic (R7-3, R7-1 b4). - The package-patch scan treats an import above the scope as the read's boundary, as every import does. It skips an optional module imported under `except ImportError`. A link, or a missing module that is not optional, still stops (R7-2). - Docs: drop a duplicated sentence. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> * fix(#864): code that runs first is checked, and unread code keeps a binding named (review round 8) - Every module on a resolution chain (the agent's own file included), the `__init__.py` of every package enclosing one, and every in-scope module those import are checked for a reassignment of an attribute named like a step of the chain. `import patches` in the agent's file, or in a module the chain re-exports through, is now a named stop (R8-2). The defining module handing its function on (`registry.lookup = lookup`) does not count, and every location is kept, so one never hides another. - A relative import in any of those modules that climbs above the scope runs code that is not read. It no longer disappears: the tool is named and its row is `not_established` with that import in the reason, in both readers; for `scan` the ADK module stays at medium (R8-1, which round 7 had made silent). An import under `if TYPE_CHECKING:` never runs and is skipped. - `sys.modules`, however spelled, and importing a module by `__name__` reach its lists like `globals()` does: an SDK list there is dynamic (R8-3). - A redirecting package `__getattr__` stays a documented residual (R8-4): a gap for every hook on a chain would make #864's own attest rows `not_established`. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> * fix(#864): every spelling of unread code keeps the caveat; a lazy package hook is proven (review round 9) - A caveat survives a `FunctionTool(func=f)` wrapper the imported module builds (R9-1), and an absolute import spelled through a directory above the scope (`from svc.patches import ...` with scope `svc/app`) is a caveat like the relative one (R9-2). - The defining module's hand-on exemption no longer covers a patch through its own import (`import tools as _me; _me.lookup = ...`, R9-3), and only `typing`'s `TYPE_CHECKING` skips a block (R9-4). - Ordinary code no longer stops a resolution (R9-5): every location an ambiguous import could mean is read, a generated `*_pb2` module is the boundary and any other missing relative module a caveat, and the patch scan reads up to 1024 modules. - A package `__getattr__` is established only when every return it can reach for the name gives that submodule (`import_module(f".{name}", __name__)`, `from . import name`); attest's two lazy loaders stay established, a redirecting hook is named (R8-4). Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> * fix(#864): the repository says what is its own code; harden the lazy-hook idiom (review round 10) - Whether an absolute import no file in the scope provides is the application's own code is read from the repository: the compared commit's tree for `diff --application` (set per side), the checkout for `scan`. A module or regular package at the root or under `src/` is; a directory without `__init__.py` only when it holds the named submodule. `from common.patches import ...` with scope `svc/app` is a caveat (R10-1), and SDK apps under `agents/` import the SDK again (R10-3: the round-9 ancestor-name rule had caveated every tool there). - The scope spelled from the repository root (`svc.app.tools` with scope `svc/app`) is read inside the scope, so it resolves and a patch module it names is checked instead of guessed. - The lazy-hook idiom requires an undecorated hook that never rebinds its parameter, `importlib` / `import_module` bound only by importing them, and a returned local bound exactly once, counting `for`, `with`, walrus and `except` targets (R10-2). - A generated `_version` module is the boundary like `*_pb2` (R10-4); the patch-scan budget message names what it scans. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> * fix(#864): every directory above the scope is an import root; unread links and subverted hooks are named (review round 11) - An absolute import is looked up at the repository root, under `src/`, and under every directory between the root and the scope, so `backend/app` importing `common` from `backend/common` is a caveat (R11-1). - A root entry that is a symbolic link or a submodule, spelled by the import, is unread code: a caveat (R11-2). `scan` outside a checkout reads the three directories above the scope instead of nothing (R11-3). - A store into `sys.modules` in code that runs first is a named stop, and a package hook is not trusted when the package rebinds `__name__`, patches `importlib`, or stores into `sys.modules` or `globals()` other than the idiom's own cache (R11-4). - The scope spelled from an import root is read inside the scope only through a regular package (R11-5). Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> * fix(#864): a package is not an import root; a sys.modules store is read by its key (review round 12) - A directory between the repository root and the scope that holds an `__init__.py` is imported through its parent, never from the path: SDK apps under a regular `app/agents/` package import the SDK again, and `app/types.py` no longer shadows the standard library (R12-1, a round-11 regression). - A store into `sys.modules` (`[...] =`, `setdefault`, `__setitem__`, `update`) is a named stop when its key names a module on the chain or is built on `__name__`, a caveat when it is computed (a plugin loader's `spec.name`), and nothing when it names another module (R12-2, R12-3). - `globals().update(...)` / `setdefault` in a package disqualifies its hook (R12-3). Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> * fix(#864): packages above the scope are read; sys.modules and globals() by allow-list (review round 13) - Every directory between the repository root and the scope is an import root again, package or not: a service run from `backend/` with a stray `backend/__init__.py` names `common.patches` (R13-1, a round-12 regression). A standard-library name, and the scope's own package on the way to it unless it holds the name imported, are not the repository's code, so SDK apps under `app/agents/` still import the SDK (R13-5). - The `__init__.py` of every package above the scope, and each module it imports, are read through the repository layout (the commit's tree, or the checkout) for the same reassignments (R13-2). - `sys.modules` and `globals()` are read by allow-list. A store whose key names a module on the chain, a package above one, or the framework's own modules is a named stop; `__name__` plus a literal is that module's own name; any other use is a caveat; a module rebinding its own name through them is a reassignment; a change to `__path__` is a caveat (R13-3, R13-4, R13-6). A package that rebinds `__getattr__` or `__path__` does not have a trusted hook. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> * Read what the packages above the scope import, and every spelling of a module's own object (#879 review, round 14) - Above the scope, follow each import the way the in-scope reader does: every package on the way to the module and each submodule named (`from .hooks import patches`, `import svc.lib.util`). A reassignment there counts only when rooted at the scope or at something that cannot be found; guarded and generated imports are exempt. - A standard-library name is exempt only at a root that is a regular package; interpreter-preloaded modules are always exempt. - Allow-list reads: comparisons, spreads, iteration, pkgutil.iter_modules(__path__), namespace keywords (get_type_hints(globalns=globals())), and patch.dict/setitem by key. - The module's own object through an alias, sys.modules.get, import_module(__name__), __dict__/vars() stores, a computed setattr, or handed to a function is a reassignment or caveat, never silent. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> * A local named vars or globals is a variable, not the builtin namespace (#879 review, round 14 corpus) siada-cli's DebugUtils.dump binds `vars = stack[-2][-3]` and iterates it; the bare-name rule read it as the builtin handed on and named a real definition change not_established. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> * Only a two-part patch on another module is exempt above the scope; read the scope's modules an ancestor imports and the module object anywhere (#879 review, round 15) - R15-1: above the scope, a patch is exempt only when it sets one attribute of another module file outside the scope that its package binds nothing else under. A longer path, a name imported from a module, a package attribute shadowing the submodule, or a module alias (`_t = tools`) counts, in both readers. - R15-2: a scope module that an ancestor __init__.py imports is read like the chain's own. - R15-3: a module file wins over a same-named directory without __init__.py. - R15-4: a link's blob text is never read as source. - R15-5: the module object used anywhere but an attribute, a plain alias, a comparison or a reader is a caveat. So are __dict__/vars().update, getattr(m, "__dict__") stores, f_globals, builtins.globals under another name, a reader of the module's own under a reader's name, `globals = globals`, and __path__ changes above the scope (pkgutil.extend_path aside). get_type_hints(fn, globals()) is a read. - R15-6: above-scope files are read in batched `cat-file --batch` per directory (1100 modules: 76 s -> 3 s). The read bound is named once. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> * Read a frame's namespace by allow-list; batch a directory only while it is small (#879 review, round 16) - R16-1: `sys._getframe(1).f_globals.get("__name__")` in a logging helper is a read. A frame's namespace is read by the same allow-list as globals(); any other use is a caveat. - R16-2: above-scope reads batch a directory's modules in one `cat-file --batch` only while the directory holds at most 16 MB of Python. A directory of generated or vendored modules is read file by file, as asked (peak RSS 305 MB -> 87 MB on the 60 MB case). - Docs: `sys.path` / `sys.meta_path` changes that make a same-named module elsewhere the one imported are not followed. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> * Bind the located definition, prove readers by their binding, check the lazy hook's import (#879 PR review) - The resolved locator decides between same-named definitions even when the edge's own source holds exactly one. After source deduplication, a `tools.lookup` imported as `shared_lookup` no longer binds the local `lookup`. - A read-only reader is proven by its binding: - a bare builtin only when nothing binds it and no star import may; - a standard-library reader only when imported from that module (`json.dumps`, `from pprint import pprint`); - a logging method only on `logging` or a `getLogger()` logger. A `print` imported from the application, or an `.info()` on its own object, makes the list dynamic. - A lazy `__getattr__` is the submodule idiom only when its alias imports the submodule asked for: `from . import alternate as memory` is not. perf: read materialized Git tree blobs through batched cat-file (#878) _materialize_isolated_tree spawned one `git cat-file blob <oid>` per file in both its entries and links loops. Once `diff --application --scope .` materializes the whole tree (#877), that dominated the run: one side of TencentCloud/CubeSandbox took ~180 s to archive (the #686 cost class). Blobs are now read by _isolated_blobs: one `cat-file --batch-check` types and sizes every object (a missing or non-blob object refuses with a ConfigError before any content is read), then `cat-file --batch` runs of at most 64 MiB each return the content, with strict framing checks. Only full object IDs reach the batch. The link-text reads in _scope_through_boundary_links use the same reader. Every blob is still hashed against its tree entry's oid before it is written; path handling, containment and escape checks, blobs-before-links order, placeholders and the final digest comparison are unchanged. The reads go through _run_git_dir (now accepting input=) and the existing _run_process boundary, so no subprocess call site is added; only the two line pins in test_adapter_static_only.py moved. diff --application --scope . with #877: CubeSandbox 412 s -> 46 s, dlt 203 s -> 45 s, identical rows. fix: compare application wiring past unrelated links and submodules; read google.adk.Agent (#877) * fix: compare application wiring past unrelated links and submodules `diff --application` at the root scope took the unscoped archive route, which refuses every symlink and gitlink. On 2026-09-25, 15 of the first 27 runs against open third-party SDK/ADK PRs exited 2 on a path the reader never opens: `CLAUDE.md -> AGENTS.md`, a linked skill directory, a `VERSION` link leaving the tree, a vendored submodule. - Materialize every scope, the root included, through the scoped verified materializer, which recreates links rather than refusing them and packs the tree instead of the history. - Never read a Python input through a link: a target this scope already reads is compared at its own path, and any other in-scope target is a gap over the link's path. Reading the alias made one agent two ambiguous ones and hid the real file's change. - Record gitlinks behind an opt-in `archive_tree(record_gitlinks=True)`, materialized as the empty directory an unpopulated checkout leaves. An unchanged gitlink commit is named in limits; any other is a coverage gap over its path. Other archive callers still refuse. - Stop turning the host-configuration census's link count into application coverage gaps. - Read `google.adk.Agent`, the package-root re-export, as an ADK agent constructor (`from google.adk import Agent`). Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> * fix: keep unread links and scope-level submodules from reading as removals PR #877 review found two cases where unread input became conclusive negative evidence: - Discovery drops a path that does not resolve, so replacing `agent.py` with a dangling link read as a definite removal, `compared`, with no gap. The dropped host-census gap had been the only fallback, and it also covered a source directory replaced by an absolute or dangling link. Census every link under the scope directly, without following it: gap a `*.py` link that does not alias an input the scope already reads, a directory link holding Python outside the scope, and a link resolving to nothing in the tree where the other side reads source at or beneath it. An unchanged `agent/VERSION -> ../../VERSION` still changes nothing. - A gitlink at the selected scope itself gave a gap with source ".", which covers no relative binding path, so `agent.py` read as a definite addition or removal. Make that gap scope-wide. fix(diff --application): key SDK identity on imports; unobserved agents are not removals (#873) * fix(diff --application): key SDK identity on imports; unobserved agents are not removals On speechmatics/speechmatics-academy#142, a LiveKit voice agent moved its tools into an Agent subclass that passes them through super().__init__. `diff --application` printed `compared` with four false REMOVED rows. Two defects: - Framework identity. Discovery counted any bare `@function_tool` as the OpenAI Agents SDK, and the SDK reader recognized `function_tool` and `Agent` by spelling alone, so `livekit.agents` symbols were read as the SDK's. Both now key on import provenance: the absolute `agents` / `openai_agents` package, or an unimported name (unless a foreign wildcard could supply it). Relative imports are no framework's signal. - Unobserved is not removed. An agent observed on one side whose file on the other side still assigns its name (or passes it as `name=`), via a construction no reader supports (subclass, factory, Agent[Ctx], clone), now records a scoped coverage gap for that agent: `partial`, rows `not_established`. Handoff-only references do not count as an observed construction. A genuinely deleted agent stays an established removal. Regression tests fail on the unfixed tree. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> * fix(diff --application): resolve SDK names per scope; imports keep an agent named Addresses the two P2 findings on #873. - Import provenance was flattened across lexical scopes: one file-wide import map, accepted if any origin was the SDK, so a LiveKit `Builder(...)` became an SDK agent when a sibling function imported the SDK under the same alias. `_SdkNames` now resolves each spelling in the scope that uses it (nearest binding scope, class bodies skipped for nested code, global/nonlocal obeyed; decorators and defaults in the enclosing scope). An import there must be the SDK's; a parameter or local assignment is not. `function_tool` is decided per decorator node, not as a file-wide set of spellings. The walk is iterative. - The missing-agent check ignored import bindings, so `from agent_factory import agent` (or `exported as agent`) still gave a definite removal. Import aliases now count as the file still naming the agent. Regression tests for both fail on the previous head. feat: compare application agent wiring without prior setup (#871) * feat: compare application agent wiring without prior setup * fix: address application comparison review findings * test: normalize colored CLI errors in review regressions fix(ci): prevent checkout module shadowing in non-GitHub recipes (#870) * fix(ci): prevent checkout module shadowing in installation recipes * test(ci): cover quoted and qualified Python installation commands Read agent launches in CI workflows (#823) (#850) * Read agent launches in CI workflows (#823) The workflow grant read triggers, token permissions, reusable-workflow secrets (#693) and step `uses:` references (#771), and nothing about how a coding agent is launched inside a job. Changing a claude-code-action's `claude_args` from `--allowedTools "Read"` to `--permission-mode bypassPermissions --allowedTools "Bash(*)"`, adding a `claude -p --permission-mode acceptEdits` run step, or checking out the pull request head in a `pull_request_target` job each gave "No static host-grant changes detected", and a move to `issue_comment` with `pull-requests: write` gave a row with no agent context. The workflow grant now lists, as text that is never executed, fetched or evaluated: - `agent_launches[]`: a step whose `uses:` is anthropics/claude-code-action, anthropics/claude-code-base-action or openai/codex-action (any ref, any case) with the documented permission inputs it sets, or a `run:` that is one literal simple command starting with `claude -p/--print` or `codex exec`, with its documented permission flags under their primary spelling. The prompt, `--model` and undocumented flags are not compared. `job_secrets` names the secrets the job references, as context only. - `checkout_refs[]`: each actions/checkout step's `with.ref`, null for the default. Each job's multiset of launches and refs is compared, never the step label, so a rename or reorder is quiet; a difference is one `changed` row on the existing workflow row naming `job/step` and both values. Direction is claimed only by documented rules a job's launches gain, read from literal values: bypassed permission checks (either spelling), bypassed approvals and sandbox, a danger-full-access sandbox, `safety-strategy: unsafe`, or a `*` user gate. Those raise `workflow_agent_widened_<added|changed>` and make the row widened; every other edit, including `--allowedTools "Bash(*)"` (#824's to rate), is `changed`. A workflow row whose workflow runs an agent ends its `why` with the untrusted-input trigger, write scopes, secrets and pull request checkout beside each agent step; it is a note and moves no direction. A compound `run:`, an expansion or an expression is `unresolved`, publishes none of its text and records a non-blocking coverage issue naming `job/step`, as an unread secret value does (#693): coverage stays complete, adding one is a row that claims no effect, and an edit inside one is not reported. Scripts, composite (#701) and unknown actions, and agents reached through npx/timeout/sudo are listed as unread surfaces. Every value, ref and secret name goes through the #802 label redaction; a rewritten one is null with `redacted` and is neither published nor compared. Contract 40 and host-grants 0.6 shipped in 1.1.0, so host-grants inventory, baseline and drift move to 0.7 (the 0.6 schema files are untouched) and the runtime contract to 41. A 0.4-0.6 baseline holding a workflow grant is incomparable (`baseline_workflow_agent_launches_unavailable`); one without a workflow stays comparable. Verifier 0.20 and capability diff 0.3 do not move. No check id is added or removed and `check` decides as before. Docs: the support page (tables, rules, note, limits, a new unread-surfaces bullet), a STABILITY "Migration Note: Unreleased", CHANGELOG `## Unreleased` above 1.1.0, the agent contract page, the distribution-surfaces `capability_diff` row and its parity comment, the Stop hook note, version tables and pins, and a rebuilt llms-full.txt. The pilot ledger's source-tree column was re-measured on the Route H fixture: identical cells to the 1.1.0 engine apart from contract 41 and inventory 0.7, with byte-identical diff rows. Tests: new tests/test_workflow_agent_launches.py covers the four reproduction cases, what is and is not read, every unsupported shape, direction rules and non-rules, the acceptance's negative controls, redaction with a CLI canary sweep across every published output, the 0.6 baseline migration, schema validation, and the same row on diff, verify, the PR comment, check, the control envelope and the Stop hook. Existing tests move to 0.7/41. The host-config and cold-start replays reproduce their committed outcomes and run-of-record scores, and the sample goldens are unchanged. Closes #823 * Read no widening rule from an agent input that holds an expression (#823) The support page and STABILITY say a documented rule is read from literal values only, so a value holding `${{ }}` never meets one. The `*` user gate did not apply that: `allowed_non_write_users: "${{ vars.USERS }}, *"` raised workflow_agent_widened_changed. Every rule now skips a value holding an expression, as the claude_args and codex-args rules already did, and a test pins the gate case. * Address review cycle 1 on agent launches in CI (#823) F1. `claude_args` and `codex-args` were split with the `run:` shell tokenizer, which gave up on a newline or an unquoted `(`, so a widening written the way the actions document it gave `changed`: `claude_args: |` on several lines, `--allowedTools Bash(git:*) --dangerously-skip-permissions`, a `# comment` line, or a multi-line `codex-args`. Each input is now split the way its action splits it. The Claude actions' parse-sdk-options.ts drops full `#` lines, makes `()|&;<>` literal and reads the rest with shell-quote (newlines are whitespace, `$NAME` is empty, an unquoted `#` ends the input, a `--` word is always a flag); codex-action reads a JSON array of strings or string-argv. The `run:` tokenizer is kept for `run:` steps only. The published `claude_args` is the text the action parses, so a full-line comment is neither published nor compared. F2. Any URL path made a whole setting `redacted`, so a `--dangerously-skip-permissions` beside `https://example.com/style-guide` gave no row, and every `plugin_marketplaces` value was never compared, under a limit that wrongly called it credential-shaped. The documented rules are now decided from the declared text when the workflow is read, before anything is withheld, and published on the launch as `widening_rules` (rule and setting), which the comparator keys on, so redaction never hides a rule. A URL publishes its scheme and host with `<redacted-path>`, as an MCP server URL does (#723), and the rest of the value is published and compared. Other credential-shaped text in a setting or checkout ref (a token shape, an assignment, a bearer or header value, URL userinfo) is published redacted and makes the workflow a blocking limit through `_uncompared_workflow_text`, as a redacted step reference does (#767). F3. JSON in `settings`, `mcp_config`, `--settings`, `--mcp-config` or any argument word was published verbatim, env values and apiKeyHelper included. A JSON object now publishes what `.claude/settings.json` and `.mcp.json` publish: key names, with `env`/`headers` values, apiKeyHelper and secret-named values `<redacted>`, as canonical JSON. A codex `--config` override under env, headers or a secret-named key publishes `<redacted>`. Text that starts like JSON and does not parse is withheld (`unparsed_json`, a non-blocking limit). The agent-launch canary sweep now carries JSON-shaped canaries and their digests across diff, audit, check, verify, the PR comment and verifier.json, and a second sweep covers the refused credential-shaped case. Nonblocking: a rule gained where the job launched that agent before only in an unread form is named and not claimed (the `unknown_before` rule); a quoted word starting with `#` no longer makes a `run:` compound; `anthropics/claude-code-action/base-action` is read as the base action; the STABILITY note says a prompt after a variadic flag is compared; and `__all__ =[` is spaced. Host-grants 0.7 is unreleased, so `widening_rules`, `unparsed_json` and the base-action agent value extend it in place; the 0.7 schema files are regenerated. The support page, STABILITY migration note, CHANGELOG Unreleased entry, contract summary, integrations Stop-hook sentence and the capability_diff distribution-surface row say the same. * Address review cycle 1 on agent launches in CI (#823) The second review of #850 at 25c13ce8 found two P1 and four P2 defects. The branch is rebased onto origin/main daa4ad5f. Attached JSON values are withheld (F1). A --settings={...} or --mcp-config={...} word inside claude_args, and codex's attached -c<override> / -c=<override>, published their env, header and apiKeyHelper values verbatim, because only a word starting with "{" was withheld. _withheld_words now splits a --name=value word and withholds the value, and reads codex's attached -c through _withheld_config, as clap reads it. The canary sweep carries both spellings. Redacted prose no longer refuses the comparison (F2). The #802 label redaction rewrites ordinary prose ("never print bearer tokens", "Authorization: headers"), and a redacted agent setting made the workflow a blocking limit. That hid every row beside it and made every check incomparable while the workflow existed. The rules are already read from the declared text, so a redacted setting is now compared by its published text and its widening_rules. uncompared_agent_launch_texts names it as a non-blocking limit, and the row cell shows the redacted text. A redacted checkout ref still refuses, as a redacted step reference does (#767), because it names the code a job runs. A renamed job's launch moves its rules (F3). _agent_rule_gains keyed a rule on its job, so renaming a job that launches a bypassing agent was a widening. agent_rule_gains now pairs a rule one job gains with the same rule another job lost, when the launch that met it left that job: the job no longer launches that agent, or the same launch (agent_launch_key less the job) now runs in the gaining job. The why names the move. A second job gaining a rule, or a different launch gaining one while the first job still launches that agent, still widens. An expression no longer turns off every rule, and the row says what it leaves unread (F4). Rules are read from literal text a ${{ }} expression cannot reach: - the words of claude_args or codex-args before the first expression, less the word it touches and any quoted run still open at it; - the elements of a JSON-array codex-args before the one holding it; - the gate entries that hold none. A setting holding an expression is published with holds_expression (the unreleased 0.7 schema extends in place), and a row that changes it says the text the expression reaches is not read. A gain where the job's launch held an expression before, in the input the rule is read from, is named and not claimed, as unknown_before is. That also fixes a false widening at 25c13ce8: replacing --model ${{ vars.CLAUDE_MODEL }} with --model opus beside --dangerously-skip-permissions was reported as gaining the bypass. Two non-blocking fixes: - _published_value reads each expression as one word, so an expression in a URL's userinfo is withheld with it rather than garbling the URL and publishing its path. - A codex --config value that starts like a table and does not parse, as string-argv leaves a quoted one, is withheld as unparsed_json. Rebase (F5). CHANGELOG keeps #853's #778 line beside #823's under github_action row. llms.txt and ai-search-summary state the source tree as contract 41, unreleased, ahead of the published v1.1.0 (contract 40), as test_public_surface_contract requires while the two differ. The pilot ledger's source-tree column was re-taken on the rebased tree, through ./shipgate beside the engine of e3c6cb0c, the commit v1.1.0 was cut from: - the only differences are contract 40 -> 41 and inventory schema 0.6 -> 0.7; - diff rows are byte-identical; - check JSON differs only in the launcher path its next action names. The support page, the STABILITY preamble and migration note, the CHANGELOG entry, agent-contract-current, the Stop hook sentence in integrations, the capability_diff row in distribution-surfaces with its parity comment, the 0.7 schema files and llms-full.txt are updated to match. * Address review cycle 2 on agent launches in CI (#823) The third review of #850 at c7f67550 found one P1 and one P2 defect and four P3 notes. The branch is on origin/main 44b9e05d; no rebase was needed. A JSON setting publishes its shape, not its free text (C2-F1). _withheld_json published the whole _redact_secret_values tree. That tree is the host readers' digest input, not what they publish: it keeps every string outside env, headers and secret-named keys. So an mcp-remote --header "Authorization: Bearer ..." argument in --mcp-config, and a hook's curl command in settings, reached diff text, diff --json, audit --host --json, the PR comment and verifier.json. _json_shape now keeps key names, numbers, booleans and null, redacts what the host readers redact, and replaces each other string with <withheld:...>, a 12-hex digest of redacted_config_sha256 for that string, so an edit to it is still a changed row. It keeps only the strings a host reader publishes: - a permissions.allow/ask/deny rule and a documented Claude Code setting's value (defaultMode, the switches, enabledMcpjsonServers entries); - an MCP server's command name and its URL's scheme and host, followed by the digest when they drop a command's arguments or a URL's query. A codex --config table or array is read under its key path, so mcp_servers.gh={command="gh", ...} keeps its command name. The canary sweep adds the mcp-remote header and the hook command in every spelling (action input, claude_args, CLI flag, codex -c table). Re-running the reviewer's two repositories through ./shipgate gives 0 canaries in every output. The documented bypasses written through listed inputs widen (C2-F2). - Claude Code settings written as JSON meet bypass_permissions when their defaultMode is bypassPermissions, read by claude_setting_values as the settings reader reads .claude/settings.json. This covers the action's settings input, a --settings value in claude_args, and the CLI's --settings flag. A path is not read, and a settings value holding an expression meets none. - openai/codex-action's permission-profile: :danger-full-access meets danger_full_access. ":danger-full-access" is Codex's reserved name for its built-in full-access profile (BUILT_IN_PERMISSION_PROFILE_DANGER_FULL_ACCESS). Mode inputs are now (input, value, rule) triples. settings and permission-profile join _RULE_SETTINGS, so a gain after an expression in either is named and not claimed, as for claude_args. P3 notes: - allowed_bots: "*" now reads "accepts runs triggered by any bot" through agent_rule_text. - Which of two gaining jobs a moved rule goes to no longer depends on declaration order: the same launch arriving is matched before a job that merely stopped launching the agent. The support page, the STABILITY migration note, the contract summary, the schema docstrings and the CHANGELOG entry now say what a structured value publishes instead of claiming it publishes what the host readers would, and list the two rules. host-grants 0.7 is extended in place; it is unreleased. * Address review cycle 3 on agent launches in CI (#823) Rebased onto origin/main 01777037 (#852, #821), which had already moved the unreleased runtime contract 40 -> 41, verifier 0.20 -> 0.21 and capability diff 0.3 -> 0.4. This change now extends contract 41 in place instead of minting it: one comment in schemas/contract.py names #821 and #823, and the texts that said verifier 0.20 and capability diff 0.3 were unchanged now say capability_diff registry row keeps #852's two added roots and both paragraphs; STABILITY keeps the #821, #823 and #827 notes, with #821's "host-grants stays 0.6" and #827's version sentence corrected for a tree that also carries host-grants 0.7; CHANGELOG keeps both Unreleased entries. llms-full.txt is rebuilt. The pilot ledger's source-tree column was re-measured on the rebased tree beside e3c6cb0c: identical cells except contract 40/41 and inventory schema 0.6/0.7, identical diff text, rows, check JSON and inventory apart from its schema version, and diff --json / verifier.json differing only in #821's schema versions and coverage members. A codex exec step that selects the full-access sandbox through --config was a changed row while --sandbox danger-full-access widened, and -sdanger-full-access was not read at all. The flag reader now reads a short flag's attached value as clap does (-s<mode>, -s=<mode>, -c<override>, -c=<override>), and a --config override that sets sandbox_mode to danger-full-access, or default_permissions to :danger-full-access (what codex-action's permission-profile input passes the CLI), meets danger_full_access, its value read as the CLI's parse_overrides reads it. As the CLI resolves them, the last override of a key counts, default_permissions outranks sandbox_mode, and a --sandbox flag outranks both, so an override beside --sandbox meets none. In codex-args such an override meets none, because the action appends its own --sandbox or default_permissions selection after codex-args. The support page and the STABILITY Direction bullet list the spellings. Withholding a quoted URL inside a JSON-array codex-args element no longer drops the element's escaped closing quote, so the published value stays the JSON array the support page describes. * Address review cycle 2 on agent launches in CI (#823) An agent CLI inside a double-quoted $(...) or a backtick substitution, or after a shell reserved word, got no launch, no row and no coverage issue: gh pr comment --body "$(claude -p --dangerously-skip-permissions ...)" read as "No static host-grant changes detected", while the support page said such a run: is listed as unresolved once for each agent CLI it starts at the head of a command. The word splitter reads a double-quoted substitution as one word and a backtick as no operator, and the head reader took then, do, { and ! for the command. _command_substitutions now lists the text of each $(...) and backtick substitution outside single quotes, double-quoted ones included, and an agent CLI heading a command inside one is an unresolved launch (shell_expansion for a single top-level command, compound_command otherwise), publishing none of its text. An escaped \$, a single-quoted '$(...)', a comment and $((...)) arithmetic are not substitutions, so echo "claude -p ..." stays unlisted. The scanner keeps an explicit stack and reads each nested substitution as `_` in the one around it, so no nesting depth recurses or re-splits text. When the run's quoting does not balance, each line's substitutions are read as well. A command's head is now read after shell reserved words (!, time [-p], {, if, then, elif, else, while, until, do, function NAME), and a command after one is not a simple command, so its launch is compound_command and never read. The compound_command and shell_expansion limit phrases, the launch schema description (host-grants 0.7 schemas regenerated), the support page and the STABILITY bullet say so. A read launch that became a form this audit does not recognise (npx, a path, codex options before exec) was worded "a step no longer launches an agent", though the step still starts one. The removed case now reads "a step no longer declares an agent launch this audit reads", and a row whose workflow still exists adds that the step may still start an agent in a way this audit does not read, naming those forms. codex with an option before exec is listed under Known unread surfaces and in STABILITY's "What is not read". Tests cover each shape the review named, the negative controls, the diff and audit --host route of the reproduction, the reworded removal for npx, a path and codex root options, and 20000 nested substitutions. * Address review cycle 3 on agent launches in CI (#823) A step that drafts a PR comment in a quoted here-doc, such as cat > comment.md <<'EOF' with "Reproduce locally with `claude -p ...`" in its body, was listed as an unresolved agent launch: a "high changed" row saying a step now launches an agent, and a GitHub coverage issue. The shell passes a here-doc's body to its command as input and, with a quoted delimiter, expands nothing in it, so the step starts no agent. The phantom also hid a real gain: a --dangerously-skip-permissions step added beside it was "not counted as a widening" because the job "launched the agent in a form this audit does not read" before. The substitution scanner read backticks and $( across the whole run:, and the word splitter read a body line starting with claude -p as a command head. _here_documents now removes each here-doc's body and closing line before the run: is split or scanned. <<, <<- and a quoted, escaped or partly quoted delimiter are read outside quotes, comments and $((...)) arithmetic, inside a $(...) or backtick substitution too; <<< is a here-string; bodies start after the opening line and follow in order when one line opens several. A body is never a command. Only an unquoted <<EOF body's substitutions, which the shell runs, are read, with quotes and # in it literal, so '$(claude -p ...)' there is listed and was missed before. A here-doc with no closing line is left in the text and read as ordinary lines, so a misread << hides no later command. Closing lines are looked up by content, so many here-docs cost one pass. Every shape was checked against bash with stub claude and codex functions. A launch that only became a form this audit does not read no longer counts as having left its job. The moved-between-jobs rule now takes a rule as moved only when the job that met it no longer exists, or when the same launch now runs in the gaining job and that job's own launches of the agent still run there or in the job it left (a moved step, a swap). Job a going from claude -p --dangerously-skip-permissions to npx @anthropic-ai/claude-code -p while job b's launch is edited into a bypassing one is now widened, where it was "moved between jobs". A step moved and edited while its job remains is claimed too. A checkout step on one side only, such as a default checkout in an added job, is worded "a step now declares a checkout" or "a step no longer declares a checkout" rather than "a checkout's declared ref changed". The support page, STABILITY's migration note and the CHANGELOG entry say which here-doc text is read, name the two shapes still read as lines, and state the narrowed move rule. * Name an unchanged instruction limit reached through an in-tree link instead of refusing (#822) An instruction file whose limit this entry cannot resolve, such as a SKILL.md whose metadata holds a non-string value, is named as an unchanged limit when a change leaves it alone (#721). Reached through an in-tree link the reader reads through (#700) - `.claude/skills -> ../.agents/skills`, a per-skill link or a file link - the same untouched file refused the whole comparison: diff, verify and the manifest-free PR comment printed `base_inventory_incomplete; head_inventory_incomplete` and no row, hiding a removed deny rule beside it. `unchanged_limits` asked blob_path_unchanged whether the source, the path the link is read under, was one regular file in Git, and a path through a link never is. blob_path_unchanged now resolves the path on each side the way the reader reaches it: from Git tree entries for the base and a commit head, and for a working-tree head without following any link (each component's own entry, each link's own text, the file's unfiltered hash). It holds only when both resolutions are equal - every link at the same path with the same text, every other component a directory - and the file they land on is the same regular-file blob at the same in-tree path. A link is followed only under the rules the reader and the base archive already use (#700, #711): a relative text that lands inside the tree after normalization, directories above where it lands, at most eight links, and links at one component of the path. A path no link reaches still takes the one `ls-tree` it took before, and blob object IDs are still compared, so no filter or textconv can make two byte sequences equal. Everything that asks the proof moves together: unchanged_limits in `diff --json` and verifier.json, the text and PR comment, the unchanged limits a partial comparison may carry (#808) and the shared plugin-reference limits check leaves out (#714). check's boundary result still cannot name a limit, so it now refuses these comparisons with unchanged_limits_not_representable, as it does for a limit at its own path; its rows, decision and violations do not move. No schema, member, reason code or check id is added. The metadata value is not coerced. tests/test_linked_unchanged_limits.py holds every layout (directory link, per-skill link, file link, a chain of file links) to the direct result on diff, verify, the PR comment and check; keeps the negative controls refused (skill added or edited behind the link, link retargeted to an identical copy, link text rewritten to land on the same file, link replaced by a directory or the reverse, a later hop retargeted, working-tree-only retargets and edits); and holds the proof to exactly the links the reader reads through, across dangling, looping, absolute, escaping, over-long, intermediate and nested links. The #812 coverage case that pinned the refusal is retired, and the static-only allowlist follows its two call sites down the file. * Address review cycle 1 on linked unchanged limits (#822) The STABILITY migration note said that where the other routes now compare past a limit reached through an in-tree link, `check` refuses with `unchanged_limits_not_representable` and its rows stay empty. That holds for an instruction file's limit, which `check` cannot leave out, but not for a plugin-reference limit: `_without_shared_plugin_reference_limits` asks the same unchanged proof, so a `parse_failed` `plugin.json` that is a file link to an unchanged target is now left out exactly as one at its own path has been since #714. `check` then compares and publishes the removed `deny` row in its boundary result and its control envelope's `capability_rows`, where it refused with `base_inventory_incomplete` / `head_inventory_incomplete` and no row, and `diff`, `verify` and the PR comment, which withheld `plugins/demo` as `partial` (#808), are `comparable` with the limit in `unchanged_limits`. `check`'s decision, violations and control state, and `verify`'s control state and next action, do not move. The note now splits `check` by whether it may leave the limit out, adds the `partial` to `comparable` move, and says "every row's value" where it said "every row". The CHANGELOG entry mirrors it, the distribution-surfaces row and `docs/host-boundary-support.md` no longer say any change behind the link refuses the comparison (an edited plugin manifest keeps it `partial`), and `tests/test_linked_unchanged_limits.py` pins the plugin manifest at its own path and behind a file link, on every route, with the edited-target control. The `blob_path_unchanged` and `_reader_path` docstrings now state the proof as necessary, not sufficient: it follows the links on the way to the path, not every condition the reader puts on reading a whole linked directory, and a side whose reader does not read the path carries no limit there. A test holds a head that adds a link inside the linked directory to a refusal on every route. The two pinned call-site lines in `cli/verify/git.py` move with the docstrings. * Address review cycle 4 on agent launches in CI (#823) Four review cycles each found a new shell form (quoted words, $(...), here-docs, comments, reserved words) that the run: reader mis-read. The cycle-4 finding was the fourth: a `#` comment was split as commands, so an apostrophe in a comment hid a launch and a command line in a comment invented one. Parsing arbitrary shell cannot converge, so the reader now claims only forms it reads exactly and names every other one as a limit. - A run: step is an agent launch only when it is one line of plain words (letters, digits and `_ . / : = , % + -`, separated by spaces or tabs), run by bash, sh or no declared shell:, whose program's file name is `claude` with -p/--print or `codex` followed by exec. Every POSIX shell runs such text as exactly those words; a cross-check of 2251 generated texts in sh, bash, dash and ksh agrees on every one. - Any other run: that mentions claude or codex as a word of its own is an `unread_agent_runs[]` entry (job, step, agent) on the workflow grant: a non-blocking coverage issue in audit --host that publishes none of its text, is never compared, so it gives no row, and never says whether the step starts an agent. It takes no gain from another launch; only when an unread step goes and a read launch is added in the same job is a rule that launch meets named and not claimed, since it may be that step rewritten. - claude_args and codex-args are read only as a plain list of words (the same characters and parentheses, across blanks and newlines, with no --settings or --mcp-config flag). Any other value, a ${{ }} expression included, is `unread_arguments`: published only as a digest, so an edit is a changed row, and read for no rule; a rule a launch gains where that input was unread before is named and not claimed. - A codex --config override publishes its key; its value is <redacted> under env, headers or a secret-named key, as written for sandbox_mode, default_permissions, approval_policy and model, and a digest otherwise. The word after a secret-named word such as --token is <redacted> and the value is then published redacted, a named limit. - The shell tokenizer, the command-substitution and here-doc scanners, the reserved-word reader and the shell-quote and string-argv emulations are removed, with the expression-prefix reading of argument inputs. JSON is read only in the settings and mcp_config inputs. Host-grants 0.7 is unreleased, so its schema changes in place: the launch `unresolved_reason` is only inputs_not_a_mapping, a setting may be unread_arguments, and the workflow grant adds unread_agent_runs. The support page, STABILITY, CHANGELOG, the current contract page and llms-full.txt describe the tightened reader. * Address review cycle 5 on agent launches in CI (#823) C5-F1: the move rule took a launch that only stopped being read to have left its job. With job a running `claude -p --dangerously-skip-permissions Review` in the base, quoting its prompt (an unread step) or running it through `npx` while job b added the same plain launch gave a `changed` row saying the launch "moved between jobs (a → b)" and "already met that rule in the job it left", with no widening signal. Job a still runs it. - `same_launch` now refuses a move while the losing job may still run the launch in a form this reader does not read: an unread step of that agent stands at a step label the lost launch held, the job has more unread steps of that agent than before (the launch may have moved to another index), or one of its read launches of that agent holds an expression or an unread argument input in an input the rule is read from. The review's M5 and M6 are now `widened` and name "a step no longer declares an agent launch this audit reads (a/steps[0])"; a real move, a rename, a swap, and a move beside an unread step the job already had stay moves. C5-F2: docs/integrations.md said an unread `run:` is a non-widening row that diff and the PR comment show. It gives no row; only an unread argument input is a row. The sentence is split accordingly. Nonblocking items fixed: - `_job_secrets` walks each container once, without recursion. A job `env` holding itself through a YAML alias raised RecursionError, and a ten-way fan-out eight levels deep did not finish; both now read in well under a second. - A `settings` or `mcp_config` value that is neither a JSON object nor a plain file path (path characters, and a `${{ }}` only as a plain context reference) publishes only a `<withheld:…>` digest. A comment line before JSON published the `env` value that JSON held. - The last of a repeated `--permission-mode` counts, in `claude_args` and in a plain `run:`, as parse-sdk-options and the CLI keep it. - The docs/distribution-surfaces.md capability_diff row no longer says an unread argument input or an unresolved launch is named only by the inventory and audit --host: a row reporting its launch names it too. The support page, STABILITY (Direction and What is withheld) and the CHANGELOG entry say the same. Tests: tests/test_workflow_agent_launches.py grows to 290 cases. The five new move-rule widening cases, the alias test, the two json-or-path cases and the four --permission-mode cases fail on the previous head; the moved-beside-an-unread-step case and an end-to-end diff/audit test of M5 and M6 are added beside them. * Address review cycle 6 on agent launches in CI (#823) C6-F1: the cycle 5 move guard only refused a move when an unread step of that agent stood at a step label the lost launch held, or the job had more unread steps than before. Merging job a's `npm i -g @anthropic-ai/claude-code` step into `claude -p --dangerously-skip-permissions Review` while job b added that plain step (V1), or removing an unread `echo "claude"` step while the launch became a quoted, unread step at another index (V2), kept the unread count and missed every held label, so the row said the bypass "moved between jobs (a/steps[1] -> b/steps[0])" and gave no widening signal. Job a may still run the launch. An unread step carries no text that tells which launch it is, so `may_still_meet` now refuses the move whenever the losing job has any unread step of that agent, wherever it stands and whether or not it was there before (option (a) of the review). V1 and V2 are now `widened`, with the "may still start an agent" caveat, and name "a step no longer declares an agent launch this audit reads". The cycle 5 guard that kept a move beside an unread step the job already had (`claude mcp add x`) now widens, in the safe direction; it moves to the widening cases, and a move beside a step that names no agent stays a move. C6-F2: `_codex_override` caught only `TOMLDecodeError` and `RecursionError`, and `tomllib` raises a plain `ValueError` for an integer past Python's 4300-digit limit, so `run: codex exec -c sandbox_mode=<5000 digits> Review` crashed `diff`, `audit --host` and `check` (exit 1) and `verify` (exit 4, internal_error). It now catches `ValueError`, which `TOMLDecodeError` subclasses; such a value is read as text and selects no sandbox, so the row is `changed`. `-c default_permissions=<4400 digits>` and `codex e --config=sandbox_mode=<4301 digits>` are covered too. Wording: the support page, STABILITY (Direction and What is not read) and the CHANGELOG say a launch has not left a job that keeps any named unread step of that agent or a launch of it with an unread input the rule is read from. The support page and STABILITY also name the one case where a launch that stops being read is worded as moved rather than as no longer declaring a launch: another job adds the same launch while the job it left keeps none of those, as when the launch became a script in the same change. Nonblocking items fixed: - docs/integrations.md: "One that gains a rule" read as the unread `run:` of the sentence before it; it now says "An agent launch that gains a rule". - The HostWorkflowAgentSettingV7 docstring, published as the 0.7 schema's description, states that a `settings` or `mcp_config` value that neither starts like a JSON object nor is a plain file path publishes only a `<withheld:...>` digest. The 0.7 inventory and baseline schema files are regenerated. * Address review cycle 7 on agent launches in CI (#823) C7-F1: `may_still_meet` looked for unread steps only under the losing job's name, which holds nothing once that job is renamed or removed, and `job_left` accepted any job that no longer exists. So renaming job a to a2 while quoting its `claude -p --dangerously-skip-permissions Review` (R6), or running it throu…
Closes #823
Problem
The workflow grant read triggers, token permissions, reusable-workflow secrets (#693) and step
uses:references (#771), and nothing about how a coding agent is launched inside a job. Onmain(44b9e05d), running the issue's reproduction:mainargs:claude_args--allowedTools "Read"→--permission-mode bypassPermissions --allowedTools "Bash(*)"No static host-grant changes detected.run: addsclaude -p --permission-mode acceptEdits --allowedTools "Bash(*)" "Summarize this change"No static host-grant changes detected.checkout:pull_request_targetjob starts checking out${{ github.event.pull_request.head.sha }}No static host-grant changes detected.trigger:pull_request→issue_commentwithpull-requests: writegrants write permissions to workflow jobs, no agent contextand
docs/host-boundary-support.mddid not list agent launch inputs as unread, so each zero-row result read as covered.Design
A bounded extension of the existing workflow grant, following the #771 precedent: values are compared as text and never executed, fetched or evaluated.
agent_launches[]on the workflow grant. A step whoseuses:isanthropics/claude-code-action,anthropics/claude-code-base-action(alsoanthropics/claude-code-action/base-action) oropenai/codex-action(any ref, any case) lists the documented permission/reach inputs it sets, from a table taken from each action's ownaction.yml. Arun:is a launch only when it is one line of plain words (letters, digits and_ . / : = , % + -, separated by blanks), run bybash,shor no declaredshell:(a template only when it runs the script alone: option words, then{0}last; neverbash -c '…' {0}), whose program after anyNAME=valueassignments isclaudewith-p/--print, orcodexfollowed byexec/e. It lists its documented permission flags under their primary spelling (--allowed-tools→--allowedTools,-s→--sandbox,--yolo→--dangerously-bypass-approvals-and-sandbox). The prompt,--modeland undocumented flags are not compared.job_secretsnames the secrets the agent's job references, as context only.claude_argsandcodex-argsare read only as a plain list of words: the same characters plus parentheses, across blanks and newlines, with no--settingsor--mcp-configflag. Any other value isunresolved_reason: unread_arguments. That covers a quote,${{ }},$, a backtick,#, a shell operator, a glob, JSON,--settings/--mcp-configand non-ASCII text. It publishes only a<withheld:…>digest, is compared by that digest (so an edit is achangedrow), and is read for no rule.checkout_refs[]: eachactions/checkoutstep'swith.refas text,nullfor the default. A checkout step on one side only, such as one in an added job, is worded as added or removed, not as a changed ref.changedrow on the existing workflow row namingjob/stepand the value on each side.--dangerously-skip-permissions/--permission-mode bypassPermissions, the last of a repeated--permission-modecounting, or JSON in thesettingsinput whosedefaultModeisbypassPermissions, one rule),--dangerously-bypass-approvals-and-sandbox, adanger-full-accesssandbox (sandbox,--sandboxin any spelling clap reads,permission-profile: :danger-full-access, or, in acodex execstep passing no--sandbox, a--configoverride ofsandbox_modetodanger-full-accessordefault_permissionsto:danger-full-access, resolved as the CLI resolves them; never such an override incodex-args, after which the action appends its own selection),safety-strategy: unsafe, or a*entry inallowed_bots(any bot) /allowed_non_write_users/allow-users(any user). The rules are decided from the declared text when the workflow is read and published aswidening_rules, so redaction never hides one. Gaining one raisesworkflow_agent_widened_<added|changed>and makes the rowwidened. Every other edit ischanged.access/riskare untouched.${{ }}expression before the action reads the input. So an argument input holding one is not read at all, and a mode orsettingsinput holding one meets none. A user gate's entries that hold none are still read ("${{ vars.USERS }}, *"opens the gate). A setting holding an expression is published withholds_expression: true, and a row that changes it says the text the expression reaches is not read for a rule.whyand not claimed.${{ }}expression or an unread argument input in the input the rule is read from. It may already have met the rule.uses:from a pinned SHA to@mainproduces no row #771: a step reference moved between jobs adds no scope.npx, quoting), so a launch that only stops being read has not left it. Nor has one while that job has, at the head, any unread step of that agent, wherever it stands (cycle 6), or a launch of it holding an expression or an unread argument input the rule is read from (cycle 5). Nor has one while any job other than the receiving one has more such steps and launches of that agent than it had before, a job new at the head holding any, so a job renamed while it quotes its launch or runs it throughnpxdoes not hide that launch under its new name (cycle 7). Two jobs renamed at once, one holding an unread step, therefore claim the gain, in the safe direction.whywith the job facts beside each agent step: an untrusted-input trigger (issue_comment,issues,pull_request_target,workflow_run), the job's write scopes, the job's secrets, and a checkout of pull request code. It is a note on the existing row. It is not a verdict and moves no direction.run:that mentionsclaudeorcodexas a word of its own becomes anunread_agent_runs[]entry holding only{job, step, agent}. Examples: a compound or multi-line command, quoting, an expansion, a redirection, a comment, a here-doc,npx,codexwith an option beforeexec, a script named after the CLI. Such an entry:job/stepinaudit --host, as Named reusable-workflow secret remapping is invisible to host capability diffs #693 does for an unread secret value;with:is not a mapping are non-blocking limits too..github/actions/*/action.ymlchanged on 5 of 456 real steps with no row #701), an agent CLI reached through a variable or a function, and a step'senv:/if:. They are listed on the support page under Known unread surfaces.settingsormcp_configinput publishes its key names, numbers and booleans.envandheadersvalues,apiKeyHelperand secret-named values are<redacted>. Every other string is a<withheld:…>digest, except the strings a host reader publishes: a permission rule, a documented setting's value, an MCP server's command name and URL host. So an MCP server'sargsand a hook's command are compared but never published.unparsed_json, a non-blocking limit).claude_args,codex-argsor arun:is never read, so it is never published.--configoverride in plain words publishes its key. Its value is<redacted>underenv/headers/a secret-named key; as written forsandbox_mode,default_permissions,approval_policyandmodel; and a digest otherwise.--tokenis<redacted>.unresolved_reason: redacted. In a setting it is compared as published, beside its rules, and named as a non-blocking limit, so a permission change or a rule gained beside it keeps its row. A checkout ref names the code a job runs, so a redacted ref refuses as a redacted step reference does (Preserve permission-change semantics when redacting host comparison rows #767).Versions. Host-grants 0.6 shipped in 1.1.0, so this bumps it instead of extending in place: host-grants inventory/baseline/drift
0.7(new schema files,0.6files untouched). Runtime contract41was minted after 1.1.0 by #821 (#852) and has not shipped, so this extends it in place;schemas/contract.pyhas one v41 comment naming #821 and #823. A0.4–0.6baseline holding a workflow grant is incomparable (baseline_workflow_agent_launches_unavailable), and one without a workflow stays comparable, by the #771 rule. #821's verifier0.21and capability diff0.4do not move here, because rows keep their shape.minimum_control_contract_versionstays 21. No check id is added or removed.widening_rules,unparsed_json,holds_expression,unread_argumentsandunread_agent_runsextend the unreleased0.7in place.Surface discipline (CONTRIBUTING § Surface discipline):
args,checkout) went from no row to a row, andrunbecame a named limit. Written as plain words,argsandrunalso give rows.expansion_signalsand the existing capability rows, and every route (diff,check, manifest-freeverify,verifier.json, the PR comment, the control envelope, the Stop hook) inherits the rows unchanged. The only versioned move is the grant schema, which the new members require.Docs and parity:
docs/host-boundary-support.md(support table, the input/flag tables, which argument text is read, direction rules and the three unclaimed gains, the note, what is withheld, the limits, and a Known unread surfaces bullet),STABILITY.md(preamble +Migration Note: Unreleased — workflow agent launches, beside #821's and #827's notes, whose version sentences now account for host-grants0.7),CHANGELOG.md(## Unreleasedabove## 1.1.0, which is untouched, beside #821's entry),docs/agent-contract-current.md,docs/distribution-surfaces.mdcapability_diffrow (one row, keeping #852's two added roots and its paragraph) with its comment intests/test_distribution_surface_parity.py(no claim added: direction is the engine's expansion signal, the unclaimed-gain sentences read the sameagent_rule_gainsthe signal is computed from, and the note reads published grant facts),docs/integrations.md(Stop hook), version tables/pins inAGENTS.md,llms.txt,.well-known/agents-shipgate.json,docs/INDEX.md,docs/architecture.md,docs/ai-search-summary.md,docs/passed-verdict-contract.md, and a rebuiltllms-full.txt.llms.txtanddocs/ai-search-summary.mdstate the source tree as contract 41, unreleased, ahead of the publishedv1.1.0(contract 40), astest_public_surface_contractrequires while the two differ.Evidence (measured before/after)
All measured with
./shipgatebesideorigin/main44b9e05d's engine, on synthetic two-commit repositories, before the rebases onto01777037(#852),997e9260(#860),eff60d97(#862) and8269922b(#861); themaincolumn reads the same ondaa4ad5f. The rebases move none of these rows: #852 adds coverage items, #860 routes enabled plugin hooks, #862 changes what averify --previewpointer binds and #861 publishes a partial comparison around an unreadable plugin directory, and none of them touches a workflow row.The reproduction, with the tightened reader (cycle 5 tree,
diff --base mainandaudit --hoston synthetic two-commit repositories):argsas the issue quotes it:--allowedTools "Read"→--permission-mode bypassPermissions --allowedTools "Bash(*)"islow changed,expands: false. The cell readsclaude_args (not read; digest <withheld:…>)and the why says the input is not a plain list of words. There is one non-blocking GitHub limit. The plain spelling--allowedTools Read→--permission-mode bypassPermissions --allowedTools Bashis⚠ low widened, "an agent launch now skips permission checks (bypassPermissions) (review/steps[1])".runas the issue quotes it (claude -p --permission-mode acceptEdits --allowedTools "Bash(*)" "Summarize this change") is an unread step. It gives no row, andaudit --hostnames it as one non-blocking limit. The plainclaude -p --permission-mode acceptEdits --allowedTools Edit Summarizeislow changed, "a step now launches an agent (review/steps[2])".checkoutandtriggerare unchanged from the lines below:critical changedand⚠ high widened.History, measured before cycle 4, when
run:and argument inputs were parsed as shell:args:⚠ low widened…review/steps[1]: runs anthropics/claude-code-action with claude_args: --allowedTools "Read"→… claude_args: --permission-mode bypassPermissions --allowedTools "Bash(*)"; whyan agent launch now skips permission checks (bypassPermissions) (review/steps[1]).run:low changed…→ … review/steps[2]: runs claude -p with --allowedTools 'Bash(*)' 'Summarize this change'; --permission-mode acceptEdits(the trailing prompt is a--allowedToolsvalue because the CLI's variadic option reads it that way, and the cell says so).checkout:critical changed…review/steps[0]: checkout of the default ref→review/steps[0]: checkout of ref ${{ github.event.pull_request.head.sha }}; why endsan agent runs at review/steps[1] (anthropics/claude-code-action) beside the untrusted-input trigger pull_request_target and a checkout of pull request code (review/steps[0]).trigger: same⚠ high widenedrow; why now endsan agent runs at review/steps[1] (anthropics/claude-code-action) beside the untrusted-input trigger issue_comment and the write scope pull-requests.Controls and unsupported shapes (
diffrows; GitHub coverage; non-blocking limits):mainwith: claude_argseditedrun: echo "claude -p --dangerously-skip-permissions"pull_requestworkflow, default checkout,pull-requests: writenpm ci && claude -p --dangerously-skip-permissionsaddedclaude -p $CLAUDE_FLAGSaddedrun: ./scripts/claude-review.sh --dangerously-skip-permissions./.github/actions/claudewithclaude_argsRows marked since cycle 4 were re-measured on the cycle-5 tree. As measured before cycle 4, in all twelve scenarios
check --format agent-control-json'scontrol_stateanddecisionare identical on both engines, and GitHub coverage stayscomplete.Review cycle history. The tables below record each earlier cycle's before and after. The shell forms they cover (here-docs,
$(…), quoting, compound commands) are unread steps since cycle 4. They give no row and are named inaudit --host. The cycle-4 and cycle-5 address comments give the current results.Review cycle findings (the second review at
25c13ce8; before =25c13ce8, after = this head):25c13ce8--settings='{…env…}' --mcp-config='{…env, headers…}'inclaude_argsdiff,audit --host,check,pr-comment.md,verifier.json--append-system-prompt "Never print bearer tokens…"besidepull-requests: read→writeincomparable(base and head incomplete), no rows, PR comment "unavailable"comparable,high widened"grants write permissions"; verifier comparable, 1 row.mcp.jsongains a servercheckincomparable, blocking audit issuecheckcomparable with theaddedrow; non-blocking issuereviewrenamedcode-review, same bypassing launchwidened,widenings: 1changed,widenings: 0, "moved between jobs … not counted as a widening"--model ${{ vars.CLAUDE_MODEL }}then--allowedTools Read→--dangerously-skip-permissionschanged, why claims no rule gainedchanged, why says the setting holds an expression and the text it reaches is not readallowed_non_write_users: ${{ vars.EXTRA_USERS }}→"${{ vars.EXTRA_USERS }}, *"changed, why claims no rule gainedchanged, why names the opened gate and why it is not claimed--allowedTools Read→--dangerously-skip-permissions --model ${{ vars.CLAUDE_MODEL }}changedwidened,widenings: 1--dangerously-skip-permissions --model ${{ vars.CLAUDE_MODEL }}→… --model opuswidened(false: the rule was already there)changed,widenings: 0Review cycle 3 (before =
6294321d, after = this head; baserun: codex exec -s workspace-write 'review',./shipgate diff --base HEAD~1):run:6294321dcodex exec -c sandbox_mode="danger-full-access"low changed⚠ low widened, "runs without a sandbox (danger-full-access)"codex exec -c default_permissions=":danger-full-access"low changed⚠ low widenedcodex exec -sdanger-full-accesschanged, "no permission flags"⚠ low widened,--sandbox danger-full-accesscodex exec -s workspace-write -c sandbox_mode="danger-full-access"low changedlow changed(--sandboxtakes precedence)A quoted URL inside a JSON-array
codex-argselement now publishes as valid JSON (the escaped closing quote is kept).Here-doc bodies and moved launches (the review at
b1073829; before =b1073829, after =846763fd; jobreviewwithcontents: read,pull-requests: writeand a checkout;./shipgate diff --base HEAD~1andaudit --host):b1073829846763fdrun: cat > comment.md <<'EOF'/Reproduce locally with `claude -p "review this change"`./EOFhigh changed, "a step now launches an agent (review/steps[1])"; 1 GitHub coverage issueNo static host-grant changes detected; no issue<<"EOF",<<\EOF,<<-'EOF'(backtickedcodex exec), aclaude -pbody line, or$(claude -p …)in a quoted bodyhigh changed; 1 issueclaude -p --dangerously-skip-permissions "Review"high changed,expands: false, "not counted as a widening: before, this job launched the agent in a form this audit does not read"⚠ high widened,expands: true, "an agent launch now skips permission checks (bypassPermissions) (review/steps[2])"cat <<EOF/$(claude -p x)/EOF(unquoted: the shell runs it)high changed; 1 issuecat <<EOF/claude -p x/EOF(a body line is input)high changed; 1 issuea(contents: read)claude -p --dangerously-skip-permissions 'Review'→npx @anthropic-ai/claude-code -p …, and jobb(contents: write)--allowedTools Read→--dangerously-skip-permissionshigh changed, "moved between jobs (a/steps[0] → b/steps[0]) … not counted as a widening"⚠ high widened, "an agent launch now skips permission checks (bypassPermissions) (b/steps[0]); a step no longer declares an agent launch this audit reads (a/steps[0])"lintwith a defaultactions/checkout@v4The 2026-09-15 23-PR corpus. The per-PR clones had been reaped, so at
0bcc6878I re-fetched only each PR's changed.github/workflows/*files through the GitHub contents API (file text only, nothing executed) and diffed them with both engines. 5 workflow files changed in 4 of the 23 PRs. One of them launches an agent: BerriAI/litellm#40935 adds a Codex-action workflow onissueswithallow-users: "*". I re-fetched its two workflow blobs in a later cycle: its rows are byte-identical to25c13ce8's, and againstmainits row count and direction are unchanged (added, already widening byissues: write). Re-fetched once more (text only) for this cycle, its rows read byb1073829and by846763fdare byte-identical: none of itsrun:steps namesclaudeorcodex, and the workflow is added, so neither the here-doc reading nor the move rule reaches it. Itswhyaddsan agent launch now accepts runs triggered by any user (allow-users: *) (classify/codex); an agent runs at classify/codex (openai/codex-action) beside the untrusted-input trigger issues and the secrets GITHUB_TOKEN and LITELLM_API_KEY. That is accurate, but the author explains the*in a YAML comment, so it is the 1/23 "already documented by the author" case the issue names, and automatic author-actionable yield stays 0/23. The other agent-launch PR in the corpus, ROCm/FlyDSL#1106, launchesclaudefrom a skill script and changes no workflow; it stays out of scope pending #828. The other four workflow files launch no agent; they were compared at0bcc6878(byte-identical rows) and were not re-fetched.Pilot ledger. The contract and inventory bumps require re-measuring the source-tree column of
docs/design-partner-pilot-results.md(guarded bytests/test_design_partner_pilot.py). On the tree rebased onto997e9260(with #821 and #860) I re-ran its Route H dry run with this tree's engine and with the engine ofe3c6cb0c, the commitv1.1.0was cut from. The cells are identical except runtime contract 40 → 41 and inventory schema 0.6 → 0.7:checkblock/criticalwith 4 violations and visible coverage, itsagent-boundary-jsonidentical apart from the workspace path, the host-onlyinithandoff writing no manifest or workflow, manifest-freeverifyexit 0 with 6 rows, drift with 4 expansion signals, anddiffcomparablewith byte-identical rows, 4 of 6 widening. Thedifftext is identical apart from commit ids and the inventory apart from its schema version;diff --jsonandverifier.jsondiffer only in #821's schema versions and coverage members. The published column stays #853's PyPI1.1.0measurement.Benchmarks. On the rebased head,
tests/test_host_config_replay.py,tests/test_cold_start_replay.py,tests/test_governance_benchmark.py,tests/test_governance_benchmark_baseline.pyandtests/test_semantic_cold_start.pypass: every committed replay outcome is reproduced and the runs of record re-score to their published values (row counts on the vendoredopenbootdotdev/openboot#136claude-code-review workflow are unchanged; only itswhygains the note).scripts/regenerate_goldens.py --check: 24 artifacts, 0 changed.scripts/generate_schemas.py --check: clean.Tests
tests/test_workflow_agent_launches.py(323 cases):argsgaining the bypass is one widened row, and the quotedargsa changed row that publishes only digests. A plainrunis one changed row, a head-ref checkout is one changed row, andtriggernames the agent step. Each row namesjob/stepand both values. The quotedrunis an unread form, covered by the unread-step cases;codex e, literal assignments skipped, and a plain list of words inrun:and inclaude_args/codex-args(several lines, parentheses);run:(each shell form earlier cycles found among them) is an unread step. It publishes no text, names a non-blocking limit and gives no row. Every other argument input isunread_arguments: a digest, achangedrow, no rule;codex execspelling of the full-access sandbox and the last of a repeated--permission-mode;Bash(*), Rate exec-equivalent Bash allow rules (interpreter -c/-e, npx, uv run, docker exec, xargs, env) as reaching arbitrary code #824's),acceptEditsand a replaced--permission-modearechanged;npx, at another step, an unread argument input, a settings expression), and while another job gains an unread step of that agent: the first job renamed while quoting its launch or running it throughnpx, or removed while a third job gains it quoted (cycle 7). The review's M5/M6 case is also checked end to end indiffandaudit --host;with:, apull_requestworkflow with a default checkout, a composite action;settings/mcp_configvalues, and a value there that is neither JSON nor a plain path;diffJSON/text,audit --host,check,verify,pr-comment.mdandverifier.json;${{, in linear time (cycle 7);shell:template that may run a command of its own (bash -c '…' {0},-s,-n,--rcfile, words after{0}) leaves the step unread (cycle 7);apiKeyHelper(cycle 7);0.6baseline with a workflow is incomparable. Without one it is comparable, and--save-baselineover it exits 2. The documented migration runs end to end. A0.7baseline compares launches and validates against the new schema files;diff, manifest-freeverify+ PR comment +verifier.json,check+ control envelope, and the Stop hook announcing the widening.Updated for the version bump:
tests/test_workflow_step_action_references.pyandtests/test_reusable_workflow_secret_mappings.py(legacy baselines drop the new members and name the new reason),tests/test_host_audit.py,tests/test_local_contract.py,tests/test_agent_instructions_{apply,renderers}.py,tests/test_org_governance.py,tests/test_instruction_structure_contracts.py,tests/test_host_input_recovery.py.Run locally at cycle 7 (on
origin/main8269922b):tests/test_workflow_agent_launches.pyand the files that read what the cycle touched, listed in the cycle-7 address comment. Earlier runs, at846763fd:tests/test_workflow_agent_launches.py(289 cases) and 36 other test files: those that read what this cycle touched (workflow, host diff, capability rows, coverage, distribution-surface parity, docs links, unread surfaces, public surface, schema round-trip, install hooks, control envelope, release decision, the pilot ledger, Stop verify --preview in a configured repository publishing human_review_required for input capture (#807) #862's preview currency, Keep independently established host changes when one scope is incomparable (#808) #861's partial comparison and hook-loading evidence), and the replay and benchmark files (test_host_config_replay,test_cold_start_replay,test_semantic_cold_start,test_governance_benchmark,test_governance_benchmark_baseline,test_benchmark_results_privacy): all pass. The benchmark runs of record reproduce unchanged.scripts/generate_schemas.py --checkis clean,scripts/regenerate_goldens.py --checkreports 24 artifacts and 0 changed,ruff checkon the changed files is clean, and rebuildingllms-full.txtchanges nothing.846763fd:suite (1/2/3),coverage,test,verify,verify-self,mcp extra (floor/newest),clean-checkout-launcherandwindows-launcherall pass;release-tag-consistencyis skipped, as on every PR. GitHub reports the PRMERGEABLE,CLEANagainstorigin/main8269922b.Deviations from the issue, and why
--allowedToolsvalues are not rated by its exec-equivalence tier:--allowedTools "Read"→"Bash(*)"ischanged, not widened, and the support page says rating a tool rule's reach is Rate exec-equivalent Bash allow rules (interpreter -c/-e, npx, uv run, docker exec, xargs, env) as reaching arbitrary code #824's.run:that mentions an agent CLI in a form this reader does not read is anunread_agent_runsentry (PM scope decision 2026-09-23). It is named as a non-blocking limit, following Named reusable-workflow secret remapping is invisible to host capability diffs #693's precedent, and is never compared, so it gives no row and claims no effect. Shell is not parsed: four review cycles each found a new shell form the parser misread. Launches that mention no agent CLI (a script, a composite action, an unknown action) are documented unread surfaces with no row and no limit, because naming every script or unknown action would be noise.actions/checkoutstep, not only in agent jobs, as the issue's "record it as text on the step" says. The pull-request-code classification is used only for the note and never for direction.danger-full-accessas a sandbox or as the built-in:danger-full-accesspermission profile,safety-strategy: unsafe, bypassing approvals and sandbox) are the "sandbox or approval mode on other agent actions" the issue names, taken fromopenai/codex-action'saction.ymland the Codex CLI's option definitions.