| Prompt injection (direct/indirect) |
control + documented residual |
Tool results on the wire are wrapped in TOOL_RESULT_BEGIN/END fences with inner fence markers neutralized (this fix); secrets scrubbed (E11). Residual: an injected instruction inside data can still steer the model — no framework fully solves this; delimiters + user-visible tool audit lines (what: …) limit blast radius. |
| Insecure output handling |
control |
Same fence treatment; result sizes capped (read_file 1 MiB cap, run_command 4000/2000 chars, job_poll 8 KiB window). |
| Excessive agency / tool misuse |
control (defense-in-depth) + documented residual |
Primary control: the per-call approval gate (every non-ReadOnly spawn needs consent). Secondary: check_destructive_argv blocklist refuses headline catastrophic patterns (teardown tools, recursive rm escaping the workspace, wrapper-nested variants) in BOTH run_command and job_start. The blocklist is best-effort by nature — wrapper/interpreter bypasses (sudo apt ..., timeout 5 dd ..., find / -delete, python -c rmtree) are expected and accepted residual risk; a real argv sandbox is a different product and explicitly out of scope. |
| Sensitive data disclosure (env) |
gap → fixed here (ENV-1) |
Children now inherit a scrubbed env: any var whose name contains API_KEY, _SECRET, _TOKEN, or PASSWORD is removed (GITHUB_TOKEN kept for gh/git auth — documented trade-off). |
| Supply chain (deps) |
control |
Cargo Audit + Trivy FS Scan required by rulesets on every PR and push. |
| Resource exhaustion |
control + fixed here (JOB-2) |
Provider 120 s timeout, subprocess 30 s ceiling, step budget 32, loop guard, read caps; job logs are read via bounded windows on poll (never a whole-log load). Since the 2026-09-18 scan triage, a ReadOnly poll no longer rewrites the log file — runaway log disk growth is bounded by the job's lifetime and clearing a runaway log is an operator action. |
| Session/message tampering |
control |
Append-only JSONL, CAS state transitions, seq monotonicity, torn-line recovery with truncation (#68); compaction never drops entries. |
| Injection via config |
control (MCP shipped) |
.cora.yaml/session headers are owner-controlled. MCP servers (#74) add an untrusted surface: server-supplied tool metadata is never trusted for risk tier (everything is Write → approval gate), tool results flow through the same fence/scrub path as native results, and the connection env is scrubbed. Residual: a malicious MCP server controls its own tool descriptions (model-visible) — treat server config as operator trust. |
| Identity & authz of sub-agents |
N/A |
tole v0 is single-agent, single-tenant. MCP servers are external tools, not sub-agents. |
| Human oversight |
control |
Risk-tiered approval gate (every non-ReadOnly call; Destructive never auto-allowed), #68 closed the replay-without-consent hole. |