feat: the Toolbox deployment shape — Argus runs without a Provider - #88
Merged
Conversation
Record the Deployment shape vocabulary (Toolbox, Colleague), the redrawn MEMORY/CONTEXT boundary and the platform-aware/platform-blind CodeHost distinction in CONTEXT.md. ADR 0023 conditions the MCP surface on whether the daemon can reason, amending ADR 0011. ADR 0024 records that credentials do not cross the network, which is why the Toolbox is the only shape deployable remotely.
A daemon with no LLM Provider configured no longer fails to start. The Deployment shape (ADR 0023) is derived at Build from one condition — is a Provider configured — and carried on the daemon Context as `Shape`, so every surface whose extent depends on whether Argus can reason reads one value instead of asking again. No configuration key selects it. The vocabulary lives in a new leaf package, pkg/deployment, because the shape is read by packages that have no business depending on each other: the daemon that derives it, the Channels, and `argus doctor`, whose checks stay cheap. Its zero value is the Toolbox — the floor both shapes stand on — so a Context assembled by hand promises reasoning only when it says so. What changes for an operator: - Startup prints the shape, and a Toolbox is told what it does not have and where the capabilities went. - `argus doctor` reports the shape and stops calling a Toolbox broken for having no Provider: the config and credential rows become informational. - Configuring the GitHub Channel's automatic review path in a Toolbox fails at startup, naming the reason — a channel that cannot work is a configuration error, never a silently dead one. The TUI Channel is deliberately not a startup error; a conversational turn requested there is answered with an explanation of the shape. - A Provider that IS configured must resolve at startup. Now that a missing Provider is a shape rather than a fault, a broken one has to fail loudly, or a typo reads like a working Toolbox. The environment-variable fallback still yields a Colleague, so an install that exported a key and never ran `argus init` does not silently degrade — .env is applied before the shape is derived so both spellings agree. Sessions still exist in a Toolbox (the deterministic capabilities are Session-scoped); they acquire no Provider, refuse the agent loop at its single entry point, and skip end-of-session memory curation. The MCP Channel's surface is untouched here. Closes #81
An external AI connected to a Toolbox is offered only what that daemon can actually serve. Review and Consult both start Argus's own agent loop, so where there is no Provider they do not exist: they are absent from `tools/list` rather than present-and-failing, and naming one — or one of the `start_review_*` family that drives the same loop — is answered with the reason instead of a bare "unknown tool". A client never spends a round trip discovering a capability does not work, and never offers the user something that will fail. The extent is derived from `Context.Shape` at request time, not captured when the Channel is constructed, so a daemon that has a Provider configured after it started serves the larger surface to the next client that connects — no rebuild. The handshake advertises the capability set that request will actually be served: Resources in both shapes, since the organization's knowledge needs no Provider, and tools when this shape has any. The new surface.go is the one place that answers "what does this shape expose?": a catalog carrying, per capability, its declaration, its handler and an explicit admission decision. `tools/list`, `tools/call` and the handshake all read that one answer rather than asking about the shape themselves, so a capability is exposed because someone decided it should be and never by default. A Colleague's surface is unchanged in every respect, and Resources behave identically in both shapes. Closes #82
The organization's CONTEXT reaches an external AI over MCP from a Provider-less Argus: list_context, read_context and write_context are on the surface, and what the caller and Argus worked out in one session is there for the next. In a Toolbox that knowledge is pulled by the caller, because Argus has no system prompt of its own to push it into. The important half is structural. The per-Session tool Registry the daemon already builds is now the one place that knows which Tools exist, and the deterministic surface is a filtered projection of it rather than a second hand-written declaration list that would drift the day a scanner is added. A Tool's wire declaration is its own Name, Description and Schema, so the two surfaces cannot describe a capability differently. The filter is an explicit decision per Tool, carried by the Registry itself: Register admits a Tool to Argus's own agent loop and nothing else, Expose also puts it on the surface. There is no default that exposes, so a Tool added later cannot leak outwards by being registered — proven from the protocol, where a Toolbox's listing is exactly the admitted Tools and none of the dozen others. The admission rule is ADR 0023's, inherited from ADR 0011: expose only what the caller does not already have. read_file, grep and list_files are therefore absent in both shapes — the calling agent has them already and better, and routing content through Argus makes the client pay for it twice. The deterministic surface is served in both shapes: a Colleague keeps review and consult above it. The Toolbox is the floor, not a degraded mode. This carries the first knowledge write on the surface, so it also settles where Role enforcement lives: at the MCP Channel, the way review's already is, because CONTEXT.md puts enforcement in the Channels and not in the shared Registry that Argus's own loop runs against. The policy is stated as the reads a viewer may make, so a capability admitted later is refused to a read-only caller until somebody decides otherwise — forgetting the list costs a viewer a read rather than costing the organization a write. Every call on the surface is audited and attributed to the Principal the bearer token resolved to. Closes #83
Argus remembers in both Deployment shapes. save_memory and mark_false_positive are ordinary deterministic Tools on the MCP surface, so a Provider-less daemon can be told something in one session and still know it in the next — which is the reason to install Argus rather than wire three scanners into your own agent in an afternoon, and the one thing the memory curator being a subagent would otherwise have denied a Toolbox. The important half is that this is not a second, weaker memory for the shape that has no curator. One Store owns MEMORY.md, and everything that writes it goes through that: the two new Tools, the curator's end-of-session rewrite, and the advisory a teammate accepts on a pull request. The curator stops being *the* way MEMORY is written and becomes one of two callers of the same write — it is handed the Store rather than a path, holds its lock across the run because its read-modify-write spans an LLM call, and its end-of-session behaviour is otherwise exactly what it was. The daemon's own out-of-band appender, which used to open the file itself under a lock of its own, is now that same call. There is no second write path left to keep in step. MEMORY gains a size ceiling, because a deterministic append-only memory written by an external client has none of the self-pruning the curator provided, and MEMORY is loaded into every call — unbounded growth is a fixed token cost on every single request. The ceiling is a signal, not an enforcement: past it a write still lands, in full, and returns an explanation that MEMORY is full, that nothing was truncated or dropped, and that material should migrate to a CONTEXT document. Both callers get the same signal on the same terms. This is the one place the Toolbox is worse than the Colleague, and this is how it is paid for rather than ignored. MEMORY joins SOUL as a read-only Resource. Without it a Toolbox's memory would be write-only: there is no Argus system prompt to load it into, so the only way what Argus remembers can reach the reasoning is for the caller to pull it, and writing it would be worth nothing at all. What mark_false_positive records keeps its glossary semantics — advisory, not a global mute. The record says so in the text that is read back into a system prompt, and a test pins the behaviour end to end: an accepted false positive does not stop a finding carrying the very same rule from coming back to the caller from a review elsewhere. A content-stable match is not sufficient grounds to silently drop a result. Both Tools are writes, so neither is a viewer read: a read-only caller is refused at the Channel where review's refusal already lives, and the refusal is observable — MEMORY never comes into existence. Closes #85
The organization's security workflow reaches an external AI from a Provider-less Argus, twice over. As Tools — list_skills, read_skill and read_skill_file are admitted onto the surface — for an agent that goes looking by name. And as MCP prompts, listable and retrievable through the protocol's own prompts methods, for a user who does not know Argus's tool names and should not have to: the Skills show up in their client's own prompt menu and they pick one. The second half is the point. A Toolbox that the client never thinks to call decays to nothing, so discoverability is a feature here rather than a nicety. The handshake advertises the prompts capability in both shapes, next to Resources, because the Skills need no Provider either. There is no second store. The prompts surface reads the same skill Catalog the Tools read, so the whole-bundle override is the one the daemon already applies: a user-curated bundle claiming a built-in's name wins body and supporting files together, on both surfaces at once, and supporting files stay sandboxed within their own bundle exactly as before. Proven from the protocol, where the names the prompt menu offers and the names list_skills reports are one list, and a retrieved body is byte for byte the one read_skill returns. Reading a Skill mutates nothing, so it is a viewer read on both surfaces — the read-only Role means the same thing whichever way the catalog is reached — and every list and retrieval is attributed to the Principal the bearer token resolved to. That settles a question CONTEXT left open: ADR 0005 intended skills as analyst+, the Tool layer still does not enforce it, and the first Channel that does gate has deliberately answered viewer. Where a Role bites is following a skill, on whatever capability the skill reaches for — which is the ADR's own "a skill cannot escalate the caller's permissions" arrived at from the other end. The guide, which claimed the analyst+ gate as fact, is corrected to match. One consequence is recorded where Skill authors are addressed rather than enforced at runtime: a Skill body now has two possible consumers, Argus's own agent and an external agent holding the client's tools plus Argus's exposed ones, so it cannot assume the Tools it names are present. It never could; what is new is that what is missing depends on who is reading. Nothing validates a skill's tool references, and nothing should. Closes #86
An external AI can now run Argus's real scanners against code on the daemon host. `run_semgrep`, `run_gitleaks` and `run_osv_scanner` are admitted onto the MCP surface, so a Provider-less Argus answers the question a developer installs it for: the scanners wired up and verified, driven by whoever reasons. The caller names the target — an absolute path on the Argus host, ADR 0024's same-machine constraint read from the scanner's end — because there is no Review to inherit one from. A path that is relative, absent, not a directory or not readable is refused by naming the problem: a scan that quietly found nothing because it ran nowhere is the worst answer a security tool can give, and the AI relaying it to a developer has to be able to say what to fix. The Snapshot workspace the MCP Session maintains stays the answer for the remote deployment and is deliberately not wired in here. Naming a target is the surface's freedom and not the agent loop's. Where a Review has pinned a checkout, a named path must be inside it, exactly as every other file-scoped Tool's argument is — so the argument cannot become a route for content inside a repository under review to talk the model into scanning the rest of the daemon's disk, and the confinement Argus promises a reviewed repository still holds. Omitting the path during a review keeps meaning "all of the checkout", so the review paths are unchanged. The prefactor is what made the rest easy. The command Runner the scanners shell out through was fixed where the Registry is built, so the surface projected from that Registry could not reach it and a test of a scan would have shelled out to a real binary. It moves onto the daemon Context as an injectable dependency defaulting to the real executor, and both Registry callers take it from there. No new abstraction: the existing Runner is relocated to where both of them can see it. The scanners are deliberately NOT viewer reads. A scan reads nothing of Argus's knowledge — it makes the daemon execute a binary against a directory on its host that the caller names, and gitleaks answers with the secrets it finds there. That is a capability an admin grants, not a consequence of being allowed to read. The Channel's Role policy is fail-closed by design, and this is precisely the case it was written to catch, so nothing was added to it. `argus doctor` goes on verifying semgrep, gitleaks and osv-scanner in a Toolbox, which is the shape that most depends on them. Closes #84
# Conflicts: # pkg/channel/mcp/surface_test.go # pkg/channel/mcp/toolbox_test.go # pkg/daemon/session.go
Review of the previous commit found the writers' lock reaching further than it should. A curation holds it across an LLM call, and MEMORY's snapshot was being taken under it, so every new Session — an MCP review or consult, a TUI turn, a PR review — and every save_memory would have blocked for the length of somebody else's model call. Reading MEMORY now takes no lock at all. What makes that safe is that every write lands by rename, so a reader sees one whole version of MEMORY or the other and never half of each; the truncate-in-place write that made the lock necessary is gone, which is also what the curator's `update_memory` always claimed to do. Two consequences follow. The MEMORY Resource reads through the mechanism rather than opening the file itself, so it cannot observe a rewrite in progress and the path is known in one place. And appending is now load-join-write, which is what lets it notice that a previous writer left MEMORY mid-line — a curator rewrite is whatever the model produced — instead of starting a list item on the end of somebody else's sentence. The other caller of the write is told about the ceiling too: the advisory a teammate accepts on a pull request carries the signal back into the thread, so "the caller is told, not silently truncated" holds on every path and not only over MCP. Also drops an exported Replace nothing called, and names the write's result for what it is — the state MEMORY is in — rather than for the verb that produced it.
# Conflicts: # pkg/channel/mcp/surface_test.go # pkg/channel/mcp/toolbox_test.go
There are now two answers to "what is Argus?", and the guide carried only one of them. A new page holds both: the shape is derived from whether a provider is configured, the toolbox is the floor the colleague stands on, and adding a provider is an upgrade to a running installation rather than a migration. The MCP channel page is rewritten around it. It no longer opens with "as a colleague, not a toolbox", and its list of what is exposed now names what actually ships: the three scanners and their absolute-path target, the CONTEXT tools, the memory tools with the 8 KiB ceiling and what the signal past it says, the Skill tools and the Skills as prompts, and the role that each of those needs. The rest of the guide stops implying a model is required to start: the landing page, getting started, configuration and the providers page, plus the GitHub channel's automatic reviews (which fail loudly at startup with no provider unless auto-enrolment is turned off) and the Kubernetes guide, which deploys a colleague. Closes #87
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Closes #80. Closes #81. Closes #82. Closes #83. Closes #84. Closes #85. Closes #86. Closes #87.
Argus refused to start without an LLM Provider. The first thing it asked of
anyone evaluating it was a token-billed plan, and for a large group of
developers that ask is a wall rather than a step: they already pay for a
coding-agent subscription — Claude Code, Codex, opencode — and cannot justify a
second, per-token plan on top of it. That subscription cannot be handed to
Argus either: it is authenticated in a CLI on the developer's own machine under
a named user, so it is not an endpoint a daemon can be pointed at, and a shared
daemon holding one such login would attribute everybody's work to one person's
account.
What those people were being denied was never the reasoning — their own agent
reasons perfectly well. It was everything around it: the real scanners wired up
and verified, the organization's knowledge, the memory of what was already
decided, the Skills.
Argus now runs without a Provider, as a Toolbox. The daemon starts, serves
the MCP Channel with its deterministic capabilities, and the reasoning is
supplied by whoever calls it. Configuring the first Provider turns the same
daemon into a Colleague: it reasons on its own behalf again — Review,
Consult, the memory curator, automatic PR review — and keeps the whole Toolbox
as well.
The two are not alternatives, and no setting chooses between them. The Toolbox
is the floor; the Colleague is a storey above it. The presence of a Provider is
the switch.
Governed by ADR 0023 (the surface is conditioned on reasoning, amending ADR
0011) and ADR 0024 (credentials do not cross the network — which is why the
Toolbox is the only one of the three candidate shapes that can be deployed
remotely at all). The vocabulary is in
CONTEXT.md.What's in it
Deployment shape, derived and never declared
pkg/deployment— a dependency-free vocabulary package holdingShape.Toolboxis the zero value, so a shape nobody set claims the smallerpromise rather than a reasoning it may not be able to do.
pkg/daemon/shape.go—ShapeOfis the one place the question is asked:Colleague as soon as a Provider is reachable, Toolbox otherwise. The
no-
providers:environment-variable fallback still counts as a Provider, soan install that exported a key and never ran
argus initreasons exactly asit did before instead of silently degrading.
configured but wrong still fails loudly, because a broken Colleague that
quietly downgrades reads like a working Toolbox.
GitHub Channel's automatic review path today, Slack the day that Channel
lands (a test fails on that day to say so). The TUI is never refused:
possession of the socket is how an operator administers the daemon, so a
conversational turn there answers with an explanation instead.
The MCP Channel: one endpoint, two extents
Toolbox,
review,consultand thestart_review_*family are absentfrom
tools/list, and naming one is a tool error that says why rather than"unknown tool".
a Provider on a running daemon and the next client that connects sees the
larger surface, with no rebuild and no restart.
included.
The Registry is the single source of truth
per-Session tool Registry the daemon already builds — never a second
hand-maintained list that drifts the day a Tool is added.
exposed by default and cannot leak onto the surface by accident. The rule,
inherited from ADR 0011: expose only what the caller does not already
have.
read_file,grepandlist_filesstay out — the calling agent hasthose already and better, and routing content through Argus makes the client
pay for it twice.
review's already does. Aviewer gets reads; every write is refused by name so the external AI can
relay it.
Scanners on the Toolbox surface
semgrep,gitleaksandosvare callable over MCP against an absolutepath on the daemon host, validated and rejected clearly when it is not
readable — consistent with the same-machine constraint of ADR 0024.
as an injectable dependency defaulting to the real executor, so both callers
take it from one place. An existing abstraction relocated, not a new one.
MEMORY: one mechanism, two callers, with a ceiling
save_memoryandmark_false_positiveare ordinary deterministic Tools.Argus's own curator subagent stops being the way MEMORY is written and
becomes one of two callers of the same mechanism; its end-of-session
behaviour in a Colleague is unchanged.
caller is told MEMORY is full and that material should migrate to CONTEXT —
told, never silently truncated. A deterministic append-only MEMORY written by
an external client has none of the self-pruning the curator provided, and
MEMORY is paid for on every call.
mark_false_positivekeeps its glossary semantics: advisory, not a globalmute.
Skills, twice
list_skills,read_skill,read_skill_file) and asMCP prompts, listable and retrievable through the protocol's own methods.
Same catalog, same bodies, no second store.
Argus's tool names can still invoke the organization's security workflow from
its own prompt menu.
Operator surfaces
what that means and what to do about it.
argus doctorreports the shape and checks only what applies to it. AToolbox is a deployment shape, not a broken Colleague — it is never
reported as an environment that is not ready for lacking a Provider.
Docs
CONTEXT.mdcarries the new vocabulary — Toolbox,Colleague, Deployment shape, platform-aware/platform-blind CodeHost, and the
redrawn MEMORY/CONTEXT boundary.
deployment-shapes.md, plus passes over getting-started,configuration, the MCP and GitHub Channel pages, Skills, LLM providers and
Kubernetes, so the guide carries both promises without either reading as
a degraded version of the other.
Tests
68 new tests, driven through the seams a client can actually observe.
The primary one is the MCP Channel over HTTP: real JSON-RPC bodies posted
to a real server. Handshake capabilities in both shapes,
tools/listcontentsin both shapes, absent capabilities failing by name, scanner invocation and
output, CONTEXT read/write, memory written through one capability and read back
through another, the ceiling signal, Role refusals, prompts listing and
retrieval.
The secondary one is daemon construction: builds with no Provider, builds
with one, preserves the environment-variable fallback, still fails loudly on a
Provider configured wrongly, refuses a Channel that needs reasoning.
No test shells out to a real scanner binary — execution is stubbed through the
Runner that this work relocated.
go build ./... && go test ./...is green.Out of scope
Recorded as decisions, not omissions: the platform-blind CodeHost and the
remote Toolbox (ADR 0024 makes them a deliberate second step); automatic
triggers in a Toolbox (impossible by construction — nothing can reason when
nobody has asked); the Provider CLI and the Runner-delegated shape (both
rejected in ADR 0023); advertising third-party subscription proxies; closing
the Tool-layer RBAC gap — this work gates the new writes at the Channel and
declines to widen it; multi-daemon tooling; spend control.