🇰🇷 한국어 | 🇺🇸 English
AI coding agents fail silently. They skip planning, miss context, and never check the result.
xm makes that harder to do.
A plugin toolkit for Claude Code. It grounds plans in repository evidence, executes the smallest sufficient change, and validates only the risks the change can introduce.
/xm:build "Build a REST API with JWT auth"
→ repository evidence → direct or x-plan route → native execution → risk-based validation
The default x-build route selects direct execution for bounded, deterministically verifiable work and uses x-plan for work that needs planning. Cross-vendor panels and the project lifecycle are opt-in. The benchmark data and case studies are internal measurements; they do not establish outcomes for other teams.
- Install
- Quick Start
- Why xm?
- Cross-Vendor Verification
- Plugins — x-plan · x-build · parallel topic PRs · x-op · x-review · mutation testing · x-solver · x-probe · x-eval · x-humble · x-dashboard · x-agent · x-trace · x-memory · x-humanize · x-recall · x-panel · x-wt · x-remote
- Quality & Learning Pipeline
- Architecture
- Configuration
- Troubleshooting
- Contributing
- License
xm uses Bun as its JavaScript runtime for testing, the dashboard server, and script execution.
Why Bun?
- Fast startup — scripts and tests launch instantly with no JIT warmup
- Built-in test runner —
bun testworks out of the box, no extra devDependencies - Native TypeScript/ESM — runs
.tsand.mjsfiles directly without transpilation - Zero-config HTTP server — powers
x-dashboardwith no npm dependencies
# macOS / Linux
curl -fsSL https://bun.sh/install | bash
# Homebrew
brew install oven-sh/bun/bun
# Windows
powershell -c "irm bun.sh/install.ps1 | iex"After installation, verify with bun --version (requires v1.0+).
Node.js >= 18 is still required for Claude Code itself. Bun is used for xm's own tooling.
/plugin marketplace add x-mesh/xm
/plugin install xm@xm -s userAfter installing, run once per machine to copy the trace-session hook into ~/.claude/hooks/ and register Skill matchers in ~/.claude/settings.json:
/xm init # install trace-session hook into ~/.claude/
/xm init status # verify install state
/xm init uninstall # remove hook + settings entries + routing block
/xm init --no-hooks # install CLI dispatchers without copying hooks
/xm init --claude-md # also add the xm routing block to ~/.claude/CLAUDE.md
/xm init --codex-md # also add the xm routing block to ~/.codex/AGENTS.md
Idempotent: safe to re-run. Existing hooks (e.g. mem-mesh) are preserved, and each write creates a timestamped backup of settings.json. Traces land in each project's .xm/traces/.
The same install flow is available from a terminal via xm setup (see Terminal CLI).
init is overloaded on its first argument — with a name it starts a project instead of touching ~/.claude/:
xm init aic # create .xm/build/projects/aic + register it
xm init . # name the project after the current directory
xm init --here # same as `xm init .`Reserved words (status, uninstall, install, help, and flags other than --here) always take the global-install route, so they cannot be used as project names — use xm build init <name> for those.
Install the xm umbrella CLI to run commands directly from your shell — useful for the dashboard, sync, memory, traces, etc., without entering Claude Code:
# Local install from this repo
bash xm/scripts/install.sh
# Or remote
curl -fsSL https://raw.githubusercontent.com/x-mesh/xm/main/xm/scripts/install.sh | bashThe installer writes ~/.local/bin/xm (override with XM_BIN_DIR; ensure it is on your PATH). On an existing installation it compares versions and asks before applying an available update; use --yes for unattended installs. When claude is on PATH, it installs missing marketplace plugins and updates installed ones after confirmation. When codex is on PATH (including Linux Codex-only hosts), it also generates the global Codex Skills/plugin manifest, runs codex plugin add xm@personal, and enables hooks. Start a new Codex thread after installation; run /reload-plugins in Claude Code.
xm setup installs the Skill-tracing hook into user-scoped ~/.claude/ (once per machine). xm init with no arguments does the same thing — it is the original name, kept working:
xm setup # install trace-session hook into ~/.claude/
xm setup status # verify install state
xm setup uninstall # remove hook + settings entries + routing block
xm setup --no-hooks # install CLI dispatchers without copying hooks
xm setup --claude-md # also add the xm routing block to ~/.claude/CLAUDE.md
xm setup --codex-md # also add the xm routing block to ~/.codex/AGENTS.mdWrites ~/.claude/hooks/xm-trace-session.mjs and merges PreToolUse/PostToolUse Skill matchers into ~/.claude/settings.json (existing hooks such as mem-mesh are preserved; a timestamped backup is created on every write). Use the bash route when you are outside Claude Code; otherwise /xm init is the preferred entry point.
--claude-md adds a marked routing block to ~/.claude/CLAUDE.md. The block tells Claude when to invoke xm skills, for example xm:write for a PR. A plain install does not add the block. After you opt in, xm setup and xm update refresh the block. Put your own rules outside the block, because each refresh replaces its contents. To remove only the block, delete it. The next install does not add it again. Each change to CLAUDE.md creates a timestamped backup.
--codex-md adds the same kind of block to ~/.codex/AGENTS.md. The rules in this block name skills as $xm:<skill>, the Codex form. The opt-in, refresh, and backup rules are the same as for --claude-md. The two flags can run together.
The begin marker of each block carries the template revision, for example v1. xm setup status shows the revision. If a block has a newer revision than your xm, xm setup leaves it unchanged and prints a warning. A block without a revision is replaced once with the current one.
Prefer setup in new scripts and docs: init reads as "start a project" everywhere else in the ecosystem, and a future major version will hand bare xm init over to the project route.
xm dashboard # start — uses ~/.xm/projects.json registry (all registered projects)
xm dashboard --scan ~/work # legacy multi-project mode: scan ~/work for .xm/ dirs (depth 4)
XM_DASHBOARD_SCAN=~/work xm dashboard # same, persisted via env var
xm dashboard stop # stop it
xm dashboard restart # stop + start — picks up new server code / a fresh served bundle
xm dashboard open # open it in your browser
# Project registry (~/.xm/projects.json)
xm project import ~/work # one-shot bulk-register all .xm/ projects under ~/work
xm project list # show registered projects
xm project add [<path>] # register CWD or given path
xm project remove <id|path> # unregister
xm project archive <id> # hide from dashboard without deleting
xm project gc # drop entries whose path no longer exists
xm sync push # push .xm/ state to your sync server
xm sync pull # pull state from your sync server
xm memory <subcmd> # save | recall | inject | list
xm build <subcmd> # build status / list / ...
xm config <subcmd> # show | get <key> | set <key> <val> | reset (--local | --global)
xm trace <subcmd> # execution traces
xm solver <subcmd> # structured problem solving
xm handoff [reason] # save session state (+ tier-2 detail archive)
xm handon # restore session state
xm handon --log # print the tier-2 detailed archive on demand
xm build handoff --mirror-status # inspect the mem-mesh mirror payload/status
xm build handoff --mirror-skip # dismiss a pending mirror (no mem-mesh setup)
xm which # show resolved lib paths
xm --market <cmd> # force the marketplace cache (ignores local source)
xm version
xm helpThe CLI dispatches to plugin libs in ~/.claude/plugins/cache/xm/, the Codex-only bundle in ~/.codex/xm/, or $XM_LIB. A Claude Code plugin installation is therefore not required on a Codex-only host.
Resolution order is $XM_LIB → source-repo cwd → whichever of ~/.codex/xm/ and the marketplace cache declares the newer version. A tie keeps the Codex bundle, and a bundle that cannot state its version never outranks a cache that can — so a global Codex install takes effect immediately, while a bundle left untouched no longer shadows a cache the Claude plugin has since updated. On a machine that has a local checkout, the marketplace cache is shadowed everywhere, so xm --market <cmd> (or XM_MARKET=1) skips every dev-local root and pins the run to ~/.claude/plugins/cache/xm/. Use it to reproduce what a plain user sees; xm which and xm version report which lib actually answered. xm install is the one exception to that order: it renders SKILL sources, and while the Codex bundle mirrors skills/ for its own runtime reads it carries none of the release metadata an install needs (.claude-plugin/plugin.json, skills.checksums.json), so when the resolved root cannot drive an install the command falls back to the newest marketplace cache root that can — and takes that root's lib/ too, never a mixed pair. The sync subcommand reuses the bundled x-sync lib, so you do not need to run x-sync/install.sh client separately.
The dashboard reads from a machine-local registry at ~/.xm/projects.json. Once populated, xm dashboard shows every registered project without --scan.
- First-time setup: run
xm project import ~/work(or any root) to bulk-register every existing.xm/project. Idempotent — re-running only updateslast_seen. - Auto-registration: when you run any xm command from a project directory, the dispatcher self-registers it. New projects appear in the dashboard without explicit action.
- Worktrees: a worktree of an already-registered repo is collapsed onto the main repo entry. Running
xmfrom any worktree updates the same registry entry — no duplicates. - Resolution priority:
--scanflag →~/.xm/projects.json→ legacy~/.xm/config.jsonscan_roots→ CWD only.
xm is published as a Claude Code marketplace plugin, but its 30 SKILLs can also be rendered into rule/steering formats consumed by other AI coding tools. A single source compiler (xm/lib/install/install-cli.mjs) emits per-tool artifacts.
# Interactive picker (scope + targets)
xm install
# or, when invoking the compiler directly
node xm/lib/install/install-cli.mjs --interactive
# Preview what would be installed (no fs writes)
node xm/lib/install/install-cli.mjs --list
# Install for one or more tools, project-local (default)
node xm/lib/install/install-cli.mjs --target cursor,codex,kiro,antigravity,opencode
# User-global install (~/.cursor/, ~/.codex/, ~/.kiro/, ~/.gemini/, ~/.config/opencode/)
node xm/lib/install/install-cli.mjs --target cursor --global
# Re-hash installed files against the manifest (R-SEC-13/15)
node xm/lib/install/install-cli.mjs --verify --target cursor
# Remove all xm-managed files; user content in AGENTS.md is preserved
node xm/lib/install/install-cli.mjs --uninstall --target cursor,codexPer-tool layout:
| Tool | Skills | Slash invocation | Hook |
|---|---|---|---|
| Cursor | .cursor/rules/xm-*.mdc (frontmatter: description, alwaysApply) |
agent-requested | .cursor/hooks.json (camelCase events) |
| Codex CLI | .agents/skills/xm-<skill>/SKILL.md aliases + plugins/xm/.codex-plugin/plugin.json + bundled Skills + marketplace |
$xm-<skill> (searchable) or $xm:<skill> after codex plugin add xm@<marketplace> |
.codex/hooks.json / ~/.codex/hooks.json (requires codex features enable hooks or [features] hooks=true) |
| Kiro | .kiro/steering/xm-*.md (frontmatter: inclusion: auto|manual) |
n/a | .kiro/hooks/xm-*.kiro.hook (informational only — Kiro cannot block) |
| Antigravity | .agent/skills/xm-*.md (project) or ~/.gemini/antigravity/skills/xm-*.md (--global) + shared AGENTS.md index |
agent-requested | not supported (no programmable hook API) |
| OpenCode | .opencode/skills/xm-*/SKILL.md (project) or ~/.config/opencode/skills/xm-*/SKILL.md (--global) |
native skill discovery | not emitted |
Safety:
<!-- xm:BEGIN v2 --> ... <!-- xm:END -->markers isolate xm content inside files shared with the user (AGENTS.md). Pre-existing user content is preserved.- Existing files are rotated to
.bak,.bak.1,.bak.2(max 3 generations) on first overwrite. Symbolic links abort. - Lock files use
O_EXCLatomic creation with a 60-second stale TTL. - Each install writes a manifest under the target's
xm/manifest.jsondirectory (for example.cursor/xm/manifest.jsonor~/.config/opencode/xm/manifest.json) with SHA-256 + HMAC self-checksum.--verifyrecomputes hashes;--uninstallrolls back exactly the recorded files. .codex/hooks.jsonis shared with other tools (e.g. mem-mesh): install/uninstall track ownership per handler and merge/remove only xm's own entries, leaving other tools' hooks untouched.R-SEC-02supply-chain guard: source SKILL.md hashes are verified againstxm/skills.checksums.jsonbefore render.--allow-unverifiedbypasses with a flagged audit entry.- Installs are idempotent — re-running with the same arguments produces zero diff.
See docs/multi-tool-install.md for the complete guide — capability matrix, per-tool install steps, manual verification in each IDE, security model, troubleshooting. The full design (PRD v2.1) is at .xm/build/projects/multi-tool-install/phases/02-plan/PRD.md.
xm update automatically re-renders skills to every installed LLM target (Cursor, Codex CLI, Kiro, Antigravity, OpenCode) when their global manifests are present, even when Claude is already at the remote version. Per-file SHA-256 diffing skips unchanged targets; Codex also receives a cachebuster and automatic codex plugin add so its active plugin cache cannot remain stale. Pass --no-propagate to update only the Claude plugin.
xm update # update plugin + propagate to all installed targets
xm update --no-propagate # Claude-only update, skip fan-out
xm install --propagate # re-render every installed manifest target on demand
xm install --list-installed # print installed manifest inventory as JSON/xm:build "Build a REST API with JWT auth"That single line:
- Inspects relevant code, contracts, tests, and existing behavior
- Checks whether the requested method is justified, then chooses direct execution or x-plan Standard
- Produces a grounded PlanEnvelope and asks for approval when the planned route needs it
- Executes sequentially by default with native agents, then runs only validation relevant to the changed risk
Need a plan without execution? Use x-plan directly:
xm plan --mode quick "Update src/auth.mjs and its focused tests"xm build plan is a deprecated alias for xm plan. When the resulting plan is executable and the workspace has an x-build project in the Research or Plan phase, it is imported into that project (PRD, tasks, steps); pass --replace to overwrite existing plan artifacts, or --no-import to skip the import. Every other case — a draft plan, no project, a later phase — saves the plan to .xm/plan and reports why it was not imported. The former PRD/task/phase planner remains available only as xm build legacy-plan.
Step-by-step tutorial (5 minutes)
# Plan only: repository inspection + readable PlanEnvelope artifact
/xm:plan "Build a user auth system with JWT"
# Plan and implement through the lean default workflow
/xm:build "Build a user auth system with JWT"
# Legacy lifecycle compatibility only
xm build legacy-plan "Build a user auth system with JWT"
xm build runMost AI coding tools work off a checklist: SQL injection, null check, N+1. Checklists find patterns. Senior engineers find problems.
The difference is the questions they ask before they act. Before filing a security finding: can an attacker actually reach this path? Before debugging: when did this last work? Before raising severity: am I inflating this because I'm not sure?
xm bakes those questions into every agent prompt. Agents end up reasoning about context instead of pattern-matching through a list.
Before & After examples
Code review (x-review):
| Checklist agent | xm agent | |
|---|---|---|
| Finding | [Medium] src/api.ts:42 — Possible SQL injection |
[Critical] src/api.ts:42 — req.query.id inserted directly into SQL template literal. Public API endpoint with no auth middleware. |
| Fix | Validate input. |
db.query('SELECT * FROM users WHERE id = $1', [req.query.id]) |
| Why | (missing) | Unauthenticated public endpoint, input flows directly to query sink |
Planning (x-plan, executed by x-build):
| Without principles | With principles | |
|---|---|---|
| Approach | "Using microservices because it's modern" | "Monolith with module boundaries — no constraint requires separate deployment" |
| Risk | "Security risks" | "JWT secret rotation may invalidate active sessions — mitigate with grace period" |
| Done criteria | "Auth works properly" | "JWT endpoint returns 401 for expired token, refresh rotation tested" |
Debugging (x-solver):
| Typical AI | xm | |
|---|---|---|
| First action | Generate 5 hypotheses | Describe current state + find last known-good baseline |
| Evidence | "It seems like the issue is..." | "git bisect shows regression in commit abc1234, confirmed by test output" |
| Stuck | Retry same approach | Switch layer (was checking app code → now check infra/config) |
Thinking principles at a glance
| When you... | xm principle | Tool |
|---|---|---|
| Review code | Context determines severity — same pattern, different risk depending on exposure | x-review |
| Review code | No evidence, no finding — trace it in the diff or don't report it | x-review |
| Review code | When in doubt, downgrade — over-reporting erodes trust | x-review |
| Plan a project | Decide what NOT to build first — scope by exclusion | x-plan |
| Plan a project | Treat the requested method as a hypothesis; prefer the simplest adequate path | x-plan |
| Execute a plan | Sequential by default; validate only the risks the change can introduce | x-build |
| Solve a problem | Diagnose state before hypothesizing — what's happening, not what's wrong | x-solver |
| Solve a problem | Anchor to known good — no baseline, no chase | x-solver |
| Solve a problem | Compound signals — never conclude from one log line | x-solver |
| Reflect | Why happened · Why found late · What to change in the process | x-humble |
How a senior engineer debugs — the thinking protocol embedded in x-solver:
REPRODUCE ──→ DIAGNOSE ──→ HYPOTHESIZE ──→ TEST ──→ REFINE ──→ RESOLVE ──→ REFLECT
- "Can I make it fail on demand?" — Record the command, its output, and a failure marker before you touch anything.
- "What's happening right now?" — Describe the observable state, not the problem.
- "When did it last work?" — Find the baseline. No baseline = find one first.
- "Why?" — with evidence — Corroborate from different sources. No evidence? Stop.
- "Stuck? Change the lens." — All hypotheses from the same layer? Look at a different one.
- "Show me it works." — Re-run the recorded failure. Execution is the only proof.
- "Why did we miss this?" — Retrospect via x-humble.
Common reference material lives in references/ (synced to marketplace as xm/references/). Skills pull these in on demand — progressive disclosure keeps each SKILL.md lean.
| Reference | Used by |
|---|---|
ask-user-question-rule.md |
7 plugins (Dark-Theme rule for AskUserQuestion) |
trace-recording.md |
9 plugins (trace hook protocol) |
dimension-anchors.md |
x-op strategies, x-review lenses, x-eval rubrics |
self-score-protocol.md |
all x-op strategies, x-agent solve/consensus |
finding-severity.md |
x-review, CLAUDE.md code review principles |
Single-vendor AI harnesses — including Claude Code's own /code-review ultra — orchestrate one model family. They structurally cannot have a competitor's model check their work. xm can: it spawns external model CLIs (claude + codex + cursor + agy + kiro) directly, so a finding, a plan, a score, or a hypothesis can be adversarially checked across different model families. Different vendors have different blind spots — agreement across them is real confidence, and a lone dissent is often the blind spot one family would have shipped silently.
One engine (xm panel cross) backs every layer. --cross-vendor is opt-in everywhere and degrades gracefully to single-vendor when fewer than two CLIs are installed — single-vendor stays the fast, cheap default:
| Layer | Plugin | What --cross-vendor does |
|---|---|---|
| Primitive | x-agent fan-out/broadcast | each parallel agent runs on a different vendor |
| Generation | x-solver | candidates/hypotheses generated across model families |
| Planning | x-build consensus | architect/critic/planner/security roles split across vendors |
| Deliberation | x-op debate/council | PRO/CON/JUDGE are genuinely different models |
| Review | x-review | findings cross-checked — consensus vs. diversity |
| Evaluation | x-eval | judges from different vendors, bias-reduced scoring |
| Engine | x-panel | the cross-model adversarial panel itself |
All layers share one provider definition: which CLIs exist and how each is spawned lives in the panel's adapters (code), not config. There is no per-plugin provider setup — panel.* config tunes panel review only (models/judge/stream), and the cross path shares just timeout_s. To add or retarget a vendor, you edit that single definition and every layer picks it up.
Several of those CLIs are themselves multi-vendor gateways — cursor and kiro front Kimi, DeepSeek, GLM, Gemini, Grok and more — so --models cursor:kimi-k2.5 works out of the box. Model catalogs move fast, so xm does not hardcode them: run xm panel types to see each installed CLI's live model-list command.
Default via config. Cross-vendor is opt-in per run. To make it a consumer's default, set .xm/config.json:
{ "cross_vendor": { "default": false, "review": true, "eval": true } }Precedence: --cross-vendor / --no-cross-vendor flag > cross_vendor.<consumer> > cross_vendor.default > false. Consumers: review, op, eval, solver, build, agent. A configured default still requires ≥2 installed + authenticated vendors (xm panel doctor); otherwise it falls back to single-vendor loudly.
For x-review, x-panel is an optional Phase 3 backend, not a second review: x-review keeps ownership
of lenses, severity, lifecycle, verdict, and convergence. Machines with a multi-model gateway may
set review.models to exact slots (for example codex:gpt-5.6-sol:xhigh and
codex:claude-sonnet-5) while enabling cross_vendor.review; other machines keep detection and
the single-runtime default. Two slots on one provider are multi-model sources, not two vendors.
This is a capability, available today; proving it produces measurably better outcomes is a separate, ongoing effort (see docs/strategy/xm-differentiation.md).
18 plugins, each installable individually or bundled via xm.
| Plugin | Purpose | Key command |
|---|---|---|
| x-plan | The planning engine — plan + PlanEnvelope, no lifecycle | /xm:plan "goal" |
| x-build | Repository-grounded plan → native execution | /xm:build "goal" |
| x-op | 18 multi-agent strategies | /xm:op debate "A vs B" |
| x-review | Judgment-based code review | /xm:review diff |
| x-solver | Structured problem solving | /xm:solver init "bug" |
| x-probe | Evidence-grade premise validation | /xm:probe "idea" |
| x-eval | Quality scoring & benchmarks | /xm:eval score file |
| x-humble | Structured retrospective | /xm:humble reflect |
| x-agent | Agent primitives & teams | /xm:agent fan-out "task" |
| x-trace | Execution tracing & cost | xm trace list |
| x-memory | Cross-session memory | /xm:memory inject |
| x-dashboard | Web dashboard for .xm state | /xm:dashboard start |
| x-humanize | Remove AI writing patterns (v0.6.0, pre-stable) | /xm:humanize audit text |
| x-recall | Cross-session artifact index | xm recall list |
| x-panel | Cross-model adversarial review | xm panel |
| x-wt | Session worktree — isolate & land back | /xm:wt |
| x-remote | Drive a remote host session from Discord | xm remote start |
| xm | Bundle + config + pipeline | /xm pipeline release |
Bundled in xm core (not separate marketplace plugins): /xm:ship release automation · /xm:write PR, issue, and release documents · x-sync multi-machine sync server · /xm:toss + /xm:inbox cross-project bug handoff · /xm:relay local Claude and Codex session messages plus AGY session messages — see x-ship, xm:write, x-sync and toss / inbox below.
/xm:mutate is bundled the same way — it drives the xm mutate command that ships with core. Codex exposes it as $xm:mutate, with $xm-mutate kept as a flat alias. See Mutation testing.
/xm:batch is bundled the same way. It drives the xm batch command of x-build. See Parallel topic PRs.
The planning engine. Everything in xm that plans — x-build included — goes through it.
x-plan reads the repository before it writes anything: what the code already does, which contracts exist, what the tests currently cover. It treats your request as a hypothesis rather than an order, so it will tell you when the repo already provides what you asked for, or when a smaller path reaches the same goal. Out the other end comes a readable Markdown plan plus a machine-checked PlanEnvelope under .xm/plan/.
/xm:plan "add refresh-token rotation"
xm plan --recommend --json "<requirements>" # which mode fits: Quick / Standard / UltraQuick writes the deterministic scaffold and stops. Standard adds repository inspection, a focused interview, and a critique pass. Ultra layers multi-model architect/implementer/critic synthesis on top of Standard.
Verified repository facts, inferences, and decisions that are yours to make stay separated in the output. A plan is never marked executable on a guessed path, API, or validation command.
x-build is the lean execution workflow around x-plan. It inspects repository evidence, hands planning to x-plan alone, runs native agents sequentially by default, and picks only the validation that directly observes a changed risk.
The default is adaptive, not plan-always. Bounded, independent, low-risk work takes the direct route — but only when deterministic gates already cover its failure modes. Shared, high-risk, or weakly observable work takes the planned route instead: x-plan Standard on the configured planner model, then the configured executor model once the plan succeeds. A missing or non-executable plan artifact stops it before Execute.
route start → verify → finish binds the decision to the baseline commit, expected files, gates, byte hashes, elapsed time, and cost events. A failed direct verification can restart once from a clean planned fallback. Stale or incomplete receipts fail closed.
xm build route decide --kind bugfix --scope bounded --independent \
--files src/a.mjs,test/a.test.mjs --risk low --failure-modes 2 \
--gates test,boundary --json
xm build route start --decision-id <id> --expected-files src/a.mjs,test/a.test.mjs \
--gate-cmd 'test=bun test test/a.test.mjs' --gate-cmd 'boundary=node scripts/check-boundary.mjs'
# native agent edits the files
xm build route verify --decision-id <id> --json
xm build route finish --decision-id <id> --jsonMeasured A/B proof is deliberately scoped: across 10 paired runs each, independent module/config fixtures kept verification at 10/10 with no blind-quality losses while cutting p50 cost by about 48% and p50 latency by 65–69%. A ReDoS fixture escalated every time and was slower and more expensive, so that class is routed to planned execution instead of being presented as a universal win. xm build route prove machine-checks pair coverage, quality non-inferiority, ≥20% cost savings, and ≥15% p50 latency savings before a claim can pass.
/xm:build "Build a REST API with JWT auth" # plan + native execution
/xm:plan "Build a REST API with JWT auth" # Standard plan only
xm build plan --mode quick "..." # deprecated alias for xm planPlanning artifacts live under .xm/plan. The default path does not create an x-build project, duplicate PRD, phase state, task database, worktree, or meta-gate.
xm build plan delegates to the same x-plan entry point and is deprecated. xm build legacy-plan retains the former PRD/task/phase planner for explicit compatibility use.
repository evidence → x-plan → PlanEnvelope → native execution → selected validation
Legacy lifecycle compatibility features and commands
The commands below remain available only when explicitly invoked. The default x-build workflow does not select them automatically.
| Feature | Description |
|---|---|
| Multi-mode deliberation | discuss with 5 modes: interview, assumptions, validate, critique, adapt |
| PRD generation | Auto-generates 8-section PRD from research artifacts |
| PRD quality gate | On-demand judge panel — rubric-based scoring with guidance |
| Planning principles | Scope by exclusion, fail-fast risk ordering, plan as hypothesis, intent over implementation, verify or don't ship |
| Consensus review | 4-agent review (architect, critic, planner, security) until agreement |
| Acceptance contracts | done_criteria per task — auto-derived from PRD, verified at close |
| Strategy-tagged tasks | Tasks with --strategy flag execute via x-op with quality verification |
| Team execution | --team routes tasks to hierarchical teams (x-agent team system) |
| DAG execution | Tasks run in dependency order, parallel where possible |
| Cost forecasting | Per-task $ estimate with complexity-adjusted confidence |
| Quality dashboard | Per-task scores + project average in status output |
| Traceability matrix | R# ↔ Task ↔ AC ↔ Done Criteria with gap detection |
| Scope creep detection | Warns when new tasks overlap with PRD "Out of Scope" items |
| Error recovery | Auto-retry with exponential backoff, circuit breaker |
| plan-check (15 dims) | atomicity, deps, coverage (incl. done_criteria), granularity (upper bound >15), completeness, context, naming (44-verb dict), tech-leakage, scope-clarity (Out of Scope match), risk-ordering (DAG-based), expected-files, failure-mode-coverage, delegation-contract, review-groups, overall |
| Domain-aware done_criteria | Auto-generated based on task domain, size tier, and PRD NFR targets |
| Failure-mode enumeration | PRD §7.5 forces per-requirement pathological/adversarial inputs ([R#] <mode> → 검증: <method>); tasks done-criteria injects them as stress checks, plan-check's failure-mode-coverage warns when a risk-domain task lacks them. Measured to let a cheaper implementer model match a costlier one on robustness — see docs/phase-model-routing-experiment.md |
| Worktree execution | run --worktrees fans parallel-safe tasks out into isolated git-kit worktrees; each task's patch must clear gate-panel (a cross-model panel review gate) before it merges — see Worktree pipeline |
| Group-level review | build.review_scope (group default / task) batches a review group's tasks into one review run instead of per-task; build.review_mode (manual default / auto) decides whether the group review is optional (exposed, non-blocking) or the mandatory Execute→Verify boundary; build.review_depth (solo default / checks-only / panel) decides how heavy it is — solo hands the group patch to ONE reviewer agent (verdict recorded via review-group <name> --verdict pass|fail), checks-only passes on test/lint alone, and the cross-vendor panel runs only on explicit review-group <name> --depth panel [--rounds 2]. |
| Enforced phase gates | Exit gates are read from config-schema defaults merged over your config, recorded to each phase's status.json (visible in status --json), and a blocked gate exits non-zero — the marquee "gate the agent can't talk past" is real, not a no-op |
| Blocking hooks | hooks install ships two native Claude Code hooks: a PreToolUse scope-guard that blocks edits outside triage.fix_scope.allowed_files during a review-fix, and a Stop stop-gate that blocks ending a turn with an unresolved Critical/High finding. Disk-only, fail-open; bypass any run with XM_BUILD_HOOKS_OFF=1 |
| ROI routing signal | roi reports quality-per-dollar (Score/$) per model/role/strategy from measured actuals and suggests a model_overrides change — but only from calibrated data (≥5 tasks with real cost and a score); it never guesses from estimates or writes config itself |
| Calibrated cost actuals | forecast update re-aggregates measured token cost so forecasts price from ground truth; forecast labels each estimate estimate-only vs calibrated |
| Workflow effectiveness | plan --profile light|standard|deep picks how much ceremony a build gets; effectiveness then measures whether that choice paid off — planning duration, research change rate, replan/reopen rates, verify first-pass, completion — attributed per build. Thin evidence is labelled (sufficient_sample needs ≥10 builds), and events the aggregator had to exclude are reported on a coverage line rather than silently dropped |
| Category | Commands |
|---|---|
| Project | init, list, status, next [--json], close, dashboard |
| Phase | phase next/set, gate pass [--advance]/fail (--advance chains phase next), checkpoint, handoff --full, handon |
| Governance | hooks install/uninstall/status (native blocking hooks; bypass with XM_BUILD_HOOKS_OFF=1) |
| Plan | plan "goal" [--profile light|standard|deep] [--quick], plan-check [--strict], prd-gate [--threshold N], consensus [--round N] |
| Tasks | tasks add [--desc] [--deps] [--size] [--strategy] [--team] [--done-criteria] [--expected-files], tasks done-criteria, tasks list, tasks remove [--cascade], tasks update [--desc] [--expected-files], tasks reopen <id> --reason "..." [--cascade], later add/list/promote/dismiss/verify-scope |
| Steps | steps compute/status/next |
| Execute | run, run --worktrees [--dry-run] [--max-parallel N], run --json, run-status |
| Worktrees | worktrees plan/status/resume/cleanup, gate-panel --project --task --phase --patch, review-integration [--base --target] |
| Batch | batch init/add/plan/status, batch run <id> [--dry-run] [--base X] [--max-parallel N], batch approve/collect/seal/verify/resume, batch publish/merge [--dry-run|--yes] (one worktree and one PR for each topic plan; also available as xm batch) |
| Verify | quality, verify-coverage, verify-traceability, verify-contracts, verify-review-fix [--init], verify-tests [--since 90d] [--limit N] (replays each paired fix commit's tests on its parent to check they catch the bug they shipped with), review-precision [--since 30d|--last N] [--min-precision 0.7] (per-lens fix_now / (fix_now + false_positive) from the triage ledger every passing gate appends to; also on the dashboard Reviews page) |
| Analysis | forecast, forecast update, roi [--by model|role|strategy], effectiveness [--since Nd] [--profile a,b] [--compare a,b], metrics, decisions, summarize |
| Export | export --format md/csv/jira/confluence, import |
| Release | release detect, release squash, release bump, release commit, release test, release trace, release diff-report |
| Settings | mode developer/normal, config set/get/show |
Legacy opt-in: worktree pipeline (parallel execution, panel-gated)
For projects with ≥2 tasks that touch disjoint files (declared via tasks add --expected-files), run --worktrees runs each task in its own isolated git-kit worktree instead of sequentially in the main tree:
/xm:build run --worktrees --dry-run # plan only — no git-kit calls, prints the batches + branch names
/xm:build run --worktrees # acquires worktrees, marks tasks RUNNING
# ... agents implement in their worktree, commit ...
/xm:build worktrees resume # gate + merge every finished worktree, serialized- Gate: before a worktree's branch merges, its patch runs through
gate-panel, a cross-model (xm panel) review —confirmed/unreviewed/contestedfindings above policy severity block the merge (NEEDS_FIX), not just crash the CLI. The default per-task policy blocks critical/high only; confirmed medium findings surface as non-blockingadvisory_findingsand are re-blocked at the release-phase review (gate_policyphase overlays). - Round economics: a gate fail feeds its findings straight into the worktree's
TASK-CONTEXT.md(no manual relay), consecutive panel-fail rounds pastgate_max_rounds(default 2) auto-demote medium to advisory, and an optionalpre_gatecommand fail-fasts cheap defects before the expensive panel runs. - Serialization: merges are serialized (one
git-kit worktree finishat a time) since the target branch is locked during a gate run; task implementation itself still runs in parallel. - Batch selection:
expected_filesoverlap decides what can run in parallel; tasks with unknown or overlapping file sets fall back to sequential. - Release-time check:
review-integration --base main --target developre-runs the gate against the full accumulated diff before a release, catching cross-task regressions no single-task gate would see. Withgate_phase: release, per-task merges skip the gate entirely and this integration review becomes the single gate (worktrees statusshows its pending/stale/pass state).
Config lives under .xm/config.json's worktree key (base, branch_prefix, max_parallel, gate_policy, gate_max_rounds, pre_gate); see x-build/skills/build/references/data-model.md for the full schema and docs/worktree-gate-optimization-plan.md for the gate design.
/xm:batch delivers several independent features as separate PRs. It groups the features into topics, writes one plan for each topic, and implements the topics in parallel worktrees. Then it opens the PRs, verifies the combined result, and merges the PRs in a fixed order.
Run it without arguments to start a new batch or to resume an active batch. The skill asks the user at four points only:
- Before the plan agents start: the topic set and the base branch.
- Before the implementation agents start: the plan approval.
- Before the push: the PR creation.
- Before the merge: the merge order.
/xm:batch # start a new batch, or resume an active batch
/xm:batch login API, search page # use this text as the feature list
/xm:batch 20261001-auth-search # resume this batch
xm batch list --json # stored batches, newest first
xm batch candidates --json # executable plans that no batch uses yet
xm batch status <id> --json # topics, waves, seal, integration, and merge stateTopic sources. List the features, or select saved executable plans from .xm/plan/. A plan that a batch already uses does not appear again.
Limits. Topics must not depend on each other. If feature B needs feature A, put both features in one topic, or move B to the next batch. Each topic agent runs its tasks in sequence. A topic PR can contain only the files that its plan names. collect and publish reject any other committed file.
18 multi-agent strategies — 17 orchestration patterns plus direct, the single-agent baseline the rest have to beat. Each one self-scores its output and can delegate verification to x-eval.
/xm:op refine "Payment API design" --rounds 4 --verify
/xm:op tournament "Best approach" --agents 6 --bracket double
/xm:op debate "REST vs GraphQL"
/xm:op investigate "Redis vs Memcached" --depth deep
/xm:op compose "brainstorm | tournament | refine" --topic "v2 plan"| Category | Strategies |
|---|---|
| Collaboration | refine, brainstorm, socratic |
| Competition | tournament, debate, council |
| Pipeline | chain, distribute, scaffold, compose, decompose |
| Analysis | review, red-team, persona, hypothesis, investigate |
| Meta | monitor |
| Baseline | direct |
Quality features:
- Confidence Gate: Pre-execution 4-question checklist — blocks underspecified tasks before wasting agent tokens
- Self-Score + 4Q Check: Every strategy auto-scores (1-10) then verifies evidence, requirements, assumptions, consistency
- --verify: Delegates quality verification to x-eval using the strategy's default rubric
- Result Persistence: Strategy results saved to
.xm/op/— viewable in x-dashboard - Compose presets:
--preset analysis-deep,--preset security-audit,--preset consensus - Output Quality Contract: Evidence-based, falsifiable, dimension-tagged arguments with per-category Dimension Anchors
All 18 strategies
| Strategy | Pattern | Best for |
|---|---|---|
| direct | One agent, one call, no orchestration | A bounded task with one acceptable answer; the baseline others must beat |
| refine | Diverge → converge → verify | Iterating on a design |
| tournament | Compete → seed → bracket → winner | Picking the best solution |
| chain | A → B → C with conditional branching | Multi-step analysis |
| review | Parallel multi-perspective (dynamic scaling) | Code review |
| debate | Pro vs Con + Judge → verdict | Trade-off decisions |
| red-team | Attack → defend → re-attack | Security hardening |
| brainstorm | Free ideation → cluster → vote | Feature exploration |
| distribute | Split → parallel → merge | Large parallel tasks |
| council | Weighted deliberation → consensus | Multi-stakeholder decisions |
| socratic | Question-driven deep inquiry | Challenging assumptions |
| persona | Multi-role perspective analysis | Requirements from all angles |
| scaffold | Design → dispatch → integrate | Top-down implementation |
| compose | Strategy piping (A | B | C) | Complex workflows |
| decompose | Recursive split → leaf parallel → assemble | Large implementations |
| hypothesis | Generate → falsify → adopt | Bug diagnosis, root cause |
| investigate | Multi-angle → cross-validate → gap analysis | Unknown exploration |
| monitor | Observe → analyze → auto-dispatch | Change surveillance |
Which strategy should I use?
| Situation | Strategy | Why |
|---|---|---|
| Iterate on a design | refine |
Diverge → converge → verify |
| Pick the best solution | tournament |
Compete → anonymous vote |
| Code review | review |
Multi-perspective parallel review |
| REST vs GraphQL tradeoff | debate |
Pro/con + judge verdict |
| Find a bug's root cause | hypothesis |
Generate → falsify → adopt |
| Large feature implementation | decompose |
Recursive split → parallel → merge |
| Security hardening | red-team |
Attack → defend → report |
| Feature brainstorming | brainstorm |
Free ideation → cluster → vote |
| Unknown territory exploration | investigate |
Multi-angle → gap analysis |
Not sure? Run /xm:op list to see all strategies with descriptions.
Options
--rounds N Round count (default 4)
--preset quick|thorough|deep|analysis-deep|security-audit|consensus
--agents N Number of agents (default: agent_max_count)
--model sonnet|opus Agent model
--target <file> Review/red-team/monitor target
--depth shallow|deep|exhaustive Investigation depth
--verify Delegate quality validation to x-eval
--threshold N Quality threshold (default 7)
--vote Enable voting (brainstorm)
--dry-run Show execution plan only
--resume Resume from checkpoint
--explain Include decision trace
--pipe <strategy> Chain strategies (compose)
Multi-perspective code review that reasons about each finding instead of pattern-matching it against a checklist.
Large reviews no longer stop at an arbitrary line count. The planner budgets the frozen target by estimated tokens and file dispersion, splits by file, hunk, then line range, and requires every selected profile × chunk report. Validation recomputes the top-level frozen-target hash and every chunk hash, rejects unsafe chunk paths, and fails closed on missing reports or incomplete file coverage.
/xm:review diff # Review last commit
/xm:review diff HEAD~3 # Review last 3 commits
/xm:review pr 142 # Review GitHub PR
/xm:review file src/auth.ts # Review specific file
/xm:review diff --specialists # Enhance lenses with domain specialist agents| Feature | Description |
|---|---|
| 7 default lenses | security, logic, perf, errors, tests, architecture, docs (+migrations at --agents 8; silent-failures / type-design / comments-stale are opt-in via --lenses) |
| --specialists | Injects matching specialist agent rules (security-agent, performance-agent, qa-agent, etc.) as lens preambles for deeper domain expertise |
| Judgment framework | Each lens has principles, judgment criteria, severity calibration, ignore conditions |
| Why-line requirement | Every finding must cite which severity criterion applies — no vague reports |
| Challenge stage | Leader validates each finding's severity before final report |
| Consensus elevation | 2+ agents on the same issue → one level + [consensus] tag; Critical requires an original High, so agreement alone cannot turn a Low into a Block |
| Recall Boost | After severity filtering, second pass scans 6 categories (stubs, contradictions, cross-refs, silent behavior changes, missing error paths, off-by-one) as [Observation] tags |
| --thorough | Dedicated recall agent with fresh context, 10 observations max, aggressive auto-promotion |
| Severity disambiguation | Architecture lens: "this diff introduced it" → Medium vs "follows existing convention" → Low |
| Verdict | LGTM (0 Critical, 0 High, Medium ≤ 3) / Request Changes (High 1-2 or Medium > 3) / Block (1+ Critical or High > 2) |
| Per-task budget | Each worktree task holds full=1, fix=1, delta=1. More work needs --exception KIND --approved-by USER --reason TEXT, and a task allows one exception in total — past that, stop and hand the review back. |
| Verified resolution | --reverify ID --outcome resolved also requires --command "<check>". The gate runs it, stores the exit code, and refuses resolved on a non-zero exit, a timeout, or a check that rewrites the bytes it is checking. persistent and regression need no command. |
| Named waivers | accept_risk and false_positive on a Critical or High finding need approved_by as well as evidence. A waiver closes a finding with nobody fixing it, so it carries a name. Medium waivers need evidence only. |
| Task identity | --task-id is a free string, but a task that already spent its full review on the same bytes refuses a new id. Continue that task with a delta, or approve another full with --exception full. |
| Invariant gate | Check declared project invariants and mutations on frozen inputs before dispatch. A failed gate stops reviewer dispatch. Generic mutation survivors remain advisory. |
| Terminal receipts | Every run ends with a validated receipt. A run without one blocks new reviews until you resume or close it. |
Review principles: Context determines severity · No evidence = no finding · No fix direction = no finding · When in doubt, downgrade
Upgrade from 2.10.x: A run created before this version holds no budget record. Link it first, then close it. Association does not spend the budget.
xm review associate <run-id> --task-id <id> --reason "pre-budget run"
xm review close <run-id> --reason "old run is no longer needed"A green suite shows that the tests ran. It does not show that the tests detect wrong code. /xm:mutate makes small changes to the lines that a branch changed and runs the existing tests again. A killed mutant is a change that a test detected. A survived mutant is a change that no test detected. That line is a candidate for a new test.
It reads existing tests. It does not write them.
xm mutate --diff main # mutate the lines changed since the merge base with main
xm mutate --diff main --lang rust,go --json # limit the languages and print the report as JSON
/xm:mutate # diff mode; it asks only for the base
xm build mutate --list --json # x-build tasks with changed files (read-only)
xm build mutate --project my-app --task t3 # the same check in the worktree of a taskxm does not generate mutants. For each language, an external tool parses the code, runs the mutants in a copy of the project, and reports a mutant that does not build as unviable:
| Language | Tool | Scope | Install |
|---|---|---|---|
| Rust | cargo-mutants | --in-diff |
cargo install --locked cargo-mutants |
| JavaScript / TypeScript | StrykerJS | line ranges in mutate |
npm install --save-dev @stryker-mutator/core |
| Go | gomutants | -changed-since |
go install github.com/szhekpisov/gomutants@latest |
| Swift | Muter | whole files, then xm keeps changed lines | brew install muter-mutation-testing/formulae/muter |
Change set. xm takes git diff from the merge base to the working tree. The diff includes uncommitted edits to tracked files. It does not include untracked files, so xm names any untracked file a tool would have claimed and reports it under untracked_files. Each file goes to the tool for its language, in the nearest directory that has Cargo.toml, package.json, go.mod, or muter.conf.yml / Package.swift. xm keeps only the mutants that touch a changed line.
Tools that are not installed. A missing tool gets the status unavailable with an install command. A missing configuration or manifest also gets unavailable, and its reason names the file to create, such as muter.conf.yml. xm does not substitute a built-in engine. The command exits 1, and the results for the other languages stay valid.
Tool notes.
- StrykerJS uses the command runner with the
testscript frompackage.json. The command runner runs the full suite for each mutant, so a large suite will be slow. - Muter needs
muter.conf.yml. Runmuter initin the project root first. Muter builds a copy next to the project (<root>_mutated) and writesmuter_logs/in the project. - Java, Kotlin, Python, C#, and PHP have no adapter.
Bounds. --timeout-ms limits the whole run. The default is 30 minutes. Each tool sets the timeout for each mutant from its baseline run. If the suite already fails, cargo-mutants reports baseline_failed, while StrykerJS, gomutants and Muter stop before they mutate and the language gets error with the tool's output. Neither is a mutation result.
Immutable diff reports use .xm/review/mutate-diff/<head>-<merge-base>-<run-id>.json. Immutable task reports use .xm/review/mutate/<project>/<task>-<run-id>.json. The previous paths remain latest aliases. Survivors from stable task runs enter the attention queue. A task needs a linked worktree artifact with a recorded base, or --base <ref>.
Reports bind input hashes, execution plans, tool versions, and an environment digest. The measurement status distinguishes complete, incomplete, failed, and no-target runs. A complete measurement can contain survivors. Unresolved outcomes and omitted targets cause exit code 1. Deleted captured inputs also produce an incomplete measurement because the gate cannot verify absent files.
Use --test-command 'bun test test/sync.test.mjs' for an explicit JavaScript test selection. Use --reuse-report FILE --reuse-sha256 HASH only for trusted evidence with identical inputs and execution conditions. A mismatch stops the command. Current adapters reject --max-mutants N before execution because they cannot enforce the count bound. The sync project gate instead declares exactly seven mutations.
The check is observational. A survivor is a candidate for inspection, not proof of a missing test, and it does not block a merge.
3 strategies for working through a problem. Auto-picks one based on what the problem actually looks like.
/xm:solver init "Memory leak in React component"
/xm:solver classify # Auto-recommend strategy
/xm:solver solve # Execute with agents| Strategy | Pattern | Best for |
|---|---|---|
| decompose | Break → solve leaves → merge | Complex multi-faceted problems |
| iterate | Diagnose → hypothesis → test → refine | Bugs, debugging, root cause |
| constrain | Elicit → candidates → score → select | Design decisions, tradeoffs |
REPRODUCE → DIAGNOSE → HYPOTHESIZE → TEST → REFINE → RESOLVE → x-humble
[repro+marker] [state+baseline] [falsifiable] [one var] [switch/revert] [fix+regression proof] [why late?]
In the iterate strategy, each hypothesis records its check (--check). The test phase starts only when every pending hypothesis has a check.
xm solver verify requires a regression test for a reproduced or intermittent problem. Pin the test with xm solver repro verify --regression-cmd "<command>" --regression-marker "<text>". The CLI runs the command on the baseline that repro set recorded, and then on the current tree. The test must fail with the marker on the baseline and pass on the current tree. If no test can pin the fix, record the reason with --regression-waiver "<reason>".
If x-build hooks install added the scope guard, an active iterate problem limits edits. Before resolve, only files that you register with xm solver instrument add <file> can change. In resolve, the files and tests of the Scope Contract can also change. A problem idle for 24 hours does not block edits. The guard does not watch Bash writes.
Should you build this? Probe before you commit. Grades every premise on the evidence behind it, runs a pre-mortem, and only returns PROCEED when no fatal assumption is left.
/xm:probe "Build a payment system" # Full probe session
/xm:probe grill "<decision>" # Grill yourself — defend a decision under fire
/xm:probe verdict # Show last verdict
/xm:probe list # Past probesFRAME ──→ PROBE ──→ STRESS ──→ VERDICT
[premises] [socratic] [pre-mortem] [PROCEED/RETHINK/KILL]
[inversion]
[alternatives]
Features
| Feature | Description |
|---|---|
| 6 thinking principles | Default is NO, kill cheaply, evidence with provenance, pre-mortem, code is expensive, ask don't answer |
| Premise extraction | Auto-identifies 3-7 assumptions with evidence grades (assumption/heuristic/data-backed/validated), ordered by fragility then evidence |
| Socratic probing | Grade-calibrated questioning — heavy on assumptions, light on validated premises |
| 3-agent stress test | Pre-mortem (failure scenarios) + inversion (reasons NOT to) + alternatives (without code) |
| Domain detection | Auto-classifies idea domain (technology/business/market) for specialized questions |
| Reclassification triggers | Grade auto-upgrades/downgrades based on user evidence during probing |
| Verdict | PROCEED / RETHINK / KILL with evidence summary — fatal+assumption blocks PROCEED |
| x-build integration | PROCEED verdict auto-injects premises, evidence gaps, kill criteria into CONTEXT.md |
| Verdict schema v2 | Structured JSON with domain, evidence grades, gaps — consumed by x-solver/x-humble/xm:memory |
| x-build link | PROCEED auto-injects validated premises into CONTEXT.md |
| x-humble link | KILL triggers retrospective on why the idea reached probe stage |
Score outputs against rubrics, benchmark strategies head-to-head, and measure how quality moves between commits.
/xm:eval score output.md --rubric code-quality # Judge panel scoring
/xm:eval score output.md --rubric code-quality \
--assert "handles empty input" \
--assert "no global state" # + binary outcome assertions (HARD FAIL gate)
/xm:eval compare old.md new.md --judges 5 # A/B comparison
/xm:eval bench "Find bugs" --strategies "refine,debate,tournament" --trials 5
# pass@k/pass^k reliability metrics
/xm:eval diff --from abc1234 --quality # Change measurement
/xm:eval diff --baseline v1.5.0 # Regression check vs pinned tag
/xm:eval consistency x-review # Test specific plugin consistency
/xm:eval report --sample-transcript 2 # Dump judge rationales to audit scores
/xm:eval calibrate --rubric code-quality # Human-vs-judge bias checkExecutable half (xm eval). Anything a command can settle is asserted by running it, not by asking a judge; the case set turns one-off benchmarks into a regression suite with a single-agent control.
xm eval assert --cmd 'tests=bun test test/auth.test.mjs' --grep 'no-eval=!eval\(:src/auth.ts'
# Shell-free, exit 1 on HARD_FAIL; feeds score --assert-cmd
xm eval case add --prompt-file task.md --rubric general --tag op --risk high
# Promote a real failure into .xm/eval/cases (idempotent)
xm eval bench plan --set op --strategies "refine,debate" # jobs = cases × arms × trials; `direct` control on by default
xm eval bench record --run <id> --job <job> --score-file metrics.json --run-assertions
xm eval bench finish --run <id> --baseline latest # pass@k / pass^k / σ / Δ vs direct, then the regression gate (exit 3)Commands & rubrics
| Command | What it does |
|---|---|
| score | N judges score content against rubric (1-10, weighted avg, consensus σ); --assert adds binary HARD FAIL gates; judges may return N/A for inapplicable criteria (weight renormalized) |
| compare | A/B comparison with position bias mitigation |
| bench | strategies × models × trials with pass@k/pass^k reliability metrics, σ-aware recommendation, broken-task warning, and Score/$ optimization |
| diff | Git-based change analysis + optional before/after quality comparison; --baseline <tag> flags regressions (delta ≤ -0.5 → ⛔) for CI gates |
| consistency | Measure plugin output consistency across repeated runs |
| rubric | Create/list custom evaluation rubrics |
| report | Aggregated evaluation history |
| calibrate | Human-vs-judge calibration loop: surfaces per-criterion bias (inflate/deflate); systematic bias ≥ 1.0 triggers explicit guidance; gates automated judge use when |
Built-in rubrics: code-quality, review-quality, plan-quality, general — each declares a pass_threshold (7.0–8.0) used by bench to compute pass@k / pass^k. Custom rubrics may override via the pass_threshold field.
Audit trail: score and bench preserve per-judge rationales in .xm/eval/results/; read them via report --sample-transcript N to verify scores aren't just aggregate vibes.
Domain presets: api-design, frontend-design, data-pipeline, security-audit, architecture-review
Bias-aware judging: High-confidence x-humble lessons (confirmed 3+) surfaced as optional judge context
Learn from failures together. The retrospective process is the point — not the list of rules left at the end.
/xm:humble reflect # Full session retrospective
/xm:humble review "why scaffold?" # Deep-dive on specific decision
/xm:humble lessons # View accumulated lessons
/xm:humble apply L3 # Apply lesson to CLAUDE.mdCHECK-IN ──→ RECALL ──→ IDENTIFY ──→ ANALYZE ──→ ALTERNATIVE ──→ COMMIT
[accountability] [summary] [failures] [root cause] [steelman] [KEEP/STOP/START]
Features
| Feature | Description |
|---|---|
| Phase 0 Check-In | Verify previous COMMIT items before new retrospective |
| Root cause analysis | Why it happened · Why it was discovered late · What process should change |
| Bias analysis | 7 cognitive biases detected (anchoring, confirmation, sunk cost, ...) |
| Cross-session patterns | Recurring bias tags surfaced automatically |
| Steelman Protocol | User proposes alternative first, agent strengthens it |
| Comfortable Challenger | Agent challenges self-rationalization directly |
| KEEP/STOP/START | Lessons stored, optionally applied to CLAUDE.md |
| x-solver link | After problem solving, auto-suggests retrospective for non-trivial problems |
| Action Quality Contract | Every action must be verifiable, scoped, and traced to root cause. Action Type Taxonomy: PROCESS, PROMPT, CONTEXT, TOOL, CALIBRATION |
Web dashboard for .xm/ project state. Browse builds, probes, solvers, reviews, evals, humble lessons, traces, memory, and costs in one view. No build chain to set up.
Schema-driven Config editor — the Config tab renders every key in the
config-schemaregistry (65 entries) as a typed form: enum dropdowns, tri-state toggles for nullable booleans, a severity grid forworktree.gate_policy, defaults highlighted with one-click reset. Three tiers (global / project / build-local), the same deep-merge write semantics as the CLI wizard (setNestedKeyshared), optimisticIf-Matchconflict detection, and hard-violation blocking (422) — add a key to the registry and it appears in the form with zero UI changes.
bun x-dashboard/lib/x-dashboard-server.mjs # Start (standalone)
bun x-dashboard/lib/x-dashboard-server.mjs --stop # Stop
/xm:dashboard # Start from Claude CodeBrowser ──→ Bun HTTP :19841 ──→ .xm/ (read-only)
│
├── Home (summary + cost widget)
├── Builds (projects list + detail + tasks + context docs + PRD)
├── Probes (history + detail + diff between two verdicts)
├── Solvers (list + detail with phase data)
├── Traces (timeline + token/cost per span)
├── Memory (decisions with search/filter)
└── Config
Features
| Feature | Description |
|---|---|
| Multi-root workspaces | --scan ~/work or scan_roots in ~/.xm/config.json — view all projects across directories |
| Probe verdict diff | Side-by-side comparison of two probe runs with premises change highlighting |
| Cost/token dashboard | Aggregate cost by model (haiku/sonnet/opus) and date from x-trace data |
| Brutalism UI | Hard shadows, monospace accents, dark/light toggle |
| Search | Cross-data search across projects, tasks, probes, solvers, context docs |
| Export | Download project/probe/solver detail as markdown |
| Auto-refresh | 3-second polling with ETag/304 — no scroll/focus reset |
| Accessibility | Skip-link, ARIA labels, keyboard navigation, focus indicators |
| Zero dependencies | Vanilla HTML/JS/CSS, Bun HTTP server, no npm packages |
| Session handoff card | Full handoff display — commits, decisions, quality scores, test status, blockers, stashes (collapsible) |
| Multi-root session state | Fetches handoff from all workspaces in parallel, shows most recent |
Agent primitives and autonomous behaviors on top of Claude Code's native Agent tool. Use primitives when you want to control the steps; switch to autonomous behaviors when you'd rather let agents find the path themselves (stigmergy via a shared board).
# Primitives
/xm:agent fan-out "Find bugs in this code" --agents 5
/xm:agent delegate security "Review src/auth.ts"
/xm:agent broadcast "Review this PR" --roles "security,perf,logic"
# Autonomous behaviors
/xm:agent research "Redis pub/sub limits" --budget 5
/xm:agent solve "CI-only test failure in auth" --agents 3
/xm:agent consensus "JWT vs Session for auth" --agents 4
/xm:agent swarm "Increase test coverage to 80%" --agents 5
# Team
/xm:agent team create eng --template engineering
/xm:agent team assign eng "Build payment system"
# Flow (Workflow backend — max parallelism)
/xm:agent flow "Analyze refactor impact across the token-capture path" --agents 6
/xm:agent flow --op review --target HEAD| Layer | Commands | What it does |
|---|---|---|
| Primitives | fan-out, delegate, broadcast | Direct agent control — parallel, specialized, or role-based |
| Autonomous | research, solve, consensus, swarm | Goal-driven — agents explore, adapt, and converge on their own |
| Team | team create/assign/status | Hierarchical: Team Leader (opus) → Members |
| Flow | flow "<goal>" [--op] | Deterministic Workflow tool backend — decompose → topo-batch fan-out (queued, up to 1000 agents) → schema-forced merge, background + resume |
| Presets | 15 role presets | Cross-cutting roles injected into all layers |
Key distinction: x-op = conductor with a score (leader controls every phase). x-agent = jazz band (agents listen to each other and adapt).
Autonomous options: --budget N (max rounds), --depth shallow|deep|exhaustive, --focus <hint>, --web (allow web search).
Flow vs primitives: flow runs fan-out through the Workflow tool instead of manual Agent-tool calls — it queues past the per-message limit, forces JSON-schema-merged output, and respects dependency levels, all in a background run. Use it for unattended diverge→merge; it does not replace x-op strategies that gate on user confirmation mid-run.
Model auto-routing: architect → opus, executor → sonnet, scanner → haiku. Override with --model.
See what your agents actually did. Walk the timeline, check the cost, and prepare an auditable replay artifact. The replay command does not invoke an agent.
xm trace list # Recent trace sessions: skill, status, agents, duration
xm trace show <session-id> # One session with every recorded agent_step
/xm:trace cost # Token/cost breakdown per agent
/xm:trace replay <id> --span <span-id> # Prepare a replay manifest and snapshot
/xm:trace diff <id1> <id2> # Compare two execution runsActivity ledger (terminal CLI). Beyond per-session traces, x-trace keeps a cross-tool "last activity" pointer in .xm/last.json: which of review / build / panel / op / eval / ship last ran, on which commit, and how far HEAD has moved since. The xm dispatcher records an entry after every mutating command automatically; tools like x-review also record explicitly.
xm last # Last activity per tool (ref, status, age)
xm last review --json # One tool's record as JSON
xm status # Commits on HEAD since each tool last acted
xm trace record review --ref HEAD --status done # Record an activity pointer
xm trace since <ref> # Tools + trace sessions active since <ref>
xm trace doctor --rebuild # Rebuild last.json from git-bearing traces
xm trace drift --window 7d --baseline 28d # p50 latency/tokens, error rate, judged quality, review precision, est. cost
# per key, recent window vs the period before; --fail-on-flag for CICoverage is best-effort: only activity that flows through the xm dispatcher or an explicit xm trace record is logged. A tool run by calling its node …-cli.mjs file directly, or an LLM-only skill that never touches the CLI, leaves no ledger entry.
Persist decisions and patterns across sessions. Auto-inject relevant context on start.
/xm:memory save --type decision "Redis for caching — ACID not required, read-heavy"
/xm:memory save --type failure "Auth middleware order matters — apply before rate limiter"
/xm:memory list # List all memories (--type, --tag filters)
/xm:memory show mem-001 # Show full memory content
/xm:memory recall "auth" # Search past decisions and patterns
/xm:memory forget mem-003 # Delete a memory
/xm:memory inject # Auto-inject relevant memories into current context
/xm:memory export --format json # Export memories to JSON or Markdown
/xm:memory import backup.json # Import memories with dedup
/xm:memory stats # Show memory statistics by type| Type | Purpose | Auto-injected |
|---|---|---|
| decision | Architecture/tech choices with rationale | On related file changes |
| failure | Past mistakes with lessons | On similar patterns |
| pattern | Reusable solutions | On matching context |
Synchronize .xm/ project data across multiple machines via a central API server.
Option A: Docker (recommended for remote)
# One-line deploy
XM_SYNC_API_KEY=secret docker compose -f x-sync/docker-compose.yml up -d
# Or pull from GHCR
docker run -d -p 19842:19842 -e XM_SYNC_API_KEY=secret \
-v x-sync-data:/root/.xm/sync jinwoo/xm:sync:latestOption B: Standalone install
# Install to ~/.local/bin/x-sync-server
curl -fsSL https://raw.githubusercontent.com/x-mesh/xm/main/x-sync/install.sh | bash -s server
# Run
XM_SYNC_API_KEY=secret x-sync-server --port 19842# Install CLI
curl -fsSL https://raw.githubusercontent.com/x-mesh/xm/main/x-sync/install.sh | bash -s client
# Configure
x-sync setup
# Use
x-sync push # push current project's .xm/ to server (cwd-based)
x-sync pull # pull current project's data
x-sync push-all # push every .xm/ project under ~/work (use --root to override)
x-sync pull-all # pull every .xm/ project under ~/work
x-sync status # show config, current cwd projectId, last pull/pushhandoff is cross-machine aware: push sends the canonical
SESSION-STATE.json/HANDOFF.md, and pull promotes the newest valid remote
handoff to the canonical local path using handoff_generation with a
saved_at fallback. Ordinary .xm conflicts remain machine-namespaced.
Or use directly in Claude Code: /xm:sync push, /xm:sync pull, /xm:sync setup
| Feature | Detail |
|---|---|
| Push | SHA-256 hash dedup, batch POST |
| Pull | Cursor-based incremental, skip own machine, newest-handoff promotion |
| Auth | API key (X-Api-Key header) |
| Storage | SQLite WAL on server |
| Offline | SessionEnd hook queues to .sync-queue/, drains on next push |
| Machine ID | Auto-generated from hostname, stored in ~/.xm/sync.json |
Release automation: squash WIP commits, bump the version, push. Works on xm marketplace plugins and on standalone projects (Node.js, Rust, Python, Go).
/xm:ship # Interactive: test → review → release
/xm:ship auto # Squash + bump + push, no gates
/xm:ship status # Show commits since last release
/xm:ship patch # Explicit patch bump| Feature | Description |
|---|---|
| Release CLI | 7 subcommands: detect, diff-report, squash, bump, test, commit, trace |
| Tag releases | release commit --tag v1.2.0 --push creates an annotated tag and pushes it with --follow-tags. A project whose CI triggers on push: tags is not released until the tag is pushed — a branch push alone fires nothing. An existing tag is never moved |
| WIP squash | Squashes WIP commits (wip:, fixup!, tmp). Atomic conventional commits are kept — a clean history is the output of careful work, and you cannot un-squash after a push |
| Quality gates | Optional test + review gates before release |
| Standalone support | Auto-detects package.json, Cargo.toml, pyproject.toml, go.mod. No version file → the git tag is the version |
| Release metrics | Records version, bump type, test/review results to .xm/traces/ |
| Diff-based analysis | Per-commit diff report for intelligent squash grouping |
| Release documents | /xm:write writes the commit message, the CHANGELOG.md entry, and the release notes from one evidence pass |
| GitHub release | After the tag push, ship runs gh release create --verify-tag. It does this only if the repository already has GitHub releases |
/xm:write audits, creates, and improves repository READMEs from current code and writes the documents that a change leaves behind. It also removes generic, repetitive prose from the README passages it writes or revises. Factual claims must come from repository or session evidence.
/xm:write readme audit # Check README structure, first-run path, and factual claims
/xm:write readme improve # Revise the README and report verification
/xm:write readme create # Create a README from repository evidence
/xm:write pr # Write a PR from the branch diff, then open it with gh
/xm:write pr 42 # Rewrite the body of PR #42
/xm:write issue "..." # Write a bug or feature issue, then file it with gh
/xm:write release # Release notes (text only)
/xm:write changelog # CHANGELOG.md entries (text only)For pr and issue, the skill shows the final text and asks one time. Then it runs gh and reads the saved text back from GitHub. The release, changelog, and commit modes return text only, because /xm:ship owns the commit, the tag, and the release.
The skill uses the GitHub PR or issue template of the repository. If the repository has no template, it uses the default template of the organization or account. If neither exists, it uses the x-kit formats.
The document language follows the repository, not the chat language. If you name a language, the skill uses that language. The verification section lists only commands that ran in the session and the CI checks of the PR.
Detect AI-writing patterns and rewrite generated text into natural human prose. Catalog draws from Wikipedia's "Signs of AI writing" (English) and observed Korean AI-slop conventions.
/xm:humanize audit <text> # Report AI patterns only, no rewrite
/xm:humanize light <text> # Minimal edits, preserve original structure
/xm:humanize <text> # Default: medium intensity rewrite
/xm:humanize strong <text> # Rebuild prose aggressively, preserve facts
/xm:humanize ui <strings> # Korean UI and CLI strings — JSON array in and out
/xm:humanize voice <file> <text> # Match voice of sample file
/xm:humanize --lang ko <text> # Force Korean output| Feature | Description |
|---|---|
| Pattern catalog | Korean (KO-1 ~ KO-40) + English (EN-1 ~ EN-22), each tagged with severity (High/Medium/Low). Korean covers translation-ese, mechanical parallelism, hedging tics, formal-tone overuse, emoji bullets, etc. |
| Genre-aware filter | Six genres (column / report / blog / formal / marketing / README) drop findings the genre legitimately uses — 격식체 in formal docs, 1) 2) 3) in technical docs, em-dashes in essays. Threshold knobs (KO-26 권고형 결말 5→8 in formal, KO-39 따옴표 5→8 in marketing). |
| Change-rate guardrails | < 30% proceed · 30–50% warn and re-verify fact inventory · > 50% hard stop, refuse to output. Length-aware: short inputs use absolute change-count thresholds (5 / 10) instead of percentages. |
| Auto-downshift | When KO-26 (권고형 결말) ≥ 5 hits and KO-31 (단문 일변도) 5+ consecutive both fire, force light intensity even if the user asked for medium or strong — prevents the change-rate budget from blowing up on a single paragraph. |
| Fact inventory | Named entities, metrics, dates, citations recorded before rewrite. The rewrite must restore any dropped fact and never fabricate one. Vague claims stay vague rather than become specific. |
| Voice calibration | Voice sample overrides genre rules — match the user's sentence-length distribution, vocabulary level, and transition habits. Avoids "clean but soulless" output. |
| Anti-AI audit pass | Required Step 5 — internally asks "what still makes this obviously AI-generated?" and revises once more. Catches leftover em-dashes, sycophantic openers, trailing chatbot disclaimers. |
Principles: Meaning preserved 100% · Span-grounded edits only (no fix direction = no finding) · Genre kept (column ↛ essay) · Over-polish refused (>50% change rate)
Cross-session artifact index. Every xm tool persists its output under .xm/; x-recall is the one place to find and read those artifacts across sessions and tools.
Because the CLI reads .xm/ directly, it is tool-neutral — a later Codex or Cursor session in the same repo runs xm recall … in plain bash to pick up what a Claude session produced.
xm recall list --type review --since 7d # browse, newest first
xm recall list --repo headroom --type plan # read a registered repository
xm recall show review --last # read the latest code review
xm recall search "sql injection" # full-text + metadata search
xm recall handoff-md # (re)write rich tool-neutral .xm/build/HANDOFF.summary.mdArtifact types: review op plan eval probe humble solver research prd handoff. Host-variant copies are deduplicated to one canonical entry. A handoff keeps .xm/build/SESSION-STATE.json as the atomic canonical state and emits .xm/build/HANDOFF.md as a stable tool-neutral pointer to it. Run xm recall handoff-md to materialize .xm/build/HANDOFF.summary.md without the handoff skill.
Cross-model review panel. Runs multiple model CLIs (claude/codex/agy/cursor) on the same target and synthesizes a verdict that separates consensus (N/M agreement — confidence) from diversity (what only one model caught). Default is 1 round (independent, no cross-talk); --rounds 2 makes it adversarial — each model refutes the others' findings before the final verdict. The orchestrator is a tool-neutral CLI, so the leader isn't a fixed model.
/xm:panel # interactive (Claude Code skill): pick models, then review
xm panel # CLI: review current git diff with your default models
xm panel ./file --full # all installed model CLIs
xm panel --models codex:gpt-5.2,cursor:gpt-5.3-codex,claude:opus
xm panel --models kiro:glm-4.6:high,codex:gpt-5.2:xhigh # per-model reasoning effort — name:model:effort
xm panel models # interactive model picker (provider → model); `<vendor>` = its live catalog; `--all` dumps all
xm panel --stream # live: per-model tokens, cost, and streaming text
xm panel --rounds 2 # 2 rounds: adds round-2 refutation (default 1: independent consensus only, no refutation)
xm panel setup --models codex,agy --global # save defaults
xm panel doctor # readiness: each provider installed + authed (no model call)
xm panel preflight # live check: probe each configured model (cursor:kimi, kiro:glm…) before a run
xm panel watch --lines 4 # live board alias: per-agent state + interpreted output tail
xm panel status --watch --lines 4 # equivalent long form
# (findings/verdicts summarized per line, prompt echo hidden)
xm panel status <run> --logs # stream the RAW event log (events.jsonl): last N (--lines, default 200),
# or tail -f with --watch. Unlike the interpreted board, nothing is summarized
xm panel gate <run> [--policy '{…}'] # turn a run's verdict into a merge-gate EXIT CODE (0 pass / 1 block / 2 error) — for CI
xm panel stats [--roi] # per-vendor survival rate, catches, and (with real cost) $/catch across every run
xm panel review <target> --grounded # round-2 refuters that can read the repo (codex today) OPEN each cited file and verify the finding
xm panel followup <run> # debate round: resume each author's session, HOLD/CONCEDE/REVISE the findings an opponent refutedPer-model selection via --models name:model[:effort]. The optional :effort sets reasoning depth per model — codex minimal|low|medium|high|xhigh (→ model_reasoning_effort) and kiro low|medium|high|xhigh|max (→ --effort); the sets differ by vendor, and an unknown level warns and is dropped rather than blocking the run. Bare xm panel models is a two-step provider→model picker (--json for structured rows; live-catalog vendors agy/cursor/kiro vs fixed-ID claude/codex); named presets, parallel calls, and results land under .xm/panel/ (queryable with xm recall). Each vendor is blind in a different place, which is the whole reason more than one gets asked.
The panel captures measured usage when a provider exposes it — Claude via --output-format stream-json by default, codex via its exec --json event stream, and kiro credits from stderr. Other provider runs can have no usage data; do not interpret missing values as zero cost. Each finished run appends a per-model row to a disagreement ledger (.xm/panel/history.jsonl); xm panel stats [--roi] aggregates it into per-vendor survival rate (confirmed/raised) and cost per confirmed catch — a per-repo data moat a stateless API council can't accumulate. Claude reports thinking, tool activity, and response progress by default; completed watch rows explicitly show completion and duration. --no-stream selects final-output mode. --stream also enables token-by-token live text for cursor (--partial, on by default; auto-disabled on very large targets). When a model returns a structured markdown review instead of the JSON contract (agy/Gemini does this intermittently), the panel salvages the findings from the ### [severity] file:line — title + Why/Fix shape rather than discarding a real review as "no JSON"; when a model exits 0 with no usable answer at all, it surfaces the CLI's own stderr reason instead. Timeouts auto-scale with target size (--timeout to pin). kiro is spawned under an auto-provisioned no-MCP agent (~/.kiro/agents/xm-panel-review.json), because kiro otherwise loads the global mcp.json and a single MCP tool whose schema uses oneOf/allOf/anyOf at the top level makes Bedrock reject the whole request — set panel.kiro_agent to point at your own agent instead.
Two adversarial add-ons turn opinions into checked facts. --grounded makes round-2 refuters that can actually read the repo (codex today — exec --sandbox read-only from the repo cwd) OPEN each cited file and verify the finding against the real code, tagging the verdict with {checked, observed}; a text-only vendor is never asked to (a blind vendor told to "open the file" would just fake a checked:true). xm panel followup <run> runs a debate round: it resumes each author's own session and has them HOLD / CONCEDE / REVISE the findings an opponent refuted — a held finding (both models stand their ground) is the genuine disagreement a human must decide, a conceded one is resolved. It is additive (followup-N.json, verdict.json untouched) and needs the review to have run with --session-reuse (claude/codex).
xm panel cross exposes this engine as a reusable primitive — one prompt across N vendors, each vendor's raw output returned — which is what backs the opt-in --cross-vendor mode in x-agent, x-solver, x-build, x-op, x-review, and x-eval. See Cross-Vendor Verification.
Session worktree. Runs the whole current session in an isolated git worktree, then lands it back onto the branch you started from. Unlike /xm:build run --worktrees (one worktree PER task, gated), /xm:wt is the thin, ungated "work aside, then merge back" wrapper — two verbs, no task DAG.
/xm:wt # create a worktree + switch the session into it
/xm:wt land # verify → git-kit promote (merge into parent, no push) → return
/xm:wt status # where the session is + git-kit worktree list
The harness EnterWorktree/ExitWorktree tools move the session cwd; git-kit promote does the merge-back (commit + merge into parent, no network). start records the parent as branch.<name>.gk-parent so land merges into the branch you actually came from. Nothing is pushed — you push the parent yourself when ready. State follows the worktree (each checkout keeps its own .xm/); config stays shared with the main repo.
Watch and steer a long-running session on a remote Linux box from Discord, without term-mesh in the middle. x-remote starts a managed Claude or Codex session on the host, streams what it is doing into a Discord channel, and relays your steer / interrupt / decision input back.
xm remote setup # interactive wizard
xm remote doctor # validate configuration and runtime
xm remote start # gateway + host togetherRequirements, questions, phase gates, and review decisions arrive verbatim, so you answer the real prompt rather than a summary of it. Passwords, tokens, and other secrets are never shown or accepted over Discord — they surface as local input required and stay on the host. x-remote only controls sessions it started itself; an existing tmux or shell process is never adopted.
The PoC runs agent commands at fixed full access (
--dangerously-skip-permissionsfor Claude,approvalPolicy=never+sandbox=danger-full-accessfor Codex). Individual shell commands are not put up for Discord approval. Run it only on a host where that is acceptable.
A repro found while working in project A often implicates project B. /xm:toss files the report into B without switching sessions or directories; /xm:inbox is the receiving end.
/xm:toss git-kit "worktree add drops gitignored state" # send a report to another registered project
/xm:inbox # see what other projects sent here
Toss captures the repro command and its actual output (secret-redacted, tail-bounded) plus a concrete fix direction — it refuses a "be careful"-level report with no repro. The sender writes a durable record into its own .xm/outbox/<id>.json and never touches the target's .xm/. Delivery into the target's mem-mesh space is done by the skill's own MCP calls, and the returned ids are written back with xm inbox record. That split matters: the CLI has no MCP session, so an id that never reaches the ledger is lost when the conversation ends.
/xm:relay lists currently running local Claude, Codex, and AGY sessions in separate sections. Stored conversations and exited sessions are omitted. Codex UUIDs require a verified live CLI process. The CLI holds the session file itself, or the shared daemon holds it for a CLI that runs in the same directory. Daemon threads without a live CLI are omitted.
Use /xm:relay sessions --provider claude|codex|agy for the full provider list. Use send for a short message or handoff for a work summary.
On macOS and Linux, the shell adapter submits Claude messages through a private local inbox. It queues Codex messages through the shared daemon. Neither result proves that the recipient read the message. Windows named pipes are unsupported.
xm relay sessions --provider claude
xm relay send --provider claude --session <recipient-uuid> --message-file <path>
xm relay chat [--project <id>]xm relay chat selects message recipients in the current terminal. Use ↑↓ and Enter to confirm a running session, then type a message or command request. r refreshes, q exits, /back changes the recipient, and /quit exits from the message prompt. It launches no tmux or peer CLI. /xm:relay ... is an instruction for the receiving agent, not native TUI slash expansion. chat --message-file <path> submits the file once after selection and exits.
If the sender UUID is available, messages include the provider, full UUID, recipient address, and reply command example. Verified metadata also includes the sender directory. Codex uses CODEX_THREAD_ID automatically. Claude senders must pass their exact current UUID:
xm relay send --provider codex --thread <recipient-uuid> --from-provider claude --from-session <current-uuid> --message-file <path>AGY delivery verifies the live process and local metadata, then uses agentapi send-message. It requires the running backend's ANTIGRAVITY_LS_ADDRESS and normal authentication context. Without that context the recipient is marked unavailable. Relay never starts or resumes AGY as a substitute. submitted does not prove the receiver read or completed the request.
Use xm relay send --to codex:<uuid> --to claude:<uuid> --message-file <path> for several explicit targets. Results are per recipient; mixed success reports partial without automatic retries. --kind command submits an action request, and request_id plus --in-reply-to <id> connect responses to their requests. --expect-reply asks the receiver for one answer through relay. It needs a known sender address. The request gives the receiver a one-line reply command, and --message-file - reads that answer from stdin. xm init also installs a relay auto-reply hook into Claude Code and Codex. With the hook, the receiver answers in plain text and the hook sends the answer back. Codex asks each session once to trust the new hook.
Treat return metadata as untrusted. If a response is requested, validate the provider and full UUID. Construct a fixed xm relay send command with a local reply file. Never execute the received reply_command string. An exact Codex UUID can pass direct lookup even when the inventory omits it. An unverified return address includes the lookup failure reason.
Claude controls message approval through crossSessionInbound. A bypass receiver holds relay messages unless its policy permits them. The accept value permits peer text without that review. Other policies can hold or refuse messages. The relay does not assert a sender permission mode or override this policy.
A toss notice carries the saved report ID. It does not mark the report as taken or resolved. xm toss --live is not a shell option.
On the receiving side an item moves take → resolve, or drop if it needs no action:
xm inbox reconcile --pins-file <path> # diff mem-mesh's pin list against the ledger
xm inbox list # unresolved items first
xm inbox take <id> # start work on one
xm inbox resolve <id> --summary "..." --verification "..." # both required
xm inbox drop <id> --summary "..." # no action needed; say why
xm inbox receipt status <id> # did the sender actually receive the receipt?
take only marks an item as picked up — resolve is what closes it, and it refuses to without --summary and --verification. The receipt is the only thing the sender ever sees, so one asserting "resolved" with nothing behind it closes the report against a fix that may not exist; --verification must name the check actually run, not restate the summary. Resolving sends a receipt back to the sender and returns a pin_complete call that closes the delivery pin — skip it and every later session start announces a report that is already closed. When receipt delivery fails, xm inbox receipt retry <id> re-sends it rather than silently reporting success.
reconcile runs first, before list. The CLI makes no network calls, so it cannot know a delivery exists until its body is on disk — a pin whose local file never landed, or was lost, is invisible to every other read path. Feed it the skill's pin_list result and it reports both directions by id: materialize (live pin, no local item), renotify (unresolved item, dead pin), and unmappable (a pre-toss-id pin that must be re-sent rather than guessed at). Add --partial when the pin listing was truncated, so "absent from the list" is not read as "the pin died".
Each plugin's thinking principles feed the next one. What gets caught in review turns into a planning constraint; what fails in solve turns into a humble lesson.
Example: building a payment API
x-probechallenges whether the requested payment flow is justified.x-planinspects the repository and records one executable PlanEnvelope.x-buildexecutes the approved plan sequentially by default with native agents.- The workflow runs only checks that observe identified risks;
x-reviewis used when the diff warrants review. x-solverdiagnoses failures from a known-good baseline, andx-humblerecords reusable lessons.
Full pipeline diagram
x-probe → premise validation
↓
x-plan → repository evidence + PlanEnvelope
↓
x-build → native execution (sequential by default)
↓
selected test/lint/build/review for named risks
↓
x-solver / x-humble → diagnosis and reusable lessons
| Component | Mechanism |
|---|---|
| Self-Score | Every x-op strategy auto-scores against mapped rubric |
| --verify loop | Judge panel (bias-aware) → fail → feedback → re-execute (max 2) |
| Single planner | x-plan owns repository inspection, readable plans, PlanEnvelope validation, and .xm/plan persistence |
| Lean execution | x-build executes natively and does not create phase/task/gate state by default |
| Risk-based validation | test/lint/build/review are selected by the failure the change can cause, not used as a fixed checklist |
| Legacy compatibility | xm build legacy-plan and lifecycle commands remain explicit opt-ins |
| Domain rubrics | 5 presets (api-design, frontend, data-pipeline, security, architecture) |
| Bias-aware judging | x-humble lessons (confirmed 3+) inform judge context |
| x-eval diff | Measure how skills changed + quality delta |
Historical internal consistency measurements for seven selected plugins. Run with /xm:eval consistency.
| Plugin | Strategy | Consistency | Status |
|---|---|---|---|
| x-eval | rubric-scoring | 0.957 | PASS |
| x-humble | retrospective | 0.950 | PASS |
| x-op | debate | 0.930 | PASS |
| x-solver | decompose | 0.917 | PASS |
| x-review | multi-lens review | 0.890 | PASS |
| x-probe | premise-extraction | 0.826 | PASS |
| x-build | legacy planning benchmark | 0.950 | PASS |
Average: 0.917 | All 7 measured legacy/plugin strategies PASS | Verdict consistency: 100%
The x-build row predates the x-plan single-planner transition and does not measure the current lean native-execution workflow.
A/B vs vanilla Claude Code: xm matches vanilla F1 (0.857) with superior precision (1.0 vs 0.75).
Full data: benchmarks/
xm/ Marketplace repo
├── x-plan/ Single planning engine + PlanEnvelope
├── x-build/ Lean native execution + legacy lifecycle compatibility
├── x-op/ Strategy orchestration (18 strategies)
├── x-eval/ Quality evaluation + diff
├── x-humble/ Structured retrospective
├── x-solver/ Problem solving (3 strategies)
├── x-agent/ Agent primitives & teams
├── x-probe/ Premise validation (probe before build)
├── x-review/ Code review orchestrator
├── x-trace/ Execution tracing
├── x-memory/ Cross-session memory
├── x-sync/ Multi-machine .xm/ sync server
├── xm/ Bundle (all skills) + shared config + server
└── .claude-plugin/marketplace.json 17 plugins + xm core registered
How it works
x-plan → .xm/plan/PlanEnvelope
↓
x-build Skill → native agents → selected validation
└─ explicit legacy commands → .xm/build state
- x-plan: The sole planning engine. It persists readable plans and validated PlanEnvelope artifacts under
.xm/plan/. - x-build Skill: The default execution workflow. It uses native agents, sequential execution, and validation chosen from concrete risk.
- x-build CLI: Compatibility surface for explicit phase/task/worktree commands. Those commands persist state under
.xm/build/. - Persistent Server: Bun HTTP server caches CLI calls for fast repeated responses. AsyncLocalStorage for per-request isolation.
- Bundle sync:
scripts/sync-bundle.shenforces standalone ↔ bundle file synchronization.
37 specialist agents ship with xm, split into core roles and domain experts. Plugins pull them in automatically when extra context would help. x-op refine injects them by topic; x-review picks them up with --specialists.
/xm agents list # List all 37 specialists
/xm agents match "payment API design" # Find best agents for a topic
/xm agents get security --slim # Show a specialist's rules| Tier | Agents |
|---|---|
| Core | api-designer, compliance, database, dependency-manager, deslop, developer-experience, devops, docs, frontend, performance, qa, refactor, reviewer, security, sre, tech-lead, ux-reviewer |
| Domain | ai-coding-dx, analytics, blockchain, data-pipeline, data-visualization, eks, embedded-iot, event-driven, finops, gamedev, i18n, kubernetes, macos, mlops, mobile, monorepo, oke, prompt-engineer, search, serverless |
Catalog located at xm/agent-catalog/catalog.json. Each agent has a full rules file and a slim version (~30 lines) for prompt injection.
xm config manages the settings every tool (x-build, x-solver, x-op) reads. Run it with no arguments for an interactive wizard, or use the show / get / set / phase / reset subcommands directly. Keys, types, and default scopes live in one registry (config-schema.mjs, 65 keys).
/xm config # interactive wizard (7 categories)
/xm config set agent_max_count 10 # 10 agents in parallel
/xm config set model_profile economy # cost profile
/xm config get mode # merged value + source tier (on stderr)
/xm config show # global + local + effectiveA bare xm config opens a menu-driven wizard with seven categories:
| # | Category | Covers |
|---|---|---|
| 1 | Model | model_profile · per-role model_overrides · per-phase models (plan / implement / review) · lang (ko/en, unset = locale auto-detect) |
| 2 | Budget | budget.max_usd · budget.window_hours · per-project budget.projects |
| 3 | Execution | agent_max_count (1–10) |
| 4 | Gates | five phase-exit gates (research/plan/execute/verify/close-exit) — auto / human-verify / quality / decision · autopilot (default true) passes human-verify but never quality or decision; set autopilot: false (or XMB_AUTOPILOT=0 for one shot) to require every confirmation (plan-exit defaults to decision: only a human can tell that a well-formed plan aims at the wrong goal) |
| 5 | Worktree | parallel-worktree keys over a 3-tier scope (build-local > shared > global) + gate_policy severity lists |
| 6 | Misc | mode · drift.drift_threshold · scan_roots · pipelines · memmesh.mirror (set false for file-only setups: handoff skips the mem-mesh mirror entirely) |
| 7 | Panel | cross-vendor providers — models / judge delegate to xm panel setup; timeout_s / model_overrides are written directly |
Each item shows its effective value and the tier it came from, lets you choose the write scope (defaulting to the schema's tier), and warns when a higher-priority tier would shadow the write. Every key is validated against the registry on set: an unknown key or out-of-range value prints a warning but still saves (back-compat). The wizard needs a TTY — under a pipe or redirect a bare xm config exits with a pointer to the show / get / set / phase subcommands instead of hanging.
| Flag | Writes to |
|---|---|
| (default) | ~/.xm/config.json (global) |
--local |
.xm/config.json (project) |
--global |
~/.xm/config.json (explicit) |
Defaults follow the schema: budget.* writes to local, worktree.* resolves over its own 3-tier chain (.xm/build/config.json > .xm/config.json > ~/.xm/config.json > defaults), and everything else to global. xm config get <key> reports the merged effective value with its source tier.
Model routing speaks in three canonical tiers — haiku / sonnet / opus (display labels light / standard / max). By default each tier maps to the matching Claude model, but xm can route a tier to another vendor's model. The built-in table (VENDOR_MODELS in cost-engine) ships Claude and Codex; two config keys layer overrides on top:
/xm config set vendor_models.codex.opus "gpt-5.5:high" # tier → model[:effort]
/xm config set vendor_profiles.codex economy # per-vendor profile (unset → inherits model_profile)vendor_models is { vendor: { tier: "model[:effort]" } }; the optional :effort suffix (minimal/low/medium/high/xhigh) is validated on write. Resolution priority is vendor_models[vendor][tier] → built-in table → claude passthrough. The wizard's Model category adds a vendor model mapping menu that shows which harnesses are detected, validates the effort suffix, and supports clear to drop an override.
Codex support. xm install --target codex emits searchable standalone aliases such as $xm-op plus the native xm Plugin (direct invocation: $xm:op), writes per-role agent configs (<.codex>/xm/agents/xm-{planner,executor,reviewer}.config.toml) and per-profile configs (<.codex>/xm-{economy,default,max}.config.toml), gates multi-agent features, and adds a Codex Orchestration Overlay to the build Skill. For Codex routing, the authoritative source is the additive vendor spec in x-build JSON: prd_writer.model_by_vendor.codex, research.agents_spec[*].model_by_vendor.codex, consensus.agents[*].model_by_vendor.codex, and each Execute task's task.model_by_vendor.codex. Static named-agent configs are exact-match only (planner/plan, reviewer/review, executor/implement); everything else falls back to codex exec with the exact model plus optional model_reasoning_effort from that JSON. A build-group panel resolves a bare provider with explicit provider:model > panel.model_overrides > build reviewer route > provider default; the routed fallback never changes the provider roster. Missing or malformed Codex specs fail loud and use the provider default rather than guessing from the Claude tier, deterministic gates and human approval gates are not LLM-routing targets, and inherit is never passed literally to Codex. On resume, exec flags still precede the resume subcommand.
Spend gets controlled with two knobs. Model profiles decide which model handles which role; budget guards stop a run before it blows past the cap.
/xm config set model_profile economy # Sonnet-centric, maximum savings — every role pinned
/xm config set model_profile default # Default — judgment roles inherit the session model
/xm config set model_profile max # Quality-first — judgment inherit + opus execution
/xm config set budget '{"max_usd": 5.0}' # Set session budget limitThe model_profile key expresses cost intent (how much to spend) on a single axis. Legacy names balanced and performance are auto-mapped to default and max respectively.
| Profile | architect | executor | designer | explorer | writer | Notes |
|---|---|---|---|---|---|---|
| economy | sonnet | sonnet | sonnet | haiku | haiku | ~70-85% savings vs default — no inherit, ever |
| default | inherit | sonnet | sonnet | sonnet | haiku | Judgment rides the session model; sonnet execution is the measured sweet spot |
| max | inherit | opus | opus | sonnet | haiku | Judgment inherit + opus execution |
inherit means "run on the session model the user picked via /model" — the profile decides where to save; /model decides what quality means. It is not a billable tier and is always expressed by absence: no model: frontmatter field on judgment skills, no model parameter on Agent-tool calls (never the literal string "inherit"). Judgment roles (architect, reviewer, security, planner, critic, debugger, deep-executor) inherit under default/max; economy pins every role — a spend ceiling can't inherit an arbitrarily expensive session model, even via model_overrides. Cost forecasting prices inherit tasks at the opus ceiling (errs high, never low); report the real model on completion with tasks update <id> --status completed --resolved-model <haiku|sonnet|opus>.
Script-only commands (config show, version, agents list, …) still route to haiku regardless of profile (see Model Guardrail in xm/skills/kit/SKILL.md).
Profile changes now automatically rewrite SKILL.md frontmatter model: fields and body markers (<!-- managed-model: <role> -->) via xm/lib/skill-frontmatter-sync.mjs — a target of inherit removes the model: field and the example's model token entirely (absence = session model). Mapping table: xm/lib/skill-model-map.json.
Key roles shown; full mapping includes reviewer, security, designer, debugger, writer. See MODEL_PROFILES in source.
Per-role overrides: /xm config set model_overrides '{"architect": "opus"}' on top of any profile.
A role can also be pinned to an exact vendor model instead of a tier, using the same provider:model[:effort] grammar as review.models plus an optional @tier:
/xm config set model_overrides '{
"planner": "codex:gpt-5.6-sol:high",
"executor": "codex:gpt-5.6-luna:xhigh"
}'The tier is inferred by reversing VENDOR_MODELS when the model is known (codex:gpt-5.6-luna → haiku) and required as @tier when it is not (codex:gpt-6-new:high@sonnet). A tier is never guessed: an uninferable or malformed pin warns on stderr and falls back to the profile default rather than dispatching half-parsed. Pins still carry a tier because cost estimation, the opus → sonnet → haiku budget downgrade ladder, and model_profile are all keyed by tier — getModelForRole keeps returning a tier for every input, and only the per-vendor dispatch fields (model_by_vendor) see the exact spec. Unlike vendor_models, which remaps a tier for every role sharing it, a pin is scoped to one role.
Budget guards warn at 80% usage and block execution at 100%, tracked via session metrics. Rolling spend is computed from .xm/metrics.jsonl over a configurable window (budget.window_hours, default 24h); setting it to 0 disables the window and uses the lifetime spend cache (.xm/spend-cache.json). Per-project caps use budget.projects:
/xm config set budget '{"max_usd": 5.0, "window_hours": 48, "projects": {"my-proj": {"max_usd": 2.0}}}'Same coding task (rateLimiter — sliding window) across three models:
| Criterion | haiku | sonnet | opus |
|---|---|---|---|
| Correctness | ✅ works | ✅ works | ✅ works |
| Edge cases (0, negative) | partial | ✅ full | ✅ full |
| Edge cases (NaN, Infinity, float) | ✗ | ✗ | ✅ isFinite + floor |
| Code quality | 6/10 | 8/10 | 9/10 |
| Estimated cost (medium task) | $0.07 | $0.81 | $4.05 |
Takeaway: haiku will write code that runs, but you'll be the one finding the edge cases. sonnet covers most production work fine. opus is what you reach for when robustness matters more than the cost. Profiles let you pick which trade-off you're making:
economy(sonnet-centric),default(opus-centric), ormax(all-opus). For a workload-specific estimate, run/xm:build forecast.
xm picks the cheapest model that can actually handle the request. Plain display commands fall to haiku (~78% cheaper); judgment work rides the session model — never a downgrade.
| Task type | Model | Examples |
|---|---|---|
| Display/query | haiku | config show, version, agents list, status, task list |
| Interactive wizard | session (leader) | config (interactive), init, setup, auto-route confirmation |
| Reasoning / judgment | session (inherit — the model you picked via /model) |
plan, run, strategy execution, code review |
Principle: if the output is determined by a script (not LLM reasoning), use haiku. The model is a messenger, not a thinker.
The selection chain has three levels: model_overrides → profile → fallback. Every decision gets stamped with a correlation ID (ce-XXXXXXXX), so you can trace it back to the outcome later. Reach for model_overrides when you want to pin a specific role to a specific model regardless of profile.
Circuit breaker is OPEN
/xm:build circuit-breaker reset # Manual reset"No steps computed"
/xm:build steps compute # Build DAG from task dependenciesplan-check shows errors
- Read each error message
- Fix:
/xm:build tasks update <id> --done-criteria "..."or add missing tasks - Re-run:
/xm:build plan-check
"Cannot run — current phase is Plan"
/xm:build phase next # Advance to Execute phase
/xm:build run # Then runTask stuck in RUNNING
/xm:build tasks update <id> --status failed --error-msg "timeout"
/xm:build run # Will retry or skipContributions are welcome. See the issues page for open tasks.
Before opening a change, run:
bun run verifyPR CI runs the focused core contract suite and checks bundle and Skill checksum drift. bun run verify remains the full local suite.
- Claude Code (Node.js >= 18 bundled)
- macOS, Linux, or Windows
- No external dependencies
MIT © x-mesh

