diff --git a/.dockerignore b/.dockerignore new file mode 100644 index 00000000..6d855d1c --- /dev/null +++ b/.dockerignore @@ -0,0 +1,20 @@ +# Keep the image build context lean. Runtime corpus is mounted, not copied. +.git +target +example +**/node_modules +dashboard/dist/node_modules +dashboard/node_modules +**/.rgctl +**/.rgctl-diff +docs +openspec +rgctl-tests +fuzz +.cursor +.claude +.idea +.vscode +*.log +*.profraw +*.profdata diff --git a/.gitignore b/.gitignore index f6a8fc4c..9cd55cac 100644 --- a/.gitignore +++ b/.gitignore @@ -3,6 +3,7 @@ .claude/ .opencode/ .agent +.reports # OpenSpec change proposals (local only — do not commit) openspec/ .scratch/ diff --git a/AGENTS.md b/AGENTS.md index 562f0485..1593b766 100644 --- a/AGENTS.md +++ b/AGENTS.md @@ -6,7 +6,7 @@ **Your goal when contributing here:** preserve ingest scale, query correctness, memory discipline, and deterministic artifacts under `.rgctl/` — not add convenience at the cost of Tokio blocking, whole-repo clones, or ungated cold regressions. -> **Looking for how to *use* rgctl on another codebase?** Install skills (`rgctl install --skill --with-commands`) or copy [docs/agents/USER_AGENTS_TEMPLATE.md](docs/agents/USER_AGENTS_TEMPLATE.md) into *that* repo’s `AGENTS.md`. See [docs/guides/agent-commands.md](docs/guides/agent-commands.md). +> **Looking for how to *use* rgctl on another codebase?** Install skills (`rgctl install --skill`) or copy [docs/agents/USER_AGENTS_TEMPLATE.md](docs/agents/USER_AGENTS_TEMPLATE.md) into *that* repo’s `AGENTS.md`. See [docs/guides/agent-skill.md](docs/guides/agent-skill.md). --- @@ -16,12 +16,14 @@ - **Parallel ingest:** Per-file plugin extraction runs on the discover worker pool. Do not replace with a serial whole-repo walk when parallel ingest exists. - **Streaming commits:** Emit symbols/relations file-by-file; avoid unbounded `Vec` / whole-repo ASTs before commit (`rgctl-extraction` spill patterns). - **Clone hygiene:** Prefer `&[u8]` / `Cow` / borrows in tree-sitter walkers; `Vec::with_capacity` when sizes are known; no `unwrap()` in library paths. +- **Ingest hot path:** Follow **Ingest hot-path practices** below (no per-symbol heap strings, hash/prep once on workers, spill scratch reuse, tracker mapping without re-scan). - **Typed graph:** Respect `EdgeType` / node kinds; do not invent ad-hoc string edges for hot paths. - **Artifacts:** Session data lives in `{repo}/.rgctl/`. Warm caches invalidate wall-time claims. +- **Constrained discover (opt-in):** Prefer `rgctl discover . --with-limits max-mem-mb=4096,threads=1` (or `RGCTL_WITH_LIMITS`) in containers / cgroups — do **not** change default desktop discover for memory. Soft RSS tripwire at ~95%; smaller spill sort-runs and stream channel when `max-mem-mb` is set. Container smoke: `./scripts/run-container-with-limits-smoke.sh` (`tests/Containerfile`, mounts `example/linux`). - **Features:** Default semantic embedder is compiled **vocab**. Do not require ONNX / Python ML unless behind an explicit feature (e.g. `semantic-onnx` / code-daemon + Git LFS). - **OpenSpec language work:** Still cite [openspec/changes/_shared/starting-context.md](openspec/changes/_shared/starting-context.md) (pointer here); follow the sections below. - **Grammar bumps:** When you bump a tree-sitter grammar pin, update that language’s `*-ast-coverage.json` (and add the language to `rgctl-ast-coverage::bundled_specs` for new languages). Unit tests hard-fail the same drift; `cargo check -p rgctl-languages` warns (`RGCTL_AST_COVERAGE_STRICT=1` fails). The website `/docs/languages/` pages are generated from those JSON files — do not maintain parallel tables under `docs/languages/`. - +- **Releases:** Follow **Releases** below (and [docs/releasing.md](docs/releasing.md)). Do not hand-edit dozens of crate `version =` lines. --- ## Context & architecture @@ -44,6 +46,25 @@ Applies to all extraction / language / discover hot-path work (and OpenSpec `*-e 3. **Streaming** — incremental graph commit; match extraction spill/channel patterns. 4. **Idiomatic Rust** — `Result` + `thiserror`; follow `rgctl-lang-java` / `rgctl-extraction` conventions. +### Ingest hot-path practices + +Rules distilled from linux cold-discover work (`index_extract` / pass-1 / spill / `save_tracker`). Breaking these usually shows up as Gate A wall or RSS regressions — treat O(files)×O(symbols) heap work as a bug. + +| Practice | Do | Don't | +|----------|----|-------| +| **No heap strings per symbol** | Pass `&str` / slices into pass-1 (`add_symbol_with_prep`); borrow file bytes | `String::from` / `to_string()` for every symbol body or path key on the merge thread | +| **Prep on workers** | Compute line offsets, BLAKE3 `code_hash`, token bloom in `SymbolPass1Prep` on extract workers | Re-walk source / re-hash on the sequential pass-1 thread | +| **Hash once** | Set `FileExtraction.file_hash` from bytes already in memory; thread through `StreamStats` / `PipelineStats` into `FileTracker::index_files_with_mapping` | Re-`fs::read` + BLAKE3 all files in `save_tracker` after extract already hashed them | +| **Empty-tracker short-circuit** | When `file_hashes.json` is empty, `detect_changes` marks all paths **added** without hashing (keeps ChangeSet non-empty so a stale snapshot is not reused) | Hash the whole tree twice on cold discover (detect + index) | +| **Normalize / map once per file** | `begin_file_batch` → push into `active_tracker_ids`; flush once in `end_file_batch`; accumulate mapping at commit | Per-symbol `HashMap::get_mut` / `normalize_path_str(...).into_owned()`; full mmap node scan/sort just to rebuild file→node ids | +| **Reuse tree-sitter parsers** | Call `rgctl_plugin_helpers::parse_source` (thread-local `Parser` per Rayon worker); never `Parser::new()` per file on the extract hot path | Fresh `Parser::new` + `set_language` inside `extract_*` / `parse` for every file (linux: ~71k×) | +| **Spill alloc reuse** | `SegmentedSpill` scratch `Vec` + `bincode::serialize_into`; keep sort runs at `DEFAULT_SORT_RUN_BYTES` (256 MiB) unless profiling says otherwise | Fresh `bincode::serialize` → new `Vec` per node/edge; shrinking sort runs without a cold gate | +| **CodeIndex bodies off by default** | Default discover: no body-storing `CodeIndex` (no multi-GB `code_index.json`); nodes still get `code_hash` from prep | Attach a full CodeIndex on the cold path “for convenience” | + +When adding extract or graph-commit code, ask: *does this allocate or re-read once per symbol/file on the sequential merge thread?* If yes, move it to workers or reuse an existing buffer/key. + +Stage meanings and current linux notes: [docs/internal/profile.md](docs/internal/profile.md). + ### Cold profile (mandatory for scale / perf claims) 1. **Release binary only:** `cargo build --release --bin rgctl` @@ -138,10 +159,27 @@ Baselines and notes: [docs/internal/profile.md](docs/internal/profile.md#snapsho | [CONTRIBUTING.md](CONTRIBUTING.md) | Setup, tests, PR norms | | [docs/contributor-checklist.md](docs/contributor-checklist.md) | Language / feature checklist | | [docs/guides/semantic-search.md](docs/guides/semantic-search.md) | Embedders (if touching semantic) | +| [docs/releasing.md](docs/releasing.md) | Version bump + GitHub Release tags | | [openspec/changes/_shared/starting-context.md](openspec/changes/_shared/starting-context.md) | OpenSpec pointer (canonical policy is this file) | --- +## Releases + +When asked to cut or bump a release, use the lockstep tooling — full detail: [docs/releasing.md](docs/releasing.md). + +| Rule | Detail | +|------|--------| +| **One version** | SSOT is `[workspace.package] version` in root `Cargo.toml`. Crates use `version.workspace = true`. Do **not** sed/`version =` across every crate by hand. | +| **Bump TOMLs only** | `./scripts/bump-version.sh patch` (or `minor` / `major` / `X.Y.Z`). Syncs workspace version, `[workspace.dependencies]` path pins, and README release links. | +| **Bump + tag + push** | `cargo release patch --workspace` (dry-run), then `--execute` when the user wants commit/tag/push. Config: [`release.toml`](release.toml) (`shared-version`, `publish = false`, tag `v{{version}}`). Needs a **clean** git tree. | +| **Tools** | `cargo install cargo-edit cargo-release --locked` if missing. | +| **GitHub Release** | Pushing `v*` runs [`.github/workflows/release.yml`](.github/workflows/release.yml) (binaries). Add `docs/releases/vX.Y.Z.md` for curated notes. | +| **No crates.io** | `publish = false` — do not `cargo publish` unless the user explicitly asks to enable it. | +| **Commits / tags / push** | Only when the user explicitly requests them (same standing rule as other git ops). Prefer preparing the bump + release notes and stopping for the user to commit/sign/tag if they GPG-sign locally. | + +--- + ## Build and day-to-day commands ```bash @@ -159,7 +197,7 @@ cargo build --release Code-daemon / ONNX weights: `git lfs pull` when using that embedder feature. -Dogfood fixtures: `rgctl-tests/` (e.g. ecommerce-*). Consumer agent pack: `rgctl install --skill --with-commands --tools cursor`. +Dogfood fixtures: `rgctl-tests/` (e.g. ecommerce-*). Consumer agent pack: `rgctl install --skill --tools cursor`. --- diff --git a/CONTRIBUTING.md b/CONTRIBUTING.md index f2293ef0..89d728dd 100644 --- a/CONTRIBUTING.md +++ b/CONTRIBUTING.md @@ -89,6 +89,13 @@ Full map: [docs/Code_structure.md](docs/Code_structure.md) --- +## Releasing (version bump) + +Lockstep workspace version via `[workspace.package]` + `version.workspace = true`. +See **[docs/releasing.md](docs/releasing.md)** for `./scripts/bump-version.sh` and `cargo release`. + +--- + ## Adding or improving a language / feature Use the hub checklist for path choice, test matrices, and pre-PR commands: @@ -103,7 +110,7 @@ Tier 1 depth (Layers A–F): [docs/tier-1-language-support.md](docs/tier-1-langu - **User-facing:** `docs/Introduction.md`, `docs/user-guide.md`, `docs/dashboard-user-guide.md` - **Agents (contribute to rgctl):** root [`AGENTS.md`](AGENTS.md) (starting-context, profiles/tests/benches) · [`docs/json-api.md`](docs/json-api.md) -- **Agents (use rgctl elsewhere):** [`docs/guides/agent-commands.md`](docs/guides/agent-commands.md) · [`docs/agents/USER_AGENTS_TEMPLATE.md`](docs/agents/USER_AGENTS_TEMPLATE.md) · [`docs/agent-recipes.md`](docs/agent-recipes.md) +- **Agents (use rgctl elsewhere):** [`docs/guides/agent-skill.md`](docs/guides/agent-skill.md) · [`docs/agents/USER_AGENTS_TEMPLATE.md`](docs/agents/USER_AGENTS_TEMPLATE.md) · [`docs/agent-recipes.md`](docs/agent-recipes.md) - **Accuracy:** keep CLI examples aligned with `dashboard/scripts/validate-guide-cli-gbuilder.sh` where possible --- diff --git a/Cargo.toml b/Cargo.toml index 249ec690..c5a1fbbd 100644 --- a/Cargo.toml +++ b/Cargo.toml @@ -48,7 +48,8 @@ resolver = "2" [workspace.package] edition = "2024" -rust-version = "1.88" +rust-version = "1.99" +version = "0.4.17" [workspace.dependencies] rgctl-plugin-api = { path = "crates/rgctl-plugin-api", version = "0.4.17" } @@ -95,6 +96,12 @@ rgctl-lang-groovy = { path = "crates/rgctl-lang-groovy", version = "0.4.17" } rgctl-ast-coverage = { path = "crates/rgctl-ast-coverage", version = "0.4.17" } rgctl-languages = { path = "crates/rgctl-languages", version = "0.4.17" } tree-sitter = "0.25" +# OSV schema parse (no client); OpenVEX types for later VEX emission +osv = { version = "0.3", default-features = false, features = ["schema"] } +openvex = "0.1.1" +walkdir = "2" +zip = { version = "2", default-features = false, features = ["deflate"] } +thiserror = "1" [workspace.lints.rust] unsafe_code = "warn" @@ -124,7 +131,7 @@ expect_used = "warn" [package] name = "rgctl" -version = "0.4.17" +version.workspace = true edition.workspace = true rust-version.workspace = true authors = ["rgctl Contributors"] diff --git a/README.md b/README.md index 800874a1..df09b6c4 100644 --- a/README.md +++ b/README.md @@ -7,10 +7,10 @@ [![Docs](https://img.shields.io/badge/docs-shaaf.dev%2Frgctl-2563eb?style=flat-square&logo=readthedocs&logoColor=white)](https://shaaf.dev/rgctl) [![Website](https://img.shields.io/github/actions/workflow/status/sshaaf/rgctl/website.yml?branch=main&style=flat-square&label=website)](https://shaaf.dev/rgctl) -[![Rust](https://img.shields.io/badge/rust-1.88%2B-orange?style=flat-square&logo=rust)](https://www.rust-lang.org/) +[![Rust](https://img.shields.io/badge/rust-1.99%2B-orange?style=flat-square&logo=rust)](https://www.rust-lang.org/) [![Platforms](https://img.shields.io/badge/platform-macOS%20%7C%20Linux%20%7C%20Windows-555?style=flat-square)](https://github.com/sshaaf/rgctl/releases/latest) [![tree-sitter](https://img.shields.io/badge/parser-tree--sitter-brightgreen?style=flat-square)](https://tree-sitter.github.io/tree-sitter/) -[![Agents](https://img.shields.io/badge/agents-Cursor%20%7C%20Claude%20%7C%20Codex-111827?style=flat-square)](https://shaaf.dev/rgctl/docs/guides/agent-commands/) +[![Agents](https://img.shields.io/badge/agents-Cursor%20%7C%20Claude%20%7C%20Codex-111827?style=flat-square)](https://shaaf.dev/rgctl/docs/guides/agent-skill/) [![Tier 1](https://img.shields.io/badge/languages-14%20Tier%201-8b5cf6?style=flat-square)](https://shaaf.dev/rgctl/docs/languages/) [![C](https://img.shields.io/badge/C-A8B9CC?style=flat-square&logo=c&logoColor=black)](docs/languages/README.md) @@ -39,7 +39,7 @@ rgctl -f json blast-radius MyService rgctl -f json gql 'MATCH (a:Function)-[:CALLS]->(b) RETURN a,b LIMIT 20' # Use with your favorite LLM agent -rgctl install --skill --with-commands --tools cursor,claude,codex,agents +rgctl install --skill --tools cursor,claude,codex,agents ``` https://github.com/user-attachments/assets/15ec6d91-f716-4cbd-a873-e982ba3c6dca @@ -57,7 +57,7 @@ https://github.com/user-attachments/assets/15ec6d91-f716-4cbd-a873-e982ba3c6dca rgctl --version ``` -**Or build from source** (Rust **1.88+**): +**Or build from source** (Rust **1.99+**): ```bash git clone https://github.com/sshaaf/rgctl.git @@ -97,13 +97,13 @@ Always prefer **`-f json`** for agents and scripts ([JSON API](docs/json-api.md) ## Use with coding agents -Install the bundled pack (skills + slash commands) into your IDE tooling: +Install the bundled pack (skills) into your IDE tooling: ```bash -rgctl install --skill --with-commands --tools cursor,claude,codex,agents +rgctl install --skill --tools cursor,claude,codex,agents ``` -Then: **discover once → query with `-f json`**. See [Agent commands](docs/guides/agent-commands.md). +Then: **discover once → query with `-f json`**. See [Agent pack](docs/guides/agent-skill.md). For *your* application repo, optionally paste [USER_AGENTS_TEMPLATE.md](docs/agents/USER_AGENTS_TEMPLATE.md) as `AGENTS.md`. --- diff --git a/agent-pack/agents/registry.toml b/agent-pack/agents/registry.toml index 82583200..96cdbcf1 100644 --- a/agent-pack/agents/registry.toml +++ b/agent-pack/agents/registry.toml @@ -1,4 +1,4 @@ -# Agent adapters (OpenSpec-aligned paths). See docs/guides/agent-commands.md +# Agent adapters (OpenSpec-aligned paths). See docs/guides/agent-skill.md # command_style: hyphen | colon | dollar | slash # command_extension: md (default), prompt, prompt.md, toml # global_skills / global_commands: reserved for future -g path wiring (install uses agent_dir under prefix today) diff --git a/agent-pack/manifest.yaml b/agent-pack/manifest.yaml index c3b34271..0880f504 100644 --- a/agent-pack/manifest.yaml +++ b/agent-pack/manifest.yaml @@ -9,14 +9,14 @@ workflows: title: Data flow and slices - id: search title: Semantic and structural search - - id: gql - title: Graph query language - id: migrate title: Migration roadmap - id: kantra title: Konveyor Kantra rules - id: gate title: CI and policy gates + - id: vuln + title: OSV triage and deps check meta_skills: - id: rgctl - title: rgctl router + title: rgctl diff --git a/agent-pack/out/agents/host-agents/skills/rgctl-discover/SKILL.md b/agent-pack/out/agents/host-agents/skills/rgctl-discover/SKILL.md deleted file mode 100644 index 62094801..00000000 --- a/agent-pack/out/agents/host-agents/skills/rgctl-discover/SKILL.md +++ /dev/null @@ -1,38 +0,0 @@ ---- -name: rgctl-discover -description: "Index and discover. Use for rgctl discover workflow. Spawn rgctl -f json; parse schema_version from stdout." -rgctl-managed: true -metadata: - generatedBy: "rgctl 0.4.13" ---- - -# Discover workflow - -**When:** First use, rebuild after large changes, or incremental `--files` update. - -| Intent | Command | -|--------|---------| -| Index repo | `cd "$REPO" && rgctl discover .` or `rgctl -r "$REPO" discover` | -| Full pipeline | `discover . --full` | -| Incremental | `discover --files path1,path2` (requires existing `.rgctl/`) | - -**Fast path:** If `.rgctl/` exists and the user did not ask to rebuild, do **not** re-run discover. - -Common flags: `--with-cfg`, `--with-kantra`, `--export-migration-hints` (migration plan is the **migrate** workflow, not discover alone). - - -## Agent loop - -1. Parse the user question (natural language). -2. Run `rgctl -f json …` (or `rgctl serve` + HTTP for repeated queries). -3. Parse `schema_version` and payload from **stdout** only. -4. Summarize facts; do not dump raw JSON. -5. Re-query if the graph may be stale after edits. - -**Never** redirect stderr to `/dev/null`. If `.rgctl/` exists and the question is structural, use rgctl before ripgrep or bulk file reads. - -```bash -export REPO=/path/to/repo -rgctl -r "$REPO" -f json -``` - diff --git a/agent-pack/out/agents/host-agents/skills/rgctl-flow/SKILL.md b/agent-pack/out/agents/host-agents/skills/rgctl-flow/SKILL.md deleted file mode 100644 index d620f4bd..00000000 --- a/agent-pack/out/agents/host-agents/skills/rgctl-flow/SKILL.md +++ /dev/null @@ -1,35 +0,0 @@ ---- -name: rgctl-flow -description: "Data flow and slices. Use for rgctl flow workflow. Spawn rgctl -f json; parse schema_version from stdout." -rgctl-managed: true -metadata: - generatedBy: "rgctl 0.4.13" ---- - -# Flow workflow - -**When:** Slices, PDG, taint, CPG data flows. Requires `discover --with-cfg`. - -```bash -rgctl -r "$REPO" -f json slice FILE --line N --variable V [--function F] [--direction backward|forward] -rgctl -r "$REPO" -f json cpg flows FILE --line N --variable V --function F -``` - -Check readiness: `rgctl -f json cpg status`. - - -## Agent loop - -1. Parse the user question (natural language). -2. Run `rgctl -f json …` (or `rgctl serve` + HTTP for repeated queries). -3. Parse `schema_version` and payload from **stdout** only. -4. Summarize facts; do not dump raw JSON. -5. Re-query if the graph may be stale after edits. - -**Never** redirect stderr to `/dev/null`. If `.rgctl/` exists and the question is structural, use rgctl before ripgrep or bulk file reads. - -```bash -export REPO=/path/to/repo -rgctl -r "$REPO" -f json -``` - diff --git a/agent-pack/out/agents/host-agents/skills/rgctl-gate/SKILL.md b/agent-pack/out/agents/host-agents/skills/rgctl-gate/SKILL.md deleted file mode 100644 index 07ab7e10..00000000 --- a/agent-pack/out/agents/host-agents/skills/rgctl-gate/SKILL.md +++ /dev/null @@ -1,36 +0,0 @@ ---- -name: rgctl-gate -description: "CI and policy gates. Use for rgctl gate workflow. Spawn rgctl -f json; parse schema_version from stdout." -rgctl-managed: true -metadata: - generatedBy: "rgctl 0.4.13" ---- - -# Gate workflow - -**When:** Policy checks and temporal PR gates. - -```bash -rgctl -r "$REPO" -f json check --policy-file policy.json -rgctl -r "$REPO" -f json pr-check --policy-file rgctl-pr-policy.json --base-ref origin/main --head-ref HEAD --strict -rgctl -r "$REPO" -f json check --temporal --policy-file policy.json --base-ref origin/main --head-ref HEAD -``` - -Exit code 1 means violations. Parse JSON for violation details. - - -## Agent loop - -1. Parse the user question (natural language). -2. Run `rgctl -f json …` (or `rgctl serve` + HTTP for repeated queries). -3. Parse `schema_version` and payload from **stdout** only. -4. Summarize facts; do not dump raw JSON. -5. Re-query if the graph may be stale after edits. - -**Never** redirect stderr to `/dev/null`. If `.rgctl/` exists and the question is structural, use rgctl before ripgrep or bulk file reads. - -```bash -export REPO=/path/to/repo -rgctl -r "$REPO" -f json -``` - diff --git a/agent-pack/out/agents/host-agents/skills/rgctl-gql/SKILL.md b/agent-pack/out/agents/host-agents/skills/rgctl-gql/SKILL.md deleted file mode 100644 index da90d1d6..00000000 --- a/agent-pack/out/agents/host-agents/skills/rgctl-gql/SKILL.md +++ /dev/null @@ -1,35 +0,0 @@ ---- -name: rgctl-gql -description: "Graph query language. Use for rgctl gql workflow. Spawn rgctl -f json; parse schema_version from stdout." -rgctl-managed: true -metadata: - generatedBy: "rgctl 0.4.13" ---- - -# GQL workflow - -**When:** Ad-hoc graph queries, inventories, call neighborhoods. - -```bash -rgctl -r "$REPO" -f json gql 'MATCH (n:Function) WHERE n.name LIKE "*Service*" RETURN n LIMIT 20' -rgctl -r "$REPO" -f json gql --macro-name all_functions unused -``` - -Use **qualified_name** / FQN for classes, not bare `n.name` when disambiguating. Always use **LIMIT** on broad patterns. Explain macros before inventing raw GQL. - - -## Agent loop - -1. Parse the user question (natural language). -2. Run `rgctl -f json …` (or `rgctl serve` + HTTP for repeated queries). -3. Parse `schema_version` and payload from **stdout** only. -4. Summarize facts; do not dump raw JSON. -5. Re-query if the graph may be stale after edits. - -**Never** redirect stderr to `/dev/null`. If `.rgctl/` exists and the question is structural, use rgctl before ripgrep or bulk file reads. - -```bash -export REPO=/path/to/repo -rgctl -r "$REPO" -f json -``` - diff --git a/agent-pack/out/agents/host-agents/skills/rgctl-impact/SKILL.md b/agent-pack/out/agents/host-agents/skills/rgctl-impact/SKILL.md deleted file mode 100644 index 3fa5b702..00000000 --- a/agent-pack/out/agents/host-agents/skills/rgctl-impact/SKILL.md +++ /dev/null @@ -1,34 +0,0 @@ ---- -name: rgctl-impact -description: "Blast radius and impact. Use for rgctl impact workflow. Spawn rgctl -f json; parse schema_version from stdout." -rgctl-managed: true -metadata: - generatedBy: "rgctl 0.4.13" ---- - -# Impact workflow - -**When:** Before refactors, renames, or API changes. - -```bash -rgctl -r "$REPO" -f json blast-radius SYMBOL [--depth N] [--class NAME] [--file PATH] -``` - -Disambiguate symbols with `--class` or `--file` when names collide. Report hop depth and top callers/callees from JSON payload. - - -## Agent loop - -1. Parse the user question (natural language). -2. Run `rgctl -f json …` (or `rgctl serve` + HTTP for repeated queries). -3. Parse `schema_version` and payload from **stdout** only. -4. Summarize facts; do not dump raw JSON. -5. Re-query if the graph may be stale after edits. - -**Never** redirect stderr to `/dev/null`. If `.rgctl/` exists and the question is structural, use rgctl before ripgrep or bulk file reads. - -```bash -export REPO=/path/to/repo -rgctl -r "$REPO" -f json -``` - diff --git a/agent-pack/out/agents/host-agents/skills/rgctl-kantra/SKILL.md b/agent-pack/out/agents/host-agents/skills/rgctl-kantra/SKILL.md deleted file mode 100644 index d200c8c1..00000000 --- a/agent-pack/out/agents/host-agents/skills/rgctl-kantra/SKILL.md +++ /dev/null @@ -1,40 +0,0 @@ ---- -name: rgctl-kantra -description: "Konveyor Kantra rules. Use for rgctl kantra workflow. Spawn rgctl -f json; parse schema_version from stdout." -rgctl-managed: true -metadata: - generatedBy: "rgctl 0.4.13" ---- - -# Kantra workflow - -**Primary output:** `.rgctl/kantra_findings.json` and `KantraRule` / `VIOLATES` in the graph. - -**This workflow is not migration roadmap export.** Do not present `migration_plan.json` as the main Kantra deliverable. - -```bash -rgctl -r "$REPO" discover . -l java --with-kantra -rgctl -r "$REPO" discover . -l java --with-kantra --kantra-target quarkus -rgctl -r "$REPO" -f json gql 'MATCH (r:KantraRule) RETURN r LIMIT 20' -``` - -Overrides: `--kantra-rules DIR`, `--kantra-catalog ROOT`, `--kantra-index-only` (index without eval). - -For extraction ordering after violations, use the **migrate** workflow separately. - - -## Agent loop - -1. Parse the user question (natural language). -2. Run `rgctl -f json …` (or `rgctl serve` + HTTP for repeated queries). -3. Parse `schema_version` and payload from **stdout** only. -4. Summarize facts; do not dump raw JSON. -5. Re-query if the graph may be stale after edits. - -**Never** redirect stderr to `/dev/null`. If `.rgctl/` exists and the question is structural, use rgctl before ripgrep or bulk file reads. - -```bash -export REPO=/path/to/repo -rgctl -r "$REPO" -f json -``` - diff --git a/agent-pack/out/agents/host-agents/skills/rgctl-migrate/SKILL.md b/agent-pack/out/agents/host-agents/skills/rgctl-migrate/SKILL.md deleted file mode 100644 index 8395c20b..00000000 --- a/agent-pack/out/agents/host-agents/skills/rgctl-migrate/SKILL.md +++ /dev/null @@ -1,41 +0,0 @@ ---- -name: rgctl-migrate -description: "Migration roadmap. Use for rgctl migrate workflow. Spawn rgctl -f json; parse schema_version from stdout." -rgctl-managed: true -metadata: - generatedBy: "rgctl 0.4.13" ---- - -# Migrate workflow - -**Primary output:** `.rgctl/migration_plan.json` (and dashboard migration view via `serve --open`). - -**This workflow is not Kantra.** Do not treat `--with-kantra` or `kantra_findings.json` as the main deliverable here. - -```bash -rgctl -r "$REPO" discover . --with-cfg --with-harmonic --export-migration-hints \ - --migration-preset hybrid_default --migration-order scheduled -rgctl -r "$REPO" -f json metrics --pagerank -``` - -Presets: `hybrid_default`, `foundational_first`, `dense_cluster`, `risk_mitigation`. -Orders: `scheduled` (dependency-aware), `priority` (score rank). - -Report plan path, preset/order, and top packages — not raw discover telemetry. - - -## Agent loop - -1. Parse the user question (natural language). -2. Run `rgctl -f json …` (or `rgctl serve` + HTTP for repeated queries). -3. Parse `schema_version` and payload from **stdout** only. -4. Summarize facts; do not dump raw JSON. -5. Re-query if the graph may be stale after edits. - -**Never** redirect stderr to `/dev/null`. If `.rgctl/` exists and the question is structural, use rgctl before ripgrep or bulk file reads. - -```bash -export REPO=/path/to/repo -rgctl -r "$REPO" -f json -``` - diff --git a/agent-pack/out/agents/host-agents/skills/rgctl-search/SKILL.md b/agent-pack/out/agents/host-agents/skills/rgctl-search/SKILL.md deleted file mode 100644 index 06385e82..00000000 --- a/agent-pack/out/agents/host-agents/skills/rgctl-search/SKILL.md +++ /dev/null @@ -1,36 +0,0 @@ ---- -name: rgctl-search -description: "Semantic and structural search. Use for rgctl search workflow. Spawn rgctl -f json; parse schema_version from stdout." -rgctl-managed: true -metadata: - generatedBy: "rgctl 0.4.13" ---- - -# Search workflow - -**When:** Natural-language or intent-based code location (requires `semantic index`). - -```bash -rgctl -r "$REPO" -f json semantic query "…" [--limit 10] -rgctl -r "$REPO" -f json semantic query "…" --scope community --limit 10 -rgctl -r "$REPO" -f json gql --macro-name all_communities unused -``` - -Fusion is on by default for semantic query; use GQL for exact graph patterns. - - -## Agent loop - -1. Parse the user question (natural language). -2. Run `rgctl -f json …` (or `rgctl serve` + HTTP for repeated queries). -3. Parse `schema_version` and payload from **stdout** only. -4. Summarize facts; do not dump raw JSON. -5. Re-query if the graph may be stale after edits. - -**Never** redirect stderr to `/dev/null`. If `.rgctl/` exists and the question is structural, use rgctl before ripgrep or bulk file reads. - -```bash -export REPO=/path/to/repo -rgctl -r "$REPO" -f json -``` - diff --git a/agent-pack/out/agents/host-agents/skills/rgctl/README.md b/agent-pack/out/agents/host-agents/skills/rgctl/README.md index 3d832943..c0fedf57 100644 --- a/agent-pack/out/agents/host-agents/skills/rgctl/README.md +++ b/agent-pack/out/agents/host-agents/skills/rgctl/README.md @@ -4,101 +4,85 @@ A skill for answering structural questions about codebases using the rgctl CLI g ## Quick Stats -- **Main skill:** 352 lines (57% reduction from original 814 lines) -- **Total documentation:** 1,505 lines (85% more comprehensive coverage) -- **Files:** 5 reference files + main skill -- **Workflow families:** 6 + Kantra rules -- **NL routing examples:** 25+ common user utterances +- **One skill:** `rgctl` (router + structured verb tables + references) +- **Reference files:** command encyclopedia, workflows (assembled), communities & policy +- **Workflow families (docs only):** discover, impact, flow, search, migrate, kantra, gate +- **Agent query path:** `find` / `callers` / `callees` / `relations` / `inventory` / `status` (no Cypher) ## Structure ``` skills/rgctl/ -├── SKILL.md # Main skill (352 lines) +├── SKILL.md # Main skill ├── README.md # This file +├── workflows/ # Fragments assembled into references/workflows.md at build └── references/ - ├── command-encyclopedia.md # All commands with JSON samples (19KB) - ├── workflows.md # Migration, Kantra rules, refactor, audit scenarios - ├── gql-reference.md # GQL patterns & limitations (4.7KB) - └── communities-and-policy.md # Community detection + CI policy (13KB) + ├── command-encyclopedia.md # All commands with JSON samples + ├── workflows.md # Generated from workflows/ at build (do not edit by hand) + └── communities-and-policy.md # Community detection + CI policy ``` +`rgctl install --skill` writes **only** this skill tree (e.g. `.cursor/skills/rgctl/`). It does **not** create separate `rgctl-discover` / `rgctl-impact` / … skill directories. + ## What's Covered ### Main SKILL.md (Always Loaded) - When to use rgctl - **CLI subprocess workflow** — spawn `rgctl -f json` for agents -- **6 workflow families:** +- **Workflow families:** 1. Discovery & Indexing 1b. Konveyor Kantra rules (`--with-kantra`) - 2. Query & Search (includes communities + KantraRule GQL) + 2. Query & Search (structured verbs + communities) 3. Impact & Safety (includes policy checks) 4. Metrics & Analysis 5. Code Analysis (CFG/PDG/slicing) 6. Export & Visualization -- **NL routing table** (20+ user utterances → commands) -- Common scenarios (migration, pre-refactor safety) +- **NL routing table** (user utterances → commands) - Failure playbook ### References (Loaded On-Demand) #### command-encyclopedia.md -- All 15+ commands with full details +- Structured query verbs and domain commands - JSON sample responses - Prerequisites and pitfalls - "What to report" guidelines -#### workflows.md -- Migration & audit workflows (incl. Konveyor Kantra `--with-kantra`) -- Intent discovery & subsystem mapping -- Pre-refactor safety analysis -- CI gates & policy -- Advanced patterns - -#### gql-reference.md -- Cypher subset capabilities -- Macros (all_functions, all_communities) -- Valid edge types -- LIKE pattern matching limitations -- Common patterns & troubleshooting +#### workflows.md (generated) +- Assembled from `workflows/*.md` when the agent pack is built (`cargo build`) +- Edit fragments under `workflows/` (e.g. `migrate.md`, `kantra.md`); order comes from `agent-pack/manifest.yaml` #### communities-and-policy.md -- **Community Detection:** - - What communities are (implicit architecture) - - Commands (list, query, label, semantic scope) - - Use cases (microservice extraction, ownership) - - 5 complete workflows -- **CI Policy Checks:** - - Policy schema (max_impact_nodes, centrality, forbidden_crossings) - - CI integration (GitHub Actions, GitLab) - - Crafting policies (calibration, gradual tightening) - - 4 complete workflows -- Combined workflows using both features +- **Community Detection:** list, semantic scope, ownership workflows +- **CI Policy Checks:** schema, CI integration, calibration ## Design Principles -✅ **Progressive disclosure** - Main skill <500 lines, details in references +✅ **Progressive disclosure** - Main skill lean, details in references ✅ **Workflow-centric** - Organized by user intent, not commands -✅ **CLI-first** - Agents use `rgctl -f json` subprocesses (optional `serve` HTTP) -✅ **Clear routing** - Natural language → tool mapping -✅ **Comprehensive** - All features documented with examples -✅ **Integration** - Shows how features work together +✅ **CLI-first** - Agents use `rgctl -f json` structured verbs +✅ **Clear routing** - Natural language → command mapping +✅ **No Cypher in skill surface** - Agents must not invent MATCH strings ## Installation -From another repo: +From a target repository (not the rgctl source tree unless you are dogfooding): ```bash -rgctl install --skill +rgctl install --skill --tools cursor,claude,codex,antigravity,agents ``` -This writes `.claude/skills/rgctl/`, `.agents/skills/rgctl/`, and `.cursor/skills/rgctl/` from the embedded skill. +Installs a single skill named `rgctl`. See [Agent pack walkthrough](../../docs/guides/agent-skill.md). + +If an older pack left `rgctl-discover` / `rgctl-impact` / … directories behind, delete them manually — install no longer writes those folders. + +**Maintainers:** edit workflow bodies under `workflows/`; regenerate `references/workflows.md` with `assemble_workflows_reference` (see `rgctl-agent-pack-codegen` test `workflows_reference_matches_fragments`). ## See Also - [User Guide](../../docs/user-guide.md) - Complete CLI tutorial - [Agent recipes](../../docs/agent-recipes.md) - Copy-paste CLI workflows -- [HTTP API](../../docs/http-api.md) - Optional `rgctl serve` for repeated queries +- [HTTP Server and Dashboard](../../docs/guides/http-server-and-dashboard.md) - Optional `rgctl serve` for dashboard - [JSON API](../../docs/json-api.md) - Schema specifications - [All Guides](../../docs/guides/README.md) - Feature-specific guides diff --git a/agent-pack/out/agents/host-agents/skills/rgctl/SKILL.md b/agent-pack/out/agents/host-agents/skills/rgctl/SKILL.md index f0a8b0b0..1da326ef 100644 --- a/agent-pack/out/agents/host-agents/skills/rgctl/SKILL.md +++ b/agent-pack/out/agents/host-agents/skills/rgctl/SKILL.md @@ -8,6 +8,7 @@ description: >- impact of changing a symbol, where data flows, migration rule violations, Konveyor/quarkus/spring targets, repo structure/hotspots, or when `.rgctl/` exists — treat natural-language codebase questions as rgctl queries first. +rgctl-managed: true --- # rgctl @@ -37,15 +38,15 @@ rgctl -r "$REPO" -f json … **Critical:** Parse `schema_version` + payload from **stdout**. **Never use `2>/dev/null`** — it swallows rgctl errors. -For many queries in one session, optional: `rgctl serve` + `POST /api/query` (see [HTTP API](../../docs/http-api.md)). +For interactive exploration, optional: `rgctl serve --open` (dashboard). Agents should still spawn CLI structured verbs (`find` / `callers` / `relations` / `inventory` / …). -Legacy daemon cache: `rgctl migrate-cache` copies `~/.rgctl/cache/{name}/.rgctl/` into the repo. +Legacy daemon cache under `~/.rgctl/cache/` is obsolete; run `rgctl discover .` in the repo to build `{repo}/.rgctl/`. ## Agent Loop ```text 1. USER PROMPT → natural language (not a CLI string) -2. SUBPROCESS → rgctl -f json (or HTTP /api/query) +2. SUBPROCESS → rgctl -f json 3. GRAPH FACTS → parse schema_version + payload 4. LLM REASONING → summarize using "what to report" guidelines 5. ACTION → edit / plan / check — re-query if graph may be stale @@ -90,30 +91,44 @@ Legacy daemon cache: `rgctl migrate-cache` copies `~/.rgctl/cache/{name}/.rgctl/ | User Intent | CLI Command | |-------------|-------------| -| Evaluate migration rules | `discover . --with-kantra` | -| Filter by migration target | `discover . --with-kantra --kantra-target quarkus` | -| CI / custom ruleset | `discover . --with-kantra --kantra-rules PATH` | -| List indexed rules (GQL) | `gql "MATCH (r:KantraRule) RETURN r LIMIT 20"` | -| Rules for one target label | `gql` with `` r.`konveyor.io/target` `` property (backticks) | +| Evaluate migration rules | `discover . --with-kantra` or `rules run ./rules/` | +| Filter by migration target | `discover . --with-kantra --kantra-target quarkus` / `rules run ./rules/ --target quarkus` | +| CI / custom ruleset | `discover . --with-kantra --kantra-rules PATH` / `rules run PATH` | +| Index rules only | `discover . --with-kantra --kantra-index-only` | +| List indexed rules | `find --type kantrarule --limit 50` / `inventory --by type` | +| Rule → code links | `relations --edge violates --from-type kantrarule` | | Read violations artifact | `.rgctl/kantra_findings.json` | **See:** [User guide — Kantra](../../docs/user-guide.md#kantra-migration-rules---with-kantra), [JSON — kantra_findings](../../docs/json-api.md#kantra_findingsjson) ### 2. Query & Search +**Prefer structured verbs** (mmap; no Cypher). Parse `-f json` from **stdout** (`schema_version`); never `2>/dev/null`. + | User Intent | CLI Command | |-------------|-------------| -| Inventory functions | `gql --macro-name all_functions unused` | -| Find callers/callees | `gql "MATCH (a)-[:CALLS]->(b) WHERE ..."` | +| Session / index freshness | `status` | +| Schema / counts (incl. zeros) | `inventory --by type` or `inventory --by edge` | +| Import prefix census | `inventory --by import-prefix` | +| Count functions | `find --type function --count-only` | +| Find by name/type | `find "User*" --type class --limit 50` | +| Suffix scan (MDB / Remote) | `find '*MDB*' --type class` | +| Classes with annotation | `find --annotation @MessageDriven --type class` | +| javax import worklist | `find "import javax*" --type import --scope ` | +| Annotation pairs (seedless) | `relations --edge annotatedwith --from-type function --to-type annotation --scope ` | +| Find callers/callees | `callers --depth 1` / `callees ` | +| Outside callers of a module | `callers --scope --scope-mode outside` | +| EXTENDS / IMPLEMENTS inventory | `relations --edge extends --from-type class` (omit SYMBOL) | | Natural-language search | `semantic query "checkout flow"` | | List communities | `communities list` | -| Community members | `gql "MATCH (f) WHERE f.community_id='12'"` | | Subsystem ownership | `semantic query "X" --scope community` | | Refresh community labels | `communities label --write` | +| Community census | `inventory --by community` | -**GQL limitations:** no `COUNT`/`ORDER BY`; LIKE prefix/suffix only; CALLS misses dynamic dispatch; Konveyor labels need backticks in `WHERE`. +**Migration probe order:** `status` → `inventory --by import-prefix` → `find --annotation …` / suffix globs → `rules run` / `--with-kantra` → `callers InitialContext`. -**See:** [GQL Reference](references/gql-reference.md), [Semantic Search Guide](../../docs/guides/semantic-search.md) +**Complexity honesty:** exact name = hash index; prefix/`*mid*`/`--scope` may scan keys/columns until better indexes land. Module re-index is still a strong speed lever. Annotation **arguments** (e.g. `@Path("/x")`) need `--show-attributes` when `annotation_args.json` is present. +**See:** [Command Encyclopedia](references/command-encyclopedia.md) (find/callers/relations/inventory/status), [Semantic Search Guide](../../docs/guides/semantic-search.md) ### 3. Impact & Safety @@ -164,19 +179,21 @@ Needs `discover --with-cfg`. `--function` is method name, not class. | "Where is checkout flow?" | `semantic query "checkout flow" --limit 10` | | "Impact if I change X" | `blast-radius X --depth 2` | | "Validate against policy" | `check --policy-file policy.json` | -| "Who calls X" | `gql "MATCH (a)-[:CALLS*1..3]->(b) WHERE a.name='X' RETURN a,b"` | +| "Who calls X" | `callers X --depth 2` (impact → `blast-radius X`) | +| "javax imports / annotations" | `find "import javax*" --type import`; `relations --edge annotatedwith --from-type function --to-type annotation` | | "Where is X mutated?" | `cpg mutations --type X --exclude-ctors` | ## Failure Playbook | Symptom | Fix | |---------|-----| -| No `.rgctl/` in repo | Run `cd repo && rgctl discover .`; or `rgctl migrate-cache` from legacy daemon cache | +| No `.rgctl/` in repo | Run `cd repo && rgctl discover .` | | slice/inspect/cpg fails | Re-discover with `--with-cfg` | | semantic query fails | `semantic index` | -| Ambiguous symbol | Add `--class` or `--file`; disambiguate via GQL | +| Ambiguous symbol | Add `--class` or `--file` on callers/find | | `check` exit 1 | Report violations (JSON still on stdout) | -| GQL LIKE returns 0 | Try `communities list`, `semantic query`, or broader type patterns | +| find/relations empty | Run `inventory --by type` / `--by edge` (zeros mean unpopulated schema); check `--scope` | +| Name glob returns 0 | Try `semantic query` / `communities list` / broader `find '*X*'` types | ## Artifacts @@ -205,7 +222,6 @@ rgctl -r "$REPO" -f json … - **[Command Encyclopedia](references/command-encyclopedia.md)** — Full command reference - **[Workflows](references/workflows.md)** — Worked scenarios -- **[GQL Reference](references/gql-reference.md)** — GQL patterns - **[Communities & Policy](references/communities-and-policy.md)** — CI policy checks ## External Documentation @@ -213,30 +229,14 @@ rgctl -r "$REPO" -f json … - [User Guide](../../docs/user-guide.md) — Complete CLI tutorial - [JSON API](../../docs/json-api.md) — Schema specifications - [Agent Recipes](../../docs/agent-recipes.md) — Copy-paste recipes -- [AGENTS.md](../../AGENTS.md) — Minimal agent contract +- [USER_AGENTS_TEMPLATE.md](../../docs/agents/USER_AGENTS_TEMPLATE.md) — paste into consumer repos +- [AGENTS.md](../../AGENTS.md) — contributor agent README (rgctl source tree) - [Policy Format](../../docs/policy-format.md) — CI policy schema ## Installation ```bash -rgctl install --skill +rgctl install --skill --tools cursor,claude,codex,antigravity,agents ``` -Writes `.claude/skills/rgctl/`, `.agents/skills/rgctl/`, and `.cursor/skills/rgctl/` from the embedded skill in the binary. - - -## Workflow slash commands (generated) - -| Intent | Command | -|--------|---------| -| Index and discover | `/rgctl-discover` | -| Blast radius and impact | `/rgctl-impact` | -| Data flow and slices | `/rgctl-flow` | -| Semantic and structural search | `/rgctl-search` | -| Graph query language | `/rgctl-gql` | -| Migration roadmap | `/rgctl-migrate` | -| Konveyor Kantra rules | `/rgctl-kantra` | -| CI and policy gates | `/rgctl-gate` | - -**Migrate** (roadmap / `migration_plan.json`) and **Kantra** (rules / `kantra_findings.json`) are separate workflows — do not conflate. -rgctl-managed router note generatedBy rgctl 0.4.13 +Installs the single skill `rgctl` (with `references/`) for each selected adapter. Omitting `--tools` installs **cursor, claude, codex, agents, antigravity**; use `--tools all` for the full registry. See [docs/guides/agent-skill.md](../../docs/guides/agent-skill.md). Scenario prose lives under `workflows/` and is assembled into `references/workflows.md` (`cargo test -p rgctl-agent-pack-codegen workflows_reference_matches_fragments`). diff --git a/agent-pack/out/agents/host-agents/skills/rgctl/references/command-encyclopedia.md b/agent-pack/out/agents/host-agents/skills/rgctl/references/command-encyclopedia.md index 5e48b95c..857fe2af 100644 --- a/agent-pack/out/agents/host-agents/skills/rgctl/references/command-encyclopedia.md +++ b/agent-pack/out/agents/host-agents/skills/rgctl/references/command-encyclopedia.md @@ -7,7 +7,7 @@ Samples below are truncated where noted. Field names match live CLI / `docs/json ## Table of Contents - [discover](#discover) -- [gql](#gql) +- [find / callers / callees / relations / inventory](#find--callers--callees--relations--inventory) - [blast-radius](#blast-radius) - [slice](#slice) - [inspect](#inspect) @@ -90,9 +90,9 @@ Samples below are truncated where noted. Field names match live CLI / `docs/json } ``` -**GQL companion:** After discover, query indexed rules with `gql "MATCH (r:KantraRule) RETURN r"`. Konveyor labels are properties — filter with backticks: `` r.`konveyor.io/target` ``. +**Structured companion:** After discover, list indexed rules with `find --type kantrarule` / `inventory --by type`. Rule→code links after full eval: `relations --edge violates --from-type kantrarule`. Prefer `kantra_findings.json` for line-level violations. -**Pitfalls:** Full embedded catalog skips many rules (unsupported providers, Windup regex). Use `kantra_findings.json` for violation details; GQL `VIOLATES` edges link rules to code nodes after full eval (not `--kantra-index-only`). +**Pitfalls:** Full embedded catalog skips many rules (unsupported providers, Windup regex). Use `kantra_findings.json` for violation details; `VIOLATES` edges exist after full eval (not `--kantra-index-only`). **Agent should report:** `catalog_id`, `target_filter`, violation count, representative hits, skip summary — not full JSON dump. @@ -100,44 +100,49 @@ Samples below are truncated where noted. Field names match live CLI / `docs/json --- -## gql +## find / callers / callees / relations / inventory -**Command:** `rgctl -f json gql ''` or `rgctl -f json gql --macro-name unused` +**Commands:** -**Purpose:** Inventory, callers/callees, communities, path/relationship queries. +```bash +rgctl -f json find [PATTERN] --type function --scope pkg --limit 50 +rgctl -f json find --type function --count-only +rgctl -f json find --annotation @MessageDriven --type class +rgctl -f json find --annotation @Stateful,@Stateless,@Singleton --type class +rgctl -f json find '*MDB*' --type class --limit 50 # bare-name suffix scan +rgctl -f json callers --depth 1 --file PATH --class NAME --line N +rgctl -f json callees --depth 1 +rgctl -f json relations [SYMBOL] --edge annotatedwith --from-type function --to-type annotation --scope pkg +rgctl -f json relations --edge extends --from-type class # seedless +rgctl -f json inventory --by type # includes zero-count kinds +rgctl -f json inventory --by edge +rgctl -f json inventory --by import-prefix # javax.ejb / javax.jms / org.eclipse … +rgctl -f json status # snapshot presence, digest, node/edge counts +rgctl discover . --find '*coolstore*' # locate candidate project roots (no index) +rgctl -f json rules run ./rules/ [--target quarkus] # post-index Kantra eval +rgctl -f json find --annotation @Resource --show-attributes # needs annotation_args.json from discover +rgctl -f json query find … # alias namespace +``` -**Prerequisites:** `discover` done. Virtual `:Community` needs analysis overlay from discover. +**Purpose:** Deterministic mmap structured query (no Cypher, no `MemoryBackend` hydrate). **Agents must use these verbs** — do not invent MATCH strings. Relations `total` is distinct `(source,target,edge)`; duplicates collapse with `occurrences` (`schema_version` ≥ 2). `inventory --by edge` uses the same rule: `count` = distinct, `occurrences` = raw stored edges. -**Sample** (macro `all_functions`): +**Migration probes (Coolstore-shaped):** +1. `status` — is `.rgctl/` fresh? +2. `inventory --by import-prefix` — EE surface census +3. `find --annotation @MessageDriven|@SessionScoped|…` — blockers without package guess +4. `find '*MDB*'` / `'*Remote*'` — suffix scan before reading files +5. `rules run ./rules/` or `discover --with-kantra` — fire `when:` catalog (M2) +6. `callers InitialContext` — JNDI usage sites -```json -{ - "schema_version": 1, - "count": 260, - "rows": [ - [{ "binding": "f", "node": "addItem", "type": "Function", - "file": "…/controller/CartController.java" }] - ], - "explain": false -} -``` +**Prerequisites:** `discover` done (columnar `graph.snapshot.bin`). `status` does not rediscover. `rules run` requires a snapshot; Kantra stays opt-in. -**Useful patterns:** +**Flags:** `--annotation` inverts `AnnotatedWith` (OR list; `@` optional). `--show-attributes` needs annotation-arg indexing (errors honestly until indexed). `--scope` + `--scope-mode inside|outside|crossing` (or `--exclude-scope`). `--file` / `--class` / `--line` disambiguate. Edge rows use keyed `source`/`target` (never positional). Omit `SYMBOL` on `relations` for set-wide typed-edge scans. -```bash -# Incoming callers of X -rgctl -f json gql "MATCH (a:Function)-[:CALLS]->(b:Function) WHERE b.name = 'checkout' RETURN a,b LIMIT 20" -# Outgoing callees of X -rgctl -f json gql "MATCH (a:Function)-[:CALLS]->(b:Function) WHERE a.name = 'checkout' RETURN a,b LIMIT 20" -# Name search (prefix or suffix only — *middle* silently returns 0) -rgctl -f json gql "MATCH (n:Function) WHERE n.name LIKE '*Service' RETURN n LIMIT 20" -# Communities macro -rgctl -f json gql --macro-name all_communities unused -``` +**Pitfalls:** Exact name is O(1) hash; prefix/contains/`--scope` may scan. Ambiguous symbols emit candidates (`error: ambiguous_symbol` JSON under `-f json`). Annotation argument values are not in the graph yet. Warm caches invalidate wall-time claims — label cold vs warm. Do not scrape stderr; parse `schema_version` on stdout. Do not treat Kantra as the only search path — use annotation/import first. -**Pitfalls:** `--macro-name` still needs a positional query arg — pass `unused`. `--explain` plan is text-mode only. rgctl GQL is a **subset of Cypher** — no `COUNT`, `ORDER BY`, `GROUP BY`, or aggregation functions. CALLS edges are static — interface / dynamic dispatch (receiver methods, virtual calls, trait impls) may not appear; if a `CALLS*1..N` query returns 0 edges for a method you know is called, fall back to `grep` for call sites. If LIKE on function names returns 0 for a concept (e.g. "ingress", "gateway"), it likely lives in package/directory names, type names, or community labels — try `communities list`, `semantic query`, or broaden the LIKE to non-Function node types before concluding nothing exists. +**Agent should report:** counts, lean names/files, keyed edge pairs — not full node dumps. -**Agent should report:** matching symbols, files, hop relationships — not raw row dumps. +**See:** OpenSpec `add-migration-search-primitives` (+ `add-structured-query-cli`). --- @@ -247,7 +252,7 @@ rgctl -f json slice src/main/java/com/example/ecommerce/service/CartService.java **Purpose:** Raw CFG / PDG / dominator view for one function. -**Prerequisites:** `discover --with-cfg`. Symbol only — **no** `--class` (disambiguate via blast-radius / GQL first). +**Prerequisites:** `discover --with-cfg`. Symbol only — **no** `--class` (disambiguate via `find` / `blast-radius` / `callers` first). **Layer flags:** `cfg --prune` drops unreachable blocks before display. `pdg --edge-layer data|control` filters to one dependence type (default `all`); `--def-use` adds def-use variable lists per node. `dom --frontiers` prints dominance frontiers instead of just the tree. @@ -322,7 +327,7 @@ rgctl -f json cpg function '' rgctl -f json blast-radius '' ``` -Loop over `top[]` UUIDs and resolve each. GQL `WHERE n.id = ''` does **not** work (node id is not a queryable property). +Loop over `top[]` UUIDs and resolve each with `cpg function` / `blast-radius` (node id is not a `find` name). **Agent should report:** top hotspot symbols (resolve UUIDs first), modularity/community count when requested. @@ -337,7 +342,7 @@ rgctl semantic index [--embedder vocab|hash|onnx|code-daemon] [--embed-bodies] [ [--dimensions N] [--incremental] [--diffuse] [--diffuse-alpha F] [--diffuse-iters N] [--diffuse-bidirectional] rgctl semantic distill --matrix PATH [--embedder code-daemon|hash|onnx] [--tokens PATH] [--dimensions N] rgctl -f json semantic query "…" [--limit N] [--scope function|community] \ - [--expand neighbors|blast|gql|all] [--expand-depth N] [--no-fusion] [--candidate-pool N] [--keyword-and] + [--expand neighbors|blast|all] [--expand-depth N] [--no-fusion] [--candidate-pool N] [--keyword-and] ``` **Purpose:** Natural-language / keyword find of functions (and community-scoped search), with optional one-shot expansion into graph context. @@ -346,7 +351,7 @@ rgctl -f json semantic query "…" [--limit N] [--scope function|community] \ **Index tuning:** `--dimensions` (default 256, multiple of 8) trades index size for precision. `--incremental` (default true) reuses embeddings for unchanged `code_hash`. `--diffuse` blends each embedding toward its call-graph neighbors' mean (Jacobi iterations via `--diffuse-alpha`/`--diffuse-iters`; `--diffuse-bidirectional` includes callers, not just callees) — useful when bare-name/docstring signal is weak and callers/callees disambiguate intent; `--no-diffuse` forces it off. -**Query expansion:** `--expand neighbors` pulls CALLS neighbors of top hits, `--expand blast` runs blast-radius on top hits, `--expand gql` returns a ready GQL query, `--expand all` does all three — use when the user's NL query implies "and show me what's connected," so you skip a manual follow-up call. `--expand-depth` controls hop depth for `neighbors`/`gql` expansion (default 1). `--no-fusion` returns pure Hamming top-k (skip late-fusion re-ranking — rarely needed). `--candidate-pool` widens/narrows the pre-fusion candidate set (default 256). `--keyword-and` requires all query keywords to match entry metadata (stricter than default OR). +**Query expansion:** `--expand neighbors` pulls CALLS neighbors of top hits, `--expand blast` runs blast-radius on top hits, `--expand all` combines those — use when the user's NL query implies "and show me what's connected," so you skip a manual follow-up call. `--expand-depth` controls hop depth for `neighbors` expansion (default 1). Prefer follow-up `callers` / `blast-radius` over any Cypher expand mode. `--no-fusion` returns pure Hamming top-k (skip late-fusion re-ranking — rarely needed). `--candidate-pool` widens/narrows the pre-fusion candidate set (default 256). `--keyword-and` requires all query keywords to match entry metadata (stricter than default OR). **Sample** (default vocab, query `checkout cart`): @@ -374,7 +379,7 @@ rgctl -f json semantic query "…" [--limit N] [--scope function|community] \ **Pitfalls:** Query without index fails. Restart `serve` after rebuilding index for dashboard search. **Large repos (100K+ nodes):** `--scope community` may return only singleton communities because label-propagation produces very granular clusters. For subsystem ownership on large repos, prefer `communities list` + grep labels over `--scope community`. -**Agent should report:** top hit names, files, scores (`score` / `fused_score`); keep `node_id` for follow-up GQL — not every hit. +**Agent should report:** top hit names, files, scores (`score` / `fused_score`); keep `node_id` for follow-up `callers` / `blast-radius` — not every hit. --- @@ -399,7 +404,7 @@ rgctl -f json semantic query "…" [--limit N] [--scope function|community] \ } ``` -**Agent should report:** top labels + sizes; use GQL `community_id` for members. +**Agent should report:** top labels + sizes; use `inventory --by community` for census; explore ownership via `semantic query --scope community` / `blast-radius` (responses may include `community_id`). --- @@ -523,7 +528,7 @@ rgctl cpg export --format graphson --output cpg.json [--path-contains src/] \ **Prerequisites:** `discover` done. -**Pitfalls (critical):** `--query` uses **filter** syntax — `name:Foo`, `type:Function`, `all` — **not** GQL `MATCH … RETURN`. Agents must not pass MATCH strings to `--query`. +**Pitfalls (critical):** `--query` uses **filter** syntax — `name:Foo`, `type:Function`, `all` — **not** Cypher `MATCH … RETURN`. Agents must not pass MATCH strings to `--query`. **Agent should report:** output path + format; confirm filter used. diff --git a/agent-pack/out/agents/host-agents/skills/rgctl/references/communities-and-policy.md b/agent-pack/out/agents/host-agents/skills/rgctl/references/communities-and-policy.md index 1838f662..1804f37c 100644 --- a/agent-pack/out/agents/host-agents/skills/rgctl/references/communities-and-policy.md +++ b/agent-pack/out/agents/host-agents/skills/rgctl/references/communities-and-policy.md @@ -65,11 +65,15 @@ rgctl -f json communities list #### Query Community Members -Once you have a community ID, list its members: +Once you have a community ID from `communities list`, explore ownership and members via structured verbs: **CLI:** ```bash -rgctl -f json gql "MATCH (f:Function) WHERE f.community_id = '12715' RETURN f LIMIT 20" +rgctl -f json inventory --by community +rgctl -f json semantic query "