Skip to content

chore(specs): gardener checkbox sync - #1333

Merged
Dumbris merged 20 commits into
mainfrom
claude/spec-gardener
Oct 6, 2026
Merged

Dumbris merged 20 commits into
mainfrom
claude/spec-gardener

Conversation

@Dumbris

@Dumbris Dumbris commented Sep 21, 2026 •

Copy link
Copy Markdown
Member

5 ticks applied, 0 un-ticks proposed (one already-flagged candidate intentionally not re-proposed a third time), ~380 candidates examined and dropped, run date 2026-10-05.

This run continues the open gardener PR. It merged the branch forward to the current main (648a618) — the previous rebase attempt hit a catastrophic conflict against unrelated pre-history reachable only from this branch, so a merge commit was used instead, exactly as the prior run did. That merge brought in 16 tasks (15 in 109-ux-navigation-consistency + T011a) that the previous run had already verified but couldn't apply because it hit the 40-tick cap — they turned out to already be ticked directly on main by the feature PRs that landed them, so no gardener action was needed on those.

Scope of this run: scripts/check-spec-evidence.py (run against the merged, current tree) flagged 186 possibly_built tick candidates and 92 currently-ticked tasks with unresolved evidence (70 UNRESOLVED + 22 REMOVED) across 47 specs, plus 47 RELOCATED informational findings. Every candidate was hand-verified against current code by 8 parallel review passes (grouped by spec), each applying the tick test (behavior over keyword-matching, real vs. stub/unwired, a named test must actually assert the behavior, an adversarial "argue the opposite" pass) before anything changed. I independently re-verified each of the 5 ticks below myself (reading the cited file:line directly) before applying them.

No cap issue this run — only 5 ticks survived verification, well under the 40-tick limit.

Applied ticks

Task file:line Decisive evidence
004-management-health-refactor T086 internal/logs/sanitizer.go:12,26 (SecretSanitizer, NewSecretSanitizer) Token-redaction core wired into every logger path: internal/logs/logger.go:94,381,449,494 all call NewSecretSanitizer(core). Not the literal redactToken/sanitizeAuthHeader names, but the same masking behavior, real and globally applied.
004-management-health-refactor T087 internal/logs/sanitizer_test.go:17,40 (TestSanitizer_GitHubTokens, TestSanitizer_LongStatelessGitHubToken) Real tests asserting the sanitizer masks tokens — companion test task to T086.
009-proactive-oauth-refresh T054 frontend/src/utils/health.ts:225-233 (oauthSignInState) + frontend/src/components/ServerCard.vue:376-416 No literal oauthExpired computed property, but the exact promised behavior ("connected server with expired OAuth token ⇒ Login button visible") is implemented end-to-end via Spec 109's unified health-action system (health.action === Login → sign-in CTA, rendered by ServerCard.vue's primaryAction).
009-proactive-oauth-refresh T060 frontend/src/components/ServerCard.vue:597-619 (canLogout computed, gates the Logout menu item at line 110) canLogout is the authentication-gated computed property the task asks for under a different name, and it drives the real Logout button.
073-activity-size-retention T012 docs/features/activity-log.md:466,478 activity_max_size_mb default/semantics fully documented (JSON example + table row + rationale paragraph). docs/configuration.md has no Activity-Log section at all for any of this spec's fields, consistent with existing project convention rather than a gap specific to this task.

Already-ticked-on-main (not gardener-applied, noted for the record)

These 16 tasks were verified-but-deferred by the previous gardener run (cap-limited) and this run found them already [x] on main after the forward-merge — landed directly by their feature PRs, not by this routine. No diff, no action, listed here only so the history is clear: 109-ux-navigation-consistency T047, T097, T098, T099, T100, T101, T104, T105, T107, T109a, T114, T115, T116, T117, T121, T011a.

Proposed un-ticks (NOT applied)

None new this run. 102-schema-deferred T026 (regression test for the R14 race, internal/server/server.go:382-395's SubscribeEvents() hoist) was independently re-found by this run's reviewer as still missing a dedicated test — but it has already been proposed twice in this PR's history without being applied, so it is not re-proposed a third time. It remains an open item for a human to action directly.

Worth a human glance (not proposed as an un-tick)

Carried forward from last run, still unresolved: 107-server-edition-sso-hardening T101 — the caller-kind-derivation half is real and tested, but the large combinatorial surfaces×caller-kinds×situations matrix test the task also promises could not be found at the scale described (closest match, audit_funnel_test.go, is ~40-50x smaller). Flagging again for a maintainer to confirm nothing was silently dropped.

Dropped by verification

Tick candidates rejected — real underlying code/feature exists, but the specific promised behavior (a named test, a doc section, a UI element, masking, CLI flag) does not, or is only partially built:

  • 001-code-execution (13 of 13): all are TDD placeholder tests never actually written — nested-input serialization, execution_id uniqueness assertion, upstream-failure integration test, CLI flag tests (T092-96) test different plumbing (client-timeout/ping-fallback) than what's named.
  • 001-oas-endpoint-documentation (6 of 6): /secrets/refs, /secrets/config, /secrets/migrate, /code/exec, /events (GET+HEAD) all lack swag annotations and are absent from oas/swagger.yaml.
  • 001-update-version-display (6 of 6): tray version-item test, update-field E2E assertion, startup-check/periodic-ticker tests, "Check for Updates" removal test all absent (the underlying removal/display code is real, just untested).
  • 003-tool-annotations-webui (2 of 2): no sessionId param on getToolCalls; no session-lifecycle SSE events.
  • 004-management-health-refactor (9 of 9 remaining): no restart E2E test, no logHTTPRequest, Docker log streaming is intentionally disabled (task wanted the opposite), no CLI→REST mapping doc table, no architecture diagram in plan.md.
  • 005-rest-management-integration: 0 rejected — all 4 UNRESOLVED findings confirmed as clean relocations into internal/management/service.go.
  • 006-oauth-extra-params (7 of 7 remaining): masking, redirect-URI display, last-refresh display, login preview/summary, post-success verification, example config snippets all still missing from auth_cmd.go/doctor_cmd.go.
  • 007-oauth-e2e-testing (4 of 4): masking (same gap as 006), provider-URL+grant_type logging, discovery-endpoint reachability check, auth-status-format tests all absent.
  • 008-oauth-token-refresh (6 of 6): correlation-ID propagation reaches exactly 2 functions (already ticked); CreateOAuthConfig, handleCallback, discovery.go, persistent_token_store.go, handleOAuthAuthorization, managed/client.go have zero correlation-ID references.
  • 009-proactive-oauth-refresh (19 of 21 remaining): expiry badges, distinct "Token Expired"/"Auth Error" badges, confirmation dialog on logout, FormatRelativeTime, logout unit tests, --all flag, 400/404 contract tests all absent (2 of 21 ticked, see above).
  • 011-resource-auto-detect (1 of 1): doctor output never surfaces the auto-detected resource parameter (unlike auth status, correctly ticked elsewhere).
  • 014-cli-output-formatting, 016-activity-log-backend, 017-activity-cli-commands, 021-request-id-logging, 022-oauth-redirect-uri-persistence: same pattern — doctor_cmd.go never migrated off raw fmt.Print*; CSV activity export untested (JSON is); no ULID validation, no summary/show output-format tests, no storage-layer GetActivitySummary, no export path-validation test; the only CLI-flag E2E test for request-id actually only covers the API query param, not the flag; 022 T007 confirmed a correct ticked no-op (repo has no Storage interface by design).
  • 026-pii-detection (5 of 5 remaining; 24 of 29 originally-flagged UNRESOLVED findings confirmed as clean relocations into internal/security/paths.go, entropy.go, activity.go, Activity.vue, activity_cmd.go): custom-pattern E2E wiring, sensitive-data E2E script scenarios, SSE-on-detection test all absent.
  • 028-agent-tokens (1 of 1 remaining): revoke/regenerate behavior itself untested (only command-registration and table-output are); 3 UNRESOLVED findings confirmed clean relocations (mcp_auth_scope_test.go, mcp_activity_agent_test.go, inlined into AgentTokens.vue).
  • 029-mcpproxy-teams (1 of 1): Linux server-edition tarball CI matrix intentionally commented out pending "server MVP ready" (matches CLAUDE.md: Docker-only distribution).
  • 040-server-ux (4 of 4): the whole AddServerView.swift import flow lacks preview/per-server breakdown/Browse-file/extended-timeout, confirmed still true after the macOS-UI code that merged in from main this run (that code is Profiles-v3, a different feature).
  • 042-telemetry-tier2 (10 of 10 remaining): production wiring is real and comprehensive almost everywhere, but the specific named tests (integration-style, not just registry-unit-level) for surface/builtin-tool recording, OAuth-refresh-failure counter, quarantine-blocked counter, doctor-run counter, show-payload command, and DO_NOT_TRACK/CI-autodisable docs are all missing; ErrCatOAuthTokenExpired is dead code (defined, never recorded). 7 UNRESOLVED/RELOCATED findings confirmed clean (startup-outcome, upgrade-funnel, ID-rotation, first-run-notice, surface-classifier middleware all relocated correctly).
  • 044-diagnostics-taxonomy (9 of 9): no go:generate wiring, no RecordFixAttempt, the E2E script never actually starts a server/curls the diagnostics endpoint despite the task's name, link-checker exists but isn't wired into run-all-tests.sh, no diagnostics sub-object test on the v3 telemetry payload, no dedicated docs sections. 5 UNRESOLVED/REMOVED findings (macOS: AppState.swift/TrayPresentation.swift/doctor_cmd.go/doctor_listcodes_cmd.go) confirmed clean relocations; 1 (FixIssuesMenu) confirmed genuinely absent — flagged as worth a look but not formally proposed as an un-tick this run.
  • 044-retention-telemetry-v3 (3 of 3): placeholder doc section headers never added; the 5-field doc update (decision tree, bucketing, opt-out) never written even though the code is real; quickstart walkthrough is an unverifiable manual-process claim. 2 findings confirmed clean relocations (runtime.go, AutoStartService.swift).
  • 047-cpu-hotpath-fix: 0 rejected this run; the 1 untick-candidate finding confirmed a clean relocation (stores/servers.ts).
  • 056-output-schema-validation (1 of 1): E2E curl scenario for output-schema validation never added, despite the feature itself being fully built and unit-tested.
  • 057-in-proxy-profiles (1 of 1): 2 of 3 required doc touches (CLAUDE.md, README.md) missing; only docs/features/profiles.md was done.
  • 058-mcp-2026-upgrade (40 of 40): confirmed again — every candidate is a false positive from the deterministic script matching either unrelated earlier-phase code or a symbol that only exists in spec markdown. Phases 3-8 (protocol-era-aware session handling, -32022, input_required detection, cache hints, trace-context capture, etc.) remain entirely unbuilt.
  • 069-observability-usage-graphs, 070-registry-easy-upstream-add, 088-scanner-trust-ui, 090-tray-glance-v2: all "removed" findings confirmed as deliberate, documented supersessions by the Spec 109 catalog-first/UX redesign (Repositories.vue→CatalogSearch.vue etc.), not regressions.
  • 076-deterministic-tool-scanner, 077-scanner-simplification, 083-discovery-profiler, 096-batched-call-tools, 097-stored-scripts, 098-tools-preflight (8 total): all are "run the full lint/test/e2e suite" process-gate checkboxes — existence of the referenced script/config proves nothing about whether the run was ever actually executed and green; left unticked per the uncertainty rule.
  • 105-agent-scope-hardening (1 of 1): the one candidate is a documented, still-open, unticked follow-up describing a known unfixed security gap (REST replay bypassing the tool gate) — ticking would misrepresent an open vulnerability as resolved.
  • 107-server-edition-sso-hardening: 0 possibly_built; all 11 UNRESOLVED/REMOVED findings confirmed clean — two are deliberate security-motivated deletions (IdP-token-storage leak surface, documented in docs/features/idp-token-storage.md), the rest are clean relocations/consolidations. One item flagged above for a human glance (T101's matrix-test scale).
  • 108-profiles-v3 (14 of 14 possibly_built rejected): this is late-stage (US3-5) work that genuinely hasn't landed — no profile/client fields on session/activity records, no CLI flags, no profile-management service, no REST docs, no binding subcommands, no doctor checks, no Vue/Swift binding-control components mounted. (New macOS Profiles-v3 UI files merged in from main this run — ProfileEditorView.swift, ClientBindingControls.swift, TokenCreateModel.swift, etc. — are a different surface than what these specific unticked tasks name; re-ran the evidence checker after the merge and confirmed 0 new possibly_built candidates appeared for this spec.)
  • 109-ux-navigation-consistency (2 of 2 remaining possibly_built rejected): CLI docs (docs/cli/attention-command.md etc.) referenced by CLAUDE.md's own "Recent Changes" line don't exist on disk — flagging as a stale doc reference worth a maintainer's attention; attention.go doesn't consume Spec-108 clients-service warnings/token-expiry yet.

Un-tick candidates dropped — ticked correctly, artifact just relocated/renamed/consolidated, or a deliberate, documented removal (not a regression):

The dominant pattern across nearly every spec above: per-domain files consolidated into internal/contracts/types.go or internal/management/service.go; tests renamed to a sibling file in the same package; Vue/Swift views renamed or merged during the Spec 109 UX redesign (Dashboard.vue→Home.vue+AttentionList.vue, DashboardView.swift→HomeView.swift). Every one of these was independently confirmed by reading the actual replacement code, not inferred from the script's own medium-confidence guess. Full list of specs with clean relocations this run: 001-oas-endpoint-documentation, 001-update-version-display, 004-management-health-refactor, 005-rest-management-integration, 009-proactive-oauth-refresh, 011-resource-auto-detect, 012-docusaurus-docs-site, 012-unified-health-status, 013-structured-server-state, 016-activity-log-backend, 022-oauth-redirect-uri-persistence, 026-pii-detection, 028-agent-tokens, 029-mcpproxy-teams, 040-server-ux, 042-telemetry-tier2, 044-diagnostics-taxonomy, 044-retention-telemetry-v3, 047-cpu-hotpath-fix, 058-mcp-2026-upgrade, 069-observability-usage-graphs, 070-registry-easy-upstream-add, 088-scanner-trust-ui, 090-tray-glance-v2, 102-schema-deferred, 107-server-edition-sso-hardening, 108-profiles-v3, 109-ux-navigation-consistency.

Notes

  • Verification was performed by 8 parallel code-reading passes (grouped by spec cluster), each independently applying the tick test before anything was ticked, followed by my own direct re-read of the cited file:line for every one of the 5 ticks actually applied.
  • The branch had diverged from main the same way the previous run described (an old, unrelated pre-history commit chain made git rebase origin/main fail catastrophically). Resolved the same way: git merge origin/main (clean, no conflicts) instead of rebase.
  • ROADMAP.md was regenerated (python3 scripts/gen-roadmap.py) in a separate commit after the checkbox commit, per the pre-commit hook requirement — 004-management-health-refactor moved 73/101→75/101, 009-proactive-oauth-refresh 47/87→49/87, 073-activity-size-retention 13/14 (93%, in-flight)→14/14 (100%, shipped). ROADMAP.md and roadmap.yaml were not hand-edited.
  • Cap: 5 of 40 allowed applied ticks used this run.

@cloudflare-workers-and-pages

cloudflare-workers-and-pages Bot commented Sep 21, 2026 •

Copy link
Copy Markdown

Deploying mcpproxy-docs with  Cloudflare Pages  Cloudflare Pages

Latest commit: 17a3e7d
Status: ✅  Deploy successful!
Preview URL: https://19b6a725.mcpproxy-docs.pages.dev
Branch Preview URL: https://claude-spec-gardener.mcpproxy-docs.pages.dev

View logs

@codecov-commenter

Copy link
Copy Markdown

⚠️ Please install the 'codecov app svg image' to ensure uploads and comments are reliably processed by Codecov.

Codecov Report

✅ All modified and coverable lines are covered by tests.

📢 Thoughts on this report? Let us know!

@github-actions

github-actions Bot commented Sep 21, 2026 •

Copy link
Copy Markdown
Contributor

📦 Build Artifacts

Workflow Run: View Run
Branch: claude/spec-gardener

Available Artifacts

  • archive-darwin-amd64 (31 MB)
  • archive-darwin-arm64 (28 MB)
  • archive-linux-amd64 (19 MB)
  • archive-linux-arm64 (17 MB)
  • archive-windows-amd64 (31 MB)
  • archive-windows-arm64 (27 MB)
  • frontend-dist-pr (0 MB)
  • installer-dmg-darwin-amd64 (27 MB)
  • installer-dmg-darwin-arm64 (24 MB)
  • smart-mcp-proxymcpproxy-goH8IXGG.dockerbuild (0 MB)

How to Download

Option 1: GitHub Web UI (easiest)

  1. Go to the workflow run page linked above
  2. Scroll to the bottom "Artifacts" section
  3. Click on the artifact you want to download

Option 2: GitHub CLI

gh run download 37273036475 --repo smart-mcp-proxy/mcpproxy-go

Note: Artifacts expire in 14 days.

Dumbris commented Sep 28, 2026

Copy link
Copy Markdown
Member Author

CI failure on this run: Build Binaries (macos-15, darwin, arm64) — the "Run tests (skip binary E2E tests...)" step failed (exit code 1) after a git merge origin/main fast-forwarded this branch's diff base.

This is not this PR's failure: the PR's only diff is two markdown files (specs/105-agent-scope-hardening/tasks.md checkbox flips + a regenerated ROADMAP.md), which cannot affect Go test behavior on any platform. The equivalent Linux and Windows "Build Binaries" jobs on this same commit passed, and every other recent PR Build run against current main succeeded, so this isn't a pre-existing base-branch break either. I couldn't pull the specific failing test name from the job log (the log-retrieval tool truncates before reaching the earlier failure in a ~7-minute, very verbose go test ./... run), but given the diff shape this is most consistent with a flaky/platform-specific test on the macOS arm64 runner.

Re-running the failed jobs once to confirm.


Generated by Claude Code


Generated by Claude Code

Dumbris commented Sep 28, 2026

Copy link
Copy Markdown
Member Author

Re-run attempt failed: rerun-failed-jobs returned 403 (resource not accessible by this integration) — I don't have permission to re-run CI jobs on this repo. Since the failing test is unrelated to this PR's markdown-only diff, there's no in-scope fix to push. Keeping this PR watched; will re-check on the next scheduled check-in and act on anything that changes (a maintainer re-run, a new push, or a genuine regression surfacing elsewhere).


Generated by Claude Code


Generated by Claude Code

Dumbris commented Sep 29, 2026

Copy link
Copy Markdown
Member Author

CI note: Build Binaries (windows-latest, windows, amd64, ...) failed on head 08f41b2a (run). Judging this a flake, not a regression from this PR:

  • This PR's diff is markdown-only (specs/*/tasks.md, ROADMAP.md — spec-tracking files), which cannot affect Go test/build behavior.
  • The equivalent windows-arm64 job, and both linux-amd64/linux-arm64 jobs, all passed on this exact commit.
  • The retrievable tail of the failing job's log shows a managed/client.go connection state transition to Error for test-headers-server right before the aggregate FAIL, consistent with network/timing flakiness on the Windows runner rather than a code issue; the specific failing test's own output was further back in the log than the log-retrieval tool could reach.

Attempted to re-run the failed job via rerun-failed-jobs and got a 403 ("Resource not accessible by integration") — I don't have permission to trigger a re-run from here. Leaving the PR watched; will re-check on the next scheduled pass.


Generated by Claude Code

Dumbris commented Sep 29, 2026

Copy link
Copy Markdown
Member Author

CI note: Build Binaries (windows-latest, windows, amd64, ...) failed again on head 1970d608 (run) — same job that failed on the previous head (08f41b2a, see earlier comment on this PR).

Judging this a flake, not a regression from this PR, for the same reasons as before:

  • This PR's diff is still markdown-only (specs/*/tasks.md, ROADMAP.md), which cannot affect Go test/build behavior.
  • linux-amd64 passed on this exact commit; other platform jobs were still in progress or got cancelled as a side effect of this job's failure (not independent failures).
  • The retrievable tail of the log again ends in a bare aggregate FAIL right after the web package's tests pass, with no visible --- FAIL: for the specific failing test/package — same signature as the prior occurrence, and again further back in the log than the log-retrieval tool could reach.

This is now two consecutive occurrences of the same job failing across two different merge commits, which suggests this might be a persistent flake in that specific CI leg (windows-amd64) rather than a one-off — worth a maintainer's attention if it keeps recurring, since it's outside what a markdown-only gardener PR can diagnose or fix.

Attempted to re-run the failed job via rerun-failed-jobs again and got the same 403 ("Resource not accessible by integration") — still no permission to trigger a re-run from here. Leaving the PR watched; will re-check on the next scheduled pass.


Generated by Claude Code

# Conflicts:
#	ROADMAP.md
#	specs/108-profiles-v3/tasks.md

Dumbris commented Sep 30, 2026

Copy link
Copy Markdown
Member Author

CI failure on head d6328f76: Build Binaries (macos-15, darwin, arm64) (job 109813161727) — a different job than the previously-reported windows-amd64 flake.

This PR's diff is still markdown-only (ROADMAP.md + 4 specs/*/tasks.md files), so it cannot be the cause. On this same run, windows-amd64 and linux-amd64/linux-arm64 Build Binaries all passed; darwin-amd64 was still in progress at the time of this check.

The retrievable log window shows the same truncation pattern seen before: it ends right after ok github.com/smart-mcp-proxy/mcpproxy-go/web with a bare FAIL and no individual --- FAIL: line visible — the actual failing test is further back in the log than this tooling can retrieve.

Attempted rerun_failed_jobs on run 36692583703 → 403 Resource not accessible by integration (same permission gap as before, no access to trigger re-runs from this integration).

Judging this a flake on the same basis as the earlier windows-amd64 one (unrelated diff, other platforms green, truncated-log signature), but noting it's now a second distinct job exhibiting this pattern (previously windows-amd64, now darwin-arm64) rather than a repeat of the same one — worth a maintainer's attention if it keeps recurring across different jobs.


Generated by Claude Code

@Dumbris
Dumbris marked this pull request as ready for review October 6, 2026 17:23
@Dumbris
Dumbris merged commit 0c0fc00 into main Oct 6, 2026
45 checks passed
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants