Skip to content

feat(copilot): GitHub Copilot CLI as a provider, and wire the stack into it - #21

Merged
ousamabenyounes merged 1 commit into
mainfrom
feat/copilot-provider
Sep 4, 2026
Merged

ousamabenyounes merged 1 commit into
mainfrom
feat/copilot-provider

Conversation

@ousamabenyounes

Copy link
Copy Markdown
Collaborator

Why

tokenwar tracked five providers; GitHub Copilot CLI was not one of them, and none of the seven tools reached it.

Two separate gaps, both closed here:

  1. Copilot as a provider. It ships real, local token telemetry — ~/.copilot/session-store.db → assistant_usage_events — with a per-call token breakdown and total_nano_aiu, the AI-credit cost GitHub actually bills. That is the same class of native source tokenwar already reads for Codex and opencode, so there was no reason to leave it as an N/A.
  2. The stack inside Copilot. Being a tracked provider only gets you the numbers. The tools are published for Claude Code and reach Copilot only if you point them at Copilot's own extension points, of which there are exactly three: hooks (~/.copilot/hooks/*.json), skills (~/.copilot/skills/<name>/SKILL.md), MCP (~/.copilot/mcp-config.json).

Provider

Registry PROVIDER_IDX_COPILOT=5, PROVIDER_COUNT=6 in scripts/lib/providers.sh
Telemetry ~/.copilot/session-store.db → assistant_usage_events (tokens + total_nano_aiu), total and monthly
Config dir $COPILOT_HOME (Copilot's own override), default ~/.copilot
Wrapper copilot joins codex/gemini/kimi/opencode in WRAPPED_PROVIDER_CLIS
Scan copilot client, root ~/.copilot, override TOKENWAR_COPILOT_LOG_ROOT
  provider        tokens      note
  ─────────────────────────────────────────────────────────────
  Codex           8713.9M     1300 Codex sessions (real tokens_used)
  opencode        105.7K      11 opencode sessions (real token cols)
  Copilot CLI     13.3K       1 Copilot sessions (real assistant_usage_events) - 0.24 AI credits billed

0.24 AI credits matches what the CLI itself prints in its exit summary for that session — same number, read from its own store.

Three details that needed care:

  • Version parsing. Copilot prints GitHub Copilot CLI 1.0.83. — trailing full stop, then a second line about updates. The shared parser yielded the literal 1.0.83., which can never compare equal to a registry version. provider_version now strips trailing punctuation (harmless for every other provider, tested against the real two-line shape).
  • Copilot is not billed per token. It is a seat subscription plus AI credits, and GitHub publishes no per-token list price. The $ column is therefore an API-equivalent valuation at a GPT-5-class input rate — never an invoice — and the unit that is billed appears in the telemetry note. That constraint is written into the registry comment and into SKILL.md so nobody later "fixes" it into a fake bill.
  • No update check. Copilot self-updates (autoUpdate defaults true). tokenwar reports the version it finds rather than duplicating and racing that mechanism.

The stack inside Copilot — tokenwar copilot

New scripts/copilot.sh, wired into the dispatcher as tokenwar copilot [check|wire] and into install.sh as --with-copilot (folded into --all).

Tool Via Wiring
rtk hook rtk init -g --copilot → PreToolUse hook + user-level instructions
graphify skill graphify copilot install (its own native command)
caveman skill copilot skill add on the plugin's SKILL.md
ponytail skill copilot skill add on the plugin's SKILL.md
claude-mem MCP its own .mcp.json definition, re-registered via copilot mcp add
context-mode — n/a: its plugin manifest pins an absolute, version-specific interpreter path — the registration would break on the next upgrade
pxpipe — n/a: a proxy on the Anthropic-compatible API path; Copilot talks to GitHub's endpoint
# /tokenwar copilot

  ·  tool          via       state             note
  ─────────────────────────────────────────────────────────────────
  ✓  rtk           hook      wired             ~/.copilot/hooks/rtk-rewrite.json
  ✓  graphify      skill     wired             ~/.copilot/skills/graphify
  ✓  caveman       skill     wired             ~/.copilot/skills/caveman
  ✓  ponytail      skill     wired             ~/.copilot/skills/ponytail
  ✓  claude-mem    MCP       wired             ~/.copilot/mcp-config.json → claude-mem

Two decisions worth reviewing

1. claude-mem is registered from its own .mcp.json, never a hardcoded path. That published definition wraps a locator resolving the current plugin version at runtime, so the Copilot registration survives claude plugin update. Pointing Copilot at .../claude-mem/13.6.1/scripts/mcp-server.cjs works exactly until the next upgrade. A test asserts the real args (including the inline locator script) are what reach copilot mcp add.

2. The raised timeouts are load-bearing, not padding. claude-mem's MCP server calls a local worker over HTTP and aborts at CLAUDE_MEM_API_TIMEOUT_MS, 30s by default. The first search after a cold worker path builds an index over the whole memory DB. Measured here:

✗ search (MCP: claude-mem) · query: "tokenwar"                     30s
  └ MCP server 'claude-mem': Error calling Worker API: Request timed out after 30000ms

Reproduced outside Copilot too (env -i + a raw stdio JSON-RPC probe), so it is claude-mem's cold path, not a Copilot bug — and the second call always succeeds. With the defaults, the first claude-mem call of a Copilot session always fails, which reads as "claude-mem is broken under Copilot" when it is not. The wiring raises claude-mem's own variable and Copilot's per-tool timeout, after which:

● search (MCP: claude-mem) · query: "tokenwar"                     2m 2s
  └ Found 120 result(s) matching "tokenwar" (50 obs, 50 sessions, 20 prompts)

Live verification (real Copilot CLI 1.0.83, this machine)

Not just unit tests — every wired tool was exercised inside a real copilot -p session:

Tool Evidence
rtk asked for git status; the agent executed rtk git status
graphify ● skill(graphify) listed by copilot skill list; graphify god-nodes --top 3 ran and returned the graph's hubs
caveman ● skill(caveman) loaded, reported its default intensity full, answered in caveman style
ponytail ● skill(ponytail) loaded, reported default full and the first rung of its YAGNI ladder
claude-mem search over MCP returned 120 results

Test verification (RED → GREEN)

RED — upstream main scripts (implementation stashed, scripts/copilot.sh moved aside), new tests applied:

not ok  1-13  tests/copilot.bats — all 13 unrunnable: copilot.sh does not exist
not ok  24    status.sh detects Copilot CLI and strips its trailing full stop
not ok  25    status.sh names the Copilot telemetry source
not ok  26    gain.sh shows Copilot N/A when its session store is absent
not ok  27    gain.sh reads REAL Copilot token telemetry and AI credits from session-store.db
not ok  59    --with-copilot delegates to scripts/copilot.sh rather than re-implementing it
not ok  60    --with-copilot warns and skips when the Copilot CLI is absent
not ok  69-71 scan detects / honours TOKENWAR_COPILOT_LOG_ROOT / --json exposes copilot
22 not ok / 49 ok

and, separately, for the launch filter under a real pty:

not ok 10 copilot --acp (Agent Client Protocol server) → no banner on a tty
not ok 11 copilot management subcommands → no banner on a tty
not ok 12 copilot --version → no banner on a tty

Those three are driven through script -qfec on purpose. Without a pty [[ -t 1 ]] is false and the banner is suppressed whatever the filter says — the first version of these tests passed against a launcher that knew nothing about Copilot, which made them worthless. They now fail on main and pass here.

GREEN — with the implementation: see the run below.

CI

  • tokenwar-launch.sh copilot and copilot.sh check added to the smoke step.
  • The JSON contract smoke now also asserts every provider is present in status.sh --json and gain.sh --json. Providers are registry-driven, so one added to lib/providers.sh but forgotten in a note map or a telemetry switch used to surface only as a blank column at runtime.

Known drift, not addressed here

index.html (the landing page) still says five tools — it was already two behind before the graphify PR. Bringing it to seven tools + six providers is a self-contained follow-up.

🤖 Generated with Claude Code

https://claude.ai/code/session_01NXot81hzCH5ViVyXJDwX4W

…nto it

Two gaps, both closed here.

1. Copilot as a tracked provider. It ships real local telemetry —
   ~/.copilot/session-store.db -> assistant_usage_events — with a
   per-call token breakdown AND total_nano_aiu, the AI-credit cost
   GitHub actually bills. Same class of native source tokenwar already
   reads for Codex and opencode, so there was no reason to leave it N/A.

2. The stack INSIDE Copilot. Being a tracked provider only gets you the
   numbers. The tools are published for Claude Code and reach Copilot
   only if pointed at its own extension points, of which there are three:
   hooks (~/.copilot/hooks/*.json), skills (~/.copilot/skills/<n>/SKILL.md)
   and MCP (~/.copilot/mcp-config.json).

Provider side:
  lib/providers.sh  PROVIDER_IDX_COPILOT=5, COUNT=6, telemetry total +
                    monthly, config dir via Copilot's own COPILOT_HOME.
  provider_version  strips trailing punctuation: Copilot prints
                    "GitHub Copilot CLI 1.0.83." and the literal "1.0.83."
                    can never compare equal to a registry version.
  pricing           Copilot is NOT billed per token — seat + AI credits,
                    no published per-token price. The $ column is an
                    API-equivalent valuation, never an invoice; the unit
                    that IS billed is read from total_nano_aiu into the
                    note. That constraint is written down so nobody later
                    "fixes" it into a fake bill.
  check-updates     no update check: Copilot self-updates (autoUpdate
                    defaults true) and racing it would only report noise.
  launch/install    `copilot` joins the wrapped CLIs; its non-interactive
                    subcommands and --acp/--version are banner-silent.
  scan              `copilot` client on ~/.copilot.

New scripts/copilot.sh — `tokenwar copilot [check|wire]`, also reachable
as install.sh --with-copilot (folded into --all, delegating so there is
one implementation, not two that drift):

  rtk        -> hook   rtk init -g --copilot
  graphify   -> skill  graphify copilot install
  caveman    -> skill  copilot skill add <plugin cache SKILL.md>
  ponytail   -> skill  copilot skill add <plugin cache SKILL.md>
  claude-mem -> MCP    its OWN .mcp.json definition, re-registered

  context-mode and pxpipe are reported n/a WITH a reason and left alone:
  context-mode's manifest pins an absolute version-specific interpreter
  path, and pxpipe proxies the Anthropic-compatible API path that Copilot
  does not use.

Two decisions worth reading:

  claude-mem is registered from its published .mcp.json, never from a
  hardcoded path. That definition wraps a locator resolving the current
  plugin version at runtime, so the registration survives
  `claude plugin update`; a path to .../claude-mem/13.6.1/... works right
  up until the next upgrade.

  The raised timeouts are load-bearing. claude-mem's MCP server calls a
  local worker and aborts at CLAUDE_MEM_API_TIMEOUT_MS (30s default). The
  first search after a cold worker path indexes the whole memory DB —
  measured 2m02s. Reproduced outside Copilot with `env -i` and a raw
  stdio JSON-RPC probe, so it is claude-mem's cold path, not a Copilot
  bug. With the defaults the FIRST call of every Copilot session fails
  and reads as a broken integration.

Live verification on real Copilot CLI 1.0.83 (copilot -p sessions):
  rtk         agent executed `rtk git status`, not `git status`
  graphify    skill listed; `graphify god-nodes --top 3` returned hubs
  caveman     skill(caveman) loaded, default intensity `full`
  ponytail    skill(ponytail) loaded, first YAGNI rung correct
  claude-mem  MCP `search` returned 120 results

Test verification (RED -> GREEN)

RED — upstream main scripts, copilot.sh moved aside, new tests applied:
  not ok 1-13   tests/copilot.bats, all unrunnable: copilot.sh absent
  not ok 24-27  Copilot detection, version parse, telemetry, AI credits
  not ok 59-60  install --with-copilot delegation + absent-CLI warning
  not ok 69-71  scan client, log-root override, --json contract
  22 not ok / 49 ok
and under a real pty for the launch filter:
  not ok 10 copilot --acp -> no banner on a tty
  not ok 11 copilot management subcommands -> no banner on a tty
  not ok 12 copilot --version -> no banner on a tty

Those three run through `script -qfec` deliberately: without a pty
[[ -t 1 ]] is false and the banner is suppressed whatever the filter
says, so the first version of these tests passed against a launcher that
knew nothing about Copilot.

GREEN — with the implementation:
  171/171 bats (was 144), and 171/171 again under a CI-like PATH with
  copilot, graphify, rtk and claude all absent.
  shellcheck -S warning scripts/*.sh scripts/lib/*.sh install.sh
  uninstall.sh -> clean

CI also gains a provider contract assertion: providers are
registry-driven, so one added to lib/providers.sh but forgotten in a note
map or a telemetry switch used to surface only as a blank runtime column.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01NXot81hzCH5ViVyXJDwX4W
@ousamabenyounes
ousamabenyounes merged commit f410d2c into main Sep 4, 2026
1 check passed
@ousamabenyounes
ousamabenyounes deleted the feat/copilot-provider branch September 4, 2026 19:06
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant