Skip to content

feat(scan): structured log parser, cache-aware cost model, multi-client adapters - #22

Merged
ousamabenyounes merged 4 commits into
mainfrom
feat/scan-structured-engine
Sep 13, 2026
Merged

ousamabenyounes merged 4 commits into
mainfrom
feat/scan-structured-engine

Conversation

@ousamabenyounes

Copy link
Copy Markdown
Collaborator

Why

tokenwar scan counted log lines matching regexes like /\b(git|grep|find|cat)\b/i. A JSONL line is a whole message object carrying several tool_use blocks, so those counts were neither tool calls nor tokens. Prose mentions scored as executions, every call was double counted via its result, find matched findViewById. The "estimated avoidable tokens" were then derived from invented constants (900 tokens per shell hit, a flat 12% ratio), producing false precision.

It was also Claude-only in practice: Codex logs were found but none parsed.

What changed

Structured parsing. Iterate message.content[], deduplicate on tool_use.id, separate sidechain traffic, and recover argv[0] by tokenizing the command — unwrapping sudo, env, timeout and leading VAR=value assignments.

Corrected cost model. A static skill listing sits in the prompt prefix, so after turn 1 it is served from cache at 0.1x input, not 1.0x. Multiplying the listing by turn count overstates cost ~8.4x; on the logs used to develop this, caching already absorbed ~88% of the theoretical waste. The report now shows both figures and names the gap, because the pre-cache number is one a user's billing page disproves in a minute.

The costs that survive caching are reported instead:

  • Window occupancy — caching discounts price, not space.
  • Prefix invalidation — a 12.5x step on every inventory change. This is why tokenwar bundle is session-start only.

Multi-client adapters. Codex sessions live under sessions/YYYY/MM/DD/rollout-*.jsonl with a different schema; its token_count events are cumulative, so a turn is the delta between snapshots, and input_tokens already includes the cached portion (subtracted to avoid double counting). 104 Codex sessions now parse, up from 0. The report states which clients were read, separating "no sessions" from "format unsupported" — an unparsed client must not look like a frugal one.

Inventory fixes. Plugin skills were counted once per vendored copy (.junie/, .codex/), and folded YAML description: > blocks were scored at a few tokens. Both fixed.

Pricing fix. Model selection took the first session naming one — a short Haiku subagent — pricing an Opus workload ~15x too low.

New commands

  • tokenwar prune — capabilities that load every request but were never invoked. Deletes nothing; infrequent use is not disuse.
  • tokenwar bundle <dev|devops|architect|testing> — session-start bundles, with --dry-run.
  • tokenwar scan --html — graded HTML report.

Honesty constraints enforced in code

  • Savings are never summed across overlapping lanes (RTK/context-mode/caveman).
  • Every recommendation carries a signal, a counter-signal, its own cost and a break-even rule. graphify and openwiki can report NOT YET; pxpipe reports AVOID.
  • Profile inference returns a distribution with confidence, never a single label, and declares SEO/product/design not inferrable from coding-agent logs.

Tests

29 new (tests/parse.bats 10, tests/scan.bats 19), all passing, covering the tokenizer, the cache model, the Codex adapter and client coverage reporting. dispatcher.bats and toggle.bats still pass.

check.bats test 7 fails on this machine because the graphify CLI is absent — verified pre-existing via git stash.

Credit

Report shape inspired by gmetais/YellowLabTools.

🤖 Generated with Claude Code

https://claude.ai/code/session_01A7T2MNqFo5hfDvF2ovoTig

ousamabenyounes and others added 4 commits September 6, 2026 17:54
The scanner counted log LINES matching regexes like /\b(git|grep|find)\b/i.
A JSONL line is a whole message object carrying several tool_use blocks, so
those counts were neither tool calls nor tokens: prose mentions scored as
executions, every call was double counted via its result, `find` matched
findViewById, and reading only the last 512KB sampled session ends. The
"estimated avoidable tokens" were then derived from invented constants
(900 tokens per shell hit, a flat 12% ratio), producing false precision.

Replace it with a parser that iterates message.content[], deduplicates on
tool_use.id, separates sidechain traffic, and recovers argv[0] by tokenizing
the command (unwrapping sudo/env/timeout and VAR=value assignments).

Correct the cost model. A static skill listing sits in the prompt prefix, so
after turn 1 it is served from cache at 0.1x input, not 1.0x. Multiplying the
listing size by the turn count overstates cost 8-9x; on the logs used to
develop this, caching already absorbed ~88% of the theoretical waste. The
report now shows both figures and names the gap, because the pre-cache number
is one a user's billing page disproves in a minute. The costs that survive
caching are stated instead: context-window occupancy, which caching does not
discount, and prefix invalidation, a 12.5x step on every inventory change.

Profile inference reports a distribution over modes with confidence, never a
single label, and declares SEO, product and design not inferrable from
coding-agent logs rather than guessing them.

Recommendations each carry a signal, a counter-signal, their own cost and a
break-even rule; graphify's build cost and openwiki's ~50:1 adverse
maintenance ratio mean both can be reported as NOT YET, and pxpipe as AVOID.
Overlapping lanes are flagged so savings are never summed.

Add `tokenwar prune` (review list, deletes nothing) and `tokenwar bundle`
(session-start only, since mid-session switching invalidates the prefix),
an HTML report, and 25 tests covering the tokenizer and the cost model.

Report shape inspired by gmetais/YellowLabTools.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01A7T2MNqFo5hfDvF2ovoTig
Two inventory bugs, both found by reading the output rather than the code.

Plugin skills were counted once per vendored copy: a plugin that ships its
skills for several agents (.junie/, .codex/, .cursor/) had each copy scored,
so cavecrew appeared eight times. Only the copy the client loads occupies
prompt space, so skip vendored paths and keep one entry per name.

Folded YAML descriptions were scored at a few tokens. `description: >` puts
the text on following indented lines, and the frontmatter reader matched only
to end of line, so those skills looked nearly free. Read block scalars to the
next top-level key. The two fixes move the count from a wrong 67-of-69 to 59
of 61, and the listing cost from 4.0K to 5.0K tokens.

The scan was also Claude-only in practice. Codex logs were found (109 files)
but none parsed: its sessions live under sessions/YYYY/MM/DD/rollout-*.jsonl,
below the old depth limit, and use a different schema. Add an adapter layer.
Codex reports cumulative totals in periodic token_count events, so a turn is
the delta between snapshots, and its input_tokens already includes the cached
portion, which is subtracted so tokens are not counted twice. 104 Codex
sessions now parse.

Report which clients were actually read, separating "no sessions" from
"format unsupported" — an unparsed client must not look like a frugal one.
Gemini's logs carry no usage fields, so it contributes tool evidence and is
excluded from cost. Copilot and opencode are detected pending real logs.

Price the skill listing against Claude turns only. Skills are a Claude Code
capability, and charging them to Codex turns attributed cost to sessions that
never carried them.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01A7T2MNqFo5hfDvF2ovoTig
Model selection took the first session that named one, which on this machine
was a short Haiku subagent transcript. An Opus workload was therefore priced
at Haiku rates, understating every dollar figure roughly fifteenfold. Pick the
model with the largest token volume instead, and keep --model as an override.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01A7T2MNqFo5hfDvF2ovoTig
Two scan tests read the developer's real ~/.claude/skills, so they passed
locally and failed on CI, where nothing is installed: with no skills the dead
listing is zero tokens and both the cached and uncached costs collapse to $0,
so the assertion that one is below the other could not hold.

Point the skills directory, plugin cache and MCP config at fixtures via
TOKENWAR_SKILLS_DIR, TOKENWAR_PLUGIN_CACHE_DIR and TOKENWAR_MCP_CONFIG. The
fixture includes a folded `description: >` block so the block-scalar parse
stays covered.

The scan itself was correct; only the tests were reading host state.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01A7T2MNqFo5hfDvF2ovoTig
@ousamabenyounes
ousamabenyounes merged commit d3e681a into main Sep 13, 2026
1 check passed
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant