feat(scan): structured log parser, cache-aware cost model, multi-client adapters - #22
Merged
Merged
Conversation
The scanner counted log LINES matching regexes like /\b(git|grep|find)\b/i. A JSONL line is a whole message object carrying several tool_use blocks, so those counts were neither tool calls nor tokens: prose mentions scored as executions, every call was double counted via its result, `find` matched findViewById, and reading only the last 512KB sampled session ends. The "estimated avoidable tokens" were then derived from invented constants (900 tokens per shell hit, a flat 12% ratio), producing false precision. Replace it with a parser that iterates message.content[], deduplicates on tool_use.id, separates sidechain traffic, and recovers argv[0] by tokenizing the command (unwrapping sudo/env/timeout and VAR=value assignments). Correct the cost model. A static skill listing sits in the prompt prefix, so after turn 1 it is served from cache at 0.1x input, not 1.0x. Multiplying the listing size by the turn count overstates cost 8-9x; on the logs used to develop this, caching already absorbed ~88% of the theoretical waste. The report now shows both figures and names the gap, because the pre-cache number is one a user's billing page disproves in a minute. The costs that survive caching are stated instead: context-window occupancy, which caching does not discount, and prefix invalidation, a 12.5x step on every inventory change. Profile inference reports a distribution over modes with confidence, never a single label, and declares SEO, product and design not inferrable from coding-agent logs rather than guessing them. Recommendations each carry a signal, a counter-signal, their own cost and a break-even rule; graphify's build cost and openwiki's ~50:1 adverse maintenance ratio mean both can be reported as NOT YET, and pxpipe as AVOID. Overlapping lanes are flagged so savings are never summed. Add `tokenwar prune` (review list, deletes nothing) and `tokenwar bundle` (session-start only, since mid-session switching invalidates the prefix), an HTML report, and 25 tests covering the tokenizer and the cost model. Report shape inspired by gmetais/YellowLabTools. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01A7T2MNqFo5hfDvF2ovoTig
Two inventory bugs, both found by reading the output rather than the code. Plugin skills were counted once per vendored copy: a plugin that ships its skills for several agents (.junie/, .codex/, .cursor/) had each copy scored, so cavecrew appeared eight times. Only the copy the client loads occupies prompt space, so skip vendored paths and keep one entry per name. Folded YAML descriptions were scored at a few tokens. `description: >` puts the text on following indented lines, and the frontmatter reader matched only to end of line, so those skills looked nearly free. Read block scalars to the next top-level key. The two fixes move the count from a wrong 67-of-69 to 59 of 61, and the listing cost from 4.0K to 5.0K tokens. The scan was also Claude-only in practice. Codex logs were found (109 files) but none parsed: its sessions live under sessions/YYYY/MM/DD/rollout-*.jsonl, below the old depth limit, and use a different schema. Add an adapter layer. Codex reports cumulative totals in periodic token_count events, so a turn is the delta between snapshots, and its input_tokens already includes the cached portion, which is subtracted so tokens are not counted twice. 104 Codex sessions now parse. Report which clients were actually read, separating "no sessions" from "format unsupported" — an unparsed client must not look like a frugal one. Gemini's logs carry no usage fields, so it contributes tool evidence and is excluded from cost. Copilot and opencode are detected pending real logs. Price the skill listing against Claude turns only. Skills are a Claude Code capability, and charging them to Codex turns attributed cost to sessions that never carried them. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01A7T2MNqFo5hfDvF2ovoTig
Model selection took the first session that named one, which on this machine was a short Haiku subagent transcript. An Opus workload was therefore priced at Haiku rates, understating every dollar figure roughly fifteenfold. Pick the model with the largest token volume instead, and keep --model as an override. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01A7T2MNqFo5hfDvF2ovoTig
Two scan tests read the developer's real ~/.claude/skills, so they passed locally and failed on CI, where nothing is installed: with no skills the dead listing is zero tokens and both the cached and uncached costs collapse to $0, so the assertion that one is below the other could not hold. Point the skills directory, plugin cache and MCP config at fixtures via TOKENWAR_SKILLS_DIR, TOKENWAR_PLUGIN_CACHE_DIR and TOKENWAR_MCP_CONFIG. The fixture includes a folded `description: >` block so the block-scalar parse stays covered. The scan itself was correct; only the tests were reading host state. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01A7T2MNqFo5hfDvF2ovoTig
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Why
tokenwar scancounted log lines matching regexes like/\b(git|grep|find|cat)\b/i. A JSONL line is a whole message object carrying severaltool_useblocks, so those counts were neither tool calls nor tokens. Prose mentions scored as executions, every call was double counted via its result,findmatchedfindViewById. The "estimated avoidable tokens" were then derived from invented constants (900 tokens per shell hit, a flat 12% ratio), producing false precision.It was also Claude-only in practice: Codex logs were found but none parsed.
What changed
Structured parsing. Iterate
message.content[], deduplicate ontool_use.id, separate sidechain traffic, and recoverargv[0]by tokenizing the command — unwrappingsudo,env,timeoutand leadingVAR=valueassignments.Corrected cost model. A static skill listing sits in the prompt prefix, so after turn 1 it is served from cache at 0.1x input, not 1.0x. Multiplying the listing by turn count overstates cost ~8.4x; on the logs used to develop this, caching already absorbed ~88% of the theoretical waste. The report now shows both figures and names the gap, because the pre-cache number is one a user's billing page disproves in a minute.
The costs that survive caching are reported instead:
tokenwar bundleis session-start only.Multi-client adapters. Codex sessions live under
sessions/YYYY/MM/DD/rollout-*.jsonlwith a different schema; itstoken_countevents are cumulative, so a turn is the delta between snapshots, andinput_tokensalready includes the cached portion (subtracted to avoid double counting). 104 Codex sessions now parse, up from 0. The report states which clients were read, separating "no sessions" from "format unsupported" — an unparsed client must not look like a frugal one.Inventory fixes. Plugin skills were counted once per vendored copy (
.junie/,.codex/), and folded YAMLdescription: >blocks were scored at a few tokens. Both fixed.Pricing fix. Model selection took the first session naming one — a short Haiku subagent — pricing an Opus workload ~15x too low.
New commands
tokenwar prune— capabilities that load every request but were never invoked. Deletes nothing; infrequent use is not disuse.tokenwar bundle <dev|devops|architect|testing>— session-start bundles, with--dry-run.tokenwar scan --html— graded HTML report.Honesty constraints enforced in code
NOT YET; pxpipe reportsAVOID.Tests
29 new (
tests/parse.bats10,tests/scan.bats19), all passing, covering the tokenizer, the cache model, the Codex adapter and client coverage reporting.dispatcher.batsandtoggle.batsstill pass.check.batstest 7 fails on this machine because the graphify CLI is absent — verified pre-existing viagit stash.Credit
Report shape inspired by gmetais/YellowLabTools.
🤖 Generated with Claude Code
https://claude.ai/code/session_01A7T2MNqFo5hfDvF2ovoTig