fix(mothership): close abandoned tool meters once instead of alarming every tick - #8479
Conversation
… every tick A tool meter row (cost unknown) stays open when the process that owned the tool ends mid-execution, and nothing ever closed it. The replay tick counted those rows and logged "Service usage requires reconciliation" at ERROR on every tick in every process, forever. Its 5-minute threshold also flagged tools that were still legitimately running. The replay tick now closes meters older than twice the longest tool watchdog, keeping a pricing failure's error or recording that the tool never finished, and logs each closed meter once with its stream, tool call and reason. Known spend is unaffected: it is saved and delivered as separate receipts.
|
The latest updates on your projects. Learn more about Vercel for GitHub. |
|
@cubic-dev-ai review this PR |
@waleedlatif1 I have started the AI code review. It will take a few minutes to complete. |
|
There was a problem hiding this comment.
All reported issues were addressed across 3 files
Reply with feedback, questions, or to request a fix.
Fix all with cubic | Re-trigger cubic
|
@cubic-dev-ai review this PR |
@waleedlatif1 I have started the AI code review. It will take a few minutes to complete. |
Summary
Problem. The service-usage replay tick logged
Service usage requires reconciliationat ERROR on every tick, in every process, indefinitely. The alarm never cleared and carried no information about which meter was stuck.Root cause. A tool meter row (
_tool_execution,cost_usd IS NULL) is opened when a tool starts and closed byfinishServiceUsagewhen it ends. When the process that owns the tool dies mid-execution, nothing ever closes that row.serviceMeteringHealthcounted every such row older than 5 minutes, so a single abandoned meter produced a permanent, repeating error. The 5-minute threshold also flagged tools that were still legitimately running under the long-running watchdog.Fix. Replace the health count with
closeAbandonedServiceMeters: the replay tick closes open meters older than twice the longest tool watchdog (TOOL_WATCHDOG_LONG_RUNNING_MS), claimed withFOR UPDATE SKIP LOCKEDso concurrent ticks close each meter exactly once. A pricing failure keeps its recorded error; otherwise the row recordsTool execution never finished. Each closed meter is logged once, at WARN, with its stream, tool call, creation time and reason.Behaviour changes
Service usage requires reconciliationERROR is gone. Each abandoned meter now yields one WARN line naming its stream and tool call, then never again.delivered_atset). Meters younger than that are left alone, so in-flight tools are no longer flagged.Test plan
service-store.integration.ts(real PostgreSQL 17 + pgvector, migrated withdb:migrate): a new case covers an abandoned meter, an abandoned meter with a pricing error, a meter just past the watchdog that must stay open, and a priced receipt that must stay deliverable. Two concurrent closes return each abandoned meter once, a third returns nothing, and the receipt is still claimable.vitest run lib/mothership/billing(unit) passes.bun run type-check(apps/sim) and Biome pass.