Skip to content

Let a bill's verdict keep up with the bill - #91

Open
mikaalnaik wants to merge 1 commit into
mainfrom
mikaal/builder-mp-improvements
Open

mikaalnaik wants to merge 1 commit into
mainfrom
mikaal/builder-mp-improvements

Conversation

@mikaalnaik

Copy link
Copy Markdown
Contributor

Builder MP publishes a pro-growth verdict on every bill before the 45th Parliament. Three things were wrong with how those verdicts were produced and shown, and they compounded.

What was broken

Analyses were write-once. fromCivicsProjectApiBill was the only caller of summarizeBillText, and [id]/page.tsx only reached it when a bill was absent from Mongo. So existingBill inside that function was always null, and the sourceChanged / countChanged regeneration guard could never fire. A bill analyzed at first reading kept that verdict through every committee amendment. Only a manual /reprocess updated it — there was no cron, no queue, no worker.

The list and the detail page disagreed. The list merged fresh Civics API data over the stored document; the detail page rendered the stored document alone, whose status and stages froze at ingestion. A bill could read "Royal Assent" in the list and "Introduced" on its own page.

The prompt and the parser disagreed. steel_man, needs_more_info and missing_details were read off the response but never requested, so they were empty on every bill in the database. The output-format block showed the model invalid JSON (unescaped nested quotes, a trailing comma) while the parser did a bare JSON.parse; anything fenced or malformed fell to a regex that breaks on apostrophes and yielded a permanent abstain. And is_social_issue was requested and discarded, while a second gpt-5 call answered the same question from the first 8000 characters, swallowed its errors to false, and was never allowed to change the verdict — the drift migrations/1.ts existed to repair by hand.

What this changes

Freshness — the sweep

services/refresh.ts is the job that didn't exist. Two passes of deliberately unequal cost:

  • Facts (status, stages, sponsor, genres) refresh for every stored bill on every sweep. No LLM, so it's cheap, and it's what stops a current verdict sitting next to a months-old status.
  • Analysis re-runs only for bills that are new or whose text actually changed, up to analysisBudget (default 10) per sweep. The decision lives in services/refresh-decision.ts and is unit tested.

It takes a Mongo lease (models/JobLock.ts) before doing anything, so running more than one replica is safe rather than expensive. src/instrumentation.ts schedules it in-process (off unless BILLS_REFRESH_ENABLED=true); POST /bills/api/refresh behind BILLS_CRON_SECRET is the manual kick and the seam for moving the schedule out later.

No page view can trigger an OpenAI call any more — fromCivicsProjectApiBill is now a pure converter. A bill the sweep hasn't reached shows its real facts with "Analysis pending".

Trust — one set of facts

utils/merge-bill.ts holds the precedence rule once (API wins on facts, database wins on the verdict) and both pages call it. Bills carry analysisGeneratedAt and analysisSourceRef, surfaced to readers:

Assessed 12 March 2026 from the bill text published at that date. Later amendments may not be reflected.

The steel man — on the model and the admin form since the start, but never generated and never rendered — is now both, as "The other side".

Quality — one call, a shape it can't get wrong

Structured Outputs against prompt/analysis-schema.ts. That retires the invalid-JSON example, the enum re-casing, the regex parse fallback and ~65 lines of prompt. social-issue-grader.ts is deleted; the main call answers is_social_issue, and fromRawAnalysis forces abstain when it's true. Tenet titles now come from TENETS by id, ending a three-way drift between TENETS, FALLBACK_TENET_TITLES, and an inline copy in the converter.

Also fixed in passing: page.tsx read process.env.MONGO_URI directly and missed the MONGODB_URI fallback env.ts provides (a deploy setting only MONGODB_URI renders an empty list while detail pages work); CIVICS_PROJECT_BASE_URL was documented as configurable but always ignored its env var; the reprocess route now shares the one write path; shipped typos ("crticial" ×3, "tenents" ×2).

Verified

  • Live sweep against the real Civics API and Mongo: 185 bills scanned, 185 facts written, 0 queued for re-analysis, 0 errors. Two concurrent sweeps — exactly one ran, the other reported locked.
  • 18 unit tests (pnpm test:bills) covering merge precedence, the re-analysis decision, and schema validity under strict mode.
  • tsc clean; eslint 0 errors (3 warnings, all pre-existing). /bills and /bills/[id] serve 200 in ~250ms.

Two problems the live run caught:

  1. Every legacy row looked unanalyzed (no analysisGeneratedAt), which would have re-paid for all 185 analyses. analysisReason now keys off whether a verdict exists; deferred went 185 → 0.
  2. minItems/maxItems are rejected by OpenAI strict mode — leaving them in would have 400'd every request and silently degraded every bill to a fallback abstain. There's a test guarding the banned-keyword list.

Not verified

The live model call. The local OPENAI_API_KEY is rejected (invalid_api_key), so I could not run a real eval. pnpm eval:bills --refresh needs a run with a working key before this deploys. The zero-token --fallback run passes 110/120; the 10 failures are all qp-count on the degraded path and pre-date this work.

Deploying

New env vars for Dokploy — all optional, all off by default:

Var Effect
BILLS_REFRESH_ENABLED true starts the scheduled sweep
BILLS_REFRESH_INTERVAL_MINUTES Defaults to 60
BILLS_CRON_SECRET Bearer token for POST /bills/api/refresh; route returns 503 without it

Existing rows show no provenance line or steel man until the sweep re-analyzes them — the components render nothing rather than something false, since we genuinely don't know when those verdicts were reached.

🤖 Generated with Claude Code

Three things were wrong, and they compounded.

Analyses were write-once. fromCivicsProjectApiBill was the only caller of
summarizeBillText, and [id]/page.tsx only reached it when a bill was absent
from Mongo — so existingBill inside it was always null and the
sourceChanged/countChanged guard could never fire. A bill analyzed at first
reading kept that verdict through every amendment.

The list and the detail page disagreed. The list merged fresh Civics data over
the stored document; the detail page rendered the stored document alone, whose
status and stages froze at ingestion. A bill could read "Royal Assent" in the
list and "Introduced" on its own page.

The prompt and the parser disagreed. steel_man, needs_more_info and
missing_details were read off the response but never asked for, so they were
empty on every bill in the database. The output-format block showed the model
invalid JSON while the parser did a bare JSON.parse. And is_social_issue was
requested and discarded while a second gpt-5 call answered the same question
from the first 8000 characters, swallowed its errors to false, and was never
allowed to change the verdict — the drift migrations/1.ts existed to repair.

Now: a scheduled sweep (services/refresh.ts) refreshes every bill's facts each
run with no LLM, and re-analyzes only bills that are new or whose text changed,
against a per-sweep budget. It takes a Mongo lease so a second replica is safe.
src/instrumentation.ts schedules it in-process; POST /bills/api/refresh is the
manual kick. No page view can trigger an OpenAI call any more —
fromCivicsProjectApiBill is a pure converter.

utils/merge-bill.ts holds the precedence rule — API wins on facts, database
wins on the verdict — and both pages call it, so they cannot disagree. Bills
carry analysisGeneratedAt and analysisSourceRef, shown to readers as the text
and date that were actually judged. The steel man is generated and rendered.

One LLM call, with Structured Outputs against prompt/analysis-schema.ts. That
retires the invalid JSON example, the enum re-casing, the regex parse fallback
and ~65 lines of prompt; social-issue-grader.ts is gone, and a social issue now
forces abstain at write time. Tenet titles come from TENETS by id, ending a
three-way drift between TENETS, FALLBACK_TENET_TITLES and an inline copy.

Verified against the live API and Mongo: 185 bills scanned, 185 facts written,
nothing queued for re-analysis, no errors; two concurrent sweeps and exactly one
ran. 18 unit tests cover the merge precedence, the re-analysis decision, and the
schema's validity under strict mode — including the keywords strict mode
rejects, which would 400 every request and degrade every bill to a fallback.

Not verified: the live model call. The local OPENAI_API_KEY is rejected, so
`pnpm eval:bills --refresh` still needs a run with a working key.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant