Let a bill's verdict keep up with the bill - #91
Open
mikaalnaik wants to merge 1 commit into
Open
mikaalnaik wants to merge 1 commit into
mikaalnaik wants to merge 1 commit into
Conversation
Three things were wrong, and they compounded. Analyses were write-once. fromCivicsProjectApiBill was the only caller of summarizeBillText, and [id]/page.tsx only reached it when a bill was absent from Mongo — so existingBill inside it was always null and the sourceChanged/countChanged guard could never fire. A bill analyzed at first reading kept that verdict through every amendment. The list and the detail page disagreed. The list merged fresh Civics data over the stored document; the detail page rendered the stored document alone, whose status and stages froze at ingestion. A bill could read "Royal Assent" in the list and "Introduced" on its own page. The prompt and the parser disagreed. steel_man, needs_more_info and missing_details were read off the response but never asked for, so they were empty on every bill in the database. The output-format block showed the model invalid JSON while the parser did a bare JSON.parse. And is_social_issue was requested and discarded while a second gpt-5 call answered the same question from the first 8000 characters, swallowed its errors to false, and was never allowed to change the verdict — the drift migrations/1.ts existed to repair. Now: a scheduled sweep (services/refresh.ts) refreshes every bill's facts each run with no LLM, and re-analyzes only bills that are new or whose text changed, against a per-sweep budget. It takes a Mongo lease so a second replica is safe. src/instrumentation.ts schedules it in-process; POST /bills/api/refresh is the manual kick. No page view can trigger an OpenAI call any more — fromCivicsProjectApiBill is a pure converter. utils/merge-bill.ts holds the precedence rule — API wins on facts, database wins on the verdict — and both pages call it, so they cannot disagree. Bills carry analysisGeneratedAt and analysisSourceRef, shown to readers as the text and date that were actually judged. The steel man is generated and rendered. One LLM call, with Structured Outputs against prompt/analysis-schema.ts. That retires the invalid JSON example, the enum re-casing, the regex parse fallback and ~65 lines of prompt; social-issue-grader.ts is gone, and a social issue now forces abstain at write time. Tenet titles come from TENETS by id, ending a three-way drift between TENETS, FALLBACK_TENET_TITLES and an inline copy. Verified against the live API and Mongo: 185 bills scanned, 185 facts written, nothing queued for re-analysis, no errors; two concurrent sweeps and exactly one ran. 18 unit tests cover the merge precedence, the re-analysis decision, and the schema's validity under strict mode — including the keywords strict mode rejects, which would 400 every request and degrade every bill to a fallback. Not verified: the live model call. The local OPENAI_API_KEY is rejected, so `pnpm eval:bills --refresh` still needs a run with a working key. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Builder MP publishes a pro-growth verdict on every bill before the 45th Parliament. Three things were wrong with how those verdicts were produced and shown, and they compounded.
What was broken
Analyses were write-once.
fromCivicsProjectApiBillwas the only caller ofsummarizeBillText, and[id]/page.tsxonly reached it when a bill was absent from Mongo. SoexistingBillinside that function was alwaysnull, and thesourceChanged/countChangedregeneration guard could never fire. A bill analyzed at first reading kept that verdict through every committee amendment. Only a manual/reprocessupdated it — there was no cron, no queue, no worker.The list and the detail page disagreed. The list merged fresh Civics API data over the stored document; the detail page rendered the stored document alone, whose
statusandstagesfroze at ingestion. A bill could read "Royal Assent" in the list and "Introduced" on its own page.The prompt and the parser disagreed.
steel_man,needs_more_infoandmissing_detailswere read off the response but never requested, so they were empty on every bill in the database. The output-format block showed the model invalid JSON (unescaped nested quotes, a trailing comma) while the parser did a bareJSON.parse; anything fenced or malformed fell to a regex that breaks on apostrophes and yielded a permanentabstain. Andis_social_issuewas requested and discarded, while a secondgpt-5call answered the same question from the first 8000 characters, swallowed its errors tofalse, and was never allowed to change the verdict — the driftmigrations/1.tsexisted to repair by hand.What this changes
Freshness — the sweep
services/refresh.tsis the job that didn't exist. Two passes of deliberately unequal cost:analysisBudget(default 10) per sweep. The decision lives inservices/refresh-decision.tsand is unit tested.It takes a Mongo lease (
models/JobLock.ts) before doing anything, so running more than one replica is safe rather than expensive.src/instrumentation.tsschedules it in-process (off unlessBILLS_REFRESH_ENABLED=true);POST /bills/api/refreshbehindBILLS_CRON_SECRETis the manual kick and the seam for moving the schedule out later.No page view can trigger an OpenAI call any more —
fromCivicsProjectApiBillis now a pure converter. A bill the sweep hasn't reached shows its real facts with "Analysis pending".Trust — one set of facts
utils/merge-bill.tsholds the precedence rule once (API wins on facts, database wins on the verdict) and both pages call it. Bills carryanalysisGeneratedAtandanalysisSourceRef, surfaced to readers:The steel man — on the model and the admin form since the start, but never generated and never rendered — is now both, as "The other side".
Quality — one call, a shape it can't get wrong
Structured Outputs against
prompt/analysis-schema.ts. That retires the invalid-JSON example, the enum re-casing, the regex parse fallback and ~65 lines of prompt.social-issue-grader.tsis deleted; the main call answersis_social_issue, andfromRawAnalysisforcesabstainwhen it's true. Tenet titles now come fromTENETSby id, ending a three-way drift betweenTENETS,FALLBACK_TENET_TITLES, and an inline copy in the converter.Also fixed in passing:
page.tsxreadprocess.env.MONGO_URIdirectly and missed theMONGODB_URIfallbackenv.tsprovides (a deploy setting onlyMONGODB_URIrenders an empty list while detail pages work);CIVICS_PROJECT_BASE_URLwas documented as configurable but always ignored its env var; the reprocess route now shares the one write path; shipped typos ("crticial" ×3, "tenents" ×2).Verified
locked.pnpm test:bills) covering merge precedence, the re-analysis decision, and schema validity under strict mode.tscclean;eslint0 errors (3 warnings, all pre-existing)./billsand/bills/[id]serve 200 in ~250ms.Two problems the live run caught:
analysisGeneratedAt), which would have re-paid for all 185 analyses.analysisReasonnow keys off whether a verdict exists;deferredwent 185 → 0.minItems/maxItemsare rejected by OpenAI strict mode — leaving them in would have 400'd every request and silently degraded every bill to a fallbackabstain. There's a test guarding the banned-keyword list.Not verified
The live model call. The local
OPENAI_API_KEYis rejected (invalid_api_key), so I could not run a real eval.pnpm eval:bills --refreshneeds a run with a working key before this deploys. The zero-token--fallbackrun passes 110/120; the 10 failures are allqp-counton the degraded path and pre-date this work.Deploying
New env vars for Dokploy — all optional, all off by default:
BILLS_REFRESH_ENABLEDtruestarts the scheduled sweepBILLS_REFRESH_INTERVAL_MINUTESBILLS_CRON_SECRETPOST /bills/api/refresh; route returns 503 without itExisting rows show no provenance line or steel man until the sweep re-analyzes them — the components render nothing rather than something false, since we genuinely don't know when those verdicts were reached.
🤖 Generated with Claude Code