Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
236 changes: 236 additions & 0 deletions PRDs/hub-047-answer-completeness-scoring.md
Original file line number Diff line number Diff line change
@@ -0,0 +1,236 @@
# HUB-047 — Scoring design for end-to-end answer quality

Status: DESIGN ONLY, 2026-08-13. No question set is built and no code is
written. This document settles *how a report is scored* first, because the
scoring scheme determines what the question set must contain, and building
the set first would force the scoring to fit whatever was collected.

Prior art reviewed: 3 arXiv abstracts fetched and read 2026-08-13 for this
design, plus 16 read for the HUB-049 pass (citations at the end; every claim
below is attributed to a fetched abstract or marked as this design's own
choice).

## Problem

The system verifies every claim it publishes and measures nothing about what
it failed to say.

The claim gate checks each drafted claim against its cited evidence, with
per-reference necessity, and rejects padding. `benchmark_report_retrieval.py`
checks citation validity. `benchmark_retrieval_exact_terms.py` checks that
needles are retrievable. HUB-044 measures retrieval breadth. Every one of
these is a **precision-side** instrument.

Consequence: a report that states three true, trivially-supported facts and
omits everything that mattered passes every gate this system has, at
`citation_validity = 1.0`. Nothing notices.

That gap now blocks work. HUB-049 is implemented and cannot be adopted;
HUB-045 and HUB-048 would hit the same wall. Two candidate proxies are ruled
out by the prior art rather than by opinion: retrieval recall stops predicting
accuracy once a context budget forces evidence to be discarded (2607.00725),
which is this system's measured condition, and source coverage does not stand
in either (2603.22633 measured large coverage gains worth roughly 0.01 F1).

## The scoring scheme

### Two axes, never merged

Precision and completeness are scored separately and reported separately
(2603.09891 treats completeness as its own axis, measured independently of
whether the cited support checks out; 2509.26184 decomposes a cited report
along the same two axes).

| Axis | Question | Instrument | Status |
|---|---|---|---|
| Precision | Is what the report said supported by what it cited? | claim gate, `citation_validity` | **exists** |
| Completeness | Did the report say what a good answer must say? | nugget recall (below) | **missing — this item** |

A single blended score is explicitly rejected. The two axes trade against each
other — a report can buy completeness with unsupported claims — and blending
them hides exactly the trade a change needs to be judged on. This is the same
discipline HUB-044 already applies to coverage versus recall.

### The completeness metric: nugget recall

Nugget evaluation comes from the TREC QA track and was revived for RAG by the
AutoNuggetizer framework (2504.15068): a question is distilled into a short
list of **atomic facts a good answer ought to contain**, each labelled
**vital** (must be present) or **okay** (worth having, not essential); a
report is then scored by which nuggets it contains.

Per question, per report:

- Each nugget is assigned **supported / partial / absent** against the report
text (the three-way assignment of 2504.15068 — a binary label discards the
common case of a nugget stated incompletely).
- `vital_nugget_recall` — **the primary number.** Fraction of vital nuggets
assigned supported.
- `weighted_nugget_recall` — vital and okay pooled, vital counted double.
*(The 2:1 weight is this design's choice, not an attributed value; it is
reported alongside the unweighted vital figure so the weight can never be
the thing that moves a decision.)*
- `all_nugget_recall` — unweighted, for the record.
- Partial assignments count as 0 in every recall figure and are reported as a
separate `partial_rate`. Crediting half a fact would make a report that
gestures at everything beat one that establishes anything.

Retrieval coverage@k (HUB-044) continues to be reported beside these and
continues to gate nothing.

### A nugget is only credited when a *verified* claim carries it

A nugget is assigned supported only if the report text carrying it is a claim
that passed the claim gate (2509.26184 scores completeness over the report's
*supported* claims specifically). Otherwise completeness could be raised by
emitting unverified assertions, which is the failure the gate exists to
prevent.

## Construction: where nuggets come from

This is where the design can quietly destroy its own validity, so it is
specified tightly.

### Nuggets are mined from documents, not from retrieval

**The trap:** nuggets mined from what retrieval surfaced can only ever reward
facts retrieval already surfaced (2504.15068's own stated limitation — whole
corpus blind spots stay hidden). The same self-reference already flattered a
published study in the neighbouring literature: SciRet's relevance labels were
generated by the very hybrid system being scored, mechanically producing
perfect recall for that configuration (2608.03860).

**The rule:** candidate nuggets are extracted from the **retained documents
resolved by scope** — `documents_for_job` for a job-scoped question,
`documents_matching` for a filtered corpus question — never from
`ScopedRetrievalService` output. The pipeline under test must not choose the
yardstick it is measured against. This is testable and will be asserted: the
construction command must not import the retrieval service.

### The metered judge is not spent on construction

Drafting runs on the **local** model (`qwen3.5:9b`); the quality bar is human
verification, not the drafter. The sealed v4 protocol already in this repo
established that shape — draft with a model, verify blind by hand, seal the
result — and the drafter's cost is therefore irrelevant to the set's validity.

**Metered judge calls required by HUB-047: zero.** The only metered calls in
the loop are the claim gate's, which a report generation spends anyway. This
is what makes the item affordable, and it should stay true: any future design
that needs the metered judge per evaluation run has regressed.

### The expensive resource is operator time, not compute

Approximately 15 questions × ~12 candidate nuggets ≈ **180 nuggets to verify
and label** vital/okay by hand. That is the real cost of this item and it is
the operator's own hours. It cannot be delegated to the drafter without
reintroducing the circularity the whole design exists to avoid.

Mitigation for the residual blind spot: for a stated subset of questions the
verifier reads the scope's source documents directly and *adds* nuggets the
drafter missed. Deletion-only verification can prune a bad list but cannot
discover an omission, and omissions are the thing being measured.

### Two sets, two different disciplines

These are easy to conflate and the standing v4 obligation makes the confusion
costly:

| Set | Purpose | Discipline |
|---|---|---|
| **Calibration set** | human labels validating the local assigner | blind, sealed, **consumed — never reused** |
| **Evaluation set** | questions + verified nuggets | **stable, versioned, re-run freely** |

A regression benchmark that could only be run once would be useless; a
calibration set reused after consumption invalidates the agreement number.
Both statements are true simultaneously and the two artefacts must be stored
and named separately.

### The local assigner must be calibrated before it is trusted

Nugget assignment is done locally and unmetered, so its agreement with human
labels is the load-bearing assumption. Calibrate exactly as v4 did: blind
human assignment on the calibration set, then agreement against the local
assigner, recorded before any score is quoted. If agreement fails to clear a
bar stated in advance, the assigner is the finding and the metric does not
ship. Lexical-only matching is expected to be insufficient here — it silently
misses paraphrased-but-correct coverage (2603.09891) — which is precisely why
the assigner is a model and why it must be measured.

## What the instrument is, and is not, allowed to be used for

All three fetched abstracts converge on the same caveat: nugget-style scoring
correlates with human judgment **at the system level**, while per-topic and
per-report agreement is noticeably noisier (2504.15068, 2509.26184).

Therefore:

- **Allowed:** ranking configurations. Is `EVIDENCE_PACKING=marginal_gain`
better than `rank`? Does late chunking (HUB-045) help? Does a knowledge
graph earn its keep (HUB-048)?
- **Not allowed:** diagnosing a single report, or gating a single report's
publication. A per-report number from this instrument is noise.

This also sizes the question set: enough questions that a system-level
comparison is meaningful, not enough to make hand-verification impossible.
**15 questions is the design's starting point**, split across the two scopes
HUB-043 created — job-topic questions exercising the report path, and
cross-job corpus questions exercising `/query` and `/rag` — because a metric
that only covers one scope would leave the other unjudgeable.

## Acceptance criteria

1. Two axes reported separately; no blended score anywhere.
2. Nuggets provably mined from document scope, not retrieval output —
asserted by test, not by convention.
3. Zero metered judge calls per evaluation run.
4. Local assigner agreement with blind human labels recorded before any
score is quoted, against a bar set in advance.
5. Calibration and evaluation sets stored separately with their different
reuse rules stated in the files themselves.
6. A baseline recorded for the current deployed configuration, so HUB-049's
flag becomes a decidable question.

## Known limits, recorded rather than discovered later

- **Recall is against a self-authored reference.** Omissions outside the
nugget lists are invisible (2509.26184, 2603.09891). Directly mitigated only
by the document-reading subset above, and never fully.
- **Per-report scores are noise.** See above; enforced by how the instrument
is allowed to be used.
- **Private corpus, no assessor pool.** TREC-style reference lists rest on
pooled human assessment over a shared public corpus; here there is one
annotator and no inter-assessor agreement to report (2603.09891). Single
annotator is a stated limitation of every number this produces.
- **The corpus grows.** A nugget verified against today's 679 documents may
become answerable in new ways later. The evaluation set is versioned so a
score is always quoted against a corpus state.

## Citations

Fetched and read for this design, 2026-08-13:

- **arXiv:2504.15068** — *The Great Nugget Recall: Automating Fact Extraction
and RAG Evaluation with Large Language Models*. Primary mechanism: vital/okay
nuggets, three-way assignment, weighted recall; automation calibrated
against manual nugget creation on the TREC 2024 RAG Track. Also the source
of the pool-blind-spot limitation.
<https://arxiv.org/abs/2504.15068>
- **arXiv:2509.26184** — *Auto-ARGUE: LLM-Based Report Generation Evaluation*.
Supporting evidence: decomposes a citation-backed report into sentence-level
claims and scores attribution and completeness as independent axes, with
completeness measured over supported claims.
<https://arxiv.org/abs/2509.26184>
- **arXiv:2603.09891** — *Overview of the TREC 2025 Retrieval Augmented
Generation (RAG) Track*. Supporting evidence for separating completeness
from attribution; source of the private-corpus and lexical-matching limits.
<https://arxiv.org/abs/2603.09891>

Carried from the HUB-049 pass, 2026-08-13:

- **arXiv:2607.00725** — recall stops predicting accuracy under a context
budget; the reason a retrieval-side proxy cannot stand in here.
- **arXiv:2603.22633** — large source-coverage gains worth roughly 0.01 F1;
the reason coverage cannot stand in either.
- **arXiv:2608.03860** — relevance labels generated by the system under test
mechanically flatter it; the failure mode the construction rule avoids.
26 changes: 20 additions & 6 deletions backlog.md
Original file line number Diff line number Diff line change
Expand Up @@ -1521,12 +1521,26 @@ scored on answer correctness and source coverage. HUB-044 is its first
metric. This is what would let HUB-045, HUB-048 and HUB-049's flag be decided
by measurement rather than argument.

**Open design question, unanswered by the prior art.** Every paper read
validates on short-answer QA with gold spans. This system's output is
long-form cited synthesis over a private, growing corpus, where "correct" is
not a string match. Settle the scoring design before building the set — the
existing sealed judge protocol (v4) is the closest thing already in the repo
and is the obvious starting point.
**Scoring design settled 2026-08-13** — `PRDs/hub-047-answer-completeness-scoring.md`,
grounded in three further fetched abstracts. Stage 1 of the item is done; no
question set is built yet. The design in one paragraph: score precision and
completeness as **two axes that are never blended**, keep the existing claim
gate as the precision axis, and add **nugget recall** as the completeness axis
— per question a short list of atomic facts a good answer must contain,
labelled vital or okay, assigned three-way against the report, and credited
only when carried by a claim that passed the gate. Nuggets are mined from
**document scope, never from retrieval output**, or the pipeline picks its own
yardstick. Drafting is local and human verification is the quality bar, so
**HUB-047 costs zero metered judge calls** — the expensive resource is roughly
180 nuggets of operator hand-verification. Calibration and evaluation sets are
kept separate because one is consumed on use and the other must be re-runnable.

The instrument ranks **configurations**, not reports: nugget-style scoring
agrees with human judgment at the system level and is noisy per report
(arXiv:2504.15068, arXiv:2509.26184). It may never gate a single report.

**Next stage:** build the 15-question set against that design — the operator
hand-verification pass is the gating cost and needs scheduling, not code.

### HUB-048 — Knowledge-graph go/no-go, decided by measurement

Expand Down
Loading