Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension


Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
97 changes: 70 additions & 27 deletions README.md
Original file line number Diff line number Diff line change
Expand Up @@ -35,7 +35,7 @@ Node 22.20+ or 24+ and Git are required. Glob matching uses the small `picomatch
Install the published package from npm.

```sh
npm install --save-dev --save-exact @roo-code/judgement@0.2.0
npm install --save-dev --save-exact @roo-code/judgement@0.3.0
export TYPESAFE_API_KEY=...
./node_modules/.bin/judgement check --staged
```
Expand Down Expand Up @@ -92,7 +92,7 @@ incomplete hook runs from completed ones. `--dry-run` sends no model requests.
`--verbose` (or `-v`) streams file/rule progress, request sizes, cache hits,
model answers, and timings to stderr. It does not print source contents or
credentials. JSON reports remain on stdout. Model answers are intermediate;
the final report applies confidence thresholds and coverage requirements.
the final report applies violation-probability thresholds and coverage requirements.

## Rules

Expand All @@ -103,8 +103,8 @@ the final report applies confidence thresholds and coverage requirements.
- `files`: optional repository-relative globs selecting changes that activate
the rule. Other files can still be supporting evidence.
- `context`: optional repository-relative globs adding supporting source files from the proposed snapshot.
- `threshold`: confidence cutoff between 0 and 1; defaults to 0.85. It is not a
measured probability that the change is correct.
- `threshold`: violation-probability cutoff between 0 and 1; defaults to 0.85.
A score at or above it flags a violation. A score below it does not flag one.

Globs use `picomatch` syntax and include dotfiles and hidden directories. Bare names also match path components;
trailing `/` selects a directory. Absolute paths, `..`, negation, and backslashes
Expand Down Expand Up @@ -237,12 +237,53 @@ console.log(formatCalibrationReport(report));

The [wording example](examples/wording/.judgement/) includes a policy and fixtures.

## Calibrating a rule's confidence threshold
## Inspect example requests

Choose a threshold from labeled examples of your rule. Confidence measures how
strongly the model favors its answer; it is not a measured accuracy rate for your
repository. See TypeSafe's [confidence guide](https://docs.typesafe.ai/confidence)
for the distinction. The default `0.85` is a starting point, not a universal cutoff.
Use `prepareExamples` to load labeled fixtures into a model tester:

```js
import { prepareExamples } from '@roo-code/judgement';

const prepared = await prepareExamples({
cwd: process.cwd(),
ruleId: 'wording',
});
const example = prepared.examples[0];
const packet = example.packets[0];
// Send only packet.state and packet.questions to your inference backend.
console.log(example.name, example.expected, example.threshold, packet.stage);
```

Preparation uses disposable Git repositories and the checker's evidence planner.
It makes no inference calls and leaves the real index untouched. Packets include
initial evidence, expanded evidence for manual inspection, and partial screens
where needed. The default binary evaluator uses the initial packet; it does not
expand evidence merely because a probability is near the cutoff. Expected labels remain outside model inputs. The output
includes policy and fixture hashes for detecting stale presets. Options include
`ruleId`, `examplesPath`, `examplesDirectory`, `deadlineMs` (30 seconds per example),
and `signal`.

Replay packets to inspect violation probabilities and latency.
A packet answer is not a full check result: partial screens cannot approve a file,
and unresolved context still prevents approval. Testers with longer timeouts also
do not establish hook performance. Confirm improvements through `testRules` or
`calibrate`, with held-out examples and the production deadline.

The model answers one boolean question: does this change violate the rule?
TypeSafe's [Noul primitive](https://docs.typesafe.ai/primitives/noul) returns the
probability of yes, without a separate confidence score. Judgement flags a violation
at or above the rule's threshold. Every lower score, including 0.5, produces no
finding. Missing required context, unsupported or partial evidence, and request
failures still make the check incomplete. A completed check with no findings is
reported as `pass`; this is not a guarantee that the changes contain no violations.

## Calibrating a rule's violation cutoff

Choose a cutoff from labeled examples of your rule. A higher cutoff requires
stronger evidence to flag a violation; it can also miss more real violations.
The default `0.85` is a starting point. Recalibrate when changing the question,
primitive, or model: a Choice confidence cutoff does not transfer to a Noul
probability cutoff.

1. **Build a small labeled set before looking at scores.** Include clear
violations, valid changes, and valid exceptions that resemble violations.
Expand Down Expand Up @@ -270,7 +311,7 @@ for the distinction. The default `0.85` is a starting point, not a universal cut
several times per example, such as three to five runs, with caching disabled.
When Judgement is embedded in an application, use its evaluator and settings
rather than accidentally testing a different standalone backend. Record the
backend/model version, outcomes, confidence scores, final statuses, and time.
backend/model version, outcomes, violation probabilities, final statuses, and time.
Avoid concurrent checks against the same index; use separate fixtures or
run their repetitions sequentially.

Expand All @@ -279,19 +320,17 @@ for the distinction. The default `0.85` is a starting point, not a universal cut
Also track outright incorrect passes. A valid change reported as incomplete
is not a successful pass: hook mode permits it, but strict mode blocks it.
Likewise, an incomplete violation is a missed block in hook mode.
5. **Verify the chosen threshold through the full checker.** Raw scores are
useful for an initial comparison, but changing the threshold can trigger
context expansion and a different answer. Check final reports, held-out
5. **Verify the chosen threshold through the full checker.** Raw probabilities are useful for comparing cutoffs. Repeat the full check
to verify behavior on actual evidence and measure variation. Check final reports, held-out
examples, and hook deadlines before adopting it. Recalibrate after changing
the rule wording, evidence selection, model, or backend.

The same threshold applies to `pass`, `not_applicable`, and `violation` answers.
An answer below the threshold remains incomplete; `unclear` remains incomplete
regardless of confidence. Lowering the threshold cannot fix missing credentials,
timeouts, unsupported evidence, or an ambiguous rule. If false blocks and true
violations have overlapping scores, clarify the rule or supply the needed
context and test again. Choose the tradeoff per rule: missing a wording issue
and missing an authorization flaw have different consequences.
For binary answers, flag only when `violationProbability >= threshold`. Lowering
the cutoff can recover missed violations but may introduce false blocks. It cannot
fix unavailable inference or missing required evidence. If valid and violating
examples have overlapping scores, clarify the rule or improve the evidence before
selecting a cutoff. Report any rule without positive examples as uncalibrated for
detection; valid examples alone do not establish recall.

## Evidence and large changes

Expand All @@ -305,13 +344,16 @@ selected by `context`. Write localized rules: UI wording, use of a shared
component, or a guard around an operation in that file. Supply the helper or
convention file explicitly when the rule needs it. Repository-wide inventories,
uniqueness, parity, and aggregate rules are outside v1's supported scope; the
model is instructed to return unclear for them.
single-file evidence cannot establish those guarantees.

Checks start with every diff hunk and 12 nearby unchanged lines, plus explicit
`context` files. A small edit to a large file can finish in one request. If the
answer is unclear or below the confidence threshold, Judgement tries the full
staged file; if that will not fit, it tries 80 lines around every hunk. Expansion
uses at most one additional model request, within the same execution deadline.
`context` files. The default binary judgment makes one request for this packet.
A below-threshold score completes without a finding. Supply required supporting
files through `context`; the model cannot request missing semantic context with a
separate answer. `prepareExamples` also provides expanded packets for inspection.
Custom outcome/confidence evaluators can request one expansion by returning
`unclear` or a below-threshold confidence. Expansion includes the full staged file
when it fits, or 80 lines around every hunk, within the same execution deadline.

Large commits run file checks in parallel with bounded concurrency. Oversized
diffs or explicit context are screened in bounded chunks. Every text chunk is visited in a
Expand Down Expand Up @@ -345,10 +387,11 @@ process.exitCode = exitCode(report, { hook: true });
```

An `evaluate(request, signal)` function can supply another transport. It returns
`{ outcome, confidence }`. Use exported `question(request)` to retain the same
`{ violationProbability }`, a finite number between 0 and 1. Use exported `question(request)` to retain the same
rubric, honor the AbortSignal, and set a stable `cacheIdentity` that changes with
the model/backend configuration. Custom evaluators have caching disabled unless
an identity is supplied. Type declarations ship with the package.
an identity is supplied. Type declarations ship with the package. Custom evaluators returning
`{ outcome, confidence }` retain their existing confidence and uncertainty semantics.

`installGitHook({ cwd, command })` is available for managed runtimes. It stores
wrappers in Git metadata, runs the original pre-commit hook first, and preserves
Expand Down
24 changes: 21 additions & 3 deletions docs/agents.md
Original file line number Diff line number Diff line change
Expand Up @@ -38,7 +38,7 @@ Save rules in `.judgement/rules.json`:
- Scope `files` to relevant paths. Globs are relative to the repository root and
match hidden paths too. `context` selects unchanged supporting files from the
same Git snapshot. Keep it focused; it is not a request to inspect the whole repo.
- A threshold is a decision cutoff, not proof of accuracy. The default 0.85 is a
- A threshold is a violation-probability cutoff, not proof of accuracy. The default 0.85 is a
starting point. Do not weaken the rule or lower its threshold just to get green.

## Save labeled examples
Expand Down Expand Up @@ -108,8 +108,8 @@ Inspect false blocks, caught violations, incorrect passes, incomplete results,
operational failures, scores, and timing separately. When a case fails:

1. Check its label, path, rule wording, and selected supporting context.
2. Distinguish an uncertain correct answer from a wrong answer. Lowering a threshold
cannot repair a wrong outcome and may turn uncertainty into a false block.
2. Distinguish an uncertain correct answer from a wrong answer. Lowering a cutoff can catch more violations and also introduce false blocks.
If valid and violating examples overlap, clarify the rule or its evidence.
3. Repeat the run. Never present one score as a guarantee.
4. Validate a candidate against separately labeled held-out examples before adopting
it. Use `--examples` for one rule, or `test --examples-dir` / `calibrate --all
Expand All @@ -132,3 +132,21 @@ missing context, or every possible violation will be handled correctly.
Applications with existing credentials should supply their evaluator to `testRules`
or `calibrateRules`. Keep that adapter thin: do not copy the runner, fixture loading,
isolated repository setup, threshold logic, or reporting into each application.

## Inspect uncertainty in a model tester

Use `prepareExamples` to obtain exact checker request packets from the project's
fixtures. Send each packet's `state` and `questions` to the configured backend;
keep `expected`, names, and thresholds as tester metadata. Inspect both initial
and expanded packets, with raw violation probabilities visible.
Preparation makes no inference calls. Generated inputs can be bundled as presets;
regenerate them when their policy, fixture, or package version changes.

The model answers whether a change violates the rule. Its Noul value is the
probability of a violation. Scores at or above the cutoff flag a violation;
every lower score produces no finding, including an ambiguous score of 0.5.
Missing required evidence, partial screens, unsupported files, and failed requests
remain incomplete. The default binary evaluator does not request context expansion
based on its probability. Use explicit context and compare expanded packets when
investigating a missed violation. Recalibrate when changing question types; Choice
confidence and Noul probability use different meanings and scales.
4 changes: 2 additions & 2 deletions package-lock.json

Some generated files are not rendered by default. Learn more about how customized files appear on GitHub.

2 changes: 1 addition & 1 deletion package.json
Original file line number Diff line number Diff line change
@@ -1,6 +1,6 @@
{
"name": "@roo-code/judgement",
"version": "0.2.0",
"version": "0.3.0",
"description": "Fast Git checks for plain-language business rules, powered by Jev",
"type": "module",
"license": "MIT",
Expand Down
7 changes: 4 additions & 3 deletions src/calibrate.js
Original file line number Diff line number Diff line change
Expand Up @@ -13,6 +13,7 @@ import {
createJevEvaluator,
MODEL,
PROTOCOL_VERSION,
formatAnswer,
validateAnswer,
} from './model.js';

Expand Down Expand Up @@ -312,7 +313,7 @@ export async function calibrate(options = {}) {
operationalFailure,
});
diagnostics(
`${JSON.stringify(example.name)} threshold ${threshold} run ${repetition}: ${result.status}; ${answers.map((answer) => `${answer.outcome} ${answer.confidence.toFixed(2)}`).join(', ') || 'no judgments'}`,
`${JSON.stringify(example.name)} threshold ${threshold} run ${repetition}: ${result.status}; ${answers.map((answer) => formatAnswer(answer)).join(', ') || 'no judgments'}`,
);
}
}
Expand Down Expand Up @@ -391,10 +392,10 @@ export function formatCalibrationReport(report) {
),
`Recommendation: ${report.recommendation ?? 'none'}. ${report.recommendationReason}`,
'Incomplete checks permit hook commits but fail strict checks. No policy files were changed.',
'Example results (confidence may include context expansion):',
'Example results (raw model answers):',
...report.results.map(
(run) =>
` ${JSON.stringify(run.example)} @ ${run.threshold}, run ${run.repetition}: ${run.status}; ${run.answers.map((answer) => `${answer.outcome} ${answer.confidence.toFixed(2)}${answer.complete ? '' : ' [partial]'}`).join(', ') || 'no judgments'}`,
` ${JSON.stringify(run.example)} @ ${run.threshold}, run ${run.repetition}: ${run.status}; ${run.answers.map((answer) => `${formatAnswer(answer)}${answer.complete ? '' : ' [partial]'}`).join(', ') || 'no judgments'}`,
),
];
return lines.join('\n');
Expand Down
35 changes: 27 additions & 8 deletions src/check.js
Original file line number Diff line number Diff line change
Expand Up @@ -23,6 +23,7 @@ import {
MAX_EVIDENCE_BYTES,
MODEL,
PROTOCOL_VERSION,
formatAnswer,
validateAnswer,
} from './model.js';

Expand Down Expand Up @@ -206,9 +207,7 @@ async function run(options, result, signal, stopped) {
result.coverage.cached++;
result.coverage.judged++;
}
trace(
`cache hit: ${label}; answer ${cached.outcome}, confidence ${cached.confidence.toFixed(2)}`,
);
trace(`cache hit: ${label}; ${formatAnswer(cached)}`);
return cached;
} catch {
/* Missing or invalid cache entries are evaluated again. */
Expand All @@ -226,9 +225,12 @@ async function run(options, result, signal, stopped) {
signal.throwIfAborted();
if (!stopped()) result.coverage.judged++;
trace(
`answer: ${label}; ${answer.outcome}, confidence ${answer.confidence.toFixed(2)} (${Math.round(performance.now() - started)} ms)`,
`answer: ${label}; ${formatAnswer(answer)} (${Math.round(performance.now() - started)} ms)`,
);
if (target && !['unclear'].includes(answer.outcome)) {
if (
target &&
('violationProbability' in answer || answer.outcome !== 'unclear')
) {
// Cache failures cannot turn an otherwise completed check into a failure.
try {
await mkdir(cacheDir, { recursive: true, mode: 0o700 });
Expand Down Expand Up @@ -338,6 +340,22 @@ async function run(options, result, signal, stopped) {
}
const contextUnresolved = [...record.unresolved];
const consume = (answer, request) => {
if ('violationProbability' in answer) {
if (answer.violationProbability >= criterion.threshold) {
record.status = 'violation';
record.findings.push({
paths: request.focusPaths,
confidence: answer.violationProbability,
violationProbability: answer.violationProbability,
sources: request.evidence.map(({ path, kind, line }) => ({
path,
kind,
line,
})),
});
}
return;
}
if (
answer.outcome === 'violation' &&
answer.confidence >= criterion.threshold
Expand Down Expand Up @@ -395,8 +413,9 @@ async function run(options, result, signal, stopped) {
if (bytes(request) <= MAX_EVIDENCE_BYTES) {
let answer = await ask(request, criterion.id);
const uncertain =
answer.outcome === 'unclear' ||
answer.confidence < criterion.threshold;
!('violationProbability' in answer) &&
(answer.outcome === 'unclear' ||
answer.confidence < criterion.threshold);
if (uncertain && !options.dryRun) {
// Expand only when needed; keep every hunk in either request.
const parts =
Expand Down Expand Up @@ -530,7 +549,7 @@ export function formatReport(report) {
lines.push(`${rule.status}: ${rule.id}: ${rule.rule}`);
for (const finding of rule.findings)
lines.push(
` ${finding.paths.join(', ')} (confidence ${finding.confidence.toFixed(2)})`,
` ${finding.paths.join(', ')} (${finding.violationProbability === undefined ? 'confidence' : 'violation probability'} ${finding.confidence.toFixed(2)})`,
);
for (const unresolved of new Set(rule.unresolved))
lines.push(` ${unresolved}`);
Expand Down
15 changes: 10 additions & 5 deletions src/examples.js
Original file line number Diff line number Diff line change
Expand Up @@ -12,11 +12,7 @@ import { ConfigurationError, readPolicy, POLICY_PATH } from './policy.js';

const hash = (text) => createHash('sha256').update(text).digest('hex');

async function runExamples(options, mode) {
if (mode === 'test' && options.thresholds !== undefined)
throw new ConfigurationError(
'Tests use configured thresholds. Use calibrate to compare candidates.',
);
export async function loadExampleInputs(options = {}) {
const policy = await readPolicy(options.cwd ?? process.cwd());
const rules = policy.criteria.filter(
(rule) => !options.ruleId || rule.id === options.ruleId,
Expand Down Expand Up @@ -56,6 +52,15 @@ async function runExamples(options, mode) {
return { rule, text };
}),
);
return { policy, fixtures };
}

async function runExamples(options, mode) {
if (mode === 'test' && options.thresholds !== undefined)
throw new ConfigurationError(
'Tests use configured thresholds. Use calibrate to compare candidates.',
);
const { policy, fixtures } = await loadExampleInputs(options);
const policyText = JSON.stringify(policy);
const cwd = await mkdtemp(join(tmpdir(), 'judgement-examples-'));
const reports = [];
Expand Down
Loading
Loading