Skip to content

[Skills] Add deterministic preflight to the review runner - #1124

Merged
jhinpan merged 1 commit into
add-flydsl-code-review-skillfrom
integrate/review-checks-into-1106
Sep 11, 2026
Merged

jhinpan merged 1 commit into
add-flydsl-code-review-skillfrom
integrate/review-checks-into-1106

Conversation

@jhinpan

@jhinpan jhinpan commented Sep 11, 2026

Copy link
Copy Markdown
Collaborator

This PR targets #1106's add-flydsl-code-review-skill branch. It distills the two useful deterministic checks from #1047 into the existing review runner, so #1106 carries the combined review capability into main.

The pinned diff is scanned for added kernel legacy-API spellings and added tests without a statically visible script entry path. Raw observations go to Angles G and I for contextual adjudication; promoted candidates retain the existing independent verifier, challenger, and ranked-report path. Both preflight stages are checkpointed and required: exit 0/1 permits semantic review, while errors/timeouts yield INCOMPLETE. Scanner implementation changes invalidate resume, and the saved diff is checked against the pinned hash. Artifact schema v2 requires valid preflight outputs before publication.

The source scanners and 38 original regressions are adapted from #1047 at 2480f877aa7f74a24d1ce857b48ba92ff0699cf5. The integration also fixes four reproduced scanner edge cases: Python 3.12 f-string literal false positives, unconsumed generator expressions, pattern-match shadowing, and rebound pytest test functions/classes/methods. The selection evidence stays linked to the source revision; no second review skill or combined standalone wrapper is added.

This also fixes #1106's remaining CLI compatibility finding: verbose Claude Code output can be a transcript array. The runner consumes the last result record and preserves all error, permission-denial, and structured-output checks.

Validation:

  • 111 CPU regressions passed: 67 runner/publication tests and 44 scanner tests. Coverage includes G/I routing, scanner failure and resume, pinned-diff integrity, malformed artifact rejection, and verbose CLI success/failure shapes.
  • Real Claude Code 2.1.263, with the existing verbose: true setting: the exact cli_agent invocation returned an array-shaped transcript and completed successfully.
  • Frozen [Feat] Align quant and fused rmsnorm kernels with aiter/triton #481 replay at 7e117e29c9c4582158aed33da4ee6dfe14d2c76e: the entry-path scanner finds the three omitted fused/quant tests at lines 632, 884, and 924.
  • scripts/check_repo.py, whole-tree agent-doc validation, the repository Python style gate, and explicit black/ruff checks on the skill scripts passed.
  • Independent source and final integration reviews completed; all reproduced findings were corrected and retested.

These checks establish scanner and runner behavior. Full model-review corpus precision/recall and GPU execution were not measured. The existing per-angle candidate limit remains; raw scanner leads are preserved, not individually certified. Standard repository CI is configured for PRs targeting main, so the combined branch will run that CI through #1106 after this child PR merges.

Adapt the two source scanners and their regressions from #1047 at 2480f87. Feed raw leads through the existing G/I verification path, preserve preflight identity and failures, and accept verbose CLI result arrays.
@jhinpan
jhinpan merged commit bc5c97a into add-flydsl-code-review-skill Sep 11, 2026
1 check passed
@jhinpan
jhinpan deleted the integrate/review-checks-into-1106 branch September 11, 2026 17:22
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant