Skip to content

ci(pr-automation): let general review repair converge - #2062

Merged
Benoît Cortier (CBenoit) merged 5 commits into
masterfrom
claude/ironrdp-specialist-validation-b4f225
Sep 30, 2026
Merged

Benoît Cortier (CBenoit) merged 5 commits into
masterfrom
claude/ironrdp-specialist-validation-b4f225

Conversation

@CBenoit

@CBenoit Benoît Cortier (CBenoit) commented Sep 30, 2026 •

Copy link
Copy Markdown
Member

A general review failed after exhausting its repairs because the model never had the specialist candidate IDs (https://github.com/Devolutions/IronRDP/actions/runs/36724237360). It finalized with no dispositions, its one repair lookup hit a rejected ./ path, and the last repair invented nine IDs from the positional diagnostic. The published reason kept only the final attempt's counts, so none of this was visible.

Validators may now return a content-free detail, which replaces the short reason in repair feedback and is kept per attempt, and repair-only guidance, which may quote trusted identifiers and is never logged. The final-review validator uses them to report every coordinate and the exact reviewer and finding_id pairs to copy: an entry to add for absent candidates, and the existing entry to correct when only its rationale is unusable. This deliberately relaxes the earlier rule that repair feedback never quotes candidate IDs; they come from the validated aggregate and are restricted to ^[a-z][a-z0-9-]{0,63}$.

The general stage also receives the candidate ID index through a new trusted prompt-context input, and each stage's rejected attempts reach the workflow summary in a new "Rejected output attempts" table. The openai-agent intent records the prompt context trust rule and the detail and guidance contract.

The sandbox now drops empty and . tool path segments, lists the allowed capabilities for the workspace root, and truncates lines over the per-line limit instead of failing the read, which makes an oversized pull request body readable again. Search results are bounded by the serialized tool-result budget, so many long matching lines truncate the result instead of failing the call.

🤖 Generated with Claude Code

A general review failed after exhausting its repairs because the model
never had the specialist candidate IDs. It finalized with no
dispositions, its one repair lookup hit a rejected `./` path, and the
last repair invented nine IDs from the positional diagnostic. The
published reason kept only the final attempt's counts, so none of this
was visible.

Validators may now return a content-free detail, which replaces the
short reason in repair feedback and is kept per attempt, and repair-only
guidance, which may quote trusted identifiers and is never logged. The
final-review validator uses them to report every coordinate and the
exact reviewer and finding_id pairs still missing or uncited.

The general stage also receives the candidate ID index through a new
trusted prompt-context input, and each stage's rejected attempts reach
the workflow summary.

The sandbox now drops empty and `.` tool path segments, lists the
allowed capabilities for the workspace root, and truncates lines over
the per-line limit instead of failing the read. This makes an oversized
pull request body readable again.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Copilot AI balanced review requested due to automatic review settings September 30, 2026 16:37
@github-actions github-actions Bot added risk/low Self-contained change with no cross-crate behavioral effect scope/tooling Build, CI, release, or developer tooling size/XL Size: up to 1299 counted lines and 49 files; exceeds L in either measure labels Sep 30, 2026

Copilot AI left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Copilot review overview

🔵 Needs a closer look

The cross-cutting repair and sandbox changes need human review, and unresolved search-result and repair-guidance issues remain.

Review effort: Balanced
Findings: 2 Medium severity

Open (2)
What changed in this PR

This PR helps IronRDP’s automated general review recover from invalid output by supplying specialist candidate IDs up front and making rejected attempts easier to diagnose.

Changes:

  • Add candidate IDs to the general reviewer’s prompt and improve repair feedback.
  • Carry rejected-attempt details into the workflow summary.
  • Make sandbox paths and overlong lines easier for the reviewer’s read-only tools to handle.
File Description
.github/​workflows/​review-pipeline.yml Passes the candidate index and rejection diagnostics between stages.
.github/​pr-automation/​validate-final-review.js Produces candidate indexes and fuller repair diagnostics.
.github/​pr-automation/​review-report.js Preserves bounded rejection records.
.github/​pr-automation/​review-report-summary.js Displays rejected attempts in the workflow summary.
.github/​pr-automation/​review-pipeline.js Extracts rejection records from agent diagnostics.
.github/​pr-automation/​prompts/​general-reviewer.md Directs the reviewer to use the supplied candidate index.
.github/​pr-automation/​automation.test.js Tests indexing, diagnostics, and reporting.
.github/​pr-automation/​agent-validator.js Forwards repair detail and guidance.
.github/​actions/​openai-agent/​test/​validator.test.js Tests rejection field limits.
.github/​actions/​openai-agent/​test/​sandbox.test.js Tests path handling and long-line behavior.
.github/​actions/​openai-agent/​test/​main.test.js Tests prompt context and repair feedback.
.github/​actions/​openai-agent/​test/​config.test.js Tests prompt-context loading and bounds.
.github/​actions/​openai-agent/​test/​agent.test.js Tests repair feedback and diagnostic separation.
.github/​actions/​openai-agent/​src/​validator.js Validates optional rejection detail and guidance.
.github/​actions/​openai-agent/​src/​sandbox.js Normalizes tool paths and truncates long-line excerpts.
.github/​actions/​openai-agent/​src/​provider.js Records bounded rejection detail.
.github/​actions/​openai-agent/​src/​main.js Reads the prompt-context input.
.github/​actions/​openai-agent/​src/​limits.js Defines new context and feedback limits.
.github/​actions/​openai-agent/​src/​config.js Appends bounded context to the prompt.
.github/​actions/​openai-agent/​src/​agent.js Sends fuller feedback during repair.
.github/​actions/​openai-agent/​INTENT.md Documents the new runtime behavior.
.github/​actions/​openai-agent/​action.yml Declares the prompt-context input.

Comment thread .github/actions/openai-agent/src/sandbox.js Outdated
Comment thread .github/pr-automation/validate-final-review.js Outdated
Long matching lines each keep up to 8 KiB, so a search over enough of
them exceeded the tool-result limit and failed outright. Matches are now
collected against the serialized budget, with the envelope reserved at
its largest shape, and the search reports truncated when it stops.

A candidate whose only entry has an unusable rationale was listed as
needing an entry, which invited a duplicate. Guidance now asks to add an
entry only for absent candidates and to correct the existing entry
otherwise.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Intent files are human-owned, so the earlier edits to the openai-agent
INTENT.md are withdrawn. The behavior they described is proposed for
human review on the pull request instead.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Apply the maintainer-approved intent for the new prompt context input
and the validator detail and guidance contract.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
@CBenoit
Benoît Cortier (CBenoit) merged commit 0c5419e into master Sep 30, 2026
50 of 54 checks passed
@CBenoit
Benoît Cortier (CBenoit) deleted the claude/ironrdp-specialist-validation-b4f225 branch September 30, 2026 17:29

This branch was successfully deployed

1 active deployment
llm-providers — b82187ac Deployed Sep 30, 2026 by CBenoit via Classify pull request #1412
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

risk/low Self-contained change with no cross-crate behavioral effect scope/tooling Build, CI, release, or developer tooling size/XL Size: up to 1299 counted lines and 49 files; exceeds L in either measure

Development

Successfully merging this pull request may close these issues.

2 participants