ci(pr-automation): let general review repair converge - #2062
Merged
Benoît Cortier (CBenoit) merged 5 commits intoSep 30, 2026
Merged
Benoît Cortier (CBenoit) merged 5 commits into
Benoît Cortier (CBenoit) merged 5 commits into
Conversation
A general review failed after exhausting its repairs because the model never had the specialist candidate IDs. It finalized with no dispositions, its one repair lookup hit a rejected `./` path, and the last repair invented nine IDs from the positional diagnostic. The published reason kept only the final attempt's counts, so none of this was visible. Validators may now return a content-free detail, which replaces the short reason in repair feedback and is kept per attempt, and repair-only guidance, which may quote trusted identifiers and is never logged. The final-review validator uses them to report every coordinate and the exact reviewer and finding_id pairs still missing or uncited. The general stage also receives the candidate ID index through a new trusted prompt-context input, and each stage's rejected attempts reach the workflow summary. The sandbox now drops empty and `.` tool path segments, lists the allowed capabilities for the workspace root, and truncates lines over the per-line limit instead of failing the read. This makes an oversized pull request body readable again. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…ialist-validation-b4f225
Benoît Cortier (CBenoit)
deployed
to
llm-providers
September 30, 2026 16:37 — with
GitHub Actions
Active
Copilot started reviewing on behalf of
Benoît Cortier (CBenoit)
September 30, 2026 16:38
View session
Contributor
There was a problem hiding this comment.
Copilot review overview
🔵 Needs a closer look
The cross-cutting repair and sandbox changes need human review, and unresolved search-result and repair-guidance issues remain.
Review effort: Balanced
Findings: 2
Open (2)
What changed in this PR
This PR helps IronRDP’s automated general review recover from invalid output by supplying specialist candidate IDs up front and making rejected attempts easier to diagnose.
Changes:
- Add candidate IDs to the general reviewer’s prompt and improve repair feedback.
- Carry rejected-attempt details into the workflow summary.
- Make sandbox paths and overlong lines easier for the reviewer’s read-only tools to handle.
| File | Description |
|---|---|
.github/workflows/review-pipeline.yml |
Passes the candidate index and rejection diagnostics between stages. |
.github/pr-automation/validate-final-review.js |
Produces candidate indexes and fuller repair diagnostics. |
.github/pr-automation/review-report.js |
Preserves bounded rejection records. |
.github/pr-automation/review-report-summary.js |
Displays rejected attempts in the workflow summary. |
.github/pr-automation/review-pipeline.js |
Extracts rejection records from agent diagnostics. |
.github/pr-automation/prompts/general-reviewer.md |
Directs the reviewer to use the supplied candidate index. |
.github/pr-automation/automation.test.js |
Tests indexing, diagnostics, and reporting. |
.github/pr-automation/agent-validator.js |
Forwards repair detail and guidance. |
.github/actions/openai-agent/test/validator.test.js |
Tests rejection field limits. |
.github/actions/openai-agent/test/sandbox.test.js |
Tests path handling and long-line behavior. |
.github/actions/openai-agent/test/main.test.js |
Tests prompt context and repair feedback. |
.github/actions/openai-agent/test/config.test.js |
Tests prompt-context loading and bounds. |
.github/actions/openai-agent/test/agent.test.js |
Tests repair feedback and diagnostic separation. |
.github/actions/openai-agent/src/validator.js |
Validates optional rejection detail and guidance. |
.github/actions/openai-agent/src/sandbox.js |
Normalizes tool paths and truncates long-line excerpts. |
.github/actions/openai-agent/src/provider.js |
Records bounded rejection detail. |
.github/actions/openai-agent/src/main.js |
Reads the prompt-context input. |
.github/actions/openai-agent/src/limits.js |
Defines new context and feedback limits. |
.github/actions/openai-agent/src/config.js |
Appends bounded context to the prompt. |
.github/actions/openai-agent/src/agent.js |
Sends fuller feedback during repair. |
.github/actions/openai-agent/INTENT.md |
Documents the new runtime behavior. |
.github/actions/openai-agent/action.yml |
Declares the prompt-context input. |
Long matching lines each keep up to 8 KiB, so a search over enough of them exceeded the tool-result limit and failed outright. Matches are now collected against the serialized budget, with the envelope reserved at its largest shape, and the search reports truncated when it stops. A candidate whose only entry has an unusable rationale was listed as needing an entry, which invited a duplicate. Guidance now asks to add an entry only for absent candidates and to correct the existing entry otherwise. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Benoît Cortier (CBenoit)
deployed
to
llm-providers
September 30, 2026 16:48 — with
GitHub Actions
Active
Intent files are human-owned, so the earlier edits to the openai-agent INTENT.md are withdrawn. The behavior they described is proposed for human review on the pull request instead. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Benoît Cortier (CBenoit)
deployed
to
llm-providers
September 30, 2026 16:54 — with
GitHub Actions
Active
Apply the maintainer-approved intent for the new prompt context input and the validator detail and guidance contract. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Benoît Cortier (CBenoit)
had a problem deploying
to
llm-providers
September 30, 2026 17:14 — with
GitHub Actions
Error
Benoît Cortier (CBenoit)
deployed
to
llm-providers
September 30, 2026 17:16 — with
GitHub Actions
Active
Benoît Cortier (CBenoit)
deleted the
claude/ironrdp-specialist-validation-b4f225
branch
September 30, 2026 17:29
This branch was successfully deployed
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.

A general review failed after exhausting its repairs because the model never had the specialist candidate IDs (https://github.com/Devolutions/IronRDP/actions/runs/36724237360). It finalized with no dispositions, its one repair lookup hit a rejected
./path, and the last repair invented nine IDs from the positional diagnostic. The published reason kept only the final attempt's counts, so none of this was visible.Validators may now return a content-free detail, which replaces the short reason in repair feedback and is kept per attempt, and repair-only guidance, which may quote trusted identifiers and is never logged. The final-review validator uses them to report every coordinate and the exact
reviewerandfinding_idpairs to copy: an entry to add for absent candidates, and the existing entry to correct when only its rationale is unusable. This deliberately relaxes the earlier rule that repair feedback never quotes candidate IDs; they come from the validated aggregate and are restricted to^[a-z][a-z0-9-]{0,63}$.The general stage also receives the candidate ID index through a new trusted
prompt-contextinput, and each stage's rejected attempts reach the workflow summary in a new "Rejected output attempts" table. The openai-agent intent records the prompt context trust rule and the detail and guidance contract.The sandbox now drops empty and
.tool path segments, lists the allowed capabilities for the workspace root, and truncates lines over the per-line limit instead of failing the read, which makes an oversized pull request body readable again. Search results are bounded by the serialized tool-result budget, so many long matching lines truncate the result instead of failing the call.🤖 Generated with Claude Code