Skip to content

Simplify grade-tests decisions for PR feedback - #1258

Merged
Evangelink merged 4 commits into
mainfrom
dev/amauryleve/grade-tests-agent
Oct 5, 2026
Merged

Evangelink merged 4 commits into
mainfrom
dev/amauryleve/grade-tests-agent

Conversation

@Evangelink

Copy link
Copy Markdown
Member

Summary

  • Make Pass, Failed, Uncertain, and Not applicable the primary grade-tests decisions while retaining A-F quality as supporting detail.
  • Define Failed as any evidence-backed actionable improvement, independent of the letter grade, and reserve Uncertain for evidence gaps requiring human review.
  • Route explicit curated-test requests from test-quality-auditor to grade-tests without adding grading to broad audit pipelines or introducing another agent.
  • Update documentation and evals for framework-specific decisions, unresolved methods, missing production context, large reports, and valid scopes with no tests.

Related issue

Fixes #1256

Validation

  • dotnet run --project eng/skill-validator/src/SkillValidator.csproj -- check --plugin ./plugins/dotnet-test — passed: 22 skills, 10 agents, and 1 plugin validated.
  • python eng/eval-quality/check_eval_quality.py — passed with no errors; repository-wide advisory warnings remain.
  • git diff --check — passed.

Checklist

  • I searched existing issues and pull requests to avoid duplicates.
  • I kept this pull request focused and avoided unrelated refactors.
  • I added or updated tests, evals, or documentation when changing skill or agent behavior.
  • I updated CODEOWNERS when adding or moving owned content.
  • I updated all marketplace manifests when plugin metadata changed.
  • I updated eng/known-domains.txt for any new external domains referenced by skill content.

Evangelink and others added 2 commits October 4, 2026 14:35
Co-authored-by: Copilot App <223556219+Copilot@users.noreply.github.com>
Co-authored-by: Copilot App <223556219+Copilot@users.noreply.github.com>
Copilot AI balanced review requested due to automatic review settings October 4, 2026 12:38
@Evangelink
Evangelink requested a review from a team as a code owner October 4, 2026 12:38
@github-actions

github-actions Bot commented Oct 4, 2026 •

Copy link
Copy Markdown
Contributor

Skill Coverage Report

Plugin Skill Covered Coverage
✅ dotnet-test assertion-quality 23/24 95.8%
✅ dotnet-test code-testing-agent 4/4 100%
ℹ️ dotnet-test code-testing-extensions - N/A (reference-only)
✅ dotnet-test crap-score 6/6 100%
✅ dotnet-test detect-static-dependencies 25/25 100%
ℹ️ dotnet-test filter-syntax - N/A (reference-only)
✅ dotnet-test generate-testability-wrappers 28/28 100%
✅ dotnet-test grade-tests 30/34 88.2%
✅ dotnet-test migrate-static-to-wrapper 32/34 94.1%
⚠️ dotnet-test mtp-hot-reload 12/16 75%
ℹ️ dotnet-test test-analysis-extensions - N/A (reference-only)
✅ dotnet-test test-anti-patterns 28/28 100%
✅ dotnet-test test-gap-analysis 10/10 100%
✅ dotnet-test test-tagging 28/31 90.3%
⚠️ dotnet-test testability-obstacle 18/23 78.3%
⚠️ dotnet-test writing-mstest-tests 32/41 78%
Uncovered: dotnet-test/assertion-quality
  • [Validation] Jest matcher semantics are precise (toBeDefined versus undefined; (line 188)
Uncovered: dotnet-test/grade-tests
  • [Validation] Every grade is justified by at least one observable signal in the (line 330)
  • [Pitfall] Inflating deductions to justify the grade (line 354)
  • [Pitfall] Using a fake-precise score (e.g., 87/100) (line 362)
  • [Pitfall] Spilling a 500-row table into a PR comment (line 363)
Uncovered: dotnet-test/migrate-static-to-wrapper
  • [Validation] A before/after exact-member search proves the in-scope occurrence count (line 306)
  • [CodePattern] sealed (line 159)
Uncovered: dotnet-test/mtp-hot-reload
  • [Validation] Microsoft.Testing.Extensions.HotReload package is installed (line 203)
  • [Validation] TESTINGPLATFORM_HOTRELOAD_ENABLED environment variable is set to 1 (line 204)
  • [Validation] Code changes are picked up without manual restart (line 206)
  • [Pitfall] Forgetting to set the environment variable (line 214)
Uncovered: dotnet-test/test-tagging
  • [WorkflowStep] Step 6: Verify edits before reporting (line 316)
  • [CodePattern] [negative] (line 270)
  • [CodePattern] [boundary] (line 270)
Uncovered: dotnet-test/testability-obstacle
  • [Validation] The original obstacle was concrete and in the requested path. (line 250)
  • [Validation] An existing seam was reused when available. (line 251)
  • [Validation] The new abstraction exposes only members required by the target behavior. (line 252)
  • [Validation] The seam did not enlarge the public API when an internal test seam was sufficient. (line 254)
  • [Pitfall] Wrapping an entire static API (line 265)
Uncovered: dotnet-test/writing-mstest-tests
  • [CodePattern] Assert.AreEqual (line 184)
  • [CodePattern] Assert.IsInRange (line 332)
  • [CodePattern] TestDataRow (line 386)
  • [CodePattern] [TestClass] (line 184)
  • [CodePattern] readonly (line 402)
  • [CodePattern] [DataRow] (line 342)
  • [CodePattern] Assert.Contains (line 267)
  • [CodePattern] Assert.IsNotEmpty (line 267)
  • [CodePattern] [TestMethod] (line 184)

@Evangelink
Evangelink enabled auto-merge (squash) October 4, 2026 12:41

Copilot AI left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Copilot review overview

🟡 Changes recommended

The documented output contract is inconsistent, and key grade-independence and aggregation cases lack eval coverage.

Review effort: Balanced
Findings: 2 Medium severity · 2 Low severity

Open (4)
What changed in this PR

Introduces four-state test-quality decisions while retaining A–F grades as supporting detail.

Changes:

  • Defines Pass, Failed, Uncertain, and Not applicable outcomes.
  • Expands eval coverage for decision scenarios and auditor routing.
  • Updates auditor routing and plugin documentation.
File Description
plugins/​dotnet-test/​skills/​grade-tests/​SKILL.md Defines decision and reporting rules.
plugins/​dotnet-test/​agents/​test-quality-auditor.agent.md Routes curated reviews to grade-tests.
plugins/​dotnet-test/​README.md Documents the revised behavior.
tests/​dotnet-test/​grade-tests/​eval.yaml Updates decision-focused evals.
tests/​dotnet-test/​agent.test-quality-auditor/​eval.yaml Tests curated-list routing.

💡 Configure MCP servers for context-aware, tailored reviews. Learn more in the docs.

Comment thread plugins/dotnet-test/skills/grade-tests/SKILL.md
Comment thread tests/dotnet-test/grade-tests/eval.yaml Outdated
Comment thread plugins/dotnet-test/README.md Outdated
Comment thread plugins/dotnet-test/skills/grade-tests/SKILL.md Outdated
Co-authored-by: Copilot App <223556219+Copilot@users.noreply.github.com>
Copilot AI balanced review requested due to automatic review settings October 4, 2026 12:47

@Evangelink Evangelink left a comment

Copy link
Copy Markdown
Member Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

/evaluate

@Evangelink
Evangelink deployed to copilot-pat-pool October 4, 2026 12:50 — with GitHub Actions Active
@Evangelink
Evangelink deployed to copilot-pat-pool October 4, 2026 12:50 — with GitHub Actions Active
@Evangelink
Evangelink deployed to copilot-pat-pool October 4, 2026 12:50 — with GitHub Actions Active
@Evangelink
Evangelink deployed to copilot-pat-pool October 4, 2026 12:50 — with GitHub Actions Active
@Evangelink
Evangelink deployed to copilot-pat-pool October 4, 2026 12:50 — with GitHub Actions Active
@Evangelink
Evangelink deployed to copilot-pat-pool October 4, 2026 12:50 — with GitHub Actions Active
@Evangelink
Evangelink deployed to copilot-pat-pool October 4, 2026 12:50 — with GitHub Actions Active
@Evangelink
Evangelink deployed to copilot-pat-pool October 4, 2026 12:50 — with GitHub Actions Active

Copilot AI left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Copilot review overview

🟡 Changes recommended

Core aggregation and broad-audit exclusion behavior are not fully exercised by the evals.

Review effort: Balanced
Findings: 1 Medium severity

Open (1)
Resolved since last review (4)
Previously missed (2)

In code that hasn't changed since last review

Medium severity Add coverage for Uncertain and Pass aggregation outcomes

tests/​dotnet-test/​grade-tests/​eval.yaml:90

This scenario only proves Failed outranks Uncertain. None of the nine stimuli requires an overall Uncertain result for a Pass+Uncertain set or an overall Pass for an all-clean set, so regressions in the rest of the new aggregation precedence can still pass. Add uncued cases with hard assertions for **Result: Uncertain** and **Result: Pass**.

Low severity Rename zero-finding decision to four-state decision

plugins/​dotnet-test/​agents/​test-quality-auditor.agent.md:93

“Zero-finding” mischaracterizes this decision model, which can report actionable failures and evidence gaps. Call it a “four-state” decision to match the contract used elsewhere.

Comment thread tests/dotnet-test/agent.test-quality-auditor/eval.yaml
@github-actions

github-actions Bot commented Oct 4, 2026

Copy link
Copy Markdown
Contributor

📊 Skill and Agent Evaluation Results

8 model/target results across 4 targets and 2 models — ✅ 2 improved, ➖ 6 results without a clear winner, ⚠️ 0 invalid or underpowered, ⛔ 0 activation contract failures, 📉 0 preference losses (report only).

Measurement identity: evaluated commit 2b95e83d524b06386ef522c9896de09bff63d2a9; 2 judge models.

Measurement health: 8 expected / 8 observed / 8 written; 0 missing, 0 unexpected, 0 invalid; 0 recovered comparison error slots and 0 unresolved comparison error slots.

Objective completion gate: not enabled. Aggregate completion transitions are telemetry only, so this report does not claim that zero objective regressions were proven.

A result passes only when preference-eligible distinct-stimulus votes have aggregate net win of at least 20% and an exact one-sided sign-test result of p ≤ 0.05, and every explicit dormancy activation contract passes. Repeated runs measure reliability only.

Target Model Verdict Gate evidence Overfit Warnings Next action
agent.code-testing-generator claude-sonnet-5 ➖ Baseline signal, unproven n=5; 2W/0T/3L; d=5; p=0.500; net -20.0% — — Evidence leans baseline but is not credible; inspect losing scenarios for recurring defects.
agent.code-testing-generator gpt-5.6-luna ➖ Baseline signal, tie-limited n=5; 1W/2T/2L; d=3; p=0.500; net -20.0% — — Evidence leans baseline but is not credible; inspect losses and make tied tasks discriminate.
agent.test-quality-auditor claude-sonnet-5 ➖ Improvement signal, tie-limited n=6; 3W/2T/1L; d=4; p=0.312; net +33.3%; 1 dormancy excluded — — The signal favors the target, but ties leave too few discordant tasks; inspect ties and predeclare more discriminating breadth.
agent.test-quality-auditor gpt-5.6-luna ➖ Improvement signal, unproven n=6; 4W/0T/2L; d=6; p=0.344; net +33.3%; 1 dormancy excluded — — The signal favors the target but is inconsistent; inspect tied or lost scenarios and fix weak behavior.
agent.testability-migration claude-sonnet-5 ➖ Improvement signal, unproven n=5; 3W/0T/2L; d=5; p=0.500; net +20.0% — — The signal favors the target but is inconsistent; inspect tied or lost scenarios and fix weak behavior.
agent.testability-migration gpt-5.6-luna ➖ Improvement signal, tie-limited n=5; 2W/3T/0L; d=2; p=0.250; net +40.0% — — The signal favors the target, but ties leave too few discordant tasks; inspect ties and predeclare more discriminating breadth.
grade-tests claude-sonnet-5 ✅ Improved n=9; 7W/1T/1L; d=8; p=0.035; net +66.7% 🔴 0.58 — Review overfit evidence.
grade-tests gpt-5.6-luna ✅ Improved n=9; 7W/2T/0L; d=7; p=0.008; net +77.8% 🟡 0.38 — Review overfit evidence.
ℹ️ How to read this report
  • ✅ Improved — the result passed both the statistical gate and the 20% practical net-win floor.
  • ➖ No clear winner — the result is valid but did not pass both gates. The label distinguishes all ties, mixed evidence, directional but unproven evidence, and credible effects below the practical floor.
  • ⚠️ Invalid / underpowered — the gate withheld a quality verdict. Fix the measurement before judging the target.
  • ⛔ Activation contract failed — the isolated target activated on an explicit dormancy scenario. Dormancy preference is excluded, but this routing failure still blocks a pass.
  • 📉 Preference loss — the LLM judge credibly preferred baseline. It is report-only, not objective completion proof.
  • Gate evidence — n preference-eligible distinct-stimulus votes, W/T/L stimulus votes, d discordant votes, exact one-sided p, net win, and the count of separately retained dormancy stimuli. The p value applies to one model/target result; no matrix-wide multiple-comparison correction is applied.
  • Overfit — overfitting-judge severity (✅ Low, 🟡 Moderate, 🔴 High, — none) and score.
  • Warnings — activation, timeout, retry recovery, or unresolved comparison conditions that need attention.
  • Do not add repeated runs to increase statistical power. Do not add stimuli after seeing a near-pass unless the new breadth is predeclared for a new experiment.
➖ Baseline signal, unproven — agent.code-testing-generator (claude-sonnet-5)

Why: Net win -20.0% (2W/0T/3L over 5 preference-eligible stimulus vote(s), sign test p=0.500), mean preference -20.0% across 5 paired run(s) — no improvement

Next action: Evidence leans baseline but is not credible; inspect losing scenarios for recurring defects.

State: VALID_NO_CHANGE (no_credible_preference_change)

Gate evidence: n=5; 2W/0T/3L; d=5; p=0.500; net -20.0%

Repeated-run reliability (not used by the gate): 5 paired runs (2W/0T/3L).

Weak or warning scenarios:

Scenario Preference gate Net win Δ Pref Runs (W/T/L)
▼ Generate a project-wide pytest suite across modules Eligible -100.0% -40.0% 0/0/1
▼ Generate layered Vitest coverage for an async cart Eligible -100.0% -100.0% 0/0/1
▼ Preserve a classic MSTest project while adding broad coverage Eligible -100.0% -40.0% 0/0/1

Illustrative judge evidence:

  • Generate a project-wide pytest suite across modules: Both responses appear to have produced comprehensive suites and honestly disclosed that pytest could not run. A is stronger because its final answer gives substantially more auditable, concrete evidence of requested behavior coverage.

This is one example, not the aggregate verdict. Open Full Results for every judgment.

➖ Baseline signal, tie-limited — agent.code-testing-generator (gpt-5.6-luna)

Why: Net win -20.0% (1W/2T/2L over 5 preference-eligible stimulus vote(s), sign test p=0.500), mean preference -8.0% across 5 paired run(s) — no improvement

Next action: Evidence leans baseline but is not credible; inspect losses and make tied tasks discriminate.

State: VALID_NO_CHANGE (no_credible_preference_change)

Gate evidence: n=5; 1W/2T/2L; d=3; p=0.500; net -20.0%

Repeated-run reliability (not used by the gate): 5 paired runs (1W/2T/2L).

Weak or warning scenarios:

Scenario Preference gate Net win Δ Pref Runs (W/T/L)
▼ Generate collaborating Go package tests Eligible -100.0% -40.0% 0/0/1
= Generate layered Vitest coverage for an async cart Eligible +0.0% +0.0% 0/1/0
▼ Generate project-wide xUnit tests for a .NET library Eligible -100.0% -40.0% 0/0/1
= Preserve a classic MSTest project while adding broad coverage Eligible +0.0% +0.0% 0/1/0

Illustrative judge evidence:

  • Generate collaborating Go package tests: Response A delivered a cleaner, more efficient execution with significantly fewer errors (13 vs 28) and tool calls (34 vs 116). It showed clear views of concrete test implementations and took a direct approach to creating the test files. Response B had more verbose coverage de...

This is one example, not the aggregate verdict. Open Full Results for every judgment.

➖ Improvement signal, tie-limited — agent.test-quality-auditor (claude-sonnet-5)

Why: Net win +33.3% (3W/2T/1L over 6 preference-eligible stimulus vote(s), sign test p=0.312), mean preference +20.0% across 7 paired run(s), 1 dormancy stimulus/stimuli excluded from preference — not credible — 2 of 6 preference-eligible stimulus vote(s) tied, leaving only 4 discordant preference vote(s). The sign test conditions on non-tie stimulus votes and cannot reach 0.05 below 5, so no record could have passed here — this is not a measured null. Either the agent is inert on these scenarios (make them discriminate) or the eval needs more distinct stimuli to clear the ties

Next action: The signal favors the target, but ties leave too few discordant tasks; inspect ties and predeclare more discriminating breadth.

State: VALID_NO_CHANGE (no_credible_preference_change)

Gate evidence: n=6; 3W/2T/1L; d=4; p=0.312; net +33.3%; 1 dormancy excluded

Repeated-run reliability (not used by the gate): 7 paired runs (3W/3T/1L).

Weak or warning scenarios:

Scenario Preference gate Net win Δ Pref Runs (W/T/L)
= Decline request to generate new tests Excluded (activation contract) +0.0% +0.0% 0/1/0
= Diagnose test smells and propose a repair order Eligible +0.0% +0.0% 0/1/0
= Route a curated test list to per-test decisions Eligible +0.0% +0.0% 0/1/0
▼ Targeted anti-pattern review Eligible -100.0% -40.0% 0/0/1

Illustrative judge evidence:

  • Diagnose test smells and propose a repair order: Position-swap inconsistent (forward: skill, reverse: baseline). Defaulting to tie.

This is one example, not the aggregate verdict. Open Full Results for every judgment.

➖ Improvement signal, unproven — agent.test-quality-auditor (gpt-5.6-luna)

Why: Net win +33.3% (4W/0T/2L over 6 preference-eligible stimulus vote(s), sign test p=0.344), mean preference +20.0% across 7 paired run(s), 1 dormancy stimulus/stimuli excluded from preference — not credible (sign test p=0.344 > 0.05)

Next action: The signal favors the target but is inconsistent; inspect tied or lost scenarios and fix weak behavior.

State: VALID_NO_CHANGE (no_credible_preference_change)

Gate evidence: n=6; 4W/0T/2L; d=6; p=0.344; net +33.3%; 1 dormancy excluded

Repeated-run reliability (not used by the gate): 7 paired runs (4W/1T/2L).

Weak or warning scenarios:

Scenario Preference gate Net win Δ Pref Runs (W/T/L)
= Decline request to generate new tests Excluded (activation contract) +0.0% +0.0% 0/1/0
▼ Diagnose test smells and propose a repair order Eligible -100.0% -40.0% 0/0/1
▼ Route a curated test list to per-test decisions Eligible -100.0% -40.0% 0/0/1

Illustrative judge evidence:

  • Diagnose test smells and propose a repair order: Response A is more thorough (8 vs. 5 issues identified), more methodical (repair order fixes false-positives as a coherent block), cleaner execution (0 errors vs. 4), and includes important structural concerns beyond false-positive tests. Both identify the critical reliability...

This is one example, not the aggregate verdict. Open Full Results for every judgment.

➖ Improvement signal, unproven — agent.testability-migration (claude-sonnet-5)

Why: Net win +20.0% (3W/0T/2L over 5 preference-eligible stimulus vote(s), sign test p=0.500), mean preference +20.0% across 5 paired run(s) — not credible (sign test p=0.500 > 0.05)

Next action: The signal favors the target but is inconsistent; inspect tied or lost scenarios and fix weak behavior.

State: VALID_NO_CHANGE (no_credible_preference_change)

Gate evidence: n=5; 3W/0T/2L; d=5; p=0.500; net +20.0%

Repeated-run reliability (not used by the gate): 5 paired runs (3W/0T/2L).

Weak or warning scenarios:

Scenario Preference gate Net win Δ Pref Runs (W/T/L)
▼ Inventory static dependencies without modifying the project Eligible -100.0% -40.0% 0/0/1
▼ Replace filesystem statics without touching unrelated dependencies Eligible -100.0% -40.0% 0/0/1

Illustrative judge evidence:

  • Inventory static dependencies without modifying the project: A is concise, complete, and internally consistent. B adds useful detail, but its contradictory occurrence accounting and incorrect member-count statement reduce confidence in an otherwise correct analysis.

This is one example, not the aggregate verdict. Open Full Results for every judgment.

➖ Improvement signal, tie-limited — agent.testability-migration (gpt-5.6-luna)

Why: Net win +40.0% (2W/3T/0L over 5 preference-eligible stimulus vote(s), sign test p=0.250), mean preference +28.0% across 5 paired run(s) — not credible — 3 of 5 preference-eligible stimulus vote(s) tied, leaving only 2 discordant preference vote(s). The sign test conditions on non-tie stimulus votes and cannot reach 0.05 below 5, so no record could have passed here — this is not a measured null. Either the agent is inert on these scenarios (make them discriminate) or the eval needs more distinct stimuli to clear the ties

Next action: The signal favors the target, but ties leave too few discordant tasks; inspect ties and predeclare more discriminating breadth.

State: VALID_NO_CHANGE (no_credible_preference_change)

Gate evidence: n=5; 2W/3T/0L; d=2; p=0.250; net +40.0%

Repeated-run reliability (not used by the gate): 5 paired runs (2W/3T/0L).

Weak or warning scenarios:

Scenario Preference gate Net win Δ Pref Runs (W/T/L)
= Full pipeline: detect statics and recommend migration plan Eligible +0.0% +0.0% 0/1/0
= Inventory static dependencies without modifying the project Eligible +0.0% +0.0% 0/1/0
= Migrate time dependencies and add deterministic tests Eligible +0.0% +0.0% 0/1/0

Illustrative judge evidence:

  • Full pipeline: detect statics and recommend migration plan: Position-swap inconsistent (forward: skill, reverse: baseline). Defaulting to tie.

This is one example, not the aggregate verdict. Open Full Results for every judgment.

✅ Improved — grade-tests (claude-sonnet-5)

Why: Net win +66.7% (7W/1T/1L over 9 preference-eligible stimulus vote(s), sign test p=0.035), mean preference +46.7% across 27 paired run(s) — credibly better

Next action: Review overfit evidence.

State: VALID_PASS (credible_preference_improvement)

Gate evidence: n=9; 7W/1T/1L; d=8; p=0.035; net +66.7%

Overfit: High (score 0.58)

Repeated-run reliability (not used by the gate): 27 paired runs (20W/2T/5L).

Weak or warning scenarios:

Scenario Preference gate Net win Δ Pref Runs (W/T/L)
= Decide C# test quality against available production code Eligible +0.0% +0.0% 1/1/1
▼ Report an unresolved method as requiring human review Eligible -100.0% -40.0% 0/0/3

Illustrative judge evidence:

  • Decide C# test quality against available production code: The substantive judgments are nearly identical, including the same important mistake of passing the positive-deposit test despite actionable debug output and the same overly harsh F for the self-comparison test. A is slightly better presented: it is a direct compact per-test r...

This is one example, not the aggregate verdict. Open Full Results for every judgment.

✅ Improved — grade-tests (gpt-5.6-luna)

Why: Net win +77.8% (7W/2T/0L over 9 preference-eligible stimulus vote(s), sign test p=0.008), mean preference +54.1% across 27 paired run(s) — credibly better

Next action: Review overfit evidence.

State: VALID_PASS (credible_preference_improvement)

Gate evidence: n=9; 7W/2T/0L; d=7; p=0.008; net +77.8%

Overfit: Moderate (score 0.38)

Repeated-run reliability (not used by the gate): 27 paired runs (20W/4T/3L).

Weak or warning scenarios:

Scenario Preference gate Net win Δ Pref Runs (W/T/L)
= Decide C# test quality against available production code Eligible +0.0% +0.0% 1/1/1
= Decide test quality when production code is unavailable Eligible +0.0% +20.0% 1/1/1

Illustrative judge evidence:

  • Decide C# test quality against available production code: Response A more closely aligns with the task's terminology ('Passes quality review' vs 'Fails: actionable improvement') and provides a more thorough, contextual summary that explicitly addresses all three result categories from the task requirements. Its technical analysis is ...

This is one example, not the aggregate verdict. Open Full Results for every judgment.

🔍 Full Results - all metrics and investigation details

To investigate non-passing or warning results, paste this to your AI coding agent:

For PR 1258 in dotnet/skills, download eval artifacts with gh run download 37203434392 --repo dotnet/skills --pattern "vally-results-*" --dir ./eval-results, then fetch https://raw.githubusercontent.com/dotnet/skills/2b95e83d524b06386ef522c9896de09bff63d2a9/eng/vally-adapter/InvestigatingResults.md and follow it. Classify each result as measurement-invalid, underpowered, not-proven, preference-loss, or passing-with-warning. Use stateReason, result accounting, weak scenarios, and judge evidence to give the cause and exact next fix.

⚠️ Session replay telemetry was not published because the auxiliary dotnet/skills-data publisher failed. The evaluation verdicts above remain authoritative; maintainers must repair SKILLS_DATA_TOKEN or the publisher before replay links are available.

github-actions Bot added a commit that referenced this pull request Oct 4, 2026
Co-authored-by: Copilot App <223556219+Copilot@users.noreply.github.com>
Copilot AI balanced review requested due to automatic review settings October 4, 2026 13:25

@Evangelink Evangelink left a comment

Copy link
Copy Markdown
Member Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

/evaluate

@Evangelink
Evangelink deployed to copilot-pat-pool October 4, 2026 13:27 — with GitHub Actions Active
@Evangelink
Evangelink deployed to copilot-pat-pool October 4, 2026 13:27 — with GitHub Actions Active
@Evangelink
Evangelink deployed to copilot-pat-pool October 4, 2026 13:27 — with GitHub Actions Active
@Evangelink
Evangelink deployed to copilot-pat-pool October 4, 2026 13:27 — with GitHub Actions Active
@Evangelink
Evangelink deployed to copilot-pat-pool October 4, 2026 13:27 — with GitHub Actions Active

Copilot AI left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Copilot review overview

🔵 Needs a closer look

The isolated routing eval omits a required skill dependency, and the Uncertain validation contract remains internally inconsistent.

Review effort: Balanced
Findings: None

Resolved since last review (1)
Previously missed (2)

In code that hasn't changed since last review

Medium severity Stage required test-analysis-extensions skill dependency

tests/​dotnet-test/​agent.test-quality-auditor/​eval.yaml:299

The isolated agent lane stages only the skills listed here, but grade-tests requires test-analysis-extensions before framework-specific scoring (plugins/dotnet-test/skills/grade-tests/SKILL.md:23-29). Without that dependency, this scenario exercises a fallback/non-production path and cannot verify the intended routing end to end. Include the reference skill alongside grade-tests.

Low severity Align validation rule with permitted Uncertain outcomes

plugins/​dotnet-test/​skills/​grade-tests/​SKILL.md:328

This validation rule conflicts with Step 4, which permits Uncertain when a located body uses an unsupported construct or lacks an essential contract. Such a test is resolved in Step 2 but must not be forced to Pass/Failed or given a grade. Define the rule in terms of sufficient evidence instead.

@github-actions

github-actions Bot commented Oct 4, 2026

Copy link
Copy Markdown
Contributor

📊 Skill and Agent Evaluation Results

8 model/target results across 4 targets and 2 models — ✅ 2 improved, ➖ 6 results without a clear winner, ⚠️ 0 invalid or underpowered, ⛔ 0 activation contract failures, 📉 0 preference losses (report only).

Measurement identity: evaluated commit ed146544d70b71784169c1287fab67d4f3004b0c; 2 judge models.

Measurement health: 8 expected / 8 observed / 8 written; 0 missing, 0 unexpected, 0 invalid; 0 recovered comparison error slots and 0 unresolved comparison error slots.

Objective completion gate: not enabled. Aggregate completion transitions are telemetry only, so this report does not claim that zero objective regressions were proven.

A result passes only when preference-eligible distinct-stimulus votes have aggregate net win of at least 20% and an exact one-sided sign-test result of p ≤ 0.05, and every explicit dormancy activation contract passes. Repeated runs measure reliability only.

Target Model Verdict Gate evidence Overfit Warnings Next action
agent.code-testing-generator claude-sonnet-5 ➖ Improvement signal, tie-limited n=5; 2W/3T/0L; d=2; p=0.250; net +40.0% — — The signal favors the target, but ties leave too few discordant tasks; inspect ties and predeclare more discriminating breadth.
agent.code-testing-generator gpt-5.6-luna ➖ Baseline signal, unproven n=5; 1W/0T/4L; d=5; p=0.188; net -60.0% — — Evidence leans baseline but is not credible; inspect losing scenarios for recurring defects.
agent.test-quality-auditor claude-sonnet-5 ➖ Baseline signal, unproven n=6; 2W/0T/4L; d=6; p=0.344; net -33.3%; 1 dormancy excluded — — Evidence leans baseline but is not credible; inspect losing scenarios for recurring defects.
agent.test-quality-auditor gpt-5.6-luna ➖ Mixed evidence n=6; 2W/2T/2L; d=4; p=0.687; net +0.0%; 1 dormancy excluded — — Compare winning and losing scenarios to isolate where the target helps versus hurts.
agent.testability-migration claude-sonnet-5 ➖ Improvement signal, tie-limited n=5; 3W/1T/1L; d=4; p=0.312; net +40.0% — — The signal favors the target, but ties leave too few discordant tasks; inspect ties and predeclare more discriminating breadth.
agent.testability-migration gpt-5.6-luna ➖ Baseline signal, tie-limited n=5; 1W/2T/2L; d=3; p=0.500; net -20.0% — — Evidence leans baseline but is not credible; inspect losses and make tied tasks discriminate.
grade-tests claude-sonnet-5 ✅ Improved n=9; 8W/0T/1L; d=9; p=0.020; net +77.8% 🔴 0.63 — Review overfit evidence.
grade-tests gpt-5.6-luna ✅ Improved n=9; 8W/1T/0L; d=8; p=0.004; net +88.9% 🟡 0.44 — Review overfit evidence.
ℹ️ How to read this report
  • ✅ Improved — the result passed both the statistical gate and the 20% practical net-win floor.
  • ➖ No clear winner — the result is valid but did not pass both gates. The label distinguishes all ties, mixed evidence, directional but unproven evidence, and credible effects below the practical floor.
  • ⚠️ Invalid / underpowered — the gate withheld a quality verdict. Fix the measurement before judging the target.
  • ⛔ Activation contract failed — the isolated target activated on an explicit dormancy scenario. Dormancy preference is excluded, but this routing failure still blocks a pass.
  • 📉 Preference loss — the LLM judge credibly preferred baseline. It is report-only, not objective completion proof.
  • Gate evidence — n preference-eligible distinct-stimulus votes, W/T/L stimulus votes, d discordant votes, exact one-sided p, net win, and the count of separately retained dormancy stimuli. The p value applies to one model/target result; no matrix-wide multiple-comparison correction is applied.
  • Overfit — overfitting-judge severity (✅ Low, 🟡 Moderate, 🔴 High, — none) and score.
  • Warnings — activation, timeout, retry recovery, or unresolved comparison conditions that need attention.
  • Do not add repeated runs to increase statistical power. Do not add stimuli after seeing a near-pass unless the new breadth is predeclared for a new experiment.
➖ Improvement signal, tie-limited — agent.code-testing-generator (claude-sonnet-5)

Why: Net win +40.0% (2W/3T/0L over 5 preference-eligible stimulus vote(s), sign test p=0.250), mean preference +28.0% across 5 paired run(s) — not credible — 3 of 5 preference-eligible stimulus vote(s) tied, leaving only 2 discordant preference vote(s). The sign test conditions on non-tie stimulus votes and cannot reach 0.05 below 5, so no record could have passed here — this is not a measured null. Either the agent is inert on these scenarios (make them discriminate) or the eval needs more distinct stimuli to clear the ties

Next action: The signal favors the target, but ties leave too few discordant tasks; inspect ties and predeclare more discriminating breadth.

State: VALID_NO_CHANGE (no_credible_preference_change)

Gate evidence: n=5; 2W/3T/0L; d=2; p=0.250; net +40.0%

Repeated-run reliability (not used by the gate): 5 paired runs (2W/3T/0L).

Weak or warning scenarios:

Scenario Preference gate Net win Δ Pref Runs (W/T/L)
= Generate a project-wide pytest suite across modules Eligible +0.0% +0.0% 0/1/0
= Generate collaborating Go package tests Eligible +0.0% +0.0% 0/1/0
= Generate project-wide xUnit tests for a .NET library Eligible +0.0% +0.0% 0/1/0

Illustrative judge evidence:

  • Generate a project-wide pytest suite across modules: Position-swap inconsistent (forward: skill, reverse: baseline). Defaulting to tie.

This is one example, not the aggregate verdict. Open Full Results for every judgment.

➖ Baseline signal, unproven — agent.code-testing-generator (gpt-5.6-luna)

Why: Net win -60.0% (1W/0T/4L over 5 preference-eligible stimulus vote(s), sign test p=0.188), mean preference -12.0% across 5 paired run(s) — no improvement

Next action: Evidence leans baseline but is not credible; inspect losing scenarios for recurring defects.

State: VALID_NO_CHANGE (no_credible_preference_change)

Gate evidence: n=5; 1W/0T/4L; d=5; p=0.188; net -60.0%

Repeated-run reliability (not used by the gate): 5 paired runs (1W/0T/4L).

Weak or warning scenarios:

Scenario Preference gate Net win Δ Pref Runs (W/T/L)
▼ Generate collaborating Go package tests Eligible -100.0% -40.0% 0/0/1
▼ Generate layered Vitest coverage for an async cart Eligible -100.0% -40.0% 0/0/1
▼ Generate project-wide xUnit tests for a .NET library Eligible -100.0% -40.0% 0/0/1
▼ Preserve a classic MSTest project while adding broad coverage Eligible -100.0% -40.0% 0/0/1

Illustrative judge evidence:

  • Generate collaborating Go package tests: While Response B created slightly better-organized test function names and potentially more comprehensive coverage structure, it did so at a significant cost in methodology clarity and efficiency. Response A achieved comparable test file generation with fewer tools (7 vs 35+),...

This is one example, not the aggregate verdict. Open Full Results for every judgment.

➖ Baseline signal, unproven — agent.test-quality-auditor (claude-sonnet-5)

Why: Net win -33.3% (2W/0T/4L over 6 preference-eligible stimulus vote(s), sign test p=0.344), mean preference -11.4% across 7 paired run(s), 1 dormancy stimulus/stimuli excluded from preference — no improvement

Next action: Evidence leans baseline but is not credible; inspect losing scenarios for recurring defects.

State: VALID_NO_CHANGE (no_credible_preference_change)

Gate evidence: n=6; 2W/0T/4L; d=6; p=0.344; net -33.3%; 1 dormancy excluded

Repeated-run reliability (not used by the gate): 7 paired runs (2W/1T/4L).

Weak or warning scenarios:

Scenario Preference gate Net win Δ Pref Runs (W/T/L)
▼ Assertion quality analysis Eligible -100.0% -100.0% 0/0/1
= Decline request to generate new tests Excluded (activation contract) +0.0% +0.0% 0/1/0
▼ Diagnose test smells and propose a repair order Eligible -100.0% -40.0% 0/0/1
▼ Identify behavior gaps that existing tests would miss Eligible -100.0% -100.0% 0/0/1
▼ Route a curated test list to per-test decisions Eligible -100.0% -40.0% 0/0/1

Illustrative judge evidence:

  • Assertion quality analysis: A recovered from the blocked broad filesystem search, located and inspected the relevant files, and delivered a concrete, accurate assessment aligned with every rubric criterion. B stopped after the blocked search and gave no substantive answer.

This is one example, not the aggregate verdict. Open Full Results for every judgment.

➖ Mixed evidence — agent.test-quality-auditor (gpt-5.6-luna)

Why: Net win +0.0% (2W/2T/2L over 6 preference-eligible stimulus vote(s), sign test p=0.687), mean preference +14.3% across 7 paired run(s), 1 dormancy stimulus/stimuli excluded from preference — no improvement

Next action: Compare winning and losing scenarios to isolate where the target helps versus hurts.

State: VALID_NO_CHANGE (no_credible_preference_change)

Gate evidence: n=6; 2W/2T/2L; d=4; p=0.687; net +0.0%; 1 dormancy excluded

Repeated-run reliability (not used by the gate): 7 paired runs (3W/2T/2L).

Weak or warning scenarios:

Scenario Preference gate Net win Δ Pref Runs (W/T/L)
= Assertion quality analysis Eligible +0.0% +0.0% 0/1/0
▼ Diagnose test smells and propose a repair order Eligible -100.0% -40.0% 0/0/1
▼ Identify behavior gaps that existing tests would miss Eligible -100.0% -40.0% 0/0/1
= Route a curated test list to per-test decisions Eligible +0.0% +0.0% 0/1/0

Illustrative judge evidence:

  • Assertion quality analysis: Position-swap inconsistent (forward: skill, reverse: tie). Defaulting to tie.

This is one example, not the aggregate verdict. Open Full Results for every judgment.

➖ Improvement signal, tie-limited — agent.testability-migration (claude-sonnet-5)

Why: Net win +40.0% (3W/1T/1L over 5 preference-eligible stimulus vote(s), sign test p=0.312), mean preference +28.0% across 5 paired run(s) — not credible — 1 of 5 preference-eligible stimulus vote(s) tied, leaving only 4 discordant preference vote(s). The sign test conditions on non-tie stimulus votes and cannot reach 0.05 below 5, so no record could have passed here — this is not a measured null. Either the agent is inert on these scenarios (make them discriminate) or the eval needs more distinct stimuli to clear the ties

Next action: The signal favors the target, but ties leave too few discordant tasks; inspect ties and predeclare more discriminating breadth.

State: VALID_NO_CHANGE (no_credible_preference_change)

Gate evidence: n=5; 3W/1T/1L; d=4; p=0.312; net +40.0%

Repeated-run reliability (not used by the gate): 5 paired runs (3W/1T/1L).

Weak or warning scenarios:

Scenario Preference gate Net win Δ Pref Runs (W/T/L)
▼ Full pipeline: detect statics and recommend migration plan Eligible -100.0% -40.0% 0/0/1
= Migrate time dependencies and add deterministic tests Eligible +0.0% +0.0% 0/1/0

Illustrative judge evidence:

  • Full pipeline: detect statics and recommend migration plan: Both responses satisfy the requested analysis-only bounded plan. B is stronger on concrete repository-derived detail and TimeProvider usage, but contains an internal call-site-count error and trends toward a broader, more speculative DI/package migration. A provides the more f...

This is one example, not the aggregate verdict. Open Full Results for every judgment.

➖ Baseline signal, tie-limited — agent.testability-migration (gpt-5.6-luna)

Why: Net win -20.0% (1W/2T/2L over 5 preference-eligible stimulus vote(s), sign test p=0.500), mean preference -8.0% across 5 paired run(s) — no improvement

Next action: Evidence leans baseline but is not credible; inspect losses and make tied tasks discriminate.

State: VALID_NO_CHANGE (no_credible_preference_change)

Gate evidence: n=5; 1W/2T/2L; d=3; p=0.500; net -20.0%

Repeated-run reliability (not used by the gate): 5 paired runs (1W/2T/2L).

Weak or warning scenarios:

Scenario Preference gate Net win Δ Pref Runs (W/T/L)
= Full pipeline: detect statics and recommend migration plan Eligible +0.0% +0.0% 0/1/0
▼ Inventory static dependencies without modifying the project Eligible -100.0% -40.0% 0/0/1
▼ Migrate time dependencies and add deterministic tests Eligible -100.0% -40.0% 0/0/1
= Targeted request: just migrate DateTime to TimeProvider Eligible +0.0% +0.0% 0/1/0

Illustrative judge evidence:

  • Full pipeline: detect statics and recommend migration plan: Position-swap inconsistent (forward: skill, reverse: baseline). Defaulting to tie.

This is one example, not the aggregate verdict. Open Full Results for every judgment.

✅ Improved — grade-tests (claude-sonnet-5)

Why: Net win +77.8% (8W/0T/1L over 9 preference-eligible stimulus vote(s), sign test p=0.020), mean preference +48.9% across 27 paired run(s) — credibly better

Next action: Review overfit evidence.

State: VALID_PASS (credible_preference_improvement)

Gate evidence: n=9; 8W/0T/1L; d=9; p=0.020; net +77.8%

Overfit: High (score 0.63)

Repeated-run reliability (not used by the gate): 27 paired runs (21W/3T/3L).

Weak or warning scenarios:

Scenario Preference gate Net win Δ Pref Runs (W/T/L)
▼ Report an unresolved method as requiring human review Eligible -100.0% -40.0% 0/0/3

Illustrative judge evidence:

  • Report an unresolved method as requiring human review: The reports are substantively aligned and correct. A is modestly stronger because it provides the requested supporting A–F quality detail more completely for both assessable tests, while retaining the same appropriate treatment of the missing transfer test.

This is one example, not the aggregate verdict. Open Full Results for every judgment.

✅ Improved — grade-tests (gpt-5.6-luna)

Why: Net win +88.9% (8W/1T/0L over 9 preference-eligible stimulus vote(s), sign test p=0.004), mean preference +57.8% across 27 paired run(s) — credibly better

Next action: Review overfit evidence.

State: VALID_PASS (credible_preference_improvement)

Gate evidence: n=9; 8W/1T/0L; d=8; p=0.004; net +88.9%

Overfit: Moderate (score 0.44)

Repeated-run reliability (not used by the gate): 27 paired runs (19W/7T/1L).

Weak or warning scenarios:

Scenario Preference gate Net win Δ Pref Runs (W/T/L)
= Decide C# test quality against available production code Eligible +0.0% +20.0% 1/1/1

Illustrative judge evidence:

  • Decide C# test quality against available production code: Both responses perform strong analytical work, correctly identifying all five tests' quality issues. However, Response A delivers a factually accurate and task-compliant result, while Response B contains a critical factual error in its summary (claiming 1 pass and 4 actionable...

This is one example, not the aggregate verdict. Open Full Results for every judgment.

🔍 Full Results - all metrics and investigation details

To investigate non-passing or warning results, paste this to your AI coding agent:

For PR 1258 in dotnet/skills, download eval artifacts with gh run download 37205615576 --repo dotnet/skills --pattern "vally-results-*" --dir ./eval-results, then fetch https://raw.githubusercontent.com/dotnet/skills/ed146544d70b71784169c1287fab67d4f3004b0c/eng/vally-adapter/InvestigatingResults.md and follow it. Classify each result as measurement-invalid, underpowered, not-proven, preference-loss, or passing-with-warning. Use stateReason, result accounting, weak scenarios, and judge evidence to give the cause and exact next fix.

⚠️ Session replay telemetry was not published because the auxiliary dotnet/skills-data publisher failed. The evaluation verdicts above remain authoritative; maintainers must repair SKILLS_DATA_TOKEN or the publisher before replay links are available.

github-actions Bot added a commit that referenced this pull request Oct 4, 2026
@github-actions github-actions Bot added the waiting-on-review PR state label label Oct 4, 2026
@github-actions

github-actions Bot commented Oct 4, 2026

Copy link
Copy Markdown
Contributor

✅ Evaluation passed for ed14654. cc @dotnet/dotnet-testing — please review.

@Evangelink
Evangelink merged commit 4bd1c81 into main Oct 5, 2026
78 checks passed
@Evangelink
Evangelink deleted the dev/amauryleve/grade-tests-agent branch October 5, 2026 14:06

This branch was successfully deployed

1 active deployment
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

waiting-on-review PR state label

Projects

None yet

Development

Successfully merging this pull request may close these issues.

Simplify grade-tests output to actionable decisions

3 participants