You signed in with another tab or window. Reload to refresh your session.You signed out in another tab or window. Reload to refresh your session.You switched accounts on another tab or window. Reload to refresh your session.Dismiss alert
Make Pass, Failed, Uncertain, and Not applicable the primary grade-tests decisions while retaining A-F quality as supporting detail.
Define Failed as any evidence-backed actionable improvement, independent of the letter grade, and reserve Uncertain for evidence gaps requiring human review.
Route explicit curated-test requests from test-quality-auditor to grade-tests without adding grading to broad audit pipelines or introducing another agent.
Update documentation and evals for framework-specific decisions, unresolved methods, missing production context, large reports, and valid scopes with no tests.
Add coverage for Uncertain and Pass aggregation outcomes
tests/dotnet-test/grade-tests/eval.yaml:90
This scenario only proves Failed outranks Uncertain. None of the nine stimuli requires an overall Uncertain result for a Pass+Uncertain set or an overall Pass for an all-clean set, so regressions in the rest of the new aggregation precedence can still pass. Add uncued cases with hard assertions for **Result: Uncertain** and **Result: Pass**.
Rename zero-finding decision to four-state decision
“Zero-finding” mischaracterizes this decision model, which can report actionable failures and evidence gaps. Call it a “four-state” decision to match the contract used elsewhere.
Objective completion gate: not enabled. Aggregate completion transitions are telemetry only, so this report does not claim that zero objective regressions were proven.
A result passes only when preference-eligible distinct-stimulus votes have aggregate net win of at least 20% and an exact one-sided sign-test result of p ≤ 0.05, and every explicit dormancy activation contract passes. Repeated runs measure reliability only.
Target
Model
Verdict
Gate evidence
Overfit
Warnings
Next action
agent.code-testing-generator
claude-sonnet-5
➖ Baseline signal, unproven
n=5; 2W/0T/3L; d=5; p=0.500; net -20.0%
—
—
Evidence leans baseline but is not credible; inspect losing scenarios for recurring defects.
agent.code-testing-generator
gpt-5.6-luna
➖ Baseline signal, tie-limited
n=5; 1W/2T/2L; d=3; p=0.500; net -20.0%
—
—
Evidence leans baseline but is not credible; inspect losses and make tied tasks discriminate.
agent.test-quality-auditor
claude-sonnet-5
➖ Improvement signal, tie-limited
n=6; 3W/2T/1L; d=4; p=0.312; net +33.3%; 1 dormancy excluded
—
—
The signal favors the target, but ties leave too few discordant tasks; inspect ties and predeclare more discriminating breadth.
agent.test-quality-auditor
gpt-5.6-luna
➖ Improvement signal, unproven
n=6; 4W/0T/2L; d=6; p=0.344; net +33.3%; 1 dormancy excluded
—
—
The signal favors the target but is inconsistent; inspect tied or lost scenarios and fix weak behavior.
agent.testability-migration
claude-sonnet-5
➖ Improvement signal, unproven
n=5; 3W/0T/2L; d=5; p=0.500; net +20.0%
—
—
The signal favors the target but is inconsistent; inspect tied or lost scenarios and fix weak behavior.
agent.testability-migration
gpt-5.6-luna
➖ Improvement signal, tie-limited
n=5; 2W/3T/0L; d=2; p=0.250; net +40.0%
—
—
The signal favors the target, but ties leave too few discordant tasks; inspect ties and predeclare more discriminating breadth.
grade-tests
claude-sonnet-5
✅ Improved
n=9; 7W/1T/1L; d=8; p=0.035; net +66.7%
🔴 0.58
—
Review overfit evidence.
grade-tests
gpt-5.6-luna
✅ Improved
n=9; 7W/2T/0L; d=7; p=0.008; net +77.8%
🟡 0.38
—
Review overfit evidence.
ℹ️ How to read this report
✅ Improved — the result passed both the statistical gate and the 20% practical net-win floor.
➖ No clear winner — the result is valid but did not pass both gates. The label distinguishes all ties, mixed evidence, directional but unproven evidence, and credible effects below the practical floor.
⚠️ Invalid / underpowered — the gate withheld a quality verdict. Fix the measurement before judging the target.
⛔ Activation contract failed — the isolated target activated on an explicit dormancy scenario. Dormancy preference is excluded, but this routing failure still blocks a pass.
📉 Preference loss — the LLM judge credibly preferred baseline. It is report-only, not objective completion proof.
Gate evidence — n preference-eligible distinct-stimulus votes, W/T/L stimulus votes, d discordant votes, exact one-sided p, net win, and the count of separately retained dormancy stimuli. The p value applies to one model/target result; no matrix-wide multiple-comparison correction is applied.
Warnings — activation, timeout, retry recovery, or unresolved comparison conditions that need attention.
Do not add repeated runs to increase statistical power. Do not add stimuli after seeing a near-pass unless the new breadth is predeclared for a new experiment.
Why: Net win -20.0% (2W/0T/3L over 5 preference-eligible stimulus vote(s), sign test p=0.500), mean preference -20.0% across 5 paired run(s) — no improvement
Next action: Evidence leans baseline but is not credible; inspect losing scenarios for recurring defects.
Gate evidence: n=5; 2W/0T/3L; d=5; p=0.500; net -20.0%
Repeated-run reliability (not used by the gate): 5 paired runs (2W/0T/3L).
Weak or warning scenarios:
Scenario
Preference gate
Net win
Δ Pref
Runs (W/T/L)
▼ Generate a project-wide pytest suite across modules
Eligible
-100.0%
-40.0%
0/0/1
▼ Generate layered Vitest coverage for an async cart
Eligible
-100.0%
-100.0%
0/0/1
▼ Preserve a classic MSTest project while adding broad coverage
Eligible
-100.0%
-40.0%
0/0/1
Illustrative judge evidence:
Generate a project-wide pytest suite across modules: Both responses appear to have produced comprehensive suites and honestly disclosed that pytest could not run. A is stronger because its final answer gives substantially more auditable, concrete evidence of requested behavior coverage.
This is one example, not the aggregate verdict. Open Full Results for every judgment.
Why: Net win -20.0% (1W/2T/2L over 5 preference-eligible stimulus vote(s), sign test p=0.500), mean preference -8.0% across 5 paired run(s) — no improvement
Next action: Evidence leans baseline but is not credible; inspect losses and make tied tasks discriminate.
Gate evidence: n=5; 1W/2T/2L; d=3; p=0.500; net -20.0%
Repeated-run reliability (not used by the gate): 5 paired runs (1W/2T/2L).
Weak or warning scenarios:
Scenario
Preference gate
Net win
Δ Pref
Runs (W/T/L)
▼ Generate collaborating Go package tests
Eligible
-100.0%
-40.0%
0/0/1
= Generate layered Vitest coverage for an async cart
Eligible
+0.0%
+0.0%
0/1/0
▼ Generate project-wide xUnit tests for a .NET library
Eligible
-100.0%
-40.0%
0/0/1
= Preserve a classic MSTest project while adding broad coverage
Eligible
+0.0%
+0.0%
0/1/0
Illustrative judge evidence:
Generate collaborating Go package tests: Response A delivered a cleaner, more efficient execution with significantly fewer errors (13 vs 28) and tool calls (34 vs 116). It showed clear views of concrete test implementations and took a direct approach to creating the test files. Response B had more verbose coverage de...
This is one example, not the aggregate verdict. Open Full Results for every judgment.
Why: Net win +33.3% (3W/2T/1L over 6 preference-eligible stimulus vote(s), sign test p=0.312), mean preference +20.0% across 7 paired run(s), 1 dormancy stimulus/stimuli excluded from preference — not credible — 2 of 6 preference-eligible stimulus vote(s) tied, leaving only 4 discordant preference vote(s). The sign test conditions on non-tie stimulus votes and cannot reach 0.05 below 5, so no record could have passed here — this is not a measured null. Either the agent is inert on these scenarios (make them discriminate) or the eval needs more distinct stimuli to clear the ties
Next action: The signal favors the target, but ties leave too few discordant tasks; inspect ties and predeclare more discriminating breadth.
Why: Net win +33.3% (4W/0T/2L over 6 preference-eligible stimulus vote(s), sign test p=0.344), mean preference +20.0% across 7 paired run(s), 1 dormancy stimulus/stimuli excluded from preference — not credible (sign test p=0.344 > 0.05)
Next action: The signal favors the target but is inconsistent; inspect tied or lost scenarios and fix weak behavior.
Repeated-run reliability (not used by the gate): 7 paired runs (4W/1T/2L).
Weak or warning scenarios:
Scenario
Preference gate
Net win
Δ Pref
Runs (W/T/L)
= Decline request to generate new tests
Excluded (activation contract)
+0.0%
+0.0%
0/1/0
▼ Diagnose test smells and propose a repair order
Eligible
-100.0%
-40.0%
0/0/1
▼ Route a curated test list to per-test decisions
Eligible
-100.0%
-40.0%
0/0/1
Illustrative judge evidence:
Diagnose test smells and propose a repair order: Response A is more thorough (8 vs. 5 issues identified), more methodical (repair order fixes false-positives as a coherent block), cleaner execution (0 errors vs. 4), and includes important structural concerns beyond false-positive tests. Both identify the critical reliability...
This is one example, not the aggregate verdict. Open Full Results for every judgment.
Why: Net win +20.0% (3W/0T/2L over 5 preference-eligible stimulus vote(s), sign test p=0.500), mean preference +20.0% across 5 paired run(s) — not credible (sign test p=0.500 > 0.05)
Next action: The signal favors the target but is inconsistent; inspect tied or lost scenarios and fix weak behavior.
Gate evidence: n=5; 3W/0T/2L; d=5; p=0.500; net +20.0%
Repeated-run reliability (not used by the gate): 5 paired runs (3W/0T/2L).
Weak or warning scenarios:
Scenario
Preference gate
Net win
Δ Pref
Runs (W/T/L)
▼ Inventory static dependencies without modifying the project
Eligible
-100.0%
-40.0%
0/0/1
▼ Replace filesystem statics without touching unrelated dependencies
Eligible
-100.0%
-40.0%
0/0/1
Illustrative judge evidence:
Inventory static dependencies without modifying the project: A is concise, complete, and internally consistent. B adds useful detail, but its contradictory occurrence accounting and incorrect member-count statement reduce confidence in an otherwise correct analysis.
This is one example, not the aggregate verdict. Open Full Results for every judgment.
Why: Net win +40.0% (2W/3T/0L over 5 preference-eligible stimulus vote(s), sign test p=0.250), mean preference +28.0% across 5 paired run(s) — not credible — 3 of 5 preference-eligible stimulus vote(s) tied, leaving only 2 discordant preference vote(s). The sign test conditions on non-tie stimulus votes and cannot reach 0.05 below 5, so no record could have passed here — this is not a measured null. Either the agent is inert on these scenarios (make them discriminate) or the eval needs more distinct stimuli to clear the ties
Next action: The signal favors the target, but ties leave too few discordant tasks; inspect ties and predeclare more discriminating breadth.
Gate evidence: n=5; 2W/3T/0L; d=2; p=0.250; net +40.0%
Repeated-run reliability (not used by the gate): 5 paired runs (2W/3T/0L).
Weak or warning scenarios:
Scenario
Preference gate
Net win
Δ Pref
Runs (W/T/L)
= Full pipeline: detect statics and recommend migration plan
Eligible
+0.0%
+0.0%
0/1/0
= Inventory static dependencies without modifying the project
Eligible
+0.0%
+0.0%
0/1/0
= Migrate time dependencies and add deterministic tests
Eligible
+0.0%
+0.0%
0/1/0
Illustrative judge evidence:
Full pipeline: detect statics and recommend migration plan: Position-swap inconsistent (forward: skill, reverse: baseline). Defaulting to tie.
This is one example, not the aggregate verdict. Open Full Results for every judgment.
✅ Improved — grade-tests (claude-sonnet-5)
Why: Net win +66.7% (7W/1T/1L over 9 preference-eligible stimulus vote(s), sign test p=0.035), mean preference +46.7% across 27 paired run(s) — credibly better
Gate evidence: n=9; 7W/1T/1L; d=8; p=0.035; net +66.7%
Overfit: High (score 0.58)
Repeated-run reliability (not used by the gate): 27 paired runs (20W/2T/5L).
Weak or warning scenarios:
Scenario
Preference gate
Net win
Δ Pref
Runs (W/T/L)
= Decide C# test quality against available production code
Eligible
+0.0%
+0.0%
1/1/1
▼ Report an unresolved method as requiring human review
Eligible
-100.0%
-40.0%
0/0/3
Illustrative judge evidence:
Decide C# test quality against available production code: The substantive judgments are nearly identical, including the same important mistake of passing the positive-deposit test despite actionable debug output and the same overly harsh F for the self-comparison test. A is slightly better presented: it is a direct compact per-test r...
This is one example, not the aggregate verdict. Open Full Results for every judgment.
✅ Improved — grade-tests (gpt-5.6-luna)
Why: Net win +77.8% (7W/2T/0L over 9 preference-eligible stimulus vote(s), sign test p=0.008), mean preference +54.1% across 27 paired run(s) — credibly better
Gate evidence: n=9; 7W/2T/0L; d=7; p=0.008; net +77.8%
Overfit: Moderate (score 0.38)
Repeated-run reliability (not used by the gate): 27 paired runs (20W/4T/3L).
Weak or warning scenarios:
Scenario
Preference gate
Net win
Δ Pref
Runs (W/T/L)
= Decide C# test quality against available production code
Eligible
+0.0%
+0.0%
1/1/1
= Decide test quality when production code is unavailable
Eligible
+0.0%
+20.0%
1/1/1
Illustrative judge evidence:
Decide C# test quality against available production code: Response A more closely aligns with the task's terminology ('Passes quality review' vs 'Fails: actionable improvement') and provides a more thorough, contextual summary that explicitly addresses all three result categories from the task requirements. Its technical analysis is ...
This is one example, not the aggregate verdict. Open Full Results for every judgment.
To investigate non-passing or warning results, paste this to your AI coding agent:
For PR 1258 in dotnet/skills, download eval artifacts with gh run download 37203434392 --repo dotnet/skills --pattern "vally-results-*" --dir ./eval-results, then fetch https://raw.githubusercontent.com/dotnet/skills/2b95e83d524b06386ef522c9896de09bff63d2a9/eng/vally-adapter/InvestigatingResults.md and follow it. Classify each result as measurement-invalid, underpowered, not-proven, preference-loss, or passing-with-warning. Use stateReason, result accounting, weak scenarios, and judge evidence to give the cause and exact next fix.
⚠️ Session replay telemetry was not published because the auxiliary dotnet/skills-data publisher failed. The evaluation verdicts above remain authoritative; maintainers must repair SKILLS_DATA_TOKEN or the publisher before replay links are available.
The isolated agent lane stages only the skills listed here, but grade-tests requires test-analysis-extensions before framework-specific scoring (plugins/dotnet-test/skills/grade-tests/SKILL.md:23-29). Without that dependency, this scenario exercises a fallback/non-production path and cannot verify the intended routing end to end. Include the reference skill alongside grade-tests.
Align validation rule with permitted Uncertain outcomes
This validation rule conflicts with Step 4, which permits Uncertain when a located body uses an unsupported construct or lacks an essential contract. Such a test is resolved in Step 2 but must not be forced to Pass/Failed or given a grade. Define the rule in terms of sufficient evidence instead.
Objective completion gate: not enabled. Aggregate completion transitions are telemetry only, so this report does not claim that zero objective regressions were proven.
A result passes only when preference-eligible distinct-stimulus votes have aggregate net win of at least 20% and an exact one-sided sign-test result of p ≤ 0.05, and every explicit dormancy activation contract passes. Repeated runs measure reliability only.
Target
Model
Verdict
Gate evidence
Overfit
Warnings
Next action
agent.code-testing-generator
claude-sonnet-5
➖ Improvement signal, tie-limited
n=5; 2W/3T/0L; d=2; p=0.250; net +40.0%
—
—
The signal favors the target, but ties leave too few discordant tasks; inspect ties and predeclare more discriminating breadth.
agent.code-testing-generator
gpt-5.6-luna
➖ Baseline signal, unproven
n=5; 1W/0T/4L; d=5; p=0.188; net -60.0%
—
—
Evidence leans baseline but is not credible; inspect losing scenarios for recurring defects.
agent.test-quality-auditor
claude-sonnet-5
➖ Baseline signal, unproven
n=6; 2W/0T/4L; d=6; p=0.344; net -33.3%; 1 dormancy excluded
—
—
Evidence leans baseline but is not credible; inspect losing scenarios for recurring defects.
agent.test-quality-auditor
gpt-5.6-luna
➖ Mixed evidence
n=6; 2W/2T/2L; d=4; p=0.687; net +0.0%; 1 dormancy excluded
—
—
Compare winning and losing scenarios to isolate where the target helps versus hurts.
agent.testability-migration
claude-sonnet-5
➖ Improvement signal, tie-limited
n=5; 3W/1T/1L; d=4; p=0.312; net +40.0%
—
—
The signal favors the target, but ties leave too few discordant tasks; inspect ties and predeclare more discriminating breadth.
agent.testability-migration
gpt-5.6-luna
➖ Baseline signal, tie-limited
n=5; 1W/2T/2L; d=3; p=0.500; net -20.0%
—
—
Evidence leans baseline but is not credible; inspect losses and make tied tasks discriminate.
grade-tests
claude-sonnet-5
✅ Improved
n=9; 8W/0T/1L; d=9; p=0.020; net +77.8%
🔴 0.63
—
Review overfit evidence.
grade-tests
gpt-5.6-luna
✅ Improved
n=9; 8W/1T/0L; d=8; p=0.004; net +88.9%
🟡 0.44
—
Review overfit evidence.
ℹ️ How to read this report
✅ Improved — the result passed both the statistical gate and the 20% practical net-win floor.
➖ No clear winner — the result is valid but did not pass both gates. The label distinguishes all ties, mixed evidence, directional but unproven evidence, and credible effects below the practical floor.
⚠️ Invalid / underpowered — the gate withheld a quality verdict. Fix the measurement before judging the target.
⛔ Activation contract failed — the isolated target activated on an explicit dormancy scenario. Dormancy preference is excluded, but this routing failure still blocks a pass.
📉 Preference loss — the LLM judge credibly preferred baseline. It is report-only, not objective completion proof.
Gate evidence — n preference-eligible distinct-stimulus votes, W/T/L stimulus votes, d discordant votes, exact one-sided p, net win, and the count of separately retained dormancy stimuli. The p value applies to one model/target result; no matrix-wide multiple-comparison correction is applied.
Warnings — activation, timeout, retry recovery, or unresolved comparison conditions that need attention.
Do not add repeated runs to increase statistical power. Do not add stimuli after seeing a near-pass unless the new breadth is predeclared for a new experiment.
Why: Net win +40.0% (2W/3T/0L over 5 preference-eligible stimulus vote(s), sign test p=0.250), mean preference +28.0% across 5 paired run(s) — not credible — 3 of 5 preference-eligible stimulus vote(s) tied, leaving only 2 discordant preference vote(s). The sign test conditions on non-tie stimulus votes and cannot reach 0.05 below 5, so no record could have passed here — this is not a measured null. Either the agent is inert on these scenarios (make them discriminate) or the eval needs more distinct stimuli to clear the ties
Next action: The signal favors the target, but ties leave too few discordant tasks; inspect ties and predeclare more discriminating breadth.
Why: Net win -60.0% (1W/0T/4L over 5 preference-eligible stimulus vote(s), sign test p=0.188), mean preference -12.0% across 5 paired run(s) — no improvement
Next action: Evidence leans baseline but is not credible; inspect losing scenarios for recurring defects.
Gate evidence: n=5; 1W/0T/4L; d=5; p=0.188; net -60.0%
Repeated-run reliability (not used by the gate): 5 paired runs (1W/0T/4L).
Weak or warning scenarios:
Scenario
Preference gate
Net win
Δ Pref
Runs (W/T/L)
▼ Generate collaborating Go package tests
Eligible
-100.0%
-40.0%
0/0/1
▼ Generate layered Vitest coverage for an async cart
Eligible
-100.0%
-40.0%
0/0/1
▼ Generate project-wide xUnit tests for a .NET library
Eligible
-100.0%
-40.0%
0/0/1
▼ Preserve a classic MSTest project while adding broad coverage
Eligible
-100.0%
-40.0%
0/0/1
Illustrative judge evidence:
Generate collaborating Go package tests: While Response B created slightly better-organized test function names and potentially more comprehensive coverage structure, it did so at a significant cost in methodology clarity and efficiency. Response A achieved comparable test file generation with fewer tools (7 vs 35+),...
This is one example, not the aggregate verdict. Open Full Results for every judgment.
Why: Net win -33.3% (2W/0T/4L over 6 preference-eligible stimulus vote(s), sign test p=0.344), mean preference -11.4% across 7 paired run(s), 1 dormancy stimulus/stimuli excluded from preference — no improvement
Next action: Evidence leans baseline but is not credible; inspect losing scenarios for recurring defects.
Repeated-run reliability (not used by the gate): 7 paired runs (2W/1T/4L).
Weak or warning scenarios:
Scenario
Preference gate
Net win
Δ Pref
Runs (W/T/L)
▼ Assertion quality analysis
Eligible
-100.0%
-100.0%
0/0/1
= Decline request to generate new tests
Excluded (activation contract)
+0.0%
+0.0%
0/1/0
▼ Diagnose test smells and propose a repair order
Eligible
-100.0%
-40.0%
0/0/1
▼ Identify behavior gaps that existing tests would miss
Eligible
-100.0%
-100.0%
0/0/1
▼ Route a curated test list to per-test decisions
Eligible
-100.0%
-40.0%
0/0/1
Illustrative judge evidence:
Assertion quality analysis: A recovered from the blocked broad filesystem search, located and inspected the relevant files, and delivered a concrete, accurate assessment aligned with every rubric criterion. B stopped after the blocked search and gave no substantive answer.
This is one example, not the aggregate verdict. Open Full Results for every judgment.
Why: Net win +0.0% (2W/2T/2L over 6 preference-eligible stimulus vote(s), sign test p=0.687), mean preference +14.3% across 7 paired run(s), 1 dormancy stimulus/stimuli excluded from preference — no improvement
Next action: Compare winning and losing scenarios to isolate where the target helps versus hurts.
Why: Net win +40.0% (3W/1T/1L over 5 preference-eligible stimulus vote(s), sign test p=0.312), mean preference +28.0% across 5 paired run(s) — not credible — 1 of 5 preference-eligible stimulus vote(s) tied, leaving only 4 discordant preference vote(s). The sign test conditions on non-tie stimulus votes and cannot reach 0.05 below 5, so no record could have passed here — this is not a measured null. Either the agent is inert on these scenarios (make them discriminate) or the eval needs more distinct stimuli to clear the ties
Next action: The signal favors the target, but ties leave too few discordant tasks; inspect ties and predeclare more discriminating breadth.
Gate evidence: n=5; 3W/1T/1L; d=4; p=0.312; net +40.0%
Repeated-run reliability (not used by the gate): 5 paired runs (3W/1T/1L).
Weak or warning scenarios:
Scenario
Preference gate
Net win
Δ Pref
Runs (W/T/L)
▼ Full pipeline: detect statics and recommend migration plan
Eligible
-100.0%
-40.0%
0/0/1
= Migrate time dependencies and add deterministic tests
Eligible
+0.0%
+0.0%
0/1/0
Illustrative judge evidence:
Full pipeline: detect statics and recommend migration plan: Both responses satisfy the requested analysis-only bounded plan. B is stronger on concrete repository-derived detail and TimeProvider usage, but contains an internal call-site-count error and trends toward a broader, more speculative DI/package migration. A provides the more f...
This is one example, not the aggregate verdict. Open Full Results for every judgment.
Why: Net win -20.0% (1W/2T/2L over 5 preference-eligible stimulus vote(s), sign test p=0.500), mean preference -8.0% across 5 paired run(s) — no improvement
Next action: Evidence leans baseline but is not credible; inspect losses and make tied tasks discriminate.
Gate evidence: n=5; 1W/2T/2L; d=3; p=0.500; net -20.0%
Repeated-run reliability (not used by the gate): 5 paired runs (1W/2T/2L).
Weak or warning scenarios:
Scenario
Preference gate
Net win
Δ Pref
Runs (W/T/L)
= Full pipeline: detect statics and recommend migration plan
Eligible
+0.0%
+0.0%
0/1/0
▼ Inventory static dependencies without modifying the project
Eligible
-100.0%
-40.0%
0/0/1
▼ Migrate time dependencies and add deterministic tests
Eligible
-100.0%
-40.0%
0/0/1
= Targeted request: just migrate DateTime to TimeProvider
Eligible
+0.0%
+0.0%
0/1/0
Illustrative judge evidence:
Full pipeline: detect statics and recommend migration plan: Position-swap inconsistent (forward: skill, reverse: baseline). Defaulting to tie.
This is one example, not the aggregate verdict. Open Full Results for every judgment.
✅ Improved — grade-tests (claude-sonnet-5)
Why: Net win +77.8% (8W/0T/1L over 9 preference-eligible stimulus vote(s), sign test p=0.020), mean preference +48.9% across 27 paired run(s) — credibly better
Gate evidence: n=9; 8W/0T/1L; d=9; p=0.020; net +77.8%
Overfit: High (score 0.63)
Repeated-run reliability (not used by the gate): 27 paired runs (21W/3T/3L).
Weak or warning scenarios:
Scenario
Preference gate
Net win
Δ Pref
Runs (W/T/L)
▼ Report an unresolved method as requiring human review
Eligible
-100.0%
-40.0%
0/0/3
Illustrative judge evidence:
Report an unresolved method as requiring human review: The reports are substantively aligned and correct. A is modestly stronger because it provides the requested supporting A–F quality detail more completely for both assessable tests, while retaining the same appropriate treatment of the missing transfer test.
This is one example, not the aggregate verdict. Open Full Results for every judgment.
✅ Improved — grade-tests (gpt-5.6-luna)
Why: Net win +88.9% (8W/1T/0L over 9 preference-eligible stimulus vote(s), sign test p=0.004), mean preference +57.8% across 27 paired run(s) — credibly better
Gate evidence: n=9; 8W/1T/0L; d=8; p=0.004; net +88.9%
Overfit: Moderate (score 0.44)
Repeated-run reliability (not used by the gate): 27 paired runs (19W/7T/1L).
Weak or warning scenarios:
Scenario
Preference gate
Net win
Δ Pref
Runs (W/T/L)
= Decide C# test quality against available production code
Eligible
+0.0%
+20.0%
1/1/1
Illustrative judge evidence:
Decide C# test quality against available production code: Both responses perform strong analytical work, correctly identifying all five tests' quality issues. However, Response A delivers a factually accurate and task-compliant result, while Response B contains a critical factual error in its summary (claiming 1 pass and 4 actionable...
This is one example, not the aggregate verdict. Open Full Results for every judgment.
To investigate non-passing or warning results, paste this to your AI coding agent:
For PR 1258 in dotnet/skills, download eval artifacts with gh run download 37205615576 --repo dotnet/skills --pattern "vally-results-*" --dir ./eval-results, then fetch https://raw.githubusercontent.com/dotnet/skills/ed146544d70b71784169c1287fab67d4f3004b0c/eng/vally-adapter/InvestigatingResults.md and follow it. Classify each result as measurement-invalid, underpowered, not-proven, preference-loss, or passing-with-warning. Use stateReason, result accounting, weak scenarios, and judge evidence to give the cause and exact next fix.
⚠️ Session replay telemetry was not published because the auxiliary dotnet/skills-data publisher failed. The evaluation verdicts above remain authoritative; maintainers must repair SKILLS_DATA_TOKEN or the publisher before replay links are available.
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Summary
Pass,Failed,Uncertain, andNot applicablethe primarygrade-testsdecisions while retaining A-F quality as supporting detail.Failedas any evidence-backed actionable improvement, independent of the letter grade, and reserveUncertainfor evidence gaps requiring human review.test-quality-auditortograde-testswithout adding grading to broad audit pipelines or introducing another agent.Related issue
Fixes #1256
Validation
dotnet run --project eng/skill-validator/src/SkillValidator.csproj -- check --plugin ./plugins/dotnet-test— passed: 22 skills, 10 agents, and 1 plugin validated.python eng/eval-quality/check_eval_quality.py— passed with no errors; repository-wide advisory warnings remain.git diff --check— passed.Checklist
eng/known-domains.txtfor any new external domains referenced by skill content.