You signed in with another tab or window. Reload to refresh your session.You signed out in another tab or window. Reload to refresh your session.You switched accounts on another tab or window. Reload to refresh your session.Dismiss alert
Objective completion gate: not enabled. Aggregate completion transitions are telemetry only, so this report does not claim that zero objective regressions were proven.
A result passes only when preference-eligible distinct-stimulus votes have aggregate net win of at least 20% and an exact one-sided sign-test result of p ≤ 0.05, and every explicit dormancy activation contract passes. Repeated runs measure reliability only.
Target
Model
Verdict
Gate evidence
Overfit
Warnings
Next action
csharp-expert
claude-sonnet-5
⛔ Activation contract failed
n=18; 3W/14T/1L; d=4; p=0.312; net +11.1%; 1 dormancy excluded
Narrow skill routing so the listed off-target scenarios stay dormant.
csharp-expert
gpt-5.6-luna
⛔ Activation contract failed
n=18; 1W/16T/1L; d=2; p=0.750; net +0.0%; 1 dormancy excluded
🟡 0.45
Dormancy contract: 1 unexpected activation(s)
Narrow skill routing so the listed off-target scenarios stay dormant.
ℹ️ How to read this report
✅ Improved — the result passed both the statistical gate and the 20% practical net-win floor.
➖ Not proven improved — the result is valid but did not pass both gates. This is not automatically a regression.
⚠️ Invalid / underpowered — the gate withheld a quality verdict. Fix the measurement before judging the target.
⛔ Activation contract failed — the isolated target activated on an explicit dormancy scenario. Dormancy preference is excluded, but this routing failure still blocks a pass.
📉 Preference loss — the LLM judge credibly preferred baseline. It is report-only, not objective completion proof.
Gate evidence — n preference-eligible distinct-stimulus votes, W/T/L stimulus votes, d discordant votes, exact one-sided p, net win, and the count of separately retained dormancy stimuli. The p value applies to one model/target result; no matrix-wide multiple-comparison correction is applied.
Warnings — activation, timeout, retry recovery, or unresolved comparison conditions that need attention.
Do not add repeated runs to increase statistical power. Do not add stimuli after seeing a near-pass unless the new breadth is predeclared for a new experiment.
Repeated-run reliability (not used by the gate): 19 paired runs (3W/14T/2L).
Weak or warning scenarios:
Scenario
Preference gate
Net win
Δ Pref
Runs (W/T/L)
= Correct percentage arithmetic without adding architecture
Eligible
+0.0%
+0.0%
0/1/0
= Enforce a TryParse style port contract
Eligible
+0.0%
+0.0%
0/1/0
▲ Evaluate a proposed optimization with runtime evidence
Eligible
+100.0%
+40.0%
1/0/0
= Give identifier objects value equality
Eligible
+0.0%
+0.0%
0/1/0
= Handle nullable lookup results explicitly
Eligible
+0.0%
+0.0%
0/1/0
= Keep an invoice bug fix structurally focused
Eligible
+0.0%
+0.0%
0/1/0
▲ Keep deferred file results valid after return
Eligible
+100.0%
+40.0%
1/0/0
▲ Keep returned memory valid after returning a pooled buffer
Eligible
+100.0%
+40.0%
1/0/0
= Leave a simple asynchronous API in its safe form
Eligible
+0.0%
+0.0%
0/1/0
= Leave an already correct cancellation implementation unchanged
Eligible
+0.0%
+0.0%
0/1/0
= Make concurrent increments atomic
Eligible
+0.0%
+0.0%
0/1/0
= Preserve ownership of a caller supplied stream
Eligible
+0.0%
+0.0%
0/1/0
= Propagate cancellation into asynchronous work
Eligible
+0.0%
+0.0%
0/1/0
= Reject invalid values without hiding type evidence
Eligible
+0.0%
+0.0%
0/1/0
= Remove hot path allocations from word counting
Eligible
+0.0%
+0.0%
0/1/0
▼ Repair code within the configured language version
Eligible
-100.0%
-100.0%
0/0/1
▼ Stay dormant for a behavior-preserving rename
Excluded (activation contract)
-100.0%
-40.0%
0/0/1
= Stop timed-out work before returning
Eligible
+0.0%
+0.0%
0/1/0
= Validate behavior on every supported target
Eligible
+0.0%
+0.0%
0/1/0
Illustrative judge evidence:
Correct percentage arithmetic without adding architecture: The resulting code and verification are materially identical: both correctly fix the defect, preserve zero-discount behavior through the formula and passing run, and successfully build and execute the project.
This is one example, not the aggregate verdict. Open Full Results for every judgment.
Repeated-run reliability (not used by the gate): 19 paired runs (1W/16T/2L).
Weak or warning scenarios:
Scenario
Preference gate
Net win
Δ Pref
Runs (W/T/L)
= Correct percentage arithmetic without adding architecture
Eligible
+0.0%
+0.0%
0/1/0
= Enforce a TryParse style port contract
Eligible
+0.0%
+0.0%
0/1/0
= Evaluate a proposed optimization with runtime evidence
Eligible
+0.0%
+0.0%
0/1/0
= Give identifier objects value equality
Eligible
+0.0%
+0.0%
0/1/0
= Handle nullable lookup results explicitly
Eligible
+0.0%
+0.0%
0/1/0
= Keep an invoice bug fix structurally focused
Eligible
+0.0%
+0.0%
0/1/0
▼ Keep deferred file results valid after return
Eligible
-100.0%
-40.0%
0/0/1
= Keep returned memory valid after returning a pooled buffer
Eligible
+0.0%
+0.0%
0/1/0
= Leave a simple asynchronous API in its safe form
Eligible
+0.0%
+0.0%
0/1/0
= Leave an already correct cancellation implementation unchanged
Eligible
+0.0%
+0.0%
0/1/0
= Make concurrent increments atomic
Eligible
+0.0%
+0.0%
0/1/0
= Propagate cancellation into asynchronous work
Eligible
+0.0%
+0.0%
0/1/0
= Reject invalid values without hiding type evidence
Eligible
+0.0%
+0.0%
0/1/0
= Remove hot path allocations from word counting
Eligible
+0.0%
+0.0%
0/1/0
= Repair code within the configured language version
Eligible
+0.0%
+0.0%
0/1/0
▼ Stay dormant for a behavior-preserving rename
Excluded (activation contract)
-100.0%
-40.0%
0/0/1
= Stop timed-out work before returning
Eligible
+0.0%
+0.0%
0/1/0
= Validate behavior on every supported target
Eligible
+0.0%
+0.0%
0/1/0
Illustrative judge evidence:
Correct percentage arithmetic without adding architecture: Both responses correctly identify and fix the identical bug in the ApplyDiscount method, producing the same correct result (180 for 200 with 10% discount) while preserving zero-discount behavior. Both preserve the API contract, avoid unnecessary abstractions, and successfully ...
This is one example, not the aggregate verdict. Open Full Results for every judgment.
To investigate non-passing or warning results, paste this to your AI coding agent:
For PR 1207 in dotnet/skills, download eval artifacts with gh run download 35928345392 --repo dotnet/skills --pattern "vally-results-*" --dir ./eval-results, then fetch https://raw.githubusercontent.com/dotnet/skills/d7b47a51d065e795641539c29d362a35a060e1a4/eng/vally-adapter/InvestigatingResults.md and follow it. Classify each result as measurement-invalid, underpowered, not-proven, preference-loss, or passing-with-warning. Use stateReason, result accounting, weak scenarios, and judge evidence to give the cause and exact next fix.
👋 @webreidi — this PR has 1 unresolved review thread(s). When you're ready, please address the feedback and push an update; the triage bot will pick up the next state automatically. (Add the no-stale label to silence further pings.)
👋 @webreidi — this PR has 1 unresolved review thread(s). When you're ready, please address the feedback and push an update; the triage bot will pick up the next state automatically. (Add the no-stale label to silence further pings.)
The scenario never exercises the caller-cancellation branch it promises to preserve: it always passes CancellationToken.None. An implementation can therefore return timed-out correctly while mishandling a caller token, and still pass the run-command grader despite rubric item 67. Add a second call with a caller token canceled before the timeout, assert OperationCanceledException, and verify that worker state does not change afterward.
👋 @webreidi — this PR has 1 unresolved review thread(s). When you're ready, please address the feedback and push an update; the triage bot will pick up the next state automatically. (Add the no-stale label to silence further pings.)
- Deleted test files for allocation hot path, already correct, API contract, async cancellation, bug fix boundary, concurrent counter, error evidence, focused bug fix, language version, lazy disposal, multi-target validation, nullable flow, performance review, pooled buffer lifetime, simple safe code, stream ownership, timeout cancels work, and value equality.
- Added new project files for ASP.NET minimal, Blazor components, and MAUI binding with basic implementations.
- Cleaned up unused code and ensured consistency across test fixtures.
This verifier only invokes RunWithTimeoutAsync with CancellationToken.None, so it never observes the rubric's caller-cancellation contract. An implementation that catches OperationCanceledException and returns timed-out for caller cancellation would still pass this run; add a second invocation with a pre-canceled or promptly canceled caller token and assert OperationCanceledException (while keeping the existing timeout assertion).
Drop the unrelated textkit.egg-info build output that entered the C# expert PR through merged branch history.
Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
Objective completion gate: not enabled. Aggregate completion transitions are telemetry only, so this report does not claim that zero objective regressions were proven.
A result passes only when preference-eligible distinct-stimulus votes have aggregate net win of at least 20% and an exact one-sided sign-test result of p ≤ 0.05, and every explicit dormancy activation contract passes. Repeated runs measure reliability only.
Target
Model
Verdict
Gate evidence
Overfit
Warnings
Next action
csharp-expert
claude-sonnet-5
➖ Improvement signal, unproven
n=13; 8W/2T/3L; d=11; p=0.113; net +38.5%; 1 dormancy excluded
Inspect activation-only failed runs before rewriting skill content; the model stopped after loading a skill.
ℹ️ How to read this report
✅ Improved — the result passed both the statistical gate and the 20% practical net-win floor.
➖ No clear winner — the result is valid but did not pass both gates. The label distinguishes all ties, mixed evidence, directional but unproven evidence, and credible effects below the practical floor.
⚠️ Invalid / underpowered — the gate withheld a quality verdict. Fix the measurement before judging the target.
⛔ Activation contract failed — the isolated target activated on an explicit dormancy scenario. Dormancy preference is excluded, but this routing failure still blocks a pass.
📉 Preference loss — the LLM judge credibly preferred baseline. It is report-only, not objective completion proof.
Gate evidence — n preference-eligible distinct-stimulus votes, W/T/L stimulus votes, d discordant votes, exact one-sided p, net win, and the count of separately retained dormancy stimuli. The p value applies to one model/target result; no matrix-wide multiple-comparison correction is applied.
Warnings — activation, timeout, retry recovery, or unresolved comparison conditions that need attention.
Do not add repeated runs to increase statistical power. Do not add stimuli after seeing a near-pass unless the new breadth is predeclared for a new experiment.
Why: Net win +38.5% (8W/2T/3L over 13 preference-eligible stimulus vote(s), sign test p=0.113), mean preference +34.3% across 14 paired run(s), 1 dormancy stimulus/stimuli excluded from preference — not credible (sign test p=0.113 > 0.05)
Next action: Inspect activation-only failed runs before rewriting skill content; the model stopped after loading a skill.
Repeated-run reliability (not used by the gate): 14 paired runs (9W/2T/3L).
Weak or warning scenarios:
Scenario
Preference gate
Net win
Δ Pref
Runs (W/T/L)
▼ Keep a language semantic fix in the fallback
Eligible
-100.0%
-40.0%
0/0/1
= Route a framework upgrade
Eligible
+0.0%
+0.0%
0/1/0
= Route an ASP.NET Core endpoint request
Eligible
+0.0%
+0.0%
0/1/0
▲ Route an EF Core query optimization
Eligible
+100.0%
+40.0%
1/0/0
▼ Route an MSBuild binlog failure
Eligible
-100.0%
-40.0%
0/0/1
▼ Sequence distinct template and component phases
Eligible
-100.0%
-40.0%
0/0/1
Illustrative judge evidence:
Keep a language semantic fix in the fallback: Both answers are correct and safely solve the issue. A is marginally better because it is more focused and minimally preserves behavior, whereas B includes opaque marketplace-process wording and an unnecessary null-to-empty behavior change.
This is one example, not the aggregate verdict. Open Full Results for every judgment.
Why: Net win +15.4% (6W/3T/4L over 13 preference-eligible stimulus vote(s), sign test p=0.377), mean preference +7.1% across 14 paired run(s), 1 dormancy stimulus/stimuli excluded from preference — not credible (sign test p=0.377 > 0.05)
Next action: Inspect activation-only failed runs before rewriting skill content; the model stopped after loading a skill.
Repeated-run reliability (not used by the gate): 14 paired runs (6W/3T/5L).
Weak or warning scenarios:
Scenario
Preference gate
Net win
Δ Pref
Runs (W/T/L)
▼ Keep a language semantic fix in the fallback
Eligible
-100.0%
-40.0%
0/0/1
▼ Route a test framework migration
Eligible
-100.0%
-40.0%
0/0/1
▼ Route an ASP.NET Core endpoint request
Eligible
-100.0%
-40.0%
0/0/1
▼ Route an MSBuild binlog failure
Eligible
-100.0%
-40.0%
0/0/1
= Route runtime performance evidence
Eligible
+0.0%
+0.0%
0/1/0
= Sequence distinct template and component phases
Eligible
+0.0%
+0.0%
0/1/0
▼ Stay dormant for a Python request
Excluded (activation contract)
-100.0%
-40.0%
0/0/1
= Use the Codex marketplace flow
Eligible
+0.0%
+0.0%
0/1/0
Illustrative judge evidence:
Keep a language semantic fix in the fallback: Both responses provide the correct technical solution with identical code. However, Response A delivers a clearer explanation of the underlying ownership semantics—explicitly stating that the helper disposes the reader while the caller retains ownership of the stream. This ped...
This is one example, not the aggregate verdict. Open Full Results for every judgment.
To investigate non-passing or warning results, paste this to your AI coding agent:
For PR 1207 in dotnet/skills, download eval artifacts with gh run download 37357574774 --repo dotnet/skills --pattern "vally-results-*" --dir ./eval-results, then fetch https://raw.githubusercontent.com/dotnet/skills/b7267fc09846e481c99584b4cd12da3a7f009e89/eng/vally-adapter/InvestigatingResults.md and follow it. Classify each result as measurement-invalid, underpowered, not-proven, preference-loss, or passing-with-warning. Use stateReason, result accounting, weak scenarios, and judge evidence to give the cause and exact next fix.
⚠️ Session replay telemetry was not published because the auxiliary dotnet/skills-data publisher failed. The evaluation verdicts above remain authoritative; maintainers must repair SKILLS_DATA_TOKEN or the publisher before replay links are available.
Include SDK pin in copied fixtures for reproducible evaluations
tests/dotnet/csharp-expert/global.json:1
This SDK pin is not present in any evaluation workspace: every environment.files entry in tests/dotnet/csharp-expert/eval.yaml copies a fixtures/... directory, and no stimulus copies tests/dotnet/csharp-expert/global.json. The net10.0 fixtures therefore resolve whatever SDK the evaluator happens to provide, so this file does not make the new build/test scenarios reproducible; include the pin in each relevant fixture/environment or move it under the copied fixture root.
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Summary
Add the
csharp-expertskill converted from the existing custom agent, including focused evaluation scenarios and fixtures.This work was split from #1187 so each expert skill can be reviewed independently.
Validation
python eng/eval-quality/check_eval_quality.py(passed in a clean worktree)Checklist