Skip to content

Add C# expert skill - #1207

Open
webreidi wants to merge 24 commits into
mainfrom
webreidi/csharp-expert-skill
Open

webreidi wants to merge 24 commits into
mainfrom
webreidi/csharp-expert-skill

Conversation

@webreidi

Copy link
Copy Markdown
Collaborator

Summary

Add the csharp-expert skill converted from the existing custom agent, including focused evaluation scenarios and fixtures.

This work was split from #1187 so each expert skill can be reviewed independently.

Validation

  • python eng/eval-quality/check_eval_quality.py (passed in a clean worktree)

Checklist

  • I searched existing issues and pull requests to avoid duplicates.
  • I kept this pull request focused and avoided unrelated refactors.
  • I added or updated tests, evals, or documentation when changing skill or agent behavior.
  • I updated CODEOWNERS when adding or moving owned content.
  • No plugin marketplace metadata changed.
  • No new external domains were added.

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
Copilot AI lite review requested due to automatic review settings September 23, 2026 22:25
@webreidi webreidi added the waiting-on-review PR state label label Sep 23, 2026
@webreidi webreidi mentioned this pull request Sep 23, 2026
6 tasks done
@github-actions github-actions Bot added pr-state/ready-for-eval PR is mergeable and awaiting evaluation and removed waiting-on-review PR state label labels Sep 23, 2026
@github-actions github-actions Bot added pr-state/evals-in-progress PR evaluations are in progress and removed pr-state/ready-for-eval PR is mergeable and awaiting evaluation labels Sep 23, 2026
@github-actions

github-actions Bot commented Sep 23, 2026 •

Copy link
Copy Markdown
Contributor

Skill Coverage Report

Plugin Skill Covered Coverage
⚠️ dotnet setup-local-sdk 4/8 50%
Uncovered: dotnet/setup-local-sdk
  • [Pitfall] dotnet app.dll wrong runtime (line 386)
  • [CodePattern] [ordered] (line 301)
  • [CodePattern] [pscustomobject] (line 301)
  • [CodePattern] [guid] (line 104)

Copilot AI left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Copilot review overview

🔵 Needs a closer look

Unresolved routing and rubric findings remain, and the supplied approval assessments are mixed.

Review effort: Lite
Findings: 1 Low severity

Open (1)
What changed in this PR

Adds the csharp-expert skill with routing guidance, evaluation scenarios, fixtures, documentation, and ownership updates.

Changes:

  • Adds C# correctness, lifetime, async, concurrency, and performance guidance.
  • Adds 18 focused evaluation fixture projects and SDK configuration.
  • Updates plugin documentation and CODEOWNERS.
File Description
tests/​dotnet/​csharp-expert/​global.json Pins the .NET SDK.
tests/​dotnet/​csharp-expert/​fixtures/​value-equality/​Program.cs Value-equality fixture entry point.
tests/​dotnet/​csharp-expert/​fixtures/​value-equality/​Implementation.cs Value-equality fixture implementation.
tests/​dotnet/​csharp-expert/​fixtures/​value-equality/​Fixture.csproj Value-equality fixture project.
tests/​dotnet/​csharp-expert/​fixtures/​timeout-cancels-work/​Program.cs Timeout cancellation fixture entry point.
tests/​dotnet/​csharp-expert/​fixtures/​timeout-cancels-work/​Implementation.cs Timeout cancellation fixture implementation.
tests/​dotnet/​csharp-expert/​fixtures/​timeout-cancels-work/​Fixture.csproj Timeout cancellation fixture project.
tests/​dotnet/​csharp-expert/​fixtures/​stream-ownership/​Program.cs Stream ownership fixture entry point.
tests/​dotnet/​csharp-expert/​fixtures/​stream-ownership/​Implementation.cs Stream ownership fixture implementation.
tests/​dotnet/​csharp-expert/​fixtures/​stream-ownership/​Fixture.csproj Stream ownership fixture project.
tests/​dotnet/​csharp-expert/​fixtures/​simple-safe-code/​Program.cs Safe-code fixture entry point.
tests/​dotnet/​csharp-expert/​fixtures/​simple-safe-code/​Implementation.cs Safe-code fixture implementation.
tests/​dotnet/​csharp-expert/​fixtures/​simple-safe-code/​Fixture.csproj Safe-code fixture project.
tests/​dotnet/​csharp-expert/​fixtures/​pooled-buffer-lifetime/​Program.cs Pooled-buffer fixture entry point.
tests/​dotnet/​csharp-expert/​fixtures/​pooled-buffer-lifetime/​Implementation.cs Pooled-buffer fixture implementation.
tests/​dotnet/​csharp-expert/​fixtures/​pooled-buffer-lifetime/​Fixture.csproj Pooled-buffer fixture project.
tests/​dotnet/​csharp-expert/​fixtures/​performance-review/​Program.cs Performance fixture entry point.
tests/​dotnet/​csharp-expert/​fixtures/​performance-review/​Implementation.cs Performance fixture implementation.
tests/​dotnet/​csharp-expert/​fixtures/​performance-review/​Fixture.csproj Performance fixture project.
tests/​dotnet/​csharp-expert/​fixtures/​nullable-flow/​Program.cs Nullable-flow fixture entry point.
tests/​dotnet/​csharp-expert/​fixtures/​nullable-flow/​Implementation.cs Nullable-flow fixture implementation.
tests/​dotnet/​csharp-expert/​fixtures/​nullable-flow/​Fixture.csproj Nullable-flow fixture project.
tests/​dotnet/​csharp-expert/​fixtures/​multi-target-validation/​Program.cs Multi-target fixture entry point.
tests/​dotnet/​csharp-expert/​fixtures/​multi-target-validation/​Implementation.cs Multi-target fixture implementation.
tests/​dotnet/​csharp-expert/​fixtures/​multi-target-validation/​Fixture.csproj Multi-target fixture project.
tests/​dotnet/​csharp-expert/​fixtures/​lazy-disposal/​Program.cs Lazy-disposal fixture entry point.
tests/​dotnet/​csharp-expert/​fixtures/​lazy-disposal/​Implementation.cs Lazy-disposal fixture implementation.
tests/​dotnet/​csharp-expert/​fixtures/​lazy-disposal/​Fixture.csproj Lazy-disposal fixture project.
tests/​dotnet/​csharp-expert/​fixtures/​language-version/​Program.cs Language-version fixture entry point.
tests/​dotnet/​csharp-expert/​fixtures/​language-version/​Implementation.cs Language-version fixture implementation.
tests/​dotnet/​csharp-expert/​fixtures/​language-version/​Fixture.csproj Language-version fixture project.
tests/​dotnet/​csharp-expert/​fixtures/​focused-bug-fix/​Program.cs Focused bug-fix fixture entry point.
tests/​dotnet/​csharp-expert/​fixtures/​focused-bug-fix/​Implementation.cs Focused bug-fix fixture implementation.
tests/​dotnet/​csharp-expert/​fixtures/​focused-bug-fix/​Fixture.csproj Focused bug-fix fixture project.
tests/​dotnet/​csharp-expert/​fixtures/​error-evidence/​Program.cs Error-evidence fixture entry point.
tests/​dotnet/​csharp-expert/​fixtures/​error-evidence/​Implementation.cs Error-evidence fixture implementation.
tests/​dotnet/​csharp-expert/​fixtures/​error-evidence/​Fixture.csproj Error-evidence fixture project.
tests/​dotnet/​csharp-expert/​fixtures/​concurrent-counter/​Program.cs Concurrent-counter fixture entry point.
tests/​dotnet/​csharp-expert/​fixtures/​concurrent-counter/​Implementation.cs Concurrent-counter fixture implementation.
tests/​dotnet/​csharp-expert/​fixtures/​concurrent-counter/​Fixture.csproj Concurrent-counter fixture project.
tests/​dotnet/​csharp-expert/​fixtures/​bug-fix-boundary/​Program.cs Bug-fix boundary fixture entry point.
tests/​dotnet/​csharp-expert/​fixtures/​bug-fix-boundary/​Implementation.cs Bug-fix boundary fixture implementation.
tests/​dotnet/​csharp-expert/​fixtures/​bug-fix-boundary/​Fixture.csproj Bug-fix boundary fixture project.
tests/​dotnet/​csharp-expert/​fixtures/​async-cancellation/​Program.cs Async-cancellation fixture entry point.
tests/​dotnet/​csharp-expert/​fixtures/​async-cancellation/​Implementation.cs Async-cancellation fixture implementation.
tests/​dotnet/​csharp-expert/​fixtures/​async-cancellation/​Fixture.csproj Async-cancellation fixture project.
tests/​dotnet/​csharp-expert/​fixtures/​api-contract/​Program.cs API-contract fixture entry point.
tests/​dotnet/​csharp-expert/​fixtures/​api-contract/​Implementation.cs API-contract fixture implementation.
tests/​dotnet/​csharp-expert/​fixtures/​api-contract/​Fixture.csproj API-contract fixture project.
tests/​dotnet/​csharp-expert/​fixtures/​already-correct/​Program.cs Already-correct fixture entry point.
tests/​dotnet/​csharp-expert/​fixtures/​already-correct/​Implementation.cs Already-correct fixture implementation.
tests/​dotnet/​csharp-expert/​fixtures/​already-correct/​Fixture.csproj Already-correct fixture project.
tests/​dotnet/​csharp-expert/​fixtures/​allocation-hot-path/​Program.cs Allocation hot-path fixture entry point.
tests/​dotnet/​csharp-expert/​fixtures/​allocation-hot-path/​Implementation.cs Allocation hot-path fixture implementation.
tests/​dotnet/​csharp-expert/​fixtures/​allocation-hot-path/​Fixture.csproj Allocation hot-path fixture project.
tests/​dotnet/​csharp-expert/​eval.yaml Defines evaluation scenarios and graders.
plugins/​dotnet/​skills/​csharp-expert/​SKILL.md Adds C# expert guidance and routing.
plugins/​dotnet/​README.md Registers the skill.
.github/​CODEOWNERS Assigns ownership.

💡 Configure MCP servers for context-aware, tailored reviews. Learn more in the docs.

Comment thread tests/dotnet/csharp-expert/eval.yaml Outdated
github-actions Bot added a commit that referenced this pull request Sep 23, 2026
@github-actions

Copy link
Copy Markdown
Contributor

📊 Skill and Agent Evaluation Results

2 model/target results across 1 target and 2 models — ✅ 0 improved, ➖ 0 not proven improved, ⚠️ 0 invalid or underpowered, ⛔ 2 activation contract failures, 📉 0 preference losses (report only).

Measurement identity: evaluated commit d7b47a51d065e795641539c29d362a35a060e1a4; 2 judge models.

Measurement health: 2 expected / 2 observed / 2 written; 0 missing, 0 unexpected, 0 invalid; 0 recovered comparison error slots and 0 unresolved comparison error slots.

Objective completion gate: not enabled. Aggregate completion transitions are telemetry only, so this report does not claim that zero objective regressions were proven.

A result passes only when preference-eligible distinct-stimulus votes have aggregate net win of at least 20% and an exact one-sided sign-test result of p ≤ 0.05, and every explicit dormancy activation contract passes. Repeated runs measure reliability only.

Target Model Verdict Gate evidence Overfit Warnings Next action
csharp-expert claude-sonnet-5 ⛔ Activation contract failed n=18; 3W/14T/1L; d=4; p=0.312; net +11.1%; 1 dormancy excluded 🟡 0.29 Dormancy contract: 1 unexpected activation(s); Activation: isolated 4/18; plugin 8/18 Narrow skill routing so the listed off-target scenarios stay dormant.
csharp-expert gpt-5.6-luna ⛔ Activation contract failed n=18; 1W/16T/1L; d=2; p=0.750; net +0.0%; 1 dormancy excluded 🟡 0.45 Dormancy contract: 1 unexpected activation(s) Narrow skill routing so the listed off-target scenarios stay dormant.
ℹ️ How to read this report
  • ✅ Improved — the result passed both the statistical gate and the 20% practical net-win floor.
  • ➖ Not proven improved — the result is valid but did not pass both gates. This is not automatically a regression.
  • ⚠️ Invalid / underpowered — the gate withheld a quality verdict. Fix the measurement before judging the target.
  • ⛔ Activation contract failed — the isolated target activated on an explicit dormancy scenario. Dormancy preference is excluded, but this routing failure still blocks a pass.
  • 📉 Preference loss — the LLM judge credibly preferred baseline. It is report-only, not objective completion proof.
  • Gate evidence — n preference-eligible distinct-stimulus votes, W/T/L stimulus votes, d discordant votes, exact one-sided p, net win, and the count of separately retained dormancy stimuli. The p value applies to one model/target result; no matrix-wide multiple-comparison correction is applied.
  • Overfit — overfitting-judge severity (✅ Low, 🟡 Moderate, 🔴 High, — none) and score.
  • Warnings — activation, timeout, retry recovery, or unresolved comparison conditions that need attention.
  • Do not add repeated runs to increase statistical power. Do not add stimuli after seeing a near-pass unless the new breadth is predeclared for a new experiment.
⛔ Activation contract failed — csharp-expert (claude-sonnet-5)

Why: Net win +11.1% (3W/14T/1L over 18 preference-eligible stimulus vote(s), sign test p=0.312), mean preference -1.1% across 19 paired run(s), 1 dormancy stimulus/stimuli excluded from preference — activation contract failed (1 explicit dormancy scenario(s) activated the isolated target skill)

Next action: Narrow skill routing so the listed off-target scenarios stay dormant.

State: VALID_NO_CHANGE (activation_contract_failed)

Gate evidence: n=18; 3W/14T/1L; d=4; p=0.312; net +11.1%; 1 dormancy excluded

Warnings: Dormancy contract: 1 unexpected activation(s); Activation: isolated 4/18; plugin 8/18

Overfit: Moderate (score 0.29)

Repeated-run reliability (not used by the gate): 19 paired runs (3W/14T/2L).

Weak or warning scenarios:

Scenario Preference gate Net win Δ Pref Runs (W/T/L)
= Correct percentage arithmetic without adding architecture Eligible +0.0% +0.0% 0/1/0
= Enforce a TryParse style port contract Eligible +0.0% +0.0% 0/1/0
▲ Evaluate a proposed optimization with runtime evidence Eligible +100.0% +40.0% 1/0/0
= Give identifier objects value equality Eligible +0.0% +0.0% 0/1/0
= Handle nullable lookup results explicitly Eligible +0.0% +0.0% 0/1/0
= Keep an invoice bug fix structurally focused Eligible +0.0% +0.0% 0/1/0
▲ Keep deferred file results valid after return Eligible +100.0% +40.0% 1/0/0
▲ Keep returned memory valid after returning a pooled buffer Eligible +100.0% +40.0% 1/0/0
= Leave a simple asynchronous API in its safe form Eligible +0.0% +0.0% 0/1/0
= Leave an already correct cancellation implementation unchanged Eligible +0.0% +0.0% 0/1/0
= Make concurrent increments atomic Eligible +0.0% +0.0% 0/1/0
= Preserve ownership of a caller supplied stream Eligible +0.0% +0.0% 0/1/0
= Propagate cancellation into asynchronous work Eligible +0.0% +0.0% 0/1/0
= Reject invalid values without hiding type evidence Eligible +0.0% +0.0% 0/1/0
= Remove hot path allocations from word counting Eligible +0.0% +0.0% 0/1/0
▼ Repair code within the configured language version Eligible -100.0% -100.0% 0/0/1
▼ Stay dormant for a behavior-preserving rename Excluded (activation contract) -100.0% -40.0% 0/0/1
= Stop timed-out work before returning Eligible +0.0% +0.0% 0/1/0
= Validate behavior on every supported target Eligible +0.0% +0.0% 0/1/0

Illustrative judge evidence:

  • Correct percentage arithmetic without adding architecture: The resulting code and verification are materially identical: both correctly fix the defect, preserve zero-discount behavior through the formula and passing run, and successfully build and execute the project.

This is one example, not the aggregate verdict. Open Full Results for every judgment.

⛔ Activation contract failed — csharp-expert (gpt-5.6-luna)

Why: Net win +0.0% (1W/16T/1L over 18 preference-eligible stimulus vote(s), sign test p=0.750), mean preference -2.1% across 19 paired run(s), 1 dormancy stimulus/stimuli excluded from preference — activation contract failed (1 explicit dormancy scenario(s) activated the isolated target skill)

Next action: Narrow skill routing so the listed off-target scenarios stay dormant.

State: VALID_NO_CHANGE (activation_contract_failed)

Gate evidence: n=18; 1W/16T/1L; d=2; p=0.750; net +0.0%; 1 dormancy excluded

Warnings: Dormancy contract: 1 unexpected activation(s)

Overfit: Moderate (score 0.45)

Repeated-run reliability (not used by the gate): 19 paired runs (1W/16T/2L).

Weak or warning scenarios:

Scenario Preference gate Net win Δ Pref Runs (W/T/L)
= Correct percentage arithmetic without adding architecture Eligible +0.0% +0.0% 0/1/0
= Enforce a TryParse style port contract Eligible +0.0% +0.0% 0/1/0
= Evaluate a proposed optimization with runtime evidence Eligible +0.0% +0.0% 0/1/0
= Give identifier objects value equality Eligible +0.0% +0.0% 0/1/0
= Handle nullable lookup results explicitly Eligible +0.0% +0.0% 0/1/0
= Keep an invoice bug fix structurally focused Eligible +0.0% +0.0% 0/1/0
▼ Keep deferred file results valid after return Eligible -100.0% -40.0% 0/0/1
= Keep returned memory valid after returning a pooled buffer Eligible +0.0% +0.0% 0/1/0
= Leave a simple asynchronous API in its safe form Eligible +0.0% +0.0% 0/1/0
= Leave an already correct cancellation implementation unchanged Eligible +0.0% +0.0% 0/1/0
= Make concurrent increments atomic Eligible +0.0% +0.0% 0/1/0
= Propagate cancellation into asynchronous work Eligible +0.0% +0.0% 0/1/0
= Reject invalid values without hiding type evidence Eligible +0.0% +0.0% 0/1/0
= Remove hot path allocations from word counting Eligible +0.0% +0.0% 0/1/0
= Repair code within the configured language version Eligible +0.0% +0.0% 0/1/0
▼ Stay dormant for a behavior-preserving rename Excluded (activation contract) -100.0% -40.0% 0/0/1
= Stop timed-out work before returning Eligible +0.0% +0.0% 0/1/0
= Validate behavior on every supported target Eligible +0.0% +0.0% 0/1/0

Illustrative judge evidence:

  • Correct percentage arithmetic without adding architecture: Both responses correctly identify and fix the identical bug in the ApplyDiscount method, producing the same correct result (180 for 200 with 10% discount) while preserving zero-discount behavior. Both preserve the API contract, avoid unnecessary abstractions, and successfully ...

This is one example, not the aggregate verdict. Open Full Results for every judgment.

🔍 Full Results - all metrics and investigation details

To investigate non-passing or warning results, paste this to your AI coding agent:

For PR 1207 in dotnet/skills, download eval artifacts with gh run download 35928345392 --repo dotnet/skills --pattern "vally-results-*" --dir ./eval-results, then fetch https://raw.githubusercontent.com/dotnet/skills/d7b47a51d065e795641539c29d362a35a060e1a4/eng/vally-adapter/InvestigatingResults.md and follow it. Classify each result as measurement-invalid, underpowered, not-proven, preference-loss, or passing-with-warning. Use stateReason, result accounting, weak scenarios, and judge evidence to give the cause and exact next fix.

@github-actions github-actions Bot added waiting-on-author PR state label and removed pr-state/evals-in-progress PR evaluations are in progress labels Sep 23, 2026
@github-actions

Copy link
Copy Markdown
Contributor

👋 @webreidi — this PR has 1 unresolved review thread(s). When you're ready, please address the feedback and push an update; the triage bot will pick up the next state automatically. (Add the no-stale label to silence further pings.)

@github-actions

Copy link
Copy Markdown
Contributor

👋 @webreidi — this PR has 1 unresolved review thread(s). When you're ready, please address the feedback and push an update; the triage bot will pick up the next state automatically. (Add the no-stale label to silence further pings.)

Added steps for using LSP diagnostics and navigation in project constraints and validation processes.
Copilot AI review requested due to automatic review settings September 29, 2026 16:14

Copilot AI left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Copilot review overview

🔵 Needs a closer look

Moderate findings remain in pooled-buffer lifetime and timeout-cancellation coverage.

Review effort: Lite
Findings: 1 Low severity

Open (1)
Previously missed (1)

In code that hasn't changed since last review

Medium severity Verifier does not test caller cancellation handling

tests/​dotnet/​csharp-expert/​fixtures/​timeout-cancels-work/​Program.cs:4

The scenario never exercises the caller-cancellation branch it promises to preserve: it always passes CancellationToken.None. An implementation can therefore return timed-out correctly while mishandling a caller token, and still pass the run-command grader despite rubric item 67. Add a second call with a caller token canceled before the timeout, assert OperationCanceledException, and verify that worker state does not change afterward.

@github-actions

github-actions Bot commented Oct 1, 2026

Copy link
Copy Markdown
Contributor

👋 @webreidi — this PR has 1 unresolved review thread(s). When you're ready, please address the feedback and push an update; the triage bot will pick up the next state automatically. (Add the no-stale label to silence further pings.)

webreidi and others added 2 commits October 2, 2026 12:43
- Deleted test files for allocation hot path, already correct, API contract, async cancellation, bug fix boundary, concurrent counter, error evidence, focused bug fix, language version, lazy disposal, multi-target validation, nullable flow, performance review, pooled buffer lifetime, simple safe code, stream ownership, timeout cancels work, and value equality.
- Added new project files for ASP.NET minimal, Blazor components, and MAUI binding with basic implementations.
- Cleaned up unused code and ensured consistency across test fixtures.
Co-authored-by: Copilot Autofix powered by AI <175728472+Copilot@users.noreply.github.com>
Copilot AI lite review requested due to automatic review settings October 2, 2026 19:58

Copilot AI left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Copilot review overview

🟡 Changes recommended

The evaluation suite has a critical integrity-digest mismatch and incomplete caller-cancellation coverage.

Review effort: Lite
Findings: 1 High severity

Open (1)
Resolved since last review (1)
Previously missed (1)

In code that hasn't changed since last review

Medium severity Verify caller cancellation is distinct from timeout cancellation

tests/​dotnet/​csharp-expert/​fixtures/​timeout-cancels-work/​Program.cs:4

This verifier only invokes RunWithTimeoutAsync with CancellationToken.None, so it never observes the rubric's caller-cancellation contract. An implementation that catches OperationCanceledException and returns timed-out for caller cancellation would still pass this run; add a second invocation with a pre-canceled or promptly canceled caller token and assert OperationCanceledException (while keeping the existing timeout assertion).

Comment thread tests/dotnet/csharp-expert/eval.yaml Outdated
webreidi and others added 2 commits October 5, 2026 10:35
Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
Copilot AI lite review requested due to automatic review settings October 5, 2026 17:38
@webreidi
webreidi requested a review from a team as a code owner October 5, 2026 17:38

Copilot AI left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Copilot review overview

🟡 Changes recommended

Narrow the skill routing, strengthen evaluation coverage and SDK staging, and remove unrelated generated metadata.

Review effort: Lite
Findings: 2 Low severity

Open (2)
Resolved since last review (1)

Drop the unrelated textkit.egg-info build output that entered the C# expert PR through merged branch history.

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
Copilot AI lite review requested due to automatic review settings October 5, 2026 17:53

Copilot AI left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Copilot review overview

🔵 Needs a closer look

Moderate evaluation gaps remain around skill-tool delegation and SDK pinning in the EF Core scenario.

Review effort: Lite
Findings: None

Resolved since last review (2)

@github-actions github-actions Bot added pr-state/ready-for-eval PR is mergeable and awaiting evaluation pr-state/evals-in-progress PR evaluations are in progress and removed waiting-on-author PR state label pr-state/ready-for-eval PR is mergeable and awaiting evaluation labels Oct 5, 2026
@github-actions

github-actions Bot commented Oct 5, 2026

Copy link
Copy Markdown
Contributor

📊 Skill and Agent Evaluation Results

2 model/target results across 1 target and 2 models — ✅ 0 improved, ➖ 2 results without a clear winner, ⚠️ 0 invalid or underpowered, ⛔ 0 activation contract failures, 📉 0 preference losses (report only).

Measurement identity: evaluated commit b7267fc09846e481c99584b4cd12da3a7f009e89; 2 judge models.

Measurement health: 2 expected / 2 observed / 2 written; 0 missing, 0 unexpected, 0 invalid; 0 recovered comparison error slots and 0 unresolved comparison error slots.

Objective completion gate: not enabled. Aggregate completion transitions are telemetry only, so this report does not claim that zero objective regressions were proven.

A result passes only when preference-eligible distinct-stimulus votes have aggregate net win of at least 20% and an exact one-sided sign-test result of p ≤ 0.05, and every explicit dormancy activation contract passes. Repeated runs measure reliability only.

Target Model Verdict Gate evidence Overfit Warnings Next action
csharp-expert claude-sonnet-5 ➖ Improvement signal, unproven n=13; 8W/2T/3L; d=11; p=0.113; net +38.5%; 1 dormancy excluded 🔴 0.64 Activation: isolated 11/13; plugin 11/13; Activation-only stop: plugin 1 failed run Inspect activation-only failed runs before rewriting skill content; the model stopped after loading a skill.
csharp-expert gpt-5.6-luna ➖ Improvement signal, unproven n=13; 6W/3T/4L; d=10; p=0.377; net +15.4%; 1 dormancy excluded 🔴 0.55 Activation: isolated 13/13; plugin 12/13; Activation-only stop: plugin 1 failed run Inspect activation-only failed runs before rewriting skill content; the model stopped after loading a skill.
ℹ️ How to read this report
  • ✅ Improved — the result passed both the statistical gate and the 20% practical net-win floor.
  • ➖ No clear winner — the result is valid but did not pass both gates. The label distinguishes all ties, mixed evidence, directional but unproven evidence, and credible effects below the practical floor.
  • ⚠️ Invalid / underpowered — the gate withheld a quality verdict. Fix the measurement before judging the target.
  • ⛔ Activation contract failed — the isolated target activated on an explicit dormancy scenario. Dormancy preference is excluded, but this routing failure still blocks a pass.
  • 📉 Preference loss — the LLM judge credibly preferred baseline. It is report-only, not objective completion proof.
  • Gate evidence — n preference-eligible distinct-stimulus votes, W/T/L stimulus votes, d discordant votes, exact one-sided p, net win, and the count of separately retained dormancy stimuli. The p value applies to one model/target result; no matrix-wide multiple-comparison correction is applied.
  • Overfit — overfitting-judge severity (✅ Low, 🟡 Moderate, 🔴 High, — none) and score.
  • Warnings — activation, timeout, retry recovery, or unresolved comparison conditions that need attention.
  • Do not add repeated runs to increase statistical power. Do not add stimuli after seeing a near-pass unless the new breadth is predeclared for a new experiment.
➖ Improvement signal, unproven — csharp-expert (claude-sonnet-5)

Why: Net win +38.5% (8W/2T/3L over 13 preference-eligible stimulus vote(s), sign test p=0.113), mean preference +34.3% across 14 paired run(s), 1 dormancy stimulus/stimuli excluded from preference — not credible (sign test p=0.113 > 0.05)

Next action: Inspect activation-only failed runs before rewriting skill content; the model stopped after loading a skill.

State: VALID_NO_CHANGE (no_credible_preference_change)

Gate evidence: n=13; 8W/2T/3L; d=11; p=0.113; net +38.5%; 1 dormancy excluded

Warnings: Activation: isolated 11/13; plugin 11/13; Activation-only stop: plugin 1 failed run

Overfit: High (score 0.64)

Repeated-run reliability (not used by the gate): 14 paired runs (9W/2T/3L).

Weak or warning scenarios:

Scenario Preference gate Net win Δ Pref Runs (W/T/L)
▼ Keep a language semantic fix in the fallback Eligible -100.0% -40.0% 0/0/1
= Route a framework upgrade Eligible +0.0% +0.0% 0/1/0
= Route an ASP.NET Core endpoint request Eligible +0.0% +0.0% 0/1/0
▲ Route an EF Core query optimization Eligible +100.0% +40.0% 1/0/0
▼ Route an MSBuild binlog failure Eligible -100.0% -40.0% 0/0/1
▼ Sequence distinct template and component phases Eligible -100.0% -40.0% 0/0/1

Illustrative judge evidence:

  • Keep a language semantic fix in the fallback: Both answers are correct and safely solve the issue. A is marginally better because it is more focused and minimally preserves behavior, whereas B includes opaque marketplace-process wording and an unnecessary null-to-empty behavior change.

This is one example, not the aggregate verdict. Open Full Results for every judgment.

➖ Improvement signal, unproven — csharp-expert (gpt-5.6-luna)

Why: Net win +15.4% (6W/3T/4L over 13 preference-eligible stimulus vote(s), sign test p=0.377), mean preference +7.1% across 14 paired run(s), 1 dormancy stimulus/stimuli excluded from preference — not credible (sign test p=0.377 > 0.05)

Next action: Inspect activation-only failed runs before rewriting skill content; the model stopped after loading a skill.

State: VALID_NO_CHANGE (no_credible_preference_change)

Gate evidence: n=13; 6W/3T/4L; d=10; p=0.377; net +15.4%; 1 dormancy excluded

Warnings: Activation: isolated 13/13; plugin 12/13; Activation-only stop: plugin 1 failed run

Overfit: High (score 0.55)

Repeated-run reliability (not used by the gate): 14 paired runs (6W/3T/5L).

Weak or warning scenarios:

Scenario Preference gate Net win Δ Pref Runs (W/T/L)
▼ Keep a language semantic fix in the fallback Eligible -100.0% -40.0% 0/0/1
▼ Route a test framework migration Eligible -100.0% -40.0% 0/0/1
▼ Route an ASP.NET Core endpoint request Eligible -100.0% -40.0% 0/0/1
▼ Route an MSBuild binlog failure Eligible -100.0% -40.0% 0/0/1
= Route runtime performance evidence Eligible +0.0% +0.0% 0/1/0
= Sequence distinct template and component phases Eligible +0.0% +0.0% 0/1/0
▼ Stay dormant for a Python request Excluded (activation contract) -100.0% -40.0% 0/0/1
= Use the Codex marketplace flow Eligible +0.0% +0.0% 0/1/0

Illustrative judge evidence:

  • Keep a language semantic fix in the fallback: Both responses provide the correct technical solution with identical code. However, Response A delivers a clearer explanation of the underlying ownership semantics—explicitly stating that the helper disposes the reader while the caller retains ownership of the stream. This ped...

This is one example, not the aggregate verdict. Open Full Results for every judgment.

🔍 Full Results - all metrics and investigation details

To investigate non-passing or warning results, paste this to your AI coding agent:

For PR 1207 in dotnet/skills, download eval artifacts with gh run download 37357574774 --repo dotnet/skills --pattern "vally-results-*" --dir ./eval-results, then fetch https://raw.githubusercontent.com/dotnet/skills/b7267fc09846e481c99584b4cd12da3a7f009e89/eng/vally-adapter/InvestigatingResults.md and follow it. Classify each result as measurement-invalid, underpowered, not-proven, preference-loss, or passing-with-warning. Use stateReason, result accounting, weak scenarios, and judge evidence to give the cause and exact next fix.

⚠️ Session replay telemetry was not published because the auxiliary dotnet/skills-data publisher failed. The evaluation verdicts above remain authoritative; maintainers must repair SKILLS_DATA_TOKEN or the publisher before replay links are available.

github-actions Bot added a commit that referenced this pull request Oct 5, 2026
Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
Copilot AI lite review requested due to automatic review settings October 5, 2026 20:01

Copilot AI left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Copilot review overview

🔵 Needs a closer look

Two moderate routing issues in SKILL.md remain unresolved.

Review effort: Lite
Findings: None

Copilot AI lite review requested due to automatic review settings October 5, 2026 20:35

Copilot AI left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Note

Copilot was unable to run its full agentic suite in this review.

Copilot review overview

Review effort: Lite
Findings: 2 Medium severity

Open (2)

Comment thread plugins/dotnet/skills/csharp-expert/references/dotnet-skills-marketplace.md Outdated
Comment thread plugins/dotnet/skills/csharp-expert/references/dotnet-skills-marketplace.md Outdated
Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
Copilot AI lite review requested due to automatic review settings October 5, 2026 21:10

Copilot AI left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Copilot review overview

🟡 Changes recommended

Narrow the skill’s routing and invocation rules, and ensure the SDK pin is included in evaluation workspaces.

Review effort: Lite
Findings: 1 High severity

Open (1)
Resolved since last review (2)
Previously missed (1)

In code that hasn't changed since last review

Medium severity Include SDK pin in copied fixtures for reproducible evaluations

tests/​dotnet/​csharp-expert/​global.json:1

This SDK pin is not present in any evaluation workspace: every environment.files entry in tests/dotnet/csharp-expert/eval.yaml copies a fixtures/... directory, and no stimulus copies tests/dotnet/csharp-expert/global.json. The net10.0 fixtures therefore resolve whatever SDK the evaluator happens to provide, so this file does not make the new build/test scenarios reproducible; include the pin in each relevant fixture/environment or move it under the copied fixture root.

Comment on lines +3 to +12
description: >-
Route C# and .NET requests to the most specific skill by combining the user's intent with solution
evidence, and help obtain missing specialists from the dotnet/skills plugin marketplace. USE FOR:
any C#/.NET task where ASP.NET Core, Blazor, MAUI, WinForms, EF Core, testing, MSBuild, NuGet,
diagnostics, upgrades, interop, performance, templates, AI, or C# semantics may determine the
workflow; ambiguous prompts such as "fix this", "upgrade", "make it faster", or "write tests";
identifying which dotnet-agent-skills plugin to install; one-file C# apps with no project;
installed plugins missing from `/skills`; and mixed .NET solutions needing multiple specialists.
DO NOT USE FOR: requests clearly unrelated to C# or .NET, or when the user explicitly selected an
available specialist skill and no routing or marketplace decision remains.
@github-actions github-actions Bot added waiting-on-author PR state label and removed pr-state/evals-in-progress PR evaluations are in progress labels Oct 5, 2026

This branch was successfully deployed

1 active (outdated) deployment
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

waiting-on-author PR state label

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants