You signed in with another tab or window. Reload to refresh your session.You signed out in another tab or window. Reload to refresh your session.You switched accounts on another tab or window. Reload to refresh your session.Dismiss alert
Add the winforms-expert skill as a dedicated dotnet-winforms plugin, including marketplace registration, ownership, evaluation scenarios, fixtures, and validation tooling.
The C# expert skill was split into #1207 so each skill can be reviewed independently.
Validation
python eng/eval-quality/check_eval_quality.py (passed in a clean worktree)
Checklist
I searched existing issues and pull requests to avoid duplicates.
I kept this pull request focused and avoided unrelated refactors.
I added or updated tests, evals, or documentation when changing skill or agent behavior.
I updated CODEOWNERS when adding or moving owned content.
I updated all marketplace manifests for the new plugin.
👋 @webreidi — this PR has 10 unresolved review thread(s). When you're ready, please address the feedback and push an update; the triage bot will pick up the next state automatically. (Add the no-stale label to silence further pings.)
On Unix runners, Path.GetRelativePath returns /-separated paths, but this split only recognizes \\. After the required dotnet build, generated .cs, .props, and .targets files under obj/bin remain in actualPaths, so the no-op and VB preservation checks fail with unexpected-file errors. Split on both separators (or use a platform-independent filter).
Objective completion gate: not enabled. Aggregate completion transitions are telemetry only, so this report does not claim that zero objective regressions were proven.
A result passes only when preference-eligible distinct-stimulus votes have aggregate net win of at least 20% and an exact one-sided sign-test result of p ≤ 0.05, and every explicit dormancy activation contract passes. Repeated runs measure reliability only.
Target
Model
Verdict
Gate evidence
Overfit
Warnings
Next action
winforms-expert
claude-sonnet-5
➖ Not proven improved
n=15; 7W/6T/2L; d=9; p=0.090; net +33.3%; 1 dormancy excluded
🟡 0.48
Activation: isolated 15/15; plugin 14/15
Inspect tied or lost stimuli and fix inconsistent skill behavior.
winforms-expert
gpt-5.6-luna
⛔ Activation contract failed
n=15; 11W/2T/2L; d=13; p=0.011; net +60.0%; 1 dormancy excluded
🔴 0.50
Dormancy contract: 1 unexpected activation(s)
Narrow skill routing so the listed off-target scenarios stay dormant.
ℹ️ How to read this report
✅ Improved — the result passed both the statistical gate and the 20% practical net-win floor.
➖ Not proven improved — the result is valid but did not pass both gates. This is not automatically a regression.
⚠️ Invalid / underpowered — the gate withheld a quality verdict. Fix the measurement before judging the target.
⛔ Activation contract failed — the isolated target activated on an explicit dormancy scenario. Dormancy preference is excluded, but this routing failure still blocks a pass.
📉 Preference loss — the LLM judge credibly preferred baseline. It is report-only, not objective completion proof.
Gate evidence — n preference-eligible distinct-stimulus votes, W/T/L stimulus votes, d discordant votes, exact one-sided p, net win, and the count of separately retained dormancy stimuli. The p value applies to one model/target result; no matrix-wide multiple-comparison correction is applied.
Warnings — activation, timeout, retry recovery, or unresolved comparison conditions that need attention.
Do not add repeated runs to increase statistical power. Do not add stimuli after seeing a near-pass unless the new breadth is predeclared for a new experiment.
Repeated-run reliability (not used by the gate): 16 paired runs (11W/2T/3L).
Weak or warning scenarios:
Scenario
Preference gate
Net win
Δ Pref
Runs (W/T/L)
▼ Add VB application event handling without a new entry point
Eligible
-100.0%
-40.0%
0/0/1
▼ Add a button with a named event handler
Eligible
-100.0%
-40.0%
0/0/1
= Add safe serialization metadata to a custom control
Eligible
+0.0%
+0.0%
0/1/0
= Audit and repair SQL injection risks across a WinForms solution
Eligible
+0.0%
+0.0%
0/1/0
▼ Stay dormant for a WPF command binding request
Excluded (activation contract)
-100.0%
-40.0%
0/0/1
Illustrative judge evidence:
Add VB application event handling without a new entry point: Both responses successfully implement the required functionality (startup restoration and exception logging through ApplicationEvents.vb) and both achieve successful builds. However, Response A executed the implementation cleanly on the first attempt with proper imports includ...
This is one example, not the aggregate verdict. Open Full Results for every judgment.
➖ Not proven improved — winforms-expert (claude-sonnet-5)
Why: Net win +33.3% (7W/6T/2L over 15 preference-eligible stimulus vote(s), sign test p=0.090), mean preference +16.2% across 16 paired run(s), 1 dormancy stimulus/stimuli excluded from preference — not credible (sign test p=0.090 > 0.05)
Next action: Inspect tied or lost stimuli and fix inconsistent skill behavior.
To investigate non-passing or warning results, paste this to your AI coding agent:
For PR 1187 in dotnet/skills, download eval artifacts with gh run download 36753671540 --repo dotnet/skills --pattern "vally-results-*" --dir ./eval-results, then fetch https://raw.githubusercontent.com/dotnet/skills/1a7e1fa6af7cda28c3ab2021056a69ab9a6a0f71/eng/vally-adapter/InvestigatingResults.md and follow it. Classify each result as measurement-invalid, underpowered, not-proven, preference-loss, or passing-with-warning. Use stateReason, result accounting, weak scenarios, and judge evidence to give the cause and exact next fix.
⚠️ Session replay telemetry was not published because the auxiliary dotnet/skills-data publisher failed. The evaluation verdicts above remain authoritative; maintainers must repair SKILLS_DATA_TOKEN or the publisher before replay links are available.
Objective completion gate: not enabled. Aggregate completion transitions are telemetry only, so this report does not claim that zero objective regressions were proven.
A result passes only when preference-eligible distinct-stimulus votes have aggregate net win of at least 20% and an exact one-sided sign-test result of p ≤ 0.05, and every explicit dormancy activation contract passes. Repeated runs measure reliability only.
Target
Model
Verdict
Gate evidence
Overfit
Warnings
Next action
winforms-expert
claude-sonnet-5
✅ Improved
n=15; 8W/7T/0L; d=8; p=0.004; net +53.3%; 1 dormancy excluded
🟡 0.40
Activation: isolated 15/15; plugin 14/15
Fix activation gaps; Review overfit evidence.
winforms-expert
gpt-5.6-luna
✅ Improved
n=15; 9W/6T/0L; d=9; p=0.002; net +60.0%; 1 dormancy excluded
—
—
None.
ℹ️ How to read this report
✅ Improved — the result passed both the statistical gate and the 20% practical net-win floor.
➖ Not proven improved — the result is valid but did not pass both gates. This is not automatically a regression.
⚠️ Invalid / underpowered — the gate withheld a quality verdict. Fix the measurement before judging the target.
⛔ Activation contract failed — the isolated target activated on an explicit dormancy scenario. Dormancy preference is excluded, but this routing failure still blocks a pass.
📉 Preference loss — the LLM judge credibly preferred baseline. It is report-only, not objective completion proof.
Gate evidence — n preference-eligible distinct-stimulus votes, W/T/L stimulus votes, d discordant votes, exact one-sided p, net win, and the count of separately retained dormancy stimuli. The p value applies to one model/target result; no matrix-wide multiple-comparison correction is applied.
Warnings — activation, timeout, retry recovery, or unresolved comparison conditions that need attention.
Do not add repeated runs to increase statistical power. Do not add stimuli after seeing a near-pass unless the new breadth is predeclared for a new experiment.
✅ Improved — winforms-expert (claude-sonnet-5)
Why: Net win +53.3% (8W/7T/0L over 15 preference-eligible stimulus vote(s), sign test p=0.004), mean preference +26.2% across 16 paired run(s), 1 dormancy stimulus/stimuli excluded from preference — credibly better
Next action: Fix activation gaps; Review overfit evidence.
⚠️ Session replay telemetry was not published because the auxiliary dotnet/skills-data publisher failed. The evaluation verdicts above remain authoritative; maintainers must repair SKILLS_DATA_TOKEN or the publisher before replay links are available.
The reason will be displayed to describe this comment to others. Learn more.
Thanks for the updates here. The additional VB coverage and stronger SQL construction checks are good improvements, but I still found several grader issues that can change the eval result independently of solution quality.
The main blocker is the SQL fixture layout. The eval copies the contents of fixtures/TimeTracking to the work-directory root, but the validator and build command now add another TimeTracking/ prefix. I reproduced this with a fully repaired solution; validation failed before it could inspect the code because TimeTracking/TimeTracking/FrmMain.cs does not exist.
I also verified that the SQL validator accepts placeholder and parameter-name mismatches, accepts a repair that loses the original contains-search behavior, and rejects valid typed or derived parameter-value patterns. The async validator is source-order dependent: the new skill-prescribed RefreshAsync extraction passes or fails based only on whether the helper appears before or after the event handler.
Could we correct the harness paths and make both validators follow the actual command or method structure rather than independent text fragments? I would also move the general SQL audit out of this eval, or add corresponding WinForms data-access guidance so the case measures skill lift.
The merge-readiness verdict is request changes, but I am submitting this review with GitHub status Comment as requested.
👋 @webreidi — this PR has 5 unresolved review thread(s). When you're ready, please address the feedback and push an update; the triage bot will pick up the next state automatically. (Add the no-stale label to silence further pings.)
Objective completion gate: not enabled. Aggregate completion transitions are telemetry only, so this report does not claim that zero objective regressions were proven.
A result passes only when preference-eligible distinct-stimulus votes have aggregate net win of at least 20% and an exact one-sided sign-test result of p ≤ 0.05, and every explicit dormancy activation contract passes. Repeated runs measure reliability only.
Target
Model
Verdict
Gate evidence
Overfit
Warnings
Next action
winforms-expert
claude-sonnet-5
➖ Not proven improved
n=15; 5W/9T/1L; d=6; p=0.109; net +26.7%; 1 dormancy excluded
🟡 0.46
Activation: isolated 14/15; plugin 15/15
Inspect tied or lost stimuli and fix inconsistent skill behavior.
winforms-expert
gpt-5.6-luna
✅ Improved
n=15; 10W/5T/0L; d=10; p=0.001; net +66.7%; 1 dormancy excluded
—
—
None.
ℹ️ How to read this report
✅ Improved — the result passed both the statistical gate and the 20% practical net-win floor.
➖ Not proven improved — the result is valid but did not pass both gates. This is not automatically a regression.
⚠️ Invalid / underpowered — the gate withheld a quality verdict. Fix the measurement before judging the target.
⛔ Activation contract failed — the isolated target activated on an explicit dormancy scenario. Dormancy preference is excluded, but this routing failure still blocks a pass.
📉 Preference loss — the LLM judge credibly preferred baseline. It is report-only, not objective completion proof.
Gate evidence — n preference-eligible distinct-stimulus votes, W/T/L stimulus votes, d discordant votes, exact one-sided p, net win, and the count of separately retained dormancy stimuli. The p value applies to one model/target result; no matrix-wide multiple-comparison correction is applied.
Warnings — activation, timeout, retry recovery, or unresolved comparison conditions that need attention.
Do not add repeated runs to increase statistical power. Do not add stimuli after seeing a near-pass unless the new breadth is predeclared for a new experiment.
➖ Not proven improved — winforms-expert (claude-sonnet-5)
Why: Net win +26.7% (5W/9T/1L over 15 preference-eligible stimulus vote(s), sign test p=0.109), mean preference +16.2% across 16 paired run(s), 1 dormancy stimulus/stimuli excluded from preference — not credible (sign test p=0.109 > 0.05)
Next action: Inspect tied or lost stimuli and fix inconsistent skill behavior.
To investigate non-passing or warning results, paste this to your AI coding agent:
For PR 1187 in dotnet/skills, download eval artifacts with gh run download 36892769472 --repo dotnet/skills --pattern "vally-results-*" --dir ./eval-results, then fetch https://raw.githubusercontent.com/dotnet/skills/8f00c36c366cb479d4845672e5252e25abcbaaad/eng/vally-adapter/InvestigatingResults.md and follow it. Classify each result as measurement-invalid, underpowered, not-proven, preference-loss, or passing-with-warning. Use stateReason, result accounting, weak scenarios, and judge evidence to give the cause and exact next fix.
⚠️ Session replay telemetry was not published because the auxiliary dotnet/skills-data publisher failed. The evaluation verdicts above remain authoritative; maintainers must repair SKILLS_DATA_TOKEN or the publisher before replay links are available.
Objective completion gate: not enabled. Aggregate completion transitions are telemetry only, so this report does not claim that zero objective regressions were proven.
A result passes only when preference-eligible distinct-stimulus votes have aggregate net win of at least 20% and an exact one-sided sign-test result of p ≤ 0.05, and every explicit dormancy activation contract passes. Repeated runs measure reliability only.
Target
Model
Verdict
Gate evidence
Overfit
Warnings
Next action
winforms-expert
claude-sonnet-5
✅ Improved
n=15; 11W/3T/1L; d=12; p=0.003; net +66.7%; 1 dormancy excluded
🟡 0.42
—
Review overfit evidence.
winforms-expert
gpt-5.6-luna
✅ Improved
n=15; 11W/3T/1L; d=12; p=0.003; net +66.7%; 1 dormancy excluded
—
—
None.
ℹ️ How to read this report
✅ Improved — the result passed both the statistical gate and the 20% practical net-win floor.
➖ Not proven improved — the result is valid but did not pass both gates. This is not automatically a regression.
⚠️ Invalid / underpowered — the gate withheld a quality verdict. Fix the measurement before judging the target.
⛔ Activation contract failed — the isolated target activated on an explicit dormancy scenario. Dormancy preference is excluded, but this routing failure still blocks a pass.
📉 Preference loss — the LLM judge credibly preferred baseline. It is report-only, not objective completion proof.
Gate evidence — n preference-eligible distinct-stimulus votes, W/T/L stimulus votes, d discordant votes, exact one-sided p, net win, and the count of separately retained dormancy stimuli. The p value applies to one model/target result; no matrix-wide multiple-comparison correction is applied.
Warnings — activation, timeout, retry recovery, or unresolved comparison conditions that need attention.
Do not add repeated runs to increase statistical power. Do not add stimuli after seeing a near-pass unless the new breadth is predeclared for a new experiment.
✅ Improved — winforms-expert (claude-sonnet-5)
Why: Net win +66.7% (11W/3T/1L over 15 preference-eligible stimulus vote(s), sign test p=0.003), mean preference +38.8% across 16 paired run(s), 1 dormancy stimulus/stimuli excluded from preference — credibly better
Repeated-run reliability (not used by the gate): 16 paired runs (12W/3T/1L).
Weak or warning scenarios:
Scenario
Preference gate
Net win
Δ Pref
Runs (W/T/L)
▼ Add VB application event handling without a new entry point
Eligible
-100.0%
-40.0%
0/0/1
= Await background refresh and restore UI state
Eligible
+0.0%
+0.0%
0/1/0
= Localize form and button text through resources
Eligible
+0.0%
+0.0%
0/1/0
= Restore designer ownership of a form timer
Eligible
+0.0%
+0.0%
0/1/0
Illustrative judge evidence:
Add VB application event handling without a new entry point: The implementations appear substantively equivalent and both compile cleanly. B describes a slightly more defensive logging implementation, but A supplies stronger project-specific validation evidence, including the provided scenario and designer-safety validation passing.
This is one example, not the aggregate verdict. Open Full Results for every judgment.
Routine passing details for 1 result are in Full Results.
⚠️ Session replay telemetry was not published because the auxiliary dotnet/skills-data publisher failed. The evaluation verdicts above remain authoritative; maintainers must repair SKILLS_DATA_TOKEN or the publisher before replay links are available.
The reason will be displayed to describe this comment to others. Learn more.
Thanks for the substantial update here. I verified that the earlier harness-path, integrity-hash, placeholder-name, wildcard, async source-order, and cancellation-token issues are now fixed.
There are a few remaining cases where the validators can still give the wrong result. Could we take another look at these as follow-up improvements?
The async validator extracts finally with a lazy regex. A valid finally containing a nested if before _refreshButton.Enabled = true builds successfully but fails validation because the capture stops at the inner }. I included an apply-ready suggestion that uses the existing brace-aware helper.
The SQL binding check accepts any placeholder that carries userSearchText. I bound the real LIKE @search parameter to a constant and the input to a tautological @decoy = @decoy; validation and build both passed, although the search behavior was lost.
The whole-source SQL concatenation check can reject safe parameter-value locals. A derived %term% value assigned to a typed SqlParameter built successfully but was reported as dynamic SQL.
The new C# delimiter scanner still counts braces and semicolons inside strings and comments. This can truncate an otherwise correct method based only on its string contents.
Verbatim SQL strings such as @"SELECT ..." are not recognized by the query extractor.
The SQL audit still does not map to guidance in winforms-expert/SKILL.md. It might be better to add the guidance this case is intended to measure or move the stimulus to a skill that owns SQL security.
The first item has a direct suggestion. The remaining items need a small design choice, so I left the reason and reproduction evidence rather than prescribing one implementation.
Thanks again for working through these validator cases.
- Introduced functions to handle C# statement parsing, including Get-CSharpStatementEnd, Get-CSharpCodeMask, and Get-CSharpAssignmentExpressions for better readability and maintainability.
- Enhanced Get-QueryInfo to support various SQL literal formats and improved regex patterns for detecting SQL commands.
- Added Get-PredicatePlaceholders to identify parameterized column predicates in SQL queries.
- Simplified Get-InitializeComponentBody by utilizing Get-CSharpBlockBody for improved brace handling.
- Removed redundant regex patterns and assertions related to SQL string construction.
- Improved SQL statement end detection in C# by allowing both ';' and ',' as valid statement terminators when no parentheses, braces, or brackets are open.
- Added a new function `Get-VisualBasicCodeMask` to handle Visual Basic code, distinguishing between code, strings, and comments.
- Enhanced `Get-QueryInfo` to support object initializer syntax for `SqlCommand` and improved regex patterns for command text extraction.
- Updated `Test-ExpressionUsesInput` to utilize the new Visual Basic code mask for input validation.
- Refined regex patterns in `validate-winforms.ps1` to accommodate async methods and additional property attributes in Visual Basic.
- Ensured that the parameterless constructor for `MainForm` initializes designer controls correctly.
- Implemented Get-SqlCodeMask function to mask SQL code elements such as strings, comments, and quoted identifiers.
- Added Get-CSharpInterpolationExpressions function to extract and mask interpolation expressions from C# strings.
- Updated Test-ExpressionUsesInput function to include handling of C# interpolation expressions when not using Visual Basic.
- Modified Get-PredicatePlaceholders to utilize the new SQL code masking functionality.
Objective completion gate: not enabled. Aggregate completion transitions are telemetry only, so this report does not claim that zero objective regressions were proven.
A result passes only when preference-eligible distinct-stimulus votes have aggregate net win of at least 20% and an exact one-sided sign-test result of p ≤ 0.05, and every explicit dormancy activation contract passes. Repeated runs measure reliability only.
Target
Model
Verdict
Gate evidence
Overfit
Warnings
Next action
winforms-expert
claude-sonnet-5
✅ Improved
n=15; 11W/4T/0L; d=11; p=0.000; net +73.3%; 1 dormancy excluded
🟡 0.47
Activation: isolated 14/15; plugin 15/15
Fix activation gaps; Review overfit evidence.
winforms-expert
gpt-5.6-luna
✅ Improved
n=15; 14W/1T/0L; d=14; p=0.000; net +93.3%; 1 dormancy excluded
—
—
None.
ℹ️ How to read this report
✅ Improved — the result passed both the statistical gate and the 20% practical net-win floor.
➖ No clear winner — the result is valid but did not pass both gates. The label distinguishes all ties, mixed evidence, directional but unproven evidence, and credible effects below the practical floor.
⚠️ Invalid / underpowered — the gate withheld a quality verdict. Fix the measurement before judging the target.
⛔ Activation contract failed — the isolated target activated on an explicit dormancy scenario. Dormancy preference is excluded, but this routing failure still blocks a pass.
📉 Preference loss — the LLM judge credibly preferred baseline. It is report-only, not objective completion proof.
Gate evidence — n preference-eligible distinct-stimulus votes, W/T/L stimulus votes, d discordant votes, exact one-sided p, net win, and the count of separately retained dormancy stimuli. The p value applies to one model/target result; no matrix-wide multiple-comparison correction is applied.
Warnings — activation, timeout, retry recovery, or unresolved comparison conditions that need attention.
Do not add repeated runs to increase statistical power. Do not add stimuli after seeing a near-pass unless the new breadth is predeclared for a new experiment.
✅ Improved — winforms-expert (claude-sonnet-5)
Why: Net win +73.3% (11W/4T/0L over 15 preference-eligible stimulus vote(s), sign test p=0.000), mean preference +37.5% across 16 paired run(s), 1 dormancy stimulus/stimuli excluded from preference — credibly better
Next action: Fix activation gaps; Review overfit evidence.
⚠️ Session replay telemetry was not published because the auxiliary dotnet/skills-data publisher failed. The evaluation verdicts above remain authoritative; maintainers must repair SKILLS_DATA_TOKEN or the publisher before replay links are available.
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Summary
Add the
winforms-expertskill as a dedicateddotnet-winformsplugin, including marketplace registration, ownership, evaluation scenarios, fixtures, and validation tooling.The C# expert skill was split into #1207 so each skill can be reviewed independently.
Validation
python eng/eval-quality/check_eval_quality.py(passed in a clean worktree)Checklist