You signed in with another tab or window. Reload to refresh your session.You signed out in another tab or window. Reload to refresh your session.You switched accounts on another tab or window. Reload to refresh your session.Dismiss alert
This stacked change sharpens 14 MSBuild skill guides on top of #1214. It fixes command syntax, corrects build-parallelism guidance, and defines clearer routing boundaries among shared-file organization, property changes, imports and hooks, F# project ordering, generated files, item operations, and performance workflows.
Why
The PR1 scenarios exposed three guidance groups that needed correction: reliable PowerShell and file-logger syntax; accurate node and dependency advice; and routing descriptions that separate sibling skills by the requested outcome.
The final boundaries are explicit. directory-build-organization owns existing Directory.Build.* discovery, hierarchy/import timing, and props-versus-targets relocation, while property-patterns owns concrete value changes that stay within the current layout. extension-points owns NuGet auto-import and packed-layout discovery, while msbuild-antipatterns owns proven unsafe imports during audits and directly claims .fs, .fsi, and FS0039 compile-order tasks.
Impact
Agents receive commands that preserve literal braces and semicolon-delimited logger arguments, avoid the incorrect claim that plain dotnet build is sequential, and avoid deleting valid ProjectReference edges. Routing now keeps lone-project shared-file introduction dormant without excluding diagnosis of an existing one-project Directory.Build.* timing defect.
All descriptions remain below 1,024 characters. The dotnet-msbuild skill-menu description total is 10,943 characters, still 1,256 characters below the stacked base.
Validation
skill-validator check --plugin ./plugins/dotnet-msbuild: passed for 18 skills and 3 agents
python eng/eval-quality/check_eval_quality.py: passed with no errors
Vally lint: passed for all 14 affected skill/eval pairs
Static routing contract probes: passed for hierarchy-clobber, TargetFramework timing, lone-project dormancy, all three F# scenarios, and reciprocal NuGet import boundaries
Description limits: all plugin descriptions remain below 1,024 characters
👋 @AbhitejJohn — this PR has 1 unresolved review thread(s). When you're ready, please address the feedback and push an update; the triage bot will pick up the next state automatically. (Add the no-stale label to silence further pings.)
👋 @AbhitejJohn — this PR has 1 unresolved review thread(s). When you're ready, please address the feedback and push an update; the triage bot will pick up the next state automatically. (Add the no-stale label to silence further pings.)
Objective completion gate: not enabled. Aggregate completion transitions are telemetry only, so this report does not claim that zero objective regressions were proven.
A result passes only when preference-eligible distinct-stimulus votes have aggregate net win of at least 20% and an exact one-sided sign-test result of p ≤ 0.05, and every explicit dormancy activation contract passes. Repeated runs measure reliability only.
Target
Model
Verdict
Gate evidence
Overfit
Warnings
Next action
agent.msbuild
claude-sonnet-5
➖ Improvement signal, unproven
n=5; 4W/0T/1L; d=5; p=0.188; net +60.0%
—
—
The signal favors the target but is inconsistent; inspect tied or lost scenarios and fix weak behavior.
agent.msbuild
gpt-5.6-luna
➖ Improvement signal, tie-limited
n=5; 3W/1T/1L; d=4; p=0.312; net +40.0%
—
—
The signal favors the target, but ties leave too few discordant tasks; inspect ties and predeclare more discriminating breadth.
binlog-failure-analysis
claude-sonnet-5
➖ Baseline signal, tie-limited
n=6; 1W/3T/2L; d=3; p=0.500; net -16.7%; 1 dormancy excluded
🟡 0.37
Activation: isolated 6/7; plugin 5/7
Evidence leans baseline but is not credible; inspect losses and make tied tasks discriminate.
binlog-failure-analysis
gpt-5.6-luna
➖ Improvement signal, tie-limited
n=6; 4W/2T/0L; d=4; p=0.063; net +66.7%; 1 dormancy excluded
🟡 0.31
Activation: isolated 6/7; plugin 6/7
The signal favors the target, but ties leave too few discordant tasks; inspect ties and predeclare more discriminating breadth.
binlog-generation
claude-sonnet-5
✅ Improved
n=7; 5W/2T/0L; d=5; p=0.031; net +71.4%; 1 dormancy excluded
🔴 0.51
Activation: isolated 5/7; plugin 5/7
Fix activation gaps; Review overfit evidence.
binlog-generation
gpt-5.6-luna
➖ Improvement signal, unproven
n=7; 4W/2T/1L; d=5; p=0.188; net +42.9%; 1 dormancy excluded
🟡 0.39
Activation: isolated 7/7; plugin 6/7
The signal favors the target but is inconsistent; inspect tied or lost scenarios and fix weak behavior.
build-parallelism
claude-sonnet-5
➖ Improvement signal, tie-limited
n=6; 4W/2T/0L; d=4; p=0.063; net +66.7%; 1 dormancy excluded
Narrow skill routing so the listed off-target scenarios stay dormant.
build-perf-diagnostics
claude-sonnet-5
➖ Improvement signal, unproven
n=7; 4W/1T/2L; d=6; p=0.344; net +28.6%; 1 dormancy excluded
🟡 0.43
Activation: isolated 7/7; plugin 2/7
The signal favors the target but is inconsistent; inspect tied or lost scenarios and fix weak behavior.
build-perf-diagnostics
gpt-5.6-luna
➖ Improvement signal, unproven
n=7; 5W/0T/2L; d=7; p=0.227; net +42.9%; 1 dormancy excluded
🟡 0.31
Activation: isolated 7/7; plugin 3/7
The signal favors the target but is inconsistent; inspect tied or lost scenarios and fix weak behavior.
check-bin-obj-clash
claude-sonnet-5
➖ Mixed evidence
n=7; 1W/5T/1L; d=2; p=0.750; net +0.0%
🟡 0.23
Activation: isolated 7/7; plugin 6/7
Compare winning and losing scenarios to isolate where the target helps versus hurts.
check-bin-obj-clash
gpt-5.6-luna
➖ Improvement signal, tie-limited
n=7; 3W/4T/0L; d=3; p=0.125; net +42.9%
🟡 0.20
—
The signal favors the target, but ties leave too few discordant tasks; inspect ties and predeclare more discriminating breadth.
directory-build-organization
claude-sonnet-5
➖ Baseline signal, unproven
n=6; 2W/1T/3L; d=5; p=0.500; net -16.7%; 1 dormancy excluded
🟡 0.37
Activation: isolated 4/6; plugin 4/6
Evidence leans baseline but is not credible; inspect losing scenarios for recurring defects.
directory-build-organization
gpt-5.6-luna
➖ Improvement signal, tie-limited
n=6; 3W/2T/1L; d=4; p=0.312; net +33.3%; 1 dormancy excluded
🟡 0.35
Activation: isolated 6/6; plugin 4/6
The signal favors the target, but ties leave too few discordant tasks; inspect ties and predeclare more discriminating breadth.
eval-performance
claude-sonnet-5
➖ Mixed evidence
n=8; 3W/2T/3L; d=6; p=0.656; net +0.0%
🟡 0.39
Activation: isolated 6/8; plugin 4/8
Compare winning and losing scenarios to isolate where the target helps versus hurts.
eval-performance
gpt-5.6-luna
➖ Improvement signal, tie-limited
n=8; 4W/4T/0L; d=4; p=0.063; net +50.0%
🟡 0.40
—
The signal favors the target, but ties leave too few discordant tasks; inspect ties and predeclare more discriminating breadth.
extension-points
claude-sonnet-5
➖ Improvement signal, tie-limited
n=7; 3W/3T/1L; d=4; p=0.312; net +28.6%; 2 dormancy excluded
✅ 0.19
Activation: isolated 4/7; plugin 6/7
The signal favors the target, but ties leave too few discordant tasks; inspect ties and predeclare more discriminating breadth.
extension-points
gpt-5.6-luna
➖ Improvement signal, unproven
n=7; 4W/2T/1L; d=5; p=0.188; net +42.9%; 2 dormancy excluded
🟡 0.31
Activation: isolated 7/7; plugin 6/7
The signal favors the target but is inconsistent; inspect tied or lost scenarios and fix weak behavior.
including-generated-files
claude-sonnet-5
➖ Mixed evidence
n=7; 3W/1T/3L; d=6; p=0.656; net +0.0%; 1 dormancy excluded
🟡 0.38
Activation: isolated 4/7; plugin 6/7
Compare winning and losing scenarios to isolate where the target helps versus hurts.
including-generated-files
gpt-5.6-luna
➖ Baseline signal, tie-limited
n=7; 1W/3T/3L; d=4; p=0.312; net -28.6%; 1 dormancy excluded
🟡 0.36
—
Evidence leans baseline but is not credible; inspect losses and make tied tasks discriminate.
incremental-build
claude-sonnet-5
➖ Improvement signal, unproven
n=9; 4W/2T/3L; d=7; p=0.500; net +11.1%
🟡 0.39
Activation: isolated 4/9; plugin 4/9
The signal favors the target but is inconsistent; inspect tied or lost scenarios and fix weak behavior.
incremental-build
gpt-5.6-luna
✅ Improved
n=9; 7W/2T/0L; d=7; p=0.008; net +77.8%
🟡 0.36
Activation: isolated 8/9; plugin 6/9
Fix activation gaps; Review overfit evidence.
item-management
claude-sonnet-5
➖ Improvement signal, tie-limited
n=6; 2W/3T/1L; d=3; p=0.500; net +16.7%; 1 dormancy excluded
🟡 0.37
Activation: isolated 2/6; plugin 3/6
The signal favors the target, but ties leave too few discordant tasks; inspect ties and predeclare more discriminating breadth.
item-management
gpt-5.6-luna
➖ Mixed evidence
n=6; 2W/2T/2L; d=4; p=0.687; net +0.0%; 1 dormancy excluded
🟡 0.49
Activation: isolated 6/6; plugin 5/6
Compare winning and losing scenarios to isolate where the target helps versus hurts.
msbuild-antipatterns
claude-sonnet-5
➖ Improvement signal, tie-limited
n=7; 2W/4T/1L; d=3; p=0.500; net +14.3%; 1 dormancy excluded
🟡 0.30
Activation: isolated 3/7; plugin 3/7
The signal favors the target, but ties leave too few discordant tasks; inspect ties and predeclare more discriminating breadth.
msbuild-antipatterns
gpt-5.6-luna
➖ Baseline signal, tie-limited
n=7; 1W/3T/3L; d=4; p=0.312; net -28.6%; 1 dormancy excluded
🟡 0.24
Activation: isolated 5/7; plugin 4/7
Evidence leans baseline but is not credible; inspect losses and make tied tasks discriminate.
property-patterns
claude-sonnet-5
➖ Improvement signal, unproven
n=7; 3W/2T/2L; d=5; p=0.500; net +14.3%
🟡 0.24
Activation: isolated 6/7; plugin 6/7
The signal favors the target but is inconsistent; inspect tied or lost scenarios and fix weak behavior.
property-patterns
gpt-5.6-luna
➖ Improvement signal, tie-limited
n=7; 3W/4T/0L; d=3; p=0.125; net +42.9%
🟡 0.44
Activation: isolated 6/7; plugin 5/7
The signal favors the target, but ties leave too few discordant tasks; inspect ties and predeclare more discriminating breadth.
ℹ️ How to read this report
✅ Improved — the result passed both the statistical gate and the 20% practical net-win floor.
➖ No clear winner — the result is valid but did not pass both gates. The label distinguishes all ties, mixed evidence, directional but unproven evidence, and credible effects below the practical floor.
⚠️ Invalid / underpowered — the gate withheld a quality verdict. Fix the measurement before judging the target.
⛔ Activation contract failed — the isolated target activated on an explicit dormancy scenario. Dormancy preference is excluded, but this routing failure still blocks a pass.
📉 Preference loss — the LLM judge credibly preferred baseline. It is report-only, not objective completion proof.
Gate evidence — n preference-eligible distinct-stimulus votes, W/T/L stimulus votes, d discordant votes, exact one-sided p, net win, and the count of separately retained dormancy stimuli. The p value applies to one model/target result; no matrix-wide multiple-comparison correction is applied.
Warnings — activation, timeout, retry recovery, or unresolved comparison conditions that need attention.
Do not add repeated runs to increase statistical power. Do not add stimuli after seeing a near-pass unless the new breadth is predeclared for a new experiment.
Why: Net win +60.0% (4W/0T/1L over 5 preference-eligible stimulus vote(s), sign test p=0.188), mean preference +36.0% across 5 paired run(s) — not credible (sign test p=0.188 > 0.05)
Next action: The signal favors the target but is inconsistent; inspect tied or lost scenarios and fix weak behavior.
Gate evidence: n=5; 4W/0T/1L; d=5; p=0.188; net +60.0%
Repeated-run reliability (not used by the gate): 5 paired runs (4W/0T/1L).
Weak or warning scenarios:
Scenario
Preference gate
Net win
Δ Pref
Runs (W/T/L)
▼ Route a slow build to performance analysis
Eligible
-100.0%
-40.0%
0/0/1
Illustrative judge evidence:
Route a slow build to performance analysis: A is more comprehensive and directly actionable for identifying expensive MSBuild targets, tasks, analyzers, restore costs, incrementality failures, and critical-path constraints. B has a stronger statistical baseline prescription, but its unexplained skill-routing language an...
This is one example, not the aggregate verdict. Open Full Results for every judgment.
Why: Net win +40.0% (3W/1T/1L over 5 preference-eligible stimulus vote(s), sign test p=0.312), mean preference +16.0% across 5 paired run(s) — not credible — 1 of 5 preference-eligible stimulus vote(s) tied, leaving only 4 discordant preference vote(s). The sign test conditions on non-tie stimulus votes and cannot reach 0.05 below 5, so no record could have passed here — this is not a measured null. Either the agent is inert on these scenarios (make them discriminate) or the eval needs more distinct stimuli to clear the ties
Next action: The signal favors the target, but ties leave too few discordant tasks; inspect ties and predeclare more discriminating breadth.
Gate evidence: n=5; 3W/1T/1L; d=4; p=0.312; net +40.0%
Repeated-run reliability (not used by the gate): 5 paired runs (3W/1T/1L).
Weak or warning scenarios:
Scenario
Preference gate
Net win
Δ Pref
Runs (W/T/L)
▼ Route a slow build to performance analysis
Eligible
-100.0%
-40.0%
0/0/1
= Triage a build failure and route to appropriate analysis
Eligible
+0.0%
+0.0%
0/1/0
Illustrative judge evidence:
Route a slow build to performance analysis: Response A provides a more systematically organized, measurement-first framework with greater depth on evaluation profiling, suspicious patterns, and validation methodology. It explicitly frames the entire task as a measurement problem before tuning. Response B has strengths i...
This is one example, not the aggregate verdict. Open Full Results for every judgment.
Why: Net win -16.7% (1W/3T/2L over 6 preference-eligible stimulus vote(s), sign test p=0.500), mean preference -20.0% across 7 paired run(s), 1 dormancy stimulus/stimuli excluded from preference — no improvement
Next action: Evidence leans baseline but is not credible; inspect losses and make tied tasks discriminate.
Repeated-run reliability (not used by the gate): 7 paired runs (1W/3T/3L).
Weak or warning scenarios:
Scenario
Preference gate
Net win
Δ Pref
Runs (W/T/L)
▼ Assess a requested binlog investigation when the current build is healthy
Eligible
-100.0%
-100.0%
0/0/1
= Confirm the actual resolved target framework and package version from a binlog
Eligible
+0.0%
+0.0%
0/0/0
= Determine whether a quiet second build actually failed
Eligible
+0.0%
+0.0%
0/1/0
= Diagnose a warning behind a build that actually succeeded
Eligible
+0.0%
+0.0%
0/1/0
= Diagnose build failures from binlog only (no source files)
Eligible
+0.0%
+0.0%
0/1/0
▼ Stay dormant for a non-MSBuild build failure log
Excluded (activation contract)
-100.0%
-40.0%
0/0/1
▼ Trace why a generated source file is missing at compile time
Eligible
-100.0%
-100.0%
0/0/1
▲ Use capture-time text logs when the binlog reader is unavailable
Eligible
+100.0%
+100.0%
1/0/0
Illustrative judge evidence:
Assess a requested binlog investigation when the current build is healthy: A is aligned with the observed evidence and the task's key premise: the current project builds successfully, so no failure root cause exists. It captures and reports the binlog path without inventing a defect. B's additional log inspection is not inherently bad, but its confid...
This is one example, not the aggregate verdict. Open Full Results for every judgment.
Why: Net win +66.7% (4W/2T/0L over 6 preference-eligible stimulus vote(s), sign test p=0.063), mean preference +17.1% across 7 paired run(s), 1 dormancy stimulus/stimuli excluded from preference — not credible — 2 of 6 preference-eligible stimulus vote(s) tied, leaving only 4 discordant preference vote(s). The sign test conditions on non-tie stimulus votes and cannot reach 0.05 below 5, so no record could have passed here — this is not a measured null. Either the skill is inert on these scenarios (make them discriminate) or the eval needs more distinct stimuli to clear the ties
Next action: The signal favors the target, but ties leave too few discordant tasks; inspect ties and predeclare more discriminating breadth.
Why: Net win +42.9% (4W/2T/1L over 7 preference-eligible stimulus vote(s), sign test p=0.188), mean preference +37.5% across 8 paired run(s), 1 dormancy stimulus/stimuli excluded from preference — not credible (sign test p=0.188 > 0.05)
Next action: The signal favors the target but is inconsistent; inspect tied or lost scenarios and fix weak behavior.
Why: Net win +66.7% (4W/2T/0L over 6 preference-eligible stimulus vote(s), sign test p=0.063), mean preference +25.7% across 7 paired run(s), 1 dormancy stimulus/stimuli excluded from preference — not credible — 2 of 6 preference-eligible stimulus vote(s) tied, leaving only 4 discordant preference vote(s). The sign test conditions on non-tie stimulus votes and cannot reach 0.05 below 5, so no record could have passed here — this is not a measured null. Either the skill is inert on these scenarios (make them discriminate) or the eval needs more distinct stimuli to clear the ties
Next action: Inspect activation-only failed runs before rewriting skill content; the model stopped after loading a skill.
Why: Net win +16.7% (3W/1T/2L over 6 preference-eligible stimulus vote(s), sign test p=0.500), mean preference +11.4% across 7 paired run(s), 1 dormancy stimulus/stimuli excluded from preference — not credible (sign test p=0.500 > 0.05)
Next action: Inspect activation-only failed runs before rewriting skill content; the model stopped after loading a skill.
Repeated-run reliability (not used by the gate): 7 paired runs (4W/1T/2L).
Weak or warning scenarios:
Scenario
Preference gate
Net win
Δ Pref
Runs (W/T/L)
▼ Decline graph build for runtime-discovered projects
Eligible
-100.0%
-40.0%
0/0/1
= Enable parallel nodes for a wide project graph
Eligible
+0.0%
+0.0%
0/1/0
▼ Reduce CI build scope with a solution filter
Eligible
-100.0%
-40.0%
0/0/1
Illustrative judge evidence:
Decline graph build for runtime-discovered projects: Both responses are technically sound and cover the core issues well. Response A edges ahead with a clearer three-option structure that directly addresses the user's scenario ('if discovery must remain, don't use /graph'). Response B offers equivalent technical accuracy and add...
This is one example, not the aggregate verdict. Open Full Results for every judgment.
Why: Net win +28.6% (4W/1T/2L over 7 preference-eligible stimulus vote(s), sign test p=0.344), mean preference +15.0% across 8 paired run(s), 1 dormancy stimulus/stimuli excluded from preference — not credible (sign test p=0.344 > 0.05)
Next action: The signal favors the target but is inconsistent; inspect tied or lost scenarios and fix weak behavior.
Why: Net win +42.9% (5W/0T/2L over 7 preference-eligible stimulus vote(s), sign test p=0.227), mean preference +10.0% across 8 paired run(s), 1 dormancy stimulus/stimuli excluded from preference — not credible (sign test p=0.227 > 0.05)
Next action: The signal favors the target but is inconsistent; inspect tied or lost scenarios and fix weak behavior.
Repeated-run reliability (not used by the gate): 8 paired runs (5W/0T/3L).
Weak or warning scenarios:
Scenario
Preference gate
Net win
Δ Pref
Runs (W/T/L)
▼ Decline a runtime latency request that is not a build performance issue
Excluded (activation contract)
-100.0%
-40.0%
0/0/1
▼ Diagnose a Copy task dominating build time
Eligible
-100.0%
-40.0%
0/0/1
▲ Diagnose a pathological ResolveAssemblyReference time
Eligible
+100.0%
+40.0%
1/0/0
▲ Diagnose a single custom target dominating one project's build
Eligible
+100.0%
+40.0%
1/0/0
▲ Diagnose evaluation overhead before any target runs
Eligible
+100.0%
+40.0%
1/0/0
▼ Diagnose per-project overhead across many small projects
Eligible
-100.0%
-40.0%
0/0/1
Illustrative judge evidence:
Diagnose a Copy task dominating build time: Response A delivers a more technically sound diagnosis and avoids recommending an inapplicable MSBuild setting. Its explicit rejection of CreateHardLinksForAdditionalFilesIfPossible (with correct reasoning about scope) demonstrates better MSBuild semantics understanding than R...
This is one example, not the aggregate verdict. Open Full Results for every judgment.
Why: Net win +0.0% (1W/5T/1L over 7 preference-eligible stimulus vote(s), sign test p=0.750), mean preference +0.0% across 7 paired run(s) — no improvement
Next action: Compare winning and losing scenarios to isolate where the target helps versus hurts.
Gate evidence: n=7; 1W/5T/1L; d=2; p=0.750; net +0.0%
Warnings: Activation: isolated 7/7; plugin 6/7
Overfit: Moderate (score 0.23)
Repeated-run reliability (not used by the gate): 7 paired runs (1W/5T/1L).
Weak or warning scenarios:
Scenario
Preference gate
Net win
Δ Pref
Runs (W/T/L)
▼ Audit the mixed solution and separate safe projects from unsafe ones
Eligible
-100.0%
-100.0%
0/0/1
= Avoid a false clash report when projects share only the top-level artifacts root
Eligible
+0.0%
+0.0%
0/1/0
= Decline output-clash remediation for separate projects using the default SDK layout
Eligible
+0.0%
+0.0%
0/1/0
= Diagnose multi-targeting outputs that collapse into one path
Eligible
+0.0%
+0.0%
0/1/0
= Diagnose redundant project reference metadata that forks a same-path build
Eligible
+0.0%
+0.0%
0/1/0
= Diagnose shared output and intermediate path collision
Eligible
+0.0%
+0.0%
0/1/0
Illustrative judge evidence:
Audit the mixed solution and separate safe projects from unsafe ones: A delivers a complete, verified audit covering every actual hazard and the safe projects. B gets the shared-directory and redundant-reference issues right but fails the central MultiTargetLib determination by labeling it safe pending a check it did not perform.
This is one example, not the aggregate verdict. Open Full Results for every judgment.
Why: Net win +42.9% (3W/4T/0L over 7 preference-eligible stimulus vote(s), sign test p=0.125), mean preference +17.1% across 7 paired run(s) — not credible — 4 of 7 preference-eligible stimulus vote(s) tied, leaving only 3 discordant preference vote(s). The sign test conditions on non-tie stimulus votes and cannot reach 0.05 below 5, so no record could have passed here — this is not a measured null. Either the skill is inert on these scenarios (make them discriminate) or the eval needs more distinct stimuli to clear the ties
Next action: The signal favors the target, but ties leave too few discordant tasks; inspect ties and predeclare more discriminating breadth.
Why: Net win -16.7% (2W/1T/3L over 6 preference-eligible stimulus vote(s), sign test p=0.500), mean preference -5.7% across 7 paired run(s), 1 dormancy stimulus/stimuli excluded from preference — no improvement
Next action: Evidence leans baseline but is not credible; inspect losing scenarios for recurring defects.
Why: Net win +33.3% (3W/2T/1L over 6 preference-eligible stimulus vote(s), sign test p=0.312), mean preference +11.4% across 7 paired run(s), 1 dormancy stimulus/stimuli excluded from preference — not credible — 2 of 6 preference-eligible stimulus vote(s) tied, leaving only 4 discordant preference vote(s). The sign test conditions on non-tie stimulus votes and cannot reach 0.05 below 5, so no record could have passed here — this is not a measured null. Either the skill is inert on these scenarios (make them discriminate) or the eval needs more distinct stimuli to clear the ties
Next action: The signal favors the target, but ties leave too few discordant tasks; inspect ties and predeclare more discriminating breadth.
Repeated-run reliability (not used by the gate): 7 paired runs (3W/3T/1L).
Weak or warning scenarios:
Scenario
Preference gate
Net win
Δ Pref
Runs (W/T/L)
▼ Apply repo-level build organization cleanup
Eligible
-100.0%
-100.0%
0/0/1
= Decline adding shared build files to a lone project
Excluded (activation contract)
+0.0%
+0.0%
0/1/0
▲ Diagnose a TargetFramework condition that silently skips in props
Eligible
+100.0%
+40.0%
1/0/0
= Diagnose an inner shared-props file that overwrites its own override
Eligible
+0.0%
+0.0%
0/1/0
= Preserve an intentional project-specific exception while centralizing
Eligible
+0.0%
+0.0%
0/1/0
Illustrative judge evidence:
Apply repo-level build organization cleanup: Response A delivered a complete, working solution: all shared build files were created, settings properly centralized, test configuration isolated, and the final build and tests pass successfully. Response B, despite using a specialized skill and attempting similar changes, le...
This is one example, not the aggregate verdict. Open Full Results for every judgment.
Why: Net win +0.0% (3W/2T/3L over 8 preference-eligible stimulus vote(s), sign test p=0.656), mean preference +0.0% across 8 paired run(s) — no improvement
Next action: Compare winning and losing scenarios to isolate where the target helps versus hurts.
Gate evidence: n=8; 3W/2T/3L; d=6; p=0.656; net +0.0%
Warnings: Activation: isolated 6/8; plugin 4/8
Overfit: Moderate (score 0.39)
Repeated-run reliability (not used by the gate): 8 paired runs (3W/2T/3L).
Weak or warning scenarios:
Scenario
Preference gate
Net win
Δ Pref
Runs (W/T/L)
▲ Decline to invent evaluation problems in an already-clean project
Eligible
+100.0%
+40.0%
1/0/0
▼ Detect a project evaluated twice under different global properties
Eligible
-100.0%
-40.0%
0/0/1
= Gather a measurement before proposing evaluation fixes for an unremarkable project
Eligible
+0.0%
+0.0%
0/1/0
▼ Recognize TreatAsLocalProperty overuse versus one justified entry
Eligible
-100.0%
-40.0%
0/0/1
▼ Redirect an incremental-rebuild complaint mistakenly framed as an evaluation problem
Eligible
-100.0%
-40.0%
0/0/1
= Triage which of two property functions actually costs evaluation time
Eligible
+0.0%
+0.0%
0/1/0
Illustrative judge evidence:
Detect a project evaluated twice under different global properties: A gives the more actionable and better-supported diagnosis: it verifies the differing property appears unused and that both calls generate the same output, then offers the appropriate remove-the-duplicate-or-wire-the-feature alternatives. B reaches the core finding but oversta...
This is one example, not the aggregate verdict. Open Full Results for every judgment.
Why: Net win +50.0% (4W/4T/0L over 8 preference-eligible stimulus vote(s), sign test p=0.063), mean preference +35.0% across 8 paired run(s) — not credible — 4 of 8 preference-eligible stimulus vote(s) tied, leaving only 4 discordant preference vote(s). The sign test conditions on non-tie stimulus votes and cannot reach 0.05 below 5, so no record could have passed here — this is not a measured null. Either the skill is inert on these scenarios (make them discriminate) or the eval needs more distinct stimuli to clear the ties
Next action: The signal favors the target, but ties leave too few discordant tasks; inspect ties and predeclare more discriminating breadth.
Why: Net win +28.6% (3W/3T/1L over 7 preference-eligible stimulus vote(s), sign test p=0.312), mean preference +8.9% across 9 paired run(s), 2 dormancy stimulus/stimuli excluded from preference — not credible — 3 of 7 preference-eligible stimulus vote(s) tied, leaving only 4 discordant preference vote(s). The sign test conditions on non-tie stimulus votes and cannot reach 0.05 below 5, so no record could have passed here — this is not a measured null. Either the skill is inert on these scenarios (make them discriminate) or the eval needs more distinct stimuli to clear the ties
Next action: The signal favors the target, but ties leave too few discordant tasks; inspect ties and predeclare more discriminating breadth.
Why: Net win +42.9% (4W/2T/1L over 7 preference-eligible stimulus vote(s), sign test p=0.188), mean preference +24.4% across 9 paired run(s), 2 dormancy stimulus/stimuli excluded from preference — not credible (sign test p=0.188 > 0.05)
Next action: The signal favors the target but is inconsistent; inspect tied or lost scenarios and fix weak behavior.
Why: Net win +0.0% (3W/1T/3L over 7 preference-eligible stimulus vote(s), sign test p=0.656), mean preference +0.0% across 8 paired run(s), 1 dormancy stimulus/stimuli excluded from preference — no improvement
Next action: Compare winning and losing scenarios to isolate where the target helps versus hurts.
Why: Net win -28.6% (1W/3T/3L over 7 preference-eligible stimulus vote(s), sign test p=0.312), mean preference -10.0% across 8 paired run(s), 1 dormancy stimulus/stimuli excluded from preference — no improvement
Next action: Evidence leans baseline but is not credible; inspect losses and make tied tasks discriminate.
Why: Net win +11.1% (4W/2T/3L over 9 preference-eligible stimulus vote(s), sign test p=0.500), mean preference +4.4% across 9 paired run(s) — not credible (sign test p=0.500 > 0.05)
Next action: The signal favors the target but is inconsistent; inspect tied or lost scenarios and fix weak behavior.
Why: Net win +16.7% (2W/3T/1L over 6 preference-eligible stimulus vote(s), sign test p=0.500), mean preference +5.7% across 7 paired run(s), 1 dormancy stimulus/stimuli excluded from preference — not credible — 3 of 6 preference-eligible stimulus vote(s) tied, leaving only 3 discordant preference vote(s). The sign test conditions on non-tie stimulus votes and cannot reach 0.05 below 5, so no record could have passed here — this is not a measured null. Either the skill is inert on these scenarios (make them discriminate) or the eval needs more distinct stimuli to clear the ties
Next action: The signal favors the target, but ties leave too few discordant tasks; inspect ties and predeclare more discriminating breadth.
Repeated-run reliability (not used by the gate): 7 paired runs (2W/4T/1L).
Weak or warning scenarios:
Scenario
Preference gate
Net win
Δ Pref
Runs (W/T/L)
= Decline JavaScript configuration merging
Excluded (activation contract)
+0.0%
+0.0%
0/1/0
▲ Diagnose an ineffective Compile Remove that does not match the glob
Eligible
+100.0%
+40.0%
1/0/0
▲ Diagnose item group and batching issues
Eligible
+100.0%
+40.0%
1/0/0
▼ Diagnose real and claimed item problems in a code generation pipeline
Eligible
-100.0%
-40.0%
0/0/1
= Fix item management anti-patterns
Eligible
+0.0%
+0.0%
0/1/0
= Leave already-correct item management unchanged
Eligible
+0.0%
+0.0%
0/1/0
= Leave correct single-list batching unchanged
Eligible
+0.0%
+0.0%
0/1/0
Illustrative judge evidence:
Diagnose real and claimed item problems in a code generation pipeline: Both fix the visible duplicate-warning and validation symptoms and verify build/clean behavior, but both miss two requested structural corrections: an exact evaluation-time generated-source declaration and removal of CleanGeneratedCode. A is marginally preferable because it pr...
This is one example, not the aggregate verdict. Open Full Results for every judgment.
➖ Mixed evidence — item-management (gpt-5.6-luna)
Why: Net win +0.0% (2W/2T/2L over 6 preference-eligible stimulus vote(s), sign test p=0.687), mean preference +5.7% across 7 paired run(s), 1 dormancy stimulus/stimuli excluded from preference — no improvement
Next action: Compare winning and losing scenarios to isolate where the target helps versus hurts.
Repeated-run reliability (not used by the gate): 7 paired runs (3W/2T/2L).
Weak or warning scenarios:
Scenario
Preference gate
Net win
Δ Pref
Runs (W/T/L)
▼ Diagnose an ineffective Compile Remove that does not match the glob
Eligible
-100.0%
-100.0%
0/0/1
= Diagnose item group and batching issues
Eligible
+0.0%
+0.0%
0/1/0
= Diagnose real and claimed item problems in a code generation pipeline
Eligible
+0.0%
+0.0%
0/1/0
▼ Leave already-correct item management unchanged
Eligible
-100.0%
-40.0%
0/0/1
Illustrative judge evidence:
Diagnose an ineffective Compile Remove that does not match the glob: Response A provides the correct, precise solution to the glob-matching problem by changing the pattern to Generated\.g.cs, which directly matches the actual file naming convention. Response B's solution using Generated\**\.cs is overly broad and would exclude ALL C# files...
This is one example, not the aggregate verdict. Open Full Results for every judgment.
Why: Net win +14.3% (2W/4T/1L over 7 preference-eligible stimulus vote(s), sign test p=0.500), mean preference +0.0% across 8 paired run(s), 1 dormancy stimulus/stimuli excluded from preference — not credible — 4 of 7 preference-eligible stimulus vote(s) tied, leaving only 3 discordant preference vote(s). The sign test conditions on non-tie stimulus votes and cannot reach 0.05 below 5, so no record could have passed here — this is not a measured null. Either the skill is inert on these scenarios (make them discriminate) or the eval needs more distinct stimuli to clear the ties
Next action: The signal favors the target, but ties leave too few discordant tasks; inspect ties and predeclare more discriminating breadth.
Repeated-run reliability (not used by the gate): 8 paired runs (2W/4T/2L).
Weak or warning scenarios:
Scenario
Preference gate
Net win
Δ Pref
Runs (W/T/L)
▲ Add a module to an F# project
Eligible
+100.0%
+40.0%
1/0/0
= Add a signature file to define public API
Eligible
+0.0%
+0.0%
0/1/0
= Distinguish a style backslash from a real cross-platform backslash bug
Eligible
+0.0%
+0.0%
0/1/0
= Fix broken file order causing FS0039
Eligible
+0.0%
+0.0%
0/1/0
▲ Judge an unguarded import inside a NuGet package build folder
Eligible
+100.0%
+40.0%
1/0/0
= Leave a clean project without inventing anti-patterns
Eligible
+0.0%
+0.0%
0/1/0
▼ Non-activation: migrate a legacy project to SDK style
Excluded (activation contract)
-100.0%
-40.0%
0/0/1
▼ Review MSBuild files for anti-patterns and style issues
Eligible
-100.0%
-40.0%
0/0/1
Illustrative judge evidence:
Add a signature file to define public API: The two solutions are substantively equivalent and fully satisfy the task: each adds a correct public type signature, compiles it in the required order, and verifies a clean build.
This is one example, not the aggregate verdict. Open Full Results for every judgment.
Why: Net win -28.6% (1W/3T/3L over 7 preference-eligible stimulus vote(s), sign test p=0.312), mean preference -10.0% across 8 paired run(s), 1 dormancy stimulus/stimuli excluded from preference — no improvement
Next action: Evidence leans baseline but is not credible; inspect losses and make tied tasks discriminate.
Repeated-run reliability (not used by the gate): 8 paired runs (1W/4T/3L).
Weak or warning scenarios:
Scenario
Preference gate
Net win
Δ Pref
Runs (W/T/L)
▼ Add a module to an F# project
Eligible
-100.0%
-40.0%
0/0/1
= Add a signature file to define public API
Eligible
+0.0%
+0.0%
0/1/0
= Distinguish a style backslash from a real cross-platform backslash bug
Eligible
+0.0%
+0.0%
0/1/0
▼ Fix broken file order causing FS0039
Eligible
-100.0%
-40.0%
0/0/1
▼ Judge an unguarded import inside a NuGet package build folder
Eligible
-100.0%
-40.0%
0/0/1
= Leave a clean project without inventing anti-patterns
Eligible
+0.0%
+0.0%
0/1/0
= Non-activation: migrate a legacy project to SDK style
Excluded (activation contract)
+0.0%
+0.0%
0/1/0
Illustrative judge evidence:
Add a module to an F# project: Both responses successfully completed the core task, but Response A provides superior transparency and evidence of implementation. Response A explicitly shows the error handling code with match expressions and result type handling, with multiple iterations demonstrating thorou...
This is one example, not the aggregate verdict. Open Full Results for every judgment.
Why: Net win +14.3% (3W/2T/2L over 7 preference-eligible stimulus vote(s), sign test p=0.500), mean preference +5.7% across 7 paired run(s) — not credible (sign test p=0.500 > 0.05)
Next action: The signal favors the target but is inconsistent; inspect tied or lost scenarios and fix weak behavior.
Why: Net win +42.9% (3W/4T/0L over 7 preference-eligible stimulus vote(s), sign test p=0.125), mean preference +17.1% across 7 paired run(s) — not credible — 4 of 7 preference-eligible stimulus vote(s) tied, leaving only 3 discordant preference vote(s). The sign test conditions on non-tie stimulus votes and cannot reach 0.05 below 5, so no record could have passed here — this is not a measured null. Either the skill is inert on these scenarios (make them discriminate) or the eval needs more distinct stimuli to clear the ties
Next action: The signal favors the target, but ties leave too few discordant tasks; inspect ties and predeclare more discriminating breadth.
To investigate non-passing or warning results, paste this to your AI coding agent:
For PR 1215 in dotnet/skills, download eval artifacts with gh run download 37048637618 --repo dotnet/skills --pattern "vally-results-*" --dir ./eval-results, then fetch https://raw.githubusercontent.com/dotnet/skills/970e0dc1bd6bfcd78f4b1263252d049cda277641/eng/vally-adapter/InvestigatingResults.md and follow it. Classify each result as measurement-invalid, underpowered, not-proven, preference-loss, or passing-with-warning. Use stateReason, result accounting, weak scenarios, and judge evidence to give the cause and exact next fix.
⚠️ Session replay telemetry was not published because the auxiliary dotnet/skills-data publisher failed. The evaluation verdicts above remain authoritative; maintainers must repair SKILLS_DATA_TOKEN or the publisher before replay links are available.
This standalone fallback checklist starts at step 2 and continues with steps 3–5, even though the preceding MCP checklist is a separate section. Renumber these items 1–4 so the fallback workflow is not presented with a missing first step.
Objective completion gate: not enabled. Aggregate completion transitions are telemetry only, so this report does not claim that zero objective regressions were proven.
A result passes only when preference-eligible distinct-stimulus votes have aggregate net win of at least 20% and an exact one-sided sign-test result of p ≤ 0.05, and every explicit dormancy activation contract passes. Repeated runs measure reliability only.
Target
Model
Verdict
Gate evidence
Overfit
Warnings
Next action
agent.msbuild
claude-sonnet-5
➖ Improvement signal, tie-limited
n=5; 3W/1T/1L; d=4; p=0.312; net +40.0%
—
—
The signal favors the target, but ties leave too few discordant tasks; inspect ties and predeclare more discriminating breadth.
agent.msbuild
gpt-5.6-luna
➖ Improvement signal, tie-limited
n=5; 3W/2T/0L; d=3; p=0.125; net +60.0%
—
—
The signal favors the target, but ties leave too few discordant tasks; inspect ties and predeclare more discriminating breadth.
binlog-failure-analysis
claude-sonnet-5
➖ Improvement signal, tie-limited
n=6; 3W/2T/1L; d=4; p=0.312; net +33.3%; 1 dormancy excluded
🟡 0.30
Activation: isolated 6/7; plugin 6/7
The signal favors the target, but ties leave too few discordant tasks; inspect ties and predeclare more discriminating breadth.
binlog-failure-analysis
gpt-5.6-luna
➖ Improvement signal, unproven
n=6; 4W/1T/1L; d=5; p=0.188; net +50.0%; 1 dormancy excluded
🟡 0.26
Activation: isolated 6/7; plugin 6/7
The signal favors the target but is inconsistent; inspect tied or lost scenarios and fix weak behavior.
binlog-generation
claude-sonnet-5
➖ Improvement signal, unproven
n=7; 4W/1T/2L; d=6; p=0.344; net +28.6%; 1 dormancy excluded
🔴 0.50
Activation: isolated 5/7; plugin 6/7
The signal favors the target but is inconsistent; inspect tied or lost scenarios and fix weak behavior.
binlog-generation
gpt-5.6-luna
➖ Improvement signal, unproven
n=7; 5W/1T/1L; d=6; p=0.109; net +57.1%; 1 dormancy excluded
🟡 0.36
Activation: isolated 7/7; plugin 5/7
The signal favors the target but is inconsistent; inspect tied or lost scenarios and fix weak behavior.
build-parallelism
claude-sonnet-5
➖ Baseline signal, tie-limited
n=6; 1W/3T/2L; d=3; p=0.500; net -16.7%; 1 dormancy excluded
Inspect activation-only failed runs before rewriting skill content; the model stopped after loading a skill.
build-perf-baseline
claude-sonnet-5
➖ Mixed evidence
n=6; 3W/0T/3L; d=6; p=0.656; net +0.0%; 1 dormancy excluded
🟡 0.44
Activation: isolated 5/6; plugin 5/6
Compare winning and losing scenarios to isolate where the target helps versus hurts.
build-perf-baseline
gpt-5.6-luna
➖ Improvement signal, tie-limited
n=6; 3W/2T/1L; d=4; p=0.312; net +33.3%; 1 dormancy excluded
🟡 0.24
Activation: isolated 5/6; plugin 4/6
The signal favors the target, but ties leave too few discordant tasks; inspect ties and predeclare more discriminating breadth.
build-perf-diagnostics
claude-sonnet-5
➖ Improvement signal, unproven
n=7; 3W/2T/2L; d=5; p=0.500; net +14.3%; 1 dormancy excluded
🟡 0.47
Activation: isolated 7/7; plugin 2/7
The signal favors the target but is inconsistent; inspect tied or lost scenarios and fix weak behavior.
build-perf-diagnostics
gpt-5.6-luna
➖ Improvement signal, unproven
n=7; 4W/2T/1L; d=5; p=0.188; net +42.9%; 1 dormancy excluded
🟡 0.43
Activation: isolated 7/7; plugin 3/7
The signal favors the target but is inconsistent; inspect tied or lost scenarios and fix weak behavior.
check-bin-obj-clash
claude-sonnet-5
➖ Improvement signal, tie-limited
n=7; 3W/4T/0L; d=3; p=0.125; net +42.9%
✅ 0.16
Activation: isolated 7/7; plugin 5/7
The signal favors the target, but ties leave too few discordant tasks; inspect ties and predeclare more discriminating breadth.
check-bin-obj-clash
gpt-5.6-luna
➖ Improvement signal, tie-limited
n=7; 2W/4T/1L; d=3; p=0.500; net +14.3%
🟡 0.31
Activation: isolated 7/7; plugin 6/7
The signal favors the target, but ties leave too few discordant tasks; inspect ties and predeclare more discriminating breadth.
directory-build-organization
claude-sonnet-5
✅ Improved
n=6; 5W/1T/0L; d=5; p=0.031; net +83.3%; 1 dormancy excluded
🟡 0.38
Activation: isolated 4/6; plugin 4/6
Fix activation gaps; Review overfit evidence.
directory-build-organization
gpt-5.6-luna
➖ Baseline signal, tie-limited
n=6; 1W/3T/2L; d=3; p=0.500; net -16.7%; 1 dormancy excluded
🟡 0.31
Activation: isolated 6/6; plugin 4/6
Evidence leans baseline but is not credible; inspect losses and make tied tasks discriminate.
eval-performance
claude-sonnet-5
➖ Improvement signal, unproven
n=8; 5W/2T/1L; d=6; p=0.109; net +50.0%
🟡 0.27
Activation: isolated 5/8; plugin 5/8
The signal favors the target but is inconsistent; inspect tied or lost scenarios and fix weak behavior.
eval-performance
gpt-5.6-luna
➖ Improvement signal, unproven
n=8; 5W/2T/1L; d=6; p=0.109; net +50.0%
🟡 0.38
Activation: isolated 8/8; plugin 7/8
The signal favors the target but is inconsistent; inspect tied or lost scenarios and fix weak behavior.
extension-points
claude-sonnet-5
➖ Improvement signal, tie-limited
n=7; 3W/3T/1L; d=4; p=0.312; net +28.6%; 2 dormancy excluded
🟡 0.27
Activation: isolated 6/7; plugin 4/7
The signal favors the target, but ties leave too few discordant tasks; inspect ties and predeclare more discriminating breadth.
extension-points
gpt-5.6-luna
➖ Improvement signal, tie-limited
n=7; 2W/4T/1L; d=3; p=0.500; net +14.3%; 2 dormancy excluded
🟡 0.39
Activation: isolated 7/7; plugin 5/7
The signal favors the target, but ties leave too few discordant tasks; inspect ties and predeclare more discriminating breadth.
including-generated-files
claude-sonnet-5
➖ Improvement signal, tie-limited
n=7; 2W/4T/1L; d=3; p=0.500; net +14.3%; 1 dormancy excluded
🟡 0.35
Activation: isolated 4/7; plugin 4/7
The signal favors the target, but ties leave too few discordant tasks; inspect ties and predeclare more discriminating breadth.
including-generated-files
gpt-5.6-luna
➖ Mixed evidence
n=7; 3W/1T/3L; d=6; p=0.656; net +0.0%; 1 dormancy excluded
🟡 0.40
—
Compare winning and losing scenarios to isolate where the target helps versus hurts.
incremental-build
claude-sonnet-5
➖ Improvement signal, tie-limited
n=9; 3W/5T/1L; d=4; p=0.312; net +22.2%
🟡 0.24
Activation: isolated 3/9; plugin 5/9
The signal favors the target, but ties leave too few discordant tasks; inspect ties and predeclare more discriminating breadth.
incremental-build
gpt-5.6-luna
➖ Improvement signal, unproven
n=9; 4W/4T/1L; d=5; p=0.188; net +33.3%
🟡 0.22
Activation: isolated 8/9; plugin 6/9
The signal favors the target but is inconsistent; inspect tied or lost scenarios and fix weak behavior.
item-management
claude-sonnet-5
➖ Mixed evidence
n=6; 2W/2T/2L; d=4; p=0.687; net +0.0%; 1 dormancy excluded
🟡 0.33
Activation: isolated 4/6; plugin 3/6
Compare winning and losing scenarios to isolate where the target helps versus hurts.
item-management
gpt-5.6-luna
➖ Mixed evidence
n=6; 1W/4T/1L; d=2; p=0.750; net +0.0%; 1 dormancy excluded
🟡 0.30
Activation: isolated 6/6; plugin 5/6
Compare winning and losing scenarios to isolate where the target helps versus hurts.
msbuild-antipatterns
claude-sonnet-5
➖ Mixed evidence
n=7; 2W/3T/2L; d=4; p=0.687; net +0.0%; 1 dormancy excluded
🟡 0.30
Activation: isolated 3/7; plugin 2/7
Compare winning and losing scenarios to isolate where the target helps versus hurts.
msbuild-antipatterns
gpt-5.6-luna
➖ Mixed evidence
n=7; 1W/5T/1L; d=2; p=0.750; net +0.0%; 1 dormancy excluded
🟡 0.29
Activation: isolated 5/7; plugin 3/7
Compare winning and losing scenarios to isolate where the target helps versus hurts.
property-patterns
claude-sonnet-5
➖ Baseline signal, unproven
n=7; 2W/0T/5L; d=7; p=0.227; net -42.9%
🟡 0.27
Activation: isolated 6/7; plugin 6/7
Evidence leans baseline but is not credible; inspect losing scenarios for recurring defects.
property-patterns
gpt-5.6-luna
➖ Improvement signal, unproven
n=7; 3W/2T/2L; d=5; p=0.500; net +14.3%
🟡 0.37
Activation: isolated 6/7; plugin 6/7
The signal favors the target but is inconsistent; inspect tied or lost scenarios and fix weak behavior.
ℹ️ How to read this report
✅ Improved — the result passed both the statistical gate and the 20% practical net-win floor.
➖ No clear winner — the result is valid but did not pass both gates. The label distinguishes all ties, mixed evidence, directional but unproven evidence, and credible effects below the practical floor.
⚠️ Invalid / underpowered — the gate withheld a quality verdict. Fix the measurement before judging the target.
⛔ Activation contract failed — the isolated target activated on an explicit dormancy scenario. Dormancy preference is excluded, but this routing failure still blocks a pass.
📉 Preference loss — the LLM judge credibly preferred baseline. It is report-only, not objective completion proof.
Gate evidence — n preference-eligible distinct-stimulus votes, W/T/L stimulus votes, d discordant votes, exact one-sided p, net win, and the count of separately retained dormancy stimuli. The p value applies to one model/target result; no matrix-wide multiple-comparison correction is applied.
Warnings — activation, timeout, retry recovery, or unresolved comparison conditions that need attention.
Do not add repeated runs to increase statistical power. Do not add stimuli after seeing a near-pass unless the new breadth is predeclared for a new experiment.
Why: Net win +40.0% (3W/1T/1L over 5 preference-eligible stimulus vote(s), sign test p=0.312), mean preference +4.0% across 5 paired run(s) — not credible — 1 of 5 preference-eligible stimulus vote(s) tied, leaving only 4 discordant preference vote(s). The sign test conditions on non-tie stimulus votes and cannot reach 0.05 below 5, so no record could have passed here — this is not a measured null. Either the agent is inert on these scenarios (make them discriminate) or the eval needs more distinct stimuli to clear the ties
Next action: The signal favors the target, but ties leave too few discordant tasks; inspect ties and predeclare more discriminating breadth.
Why: Net win +60.0% (3W/2T/0L over 5 preference-eligible stimulus vote(s), sign test p=0.125), mean preference +36.0% across 5 paired run(s) — not credible — 2 of 5 preference-eligible stimulus vote(s) tied, leaving only 3 discordant preference vote(s). The sign test conditions on non-tie stimulus votes and cannot reach 0.05 below 5, so no record could have passed here — this is not a measured null. Either the agent is inert on these scenarios (make them discriminate) or the eval needs more distinct stimuli to clear the ties
Next action: The signal favors the target, but ties leave too few discordant tasks; inspect ties and predeclare more discriminating breadth.
Why: Net win +33.3% (3W/2T/1L over 6 preference-eligible stimulus vote(s), sign test p=0.312), mean preference +17.1% across 7 paired run(s), 1 dormancy stimulus/stimuli excluded from preference — not credible — 2 of 6 preference-eligible stimulus vote(s) tied, leaving only 4 discordant preference vote(s). The sign test conditions on non-tie stimulus votes and cannot reach 0.05 below 5, so no record could have passed here — this is not a measured null. Either the skill is inert on these scenarios (make them discriminate) or the eval needs more distinct stimuli to clear the ties
Next action: The signal favors the target, but ties leave too few discordant tasks; inspect ties and predeclare more discriminating breadth.
Repeated-run reliability (not used by the gate): 7 paired runs (4W/2T/1L).
Weak or warning scenarios:
Scenario
Preference gate
Net win
Δ Pref
Runs (W/T/L)
= Confirm the actual resolved target framework and package version from a binlog
Eligible
+0.0%
+0.0%
0/0/0
▼ Determine whether a quiet second build actually failed
Eligible
-100.0%
-40.0%
0/0/1
= Diagnose a warning behind a build that actually succeeded
Eligible
+0.0%
+0.0%
0/1/0
= Trace why a generated source file is missing at compile time
Eligible
+0.0%
+0.0%
0/1/0
Illustrative judge evidence:
Determine whether a quiet second build actually failed: Both answers are correct, concise, and directly responsive. A is marginally stronger because it supplies a more precise, directly quoted binlog citation (including line numbers) for both success and the CoreCompile up-to-date skip; B's optional rebuild advice is useful but doe...
This is one example, not the aggregate verdict. Open Full Results for every judgment.
Why: Net win +50.0% (4W/1T/1L over 6 preference-eligible stimulus vote(s), sign test p=0.188), mean preference +14.3% across 7 paired run(s), 1 dormancy stimulus/stimuli excluded from preference — not credible (sign test p=0.188 > 0.05)
Next action: The signal favors the target but is inconsistent; inspect tied or lost scenarios and fix weak behavior.
Repeated-run reliability (not used by the gate): 7 paired runs (5W/1T/1L).
Weak or warning scenarios:
Scenario
Preference gate
Net win
Δ Pref
Runs (W/T/L)
▼ Assess a requested binlog investigation when the current build is healthy
Eligible
-100.0%
-100.0%
0/0/1
= Confirm the actual resolved target framework and package version from a binlog
Eligible
+0.0%
+0.0%
0/0/0
= Diagnose build failures from binlog only (no source files)
Eligible
+0.0%
+0.0%
0/1/0
Illustrative judge evidence:
Assess a requested binlog investigation when the current build is healthy: Response A correctly identifies that InventorySync builds successfully with 0 errors and 0 warnings, captures binlogs as evidence, and identifies only a reproducibility risk (missing global.json). Response B violates the core rubric by artificially manufacturing a build failur...
This is one example, not the aggregate verdict. Open Full Results for every judgment.
Why: Net win +28.6% (4W/1T/2L over 7 preference-eligible stimulus vote(s), sign test p=0.344), mean preference +25.0% across 8 paired run(s), 1 dormancy stimulus/stimuli excluded from preference — not credible (sign test p=0.344 > 0.05)
Next action: The signal favors the target but is inconsistent; inspect tied or lost scenarios and fix weak behavior.
Why: Net win +57.1% (5W/1T/1L over 7 preference-eligible stimulus vote(s), sign test p=0.109), mean preference +50.0% across 8 paired run(s), 1 dormancy stimulus/stimuli excluded from preference — not credible (sign test p=0.109 > 0.05)
Next action: The signal favors the target but is inconsistent; inspect tied or lost scenarios and fix weak behavior.
Why: Net win -16.7% (1W/3T/2L over 6 preference-eligible stimulus vote(s), sign test p=0.500), mean preference +2.9% across 7 paired run(s), 1 dormancy stimulus/stimuli excluded from preference — no improvement
Next action: Inspect activation-only failed runs before rewriting skill content; the model stopped after loading a skill.
Repeated-run reliability (not used by the gate): 7 paired runs (1W/4T/2L).
Weak or warning scenarios:
Scenario
Preference gate
Net win
Δ Pref
Runs (W/T/L)
▼ Analyze build parallelism bottlenecks
Eligible
-100.0%
-40.0%
0/0/1
= Decline database query tuning request
Excluded (activation contract)
+0.0%
+0.0%
0/1/0
= Decline graph build for runtime-discovered projects
Eligible
+0.0%
+0.0%
0/1/0
▼ Enable BuildInParallel on a custom MSBuild task
Eligible
-100.0%
-40.0%
0/0/1
= Enable parallel nodes for a wide project graph
Eligible
+0.0%
+0.0%
0/1/0
= Reduce CI build scope with a solution filter
Eligible
+0.0%
+0.0%
0/1/0
Illustrative judge evidence:
Analyze build parallelism bottlenecks: The core dependency and redundancy conclusions are equally correct, but A better fulfills the explicit binlog-based analysis requirement with concrete execution evidence.
This is one example, not the aggregate verdict. Open Full Results for every judgment.
Why: Net win +66.7% (4W/2T/0L over 6 preference-eligible stimulus vote(s), sign test p=0.063), mean preference +31.4% across 7 paired run(s), 1 dormancy stimulus/stimuli excluded from preference — not credible — 2 of 6 preference-eligible stimulus vote(s) tied, leaving only 4 discordant preference vote(s). The sign test conditions on non-tie stimulus votes and cannot reach 0.05 below 5, so no record could have passed here — this is not a measured null. Either the skill is inert on these scenarios (make them discriminate) or the eval needs more distinct stimuli to clear the ties
Next action: Inspect activation-only failed runs before rewriting skill content; the model stopped after loading a skill.
Why: Net win +0.0% (3W/0T/3L over 6 preference-eligible stimulus vote(s), sign test p=0.656), mean preference -8.6% across 7 paired run(s), 1 dormancy stimulus/stimuli excluded from preference — no improvement
Next action: Compare winning and losing scenarios to isolate where the target helps versus hurts.
Repeated-run reliability (not used by the gate): 7 paired runs (3W/1T/3L).
Weak or warning scenarios:
Scenario
Preference gate
Net win
Δ Pref
Runs (W/T/L)
▼ Configure deterministic, cache-safe CI builds
Eligible
-100.0%
-40.0%
0/0/1
= Decline a non-MSBuild build performance request
Excluded (activation contract)
+0.0%
+0.0%
0/1/0
▼ Leave an already-optimized build unchanged
Eligible
-100.0%
-100.0%
0/0/1
▼ Route a broken no-op rebuild away from generic optimization
Eligible
-100.0%
-40.0%
0/0/1
▲ Route a restore-bound cold build away from architecture changes
Eligible
+100.0%
+40.0%
1/0/0
Illustrative judge evidence:
Configure deterministic, cache-safe CI builds: Both make the essential CI-gated ContinuousIntegrationBuild change, and both unnecessarily make Deterministic explicit. A is modestly stronger because it investigated the actual cross-agent discrepancy, adds path normalization, and validates the resulting artifact byte-for-byt...
This is one example, not the aggregate verdict. Open Full Results for every judgment.
Why: Net win +33.3% (3W/2T/1L over 6 preference-eligible stimulus vote(s), sign test p=0.312), mean preference +22.9% across 7 paired run(s), 1 dormancy stimulus/stimuli excluded from preference — not credible — 2 of 6 preference-eligible stimulus vote(s) tied, leaving only 4 discordant preference vote(s). The sign test conditions on non-tie stimulus votes and cannot reach 0.05 below 5, so no record could have passed here — this is not a measured null. Either the skill is inert on these scenarios (make them discriminate) or the eval needs more distinct stimuli to clear the ties
Next action: The signal favors the target, but ties leave too few discordant tasks; inspect ties and predeclare more discriminating breadth.
Repeated-run reliability (not used by the gate): 7 paired runs (3W/2T/2L).
Weak or warning scenarios:
Scenario
Preference gate
Net win
Δ Pref
Runs (W/T/L)
▲ Configure deterministic, cache-safe CI builds
Eligible
+100.0%
+100.0%
1/0/0
▼ Decline a non-MSBuild build performance request
Excluded (activation contract)
-100.0%
-40.0%
0/0/1
▼ Establish build performance baseline and recommend optimizations
Eligible
-100.0%
-100.0%
0/0/1
= Route a broken no-op rebuild away from generic optimization
Eligible
+0.0%
+0.0%
0/1/0
= Route a restore-bound cold build away from architecture changes
Eligible
+0.0%
+0.0%
0/1/0
Illustrative judge evidence:
Establish build performance baseline and recommend optimizations: Response A completed the full task by delivering specific, actionable build optimization recommendations grounded in measured baseline data. It identified redundant project references by name, recommended conditionalizing unnecessary build overhead, suggested advanced techniqu...
This is one example, not the aggregate verdict. Open Full Results for every judgment.
Why: Net win +14.3% (3W/2T/2L over 7 preference-eligible stimulus vote(s), sign test p=0.500), mean preference +5.0% across 8 paired run(s), 1 dormancy stimulus/stimuli excluded from preference — not credible (sign test p=0.500 > 0.05)
Next action: The signal favors the target but is inconsistent; inspect tied or lost scenarios and fix weak behavior.
Repeated-run reliability (not used by the gate): 8 paired runs (3W/3T/2L).
Weak or warning scenarios:
Scenario
Preference gate
Net win
Δ Pref
Runs (W/T/L)
= Decline a runtime latency request that is not a build performance issue
Excluded (activation contract)
+0.0%
+0.0%
0/1/0
▼ Diagnose a Copy task dominating build time
Eligible
-100.0%
-40.0%
0/0/1
▲ Diagnose a pathological ResolveAssemblyReference time
Eligible
+100.0%
+40.0%
1/0/0
= Diagnose a single custom target dominating one project's build
Eligible
+0.0%
+0.0%
0/1/0
= Diagnose evaluation overhead before any target runs
Eligible
+0.0%
+0.0%
0/1/0
▼ Diagnose per-project overhead across many small projects
Eligible
-100.0%
-40.0%
0/0/1
▲ Diagnose slow build for a small project
Eligible
+100.0%
+40.0%
1/0/0
Illustrative judge evidence:
Diagnose a Copy task dominating build time: The central diagnosis and primary PreserveNewest fix are strong in both responses. A is more reliable because it supplies the applicable Always-copy skip property rather than B's dubious SkipCopyUnchangedFiles recommendation. B's more explicit hardlink snippet is useful but ...
This is one example, not the aggregate verdict. Open Full Results for every judgment.
Why: Net win +42.9% (4W/2T/1L over 7 preference-eligible stimulus vote(s), sign test p=0.188), mean preference +15.0% across 8 paired run(s), 1 dormancy stimulus/stimuli excluded from preference — not credible (sign test p=0.188 > 0.05)
Next action: The signal favors the target but is inconsistent; inspect tied or lost scenarios and fix weak behavior.
Why: Net win +42.9% (3W/4T/0L over 7 preference-eligible stimulus vote(s), sign test p=0.125), mean preference +25.7% across 7 paired run(s) — not credible — 4 of 7 preference-eligible stimulus vote(s) tied, leaving only 3 discordant preference vote(s). The sign test conditions on non-tie stimulus votes and cannot reach 0.05 below 5, so no record could have passed here — this is not a measured null. Either the skill is inert on these scenarios (make them discriminate) or the eval needs more distinct stimuli to clear the ties
Next action: The signal favors the target, but ties leave too few discordant tasks; inspect ties and predeclare more discriminating breadth.
Why: Net win +14.3% (2W/4T/1L over 7 preference-eligible stimulus vote(s), sign test p=0.500), mean preference +5.7% across 7 paired run(s) — not credible — 4 of 7 preference-eligible stimulus vote(s) tied, leaving only 3 discordant preference vote(s). The sign test conditions on non-tie stimulus votes and cannot reach 0.05 below 5, so no record could have passed here — this is not a measured null. Either the skill is inert on these scenarios (make them discriminate) or the eval needs more distinct stimuli to clear the ties
Next action: The signal favors the target, but ties leave too few discordant tasks; inspect ties and predeclare more discriminating breadth.
Why: Net win -16.7% (1W/3T/2L over 6 preference-eligible stimulus vote(s), sign test p=0.500), mean preference -5.7% across 7 paired run(s), 1 dormancy stimulus/stimuli excluded from preference — no improvement
Next action: Evidence leans baseline but is not credible; inspect losses and make tied tasks discriminate.
Why: Net win +50.0% (5W/2T/1L over 8 preference-eligible stimulus vote(s), sign test p=0.109), mean preference +20.0% across 8 paired run(s) — not credible (sign test p=0.109 > 0.05)
Next action: The signal favors the target but is inconsistent; inspect tied or lost scenarios and fix weak behavior.
Why: Net win +50.0% (5W/2T/1L over 8 preference-eligible stimulus vote(s), sign test p=0.109), mean preference +35.0% across 8 paired run(s) — not credible (sign test p=0.109 > 0.05)
Next action: The signal favors the target but is inconsistent; inspect tied or lost scenarios and fix weak behavior.
▲ Recognize TreatAsLocalProperty overuse versus one justified entry
Eligible
+100.0%
+40.0%
1/0/0
= Redirect a compile-time slowdown mistakenly framed as an evaluation problem
Eligible
+0.0%
+0.0%
0/1/0
= Triage which of two property functions actually costs evaluation time
Eligible
+0.0%
+0.0%
0/1/0
Illustrative judge evidence:
Diagnose stacked evaluation-time patterns (deep imports, broad glob, file-I/O property function): Response A better addresses the rubric criteria, particularly in providing specific fix recommendations (restricting globs, moving I/O to execution phase) and explaining evaluation overhead. While Response B takes a more empirical approach with measurements showing 305 ms eval...
This is one example, not the aggregate verdict. Open Full Results for every judgment.
Why: Net win +28.6% (3W/3T/1L over 7 preference-eligible stimulus vote(s), sign test p=0.312), mean preference +15.6% across 9 paired run(s), 2 dormancy stimulus/stimuli excluded from preference — not credible — 3 of 7 preference-eligible stimulus vote(s) tied, leaving only 4 discordant preference vote(s). The sign test conditions on non-tie stimulus votes and cannot reach 0.05 below 5, so no record could have passed here — this is not a measured null. Either the skill is inert on these scenarios (make them discriminate) or the eval needs more distinct stimuli to clear the ties
Next action: The signal favors the target, but ties leave too few discordant tasks; inspect ties and predeclare more discriminating breadth.
Why: Net win +14.3% (2W/4T/1L over 7 preference-eligible stimulus vote(s), sign test p=0.500), mean preference +4.4% across 9 paired run(s), 2 dormancy stimulus/stimuli excluded from preference — not credible — 4 of 7 preference-eligible stimulus vote(s) tied, leaving only 3 discordant preference vote(s). The sign test conditions on non-tie stimulus votes and cannot reach 0.05 below 5, so no record could have passed here — this is not a measured null. Either the skill is inert on these scenarios (make them discriminate) or the eval needs more distinct stimuli to clear the ties
Next action: The signal favors the target, but ties leave too few discordant tasks; inspect ties and predeclare more discriminating breadth.
Why: Net win +14.3% (2W/4T/1L over 7 preference-eligible stimulus vote(s), sign test p=0.500), mean preference +0.0% across 8 paired run(s), 1 dormancy stimulus/stimuli excluded from preference — not credible — 4 of 7 preference-eligible stimulus vote(s) tied, leaving only 3 discordant preference vote(s). The sign test conditions on non-tie stimulus votes and cannot reach 0.05 below 5, so no record could have passed here — this is not a measured null. Either the skill is inert on these scenarios (make them discriminate) or the eval needs more distinct stimuli to clear the ties
Next action: The signal favors the target, but ties leave too few discordant tasks; inspect ties and predeclare more discriminating breadth.
Why: Net win +0.0% (3W/1T/3L over 7 preference-eligible stimulus vote(s), sign test p=0.656), mean preference +0.0% across 8 paired run(s), 1 dormancy stimulus/stimuli excluded from preference — no improvement
Next action: Compare winning and losing scenarios to isolate where the target helps versus hurts.
Repeated-run reliability (not used by the gate): 8 paired runs (3W/2T/3L).
Weak or warning scenarios:
Scenario
Preference gate
Net win
Δ Pref
Runs (W/T/L)
= Decline a Python runtime report-path problem
Excluded (activation contract)
+0.0%
+0.0%
0/1/0
▼ Diagnose missing clean tracking for generated source
Eligible
-100.0%
-40.0%
0/0/1
▼ Diagnose missing output registration for generated non-code file
Eligible
-100.0%
-40.0%
0/0/1
= Diagnose project-level glob for generated source
Eligible
+0.0%
+0.0%
0/1/0
▼ Diagnose wrong hook for generated source files
Eligible
-100.0%
-40.0%
0/0/1
Illustrative judge evidence:
Diagnose missing clean tracking for generated source: Both responses provide the correct solution and explanation. However, Response A offers slightly greater technical depth by explaining the MSBuild mechanism more explicitly (_CleanRecordFileWrites, the file list persistence location, and how clean reads that list). While Respo...
This is one example, not the aggregate verdict. Open Full Results for every judgment.
Why: Net win +22.2% (3W/5T/1L over 9 preference-eligible stimulus vote(s), sign test p=0.312), mean preference +8.9% across 9 paired run(s) — not credible — 5 of 9 preference-eligible stimulus vote(s) tied, leaving only 4 discordant preference vote(s). The sign test conditions on non-tie stimulus votes and cannot reach 0.05 below 5, so no record could have passed here — this is not a measured null. Either the skill is inert on these scenarios (make them discriminate) or the eval needs more distinct stimuli to clear the ties
Next action: The signal favors the target, but ties leave too few discordant tasks; inspect ties and predeclare more discriminating breadth.
= Identify a volatile output path that defeats incrementality
Eligible
+0.0%
+0.0%
0/1/0
= Read a diagnostic log to find the stale input
Eligible
+0.0%
+0.0%
0/1/0
Illustrative judge evidence:
Diagnose custom targets that always rerun: Both answer the requested diagnosis and fix correctly. A is marginally better because it avoids B's unnecessary and potentially incorrect FileWrites claim.
This is one example, not the aggregate verdict. Open Full Results for every judgment.
Why: Net win +33.3% (4W/4T/1L over 9 preference-eligible stimulus vote(s), sign test p=0.188), mean preference +13.3% across 9 paired run(s) — not credible (sign test p=0.188 > 0.05)
Next action: The signal favors the target but is inconsistent; inspect tied or lost scenarios and fix weak behavior.
Why: Net win +0.0% (2W/2T/2L over 6 preference-eligible stimulus vote(s), sign test p=0.687), mean preference +5.7% across 7 paired run(s), 1 dormancy stimulus/stimuli excluded from preference — no improvement
Next action: Compare winning and losing scenarios to isolate where the target helps versus hurts.
Repeated-run reliability (not used by the gate): 7 paired runs (3W/2T/2L).
Weak or warning scenarios:
Scenario
Preference gate
Net win
Δ Pref
Runs (W/T/L)
= Diagnose an ineffective Compile Remove that does not match the glob
Eligible
+0.0%
+0.0%
0/1/0
▲ Diagnose real and claimed item problems in a code generation pipeline
Eligible
+100.0%
+40.0%
1/0/0
▼ Fix item management anti-patterns
Eligible
-100.0%
-100.0%
0/0/1
= Leave already-correct item management unchanged
Eligible
+0.0%
+0.0%
0/1/0
▼ Leave correct single-list batching unchanged
Eligible
-100.0%
-40.0%
0/0/1
Illustrative judge evidence:
Diagnose an ineffective Compile Remove that does not match the glob: Position-swap inconsistent (forward: tie, reverse: A). Defaulting to tie.
This is one example, not the aggregate verdict. Open Full Results for every judgment.
➖ Mixed evidence — item-management (gpt-5.6-luna)
Why: Net win +0.0% (1W/4T/1L over 6 preference-eligible stimulus vote(s), sign test p=0.750), mean preference +8.6% across 7 paired run(s), 1 dormancy stimulus/stimuli excluded from preference — no improvement
Next action: Compare winning and losing scenarios to isolate where the target helps versus hurts.
Repeated-run reliability (not used by the gate): 7 paired runs (1W/5T/1L).
Weak or warning scenarios:
Scenario
Preference gate
Net win
Δ Pref
Runs (W/T/L)
= Decline JavaScript configuration merging
Excluded (activation contract)
+0.0%
+0.0%
0/1/0
▼ Diagnose an ineffective Compile Remove that does not match the glob
Eligible
-100.0%
-40.0%
0/0/1
= Diagnose item group and batching issues
Eligible
+0.0%
+0.0%
0/1/0
= Diagnose real and claimed item problems in a code generation pipeline
Eligible
+0.0%
+0.0%
0/1/0
= Fix item management anti-patterns
Eligible
+0.0%
+0.0%
0/1/0
= Leave correct single-list batching unchanged
Eligible
+0.0%
+0.0%
0/1/0
Illustrative judge evidence:
Diagnose an ineffective Compile Remove that does not match the glob: Both responses correctly identify the core problem (Remove pattern mismatch) and both provide working solutions. However, Response A's solution is architecturally superior: it targets *.g.cs specifically, matching the actual naming convention of the generated files, whereas ...
This is one example, not the aggregate verdict. Open Full Results for every judgment.
Why: Net win +0.0% (2W/3T/2L over 7 preference-eligible stimulus vote(s), sign test p=0.687), mean preference +5.0% across 8 paired run(s), 1 dormancy stimulus/stimuli excluded from preference — no improvement
Next action: Compare winning and losing scenarios to isolate where the target helps versus hurts.
Repeated-run reliability (not used by the gate): 8 paired runs (3W/3T/2L).
Weak or warning scenarios:
Scenario
Preference gate
Net win
Δ Pref
Runs (W/T/L)
= Add a module to an F# project
Eligible
+0.0%
+0.0%
0/1/0
= Add a signature file to define public API
Eligible
+0.0%
+0.0%
0/1/0
▼ Distinguish a style backslash from a real cross-platform backslash bug
Eligible
-100.0%
-40.0%
0/0/1
= Fix broken file order causing FS0039
Eligible
+0.0%
+0.0%
0/1/0
▲ Judge an unguarded import inside a NuGet package build folder
Eligible
+100.0%
+40.0%
1/0/0
▼ Leave a clean project without inventing anti-patterns
Eligible
-100.0%
-40.0%
0/0/1
Illustrative judge evidence:
Add a module to an F# project: The implementations and verification outcomes are materially equivalent: both fulfill the requested validation, project wiring, pre-processing control flow, and successful compilation/run.
This is one example, not the aggregate verdict. Open Full Results for every judgment.
Why: Net win +14.3% (3W/2T/2L over 7 preference-eligible stimulus vote(s), sign test p=0.500), mean preference -2.9% across 7 paired run(s) — not credible (sign test p=0.500 > 0.05)
Next action: The signal favors the target but is inconsistent; inspect tied or lost scenarios and fix weak behavior.
To investigate non-passing or warning results, paste this to your AI coding agent:
For PR 1215 in dotnet/skills, download eval artifacts with gh run download 37303425464 --repo dotnet/skills --pattern "vally-results-*" --dir ./eval-results, then fetch https://raw.githubusercontent.com/dotnet/skills/b97f9c542191718f6b4d15e24c1be7af9908b9e8/eng/vally-adapter/InvestigatingResults.md and follow it. Classify each result as measurement-invalid, underpowered, not-proven, preference-loss, or passing-with-warning. Use stateReason, result accounting, weak scenarios, and judge evidence to give the cause and exact next fix.
⚠️ Session replay telemetry was not published because the auxiliary dotnet/skills-data publisher failed. The evaluation verdicts above remain authoritative; maintainers must repair SKILLS_DATA_TOKEN or the publisher before replay links are available.
Objective completion gate: not enabled. Aggregate completion transitions are telemetry only, so this report does not claim that zero objective regressions were proven.
A result passes only when preference-eligible distinct-stimulus votes have aggregate net win of at least 20% and an exact one-sided sign-test result of p ≤ 0.05, and every explicit dormancy activation contract passes. Repeated runs measure reliability only.
Target
Model
Verdict
Gate evidence
Overfit
Warnings
Next action
agent.msbuild
claude-sonnet-5
➖ Improvement signal, tie-limited
n=5; 2W/3T/0L; d=2; p=0.250; net +40.0%
—
—
The signal favors the target, but ties leave too few discordant tasks; inspect ties and predeclare more discriminating breadth.
agent.msbuild
gpt-5.6-luna
➖ Improvement signal, tie-limited
n=5; 3W/1T/1L; d=4; p=0.312; net +40.0%
—
—
The signal favors the target, but ties leave too few discordant tasks; inspect ties and predeclare more discriminating breadth.
binlog-failure-analysis
claude-sonnet-5
➖ Improvement signal, tie-limited
n=6; 2W/3T/1L; d=3; p=0.500; net +16.7%; 1 dormancy excluded
🟡 0.42
Activation: isolated 6/7; plugin 6/7
The signal favors the target, but ties leave too few discordant tasks; inspect ties and predeclare more discriminating breadth.
binlog-failure-analysis
gpt-5.6-luna
➖ Improvement signal, tie-limited
n=6; 3W/3T/0L; d=3; p=0.125; net +50.0%; 1 dormancy excluded
🟡 0.35
Activation: isolated 6/7; plugin 6/7
The signal favors the target, but ties leave too few discordant tasks; inspect ties and predeclare more discriminating breadth.
binlog-generation
claude-sonnet-5
➖ Improvement signal, unproven
n=7; 5W/1T/1L; d=6; p=0.109; net +57.1%; 1 dormancy excluded
🟡 0.48
Activation: isolated 6/7; plugin 5/7
The signal favors the target but is inconsistent; inspect tied or lost scenarios and fix weak behavior.
binlog-generation
gpt-5.6-luna
✅ Improved
n=7; 5W/2T/0L; d=5; p=0.031; net +71.4%; 1 dormancy excluded
🟡 0.26
Activation: isolated 6/7; plugin 6/7
Fix activation gaps; Review overfit evidence.
build-parallelism
claude-sonnet-5
➖ Mixed evidence
n=6; 2W/2T/2L; d=4; p=0.687; net +0.0%; 1 dormancy excluded
🟡 0.47
Activation: isolated 6/6; plugin 4/6
Compare winning and losing scenarios to isolate where the target helps versus hurts.
build-parallelism
gpt-5.6-luna
➖ Baseline signal, tie-limited
n=6; 1W/2T/3L; d=4; p=0.312; net -33.3%; 1 dormancy excluded
Inspect activation-only failed runs before rewriting skill content; the model stopped after loading a skill.
build-perf-baseline
claude-sonnet-5
➖ Baseline signal, tie-limited
n=6; 1W/3T/2L; d=3; p=0.500; net -16.7%; 1 dormancy excluded
🔴 0.51
Activation: isolated 6/6; plugin 5/6
Evidence leans baseline but is not credible; inspect losses and make tied tasks discriminate.
build-perf-baseline
gpt-5.6-luna
➖ Improvement signal, tie-limited
n=6; 4W/2T/0L; d=4; p=0.063; net +66.7%; 1 dormancy excluded
🟡 0.28
Activation: isolated 6/6; plugin 3/6
The signal favors the target, but ties leave too few discordant tasks; inspect ties and predeclare more discriminating breadth.
build-perf-diagnostics
claude-sonnet-5
➖ Improvement signal, tie-limited
n=7; 2W/4T/1L; d=3; p=0.500; net +14.3%; 1 dormancy excluded
🟡 0.35
Activation: isolated 7/7; plugin 3/7
The signal favors the target, but ties leave too few discordant tasks; inspect ties and predeclare more discriminating breadth.
build-perf-diagnostics
gpt-5.6-luna
✅ Improved
n=7; 5W/2T/0L; d=5; p=0.031; net +71.4%; 1 dormancy excluded
🟡 0.45
Activation: isolated 7/7; plugin 3/7
Fix activation gaps; Review overfit evidence.
check-bin-obj-clash
claude-sonnet-5
➖ Baseline signal, unproven
n=7; 1W/2T/4L; d=5; p=0.188; net -42.9%
✅ 0.17
Activation: isolated 7/7; plugin 4/7
Evidence leans baseline but is not credible; inspect losing scenarios for recurring defects.
check-bin-obj-clash
gpt-5.6-luna
➖ Improvement signal, unproven
n=7; 4W/2T/1L; d=5; p=0.188; net +42.9%
🟡 0.30
—
The signal favors the target but is inconsistent; inspect tied or lost scenarios and fix weak behavior.
directory-build-organization
claude-sonnet-5
➖ Improvement signal, unproven
n=6; 3W/1T/2L; d=5; p=0.500; net +16.7%; 1 dormancy excluded
🟡 0.43
Activation: isolated 3/6; plugin 4/6
The signal favors the target but is inconsistent; inspect tied or lost scenarios and fix weak behavior.
directory-build-organization
gpt-5.6-luna
➖ Mixed evidence
n=6; 2W/2T/2L; d=4; p=0.687; net +0.0%; 1 dormancy excluded
🔴 0.51
Activation: isolated 6/6; plugin 4/6
Compare winning and losing scenarios to isolate where the target helps versus hurts.
eval-performance
claude-sonnet-5
✅ Improved
n=8; 5W/3T/0L; d=5; p=0.031; net +62.5%
🟡 0.32
Activation: isolated 5/8; plugin 5/8
Fix activation gaps; Review overfit evidence.
eval-performance
gpt-5.6-luna
✅ Improved
n=8; 5W/3T/0L; d=5; p=0.031; net +62.5%
🟡 0.32
—
Review overfit evidence.
extension-points
claude-sonnet-5
➖ Improvement signal, unproven
n=7; 4W/1T/2L; d=6; p=0.344; net +28.6%; 2 dormancy excluded
🟡 0.23
Activation: isolated 5/7; plugin 7/7
The signal favors the target but is inconsistent; inspect tied or lost scenarios and fix weak behavior.
extension-points
gpt-5.6-luna
➖ Improvement signal, tie-limited
n=7; 2W/4T/1L; d=3; p=0.500; net +14.3%; 2 dormancy excluded
🟡 0.47
Activation: isolated 7/7; plugin 6/7
The signal favors the target, but ties leave too few discordant tasks; inspect ties and predeclare more discriminating breadth.
including-generated-files
claude-sonnet-5
➖ Improvement signal, tie-limited
n=7; 3W/3T/1L; d=4; p=0.312; net +28.6%; 1 dormancy excluded
🟡 0.39
Activation: isolated 5/7; plugin 4/7
The signal favors the target, but ties leave too few discordant tasks; inspect ties and predeclare more discriminating breadth.
including-generated-files
gpt-5.6-luna
➖ Improvement signal, tie-limited
n=7; 2W/4T/1L; d=3; p=0.500; net +14.3%; 1 dormancy excluded
🔴 0.53
Activation: isolated 7/7; plugin 6/7
The signal favors the target, but ties leave too few discordant tasks; inspect ties and predeclare more discriminating breadth.
incremental-build
claude-sonnet-5
✅ Improved
n=9; 6W/3T/0L; d=6; p=0.016; net +66.7%
🟡 0.35
Activation: isolated 4/9; plugin 4/9
Fix activation gaps; Review overfit evidence.
incremental-build
gpt-5.6-luna
➖ Improvement signal, tie-limited
n=9; 4W/5T/0L; d=4; p=0.063; net +44.4%
🟡 0.21
Activation: isolated 8/9; plugin 7/9
The signal favors the target, but ties leave too few discordant tasks; inspect ties and predeclare more discriminating breadth.
item-management
claude-sonnet-5
➖ Improvement signal, tie-limited
n=6; 3W/3T/0L; d=3; p=0.125; net +50.0%; 1 dormancy excluded
🟡 0.29
Activation: isolated 5/6; plugin 4/6
The signal favors the target, but ties leave too few discordant tasks; inspect ties and predeclare more discriminating breadth.
item-management
gpt-5.6-luna
➖ Improvement signal, tie-limited
n=6; 4W/2T/0L; d=4; p=0.063; net +66.7%; 1 dormancy excluded
🟡 0.41
—
The signal favors the target, but ties leave too few discordant tasks; inspect ties and predeclare more discriminating breadth.
msbuild-antipatterns
claude-sonnet-5
➖ Improvement signal, tie-limited
n=7; 1W/6T/0L; d=1; p=0.500; net +14.3%; 1 dormancy excluded
🟡 0.31
Activation: isolated 3/7; plugin 3/7
The signal favors the target, but ties leave too few discordant tasks; inspect ties and predeclare more discriminating breadth.
msbuild-antipatterns
gpt-5.6-luna
➖ Improvement signal, tie-limited
n=7; 3W/4T/0L; d=3; p=0.125; net +42.9%; 1 dormancy excluded
🟡 0.49
Activation: isolated 5/7; plugin 4/7
The signal favors the target, but ties leave too few discordant tasks; inspect ties and predeclare more discriminating breadth.
property-patterns
claude-sonnet-5
➖ Baseline signal, tie-limited
n=7; 0W/4T/3L; d=3; p=0.125; net -42.9%
✅ 0.17
Activation: isolated 5/7; plugin 6/7
Evidence leans baseline but is not credible; inspect losses and make tied tasks discriminate.
property-patterns
gpt-5.6-luna
➖ Mixed evidence
n=7; 2W/3T/2L; d=4; p=0.687; net +0.0%
🟡 0.23
Activation: isolated 6/7; plugin 6/7
Compare winning and losing scenarios to isolate where the target helps versus hurts.
ℹ️ How to read this report
✅ Improved — the result passed both the statistical gate and the 20% practical net-win floor.
➖ No clear winner — the result is valid but did not pass both gates. The label distinguishes all ties, mixed evidence, directional but unproven evidence, and credible effects below the practical floor.
⚠️ Invalid / underpowered — the gate withheld a quality verdict. Fix the measurement before judging the target.
⛔ Activation contract failed — the isolated target activated on an explicit dormancy scenario. Dormancy preference is excluded, but this routing failure still blocks a pass.
📉 Preference loss — the LLM judge credibly preferred baseline. It is report-only, not objective completion proof.
Gate evidence — n preference-eligible distinct-stimulus votes, W/T/L stimulus votes, d discordant votes, exact one-sided p, net win, and the count of separately retained dormancy stimuli. The p value applies to one model/target result; no matrix-wide multiple-comparison correction is applied.
Warnings — activation, timeout, retry recovery, or unresolved comparison conditions that need attention.
Do not add repeated runs to increase statistical power. Do not add stimuli after seeing a near-pass unless the new breadth is predeclared for a new experiment.
Why: Net win +40.0% (2W/3T/0L over 5 preference-eligible stimulus vote(s), sign test p=0.250), mean preference +28.0% across 5 paired run(s) — not credible — 3 of 5 preference-eligible stimulus vote(s) tied, leaving only 2 discordant preference vote(s). The sign test conditions on non-tie stimulus votes and cannot reach 0.05 below 5, so no record could have passed here — this is not a measured null. Either the agent is inert on these scenarios (make them discriminate) or the eval needs more distinct stimuli to clear the ties
Next action: The signal favors the target, but ties leave too few discordant tasks; inspect ties and predeclare more discriminating breadth.
Why: Net win +40.0% (3W/1T/1L over 5 preference-eligible stimulus vote(s), sign test p=0.312), mean preference +16.0% across 5 paired run(s) — not credible — 1 of 5 preference-eligible stimulus vote(s) tied, leaving only 4 discordant preference vote(s). The sign test conditions on non-tie stimulus votes and cannot reach 0.05 below 5, so no record could have passed here — this is not a measured null. Either the agent is inert on these scenarios (make them discriminate) or the eval needs more distinct stimuli to clear the ties
Next action: The signal favors the target, but ties leave too few discordant tasks; inspect ties and predeclare more discriminating breadth.
Why: Net win +16.7% (2W/3T/1L over 6 preference-eligible stimulus vote(s), sign test p=0.500), mean preference +11.4% across 7 paired run(s), 1 dormancy stimulus/stimuli excluded from preference — not credible — 3 of 6 preference-eligible stimulus vote(s) tied, leaving only 3 discordant preference vote(s). The sign test conditions on non-tie stimulus votes and cannot reach 0.05 below 5, so no record could have passed here — this is not a measured null. Either the skill is inert on these scenarios (make them discriminate) or the eval needs more distinct stimuli to clear the ties
Next action: The signal favors the target, but ties leave too few discordant tasks; inspect ties and predeclare more discriminating breadth.
Why: Net win +50.0% (3W/3T/0L over 6 preference-eligible stimulus vote(s), sign test p=0.125), mean preference +22.9% across 7 paired run(s), 1 dormancy stimulus/stimuli excluded from preference — not credible — 3 of 6 preference-eligible stimulus vote(s) tied, leaving only 3 discordant preference vote(s). The sign test conditions on non-tie stimulus votes and cannot reach 0.05 below 5, so no record could have passed here — this is not a measured null. Either the skill is inert on these scenarios (make them discriminate) or the eval needs more distinct stimuli to clear the ties
Next action: The signal favors the target, but ties leave too few discordant tasks; inspect ties and predeclare more discriminating breadth.
Why: Net win +57.1% (5W/1T/1L over 7 preference-eligible stimulus vote(s), sign test p=0.109), mean preference +42.5% across 8 paired run(s), 1 dormancy stimulus/stimuli excluded from preference — not credible (sign test p=0.109 > 0.05)
Next action: The signal favors the target but is inconsistent; inspect tied or lost scenarios and fix weak behavior.
Repeated-run reliability (not used by the gate): 8 paired runs (5W/2T/1L).
Weak or warning scenarios:
Scenario
Preference gate
Net win
Δ Pref
Runs (W/T/L)
= Choose a predictable, non-colliding binlog name for a CI upload step
Eligible
+0.0%
+0.0%
0/1/0
= Decline Maven compiler configuration
Excluded (activation contract)
+0.0%
+0.0%
0/1/0
▲ Preserve binlog history while cleaning stale build output
Eligible
+100.0%
+40.0%
1/0/0
▼ Recognize that a failed build produced no binlog at all
Eligible
-100.0%
-40.0%
0/0/1
Illustrative judge evidence:
Choose a predictable, non-colliding binlog name for a CI upload step: Both responses fully satisfy the task and every rubric requirement, reporting the same valid, explicitly named non-colliding Debug and Release binlogs after successful builds.
This is one example, not the aggregate verdict. Open Full Results for every judgment.
Why: Net win +0.0% (2W/2T/2L over 6 preference-eligible stimulus vote(s), sign test p=0.687), mean preference +8.6% across 7 paired run(s), 1 dormancy stimulus/stimuli excluded from preference — no improvement
Next action: Compare winning and losing scenarios to isolate where the target helps versus hurts.
Repeated-run reliability (not used by the gate): 7 paired runs (2W/3T/2L).
Weak or warning scenarios:
Scenario
Preference gate
Net win
Δ Pref
Runs (W/T/L)
▼ Analyze build parallelism bottlenecks
Eligible
-100.0%
-40.0%
0/0/1
= Decline database query tuning request
Excluded (activation contract)
+0.0%
+0.0%
0/1/0
= Enable BuildInParallel on a custom MSBuild task
Eligible
+0.0%
+0.0%
0/1/0
= Enable parallel nodes for a wide project graph
Eligible
+0.0%
+0.0%
0/1/0
▼ Reduce CI build scope with a solution filter
Eligible
-100.0%
-40.0%
0/0/1
Illustrative judge evidence:
Analyze build parallelism bottlenecks: The substantive conclusions are the same and correct, but A better connects its conclusion to the requested binlog evidence and more explicitly explains why parallel workers cannot break the serial chain.
This is one example, not the aggregate verdict. Open Full Results for every judgment.
Why: Net win -33.3% (1W/2T/3L over 6 preference-eligible stimulus vote(s), sign test p=0.312), mean preference -17.1% across 7 paired run(s), 1 dormancy stimulus/stimuli excluded from preference — no improvement
Next action: Inspect activation-only failed runs before rewriting skill content; the model stopped after loading a skill.
Repeated-run reliability (not used by the gate): 7 paired runs (1W/2T/4L).
Weak or warning scenarios:
Scenario
Preference gate
Net win
Δ Pref
Runs (W/T/L)
▼ Analyze build parallelism bottlenecks
Eligible
-100.0%
-40.0%
0/0/1
▼ Decline database query tuning request
Excluded (activation contract)
-100.0%
-40.0%
0/0/1
= Decline graph build for runtime-discovered projects
Eligible
+0.0%
+0.0%
0/1/0
= Enable BuildInParallel on a custom MSBuild task
Eligible
+0.0%
+0.0%
0/1/0
▼ Enable parallel nodes for a wide project graph
Eligible
-100.0%
-40.0%
0/0/1
▼ Reduce CI build scope with a solution filter
Eligible
-100.0%
-40.0%
0/0/1
Illustrative judge evidence:
Analyze build parallelism bottlenecks: Both responses correctly analyze the build parallelism and identify the same core findings: the Core → Api → Web → Tests critical path limits parallelism, and Tests → Api is a redundant but non-critical-path reference. Response A edges ahead by providing quantitative evidence—...
This is one example, not the aggregate verdict. Open Full Results for every judgment.
Why: Net win -16.7% (1W/3T/2L over 6 preference-eligible stimulus vote(s), sign test p=0.500), mean preference -5.7% across 7 paired run(s), 1 dormancy stimulus/stimuli excluded from preference — no improvement
Next action: Evidence leans baseline but is not credible; inspect losses and make tied tasks discriminate.
Why: Net win +66.7% (4W/2T/0L over 6 preference-eligible stimulus vote(s), sign test p=0.063), mean preference +25.7% across 7 paired run(s), 1 dormancy stimulus/stimuli excluded from preference — not credible — 2 of 6 preference-eligible stimulus vote(s) tied, leaving only 4 discordant preference vote(s). The sign test conditions on non-tie stimulus votes and cannot reach 0.05 below 5, so no record could have passed here — this is not a measured null. Either the skill is inert on these scenarios (make them discriminate) or the eval needs more distinct stimuli to clear the ties
Next action: The signal favors the target, but ties leave too few discordant tasks; inspect ties and predeclare more discriminating breadth.
Why: Net win +14.3% (2W/4T/1L over 7 preference-eligible stimulus vote(s), sign test p=0.500), mean preference +5.0% across 8 paired run(s), 1 dormancy stimulus/stimuli excluded from preference — not credible — 4 of 7 preference-eligible stimulus vote(s) tied, leaving only 3 discordant preference vote(s). The sign test conditions on non-tie stimulus votes and cannot reach 0.05 below 5, so no record could have passed here — this is not a measured null. Either the skill is inert on these scenarios (make them discriminate) or the eval needs more distinct stimuli to clear the ties
Next action: The signal favors the target, but ties leave too few discordant tasks; inspect ties and predeclare more discriminating breadth.
Why: Net win -42.9% (1W/2T/4L over 7 preference-eligible stimulus vote(s), sign test p=0.188), mean preference -25.7% across 7 paired run(s) — no improvement
Next action: Evidence leans baseline but is not credible; inspect losing scenarios for recurring defects.
Gate evidence: n=7; 1W/2T/4L; d=5; p=0.188; net -42.9%
Warnings: Activation: isolated 7/7; plugin 4/7
Overfit: Low (score 0.17)
Repeated-run reliability (not used by the gate): 7 paired runs (1W/2T/4L).
Weak or warning scenarios:
Scenario
Preference gate
Net win
Δ Pref
Runs (W/T/L)
▼ Audit the mixed solution and separate safe projects from unsafe ones
Eligible
-100.0%
-100.0%
0/0/1
▼ Avoid a false clash report when projects share only the top-level artifacts root
Eligible
-100.0%
-40.0%
0/0/1
▼ Decline output-clash remediation for separate projects using the default SDK layout
Eligible
-100.0%
-40.0%
0/0/1
= Diagnose multi-targeting outputs that collapse into one path
Eligible
+0.0%
+0.0%
0/1/0
▲ Diagnose redundant project reference metadata that forks a same-path build
Eligible
+100.0%
+40.0%
1/0/0
▼ Diagnose shared output and intermediate path collision
Eligible
-100.0%
-40.0%
0/0/1
= Fix all clash mechanisms in the mixed solution
Eligible
+0.0%
+0.0%
0/1/0
Illustrative judge evidence:
Audit the mixed solution and separate safe projects from unsafe ones: A gets the central MultiTargetLib diagnosis right as well as the LibraryA/LibraryB collision and identifies the ConsumerApp reference issue. B's erroneous declaration that MultiTargetLib is safe misses a major required unsafe case, outweighing its slightly clearer unsafe label...
This is one example, not the aggregate verdict. Open Full Results for every judgment.
Why: Net win +42.9% (4W/2T/1L over 7 preference-eligible stimulus vote(s), sign test p=0.188), mean preference +17.1% across 7 paired run(s) — not credible (sign test p=0.188 > 0.05)
Next action: The signal favors the target but is inconsistent; inspect tied or lost scenarios and fix weak behavior.
Gate evidence: n=7; 4W/2T/1L; d=5; p=0.188; net +42.9%
Overfit: Moderate (score 0.30)
Repeated-run reliability (not used by the gate): 7 paired runs (4W/2T/1L).
Weak or warning scenarios:
Scenario
Preference gate
Net win
Δ Pref
Runs (W/T/L)
▼ Audit the mixed solution and separate safe projects from unsafe ones
Eligible
-100.0%
-40.0%
0/0/1
= Diagnose redundant project reference metadata that forks a same-path build
Eligible
+0.0%
+0.0%
0/1/0
= Diagnose shared output and intermediate path collision
Eligible
+0.0%
+0.0%
0/1/0
Illustrative judge evidence:
Audit the mixed solution and separate safe projects from unsafe ones: Response A provides a more nuanced and clearly-structured analysis. Its key advantage is treating the ConsumerApp→ToolLib reference issue as a separate unsafe item, while explicitly confirming that ToolLib's own project configuration is safe. This distinction between inherent ...
This is one example, not the aggregate verdict. Open Full Results for every judgment.
Why: Net win +16.7% (3W/1T/2L over 6 preference-eligible stimulus vote(s), sign test p=0.500), mean preference +2.9% across 7 paired run(s), 1 dormancy stimulus/stimuli excluded from preference — not credible (sign test p=0.500 > 0.05)
Next action: The signal favors the target but is inconsistent; inspect tied or lost scenarios and fix weak behavior.
Why: Net win +0.0% (2W/2T/2L over 6 preference-eligible stimulus vote(s), sign test p=0.687), mean preference +8.6% across 7 paired run(s), 1 dormancy stimulus/stimuli excluded from preference — no improvement
Next action: Compare winning and losing scenarios to isolate where the target helps versus hurts.
Why: Net win +28.6% (4W/1T/2L over 7 preference-eligible stimulus vote(s), sign test p=0.344), mean preference +2.2% across 9 paired run(s), 2 dormancy stimulus/stimuli excluded from preference — not credible (sign test p=0.344 > 0.05)
Next action: The signal favors the target but is inconsistent; inspect tied or lost scenarios and fix weak behavior.
Why: Net win +14.3% (2W/4T/1L over 7 preference-eligible stimulus vote(s), sign test p=0.500), mean preference +6.7% across 9 paired run(s), 2 dormancy stimulus/stimuli excluded from preference — not credible — 4 of 7 preference-eligible stimulus vote(s) tied, leaving only 3 discordant preference vote(s). The sign test conditions on non-tie stimulus votes and cannot reach 0.05 below 5, so no record could have passed here — this is not a measured null. Either the skill is inert on these scenarios (make them discriminate) or the eval needs more distinct stimuli to clear the ties
Next action: The signal favors the target, but ties leave too few discordant tasks; inspect ties and predeclare more discriminating breadth.
= Diagnose NuGet package and repo extension conflicts
Eligible
+0.0%
+0.0%
0/1/0
= Diagnose build extension point failures
Eligible
+0.0%
+0.0%
0/1/0
= Fix extension point anti-patterns
Eligible
+0.0%
+0.0%
0/1/0
▼ Leave incremental target authoring to its owning workflow
Excluded (activation contract)
-100.0%
-40.0%
0/0/1
= Review packed layout without a false missing-file bug
Eligible
+0.0%
+0.0%
0/1/0
Illustrative judge evidence:
Create extensibility hooks for a custom SDK target file: Both responses successfully completed the task and verified working solutions. Response A is the better solution because it follows the principle of minimal, elegant design—it splits the monolithic target into CoreMySDKBuild and MySDKBuild, enabling the existing BeforeTargets/...
This is one example, not the aggregate verdict. Open Full Results for every judgment.
Why: Net win +28.6% (3W/3T/1L over 7 preference-eligible stimulus vote(s), sign test p=0.312), mean preference +10.0% across 8 paired run(s), 1 dormancy stimulus/stimuli excluded from preference — not credible — 3 of 7 preference-eligible stimulus vote(s) tied, leaving only 4 discordant preference vote(s). The sign test conditions on non-tie stimulus votes and cannot reach 0.05 below 5, so no record could have passed here — this is not a measured null. Either the skill is inert on these scenarios (make them discriminate) or the eval needs more distinct stimuli to clear the ties
Next action: The signal favors the target, but ties leave too few discordant tasks; inspect ties and predeclare more discriminating breadth.
Why: Net win +14.3% (2W/4T/1L over 7 preference-eligible stimulus vote(s), sign test p=0.500), mean preference +12.5% across 8 paired run(s), 1 dormancy stimulus/stimuli excluded from preference — not credible — 4 of 7 preference-eligible stimulus vote(s) tied, leaving only 3 discordant preference vote(s). The sign test conditions on non-tie stimulus votes and cannot reach 0.05 below 5, so no record could have passed here — this is not a measured null. Either the skill is inert on these scenarios (make them discriminate) or the eval needs more distinct stimuli to clear the ties
Next action: The signal favors the target, but ties leave too few discordant tasks; inspect ties and predeclare more discriminating breadth.
Why: Net win +44.4% (4W/5T/0L over 9 preference-eligible stimulus vote(s), sign test p=0.063), mean preference +17.8% across 9 paired run(s) — not credible — 5 of 9 preference-eligible stimulus vote(s) tied, leaving only 4 discordant preference vote(s). The sign test conditions on non-tie stimulus votes and cannot reach 0.05 below 5, so no record could have passed here — this is not a measured null. Either the skill is inert on these scenarios (make them discriminate) or the eval needs more distinct stimuli to clear the ties
Next action: The signal favors the target, but ties leave too few discordant tasks; inspect ties and predeclare more discriminating breadth.
Why: Net win +50.0% (3W/3T/0L over 6 preference-eligible stimulus vote(s), sign test p=0.125), mean preference +31.4% across 7 paired run(s), 1 dormancy stimulus/stimuli excluded from preference — not credible — 3 of 6 preference-eligible stimulus vote(s) tied, leaving only 3 discordant preference vote(s). The sign test conditions on non-tie stimulus votes and cannot reach 0.05 below 5, so no record could have passed here — this is not a measured null. Either the skill is inert on these scenarios (make them discriminate) or the eval needs more distinct stimuli to clear the ties
Next action: The signal favors the target, but ties leave too few discordant tasks; inspect ties and predeclare more discriminating breadth.
Why: Net win +66.7% (4W/2T/0L over 6 preference-eligible stimulus vote(s), sign test p=0.063), mean preference +31.4% across 7 paired run(s), 1 dormancy stimulus/stimuli excluded from preference — not credible — 2 of 6 preference-eligible stimulus vote(s) tied, leaving only 4 discordant preference vote(s). The sign test conditions on non-tie stimulus votes and cannot reach 0.05 below 5, so no record could have passed here — this is not a measured null. Either the skill is inert on these scenarios (make them discriminate) or the eval needs more distinct stimuli to clear the ties
Next action: The signal favors the target, but ties leave too few discordant tasks; inspect ties and predeclare more discriminating breadth.
Why: Net win +14.3% (1W/6T/0L over 7 preference-eligible stimulus vote(s), sign test p=0.500), mean preference +17.5% across 8 paired run(s), 1 dormancy stimulus/stimuli excluded from preference — not credible — 6 of 7 preference-eligible stimulus vote(s) tied, leaving only 1 discordant preference vote(s). The sign test conditions on non-tie stimulus votes and cannot reach 0.05 below 5, so no record could have passed here — this is not a measured null. Either the skill is inert on these scenarios (make them discriminate) or the eval needs more distinct stimuli to clear the ties
Next action: The signal favors the target, but ties leave too few discordant tasks; inspect ties and predeclare more discriminating breadth.
Repeated-run reliability (not used by the gate): 8 paired runs (2W/6T/0L).
Weak or warning scenarios:
Scenario
Preference gate
Net win
Δ Pref
Runs (W/T/L)
= Add a module to an F# project
Eligible
+0.0%
+0.0%
0/1/0
= Add a signature file to define public API
Eligible
+0.0%
+0.0%
0/1/0
= Fix broken file order causing FS0039
Eligible
+0.0%
+0.0%
0/1/0
= Judge an unguarded import inside a NuGet package build folder
Eligible
+0.0%
+0.0%
0/1/0
= Leave a clean project without inventing anti-patterns
Eligible
+0.0%
+0.0%
0/1/0
= Review MSBuild files for anti-patterns and style issues
Eligible
+0.0%
+0.0%
0/1/0
Illustrative judge evidence:
Add a module to an F# project: Both implementations satisfy the requested functionality and compile/run successfully. B's explicit failure-path sanity check is a modest verification advantage, but it does not establish a materially better final code result.
This is one example, not the aggregate verdict. Open Full Results for every judgment.
Why: Net win +42.9% (3W/4T/0L over 7 preference-eligible stimulus vote(s), sign test p=0.125), mean preference +20.0% across 8 paired run(s), 1 dormancy stimulus/stimuli excluded from preference — not credible — 4 of 7 preference-eligible stimulus vote(s) tied, leaving only 3 discordant preference vote(s). The sign test conditions on non-tie stimulus votes and cannot reach 0.05 below 5, so no record could have passed here — this is not a measured null. Either the skill is inert on these scenarios (make them discriminate) or the eval needs more distinct stimuli to clear the ties
Next action: The signal favors the target, but ties leave too few discordant tasks; inspect ties and predeclare more discriminating breadth.
Why: Net win -42.9% (0W/4T/3L over 7 preference-eligible stimulus vote(s), sign test p=0.125), mean preference -25.7% across 7 paired run(s) — no improvement
Next action: Evidence leans baseline but is not credible; inspect losses and make tied tasks discriminate.
Choose the correct shared build file for a post-build target: Both are substantively correct and answer the organizational question well. A is modestly stronger because it gives the same useful placement rule while actually adding and validating the requested post-build target, with the existing props content left intact.
This is one example, not the aggregate verdict. Open Full Results for every judgment.
Why: Net win +0.0% (2W/3T/2L over 7 preference-eligible stimulus vote(s), sign test p=0.687), mean preference -8.6% across 7 paired run(s) — no improvement
Next action: Compare winning and losing scenarios to isolate where the target helps versus hurts.
Choose the right OS detection for a cross-platform property: Response A edges ahead by successfully implementing and verifying the fix—the build completed with zero errors, proving the solution works in practice. While Response B has marginally better capitalization convention and more complete XML formatting, practical validation is mo...
This is one example, not the aggregate verdict. Open Full Results for every judgment.
✅ Improved — binlog-generation (gpt-5.6-luna)
Why: Net win +71.4% (5W/2T/0L over 7 preference-eligible stimulus vote(s), sign test p=0.031), mean preference +55.0% across 8 paired run(s), 1 dormancy stimulus/stimuli excluded from preference — credibly better
Next action: Fix activation gaps; Review overfit evidence.
Repeated-run reliability (not used by the gate): 8 paired runs (5W/3T/0L).
Weak or warning scenarios:
Scenario
Preference gate
Net win
Δ Pref
Runs (W/T/L)
= Choose a predictable, non-colliding binlog name for a CI upload step
Eligible
+0.0%
+0.0%
0/1/0
= Decline Maven compiler configuration
Excluded (activation contract)
+0.0%
+0.0%
0/1/0
= Fix a CI build script that reuses the same binlog on every run
Eligible
+0.0%
+0.0%
0/1/0
▲ Preserve binlog history while cleaning stale build output
Eligible
+100.0%
+100.0%
1/0/0
Illustrative judge evidence:
Choose a predictable, non-colliding binlog name for a CI upload step: Position-swap inconsistent (forward: A, reverse: tie). Defaulting to tie.
This is one example, not the aggregate verdict. Open Full Results for every judgment.
✅ Improved — eval-performance (gpt-5.6-luna)
Why: Net win +62.5% (5W/3T/0L over 8 preference-eligible stimulus vote(s), sign test p=0.031), mean preference +47.5% across 8 paired run(s) — credibly better
To investigate non-passing or warning results, paste this to your AI coding agent:
For PR 1215 in dotnet/skills, download eval artifacts with gh run download 37311852075 --repo dotnet/skills --pattern "vally-results-*" --dir ./eval-results, then fetch https://raw.githubusercontent.com/dotnet/skills/b9052fd631b50ceda0f8f4d3033bc3b02fb97671/eng/vally-adapter/InvestigatingResults.md and follow it. Classify each result as measurement-invalid, underpowered, not-proven, preference-loss, or passing-with-warning. Use stateReason, result accounting, weak scenarios, and judge evidence to give the cause and exact next fix.
⚠️ Session replay telemetry was not published because the auxiliary dotnet/skills-data publisher failed. The evaluation verdicts above remain authoritative; maintainers must repair SKILLS_DATA_TOKEN or the publisher before replay links are available.
✅ Approved by @YuliiaKovalova. cc @dotnet/skills-merge-approvers — ready to merge.
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Summary
This stacked change sharpens 14 MSBuild skill guides on top of #1214. It fixes command syntax, corrects build-parallelism guidance, and defines clearer routing boundaries among shared-file organization, property changes, imports and hooks, F# project ordering, generated files, item operations, and performance workflows.
Why
The PR1 scenarios exposed three guidance groups that needed correction: reliable PowerShell and file-logger syntax; accurate node and dependency advice; and routing descriptions that separate sibling skills by the requested outcome.
The final boundaries are explicit.
directory-build-organizationowns existingDirectory.Build.*discovery, hierarchy/import timing, and props-versus-targets relocation, whileproperty-patternsowns concrete value changes that stay within the current layout.extension-pointsowns NuGet auto-import and packed-layout discovery, whilemsbuild-antipatternsowns proven unsafe imports during audits and directly claims.fs,.fsi, andFS0039compile-order tasks.Impact
Agents receive commands that preserve literal braces and semicolon-delimited logger arguments, avoid the incorrect claim that plain
dotnet buildis sequential, and avoid deleting validProjectReferenceedges. Routing now keeps lone-project shared-file introduction dormant without excluding diagnosis of an existing one-projectDirectory.Build.*timing defect.All descriptions remain below 1,024 characters. The dotnet-msbuild skill-menu description total is 10,943 characters, still 1,256 characters below the stacked base.
Validation
skill-validator check --plugin ./plugins/dotnet-msbuild: passed for 18 skills and 3 agentspython eng/eval-quality/check_eval_quality.py: passed with no errorsSKILL.mdfiles;git diff --checkpassedNo evaluation run was triggered by this update.