You signed in with another tab or window. Reload to refresh your session.You signed out in another tab or window. Reload to refresh your session.You switched accounts on another tab or window. Reload to refresh your session.Dismiss alert
This PR strengthens the dotnet-msbuild evaluation surface from 33 to 139 distinct stimuli across 19 suites. It adds realistic failures, routing boundaries, valid no-op cases, replayable golden patches, substantial ATIF references, deterministic graders, and self-contained fixtures.
It also hardens the shared eval-quality gate, native evaluator, and Vally result adaptation so activation, executor failures, permissions, session persistence, and result completeness fail closed.
Why
Earlier CI reports mixed real routing defects with fixture defects, harness limitations, statistical power, reliability failures, and trusted-workflow bootstrap behavior. This change separates those causes and fixes only defects supported by executable or multi-model evidence.
The base PR keeps the three routing fixes required by its own evidence: the msbuild custom agent, the non-MSBuild boundary for build-perf-baseline, and the incremental/FileWrites boundary for extension-points. Broader production guidance fixes are isolated in #1182.
Evaluations cover customer outcomes, routing, dormancy, and valid no-op behavior instead of prompt variations.
Dormancy scenarios cannot remove their target skill, are excluded from preference inference, and fail when the isolated target activates unexpectedly or disappears from executor output.
Fixtures, references, graders, and setup artifacts are contained, tracked, deterministic, and checked against their stated paths.
The native evaluator uses process-private storage, owner-only permissions, staged plugin content, anchored no-follow filesystem operations, scenario-scoped access, and fail-closed SDK permission handling.
GitHub token aliases and sensitive child-process environment variables are removed.
MCP execution is limited to the pinned Binlog MCP package through validator-owned NuGet configuration and private caches.
Session schema v5 preserves unknown legacy activation expectations and recovers them from current eval specs during rejudge.
Partial and terminal execution failures remain visible; rejudge cannot silently select only successful runs.
Vally plugin and target-agent activation attribution is target-specific, and missing expected stimuli remain explicit contract failures.
Validation
Current head: 59bc36e2dbeaf5bb059f956a313a11cbf24b4897, current with main and conflict-free.
Full Skill Validator suite: 789 tests passed on the coordinated evaluator hardening changes.
Cross-platform focused security and filesystem tests passed on Windows and self-contained .NET 10 Linux.
Vally adapter suites passed, including target-specific activation, missing-dormancy, retry, and native-agent result classification coverage.
Eval-quality gate passed with no errors; all 123 self-tests passed.
Skill Validator static check passed for 18 skills, 3 agents, and 1 plugin.
All review threads are resolved.
Earlier exact ordinary evidence remains successful and applicable to the MSBuild payload:
Run 35117124380 produced complete default, medium, and heavy evidence with zero executor errors, unmatched trials, or unresolved retries after its targeted reliability retry.
The latest exact-head evaluation 35246150186 proved two trusted-harness defects: PAT quota/no-auth responses were not classified as pool-candidate failures, and native-agent adaptation treated diagnostic errorCount as terminal measurement failure. The fixes exist on this branch but cannot affect its own run because evaluation-run.yml checks out trusted workflow source and the validator archive from default branch.
Prerequisite #1190 contains only those six trusted harness/docs/test files against main. After #1190 merges, update this branch from main, remove the harness files from the effective PR diff, and run one exact-head evaluation. A fresh human approval is also required after the latest head changes.
Scope
This PR contains the strengthened MSBuild evals and fixtures, shared eval-quality checks and authoring guidance, native evaluator isolation and correctness fixes, permission and token hardening, rejudge role/activation support, target-specific Vally attribution, and only the three evidence-backed routing changes listed above. The trusted harness classification subset is split into #1190 so it can land on main before this PR evaluates.
Clear NuGet fallback folders so the intentional NU1101 case cannot resolve from machine state.
Co-authored-by: Copilot App <223556219+Copilot@users.noreply.github.com>
Copilot-Session: eb11c85b-4046-4dbc-94b3-3bdbd4839a0d
whether the 19 eval suites cover useful customer decisions and routing boundaries rather than prompt variations;
whether the balance of deterministic graders, golden patches, and substantial ATIF references is reliable and easy to maintain;
whether the changed-suite quality ratchet is the right way to improve new work without blocking on unrelated legacy debt.
For repository-size context, this PR adds 290 MSBuild artifact files totaling 176,115 bytes (172 KiB) uncompressed, or about 145,041 bytes (142 KiB) as a standalone ZIP: 83 KiB of fixtures/support files, 56 KiB of ATIF JSON, and 33 KiB of golden patches. The PR description now has a row-by-row review map for every changed skill eval and the TDD/Vally-guided design principles used to create them.
❌ Evaluation did not complete successfully (the evaluate job reported failure). Check the workflow run logs, then comment /evaluate 0d63a7748d1941554011285e0ad8c5916aa5dae4 to retry this exact commit.
35 partial result file(s) were preserved for diagnosis but were not consolidated because the full matrix did not complete.
Objective completion gate: not enabled. Aggregate completion transitions are telemetry only, so this report does not claim that zero objective regressions were proven.
A result passes only when preference-eligible distinct-stimulus votes have aggregate net win of at least 20% and an exact one-sided sign-test result of p ≤ 0.05, and every explicit dormancy activation contract passes. Repeated runs measure reliability only.
Skill
Model
Verdict
Gate evidence
Overfit
Warnings
Next action
binlog-failure-analysis
claude-sonnet-4.6
⛔ Activation contract failed
n=6; 3W/2T/1L; d=4; p=0.312; net +33.3%; 2 dormancy excluded
🟡 0.26
Dormancy contract: 1 unexpected activation(s)
Narrow skill routing so the listed off-target scenarios stay dormant.
binlog-failure-analysis
gpt-5.6-luna
⛔ Activation contract failed
n=6; 2W/3T/1L; d=3; p=0.500; net +16.7%; 2 dormancy excluded
✅ 0.14
Dormancy contract: 1 unexpected activation(s)
Narrow skill routing so the listed off-target scenarios stay dormant.
binlog-generation
claude-sonnet-4.6
➖ Not proven improved
n=7; 3W/3T/1L; d=4; p=0.312; net +28.6%; 1 dormancy excluded
🔴 0.58
Activation: isolated 5/7; plugin 5/7
Inspect tied or lost stimuli; predeclare added breadth before a new experiment.
binlog-generation
gpt-5.6-luna
➖ Not proven improved
n=7; 3W/2T/2L; d=5; p=0.500; net +14.3%; 1 dormancy excluded
🟡 0.21
Activation: isolated 7/7; plugin 6/7
Inspect tied or lost stimuli and fix inconsistent skill behavior.
build-parallelism
claude-sonnet-4.6
⛔ Activation contract failed
n=6; 4W/2T/0L; d=4; p=0.063; net +66.7%; 1 dormancy excluded
Narrow skill routing so the listed off-target scenarios stay dormant.
target-authoring
gpt-5.6-luna
➖ Not proven improved
n=6; 2W/2T/2L; d=4; p=0.687; net +0.0%; 1 dormancy excluded
✅ 0.13
—
Inspect tied or lost stimuli; predeclare added breadth before a new experiment.
ℹ️ How to read this report
✅ Improved — the result passed both the statistical gate and the 20% practical net-win floor.
➖ Not proven improved — the result is valid but did not pass both gates. This is not automatically a regression.
⚠️ Invalid / underpowered — the gate withheld a quality verdict. Fix the measurement before judging the skill.
⛔ Activation contract failed — the isolated target skill activated on an explicit dormancy scenario. Dormancy preference is excluded, but this routing failure still blocks a pass.
📉 Preference loss — the LLM judge credibly preferred baseline. It is report-only, not objective completion proof.
Gate evidence — n preference-eligible distinct-stimulus votes, W/T/L stimulus votes, d discordant votes, exact one-sided p, net win, and the count of separately retained dormancy stimuli. The p value applies to one model/skill result; no matrix-wide multiple-comparison correction is applied.
Warnings — activation, timeout, retry recovery, or unresolved comparison conditions that need attention.
Do not add repeated runs to increase statistical power. Do not add stimuli after seeing a near-pass unless the new breadth is predeclared for a new experiment.
Repeated-run reliability (not used by the gate): 8 paired runs (5W/2T/1L).
Weak or warning scenarios:
Scenario
Preference gate
Net win
Δ Pref
Runs (W/T/L)
= Confirm the actual resolved target framework and package version from a binlog
Eligible
+0.0%
+0.0%
0/1/0
▲ Decline to "diagnose" a build that has not actually failed yet
Excluded (activation contract)
+100.0%
+100.0%
1/0/0
= Diagnose a warning behind a build that actually succeeded
Eligible
+0.0%
+0.0%
0/1/0
▼ Trace why a generated source file is missing at compile time
Eligible
-100.0%
-40.0%
0/0/1
Illustrative judge evidence:
Confirm the actual resolved target framework and package version from a binlog: Position-swap inconsistent (forward: B, reverse: tie). Defaulting to tie.
This is one example, not the aggregate verdict. Open Full Results for every judgment.
Repeated-run reliability (not used by the gate): 8 paired runs (2W/5T/1L).
Weak or warning scenarios:
Scenario
Preference gate
Net win
Δ Pref
Runs (W/T/L)
= Confirm the actual resolved target framework and package version from a binlog
Eligible
+0.0%
+0.0%
0/1/0
= Decline to "diagnose" a build that has not actually failed yet
Excluded (activation contract)
+0.0%
+0.0%
0/1/0
= Diagnose build failures from binlog only (no source files)
Eligible
+0.0%
+0.0%
0/1/0
= Fall back to command-line log replay when the usual binlog tool is unavailable
Eligible
+0.0%
+0.0%
0/1/0
= Stay dormant for a non-MSBuild build failure log
Excluded (activation contract)
+0.0%
+0.0%
0/1/0
▼ Trace why a generated source file is missing at compile time
Eligible
-100.0%
-40.0%
0/0/1
Illustrative judge evidence:
Confirm the actual resolved target framework and package version from a binlog: Position-swap inconsistent (forward: B, reverse: tie). Defaulting to tie.
This is one example, not the aggregate verdict. Open Full Results for every judgment.
Repeated-run reliability (not used by the gate): 7 paired runs (5W/1T/1L).
Weak or warning scenarios:
Scenario
Preference gate
Net win
Δ Pref
Runs (W/T/L)
▲ Configure deterministic, cache-safe CI builds
Eligible
+100.0%
+100.0%
1/0/0
▲ Decline a non-MSBuild build performance request
Excluded (activation contract)
+100.0%
+40.0%
1/0/0
▼ Leave an already-optimized build unchanged
Eligible
-100.0%
-40.0%
0/0/1
= Route a restore-bound cold build away from architecture changes
Eligible
+0.0%
+0.0%
0/1/0
Illustrative judge evidence:
Leave an already-optimized build unchanged: Both fail the central task by changing a repository with no applicable checklist anti-pattern. A at least locates and reviews the requested App/Core project, whereas B works on a different WebApi/BusinessLogic repository and makes unrelated reference edits; A is therefore slig...
This is one example, not the aggregate verdict. Open Full Results for every judgment.
Repeated-run reliability (not used by the gate): 7 paired runs (3W/2T/2L).
Weak or warning scenarios:
Scenario
Preference gate
Net win
Δ Pref
Runs (W/T/L)
▲ Configure deterministic, cache-safe CI builds
Eligible
+100.0%
+100.0%
1/0/0
▼ Decline a non-MSBuild build performance request
Excluded (activation contract)
-100.0%
-40.0%
0/0/1
▼ Leave an already-optimized build unchanged
Eligible
-100.0%
-40.0%
0/0/1
= Route a broken no-op rebuild away from generic optimization
Eligible
+0.0%
+0.0%
0/1/0
= Route a restore-bound cold build away from architecture changes
Eligible
+0.0%
+0.0%
0/1/0
Illustrative judge evidence:
Leave an already-optimized build unchanged: The ideal outcome was to make no changes, as no baseline anti-pattern was present. Both responses failed this and invented a change. However, B committed the specific anti-pattern the rubric explicitly warns against (adding UseArtifactsOutput to a two-project solution), and it...
This is one example, not the aggregate verdict. Open Full Results for every judgment.
Repeated-run reliability (not used by the gate): 7 paired runs (0W/4T/3L).
Weak or warning scenarios:
Scenario
Preference gate
Net win
Δ Pref
Runs (W/T/L)
▼ Audit the mixed solution and separate safe projects from unsafe ones
Eligible
-100.0%
-40.0%
0/0/1
= Avoid a false clash report when projects share only the top-level artifacts root
Eligible
+0.0%
+0.0%
0/1/0
▼ Decline output-clash remediation for separate projects using the default SDK layout
Excluded (activation contract)
-100.0%
-40.0%
0/0/1
= Diagnose multi-targeting outputs that collapse into one path
Eligible
+0.0%
+0.0%
0/1/0
= Diagnose redundant project reference metadata that forks a same-path build
Eligible
+0.0%
+0.0%
0/1/0
= Diagnose shared output and intermediate path collision
Eligible
+0.0%
+0.0%
0/1/0
▼ Fix all clash mechanisms in the mixed solution
Eligible
-100.0%
-40.0%
0/0/1
Illustrative judge evidence:
Audit the mixed solution and separate safe projects from unsafe ones: Both responses correctly identify the core clashes (LibraryA/B sharing, MultiTargetLib TFM overwrite) and the ToolLib redundant-build mechanism. However, A aligns better with the rubric's intended classification: it flags ConsumerApp as unsafe (the source of the redundant meta...
This is one example, not the aggregate verdict. Open Full Results for every judgment.
Repeated-run reliability (not used by the gate): 8 paired runs (4W/2T/2L).
Weak or warning scenarios:
Scenario
Preference gate
Net win
Δ Pref
Runs (W/T/L)
▼ Decline manual wiring for Roslyn source generators
Excluded (activation contract)
-100.0%
-40.0%
0/0/1
= Diagnose hardcoded obj path for generated source
Eligible
+0.0%
+0.0%
0/1/0
▲ Diagnose missing clean tracking for generated source
Eligible
+100.0%
+40.0%
1/0/0
▼ Diagnose missing output registration for generated non-code file
Eligible
-100.0%
-40.0%
0/0/1
= Diagnose wrong hook for generated source files
Eligible
+0.0%
+0.0%
0/1/0
Illustrative judge evidence:
Diagnose hardcoded obj path for generated source: The outputs are substantively equivalent, accurate, and directly answer the requested root cause and correct path pattern. B is marginally more concise, while A is equally clear; neither has a meaningful quality advantage.
This is one example, not the aggregate verdict. Open Full Results for every judgment.
Repeated-run reliability (not used by the gate): 7 paired runs (1W/3T/3L).
Weak or warning scenarios:
Scenario
Preference gate
Net win
Δ Pref
Runs (W/T/L)
▼ Decline a target ordering problem with no item-group defect
Excluded (activation contract)
-100.0%
-40.0%
0/0/1
= Diagnose an ineffective Compile Remove that does not match the glob
Eligible
+0.0%
+0.0%
0/1/0
= Diagnose real and claimed item problems in a code generation pipeline
Eligible
+0.0%
+0.0%
0/1/0
▼ Fix item management anti-patterns
Eligible
-100.0%
-40.0%
0/0/1
= Leave already-correct item management unchanged
Eligible
+0.0%
+0.0%
0/1/0
▼ Leave correct single-list batching unchanged
Eligible
-100.0%
-40.0%
0/0/1
Illustrative judge evidence:
Diagnose an ineffective Compile Remove that does not match the glob: The final answers are substantively equivalent, concise, correct, and fully address both the cause and correction.
This is one example, not the aggregate verdict. Open Full Results for every judgment.
Repeated-run reliability (not used by the gate): 8 paired runs (2W/6T/0L).
Weak or warning scenarios:
Scenario
Preference gate
Net win
Δ Pref
Runs (W/T/L)
= Add a module to an F# project
Eligible
+0.0%
+0.0%
0/1/0
= Add a signature file to define public API
Eligible
+0.0%
+0.0%
0/1/0
= Distinguish a style backslash from a real cross-platform backslash bug
Eligible
+0.0%
+0.0%
0/1/0
= Fix broken file order causing FS0039
Eligible
+0.0%
+0.0%
0/1/0
= Judge an unguarded import inside a NuGet package build folder
Eligible
+0.0%
+0.0%
0/1/0
= Leave a clean project without inventing anti-patterns
Eligible
+0.0%
+0.0%
0/1/0
▲ Non-activation: migrate a legacy project to SDK style
Excluded (activation contract)
+100.0%
+40.0%
1/0/0
Illustrative judge evidence:
Add a module to an F# project: The implementations and demonstrated behavior are substantively identical: both meet the requested validation, project ordering, integration, and successful build/run requirements.
This is one example, not the aggregate verdict. Open Full Results for every judgment.
Repeated-run reliability (not used by the gate): 8 paired runs (0W/4T/4L).
Weak or warning scenarios:
Scenario
Preference gate
Net win
Δ Pref
Runs (W/T/L)
= Add a module to an F# project
Eligible
+0.0%
+0.0%
0/1/0
= Add a signature file to define public API
Eligible
+0.0%
+0.0%
0/1/0
= Distinguish a style backslash from a real cross-platform backslash bug
Eligible
+0.0%
+0.0%
0/1/0
▼ Fix broken file order causing FS0039
Eligible
-100.0%
-40.0%
0/0/1
= Judge an unguarded import inside a NuGet package build folder
Eligible
+0.0%
+0.0%
0/1/0
▼ Leave a clean project without inventing anti-patterns
Eligible
-100.0%
-40.0%
0/0/1
▼ Non-activation: migrate a legacy project to SDK style
Excluded (activation contract)
-100.0%
-40.0%
0/0/1
▼ Review MSBuild files for anti-patterns and style issues
Eligible
-100.0%
-40.0%
0/0/1
Illustrative judge evidence:
Add a module to an F# project: Both responses produced functionally identical results: an equivalent Validation.fs, updated Program.fs, correct compile ordering (evidenced by successful builds requiring Validation before Program), and successful dotnet run. B briefly hit a NuGet restore error but recovered ...
This is one example, not the aggregate verdict. Open Full Results for every judgment.
Repeated-run reliability (not used by the gate): 8 paired runs (4W/1T/3L).
Weak or warning scenarios:
Scenario
Preference gate
Net win
Δ Pref
Runs (W/T/L)
= Confirm the MSBuild Server is actually improving build times before declaring success
Eligible
+0.0%
+0.0%
0/1/0
▼ Decline MSBuild Server for a single one-off release build
Excluded (activation contract)
-100.0%
-40.0%
0/0/1
▼ Explain a background MSBuild Server process consuming memory
Eligible
-100.0%
-40.0%
0/0/1
▼ Recommend MSBuild Server for slow CLI incremental builds
Eligible
-100.0%
-40.0%
0/0/1
▲ Stay dormant for a purely IDE-side build slowdown
Excluded (activation contract)
+100.0%
+100.0%
1/0/0
Illustrative judge evidence:
Confirm the MSBuild Server is actually improving build times before declaring success: Position-swap inconsistent (forward: A, reverse: B). Defaulting to tie.
This is one example, not the aggregate verdict. Open Full Results for every judgment.
Repeated-run reliability (not used by the gate): 8 paired runs (5W/2T/1L).
Weak or warning scenarios:
Scenario
Preference gate
Net win
Δ Pref
Runs (W/T/L)
▼ Configure MSBuild Server persistently across new Windows terminals
Eligible
-100.0%
-40.0%
0/0/1
= Decline MSBuild Server for a single one-off release build
Excluded (activation contract)
+0.0%
+0.0%
0/1/0
= Diagnose stale build output after enabling MSBuild Server
Eligible
+0.0%
+0.0%
0/1/0
Illustrative judge evidence:
Configure MSBuild Server persistently across new Windows terminals: Both responses correctly recommend a persistent user-level Windows environment variable and instruct closing/reopening terminals. The decisive difference is the variable name: Response A uses DOTNET_CLI_USE_MSBUILD_SERVER, which matches the rubric's designated correct variable...
This is one example, not the aggregate verdict. Open Full Results for every judgment.
Repeated-run reliability (not used by the gate): 8 paired runs (6W/0T/2L).
Weak or warning scenarios:
Scenario
Preference gate
Net win
Δ Pref
Runs (W/T/L)
▼ Rank Copy ahead of Csc when task self-time is higher
Eligible
-100.0%
-100.0%
0/0/1
▼ Request diagnostic artifact before analysis
Excluded (activation contract)
-100.0%
-40.0%
0/0/1
Illustrative judge evidence:
Rank Copy ahead of Csc when task self-time is higher: A directly and correctly answers the question using the report's task-level evidence, explains why the target-level leader is not the optimization target, and gives a relevant copy optimization direction. B has the general diagnostic principle right but fails to use the suppli...
This is one example, not the aggregate verdict. Open Full Results for every judgment.
Repeated-run reliability (not used by the gate): 7 paired runs (3W/3T/1L).
Weak or warning scenarios:
Scenario
Preference gate
Net win
Δ Pref
Runs (W/T/L)
= Decline an incremental-build tuning request
Excluded (activation contract)
+0.0%
+0.0%
0/1/0
▲ Diagnose a target hooked to Build that misses direct compile
Eligible
+100.0%
+40.0%
1/0/0
= Diagnose broken SDK target chain across files
Eligible
+0.0%
+0.0%
0/1/0
= Diagnose custom target reliability issues
Eligible
+0.0%
+0.0%
0/1/0
▼ Fix a query target that uses Outputs instead of Returns
Eligible
-100.0%
-40.0%
0/0/1
Illustrative judge evidence:
Diagnose broken SDK target chain across files: Position-swap inconsistent (forward: tie, reverse: A). Defaulting to tie.
This is one example, not the aggregate verdict. Open Full Results for every judgment.
➖ Not proven improved — binlog-generation (claude-sonnet-4.6)
Why: Net win +28.6% (3W/3T/1L over 7 preference-eligible stimulus vote(s), sign test p=0.312), mean preference +20.0% across 8 paired run(s), 1 dormancy stimulus/stimuli excluded from preference — not credible — 3 of 7 preference-eligible stimulus vote(s) tied, leaving only 4 discordant preference vote(s). The sign test conditions on non-tie stimulus votes and cannot reach 0.05 below 5, so no record could have passed here — this is not a measured null. Either the skill is inert on these scenarios (make them discriminate) or the eval needs more distinct stimuli to clear the ties
Next action: Inspect tied or lost stimuli; predeclare added breadth before a new experiment.
Repeated-run reliability (not used by the gate): 8 paired runs (3W/3T/2L).
Weak or warning scenarios:
Scenario
Preference gate
Net win
Δ Pref
Runs (W/T/L)
= Build with /bl in PowerShell
Eligible
+0.0%
+0.0%
0/1/0
= Choose a predictable, non-colliding binlog name for a CI upload step
Eligible
+0.0%
+0.0%
0/1/0
▼ Decline a binlog request for a non-MSBuild Java build
Excluded (activation contract)
-100.0%
-40.0%
0/0/1
= Fix a CI build script that reuses the same binlog on every run
Eligible
+0.0%
+0.0%
0/1/0
▼ Preserve binlog history while cleaning stale build output
Eligible
-100.0%
-40.0%
0/0/1
Illustrative judge evidence:
Build with /bl in PowerShell: Position-swap inconsistent (forward: B, reverse: A). Defaulting to tie.
This is one example, not the aggregate verdict. Open Full Results for every judgment.
➖ Not proven improved — binlog-generation (gpt-5.6-luna)
Why: Net win +14.3% (3W/2T/2L over 7 preference-eligible stimulus vote(s), sign test p=0.500), mean preference +27.5% across 8 paired run(s), 1 dormancy stimulus/stimuli excluded from preference — not credible (sign test p=0.500 > 0.05)
Next action: Inspect tied or lost stimuli and fix inconsistent skill behavior.
Repeated-run reliability (not used by the gate): 8 paired runs (3W/3T/2L).
Weak or warning scenarios:
Scenario
Preference gate
Net win
Δ Pref
Runs (W/T/L)
▼ Build with /bl in PowerShell
Eligible
-100.0%
-40.0%
0/0/1
= Choose a predictable, non-colliding binlog name for a CI upload step
Eligible
+0.0%
+0.0%
0/1/0
= Decline a binlog request for a non-MSBuild Java build
Excluded (activation contract)
+0.0%
+0.0%
0/1/0
= Fix a CI build script that reuses the same binlog on every run
Eligible
+0.0%
+0.0%
0/1/0
▼ Recognize that a failed build produced no binlog at all
Eligible
-100.0%
-40.0%
0/0/1
Illustrative judge evidence:
Build with /bl in PowerShell: Both successfully built the project via PowerShell and produced a valid binlog with 0 errors. A's approach used an explicit quoted filename, which is unambiguous and clean. B relied on the skill's {{}} template syntax which, while it produced a valid unique filename, uses the ...
This is one example, not the aggregate verdict. Open Full Results for every judgment.
➖ Not proven improved — build-parallelism (gpt-5.6-luna)
Why: Net win +0.0% (3W/0T/3L over 6 preference-eligible stimulus vote(s), sign test p=0.656), mean preference -5.7% across 7 paired run(s), 1 dormancy stimulus/stimuli excluded from preference — no improvement
Next action: Inspect tied or lost stimuli and fix inconsistent skill behavior.
Repeated-run reliability (not used by the gate): 7 paired runs (3W/0T/4L).
Weak or warning scenarios:
Scenario
Preference gate
Net win
Δ Pref
Runs (W/T/L)
▼ Enable parallel nodes for a wide project graph
Eligible
-100.0%
-40.0%
0/0/1
▼ Non-activation: make one custom target incremental
Excluded (activation contract)
-100.0%
-40.0%
0/0/1
▼ Preserve valid dependencies and optimize the slow critical-path project
Eligible
-100.0%
-40.0%
0/0/1
▼ Reduce CI build scope with a solution filter
Eligible
-100.0%
-40.0%
0/0/1
Illustrative judge evidence:
Enable parallel nodes for a wide project graph: Both responses are correct and nearly identical in substance, correctly diagnosing the single-node default and recommending -m. Neither suggested confirming via binlog. A slightly edges out by also offering the explicit -m:4 option, better matching the -m:N criterion.
This is one example, not the aggregate verdict. Open Full Results for every judgment.
➖ Not proven improved — build-perf-diagnostics (claude-sonnet-4.6)
Why: Net win -14.3% (1W/4T/2L over 7 preference-eligible stimulus vote(s), sign test p=0.500), mean preference -17.5% across 8 paired run(s), 1 dormancy stimulus/stimuli excluded from preference — no improvement
Next action: Inspect tied or lost stimuli; predeclare added breadth before a new experiment.
Repeated-run reliability (not used by the gate): 8 paired runs (1W/4T/3L).
Weak or warning scenarios:
Scenario
Preference gate
Net win
Δ Pref
Runs (W/T/L)
▼ Decline a runtime latency request that is not a build performance issue
Excluded (activation contract)
-100.0%
-100.0%
0/0/1
▼ Diagnose NuGet restore running redundantly across CI stages
Eligible
-100.0%
-40.0%
0/0/1
= Diagnose a Copy task dominating build time
Eligible
+0.0%
+0.0%
0/1/0
▼ Diagnose a single custom target dominating one project's build
Eligible
-100.0%
-40.0%
0/0/1
= Diagnose evaluation overhead before any target runs
Eligible
+0.0%
+0.0%
0/1/0
= Diagnose per-project overhead across many small projects
Eligible
+0.0%
+0.0%
0/1/0
= Diagnose slow build for a small project
Eligible
+0.0%
+0.0%
0/1/0
Illustrative judge evidence:
Diagnose NuGet restore running redundantly across CI stages: Both reach the central diagnosis and recommended pipeline change, but A gives a more accurate accounting of avoidable work and avoids B's misleading claims about static graph evaluation and post-change stage timing.
This is one example, not the aggregate verdict. Open Full Results for every judgment.
Routine passing details for 1 result are in Full Results.
Details for 8 results were omitted to keep this comment under GitHub's limit. Open the workflow summary or Full Results for the complete breakdown.
To investigate non-passing or warning results, paste this to your AI coding agent:
For PR 1118 in dotnet/skills, download eval artifacts with gh run download 33898074788 --repo dotnet/skills --pattern "vally-results-*" --dir ./eval-results, then fetch https://raw.githubusercontent.com/dotnet/skills/0d63a7748d1941554011285e0ad8c5916aa5dae4/eng/vally-adapter/InvestigatingResults.md and follow it. Classify each result as measurement-invalid, underpowered, not-proven, preference-loss, or passing-with-warning. Use stateReason, result accounting, weak scenarios, and judge evidence to give the cause and exact next fix.
❌ Evaluation did not complete successfully (the evaluate job reported failure). Check the workflow run logs, then comment /evaluate 34de668c6602fbf6ec12ec19da961e9f11e6e2e1 to retry this exact commit.
36 partial result file(s) were preserved for diagnosis but were not consolidated because the full matrix did not complete.
Objective completion gate: not enabled. Aggregate completion transitions are telemetry only, so this report does not claim that zero objective regressions were proven.
A result passes only when preference-eligible distinct-stimulus votes have aggregate net win of at least 20% and an exact one-sided sign-test result of p ≤ 0.05, and every explicit dormancy activation contract passes. Repeated runs measure reliability only.
Skill
Model
Verdict
Gate evidence
Overfit
Warnings
Next action
binlog-failure-analysis
claude-sonnet-4.6
➖ Not proven improved
n=7; 5W/1T/1L; d=6; p=0.109; net +57.1%; 1 dormancy excluded
🟡 0.24
—
Inspect tied or lost stimuli and fix inconsistent skill behavior.
binlog-failure-analysis
gpt-5.6-luna
➖ Not proven improved
n=7; 4W/1T/2L; d=6; p=0.344; net +28.6%; 1 dormancy excluded
✅ 0.10
—
Inspect tied or lost stimuli and fix inconsistent skill behavior.
binlog-generation
claude-sonnet-4.6
✅ Improved
n=7; 5W/2T/0L; d=5; p=0.031; net +71.4%; 1 dormancy excluded
🟡 0.49
—
Review overfit evidence.
binlog-generation
gpt-5.6-luna
➖ Not proven improved
n=7; 4W/3T/0L; d=4; p=0.063; net +57.1%; 1 dormancy excluded
🟡 0.28
—
Inspect tied or lost stimuli; predeclare added breadth before a new experiment.
build-parallelism
claude-sonnet-4.6
⛔ Activation contract failed
n=6; 3W/3T/0L; d=3; p=0.125; net +50.0%; 1 dormancy excluded
Narrow skill routing so the listed off-target scenarios stay dormant.
property-patterns
gpt-5.6-luna
➖ Not proven improved
n=6; 3W/1T/2L; d=5; p=0.500; net +16.7%; 1 dormancy excluded
✅ 0.06
—
Inspect tied or lost stimuli and fix inconsistent skill behavior.
resolve-project-references
claude-sonnet-4.6
✅ Improved
n=6; 6W/0T/0L; d=6; p=0.016; net +100.0%; 2 dormancy excluded
🟡 0.36
Activation: isolated 6/6; plugin 5/6
Fix activation gaps; Review overfit evidence.
resolve-project-references
gpt-5.6-luna
➖ Not proven improved
n=6; 2W/3T/1L; d=3; p=0.500; net +16.7%; 2 dormancy excluded
✅ 0.15
Activation: isolated 4/6; plugin 5/6
Inspect tied or lost stimuli; predeclare added breadth before a new experiment.
target-authoring
claude-sonnet-4.6
➖ Not proven improved
n=6; 2W/1T/3L; d=5; p=0.500; net -16.7%; 1 dormancy excluded
🟡 0.31
Activation: isolated 6/6; plugin 5/6
Inspect tied or lost stimuli and fix inconsistent skill behavior.
target-authoring
gpt-5.6-luna
➖ Not proven improved
n=6; 2W/3T/1L; d=3; p=0.500; net +16.7%; 1 dormancy excluded
✅ 0.10
—
Inspect tied or lost stimuli; predeclare added breadth before a new experiment.
ℹ️ How to read this report
✅ Improved — the result passed both the statistical gate and the 20% practical net-win floor.
➖ Not proven improved — the result is valid but did not pass both gates. This is not automatically a regression.
⚠️ Invalid / underpowered — the gate withheld a quality verdict. Fix the measurement before judging the skill.
⛔ Activation contract failed — the isolated target skill activated on an explicit dormancy scenario. Dormancy preference is excluded, but this routing failure still blocks a pass.
📉 Preference loss — the LLM judge credibly preferred baseline. It is report-only, not objective completion proof.
Gate evidence — n preference-eligible distinct-stimulus votes, W/T/L stimulus votes, d discordant votes, exact one-sided p, net win, and the count of separately retained dormancy stimuli. The p value applies to one model/skill result; no matrix-wide multiple-comparison correction is applied.
Warnings — activation, timeout, retry recovery, or unresolved comparison conditions that need attention.
Do not add repeated runs to increase statistical power. Do not add stimuli after seeing a near-pass unless the new breadth is predeclared for a new experiment.
Repeated-run reliability (not used by the gate): 7 paired runs (4W/0T/3L).
Weak or warning scenarios:
Scenario
Preference gate
Net win
Δ Pref
Runs (W/T/L)
▼ Configure deterministic, cache-safe CI builds
Eligible
-100.0%
-100.0%
0/0/1
▼ Decline a non-MSBuild build performance request
Excluded (activation contract)
-100.0%
-40.0%
0/0/1
▼ Leave an already-optimized build unchanged
Eligible
-100.0%
-40.0%
0/0/1
Illustrative judge evidence:
Configure deterministic, cache-safe CI builds: A makes the essential requested, CI-guarded ContinuousIntegrationBuild change in the relevant Directory.Build.props and leaves valid configuration. Its unnecessary Deterministic addition and inaccurate default claim are flaws, but B misses the central required property and edi...
This is one example, not the aggregate verdict. Open Full Results for every judgment.
Repeated-run reliability (not used by the gate): 7 paired runs (3W/3T/1L).
Weak or warning scenarios:
Scenario
Preference gate
Net win
Δ Pref
Runs (W/T/L)
= Decline adding shared build files to a lone project
Excluded (activation contract)
+0.0%
+0.0%
0/1/0
▼ Diagnose a package downgrade chain and reorganize version management
Eligible
-100.0%
-40.0%
0/0/1
= Diagnose an inner shared-props file that overwrites its own override
Eligible
+0.0%
+0.0%
0/1/0
= Preserve an intentional project-specific exception while centralizing
Eligible
+0.0%
+0.0%
0/1/0
Illustrative judge evidence:
Diagnose a package downgrade chain and reorganize version management: Both responses reach the same correct conclusion and implementation. However, Response A verified its fix more thoroughly by running both dotnet restore AND dotnet test successfully, and normalized the project XML. Response B ran a build that surfaced unrelated CS1591 errors a...
This is one example, not the aggregate verdict. Open Full Results for every judgment.
Repeated-run reliability (not used by the gate): 8 paired runs (3W/1T/4L).
Weak or warning scenarios:
Scenario
Preference gate
Net win
Δ Pref
Runs (W/T/L)
= Decline manual wiring for Roslyn source generators
Excluded (activation contract)
+0.0%
+0.0%
0/1/0
▼ Diagnose hardcoded obj path for generated source
Eligible
-100.0%
-100.0%
0/0/1
▲ Diagnose missing clean tracking for generated source
Eligible
+100.0%
+40.0%
1/0/0
▼ Diagnose missing output registration for generated non-code file
Eligible
-100.0%
-40.0%
0/0/1
▼ Diagnose project-level glob for generated source
Eligible
-100.0%
-40.0%
0/0/1
▼ Diagnose wrong hook for generated source files
Eligible
-100.0%
-40.0%
0/0/1
Illustrative judge evidence:
Diagnose hardcoded obj path for generated source: A is precise, consistent with the stated behavior, and supplies the correct reusable pattern. B's recommendation is broadly right, but its root-cause narrative is internally inconsistent and reverses the key path relationship.
This is one example, not the aggregate verdict. Open Full Results for every judgment.
Repeated-run reliability (not used by the gate): 7 paired runs (5W/2T/0L).
Weak or warning scenarios:
Scenario
Preference gate
Net win
Δ Pref
Runs (W/T/L)
▲ Decline a target ordering problem with no item-group defect
Excluded (activation contract)
+100.0%
+40.0%
1/0/0
▲ Diagnose an ineffective Compile Remove that does not match the glob
Eligible
+100.0%
+40.0%
1/0/0
▲ Diagnose real and claimed item problems in a code generation pipeline
Eligible
+100.0%
+40.0%
1/0/0
= Fix item management anti-patterns
Eligible
+0.0%
+0.0%
0/1/0
= Leave already-correct item management unchanged
Eligible
+0.0%
+0.0%
0/1/0
▲ Leave correct single-list batching unchanged
Eligible
+100.0%
+40.0%
1/0/0
Illustrative judge evidence:
Fix item management anti-patterns: The two final project edits implement the same four substantive corrections with no material correctness difference evident in the recorded outputs.
This is one example, not the aggregate verdict. Open Full Results for every judgment.
Repeated-run reliability (not used by the gate): 7 paired runs (3W/2T/2L).
Weak or warning scenarios:
Scenario
Preference gate
Net win
Δ Pref
Runs (W/T/L)
▼ Choose the right OS detection for a cross-platform property
Eligible
-100.0%
-40.0%
0/0/1
▼ Decline a props-versus-targets placement question
Excluded (activation contract)
-100.0%
-40.0%
0/0/1
▲ Diagnose a TargetFramework condition that never applies in props
Eligible
+100.0%
+40.0%
1/0/0
= Diagnose multi-level property hierarchy bugs
Eligible
+0.0%
+0.0%
0/1/0
= Fix shared property configuration
Eligible
+0.0%
+0.0%
0/1/0
Illustrative judge evidence:
Choose the right OS detection for a cross-platform property: Both responses correctly identify the defect and propose explicit Windows/macOS/Linux detection. A is slightly clearer and more directly actionable by naming the MSBuild intrinsic and its required platform names.
This is one example, not the aggregate verdict. Open Full Results for every judgment.
➖ Not proven improved — binlog-failure-analysis (claude-sonnet-4.6)
Why: Net win +57.1% (5W/1T/1L over 7 preference-eligible stimulus vote(s), sign test p=0.109), mean preference +42.5% across 8 paired run(s), 1 dormancy stimulus/stimuli excluded from preference — not credible (sign test p=0.109 > 0.05)
Next action: Inspect tied or lost stimuli and fix inconsistent skill behavior.
Repeated-run reliability (not used by the gate): 8 paired runs (5W/2T/1L).
Weak or warning scenarios:
Scenario
Preference gate
Net win
Δ Pref
Runs (W/T/L)
▼ Determine whether a quiet second build actually failed
Eligible
-100.0%
-40.0%
0/0/1
= Diagnose a warning behind a build that actually succeeded
Eligible
+0.0%
+0.0%
0/1/0
= Stay dormant for a non-MSBuild build failure log
Excluded (activation contract)
+0.0%
+0.0%
0/1/0
Illustrative judge evidence:
Determine whether a quiet second build actually failed: Both answers are correct, concise, and address the user's concern. A is marginally stronger because it presents the key skip/up-to-date evidence directly and adds practical ways to force recompilation.
This is one example, not the aggregate verdict. Open Full Results for every judgment.
➖ Not proven improved — binlog-failure-analysis (gpt-5.6-luna)
Why: Net win +28.6% (4W/1T/2L over 7 preference-eligible stimulus vote(s), sign test p=0.344), mean preference +2.5% across 8 paired run(s), 1 dormancy stimulus/stimuli excluded from preference — not credible (sign test p=0.344 > 0.05)
Next action: Inspect tied or lost stimuli and fix inconsistent skill behavior.
Repeated-run reliability (not used by the gate): 8 paired runs (4W/2T/2L).
Weak or warning scenarios:
Scenario
Preference gate
Net win
Δ Pref
Runs (W/T/L)
▼ Assess a requested binlog investigation when the current build is healthy
Eligible
-100.0%
-100.0%
0/0/1
= Confirm the actual resolved target framework and package version from a binlog
Eligible
+0.0%
+0.0%
0/1/0
▼ Fall back to command-line log replay when the usual binlog tool is unavailable
Eligible
-100.0%
-40.0%
0/0/1
= Stay dormant for a non-MSBuild build failure log
Excluded (activation contract)
+0.0%
+0.0%
0/1/0
Illustrative judge evidence:
Assess a requested binlog investigation when the current build is healthy: The staged project builds successfully; there is no genuine failure to diagnose. Response A captured a binlog and correctly concluded the build succeeds with 0 warnings/errors, noting only benign SDK-environment diagnostics. Response B, despite using the dedicated skill, manuf...
This is one example, not the aggregate verdict. Open Full Results for every judgment.
➖ Not proven improved — binlog-generation (gpt-5.6-luna)
Why: Net win +57.1% (4W/3T/0L over 7 preference-eligible stimulus vote(s), sign test p=0.063), mean preference +37.5% across 8 paired run(s), 1 dormancy stimulus/stimuli excluded from preference — not credible — 3 of 7 preference-eligible stimulus vote(s) tied, leaving only 4 discordant preference vote(s). The sign test conditions on non-tie stimulus votes and cannot reach 0.05 below 5, so no record could have passed here — this is not a measured null. Either the skill is inert on these scenarios (make them discriminate) or the eval needs more distinct stimuli to clear the ties
Next action: Inspect tied or lost stimuli; predeclare added breadth before a new experiment.
Repeated-run reliability (not used by the gate): 8 paired runs (4W/3T/1L).
Weak or warning scenarios:
Scenario
Preference gate
Net win
Δ Pref
Runs (W/T/L)
= Build with /bl in PowerShell
Eligible
+0.0%
+0.0%
0/1/0
= Choose a predictable, non-colliding binlog name for a CI upload step
Eligible
+0.0%
+0.0%
0/1/0
▼ Decline a binlog request for a non-MSBuild Java build
Excluded (activation contract)
-100.0%
-40.0%
0/0/1
= Recognize that a failed build produced no binlog at all
Eligible
+0.0%
+0.0%
0/1/0
Illustrative judge evidence:
Build with /bl in PowerShell: Position-swap inconsistent (forward: tie, reverse: A). Defaulting to tie.
This is one example, not the aggregate verdict. Open Full Results for every judgment.
➖ Not proven improved — build-parallelism (gpt-5.6-luna)
Why: Net win -50.0% (0W/3T/3L over 6 preference-eligible stimulus vote(s), sign test p=0.125), mean preference -25.7% across 7 paired run(s), 1 dormancy stimulus/stimuli excluded from preference — no improvement
Next action: Inspect tied or lost stimuli; predeclare added breadth before a new experiment.
Repeated-run reliability (not used by the gate): 7 paired runs (0W/4T/3L).
Weak or warning scenarios:
Scenario
Preference gate
Net win
Δ Pref
Runs (W/T/L)
▼ Analyze build parallelism bottlenecks
Eligible
-100.0%
-40.0%
0/0/1
= Decline graph build for runtime-discovered projects
Eligible
+0.0%
+0.0%
0/1/0
= Enable BuildInParallel on a custom MSBuild task
Eligible
+0.0%
+0.0%
0/1/0
= Enable parallel nodes for a wide project graph
Eligible
+0.0%
+0.0%
0/1/0
= Non-activation: make one custom target incremental
Excluded (activation contract)
+0.0%
+0.0%
0/1/0
▼ Preserve valid dependencies and optimize the slow critical-path project
Eligible
-100.0%
-100.0%
0/0/1
▼ Reduce CI build scope with a solution filter
Eligible
-100.0%
-40.0%
0/0/1
Illustrative judge evidence:
Analyze build parallelism bottlenecks: Both responses reach the same correct conclusions on all substantive rubric points (serial chain, minimum build time, redundant Tests->Api reference, graph-hygiene-only). They are essentially tied. A edges ahead slightly because its evidence (per-project performance summary wi...
This is one example, not the aggregate verdict. Open Full Results for every judgment.
➖ Not proven improved — build-perf-baseline (gpt-5.6-luna)
Why: Net win +33.3% (4W/0T/2L over 6 preference-eligible stimulus vote(s), sign test p=0.344), mean preference +25.7% across 7 paired run(s), 1 dormancy stimulus/stimuli excluded from preference — not credible (sign test p=0.344 > 0.05)
Next action: Inspect tied or lost stimuli and fix inconsistent skill behavior.
Repeated-run reliability (not used by the gate): 7 paired runs (5W/0T/2L).
Weak or warning scenarios:
Scenario
Preference gate
Net win
Δ Pref
Runs (W/T/L)
▲ Configure deterministic, cache-safe CI builds
Eligible
+100.0%
+100.0%
1/0/0
▼ Leave an already-optimized build unchanged
Eligible
-100.0%
-40.0%
0/0/1
▼ Route a broken no-op rebuild away from generic optimization
Eligible
-100.0%
-40.0%
0/0/1
Illustrative judge evidence:
Leave an already-optimized build unchanged: The ideal answer was to conclude no changes were needed since the project is already optimized (minimal refs, CI-conditioned docs). Both responses failed this by inventing a change. However, B committed exactly the anti-pattern the rubric explicitly calls out (adding UseArtifa...
This is one example, not the aggregate verdict. Open Full Results for every judgment.
➖ Not proven improved — build-perf-diagnostics (claude-sonnet-4.6)
Why: Net win +0.0% (3W/1T/3L over 7 preference-eligible stimulus vote(s), sign test p=0.656), mean preference -20.0% across 8 paired run(s), 1 dormancy stimulus/stimuli excluded from preference — no improvement
Next action: Inspect tied or lost stimuli and fix inconsistent skill behavior.
Repeated-run reliability (not used by the gate): 8 paired runs (3W/1T/4L).
Weak or warning scenarios:
Scenario
Preference gate
Net win
Δ Pref
Runs (W/T/L)
▼ Decline a runtime latency request that is not a build performance issue
Excluded (activation contract)
-100.0%
-100.0%
0/0/1
▼ Diagnose NuGet restore running redundantly across CI stages
Eligible
-100.0%
-100.0%
0/0/1
▼ Diagnose a single custom target dominating one project's build
Eligible
-100.0%
-40.0%
0/0/1
▼ Diagnose evaluation overhead before any target runs
Eligible
-100.0%
-40.0%
0/0/1
= Diagnose slow build for a small project
Eligible
+0.0%
+0.0%
0/1/0
Illustrative judge evidence:
Diagnose NuGet restore running redundantly across CI stages: A correctly diagnoses the repeated restore cost and supplies all requested concrete CI and MSBuild changes; B fails to locate the available file and produces no substantive answer.
This is one example, not the aggregate verdict. Open Full Results for every judgment.
➖ Not proven improved — build-perf-diagnostics (gpt-5.6-luna)
Why: Net win +28.6% (3W/3T/1L over 7 preference-eligible stimulus vote(s), sign test p=0.312), mean preference +10.0% across 8 paired run(s), 1 dormancy stimulus/stimuli excluded from preference — not credible — 3 of 7 preference-eligible stimulus vote(s) tied, leaving only 4 discordant preference vote(s). The sign test conditions on non-tie stimulus votes and cannot reach 0.05 below 5, so no record could have passed here — this is not a measured null. Either the skill is inert on these scenarios (make them discriminate) or the eval needs more distinct stimuli to clear the ties
Next action: Inspect tied or lost stimuli; predeclare added breadth before a new experiment.
Repeated-run reliability (not used by the gate): 8 paired runs (3W/4T/1L).
Weak or warning scenarios:
Scenario
Preference gate
Net win
Δ Pref
Runs (W/T/L)
= Decline a runtime latency request that is not a build performance issue
Excluded (activation contract)
+0.0%
+0.0%
0/1/0
= Diagnose NuGet restore running redundantly across CI stages
Eligible
+0.0%
+0.0%
0/1/0
▼ Diagnose a pathological ResolveAssemblyReference time
Eligible
-100.0%
-40.0%
0/0/1
= Diagnose per-project overhead across many small projects
Eligible
+0.0%
+0.0%
0/1/0
= Diagnose slow build for a small project
Eligible
+0.0%
+0.0%
0/1/0
Illustrative judge evidence:
Diagnose NuGet restore running redundantly across CI stages: Position-swap inconsistent (forward: B, reverse: tie). Defaulting to tie.
This is one example, not the aggregate verdict. Open Full Results for every judgment.
➖ Not proven improved — check-bin-obj-clash (claude-sonnet-4.6)
Why: Net win +28.6% (4W/1T/2L over 7 preference-eligible stimulus vote(s), sign test p=0.344), mean preference +20.0% across 7 paired run(s) — not credible (sign test p=0.344 > 0.05)
Next action: Inspect tied or lost stimuli and fix inconsistent skill behavior.
Gate evidence: n=7; 4W/1T/2L; d=6; p=0.344; net +28.6%
Warnings: Activation: isolated 6/7; plugin 6/7
Overfit: Low (score 0.19)
Repeated-run reliability (not used by the gate): 7 paired runs (4W/1T/2L).
Weak or warning scenarios:
Scenario
Preference gate
Net win
Δ Pref
Runs (W/T/L)
▼ Audit the mixed solution and separate safe projects from unsafe ones
Eligible
-100.0%
-40.0%
0/0/1
▼ Diagnose multi-targeting outputs that collapse into one path
Eligible
-100.0%
-40.0%
0/0/1
= Fix all clash mechanisms in the mixed solution
Eligible
+0.0%
+0.0%
0/1/0
Illustrative judge evidence:
Audit the mixed solution and separate safe projects from unsafe ones: The audits are substantively very similar and correctly classify the risks. A is slightly stronger overall because its LibraryA/LibraryB remediation actually separates both colliding output and intermediate directories; B's fix addresses only obj and its "Unsafe Projects (3)" ...
This is one example, not the aggregate verdict. Open Full Results for every judgment.
➖ Not proven improved — check-bin-obj-clash (gpt-5.6-luna)
Why: Net win -14.3% (0W/6T/1L over 7 preference-eligible stimulus vote(s), sign test p=0.500), mean preference -5.7% across 7 paired run(s) — no improvement
Next action: Inspect tied or lost stimuli; predeclare added breadth before a new experiment.
Gate evidence: n=7; 0W/6T/1L; d=1; p=0.500; net -14.3%
Overfit: Low (score 0.13)
Repeated-run reliability (not used by the gate): 7 paired runs (0W/6T/1L).
Weak or warning scenarios:
Scenario
Preference gate
Net win
Δ Pref
Runs (W/T/L)
▼ Audit the mixed solution and separate safe projects from unsafe ones
Eligible
-100.0%
-40.0%
0/0/1
= Avoid a false clash report when projects share only the top-level artifacts root
Eligible
+0.0%
+0.0%
0/1/0
= Decline output-clash remediation for separate projects using the default SDK layout
Eligible
+0.0%
+0.0%
0/1/0
= Diagnose multi-targeting outputs that collapse into one path
Eligible
+0.0%
+0.0%
0/1/0
= Diagnose redundant project reference metadata that forks a same-path build
Eligible
+0.0%
+0.0%
0/1/0
= Diagnose shared output and intermediate path collision
Eligible
+0.0%
+0.0%
0/1/0
= Fix all clash mechanisms in the mixed solution
Eligible
+0.0%
+0.0%
0/1/0
Illustrative judge evidence:
Audit the mixed solution and separate safe projects from unsafe ones: Both responses reach substantively correct and thorough conclusions with the same technical verification (msbuild property inspection, multi-target checks). The key differentiator is causal attribution: the rubric wants ConsumerApp flagged as unsafe (its reference metadata is ...
This is one example, not the aggregate verdict. Open Full Results for every judgment.
➖ Not proven improved — directory-build-organization (claude-sonnet-4.6)
Why: Net win +33.3% (3W/2T/1L over 6 preference-eligible stimulus vote(s), sign test p=0.312), mean preference +5.7% across 7 paired run(s), 1 dormancy stimulus/stimuli excluded from preference — not credible — 2 of 6 preference-eligible stimulus vote(s) tied, leaving only 4 discordant preference vote(s). The sign test conditions on non-tie stimulus votes and cannot reach 0.05 below 5, so no record could have passed here — this is not a measured null. Either the skill is inert on these scenarios (make them discriminate) or the eval needs more distinct stimuli to clear the ties
Next action: Inspect tied or lost stimuli; predeclare added breadth before a new experiment.
Repeated-run reliability (not used by the gate): 7 paired runs (3W/2T/2L).
Weak or warning scenarios:
Scenario
Preference gate
Net win
Δ Pref
Runs (W/T/L)
▼ Apply repo-level build organization cleanup
Eligible
-100.0%
-100.0%
0/0/1
▼ Decline adding shared build files to a lone project
Excluded (activation contract)
-100.0%
-40.0%
0/0/1
= Diagnose an inner shared-props file that overwrites its own override
Eligible
+0.0%
+0.0%
0/1/0
= Preserve an intentional project-specific exception while centralizing
Eligible
+0.0%
+0.0%
0/1/0
Illustrative judge evidence:
Apply repo-level build organization cleanup: A completed the requested build-layout refactor and verified the created shared files, while B only reported an unavailable path and delivered no task result.
This is one example, not the aggregate verdict. Open Full Results for every judgment.
➖ Not proven improved — eval-performance (claude-sonnet-4.6)
Why: Net win +62.5% (6W/1T/1L over 8 preference-eligible stimulus vote(s), sign test p=0.063), mean preference +40.0% across 8 paired run(s) — not credible (sign test p=0.063 > 0.05)
Next action: Inspect tied or lost stimuli and fix inconsistent skill behavior.
Gate evidence: n=8; 6W/1T/1L; d=7; p=0.063; net +62.5%
Warnings: Activation: isolated 8/8; plugin 7/8
Overfit: Moderate (score 0.27)
Repeated-run reliability (not used by the gate): 8 paired runs (6W/1T/1L).
Weak or warning scenarios:
Scenario
Preference gate
Net win
Δ Pref
Runs (W/T/L)
= Redirect a compile-time slowdown mistakenly framed as an evaluation problem
Eligible
+0.0%
+0.0%
0/1/0
▼ Triage which of two property functions actually costs evaluation time
Eligible
-100.0%
-40.0%
0/0/1
Illustrative judge evidence:
Redirect a compile-time slowdown mistakenly framed as an evaluation problem: Position-swap inconsistent (forward: A, reverse: B). Defaulting to tie.
This is one example, not the aggregate verdict. Open Full Results for every judgment.
➖ Not proven improved — eval-performance (gpt-5.6-luna)
Why: Net win +50.0% (5W/2T/1L over 8 preference-eligible stimulus vote(s), sign test p=0.109), mean preference +35.0% across 8 paired run(s) — not credible (sign test p=0.109 > 0.05)
Next action: Inspect tied or lost stimuli and fix inconsistent skill behavior.
= Triage which of two property functions actually costs evaluation time
Eligible
+0.0%
+0.0%
0/1/0
Illustrative judge evidence:
Detect a project evaluated twice under different global properties: Position-swap inconsistent (forward: B, reverse: A). Defaulting to tie.
This is one example, not the aggregate verdict. Open Full Results for every judgment.
➖ Not proven improved — extension-points (claude-sonnet-4.6)
Why: Net win +14.3% (3W/2T/2L over 7 preference-eligible stimulus vote(s), sign test p=0.500), mean preference +5.0% across 8 paired run(s), 1 dormancy stimulus/stimuli excluded from preference — not credible (sign test p=0.500 > 0.05)
Next action: Inspect tied or lost stimuli and fix inconsistent skill behavior.
Repeated-run reliability (not used by the gate): 8 paired runs (3W/3T/2L).
Weak or warning scenarios:
Scenario
Preference gate
Net win
Δ Pref
Runs (W/T/L)
▲ Diagnose a broken per-TFM forwarder
Eligible
+100.0%
+40.0%
1/0/0
▼ Diagnose a package ID and file-name mismatch
Eligible
-100.0%
-40.0%
0/0/1
▼ Diagnose build extension point failures
Eligible
-100.0%
-100.0%
0/0/1
= Fix extension point anti-patterns
Eligible
+0.0%
+0.0%
0/1/0
= Non-activation: repair an incremental custom target
Excluded (activation contract)
+0.0%
+0.0%
0/1/0
= Review packed layout without a false missing-file bug
Eligible
+0.0%
+0.0%
0/1/0
Illustrative judge evidence:
Diagnose a package ID and file-name mismatch: Both are correct, concise, and directly address the silent convention-based import failure. A is marginally stronger as a standalone final result because it includes exact corrected filenames, nuspec entries, and the necessary repackaging step; B's claimed applied fix is usefu...
This is one example, not the aggregate verdict. Open Full Results for every judgment.
➖ Not proven improved — extension-points (gpt-5.6-luna)
Why: Net win +57.1% (4W/3T/0L over 7 preference-eligible stimulus vote(s), sign test p=0.063), mean preference +22.5% across 8 paired run(s), 1 dormancy stimulus/stimuli excluded from preference — not credible — 3 of 7 preference-eligible stimulus vote(s) tied, leaving only 4 discordant preference vote(s). The sign test conditions on non-tie stimulus votes and cannot reach 0.05 below 5, so no record could have passed here — this is not a measured null. Either the skill is inert on these scenarios (make them discriminate) or the eval needs more distinct stimuli to clear the ties
Next action: Inspect tied or lost stimuli; predeclare added breadth before a new experiment.
Repeated-run reliability (not used by the gate): 8 paired runs (4W/3T/1L).
Weak or warning scenarios:
Scenario
Preference gate
Net win
Δ Pref
Runs (W/T/L)
= Diagnose build extension point failures
Eligible
+0.0%
+0.0%
0/1/0
= Fix extension point anti-patterns
Eligible
+0.0%
+0.0%
0/1/0
▼ Non-activation: repair an incremental custom target
Excluded (activation contract)
-100.0%
-40.0%
0/0/1
= Review packed layout without a false missing-file bug
Eligible
+0.0%
+0.0%
0/1/0
Illustrative judge evidence:
Diagnose build extension point failures: Position-swap inconsistent (forward: B, reverse: A). Defaulting to tie.
This is one example, not the aggregate verdict. Open Full Results for every judgment.
➖ Not proven improved — including-generated-files (gpt-5.6-luna)
Why: Net win +42.9% (4W/2T/1L over 7 preference-eligible stimulus vote(s), sign test p=0.188), mean preference +15.0% across 8 paired run(s), 1 dormancy stimulus/stimuli excluded from preference — not credible (sign test p=0.188 > 0.05)
Next action: Inspect tied or lost stimuli and fix inconsistent skill behavior.
Repeated-run reliability (not used by the gate): 8 paired runs (4W/3T/1L).
Weak or warning scenarios:
Scenario
Preference gate
Net win
Δ Pref
Runs (W/T/L)
= Decline manual wiring for Roslyn source generators
Excluded (activation contract)
+0.0%
+0.0%
0/1/0
= Diagnose missing clean tracking for generated source
Eligible
+0.0%
+0.0%
0/1/0
= Diagnose missing generated source inclusion
Eligible
+0.0%
+0.0%
0/1/0
▼ Diagnose missing output registration for generated non-code file
Eligible
-100.0%
-40.0%
0/0/1
Illustrative judge evidence:
Diagnose missing clean tracking for generated source: Position-swap inconsistent (forward: tie, reverse: A). Defaulting to tie.
This is one example, not the aggregate verdict. Open Full Results for every judgment.
➖ Not proven improved — incremental-build (claude-sonnet-4.6)
Why: Net win +33.3% (3W/6T/0L over 9 preference-eligible stimulus vote(s), sign test p=0.125), mean preference +13.3% across 9 paired run(s) — not credible — 6 of 9 preference-eligible stimulus vote(s) tied, leaving only 3 discordant preference vote(s). The sign test conditions on non-tie stimulus votes and cannot reach 0.05 below 5, so no record could have passed here — this is not a measured null. Either the skill is inert on these scenarios (make them discriminate) or the eval needs more distinct stimuli to clear the ties
Next action: Inspect tied or lost stimuli; predeclare added breadth before a new experiment.
= Fix broken incremental targets and clean tracking
Eligible
+0.0%
+0.0%
0/1/0
= Identify a volatile output path that defeats incrementality
Eligible
+0.0%
+0.0%
0/1/0
= Read a diagnostic log to find the stale input
Eligible
+0.0%
+0.0%
0/1/0
Illustrative judge evidence:
Diagnose custom targets that always rerun: Position-swap inconsistent (forward: B, reverse: A). Defaulting to tie.
This is one example, not the aggregate verdict. Open Full Results for every judgment.
➖ Not proven improved — incremental-build (gpt-5.6-luna)
Why: Net win +33.3% (4W/4T/1L over 9 preference-eligible stimulus vote(s), sign test p=0.188), mean preference +13.3% across 9 paired run(s) — not credible (sign test p=0.188 > 0.05)
Next action: Inspect tied or lost stimuli and fix inconsistent skill behavior.
Gate evidence: n=9; 4W/4T/1L; d=5; p=0.188; net +33.3%
Warnings: Activation: isolated 9/9; plugin 8/9
Overfit: Low (score 0.18)
Repeated-run reliability (not used by the gate): 9 paired runs (4W/4T/1L).
Weak or warning scenarios:
Scenario
Preference gate
Net win
Δ Pref
Runs (W/T/L)
▼ Diagnose custom targets that always rerun
Eligible
-100.0%
-40.0%
0/0/1
= Distinguish a cold first build from broken incrementality
Eligible
+0.0%
+0.0%
0/1/0
▲ Explain slower builds when MSBuild skipped everything
Eligible
+100.0%
+40.0%
1/0/0
= Fix broken incremental targets and clean tracking
Eligible
+0.0%
+0.0%
0/1/0
= Identify a volatile output path that defeats incrementality
Eligible
+0.0%
+0.0%
0/1/0
= Read a diagnostic log to find the stale input
Eligible
+0.0%
+0.0%
0/1/0
Illustrative judge evidence:
Diagnose custom targets that always rerun: Both responses correctly diagnose the problem and identify the two targets. A is slightly better because it provides concrete corrected XML with explicit Inputs/Outputs, making the fix actionable, while B stays at a prose description.
This is one example, not the aggregate verdict. Open Full Results for every judgment.
➖ Not proven improved — item-management (gpt-5.6-luna)
Why: Net win -16.7% (0W/5T/1L over 6 preference-eligible stimulus vote(s), sign test p=0.500), mean preference -5.7% across 7 paired run(s), 1 dormancy stimulus/stimuli excluded from preference — no improvement
Next action: Inspect tied or lost stimuli; predeclare added breadth before a new experiment.
Repeated-run reliability (not used by the gate): 7 paired runs (0W/6T/1L).
Weak or warning scenarios:
Scenario
Preference gate
Net win
Δ Pref
Runs (W/T/L)
= Decline a target ordering problem with no item-group defect
Excluded (activation contract)
+0.0%
+0.0%
0/1/0
= Diagnose an ineffective Compile Remove that does not match the glob
Eligible
+0.0%
+0.0%
0/1/0
= Diagnose item group and batching issues
Eligible
+0.0%
+0.0%
0/1/0
▼ Diagnose real and claimed item problems in a code generation pipeline
Eligible
-100.0%
-40.0%
0/0/1
= Fix item management anti-patterns
Eligible
+0.0%
+0.0%
0/1/0
= Leave already-correct item management unchanged
Eligible
+0.0%
+0.0%
0/1/0
= Leave correct single-list batching unchanged
Eligible
+0.0%
+0.0%
0/1/0
Illustrative judge evidence:
Diagnose an ineffective Compile Remove that does not match the glob: Position-swap inconsistent (forward: B, reverse: tie). Defaulting to tie.
This is one example, not the aggregate verdict. Open Full Results for every judgment.
➖ Not proven improved — msbuild-antipatterns (claude-sonnet-4.6)
Why: Net win -14.3% (1W/4T/2L over 7 preference-eligible stimulus vote(s), sign test p=0.500), mean preference -17.5% across 8 paired run(s), 1 dormancy stimulus/stimuli excluded from preference — no improvement
Next action: Inspect tied or lost stimuli; predeclare added breadth before a new experiment.
Repeated-run reliability (not used by the gate): 8 paired runs (1W/4T/3L).
Weak or warning scenarios:
Scenario
Preference gate
Net win
Δ Pref
Runs (W/T/L)
= Add a module to an F# project
Eligible
+0.0%
+0.0%
0/1/0
= Add a signature file to define public API
Eligible
+0.0%
+0.0%
0/1/0
▼ Distinguish a style backslash from a real cross-platform backslash bug
Eligible
-100.0%
-100.0%
0/0/1
▼ Fix broken file order causing FS0039
Eligible
-100.0%
-40.0%
0/0/1
= Judge an unguarded import inside a NuGet package build folder
Eligible
+0.0%
+0.0%
0/1/0
= Leave a clean project without inventing anti-patterns
Eligible
+0.0%
+0.0%
0/1/0
▼ Non-activation: migrate a legacy project to SDK style
Excluded (activation contract)
-100.0%
-40.0%
0/0/1
Illustrative judge evidence:
Add a module to an F# project: Both runs implement the requested functionality correctly, include the new file in valid F# compilation order, update processing flow for validation failures, and successfully build/run the application. B's placement directly after Domain.fs is marginally tidier, but A's place...
This is one example, not the aggregate verdict. Open Full Results for every judgment.
➖ Not proven improved — msbuild-antipatterns (gpt-5.6-luna)
Why: Net win -14.3% (2W/2T/3L over 7 preference-eligible stimulus vote(s), sign test p=0.500), mean preference -2.5% across 8 paired run(s), 1 dormancy stimulus/stimuli excluded from preference — no improvement
Next action: Inspect tied or lost stimuli and fix inconsistent skill behavior.
Repeated-run reliability (not used by the gate): 8 paired runs (2W/2T/4L).
Weak or warning scenarios:
Scenario
Preference gate
Net win
Δ Pref
Runs (W/T/L)
▲ Add a module to an F# project
Eligible
+100.0%
+40.0%
1/0/0
= Add a signature file to define public API
Eligible
+0.0%
+0.0%
0/1/0
= Distinguish a style backslash from a real cross-platform backslash bug
Eligible
+0.0%
+0.0%
0/1/0
▼ Fix broken file order causing FS0039
Eligible
-100.0%
-40.0%
0/0/1
▼ Judge an unguarded import inside a NuGet package build folder
Eligible
-100.0%
-40.0%
0/0/1
▼ Non-activation: migrate a legacy project to SDK style
Excluded (activation contract)
-100.0%
-40.0%
0/0/1
▼ Review MSBuild files for anti-patterns and style issues
Eligible
-100.0%
-40.0%
0/0/1
Illustrative judge evidence:
Add a signature file to define public API: Both responses produced functionally identical, correct results: an appropriate Domain.fsi with all types, placed before Domain.fs in the project file, with a verified clean build. Their approaches, recovery from the apply_patch failure, and final verification were essentially...
This is one example, not the aggregate verdict. Open Full Results for every judgment.
➖ Not proven improved — msbuild-modernization (claude-sonnet-4.6)
Why: Net win -16.7% (1W/3T/2L over 6 preference-eligible stimulus vote(s), sign test p=0.500), mean preference +2.9% across 7 paired run(s), 1 dormancy stimulus/stimuli excluded from preference — no improvement
Next action: Inspect tied or lost stimuli; predeclare added breadth before a new experiment.
Repeated-run reliability (not used by the gate): 7 paired runs (1W/4T/2L).
Weak or warning scenarios:
Scenario
Preference gate
Net win
Δ Pref
Runs (W/T/L)
▼ Apply the migration to SDK-style
Eligible
-100.0%
-40.0%
0/0/1
= Consolidate duplicated projects into a multi-targeting SDK-style project
Eligible
+0.0%
+0.0%
0/1/0
= Decline modernizing a non-.NET build
Excluded (activation contract)
+0.0%
+0.0%
0/1/0
▼ Identify legacy patterns for SDK-style migration
Eligible
-100.0%
-40.0%
0/0/1
= Modernize a single project without introducing Central Package Management
Eligible
+0.0%
+0.0%
0/1/0
▲ Modernize further without introducing a nondeterministic language version
Eligible
+100.0%
+100.0%
1/0/0
= Recognize an already-modern project needs no migration
Eligible
+0.0%
+0.0%
0/1/0
Illustrative judge evidence:
Apply the migration to SDK-style: Both conversions satisfy the core SDK-style migration and should avoid duplicate assembly attributes. A is marginally safer and less behavior-changing because it preserves the complete existing AssemblyInfo source rather than relying on a manual migration of its attributes.
This is one example, not the aggregate verdict. Open Full Results for every judgment.
➖ Not proven improved — msbuild-modernization (gpt-5.6-luna)
Why: Net win +16.7% (1W/5T/0L over 6 preference-eligible stimulus vote(s), sign test p=0.500), mean preference +0.0% across 7 paired run(s), 1 dormancy stimulus/stimuli excluded from preference — not credible — 5 of 6 preference-eligible stimulus vote(s) tied, leaving only 1 discordant preference vote(s). The sign test conditions on non-tie stimulus votes and cannot reach 0.05 below 5, so no record could have passed here — this is not a measured null. Either the skill is inert on these scenarios (make them discriminate) or the eval needs more distinct stimuli to clear the ties
Next action: Inspect tied or lost stimuli; predeclare added breadth before a new experiment.
To investigate non-passing or warning results, paste this to your AI coding agent:
For PR 1118 in dotnet/skills, download eval artifacts with gh run download 34273037202 --repo dotnet/skills --pattern "vally-results-*" --dir ./eval-results, then fetch https://raw.githubusercontent.com/dotnet/skills/6866836a9b096ae2e9aefefe6dbfdab67405f37f/eng/vally-adapter/InvestigatingResults.md and follow it. Classify each result as measurement-invalid, underpowered, not-proven, preference-loss, or passing-with-warning. Use stateReason, result accounting, weak scenarios, and judge evidence to give the cause and exact next fix.
Put required input constraints before overlapping performance, generated-file, item, and property vocabulary identified by the second CI evaluation.
Co-authored-by: Copilot App <223556219+Copilot@users.noreply.github.com>
Copilot-Session: 2d0a0a20-fd83-4343-8434-4763f8120503
❌ Evaluation did not complete successfully (the evaluate job reported failure). Check the workflow run logs, then comment /evaluate 26b6a729a5c819dc4dd3fb45a8beb43ee14c3ba1 to retry this exact commit.
36 partial result file(s) were preserved for diagnosis but were not consolidated because the full matrix did not complete.
Preserve the off-target contracts with leading exclusion rules and retry transient binlog setup failures that otherwise drop one comparison arm.
Co-authored-by: Copilot App <223556219+Copilot@users.noreply.github.com>
Copilot-Session: 2d0a0a20-fd83-4343-8434-4763f8120503
❌ Evaluation did not complete successfully (the evaluate job reported failure). Check the workflow run logs, then comment /evaluate b37f36513cac14d8fd738759cfee7678f3b65483 to retry this exact commit.
36 partial result file(s) were preserved for diagnosis but were not consolidated because the full matrix did not complete.
Describe the positive applicability test for the three remaining Sonnet dormancy misses without leading with overlapping off-target vocabulary.
Co-authored-by: Copilot App <223556219+Copilot@users.noreply.github.com>
Copilot-Session: 2d0a0a20-fd83-4343-8434-4763f8120503
Objective completion gate: not enabled. Aggregate completion transitions are telemetry only, so this report does not claim that zero objective regressions were proven.
A result passes only when preference-eligible distinct-stimulus votes have aggregate net win of at least 20% and an exact one-sided sign-test result of p ≤ 0.05, and every explicit dormancy activation contract passes. Repeated runs measure reliability only.
Skill
Model
Verdict
Gate evidence
Overfit
Warnings
Next action
binlog-failure-analysis
claude-sonnet-4.6
✅ Improved
n=7; 5W/2T/0L; d=5; p=0.031; net +71.4%; 1 dormancy excluded
✅ 0.19
—
None.
binlog-failure-analysis
gpt-5.6-luna
✅ Improved
n=7; 6W/1T/0L; d=6; p=0.016; net +85.7%; 1 dormancy excluded
✅ 0.11
—
None.
binlog-generation
claude-sonnet-4.6
✅ Improved
n=7; 6W/1T/0L; d=6; p=0.016; net +85.7%; 1 dormancy excluded
🔴 0.56
Activation: isolated 7/7; plugin 6/7
Fix activation gaps; Review overfit evidence.
binlog-generation
gpt-5.6-luna
✅ Improved
n=7; 5W/2T/0L; d=5; p=0.031; net +71.4%; 1 dormancy excluded
🟡 0.26
—
Review overfit evidence.
build-parallelism
claude-sonnet-4.6
⛔ Activation contract failed
n=6; 2W/3T/1L; d=3; p=0.500; net +16.7%; 1 dormancy excluded
Narrow skill routing so the listed off-target scenarios stay dormant.
property-patterns
gpt-5.6-luna
➖ Not proven improved
n=6; 3W/2T/1L; d=4; p=0.312; net +33.3%; 1 dormancy excluded
✅ 0.08
—
Inspect tied or lost stimuli; predeclare added breadth before a new experiment.
resolve-project-references
claude-sonnet-4.6
➖ Not proven improved
n=6; 5W/0T/1L; d=6; p=0.109; net +66.7%; 2 dormancy excluded
🟡 0.36
—
Inspect tied or lost stimuli and fix inconsistent skill behavior.
resolve-project-references
gpt-5.6-luna
➖ Not proven improved
n=6; 2W/3T/1L; d=3; p=0.500; net +16.7%; 2 dormancy excluded
✅ 0.15
Activation: isolated 4/6; plugin 6/6
Inspect tied or lost stimuli; predeclare added breadth before a new experiment.
target-authoring
claude-sonnet-4.6
➖ Not proven improved
n=6; 3W/2T/1L; d=4; p=0.312; net +33.3%; 1 dormancy excluded
🟡 0.35
—
Inspect tied or lost stimuli; predeclare added breadth before a new experiment.
target-authoring
gpt-5.6-luna
➖ Not proven improved
n=6; 3W/1T/2L; d=5; p=0.500; net +16.7%; 1 dormancy excluded
✅ 0.12
—
Inspect tied or lost stimuli and fix inconsistent skill behavior.
ℹ️ How to read this report
✅ Improved — the result passed both the statistical gate and the 20% practical net-win floor.
➖ Not proven improved — the result is valid but did not pass both gates. This is not automatically a regression.
⚠️ Invalid / underpowered — the gate withheld a quality verdict. Fix the measurement before judging the skill.
⛔ Activation contract failed — the isolated target skill activated on an explicit dormancy scenario. Dormancy preference is excluded, but this routing failure still blocks a pass.
📉 Preference loss — the LLM judge credibly preferred baseline. It is report-only, not objective completion proof.
Gate evidence — n preference-eligible distinct-stimulus votes, W/T/L stimulus votes, d discordant votes, exact one-sided p, net win, and the count of separately retained dormancy stimuli. The p value applies to one model/skill result; no matrix-wide multiple-comparison correction is applied.
Warnings — activation, timeout, retry recovery, or unresolved comparison conditions that need attention.
Do not add repeated runs to increase statistical power. Do not add stimuli after seeing a near-pass unless the new breadth is predeclared for a new experiment.
Repeated-run reliability (not used by the gate): 7 paired runs (5W/2T/0L).
Weak or warning scenarios:
Scenario
Preference gate
Net win
Δ Pref
Runs (W/T/L)
= Choose the right OS detection for a cross-platform property
Eligible
+0.0%
+0.0%
0/1/0
▲ Decline a props-versus-targets placement question
Excluded (activation contract)
+100.0%
+40.0%
1/0/0
▲ Diagnose a TargetFramework condition that never applies in props
Eligible
+100.0%
+40.0%
1/0/0
= Diagnose multi-level property hierarchy bugs
Eligible
+0.0%
+0.0%
0/1/0
Illustrative judge evidence:
Choose the right OS detection for a cross-platform property: Position-swap inconsistent (forward: tie, reverse: B). Defaulting to tie.
This is one example, not the aggregate verdict. Open Full Results for every judgment.
➖ Not proven improved — build-parallelism (gpt-5.6-luna)
Why: Net win +16.7% (2W/3T/1L over 6 preference-eligible stimulus vote(s), sign test p=0.500), mean preference +14.3% across 7 paired run(s), 1 dormancy stimulus/stimuli excluded from preference — not credible — 3 of 6 preference-eligible stimulus vote(s) tied, leaving only 3 discordant preference vote(s). The sign test conditions on non-tie stimulus votes and cannot reach 0.05 below 5, so no record could have passed here — this is not a measured null. Either the skill is inert on these scenarios (make them discriminate) or the eval needs more distinct stimuli to clear the ties
Next action: Inspect tied or lost stimuli; predeclare added breadth before a new experiment.
This is one example, not the aggregate verdict. Open Full Results for every judgment.
➖ Not proven improved — build-perf-baseline (claude-sonnet-4.6)
Why: Net win -33.3% (2W/0T/4L over 6 preference-eligible stimulus vote(s), sign test p=0.344), mean preference -25.7% across 7 paired run(s), 1 dormancy stimulus/stimuli excluded from preference — no improvement
Next action: Inspect tied or lost stimuli and fix inconsistent skill behavior.
Repeated-run reliability (not used by the gate): 7 paired runs (2W/0T/5L).
Weak or warning scenarios:
Scenario
Preference gate
Net win
Δ Pref
Runs (W/T/L)
▼ Configure deterministic, cache-safe CI builds
Eligible
-100.0%
-40.0%
0/0/1
▼ Decline a non-MSBuild build performance request
Excluded (activation contract)
-100.0%
-100.0%
0/0/1
▼ Leave an already-optimized build unchanged
Eligible
-100.0%
-100.0%
0/0/1
▼ Route a broken no-op rebuild away from generic optimization
Eligible
-100.0%
-40.0%
0/0/1
▼ Route a restore-bound cold build away from architecture changes
Eligible
-100.0%
-40.0%
0/0/1
Illustrative judge evidence:
Configure deterministic, cache-safe CI builds: Both miss the required CI-guarded ContinuousIntegrationBuild change and give materially misleading root-cause explanations. A is only marginally better because it acknowledges the SDK default for Deterministic, even though its chosen fix remains unnecessary and insufficient.
This is one example, not the aggregate verdict. Open Full Results for every judgment.
➖ Not proven improved — build-perf-baseline (gpt-5.6-luna)
Why: Net win +50.0% (4W/1T/1L over 6 preference-eligible stimulus vote(s), sign test p=0.188), mean preference +31.4% across 7 paired run(s), 1 dormancy stimulus/stimuli excluded from preference — not credible (sign test p=0.188 > 0.05)
Next action: Inspect tied or lost stimuli and fix inconsistent skill behavior.
Repeated-run reliability (not used by the gate): 7 paired runs (5W/1T/1L).
Weak or warning scenarios:
Scenario
Preference gate
Net win
Δ Pref
Runs (W/T/L)
▼ Configure deterministic, cache-safe CI builds
Eligible
-100.0%
-40.0%
0/0/1
= Route a broken no-op rebuild away from generic optimization
Eligible
+0.0%
+0.0%
0/1/0
▲ Route a restore-bound cold build away from architecture changes
Eligible
+100.0%
+40.0%
1/0/0
Illustrative judge evidence:
Configure deterministic, cache-safe CI builds: The responses are nearly identical in approach, change, and outcome. Both added ContinuousIntegrationBuild unconditionally (missing the CI-only guard), both verified the build, both explained the rationale briefly. A slightly edges out by additionally verifying DeterministicSo...
This is one example, not the aggregate verdict. Open Full Results for every judgment.
➖ Not proven improved — build-perf-diagnostics (claude-sonnet-4.6)
Why: Net win +28.6% (4W/1T/2L over 7 preference-eligible stimulus vote(s), sign test p=0.344), mean preference +5.0% across 8 paired run(s), 1 dormancy stimulus/stimuli excluded from preference — not credible (sign test p=0.344 > 0.05)
Next action: Inspect tied or lost stimuli and fix inconsistent skill behavior.
Repeated-run reliability (not used by the gate): 8 paired runs (4W/1T/3L).
Weak or warning scenarios:
Scenario
Preference gate
Net win
Δ Pref
Runs (W/T/L)
▼ Decline a runtime latency request that is not a build performance issue
Excluded (activation contract)
-100.0%
-40.0%
0/0/1
▼ Diagnose NuGet restore running redundantly across CI stages
Eligible
-100.0%
-40.0%
0/0/1
▼ Diagnose evaluation overhead before any target runs
Eligible
-100.0%
-40.0%
0/0/1
= Diagnose per-project overhead across many small projects
Eligible
+0.0%
+0.0%
0/1/0
Illustrative judge evidence:
Diagnose NuGet restore running redundantly across CI stages: The core diagnosis and prescribed changes are essentially the same, but A is more concise and avoids B's notably misleading after-time table: a dedicated restore still costs about 9.7s, and test/publish have work beyond compilation, so a 4–5s total CI estimate is unsupported. ...
This is one example, not the aggregate verdict. Open Full Results for every judgment.
➖ Not proven improved — build-perf-diagnostics (gpt-5.6-luna)
Why: Net win +57.1% (5W/1T/1L over 7 preference-eligible stimulus vote(s), sign test p=0.109), mean preference +20.0% across 8 paired run(s), 1 dormancy stimulus/stimuli excluded from preference — not credible (sign test p=0.109 > 0.05)
Next action: Inspect tied or lost stimuli and fix inconsistent skill behavior.
Repeated-run reliability (not used by the gate): 8 paired runs (5W/2T/1L).
Weak or warning scenarios:
Scenario
Preference gate
Net win
Δ Pref
Runs (W/T/L)
= Decline a runtime latency request that is not a build performance issue
Excluded (activation contract)
+0.0%
+0.0%
0/1/0
▼ Diagnose NuGet restore running redundantly across CI stages
Eligible
-100.0%
-40.0%
0/0/1
= Diagnose evaluation overhead before any target runs
Eligible
+0.0%
+0.0%
0/1/0
Illustrative judge evidence:
Diagnose NuGet restore running redundantly across CI stages: Both responses are essentially correct and equivalent on the core diagnosis and main recommendations, covering all three rubric points. A edges ahead slightly by adding --locked-mode (appropriate given the committed lock file) and explicitly noting the need to preserve restore...
This is one example, not the aggregate verdict. Open Full Results for every judgment.
➖ Not proven improved — check-bin-obj-clash (claude-sonnet-4.6)
Why: Net win +28.6% (3W/3T/1L over 7 preference-eligible stimulus vote(s), sign test p=0.312), mean preference +20.0% across 7 paired run(s) — not credible — 3 of 7 preference-eligible stimulus vote(s) tied, leaving only 4 discordant preference vote(s). The sign test conditions on non-tie stimulus votes and cannot reach 0.05 below 5, so no record could have passed here — this is not a measured null. Either the skill is inert on these scenarios (make them discriminate) or the eval needs more distinct stimuli to clear the ties
Next action: Inspect tied or lost stimuli; predeclare added breadth before a new experiment.
Gate evidence: n=7; 3W/3T/1L; d=4; p=0.312; net +28.6%
Warnings: Activation: isolated 6/7; plugin 6/7
Overfit: Low (score 0.18)
Repeated-run reliability (not used by the gate): 7 paired runs (3W/3T/1L).
Weak or warning scenarios:
Scenario
Preference gate
Net win
Δ Pref
Runs (W/T/L)
= Audit the mixed solution and separate safe projects from unsafe ones
Eligible
+0.0%
+0.0%
0/1/0
= Decline output-clash remediation for separate projects using the default SDK layout
Eligible
+0.0%
+0.0%
0/1/0
= Diagnose multi-targeting outputs that collapse into one path
Eligible
+0.0%
+0.0%
0/1/0
▼ Fix all clash mechanisms in the mixed solution
Eligible
-100.0%
-40.0%
0/0/1
Illustrative judge evidence:
Audit the mixed solution and separate safe projects from unsafe ones: Position-swap inconsistent (forward: B, reverse: A). Defaulting to tie.
This is one example, not the aggregate verdict. Open Full Results for every judgment.
➖ Not proven improved — check-bin-obj-clash (gpt-5.6-luna)
Why: Net win +57.1% (4W/3T/0L over 7 preference-eligible stimulus vote(s), sign test p=0.063), mean preference +22.9% across 7 paired run(s) — not credible — 3 of 7 preference-eligible stimulus vote(s) tied, leaving only 4 discordant preference vote(s). The sign test conditions on non-tie stimulus votes and cannot reach 0.05 below 5, so no record could have passed here — this is not a measured null. Either the skill is inert on these scenarios (make them discriminate) or the eval needs more distinct stimuli to clear the ties
Next action: Inspect tied or lost stimuli; predeclare added breadth before a new experiment.
Gate evidence: n=7; 4W/3T/0L; d=4; p=0.063; net +57.1%
Overfit: Low (score 0.11)
Repeated-run reliability (not used by the gate): 7 paired runs (4W/3T/0L).
Weak or warning scenarios:
Scenario
Preference gate
Net win
Δ Pref
Runs (W/T/L)
= Audit the mixed solution and separate safe projects from unsafe ones
Eligible
+0.0%
+0.0%
0/1/0
= Diagnose multi-targeting outputs that collapse into one path
Eligible
+0.0%
+0.0%
0/1/0
= Diagnose shared output and intermediate path collision
Eligible
+0.0%
+0.0%
0/1/0
Illustrative judge evidence:
Audit the mixed solution and separate safe projects from unsafe ones: Position-swap inconsistent (forward: tie, reverse: B). Defaulting to tie.
This is one example, not the aggregate verdict. Open Full Results for every judgment.
➖ Not proven improved — directory-build-organization (claude-sonnet-4.6)
Why: Net win +33.3% (2W/4T/0L over 6 preference-eligible stimulus vote(s), sign test p=0.250), mean preference +20.0% across 7 paired run(s), 1 dormancy stimulus/stimuli excluded from preference — not credible — 4 of 6 preference-eligible stimulus vote(s) tied, leaving only 2 discordant preference vote(s). The sign test conditions on non-tie stimulus votes and cannot reach 0.05 below 5, so no record could have passed here — this is not a measured null. Either the skill is inert on these scenarios (make them discriminate) or the eval needs more distinct stimuli to clear the ties
Next action: Inspect tied or lost stimuli; predeclare added breadth before a new experiment.
Repeated-run reliability (not used by the gate): 7 paired runs (2W/5T/0L).
Weak or warning scenarios:
Scenario
Preference gate
Net win
Δ Pref
Runs (W/T/L)
= Decline adding shared build files to a lone project
Excluded (activation contract)
+0.0%
+0.0%
0/1/0
= Diagnose a TargetFramework condition that silently skips in props
Eligible
+0.0%
+0.0%
0/1/0
= Diagnose a package downgrade chain and reorganize version management
Eligible
+0.0%
+0.0%
0/1/0
= Diagnose an inner shared-props file that overwrites its own override
Eligible
+0.0%
+0.0%
0/1/0
= Preserve an intentional project-specific exception while centralizing
Eligible
+0.0%
+0.0%
0/1/0
Illustrative judge evidence:
Diagnose a TargetFramework condition that silently skips in props: Position-swap inconsistent (forward: tie, reverse: B). Defaulting to tie.
This is one example, not the aggregate verdict. Open Full Results for every judgment.
➖ Not proven improved — directory-build-organization (gpt-5.6-luna)
Why: Net win +16.7% (1W/5T/0L over 6 preference-eligible stimulus vote(s), sign test p=0.500), mean preference +5.7% across 7 paired run(s), 1 dormancy stimulus/stimuli excluded from preference — not credible — 5 of 6 preference-eligible stimulus vote(s) tied, leaving only 1 discordant preference vote(s). The sign test conditions on non-tie stimulus votes and cannot reach 0.05 below 5, so no record could have passed here — this is not a measured null. Either the skill is inert on these scenarios (make them discriminate) or the eval needs more distinct stimuli to clear the ties
Next action: Inspect tied or lost stimuli; predeclare added breadth before a new experiment.
This is one example, not the aggregate verdict. Open Full Results for every judgment.
➖ Not proven improved — eval-performance (claude-sonnet-4.6)
Why: Net win +37.5% (5W/1T/2L over 8 preference-eligible stimulus vote(s), sign test p=0.227), mean preference +30.0% across 8 paired run(s) — not credible (sign test p=0.227 > 0.05)
Next action: Inspect tied or lost stimuli and fix inconsistent skill behavior.
Gate evidence: n=8; 5W/1T/2L; d=7; p=0.227; net +37.5%
Warnings: Activation: isolated 8/8; plugin 7/8
Overfit: Moderate (score 0.36)
Repeated-run reliability (not used by the gate): 8 paired runs (5W/1T/2L).
Weak or warning scenarios:
Scenario
Preference gate
Net win
Δ Pref
Runs (W/T/L)
▼ Recognize TreatAsLocalProperty overuse versus one justified entry
Eligible
-100.0%
-40.0%
0/0/1
▼ Redirect a compile-time slowdown mistakenly framed as an evaluation problem
Eligible
-100.0%
-40.0%
0/0/1
= Triage which of two property functions actually costs evaluation time
Eligible
+0.0%
+0.0%
0/1/0
Illustrative judge evidence:
Recognize TreatAsLocalProperty overuse versus one justified entry: Both satisfy every requested rubric item and give the right change. A is marginally stronger because it avoids B's questionable claims about blocking parent overrides and inheritance behavior.
This is one example, not the aggregate verdict. Open Full Results for every judgment.
➖ Not proven improved — extension-points (claude-sonnet-4.6)
Why: Net win +0.0% (2W/3T/2L over 7 preference-eligible stimulus vote(s), sign test p=0.687), mean preference -12.5% across 8 paired run(s), 1 dormancy stimulus/stimuli excluded from preference — no improvement
Next action: Inspect tied or lost stimuli; predeclare added breadth before a new experiment.
Repeated-run reliability (not used by the gate): 8 paired runs (2W/3T/3L).
Weak or warning scenarios:
Scenario
Preference gate
Net win
Δ Pref
Runs (W/T/L)
= Create extensibility hooks for a custom SDK target file
Eligible
+0.0%
+0.0%
0/1/0
▼ Diagnose a broken per-TFM forwarder
Eligible
-100.0%
-40.0%
0/0/1
= Diagnose a package ID and file-name mismatch
Eligible
+0.0%
+0.0%
0/1/0
▼ Fix extension point anti-patterns
Eligible
-100.0%
-100.0%
0/0/1
▼ Non-activation: repair an incremental custom target
Excluded (activation contract)
-100.0%
-40.0%
0/0/1
= Review packed layout without a false missing-file bug
Eligible
+0.0%
+0.0%
0/1/0
Illustrative judge evidence:
Create extensibility hooks for a custom SDK target file: Position-swap inconsistent (forward: B, reverse: A). Defaulting to tie.
This is one example, not the aggregate verdict. Open Full Results for every judgment.
➖ Not proven improved — extension-points (gpt-5.6-luna)
Why: Net win +28.6% (3W/3T/1L over 7 preference-eligible stimulus vote(s), sign test p=0.312), mean preference +22.5% across 8 paired run(s), 1 dormancy stimulus/stimuli excluded from preference — not credible — 3 of 7 preference-eligible stimulus vote(s) tied, leaving only 4 discordant preference vote(s). The sign test conditions on non-tie stimulus votes and cannot reach 0.05 below 5, so no record could have passed here — this is not a measured null. Either the skill is inert on these scenarios (make them discriminate) or the eval needs more distinct stimuli to clear the ties
Next action: Inspect tied or lost stimuli; predeclare added breadth before a new experiment.
Repeated-run reliability (not used by the gate): 8 paired runs (4W/3T/1L).
Weak or warning scenarios:
Scenario
Preference gate
Net win
Δ Pref
Runs (W/T/L)
= Diagnose NuGet package and repo extension conflicts
Eligible
+0.0%
+0.0%
0/1/0
▼ Diagnose build extension point failures
Eligible
-100.0%
-40.0%
0/0/1
= Fix extension point anti-patterns
Eligible
+0.0%
+0.0%
0/1/0
= Review packed layout without a false missing-file bug
Eligible
+0.0%
+0.0%
0/1/0
Illustrative judge evidence:
Diagnose NuGet package and repo extension conflicts: Position-swap inconsistent (forward: tie, reverse: A). Defaulting to tie.
This is one example, not the aggregate verdict. Open Full Results for every judgment.
➖ Not proven improved — including-generated-files (gpt-5.6-luna)
Why: Net win +0.0% (1W/5T/1L over 7 preference-eligible stimulus vote(s), sign test p=0.750), mean preference +0.0% across 8 paired run(s), 1 dormancy stimulus/stimuli excluded from preference — no improvement
Next action: Inspect tied or lost stimuli; predeclare added breadth before a new experiment.
Repeated-run reliability (not used by the gate): 8 paired runs (1W/6T/1L).
Weak or warning scenarios:
Scenario
Preference gate
Net win
Δ Pref
Runs (W/T/L)
= Decline manual wiring for Roslyn source generators
Excluded (activation contract)
+0.0%
+0.0%
0/1/0
= Diagnose hardcoded obj path for generated source
Eligible
+0.0%
+0.0%
0/1/0
▼ Diagnose missing clean tracking for generated source
Eligible
-100.0%
-40.0%
0/0/1
= Diagnose missing generated source inclusion
Eligible
+0.0%
+0.0%
0/1/0
= Diagnose missing output registration for generated non-code file
Eligible
+0.0%
+0.0%
0/1/0
= Diagnose project-level glob for generated source
Eligible
+0.0%
+0.0%
0/1/0
= Diagnose wrong hook for generated source files
Eligible
+0.0%
+0.0%
0/1/0
Illustrative judge evidence:
Diagnose hardcoded obj path for generated source: Position-swap inconsistent (forward: B, reverse: tie). Defaulting to tie.
This is one example, not the aggregate verdict. Open Full Results for every judgment.
➖ Not proven improved — incremental-build (claude-sonnet-4.6)
Why: Net win +44.4% (4W/5T/0L over 9 preference-eligible stimulus vote(s), sign test p=0.063), mean preference +17.8% across 9 paired run(s) — not credible — 5 of 9 preference-eligible stimulus vote(s) tied, leaving only 4 discordant preference vote(s). The sign test conditions on non-tie stimulus votes and cannot reach 0.05 below 5, so no record could have passed here — this is not a measured null. Either the skill is inert on these scenarios (make them discriminate) or the eval needs more distinct stimuli to clear the ties
Next action: Inspect tied or lost stimuli; predeclare added breadth before a new experiment.
= Fix broken incremental targets and clean tracking
Eligible
+0.0%
+0.0%
0/1/0
= Identify a volatile output path that defeats incrementality
Eligible
+0.0%
+0.0%
0/1/0
▲ Read a diagnostic log to find the stale input
Eligible
+100.0%
+40.0%
1/0/0
Illustrative judge evidence:
Diagnose custom targets that always rerun: Position-swap inconsistent (forward: B, reverse: tie). Defaulting to tie.
This is one example, not the aggregate verdict. Open Full Results for every judgment.
➖ Not proven improved — incremental-build (gpt-5.6-luna)
Why: Net win +22.2% (3W/5T/1L over 9 preference-eligible stimulus vote(s), sign test p=0.312), mean preference +15.6% across 9 paired run(s) — not credible — 5 of 9 preference-eligible stimulus vote(s) tied, leaving only 4 discordant preference vote(s). The sign test conditions on non-tie stimulus votes and cannot reach 0.05 below 5, so no record could have passed here — this is not a measured null. Either the skill is inert on these scenarios (make them discriminate) or the eval needs more distinct stimuli to clear the ties
Next action: Inspect tied or lost stimuli; predeclare added breadth before a new experiment.
Gate evidence: n=9; 3W/5T/1L; d=4; p=0.312; net +22.2%
Warnings: Activation: isolated 9/9; plugin 7/9
Overfit: Low (score 0.14)
Repeated-run reliability (not used by the gate): 9 paired runs (3W/5T/1L).
Weak or warning scenarios:
Scenario
Preference gate
Net win
Δ Pref
Runs (W/T/L)
= Diagnose custom targets that always rerun
Eligible
+0.0%
+0.0%
0/1/0
= Explain slower builds when MSBuild skipped everything
Eligible
+0.0%
+0.0%
0/1/0
▼ Explain why Visual Studio keeps rebuilding an up-to-date project
Eligible
-100.0%
-40.0%
0/0/1
= Fix broken incremental targets and clean tracking
Eligible
+0.0%
+0.0%
0/1/0
= Identify a volatile output path that defeats incrementality
Eligible
+0.0%
+0.0%
0/1/0
= Read a diagnostic log to find the stale input
Eligible
+0.0%
+0.0%
0/1/0
Illustrative judge evidence:
Diagnose custom targets that always rerun: Position-swap inconsistent (forward: B, reverse: tie). Defaulting to tie.
This is one example, not the aggregate verdict. Open Full Results for every judgment.
➖ Not proven improved — item-management (claude-sonnet-4.6)
Why: Net win +0.0% (1W/4T/1L over 6 preference-eligible stimulus vote(s), sign test p=0.750), mean preference -8.6% across 7 paired run(s), 1 dormancy stimulus/stimuli excluded from preference — no improvement
Next action: Inspect tied or lost stimuli; predeclare added breadth before a new experiment.
Repeated-run reliability (not used by the gate): 7 paired runs (1W/5T/1L).
Weak or warning scenarios:
Scenario
Preference gate
Net win
Δ Pref
Runs (W/T/L)
= Decline a target ordering problem with no item-group defect
Excluded (activation contract)
+0.0%
+0.0%
0/1/0
▼ Diagnose an ineffective Compile Remove that does not match the glob
Eligible
-100.0%
-100.0%
0/0/1
= Diagnose real and claimed item problems in a code generation pipeline
Eligible
+0.0%
+0.0%
0/1/0
= Fix item management anti-patterns
Eligible
+0.0%
+0.0%
0/1/0
= Leave already-correct item management unchanged
Eligible
+0.0%
+0.0%
0/1/0
= Leave correct single-list batching unchanged
Eligible
+0.0%
+0.0%
0/1/0
Illustrative judge evidence:
Diagnose an ineffective Compile Remove that does not match the glob: A directly and accurately answers the actual cause and fix. B supplies a misleading, technically incorrect separator-based root cause, failing the core requested diagnosis despite ending with a usable-looking glob.
This is one example, not the aggregate verdict. Open Full Results for every judgment.
➖ Not proven improved — item-management (gpt-5.6-luna)
Why: Net win +16.7% (2W/3T/1L over 6 preference-eligible stimulus vote(s), sign test p=0.500), mean preference +5.7% across 7 paired run(s), 1 dormancy stimulus/stimuli excluded from preference — not credible — 3 of 6 preference-eligible stimulus vote(s) tied, leaving only 3 discordant preference vote(s). The sign test conditions on non-tie stimulus votes and cannot reach 0.05 below 5, so no record could have passed here — this is not a measured null. Either the skill is inert on these scenarios (make them discriminate) or the eval needs more distinct stimuli to clear the ties
Next action: Inspect tied or lost stimuli; predeclare added breadth before a new experiment.
Repeated-run reliability (not used by the gate): 7 paired runs (2W/4T/1L).
Weak or warning scenarios:
Scenario
Preference gate
Net win
Δ Pref
Runs (W/T/L)
= Decline a target ordering problem with no item-group defect
Excluded (activation contract)
+0.0%
+0.0%
0/1/0
▼ Diagnose an ineffective Compile Remove that does not match the glob
Eligible
-100.0%
-40.0%
0/0/1
= Fix item management anti-patterns
Eligible
+0.0%
+0.0%
0/1/0
= Leave already-correct item management unchanged
Eligible
+0.0%
+0.0%
0/1/0
= Leave correct single-list batching unchanged
Eligible
+0.0%
+0.0%
0/1/0
Illustrative judge evidence:
Diagnose an ineffective Compile Remove that does not match the glob: Both reach the same correct diagnosis and avoid recommending Exclude. However, Response A actually applied the fix to the file (recovering from the failed apply_patch by using sed) and verified via dotnet msbuild -getItem:Compile that Api.g.cs is now excluded while Service.c...
This is one example, not the aggregate verdict. Open Full Results for every judgment.
➖ Not proven improved — msbuild-antipatterns (claude-sonnet-4.6)
Why: Net win +0.0% (1W/5T/1L over 7 preference-eligible stimulus vote(s), sign test p=0.750), mean preference +0.0% across 8 paired run(s), 1 dormancy stimulus/stimuli excluded from preference — no improvement
Next action: Inspect tied or lost stimuli; predeclare added breadth before a new experiment.
Repeated-run reliability (not used by the gate): 8 paired runs (1W/6T/1L).
Weak or warning scenarios:
Scenario
Preference gate
Net win
Δ Pref
Runs (W/T/L)
= Add a module to an F# project
Eligible
+0.0%
+0.0%
0/1/0
= Add a signature file to define public API
Eligible
+0.0%
+0.0%
0/1/0
= Distinguish a style backslash from a real cross-platform backslash bug
Eligible
+0.0%
+0.0%
0/1/0
= Fix broken file order causing FS0039
Eligible
+0.0%
+0.0%
0/1/0
▼ Judge an unguarded import inside a NuGet package build folder
Eligible
-100.0%
-40.0%
0/0/1
= Non-activation: migrate a legacy project to SDK style
Excluded (activation contract)
+0.0%
+0.0%
0/1/0
= Review MSBuild files for anti-patterns and style issues
Eligible
+0.0%
+0.0%
0/1/0
Illustrative judge evidence:
Add a module to an F# project: The implementations and verification outcomes are materially identical, fully addressing the requested validation, compilation ordering, and guarded processing behavior.
This is one example, not the aggregate verdict. Open Full Results for every judgment.
➖ Not proven improved — msbuild-antipatterns (gpt-5.6-luna)
Why: Net win -14.3% (0W/6T/1L over 7 preference-eligible stimulus vote(s), sign test p=0.500), mean preference -10.0% across 8 paired run(s), 1 dormancy stimulus/stimuli excluded from preference — no improvement
Next action: Inspect tied or lost stimuli; predeclare added breadth before a new experiment.
Repeated-run reliability (not used by the gate): 8 paired runs (0W/6T/2L).
Weak or warning scenarios:
Scenario
Preference gate
Net win
Δ Pref
Runs (W/T/L)
= Add a module to an F# project
Eligible
+0.0%
+0.0%
0/1/0
= Add a signature file to define public API
Eligible
+0.0%
+0.0%
0/1/0
= Distinguish a style backslash from a real cross-platform backslash bug
Eligible
+0.0%
+0.0%
0/1/0
= Fix broken file order causing FS0039
Eligible
+0.0%
+0.0%
0/1/0
▼ Judge an unguarded import inside a NuGet package build folder
Eligible
-100.0%
-40.0%
0/0/1
= Leave a clean project without inventing anti-patterns
Eligible
+0.0%
+0.0%
0/1/0
▼ Non-activation: migrate a legacy project to SDK style
Excluded (activation contract)
-100.0%
-40.0%
0/0/1
= Review MSBuild files for anti-patterns and style issues
Eligible
+0.0%
+0.0%
0/1/0
Illustrative judge evidence:
Add a module to an F# project: Position-swap inconsistent (forward: tie, reverse: B). Defaulting to tie.
This is one example, not the aggregate verdict. Open Full Results for every judgment.
➖ Not proven improved — msbuild-modernization (claude-sonnet-4.6)
Why: Net win +0.0% (1W/4T/1L over 6 preference-eligible stimulus vote(s), sign test p=0.750), mean preference +22.9% across 7 paired run(s), 1 dormancy stimulus/stimuli excluded from preference — no improvement
Next action: Inspect tied or lost stimuli; predeclare added breadth before a new experiment.
Repeated-run reliability (not used by the gate): 7 paired runs (2W/4T/1L).
Weak or warning scenarios:
Scenario
Preference gate
Net win
Δ Pref
Runs (W/T/L)
▼ Apply the migration to SDK-style
Eligible
-100.0%
-40.0%
0/0/1
= Consolidate duplicated projects into a multi-targeting SDK-style project
Eligible
+0.0%
+0.0%
0/1/0
= Identify legacy patterns for SDK-style migration
Eligible
+0.0%
+0.0%
0/1/0
= Modernize a single project without introducing Central Package Management
Eligible
+0.0%
+0.0%
0/1/0
= Recognize an already-modern project needs no migration
Eligible
+0.0%
+0.0%
0/1/0
Illustrative judge evidence:
Apply the migration to SDK-style: Both conversions satisfy the core SDK-style migration requirements. A is marginally safer for fidelity because it retains all preexisting assembly metadata directly rather than requiring every attribute to be accurately translated before deleting AssemblyInfo.cs.
This is one example, not the aggregate verdict. Open Full Results for every judgment.
➖ Not proven improved — msbuild-modernization (gpt-5.6-luna)
Why: Net win +33.3% (2W/4T/0L over 6 preference-eligible stimulus vote(s), sign test p=0.250), mean preference +11.4% across 7 paired run(s), 1 dormancy stimulus/stimuli excluded from preference — not credible — 4 of 6 preference-eligible stimulus vote(s) tied, leaving only 2 discordant preference vote(s). The sign test conditions on non-tie stimulus votes and cannot reach 0.05 below 5, so no record could have passed here — this is not a measured null. Either the skill is inert on these scenarios (make them discriminate) or the eval needs more distinct stimuli to clear the ties
Next action: Inspect tied or lost stimuli; predeclare added breadth before a new experiment.
Repeated-run reliability (not used by the gate): 7 paired runs (2W/5T/0L).
Weak or warning scenarios:
Scenario
Preference gate
Net win
Δ Pref
Runs (W/T/L)
= Consolidate duplicated projects into a multi-targeting SDK-style project
Eligible
+0.0%
+0.0%
0/1/0
= Decline modernizing a non-.NET build
Excluded (activation contract)
+0.0%
+0.0%
0/1/0
= Identify legacy patterns for SDK-style migration
Eligible
+0.0%
+0.0%
0/1/0
= Modernize a single project without introducing Central Package Management
Eligible
+0.0%
+0.0%
0/1/0
= Recognize an already-modern project needs no migration
Eligible
+0.0%
+0.0%
0/1/0
Illustrative judge evidence:
Consolidate duplicated projects into a multi-targeting SDK-style project: Position-swap inconsistent (forward: tie, reverse: A). Defaulting to tie.
This is one example, not the aggregate verdict. Open Full Results for every judgment.
➖ Not proven improved — msbuild-server (gpt-5.6-luna)
Why: Net win +62.5% (6W/1T/1L over 8 preference-eligible stimulus vote(s), sign test p=0.063), mean preference +40.0% across 8 paired run(s) — not credible (sign test p=0.063 > 0.05)
Next action: Inspect tied or lost stimuli and fix inconsistent skill behavior.
Gate evidence: n=8; 6W/1T/1L; d=7; p=0.063; net +62.5%
Overfit: Low (score 0.19)
Repeated-run reliability (not used by the gate): 8 paired runs (6W/1T/1L).
Weak or warning scenarios:
Scenario
Preference gate
Net win
Δ Pref
Runs (W/T/L)
▼ Confirm the MSBuild Server is actually improving build times before declaring success
Eligible
-100.0%
-40.0%
0/0/1
= Decline MSBuild Server for a single one-off release build
Eligible
+0.0%
+0.0%
0/1/0
Illustrative judge evidence:
Confirm the MSBuild Server is actually improving build times before declaring success: Both responses give solid, empirically-grounded advice with cold/warm comparison and A/B testing. A adds value by investigating the actual project (discovering it's a tiny LoopApp), warning that the trivial project gives no meaningful signal, and providing a diagnostic log gre...
This is one example, not the aggregate verdict. Open Full Results for every judgment.
➖ Not proven improved — property-patterns (gpt-5.6-luna)
Why: Net win +33.3% (3W/2T/1L over 6 preference-eligible stimulus vote(s), sign test p=0.312), mean preference +11.4% across 7 paired run(s), 1 dormancy stimulus/stimuli excluded from preference — not credible — 2 of 6 preference-eligible stimulus vote(s) tied, leaving only 4 discordant preference vote(s). The sign test conditions on non-tie stimulus votes and cannot reach 0.05 below 5, so no record could have passed here — this is not a measured null. Either the skill is inert on these scenarios (make them discriminate) or the eval needs more distinct stimuli to clear the ties
Next action: Inspect tied or lost stimuli; predeclare added breadth before a new experiment.
Repeated-run reliability (not used by the gate): 7 paired runs (3W/3T/1L).
Weak or warning scenarios:
Scenario
Preference gate
Net win
Δ Pref
Runs (W/T/L)
= Decline a props-versus-targets placement question
Excluded (activation contract)
+0.0%
+0.0%
0/1/0
= Diagnose a TargetFramework condition that never applies in props
Eligible
+0.0%
+0.0%
0/1/0
= Diagnose shared build property issues
Eligible
+0.0%
+0.0%
0/1/0
▼ Fix shared property configuration
Eligible
-100.0%
-40.0%
0/0/1
Illustrative judge evidence:
Diagnose a TargetFramework condition that never applies in props: Position-swap inconsistent (forward: B, reverse: A). Defaulting to tie.
This is one example, not the aggregate verdict. Open Full Results for every judgment.
➖ Not proven improved — resolve-project-references (claude-sonnet-4.6)
Why: Net win +66.7% (5W/0T/1L over 6 preference-eligible stimulus vote(s), sign test p=0.109), mean preference +37.5% across 8 paired run(s), 2 dormancy stimulus/stimuli excluded from preference — not credible (sign test p=0.109 > 0.05)
Next action: Inspect tied or lost stimuli and fix inconsistent skill behavior.
Repeated-run reliability (not used by the gate): 8 paired runs (7W/0T/1L).
Weak or warning scenarios:
Scenario
Preference gate
Net win
Δ Pref
Runs (W/T/L)
▼ Distinguish wait time from a real serial dependency chain
Eligible
-100.0%
-40.0%
0/0/1
Illustrative judge evidence:
Distinguish wait time from a real serial dependency chain: Both are strong and satisfy every requested point. A is marginally more technically precise and avoids conflating the inflated target accounting with a direct bottleneck, while still clearly identifying the serial graph and remedy.
This is one example, not the aggregate verdict. Open Full Results for every judgment.
➖ Not proven improved — resolve-project-references (gpt-5.6-luna)
Why: Net win +16.7% (2W/3T/1L over 6 preference-eligible stimulus vote(s), sign test p=0.500), mean preference +17.5% across 8 paired run(s), 2 dormancy stimulus/stimuli excluded from preference — not credible — 3 of 6 preference-eligible stimulus vote(s) tied, leaving only 3 discordant preference vote(s). The sign test conditions on non-tie stimulus votes and cannot reach 0.05 below 5, so no record could have passed here — this is not a measured null. Either the skill is inert on these scenarios (make them discriminate) or the eval needs more distinct stimuli to clear the ties
Next action: Inspect tied or lost stimuli; predeclare added breadth before a new experiment.
To investigate non-passing or warning results, paste this to your AI coding agent:
For PR 1118 in dotnet/skills, download eval artifacts with gh run download 34291893629 --repo dotnet/skills --pattern "vally-results-*" --dir ./eval-results, then fetch https://raw.githubusercontent.com/dotnet/skills/8bc2f2078f7640a8ab78eff8e439d1370c816923/eng/vally-adapter/InvestigatingResults.md and follow it. Classify each result as measurement-invalid, underpowered, not-proven, preference-loss, or passing-with-warning. Use stateReason, result accounting, weak scenarios, and judge evidence to give the cause and exact next fix.
Remove off-target vocabulary from the three descriptions that Sonnet selected before any workspace inspection, while retaining full boundary guidance in the skill bodies.
Co-authored-by: Copilot App <223556219+Copilot@users.noreply.github.com>
Copilot-Session: 2d0a0a20-fd83-4343-8434-4763f8120503
❌ Evaluation did not complete successfully (the evaluate job reported failure). Check the workflow run logs, then comment /evaluate a37f8fc84c157757eb17ba8a2ee79a98799bb321 to retry this exact commit.
70 partial result file(s) were preserved for diagnosis but were not consolidated because the full matrix did not complete.
👋 @AbhitejJohn — this PR has merge conflict. When you're ready, please address the feedback and push an update; the triage bot will pick up the next state automatically. (Add the no-stale label to silence further pings.)
Bring in the Markdown linter startup fix before the final exact-head evaluation.
Co-authored-by: Copilot App <223556219+Copilot@users.noreply.github.com>
Copilot-Session: eb11c85b-4046-4dbc-94b3-3bdbd4839a0d
Read executionShard only from suite-level tags and require isolated-arm evidence for dormancy contracts.
Co-authored-by: Copilot App <223556219+Copilot@users.noreply.github.com>
Copilot-Session: eb11c85b-4046-4dbc-94b3-3bdbd4839a0d
❌ Evaluation did not complete successfully (the evaluate job reported failure). Check the workflow run logs, then comment /evaluate 69b62d37c8de23d24b1dcb440785ff4aa2e28bef to retry this exact commit.
44 partial result file(s) were preserved for diagnosis but were not consolidated because the full matrix did not complete.
Bring in the preference-based skill value dashboard before final evaluation.
Co-authored-by: Copilot App <223556219+Copilot@users.noreply.github.com>
Copilot-Session: eb11c85b-4046-4dbc-94b3-3bdbd4839a0d
Apply the standard manifest allowlist to Codex and enforce exact parity in validator tests.
Co-authored-by: Copilot App <223556219+Copilot@users.noreply.github.com>
Copilot-Session: eb11c85b-4046-4dbc-94b3-3bdbd4839a0d
Bring in Claude manifest version synchronization fixes before final evaluation.
Co-authored-by: Copilot App <223556219+Copilot@users.noreply.github.com>
Copilot-Session: eb11c85b-4046-4dbc-94b3-3bdbd4839a0d
The new explicit allowlist removes binlog_overview. The repository's Codex smoke contract calls that tool directly (CONTRIBUTING.md:147-150), so the smoke call will no longer be exposed through this MCP server after this change. Keep binlog_overview in the allowlist, or update the smoke contract and every host manifest together.
Keep binlog_overview in the explicit allowlist
plugins/dotnet-msbuild/plugin.json:25
The new explicit allowlist removes binlog_overview. The repository's Codex smoke contract calls that tool directly (CONTRIBUTING.md:147-150), so the smoke call will no longer be exposed through this MCP server after this change. Keep binlog_overview in the allowlist, or update the smoke contract and every host manifest together.
Exercise current allowlisted load and diagnostics operations instead of the removed binlog_overview tool.
Co-authored-by: Copilot App <223556219+Copilot@users.noreply.github.com>
Copilot-Session: eb11c85b-4046-4dbc-94b3-3bdbd4839a0d
Use the live read-only tool inventory across all hosts and restore the credential-free overview smoke.
Co-authored-by: Copilot App <223556219+Copilot@users.noreply.github.com>
Copilot-Session: eb11c85b-4046-4dbc-94b3-3bdbd4839a0d
❌ Evaluation did not complete successfully (the evaluate job reported failure). Check the workflow run logs, then comment /evaluate 5a0ac39c6d4c1fd3cf30290b0075dab55baeb381 to retry this exact commit.
53 partial result file(s) were preserved for diagnosis but were not consolidated because the full matrix did not complete.
⚠️ Session replay telemetry was not published because the auxiliary dotnet/skills-data publisher failed. The evaluation verdicts above remain authoritative; maintainers must repair SKILLS_DATA_TOKEN or the publisher before replay links are available.
❌ Evaluation did not complete successfully (the evaluate job reported failure). Check the workflow run logs, then comment /evaluate 417727600696a84e7b6846f0a13a4da4a7b8f195 to retry this exact commit.
69 partial result file(s) were preserved for diagnosis but were not consolidated because the full matrix did not complete.
⚠️ Session replay telemetry was not published because the auxiliary dotnet/skills-data publisher failed. The evaluation verdicts above remain authoritative; maintainers must repair SKILLS_DATA_TOKEN or the publisher before replay links are available.
Objective completion gate: not enabled. Aggregate completion transitions are telemetry only, so this report does not claim that zero objective regressions were proven.
A result passes only when preference-eligible distinct-stimulus votes have aggregate net win of at least 20% and an exact one-sided sign-test result of p ≤ 0.05, and every explicit dormancy activation contract passes. Repeated runs measure reliability only.
Target
Model
Verdict
Gate evidence
Overfit
Warnings
Next action
agent.code-testing-generator
claude-sonnet-5
➖ Not proven improved
n=5; 3W/1T/1L; d=4; p=0.312; net +40.0%
—
Activation: isolated 0/5; plugin 2/5
Inspect tied or lost stimuli; predeclare added breadth before a new experiment.
agent.code-testing-generator
gpt-5.6-luna
➖ Not proven improved
n=5; 3W/1T/1L; d=4; p=0.312; net +40.0%
—
—
Inspect tied or lost stimuli; predeclare added breadth before a new experiment.
agent.msbuild
claude-sonnet-5
➖ Not proven improved
n=5; 1W/3T/1L; d=2; p=0.750; net +0.0%
—
Activation: isolated 2/5; plugin 0/5
Inspect tied or lost stimuli; predeclare added breadth before a new experiment.
agent.msbuild
gpt-5.6-luna
➖ Not proven improved
n=5; 1W/3T/1L; d=2; p=0.750; net +0.0%
—
Activation: isolated 2/5; plugin 1/5
Inspect tied or lost stimuli; predeclare added breadth before a new experiment.
agent.test-quality-auditor
claude-sonnet-5
➖ Not proven improved
n=5; 5W/0T/0L; d=5; p=0.031; net +100.0%; 1 dormancy excluded
—
Activation: isolated 0/5; plugin 0/5
Inspect tied or lost stimuli and fix inconsistent skill behavior.
agent.test-quality-auditor
gpt-5.6-luna
➖ Not proven improved
n=5; 1W/3T/1L; d=2; p=0.750; net +0.0%; 1 dormancy excluded
—
Activation: isolated 1/5; plugin 1/5
Inspect tied or lost stimuli; predeclare added breadth before a new experiment.
agent.testability-migration
claude-sonnet-5
➖ Not proven improved
n=5; 3W/2T/0L; d=3; p=0.125; net +60.0%
—
Activation: isolated 0/5; plugin 0/5
Inspect tied or lost stimuli; predeclare added breadth before a new experiment.
agent.testability-migration
gpt-5.6-luna
➖ Not proven improved
n=5; 3W/1T/1L; d=4; p=0.312; net +40.0%
—
Activation: isolated 0/5; plugin 0/5
Inspect tied or lost stimuli; predeclare added breadth before a new experiment.
binlog-failure-analysis
claude-sonnet-5
➖ Not proven improved
n=6; 0W/2T/4L; d=4; p=0.063; net -66.7%; 1 dormancy excluded
🟡 0.40
Activation: isolated 6/7; plugin 6/7
Inspect tied or lost stimuli; predeclare added breadth before a new experiment.
binlog-failure-analysis
gpt-5.6-luna
➖ Not proven improved
n=6; 2W/4T/0L; d=2; p=0.250; net +33.3%; 1 dormancy excluded
🟡 0.22
Activation: isolated 6/7; plugin 6/7
Inspect tied or lost stimuli; predeclare added breadth before a new experiment.
binlog-generation
claude-sonnet-5
➖ Not proven improved
n=7; 4W/2T/1L; d=5; p=0.188; net +42.9%; 1 dormancy excluded
🟡 0.45
Activation: isolated 5/7; plugin 6/7
Inspect tied or lost stimuli and fix inconsistent skill behavior.
binlog-generation
gpt-5.6-luna
➖ Not proven improved
n=7; 4W/2T/1L; d=5; p=0.188; net +42.9%; 1 dormancy excluded
Inspect activation-only failed runs before rewriting skill content; the model stopped after loading a skill.
writing-mstest-tests
gpt-5.6-luna
➖ Not proven improved
n=15; 8W/4T/3L; d=11; p=0.113; net +33.3%
✅ 0.00
Activation-only stop: isolated 3 failed runs
Inspect activation-only failed runs before rewriting skill content; the model stopped after loading a skill.
ℹ️ How to read this report
✅ Improved — the result passed both the statistical gate and the 20% practical net-win floor.
➖ Not proven improved — the result is valid but did not pass both gates. This is not automatically a regression.
⚠️ Invalid / underpowered — the gate withheld a quality verdict. Fix the measurement before judging the target.
⛔ Activation contract failed — the isolated target activated on an explicit dormancy scenario. Dormancy preference is excluded, but this routing failure still blocks a pass.
📉 Preference loss — the LLM judge credibly preferred baseline. It is report-only, not objective completion proof.
Gate evidence — n preference-eligible distinct-stimulus votes, W/T/L stimulus votes, d discordant votes, exact one-sided p, net win, and the count of separately retained dormancy stimuli. The p value applies to one model/target result; no matrix-wide multiple-comparison correction is applied.
Warnings — activation, timeout, retry recovery, or unresolved comparison conditions that need attention.
Do not add repeated runs to increase statistical power. Do not add stimuli after seeing a near-pass unless the new breadth is predeclared for a new experiment.
Repeated-run reliability (not used by the gate): 9 paired runs (6W/1T/2L).
Weak or warning scenarios:
Scenario
Preference gate
Net win
Δ Pref
Runs (W/T/L)
▼ Audit a pytest suite using Python-specific anti-pattern markers
Eligible
-100.0%
-40.0%
0/0/1
= Detect coverage-touching pattern across a service facade
Eligible
+0.0%
+0.0%
0/1/0
▼ Stay dormant for test trait distribution
Excluded (activation contract)
-100.0%
-40.0%
0/0/1
Illustrative judge evidence:
Audit a pytest suite using Python-specific anti-pattern markers: Response A is more complete and directly addresses the task's core ask: ranking findings by how badly they hurt. It identifies the pytest dependency issue (a tier-1 blocker preventing test execution), applies explicit severity ranking, and accounts for all 9 distinct problems ...
This is one example, not the aggregate verdict. Open Full Results for every judgment.
➖ Not proven improved — agent.code-testing-generator (claude-sonnet-5)
Why: Net win +40.0% (3W/1T/1L over 5 preference-eligible stimulus vote(s), sign test p=0.312), mean preference +16.0% across 5 paired run(s) — not credible — 1 of 5 preference-eligible stimulus vote(s) tied, leaving only 4 discordant preference vote(s). The sign test conditions on non-tie stimulus votes and cannot reach 0.05 below 5, so no record could have passed here — this is not a measured null. Either the agent is inert on these scenarios (make them discriminate) or the eval needs more distinct stimuli to clear the ties — native evaluator reported that the target agent did not activate
Next action: Inspect tied or lost stimuli; predeclare added breadth before a new experiment.
Gate evidence: n=5; 3W/1T/1L; d=4; p=0.312; net +40.0%
Warnings: Activation: isolated 0/5; plugin 2/5
Repeated-run reliability (not used by the gate): 5 paired runs (3W/1T/1L).
Weak or warning scenarios:
Scenario
Preference gate
Net win
Δ Pref
Runs (W/T/L)
▲ Generate a project-wide pytest suite across modules
Eligible
+100.0%
+40.0%
1/0/0
▲ Generate collaborating Go package tests
Eligible
+100.0%
+40.0%
1/0/0
= Generate layered Vitest coverage for an async cart
Eligible
+0.0%
+0.0%
0/1/0
▼ Generate project-wide xUnit tests for a .NET library
Eligible
-100.0%
-40.0%
0/0/1
▲ Preserve a classic MSTest project while adding broad coverage
Eligible
+100.0%
+40.0%
1/0/0
Illustrative judge evidence:
Generate layered Vitest coverage for an async cart: Position-swap inconsistent (forward: skill, reverse: baseline). Defaulting to tie.
This is one example, not the aggregate verdict. Open Full Results for every judgment.
➖ Not proven improved — agent.code-testing-generator (gpt-5.6-luna)
Why: Net win +40.0% (3W/1T/1L over 5 preference-eligible stimulus vote(s), sign test p=0.312), mean preference +16.0% across 5 paired run(s) — not credible — 1 of 5 preference-eligible stimulus vote(s) tied, leaving only 4 discordant preference vote(s). The sign test conditions on non-tie stimulus votes and cannot reach 0.05 below 5, so no record could have passed here — this is not a measured null. Either the agent is inert on these scenarios (make them discriminate) or the eval needs more distinct stimuli to clear the ties
Next action: Inspect tied or lost stimuli; predeclare added breadth before a new experiment.
Gate evidence: n=5; 3W/1T/1L; d=4; p=0.312; net +40.0%
Repeated-run reliability (not used by the gate): 5 paired runs (3W/1T/1L).
Weak or warning scenarios:
Scenario
Preference gate
Net win
Δ Pref
Runs (W/T/L)
= Generate collaborating Go package tests
Eligible
+0.0%
+0.0%
0/1/0
▼ Preserve a classic MSTest project while adding broad coverage
Eligible
-100.0%
-40.0%
0/0/1
Illustrative judge evidence:
Generate collaborating Go package tests: Position-swap inconsistent (forward: tie, reverse: baseline). Defaulting to tie.
This is one example, not the aggregate verdict. Open Full Results for every judgment.
➖ Not proven improved — agent.msbuild (claude-sonnet-5)
Why: Net win +0.0% (1W/3T/1L over 5 preference-eligible stimulus vote(s), sign test p=0.750), mean preference +0.0% across 5 paired run(s) — no improvement — native evaluator reported that the target agent did not activate
Next action: Inspect tied or lost stimuli; predeclare added breadth before a new experiment.
Gate evidence: n=5; 1W/3T/1L; d=2; p=0.750; net +0.0%
Warnings: Activation: isolated 2/5; plugin 0/5
Repeated-run reliability (not used by the gate): 5 paired runs (1W/3T/1L).
Weak or warning scenarios:
Scenario
Preference gate
Net win
Δ Pref
Runs (W/T/L)
= Advise on project file organization
Eligible
+0.0%
+0.0%
0/1/0
= Diagnose broken incremental build behavior
Eligible
+0.0%
+0.0%
0/1/0
▲ Review a project file for maintainability risks
Eligible
+100.0%
+40.0%
1/0/0
▼ Route a slow build to performance analysis
Eligible
-100.0%
-40.0%
0/0/1
= Triage a build failure and route to appropriate analysis
Eligible
+0.0%
+0.0%
0/1/0
Illustrative judge evidence:
Advise on project file organization: Position-swap inconsistent (forward: baseline, reverse: skill). Defaulting to tie.
This is one example, not the aggregate verdict. Open Full Results for every judgment.
➖ Not proven improved — agent.msbuild (gpt-5.6-luna)
Why: Net win +0.0% (1W/3T/1L over 5 preference-eligible stimulus vote(s), sign test p=0.750), mean preference -12.0% across 5 paired run(s) — no improvement — native evaluator reported that the target agent did not activate
Next action: Inspect tied or lost stimuli; predeclare added breadth before a new experiment.
Gate evidence: n=5; 1W/3T/1L; d=2; p=0.750; net +0.0%
Warnings: Activation: isolated 2/5; plugin 1/5
Repeated-run reliability (not used by the gate): 5 paired runs (1W/3T/1L).
Weak or warning scenarios:
Scenario
Preference gate
Net win
Δ Pref
Runs (W/T/L)
= Advise on project file organization
Eligible
+0.0%
+0.0%
0/1/0
▲ Diagnose broken incremental build behavior
Eligible
+100.0%
+40.0%
1/0/0
▼ Review a project file for maintainability risks
Eligible
-100.0%
-100.0%
0/0/1
= Route a slow build to performance analysis
Eligible
+0.0%
+0.0%
0/1/0
= Triage a build failure and route to appropriate analysis
Eligible
+0.0%
+0.0%
0/1/0
Illustrative judge evidence:
Advise on project file organization: Position-swap inconsistent (forward: tie, reverse: baseline). Defaulting to tie.
This is one example, not the aggregate verdict. Open Full Results for every judgment.
➖ Not proven improved — agent.test-quality-auditor (claude-sonnet-5)
Why: Net win +100.0% (5W/0T/0L over 5 preference-eligible stimulus vote(s), sign test p=0.031), mean preference +26.7% across 6 paired run(s), 1 dormancy stimulus/stimuli excluded from preference — credibly better — native evaluator reported that the target agent did not activate
Next action: Inspect tied or lost stimuli and fix inconsistent skill behavior.
Repeated-run reliability (not used by the gate): 6 paired runs (5W/0T/1L).
Weak or warning scenarios:
Scenario
Preference gate
Net win
Δ Pref
Runs (W/T/L)
▲ Assertion quality analysis
Eligible
+100.0%
+40.0%
1/0/0
▲ Comprehensive test quality audit of weak test suite
Eligible
+100.0%
+40.0%
1/0/0
▼ Decline request to generate new tests
Excluded (activation contract)
-100.0%
-100.0%
0/0/1
▲ Diagnose test smells and propose a repair order
Eligible
+100.0%
+40.0%
1/0/0
▲ Identify behavior gaps that existing tests would miss
Eligible
+100.0%
+40.0%
1/0/0
▲ Targeted anti-pattern review
Eligible
+100.0%
+100.0%
1/0/0
Illustrative judge evidence:
Decline request to generate new tests: A demonstrated substantially better contextual judgment by preserving the audit fixture while still producing and validating a separate comprehensive suite. B's tests may be reasonable, but modifying the deliberately weak fixture and its csproj is inappropriate for this reposi...
This is one example, not the aggregate verdict. Open Full Results for every judgment.
➖ Not proven improved — agent.test-quality-auditor (gpt-5.6-luna)
Why: Net win +0.0% (1W/3T/1L over 5 preference-eligible stimulus vote(s), sign test p=0.750), mean preference +10.0% across 6 paired run(s), 1 dormancy stimulus/stimuli excluded from preference — no improvement — native evaluator reported that the target agent did not activate
Next action: Inspect tied or lost stimuli; predeclare added breadth before a new experiment.
This is one example, not the aggregate verdict. Open Full Results for every judgment.
➖ Not proven improved — agent.testability-migration (claude-sonnet-5)
Why: Net win +60.0% (3W/2T/0L over 5 preference-eligible stimulus vote(s), sign test p=0.125), mean preference +24.0% across 5 paired run(s) — not credible — 2 of 5 preference-eligible stimulus vote(s) tied, leaving only 3 discordant preference vote(s). The sign test conditions on non-tie stimulus votes and cannot reach 0.05 below 5, so no record could have passed here — this is not a measured null. Either the agent is inert on these scenarios (make them discriminate) or the eval needs more distinct stimuli to clear the ties — native evaluator reported that the target agent did not activate
Next action: Inspect tied or lost stimuli; predeclare added breadth before a new experiment.
Gate evidence: n=5; 3W/2T/0L; d=3; p=0.125; net +60.0%
Warnings: Activation: isolated 0/5; plugin 0/5
Repeated-run reliability (not used by the gate): 5 paired runs (3W/2T/0L).
Weak or warning scenarios:
Scenario
Preference gate
Net win
Δ Pref
Runs (W/T/L)
▲ Full pipeline: detect statics and recommend migration plan
Eligible
+100.0%
+40.0%
1/0/0
▲ Inventory static dependencies without modifying the project
Eligible
+100.0%
+40.0%
1/0/0
= Migrate time dependencies and add deterministic tests
Eligible
+0.0%
+0.0%
0/1/0
= Replace filesystem statics without touching unrelated dependencies
Eligible
+0.0%
+0.0%
0/1/0
▲ Targeted request: just migrate DateTime to TimeProvider
Eligible
+100.0%
+40.0%
1/0/0
Illustrative judge evidence:
Migrate time dependencies and add deterministic tests: Position-swap inconsistent (forward: skill, reverse: baseline). Defaulting to tie.
This is one example, not the aggregate verdict. Open Full Results for every judgment.
➖ Not proven improved — agent.testability-migration (gpt-5.6-luna)
Why: Net win +40.0% (3W/1T/1L over 5 preference-eligible stimulus vote(s), sign test p=0.312), mean preference +28.0% across 5 paired run(s) — not credible — 1 of 5 preference-eligible stimulus vote(s) tied, leaving only 4 discordant preference vote(s). The sign test conditions on non-tie stimulus votes and cannot reach 0.05 below 5, so no record could have passed here — this is not a measured null. Either the agent is inert on these scenarios (make them discriminate) or the eval needs more distinct stimuli to clear the ties — native evaluator reported that the target agent did not activate
Next action: Inspect tied or lost stimuli; predeclare added breadth before a new experiment.
Gate evidence: n=5; 3W/1T/1L; d=4; p=0.312; net +40.0%
Warnings: Activation: isolated 0/5; plugin 0/5
Repeated-run reliability (not used by the gate): 5 paired runs (3W/1T/1L).
Weak or warning scenarios:
Scenario
Preference gate
Net win
Δ Pref
Runs (W/T/L)
▼ Full pipeline: detect statics and recommend migration plan
Eligible
-100.0%
-40.0%
0/0/1
= Inventory static dependencies without modifying the project
Eligible
+0.0%
+0.0%
0/1/0
▲ Migrate time dependencies and add deterministic tests
Eligible
+100.0%
+100.0%
1/0/0
▲ Replace filesystem statics without touching unrelated dependencies
Eligible
+100.0%
+40.0%
1/0/0
▲ Targeted request: just migrate DateTime to TimeProvider
Eligible
+100.0%
+40.0%
1/0/0
Illustrative judge evidence:
Full pipeline: detect statics and recommend migration plan: While Response B excels at structured analysis with precise line numbers and recommends pragmatic standard libraries (TimeProvider, System.IO.Abstractions), Response A better serves the specific request for a 'bounded' migration plan. Response A identifies the dual-read clock ...
This is one example, not the aggregate verdict. Open Full Results for every judgment.
➖ Not proven improved — binlog-failure-analysis (claude-sonnet-5)
Why: Net win -66.7% (0W/2T/4L over 6 preference-eligible stimulus vote(s), sign test p=0.063), mean preference -22.9% across 7 paired run(s), 1 dormancy stimulus/stimuli excluded from preference — no improvement
Next action: Inspect tied or lost stimuli; predeclare added breadth before a new experiment.
Repeated-run reliability (not used by the gate): 7 paired runs (0W/3T/4L).
Weak or warning scenarios:
Scenario
Preference gate
Net win
Δ Pref
Runs (W/T/L)
▼ Assess a requested binlog investigation when the current build is healthy
Eligible
-100.0%
-40.0%
0/0/1
= Confirm the actual resolved target framework and package version from a binlog
Eligible
+0.0%
+0.0%
0/0/0
▼ Determine whether a quiet second build actually failed
Eligible
-100.0%
-40.0%
0/0/1
= Diagnose a warning behind a build that actually succeeded
Eligible
+0.0%
+0.0%
0/1/0
= Diagnose build failures from binlog only (no source files)
Eligible
+0.0%
+0.0%
0/1/0
= Stay dormant for a non-MSBuild build failure log
Excluded (activation contract)
+0.0%
+0.0%
0/1/0
▼ Trace why a generated source file is missing at compile time
Eligible
-100.0%
-40.0%
0/0/1
▼ Use capture-time text logs when the binlog reader is unavailable
Eligible
-100.0%
-40.0%
0/0/1
Illustrative judge evidence:
Assess a requested binlog investigation when the current build is healthy: Both are correct and satisfy the central request by capturing a binlog and establishing that the currently staged project is healthy. A is modestly stronger because its final answer gives a more substantive binlog-based analysis, including separate error/warning checks and tim...
This is one example, not the aggregate verdict. Open Full Results for every judgment.
➖ Not proven improved — binlog-failure-analysis (gpt-5.6-luna)
Why: Net win +33.3% (2W/4T/0L over 6 preference-eligible stimulus vote(s), sign test p=0.250), mean preference +11.4% across 7 paired run(s), 1 dormancy stimulus/stimuli excluded from preference — not credible — 4 of 6 preference-eligible stimulus vote(s) tied, leaving only 2 discordant preference vote(s). The sign test conditions on non-tie stimulus votes and cannot reach 0.05 below 5, so no record could have passed here — this is not a measured null. Either the skill is inert on these scenarios (make them discriminate) or the eval needs more distinct stimuli to clear the ties
Next action: Inspect tied or lost stimuli; predeclare added breadth before a new experiment.
Repeated-run reliability (not used by the gate): 7 paired runs (2W/5T/0L).
Weak or warning scenarios:
Scenario
Preference gate
Net win
Δ Pref
Runs (W/T/L)
= Assess a requested binlog investigation when the current build is healthy
Eligible
+0.0%
+0.0%
0/1/0
= Confirm the actual resolved target framework and package version from a binlog
Eligible
+0.0%
+0.0%
0/0/0
= Determine whether a quiet second build actually failed
Eligible
+0.0%
+0.0%
0/1/0
= Diagnose a warning behind a build that actually succeeded
Eligible
+0.0%
+0.0%
0/1/0
= Diagnose build failures from binlog only (no source files)
Eligible
+0.0%
+0.0%
0/1/0
= Stay dormant for a non-MSBuild build failure log
Excluded (activation contract)
+0.0%
+0.0%
0/1/0
Illustrative judge evidence:
Assess a requested binlog investigation when the current build is healthy: Position-swap inconsistent (forward: tie, reverse: A). Defaulting to tie.
This is one example, not the aggregate verdict. Open Full Results for every judgment.
➖ Not proven improved — binlog-generation (claude-sonnet-5)
Why: Net win +42.9% (4W/2T/1L over 7 preference-eligible stimulus vote(s), sign test p=0.188), mean preference +37.5% across 8 paired run(s), 1 dormancy stimulus/stimuli excluded from preference — not credible (sign test p=0.188 > 0.05)
Next action: Inspect tied or lost stimuli and fix inconsistent skill behavior.
Repeated-run reliability (not used by the gate): 8 paired runs (4W/3T/1L).
Weak or warning scenarios:
Scenario
Preference gate
Net win
Δ Pref
Runs (W/T/L)
= Choose a predictable, non-colliding binlog name for a CI upload step
Eligible
+0.0%
+0.0%
0/1/0
= Decline Maven compiler configuration
Excluded (activation contract)
+0.0%
+0.0%
0/1/0
▼ Preserve binlog history while cleaning stale build output
Eligible
-100.0%
-40.0%
0/0/1
= Recognize that a failed build produced no binlog at all
Eligible
+0.0%
+0.0%
0/1/0
Illustrative judge evidence:
Choose a predictable, non-colliding binlog name for a CI upload step: Both responses fully satisfy the task with the same valid, explicit non-colliding filenames and successful Debug and Release binlog builds. B provides an absolute directory in its final message, but that is not a material quality advantage over A's equally clear result.
This is one example, not the aggregate verdict. Open Full Results for every judgment.
➖ Not proven improved — binlog-generation (gpt-5.6-luna)
Why: Net win +42.9% (4W/2T/1L over 7 preference-eligible stimulus vote(s), sign test p=0.188), mean preference +37.5% across 8 paired run(s), 1 dormancy stimulus/stimuli excluded from preference — not credible (sign test p=0.188 > 0.05)
Next action: Inspect activation-only failed runs before rewriting skill content; the model stopped after loading a skill.
Repeated-run reliability (not used by the gate): 8 paired runs (4W/3T/1L).
Weak or warning scenarios:
Scenario
Preference gate
Net win
Δ Pref
Runs (W/T/L)
= Choose a predictable, non-colliding binlog name for a CI upload step
Eligible
+0.0%
+0.0%
0/1/0
= Decline Maven compiler configuration
Excluded (activation contract)
+0.0%
+0.0%
0/1/0
▼ Preserve binlog history while cleaning stale build output
Eligible
-100.0%
-40.0%
0/0/1
= Recognize that a failed build produced no binlog at all
Eligible
+0.0%
+0.0%
0/1/0
Illustrative judge evidence:
Choose a predictable, non-colliding binlog name for a CI upload step: Both responses successfully complete the task with identical, correct final outputs: two non-colliding binlogs (debug-4.binlog and release-4.binlog) that follow the manifest's naming scheme and are ready for CI use. Response B is marginally more explicit in confirming the mani...
This is one example, not the aggregate verdict. Open Full Results for every judgment.
➖ Not proven improved — build-parallelism (claude-sonnet-5)
Why: Net win -16.7% (1W/3T/2L over 6 preference-eligible stimulus vote(s), sign test p=0.500), mean preference +0.0% across 7 paired run(s), 1 dormancy stimulus/stimuli excluded from preference — no improvement
Next action: Inspect activation-only failed runs before rewriting skill content; the model stopped after loading a skill.
Repeated-run reliability (not used by the gate): 7 paired runs (2W/3T/2L).
Weak or warning scenarios:
Scenario
Preference gate
Net win
Δ Pref
Runs (W/T/L)
▼ Analyze build parallelism bottlenecks
Eligible
-100.0%
-40.0%
0/0/1
= Decline graph build for runtime-discovered projects
Eligible
+0.0%
+0.0%
0/1/0
= Enable BuildInParallel on a custom MSBuild task
Eligible
+0.0%
+0.0%
0/1/0
= Enable parallel nodes for a wide project graph
Eligible
+0.0%
+0.0%
0/1/0
▼ Reduce CI build scope with a solution filter
Eligible
-100.0%
-40.0%
0/0/1
Illustrative judge evidence:
Analyze build parallelism bottlenecks: The substantive conclusions are equally correct, but A better substantiates the critical-path claim with concrete timing/scheduling evidence and more explicitly addresses why parallel nodes do not help.
This is one example, not the aggregate verdict. Open Full Results for every judgment.
➖ Not proven improved — build-parallelism (gpt-5.6-luna)
Why: Net win +66.7% (4W/2T/0L over 6 preference-eligible stimulus vote(s), sign test p=0.063), mean preference +25.7% across 7 paired run(s), 1 dormancy stimulus/stimuli excluded from preference — not credible — 2 of 6 preference-eligible stimulus vote(s) tied, leaving only 4 discordant preference vote(s). The sign test conditions on non-tie stimulus votes and cannot reach 0.05 below 5, so no record could have passed here — this is not a measured null. Either the skill is inert on these scenarios (make them discriminate) or the eval needs more distinct stimuli to clear the ties
Next action: Inspect activation-only failed runs before rewriting skill content; the model stopped after loading a skill.
Repeated-run reliability (not used by the gate): 7 paired runs (4W/2T/1L).
Weak or warning scenarios:
Scenario
Preference gate
Net win
Δ Pref
Runs (W/T/L)
▼ Decline database query tuning request
Excluded (activation contract)
-100.0%
-40.0%
0/0/1
= Decline graph build for runtime-discovered projects
Eligible
+0.0%
+0.0%
0/1/0
▲ Enable parallel nodes for a wide project graph
Eligible
+100.0%
+40.0%
1/0/0
= Reduce CI build scope with a solution filter
Eligible
+0.0%
+0.0%
0/1/0
Illustrative judge evidence:
Decline graph build for runtime-discovered projects: Position-swap inconsistent (forward: B, reverse: A). Defaulting to tie.
This is one example, not the aggregate verdict. Open Full Results for every judgment.
➖ Not proven improved — build-perf-baseline (claude-sonnet-5)
Why: Net win +33.3% (3W/2T/1L over 6 preference-eligible stimulus vote(s), sign test p=0.312), mean preference +5.7% across 7 paired run(s), 1 dormancy stimulus/stimuli excluded from preference — not credible — 2 of 6 preference-eligible stimulus vote(s) tied, leaving only 4 discordant preference vote(s). The sign test conditions on non-tie stimulus votes and cannot reach 0.05 below 5, so no record could have passed here — this is not a measured null. Either the skill is inert on these scenarios (make them discriminate) or the eval needs more distinct stimuli to clear the ties
Next action: Inspect tied or lost stimuli; predeclare added breadth before a new experiment.
Repeated-run reliability (not used by the gate): 7 paired runs (3W/2T/2L).
Weak or warning scenarios:
Scenario
Preference gate
Net win
Δ Pref
Runs (W/T/L)
▼ Adopt UseArtifactsOutput while declining graph build for a small solution
Eligible
-100.0%
-40.0%
0/0/1
= Configure deterministic, cache-safe CI builds
Eligible
+0.0%
+0.0%
0/1/0
▼ Decline a non-MSBuild build performance request
Excluded (activation contract)
-100.0%
-40.0%
0/0/1
= Route a restore-bound cold build away from architecture changes
Eligible
+0.0%
+0.0%
0/1/0
Illustrative judge evidence:
Adopt UseArtifactsOutput while declining graph build for a small solution: Both provide the essential artifacts-layout recommendation, but A is more internally consistent and directly answers that /graph should not be enabled merely alongside artifacts output for a five-project solution. B's subsequent caveat improves its advice, but conflicts with i...
This is one example, not the aggregate verdict. Open Full Results for every judgment.
➖ Not proven improved — build-perf-baseline (gpt-5.6-luna)
Why: Net win +16.7% (3W/1T/2L over 6 preference-eligible stimulus vote(s), sign test p=0.500), mean preference +8.6% across 7 paired run(s), 1 dormancy stimulus/stimuli excluded from preference — not credible (sign test p=0.500 > 0.05)
Next action: Inspect tied or lost stimuli and fix inconsistent skill behavior.
Repeated-run reliability (not used by the gate): 7 paired runs (3W/1T/3L).
Weak or warning scenarios:
Scenario
Preference gate
Net win
Δ Pref
Runs (W/T/L)
= Configure deterministic, cache-safe CI builds
Eligible
+0.0%
+0.0%
0/1/0
▼ Decline a non-MSBuild build performance request
Excluded (activation contract)
-100.0%
-40.0%
0/0/1
▼ Leave an already-optimized build unchanged
Eligible
-100.0%
-40.0%
0/0/1
▼ Route a restore-bound cold build away from architecture changes
Eligible
-100.0%
-40.0%
0/0/1
Illustrative judge evidence:
Configure deterministic, cache-safe CI builds: Position-swap inconsistent (forward: B, reverse: A). Defaulting to tie.
This is one example, not the aggregate verdict. Open Full Results for every judgment.
➖ Not proven improved — build-perf-diagnostics (claude-sonnet-5)
Why: Net win +42.9% (3W/4T/0L over 7 preference-eligible stimulus vote(s), sign test p=0.125), mean preference +20.0% across 8 paired run(s), 1 dormancy stimulus/stimuli excluded from preference — not credible — 4 of 7 preference-eligible stimulus vote(s) tied, leaving only 3 discordant preference vote(s). The sign test conditions on non-tie stimulus votes and cannot reach 0.05 below 5, so no record could have passed here — this is not a measured null. Either the skill is inert on these scenarios (make them discriminate) or the eval needs more distinct stimuli to clear the ties
Next action: Inspect tied or lost stimuli; predeclare added breadth before a new experiment.
Repeated-run reliability (not used by the gate): 8 paired runs (4W/4T/0L).
Weak or warning scenarios:
Scenario
Preference gate
Net win
Δ Pref
Runs (W/T/L)
= Diagnose NuGet restore running redundantly across CI stages
Eligible
+0.0%
+0.0%
0/1/0
= Diagnose a Copy task dominating build time
Eligible
+0.0%
+0.0%
0/1/0
= Diagnose a single custom target dominating one project's build
Eligible
+0.0%
+0.0%
0/1/0
= Diagnose per-project overhead across many small projects
Eligible
+0.0%
+0.0%
0/1/0
Illustrative judge evidence:
Diagnose NuGet restore running redundantly across CI stages: Position-swap inconsistent (forward: tie, reverse: B). Defaulting to tie.
This is one example, not the aggregate verdict. Open Full Results for every judgment.
➖ Not proven improved — build-perf-diagnostics (gpt-5.6-luna)
Why: Net win +14.3% (3W/2T/2L over 7 preference-eligible stimulus vote(s), sign test p=0.500), mean preference +5.0% across 8 paired run(s), 1 dormancy stimulus/stimuli excluded from preference — not credible (sign test p=0.500 > 0.05)
Next action: Inspect tied or lost stimuli and fix inconsistent skill behavior.
Repeated-run reliability (not used by the gate): 8 paired runs (3W/3T/2L).
Weak or warning scenarios:
Scenario
Preference gate
Net win
Δ Pref
Runs (W/T/L)
= Decline a runtime latency request that is not a build performance issue
Excluded (activation contract)
+0.0%
+0.0%
0/1/0
= Diagnose NuGet restore running redundantly across CI stages
Eligible
+0.0%
+0.0%
0/1/0
= Diagnose a Copy task dominating build time
Eligible
+0.0%
+0.0%
0/1/0
▼ Diagnose a pathological ResolveAssemblyReference time
Eligible
-100.0%
-40.0%
0/0/1
▼ Diagnose per-project overhead across many small projects
Eligible
-100.0%
-40.0%
0/0/1
Illustrative judge evidence:
Diagnose NuGet restore running redundantly across CI stages: Position-swap inconsistent (forward: tie, reverse: A). Defaulting to tie.
This is one example, not the aggregate verdict. Open Full Results for every judgment.
➖ Not proven improved — check-bin-obj-clash (claude-sonnet-5)
Why: Net win +14.3% (2W/4T/1L over 7 preference-eligible stimulus vote(s), sign test p=0.500), mean preference +5.7% across 7 paired run(s) — not credible — 4 of 7 preference-eligible stimulus vote(s) tied, leaving only 3 discordant preference vote(s). The sign test conditions on non-tie stimulus votes and cannot reach 0.05 below 5, so no record could have passed here — this is not a measured null. Either the skill is inert on these scenarios (make them discriminate) or the eval needs more distinct stimuli to clear the ties
Next action: Inspect tied or lost stimuli; predeclare added breadth before a new experiment.
Gate evidence: n=7; 2W/4T/1L; d=3; p=0.500; net +14.3%
Warnings: Activation: isolated 7/7; plugin 5/7
Overfit: Moderate (score 0.21)
Repeated-run reliability (not used by the gate): 7 paired runs (2W/4T/1L).
Weak or warning scenarios:
Scenario
Preference gate
Net win
Δ Pref
Runs (W/T/L)
= Avoid a false clash report when projects share only the top-level artifacts root
Eligible
+0.0%
+0.0%
0/1/0
= Decline output-clash remediation for separate projects using the default SDK layout
Eligible
+0.0%
+0.0%
0/1/0
= Diagnose multi-targeting outputs that collapse into one path
Eligible
+0.0%
+0.0%
0/1/0
▼ Diagnose redundant project reference metadata that forks a same-path build
Eligible
-100.0%
-40.0%
0/0/1
= Fix all clash mechanisms in the mixed solution
Eligible
+0.0%
+0.0%
0/1/0
Illustrative judge evidence:
Avoid a false clash report when projects share only the top-level artifacts root: Position-swap inconsistent (forward: B, reverse: A). Defaulting to tie.
This is one example, not the aggregate verdict. Open Full Results for every judgment.
➖ Not proven improved — check-bin-obj-clash (gpt-5.6-luna)
Why: Net win +28.6% (2W/5T/0L over 7 preference-eligible stimulus vote(s), sign test p=0.250), mean preference +11.4% across 7 paired run(s) — not credible — 5 of 7 preference-eligible stimulus vote(s) tied, leaving only 2 discordant preference vote(s). The sign test conditions on non-tie stimulus votes and cannot reach 0.05 below 5, so no record could have passed here — this is not a measured null. Either the skill is inert on these scenarios (make them discriminate) or the eval needs more distinct stimuli to clear the ties
Next action: Inspect tied or lost stimuli; predeclare added breadth before a new experiment.
Gate evidence: n=7; 2W/5T/0L; d=2; p=0.250; net +28.6%
Overfit: Low (score 0.18)
Repeated-run reliability (not used by the gate): 7 paired runs (2W/5T/0L).
Weak or warning scenarios:
Scenario
Preference gate
Net win
Δ Pref
Runs (W/T/L)
= Audit the mixed solution and separate safe projects from unsafe ones
Eligible
+0.0%
+0.0%
0/1/0
= Decline output-clash remediation for separate projects using the default SDK layout
Eligible
+0.0%
+0.0%
0/1/0
= Diagnose multi-targeting outputs that collapse into one path
Eligible
+0.0%
+0.0%
0/1/0
= Diagnose redundant project reference metadata that forks a same-path build
Eligible
+0.0%
+0.0%
0/1/0
= Diagnose shared output and intermediate path collision
Eligible
+0.0%
+0.0%
0/1/0
Illustrative judge evidence:
Audit the mixed solution and separate safe projects from unsafe ones: Position-swap inconsistent (forward: B, reverse: A). Defaulting to tie.
This is one example, not the aggregate verdict. Open Full Results for every judgment.
➖ Not proven improved — code-testing-agent (claude-sonnet-5)
Why: Net win +33.3% (5W/2T/2L over 9 preference-eligible stimulus vote(s), sign test p=0.227), mean preference +15.6% across 18 paired run(s) — not credible (sign test p=0.227 > 0.05)
Next action: Inspect tied or lost stimuli and fix inconsistent skill behavior.
Gate evidence: n=9; 5W/2T/2L; d=7; p=0.227; net +33.3%
Warnings: Activation: isolated 5/9; plugin 5/9
Overfit: Moderate (score 0.45)
Repeated-run reliability (not used by the gate): 18 paired runs (11W/3T/4L).
Weak or warning scenarios:
Scenario
Preference gate
Net win
Δ Pref
Runs (W/T/L)
= Add focused xUnit tests for one reservation class
Eligible
+0.0%
+0.0%
0/2/0
▲ Expand a healthy existing pytest suite to every ledger boundary
Eligible
+100.0%
+40.0%
2/0/0
= Generate a layered Vitest suite for an async shopping cart
Eligible
+0.0%
+0.0%
1/0/1
▼ Generate a project-wide Go suite across collaborating packages
Eligible
-50.0%
-20.0%
0/1/1
▼ Generate project-wide tests for an SDK-style xUnit library
Eligible
-100.0%
-40.0%
0/0/2
Illustrative judge evidence:
Add focused xUnit tests for one reservation class: The two responses are substantively equivalent: both produced focused passing ReservationWindow boundary and validation tests, avoided forbidden changes, and gave concise completion summaries. Neither has a meaningful quality advantage.
This is one example, not the aggregate verdict. Open Full Results for every judgment.
➖ Not proven improved — detect-static-dependencies (gpt-5.6-luna)
Why: Net win +42.9% (4W/2T/1L over 7 preference-eligible stimulus vote(s), sign test p=0.188), mean preference +22.5% across 8 paired run(s), 1 dormancy stimulus/stimuli excluded from preference — not credible (sign test p=0.188 > 0.05)
Next action: Inspect tied or lost stimuli and fix inconsistent skill behavior.
Repeated-run reliability (not used by the gate): 8 paired runs (4W/3T/1L).
Weak or warning scenarios:
Scenario
Preference gate
Net win
Δ Pref
Runs (W/T/L)
= Detect time-related statics and recommend TimeProvider
Eligible
+0.0%
+0.0%
0/1/0
= Exclude obj and bin directories from the scan
Eligible
+0.0%
+0.0%
0/1/0
▼ Keep one authoritative total with file line locations and no seam for pure helpers
Eligible
-100.0%
-40.0%
0/0/1
= Stay dormant for Python timezone review
Excluded (activation contract)
+0.0%
+0.0%
0/1/0
Illustrative judge evidence:
Detect time-related statics and recommend TimeProvider: Position-swap inconsistent (forward: B, reverse: A). Defaulting to tie.
This is one example, not the aggregate verdict. Open Full Results for every judgment.
➖ Not proven improved — eval-performance (claude-sonnet-5)
Why: Net win +12.5% (3W/3T/2L over 8 preference-eligible stimulus vote(s), sign test p=0.500), mean preference +12.5% across 8 paired run(s) — not credible (sign test p=0.500 > 0.05)
Next action: Inspect tied or lost stimuli and fix inconsistent skill behavior.
To investigate non-passing or warning results, paste this to your AI coding agent:
For PR 1118 in dotnet/skills, download eval artifacts with gh run download 36034404215 --repo dotnet/skills --pattern "vally-results-*" --dir ./eval-results, then fetch https://raw.githubusercontent.com/dotnet/skills/be46d44caa481ba5dab779b84f8e16377f83e12b/eng/vally-adapter/InvestigatingResults.md and follow it. Classify each result as measurement-invalid, underpowered, not-proven, preference-loss, or passing-with-warning. Use stateReason, result accounting, weak scenarios, and judge evidence to give the cause and exact next fix.
⚠️ Session replay telemetry was not published because the auxiliary dotnet/skills-data publisher failed. The evaluation verdicts above remain authoritative; maintainers must repair SKILLS_DATA_TOKEN or the publisher before replay links are available.
Superseded by the clean replacement set: #1214 for MSBuild eval tests and shared infrastructure, stacked #1215 for MSBuild SKILL.md guidance, and #1213 for dotnet-test fixes. The replacement diffs were independently reviewed and verified against the frozen source ledger.
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Summary
This PR strengthens the
dotnet-msbuildevaluation surface from 33 to 139 distinct stimuli across 19 suites. It adds realistic failures, routing boundaries, valid no-op cases, replayable golden patches, substantial ATIF references, deterministic graders, and self-contained fixtures.It also hardens the shared eval-quality gate, native evaluator, and Vally result adaptation so activation, executor failures, permissions, session persistence, and result completeness fail closed.
Why
Earlier CI reports mixed real routing defects with fixture defects, harness limitations, statistical power, reliability failures, and trusted-workflow bootstrap behavior. This change separates those causes and fixes only defects supported by executable or multi-model evidence.
The base PR keeps the three routing fixes required by its own evidence: the
msbuildcustom agent, the non-MSBuild boundary forbuild-perf-baseline, and the incremental/FileWrites boundary forextension-points. Broader production guidance fixes are isolated in #1182.Related issue: #896.
Impact
Validation
Current head:
59bc36e2dbeaf5bb059f956a313a11cbf24b4897, current withmainand conflict-free.The latest exact-head evaluation 35246150186 proved two trusted-harness defects: PAT quota/no-auth responses were not classified as pool-candidate failures, and native-agent adaptation treated diagnostic
errorCountas terminal measurement failure. The fixes exist on this branch but cannot affect its own run becauseevaluation-run.ymlchecks out trusted workflow source and the validator archive from default branch.Prerequisite #1190 contains only those six trusted harness/docs/test files against
main. After #1190 merges, update this branch frommain, remove the harness files from the effective PR diff, and run one exact-head evaluation. A fresh human approval is also required after the latest head changes.Scope
This PR contains the strengthened MSBuild evals and fixtures, shared eval-quality checks and authoring guidance, native evaluator isolation and correctness fixes, permission and token hardening, rejudge role/activation support, target-specific Vally attribution, and only the three evidence-backed routing changes listed above. The trusted harness classification subset is split into #1190 so it can land on
mainbefore this PR evaluates.