Skip to content

test(msbuild): strengthen eval quality - #1118

Closed
AbhitejJohn wants to merge 133 commits into
mainfrom
abhitejjohn-improve-eval-quality
Closed

AbhitejJohn wants to merge 133 commits into
mainfrom
abhitejjohn-improve-eval-quality

Conversation

@AbhitejJohn

@AbhitejJohn AbhitejJohn commented Sep 4, 2026 •

Copy link
Copy Markdown
Collaborator

Summary

This PR strengthens the dotnet-msbuild evaluation surface from 33 to 139 distinct stimuli across 19 suites. It adds realistic failures, routing boundaries, valid no-op cases, replayable golden patches, substantial ATIF references, deterministic graders, and self-contained fixtures.

It also hardens the shared eval-quality gate, native evaluator, and Vally result adaptation so activation, executor failures, permissions, session persistence, and result completeness fail closed.

Why

Earlier CI reports mixed real routing defects with fixture defects, harness limitations, statistical power, reliability failures, and trusted-workflow bootstrap behavior. This change separates those causes and fixes only defects supported by executable or multi-model evidence.

The base PR keeps the three routing fixes required by its own evidence: the msbuild custom agent, the non-MSBuild boundary for build-perf-baseline, and the incremental/FileWrites boundary for extension-points. Broader production guidance fixes are isolated in #1182.

Related issue: #896.

Impact

  • Evaluations cover customer outcomes, routing, dormancy, and valid no-op behavior instead of prompt variations.
  • Dormancy scenarios cannot remove their target skill, are excluded from preference inference, and fail when the isolated target activates unexpectedly or disappears from executor output.
  • Fixtures, references, graders, and setup artifacts are contained, tracked, deterministic, and checked against their stated paths.
  • The native evaluator uses process-private storage, owner-only permissions, staged plugin content, anchored no-follow filesystem operations, scenario-scoped access, and fail-closed SDK permission handling.
  • GitHub token aliases and sensitive child-process environment variables are removed.
  • MCP execution is limited to the pinned Binlog MCP package through validator-owned NuGet configuration and private caches.
  • Session schema v5 preserves unknown legacy activation expectations and recovers them from current eval specs during rejudge.
  • Partial and terminal execution failures remain visible; rejudge cannot silently select only successful runs.
  • Vally plugin and target-agent activation attribution is target-specific, and missing expected stimuli remain explicit contract failures.

Validation

Current head: 59bc36e2dbeaf5bb059f956a313a11cbf24b4897, current with main and conflict-free.

  • Full Skill Validator suite: 789 tests passed on the coordinated evaluator hardening changes.
  • Cross-platform focused security and filesystem tests passed on Windows and self-contained .NET 10 Linux.
  • Vally adapter suites passed, including target-specific activation, missing-dormancy, retry, and native-agent result classification coverage.
  • Eval-quality gate passed with no errors; all 123 self-tests passed.
  • Skill Validator static check passed for 18 skills, 3 agents, and 1 plugin.
  • All review threads are resolved.
  • Earlier exact ordinary evidence remains successful and applicable to the MSBuild payload:

The latest exact-head evaluation 35246150186 proved two trusted-harness defects: PAT quota/no-auth responses were not classified as pool-candidate failures, and native-agent adaptation treated diagnostic errorCount as terminal measurement failure. The fixes exist on this branch but cannot affect its own run because evaluation-run.yml checks out trusted workflow source and the validator archive from default branch.

Prerequisite #1190 contains only those six trusted harness/docs/test files against main. After #1190 merges, update this branch from main, remove the harness files from the effective PR diff, and run one exact-head evaluation. A fresh human approval is also required after the latest head changes.

Scope

This PR contains the strengthened MSBuild evals and fixtures, shared eval-quality checks and authoring guidance, native evaluator isolation and correctness fixes, permission and token hardening, rejudge role/activation support, target-specific Vally attribution, and only the three evidence-backed routing changes listed above. The trusted harness classification subset is split into #1190 so it can land on main before this PR evaluates.

Expand MSBuild eval coverage with realistic fixtures, replayable references, and a changed-suite quality ratchet.

Co-authored-by: Copilot App <223556219+Copilot@users.noreply.github.com>

Copilot-Session: eb11c85b-4046-4dbc-94b3-3bdbd4839a0d
Clear NuGet fallback folders so the intentional NU1101 case cannot resolve from machine state.

Co-authored-by: Copilot App <223556219+Copilot@users.noreply.github.com>

Copilot-Session: eb11c85b-4046-4dbc-94b3-3bdbd4839a0d
@AbhitejJohn

Copy link
Copy Markdown
Collaborator Author

@JanKrivanek @YuliiaKovalova, I would appreciate your thoughts and feedback on this draft, especially on:

  • whether the 19 eval suites cover useful customer decisions and routing boundaries rather than prompt variations;
  • whether the balance of deterministic graders, golden patches, and substantial ATIF references is reliable and easy to maintain;
  • whether the changed-suite quality ratchet is the right way to improve new work without blocking on unrelated legacy debt.

For repository-size context, this PR adds 290 MSBuild artifact files totaling 176,115 bytes (172 KiB) uncompressed, or about 145,041 bytes (142 KiB) as a standalone ZIP: 83 KiB of fixtures/support files, 56 KiB of ATIF JSON, and 33 KiB of golden patches. The PR description now has a row-by-row review map for every changed skill eval and the TDD/Vally-guided design principles used to create them.

@AbhitejJohn AbhitejJohn left a comment

Copy link
Copy Markdown
Collaborator Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

/evaluate

@AbhitejJohn

Copy link
Copy Markdown
Collaborator Author

/evaluate 0d63a77

github-actions Bot added a commit that referenced this pull request Sep 4, 2026
@github-actions

github-actions Bot commented Sep 4, 2026

Copy link
Copy Markdown
Contributor

❌ Evaluation did not complete successfully (the evaluate job reported failure). Check the workflow run logs, then comment /evaluate 0d63a7748d1941554011285e0ad8c5916aa5dae4 to retry this exact commit.

35 partial result file(s) were preserved for diagnosis but were not consolidated because the full matrix did not complete.

@AbhitejJohn

Copy link
Copy Markdown
Collaborator Author

/evaluate 0d63a77

github-actions Bot added a commit that referenced this pull request Sep 4, 2026
@github-actions

github-actions Bot commented Sep 4, 2026

Copy link
Copy Markdown
Contributor

📊 Skill Evaluation Results

36 model/skill results across 18 skills and 2 models — ✅ 1 improved, ➖ 12 not proven improved, ⚠️ 0 invalid or underpowered, ⛔ 23 activation contract failures, 📉 0 preference losses (report only).

Measurement identity: evaluated commit 0d63a7748d1941554011285e0ad8c5916aa5dae4; 2 judge models.

Measurement health: 36 expected / 36 observed / 36 written; 0 missing, 0 unexpected, 0 invalid; 0 recovered comparison error slots and 0 unresolved comparison error slots.

Objective completion gate: not enabled. Aggregate completion transitions are telemetry only, so this report does not claim that zero objective regressions were proven.

A result passes only when preference-eligible distinct-stimulus votes have aggregate net win of at least 20% and an exact one-sided sign-test result of p ≤ 0.05, and every explicit dormancy activation contract passes. Repeated runs measure reliability only.

Skill Model Verdict Gate evidence Overfit Warnings Next action
binlog-failure-analysis claude-sonnet-4.6 ⛔ Activation contract failed n=6; 3W/2T/1L; d=4; p=0.312; net +33.3%; 2 dormancy excluded 🟡 0.26 Dormancy contract: 1 unexpected activation(s) Narrow skill routing so the listed off-target scenarios stay dormant.
binlog-failure-analysis gpt-5.6-luna ⛔ Activation contract failed n=6; 2W/3T/1L; d=3; p=0.500; net +16.7%; 2 dormancy excluded ✅ 0.14 Dormancy contract: 1 unexpected activation(s) Narrow skill routing so the listed off-target scenarios stay dormant.
binlog-generation claude-sonnet-4.6 ➖ Not proven improved n=7; 3W/3T/1L; d=4; p=0.312; net +28.6%; 1 dormancy excluded 🔴 0.58 Activation: isolated 5/7; plugin 5/7 Inspect tied or lost stimuli; predeclare added breadth before a new experiment.
binlog-generation gpt-5.6-luna ➖ Not proven improved n=7; 3W/2T/2L; d=5; p=0.500; net +14.3%; 1 dormancy excluded 🟡 0.21 Activation: isolated 7/7; plugin 6/7 Inspect tied or lost stimuli and fix inconsistent skill behavior.
build-parallelism claude-sonnet-4.6 ⛔ Activation contract failed n=6; 4W/2T/0L; d=4; p=0.063; net +66.7%; 1 dormancy excluded 🟡 0.34 Dormancy contract: 1 unexpected activation(s); Activation: isolated 5/6; plugin 5/6 Narrow skill routing so the listed off-target scenarios stay dormant.
build-parallelism gpt-5.6-luna ➖ Not proven improved n=6; 3W/0T/3L; d=6; p=0.656; net +0.0%; 1 dormancy excluded 🟡 0.22 — Inspect tied or lost stimuli and fix inconsistent skill behavior.
build-perf-baseline claude-sonnet-4.6 ⛔ Activation contract failed n=6; 4W/1T/1L; d=5; p=0.188; net +50.0%; 1 dormancy excluded 🔴 0.52 Dormancy contract: 1 unexpected activation(s); Activation: isolated 4/6; plugin 6/6 Narrow skill routing so the listed off-target scenarios stay dormant.
build-perf-baseline gpt-5.6-luna ⛔ Activation contract failed n=6; 3W/2T/1L; d=4; p=0.312; net +33.3%; 1 dormancy excluded 🟡 0.23 Dormancy contract: 1 unexpected activation(s); Activation: isolated 5/6; plugin 6/6 Narrow skill routing so the listed off-target scenarios stay dormant.
build-perf-diagnostics claude-sonnet-4.6 ➖ Not proven improved n=7; 1W/4T/2L; d=3; p=0.500; net -14.3%; 1 dormancy excluded 🟡 0.39 — Inspect tied or lost stimuli; predeclare added breadth before a new experiment.
build-perf-diagnostics gpt-5.6-luna ✅ Improved n=7; 5W/2T/0L; d=5; p=0.031; net +71.4%; 1 dormancy excluded ✅ 0.16 — None.
check-bin-obj-clash claude-sonnet-4.6 ⛔ Activation contract failed n=6; 3W/2T/1L; d=4; p=0.312; net +33.3%; 1 dormancy excluded 🟡 0.22 Dormancy contract: 1 unexpected activation(s); Activation: isolated 5/6; plugin 5/6 Narrow skill routing so the listed off-target scenarios stay dormant.
check-bin-obj-clash gpt-5.6-luna ⛔ Activation contract failed n=6; 0W/4T/2L; d=2; p=0.250; net -33.3%; 1 dormancy excluded ✅ 0.11 Dormancy contract: 1 unexpected activation(s); Activation: isolated 6/6; plugin 5/6 Narrow skill routing so the listed off-target scenarios stay dormant.
directory-build-organization claude-sonnet-4.6 ➖ Not proven improved n=6; 2W/3T/1L; d=3; p=0.500; net +16.7%; 1 dormancy excluded 🟡 0.33 Activation: isolated 5/6; plugin 5/6 Inspect tied or lost stimuli; predeclare added breadth before a new experiment.
directory-build-organization gpt-5.6-luna ⛔ Activation contract failed n=6; 1W/5T/0L; d=1; p=0.500; net +16.7%; 1 dormancy excluded ✅ 0.07 Dormancy contract: 1 unexpected activation(s) Narrow skill routing so the listed off-target scenarios stay dormant.
eval-performance claude-sonnet-4.6 ⛔ Activation contract failed n=6; 3W/3T/0L; d=3; p=0.125; net +50.0%; 2 dormancy excluded 🟡 0.38 Dormancy contract: 2 unexpected activation(s); Activation: isolated 4/6; plugin 5/6 Narrow skill routing so the listed off-target scenarios stay dormant.
eval-performance gpt-5.6-luna ⛔ Activation contract failed n=6; 2W/4T/0L; d=2; p=0.250; net +33.3%; 2 dormancy excluded ✅ 0.09 Dormancy contract: 2 unexpected activation(s); Activation: isolated 6/6; plugin 5/6 Narrow skill routing so the listed off-target scenarios stay dormant.
extension-points claude-sonnet-4.6 ➖ Not proven improved n=7; 4W/2T/1L; d=5; p=0.188; net +42.9%; 1 dormancy excluded 🟡 0.26 Activation: isolated 5/7; plugin 6/7 Inspect tied or lost stimuli and fix inconsistent skill behavior.
extension-points gpt-5.6-luna ➖ Not proven improved n=7; 2W/3T/2L; d=4; p=0.687; net +0.0%; 1 dormancy excluded 🟡 0.26 Activation: isolated 7/7; plugin 6/7 Inspect tied or lost stimuli; predeclare added breadth before a new experiment.
including-generated-files claude-sonnet-4.6 ⛔ Activation contract failed n=7; 4W/2T/1L; d=5; p=0.188; net +42.9%; 1 dormancy excluded 🟡 0.41 Dormancy contract: 1 unexpected activation(s); Activation: isolated 4/7; plugin 6/7 Narrow skill routing so the listed off-target scenarios stay dormant.
including-generated-files gpt-5.6-luna ➖ Not proven improved n=7; 2W/2T/3L; d=5; p=0.500; net -14.3%; 1 dormancy excluded ✅ 0.15 — Inspect tied or lost stimuli and fix inconsistent skill behavior.
incremental-build claude-sonnet-4.6 ⛔ Activation contract failed n=8; 2W/5T/1L; d=3; p=0.500; net +12.5%; 1 dormancy excluded 🟡 0.36 Dormancy contract: 1 unexpected activation(s); Activation: isolated 4/8; plugin 4/8 Narrow skill routing so the listed off-target scenarios stay dormant.
incremental-build gpt-5.6-luna ⛔ Activation contract failed n=8; 4W/4T/0L; d=4; p=0.063; net +50.0%; 1 dormancy excluded ✅ 0.10 Dormancy contract: 1 unexpected activation(s); Activation: isolated 6/8; plugin 8/8 Narrow skill routing so the listed off-target scenarios stay dormant.
item-management claude-sonnet-4.6 ⛔ Activation contract failed n=6; 1W/3T/2L; d=3; p=0.500; net -16.7%; 1 dormancy excluded 🟡 0.38 Dormancy contract: 1 unexpected activation(s); Activation: isolated 5/6; plugin 4/6 Narrow skill routing so the listed off-target scenarios stay dormant.
item-management gpt-5.6-luna ➖ Not proven improved n=6; 1W/4T/1L; d=2; p=0.750; net +0.0%; 1 dormancy excluded ✅ 0.11 — Inspect tied or lost stimuli; predeclare added breadth before a new experiment.
msbuild-antipatterns claude-sonnet-4.6 ⛔ Activation contract failed n=7; 1W/6T/0L; d=1; p=0.500; net +14.3%; 1 dormancy excluded 🟡 0.36 Dormancy contract: 1 unexpected activation(s); Activation: isolated 2/7; plugin 2/7 Narrow skill routing so the listed off-target scenarios stay dormant.
msbuild-antipatterns gpt-5.6-luna ⛔ Activation contract failed n=7; 0W/4T/3L; d=3; p=0.125; net -42.9%; 1 dormancy excluded ✅ 0.09 Dormancy contract: 1 unexpected activation(s); Activation: isolated 5/7; plugin 4/7 Narrow skill routing so the listed off-target scenarios stay dormant.
msbuild-modernization claude-sonnet-4.6 ➖ Not proven improved n=6; 4W/2T/0L; d=4; p=0.063; net +66.7%; 1 dormancy excluded 🟡 0.29 Activation: isolated 6/6; plugin 5/6 Inspect tied or lost stimuli; predeclare added breadth before a new experiment.
msbuild-modernization gpt-5.6-luna ➖ Not proven improved n=6; 2W/2T/2L; d=4; p=0.687; net +0.0%; 1 dormancy excluded ✅ 0.05 Activation: isolated 5/6; plugin 6/6 Inspect tied or lost stimuli; predeclare added breadth before a new experiment.
msbuild-server claude-sonnet-4.6 ⛔ Activation contract failed n=6; 3W/1T/2L; d=5; p=0.500; net +16.7%; 2 dormancy excluded 🟡 0.47 Dormancy contract: 2 unexpected activation(s) Narrow skill routing so the listed off-target scenarios stay dormant.
msbuild-server gpt-5.6-luna ⛔ Activation contract failed n=6; 4W/1T/1L; d=5; p=0.188; net +50.0%; 2 dormancy excluded 🟡 0.22 Dormancy contract: 1 unexpected activation(s); Activation: isolated 5/6; plugin 6/6 Narrow skill routing so the listed off-target scenarios stay dormant.
property-patterns claude-sonnet-4.6 ⛔ Activation contract failed n=6; 4W/1T/1L; d=5; p=0.188; net +50.0%; 1 dormancy excluded 🟡 0.23 Dormancy contract: 1 unexpected activation(s); Activation: isolated 4/6; plugin 5/6 Narrow skill routing so the listed off-target scenarios stay dormant.
property-patterns gpt-5.6-luna ⛔ Activation contract failed n=6; 3W/3T/0L; d=3; p=0.125; net +50.0%; 1 dormancy excluded ✅ 0.06 Dormancy contract: 1 unexpected activation(s); Activation: isolated 6/6; plugin 5/6 Narrow skill routing so the listed off-target scenarios stay dormant.
resolve-project-references claude-sonnet-4.6 ⛔ Activation contract failed n=5; 4W/0T/1L; d=5; p=0.188; net +60.0%; 3 dormancy excluded 🟡 0.32 Dormancy contract: 1 unexpected activation(s) Narrow skill routing so the listed off-target scenarios stay dormant.
resolve-project-references gpt-5.6-luna ⛔ Activation contract failed n=5; 2W/3T/0L; d=2; p=0.250; net +40.0%; 3 dormancy excluded ✅ 0.14 Dormancy contract: 1 unexpected activation(s) Narrow skill routing so the listed off-target scenarios stay dormant.
target-authoring claude-sonnet-4.6 ⛔ Activation contract failed n=6; 3W/2T/1L; d=4; p=0.312; net +33.3%; 1 dormancy excluded 🟡 0.29 Dormancy contract: 1 unexpected activation(s); Activation: isolated 6/6; plugin 4/6 Narrow skill routing so the listed off-target scenarios stay dormant.
target-authoring gpt-5.6-luna ➖ Not proven improved n=6; 2W/2T/2L; d=4; p=0.687; net +0.0%; 1 dormancy excluded ✅ 0.13 — Inspect tied or lost stimuli; predeclare added breadth before a new experiment.
ℹ️ How to read this report
  • ✅ Improved — the result passed both the statistical gate and the 20% practical net-win floor.
  • ➖ Not proven improved — the result is valid but did not pass both gates. This is not automatically a regression.
  • ⚠️ Invalid / underpowered — the gate withheld a quality verdict. Fix the measurement before judging the skill.
  • ⛔ Activation contract failed — the isolated target skill activated on an explicit dormancy scenario. Dormancy preference is excluded, but this routing failure still blocks a pass.
  • 📉 Preference loss — the LLM judge credibly preferred baseline. It is report-only, not objective completion proof.
  • Gate evidence — n preference-eligible distinct-stimulus votes, W/T/L stimulus votes, d discordant votes, exact one-sided p, net win, and the count of separately retained dormancy stimuli. The p value applies to one model/skill result; no matrix-wide multiple-comparison correction is applied.
  • Overfit — overfitting-judge severity (✅ Low, 🟡 Moderate, 🔴 High, — none) and score.
  • Warnings — activation, timeout, retry recovery, or unresolved comparison conditions that need attention.
  • Do not add repeated runs to increase statistical power. Do not add stimuli after seeing a near-pass unless the new breadth is predeclared for a new experiment.
⛔ Activation contract failed — binlog-failure-analysis (claude-sonnet-4.6)

Why: Net win +33.3% (3W/2T/1L over 6 preference-eligible stimulus vote(s), sign test p=0.312), mean preference +35.0% across 8 paired run(s), 2 dormancy stimulus/stimuli excluded from preference — activation contract failed (1 explicit dormancy scenario(s) activated the isolated target skill)

Next action: Narrow skill routing so the listed off-target scenarios stay dormant.

State: VALID_NO_CHANGE (activation_contract_failed)

Gate evidence: n=6; 3W/2T/1L; d=4; p=0.312; net +33.3%; 2 dormancy excluded

Warnings: Dormancy contract: 1 unexpected activation(s)

Overfit: Moderate (score 0.26)

Repeated-run reliability (not used by the gate): 8 paired runs (5W/2T/1L).

Weak or warning scenarios:

Scenario Preference gate Net win Δ Pref Runs (W/T/L)
= Confirm the actual resolved target framework and package version from a binlog Eligible +0.0% +0.0% 0/1/0
▲ Decline to "diagnose" a build that has not actually failed yet Excluded (activation contract) +100.0% +100.0% 1/0/0
= Diagnose a warning behind a build that actually succeeded Eligible +0.0% +0.0% 0/1/0
▼ Trace why a generated source file is missing at compile time Eligible -100.0% -40.0% 0/0/1

Illustrative judge evidence:

  • Confirm the actual resolved target framework and package version from a binlog: Position-swap inconsistent (forward: B, reverse: tie). Defaulting to tie.

This is one example, not the aggregate verdict. Open Full Results for every judgment.

⛔ Activation contract failed — binlog-failure-analysis (gpt-5.6-luna)

Why: Net win +16.7% (2W/3T/1L over 6 preference-eligible stimulus vote(s), sign test p=0.500), mean preference +5.0% across 8 paired run(s), 2 dormancy stimulus/stimuli excluded from preference — activation contract failed (1 explicit dormancy scenario(s) activated the isolated target skill)

Next action: Narrow skill routing so the listed off-target scenarios stay dormant.

State: VALID_NO_CHANGE (activation_contract_failed)

Gate evidence: n=6; 2W/3T/1L; d=3; p=0.500; net +16.7%; 2 dormancy excluded

Warnings: Dormancy contract: 1 unexpected activation(s)

Overfit: Low (score 0.14)

Repeated-run reliability (not used by the gate): 8 paired runs (2W/5T/1L).

Weak or warning scenarios:

Scenario Preference gate Net win Δ Pref Runs (W/T/L)
= Confirm the actual resolved target framework and package version from a binlog Eligible +0.0% +0.0% 0/1/0
= Decline to "diagnose" a build that has not actually failed yet Excluded (activation contract) +0.0% +0.0% 0/1/0
= Diagnose build failures from binlog only (no source files) Eligible +0.0% +0.0% 0/1/0
= Fall back to command-line log replay when the usual binlog tool is unavailable Eligible +0.0% +0.0% 0/1/0
= Stay dormant for a non-MSBuild build failure log Excluded (activation contract) +0.0% +0.0% 0/1/0
▼ Trace why a generated source file is missing at compile time Eligible -100.0% -40.0% 0/0/1

Illustrative judge evidence:

  • Confirm the actual resolved target framework and package version from a binlog: Position-swap inconsistent (forward: B, reverse: tie). Defaulting to tie.

This is one example, not the aggregate verdict. Open Full Results for every judgment.

⛔ Activation contract failed — build-parallelism (claude-sonnet-4.6)

Why: Net win +66.7% (4W/2T/0L over 6 preference-eligible stimulus vote(s), sign test p=0.063), mean preference +17.1% across 7 paired run(s), 1 dormancy stimulus/stimuli excluded from preference — activation contract failed (1 explicit dormancy scenario(s) activated the isolated target skill)

Next action: Narrow skill routing so the listed off-target scenarios stay dormant.

State: VALID_NO_CHANGE (activation_contract_failed)

Gate evidence: n=6; 4W/2T/0L; d=4; p=0.063; net +66.7%; 1 dormancy excluded

Warnings: Dormancy contract: 1 unexpected activation(s); Activation: isolated 5/6; plugin 5/6

Overfit: Moderate (score 0.34)

Repeated-run reliability (not used by the gate): 7 paired runs (4W/2T/1L).

Weak or warning scenarios:

Scenario Preference gate Net win Δ Pref Runs (W/T/L)
= Enable BuildInParallel on a custom MSBuild task Eligible +0.0% +0.0% 0/1/0
= Enable parallel nodes for a wide project graph Eligible +0.0% +0.0% 0/1/0
▼ Non-activation: make one custom target incremental Excluded (activation contract) -100.0% -40.0% 0/0/1

Illustrative judge evidence:

  • Enable BuildInParallel on a custom MSBuild task: The resulting edits and final explanations are substantively identical and fully satisfy the task.

This is one example, not the aggregate verdict. Open Full Results for every judgment.

⛔ Activation contract failed — build-perf-baseline (claude-sonnet-4.6)

Why: Net win +50.0% (4W/1T/1L over 6 preference-eligible stimulus vote(s), sign test p=0.188), mean preference +40.0% across 7 paired run(s), 1 dormancy stimulus/stimuli excluded from preference — activation contract failed (1 explicit dormancy scenario(s) activated the isolated target skill)

Next action: Narrow skill routing so the listed off-target scenarios stay dormant.

State: VALID_NO_CHANGE (activation_contract_failed)

Gate evidence: n=6; 4W/1T/1L; d=5; p=0.188; net +50.0%; 1 dormancy excluded

Warnings: Dormancy contract: 1 unexpected activation(s); Activation: isolated 4/6; plugin 6/6

Overfit: High (score 0.52)

Repeated-run reliability (not used by the gate): 7 paired runs (5W/1T/1L).

Weak or warning scenarios:

Scenario Preference gate Net win Δ Pref Runs (W/T/L)
▲ Configure deterministic, cache-safe CI builds Eligible +100.0% +100.0% 1/0/0
▲ Decline a non-MSBuild build performance request Excluded (activation contract) +100.0% +40.0% 1/0/0
▼ Leave an already-optimized build unchanged Eligible -100.0% -40.0% 0/0/1
= Route a restore-bound cold build away from architecture changes Eligible +0.0% +0.0% 0/1/0

Illustrative judge evidence:

  • Leave an already-optimized build unchanged: Both fail the central task by changing a repository with no applicable checklist anti-pattern. A at least locates and reviews the requested App/Core project, whereas B works on a different WebApi/BusinessLogic repository and makes unrelated reference edits; A is therefore slig...

This is one example, not the aggregate verdict. Open Full Results for every judgment.

⛔ Activation contract failed — build-perf-baseline (gpt-5.6-luna)

Why: Net win +33.3% (3W/2T/1L over 6 preference-eligible stimulus vote(s), sign test p=0.312), mean preference +14.3% across 7 paired run(s), 1 dormancy stimulus/stimuli excluded from preference — activation contract failed (1 explicit dormancy scenario(s) activated the isolated target skill)

Next action: Narrow skill routing so the listed off-target scenarios stay dormant.

State: VALID_NO_CHANGE (activation_contract_failed)

Gate evidence: n=6; 3W/2T/1L; d=4; p=0.312; net +33.3%; 1 dormancy excluded

Warnings: Dormancy contract: 1 unexpected activation(s); Activation: isolated 5/6; plugin 6/6

Overfit: Moderate (score 0.23)

Repeated-run reliability (not used by the gate): 7 paired runs (3W/2T/2L).

Weak or warning scenarios:

Scenario Preference gate Net win Δ Pref Runs (W/T/L)
▲ Configure deterministic, cache-safe CI builds Eligible +100.0% +100.0% 1/0/0
▼ Decline a non-MSBuild build performance request Excluded (activation contract) -100.0% -40.0% 0/0/1
▼ Leave an already-optimized build unchanged Eligible -100.0% -40.0% 0/0/1
= Route a broken no-op rebuild away from generic optimization Eligible +0.0% +0.0% 0/1/0
= Route a restore-bound cold build away from architecture changes Eligible +0.0% +0.0% 0/1/0

Illustrative judge evidence:

  • Leave an already-optimized build unchanged: The ideal outcome was to make no changes, as no baseline anti-pattern was present. Both responses failed this and invented a change. However, B committed the specific anti-pattern the rubric explicitly warns against (adding UseArtifactsOutput to a two-project solution), and it...

This is one example, not the aggregate verdict. Open Full Results for every judgment.

⛔ Activation contract failed — check-bin-obj-clash (claude-sonnet-4.6)

Why: Net win +33.3% (3W/2T/1L over 6 preference-eligible stimulus vote(s), sign test p=0.312), mean preference +25.7% across 7 paired run(s), 1 dormancy stimulus/stimuli excluded from preference — activation contract failed (1 explicit dormancy scenario(s) activated the isolated target skill)

Next action: Narrow skill routing so the listed off-target scenarios stay dormant.

State: VALID_NO_CHANGE (activation_contract_failed)

Gate evidence: n=6; 3W/2T/1L; d=4; p=0.312; net +33.3%; 1 dormancy excluded

Warnings: Dormancy contract: 1 unexpected activation(s); Activation: isolated 5/6; plugin 5/6

Overfit: Moderate (score 0.22)

Repeated-run reliability (not used by the gate): 7 paired runs (4W/2T/1L).

Weak or warning scenarios:

Scenario Preference gate Net win Δ Pref Runs (W/T/L)
= Audit the mixed solution and separate safe projects from unsafe ones Eligible +0.0% +0.0% 0/1/0
▲ Decline output-clash remediation for separate projects using the default SDK layout Excluded (activation contract) +100.0% +40.0% 1/0/0
= Diagnose multi-targeting outputs that collapse into one path Eligible +0.0% +0.0% 0/1/0
▼ Fix all clash mechanisms in the mixed solution Eligible -100.0% -40.0% 0/0/1

Illustrative judge evidence:

  • Audit the mixed solution and separate safe projects from unsafe ones: Position-swap inconsistent (forward: B, reverse: A). Defaulting to tie.

This is one example, not the aggregate verdict. Open Full Results for every judgment.

⛔ Activation contract failed — check-bin-obj-clash (gpt-5.6-luna)

Why: Net win -33.3% (0W/4T/2L over 6 preference-eligible stimulus vote(s), sign test p=0.250), mean preference -17.1% across 7 paired run(s), 1 dormancy stimulus/stimuli excluded from preference — activation contract failed (1 explicit dormancy scenario(s) activated the isolated target skill)

Next action: Narrow skill routing so the listed off-target scenarios stay dormant.

State: VALID_NO_CHANGE (activation_contract_failed)

Gate evidence: n=6; 0W/4T/2L; d=2; p=0.250; net -33.3%; 1 dormancy excluded

Warnings: Dormancy contract: 1 unexpected activation(s); Activation: isolated 6/6; plugin 5/6

Overfit: Low (score 0.11)

Repeated-run reliability (not used by the gate): 7 paired runs (0W/4T/3L).

Weak or warning scenarios:

Scenario Preference gate Net win Δ Pref Runs (W/T/L)
▼ Audit the mixed solution and separate safe projects from unsafe ones Eligible -100.0% -40.0% 0/0/1
= Avoid a false clash report when projects share only the top-level artifacts root Eligible +0.0% +0.0% 0/1/0
▼ Decline output-clash remediation for separate projects using the default SDK layout Excluded (activation contract) -100.0% -40.0% 0/0/1
= Diagnose multi-targeting outputs that collapse into one path Eligible +0.0% +0.0% 0/1/0
= Diagnose redundant project reference metadata that forks a same-path build Eligible +0.0% +0.0% 0/1/0
= Diagnose shared output and intermediate path collision Eligible +0.0% +0.0% 0/1/0
▼ Fix all clash mechanisms in the mixed solution Eligible -100.0% -40.0% 0/0/1

Illustrative judge evidence:

  • Audit the mixed solution and separate safe projects from unsafe ones: Both responses correctly identify the core clashes (LibraryA/B sharing, MultiTargetLib TFM overwrite) and the ToolLib redundant-build mechanism. However, A aligns better with the rubric's intended classification: it flags ConsumerApp as unsafe (the source of the redundant meta...

This is one example, not the aggregate verdict. Open Full Results for every judgment.

⛔ Activation contract failed — directory-build-organization (gpt-5.6-luna)

Why: Net win +16.7% (1W/5T/0L over 6 preference-eligible stimulus vote(s), sign test p=0.500), mean preference +5.7% across 7 paired run(s), 1 dormancy stimulus/stimuli excluded from preference — activation contract failed (1 explicit dormancy scenario(s) activated the isolated target skill)

Next action: Narrow skill routing so the listed off-target scenarios stay dormant.

State: VALID_NO_CHANGE (activation_contract_failed)

Gate evidence: n=6; 1W/5T/0L; d=1; p=0.500; net +16.7%; 1 dormancy excluded

Warnings: Dormancy contract: 1 unexpected activation(s)

Overfit: Low (score 0.07)

Repeated-run reliability (not used by the gate): 7 paired runs (1W/6T/0L).

Weak or warning scenarios:

Scenario Preference gate Net win Δ Pref Runs (W/T/L)
= Apply repo-level build organization cleanup Eligible +0.0% +0.0% 0/1/0
= Decline adding shared build files to a lone project Excluded (activation contract) +0.0% +0.0% 0/1/0
= Diagnose a TargetFramework condition that silently skips in props Eligible +0.0% +0.0% 0/1/0
= Diagnose a package downgrade chain and reorganize version management Eligible +0.0% +0.0% 0/1/0
= Diagnose an inner shared-props file that overwrites its own override Eligible +0.0% +0.0% 0/1/0
= Preserve an intentional project-specific exception while centralizing Eligible +0.0% +0.0% 0/1/0

Illustrative judge evidence:

  • Apply repo-level build organization cleanup: Position-swap inconsistent (forward: B, reverse: A). Defaulting to tie.

This is one example, not the aggregate verdict. Open Full Results for every judgment.

⛔ Activation contract failed — eval-performance (claude-sonnet-4.6)

Why: Net win +50.0% (3W/3T/0L over 6 preference-eligible stimulus vote(s), sign test p=0.125), mean preference +32.5% across 8 paired run(s), 2 dormancy stimulus/stimuli excluded from preference — activation contract failed (2 explicit dormancy scenario(s) activated the isolated target skill)

Next action: Narrow skill routing so the listed off-target scenarios stay dormant.

State: VALID_NO_CHANGE (activation_contract_failed)

Gate evidence: n=6; 3W/3T/0L; d=3; p=0.125; net +50.0%; 2 dormancy excluded

Warnings: Dormancy contract: 2 unexpected activation(s); Activation: isolated 4/6; plugin 5/6

Overfit: Moderate (score 0.38)

Repeated-run reliability (not used by the gate): 8 paired runs (5W/3T/0L).

Weak or warning scenarios:

Scenario Preference gate Net win Δ Pref Runs (W/T/L)
= Detect a project evaluated twice under different global properties Eligible +0.0% +0.0% 0/1/0
= Recognize TreatAsLocalProperty overuse versus one justified entry Eligible +0.0% +0.0% 0/1/0
▲ Redirect a compile-time slowdown mistakenly framed as an evaluation problem Excluded (activation contract) +100.0% +40.0% 1/0/0
▲ Redirect an incremental-rebuild complaint mistakenly framed as an evaluation problem Excluded (activation contract) +100.0% +40.0% 1/0/0
= Triage which of two property functions actually costs evaluation time Eligible +0.0% +0.0% 0/1/0

Illustrative judge evidence:

  • Detect a project evaluated twice under different global properties: Position-swap inconsistent (forward: A, reverse: B). Defaulting to tie.

This is one example, not the aggregate verdict. Open Full Results for every judgment.

⛔ Activation contract failed — eval-performance (gpt-5.6-luna)

Why: Net win +33.3% (2W/4T/0L over 6 preference-eligible stimulus vote(s), sign test p=0.250), mean preference +25.0% across 8 paired run(s), 2 dormancy stimulus/stimuli excluded from preference — activation contract failed (2 explicit dormancy scenario(s) activated the isolated target skill)

Next action: Narrow skill routing so the listed off-target scenarios stay dormant.

State: VALID_NO_CHANGE (activation_contract_failed)

Gate evidence: n=6; 2W/4T/0L; d=2; p=0.250; net +33.3%; 2 dormancy excluded

Warnings: Dormancy contract: 2 unexpected activation(s); Activation: isolated 6/6; plugin 5/6

Overfit: Low (score 0.09)

Repeated-run reliability (not used by the gate): 8 paired runs (3W/4T/1L).

Weak or warning scenarios:

Scenario Preference gate Net win Δ Pref Runs (W/T/L)
= Detect a project evaluated twice under different global properties Eligible +0.0% +0.0% 0/1/0
= Diagnose stacked evaluation-time patterns (deep imports, broad glob, file-I/O property function) Eligible +0.0% +0.0% 0/1/0
= Recognize TreatAsLocalProperty overuse versus one justified entry Eligible +0.0% +0.0% 0/1/0
▼ Redirect a compile-time slowdown mistakenly framed as an evaluation problem Excluded (activation contract) -100.0% -40.0% 0/0/1
▲ Redirect an incremental-rebuild complaint mistakenly framed as an evaluation problem Excluded (activation contract) +100.0% +100.0% 1/0/0
= Triage which of two property functions actually costs evaluation time Eligible +0.0% +0.0% 0/1/0

Illustrative judge evidence:

  • Detect a project evaluated twice under different global properties: Position-swap inconsistent (forward: tie, reverse: A). Defaulting to tie.

This is one example, not the aggregate verdict. Open Full Results for every judgment.

⛔ Activation contract failed — including-generated-files (claude-sonnet-4.6)

Why: Net win +42.9% (4W/2T/1L over 7 preference-eligible stimulus vote(s), sign test p=0.188), mean preference +10.0% across 8 paired run(s), 1 dormancy stimulus/stimuli excluded from preference — activation contract failed (1 explicit dormancy scenario(s) activated the isolated target skill)

Next action: Narrow skill routing so the listed off-target scenarios stay dormant.

State: VALID_NO_CHANGE (activation_contract_failed)

Gate evidence: n=7; 4W/2T/1L; d=5; p=0.188; net +42.9%; 1 dormancy excluded

Warnings: Dormancy contract: 1 unexpected activation(s); Activation: isolated 4/7; plugin 6/7

Overfit: Moderate (score 0.41)

Repeated-run reliability (not used by the gate): 8 paired runs (4W/2T/2L).

Weak or warning scenarios:

Scenario Preference gate Net win Δ Pref Runs (W/T/L)
▼ Decline manual wiring for Roslyn source generators Excluded (activation contract) -100.0% -40.0% 0/0/1
= Diagnose hardcoded obj path for generated source Eligible +0.0% +0.0% 0/1/0
▲ Diagnose missing clean tracking for generated source Eligible +100.0% +40.0% 1/0/0
▼ Diagnose missing output registration for generated non-code file Eligible -100.0% -40.0% 0/0/1
= Diagnose wrong hook for generated source files Eligible +0.0% +0.0% 0/1/0

Illustrative judge evidence:

  • Diagnose hardcoded obj path for generated source: The outputs are substantively equivalent, accurate, and directly answer the requested root cause and correct path pattern. B is marginally more concise, while A is equally clear; neither has a meaningful quality advantage.

This is one example, not the aggregate verdict. Open Full Results for every judgment.

⛔ Activation contract failed — incremental-build (claude-sonnet-4.6)

Why: Net win +12.5% (2W/5T/1L over 8 preference-eligible stimulus vote(s), sign test p=0.500), mean preference +8.9% across 9 paired run(s), 1 dormancy stimulus/stimuli excluded from preference — activation contract failed (1 explicit dormancy scenario(s) activated the isolated target skill)

Next action: Narrow skill routing so the listed off-target scenarios stay dormant.

State: VALID_NO_CHANGE (activation_contract_failed)

Gate evidence: n=8; 2W/5T/1L; d=3; p=0.500; net +12.5%; 1 dormancy excluded

Warnings: Dormancy contract: 1 unexpected activation(s); Activation: isolated 4/8; plugin 4/8

Overfit: Moderate (score 0.36)

Repeated-run reliability (not used by the gate): 9 paired runs (3W/5T/1L).

Weak or warning scenarios:

Scenario Preference gate Net win Δ Pref Runs (W/T/L)
= Correct the assumption that Outputs alone enables incremental skipping Eligible +0.0% +0.0% 0/1/0
▲ Decline treating a cold first build as broken incrementality Excluded (activation contract) +100.0% +40.0% 1/0/0
= Diagnose custom targets that always rerun Eligible +0.0% +0.0% 0/1/0
▼ Explain slower builds when MSBuild skipped everything Eligible -100.0% -40.0% 0/0/1
= Explain why Visual Studio keeps rebuilding an up-to-date project Eligible +0.0% +0.0% 0/1/0
= Explain why clean leaves generated hash source behind Eligible +0.0% +0.0% 0/1/0
= Read a diagnostic log to find the stale input Eligible +0.0% +0.0% 0/1/0

Illustrative judge evidence:

  • Correct the assumption that Outputs alone enables incremental skipping: Position-swap inconsistent (forward: A, reverse: tie). Defaulting to tie.

This is one example, not the aggregate verdict. Open Full Results for every judgment.

⛔ Activation contract failed — incremental-build (gpt-5.6-luna)

Why: Net win +50.0% (4W/4T/0L over 8 preference-eligible stimulus vote(s), sign test p=0.063), mean preference +22.2% across 9 paired run(s), 1 dormancy stimulus/stimuli excluded from preference — activation contract failed (1 explicit dormancy scenario(s) activated the isolated target skill)

Next action: Narrow skill routing so the listed off-target scenarios stay dormant.

State: VALID_NO_CHANGE (activation_contract_failed)

Gate evidence: n=8; 4W/4T/0L; d=4; p=0.063; net +50.0%; 1 dormancy excluded

Warnings: Dormancy contract: 1 unexpected activation(s); Activation: isolated 6/8; plugin 8/8

Overfit: Low (score 0.10)

Repeated-run reliability (not used by the gate): 9 paired runs (5W/4T/0L).

Weak or warning scenarios:

Scenario Preference gate Net win Δ Pref Runs (W/T/L)
▲ Correct the assumption that Outputs alone enables incremental skipping Eligible +100.0% +40.0% 1/0/0
▲ Decline treating a cold first build as broken incrementality Excluded (activation contract) +100.0% +40.0% 1/0/0
= Explain slower builds when MSBuild skipped everything Eligible +0.0% +0.0% 0/1/0
= Explain why clean leaves generated hash source behind Eligible +0.0% +0.0% 0/1/0
= Identify a volatile output path that defeats incrementality Eligible +0.0% +0.0% 0/1/0
= Read a diagnostic log to find the stale input Eligible +0.0% +0.0% 0/1/0

Illustrative judge evidence:

  • Explain slower builds when MSBuild skipped everything: Position-swap inconsistent (forward: tie, reverse: A). Defaulting to tie.

This is one example, not the aggregate verdict. Open Full Results for every judgment.

⛔ Activation contract failed — item-management (claude-sonnet-4.6)

Why: Net win -16.7% (1W/3T/2L over 6 preference-eligible stimulus vote(s), sign test p=0.500), mean preference -11.4% across 7 paired run(s), 1 dormancy stimulus/stimuli excluded from preference — activation contract failed (1 explicit dormancy scenario(s) activated the isolated target skill)

Next action: Narrow skill routing so the listed off-target scenarios stay dormant.

State: VALID_NO_CHANGE (activation_contract_failed)

Gate evidence: n=6; 1W/3T/2L; d=3; p=0.500; net -16.7%; 1 dormancy excluded

Warnings: Dormancy contract: 1 unexpected activation(s); Activation: isolated 5/6; plugin 4/6

Overfit: Moderate (score 0.38)

Repeated-run reliability (not used by the gate): 7 paired runs (1W/3T/3L).

Weak or warning scenarios:

Scenario Preference gate Net win Δ Pref Runs (W/T/L)
▼ Decline a target ordering problem with no item-group defect Excluded (activation contract) -100.0% -40.0% 0/0/1
= Diagnose an ineffective Compile Remove that does not match the glob Eligible +0.0% +0.0% 0/1/0
= Diagnose real and claimed item problems in a code generation pipeline Eligible +0.0% +0.0% 0/1/0
▼ Fix item management anti-patterns Eligible -100.0% -40.0% 0/0/1
= Leave already-correct item management unchanged Eligible +0.0% +0.0% 0/1/0
▼ Leave correct single-list batching unchanged Eligible -100.0% -40.0% 0/0/1

Illustrative judge evidence:

  • Diagnose an ineffective Compile Remove that does not match the glob: The final answers are substantively equivalent, concise, correct, and fully address both the cause and correction.

This is one example, not the aggregate verdict. Open Full Results for every judgment.

⛔ Activation contract failed — msbuild-antipatterns (claude-sonnet-4.6)

Why: Net win +14.3% (1W/6T/0L over 7 preference-eligible stimulus vote(s), sign test p=0.500), mean preference +10.0% across 8 paired run(s), 1 dormancy stimulus/stimuli excluded from preference — activation contract failed (1 explicit dormancy scenario(s) activated the isolated target skill)

Next action: Narrow skill routing so the listed off-target scenarios stay dormant.

State: VALID_NO_CHANGE (activation_contract_failed)

Gate evidence: n=7; 1W/6T/0L; d=1; p=0.500; net +14.3%; 1 dormancy excluded

Warnings: Dormancy contract: 1 unexpected activation(s); Activation: isolated 2/7; plugin 2/7

Overfit: Moderate (score 0.36)

Repeated-run reliability (not used by the gate): 8 paired runs (2W/6T/0L).

Weak or warning scenarios:

Scenario Preference gate Net win Δ Pref Runs (W/T/L)
= Add a module to an F# project Eligible +0.0% +0.0% 0/1/0
= Add a signature file to define public API Eligible +0.0% +0.0% 0/1/0
= Distinguish a style backslash from a real cross-platform backslash bug Eligible +0.0% +0.0% 0/1/0
= Fix broken file order causing FS0039 Eligible +0.0% +0.0% 0/1/0
= Judge an unguarded import inside a NuGet package build folder Eligible +0.0% +0.0% 0/1/0
= Leave a clean project without inventing anti-patterns Eligible +0.0% +0.0% 0/1/0
▲ Non-activation: migrate a legacy project to SDK style Excluded (activation contract) +100.0% +40.0% 1/0/0

Illustrative judge evidence:

  • Add a module to an F# project: The implementations and demonstrated behavior are substantively identical: both meet the requested validation, project ordering, integration, and successful build/run requirements.

This is one example, not the aggregate verdict. Open Full Results for every judgment.

⛔ Activation contract failed — msbuild-antipatterns (gpt-5.6-luna)

Why: Net win -42.9% (0W/4T/3L over 7 preference-eligible stimulus vote(s), sign test p=0.125), mean preference -20.0% across 8 paired run(s), 1 dormancy stimulus/stimuli excluded from preference — activation contract failed (1 explicit dormancy scenario(s) activated the isolated target skill)

Next action: Narrow skill routing so the listed off-target scenarios stay dormant.

State: VALID_NO_CHANGE (activation_contract_failed)

Gate evidence: n=7; 0W/4T/3L; d=3; p=0.125; net -42.9%; 1 dormancy excluded

Warnings: Dormancy contract: 1 unexpected activation(s); Activation: isolated 5/7; plugin 4/7

Overfit: Low (score 0.09)

Repeated-run reliability (not used by the gate): 8 paired runs (0W/4T/4L).

Weak or warning scenarios:

Scenario Preference gate Net win Δ Pref Runs (W/T/L)
= Add a module to an F# project Eligible +0.0% +0.0% 0/1/0
= Add a signature file to define public API Eligible +0.0% +0.0% 0/1/0
= Distinguish a style backslash from a real cross-platform backslash bug Eligible +0.0% +0.0% 0/1/0
▼ Fix broken file order causing FS0039 Eligible -100.0% -40.0% 0/0/1
= Judge an unguarded import inside a NuGet package build folder Eligible +0.0% +0.0% 0/1/0
▼ Leave a clean project without inventing anti-patterns Eligible -100.0% -40.0% 0/0/1
▼ Non-activation: migrate a legacy project to SDK style Excluded (activation contract) -100.0% -40.0% 0/0/1
▼ Review MSBuild files for anti-patterns and style issues Eligible -100.0% -40.0% 0/0/1

Illustrative judge evidence:

  • Add a module to an F# project: Both responses produced functionally identical results: an equivalent Validation.fs, updated Program.fs, correct compile ordering (evidenced by successful builds requiring Validation before Program), and successful dotnet run. B briefly hit a NuGet restore error but recovered ...

This is one example, not the aggregate verdict. Open Full Results for every judgment.

⛔ Activation contract failed — msbuild-server (claude-sonnet-4.6)

Why: Net win +16.7% (3W/1T/2L over 6 preference-eligible stimulus vote(s), sign test p=0.500), mean preference +20.0% across 8 paired run(s), 2 dormancy stimulus/stimuli excluded from preference — activation contract failed (2 explicit dormancy scenario(s) activated the isolated target skill)

Next action: Narrow skill routing so the listed off-target scenarios stay dormant.

State: VALID_NO_CHANGE (activation_contract_failed)

Gate evidence: n=6; 3W/1T/2L; d=5; p=0.500; net +16.7%; 2 dormancy excluded

Warnings: Dormancy contract: 2 unexpected activation(s)

Overfit: Moderate (score 0.47)

Repeated-run reliability (not used by the gate): 8 paired runs (4W/1T/3L).

Weak or warning scenarios:

Scenario Preference gate Net win Δ Pref Runs (W/T/L)
= Confirm the MSBuild Server is actually improving build times before declaring success Eligible +0.0% +0.0% 0/1/0
▼ Decline MSBuild Server for a single one-off release build Excluded (activation contract) -100.0% -40.0% 0/0/1
▼ Explain a background MSBuild Server process consuming memory Eligible -100.0% -40.0% 0/0/1
▼ Recommend MSBuild Server for slow CLI incremental builds Eligible -100.0% -40.0% 0/0/1
▲ Stay dormant for a purely IDE-side build slowdown Excluded (activation contract) +100.0% +100.0% 1/0/0

Illustrative judge evidence:

  • Confirm the MSBuild Server is actually improving build times before declaring success: Position-swap inconsistent (forward: A, reverse: B). Defaulting to tie.

This is one example, not the aggregate verdict. Open Full Results for every judgment.

⛔ Activation contract failed — msbuild-server (gpt-5.6-luna)

Why: Net win +50.0% (4W/1T/1L over 6 preference-eligible stimulus vote(s), sign test p=0.188), mean preference +35.0% across 8 paired run(s), 2 dormancy stimulus/stimuli excluded from preference — activation contract failed (1 explicit dormancy scenario(s) activated the isolated target skill)

Next action: Narrow skill routing so the listed off-target scenarios stay dormant.

State: VALID_NO_CHANGE (activation_contract_failed)

Gate evidence: n=6; 4W/1T/1L; d=5; p=0.188; net +50.0%; 2 dormancy excluded

Warnings: Dormancy contract: 1 unexpected activation(s); Activation: isolated 5/6; plugin 6/6

Overfit: Moderate (score 0.22)

Repeated-run reliability (not used by the gate): 8 paired runs (5W/2T/1L).

Weak or warning scenarios:

Scenario Preference gate Net win Δ Pref Runs (W/T/L)
▼ Configure MSBuild Server persistently across new Windows terminals Eligible -100.0% -40.0% 0/0/1
= Decline MSBuild Server for a single one-off release build Excluded (activation contract) +0.0% +0.0% 0/1/0
= Diagnose stale build output after enabling MSBuild Server Eligible +0.0% +0.0% 0/1/0

Illustrative judge evidence:

  • Configure MSBuild Server persistently across new Windows terminals: Both responses correctly recommend a persistent user-level Windows environment variable and instruct closing/reopening terminals. The decisive difference is the variable name: Response A uses DOTNET_CLI_USE_MSBUILD_SERVER, which matches the rubric's designated correct variable...

This is one example, not the aggregate verdict. Open Full Results for every judgment.

⛔ Activation contract failed — property-patterns (claude-sonnet-4.6)

Why: Net win +50.0% (4W/1T/1L over 6 preference-eligible stimulus vote(s), sign test p=0.188), mean preference +31.4% across 7 paired run(s), 1 dormancy stimulus/stimuli excluded from preference — activation contract failed (1 explicit dormancy scenario(s) activated the isolated target skill)

Next action: Narrow skill routing so the listed off-target scenarios stay dormant.

State: VALID_NO_CHANGE (activation_contract_failed)

Gate evidence: n=6; 4W/1T/1L; d=5; p=0.188; net +50.0%; 1 dormancy excluded

Warnings: Dormancy contract: 1 unexpected activation(s); Activation: isolated 4/6; plugin 5/6

Overfit: Moderate (score 0.23)

Repeated-run reliability (not used by the gate): 7 paired runs (5W/1T/1L).

Weak or warning scenarios:

Scenario Preference gate Net win Δ Pref Runs (W/T/L)
▲ Choose the right OS detection for a cross-platform property Eligible +100.0% +40.0% 1/0/0
▲ Decline a props-versus-targets placement question Excluded (activation contract) +100.0% +40.0% 1/0/0
= Diagnose a TargetFramework condition that never applies in props Eligible +0.0% +0.0% 0/1/0
▼ Fix shared property configuration Eligible -100.0% -40.0% 0/0/1

Illustrative judge evidence:

  • Diagnose a TargetFramework condition that never applies in props: Position-swap inconsistent (forward: B, reverse: tie). Defaulting to tie.

This is one example, not the aggregate verdict. Open Full Results for every judgment.

⛔ Activation contract failed — property-patterns (gpt-5.6-luna)

Why: Net win +50.0% (3W/3T/0L over 6 preference-eligible stimulus vote(s), sign test p=0.125), mean preference +17.1% across 7 paired run(s), 1 dormancy stimulus/stimuli excluded from preference — activation contract failed (1 explicit dormancy scenario(s) activated the isolated target skill)

Next action: Narrow skill routing so the listed off-target scenarios stay dormant.

State: VALID_NO_CHANGE (activation_contract_failed)

Gate evidence: n=6; 3W/3T/0L; d=3; p=0.125; net +50.0%; 1 dormancy excluded

Warnings: Dormancy contract: 1 unexpected activation(s); Activation: isolated 6/6; plugin 5/6

Overfit: Low (score 0.06)

Repeated-run reliability (not used by the gate): 7 paired runs (3W/4T/0L).

Weak or warning scenarios:

Scenario Preference gate Net win Δ Pref Runs (W/T/L)
▲ Choose the right OS detection for a cross-platform property Eligible +100.0% +40.0% 1/0/0
= Decline a props-versus-targets placement question Excluded (activation contract) +0.0% +0.0% 0/1/0
= Diagnose a TargetFramework condition that never applies in props Eligible +0.0% +0.0% 0/1/0
= Diagnose multi-level property hierarchy bugs Eligible +0.0% +0.0% 0/1/0
= Diagnose shared build property issues Eligible +0.0% +0.0% 0/1/0

Illustrative judge evidence:

  • Diagnose a TargetFramework condition that never applies in props: Position-swap inconsistent (forward: B, reverse: A). Defaulting to tie.

This is one example, not the aggregate verdict. Open Full Results for every judgment.

⛔ Activation contract failed — resolve-project-references (claude-sonnet-4.6)

Why: Net win +60.0% (4W/0T/1L over 5 preference-eligible stimulus vote(s), sign test p=0.188), mean preference +27.5% across 8 paired run(s), 3 dormancy stimulus/stimuli excluded from preference — activation contract failed (1 explicit dormancy scenario(s) activated the isolated target skill)

Next action: Narrow skill routing so the listed off-target scenarios stay dormant.

State: VALID_NO_CHANGE (activation_contract_failed)

Gate evidence: n=5; 4W/0T/1L; d=5; p=0.188; net +60.0%; 3 dormancy excluded

Warnings: Dormancy contract: 1 unexpected activation(s)

Overfit: Moderate (score 0.32)

Repeated-run reliability (not used by the gate): 8 paired runs (6W/0T/2L).

Weak or warning scenarios:

Scenario Preference gate Net win Δ Pref Runs (W/T/L)
▼ Rank Copy ahead of Csc when task self-time is higher Eligible -100.0% -100.0% 0/0/1
▼ Request diagnostic artifact before analysis Excluded (activation contract) -100.0% -40.0% 0/0/1

Illustrative judge evidence:

  • Rank Copy ahead of Csc when task self-time is higher: A directly and correctly answers the question using the report's task-level evidence, explains why the target-level leader is not the optimization target, and gives a relevant copy optimization direction. B has the general diagnostic principle right but fails to use the suppli...

This is one example, not the aggregate verdict. Open Full Results for every judgment.

⛔ Activation contract failed — resolve-project-references (gpt-5.6-luna)

Why: Net win +40.0% (2W/3T/0L over 5 preference-eligible stimulus vote(s), sign test p=0.250), mean preference +27.5% across 8 paired run(s), 3 dormancy stimulus/stimuli excluded from preference — activation contract failed (1 explicit dormancy scenario(s) activated the isolated target skill)

Next action: Narrow skill routing so the listed off-target scenarios stay dormant.

State: VALID_NO_CHANGE (activation_contract_failed)

Gate evidence: n=5; 2W/3T/0L; d=2; p=0.250; net +40.0%; 3 dormancy excluded

Warnings: Dormancy contract: 1 unexpected activation(s)

Overfit: Low (score 0.14)

Repeated-run reliability (not used by the gate): 8 paired runs (4W/4T/0L).

Weak or warning scenarios:

Scenario Preference gate Net win Δ Pref Runs (W/T/L)
= Distinguish wait time from a real serial dependency chain Eligible +0.0% +0.0% 0/1/0
= Rank Copy ahead of Csc when task self-time is higher Eligible +0.0% +0.0% 0/1/0
= Redirect from misleading target summary to Csc self-time Eligible +0.0% +0.0% 0/1/0
▲ Request diagnostic artifact before analysis Excluded (activation contract) +100.0% +40.0% 1/0/0
= Stay dormant when Csc is already the obvious bottleneck Excluded (activation contract) +0.0% +0.0% 0/1/0

Illustrative judge evidence:

  • Distinguish wait time from a real serial dependency chain: Position-swap inconsistent (forward: B, reverse: tie). Defaulting to tie.

This is one example, not the aggregate verdict. Open Full Results for every judgment.

⛔ Activation contract failed — target-authoring (claude-sonnet-4.6)

Why: Net win +33.3% (3W/2T/1L over 6 preference-eligible stimulus vote(s), sign test p=0.312), mean preference +20.0% across 7 paired run(s), 1 dormancy stimulus/stimuli excluded from preference — activation contract failed (1 explicit dormancy scenario(s) activated the isolated target skill)

Next action: Narrow skill routing so the listed off-target scenarios stay dormant.

State: VALID_NO_CHANGE (activation_contract_failed)

Gate evidence: n=6; 3W/2T/1L; d=4; p=0.312; net +33.3%; 1 dormancy excluded

Warnings: Dormancy contract: 1 unexpected activation(s); Activation: isolated 6/6; plugin 4/6

Overfit: Moderate (score 0.29)

Repeated-run reliability (not used by the gate): 7 paired runs (3W/3T/1L).

Weak or warning scenarios:

Scenario Preference gate Net win Δ Pref Runs (W/T/L)
= Decline an incremental-build tuning request Excluded (activation contract) +0.0% +0.0% 0/1/0
▲ Diagnose a target hooked to Build that misses direct compile Eligible +100.0% +40.0% 1/0/0
= Diagnose broken SDK target chain across files Eligible +0.0% +0.0% 0/1/0
= Diagnose custom target reliability issues Eligible +0.0% +0.0% 0/1/0
▼ Fix a query target that uses Outputs instead of Returns Eligible -100.0% -40.0% 0/0/1

Illustrative judge evidence:

  • Diagnose broken SDK target chain across files: Position-swap inconsistent (forward: tie, reverse: A). Defaulting to tie.

This is one example, not the aggregate verdict. Open Full Results for every judgment.

➖ Not proven improved — binlog-generation (claude-sonnet-4.6)

Why: Net win +28.6% (3W/3T/1L over 7 preference-eligible stimulus vote(s), sign test p=0.312), mean preference +20.0% across 8 paired run(s), 1 dormancy stimulus/stimuli excluded from preference — not credible — 3 of 7 preference-eligible stimulus vote(s) tied, leaving only 4 discordant preference vote(s). The sign test conditions on non-tie stimulus votes and cannot reach 0.05 below 5, so no record could have passed here — this is not a measured null. Either the skill is inert on these scenarios (make them discriminate) or the eval needs more distinct stimuli to clear the ties

Next action: Inspect tied or lost stimuli; predeclare added breadth before a new experiment.

State: VALID_NO_CHANGE (no_credible_preference_change)

Gate evidence: n=7; 3W/3T/1L; d=4; p=0.312; net +28.6%; 1 dormancy excluded

Warnings: Activation: isolated 5/7; plugin 5/7

Overfit: High (score 0.58)

Repeated-run reliability (not used by the gate): 8 paired runs (3W/3T/2L).

Weak or warning scenarios:

Scenario Preference gate Net win Δ Pref Runs (W/T/L)
= Build with /bl in PowerShell Eligible +0.0% +0.0% 0/1/0
= Choose a predictable, non-colliding binlog name for a CI upload step Eligible +0.0% +0.0% 0/1/0
▼ Decline a binlog request for a non-MSBuild Java build Excluded (activation contract) -100.0% -40.0% 0/0/1
= Fix a CI build script that reuses the same binlog on every run Eligible +0.0% +0.0% 0/1/0
▼ Preserve binlog history while cleaning stale build output Eligible -100.0% -40.0% 0/0/1

Illustrative judge evidence:

  • Build with /bl in PowerShell: Position-swap inconsistent (forward: B, reverse: A). Defaulting to tie.

This is one example, not the aggregate verdict. Open Full Results for every judgment.

➖ Not proven improved — binlog-generation (gpt-5.6-luna)

Why: Net win +14.3% (3W/2T/2L over 7 preference-eligible stimulus vote(s), sign test p=0.500), mean preference +27.5% across 8 paired run(s), 1 dormancy stimulus/stimuli excluded from preference — not credible (sign test p=0.500 > 0.05)

Next action: Inspect tied or lost stimuli and fix inconsistent skill behavior.

State: VALID_NO_CHANGE (no_credible_preference_change)

Gate evidence: n=7; 3W/2T/2L; d=5; p=0.500; net +14.3%; 1 dormancy excluded

Warnings: Activation: isolated 7/7; plugin 6/7

Overfit: Moderate (score 0.21)

Repeated-run reliability (not used by the gate): 8 paired runs (3W/3T/2L).

Weak or warning scenarios:

Scenario Preference gate Net win Δ Pref Runs (W/T/L)
▼ Build with /bl in PowerShell Eligible -100.0% -40.0% 0/0/1
= Choose a predictable, non-colliding binlog name for a CI upload step Eligible +0.0% +0.0% 0/1/0
= Decline a binlog request for a non-MSBuild Java build Excluded (activation contract) +0.0% +0.0% 0/1/0
= Fix a CI build script that reuses the same binlog on every run Eligible +0.0% +0.0% 0/1/0
▼ Recognize that a failed build produced no binlog at all Eligible -100.0% -40.0% 0/0/1

Illustrative judge evidence:

  • Build with /bl in PowerShell: Both successfully built the project via PowerShell and produced a valid binlog with 0 errors. A's approach used an explicit quoted filename, which is unambiguous and clean. B relied on the skill's {{}} template syntax which, while it produced a valid unique filename, uses the ...

This is one example, not the aggregate verdict. Open Full Results for every judgment.

➖ Not proven improved — build-parallelism (gpt-5.6-luna)

Why: Net win +0.0% (3W/0T/3L over 6 preference-eligible stimulus vote(s), sign test p=0.656), mean preference -5.7% across 7 paired run(s), 1 dormancy stimulus/stimuli excluded from preference — no improvement

Next action: Inspect tied or lost stimuli and fix inconsistent skill behavior.

State: VALID_NO_CHANGE (no_credible_preference_change)

Gate evidence: n=6; 3W/0T/3L; d=6; p=0.656; net +0.0%; 1 dormancy excluded

Overfit: Moderate (score 0.22)

Repeated-run reliability (not used by the gate): 7 paired runs (3W/0T/4L).

Weak or warning scenarios:

Scenario Preference gate Net win Δ Pref Runs (W/T/L)
▼ Enable parallel nodes for a wide project graph Eligible -100.0% -40.0% 0/0/1
▼ Non-activation: make one custom target incremental Excluded (activation contract) -100.0% -40.0% 0/0/1
▼ Preserve valid dependencies and optimize the slow critical-path project Eligible -100.0% -40.0% 0/0/1
▼ Reduce CI build scope with a solution filter Eligible -100.0% -40.0% 0/0/1

Illustrative judge evidence:

  • Enable parallel nodes for a wide project graph: Both responses are correct and nearly identical in substance, correctly diagnosing the single-node default and recommending -m. Neither suggested confirming via binlog. A slightly edges out by also offering the explicit -m:4 option, better matching the -m:N criterion.

This is one example, not the aggregate verdict. Open Full Results for every judgment.

➖ Not proven improved — build-perf-diagnostics (claude-sonnet-4.6)

Why: Net win -14.3% (1W/4T/2L over 7 preference-eligible stimulus vote(s), sign test p=0.500), mean preference -17.5% across 8 paired run(s), 1 dormancy stimulus/stimuli excluded from preference — no improvement

Next action: Inspect tied or lost stimuli; predeclare added breadth before a new experiment.

State: VALID_NO_CHANGE (no_credible_preference_change)

Gate evidence: n=7; 1W/4T/2L; d=3; p=0.500; net -14.3%; 1 dormancy excluded

Overfit: Moderate (score 0.39)

Repeated-run reliability (not used by the gate): 8 paired runs (1W/4T/3L).

Weak or warning scenarios:

Scenario Preference gate Net win Δ Pref Runs (W/T/L)
▼ Decline a runtime latency request that is not a build performance issue Excluded (activation contract) -100.0% -100.0% 0/0/1
▼ Diagnose NuGet restore running redundantly across CI stages Eligible -100.0% -40.0% 0/0/1
= Diagnose a Copy task dominating build time Eligible +0.0% +0.0% 0/1/0
▼ Diagnose a single custom target dominating one project's build Eligible -100.0% -40.0% 0/0/1
= Diagnose evaluation overhead before any target runs Eligible +0.0% +0.0% 0/1/0
= Diagnose per-project overhead across many small projects Eligible +0.0% +0.0% 0/1/0
= Diagnose slow build for a small project Eligible +0.0% +0.0% 0/1/0

Illustrative judge evidence:

  • Diagnose NuGet restore running redundantly across CI stages: Both reach the central diagnosis and recommended pipeline change, but A gives a more accurate accounting of avoidable work and avoids B's misleading claims about static graph evaluation and post-change stage timing.

This is one example, not the aggregate verdict. Open Full Results for every judgment.

Routine passing details for 1 result are in Full Results.

Details for 8 results were omitted to keep this comment under GitHub's limit. Open the workflow summary or Full Results for the complete breakdown.

🔍 Full Results - all metrics and investigation details

To investigate non-passing or warning results, paste this to your AI coding agent:

For PR 1118 in dotnet/skills, download eval artifacts with gh run download 33898074788 --repo dotnet/skills --pattern "vally-results-*" --dir ./eval-results, then fetch https://raw.githubusercontent.com/dotnet/skills/0d63a7748d1941554011285e0ad8c5916aa5dae4/eng/vally-adapter/InvestigatingResults.md and follow it. Classify each result as measurement-invalid, underpowered, not-proven, preference-loss, or passing-with-warning. Use stateReason, result accounting, weak scenarios, and judge evidence to give the cause and exact next fix.

▶ Sessions Visualisation -- interactive replay of all evaluation sessions
📊 Session Analytics (preview) -- aggregated metrics across evaluation sessions

@JanKrivanek JanKrivanek left a comment

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Do we want to split the infra changes and msbuild eval changes? Both are quite loaded by themselves

Comment thread eng/eval-quality/check_eval_quality.py
Restore dormancy constraint protection, correct in-scope activation expectations, and strengthen MSBuild skill routing from CI evidence.

Co-authored-by: Copilot App <223556219+Copilot@users.noreply.github.com>

Copilot-Session: 2d0a0a20-fd83-4343-8434-4763f8120503

@AbhitejJohn AbhitejJohn left a comment

Copy link
Copy Markdown
Collaborator Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

/evaluate

github-actions Bot added a commit that referenced this pull request Sep 8, 2026
@github-actions

github-actions Bot commented Sep 8, 2026

Copy link
Copy Markdown
Contributor

❌ Evaluation did not complete successfully (the evaluate job reported failure). Check the workflow run logs, then comment /evaluate 34de668c6602fbf6ec12ec19da961e9f11e6e2e1 to retry this exact commit.

36 partial result file(s) were preserved for diagnosis but were not consolidated because the full matrix did not complete.

Reclassify in-scope suitability checks, sharpen true routing boundaries, and isolate binlog fixture builds from shared build-server state.

Co-authored-by: Copilot App <223556219+Copilot@users.noreply.github.com>

Copilot-Session: 2d0a0a20-fd83-4343-8434-4763f8120503
github-actions Bot added a commit that referenced this pull request Sep 8, 2026
@github-actions

github-actions Bot commented Sep 8, 2026

Copy link
Copy Markdown
Contributor

📊 Skill Evaluation Results

36 model/skill results across 18 skills and 2 models — ✅ 3 improved, ➖ 27 not proven improved, ⚠️ 0 invalid or underpowered, ⛔ 6 activation contract failures, 📉 0 preference losses (report only).

Measurement identity: evaluated commit 6866836a9b096ae2e9aefefe6dbfdab67405f37f; 2 judge models.

Measurement health: 36 expected / 36 observed / 36 written; 0 missing, 0 unexpected, 0 invalid; 0 recovered comparison error slots and 0 unresolved comparison error slots.

Objective completion gate: not enabled. Aggregate completion transitions are telemetry only, so this report does not claim that zero objective regressions were proven.

A result passes only when preference-eligible distinct-stimulus votes have aggregate net win of at least 20% and an exact one-sided sign-test result of p ≤ 0.05, and every explicit dormancy activation contract passes. Repeated runs measure reliability only.

Skill Model Verdict Gate evidence Overfit Warnings Next action
binlog-failure-analysis claude-sonnet-4.6 ➖ Not proven improved n=7; 5W/1T/1L; d=6; p=0.109; net +57.1%; 1 dormancy excluded 🟡 0.24 — Inspect tied or lost stimuli and fix inconsistent skill behavior.
binlog-failure-analysis gpt-5.6-luna ➖ Not proven improved n=7; 4W/1T/2L; d=6; p=0.344; net +28.6%; 1 dormancy excluded ✅ 0.10 — Inspect tied or lost stimuli and fix inconsistent skill behavior.
binlog-generation claude-sonnet-4.6 ✅ Improved n=7; 5W/2T/0L; d=5; p=0.031; net +71.4%; 1 dormancy excluded 🟡 0.49 — Review overfit evidence.
binlog-generation gpt-5.6-luna ➖ Not proven improved n=7; 4W/3T/0L; d=4; p=0.063; net +57.1%; 1 dormancy excluded 🟡 0.28 — Inspect tied or lost stimuli; predeclare added breadth before a new experiment.
build-parallelism claude-sonnet-4.6 ⛔ Activation contract failed n=6; 3W/3T/0L; d=3; p=0.125; net +50.0%; 1 dormancy excluded 🟡 0.43 Dormancy contract: 1 unexpected activation(s); Activation: isolated 3/6; plugin 5/6 Narrow skill routing so the listed off-target scenarios stay dormant.
build-parallelism gpt-5.6-luna ➖ Not proven improved n=6; 0W/3T/3L; d=3; p=0.125; net -50.0%; 1 dormancy excluded ✅ 0.11 Activation: isolated 5/6; plugin 6/6 Inspect tied or lost stimuli; predeclare added breadth before a new experiment.
build-perf-baseline claude-sonnet-4.6 ⛔ Activation contract failed n=6; 4W/0T/2L; d=6; p=0.344; net +33.3%; 1 dormancy excluded 🟡 0.43 Dormancy contract: 1 unexpected activation(s); Activation: isolated 5/6; plugin 5/6 Narrow skill routing so the listed off-target scenarios stay dormant.
build-perf-baseline gpt-5.6-luna ➖ Not proven improved n=6; 4W/0T/2L; d=6; p=0.344; net +33.3%; 1 dormancy excluded 🟡 0.23 Activation: isolated 6/6; plugin 5/6 Inspect tied or lost stimuli and fix inconsistent skill behavior.
build-perf-diagnostics claude-sonnet-4.6 ➖ Not proven improved n=7; 3W/1T/3L; d=6; p=0.656; net +0.0%; 1 dormancy excluded 🟡 0.45 Activation: isolated 6/7; plugin 7/7 Inspect tied or lost stimuli and fix inconsistent skill behavior.
build-perf-diagnostics gpt-5.6-luna ➖ Not proven improved n=7; 3W/3T/1L; d=4; p=0.312; net +28.6%; 1 dormancy excluded ✅ 0.14 — Inspect tied or lost stimuli; predeclare added breadth before a new experiment.
check-bin-obj-clash claude-sonnet-4.6 ➖ Not proven improved n=7; 4W/1T/2L; d=6; p=0.344; net +28.6% ✅ 0.19 Activation: isolated 6/7; plugin 6/7 Inspect tied or lost stimuli and fix inconsistent skill behavior.
check-bin-obj-clash gpt-5.6-luna ➖ Not proven improved n=7; 0W/6T/1L; d=1; p=0.500; net -14.3% ✅ 0.13 — Inspect tied or lost stimuli; predeclare added breadth before a new experiment.
directory-build-organization claude-sonnet-4.6 ➖ Not proven improved n=6; 3W/2T/1L; d=4; p=0.312; net +33.3%; 1 dormancy excluded 🟡 0.33 Activation: isolated 5/6; plugin 5/6 Inspect tied or lost stimuli; predeclare added breadth before a new experiment.
directory-build-organization gpt-5.6-luna ⛔ Activation contract failed n=6; 3W/2T/1L; d=4; p=0.312; net +33.3%; 1 dormancy excluded ✅ 0.10 Dormancy contract: 1 unexpected activation(s) Narrow skill routing so the listed off-target scenarios stay dormant.
eval-performance claude-sonnet-4.6 ➖ Not proven improved n=8; 6W/1T/1L; d=7; p=0.063; net +62.5% 🟡 0.27 Activation: isolated 8/8; plugin 7/8 Inspect tied or lost stimuli and fix inconsistent skill behavior.
eval-performance gpt-5.6-luna ➖ Not proven improved n=8; 5W/2T/1L; d=6; p=0.109; net +50.0% ✅ 0.09 — Inspect tied or lost stimuli and fix inconsistent skill behavior.
extension-points claude-sonnet-4.6 ➖ Not proven improved n=7; 3W/2T/2L; d=5; p=0.500; net +14.3%; 1 dormancy excluded 🟡 0.26 Activation: isolated 4/7; plugin 5/7 Inspect tied or lost stimuli and fix inconsistent skill behavior.
extension-points gpt-5.6-luna ➖ Not proven improved n=7; 4W/3T/0L; d=4; p=0.063; net +57.1%; 1 dormancy excluded 🟡 0.24 — Inspect tied or lost stimuli; predeclare added breadth before a new experiment.
including-generated-files claude-sonnet-4.6 ⛔ Activation contract failed n=7; 3W/0T/4L; d=7; p=0.500; net -14.3%; 1 dormancy excluded 🟡 0.37 Dormancy contract: 1 unexpected activation(s); Activation: isolated 5/7; plugin 4/7 Narrow skill routing so the listed off-target scenarios stay dormant.
including-generated-files gpt-5.6-luna ➖ Not proven improved n=7; 4W/2T/1L; d=5; p=0.188; net +42.9%; 1 dormancy excluded ✅ 0.17 — Inspect tied or lost stimuli and fix inconsistent skill behavior.
incremental-build claude-sonnet-4.6 ➖ Not proven improved n=9; 3W/6T/0L; d=3; p=0.125; net +33.3% 🟡 0.33 Activation: isolated 8/9; plugin 2/9 Inspect tied or lost stimuli; predeclare added breadth before a new experiment.
incremental-build gpt-5.6-luna ➖ Not proven improved n=9; 4W/4T/1L; d=5; p=0.188; net +33.3% ✅ 0.18 Activation: isolated 9/9; plugin 8/9 Inspect tied or lost stimuli and fix inconsistent skill behavior.
item-management claude-sonnet-4.6 ⛔ Activation contract failed n=6; 4W/2T/0L; d=4; p=0.063; net +66.7%; 1 dormancy excluded 🟡 0.34 Dormancy contract: 1 unexpected activation(s); Activation: isolated 6/6; plugin 3/6 Narrow skill routing so the listed off-target scenarios stay dormant.
item-management gpt-5.6-luna ➖ Not proven improved n=6; 0W/5T/1L; d=1; p=0.500; net -16.7%; 1 dormancy excluded ✅ 0.11 — Inspect tied or lost stimuli; predeclare added breadth before a new experiment.
msbuild-antipatterns claude-sonnet-4.6 ➖ Not proven improved n=7; 1W/4T/2L; d=3; p=0.500; net -14.3%; 1 dormancy excluded 🟡 0.29 Activation: isolated 3/7; plugin 2/7 Inspect tied or lost stimuli; predeclare added breadth before a new experiment.
msbuild-antipatterns gpt-5.6-luna ➖ Not proven improved n=7; 2W/2T/3L; d=5; p=0.500; net -14.3%; 1 dormancy excluded ✅ 0.09 Activation: isolated 6/7; plugin 4/7 Inspect tied or lost stimuli and fix inconsistent skill behavior.
msbuild-modernization claude-sonnet-4.6 ➖ Not proven improved n=6; 1W/3T/2L; d=3; p=0.500; net -16.7%; 1 dormancy excluded 🟡 0.32 Activation: isolated 6/6; plugin 5/6 Inspect tied or lost stimuli; predeclare added breadth before a new experiment.
msbuild-modernization gpt-5.6-luna ➖ Not proven improved n=6; 1W/5T/0L; d=1; p=0.500; net +16.7%; 1 dormancy excluded ✅ 0.05 — Inspect tied or lost stimuli; predeclare added breadth before a new experiment.
msbuild-server claude-sonnet-4.6 ✅ Improved n=8; 8W/0T/0L; d=8; p=0.004; net +100.0% 🔴 0.52 Activation: isolated 8/8; plugin 7/8 Fix activation gaps; Review overfit evidence.
msbuild-server gpt-5.6-luna ➖ Not proven improved n=8; 5W/2T/1L; d=6; p=0.109; net +50.0% ✅ 0.15 — Inspect tied or lost stimuli and fix inconsistent skill behavior.
property-patterns claude-sonnet-4.6 ⛔ Activation contract failed n=6; 3W/2T/1L; d=4; p=0.312; net +33.3%; 1 dormancy excluded 🟡 0.26 Dormancy contract: 1 unexpected activation(s); Activation: isolated 5/6; plugin 4/6 Narrow skill routing so the listed off-target scenarios stay dormant.
property-patterns gpt-5.6-luna ➖ Not proven improved n=6; 3W/1T/2L; d=5; p=0.500; net +16.7%; 1 dormancy excluded ✅ 0.06 — Inspect tied or lost stimuli and fix inconsistent skill behavior.
resolve-project-references claude-sonnet-4.6 ✅ Improved n=6; 6W/0T/0L; d=6; p=0.016; net +100.0%; 2 dormancy excluded 🟡 0.36 Activation: isolated 6/6; plugin 5/6 Fix activation gaps; Review overfit evidence.
resolve-project-references gpt-5.6-luna ➖ Not proven improved n=6; 2W/3T/1L; d=3; p=0.500; net +16.7%; 2 dormancy excluded ✅ 0.15 Activation: isolated 4/6; plugin 5/6 Inspect tied or lost stimuli; predeclare added breadth before a new experiment.
target-authoring claude-sonnet-4.6 ➖ Not proven improved n=6; 2W/1T/3L; d=5; p=0.500; net -16.7%; 1 dormancy excluded 🟡 0.31 Activation: isolated 6/6; plugin 5/6 Inspect tied or lost stimuli and fix inconsistent skill behavior.
target-authoring gpt-5.6-luna ➖ Not proven improved n=6; 2W/3T/1L; d=3; p=0.500; net +16.7%; 1 dormancy excluded ✅ 0.10 — Inspect tied or lost stimuli; predeclare added breadth before a new experiment.
ℹ️ How to read this report
  • ✅ Improved — the result passed both the statistical gate and the 20% practical net-win floor.
  • ➖ Not proven improved — the result is valid but did not pass both gates. This is not automatically a regression.
  • ⚠️ Invalid / underpowered — the gate withheld a quality verdict. Fix the measurement before judging the skill.
  • ⛔ Activation contract failed — the isolated target skill activated on an explicit dormancy scenario. Dormancy preference is excluded, but this routing failure still blocks a pass.
  • 📉 Preference loss — the LLM judge credibly preferred baseline. It is report-only, not objective completion proof.
  • Gate evidence — n preference-eligible distinct-stimulus votes, W/T/L stimulus votes, d discordant votes, exact one-sided p, net win, and the count of separately retained dormancy stimuli. The p value applies to one model/skill result; no matrix-wide multiple-comparison correction is applied.
  • Overfit — overfitting-judge severity (✅ Low, 🟡 Moderate, 🔴 High, — none) and score.
  • Warnings — activation, timeout, retry recovery, or unresolved comparison conditions that need attention.
  • Do not add repeated runs to increase statistical power. Do not add stimuli after seeing a near-pass unless the new breadth is predeclared for a new experiment.
⛔ Activation contract failed — build-parallelism (claude-sonnet-4.6)

Why: Net win +50.0% (3W/3T/0L over 6 preference-eligible stimulus vote(s), sign test p=0.125), mean preference +22.9% across 7 paired run(s), 1 dormancy stimulus/stimuli excluded from preference — activation contract failed (1 explicit dormancy scenario(s) activated the isolated target skill)

Next action: Narrow skill routing so the listed off-target scenarios stay dormant.

State: VALID_NO_CHANGE (activation_contract_failed)

Gate evidence: n=6; 3W/3T/0L; d=3; p=0.125; net +50.0%; 1 dormancy excluded

Warnings: Dormancy contract: 1 unexpected activation(s); Activation: isolated 3/6; plugin 5/6

Overfit: Moderate (score 0.43)

Repeated-run reliability (not used by the gate): 7 paired runs (4W/3T/0L).

Weak or warning scenarios:

Scenario Preference gate Net win Δ Pref Runs (W/T/L)
= Enable BuildInParallel on a custom MSBuild task Eligible +0.0% +0.0% 0/1/0
= Enable parallel nodes for a wide project graph Eligible +0.0% +0.0% 0/1/0
▲ Non-activation: make one custom target incremental Excluded (activation contract) +100.0% +40.0% 1/0/0
▲ Preserve valid dependencies and optimize the slow critical-path project Eligible +100.0% +40.0% 1/0/0
= Reduce CI build scope with a solution filter Eligible +0.0% +0.0% 0/1/0

Illustrative judge evidence:

  • Enable BuildInParallel on a custom MSBuild task: The responses make the same correct change and communicate it with essentially equivalent clarity.

This is one example, not the aggregate verdict. Open Full Results for every judgment.

⛔ Activation contract failed — build-perf-baseline (claude-sonnet-4.6)

Why: Net win +33.3% (4W/0T/2L over 6 preference-eligible stimulus vote(s), sign test p=0.344), mean preference +5.7% across 7 paired run(s), 1 dormancy stimulus/stimuli excluded from preference — activation contract failed (1 explicit dormancy scenario(s) activated the isolated target skill)

Next action: Narrow skill routing so the listed off-target scenarios stay dormant.

State: VALID_NO_CHANGE (activation_contract_failed)

Gate evidence: n=6; 4W/0T/2L; d=6; p=0.344; net +33.3%; 1 dormancy excluded

Warnings: Dormancy contract: 1 unexpected activation(s); Activation: isolated 5/6; plugin 5/6

Overfit: Moderate (score 0.43)

Repeated-run reliability (not used by the gate): 7 paired runs (4W/0T/3L).

Weak or warning scenarios:

Scenario Preference gate Net win Δ Pref Runs (W/T/L)
▼ Configure deterministic, cache-safe CI builds Eligible -100.0% -100.0% 0/0/1
▼ Decline a non-MSBuild build performance request Excluded (activation contract) -100.0% -40.0% 0/0/1
▼ Leave an already-optimized build unchanged Eligible -100.0% -40.0% 0/0/1

Illustrative judge evidence:

  • Configure deterministic, cache-safe CI builds: A makes the essential requested, CI-guarded ContinuousIntegrationBuild change in the relevant Directory.Build.props and leaves valid configuration. Its unnecessary Deterministic addition and inaccurate default claim are flaws, but B misses the central required property and edi...

This is one example, not the aggregate verdict. Open Full Results for every judgment.

⛔ Activation contract failed — directory-build-organization (gpt-5.6-luna)

Why: Net win +33.3% (3W/2T/1L over 6 preference-eligible stimulus vote(s), sign test p=0.312), mean preference +11.4% across 7 paired run(s), 1 dormancy stimulus/stimuli excluded from preference — activation contract failed (1 explicit dormancy scenario(s) activated the isolated target skill)

Next action: Narrow skill routing so the listed off-target scenarios stay dormant.

State: VALID_NO_CHANGE (activation_contract_failed)

Gate evidence: n=6; 3W/2T/1L; d=4; p=0.312; net +33.3%; 1 dormancy excluded

Warnings: Dormancy contract: 1 unexpected activation(s)

Overfit: Low (score 0.10)

Repeated-run reliability (not used by the gate): 7 paired runs (3W/3T/1L).

Weak or warning scenarios:

Scenario Preference gate Net win Δ Pref Runs (W/T/L)
= Decline adding shared build files to a lone project Excluded (activation contract) +0.0% +0.0% 0/1/0
▼ Diagnose a package downgrade chain and reorganize version management Eligible -100.0% -40.0% 0/0/1
= Diagnose an inner shared-props file that overwrites its own override Eligible +0.0% +0.0% 0/1/0
= Preserve an intentional project-specific exception while centralizing Eligible +0.0% +0.0% 0/1/0

Illustrative judge evidence:

  • Diagnose a package downgrade chain and reorganize version management: Both responses reach the same correct conclusion and implementation. However, Response A verified its fix more thoroughly by running both dotnet restore AND dotnet test successfully, and normalized the project XML. Response B ran a build that surfaced unrelated CS1591 errors a...

This is one example, not the aggregate verdict. Open Full Results for every judgment.

⛔ Activation contract failed — including-generated-files (claude-sonnet-4.6)

Why: Net win -14.3% (3W/0T/4L over 7 preference-eligible stimulus vote(s), sign test p=0.500), mean preference -12.5% across 8 paired run(s), 1 dormancy stimulus/stimuli excluded from preference — activation contract failed (1 explicit dormancy scenario(s) activated the isolated target skill)

Next action: Narrow skill routing so the listed off-target scenarios stay dormant.

State: VALID_NO_CHANGE (activation_contract_failed)

Gate evidence: n=7; 3W/0T/4L; d=7; p=0.500; net -14.3%; 1 dormancy excluded

Warnings: Dormancy contract: 1 unexpected activation(s); Activation: isolated 5/7; plugin 4/7

Overfit: Moderate (score 0.37)

Repeated-run reliability (not used by the gate): 8 paired runs (3W/1T/4L).

Weak or warning scenarios:

Scenario Preference gate Net win Δ Pref Runs (W/T/L)
= Decline manual wiring for Roslyn source generators Excluded (activation contract) +0.0% +0.0% 0/1/0
▼ Diagnose hardcoded obj path for generated source Eligible -100.0% -100.0% 0/0/1
▲ Diagnose missing clean tracking for generated source Eligible +100.0% +40.0% 1/0/0
▼ Diagnose missing output registration for generated non-code file Eligible -100.0% -40.0% 0/0/1
▼ Diagnose project-level glob for generated source Eligible -100.0% -40.0% 0/0/1
▼ Diagnose wrong hook for generated source files Eligible -100.0% -40.0% 0/0/1

Illustrative judge evidence:

  • Diagnose hardcoded obj path for generated source: A is precise, consistent with the stated behavior, and supplies the correct reusable pattern. B's recommendation is broadly right, but its root-cause narrative is internally inconsistent and reverses the key path relationship.

This is one example, not the aggregate verdict. Open Full Results for every judgment.

⛔ Activation contract failed — item-management (claude-sonnet-4.6)

Why: Net win +66.7% (4W/2T/0L over 6 preference-eligible stimulus vote(s), sign test p=0.063), mean preference +37.1% across 7 paired run(s), 1 dormancy stimulus/stimuli excluded from preference — activation contract failed (1 explicit dormancy scenario(s) activated the isolated target skill)

Next action: Narrow skill routing so the listed off-target scenarios stay dormant.

State: VALID_NO_CHANGE (activation_contract_failed)

Gate evidence: n=6; 4W/2T/0L; d=4; p=0.063; net +66.7%; 1 dormancy excluded

Warnings: Dormancy contract: 1 unexpected activation(s); Activation: isolated 6/6; plugin 3/6

Overfit: Moderate (score 0.34)

Repeated-run reliability (not used by the gate): 7 paired runs (5W/2T/0L).

Weak or warning scenarios:

Scenario Preference gate Net win Δ Pref Runs (W/T/L)
▲ Decline a target ordering problem with no item-group defect Excluded (activation contract) +100.0% +40.0% 1/0/0
▲ Diagnose an ineffective Compile Remove that does not match the glob Eligible +100.0% +40.0% 1/0/0
▲ Diagnose real and claimed item problems in a code generation pipeline Eligible +100.0% +40.0% 1/0/0
= Fix item management anti-patterns Eligible +0.0% +0.0% 0/1/0
= Leave already-correct item management unchanged Eligible +0.0% +0.0% 0/1/0
▲ Leave correct single-list batching unchanged Eligible +100.0% +40.0% 1/0/0

Illustrative judge evidence:

  • Fix item management anti-patterns: The two final project edits implement the same four substantive corrections with no material correctness difference evident in the recorded outputs.

This is one example, not the aggregate verdict. Open Full Results for every judgment.

⛔ Activation contract failed — property-patterns (claude-sonnet-4.6)

Why: Net win +33.3% (3W/2T/1L over 6 preference-eligible stimulus vote(s), sign test p=0.312), mean preference +14.3% across 7 paired run(s), 1 dormancy stimulus/stimuli excluded from preference — activation contract failed (1 explicit dormancy scenario(s) activated the isolated target skill)

Next action: Narrow skill routing so the listed off-target scenarios stay dormant.

State: VALID_NO_CHANGE (activation_contract_failed)

Gate evidence: n=6; 3W/2T/1L; d=4; p=0.312; net +33.3%; 1 dormancy excluded

Warnings: Dormancy contract: 1 unexpected activation(s); Activation: isolated 5/6; plugin 4/6

Overfit: Moderate (score 0.26)

Repeated-run reliability (not used by the gate): 7 paired runs (3W/2T/2L).

Weak or warning scenarios:

Scenario Preference gate Net win Δ Pref Runs (W/T/L)
▼ Choose the right OS detection for a cross-platform property Eligible -100.0% -40.0% 0/0/1
▼ Decline a props-versus-targets placement question Excluded (activation contract) -100.0% -40.0% 0/0/1
▲ Diagnose a TargetFramework condition that never applies in props Eligible +100.0% +40.0% 1/0/0
= Diagnose multi-level property hierarchy bugs Eligible +0.0% +0.0% 0/1/0
= Fix shared property configuration Eligible +0.0% +0.0% 0/1/0

Illustrative judge evidence:

  • Choose the right OS detection for a cross-platform property: Both responses correctly identify the defect and propose explicit Windows/macOS/Linux detection. A is slightly clearer and more directly actionable by naming the MSBuild intrinsic and its required platform names.

This is one example, not the aggregate verdict. Open Full Results for every judgment.

➖ Not proven improved — binlog-failure-analysis (claude-sonnet-4.6)

Why: Net win +57.1% (5W/1T/1L over 7 preference-eligible stimulus vote(s), sign test p=0.109), mean preference +42.5% across 8 paired run(s), 1 dormancy stimulus/stimuli excluded from preference — not credible (sign test p=0.109 > 0.05)

Next action: Inspect tied or lost stimuli and fix inconsistent skill behavior.

State: VALID_NO_CHANGE (no_credible_preference_change)

Gate evidence: n=7; 5W/1T/1L; d=6; p=0.109; net +57.1%; 1 dormancy excluded

Overfit: Moderate (score 0.24)

Repeated-run reliability (not used by the gate): 8 paired runs (5W/2T/1L).

Weak or warning scenarios:

Scenario Preference gate Net win Δ Pref Runs (W/T/L)
▼ Determine whether a quiet second build actually failed Eligible -100.0% -40.0% 0/0/1
= Diagnose a warning behind a build that actually succeeded Eligible +0.0% +0.0% 0/1/0
= Stay dormant for a non-MSBuild build failure log Excluded (activation contract) +0.0% +0.0% 0/1/0

Illustrative judge evidence:

  • Determine whether a quiet second build actually failed: Both answers are correct, concise, and address the user's concern. A is marginally stronger because it presents the key skip/up-to-date evidence directly and adds practical ways to force recompilation.

This is one example, not the aggregate verdict. Open Full Results for every judgment.

➖ Not proven improved — binlog-failure-analysis (gpt-5.6-luna)

Why: Net win +28.6% (4W/1T/2L over 7 preference-eligible stimulus vote(s), sign test p=0.344), mean preference +2.5% across 8 paired run(s), 1 dormancy stimulus/stimuli excluded from preference — not credible (sign test p=0.344 > 0.05)

Next action: Inspect tied or lost stimuli and fix inconsistent skill behavior.

State: VALID_NO_CHANGE (no_credible_preference_change)

Gate evidence: n=7; 4W/1T/2L; d=6; p=0.344; net +28.6%; 1 dormancy excluded

Overfit: Low (score 0.10)

Repeated-run reliability (not used by the gate): 8 paired runs (4W/2T/2L).

Weak or warning scenarios:

Scenario Preference gate Net win Δ Pref Runs (W/T/L)
▼ Assess a requested binlog investigation when the current build is healthy Eligible -100.0% -100.0% 0/0/1
= Confirm the actual resolved target framework and package version from a binlog Eligible +0.0% +0.0% 0/1/0
▼ Fall back to command-line log replay when the usual binlog tool is unavailable Eligible -100.0% -40.0% 0/0/1
= Stay dormant for a non-MSBuild build failure log Excluded (activation contract) +0.0% +0.0% 0/1/0

Illustrative judge evidence:

  • Assess a requested binlog investigation when the current build is healthy: The staged project builds successfully; there is no genuine failure to diagnose. Response A captured a binlog and correctly concluded the build succeeds with 0 warnings/errors, noting only benign SDK-environment diagnostics. Response B, despite using the dedicated skill, manuf...

This is one example, not the aggregate verdict. Open Full Results for every judgment.

➖ Not proven improved — binlog-generation (gpt-5.6-luna)

Why: Net win +57.1% (4W/3T/0L over 7 preference-eligible stimulus vote(s), sign test p=0.063), mean preference +37.5% across 8 paired run(s), 1 dormancy stimulus/stimuli excluded from preference — not credible — 3 of 7 preference-eligible stimulus vote(s) tied, leaving only 4 discordant preference vote(s). The sign test conditions on non-tie stimulus votes and cannot reach 0.05 below 5, so no record could have passed here — this is not a measured null. Either the skill is inert on these scenarios (make them discriminate) or the eval needs more distinct stimuli to clear the ties

Next action: Inspect tied or lost stimuli; predeclare added breadth before a new experiment.

State: VALID_NO_CHANGE (no_credible_preference_change)

Gate evidence: n=7; 4W/3T/0L; d=4; p=0.063; net +57.1%; 1 dormancy excluded

Overfit: Moderate (score 0.28)

Repeated-run reliability (not used by the gate): 8 paired runs (4W/3T/1L).

Weak or warning scenarios:

Scenario Preference gate Net win Δ Pref Runs (W/T/L)
= Build with /bl in PowerShell Eligible +0.0% +0.0% 0/1/0
= Choose a predictable, non-colliding binlog name for a CI upload step Eligible +0.0% +0.0% 0/1/0
▼ Decline a binlog request for a non-MSBuild Java build Excluded (activation contract) -100.0% -40.0% 0/0/1
= Recognize that a failed build produced no binlog at all Eligible +0.0% +0.0% 0/1/0

Illustrative judge evidence:

  • Build with /bl in PowerShell: Position-swap inconsistent (forward: tie, reverse: A). Defaulting to tie.

This is one example, not the aggregate verdict. Open Full Results for every judgment.

➖ Not proven improved — build-parallelism (gpt-5.6-luna)

Why: Net win -50.0% (0W/3T/3L over 6 preference-eligible stimulus vote(s), sign test p=0.125), mean preference -25.7% across 7 paired run(s), 1 dormancy stimulus/stimuli excluded from preference — no improvement

Next action: Inspect tied or lost stimuli; predeclare added breadth before a new experiment.

State: VALID_NO_CHANGE (no_credible_preference_change)

Gate evidence: n=6; 0W/3T/3L; d=3; p=0.125; net -50.0%; 1 dormancy excluded

Warnings: Activation: isolated 5/6; plugin 6/6

Overfit: Low (score 0.11)

Repeated-run reliability (not used by the gate): 7 paired runs (0W/4T/3L).

Weak or warning scenarios:

Scenario Preference gate Net win Δ Pref Runs (W/T/L)
▼ Analyze build parallelism bottlenecks Eligible -100.0% -40.0% 0/0/1
= Decline graph build for runtime-discovered projects Eligible +0.0% +0.0% 0/1/0
= Enable BuildInParallel on a custom MSBuild task Eligible +0.0% +0.0% 0/1/0
= Enable parallel nodes for a wide project graph Eligible +0.0% +0.0% 0/1/0
= Non-activation: make one custom target incremental Excluded (activation contract) +0.0% +0.0% 0/1/0
▼ Preserve valid dependencies and optimize the slow critical-path project Eligible -100.0% -100.0% 0/0/1
▼ Reduce CI build scope with a solution filter Eligible -100.0% -40.0% 0/0/1

Illustrative judge evidence:

  • Analyze build parallelism bottlenecks: Both responses reach the same correct conclusions on all substantive rubric points (serial chain, minimum build time, redundant Tests->Api reference, graph-hygiene-only). They are essentially tied. A edges ahead slightly because its evidence (per-project performance summary wi...

This is one example, not the aggregate verdict. Open Full Results for every judgment.

➖ Not proven improved — build-perf-baseline (gpt-5.6-luna)

Why: Net win +33.3% (4W/0T/2L over 6 preference-eligible stimulus vote(s), sign test p=0.344), mean preference +25.7% across 7 paired run(s), 1 dormancy stimulus/stimuli excluded from preference — not credible (sign test p=0.344 > 0.05)

Next action: Inspect tied or lost stimuli and fix inconsistent skill behavior.

State: VALID_NO_CHANGE (no_credible_preference_change)

Gate evidence: n=6; 4W/0T/2L; d=6; p=0.344; net +33.3%; 1 dormancy excluded

Warnings: Activation: isolated 6/6; plugin 5/6

Overfit: Moderate (score 0.23)

Repeated-run reliability (not used by the gate): 7 paired runs (5W/0T/2L).

Weak or warning scenarios:

Scenario Preference gate Net win Δ Pref Runs (W/T/L)
▲ Configure deterministic, cache-safe CI builds Eligible +100.0% +100.0% 1/0/0
▼ Leave an already-optimized build unchanged Eligible -100.0% -40.0% 0/0/1
▼ Route a broken no-op rebuild away from generic optimization Eligible -100.0% -40.0% 0/0/1

Illustrative judge evidence:

  • Leave an already-optimized build unchanged: The ideal answer was to conclude no changes were needed since the project is already optimized (minimal refs, CI-conditioned docs). Both responses failed this by inventing a change. However, B committed exactly the anti-pattern the rubric explicitly calls out (adding UseArtifa...

This is one example, not the aggregate verdict. Open Full Results for every judgment.

➖ Not proven improved — build-perf-diagnostics (claude-sonnet-4.6)

Why: Net win +0.0% (3W/1T/3L over 7 preference-eligible stimulus vote(s), sign test p=0.656), mean preference -20.0% across 8 paired run(s), 1 dormancy stimulus/stimuli excluded from preference — no improvement

Next action: Inspect tied or lost stimuli and fix inconsistent skill behavior.

State: VALID_NO_CHANGE (no_credible_preference_change)

Gate evidence: n=7; 3W/1T/3L; d=6; p=0.656; net +0.0%; 1 dormancy excluded

Warnings: Activation: isolated 6/7; plugin 7/7

Overfit: Moderate (score 0.45)

Repeated-run reliability (not used by the gate): 8 paired runs (3W/1T/4L).

Weak or warning scenarios:

Scenario Preference gate Net win Δ Pref Runs (W/T/L)
▼ Decline a runtime latency request that is not a build performance issue Excluded (activation contract) -100.0% -100.0% 0/0/1
▼ Diagnose NuGet restore running redundantly across CI stages Eligible -100.0% -100.0% 0/0/1
▼ Diagnose a single custom target dominating one project's build Eligible -100.0% -40.0% 0/0/1
▼ Diagnose evaluation overhead before any target runs Eligible -100.0% -40.0% 0/0/1
= Diagnose slow build for a small project Eligible +0.0% +0.0% 0/1/0

Illustrative judge evidence:

  • Diagnose NuGet restore running redundantly across CI stages: A correctly diagnoses the repeated restore cost and supplies all requested concrete CI and MSBuild changes; B fails to locate the available file and produces no substantive answer.

This is one example, not the aggregate verdict. Open Full Results for every judgment.

➖ Not proven improved — build-perf-diagnostics (gpt-5.6-luna)

Why: Net win +28.6% (3W/3T/1L over 7 preference-eligible stimulus vote(s), sign test p=0.312), mean preference +10.0% across 8 paired run(s), 1 dormancy stimulus/stimuli excluded from preference — not credible — 3 of 7 preference-eligible stimulus vote(s) tied, leaving only 4 discordant preference vote(s). The sign test conditions on non-tie stimulus votes and cannot reach 0.05 below 5, so no record could have passed here — this is not a measured null. Either the skill is inert on these scenarios (make them discriminate) or the eval needs more distinct stimuli to clear the ties

Next action: Inspect tied or lost stimuli; predeclare added breadth before a new experiment.

State: VALID_NO_CHANGE (no_credible_preference_change)

Gate evidence: n=7; 3W/3T/1L; d=4; p=0.312; net +28.6%; 1 dormancy excluded

Overfit: Low (score 0.14)

Repeated-run reliability (not used by the gate): 8 paired runs (3W/4T/1L).

Weak or warning scenarios:

Scenario Preference gate Net win Δ Pref Runs (W/T/L)
= Decline a runtime latency request that is not a build performance issue Excluded (activation contract) +0.0% +0.0% 0/1/0
= Diagnose NuGet restore running redundantly across CI stages Eligible +0.0% +0.0% 0/1/0
▼ Diagnose a pathological ResolveAssemblyReference time Eligible -100.0% -40.0% 0/0/1
= Diagnose per-project overhead across many small projects Eligible +0.0% +0.0% 0/1/0
= Diagnose slow build for a small project Eligible +0.0% +0.0% 0/1/0

Illustrative judge evidence:

  • Diagnose NuGet restore running redundantly across CI stages: Position-swap inconsistent (forward: B, reverse: tie). Defaulting to tie.

This is one example, not the aggregate verdict. Open Full Results for every judgment.

➖ Not proven improved — check-bin-obj-clash (claude-sonnet-4.6)

Why: Net win +28.6% (4W/1T/2L over 7 preference-eligible stimulus vote(s), sign test p=0.344), mean preference +20.0% across 7 paired run(s) — not credible (sign test p=0.344 > 0.05)

Next action: Inspect tied or lost stimuli and fix inconsistent skill behavior.

State: VALID_NO_CHANGE (no_credible_preference_change)

Gate evidence: n=7; 4W/1T/2L; d=6; p=0.344; net +28.6%

Warnings: Activation: isolated 6/7; plugin 6/7

Overfit: Low (score 0.19)

Repeated-run reliability (not used by the gate): 7 paired runs (4W/1T/2L).

Weak or warning scenarios:

Scenario Preference gate Net win Δ Pref Runs (W/T/L)
▼ Audit the mixed solution and separate safe projects from unsafe ones Eligible -100.0% -40.0% 0/0/1
▼ Diagnose multi-targeting outputs that collapse into one path Eligible -100.0% -40.0% 0/0/1
= Fix all clash mechanisms in the mixed solution Eligible +0.0% +0.0% 0/1/0

Illustrative judge evidence:

  • Audit the mixed solution and separate safe projects from unsafe ones: The audits are substantively very similar and correctly classify the risks. A is slightly stronger overall because its LibraryA/LibraryB remediation actually separates both colliding output and intermediate directories; B's fix addresses only obj and its "Unsafe Projects (3)" ...

This is one example, not the aggregate verdict. Open Full Results for every judgment.

➖ Not proven improved — check-bin-obj-clash (gpt-5.6-luna)

Why: Net win -14.3% (0W/6T/1L over 7 preference-eligible stimulus vote(s), sign test p=0.500), mean preference -5.7% across 7 paired run(s) — no improvement

Next action: Inspect tied or lost stimuli; predeclare added breadth before a new experiment.

State: VALID_NO_CHANGE (no_credible_preference_change)

Gate evidence: n=7; 0W/6T/1L; d=1; p=0.500; net -14.3%

Overfit: Low (score 0.13)

Repeated-run reliability (not used by the gate): 7 paired runs (0W/6T/1L).

Weak or warning scenarios:

Scenario Preference gate Net win Δ Pref Runs (W/T/L)
▼ Audit the mixed solution and separate safe projects from unsafe ones Eligible -100.0% -40.0% 0/0/1
= Avoid a false clash report when projects share only the top-level artifacts root Eligible +0.0% +0.0% 0/1/0
= Decline output-clash remediation for separate projects using the default SDK layout Eligible +0.0% +0.0% 0/1/0
= Diagnose multi-targeting outputs that collapse into one path Eligible +0.0% +0.0% 0/1/0
= Diagnose redundant project reference metadata that forks a same-path build Eligible +0.0% +0.0% 0/1/0
= Diagnose shared output and intermediate path collision Eligible +0.0% +0.0% 0/1/0
= Fix all clash mechanisms in the mixed solution Eligible +0.0% +0.0% 0/1/0

Illustrative judge evidence:

  • Audit the mixed solution and separate safe projects from unsafe ones: Both responses reach substantively correct and thorough conclusions with the same technical verification (msbuild property inspection, multi-target checks). The key differentiator is causal attribution: the rubric wants ConsumerApp flagged as unsafe (its reference metadata is ...

This is one example, not the aggregate verdict. Open Full Results for every judgment.

➖ Not proven improved — directory-build-organization (claude-sonnet-4.6)

Why: Net win +33.3% (3W/2T/1L over 6 preference-eligible stimulus vote(s), sign test p=0.312), mean preference +5.7% across 7 paired run(s), 1 dormancy stimulus/stimuli excluded from preference — not credible — 2 of 6 preference-eligible stimulus vote(s) tied, leaving only 4 discordant preference vote(s). The sign test conditions on non-tie stimulus votes and cannot reach 0.05 below 5, so no record could have passed here — this is not a measured null. Either the skill is inert on these scenarios (make them discriminate) or the eval needs more distinct stimuli to clear the ties

Next action: Inspect tied or lost stimuli; predeclare added breadth before a new experiment.

State: VALID_NO_CHANGE (no_credible_preference_change)

Gate evidence: n=6; 3W/2T/1L; d=4; p=0.312; net +33.3%; 1 dormancy excluded

Warnings: Activation: isolated 5/6; plugin 5/6

Overfit: Moderate (score 0.33)

Repeated-run reliability (not used by the gate): 7 paired runs (3W/2T/2L).

Weak or warning scenarios:

Scenario Preference gate Net win Δ Pref Runs (W/T/L)
▼ Apply repo-level build organization cleanup Eligible -100.0% -100.0% 0/0/1
▼ Decline adding shared build files to a lone project Excluded (activation contract) -100.0% -40.0% 0/0/1
= Diagnose an inner shared-props file that overwrites its own override Eligible +0.0% +0.0% 0/1/0
= Preserve an intentional project-specific exception while centralizing Eligible +0.0% +0.0% 0/1/0

Illustrative judge evidence:

  • Apply repo-level build organization cleanup: A completed the requested build-layout refactor and verified the created shared files, while B only reported an unavailable path and delivered no task result.

This is one example, not the aggregate verdict. Open Full Results for every judgment.

➖ Not proven improved — eval-performance (claude-sonnet-4.6)

Why: Net win +62.5% (6W/1T/1L over 8 preference-eligible stimulus vote(s), sign test p=0.063), mean preference +40.0% across 8 paired run(s) — not credible (sign test p=0.063 > 0.05)

Next action: Inspect tied or lost stimuli and fix inconsistent skill behavior.

State: VALID_NO_CHANGE (no_credible_preference_change)

Gate evidence: n=8; 6W/1T/1L; d=7; p=0.063; net +62.5%

Warnings: Activation: isolated 8/8; plugin 7/8

Overfit: Moderate (score 0.27)

Repeated-run reliability (not used by the gate): 8 paired runs (6W/1T/1L).

Weak or warning scenarios:

Scenario Preference gate Net win Δ Pref Runs (W/T/L)
= Redirect a compile-time slowdown mistakenly framed as an evaluation problem Eligible +0.0% +0.0% 0/1/0
▼ Triage which of two property functions actually costs evaluation time Eligible -100.0% -40.0% 0/0/1

Illustrative judge evidence:

  • Redirect a compile-time slowdown mistakenly framed as an evaluation problem: Position-swap inconsistent (forward: A, reverse: B). Defaulting to tie.

This is one example, not the aggregate verdict. Open Full Results for every judgment.

➖ Not proven improved — eval-performance (gpt-5.6-luna)

Why: Net win +50.0% (5W/2T/1L over 8 preference-eligible stimulus vote(s), sign test p=0.109), mean preference +35.0% across 8 paired run(s) — not credible (sign test p=0.109 > 0.05)

Next action: Inspect tied or lost stimuli and fix inconsistent skill behavior.

State: VALID_NO_CHANGE (no_credible_preference_change)

Gate evidence: n=8; 5W/2T/1L; d=6; p=0.109; net +50.0%

Overfit: Low (score 0.09)

Repeated-run reliability (not used by the gate): 8 paired runs (5W/2T/1L).

Weak or warning scenarios:

Scenario Preference gate Net win Δ Pref Runs (W/T/L)
= Detect a project evaluated twice under different global properties Eligible +0.0% +0.0% 0/1/0
▼ Diagnose stacked evaluation-time patterns (deep imports, broad glob, file-I/O property function) Eligible -100.0% -40.0% 0/0/1
= Triage which of two property functions actually costs evaluation time Eligible +0.0% +0.0% 0/1/0

Illustrative judge evidence:

  • Detect a project evaluated twice under different global properties: Position-swap inconsistent (forward: B, reverse: A). Defaulting to tie.

This is one example, not the aggregate verdict. Open Full Results for every judgment.

➖ Not proven improved — extension-points (claude-sonnet-4.6)

Why: Net win +14.3% (3W/2T/2L over 7 preference-eligible stimulus vote(s), sign test p=0.500), mean preference +5.0% across 8 paired run(s), 1 dormancy stimulus/stimuli excluded from preference — not credible (sign test p=0.500 > 0.05)

Next action: Inspect tied or lost stimuli and fix inconsistent skill behavior.

State: VALID_NO_CHANGE (no_credible_preference_change)

Gate evidence: n=7; 3W/2T/2L; d=5; p=0.500; net +14.3%; 1 dormancy excluded

Warnings: Activation: isolated 4/7; plugin 5/7

Overfit: Moderate (score 0.26)

Repeated-run reliability (not used by the gate): 8 paired runs (3W/3T/2L).

Weak or warning scenarios:

Scenario Preference gate Net win Δ Pref Runs (W/T/L)
▲ Diagnose a broken per-TFM forwarder Eligible +100.0% +40.0% 1/0/0
▼ Diagnose a package ID and file-name mismatch Eligible -100.0% -40.0% 0/0/1
▼ Diagnose build extension point failures Eligible -100.0% -100.0% 0/0/1
= Fix extension point anti-patterns Eligible +0.0% +0.0% 0/1/0
= Non-activation: repair an incremental custom target Excluded (activation contract) +0.0% +0.0% 0/1/0
= Review packed layout without a false missing-file bug Eligible +0.0% +0.0% 0/1/0

Illustrative judge evidence:

  • Diagnose a package ID and file-name mismatch: Both are correct, concise, and directly address the silent convention-based import failure. A is marginally stronger as a standalone final result because it includes exact corrected filenames, nuspec entries, and the necessary repackaging step; B's claimed applied fix is usefu...

This is one example, not the aggregate verdict. Open Full Results for every judgment.

➖ Not proven improved — extension-points (gpt-5.6-luna)

Why: Net win +57.1% (4W/3T/0L over 7 preference-eligible stimulus vote(s), sign test p=0.063), mean preference +22.5% across 8 paired run(s), 1 dormancy stimulus/stimuli excluded from preference — not credible — 3 of 7 preference-eligible stimulus vote(s) tied, leaving only 4 discordant preference vote(s). The sign test conditions on non-tie stimulus votes and cannot reach 0.05 below 5, so no record could have passed here — this is not a measured null. Either the skill is inert on these scenarios (make them discriminate) or the eval needs more distinct stimuli to clear the ties

Next action: Inspect tied or lost stimuli; predeclare added breadth before a new experiment.

State: VALID_NO_CHANGE (no_credible_preference_change)

Gate evidence: n=7; 4W/3T/0L; d=4; p=0.063; net +57.1%; 1 dormancy excluded

Overfit: Moderate (score 0.24)

Repeated-run reliability (not used by the gate): 8 paired runs (4W/3T/1L).

Weak or warning scenarios:

Scenario Preference gate Net win Δ Pref Runs (W/T/L)
= Diagnose build extension point failures Eligible +0.0% +0.0% 0/1/0
= Fix extension point anti-patterns Eligible +0.0% +0.0% 0/1/0
▼ Non-activation: repair an incremental custom target Excluded (activation contract) -100.0% -40.0% 0/0/1
= Review packed layout without a false missing-file bug Eligible +0.0% +0.0% 0/1/0

Illustrative judge evidence:

  • Diagnose build extension point failures: Position-swap inconsistent (forward: B, reverse: A). Defaulting to tie.

This is one example, not the aggregate verdict. Open Full Results for every judgment.

➖ Not proven improved — including-generated-files (gpt-5.6-luna)

Why: Net win +42.9% (4W/2T/1L over 7 preference-eligible stimulus vote(s), sign test p=0.188), mean preference +15.0% across 8 paired run(s), 1 dormancy stimulus/stimuli excluded from preference — not credible (sign test p=0.188 > 0.05)

Next action: Inspect tied or lost stimuli and fix inconsistent skill behavior.

State: VALID_NO_CHANGE (no_credible_preference_change)

Gate evidence: n=7; 4W/2T/1L; d=5; p=0.188; net +42.9%; 1 dormancy excluded

Overfit: Low (score 0.17)

Repeated-run reliability (not used by the gate): 8 paired runs (4W/3T/1L).

Weak or warning scenarios:

Scenario Preference gate Net win Δ Pref Runs (W/T/L)
= Decline manual wiring for Roslyn source generators Excluded (activation contract) +0.0% +0.0% 0/1/0
= Diagnose missing clean tracking for generated source Eligible +0.0% +0.0% 0/1/0
= Diagnose missing generated source inclusion Eligible +0.0% +0.0% 0/1/0
▼ Diagnose missing output registration for generated non-code file Eligible -100.0% -40.0% 0/0/1

Illustrative judge evidence:

  • Diagnose missing clean tracking for generated source: Position-swap inconsistent (forward: tie, reverse: A). Defaulting to tie.

This is one example, not the aggregate verdict. Open Full Results for every judgment.

➖ Not proven improved — incremental-build (claude-sonnet-4.6)

Why: Net win +33.3% (3W/6T/0L over 9 preference-eligible stimulus vote(s), sign test p=0.125), mean preference +13.3% across 9 paired run(s) — not credible — 6 of 9 preference-eligible stimulus vote(s) tied, leaving only 3 discordant preference vote(s). The sign test conditions on non-tie stimulus votes and cannot reach 0.05 below 5, so no record could have passed here — this is not a measured null. Either the skill is inert on these scenarios (make them discriminate) or the eval needs more distinct stimuli to clear the ties

Next action: Inspect tied or lost stimuli; predeclare added breadth before a new experiment.

State: VALID_NO_CHANGE (no_credible_preference_change)

Gate evidence: n=9; 3W/6T/0L; d=3; p=0.125; net +33.3%

Warnings: Activation: isolated 8/9; plugin 2/9

Overfit: Moderate (score 0.33)

Repeated-run reliability (not used by the gate): 9 paired runs (3W/6T/0L).

Weak or warning scenarios:

Scenario Preference gate Net win Δ Pref Runs (W/T/L)
▲ Correct the assumption that Outputs alone enables incremental skipping Eligible +100.0% +40.0% 1/0/0
= Diagnose custom targets that always rerun Eligible +0.0% +0.0% 0/1/0
▲ Explain slower builds when MSBuild skipped everything Eligible +100.0% +40.0% 1/0/0
= Explain why Visual Studio keeps rebuilding an up-to-date project Eligible +0.0% +0.0% 0/1/0
= Explain why clean leaves generated hash source behind Eligible +0.0% +0.0% 0/1/0
= Fix broken incremental targets and clean tracking Eligible +0.0% +0.0% 0/1/0
= Identify a volatile output path that defeats incrementality Eligible +0.0% +0.0% 0/1/0
= Read a diagnostic log to find the stale input Eligible +0.0% +0.0% 0/1/0

Illustrative judge evidence:

  • Diagnose custom targets that always rerun: Position-swap inconsistent (forward: B, reverse: A). Defaulting to tie.

This is one example, not the aggregate verdict. Open Full Results for every judgment.

➖ Not proven improved — incremental-build (gpt-5.6-luna)

Why: Net win +33.3% (4W/4T/1L over 9 preference-eligible stimulus vote(s), sign test p=0.188), mean preference +13.3% across 9 paired run(s) — not credible (sign test p=0.188 > 0.05)

Next action: Inspect tied or lost stimuli and fix inconsistent skill behavior.

State: VALID_NO_CHANGE (no_credible_preference_change)

Gate evidence: n=9; 4W/4T/1L; d=5; p=0.188; net +33.3%

Warnings: Activation: isolated 9/9; plugin 8/9

Overfit: Low (score 0.18)

Repeated-run reliability (not used by the gate): 9 paired runs (4W/4T/1L).

Weak or warning scenarios:

Scenario Preference gate Net win Δ Pref Runs (W/T/L)
▼ Diagnose custom targets that always rerun Eligible -100.0% -40.0% 0/0/1
= Distinguish a cold first build from broken incrementality Eligible +0.0% +0.0% 0/1/0
▲ Explain slower builds when MSBuild skipped everything Eligible +100.0% +40.0% 1/0/0
= Fix broken incremental targets and clean tracking Eligible +0.0% +0.0% 0/1/0
= Identify a volatile output path that defeats incrementality Eligible +0.0% +0.0% 0/1/0
= Read a diagnostic log to find the stale input Eligible +0.0% +0.0% 0/1/0

Illustrative judge evidence:

  • Diagnose custom targets that always rerun: Both responses correctly diagnose the problem and identify the two targets. A is slightly better because it provides concrete corrected XML with explicit Inputs/Outputs, making the fix actionable, while B stays at a prose description.

This is one example, not the aggregate verdict. Open Full Results for every judgment.

➖ Not proven improved — item-management (gpt-5.6-luna)

Why: Net win -16.7% (0W/5T/1L over 6 preference-eligible stimulus vote(s), sign test p=0.500), mean preference -5.7% across 7 paired run(s), 1 dormancy stimulus/stimuli excluded from preference — no improvement

Next action: Inspect tied or lost stimuli; predeclare added breadth before a new experiment.

State: VALID_NO_CHANGE (no_credible_preference_change)

Gate evidence: n=6; 0W/5T/1L; d=1; p=0.500; net -16.7%; 1 dormancy excluded

Overfit: Low (score 0.11)

Repeated-run reliability (not used by the gate): 7 paired runs (0W/6T/1L).

Weak or warning scenarios:

Scenario Preference gate Net win Δ Pref Runs (W/T/L)
= Decline a target ordering problem with no item-group defect Excluded (activation contract) +0.0% +0.0% 0/1/0
= Diagnose an ineffective Compile Remove that does not match the glob Eligible +0.0% +0.0% 0/1/0
= Diagnose item group and batching issues Eligible +0.0% +0.0% 0/1/0
▼ Diagnose real and claimed item problems in a code generation pipeline Eligible -100.0% -40.0% 0/0/1
= Fix item management anti-patterns Eligible +0.0% +0.0% 0/1/0
= Leave already-correct item management unchanged Eligible +0.0% +0.0% 0/1/0
= Leave correct single-list batching unchanged Eligible +0.0% +0.0% 0/1/0

Illustrative judge evidence:

  • Diagnose an ineffective Compile Remove that does not match the glob: Position-swap inconsistent (forward: B, reverse: tie). Defaulting to tie.

This is one example, not the aggregate verdict. Open Full Results for every judgment.

➖ Not proven improved — msbuild-antipatterns (claude-sonnet-4.6)

Why: Net win -14.3% (1W/4T/2L over 7 preference-eligible stimulus vote(s), sign test p=0.500), mean preference -17.5% across 8 paired run(s), 1 dormancy stimulus/stimuli excluded from preference — no improvement

Next action: Inspect tied or lost stimuli; predeclare added breadth before a new experiment.

State: VALID_NO_CHANGE (no_credible_preference_change)

Gate evidence: n=7; 1W/4T/2L; d=3; p=0.500; net -14.3%; 1 dormancy excluded

Warnings: Activation: isolated 3/7; plugin 2/7

Overfit: Moderate (score 0.29)

Repeated-run reliability (not used by the gate): 8 paired runs (1W/4T/3L).

Weak or warning scenarios:

Scenario Preference gate Net win Δ Pref Runs (W/T/L)
= Add a module to an F# project Eligible +0.0% +0.0% 0/1/0
= Add a signature file to define public API Eligible +0.0% +0.0% 0/1/0
▼ Distinguish a style backslash from a real cross-platform backslash bug Eligible -100.0% -100.0% 0/0/1
▼ Fix broken file order causing FS0039 Eligible -100.0% -40.0% 0/0/1
= Judge an unguarded import inside a NuGet package build folder Eligible +0.0% +0.0% 0/1/0
= Leave a clean project without inventing anti-patterns Eligible +0.0% +0.0% 0/1/0
▼ Non-activation: migrate a legacy project to SDK style Excluded (activation contract) -100.0% -40.0% 0/0/1

Illustrative judge evidence:

  • Add a module to an F# project: Both runs implement the requested functionality correctly, include the new file in valid F# compilation order, update processing flow for validation failures, and successfully build/run the application. B's placement directly after Domain.fs is marginally tidier, but A's place...

This is one example, not the aggregate verdict. Open Full Results for every judgment.

➖ Not proven improved — msbuild-antipatterns (gpt-5.6-luna)

Why: Net win -14.3% (2W/2T/3L over 7 preference-eligible stimulus vote(s), sign test p=0.500), mean preference -2.5% across 8 paired run(s), 1 dormancy stimulus/stimuli excluded from preference — no improvement

Next action: Inspect tied or lost stimuli and fix inconsistent skill behavior.

State: VALID_NO_CHANGE (no_credible_preference_change)

Gate evidence: n=7; 2W/2T/3L; d=5; p=0.500; net -14.3%; 1 dormancy excluded

Warnings: Activation: isolated 6/7; plugin 4/7

Overfit: Low (score 0.09)

Repeated-run reliability (not used by the gate): 8 paired runs (2W/2T/4L).

Weak or warning scenarios:

Scenario Preference gate Net win Δ Pref Runs (W/T/L)
▲ Add a module to an F# project Eligible +100.0% +40.0% 1/0/0
= Add a signature file to define public API Eligible +0.0% +0.0% 0/1/0
= Distinguish a style backslash from a real cross-platform backslash bug Eligible +0.0% +0.0% 0/1/0
▼ Fix broken file order causing FS0039 Eligible -100.0% -40.0% 0/0/1
▼ Judge an unguarded import inside a NuGet package build folder Eligible -100.0% -40.0% 0/0/1
▼ Non-activation: migrate a legacy project to SDK style Excluded (activation contract) -100.0% -40.0% 0/0/1
▼ Review MSBuild files for anti-patterns and style issues Eligible -100.0% -40.0% 0/0/1

Illustrative judge evidence:

  • Add a signature file to define public API: Both responses produced functionally identical, correct results: an appropriate Domain.fsi with all types, placed before Domain.fs in the project file, with a verified clean build. Their approaches, recovery from the apply_patch failure, and final verification were essentially...

This is one example, not the aggregate verdict. Open Full Results for every judgment.

➖ Not proven improved — msbuild-modernization (claude-sonnet-4.6)

Why: Net win -16.7% (1W/3T/2L over 6 preference-eligible stimulus vote(s), sign test p=0.500), mean preference +2.9% across 7 paired run(s), 1 dormancy stimulus/stimuli excluded from preference — no improvement

Next action: Inspect tied or lost stimuli; predeclare added breadth before a new experiment.

State: VALID_NO_CHANGE (no_credible_preference_change)

Gate evidence: n=6; 1W/3T/2L; d=3; p=0.500; net -16.7%; 1 dormancy excluded

Warnings: Activation: isolated 6/6; plugin 5/6

Overfit: Moderate (score 0.32)

Repeated-run reliability (not used by the gate): 7 paired runs (1W/4T/2L).

Weak or warning scenarios:

Scenario Preference gate Net win Δ Pref Runs (W/T/L)
▼ Apply the migration to SDK-style Eligible -100.0% -40.0% 0/0/1
= Consolidate duplicated projects into a multi-targeting SDK-style project Eligible +0.0% +0.0% 0/1/0
= Decline modernizing a non-.NET build Excluded (activation contract) +0.0% +0.0% 0/1/0
▼ Identify legacy patterns for SDK-style migration Eligible -100.0% -40.0% 0/0/1
= Modernize a single project without introducing Central Package Management Eligible +0.0% +0.0% 0/1/0
▲ Modernize further without introducing a nondeterministic language version Eligible +100.0% +100.0% 1/0/0
= Recognize an already-modern project needs no migration Eligible +0.0% +0.0% 0/1/0

Illustrative judge evidence:

  • Apply the migration to SDK-style: Both conversions satisfy the core SDK-style migration and should avoid duplicate assembly attributes. A is marginally safer and less behavior-changing because it preserves the complete existing AssemblyInfo source rather than relying on a manual migration of its attributes.

This is one example, not the aggregate verdict. Open Full Results for every judgment.

➖ Not proven improved — msbuild-modernization (gpt-5.6-luna)

Why: Net win +16.7% (1W/5T/0L over 6 preference-eligible stimulus vote(s), sign test p=0.500), mean preference +0.0% across 7 paired run(s), 1 dormancy stimulus/stimuli excluded from preference — not credible — 5 of 6 preference-eligible stimulus vote(s) tied, leaving only 1 discordant preference vote(s). The sign test conditions on non-tie stimulus votes and cannot reach 0.05 below 5, so no record could have passed here — this is not a measured null. Either the skill is inert on these scenarios (make them discriminate) or the eval needs more distinct stimuli to clear the ties

Next action: Inspect tied or lost stimuli; predeclare added breadth before a new experiment.

State: VALID_NO_CHANGE (no_credible_preference_change)

Gate evidence: n=6; 1W/5T/0L; d=1; p=0.500; net +16.7%; 1 dormancy excluded

Overfit: Low (score 0.05)

Repeated-run reliability (not used by the gate): 7 paired runs (1W/5T/1L).

Weak or warning scenarios:

Scenario Preference gate Net win Δ Pref Runs (W/T/L)
= Apply the migration to SDK-style Eligible +0.0% +0.0% 0/1/0
= Consolidate duplicated projects into a multi-targeting SDK-style project Eligible +0.0% +0.0% 0/1/0
▼ Decline modernizing a non-.NET build Excluded (activation contract) -100.0% -40.0% 0/0/1
= Identify legacy patterns for SDK-style migration Eligible +0.0% +0.0% 0/1/0
= Modernize a single project without introducing Central Package Management Eligible +0.0% +0.0% 0/1/0
= Recognize an already-modern project needs no migration Eligible +0.0% +0.0% 0/1/0

Illustrative judge evidence:

  • Apply the migration to SDK-style: Position-swap inconsistent (forward: tie, reverse: A). Defaulting to tie.

This is one example, not the aggregate verdict. Open Full Results for every judgment.

Details for 8 results were omitted to keep this comment under GitHub's limit. Open the workflow summary or Full Results for the complete breakdown.

🔍 Full Results - all metrics and investigation details

To investigate non-passing or warning results, paste this to your AI coding agent:

For PR 1118 in dotnet/skills, download eval artifacts with gh run download 34273037202 --repo dotnet/skills --pattern "vally-results-*" --dir ./eval-results, then fetch https://raw.githubusercontent.com/dotnet/skills/6866836a9b096ae2e9aefefe6dbfdab67405f37f/eng/vally-adapter/InvestigatingResults.md and follow it. Classify each result as measurement-invalid, underpowered, not-proven, preference-loss, or passing-with-warning. Use stateReason, result accounting, weak scenarios, and judge evidence to give the cause and exact next fix.

Put required input constraints before overlapping performance, generated-file, item, and property vocabulary identified by the second CI evaluation.

Co-authored-by: Copilot App <223556219+Copilot@users.noreply.github.com>

Copilot-Session: 2d0a0a20-fd83-4343-8434-4763f8120503
github-actions Bot added a commit that referenced this pull request Sep 8, 2026
@github-actions

github-actions Bot commented Sep 8, 2026

Copy link
Copy Markdown
Contributor

❌ Evaluation did not complete successfully (the evaluate job reported failure). Check the workflow run logs, then comment /evaluate 26b6a729a5c819dc4dd3fb45a8beb43ee14c3ba1 to retry this exact commit.

36 partial result file(s) were preserved for diagnosis but were not consolidated because the full matrix did not complete.

Preserve the off-target contracts with leading exclusion rules and retry transient binlog setup failures that otherwise drop one comparison arm.

Co-authored-by: Copilot App <223556219+Copilot@users.noreply.github.com>

Copilot-Session: 2d0a0a20-fd83-4343-8434-4763f8120503
github-actions Bot added a commit that referenced this pull request Sep 8, 2026
@github-actions

github-actions Bot commented Sep 8, 2026

Copy link
Copy Markdown
Contributor

❌ Evaluation did not complete successfully (the evaluate job reported failure). Check the workflow run logs, then comment /evaluate b37f36513cac14d8fd738759cfee7678f3b65483 to retry this exact commit.

36 partial result file(s) were preserved for diagnosis but were not consolidated because the full matrix did not complete.

Describe the positive applicability test for the three remaining Sonnet dormancy misses without leading with overlapping off-target vocabulary.

Co-authored-by: Copilot App <223556219+Copilot@users.noreply.github.com>

Copilot-Session: 2d0a0a20-fd83-4343-8434-4763f8120503
github-actions Bot added a commit that referenced this pull request Sep 9, 2026
@github-actions

github-actions Bot commented Sep 9, 2026

Copy link
Copy Markdown
Contributor

📊 Skill Evaluation Results

36 model/skill results across 18 skills and 2 models — ✅ 6 improved, ➖ 27 not proven improved, ⚠️ 0 invalid or underpowered, ⛔ 3 activation contract failures, 📉 0 preference losses (report only).

Measurement identity: evaluated commit 8bc2f2078f7640a8ab78eff8e439d1370c816923; 2 judge models.

Measurement health: 36 expected / 36 observed / 36 written; 0 missing, 0 unexpected, 0 invalid; 0 recovered comparison error slots and 0 unresolved comparison error slots.

Objective completion gate: not enabled. Aggregate completion transitions are telemetry only, so this report does not claim that zero objective regressions were proven.

A result passes only when preference-eligible distinct-stimulus votes have aggregate net win of at least 20% and an exact one-sided sign-test result of p ≤ 0.05, and every explicit dormancy activation contract passes. Repeated runs measure reliability only.

Skill Model Verdict Gate evidence Overfit Warnings Next action
binlog-failure-analysis claude-sonnet-4.6 ✅ Improved n=7; 5W/2T/0L; d=5; p=0.031; net +71.4%; 1 dormancy excluded ✅ 0.19 — None.
binlog-failure-analysis gpt-5.6-luna ✅ Improved n=7; 6W/1T/0L; d=6; p=0.016; net +85.7%; 1 dormancy excluded ✅ 0.11 — None.
binlog-generation claude-sonnet-4.6 ✅ Improved n=7; 6W/1T/0L; d=6; p=0.016; net +85.7%; 1 dormancy excluded 🔴 0.56 Activation: isolated 7/7; plugin 6/7 Fix activation gaps; Review overfit evidence.
binlog-generation gpt-5.6-luna ✅ Improved n=7; 5W/2T/0L; d=5; p=0.031; net +71.4%; 1 dormancy excluded 🟡 0.26 — Review overfit evidence.
build-parallelism claude-sonnet-4.6 ⛔ Activation contract failed n=6; 2W/3T/1L; d=3; p=0.500; net +16.7%; 1 dormancy excluded 🟡 0.32 Dormancy contract: 1 unexpected activation(s); Activation: isolated 6/6; plugin 5/6 Narrow skill routing so the listed off-target scenarios stay dormant.
build-parallelism gpt-5.6-luna ➖ Not proven improved n=6; 2W/3T/1L; d=3; p=0.500; net +16.7%; 1 dormancy excluded ✅ 0.11 — Inspect tied or lost stimuli; predeclare added breadth before a new experiment.
build-perf-baseline claude-sonnet-4.6 ➖ Not proven improved n=6; 2W/0T/4L; d=6; p=0.344; net -33.3%; 1 dormancy excluded 🟡 0.35 Activation: isolated 5/6; plugin 5/6 Inspect tied or lost stimuli and fix inconsistent skill behavior.
build-perf-baseline gpt-5.6-luna ➖ Not proven improved n=6; 4W/1T/1L; d=5; p=0.188; net +50.0%; 1 dormancy excluded 🟡 0.25 Activation: isolated 6/6; plugin 5/6 Inspect tied or lost stimuli and fix inconsistent skill behavior.
build-perf-diagnostics claude-sonnet-4.6 ➖ Not proven improved n=7; 4W/1T/2L; d=6; p=0.344; net +28.6%; 1 dormancy excluded 🔴 0.50 — Inspect tied or lost stimuli and fix inconsistent skill behavior.
build-perf-diagnostics gpt-5.6-luna ➖ Not proven improved n=7; 5W/1T/1L; d=6; p=0.109; net +57.1%; 1 dormancy excluded ✅ 0.16 — Inspect tied or lost stimuli and fix inconsistent skill behavior.
check-bin-obj-clash claude-sonnet-4.6 ➖ Not proven improved n=7; 3W/3T/1L; d=4; p=0.312; net +28.6% ✅ 0.18 Activation: isolated 6/7; plugin 6/7 Inspect tied or lost stimuli; predeclare added breadth before a new experiment.
check-bin-obj-clash gpt-5.6-luna ➖ Not proven improved n=7; 4W/3T/0L; d=4; p=0.063; net +57.1% ✅ 0.11 — Inspect tied or lost stimuli; predeclare added breadth before a new experiment.
directory-build-organization claude-sonnet-4.6 ➖ Not proven improved n=6; 2W/4T/0L; d=2; p=0.250; net +33.3%; 1 dormancy excluded 🟡 0.36 Activation: isolated 4/6; plugin 5/6 Inspect tied or lost stimuli; predeclare added breadth before a new experiment.
directory-build-organization gpt-5.6-luna ➖ Not proven improved n=6; 1W/5T/0L; d=1; p=0.500; net +16.7%; 1 dormancy excluded ✅ 0.07 Activation: isolated 6/6; plugin 5/6 Inspect tied or lost stimuli; predeclare added breadth before a new experiment.
eval-performance claude-sonnet-4.6 ➖ Not proven improved n=8; 5W/1T/2L; d=7; p=0.227; net +37.5% 🟡 0.36 Activation: isolated 8/8; plugin 7/8 Inspect tied or lost stimuli and fix inconsistent skill behavior.
eval-performance gpt-5.6-luna ✅ Improved n=8; 7W/1T/0L; d=7; p=0.008; net +87.5% ✅ 0.09 — None.
extension-points claude-sonnet-4.6 ➖ Not proven improved n=7; 2W/3T/2L; d=4; p=0.687; net +0.0%; 1 dormancy excluded 🟡 0.41 Activation: isolated 4/7; plugin 4/7 Inspect tied or lost stimuli; predeclare added breadth before a new experiment.
extension-points gpt-5.6-luna ➖ Not proven improved n=7; 3W/3T/1L; d=4; p=0.312; net +28.6%; 1 dormancy excluded 🟡 0.29 — Inspect tied or lost stimuli; predeclare added breadth before a new experiment.
including-generated-files claude-sonnet-4.6 ⛔ Activation contract failed n=7; 1W/4T/2L; d=3; p=0.500; net -14.3%; 1 dormancy excluded 🟡 0.46 Dormancy contract: 1 unexpected activation(s); Activation: isolated 5/7; plugin 6/7 Narrow skill routing so the listed off-target scenarios stay dormant.
including-generated-files gpt-5.6-luna ➖ Not proven improved n=7; 1W/5T/1L; d=2; p=0.750; net +0.0%; 1 dormancy excluded ✅ 0.17 — Inspect tied or lost stimuli; predeclare added breadth before a new experiment.
incremental-build claude-sonnet-4.6 ➖ Not proven improved n=9; 4W/5T/0L; d=4; p=0.063; net +44.4% 🟡 0.45 Activation: isolated 8/9; plugin 4/9 Inspect tied or lost stimuli; predeclare added breadth before a new experiment.
incremental-build gpt-5.6-luna ➖ Not proven improved n=9; 3W/5T/1L; d=4; p=0.312; net +22.2% ✅ 0.14 Activation: isolated 9/9; plugin 7/9 Inspect tied or lost stimuli; predeclare added breadth before a new experiment.
item-management claude-sonnet-4.6 ➖ Not proven improved n=6; 1W/4T/1L; d=2; p=0.750; net +0.0%; 1 dormancy excluded 🟡 0.32 Activation: isolated 4/6; plugin 3/6 Inspect tied or lost stimuli; predeclare added breadth before a new experiment.
item-management gpt-5.6-luna ➖ Not proven improved n=6; 2W/3T/1L; d=3; p=0.500; net +16.7%; 1 dormancy excluded ✅ 0.11 — Inspect tied or lost stimuli; predeclare added breadth before a new experiment.
msbuild-antipatterns claude-sonnet-4.6 ➖ Not proven improved n=7; 1W/5T/1L; d=2; p=0.750; net +0.0%; 1 dormancy excluded 🟡 0.32 Activation: isolated 2/7; plugin 2/7 Inspect tied or lost stimuli; predeclare added breadth before a new experiment.
msbuild-antipatterns gpt-5.6-luna ➖ Not proven improved n=7; 0W/6T/1L; d=1; p=0.500; net -14.3%; 1 dormancy excluded ✅ 0.08 Activation: isolated 6/7; plugin 5/7 Inspect tied or lost stimuli; predeclare added breadth before a new experiment.
msbuild-modernization claude-sonnet-4.6 ➖ Not proven improved n=6; 1W/4T/1L; d=2; p=0.750; net +0.0%; 1 dormancy excluded 🟡 0.29 — Inspect tied or lost stimuli; predeclare added breadth before a new experiment.
msbuild-modernization gpt-5.6-luna ➖ Not proven improved n=6; 2W/4T/0L; d=2; p=0.250; net +33.3%; 1 dormancy excluded ✅ 0.06 — Inspect tied or lost stimuli; predeclare added breadth before a new experiment.
msbuild-server claude-sonnet-4.6 ✅ Improved n=8; 7W/1T/0L; d=7; p=0.008; net +87.5% 🟡 0.43 — Review overfit evidence.
msbuild-server gpt-5.6-luna ➖ Not proven improved n=8; 6W/1T/1L; d=7; p=0.063; net +62.5% ✅ 0.19 — Inspect tied or lost stimuli and fix inconsistent skill behavior.
property-patterns claude-sonnet-4.6 ⛔ Activation contract failed n=6; 4W/2T/0L; d=4; p=0.063; net +66.7%; 1 dormancy excluded 🟡 0.24 Dormancy contract: 1 unexpected activation(s); Activation: isolated 5/6; plugin 4/6 Narrow skill routing so the listed off-target scenarios stay dormant.
property-patterns gpt-5.6-luna ➖ Not proven improved n=6; 3W/2T/1L; d=4; p=0.312; net +33.3%; 1 dormancy excluded ✅ 0.08 — Inspect tied or lost stimuli; predeclare added breadth before a new experiment.
resolve-project-references claude-sonnet-4.6 ➖ Not proven improved n=6; 5W/0T/1L; d=6; p=0.109; net +66.7%; 2 dormancy excluded 🟡 0.36 — Inspect tied or lost stimuli and fix inconsistent skill behavior.
resolve-project-references gpt-5.6-luna ➖ Not proven improved n=6; 2W/3T/1L; d=3; p=0.500; net +16.7%; 2 dormancy excluded ✅ 0.15 Activation: isolated 4/6; plugin 6/6 Inspect tied or lost stimuli; predeclare added breadth before a new experiment.
target-authoring claude-sonnet-4.6 ➖ Not proven improved n=6; 3W/2T/1L; d=4; p=0.312; net +33.3%; 1 dormancy excluded 🟡 0.35 — Inspect tied or lost stimuli; predeclare added breadth before a new experiment.
target-authoring gpt-5.6-luna ➖ Not proven improved n=6; 3W/1T/2L; d=5; p=0.500; net +16.7%; 1 dormancy excluded ✅ 0.12 — Inspect tied or lost stimuli and fix inconsistent skill behavior.
ℹ️ How to read this report
  • ✅ Improved — the result passed both the statistical gate and the 20% practical net-win floor.
  • ➖ Not proven improved — the result is valid but did not pass both gates. This is not automatically a regression.
  • ⚠️ Invalid / underpowered — the gate withheld a quality verdict. Fix the measurement before judging the skill.
  • ⛔ Activation contract failed — the isolated target skill activated on an explicit dormancy scenario. Dormancy preference is excluded, but this routing failure still blocks a pass.
  • 📉 Preference loss — the LLM judge credibly preferred baseline. It is report-only, not objective completion proof.
  • Gate evidence — n preference-eligible distinct-stimulus votes, W/T/L stimulus votes, d discordant votes, exact one-sided p, net win, and the count of separately retained dormancy stimuli. The p value applies to one model/skill result; no matrix-wide multiple-comparison correction is applied.
  • Overfit — overfitting-judge severity (✅ Low, 🟡 Moderate, 🔴 High, — none) and score.
  • Warnings — activation, timeout, retry recovery, or unresolved comparison conditions that need attention.
  • Do not add repeated runs to increase statistical power. Do not add stimuli after seeing a near-pass unless the new breadth is predeclared for a new experiment.
⛔ Activation contract failed — build-parallelism (claude-sonnet-4.6)

Why: Net win +16.7% (2W/3T/1L over 6 preference-eligible stimulus vote(s), sign test p=0.500), mean preference +20.0% across 7 paired run(s), 1 dormancy stimulus/stimuli excluded from preference — activation contract failed (1 explicit dormancy scenario(s) activated the isolated target skill)

Next action: Narrow skill routing so the listed off-target scenarios stay dormant.

State: VALID_NO_CHANGE (activation_contract_failed)

Gate evidence: n=6; 2W/3T/1L; d=3; p=0.500; net +16.7%; 1 dormancy excluded

Warnings: Dormancy contract: 1 unexpected activation(s); Activation: isolated 6/6; plugin 5/6

Overfit: Moderate (score 0.32)

Repeated-run reliability (not used by the gate): 7 paired runs (3W/3T/1L).

Weak or warning scenarios:

Scenario Preference gate Net win Δ Pref Runs (W/T/L)
= Decline graph build for runtime-discovered projects Eligible +0.0% +0.0% 0/1/0
= Enable BuildInParallel on a custom MSBuild task Eligible +0.0% +0.0% 0/1/0
= Enable parallel nodes for a wide project graph Eligible +0.0% +0.0% 0/1/0
▲ Non-activation: make one custom target incremental Excluded (activation contract) +100.0% +40.0% 1/0/0
▼ Reduce CI build scope with a solution filter Eligible -100.0% -40.0% 0/0/1

Illustrative judge evidence:

  • Decline graph build for runtime-discovered projects: Position-swap inconsistent (forward: B, reverse: A). Defaulting to tie.

This is one example, not the aggregate verdict. Open Full Results for every judgment.

⛔ Activation contract failed — including-generated-files (claude-sonnet-4.6)

Why: Net win -14.3% (1W/4T/2L over 7 preference-eligible stimulus vote(s), sign test p=0.500), mean preference -17.5% across 8 paired run(s), 1 dormancy stimulus/stimuli excluded from preference — activation contract failed (1 explicit dormancy scenario(s) activated the isolated target skill)

Next action: Narrow skill routing so the listed off-target scenarios stay dormant.

State: VALID_NO_CHANGE (activation_contract_failed)

Gate evidence: n=7; 1W/4T/2L; d=3; p=0.500; net -14.3%; 1 dormancy excluded

Warnings: Dormancy contract: 1 unexpected activation(s); Activation: isolated 5/7; plugin 6/7

Overfit: Moderate (score 0.46)

Repeated-run reliability (not used by the gate): 8 paired runs (1W/4T/3L).

Weak or warning scenarios:

Scenario Preference gate Net win Δ Pref Runs (W/T/L)
▼ Decline manual wiring for Roslyn source generators Excluded (activation contract) -100.0% -40.0% 0/0/1
= Diagnose missing clean tracking for generated source Eligible +0.0% +0.0% 0/1/0
= Diagnose missing generated source inclusion Eligible +0.0% +0.0% 0/1/0
▼ Diagnose missing output registration for generated non-code file Eligible -100.0% -100.0% 0/0/1
= Diagnose project-level glob for generated source Eligible +0.0% +0.0% 0/1/0
▼ Diagnose wrong hook for generated source files Eligible -100.0% -40.0% 0/0/1
= Fix generated source inclusion and clean tracking Eligible +0.0% +0.0% 0/1/0

Illustrative judge evidence:

  • Diagnose missing clean tracking for generated source: Position-swap inconsistent (forward: tie, reverse: B). Defaulting to tie.

This is one example, not the aggregate verdict. Open Full Results for every judgment.

⛔ Activation contract failed — property-patterns (claude-sonnet-4.6)

Why: Net win +66.7% (4W/2T/0L over 6 preference-eligible stimulus vote(s), sign test p=0.063), mean preference +45.7% across 7 paired run(s), 1 dormancy stimulus/stimuli excluded from preference — activation contract failed (1 explicit dormancy scenario(s) activated the isolated target skill)

Next action: Narrow skill routing so the listed off-target scenarios stay dormant.

State: VALID_NO_CHANGE (activation_contract_failed)

Gate evidence: n=6; 4W/2T/0L; d=4; p=0.063; net +66.7%; 1 dormancy excluded

Warnings: Dormancy contract: 1 unexpected activation(s); Activation: isolated 5/6; plugin 4/6

Overfit: Moderate (score 0.24)

Repeated-run reliability (not used by the gate): 7 paired runs (5W/2T/0L).

Weak or warning scenarios:

Scenario Preference gate Net win Δ Pref Runs (W/T/L)
= Choose the right OS detection for a cross-platform property Eligible +0.0% +0.0% 0/1/0
▲ Decline a props-versus-targets placement question Excluded (activation contract) +100.0% +40.0% 1/0/0
▲ Diagnose a TargetFramework condition that never applies in props Eligible +100.0% +40.0% 1/0/0
= Diagnose multi-level property hierarchy bugs Eligible +0.0% +0.0% 0/1/0

Illustrative judge evidence:

  • Choose the right OS detection for a cross-platform property: Position-swap inconsistent (forward: tie, reverse: B). Defaulting to tie.

This is one example, not the aggregate verdict. Open Full Results for every judgment.

➖ Not proven improved — build-parallelism (gpt-5.6-luna)

Why: Net win +16.7% (2W/3T/1L over 6 preference-eligible stimulus vote(s), sign test p=0.500), mean preference +14.3% across 7 paired run(s), 1 dormancy stimulus/stimuli excluded from preference — not credible — 3 of 6 preference-eligible stimulus vote(s) tied, leaving only 3 discordant preference vote(s). The sign test conditions on non-tie stimulus votes and cannot reach 0.05 below 5, so no record could have passed here — this is not a measured null. Either the skill is inert on these scenarios (make them discriminate) or the eval needs more distinct stimuli to clear the ties

Next action: Inspect tied or lost stimuli; predeclare added breadth before a new experiment.

State: VALID_NO_CHANGE (no_credible_preference_change)

Gate evidence: n=6; 2W/3T/1L; d=3; p=0.500; net +16.7%; 1 dormancy excluded

Overfit: Low (score 0.11)

Repeated-run reliability (not used by the gate): 7 paired runs (2W/4T/1L).

Weak or warning scenarios:

Scenario Preference gate Net win Δ Pref Runs (W/T/L)
= Analyze build parallelism bottlenecks Eligible +0.0% +0.0% 0/1/0
▼ Decline graph build for runtime-discovered projects Eligible -100.0% -40.0% 0/0/1
= Enable BuildInParallel on a custom MSBuild task Eligible +0.0% +0.0% 0/1/0
= Enable parallel nodes for a wide project graph Eligible +0.0% +0.0% 0/1/0
= Non-activation: make one custom target incremental Excluded (activation contract) +0.0% +0.0% 0/1/0

Illustrative judge evidence:

  • Analyze build parallelism bottlenecks: Position-swap inconsistent (forward: B, reverse: tie). Defaulting to tie.

This is one example, not the aggregate verdict. Open Full Results for every judgment.

➖ Not proven improved — build-perf-baseline (claude-sonnet-4.6)

Why: Net win -33.3% (2W/0T/4L over 6 preference-eligible stimulus vote(s), sign test p=0.344), mean preference -25.7% across 7 paired run(s), 1 dormancy stimulus/stimuli excluded from preference — no improvement

Next action: Inspect tied or lost stimuli and fix inconsistent skill behavior.

State: VALID_NO_CHANGE (no_credible_preference_change)

Gate evidence: n=6; 2W/0T/4L; d=6; p=0.344; net -33.3%; 1 dormancy excluded

Warnings: Activation: isolated 5/6; plugin 5/6

Overfit: Moderate (score 0.35)

Repeated-run reliability (not used by the gate): 7 paired runs (2W/0T/5L).

Weak or warning scenarios:

Scenario Preference gate Net win Δ Pref Runs (W/T/L)
▼ Configure deterministic, cache-safe CI builds Eligible -100.0% -40.0% 0/0/1
▼ Decline a non-MSBuild build performance request Excluded (activation contract) -100.0% -100.0% 0/0/1
▼ Leave an already-optimized build unchanged Eligible -100.0% -100.0% 0/0/1
▼ Route a broken no-op rebuild away from generic optimization Eligible -100.0% -40.0% 0/0/1
▼ Route a restore-bound cold build away from architecture changes Eligible -100.0% -40.0% 0/0/1

Illustrative judge evidence:

  • Configure deterministic, cache-safe CI builds: Both miss the required CI-guarded ContinuousIntegrationBuild change and give materially misleading root-cause explanations. A is only marginally better because it acknowledges the SDK default for Deterministic, even though its chosen fix remains unnecessary and insufficient.

This is one example, not the aggregate verdict. Open Full Results for every judgment.

➖ Not proven improved — build-perf-baseline (gpt-5.6-luna)

Why: Net win +50.0% (4W/1T/1L over 6 preference-eligible stimulus vote(s), sign test p=0.188), mean preference +31.4% across 7 paired run(s), 1 dormancy stimulus/stimuli excluded from preference — not credible (sign test p=0.188 > 0.05)

Next action: Inspect tied or lost stimuli and fix inconsistent skill behavior.

State: VALID_NO_CHANGE (no_credible_preference_change)

Gate evidence: n=6; 4W/1T/1L; d=5; p=0.188; net +50.0%; 1 dormancy excluded

Warnings: Activation: isolated 6/6; plugin 5/6

Overfit: Moderate (score 0.25)

Repeated-run reliability (not used by the gate): 7 paired runs (5W/1T/1L).

Weak or warning scenarios:

Scenario Preference gate Net win Δ Pref Runs (W/T/L)
▼ Configure deterministic, cache-safe CI builds Eligible -100.0% -40.0% 0/0/1
= Route a broken no-op rebuild away from generic optimization Eligible +0.0% +0.0% 0/1/0
▲ Route a restore-bound cold build away from architecture changes Eligible +100.0% +40.0% 1/0/0

Illustrative judge evidence:

  • Configure deterministic, cache-safe CI builds: The responses are nearly identical in approach, change, and outcome. Both added ContinuousIntegrationBuild unconditionally (missing the CI-only guard), both verified the build, both explained the rationale briefly. A slightly edges out by additionally verifying DeterministicSo...

This is one example, not the aggregate verdict. Open Full Results for every judgment.

➖ Not proven improved — build-perf-diagnostics (claude-sonnet-4.6)

Why: Net win +28.6% (4W/1T/2L over 7 preference-eligible stimulus vote(s), sign test p=0.344), mean preference +5.0% across 8 paired run(s), 1 dormancy stimulus/stimuli excluded from preference — not credible (sign test p=0.344 > 0.05)

Next action: Inspect tied or lost stimuli and fix inconsistent skill behavior.

State: VALID_NO_CHANGE (no_credible_preference_change)

Gate evidence: n=7; 4W/1T/2L; d=6; p=0.344; net +28.6%; 1 dormancy excluded

Overfit: High (score 0.50)

Repeated-run reliability (not used by the gate): 8 paired runs (4W/1T/3L).

Weak or warning scenarios:

Scenario Preference gate Net win Δ Pref Runs (W/T/L)
▼ Decline a runtime latency request that is not a build performance issue Excluded (activation contract) -100.0% -40.0% 0/0/1
▼ Diagnose NuGet restore running redundantly across CI stages Eligible -100.0% -40.0% 0/0/1
▼ Diagnose evaluation overhead before any target runs Eligible -100.0% -40.0% 0/0/1
= Diagnose per-project overhead across many small projects Eligible +0.0% +0.0% 0/1/0

Illustrative judge evidence:

  • Diagnose NuGet restore running redundantly across CI stages: The core diagnosis and prescribed changes are essentially the same, but A is more concise and avoids B's notably misleading after-time table: a dedicated restore still costs about 9.7s, and test/publish have work beyond compilation, so a 4–5s total CI estimate is unsupported. ...

This is one example, not the aggregate verdict. Open Full Results for every judgment.

➖ Not proven improved — build-perf-diagnostics (gpt-5.6-luna)

Why: Net win +57.1% (5W/1T/1L over 7 preference-eligible stimulus vote(s), sign test p=0.109), mean preference +20.0% across 8 paired run(s), 1 dormancy stimulus/stimuli excluded from preference — not credible (sign test p=0.109 > 0.05)

Next action: Inspect tied or lost stimuli and fix inconsistent skill behavior.

State: VALID_NO_CHANGE (no_credible_preference_change)

Gate evidence: n=7; 5W/1T/1L; d=6; p=0.109; net +57.1%; 1 dormancy excluded

Overfit: Low (score 0.16)

Repeated-run reliability (not used by the gate): 8 paired runs (5W/2T/1L).

Weak or warning scenarios:

Scenario Preference gate Net win Δ Pref Runs (W/T/L)
= Decline a runtime latency request that is not a build performance issue Excluded (activation contract) +0.0% +0.0% 0/1/0
▼ Diagnose NuGet restore running redundantly across CI stages Eligible -100.0% -40.0% 0/0/1
= Diagnose evaluation overhead before any target runs Eligible +0.0% +0.0% 0/1/0

Illustrative judge evidence:

  • Diagnose NuGet restore running redundantly across CI stages: Both responses are essentially correct and equivalent on the core diagnosis and main recommendations, covering all three rubric points. A edges ahead slightly by adding --locked-mode (appropriate given the committed lock file) and explicitly noting the need to preserve restore...

This is one example, not the aggregate verdict. Open Full Results for every judgment.

➖ Not proven improved — check-bin-obj-clash (claude-sonnet-4.6)

Why: Net win +28.6% (3W/3T/1L over 7 preference-eligible stimulus vote(s), sign test p=0.312), mean preference +20.0% across 7 paired run(s) — not credible — 3 of 7 preference-eligible stimulus vote(s) tied, leaving only 4 discordant preference vote(s). The sign test conditions on non-tie stimulus votes and cannot reach 0.05 below 5, so no record could have passed here — this is not a measured null. Either the skill is inert on these scenarios (make them discriminate) or the eval needs more distinct stimuli to clear the ties

Next action: Inspect tied or lost stimuli; predeclare added breadth before a new experiment.

State: VALID_NO_CHANGE (no_credible_preference_change)

Gate evidence: n=7; 3W/3T/1L; d=4; p=0.312; net +28.6%

Warnings: Activation: isolated 6/7; plugin 6/7

Overfit: Low (score 0.18)

Repeated-run reliability (not used by the gate): 7 paired runs (3W/3T/1L).

Weak or warning scenarios:

Scenario Preference gate Net win Δ Pref Runs (W/T/L)
= Audit the mixed solution and separate safe projects from unsafe ones Eligible +0.0% +0.0% 0/1/0
= Decline output-clash remediation for separate projects using the default SDK layout Eligible +0.0% +0.0% 0/1/0
= Diagnose multi-targeting outputs that collapse into one path Eligible +0.0% +0.0% 0/1/0
▼ Fix all clash mechanisms in the mixed solution Eligible -100.0% -40.0% 0/0/1

Illustrative judge evidence:

  • Audit the mixed solution and separate safe projects from unsafe ones: Position-swap inconsistent (forward: B, reverse: A). Defaulting to tie.

This is one example, not the aggregate verdict. Open Full Results for every judgment.

➖ Not proven improved — check-bin-obj-clash (gpt-5.6-luna)

Why: Net win +57.1% (4W/3T/0L over 7 preference-eligible stimulus vote(s), sign test p=0.063), mean preference +22.9% across 7 paired run(s) — not credible — 3 of 7 preference-eligible stimulus vote(s) tied, leaving only 4 discordant preference vote(s). The sign test conditions on non-tie stimulus votes and cannot reach 0.05 below 5, so no record could have passed here — this is not a measured null. Either the skill is inert on these scenarios (make them discriminate) or the eval needs more distinct stimuli to clear the ties

Next action: Inspect tied or lost stimuli; predeclare added breadth before a new experiment.

State: VALID_NO_CHANGE (no_credible_preference_change)

Gate evidence: n=7; 4W/3T/0L; d=4; p=0.063; net +57.1%

Overfit: Low (score 0.11)

Repeated-run reliability (not used by the gate): 7 paired runs (4W/3T/0L).

Weak or warning scenarios:

Scenario Preference gate Net win Δ Pref Runs (W/T/L)
= Audit the mixed solution and separate safe projects from unsafe ones Eligible +0.0% +0.0% 0/1/0
= Diagnose multi-targeting outputs that collapse into one path Eligible +0.0% +0.0% 0/1/0
= Diagnose shared output and intermediate path collision Eligible +0.0% +0.0% 0/1/0

Illustrative judge evidence:

  • Audit the mixed solution and separate safe projects from unsafe ones: Position-swap inconsistent (forward: tie, reverse: B). Defaulting to tie.

This is one example, not the aggregate verdict. Open Full Results for every judgment.

➖ Not proven improved — directory-build-organization (claude-sonnet-4.6)

Why: Net win +33.3% (2W/4T/0L over 6 preference-eligible stimulus vote(s), sign test p=0.250), mean preference +20.0% across 7 paired run(s), 1 dormancy stimulus/stimuli excluded from preference — not credible — 4 of 6 preference-eligible stimulus vote(s) tied, leaving only 2 discordant preference vote(s). The sign test conditions on non-tie stimulus votes and cannot reach 0.05 below 5, so no record could have passed here — this is not a measured null. Either the skill is inert on these scenarios (make them discriminate) or the eval needs more distinct stimuli to clear the ties

Next action: Inspect tied or lost stimuli; predeclare added breadth before a new experiment.

State: VALID_NO_CHANGE (no_credible_preference_change)

Gate evidence: n=6; 2W/4T/0L; d=2; p=0.250; net +33.3%; 1 dormancy excluded

Warnings: Activation: isolated 4/6; plugin 5/6

Overfit: Moderate (score 0.36)

Repeated-run reliability (not used by the gate): 7 paired runs (2W/5T/0L).

Weak or warning scenarios:

Scenario Preference gate Net win Δ Pref Runs (W/T/L)
= Decline adding shared build files to a lone project Excluded (activation contract) +0.0% +0.0% 0/1/0
= Diagnose a TargetFramework condition that silently skips in props Eligible +0.0% +0.0% 0/1/0
= Diagnose a package downgrade chain and reorganize version management Eligible +0.0% +0.0% 0/1/0
= Diagnose an inner shared-props file that overwrites its own override Eligible +0.0% +0.0% 0/1/0
= Preserve an intentional project-specific exception while centralizing Eligible +0.0% +0.0% 0/1/0

Illustrative judge evidence:

  • Diagnose a TargetFramework condition that silently skips in props: Position-swap inconsistent (forward: tie, reverse: B). Defaulting to tie.

This is one example, not the aggregate verdict. Open Full Results for every judgment.

➖ Not proven improved — directory-build-organization (gpt-5.6-luna)

Why: Net win +16.7% (1W/5T/0L over 6 preference-eligible stimulus vote(s), sign test p=0.500), mean preference +5.7% across 7 paired run(s), 1 dormancy stimulus/stimuli excluded from preference — not credible — 5 of 6 preference-eligible stimulus vote(s) tied, leaving only 1 discordant preference vote(s). The sign test conditions on non-tie stimulus votes and cannot reach 0.05 below 5, so no record could have passed here — this is not a measured null. Either the skill is inert on these scenarios (make them discriminate) or the eval needs more distinct stimuli to clear the ties

Next action: Inspect tied or lost stimuli; predeclare added breadth before a new experiment.

State: VALID_NO_CHANGE (no_credible_preference_change)

Gate evidence: n=6; 1W/5T/0L; d=1; p=0.500; net +16.7%; 1 dormancy excluded

Warnings: Activation: isolated 6/6; plugin 5/6

Overfit: Low (score 0.07)

Repeated-run reliability (not used by the gate): 7 paired runs (1W/6T/0L).

Weak or warning scenarios:

Scenario Preference gate Net win Δ Pref Runs (W/T/L)
= Apply repo-level build organization cleanup Eligible +0.0% +0.0% 0/1/0
= Decline adding shared build files to a lone project Excluded (activation contract) +0.0% +0.0% 0/1/0
= Diagnose a TargetFramework condition that silently skips in props Eligible +0.0% +0.0% 0/1/0
= Diagnose a package downgrade chain and reorganize version management Eligible +0.0% +0.0% 0/1/0
= Diagnose an inner shared-props file that overwrites its own override Eligible +0.0% +0.0% 0/1/0
= Preserve an intentional project-specific exception while centralizing Eligible +0.0% +0.0% 0/1/0

Illustrative judge evidence:

  • Apply repo-level build organization cleanup: Position-swap inconsistent (forward: tie, reverse: A). Defaulting to tie.

This is one example, not the aggregate verdict. Open Full Results for every judgment.

➖ Not proven improved — eval-performance (claude-sonnet-4.6)

Why: Net win +37.5% (5W/1T/2L over 8 preference-eligible stimulus vote(s), sign test p=0.227), mean preference +30.0% across 8 paired run(s) — not credible (sign test p=0.227 > 0.05)

Next action: Inspect tied or lost stimuli and fix inconsistent skill behavior.

State: VALID_NO_CHANGE (no_credible_preference_change)

Gate evidence: n=8; 5W/1T/2L; d=7; p=0.227; net +37.5%

Warnings: Activation: isolated 8/8; plugin 7/8

Overfit: Moderate (score 0.36)

Repeated-run reliability (not used by the gate): 8 paired runs (5W/1T/2L).

Weak or warning scenarios:

Scenario Preference gate Net win Δ Pref Runs (W/T/L)
▼ Recognize TreatAsLocalProperty overuse versus one justified entry Eligible -100.0% -40.0% 0/0/1
▼ Redirect a compile-time slowdown mistakenly framed as an evaluation problem Eligible -100.0% -40.0% 0/0/1
= Triage which of two property functions actually costs evaluation time Eligible +0.0% +0.0% 0/1/0

Illustrative judge evidence:

  • Recognize TreatAsLocalProperty overuse versus one justified entry: Both satisfy every requested rubric item and give the right change. A is marginally stronger because it avoids B's questionable claims about blocking parent overrides and inheritance behavior.

This is one example, not the aggregate verdict. Open Full Results for every judgment.

➖ Not proven improved — extension-points (claude-sonnet-4.6)

Why: Net win +0.0% (2W/3T/2L over 7 preference-eligible stimulus vote(s), sign test p=0.687), mean preference -12.5% across 8 paired run(s), 1 dormancy stimulus/stimuli excluded from preference — no improvement

Next action: Inspect tied or lost stimuli; predeclare added breadth before a new experiment.

State: VALID_NO_CHANGE (no_credible_preference_change)

Gate evidence: n=7; 2W/3T/2L; d=4; p=0.687; net +0.0%; 1 dormancy excluded

Warnings: Activation: isolated 4/7; plugin 4/7

Overfit: Moderate (score 0.41)

Repeated-run reliability (not used by the gate): 8 paired runs (2W/3T/3L).

Weak or warning scenarios:

Scenario Preference gate Net win Δ Pref Runs (W/T/L)
= Create extensibility hooks for a custom SDK target file Eligible +0.0% +0.0% 0/1/0
▼ Diagnose a broken per-TFM forwarder Eligible -100.0% -40.0% 0/0/1
= Diagnose a package ID and file-name mismatch Eligible +0.0% +0.0% 0/1/0
▼ Fix extension point anti-patterns Eligible -100.0% -100.0% 0/0/1
▼ Non-activation: repair an incremental custom target Excluded (activation contract) -100.0% -40.0% 0/0/1
= Review packed layout without a false missing-file bug Eligible +0.0% +0.0% 0/1/0

Illustrative judge evidence:

  • Create extensibility hooks for a custom SDK target file: Position-swap inconsistent (forward: B, reverse: A). Defaulting to tie.

This is one example, not the aggregate verdict. Open Full Results for every judgment.

➖ Not proven improved — extension-points (gpt-5.6-luna)

Why: Net win +28.6% (3W/3T/1L over 7 preference-eligible stimulus vote(s), sign test p=0.312), mean preference +22.5% across 8 paired run(s), 1 dormancy stimulus/stimuli excluded from preference — not credible — 3 of 7 preference-eligible stimulus vote(s) tied, leaving only 4 discordant preference vote(s). The sign test conditions on non-tie stimulus votes and cannot reach 0.05 below 5, so no record could have passed here — this is not a measured null. Either the skill is inert on these scenarios (make them discriminate) or the eval needs more distinct stimuli to clear the ties

Next action: Inspect tied or lost stimuli; predeclare added breadth before a new experiment.

State: VALID_NO_CHANGE (no_credible_preference_change)

Gate evidence: n=7; 3W/3T/1L; d=4; p=0.312; net +28.6%; 1 dormancy excluded

Overfit: Moderate (score 0.29)

Repeated-run reliability (not used by the gate): 8 paired runs (4W/3T/1L).

Weak or warning scenarios:

Scenario Preference gate Net win Δ Pref Runs (W/T/L)
= Diagnose NuGet package and repo extension conflicts Eligible +0.0% +0.0% 0/1/0
▼ Diagnose build extension point failures Eligible -100.0% -40.0% 0/0/1
= Fix extension point anti-patterns Eligible +0.0% +0.0% 0/1/0
= Review packed layout without a false missing-file bug Eligible +0.0% +0.0% 0/1/0

Illustrative judge evidence:

  • Diagnose NuGet package and repo extension conflicts: Position-swap inconsistent (forward: tie, reverse: A). Defaulting to tie.

This is one example, not the aggregate verdict. Open Full Results for every judgment.

➖ Not proven improved — including-generated-files (gpt-5.6-luna)

Why: Net win +0.0% (1W/5T/1L over 7 preference-eligible stimulus vote(s), sign test p=0.750), mean preference +0.0% across 8 paired run(s), 1 dormancy stimulus/stimuli excluded from preference — no improvement

Next action: Inspect tied or lost stimuli; predeclare added breadth before a new experiment.

State: VALID_NO_CHANGE (no_credible_preference_change)

Gate evidence: n=7; 1W/5T/1L; d=2; p=0.750; net +0.0%; 1 dormancy excluded

Overfit: Low (score 0.17)

Repeated-run reliability (not used by the gate): 8 paired runs (1W/6T/1L).

Weak or warning scenarios:

Scenario Preference gate Net win Δ Pref Runs (W/T/L)
= Decline manual wiring for Roslyn source generators Excluded (activation contract) +0.0% +0.0% 0/1/0
= Diagnose hardcoded obj path for generated source Eligible +0.0% +0.0% 0/1/0
▼ Diagnose missing clean tracking for generated source Eligible -100.0% -40.0% 0/0/1
= Diagnose missing generated source inclusion Eligible +0.0% +0.0% 0/1/0
= Diagnose missing output registration for generated non-code file Eligible +0.0% +0.0% 0/1/0
= Diagnose project-level glob for generated source Eligible +0.0% +0.0% 0/1/0
= Diagnose wrong hook for generated source files Eligible +0.0% +0.0% 0/1/0

Illustrative judge evidence:

  • Diagnose hardcoded obj path for generated source: Position-swap inconsistent (forward: B, reverse: tie). Defaulting to tie.

This is one example, not the aggregate verdict. Open Full Results for every judgment.

➖ Not proven improved — incremental-build (claude-sonnet-4.6)

Why: Net win +44.4% (4W/5T/0L over 9 preference-eligible stimulus vote(s), sign test p=0.063), mean preference +17.8% across 9 paired run(s) — not credible — 5 of 9 preference-eligible stimulus vote(s) tied, leaving only 4 discordant preference vote(s). The sign test conditions on non-tie stimulus votes and cannot reach 0.05 below 5, so no record could have passed here — this is not a measured null. Either the skill is inert on these scenarios (make them discriminate) or the eval needs more distinct stimuli to clear the ties

Next action: Inspect tied or lost stimuli; predeclare added breadth before a new experiment.

State: VALID_NO_CHANGE (no_credible_preference_change)

Gate evidence: n=9; 4W/5T/0L; d=4; p=0.063; net +44.4%

Warnings: Activation: isolated 8/9; plugin 4/9

Overfit: Moderate (score 0.45)

Repeated-run reliability (not used by the gate): 9 paired runs (4W/5T/0L).

Weak or warning scenarios:

Scenario Preference gate Net win Δ Pref Runs (W/T/L)
▲ Correct the assumption that Outputs alone enables incremental skipping Eligible +100.0% +40.0% 1/0/0
= Diagnose custom targets that always rerun Eligible +0.0% +0.0% 0/1/0
= Explain slower builds when MSBuild skipped everything Eligible +0.0% +0.0% 0/1/0
= Explain why Visual Studio keeps rebuilding an up-to-date project Eligible +0.0% +0.0% 0/1/0
▲ Explain why clean leaves generated hash source behind Eligible +100.0% +40.0% 1/0/0
= Fix broken incremental targets and clean tracking Eligible +0.0% +0.0% 0/1/0
= Identify a volatile output path that defeats incrementality Eligible +0.0% +0.0% 0/1/0
▲ Read a diagnostic log to find the stale input Eligible +100.0% +40.0% 1/0/0

Illustrative judge evidence:

  • Diagnose custom targets that always rerun: Position-swap inconsistent (forward: B, reverse: tie). Defaulting to tie.

This is one example, not the aggregate verdict. Open Full Results for every judgment.

➖ Not proven improved — incremental-build (gpt-5.6-luna)

Why: Net win +22.2% (3W/5T/1L over 9 preference-eligible stimulus vote(s), sign test p=0.312), mean preference +15.6% across 9 paired run(s) — not credible — 5 of 9 preference-eligible stimulus vote(s) tied, leaving only 4 discordant preference vote(s). The sign test conditions on non-tie stimulus votes and cannot reach 0.05 below 5, so no record could have passed here — this is not a measured null. Either the skill is inert on these scenarios (make them discriminate) or the eval needs more distinct stimuli to clear the ties

Next action: Inspect tied or lost stimuli; predeclare added breadth before a new experiment.

State: VALID_NO_CHANGE (no_credible_preference_change)

Gate evidence: n=9; 3W/5T/1L; d=4; p=0.312; net +22.2%

Warnings: Activation: isolated 9/9; plugin 7/9

Overfit: Low (score 0.14)

Repeated-run reliability (not used by the gate): 9 paired runs (3W/5T/1L).

Weak or warning scenarios:

Scenario Preference gate Net win Δ Pref Runs (W/T/L)
= Diagnose custom targets that always rerun Eligible +0.0% +0.0% 0/1/0
= Explain slower builds when MSBuild skipped everything Eligible +0.0% +0.0% 0/1/0
▼ Explain why Visual Studio keeps rebuilding an up-to-date project Eligible -100.0% -40.0% 0/0/1
= Fix broken incremental targets and clean tracking Eligible +0.0% +0.0% 0/1/0
= Identify a volatile output path that defeats incrementality Eligible +0.0% +0.0% 0/1/0
= Read a diagnostic log to find the stale input Eligible +0.0% +0.0% 0/1/0

Illustrative judge evidence:

  • Diagnose custom targets that always rerun: Position-swap inconsistent (forward: B, reverse: tie). Defaulting to tie.

This is one example, not the aggregate verdict. Open Full Results for every judgment.

➖ Not proven improved — item-management (claude-sonnet-4.6)

Why: Net win +0.0% (1W/4T/1L over 6 preference-eligible stimulus vote(s), sign test p=0.750), mean preference -8.6% across 7 paired run(s), 1 dormancy stimulus/stimuli excluded from preference — no improvement

Next action: Inspect tied or lost stimuli; predeclare added breadth before a new experiment.

State: VALID_NO_CHANGE (no_credible_preference_change)

Gate evidence: n=6; 1W/4T/1L; d=2; p=0.750; net +0.0%; 1 dormancy excluded

Warnings: Activation: isolated 4/6; plugin 3/6

Overfit: Moderate (score 0.32)

Repeated-run reliability (not used by the gate): 7 paired runs (1W/5T/1L).

Weak or warning scenarios:

Scenario Preference gate Net win Δ Pref Runs (W/T/L)
= Decline a target ordering problem with no item-group defect Excluded (activation contract) +0.0% +0.0% 0/1/0
▼ Diagnose an ineffective Compile Remove that does not match the glob Eligible -100.0% -100.0% 0/0/1
= Diagnose real and claimed item problems in a code generation pipeline Eligible +0.0% +0.0% 0/1/0
= Fix item management anti-patterns Eligible +0.0% +0.0% 0/1/0
= Leave already-correct item management unchanged Eligible +0.0% +0.0% 0/1/0
= Leave correct single-list batching unchanged Eligible +0.0% +0.0% 0/1/0

Illustrative judge evidence:

  • Diagnose an ineffective Compile Remove that does not match the glob: A directly and accurately answers the actual cause and fix. B supplies a misleading, technically incorrect separator-based root cause, failing the core requested diagnosis despite ending with a usable-looking glob.

This is one example, not the aggregate verdict. Open Full Results for every judgment.

➖ Not proven improved — item-management (gpt-5.6-luna)

Why: Net win +16.7% (2W/3T/1L over 6 preference-eligible stimulus vote(s), sign test p=0.500), mean preference +5.7% across 7 paired run(s), 1 dormancy stimulus/stimuli excluded from preference — not credible — 3 of 6 preference-eligible stimulus vote(s) tied, leaving only 3 discordant preference vote(s). The sign test conditions on non-tie stimulus votes and cannot reach 0.05 below 5, so no record could have passed here — this is not a measured null. Either the skill is inert on these scenarios (make them discriminate) or the eval needs more distinct stimuli to clear the ties

Next action: Inspect tied or lost stimuli; predeclare added breadth before a new experiment.

State: VALID_NO_CHANGE (no_credible_preference_change)

Gate evidence: n=6; 2W/3T/1L; d=3; p=0.500; net +16.7%; 1 dormancy excluded

Overfit: Low (score 0.11)

Repeated-run reliability (not used by the gate): 7 paired runs (2W/4T/1L).

Weak or warning scenarios:

Scenario Preference gate Net win Δ Pref Runs (W/T/L)
= Decline a target ordering problem with no item-group defect Excluded (activation contract) +0.0% +0.0% 0/1/0
▼ Diagnose an ineffective Compile Remove that does not match the glob Eligible -100.0% -40.0% 0/0/1
= Fix item management anti-patterns Eligible +0.0% +0.0% 0/1/0
= Leave already-correct item management unchanged Eligible +0.0% +0.0% 0/1/0
= Leave correct single-list batching unchanged Eligible +0.0% +0.0% 0/1/0

Illustrative judge evidence:

  • Diagnose an ineffective Compile Remove that does not match the glob: Both reach the same correct diagnosis and avoid recommending Exclude. However, Response A actually applied the fix to the file (recovering from the failed apply_patch by using sed) and verified via dotnet msbuild -getItem:Compile that Api.g.cs is now excluded while Service.c...

This is one example, not the aggregate verdict. Open Full Results for every judgment.

➖ Not proven improved — msbuild-antipatterns (claude-sonnet-4.6)

Why: Net win +0.0% (1W/5T/1L over 7 preference-eligible stimulus vote(s), sign test p=0.750), mean preference +0.0% across 8 paired run(s), 1 dormancy stimulus/stimuli excluded from preference — no improvement

Next action: Inspect tied or lost stimuli; predeclare added breadth before a new experiment.

State: VALID_NO_CHANGE (no_credible_preference_change)

Gate evidence: n=7; 1W/5T/1L; d=2; p=0.750; net +0.0%; 1 dormancy excluded

Warnings: Activation: isolated 2/7; plugin 2/7

Overfit: Moderate (score 0.32)

Repeated-run reliability (not used by the gate): 8 paired runs (1W/6T/1L).

Weak or warning scenarios:

Scenario Preference gate Net win Δ Pref Runs (W/T/L)
= Add a module to an F# project Eligible +0.0% +0.0% 0/1/0
= Add a signature file to define public API Eligible +0.0% +0.0% 0/1/0
= Distinguish a style backslash from a real cross-platform backslash bug Eligible +0.0% +0.0% 0/1/0
= Fix broken file order causing FS0039 Eligible +0.0% +0.0% 0/1/0
▼ Judge an unguarded import inside a NuGet package build folder Eligible -100.0% -40.0% 0/0/1
= Non-activation: migrate a legacy project to SDK style Excluded (activation contract) +0.0% +0.0% 0/1/0
= Review MSBuild files for anti-patterns and style issues Eligible +0.0% +0.0% 0/1/0

Illustrative judge evidence:

  • Add a module to an F# project: The implementations and verification outcomes are materially identical, fully addressing the requested validation, compilation ordering, and guarded processing behavior.

This is one example, not the aggregate verdict. Open Full Results for every judgment.

➖ Not proven improved — msbuild-antipatterns (gpt-5.6-luna)

Why: Net win -14.3% (0W/6T/1L over 7 preference-eligible stimulus vote(s), sign test p=0.500), mean preference -10.0% across 8 paired run(s), 1 dormancy stimulus/stimuli excluded from preference — no improvement

Next action: Inspect tied or lost stimuli; predeclare added breadth before a new experiment.

State: VALID_NO_CHANGE (no_credible_preference_change)

Gate evidence: n=7; 0W/6T/1L; d=1; p=0.500; net -14.3%; 1 dormancy excluded

Warnings: Activation: isolated 6/7; plugin 5/7

Overfit: Low (score 0.08)

Repeated-run reliability (not used by the gate): 8 paired runs (0W/6T/2L).

Weak or warning scenarios:

Scenario Preference gate Net win Δ Pref Runs (W/T/L)
= Add a module to an F# project Eligible +0.0% +0.0% 0/1/0
= Add a signature file to define public API Eligible +0.0% +0.0% 0/1/0
= Distinguish a style backslash from a real cross-platform backslash bug Eligible +0.0% +0.0% 0/1/0
= Fix broken file order causing FS0039 Eligible +0.0% +0.0% 0/1/0
▼ Judge an unguarded import inside a NuGet package build folder Eligible -100.0% -40.0% 0/0/1
= Leave a clean project without inventing anti-patterns Eligible +0.0% +0.0% 0/1/0
▼ Non-activation: migrate a legacy project to SDK style Excluded (activation contract) -100.0% -40.0% 0/0/1
= Review MSBuild files for anti-patterns and style issues Eligible +0.0% +0.0% 0/1/0

Illustrative judge evidence:

  • Add a module to an F# project: Position-swap inconsistent (forward: tie, reverse: B). Defaulting to tie.

This is one example, not the aggregate verdict. Open Full Results for every judgment.

➖ Not proven improved — msbuild-modernization (claude-sonnet-4.6)

Why: Net win +0.0% (1W/4T/1L over 6 preference-eligible stimulus vote(s), sign test p=0.750), mean preference +22.9% across 7 paired run(s), 1 dormancy stimulus/stimuli excluded from preference — no improvement

Next action: Inspect tied or lost stimuli; predeclare added breadth before a new experiment.

State: VALID_NO_CHANGE (no_credible_preference_change)

Gate evidence: n=6; 1W/4T/1L; d=2; p=0.750; net +0.0%; 1 dormancy excluded

Overfit: Moderate (score 0.29)

Repeated-run reliability (not used by the gate): 7 paired runs (2W/4T/1L).

Weak or warning scenarios:

Scenario Preference gate Net win Δ Pref Runs (W/T/L)
▼ Apply the migration to SDK-style Eligible -100.0% -40.0% 0/0/1
= Consolidate duplicated projects into a multi-targeting SDK-style project Eligible +0.0% +0.0% 0/1/0
= Identify legacy patterns for SDK-style migration Eligible +0.0% +0.0% 0/1/0
= Modernize a single project without introducing Central Package Management Eligible +0.0% +0.0% 0/1/0
= Recognize an already-modern project needs no migration Eligible +0.0% +0.0% 0/1/0

Illustrative judge evidence:

  • Apply the migration to SDK-style: Both conversions satisfy the core SDK-style migration requirements. A is marginally safer for fidelity because it retains all preexisting assembly metadata directly rather than requiring every attribute to be accurately translated before deleting AssemblyInfo.cs.

This is one example, not the aggregate verdict. Open Full Results for every judgment.

➖ Not proven improved — msbuild-modernization (gpt-5.6-luna)

Why: Net win +33.3% (2W/4T/0L over 6 preference-eligible stimulus vote(s), sign test p=0.250), mean preference +11.4% across 7 paired run(s), 1 dormancy stimulus/stimuli excluded from preference — not credible — 4 of 6 preference-eligible stimulus vote(s) tied, leaving only 2 discordant preference vote(s). The sign test conditions on non-tie stimulus votes and cannot reach 0.05 below 5, so no record could have passed here — this is not a measured null. Either the skill is inert on these scenarios (make them discriminate) or the eval needs more distinct stimuli to clear the ties

Next action: Inspect tied or lost stimuli; predeclare added breadth before a new experiment.

State: VALID_NO_CHANGE (no_credible_preference_change)

Gate evidence: n=6; 2W/4T/0L; d=2; p=0.250; net +33.3%; 1 dormancy excluded

Overfit: Low (score 0.06)

Repeated-run reliability (not used by the gate): 7 paired runs (2W/5T/0L).

Weak or warning scenarios:

Scenario Preference gate Net win Δ Pref Runs (W/T/L)
= Consolidate duplicated projects into a multi-targeting SDK-style project Eligible +0.0% +0.0% 0/1/0
= Decline modernizing a non-.NET build Excluded (activation contract) +0.0% +0.0% 0/1/0
= Identify legacy patterns for SDK-style migration Eligible +0.0% +0.0% 0/1/0
= Modernize a single project without introducing Central Package Management Eligible +0.0% +0.0% 0/1/0
= Recognize an already-modern project needs no migration Eligible +0.0% +0.0% 0/1/0

Illustrative judge evidence:

  • Consolidate duplicated projects into a multi-targeting SDK-style project: Position-swap inconsistent (forward: tie, reverse: A). Defaulting to tie.

This is one example, not the aggregate verdict. Open Full Results for every judgment.

➖ Not proven improved — msbuild-server (gpt-5.6-luna)

Why: Net win +62.5% (6W/1T/1L over 8 preference-eligible stimulus vote(s), sign test p=0.063), mean preference +40.0% across 8 paired run(s) — not credible (sign test p=0.063 > 0.05)

Next action: Inspect tied or lost stimuli and fix inconsistent skill behavior.

State: VALID_NO_CHANGE (no_credible_preference_change)

Gate evidence: n=8; 6W/1T/1L; d=7; p=0.063; net +62.5%

Overfit: Low (score 0.19)

Repeated-run reliability (not used by the gate): 8 paired runs (6W/1T/1L).

Weak or warning scenarios:

Scenario Preference gate Net win Δ Pref Runs (W/T/L)
▼ Confirm the MSBuild Server is actually improving build times before declaring success Eligible -100.0% -40.0% 0/0/1
= Decline MSBuild Server for a single one-off release build Eligible +0.0% +0.0% 0/1/0

Illustrative judge evidence:

  • Confirm the MSBuild Server is actually improving build times before declaring success: Both responses give solid, empirically-grounded advice with cold/warm comparison and A/B testing. A adds value by investigating the actual project (discovering it's a tiny LoopApp), warning that the trivial project gives no meaningful signal, and providing a diagnostic log gre...

This is one example, not the aggregate verdict. Open Full Results for every judgment.

➖ Not proven improved — property-patterns (gpt-5.6-luna)

Why: Net win +33.3% (3W/2T/1L over 6 preference-eligible stimulus vote(s), sign test p=0.312), mean preference +11.4% across 7 paired run(s), 1 dormancy stimulus/stimuli excluded from preference — not credible — 2 of 6 preference-eligible stimulus vote(s) tied, leaving only 4 discordant preference vote(s). The sign test conditions on non-tie stimulus votes and cannot reach 0.05 below 5, so no record could have passed here — this is not a measured null. Either the skill is inert on these scenarios (make them discriminate) or the eval needs more distinct stimuli to clear the ties

Next action: Inspect tied or lost stimuli; predeclare added breadth before a new experiment.

State: VALID_NO_CHANGE (no_credible_preference_change)

Gate evidence: n=6; 3W/2T/1L; d=4; p=0.312; net +33.3%; 1 dormancy excluded

Overfit: Low (score 0.08)

Repeated-run reliability (not used by the gate): 7 paired runs (3W/3T/1L).

Weak or warning scenarios:

Scenario Preference gate Net win Δ Pref Runs (W/T/L)
= Decline a props-versus-targets placement question Excluded (activation contract) +0.0% +0.0% 0/1/0
= Diagnose a TargetFramework condition that never applies in props Eligible +0.0% +0.0% 0/1/0
= Diagnose shared build property issues Eligible +0.0% +0.0% 0/1/0
▼ Fix shared property configuration Eligible -100.0% -40.0% 0/0/1

Illustrative judge evidence:

  • Diagnose a TargetFramework condition that never applies in props: Position-swap inconsistent (forward: B, reverse: A). Defaulting to tie.

This is one example, not the aggregate verdict. Open Full Results for every judgment.

➖ Not proven improved — resolve-project-references (claude-sonnet-4.6)

Why: Net win +66.7% (5W/0T/1L over 6 preference-eligible stimulus vote(s), sign test p=0.109), mean preference +37.5% across 8 paired run(s), 2 dormancy stimulus/stimuli excluded from preference — not credible (sign test p=0.109 > 0.05)

Next action: Inspect tied or lost stimuli and fix inconsistent skill behavior.

State: VALID_NO_CHANGE (no_credible_preference_change)

Gate evidence: n=6; 5W/0T/1L; d=6; p=0.109; net +66.7%; 2 dormancy excluded

Overfit: Moderate (score 0.36)

Repeated-run reliability (not used by the gate): 8 paired runs (7W/0T/1L).

Weak or warning scenarios:

Scenario Preference gate Net win Δ Pref Runs (W/T/L)
▼ Distinguish wait time from a real serial dependency chain Eligible -100.0% -40.0% 0/0/1

Illustrative judge evidence:

  • Distinguish wait time from a real serial dependency chain: Both are strong and satisfy every requested point. A is marginally more technically precise and avoids conflating the inflated target accounting with a direct bottleneck, while still clearly identifying the serial graph and remedy.

This is one example, not the aggregate verdict. Open Full Results for every judgment.

➖ Not proven improved — resolve-project-references (gpt-5.6-luna)

Why: Net win +16.7% (2W/3T/1L over 6 preference-eligible stimulus vote(s), sign test p=0.500), mean preference +17.5% across 8 paired run(s), 2 dormancy stimulus/stimuli excluded from preference — not credible — 3 of 6 preference-eligible stimulus vote(s) tied, leaving only 3 discordant preference vote(s). The sign test conditions on non-tie stimulus votes and cannot reach 0.05 below 5, so no record could have passed here — this is not a measured null. Either the skill is inert on these scenarios (make them discriminate) or the eval needs more distinct stimuli to clear the ties

Next action: Inspect tied or lost stimuli; predeclare added breadth before a new experiment.

State: VALID_NO_CHANGE (no_credible_preference_change)

Gate evidence: n=6; 2W/3T/1L; d=3; p=0.500; net +16.7%; 2 dormancy excluded

Warnings: Activation: isolated 4/6; plugin 6/6

Overfit: Low (score 0.15)

Repeated-run reliability (not used by the gate): 8 paired runs (3W/4T/1L).

Weak or warning scenarios:

Scenario Preference gate Net win Δ Pref Runs (W/T/L)
= Distinguish wait time from a real serial dependency chain Eligible +0.0% +0.0% 0/1/0
▲ Give the exact command to replay task self-time Eligible +100.0% +100.0% 1/0/0
= Rank Copy ahead of Csc when task self-time is higher Eligible +0.0% +0.0% 0/1/0
= Redirect from misleading target summary to Csc self-time Eligible +0.0% +0.0% 0/1/0
▼ Require diagnostic evidence before analyzing project-reference time Eligible -100.0% -40.0% 0/0/1
= Stay dormant when Csc is already the obvious bottleneck Excluded (activation contract) +0.0% +0.0% 0/1/0

Illustrative judge evidence:

  • Distinguish wait time from a real serial dependency chain: Position-swap inconsistent (forward: tie, reverse: A). Defaulting to tie.

This is one example, not the aggregate verdict. Open Full Results for every judgment.

Routine passing details for 3 results are in Full Results.

Details for 5 results were omitted to keep this comment under GitHub's limit. Open the workflow summary or Full Results for the complete breakdown.

🔍 Full Results - all metrics and investigation details

To investigate non-passing or warning results, paste this to your AI coding agent:

For PR 1118 in dotnet/skills, download eval artifacts with gh run download 34291893629 --repo dotnet/skills --pattern "vally-results-*" --dir ./eval-results, then fetch https://raw.githubusercontent.com/dotnet/skills/8bc2f2078f7640a8ab78eff8e439d1370c816923/eng/vally-adapter/InvestigatingResults.md and follow it. Classify each result as measurement-invalid, underpowered, not-proven, preference-loss, or passing-with-warning. Use stateReason, result accounting, weak scenarios, and judge evidence to give the cause and exact next fix.

Remove off-target vocabulary from the three descriptions that Sonnet selected before any workspace inspection, while retaining full boundary guidance in the skill bodies.

Co-authored-by: Copilot App <223556219+Copilot@users.noreply.github.com>

Copilot-Session: 2d0a0a20-fd83-4343-8434-4763f8120503
@github-actions

Copy link
Copy Markdown
Contributor

❌ Evaluation did not complete successfully (the evaluate job reported failure). Check the workflow run logs, then comment /evaluate a37f8fc84c157757eb17ba8a2ee79a98799bb321 to retry this exact commit.

70 partial result file(s) were preserved for diagnosis but were not consolidated because the full matrix did not complete.

Resolve the MSTest migration conflicts while preserving #1118 evaluation reliability coverage.

Co-authored-by: Copilot App <223556219+Copilot@users.noreply.github.com>

Copilot-Session: eb11c85b-4046-4dbc-94b3-3bdbd4839a0d

Copilot AI left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Copilot review overview

🔵 Needs a closer look

The broad evaluator and fixture changes require the documented exact-head validation and fresh human approval.

Review effort: Lite
Findings: None

Resolved since last review (1)

@github-actions

Copy link
Copy Markdown
Contributor

👋 @AbhitejJohn — this PR has merge conflict. When you're ready, please address the feedback and push an update; the triage bot will pick up the next state automatically. (Add the no-stale label to silence further pings.)

Keep the merged #1208 recovery authoritative while preserving #1118-specific evaluation semantics and current MSTest coverage.

Co-authored-by: Copilot App <223556219+Copilot@users.noreply.github.com>

Copilot-Session: eb11c85b-4046-4dbc-94b3-3bdbd4839a0d
Bring in the Markdown linter startup fix before the final exact-head evaluation.

Co-authored-by: Copilot App <223556219+Copilot@users.noreply.github.com>

Copilot-Session: eb11c85b-4046-4dbc-94b3-3bdbd4839a0d

Copilot AI left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Copilot review overview

🟡 Changes recommended

The serial-chain fixture is missing the direct Tests -> Api reference required by its golden trajectory and rubric.

Get a fresh assessment by requesting another Copilot review.

Review effort: Lite
Findings: 1 High severity

Open (1)

Copilot AI left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Copilot review overview

🟡 Changes recommended

Unresolved shard matching and dormancy contract accounting issues can produce incorrect or passing evaluation results.

Get a fresh assessment by requesting another Copilot review.

Review effort: Lite
Findings: 2 High severity

Open (2)
Resolved since last review (1)

Comment thread eng/evaluation/find-targets.ps1 Outdated
Comment thread eng/vally-adapter/adapt.mjs Outdated
Read executionShard only from suite-level tags and require isolated-arm evidence for dormancy contracts.

Co-authored-by: Copilot App <223556219+Copilot@users.noreply.github.com>

Copilot-Session: eb11c85b-4046-4dbc-94b3-3bdbd4839a0d

Copilot AI left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Copilot review overview

🔵 Needs a closer look

MCP allowlists omit documented tools, and routing evals remain inconsistent with the updated skill boundaries.

Review effort: Lite
Findings: None

Resolved since last review (2)

@github-actions

Copy link
Copy Markdown
Contributor

❌ Evaluation did not complete successfully (the evaluate job reported failure). Check the workflow run logs, then comment /evaluate 69b62d37c8de23d24b1dcb440785ff4aa2e28bef to retry this exact commit.

44 partial result file(s) were preserved for diagnosis but were not consolidated because the full matrix did not complete.

Bring in the preference-based skill value dashboard before final evaluation.

Co-authored-by: Copilot App <223556219+Copilot@users.noreply.github.com>

Copilot-Session: eb11c85b-4046-4dbc-94b3-3bdbd4839a0d

Copilot AI left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Copilot review overview

🟡 Changes recommended

The Codex companion manifest still lacks equivalent MCP tool restrictions, leaving access unrestricted.

Get a fresh assessment by requesting another Copilot review.

Review effort: Lite
Findings: 1 High severity

Open (1)

Comment thread plugins/dotnet-msbuild/plugin.json
Apply the standard manifest allowlist to Codex and enforce exact parity in validator tests.

Co-authored-by: Copilot App <223556219+Copilot@users.noreply.github.com>

Copilot-Session: eb11c85b-4046-4dbc-94b3-3bdbd4839a0d
Bring in Claude manifest version synchronization fixes before final evaluation.

Co-authored-by: Copilot App <223556219+Copilot@users.noreply.github.com>

Copilot-Session: eb11c85b-4046-4dbc-94b3-3bdbd4839a0d

Copilot AI left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Copilot review overview

🟡 Changes recommended

One or more issues must be addressed before approval.

Get a fresh assessment by requesting another Copilot review.

Review effort: Lite
Findings: 1 Medium severity

Open (1)
Resolved since last review (1)
Previously missed (2)

In code that hasn't changed since last review

Medium severity Keep binlog_overview in the explicit allowlist

plugins/​dotnet-msbuild/​.claude-plugin/​plugin.json:25

The new explicit allowlist removes binlog_overview. The repository's Codex smoke contract calls that tool directly (CONTRIBUTING.md:147-150), so the smoke call will no longer be exposed through this MCP server after this change. Keep binlog_overview in the allowlist, or update the smoke contract and every host manifest together.

Medium severity Keep binlog_overview in the explicit allowlist

plugins/​dotnet-msbuild/​plugin.json:25

The new explicit allowlist removes binlog_overview. The repository's Codex smoke contract calls that tool directly (CONTRIBUTING.md:147-150), so the smoke call will no longer be exposed through this MCP server after this change. Keep binlog_overview in the allowlist, or update the smoke contract and every host manifest together.

Comment thread plugins/dotnet-msbuild/.codex-plugin/plugin.json Outdated
Exercise current allowlisted load and diagnostics operations instead of the removed binlog_overview tool.

Co-authored-by: Copilot App <223556219+Copilot@users.noreply.github.com>

Copilot-Session: eb11c85b-4046-4dbc-94b3-3bdbd4839a0d
Use the live read-only tool inventory across all hosts and restore the credential-free overview smoke.

Co-authored-by: Copilot App <223556219+Copilot@users.noreply.github.com>

Copilot-Session: eb11c85b-4046-4dbc-94b3-3bdbd4839a0d

Copilot AI left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Copilot review overview

🟡 Changes recommended

Unresolved MCP allow-list, cleanup, and fixture documentation findings remain.

Get a fresh assessment by requesting another Copilot review.

Review effort: Lite
Findings: 2 High severity

Open (2)
Resolved since last review (1)

Comment thread plugins/dotnet-msbuild/.claude-plugin/plugin.json Outdated
Comment thread plugins/dotnet-msbuild/plugin.json Outdated

Copilot AI left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Copilot review overview

🟡 Changes recommended

The unresolved critical routing conflict and moderate cleanup issue block approval.

Get a fresh assessment by requesting another Copilot review.

Review effort: Lite
Findings: 1 High severity

Open (1)
Resolved since last review (2)

Comment thread plugins/dotnet-msbuild/skills/directory-build-organization/SKILL.md Outdated
Allow diagnosis of existing Directory.Build import-timing defects without recommending new shared files for lone projects.

Co-authored-by: Copilot App <223556219+Copilot@users.noreply.github.com>

Copilot-Session: eb11c85b-4046-4dbc-94b3-3bdbd4839a0d
@github-actions

Copy link
Copy Markdown
Contributor

❌ Evaluation did not complete successfully (the evaluate job reported failure). Check the workflow run logs, then comment /evaluate 5a0ac39c6d4c1fd3cf30290b0075dab55baeb381 to retry this exact commit.

53 partial result file(s) were preserved for diagnosis but were not consolidated because the full matrix did not complete.

⚠️ Session replay telemetry was not published because the auxiliary dotnet/skills-data publisher failed. The evaluation verdicts above remain authoritative; maintainers must repair SKILLS_DATA_TOKEN or the publisher before replay links are available.

Copilot AI left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Copilot review overview

🔵 Needs a closer look

The build-perf-baseline routing boundary must include .slnx files.

Review effort: Lite
Findings: None

Resolved since last review (1)

Avoid duplicating BinlogMcp tool names while preserving overview and unsafe-tool permission assertions.

Co-authored-by: Copilot App <223556219+Copilot@users.noreply.github.com>

Copilot-Session: eb11c85b-4046-4dbc-94b3-3bdbd4839a0d

Copilot AI left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Copilot review overview

🔵 Needs a closer look

The broad change set still has an unresolved workspace-cleanup issue and a scope discrepancy in routing guidance.

Review effort: Lite
Findings: None

@github-actions

Copy link
Copy Markdown
Contributor

❌ Evaluation did not complete successfully (the evaluate job reported failure). Check the workflow run logs, then comment /evaluate 417727600696a84e7b6846f0a13a4da4a7b8f195 to retry this exact commit.

69 partial result file(s) were preserved for diagnosis but were not consolidated because the full matrix did not complete.

⚠️ Session replay telemetry was not published because the auxiliary dotnet/skills-data publisher failed. The evaluation verdicts above remain authoritative; maintainers must repair SKILLS_DATA_TOKEN or the publisher before replay links are available.

@github-actions

Copy link
Copy Markdown
Contributor

📊 Skill and Agent Evaluation Results

70 model/target results across 35 targets and 2 models — ✅ 21 improved, ➖ 48 not proven improved, ⚠️ 0 invalid or underpowered, ⛔ 1 activation contract failure, 📉 0 preference losses (report only).

Measurement identity: evaluated commit be46d44caa481ba5dab779b84f8e16377f83e12b; 2 judge models.

Measurement health: 70 expected / 70 observed / 70 written; 0 missing, 0 unexpected, 0 invalid; 1 recovered comparison error slot and 0 unresolved comparison error slots.

Objective completion gate: not enabled. Aggregate completion transitions are telemetry only, so this report does not claim that zero objective regressions were proven.

A result passes only when preference-eligible distinct-stimulus votes have aggregate net win of at least 20% and an exact one-sided sign-test result of p ≤ 0.05, and every explicit dormancy activation contract passes. Repeated runs measure reliability only.

Target Model Verdict Gate evidence Overfit Warnings Next action
agent.code-testing-generator claude-sonnet-5 ➖ Not proven improved n=5; 3W/1T/1L; d=4; p=0.312; net +40.0% — Activation: isolated 0/5; plugin 2/5 Inspect tied or lost stimuli; predeclare added breadth before a new experiment.
agent.code-testing-generator gpt-5.6-luna ➖ Not proven improved n=5; 3W/1T/1L; d=4; p=0.312; net +40.0% — — Inspect tied or lost stimuli; predeclare added breadth before a new experiment.
agent.msbuild claude-sonnet-5 ➖ Not proven improved n=5; 1W/3T/1L; d=2; p=0.750; net +0.0% — Activation: isolated 2/5; plugin 0/5 Inspect tied or lost stimuli; predeclare added breadth before a new experiment.
agent.msbuild gpt-5.6-luna ➖ Not proven improved n=5; 1W/3T/1L; d=2; p=0.750; net +0.0% — Activation: isolated 2/5; plugin 1/5 Inspect tied or lost stimuli; predeclare added breadth before a new experiment.
agent.test-quality-auditor claude-sonnet-5 ➖ Not proven improved n=5; 5W/0T/0L; d=5; p=0.031; net +100.0%; 1 dormancy excluded — Activation: isolated 0/5; plugin 0/5 Inspect tied or lost stimuli and fix inconsistent skill behavior.
agent.test-quality-auditor gpt-5.6-luna ➖ Not proven improved n=5; 1W/3T/1L; d=2; p=0.750; net +0.0%; 1 dormancy excluded — Activation: isolated 1/5; plugin 1/5 Inspect tied or lost stimuli; predeclare added breadth before a new experiment.
agent.testability-migration claude-sonnet-5 ➖ Not proven improved n=5; 3W/2T/0L; d=3; p=0.125; net +60.0% — Activation: isolated 0/5; plugin 0/5 Inspect tied or lost stimuli; predeclare added breadth before a new experiment.
agent.testability-migration gpt-5.6-luna ➖ Not proven improved n=5; 3W/1T/1L; d=4; p=0.312; net +40.0% — Activation: isolated 0/5; plugin 0/5 Inspect tied or lost stimuli; predeclare added breadth before a new experiment.
binlog-failure-analysis claude-sonnet-5 ➖ Not proven improved n=6; 0W/2T/4L; d=4; p=0.063; net -66.7%; 1 dormancy excluded 🟡 0.40 Activation: isolated 6/7; plugin 6/7 Inspect tied or lost stimuli; predeclare added breadth before a new experiment.
binlog-failure-analysis gpt-5.6-luna ➖ Not proven improved n=6; 2W/4T/0L; d=2; p=0.250; net +33.3%; 1 dormancy excluded 🟡 0.22 Activation: isolated 6/7; plugin 6/7 Inspect tied or lost stimuli; predeclare added breadth before a new experiment.
binlog-generation claude-sonnet-5 ➖ Not proven improved n=7; 4W/2T/1L; d=5; p=0.188; net +42.9%; 1 dormancy excluded 🟡 0.45 Activation: isolated 5/7; plugin 6/7 Inspect tied or lost stimuli and fix inconsistent skill behavior.
binlog-generation gpt-5.6-luna ➖ Not proven improved n=7; 4W/2T/1L; d=5; p=0.188; net +42.9%; 1 dormancy excluded 🟡 0.37 Activation: isolated 6/7; plugin 7/7; Activation-only stop: plugin 1 failed run Inspect activation-only failed runs before rewriting skill content; the model stopped after loading a skill.
build-parallelism claude-sonnet-5 ➖ Not proven improved n=6; 1W/3T/2L; d=3; p=0.500; net -16.7%; 1 dormancy excluded 🟡 0.20 Activation: isolated 5/6; plugin 5/6; Activation-only stop: isolated 1 failed run; Activation-only stop: plugin 1 failed run Inspect activation-only failed runs before rewriting skill content; the model stopped after loading a skill.
build-parallelism gpt-5.6-luna ➖ Not proven improved n=6; 4W/2T/0L; d=4; p=0.063; net +66.7%; 1 dormancy excluded 🟡 0.26 Activation-only stop: isolated 1 failed run; Activation-only stop: plugin 1 failed run Inspect activation-only failed runs before rewriting skill content; the model stopped after loading a skill.
build-perf-baseline claude-sonnet-5 ➖ Not proven improved n=6; 3W/2T/1L; d=4; p=0.312; net +33.3%; 1 dormancy excluded 🟡 0.34 Activation: isolated 4/6; plugin 5/6 Inspect tied or lost stimuli; predeclare added breadth before a new experiment.
build-perf-baseline gpt-5.6-luna ➖ Not proven improved n=6; 3W/1T/2L; d=5; p=0.500; net +16.7%; 1 dormancy excluded 🟡 0.38 Activation: isolated 5/6; plugin 6/6 Inspect tied or lost stimuli and fix inconsistent skill behavior.
build-perf-diagnostics claude-sonnet-5 ➖ Not proven improved n=7; 3W/4T/0L; d=3; p=0.125; net +42.9%; 1 dormancy excluded 🟡 0.45 — Inspect tied or lost stimuli; predeclare added breadth before a new experiment.
build-perf-diagnostics gpt-5.6-luna ➖ Not proven improved n=7; 3W/2T/2L; d=5; p=0.500; net +14.3%; 1 dormancy excluded 🟡 0.46 Activation: isolated 6/7; plugin 7/7 Inspect tied or lost stimuli and fix inconsistent skill behavior.
check-bin-obj-clash claude-sonnet-5 ➖ Not proven improved n=7; 2W/4T/1L; d=3; p=0.500; net +14.3% 🟡 0.21 Activation: isolated 7/7; plugin 5/7 Inspect tied or lost stimuli; predeclare added breadth before a new experiment.
check-bin-obj-clash gpt-5.6-luna ➖ Not proven improved n=7; 2W/5T/0L; d=2; p=0.250; net +28.6% ✅ 0.18 — Inspect tied or lost stimuli; predeclare added breadth before a new experiment.
code-testing-agent claude-sonnet-5 ➖ Not proven improved n=9; 5W/2T/2L; d=7; p=0.227; net +33.3% 🟡 0.45 Activation: isolated 5/9; plugin 5/9 Inspect tied or lost stimuli and fix inconsistent skill behavior.
code-testing-agent gpt-5.6-luna ✅ Improved n=9; 8W/1T/0L; d=8; p=0.004; net +88.9% — — None.
coverage-analysis claude-sonnet-5 ✅ Improved n=8; 7W/1T/0L; d=7; p=0.008; net +87.5%; 4 dormancy excluded 🟡 0.30 — Review overfit evidence.
coverage-analysis gpt-5.6-luna ✅ Improved n=8; 7W/0T/1L; d=8; p=0.035; net +75.0%; 4 dormancy excluded — Activation: isolated 7/8; plugin 7/8 Fix activation gaps.
crap-score claude-sonnet-5 ✅ Improved n=9; 5W/4T/0L; d=5; p=0.031; net +55.6% 🟡 0.32 — Review overfit evidence.
crap-score gpt-5.6-luna ✅ Improved n=9; 7W/1T/1L; d=8; p=0.035; net +66.7% 🟡 0.41 — Review overfit evidence.
detect-static-dependencies claude-sonnet-5 ✅ Improved n=7; 6W/1T/0L; d=6; p=0.016; net +85.7%; 1 dormancy excluded 🔴 0.53 — Review overfit evidence.
detect-static-dependencies gpt-5.6-luna ➖ Not proven improved n=7; 4W/2T/1L; d=5; p=0.188; net +42.9%; 1 dormancy excluded — 1 judge slot recovered Inspect tied or lost stimuli and fix inconsistent skill behavior.
directory-build-organization claude-sonnet-5 ➖ Not proven improved n=6; 2W/1T/3L; d=5; p=0.500; net -16.7%; 1 dormancy excluded 🟡 0.42 Activation: isolated 5/6; plugin 5/6 Inspect tied or lost stimuli and fix inconsistent skill behavior.
directory-build-organization gpt-5.6-luna ➖ Not proven improved n=6; 3W/2T/1L; d=4; p=0.312; net +33.3%; 1 dormancy excluded 🟡 0.43 — Inspect tied or lost stimuli; predeclare added breadth before a new experiment.
eval-performance claude-sonnet-5 ➖ Not proven improved n=8; 3W/3T/2L; d=5; p=0.500; net +12.5% 🟡 0.36 Activation: isolated 7/8; plugin 7/8 Inspect tied or lost stimuli and fix inconsistent skill behavior.
eval-performance gpt-5.6-luna ➖ Not proven improved n=8; 4W/3T/1L; d=5; p=0.188; net +37.5% 🟡 0.23 Activation: isolated 8/8; plugin 7/8 Inspect tied or lost stimuli and fix inconsistent skill behavior.
extension-points claude-sonnet-5 ➖ Not proven improved n=7; 4W/2T/1L; d=5; p=0.188; net +42.9%; 2 dormancy excluded 🟡 0.28 Activation: isolated 6/7; plugin 4/7 Inspect tied or lost stimuli and fix inconsistent skill behavior.
extension-points gpt-5.6-luna ➖ Not proven improved n=7; 2W/4T/1L; d=3; p=0.500; net +14.3%; 2 dormancy excluded 🟡 0.49 — Inspect tied or lost stimuli; predeclare added breadth before a new experiment.
find-untested-sources claude-sonnet-5 ✅ Improved n=9; 7W/1T/1L; d=8; p=0.035; net +66.7% 🔴 0.52 Activation: isolated 9/9; plugin 8/9 Fix activation gaps; Review overfit evidence.
find-untested-sources gpt-5.6-luna ✅ Improved n=9; 9W/0T/0L; d=9; p=0.002; net +100.0% — — None.
grade-tests claude-sonnet-5 ✅ Improved n=8; 8W/0T/0L; d=8; p=0.004; net +100.0% 🟡 0.29 — Review overfit evidence.
grade-tests gpt-5.6-luna ✅ Improved n=8; 7W/0T/1L; d=8; p=0.035; net +75.0% 🟡 0.47 — Review overfit evidence.
including-generated-files claude-sonnet-5 ➖ Not proven improved n=7; 3W/4T/0L; d=3; p=0.125; net +42.9%; 1 dormancy excluded 🟡 0.32 Activation: isolated 4/7; plugin 4/7 Inspect tied or lost stimuli; predeclare added breadth before a new experiment.
including-generated-files gpt-5.6-luna ➖ Not proven improved n=7; 4W/3T/0L; d=4; p=0.063; net +57.1%; 1 dormancy excluded 🟡 0.28 Activation: isolated 7/7; plugin 6/7 Inspect tied or lost stimuli; predeclare added breadth before a new experiment.
incremental-build claude-sonnet-5 ✅ Improved n=9; 5W/4T/0L; d=5; p=0.031; net +55.6% 🟡 0.31 Activation: isolated 5/9; plugin 4/9 Fix activation gaps; Review overfit evidence.
incremental-build gpt-5.6-luna ➖ Not proven improved n=9; 4W/2T/3L; d=7; p=0.500; net +11.1% 🟡 0.44 Activation: isolated 7/9; plugin 8/9 Inspect tied or lost stimuli and fix inconsistent skill behavior.
item-management claude-sonnet-5 ➖ Not proven improved n=6; 0W/4T/2L; d=2; p=0.250; net -33.3%; 1 dormancy excluded 🟡 0.35 Activation: isolated 4/6; plugin 4/6 Inspect tied or lost stimuli; predeclare added breadth before a new experiment.
item-management gpt-5.6-luna ➖ Not proven improved n=6; 1W/4T/1L; d=2; p=0.750; net +0.0%; 1 dormancy excluded ✅ 0.14 — Inspect tied or lost stimuli; predeclare added breadth before a new experiment.
migrate-static-to-wrapper claude-sonnet-5 ✅ Improved n=10; 7W/2T/1L; d=8; p=0.035; net +60.0% 🟡 0.31 — Review overfit evidence.
migrate-static-to-wrapper gpt-5.6-luna ✅ Improved n=10; 7W/3T/0L; d=7; p=0.008; net +70.0% 🔴 0.54 — Review overfit evidence.
msbuild-antipatterns claude-sonnet-5 ➖ Not proven improved n=7; 3W/3T/1L; d=4; p=0.312; net +28.6%; 1 dormancy excluded 🟡 0.38 Activation: isolated 3/7; plugin 3/7 Inspect tied or lost stimuli; predeclare added breadth before a new experiment.
msbuild-antipatterns gpt-5.6-luna ➖ Not proven improved n=7; 2W/3T/2L; d=4; p=0.687; net +0.0%; 1 dormancy excluded 🟡 0.47 Activation: isolated 5/7; plugin 2/7 Inspect tied or lost stimuli; predeclare added breadth before a new experiment.
msbuild-modernization claude-sonnet-5 ➖ Not proven improved n=6; 2W/4T/0L; d=2; p=0.250; net +33.3%; 1 dormancy excluded 🟡 0.23 — Inspect tied or lost stimuli; predeclare added breadth before a new experiment.
msbuild-modernization gpt-5.6-luna ➖ Not proven improved n=6; 1W/3T/2L; d=3; p=0.500; net -16.7%; 1 dormancy excluded ✅ 0.16 Activation: isolated 6/6; plugin 5/6 Inspect tied or lost stimuli; predeclare added breadth before a new experiment.
mtp-hot-reload claude-sonnet-5 ✅ Improved n=10; 9W/1T/0L; d=9; p=0.002; net +90.0%; 1 dormancy excluded 🟡 0.35 Activation-only stop: isolated 2 failed runs; Activation-only stop: plugin 2 failed runs Inspect activation-only failed runs before rewriting skill content; the model stopped after loading a skill.
mtp-hot-reload gpt-5.6-luna ✅ Improved n=10; 7W/3T/0L; d=7; p=0.008; net +70.0%; 1 dormancy excluded 🟡 0.25 Activation-only stop: isolated 2 failed runs Inspect activation-only failed runs before rewriting skill content; the model stopped after loading a skill.
platform-detection claude-sonnet-5 ✅ Improved n=15; 10W/4T/1L; d=11; p=0.006; net +60.0% 🟡 0.22 — Review overfit evidence.
platform-detection gpt-5.6-luna ➖ Not proven improved n=15; 5W/7T/3L; d=8; p=0.363; net +13.3% — — Inspect tied or lost stimuli and fix inconsistent skill behavior.
property-patterns claude-sonnet-5 ➖ Not proven improved n=7; 3W/1T/3L; d=6; p=0.656; net +0.0% 🟡 0.21 Activation: isolated 6/7; plugin 7/7 Inspect tied or lost stimuli and fix inconsistent skill behavior.
property-patterns gpt-5.6-luna ➖ Not proven improved n=7; 3W/2T/2L; d=5; p=0.500; net +14.3% ✅ 0.17 Activation: isolated 6/7; plugin 7/7 Inspect tied or lost stimuli and fix inconsistent skill behavior.
resolve-project-references claude-sonnet-5 ➖ Not proven improved n=6; 1W/4T/1L; d=2; p=0.750; net +0.0%; 2 dormancy excluded 🟡 0.36 Activation: isolated 3/6; plugin 4/6; Activation-only stop: plugin 1 failed run Inspect activation-only failed runs before rewriting skill content; the model stopped after loading a skill.
resolve-project-references gpt-5.6-luna ➖ Not proven improved n=6; 2W/3T/1L; d=3; p=0.500; net +16.7%; 2 dormancy excluded 🟡 0.44 Activation: isolated 5/6; plugin 6/6; Activation-only stop: plugin 2 failed runs Inspect activation-only failed runs before rewriting skill content; the model stopped after loading a skill.
run-tests claude-sonnet-5 ✅ Improved n=22; 12W/8T/2L; d=14; p=0.006; net +45.5% 🟡 0.42 Activation-only stop: isolated 2 failed runs; Activation-only stop: plugin 2 failed runs Inspect activation-only failed runs before rewriting skill content; the model stopped after loading a skill.
run-tests gpt-5.6-luna ➖ Not proven improved n=22; 7W/10T/5L; d=12; p=0.387; net +9.1% — Activation-only stop: isolated 3 failed runs; Activation-only stop: plugin 2 failed runs Inspect activation-only failed runs before rewriting skill content; the model stopped after loading a skill.
target-authoring claude-sonnet-5 ➖ Not proven improved n=6; 1W/1T/4L; d=5; p=0.188; net -50.0%; 1 dormancy excluded 🟡 0.33 Activation: isolated 6/6; plugin 5/6 Inspect tied or lost stimuli and fix inconsistent skill behavior.
target-authoring gpt-5.6-luna ➖ Not proven improved n=6; 2W/4T/0L; d=2; p=0.250; net +33.3%; 1 dormancy excluded 🟡 0.28 — Inspect tied or lost stimuli; predeclare added breadth before a new experiment.
test-anti-patterns claude-sonnet-5 ✅ Improved n=8; 7W/1T/0L; d=7; p=0.008; net +87.5%; 1 dormancy excluded 🟡 0.22 — Review overfit evidence.
test-anti-patterns gpt-5.6-luna ⛔ Activation contract failed n=8; 6W/1T/1L; d=7; p=0.063; net +62.5%; 1 dormancy excluded 🔴 0.52 Dormancy contract: 1 unexpected activation(s) Narrow skill routing so the listed off-target scenarios stay dormant.
test-gap-analysis claude-sonnet-5 ✅ Improved n=8; 6W/2T/0L; d=6; p=0.016; net +75.0%; 2 dormancy excluded 🟡 0.42 — Review overfit evidence.
test-gap-analysis gpt-5.6-luna ➖ Not proven improved n=8; 5W/2T/1L; d=6; p=0.109; net +50.0%; 2 dormancy excluded 🔴 0.56 — Inspect tied or lost stimuli and fix inconsistent skill behavior.
test-smell-detection claude-sonnet-5 ✅ Improved n=10; 8W/1T/1L; d=9; p=0.020; net +70.0% 🟡 0.31 Activation-only stop: isolated 1 failed run Inspect activation-only failed runs before rewriting skill content; the model stopped after loading a skill.
test-smell-detection gpt-5.6-luna ➖ Not proven improved n=10; 5W/4T/1L; d=6; p=0.109; net +40.0% — — Inspect tied or lost stimuli and fix inconsistent skill behavior.
writing-mstest-tests claude-sonnet-5 ✅ Improved n=15; 11W/3T/1L; d=12; p=0.003; net +66.7% 🔴 0.53 Activation-only stop: isolated 2 failed runs; Activation-only stop: plugin 3 failed runs Inspect activation-only failed runs before rewriting skill content; the model stopped after loading a skill.
writing-mstest-tests gpt-5.6-luna ➖ Not proven improved n=15; 8W/4T/3L; d=11; p=0.113; net +33.3% ✅ 0.00 Activation-only stop: isolated 3 failed runs Inspect activation-only failed runs before rewriting skill content; the model stopped after loading a skill.
ℹ️ How to read this report
  • ✅ Improved — the result passed both the statistical gate and the 20% practical net-win floor.
  • ➖ Not proven improved — the result is valid but did not pass both gates. This is not automatically a regression.
  • ⚠️ Invalid / underpowered — the gate withheld a quality verdict. Fix the measurement before judging the target.
  • ⛔ Activation contract failed — the isolated target activated on an explicit dormancy scenario. Dormancy preference is excluded, but this routing failure still blocks a pass.
  • 📉 Preference loss — the LLM judge credibly preferred baseline. It is report-only, not objective completion proof.
  • Gate evidence — n preference-eligible distinct-stimulus votes, W/T/L stimulus votes, d discordant votes, exact one-sided p, net win, and the count of separately retained dormancy stimuli. The p value applies to one model/target result; no matrix-wide multiple-comparison correction is applied.
  • Overfit — overfitting-judge severity (✅ Low, 🟡 Moderate, 🔴 High, — none) and score.
  • Warnings — activation, timeout, retry recovery, or unresolved comparison conditions that need attention.
  • Do not add repeated runs to increase statistical power. Do not add stimuli after seeing a near-pass unless the new breadth is predeclared for a new experiment.
⛔ Activation contract failed — test-anti-patterns (gpt-5.6-luna)

Why: Net win +62.5% (6W/1T/1L over 8 preference-eligible stimulus vote(s), sign test p=0.063), mean preference +17.8% across 9 paired run(s), 1 dormancy stimulus/stimuli excluded from preference — activation contract failed (1 explicit dormancy scenario(s) activated the isolated target skill)

Next action: Narrow skill routing so the listed off-target scenarios stay dormant.

State: VALID_NO_CHANGE (activation_contract_failed)

Gate evidence: n=8; 6W/1T/1L; d=7; p=0.063; net +62.5%; 1 dormancy excluded

Warnings: Dormancy contract: 1 unexpected activation(s)

Overfit: High (score 0.52)

Repeated-run reliability (not used by the gate): 9 paired runs (6W/1T/2L).

Weak or warning scenarios:

Scenario Preference gate Net win Δ Pref Runs (W/T/L)
▼ Audit a pytest suite using Python-specific anti-pattern markers Eligible -100.0% -40.0% 0/0/1
= Detect coverage-touching pattern across a service facade Eligible +0.0% +0.0% 0/1/0
▼ Stay dormant for test trait distribution Excluded (activation contract) -100.0% -40.0% 0/0/1

Illustrative judge evidence:

  • Audit a pytest suite using Python-specific anti-pattern markers: Response A is more complete and directly addresses the task's core ask: ranking findings by how badly they hurt. It identifies the pytest dependency issue (a tier-1 blocker preventing test execution), applies explicit severity ranking, and accounts for all 9 distinct problems ...

This is one example, not the aggregate verdict. Open Full Results for every judgment.

➖ Not proven improved — agent.code-testing-generator (claude-sonnet-5)

Why: Net win +40.0% (3W/1T/1L over 5 preference-eligible stimulus vote(s), sign test p=0.312), mean preference +16.0% across 5 paired run(s) — not credible — 1 of 5 preference-eligible stimulus vote(s) tied, leaving only 4 discordant preference vote(s). The sign test conditions on non-tie stimulus votes and cannot reach 0.05 below 5, so no record could have passed here — this is not a measured null. Either the agent is inert on these scenarios (make them discriminate) or the eval needs more distinct stimuli to clear the ties — native evaluator reported that the target agent did not activate

Next action: Inspect tied or lost stimuli; predeclare added breadth before a new experiment.

State: VALID_NO_CHANGE (target_agent_not_activated)

Gate evidence: n=5; 3W/1T/1L; d=4; p=0.312; net +40.0%

Warnings: Activation: isolated 0/5; plugin 2/5

Repeated-run reliability (not used by the gate): 5 paired runs (3W/1T/1L).

Weak or warning scenarios:

Scenario Preference gate Net win Δ Pref Runs (W/T/L)
▲ Generate a project-wide pytest suite across modules Eligible +100.0% +40.0% 1/0/0
▲ Generate collaborating Go package tests Eligible +100.0% +40.0% 1/0/0
= Generate layered Vitest coverage for an async cart Eligible +0.0% +0.0% 0/1/0
▼ Generate project-wide xUnit tests for a .NET library Eligible -100.0% -40.0% 0/0/1
▲ Preserve a classic MSTest project while adding broad coverage Eligible +100.0% +40.0% 1/0/0

Illustrative judge evidence:

  • Generate layered Vitest coverage for an async cart: Position-swap inconsistent (forward: skill, reverse: baseline). Defaulting to tie.

This is one example, not the aggregate verdict. Open Full Results for every judgment.

➖ Not proven improved — agent.code-testing-generator (gpt-5.6-luna)

Why: Net win +40.0% (3W/1T/1L over 5 preference-eligible stimulus vote(s), sign test p=0.312), mean preference +16.0% across 5 paired run(s) — not credible — 1 of 5 preference-eligible stimulus vote(s) tied, leaving only 4 discordant preference vote(s). The sign test conditions on non-tie stimulus votes and cannot reach 0.05 below 5, so no record could have passed here — this is not a measured null. Either the agent is inert on these scenarios (make them discriminate) or the eval needs more distinct stimuli to clear the ties

Next action: Inspect tied or lost stimuli; predeclare added breadth before a new experiment.

State: VALID_NO_CHANGE (no_credible_preference_change)

Gate evidence: n=5; 3W/1T/1L; d=4; p=0.312; net +40.0%

Repeated-run reliability (not used by the gate): 5 paired runs (3W/1T/1L).

Weak or warning scenarios:

Scenario Preference gate Net win Δ Pref Runs (W/T/L)
= Generate collaborating Go package tests Eligible +0.0% +0.0% 0/1/0
▼ Preserve a classic MSTest project while adding broad coverage Eligible -100.0% -40.0% 0/0/1

Illustrative judge evidence:

  • Generate collaborating Go package tests: Position-swap inconsistent (forward: tie, reverse: baseline). Defaulting to tie.

This is one example, not the aggregate verdict. Open Full Results for every judgment.

➖ Not proven improved — agent.msbuild (claude-sonnet-5)

Why: Net win +0.0% (1W/3T/1L over 5 preference-eligible stimulus vote(s), sign test p=0.750), mean preference +0.0% across 5 paired run(s) — no improvement — native evaluator reported that the target agent did not activate

Next action: Inspect tied or lost stimuli; predeclare added breadth before a new experiment.

State: VALID_NO_CHANGE (target_agent_not_activated)

Gate evidence: n=5; 1W/3T/1L; d=2; p=0.750; net +0.0%

Warnings: Activation: isolated 2/5; plugin 0/5

Repeated-run reliability (not used by the gate): 5 paired runs (1W/3T/1L).

Weak or warning scenarios:

Scenario Preference gate Net win Δ Pref Runs (W/T/L)
= Advise on project file organization Eligible +0.0% +0.0% 0/1/0
= Diagnose broken incremental build behavior Eligible +0.0% +0.0% 0/1/0
▲ Review a project file for maintainability risks Eligible +100.0% +40.0% 1/0/0
▼ Route a slow build to performance analysis Eligible -100.0% -40.0% 0/0/1
= Triage a build failure and route to appropriate analysis Eligible +0.0% +0.0% 0/1/0

Illustrative judge evidence:

  • Advise on project file organization: Position-swap inconsistent (forward: baseline, reverse: skill). Defaulting to tie.

This is one example, not the aggregate verdict. Open Full Results for every judgment.

➖ Not proven improved — agent.msbuild (gpt-5.6-luna)

Why: Net win +0.0% (1W/3T/1L over 5 preference-eligible stimulus vote(s), sign test p=0.750), mean preference -12.0% across 5 paired run(s) — no improvement — native evaluator reported that the target agent did not activate

Next action: Inspect tied or lost stimuli; predeclare added breadth before a new experiment.

State: VALID_NO_CHANGE (target_agent_not_activated)

Gate evidence: n=5; 1W/3T/1L; d=2; p=0.750; net +0.0%

Warnings: Activation: isolated 2/5; plugin 1/5

Repeated-run reliability (not used by the gate): 5 paired runs (1W/3T/1L).

Weak or warning scenarios:

Scenario Preference gate Net win Δ Pref Runs (W/T/L)
= Advise on project file organization Eligible +0.0% +0.0% 0/1/0
▲ Diagnose broken incremental build behavior Eligible +100.0% +40.0% 1/0/0
▼ Review a project file for maintainability risks Eligible -100.0% -100.0% 0/0/1
= Route a slow build to performance analysis Eligible +0.0% +0.0% 0/1/0
= Triage a build failure and route to appropriate analysis Eligible +0.0% +0.0% 0/1/0

Illustrative judge evidence:

  • Advise on project file organization: Position-swap inconsistent (forward: tie, reverse: baseline). Defaulting to tie.

This is one example, not the aggregate verdict. Open Full Results for every judgment.

➖ Not proven improved — agent.test-quality-auditor (claude-sonnet-5)

Why: Net win +100.0% (5W/0T/0L over 5 preference-eligible stimulus vote(s), sign test p=0.031), mean preference +26.7% across 6 paired run(s), 1 dormancy stimulus/stimuli excluded from preference — credibly better — native evaluator reported that the target agent did not activate

Next action: Inspect tied or lost stimuli and fix inconsistent skill behavior.

State: VALID_NO_CHANGE (target_agent_not_activated)

Gate evidence: n=5; 5W/0T/0L; d=5; p=0.031; net +100.0%; 1 dormancy excluded

Warnings: Activation: isolated 0/5; plugin 0/5

Repeated-run reliability (not used by the gate): 6 paired runs (5W/0T/1L).

Weak or warning scenarios:

Scenario Preference gate Net win Δ Pref Runs (W/T/L)
▲ Assertion quality analysis Eligible +100.0% +40.0% 1/0/0
▲ Comprehensive test quality audit of weak test suite Eligible +100.0% +40.0% 1/0/0
▼ Decline request to generate new tests Excluded (activation contract) -100.0% -100.0% 0/0/1
▲ Diagnose test smells and propose a repair order Eligible +100.0% +40.0% 1/0/0
▲ Identify behavior gaps that existing tests would miss Eligible +100.0% +40.0% 1/0/0
▲ Targeted anti-pattern review Eligible +100.0% +100.0% 1/0/0

Illustrative judge evidence:

  • Decline request to generate new tests: A demonstrated substantially better contextual judgment by preserving the audit fixture while still producing and validating a separate comprehensive suite. B's tests may be reasonable, but modifying the deliberately weak fixture and its csproj is inappropriate for this reposi...

This is one example, not the aggregate verdict. Open Full Results for every judgment.

➖ Not proven improved — agent.test-quality-auditor (gpt-5.6-luna)

Why: Net win +0.0% (1W/3T/1L over 5 preference-eligible stimulus vote(s), sign test p=0.750), mean preference +10.0% across 6 paired run(s), 1 dormancy stimulus/stimuli excluded from preference — no improvement — native evaluator reported that the target agent did not activate

Next action: Inspect tied or lost stimuli; predeclare added breadth before a new experiment.

State: VALID_NO_CHANGE (target_agent_not_activated)

Gate evidence: n=5; 1W/3T/1L; d=2; p=0.750; net +0.0%; 1 dormancy excluded

Warnings: Activation: isolated 1/5; plugin 1/5

Repeated-run reliability (not used by the gate): 6 paired runs (1W/4T/1L).

Weak or warning scenarios:

Scenario Preference gate Net win Δ Pref Runs (W/T/L)
= Assertion quality analysis Eligible +0.0% +0.0% 0/1/0
▼ Comprehensive test quality audit of weak test suite Eligible -100.0% -40.0% 0/0/1
= Decline request to generate new tests Excluded (activation contract) +0.0% +0.0% 0/1/0
= Diagnose test smells and propose a repair order Eligible +0.0% +0.0% 0/1/0
= Identify behavior gaps that existing tests would miss Eligible +0.0% +0.0% 0/1/0
▲ Targeted anti-pattern review Eligible +100.0% +100.0% 1/0/0

Illustrative judge evidence:

  • Assertion quality analysis: Position-swap inconsistent (forward: skill, reverse: baseline). Defaulting to tie.

This is one example, not the aggregate verdict. Open Full Results for every judgment.

➖ Not proven improved — agent.testability-migration (claude-sonnet-5)

Why: Net win +60.0% (3W/2T/0L over 5 preference-eligible stimulus vote(s), sign test p=0.125), mean preference +24.0% across 5 paired run(s) — not credible — 2 of 5 preference-eligible stimulus vote(s) tied, leaving only 3 discordant preference vote(s). The sign test conditions on non-tie stimulus votes and cannot reach 0.05 below 5, so no record could have passed here — this is not a measured null. Either the agent is inert on these scenarios (make them discriminate) or the eval needs more distinct stimuli to clear the ties — native evaluator reported that the target agent did not activate

Next action: Inspect tied or lost stimuli; predeclare added breadth before a new experiment.

State: VALID_NO_CHANGE (target_agent_not_activated)

Gate evidence: n=5; 3W/2T/0L; d=3; p=0.125; net +60.0%

Warnings: Activation: isolated 0/5; plugin 0/5

Repeated-run reliability (not used by the gate): 5 paired runs (3W/2T/0L).

Weak or warning scenarios:

Scenario Preference gate Net win Δ Pref Runs (W/T/L)
▲ Full pipeline: detect statics and recommend migration plan Eligible +100.0% +40.0% 1/0/0
▲ Inventory static dependencies without modifying the project Eligible +100.0% +40.0% 1/0/0
= Migrate time dependencies and add deterministic tests Eligible +0.0% +0.0% 0/1/0
= Replace filesystem statics without touching unrelated dependencies Eligible +0.0% +0.0% 0/1/0
▲ Targeted request: just migrate DateTime to TimeProvider Eligible +100.0% +40.0% 1/0/0

Illustrative judge evidence:

  • Migrate time dependencies and add deterministic tests: Position-swap inconsistent (forward: skill, reverse: baseline). Defaulting to tie.

This is one example, not the aggregate verdict. Open Full Results for every judgment.

➖ Not proven improved — agent.testability-migration (gpt-5.6-luna)

Why: Net win +40.0% (3W/1T/1L over 5 preference-eligible stimulus vote(s), sign test p=0.312), mean preference +28.0% across 5 paired run(s) — not credible — 1 of 5 preference-eligible stimulus vote(s) tied, leaving only 4 discordant preference vote(s). The sign test conditions on non-tie stimulus votes and cannot reach 0.05 below 5, so no record could have passed here — this is not a measured null. Either the agent is inert on these scenarios (make them discriminate) or the eval needs more distinct stimuli to clear the ties — native evaluator reported that the target agent did not activate

Next action: Inspect tied or lost stimuli; predeclare added breadth before a new experiment.

State: VALID_NO_CHANGE (target_agent_not_activated)

Gate evidence: n=5; 3W/1T/1L; d=4; p=0.312; net +40.0%

Warnings: Activation: isolated 0/5; plugin 0/5

Repeated-run reliability (not used by the gate): 5 paired runs (3W/1T/1L).

Weak or warning scenarios:

Scenario Preference gate Net win Δ Pref Runs (W/T/L)
▼ Full pipeline: detect statics and recommend migration plan Eligible -100.0% -40.0% 0/0/1
= Inventory static dependencies without modifying the project Eligible +0.0% +0.0% 0/1/0
▲ Migrate time dependencies and add deterministic tests Eligible +100.0% +100.0% 1/0/0
▲ Replace filesystem statics without touching unrelated dependencies Eligible +100.0% +40.0% 1/0/0
▲ Targeted request: just migrate DateTime to TimeProvider Eligible +100.0% +40.0% 1/0/0

Illustrative judge evidence:

  • Full pipeline: detect statics and recommend migration plan: While Response B excels at structured analysis with precise line numbers and recommends pragmatic standard libraries (TimeProvider, System.IO.Abstractions), Response A better serves the specific request for a 'bounded' migration plan. Response A identifies the dual-read clock ...

This is one example, not the aggregate verdict. Open Full Results for every judgment.

➖ Not proven improved — binlog-failure-analysis (claude-sonnet-5)

Why: Net win -66.7% (0W/2T/4L over 6 preference-eligible stimulus vote(s), sign test p=0.063), mean preference -22.9% across 7 paired run(s), 1 dormancy stimulus/stimuli excluded from preference — no improvement

Next action: Inspect tied or lost stimuli; predeclare added breadth before a new experiment.

State: VALID_NO_CHANGE (no_credible_preference_change)

Gate evidence: n=6; 0W/2T/4L; d=4; p=0.063; net -66.7%; 1 dormancy excluded

Warnings: Activation: isolated 6/7; plugin 6/7

Overfit: Moderate (score 0.40)

Repeated-run reliability (not used by the gate): 7 paired runs (0W/3T/4L).

Weak or warning scenarios:

Scenario Preference gate Net win Δ Pref Runs (W/T/L)
▼ Assess a requested binlog investigation when the current build is healthy Eligible -100.0% -40.0% 0/0/1
= Confirm the actual resolved target framework and package version from a binlog Eligible +0.0% +0.0% 0/0/0
▼ Determine whether a quiet second build actually failed Eligible -100.0% -40.0% 0/0/1
= Diagnose a warning behind a build that actually succeeded Eligible +0.0% +0.0% 0/1/0
= Diagnose build failures from binlog only (no source files) Eligible +0.0% +0.0% 0/1/0
= Stay dormant for a non-MSBuild build failure log Excluded (activation contract) +0.0% +0.0% 0/1/0
▼ Trace why a generated source file is missing at compile time Eligible -100.0% -40.0% 0/0/1
▼ Use capture-time text logs when the binlog reader is unavailable Eligible -100.0% -40.0% 0/0/1

Illustrative judge evidence:

  • Assess a requested binlog investigation when the current build is healthy: Both are correct and satisfy the central request by capturing a binlog and establishing that the currently staged project is healthy. A is modestly stronger because its final answer gives a more substantive binlog-based analysis, including separate error/warning checks and tim...

This is one example, not the aggregate verdict. Open Full Results for every judgment.

➖ Not proven improved — binlog-failure-analysis (gpt-5.6-luna)

Why: Net win +33.3% (2W/4T/0L over 6 preference-eligible stimulus vote(s), sign test p=0.250), mean preference +11.4% across 7 paired run(s), 1 dormancy stimulus/stimuli excluded from preference — not credible — 4 of 6 preference-eligible stimulus vote(s) tied, leaving only 2 discordant preference vote(s). The sign test conditions on non-tie stimulus votes and cannot reach 0.05 below 5, so no record could have passed here — this is not a measured null. Either the skill is inert on these scenarios (make them discriminate) or the eval needs more distinct stimuli to clear the ties

Next action: Inspect tied or lost stimuli; predeclare added breadth before a new experiment.

State: VALID_NO_CHANGE (no_credible_preference_change)

Gate evidence: n=6; 2W/4T/0L; d=2; p=0.250; net +33.3%; 1 dormancy excluded

Warnings: Activation: isolated 6/7; plugin 6/7

Overfit: Moderate (score 0.22)

Repeated-run reliability (not used by the gate): 7 paired runs (2W/5T/0L).

Weak or warning scenarios:

Scenario Preference gate Net win Δ Pref Runs (W/T/L)
= Assess a requested binlog investigation when the current build is healthy Eligible +0.0% +0.0% 0/1/0
= Confirm the actual resolved target framework and package version from a binlog Eligible +0.0% +0.0% 0/0/0
= Determine whether a quiet second build actually failed Eligible +0.0% +0.0% 0/1/0
= Diagnose a warning behind a build that actually succeeded Eligible +0.0% +0.0% 0/1/0
= Diagnose build failures from binlog only (no source files) Eligible +0.0% +0.0% 0/1/0
= Stay dormant for a non-MSBuild build failure log Excluded (activation contract) +0.0% +0.0% 0/1/0

Illustrative judge evidence:

  • Assess a requested binlog investigation when the current build is healthy: Position-swap inconsistent (forward: tie, reverse: A). Defaulting to tie.

This is one example, not the aggregate verdict. Open Full Results for every judgment.

➖ Not proven improved — binlog-generation (claude-sonnet-5)

Why: Net win +42.9% (4W/2T/1L over 7 preference-eligible stimulus vote(s), sign test p=0.188), mean preference +37.5% across 8 paired run(s), 1 dormancy stimulus/stimuli excluded from preference — not credible (sign test p=0.188 > 0.05)

Next action: Inspect tied or lost stimuli and fix inconsistent skill behavior.

State: VALID_NO_CHANGE (no_credible_preference_change)

Gate evidence: n=7; 4W/2T/1L; d=5; p=0.188; net +42.9%; 1 dormancy excluded

Warnings: Activation: isolated 5/7; plugin 6/7

Overfit: Moderate (score 0.45)

Repeated-run reliability (not used by the gate): 8 paired runs (4W/3T/1L).

Weak or warning scenarios:

Scenario Preference gate Net win Δ Pref Runs (W/T/L)
= Choose a predictable, non-colliding binlog name for a CI upload step Eligible +0.0% +0.0% 0/1/0
= Decline Maven compiler configuration Excluded (activation contract) +0.0% +0.0% 0/1/0
▼ Preserve binlog history while cleaning stale build output Eligible -100.0% -40.0% 0/0/1
= Recognize that a failed build produced no binlog at all Eligible +0.0% +0.0% 0/1/0

Illustrative judge evidence:

  • Choose a predictable, non-colliding binlog name for a CI upload step: Both responses fully satisfy the task with the same valid, explicit non-colliding filenames and successful Debug and Release binlog builds. B provides an absolute directory in its final message, but that is not a material quality advantage over A's equally clear result.

This is one example, not the aggregate verdict. Open Full Results for every judgment.

➖ Not proven improved — binlog-generation (gpt-5.6-luna)

Why: Net win +42.9% (4W/2T/1L over 7 preference-eligible stimulus vote(s), sign test p=0.188), mean preference +37.5% across 8 paired run(s), 1 dormancy stimulus/stimuli excluded from preference — not credible (sign test p=0.188 > 0.05)

Next action: Inspect activation-only failed runs before rewriting skill content; the model stopped after loading a skill.

State: VALID_NO_CHANGE (no_credible_preference_change)

Gate evidence: n=7; 4W/2T/1L; d=5; p=0.188; net +42.9%; 1 dormancy excluded

Warnings: Activation: isolated 6/7; plugin 7/7; Activation-only stop: plugin 1 failed run

Overfit: Moderate (score 0.37)

Repeated-run reliability (not used by the gate): 8 paired runs (4W/3T/1L).

Weak or warning scenarios:

Scenario Preference gate Net win Δ Pref Runs (W/T/L)
= Choose a predictable, non-colliding binlog name for a CI upload step Eligible +0.0% +0.0% 0/1/0
= Decline Maven compiler configuration Excluded (activation contract) +0.0% +0.0% 0/1/0
▼ Preserve binlog history while cleaning stale build output Eligible -100.0% -40.0% 0/0/1
= Recognize that a failed build produced no binlog at all Eligible +0.0% +0.0% 0/1/0

Illustrative judge evidence:

  • Choose a predictable, non-colliding binlog name for a CI upload step: Both responses successfully complete the task with identical, correct final outputs: two non-colliding binlogs (debug-4.binlog and release-4.binlog) that follow the manifest's naming scheme and are ready for CI use. Response B is marginally more explicit in confirming the mani...

This is one example, not the aggregate verdict. Open Full Results for every judgment.

➖ Not proven improved — build-parallelism (claude-sonnet-5)

Why: Net win -16.7% (1W/3T/2L over 6 preference-eligible stimulus vote(s), sign test p=0.500), mean preference +0.0% across 7 paired run(s), 1 dormancy stimulus/stimuli excluded from preference — no improvement

Next action: Inspect activation-only failed runs before rewriting skill content; the model stopped after loading a skill.

State: VALID_NO_CHANGE (no_credible_preference_change)

Gate evidence: n=6; 1W/3T/2L; d=3; p=0.500; net -16.7%; 1 dormancy excluded

Warnings: Activation: isolated 5/6; plugin 5/6; Activation-only stop: isolated 1 failed run; Activation-only stop: plugin 1 failed run

Overfit: Moderate (score 0.20)

Repeated-run reliability (not used by the gate): 7 paired runs (2W/3T/2L).

Weak or warning scenarios:

Scenario Preference gate Net win Δ Pref Runs (W/T/L)
▼ Analyze build parallelism bottlenecks Eligible -100.0% -40.0% 0/0/1
= Decline graph build for runtime-discovered projects Eligible +0.0% +0.0% 0/1/0
= Enable BuildInParallel on a custom MSBuild task Eligible +0.0% +0.0% 0/1/0
= Enable parallel nodes for a wide project graph Eligible +0.0% +0.0% 0/1/0
▼ Reduce CI build scope with a solution filter Eligible -100.0% -40.0% 0/0/1

Illustrative judge evidence:

  • Analyze build parallelism bottlenecks: The substantive conclusions are equally correct, but A better substantiates the critical-path claim with concrete timing/scheduling evidence and more explicitly addresses why parallel nodes do not help.

This is one example, not the aggregate verdict. Open Full Results for every judgment.

➖ Not proven improved — build-parallelism (gpt-5.6-luna)

Why: Net win +66.7% (4W/2T/0L over 6 preference-eligible stimulus vote(s), sign test p=0.063), mean preference +25.7% across 7 paired run(s), 1 dormancy stimulus/stimuli excluded from preference — not credible — 2 of 6 preference-eligible stimulus vote(s) tied, leaving only 4 discordant preference vote(s). The sign test conditions on non-tie stimulus votes and cannot reach 0.05 below 5, so no record could have passed here — this is not a measured null. Either the skill is inert on these scenarios (make them discriminate) or the eval needs more distinct stimuli to clear the ties

Next action: Inspect activation-only failed runs before rewriting skill content; the model stopped after loading a skill.

State: VALID_NO_CHANGE (no_credible_preference_change)

Gate evidence: n=6; 4W/2T/0L; d=4; p=0.063; net +66.7%; 1 dormancy excluded

Warnings: Activation-only stop: isolated 1 failed run; Activation-only stop: plugin 1 failed run

Overfit: Moderate (score 0.26)

Repeated-run reliability (not used by the gate): 7 paired runs (4W/2T/1L).

Weak or warning scenarios:

Scenario Preference gate Net win Δ Pref Runs (W/T/L)
▼ Decline database query tuning request Excluded (activation contract) -100.0% -40.0% 0/0/1
= Decline graph build for runtime-discovered projects Eligible +0.0% +0.0% 0/1/0
▲ Enable parallel nodes for a wide project graph Eligible +100.0% +40.0% 1/0/0
= Reduce CI build scope with a solution filter Eligible +0.0% +0.0% 0/1/0

Illustrative judge evidence:

  • Decline graph build for runtime-discovered projects: Position-swap inconsistent (forward: B, reverse: A). Defaulting to tie.

This is one example, not the aggregate verdict. Open Full Results for every judgment.

➖ Not proven improved — build-perf-baseline (claude-sonnet-5)

Why: Net win +33.3% (3W/2T/1L over 6 preference-eligible stimulus vote(s), sign test p=0.312), mean preference +5.7% across 7 paired run(s), 1 dormancy stimulus/stimuli excluded from preference — not credible — 2 of 6 preference-eligible stimulus vote(s) tied, leaving only 4 discordant preference vote(s). The sign test conditions on non-tie stimulus votes and cannot reach 0.05 below 5, so no record could have passed here — this is not a measured null. Either the skill is inert on these scenarios (make them discriminate) or the eval needs more distinct stimuli to clear the ties

Next action: Inspect tied or lost stimuli; predeclare added breadth before a new experiment.

State: VALID_NO_CHANGE (no_credible_preference_change)

Gate evidence: n=6; 3W/2T/1L; d=4; p=0.312; net +33.3%; 1 dormancy excluded

Warnings: Activation: isolated 4/6; plugin 5/6

Overfit: Moderate (score 0.34)

Repeated-run reliability (not used by the gate): 7 paired runs (3W/2T/2L).

Weak or warning scenarios:

Scenario Preference gate Net win Δ Pref Runs (W/T/L)
▼ Adopt UseArtifactsOutput while declining graph build for a small solution Eligible -100.0% -40.0% 0/0/1
= Configure deterministic, cache-safe CI builds Eligible +0.0% +0.0% 0/1/0
▼ Decline a non-MSBuild build performance request Excluded (activation contract) -100.0% -40.0% 0/0/1
= Route a restore-bound cold build away from architecture changes Eligible +0.0% +0.0% 0/1/0

Illustrative judge evidence:

  • Adopt UseArtifactsOutput while declining graph build for a small solution: Both provide the essential artifacts-layout recommendation, but A is more internally consistent and directly answers that /graph should not be enabled merely alongside artifacts output for a five-project solution. B's subsequent caveat improves its advice, but conflicts with i...

This is one example, not the aggregate verdict. Open Full Results for every judgment.

➖ Not proven improved — build-perf-baseline (gpt-5.6-luna)

Why: Net win +16.7% (3W/1T/2L over 6 preference-eligible stimulus vote(s), sign test p=0.500), mean preference +8.6% across 7 paired run(s), 1 dormancy stimulus/stimuli excluded from preference — not credible (sign test p=0.500 > 0.05)

Next action: Inspect tied or lost stimuli and fix inconsistent skill behavior.

State: VALID_NO_CHANGE (no_credible_preference_change)

Gate evidence: n=6; 3W/1T/2L; d=5; p=0.500; net +16.7%; 1 dormancy excluded

Warnings: Activation: isolated 5/6; plugin 6/6

Overfit: Moderate (score 0.38)

Repeated-run reliability (not used by the gate): 7 paired runs (3W/1T/3L).

Weak or warning scenarios:

Scenario Preference gate Net win Δ Pref Runs (W/T/L)
= Configure deterministic, cache-safe CI builds Eligible +0.0% +0.0% 0/1/0
▼ Decline a non-MSBuild build performance request Excluded (activation contract) -100.0% -40.0% 0/0/1
▼ Leave an already-optimized build unchanged Eligible -100.0% -40.0% 0/0/1
▼ Route a restore-bound cold build away from architecture changes Eligible -100.0% -40.0% 0/0/1

Illustrative judge evidence:

  • Configure deterministic, cache-safe CI builds: Position-swap inconsistent (forward: B, reverse: A). Defaulting to tie.

This is one example, not the aggregate verdict. Open Full Results for every judgment.

➖ Not proven improved — build-perf-diagnostics (claude-sonnet-5)

Why: Net win +42.9% (3W/4T/0L over 7 preference-eligible stimulus vote(s), sign test p=0.125), mean preference +20.0% across 8 paired run(s), 1 dormancy stimulus/stimuli excluded from preference — not credible — 4 of 7 preference-eligible stimulus vote(s) tied, leaving only 3 discordant preference vote(s). The sign test conditions on non-tie stimulus votes and cannot reach 0.05 below 5, so no record could have passed here — this is not a measured null. Either the skill is inert on these scenarios (make them discriminate) or the eval needs more distinct stimuli to clear the ties

Next action: Inspect tied or lost stimuli; predeclare added breadth before a new experiment.

State: VALID_NO_CHANGE (no_credible_preference_change)

Gate evidence: n=7; 3W/4T/0L; d=3; p=0.125; net +42.9%; 1 dormancy excluded

Overfit: Moderate (score 0.45)

Repeated-run reliability (not used by the gate): 8 paired runs (4W/4T/0L).

Weak or warning scenarios:

Scenario Preference gate Net win Δ Pref Runs (W/T/L)
= Diagnose NuGet restore running redundantly across CI stages Eligible +0.0% +0.0% 0/1/0
= Diagnose a Copy task dominating build time Eligible +0.0% +0.0% 0/1/0
= Diagnose a single custom target dominating one project's build Eligible +0.0% +0.0% 0/1/0
= Diagnose per-project overhead across many small projects Eligible +0.0% +0.0% 0/1/0

Illustrative judge evidence:

  • Diagnose NuGet restore running redundantly across CI stages: Position-swap inconsistent (forward: tie, reverse: B). Defaulting to tie.

This is one example, not the aggregate verdict. Open Full Results for every judgment.

➖ Not proven improved — build-perf-diagnostics (gpt-5.6-luna)

Why: Net win +14.3% (3W/2T/2L over 7 preference-eligible stimulus vote(s), sign test p=0.500), mean preference +5.0% across 8 paired run(s), 1 dormancy stimulus/stimuli excluded from preference — not credible (sign test p=0.500 > 0.05)

Next action: Inspect tied or lost stimuli and fix inconsistent skill behavior.

State: VALID_NO_CHANGE (no_credible_preference_change)

Gate evidence: n=7; 3W/2T/2L; d=5; p=0.500; net +14.3%; 1 dormancy excluded

Warnings: Activation: isolated 6/7; plugin 7/7

Overfit: Moderate (score 0.46)

Repeated-run reliability (not used by the gate): 8 paired runs (3W/3T/2L).

Weak or warning scenarios:

Scenario Preference gate Net win Δ Pref Runs (W/T/L)
= Decline a runtime latency request that is not a build performance issue Excluded (activation contract) +0.0% +0.0% 0/1/0
= Diagnose NuGet restore running redundantly across CI stages Eligible +0.0% +0.0% 0/1/0
= Diagnose a Copy task dominating build time Eligible +0.0% +0.0% 0/1/0
▼ Diagnose a pathological ResolveAssemblyReference time Eligible -100.0% -40.0% 0/0/1
▼ Diagnose per-project overhead across many small projects Eligible -100.0% -40.0% 0/0/1

Illustrative judge evidence:

  • Diagnose NuGet restore running redundantly across CI stages: Position-swap inconsistent (forward: tie, reverse: A). Defaulting to tie.

This is one example, not the aggregate verdict. Open Full Results for every judgment.

➖ Not proven improved — check-bin-obj-clash (claude-sonnet-5)

Why: Net win +14.3% (2W/4T/1L over 7 preference-eligible stimulus vote(s), sign test p=0.500), mean preference +5.7% across 7 paired run(s) — not credible — 4 of 7 preference-eligible stimulus vote(s) tied, leaving only 3 discordant preference vote(s). The sign test conditions on non-tie stimulus votes and cannot reach 0.05 below 5, so no record could have passed here — this is not a measured null. Either the skill is inert on these scenarios (make them discriminate) or the eval needs more distinct stimuli to clear the ties

Next action: Inspect tied or lost stimuli; predeclare added breadth before a new experiment.

State: VALID_NO_CHANGE (no_credible_preference_change)

Gate evidence: n=7; 2W/4T/1L; d=3; p=0.500; net +14.3%

Warnings: Activation: isolated 7/7; plugin 5/7

Overfit: Moderate (score 0.21)

Repeated-run reliability (not used by the gate): 7 paired runs (2W/4T/1L).

Weak or warning scenarios:

Scenario Preference gate Net win Δ Pref Runs (W/T/L)
= Avoid a false clash report when projects share only the top-level artifacts root Eligible +0.0% +0.0% 0/1/0
= Decline output-clash remediation for separate projects using the default SDK layout Eligible +0.0% +0.0% 0/1/0
= Diagnose multi-targeting outputs that collapse into one path Eligible +0.0% +0.0% 0/1/0
▼ Diagnose redundant project reference metadata that forks a same-path build Eligible -100.0% -40.0% 0/0/1
= Fix all clash mechanisms in the mixed solution Eligible +0.0% +0.0% 0/1/0

Illustrative judge evidence:

  • Avoid a false clash report when projects share only the top-level artifacts root: Position-swap inconsistent (forward: B, reverse: A). Defaulting to tie.

This is one example, not the aggregate verdict. Open Full Results for every judgment.

➖ Not proven improved — check-bin-obj-clash (gpt-5.6-luna)

Why: Net win +28.6% (2W/5T/0L over 7 preference-eligible stimulus vote(s), sign test p=0.250), mean preference +11.4% across 7 paired run(s) — not credible — 5 of 7 preference-eligible stimulus vote(s) tied, leaving only 2 discordant preference vote(s). The sign test conditions on non-tie stimulus votes and cannot reach 0.05 below 5, so no record could have passed here — this is not a measured null. Either the skill is inert on these scenarios (make them discriminate) or the eval needs more distinct stimuli to clear the ties

Next action: Inspect tied or lost stimuli; predeclare added breadth before a new experiment.

State: VALID_NO_CHANGE (no_credible_preference_change)

Gate evidence: n=7; 2W/5T/0L; d=2; p=0.250; net +28.6%

Overfit: Low (score 0.18)

Repeated-run reliability (not used by the gate): 7 paired runs (2W/5T/0L).

Weak or warning scenarios:

Scenario Preference gate Net win Δ Pref Runs (W/T/L)
= Audit the mixed solution and separate safe projects from unsafe ones Eligible +0.0% +0.0% 0/1/0
= Decline output-clash remediation for separate projects using the default SDK layout Eligible +0.0% +0.0% 0/1/0
= Diagnose multi-targeting outputs that collapse into one path Eligible +0.0% +0.0% 0/1/0
= Diagnose redundant project reference metadata that forks a same-path build Eligible +0.0% +0.0% 0/1/0
= Diagnose shared output and intermediate path collision Eligible +0.0% +0.0% 0/1/0

Illustrative judge evidence:

  • Audit the mixed solution and separate safe projects from unsafe ones: Position-swap inconsistent (forward: B, reverse: A). Defaulting to tie.

This is one example, not the aggregate verdict. Open Full Results for every judgment.

➖ Not proven improved — code-testing-agent (claude-sonnet-5)

Why: Net win +33.3% (5W/2T/2L over 9 preference-eligible stimulus vote(s), sign test p=0.227), mean preference +15.6% across 18 paired run(s) — not credible (sign test p=0.227 > 0.05)

Next action: Inspect tied or lost stimuli and fix inconsistent skill behavior.

State: VALID_NO_CHANGE (no_credible_preference_change)

Gate evidence: n=9; 5W/2T/2L; d=7; p=0.227; net +33.3%

Warnings: Activation: isolated 5/9; plugin 5/9

Overfit: Moderate (score 0.45)

Repeated-run reliability (not used by the gate): 18 paired runs (11W/3T/4L).

Weak or warning scenarios:

Scenario Preference gate Net win Δ Pref Runs (W/T/L)
= Add focused xUnit tests for one reservation class Eligible +0.0% +0.0% 0/2/0
▲ Expand a healthy existing pytest suite to every ledger boundary Eligible +100.0% +40.0% 2/0/0
= Generate a layered Vitest suite for an async shopping cart Eligible +0.0% +0.0% 1/0/1
▼ Generate a project-wide Go suite across collaborating packages Eligible -50.0% -20.0% 0/1/1
▼ Generate project-wide tests for an SDK-style xUnit library Eligible -100.0% -40.0% 0/0/2

Illustrative judge evidence:

  • Add focused xUnit tests for one reservation class: The two responses are substantively equivalent: both produced focused passing ReservationWindow boundary and validation tests, avoided forbidden changes, and gave concise completion summaries. Neither has a meaningful quality advantage.

This is one example, not the aggregate verdict. Open Full Results for every judgment.

➖ Not proven improved — detect-static-dependencies (gpt-5.6-luna)

Why: Net win +42.9% (4W/2T/1L over 7 preference-eligible stimulus vote(s), sign test p=0.188), mean preference +22.5% across 8 paired run(s), 1 dormancy stimulus/stimuli excluded from preference — not credible (sign test p=0.188 > 0.05)

Next action: Inspect tied or lost stimuli and fix inconsistent skill behavior.

State: VALID_NO_CHANGE (no_credible_preference_change)

Gate evidence: n=7; 4W/2T/1L; d=5; p=0.188; net +42.9%; 1 dormancy excluded

Warnings: 1 judge slot recovered

Retry recovery: 1 errored judgment slot recovered; successful first-attempt judgments stayed fixed.

Repeated-run reliability (not used by the gate): 8 paired runs (4W/3T/1L).

Weak or warning scenarios:

Scenario Preference gate Net win Δ Pref Runs (W/T/L)
= Detect time-related statics and recommend TimeProvider Eligible +0.0% +0.0% 0/1/0
= Exclude obj and bin directories from the scan Eligible +0.0% +0.0% 0/1/0
▼ Keep one authoritative total with file line locations and no seam for pure helpers Eligible -100.0% -40.0% 0/0/1
= Stay dormant for Python timezone review Excluded (activation contract) +0.0% +0.0% 0/1/0

Illustrative judge evidence:

  • Detect time-related statics and recommend TimeProvider: Position-swap inconsistent (forward: B, reverse: A). Defaulting to tie.

This is one example, not the aggregate verdict. Open Full Results for every judgment.

➖ Not proven improved — eval-performance (claude-sonnet-5)

Why: Net win +12.5% (3W/3T/2L over 8 preference-eligible stimulus vote(s), sign test p=0.500), mean preference +12.5% across 8 paired run(s) — not credible (sign test p=0.500 > 0.05)

Next action: Inspect tied or lost stimuli and fix inconsistent skill behavior.

State: VALID_NO_CHANGE (no_credible_preference_change)

Gate evidence: n=8; 3W/3T/2L; d=5; p=0.500; net +12.5%

Warnings: Activation: isolated 7/8; plugin 7/8

Overfit: Moderate (score 0.36)

Repeated-run reliability (not used by the gate): 8 paired runs (3W/3T/2L).

Weak or warning scenarios:

Scenario Preference gate Net win Δ Pref Runs (W/T/L)
= Decline to invent evaluation problems in an already-clean project Eligible +0.0% +0.0% 0/1/0
▼ Detect a project evaluated twice under different global properties Eligible -100.0% -40.0% 0/0/1
▼ Gather a measurement before proposing evaluation fixes for an unremarkable project Eligible -100.0% -40.0% 0/0/1
= Redirect a compile-time slowdown mistakenly framed as an evaluation problem Eligible +0.0% +0.0% 0/1/0
= Triage which of two property functions actually costs evaluation time Eligible +0.0% +0.0% 0/1/0

Illustrative judge evidence:

  • Decline to invent evaluation problems in an already-clean project: Position-swap inconsistent (forward: B, reverse: A). Defaulting to tie.

This is one example, not the aggregate verdict. Open Full Results for every judgment.

Routine passing details for 2 results are in Full Results.

Details for 44 results were omitted to keep this comment under GitHub's limit. Open the workflow summary or Full Results for the complete breakdown.

🔍 Full Results - all metrics and investigation details

To investigate non-passing or warning results, paste this to your AI coding agent:

For PR 1118 in dotnet/skills, download eval artifacts with gh run download 36034404215 --repo dotnet/skills --pattern "vally-results-*" --dir ./eval-results, then fetch https://raw.githubusercontent.com/dotnet/skills/be46d44caa481ba5dab779b84f8e16377f83e12b/eng/vally-adapter/InvestigatingResults.md and follow it. Classify each result as measurement-invalid, underpowered, not-proven, preference-loss, or passing-with-warning. Use stateReason, result accounting, weak scenarios, and judge evidence to give the cause and exact next fix.

⚠️ Session replay telemetry was not published because the auxiliary dotnet/skills-data publisher failed. The evaluation verdicts above remain authoritative; maintainers must repair SKILLS_DATA_TOKEN or the publisher before replay links are available.

@AbhitejJohn

Copy link
Copy Markdown
Collaborator Author

Superseded by the clean replacement set: #1214 for MSBuild eval tests and shared infrastructure, stacked #1215 for MSBuild SKILL.md guidance, and #1213 for dotnet-test fixes. The replacement diffs were independently reviewed and verified against the frozen source ledger.

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

waiting-on-review PR state label

Projects

None yet

Development

Successfully merging this pull request may close these issues.

4 participants