Skip to content

fix(msbuild): sharpen skill guidance - #1215

Merged
AbhitejJohn merged 12 commits into
mainfrom
abhitejjohn-msbuild-skill-guidance
Oct 5, 2026
Merged

AbhitejJohn merged 12 commits into
mainfrom
abhitejjohn-msbuild-skill-guidance

Conversation

@AbhitejJohn

@AbhitejJohn AbhitejJohn commented Sep 25, 2026 •

Copy link
Copy Markdown
Collaborator

Summary

This stacked change sharpens 14 MSBuild skill guides on top of #1214. It fixes command syntax, corrects build-parallelism guidance, and defines clearer routing boundaries among shared-file organization, property changes, imports and hooks, F# project ordering, generated files, item operations, and performance workflows.

Why

The PR1 scenarios exposed three guidance groups that needed correction: reliable PowerShell and file-logger syntax; accurate node and dependency advice; and routing descriptions that separate sibling skills by the requested outcome.

The final boundaries are explicit. directory-build-organization owns existing Directory.Build.* discovery, hierarchy/import timing, and props-versus-targets relocation, while property-patterns owns concrete value changes that stay within the current layout. extension-points owns NuGet auto-import and packed-layout discovery, while msbuild-antipatterns owns proven unsafe imports during audits and directly claims .fs, .fsi, and FS0039 compile-order tasks.

Impact

Agents receive commands that preserve literal braces and semicolon-delimited logger arguments, avoid the incorrect claim that plain dotnet build is sequential, and avoid deleting valid ProjectReference edges. Routing now keeps lone-project shared-file introduction dormant without excluding diagnosis of an existing one-project Directory.Build.* timing defect.

All descriptions remain below 1,024 characters. The dotnet-msbuild skill-menu description total is 10,943 characters, still 1,256 characters below the stacked base.

Validation

  • skill-validator check --plugin ./plugins/dotnet-msbuild: passed for 18 skills and 3 agents
  • python eng/eval-quality/check_eval_quality.py: passed with no errors
  • Vally lint: passed for all 14 affected skill/eval pairs
  • Static routing contract probes: passed for hierarchy-clobber, TargetFramework timing, lone-project dormancy, all three F# scenarios, and reciprocal NuGet import boundaries
  • Description limits: all plugin descriptions remain below 1,024 characters
  • Scope: cumulative PR diff contains exactly the 14 approved SKILL.md files; git diff --check passed

No evaluation run was triggered by this update.

@github-actions

github-actions Bot commented Sep 25, 2026 •

Copy link
Copy Markdown
Contributor

Skill Coverage Report

Plugin Skill Covered Coverage
✅ dotnet-msbuild check-bin-obj-clash 4/5 80%
❌ dotnet-msbuild directory-build-organization 0/1 0%
❌ dotnet-msbuild extension-points 0/1 0%
✅ dotnet-msbuild property-patterns 1/1 100%
Uncovered: dotnet-msbuild/check-bin-obj-clash
  • [WorkflowStep] Step 2: Get an overview and list projects (line 48)
Uncovered: dotnet-msbuild/directory-build-organization
  • [CodePattern] [MSBuild] (line 143)
Uncovered: dotnet-msbuild/extension-points
  • [CodePattern] [MSBuild] (line 195)

@AbhitejJohn
AbhitejJohn added this pull request to stack #1216 September 25, 2026 08:14
Copilot AI lite review requested due to automatic review settings September 25, 2026 08:30

Copilot AI left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Copilot review overview

🟡 Changes recommended

One or more issues must be addressed before approval.

Get a fresh assessment by requesting another Copilot review.

Review effort: Lite
Findings: 1 Medium severity

Open (1)
What changed in this PR

Sharpens 14 MSBuild skill guides with corrected command syntax, parallelism guidance, and clearer routing boundaries.

Changes:

  • Corrects PowerShell and semicolon-delimited logger syntax.
  • Clarifies dotnet build parallelism and preserves valid project dependencies.
  • Refines ownership boundaries among shared-file, property, import, F#, generated-file, and performance skills.
File Description
plugins/​dotnet-msbuild/​skills/​property-patterns/​SKILL.md Updated as part of this pull request.
plugins/​dotnet-msbuild/​skills/​msbuild-antipatterns/​SKILL.md Updated as part of this pull request.
plugins/​dotnet-msbuild/​skills/​item-management/​SKILL.md Updated as part of this pull request.
plugins/​dotnet-msbuild/​skills/​incremental-build/​SKILL.md Updated as part of this pull request.
plugins/​dotnet-msbuild/​skills/​including-generated-files/​SKILL.md Updated as part of this pull request.
plugins/​dotnet-msbuild/​skills/​extension-points/​SKILL.md Updated as part of this pull request.
plugins/​dotnet-msbuild/​skills/​eval-performance/​SKILL.md Updated as part of this pull request.
plugins/​dotnet-msbuild/​skills/​directory-build-organization/​SKILL.md Updated as part of this pull request.
plugins/​dotnet-msbuild/​skills/​check-bin-obj-clash/​SKILL.md Updated as part of this pull request.
plugins/​dotnet-msbuild/​skills/​build-perf-diagnostics/​SKILL.md Updated as part of this pull request.
plugins/​dotnet-msbuild/​skills/​build-perf-baseline/​SKILL.md Updated as part of this pull request.
plugins/​dotnet-msbuild/​skills/​build-parallelism/​SKILL.md Updated as part of this pull request.
plugins/​dotnet-msbuild/​skills/​binlog-generation/​SKILL.md Updated as part of this pull request.
plugins/​dotnet-msbuild/​skills/​binlog-failure-analysis/​SKILL.md Updated as part of this pull request.

💡 Configure MCP servers for context-aware, tailored reviews. Learn more in the docs.

Comment thread plugins/dotnet-msbuild/skills/build-perf-baseline/SKILL.md Outdated
@github-actions github-actions Bot added the waiting-on-author PR state label label Sep 25, 2026
@github-actions

Copy link
Copy Markdown
Contributor

👋 @AbhitejJohn — this PR has 1 unresolved review thread(s). When you're ready, please address the feedback and push an update; the triage bot will pick up the next state automatically. (Add the no-stale label to silence further pings.)

@github-actions

Copy link
Copy Markdown
Contributor

👋 @AbhitejJohn — this PR has 1 unresolved review thread(s). When you're ready, please address the feedback and push an update; the triage bot will pick up the next state automatically. (Add the no-stale label to silence further pings.)

Copilot AI lite review requested due to automatic review settings September 29, 2026 22:46

Copilot AI left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Copilot review overview

🔵 Needs a closer look

Two moderate routing conflicts can prevent valid item-management and import-safety reviews from activating.

Review effort: Lite
Findings: None

Resolved since last review (1)

@AbhitejJohn
AbhitejJohn marked this pull request as draft September 29, 2026 23:01
Base automatically changed from abhitejjohn-msbuild-eval-infrastructure to main September 30, 2026 17:31
AbhitejJohn and others added 4 commits September 30, 2026 10:31
Co-authored-by: Copilot App <223556219+Copilot@users.noreply.github.com>
Co-authored-by: Copilot App <223556219+Copilot@users.noreply.github.com>
Co-authored-by: Copilot App <223556219+Copilot@users.noreply.github.com>
Co-authored-by: Copilot App <223556219+Copilot@users.noreply.github.com>
@AbhitejJohn
AbhitejJohn force-pushed the abhitejjohn-msbuild-skill-guidance branch from edaaf71 to 2b5f5a4 Compare September 30, 2026 17:31
@AbhitejJohn
AbhitejJohn marked this pull request as ready for review September 30, 2026 17:49
Copilot AI lite review requested due to automatic review settings September 30, 2026 17:49
@AbhitejJohn

Copy link
Copy Markdown
Collaborator Author

/evaluate 2b5f5a4

Copilot AI left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Copilot review overview

🟡 Changes recommended

One or more issues must be addressed before approval.

Review effort: Lite
Findings: 1 Medium severity

Open (1)

Comment thread plugins/dotnet-msbuild/skills/property-patterns/SKILL.md Outdated
Co-authored-by: Copilot App <223556219+Copilot@users.noreply.github.com>
Copilot AI lite review requested due to automatic review settings September 30, 2026 18:08
Copilot AI lite review requested due to automatic review settings October 2, 2026 18:15

Copilot AI left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Copilot review overview

🟡 Changes recommended

One or more issues must be addressed before approval.

Review effort: Lite
Findings: 2 Medium severity

Open (2)

Comment thread plugins/dotnet-msbuild/skills/build-perf-baseline/SKILL.md Outdated
Comment thread plugins/dotnet-msbuild/skills/incremental-build/SKILL.md Outdated
Co-authored-by: Copilot App <223556219+Copilot@users.noreply.github.com>
Copilot AI lite review requested due to automatic review settings October 2, 2026 18:22

Copilot AI left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Copilot review overview

🔵 Needs a closer look

One or more issues must be addressed before approval.

Review effort: Lite
Findings: None

Resolved since last review (2)

@github-actions github-actions Bot added pr-state/ready-for-eval PR is mergeable and awaiting evaluation pr-state/evals-in-progress PR evaluations are in progress and removed waiting-on-author PR state label pr-state/ready-for-eval PR is mergeable and awaiting evaluation labels Oct 2, 2026
@github-actions

github-actions Bot commented Oct 2, 2026

Copy link
Copy Markdown
Contributor

📊 Skill and Agent Evaluation Results

30 model/target results across 15 targets and 2 models — ✅ 3 improved, ➖ 26 results without a clear winner, ⚠️ 0 invalid or underpowered, ⛔ 1 activation contract failure, 📉 0 preference losses (report only).

Measurement identity: evaluated commit 970e0dc1bd6bfcd78f4b1263252d049cda277641; 2 judge models.

Measurement health: 30 expected / 30 observed / 30 written; 0 missing, 0 unexpected, 0 invalid; 0 recovered comparison error slots and 0 unresolved comparison error slots.

Objective completion gate: not enabled. Aggregate completion transitions are telemetry only, so this report does not claim that zero objective regressions were proven.

A result passes only when preference-eligible distinct-stimulus votes have aggregate net win of at least 20% and an exact one-sided sign-test result of p ≤ 0.05, and every explicit dormancy activation contract passes. Repeated runs measure reliability only.

Target Model Verdict Gate evidence Overfit Warnings Next action
agent.msbuild claude-sonnet-5 ➖ Improvement signal, unproven n=5; 4W/0T/1L; d=5; p=0.188; net +60.0% — — The signal favors the target but is inconsistent; inspect tied or lost scenarios and fix weak behavior.
agent.msbuild gpt-5.6-luna ➖ Improvement signal, tie-limited n=5; 3W/1T/1L; d=4; p=0.312; net +40.0% — — The signal favors the target, but ties leave too few discordant tasks; inspect ties and predeclare more discriminating breadth.
binlog-failure-analysis claude-sonnet-5 ➖ Baseline signal, tie-limited n=6; 1W/3T/2L; d=3; p=0.500; net -16.7%; 1 dormancy excluded 🟡 0.37 Activation: isolated 6/7; plugin 5/7 Evidence leans baseline but is not credible; inspect losses and make tied tasks discriminate.
binlog-failure-analysis gpt-5.6-luna ➖ Improvement signal, tie-limited n=6; 4W/2T/0L; d=4; p=0.063; net +66.7%; 1 dormancy excluded 🟡 0.31 Activation: isolated 6/7; plugin 6/7 The signal favors the target, but ties leave too few discordant tasks; inspect ties and predeclare more discriminating breadth.
binlog-generation claude-sonnet-5 ✅ Improved n=7; 5W/2T/0L; d=5; p=0.031; net +71.4%; 1 dormancy excluded 🔴 0.51 Activation: isolated 5/7; plugin 5/7 Fix activation gaps; Review overfit evidence.
binlog-generation gpt-5.6-luna ➖ Improvement signal, unproven n=7; 4W/2T/1L; d=5; p=0.188; net +42.9%; 1 dormancy excluded 🟡 0.39 Activation: isolated 7/7; plugin 6/7 The signal favors the target but is inconsistent; inspect tied or lost scenarios and fix weak behavior.
build-parallelism claude-sonnet-5 ➖ Improvement signal, tie-limited n=6; 4W/2T/0L; d=4; p=0.063; net +66.7%; 1 dormancy excluded 🟡 0.41 Activation: isolated 5/6; plugin 5/6; Activation-only stop: plugin 1 failed run Inspect activation-only failed runs before rewriting skill content; the model stopped after loading a skill.
build-parallelism gpt-5.6-luna ➖ Improvement signal, unproven n=6; 3W/1T/2L; d=5; p=0.500; net +16.7%; 1 dormancy excluded 🟡 0.26 Activation-only stop: isolated 1 failed run; Activation-only stop: plugin 1 failed run Inspect activation-only failed runs before rewriting skill content; the model stopped after loading a skill.
build-perf-baseline claude-sonnet-5 ✅ Improved n=6; 5W/1T/0L; d=5; p=0.031; net +83.3%; 1 dormancy excluded 🟡 0.45 Activation: isolated 4/6; plugin 4/6 Fix activation gaps; Review overfit evidence.
build-perf-baseline gpt-5.6-luna ⛔ Activation contract failed n=6; 3W/2T/1L; d=4; p=0.312; net +33.3%; 1 dormancy excluded 🟡 0.33 Dormancy contract: 1 unexpected activation(s); Activation: isolated 5/6; plugin 3/6 Narrow skill routing so the listed off-target scenarios stay dormant.
build-perf-diagnostics claude-sonnet-5 ➖ Improvement signal, unproven n=7; 4W/1T/2L; d=6; p=0.344; net +28.6%; 1 dormancy excluded 🟡 0.43 Activation: isolated 7/7; plugin 2/7 The signal favors the target but is inconsistent; inspect tied or lost scenarios and fix weak behavior.
build-perf-diagnostics gpt-5.6-luna ➖ Improvement signal, unproven n=7; 5W/0T/2L; d=7; p=0.227; net +42.9%; 1 dormancy excluded 🟡 0.31 Activation: isolated 7/7; plugin 3/7 The signal favors the target but is inconsistent; inspect tied or lost scenarios and fix weak behavior.
check-bin-obj-clash claude-sonnet-5 ➖ Mixed evidence n=7; 1W/5T/1L; d=2; p=0.750; net +0.0% 🟡 0.23 Activation: isolated 7/7; plugin 6/7 Compare winning and losing scenarios to isolate where the target helps versus hurts.
check-bin-obj-clash gpt-5.6-luna ➖ Improvement signal, tie-limited n=7; 3W/4T/0L; d=3; p=0.125; net +42.9% 🟡 0.20 — The signal favors the target, but ties leave too few discordant tasks; inspect ties and predeclare more discriminating breadth.
directory-build-organization claude-sonnet-5 ➖ Baseline signal, unproven n=6; 2W/1T/3L; d=5; p=0.500; net -16.7%; 1 dormancy excluded 🟡 0.37 Activation: isolated 4/6; plugin 4/6 Evidence leans baseline but is not credible; inspect losing scenarios for recurring defects.
directory-build-organization gpt-5.6-luna ➖ Improvement signal, tie-limited n=6; 3W/2T/1L; d=4; p=0.312; net +33.3%; 1 dormancy excluded 🟡 0.35 Activation: isolated 6/6; plugin 4/6 The signal favors the target, but ties leave too few discordant tasks; inspect ties and predeclare more discriminating breadth.
eval-performance claude-sonnet-5 ➖ Mixed evidence n=8; 3W/2T/3L; d=6; p=0.656; net +0.0% 🟡 0.39 Activation: isolated 6/8; plugin 4/8 Compare winning and losing scenarios to isolate where the target helps versus hurts.
eval-performance gpt-5.6-luna ➖ Improvement signal, tie-limited n=8; 4W/4T/0L; d=4; p=0.063; net +50.0% 🟡 0.40 — The signal favors the target, but ties leave too few discordant tasks; inspect ties and predeclare more discriminating breadth.
extension-points claude-sonnet-5 ➖ Improvement signal, tie-limited n=7; 3W/3T/1L; d=4; p=0.312; net +28.6%; 2 dormancy excluded ✅ 0.19 Activation: isolated 4/7; plugin 6/7 The signal favors the target, but ties leave too few discordant tasks; inspect ties and predeclare more discriminating breadth.
extension-points gpt-5.6-luna ➖ Improvement signal, unproven n=7; 4W/2T/1L; d=5; p=0.188; net +42.9%; 2 dormancy excluded 🟡 0.31 Activation: isolated 7/7; plugin 6/7 The signal favors the target but is inconsistent; inspect tied or lost scenarios and fix weak behavior.
including-generated-files claude-sonnet-5 ➖ Mixed evidence n=7; 3W/1T/3L; d=6; p=0.656; net +0.0%; 1 dormancy excluded 🟡 0.38 Activation: isolated 4/7; plugin 6/7 Compare winning and losing scenarios to isolate where the target helps versus hurts.
including-generated-files gpt-5.6-luna ➖ Baseline signal, tie-limited n=7; 1W/3T/3L; d=4; p=0.312; net -28.6%; 1 dormancy excluded 🟡 0.36 — Evidence leans baseline but is not credible; inspect losses and make tied tasks discriminate.
incremental-build claude-sonnet-5 ➖ Improvement signal, unproven n=9; 4W/2T/3L; d=7; p=0.500; net +11.1% 🟡 0.39 Activation: isolated 4/9; plugin 4/9 The signal favors the target but is inconsistent; inspect tied or lost scenarios and fix weak behavior.
incremental-build gpt-5.6-luna ✅ Improved n=9; 7W/2T/0L; d=7; p=0.008; net +77.8% 🟡 0.36 Activation: isolated 8/9; plugin 6/9 Fix activation gaps; Review overfit evidence.
item-management claude-sonnet-5 ➖ Improvement signal, tie-limited n=6; 2W/3T/1L; d=3; p=0.500; net +16.7%; 1 dormancy excluded 🟡 0.37 Activation: isolated 2/6; plugin 3/6 The signal favors the target, but ties leave too few discordant tasks; inspect ties and predeclare more discriminating breadth.
item-management gpt-5.6-luna ➖ Mixed evidence n=6; 2W/2T/2L; d=4; p=0.687; net +0.0%; 1 dormancy excluded 🟡 0.49 Activation: isolated 6/6; plugin 5/6 Compare winning and losing scenarios to isolate where the target helps versus hurts.
msbuild-antipatterns claude-sonnet-5 ➖ Improvement signal, tie-limited n=7; 2W/4T/1L; d=3; p=0.500; net +14.3%; 1 dormancy excluded 🟡 0.30 Activation: isolated 3/7; plugin 3/7 The signal favors the target, but ties leave too few discordant tasks; inspect ties and predeclare more discriminating breadth.
msbuild-antipatterns gpt-5.6-luna ➖ Baseline signal, tie-limited n=7; 1W/3T/3L; d=4; p=0.312; net -28.6%; 1 dormancy excluded 🟡 0.24 Activation: isolated 5/7; plugin 4/7 Evidence leans baseline but is not credible; inspect losses and make tied tasks discriminate.
property-patterns claude-sonnet-5 ➖ Improvement signal, unproven n=7; 3W/2T/2L; d=5; p=0.500; net +14.3% 🟡 0.24 Activation: isolated 6/7; plugin 6/7 The signal favors the target but is inconsistent; inspect tied or lost scenarios and fix weak behavior.
property-patterns gpt-5.6-luna ➖ Improvement signal, tie-limited n=7; 3W/4T/0L; d=3; p=0.125; net +42.9% 🟡 0.44 Activation: isolated 6/7; plugin 5/7 The signal favors the target, but ties leave too few discordant tasks; inspect ties and predeclare more discriminating breadth.
ℹ️ How to read this report
  • ✅ Improved — the result passed both the statistical gate and the 20% practical net-win floor.
  • ➖ No clear winner — the result is valid but did not pass both gates. The label distinguishes all ties, mixed evidence, directional but unproven evidence, and credible effects below the practical floor.
  • ⚠️ Invalid / underpowered — the gate withheld a quality verdict. Fix the measurement before judging the target.
  • ⛔ Activation contract failed — the isolated target activated on an explicit dormancy scenario. Dormancy preference is excluded, but this routing failure still blocks a pass.
  • 📉 Preference loss — the LLM judge credibly preferred baseline. It is report-only, not objective completion proof.
  • Gate evidence — n preference-eligible distinct-stimulus votes, W/T/L stimulus votes, d discordant votes, exact one-sided p, net win, and the count of separately retained dormancy stimuli. The p value applies to one model/target result; no matrix-wide multiple-comparison correction is applied.
  • Overfit — overfitting-judge severity (✅ Low, 🟡 Moderate, 🔴 High, — none) and score.
  • Warnings — activation, timeout, retry recovery, or unresolved comparison conditions that need attention.
  • Do not add repeated runs to increase statistical power. Do not add stimuli after seeing a near-pass unless the new breadth is predeclared for a new experiment.
⛔ Activation contract failed — build-perf-baseline (gpt-5.6-luna)

Why: Net win +33.3% (3W/2T/1L over 6 preference-eligible stimulus vote(s), sign test p=0.312), mean preference +5.7% across 7 paired run(s), 1 dormancy stimulus/stimuli excluded from preference — activation contract failed (1 explicit dormancy scenario(s) activated the isolated target skill)

Next action: Narrow skill routing so the listed off-target scenarios stay dormant.

State: VALID_NO_CHANGE (activation_contract_failed)

Gate evidence: n=6; 3W/2T/1L; d=4; p=0.312; net +33.3%; 1 dormancy excluded

Warnings: Dormancy contract: 1 unexpected activation(s); Activation: isolated 5/6; plugin 3/6

Overfit: Moderate (score 0.33)

Repeated-run reliability (not used by the gate): 7 paired runs (3W/2T/2L).

Weak or warning scenarios:

Scenario Preference gate Net win Δ Pref Runs (W/T/L)
▲ Adopt UseArtifactsOutput while declining graph build for a small solution Eligible +100.0% +40.0% 1/0/0
= Configure deterministic, cache-safe CI builds Eligible +0.0% +0.0% 0/1/0
▼ Decline a non-MSBuild build performance request Excluded (activation contract) -100.0% -40.0% 0/0/1
▼ Route a broken no-op rebuild away from generic optimization Eligible -100.0% -40.0% 0/0/1
= Route a restore-bound cold build away from architecture changes Eligible +0.0% +0.0% 0/1/0

Illustrative judge evidence:

  • Configure deterministic, cache-safe CI builds: Position-swap inconsistent (forward: B, reverse: A). Defaulting to tie.

This is one example, not the aggregate verdict. Open Full Results for every judgment.

➖ Improvement signal, unproven — agent.msbuild (claude-sonnet-5)

Why: Net win +60.0% (4W/0T/1L over 5 preference-eligible stimulus vote(s), sign test p=0.188), mean preference +36.0% across 5 paired run(s) — not credible (sign test p=0.188 > 0.05)

Next action: The signal favors the target but is inconsistent; inspect tied or lost scenarios and fix weak behavior.

State: VALID_NO_CHANGE (no_credible_preference_change)

Gate evidence: n=5; 4W/0T/1L; d=5; p=0.188; net +60.0%

Repeated-run reliability (not used by the gate): 5 paired runs (4W/0T/1L).

Weak or warning scenarios:

Scenario Preference gate Net win Δ Pref Runs (W/T/L)
▼ Route a slow build to performance analysis Eligible -100.0% -40.0% 0/0/1

Illustrative judge evidence:

  • Route a slow build to performance analysis: A is more comprehensive and directly actionable for identifying expensive MSBuild targets, tasks, analyzers, restore costs, incrementality failures, and critical-path constraints. B has a stronger statistical baseline prescription, but its unexplained skill-routing language an...

This is one example, not the aggregate verdict. Open Full Results for every judgment.

➖ Improvement signal, tie-limited — agent.msbuild (gpt-5.6-luna)

Why: Net win +40.0% (3W/1T/1L over 5 preference-eligible stimulus vote(s), sign test p=0.312), mean preference +16.0% across 5 paired run(s) — not credible — 1 of 5 preference-eligible stimulus vote(s) tied, leaving only 4 discordant preference vote(s). The sign test conditions on non-tie stimulus votes and cannot reach 0.05 below 5, so no record could have passed here — this is not a measured null. Either the agent is inert on these scenarios (make them discriminate) or the eval needs more distinct stimuli to clear the ties

Next action: The signal favors the target, but ties leave too few discordant tasks; inspect ties and predeclare more discriminating breadth.

State: VALID_NO_CHANGE (no_credible_preference_change)

Gate evidence: n=5; 3W/1T/1L; d=4; p=0.312; net +40.0%

Repeated-run reliability (not used by the gate): 5 paired runs (3W/1T/1L).

Weak or warning scenarios:

Scenario Preference gate Net win Δ Pref Runs (W/T/L)
▼ Route a slow build to performance analysis Eligible -100.0% -40.0% 0/0/1
= Triage a build failure and route to appropriate analysis Eligible +0.0% +0.0% 0/1/0

Illustrative judge evidence:

  • Route a slow build to performance analysis: Response A provides a more systematically organized, measurement-first framework with greater depth on evaluation profiling, suspicious patterns, and validation methodology. It explicitly frames the entire task as a measurement problem before tuning. Response B has strengths i...

This is one example, not the aggregate verdict. Open Full Results for every judgment.

➖ Baseline signal, tie-limited — binlog-failure-analysis (claude-sonnet-5)

Why: Net win -16.7% (1W/3T/2L over 6 preference-eligible stimulus vote(s), sign test p=0.500), mean preference -20.0% across 7 paired run(s), 1 dormancy stimulus/stimuli excluded from preference — no improvement

Next action: Evidence leans baseline but is not credible; inspect losses and make tied tasks discriminate.

State: VALID_NO_CHANGE (no_credible_preference_change)

Gate evidence: n=6; 1W/3T/2L; d=3; p=0.500; net -16.7%; 1 dormancy excluded

Warnings: Activation: isolated 6/7; plugin 5/7

Overfit: Moderate (score 0.37)

Repeated-run reliability (not used by the gate): 7 paired runs (1W/3T/3L).

Weak or warning scenarios:

Scenario Preference gate Net win Δ Pref Runs (W/T/L)
▼ Assess a requested binlog investigation when the current build is healthy Eligible -100.0% -100.0% 0/0/1
= Confirm the actual resolved target framework and package version from a binlog Eligible +0.0% +0.0% 0/0/0
= Determine whether a quiet second build actually failed Eligible +0.0% +0.0% 0/1/0
= Diagnose a warning behind a build that actually succeeded Eligible +0.0% +0.0% 0/1/0
= Diagnose build failures from binlog only (no source files) Eligible +0.0% +0.0% 0/1/0
▼ Stay dormant for a non-MSBuild build failure log Excluded (activation contract) -100.0% -40.0% 0/0/1
▼ Trace why a generated source file is missing at compile time Eligible -100.0% -100.0% 0/0/1
▲ Use capture-time text logs when the binlog reader is unavailable Eligible +100.0% +100.0% 1/0/0

Illustrative judge evidence:

  • Assess a requested binlog investigation when the current build is healthy: A is aligned with the observed evidence and the task's key premise: the current project builds successfully, so no failure root cause exists. It captures and reports the binlog path without inventing a defect. B's additional log inspection is not inherently bad, but its confid...

This is one example, not the aggregate verdict. Open Full Results for every judgment.

➖ Improvement signal, tie-limited — binlog-failure-analysis (gpt-5.6-luna)

Why: Net win +66.7% (4W/2T/0L over 6 preference-eligible stimulus vote(s), sign test p=0.063), mean preference +17.1% across 7 paired run(s), 1 dormancy stimulus/stimuli excluded from preference — not credible — 2 of 6 preference-eligible stimulus vote(s) tied, leaving only 4 discordant preference vote(s). The sign test conditions on non-tie stimulus votes and cannot reach 0.05 below 5, so no record could have passed here — this is not a measured null. Either the skill is inert on these scenarios (make them discriminate) or the eval needs more distinct stimuli to clear the ties

Next action: The signal favors the target, but ties leave too few discordant tasks; inspect ties and predeclare more discriminating breadth.

State: VALID_NO_CHANGE (no_credible_preference_change)

Gate evidence: n=6; 4W/2T/0L; d=4; p=0.063; net +66.7%; 1 dormancy excluded

Warnings: Activation: isolated 6/7; plugin 6/7

Overfit: Moderate (score 0.31)

Repeated-run reliability (not used by the gate): 7 paired runs (4W/2T/1L).

Weak or warning scenarios:

Scenario Preference gate Net win Δ Pref Runs (W/T/L)
= Assess a requested binlog investigation when the current build is healthy Eligible +0.0% +0.0% 0/1/0
= Confirm the actual resolved target framework and package version from a binlog Eligible +0.0% +0.0% 0/0/0
▼ Stay dormant for a non-MSBuild build failure log Excluded (activation contract) -100.0% -40.0% 0/0/1
= Trace why a generated source file is missing at compile time Eligible +0.0% +0.0% 0/1/0

Illustrative judge evidence:

  • Assess a requested binlog investigation when the current build is healthy: Position-swap inconsistent (forward: A, reverse: B). Defaulting to tie.

This is one example, not the aggregate verdict. Open Full Results for every judgment.

➖ Improvement signal, unproven — binlog-generation (gpt-5.6-luna)

Why: Net win +42.9% (4W/2T/1L over 7 preference-eligible stimulus vote(s), sign test p=0.188), mean preference +37.5% across 8 paired run(s), 1 dormancy stimulus/stimuli excluded from preference — not credible (sign test p=0.188 > 0.05)

Next action: The signal favors the target but is inconsistent; inspect tied or lost scenarios and fix weak behavior.

State: VALID_NO_CHANGE (no_credible_preference_change)

Gate evidence: n=7; 4W/2T/1L; d=5; p=0.188; net +42.9%; 1 dormancy excluded

Warnings: Activation: isolated 7/7; plugin 6/7

Overfit: Moderate (score 0.39)

Repeated-run reliability (not used by the gate): 8 paired runs (4W/3T/1L).

Weak or warning scenarios:

Scenario Preference gate Net win Δ Pref Runs (W/T/L)
= Build multiple configurations with unique binlogs Eligible +0.0% +0.0% 0/1/0
= Choose a predictable, non-colliding binlog name for a CI upload step Eligible +0.0% +0.0% 0/1/0
= Decline Maven compiler configuration Excluded (activation contract) +0.0% +0.0% 0/1/0
▲ Fix a CI build script that reuses the same binlog on every run Eligible +100.0% +100.0% 1/0/0
▼ Recognize that a failed build produced no binlog at all Eligible -100.0% -40.0% 0/0/1

Illustrative judge evidence:

  • Build multiple configurations with unique binlogs: Position-swap inconsistent (forward: tie, reverse: B). Defaulting to tie.

This is one example, not the aggregate verdict. Open Full Results for every judgment.

➖ Improvement signal, tie-limited — build-parallelism (claude-sonnet-5)

Why: Net win +66.7% (4W/2T/0L over 6 preference-eligible stimulus vote(s), sign test p=0.063), mean preference +25.7% across 7 paired run(s), 1 dormancy stimulus/stimuli excluded from preference — not credible — 2 of 6 preference-eligible stimulus vote(s) tied, leaving only 4 discordant preference vote(s). The sign test conditions on non-tie stimulus votes and cannot reach 0.05 below 5, so no record could have passed here — this is not a measured null. Either the skill is inert on these scenarios (make them discriminate) or the eval needs more distinct stimuli to clear the ties

Next action: Inspect activation-only failed runs before rewriting skill content; the model stopped after loading a skill.

State: VALID_NO_CHANGE (no_credible_preference_change)

Gate evidence: n=6; 4W/2T/0L; d=4; p=0.063; net +66.7%; 1 dormancy excluded

Warnings: Activation: isolated 5/6; plugin 5/6; Activation-only stop: plugin 1 failed run

Overfit: Moderate (score 0.41)

Repeated-run reliability (not used by the gate): 7 paired runs (4W/2T/1L).

Weak or warning scenarios:

Scenario Preference gate Net win Δ Pref Runs (W/T/L)
▼ Decline database query tuning request Excluded (activation contract) -100.0% -40.0% 0/0/1
= Enable BuildInParallel on a custom MSBuild task Eligible +0.0% +0.0% 0/1/0
= Enable parallel nodes for a wide project graph Eligible +0.0% +0.0% 0/1/0
▲ Reduce CI build scope with a solution filter Eligible +100.0% +40.0% 1/0/0

Illustrative judge evidence:

  • Enable BuildInParallel on a custom MSBuild task: The resulting edits and final explanations are substantively equivalent and fully satisfy the task.

This is one example, not the aggregate verdict. Open Full Results for every judgment.

➖ Improvement signal, unproven — build-parallelism (gpt-5.6-luna)

Why: Net win +16.7% (3W/1T/2L over 6 preference-eligible stimulus vote(s), sign test p=0.500), mean preference +11.4% across 7 paired run(s), 1 dormancy stimulus/stimuli excluded from preference — not credible (sign test p=0.500 > 0.05)

Next action: Inspect activation-only failed runs before rewriting skill content; the model stopped after loading a skill.

State: VALID_NO_CHANGE (no_credible_preference_change)

Gate evidence: n=6; 3W/1T/2L; d=5; p=0.500; net +16.7%; 1 dormancy excluded

Warnings: Activation-only stop: isolated 1 failed run; Activation-only stop: plugin 1 failed run

Overfit: Moderate (score 0.26)

Repeated-run reliability (not used by the gate): 7 paired runs (4W/1T/2L).

Weak or warning scenarios:

Scenario Preference gate Net win Δ Pref Runs (W/T/L)
▼ Decline graph build for runtime-discovered projects Eligible -100.0% -40.0% 0/0/1
= Enable parallel nodes for a wide project graph Eligible +0.0% +0.0% 0/1/0
▼ Reduce CI build scope with a solution filter Eligible -100.0% -40.0% 0/0/1

Illustrative judge evidence:

  • Decline graph build for runtime-discovered projects: Both responses are technically sound and cover the core issues well. Response A edges ahead with a clearer three-option structure that directly addresses the user's scenario ('if discovery must remain, don't use /graph'). Response B offers equivalent technical accuracy and add...

This is one example, not the aggregate verdict. Open Full Results for every judgment.

➖ Improvement signal, unproven — build-perf-diagnostics (claude-sonnet-5)

Why: Net win +28.6% (4W/1T/2L over 7 preference-eligible stimulus vote(s), sign test p=0.344), mean preference +15.0% across 8 paired run(s), 1 dormancy stimulus/stimuli excluded from preference — not credible (sign test p=0.344 > 0.05)

Next action: The signal favors the target but is inconsistent; inspect tied or lost scenarios and fix weak behavior.

State: VALID_NO_CHANGE (no_credible_preference_change)

Gate evidence: n=7; 4W/1T/2L; d=6; p=0.344; net +28.6%; 1 dormancy excluded

Warnings: Activation: isolated 7/7; plugin 2/7

Overfit: Moderate (score 0.43)

Repeated-run reliability (not used by the gate): 8 paired runs (5W/1T/2L).

Weak or warning scenarios:

Scenario Preference gate Net win Δ Pref Runs (W/T/L)
= Diagnose NuGet restore running redundantly across CI stages Eligible +0.0% +0.0% 0/1/0
▼ Diagnose a Copy task dominating build time Eligible -100.0% -40.0% 0/0/1
▲ Diagnose a single custom target dominating one project's build Eligible +100.0% +40.0% 1/0/0
▲ Diagnose evaluation overhead before any target runs Eligible +100.0% +40.0% 1/0/0
▼ Diagnose per-project overhead across many small projects Eligible -100.0% -40.0% 0/0/1
▲ Diagnose slow build for a small project Eligible +100.0% +40.0% 1/0/0

Illustrative judge evidence:

  • Diagnose NuGet restore running redundantly across CI stages: Position-swap inconsistent (forward: A, reverse: B). Defaulting to tie.

This is one example, not the aggregate verdict. Open Full Results for every judgment.

➖ Improvement signal, unproven — build-perf-diagnostics (gpt-5.6-luna)

Why: Net win +42.9% (5W/0T/2L over 7 preference-eligible stimulus vote(s), sign test p=0.227), mean preference +10.0% across 8 paired run(s), 1 dormancy stimulus/stimuli excluded from preference — not credible (sign test p=0.227 > 0.05)

Next action: The signal favors the target but is inconsistent; inspect tied or lost scenarios and fix weak behavior.

State: VALID_NO_CHANGE (no_credible_preference_change)

Gate evidence: n=7; 5W/0T/2L; d=7; p=0.227; net +42.9%; 1 dormancy excluded

Warnings: Activation: isolated 7/7; plugin 3/7

Overfit: Moderate (score 0.31)

Repeated-run reliability (not used by the gate): 8 paired runs (5W/0T/3L).

Weak or warning scenarios:

Scenario Preference gate Net win Δ Pref Runs (W/T/L)
▼ Decline a runtime latency request that is not a build performance issue Excluded (activation contract) -100.0% -40.0% 0/0/1
▼ Diagnose a Copy task dominating build time Eligible -100.0% -40.0% 0/0/1
▲ Diagnose a pathological ResolveAssemblyReference time Eligible +100.0% +40.0% 1/0/0
▲ Diagnose a single custom target dominating one project's build Eligible +100.0% +40.0% 1/0/0
▲ Diagnose evaluation overhead before any target runs Eligible +100.0% +40.0% 1/0/0
▼ Diagnose per-project overhead across many small projects Eligible -100.0% -40.0% 0/0/1

Illustrative judge evidence:

  • Diagnose a Copy task dominating build time: Response A delivers a more technically sound diagnosis and avoids recommending an inapplicable MSBuild setting. Its explicit rejection of CreateHardLinksForAdditionalFilesIfPossible (with correct reasoning about scope) demonstrates better MSBuild semantics understanding than R...

This is one example, not the aggregate verdict. Open Full Results for every judgment.

➖ Mixed evidence — check-bin-obj-clash (claude-sonnet-5)

Why: Net win +0.0% (1W/5T/1L over 7 preference-eligible stimulus vote(s), sign test p=0.750), mean preference +0.0% across 7 paired run(s) — no improvement

Next action: Compare winning and losing scenarios to isolate where the target helps versus hurts.

State: VALID_NO_CHANGE (no_credible_preference_change)

Gate evidence: n=7; 1W/5T/1L; d=2; p=0.750; net +0.0%

Warnings: Activation: isolated 7/7; plugin 6/7

Overfit: Moderate (score 0.23)

Repeated-run reliability (not used by the gate): 7 paired runs (1W/5T/1L).

Weak or warning scenarios:

Scenario Preference gate Net win Δ Pref Runs (W/T/L)
▼ Audit the mixed solution and separate safe projects from unsafe ones Eligible -100.0% -100.0% 0/0/1
= Avoid a false clash report when projects share only the top-level artifacts root Eligible +0.0% +0.0% 0/1/0
= Decline output-clash remediation for separate projects using the default SDK layout Eligible +0.0% +0.0% 0/1/0
= Diagnose multi-targeting outputs that collapse into one path Eligible +0.0% +0.0% 0/1/0
= Diagnose redundant project reference metadata that forks a same-path build Eligible +0.0% +0.0% 0/1/0
= Diagnose shared output and intermediate path collision Eligible +0.0% +0.0% 0/1/0

Illustrative judge evidence:

  • Audit the mixed solution and separate safe projects from unsafe ones: A delivers a complete, verified audit covering every actual hazard and the safe projects. B gets the shared-directory and redundant-reference issues right but fails the central MultiTargetLib determination by labeling it safe pending a check it did not perform.

This is one example, not the aggregate verdict. Open Full Results for every judgment.

➖ Improvement signal, tie-limited — check-bin-obj-clash (gpt-5.6-luna)

Why: Net win +42.9% (3W/4T/0L over 7 preference-eligible stimulus vote(s), sign test p=0.125), mean preference +17.1% across 7 paired run(s) — not credible — 4 of 7 preference-eligible stimulus vote(s) tied, leaving only 3 discordant preference vote(s). The sign test conditions on non-tie stimulus votes and cannot reach 0.05 below 5, so no record could have passed here — this is not a measured null. Either the skill is inert on these scenarios (make them discriminate) or the eval needs more distinct stimuli to clear the ties

Next action: The signal favors the target, but ties leave too few discordant tasks; inspect ties and predeclare more discriminating breadth.

State: VALID_NO_CHANGE (no_credible_preference_change)

Gate evidence: n=7; 3W/4T/0L; d=3; p=0.125; net +42.9%

Overfit: Moderate (score 0.20)

Repeated-run reliability (not used by the gate): 7 paired runs (3W/4T/0L).

Weak or warning scenarios:

Scenario Preference gate Net win Δ Pref Runs (W/T/L)
= Audit the mixed solution and separate safe projects from unsafe ones Eligible +0.0% +0.0% 0/1/0
= Avoid a false clash report when projects share only the top-level artifacts root Eligible +0.0% +0.0% 0/1/0
= Diagnose multi-targeting outputs that collapse into one path Eligible +0.0% +0.0% 0/1/0
= Diagnose redundant project reference metadata that forks a same-path build Eligible +0.0% +0.0% 0/1/0

Illustrative judge evidence:

  • Audit the mixed solution and separate safe projects from unsafe ones: Position-swap inconsistent (forward: B, reverse: A). Defaulting to tie.

This is one example, not the aggregate verdict. Open Full Results for every judgment.

➖ Baseline signal, unproven — directory-build-organization (claude-sonnet-5)

Why: Net win -16.7% (2W/1T/3L over 6 preference-eligible stimulus vote(s), sign test p=0.500), mean preference -5.7% across 7 paired run(s), 1 dormancy stimulus/stimuli excluded from preference — no improvement

Next action: Evidence leans baseline but is not credible; inspect losing scenarios for recurring defects.

State: VALID_NO_CHANGE (no_credible_preference_change)

Gate evidence: n=6; 2W/1T/3L; d=5; p=0.500; net -16.7%; 1 dormancy excluded

Warnings: Activation: isolated 4/6; plugin 4/6

Overfit: Moderate (score 0.37)

Repeated-run reliability (not used by the gate): 7 paired runs (2W/2T/3L).

Weak or warning scenarios:

Scenario Preference gate Net win Δ Pref Runs (W/T/L)
= Apply repo-level build organization cleanup Eligible +0.0% +0.0% 0/1/0
= Decline adding shared build files to a lone project Excluded (activation contract) +0.0% +0.0% 0/1/0
▲ Diagnose a TargetFramework condition that silently skips in props Eligible +100.0% +40.0% 1/0/0
▼ Diagnose a package downgrade chain and reorganize version management Eligible -100.0% -40.0% 0/0/1
▼ Diagnose an inner shared-props file that overwrites its own override Eligible -100.0% -40.0% 0/0/1
▼ Preserve an intentional project-specific exception while centralizing Eligible -100.0% -40.0% 0/0/1

Illustrative judge evidence:

  • Apply repo-level build organization cleanup: Position-swap inconsistent (forward: B, reverse: A). Defaulting to tie.

This is one example, not the aggregate verdict. Open Full Results for every judgment.

➖ Improvement signal, tie-limited — directory-build-organization (gpt-5.6-luna)

Why: Net win +33.3% (3W/2T/1L over 6 preference-eligible stimulus vote(s), sign test p=0.312), mean preference +11.4% across 7 paired run(s), 1 dormancy stimulus/stimuli excluded from preference — not credible — 2 of 6 preference-eligible stimulus vote(s) tied, leaving only 4 discordant preference vote(s). The sign test conditions on non-tie stimulus votes and cannot reach 0.05 below 5, so no record could have passed here — this is not a measured null. Either the skill is inert on these scenarios (make them discriminate) or the eval needs more distinct stimuli to clear the ties

Next action: The signal favors the target, but ties leave too few discordant tasks; inspect ties and predeclare more discriminating breadth.

State: VALID_NO_CHANGE (no_credible_preference_change)

Gate evidence: n=6; 3W/2T/1L; d=4; p=0.312; net +33.3%; 1 dormancy excluded

Warnings: Activation: isolated 6/6; plugin 4/6

Overfit: Moderate (score 0.35)

Repeated-run reliability (not used by the gate): 7 paired runs (3W/3T/1L).

Weak or warning scenarios:

Scenario Preference gate Net win Δ Pref Runs (W/T/L)
▼ Apply repo-level build organization cleanup Eligible -100.0% -100.0% 0/0/1
= Decline adding shared build files to a lone project Excluded (activation contract) +0.0% +0.0% 0/1/0
▲ Diagnose a TargetFramework condition that silently skips in props Eligible +100.0% +40.0% 1/0/0
= Diagnose an inner shared-props file that overwrites its own override Eligible +0.0% +0.0% 0/1/0
= Preserve an intentional project-specific exception while centralizing Eligible +0.0% +0.0% 0/1/0

Illustrative judge evidence:

  • Apply repo-level build organization cleanup: Response A delivered a complete, working solution: all shared build files were created, settings properly centralized, test configuration isolated, and the final build and tests pass successfully. Response B, despite using a specialized skill and attempting similar changes, le...

This is one example, not the aggregate verdict. Open Full Results for every judgment.

➖ Mixed evidence — eval-performance (claude-sonnet-5)

Why: Net win +0.0% (3W/2T/3L over 8 preference-eligible stimulus vote(s), sign test p=0.656), mean preference +0.0% across 8 paired run(s) — no improvement

Next action: Compare winning and losing scenarios to isolate where the target helps versus hurts.

State: VALID_NO_CHANGE (no_credible_preference_change)

Gate evidence: n=8; 3W/2T/3L; d=6; p=0.656; net +0.0%

Warnings: Activation: isolated 6/8; plugin 4/8

Overfit: Moderate (score 0.39)

Repeated-run reliability (not used by the gate): 8 paired runs (3W/2T/3L).

Weak or warning scenarios:

Scenario Preference gate Net win Δ Pref Runs (W/T/L)
▲ Decline to invent evaluation problems in an already-clean project Eligible +100.0% +40.0% 1/0/0
▼ Detect a project evaluated twice under different global properties Eligible -100.0% -40.0% 0/0/1
= Gather a measurement before proposing evaluation fixes for an unremarkable project Eligible +0.0% +0.0% 0/1/0
▼ Recognize TreatAsLocalProperty overuse versus one justified entry Eligible -100.0% -40.0% 0/0/1
▼ Redirect an incremental-rebuild complaint mistakenly framed as an evaluation problem Eligible -100.0% -40.0% 0/0/1
= Triage which of two property functions actually costs evaluation time Eligible +0.0% +0.0% 0/1/0

Illustrative judge evidence:

  • Detect a project evaluated twice under different global properties: A gives the more actionable and better-supported diagnosis: it verifies the differing property appears unused and that both calls generate the same output, then offers the appropriate remove-the-duplicate-or-wire-the-feature alternatives. B reaches the core finding but oversta...

This is one example, not the aggregate verdict. Open Full Results for every judgment.

➖ Improvement signal, tie-limited — eval-performance (gpt-5.6-luna)

Why: Net win +50.0% (4W/4T/0L over 8 preference-eligible stimulus vote(s), sign test p=0.063), mean preference +35.0% across 8 paired run(s) — not credible — 4 of 8 preference-eligible stimulus vote(s) tied, leaving only 4 discordant preference vote(s). The sign test conditions on non-tie stimulus votes and cannot reach 0.05 below 5, so no record could have passed here — this is not a measured null. Either the skill is inert on these scenarios (make them discriminate) or the eval needs more distinct stimuli to clear the ties

Next action: The signal favors the target, but ties leave too few discordant tasks; inspect ties and predeclare more discriminating breadth.

State: VALID_NO_CHANGE (no_credible_preference_change)

Gate evidence: n=8; 4W/4T/0L; d=4; p=0.063; net +50.0%

Overfit: Moderate (score 0.40)

Repeated-run reliability (not used by the gate): 8 paired runs (4W/4T/0L).

Weak or warning scenarios:

Scenario Preference gate Net win Δ Pref Runs (W/T/L)
= Detect a project evaluated twice under different global properties Eligible +0.0% +0.0% 0/1/0
= Diagnose stacked evaluation-time patterns (deep imports, broad glob, file-I/O property function) Eligible +0.0% +0.0% 0/1/0
= Recognize TreatAsLocalProperty overuse versus one justified entry Eligible +0.0% +0.0% 0/1/0
= Redirect a compile-time slowdown mistakenly framed as an evaluation problem Eligible +0.0% +0.0% 0/1/0

Illustrative judge evidence:

  • Detect a project evaluated twice under different global properties: Position-swap inconsistent (forward: B, reverse: A). Defaulting to tie.

This is one example, not the aggregate verdict. Open Full Results for every judgment.

➖ Improvement signal, tie-limited — extension-points (claude-sonnet-5)

Why: Net win +28.6% (3W/3T/1L over 7 preference-eligible stimulus vote(s), sign test p=0.312), mean preference +8.9% across 9 paired run(s), 2 dormancy stimulus/stimuli excluded from preference — not credible — 3 of 7 preference-eligible stimulus vote(s) tied, leaving only 4 discordant preference vote(s). The sign test conditions on non-tie stimulus votes and cannot reach 0.05 below 5, so no record could have passed here — this is not a measured null. Either the skill is inert on these scenarios (make them discriminate) or the eval needs more distinct stimuli to clear the ties

Next action: The signal favors the target, but ties leave too few discordant tasks; inspect ties and predeclare more discriminating breadth.

State: VALID_NO_CHANGE (no_credible_preference_change)

Gate evidence: n=7; 3W/3T/1L; d=4; p=0.312; net +28.6%; 2 dormancy excluded

Warnings: Activation: isolated 4/7; plugin 6/7

Overfit: Low (score 0.19)

Repeated-run reliability (not used by the gate): 9 paired runs (3W/5T/1L).

Weak or warning scenarios:

Scenario Preference gate Net win Δ Pref Runs (W/T/L)
= Decline Kubernetes readiness probe configuration Excluded (activation contract) +0.0% +0.0% 0/1/0
= Diagnose NuGet package and repo extension conflicts Eligible +0.0% +0.0% 0/1/0
= Diagnose a broken per-TFM forwarder Eligible +0.0% +0.0% 0/1/0
▲ Diagnose a package ID and file-name mismatch Eligible +100.0% +40.0% 1/0/0
▼ Diagnose build extension point failures Eligible -100.0% -40.0% 0/0/1
= Leave incremental target authoring to its owning workflow Excluded (activation contract) +0.0% +0.0% 0/1/0
= Review packed layout without a false missing-file bug Eligible +0.0% +0.0% 0/1/0

Illustrative judge evidence:

  • Diagnose NuGet package and repo extension conflicts: Position-swap inconsistent (forward: A, reverse: B). Defaulting to tie.

This is one example, not the aggregate verdict. Open Full Results for every judgment.

➖ Improvement signal, unproven — extension-points (gpt-5.6-luna)

Why: Net win +42.9% (4W/2T/1L over 7 preference-eligible stimulus vote(s), sign test p=0.188), mean preference +24.4% across 9 paired run(s), 2 dormancy stimulus/stimuli excluded from preference — not credible (sign test p=0.188 > 0.05)

Next action: The signal favors the target but is inconsistent; inspect tied or lost scenarios and fix weak behavior.

State: VALID_NO_CHANGE (no_credible_preference_change)

Gate evidence: n=7; 4W/2T/1L; d=5; p=0.188; net +42.9%; 2 dormancy excluded

Warnings: Activation: isolated 7/7; plugin 6/7

Overfit: Moderate (score 0.31)

Repeated-run reliability (not used by the gate): 9 paired runs (5W/3T/1L).

Weak or warning scenarios:

Scenario Preference gate Net win Δ Pref Runs (W/T/L)
▲ Create extensibility hooks for a custom SDK target file Eligible +100.0% +100.0% 1/0/0
= Decline Kubernetes readiness probe configuration Excluded (activation contract) +0.0% +0.0% 0/1/0
= Diagnose NuGet package and repo extension conflicts Eligible +0.0% +0.0% 0/1/0
▼ Diagnose a package ID and file-name mismatch Eligible -100.0% -40.0% 0/0/1
= Review packed layout without a false missing-file bug Eligible +0.0% +0.0% 0/1/0

Illustrative judge evidence:

  • Diagnose NuGet package and repo extension conflicts: Position-swap inconsistent (forward: B, reverse: A). Defaulting to tie.

This is one example, not the aggregate verdict. Open Full Results for every judgment.

➖ Mixed evidence — including-generated-files (claude-sonnet-5)

Why: Net win +0.0% (3W/1T/3L over 7 preference-eligible stimulus vote(s), sign test p=0.656), mean preference +0.0% across 8 paired run(s), 1 dormancy stimulus/stimuli excluded from preference — no improvement

Next action: Compare winning and losing scenarios to isolate where the target helps versus hurts.

State: VALID_NO_CHANGE (no_credible_preference_change)

Gate evidence: n=7; 3W/1T/3L; d=6; p=0.656; net +0.0%; 1 dormancy excluded

Warnings: Activation: isolated 4/7; plugin 6/7

Overfit: Moderate (score 0.38)

Repeated-run reliability (not used by the gate): 8 paired runs (3W/2T/3L).

Weak or warning scenarios:

Scenario Preference gate Net win Δ Pref Runs (W/T/L)
= Decline a Python runtime report-path problem Excluded (activation contract) +0.0% +0.0% 0/1/0
▲ Diagnose hardcoded obj path for generated source Eligible +100.0% +40.0% 1/0/0
= Diagnose missing clean tracking for generated source Eligible +0.0% +0.0% 0/1/0
▼ Diagnose missing generated source inclusion Eligible -100.0% -40.0% 0/0/1
▼ Diagnose project-level glob for generated source Eligible -100.0% -40.0% 0/0/1
▼ Diagnose wrong hook for generated source files Eligible -100.0% -40.0% 0/0/1

Illustrative judge evidence:

  • Diagnose missing clean tracking for generated source: Position-swap inconsistent (forward: B, reverse: tie). Defaulting to tie.

This is one example, not the aggregate verdict. Open Full Results for every judgment.

➖ Baseline signal, tie-limited — including-generated-files (gpt-5.6-luna)

Why: Net win -28.6% (1W/3T/3L over 7 preference-eligible stimulus vote(s), sign test p=0.312), mean preference -10.0% across 8 paired run(s), 1 dormancy stimulus/stimuli excluded from preference — no improvement

Next action: Evidence leans baseline but is not credible; inspect losses and make tied tasks discriminate.

State: VALID_NO_CHANGE (no_credible_preference_change)

Gate evidence: n=7; 1W/3T/3L; d=4; p=0.312; net -28.6%; 1 dormancy excluded

Overfit: Moderate (score 0.36)

Repeated-run reliability (not used by the gate): 8 paired runs (1W/4T/3L).

Weak or warning scenarios:

Scenario Preference gate Net win Δ Pref Runs (W/T/L)
= Decline a Python runtime report-path problem Excluded (activation contract) +0.0% +0.0% 0/1/0
= Diagnose hardcoded obj path for generated source Eligible +0.0% +0.0% 0/1/0
▼ Diagnose missing clean tracking for generated source Eligible -100.0% -40.0% 0/0/1
▼ Diagnose missing generated source inclusion Eligible -100.0% -40.0% 0/0/1
= Diagnose missing output registration for generated non-code file Eligible +0.0% +0.0% 0/1/0
▼ Diagnose project-level glob for generated source Eligible -100.0% -40.0% 0/0/1
= Diagnose wrong hook for generated source files Eligible +0.0% +0.0% 0/1/0

Illustrative judge evidence:

  • Diagnose hardcoded obj path for generated source: Position-swap inconsistent (forward: B, reverse: A). Defaulting to tie.

This is one example, not the aggregate verdict. Open Full Results for every judgment.

➖ Improvement signal, unproven — incremental-build (claude-sonnet-5)

Why: Net win +11.1% (4W/2T/3L over 9 preference-eligible stimulus vote(s), sign test p=0.500), mean preference +4.4% across 9 paired run(s) — not credible (sign test p=0.500 > 0.05)

Next action: The signal favors the target but is inconsistent; inspect tied or lost scenarios and fix weak behavior.

State: VALID_NO_CHANGE (no_credible_preference_change)

Gate evidence: n=9; 4W/2T/3L; d=7; p=0.500; net +11.1%

Warnings: Activation: isolated 4/9; plugin 4/9

Overfit: Moderate (score 0.39)

Repeated-run reliability (not used by the gate): 9 paired runs (4W/2T/3L).

Weak or warning scenarios:

Scenario Preference gate Net win Δ Pref Runs (W/T/L)
= Distinguish a cold first build from broken incrementality Eligible +0.0% +0.0% 0/1/0
▼ Explain slower builds when MSBuild skipped everything Eligible -100.0% -40.0% 0/0/1
▼ Explain why clean leaves generated hash source behind Eligible -100.0% -40.0% 0/0/1
▲ Fix broken incremental targets and clean tracking Eligible +100.0% +40.0% 1/0/0
= Identify a volatile output path that defeats incrementality Eligible +0.0% +0.0% 0/1/0
▼ Read a diagnostic log to find the stale input Eligible -100.0% -40.0% 0/0/1

Illustrative judge evidence:

  • Distinguish a cold first build from broken incrementality: Position-swap inconsistent (forward: tie, reverse: B). Defaulting to tie.

This is one example, not the aggregate verdict. Open Full Results for every judgment.

➖ Improvement signal, tie-limited — item-management (claude-sonnet-5)

Why: Net win +16.7% (2W/3T/1L over 6 preference-eligible stimulus vote(s), sign test p=0.500), mean preference +5.7% across 7 paired run(s), 1 dormancy stimulus/stimuli excluded from preference — not credible — 3 of 6 preference-eligible stimulus vote(s) tied, leaving only 3 discordant preference vote(s). The sign test conditions on non-tie stimulus votes and cannot reach 0.05 below 5, so no record could have passed here — this is not a measured null. Either the skill is inert on these scenarios (make them discriminate) or the eval needs more distinct stimuli to clear the ties

Next action: The signal favors the target, but ties leave too few discordant tasks; inspect ties and predeclare more discriminating breadth.

State: VALID_NO_CHANGE (no_credible_preference_change)

Gate evidence: n=6; 2W/3T/1L; d=3; p=0.500; net +16.7%; 1 dormancy excluded

Warnings: Activation: isolated 2/6; plugin 3/6

Overfit: Moderate (score 0.37)

Repeated-run reliability (not used by the gate): 7 paired runs (2W/4T/1L).

Weak or warning scenarios:

Scenario Preference gate Net win Δ Pref Runs (W/T/L)
= Decline JavaScript configuration merging Excluded (activation contract) +0.0% +0.0% 0/1/0
▲ Diagnose an ineffective Compile Remove that does not match the glob Eligible +100.0% +40.0% 1/0/0
▲ Diagnose item group and batching issues Eligible +100.0% +40.0% 1/0/0
▼ Diagnose real and claimed item problems in a code generation pipeline Eligible -100.0% -40.0% 0/0/1
= Fix item management anti-patterns Eligible +0.0% +0.0% 0/1/0
= Leave already-correct item management unchanged Eligible +0.0% +0.0% 0/1/0
= Leave correct single-list batching unchanged Eligible +0.0% +0.0% 0/1/0

Illustrative judge evidence:

  • Diagnose real and claimed item problems in a code generation pipeline: Both fix the visible duplicate-warning and validation symptoms and verify build/clean behavior, but both miss two requested structural corrections: an exact evaluation-time generated-source declaration and removal of CleanGeneratedCode. A is marginally preferable because it pr...

This is one example, not the aggregate verdict. Open Full Results for every judgment.

➖ Mixed evidence — item-management (gpt-5.6-luna)

Why: Net win +0.0% (2W/2T/2L over 6 preference-eligible stimulus vote(s), sign test p=0.687), mean preference +5.7% across 7 paired run(s), 1 dormancy stimulus/stimuli excluded from preference — no improvement

Next action: Compare winning and losing scenarios to isolate where the target helps versus hurts.

State: VALID_NO_CHANGE (no_credible_preference_change)

Gate evidence: n=6; 2W/2T/2L; d=4; p=0.687; net +0.0%; 1 dormancy excluded

Warnings: Activation: isolated 6/6; plugin 5/6

Overfit: Moderate (score 0.49)

Repeated-run reliability (not used by the gate): 7 paired runs (3W/2T/2L).

Weak or warning scenarios:

Scenario Preference gate Net win Δ Pref Runs (W/T/L)
▼ Diagnose an ineffective Compile Remove that does not match the glob Eligible -100.0% -100.0% 0/0/1
= Diagnose item group and batching issues Eligible +0.0% +0.0% 0/1/0
= Diagnose real and claimed item problems in a code generation pipeline Eligible +0.0% +0.0% 0/1/0
▼ Leave already-correct item management unchanged Eligible -100.0% -40.0% 0/0/1

Illustrative judge evidence:

  • Diagnose an ineffective Compile Remove that does not match the glob: Response A provides the correct, precise solution to the glob-matching problem by changing the pattern to Generated\.g.cs, which directly matches the actual file naming convention. Response B's solution using Generated\**\.cs is overly broad and would exclude ALL C# files...

This is one example, not the aggregate verdict. Open Full Results for every judgment.

➖ Improvement signal, tie-limited — msbuild-antipatterns (claude-sonnet-5)

Why: Net win +14.3% (2W/4T/1L over 7 preference-eligible stimulus vote(s), sign test p=0.500), mean preference +0.0% across 8 paired run(s), 1 dormancy stimulus/stimuli excluded from preference — not credible — 4 of 7 preference-eligible stimulus vote(s) tied, leaving only 3 discordant preference vote(s). The sign test conditions on non-tie stimulus votes and cannot reach 0.05 below 5, so no record could have passed here — this is not a measured null. Either the skill is inert on these scenarios (make them discriminate) or the eval needs more distinct stimuli to clear the ties

Next action: The signal favors the target, but ties leave too few discordant tasks; inspect ties and predeclare more discriminating breadth.

State: VALID_NO_CHANGE (no_credible_preference_change)

Gate evidence: n=7; 2W/4T/1L; d=3; p=0.500; net +14.3%; 1 dormancy excluded

Warnings: Activation: isolated 3/7; plugin 3/7

Overfit: Moderate (score 0.30)

Repeated-run reliability (not used by the gate): 8 paired runs (2W/4T/2L).

Weak or warning scenarios:

Scenario Preference gate Net win Δ Pref Runs (W/T/L)
▲ Add a module to an F# project Eligible +100.0% +40.0% 1/0/0
= Add a signature file to define public API Eligible +0.0% +0.0% 0/1/0
= Distinguish a style backslash from a real cross-platform backslash bug Eligible +0.0% +0.0% 0/1/0
= Fix broken file order causing FS0039 Eligible +0.0% +0.0% 0/1/0
▲ Judge an unguarded import inside a NuGet package build folder Eligible +100.0% +40.0% 1/0/0
= Leave a clean project without inventing anti-patterns Eligible +0.0% +0.0% 0/1/0
▼ Non-activation: migrate a legacy project to SDK style Excluded (activation contract) -100.0% -40.0% 0/0/1
▼ Review MSBuild files for anti-patterns and style issues Eligible -100.0% -40.0% 0/0/1

Illustrative judge evidence:

  • Add a signature file to define public API: The two solutions are substantively equivalent and fully satisfy the task: each adds a correct public type signature, compiles it in the required order, and verifies a clean build.

This is one example, not the aggregate verdict. Open Full Results for every judgment.

➖ Baseline signal, tie-limited — msbuild-antipatterns (gpt-5.6-luna)

Why: Net win -28.6% (1W/3T/3L over 7 preference-eligible stimulus vote(s), sign test p=0.312), mean preference -10.0% across 8 paired run(s), 1 dormancy stimulus/stimuli excluded from preference — no improvement

Next action: Evidence leans baseline but is not credible; inspect losses and make tied tasks discriminate.

State: VALID_NO_CHANGE (no_credible_preference_change)

Gate evidence: n=7; 1W/3T/3L; d=4; p=0.312; net -28.6%; 1 dormancy excluded

Warnings: Activation: isolated 5/7; plugin 4/7

Overfit: Moderate (score 0.24)

Repeated-run reliability (not used by the gate): 8 paired runs (1W/4T/3L).

Weak or warning scenarios:

Scenario Preference gate Net win Δ Pref Runs (W/T/L)
▼ Add a module to an F# project Eligible -100.0% -40.0% 0/0/1
= Add a signature file to define public API Eligible +0.0% +0.0% 0/1/0
= Distinguish a style backslash from a real cross-platform backslash bug Eligible +0.0% +0.0% 0/1/0
▼ Fix broken file order causing FS0039 Eligible -100.0% -40.0% 0/0/1
▼ Judge an unguarded import inside a NuGet package build folder Eligible -100.0% -40.0% 0/0/1
= Leave a clean project without inventing anti-patterns Eligible +0.0% +0.0% 0/1/0
= Non-activation: migrate a legacy project to SDK style Excluded (activation contract) +0.0% +0.0% 0/1/0

Illustrative judge evidence:

  • Add a module to an F# project: Both responses successfully completed the core task, but Response A provides superior transparency and evidence of implementation. Response A explicitly shows the error handling code with match expressions and result type handling, with multiple iterations demonstrating thorou...

This is one example, not the aggregate verdict. Open Full Results for every judgment.

➖ Improvement signal, unproven — property-patterns (claude-sonnet-5)

Why: Net win +14.3% (3W/2T/2L over 7 preference-eligible stimulus vote(s), sign test p=0.500), mean preference +5.7% across 7 paired run(s) — not credible (sign test p=0.500 > 0.05)

Next action: The signal favors the target but is inconsistent; inspect tied or lost scenarios and fix weak behavior.

State: VALID_NO_CHANGE (no_credible_preference_change)

Gate evidence: n=7; 3W/2T/2L; d=5; p=0.500; net +14.3%

Warnings: Activation: isolated 6/7; plugin 6/7

Overfit: Moderate (score 0.24)

Repeated-run reliability (not used by the gate): 7 paired runs (3W/2T/2L).

Weak or warning scenarios:

Scenario Preference gate Net win Δ Pref Runs (W/T/L)
= Choose the correct shared build file for a post-build target Eligible +0.0% +0.0% 0/1/0
▼ Choose the right OS detection for a cross-platform property Eligible -100.0% -40.0% 0/0/1
= Diagnose a TargetFramework condition that never applies in props Eligible +0.0% +0.0% 0/1/0
▼ Leave already-correct shared properties unchanged Eligible -100.0% -100.0% 0/0/1

Illustrative judge evidence:

  • Choose the correct shared build file for a post-build target: Position-swap inconsistent (forward: A, reverse: B). Defaulting to tie.

This is one example, not the aggregate verdict. Open Full Results for every judgment.

➖ Improvement signal, tie-limited — property-patterns (gpt-5.6-luna)

Why: Net win +42.9% (3W/4T/0L over 7 preference-eligible stimulus vote(s), sign test p=0.125), mean preference +17.1% across 7 paired run(s) — not credible — 4 of 7 preference-eligible stimulus vote(s) tied, leaving only 3 discordant preference vote(s). The sign test conditions on non-tie stimulus votes and cannot reach 0.05 below 5, so no record could have passed here — this is not a measured null. Either the skill is inert on these scenarios (make them discriminate) or the eval needs more distinct stimuli to clear the ties

Next action: The signal favors the target, but ties leave too few discordant tasks; inspect ties and predeclare more discriminating breadth.

State: VALID_NO_CHANGE (no_credible_preference_change)

Gate evidence: n=7; 3W/4T/0L; d=3; p=0.125; net +42.9%

Warnings: Activation: isolated 6/7; plugin 5/7

Overfit: Moderate (score 0.44)

Repeated-run reliability (not used by the gate): 7 paired runs (3W/4T/0L).

Weak or warning scenarios:

Scenario Preference gate Net win Δ Pref Runs (W/T/L)
= Choose the correct shared build file for a post-build target Eligible +0.0% +0.0% 0/1/0
▲ Choose the right OS detection for a cross-platform property Eligible +100.0% +40.0% 1/0/0
= Diagnose a TargetFramework condition that never applies in props Eligible +0.0% +0.0% 0/1/0
= Diagnose multi-level property hierarchy bugs Eligible +0.0% +0.0% 0/1/0
= Leave already-correct shared properties unchanged Eligible +0.0% +0.0% 0/1/0

Illustrative judge evidence:

  • Choose the correct shared build file for a post-build target: Position-swap inconsistent (forward: B, reverse: A). Defaulting to tie.

This is one example, not the aggregate verdict. Open Full Results for every judgment.

Details for 3 results were omitted to keep this comment under GitHub's limit. Open the workflow summary or Full Results for the complete breakdown.

🔍 Full Results - all metrics and investigation details

To investigate non-passing or warning results, paste this to your AI coding agent:

For PR 1215 in dotnet/skills, download eval artifacts with gh run download 37048637618 --repo dotnet/skills --pattern "vally-results-*" --dir ./eval-results, then fetch https://raw.githubusercontent.com/dotnet/skills/970e0dc1bd6bfcd78f4b1263252d049cda277641/eng/vally-adapter/InvestigatingResults.md and follow it. Classify each result as measurement-invalid, underpowered, not-proven, preference-loss, or passing-with-warning. Use stateReason, result accounting, weak scenarios, and judge evidence to give the cause and exact next fix.

⚠️ Session replay telemetry was not published because the auxiliary dotnet/skills-data publisher failed. The evaluation verdicts above remain authoritative; maintainers must repair SKILLS_DATA_TOKEN or the publisher before replay links are available.

github-actions Bot added a commit that referenced this pull request Oct 2, 2026
@github-actions github-actions Bot added waiting-on-review PR state label and removed pr-state/evals-in-progress PR evaluations are in progress labels Oct 2, 2026
@github-actions

github-actions Bot commented Oct 2, 2026

Copy link
Copy Markdown
Contributor

✅ Evaluation passed for 970e0dc. cc @dotnet/msbuild @JanKrivanek @YuliiaKovalova — please review.

Comment thread plugins/dotnet-msbuild/skills/build-parallelism/SKILL.md Outdated
Co-authored-by: Copilot App <223556219+Copilot@users.noreply.github.com>
Copilot AI lite review requested due to automatic review settings October 5, 2026 11:21

Copilot AI left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Copilot review overview

🔵 Needs a closer look

One or more issues must be addressed before approval.

Review effort: Lite
Findings: None

@github-actions github-actions Bot added pr-state/ready-for-eval PR is mergeable and awaiting evaluation and removed waiting-on-review PR state label labels Oct 5, 2026
Comment thread plugins/dotnet-msbuild/skills/build-parallelism/SKILL.md Outdated
Co-authored-by: Copilot App <223556219+Copilot@users.noreply.github.com>

Copilot AI left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Copilot review overview

🔵 Needs a closer look

Three moderate issues remain in binlog replay and routing guidance, plus one workflow numbering nit.

Review effort: Lite
Findings: None

Previously missed (1)

In code that hasn't changed since last review

Low severity Fallback checklist numbering starts at step 2

plugins/​dotnet-msbuild/​skills/​incremental-build/​SKILL.md:108

This standalone fallback checklist starts at step 2 and continues with steps 3–5, even though the preceding MCP checklist is a separate section. Renumber these items 1–4 so the fallback workflow is not presented with a missing first step.

@github-actions

github-actions Bot commented Oct 5, 2026

Copy link
Copy Markdown
Contributor

📊 Skill and Agent Evaluation Results

30 model/target results across 15 targets and 2 models — ✅ 1 improved, ➖ 29 results without a clear winner, ⚠️ 0 invalid or underpowered, ⛔ 0 activation contract failures, 📉 0 preference losses (report only).

Measurement identity: evaluated commit b97f9c542191718f6b4d15e24c1be7af9908b9e8; 2 judge models.

Measurement health: 30 expected / 30 observed / 30 written; 0 missing, 0 unexpected, 0 invalid; 0 recovered comparison error slots and 0 unresolved comparison error slots.

Objective completion gate: not enabled. Aggregate completion transitions are telemetry only, so this report does not claim that zero objective regressions were proven.

A result passes only when preference-eligible distinct-stimulus votes have aggregate net win of at least 20% and an exact one-sided sign-test result of p ≤ 0.05, and every explicit dormancy activation contract passes. Repeated runs measure reliability only.

Target Model Verdict Gate evidence Overfit Warnings Next action
agent.msbuild claude-sonnet-5 ➖ Improvement signal, tie-limited n=5; 3W/1T/1L; d=4; p=0.312; net +40.0% — — The signal favors the target, but ties leave too few discordant tasks; inspect ties and predeclare more discriminating breadth.
agent.msbuild gpt-5.6-luna ➖ Improvement signal, tie-limited n=5; 3W/2T/0L; d=3; p=0.125; net +60.0% — — The signal favors the target, but ties leave too few discordant tasks; inspect ties and predeclare more discriminating breadth.
binlog-failure-analysis claude-sonnet-5 ➖ Improvement signal, tie-limited n=6; 3W/2T/1L; d=4; p=0.312; net +33.3%; 1 dormancy excluded 🟡 0.30 Activation: isolated 6/7; plugin 6/7 The signal favors the target, but ties leave too few discordant tasks; inspect ties and predeclare more discriminating breadth.
binlog-failure-analysis gpt-5.6-luna ➖ Improvement signal, unproven n=6; 4W/1T/1L; d=5; p=0.188; net +50.0%; 1 dormancy excluded 🟡 0.26 Activation: isolated 6/7; plugin 6/7 The signal favors the target but is inconsistent; inspect tied or lost scenarios and fix weak behavior.
binlog-generation claude-sonnet-5 ➖ Improvement signal, unproven n=7; 4W/1T/2L; d=6; p=0.344; net +28.6%; 1 dormancy excluded 🔴 0.50 Activation: isolated 5/7; plugin 6/7 The signal favors the target but is inconsistent; inspect tied or lost scenarios and fix weak behavior.
binlog-generation gpt-5.6-luna ➖ Improvement signal, unproven n=7; 5W/1T/1L; d=6; p=0.109; net +57.1%; 1 dormancy excluded 🟡 0.36 Activation: isolated 7/7; plugin 5/7 The signal favors the target but is inconsistent; inspect tied or lost scenarios and fix weak behavior.
build-parallelism claude-sonnet-5 ➖ Baseline signal, tie-limited n=6; 1W/3T/2L; d=3; p=0.500; net -16.7%; 1 dormancy excluded 🟡 0.36 Activation: isolated 5/6; plugin 4/6; Activation-only stop: isolated 1 failed run Inspect activation-only failed runs before rewriting skill content; the model stopped after loading a skill.
build-parallelism gpt-5.6-luna ➖ Improvement signal, tie-limited n=6; 4W/2T/0L; d=4; p=0.063; net +66.7%; 1 dormancy excluded ✅ 0.06 Activation-only stop: isolated 1 failed run; Activation-only stop: plugin 1 failed run Inspect activation-only failed runs before rewriting skill content; the model stopped after loading a skill.
build-perf-baseline claude-sonnet-5 ➖ Mixed evidence n=6; 3W/0T/3L; d=6; p=0.656; net +0.0%; 1 dormancy excluded 🟡 0.44 Activation: isolated 5/6; plugin 5/6 Compare winning and losing scenarios to isolate where the target helps versus hurts.
build-perf-baseline gpt-5.6-luna ➖ Improvement signal, tie-limited n=6; 3W/2T/1L; d=4; p=0.312; net +33.3%; 1 dormancy excluded 🟡 0.24 Activation: isolated 5/6; plugin 4/6 The signal favors the target, but ties leave too few discordant tasks; inspect ties and predeclare more discriminating breadth.
build-perf-diagnostics claude-sonnet-5 ➖ Improvement signal, unproven n=7; 3W/2T/2L; d=5; p=0.500; net +14.3%; 1 dormancy excluded 🟡 0.47 Activation: isolated 7/7; plugin 2/7 The signal favors the target but is inconsistent; inspect tied or lost scenarios and fix weak behavior.
build-perf-diagnostics gpt-5.6-luna ➖ Improvement signal, unproven n=7; 4W/2T/1L; d=5; p=0.188; net +42.9%; 1 dormancy excluded 🟡 0.43 Activation: isolated 7/7; plugin 3/7 The signal favors the target but is inconsistent; inspect tied or lost scenarios and fix weak behavior.
check-bin-obj-clash claude-sonnet-5 ➖ Improvement signal, tie-limited n=7; 3W/4T/0L; d=3; p=0.125; net +42.9% ✅ 0.16 Activation: isolated 7/7; plugin 5/7 The signal favors the target, but ties leave too few discordant tasks; inspect ties and predeclare more discriminating breadth.
check-bin-obj-clash gpt-5.6-luna ➖ Improvement signal, tie-limited n=7; 2W/4T/1L; d=3; p=0.500; net +14.3% 🟡 0.31 Activation: isolated 7/7; plugin 6/7 The signal favors the target, but ties leave too few discordant tasks; inspect ties and predeclare more discriminating breadth.
directory-build-organization claude-sonnet-5 ✅ Improved n=6; 5W/1T/0L; d=5; p=0.031; net +83.3%; 1 dormancy excluded 🟡 0.38 Activation: isolated 4/6; plugin 4/6 Fix activation gaps; Review overfit evidence.
directory-build-organization gpt-5.6-luna ➖ Baseline signal, tie-limited n=6; 1W/3T/2L; d=3; p=0.500; net -16.7%; 1 dormancy excluded 🟡 0.31 Activation: isolated 6/6; plugin 4/6 Evidence leans baseline but is not credible; inspect losses and make tied tasks discriminate.
eval-performance claude-sonnet-5 ➖ Improvement signal, unproven n=8; 5W/2T/1L; d=6; p=0.109; net +50.0% 🟡 0.27 Activation: isolated 5/8; plugin 5/8 The signal favors the target but is inconsistent; inspect tied or lost scenarios and fix weak behavior.
eval-performance gpt-5.6-luna ➖ Improvement signal, unproven n=8; 5W/2T/1L; d=6; p=0.109; net +50.0% 🟡 0.38 Activation: isolated 8/8; plugin 7/8 The signal favors the target but is inconsistent; inspect tied or lost scenarios and fix weak behavior.
extension-points claude-sonnet-5 ➖ Improvement signal, tie-limited n=7; 3W/3T/1L; d=4; p=0.312; net +28.6%; 2 dormancy excluded 🟡 0.27 Activation: isolated 6/7; plugin 4/7 The signal favors the target, but ties leave too few discordant tasks; inspect ties and predeclare more discriminating breadth.
extension-points gpt-5.6-luna ➖ Improvement signal, tie-limited n=7; 2W/4T/1L; d=3; p=0.500; net +14.3%; 2 dormancy excluded 🟡 0.39 Activation: isolated 7/7; plugin 5/7 The signal favors the target, but ties leave too few discordant tasks; inspect ties and predeclare more discriminating breadth.
including-generated-files claude-sonnet-5 ➖ Improvement signal, tie-limited n=7; 2W/4T/1L; d=3; p=0.500; net +14.3%; 1 dormancy excluded 🟡 0.35 Activation: isolated 4/7; plugin 4/7 The signal favors the target, but ties leave too few discordant tasks; inspect ties and predeclare more discriminating breadth.
including-generated-files gpt-5.6-luna ➖ Mixed evidence n=7; 3W/1T/3L; d=6; p=0.656; net +0.0%; 1 dormancy excluded 🟡 0.40 — Compare winning and losing scenarios to isolate where the target helps versus hurts.
incremental-build claude-sonnet-5 ➖ Improvement signal, tie-limited n=9; 3W/5T/1L; d=4; p=0.312; net +22.2% 🟡 0.24 Activation: isolated 3/9; plugin 5/9 The signal favors the target, but ties leave too few discordant tasks; inspect ties and predeclare more discriminating breadth.
incremental-build gpt-5.6-luna ➖ Improvement signal, unproven n=9; 4W/4T/1L; d=5; p=0.188; net +33.3% 🟡 0.22 Activation: isolated 8/9; plugin 6/9 The signal favors the target but is inconsistent; inspect tied or lost scenarios and fix weak behavior.
item-management claude-sonnet-5 ➖ Mixed evidence n=6; 2W/2T/2L; d=4; p=0.687; net +0.0%; 1 dormancy excluded 🟡 0.33 Activation: isolated 4/6; plugin 3/6 Compare winning and losing scenarios to isolate where the target helps versus hurts.
item-management gpt-5.6-luna ➖ Mixed evidence n=6; 1W/4T/1L; d=2; p=0.750; net +0.0%; 1 dormancy excluded 🟡 0.30 Activation: isolated 6/6; plugin 5/6 Compare winning and losing scenarios to isolate where the target helps versus hurts.
msbuild-antipatterns claude-sonnet-5 ➖ Mixed evidence n=7; 2W/3T/2L; d=4; p=0.687; net +0.0%; 1 dormancy excluded 🟡 0.30 Activation: isolated 3/7; plugin 2/7 Compare winning and losing scenarios to isolate where the target helps versus hurts.
msbuild-antipatterns gpt-5.6-luna ➖ Mixed evidence n=7; 1W/5T/1L; d=2; p=0.750; net +0.0%; 1 dormancy excluded 🟡 0.29 Activation: isolated 5/7; plugin 3/7 Compare winning and losing scenarios to isolate where the target helps versus hurts.
property-patterns claude-sonnet-5 ➖ Baseline signal, unproven n=7; 2W/0T/5L; d=7; p=0.227; net -42.9% 🟡 0.27 Activation: isolated 6/7; plugin 6/7 Evidence leans baseline but is not credible; inspect losing scenarios for recurring defects.
property-patterns gpt-5.6-luna ➖ Improvement signal, unproven n=7; 3W/2T/2L; d=5; p=0.500; net +14.3% 🟡 0.37 Activation: isolated 6/7; plugin 6/7 The signal favors the target but is inconsistent; inspect tied or lost scenarios and fix weak behavior.
ℹ️ How to read this report
  • ✅ Improved — the result passed both the statistical gate and the 20% practical net-win floor.
  • ➖ No clear winner — the result is valid but did not pass both gates. The label distinguishes all ties, mixed evidence, directional but unproven evidence, and credible effects below the practical floor.
  • ⚠️ Invalid / underpowered — the gate withheld a quality verdict. Fix the measurement before judging the target.
  • ⛔ Activation contract failed — the isolated target activated on an explicit dormancy scenario. Dormancy preference is excluded, but this routing failure still blocks a pass.
  • 📉 Preference loss — the LLM judge credibly preferred baseline. It is report-only, not objective completion proof.
  • Gate evidence — n preference-eligible distinct-stimulus votes, W/T/L stimulus votes, d discordant votes, exact one-sided p, net win, and the count of separately retained dormancy stimuli. The p value applies to one model/target result; no matrix-wide multiple-comparison correction is applied.
  • Overfit — overfitting-judge severity (✅ Low, 🟡 Moderate, 🔴 High, — none) and score.
  • Warnings — activation, timeout, retry recovery, or unresolved comparison conditions that need attention.
  • Do not add repeated runs to increase statistical power. Do not add stimuli after seeing a near-pass unless the new breadth is predeclared for a new experiment.
➖ Improvement signal, tie-limited — agent.msbuild (claude-sonnet-5)

Why: Net win +40.0% (3W/1T/1L over 5 preference-eligible stimulus vote(s), sign test p=0.312), mean preference +4.0% across 5 paired run(s) — not credible — 1 of 5 preference-eligible stimulus vote(s) tied, leaving only 4 discordant preference vote(s). The sign test conditions on non-tie stimulus votes and cannot reach 0.05 below 5, so no record could have passed here — this is not a measured null. Either the agent is inert on these scenarios (make them discriminate) or the eval needs more distinct stimuli to clear the ties

Next action: The signal favors the target, but ties leave too few discordant tasks; inspect ties and predeclare more discriminating breadth.

State: VALID_NO_CHANGE (no_credible_preference_change)

Gate evidence: n=5; 3W/1T/1L; d=4; p=0.312; net +40.0%

Repeated-run reliability (not used by the gate): 5 paired runs (3W/1T/1L).

Weak or warning scenarios:

Scenario Preference gate Net win Δ Pref Runs (W/T/L)
= Advise on project file organization Eligible +0.0% +0.0% 0/1/0
▼ Triage a build failure and route to appropriate analysis Eligible -100.0% -100.0% 0/0/1

Illustrative judge evidence:

  • Advise on project file organization: Position-swap inconsistent (forward: baseline, reverse: skill). Defaulting to tie.

This is one example, not the aggregate verdict. Open Full Results for every judgment.

➖ Improvement signal, tie-limited — agent.msbuild (gpt-5.6-luna)

Why: Net win +60.0% (3W/2T/0L over 5 preference-eligible stimulus vote(s), sign test p=0.125), mean preference +36.0% across 5 paired run(s) — not credible — 2 of 5 preference-eligible stimulus vote(s) tied, leaving only 3 discordant preference vote(s). The sign test conditions on non-tie stimulus votes and cannot reach 0.05 below 5, so no record could have passed here — this is not a measured null. Either the agent is inert on these scenarios (make them discriminate) or the eval needs more distinct stimuli to clear the ties

Next action: The signal favors the target, but ties leave too few discordant tasks; inspect ties and predeclare more discriminating breadth.

State: VALID_NO_CHANGE (no_credible_preference_change)

Gate evidence: n=5; 3W/2T/0L; d=3; p=0.125; net +60.0%

Repeated-run reliability (not used by the gate): 5 paired runs (3W/2T/0L).

Weak or warning scenarios:

Scenario Preference gate Net win Δ Pref Runs (W/T/L)
= Advise on project file organization Eligible +0.0% +0.0% 0/1/0
= Route a slow build to performance analysis Eligible +0.0% +0.0% 0/1/0

Illustrative judge evidence:

  • Advise on project file organization: Position-swap inconsistent (forward: tie, reverse: baseline). Defaulting to tie.

This is one example, not the aggregate verdict. Open Full Results for every judgment.

➖ Improvement signal, tie-limited — binlog-failure-analysis (claude-sonnet-5)

Why: Net win +33.3% (3W/2T/1L over 6 preference-eligible stimulus vote(s), sign test p=0.312), mean preference +17.1% across 7 paired run(s), 1 dormancy stimulus/stimuli excluded from preference — not credible — 2 of 6 preference-eligible stimulus vote(s) tied, leaving only 4 discordant preference vote(s). The sign test conditions on non-tie stimulus votes and cannot reach 0.05 below 5, so no record could have passed here — this is not a measured null. Either the skill is inert on these scenarios (make them discriminate) or the eval needs more distinct stimuli to clear the ties

Next action: The signal favors the target, but ties leave too few discordant tasks; inspect ties and predeclare more discriminating breadth.

State: VALID_NO_CHANGE (no_credible_preference_change)

Gate evidence: n=6; 3W/2T/1L; d=4; p=0.312; net +33.3%; 1 dormancy excluded

Warnings: Activation: isolated 6/7; plugin 6/7

Overfit: Moderate (score 0.30)

Repeated-run reliability (not used by the gate): 7 paired runs (4W/2T/1L).

Weak or warning scenarios:

Scenario Preference gate Net win Δ Pref Runs (W/T/L)
= Confirm the actual resolved target framework and package version from a binlog Eligible +0.0% +0.0% 0/0/0
▼ Determine whether a quiet second build actually failed Eligible -100.0% -40.0% 0/0/1
= Diagnose a warning behind a build that actually succeeded Eligible +0.0% +0.0% 0/1/0
= Trace why a generated source file is missing at compile time Eligible +0.0% +0.0% 0/1/0

Illustrative judge evidence:

  • Determine whether a quiet second build actually failed: Both answers are correct, concise, and directly responsive. A is marginally stronger because it supplies a more precise, directly quoted binlog citation (including line numbers) for both success and the CoreCompile up-to-date skip; B's optional rebuild advice is useful but doe...

This is one example, not the aggregate verdict. Open Full Results for every judgment.

➖ Improvement signal, unproven — binlog-failure-analysis (gpt-5.6-luna)

Why: Net win +50.0% (4W/1T/1L over 6 preference-eligible stimulus vote(s), sign test p=0.188), mean preference +14.3% across 7 paired run(s), 1 dormancy stimulus/stimuli excluded from preference — not credible (sign test p=0.188 > 0.05)

Next action: The signal favors the target but is inconsistent; inspect tied or lost scenarios and fix weak behavior.

State: VALID_NO_CHANGE (no_credible_preference_change)

Gate evidence: n=6; 4W/1T/1L; d=5; p=0.188; net +50.0%; 1 dormancy excluded

Warnings: Activation: isolated 6/7; plugin 6/7

Overfit: Moderate (score 0.26)

Repeated-run reliability (not used by the gate): 7 paired runs (5W/1T/1L).

Weak or warning scenarios:

Scenario Preference gate Net win Δ Pref Runs (W/T/L)
▼ Assess a requested binlog investigation when the current build is healthy Eligible -100.0% -100.0% 0/0/1
= Confirm the actual resolved target framework and package version from a binlog Eligible +0.0% +0.0% 0/0/0
= Diagnose build failures from binlog only (no source files) Eligible +0.0% +0.0% 0/1/0

Illustrative judge evidence:

  • Assess a requested binlog investigation when the current build is healthy: Response A correctly identifies that InventorySync builds successfully with 0 errors and 0 warnings, captures binlogs as evidence, and identifies only a reproducibility risk (missing global.json). Response B violates the core rubric by artificially manufacturing a build failur...

This is one example, not the aggregate verdict. Open Full Results for every judgment.

➖ Improvement signal, unproven — binlog-generation (claude-sonnet-5)

Why: Net win +28.6% (4W/1T/2L over 7 preference-eligible stimulus vote(s), sign test p=0.344), mean preference +25.0% across 8 paired run(s), 1 dormancy stimulus/stimuli excluded from preference — not credible (sign test p=0.344 > 0.05)

Next action: The signal favors the target but is inconsistent; inspect tied or lost scenarios and fix weak behavior.

State: VALID_NO_CHANGE (no_credible_preference_change)

Gate evidence: n=7; 4W/1T/2L; d=6; p=0.344; net +28.6%; 1 dormancy excluded

Warnings: Activation: isolated 5/7; plugin 6/7

Overfit: High (score 0.50)

Repeated-run reliability (not used by the gate): 8 paired runs (4W/2T/2L).

Weak or warning scenarios:

Scenario Preference gate Net win Δ Pref Runs (W/T/L)
= Choose a predictable, non-colliding binlog name for a CI upload step Eligible +0.0% +0.0% 0/1/0
= Decline Maven compiler configuration Excluded (activation contract) +0.0% +0.0% 0/1/0
▼ Preserve binlog history while cleaning stale build output Eligible -100.0% -40.0% 0/0/1
▼ Recognize that a failed build produced no binlog at all Eligible -100.0% -40.0% 0/0/1

Illustrative judge evidence:

  • Choose a predictable, non-colliding binlog name for a CI upload step: Position-swap inconsistent (forward: tie, reverse: A). Defaulting to tie.

This is one example, not the aggregate verdict. Open Full Results for every judgment.

➖ Improvement signal, unproven — binlog-generation (gpt-5.6-luna)

Why: Net win +57.1% (5W/1T/1L over 7 preference-eligible stimulus vote(s), sign test p=0.109), mean preference +50.0% across 8 paired run(s), 1 dormancy stimulus/stimuli excluded from preference — not credible (sign test p=0.109 > 0.05)

Next action: The signal favors the target but is inconsistent; inspect tied or lost scenarios and fix weak behavior.

State: VALID_NO_CHANGE (no_credible_preference_change)

Gate evidence: n=7; 5W/1T/1L; d=6; p=0.109; net +57.1%; 1 dormancy excluded

Warnings: Activation: isolated 7/7; plugin 5/7

Overfit: Moderate (score 0.36)

Repeated-run reliability (not used by the gate): 8 paired runs (5W/2T/1L).

Weak or warning scenarios:

Scenario Preference gate Net win Δ Pref Runs (W/T/L)
= Choose a predictable, non-colliding binlog name for a CI upload step Eligible +0.0% +0.0% 0/1/0
= Decline Maven compiler configuration Excluded (activation contract) +0.0% +0.0% 0/1/0
▲ Fix a CI build script that reuses the same binlog on every run Eligible +100.0% +40.0% 1/0/0
▲ Preserve binlog history while cleaning stale build output Eligible +100.0% +100.0% 1/0/0
▼ Recognize that a failed build produced no binlog at all Eligible -100.0% -40.0% 0/0/1

Illustrative judge evidence:

  • Choose a predictable, non-colliding binlog name for a CI upload step: Position-swap inconsistent (forward: B, reverse: A). Defaulting to tie.

This is one example, not the aggregate verdict. Open Full Results for every judgment.

➖ Baseline signal, tie-limited — build-parallelism (claude-sonnet-5)

Why: Net win -16.7% (1W/3T/2L over 6 preference-eligible stimulus vote(s), sign test p=0.500), mean preference +2.9% across 7 paired run(s), 1 dormancy stimulus/stimuli excluded from preference — no improvement

Next action: Inspect activation-only failed runs before rewriting skill content; the model stopped after loading a skill.

State: VALID_NO_CHANGE (no_credible_preference_change)

Gate evidence: n=6; 1W/3T/2L; d=3; p=0.500; net -16.7%; 1 dormancy excluded

Warnings: Activation: isolated 5/6; plugin 4/6; Activation-only stop: isolated 1 failed run

Overfit: Moderate (score 0.36)

Repeated-run reliability (not used by the gate): 7 paired runs (1W/4T/2L).

Weak or warning scenarios:

Scenario Preference gate Net win Δ Pref Runs (W/T/L)
▼ Analyze build parallelism bottlenecks Eligible -100.0% -40.0% 0/0/1
= Decline database query tuning request Excluded (activation contract) +0.0% +0.0% 0/1/0
= Decline graph build for runtime-discovered projects Eligible +0.0% +0.0% 0/1/0
▼ Enable BuildInParallel on a custom MSBuild task Eligible -100.0% -40.0% 0/0/1
= Enable parallel nodes for a wide project graph Eligible +0.0% +0.0% 0/1/0
= Reduce CI build scope with a solution filter Eligible +0.0% +0.0% 0/1/0

Illustrative judge evidence:

  • Analyze build parallelism bottlenecks: The core dependency and redundancy conclusions are equally correct, but A better fulfills the explicit binlog-based analysis requirement with concrete execution evidence.

This is one example, not the aggregate verdict. Open Full Results for every judgment.

➖ Improvement signal, tie-limited — build-parallelism (gpt-5.6-luna)

Why: Net win +66.7% (4W/2T/0L over 6 preference-eligible stimulus vote(s), sign test p=0.063), mean preference +31.4% across 7 paired run(s), 1 dormancy stimulus/stimuli excluded from preference — not credible — 2 of 6 preference-eligible stimulus vote(s) tied, leaving only 4 discordant preference vote(s). The sign test conditions on non-tie stimulus votes and cannot reach 0.05 below 5, so no record could have passed here — this is not a measured null. Either the skill is inert on these scenarios (make them discriminate) or the eval needs more distinct stimuli to clear the ties

Next action: Inspect activation-only failed runs before rewriting skill content; the model stopped after loading a skill.

State: VALID_NO_CHANGE (no_credible_preference_change)

Gate evidence: n=6; 4W/2T/0L; d=4; p=0.063; net +66.7%; 1 dormancy excluded

Warnings: Activation-only stop: isolated 1 failed run; Activation-only stop: plugin 1 failed run

Overfit: Low (score 0.06)

Repeated-run reliability (not used by the gate): 7 paired runs (4W/3T/0L).

Weak or warning scenarios:

Scenario Preference gate Net win Δ Pref Runs (W/T/L)
= Decline database query tuning request Excluded (activation contract) +0.0% +0.0% 0/1/0
= Decline graph build for runtime-discovered projects Eligible +0.0% +0.0% 0/1/0
▲ Enable parallel nodes for a wide project graph Eligible +100.0% +40.0% 1/0/0
= Reduce CI build scope with a solution filter Eligible +0.0% +0.0% 0/1/0

Illustrative judge evidence:

  • Decline graph build for runtime-discovered projects: Position-swap inconsistent (forward: B, reverse: A). Defaulting to tie.

This is one example, not the aggregate verdict. Open Full Results for every judgment.

➖ Mixed evidence — build-perf-baseline (claude-sonnet-5)

Why: Net win +0.0% (3W/0T/3L over 6 preference-eligible stimulus vote(s), sign test p=0.656), mean preference -8.6% across 7 paired run(s), 1 dormancy stimulus/stimuli excluded from preference — no improvement

Next action: Compare winning and losing scenarios to isolate where the target helps versus hurts.

State: VALID_NO_CHANGE (no_credible_preference_change)

Gate evidence: n=6; 3W/0T/3L; d=6; p=0.656; net +0.0%; 1 dormancy excluded

Warnings: Activation: isolated 5/6; plugin 5/6

Overfit: Moderate (score 0.44)

Repeated-run reliability (not used by the gate): 7 paired runs (3W/1T/3L).

Weak or warning scenarios:

Scenario Preference gate Net win Δ Pref Runs (W/T/L)
▼ Configure deterministic, cache-safe CI builds Eligible -100.0% -40.0% 0/0/1
= Decline a non-MSBuild build performance request Excluded (activation contract) +0.0% +0.0% 0/1/0
▼ Leave an already-optimized build unchanged Eligible -100.0% -100.0% 0/0/1
▼ Route a broken no-op rebuild away from generic optimization Eligible -100.0% -40.0% 0/0/1
▲ Route a restore-bound cold build away from architecture changes Eligible +100.0% +40.0% 1/0/0

Illustrative judge evidence:

  • Configure deterministic, cache-safe CI builds: Both make the essential CI-gated ContinuousIntegrationBuild change, and both unnecessarily make Deterministic explicit. A is modestly stronger because it investigated the actual cross-agent discrepancy, adds path normalization, and validates the resulting artifact byte-for-byt...

This is one example, not the aggregate verdict. Open Full Results for every judgment.

➖ Improvement signal, tie-limited — build-perf-baseline (gpt-5.6-luna)

Why: Net win +33.3% (3W/2T/1L over 6 preference-eligible stimulus vote(s), sign test p=0.312), mean preference +22.9% across 7 paired run(s), 1 dormancy stimulus/stimuli excluded from preference — not credible — 2 of 6 preference-eligible stimulus vote(s) tied, leaving only 4 discordant preference vote(s). The sign test conditions on non-tie stimulus votes and cannot reach 0.05 below 5, so no record could have passed here — this is not a measured null. Either the skill is inert on these scenarios (make them discriminate) or the eval needs more distinct stimuli to clear the ties

Next action: The signal favors the target, but ties leave too few discordant tasks; inspect ties and predeclare more discriminating breadth.

State: VALID_NO_CHANGE (no_credible_preference_change)

Gate evidence: n=6; 3W/2T/1L; d=4; p=0.312; net +33.3%; 1 dormancy excluded

Warnings: Activation: isolated 5/6; plugin 4/6

Overfit: Moderate (score 0.24)

Repeated-run reliability (not used by the gate): 7 paired runs (3W/2T/2L).

Weak or warning scenarios:

Scenario Preference gate Net win Δ Pref Runs (W/T/L)
▲ Configure deterministic, cache-safe CI builds Eligible +100.0% +100.0% 1/0/0
▼ Decline a non-MSBuild build performance request Excluded (activation contract) -100.0% -40.0% 0/0/1
▼ Establish build performance baseline and recommend optimizations Eligible -100.0% -100.0% 0/0/1
= Route a broken no-op rebuild away from generic optimization Eligible +0.0% +0.0% 0/1/0
= Route a restore-bound cold build away from architecture changes Eligible +0.0% +0.0% 0/1/0

Illustrative judge evidence:

  • Establish build performance baseline and recommend optimizations: Response A completed the full task by delivering specific, actionable build optimization recommendations grounded in measured baseline data. It identified redundant project references by name, recommended conditionalizing unnecessary build overhead, suggested advanced techniqu...

This is one example, not the aggregate verdict. Open Full Results for every judgment.

➖ Improvement signal, unproven — build-perf-diagnostics (claude-sonnet-5)

Why: Net win +14.3% (3W/2T/2L over 7 preference-eligible stimulus vote(s), sign test p=0.500), mean preference +5.0% across 8 paired run(s), 1 dormancy stimulus/stimuli excluded from preference — not credible (sign test p=0.500 > 0.05)

Next action: The signal favors the target but is inconsistent; inspect tied or lost scenarios and fix weak behavior.

State: VALID_NO_CHANGE (no_credible_preference_change)

Gate evidence: n=7; 3W/2T/2L; d=5; p=0.500; net +14.3%; 1 dormancy excluded

Warnings: Activation: isolated 7/7; plugin 2/7

Overfit: Moderate (score 0.47)

Repeated-run reliability (not used by the gate): 8 paired runs (3W/3T/2L).

Weak or warning scenarios:

Scenario Preference gate Net win Δ Pref Runs (W/T/L)
= Decline a runtime latency request that is not a build performance issue Excluded (activation contract) +0.0% +0.0% 0/1/0
▼ Diagnose a Copy task dominating build time Eligible -100.0% -40.0% 0/0/1
▲ Diagnose a pathological ResolveAssemblyReference time Eligible +100.0% +40.0% 1/0/0
= Diagnose a single custom target dominating one project's build Eligible +0.0% +0.0% 0/1/0
= Diagnose evaluation overhead before any target runs Eligible +0.0% +0.0% 0/1/0
▼ Diagnose per-project overhead across many small projects Eligible -100.0% -40.0% 0/0/1
▲ Diagnose slow build for a small project Eligible +100.0% +40.0% 1/0/0

Illustrative judge evidence:

  • Diagnose a Copy task dominating build time: The central diagnosis and primary PreserveNewest fix are strong in both responses. A is more reliable because it supplies the applicable Always-copy skip property rather than B's dubious SkipCopyUnchangedFiles recommendation. B's more explicit hardlink snippet is useful but ...

This is one example, not the aggregate verdict. Open Full Results for every judgment.

➖ Improvement signal, unproven — build-perf-diagnostics (gpt-5.6-luna)

Why: Net win +42.9% (4W/2T/1L over 7 preference-eligible stimulus vote(s), sign test p=0.188), mean preference +15.0% across 8 paired run(s), 1 dormancy stimulus/stimuli excluded from preference — not credible (sign test p=0.188 > 0.05)

Next action: The signal favors the target but is inconsistent; inspect tied or lost scenarios and fix weak behavior.

State: VALID_NO_CHANGE (no_credible_preference_change)

Gate evidence: n=7; 4W/2T/1L; d=5; p=0.188; net +42.9%; 1 dormancy excluded

Warnings: Activation: isolated 7/7; plugin 3/7

Overfit: Moderate (score 0.43)

Repeated-run reliability (not used by the gate): 8 paired runs (4W/3T/1L).

Weak or warning scenarios:

Scenario Preference gate Net win Δ Pref Runs (W/T/L)
= Decline a runtime latency request that is not a build performance issue Excluded (activation contract) +0.0% +0.0% 0/1/0
= Diagnose NuGet restore running redundantly across CI stages Eligible +0.0% +0.0% 0/1/0
▼ Diagnose a single custom target dominating one project's build Eligible -100.0% -40.0% 0/0/1
▲ Diagnose evaluation overhead before any target runs Eligible +100.0% +40.0% 1/0/0
= Diagnose per-project overhead across many small projects Eligible +0.0% +0.0% 0/1/0

Illustrative judge evidence:

  • Diagnose NuGet restore running redundantly across CI stages: Position-swap inconsistent (forward: tie, reverse: A). Defaulting to tie.

This is one example, not the aggregate verdict. Open Full Results for every judgment.

➖ Improvement signal, tie-limited — check-bin-obj-clash (claude-sonnet-5)

Why: Net win +42.9% (3W/4T/0L over 7 preference-eligible stimulus vote(s), sign test p=0.125), mean preference +25.7% across 7 paired run(s) — not credible — 4 of 7 preference-eligible stimulus vote(s) tied, leaving only 3 discordant preference vote(s). The sign test conditions on non-tie stimulus votes and cannot reach 0.05 below 5, so no record could have passed here — this is not a measured null. Either the skill is inert on these scenarios (make them discriminate) or the eval needs more distinct stimuli to clear the ties

Next action: The signal favors the target, but ties leave too few discordant tasks; inspect ties and predeclare more discriminating breadth.

State: VALID_NO_CHANGE (no_credible_preference_change)

Gate evidence: n=7; 3W/4T/0L; d=3; p=0.125; net +42.9%

Warnings: Activation: isolated 7/7; plugin 5/7

Overfit: Low (score 0.16)

Repeated-run reliability (not used by the gate): 7 paired runs (3W/4T/0L).

Weak or warning scenarios:

Scenario Preference gate Net win Δ Pref Runs (W/T/L)
= Audit the mixed solution and separate safe projects from unsafe ones Eligible +0.0% +0.0% 0/1/0
= Decline output-clash remediation for separate projects using the default SDK layout Eligible +0.0% +0.0% 0/1/0
= Diagnose multi-targeting outputs that collapse into one path Eligible +0.0% +0.0% 0/1/0
= Diagnose redundant project reference metadata that forks a same-path build Eligible +0.0% +0.0% 0/1/0

Illustrative judge evidence:

  • Audit the mixed solution and separate safe projects from unsafe ones: Position-swap inconsistent (forward: A, reverse: B). Defaulting to tie.

This is one example, not the aggregate verdict. Open Full Results for every judgment.

➖ Improvement signal, tie-limited — check-bin-obj-clash (gpt-5.6-luna)

Why: Net win +14.3% (2W/4T/1L over 7 preference-eligible stimulus vote(s), sign test p=0.500), mean preference +5.7% across 7 paired run(s) — not credible — 4 of 7 preference-eligible stimulus vote(s) tied, leaving only 3 discordant preference vote(s). The sign test conditions on non-tie stimulus votes and cannot reach 0.05 below 5, so no record could have passed here — this is not a measured null. Either the skill is inert on these scenarios (make them discriminate) or the eval needs more distinct stimuli to clear the ties

Next action: The signal favors the target, but ties leave too few discordant tasks; inspect ties and predeclare more discriminating breadth.

State: VALID_NO_CHANGE (no_credible_preference_change)

Gate evidence: n=7; 2W/4T/1L; d=3; p=0.500; net +14.3%

Warnings: Activation: isolated 7/7; plugin 6/7

Overfit: Moderate (score 0.31)

Repeated-run reliability (not used by the gate): 7 paired runs (2W/4T/1L).

Weak or warning scenarios:

Scenario Preference gate Net win Δ Pref Runs (W/T/L)
= Audit the mixed solution and separate safe projects from unsafe ones Eligible +0.0% +0.0% 0/1/0
= Decline output-clash remediation for separate projects using the default SDK layout Eligible +0.0% +0.0% 0/1/0
▲ Diagnose multi-targeting outputs that collapse into one path Eligible +100.0% +40.0% 1/0/0
= Diagnose redundant project reference metadata that forks a same-path build Eligible +0.0% +0.0% 0/1/0
= Diagnose shared output and intermediate path collision Eligible +0.0% +0.0% 0/1/0
▼ Fix all clash mechanisms in the mixed solution Eligible -100.0% -40.0% 0/0/1

Illustrative judge evidence:

  • Audit the mixed solution and separate safe projects from unsafe ones: Position-swap inconsistent (forward: B, reverse: A). Defaulting to tie.

This is one example, not the aggregate verdict. Open Full Results for every judgment.

➖ Baseline signal, tie-limited — directory-build-organization (gpt-5.6-luna)

Why: Net win -16.7% (1W/3T/2L over 6 preference-eligible stimulus vote(s), sign test p=0.500), mean preference -5.7% across 7 paired run(s), 1 dormancy stimulus/stimuli excluded from preference — no improvement

Next action: Evidence leans baseline but is not credible; inspect losses and make tied tasks discriminate.

State: VALID_NO_CHANGE (no_credible_preference_change)

Gate evidence: n=6; 1W/3T/2L; d=3; p=0.500; net -16.7%; 1 dormancy excluded

Warnings: Activation: isolated 6/6; plugin 4/6

Overfit: Moderate (score 0.31)

Repeated-run reliability (not used by the gate): 7 paired runs (1W/4T/2L).

Weak or warning scenarios:

Scenario Preference gate Net win Δ Pref Runs (W/T/L)
= Apply repo-level build organization cleanup Eligible +0.0% +0.0% 0/1/0
= Decline adding shared build files to a lone project Excluded (activation contract) +0.0% +0.0% 0/1/0
= Diagnose a TargetFramework condition that silently skips in props Eligible +0.0% +0.0% 0/1/0
= Diagnose a package downgrade chain and reorganize version management Eligible +0.0% +0.0% 0/1/0
▼ Diagnose an inner shared-props file that overwrites its own override Eligible -100.0% -40.0% 0/0/1
▼ Preserve an intentional project-specific exception while centralizing Eligible -100.0% -40.0% 0/0/1

Illustrative judge evidence:

  • Apply repo-level build organization cleanup: Position-swap inconsistent (forward: tie, reverse: A). Defaulting to tie.

This is one example, not the aggregate verdict. Open Full Results for every judgment.

➖ Improvement signal, unproven — eval-performance (claude-sonnet-5)

Why: Net win +50.0% (5W/2T/1L over 8 preference-eligible stimulus vote(s), sign test p=0.109), mean preference +20.0% across 8 paired run(s) — not credible (sign test p=0.109 > 0.05)

Next action: The signal favors the target but is inconsistent; inspect tied or lost scenarios and fix weak behavior.

State: VALID_NO_CHANGE (no_credible_preference_change)

Gate evidence: n=8; 5W/2T/1L; d=6; p=0.109; net +50.0%

Warnings: Activation: isolated 5/8; plugin 5/8

Overfit: Moderate (score 0.27)

Repeated-run reliability (not used by the gate): 8 paired runs (5W/2T/1L).

Weak or warning scenarios:

Scenario Preference gate Net win Δ Pref Runs (W/T/L)
= Decline to invent evaluation problems in an already-clean project Eligible +0.0% +0.0% 0/1/0
▼ Diagnose stacked evaluation-time patterns (deep imports, broad glob, file-I/O property function) Eligible -100.0% -40.0% 0/0/1
▲ Recognize TreatAsLocalProperty overuse versus one justified entry Eligible +100.0% +40.0% 1/0/0
▲ Redirect an incremental-rebuild complaint mistakenly framed as an evaluation problem Eligible +100.0% +40.0% 1/0/0
= Triage which of two property functions actually costs evaluation time Eligible +0.0% +0.0% 0/1/0

Illustrative judge evidence:

  • Decline to invent evaluation problems in an already-clean project: Position-swap inconsistent (forward: B, reverse: A). Defaulting to tie.

This is one example, not the aggregate verdict. Open Full Results for every judgment.

➖ Improvement signal, unproven — eval-performance (gpt-5.6-luna)

Why: Net win +50.0% (5W/2T/1L over 8 preference-eligible stimulus vote(s), sign test p=0.109), mean preference +35.0% across 8 paired run(s) — not credible (sign test p=0.109 > 0.05)

Next action: The signal favors the target but is inconsistent; inspect tied or lost scenarios and fix weak behavior.

State: VALID_NO_CHANGE (no_credible_preference_change)

Gate evidence: n=8; 5W/2T/1L; d=6; p=0.109; net +50.0%

Warnings: Activation: isolated 8/8; plugin 7/8

Overfit: Moderate (score 0.38)

Repeated-run reliability (not used by the gate): 8 paired runs (5W/2T/1L).

Weak or warning scenarios:

Scenario Preference gate Net win Δ Pref Runs (W/T/L)
▼ Diagnose stacked evaluation-time patterns (deep imports, broad glob, file-I/O property function) Eligible -100.0% -40.0% 0/0/1
▲ Recognize TreatAsLocalProperty overuse versus one justified entry Eligible +100.0% +40.0% 1/0/0
= Redirect a compile-time slowdown mistakenly framed as an evaluation problem Eligible +0.0% +0.0% 0/1/0
= Triage which of two property functions actually costs evaluation time Eligible +0.0% +0.0% 0/1/0

Illustrative judge evidence:

  • Diagnose stacked evaluation-time patterns (deep imports, broad glob, file-I/O property function): Response A better addresses the rubric criteria, particularly in providing specific fix recommendations (restricting globs, moving I/O to execution phase) and explaining evaluation overhead. While Response B takes a more empirical approach with measurements showing 305 ms eval...

This is one example, not the aggregate verdict. Open Full Results for every judgment.

➖ Improvement signal, tie-limited — extension-points (claude-sonnet-5)

Why: Net win +28.6% (3W/3T/1L over 7 preference-eligible stimulus vote(s), sign test p=0.312), mean preference +15.6% across 9 paired run(s), 2 dormancy stimulus/stimuli excluded from preference — not credible — 3 of 7 preference-eligible stimulus vote(s) tied, leaving only 4 discordant preference vote(s). The sign test conditions on non-tie stimulus votes and cannot reach 0.05 below 5, so no record could have passed here — this is not a measured null. Either the skill is inert on these scenarios (make them discriminate) or the eval needs more distinct stimuli to clear the ties

Next action: The signal favors the target, but ties leave too few discordant tasks; inspect ties and predeclare more discriminating breadth.

State: VALID_NO_CHANGE (no_credible_preference_change)

Gate evidence: n=7; 3W/3T/1L; d=4; p=0.312; net +28.6%; 2 dormancy excluded

Warnings: Activation: isolated 6/7; plugin 4/7

Overfit: Moderate (score 0.27)

Repeated-run reliability (not used by the gate): 9 paired runs (3W/5T/1L).

Weak or warning scenarios:

Scenario Preference gate Net win Δ Pref Runs (W/T/L)
▲ Create extensibility hooks for a custom SDK target file Eligible +100.0% +40.0% 1/0/0
= Decline Kubernetes readiness probe configuration Excluded (activation contract) +0.0% +0.0% 0/1/0
= Diagnose NuGet package and repo extension conflicts Eligible +0.0% +0.0% 0/1/0
▲ Diagnose a broken per-TFM forwarder Eligible +100.0% +100.0% 1/0/0
= Diagnose a package ID and file-name mismatch Eligible +0.0% +0.0% 0/1/0
▼ Diagnose build extension point failures Eligible -100.0% -40.0% 0/0/1
▲ Fix extension point anti-patterns Eligible +100.0% +40.0% 1/0/0
= Leave incremental target authoring to its owning workflow Excluded (activation contract) +0.0% +0.0% 0/1/0
= Review packed layout without a false missing-file bug Eligible +0.0% +0.0% 0/1/0

Illustrative judge evidence:

  • Diagnose NuGet package and repo extension conflicts: Position-swap inconsistent (forward: B, reverse: A). Defaulting to tie.

This is one example, not the aggregate verdict. Open Full Results for every judgment.

➖ Improvement signal, tie-limited — extension-points (gpt-5.6-luna)

Why: Net win +14.3% (2W/4T/1L over 7 preference-eligible stimulus vote(s), sign test p=0.500), mean preference +4.4% across 9 paired run(s), 2 dormancy stimulus/stimuli excluded from preference — not credible — 4 of 7 preference-eligible stimulus vote(s) tied, leaving only 3 discordant preference vote(s). The sign test conditions on non-tie stimulus votes and cannot reach 0.05 below 5, so no record could have passed here — this is not a measured null. Either the skill is inert on these scenarios (make them discriminate) or the eval needs more distinct stimuli to clear the ties

Next action: The signal favors the target, but ties leave too few discordant tasks; inspect ties and predeclare more discriminating breadth.

State: VALID_NO_CHANGE (no_credible_preference_change)

Gate evidence: n=7; 2W/4T/1L; d=3; p=0.500; net +14.3%; 2 dormancy excluded

Warnings: Activation: isolated 7/7; plugin 5/7

Overfit: Moderate (score 0.39)

Repeated-run reliability (not used by the gate): 9 paired runs (3W/4T/2L).

Weak or warning scenarios:

Scenario Preference gate Net win Δ Pref Runs (W/T/L)
▲ Create extensibility hooks for a custom SDK target file Eligible +100.0% +40.0% 1/0/0
= Diagnose NuGet package and repo extension conflicts Eligible +0.0% +0.0% 0/1/0
= Diagnose a package ID and file-name mismatch Eligible +0.0% +0.0% 0/1/0
= Diagnose build extension point failures Eligible +0.0% +0.0% 0/1/0
▼ Fix extension point anti-patterns Eligible -100.0% -40.0% 0/0/1
▼ Leave incremental target authoring to its owning workflow Excluded (activation contract) -100.0% -40.0% 0/0/1
= Review packed layout without a false missing-file bug Eligible +0.0% +0.0% 0/1/0

Illustrative judge evidence:

  • Diagnose NuGet package and repo extension conflicts: Position-swap inconsistent (forward: B, reverse: A). Defaulting to tie.

This is one example, not the aggregate verdict. Open Full Results for every judgment.

➖ Improvement signal, tie-limited — including-generated-files (claude-sonnet-5)

Why: Net win +14.3% (2W/4T/1L over 7 preference-eligible stimulus vote(s), sign test p=0.500), mean preference +0.0% across 8 paired run(s), 1 dormancy stimulus/stimuli excluded from preference — not credible — 4 of 7 preference-eligible stimulus vote(s) tied, leaving only 3 discordant preference vote(s). The sign test conditions on non-tie stimulus votes and cannot reach 0.05 below 5, so no record could have passed here — this is not a measured null. Either the skill is inert on these scenarios (make them discriminate) or the eval needs more distinct stimuli to clear the ties

Next action: The signal favors the target, but ties leave too few discordant tasks; inspect ties and predeclare more discriminating breadth.

State: VALID_NO_CHANGE (no_credible_preference_change)

Gate evidence: n=7; 2W/4T/1L; d=3; p=0.500; net +14.3%; 1 dormancy excluded

Warnings: Activation: isolated 4/7; plugin 4/7

Overfit: Moderate (score 0.35)

Repeated-run reliability (not used by the gate): 8 paired runs (2W/4T/2L).

Weak or warning scenarios:

Scenario Preference gate Net win Δ Pref Runs (W/T/L)
▼ Decline a Python runtime report-path problem Excluded (activation contract) -100.0% -40.0% 0/0/1
▲ Diagnose hardcoded obj path for generated source Eligible +100.0% +40.0% 1/0/0
▲ Diagnose missing clean tracking for generated source Eligible +100.0% +40.0% 1/0/0
= Diagnose missing generated source inclusion Eligible +0.0% +0.0% 0/1/0
= Diagnose missing output registration for generated non-code file Eligible +0.0% +0.0% 0/1/0
▼ Diagnose project-level glob for generated source Eligible -100.0% -40.0% 0/0/1
= Diagnose wrong hook for generated source files Eligible +0.0% +0.0% 0/1/0
= Fix generated source inclusion and clean tracking Eligible +0.0% +0.0% 0/1/0

Illustrative judge evidence:

  • Diagnose missing generated source inclusion: Position-swap inconsistent (forward: B, reverse: tie). Defaulting to tie.

This is one example, not the aggregate verdict. Open Full Results for every judgment.

➖ Mixed evidence — including-generated-files (gpt-5.6-luna)

Why: Net win +0.0% (3W/1T/3L over 7 preference-eligible stimulus vote(s), sign test p=0.656), mean preference +0.0% across 8 paired run(s), 1 dormancy stimulus/stimuli excluded from preference — no improvement

Next action: Compare winning and losing scenarios to isolate where the target helps versus hurts.

State: VALID_NO_CHANGE (no_credible_preference_change)

Gate evidence: n=7; 3W/1T/3L; d=6; p=0.656; net +0.0%; 1 dormancy excluded

Overfit: Moderate (score 0.40)

Repeated-run reliability (not used by the gate): 8 paired runs (3W/2T/3L).

Weak or warning scenarios:

Scenario Preference gate Net win Δ Pref Runs (W/T/L)
= Decline a Python runtime report-path problem Excluded (activation contract) +0.0% +0.0% 0/1/0
▼ Diagnose missing clean tracking for generated source Eligible -100.0% -40.0% 0/0/1
▼ Diagnose missing output registration for generated non-code file Eligible -100.0% -40.0% 0/0/1
= Diagnose project-level glob for generated source Eligible +0.0% +0.0% 0/1/0
▼ Diagnose wrong hook for generated source files Eligible -100.0% -40.0% 0/0/1

Illustrative judge evidence:

  • Diagnose missing clean tracking for generated source: Both responses provide the correct solution and explanation. However, Response A offers slightly greater technical depth by explaining the MSBuild mechanism more explicitly (_CleanRecordFileWrites, the file list persistence location, and how clean reads that list). While Respo...

This is one example, not the aggregate verdict. Open Full Results for every judgment.

➖ Improvement signal, tie-limited — incremental-build (claude-sonnet-5)

Why: Net win +22.2% (3W/5T/1L over 9 preference-eligible stimulus vote(s), sign test p=0.312), mean preference +8.9% across 9 paired run(s) — not credible — 5 of 9 preference-eligible stimulus vote(s) tied, leaving only 4 discordant preference vote(s). The sign test conditions on non-tie stimulus votes and cannot reach 0.05 below 5, so no record could have passed here — this is not a measured null. Either the skill is inert on these scenarios (make them discriminate) or the eval needs more distinct stimuli to clear the ties

Next action: The signal favors the target, but ties leave too few discordant tasks; inspect ties and predeclare more discriminating breadth.

State: VALID_NO_CHANGE (no_credible_preference_change)

Gate evidence: n=9; 3W/5T/1L; d=4; p=0.312; net +22.2%

Warnings: Activation: isolated 3/9; plugin 5/9

Overfit: Moderate (score 0.24)

Repeated-run reliability (not used by the gate): 9 paired runs (3W/5T/1L).

Weak or warning scenarios:

Scenario Preference gate Net win Δ Pref Runs (W/T/L)
▲ Correct the assumption that Outputs alone enables incremental skipping Eligible +100.0% +40.0% 1/0/0
▼ Diagnose custom targets that always rerun Eligible -100.0% -40.0% 0/0/1
= Distinguish a cold first build from broken incrementality Eligible +0.0% +0.0% 0/1/0
= Explain slower builds when MSBuild skipped everything Eligible +0.0% +0.0% 0/1/0
= Explain why Visual Studio keeps rebuilding an up-to-date project Eligible +0.0% +0.0% 0/1/0
▲ Explain why clean leaves generated hash source behind Eligible +100.0% +40.0% 1/0/0
= Identify a volatile output path that defeats incrementality Eligible +0.0% +0.0% 0/1/0
= Read a diagnostic log to find the stale input Eligible +0.0% +0.0% 0/1/0

Illustrative judge evidence:

  • Diagnose custom targets that always rerun: Both answer the requested diagnosis and fix correctly. A is marginally better because it avoids B's unnecessary and potentially incorrect FileWrites claim.

This is one example, not the aggregate verdict. Open Full Results for every judgment.

➖ Improvement signal, unproven — incremental-build (gpt-5.6-luna)

Why: Net win +33.3% (4W/4T/1L over 9 preference-eligible stimulus vote(s), sign test p=0.188), mean preference +13.3% across 9 paired run(s) — not credible (sign test p=0.188 > 0.05)

Next action: The signal favors the target but is inconsistent; inspect tied or lost scenarios and fix weak behavior.

State: VALID_NO_CHANGE (no_credible_preference_change)

Gate evidence: n=9; 4W/4T/1L; d=5; p=0.188; net +33.3%

Warnings: Activation: isolated 8/9; plugin 6/9

Overfit: Moderate (score 0.22)

Repeated-run reliability (not used by the gate): 9 paired runs (4W/4T/1L).

Weak or warning scenarios:

Scenario Preference gate Net win Δ Pref Runs (W/T/L)
▲ Correct the assumption that Outputs alone enables incremental skipping Eligible +100.0% +40.0% 1/0/0
= Distinguish a cold first build from broken incrementality Eligible +0.0% +0.0% 0/1/0
= Explain slower builds when MSBuild skipped everything Eligible +0.0% +0.0% 0/1/0
= Explain why clean leaves generated hash source behind Eligible +0.0% +0.0% 0/1/0
= Identify a volatile output path that defeats incrementality Eligible +0.0% +0.0% 0/1/0
▼ Read a diagnostic log to find the stale input Eligible -100.0% -40.0% 0/0/1

Illustrative judge evidence:

  • Distinguish a cold first build from broken incrementality: Position-swap inconsistent (forward: A, reverse: B). Defaulting to tie.

This is one example, not the aggregate verdict. Open Full Results for every judgment.

➖ Mixed evidence — item-management (claude-sonnet-5)

Why: Net win +0.0% (2W/2T/2L over 6 preference-eligible stimulus vote(s), sign test p=0.687), mean preference +5.7% across 7 paired run(s), 1 dormancy stimulus/stimuli excluded from preference — no improvement

Next action: Compare winning and losing scenarios to isolate where the target helps versus hurts.

State: VALID_NO_CHANGE (no_credible_preference_change)

Gate evidence: n=6; 2W/2T/2L; d=4; p=0.687; net +0.0%; 1 dormancy excluded

Warnings: Activation: isolated 4/6; plugin 3/6

Overfit: Moderate (score 0.33)

Repeated-run reliability (not used by the gate): 7 paired runs (3W/2T/2L).

Weak or warning scenarios:

Scenario Preference gate Net win Δ Pref Runs (W/T/L)
= Diagnose an ineffective Compile Remove that does not match the glob Eligible +0.0% +0.0% 0/1/0
▲ Diagnose real and claimed item problems in a code generation pipeline Eligible +100.0% +40.0% 1/0/0
▼ Fix item management anti-patterns Eligible -100.0% -100.0% 0/0/1
= Leave already-correct item management unchanged Eligible +0.0% +0.0% 0/1/0
▼ Leave correct single-list batching unchanged Eligible -100.0% -40.0% 0/0/1

Illustrative judge evidence:

  • Diagnose an ineffective Compile Remove that does not match the glob: Position-swap inconsistent (forward: tie, reverse: A). Defaulting to tie.

This is one example, not the aggregate verdict. Open Full Results for every judgment.

➖ Mixed evidence — item-management (gpt-5.6-luna)

Why: Net win +0.0% (1W/4T/1L over 6 preference-eligible stimulus vote(s), sign test p=0.750), mean preference +8.6% across 7 paired run(s), 1 dormancy stimulus/stimuli excluded from preference — no improvement

Next action: Compare winning and losing scenarios to isolate where the target helps versus hurts.

State: VALID_NO_CHANGE (no_credible_preference_change)

Gate evidence: n=6; 1W/4T/1L; d=2; p=0.750; net +0.0%; 1 dormancy excluded

Warnings: Activation: isolated 6/6; plugin 5/6

Overfit: Moderate (score 0.30)

Repeated-run reliability (not used by the gate): 7 paired runs (1W/5T/1L).

Weak or warning scenarios:

Scenario Preference gate Net win Δ Pref Runs (W/T/L)
= Decline JavaScript configuration merging Excluded (activation contract) +0.0% +0.0% 0/1/0
▼ Diagnose an ineffective Compile Remove that does not match the glob Eligible -100.0% -40.0% 0/0/1
= Diagnose item group and batching issues Eligible +0.0% +0.0% 0/1/0
= Diagnose real and claimed item problems in a code generation pipeline Eligible +0.0% +0.0% 0/1/0
= Fix item management anti-patterns Eligible +0.0% +0.0% 0/1/0
= Leave correct single-list batching unchanged Eligible +0.0% +0.0% 0/1/0

Illustrative judge evidence:

  • Diagnose an ineffective Compile Remove that does not match the glob: Both responses correctly identify the core problem (Remove pattern mismatch) and both provide working solutions. However, Response A's solution is architecturally superior: it targets *.g.cs specifically, matching the actual naming convention of the generated files, whereas ...

This is one example, not the aggregate verdict. Open Full Results for every judgment.

➖ Mixed evidence — msbuild-antipatterns (claude-sonnet-5)

Why: Net win +0.0% (2W/3T/2L over 7 preference-eligible stimulus vote(s), sign test p=0.687), mean preference +5.0% across 8 paired run(s), 1 dormancy stimulus/stimuli excluded from preference — no improvement

Next action: Compare winning and losing scenarios to isolate where the target helps versus hurts.

State: VALID_NO_CHANGE (no_credible_preference_change)

Gate evidence: n=7; 2W/3T/2L; d=4; p=0.687; net +0.0%; 1 dormancy excluded

Warnings: Activation: isolated 3/7; plugin 2/7

Overfit: Moderate (score 0.30)

Repeated-run reliability (not used by the gate): 8 paired runs (3W/3T/2L).

Weak or warning scenarios:

Scenario Preference gate Net win Δ Pref Runs (W/T/L)
= Add a module to an F# project Eligible +0.0% +0.0% 0/1/0
= Add a signature file to define public API Eligible +0.0% +0.0% 0/1/0
▼ Distinguish a style backslash from a real cross-platform backslash bug Eligible -100.0% -40.0% 0/0/1
= Fix broken file order causing FS0039 Eligible +0.0% +0.0% 0/1/0
▲ Judge an unguarded import inside a NuGet package build folder Eligible +100.0% +40.0% 1/0/0
▼ Leave a clean project without inventing anti-patterns Eligible -100.0% -40.0% 0/0/1

Illustrative judge evidence:

  • Add a module to an F# project: The implementations and verification outcomes are materially equivalent: both fulfill the requested validation, project wiring, pre-processing control flow, and successful compilation/run.

This is one example, not the aggregate verdict. Open Full Results for every judgment.

➖ Improvement signal, unproven — property-patterns (gpt-5.6-luna)

Why: Net win +14.3% (3W/2T/2L over 7 preference-eligible stimulus vote(s), sign test p=0.500), mean preference -2.9% across 7 paired run(s) — not credible (sign test p=0.500 > 0.05)

Next action: The signal favors the target but is inconsistent; inspect tied or lost scenarios and fix weak behavior.

State: VALID_NO_CHANGE (no_credible_preference_change)

Gate evidence: n=7; 3W/2T/2L; d=5; p=0.500; net +14.3%

Warnings: Activation: isolated 6/7; plugin 6/7

Overfit: Moderate (score 0.37)

Repeated-run reliability (not used by the gate): 7 paired runs (3W/2T/2L).

Weak or warning scenarios:

Scenario Preference gate Net win Δ Pref Runs (W/T/L)
▲ Choose the correct shared build file for a post-build target Eligible +100.0% +40.0% 1/0/0
= Diagnose multi-level property hierarchy bugs Eligible +0.0% +0.0% 0/1/0
▼ Diagnose shared build property issues Eligible -100.0% -40.0% 0/0/1
= Fix shared property configuration Eligible +0.0% +0.0% 0/1/0
▼ Leave already-correct shared properties unchanged Eligible -100.0% -100.0% 0/0/1

Illustrative judge evidence:

  • Diagnose multi-level property hierarchy bugs: Position-swap inconsistent (forward: B, reverse: A). Defaulting to tie.

This is one example, not the aggregate verdict. Open Full Results for every judgment.

Details for 3 results were omitted to keep this comment under GitHub's limit. Open the workflow summary or Full Results for the complete breakdown.

🔍 Full Results - all metrics and investigation details

To investigate non-passing or warning results, paste this to your AI coding agent:

For PR 1215 in dotnet/skills, download eval artifacts with gh run download 37303425464 --repo dotnet/skills --pattern "vally-results-*" --dir ./eval-results, then fetch https://raw.githubusercontent.com/dotnet/skills/b97f9c542191718f6b4d15e24c1be7af9908b9e8/eng/vally-adapter/InvestigatingResults.md and follow it. Classify each result as measurement-invalid, underpowered, not-proven, preference-loss, or passing-with-warning. Use stateReason, result accounting, weak scenarios, and judge evidence to give the cause and exact next fix.

⚠️ Session replay telemetry was not published because the auxiliary dotnet/skills-data publisher failed. The evaluation verdicts above remain authoritative; maintainers must repair SKILLS_DATA_TOKEN or the publisher before replay links are available.

@github-actions

github-actions Bot commented Oct 5, 2026

Copy link
Copy Markdown
Contributor

📊 Skill and Agent Evaluation Results

30 model/target results across 15 targets and 2 models — ✅ 5 improved, ➖ 25 results without a clear winner, ⚠️ 0 invalid or underpowered, ⛔ 0 activation contract failures, 📉 0 preference losses (report only).

Measurement identity: evaluated commit b9052fd631b50ceda0f8f4d3033bc3b02fb97671; 2 judge models.

Measurement health: 30 expected / 30 observed / 30 written; 0 missing, 0 unexpected, 0 invalid; 0 recovered comparison error slots and 0 unresolved comparison error slots.

Objective completion gate: not enabled. Aggregate completion transitions are telemetry only, so this report does not claim that zero objective regressions were proven.

A result passes only when preference-eligible distinct-stimulus votes have aggregate net win of at least 20% and an exact one-sided sign-test result of p ≤ 0.05, and every explicit dormancy activation contract passes. Repeated runs measure reliability only.

Target Model Verdict Gate evidence Overfit Warnings Next action
agent.msbuild claude-sonnet-5 ➖ Improvement signal, tie-limited n=5; 2W/3T/0L; d=2; p=0.250; net +40.0% — — The signal favors the target, but ties leave too few discordant tasks; inspect ties and predeclare more discriminating breadth.
agent.msbuild gpt-5.6-luna ➖ Improvement signal, tie-limited n=5; 3W/1T/1L; d=4; p=0.312; net +40.0% — — The signal favors the target, but ties leave too few discordant tasks; inspect ties and predeclare more discriminating breadth.
binlog-failure-analysis claude-sonnet-5 ➖ Improvement signal, tie-limited n=6; 2W/3T/1L; d=3; p=0.500; net +16.7%; 1 dormancy excluded 🟡 0.42 Activation: isolated 6/7; plugin 6/7 The signal favors the target, but ties leave too few discordant tasks; inspect ties and predeclare more discriminating breadth.
binlog-failure-analysis gpt-5.6-luna ➖ Improvement signal, tie-limited n=6; 3W/3T/0L; d=3; p=0.125; net +50.0%; 1 dormancy excluded 🟡 0.35 Activation: isolated 6/7; plugin 6/7 The signal favors the target, but ties leave too few discordant tasks; inspect ties and predeclare more discriminating breadth.
binlog-generation claude-sonnet-5 ➖ Improvement signal, unproven n=7; 5W/1T/1L; d=6; p=0.109; net +57.1%; 1 dormancy excluded 🟡 0.48 Activation: isolated 6/7; plugin 5/7 The signal favors the target but is inconsistent; inspect tied or lost scenarios and fix weak behavior.
binlog-generation gpt-5.6-luna ✅ Improved n=7; 5W/2T/0L; d=5; p=0.031; net +71.4%; 1 dormancy excluded 🟡 0.26 Activation: isolated 6/7; plugin 6/7 Fix activation gaps; Review overfit evidence.
build-parallelism claude-sonnet-5 ➖ Mixed evidence n=6; 2W/2T/2L; d=4; p=0.687; net +0.0%; 1 dormancy excluded 🟡 0.47 Activation: isolated 6/6; plugin 4/6 Compare winning and losing scenarios to isolate where the target helps versus hurts.
build-parallelism gpt-5.6-luna ➖ Baseline signal, tie-limited n=6; 1W/2T/3L; d=4; p=0.312; net -33.3%; 1 dormancy excluded 🟡 0.22 Activation-only stop: isolated 1 failed run; Activation-only stop: plugin 1 failed run Inspect activation-only failed runs before rewriting skill content; the model stopped after loading a skill.
build-perf-baseline claude-sonnet-5 ➖ Baseline signal, tie-limited n=6; 1W/3T/2L; d=3; p=0.500; net -16.7%; 1 dormancy excluded 🔴 0.51 Activation: isolated 6/6; plugin 5/6 Evidence leans baseline but is not credible; inspect losses and make tied tasks discriminate.
build-perf-baseline gpt-5.6-luna ➖ Improvement signal, tie-limited n=6; 4W/2T/0L; d=4; p=0.063; net +66.7%; 1 dormancy excluded 🟡 0.28 Activation: isolated 6/6; plugin 3/6 The signal favors the target, but ties leave too few discordant tasks; inspect ties and predeclare more discriminating breadth.
build-perf-diagnostics claude-sonnet-5 ➖ Improvement signal, tie-limited n=7; 2W/4T/1L; d=3; p=0.500; net +14.3%; 1 dormancy excluded 🟡 0.35 Activation: isolated 7/7; plugin 3/7 The signal favors the target, but ties leave too few discordant tasks; inspect ties and predeclare more discriminating breadth.
build-perf-diagnostics gpt-5.6-luna ✅ Improved n=7; 5W/2T/0L; d=5; p=0.031; net +71.4%; 1 dormancy excluded 🟡 0.45 Activation: isolated 7/7; plugin 3/7 Fix activation gaps; Review overfit evidence.
check-bin-obj-clash claude-sonnet-5 ➖ Baseline signal, unproven n=7; 1W/2T/4L; d=5; p=0.188; net -42.9% ✅ 0.17 Activation: isolated 7/7; plugin 4/7 Evidence leans baseline but is not credible; inspect losing scenarios for recurring defects.
check-bin-obj-clash gpt-5.6-luna ➖ Improvement signal, unproven n=7; 4W/2T/1L; d=5; p=0.188; net +42.9% 🟡 0.30 — The signal favors the target but is inconsistent; inspect tied or lost scenarios and fix weak behavior.
directory-build-organization claude-sonnet-5 ➖ Improvement signal, unproven n=6; 3W/1T/2L; d=5; p=0.500; net +16.7%; 1 dormancy excluded 🟡 0.43 Activation: isolated 3/6; plugin 4/6 The signal favors the target but is inconsistent; inspect tied or lost scenarios and fix weak behavior.
directory-build-organization gpt-5.6-luna ➖ Mixed evidence n=6; 2W/2T/2L; d=4; p=0.687; net +0.0%; 1 dormancy excluded 🔴 0.51 Activation: isolated 6/6; plugin 4/6 Compare winning and losing scenarios to isolate where the target helps versus hurts.
eval-performance claude-sonnet-5 ✅ Improved n=8; 5W/3T/0L; d=5; p=0.031; net +62.5% 🟡 0.32 Activation: isolated 5/8; plugin 5/8 Fix activation gaps; Review overfit evidence.
eval-performance gpt-5.6-luna ✅ Improved n=8; 5W/3T/0L; d=5; p=0.031; net +62.5% 🟡 0.32 — Review overfit evidence.
extension-points claude-sonnet-5 ➖ Improvement signal, unproven n=7; 4W/1T/2L; d=6; p=0.344; net +28.6%; 2 dormancy excluded 🟡 0.23 Activation: isolated 5/7; plugin 7/7 The signal favors the target but is inconsistent; inspect tied or lost scenarios and fix weak behavior.
extension-points gpt-5.6-luna ➖ Improvement signal, tie-limited n=7; 2W/4T/1L; d=3; p=0.500; net +14.3%; 2 dormancy excluded 🟡 0.47 Activation: isolated 7/7; plugin 6/7 The signal favors the target, but ties leave too few discordant tasks; inspect ties and predeclare more discriminating breadth.
including-generated-files claude-sonnet-5 ➖ Improvement signal, tie-limited n=7; 3W/3T/1L; d=4; p=0.312; net +28.6%; 1 dormancy excluded 🟡 0.39 Activation: isolated 5/7; plugin 4/7 The signal favors the target, but ties leave too few discordant tasks; inspect ties and predeclare more discriminating breadth.
including-generated-files gpt-5.6-luna ➖ Improvement signal, tie-limited n=7; 2W/4T/1L; d=3; p=0.500; net +14.3%; 1 dormancy excluded 🔴 0.53 Activation: isolated 7/7; plugin 6/7 The signal favors the target, but ties leave too few discordant tasks; inspect ties and predeclare more discriminating breadth.
incremental-build claude-sonnet-5 ✅ Improved n=9; 6W/3T/0L; d=6; p=0.016; net +66.7% 🟡 0.35 Activation: isolated 4/9; plugin 4/9 Fix activation gaps; Review overfit evidence.
incremental-build gpt-5.6-luna ➖ Improvement signal, tie-limited n=9; 4W/5T/0L; d=4; p=0.063; net +44.4% 🟡 0.21 Activation: isolated 8/9; plugin 7/9 The signal favors the target, but ties leave too few discordant tasks; inspect ties and predeclare more discriminating breadth.
item-management claude-sonnet-5 ➖ Improvement signal, tie-limited n=6; 3W/3T/0L; d=3; p=0.125; net +50.0%; 1 dormancy excluded 🟡 0.29 Activation: isolated 5/6; plugin 4/6 The signal favors the target, but ties leave too few discordant tasks; inspect ties and predeclare more discriminating breadth.
item-management gpt-5.6-luna ➖ Improvement signal, tie-limited n=6; 4W/2T/0L; d=4; p=0.063; net +66.7%; 1 dormancy excluded 🟡 0.41 — The signal favors the target, but ties leave too few discordant tasks; inspect ties and predeclare more discriminating breadth.
msbuild-antipatterns claude-sonnet-5 ➖ Improvement signal, tie-limited n=7; 1W/6T/0L; d=1; p=0.500; net +14.3%; 1 dormancy excluded 🟡 0.31 Activation: isolated 3/7; plugin 3/7 The signal favors the target, but ties leave too few discordant tasks; inspect ties and predeclare more discriminating breadth.
msbuild-antipatterns gpt-5.6-luna ➖ Improvement signal, tie-limited n=7; 3W/4T/0L; d=3; p=0.125; net +42.9%; 1 dormancy excluded 🟡 0.49 Activation: isolated 5/7; plugin 4/7 The signal favors the target, but ties leave too few discordant tasks; inspect ties and predeclare more discriminating breadth.
property-patterns claude-sonnet-5 ➖ Baseline signal, tie-limited n=7; 0W/4T/3L; d=3; p=0.125; net -42.9% ✅ 0.17 Activation: isolated 5/7; plugin 6/7 Evidence leans baseline but is not credible; inspect losses and make tied tasks discriminate.
property-patterns gpt-5.6-luna ➖ Mixed evidence n=7; 2W/3T/2L; d=4; p=0.687; net +0.0% 🟡 0.23 Activation: isolated 6/7; plugin 6/7 Compare winning and losing scenarios to isolate where the target helps versus hurts.
ℹ️ How to read this report
  • ✅ Improved — the result passed both the statistical gate and the 20% practical net-win floor.
  • ➖ No clear winner — the result is valid but did not pass both gates. The label distinguishes all ties, mixed evidence, directional but unproven evidence, and credible effects below the practical floor.
  • ⚠️ Invalid / underpowered — the gate withheld a quality verdict. Fix the measurement before judging the target.
  • ⛔ Activation contract failed — the isolated target activated on an explicit dormancy scenario. Dormancy preference is excluded, but this routing failure still blocks a pass.
  • 📉 Preference loss — the LLM judge credibly preferred baseline. It is report-only, not objective completion proof.
  • Gate evidence — n preference-eligible distinct-stimulus votes, W/T/L stimulus votes, d discordant votes, exact one-sided p, net win, and the count of separately retained dormancy stimuli. The p value applies to one model/target result; no matrix-wide multiple-comparison correction is applied.
  • Overfit — overfitting-judge severity (✅ Low, 🟡 Moderate, 🔴 High, — none) and score.
  • Warnings — activation, timeout, retry recovery, or unresolved comparison conditions that need attention.
  • Do not add repeated runs to increase statistical power. Do not add stimuli after seeing a near-pass unless the new breadth is predeclared for a new experiment.
➖ Improvement signal, tie-limited — agent.msbuild (claude-sonnet-5)

Why: Net win +40.0% (2W/3T/0L over 5 preference-eligible stimulus vote(s), sign test p=0.250), mean preference +28.0% across 5 paired run(s) — not credible — 3 of 5 preference-eligible stimulus vote(s) tied, leaving only 2 discordant preference vote(s). The sign test conditions on non-tie stimulus votes and cannot reach 0.05 below 5, so no record could have passed here — this is not a measured null. Either the agent is inert on these scenarios (make them discriminate) or the eval needs more distinct stimuli to clear the ties

Next action: The signal favors the target, but ties leave too few discordant tasks; inspect ties and predeclare more discriminating breadth.

State: VALID_NO_CHANGE (no_credible_preference_change)

Gate evidence: n=5; 2W/3T/0L; d=2; p=0.250; net +40.0%

Repeated-run reliability (not used by the gate): 5 paired runs (2W/3T/0L).

Weak or warning scenarios:

Scenario Preference gate Net win Δ Pref Runs (W/T/L)
= Advise on project file organization Eligible +0.0% +0.0% 0/1/0
= Route a slow build to performance analysis Eligible +0.0% +0.0% 0/1/0
= Triage a build failure and route to appropriate analysis Eligible +0.0% +0.0% 0/1/0

Illustrative judge evidence:

  • Advise on project file organization: Position-swap inconsistent (forward: baseline, reverse: skill). Defaulting to tie.

This is one example, not the aggregate verdict. Open Full Results for every judgment.

➖ Improvement signal, tie-limited — agent.msbuild (gpt-5.6-luna)

Why: Net win +40.0% (3W/1T/1L over 5 preference-eligible stimulus vote(s), sign test p=0.312), mean preference +16.0% across 5 paired run(s) — not credible — 1 of 5 preference-eligible stimulus vote(s) tied, leaving only 4 discordant preference vote(s). The sign test conditions on non-tie stimulus votes and cannot reach 0.05 below 5, so no record could have passed here — this is not a measured null. Either the agent is inert on these scenarios (make them discriminate) or the eval needs more distinct stimuli to clear the ties

Next action: The signal favors the target, but ties leave too few discordant tasks; inspect ties and predeclare more discriminating breadth.

State: VALID_NO_CHANGE (no_credible_preference_change)

Gate evidence: n=5; 3W/1T/1L; d=4; p=0.312; net +40.0%

Repeated-run reliability (not used by the gate): 5 paired runs (3W/1T/1L).

Weak or warning scenarios:

Scenario Preference gate Net win Δ Pref Runs (W/T/L)
= Diagnose broken incremental build behavior Eligible +0.0% +0.0% 0/1/0
▼ Route a slow build to performance analysis Eligible -100.0% -40.0% 0/0/1

Illustrative judge evidence:

  • Diagnose broken incremental build behavior: Position-swap inconsistent (forward: skill, reverse: baseline). Defaulting to tie.

This is one example, not the aggregate verdict. Open Full Results for every judgment.

➖ Improvement signal, tie-limited — binlog-failure-analysis (claude-sonnet-5)

Why: Net win +16.7% (2W/3T/1L over 6 preference-eligible stimulus vote(s), sign test p=0.500), mean preference +11.4% across 7 paired run(s), 1 dormancy stimulus/stimuli excluded from preference — not credible — 3 of 6 preference-eligible stimulus vote(s) tied, leaving only 3 discordant preference vote(s). The sign test conditions on non-tie stimulus votes and cannot reach 0.05 below 5, so no record could have passed here — this is not a measured null. Either the skill is inert on these scenarios (make them discriminate) or the eval needs more distinct stimuli to clear the ties

Next action: The signal favors the target, but ties leave too few discordant tasks; inspect ties and predeclare more discriminating breadth.

State: VALID_NO_CHANGE (no_credible_preference_change)

Gate evidence: n=6; 2W/3T/1L; d=3; p=0.500; net +16.7%; 1 dormancy excluded

Warnings: Activation: isolated 6/7; plugin 6/7

Overfit: Moderate (score 0.42)

Repeated-run reliability (not used by the gate): 7 paired runs (3W/3T/1L).

Weak or warning scenarios:

Scenario Preference gate Net win Δ Pref Runs (W/T/L)
= Assess a requested binlog investigation when the current build is healthy Eligible +0.0% +0.0% 0/1/0
= Confirm the actual resolved target framework and package version from a binlog Eligible +0.0% +0.0% 0/0/0
= Determine whether a quiet second build actually failed Eligible +0.0% +0.0% 0/1/0
▼ Diagnose build failures from binlog only (no source files) Eligible -100.0% -40.0% 0/0/1
= Trace why a generated source file is missing at compile time Eligible +0.0% +0.0% 0/1/0

Illustrative judge evidence:

  • Assess a requested binlog investigation when the current build is healthy: Position-swap inconsistent (forward: B, reverse: tie). Defaulting to tie.

This is one example, not the aggregate verdict. Open Full Results for every judgment.

➖ Improvement signal, tie-limited — binlog-failure-analysis (gpt-5.6-luna)

Why: Net win +50.0% (3W/3T/0L over 6 preference-eligible stimulus vote(s), sign test p=0.125), mean preference +22.9% across 7 paired run(s), 1 dormancy stimulus/stimuli excluded from preference — not credible — 3 of 6 preference-eligible stimulus vote(s) tied, leaving only 3 discordant preference vote(s). The sign test conditions on non-tie stimulus votes and cannot reach 0.05 below 5, so no record could have passed here — this is not a measured null. Either the skill is inert on these scenarios (make them discriminate) or the eval needs more distinct stimuli to clear the ties

Next action: The signal favors the target, but ties leave too few discordant tasks; inspect ties and predeclare more discriminating breadth.

State: VALID_NO_CHANGE (no_credible_preference_change)

Gate evidence: n=6; 3W/3T/0L; d=3; p=0.125; net +50.0%; 1 dormancy excluded

Warnings: Activation: isolated 6/7; plugin 6/7

Overfit: Moderate (score 0.35)

Repeated-run reliability (not used by the gate): 7 paired runs (4W/3T/0L).

Weak or warning scenarios:

Scenario Preference gate Net win Δ Pref Runs (W/T/L)
= Assess a requested binlog investigation when the current build is healthy Eligible +0.0% +0.0% 0/1/0
= Confirm the actual resolved target framework and package version from a binlog Eligible +0.0% +0.0% 0/0/0
= Determine whether a quiet second build actually failed Eligible +0.0% +0.0% 0/1/0
= Trace why a generated source file is missing at compile time Eligible +0.0% +0.0% 0/1/0

Illustrative judge evidence:

  • Assess a requested binlog investigation when the current build is healthy: Position-swap inconsistent (forward: B, reverse: A). Defaulting to tie.

This is one example, not the aggregate verdict. Open Full Results for every judgment.

➖ Improvement signal, unproven — binlog-generation (claude-sonnet-5)

Why: Net win +57.1% (5W/1T/1L over 7 preference-eligible stimulus vote(s), sign test p=0.109), mean preference +42.5% across 8 paired run(s), 1 dormancy stimulus/stimuli excluded from preference — not credible (sign test p=0.109 > 0.05)

Next action: The signal favors the target but is inconsistent; inspect tied or lost scenarios and fix weak behavior.

State: VALID_NO_CHANGE (no_credible_preference_change)

Gate evidence: n=7; 5W/1T/1L; d=6; p=0.109; net +57.1%; 1 dormancy excluded

Warnings: Activation: isolated 6/7; plugin 5/7

Overfit: Moderate (score 0.48)

Repeated-run reliability (not used by the gate): 8 paired runs (5W/2T/1L).

Weak or warning scenarios:

Scenario Preference gate Net win Δ Pref Runs (W/T/L)
= Choose a predictable, non-colliding binlog name for a CI upload step Eligible +0.0% +0.0% 0/1/0
= Decline Maven compiler configuration Excluded (activation contract) +0.0% +0.0% 0/1/0
▲ Preserve binlog history while cleaning stale build output Eligible +100.0% +40.0% 1/0/0
▼ Recognize that a failed build produced no binlog at all Eligible -100.0% -40.0% 0/0/1

Illustrative judge evidence:

  • Choose a predictable, non-colliding binlog name for a CI upload step: Both responses fully satisfy the task and every rubric requirement, reporting the same valid, explicitly named non-colliding Debug and Release binlogs after successful builds.

This is one example, not the aggregate verdict. Open Full Results for every judgment.

➖ Mixed evidence — build-parallelism (claude-sonnet-5)

Why: Net win +0.0% (2W/2T/2L over 6 preference-eligible stimulus vote(s), sign test p=0.687), mean preference +8.6% across 7 paired run(s), 1 dormancy stimulus/stimuli excluded from preference — no improvement

Next action: Compare winning and losing scenarios to isolate where the target helps versus hurts.

State: VALID_NO_CHANGE (no_credible_preference_change)

Gate evidence: n=6; 2W/2T/2L; d=4; p=0.687; net +0.0%; 1 dormancy excluded

Warnings: Activation: isolated 6/6; plugin 4/6

Overfit: Moderate (score 0.47)

Repeated-run reliability (not used by the gate): 7 paired runs (2W/3T/2L).

Weak or warning scenarios:

Scenario Preference gate Net win Δ Pref Runs (W/T/L)
▼ Analyze build parallelism bottlenecks Eligible -100.0% -40.0% 0/0/1
= Decline database query tuning request Excluded (activation contract) +0.0% +0.0% 0/1/0
= Enable BuildInParallel on a custom MSBuild task Eligible +0.0% +0.0% 0/1/0
= Enable parallel nodes for a wide project graph Eligible +0.0% +0.0% 0/1/0
▼ Reduce CI build scope with a solution filter Eligible -100.0% -40.0% 0/0/1

Illustrative judge evidence:

  • Analyze build parallelism bottlenecks: The substantive conclusions are the same and correct, but A better connects its conclusion to the requested binlog evidence and more explicitly explains why parallel workers cannot break the serial chain.

This is one example, not the aggregate verdict. Open Full Results for every judgment.

➖ Baseline signal, tie-limited — build-parallelism (gpt-5.6-luna)

Why: Net win -33.3% (1W/2T/3L over 6 preference-eligible stimulus vote(s), sign test p=0.312), mean preference -17.1% across 7 paired run(s), 1 dormancy stimulus/stimuli excluded from preference — no improvement

Next action: Inspect activation-only failed runs before rewriting skill content; the model stopped after loading a skill.

State: VALID_NO_CHANGE (no_credible_preference_change)

Gate evidence: n=6; 1W/2T/3L; d=4; p=0.312; net -33.3%; 1 dormancy excluded

Warnings: Activation-only stop: isolated 1 failed run; Activation-only stop: plugin 1 failed run

Overfit: Moderate (score 0.22)

Repeated-run reliability (not used by the gate): 7 paired runs (1W/2T/4L).

Weak or warning scenarios:

Scenario Preference gate Net win Δ Pref Runs (W/T/L)
▼ Analyze build parallelism bottlenecks Eligible -100.0% -40.0% 0/0/1
▼ Decline database query tuning request Excluded (activation contract) -100.0% -40.0% 0/0/1
= Decline graph build for runtime-discovered projects Eligible +0.0% +0.0% 0/1/0
= Enable BuildInParallel on a custom MSBuild task Eligible +0.0% +0.0% 0/1/0
▼ Enable parallel nodes for a wide project graph Eligible -100.0% -40.0% 0/0/1
▼ Reduce CI build scope with a solution filter Eligible -100.0% -40.0% 0/0/1

Illustrative judge evidence:

  • Analyze build parallelism bottlenecks: Both responses correctly analyze the build parallelism and identify the same core findings: the Core → Api → Web → Tests critical path limits parallelism, and Tests → Api is a redundant but non-critical-path reference. Response A edges ahead by providing quantitative evidence—...

This is one example, not the aggregate verdict. Open Full Results for every judgment.

➖ Baseline signal, tie-limited — build-perf-baseline (claude-sonnet-5)

Why: Net win -16.7% (1W/3T/2L over 6 preference-eligible stimulus vote(s), sign test p=0.500), mean preference -5.7% across 7 paired run(s), 1 dormancy stimulus/stimuli excluded from preference — no improvement

Next action: Evidence leans baseline but is not credible; inspect losses and make tied tasks discriminate.

State: VALID_NO_CHANGE (no_credible_preference_change)

Gate evidence: n=6; 1W/3T/2L; d=3; p=0.500; net -16.7%; 1 dormancy excluded

Warnings: Activation: isolated 6/6; plugin 5/6

Overfit: High (score 0.51)

Repeated-run reliability (not used by the gate): 7 paired runs (1W/4T/2L).

Weak or warning scenarios:

Scenario Preference gate Net win Δ Pref Runs (W/T/L)
= Adopt UseArtifactsOutput while declining graph build for a small solution Eligible +0.0% +0.0% 0/1/0
▼ Configure deterministic, cache-safe CI builds Eligible -100.0% -40.0% 0/0/1
= Decline a non-MSBuild build performance request Excluded (activation contract) +0.0% +0.0% 0/1/0
= Establish build performance baseline and recommend optimizations Eligible +0.0% +0.0% 0/1/0
▼ Route a broken no-op rebuild away from generic optimization Eligible -100.0% -40.0% 0/0/1
= Route a restore-bound cold build away from architecture changes Eligible +0.0% +0.0% 0/1/0

Illustrative judge evidence:

  • Adopt UseArtifactsOutput while declining graph build for a small solution: Position-swap inconsistent (forward: B, reverse: A). Defaulting to tie.

This is one example, not the aggregate verdict. Open Full Results for every judgment.

➖ Improvement signal, tie-limited — build-perf-baseline (gpt-5.6-luna)

Why: Net win +66.7% (4W/2T/0L over 6 preference-eligible stimulus vote(s), sign test p=0.063), mean preference +25.7% across 7 paired run(s), 1 dormancy stimulus/stimuli excluded from preference — not credible — 2 of 6 preference-eligible stimulus vote(s) tied, leaving only 4 discordant preference vote(s). The sign test conditions on non-tie stimulus votes and cannot reach 0.05 below 5, so no record could have passed here — this is not a measured null. Either the skill is inert on these scenarios (make them discriminate) or the eval needs more distinct stimuli to clear the ties

Next action: The signal favors the target, but ties leave too few discordant tasks; inspect ties and predeclare more discriminating breadth.

State: VALID_NO_CHANGE (no_credible_preference_change)

Gate evidence: n=6; 4W/2T/0L; d=4; p=0.063; net +66.7%; 1 dormancy excluded

Warnings: Activation: isolated 6/6; plugin 3/6

Overfit: Moderate (score 0.28)

Repeated-run reliability (not used by the gate): 7 paired runs (4W/2T/1L).

Weak or warning scenarios:

Scenario Preference gate Net win Δ Pref Runs (W/T/L)
▲ Adopt UseArtifactsOutput while declining graph build for a small solution Eligible +100.0% +40.0% 1/0/0
▲ Configure deterministic, cache-safe CI builds Eligible +100.0% +100.0% 1/0/0
▼ Decline a non-MSBuild build performance request Excluded (activation contract) -100.0% -40.0% 0/0/1
= Establish build performance baseline and recommend optimizations Eligible +0.0% +0.0% 0/1/0
▲ Route a broken no-op rebuild away from generic optimization Eligible +100.0% +40.0% 1/0/0
= Route a restore-bound cold build away from architecture changes Eligible +0.0% +0.0% 0/1/0

Illustrative judge evidence:

  • Establish build performance baseline and recommend optimizations: Position-swap inconsistent (forward: A, reverse: B). Defaulting to tie.

This is one example, not the aggregate verdict. Open Full Results for every judgment.

➖ Improvement signal, tie-limited — build-perf-diagnostics (claude-sonnet-5)

Why: Net win +14.3% (2W/4T/1L over 7 preference-eligible stimulus vote(s), sign test p=0.500), mean preference +5.0% across 8 paired run(s), 1 dormancy stimulus/stimuli excluded from preference — not credible — 4 of 7 preference-eligible stimulus vote(s) tied, leaving only 3 discordant preference vote(s). The sign test conditions on non-tie stimulus votes and cannot reach 0.05 below 5, so no record could have passed here — this is not a measured null. Either the skill is inert on these scenarios (make them discriminate) or the eval needs more distinct stimuli to clear the ties

Next action: The signal favors the target, but ties leave too few discordant tasks; inspect ties and predeclare more discriminating breadth.

State: VALID_NO_CHANGE (no_credible_preference_change)

Gate evidence: n=7; 2W/4T/1L; d=3; p=0.500; net +14.3%; 1 dormancy excluded

Warnings: Activation: isolated 7/7; plugin 3/7

Overfit: Moderate (score 0.35)

Repeated-run reliability (not used by the gate): 8 paired runs (2W/5T/1L).

Weak or warning scenarios:

Scenario Preference gate Net win Δ Pref Runs (W/T/L)
= Decline a runtime latency request that is not a build performance issue Excluded (activation contract) +0.0% +0.0% 0/1/0
= Diagnose NuGet restore running redundantly across CI stages Eligible +0.0% +0.0% 0/1/0
= Diagnose a Copy task dominating build time Eligible +0.0% +0.0% 0/1/0
▼ Diagnose a pathological ResolveAssemblyReference time Eligible -100.0% -40.0% 0/0/1
▲ Diagnose a single custom target dominating one project's build Eligible +100.0% +40.0% 1/0/0
= Diagnose evaluation overhead before any target runs Eligible +0.0% +0.0% 0/1/0
= Diagnose per-project overhead across many small projects Eligible +0.0% +0.0% 0/1/0
▲ Diagnose slow build for a small project Eligible +100.0% +40.0% 1/0/0

Illustrative judge evidence:

  • Diagnose NuGet restore running redundantly across CI stages: Position-swap inconsistent (forward: tie, reverse: A). Defaulting to tie.

This is one example, not the aggregate verdict. Open Full Results for every judgment.

➖ Baseline signal, unproven — check-bin-obj-clash (claude-sonnet-5)

Why: Net win -42.9% (1W/2T/4L over 7 preference-eligible stimulus vote(s), sign test p=0.188), mean preference -25.7% across 7 paired run(s) — no improvement

Next action: Evidence leans baseline but is not credible; inspect losing scenarios for recurring defects.

State: VALID_NO_CHANGE (no_credible_preference_change)

Gate evidence: n=7; 1W/2T/4L; d=5; p=0.188; net -42.9%

Warnings: Activation: isolated 7/7; plugin 4/7

Overfit: Low (score 0.17)

Repeated-run reliability (not used by the gate): 7 paired runs (1W/2T/4L).

Weak or warning scenarios:

Scenario Preference gate Net win Δ Pref Runs (W/T/L)
▼ Audit the mixed solution and separate safe projects from unsafe ones Eligible -100.0% -100.0% 0/0/1
▼ Avoid a false clash report when projects share only the top-level artifacts root Eligible -100.0% -40.0% 0/0/1
▼ Decline output-clash remediation for separate projects using the default SDK layout Eligible -100.0% -40.0% 0/0/1
= Diagnose multi-targeting outputs that collapse into one path Eligible +0.0% +0.0% 0/1/0
▲ Diagnose redundant project reference metadata that forks a same-path build Eligible +100.0% +40.0% 1/0/0
▼ Diagnose shared output and intermediate path collision Eligible -100.0% -40.0% 0/0/1
= Fix all clash mechanisms in the mixed solution Eligible +0.0% +0.0% 0/1/0

Illustrative judge evidence:

  • Audit the mixed solution and separate safe projects from unsafe ones: A gets the central MultiTargetLib diagnosis right as well as the LibraryA/LibraryB collision and identifies the ConsumerApp reference issue. B's erroneous declaration that MultiTargetLib is safe misses a major required unsafe case, outweighing its slightly clearer unsafe label...

This is one example, not the aggregate verdict. Open Full Results for every judgment.

➖ Improvement signal, unproven — check-bin-obj-clash (gpt-5.6-luna)

Why: Net win +42.9% (4W/2T/1L over 7 preference-eligible stimulus vote(s), sign test p=0.188), mean preference +17.1% across 7 paired run(s) — not credible (sign test p=0.188 > 0.05)

Next action: The signal favors the target but is inconsistent; inspect tied or lost scenarios and fix weak behavior.

State: VALID_NO_CHANGE (no_credible_preference_change)

Gate evidence: n=7; 4W/2T/1L; d=5; p=0.188; net +42.9%

Overfit: Moderate (score 0.30)

Repeated-run reliability (not used by the gate): 7 paired runs (4W/2T/1L).

Weak or warning scenarios:

Scenario Preference gate Net win Δ Pref Runs (W/T/L)
▼ Audit the mixed solution and separate safe projects from unsafe ones Eligible -100.0% -40.0% 0/0/1
= Diagnose redundant project reference metadata that forks a same-path build Eligible +0.0% +0.0% 0/1/0
= Diagnose shared output and intermediate path collision Eligible +0.0% +0.0% 0/1/0

Illustrative judge evidence:

  • Audit the mixed solution and separate safe projects from unsafe ones: Response A provides a more nuanced and clearly-structured analysis. Its key advantage is treating the ConsumerApp→ToolLib reference issue as a separate unsafe item, while explicitly confirming that ToolLib's own project configuration is safe. This distinction between inherent ...

This is one example, not the aggregate verdict. Open Full Results for every judgment.

➖ Improvement signal, unproven — directory-build-organization (claude-sonnet-5)

Why: Net win +16.7% (3W/1T/2L over 6 preference-eligible stimulus vote(s), sign test p=0.500), mean preference +2.9% across 7 paired run(s), 1 dormancy stimulus/stimuli excluded from preference — not credible (sign test p=0.500 > 0.05)

Next action: The signal favors the target but is inconsistent; inspect tied or lost scenarios and fix weak behavior.

State: VALID_NO_CHANGE (no_credible_preference_change)

Gate evidence: n=6; 3W/1T/2L; d=5; p=0.500; net +16.7%; 1 dormancy excluded

Warnings: Activation: isolated 3/6; plugin 4/6

Overfit: Moderate (score 0.43)

Repeated-run reliability (not used by the gate): 7 paired runs (4W/1T/2L).

Weak or warning scenarios:

Scenario Preference gate Net win Δ Pref Runs (W/T/L)
= Audit a chaotic multi-project repo Eligible +0.0% +0.0% 0/1/0
▲ Diagnose a TargetFramework condition that silently skips in props Eligible +100.0% +40.0% 1/0/0
▼ Diagnose a package downgrade chain and reorganize version management Eligible -100.0% -100.0% 0/0/1
▼ Diagnose an inner shared-props file that overwrites its own override Eligible -100.0% -40.0% 0/0/1

Illustrative judge evidence:

  • Audit a chaotic multi-project repo: Position-swap inconsistent (forward: B, reverse: A). Defaulting to tie.

This is one example, not the aggregate verdict. Open Full Results for every judgment.

➖ Mixed evidence — directory-build-organization (gpt-5.6-luna)

Why: Net win +0.0% (2W/2T/2L over 6 preference-eligible stimulus vote(s), sign test p=0.687), mean preference +8.6% across 7 paired run(s), 1 dormancy stimulus/stimuli excluded from preference — no improvement

Next action: Compare winning and losing scenarios to isolate where the target helps versus hurts.

State: VALID_NO_CHANGE (no_credible_preference_change)

Gate evidence: n=6; 2W/2T/2L; d=4; p=0.687; net +0.0%; 1 dormancy excluded

Warnings: Activation: isolated 6/6; plugin 4/6

Overfit: High (score 0.51)

Repeated-run reliability (not used by the gate): 7 paired runs (2W/3T/2L).

Weak or warning scenarios:

Scenario Preference gate Net win Δ Pref Runs (W/T/L)
= Apply repo-level build organization cleanup Eligible +0.0% +0.0% 0/1/0
= Decline adding shared build files to a lone project Excluded (activation contract) +0.0% +0.0% 0/1/0
▼ Diagnose a TargetFramework condition that silently skips in props Eligible -100.0% -40.0% 0/0/1
▼ Diagnose a package downgrade chain and reorganize version management Eligible -100.0% -40.0% 0/0/1
▲ Diagnose an inner shared-props file that overwrites its own override Eligible +100.0% +40.0% 1/0/0
= Preserve an intentional project-specific exception while centralizing Eligible +0.0% +0.0% 0/1/0

Illustrative judge evidence:

  • Apply repo-level build organization cleanup: Position-swap inconsistent (forward: B, reverse: A). Defaulting to tie.

This is one example, not the aggregate verdict. Open Full Results for every judgment.

➖ Improvement signal, unproven — extension-points (claude-sonnet-5)

Why: Net win +28.6% (4W/1T/2L over 7 preference-eligible stimulus vote(s), sign test p=0.344), mean preference +2.2% across 9 paired run(s), 2 dormancy stimulus/stimuli excluded from preference — not credible (sign test p=0.344 > 0.05)

Next action: The signal favors the target but is inconsistent; inspect tied or lost scenarios and fix weak behavior.

State: VALID_NO_CHANGE (no_credible_preference_change)

Gate evidence: n=7; 4W/1T/2L; d=6; p=0.344; net +28.6%; 2 dormancy excluded

Warnings: Activation: isolated 5/7; plugin 7/7

Overfit: Moderate (score 0.23)

Repeated-run reliability (not used by the gate): 9 paired runs (4W/3T/2L).

Weak or warning scenarios:

Scenario Preference gate Net win Δ Pref Runs (W/T/L)
= Decline Kubernetes readiness probe configuration Excluded (activation contract) +0.0% +0.0% 0/1/0
= Diagnose NuGet package and repo extension conflicts Eligible +0.0% +0.0% 0/1/0
▲ Diagnose a broken per-TFM forwarder Eligible +100.0% +40.0% 1/0/0
▲ Diagnose a package ID and file-name mismatch Eligible +100.0% +40.0% 1/0/0
▼ Fix extension point anti-patterns Eligible -100.0% -100.0% 0/0/1
= Leave incremental target authoring to its owning workflow Excluded (activation contract) +0.0% +0.0% 0/1/0
▼ Review packed layout without a false missing-file bug Eligible -100.0% -40.0% 0/0/1

Illustrative judge evidence:

  • Diagnose NuGet package and repo extension conflicts: Position-swap inconsistent (forward: B, reverse: A). Defaulting to tie.

This is one example, not the aggregate verdict. Open Full Results for every judgment.

➖ Improvement signal, tie-limited — extension-points (gpt-5.6-luna)

Why: Net win +14.3% (2W/4T/1L over 7 preference-eligible stimulus vote(s), sign test p=0.500), mean preference +6.7% across 9 paired run(s), 2 dormancy stimulus/stimuli excluded from preference — not credible — 4 of 7 preference-eligible stimulus vote(s) tied, leaving only 3 discordant preference vote(s). The sign test conditions on non-tie stimulus votes and cannot reach 0.05 below 5, so no record could have passed here — this is not a measured null. Either the skill is inert on these scenarios (make them discriminate) or the eval needs more distinct stimuli to clear the ties

Next action: The signal favors the target, but ties leave too few discordant tasks; inspect ties and predeclare more discriminating breadth.

State: VALID_NO_CHANGE (no_credible_preference_change)

Gate evidence: n=7; 2W/4T/1L; d=3; p=0.500; net +14.3%; 2 dormancy excluded

Warnings: Activation: isolated 7/7; plugin 6/7

Overfit: Moderate (score 0.47)

Repeated-run reliability (not used by the gate): 9 paired runs (2W/5T/2L).

Weak or warning scenarios:

Scenario Preference gate Net win Δ Pref Runs (W/T/L)
▼ Create extensibility hooks for a custom SDK target file Eligible -100.0% -40.0% 0/0/1
= Decline Kubernetes readiness probe configuration Excluded (activation contract) +0.0% +0.0% 0/1/0
= Diagnose NuGet package and repo extension conflicts Eligible +0.0% +0.0% 0/1/0
= Diagnose build extension point failures Eligible +0.0% +0.0% 0/1/0
= Fix extension point anti-patterns Eligible +0.0% +0.0% 0/1/0
▼ Leave incremental target authoring to its owning workflow Excluded (activation contract) -100.0% -40.0% 0/0/1
= Review packed layout without a false missing-file bug Eligible +0.0% +0.0% 0/1/0

Illustrative judge evidence:

  • Create extensibility hooks for a custom SDK target file: Both responses successfully completed the task and verified working solutions. Response A is the better solution because it follows the principle of minimal, elegant design—it splits the monolithic target into CoreMySDKBuild and MySDKBuild, enabling the existing BeforeTargets/...

This is one example, not the aggregate verdict. Open Full Results for every judgment.

➖ Improvement signal, tie-limited — including-generated-files (claude-sonnet-5)

Why: Net win +28.6% (3W/3T/1L over 7 preference-eligible stimulus vote(s), sign test p=0.312), mean preference +10.0% across 8 paired run(s), 1 dormancy stimulus/stimuli excluded from preference — not credible — 3 of 7 preference-eligible stimulus vote(s) tied, leaving only 4 discordant preference vote(s). The sign test conditions on non-tie stimulus votes and cannot reach 0.05 below 5, so no record could have passed here — this is not a measured null. Either the skill is inert on these scenarios (make them discriminate) or the eval needs more distinct stimuli to clear the ties

Next action: The signal favors the target, but ties leave too few discordant tasks; inspect ties and predeclare more discriminating breadth.

State: VALID_NO_CHANGE (no_credible_preference_change)

Gate evidence: n=7; 3W/3T/1L; d=4; p=0.312; net +28.6%; 1 dormancy excluded

Warnings: Activation: isolated 5/7; plugin 4/7

Overfit: Moderate (score 0.39)

Repeated-run reliability (not used by the gate): 8 paired runs (3W/4T/1L).

Weak or warning scenarios:

Scenario Preference gate Net win Δ Pref Runs (W/T/L)
= Decline a Python runtime report-path problem Excluded (activation contract) +0.0% +0.0% 0/1/0
= Diagnose hardcoded obj path for generated source Eligible +0.0% +0.0% 0/1/0
= Diagnose missing clean tracking for generated source Eligible +0.0% +0.0% 0/1/0
▼ Diagnose missing generated source inclusion Eligible -100.0% -40.0% 0/0/1
▲ Diagnose project-level glob for generated source Eligible +100.0% +40.0% 1/0/0
= Diagnose wrong hook for generated source files Eligible +0.0% +0.0% 0/1/0

Illustrative judge evidence:

  • Diagnose hardcoded obj path for generated source: Position-swap inconsistent (forward: A, reverse: tie). Defaulting to tie.

This is one example, not the aggregate verdict. Open Full Results for every judgment.

➖ Improvement signal, tie-limited — including-generated-files (gpt-5.6-luna)

Why: Net win +14.3% (2W/4T/1L over 7 preference-eligible stimulus vote(s), sign test p=0.500), mean preference +12.5% across 8 paired run(s), 1 dormancy stimulus/stimuli excluded from preference — not credible — 4 of 7 preference-eligible stimulus vote(s) tied, leaving only 3 discordant preference vote(s). The sign test conditions on non-tie stimulus votes and cannot reach 0.05 below 5, so no record could have passed here — this is not a measured null. Either the skill is inert on these scenarios (make them discriminate) or the eval needs more distinct stimuli to clear the ties

Next action: The signal favors the target, but ties leave too few discordant tasks; inspect ties and predeclare more discriminating breadth.

State: VALID_NO_CHANGE (no_credible_preference_change)

Gate evidence: n=7; 2W/4T/1L; d=3; p=0.500; net +14.3%; 1 dormancy excluded

Warnings: Activation: isolated 7/7; plugin 6/7

Overfit: High (score 0.53)

Repeated-run reliability (not used by the gate): 8 paired runs (2W/5T/1L).

Weak or warning scenarios:

Scenario Preference gate Net win Δ Pref Runs (W/T/L)
= Decline a Python runtime report-path problem Excluded (activation contract) +0.0% +0.0% 0/1/0
▲ Diagnose hardcoded obj path for generated source Eligible +100.0% +40.0% 1/0/0
= Diagnose missing clean tracking for generated source Eligible +0.0% +0.0% 0/1/0
= Diagnose missing generated source inclusion Eligible +0.0% +0.0% 0/1/0
= Diagnose missing output registration for generated non-code file Eligible +0.0% +0.0% 0/1/0
= Diagnose project-level glob for generated source Eligible +0.0% +0.0% 0/1/0
▼ Diagnose wrong hook for generated source files Eligible -100.0% -40.0% 0/0/1

Illustrative judge evidence:

  • Diagnose missing clean tracking for generated source: Position-swap inconsistent (forward: tie, reverse: A). Defaulting to tie.

This is one example, not the aggregate verdict. Open Full Results for every judgment.

➖ Improvement signal, tie-limited — incremental-build (gpt-5.6-luna)

Why: Net win +44.4% (4W/5T/0L over 9 preference-eligible stimulus vote(s), sign test p=0.063), mean preference +17.8% across 9 paired run(s) — not credible — 5 of 9 preference-eligible stimulus vote(s) tied, leaving only 4 discordant preference vote(s). The sign test conditions on non-tie stimulus votes and cannot reach 0.05 below 5, so no record could have passed here — this is not a measured null. Either the skill is inert on these scenarios (make them discriminate) or the eval needs more distinct stimuli to clear the ties

Next action: The signal favors the target, but ties leave too few discordant tasks; inspect ties and predeclare more discriminating breadth.

State: VALID_NO_CHANGE (no_credible_preference_change)

Gate evidence: n=9; 4W/5T/0L; d=4; p=0.063; net +44.4%

Warnings: Activation: isolated 8/9; plugin 7/9

Overfit: Moderate (score 0.21)

Repeated-run reliability (not used by the gate): 9 paired runs (4W/5T/0L).

Weak or warning scenarios:

Scenario Preference gate Net win Δ Pref Runs (W/T/L)
= Diagnose custom targets that always rerun Eligible +0.0% +0.0% 0/1/0
= Distinguish a cold first build from broken incrementality Eligible +0.0% +0.0% 0/1/0
= Explain slower builds when MSBuild skipped everything Eligible +0.0% +0.0% 0/1/0
▲ Explain why clean leaves generated hash source behind Eligible +100.0% +40.0% 1/0/0
= Identify a volatile output path that defeats incrementality Eligible +0.0% +0.0% 0/1/0
= Read a diagnostic log to find the stale input Eligible +0.0% +0.0% 0/1/0

Illustrative judge evidence:

  • Diagnose custom targets that always rerun: Position-swap inconsistent (forward: B, reverse: A). Defaulting to tie.

This is one example, not the aggregate verdict. Open Full Results for every judgment.

➖ Improvement signal, tie-limited — item-management (claude-sonnet-5)

Why: Net win +50.0% (3W/3T/0L over 6 preference-eligible stimulus vote(s), sign test p=0.125), mean preference +31.4% across 7 paired run(s), 1 dormancy stimulus/stimuli excluded from preference — not credible — 3 of 6 preference-eligible stimulus vote(s) tied, leaving only 3 discordant preference vote(s). The sign test conditions on non-tie stimulus votes and cannot reach 0.05 below 5, so no record could have passed here — this is not a measured null. Either the skill is inert on these scenarios (make them discriminate) or the eval needs more distinct stimuli to clear the ties

Next action: The signal favors the target, but ties leave too few discordant tasks; inspect ties and predeclare more discriminating breadth.

State: VALID_NO_CHANGE (no_credible_preference_change)

Gate evidence: n=6; 3W/3T/0L; d=3; p=0.125; net +50.0%; 1 dormancy excluded

Warnings: Activation: isolated 5/6; plugin 4/6

Overfit: Moderate (score 0.29)

Repeated-run reliability (not used by the gate): 7 paired runs (4W/3T/0L).

Weak or warning scenarios:

Scenario Preference gate Net win Δ Pref Runs (W/T/L)
▲ Diagnose an ineffective Compile Remove that does not match the glob Eligible +100.0% +40.0% 1/0/0
= Diagnose real and claimed item problems in a code generation pipeline Eligible +0.0% +0.0% 0/1/0
= Leave already-correct item management unchanged Eligible +0.0% +0.0% 0/1/0
= Leave correct single-list batching unchanged Eligible +0.0% +0.0% 0/1/0

Illustrative judge evidence:

  • Diagnose real and claimed item problems in a code generation pipeline: Position-swap inconsistent (forward: B, reverse: A). Defaulting to tie.

This is one example, not the aggregate verdict. Open Full Results for every judgment.

➖ Improvement signal, tie-limited — item-management (gpt-5.6-luna)

Why: Net win +66.7% (4W/2T/0L over 6 preference-eligible stimulus vote(s), sign test p=0.063), mean preference +31.4% across 7 paired run(s), 1 dormancy stimulus/stimuli excluded from preference — not credible — 2 of 6 preference-eligible stimulus vote(s) tied, leaving only 4 discordant preference vote(s). The sign test conditions on non-tie stimulus votes and cannot reach 0.05 below 5, so no record could have passed here — this is not a measured null. Either the skill is inert on these scenarios (make them discriminate) or the eval needs more distinct stimuli to clear the ties

Next action: The signal favors the target, but ties leave too few discordant tasks; inspect ties and predeclare more discriminating breadth.

State: VALID_NO_CHANGE (no_credible_preference_change)

Gate evidence: n=6; 4W/2T/0L; d=4; p=0.063; net +66.7%; 1 dormancy excluded

Overfit: Moderate (score 0.41)

Repeated-run reliability (not used by the gate): 7 paired runs (4W/3T/0L).

Weak or warning scenarios:

Scenario Preference gate Net win Δ Pref Runs (W/T/L)
= Decline JavaScript configuration merging Excluded (activation contract) +0.0% +0.0% 0/1/0
= Diagnose an ineffective Compile Remove that does not match the glob Eligible +0.0% +0.0% 0/1/0
= Diagnose item group and batching issues Eligible +0.0% +0.0% 0/1/0

Illustrative judge evidence:

  • Diagnose an ineffective Compile Remove that does not match the glob: Position-swap inconsistent (forward: B, reverse: A). Defaulting to tie.

This is one example, not the aggregate verdict. Open Full Results for every judgment.

➖ Improvement signal, tie-limited — msbuild-antipatterns (claude-sonnet-5)

Why: Net win +14.3% (1W/6T/0L over 7 preference-eligible stimulus vote(s), sign test p=0.500), mean preference +17.5% across 8 paired run(s), 1 dormancy stimulus/stimuli excluded from preference — not credible — 6 of 7 preference-eligible stimulus vote(s) tied, leaving only 1 discordant preference vote(s). The sign test conditions on non-tie stimulus votes and cannot reach 0.05 below 5, so no record could have passed here — this is not a measured null. Either the skill is inert on these scenarios (make them discriminate) or the eval needs more distinct stimuli to clear the ties

Next action: The signal favors the target, but ties leave too few discordant tasks; inspect ties and predeclare more discriminating breadth.

State: VALID_NO_CHANGE (no_credible_preference_change)

Gate evidence: n=7; 1W/6T/0L; d=1; p=0.500; net +14.3%; 1 dormancy excluded

Warnings: Activation: isolated 3/7; plugin 3/7

Overfit: Moderate (score 0.31)

Repeated-run reliability (not used by the gate): 8 paired runs (2W/6T/0L).

Weak or warning scenarios:

Scenario Preference gate Net win Δ Pref Runs (W/T/L)
= Add a module to an F# project Eligible +0.0% +0.0% 0/1/0
= Add a signature file to define public API Eligible +0.0% +0.0% 0/1/0
= Fix broken file order causing FS0039 Eligible +0.0% +0.0% 0/1/0
= Judge an unguarded import inside a NuGet package build folder Eligible +0.0% +0.0% 0/1/0
= Leave a clean project without inventing anti-patterns Eligible +0.0% +0.0% 0/1/0
= Review MSBuild files for anti-patterns and style issues Eligible +0.0% +0.0% 0/1/0

Illustrative judge evidence:

  • Add a module to an F# project: Both implementations satisfy the requested functionality and compile/run successfully. B's explicit failure-path sanity check is a modest verification advantage, but it does not establish a materially better final code result.

This is one example, not the aggregate verdict. Open Full Results for every judgment.

➖ Improvement signal, tie-limited — msbuild-antipatterns (gpt-5.6-luna)

Why: Net win +42.9% (3W/4T/0L over 7 preference-eligible stimulus vote(s), sign test p=0.125), mean preference +20.0% across 8 paired run(s), 1 dormancy stimulus/stimuli excluded from preference — not credible — 4 of 7 preference-eligible stimulus vote(s) tied, leaving only 3 discordant preference vote(s). The sign test conditions on non-tie stimulus votes and cannot reach 0.05 below 5, so no record could have passed here — this is not a measured null. Either the skill is inert on these scenarios (make them discriminate) or the eval needs more distinct stimuli to clear the ties

Next action: The signal favors the target, but ties leave too few discordant tasks; inspect ties and predeclare more discriminating breadth.

State: VALID_NO_CHANGE (no_credible_preference_change)

Gate evidence: n=7; 3W/4T/0L; d=3; p=0.125; net +42.9%; 1 dormancy excluded

Warnings: Activation: isolated 5/7; plugin 4/7

Overfit: Moderate (score 0.49)

Repeated-run reliability (not used by the gate): 8 paired runs (4W/4T/0L).

Weak or warning scenarios:

Scenario Preference gate Net win Δ Pref Runs (W/T/L)
▲ Add a module to an F# project Eligible +100.0% +40.0% 1/0/0
= Add a signature file to define public API Eligible +0.0% +0.0% 0/1/0
= Distinguish a style backslash from a real cross-platform backslash bug Eligible +0.0% +0.0% 0/1/0
= Fix broken file order causing FS0039 Eligible +0.0% +0.0% 0/1/0
= Judge an unguarded import inside a NuGet package build folder Eligible +0.0% +0.0% 0/1/0

Illustrative judge evidence:

  • Add a signature file to define public API: Position-swap inconsistent (forward: tie, reverse: A). Defaulting to tie.

This is one example, not the aggregate verdict. Open Full Results for every judgment.

➖ Baseline signal, tie-limited — property-patterns (claude-sonnet-5)

Why: Net win -42.9% (0W/4T/3L over 7 preference-eligible stimulus vote(s), sign test p=0.125), mean preference -25.7% across 7 paired run(s) — no improvement

Next action: Evidence leans baseline but is not credible; inspect losses and make tied tasks discriminate.

State: VALID_NO_CHANGE (no_credible_preference_change)

Gate evidence: n=7; 0W/4T/3L; d=3; p=0.125; net -42.9%

Warnings: Activation: isolated 5/7; plugin 6/7

Overfit: Low (score 0.17)

Repeated-run reliability (not used by the gate): 7 paired runs (0W/4T/3L).

Weak or warning scenarios:

Scenario Preference gate Net win Δ Pref Runs (W/T/L)
▼ Choose the correct shared build file for a post-build target Eligible -100.0% -40.0% 0/0/1
= Choose the right OS detection for a cross-platform property Eligible +0.0% +0.0% 0/1/0
= Diagnose a TargetFramework condition that never applies in props Eligible +0.0% +0.0% 0/1/0
= Diagnose multi-level property hierarchy bugs Eligible +0.0% +0.0% 0/1/0
▼ Diagnose shared build property issues Eligible -100.0% -40.0% 0/0/1
= Fix shared property configuration Eligible +0.0% +0.0% 0/1/0
▼ Leave already-correct shared properties unchanged Eligible -100.0% -100.0% 0/0/1

Illustrative judge evidence:

  • Choose the correct shared build file for a post-build target: Both are substantively correct and answer the organizational question well. A is modestly stronger because it gives the same useful placement rule while actually adding and validating the requested post-build target, with the existing props content left intact.

This is one example, not the aggregate verdict. Open Full Results for every judgment.

➖ Mixed evidence — property-patterns (gpt-5.6-luna)

Why: Net win +0.0% (2W/3T/2L over 7 preference-eligible stimulus vote(s), sign test p=0.687), mean preference -8.6% across 7 paired run(s) — no improvement

Next action: Compare winning and losing scenarios to isolate where the target helps versus hurts.

State: VALID_NO_CHANGE (no_credible_preference_change)

Gate evidence: n=7; 2W/3T/2L; d=4; p=0.687; net +0.0%

Warnings: Activation: isolated 6/7; plugin 6/7

Overfit: Moderate (score 0.23)

Repeated-run reliability (not used by the gate): 7 paired runs (2W/3T/2L).

Weak or warning scenarios:

Scenario Preference gate Net win Δ Pref Runs (W/T/L)
▲ Choose the correct shared build file for a post-build target Eligible +100.0% +40.0% 1/0/0
▼ Choose the right OS detection for a cross-platform property Eligible -100.0% -40.0% 0/0/1
= Diagnose a TargetFramework condition that never applies in props Eligible +0.0% +0.0% 0/1/0
= Diagnose multi-level property hierarchy bugs Eligible +0.0% +0.0% 0/1/0
= Diagnose shared build property issues Eligible +0.0% +0.0% 0/1/0
▼ Leave already-correct shared properties unchanged Eligible -100.0% -100.0% 0/0/1

Illustrative judge evidence:

  • Choose the right OS detection for a cross-platform property: Response A edges ahead by successfully implementing and verifying the fix—the build completed with zero errors, proving the solution works in practice. While Response B has marginally better capitalization convention and more complete XML formatting, practical validation is mo...

This is one example, not the aggregate verdict. Open Full Results for every judgment.

✅ Improved — binlog-generation (gpt-5.6-luna)

Why: Net win +71.4% (5W/2T/0L over 7 preference-eligible stimulus vote(s), sign test p=0.031), mean preference +55.0% across 8 paired run(s), 1 dormancy stimulus/stimuli excluded from preference — credibly better

Next action: Fix activation gaps; Review overfit evidence.

State: VALID_PASS (credible_preference_improvement)

Gate evidence: n=7; 5W/2T/0L; d=5; p=0.031; net +71.4%; 1 dormancy excluded

Warnings: Activation: isolated 6/7; plugin 6/7

Overfit: Moderate (score 0.26)

Repeated-run reliability (not used by the gate): 8 paired runs (5W/3T/0L).

Weak or warning scenarios:

Scenario Preference gate Net win Δ Pref Runs (W/T/L)
= Choose a predictable, non-colliding binlog name for a CI upload step Eligible +0.0% +0.0% 0/1/0
= Decline Maven compiler configuration Excluded (activation contract) +0.0% +0.0% 0/1/0
= Fix a CI build script that reuses the same binlog on every run Eligible +0.0% +0.0% 0/1/0
▲ Preserve binlog history while cleaning stale build output Eligible +100.0% +100.0% 1/0/0

Illustrative judge evidence:

  • Choose a predictable, non-colliding binlog name for a CI upload step: Position-swap inconsistent (forward: A, reverse: tie). Defaulting to tie.

This is one example, not the aggregate verdict. Open Full Results for every judgment.

✅ Improved — eval-performance (gpt-5.6-luna)

Why: Net win +62.5% (5W/3T/0L over 8 preference-eligible stimulus vote(s), sign test p=0.031), mean preference +47.5% across 8 paired run(s) — credibly better

Next action: Review overfit evidence.

State: VALID_PASS (credible_preference_improvement)

Gate evidence: n=8; 5W/3T/0L; d=5; p=0.031; net +62.5%

Overfit: Moderate (score 0.32)

Repeated-run reliability (not used by the gate): 8 paired runs (5W/3T/0L).

Weak or warning scenarios:

Scenario Preference gate Net win Δ Pref Runs (W/T/L)
= Recognize TreatAsLocalProperty overuse versus one justified entry Eligible +0.0% +0.0% 0/1/0
= Redirect a compile-time slowdown mistakenly framed as an evaluation problem Eligible +0.0% +0.0% 0/1/0
= Triage which of two property functions actually costs evaluation time Eligible +0.0% +0.0% 0/1/0

Illustrative judge evidence:

  • Recognize TreatAsLocalProperty overuse versus one justified entry: Position-swap inconsistent (forward: B, reverse: A). Defaulting to tie.

This is one example, not the aggregate verdict. Open Full Results for every judgment.

Details for 3 results were omitted to keep this comment under GitHub's limit. Open the workflow summary or Full Results for the complete breakdown.

🔍 Full Results - all metrics and investigation details

To investigate non-passing or warning results, paste this to your AI coding agent:

For PR 1215 in dotnet/skills, download eval artifacts with gh run download 37311852075 --repo dotnet/skills --pattern "vally-results-*" --dir ./eval-results, then fetch https://raw.githubusercontent.com/dotnet/skills/b9052fd631b50ceda0f8f4d3033bc3b02fb97671/eng/vally-adapter/InvestigatingResults.md and follow it. Classify each result as measurement-invalid, underpowered, not-proven, preference-loss, or passing-with-warning. Use stateReason, result accounting, weak scenarios, and judge evidence to give the cause and exact next fix.

⚠️ Session replay telemetry was not published because the auxiliary dotnet/skills-data publisher failed. The evaluation verdicts above remain authoritative; maintainers must repair SKILLS_DATA_TOKEN or the publisher before replay links are available.

@github-actions

github-actions Bot commented Oct 5, 2026

Copy link
Copy Markdown
Contributor

✅ Approved by @YuliiaKovalova. cc @dotnet/skills-merge-approvers — ready to merge.

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

ready-to-merge PR state label

Projects

None yet

Development

Successfully merging this pull request may close these issues.

3 participants