Skip to content

test(msbuild): strengthen evaluation coverage and infrastructure - #1214

Merged
AbhitejJohn merged 22 commits into
mainfrom
abhitejjohn-msbuild-eval-infrastructure
Sep 30, 2026
Merged

AbhitejJohn merged 22 commits into
mainfrom
abhitejjohn-msbuild-eval-infrastructure

Conversation

@AbhitejJohn

@AbhitejJohn AbhitejJohn commented Sep 25, 2026 •

Copy link
Copy Markdown
Collaborator

Summary

This is a clean replacement for #1118, rebuilt from current main. It expands the MSBuild evaluation surface to 18 specs and 135 distinct stimuli with self-contained fixtures, deterministic graders, replayable patches, and activation-aware routing boundaries.

The change also includes only the shared infrastructure required to execute those evaluations safely and interpret their results correctly. It excludes MSBuild skill and agent content, host manifests, dotnet-test content, dashboard/versioning work, and model-evaluation actions.

Why

The prior evaluation surface mixed real skill behavior with underpowered suites, contradictory fixtures, unsafe evaluator file access, incomplete result accounting, and ambiguous activation evidence. Those defects could produce false regressions, hide terminal failures, or treat a sibling skill invocation as target activation.

This replacement separates those causes. Each MSBuild suite now has enough distinct scenarios to produce a credible result, and the evaluator fails closed when execution, pairing, dormancy evidence, or sandbox containment is incomplete.

Impact

MSBuild evaluation changes

Failure Fix Expected result
Underpowered suites could not pass the statistical gate Expand 18 specs to 135 distinct stimuli and remove 17 MSBuild entries from the underpowered allowlist Each changed suite can reach a conclusive sign-test result
Fixtures failed for unrelated reasons or depended on ambient state Add self-contained healthy, broken, no-op, routing, performance, binlog, and project-graph fixtures Both arms observe the same deterministic starting state
Prompt-only graders accepted unverified claims Add run-command, file, output, golden-patch, and ATIF trajectory evidence Evaluation success is tied to executable or inspectable outcomes
Dormancy and routing boundaries were ambiguous Add positive prerequisites, non-MSBuild boundaries, and explicit expect_activation: false scenarios Target activation is measured without forcing the skilled arm to equal baseline
Binlog and performance cases lacked repeatable artifacts Add setup scripts, stable binlog placement, text-log fallbacks, and realistic performance summaries Failure analysis and performance recommendations use reproducible evidence

Shared evaluation infrastructure

Failure Fix Expected result
Push or manual validation could inspect the wrong scope Select the pre-push SHA for pushes and origin/main for manual runs; reserve repository-wide audits for explicit local use Every changed suite is validated without failing on unrelated migration debt
Nested stimulus tags could silently alter execution shards Parse executionShard only from the top-level tags mapping MSBuild default, medium, and heavy shards remain deterministic
Evaluator file access was vulnerable to path races and scope leaks Use a process-private root, scenario allowlists, fail-closed permissions, and handle-relative no-follow reads, writes, enumeration, removal, and rename Agent, judge, and session files stay inside approved roots after parent-path swaps
Failed, interrupted, or duplicate runs could disappear during rejudge Persist terminal failures and activation expectations; reject nonterminal sessions; require complete, unique baseline/treatment pairing Partial or interrupted evidence cannot publish a success-shaped verdict
Every MCP-enabled arm created a cold NuGet package/cache directory Share one process-private package and HTTP cache while keeping validator-owned NuGet configuration Repeated arms reuse the pinned package without exposing host NuGet state
Primary native-agent selection was not always visible in SDK telemetry Record agent.primary_selected and deduplicate it with SDK subagent events Successful primary selection is authoritative activation evidence
Plugin activation counted sibling skills and missing dormancy evidence was only a warning Attribute named events to the target skill and fail missing dormancy contracts Activation and dormancy gates use target-specific evidence
Shared infrastructure admission ledger (37 candidates)
Path Disposition Hunk-level reason
.github/workflows/eval-quality.yml Included Selects safe PR/push bases and compares manual runs with origin/main
eng/eval-quality/README.md Included Documents the admitted deterministic quality checks
eng/eval-quality/check_eval_quality.py Included Validates fixtures, references, graders, tags, golden evidence, patches, commands, and symlink containment
eng/eval-quality/selftest_eval_quality.py Included Locks every new quality rule with regression tests
eng/eval-quality/underpowered-allowlist.txt Included Removes the MSBuild suites that now clear the power floor
eng/evaluation/find-targets.ps1 Included Restricts shard discovery to top-level tags.executionShard
eng/evaluation/test_token_failover.py Included Tests deterministic shard discovery in the workflow
eng/skill-validator/src/Evaluate/AgentRunner.cs Included Adds private roots, fail-closed permissions, staged workspaces, shared private MCP caches, and primary-agent activation evidence
eng/skill-validator/src/Evaluate/Comparator.cs Included Keeps dormant scenarios in reports without adding them to preference scoring
eng/skill-validator/src/Evaluate/EvalSchema.cs Included Preserves run-command argument arrays without shell reconstruction
eng/skill-validator/src/Evaluate/EvaluateCommand.cs Included Fails terminal execution errors and applies skill/agent activation contracts
eng/skill-validator/src/Evaluate/LlmSession.cs Included Places judge sessions in the private evaluator root with scoped file access
eng/skill-validator/src/Evaluate/LocalSessionFsHandler.cs Included Separates roots and routes enumeration, removal, and rename through secure operations
eng/skill-validator/src/Evaluate/MetricsCollector.cs Included Counts and deduplicates agent.primary_selected evidence
eng/skill-validator/src/Evaluate/Models.cs Included Adds explicit unexpected-activation and execution-error classifications
eng/skill-validator/src/Evaluate/OverfittingCommand.cs Included Moves overfitting judge work into the private evaluator root
eng/skill-validator/src/Evaluate/RejudgeCommand.cs Included Rejects nonterminal sessions, enforces exact pairing, and reapplies persisted activation/execution gates
eng/skill-validator/src/Evaluate/Reporter.cs Included Reports dormant, missing, and unexpected activation without hiding failures
eng/skill-validator/src/Evaluate/SecureFileSystem.cs Included Adds cross-platform no-follow, descriptor-safe, handle-relative and atomic filesystem operations
eng/skill-validator/src/Evaluate/SessionDatabase.cs Included Persists activation expectations and terminal failures in schema v5
eng/skill-validator/src/README.md Included Documents fail-closed rejudge pairing
eng/skill-validator/src/docs/InvestigatingResults.md Included Synchronizes result schema, activation, failure, and rejudge guidance
eng/skill-validator/tests/Evaluate/ComparatorTests.cs Included Tests dormant reporting without score contamination
eng/skill-validator/tests/Evaluate/EvalDiscoveryTests.cs Included Tests Vally stimulus and activation parsing
eng/skill-validator/tests/Evaluate/EvaluateCommandTests.cs Included Tests execution-error and activation-contract semantics
eng/skill-validator/tests/Evaluate/MetricsTests.cs Included Tests primary-agent evidence and deduplication
eng/skill-validator/tests/Evaluate/RejudgeCommandTests.cs Included Tests complete pairing, duplicate detection, dormancy recovery, and activation gates
eng/skill-validator/tests/Evaluate/RunnerTests.cs Included Tests containment, permissions, secure file races, token isolation, and hardened MCP launch; manifest-only assertion omitted
eng/skill-validator/tests/Evaluate/SessionDatabaseTests.cs Included Tests migration, activation persistence, terminal failures, and reused baselines
eng/vally-adapter/InvestigatingResults.md Included Documents target-specific activation and missing-dormancy failures
eng/vally-adapter/README.md Included Documents adapted activation-contract behavior
eng/vally-adapter/adapt.mjs Included Filters activation by target name and retains missing dormancy as a failure
eng/vally-adapter/adapt.test.mjs Included Tests sibling filtering, missing dormancy, and target activation
.agents/skills/create-skill-test/SKILL.md Excluded Authoring guidance does not affect PR1 execution, result correctness, validation, or security
.agents/skills/improve-skill-quality/SKILL.md Excluded Authoring guidance is outside replacement runtime scope
.agents/skills/improve-skill-quality/references/eval-triage.md Excluded Triage guidance is outside replacement runtime scope
eng/skill-validator/tests/Check/PluginMcpManifestTests.cs Excluded Supports excluded host/Codex manifest allowlist changes and produced a stale tool expectation without them

Validation

  • python eng/eval-quality/check_eval_quality.py --base-ref origin/main — passed for 18 changed suites of 101 specs
  • python eng/eval-quality/selftest_eval_quality.py — 127 passed
  • python -m unittest discover -s eng/evaluation -p "test_*.py" -v — 58 passed
  • node --test --test-concurrency=1 eng/vally-adapter/*.test.mjs — 203 passed
  • dotnet test eng/skill-validator/tests/SkillValidator.Tests.csproj — 834 passed
  • Managed osx-arm64 cross-build with PublishAot=false — passed with zero warnings or errors
  • NativeAOT osx-arm64 publish was not run because the .NET IL compiler does not support cross-OS native compilation from Windows; Darwin-specific runtime coverage is included for macOS CI
  • dotnet run --project eng/skill-validator/src/SkillValidator.csproj -- check --plugin ./plugins/dotnet-msbuild — passed for 18 skills, 3 agents, and 1 plugin
  • actionlint 1.7.7 with the pinned archive checksum — passed
  • Representative fixture builds — 13 expected success/failure states reproduced
  • Binlog and multi-evaluation setup scripts — 4 deterministic artifact reproductions passed
  • Scope audit — 345 changed paths, all inside the admitted test and shared-infrastructure paths

Review flow

flowchart LR
    A[Tracked MSBuild fixture] --> B[Eval-quality preflight]
    B --> C[Baseline / isolated / plugin execution]
    C --> D{Evidence complete?}
    D -- yes --> E[Exact accounting and verdict]
    D -- bounded transient failure --> F[Targeted recovery]
    F --> E
    D -- incomplete or ambiguous --> G[Fail closed]
Loading

Suggested review order

  1. Review tests/dotnet-msbuild/** for scenario ownership, neutral fixtures, executable graders, and golden evidence.
  2. Review eng/eval-quality/** and eng/evaluation/** for deterministic preflight and discovery.
  3. Review eng/skill-validator/** for sandboxing, permissions, activation, session completeness, and rejudge behavior.
  4. Review eng/vally-adapter/** for exact accounting, activation contracts, and bounded recovery.
  5. Review stacked PR fix(msbuild): sharpen skill guidance #1215 separately for production SKILL.md guidance.

AbhitejJohn and others added 4 commits September 24, 2026 23:52
Re-carve the MSBuild evaluation specs, fixtures, setup scripts, and golden evidence from the validated source branch onto current main.

Co-authored-by: Copilot App <223556219+Copilot@users.noreply.github.com>
Add deterministic eval-quality checks, secure evaluator filesystem boundaries, exact rejudge accounting, and target-specific activation evidence required by the MSBuild evaluation corpus.

Co-authored-by: Copilot App <223556219+Copilot@users.noreply.github.com>
Reject interrupted rejudge data, secure directory operations against path swaps, reuse private MCP caches, and keep manual quality runs scoped to main.

Co-authored-by: Copilot App <223556219+Copilot@users.noreply.github.com>
Use Darwin dirent and unlinkat layouts for handle-relative enumeration and recursive removal, with macOS-specific regression coverage.

Co-authored-by: Copilot App <223556219+Copilot@users.noreply.github.com>
Copilot AI lite review requested due to automatic review settings September 25, 2026 07:48
@github-actions

Copy link
Copy Markdown
Contributor

Skill Coverage Report

Plugin Skill Covered Coverage
❔ dotnet-msbuild agent.msbuild — error
✅ dotnet-msbuild check-bin-obj-clash 4/5 80%
❌ dotnet-msbuild directory-build-organization 0/1 0%
❌ dotnet-msbuild extension-points 0/1 0%
✅ dotnet-msbuild msbuild-modernization 7/7 100%
✅ dotnet-msbuild property-patterns 1/1 100%
✅ dotnet-msbuild resolve-project-references 6/6 100%
Uncovered: dotnet-msbuild/check-bin-obj-clash
  • [WorkflowStep] Step 2: Get an overview and list projects (line 48)
Uncovered: dotnet-msbuild/directory-build-organization
  • [CodePattern] [MSBuild] (line 122)
Uncovered: dotnet-msbuild/extension-points
  • [CodePattern] [MSBuild] (line 195)

Restrict Darwin dirent parsing to arm64, clear native errno correctly, and rely on the existing cross-platform filesystem regression coverage.

Co-authored-by: Copilot App <223556219+Copilot@users.noreply.github.com>

Copilot AI left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Copilot review overview

🟡 Changes recommended

Critical and moderate issues remain in workflow coverage, shard discovery, rejudge pairing, filesystem deletion, dormancy handling, and a fixture patch.

Get a fresh assessment by requesting another Copilot review.

Review effort: Lite
Findings: 5 High severity · 1 Medium severity

Open (6)
What changed in this PR

This PR expands MSBuild evaluation coverage with deterministic fixtures and strengthens shared evaluation infrastructure for isolation, activation tracking, rejudge pairing, and quality validation.

Changes:

  • Adds MSBuild fixtures, reference patches, performance artifacts, and routing scenarios.
  • Hardens evaluator filesystem, session, cache, activation, and dormancy handling.
  • Updates quality validation, shard discovery, and rejudge accounting.
File Change
tests/​dotnet-msbuild/​target-authoring/​TargetAuthoring.csproj Reviewed
tests/​dotnet-msbuild/​target-authoring/​returns-outputs/​ReturnsOutputs.csproj Reviewed
tests/​dotnet-msbuild/​target-authoring/​returns-outputs/​Placeholder.cs Reviewed
tests/​dotnet-msbuild/​target-authoring/​references/​returns-outputs.patch Reviewed
tests/​dotnet-msbuild/​target-authoring/​references/​fix-antipatterns.patch Reviewed
tests/​dotnet-msbuild/​target-authoring/​Placeholder.cs Reviewed
tests/​dotnet-msbuild/​target-authoring/​dormancy/​Placeholder.cs Reviewed
tests/​dotnet-msbuild/​target-authoring/​dormancy/​perf-summary.txt Reviewed
tests/​dotnet-msbuild/​target-authoring/​dormancy/​IncrementalTuning.csproj Reviewed
tests/​dotnet-msbuild/​target-authoring/​dormancy/​Api.schema Reviewed
tests/​dotnet-msbuild/​target-authoring/​correct-target/​schemas/​api.schema Reviewed
tests/​dotnet-msbuild/​target-authoring/​correct-target/​Placeholder.cs Reviewed
tests/​dotnet-msbuild/​target-authoring/​correct-target/​CorrectTarget.csproj Reviewed
tests/​dotnet-msbuild/​target-authoring/​before-targets/​Placeholder.cs Reviewed
tests/​dotnet-msbuild/​target-authoring/​before-targets/​BeforeTargets.csproj Reviewed
tests/​dotnet-msbuild/​target-authoring/​before-targets/​Api.schema Reviewed
tests/​dotnet-msbuild/​target-authoring/​ApiClient.schema Reviewed
tests/​dotnet-msbuild/​resolve-project-references/​wide-graph/​perf-report-wide-graph.md Reviewed
tests/​dotnet-msbuild/​resolve-project-references/​wide-graph/​Core/​CoreContracts.cs Reviewed
tests/​dotnet-msbuild/​resolve-project-references/​wide-graph/​Core/​Core.csproj Reviewed
tests/​dotnet-msbuild/​resolve-project-references/​wide-graph/​App4/​App4Feature.cs Reviewed
tests/​dotnet-msbuild/​resolve-project-references/​wide-graph/​App4/​App4.csproj Reviewed
tests/​dotnet-msbuild/​resolve-project-references/​wide-graph/​App3/​App3Feature.cs Reviewed
tests/​dotnet-msbuild/​resolve-project-references/​wide-graph/​App3/​App3.csproj Reviewed
tests/​dotnet-msbuild/​resolve-project-references/​wide-graph/​App2/​App2Feature.cs Reviewed
tests/​dotnet-msbuild/​resolve-project-references/​wide-graph/​App2/​App2.csproj Reviewed
tests/​dotnet-msbuild/​resolve-project-references/​wide-graph/​App1/​App1Feature.cs Reviewed
tests/​dotnet-msbuild/​resolve-project-references/​wide-graph/​App1/​App1.csproj Reviewed
tests/​dotnet-msbuild/​resolve-project-references/​wide-graph/​Aggregator/​AggregatorReport.cs Reviewed
tests/​dotnet-msbuild/​resolve-project-references/​wide-graph/​Aggregator/​Aggregator.csproj Reviewed
tests/​dotnet-msbuild/​resolve-project-references/​serial-chain/​Web/​WebService.cs Reviewed
tests/​dotnet-msbuild/​resolve-project-references/​serial-chain/​Web/​Web.csproj Reviewed
tests/​dotnet-msbuild/​resolve-project-references/​serial-chain/​Tests/​Tests.csproj Reviewed
tests/​dotnet-msbuild/​resolve-project-references/​serial-chain/​Tests/​SmokeTests.cs Reviewed
tests/​dotnet-msbuild/​resolve-project-references/​serial-chain/​perf-report-serial-chain.md Reviewed
tests/​dotnet-msbuild/​resolve-project-references/​serial-chain/​Core/​CoreService.cs Reviewed
tests/​dotnet-msbuild/​resolve-project-references/​serial-chain/​Core/​Core.csproj Reviewed
tests/​dotnet-msbuild/​resolve-project-references/​serial-chain/​Api/​ApiService.cs Reviewed
tests/​dotnet-msbuild/​resolve-project-references/​serial-chain/​Api/​Api.csproj Reviewed
tests/​dotnet-msbuild/​resolve-project-references/​perf-report-general-build.md Reviewed
tests/​dotnet-msbuild/​resolve-project-references/​perf-report-csc-vs-target.md Reviewed
tests/​dotnet-msbuild/​resolve-project-references/​perf-report-csc-obvious.md Reviewed
tests/​dotnet-msbuild/​resolve-project-references/​perf-report-copy-vs-csc.md Reviewed
tests/​dotnet-msbuild/​property-patterns/​tfm-timing/​TfmTiming.csproj Reviewed
tests/​dotnet-msbuild/​property-patterns/​tfm-timing/​Placeholder.cs Reviewed
tests/​dotnet-msbuild/​property-patterns/​tfm-timing/​Directory.Build.props Reviewed
tests/​dotnet-msbuild/​property-patterns/​references/​fix-shared.patch Reviewed
tests/​dotnet-msbuild/​property-patterns/​placement/​Placement.csproj Reviewed
tests/​dotnet-msbuild/​property-patterns/​placement/​Placeholder.cs Reviewed
tests/​dotnet-msbuild/​property-patterns/​placement/​Directory.Build.props Reviewed
tests/​dotnet-msbuild/​property-patterns/​os-check/​Placeholder.cs Reviewed
tests/​dotnet-msbuild/​property-patterns/​os-check/​OsCheck.csproj Reviewed
tests/​dotnet-msbuild/​property-patterns/​os-check/​Directory.Build.props Reviewed
tests/​dotnet-msbuild/​property-patterns/​correct-props/​Placeholder.cs Reviewed
tests/​dotnet-msbuild/​property-patterns/​correct-props/​Directory.Build.props Reviewed
tests/​dotnet-msbuild/​property-patterns/​correct-props/​CorrectProps.csproj Reviewed
tests/​dotnet-msbuild/​msbuild-modernization/​references/​langversion-trap.patch Reviewed
tests/​dotnet-msbuild/​msbuild-modernization/​references/​cpm-restraint.patch Reviewed
tests/​dotnet-msbuild/​msbuild-modernization/​partial-modern/​Repository.cs Reviewed
tests/​dotnet-msbuild/​msbuild-modernization/​partial-modern/​PartialModern.csproj Reviewed
tests/​dotnet-msbuild/​msbuild-modernization/​node/​webpack.config.js Reviewed
tests/​dotnet-msbuild/​msbuild-modernization/​node/​package.json Reviewed
tests/​dotnet-msbuild/​msbuild-modernization/​multitarget/​Widget.cs Reviewed
tests/​dotnet-msbuild/​msbuild-modernization/​multitarget/​MultiTarget.Net48.csproj Reviewed
tests/​dotnet-msbuild/​msbuild-modernization/​multitarget/​MultiTarget.csproj Reviewed
tests/​dotnet-msbuild/​msbuild-modernization/​cpm-single/​Serializer.cs Reviewed
tests/​dotnet-msbuild/​msbuild-modernization/​cpm-single/​packages.config Reviewed
tests/​dotnet-msbuild/​msbuild-modernization/​cpm-single/​CpmSingle.csproj Reviewed
tests/​dotnet-msbuild/​msbuild-modernization/​already-modern/​Clock.cs Reviewed
tests/​dotnet-msbuild/​msbuild-modernization/​already-modern/​AlreadyModern.csproj Reviewed
tests/​dotnet-msbuild/​msbuild-antipatterns/​references/​fsharp-signature.patch Reviewed
tests/​dotnet-msbuild/​msbuild-antipatterns/​references/​fsharp-broken-order.patch Reviewed
tests/​dotnet-msbuild/​msbuild-antipatterns/​references/​fsharp-add-module.patch Reviewed
tests/​dotnet-msbuild/​msbuild-antipatterns/​nuget-forwarder/​tools/​common/​MyPkg.Implementation.targets Reviewed
tests/​dotnet-msbuild/​msbuild-antipatterns/​nuget-forwarder/​MyPkg.nuspec Reviewed
tests/​dotnet-msbuild/​msbuild-antipatterns/​nuget-forwarder/​MyPkg.csproj Reviewed
tests/​dotnet-msbuild/​msbuild-antipatterns/​nuget-forwarder/​lib/​net8.0/​_._ Reviewed
tests/​dotnet-msbuild/​msbuild-antipatterns/​nuget-forwarder/​buildTransitive/​common/​MyPkg.targets Reviewed
tests/​dotnet-msbuild/​msbuild-antipatterns/​nuget-forwarder/​build/​net8.0/​MyPkg.props Reviewed
tests/​dotnet-msbuild/​msbuild-antipatterns/​clean-project/​Greeter.cs Reviewed
tests/​dotnet-msbuild/​msbuild-antipatterns/​clean-project/​CleanLib.csproj Reviewed
tests/​dotnet-msbuild/​msbuild-antipatterns/​backslash/​marker-output/​README.txt Reviewed
tests/​dotnet-msbuild/​msbuild-antipatterns/​backslash/​Build.targets Reviewed
tests/​dotnet-msbuild/​item-management/​references/​diagnose-hard.patch Reviewed
tests/​dotnet-msbuild/​item-management/​glob-mismatch/​Service.cs Reviewed
tests/​dotnet-msbuild/​item-management/​glob-mismatch/​GlobMismatch.csproj Reviewed
tests/​dotnet-msbuild/​item-management/​glob-mismatch/​Generated/​Api.g.cs Reviewed
tests/​dotnet-msbuild/​item-management/​false-positive/​Widget.cs Reviewed
tests/​dotnet-msbuild/​item-management/​false-positive/​FalsePositive.csproj Reviewed
tests/​dotnet-msbuild/​item-management/​already-correct/​Legacy/​LegacyHelper.cs Reviewed
tests/​dotnet-msbuild/​item-management/​already-correct/​Calculator.cs Reviewed
tests/​dotnet-msbuild/​item-management/​already-correct/​AlreadyCorrect.csproj Reviewed
tests/​dotnet-msbuild/​incremental-build/​volatile-output/​VolatileOutput.csproj Reviewed
tests/​dotnet-msbuild/​incremental-build/​volatile-output/​VersionSeed.txt Reviewed
tests/​dotnet-msbuild/​incremental-build/​volatile-output/​Program.cs Reviewed
tests/​dotnet-msbuild/​incremental-build/​returns-query/​ReturnsQuery.csproj Reviewed
tests/​dotnet-msbuild/​incremental-build/​returns-query/​Placeholder.cs Reviewed
tests/​dotnet-msbuild/​incremental-build/​Program.cs Reviewed
tests/​dotnet-msbuild/​incremental-build/​log-interpretation/​diag-excerpt.txt Reviewed
tests/​dotnet-msbuild/​incremental-build/​GitHash.txt Reviewed
tests/​dotnet-msbuild/​incremental-build/​compiler-server-cache/​diag-excerpt.txt Reviewed
tests/​dotnet-msbuild/​incremental-build/​clean-tracking/​Program.cs Reviewed
tests/​dotnet-msbuild/​incremental-build/​clean-tracking/​GitHash.txt Reviewed
tests/​dotnet-msbuild/​incremental-build/​clean-tracking/​CleanTracking.csproj Reviewed
tests/​dotnet-msbuild/​incremental-build/​BuildStamp.txt Reviewed
tests/​dotnet-msbuild/​including-generated-files/​wrong-timing/​WrongTiming.csproj Reviewed
tests/​dotnet-msbuild/​including-generated-files/​wrong-timing/​Program.cs Reviewed
tests/​dotnet-msbuild/​including-generated-files/​references/​fix-generated-source.patch Reviewed
tests/​dotnet-msbuild/​including-generated-files/​outside-target-glob/​Program.cs Reviewed
tests/​dotnet-msbuild/​including-generated-files/​outside-target-glob/​OutsideTargetGlob.csproj Reviewed
tests/​dotnet-msbuild/​including-generated-files/​non-code-output/​Program.cs Reviewed
tests/​dotnet-msbuild/​including-generated-files/​non-code-output/​NonCodeOutput.csproj Reviewed
tests/​dotnet-msbuild/​including-generated-files/​missing-filewrites/​Program.cs Reviewed
tests/​dotnet-msbuild/​including-generated-files/​missing-filewrites/​MissingFileWrites.csproj Reviewed
tests/​dotnet-msbuild/​including-generated-files/​hardcoded-obj-path/​Program.cs Reviewed
tests/​dotnet-msbuild/​including-generated-files/​hardcoded-obj-path/​HardcodedObjPath.csproj Reviewed
tests/​dotnet-msbuild/​including-generated-files/​hardcoded-obj-path/​Directory.Build.props Reviewed
tests/​dotnet-msbuild/​extension-points/​tfm-forwarder-msb4019/​MyAdapter.nuspec Reviewed
tests/​dotnet-msbuild/​extension-points/​tfm-forwarder-msb4019/​buildTransitive/​net8.0/​MyAdapter.props Reviewed
tests/​dotnet-msbuild/​extension-points/​tfm-forwarder-msb4019/​build/​net8.0/​MyAdapter.props Reviewed
tests/​dotnet-msbuild/​extension-points/​references/​fix-extension-points.patch Reviewed
tests/​dotnet-msbuild/​extension-points/​references/​create-mysdk-extension-point.patch Reviewed
tests/​dotnet-msbuild/​extension-points/​packed-layout-false-positive/​MyAdapter.nuspec Reviewed
tests/​dotnet-msbuild/​extension-points/​packed-layout-false-positive/​buildTransitive/​common/​MyAdapter.props Reviewed
tests/​dotnet-msbuild/​extension-points/​packed-layout-false-positive/​build/​net8.0/​MyAdapter.props Reviewed
tests/​dotnet-msbuild/​extension-points/​MyHook.targets Reviewed
tests/​dotnet-msbuild/​extension-points/​filename-id-mismatch/​MyAnalyzer.BuildTasks.nuspec Reviewed
tests/​dotnet-msbuild/​extension-points/​filename-id-mismatch/​build/​MyAnalyzerBuildTasks.targets Reviewed
tests/​dotnet-msbuild/​extension-points/​filename-id-mismatch/​build/​MyAnalyzerBuildTasks.props Reviewed
tests/​dotnet-msbuild/​extension-points/​ExtensionPoints.csproj Reviewed
tests/​dotnet-msbuild/​extension-points/​Directory.Build.targets Reviewed
tests/​dotnet-msbuild/​extension-points/​author-extension-point/​MySDK.targets Reviewed
tests/​dotnet-msbuild/​extension-points/​author-extension-point/​MySDK.Before.targets Reviewed
tests/​dotnet-msbuild/​extension-points/​author-extension-point/​MySDK.After.targets Reviewed
tests/​dotnet-msbuild/​extension-points/​author-extension-point/​Driver.proj Reviewed
tests/​dotnet-msbuild/​eval-performance/​treat-as-local-property/​TaLpDemo.csproj Reviewed
tests/​dotnet-msbuild/​eval-performance/​treat-as-local-property/​Program.cs Reviewed
tests/​dotnet-msbuild/​eval-performance/​setup/​prepare-multi-evaluation.mjs Reviewed
tests/​dotnet-msbuild/​eval-performance/​references/​treat-as-local-property-overuse.json Reviewed
tests/​dotnet-msbuild/​eval-performance/​references/​property-function-triage.json Reviewed
tests/​dotnet-msbuild/​eval-performance/​references/​no-op-already-fast.json Reviewed
tests/​dotnet-msbuild/​eval-performance/​references/​multi-pattern-diagnosis.json Reviewed
tests/​dotnet-msbuild/​eval-performance/​references/​measurement-first-discipline.json Reviewed
tests/​dotnet-msbuild/​eval-performance/​references/​boundary-incremental-build.json Reviewed
tests/​dotnet-msbuild/​eval-performance/​references/​boundary-compile-time.json Reviewed
tests/​dotnet-msbuild/​eval-performance/​property-function-triage/​PropFuncs.csproj Reviewed
tests/​dotnet-msbuild/​eval-performance/​property-function-triage/​Program.cs Reviewed
tests/​dotnet-msbuild/​eval-performance/​property-function-triage/​notes.txt Reviewed
tests/​dotnet-msbuild/​eval-performance/​no-op-fast/​Program.cs Reviewed
tests/​dotnet-msbuild/​eval-performance/​no-op-fast/​FastEval.csproj Reviewed
tests/​dotnet-msbuild/​eval-performance/​multi-evaluation/​Shared/​Shared.csproj Reviewed
tests/​dotnet-msbuild/​eval-performance/​multi-evaluation/​Shared/​FlavorInfo.cs Reviewed
tests/​dotnet-msbuild/​eval-performance/​multi-evaluation/​Build.proj Reviewed
tests/​dotnet-msbuild/​eval-performance/​measurement-first/​Program.cs Reviewed
tests/​dotnet-msbuild/​eval-performance/​measurement-first/​BigApp.csproj Reviewed
tests/​dotnet-msbuild/​eval-performance/​boundary-incremental-build/​RebuildComplaint.csproj Reviewed
tests/​dotnet-msbuild/​eval-performance/​boundary-incremental-build/​Program.cs Reviewed
tests/​dotnet-msbuild/​eval-performance/​boundary-compile-time/​SlowCompile.csproj Reviewed
tests/​dotnet-msbuild/​eval-performance/​boundary-compile-time/​Program.cs Reviewed
tests/​dotnet-msbuild/​eval-performance/​boundary-compile-time/​perf-summary.txt Reviewed
tests/​dotnet-msbuild/​directory-build-organization/​tfm-timing/​TfmTiming.csproj Reviewed
tests/​dotnet-msbuild/​directory-build-organization/​tfm-timing/​Placeholder.cs Reviewed
tests/​dotnet-msbuild/​directory-build-organization/​tfm-timing/​Directory.Build.props Reviewed
tests/​dotnet-msbuild/​directory-build-organization/​single-project-console/​SingleProjectConsole.csproj Reviewed
tests/​dotnet-msbuild/​directory-build-organization/​single-project-console/​Program.cs Reviewed
tests/​dotnet-msbuild/​directory-build-organization/​intentional-divergence/​src/​Worker/​Worker.csproj Reviewed
tests/​dotnet-msbuild/​directory-build-organization/​intentional-divergence/​src/​Worker/​Worker.cs Reviewed
tests/​dotnet-msbuild/​directory-build-organization/​intentional-divergence/​src/​GeneratedClient/​GeneratedClient.csproj Reviewed
tests/​dotnet-msbuild/​directory-build-organization/​intentional-divergence/​src/​GeneratedClient/​Client.cs Reviewed
tests/​dotnet-msbuild/​directory-build-organization/​hierarchy-clobber/​src/​LegacyClient/​Program.cs Reviewed
tests/​dotnet-msbuild/​directory-build-organization/​hierarchy-clobber/​src/​LegacyClient/​LegacyClient.csproj Reviewed
tests/​dotnet-msbuild/​directory-build-organization/​hierarchy-clobber/​src/​Directory.Build.props Reviewed
tests/​dotnet-msbuild/​directory-build-organization/​hierarchy-clobber/​Directory.Build.props Reviewed
tests/​dotnet-msbuild/​check-bin-obj-clash/​ToolReferenceClash.slnx Reviewed
tests/​dotnet-msbuild/​check-bin-obj-clash/​SharedPathClash.slnx Reviewed
tests/​dotnet-msbuild/​check-bin-obj-clash/​safe-shared-base/​SafeSharedBase.slnx Reviewed
tests/​dotnet-msbuild/​check-bin-obj-clash/​safe-shared-base/​LibD/​LibD.csproj Reviewed
tests/​dotnet-msbuild/​check-bin-obj-clash/​safe-shared-base/​LibD/​ClassD.cs Reviewed
tests/​dotnet-msbuild/​check-bin-obj-clash/​safe-shared-base/​LibC/​LibC.csproj Reviewed
tests/​dotnet-msbuild/​check-bin-obj-clash/​safe-shared-base/​LibC/​ClassC.cs Reviewed
tests/​dotnet-msbuild/​check-bin-obj-clash/​global.json Reviewed
tests/​dotnet-msbuild/​check-bin-obj-clash/​default-layout-safe/​DefaultLayoutSafe.slnx Reviewed
tests/​dotnet-msbuild/​check-bin-obj-clash/​default-layout-safe/​AppBeta/​Program.cs Reviewed
tests/​dotnet-msbuild/​check-bin-obj-clash/​default-layout-safe/​AppBeta/​AppBeta.csproj Reviewed
tests/​dotnet-msbuild/​check-bin-obj-clash/​default-layout-safe/​AppAlpha/​Program.cs Reviewed
tests/​dotnet-msbuild/​check-bin-obj-clash/​default-layout-safe/​AppAlpha/​AppAlpha.csproj Reviewed
tests/​dotnet-msbuild/​build-perf-diagnostics/​single-target-domination/​perf-summary.txt Reviewed
tests/​dotnet-msbuild/​build-perf-diagnostics/​runtime-perf-request/​Storefront.csproj Reviewed
tests/​dotnet-msbuild/​build-perf-diagnostics/​runtime-perf-request/​CheckoutController.cs Reviewed
tests/​dotnet-msbuild/​build-perf-diagnostics/​restore-in-build/​perf-summary.txt Reviewed
tests/​dotnet-msbuild/​build-perf-diagnostics/​rar-bottleneck/​TaskContracts/​TaskContracts.csproj Reviewed
tests/​dotnet-msbuild/​build-perf-diagnostics/​rar-bottleneck/​TaskContracts/​ManifestSpec.cs Reviewed
tests/​dotnet-msbuild/​build-perf-diagnostics/​rar-bottleneck/​perf-summary.txt Reviewed
tests/​dotnet-msbuild/​build-perf-diagnostics/​rar-bottleneck/​GatewayService/​GatewayService.csproj Reviewed
tests/​dotnet-msbuild/​build-perf-diagnostics/​rar-bottleneck/​BuildTasks/​ManifestGeneratorTask.cs Reviewed
tests/​dotnet-msbuild/​build-perf-diagnostics/​rar-bottleneck/​BuildTasks/​BuildTasks.csproj Reviewed
tests/​dotnet-msbuild/​build-perf-diagnostics/​offline-analyzers/​global.json Reviewed
tests/​dotnet-msbuild/​build-perf-diagnostics/​offline-analyzers/​AnalyzerTwo/​AnalyzerTwo.csproj Reviewed
tests/​dotnet-msbuild/​build-perf-diagnostics/​offline-analyzers/​AnalyzerTwo/​AnalyzerTwo.cs Reviewed
tests/​dotnet-msbuild/​build-perf-diagnostics/​offline-analyzers/​AnalyzerThree/​AnalyzerThree.csproj Reviewed
tests/​dotnet-msbuild/​build-perf-diagnostics/​offline-analyzers/​AnalyzerThree/​AnalyzerThree.cs Reviewed
tests/​dotnet-msbuild/​build-perf-diagnostics/​offline-analyzers/​AnalyzerOne/​AnalyzerOne.csproj Reviewed
tests/​dotnet-msbuild/​build-perf-diagnostics/​offline-analyzers/​AnalyzerOne/​AnalyzerOne.cs Reviewed
tests/​dotnet-msbuild/​build-perf-diagnostics/​many-small-projects/​perf-summary.txt Reviewed
tests/​dotnet-msbuild/​build-perf-diagnostics/​evaluation-overhead/​perf-summary.txt Reviewed
tests/​dotnet-msbuild/​build-perf-diagnostics/​Directory.Build.props Reviewed
tests/​dotnet-msbuild/​build-perf-diagnostics/​DataService.cs Reviewed
tests/​dotnet-msbuild/​build-perf-diagnostics/​copy-io-bottleneck/​perf-summary.txt Reviewed
tests/​dotnet-msbuild/​build-perf-diagnostics/​Contoso.WebApi.csproj Reviewed
tests/​dotnet-msbuild/​build-perf-baseline/​references/​deterministic-ci-cache.patch Reviewed
tests/​dotnet-msbuild/​build-perf-baseline/​noop-regression/​build-log.txt Reviewed
tests/​dotnet-msbuild/​build-perf-baseline/​node-webpack/​webpack.config.js Reviewed
tests/​dotnet-msbuild/​build-perf-baseline/​node-webpack/​package.json Reviewed
tests/​dotnet-msbuild/​build-perf-baseline/​cold-restore-bound/​build-log.txt Reviewed
tests/​dotnet-msbuild/​build-perf-baseline/​already-optimized/​Directory.Build.props Reviewed
tests/​dotnet-msbuild/​build-perf-baseline/​already-optimized/​Core/​Core.csproj Reviewed
tests/​dotnet-msbuild/​build-perf-baseline/​already-optimized/​Core/​Calculator.cs Reviewed
tests/​dotnet-msbuild/​build-perf-baseline/​already-optimized/​App/​Program.cs Reviewed
tests/​dotnet-msbuild/​build-perf-baseline/​already-optimized/​App/​App.csproj Reviewed
tests/​dotnet-msbuild/​build-parallelism/​wide-slow-leg/​perf-summary.txt Reviewed
tests/​dotnet-msbuild/​build-parallelism/​references/​enable-build-in-parallel.patch Reviewed
tests/​dotnet-msbuild/​build-parallelism/​not-parallel/​Core/​Core.csproj Reviewed
tests/​dotnet-msbuild/​build-parallelism/​not-parallel/​App4/​App4.csproj Reviewed
tests/​dotnet-msbuild/​build-parallelism/​not-parallel/​App3/​App3.csproj Reviewed
tests/​dotnet-msbuild/​build-parallelism/​not-parallel/​App2/​App2.csproj Reviewed
tests/​dotnet-msbuild/​build-parallelism/​not-parallel/​App1/​App1.csproj Reviewed
tests/​dotnet-msbuild/​build-parallelism/​dynamic-graph/​Build.proj Reviewed
tests/​dotnet-msbuild/​build-parallelism/​build-parallel-task/​LibC/​LibC.csproj Reviewed
tests/​dotnet-msbuild/​build-parallelism/​build-parallel-task/​LibB/​LibB.csproj Reviewed
tests/​dotnet-msbuild/​build-parallelism/​build-parallel-task/​LibA/​LibA.csproj Reviewed
tests/​dotnet-msbuild/​build-parallelism/​build-parallel-task/​Core/​Core.csproj Reviewed
tests/​dotnet-msbuild/​build-parallelism/​build-parallel-task/​build.proj Reviewed
tests/​dotnet-msbuild/​binlog-generation/​references/​preserve-binlog-history.json Reviewed
tests/​dotnet-msbuild/​binlog-generation/​references/​powershell-escaping.json Reviewed
tests/​dotnet-msbuild/​binlog-generation/​references/​no-binlog-on-preflight-failure.json Reviewed
tests/​dotnet-msbuild/​binlog-generation/​references/​multi-config-unique-names.json Reviewed
tests/​dotnet-msbuild/​binlog-generation/​references/​fix-build-script.patch Reviewed
tests/​dotnet-msbuild/​binlog-generation/​references/​decline-non-msbuild.json Reviewed
tests/​dotnet-msbuild/​binlog-generation/​references/​bl-placeholder-flag.json Reviewed
tests/​dotnet-msbuild/​binlog-generation/​predictable-filename/​SimpleApp.csproj Reviewed
tests/​dotnet-msbuild/​binlog-generation/​predictable-filename/​Program.cs Reviewed
tests/​dotnet-msbuild/​binlog-generation/​predictable-filename/​ci-artifacts-manifest.txt Reviewed
tests/​dotnet-msbuild/​binlog-generation/​maven-project/​src/​main/​java/​com/​contoso/​reporting/​ReportGenerator.java Reviewed
tests/​dotnet-msbuild/​binlog-generation/​maven-project/​pom.xml Reviewed
tests/​dotnet-msbuild/​binlog-generation/​fix-build-script/​SimpleApp.csproj Reviewed
tests/​dotnet-msbuild/​binlog-generation/​fix-build-script/​Program.cs Reviewed
tests/​dotnet-msbuild/​binlog-generation/​fix-build-script/​build.sh Reviewed
tests/​dotnet-msbuild/​binlog-generation/​failing-invocation/​SimpleApp.csproj Reviewed
tests/​dotnet-msbuild/​binlog-generation/​failing-invocation/​Program.cs Reviewed
tests/​dotnet-msbuild/​binlog-generation/​cleanup-preserve-binlog/​SimpleApp.csproj Reviewed
tests/​dotnet-msbuild/​binlog-generation/​cleanup-preserve-binlog/​Program.cs Reviewed
tests/​dotnet-msbuild/​binlog-failure-analysis/​warning-only-build/​Program.cs Reviewed
tests/​dotnet-msbuild/​binlog-failure-analysis/​warning-only-build/​PriceLib.csproj Reviewed
tests/​dotnet-msbuild/​binlog-failure-analysis/​target-order-tracing/​VersionedApp.csproj Reviewed
tests/​dotnet-msbuild/​binlog-failure-analysis/​target-order-tracing/​Program.cs Reviewed
tests/​dotnet-msbuild/​binlog-failure-analysis/​skip-vs-fail/​ReportTool.csproj Reviewed
tests/​dotnet-msbuild/​binlog-failure-analysis/​skip-vs-fail/​Program.cs Reviewed
tests/​dotnet-msbuild/​binlog-failure-analysis/​references/​boundary-non-msbuild.json Reviewed
tests/​dotnet-msbuild/​binlog-failure-analysis/​references/​boundary-generation-routing.json Reviewed
tests/​dotnet-msbuild/​binlog-failure-analysis/​property-query/​Program.cs Reviewed
tests/​dotnet-msbuild/​binlog-failure-analysis/​property-query/​package-source/​Contoso.Catalog.csproj Reviewed
tests/​dotnet-msbuild/​binlog-failure-analysis/​property-query/​package-source/​CatalogItem.cs Reviewed
tests/​dotnet-msbuild/​binlog-failure-analysis/​property-query/​NuGet.Config Reviewed
tests/​dotnet-msbuild/​binlog-failure-analysis/​property-query/​local-feed/​README.txt Reviewed
tests/​dotnet-msbuild/​binlog-failure-analysis/​property-query/​CatalogTool.csproj Reviewed
tests/​dotnet-msbuild/​binlog-failure-analysis/​gradle-build/​gradle-build-output.txt Reviewed
tests/​dotnet-msbuild/​binlog-failure-analysis/​generation-routing/​Program.cs Reviewed
tests/​dotnet-msbuild/​binlog-failure-analysis/​generation-routing/​InventorySync.csproj Reviewed
tests/​dotnet-msbuild/​binlog-failure-analysis/​fallback-replay/​Program.cs Reviewed
tests/​dotnet-msbuild/​binlog-failure-analysis/​fallback-replay/​OrderCalc.csproj Reviewed
tests/​dotnet-msbuild/​agent.msbuild/​NuGet.Config Reviewed
tests/​dotnet-msbuild/​agent.msbuild/​healthy/​AppTwo/​Program.cs Reviewed
tests/​dotnet-msbuild/​agent.msbuild/​healthy/​AppTwo/​AppTwo.csproj Reviewed
tests/​dotnet-msbuild/​agent.msbuild/​healthy/​AppOne/​Program.cs Reviewed
tests/​dotnet-msbuild/​agent.msbuild/​healthy/​AppOne/​AppOne.csproj Reviewed
tests/​dotnet-msbuild/​agent.msbuild/​empty-feed/​README.txt Reviewed
eng/​vally-adapter/​README.md Reviewed
eng/​skill-validator/​tests/​Evaluate/​MetricsTests.cs Reviewed
eng/​skill-validator/​tests/​Evaluate/​ComparatorTests.cs Reviewed
eng/​skill-validator/​src/​README.md Reviewed
eng/​skill-validator/​src/​Evaluate/​OverfittingCommand.cs Reviewed
eng/​skill-validator/​src/​Evaluate/​Models.cs Reviewed
eng/​skill-validator/​src/​Evaluate/​MetricsCollector.cs Reviewed
eng/​skill-validator/​src/​Evaluate/​LlmSession.cs Reviewed
eng/​skill-validator/​src/​Evaluate/​EvalSchema.cs Reviewed
eng/​eval-quality/​underpowered-allowlist.txt Reviewed
.github/​workflows/​eval-quality.yml Reviewed

💡 Add a code-review agent skill or configure MCP servers for context-aware, tailored reviews. Learn more in the docs.

Comment thread .github/workflows/eval-quality.yml
Comment thread eng/evaluation/find-targets.ps1 Outdated
Comment thread eng/skill-validator/src/Evaluate/RejudgeCommand.cs Outdated
Comment thread eng/vally-adapter/adapt.mjs
Comment thread tests/dotnet-msbuild/target-authoring/references/fix-antipatterns.patch Outdated
Comment thread eng/skill-validator/src/Evaluate/RejudgeCommand.cs Outdated
Copilot AI review requested due to automatic review settings September 25, 2026 07:56

Copilot AI left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Copilot review overview

🟡 Changes recommended

Unresolved moderate and critical findings affect shard selection, isolation, cleanup, rejudge correctness, and activation-contract enforcement.

Get a fresh assessment by requesting another Copilot review.

Review effort: Lite
Findings: 6 High severity · 1 Medium severity

Open (7)

Comment thread eng/skill-validator/src/Evaluate/AgentRunner.cs
@AbhitejJohn
AbhitejJohn added this pull request to stack #1216 September 25, 2026 08:14
@github-actions github-actions Bot added the waiting-on-author PR state label label Sep 25, 2026
@github-actions

Copy link
Copy Markdown
Contributor

👋 @AbhitejJohn — this PR has 7 unresolved review thread(s). When you're ready, please address the feedback and push an update; the triage bot will pick up the next state automatically. (Add the no-stale label to silence further pings.)

Audit checker changes across all eval specs, enforce exact rejudge pairing, isolate NuGet fallbacks, and repair shard, dormancy, and target-authoring evidence.

Co-authored-by: Copilot App <223556219+Copilot@users.noreply.github.com>
Copilot AI review requested due to automatic review settings September 25, 2026 20:35
Keep unmatched dormancy annotations as explicit failures without adding them twice to activation contract totals.

Co-authored-by: Copilot App <223556219+Copilot@users.noreply.github.com>

@AbhitejJohn AbhitejJohn left a comment

Copy link
Copy Markdown
Collaborator Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

/evaluate ba899ab

Copilot AI left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Comment thread eng/skill-validator/src/Evaluate/AgentRunner.cs Outdated
Comment thread eng/skill-validator/src/Evaluate/EvaluateCommand.cs
Comment thread eng/skill-validator/src/Evaluate/RejudgeCommand.cs Outdated
Comment thread eng/skill-validator/src/Evaluate/RejudgeCommand.cs
Comment thread eng/skill-validator/src/Evaluate/RejudgeCommand.cs
Comment thread tests/dotnet-msbuild/build-perf-diagnostics/eval.yaml
Copilot AI review requested due to automatic review settings September 25, 2026 20:43
@AbhitejJohn
AbhitejJohn deployed to copilot-pat-pool September 25, 2026 20:46 — with GitHub Actions Active
@AbhitejJohn
AbhitejJohn deployed to copilot-pat-pool September 25, 2026 20:46 — with GitHub Actions Active
@AbhitejJohn
AbhitejJohn deployed to copilot-pat-pool September 25, 2026 20:46 — with GitHub Actions Active
@github-actions

Copy link
Copy Markdown
Contributor

📊 Skill and Agent Evaluation Results

36 model/target results across 18 targets and 2 models — ✅ 2 improved, ➖ 30 not proven improved, ⚠️ 0 invalid or underpowered, ⛔ 3 activation contract failures, 📉 1 preference losses (report only).

Measurement identity: evaluated commit 1e90581d198eb68a2357f0c411be6643ee7dff40; 2 judge models.

Measurement health: 36 expected / 36 observed / 36 written; 0 missing, 0 unexpected, 0 invalid; 0 recovered comparison error slots and 0 unresolved comparison error slots.

Objective completion gate: not enabled. Aggregate completion transitions are telemetry only, so this report does not claim that zero objective regressions were proven.

A result passes only when preference-eligible distinct-stimulus votes have aggregate net win of at least 20% and an exact one-sided sign-test result of p ≤ 0.05, and every explicit dormancy activation contract passes. Repeated runs measure reliability only.

Target Model Verdict Gate evidence Overfit Warnings Next action
agent.msbuild claude-sonnet-5 ➖ Not proven improved n=5; 1W/1T/3L; d=4; p=0.312; net -40.0% — Activation: isolated 1/5; plugin 0/5 Inspect tied or lost stimuli; predeclare added breadth before a new experiment.
agent.msbuild gpt-5.6-luna ➖ Not proven improved n=5; 1W/2T/2L; d=3; p=0.500; net -20.0% — Activation: isolated 1/5; plugin 1/5 Inspect tied or lost stimuli; predeclare added breadth before a new experiment.
binlog-failure-analysis claude-sonnet-5 📉 Preference loss (report only) n=6; 0W/1T/5L; d=5; p=0.031; net -83.3%; 1 dormancy excluded 🟡 0.21 Activation: isolated 5/7; plugin 6/7 Inspect losing stimuli and fix skill behavior; this is not objective completion proof.
binlog-failure-analysis gpt-5.6-luna ✅ Improved n=6; 5W/1T/0L; d=5; p=0.031; net +83.3%; 1 dormancy excluded ✅ 0.17 Activation: isolated 6/7; plugin 6/7 Fix activation gaps.
binlog-generation claude-sonnet-5 ➖ Not proven improved n=7; 4W/1T/2L; d=6; p=0.344; net +28.6%; 1 dormancy excluded 🔴 0.50 Activation: isolated 5/7; plugin 6/7; Activation-only stop: plugin 1 failed run Inspect activation-only failed runs before rewriting skill content; the model stopped after loading a skill.
binlog-generation gpt-5.6-luna ➖ Not proven improved n=7; 4W/2T/1L; d=5; p=0.188; net +42.9%; 1 dormancy excluded 🟡 0.35 Activation: isolated 6/7; plugin 5/7 Inspect tied or lost stimuli and fix inconsistent skill behavior.
build-parallelism claude-sonnet-5 ➖ Not proven improved n=6; 0W/2T/4L; d=4; p=0.063; net -66.7%; 1 dormancy excluded ✅ 0.17 Activation: isolated 5/6; plugin 5/6; Activation-only stop: isolated 1 failed run; Activation-only stop: plugin 1 failed run Inspect activation-only failed runs before rewriting skill content; the model stopped after loading a skill.
build-parallelism gpt-5.6-luna ➖ Not proven improved n=6; 3W/1T/2L; d=5; p=0.500; net +16.7%; 1 dormancy excluded 🟡 0.32 Activation-only stop: isolated 1 failed run; Activation-only stop: plugin 1 failed run Inspect activation-only failed runs before rewriting skill content; the model stopped after loading a skill.
build-perf-baseline claude-sonnet-5 ➖ Not proven improved n=6; 4W/2T/0L; d=4; p=0.063; net +66.7%; 1 dormancy excluded 🟡 0.32 Activation: isolated 5/6; plugin 4/6 Inspect tied or lost stimuli; predeclare added breadth before a new experiment.
build-perf-baseline gpt-5.6-luna ⛔ Activation contract failed n=6; 3W/0T/3L; d=6; p=0.656; net +0.0%; 1 dormancy excluded 🔴 0.56 Dormancy contract: 1 unexpected activation(s); Activation: isolated 4/6; plugin 6/6 Narrow skill routing so the listed off-target scenarios stay dormant.
build-perf-diagnostics claude-sonnet-5 ➖ Not proven improved n=7; 4W/0T/3L; d=7; p=0.500; net +14.3%; 1 dormancy excluded 🔴 0.50 — Inspect tied or lost stimuli and fix inconsistent skill behavior.
build-perf-diagnostics gpt-5.6-luna ➖ Not proven improved n=7; 3W/3T/1L; d=4; p=0.312; net +28.6%; 1 dormancy excluded 🟡 0.40 — Inspect tied or lost stimuli; predeclare added breadth before a new experiment.
check-bin-obj-clash claude-sonnet-5 ➖ Not proven improved n=7; 2W/5T/0L; d=2; p=0.250; net +28.6% 🟡 0.24 Activation: isolated 6/7; plugin 6/7 Inspect tied or lost stimuli; predeclare added breadth before a new experiment.
check-bin-obj-clash gpt-5.6-luna ➖ Not proven improved n=7; 4W/2T/1L; d=5; p=0.188; net +42.9% 🟡 0.21 Activation: isolated 7/7; plugin 6/7 Inspect tied or lost stimuli and fix inconsistent skill behavior.
directory-build-organization claude-sonnet-5 ⛔ Activation contract failed n=6; 4W/2T/0L; d=4; p=0.063; net +66.7%; 1 dormancy excluded 🟡 0.43 Dormancy contract: 1 unexpected activation(s); Activation: isolated 5/6; plugin 5/6 Narrow skill routing so the listed off-target scenarios stay dormant.
directory-build-organization gpt-5.6-luna ➖ Not proven improved n=6; 4W/2T/0L; d=4; p=0.063; net +66.7%; 1 dormancy excluded 🔴 0.53 — Inspect tied or lost stimuli; predeclare added breadth before a new experiment.
eval-performance claude-sonnet-5 ➖ Not proven improved n=8; 4W/2T/2L; d=6; p=0.344; net +25.0% 🟡 0.35 Activation: isolated 6/8; plugin 7/8 Inspect tied or lost stimuli and fix inconsistent skill behavior.
eval-performance gpt-5.6-luna ➖ Not proven improved n=8; 3W/4T/1L; d=4; p=0.312; net +25.0% 🟡 0.49 — Inspect tied or lost stimuli; predeclare added breadth before a new experiment.
extension-points claude-sonnet-5 ➖ Not proven improved n=7; 3W/2T/2L; d=5; p=0.500; net +14.3%; 2 dormancy excluded 🟡 0.32 Activation: isolated 6/7; plugin 5/7 Inspect tied or lost stimuli and fix inconsistent skill behavior.
extension-points gpt-5.6-luna ➖ Not proven improved n=7; 3W/2T/2L; d=5; p=0.500; net +14.3%; 2 dormancy excluded 🟡 0.33 Activation: isolated 7/7; plugin 6/7 Inspect tied or lost stimuli and fix inconsistent skill behavior.
including-generated-files claude-sonnet-5 ➖ Not proven improved n=7; 2W/2T/3L; d=5; p=0.500; net -14.3%; 1 dormancy excluded 🟡 0.42 Activation: isolated 6/7; plugin 5/7 Inspect tied or lost stimuli and fix inconsistent skill behavior.
including-generated-files gpt-5.6-luna ➖ Not proven improved n=7; 2W/4T/1L; d=3; p=0.500; net +14.3%; 1 dormancy excluded 🟡 0.39 Activation: isolated 6/7; plugin 7/7 Inspect tied or lost stimuli; predeclare added breadth before a new experiment.
incremental-build claude-sonnet-5 ✅ Improved n=9; 6W/3T/0L; d=6; p=0.016; net +66.7% 🟡 0.33 Activation: isolated 3/9; plugin 7/9 Fix activation gaps; Review overfit evidence.
incremental-build gpt-5.6-luna ➖ Not proven improved n=9; 3W/4T/2L; d=5; p=0.500; net +11.1% 🟡 0.34 Activation: isolated 7/9; plugin 7/9 Inspect tied or lost stimuli and fix inconsistent skill behavior.
item-management claude-sonnet-5 ➖ Not proven improved n=6; 1W/5T/0L; d=1; p=0.500; net +16.7%; 1 dormancy excluded 🟡 0.38 Activation: isolated 3/6; plugin 2/6 Inspect tied or lost stimuli; predeclare added breadth before a new experiment.
item-management gpt-5.6-luna ➖ Not proven improved n=6; 1W/4T/1L; d=2; p=0.750; net +0.0%; 1 dormancy excluded 🟡 0.47 — Inspect tied or lost stimuli; predeclare added breadth before a new experiment.
msbuild-antipatterns claude-sonnet-5 ➖ Not proven improved n=7; 4W/1T/2L; d=6; p=0.344; net +28.6%; 1 dormancy excluded 🟡 0.29 Activation: isolated 2/7; plugin 1/7 Inspect tied or lost stimuli and fix inconsistent skill behavior.
msbuild-antipatterns gpt-5.6-luna ⛔ Activation contract failed n=7; 1W/4T/2L; d=3; p=0.500; net -14.3%; 1 dormancy excluded 🟡 0.31 Dormancy contract: 1 unexpected activation(s); Activation: isolated 5/7; plugin 4/7 Narrow skill routing so the listed off-target scenarios stay dormant.
msbuild-modernization claude-sonnet-5 ➖ Not proven improved n=6; 2W/3T/1L; d=3; p=0.500; net +16.7%; 1 dormancy excluded 🟡 0.31 Activation: isolated 6/6; plugin 5/6 Inspect tied or lost stimuli; predeclare added breadth before a new experiment.
msbuild-modernization gpt-5.6-luna ➖ Not proven improved n=6; 2W/2T/2L; d=4; p=0.687; net +0.0%; 1 dormancy excluded 🟡 0.35 Activation: isolated 5/6; plugin 6/6 Inspect tied or lost stimuli; predeclare added breadth before a new experiment.
property-patterns claude-sonnet-5 ➖ Not proven improved n=7; 1W/2T/4L; d=5; p=0.188; net -42.9% 🟡 0.31 Activation: isolated 4/7; plugin 6/7 Inspect tied or lost stimuli and fix inconsistent skill behavior.
property-patterns gpt-5.6-luna ➖ Not proven improved n=7; 0W/4T/3L; d=3; p=0.125; net -42.9% 🟡 0.27 Activation: isolated 6/7; plugin 6/7 Inspect tied or lost stimuli; predeclare added breadth before a new experiment.
resolve-project-references claude-sonnet-5 ➖ Not proven improved n=6; 4W/1T/1L; d=5; p=0.188; net +50.0%; 2 dormancy excluded 🟡 0.23 Activation: isolated 5/6; plugin 2/6 Inspect tied or lost stimuli and fix inconsistent skill behavior.
resolve-project-references gpt-5.6-luna ➖ Not proven improved n=6; 2W/4T/0L; d=2; p=0.250; net +33.3%; 2 dormancy excluded 🟡 0.36 Activation-only stop: plugin 1 failed run Inspect activation-only failed runs before rewriting skill content; the model stopped after loading a skill.
target-authoring claude-sonnet-5 ➖ Not proven improved n=6; 3W/2T/1L; d=4; p=0.312; net +33.3%; 1 dormancy excluded 🟡 0.29 Activation: isolated 5/6; plugin 4/6 Inspect tied or lost stimuli; predeclare added breadth before a new experiment.
target-authoring gpt-5.6-luna ➖ Not proven improved n=6; 1W/4T/1L; d=2; p=0.750; net +0.0%; 1 dormancy excluded ✅ 0.18 Activation: isolated 6/6; plugin 5/6 Inspect tied or lost stimuli; predeclare added breadth before a new experiment.
ℹ️ How to read this report
  • ✅ Improved — the result passed both the statistical gate and the 20% practical net-win floor.
  • ➖ Not proven improved — the result is valid but did not pass both gates. This is not automatically a regression.
  • ⚠️ Invalid / underpowered — the gate withheld a quality verdict. Fix the measurement before judging the target.
  • ⛔ Activation contract failed — the isolated target activated on an explicit dormancy scenario. Dormancy preference is excluded, but this routing failure still blocks a pass.
  • 📉 Preference loss — the LLM judge credibly preferred baseline. It is report-only, not objective completion proof.
  • Gate evidence — n preference-eligible distinct-stimulus votes, W/T/L stimulus votes, d discordant votes, exact one-sided p, net win, and the count of separately retained dormancy stimuli. The p value applies to one model/target result; no matrix-wide multiple-comparison correction is applied.
  • Overfit — overfitting-judge severity (✅ Low, 🟡 Moderate, 🔴 High, — none) and score.
  • Warnings — activation, timeout, retry recovery, or unresolved comparison conditions that need attention.
  • Do not add repeated runs to increase statistical power. Do not add stimuli after seeing a near-pass unless the new breadth is predeclared for a new experiment.
⛔ Activation contract failed — build-perf-baseline (gpt-5.6-luna)

Why: Net win +0.0% (3W/0T/3L over 6 preference-eligible stimulus vote(s), sign test p=0.656), mean preference -8.6% across 7 paired run(s), 1 dormancy stimulus/stimuli excluded from preference — activation contract failed (1 explicit dormancy scenario(s) activated the isolated target skill)

Next action: Narrow skill routing so the listed off-target scenarios stay dormant.

State: VALID_NO_CHANGE (activation_contract_failed)

Gate evidence: n=6; 3W/0T/3L; d=6; p=0.656; net +0.0%; 1 dormancy excluded

Warnings: Dormancy contract: 1 unexpected activation(s); Activation: isolated 4/6; plugin 6/6

Overfit: High (score 0.56)

Repeated-run reliability (not used by the gate): 7 paired runs (3W/1T/3L).

Weak or warning scenarios:

Scenario Preference gate Net win Δ Pref Runs (W/T/L)
▼ Configure deterministic, cache-safe CI builds Eligible -100.0% -100.0% 0/0/1
= Decline a non-MSBuild build performance request Excluded (activation contract) +0.0% +0.0% 0/1/0
▼ Leave an already-optimized build unchanged Eligible -100.0% -40.0% 0/0/1
▼ Route a restore-bound cold build away from architecture changes Eligible -100.0% -40.0% 0/0/1

Illustrative judge evidence:

  • Configure deterministic, cache-safe CI builds: Response A implements the correct solution for the task by adding ContinuousIntegrationBuild, which is specifically designed for CI reproducibility and safe cross-agent caching. Response B adds a redundant property (Deterministic=true) that is already default in modern SDK...

This is one example, not the aggregate verdict. Open Full Results for every judgment.

⛔ Activation contract failed — directory-build-organization (claude-sonnet-5)

Why: Net win +66.7% (4W/2T/0L over 6 preference-eligible stimulus vote(s), sign test p=0.063), mean preference +34.3% across 7 paired run(s), 1 dormancy stimulus/stimuli excluded from preference — activation contract failed (1 explicit dormancy scenario(s) activated the isolated target skill)

Next action: Narrow skill routing so the listed off-target scenarios stay dormant.

State: VALID_NO_CHANGE (activation_contract_failed)

Gate evidence: n=6; 4W/2T/0L; d=4; p=0.063; net +66.7%; 1 dormancy excluded

Warnings: Dormancy contract: 1 unexpected activation(s); Activation: isolated 5/6; plugin 5/6

Overfit: Moderate (score 0.43)

Repeated-run reliability (not used by the gate): 7 paired runs (4W/2T/1L).

Weak or warning scenarios:

Scenario Preference gate Net win Δ Pref Runs (W/T/L)
▼ Decline adding shared build files to a lone project Excluded (activation contract) -100.0% -40.0% 0/0/1
= Diagnose a TargetFramework condition that silently skips in props Eligible +0.0% +0.0% 0/1/0
= Diagnose an inner shared-props file that overwrites its own override Eligible +0.0% +0.0% 0/1/0

Illustrative judge evidence:

  • Diagnose a TargetFramework condition that silently skips in props: Position-swap inconsistent (forward: tie, reverse: A). Defaulting to tie.

This is one example, not the aggregate verdict. Open Full Results for every judgment.

⛔ Activation contract failed — msbuild-antipatterns (gpt-5.6-luna)

Why: Net win -14.3% (1W/4T/2L over 7 preference-eligible stimulus vote(s), sign test p=0.500), mean preference -12.5% across 8 paired run(s), 1 dormancy stimulus/stimuli excluded from preference — activation contract failed (1 explicit dormancy scenario(s) activated the isolated target skill)

Next action: Narrow skill routing so the listed off-target scenarios stay dormant.

State: VALID_NO_CHANGE (activation_contract_failed)

Gate evidence: n=7; 1W/4T/2L; d=3; p=0.500; net -14.3%; 1 dormancy excluded

Warnings: Dormancy contract: 1 unexpected activation(s); Activation: isolated 5/7; plugin 4/7

Overfit: Moderate (score 0.31)

Repeated-run reliability (not used by the gate): 8 paired runs (1W/5T/2L).

Weak or warning scenarios:

Scenario Preference gate Net win Δ Pref Runs (W/T/L)
▲ Add a module to an F# project Eligible +100.0% +40.0% 1/0/0
= Add a signature file to define public API Eligible +0.0% +0.0% 0/1/0
▼ Distinguish a style backslash from a real cross-platform backslash bug Eligible -100.0% -40.0% 0/0/1
= Fix broken file order causing FS0039 Eligible +0.0% +0.0% 0/1/0
= Judge an unguarded import inside a NuGet package build folder Eligible +0.0% +0.0% 0/1/0
= Leave a clean project without inventing anti-patterns Eligible +0.0% +0.0% 0/1/0
= Non-activation: migrate a legacy project to SDK style Excluded (activation contract) +0.0% +0.0% 0/1/0
▼ Review MSBuild files for anti-patterns and style issues Eligible -100.0% -100.0% 0/0/1

Illustrative judge evidence:

  • Add a signature file to define public API: Position-swap inconsistent (forward: tie, reverse: A). Defaulting to tie.

This is one example, not the aggregate verdict. Open Full Results for every judgment.

📉 Preference loss (report only) — binlog-failure-analysis (claude-sonnet-5)

Why: Net win -83.3% (0W/1T/5L over 6 preference-eligible stimulus vote(s), sign test p=0.031), mean preference -42.9% across 7 paired run(s), 1 dormancy stimulus/stimuli excluded from preference — credibly worse

Next action: Inspect losing stimuli and fix skill behavior; this is not objective completion proof.

State: VALID_NO_CHANGE (preference_regression_report_only)

Gate evidence: n=6; 0W/1T/5L; d=5; p=0.031; net -83.3%; 1 dormancy excluded

Warnings: Activation: isolated 5/7; plugin 6/7

Overfit: Moderate (score 0.21)

Repeated-run reliability (not used by the gate): 7 paired runs (0W/1T/6L).

Weak or warning scenarios:

Scenario Preference gate Net win Δ Pref Runs (W/T/L)
▼ Assess a requested binlog investigation when the current build is healthy Eligible -100.0% -40.0% 0/0/1
= Confirm the actual resolved target framework and package version from a binlog Eligible +0.0% +0.0% 0/0/0
= Determine whether a quiet second build actually failed Eligible +0.0% +0.0% 0/1/0
▼ Diagnose a warning behind a build that actually succeeded Eligible -100.0% -40.0% 0/0/1
▼ Diagnose build failures from binlog only (no source files) Eligible -100.0% -40.0% 0/0/1
▼ Stay dormant for a non-MSBuild build failure log Excluded (activation contract) -100.0% -40.0% 0/0/1
▼ Trace why a generated source file is missing at compile time Eligible -100.0% -100.0% 0/0/1
▼ Use capture-time text logs when the binlog reader is unavailable Eligible -100.0% -40.0% 0/0/1

Illustrative judge evidence:

  • Assess a requested binlog investigation when the current build is healthy: Both provide the essential, correct conclusion and avoid fabricating issues. A is modestly stronger because it performs and reports a substantive fallback inspection of the captured binlog rather than relying only on the build console summary.

This is one example, not the aggregate verdict. Open Full Results for every judgment.

➖ Not proven improved — agent.msbuild (claude-sonnet-5)

Why: Net win -40.0% (1W/1T/3L over 5 preference-eligible stimulus vote(s), sign test p=0.312), mean preference -16.0% across 5 paired run(s) — no improvement — native evaluator reported that the target agent did not activate

Next action: Inspect tied or lost stimuli; predeclare added breadth before a new experiment.

State: VALID_NO_CHANGE (target_agent_not_activated)

Gate evidence: n=5; 1W/1T/3L; d=4; p=0.312; net -40.0%

Warnings: Activation: isolated 1/5; plugin 0/5

Repeated-run reliability (not used by the gate): 5 paired runs (1W/1T/3L).

Weak or warning scenarios:

Scenario Preference gate Net win Δ Pref Runs (W/T/L)
= Advise on project file organization Eligible +0.0% +0.0% 0/1/0
▼ Diagnose broken incremental build behavior Eligible -100.0% -40.0% 0/0/1
▼ Review a project file for maintainability risks Eligible -100.0% -40.0% 0/0/1
▼ Route a slow build to performance analysis Eligible -100.0% -40.0% 0/0/1
▲ Triage a build failure and route to appropriate analysis Eligible +100.0% +40.0% 1/0/0

Illustrative judge evidence:

  • Advise on project file organization: Position-swap inconsistent (forward: baseline, reverse: skill). Defaulting to tie.

This is one example, not the aggregate verdict. Open Full Results for every judgment.

➖ Not proven improved — agent.msbuild (gpt-5.6-luna)

Why: Net win -20.0% (1W/2T/2L over 5 preference-eligible stimulus vote(s), sign test p=0.500), mean preference -8.0% across 5 paired run(s) — no improvement — native evaluator reported that the target agent did not activate

Next action: Inspect tied or lost stimuli; predeclare added breadth before a new experiment.

State: VALID_NO_CHANGE (target_agent_not_activated)

Gate evidence: n=5; 1W/2T/2L; d=3; p=0.500; net -20.0%

Warnings: Activation: isolated 1/5; plugin 1/5

Repeated-run reliability (not used by the gate): 5 paired runs (1W/2T/2L).

Weak or warning scenarios:

Scenario Preference gate Net win Δ Pref Runs (W/T/L)
▼ Advise on project file organization Eligible -100.0% -40.0% 0/0/1
= Diagnose broken incremental build behavior Eligible +0.0% +0.0% 0/1/0
= Review a project file for maintainability risks Eligible +0.0% +0.0% 0/1/0
▲ Route a slow build to performance analysis Eligible +100.0% +40.0% 1/0/0
▼ Triage a build failure and route to appropriate analysis Eligible -100.0% -40.0% 0/0/1

Illustrative judge evidence:

  • Advise on project file organization: While both responses arrive at identical technical solutions and equally fail to meet the educational/explanation criteria, Response A demonstrates superior process quality. Response A took a direct, efficient approach with 6 errors mostly related to tool mechanics, while Resp...

This is one example, not the aggregate verdict. Open Full Results for every judgment.

➖ Not proven improved — binlog-generation (claude-sonnet-5)

Why: Net win +28.6% (4W/1T/2L over 7 preference-eligible stimulus vote(s), sign test p=0.344), mean preference +32.5% across 8 paired run(s), 1 dormancy stimulus/stimuli excluded from preference — not credible (sign test p=0.344 > 0.05)

Next action: Inspect activation-only failed runs before rewriting skill content; the model stopped after loading a skill.

State: VALID_NO_CHANGE (no_credible_preference_change)

Gate evidence: n=7; 4W/1T/2L; d=6; p=0.344; net +28.6%; 1 dormancy excluded

Warnings: Activation: isolated 5/7; plugin 6/7; Activation-only stop: plugin 1 failed run

Overfit: High (score 0.50)

Repeated-run reliability (not used by the gate): 8 paired runs (4W/2T/2L).

Weak or warning scenarios:

Scenario Preference gate Net win Δ Pref Runs (W/T/L)
▼ Build with /bl in PowerShell Eligible -100.0% -40.0% 0/0/1
= Choose a predictable, non-colliding binlog name for a CI upload step Eligible +0.0% +0.0% 0/1/0
= Decline Maven compiler configuration Excluded (activation contract) +0.0% +0.0% 0/1/0
▲ Preserve binlog history while cleaning stale build output Eligible +100.0% +40.0% 1/0/0
▼ Recognize that a failed build produced no binlog at all Eligible -100.0% -40.0% 0/0/1

Illustrative judge evidence:

  • Build with /bl in PowerShell: Both runs successfully build through PowerShell and create diagnostic binlogs. A is better because its logged command is clear, conventional, and does not present the doubled-brace form explicitly prohibited by the rubric; it also reports a predictable artifact name. Neither r...

This is one example, not the aggregate verdict. Open Full Results for every judgment.

➖ Not proven improved — binlog-generation (gpt-5.6-luna)

Why: Net win +42.9% (4W/2T/1L over 7 preference-eligible stimulus vote(s), sign test p=0.188), mean preference +37.5% across 8 paired run(s), 1 dormancy stimulus/stimuli excluded from preference — not credible (sign test p=0.188 > 0.05)

Next action: Inspect tied or lost stimuli and fix inconsistent skill behavior.

State: VALID_NO_CHANGE (no_credible_preference_change)

Gate evidence: n=7; 4W/2T/1L; d=5; p=0.188; net +42.9%; 1 dormancy excluded

Warnings: Activation: isolated 6/7; plugin 5/7

Overfit: Moderate (score 0.35)

Repeated-run reliability (not used by the gate): 8 paired runs (4W/3T/1L).

Weak or warning scenarios:

Scenario Preference gate Net win Δ Pref Runs (W/T/L)
= Build multiple configurations with unique binlogs Eligible +0.0% +0.0% 0/1/0
= Choose a predictable, non-colliding binlog name for a CI upload step Eligible +0.0% +0.0% 0/1/0
= Decline Maven compiler configuration Excluded (activation contract) +0.0% +0.0% 0/1/0
▲ Fix a CI build script that reuses the same binlog on every run Eligible +100.0% +100.0% 1/0/0
▲ Preserve binlog history while cleaning stale build output Eligible +100.0% +100.0% 1/0/0
▼ Recognize that a failed build produced no binlog at all Eligible -100.0% -40.0% 0/0/1

Illustrative judge evidence:

  • Build multiple configurations with unique binlogs: Both responses failed to meet the specific rubric requirements by not using the {} placeholder for automatic unique binlog naming. While both successfully completed the practical task of building in Debug and Release configurations and creating separate binlog files, they both...

This is one example, not the aggregate verdict. Open Full Results for every judgment.

➖ Not proven improved — build-parallelism (claude-sonnet-5)

Why: Net win -66.7% (0W/2T/4L over 6 preference-eligible stimulus vote(s), sign test p=0.063), mean preference -22.9% across 7 paired run(s), 1 dormancy stimulus/stimuli excluded from preference — no improvement

Next action: Inspect activation-only failed runs before rewriting skill content; the model stopped after loading a skill.

State: VALID_NO_CHANGE (no_credible_preference_change)

Gate evidence: n=6; 0W/2T/4L; d=4; p=0.063; net -66.7%; 1 dormancy excluded

Warnings: Activation: isolated 5/6; plugin 5/6; Activation-only stop: isolated 1 failed run; Activation-only stop: plugin 1 failed run

Overfit: Low (score 0.17)

Repeated-run reliability (not used by the gate): 7 paired runs (0W/3T/4L).

Weak or warning scenarios:

Scenario Preference gate Net win Δ Pref Runs (W/T/L)
▼ Analyze build parallelism bottlenecks Eligible -100.0% -40.0% 0/0/1
= Decline database query tuning request Excluded (activation contract) +0.0% +0.0% 0/1/0
▼ Decline graph build for runtime-discovered projects Eligible -100.0% -40.0% 0/0/1
= Enable BuildInParallel on a custom MSBuild task Eligible +0.0% +0.0% 0/1/0
= Enable parallel nodes for a wide project graph Eligible +0.0% +0.0% 0/1/0
▼ Preserve valid dependencies and optimize the slow critical-path project Eligible -100.0% -40.0% 0/0/1
▼ Reduce CI build scope with a solution filter Eligible -100.0% -40.0% 0/0/1

Illustrative judge evidence:

  • Analyze build parallelism bottlenecks: Both answers are substantively correct and answer all requested questions. A is better because it presents precise binlog-derived task timing evidence and uses it to support the critical-path and redundant-edge conclusion; B's timing justification is less direct and its approx...

This is one example, not the aggregate verdict. Open Full Results for every judgment.

➖ Not proven improved — build-parallelism (gpt-5.6-luna)

Why: Net win +16.7% (3W/1T/2L over 6 preference-eligible stimulus vote(s), sign test p=0.500), mean preference +5.7% across 7 paired run(s), 1 dormancy stimulus/stimuli excluded from preference — not credible (sign test p=0.500 > 0.05)

Next action: Inspect activation-only failed runs before rewriting skill content; the model stopped after loading a skill.

State: VALID_NO_CHANGE (no_credible_preference_change)

Gate evidence: n=6; 3W/1T/2L; d=5; p=0.500; net +16.7%; 1 dormancy excluded

Warnings: Activation-only stop: isolated 1 failed run; Activation-only stop: plugin 1 failed run

Overfit: Moderate (score 0.32)

Repeated-run reliability (not used by the gate): 7 paired runs (3W/2T/2L).

Weak or warning scenarios:

Scenario Preference gate Net win Δ Pref Runs (W/T/L)
= Decline database query tuning request Excluded (activation contract) +0.0% +0.0% 0/1/0
▼ Enable parallel nodes for a wide project graph Eligible -100.0% -40.0% 0/0/1
= Preserve valid dependencies and optimize the slow critical-path project Eligible +0.0% +0.0% 0/1/0
▼ Reduce CI build scope with a solution filter Eligible -100.0% -40.0% 0/0/1

Illustrative judge evidence:

  • Enable parallel nodes for a wide project graph: Both responses correctly diagnose the problem (the -m:1 flag limiting to one worker node) and recommend the correct solution (using -m for automatic CPU count). Response A is slightly better because it provides two command examples (-m and -m:4), making the solution more compr...

This is one example, not the aggregate verdict. Open Full Results for every judgment.

➖ Not proven improved — build-perf-baseline (claude-sonnet-5)

Why: Net win +66.7% (4W/2T/0L over 6 preference-eligible stimulus vote(s), sign test p=0.063), mean preference +37.1% across 7 paired run(s), 1 dormancy stimulus/stimuli excluded from preference — not credible — 2 of 6 preference-eligible stimulus vote(s) tied, leaving only 4 discordant preference vote(s). The sign test conditions on non-tie stimulus votes and cannot reach 0.05 below 5, so no record could have passed here — this is not a measured null. Either the skill is inert on these scenarios (make them discriminate) or the eval needs more distinct stimuli to clear the ties

Next action: Inspect tied or lost stimuli; predeclare added breadth before a new experiment.

State: VALID_NO_CHANGE (no_credible_preference_change)

Gate evidence: n=6; 4W/2T/0L; d=4; p=0.063; net +66.7%; 1 dormancy excluded

Warnings: Activation: isolated 5/6; plugin 4/6

Overfit: Moderate (score 0.32)

Repeated-run reliability (not used by the gate): 7 paired runs (5W/2T/0L).

Weak or warning scenarios:

Scenario Preference gate Net win Δ Pref Runs (W/T/L)
▲ Configure deterministic, cache-safe CI builds Eligible +100.0% +100.0% 1/0/0
= Establish build performance baseline and recommend optimizations Eligible +0.0% +0.0% 0/1/0
▲ Route a broken no-op rebuild away from generic optimization Eligible +100.0% +40.0% 1/0/0
= Route a restore-bound cold build away from architecture changes Eligible +0.0% +0.0% 0/1/0

Illustrative judge evidence:

  • Establish build performance baseline and recommend optimizations: Position-swap inconsistent (forward: A, reverse: B). Defaulting to tie.

This is one example, not the aggregate verdict. Open Full Results for every judgment.

➖ Not proven improved — build-perf-diagnostics (claude-sonnet-5)

Why: Net win +14.3% (4W/0T/3L over 7 preference-eligible stimulus vote(s), sign test p=0.500), mean preference +5.0% across 8 paired run(s), 1 dormancy stimulus/stimuli excluded from preference — not credible (sign test p=0.500 > 0.05)

Next action: Inspect tied or lost stimuli and fix inconsistent skill behavior.

State: VALID_NO_CHANGE (no_credible_preference_change)

Gate evidence: n=7; 4W/0T/3L; d=7; p=0.500; net +14.3%; 1 dormancy excluded

Overfit: High (score 0.50)

Repeated-run reliability (not used by the gate): 8 paired runs (4W/1T/3L).

Weak or warning scenarios:

Scenario Preference gate Net win Δ Pref Runs (W/T/L)
= Decline a runtime latency request that is not a build performance issue Excluded (activation contract) +0.0% +0.0% 0/1/0
▼ Diagnose a Copy task dominating build time Eligible -100.0% -40.0% 0/0/1
▼ Diagnose a pathological ResolveAssemblyReference time Eligible -100.0% -40.0% 0/0/1
▼ Diagnose per-project overhead across many small projects Eligible -100.0% -40.0% 0/0/1

Illustrative judge evidence:

  • Diagnose a Copy task dominating build time: Both give the central diagnosis and the highest-impact PreserveNewest fix. A is overall stronger because it names the relevant additional-file hard-link setting, whereas B explicitly rejects that relevant control and recommends a different property; B's somewhat better Dev Dri...

This is one example, not the aggregate verdict. Open Full Results for every judgment.

➖ Not proven improved — build-perf-diagnostics (gpt-5.6-luna)

Why: Net win +28.6% (3W/3T/1L over 7 preference-eligible stimulus vote(s), sign test p=0.312), mean preference +15.0% across 8 paired run(s), 1 dormancy stimulus/stimuli excluded from preference — not credible — 3 of 7 preference-eligible stimulus vote(s) tied, leaving only 4 discordant preference vote(s). The sign test conditions on non-tie stimulus votes and cannot reach 0.05 below 5, so no record could have passed here — this is not a measured null. Either the skill is inert on these scenarios (make them discriminate) or the eval needs more distinct stimuli to clear the ties

Next action: Inspect tied or lost stimuli; predeclare added breadth before a new experiment.

State: VALID_NO_CHANGE (no_credible_preference_change)

Gate evidence: n=7; 3W/3T/1L; d=4; p=0.312; net +28.6%; 1 dormancy excluded

Overfit: Moderate (score 0.40)

Repeated-run reliability (not used by the gate): 8 paired runs (4W/3T/1L).

Weak or warning scenarios:

Scenario Preference gate Net win Δ Pref Runs (W/T/L)
= Diagnose NuGet restore running redundantly across CI stages Eligible +0.0% +0.0% 0/1/0
▼ Diagnose a pathological ResolveAssemblyReference time Eligible -100.0% -40.0% 0/0/1
= Diagnose per-project overhead across many small projects Eligible +0.0% +0.0% 0/1/0
= Diagnose slow build for a small project Eligible +0.0% +0.0% 0/1/0

Illustrative judge evidence:

  • Diagnose NuGet restore running redundantly across CI stages: Position-swap inconsistent (forward: B, reverse: A). Defaulting to tie.

This is one example, not the aggregate verdict. Open Full Results for every judgment.

➖ Not proven improved — check-bin-obj-clash (claude-sonnet-5)

Why: Net win +28.6% (2W/5T/0L over 7 preference-eligible stimulus vote(s), sign test p=0.250), mean preference +20.0% across 7 paired run(s) — not credible — 5 of 7 preference-eligible stimulus vote(s) tied, leaving only 2 discordant preference vote(s). The sign test conditions on non-tie stimulus votes and cannot reach 0.05 below 5, so no record could have passed here — this is not a measured null. Either the skill is inert on these scenarios (make them discriminate) or the eval needs more distinct stimuli to clear the ties

Next action: Inspect tied or lost stimuli; predeclare added breadth before a new experiment.

State: VALID_NO_CHANGE (no_credible_preference_change)

Gate evidence: n=7; 2W/5T/0L; d=2; p=0.250; net +28.6%

Warnings: Activation: isolated 6/7; plugin 6/7

Overfit: Moderate (score 0.24)

Repeated-run reliability (not used by the gate): 7 paired runs (2W/5T/0L).

Weak or warning scenarios:

Scenario Preference gate Net win Δ Pref Runs (W/T/L)
= Avoid a false clash report when projects share only the top-level artifacts root Eligible +0.0% +0.0% 0/1/0
= Decline output-clash remediation for separate projects using the default SDK layout Eligible +0.0% +0.0% 0/1/0
= Diagnose multi-targeting outputs that collapse into one path Eligible +0.0% +0.0% 0/1/0
= Diagnose redundant project reference metadata that forks a same-path build Eligible +0.0% +0.0% 0/1/0
= Fix all clash mechanisms in the mixed solution Eligible +0.0% +0.0% 0/1/0

Illustrative judge evidence:

  • Avoid a false clash report when projects share only the top-level artifacts root: Both responses are accurate, concise, and directly answer the question with the essential resolved paths and rationale. B adds an intermediate-artifact example, but it does not create a material quality advantage over A.

This is one example, not the aggregate verdict. Open Full Results for every judgment.

➖ Not proven improved — check-bin-obj-clash (gpt-5.6-luna)

Why: Net win +42.9% (4W/2T/1L over 7 preference-eligible stimulus vote(s), sign test p=0.188), mean preference +17.1% across 7 paired run(s) — not credible (sign test p=0.188 > 0.05)

Next action: Inspect tied or lost stimuli and fix inconsistent skill behavior.

State: VALID_NO_CHANGE (no_credible_preference_change)

Gate evidence: n=7; 4W/2T/1L; d=5; p=0.188; net +42.9%

Warnings: Activation: isolated 7/7; plugin 6/7

Overfit: Moderate (score 0.21)

Repeated-run reliability (not used by the gate): 7 paired runs (4W/2T/1L).

Weak or warning scenarios:

Scenario Preference gate Net win Δ Pref Runs (W/T/L)
▼ Audit the mixed solution and separate safe projects from unsafe ones Eligible -100.0% -40.0% 0/0/1
= Decline output-clash remediation for separate projects using the default SDK layout Eligible +0.0% +0.0% 0/1/0
▲ Diagnose multi-targeting outputs that collapse into one path Eligible +100.0% +40.0% 1/0/0
= Diagnose redundant project reference metadata that forks a same-path build Eligible +0.0% +0.0% 0/1/0

Illustrative judge evidence:

  • Audit the mixed solution and separate safe projects from unsafe ones: Both responses demonstrate equivalent technical accuracy and reach identical conclusions about all five projects. They correctly identify the same four unsafe projects (LibraryA, LibraryB, MultiTargetLib, ToolLib) and one safe project (ConsumerApp), with accurate explanations ...

This is one example, not the aggregate verdict. Open Full Results for every judgment.

➖ Not proven improved — directory-build-organization (gpt-5.6-luna)

Why: Net win +66.7% (4W/2T/0L over 6 preference-eligible stimulus vote(s), sign test p=0.063), mean preference +17.1% across 7 paired run(s), 1 dormancy stimulus/stimuli excluded from preference — not credible — 2 of 6 preference-eligible stimulus vote(s) tied, leaving only 4 discordant preference vote(s). The sign test conditions on non-tie stimulus votes and cannot reach 0.05 below 5, so no record could have passed here — this is not a measured null. Either the skill is inert on these scenarios (make them discriminate) or the eval needs more distinct stimuli to clear the ties

Next action: Inspect tied or lost stimuli; predeclare added breadth before a new experiment.

State: VALID_NO_CHANGE (no_credible_preference_change)

Gate evidence: n=6; 4W/2T/0L; d=4; p=0.063; net +66.7%; 1 dormancy excluded

Overfit: High (score 0.53)

Repeated-run reliability (not used by the gate): 7 paired runs (4W/2T/1L).

Weak or warning scenarios:

Scenario Preference gate Net win Δ Pref Runs (W/T/L)
▼ Decline adding shared build files to a lone project Excluded (activation contract) -100.0% -40.0% 0/0/1
= Diagnose a package downgrade chain and reorganize version management Eligible +0.0% +0.0% 0/1/0
= Preserve an intentional project-specific exception while centralizing Eligible +0.0% +0.0% 0/1/0

Illustrative judge evidence:

  • Diagnose a package downgrade chain and reorganize version management: Position-swap inconsistent (forward: A, reverse: B). Defaulting to tie.

This is one example, not the aggregate verdict. Open Full Results for every judgment.

➖ Not proven improved — eval-performance (claude-sonnet-5)

Why: Net win +25.0% (4W/2T/2L over 8 preference-eligible stimulus vote(s), sign test p=0.344), mean preference +10.0% across 8 paired run(s) — not credible (sign test p=0.344 > 0.05)

Next action: Inspect tied or lost stimuli and fix inconsistent skill behavior.

State: VALID_NO_CHANGE (no_credible_preference_change)

Gate evidence: n=8; 4W/2T/2L; d=6; p=0.344; net +25.0%

Warnings: Activation: isolated 6/8; plugin 7/8

Overfit: Moderate (score 0.35)

Repeated-run reliability (not used by the gate): 8 paired runs (4W/2T/2L).

Weak or warning scenarios:

Scenario Preference gate Net win Δ Pref Runs (W/T/L)
▼ Gather a measurement before proposing evaluation fixes for an unremarkable project Eligible -100.0% -40.0% 0/0/1
= Recognize TreatAsLocalProperty overuse versus one justified entry Eligible +0.0% +0.0% 0/1/0
▼ Redirect an incremental-rebuild complaint mistakenly framed as an evaluation problem Eligible -100.0% -40.0% 0/0/1
= Triage which of two property functions actually costs evaluation time Eligible +0.0% +0.0% 0/1/0

Illustrative judge evidence:

  • Gather a measurement before proposing evaluation fixes for an unremarkable project: Both reach the correct practical conclusion: this tiny default project has no evidenced evaluation optimization to apply. A is stronger because it provides a direct, quantitative evaluation measurement and keeps the conclusions more tightly tied to that evidence; B has minor i...

This is one example, not the aggregate verdict. Open Full Results for every judgment.

➖ Not proven improved — eval-performance (gpt-5.6-luna)

Why: Net win +25.0% (3W/4T/1L over 8 preference-eligible stimulus vote(s), sign test p=0.312), mean preference +25.0% across 8 paired run(s) — not credible — 4 of 8 preference-eligible stimulus vote(s) tied, leaving only 4 discordant preference vote(s). The sign test conditions on non-tie stimulus votes and cannot reach 0.05 below 5, so no record could have passed here — this is not a measured null. Either the skill is inert on these scenarios (make them discriminate) or the eval needs more distinct stimuli to clear the ties

Next action: Inspect tied or lost stimuli; predeclare added breadth before a new experiment.

State: VALID_NO_CHANGE (no_credible_preference_change)

Gate evidence: n=8; 3W/4T/1L; d=4; p=0.312; net +25.0%

Overfit: Moderate (score 0.49)

Repeated-run reliability (not used by the gate): 8 paired runs (3W/4T/1L).

Weak or warning scenarios:

Scenario Preference gate Net win Δ Pref Runs (W/T/L)
= Detect a project evaluated twice under different global properties Eligible +0.0% +0.0% 0/1/0
▼ Diagnose stacked evaluation-time patterns (deep imports, broad glob, file-I/O property function) Eligible -100.0% -40.0% 0/0/1
= Recognize TreatAsLocalProperty overuse versus one justified entry Eligible +0.0% +0.0% 0/1/0
= Redirect a compile-time slowdown mistakenly framed as an evaluation problem Eligible +0.0% +0.0% 0/1/0
= Triage which of two property functions actually costs evaluation time Eligible +0.0% +0.0% 0/1/0

Illustrative judge evidence:

  • Detect a project evaluated twice under different global properties: Position-swap inconsistent (forward: B, reverse: A). Defaulting to tie.

This is one example, not the aggregate verdict. Open Full Results for every judgment.

➖ Not proven improved — extension-points (claude-sonnet-5)

Why: Net win +14.3% (3W/2T/2L over 7 preference-eligible stimulus vote(s), sign test p=0.500), mean preference +17.8% across 9 paired run(s), 2 dormancy stimulus/stimuli excluded from preference — not credible (sign test p=0.500 > 0.05)

Next action: Inspect tied or lost stimuli and fix inconsistent skill behavior.

State: VALID_NO_CHANGE (no_credible_preference_change)

Gate evidence: n=7; 3W/2T/2L; d=5; p=0.500; net +14.3%; 2 dormancy excluded

Warnings: Activation: isolated 6/7; plugin 5/7

Overfit: Moderate (score 0.32)

Repeated-run reliability (not used by the gate): 9 paired runs (3W/4T/2L).

Weak or warning scenarios:

Scenario Preference gate Net win Δ Pref Runs (W/T/L)
▼ Create extensibility hooks for a custom SDK target file Eligible -100.0% -40.0% 0/0/1
= Decline Kubernetes readiness probe configuration Excluded (activation contract) +0.0% +0.0% 0/1/0
▼ Diagnose NuGet package and repo extension conflicts Eligible -100.0% -100.0% 0/0/1
= Diagnose a package ID and file-name mismatch Eligible +0.0% +0.0% 0/1/0
= Leave incremental target authoring to its owning workflow Excluded (activation contract) +0.0% +0.0% 0/1/0
= Review packed layout without a false missing-file bug Eligible +0.0% +0.0% 0/1/0

Illustrative judge evidence:

  • Create extensibility hooks for a custom SDK target file: The substantive implementation and validation are effectively equivalent; A is marginally better only because its final summary communicates the resulting structure more precisely.

This is one example, not the aggregate verdict. Open Full Results for every judgment.

➖ Not proven improved — extension-points (gpt-5.6-luna)

Why: Net win +14.3% (3W/2T/2L over 7 preference-eligible stimulus vote(s), sign test p=0.500), mean preference -2.2% across 9 paired run(s), 2 dormancy stimulus/stimuli excluded from preference — not credible (sign test p=0.500 > 0.05)

Next action: Inspect tied or lost stimuli and fix inconsistent skill behavior.

State: VALID_NO_CHANGE (no_credible_preference_change)

Gate evidence: n=7; 3W/2T/2L; d=5; p=0.500; net +14.3%; 2 dormancy excluded

Warnings: Activation: isolated 7/7; plugin 6/7

Overfit: Moderate (score 0.33)

Repeated-run reliability (not used by the gate): 9 paired runs (3W/4T/2L).

Weak or warning scenarios:

Scenario Preference gate Net win Δ Pref Runs (W/T/L)
▼ Create extensibility hooks for a custom SDK target file Eligible -100.0% -40.0% 0/0/1
= Decline Kubernetes readiness probe configuration Excluded (activation contract) +0.0% +0.0% 0/1/0
▼ Diagnose NuGet package and repo extension conflicts Eligible -100.0% -100.0% 0/0/1
= Diagnose build extension point failures Eligible +0.0% +0.0% 0/1/0
= Leave incremental target authoring to its owning workflow Excluded (activation contract) +0.0% +0.0% 0/1/0
= Review packed layout without a false missing-file bug Eligible +0.0% +0.0% 0/1/0

Illustrative judge evidence:

  • Create extensibility hooks for a custom SDK target file: Response A delivers a correct, validated, and well-aligned solution that directly addresses the task requirement. Its minimal approach using CoreMySDKBuild with BeforeTargets/AfterTargets hooks matches the existing test fixtures perfectly and was proven to work through build v...

This is one example, not the aggregate verdict. Open Full Results for every judgment.

➖ Not proven improved — including-generated-files (claude-sonnet-5)

Why: Net win -14.3% (2W/2T/3L over 7 preference-eligible stimulus vote(s), sign test p=0.500), mean preference -12.5% across 8 paired run(s), 1 dormancy stimulus/stimuli excluded from preference — no improvement

Next action: Inspect tied or lost stimuli and fix inconsistent skill behavior.

State: VALID_NO_CHANGE (no_credible_preference_change)

Gate evidence: n=7; 2W/2T/3L; d=5; p=0.500; net -14.3%; 1 dormancy excluded

Warnings: Activation: isolated 6/7; plugin 5/7

Overfit: Moderate (score 0.42)

Repeated-run reliability (not used by the gate): 8 paired runs (2W/3T/3L).

Weak or warning scenarios:

Scenario Preference gate Net win Δ Pref Runs (W/T/L)
= Decline a Python runtime report-path problem Excluded (activation contract) +0.0% +0.0% 0/1/0
▼ Diagnose hardcoded obj path for generated source Eligible -100.0% -40.0% 0/0/1
= Diagnose missing clean tracking for generated source Eligible +0.0% +0.0% 0/1/0
▼ Diagnose project-level glob for generated source Eligible -100.0% -40.0% 0/0/1
= Diagnose wrong hook for generated source files Eligible +0.0% +0.0% 0/1/0
▼ Fix generated source inclusion and clean tracking Eligible -100.0% -100.0% 0/0/1

Illustrative judge evidence:

  • Diagnose hardcoded obj path for generated source: The responses are substantively equally correct, but A's concrete, complete replacement pattern makes its answer marginally more useful while retaining the same accurate diagnosis.

This is one example, not the aggregate verdict. Open Full Results for every judgment.

➖ Not proven improved — including-generated-files (gpt-5.6-luna)

Why: Net win +14.3% (2W/4T/1L over 7 preference-eligible stimulus vote(s), sign test p=0.500), mean preference +10.0% across 8 paired run(s), 1 dormancy stimulus/stimuli excluded from preference — not credible — 4 of 7 preference-eligible stimulus vote(s) tied, leaving only 3 discordant preference vote(s). The sign test conditions on non-tie stimulus votes and cannot reach 0.05 below 5, so no record could have passed here — this is not a measured null. Either the skill is inert on these scenarios (make them discriminate) or the eval needs more distinct stimuli to clear the ties

Next action: Inspect tied or lost stimuli; predeclare added breadth before a new experiment.

State: VALID_NO_CHANGE (no_credible_preference_change)

Gate evidence: n=7; 2W/4T/1L; d=3; p=0.500; net +14.3%; 1 dormancy excluded

Warnings: Activation: isolated 6/7; plugin 7/7

Overfit: Moderate (score 0.39)

Repeated-run reliability (not used by the gate): 8 paired runs (3W/4T/1L).

Weak or warning scenarios:

Scenario Preference gate Net win Δ Pref Runs (W/T/L)
▼ Diagnose hardcoded obj path for generated source Eligible -100.0% -40.0% 0/0/1
= Diagnose missing clean tracking for generated source Eligible +0.0% +0.0% 0/1/0
= Diagnose missing output registration for generated non-code file Eligible +0.0% +0.0% 0/1/0
= Diagnose project-level glob for generated source Eligible +0.0% +0.0% 0/1/0
= Diagnose wrong hook for generated source files Eligible +0.0% +0.0% 0/1/0

Illustrative judge evidence:

  • Diagnose hardcoded obj path for generated source: Both responses reach essentially identical conclusions with correct root-cause analysis and solutions. The key difference is methodology: Response A took a more systematic, bottom-up approach by gathering source files, examining them, then actually building the project and cap...

This is one example, not the aggregate verdict. Open Full Results for every judgment.

➖ Not proven improved — incremental-build (gpt-5.6-luna)

Why: Net win +11.1% (3W/4T/2L over 9 preference-eligible stimulus vote(s), sign test p=0.500), mean preference +4.4% across 9 paired run(s) — not credible (sign test p=0.500 > 0.05)

Next action: Inspect tied or lost stimuli and fix inconsistent skill behavior.

State: VALID_NO_CHANGE (no_credible_preference_change)

Gate evidence: n=9; 3W/4T/2L; d=5; p=0.500; net +11.1%

Warnings: Activation: isolated 7/9; plugin 7/9

Overfit: Moderate (score 0.34)

Repeated-run reliability (not used by the gate): 9 paired runs (3W/4T/2L).

Weak or warning scenarios:

Scenario Preference gate Net win Δ Pref Runs (W/T/L)
= Correct the assumption that Outputs alone enables incremental skipping Eligible +0.0% +0.0% 0/1/0
▼ Diagnose custom targets that always rerun Eligible -100.0% -40.0% 0/0/1
▼ Explain slower builds when MSBuild skipped everything Eligible -100.0% -40.0% 0/0/1
= Fix broken incremental targets and clean tracking Eligible +0.0% +0.0% 0/1/0
= Identify a volatile output path that defeats incrementality Eligible +0.0% +0.0% 0/1/0
= Read a diagnostic log to find the stale input Eligible +0.0% +0.0% 0/1/0

Illustrative judge evidence:

  • Correct the assumption that Outputs alone enables incremental skipping: Position-swap inconsistent (forward: B, reverse: tie). Defaulting to tie.

This is one example, not the aggregate verdict. Open Full Results for every judgment.

➖ Not proven improved — item-management (claude-sonnet-5)

Why: Net win +16.7% (1W/5T/0L over 6 preference-eligible stimulus vote(s), sign test p=0.500), mean preference +5.7% across 7 paired run(s), 1 dormancy stimulus/stimuli excluded from preference — not credible — 5 of 6 preference-eligible stimulus vote(s) tied, leaving only 1 discordant preference vote(s). The sign test conditions on non-tie stimulus votes and cannot reach 0.05 below 5, so no record could have passed here — this is not a measured null. Either the skill is inert on these scenarios (make them discriminate) or the eval needs more distinct stimuli to clear the ties

Next action: Inspect tied or lost stimuli; predeclare added breadth before a new experiment.

State: VALID_NO_CHANGE (no_credible_preference_change)

Gate evidence: n=6; 1W/5T/0L; d=1; p=0.500; net +16.7%; 1 dormancy excluded

Warnings: Activation: isolated 3/6; plugin 2/6

Overfit: Moderate (score 0.38)

Repeated-run reliability (not used by the gate): 7 paired runs (1W/6T/0L).

Weak or warning scenarios:

Scenario Preference gate Net win Δ Pref Runs (W/T/L)
= Decline JavaScript configuration merging Excluded (activation contract) +0.0% +0.0% 0/1/0
▲ Diagnose an ineffective Compile Remove that does not match the glob Eligible +100.0% +40.0% 1/0/0
= Diagnose item group and batching issues Eligible +0.0% +0.0% 0/1/0
= Diagnose real and claimed item problems in a code generation pipeline Eligible +0.0% +0.0% 0/1/0
= Fix item management anti-patterns Eligible +0.0% +0.0% 0/1/0
= Leave already-correct item management unchanged Eligible +0.0% +0.0% 0/1/0
= Leave correct single-list batching unchanged Eligible +0.0% +0.0% 0/1/0

Illustrative judge evidence:

  • Diagnose item group and batching issues: Position-swap inconsistent (forward: A, reverse: tie). Defaulting to tie.

This is one example, not the aggregate verdict. Open Full Results for every judgment.

➖ Not proven improved — item-management (gpt-5.6-luna)

Why: Net win +0.0% (1W/4T/1L over 6 preference-eligible stimulus vote(s), sign test p=0.750), mean preference +8.6% across 7 paired run(s), 1 dormancy stimulus/stimuli excluded from preference — no improvement

Next action: Inspect tied or lost stimuli; predeclare added breadth before a new experiment.

State: VALID_NO_CHANGE (no_credible_preference_change)

Gate evidence: n=6; 1W/4T/1L; d=2; p=0.750; net +0.0%; 1 dormancy excluded

Overfit: Moderate (score 0.47)

Repeated-run reliability (not used by the gate): 7 paired runs (1W/5T/1L).

Weak or warning scenarios:

Scenario Preference gate Net win Δ Pref Runs (W/T/L)
= Decline JavaScript configuration merging Excluded (activation contract) +0.0% +0.0% 0/1/0
= Diagnose an ineffective Compile Remove that does not match the glob Eligible +0.0% +0.0% 0/1/0
= Diagnose item group and batching issues Eligible +0.0% +0.0% 0/1/0
= Diagnose real and claimed item problems in a code generation pipeline Eligible +0.0% +0.0% 0/1/0
▼ Fix item management anti-patterns Eligible -100.0% -40.0% 0/0/1
= Leave correct single-list batching unchanged Eligible +0.0% +0.0% 0/1/0

Illustrative judge evidence:

  • Diagnose an ineffective Compile Remove that does not match the glob: Position-swap inconsistent (forward: B, reverse: A). Defaulting to tie.

This is one example, not the aggregate verdict. Open Full Results for every judgment.

➖ Not proven improved — msbuild-antipatterns (claude-sonnet-5)

Why: Net win +28.6% (4W/1T/2L over 7 preference-eligible stimulus vote(s), sign test p=0.344), mean preference +10.0% across 8 paired run(s), 1 dormancy stimulus/stimuli excluded from preference — not credible (sign test p=0.344 > 0.05)

Next action: Inspect tied or lost stimuli and fix inconsistent skill behavior.

State: VALID_NO_CHANGE (no_credible_preference_change)

Gate evidence: n=7; 4W/1T/2L; d=6; p=0.344; net +28.6%; 1 dormancy excluded

Warnings: Activation: isolated 2/7; plugin 1/7

Overfit: Moderate (score 0.29)

Repeated-run reliability (not used by the gate): 8 paired runs (4W/2T/2L).

Weak or warning scenarios:

Scenario Preference gate Net win Δ Pref Runs (W/T/L)
▲ Add a module to an F# project Eligible +100.0% +40.0% 1/0/0
= Add a signature file to define public API Eligible +0.0% +0.0% 0/1/0
▲ Distinguish a style backslash from a real cross-platform backslash bug Eligible +100.0% +100.0% 1/0/0
▲ Fix broken file order causing FS0039 Eligible +100.0% +40.0% 1/0/0
▼ Judge an unguarded import inside a NuGet package build folder Eligible -100.0% -100.0% 0/0/1
▼ Leave a clean project without inventing anti-patterns Eligible -100.0% -40.0% 0/0/1
= Non-activation: migrate a legacy project to SDK style Excluded (activation contract) +0.0% +0.0% 0/1/0
▲ Review MSBuild files for anti-patterns and style issues Eligible +100.0% +40.0% 1/0/0

Illustrative judge evidence:

  • Add a signature file to define public API: The resulting changes are substantively identical and fully satisfy the task: a matching public signature file is added, included in the correct order, and successfully built.

This is one example, not the aggregate verdict. Open Full Results for every judgment.

➖ Not proven improved — msbuild-modernization (gpt-5.6-luna)

Why: Net win +0.0% (2W/2T/2L over 6 preference-eligible stimulus vote(s), sign test p=0.687), mean preference -8.6% across 7 paired run(s), 1 dormancy stimulus/stimuli excluded from preference — no improvement

Next action: Inspect tied or lost stimuli; predeclare added breadth before a new experiment.

State: VALID_NO_CHANGE (no_credible_preference_change)

Gate evidence: n=6; 2W/2T/2L; d=4; p=0.687; net +0.0%; 1 dormancy excluded

Warnings: Activation: isolated 5/6; plugin 6/6

Overfit: Moderate (score 0.35)

Repeated-run reliability (not used by the gate): 7 paired runs (2W/3T/2L).

Weak or warning scenarios:

Scenario Preference gate Net win Δ Pref Runs (W/T/L)
= Decline modernizing a non-.NET build Excluded (activation contract) +0.0% +0.0% 0/1/0
▼ Identify legacy patterns for SDK-style migration Eligible -100.0% -100.0% 0/0/1
▼ Modernize a single project without introducing Central Package Management Eligible -100.0% -40.0% 0/0/1
= Modernize further without introducing a nondeterministic language version Eligible +0.0% +0.0% 0/1/0
= Recognize an already-modern project needs no migration Eligible +0.0% +0.0% 0/1/0

Illustrative judge evidence:

  • Identify legacy patterns for SDK-style migration: Response A delivers a superior modernization outcome: (1) The project builds successfully on its target framework (net472), while Response B fails to build on net472 and requires retargeting to net8.0, undermining the modernization's validity. (2) Response A achieves genuine s...

This is one example, not the aggregate verdict. Open Full Results for every judgment.

Details for 9 results were omitted to keep this comment under GitHub's limit. Open the workflow summary or Full Results for the complete breakdown.

🔍 Full Results - all metrics and investigation details

To investigate non-passing or warning results, paste this to your AI coding agent:

For PR 1214 in dotnet/skills, download eval artifacts with gh run download 36200961326 --repo dotnet/skills --pattern "vally-results-*" --dir ./eval-results, then fetch https://raw.githubusercontent.com/dotnet/skills/1e90581d198eb68a2357f0c411be6643ee7dff40/eng/vally-adapter/InvestigatingResults.md and follow it. Classify each result as measurement-invalid, underpowered, not-proven, preference-loss, or passing-with-warning. Use stateReason, result accounting, weak scenarios, and judge evidence to give the cause and exact next fix.

⚠️ Session replay telemetry was not published because the auxiliary dotnet/skills-data publisher failed. The evaluation verdicts above remain authoritative; maintainers must repair SKILLS_DATA_TOKEN or the publisher before replay links are available.

github-actions Bot added a commit that referenced this pull request Sep 26, 2026
@github-actions github-actions Bot added waiting-on-review PR state label and removed pr-state/evals-in-progress PR evaluations are in progress labels Sep 26, 2026
@github-actions

Copy link
Copy Markdown
Contributor

✅ Evaluation passed for 1e90581. cc @AbhitejJohn @JanKrivanek @dotnet/msbuild @YuliiaKovalova — please review.

@github-actions

Copy link
Copy Markdown
Contributor

✅ Evaluation passed for 1e90581. cc @AbhitejJohn @JanKrivanek @dotnet/msbuild @YuliiaKovalova — please review.

@JanKrivanek JanKrivanek left a comment

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Please consider the findings before merging. Otherwise carrying over signoff from #1118

Comment thread eng/skill-validator/src/Evaluate/EvaluateCommand.cs Outdated
Comment thread eng/skill-validator/src/Evaluate/RejudgeCommand.cs Outdated
Persist terminal runner failures separately from recoverable tool errors and support complete three-arm databases during cross-directory rejudge.

Co-authored-by: Copilot App <223556219+Copilot@users.noreply.github.com>
Copilot AI lite review requested due to automatic review settings September 30, 2026 12:41
@github-actions github-actions Bot added pr-state/ready-for-eval PR is mergeable and awaiting evaluation pr-state/evals-in-progress PR evaluations are in progress and removed waiting-on-review PR state label pr-state/ready-for-eval PR is mergeable and awaiting evaluation labels Sep 30, 2026

Copilot AI left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Copilot review overview

🟡 Changes recommended

Rejudge role filtering and pairing leave evidence unaccounted, and the binlog grader can pass without proving a valid build artifact.

Review effort: Lite
Findings: 2 High severity

Open (2)

Comment thread eng/skill-validator/src/Evaluate/RejudgeCommand.cs
Comment thread tests/dotnet-msbuild/binlog-generation/eval.yaml
AbhitejJohn and others added 3 commits September 30, 2026 05:58
Persist terminal runner failures separately from recoverable tool errors and support complete three-arm databases during cross-directory rejudge.

Co-authored-by: Copilot App <223556219+Copilot@users.noreply.github.com>
Use merge-base change discovery as main advances and require executable MSBuild binlog evidence for single-build scenarios.

Co-authored-by: Copilot App <223556219+Copilot@users.noreply.github.com>
Copilot AI lite review requested due to automatic review settings September 30, 2026 13:05

Copilot AI left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Copilot review overview

🔵 Needs a closer look

Two moderate findings remain, and the broad evaluation infrastructure changes warrant human review.

Review effort: Lite
Findings: None

Resolved since last review (2)

@github-actions github-actions Bot added pr-state/ready-for-eval PR is mergeable and awaiting evaluation pr-state/evals-in-progress PR evaluations are in progress and removed pr-state/evals-in-progress PR evaluations are in progress pr-state/ready-for-eval PR is mergeable and awaiting evaluation labels Sep 30, 2026
@github-actions

Copy link
Copy Markdown
Contributor

❌ Evaluation did not complete successfully (the evaluate job reported failure). Check the workflow run logs, then comment /evaluate dc917f7db7b7075a8a953a28e07ed5e3f67822d2 to retry this exact commit.

36 partial result file(s) were preserved for diagnosis but were not consolidated because the full matrix did not complete.

⚠️ Session replay telemetry was not published because the auxiliary dotnet/skills-data publisher failed. The evaluation verdicts above remain authoritative; maintainers must repair SKILLS_DATA_TOKEN or the publisher before replay links are available.

github-actions Bot added a commit that referenced this pull request Sep 30, 2026
@github-actions

Copy link
Copy Markdown
Contributor

📊 Skill and Agent Evaluation Results

36 model/target results across 18 targets and 2 models — ✅ 1 improved, ➖ 31 not proven improved, ⚠️ 0 invalid or underpowered, ⛔ 4 activation contract failures, 📉 0 preference losses (report only).

Measurement identity: evaluated commit 4b6cf7540e24d8984c70e63b442ea9588520a041; 2 judge models.

Measurement health: 36 expected / 36 observed / 36 written; 0 missing, 0 unexpected, 0 invalid; 0 recovered comparison error slots and 0 unresolved comparison error slots.

Objective completion gate: not enabled. Aggregate completion transitions are telemetry only, so this report does not claim that zero objective regressions were proven.

A result passes only when preference-eligible distinct-stimulus votes have aggregate net win of at least 20% and an exact one-sided sign-test result of p ≤ 0.05, and every explicit dormancy activation contract passes. Repeated runs measure reliability only.

Target Model Verdict Gate evidence Overfit Warnings Next action
agent.msbuild claude-sonnet-5 ➖ Not proven improved n=5; 2W/1T/2L; d=4; p=0.687; net +0.0% — Activation: isolated 1/5; plugin 0/5 Inspect tied or lost stimuli; predeclare added breadth before a new experiment.
agent.msbuild gpt-5.6-luna ➖ Not proven improved n=5; 2W/2T/1L; d=3; p=0.500; net +20.0% — Activation: isolated 2/5; plugin 0/5 Inspect tied or lost stimuli; predeclare added breadth before a new experiment.
binlog-failure-analysis claude-sonnet-5 ➖ Not proven improved n=6; 3W/1T/2L; d=5; p=0.500; net +16.7%; 1 dormancy excluded 🟡 0.36 Activation: isolated 5/7; plugin 6/7 Inspect tied or lost stimuli and fix inconsistent skill behavior.
binlog-failure-analysis gpt-5.6-luna ➖ Not proven improved n=6; 2W/3T/1L; d=3; p=0.500; net +16.7%; 1 dormancy excluded ✅ 0.18 Activation: isolated 6/7; plugin 6/7 Inspect tied or lost stimuli; predeclare added breadth before a new experiment.
binlog-generation claude-sonnet-5 ➖ Not proven improved n=7; 4W/3T/0L; d=4; p=0.063; net +57.1%; 1 dormancy excluded 🟡 0.48 Activation: isolated 6/7; plugin 5/7 Inspect tied or lost stimuli; predeclare added breadth before a new experiment.
binlog-generation gpt-5.6-luna ✅ Improved n=7; 5W/2T/0L; d=5; p=0.031; net +71.4%; 1 dormancy excluded 🔴 0.51 Activation: isolated 7/7; plugin 5/7 Fix activation gaps; Review overfit evidence.
build-parallelism claude-sonnet-5 ➖ Not proven improved n=6; 1W/4T/1L; d=2; p=0.750; net +0.0%; 1 dormancy excluded ✅ 0.18 Activation: isolated 3/6; plugin 6/6; Activation-only stop: isolated 1 failed run; Activation-only stop: plugin 1 failed run Inspect activation-only failed runs before rewriting skill content; the model stopped after loading a skill.
build-parallelism gpt-5.6-luna ➖ Not proven improved n=6; 3W/2T/1L; d=4; p=0.312; net +33.3%; 1 dormancy excluded 🔴 0.55 Activation-only stop: isolated 1 failed run; Activation-only stop: plugin 1 failed run Inspect activation-only failed runs before rewriting skill content; the model stopped after loading a skill.
build-perf-baseline claude-sonnet-5 ➖ Not proven improved n=6; 5W/0T/1L; d=6; p=0.109; net +66.7%; 1 dormancy excluded 🟡 0.33 Activation: isolated 4/6; plugin 5/6 Inspect tied or lost stimuli and fix inconsistent skill behavior.
build-perf-baseline gpt-5.6-luna ⛔ Activation contract failed n=6; 5W/1T/0L; d=5; p=0.031; net +83.3%; 1 dormancy excluded 🟡 0.30 Dormancy contract: 1 unexpected activation(s); Activation: isolated 6/6; plugin 5/6 Narrow skill routing so the listed off-target scenarios stay dormant.
build-perf-diagnostics claude-sonnet-5 ➖ Not proven improved n=7; 4W/1T/2L; d=6; p=0.344; net +28.6%; 1 dormancy excluded 🔴 0.52 — Inspect tied or lost stimuli and fix inconsistent skill behavior.
build-perf-diagnostics gpt-5.6-luna ➖ Not proven improved n=7; 2W/5T/0L; d=2; p=0.250; net +28.6%; 1 dormancy excluded 🟡 0.42 — Inspect tied or lost stimuli; predeclare added breadth before a new experiment.
check-bin-obj-clash claude-sonnet-5 ➖ Not proven improved n=7; 3W/1T/3L; d=6; p=0.656; net +0.0% ✅ 0.17 Activation: isolated 7/7; plugin 6/7 Inspect tied or lost stimuli and fix inconsistent skill behavior.
check-bin-obj-clash gpt-5.6-luna ➖ Not proven improved n=7; 3W/1T/3L; d=6; p=0.656; net +0.0% 🟡 0.22 — Inspect tied or lost stimuli and fix inconsistent skill behavior.
directory-build-organization claude-sonnet-5 ⛔ Activation contract failed n=6; 1W/3T/2L; d=3; p=0.500; net -16.7%; 1 dormancy excluded 🟡 0.36 Dormancy contract: 1 unexpected activation(s); Activation: isolated 3/6; plugin 5/6 Narrow skill routing so the listed off-target scenarios stay dormant.
directory-build-organization gpt-5.6-luna ⛔ Activation contract failed n=6; 3W/3T/0L; d=3; p=0.125; net +50.0%; 1 dormancy excluded 🔴 0.54 Dormancy contract: 1 unexpected activation(s) Narrow skill routing so the listed off-target scenarios stay dormant.
eval-performance claude-sonnet-5 ➖ Not proven improved n=8; 4W/4T/0L; d=4; p=0.063; net +50.0% 🟡 0.28 Activation: isolated 6/8; plugin 6/8; Activation-only stop: plugin 1 failed run Inspect activation-only failed runs before rewriting skill content; the model stopped after loading a skill.
eval-performance gpt-5.6-luna ➖ Not proven improved n=8; 4W/4T/0L; d=4; p=0.063; net +50.0% 🟡 0.26 Activation: isolated 8/8; plugin 7/8 Inspect tied or lost stimuli; predeclare added breadth before a new experiment.
extension-points claude-sonnet-5 ➖ Not proven improved n=7; 2W/2T/3L; d=5; p=0.500; net -14.3%; 2 dormancy excluded 🟡 0.31 Activation: isolated 4/7; plugin 5/7 Inspect tied or lost stimuli and fix inconsistent skill behavior.
extension-points gpt-5.6-luna ⛔ Activation contract failed n=7; 3W/2T/2L; d=5; p=0.500; net +14.3%; 2 dormancy excluded 🟡 0.36 Dormancy contract: 1 unexpected activation(s); Activation: isolated 7/7; plugin 6/7 Narrow skill routing so the listed off-target scenarios stay dormant.
including-generated-files claude-sonnet-5 ➖ Not proven improved n=7; 2W/3T/2L; d=4; p=0.687; net +0.0%; 1 dormancy excluded 🟡 0.49 Activation: isolated 6/7; plugin 5/7; Activation-only stop: plugin 1 failed run Inspect activation-only failed runs before rewriting skill content; the model stopped after loading a skill.
including-generated-files gpt-5.6-luna ➖ Not proven improved n=7; 4W/1T/2L; d=6; p=0.344; net +28.6%; 1 dormancy excluded 🟡 0.43 — Inspect tied or lost stimuli and fix inconsistent skill behavior.
incremental-build claude-sonnet-5 ➖ Not proven improved n=9; 1W/5T/3L; d=4; p=0.312; net -22.2% 🟡 0.32 Activation: isolated 4/9; plugin 5/9 Inspect tied or lost stimuli; predeclare added breadth before a new experiment.
incremental-build gpt-5.6-luna ➖ Not proven improved n=9; 5W/3T/1L; d=6; p=0.109; net +44.4% 🟡 0.50 Activation: isolated 8/9; plugin 7/9 Inspect tied or lost stimuli and fix inconsistent skill behavior.
item-management claude-sonnet-5 ➖ Not proven improved n=6; 3W/1T/2L; d=5; p=0.500; net +16.7%; 1 dormancy excluded 🟡 0.37 Activation: isolated 4/6; plugin 3/6 Inspect tied or lost stimuli and fix inconsistent skill behavior.
item-management gpt-5.6-luna ➖ Not proven improved n=6; 2W/4T/0L; d=2; p=0.250; net +33.3%; 1 dormancy excluded 🟡 0.42 — Inspect tied or lost stimuli; predeclare added breadth before a new experiment.
msbuild-antipatterns claude-sonnet-5 ➖ Not proven improved n=7; 1W/3T/3L; d=4; p=0.312; net -28.6%; 1 dormancy excluded 🟡 0.31 Activation: isolated 3/7; plugin 1/7 Inspect tied or lost stimuli; predeclare added breadth before a new experiment.
msbuild-antipatterns gpt-5.6-luna ➖ Not proven improved n=7; 1W/5T/1L; d=2; p=0.750; net +0.0%; 1 dormancy excluded 🟡 0.39 Activation: isolated 4/7; plugin 3/7 Inspect tied or lost stimuli; predeclare added breadth before a new experiment.
msbuild-modernization claude-sonnet-5 ➖ Not proven improved n=6; 2W/3T/1L; d=3; p=0.500; net +16.7%; 1 dormancy excluded 🟡 0.35 Activation: isolated 5/6; plugin 5/6 Inspect tied or lost stimuli; predeclare added breadth before a new experiment.
msbuild-modernization gpt-5.6-luna ➖ Not proven improved n=6; 1W/2T/3L; d=4; p=0.312; net -33.3%; 1 dormancy excluded ✅ 0.14 Activation: isolated 5/6; plugin 6/6 Inspect tied or lost stimuli; predeclare added breadth before a new experiment.
property-patterns claude-sonnet-5 ➖ Not proven improved n=7; 3W/2T/2L; d=5; p=0.500; net +14.3% 🟡 0.26 Activation: isolated 6/7; plugin 5/7 Inspect tied or lost stimuli and fix inconsistent skill behavior.
property-patterns gpt-5.6-luna ➖ Not proven improved n=7; 3W/2T/2L; d=5; p=0.500; net +14.3% 🟡 0.22 Activation: isolated 6/7; plugin 7/7 Inspect tied or lost stimuli and fix inconsistent skill behavior.
resolve-project-references claude-sonnet-5 ➖ Not proven improved n=6; 2W/2T/2L; d=4; p=0.687; net +0.0%; 2 dormancy excluded 🟡 0.28 Activation: isolated 5/6; plugin 5/6; Activation-only stop: plugin 1 failed run Inspect activation-only failed runs before rewriting skill content; the model stopped after loading a skill.
resolve-project-references gpt-5.6-luna ➖ Not proven improved n=6; 1W/4T/1L; d=2; p=0.750; net +0.0%; 2 dormancy excluded 🟡 0.28 Activation-only stop: plugin 1 failed run Inspect activation-only failed runs before rewriting skill content; the model stopped after loading a skill.
target-authoring claude-sonnet-5 ➖ Not proven improved n=6; 1W/3T/2L; d=3; p=0.500; net -16.7%; 1 dormancy excluded 🟡 0.47 Activation: isolated 4/6; plugin 6/6 Inspect tied or lost stimuli; predeclare added breadth before a new experiment.
target-authoring gpt-5.6-luna ➖ Not proven improved n=6; 1W/3T/2L; d=3; p=0.500; net -16.7%; 1 dormancy excluded 🟡 0.44 — Inspect tied or lost stimuli; predeclare added breadth before a new experiment.
ℹ️ How to read this report
  • ✅ Improved — the result passed both the statistical gate and the 20% practical net-win floor.
  • ➖ Not proven improved — the result is valid but did not pass both gates. This is not automatically a regression.
  • ⚠️ Invalid / underpowered — the gate withheld a quality verdict. Fix the measurement before judging the target.
  • ⛔ Activation contract failed — the isolated target activated on an explicit dormancy scenario. Dormancy preference is excluded, but this routing failure still blocks a pass.
  • 📉 Preference loss — the LLM judge credibly preferred baseline. It is report-only, not objective completion proof.
  • Gate evidence — n preference-eligible distinct-stimulus votes, W/T/L stimulus votes, d discordant votes, exact one-sided p, net win, and the count of separately retained dormancy stimuli. The p value applies to one model/target result; no matrix-wide multiple-comparison correction is applied.
  • Overfit — overfitting-judge severity (✅ Low, 🟡 Moderate, 🔴 High, — none) and score.
  • Warnings — activation, timeout, retry recovery, or unresolved comparison conditions that need attention.
  • Do not add repeated runs to increase statistical power. Do not add stimuli after seeing a near-pass unless the new breadth is predeclared for a new experiment.
⛔ Activation contract failed — build-perf-baseline (gpt-5.6-luna)

Why: Net win +83.3% (5W/1T/0L over 6 preference-eligible stimulus vote(s), sign test p=0.031), mean preference +28.6% across 7 paired run(s), 1 dormancy stimulus/stimuli excluded from preference — activation contract failed (1 explicit dormancy scenario(s) activated the isolated target skill)

Next action: Narrow skill routing so the listed off-target scenarios stay dormant.

State: VALID_NO_CHANGE (activation_contract_failed)

Gate evidence: n=6; 5W/1T/0L; d=5; p=0.031; net +83.3%; 1 dormancy excluded

Warnings: Dormancy contract: 1 unexpected activation(s); Activation: isolated 6/6; plugin 5/6

Overfit: Moderate (score 0.30)

Repeated-run reliability (not used by the gate): 7 paired runs (5W/2T/0L).

Weak or warning scenarios:

Scenario Preference gate Net win Δ Pref Runs (W/T/L)
▲ Configure deterministic, cache-safe CI builds Eligible +100.0% +40.0% 1/0/0
= Decline a non-MSBuild build performance request Excluded (activation contract) +0.0% +0.0% 0/1/0
= Leave an already-optimized build unchanged Eligible +0.0% +0.0% 0/1/0

Illustrative judge evidence:

  • Leave an already-optimized build unchanged: Position-swap inconsistent (forward: B, reverse: A). Defaulting to tie.

This is one example, not the aggregate verdict. Open Full Results for every judgment.

⛔ Activation contract failed — directory-build-organization (claude-sonnet-5)

Why: Net win -16.7% (1W/3T/2L over 6 preference-eligible stimulus vote(s), sign test p=0.500), mean preference -5.7% across 7 paired run(s), 1 dormancy stimulus/stimuli excluded from preference — activation contract failed (1 explicit dormancy scenario(s) activated the isolated target skill)

Next action: Narrow skill routing so the listed off-target scenarios stay dormant.

State: VALID_NO_CHANGE (activation_contract_failed)

Gate evidence: n=6; 1W/3T/2L; d=3; p=0.500; net -16.7%; 1 dormancy excluded

Warnings: Dormancy contract: 1 unexpected activation(s); Activation: isolated 3/6; plugin 5/6

Overfit: Moderate (score 0.36)

Repeated-run reliability (not used by the gate): 7 paired runs (1W/4T/2L).

Weak or warning scenarios:

Scenario Preference gate Net win Δ Pref Runs (W/T/L)
= Apply repo-level build organization cleanup Eligible +0.0% +0.0% 0/1/0
= Decline adding shared build files to a lone project Excluded (activation contract) +0.0% +0.0% 0/1/0
▼ Diagnose a TargetFramework condition that silently skips in props Eligible -100.0% -40.0% 0/0/1
= Diagnose a package downgrade chain and reorganize version management Eligible +0.0% +0.0% 0/1/0
▼ Diagnose an inner shared-props file that overwrites its own override Eligible -100.0% -40.0% 0/0/1
= Preserve an intentional project-specific exception while centralizing Eligible +0.0% +0.0% 0/1/0

Illustrative judge evidence:

  • Apply repo-level build organization cleanup: Both runs deliver the requested MSBuild/CPM reorganization, recover from tooling/edit issues, validate restore, and appropriately identify the identical CS1591 build failure as pre-existing. B's extra src props import chain is harmless but not a meaningful quality advantage ov...

This is one example, not the aggregate verdict. Open Full Results for every judgment.

⛔ Activation contract failed — directory-build-organization (gpt-5.6-luna)

Why: Net win +50.0% (3W/3T/0L over 6 preference-eligible stimulus vote(s), sign test p=0.125), mean preference +17.1% across 7 paired run(s), 1 dormancy stimulus/stimuli excluded from preference — activation contract failed (1 explicit dormancy scenario(s) activated the isolated target skill)

Next action: Narrow skill routing so the listed off-target scenarios stay dormant.

State: VALID_NO_CHANGE (activation_contract_failed)

Gate evidence: n=6; 3W/3T/0L; d=3; p=0.125; net +50.0%; 1 dormancy excluded

Warnings: Dormancy contract: 1 unexpected activation(s)

Overfit: High (score 0.54)

Repeated-run reliability (not used by the gate): 7 paired runs (3W/4T/0L).

Weak or warning scenarios:

Scenario Preference gate Net win Δ Pref Runs (W/T/L)
= Apply repo-level build organization cleanup Eligible +0.0% +0.0% 0/1/0
= Decline adding shared build files to a lone project Excluded (activation contract) +0.0% +0.0% 0/1/0
= Diagnose a TargetFramework condition that silently skips in props Eligible +0.0% +0.0% 0/1/0
= Preserve an intentional project-specific exception while centralizing Eligible +0.0% +0.0% 0/1/0

Illustrative judge evidence:

  • Apply repo-level build organization cleanup: Position-swap inconsistent (forward: B, reverse: A). Defaulting to tie.

This is one example, not the aggregate verdict. Open Full Results for every judgment.

⛔ Activation contract failed — extension-points (gpt-5.6-luna)

Why: Net win +14.3% (3W/2T/2L over 7 preference-eligible stimulus vote(s), sign test p=0.500), mean preference +6.7% across 9 paired run(s), 2 dormancy stimulus/stimuli excluded from preference — activation contract failed (1 explicit dormancy scenario(s) activated the isolated target skill)

Next action: Narrow skill routing so the listed off-target scenarios stay dormant.

State: VALID_NO_CHANGE (activation_contract_failed)

Gate evidence: n=7; 3W/2T/2L; d=5; p=0.500; net +14.3%; 2 dormancy excluded

Warnings: Dormancy contract: 1 unexpected activation(s); Activation: isolated 7/7; plugin 6/7

Overfit: Moderate (score 0.36)

Repeated-run reliability (not used by the gate): 9 paired runs (3W/3T/3L).

Weak or warning scenarios:

Scenario Preference gate Net win Δ Pref Runs (W/T/L)
= Decline Kubernetes readiness probe configuration Excluded (activation contract) +0.0% +0.0% 0/1/0
= Diagnose NuGet package and repo extension conflicts Eligible +0.0% +0.0% 0/1/0
= Diagnose a package ID and file-name mismatch Eligible +0.0% +0.0% 0/1/0
▼ Diagnose build extension point failures Eligible -100.0% -40.0% 0/0/1
▼ Leave incremental target authoring to its owning workflow Excluded (activation contract) -100.0% -100.0% 0/0/1
▼ Review packed layout without a false missing-file bug Eligible -100.0% -40.0% 0/0/1

Illustrative judge evidence:

  • Diagnose NuGet package and repo extension conflicts: Position-swap inconsistent (forward: B, reverse: A). Defaulting to tie.

This is one example, not the aggregate verdict. Open Full Results for every judgment.

➖ Not proven improved — agent.msbuild (claude-sonnet-5)

Why: Net win +0.0% (2W/1T/2L over 5 preference-eligible stimulus vote(s), sign test p=0.687), mean preference +12.0% across 5 paired run(s) — no improvement — native evaluator reported that the target agent did not activate

Next action: Inspect tied or lost stimuli; predeclare added breadth before a new experiment.

State: VALID_NO_CHANGE (target_agent_not_activated)

Gate evidence: n=5; 2W/1T/2L; d=4; p=0.687; net +0.0%

Warnings: Activation: isolated 1/5; plugin 0/5

Repeated-run reliability (not used by the gate): 5 paired runs (2W/1T/2L).

Weak or warning scenarios:

Scenario Preference gate Net win Δ Pref Runs (W/T/L)
▲ Advise on project file organization Eligible +100.0% +40.0% 1/0/0
▼ Diagnose broken incremental build behavior Eligible -100.0% -40.0% 0/0/1
▲ Review a project file for maintainability risks Eligible +100.0% +100.0% 1/0/0
= Route a slow build to performance analysis Eligible +0.0% +0.0% 0/1/0
▼ Triage a build failure and route to appropriate analysis Eligible -100.0% -40.0% 0/0/1

Illustrative judge evidence:

  • Diagnose broken incremental build behavior: Both correctly diagnose the unconditional custom target and propose Inputs/Outputs. A better fulfills the request for diagnostic guidance and specific input/output/timestamp failure modes. B's FileWrites addition is useful for clean integration, but its answer is materially le...

This is one example, not the aggregate verdict. Open Full Results for every judgment.

➖ Not proven improved — agent.msbuild (gpt-5.6-luna)

Why: Net win +20.0% (2W/2T/1L over 5 preference-eligible stimulus vote(s), sign test p=0.500), mean preference +8.0% across 5 paired run(s) — not credible — 2 of 5 preference-eligible stimulus vote(s) tied, leaving only 3 discordant preference vote(s). The sign test conditions on non-tie stimulus votes and cannot reach 0.05 below 5, so no record could have passed here — this is not a measured null. Either the agent is inert on these scenarios (make them discriminate) or the eval needs more distinct stimuli to clear the ties — native evaluator reported that the target agent did not activate

Next action: Inspect tied or lost stimuli; predeclare added breadth before a new experiment.

State: VALID_NO_CHANGE (target_agent_not_activated)

Gate evidence: n=5; 2W/2T/1L; d=3; p=0.500; net +20.0%

Warnings: Activation: isolated 2/5; plugin 0/5

Repeated-run reliability (not used by the gate): 5 paired runs (2W/2T/1L).

Weak or warning scenarios:

Scenario Preference gate Net win Δ Pref Runs (W/T/L)
= Advise on project file organization Eligible +0.0% +0.0% 0/1/0
= Diagnose broken incremental build behavior Eligible +0.0% +0.0% 0/1/0
▲ Review a project file for maintainability risks Eligible +100.0% +40.0% 1/0/0
▲ Route a slow build to performance analysis Eligible +100.0% +40.0% 1/0/0
▼ Triage a build failure and route to appropriate analysis Eligible -100.0% -40.0% 0/0/1

Illustrative judge evidence:

  • Advise on project file organization: Position-swap inconsistent (forward: skill, reverse: baseline). Defaulting to tie.

This is one example, not the aggregate verdict. Open Full Results for every judgment.

➖ Not proven improved — binlog-failure-analysis (claude-sonnet-5)

Why: Net win +16.7% (3W/1T/2L over 6 preference-eligible stimulus vote(s), sign test p=0.500), mean preference +5.7% across 7 paired run(s), 1 dormancy stimulus/stimuli excluded from preference — not credible (sign test p=0.500 > 0.05)

Next action: Inspect tied or lost stimuli and fix inconsistent skill behavior.

State: VALID_NO_CHANGE (no_credible_preference_change)

Gate evidence: n=6; 3W/1T/2L; d=5; p=0.500; net +16.7%; 1 dormancy excluded

Warnings: Activation: isolated 5/7; plugin 6/7

Overfit: Moderate (score 0.36)

Repeated-run reliability (not used by the gate): 7 paired runs (3W/2T/2L).

Weak or warning scenarios:

Scenario Preference gate Net win Δ Pref Runs (W/T/L)
▼ Assess a requested binlog investigation when the current build is healthy Eligible -100.0% -40.0% 0/0/1
= Confirm the actual resolved target framework and package version from a binlog Eligible +0.0% +0.0% 0/0/0
= Diagnose a warning behind a build that actually succeeded Eligible +0.0% +0.0% 0/1/0
▼ Diagnose build failures from binlog only (no source files) Eligible -100.0% -40.0% 0/0/1
= Stay dormant for a non-MSBuild build failure log Excluded (activation contract) +0.0% +0.0% 0/1/0

Illustrative judge evidence:

  • Assess a requested binlog investigation when the current build is healthy: Both are correct and satisfy the core request for a safe binlog capture without fabricating a diagnosis. A is modestly more complete and evidentially grounded, including source/project context and a runtime verification, while B is accurate but prematurely dismisses any furthe...

This is one example, not the aggregate verdict. Open Full Results for every judgment.

➖ Not proven improved — binlog-failure-analysis (gpt-5.6-luna)

Why: Net win +16.7% (2W/3T/1L over 6 preference-eligible stimulus vote(s), sign test p=0.500), mean preference -2.9% across 7 paired run(s), 1 dormancy stimulus/stimuli excluded from preference — not credible — 3 of 6 preference-eligible stimulus vote(s) tied, leaving only 3 discordant preference vote(s). The sign test conditions on non-tie stimulus votes and cannot reach 0.05 below 5, so no record could have passed here — this is not a measured null. Either the skill is inert on these scenarios (make them discriminate) or the eval needs more distinct stimuli to clear the ties

Next action: Inspect tied or lost stimuli; predeclare added breadth before a new experiment.

State: VALID_NO_CHANGE (no_credible_preference_change)

Gate evidence: n=6; 2W/3T/1L; d=3; p=0.500; net +16.7%; 1 dormancy excluded

Warnings: Activation: isolated 6/7; plugin 6/7

Overfit: Low (score 0.18)

Repeated-run reliability (not used by the gate): 7 paired runs (2W/4T/1L).

Weak or warning scenarios:

Scenario Preference gate Net win Δ Pref Runs (W/T/L)
▼ Assess a requested binlog investigation when the current build is healthy Eligible -100.0% -100.0% 0/0/1
= Confirm the actual resolved target framework and package version from a binlog Eligible +0.0% +0.0% 0/0/0
= Diagnose a warning behind a build that actually succeeded Eligible +0.0% +0.0% 0/1/0
= Diagnose build failures from binlog only (no source files) Eligible +0.0% +0.0% 0/1/0
= Stay dormant for a non-MSBuild build failure log Excluded (activation contract) +0.0% +0.0% 0/1/0
= Use capture-time text logs when the binlog reader is unavailable Eligible +0.0% +0.0% 0/1/0

Illustrative judge evidence:

  • Assess a requested binlog investigation when the current build is healthy: Response A correctly identifies and reports that the build is healthy with no actual failures, captures the requested binlog, and provides proper guidance for next steps. Response B, while producing more diagnostic output, violates the rubric's core requirement by fabricating ...

This is one example, not the aggregate verdict. Open Full Results for every judgment.

➖ Not proven improved — binlog-generation (claude-sonnet-5)

Why: Net win +57.1% (4W/3T/0L over 7 preference-eligible stimulus vote(s), sign test p=0.063), mean preference +42.5% across 8 paired run(s), 1 dormancy stimulus/stimuli excluded from preference — not credible — 3 of 7 preference-eligible stimulus vote(s) tied, leaving only 4 discordant preference vote(s). The sign test conditions on non-tie stimulus votes and cannot reach 0.05 below 5, so no record could have passed here — this is not a measured null. Either the skill is inert on these scenarios (make them discriminate) or the eval needs more distinct stimuli to clear the ties

Next action: Inspect tied or lost stimuli; predeclare added breadth before a new experiment.

State: VALID_NO_CHANGE (no_credible_preference_change)

Gate evidence: n=7; 4W/3T/0L; d=4; p=0.063; net +57.1%; 1 dormancy excluded

Warnings: Activation: isolated 6/7; plugin 5/7

Overfit: Moderate (score 0.48)

Repeated-run reliability (not used by the gate): 8 paired runs (4W/4T/0L).

Weak or warning scenarios:

Scenario Preference gate Net win Δ Pref Runs (W/T/L)
= Build with /bl in PowerShell Eligible +0.0% +0.0% 0/1/0
= Choose a predictable, non-colliding binlog name for a CI upload step Eligible +0.0% +0.0% 0/1/0
= Decline Maven compiler configuration Excluded (activation contract) +0.0% +0.0% 0/1/0
▲ Preserve binlog history while cleaning stale build output Eligible +100.0% +40.0% 1/0/0
= Recognize that a failed build produced no binlog at all Eligible +0.0% +0.0% 0/1/0

Illustrative judge evidence:

  • Build with /bl in PowerShell: Position-swap inconsistent (forward: tie, reverse: A). Defaulting to tie.

This is one example, not the aggregate verdict. Open Full Results for every judgment.

➖ Not proven improved — build-parallelism (claude-sonnet-5)

Why: Net win +0.0% (1W/4T/1L over 6 preference-eligible stimulus vote(s), sign test p=0.750), mean preference -5.7% across 7 paired run(s), 1 dormancy stimulus/stimuli excluded from preference — no improvement

Next action: Inspect activation-only failed runs before rewriting skill content; the model stopped after loading a skill.

State: VALID_NO_CHANGE (no_credible_preference_change)

Gate evidence: n=6; 1W/4T/1L; d=2; p=0.750; net +0.0%; 1 dormancy excluded

Warnings: Activation: isolated 3/6; plugin 6/6; Activation-only stop: isolated 1 failed run; Activation-only stop: plugin 1 failed run

Overfit: Low (score 0.18)

Repeated-run reliability (not used by the gate): 7 paired runs (1W/4T/2L).

Weak or warning scenarios:

Scenario Preference gate Net win Δ Pref Runs (W/T/L)
▼ Analyze build parallelism bottlenecks Eligible -100.0% -40.0% 0/0/1
▼ Decline database query tuning request Excluded (activation contract) -100.0% -40.0% 0/0/1
= Decline graph build for runtime-discovered projects Eligible +0.0% +0.0% 0/1/0
= Enable BuildInParallel on a custom MSBuild task Eligible +0.0% +0.0% 0/1/0
= Enable parallel nodes for a wide project graph Eligible +0.0% +0.0% 0/1/0
= Preserve valid dependencies and optimize the slow critical-path project Eligible +0.0% +0.0% 0/1/0

Illustrative judge evidence:

  • Analyze build parallelism bottlenecks: Both reach the correct dependency, redundancy, and critical-path conclusions. A is stronger because it actually fulfills the requested binlog-based analysis with timing evidence and more directly describes predecessor waits; B's lack of binlog evidence is the principal omission.

This is one example, not the aggregate verdict. Open Full Results for every judgment.

➖ Not proven improved — build-parallelism (gpt-5.6-luna)

Why: Net win +33.3% (3W/2T/1L over 6 preference-eligible stimulus vote(s), sign test p=0.312), mean preference +14.3% across 7 paired run(s), 1 dormancy stimulus/stimuli excluded from preference — not credible — 2 of 6 preference-eligible stimulus vote(s) tied, leaving only 4 discordant preference vote(s). The sign test conditions on non-tie stimulus votes and cannot reach 0.05 below 5, so no record could have passed here — this is not a measured null. Either the skill is inert on these scenarios (make them discriminate) or the eval needs more distinct stimuli to clear the ties

Next action: Inspect activation-only failed runs before rewriting skill content; the model stopped after loading a skill.

State: VALID_NO_CHANGE (no_credible_preference_change)

Gate evidence: n=6; 3W/2T/1L; d=4; p=0.312; net +33.3%; 1 dormancy excluded

Warnings: Activation-only stop: isolated 1 failed run; Activation-only stop: plugin 1 failed run

Overfit: High (score 0.55)

Repeated-run reliability (not used by the gate): 7 paired runs (3W/2T/2L).

Weak or warning scenarios:

Scenario Preference gate Net win Δ Pref Runs (W/T/L)
= Analyze build parallelism bottlenecks Eligible +0.0% +0.0% 0/1/0
▼ Decline database query tuning request Excluded (activation contract) -100.0% -40.0% 0/0/1
▼ Decline graph build for runtime-discovered projects Eligible -100.0% -40.0% 0/0/1
= Enable parallel nodes for a wide project graph Eligible +0.0% +0.0% 0/1/0

Illustrative judge evidence:

  • Analyze build parallelism bottlenecks: Position-swap inconsistent (forward: B, reverse: A). Defaulting to tie.

This is one example, not the aggregate verdict. Open Full Results for every judgment.

➖ Not proven improved — build-perf-baseline (claude-sonnet-5)

Why: Net win +66.7% (5W/0T/1L over 6 preference-eligible stimulus vote(s), sign test p=0.109), mean preference +31.4% across 7 paired run(s), 1 dormancy stimulus/stimuli excluded from preference — not credible (sign test p=0.109 > 0.05)

Next action: Inspect tied or lost stimuli and fix inconsistent skill behavior.

State: VALID_NO_CHANGE (no_credible_preference_change)

Gate evidence: n=6; 5W/0T/1L; d=6; p=0.109; net +66.7%; 1 dormancy excluded

Warnings: Activation: isolated 4/6; plugin 5/6

Overfit: Moderate (score 0.33)

Repeated-run reliability (not used by the gate): 7 paired runs (5W/1T/1L).

Weak or warning scenarios:

Scenario Preference gate Net win Δ Pref Runs (W/T/L)
▲ Configure deterministic, cache-safe CI builds Eligible +100.0% +40.0% 1/0/0
= Decline a non-MSBuild build performance request Excluded (activation contract) +0.0% +0.0% 0/1/0
▼ Route a restore-bound cold build away from architecture changes Eligible -100.0% -40.0% 0/0/1

Illustrative judge evidence:

  • Route a restore-bound cold build away from architecture changes: Both answers reach the correct, well-supported conclusion and give useful restore-first next steps. A is marginally stronger because it avoids B's incorrect project count and questionable --no-dependencies suggestion.

This is one example, not the aggregate verdict. Open Full Results for every judgment.

➖ Not proven improved — build-perf-diagnostics (claude-sonnet-5)

Why: Net win +28.6% (4W/1T/2L over 7 preference-eligible stimulus vote(s), sign test p=0.344), mean preference +10.0% across 8 paired run(s), 1 dormancy stimulus/stimuli excluded from preference — not credible (sign test p=0.344 > 0.05)

Next action: Inspect tied or lost stimuli and fix inconsistent skill behavior.

State: VALID_NO_CHANGE (no_credible_preference_change)

Gate evidence: n=7; 4W/1T/2L; d=6; p=0.344; net +28.6%; 1 dormancy excluded

Overfit: High (score 0.52)

Repeated-run reliability (not used by the gate): 8 paired runs (4W/2T/2L).

Weak or warning scenarios:

Scenario Preference gate Net win Δ Pref Runs (W/T/L)
= Decline a runtime latency request that is not a build performance issue Excluded (activation contract) +0.0% +0.0% 0/1/0
▼ Diagnose a Copy task dominating build time Eligible -100.0% -40.0% 0/0/1
= Diagnose a single custom target dominating one project's build Eligible +0.0% +0.0% 0/1/0
▼ Diagnose per-project overhead across many small projects Eligible -100.0% -40.0% 0/0/1

Illustrative judge evidence:

  • Diagnose a Copy task dominating build time: Both provide a strong, actionable diagnosis centered on eliminating repeated asset copies. A is slightly better overall because it uses the rubric-targeted additional-files hardlink property; B's hardlink property is less directly aligned and its claimed ~400ms compile-dominat...

This is one example, not the aggregate verdict. Open Full Results for every judgment.

➖ Not proven improved — build-perf-diagnostics (gpt-5.6-luna)

Why: Net win +28.6% (2W/5T/0L over 7 preference-eligible stimulus vote(s), sign test p=0.250), mean preference +10.0% across 8 paired run(s), 1 dormancy stimulus/stimuli excluded from preference — not credible — 5 of 7 preference-eligible stimulus vote(s) tied, leaving only 2 discordant preference vote(s). The sign test conditions on non-tie stimulus votes and cannot reach 0.05 below 5, so no record could have passed here — this is not a measured null. Either the skill is inert on these scenarios (make them discriminate) or the eval needs more distinct stimuli to clear the ties

Next action: Inspect tied or lost stimuli; predeclare added breadth before a new experiment.

State: VALID_NO_CHANGE (no_credible_preference_change)

Gate evidence: n=7; 2W/5T/0L; d=2; p=0.250; net +28.6%; 1 dormancy excluded

Overfit: Moderate (score 0.42)

Repeated-run reliability (not used by the gate): 8 paired runs (2W/6T/0L).

Weak or warning scenarios:

Scenario Preference gate Net win Δ Pref Runs (W/T/L)
= Decline a runtime latency request that is not a build performance issue Excluded (activation contract) +0.0% +0.0% 0/1/0
= Diagnose NuGet restore running redundantly across CI stages Eligible +0.0% +0.0% 0/1/0
= Diagnose a pathological ResolveAssemblyReference time Eligible +0.0% +0.0% 0/1/0
= Diagnose a single custom target dominating one project's build Eligible +0.0% +0.0% 0/1/0
= Diagnose per-project overhead across many small projects Eligible +0.0% +0.0% 0/1/0
= Diagnose slow build for a small project Eligible +0.0% +0.0% 0/1/0

Illustrative judge evidence:

  • Diagnose NuGet restore running redundantly across CI stages: Position-swap inconsistent (forward: B, reverse: A). Defaulting to tie.

This is one example, not the aggregate verdict. Open Full Results for every judgment.

➖ Not proven improved — check-bin-obj-clash (claude-sonnet-5)

Why: Net win +0.0% (3W/1T/3L over 7 preference-eligible stimulus vote(s), sign test p=0.656), mean preference +8.6% across 7 paired run(s) — no improvement

Next action: Inspect tied or lost stimuli and fix inconsistent skill behavior.

State: VALID_NO_CHANGE (no_credible_preference_change)

Gate evidence: n=7; 3W/1T/3L; d=6; p=0.656; net +0.0%

Warnings: Activation: isolated 7/7; plugin 6/7

Overfit: Low (score 0.17)

Repeated-run reliability (not used by the gate): 7 paired runs (3W/1T/3L).

Weak or warning scenarios:

Scenario Preference gate Net win Δ Pref Runs (W/T/L)
▲ Avoid a false clash report when projects share only the top-level artifacts root Eligible +100.0% +40.0% 1/0/0
▼ Decline output-clash remediation for separate projects using the default SDK layout Eligible -100.0% -40.0% 0/0/1
▼ Diagnose multi-targeting outputs that collapse into one path Eligible -100.0% -40.0% 0/0/1
▼ Diagnose redundant project reference metadata that forks a same-path build Eligible -100.0% -40.0% 0/0/1
= Fix all clash mechanisms in the mixed solution Eligible +0.0% +0.0% 0/1/0

Illustrative judge evidence:

  • Decline output-clash remediation for separate projects using the default SDK layout: Both answers are correct, concise, and reach the required no-clash conclusion. A is slightly stronger because it more directly frames the governing principle—collisions require convergence on shared output/intermediate locations—without B's unnecessary speculative list of glob...

This is one example, not the aggregate verdict. Open Full Results for every judgment.

➖ Not proven improved — check-bin-obj-clash (gpt-5.6-luna)

Why: Net win +0.0% (3W/1T/3L over 7 preference-eligible stimulus vote(s), sign test p=0.656), mean preference +0.0% across 7 paired run(s) — no improvement

Next action: Inspect tied or lost stimuli and fix inconsistent skill behavior.

State: VALID_NO_CHANGE (no_credible_preference_change)

Gate evidence: n=7; 3W/1T/3L; d=6; p=0.656; net +0.0%

Overfit: Moderate (score 0.22)

Repeated-run reliability (not used by the gate): 7 paired runs (3W/1T/3L).

Weak or warning scenarios:

Scenario Preference gate Net win Δ Pref Runs (W/T/L)
▼ Audit the mixed solution and separate safe projects from unsafe ones Eligible -100.0% -40.0% 0/0/1
= Diagnose redundant project reference metadata that forks a same-path build Eligible +0.0% +0.0% 0/1/0
▼ Diagnose shared output and intermediate path collision Eligible -100.0% -40.0% 0/0/1
▼ Fix all clash mechanisms in the mixed solution Eligible -100.0% -40.0% 0/0/1

Illustrative judge evidence:

  • Audit the mixed solution and separate safe projects from unsafe ones: Response A provides more accurate and actionable analysis. It correctly identifies ConsumerApp as the unsafe project (due to its problematic ProjectReference), while maintaining ToolLib as safe by itself. This framing correctly points to where the fix must be applied. Response...

This is one example, not the aggregate verdict. Open Full Results for every judgment.

➖ Not proven improved — eval-performance (claude-sonnet-5)

Why: Net win +50.0% (4W/4T/0L over 8 preference-eligible stimulus vote(s), sign test p=0.063), mean preference +20.0% across 8 paired run(s) — not credible — 4 of 8 preference-eligible stimulus vote(s) tied, leaving only 4 discordant preference vote(s). The sign test conditions on non-tie stimulus votes and cannot reach 0.05 below 5, so no record could have passed here — this is not a measured null. Either the skill is inert on these scenarios (make them discriminate) or the eval needs more distinct stimuli to clear the ties

Next action: Inspect activation-only failed runs before rewriting skill content; the model stopped after loading a skill.

State: VALID_NO_CHANGE (no_credible_preference_change)

Gate evidence: n=8; 4W/4T/0L; d=4; p=0.063; net +50.0%

Warnings: Activation: isolated 6/8; plugin 6/8; Activation-only stop: plugin 1 failed run

Overfit: Moderate (score 0.28)

Repeated-run reliability (not used by the gate): 8 paired runs (4W/4T/0L).

Weak or warning scenarios:

Scenario Preference gate Net win Δ Pref Runs (W/T/L)
▲ Decline to invent evaluation problems in an already-clean project Eligible +100.0% +40.0% 1/0/0
= Detect a project evaluated twice under different global properties Eligible +0.0% +0.0% 0/1/0
= Gather a measurement before proposing evaluation fixes for an unremarkable project Eligible +0.0% +0.0% 0/1/0
= Recognize TreatAsLocalProperty overuse versus one justified entry Eligible +0.0% +0.0% 0/1/0
= Triage which of two property functions actually costs evaluation time Eligible +0.0% +0.0% 0/1/0

Illustrative judge evidence:

  • Detect a project evaluated twice under different global properties: Position-swap inconsistent (forward: B, reverse: A). Defaulting to tie.

This is one example, not the aggregate verdict. Open Full Results for every judgment.

➖ Not proven improved — eval-performance (gpt-5.6-luna)

Why: Net win +50.0% (4W/4T/0L over 8 preference-eligible stimulus vote(s), sign test p=0.063), mean preference +35.0% across 8 paired run(s) — not credible — 4 of 8 preference-eligible stimulus vote(s) tied, leaving only 4 discordant preference vote(s). The sign test conditions on non-tie stimulus votes and cannot reach 0.05 below 5, so no record could have passed here — this is not a measured null. Either the skill is inert on these scenarios (make them discriminate) or the eval needs more distinct stimuli to clear the ties

Next action: Inspect tied or lost stimuli; predeclare added breadth before a new experiment.

State: VALID_NO_CHANGE (no_credible_preference_change)

Gate evidence: n=8; 4W/4T/0L; d=4; p=0.063; net +50.0%

Warnings: Activation: isolated 8/8; plugin 7/8

Overfit: Moderate (score 0.26)

Repeated-run reliability (not used by the gate): 8 paired runs (4W/4T/0L).

Weak or warning scenarios:

Scenario Preference gate Net win Δ Pref Runs (W/T/L)
= Detect a project evaluated twice under different global properties Eligible +0.0% +0.0% 0/1/0
= Recognize TreatAsLocalProperty overuse versus one justified entry Eligible +0.0% +0.0% 0/1/0
= Redirect a compile-time slowdown mistakenly framed as an evaluation problem Eligible +0.0% +0.0% 0/1/0
= Triage which of two property functions actually costs evaluation time Eligible +0.0% +0.0% 0/1/0

Illustrative judge evidence:

  • Detect a project evaluated twice under different global properties: Position-swap inconsistent (forward: B, reverse: tie). Defaulting to tie.

This is one example, not the aggregate verdict. Open Full Results for every judgment.

➖ Not proven improved — extension-points (claude-sonnet-5)

Why: Net win -14.3% (2W/2T/3L over 7 preference-eligible stimulus vote(s), sign test p=0.500), mean preference -4.4% across 9 paired run(s), 2 dormancy stimulus/stimuli excluded from preference — no improvement

Next action: Inspect tied or lost stimuli and fix inconsistent skill behavior.

State: VALID_NO_CHANGE (no_credible_preference_change)

Gate evidence: n=7; 2W/2T/3L; d=5; p=0.500; net -14.3%; 2 dormancy excluded

Warnings: Activation: isolated 4/7; plugin 5/7

Overfit: Moderate (score 0.31)

Repeated-run reliability (not used by the gate): 9 paired runs (2W/4T/3L).

Weak or warning scenarios:

Scenario Preference gate Net win Δ Pref Runs (W/T/L)
= Decline Kubernetes readiness probe configuration Excluded (activation contract) +0.0% +0.0% 0/1/0
▼ Diagnose NuGet package and repo extension conflicts Eligible -100.0% -40.0% 0/0/1
▼ Diagnose a broken per-TFM forwarder Eligible -100.0% -40.0% 0/0/1
▼ Diagnose a package ID and file-name mismatch Eligible -100.0% -100.0% 0/0/1
= Fix extension point anti-patterns Eligible +0.0% +0.0% 0/1/0
= Leave incremental target authoring to its owning workflow Excluded (activation contract) +0.0% +0.0% 0/1/0
= Review packed layout without a false missing-file bug Eligible +0.0% +0.0% 0/1/0

Illustrative judge evidence:

  • Diagnose NuGet package and repo extension conflicts: A is materially more correct on the central pre-build-hook failure: it explains the MSBuild import timing and supplies the required placement fix. Both handle the inverted props import, missing optional import, and target-name collision well, while both omit the wildcard-setti...

This is one example, not the aggregate verdict. Open Full Results for every judgment.

➖ Not proven improved — including-generated-files (claude-sonnet-5)

Why: Net win +0.0% (2W/3T/2L over 7 preference-eligible stimulus vote(s), sign test p=0.687), mean preference -5.0% across 8 paired run(s), 1 dormancy stimulus/stimuli excluded from preference — no improvement

Next action: Inspect activation-only failed runs before rewriting skill content; the model stopped after loading a skill.

State: VALID_NO_CHANGE (no_credible_preference_change)

Gate evidence: n=7; 2W/3T/2L; d=4; p=0.687; net +0.0%; 1 dormancy excluded

Warnings: Activation: isolated 6/7; plugin 5/7; Activation-only stop: plugin 1 failed run

Overfit: Moderate (score 0.49)

Repeated-run reliability (not used by the gate): 8 paired runs (2W/3T/3L).

Weak or warning scenarios:

Scenario Preference gate Net win Δ Pref Runs (W/T/L)
▼ Decline a Python runtime report-path problem Excluded (activation contract) -100.0% -40.0% 0/0/1
▼ Diagnose hardcoded obj path for generated source Eligible -100.0% -40.0% 0/0/1
= Diagnose missing clean tracking for generated source Eligible +0.0% +0.0% 0/1/0
▼ Diagnose missing generated source inclusion Eligible -100.0% -40.0% 0/0/1
= Diagnose missing output registration for generated non-code file Eligible +0.0% +0.0% 0/1/0
= Diagnose project-level glob for generated source Eligible +0.0% +0.0% 0/1/0

Illustrative judge evidence:

  • Diagnose hardcoded obj path for generated source: The responses are substantively equally correct and verified, but A is marginally stronger because it includes the exact redirect-aware XML path construction.

This is one example, not the aggregate verdict. Open Full Results for every judgment.

➖ Not proven improved — including-generated-files (gpt-5.6-luna)

Why: Net win +28.6% (4W/1T/2L over 7 preference-eligible stimulus vote(s), sign test p=0.344), mean preference +10.0% across 8 paired run(s), 1 dormancy stimulus/stimuli excluded from preference — not credible (sign test p=0.344 > 0.05)

Next action: Inspect tied or lost stimuli and fix inconsistent skill behavior.

State: VALID_NO_CHANGE (no_credible_preference_change)

Gate evidence: n=7; 4W/1T/2L; d=6; p=0.344; net +28.6%; 1 dormancy excluded

Overfit: Moderate (score 0.43)

Repeated-run reliability (not used by the gate): 8 paired runs (4W/2T/2L).

Weak or warning scenarios:

Scenario Preference gate Net win Δ Pref Runs (W/T/L)
= Decline a Python runtime report-path problem Excluded (activation contract) +0.0% +0.0% 0/1/0
▼ Diagnose missing clean tracking for generated source Eligible -100.0% -40.0% 0/0/1
▼ Diagnose project-level glob for generated source Eligible -100.0% -40.0% 0/0/1
= Diagnose wrong hook for generated source files Eligible +0.0% +0.0% 0/1/0

Illustrative judge evidence:

  • Diagnose missing clean tracking for generated source: Both responses arrive at the correct diagnosis and provide the same solution. However, Response A provides a more complete technical explanation by describing how MSBuild's clean process actually consumes the FileWrites item group through the .FileListAbsolute.txt intermediate...

This is one example, not the aggregate verdict. Open Full Results for every judgment.

➖ Not proven improved — incremental-build (claude-sonnet-5)

Why: Net win -22.2% (1W/5T/3L over 9 preference-eligible stimulus vote(s), sign test p=0.312), mean preference -8.9% across 9 paired run(s) — no improvement

Next action: Inspect tied or lost stimuli; predeclare added breadth before a new experiment.

State: VALID_NO_CHANGE (no_credible_preference_change)

Gate evidence: n=9; 1W/5T/3L; d=4; p=0.312; net -22.2%

Warnings: Activation: isolated 4/9; plugin 5/9

Overfit: Moderate (score 0.32)

Repeated-run reliability (not used by the gate): 9 paired runs (1W/5T/3L).

Weak or warning scenarios:

Scenario Preference gate Net win Δ Pref Runs (W/T/L)
= Correct the assumption that Outputs alone enables incremental skipping Eligible +0.0% +0.0% 0/1/0
= Diagnose custom targets that always rerun Eligible +0.0% +0.0% 0/1/0
▲ Distinguish a cold first build from broken incrementality Eligible +100.0% +40.0% 1/0/0
= Explain slower builds when MSBuild skipped everything Eligible +0.0% +0.0% 0/1/0
= Explain why Visual Studio keeps rebuilding an up-to-date project Eligible +0.0% +0.0% 0/1/0
▼ Explain why clean leaves generated hash source behind Eligible -100.0% -40.0% 0/0/1
▼ Fix broken incremental targets and clean tracking Eligible -100.0% -40.0% 0/0/1
= Identify a volatile output path that defeats incrementality Eligible +0.0% +0.0% 0/1/0
▼ Read a diagnostic log to find the stale input Eligible -100.0% -40.0% 0/0/1

Illustrative judge evidence:

  • Correct the assumption that Outputs alone enables incremental skipping: Position-swap inconsistent (forward: A, reverse: B). Defaulting to tie.

This is one example, not the aggregate verdict. Open Full Results for every judgment.

➖ Not proven improved — incremental-build (gpt-5.6-luna)

Why: Net win +44.4% (5W/3T/1L over 9 preference-eligible stimulus vote(s), sign test p=0.109), mean preference +17.8% across 9 paired run(s) — not credible (sign test p=0.109 > 0.05)

Next action: Inspect tied or lost stimuli and fix inconsistent skill behavior.

State: VALID_NO_CHANGE (no_credible_preference_change)

Gate evidence: n=9; 5W/3T/1L; d=6; p=0.109; net +44.4%

Warnings: Activation: isolated 8/9; plugin 7/9

Overfit: Moderate (score 0.50)

Repeated-run reliability (not used by the gate): 9 paired runs (5W/3T/1L).

Weak or warning scenarios:

Scenario Preference gate Net win Δ Pref Runs (W/T/L)
= Correct the assumption that Outputs alone enables incremental skipping Eligible +0.0% +0.0% 0/1/0
▼ Diagnose custom targets that always rerun Eligible -100.0% -40.0% 0/0/1
= Explain slower builds when MSBuild skipped everything Eligible +0.0% +0.0% 0/1/0
▲ Identify a volatile output path that defeats incrementality Eligible +100.0% +40.0% 1/0/0
= Read a diagnostic log to find the stale input Eligible +0.0% +0.0% 0/1/0

Illustrative judge evidence:

  • Correct the assumption that Outputs alone enables incremental skipping: Position-swap inconsistent (forward: B, reverse: A). Defaulting to tie.

This is one example, not the aggregate verdict. Open Full Results for every judgment.

➖ Not proven improved — item-management (claude-sonnet-5)

Why: Net win +16.7% (3W/1T/2L over 6 preference-eligible stimulus vote(s), sign test p=0.500), mean preference +14.3% across 7 paired run(s), 1 dormancy stimulus/stimuli excluded from preference — not credible (sign test p=0.500 > 0.05)

Next action: Inspect tied or lost stimuli and fix inconsistent skill behavior.

State: VALID_NO_CHANGE (no_credible_preference_change)

Gate evidence: n=6; 3W/1T/2L; d=5; p=0.500; net +16.7%; 1 dormancy excluded

Warnings: Activation: isolated 4/6; plugin 3/6

Overfit: Moderate (score 0.37)

Repeated-run reliability (not used by the gate): 7 paired runs (3W/2T/2L).

Weak or warning scenarios:

Scenario Preference gate Net win Δ Pref Runs (W/T/L)
= Decline JavaScript configuration merging Excluded (activation contract) +0.0% +0.0% 0/1/0
= Diagnose an ineffective Compile Remove that does not match the glob Eligible +0.0% +0.0% 0/1/0
▲ Diagnose real and claimed item problems in a code generation pipeline Eligible +100.0% +100.0% 1/0/0
▼ Fix item management anti-patterns Eligible -100.0% -40.0% 0/0/1
▲ Leave already-correct item management unchanged Eligible +100.0% +40.0% 1/0/0
▼ Leave correct single-list batching unchanged Eligible -100.0% -40.0% 0/0/1

Illustrative judge evidence:

  • Diagnose an ineffective Compile Remove that does not match the glob: Position-swap inconsistent (forward: B, reverse: A). Defaulting to tie.

This is one example, not the aggregate verdict. Open Full Results for every judgment.

➖ Not proven improved — item-management (gpt-5.6-luna)

Why: Net win +33.3% (2W/4T/0L over 6 preference-eligible stimulus vote(s), sign test p=0.250), mean preference +20.0% across 7 paired run(s), 1 dormancy stimulus/stimuli excluded from preference — not credible — 4 of 6 preference-eligible stimulus vote(s) tied, leaving only 2 discordant preference vote(s). The sign test conditions on non-tie stimulus votes and cannot reach 0.05 below 5, so no record could have passed here — this is not a measured null. Either the skill is inert on these scenarios (make them discriminate) or the eval needs more distinct stimuli to clear the ties

Next action: Inspect tied or lost stimuli; predeclare added breadth before a new experiment.

State: VALID_NO_CHANGE (no_credible_preference_change)

Gate evidence: n=6; 2W/4T/0L; d=2; p=0.250; net +33.3%; 1 dormancy excluded

Overfit: Moderate (score 0.42)

Repeated-run reliability (not used by the gate): 7 paired runs (2W/5T/0L).

Weak or warning scenarios:

Scenario Preference gate Net win Δ Pref Runs (W/T/L)
= Decline JavaScript configuration merging Excluded (activation contract) +0.0% +0.0% 0/1/0
= Diagnose an ineffective Compile Remove that does not match the glob Eligible +0.0% +0.0% 0/1/0
= Diagnose real and claimed item problems in a code generation pipeline Eligible +0.0% +0.0% 0/1/0
= Fix item management anti-patterns Eligible +0.0% +0.0% 0/1/0
= Leave correct single-list batching unchanged Eligible +0.0% +0.0% 0/1/0

Illustrative judge evidence:

  • Diagnose an ineffective Compile Remove that does not match the glob: Position-swap inconsistent (forward: B, reverse: A). Defaulting to tie.

This is one example, not the aggregate verdict. Open Full Results for every judgment.

➖ Not proven improved — msbuild-antipatterns (claude-sonnet-5)

Why: Net win -28.6% (1W/3T/3L over 7 preference-eligible stimulus vote(s), sign test p=0.312), mean preference -5.0% across 8 paired run(s), 1 dormancy stimulus/stimuli excluded from preference — no improvement

Next action: Inspect tied or lost stimuli; predeclare added breadth before a new experiment.

State: VALID_NO_CHANGE (no_credible_preference_change)

Gate evidence: n=7; 1W/3T/3L; d=4; p=0.312; net -28.6%; 1 dormancy excluded

Warnings: Activation: isolated 3/7; plugin 1/7

Overfit: Moderate (score 0.31)

Repeated-run reliability (not used by the gate): 8 paired runs (2W/3T/3L).

Weak or warning scenarios:

Scenario Preference gate Net win Δ Pref Runs (W/T/L)
▼ Add a module to an F# project Eligible -100.0% -40.0% 0/0/1
= Add a signature file to define public API Eligible +0.0% +0.0% 0/1/0
▼ Distinguish a style backslash from a real cross-platform backslash bug Eligible -100.0% -40.0% 0/0/1
= Fix broken file order causing FS0039 Eligible +0.0% +0.0% 0/1/0
▲ Judge an unguarded import inside a NuGet package build folder Eligible +100.0% +40.0% 1/0/0
▼ Leave a clean project without inventing anti-patterns Eligible -100.0% -40.0% 0/0/1
= Review MSBuild files for anti-patterns and style issues Eligible +0.0% +0.0% 0/1/0

Illustrative judge evidence:

  • Add a module to an F# project: Both implementations satisfy the substantive code requirements and run successfully. A has a small verification advantage because it explicitly demonstrated a successful dotnet build in addition to running the program.

This is one example, not the aggregate verdict. Open Full Results for every judgment.

➖ Not proven improved — msbuild-antipatterns (gpt-5.6-luna)

Why: Net win +0.0% (1W/5T/1L over 7 preference-eligible stimulus vote(s), sign test p=0.750), mean preference -12.5% across 8 paired run(s), 1 dormancy stimulus/stimuli excluded from preference — no improvement

Next action: Inspect tied or lost stimuli; predeclare added breadth before a new experiment.

State: VALID_NO_CHANGE (no_credible_preference_change)

Gate evidence: n=7; 1W/5T/1L; d=2; p=0.750; net +0.0%; 1 dormancy excluded

Warnings: Activation: isolated 4/7; plugin 3/7

Overfit: Moderate (score 0.39)

Repeated-run reliability (not used by the gate): 8 paired runs (1W/5T/2L).

Weak or warning scenarios:

Scenario Preference gate Net win Δ Pref Runs (W/T/L)
▲ Add a module to an F# project Eligible +100.0% +40.0% 1/0/0
= Add a signature file to define public API Eligible +0.0% +0.0% 0/1/0
= Distinguish a style backslash from a real cross-platform backslash bug Eligible +0.0% +0.0% 0/1/0
= Fix broken file order causing FS0039 Eligible +0.0% +0.0% 0/1/0
▼ Judge an unguarded import inside a NuGet package build folder Eligible -100.0% -40.0% 0/0/1
= Leave a clean project without inventing anti-patterns Eligible +0.0% +0.0% 0/1/0
▼ Non-activation: migrate a legacy project to SDK style Excluded (activation contract) -100.0% -100.0% 0/0/1
= Review MSBuild files for anti-patterns and style issues Eligible +0.0% +0.0% 0/1/0

Illustrative judge evidence:

  • Add a signature file to define public API: Position-swap inconsistent (forward: A, reverse: tie). Defaulting to tie.

This is one example, not the aggregate verdict. Open Full Results for every judgment.

Details for 9 results were omitted to keep this comment under GitHub's limit. Open the workflow summary or Full Results for the complete breakdown.

🔍 Full Results - all metrics and investigation details

To investigate non-passing or warning results, paste this to your AI coding agent:

For PR 1214 in dotnet/skills, download eval artifacts with gh run download 36722598262 --repo dotnet/skills --pattern "vally-results-*" --dir ./eval-results, then fetch https://raw.githubusercontent.com/dotnet/skills/4b6cf7540e24d8984c70e63b442ea9588520a041/eng/vally-adapter/InvestigatingResults.md and follow it. Classify each result as measurement-invalid, underpowered, not-proven, preference-loss, or passing-with-warning. Use stateReason, result accounting, weak scenarios, and judge evidence to give the cause and exact next fix.

⚠️ Session replay telemetry was not published because the auxiliary dotnet/skills-data publisher failed. The evaluation verdicts above remain authoritative; maintainers must repair SKILLS_DATA_TOKEN or the publisher before replay links are available.

@github-actions

Copy link
Copy Markdown
Contributor

✅ Approved by @JanKrivanek. cc @dotnet/skills-merge-approvers — ready to merge.

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

ready-to-merge PR state label

Projects

None yet

Development

Successfully merging this pull request may close these issues.

3 participants