test(msbuild): strengthen evaluation coverage and infrastructure - #1214
Conversation
Re-carve the MSBuild evaluation specs, fixtures, setup scripts, and golden evidence from the validated source branch onto current main. Co-authored-by: Copilot App <223556219+Copilot@users.noreply.github.com>
Add deterministic eval-quality checks, secure evaluator filesystem boundaries, exact rejudge accounting, and target-specific activation evidence required by the MSBuild evaluation corpus. Co-authored-by: Copilot App <223556219+Copilot@users.noreply.github.com>
Reject interrupted rejudge data, secure directory operations against path swaps, reuse private MCP caches, and keep manual quality runs scoped to main. Co-authored-by: Copilot App <223556219+Copilot@users.noreply.github.com>
Use Darwin dirent and unlinkat layouts for handle-relative enumeration and recursive removal, with macOS-specific regression coverage. Co-authored-by: Copilot App <223556219+Copilot@users.noreply.github.com>
Skill Coverage Report
Uncovered:
|
Restrict Darwin dirent parsing to arm64, clear native errno correctly, and rely on the existing cross-platform filesystem regression coverage. Co-authored-by: Copilot App <223556219+Copilot@users.noreply.github.com>
There was a problem hiding this comment.
Copilot review overview
🟡 Changes recommended
Critical and moderate issues remain in workflow coverage, shard discovery, rejudge pairing, filesystem deletion, dormancy handling, and a fixture patch.
Get a fresh assessment by requesting another Copilot review.
Review effort: Lite
Findings: 5
Open (6)
Run all eval specs when checker or workflow logic changes · New Fix shard matcher ordering to recognize indented executionShard · New Require exact baseline run-index pairing · New Treat unmatched dormancy names as activation contract failures · New Retain generated file Compile include alongside FileWrites · New Reject duplicate sessions for inline rejudge roles · New
What changed in this PR
This PR expands MSBuild evaluation coverage with deterministic fixtures and strengthens shared evaluation infrastructure for isolation, activation tracking, rejudge pairing, and quality validation.
Changes:
- Adds MSBuild fixtures, reference patches, performance artifacts, and routing scenarios.
- Hardens evaluator filesystem, session, cache, activation, and dormancy handling.
- Updates quality validation, shard discovery, and rejudge accounting.
| File | Change |
|---|---|
tests/dotnet-msbuild/target-authoring/TargetAuthoring.csproj |
Reviewed |
tests/dotnet-msbuild/target-authoring/returns-outputs/ReturnsOutputs.csproj |
Reviewed |
tests/dotnet-msbuild/target-authoring/returns-outputs/Placeholder.cs |
Reviewed |
tests/dotnet-msbuild/target-authoring/references/returns-outputs.patch |
Reviewed |
tests/dotnet-msbuild/target-authoring/references/fix-antipatterns.patch |
Reviewed |
tests/dotnet-msbuild/target-authoring/Placeholder.cs |
Reviewed |
tests/dotnet-msbuild/target-authoring/dormancy/Placeholder.cs |
Reviewed |
tests/dotnet-msbuild/target-authoring/dormancy/perf-summary.txt |
Reviewed |
tests/dotnet-msbuild/target-authoring/dormancy/IncrementalTuning.csproj |
Reviewed |
tests/dotnet-msbuild/target-authoring/dormancy/Api.schema |
Reviewed |
tests/dotnet-msbuild/target-authoring/correct-target/schemas/api.schema |
Reviewed |
tests/dotnet-msbuild/target-authoring/correct-target/Placeholder.cs |
Reviewed |
tests/dotnet-msbuild/target-authoring/correct-target/CorrectTarget.csproj |
Reviewed |
tests/dotnet-msbuild/target-authoring/before-targets/Placeholder.cs |
Reviewed |
tests/dotnet-msbuild/target-authoring/before-targets/BeforeTargets.csproj |
Reviewed |
tests/dotnet-msbuild/target-authoring/before-targets/Api.schema |
Reviewed |
tests/dotnet-msbuild/target-authoring/ApiClient.schema |
Reviewed |
tests/dotnet-msbuild/resolve-project-references/wide-graph/perf-report-wide-graph.md |
Reviewed |
tests/dotnet-msbuild/resolve-project-references/wide-graph/Core/CoreContracts.cs |
Reviewed |
tests/dotnet-msbuild/resolve-project-references/wide-graph/Core/Core.csproj |
Reviewed |
tests/dotnet-msbuild/resolve-project-references/wide-graph/App4/App4Feature.cs |
Reviewed |
tests/dotnet-msbuild/resolve-project-references/wide-graph/App4/App4.csproj |
Reviewed |
tests/dotnet-msbuild/resolve-project-references/wide-graph/App3/App3Feature.cs |
Reviewed |
tests/dotnet-msbuild/resolve-project-references/wide-graph/App3/App3.csproj |
Reviewed |
tests/dotnet-msbuild/resolve-project-references/wide-graph/App2/App2Feature.cs |
Reviewed |
tests/dotnet-msbuild/resolve-project-references/wide-graph/App2/App2.csproj |
Reviewed |
tests/dotnet-msbuild/resolve-project-references/wide-graph/App1/App1Feature.cs |
Reviewed |
tests/dotnet-msbuild/resolve-project-references/wide-graph/App1/App1.csproj |
Reviewed |
tests/dotnet-msbuild/resolve-project-references/wide-graph/Aggregator/AggregatorReport.cs |
Reviewed |
tests/dotnet-msbuild/resolve-project-references/wide-graph/Aggregator/Aggregator.csproj |
Reviewed |
tests/dotnet-msbuild/resolve-project-references/serial-chain/Web/WebService.cs |
Reviewed |
tests/dotnet-msbuild/resolve-project-references/serial-chain/Web/Web.csproj |
Reviewed |
tests/dotnet-msbuild/resolve-project-references/serial-chain/Tests/Tests.csproj |
Reviewed |
tests/dotnet-msbuild/resolve-project-references/serial-chain/Tests/SmokeTests.cs |
Reviewed |
tests/dotnet-msbuild/resolve-project-references/serial-chain/perf-report-serial-chain.md |
Reviewed |
tests/dotnet-msbuild/resolve-project-references/serial-chain/Core/CoreService.cs |
Reviewed |
tests/dotnet-msbuild/resolve-project-references/serial-chain/Core/Core.csproj |
Reviewed |
tests/dotnet-msbuild/resolve-project-references/serial-chain/Api/ApiService.cs |
Reviewed |
tests/dotnet-msbuild/resolve-project-references/serial-chain/Api/Api.csproj |
Reviewed |
tests/dotnet-msbuild/resolve-project-references/perf-report-general-build.md |
Reviewed |
tests/dotnet-msbuild/resolve-project-references/perf-report-csc-vs-target.md |
Reviewed |
tests/dotnet-msbuild/resolve-project-references/perf-report-csc-obvious.md |
Reviewed |
tests/dotnet-msbuild/resolve-project-references/perf-report-copy-vs-csc.md |
Reviewed |
tests/dotnet-msbuild/property-patterns/tfm-timing/TfmTiming.csproj |
Reviewed |
tests/dotnet-msbuild/property-patterns/tfm-timing/Placeholder.cs |
Reviewed |
tests/dotnet-msbuild/property-patterns/tfm-timing/Directory.Build.props |
Reviewed |
tests/dotnet-msbuild/property-patterns/references/fix-shared.patch |
Reviewed |
tests/dotnet-msbuild/property-patterns/placement/Placement.csproj |
Reviewed |
tests/dotnet-msbuild/property-patterns/placement/Placeholder.cs |
Reviewed |
tests/dotnet-msbuild/property-patterns/placement/Directory.Build.props |
Reviewed |
tests/dotnet-msbuild/property-patterns/os-check/Placeholder.cs |
Reviewed |
tests/dotnet-msbuild/property-patterns/os-check/OsCheck.csproj |
Reviewed |
tests/dotnet-msbuild/property-patterns/os-check/Directory.Build.props |
Reviewed |
tests/dotnet-msbuild/property-patterns/correct-props/Placeholder.cs |
Reviewed |
tests/dotnet-msbuild/property-patterns/correct-props/Directory.Build.props |
Reviewed |
tests/dotnet-msbuild/property-patterns/correct-props/CorrectProps.csproj |
Reviewed |
tests/dotnet-msbuild/msbuild-modernization/references/langversion-trap.patch |
Reviewed |
tests/dotnet-msbuild/msbuild-modernization/references/cpm-restraint.patch |
Reviewed |
tests/dotnet-msbuild/msbuild-modernization/partial-modern/Repository.cs |
Reviewed |
tests/dotnet-msbuild/msbuild-modernization/partial-modern/PartialModern.csproj |
Reviewed |
tests/dotnet-msbuild/msbuild-modernization/node/webpack.config.js |
Reviewed |
tests/dotnet-msbuild/msbuild-modernization/node/package.json |
Reviewed |
tests/dotnet-msbuild/msbuild-modernization/multitarget/Widget.cs |
Reviewed |
tests/dotnet-msbuild/msbuild-modernization/multitarget/MultiTarget.Net48.csproj |
Reviewed |
tests/dotnet-msbuild/msbuild-modernization/multitarget/MultiTarget.csproj |
Reviewed |
tests/dotnet-msbuild/msbuild-modernization/cpm-single/Serializer.cs |
Reviewed |
tests/dotnet-msbuild/msbuild-modernization/cpm-single/packages.config |
Reviewed |
tests/dotnet-msbuild/msbuild-modernization/cpm-single/CpmSingle.csproj |
Reviewed |
tests/dotnet-msbuild/msbuild-modernization/already-modern/Clock.cs |
Reviewed |
tests/dotnet-msbuild/msbuild-modernization/already-modern/AlreadyModern.csproj |
Reviewed |
tests/dotnet-msbuild/msbuild-antipatterns/references/fsharp-signature.patch |
Reviewed |
tests/dotnet-msbuild/msbuild-antipatterns/references/fsharp-broken-order.patch |
Reviewed |
tests/dotnet-msbuild/msbuild-antipatterns/references/fsharp-add-module.patch |
Reviewed |
tests/dotnet-msbuild/msbuild-antipatterns/nuget-forwarder/tools/common/MyPkg.Implementation.targets |
Reviewed |
tests/dotnet-msbuild/msbuild-antipatterns/nuget-forwarder/MyPkg.nuspec |
Reviewed |
tests/dotnet-msbuild/msbuild-antipatterns/nuget-forwarder/MyPkg.csproj |
Reviewed |
tests/dotnet-msbuild/msbuild-antipatterns/nuget-forwarder/lib/net8.0/_._ |
Reviewed |
tests/dotnet-msbuild/msbuild-antipatterns/nuget-forwarder/buildTransitive/common/MyPkg.targets |
Reviewed |
tests/dotnet-msbuild/msbuild-antipatterns/nuget-forwarder/build/net8.0/MyPkg.props |
Reviewed |
tests/dotnet-msbuild/msbuild-antipatterns/clean-project/Greeter.cs |
Reviewed |
tests/dotnet-msbuild/msbuild-antipatterns/clean-project/CleanLib.csproj |
Reviewed |
tests/dotnet-msbuild/msbuild-antipatterns/backslash/marker-output/README.txt |
Reviewed |
tests/dotnet-msbuild/msbuild-antipatterns/backslash/Build.targets |
Reviewed |
tests/dotnet-msbuild/item-management/references/diagnose-hard.patch |
Reviewed |
tests/dotnet-msbuild/item-management/glob-mismatch/Service.cs |
Reviewed |
tests/dotnet-msbuild/item-management/glob-mismatch/GlobMismatch.csproj |
Reviewed |
tests/dotnet-msbuild/item-management/glob-mismatch/Generated/Api.g.cs |
Reviewed |
tests/dotnet-msbuild/item-management/false-positive/Widget.cs |
Reviewed |
tests/dotnet-msbuild/item-management/false-positive/FalsePositive.csproj |
Reviewed |
tests/dotnet-msbuild/item-management/already-correct/Legacy/LegacyHelper.cs |
Reviewed |
tests/dotnet-msbuild/item-management/already-correct/Calculator.cs |
Reviewed |
tests/dotnet-msbuild/item-management/already-correct/AlreadyCorrect.csproj |
Reviewed |
tests/dotnet-msbuild/incremental-build/volatile-output/VolatileOutput.csproj |
Reviewed |
tests/dotnet-msbuild/incremental-build/volatile-output/VersionSeed.txt |
Reviewed |
tests/dotnet-msbuild/incremental-build/volatile-output/Program.cs |
Reviewed |
tests/dotnet-msbuild/incremental-build/returns-query/ReturnsQuery.csproj |
Reviewed |
tests/dotnet-msbuild/incremental-build/returns-query/Placeholder.cs |
Reviewed |
tests/dotnet-msbuild/incremental-build/Program.cs |
Reviewed |
tests/dotnet-msbuild/incremental-build/log-interpretation/diag-excerpt.txt |
Reviewed |
tests/dotnet-msbuild/incremental-build/GitHash.txt |
Reviewed |
tests/dotnet-msbuild/incremental-build/compiler-server-cache/diag-excerpt.txt |
Reviewed |
tests/dotnet-msbuild/incremental-build/clean-tracking/Program.cs |
Reviewed |
tests/dotnet-msbuild/incremental-build/clean-tracking/GitHash.txt |
Reviewed |
tests/dotnet-msbuild/incremental-build/clean-tracking/CleanTracking.csproj |
Reviewed |
tests/dotnet-msbuild/incremental-build/BuildStamp.txt |
Reviewed |
tests/dotnet-msbuild/including-generated-files/wrong-timing/WrongTiming.csproj |
Reviewed |
tests/dotnet-msbuild/including-generated-files/wrong-timing/Program.cs |
Reviewed |
tests/dotnet-msbuild/including-generated-files/references/fix-generated-source.patch |
Reviewed |
tests/dotnet-msbuild/including-generated-files/outside-target-glob/Program.cs |
Reviewed |
tests/dotnet-msbuild/including-generated-files/outside-target-glob/OutsideTargetGlob.csproj |
Reviewed |
tests/dotnet-msbuild/including-generated-files/non-code-output/Program.cs |
Reviewed |
tests/dotnet-msbuild/including-generated-files/non-code-output/NonCodeOutput.csproj |
Reviewed |
tests/dotnet-msbuild/including-generated-files/missing-filewrites/Program.cs |
Reviewed |
tests/dotnet-msbuild/including-generated-files/missing-filewrites/MissingFileWrites.csproj |
Reviewed |
tests/dotnet-msbuild/including-generated-files/hardcoded-obj-path/Program.cs |
Reviewed |
tests/dotnet-msbuild/including-generated-files/hardcoded-obj-path/HardcodedObjPath.csproj |
Reviewed |
tests/dotnet-msbuild/including-generated-files/hardcoded-obj-path/Directory.Build.props |
Reviewed |
tests/dotnet-msbuild/extension-points/tfm-forwarder-msb4019/MyAdapter.nuspec |
Reviewed |
tests/dotnet-msbuild/extension-points/tfm-forwarder-msb4019/buildTransitive/net8.0/MyAdapter.props |
Reviewed |
tests/dotnet-msbuild/extension-points/tfm-forwarder-msb4019/build/net8.0/MyAdapter.props |
Reviewed |
tests/dotnet-msbuild/extension-points/references/fix-extension-points.patch |
Reviewed |
tests/dotnet-msbuild/extension-points/references/create-mysdk-extension-point.patch |
Reviewed |
tests/dotnet-msbuild/extension-points/packed-layout-false-positive/MyAdapter.nuspec |
Reviewed |
tests/dotnet-msbuild/extension-points/packed-layout-false-positive/buildTransitive/common/MyAdapter.props |
Reviewed |
tests/dotnet-msbuild/extension-points/packed-layout-false-positive/build/net8.0/MyAdapter.props |
Reviewed |
tests/dotnet-msbuild/extension-points/MyHook.targets |
Reviewed |
tests/dotnet-msbuild/extension-points/filename-id-mismatch/MyAnalyzer.BuildTasks.nuspec |
Reviewed |
tests/dotnet-msbuild/extension-points/filename-id-mismatch/build/MyAnalyzerBuildTasks.targets |
Reviewed |
tests/dotnet-msbuild/extension-points/filename-id-mismatch/build/MyAnalyzerBuildTasks.props |
Reviewed |
tests/dotnet-msbuild/extension-points/ExtensionPoints.csproj |
Reviewed |
tests/dotnet-msbuild/extension-points/Directory.Build.targets |
Reviewed |
tests/dotnet-msbuild/extension-points/author-extension-point/MySDK.targets |
Reviewed |
tests/dotnet-msbuild/extension-points/author-extension-point/MySDK.Before.targets |
Reviewed |
tests/dotnet-msbuild/extension-points/author-extension-point/MySDK.After.targets |
Reviewed |
tests/dotnet-msbuild/extension-points/author-extension-point/Driver.proj |
Reviewed |
tests/dotnet-msbuild/eval-performance/treat-as-local-property/TaLpDemo.csproj |
Reviewed |
tests/dotnet-msbuild/eval-performance/treat-as-local-property/Program.cs |
Reviewed |
tests/dotnet-msbuild/eval-performance/setup/prepare-multi-evaluation.mjs |
Reviewed |
tests/dotnet-msbuild/eval-performance/references/treat-as-local-property-overuse.json |
Reviewed |
tests/dotnet-msbuild/eval-performance/references/property-function-triage.json |
Reviewed |
tests/dotnet-msbuild/eval-performance/references/no-op-already-fast.json |
Reviewed |
tests/dotnet-msbuild/eval-performance/references/multi-pattern-diagnosis.json |
Reviewed |
tests/dotnet-msbuild/eval-performance/references/measurement-first-discipline.json |
Reviewed |
tests/dotnet-msbuild/eval-performance/references/boundary-incremental-build.json |
Reviewed |
tests/dotnet-msbuild/eval-performance/references/boundary-compile-time.json |
Reviewed |
tests/dotnet-msbuild/eval-performance/property-function-triage/PropFuncs.csproj |
Reviewed |
tests/dotnet-msbuild/eval-performance/property-function-triage/Program.cs |
Reviewed |
tests/dotnet-msbuild/eval-performance/property-function-triage/notes.txt |
Reviewed |
tests/dotnet-msbuild/eval-performance/no-op-fast/Program.cs |
Reviewed |
tests/dotnet-msbuild/eval-performance/no-op-fast/FastEval.csproj |
Reviewed |
tests/dotnet-msbuild/eval-performance/multi-evaluation/Shared/Shared.csproj |
Reviewed |
tests/dotnet-msbuild/eval-performance/multi-evaluation/Shared/FlavorInfo.cs |
Reviewed |
tests/dotnet-msbuild/eval-performance/multi-evaluation/Build.proj |
Reviewed |
tests/dotnet-msbuild/eval-performance/measurement-first/Program.cs |
Reviewed |
tests/dotnet-msbuild/eval-performance/measurement-first/BigApp.csproj |
Reviewed |
tests/dotnet-msbuild/eval-performance/boundary-incremental-build/RebuildComplaint.csproj |
Reviewed |
tests/dotnet-msbuild/eval-performance/boundary-incremental-build/Program.cs |
Reviewed |
tests/dotnet-msbuild/eval-performance/boundary-compile-time/SlowCompile.csproj |
Reviewed |
tests/dotnet-msbuild/eval-performance/boundary-compile-time/Program.cs |
Reviewed |
tests/dotnet-msbuild/eval-performance/boundary-compile-time/perf-summary.txt |
Reviewed |
tests/dotnet-msbuild/directory-build-organization/tfm-timing/TfmTiming.csproj |
Reviewed |
tests/dotnet-msbuild/directory-build-organization/tfm-timing/Placeholder.cs |
Reviewed |
tests/dotnet-msbuild/directory-build-organization/tfm-timing/Directory.Build.props |
Reviewed |
tests/dotnet-msbuild/directory-build-organization/single-project-console/SingleProjectConsole.csproj |
Reviewed |
tests/dotnet-msbuild/directory-build-organization/single-project-console/Program.cs |
Reviewed |
tests/dotnet-msbuild/directory-build-organization/intentional-divergence/src/Worker/Worker.csproj |
Reviewed |
tests/dotnet-msbuild/directory-build-organization/intentional-divergence/src/Worker/Worker.cs |
Reviewed |
tests/dotnet-msbuild/directory-build-organization/intentional-divergence/src/GeneratedClient/GeneratedClient.csproj |
Reviewed |
tests/dotnet-msbuild/directory-build-organization/intentional-divergence/src/GeneratedClient/Client.cs |
Reviewed |
tests/dotnet-msbuild/directory-build-organization/hierarchy-clobber/src/LegacyClient/Program.cs |
Reviewed |
tests/dotnet-msbuild/directory-build-organization/hierarchy-clobber/src/LegacyClient/LegacyClient.csproj |
Reviewed |
tests/dotnet-msbuild/directory-build-organization/hierarchy-clobber/src/Directory.Build.props |
Reviewed |
tests/dotnet-msbuild/directory-build-organization/hierarchy-clobber/Directory.Build.props |
Reviewed |
tests/dotnet-msbuild/check-bin-obj-clash/ToolReferenceClash.slnx |
Reviewed |
tests/dotnet-msbuild/check-bin-obj-clash/SharedPathClash.slnx |
Reviewed |
tests/dotnet-msbuild/check-bin-obj-clash/safe-shared-base/SafeSharedBase.slnx |
Reviewed |
tests/dotnet-msbuild/check-bin-obj-clash/safe-shared-base/LibD/LibD.csproj |
Reviewed |
tests/dotnet-msbuild/check-bin-obj-clash/safe-shared-base/LibD/ClassD.cs |
Reviewed |
tests/dotnet-msbuild/check-bin-obj-clash/safe-shared-base/LibC/LibC.csproj |
Reviewed |
tests/dotnet-msbuild/check-bin-obj-clash/safe-shared-base/LibC/ClassC.cs |
Reviewed |
tests/dotnet-msbuild/check-bin-obj-clash/global.json |
Reviewed |
tests/dotnet-msbuild/check-bin-obj-clash/default-layout-safe/DefaultLayoutSafe.slnx |
Reviewed |
tests/dotnet-msbuild/check-bin-obj-clash/default-layout-safe/AppBeta/Program.cs |
Reviewed |
tests/dotnet-msbuild/check-bin-obj-clash/default-layout-safe/AppBeta/AppBeta.csproj |
Reviewed |
tests/dotnet-msbuild/check-bin-obj-clash/default-layout-safe/AppAlpha/Program.cs |
Reviewed |
tests/dotnet-msbuild/check-bin-obj-clash/default-layout-safe/AppAlpha/AppAlpha.csproj |
Reviewed |
tests/dotnet-msbuild/build-perf-diagnostics/single-target-domination/perf-summary.txt |
Reviewed |
tests/dotnet-msbuild/build-perf-diagnostics/runtime-perf-request/Storefront.csproj |
Reviewed |
tests/dotnet-msbuild/build-perf-diagnostics/runtime-perf-request/CheckoutController.cs |
Reviewed |
tests/dotnet-msbuild/build-perf-diagnostics/restore-in-build/perf-summary.txt |
Reviewed |
tests/dotnet-msbuild/build-perf-diagnostics/rar-bottleneck/TaskContracts/TaskContracts.csproj |
Reviewed |
tests/dotnet-msbuild/build-perf-diagnostics/rar-bottleneck/TaskContracts/ManifestSpec.cs |
Reviewed |
tests/dotnet-msbuild/build-perf-diagnostics/rar-bottleneck/perf-summary.txt |
Reviewed |
tests/dotnet-msbuild/build-perf-diagnostics/rar-bottleneck/GatewayService/GatewayService.csproj |
Reviewed |
tests/dotnet-msbuild/build-perf-diagnostics/rar-bottleneck/BuildTasks/ManifestGeneratorTask.cs |
Reviewed |
tests/dotnet-msbuild/build-perf-diagnostics/rar-bottleneck/BuildTasks/BuildTasks.csproj |
Reviewed |
tests/dotnet-msbuild/build-perf-diagnostics/offline-analyzers/global.json |
Reviewed |
tests/dotnet-msbuild/build-perf-diagnostics/offline-analyzers/AnalyzerTwo/AnalyzerTwo.csproj |
Reviewed |
tests/dotnet-msbuild/build-perf-diagnostics/offline-analyzers/AnalyzerTwo/AnalyzerTwo.cs |
Reviewed |
tests/dotnet-msbuild/build-perf-diagnostics/offline-analyzers/AnalyzerThree/AnalyzerThree.csproj |
Reviewed |
tests/dotnet-msbuild/build-perf-diagnostics/offline-analyzers/AnalyzerThree/AnalyzerThree.cs |
Reviewed |
tests/dotnet-msbuild/build-perf-diagnostics/offline-analyzers/AnalyzerOne/AnalyzerOne.csproj |
Reviewed |
tests/dotnet-msbuild/build-perf-diagnostics/offline-analyzers/AnalyzerOne/AnalyzerOne.cs |
Reviewed |
tests/dotnet-msbuild/build-perf-diagnostics/many-small-projects/perf-summary.txt |
Reviewed |
tests/dotnet-msbuild/build-perf-diagnostics/evaluation-overhead/perf-summary.txt |
Reviewed |
tests/dotnet-msbuild/build-perf-diagnostics/Directory.Build.props |
Reviewed |
tests/dotnet-msbuild/build-perf-diagnostics/DataService.cs |
Reviewed |
tests/dotnet-msbuild/build-perf-diagnostics/copy-io-bottleneck/perf-summary.txt |
Reviewed |
tests/dotnet-msbuild/build-perf-diagnostics/Contoso.WebApi.csproj |
Reviewed |
tests/dotnet-msbuild/build-perf-baseline/references/deterministic-ci-cache.patch |
Reviewed |
tests/dotnet-msbuild/build-perf-baseline/noop-regression/build-log.txt |
Reviewed |
tests/dotnet-msbuild/build-perf-baseline/node-webpack/webpack.config.js |
Reviewed |
tests/dotnet-msbuild/build-perf-baseline/node-webpack/package.json |
Reviewed |
tests/dotnet-msbuild/build-perf-baseline/cold-restore-bound/build-log.txt |
Reviewed |
tests/dotnet-msbuild/build-perf-baseline/already-optimized/Directory.Build.props |
Reviewed |
tests/dotnet-msbuild/build-perf-baseline/already-optimized/Core/Core.csproj |
Reviewed |
tests/dotnet-msbuild/build-perf-baseline/already-optimized/Core/Calculator.cs |
Reviewed |
tests/dotnet-msbuild/build-perf-baseline/already-optimized/App/Program.cs |
Reviewed |
tests/dotnet-msbuild/build-perf-baseline/already-optimized/App/App.csproj |
Reviewed |
tests/dotnet-msbuild/build-parallelism/wide-slow-leg/perf-summary.txt |
Reviewed |
tests/dotnet-msbuild/build-parallelism/references/enable-build-in-parallel.patch |
Reviewed |
tests/dotnet-msbuild/build-parallelism/not-parallel/Core/Core.csproj |
Reviewed |
tests/dotnet-msbuild/build-parallelism/not-parallel/App4/App4.csproj |
Reviewed |
tests/dotnet-msbuild/build-parallelism/not-parallel/App3/App3.csproj |
Reviewed |
tests/dotnet-msbuild/build-parallelism/not-parallel/App2/App2.csproj |
Reviewed |
tests/dotnet-msbuild/build-parallelism/not-parallel/App1/App1.csproj |
Reviewed |
tests/dotnet-msbuild/build-parallelism/dynamic-graph/Build.proj |
Reviewed |
tests/dotnet-msbuild/build-parallelism/build-parallel-task/LibC/LibC.csproj |
Reviewed |
tests/dotnet-msbuild/build-parallelism/build-parallel-task/LibB/LibB.csproj |
Reviewed |
tests/dotnet-msbuild/build-parallelism/build-parallel-task/LibA/LibA.csproj |
Reviewed |
tests/dotnet-msbuild/build-parallelism/build-parallel-task/Core/Core.csproj |
Reviewed |
tests/dotnet-msbuild/build-parallelism/build-parallel-task/build.proj |
Reviewed |
tests/dotnet-msbuild/binlog-generation/references/preserve-binlog-history.json |
Reviewed |
tests/dotnet-msbuild/binlog-generation/references/powershell-escaping.json |
Reviewed |
tests/dotnet-msbuild/binlog-generation/references/no-binlog-on-preflight-failure.json |
Reviewed |
tests/dotnet-msbuild/binlog-generation/references/multi-config-unique-names.json |
Reviewed |
tests/dotnet-msbuild/binlog-generation/references/fix-build-script.patch |
Reviewed |
tests/dotnet-msbuild/binlog-generation/references/decline-non-msbuild.json |
Reviewed |
tests/dotnet-msbuild/binlog-generation/references/bl-placeholder-flag.json |
Reviewed |
tests/dotnet-msbuild/binlog-generation/predictable-filename/SimpleApp.csproj |
Reviewed |
tests/dotnet-msbuild/binlog-generation/predictable-filename/Program.cs |
Reviewed |
tests/dotnet-msbuild/binlog-generation/predictable-filename/ci-artifacts-manifest.txt |
Reviewed |
tests/dotnet-msbuild/binlog-generation/maven-project/src/main/java/com/contoso/reporting/ReportGenerator.java |
Reviewed |
tests/dotnet-msbuild/binlog-generation/maven-project/pom.xml |
Reviewed |
tests/dotnet-msbuild/binlog-generation/fix-build-script/SimpleApp.csproj |
Reviewed |
tests/dotnet-msbuild/binlog-generation/fix-build-script/Program.cs |
Reviewed |
tests/dotnet-msbuild/binlog-generation/fix-build-script/build.sh |
Reviewed |
tests/dotnet-msbuild/binlog-generation/failing-invocation/SimpleApp.csproj |
Reviewed |
tests/dotnet-msbuild/binlog-generation/failing-invocation/Program.cs |
Reviewed |
tests/dotnet-msbuild/binlog-generation/cleanup-preserve-binlog/SimpleApp.csproj |
Reviewed |
tests/dotnet-msbuild/binlog-generation/cleanup-preserve-binlog/Program.cs |
Reviewed |
tests/dotnet-msbuild/binlog-failure-analysis/warning-only-build/Program.cs |
Reviewed |
tests/dotnet-msbuild/binlog-failure-analysis/warning-only-build/PriceLib.csproj |
Reviewed |
tests/dotnet-msbuild/binlog-failure-analysis/target-order-tracing/VersionedApp.csproj |
Reviewed |
tests/dotnet-msbuild/binlog-failure-analysis/target-order-tracing/Program.cs |
Reviewed |
tests/dotnet-msbuild/binlog-failure-analysis/skip-vs-fail/ReportTool.csproj |
Reviewed |
tests/dotnet-msbuild/binlog-failure-analysis/skip-vs-fail/Program.cs |
Reviewed |
tests/dotnet-msbuild/binlog-failure-analysis/references/boundary-non-msbuild.json |
Reviewed |
tests/dotnet-msbuild/binlog-failure-analysis/references/boundary-generation-routing.json |
Reviewed |
tests/dotnet-msbuild/binlog-failure-analysis/property-query/Program.cs |
Reviewed |
tests/dotnet-msbuild/binlog-failure-analysis/property-query/package-source/Contoso.Catalog.csproj |
Reviewed |
tests/dotnet-msbuild/binlog-failure-analysis/property-query/package-source/CatalogItem.cs |
Reviewed |
tests/dotnet-msbuild/binlog-failure-analysis/property-query/NuGet.Config |
Reviewed |
tests/dotnet-msbuild/binlog-failure-analysis/property-query/local-feed/README.txt |
Reviewed |
tests/dotnet-msbuild/binlog-failure-analysis/property-query/CatalogTool.csproj |
Reviewed |
tests/dotnet-msbuild/binlog-failure-analysis/gradle-build/gradle-build-output.txt |
Reviewed |
tests/dotnet-msbuild/binlog-failure-analysis/generation-routing/Program.cs |
Reviewed |
tests/dotnet-msbuild/binlog-failure-analysis/generation-routing/InventorySync.csproj |
Reviewed |
tests/dotnet-msbuild/binlog-failure-analysis/fallback-replay/Program.cs |
Reviewed |
tests/dotnet-msbuild/binlog-failure-analysis/fallback-replay/OrderCalc.csproj |
Reviewed |
tests/dotnet-msbuild/agent.msbuild/NuGet.Config |
Reviewed |
tests/dotnet-msbuild/agent.msbuild/healthy/AppTwo/Program.cs |
Reviewed |
tests/dotnet-msbuild/agent.msbuild/healthy/AppTwo/AppTwo.csproj |
Reviewed |
tests/dotnet-msbuild/agent.msbuild/healthy/AppOne/Program.cs |
Reviewed |
tests/dotnet-msbuild/agent.msbuild/healthy/AppOne/AppOne.csproj |
Reviewed |
tests/dotnet-msbuild/agent.msbuild/empty-feed/README.txt |
Reviewed |
eng/vally-adapter/README.md |
Reviewed |
eng/skill-validator/tests/Evaluate/MetricsTests.cs |
Reviewed |
eng/skill-validator/tests/Evaluate/ComparatorTests.cs |
Reviewed |
eng/skill-validator/src/README.md |
Reviewed |
eng/skill-validator/src/Evaluate/OverfittingCommand.cs |
Reviewed |
eng/skill-validator/src/Evaluate/Models.cs |
Reviewed |
eng/skill-validator/src/Evaluate/MetricsCollector.cs |
Reviewed |
eng/skill-validator/src/Evaluate/LlmSession.cs |
Reviewed |
eng/skill-validator/src/Evaluate/EvalSchema.cs |
Reviewed |
eng/eval-quality/underpowered-allowlist.txt |
Reviewed |
.github/workflows/eval-quality.yml |
Reviewed |
💡 Add a code-review agent skill or configure MCP servers for context-aware, tailored reviews. Learn more in the docs.
There was a problem hiding this comment.
Copilot review overview
🟡 Changes recommended
Unresolved moderate and critical findings affect shard selection, isolation, cleanup, rejudge correctness, and activation-contract enforcement.
Get a fresh assessment by requesting another Copilot review.
Review effort: Lite
Findings: 6
Open (7)
NuGet fallback folders bypass package isolation · New Retain generated file Compile include alongside FileWrites Treat unmatched dormancy names as activation contract failures Require exact baseline run-index pairing Fix shard matcher ordering to recognize indented executionShard Run all eval specs when checker or workflow logic changes Reject duplicate sessions for inline rejudge roles
|
👋 @AbhitejJohn — this PR has 7 unresolved review thread(s). When you're ready, please address the feedback and push an update; the triage bot will pick up the next state automatically. (Add the |
Audit checker changes across all eval specs, enforce exact rejudge pairing, isolate NuGet fallbacks, and repair shard, dormancy, and target-authoring evidence. Co-authored-by: Copilot App <223556219+Copilot@users.noreply.github.com>
Keep unmatched dormancy annotations as explicit failures without adding them twice to activation contract totals. Co-authored-by: Copilot App <223556219+Copilot@users.noreply.github.com>
AbhitejJohn
left a comment
There was a problem hiding this comment.
/evaluate ba899ab
There was a problem hiding this comment.
Copilot review overview
🟡 Changes recommended
Critical setup, execution-error propagation, and fail-closed pairing issues remain unresolved.
Get a fresh assessment by requesting another Copilot review.
Review effort: Lite
Findings: 6
Open (6)
Fail evaluation when setup commands or launches fail · New Propagate terminal session errors before scoring · New Validate consistent activation expectations across runs · New Reject unexpected treatment roles before judging · New Reject treatment records with mismatched baseline keys · New Fail scenarios when analyzer setup fails or times out · New
Resolved since last review (7)
NuGet fallback folders bypass package isolation Retain generated file Compile include alongside FileWrites Treat unmatched dormancy names as activation contract failures Require exact baseline run-index pairing Fix shard matcher ordering to recognize indented executionShard Run all eval specs when checker or workflow logic changes Reject duplicate sessions for inline rejudge roles
📊 Skill and Agent Evaluation Results36 model/target results across 18 targets and 2 models — ✅ 2 improved, ➖ 30 not proven improved, Measurement identity: evaluated commit Measurement health: 36 expected / 36 observed / 36 written; 0 missing, 0 unexpected, 0 invalid; 0 recovered comparison error slots and 0 unresolved comparison error slots. Objective completion gate: not enabled. Aggregate completion transitions are telemetry only, so this report does not claim that zero objective regressions were proven. A result passes only when preference-eligible distinct-stimulus votes have aggregate net win of at least 20% and an exact one-sided sign-test result of
ℹ️ How to read this report
⛔ Activation contract failed — build-perf-baseline (gpt-5.6-luna)Why: Net win +0.0% (3W/0T/3L over 6 preference-eligible stimulus vote(s), sign test p=0.656), mean preference -8.6% across 7 paired run(s), 1 dormancy stimulus/stimuli excluded from preference — activation contract failed (1 explicit dormancy scenario(s) activated the isolated target skill) Next action: Narrow skill routing so the listed off-target scenarios stay dormant. State: Gate evidence: n=6; 3W/0T/3L; d=6; p=0.656; net +0.0%; 1 dormancy excluded Warnings: Dormancy contract: 1 unexpected activation(s); Activation: isolated 4/6; plugin 6/6 Overfit: High (score 0.56) Repeated-run reliability (not used by the gate): 7 paired runs (3W/1T/3L). Weak or warning scenarios:
Illustrative judge evidence:
This is one example, not the aggregate verdict. Open Full Results for every judgment. ⛔ Activation contract failed — directory-build-organization (claude-sonnet-5)Why: Net win +66.7% (4W/2T/0L over 6 preference-eligible stimulus vote(s), sign test p=0.063), mean preference +34.3% across 7 paired run(s), 1 dormancy stimulus/stimuli excluded from preference — activation contract failed (1 explicit dormancy scenario(s) activated the isolated target skill) Next action: Narrow skill routing so the listed off-target scenarios stay dormant. State: Gate evidence: n=6; 4W/2T/0L; d=4; p=0.063; net +66.7%; 1 dormancy excluded Warnings: Dormancy contract: 1 unexpected activation(s); Activation: isolated 5/6; plugin 5/6 Overfit: Moderate (score 0.43) Repeated-run reliability (not used by the gate): 7 paired runs (4W/2T/1L). Weak or warning scenarios:
Illustrative judge evidence:
This is one example, not the aggregate verdict. Open Full Results for every judgment. ⛔ Activation contract failed — msbuild-antipatterns (gpt-5.6-luna)Why: Net win -14.3% (1W/4T/2L over 7 preference-eligible stimulus vote(s), sign test p=0.500), mean preference -12.5% across 8 paired run(s), 1 dormancy stimulus/stimuli excluded from preference — activation contract failed (1 explicit dormancy scenario(s) activated the isolated target skill) Next action: Narrow skill routing so the listed off-target scenarios stay dormant. State: Gate evidence: n=7; 1W/4T/2L; d=3; p=0.500; net -14.3%; 1 dormancy excluded Warnings: Dormancy contract: 1 unexpected activation(s); Activation: isolated 5/7; plugin 4/7 Overfit: Moderate (score 0.31) Repeated-run reliability (not used by the gate): 8 paired runs (1W/5T/2L). Weak or warning scenarios:
Illustrative judge evidence:
This is one example, not the aggregate verdict. Open Full Results for every judgment. 📉 Preference loss (report only) — binlog-failure-analysis (claude-sonnet-5)Why: Net win -83.3% (0W/1T/5L over 6 preference-eligible stimulus vote(s), sign test p=0.031), mean preference -42.9% across 7 paired run(s), 1 dormancy stimulus/stimuli excluded from preference — credibly worse Next action: Inspect losing stimuli and fix skill behavior; this is not objective completion proof. State: Gate evidence: n=6; 0W/1T/5L; d=5; p=0.031; net -83.3%; 1 dormancy excluded Warnings: Activation: isolated 5/7; plugin 6/7 Overfit: Moderate (score 0.21) Repeated-run reliability (not used by the gate): 7 paired runs (0W/1T/6L). Weak or warning scenarios:
Illustrative judge evidence:
This is one example, not the aggregate verdict. Open Full Results for every judgment. ➖ Not proven improved — agent.msbuild (claude-sonnet-5)Why: Net win -40.0% (1W/1T/3L over 5 preference-eligible stimulus vote(s), sign test p=0.312), mean preference -16.0% across 5 paired run(s) — no improvement — native evaluator reported that the target agent did not activate Next action: Inspect tied or lost stimuli; predeclare added breadth before a new experiment. State: Gate evidence: n=5; 1W/1T/3L; d=4; p=0.312; net -40.0% Warnings: Activation: isolated 1/5; plugin 0/5 Repeated-run reliability (not used by the gate): 5 paired runs (1W/1T/3L). Weak or warning scenarios:
Illustrative judge evidence:
This is one example, not the aggregate verdict. Open Full Results for every judgment. ➖ Not proven improved — agent.msbuild (gpt-5.6-luna)Why: Net win -20.0% (1W/2T/2L over 5 preference-eligible stimulus vote(s), sign test p=0.500), mean preference -8.0% across 5 paired run(s) — no improvement — native evaluator reported that the target agent did not activate Next action: Inspect tied or lost stimuli; predeclare added breadth before a new experiment. State: Gate evidence: n=5; 1W/2T/2L; d=3; p=0.500; net -20.0% Warnings: Activation: isolated 1/5; plugin 1/5 Repeated-run reliability (not used by the gate): 5 paired runs (1W/2T/2L). Weak or warning scenarios:
Illustrative judge evidence:
This is one example, not the aggregate verdict. Open Full Results for every judgment. ➖ Not proven improved — binlog-generation (claude-sonnet-5)Why: Net win +28.6% (4W/1T/2L over 7 preference-eligible stimulus vote(s), sign test p=0.344), mean preference +32.5% across 8 paired run(s), 1 dormancy stimulus/stimuli excluded from preference — not credible (sign test p=0.344 > 0.05) Next action: Inspect activation-only failed runs before rewriting skill content; the model stopped after loading a skill. State: Gate evidence: n=7; 4W/1T/2L; d=6; p=0.344; net +28.6%; 1 dormancy excluded Warnings: Activation: isolated 5/7; plugin 6/7; Activation-only stop: plugin 1 failed run Overfit: High (score 0.50) Repeated-run reliability (not used by the gate): 8 paired runs (4W/2T/2L). Weak or warning scenarios:
Illustrative judge evidence:
This is one example, not the aggregate verdict. Open Full Results for every judgment. ➖ Not proven improved — binlog-generation (gpt-5.6-luna)Why: Net win +42.9% (4W/2T/1L over 7 preference-eligible stimulus vote(s), sign test p=0.188), mean preference +37.5% across 8 paired run(s), 1 dormancy stimulus/stimuli excluded from preference — not credible (sign test p=0.188 > 0.05) Next action: Inspect tied or lost stimuli and fix inconsistent skill behavior. State: Gate evidence: n=7; 4W/2T/1L; d=5; p=0.188; net +42.9%; 1 dormancy excluded Warnings: Activation: isolated 6/7; plugin 5/7 Overfit: Moderate (score 0.35) Repeated-run reliability (not used by the gate): 8 paired runs (4W/3T/1L). Weak or warning scenarios:
Illustrative judge evidence:
This is one example, not the aggregate verdict. Open Full Results for every judgment. ➖ Not proven improved — build-parallelism (claude-sonnet-5)Why: Net win -66.7% (0W/2T/4L over 6 preference-eligible stimulus vote(s), sign test p=0.063), mean preference -22.9% across 7 paired run(s), 1 dormancy stimulus/stimuli excluded from preference — no improvement Next action: Inspect activation-only failed runs before rewriting skill content; the model stopped after loading a skill. State: Gate evidence: n=6; 0W/2T/4L; d=4; p=0.063; net -66.7%; 1 dormancy excluded Warnings: Activation: isolated 5/6; plugin 5/6; Activation-only stop: isolated 1 failed run; Activation-only stop: plugin 1 failed run Overfit: Low (score 0.17) Repeated-run reliability (not used by the gate): 7 paired runs (0W/3T/4L). Weak or warning scenarios:
Illustrative judge evidence:
This is one example, not the aggregate verdict. Open Full Results for every judgment. ➖ Not proven improved — build-parallelism (gpt-5.6-luna)Why: Net win +16.7% (3W/1T/2L over 6 preference-eligible stimulus vote(s), sign test p=0.500), mean preference +5.7% across 7 paired run(s), 1 dormancy stimulus/stimuli excluded from preference — not credible (sign test p=0.500 > 0.05) Next action: Inspect activation-only failed runs before rewriting skill content; the model stopped after loading a skill. State: Gate evidence: n=6; 3W/1T/2L; d=5; p=0.500; net +16.7%; 1 dormancy excluded Warnings: Activation-only stop: isolated 1 failed run; Activation-only stop: plugin 1 failed run Overfit: Moderate (score 0.32) Repeated-run reliability (not used by the gate): 7 paired runs (3W/2T/2L). Weak or warning scenarios:
Illustrative judge evidence:
This is one example, not the aggregate verdict. Open Full Results for every judgment. ➖ Not proven improved — build-perf-baseline (claude-sonnet-5)Why: Net win +66.7% (4W/2T/0L over 6 preference-eligible stimulus vote(s), sign test p=0.063), mean preference +37.1% across 7 paired run(s), 1 dormancy stimulus/stimuli excluded from preference — not credible — 2 of 6 preference-eligible stimulus vote(s) tied, leaving only 4 discordant preference vote(s). The sign test conditions on non-tie stimulus votes and cannot reach 0.05 below 5, so no record could have passed here — this is not a measured null. Either the skill is inert on these scenarios (make them discriminate) or the eval needs more distinct stimuli to clear the ties Next action: Inspect tied or lost stimuli; predeclare added breadth before a new experiment. State: Gate evidence: n=6; 4W/2T/0L; d=4; p=0.063; net +66.7%; 1 dormancy excluded Warnings: Activation: isolated 5/6; plugin 4/6 Overfit: Moderate (score 0.32) Repeated-run reliability (not used by the gate): 7 paired runs (5W/2T/0L). Weak or warning scenarios:
Illustrative judge evidence:
This is one example, not the aggregate verdict. Open Full Results for every judgment. ➖ Not proven improved — build-perf-diagnostics (claude-sonnet-5)Why: Net win +14.3% (4W/0T/3L over 7 preference-eligible stimulus vote(s), sign test p=0.500), mean preference +5.0% across 8 paired run(s), 1 dormancy stimulus/stimuli excluded from preference — not credible (sign test p=0.500 > 0.05) Next action: Inspect tied or lost stimuli and fix inconsistent skill behavior. State: Gate evidence: n=7; 4W/0T/3L; d=7; p=0.500; net +14.3%; 1 dormancy excluded Overfit: High (score 0.50) Repeated-run reliability (not used by the gate): 8 paired runs (4W/1T/3L). Weak or warning scenarios:
Illustrative judge evidence:
This is one example, not the aggregate verdict. Open Full Results for every judgment. ➖ Not proven improved — build-perf-diagnostics (gpt-5.6-luna)Why: Net win +28.6% (3W/3T/1L over 7 preference-eligible stimulus vote(s), sign test p=0.312), mean preference +15.0% across 8 paired run(s), 1 dormancy stimulus/stimuli excluded from preference — not credible — 3 of 7 preference-eligible stimulus vote(s) tied, leaving only 4 discordant preference vote(s). The sign test conditions on non-tie stimulus votes and cannot reach 0.05 below 5, so no record could have passed here — this is not a measured null. Either the skill is inert on these scenarios (make them discriminate) or the eval needs more distinct stimuli to clear the ties Next action: Inspect tied or lost stimuli; predeclare added breadth before a new experiment. State: Gate evidence: n=7; 3W/3T/1L; d=4; p=0.312; net +28.6%; 1 dormancy excluded Overfit: Moderate (score 0.40) Repeated-run reliability (not used by the gate): 8 paired runs (4W/3T/1L). Weak or warning scenarios:
Illustrative judge evidence:
This is one example, not the aggregate verdict. Open Full Results for every judgment. ➖ Not proven improved — check-bin-obj-clash (claude-sonnet-5)Why: Net win +28.6% (2W/5T/0L over 7 preference-eligible stimulus vote(s), sign test p=0.250), mean preference +20.0% across 7 paired run(s) — not credible — 5 of 7 preference-eligible stimulus vote(s) tied, leaving only 2 discordant preference vote(s). The sign test conditions on non-tie stimulus votes and cannot reach 0.05 below 5, so no record could have passed here — this is not a measured null. Either the skill is inert on these scenarios (make them discriminate) or the eval needs more distinct stimuli to clear the ties Next action: Inspect tied or lost stimuli; predeclare added breadth before a new experiment. State: Gate evidence: n=7; 2W/5T/0L; d=2; p=0.250; net +28.6% Warnings: Activation: isolated 6/7; plugin 6/7 Overfit: Moderate (score 0.24) Repeated-run reliability (not used by the gate): 7 paired runs (2W/5T/0L). Weak or warning scenarios:
Illustrative judge evidence:
This is one example, not the aggregate verdict. Open Full Results for every judgment. ➖ Not proven improved — check-bin-obj-clash (gpt-5.6-luna)Why: Net win +42.9% (4W/2T/1L over 7 preference-eligible stimulus vote(s), sign test p=0.188), mean preference +17.1% across 7 paired run(s) — not credible (sign test p=0.188 > 0.05) Next action: Inspect tied or lost stimuli and fix inconsistent skill behavior. State: Gate evidence: n=7; 4W/2T/1L; d=5; p=0.188; net +42.9% Warnings: Activation: isolated 7/7; plugin 6/7 Overfit: Moderate (score 0.21) Repeated-run reliability (not used by the gate): 7 paired runs (4W/2T/1L). Weak or warning scenarios:
Illustrative judge evidence:
This is one example, not the aggregate verdict. Open Full Results for every judgment. ➖ Not proven improved — directory-build-organization (gpt-5.6-luna)Why: Net win +66.7% (4W/2T/0L over 6 preference-eligible stimulus vote(s), sign test p=0.063), mean preference +17.1% across 7 paired run(s), 1 dormancy stimulus/stimuli excluded from preference — not credible — 2 of 6 preference-eligible stimulus vote(s) tied, leaving only 4 discordant preference vote(s). The sign test conditions on non-tie stimulus votes and cannot reach 0.05 below 5, so no record could have passed here — this is not a measured null. Either the skill is inert on these scenarios (make them discriminate) or the eval needs more distinct stimuli to clear the ties Next action: Inspect tied or lost stimuli; predeclare added breadth before a new experiment. State: Gate evidence: n=6; 4W/2T/0L; d=4; p=0.063; net +66.7%; 1 dormancy excluded Overfit: High (score 0.53) Repeated-run reliability (not used by the gate): 7 paired runs (4W/2T/1L). Weak or warning scenarios:
Illustrative judge evidence:
This is one example, not the aggregate verdict. Open Full Results for every judgment. ➖ Not proven improved — eval-performance (claude-sonnet-5)Why: Net win +25.0% (4W/2T/2L over 8 preference-eligible stimulus vote(s), sign test p=0.344), mean preference +10.0% across 8 paired run(s) — not credible (sign test p=0.344 > 0.05) Next action: Inspect tied or lost stimuli and fix inconsistent skill behavior. State: Gate evidence: n=8; 4W/2T/2L; d=6; p=0.344; net +25.0% Warnings: Activation: isolated 6/8; plugin 7/8 Overfit: Moderate (score 0.35) Repeated-run reliability (not used by the gate): 8 paired runs (4W/2T/2L). Weak or warning scenarios:
Illustrative judge evidence:
This is one example, not the aggregate verdict. Open Full Results for every judgment. ➖ Not proven improved — eval-performance (gpt-5.6-luna)Why: Net win +25.0% (3W/4T/1L over 8 preference-eligible stimulus vote(s), sign test p=0.312), mean preference +25.0% across 8 paired run(s) — not credible — 4 of 8 preference-eligible stimulus vote(s) tied, leaving only 4 discordant preference vote(s). The sign test conditions on non-tie stimulus votes and cannot reach 0.05 below 5, so no record could have passed here — this is not a measured null. Either the skill is inert on these scenarios (make them discriminate) or the eval needs more distinct stimuli to clear the ties Next action: Inspect tied or lost stimuli; predeclare added breadth before a new experiment. State: Gate evidence: n=8; 3W/4T/1L; d=4; p=0.312; net +25.0% Overfit: Moderate (score 0.49) Repeated-run reliability (not used by the gate): 8 paired runs (3W/4T/1L). Weak or warning scenarios:
Illustrative judge evidence:
This is one example, not the aggregate verdict. Open Full Results for every judgment. ➖ Not proven improved — extension-points (claude-sonnet-5)Why: Net win +14.3% (3W/2T/2L over 7 preference-eligible stimulus vote(s), sign test p=0.500), mean preference +17.8% across 9 paired run(s), 2 dormancy stimulus/stimuli excluded from preference — not credible (sign test p=0.500 > 0.05) Next action: Inspect tied or lost stimuli and fix inconsistent skill behavior. State: Gate evidence: n=7; 3W/2T/2L; d=5; p=0.500; net +14.3%; 2 dormancy excluded Warnings: Activation: isolated 6/7; plugin 5/7 Overfit: Moderate (score 0.32) Repeated-run reliability (not used by the gate): 9 paired runs (3W/4T/2L). Weak or warning scenarios:
Illustrative judge evidence:
This is one example, not the aggregate verdict. Open Full Results for every judgment. ➖ Not proven improved — extension-points (gpt-5.6-luna)Why: Net win +14.3% (3W/2T/2L over 7 preference-eligible stimulus vote(s), sign test p=0.500), mean preference -2.2% across 9 paired run(s), 2 dormancy stimulus/stimuli excluded from preference — not credible (sign test p=0.500 > 0.05) Next action: Inspect tied or lost stimuli and fix inconsistent skill behavior. State: Gate evidence: n=7; 3W/2T/2L; d=5; p=0.500; net +14.3%; 2 dormancy excluded Warnings: Activation: isolated 7/7; plugin 6/7 Overfit: Moderate (score 0.33) Repeated-run reliability (not used by the gate): 9 paired runs (3W/4T/2L). Weak or warning scenarios:
Illustrative judge evidence:
This is one example, not the aggregate verdict. Open Full Results for every judgment. ➖ Not proven improved — including-generated-files (claude-sonnet-5)Why: Net win -14.3% (2W/2T/3L over 7 preference-eligible stimulus vote(s), sign test p=0.500), mean preference -12.5% across 8 paired run(s), 1 dormancy stimulus/stimuli excluded from preference — no improvement Next action: Inspect tied or lost stimuli and fix inconsistent skill behavior. State: Gate evidence: n=7; 2W/2T/3L; d=5; p=0.500; net -14.3%; 1 dormancy excluded Warnings: Activation: isolated 6/7; plugin 5/7 Overfit: Moderate (score 0.42) Repeated-run reliability (not used by the gate): 8 paired runs (2W/3T/3L). Weak or warning scenarios:
Illustrative judge evidence:
This is one example, not the aggregate verdict. Open Full Results for every judgment. ➖ Not proven improved — including-generated-files (gpt-5.6-luna)Why: Net win +14.3% (2W/4T/1L over 7 preference-eligible stimulus vote(s), sign test p=0.500), mean preference +10.0% across 8 paired run(s), 1 dormancy stimulus/stimuli excluded from preference — not credible — 4 of 7 preference-eligible stimulus vote(s) tied, leaving only 3 discordant preference vote(s). The sign test conditions on non-tie stimulus votes and cannot reach 0.05 below 5, so no record could have passed here — this is not a measured null. Either the skill is inert on these scenarios (make them discriminate) or the eval needs more distinct stimuli to clear the ties Next action: Inspect tied or lost stimuli; predeclare added breadth before a new experiment. State: Gate evidence: n=7; 2W/4T/1L; d=3; p=0.500; net +14.3%; 1 dormancy excluded Warnings: Activation: isolated 6/7; plugin 7/7 Overfit: Moderate (score 0.39) Repeated-run reliability (not used by the gate): 8 paired runs (3W/4T/1L). Weak or warning scenarios:
Illustrative judge evidence:
This is one example, not the aggregate verdict. Open Full Results for every judgment. ➖ Not proven improved — incremental-build (gpt-5.6-luna)Why: Net win +11.1% (3W/4T/2L over 9 preference-eligible stimulus vote(s), sign test p=0.500), mean preference +4.4% across 9 paired run(s) — not credible (sign test p=0.500 > 0.05) Next action: Inspect tied or lost stimuli and fix inconsistent skill behavior. State: Gate evidence: n=9; 3W/4T/2L; d=5; p=0.500; net +11.1% Warnings: Activation: isolated 7/9; plugin 7/9 Overfit: Moderate (score 0.34) Repeated-run reliability (not used by the gate): 9 paired runs (3W/4T/2L). Weak or warning scenarios:
Illustrative judge evidence:
This is one example, not the aggregate verdict. Open Full Results for every judgment. ➖ Not proven improved — item-management (claude-sonnet-5)Why: Net win +16.7% (1W/5T/0L over 6 preference-eligible stimulus vote(s), sign test p=0.500), mean preference +5.7% across 7 paired run(s), 1 dormancy stimulus/stimuli excluded from preference — not credible — 5 of 6 preference-eligible stimulus vote(s) tied, leaving only 1 discordant preference vote(s). The sign test conditions on non-tie stimulus votes and cannot reach 0.05 below 5, so no record could have passed here — this is not a measured null. Either the skill is inert on these scenarios (make them discriminate) or the eval needs more distinct stimuli to clear the ties Next action: Inspect tied or lost stimuli; predeclare added breadth before a new experiment. State: Gate evidence: n=6; 1W/5T/0L; d=1; p=0.500; net +16.7%; 1 dormancy excluded Warnings: Activation: isolated 3/6; plugin 2/6 Overfit: Moderate (score 0.38) Repeated-run reliability (not used by the gate): 7 paired runs (1W/6T/0L). Weak or warning scenarios:
Illustrative judge evidence:
This is one example, not the aggregate verdict. Open Full Results for every judgment. ➖ Not proven improved — item-management (gpt-5.6-luna)Why: Net win +0.0% (1W/4T/1L over 6 preference-eligible stimulus vote(s), sign test p=0.750), mean preference +8.6% across 7 paired run(s), 1 dormancy stimulus/stimuli excluded from preference — no improvement Next action: Inspect tied or lost stimuli; predeclare added breadth before a new experiment. State: Gate evidence: n=6; 1W/4T/1L; d=2; p=0.750; net +0.0%; 1 dormancy excluded Overfit: Moderate (score 0.47) Repeated-run reliability (not used by the gate): 7 paired runs (1W/5T/1L). Weak or warning scenarios:
Illustrative judge evidence:
This is one example, not the aggregate verdict. Open Full Results for every judgment. ➖ Not proven improved — msbuild-antipatterns (claude-sonnet-5)Why: Net win +28.6% (4W/1T/2L over 7 preference-eligible stimulus vote(s), sign test p=0.344), mean preference +10.0% across 8 paired run(s), 1 dormancy stimulus/stimuli excluded from preference — not credible (sign test p=0.344 > 0.05) Next action: Inspect tied or lost stimuli and fix inconsistent skill behavior. State: Gate evidence: n=7; 4W/1T/2L; d=6; p=0.344; net +28.6%; 1 dormancy excluded Warnings: Activation: isolated 2/7; plugin 1/7 Overfit: Moderate (score 0.29) Repeated-run reliability (not used by the gate): 8 paired runs (4W/2T/2L). Weak or warning scenarios:
Illustrative judge evidence:
This is one example, not the aggregate verdict. Open Full Results for every judgment. ➖ Not proven improved — msbuild-modernization (gpt-5.6-luna)Why: Net win +0.0% (2W/2T/2L over 6 preference-eligible stimulus vote(s), sign test p=0.687), mean preference -8.6% across 7 paired run(s), 1 dormancy stimulus/stimuli excluded from preference — no improvement Next action: Inspect tied or lost stimuli; predeclare added breadth before a new experiment. State: Gate evidence: n=6; 2W/2T/2L; d=4; p=0.687; net +0.0%; 1 dormancy excluded Warnings: Activation: isolated 5/6; plugin 6/6 Overfit: Moderate (score 0.35) Repeated-run reliability (not used by the gate): 7 paired runs (2W/3T/2L). Weak or warning scenarios:
Illustrative judge evidence:
This is one example, not the aggregate verdict. Open Full Results for every judgment. Details for 9 results were omitted to keep this comment under GitHub's limit. Open the workflow summary or Full Results for the complete breakdown. 🔍 Full Results - all metrics and investigation details
|
|
✅ Evaluation passed for |
|
✅ Evaluation passed for |
JanKrivanek
left a comment
There was a problem hiding this comment.
Please consider the findings before merging. Otherwise carrying over signoff from #1118
Persist terminal runner failures separately from recoverable tool errors and support complete three-arm databases during cross-directory rejudge. Co-authored-by: Copilot App <223556219+Copilot@users.noreply.github.com>
Persist terminal runner failures separately from recoverable tool errors and support complete three-arm databases during cross-directory rejudge. Co-authored-by: Copilot App <223556219+Copilot@users.noreply.github.com>
Use merge-base change discovery as main advances and require executable MSBuild binlog evidence for single-build scenarios. Co-authored-by: Copilot App <223556219+Copilot@users.noreply.github.com>
…val-infrastructure
|
❌ Evaluation did not complete successfully (the evaluate job reported 36 partial result file(s) were preserved for diagnosis but were not consolidated because the full matrix did not complete. |
📊 Skill and Agent Evaluation Results36 model/target results across 18 targets and 2 models — ✅ 1 improved, ➖ 31 not proven improved, Measurement identity: evaluated commit Measurement health: 36 expected / 36 observed / 36 written; 0 missing, 0 unexpected, 0 invalid; 0 recovered comparison error slots and 0 unresolved comparison error slots. Objective completion gate: not enabled. Aggregate completion transitions are telemetry only, so this report does not claim that zero objective regressions were proven. A result passes only when preference-eligible distinct-stimulus votes have aggregate net win of at least 20% and an exact one-sided sign-test result of
ℹ️ How to read this report
⛔ Activation contract failed — build-perf-baseline (gpt-5.6-luna)Why: Net win +83.3% (5W/1T/0L over 6 preference-eligible stimulus vote(s), sign test p=0.031), mean preference +28.6% across 7 paired run(s), 1 dormancy stimulus/stimuli excluded from preference — activation contract failed (1 explicit dormancy scenario(s) activated the isolated target skill) Next action: Narrow skill routing so the listed off-target scenarios stay dormant. State: Gate evidence: n=6; 5W/1T/0L; d=5; p=0.031; net +83.3%; 1 dormancy excluded Warnings: Dormancy contract: 1 unexpected activation(s); Activation: isolated 6/6; plugin 5/6 Overfit: Moderate (score 0.30) Repeated-run reliability (not used by the gate): 7 paired runs (5W/2T/0L). Weak or warning scenarios:
Illustrative judge evidence:
This is one example, not the aggregate verdict. Open Full Results for every judgment. ⛔ Activation contract failed — directory-build-organization (claude-sonnet-5)Why: Net win -16.7% (1W/3T/2L over 6 preference-eligible stimulus vote(s), sign test p=0.500), mean preference -5.7% across 7 paired run(s), 1 dormancy stimulus/stimuli excluded from preference — activation contract failed (1 explicit dormancy scenario(s) activated the isolated target skill) Next action: Narrow skill routing so the listed off-target scenarios stay dormant. State: Gate evidence: n=6; 1W/3T/2L; d=3; p=0.500; net -16.7%; 1 dormancy excluded Warnings: Dormancy contract: 1 unexpected activation(s); Activation: isolated 3/6; plugin 5/6 Overfit: Moderate (score 0.36) Repeated-run reliability (not used by the gate): 7 paired runs (1W/4T/2L). Weak or warning scenarios:
Illustrative judge evidence:
This is one example, not the aggregate verdict. Open Full Results for every judgment. ⛔ Activation contract failed — directory-build-organization (gpt-5.6-luna)Why: Net win +50.0% (3W/3T/0L over 6 preference-eligible stimulus vote(s), sign test p=0.125), mean preference +17.1% across 7 paired run(s), 1 dormancy stimulus/stimuli excluded from preference — activation contract failed (1 explicit dormancy scenario(s) activated the isolated target skill) Next action: Narrow skill routing so the listed off-target scenarios stay dormant. State: Gate evidence: n=6; 3W/3T/0L; d=3; p=0.125; net +50.0%; 1 dormancy excluded Warnings: Dormancy contract: 1 unexpected activation(s) Overfit: High (score 0.54) Repeated-run reliability (not used by the gate): 7 paired runs (3W/4T/0L). Weak or warning scenarios:
Illustrative judge evidence:
This is one example, not the aggregate verdict. Open Full Results for every judgment. ⛔ Activation contract failed — extension-points (gpt-5.6-luna)Why: Net win +14.3% (3W/2T/2L over 7 preference-eligible stimulus vote(s), sign test p=0.500), mean preference +6.7% across 9 paired run(s), 2 dormancy stimulus/stimuli excluded from preference — activation contract failed (1 explicit dormancy scenario(s) activated the isolated target skill) Next action: Narrow skill routing so the listed off-target scenarios stay dormant. State: Gate evidence: n=7; 3W/2T/2L; d=5; p=0.500; net +14.3%; 2 dormancy excluded Warnings: Dormancy contract: 1 unexpected activation(s); Activation: isolated 7/7; plugin 6/7 Overfit: Moderate (score 0.36) Repeated-run reliability (not used by the gate): 9 paired runs (3W/3T/3L). Weak or warning scenarios:
Illustrative judge evidence:
This is one example, not the aggregate verdict. Open Full Results for every judgment. ➖ Not proven improved — agent.msbuild (claude-sonnet-5)Why: Net win +0.0% (2W/1T/2L over 5 preference-eligible stimulus vote(s), sign test p=0.687), mean preference +12.0% across 5 paired run(s) — no improvement — native evaluator reported that the target agent did not activate Next action: Inspect tied or lost stimuli; predeclare added breadth before a new experiment. State: Gate evidence: n=5; 2W/1T/2L; d=4; p=0.687; net +0.0% Warnings: Activation: isolated 1/5; plugin 0/5 Repeated-run reliability (not used by the gate): 5 paired runs (2W/1T/2L). Weak or warning scenarios:
Illustrative judge evidence:
This is one example, not the aggregate verdict. Open Full Results for every judgment. ➖ Not proven improved — agent.msbuild (gpt-5.6-luna)Why: Net win +20.0% (2W/2T/1L over 5 preference-eligible stimulus vote(s), sign test p=0.500), mean preference +8.0% across 5 paired run(s) — not credible — 2 of 5 preference-eligible stimulus vote(s) tied, leaving only 3 discordant preference vote(s). The sign test conditions on non-tie stimulus votes and cannot reach 0.05 below 5, so no record could have passed here — this is not a measured null. Either the agent is inert on these scenarios (make them discriminate) or the eval needs more distinct stimuli to clear the ties — native evaluator reported that the target agent did not activate Next action: Inspect tied or lost stimuli; predeclare added breadth before a new experiment. State: Gate evidence: n=5; 2W/2T/1L; d=3; p=0.500; net +20.0% Warnings: Activation: isolated 2/5; plugin 0/5 Repeated-run reliability (not used by the gate): 5 paired runs (2W/2T/1L). Weak or warning scenarios:
Illustrative judge evidence:
This is one example, not the aggregate verdict. Open Full Results for every judgment. ➖ Not proven improved — binlog-failure-analysis (claude-sonnet-5)Why: Net win +16.7% (3W/1T/2L over 6 preference-eligible stimulus vote(s), sign test p=0.500), mean preference +5.7% across 7 paired run(s), 1 dormancy stimulus/stimuli excluded from preference — not credible (sign test p=0.500 > 0.05) Next action: Inspect tied or lost stimuli and fix inconsistent skill behavior. State: Gate evidence: n=6; 3W/1T/2L; d=5; p=0.500; net +16.7%; 1 dormancy excluded Warnings: Activation: isolated 5/7; plugin 6/7 Overfit: Moderate (score 0.36) Repeated-run reliability (not used by the gate): 7 paired runs (3W/2T/2L). Weak or warning scenarios:
Illustrative judge evidence:
This is one example, not the aggregate verdict. Open Full Results for every judgment. ➖ Not proven improved — binlog-failure-analysis (gpt-5.6-luna)Why: Net win +16.7% (2W/3T/1L over 6 preference-eligible stimulus vote(s), sign test p=0.500), mean preference -2.9% across 7 paired run(s), 1 dormancy stimulus/stimuli excluded from preference — not credible — 3 of 6 preference-eligible stimulus vote(s) tied, leaving only 3 discordant preference vote(s). The sign test conditions on non-tie stimulus votes and cannot reach 0.05 below 5, so no record could have passed here — this is not a measured null. Either the skill is inert on these scenarios (make them discriminate) or the eval needs more distinct stimuli to clear the ties Next action: Inspect tied or lost stimuli; predeclare added breadth before a new experiment. State: Gate evidence: n=6; 2W/3T/1L; d=3; p=0.500; net +16.7%; 1 dormancy excluded Warnings: Activation: isolated 6/7; plugin 6/7 Overfit: Low (score 0.18) Repeated-run reliability (not used by the gate): 7 paired runs (2W/4T/1L). Weak or warning scenarios:
Illustrative judge evidence:
This is one example, not the aggregate verdict. Open Full Results for every judgment. ➖ Not proven improved — binlog-generation (claude-sonnet-5)Why: Net win +57.1% (4W/3T/0L over 7 preference-eligible stimulus vote(s), sign test p=0.063), mean preference +42.5% across 8 paired run(s), 1 dormancy stimulus/stimuli excluded from preference — not credible — 3 of 7 preference-eligible stimulus vote(s) tied, leaving only 4 discordant preference vote(s). The sign test conditions on non-tie stimulus votes and cannot reach 0.05 below 5, so no record could have passed here — this is not a measured null. Either the skill is inert on these scenarios (make them discriminate) or the eval needs more distinct stimuli to clear the ties Next action: Inspect tied or lost stimuli; predeclare added breadth before a new experiment. State: Gate evidence: n=7; 4W/3T/0L; d=4; p=0.063; net +57.1%; 1 dormancy excluded Warnings: Activation: isolated 6/7; plugin 5/7 Overfit: Moderate (score 0.48) Repeated-run reliability (not used by the gate): 8 paired runs (4W/4T/0L). Weak or warning scenarios:
Illustrative judge evidence:
This is one example, not the aggregate verdict. Open Full Results for every judgment. ➖ Not proven improved — build-parallelism (claude-sonnet-5)Why: Net win +0.0% (1W/4T/1L over 6 preference-eligible stimulus vote(s), sign test p=0.750), mean preference -5.7% across 7 paired run(s), 1 dormancy stimulus/stimuli excluded from preference — no improvement Next action: Inspect activation-only failed runs before rewriting skill content; the model stopped after loading a skill. State: Gate evidence: n=6; 1W/4T/1L; d=2; p=0.750; net +0.0%; 1 dormancy excluded Warnings: Activation: isolated 3/6; plugin 6/6; Activation-only stop: isolated 1 failed run; Activation-only stop: plugin 1 failed run Overfit: Low (score 0.18) Repeated-run reliability (not used by the gate): 7 paired runs (1W/4T/2L). Weak or warning scenarios:
Illustrative judge evidence:
This is one example, not the aggregate verdict. Open Full Results for every judgment. ➖ Not proven improved — build-parallelism (gpt-5.6-luna)Why: Net win +33.3% (3W/2T/1L over 6 preference-eligible stimulus vote(s), sign test p=0.312), mean preference +14.3% across 7 paired run(s), 1 dormancy stimulus/stimuli excluded from preference — not credible — 2 of 6 preference-eligible stimulus vote(s) tied, leaving only 4 discordant preference vote(s). The sign test conditions on non-tie stimulus votes and cannot reach 0.05 below 5, so no record could have passed here — this is not a measured null. Either the skill is inert on these scenarios (make them discriminate) or the eval needs more distinct stimuli to clear the ties Next action: Inspect activation-only failed runs before rewriting skill content; the model stopped after loading a skill. State: Gate evidence: n=6; 3W/2T/1L; d=4; p=0.312; net +33.3%; 1 dormancy excluded Warnings: Activation-only stop: isolated 1 failed run; Activation-only stop: plugin 1 failed run Overfit: High (score 0.55) Repeated-run reliability (not used by the gate): 7 paired runs (3W/2T/2L). Weak or warning scenarios:
Illustrative judge evidence:
This is one example, not the aggregate verdict. Open Full Results for every judgment. ➖ Not proven improved — build-perf-baseline (claude-sonnet-5)Why: Net win +66.7% (5W/0T/1L over 6 preference-eligible stimulus vote(s), sign test p=0.109), mean preference +31.4% across 7 paired run(s), 1 dormancy stimulus/stimuli excluded from preference — not credible (sign test p=0.109 > 0.05) Next action: Inspect tied or lost stimuli and fix inconsistent skill behavior. State: Gate evidence: n=6; 5W/0T/1L; d=6; p=0.109; net +66.7%; 1 dormancy excluded Warnings: Activation: isolated 4/6; plugin 5/6 Overfit: Moderate (score 0.33) Repeated-run reliability (not used by the gate): 7 paired runs (5W/1T/1L). Weak or warning scenarios:
Illustrative judge evidence:
This is one example, not the aggregate verdict. Open Full Results for every judgment. ➖ Not proven improved — build-perf-diagnostics (claude-sonnet-5)Why: Net win +28.6% (4W/1T/2L over 7 preference-eligible stimulus vote(s), sign test p=0.344), mean preference +10.0% across 8 paired run(s), 1 dormancy stimulus/stimuli excluded from preference — not credible (sign test p=0.344 > 0.05) Next action: Inspect tied or lost stimuli and fix inconsistent skill behavior. State: Gate evidence: n=7; 4W/1T/2L; d=6; p=0.344; net +28.6%; 1 dormancy excluded Overfit: High (score 0.52) Repeated-run reliability (not used by the gate): 8 paired runs (4W/2T/2L). Weak or warning scenarios:
Illustrative judge evidence:
This is one example, not the aggregate verdict. Open Full Results for every judgment. ➖ Not proven improved — build-perf-diagnostics (gpt-5.6-luna)Why: Net win +28.6% (2W/5T/0L over 7 preference-eligible stimulus vote(s), sign test p=0.250), mean preference +10.0% across 8 paired run(s), 1 dormancy stimulus/stimuli excluded from preference — not credible — 5 of 7 preference-eligible stimulus vote(s) tied, leaving only 2 discordant preference vote(s). The sign test conditions on non-tie stimulus votes and cannot reach 0.05 below 5, so no record could have passed here — this is not a measured null. Either the skill is inert on these scenarios (make them discriminate) or the eval needs more distinct stimuli to clear the ties Next action: Inspect tied or lost stimuli; predeclare added breadth before a new experiment. State: Gate evidence: n=7; 2W/5T/0L; d=2; p=0.250; net +28.6%; 1 dormancy excluded Overfit: Moderate (score 0.42) Repeated-run reliability (not used by the gate): 8 paired runs (2W/6T/0L). Weak or warning scenarios:
Illustrative judge evidence:
This is one example, not the aggregate verdict. Open Full Results for every judgment. ➖ Not proven improved — check-bin-obj-clash (claude-sonnet-5)Why: Net win +0.0% (3W/1T/3L over 7 preference-eligible stimulus vote(s), sign test p=0.656), mean preference +8.6% across 7 paired run(s) — no improvement Next action: Inspect tied or lost stimuli and fix inconsistent skill behavior. State: Gate evidence: n=7; 3W/1T/3L; d=6; p=0.656; net +0.0% Warnings: Activation: isolated 7/7; plugin 6/7 Overfit: Low (score 0.17) Repeated-run reliability (not used by the gate): 7 paired runs (3W/1T/3L). Weak or warning scenarios:
Illustrative judge evidence:
This is one example, not the aggregate verdict. Open Full Results for every judgment. ➖ Not proven improved — check-bin-obj-clash (gpt-5.6-luna)Why: Net win +0.0% (3W/1T/3L over 7 preference-eligible stimulus vote(s), sign test p=0.656), mean preference +0.0% across 7 paired run(s) — no improvement Next action: Inspect tied or lost stimuli and fix inconsistent skill behavior. State: Gate evidence: n=7; 3W/1T/3L; d=6; p=0.656; net +0.0% Overfit: Moderate (score 0.22) Repeated-run reliability (not used by the gate): 7 paired runs (3W/1T/3L). Weak or warning scenarios:
Illustrative judge evidence:
This is one example, not the aggregate verdict. Open Full Results for every judgment. ➖ Not proven improved — eval-performance (claude-sonnet-5)Why: Net win +50.0% (4W/4T/0L over 8 preference-eligible stimulus vote(s), sign test p=0.063), mean preference +20.0% across 8 paired run(s) — not credible — 4 of 8 preference-eligible stimulus vote(s) tied, leaving only 4 discordant preference vote(s). The sign test conditions on non-tie stimulus votes and cannot reach 0.05 below 5, so no record could have passed here — this is not a measured null. Either the skill is inert on these scenarios (make them discriminate) or the eval needs more distinct stimuli to clear the ties Next action: Inspect activation-only failed runs before rewriting skill content; the model stopped after loading a skill. State: Gate evidence: n=8; 4W/4T/0L; d=4; p=0.063; net +50.0% Warnings: Activation: isolated 6/8; plugin 6/8; Activation-only stop: plugin 1 failed run Overfit: Moderate (score 0.28) Repeated-run reliability (not used by the gate): 8 paired runs (4W/4T/0L). Weak or warning scenarios:
Illustrative judge evidence:
This is one example, not the aggregate verdict. Open Full Results for every judgment. ➖ Not proven improved — eval-performance (gpt-5.6-luna)Why: Net win +50.0% (4W/4T/0L over 8 preference-eligible stimulus vote(s), sign test p=0.063), mean preference +35.0% across 8 paired run(s) — not credible — 4 of 8 preference-eligible stimulus vote(s) tied, leaving only 4 discordant preference vote(s). The sign test conditions on non-tie stimulus votes and cannot reach 0.05 below 5, so no record could have passed here — this is not a measured null. Either the skill is inert on these scenarios (make them discriminate) or the eval needs more distinct stimuli to clear the ties Next action: Inspect tied or lost stimuli; predeclare added breadth before a new experiment. State: Gate evidence: n=8; 4W/4T/0L; d=4; p=0.063; net +50.0% Warnings: Activation: isolated 8/8; plugin 7/8 Overfit: Moderate (score 0.26) Repeated-run reliability (not used by the gate): 8 paired runs (4W/4T/0L). Weak or warning scenarios:
Illustrative judge evidence:
This is one example, not the aggregate verdict. Open Full Results for every judgment. ➖ Not proven improved — extension-points (claude-sonnet-5)Why: Net win -14.3% (2W/2T/3L over 7 preference-eligible stimulus vote(s), sign test p=0.500), mean preference -4.4% across 9 paired run(s), 2 dormancy stimulus/stimuli excluded from preference — no improvement Next action: Inspect tied or lost stimuli and fix inconsistent skill behavior. State: Gate evidence: n=7; 2W/2T/3L; d=5; p=0.500; net -14.3%; 2 dormancy excluded Warnings: Activation: isolated 4/7; plugin 5/7 Overfit: Moderate (score 0.31) Repeated-run reliability (not used by the gate): 9 paired runs (2W/4T/3L). Weak or warning scenarios:
Illustrative judge evidence:
This is one example, not the aggregate verdict. Open Full Results for every judgment. ➖ Not proven improved — including-generated-files (claude-sonnet-5)Why: Net win +0.0% (2W/3T/2L over 7 preference-eligible stimulus vote(s), sign test p=0.687), mean preference -5.0% across 8 paired run(s), 1 dormancy stimulus/stimuli excluded from preference — no improvement Next action: Inspect activation-only failed runs before rewriting skill content; the model stopped after loading a skill. State: Gate evidence: n=7; 2W/3T/2L; d=4; p=0.687; net +0.0%; 1 dormancy excluded Warnings: Activation: isolated 6/7; plugin 5/7; Activation-only stop: plugin 1 failed run Overfit: Moderate (score 0.49) Repeated-run reliability (not used by the gate): 8 paired runs (2W/3T/3L). Weak or warning scenarios:
Illustrative judge evidence:
This is one example, not the aggregate verdict. Open Full Results for every judgment. ➖ Not proven improved — including-generated-files (gpt-5.6-luna)Why: Net win +28.6% (4W/1T/2L over 7 preference-eligible stimulus vote(s), sign test p=0.344), mean preference +10.0% across 8 paired run(s), 1 dormancy stimulus/stimuli excluded from preference — not credible (sign test p=0.344 > 0.05) Next action: Inspect tied or lost stimuli and fix inconsistent skill behavior. State: Gate evidence: n=7; 4W/1T/2L; d=6; p=0.344; net +28.6%; 1 dormancy excluded Overfit: Moderate (score 0.43) Repeated-run reliability (not used by the gate): 8 paired runs (4W/2T/2L). Weak or warning scenarios:
Illustrative judge evidence:
This is one example, not the aggregate verdict. Open Full Results for every judgment. ➖ Not proven improved — incremental-build (claude-sonnet-5)Why: Net win -22.2% (1W/5T/3L over 9 preference-eligible stimulus vote(s), sign test p=0.312), mean preference -8.9% across 9 paired run(s) — no improvement Next action: Inspect tied or lost stimuli; predeclare added breadth before a new experiment. State: Gate evidence: n=9; 1W/5T/3L; d=4; p=0.312; net -22.2% Warnings: Activation: isolated 4/9; plugin 5/9 Overfit: Moderate (score 0.32) Repeated-run reliability (not used by the gate): 9 paired runs (1W/5T/3L). Weak or warning scenarios:
Illustrative judge evidence:
This is one example, not the aggregate verdict. Open Full Results for every judgment. ➖ Not proven improved — incremental-build (gpt-5.6-luna)Why: Net win +44.4% (5W/3T/1L over 9 preference-eligible stimulus vote(s), sign test p=0.109), mean preference +17.8% across 9 paired run(s) — not credible (sign test p=0.109 > 0.05) Next action: Inspect tied or lost stimuli and fix inconsistent skill behavior. State: Gate evidence: n=9; 5W/3T/1L; d=6; p=0.109; net +44.4% Warnings: Activation: isolated 8/9; plugin 7/9 Overfit: Moderate (score 0.50) Repeated-run reliability (not used by the gate): 9 paired runs (5W/3T/1L). Weak or warning scenarios:
Illustrative judge evidence:
This is one example, not the aggregate verdict. Open Full Results for every judgment. ➖ Not proven improved — item-management (claude-sonnet-5)Why: Net win +16.7% (3W/1T/2L over 6 preference-eligible stimulus vote(s), sign test p=0.500), mean preference +14.3% across 7 paired run(s), 1 dormancy stimulus/stimuli excluded from preference — not credible (sign test p=0.500 > 0.05) Next action: Inspect tied or lost stimuli and fix inconsistent skill behavior. State: Gate evidence: n=6; 3W/1T/2L; d=5; p=0.500; net +16.7%; 1 dormancy excluded Warnings: Activation: isolated 4/6; plugin 3/6 Overfit: Moderate (score 0.37) Repeated-run reliability (not used by the gate): 7 paired runs (3W/2T/2L). Weak or warning scenarios:
Illustrative judge evidence:
This is one example, not the aggregate verdict. Open Full Results for every judgment. ➖ Not proven improved — item-management (gpt-5.6-luna)Why: Net win +33.3% (2W/4T/0L over 6 preference-eligible stimulus vote(s), sign test p=0.250), mean preference +20.0% across 7 paired run(s), 1 dormancy stimulus/stimuli excluded from preference — not credible — 4 of 6 preference-eligible stimulus vote(s) tied, leaving only 2 discordant preference vote(s). The sign test conditions on non-tie stimulus votes and cannot reach 0.05 below 5, so no record could have passed here — this is not a measured null. Either the skill is inert on these scenarios (make them discriminate) or the eval needs more distinct stimuli to clear the ties Next action: Inspect tied or lost stimuli; predeclare added breadth before a new experiment. State: Gate evidence: n=6; 2W/4T/0L; d=2; p=0.250; net +33.3%; 1 dormancy excluded Overfit: Moderate (score 0.42) Repeated-run reliability (not used by the gate): 7 paired runs (2W/5T/0L). Weak or warning scenarios:
Illustrative judge evidence:
This is one example, not the aggregate verdict. Open Full Results for every judgment. ➖ Not proven improved — msbuild-antipatterns (claude-sonnet-5)Why: Net win -28.6% (1W/3T/3L over 7 preference-eligible stimulus vote(s), sign test p=0.312), mean preference -5.0% across 8 paired run(s), 1 dormancy stimulus/stimuli excluded from preference — no improvement Next action: Inspect tied or lost stimuli; predeclare added breadth before a new experiment. State: Gate evidence: n=7; 1W/3T/3L; d=4; p=0.312; net -28.6%; 1 dormancy excluded Warnings: Activation: isolated 3/7; plugin 1/7 Overfit: Moderate (score 0.31) Repeated-run reliability (not used by the gate): 8 paired runs (2W/3T/3L). Weak or warning scenarios:
Illustrative judge evidence:
This is one example, not the aggregate verdict. Open Full Results for every judgment. ➖ Not proven improved — msbuild-antipatterns (gpt-5.6-luna)Why: Net win +0.0% (1W/5T/1L over 7 preference-eligible stimulus vote(s), sign test p=0.750), mean preference -12.5% across 8 paired run(s), 1 dormancy stimulus/stimuli excluded from preference — no improvement Next action: Inspect tied or lost stimuli; predeclare added breadth before a new experiment. State: Gate evidence: n=7; 1W/5T/1L; d=2; p=0.750; net +0.0%; 1 dormancy excluded Warnings: Activation: isolated 4/7; plugin 3/7 Overfit: Moderate (score 0.39) Repeated-run reliability (not used by the gate): 8 paired runs (1W/5T/2L). Weak or warning scenarios:
Illustrative judge evidence:
This is one example, not the aggregate verdict. Open Full Results for every judgment. Details for 9 results were omitted to keep this comment under GitHub's limit. Open the workflow summary or Full Results for the complete breakdown. 🔍 Full Results - all metrics and investigation details
|
|
✅ Approved by @JanKrivanek. cc @dotnet/skills-merge-approvers — ready to merge. |


Summary
This is a clean replacement for #1118, rebuilt from current
main. It expands the MSBuild evaluation surface to 18 specs and 135 distinct stimuli with self-contained fixtures, deterministic graders, replayable patches, and activation-aware routing boundaries.The change also includes only the shared infrastructure required to execute those evaluations safely and interpret their results correctly. It excludes MSBuild skill and agent content, host manifests, dotnet-test content, dashboard/versioning work, and model-evaluation actions.
Why
The prior evaluation surface mixed real skill behavior with underpowered suites, contradictory fixtures, unsafe evaluator file access, incomplete result accounting, and ambiguous activation evidence. Those defects could produce false regressions, hide terminal failures, or treat a sibling skill invocation as target activation.
This replacement separates those causes. Each MSBuild suite now has enough distinct scenarios to produce a credible result, and the evaluator fails closed when execution, pairing, dormancy evidence, or sandbox containment is incomplete.
Impact
MSBuild evaluation changes
expect_activation: falsescenariosShared evaluation infrastructure
origin/mainfor manual runs; reserve repository-wide audits for explicit local useexecutionShardonly from the top-leveltagsmappingagent.primary_selectedand deduplicate it with SDK subagent eventsShared infrastructure admission ledger (37 candidates)
.github/workflows/eval-quality.ymlorigin/maineng/eval-quality/README.mdeng/eval-quality/check_eval_quality.pyeng/eval-quality/selftest_eval_quality.pyeng/eval-quality/underpowered-allowlist.txteng/evaluation/find-targets.ps1tags.executionShardeng/evaluation/test_token_failover.pyeng/skill-validator/src/Evaluate/AgentRunner.cseng/skill-validator/src/Evaluate/Comparator.cseng/skill-validator/src/Evaluate/EvalSchema.cseng/skill-validator/src/Evaluate/EvaluateCommand.cseng/skill-validator/src/Evaluate/LlmSession.cseng/skill-validator/src/Evaluate/LocalSessionFsHandler.cseng/skill-validator/src/Evaluate/MetricsCollector.csagent.primary_selectedevidenceeng/skill-validator/src/Evaluate/Models.cseng/skill-validator/src/Evaluate/OverfittingCommand.cseng/skill-validator/src/Evaluate/RejudgeCommand.cseng/skill-validator/src/Evaluate/Reporter.cseng/skill-validator/src/Evaluate/SecureFileSystem.cseng/skill-validator/src/Evaluate/SessionDatabase.cseng/skill-validator/src/README.mdeng/skill-validator/src/docs/InvestigatingResults.mdeng/skill-validator/tests/Evaluate/ComparatorTests.cseng/skill-validator/tests/Evaluate/EvalDiscoveryTests.cseng/skill-validator/tests/Evaluate/EvaluateCommandTests.cseng/skill-validator/tests/Evaluate/MetricsTests.cseng/skill-validator/tests/Evaluate/RejudgeCommandTests.cseng/skill-validator/tests/Evaluate/RunnerTests.cseng/skill-validator/tests/Evaluate/SessionDatabaseTests.cseng/vally-adapter/InvestigatingResults.mdeng/vally-adapter/README.mdeng/vally-adapter/adapt.mjseng/vally-adapter/adapt.test.mjs.agents/skills/create-skill-test/SKILL.md.agents/skills/improve-skill-quality/SKILL.md.agents/skills/improve-skill-quality/references/eval-triage.mdeng/skill-validator/tests/Check/PluginMcpManifestTests.csValidation
python eng/eval-quality/check_eval_quality.py --base-ref origin/main— passed for 18 changed suites of 101 specspython eng/eval-quality/selftest_eval_quality.py— 127 passedpython -m unittest discover -s eng/evaluation -p "test_*.py" -v— 58 passednode --test --test-concurrency=1 eng/vally-adapter/*.test.mjs— 203 passeddotnet test eng/skill-validator/tests/SkillValidator.Tests.csproj— 834 passedosx-arm64cross-build withPublishAot=false— passed with zero warnings or errorsosx-arm64publish was not run because the .NET IL compiler does not support cross-OS native compilation from Windows; Darwin-specific runtime coverage is included for macOS CIdotnet run --project eng/skill-validator/src/SkillValidator.csproj -- check --plugin ./plugins/dotnet-msbuild— passed for 18 skills, 3 agents, and 1 pluginactionlint1.7.7 with the pinned archive checksum — passedReview flow
flowchart LR A[Tracked MSBuild fixture] --> B[Eval-quality preflight] B --> C[Baseline / isolated / plugin execution] C --> D{Evidence complete?} D -- yes --> E[Exact accounting and verdict] D -- bounded transient failure --> F[Targeted recovery] F --> E D -- incomplete or ambiguous --> G[Fail closed]Suggested review order
tests/dotnet-msbuild/**for scenario ownership, neutral fixtures, executable graders, and golden evidence.eng/eval-quality/**andeng/evaluation/**for deterministic preflight and discovery.eng/skill-validator/**for sandboxing, permissions, activation, session completeness, and rejudge behavior.eng/vally-adapter/**for exact accounting, activation contracts, and bounded recovery.SKILL.mdguidance.