Skip to content

Remove XML tags from skill descriptions for Claude compatibility - #1260

Merged
AbhitejJohn merged 2 commits into
mainfrom
ykovalova-fix-skill-description-xml-tags
Oct 5, 2026
Merged

AbhitejJohn merged 2 commits into
mainfrom
ykovalova-fix-skill-description-xml-tags

Conversation

@YuliiaKovalova

Copy link
Copy Markdown
Member

Summary

Claude (claude.ai and Claude Code marketplace sync) rejects skill descriptions that contain XML tags and stores them with the angle brackets stripped. Four skills were reported after a marketplace sync.

  • Reworded the tag-like text in the four affected descriptions in plain language (no change to routing intent):
    • dotnet-advanced/vectorization: Vector<T> -> "the generic Vector type"
    • dotnet-blazor/configure-auth: <NotAuthorized> -> "NotAuthorized"
    • dotnet-msbuild/msbuild-modernization: <Compile Include> / <Import Project=...> reworded; > 50 lines -> "over 50 lines"
    • dotnet11/system-text-json-net11: GetTypeInfo<T>() / TryGetTypeInfo<T>(out ...) -> generic GetTypeInfo / TryGetTypeInfo overloads
  • Added a skill-validator check rule that errors on XML-like tags in a skill description, so this cannot regress. The matcher requires a tag name followed only by bare words or name=value attributes (plus comments, processing instructions and CDATA), so comparisons such as >5s, a < b and i<length && count>0 are not flagged. Names may start with _.
  • Documented the rule in the create-skill authoring guide (description rules + checklist).

Descriptions that use >/< only as comparison operators (build-perf-diagnostics, eval-performance) are intentionally unchanged; they are not tags and were not reported.

Supersedes #1259, which came from a fork and so could not run the eng/ CI with the right permissions. Review feedback from that PR is already incorporated.

Related issue

Fixes #1251

Validation

  • dotnet test in eng/skill-validator: 878 passed, 0 failed (includes 14 cases for the XML-tag rule).
  • skill-validator check --plugin plugins/* with the repo's allow-lists: all checks passed (100 skills, 16 agents, 17 plugins).
  • Scanned every SKILL.md description (parsed YAML) for tag-like text: 0 remaining; all descriptions are still under 1,024 characters.

Checklist

  • I searched existing issues and pull requests to avoid duplicates.
  • I kept this pull request focused and avoided unrelated refactors.
  • I added or updated tests, evals, or documentation when changing skill or agent behavior.
  • I updated CODEOWNERS when adding or moving owned content. (N/A: no new content)
  • I updated all marketplace manifests when plugin metadata changed. (N/A: no manifest metadata changed)
  • I updated eng/known-domains.txt for any new external domains referenced by skill content. (N/A: no new domains)

YuliiaKovalova and others added 2 commits October 5, 2026 11:02
Claude (claude.ai and Claude Code marketplace sync) rejects skill
descriptions that contain XML tags and strips the angle brackets. Reword
the four affected descriptions in plain language:

- dotnet-advanced/vectorization
- dotnet-blazor/configure-auth
- dotnet-msbuild/msbuild-modernization
- dotnet11/system-text-json-net11

Add a skill-validator check that errors on XML-like tags in a skill
description so this cannot regress, with tests, and document the rule in
the create-skill authoring guide.

Fixes #1251

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
Copilot-Session: bcd44b69-8a83-470f-a43b-41606d2b67de
…arisons

Address review feedback:
- Accept '_' as an XML name start so tags like <_Root /> are flagged.
- Match a tag name followed only by bare words or name=value attributes
  (plus comments, PIs and CDATA) instead of arbitrary text between
  brackets, so unspaced comparisons such as i<length && count>0 no longer
  fail the check.

Add regression cases for both.

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
Copilot-Session: bcd44b69-8a83-470f-a43b-41606d2b67de
@github-actions

github-actions Bot commented Oct 5, 2026

Copy link
Copy Markdown
Contributor

Skill Coverage Report

Plugin Skill Covered Coverage
⚠️ dotnet-blazor configure-auth 2/3 66.7%
✅ dotnet-msbuild msbuild-modernization 7/7 100%
✅ dotnet11 system-text-json-net11 15/16 93.8%
Uncovered: dotnet-blazor/configure-auth
  • [CodePattern] [CascadingParameter] (line 51)
Uncovered: dotnet11/system-text-json-net11
  • [CodePattern] [guid] (line 48)

Copilot AI left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Copilot review overview

🟢 Approval recommended

The focused implementation matches the stated compatibility requirement and includes appropriate regression coverage.

Review effort: Balanced
Findings: None

What changed in this PR

Removes XML-like syntax from skill descriptions and prevents future Claude marketplace compatibility regressions.

Changes:

  • Reworded four affected skill descriptions without changing routing intent.
  • Added XML-like tag validation with positive and negative tests.
  • Documented the restriction in the skill-authoring guide.
File Description
plugins/​dotnet11/​skills/​system-text-json-net11/​SKILL.md Rewords generic API references.
plugins/​dotnet-msbuild/​skills/​msbuild-modernization/​SKILL.md Rewords MSBuild XML examples.
plugins/​dotnet-blazor/​skills/​configure-auth/​SKILL.md Removes tag syntax from NotAuthorized.
plugins/​dotnet-advanced/​skills/​vectorization/​SKILL.md Rewords the generic Vector reference.
eng/​skill-validator/​tests/​Check/​SkillProfileTests.cs Tests tag detection and comparison exclusions.
eng/​skill-validator/​src/​Check/​SkillProfiler.cs Adds description tag validation.
.agents/​skills/​create-skill/​SKILL.md Documents the new authoring rule.

💡 Add a code-review agent skill or configure MCP servers for context-aware, tailored reviews. Learn more in the docs.

@github-actions github-actions Bot added pr-state/ready-for-eval PR is mergeable and awaiting evaluation pr-state/evals-in-progress PR evaluations are in progress and removed pr-state/ready-for-eval PR is mergeable and awaiting evaluation labels Oct 5, 2026
@github-actions

github-actions Bot commented Oct 5, 2026

Copy link
Copy Markdown
Contributor

❌ Evaluation did not complete successfully (the evaluate job reported failure). Check the workflow run logs, then comment /evaluate 1603d1b6467800a64426107881b42de6ccc69979 to retry this exact commit.

10 partial result file(s) were preserved for diagnosis but were not consolidated because the full matrix did not complete.

⚠️ Session replay telemetry was not published because the auxiliary dotnet/skills-data publisher failed. The evaluation verdicts above remain authoritative; maintainers must repair SKILLS_DATA_TOKEN or the publisher before replay links are available.

github-actions Bot added a commit that referenced this pull request Oct 5, 2026
@github-actions github-actions Bot added pr-state/ready-for-eval PR is mergeable and awaiting evaluation and removed pr-state/evals-in-progress PR evaluations are in progress labels Oct 5, 2026
@YuliiaKovalova

Copy link
Copy Markdown
Member Author

/evaluate 1603d1b

@github-actions

github-actions Bot commented Oct 5, 2026

Copy link
Copy Markdown
Contributor

📊 Skill and Agent Evaluation Results

10 model/target results across 5 targets and 2 models — ✅ 2 improved, ➖ 5 results without a clear winner, ⚠️ 2 invalid or underpowered, ⛔ 0 activation contract failures, 📉 0 preference losses (report only), 🔻 1 objective regressions.

Measurement identity: evaluated commit 1603d1b6467800a64426107881b42de6ccc69979; 2 judge models.

Measurement health: 10 expected / 10 observed / 10 written; 0 missing, 0 unexpected, 2 invalid; 0 recovered comparison error slots and 0 unresolved comparison error slots.

Objective completion gate: enabled; 1 result shows a proven objective regression.

A result passes only when preference-eligible distinct-stimulus votes have aggregate net win of at least 20% and an exact one-sided sign-test result of p ≤ 0.05, and every explicit dormancy activation contract passes. Repeated runs measure reliability only.

Target Model Verdict Gate evidence Overfit Warnings Next action
agent.msbuild claude-sonnet-5 🔻 Objective regression n=5; 0W/4T/1L; d=1; p=0.500; net -20.0% — — Inspect objective completion losses and fix them before merge.
agent.msbuild gpt-5.6-luna ➖ Improvement signal, tie-limited n=5; 3W/1T/1L; d=4; p=0.312; net +40.0% — — The signal favors the target, but ties leave too few discordant tasks; inspect ties and predeclare more discriminating breadth.
configure-auth claude-sonnet-5 ⚠️ Underpowered n=2; 1W/0T/1L; d=2; p=0.750; net +0.0% 🔴 0.57 — Predeclare more independent, discriminating stimuli; repeated runs do not add power.
configure-auth gpt-5.6-luna ⚠️ Underpowered n=2; 0W/0T/2L; d=2; p=0.250; net -100.0% 🟡 0.36 Activation: isolated 2/2; plugin 1/2 Predeclare more independent, discriminating stimuli; repeated runs do not add power.
msbuild-modernization claude-sonnet-5 ➖ Improvement signal, unproven n=6; 3W/1T/2L; d=5; p=0.500; net +16.7%; 1 dormancy excluded 🟡 0.42 Activation: isolated 6/6; plugin 4/6 The signal favors the target but is inconsistent; inspect tied or lost scenarios and fix weak behavior.
msbuild-modernization gpt-5.6-luna ➖ Baseline signal, tie-limited n=6; 0W/3T/3L; d=3; p=0.125; net -50.0%; 1 dormancy excluded ✅ 0.20 — Evidence leans baseline but is not credible; inspect losses and make tied tasks discriminate.
system-text-json-net11 claude-sonnet-5 ✅ Improved n=6; 6W/0T/0L; d=6; p=0.016; net +100.0%; 2 dormancy excluded 🟡 0.32 — Review overfit evidence.
system-text-json-net11 gpt-5.6-luna ➖ Improvement signal, unproven n=6; 5W/0T/1L; d=6; p=0.109; net +66.7%; 2 dormancy excluded 🟡 0.36 — The signal favors the target but is inconsistent; inspect tied or lost scenarios and fix weak behavior.
vectorization claude-sonnet-5 ✅ Improved n=9; 6W/3T/0L; d=6; p=0.016; net +66.7%; 1 dormancy excluded 🟡 0.46 Activation: isolated 8/9; plugin 8/9 Fix activation gaps; Review overfit evidence.
vectorization gpt-5.6-luna ➖ Improvement signal, unproven n=9; 6W/2T/1L; d=7; p=0.063; net +55.6%; 1 dormancy excluded 🟡 0.48 Activation: isolated 8/9; plugin 5/9 The signal favors the target but is inconsistent; inspect tied or lost scenarios and fix weak behavior.
ℹ️ How to read this report
  • ✅ Improved — the result passed both the statistical gate and the 20% practical net-win floor.
  • ➖ No clear winner — the result is valid but did not pass both gates. The label distinguishes all ties, mixed evidence, directional but unproven evidence, and credible effects below the practical floor.
  • ⚠️ Invalid / underpowered — the gate withheld a quality verdict. Fix the measurement before judging the target.
  • ⛔ Activation contract failed — the isolated target activated on an explicit dormancy scenario. Dormancy preference is excluded, but this routing failure still blocks a pass.
  • 📉 Preference loss — the LLM judge credibly preferred baseline. It is report-only, not objective completion proof.
  • Gate evidence — n preference-eligible distinct-stimulus votes, W/T/L stimulus votes, d discordant votes, exact one-sided p, net win, and the count of separately retained dormancy stimuli. The p value applies to one model/target result; no matrix-wide multiple-comparison correction is applied.
  • Overfit — overfitting-judge severity (✅ Low, 🟡 Moderate, 🔴 High, — none) and score.
  • Warnings — activation, timeout, retry recovery, or unresolved comparison conditions that need attention.
  • Do not add repeated runs to increase statistical power. Do not add stimuli after seeing a near-pass unless the new breadth is predeclared for a new experiment.
🔻 Objective regression — agent.msbuild (claude-sonnet-5)

Why: Net win -20.0% (0W/4T/1L over 5 preference-eligible stimulus vote(s), sign test p=0.500), mean preference -8.0% across 5 paired run(s) — no improvement — native evaluator reported an objective task-completion regression

Next action: Inspect objective completion losses and fix them before merge.

State: VALID_REGRESSION (native_completion_regression)

Gate evidence: n=5; 0W/4T/1L; d=1; p=0.500; net -20.0%

Repeated-run reliability (not used by the gate): 5 paired runs (0W/4T/1L).

Weak or warning scenarios:

Scenario Preference gate Net win Δ Pref Runs (W/T/L)
= Advise on project file organization Eligible +0.0% +0.0% 0/1/0
= Diagnose broken incremental build behavior Eligible +0.0% +0.0% 0/1/0
= Review a project file for maintainability risks Eligible +0.0% +0.0% 0/1/0
▼ Route a slow build to performance analysis Eligible -100.0% -40.0% 0/0/1
= Triage a build failure and route to appropriate analysis Eligible +0.0% +0.0% 0/1/0

Illustrative judge evidence:

  • Advise on project file organization: Position-swap inconsistent (forward: baseline, reverse: skill). Defaulting to tie.

This is one example, not the aggregate verdict. Open Full Results for every judgment.

⚠️ Underpowered — configure-auth (claude-sonnet-5)

Why: Net win +0.0% (1W/0T/1L over 2 preference-eligible stimulus vote(s), sign test p=0.750), mean preference +0.0% across 2 paired run(s) — underpowered (2 preference-eligible stimulus vote(s); a credible verdict needs at least 5) — add distinct, discriminating stimuli; repeated runs do not increase task breadth

Next action: Predeclare more independent, discriminating stimuli; repeated runs do not add power.

State: INVALID_INCONCLUSIVE (underpowered)

Gate evidence: n=2; 1W/0T/1L; d=2; p=0.750; net +0.0%

Overfit: High (score 0.57)

Repeated-run reliability (not used by the gate): 2 paired runs (1W/0T/1L).

Weak or warning scenarios:

Scenario Preference gate Net win Δ Pref Runs (W/T/L)
▼ Login and account management in a globally interactive app Eligible -100.0% -40.0% 0/0/1

Illustrative judge evidence:

  • Login and account management in a globally interactive app: Both solutions satisfy the requested architecture and report successfully tested registration, login, logout, validation, and preserved global interactivity. A is marginally stronger because its final explanation is more precise and complete about why the routing exclusion wor...

This is one example, not the aggregate verdict. Open Full Results for every judgment.

⚠️ Underpowered — configure-auth (gpt-5.6-luna)

Why: Net win -100.0% (0W/0T/2L over 2 preference-eligible stimulus vote(s), sign test p=0.250), mean preference -70.0% across 2 paired run(s) — underpowered (2 preference-eligible stimulus vote(s); a credible verdict needs at least 5) — add distinct, discriminating stimuli; repeated runs do not increase task breadth

Next action: Predeclare more independent, discriminating stimuli; repeated runs do not add power.

State: INVALID_INCONCLUSIVE (underpowered)

Gate evidence: n=2; 0W/0T/2L; d=2; p=0.250; net -100.0%

Warnings: Activation: isolated 2/2; plugin 1/2

Overfit: Moderate (score 0.36)

Repeated-run reliability (not used by the gate): 2 paired runs (0W/0T/2L).

Weak or warning scenarios:

Scenario Preference gate Net win Δ Pref Runs (W/T/L)
▼ Login and account management in a globally interactive app Eligible -100.0% -40.0% 0/0/1
▼ Multi-tier app with WebAssembly auth Eligible -100.0% -100.0% 0/0/1

Illustrative judge evidence:

  • Login and account management in a globally interactive app: Response A provides a more complete and thoroughly tested implementation. While Response B correctly addresses the AcceptsInteractiveRouting() pattern, Response A uses the more standard [ExcludeFromInteractiveRouting] approach and—crucially—demonstrates full end-to-end functio...

This is one example, not the aggregate verdict. Open Full Results for every judgment.

➖ Improvement signal, tie-limited — agent.msbuild (gpt-5.6-luna)

Why: Net win +40.0% (3W/1T/1L over 5 preference-eligible stimulus vote(s), sign test p=0.312), mean preference +28.0% across 5 paired run(s) — not credible — 1 of 5 preference-eligible stimulus vote(s) tied, leaving only 4 discordant preference vote(s). The sign test conditions on non-tie stimulus votes and cannot reach 0.05 below 5, so no record could have passed here — this is not a measured null. Either the agent is inert on these scenarios (make them discriminate) or the eval needs more distinct stimuli to clear the ties

Next action: The signal favors the target, but ties leave too few discordant tasks; inspect ties and predeclare more discriminating breadth.

State: VALID_NO_CHANGE (no_credible_preference_change)

Gate evidence: n=5; 3W/1T/1L; d=4; p=0.312; net +40.0%

Repeated-run reliability (not used by the gate): 5 paired runs (3W/1T/1L).

Weak or warning scenarios:

Scenario Preference gate Net win Δ Pref Runs (W/T/L)
= Advise on project file organization Eligible +0.0% +0.0% 0/1/0
▼ Triage a build failure and route to appropriate analysis Eligible -100.0% -40.0% 0/0/1

Illustrative judge evidence:

  • Advise on project file organization: Position-swap inconsistent (forward: skill, reverse: baseline). Defaulting to tie.

This is one example, not the aggregate verdict. Open Full Results for every judgment.

➖ Improvement signal, unproven — msbuild-modernization (claude-sonnet-5)

Why: Net win +16.7% (3W/1T/2L over 6 preference-eligible stimulus vote(s), sign test p=0.500), mean preference +11.4% across 7 paired run(s), 1 dormancy stimulus/stimuli excluded from preference — not credible (sign test p=0.500 > 0.05)

Next action: The signal favors the target but is inconsistent; inspect tied or lost scenarios and fix weak behavior.

State: VALID_NO_CHANGE (no_credible_preference_change)

Gate evidence: n=6; 3W/1T/2L; d=5; p=0.500; net +16.7%; 1 dormancy excluded

Warnings: Activation: isolated 6/6; plugin 4/6

Overfit: Moderate (score 0.42)

Repeated-run reliability (not used by the gate): 7 paired runs (4W/1T/2L).

Weak or warning scenarios:

Scenario Preference gate Net win Δ Pref Runs (W/T/L)
▼ Consolidate duplicated projects into a multi-targeting SDK-style project Eligible -100.0% -40.0% 0/0/1
= Modernize a single project without introducing Central Package Management Eligible +0.0% +0.0% 0/1/0
▼ Modernize further without introducing a nondeterministic language version Eligible -100.0% -40.0% 0/0/1
▲ Recognize an already-modern project needs no migration Eligible +100.0% +40.0% 1/0/0

Illustrative judge evidence:

  • Consolidate duplicated projects into a multi-targeting SDK-style project: The resulting source changes are effectively identical and satisfy every rubric item, but A additionally ran and reported a successful build for both net472 and net48, providing meaningful verification that the conversion works.

This is one example, not the aggregate verdict. Open Full Results for every judgment.

➖ Baseline signal, tie-limited — msbuild-modernization (gpt-5.6-luna)

Why: Net win -50.0% (0W/3T/3L over 6 preference-eligible stimulus vote(s), sign test p=0.125), mean preference -17.1% across 7 paired run(s), 1 dormancy stimulus/stimuli excluded from preference — no improvement

Next action: Evidence leans baseline but is not credible; inspect losses and make tied tasks discriminate.

State: VALID_NO_CHANGE (no_credible_preference_change)

Gate evidence: n=6; 0W/3T/3L; d=3; p=0.125; net -50.0%; 1 dormancy excluded

Overfit: Low (score 0.20)

Repeated-run reliability (not used by the gate): 7 paired runs (0W/4T/3L).

Weak or warning scenarios:

Scenario Preference gate Net win Δ Pref Runs (W/T/L)
▼ Apply the migration to SDK-style Eligible -100.0% -40.0% 0/0/1
▼ Consolidate duplicated projects into a multi-targeting SDK-style project Eligible -100.0% -40.0% 0/0/1
= Decline modernizing a non-.NET build Excluded (activation contract) +0.0% +0.0% 0/1/0
= Identify legacy patterns for SDK-style migration Eligible +0.0% +0.0% 0/1/0
▼ Modernize a single project without introducing Central Package Management Eligible -100.0% -40.0% 0/0/1
= Modernize further without introducing a nondeterministic language version Eligible +0.0% +0.0% 0/1/0
= Recognize an already-modern project needs no migration Eligible +0.0% +0.0% 0/1/0

Illustrative judge evidence:

  • Apply the migration to SDK-style: Both responses successfully accomplish the core conversion task and meet all four rubric criteria. However, Response A uses the more standard and maintainable approach by permanently adding the Microsoft.NETFramework.ReferenceAssemblies package reference, which is the recommen...

This is one example, not the aggregate verdict. Open Full Results for every judgment.

➖ Improvement signal, unproven — system-text-json-net11 (gpt-5.6-luna)

Why: Net win +66.7% (5W/0T/1L over 6 preference-eligible stimulus vote(s), sign test p=0.109), mean preference +37.5% across 8 paired run(s), 2 dormancy stimulus/stimuli excluded from preference — not credible (sign test p=0.109 > 0.05)

Next action: The signal favors the target but is inconsistent; inspect tied or lost scenarios and fix weak behavior.

State: VALID_NO_CHANGE (no_credible_preference_change)

Gate evidence: n=6; 5W/0T/1L; d=6; p=0.109; net +66.7%; 2 dormancy excluded

Overfit: Moderate (score 0.36)

Repeated-run reliability (not used by the gate): 8 paired runs (5W/1T/2L).

Weak or warning scenarios:

Scenario Preference gate Net win Δ Pref Runs (W/T/L)
= Non-activation: Newtonsoft.Json PascalCase contract resolver Excluded (activation contract) +0.0% +0.0% 0/1/0
▼ Non-activation: camelCase JSON serialization on .NET 8 Excluded (activation contract) -100.0% -40.0% 0/0/1
▼ Type-safe JsonTypeInfo access without exceptions in .NET 11 Eligible -100.0% -40.0% 0/0/1

Illustrative judge evidence:

  • Type-safe JsonTypeInfo access without exceptions in .NET 11: Both responses successfully meet all four core rubric criteria. However, Response A better addresses the user's explicit request to "Show me a minimal .net11.0 program" by displaying the complete, runnable code. Response A also demonstrates a more realistic scenario with a c...

This is one example, not the aggregate verdict. Open Full Results for every judgment.

➖ Improvement signal, unproven — vectorization (gpt-5.6-luna)

Why: Net win +55.6% (6W/2T/1L over 9 preference-eligible stimulus vote(s), sign test p=0.063), mean preference +34.0% across 10 paired run(s), 1 dormancy stimulus/stimuli excluded from preference — not credible (sign test p=0.063 > 0.05)

Next action: The signal favors the target but is inconsistent; inspect tied or lost scenarios and fix weak behavior.

State: VALID_NO_CHANGE (no_credible_preference_change)

Gate evidence: n=9; 6W/2T/1L; d=7; p=0.063; net +55.6%; 1 dormancy excluded

Warnings: Activation: isolated 8/9; plugin 5/9

Overfit: Moderate (score 0.48)

Repeated-run reliability (not used by the gate): 10 paired runs (6W/2T/2L).

Weak or warning scenarios:

Scenario Preference gate Net win Δ Pref Runs (W/T/L)
= Detect architecture-specific behavior change Eligible +0.0% +0.0% 0/1/0
▼ Detect empty-span reference access Eligible -100.0% -40.0% 0/0/1
= Detect unsigned tail offset underflow Eligible +0.0% +0.0% 0/1/0
▲ Detect unsupported vector element type Eligible +100.0% +100.0% 1/0/0
▼ Ignore unrelated parser performance request Excluded (activation contract) -100.0% -40.0% 0/0/1
▲ Reject fragile backwards reference Eligible +100.0% +100.0% 1/0/0

Illustrative judge evidence:

  • Detect architecture-specific behavior change: Position-swap inconsistent (forward: B, reverse: tie). Defaulting to tie.

This is one example, not the aggregate verdict. Open Full Results for every judgment.

✅ Improved — system-text-json-net11 (claude-sonnet-5)

Why: Net win +100.0% (6W/0T/0L over 6 preference-eligible stimulus vote(s), sign test p=0.016), mean preference +62.5% across 8 paired run(s), 2 dormancy stimulus/stimuli excluded from preference — credibly better

Next action: Review overfit evidence.

State: VALID_PASS (credible_preference_improvement)

Gate evidence: n=6; 6W/0T/0L; d=6; p=0.016; net +100.0%; 2 dormancy excluded

Overfit: Moderate (score 0.32)

Repeated-run reliability (not used by the gate): 8 paired runs (6W/1T/1L).

Weak or warning scenarios:

Scenario Preference gate Net win Δ Pref Runs (W/T/L)
▼ Non-activation: Newtonsoft.Json PascalCase contract resolver Excluded (activation contract) -100.0% -40.0% 0/0/1
= Non-activation: camelCase JSON serialization on .NET 8 Excluded (activation contract) +0.0% +0.0% 0/1/0

Illustrative judge evidence:

  • Non-activation: Newtonsoft.Json PascalCase contract resolver: Both give a correct core Newtonsoft.Json answer, but A is more direct and robust: explicitly clearing NamingStrategy accurately preserves declared PascalCase names. B's extra custom resolver is unnecessary and does not generally transform arbitrary names into proper PascalCase.

This is one example, not the aggregate verdict. Open Full Results for every judgment.

✅ Improved — vectorization (claude-sonnet-5)

Why: Net win +66.7% (6W/3T/0L over 9 preference-eligible stimulus vote(s), sign test p=0.016), mean preference +52.0% across 10 paired run(s), 1 dormancy stimulus/stimuli excluded from preference — credibly better

Next action: Fix activation gaps; Review overfit evidence.

State: VALID_PASS (credible_preference_improvement)

Gate evidence: n=9; 6W/3T/0L; d=6; p=0.016; net +66.7%; 1 dormancy excluded

Warnings: Activation: isolated 8/9; plugin 8/9

Overfit: Moderate (score 0.46)

Repeated-run reliability (not used by the gate): 10 paired runs (7W/3T/0L).

Weak or warning scenarios:

Scenario Preference gate Net win Δ Pref Runs (W/T/L)
= Detect empty-span reference access Eligible +0.0% +0.0% 0/1/0
= Preserve product empty-input contract Eligible +0.0% +0.0% 0/1/0
= Repair vectorized reduction safely Eligible +0.0% +0.0% 0/1/0

Illustrative judge evidence:

  • Detect empty-span reference access: Position-swap inconsistent (forward: tie, reverse: A). Defaulting to tie.

This is one example, not the aggregate verdict. Open Full Results for every judgment.

🔍 Full Results - all metrics and investigation details

To investigate non-passing or warning results, paste this to your AI coding agent:

For PR 1260 in dotnet/skills, download eval artifacts with gh run download 37302010017 --repo dotnet/skills --pattern "vally-results-*" --dir ./eval-results, then fetch https://raw.githubusercontent.com/dotnet/skills/1603d1b6467800a64426107881b42de6ccc69979/eng/vally-adapter/InvestigatingResults.md and follow it. Classify each result as measurement-invalid, underpowered, not-proven, preference-loss, or passing-with-warning. Use stateReason, result accounting, weak scenarios, and judge evidence to give the cause and exact next fix.

⚠️ Session replay telemetry was not published because the auxiliary dotnet/skills-data publisher failed. The evaluation verdicts above remain authoritative; maintainers must repair SKILLS_DATA_TOKEN or the publisher before replay links are available.

github-actions Bot added a commit that referenced this pull request Oct 5, 2026
@github-actions github-actions Bot added waiting-on-review PR state label and removed pr-state/ready-for-eval PR is mergeable and awaiting evaluation labels Oct 5, 2026
@github-actions

github-actions Bot commented Oct 5, 2026

Copy link
Copy Markdown
Contributor

✅ Evaluation passed for 1603d1b. cc @webreidi @AbhitejJohn @JanKrivanek @jeffschw @artl93 @dotnet/aspnet @dotnet/msbuild @YuliiaKovalova @dotnet/skills-csharp-language-reviewers — please review.

@github-actions github-actions Bot added ready-to-merge PR state label and removed waiting-on-review PR state label labels Oct 5, 2026
@github-actions

github-actions Bot commented Oct 5, 2026

Copy link
Copy Markdown
Contributor

✅ Approved by @AbhitejJohn. cc @dotnet/skills-merge-approvers — ready to merge.

@AbhitejJohn
AbhitejJohn merged commit 6a47c0f into main Oct 5, 2026
115 of 117 checks passed
@AbhitejJohn
AbhitejJohn deleted the ykovalova-fix-skill-description-xml-tags branch October 5, 2026 18:04
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

ready-to-merge PR state label

Projects

None yet

Development

Successfully merging this pull request may close these issues.

Claude Code marketplace integration reports issues

3 participants