Remove XML tags from skill descriptions for Claude compatibility - #1260
Conversation
Claude (claude.ai and Claude Code marketplace sync) rejects skill descriptions that contain XML tags and strips the angle brackets. Reword the four affected descriptions in plain language: - dotnet-advanced/vectorization - dotnet-blazor/configure-auth - dotnet-msbuild/msbuild-modernization - dotnet11/system-text-json-net11 Add a skill-validator check that errors on XML-like tags in a skill description so this cannot regress, with tests, and document the rule in the create-skill authoring guide. Fixes #1251 Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com> Copilot-Session: bcd44b69-8a83-470f-a43b-41606d2b67de
…arisons Address review feedback: - Accept '_' as an XML name start so tags like <_Root /> are flagged. - Match a tag name followed only by bare words or name=value attributes (plus comments, PIs and CDATA) instead of arbitrary text between brackets, so unspaced comparisons such as i<length && count>0 no longer fail the check. Add regression cases for both. Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com> Copilot-Session: bcd44b69-8a83-470f-a43b-41606d2b67de
Skill Coverage Report
Uncovered:
|
There was a problem hiding this comment.
Copilot review overview
🟢 Approval recommended
The focused implementation matches the stated compatibility requirement and includes appropriate regression coverage.
Review effort: Balanced
Findings: None
What changed in this PR
Removes XML-like syntax from skill descriptions and prevents future Claude marketplace compatibility regressions.
Changes:
- Reworded four affected skill descriptions without changing routing intent.
- Added XML-like tag validation with positive and negative tests.
- Documented the restriction in the skill-authoring guide.
| File | Description |
|---|---|
plugins/dotnet11/skills/system-text-json-net11/SKILL.md |
Rewords generic API references. |
plugins/dotnet-msbuild/skills/msbuild-modernization/SKILL.md |
Rewords MSBuild XML examples. |
plugins/dotnet-blazor/skills/configure-auth/SKILL.md |
Removes tag syntax from NotAuthorized. |
plugins/dotnet-advanced/skills/vectorization/SKILL.md |
Rewords the generic Vector reference. |
eng/skill-validator/tests/Check/SkillProfileTests.cs |
Tests tag detection and comparison exclusions. |
eng/skill-validator/src/Check/SkillProfiler.cs |
Adds description tag validation. |
.agents/skills/create-skill/SKILL.md |
Documents the new authoring rule. |
💡 Add a code-review agent skill or configure MCP servers for context-aware, tailored reviews. Learn more in the docs.
|
❌ Evaluation did not complete successfully (the evaluate job reported 10 partial result file(s) were preserved for diagnosis but were not consolidated because the full matrix did not complete. |
|
/evaluate 1603d1b |
📊 Skill and Agent Evaluation Results10 model/target results across 5 targets and 2 models — ✅ 2 improved, ➖ 5 results without a clear winner, Measurement identity: evaluated commit Measurement health: 10 expected / 10 observed / 10 written; 0 missing, 0 unexpected, 2 invalid; 0 recovered comparison error slots and 0 unresolved comparison error slots. Objective completion gate: enabled; 1 result shows a proven objective regression. A result passes only when preference-eligible distinct-stimulus votes have aggregate net win of at least 20% and an exact one-sided sign-test result of
ℹ️ How to read this report
🔻 Objective regression — agent.msbuild (claude-sonnet-5)Why: Net win -20.0% (0W/4T/1L over 5 preference-eligible stimulus vote(s), sign test p=0.500), mean preference -8.0% across 5 paired run(s) — no improvement — native evaluator reported an objective task-completion regression Next action: Inspect objective completion losses and fix them before merge. State: Gate evidence: n=5; 0W/4T/1L; d=1; p=0.500; net -20.0% Repeated-run reliability (not used by the gate): 5 paired runs (0W/4T/1L). Weak or warning scenarios:
Illustrative judge evidence:
This is one example, not the aggregate verdict. Open Full Results for every judgment.
|
| Scenario | Preference gate | Net win | Δ Pref | Runs (W/T/L) |
|---|---|---|---|---|
| ▼ Login and account management in a globally interactive app | Eligible | -100.0% | -40.0% | 0/0/1 |
Illustrative judge evidence:
Login and account management in a globally interactive app:Both solutions satisfy the requested architecture and report successfully tested registration, login, logout, validation, and preserved global interactivity. A is marginally stronger because its final explanation is more precise and complete about why the routing exclusion wor...
This is one example, not the aggregate verdict. Open Full Results for every judgment.
⚠️ Underpowered — configure-auth (gpt-5.6-luna)
Why: Net win -100.0% (0W/0T/2L over 2 preference-eligible stimulus vote(s), sign test p=0.250), mean preference -70.0% across 2 paired run(s) — underpowered (2 preference-eligible stimulus vote(s); a credible verdict needs at least 5) — add distinct, discriminating stimuli; repeated runs do not increase task breadth
Next action: Predeclare more independent, discriminating stimuli; repeated runs do not add power.
State: INVALID_INCONCLUSIVE (underpowered)
Gate evidence: n=2; 0W/0T/2L; d=2; p=0.250; net -100.0%
Warnings: Activation: isolated 2/2; plugin 1/2
Overfit: Moderate (score 0.36)
Repeated-run reliability (not used by the gate): 2 paired runs (0W/0T/2L).
Weak or warning scenarios:
| Scenario | Preference gate | Net win | Δ Pref | Runs (W/T/L) |
|---|---|---|---|---|
| ▼ Login and account management in a globally interactive app | Eligible | -100.0% | -40.0% | 0/0/1 |
| ▼ Multi-tier app with WebAssembly auth | Eligible | -100.0% | -100.0% | 0/0/1 |
Illustrative judge evidence:
Login and account management in a globally interactive app:Response A provides a more complete and thoroughly tested implementation. While Response B correctly addresses the AcceptsInteractiveRouting() pattern, Response A uses the more standard [ExcludeFromInteractiveRouting] approach and—crucially—demonstrates full end-to-end functio...
This is one example, not the aggregate verdict. Open Full Results for every judgment.
➖ Improvement signal, tie-limited — agent.msbuild (gpt-5.6-luna)
Why: Net win +40.0% (3W/1T/1L over 5 preference-eligible stimulus vote(s), sign test p=0.312), mean preference +28.0% across 5 paired run(s) — not credible — 1 of 5 preference-eligible stimulus vote(s) tied, leaving only 4 discordant preference vote(s). The sign test conditions on non-tie stimulus votes and cannot reach 0.05 below 5, so no record could have passed here — this is not a measured null. Either the agent is inert on these scenarios (make them discriminate) or the eval needs more distinct stimuli to clear the ties
Next action: The signal favors the target, but ties leave too few discordant tasks; inspect ties and predeclare more discriminating breadth.
State: VALID_NO_CHANGE (no_credible_preference_change)
Gate evidence: n=5; 3W/1T/1L; d=4; p=0.312; net +40.0%
Repeated-run reliability (not used by the gate): 5 paired runs (3W/1T/1L).
Weak or warning scenarios:
| Scenario | Preference gate | Net win | Δ Pref | Runs (W/T/L) |
|---|---|---|---|---|
| = Advise on project file organization | Eligible | +0.0% | +0.0% | 0/1/0 |
| ▼ Triage a build failure and route to appropriate analysis | Eligible | -100.0% | -40.0% | 0/0/1 |
Illustrative judge evidence:
Advise on project file organization:Position-swap inconsistent (forward: skill, reverse: baseline). Defaulting to tie.
This is one example, not the aggregate verdict. Open Full Results for every judgment.
➖ Improvement signal, unproven — msbuild-modernization (claude-sonnet-5)
Why: Net win +16.7% (3W/1T/2L over 6 preference-eligible stimulus vote(s), sign test p=0.500), mean preference +11.4% across 7 paired run(s), 1 dormancy stimulus/stimuli excluded from preference — not credible (sign test p=0.500 > 0.05)
Next action: The signal favors the target but is inconsistent; inspect tied or lost scenarios and fix weak behavior.
State: VALID_NO_CHANGE (no_credible_preference_change)
Gate evidence: n=6; 3W/1T/2L; d=5; p=0.500; net +16.7%; 1 dormancy excluded
Warnings: Activation: isolated 6/6; plugin 4/6
Overfit: Moderate (score 0.42)
Repeated-run reliability (not used by the gate): 7 paired runs (4W/1T/2L).
Weak or warning scenarios:
| Scenario | Preference gate | Net win | Δ Pref | Runs (W/T/L) |
|---|---|---|---|---|
| ▼ Consolidate duplicated projects into a multi-targeting SDK-style project | Eligible | -100.0% | -40.0% | 0/0/1 |
| = Modernize a single project without introducing Central Package Management | Eligible | +0.0% | +0.0% | 0/1/0 |
| ▼ Modernize further without introducing a nondeterministic language version | Eligible | -100.0% | -40.0% | 0/0/1 |
| ▲ Recognize an already-modern project needs no migration | Eligible | +100.0% | +40.0% | 1/0/0 |
Illustrative judge evidence:
Consolidate duplicated projects into a multi-targeting SDK-style project:The resulting source changes are effectively identical and satisfy every rubric item, but A additionally ran and reported a successful build for both net472 and net48, providing meaningful verification that the conversion works.
This is one example, not the aggregate verdict. Open Full Results for every judgment.
➖ Baseline signal, tie-limited — msbuild-modernization (gpt-5.6-luna)
Why: Net win -50.0% (0W/3T/3L over 6 preference-eligible stimulus vote(s), sign test p=0.125), mean preference -17.1% across 7 paired run(s), 1 dormancy stimulus/stimuli excluded from preference — no improvement
Next action: Evidence leans baseline but is not credible; inspect losses and make tied tasks discriminate.
State: VALID_NO_CHANGE (no_credible_preference_change)
Gate evidence: n=6; 0W/3T/3L; d=3; p=0.125; net -50.0%; 1 dormancy excluded
Overfit: Low (score 0.20)
Repeated-run reliability (not used by the gate): 7 paired runs (0W/4T/3L).
Weak or warning scenarios:
| Scenario | Preference gate | Net win | Δ Pref | Runs (W/T/L) |
|---|---|---|---|---|
| ▼ Apply the migration to SDK-style | Eligible | -100.0% | -40.0% | 0/0/1 |
| ▼ Consolidate duplicated projects into a multi-targeting SDK-style project | Eligible | -100.0% | -40.0% | 0/0/1 |
| = Decline modernizing a non-.NET build | Excluded (activation contract) | +0.0% | +0.0% | 0/1/0 |
| = Identify legacy patterns for SDK-style migration | Eligible | +0.0% | +0.0% | 0/1/0 |
| ▼ Modernize a single project without introducing Central Package Management | Eligible | -100.0% | -40.0% | 0/0/1 |
| = Modernize further without introducing a nondeterministic language version | Eligible | +0.0% | +0.0% | 0/1/0 |
| = Recognize an already-modern project needs no migration | Eligible | +0.0% | +0.0% | 0/1/0 |
Illustrative judge evidence:
Apply the migration to SDK-style:Both responses successfully accomplish the core conversion task and meet all four rubric criteria. However, Response A uses the more standard and maintainable approach by permanently adding the Microsoft.NETFramework.ReferenceAssemblies package reference, which is the recommen...
This is one example, not the aggregate verdict. Open Full Results for every judgment.
➖ Improvement signal, unproven — system-text-json-net11 (gpt-5.6-luna)
Why: Net win +66.7% (5W/0T/1L over 6 preference-eligible stimulus vote(s), sign test p=0.109), mean preference +37.5% across 8 paired run(s), 2 dormancy stimulus/stimuli excluded from preference — not credible (sign test p=0.109 > 0.05)
Next action: The signal favors the target but is inconsistent; inspect tied or lost scenarios and fix weak behavior.
State: VALID_NO_CHANGE (no_credible_preference_change)
Gate evidence: n=6; 5W/0T/1L; d=6; p=0.109; net +66.7%; 2 dormancy excluded
Overfit: Moderate (score 0.36)
Repeated-run reliability (not used by the gate): 8 paired runs (5W/1T/2L).
Weak or warning scenarios:
| Scenario | Preference gate | Net win | Δ Pref | Runs (W/T/L) |
|---|---|---|---|---|
| = Non-activation: Newtonsoft.Json PascalCase contract resolver | Excluded (activation contract) | +0.0% | +0.0% | 0/1/0 |
| ▼ Non-activation: camelCase JSON serialization on .NET 8 | Excluded (activation contract) | -100.0% | -40.0% | 0/0/1 |
| ▼ Type-safe JsonTypeInfo access without exceptions in .NET 11 | Eligible | -100.0% | -40.0% | 0/0/1 |
Illustrative judge evidence:
Type-safe JsonTypeInfo access without exceptions in .NET 11:Both responses successfully meet all four core rubric criteria. However, Response A better addresses the user's explicit request to "Show me a minimal.net11.0program" by displaying the complete, runnable code. Response A also demonstrates a more realistic scenario with a c...
This is one example, not the aggregate verdict. Open Full Results for every judgment.
➖ Improvement signal, unproven — vectorization (gpt-5.6-luna)
Why: Net win +55.6% (6W/2T/1L over 9 preference-eligible stimulus vote(s), sign test p=0.063), mean preference +34.0% across 10 paired run(s), 1 dormancy stimulus/stimuli excluded from preference — not credible (sign test p=0.063 > 0.05)
Next action: The signal favors the target but is inconsistent; inspect tied or lost scenarios and fix weak behavior.
State: VALID_NO_CHANGE (no_credible_preference_change)
Gate evidence: n=9; 6W/2T/1L; d=7; p=0.063; net +55.6%; 1 dormancy excluded
Warnings: Activation: isolated 8/9; plugin 5/9
Overfit: Moderate (score 0.48)
Repeated-run reliability (not used by the gate): 10 paired runs (6W/2T/2L).
Weak or warning scenarios:
| Scenario | Preference gate | Net win | Δ Pref | Runs (W/T/L) |
|---|---|---|---|---|
| = Detect architecture-specific behavior change | Eligible | +0.0% | +0.0% | 0/1/0 |
| ▼ Detect empty-span reference access | Eligible | -100.0% | -40.0% | 0/0/1 |
| = Detect unsigned tail offset underflow | Eligible | +0.0% | +0.0% | 0/1/0 |
| ▲ Detect unsupported vector element type | Eligible | +100.0% | +100.0% | 1/0/0 |
| ▼ Ignore unrelated parser performance request | Excluded (activation contract) | -100.0% | -40.0% | 0/0/1 |
| ▲ Reject fragile backwards reference | Eligible | +100.0% | +100.0% | 1/0/0 |
Illustrative judge evidence:
Detect architecture-specific behavior change:Position-swap inconsistent (forward: B, reverse: tie). Defaulting to tie.
This is one example, not the aggregate verdict. Open Full Results for every judgment.
✅ Improved — system-text-json-net11 (claude-sonnet-5)
Why: Net win +100.0% (6W/0T/0L over 6 preference-eligible stimulus vote(s), sign test p=0.016), mean preference +62.5% across 8 paired run(s), 2 dormancy stimulus/stimuli excluded from preference — credibly better
Next action: Review overfit evidence.
State: VALID_PASS (credible_preference_improvement)
Gate evidence: n=6; 6W/0T/0L; d=6; p=0.016; net +100.0%; 2 dormancy excluded
Overfit: Moderate (score 0.32)
Repeated-run reliability (not used by the gate): 8 paired runs (6W/1T/1L).
Weak or warning scenarios:
| Scenario | Preference gate | Net win | Δ Pref | Runs (W/T/L) |
|---|---|---|---|---|
| ▼ Non-activation: Newtonsoft.Json PascalCase contract resolver | Excluded (activation contract) | -100.0% | -40.0% | 0/0/1 |
| = Non-activation: camelCase JSON serialization on .NET 8 | Excluded (activation contract) | +0.0% | +0.0% | 0/1/0 |
Illustrative judge evidence:
Non-activation: Newtonsoft.Json PascalCase contract resolver:Both give a correct core Newtonsoft.Json answer, but A is more direct and robust: explicitly clearing NamingStrategy accurately preserves declared PascalCase names. B's extra custom resolver is unnecessary and does not generally transform arbitrary names into proper PascalCase.
This is one example, not the aggregate verdict. Open Full Results for every judgment.
✅ Improved — vectorization (claude-sonnet-5)
Why: Net win +66.7% (6W/3T/0L over 9 preference-eligible stimulus vote(s), sign test p=0.016), mean preference +52.0% across 10 paired run(s), 1 dormancy stimulus/stimuli excluded from preference — credibly better
Next action: Fix activation gaps; Review overfit evidence.
State: VALID_PASS (credible_preference_improvement)
Gate evidence: n=9; 6W/3T/0L; d=6; p=0.016; net +66.7%; 1 dormancy excluded
Warnings: Activation: isolated 8/9; plugin 8/9
Overfit: Moderate (score 0.46)
Repeated-run reliability (not used by the gate): 10 paired runs (7W/3T/0L).
Weak or warning scenarios:
| Scenario | Preference gate | Net win | Δ Pref | Runs (W/T/L) |
|---|---|---|---|---|
| = Detect empty-span reference access | Eligible | +0.0% | +0.0% | 0/1/0 |
| = Preserve product empty-input contract | Eligible | +0.0% | +0.0% | 0/1/0 |
| = Repair vectorized reduction safely | Eligible | +0.0% | +0.0% | 0/1/0 |
Illustrative judge evidence:
Detect empty-span reference access:Position-swap inconsistent (forward: tie, reverse: A). Defaulting to tie.
This is one example, not the aggregate verdict. Open Full Results for every judgment.
🔍 Full Results - all metrics and investigation details
To investigate non-passing or warning results, paste this to your AI coding agent:
For PR 1260 in dotnet/skills, download eval artifacts with
gh run download 37302010017 --repo dotnet/skills --pattern "vally-results-*" --dir ./eval-results, then fetch https://raw.githubusercontent.com/dotnet/skills/1603d1b6467800a64426107881b42de6ccc69979/eng/vally-adapter/InvestigatingResults.md and follow it. Classify each result as measurement-invalid, underpowered, not-proven, preference-loss, or passing-with-warning. Use stateReason, result accounting, weak scenarios, and judge evidence to give the cause and exact next fix.
|
✅ Evaluation passed for |
|
✅ Approved by @AbhitejJohn. cc @dotnet/skills-merge-approvers — ready to merge. |
Summary
Claude (claude.ai and Claude Code marketplace sync) rejects skill descriptions that contain XML tags and stores them with the angle brackets stripped. Four skills were reported after a marketplace sync.
dotnet-advanced/vectorization:Vector<T>-> "the generic Vector type"dotnet-blazor/configure-auth:<NotAuthorized>-> "NotAuthorized"dotnet-msbuild/msbuild-modernization:<Compile Include>/<Import Project=...>reworded;> 50 lines-> "over 50 lines"dotnet11/system-text-json-net11:GetTypeInfo<T>()/TryGetTypeInfo<T>(out ...)-> genericGetTypeInfo/TryGetTypeInfooverloadsskill-validator checkrule that errors on XML-like tags in a skill description, so this cannot regress. The matcher requires a tag name followed only by bare words orname=valueattributes (plus comments, processing instructions and CDATA), so comparisons such as>5s,a < bandi<length && count>0are not flagged. Names may start with_.create-skillauthoring guide (description rules + checklist).Descriptions that use
>/<only as comparison operators (build-perf-diagnostics,eval-performance) are intentionally unchanged; they are not tags and were not reported.Supersedes #1259, which came from a fork and so could not run the
eng/CI with the right permissions. Review feedback from that PR is already incorporated.Related issue
Fixes #1251
Validation
dotnet testineng/skill-validator: 878 passed, 0 failed (includes 14 cases for the XML-tag rule).skill-validator check --plugin plugins/*with the repo's allow-lists: all checks passed (100 skills, 16 agents, 17 plugins).SKILL.mddescription (parsed YAML) for tag-like text: 0 remaining; all descriptions are still under 1,024 characters.Checklist
eng/known-domains.txtfor any new external domains referenced by skill content. (N/A: no new domains)