Skip to content

Add lightweight telemetry skill for .NET 11 console tools - #1237

Draft
AbhitejJohn wants to merge 1 commit into
mainfrom
abhitejjohn-telemetry-trusted-draft
Draft

AbhitejJohn wants to merge 1 commit into
mainfrom
abhitejjohn-telemetry-trusted-draft

Conversation

@AbhitejJohn

@AbhitejJohn AbhitejJohn commented Sep 30, 2026 •

Copy link
Copy Markdown
Collaborator

Summary

Add a dependency-free .NET 11 telemetry skill for console tools, with guidance for counters, gauges, histograms, bounded tags, and JSON-line output. Add evaluation cases and update the dotnet11 catalog. This imports the reviewed work by qapdex-maker from #1036; it leaves unrelated README website links out.

Why

Small tools need a way to report measurements without an OpenTelemetry SDK, APM package, or collector. The skill shows how to use the built-in System.Diagnostics.Metrics APIs and receive measurements before a short-lived process exits.

Impact

The new skill and evaluation cover .NET 11 console tools. This change adds no runtime dependency and does not change CI configuration.

Validation

The coordinating Windows worktree built and ran the complete inline Program.cs with the pinned .NET 11 SDK: zero warnings, zero errors, and three JSON lines, each with all seven required fields. The dotnet11 skill-validator check and the repository eval quality gate passed. git diff --check passed, and the branch differs from current main only in the five reviewed files.

Import the reviewed five-file patch by @qapdex-maker from #1036 onto current main without unrelated README links.

Co-authored-by: Copilot App <223556219+Copilot@users.noreply.github.com>
@github-actions

Copy link
Copy Markdown
Contributor

Skill Coverage Report

Plugin Skill Covered Coverage
✅ dotnet11 system-text-json-net11 15/16 93.8%
Uncovered: dotnet11/system-text-json-net11
  • [CodePattern] [guid] (line 49)

@github-actions

Copy link
Copy Markdown
Contributor

📊 Skill and Agent Evaluation Results

2 model/target results across 1 target and 2 models — ✅ 0 improved, ➖ 2 not proven improved, ⚠️ 0 invalid or underpowered, ⛔ 0 activation contract failures, 📉 0 preference losses (report only).

Measurement identity: evaluated commit 7646be66f084babc64561d1d195a9aacc9c4bb4b; 2 judge models.

Measurement health: 2 expected / 2 observed / 2 written; 0 missing, 0 unexpected, 0 invalid; 0 recovered comparison error slots and 0 unresolved comparison error slots.

Objective completion gate: not enabled. Aggregate completion transitions are telemetry only, so this report does not claim that zero objective regressions were proven.

A result passes only when preference-eligible distinct-stimulus votes have aggregate net win of at least 20% and an exact one-sided sign-test result of p ≤ 0.05, and every explicit dormancy activation contract passes. Repeated runs measure reliability only.

Target Model Verdict Gate evidence Overfit Warnings Next action
lightweight-telemetry claude-sonnet-5 ➖ Not proven improved n=10; 6W/1T/3L; d=9; p=0.254; net +30.0%; 3 dormancy excluded 🟡 0.25 Activation-only stop: isolated 1 failed run; Activation-only stop: plugin 1 failed run Inspect activation-only failed runs before rewriting skill content; the model stopped after loading a skill.
lightweight-telemetry gpt-5.6-luna ➖ Not proven improved n=10; 4W/3T/3L; d=7; p=0.500; net +10.0%; 3 dormancy excluded 🟡 0.26 Activation-only stop: isolated 1 failed run; Activation-only stop: plugin 1 failed run Inspect activation-only failed runs before rewriting skill content; the model stopped after loading a skill.
ℹ️ How to read this report
  • ✅ Improved — the result passed both the statistical gate and the 20% practical net-win floor.
  • ➖ Not proven improved — the result is valid but did not pass both gates. This is not automatically a regression.
  • ⚠️ Invalid / underpowered — the gate withheld a quality verdict. Fix the measurement before judging the target.
  • ⛔ Activation contract failed — the isolated target activated on an explicit dormancy scenario. Dormancy preference is excluded, but this routing failure still blocks a pass.
  • 📉 Preference loss — the LLM judge credibly preferred baseline. It is report-only, not objective completion proof.
  • Gate evidence — n preference-eligible distinct-stimulus votes, W/T/L stimulus votes, d discordant votes, exact one-sided p, net win, and the count of separately retained dormancy stimuli. The p value applies to one model/target result; no matrix-wide multiple-comparison correction is applied.
  • Overfit — overfitting-judge severity (✅ Low, 🟡 Moderate, 🔴 High, — none) and score.
  • Warnings — activation, timeout, retry recovery, or unresolved comparison conditions that need attention.
  • Do not add repeated runs to increase statistical power. Do not add stimuli after seeing a near-pass unless the new breadth is predeclared for a new experiment.
➖ Not proven improved — lightweight-telemetry (claude-sonnet-5)

Why: Net win +30.0% (6W/1T/3L over 10 preference-eligible stimulus vote(s), sign test p=0.254), mean preference +4.6% across 13 paired run(s), 3 dormancy stimulus/stimuli excluded from preference — not credible (sign test p=0.254 > 0.05)

Next action: Inspect activation-only failed runs before rewriting skill content; the model stopped after loading a skill.

State: VALID_NO_CHANGE (no_credible_preference_change)

Gate evidence: n=10; 6W/1T/3L; d=9; p=0.254; net +30.0%; 3 dormancy excluded

Warnings: Activation-only stop: isolated 1 failed run; Activation-only stop: plugin 1 failed run

Overfit: Moderate (score 0.25)

Repeated-run reliability (not used by the gate): 13 paired runs (6W/1T/6L).

Weak or warning scenarios:

Scenario Preference gate Net win Δ Pref Runs (W/T/L)
▼ Cloud ingestion request stays out of scope Excluded (activation contract) -100.0% -40.0% 0/0/1
▼ Cross-service tracing request stays out of scope Excluded (activation contract) -100.0% -40.0% 0/0/1
= Elapsed time measured through the framework clock abstraction Eligible +0.0% +0.0% 0/1/0
▼ Keep the measurement path cheap when nothing is listening Eligible -100.0% -40.0% 0/0/1
▼ Log shipping request stays out of scope Excluded (activation contract) -100.0% -40.0% 0/0/1
▼ Metadata that makes readings interpretable downstream Eligible -100.0% -40.0% 0/0/1
▼ Observable snapshot is collected before the tool exits Eligible -100.0% -40.0% 0/0/1

Illustrative judge evidence:

  • Elapsed time measured through the framework clock abstraction: Position-swap inconsistent (forward: A, reverse: B). Defaulting to tie.

This is one example, not the aggregate verdict. Open Full Results for every judgment.

➖ Not proven improved — lightweight-telemetry (gpt-5.6-luna)

Why: Net win +10.0% (4W/3T/3L over 10 preference-eligible stimulus vote(s), sign test p=0.500), mean preference +0.0% across 13 paired run(s), 3 dormancy stimulus/stimuli excluded from preference — not credible (sign test p=0.500 > 0.05)

Next action: Inspect activation-only failed runs before rewriting skill content; the model stopped after loading a skill.

State: VALID_NO_CHANGE (no_credible_preference_change)

Gate evidence: n=10; 4W/3T/3L; d=7; p=0.500; net +10.0%; 3 dormancy excluded

Warnings: Activation-only stop: isolated 1 failed run; Activation-only stop: plugin 1 failed run

Overfit: Moderate (score 0.26)

Repeated-run reliability (not used by the gate): 13 paired runs (4W/5T/4L).

Weak or warning scenarios:

Scenario Preference gate Net win Δ Pref Runs (W/T/L)
▼ Cloud ingestion request stays out of scope Excluded (activation contract) -100.0% -40.0% 0/0/1
= Cross-service tracing request stays out of scope Excluded (activation contract) +0.0% +0.0% 0/1/0
= Elapsed time measured through the framework clock abstraction Eligible +0.0% +0.0% 0/1/0
▼ Instrument choice for a monotonic total Eligible -100.0% -40.0% 0/0/1
= Instrument choice for an instantaneous value Eligible +0.0% +0.0% 0/1/0
▼ Keep the measurement path cheap when nothing is listening Eligible -100.0% -40.0% 0/0/1
= Log shipping request stays out of scope Excluded (activation contract) +0.0% +0.0% 0/1/0
= Metadata that makes readings interpretable downstream Eligible +0.0% +0.0% 0/1/0
▼ Observable snapshot is collected before the tool exits Eligible -100.0% -40.0% 0/0/1

Illustrative judge evidence:

  • Elapsed time measured through the framework clock abstraction: Position-swap inconsistent (forward: A, reverse: B). Defaulting to tie.

This is one example, not the aggregate verdict. Open Full Results for every judgment.

🔍 Full Results - all metrics and investigation details

To investigate non-passing or warning results, paste this to your AI coding agent:

For PR 1237 in dotnet/skills, download eval artifacts with gh run download 36676062044 --repo dotnet/skills --pattern "vally-results-*" --dir ./eval-results, then fetch https://raw.githubusercontent.com/dotnet/skills/7646be66f084babc64561d1d195a9aacc9c4bb4b/eng/vally-adapter/InvestigatingResults.md and follow it. Classify each result as measurement-invalid, underpowered, not-proven, preference-loss, or passing-with-warning. Use stateReason, result accounting, weak scenarios, and judge evidence to give the cause and exact next fix.

⚠️ Session replay telemetry was not published because the auxiliary dotnet/skills-data publisher failed. The evaluation verdicts above remain authoritative; maintainers must repair SKILLS_DATA_TOKEN or the publisher before replay links are available.

This branch has not been deployed

No deployments
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant