Add lightweight telemetry skill for .NET 11 console tools - #1237
AbhitejJohn wants to merge 1 commit into
Conversation
Import the reviewed five-file patch by @qapdex-maker from #1036 onto current main without unrelated README links. Co-authored-by: Copilot App <223556219+Copilot@users.noreply.github.com>
Skill Coverage Report
Uncovered:
|
📊 Skill and Agent Evaluation Results2 model/target results across 1 target and 2 models — ✅ 0 improved, ➖ 2 not proven improved, Measurement identity: evaluated commit Measurement health: 2 expected / 2 observed / 2 written; 0 missing, 0 unexpected, 0 invalid; 0 recovered comparison error slots and 0 unresolved comparison error slots. Objective completion gate: not enabled. Aggregate completion transitions are telemetry only, so this report does not claim that zero objective regressions were proven. A result passes only when preference-eligible distinct-stimulus votes have aggregate net win of at least 20% and an exact one-sided sign-test result of
ℹ️ How to read this report
➖ Not proven improved — lightweight-telemetry (claude-sonnet-5)Why: Net win +30.0% (6W/1T/3L over 10 preference-eligible stimulus vote(s), sign test p=0.254), mean preference +4.6% across 13 paired run(s), 3 dormancy stimulus/stimuli excluded from preference — not credible (sign test p=0.254 > 0.05) Next action: Inspect activation-only failed runs before rewriting skill content; the model stopped after loading a skill. State: Gate evidence: n=10; 6W/1T/3L; d=9; p=0.254; net +30.0%; 3 dormancy excluded Warnings: Activation-only stop: isolated 1 failed run; Activation-only stop: plugin 1 failed run Overfit: Moderate (score 0.25) Repeated-run reliability (not used by the gate): 13 paired runs (6W/1T/6L). Weak or warning scenarios:
Illustrative judge evidence:
This is one example, not the aggregate verdict. Open Full Results for every judgment. ➖ Not proven improved — lightweight-telemetry (gpt-5.6-luna)Why: Net win +10.0% (4W/3T/3L over 10 preference-eligible stimulus vote(s), sign test p=0.500), mean preference +0.0% across 13 paired run(s), 3 dormancy stimulus/stimuli excluded from preference — not credible (sign test p=0.500 > 0.05) Next action: Inspect activation-only failed runs before rewriting skill content; the model stopped after loading a skill. State: Gate evidence: n=10; 4W/3T/3L; d=7; p=0.500; net +10.0%; 3 dormancy excluded Warnings: Activation-only stop: isolated 1 failed run; Activation-only stop: plugin 1 failed run Overfit: Moderate (score 0.26) Repeated-run reliability (not used by the gate): 13 paired runs (4W/5T/4L). Weak or warning scenarios:
Illustrative judge evidence:
This is one example, not the aggregate verdict. Open Full Results for every judgment. 🔍 Full Results - all metrics and investigation details
|
Summary
Add a dependency-free .NET 11 telemetry skill for console tools, with guidance for counters, gauges, histograms, bounded tags, and JSON-line output. Add evaluation cases and update the dotnet11 catalog. This imports the reviewed work by qapdex-maker from #1036; it leaves unrelated README website links out.
Why
Small tools need a way to report measurements without an OpenTelemetry SDK, APM package, or collector. The skill shows how to use the built-in
System.Diagnostics.MetricsAPIs and receive measurements before a short-lived process exits.Impact
The new skill and evaluation cover .NET 11 console tools. This change adds no runtime dependency and does not change CI configuration.
Validation
The coordinating Windows worktree built and ran the complete inline
Program.cswith the pinned .NET 11 SDK: zero warnings, zero errors, and three JSON lines, each with all seven required fields. The dotnet11 skill-validator check and the repository eval quality gate passed.git diff --checkpassed, and the branch differs from current main only in the five reviewed files.