Add .NET 11 preview telemetry skill and local-dev (Termux/glibc) docs - #1036
qapdex-maker wants to merge 18 commits into
Conversation
|
@dotnet-policy-service agree |
…t-telemetry skill - docs/LOCAL-DEVELOPMENT.md: documents that the pinned .NET 11 preview SDK is glibc-only and cannot run on Bionic hosts (Termux/Android); lists supported environments (Codespaces, Docker, WSL2, glibc VM) and what does NOT work. - plugins/dotnet11/skills/lightweight-telemetry: small dependency-free .NET 11 telemetry sample using System.Diagnostics.Metrics + TimeProvider. - docs/slides/README.md: reserved briefing for the deferred visual manual. - README.md: link the local-dev doc and the website/dashboard.
- docs/LOCAL-DEVELOPMENT.md: add a verified walkthrough for running the .NET 11 preview SDK inside the glibc Ubuntu 24.04 guest of qapdex-maker/ubuntu-termux, including the required PRoot workaround (DOTNET_SYSTEM_GLOBALIZATION_INVARIANT, DOTNET_GCHeapHardLimit, ulimit -v) and the worked lightweight-telemetry sample run with its actual output. - plugins/dotnet11/skills/lightweight-telemetry/SKILL.md: add a "Running the sample inside ubuntu-termux (PRoot)" section so the skill carries the same workaround. All steps and output were verified on an arm64 Termux/PRoot host.
AbhitejJohn
left a comment
There was a problem hiding this comment.
Thanks for contributing. We'd want to have evals for this skill so we know how much of a lift it is providing over baseline models before we can consider it. I've linked guidance on writing evals in one of the comments. Please let us know if you have any clarifying questions.
| @@ -0,0 +1,148 @@ | |||
| # Local development on non-glibc hosts | |||
|
|
|||
| This repository pins a **.NET 11 preview SDK** in the root `global.json`: | |||
There was a problem hiding this comment.
These should all be part of the repo's readme/contributing docs. If something is missing there, please feel free to suggest a change there.
| using System.Diagnostics; | ||
| using System.Diagnostics.Metrics; | ||
| using System.Text.Json; | ||
|
|
There was a problem hiding this comment.
This would fit best as an eval over a sample as the contributing doc calls out - https://github.com/dotnet/skills/blob/main/CONTRIBUTING.md#testing-and-validation. Add an eval would help us understand how much of a lift this skill provides users over the baseline model and is required for all skills in this repo.
Adds the skill test the PR review (dotnet#1036) requested. The eval uses the Vally schema mirrored from tests/dotnet11/system-text-json-net11: - 6 distinct stimuli (>=5 floor for statistical power) - 4 activation scenarios: built-in System.Diagnostics.Metrics + Meter, TimeProvider timing, stable Meter name, MeterListener JSON sink - 2 non-activation scenarios: distributed tracing and cloud ingestion are correctly routed to OpenTelemetry / vendor SDK instead check_eval_quality.py passes with "No errors." CODEOWNERS: add explicit entries for the new skill and its test, matching the existing system-text-json-net11 pattern.
The previous 6-stimulus eval sat in the statistically fragile 5-7 band (any loss is fatal, a tie can drop below 5 discordant votes). Bump to 10 distinct stimuli for a survivable one-loss margin: - keep the 6 existing scenarios (built-in metrics, TimeProvider timing, stable Meter name, MeterListener JSON sink, two non-activation cases) - add: gauge for live scalars, tagged measurements, instrument unit/description, non-activation for log aggregation (logging != metrics) check_eval_quality.py passes with "No errors." (10 distinct stimuli).
Re: requested evals — added and validatedThanks @AbhitejJohn for the review. You asked for an eval so we can see the lift this skill provides over the baseline model. That gap is now closed. Details below, including a note on the red What was added
Eval design (Vally schema)The eval mirrors the format of the existing
Local validationRan the repo's own quality gate before pushing — this is the check the review pointed at: It reports "No errors." (exit 0). The skill's eval no longer appears in the underpowered/fragile bands; at 10 stimuli it clears the structural defect classes (missing fixtures, missing grader config, About the red
|
|
Re-check on the eval requirement: the skill now ships tests/dotnet11/lightweight-telemetry/eval.yaml (Vally schema, 10 distinct stimuli — 7 activation + 3 non-activation with expect_activation:false). Local |
- cmd_review in idun_multi.py: diff -> chunk -> race over provider ensemble (anthropic/hf/deepseek/openai/gemini/mistral, those with creds, max 3) -> merged review, optional PR comment (--post) or dry-run. - _review_providers() picks credentialed providers. - docs/code-review-options.md: detailed self-built vs Qodo comparison + decision. - Proven by dry-run on dotnet/skills#1036 (hf "KEINE FUNDE", openai 429 graceful). - Tests green (pytest exit 0).
The eval sat at 7 preference-eligible stimuli, inside the 5-7 fragile band the quality gate warns about (one loss fatal, a tie can drop below the discordant floor). It also asserted the skill's own vocabulary (CreateGauge, MeterListener, InstrumentPublished), which is technique/vocabulary overfitting rather than outcome measurement. Eval: - 10 preference stimuli (was 7) + 3 dormancy contracts, each discriminating a different decision: instrument choice for a level vs a monotonic total, dimension via tag vs instrument-per-value, unit/description metadata, cheap measurement path when nothing collects, listener lifetime in a short-lived process, testable clock seam, stable metric identity. - Rubrics rewritten as outcomes; prompts no longer leak API names, so the baseline arm is not cued. - Dormancy guards now carry explicit anti-hijack rubric items (clears the gate's dormancy warning) and answer the real question instead of only declining. - config: -> defaults: (config is the deprecated alias), timeout 6m for code-generating stimuli. Skill: added the content the new stimuli demand and the baseline gets wrong — an instrument-selection table, tag-vs-name dimensions with a cardinality warning, Instrument.Enabled guarding + TagList to keep the hot path cheap, and listener lifetime (Start before first measurement, RecordObservableInstruments before exit). Verified, not asserted: the sample and every API claim were compiled and run. No net11.0 preview SDK is available on this host, so the code was exercised on net10.0 (these System.Diagnostics.Metrics APIs are unchanged) - build succeeded with 0 warnings/0 errors and the run emits tagged JSON lines carrying unit and description. The lifetime claim is from observed behaviour: a measurement recorded before listener.Start() produced no output line; the same measurement after it produced exactly one. check_eval_quality.py reports "No errors." and its 27 self-tests pass. skill-validator could not be run here (global.json pins the net11 preview SDK).
Re: evals — resized for real statistical power, and the skill rewritten to matchThanks again for the review, @AbhitejJohn. Since my last comment I went back over 1. The eval was not as powerful as I claimedMy previous comment said "10 distinct stimuli". That was wrong in the way that matters: — still inside the 5–7 band the gate warns about, where one loss is fatal and a single tie The 10 preference stimuli each discriminate a different decision rather than re-testing one
Plus 3 dormancy contracts ( 2. The graders were measuring vocabulary, not outcomeThe previous version asserted Also switched the deprecated 3. The skill now teaches what those stimuli ask forResizing the eval exposed that the skill was thin on exactly the decisions a baseline model Validation — what I actually ran, and what I could notRan:
For the same reason the sample cannot be built at {"meter":"MyTool","instrument":"tool.step.duration","unit":"ms","description":"Duration per step","value":103.2794,"tags":{"step":"restore"}}The lifetime guidance in the skill is from observed behaviour, not from reading source: a The red
|
Follows the review on dotnet#1036. AbhitejJohn asked twice for the same thing: the sample belongs in an eval over a sample rather than in the skill, and the platform notes belong in the repo's README/contributing docs rather than in a root-level file. No skill in this repo ships a sample/ directory (find plugins -type d -name 'sample*' is empty) — examples are code fences in SKILL.md. Removed: - sample/Program.cs and sample/telemetry.csproj - docs/LOCAL-DEVELOPMENT.md - docs/slides/README.md - the README section that linked to LOCAL-DEVELOPMENT.md SKILL.md changes, in response to the quality bar in CONTRIBUTING.md: - Replaced the "Sample (runnable)" section with an explicit output contract, since the bar asks a skill to end with one: one JSON object per line, and the seven fields in order. - Documented that a non-standard unit is written in UCUM annotation form ({run}, {item}), so an agent does not "correct" it to a plain noun. - Added a verify step with the two commands and the first thing to check when run prints nothing. - Corrected the verification claim from "Verified on .NET 10" to the pinned preview SDK. The sample now builds and runs against 11.0.100-preview.3.26207.106 on net11.0: build succeeded with 0 warnings and 0 errors, and run emitted four JSON lines, one per measurement — two histogram readings split by the step tag, the counter, and the observable gauge. The platform constraint (the pinned SDK is a glibc build and cannot run on a Bionic host) is real and was verified, but it is a fact about the SDK distribution rather than about this skill, so it moves to its own issue instead of riding along in a skill PR. eval unchanged: 13 stimuli, check_eval_quality.py exits 0 with no errors, and the eval is not among the flagged files.
|
@AbhitejJohn thanks for the review, and apologies for the slow turnaround. Both of your points are addressed in c9d67b8. Evals. The sample and the platform docs. You said both twice, and you were right — I checked the repo rather than arguing. The skill now ends with a concrete output contract instead of a sample pointer: one JSON object per line and the seven fields in order. I also added the UCUM annotation detail ( On the verification claim. The previous version said "Verified on .NET 10" with a note that the APIs are unchanged. That was weak, so I rebuilt and ran against the pinned preview SDK from
The platform constraint. You offered to take platform notes as a change to the repo docs, and that is where it belongs — it is a fact about how the SDK is distributed, not about this skill. I am opening it as a separate issue so it does not hold up this PR, and I will put the verified detail there: the pinned preview is a glibc build, the ELF interpreter it wants is One question before you spend review time: does |
|
Thanks a bunch for the updates, qapdex-maker. I moved the current skill and eval to a trusted draft PR based on the latest The targeted eval completed. The workflow passed, but the skill did not show a credible improvement over baseline in either model: Claude Sonnet 5 had 6 wins, 1 tie, 3 losses (p=0.254); GPT-5.6 Luna had 4 wins, 3 ties, 3 losses (p=0.500). I'm reviewing the hot-path and observable-collection cases before a wider review. To your question, I don't think so but I'll let @adamsitnik answer it:
(Copilot, commenting on Abhitej's behalf.) |
Addressing the two inline commentsBoth inline comments point at 1. Eval over the sample ( Agreed, and that is what If you would rather have the sample back in-tree as the eval's fixture, say so and I will add it back under a path that is excluded from any repo-wide build. 2. Local-development notes ( Agreed — this belongs in the repo's own docs, not in a PR. I dropped the file for the same reason and instead put the two durable, non-.NET-specific facts in
The one genuinely .NET-specific finding — that the pinned .NET 11 preview SDK is a glibc build and does not run on Bionic-only hosts — is deliberately not in this PR. It is about how the SDK is distributed, not about this skill, and it belongs in 3. Fork-trust on Still the only red check, and still not a content problem: the run shows If there is a maintainer-side action needed to trust the branch, I am happy to do whatever is available from my side — just let me know what it is. Current state of the contribution
All three items from the |
The System.Diagnostics.Metrics owner is @tarekgh (cc @jeffhandley) |
|
Thanks for the detailed updates. I reviewed the current PR, the trusted copy in #1237, and the targeted evaluation results. To answer the question about preview status: The two unresolved inline comments point to files that have since been deleted. The repository-documentation concern was addressed by removing the local development document, and the request for an eval was addressed by adding There are still substantive issues to address:
My recommendation is to close #1036 as superseded by the trusted draft #1237, then continue the design discussion there. Before merging #1237, I think we should decide whether this belongs in a general .NET diagnostics plugin rather than |
Summary
Adds the
dotnet11/lightweight-telemetryskill: emitting structured metrics from a .NET console tool with the built-inSystem.Diagnostics.MetricsAPI, no OpenTelemetry, no APM vendor SDK, no collector.The skill is for the case where the host has nowhere to ship telemetry to — a build tool, a CLI, a local agent. It covers instrument selection (counter vs gauge vs histogram vs observable gauge), splitting a metric by a bounded dimension, keeping the measurement path cheap when nothing is listening, and listener lifetime in a short-lived process. Each reading is one JSON object on its own line, so a consumer can
tail -fand parse.The skill is scoped against its siblings on the real discriminator: distributed tracing goes to
configuring-opentelemetry-dotnet, cloud ingestion to the vendor SDK, log shipping to a logging pipeline. Thedescriptioncarries the exclusions, and the eval asserts the skill stays dormant on those three requests.What is in the PR
plugins/dotnet11/skills/lightweight-telemetry/SKILL.md— the skilltests/dotnet11/lightweight-telemetry/eval.yaml— 13 stimuli: 10 that should activate, 3 that must not (tracing, cloud ingestion, log shipping).github/CODEOWNERS— route to@dotnet/skills-csharp-language-reviewers, same as the existing dotnet11 skillVerification
Built and run against the pinned preview SDK from
global.json,11.0.100-preview.3.26207.106, targetingnet11.0:dotnet run -c Release --no-buildemitted one JSON line per measurement, four in total — two histogram readings split by thesteptag, the counter, and the observable gauge:{"meter":"MyTool","instrument":"tool.step.duration","unit":"ms","description":"Duration per step","value":83.9945,"tags":{"step":"restore"},"timestamp":"2026-09-28T11:43:02.4207328+00:00"} {"meter":"MyTool","instrument":"tool.step.duration","unit":"ms","description":"Duration per step","value":61.4548,"tags":{"step":"compile"},"timestamp":"2026-09-28T11:43:02.6133652+00:00"} {"meter":"MyTool","instrument":"tool.runs","unit":"{run}","description":"Number of executions","value":1,"tags":{},"timestamp":"2026-09-28T11:43:02.6149784+00:00"} {"meter":"MyTool","instrument":"tool.queue.depth","unit":"{item}","description":"Items currently queued","value":3,"tags":{},"timestamp":"2026-09-28T11:43:02.6184723+00:00"}The eval passes the repo's own gate:
Follow-up
The pinned .NET 11 preview SDK is a glibc build and does not run on Bionic-only hosts, which is how I hit the build problem in the first place. That is a fact about how the SDK is distributed rather than about this skill, so it is raised separately instead of riding along here.