Add dotnet11/lightweight-telemetry skill with eval - #1263
Open
qapdex-maker wants to merge 7 commits into
Open
qapdex-maker wants to merge 7 commits into
qapdex-maker wants to merge 7 commits into
Conversation
…t-telemetry skill - docs/LOCAL-DEVELOPMENT.md: documents that the pinned .NET 11 preview SDK is glibc-only and cannot run on Bionic hosts (Termux/Android); lists supported environments (Codespaces, Docker, WSL2, glibc VM) and what does NOT work. - plugins/dotnet11/skills/lightweight-telemetry: small dependency-free .NET 11 telemetry sample using System.Diagnostics.Metrics + TimeProvider. - docs/slides/README.md: reserved briefing for the deferred visual manual. - README.md: link the local-dev doc and the website/dashboard.
- docs/LOCAL-DEVELOPMENT.md: add a verified walkthrough for running the .NET 11 preview SDK inside the glibc Ubuntu 24.04 guest of qapdex-maker/ubuntu-termux, including the required PRoot workaround (DOTNET_SYSTEM_GLOBALIZATION_INVARIANT, DOTNET_GCHeapHardLimit, ulimit -v) and the worked lightweight-telemetry sample run with its actual output. - plugins/dotnet11/skills/lightweight-telemetry/SKILL.md: add a "Running the sample inside ubuntu-termux (PRoot)" section so the skill carries the same workaround. All steps and output were verified on an arm64 Termux/PRoot host.
Adds the skill test the PR review (dotnet#1036) requested. The eval uses the Vally schema mirrored from tests/dotnet11/system-text-json-net11: - 6 distinct stimuli (>=5 floor for statistical power) - 4 activation scenarios: built-in System.Diagnostics.Metrics + Meter, TimeProvider timing, stable Meter name, MeterListener JSON sink - 2 non-activation scenarios: distributed tracing and cloud ingestion are correctly routed to OpenTelemetry / vendor SDK instead check_eval_quality.py passes with "No errors." CODEOWNERS: add explicit entries for the new skill and its test, matching the existing system-text-json-net11 pattern.
The previous 6-stimulus eval sat in the statistically fragile 5-7 band (any loss is fatal, a tie can drop below 5 discordant votes). Bump to 10 distinct stimuli for a survivable one-loss margin: - keep the 6 existing scenarios (built-in metrics, TimeProvider timing, stable Meter name, MeterListener JSON sink, two non-activation cases) - add: gauge for live scalars, tagged measurements, instrument unit/description, non-activation for log aggregation (logging != metrics) check_eval_quality.py passes with "No errors." (10 distinct stimuli).
The eval sat at 7 preference-eligible stimuli, inside the 5-7 fragile band the quality gate warns about (one loss fatal, a tie can drop below the discordant floor). It also asserted the skill's own vocabulary (CreateGauge, MeterListener, InstrumentPublished), which is technique/vocabulary overfitting rather than outcome measurement. Eval: - 10 preference stimuli (was 7) + 3 dormancy contracts, each discriminating a different decision: instrument choice for a level vs a monotonic total, dimension via tag vs instrument-per-value, unit/description metadata, cheap measurement path when nothing collects, listener lifetime in a short-lived process, testable clock seam, stable metric identity. - Rubrics rewritten as outcomes; prompts no longer leak API names, so the baseline arm is not cued. - Dormancy guards now carry explicit anti-hijack rubric items (clears the gate's dormancy warning) and answer the real question instead of only declining. - config: -> defaults: (config is the deprecated alias), timeout 6m for code-generating stimuli. Skill: added the content the new stimuli demand and the baseline gets wrong — an instrument-selection table, tag-vs-name dimensions with a cardinality warning, Instrument.Enabled guarding + TagList to keep the hot path cheap, and listener lifetime (Start before first measurement, RecordObservableInstruments before exit). Verified, not asserted: the sample and every API claim were compiled and run. No net11.0 preview SDK is available on this host, so the code was exercised on net10.0 (these System.Diagnostics.Metrics APIs are unchanged) - build succeeded with 0 warnings/0 errors and the run emits tagged JSON lines carrying unit and description. The lifetime claim is from observed behaviour: a measurement recorded before listener.Start() produced no output line; the same measurement after it produced exactly one. check_eval_quality.py reports "No errors." and its 27 self-tests pass. skill-validator could not be run here (global.json pins the net11 preview SDK).
Follows the review on dotnet#1036. AbhitejJohn asked twice for the same thing: the sample belongs in an eval over a sample rather than in the skill, and the platform notes belong in the repo's README/contributing docs rather than in a root-level file. No skill in this repo ships a sample/ directory (find plugins -type d -name 'sample*' is empty) — examples are code fences in SKILL.md. Removed: - sample/Program.cs and sample/telemetry.csproj - docs/LOCAL-DEVELOPMENT.md - docs/slides/README.md - the README section that linked to LOCAL-DEVELOPMENT.md SKILL.md changes, in response to the quality bar in CONTRIBUTING.md: - Replaced the "Sample (runnable)" section with an explicit output contract, since the bar asks a skill to end with one: one JSON object per line, and the seven fields in order. - Documented that a non-standard unit is written in UCUM annotation form ({run}, {item}), so an agent does not "correct" it to a plain noun. - Added a verify step with the two commands and the first thing to check when run prints nothing. - Corrected the verification claim from "Verified on .NET 10" to the pinned preview SDK. The sample now builds and runs against 11.0.100-preview.3.26207.106 on net11.0: build succeeded with 0 warnings and 0 errors, and run emitted four JSON lines, one per measurement — two histogram readings split by the step tag, the counter, and the observable gauge. The platform constraint (the pinned SDK is a glibc build and cannot run on a Bionic host) is real and was verified, but it is a fact about the SDK distribution rather than about this skill, so it moves to its own issue instead of riding along in a skill PR. eval unchanged: 13 stimuli, check_eval_quality.py exits 0 with no errors, and the eval is not among the flagged files.
The resolved-dispatched fixture pinned correlation_date to 2026-09-16. The groom publisher expires resolved outbox rows older than 14 days using Date.now(), so the fixture silently started failing once real time passed that date (observed red on 2026-10-02, test_groom_publisher_preserves_ resolved_dispatched_rows). Default the correlation to yesterday so the row stays inside the retention window; the expiry test still pins 2000-01-01 explicitly.
qapdex-maker
requested review from
a team,
AbhitejJohn,
JanKrivanek and
webreidi
as code owners
October 5, 2026 14:19
Contributor
|
Note This PR is from a fork and modifies infrastructure files ( Changes to infrastructure typically need to be submitted from a branch in Please consider recreating this PR from an upstream branch. If you don't have push access to |
This branch has not been deployed
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Adds the
dotnet11/lightweight-telemetryskill plus its eval, and documentsthe local-development constraint discovered while authoring it.
Commits (oldest first)
docs: add local-dev constraint (glibc/.NET 11 preview) and lightweight-telemetry skilldocs: document running .NET 11 preview inside ubuntu-termux (PRoot)tests: add eval.yaml + CODEOWNERS for dotnet11/lightweight-telemetrytests: strengthen dotnet11/lightweight-telemetry eval to 10 stimulirefactor: drop the runnable sample and root docs from the skillStrengthen dotnet11/lightweight-telemetry skill and evalMake groom publisher fixture date-relative instead of hard-codedDiffstat against current
main: 4 files, +430 across the skill, its eval, theCODEOWNERS entry and the docs.
Why the sample was dropped
The runnable
sample/project was removed in commit 5. The skill documents thata .NET 11 preview cannot be executed directly inside Termux, only via PRoot, so
the sample could not be built or run by the target audience as written. The
skill text and
docs/LOCAL-DEVELOPMENT.mdcarry that constraint instead.The last commit is a test fix
evaluation-workflow-testswent red on 2026-10-02 (run 37046657914) intest_groom_publisher_preserves_resolved_dispatched_rows. The groom publisherexpires resolved outbox rows older than 14 days via
Date.now(), and the testfixture pinned
correlation_dateto the literal2026-09-16. Once real timepassed that date, the row was dropped from the rendered dashboard and the
assertion failed. The fixture now defaults to yesterday, so it stays inside the
retention window; the expiry test still pins
2000-01-01explicitly.Verification
python3 eng/evaluation/test_token_failover.pyFAILED (failures=1, errors=23)FAILED (errors=23)python3 eng/eval-quality/check_eval_quality.py→No errors.The 23 errors are all
FileNotFoundError: 'pwsh'; PowerShell is not installedon the machine this was verified on. They are unrelated to these changes and
will not appear in CI.