Skip to content

Add dotnet11/lightweight-telemetry skill with eval - #1263

Open
qapdex-maker wants to merge 7 commits into
dotnet:mainfrom
qapdex-maker:feature/dotnet11-lightweight-telemetry
Open

qapdex-maker wants to merge 7 commits into
dotnet:mainfrom
qapdex-maker:feature/dotnet11-lightweight-telemetry

Conversation

@qapdex-maker

Copy link
Copy Markdown
Contributor

Adds the dotnet11/lightweight-telemetry skill plus its eval, and documents
the local-development constraint discovered while authoring it.

Commits (oldest first)

  1. docs: add local-dev constraint (glibc/.NET 11 preview) and lightweight-telemetry skill
  2. docs: document running .NET 11 preview inside ubuntu-termux (PRoot)
  3. tests: add eval.yaml + CODEOWNERS for dotnet11/lightweight-telemetry
  4. tests: strengthen dotnet11/lightweight-telemetry eval to 10 stimuli
  5. refactor: drop the runnable sample and root docs from the skill
  6. Strengthen dotnet11/lightweight-telemetry skill and eval
  7. Make groom publisher fixture date-relative instead of hard-coded

Diffstat against current main: 4 files, +430 across the skill, its eval, the
CODEOWNERS entry and the docs.

Why the sample was dropped

The runnable sample/ project was removed in commit 5. The skill documents that
a .NET 11 preview cannot be executed directly inside Termux, only via PRoot, so
the sample could not be built or run by the target audience as written. The
skill text and docs/LOCAL-DEVELOPMENT.md carry that constraint instead.

The last commit is a test fix

evaluation-workflow-tests went red on 2026-10-02 (run 37046657914) in
test_groom_publisher_preserves_resolved_dispatched_rows. The groom publisher
expires resolved outbox rows older than 14 days via Date.now(), and the test
fixture pinned correlation_date to the literal 2026-09-16. Once real time
passed that date, the row was dropped from the rendered dashboard and the
assertion failed. The fixture now defaults to yesterday, so it stays inside the
retention window; the expiry test still pins 2000-01-01 explicitly.

Verification

python3 eng/evaluation/test_token_failover.py

  • before the fixture fix: FAILED (failures=1, errors=23)
  • after: FAILED (errors=23)

python3 eng/eval-quality/check_eval_quality.py → No errors.

The 23 errors are all FileNotFoundError: 'pwsh'; PowerShell is not installed
on the machine this was verified on. They are unrelated to these changes and
will not appear in CI.

…t-telemetry skill

- docs/LOCAL-DEVELOPMENT.md: documents that the pinned .NET 11 preview SDK is
  glibc-only and cannot run on Bionic hosts (Termux/Android); lists supported
  environments (Codespaces, Docker, WSL2, glibc VM) and what does NOT work.
- plugins/dotnet11/skills/lightweight-telemetry: small dependency-free .NET 11
  telemetry sample using System.Diagnostics.Metrics + TimeProvider.
- docs/slides/README.md: reserved briefing for the deferred visual manual.
- README.md: link the local-dev doc and the website/dashboard.
- docs/LOCAL-DEVELOPMENT.md: add a verified walkthrough for running the .NET 11
  preview SDK inside the glibc Ubuntu 24.04 guest of qapdex-maker/ubuntu-termux,
  including the required PRoot workaround (DOTNET_SYSTEM_GLOBALIZATION_INVARIANT,
  DOTNET_GCHeapHardLimit, ulimit -v) and the worked lightweight-telemetry sample
  run with its actual output.
- plugins/dotnet11/skills/lightweight-telemetry/SKILL.md: add a "Running the
  sample inside ubuntu-termux (PRoot)" section so the skill carries the same
  workaround.

All steps and output were verified on an arm64 Termux/PRoot host.
Adds the skill test the PR review (dotnet#1036) requested. The
eval uses the Vally schema mirrored from tests/dotnet11/system-text-json-net11:

- 6 distinct stimuli (>=5 floor for statistical power)
- 4 activation scenarios: built-in System.Diagnostics.Metrics + Meter,
  TimeProvider timing, stable Meter name, MeterListener JSON sink
- 2 non-activation scenarios: distributed tracing and cloud ingestion
  are correctly routed to OpenTelemetry / vendor SDK instead

check_eval_quality.py passes with "No errors."

CODEOWNERS: add explicit entries for the new skill and its test,
matching the existing system-text-json-net11 pattern.
The previous 6-stimulus eval sat in the statistically fragile 5-7 band
(any loss is fatal, a tie can drop below 5 discordant votes). Bump to
10 distinct stimuli for a survivable one-loss margin:

- keep the 6 existing scenarios (built-in metrics, TimeProvider timing,
  stable Meter name, MeterListener JSON sink, two non-activation cases)
- add: gauge for live scalars, tagged measurements, instrument unit/description,
  non-activation for log aggregation (logging != metrics)

check_eval_quality.py passes with "No errors." (10 distinct stimuli).
The eval sat at 7 preference-eligible stimuli, inside the 5-7 fragile band the
quality gate warns about (one loss fatal, a tie can drop below the discordant
floor). It also asserted the skill's own vocabulary (CreateGauge, MeterListener,
InstrumentPublished), which is technique/vocabulary overfitting rather than
outcome measurement.

Eval:
- 10 preference stimuli (was 7) + 3 dormancy contracts, each discriminating a
  different decision: instrument choice for a level vs a monotonic total,
  dimension via tag vs instrument-per-value, unit/description metadata,
  cheap measurement path when nothing collects, listener lifetime in a
  short-lived process, testable clock seam, stable metric identity.
- Rubrics rewritten as outcomes; prompts no longer leak API names, so the
  baseline arm is not cued.
- Dormancy guards now carry explicit anti-hijack rubric items (clears the
  gate's dormancy warning) and answer the real question instead of only
  declining.
- config: -> defaults: (config is the deprecated alias), timeout 6m for
  code-generating stimuli.

Skill: added the content the new stimuli demand and the baseline gets wrong —
an instrument-selection table, tag-vs-name dimensions with a cardinality
warning, Instrument.Enabled guarding + TagList to keep the hot path cheap, and
listener lifetime (Start before first measurement, RecordObservableInstruments
before exit).

Verified, not asserted: the sample and every API claim were compiled and run.
No net11.0 preview SDK is available on this host, so the code was exercised on
net10.0 (these System.Diagnostics.Metrics APIs are unchanged) - build succeeded
with 0 warnings/0 errors and the run emits tagged JSON lines carrying unit and
description. The lifetime claim is from observed behaviour: a measurement
recorded before listener.Start() produced no output line; the same measurement
after it produced exactly one. check_eval_quality.py reports "No errors." and
its 27 self-tests pass. skill-validator could not be run here (global.json
pins the net11 preview SDK).
Follows the review on dotnet#1036. AbhitejJohn asked twice for the same thing:
the sample belongs in an eval over a sample rather than in the skill, and
the platform notes belong in the repo's README/contributing docs rather
than in a root-level file. No skill in this repo ships a sample/
directory (find plugins -type d -name 'sample*' is empty) — examples are
code fences in SKILL.md.

Removed:
- sample/Program.cs and sample/telemetry.csproj
- docs/LOCAL-DEVELOPMENT.md
- docs/slides/README.md
- the README section that linked to LOCAL-DEVELOPMENT.md

SKILL.md changes, in response to the quality bar in CONTRIBUTING.md:
- Replaced the "Sample (runnable)" section with an explicit output
  contract, since the bar asks a skill to end with one: one JSON object
  per line, and the seven fields in order.
- Documented that a non-standard unit is written in UCUM annotation form
  ({run}, {item}), so an agent does not "correct" it to a plain noun.
- Added a verify step with the two commands and the first thing to check
  when run prints nothing.
- Corrected the verification claim from "Verified on .NET 10" to the
  pinned preview SDK. The sample now builds and runs against
  11.0.100-preview.3.26207.106 on net11.0: build succeeded with 0
  warnings and 0 errors, and run emitted four JSON lines, one per
  measurement — two histogram readings split by the step tag, the
  counter, and the observable gauge.

The platform constraint (the pinned SDK is a glibc build and cannot run
on a Bionic host) is real and was verified, but it is a fact about the
SDK distribution rather than about this skill, so it moves to its own
issue instead of riding along in a skill PR.

eval unchanged: 13 stimuli, check_eval_quality.py exits 0 with no
errors, and the eval is not among the flagged files.
The resolved-dispatched fixture pinned correlation_date to 2026-09-16. The
groom publisher expires resolved outbox rows older than 14 days using
Date.now(), so the fixture silently started failing once real time passed
that date (observed red on 2026-10-02, test_groom_publisher_preserves_
resolved_dispatched_rows).

Default the correlation to yesterday so the row stays inside the retention
window; the expiry test still pins 2000-01-01 explicitly.
@github-actions

github-actions Bot commented Oct 5, 2026

Copy link
Copy Markdown
Contributor

Note

This PR is from a fork and modifies infrastructure files (eng/ or .github/).

Changes to infrastructure typically need to be submitted from a branch in dotnet/skills (not a fork) so that CI workflows run with the correct permissions and secrets.

Please consider recreating this PR from an upstream branch. If you don't have push access to dotnet/skills, ask a maintainer to push your branch for you.

This branch has not been deployed

No deployments
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant