Skip to content

Harden the ratchet, tests and CI ahead of the monorepo migration (step P) - #18

Merged
tap merged 3 commits into
mainfrom
claude/sample-rate-expansion-strategies-ezqzu6
Sep 27, 2026
Merged

tap merged 3 commits into
mainfrom
claude/sample-rate-expansion-strategies-ezqzu6

Conversation

@tap

@tap tap commented Sep 27, 2026 •

Copy link
Copy Markdown
Owner

What this changes

This is RatioTap's half of step P, the pre-work before the monorepo migration. The plan lives in SampleRateTap at docs/MONOREPO_PLAN.md on claude/sample-rate-expansion-strategies-ezqzu6, and tap/SampleRateTap#48 is the companion PR.

P.2: harness and CI (df8cd61, 8550e07)

  • scripts/icount.py is changed in step with SampleRateTap's copy; the two differ only in their constants.
    • New --exact, --json-out and --compare-json options.
    • A recorded workload with no binary is now a failure, and a missing baselines file is fatal.
    • Checksums are recorded.
    • Hexagon workloads now run isolated: fixed path, fixed argv[0], empty environment.
    • Hexagon baselines are re-recorded (8550e07). The first isolated run measured every scenario 3,724–3,809 instructions lower (−0.0004% to −0.0136%). The PLAN.md §7 ledger entry records this.
  • Test names: every test gets a ratio. ctest prefix and a ratio label, including the bare-metal entry.
  • New OutputHash suite: hashes are printed for comparison, never pinned.
  • Cross-validation output: the lines now print their tolerance.
  • CI:
    • contents: read permissions.
    • Runs on main and on the migration PR are never cancelled.
    • Every action is SHA-pinned.
    • QEMU and ratchet jobs are pinned to ubuntu-24.04, with the image OS in the cache key and the versions logged.
    • --no-tests=error everywhere.
    • QEMU jobs upload their full logs.

P.3: notebooks (9c92818)

  • Pinned environment: requirements.in and requirements.lock at the repository root pin the notebook environment with hashes. The file is identical to SampleRateTap's, and it replaces the unversioned notebooks/requirements.txt.
  • Bridge rebuilds: ratiotap_py now rebuilds build_capi/ incrementally on every import, and quietly; its CMake log used to land in ratio_demo's committed output.
  • Figure digests: notebooks/figure_digest.py prints a data digest for each figure.
  • Re-execution: all three notebooks were re-executed, and every committed number reproduces exactly. The only output changes are the added digest lines and the removed build log.

Why

The migration gates key tests by unique names. They compare exact icount, output hashes, full on-target logs and notebook outputs, all of which must exist before the step-0 snapshot.

Verification

  • Host, GCC: 82/82 pass with -Werror, all named ratio.* and labelled ratio.
  • Host, Clang: clean with -Werror. Hashes are identical to GCC's and stable across runs.
  • clang-tidy-18: clean on the new and edited tests.
  • Cortex-M33 under QEMU: all 10 workloads match exactly with --exact, and the emulated suite passes in 33 s.
  • CI on the current head (9c92818): every job is green (10/10), including all three QEMU legs and the ratchet against the re-recorded Hexagon baselines. The earlier head, df8cd61, was also fully green.
  • Notebooks: executed with pip install --require-hashes -r requirements.lock on Python 3.11.15. The digests are identical across two re-executions.

Notes for the reviewer

  • Contract change: test naming only (ratio. prefix).
  • Notebooks: re-executed and committed executed.
  • Merge order: this needs to merge before the migration snapshot. Afterwards, submodules/sampleratetap is re-pinned to SampleRateTap's post-step-P main (P.4).

🤖 Generated with Claude Code

https://claude.ai/code/session_015VR1VC4SDGxHZQQsQvPBaA

- icount.py (ported in step with SampleRateTap's): --exact, --json-out
  and --compare-json for same-job A/B measurement, a failure when a
  recorded workload has no binary, a fatal error on a missing baselines
  file, and the workload checksum printed and recorded. Hexagon
  workloads run from one fixed path with a fixed argv[0] and an empty
  environment, because qemu-hexagon copies those onto the guest stack and
  static musl's startup walks them; qemu-hexagon is resolved to an
  absolute path first. Hexagon baselines are re-recorded separately.
- Tests carry a "ratio." ctest prefix and a ratio label (the bare-metal
  entry too), so names stay unique once they share a tree with
  SampleRateTap's.
- New OutputHash suite: FNV-1a hashes of the converter's output for both
  directions, every format and every profile, printed for same-job
  comparison and never pinned. About 8 s on Cortex-M33.
- The cross-validation lines now print their tolerance, so a loosened
  limit is visible in the output and not only in the source.
- CI: contents: read permissions; cancellation spares main and the
  migration PR; every action SHA-pinned (checkout v6, cache v5, as in
  SampleRateTap); ubuntu-24.04 pinned for QEMU and ratchet jobs, with the
  image OS in the plugin-qemu cache key and image/toolchain versions
  logged; --no-tests=error on every ctest; QEMU legs keep and upload
  full test logs.

Step P.2 of the monorepo migration plan (SampleRateTap
docs/MONOREPO_PLAN.md).

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_015VR1VC4SDGxHZQQsQvPBaA
The previous commit runs Hexagon workloads from a fixed path with a
fixed argv[0] and an empty environment. The first isolated CI run (PR
#18) measured every scenario 3,724 to 3,809 instructions lower
(-0.0004% to -0.0136%): the path and environment strings static musl's
startup used to walk. Inside the gate, but re-recorded so exact
comparisons start from the new harness; PLAN.md section 7 ledger entry.
M33/M55 are unaffected.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_015VR1VC4SDGxHZQQsQvPBaA
- requirements.in / requirements.lock at the repository root: the
  notebook environment pinned with hashes (numpy 2.4.6, scipy 1.17.1,
  matplotlib 3.11.2, jupyter/nbconvert; samplerate and soxr for
  SampleRateTap's comparison notebook, since the file is identical in
  both repositories). It replaces notebooks/requirements.txt, which
  named packages without versions.
- ratiotap_py rebuilds build_capi/ incrementally on every import instead
  of only when the library is missing, so a library left over from an
  older checkout can no longer be measured silently, and the build is
  quiet (its log used to land in ratio_demo's committed output; it now
  prints only on failure).
- notebooks/figure_digest.py wraps plt.show() to print a digest of each
  figure's plotted data, quantized to 9 significant digits of each
  array's peak, so a changed curve shows up as changed text; stable
  across re-executions.
- All three notebooks re-executed in the pinned environment: every
  committed number reproduces exactly; the only output changes are the
  added digest lines and the removed build log.

Step P.3 of the monorepo migration plan.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_015VR1VC4SDGxHZQQsQvPBaA
@tap
tap marked this pull request as ready for review September 27, 2026 20:39
@tap
tap merged commit dcf2abc into main Sep 27, 2026
10 checks passed
tap pushed a commit that referenced this pull request Sep 27, 2026
The previous commit runs Hexagon workloads from a fixed path with a
fixed argv[0] and an empty environment. The first isolated CI run (PR
#18) measured every scenario 3,724 to 3,809 instructions lower
(-0.0004% to -0.0136%): the path and environment strings static musl's
startup used to walk. Inside the gate, but re-recorded so exact
comparisons start from the new harness; PLAN.md section 7 ledger entry.
M33/M55 are unaffected.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_015VR1VC4SDGxHZQQsQvPBaA
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants