Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
29 changes: 28 additions & 1 deletion .github/agent/probe-issue.md
Original file line number Diff line number Diff line change
Expand Up @@ -17,7 +17,11 @@ permission:

# Doc-validation agent

You write deterministic, reviewable tests that run via CI on the PR.
Your purpose is to prepare a PR whose CI result validates or refutes a doc
claim. You are preparing for CI — CI is the ultimate test. `run_tox` is a
tool for increasing confidence that your preparation is sound (imports
resolve, types check, formatting passes). It is not the arbiter of whether
your test is correct — CI is.

## Core principle

Expand Down Expand Up @@ -50,6 +54,29 @@ are identical modulo the marker — so do not vary anything else between them.
`strict=True` matters: if the xfailed test unexpectedly passes, CI fails,
surfacing that the behavioural difference you expected does not actually exist.

### Test strategy

Choose your test type based on what the issue is about, not based on what
`run_tox` can run. If the claim is about integration test behaviour (e.g.,
what Jubilant logs during `juju.wait()`), write an integration test — even
though `run_tox` can only check that it imports and type-checks. CI will run
the integration test and determine the outcome. If the claim is about unit
test behaviour, write a unit test.

**Stay grounded in the issue's context.** The issue describes a claim in a
specific context — a particular library, tool, or test type. Your test should
engage with that context, not abstract it away. If the issue is about
Jubilant's logging, use Jubilant's logger (`jubilant.wait`), not a generic
Python logger. If the issue is about a specific library version, pin that
version and test against it. Before writing your test, verify that it
exercises the thing the issue is actually about.

**Do not be shy about integration tests.** `run_tox` runs `format,lint,unit`
only — not integration tests. But integration tests are first-class: they run
in CI after the reviewer marks the PR ready. Write them when the claim is
about integration test behaviour. Use `run_tox` to validate that they import
and type-check; let CI validate the behaviour.

## What you receive

The calling prompt supplies an issue number, a pre-created branch, and the
Expand Down
29 changes: 28 additions & 1 deletion .github/scripts/probe_issue.py
Original file line number Diff line number Diff line change
Expand Up @@ -286,7 +286,11 @@ def test_deploy(charm, juju: jubilant.Juju):
TASK_INSTRUCTIONS = """\
## Task instructions

You write deterministic, reviewable tests that run via CI on the PR.
Your purpose is to prepare a PR whose CI result validates or refutes a doc \
claim. You are preparing for CI — CI is the ultimate test. `run_tox` is a \
tool for increasing confidence that your preparation is sound (imports \
resolve, types check, formatting passes). It is not the arbiter of whether \
your test is correct — CI is.

### Core principle

Expand All @@ -308,6 +312,29 @@ def test_deploy(charm, juju: jubilant.Juju):
states what you believed, what you tested, and what the CI result means for the \
doc.

### Test strategy

Choose your test type based on what the issue is about, not based on what \
`run_tox` can run. If the claim is about integration test behaviour (e.g., \
what Jubilant logs during `juju.wait()`), write an integration test — even \
though `run_tox` can only check that it imports and type-checks. CI will run \
the integration test and determine the outcome. If the claim is about unit \
test behaviour, write a unit test.

**Stay grounded in the issue's context.** The issue describes a claim in a \
specific context — a particular library, tool, or test type. Your test should \
engage with that context, not abstract it away. If the issue is about \
Jubilant's logging, use Jubilant's logger (`jubilant.wait`), not a generic \
Python logger. If the issue is about a specific library version, pin that \
version and test against it. Before writing your test, verify that it \
exercises the thing the issue is actually about.

**Do not be shy about integration tests.** `run_tox` runs `format,lint,unit` \
only — not integration tests. But integration tests are first-class: they run \
in CI after the reviewer marks the PR ready. Write them when the claim is \
about integration test behaviour. Use `run_tox` to validate that they import \
and type-check; let CI validate the behaviour.

### Differential testing with xfail

Sometimes a claim is best tested by showing that the SAME test behaves \
Expand Down