Skip to content

feat: Rerun only the tests that failed in the last run with run-tests --rerun-failed - #3161

Merged
hatayama merged 10 commits into
mainfrom
feat/run-tests-rerun-failed
Oct 5, 2026
Merged

hatayama merged 10 commits into
mainfrom
feat/run-tests-rerun-failed

Conversation

@hatayama

@hatayama hatayama commented Oct 5, 2026 •

Copy link
Copy Markdown
Owner

Summary

  • uloop run-tests --rerun-failed reruns only the tests that failed or were inconclusive in the most recent completed run of the same test mode, in a single run.

User Impact

  • Before: after a failing run, checking a fix or telling a flaky test from a real failure meant copying failed test names into --filter-type exact one at a time, or rerunning the whole suite.
  • After: --rerun-failed reruns exactly the recorded failures (the whole fixture when its OneTimeSetUp or OneTimeTearDown failed). When there is nothing it can safely rerun, it says so and runs nothing instead of falling back to the full suite.

Behavior

  • Every completed run records its failures per test mode in .uloop/outputs/TestResults/last-run-<TestMode>.json. The record is removed when a run starts and written only once the run has a result, so a run that times out or is cancelled leaves no stale failures behind.
  • --rerun-failed returns without saving or discarding unsaved Scene and Prefab Stage changes, clearing pause points, or running tests when:
    • it is combined with --filter-type or --filter-value (ExecutionFailed)
    • no record exists for the test mode (ExecutionFailed)
    • the record cannot be read: invalid JSON, unknown format version, another test mode, missing fields (ExecutionFailed)
    • the recorded run had no failures: new status NothingToRerun with Success: true
  • Every run-tests request now resolves its filter, and with --rerun-failed the record, before the Editor-state check saves or discards unsaved changes. A request with an invalid filter therefore also returns with the Editor untouched, and a request that is invalid on both counts reports the filter or record problem before the Editor-state problem.
  • A rerun's response adds RerunTargetCount and RerunSourceCompletedAt. If every recorded test has since been renamed or removed, Message and NoTestsFoundExplanation say that instead of giving the missing-test-assembly advice.
  • With --respect-enter-play-mode-settings, a PlayMode run that reloads the domain on Play entry is recovered after the reload. That path also writes the PlayMode record, but its response carries no rerun fields; the skill's response reference documents this.
  • Unity's duplicate-name suffix ( GeneratedTestCase<n>) is stripped from recorded names so the filter matches them; a failing assembly or root suite is recorded as its leaves, because those suite names cannot be used as a filter.

Out of scope: --retries

  • Automatic retries are not added. They would need a rule for splitting the timeout across attempts and for merging counts across attempts, and a domain reload between attempts ends the retry loop. Running --rerun-failed again after a failure already tells a flaky test from a real failure.

Changes

  • Unity package: a last-run record store; rerun-target collection in the result converter; a test-names execution filter passed to Unity's testNames; the run-tests use case resolves a rerun before changing anything and records each completed run; the post-reload callback records recovered PlayMode runs.
  • Skill and catalog: --rerun-failed row and usage note; the response field list moves to references/response-fields.md (the run-tests SKILL.md shrinks from 7,894 to 5,716 bytes) with the new fields and status; RerunFailed in the embedded tool catalog; tool reference docs (English and Japanese).
  • The catalog change is structural, so both shared-inputs stamps are refreshed: this PR releases both the dispatcher and the project runner.
  • Protocol version is unchanged: one optional parameter and optional response fields keep the wire format compatible.
  • SerializableTestResultConverter.cs is now at 484 of the 500 allowed SLOC; split it before adding behavior there.

Verification

  • Unity EditMode, class-filtered on a local Editor: RunTestsUseCaseTests, RunTestsLastRunRecordStoreTests, RunTestsTestFrameworkResultTests, PlayModeTestExecuterTests, RunTestsTestNamesFilterTests, RunTestsResponseContractTests, DefaultToolsCatalogDriftTests: 153/153 passed on the final head. The 21 new use-case cases cover completed, timed-out, and cancelled runs (before start, during the run, during the cleanup wait), a record that cannot be cleared (no run, pause points untouched), the other test mode's record, a run without a result tree, NoTestsFound, filter conflicts (including whitespace-only and null values), missing, unreadable, and empty records, a rerun, recorded tests that are gone, test-mode selection, a rerun that times out, and a rerun stopped by the Editor-state check (record and pause points untouched). Every request that stops early is checked not to reach the Editor-state check.
  • Mutation checks, run after committing: in the use case, clearing pause points before removing the record, recording runs without a result tree, and sending reruns through the ordinary no-tests diagnostics fail exactly the 5 tests that cover them; in the recovery path, dropping the PlayMode-only and no-result guards fails exactly the 3 tests that cover them; keeping the asmdef explanation on a renamed-or-removed rerun and moving the Editor-state check back before the filter fail exactly the 8 tests that cover them; moving that check after the record removal fails the new untouched-state test and two existing ones. The record store, target collection, and test-names filter had the same check when they were added.
  • Manual run on the development project with temporary probe tests (not committed):
    1. Class run with one failing test: FailedCount: 1; the record lists only that test.
    2. --rerun-failed: TestCount: 1, RerunTargetCount: 1, FailedCount: 1.
    3. Fix the test, compile, --rerun-failed: Success: true, TestCount: 1.
    4. --rerun-failed again: Status: NothingToRerun, TestCount: 0.
    5. --rerun-failed --filter-type class ...: Success: false with the conflict message.
    6. --rerun-failed --test-mode PlayMode with no PlayMode record: Success: false with the no-record message.
    7. PlayMode with --respect-enter-play-mode-settings, no domain reload on Play entry: the class run records the failing test; --rerun-failed gives TestCount: 1, FailedCount: 1, RerunTargetCount: 1.
    8. The same with a domain reload on Play entry: the result is recovered after the reload (recovery Warning) and the record is rewritten with a later CompletedAt and the failing test; --rerun-failed gives TestCount: 1, FailedCount: 1 and no RerunTargetCount. Enter Play Mode settings were restored afterwards.
    9. A fixture whose OneTimeSetUp throws, with two tests: the class run fails both and the record holds only the fixture name; --rerun-failed gives TestCount: 2, RerunTargetCount: 1.
    10. Two identical [TestCase(1)] cases that fail: Unity reports the second as Fails(1 GeneratedTestCase2), and the record holds the single name Fails(1); --rerun-failed gives TestCount: 2, RerunTargetCount: 1.
  • check-skill-size, sync-tool-docs --check (no drift), CA1502 complexity at 15 (no findings), file length at 500 SLOC (no findings).
  • Go: no Go source changed. go test ./... in cli/common, cli/project-runner, and cli/release-automation passes locally except two IPC tests that the local sandbox blocks (a directory under /tmp and a Unix socket bind are denied); the build-cli job runs both of those packages and reports ok.
  • PR CI: every GitHub Actions check and CodeQL passed.

--rerun-failed has to know which tests failed in the most recent
completed run of a test mode, so each test mode gets one JSON record
under TestResults.

- Anything other than a complete record of the requested test mode reads
  as Unreadable, so a rerun never starts from a record it cannot trust.
- A failed delete propagates, so a run never starts while the old record
  could still pass as the newest one.
- Writes go through a uniquely named temporary file, so the real name
  never holds a partial record, and a write failure only logs a warning
  because the run's own result is still valid.
--rerun-failed needs, from a finished run, the names that make Unity's
testNames filter select every failed or inconclusive test again. The
response lists stay capped at ten, so the converter now collects a
separate, uncapped list.

- A suite that failed outside its tests (a thrown OneTimeSetUp or
  OneTimeTearDown) is rerun by its own name, using the same rule as
  FailedSuites; the root and a test assembly, whose names the filter
  never matches, are replaced by every leaf under them.
- The GeneratedTestCase suffix Unity adds to duplicate names is stripped,
  because the filter compares the unsuffixed NUnit name and a suffixed
  name would silently drop the failure from the rerun.
--rerun-failed reruns the recorded failures in one Test Runner call,
which the single-value filters cannot express. The new TestNames filter
passes every name to Unity's testNames, which ORs exact full-name
matches. It refuses an empty list because Unity reads empty testNames
as no filter and would run every test.
The parameter and the response shape come first so the use case can be
built against them. RerunTargetCount and RerunSourceCompletedAt appear
only on a --rerun-failed run, and NothingToRerun is a successful status
for a rerun whose recorded run had no failures, so an agent can tell
"nothing left to fix" from a failure without running anything.

The embedded tool catalog does not list RerunFailed yet; it is
regenerated together with the skill documentation.
run-tests now keeps the last-run record of each test mode current, and
--rerun-failed runs the tests that record names instead of the filter.

- The record is removed before a run starts and rewritten only once the
  run has a result, so a timed-out or cancelled run leaves no stale
  failures behind. Removal happens before pause points are cleared, so a
  record that cannot be removed stops the request with nothing changed.
- A --rerun-failed request that conflicts with a filter, finds no
  readable record, or has nothing to rerun returns before touching the
  record or the pause points.
- A rerun whose recorded tests are all gone says they were renamed or
  removed instead of giving the no-test-assembly advice.
- The use case now requires the record store, so tests cannot rewrite the
  project's own record by forgetting it.
A PlayMode run that respects Enter Play Mode settings can reload the
domain, which cancels the use case's await before it records the run.
The post-reload callback is then the run's only writer, so it now writes
the PlayMode record too.

The callback also receives the next run's RunFinished after a run that
was abandoned with its request still pending, and that run can be
EditMode, so it writes only when the result's root is a PlayMode run.
Adds the --rerun-failed row and usage note to the run-tests skill, the
RerunFailed property to the embedded tool catalog, and the rerun response
fields and NothingToRerun status to the field reference.

The response field list and the XML section move to
references/response-fields.md so the skill stays well under the 8,000-byte
injection cap with room for the new option; the skill keeps a pointer to
the fields agents usually need.

The catalog gains a property, which is a structural change to a shared
release input, so both components' shared-inputs stamps are refreshed.
Lists the new run-tests parameter and an example call in the English and
Japanese tool references, next to the other run-tests parameters.
@coderabbitai

coderabbitai Bot commented Oct 5, 2026 •

Copy link
Copy Markdown
Contributor

Review in Change Stack →

Navigate logical layers of code changes, visualize relationships, and explore their blast radius.

No actionable comments were generated in the recent review. 🎉

ℹ️ Recent review info
⚙️ Run configuration
  • Configuration used: Repository: hatayama/unity-cli-loop/.coderabbit.yaml
  • Review profile: CHILL
  • Plan: Advanced
  • Run ID: 5d7aede4-6445-4631-b827-4232946e50a5
📥 Commits

Reviewing files that changed from the base of the PR and between f3e5dfd and 915bfd8.

⛔ Files ignored due to path filters (6)
  • Assets/Tests/Editor/RunTestsLastRunRecordStoreTests.cs.meta is excluded by none and included by none
  • Assets/Tests/Editor/RunTestsTestNamesFilterTests.cs.meta is excluded by none and included by none
  • Packages/src/Editor/FirstPartyTools/RunTests/RunTestsLastRunRecord.cs.meta is excluded by none and included by none
  • Packages/src/Editor/FirstPartyTools/RunTests/RunTestsLastRunRecordStore.cs.meta is excluded by none and included by none
  • Packages/src/Editor/FirstPartyTools/RunTests/RunTestsRerunFailedResolver.cs.meta is excluded by none and included by none
  • Packages/src/Editor/FirstPartyTools/RunTests/Skill/references/response-fields.md.meta is excluded by none and included by none
📒 Files selected for processing (29)
  • .agents/skills/uloop-run-tests/SKILL.md
  • .agents/skills/uloop-run-tests/references/response-fields.md
  • .claude/skills/uloop-run-tests/SKILL.md
  • .claude/skills/uloop-run-tests/references/response-fields.md
  • Assets/Tests/Editor/PlayModeTestExecuterTests.cs
  • Assets/Tests/Editor/RunTestsLastRunRecordStoreTests.cs
  • Assets/Tests/Editor/RunTestsResponseContractTests.cs
  • Assets/Tests/Editor/RunTestsTestFrameworkResultTests.cs
  • Assets/Tests/Editor/RunTestsTestNamesFilterTests.cs
  • Assets/Tests/Editor/RunTestsUseCaseTests.cs
  • Packages/src/Documentation~/tools.md
  • Packages/src/Documentation~/tools_ja.md
  • Packages/src/Editor/FirstPartyTools/RunTests/RunTestsLastRunRecord.cs
  • Packages/src/Editor/FirstPartyTools/RunTests/RunTestsLastRunRecordStore.cs
  • Packages/src/Editor/FirstPartyTools/RunTests/RunTestsRerunFailedResolver.cs
  • Packages/src/Editor/FirstPartyTools/RunTests/RunTestsResponse.cs
  • Packages/src/Editor/FirstPartyTools/RunTests/RunTestsSchema.cs
  • Packages/src/Editor/FirstPartyTools/RunTests/RunTestsUseCase.cs
  • Packages/src/Editor/FirstPartyTools/RunTests/Skill/SKILL.md
  • Packages/src/Editor/FirstPartyTools/RunTests/Skill/references/response-fields.md
  • Packages/src/Editor/FirstPartyTools/RunTests/TestFramework/PlayModeTestExecuter.cs
  • Packages/src/Editor/FirstPartyTools/RunTests/TestFramework/RunTestsPendingRunCallback.cs
  • Packages/src/Editor/FirstPartyTools/RunTests/TestFramework/SerializableTestResultConverter.cs
  • Packages/src/Editor/FirstPartyTools/RunTests/TestRunner/SerializableTestResult.cs
  • Packages/src/Editor/FirstPartyTools/RunTests/TestRunner/TestExecutionFilter.cs
  • Packages/src/Editor/FirstPartyTools/RunTests/UnityCliLoopTestExecutionTypes.cs
  • cli/common/tools/default-tools.json
  • cli/dispatcher/shared-inputs-stamp.json
  • cli/project-runner/shared-inputs-stamp.json

Included review availability: This review used your included allowance. Your plan provides up to 4 included reviews per hour; 3 remain after this review.


📝 Walkthrough

Walkthrough

The run-tests tool adds --rerun-failed. Completed runs record rerun targets by test mode. The option resolves those targets into a test-name filter, and the response includes rerun metadata. The change also adds tests and documentation for the record lifecycle, target selection, and rerun outcomes.

Changes

Rerun failed tests

Layer / File(s) Summary
Persist completed test runs
Packages/src/Editor/FirstPartyTools/RunTests/RunTestsLastRunRecord.cs, Packages/src/Editor/FirstPartyTools/RunTests/RunTestsLastRunRecordStore.cs, Packages/src/Editor/FirstPartyTools/RunTests/TestFramework/RunTestsPendingRunCallback.cs, Assets/Tests/Editor/RunTestsLastRunRecordStoreTests.cs
Adds versioned per-mode run records with missing, found, and unreadable read statuses. The store validates, writes, replaces, and deletes records. The pending-run callback records recovered PlayMode results. Tests cover record validation, file operations, and recovery cases.
Collect and execute test-name targets
Packages/src/Editor/FirstPartyTools/RunTests/TestFramework/SerializableTestResultConverter.cs, Packages/src/Editor/FirstPartyTools/RunTests/TestRunner/SerializableTestResult.cs, Packages/src/Editor/FirstPartyTools/RunTests/TestRunner/TestExecutionFilter.cs, Packages/src/Editor/FirstPartyTools/RunTests/TestFramework/PlayModeTestExecuter.cs, Assets/Tests/Editor/RunTestsTestFrameworkResultTests.cs, Assets/Tests/Editor/RunTestsTestNamesFilterTests.cs, Assets/Tests/Editor/PlayModeTestExecuterTests.cs
Collects uncapped, deduplicated targets from failed and inconclusive results, including fixture-level failures. Adds validated test-name filters and maps them to Unity’s testNames filter. Tests cover target collection, name normalization, filter validation, and Unity filter mapping.
Resolve requests and return rerun results
Packages/src/Editor/FirstPartyTools/RunTests/RunTestsSchema.cs, Packages/src/Editor/FirstPartyTools/RunTests/RunTestsRerunFailedResolver.cs, Packages/src/Editor/FirstPartyTools/RunTests/RunTestsUseCase.cs, Packages/src/Editor/FirstPartyTools/RunTests/RunTestsResponse.cs, Packages/src/Editor/FirstPartyTools/RunTests/UnityCliLoopTestExecutionTypes.cs, cli/common/tools/default-tools.json, cli/dispatcher/shared-inputs-stamp.json, cli/project-runner/shared-inputs-stamp.json, Assets/Tests/Editor/RunTestsUseCaseTests.cs, Assets/Tests/Editor/RunTestsResponseContractTests.cs, Packages/src/Documentation~/tools.md, Packages/src/Documentation~/tools_ja.md, Packages/src/Editor/FirstPartyTools/RunTests/Skill/*, .agents/skills/uloop-run-tests/*, .claude/skills/uloop-run-tests/*
Adds RerunFailed request handling, rejects conflicting filters, and resolves records by test mode. The use case clears the requested mode’s record before execution and records completed results afterward. Responses include rerun metadata and a NothingToRerun status. Tests and documentation cover request constraints, record outcomes, response fields, and command usage.

Priority: ➖ Normal

Estimated code review effort: 4 (Complex) | ~45 minutes

Change: Feature

Sequence Diagram(s)

sequenceDiagram
  participant UseCase as RunTestsUseCase
  participant Resolver as RunTestsRerunFailedResolver
  participant Store as RunTestsLastRunRecordStore
  participant Filter as TestExecutionFilter
  participant Executor as PlayModeTestExecuter
  UseCase->>Resolver: Resolve request for test mode
  Resolver->>Store: Read recorded targets
  Store-->>Resolver: Return record
  Resolver->>Filter: Build test-name filter
  Filter-->>UseCase: Return filter
  UseCase->>Executor: Run selected tests
  Executor-->>UseCase: Return result
  UseCase->>Store: Record completed run
Loading

Merge Risk: ⚪ Minimal · up to 915bf

No actionable issue remains from the reviewed changes; the PR is mergeable after normal checks.

Security Architecture Review

Security architecture risk: 🔵 Low · up to 915bf

Reruns remain within the existing test-running capabilities. Invalid or unusable records stop execution rather than falling back to the full suite. The remaining uncertainty concerns recovery and concurrent execution behavior, not an identified privilege expansion.

Retained concerns
No architecture-level concerns identified.

Security review details

Security Blast Radius

  • inferred — A writer of the project record can influence which available tests a subsequent authorized rerun selects. In the inspected path, this does not grant execution beyond the existing ability to run all tests through the same tool and executors. Fixture and duplicate-name expansion affect selection precision, not privilege.

Trust Boundaries and Controls

  • observed — Supported-mode, framework, timeout, and execution-state checks precede rerun resolution. Missing, unreadable, malformed, wrong-version, wrong-mode, or blank-target records stop execution; an empty target list returns NothingToRerun. ByTestNames rejects empty input, preventing an empty selection from degrading into a full-suite run.

Resilience and Maintainability Implications

  • observed — Writes use a unique temporary file followed by replacement or move; file-access failures trigger best-effort cleanup and a warning. Recovery recording checks PlayMode identity and skips ExecutionFailed results. These controls contain partial-write and cross-mode recovery failures, although overlapping recovery callbacks were not runtime-verified.
🚥 Pre-merge checks | ✅ 4 | ❌ 1

❌ Failed checks (1 warning)

Check name Status Explanation Resolution
Docstring Coverage ⚠️ Warning Docstring coverage is 63.28% which is insufficient. The required threshold is 80.00%. Docstring coverage is scoped to functions touched by this diff. Analyzed 177 functions across 18 files. (11 skippe… Write docstrings for the functions missing them to satisfy the coverage threshold.
✅ Passed checks (4 passed)
Check name Status Explanation
Linked Issues check ✅ Passed Check skipped because no linked issues were found for this pull request.
Out of Scope Changes check ✅ Passed Check skipped because no linked issues were found for this pull request.
Title check ✅ Passed The title clearly identifies the main change: rerunning tests from the last run with --rerun-failed. It omits inconclusive tests, but it remains accurate and specific.
Description check ✅ Passed The description explains the rerun behavior, record handling, response changes, documentation, and verification. It is directly related to the changeset.
Full details: Docstring Coverage

Explanation

Docstring coverage is 63.28% which is insufficient. The required threshold is 80.00%. Docstring coverage is scoped to functions touched by this diff. Analyzed 177 functions across 18 files. (11 skipped: 11 unsupported.)

  • Fix all pre-merge checks with AI
✨ Finishing Touches 💡 1
📝 Generate docstrings 💡
  • Commit to this branch
  • Create a new PR
🧪 Generate unit tests (beta)
  • Commit to this branch
  • Create a new PR
  • Autopilot · Keep fixing CodeRabbit findings and required CI, and resolving merge conflicts

Thanks for using CodeRabbit! It's free for OSS, and your support helps us grow. If you like it, consider giving us a shout-out.

❤️ Share

Comment @coderabbitai help to get the list of available commands.

A rerun whose recorded tests were all renamed or removed already replaced
Message, but NoTestsFoundExplanation kept the default text that advises
adding a test assembly. The skill tells agents to read that field for
asmdef hints, so it sent them to create an .asmdef for tests that only
moved. The explanation now carries the same rerun-specific message.
The Editor-state validation saves or discards unsaved Scene and Prefab
Stage changes, and it ran before the filter and the --rerun-failed record
were resolved. A request that then stopped without running, such as a
rerun with nothing to rerun under --unsaved-changes discard, had already
thrown edits away. Resolving reads only, so it now comes first and every
rejected or empty request returns with the Editor untouched.

A request that is invalid on both counts now reports the filter or record
problem before the Editor-state problem.
@hatayama
hatayama merged commit 5671841 into main Oct 5, 2026
17 checks passed
@hatayama
hatayama deleted the feat/run-tests-rerun-failed branch October 5, 2026 23:14
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant