Repository navigation
feat: Rerun only the tests that failed in the last run with run-tests --rerun-failed - #3161
Conversation
--rerun-failed has to know which tests failed in the most recent completed run of a test mode, so each test mode gets one JSON record under TestResults. - Anything other than a complete record of the requested test mode reads as Unreadable, so a rerun never starts from a record it cannot trust. - A failed delete propagates, so a run never starts while the old record could still pass as the newest one. - Writes go through a uniquely named temporary file, so the real name never holds a partial record, and a write failure only logs a warning because the run's own result is still valid.
--rerun-failed needs, from a finished run, the names that make Unity's testNames filter select every failed or inconclusive test again. The response lists stay capped at ten, so the converter now collects a separate, uncapped list. - A suite that failed outside its tests (a thrown OneTimeSetUp or OneTimeTearDown) is rerun by its own name, using the same rule as FailedSuites; the root and a test assembly, whose names the filter never matches, are replaced by every leaf under them. - The GeneratedTestCase suffix Unity adds to duplicate names is stripped, because the filter compares the unsuffixed NUnit name and a suffixed name would silently drop the failure from the rerun.
--rerun-failed reruns the recorded failures in one Test Runner call, which the single-value filters cannot express. The new TestNames filter passes every name to Unity's testNames, which ORs exact full-name matches. It refuses an empty list because Unity reads empty testNames as no filter and would run every test.
The parameter and the response shape come first so the use case can be built against them. RerunTargetCount and RerunSourceCompletedAt appear only on a --rerun-failed run, and NothingToRerun is a successful status for a rerun whose recorded run had no failures, so an agent can tell "nothing left to fix" from a failure without running anything. The embedded tool catalog does not list RerunFailed yet; it is regenerated together with the skill documentation.
run-tests now keeps the last-run record of each test mode current, and --rerun-failed runs the tests that record names instead of the filter. - The record is removed before a run starts and rewritten only once the run has a result, so a timed-out or cancelled run leaves no stale failures behind. Removal happens before pause points are cleared, so a record that cannot be removed stops the request with nothing changed. - A --rerun-failed request that conflicts with a filter, finds no readable record, or has nothing to rerun returns before touching the record or the pause points. - A rerun whose recorded tests are all gone says they were renamed or removed instead of giving the no-test-assembly advice. - The use case now requires the record store, so tests cannot rewrite the project's own record by forgetting it.
A PlayMode run that respects Enter Play Mode settings can reload the domain, which cancels the use case's await before it records the run. The post-reload callback is then the run's only writer, so it now writes the PlayMode record too. The callback also receives the next run's RunFinished after a run that was abandoned with its request still pending, and that run can be EditMode, so it writes only when the result's root is a PlayMode run.
Adds the --rerun-failed row and usage note to the run-tests skill, the RerunFailed property to the embedded tool catalog, and the rerun response fields and NothingToRerun status to the field reference. The response field list and the XML section move to references/response-fields.md so the skill stays well under the 8,000-byte injection cap with room for the new option; the skill keeps a pointer to the fields agents usually need. The catalog gains a property, which is a structural change to a shared release input, so both components' shared-inputs stamps are refreshed.
Lists the new run-tests parameter and an example call in the English and Japanese tool references, next to the other run-tests parameters.
|
Navigate logical layers of code changes, visualize relationships, and explore their blast radius. No actionable comments were generated in the recent review. 🎉 ℹ️ Recent review info⚙️ Run configuration
⛔ Files ignored due to path filters (6)
📒 Files selected for processing (29)
Included review availability: This review used your included allowance. Your plan provides up to 4 included reviews per hour; 3 remain after this review. 📝 WalkthroughWalkthroughThe ChangesRerun failed tests
Priority: ➖ Normal Estimated code review effort: 4 (Complex) | ~45 minutes Change: Feature Sequence Diagram(s)sequenceDiagram
participant UseCase as RunTestsUseCase
participant Resolver as RunTestsRerunFailedResolver
participant Store as RunTestsLastRunRecordStore
participant Filter as TestExecutionFilter
participant Executor as PlayModeTestExecuter
UseCase->>Resolver: Resolve request for test mode
Resolver->>Store: Read recorded targets
Store-->>Resolver: Return record
Resolver->>Filter: Build test-name filter
Filter-->>UseCase: Return filter
UseCase->>Executor: Run selected tests
Executor-->>UseCase: Return result
UseCase->>Store: Record completed run
Merge Risk: ⚪ Minimal · up to No actionable issue remains from the reviewed changes; the PR is mergeable after normal checks. Security Architecture ReviewSecurity architecture risk: 🔵 Low · up to Reruns remain within the existing test-running capabilities. Invalid or unusable records stop execution rather than falling back to the full suite. The remaining uncertainty concerns recovery and concurrent execution behavior, not an identified privilege expansion. Retained concerns Security review detailsSecurity Blast Radius
Trust Boundaries and Controls
Resilience and Maintainability Implications
🚥 Pre-merge checks | ✅ 4 | ❌ 1❌ Failed checks (1 warning)
✅ Passed checks (4 passed)
Full details: Docstring CoverageExplanation Docstring coverage is 63.28% which is insufficient. The required threshold is 80.00%. Docstring coverage is scoped to functions touched by this diff. Analyzed 177 functions across 18 files. (11 skipped: 11 unsupported.)
✨ Finishing Touches 💡 1📝 Generate docstrings 💡
🧪 Generate unit tests (beta)
Thanks for using CodeRabbit! It's free for OSS, and your support helps us grow. If you like it, consider giving us a shout-out. Comment |
A rerun whose recorded tests were all renamed or removed already replaced Message, but NoTestsFoundExplanation kept the default text that advises adding a test assembly. The skill tells agents to read that field for asmdef hints, so it sent them to create an .asmdef for tests that only moved. The explanation now carries the same rerun-specific message.
The Editor-state validation saves or discards unsaved Scene and Prefab Stage changes, and it ran before the filter and the --rerun-failed record were resolved. A request that then stopped without running, such as a rerun with nothing to rerun under --unsaved-changes discard, had already thrown edits away. Resolving reads only, so it now comes first and every rejected or empty request returns with the Editor untouched. A request that is invalid on both counts now reports the filter or record problem before the Editor-state problem.
Summary
uloop run-tests --rerun-failedreruns only the tests that failed or were inconclusive in the most recent completed run of the same test mode, in a single run.User Impact
--filter-type exactone at a time, or rerunning the whole suite.--rerun-failedreruns exactly the recorded failures (the whole fixture when itsOneTimeSetUporOneTimeTearDownfailed). When there is nothing it can safely rerun, it says so and runs nothing instead of falling back to the full suite.Behavior
.uloop/outputs/TestResults/last-run-<TestMode>.json. The record is removed when a run starts and written only once the run has a result, so a run that times out or is cancelled leaves no stale failures behind.--rerun-failedreturns without saving or discarding unsaved Scene and Prefab Stage changes, clearing pause points, or running tests when:--filter-typeor--filter-value(ExecutionFailed)ExecutionFailed)ExecutionFailed)NothingToRerunwithSuccess: true--rerun-failedthe record, before the Editor-state check saves or discards unsaved changes. A request with an invalid filter therefore also returns with the Editor untouched, and a request that is invalid on both counts reports the filter or record problem before the Editor-state problem.RerunTargetCountandRerunSourceCompletedAt. If every recorded test has since been renamed or removed,MessageandNoTestsFoundExplanationsay that instead of giving the missing-test-assembly advice.--respect-enter-play-mode-settings, a PlayMode run that reloads the domain on Play entry is recovered after the reload. That path also writes the PlayMode record, but its response carries no rerun fields; the skill's response reference documents this.GeneratedTestCase<n>) is stripped from recorded names so the filter matches them; a failing assembly or root suite is recorded as its leaves, because those suite names cannot be used as a filter.Out of scope:
--retries--rerun-failedagain after a failure already tells a flaky test from a real failure.Changes
testNames; the run-tests use case resolves a rerun before changing anything and records each completed run; the post-reload callback records recovered PlayMode runs.--rerun-failedrow and usage note; the response field list moves toreferences/response-fields.md(the run-tests SKILL.md shrinks from 7,894 to 5,716 bytes) with the new fields and status;RerunFailedin the embedded tool catalog; tool reference docs (English and Japanese).SerializableTestResultConverter.csis now at 484 of the 500 allowed SLOC; split it before adding behavior there.Verification
RunTestsUseCaseTests,RunTestsLastRunRecordStoreTests,RunTestsTestFrameworkResultTests,PlayModeTestExecuterTests,RunTestsTestNamesFilterTests,RunTestsResponseContractTests,DefaultToolsCatalogDriftTests: 153/153 passed on the final head. The 21 new use-case cases cover completed, timed-out, and cancelled runs (before start, during the run, during the cleanup wait), a record that cannot be cleared (no run, pause points untouched), the other test mode's record, a run without a result tree, NoTestsFound, filter conflicts (including whitespace-only and null values), missing, unreadable, and empty records, a rerun, recorded tests that are gone, test-mode selection, a rerun that times out, and a rerun stopped by the Editor-state check (record and pause points untouched). Every request that stops early is checked not to reach the Editor-state check.FailedCount: 1; the record lists only that test.--rerun-failed:TestCount: 1,RerunTargetCount: 1,FailedCount: 1.--rerun-failed:Success: true,TestCount: 1.--rerun-failedagain:Status: NothingToRerun,TestCount: 0.--rerun-failed --filter-type class ...:Success: falsewith the conflict message.--rerun-failed --test-mode PlayModewith no PlayMode record:Success: falsewith the no-record message.--respect-enter-play-mode-settings, no domain reload on Play entry: the class run records the failing test;--rerun-failedgivesTestCount: 1,FailedCount: 1,RerunTargetCount: 1.Warning) and the record is rewritten with a laterCompletedAtand the failing test;--rerun-failedgivesTestCount: 1,FailedCount: 1and noRerunTargetCount. Enter Play Mode settings were restored afterwards.OneTimeSetUpthrows, with two tests: the class run fails both and the record holds only the fixture name;--rerun-failedgivesTestCount: 2,RerunTargetCount: 1.[TestCase(1)]cases that fail: Unity reports the second asFails(1 GeneratedTestCase2), and the record holds the single nameFails(1);--rerun-failedgivesTestCount: 2,RerunTargetCount: 1.check-skill-size,sync-tool-docs --check(no drift), CA1502 complexity at 15 (no findings), file length at 500 SLOC (no findings).go test ./...incli/common,cli/project-runner, andcli/release-automationpasses locally except two IPC tests that the local sandbox blocks (a directory under/tmpand a Unix socket bind are denied); thebuild-clijob runs both of those packages and reportsok.