Background
.github/workflows/community-ci.yml's run_on_instance() wraps the
entire python3 -m tools.brev_exec call (which runs every selected
family's container/tests, e.g. all ~130 families for scope=all) in
retry_backoff() — 6 attempts, capped exponential backoff up to 160s.
This was designed to survive transient network blips reaching
github.com/nvcr.io mid-run.
Observed in Dev Community GPU CI run 36962671513 (PR #1575): all
families that ran passed (e.g. whisper), but the newly-added family
ltx2 genuinely failed:
ERROR: Community GPU family failures: ltx2: checkpoint staging exited 1
Since tools/community_gpu_ci.py exits non-zero when any family
fails, retry_backoff could not distinguish this deterministic,
non-transient content bug from a network blip, and retried the entire
multi-family run 6 times (~8 minutes of wasted retry delay, plus
however long each full re-run takes) before finally giving up and
recreating the instance — none of which could ever fix a genuine
checkpoint-staging bug in one family's code.
Problem
retry_backoff cannot currently distinguish:
- A step that failed before any family ran (setup/network issue —
worth retrying)
- A specific family's test genuinely, deterministically failing
(content bug — retrying is pure waste, and delays the useful signal
of "this PR's code is broken")
Suggested direction (not yet implemented)
- Have
tools/community_gpu_ci.py or tools/brev_exec.py distinguish
these failure classes in its exit code or output (e.g. a specific
exit code for "N family failures, 0 setup/network errors"), so the
workflow can skip retrying when the failure is already known to be a
genuine family failure.
- Alternatively, inspect the result file for a "family failed cleanly"
marker before deciding to retry.
Not in scope here
Background
.github/workflows/community-ci.yml'srun_on_instance()wraps theentire
python3 -m tools.brev_execcall (which runs every selectedfamily's container/tests, e.g. all ~130 families for
scope=all) inretry_backoff()— 6 attempts, capped exponential backoff up to 160s.This was designed to survive transient network blips reaching
github.com/nvcr.io mid-run.
Observed in Dev Community GPU CI run 36962671513 (PR #1575): all
families that ran passed (e.g.
whisper), but the newly-added familyltx2genuinely failed:Since
tools/community_gpu_ci.pyexits non-zero when any familyfails,
retry_backoffcould not distinguish this deterministic,non-transient content bug from a network blip, and retried the entire
multi-family run 6 times (~8 minutes of wasted retry delay, plus
however long each full re-run takes) before finally giving up and
recreating the instance — none of which could ever fix a genuine
checkpoint-staging bug in one family's code.
Problem
retry_backoffcannot currently distinguish:worth retrying)
(content bug — retrying is pure waste, and delays the useful signal
of "this PR's code is broken")
Suggested direction (not yet implemented)
tools/community_gpu_ci.pyortools/brev_exec.pydistinguishthese failure classes in its exit code or output (e.g. a specific
exit code for "N family failures, 0 setup/network errors"), so the
workflow can skip retrying when the failure is already known to be a
genuine family failure.
marker before deciding to retry.
Not in scope here
ltx2's actualcheckpoint staging bug (that's PR feat(ltx2): add the LTX-2.5 text-to-audio-video family with context parallelism #1575's own problem to fix).
fix(ci): probe SSH right after instance creation to catch false-success provisioning #1569, fix(ci): longer capped backoff for content-fetch retries, keep fast probe for dead instances #1571, fix(ci): give AWS-provisioned instances a longer SSH probe budget #1573).