From 0ae29a7af45f21574855a42be0eca0647c05fcf4 Mon Sep 17 00:00:00 2001 From: Samuel Jenness Date: Sat, 1 Aug 2026 07:47:44 -0400 Subject: [PATCH] Announce a task that exhausts its restart cap MAX_RESTARTS is the last thing between a pathological task and an unbounded requeue loop, and reaching it says something definite: the doctor has stopped intervening, and the task will hold its slot until walltime producing nothing. It said so with one lowercase `restarts exhausted; leaving it` in the middle of a verbose sweep. That is the same shape of silent ending as a TIME_LIMIT kill mailed to nobody, which is what this tooling exists because of. Seen again on a fifth campaign, running the pre-2.8.3 classifier. swfcalib task 41767827_48 wedged in PSOCK worker startup and was requeued three times. Segmenting its accumulated log by attempt, since SLURM appends across requeues, gives 3 of 8 workers through package loading on the first attempt, 3 of 8 on the second, 1 of 8 on the third, then all 8 on the fourth, which ran 1.42 hours and completed. About 4.5 hours to deliver one 95-minute batch, and nothing surfaced it except reading the sweep log by hand afterwards. The cap now emits a `!! EXHAUSTED` token with the task, its classification, the node and the restart count; the sweep summary gains an `exhausted=` counter; and MAIL_TO, when set, receives one message per newly exhausted task. Once per task, not once per sweep. The doctor re-probes every ten minutes and an exhausted task stays exhausted, so an undeduplicated alert would repeat until walltime and train the reader to skip it, which would leave the state as silent as it was. The campaign-scoped ledger already in STATE_FILE is the dedup key, and its `exhausted ` rows cannot collide with the node-offense rows, which are matched on a leading node name. Behaviour is unchanged for every task below the cap, and the escalation is report-only in the sense that matters: it still declines to requeue. --- DESCRIPTION | 2 +- NEWS.md | 8 ++++++++ inst/hpc_doctor/README.md | 8 ++++++++ inst/hpc_doctor/degen_watch.sh | 29 ++++++++++++++++++++++++++--- 4 files changed, 43 insertions(+), 4 deletions(-) diff --git a/DESCRIPTION b/DESCRIPTION index 1f8ba0a..2f59500 100644 --- a/DESCRIPTION +++ b/DESCRIPTION @@ -1,5 +1,5 @@ Package: EpiModelHPC -Version: 2.9.1 +Version: 2.9.2 Date: 2026-07-28 Title: EpiModel Extensions for High-Performance Computing Description: Extension package to EpiModel to run large-scale stochastic network diff --git a/NEWS.md b/NEWS.md index f5e7045..11bb9d3 100644 --- a/NEWS.md +++ b/NEWS.md @@ -1,3 +1,11 @@ +# EpiModelHPC 2.9.2 + +## NEW FEATURES + +- `degen_watch.sh` announces a task that exhausts `MAX_RESTARTS` instead of retiring it quietly. The cap is a terminal state: the task is confirmed pathological, the doctor stops intervening, and the task then holds its slot until walltime producing nothing. It previously said so with one lowercase `restarts exhausted; leaving it` in the middle of a verbose sweep, which is the same shape of silent ending as a TIME_LIMIT kill mailed to nobody. The line is now a `!! EXHAUSTED` token carrying the task, classification, node and restart count; the sweep summary gains an `exhausted=` counter; and setting `MAIL_TO` sends one message per newly exhausted task. Alerts are emitted once per task rather than once per sweep, deduplicated through the campaign-scoped ledger already in `STATE_FILE`, since the doctor re-probes every ten minutes and an exhausted task stays exhausted. + + Observed on a `swfcalib` campaign that ran the pre-2.8.3 classifier: task `41767827_48` wedged in PSOCK worker startup three times, getting 3, 3 and 1 of its 8 workers through package loading before each requeue, and took about 4.5 hours to deliver one 95-minute batch. Nothing surfaced that except reading the sweep log by hand. + # EpiModelHPC 2.9.1 ## BUG FIXES diff --git a/inst/hpc_doctor/README.md b/inst/hpc_doctor/README.md index fe2e3bf..b535760 100644 --- a/inst/hpc_doctor/README.md +++ b/inst/hpc_doctor/README.md @@ -153,6 +153,14 @@ Three properties made it hard to see, and each is now addressed: Only `MAX_RESTARTS` stopped it: after three rounds the tasks reached their restart cap, the doctor logged `restarts exhausted; leaving it`, and the array was finally allowed to finish. That cap was doing work it was never meant to do. +### Exhausting the restart cap is a terminal state and now reads like one + +The cap is the last thing standing between a pathological task and an unbounded requeue loop, and reaching it says something definite: the doctor has given up, and the task will hold its slot until walltime producing nothing. It used to say that with one lowercase line in the middle of a verbose sweep, which is the same shape of silent ending as a TIME_LIMIT kill mailed to nobody. + +Seen again on a fifth campaign, this time running the pre-2.8.3 classifier: `swfcalib` task `41767827_48` wedged in PSOCK worker startup and was requeued three times. Segmenting its accumulated log by attempt (SLURM appends across requeues) shows 3 of 8 workers through package loading on the first attempt, 3 of 8 on the second, 1 of 8 on the third, then all 8 on the fourth, which ran 1.42 hours and completed. About 4.5 hours to deliver one 95-minute batch, and nothing surfaced it except reading the sweep log by hand afterwards. + +So the cap now emits a `!! EXHAUSTED` token carrying the task, its classification, the node and the restart count; the sweep summary gains an `exhausted=` counter alongside the other two; and `MAIL_TO`, if set, receives one message per newly exhausted task. The alert fires once per task rather than once per sweep. The doctor re-probes every ten minutes and an exhausted task stays exhausted, so an undeduplicated alert would repeat until walltime and train the reader to skip it. The campaign-scoped ledger in `STATE_FILE` is the dedup key, and its `exhausted ` rows cannot collide with the node-offense rows, which are matched on a leading node name. + The probe now reads `SLURM_ARRAY_JOB_ID` and `SLURM_ARRAY_TASK_ID` from the same environment block and emits `taskid=`, and every `squeue`/`scontrol` call in `degen_watch.sh` addresses the task by it. Taking the id from the task's own environment rather than resolving it later is the point: PSOCK workers inherit it from their master, so every process of a task agrees, and nothing has to be inferred from a `squeue` line that may describe a sibling. Verified against the successor array while it ran: `jobid=41746875 taskid=41746874_0`, where the raw `SLURM_JOB_ID` and the array id differ by one and neither is guessable from the other. A guard before the destructive call refuses any unqualified id that `scontrol` reports as an array, so the failure mode cannot return by another route. Two smaller consequences worth keeping in mind. The probe match was a bare `grep jobid=$jid`, which substring-matches any longer id sharing the prefix; it is now anchored on the field boundary, because acting on the wrong task is the one mistake this script must not make. And `Restarts=`, the age lookup, and the `ExcNodeList` read all went through the same unqualified id, so on an array they were reading whichever task `scontrol` or `squeue` happened to list first rather than the one under judgement. diff --git a/inst/hpc_doctor/degen_watch.sh b/inst/hpc_doctor/degen_watch.sh index f3bc14d..f40969f 100755 --- a/inst/hpc_doctor/degen_watch.sh +++ b/inst/hpc_doctor/degen_watch.sh @@ -215,7 +215,7 @@ done echo "--- strike 1 done; re-checking suspects in ${RECHECK}s (transient dips must not trigger) ---" sleep "$RECHECK" -acted=0; cleared=0 +acted=0; cleared=0; exhausted=0 for s in $suspects; do IFS=: read -r jid node tid <<< "$s" tid=${tid:-$jid} @@ -353,7 +353,30 @@ for s in $suspects; do fi [ "$REQUEUE" = "1" ] || continue - if [ "$restarts" -ge "$MAX_RESTARTS" ]; then echo " restarts exhausted; leaving it"; continue; fi + # Exhausting MAX_RESTARTS is a terminal state and has to read like one. The task + # is confirmed pathological, the doctor has stopped intervening, and it will now + # hold its slot until walltime producing nothing. Before this it announced + # itself with one lowercase line in the middle of a verbose sweep, which is the + # same shape of silent ending as a TIME_LIMIT kill mailed to nobody. + # + # Emitted once per task, not once per sweep: the doctor re-probes every 10 + # minutes and an exhausted task stays exhausted, so an undeduplicated alert + # would repeat until walltime and train the reader to skip it. The campaign + # ledger already in STATE_FILE is the natural dedup key, and is campaign-scoped + # for free because the filename carries the doctor's own job id. + if [ "$restarts" -ge "$MAX_RESTARTS" ]; then + if ! grep -qxF "exhausted $tid" "$STATE_FILE" 2>/dev/null; then + echo "exhausted $tid" >> "$STATE_FILE" 2>/dev/null || true + echo " !! EXHAUSTED $tid: $kind on $node after $restarts restarts (MAX_RESTARTS=$MAX_RESTARTS); no longer intervening, it will hold its slot until walltime" + exhausted=$((exhausted+1)) + if [ -n "${MAIL_TO:-}" ] && command -v mail >/dev/null 2>&1; then + printf 'task %s exhausted %s restarts and is still %s on %s.\nThe doctor has stopped intervening; the task will hold its slot until walltime.\nDoctor job: %s\nCampaign pattern: /%s/\n' \ + "$tid" "$restarts" "$kind" "$node" "${SLURM_JOB_ID:-manual}" "$PATTERN" \ + | mail -s "deploy doctor: $tid exhausted its restarts" "$MAIL_TO" 2>/dev/null || true + fi + fi + continue + fi # Last line of defence before the destructive call. An unqualified ArrayJobId # would requeue the whole array, and the failure is silent: the sibling tasks # simply vanish from the next probe and get logged as "gone (finished/moved)". @@ -380,5 +403,5 @@ for s in $suspects; do done echo "---" -echo "confirmed_starved_requeued=$acted cleared_as_transient=$cleared" +echo "confirmed_starved_requeued=$acted cleared_as_transient=$cleared exhausted=$exhausted" [ "$REQUEUE" = "1" ] || echo "report-only; pass --requeue to act"