Skip to content

Revive a dead plugin host without a call; forwarding follows the live generation - #314

Open
bittermandel wants to merge 3 commits into
mainfrom
plugin-host-revive-after-stall
Open

bittermandel wants to merge 3 commits into
mainfrom
plugin-host-revive-after-stall

Conversation

@bittermandel

@bittermandel bittermandel commented Oct 5, 2026 •

Copy link
Copy Markdown

Problem

A Lovable dev orb (oj dev on the lovable web app) came back from a VM snapshot restore with its guest starved: load average 26–49, threads in D state on kvm_async_pf. The plugin host's isolate didn't answer for 40 s, and run_hook's belt declared it gone:

oj: plugin host unresponsive for 40s running getWatchFiles (the engine stopped scheduling); treating the plugin host as gone (process rss 2969MB)
oj: watchChange failed for …/.wrangler/…sqlite-wal: plugin host exited

After that the web server stayed alive and kept accepting connections, but every request hung for more than an hour. Live capture from that process showed two defects:

  1. Revive is lazy. Only PluginHost::call() revives a dead host. The two calls right after the death fell inside the 5 s spacing (stamped by the addon keeper). After that, nothing called the host: requests are forwarded to its middleware port and never call it. No respawning line appeared until something unrelated issued a call more than an hour later.
  2. Forwarding stays pinned to the dead port. spawn_late_plugin_serve ran only when the middleware port was unknown at boot. In that orb the port was known at boot (:32959), so after the revive the new middleware on :37521 answered /health 200 in 0.125 s while :3000/health still hung, because it was forwarding to the dead :32959.

The 40 s belt fired on a real stall, not a clock jump. The guest clocksource is tsc, and CLOCK_BOOTTIME didn't advance across the sleeps (/proc/uptime 4800 s against dmesg's kvmclock-based sched_clock 6912 s at the same instant). The watchdog's time source is unchanged.

Change

  • Reviver (host.rs): each host gets a task, holding only a Weak ref, that respawns it once try_revive's spacing allows. The lifetime budget (3) and shutdown() are unchanged. try_revive now returns Live / Cooldown / Refused: the reviver retries after a cooldown, and after a refusal (orphaned napi addon, failed ignite, budget spent) it leaves the next call to retry.
  • Forwarding follows the live generation (lib.rs): spawn_late_plugin_serve always runs when a plugin host exists. It keys what it applied by engine generation, seeded from the generation captured around the boot serve_info(). A death or revive reset clears the middleware port and keeps the runner bit, so fallback renders still honour runner_dirty. Each respawn re-activates through the existing late-activation path (on_activate → runner_dirty, catch-up resync), and a death mid-resync cancels the resync. On terminal death it warns once and keeps watching, so a final respawn that is still booting gets applied. It exits on shutdown().
  • Body requests (start_dev.rs): with no middleware port, forward_with_body awaits ensure_runner_fresh before the runner serves the request.

Tests

Each test fails on origin/main 5264dfe and passes with this change. RED used the same test bodies; the only difference is that the new spawn_late_plugin_serve takes one extra argument, the boot generation.

Test Wrong implementation it catches RED on 5264dfe
a_dead_middleware_host_respawns_and_forwarding_follows_without_a_call Call-driven revive, or a revive that leaves PluginServe on the boot port. A per-generation response id avoids assuming a new port. forwarding never reached a respawned middleware (last answer None, revive attempts 0)
a_permanently_dead_host_stops_forwarding_to_its_port The observer exits on a spent budget without clearing forwarding. still forwarding to the dead Some(58033)
a_death_inside_respawn_spacing_revives_without_a_call The reviver treats a spacing cooldown as a refusal and parks. revived after the spacing, without a call: Elapsed(())

cargo test -p oj_server 189 passed, cargo test -p oj 110 passed, cargo clippy -p oj_server -p oj --all-targets -- -D warnings clean, cargo fmt --check clean (macOS arm64, rustc 1.95).

Not covered by a test: the DevServer::build wiring that always spawns the observer. It needs a full app fixture.

Not in this PR

Requests already in flight to a dead port still have no timeout on state.http. The monorepo also needs an oj pin bump in crates/oj/default.nix after release.

Follow-up commit 800a0cc: a respawn doesn't force a runner reload

On lynx, the first :3000 request after forwarding switched to a respawned middleware took 10–14 s (a timeout at 10 s, then 200 at +1.3 s and +4.2 s). Cause: the respawned generation's serve-info push counted as a late activation, because the port went from none to some. That ran on_activate, which set runner_dirty, so the next fallback request awaited a full runner reload. The observer keeps the runner bit through the down window, so the watcher's lazy path had already marked any edits from that window. Now the hook fires only when the runner bit wasn't already set; boot-time late activation is unchanged. Test reactivation_after_a_death_does_not_dirty_the_runner was RED on 6c1ec9e (hook fired once) and is GREEN after. The lynx re-measure of the first request is pending.

… generation

A plugin host declared gone after a long stall stayed dead: only a hook call
revived it, and a dev server whose requests go to the plugin middleware never
makes one. Forwarding also stayed pinned to the boot port, so even a later
revive left requests going to the dead generation.

- A per-host reviver respawns a dead host once spacing allows (lifetime budget
  and shutdown unchanged); try_revive reports Live / Cooldown / Refused.
- spawn_late_plugin_serve always runs with a plugin host, keyed by engine
  generation: a death clears the middleware port (runner bit kept), each
  respawn re-activates, a death mid-resync cancels the resync.
- forward_with_body awaits the runner refresh before serving a body request
  itself with no middleware.
@bittermandel
bittermandel requested a review from a team as a code owner October 5, 2026 23:09
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant