Repository navigation
Revive a dead plugin host without a call; forwarding follows the live generation - #314
Open
bittermandel wants to merge 3 commits into
Open
bittermandel wants to merge 3 commits into
bittermandel wants to merge 3 commits into
Conversation
… generation A plugin host declared gone after a long stall stayed dead: only a hook call revived it, and a dev server whose requests go to the plugin middleware never makes one. Forwarding also stayed pinned to the boot port, so even a later revive left requests going to the dead generation. - A per-host reviver respawns a dead host once spacing allows (lifetime budget and shutdown unchanged); try_revive reports Live / Cooldown / Refused. - spawn_late_plugin_serve always runs with a plugin host, keyed by engine generation: a death clears the middleware port (runner bit kept), each respawn re-activates, a death mid-resync cancels the resync. - forward_with_body awaits the runner refresh before serving a body request itself with no middleware.
…rough the down window
This was referenced Oct 6, 2026
Draft
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Problem
A Lovable dev orb (
oj devon the lovable web app) came back from a VM snapshot restore with its guest starved: load average 26–49, threads in D state onkvm_async_pf. The plugin host's isolate didn't answer for 40 s, andrun_hook's belt declared it gone:After that the web server stayed alive and kept accepting connections, but every request hung for more than an hour. Live capture from that process showed two defects:
PluginHost::call()revives a dead host. The two calls right after the death fell inside the 5 s spacing (stamped by the addon keeper). After that, nothing called the host: requests are forwarded to its middleware port and never call it. Norespawningline appeared until something unrelated issued a call more than an hour later.spawn_late_plugin_serveran only when the middleware port was unknown at boot. In that orb the port was known at boot (:32959), so after the revive the new middleware on:37521answered/health200 in 0.125 s while:3000/healthstill hung, because it was forwarding to the dead:32959.The 40 s belt fired on a real stall, not a clock jump. The guest clocksource is
tsc, andCLOCK_BOOTTIMEdidn't advance across the sleeps (/proc/uptime4800 s against dmesg's kvmclock-based sched_clock 6912 s at the same instant). The watchdog's time source is unchanged.Change
host.rs): each host gets a task, holding only aWeakref, that respawns it oncetry_revive's spacing allows. The lifetime budget (3) andshutdown()are unchanged.try_revivenow returnsLive / Cooldown / Refused: the reviver retries after a cooldown, and after a refusal (orphaned napi addon, failed ignite, budget spent) it leaves the next call to retry.lib.rs):spawn_late_plugin_servealways runs when a plugin host exists. It keys what it applied by engine generation, seeded from the generation captured around the bootserve_info(). A death or revive reset clears the middleware port and keeps the runner bit, so fallback renders still honourrunner_dirty. Each respawn re-activates through the existing late-activation path (on_activate→runner_dirty, catch-up resync), and a death mid-resync cancels the resync. On terminal death it warns once and keeps watching, so a final respawn that is still booting gets applied. It exits onshutdown().start_dev.rs): with no middleware port,forward_with_bodyawaitsensure_runner_freshbefore the runner serves the request.Tests
Each test fails on
origin/main5264dfe and passes with this change. RED used the same test bodies; the only difference is that the newspawn_late_plugin_servetakes one extra argument, the boot generation.a_dead_middleware_host_respawns_and_forwarding_follows_without_a_callPluginServeon the boot port. A per-generation response id avoids assuming a new port.forwarding never reached a respawned middleware (last answer None, revive attempts 0)a_permanently_dead_host_stops_forwarding_to_its_portstill forwarding to the dead Some(58033)a_death_inside_respawn_spacing_revives_without_a_callrevived after the spacing, without a call: Elapsed(())cargo test -p oj_server189 passed,cargo test -p oj110 passed,cargo clippy -p oj_server -p oj --all-targets -- -D warningsclean,cargo fmt --checkclean (macOS arm64, rustc 1.95).Not covered by a test: the
DevServer::buildwiring that always spawns the observer. It needs a full app fixture.Not in this PR
Requests already in flight to a dead port still have no timeout on
state.http. The monorepo also needs an oj pin bump incrates/oj/default.nixafter release.Follow-up commit 800a0cc: a respawn doesn't force a runner reload
On lynx, the first
:3000request after forwarding switched to a respawned middleware took 10–14 s (a timeout at 10 s, then 200 at +1.3 s and +4.2 s). Cause: the respawned generation's serve-info push counted as a late activation, because the port went from none to some. That ranon_activate, which setrunner_dirty, so the next fallback request awaited a full runner reload. The observer keeps the runner bit through the down window, so the watcher's lazy path had already marked any edits from that window. Now the hook fires only when the runner bit wasn't already set; boot-time late activation is unchanged. Testreactivation_after_a_death_does_not_dirty_the_runnerwas RED on 6c1ec9e (hook fired once) and is GREEN after. The lynx re-measure of the first request is pending.