Skip to content

Promote dev → staging - #990

Merged
harrymove-ctrl merged 7 commits into
stagingfrom
dev
Sep 22, 2026
Merged

harrymove-ctrl merged 7 commits into
stagingfrom
dev

Conversation

@harrymove-ctrl

Copy link
Copy Markdown
Collaborator

Forward merge of dev into staging after #987 landed on dev.

Heads at open

Branch SHA
dev ee5fad65
staging 0a183481

3 commits on dev not in staging. git merge-tree reports 0 conflicts. Staging-only commits are prior Merge pull request … from MystenLabs/dev promotions (no unique content).

What’s in this hop

Commit / PR What
#987 (9ee664d6) SECURITY.md (Walrus-style → security@mystenlabs.com) + retarget issue templates off /security/advisories/new (404 when PVR is off)
#931 (06d75e2f) WALM-396 — say why a recall timed out (mcp/server/sdk)

Notes

Test plan

  • CI green on this PR.
  • After merge, staging /health (or Railway relayer) reflects new tip if a deploy fires.
  • Confirm SECURITY.md exists on origin/staging and issue template contact link is /security/policy.

nikola0x0 and others added 3 commits September 22, 2026 11:52
* feat(sdk): send recall's deadline so the relayer can say where it stalled (WALM-396)

recall() now puts deadline_ms (15000, the same number it aborts at) in the
signed body. A relayer that reads it answers a recall about to miss that
deadline with a 504 RECALL_TIMEOUT naming the stuck step. Older relayers
ignore the unknown field, so deploy order does not matter; a header would
have needed the CORS allow-list first.

* feat(server): answer a recall about to miss its caller's deadline with the stuck stage (WALM-396)

When a recall carries deadline_ms, the handler runs its pipeline under a
budget one second shorter (2s floor, input capped at 600s) and, if that
runs out, answers 504 RECALL_TIMEOUT naming the stage it was in: embed,
vector_search, walrus_download or seal_decrypt. The stage lives in a
task-local marker, so fetch_batch keeps its signature and analyze, which
shares it, is unaffected.

Callers that send no deadline (Python, older SDKs) run to completion as
before. For every caller, a recall dropped before it finishes logs the
stage it was in, which is where a caller hanging up used to leave nothing.

* fix(mcp): name the cause of a timed-out tool call and check relayer health (WALM-396)

A recall the SDK gave up on reached the agent as 'Tool error: This
operation was aborted'. wrapTool now sorts failures: a relayer
RECALL_TIMEOUT names the stuck step with advice for it; an SDK timeout or a
failed connect probes the relayer's /health (2s) and reports the result.
Each message carries Cause / Relayer health / Next step. Writes are never
told a retry is safe. Every other error keeps its old wording.

* fix(mcp): check relayer health before answering a call whose reply was lost (WALM-396)

A sent call past its deadline used to be answered 'did not answer this
call ... safe to retry', which cannot tell a dead relayer from a wrong URL
from one stuck call. The sweeper now asks the relayer's /health first (3s,
MEMWAL_MCP_HEALTH_PROBE_MS) and answers with Cause / Relayer health / Next
step; writes keep their no-blind-retry text with the health line added.

Bookkeeping stays synchronous: the answer is written only if the entry is
still the one in flight when the probe settles, so a late reply, a logout
or a shutdown that answers it meanwhile wins. memwal_recall gets its own
90s ceiling, so a lost recall reply is answered at 2 minutes, not 4.

* fix(mcp,server): harden the health probe and stop blaming a healthy relayer (WALM-396)

Review follow-ups:

- A non-integer or huge MEMWAL_MCP_HEALTH_PROBE_MS made AbortSignal.timeout
  throw outside the probe's try, and the sweeper's chain had no catch: the
  first lost reply crashed the bridge. The signal is now built inside the
  try (bridge and sidecar), the value is floored and capped at 60s, and the
  sweeper answers even if the probe rejects.
- On SDK 0.1.7 the SEAL session is built on the Sui fullnode before the
  recall request, under the same 15s clock. A failure or stall there was
  reported as the relayer's. The relayer is now named only when its own
  /health also fails; otherwise the message says the relayer answered and
  names the host that failed when Node reports it.
- The recall deadline is measured from request arrival, so auth and rate
  limiting come out of the caller's budget rather than the 1s margin.
- The bridge probes once per sweep, not once per expired call, and writes
  nothing after stdin has closed. A panic no longer logs as a hang-up.
- /health's write_ready=false and writes=paused show in the health line.
- 'This call only reads' becomes 'cannot store a duplicate', which is also
  true of memwal_restore.

* fix(server,mcp): answer a spent deadline at once and keep the failed-write report off it (WALM-396)

Review follow-ups from #931:

- The 2s floor applied after subtracting time already spent, so a recall
  whose deadline went on auth still ran 2s and answered a caller that had
  given up. budget_for now returns Unbounded | Run | Exhausted: the floor
  applies to the caller's deadline only, and a spent deadline is answered
  at once with stage 'auth'.
- The failed-write report was awaited inside the deadline, so a stalled
  courtesy lookup could discard a finished recall and blame the last stage.
  recall joins it after the pipeline, bounded by the same deadline.
- The bridge's answering callback had no catch; a throw left the call
  marked probing and unanswered. It now clears probing so the next sweep
  answers.
- 'safe to retry once that is resolved' read as 'retry once'.
- Code comments keep a one-line why; ticket ids and history stay in the
  changelog.

* fix(mcp,sdk): keep a call being answered off the replay, and give the 504 room (WALM-396)

The sweeper marks a sent call whose reply is lost, then answers it once
`/health` comes back — but left the entry in `inFlight` for the whole probe.
A reconnect landing in that window replayed it, so a write with no
idempotency key ran a second time while the agent was told it never ran.
The entry is now flagged `orphaned` when the sweeper takes it, and
`reconnect()` skips those; the flag is never cleared, so the path where the
answer itself throws cannot let a replay back in either.

`recall()` sends `deadline_ms: 14000` against its own 15s abort, derived from
it so the two cannot drift. The relayer's 1s margin runs from request
arrival, so it pays for the reply's trip back but not for the connect the
SDK's timer has already been counting — on a cold TLS handshake the 504
landed after the abort it exists to beat.

---------

Co-authored-by: Le Tien Phat <91601109+Niko1444@users.noreply.github.com>
Align SECURITY.md with MystenLabs/walrus (email security@mystenlabs.com).
Retarget issue templates from /security/advisories/new to /security/policy.
docs: add security policy and fix vulnerability-reporting 404
@railway-app
railway-app Bot temporarily deployed to Walrus Memory / dev September 22, 2026 12:05 Inactive
@harrymove-ctrl
harrymove-ctrl merged commit 8fa42bc into staging Sep 22, 2026
57 of 58 checks passed

This branch was successfully deployed

1 active and 1 inactive deployments
Walrus Memory / dev — f248bd37 Deployed Sep 22, 2026 by railway-app[bot]
benchmark-dev — f248bd37 Deployed Sep 22, 2026 by harrymove-ctrl via Memory API Latency #369
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

4 participants