Skip to content

Draft-20 backlog: stream-credit refusal and shared TRACK_STATUS rounds - #118

Merged
floatdrop merged 4 commits into
draft-20from
draft20-backlog-relay-short
Sep 28, 2026
Merged

floatdrop merged 4 commits into
draft-20from
draft20-backlog-relay-short

Conversation

@floatdrop

Copy link
Copy Markdown
Owner

A short relay batch from the draft-20 backlog. Every commit was reviewed by moqt-reviewer, and each has tests that were seen failing first, either before the change or under a mutant that removes the behaviour.

Commits

  • 4725aff Stream credit gets EXCESSIVE_LOAD. When the relay can't open a SUBSCRIBE or forwarded TRACK_STATUS to a live publisher for lack of bidi-stream credit, the subscriber now gets EXCESSIVE_LOAD with the relay's ~1s jittered retry (§10.6.2). It was DOES_NOT_EXIST with Retry Interval 0 ("SHOULD NOT be retried"). It ranks as "any other refusal", above another candidate's GOING_AWAY, TIMEOUT or DOES_NOT_EXIST.
  • d1e8172 Concurrent forwarded TRACK_STATUS requests share one round. Requests for a track that arrive while a round is in flight wait for it:
    • each gets its own INCLUDE_PROPERTIES and counts against its own session's cap;
    • one requester's STOP_SENDING stops only its own wait;
    • a request from a session the relay is asking right now (possibly its own request routed back, §6.2) runs its own round, under the loop guard.
  • 5209e74 The round is relay work. It starts through relayGo from the requesting handler, so Stop joins it. It runs on a relay context that Stop cancels, so a round stuck in Discovery or waiting on an upstream is cut short. This replaces the first version's untracked singleflight goroutine, which could outlive Stop.
  • c7b434f Rounds end after the grace period. They now end at Stop's step 6, with the upstream sessions, after the downstream GOAWAY and grace (§3.6). Cancelling them as Stop began turned publisher answers given within the grace into DOES_NOT_EXIST.

Decisions

  • Maintainer decisions:
    • EXCESSIVE_LOAD with a retry, ranked as "any other refusal";
    • share one round among concurrent requests; the round runs to its bound whoever cancels.

Known gaps (added to STATUS.md)

  • A request joining a round late gets only the candidates known when the round started.
  • A TRACK_STATUS loop through two different sessions isn't detected; its round waits out the 5s bound, as before.

Verification

  • go test ./... and golangci-lint run pass.
  • go test -race ./pkg/relay/... passes.
  • go mod tidy -diff is clean.

🤖 Generated with Claude Code

floatdrop and others added 4 commits September 28, 2026 10:31
…h EXCESSIVE_LOAD (§10.6.2)

A SUBSCRIBE, or forwarded TRACK_STATUS, that the relay could not open to a
live publisher for want of bidi-stream credit was answered DOES_NOT_EXIST
with Retry Interval 0: "not available at the publisher", and "SHOULD NOT be
retried". But the publisher is there, and the relay only "cannot process
the request at this time".

upstreamRejection now answers it EXCESSIVE_LOAD with the relay's jittered
~1s retry, the code it already uses for its own caps. The maintainer chose
this. It ranks as "any other refusal", so it outranks another candidate's
GOING_AWAY, TIMEOUT or DOES_NOT_EXIST. candidateRetry reports its Retry
Interval so ties compare what is sent.

TestRendezvous_NoStreamCreditEndsHold now expects EXCESSIVE_LOAD with a
retry. TestSubscribe_NoStreamCreditOutranksNoPublisherYet covers both orders
against a GOING_AWAY, and forwarded TRACK_STATUS. All failed before this
change.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
…requests for a track

Each TRACK_STATUS for a track with no Established subscription forwarded
its own round upstream, so N concurrent requests made N upstream ones to
every candidate (§13.1 amplification). The maintainer chose to share rounds.

forwardTrackStatus now runs trackStatusUpstream through a relay-wide
singleflight.Group, keyed by track:
- requests for a track that arrive while a round is in flight wait for it;
- each requester is answered with its own INCLUDE_PROPERTIES and counts
  against its own session's cap;
- the round is bounded by its 5s timeout, not by any one requester;
- a requester's STOP_SENDING stops only its own wait, freeing its slot.

A request from a session the relay is asking about the track right now may
be the relay's own TRACK_STATUS routed back (§6.2). It does not join a round
that waits on it: it runs its own, where the loop guard skips that session.

TestTrackStatus_ConcurrentRequestsShareOneRound failed before this change
(the publisher was asked twice). TestTrackStatus_CancelReleasesRequester
replaces TestTrackStatus_CancelEndsUpstream, since a requester's cancel no
longer ends a round others may share. TestRelay_TrackStatusLoopStopsAtSecondHop
now bounds the answer time, and fails, deadlocked for the full bound, when
a routed-back request joins the round.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
…hort and joins

d1e8172 ran the shared round on singleflight's own goroutine, on a context
detached from any session. Relay.Stop neither ended nor joined it, so it
could return while the round was still in Discovery. Calling relayGo from
that goroutine instead could race Stop's handlers.Wait.

forwardTrackStatus now keeps its own map of rounds in flight (golang.org/x/sync
is no longer imported by the relay):
- a round starts through relayGo from the requesting handler's own
  goroutine, so Stop joins it;
- it runs on a relay-scoped context that Stop cancels first, so a round
  waiting on Discovery or an upstream is cut short;
- requesters wait for its result or their own STOP_SENDING, as before.
A test hook marks each request that has joined a round.

Tests, each red under a mutant removing the behaviour:
- TestTrackStatus_StopEndsRound: a round blocked in Discovery ends before
  Stop returns, within 2s;
- TestTrackStatus_LeaderCancelKeepsSharedRound: the requester that started
  a round cancelling does not end it for another;
- TestTrackStatus_ConcurrentRequestsShareOneRound: now waits on the hook
  rather than sleeping.

Also: the forwardTrackStatus doc no longer implies it catches a loop through
other sessions. STATUS.md records that, and that a request joining a round
late gets the candidates from its start. excessiveLoadRetryMax is named.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
…ons, after the grace period (§3.6)

5209e74 cancelled the relay's TRACK_STATUS rounds as soon as Stop began,
before the GOAWAY and its grace period. A request in flight then got
DOES_NOT_EXIST with no retry, though its publisher answered within the
grace. Stop drops upstream work only at step 6, after the downstream GOAWAY
and grace ("prior to unsubscribing from upstream publishers", §3.6). The
rounds now end there too, next to the upstream pool, and handlers.Wait still
joins them.

TestTrackStatus_StopKeepsRoundThroughGrace (1s grace, publisher answering
after 200ms, Stop begun mid-round) failed before this change.
TestTrackStatus_StopEndsRound still holds.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
@floatdrop
floatdrop merged commit f64426d into draft-20 Sep 28, 2026
11 checks passed
@floatdrop
floatdrop deleted the draft20-backlog-relay-short branch September 28, 2026 06:26
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant