Draft-20 backlog: stream-credit refusal and shared TRACK_STATUS rounds - #118
Merged
Merged
Conversation
…h EXCESSIVE_LOAD (§10.6.2) A SUBSCRIBE, or forwarded TRACK_STATUS, that the relay could not open to a live publisher for want of bidi-stream credit was answered DOES_NOT_EXIST with Retry Interval 0: "not available at the publisher", and "SHOULD NOT be retried". But the publisher is there, and the relay only "cannot process the request at this time". upstreamRejection now answers it EXCESSIVE_LOAD with the relay's jittered ~1s retry, the code it already uses for its own caps. The maintainer chose this. It ranks as "any other refusal", so it outranks another candidate's GOING_AWAY, TIMEOUT or DOES_NOT_EXIST. candidateRetry reports its Retry Interval so ties compare what is sent. TestRendezvous_NoStreamCreditEndsHold now expects EXCESSIVE_LOAD with a retry. TestSubscribe_NoStreamCreditOutranksNoPublisherYet covers both orders against a GOING_AWAY, and forwarded TRACK_STATUS. All failed before this change. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
…requests for a track Each TRACK_STATUS for a track with no Established subscription forwarded its own round upstream, so N concurrent requests made N upstream ones to every candidate (§13.1 amplification). The maintainer chose to share rounds. forwardTrackStatus now runs trackStatusUpstream through a relay-wide singleflight.Group, keyed by track: - requests for a track that arrive while a round is in flight wait for it; - each requester is answered with its own INCLUDE_PROPERTIES and counts against its own session's cap; - the round is bounded by its 5s timeout, not by any one requester; - a requester's STOP_SENDING stops only its own wait, freeing its slot. A request from a session the relay is asking about the track right now may be the relay's own TRACK_STATUS routed back (§6.2). It does not join a round that waits on it: it runs its own, where the loop guard skips that session. TestTrackStatus_ConcurrentRequestsShareOneRound failed before this change (the publisher was asked twice). TestTrackStatus_CancelReleasesRequester replaces TestTrackStatus_CancelEndsUpstream, since a requester's cancel no longer ends a round others may share. TestRelay_TrackStatusLoopStopsAtSecondHop now bounds the answer time, and fails, deadlocked for the full bound, when a routed-back request joins the round. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
…hort and joins d1e8172 ran the shared round on singleflight's own goroutine, on a context detached from any session. Relay.Stop neither ended nor joined it, so it could return while the round was still in Discovery. Calling relayGo from that goroutine instead could race Stop's handlers.Wait. forwardTrackStatus now keeps its own map of rounds in flight (golang.org/x/sync is no longer imported by the relay): - a round starts through relayGo from the requesting handler's own goroutine, so Stop joins it; - it runs on a relay-scoped context that Stop cancels first, so a round waiting on Discovery or an upstream is cut short; - requesters wait for its result or their own STOP_SENDING, as before. A test hook marks each request that has joined a round. Tests, each red under a mutant removing the behaviour: - TestTrackStatus_StopEndsRound: a round blocked in Discovery ends before Stop returns, within 2s; - TestTrackStatus_LeaderCancelKeepsSharedRound: the requester that started a round cancelling does not end it for another; - TestTrackStatus_ConcurrentRequestsShareOneRound: now waits on the hook rather than sleeping. Also: the forwardTrackStatus doc no longer implies it catches a loop through other sessions. STATUS.md records that, and that a request joining a round late gets the candidates from its start. excessiveLoadRetryMax is named. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
…ons, after the grace period (§3.6) 5209e74 cancelled the relay's TRACK_STATUS rounds as soon as Stop began, before the GOAWAY and its grace period. A request in flight then got DOES_NOT_EXIST with no retry, though its publisher answered within the grace. Stop drops upstream work only at step 6, after the downstream GOAWAY and grace ("prior to unsubscribing from upstream publishers", §3.6). The rounds now end there too, next to the upstream pool, and handlers.Wait still joins them. TestTrackStatus_StopKeepsRoundThroughGrace (1s grace, publisher answering after 200ms, Stop begun mid-round) failed before this change. TestTrackStatus_StopEndsRound still holds. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
A short relay batch from the draft-20 backlog. Every commit was reviewed by moqt-reviewer, and each has tests that were seen failing first, either before the change or under a mutant that removes the behaviour.
Commits
4725affStream credit gets EXCESSIVE_LOAD. When the relay can't open a SUBSCRIBE or forwarded TRACK_STATUS to a live publisher for lack of bidi-stream credit, the subscriber now gets EXCESSIVE_LOAD with the relay's ~1s jittered retry (§10.6.2). It was DOES_NOT_EXIST with Retry Interval 0 ("SHOULD NOT be retried"). It ranks as "any other refusal", above another candidate's GOING_AWAY, TIMEOUT or DOES_NOT_EXIST.d1e8172Concurrent forwarded TRACK_STATUS requests share one round. Requests for a track that arrive while a round is in flight wait for it:5209e74The round is relay work. It starts throughrelayGofrom the requesting handler, soStopjoins it. It runs on a relay context thatStopcancels, so a round stuck in Discovery or waiting on an upstream is cut short. This replaces the first version's untrackedsingleflightgoroutine, which could outliveStop.c7b434fRounds end after the grace period. They now end atStop's step 6, with the upstream sessions, after the downstream GOAWAY and grace (§3.6). Cancelling them asStopbegan turned publisher answers given within the grace into DOES_NOT_EXIST.Decisions
Known gaps (added to STATUS.md)
Verification
go test ./...andgolangci-lint runpass.go test -race ./pkg/relay/...passes.go mod tidy -diffis clean.🤖 Generated with Claude Code