Repository navigation
Conversation
- Adjust snapToWordBoundary loop to check budget-bounded indices only - Add logic to distinguish: whitespace within budget, beyond budget, or nowhere - Return 0 when whitespace exists beyond budget to wait for larger budget - Return maxLen only when no whitespace exists anywhere (guarantee progress) - Add drainFor check to break when snapToWordBoundary returns 0
When snapToWordBoundary returns 0 (word boundary exists beyond budget), the previous tick's elapsed time was discarded, creating a fixed-point where budget never grew and long unbroken tokens stalled indefinitely. Now accumulate elapsed time across ticks where no progress is made, so the budget grows monotonically until reaching the word boundary or running out of whitespace entirely. Fixes stalls on long URLs, code tokens, and unbroken CJK runs.
Removes vitest import and vi.isFakeTimers() detection from production code. Implements plain buffering of all chunks through StreamSmoother as specified in the brief, with no environment-conditional logic.
Updated the 'streaming dot appears during streaming and disappears after chunk_end' test to use fake timers and advance past one pacing tick before the first assertion. This accounts for the ~50ms delay introduced by StreamSmoother pacing, which prevents the dot from rendering until the smoother's timer fires and emits the buffered content.
npm test doesn't run tsc/eslint, so noUncheckedIndexedAccess violations and formatting issues in Task 1's files went unnoticed until `make update-dist` ran the real build/lint pipeline.
…oothing # Conflicts: # js/dist/shinychat.js # js/dist/shinychat.js.map # pkg-py/src/shinychat/www/GIT_VERSION # pkg-py/src/shinychat/www/shinychat.js # pkg-py/src/shinychat/www/shinychat.js.map # pkg-r/inst/lib/shiny/GIT_VERSION # pkg-r/inst/lib/shiny/shinychat.js # pkg-r/inst/lib/shiny/shinychat.js.map
StreamSmoother snapped every cut to whitespace and kept each pushed chunk as its own queue entry, so cuts could only land at word ends -- usually the LLM's own chunk boundaries. Text popped in whole words or chunks, with stalls of up to ~380ms waiting for a boundary. Now text is revealed at character granularity, matching databot's SmoothingController (within 2 chars on replayed streams): - Drop word-boundary snapping; cuts never stall. Cuts only extend forward past a UTF-16 surrogate pair or a run of markdown markers (` ~ * _), since the chat reducer only recognizes whole fences. - Carry the fractional character budget across ticks. - New canMerge option merges consecutive pushes with equivalent metadata (chat: same content_type, no html_deps; markdown-stream: same trust, not a segment start), so pacing isn't bound to chunk edges. - Measure elapsed time from the previous tick start. At end of stream, finish() reveals the remaining buffer within ~400ms instead of dumping it at once. chunk_end / isStreaming=false take effect after the drain. In chat, actions arriving mid-drain (history_update, update_siblings, ...) are deferred and replayed in order; new content (chunk_start, message, clear, greeting) completes the drain immediately. Stop during the drain reveals the rest without marking the finished reply cancelled.
The trailing dot was appended on every update, so with paced streaming it hopped along with the text ~20 times a second and pulsed constantly. It now appears only when it carries information: - Inline, once streamed content has sat unchanged for 1.5s (STREAM_IDLE_MS via useStreamIdle); it disappears in the same render that new content arrives. Paced streaming reveals text continuously while the model produces it, so a gap this long means the source has genuinely stalled. - In a standalone markdown stream that has no content yet (chat already covers that case with its pending indicator). hastToReact gains a streamingDot option, separate from streaming, since streaming also controls deferred finalization of pending asides and suggestion lists. The dot fades in before pulsing, and skips the pulse under prefers-reduced-motion.
Character-level cuts regularly land inside `<...>` (raw HTML, raw-html islands, inline tags), which would flash on screen as literal text for a tick. Ported from the StreamCoalescer iteration of this PR and adapted to pacing: - A cut inside a tag whose `>` is already buffered extends past it. Tags render as nothing, so revealing one whole is invisible and never stalls. - A tag whose `>` hasn't arrived is held back at its `<`, with a 1.5s stall guard (maxTagHoldMs) for stray `<` with no `>`. Unspent budget isn't banked while held, so text doesn't lurch once the tag closes. - No hold while finishing (no more input is coming) or on flush(). - A `<` counts as a tag start only before a letter, `/`, `!`, `?`, or end of buffer, so "x < y" passes through. Opt out with tagBoundaries: false.
…dle streaming dot
cpsievert
force-pushed
the
worktree-streaming-smoothing
branch
from
October 7, 2026 00:40
e032921 to
a241c7e
Compare
cpsievert
marked this pull request as ready for review
October 7, 2026 20:41
This branch has not been deployed
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Summary
Today, when a response streams into the chat, each new piece of text pops onto the screen the moment it arrives — so the reveal is only as smooth as the delivery. Pieces show up in bursts (network buffering, server pauses, tool calls), and fast providers send hundreds per second. Every arrival also makes the browser re-process the message so far — more work than it needs to do when the stream is fast.
This PR reveals streamed text at a steady pace, a few characters at a time, like typing — even when delivery is bursty. Under the hood that's two simple ideas working together: throttling (the message is re-processed at most ~20 times a second, no matter how fast pieces arrive — that's where the benchmark savings below come from) and pacing (each update reveals just enough new text to keep the reveal smooth). One behavior change to be aware of: the pulsing dot no longer trails along with incoming text. It now appears only when nothing has arrived for 1.5 seconds or no text has arrived at all, and it disappears the moment text starts moving again.
Demo
Side-by-side of the same synthetic token stream (3–8 tokens every 50–150ms, a realistic LLM cadence — the
anthropicprofile in the app below), before (main) and after (this PR):smooth-demo.mp4
Try it
R app with a synthetic token stream (no API key needed)
Click Stream sample. The default
anthropicprofile delivers 3–8 tokens every 50–150ms — a realistic LLM cadence, so the comparison againstmain(whole-chunk pops) is apples-to-apples. Thestallprofile pauses mid-stream for 6 seconds, which is when the pulsing "waiting" dot should appear; it should disappear the moment text starts moving again.Benchmarks
We recorded one real streamed response from each of five providers, preserving its exact chunk sizes and timing, then replayed each recording in headless Chrome and measured how much main-thread CPU time the page burns while the response streams. Each number below is this PR's CPU time divided by today's, for the same recording: 0.38× is 62% less work; 1.58× is 58% more. The right column reruns the recordings with Chrome's CPU throttled to quarter speed, standing in for a weaker machine.
Wall-clock time to the finished response is unchanged; the only addition is ~0.4s of "finishing typing" at the end.
The savings track delivery rate: the more pieces per second, the more re-processing the throttle eliminates. Google is the one regression — at 8 pieces per second there is little to throttle, and pacing spreads the same text across more re-processing passes, for about 1.6–1.9× today's CPU. We're accepting that cost for the smooth reveal; the pure-throttling variant explored earlier on this branch keeps Google at roughly today's cost if we ever want to switch.
Design decisions
drainRate/minCharsPerSecondinStreamSmoothertune this.STREAM_IDLE_MS) or before any text arrives, and skips the pulse underprefers-reduced-motion. The threshold is a judgement call.