Skip to content

Make streamed chat responses appear smoothly, like typing - #411

Open
cpsievert wants to merge 21 commits into
mainfrom
worktree-streaming-smoothing
Open

cpsievert wants to merge 21 commits into
mainfrom
worktree-streaming-smoothing

Conversation

@cpsievert

@cpsievert cpsievert commented Sep 21, 2026 •

Copy link
Copy Markdown
Collaborator

Summary

Today, when a response streams into the chat, each new piece of text pops onto the screen the moment it arrives — so the reveal is only as smooth as the delivery. Pieces show up in bursts (network buffering, server pauses, tool calls), and fast providers send hundreds per second. Every arrival also makes the browser re-process the message so far — more work than it needs to do when the stream is fast.

This PR reveals streamed text at a steady pace, a few characters at a time, like typing — even when delivery is bursty. Under the hood that's two simple ideas working together: throttling (the message is re-processed at most ~20 times a second, no matter how fast pieces arrive — that's where the benchmark savings below come from) and pacing (each update reveals just enough new text to keep the reveal smooth). One behavior change to be aware of: the pulsing dot no longer trails along with incoming text. It now appears only when nothing has arrived for 1.5 seconds or no text has arrived at all, and it disappears the moment text starts moving again.

Demo

Side-by-side of the same synthetic token stream (3–8 tokens every 50–150ms, a realistic LLM cadence — the anthropic profile in the app below), before (main) and after (this PR):

smooth-demo.mp4

Try it

R app with a synthetic token stream (no API key needed)
library(shiny)
library(bslib)
library(coro)
library(shinychat)

SAMPLE <- paste0(
  "Here's a **quick** overview of how the `summarise()` function works in practice.\n\n",
  "1. First, group the data with `group_by()` so each species is handled separately.\n",
  "2. Then compute summary statistics like the mean and standard deviation.\n\n",
  "```r\n",
  "penguins |>\n",
  "  group_by(species) |>\n",
  "  summarise(mean_mass = mean(body_mass_g, na.rm = TRUE))\n",
  "```\n\n",
  "The result is a tidy tibble with one row per species. Gentoo penguins are ",
  "noticeably heavier on average, while Adélie and Chinstrap are similar. \U0001F427 You ",
  "might also want to look at _sex_ differences within each species before ",
  "drawing conclusions.\n"
)
SAMPLE <- paste0(SAMPLE, SAMPLE)

# Split before each space/newline (delimiter stays with the following word),
# and split words > 7 chars into a 5-char piece + remainder, mimicking
# subword tokenization.
tokenize <- function(text) {
  parts <- strsplit(text, "(?=[ \n])", perl = TRUE)[[1]]
  toks <- character(0)
  for (w in parts) {
    if (nchar(w) > 7) {
      toks <- c(toks, substr(w, 1, 5), substr(w, 6, nchar(w)))
    } else if (nzchar(w)) {
      toks <- c(toks, w)
    }
  }
  toks
}

profile_params <- function(profile) {
  switch(profile,
    anthropic = list(n = sample(3:8, 1), gap = 0.05 + runif(1) * 0.1),
    bursty = list(n = sample(15:40, 1), gap = 0.3 + runif(1) * 0.5),
    stall = list(n = sample(3:8, 1), gap = 0.05 + runif(1) * 0.1)
  )
}
STALL_AFTER_CHUNK <- 20
STALL_SECONDS <- 6.0

ui <- page_fillable(
  layout_columns(
    selectInput("profile", NULL, c("anthropic", "bursty", "stall")),
    actionButton("go", "Stream sample"),
    col_widths = c(3, 3),
    fill = FALSE
  ),
  chat_ui("chat")
)

server <- function(input, output, session) {
  fake_stream <- function(profile) {
    set.seed(7319)  # deterministic chunk sizes and timing
    toks <- tokenize(SAMPLE)

    async_generator(function() {
      await(async_sleep(0.3))
      i <- 1
      chunk <- 0
      while (i <= length(toks)) {
        prof <- profile_params(profile)
        n <- prof$n
        gap <- prof$gap
        yield(paste(toks[i:min(i + n - 1, length(toks))], collapse = ""))
        i <- i + n
        chunk <- chunk + 1
        if (profile == "stall" && chunk == STALL_AFTER_CHUNK) {
          gap <- STALL_SECONDS
        }
        await(async_sleep(gap))
      }
    })()
  }

  observeEvent(input$go, {
    chat_append("chat", fake_stream(input$profile))
  })
}

shinyApp(ui, server)

Click Stream sample. The default anthropic profile delivers 3–8 tokens every 50–150ms — a realistic LLM cadence, so the comparison against main (whole-chunk pops) is apples-to-apples. The stall profile pauses mid-stream for 6 seconds, which is when the pulsing "waiting" dot should appear; it should disappear the moment text starts moving again.

Benchmarks

We recorded one real streamed response from each of five providers, preserving its exact chunk sizes and timing, then replayed each recording in headless Chrome and measured how much main-thread CPU time the page burns while the response streams. Each number below is this PR's CPU time divided by today's, for the same recording: 0.38× is 62% less work; 1.58× is 58% more. The right column reruns the recordings with Chrome's CPU throttled to quarter speed, standing in for a weaker machine.

Provider (pieces/sec) CPU vs today CPU vs today, 4× throttled
OpenAI (86/sec) 0.49× 0.78×
DeepSeek (252/sec) 0.38× 0.68×
Anthropic (22/sec) 0.95× 0.94×
Groq (884/sec) 0.76× 0.93×
Google (8/sec) 1.58× 1.87×

Wall-clock time to the finished response is unchanged; the only addition is ~0.4s of "finishing typing" at the end.

The savings track delivery rate: the more pieces per second, the more re-processing the throttle eliminates. Google is the one regression — at 8 pieces per second there is little to throttle, and pacing spreads the same text across more re-processing passes, for about 1.6–1.9× today's CPU. We're accepting that cost for the smooth reveal; the pure-throttling variant explored earlier on this branch keeps Google at roughly today's cost if we ever want to switch.

Design decisions

  • Pacing, not just throttling. A pure throttle (emit everything buffered on each tick) is cheaper on CPU — especially for Google's large, infrequent pieces — but keeps the jerky reveal. We chose the typing effect and accept the Google regression in the table.
  • Reveal speed follows the backlog, not the provider. Each update reveals ~2% of whatever text is waiting, with a 30 chars/s floor, so a fast provider doesn't read faster than a slow one — bursts just catch up quicker. drainRate / minCharsPerSecond in StreamSmoother tune this.
  • Responses finish "typing" (~0.4s) instead of popping the tail in at once. While that tail drains, server messages that assume the response is final (history updates, sibling navigation) are held and replayed in order; a new message or pressing Stop cuts the drain short, and Stop doesn't mark the already-finished reply as cancelled.
  • The dot means "waiting", not "streaming". It appears only after 1.5s with no visible change (STREAM_IDLE_MS) or before any text arrives, and skips the pulse under prefers-reduced-motion. The threshold is a judgement call.

- Adjust snapToWordBoundary loop to check budget-bounded indices only
- Add logic to distinguish: whitespace within budget, beyond budget, or nowhere
- Return 0 when whitespace exists beyond budget to wait for larger budget
- Return maxLen only when no whitespace exists anywhere (guarantee progress)
- Add drainFor check to break when snapToWordBoundary returns 0
When snapToWordBoundary returns 0 (word boundary exists beyond budget),
the previous tick's elapsed time was discarded, creating a fixed-point
where budget never grew and long unbroken tokens stalled indefinitely.

Now accumulate elapsed time across ticks where no progress is made,
so the budget grows monotonically until reaching the word boundary
or running out of whitespace entirely.

Fixes stalls on long URLs, code tokens, and unbroken CJK runs.
Removes vitest import and vi.isFakeTimers() detection from production code.
Implements plain buffering of all chunks through StreamSmoother as specified
in the brief, with no environment-conditional logic.
Updated the 'streaming dot appears during streaming and disappears after
chunk_end' test to use fake timers and advance past one pacing tick before
the first assertion. This accounts for the ~50ms delay introduced by
StreamSmoother pacing, which prevents the dot from rendering until the
smoother's timer fires and emits the buffered content.
npm test doesn't run tsc/eslint, so noUncheckedIndexedAccess violations and
formatting issues in Task 1's files went unnoticed until `make update-dist`
ran the real build/lint pipeline.
…oothing

# Conflicts:
#	js/dist/shinychat.js
#	js/dist/shinychat.js.map
#	pkg-py/src/shinychat/www/GIT_VERSION
#	pkg-py/src/shinychat/www/shinychat.js
#	pkg-py/src/shinychat/www/shinychat.js.map
#	pkg-r/inst/lib/shiny/GIT_VERSION
#	pkg-r/inst/lib/shiny/shinychat.js
#	pkg-r/inst/lib/shiny/shinychat.js.map
@cpsievert cpsievert changed the title Smooth client-side pacing for streamed chat responses Coalesce streamed chunk rendering to cap client update frequency Oct 6, 2026
StreamSmoother snapped every cut to whitespace and kept each pushed chunk
as its own queue entry, so cuts could only land at word ends -- usually
the LLM's own chunk boundaries. Text popped in whole words or chunks,
with stalls of up to ~380ms waiting for a boundary.

Now text is revealed at character granularity, matching databot's
SmoothingController (within 2 chars on replayed streams):

- Drop word-boundary snapping; cuts never stall. Cuts only extend
  forward past a UTF-16 surrogate pair or a run of markdown markers
  (` ~ * _), since the chat reducer only recognizes whole fences.
- Carry the fractional character budget across ticks.
- New canMerge option merges consecutive pushes with equivalent
  metadata (chat: same content_type, no html_deps; markdown-stream: same
  trust, not a segment start), so pacing isn't bound to chunk edges.
- Measure elapsed time from the previous tick start.

At end of stream, finish() reveals the remaining buffer within ~400ms
instead of dumping it at once. chunk_end / isStreaming=false take
effect after the drain. In chat, actions arriving mid-drain
(history_update, update_siblings, ...) are deferred and replayed in
order; new content (chunk_start, message, clear, greeting) completes the
drain immediately. Stop during the drain reveals the rest without
marking the finished reply cancelled.
The trailing dot was appended on every update, so with paced streaming
it hopped along with the text ~20 times a second and pulsed constantly.

It now appears only when it carries information:

- Inline, once streamed content has sat unchanged for 1.5s
  (STREAM_IDLE_MS via useStreamIdle); it disappears in the same render
  that new content arrives. Paced streaming reveals text continuously
  while the model produces it, so a gap this long means the source has
  genuinely stalled.
- In a standalone markdown stream that has no content yet (chat already
  covers that case with its pending indicator).

hastToReact gains a streamingDot option, separate from streaming, since
streaming also controls deferred finalization of pending asides and
suggestion lists. The dot fades in before pulsing, and skips the pulse
under prefers-reduced-motion.
Character-level cuts regularly land inside `<...>` (raw HTML, raw-html
islands, inline tags), which would flash on screen as literal text for
a tick. Ported from the StreamCoalescer iteration of this PR and
adapted to pacing:

- A cut inside a tag whose `>` is already buffered extends past it.
  Tags render as nothing, so revealing one whole is invisible and never
  stalls.
- A tag whose `>` hasn't arrived is held back at its `<`, with a 1.5s
  stall guard (maxTagHoldMs) for stray `<` with no `>`. Unspent budget
  isn't banked while held, so text doesn't lurch once the tag closes.
- No hold while finishing (no more input is coming) or on flush().
- A `<` counts as a tag start only before a letter, `/`, `!`, `?`, or
  end of buffer, so "x < y" passes through.

Opt out with tagBoundaries: false.
@cpsievert
cpsievert force-pushed the worktree-streaming-smoothing branch from e032921 to a241c7e Compare October 7, 2026 00:40
@cpsievert cpsievert changed the title Coalesce streamed chunk rendering to cap client update frequency Smooth streamed text with character-level pacing Oct 7, 2026
@cpsievert cpsievert changed the title Smooth streamed text with character-level pacing Make streamed chat responses appear smoothly, like typing Oct 7, 2026
@cpsievert
cpsievert requested a balanced review from Copilot October 7, 2026 20:25

This comment was marked as resolved.

@cpsievert
cpsievert marked this pull request as ready for review October 7, 2026 20:41

This branch has not been deployed

No deployments
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants