Skip to content

[docs] Update README with final CUDA serving benchmark numbers - #10

Merged
peterajhgraham merged 1 commit into
mainfrom
claude/update-cuda-benchmarks-6DzGB
May 19, 2026
Merged

peterajhgraham merged 1 commit into
mainfrom
claude/update-cuda-benchmarks-6DzGB

Conversation

@peterajhgraham

Copy link
Copy Markdown
Owner

$(cat <<'EOF'

Summary

  • Replaces MPS placeholder serving metrics with real CUDA A10 measurements from benchmarks/serving/results.md
  • Updates throughput (157.3 → 255 req/s), p50 (70.2 → 27 ms), p95 (new), p99 (358.5 → 261 ms), failure count (0/200 → 0/500)
  • Adds honest engineering note that the 261 ms p99 is a first-batch initialization spike (CUDA JIT + worker warmup) and steady-state p99 is ~28 ms
  • Updates section header from "MPS, concurrency=16" to "NVIDIA A10 (24GB), concurrency=8" to match actual test conditions

Test plan

  • Verify numbers match benchmarks/serving/results.md exactly
  • Confirm the p99 caveat note is accurate and clearly worded

https://claude.ai/code/session_011ZWNew2AiBDntmujDWfDKm
EOF
)


Generated by Claude Code

Replace MPS placeholder serving metrics with real CUDA A10 measurements:
255 req/s throughput, p50 27ms, p95 28ms, p99 261ms (500 requests, 0
failures). Add honest note that p99 reflects a first-batch initialization
spike and steady-state p99 is ~28ms. Update section header to reflect
NVIDIA A10 (24GB) hardware.

https://claude.ai/code/session_011ZWNew2AiBDntmujDWfDKm
@peterajhgraham
peterajhgraham marked this pull request as ready for review May 19, 2026 00:15
@peterajhgraham
peterajhgraham merged commit ffdca42 into main May 19, 2026
4 checks passed
@peterajhgraham
peterajhgraham deleted the claude/update-cuda-benchmarks-6DzGB branch May 19, 2026 00:15
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants