Skip to content

[docs] flagship polish: discoverability hook, mermaid diagram, BENCHMARKS.md, ROADMAP.md - #11

Merged
peterajhgraham merged 2 commits into
mainfrom
claude/cortex-engine-flagship-VtWi8
May 23, 2026
Merged

peterajhgraham merged 2 commits into
mainfrom
claude/cortex-engine-flagship-VtWi8

Conversation

@peterajhgraham

Copy link
Copy Markdown
Owner

Summary

Documentation and discoverability work only — no logic, architecture, or benchmark numbers were changed.

  • README hook — four-sentence opening above the existing intro that leads with the scientific motivation (why sub-30 ms p99 for BCI, what makes motor cortex spike trains a hard inference target) before pivoting to what the repo builds.
  • Quick Visual — mermaid architecture diagram (renders natively on GitHub) showing spike events → tokenizer → Perceiver cross-attention → self-attention stack → behavior head, with the three Triton kernels annotated at the right layers, and the serving subgraph (scheduler, paged KV cache, FastAPI) on the side.
  • Why This Matters Beyond BCI — paragraph after Architecture mapping the engineering patterns to LLM inference systems: continuous batching ↔ vLLM, paged KV cache ↔ PagedAttention, INT8 calibration ↔ LLM.int8(). Makes the repo legible to ML systems engineers who don't know BCI.
  • Honest Results callout — GitHub blockquote near the benchmark tables stating the hardware caveat plainly: the work is real and tested, the Triton speedup and full p99 numbers are pending CUDA, and everything is marked rather than estimated.
  • BENCHMARKS.md — single consolidated reference with a headline table linking to every per-area report (training, profiling, kernels, quantization, serving), plus per-area sections and a reproduction recipe.
  • ROADMAP.md — four scoped next steps (CUDA kernel benchmark sweep on A10, full trial-aligned NLB evaluation + leaderboard, multi-session / cross-subject generalization, end-to-end true INT8 matmul, self-hosted CUDA CI runner).
  • CI badge already exists and points at the correct workflow (.github/workflows/ci.yml) — no change needed.

Test plan

  • Render README on GitHub; verify the mermaid diagram and Honest Results blockquote render correctly.
  • Click every report link in BENCHMARKS.md to confirm paths resolve.
  • Set the GitHub repo description and topics from the repo Settings page (see chat for the exact strings — the MCP toolset available in this session does not expose repo-metadata mutations).

Generated by Claude Code

claude added 2 commits May 22, 2026 23:33
…MAP.md

- README: scientific-motivation hook above the existing intro, mermaid
  Quick Visual showing spike events through Perceiver to behavior head
  with Triton kernel annotations, Honest Results blockquote near the
  benchmark tables, and a Why This Matters Beyond BCI section that
  connects the engineering patterns (continuous batching, paged KV
  cache, INT8 calibration) to LLM inference systems.
- BENCHMARKS.md: single consolidated reference with a headline table
  linking to every per-area report and explicit pending CUDA markers.
- ROADMAP.md: four next steps (CUDA benchmark sweeps, full trial-aligned
  NLB evaluation, multi-session generalization, true INT8 matmul on
  CUDA, self-hosted CUDA CI runner).
@peterajhgraham
peterajhgraham marked this pull request as ready for review May 23, 2026 00:10
@peterajhgraham
peterajhgraham merged commit 57f4204 into main May 23, 2026
4 checks passed
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants