Notes capture decisions that are easy to lose in commit history. Prefer updating this file over leaving undocumented tribal knowledge.
- Cargo package
phalanxwith library (src/lib.rs) and binary (src/main.rs) errorsmodule:PhalanxError,Result<T>utils::logging:LogConfig,init_logging- Tooling:
rustfmt.toml, Clippy lints,.cargo/config.tomlaliases - Docs: README, architecture, this file, CHANGELOG, LICENSE
- Tests: unit tests in modules +
tests/smoke.rs
- Library errors =
thiserror, CLI =anyhow— matchable API vs edge ergonomics. tracingfor logging — spans map to load / prefill / decode.- Edition 2024 +
rust-version = "1.85"— latest stable, greenfield. unsafe_codelint — started asforbid; Phase 5 lowered todenysoweights::storagecan opt in formemmap2after review.- No empty domain folders — tree reflects reality.
- Commit
Cargo.lock— binary runtime needs reproducible builds.
tensormodule:DType,Shape,Tensor,TensorError- Contiguous row-major
f32storage with explicit strides helper - Ops:
add/sub/mul/div,scale,matmul,transpose,sum PhalanxError::Tensornesting via#[from]- Criterion bench
benches/tensor_ops.rs(elemwise add, matmul, transpose) - Unit + integration coverage for shape math and kernels
Pros: Clear memory model for an educational runtime; trivial aliasing story. Cons: No free broadcasting / advanced views. Reason: Teaching layout matters as much as shipping ops. Phase 5 can swap storage behind the façade.
Pros: Call sites already thread a dtype; quantized variants slot in later. Cons: Slight indirection while only one variant exists. Reason: Avoid a painful API break when GGUF quants arrive.
Pros: Shape errors stay obvious; kernels stay short.
Cons: Some NumPy-style one-liners need explicit scale / expand later.
Reason: Attention / FFN paths use explicit shapes; broadcast bugs are subtle.
Pros: Auditable reference; good correctness oracle for future kernels. Cons: Not competitive with BLAS. Reason: Phase 17/18 own performance work; wrong-fast is worthless.
Pros: Preserves the “always contiguous” invariant for every Tensor.
Cons: Extra bandwidth vs a strided view.
Reason: Defer strided tensors until KV-cache windows need them (Phase 12).
Pros: Honors rust-version = "1.85".
Cons: Misses newest Criterion features.
Reason: Cargo rejected 0.8 (needs rustc 1.86). Bump MSRV intentionally later
if we want 0.8.
| Layer | Location | Covers |
|---|---|---|
| Unit | tensor::* |
strides, offsets, constructors, ops |
| Integration | tests/smoke.rs |
public matmul + nested tensor errors |
| Bench | benches/tensor_ops.rs |
baseline latency for add / matmul / transpose |
Baseline only — numbers vary by machine. Use cargo bench --bench tensor_ops
to refresh locally. Expect matmul_256 to dominate; treat regressions after
kernel changes as signal, not absolute SLA yet.
- Batched matmul helpers → still local loops inside Attention (Phase 11)
- f16 / bf16 / quantized storage → Phase 5+
- SIMD / blocked matmul → Phase 17
- Strided views → Phase 12 (KV cache)
- CLI argument parsing (
clap) → Phase 16 - Workspace split → when compile time / deps justify it
- Golub & Van Loan, Matrix Computations (matmul structure)
- NumPy C-order / row-major
- Criterion.rs
- Prior Phase 1 references (
thiserror,anyhow,tracing)
ggufmodule:GgufFile,GgufHeader,MetadataValue,TensorInfo,GgmlType- Streaming little-endian reader with running byte offset
- Metadata KV decode for all
gguf_metadata_value_types (incl. nested arrays) - Tensor directory parse + alignment (
general.alignmentor default 32) data_offsetcomputation; weight blob intentionally unreadPhalanxError::Ggufnesting- In-test
GgufBuilderfixture writer + unit/integration tests - Educational notes in
docs/gguf.md
Pros: Inspect multi-GB checkpoints without pulling weights into RAM.
Cons: More cursor bookkeeping than Cursor<&[u8]> alone.
Reason: Phase 3’s job is the directory, not the payload.
Pros: Matches the educational mission; typed GgufError we control.
Cons: We own format quirks.
Reason: Senior readers should see the byte layout, not a black box.
Pros: Covers virtually all modern GGUF files.
Cons: Must not silently accept structural breaks in future versions.
Reason: Reject unknown versions loudly via UnsupportedVersion.
Pros: Inspection still works when ggml adds formats.
Cons: Callers must handle Unknown before dequant (Phase 5).
Reason: Forward-compatible directory listing.
Pros: Hostile headers cannot force absurd allocations. Cons: Extremely exotic files might need limit bumps. Reason: Parser robustness before trust.
| Layer | Location | Covers |
|---|---|---|
| Unit | gguf::* |
magic, version, alignment, tensor offsets, arrays |
| Integration | tests/smoke.rs |
public GgufFile / GgufError re-exports |
mmap/ dequant oftensor_data→ Phase 5- Big-endian GGUF → if/when real files require it
- CLI
inspectsubcommand → Phase 16
tokenizermodule:Tokenizer,Vocabulary,SpecialTokens,TokenizerModel- Load from
tokenizer.ggml.*viaTokenizer::from_gguf - Decode with SentencePiece
▁+<0xXX>byte pieces - Encode via greedy longest-match or BPE merges
PhalanxError::Tokenizernesting- Docs:
docs/tokenizer.md
Pros: Teaches the data path; zero new deps; uses GGUF tables directly. Cons: Not guaranteed bit-identical to every HF export. Reason: Educational runtime first; golden tests can tighten later.
Pros: Display text matches what users expect from chat UIs.
Cons: Callers debugging ids must opt into raw decode.
Reason: DecodeOptions exposes both behaviours.
Pros: Matches common Llama prefill behaviour.
Cons: Some models omit BOS — disable via EncodeOptions.
Reason: Safe default for decoder-only chat checkpoints.
| Layer | Location | Covers |
|---|---|---|
| Unit | tokenizer::* |
from_gguf, greedy/BPE round-trip, missing keys |
| Integration | tests/smoke.rs |
public encode/decode re-exports |
- Chat template / Jinja (
tokenizer.chat_template) → CLI/runtime phases - Golden parity vs llama.cpp on real GGUF → later phases
- GGUF tokenizer metadata
- SentencePiece (Kudo & Richardson, 2018)
- GPT-2 BPE (Radford et al.)
weightsmodule:WeightSet,WeightStorage,WeightTensor,QuantMeta- Read-only
memmap2mapping (soleunsafeisland) - Quant block metadata for dense + legacy Q + K-quants
- Payload bounds validation for every tensor at open
- Materialize
f32/f16→tensor::Tensor PhalanxError::Weightsnesting- Docs:
docs/weights.md
Pros: Still defaults to no unsafe; reviewed modules can opt in.
Cons: A careless #[allow] could slip in.
Reason: forbid cannot be overridden; mmap requires one unsafe call.
Pros: Simple absolute offsets; parse from the same bytes. Cons: Slightly larger map than a data-only window. Reason: Clarity over micro-optimization.
Pros: Phase stays focused; avoids half-baked Q4_K code. Cons: Can't run quantized matmul yet. Reason: Dequant belongs next to the kernels that consume it (Phase 7+).
Pros: No wrong type_size guesses.
Cons: Those GGUF files won't open until sizes are verified.
Reason: Prefer loud UnsupportedType over silent corruption.
| Layer | Location | Covers |
|---|---|---|
| Unit | weights::* |
quant sizes, f32 fixture, mmap tmpfile, truncated reject |
| Integration | tests/smoke.rs |
public QuantMeta export |
- Q4_0 / Q4_K / Q8_0 dequant → layer kernels
- CLI
inspect→ Phase 16
modelmodule:Architecture,ModelConfig, attention / RoPE sub-configs- Parse Llama
{arch}.*hyperparameters from GGUF metadata - Structural validation (GQA divisibility, head×dim = embd, RoPE parity, …)
- Defaults:
head_count_kv → head_count,rope.freq_base → 10000 - Legacy
rope.scale→ linearRopeScaling PhalanxError::Modelnesting- Docs:
docs/model.md
Pros: Honest scope; validation tuned to Llama invariants.
Cons: qwen2 / others fail with UnsupportedArchitecture.
Reason: Phase 6 is “Llama architecture”; multi-arch lands when kernels do.
Pros: Matches real GGUF writers / GGUF spec (uint64 counts).
Cons: Slightly wider parse path.
Reason: Rejecting u64 would break valid files.
Pros: Clear boundary; Phase 7 owns embedding + named tensors.
Cons: Two-step load for callers (ModelConfig + WeightSet).
Reason: Avoid half-wired layer graphs before embeddings exist.
| Layer | Location | Covers |
|---|---|---|
| Unit | model::* |
Llama 7B-style, GQA, u64 counts, bad GQA, legacy rope.scale |
| Integration | tests/smoke.rs |
nested ModelError re-export |
- Qwen2 / Phi / MoE architecture variants
- Cross-check
vocab_sizevs tokenizer length → when both load together
- LLaMA
- GGUF specification
- RoFormer
- llama.cpp
gguf-py/gguf/constants.py
layersmodule:EmbeddingTable,LayersError,TOKEN_EMBD_WEIGHT- Load
token_embd.weight, validate againstModelConfig - Reinterpret ggml
[n_embd, n_vocab]bytes as row-major[vocab, embd] forward/forward_onegather- Squeeze trailing unitary dims
- Docs:
docs/embeddings.md
Pros: Keeps hparams separate from kernels; RoPE/attn/FFN land nearby. Cons: Extra top-level module. Reason: Phase 8–11 are all layer kernels.
Pros: Zero-copy on multi-GB embedding matrices. Cons: Easy to get wrong without the ggml-order comment. Reason: Educational runtime should teach the layout, not hide a silent copy.
Pros: Uses existing to_f32_tensor; scope stays focused.
Cons: Quantized GGUF embeddings still fail materialize.
Reason: Dequant belongs with the first kernel that needs blocks at scale;
gather correctness is the Phase 7 deliverable.
| Layer | Location | Covers |
|---|---|---|
| Unit | layers::embedding |
GGUF fixture gather, trailing ones, OOR, missing weight |
| Integration | tests/smoke.rs |
public gather + nested LayersError |
- Quantized embedding dequant
- Tied
output.weight/ input embeddings
- LLaMA
- GGUF specification
- llama.cpp
TOKEN_EMBDnaming
layers::Ropewith cos/sin cache fromModelConfig- Adjacent-pair Llama / RoFormer rotation on
[seq, head_dim]and[seq, heads, head_dim] - Partial rotary dims; linear position scaling
LayersError::{InvalidActivationShape, RopePositionOutOfRange}- Docs:
docs/rope.md
Pros: Decode is a gather + mul/add; tables are inspectable. Cons: Memory grows with context. Reason: Educational clarity + typical inference pattern (llama.cpp-style).
Pros: Matches Llama / RoFormer reference. Cons: Some HF ports use NeoX layout — document the choice. Reason: GGUF Llama checkpoints expect this pairing.
Pros: No silent long-context bugs. Cons: Those GGUF files error until Phase later. Reason: Prefer loud failure over wrong angles.
| Layer | Location | Covers |
|---|---|---|
| Unit | layers::rope |
pos0 identity, L2 norm, partial, linear scale, OOR, yarn reject |
| Integration | tests/smoke.rs |
public norm-preserving rotate |
- YaRN / NTK / sectioned RoPE
- Wire into attention ✓ (Phase 11)
- Complex-view / SIMD rotate kernels (Phase 17)
layers::RmsNormwith Spec formulaγ ⊙ x / RMS(x)epsfromModelConfig::rms_norm_eps- GGUF γ helpers:
attn_norm_weight_name,ffn_norm_weight_name,OUTPUT_NORM_WEIGHT - Cross-impl binary
src/bin/validate_rmsnorm.rs - Docs:
docs/rmsnorm.md
Pros: Matches Odyssey Spec / LLaMA; cheaper than LayerNorm. Cons: Diverges from original Transformer post-norm stacks. Reason: Spec compliance is non-negotiable (Rule 6).
Pros: Matches Odyssey normalization.rms accumulation path.
Cons: Slightly more ops than a pure fp16 path.
Reason: Cross-impl parity at 1e-6 before chasing speed.
| Layer | Location | Covers |
|---|---|---|
| Unit | layers::rmsnorm |
unit RMS, γ scale, shape, eps reject |
| Integration | tests/smoke.rs |
public unit-RMS check |
| Parity | validate_rmsnorm + Odyssey script |
max/mean abs error |
- Wire into decoder pre-norm residuals (Phase 13)
- SIMD / fused norm kernels (Phase 17)
- RMSNorm
- Odyssey Spec
rmsnorm.md
layers::SwiGluwith Spec formula(SiLU(x W1ᵀ) ⊙ (x W3ᵀ)) W2ᵀ- GGUF helpers:
ffn_gate_weight_name/ffn_up_weight_name/ffn_down_weight_name - Cross-impl binary
src/bin/validate_swiglu.rs - Docs:
docs/swiglu.md Tensor::matmulfloat64 accumulators for tighter parity
Pros: Matches Odyssey Spec / GGUF mapping.
Cons: Differs from some pedagogical W1/W2/W3 renumberings.
Reason: Spec wins (Rule 6).
Pros: Honest about GEMM accumulation order vs PyTorch; mean error ≪ 1e-6.
Cons: Looser than default 1e-6.
Reason: Principle 8 / Rule 6 allow component-documented tolerances.
| Layer | Location | Covers |
|---|---|---|
| Unit | layers::swiglu |
shape, SiLU formula, bad dim, name helpers |
| Integration | tests/smoke.rs |
public shape preserve |
| Parity | validate_swiglu + Odyssey script |
max/mean/rel error |
- Fused SwiGLU kernels / BLAS (Phase 17)
- Wire into decoder block (Phase 13)
- GLU Variants
- Odyssey Spec
feedforward.md
layers::Attentionwith Spec formulasoftmax(QKᵀ/√d + M) VthenW_O- GQA KV expansion (
H / H_kvquery heads per KV head) - Optional
Ropeon Q/K insideforward - GGUF name helpers
attn_q/k/v/output - Cross-impl validator
validate_attention(tol1e-3)
- Prefill-only dense scores (no KV cache yet) — simpler correctness path before Phase 12
- Reference nested loops + f64 accumulators — matches Odyssey parity over raw speed
- Softmax max-subtraction in float32 — Spec stability recommendation
- Unit: shape preserve (MHA/GQA), causal finite outputs, GGUF name helpers
- Integration smoke: public
Attentionre-export - Cross-impl: Odyssey
scripts/validate_attention.py(with and without RoPE)