Skip to content

Latest commit

 

History

History
240 lines (201 loc) · 11.2 KB

File metadata and controls

240 lines (201 loc) · 11.2 KB

CLAUDE.md

This file provides guidance to Claude Code (claude.ai/code) when working with code in this repository.

Build Commands

# Install OCaml dependencies
opam install . --deps-only

# Build grammar libraries (requires npm)
cd grammars && ./build-grammars.sh && cd ..

# Build the project
dune build

# Run all tests
dune test

# Run specific test group (e.g., just the end-to-end Matcher tests)
dune exec tests/test_runner.exe -- test Matcher

# Run a single named test
dune exec tests/test_runner.exe -- test Matcher "find calls"

Project Overview

diffract is an OCaml library and CLI for parsing source files with tree-sitter and pattern matching. Key capabilities:

  • Parse source to S-expressions using tree-sitter grammars
  • Pattern matching with concrete syntax and metavariables
  • Change summaries (summarize): infer the spatch rules behind a before/after directory pair — rules with per-file sites, tiered after= rules over the leftovers, and per-file residual diffs for what no rule explains (docs/change-summary.md for usage, docs/change-summary-design.md for design)

Architecture

Core Library (lib/)

  • diffract.ml - Main module, re-exports submodules
  • tree.ml - Pure OCaml tree representation (eliminates FFI overhead during traversal)
  • tree_sitter_bindings.ml - Low-level ctypes FFI bindings
  • tree_sitter_helper.c - C helper layer wrapping TSNode in OCaml custom blocks (libffi can't handle 32-byte structs by value)
  • node.ml - FFI-based tree traversal (internal, used during parsing)
  • languages.ml - Static grammar registry (language name → external C binding)

Matcher (lib/) — the tokenizer-based matcher (the only matcher):

  • tokenize.ml - Parse a pattern body with tree-sitter as a lexer, keeping leaves; produces a pattern_token stream (sigil-free metavars, ellipsis, fragments)
  • cursor.ml - Abstract tree-cursor interface (Cursor.S) the matching engine runs over
  • tree_sitter_cursor.ml - Cursor.S over a real tree-sitter parse
  • stmatch.ml - The matching engine: strict/partial/field leaf-level matching with backtracking (Make functor over a Cursor.S)
  • matcher.ml - End-to-end: preamble parse → tokenize → match → transform; the public find/transform/debug_tokens/pattern_warnings API
  • text_diff.ml - Line-based unified diff (for apply's output)

Diff / change summaries (lib/)

  • tree_diff.ml - AST-level diff (GumTree-style node mapping); used by diff and as summarize's change-pair source
  • leaf_metric.ml - Token-level edit metric over tree-sitter leaf streams (Myers LCS distance); the summary safety gate's geodesic test
  • change_summary.ml - Thin facade: re-exports the public surface (summarize, format_summary, load_from_dirs, types). The summarize pipeline (propose → evaluate → select, tiered) is split across cs_*.ml modules, one per phase:
    • cs_types.ml - shared types (public API + internal pattern representation)
    • cs_config.ml - tuning constants (one documented home; internal, no CLI)
    • cs_trace.ml - diagnostics gated on the CS_TRACE env var
    • cs_pattern.ml - tree→pat_node, rendering, anti-unification, coherence predicates
    • cs_propose.ml - change-pair extraction + candidate channels (multi-level, content-extraction, deep delta chains, anchored lattice-descent, AU-intersection mining)
    • cs_cluster.ml - anti-unification dendrogram, orphan coarsening, one-sided clustering
    • cs_evaluate.ml - the per-site safety gate that defines a rule's meaning (§3.3)
    • cs_fusion.ml - conjunctive multi-section fusion of co-occurring changes
    • cs_select.ml - one tier: propose → evaluate → greedy set-cover over changed regions
    • cs_tier.ml - tiered loop, chain-effect accounting, residual emission
    • cs_io.ml - .summary formatting and the directory-pair loader
    • Golden + round-trip tests in tests/change_summary_cases/. Design: docs/change-summary-design.md (§6 Milestones is historical changelog; §1–§5 describe the current design).

Tree-sitter Integration Flow:

  1. C helper layer wraps TSNode/TSTree in OCaml custom blocks with finalizers
  2. ctypes-foreign binds to libtree-sitter and C helpers
  3. languages.ml dispatches to grammar language functions statically linked into the binary
  4. tree.ml converts FFI nodes to pure OCaml representation once during parsing

Matching Flow: a pattern's preamble (@@ sections) is parsed, its body is tokenized into a leaf stream (text, observed node type, string-interiority), and stmatch walks the source tree (via a Cursor.S) matching those tokens lexically: text must agree; a node-type disagreement is tolerated iff both sides agree on whether the leaf sits inside a string literal. A pattern body is a fragment whose re-parse assigns syntactic roles unreliably, so roles are not compared — but code never matches string contents (Tree.string_delimited, quote/heredoc detection, no per-language tables). Field mode is the exception and keeps strict node-type comparison, since its per-candidate source-context re-tokenization already assigns correct roles. Metavars are sigil-free (a leaf is a metavar iff its text equals a declared name). See docs/universal-tokenizer.md.

Adding a New Language

  1. Add C wrapper and external binding to lib/tree_sitter_helper.c and lib/languages.ml:
/* lib/tree_sitter_helper.c */
extern const TSLanguage *tree_sitter_ruby(void);
CAMLprim value tsh_ruby_language(value v_unit) {
    CAMLparam1(v_unit); CAMLreturn(caml_copy_nativeint((intnat)tree_sitter_ruby())); }
(* lib/languages.ml *)
external ruby_language : unit -> nativeint = "tsh_ruby_language"
(* add to canonical_info: ("ruby", [], ruby_language) *)
  1. Update grammars/build-grammars.sh with the compilation command:
npm install tree-sitter-ruby
cc -O2 -c -o "$TMPDIR_LOCAL/ruby_parser.o" \
  -I node_modules/tree-sitter-ruby/src \
  node_modules/tree-sitter-ruby/src/parser.c
cc -O2 -c -o "$TMPDIR_LOCAL/ruby_scanner.o" \
  -I node_modules/tree-sitter-ruby/src \
  node_modules/tree-sitter-ruby/src/scanner.c
ar rcs lib/libtree-sitter-ruby.a "$TMPDIR_LOCAL/ruby_parser.o" "$TMPDIR_LOCAL/ruby_scanner.o"
  1. Add copy rule and (foreign_archives ...) entry to lib/dune:
(rule (target libtree-sitter-ruby.a)
 (deps ../grammars/lib/libtree-sitter-ruby.a)
 (action (copy %{deps} %{target})))

And add tree-sitter-ruby to the (foreign_archives ...) list.

  1. Rebuild: cd grammars && ./build-grammars.sh && cd .. && dune build

Pattern File Format

Patterns use @@ delimiters with a required match mode and optional metavariable declarations:

@@
match: strict
metavar $obj: single
metavar $method: single
@@
$obj.$method()

Types: single (one AST node), sequence (zero or more nodes)

Ellipsis (...) can be used as anonymous sequence matching:

@@
match: strict
@@
function test() {
    ...
    echo "middle";
    ...
}

(PHP pattern bodies are bare PHP — no <?php tag: the php_only grammar variant parses fragments directly, and a tag in the body mis-tokenizes the pattern.)

  • ... matches zero or more nodes (like sequence metavars)
  • Auto-detects context: adds ; in statement position, not in argument position
  • Does NOT replace ...$var (PHP spread operator is preserved)
  • Each ... gets a unique binding name (..._0, ..._1, etc.)
  • Sequence metavars (including ...) are not supported with match: partial.
  • In transform bodies ... belongs on context lines: it is rejected anywhere on a + line (a match-side binder has nothing to bind in a replacement) and as a bare -/+ line. An inline ... within a - line's expression binds normally and the captured run is deleted with it. Rewrite one list element by marking only it between context ... lines; delete one by putting it (and its separator) on - lines between them.

Matching modes (required - must specify one):

  • match: strict - Exact positional matching (no extra children allowed, ordered). Use for function calls, arrays.
  • match: partial - Subset matching (ignores extra children, unordered). Use for object literals, JSX attributes.
  • match: field - Declaration matching that ignores optional fields the pattern omits (decorators, annotations, modifier groups, return types). The pattern's leaf stream is aligned to a subsequence of the declaration node's children: a child the pattern addresses is matched in full, a child it omits is skipped. Use for decorated/annotated definitions. (No per-language config — see docs/field-mode.md.) Transforms are surgical (all modes): a -/+ on a sub-part edits only that part and preserves the rest — context, partial's tolerated extras, field's ignored optional fields, and ...-captured source (see docs/surgical-transforms.md). Marking the whole container in partial/field mode (a body with no context line) still replaces it whole and drops those extras/fields; the CLI warns for that case only.

Conjunctive sibling sections

A pattern file with multiple @@ sections and no on $VAR directives is a conjunctive rule: every section must find at least one match for any transforms to fire. If any section finds nothing, the source is returned unchanged. Metavars of the same name across sections refer to the same binding (threaded in declaration order); a section without -/+ lines acts as a pure guard (must match but produces no edits). Matcher.find/Matcher.transform handle multi-section patterns directly.

Sequence rendering: join and foreach

When a sequence metavar is referenced inside a + replacement template, its elements are rendered and substituted in place. Two independent knobs control this:

  • join $VAR by "<sep>" (a preamble directive): the string placed between rendered elements. <sep> interprets \n, \t, \\. The default (no directive) is the empty string, so elements are concatenated.
  • foreach $VAR (a following @@ section): a per-element transform. The section matches one element of $VAR and its -/+ lines rewrite it; the rewritten elements are then joined. Without a foreach, elements render as their source text (identity).

The sequence metavar must be declared metavar $VAR: sequence and appear on the match side. Everything goes through Matcher.transform.

Example — identity render with a join separator

@@
match: strict
metavar $ELEMS: sequence
join $ELEMS by " && "
@@
- all([$ELEMS])
+ ($ELEMS)

all([x, y, z]) becomes (x && y && z).

Example — method chain (foreach per-element transform)

@@
match: strict
metavar $TAG: single
metavar $PROPS: sequence
@@
- matchExhaustive($TAG, { $PROPS });
+ match($TAG)$PROPS.exhaustive();
@@
match: strict
foreach $PROPS
metavar $KEY: single
metavar $VAL: single
@@
- $KEY: $VAL
+ .with("$KEY", $VAL)

matchExhaustive(tag, { a: f, b: g }); becomes match(tag).with("a", f).with("b", g).exhaustive(); — the foreach section rewrites each property, and $PROPS in the outer template is replaced by the joined results (default empty join here).

A foreach element with an empty replacement (a - line, no +) deletes the element and cleans up the adjacent separator — e.g. removing a deprecated property or unused argument from a list.