Skip to content

Model-side experiments: real-negative mixing, masked-character pretraining (tools + results; shipped weights unchanged) - #28

Merged
athrvk merged 3 commits into
masterfrom
claude/busy-tesla-nt8ytr
Sep 20, 2026
Merged

athrvk merged 3 commits into
masterfrom
claude/busy-tesla-nt8ytr

Conversation

@athrvk

@athrvk athrvk commented Sep 19, 2026

Copy link
Copy Markdown
Owner

Summary

Two model-side levers against real-text false positives, each run once on CPU, evaluated on synthetic and real-text gold. Tools and findings ship; weights do not change (no version bump).

Tools

  • python -m sankhya.wild negatives: 20k real no-amount sentences (balanced across the four languages, disjoint from the wild gold, CC BY-SA derived so gitignored with the regeneration command documented).
  • train.py --extra-negatives PATH… --extra-negative-ratio R: real lines mixed in as all-O examples; --init-from CKPT loads a pretrained trunk with fresh heads.
  • python -m sankhya.pretrain: masked-character pretraining of the same trunk on ~500k real sentences (15% masking, loss on masked positions only, shipped charset). ~24 min on 4 CPU cores, reusable.
  • 13 fast tests on the data plumbing; docs in python/README.

Results (int8, before the R11–R18 rules)

synthetic value acc real-text FP real-text verified FP
shipped 0.7.0 0.962 0.200 0.088
A: real negatives 15% 0.952 0.079 0.063
B: pretrained trunk + A 0.962 0.033 0.025

Strict value accuracy stays 1.000 everywhere. On the model side alone, B cuts real-text false positives 6x at no synthetic cost.

Decision

Re-evaluated under the 0.8.0 rules, B and the shipped weights are within seed noise (FP 8 vs 10 of 240, verified FP 6 vs 4, synthetic identical), and B changes the charset, so the shipped weights stay. The rules captured most of what B learned. Real-negative mixing and the pretrained trunk become the standard recipe for the next weights release (python/data_wild/REPORT.md §5).

Test plan

  • python -m pytest tests/ 389 passing
  • npm test 215 passing
  • CI on this PR

🤖 Generated with Claude Code

https://claude.ai/code/session_014m4us3Nx2Z2X5b7WbAw4iR


Generated by Claude Code

data_wild/REPORT.md measured a 20% false-positive rate on real sentences
with no amount in them (8.8% verified) against 0% on hand-written
negatives: the model has never seen real non-amount text. Two levers,
one CPU training run each.

A. `sankhya.wild negatives` builds a pool of real lines the lexicon
   scorer gives no amount signal AND the shipped weights do not span,
   balanced across the four packs and asserted disjoint from the wild
   gold. `train.py --extra-negatives/--extra-negative-ratio` mixes them
   in as all-O examples, per epoch, like --extra.

B. `sankhya.pretrain` gives the same SankhyaCNN trunk a masked-character
   reconstruction pass over 500k real sentences (15% masked, loss only
   on masked positions, shipped charset, <unk> as the mask token);
   `train.py --init-from` then loads just the trunk, heads fresh.

int8 results on wild gold: FP 0.200 -> 0.079 (A) -> 0.033 (B), strict FP
0.088 -> 0.063 -> 0.025, with synthetic value accuracy back at the
shipped 0.962 for B. Weights are NOT adopted into models/default.

The negatives pool and the pretraining corpus are derived from CC BY-SA
text, so both are gitignored; the regeneration commands are documented
in python/README.md.
…raining, wild negatives tool (no weights change)
@athrvk
athrvk merged commit ca2a9d3 into master Sep 20, 2026
2 checks passed
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants