Repository navigation
Model-side experiments: real-negative mixing, masked-character pretraining (tools + results; shipped weights unchanged) - #28
Merged
Conversation
data_wild/REPORT.md measured a 20% false-positive rate on real sentences with no amount in them (8.8% verified) against 0% on hand-written negatives: the model has never seen real non-amount text. Two levers, one CPU training run each. A. `sankhya.wild negatives` builds a pool of real lines the lexicon scorer gives no amount signal AND the shipped weights do not span, balanced across the four packs and asserted disjoint from the wild gold. `train.py --extra-negatives/--extra-negative-ratio` mixes them in as all-O examples, per epoch, like --extra. B. `sankhya.pretrain` gives the same SankhyaCNN trunk a masked-character reconstruction pass over 500k real sentences (15% masked, loss only on masked positions, shipped charset, <unk> as the mask token); `train.py --init-from` then loads just the trunk, heads fresh. int8 results on wild gold: FP 0.200 -> 0.079 (A) -> 0.033 (B), strict FP 0.088 -> 0.063 -> 0.025, with synthetic value accuracy back at the shipped 0.962 for B. Weights are NOT adopted into models/default. The negatives pool and the pretraining corpus are derived from CC BY-SA text, so both are gitignored; the regeneration commands are documented in python/README.md.
…raining, wild negatives tool (no weights change)
… recipe for next retrain)
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Summary
Two model-side levers against real-text false positives, each run once on CPU, evaluated on synthetic and real-text gold. Tools and findings ship; weights do not change (no version bump).
Tools
python -m sankhya.wild negatives: 20k real no-amount sentences (balanced across the four languages, disjoint from the wild gold, CC BY-SA derived so gitignored with the regeneration command documented).train.py --extra-negatives PATH… --extra-negative-ratio R: real lines mixed in as all-O examples;--init-from CKPTloads a pretrained trunk with fresh heads.python -m sankhya.pretrain: masked-character pretraining of the same trunk on ~500k real sentences (15% masking, loss on masked positions only, shipped charset). ~24 min on 4 CPU cores, reusable.Results (int8, before the R11–R18 rules)
Strict value accuracy stays 1.000 everywhere. On the model side alone, B cuts real-text false positives 6x at no synthetic cost.
Decision
Re-evaluated under the 0.8.0 rules, B and the shipped weights are within seed noise (FP 8 vs 10 of 240, verified FP 6 vs 4, synthetic identical), and B changes the charset, so the shipped weights stay. The rules captured most of what B learned. Real-negative mixing and the pretrained trunk become the standard recipe for the next weights release (
python/data_wild/REPORT.md§5).Test plan
python -m pytest tests/389 passingnpm test215 passing🤖 Generated with Claude Code
https://claude.ai/code/session_014m4us3Nx2Z2X5b7WbAw4iR
Generated by Claude Code