Skip to content

scripts: train_specialist, distillation for every arena plus an RLCR stage rewarding right and honest answers - #129

Merged
BANADDA merged 2 commits into
mainfrom
train-specialist
Sep 30, 2026
Merged

BANADDA merged 2 commits into
mainfrom
train-specialist

Conversation

@BANADDA

@BANADDA BANADDA commented Sep 30, 2026

Copy link
Copy Markdown
Member

Step 8 of the full system plan, part two. Stacked on #120.

scripts/train_decider.py becomes scripts/train_specialist.py (renamed, not copied), and gains an RLCR stage:

  • Distillation as before: calibrated decider training on answer probabilities for choice arenas (KL plus Brier), and fine tuning on frontier teacher completions for every other arena.
  • RLCR (--rlcr-steps N): for each task it draws --rlcr-samples answers and scores each with the arena's metric. The reward is the score minus --rlcr-weight times the squared gap between the answer's own confidence (exp of its mean token logprob) and whether it was right. Advantages are normalised within each task's group and the policy is pushed towards the better answers.
  • Reward ordering: right and sure 1.00, right and unsure 0.64, wrong and unsure −0.01, wrong and sure −0.81. Confident mistakes are what it trains out.
  • Choice arenas skip RLCR, because the Brier rule they already train on rewards honest confidence directly.
  • Export still keeps the output layer at Q8, and the temperature fold (fold_temperature.py) puts calibration into the file.

References in docs/system_miner.md and pyproject.toml move to the new name.

🤖 Generated with Claude Code

@BANADDA
BANADDA changed the base branch from generic-miner-tools to main September 30, 2026 19:48
@BANADDA
BANADDA merged commit 3ddad02 into main Sep 30, 2026
7 checks passed
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant