A fully local, autonomous reinforcement learning loop built on Apple Silicon using the MLX framework. This project implements the core concepts from the paper "Self-Rewarding Language Models" (Yuan et al., Meta AI, Jan 2024), demonstrating how a single Large Language Model can act as both the generator and the judge to bootstrap its own alignment without human intervention.
Traditional alignment methods like RLHF (Reinforcement Learning from Human Feedback) rely on expensive human annotators, while RLAIF (AI Feedback) typically requires a massive, closed-source API to grade a smaller student model.
This repository proves that a small, localized model can iteratively improve itself. We started with an open-source foundational model—specifically, Meta-Llama-3.1-8B-Instruct-4bit—and put it into a self-rewarding loop. During training, the model's ability to follow instructions improved strictly alongside its ability to judge its own quality.
- 100% Local Execution: Designed entirely around Apple's unified memory architecture using the
mlxandmlx_lmlibraries. - Direct Preference Optimization (DPO): Bypasses complex PPO architectures by directly optimizing the policy using a cross-entropy objective derived from the model's own self-generated preference pairs.
- Parameter-Efficient Fine-Tuning (PEFT): Utilizes Low-Rank Adaptation (LoRA) targeting attention projections (
q_proj,v_proj) to keep training viable on consumer hardware. - Custom Quantized Fusion ("The Identity Hack"): Features a custom fusion script that successfully dequantizes 4-bit packed MLX weights by passing identity matrices through the layers, fusing the FP16 LoRA adapters, and saving a new standalone baseline model for the next iteration.
The pipeline operates in a continuous, multi-iteration loop consisting of the following phases:
- Data Initialization: Bootstrapping from high-quality initial user queries (extracted from the OpenAssistant OASST1 dataset).
- Generation & Self-Scoring (LLM-as-a-Judge): The current model generates 4 candidate responses at a high temperature for diversity. It then switches to a strict prompt at a low temperature to grade its own answers on a 0-5 scale. The best and worst responses form the Chosen and Rejected pairs.
- Training (DPO): The model undergoes DPO fine-tuning using its own self-generated preference data.
- Adapter Fusion: The learned LoRA weights are mathematically fused back into the base model to create a permanently smarter "Iteration X" foundational model.
- Repeat: The new model generates new data, grades it, and trains itself again.
- Clone this repository:
git clone [https://github.com/yourusername/self-rewarding-lm.git](https://github.com/yourusername/self-rewarding-lm.git)
cd self-rewarding-lm- Create a virtual environment and install dependencies
python -m venv venv
source venv/bin/activate
pip install mlx mlx_lm pandas numpy tqdm-
Place your base model in the
models/base_modeldirectory -
Place your dataset files in the
data/rawdirectory
Run the scripts in sequential order to execute a full iteration of the loop. Update the configuration paths inside the scripts for subsequent iterations.
-
data_formatter.py: Filters the OASST1 dataset to extract the top seed prompts. -
generate_and_score.py: The model reads the seed prompts, generates 4 answers per prompt, scores them, and exports a DPO training dataset (dpo_pairs.jsonl). -
train_dpo.py: Trains a LoRA adapter using the self-generated DPO pairs. -
fuse_adapters.py: Custom logic to dequantize the base model, add the LoRA delta, and save a fused checkpoint. -
analyze_stats.py: Analyzes the log files to track the "Win Margin" across iterations. -
compare_models.py: Runs a static prompt across multiple model generations to qualitatively verify improvement. -
judge_utils.py: Contains the strict 0-5 additive grading rubric used during the self-evaluation phase.
During local testing over 4 full iterations, this system perfectly reproduced the theoretical phenomena associated with reinforcement learning dynamics:
-
Iteration 1: The judge was "generous" and lacked discriminative power, yielding a narrow Win Margin (0.65).
-
Iteration 2 (The Breakthrough): The model developed a sharp value system. The Win Margin doubled (1.35), and qualitative responses evolved from generic to highly structured, accurate analogies.
-
Iteration 3 (The Plateau): The model fine-tuned its preferences, maintaining a high Win Margin (1.38) while exploring new semantic spaces.
-
Iteration 4 (Reward Hacking / Model Collapse): The "Generator" learned to game the "Judge" by mimicking high-scoring formats (specific lengths, bolded headers) without improving semantic quality. The Win Margin decayed (1.19), demonstrating the empirical ceiling of small-parameter self-rewarding models.
📖 Acknowledgments Based on the paper: Self-Rewarding Language Models by Yuan et al. (Meta AI, 2024). Built using the Apple MLX framework.