Skip to content

About

No description, website, or topics provided.

Resources

Stars

0 stars

Watchers

0 watching

Forks

Latest commit

 

History

1 Commit

Folders and files

NameName
Last commit message
Last commit date
 
 
 
 
 
 
 
 

Repository files navigation

Self-Rewarding Language Model (SLRM) on Apple Silicon

A fully local, autonomous reinforcement learning loop built on Apple Silicon using the MLX framework. This project implements the core concepts from the paper "Self-Rewarding Language Models" (Yuan et al., Meta AI, Jan 2024), demonstrating how a single Large Language Model can act as both the generator and the judge to bootstrap its own alignment without human intervention.

🧠 Overview & Base Model

Traditional alignment methods like RLHF (Reinforcement Learning from Human Feedback) rely on expensive human annotators, while RLAIF (AI Feedback) typically requires a massive, closed-source API to grade a smaller student model.

This repository proves that a small, localized model can iteratively improve itself. We started with an open-source foundational model—specifically, Meta-Llama-3.1-8B-Instruct-4bit—and put it into a self-rewarding loop. During training, the model's ability to follow instructions improved strictly alongside its ability to judge its own quality.

🚀 Technical Highlights

  • 100% Local Execution: Designed entirely around Apple's unified memory architecture using the mlx and mlx_lm libraries.
  • Direct Preference Optimization (DPO): Bypasses complex PPO architectures by directly optimizing the policy using a cross-entropy objective derived from the model's own self-generated preference pairs.
  • Parameter-Efficient Fine-Tuning (PEFT): Utilizes Low-Rank Adaptation (LoRA) targeting attention projections (q_proj, v_proj) to keep training viable on consumer hardware.
  • Custom Quantized Fusion ("The Identity Hack"): Features a custom fusion script that successfully dequantizes 4-bit packed MLX weights by passing identity matrices through the layers, fusing the FP16 LoRA adapters, and saving a new standalone baseline model for the next iteration.

⚙️ Architecture & The Self-Rewarding Loop

The pipeline operates in a continuous, multi-iteration loop consisting of the following phases:

  1. Data Initialization: Bootstrapping from high-quality initial user queries (extracted from the OpenAssistant OASST1 dataset).
  2. Generation & Self-Scoring (LLM-as-a-Judge): The current model generates 4 candidate responses at a high temperature for diversity. It then switches to a strict prompt at a low temperature to grade its own answers on a 0-5 scale. The best and worst responses form the Chosen and Rejected pairs.
  3. Training (DPO): The model undergoes DPO fine-tuning using its own self-generated preference data.
  4. Adapter Fusion: The learned LoRA weights are mathematically fused back into the base model to create a permanently smarter "Iteration X" foundational model.
  5. Repeat: The new model generates new data, grades it, and trains itself again.

🛠️ Installation

  1. Clone this repository:
git clone [https://github.com/yourusername/self-rewarding-lm.git](https://github.com/yourusername/self-rewarding-lm.git)
cd self-rewarding-lm
  1. Create a virtual environment and install dependencies
python -m venv venv
source venv/bin/activate
pip install mlx mlx_lm pandas numpy tqdm
  1. Place your base model in the models/base_model directory

  2. Place your dataset files in the data/raw directory

📂 Repository Structure & Usage

Run the scripts in sequential order to execute a full iteration of the loop. Update the configuration paths inside the scripts for subsequent iterations.

  • data_formatter.py: Filters the OASST1 dataset to extract the top seed prompts.

  • generate_and_score.py: The model reads the seed prompts, generates 4 answers per prompt, scores them, and exports a DPO training dataset (dpo_pairs.jsonl).

  • train_dpo.py: Trains a LoRA adapter using the self-generated DPO pairs.

  • fuse_adapters.py: Custom logic to dequantize the base model, add the LoRA delta, and save a fused checkpoint.

  • analyze_stats.py: Analyzes the log files to track the "Win Margin" across iterations.

  • compare_models.py: Runs a static prompt across multiple model generations to qualitatively verify improvement.

  • judge_utils.py: Contains the strict 0-5 additive grading rubric used during the self-evaluation phase.

📊 Empirical Results & Observations

During local testing over 4 full iterations, this system perfectly reproduced the theoretical phenomena associated with reinforcement learning dynamics:

  • Iteration 1: The judge was "generous" and lacked discriminative power, yielding a narrow Win Margin (0.65).

  • Iteration 2 (The Breakthrough): The model developed a sharp value system. The Win Margin doubled (1.35), and qualitative responses evolved from generic to highly structured, accurate analogies.

  • Iteration 3 (The Plateau): The model fine-tuned its preferences, maintaining a high Win Margin (1.38) while exploring new semantic spaces.

  • Iteration 4 (Reward Hacking / Model Collapse): The "Generator" learned to game the "Judge" by mimicking high-scoring formats (specific lengths, bolded headers) without improving semantic quality. The Win Margin decayed (1.19), demonstrating the empirical ceiling of small-parameter self-rewarding models.

📖 Acknowledgments Based on the paper: Self-Rewarding Language Models by Yuan et al. (Meta AI, 2024). Built using the Apple MLX framework.

About

No description, website, or topics provided.

Resources

Stars

0 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages