While industry convention assumes "bigger is better," training large models fully is resource-intensive, often impractical, and costly to the environment. This project investigates how smaller models like GPT-2 can be selectively enhanced for commonsense reasoning tasks through Parameter-Efficient Fine-Tuning (PEFT) methods.
We compared four PEFT techniques: LoRA, QLoRA, Prefix Tuning, and IA³. Our results show that IA³ and Prefix Tuning substantially outperform LoRA and QLoRA, achieving up to 50% validation accuracy in commonsense categories, with noticeable reductions in perplexity. Each method, applied to the smallest GPT-2 model, outperformed the largest GPT-2 model in zero-shot commonsense reasoning tests.
- IA³ achieved the highest accuracy at 50% validation accuracy
- Prefix Tuning reached 46% validation accuracy
- Both significantly outperformed LoRA (12%) and QLoRA (15%)
- All PEFT methods on GPT-2 small outperformed zero-shot GPT-2 large (6.35%)
- Perplexity dropped dramatically from >5000 (zero-shot) to ~14 (fine-tuned)
| Method | Validation Accuracy | Perplexity |
|---|---|---|
| Zero-Shot GPT-2 (1.5B) | 6.35% | 5536.87 |
| LoRA | 12% | 14.31 |
| QLoRA | 15% | 14.45 |
| Prefix Tuning | 46% | 14.06 |
| IA³ | 50% | 16.24 |
Parameter-Efficient Fine-Tuning (PEFT) methods allow us to adapt large language models by updating only a small subset of parameters, rather than retraining the entire model. This approach:
- Reduces computational costs dramatically
- Minimizes memory requirements
- Enables rapid experimentation
- Maintains model stability
- LoRA (Low-Rank Adaptation): Introduces low-rank updates within attention weights
- QLoRA: Combines 4-bit model quantization with LoRA for memory efficiency
- Prefix Tuning: Appends learned prefix tokens to transformer layer inputs
- IA³ (Input-Adaptive Attention): Adds learnable scaling factors to key, value, and feedforward transformations
We used the CommonsenseQA dataset, which consists of multiple-choice questions testing everyday commonsense knowledge. The dataset includes:
- 785 categories of commonsense questions
- Multiple-choice format with 5 options each
- Focus on implicit world knowledge rather than surface patterns
- Clustered into 10 broad classes for analysis
Initially, we attempted mathematical reasoning with OpenMathInstruct-2, but GPT-2's limited capacity for symbolic structures caused training to stall. CommonsenseQA proved a better fit for evaluating PEFT effectiveness on reasoning tasks within GPT-2's capabilities.
- Base Model: GPT-2 Small (~125M parameters)
- Hardware: Single NVIDIA A100 GPU
- Framework: HuggingFace PEFT library
- Optimizer: AdamW
- Learning Rate: 5e-5
- Batch Size: 8
- Epochs: 10
- Evaluation: Every 100 steps
- Validation Accuracy: Percentage of correct multiple-choice answers
- Perplexity: Model confidence on validation sequences
- Category-wise Analysis: Performance across different commonsense domains
All methods showed significant improvements over zero-shot performance, but with different patterns:
- IA³ and Prefix Tuning: Rapid convergence with sustained high performance
- LoRA and QLoRA: Steady but limited improvement plateaus
Different PEFT methods showed varying strengths across commonsense categories:
- IA³ and Prefix Tuning: Excelled in abstract reasoning (Emotions, Technology, Buildings & Spaces)
- LoRA and QLoRA: More modest, uniform improvements across categories
We hypothesize that IA³ and Prefix Tuning's superior performance stems from their direct intervention in attention mechanisms, which are central to relational and inferential reasoning. In contrast, LoRA and QLoRA primarily update projection layers, which may be less effective for reshaping internal reasoning processes.
- IA³ and Prefix Tuning: Models selected semantically plausible answers even when incorrect, suggesting partial internalization of reasoning heuristics
- LoRA and QLoRA: Models frequently defaulted to seemingly random choices, implying weaker structural learning
Several exciting directions for continuation:
- Advanced PEFT Techniques: Explore newer methods and ensemble strategies
- Larger Models: Scale experiments to Mistral, Llama 4, and Qwen 3
- Complex Datasets: Test on more challenging reasoning benchmarks
- Ensemble Methods: Combine different PEFT techniques
- Multi-hop Reasoning: Evaluate on tasks requiring deeper logical chains
This work demonstrates that:
- Smaller models + PEFT can compete with larger models on specific tasks
- Targeted fine-tuning is more effective than brute-force scaling
- Environmental impact of LLM research can be significantly reduced
- Accessibility of advanced NLP capabilities can be democratized
This work builds upon several key papers in parameter-efficient fine-tuning:
- LoRA: Hu et al. (2021) - Low-Rank Adaptation of Large Language Models
- Prefix Tuning: Li & Liang (2021) - Prefix-Tuning: Optimizing Continuous Prompts
- IA³: Liu et al. (2022) - Few-Shot Parameter-Efficient Fine-Tuning
- PEFT Survey: Ding et al. (2023) - Parameter-Efficient Fine-Tuning of Large-Scale Pre-Trained Language Models
- Vishwanath Guruvayur - University of Virginia - vish@virginia.edu
- Luke Napolitano - University of Virginia - ljn5yms@virginia.edu
- Doruk Ozar - University of Virginia - bcp8dm@virginia.edu
For questions about this research, please reach out to any of the authors listed above.
This research was conducted as part of DS-6051 (Decoding Large Language Models) at the University of Virginia.