GaLLoP: Gradient-based Sparse Learning on Low-Magnitude Parameters
Anand Choudhary, Yasser Sulaıman, Lukas Mauch, Ghouthi Boukli Hacene, Fabien Cardinaux, Antoine Bosselut
TL;DR
GaLLoP tackles sparse fine-tuning of large language models by selecting a small, task-relevant parameter subset with a dual criterion: large task gradients and small pre-trained magnitudes. Implemented as a two-phase method, it updates only parameters whose gradient-to-weight ratio signals both high task relevance and preservation of pre-trained knowledge. Across eight commonsense datasets and two base models (LLaMA3 8B and Gemma 2B), GaLLoP achieves strong ID and robust OOD generalization, with 0% forget and 0% memorization rates and stable performance across seeds. The approach outperforms or matches leading PEFT and post-training editing techniques, highlighting practical impact for efficient, robust fine-tuning of expansive LLMs, while noting unstructured sparsity as a tractable area for future densification.
Abstract
Sparse fine-tuning techniques adapt LLMs to downstream tasks by only tuning a sparse subset of model parameters. However, the effectiveness of sparse adaptation depends on optimally selecting the model parameters to be fine-tuned. In this work, we introduce a novel sparse fine-tuning technique named GaLLoP: Gradient-based Sparse Learning on Low-Magnitude Parameters, which fine-tunes only those model parameters which have the largest gradient magnitudes on downstream tasks and the smallest pre-trained magnitudes, intuitively prioritizing parameters that are highly task-relevant, but minimally disruptive to pre-trained knowledge. Our experimentation with LLaMA3 8B and Gemma 2B as base models shows that GaLLoP consistently improves or matches the in-distribution as well as out-of-distribution performance obtained via the usage of other leading parameter-efficient fine-tuning techniques, including LoRA, DoRA, and SAFT. Our analysis demonstrates that GaLLoP mitigates catastrophic forgetting and memorization of task data, as important pre-trained parameters remain unchanged, and stabilizes performance relative to other fine-tuning techniques, robustly generalizing across most random seeds.
