Restoring Neural Network Plasticity for Faster Transfer Learning

Xander Coetzer; Arné Schreuder; Anna Sergeevna Bosman

Restoring Neural Network Plasticity for Faster Transfer Learning

Xander Coetzer, Arné Schreuder, Anna Sergeevna Bosman

Abstract

Transfer learning with models pretrained on ImageNet has become a standard practice in computer vision. Transfer learning refers to fine-tuning pretrained weights of a neural network on a downstream task, typically unrelated to ImageNet. However, pretrained weights can become saturated and may yield insignificant gradients, failing to adapt to the downstream task. This hinders the ability of the model to train effectively, and is commonly referred to as loss of neural plasticity. Loss of plasticity may prevent the model from fully adapting to the target domain, especially when the downstream dataset is atypical in nature. While this issue has been widely explored in continual learning, it remains relatively understudied in the context of transfer learning. In this work, we propose the use of a targeted weight re-initialization strategy to restore neural plasticity prior to fine-tuning. Our experiments show that both convolutional neural networks (CNNs) and vision transformers (ViTs) benefit from this approach, yielding higher test accuracy with faster convergence on several image classification benchmarks. Our method introduces negligible computational overhead and is compatible with common transfer learning pipelines.

Restoring Neural Network Plasticity for Faster Transfer Learning

Abstract

Paper Structure (34 sections, 2 equations, 1 figure, 6 tables)

This paper contains 34 sections, 2 equations, 1 figure, 6 tables.

Introduction
Background and Related Work
Transfer Learning
Continual Learning and Neural Plasticity
Reinitialization Algorithms
Utility Functions:
Pruning Functions:
Reinitialization Methods:
Related Work: Parameter Efficient Tuning
Guided Training:
Weight Reduction:
Summary:
Methodology
Approach
Utility Functions:
...and 19 more sections

Figures (1)

Figure 1: Comparison between 25 $M$ (a) and 25 $N$ (b) experimental cases. The blue histogram is the total weight distribution after training on the base case. The green histogram represents the total weight distribution after fine-tuning in the experimental cases. The red histogram shows the absolute difference between the blue and green histograms and the final histogram overlays the blue and green histograms for easy comparison between base and experimental cases.

Restoring Neural Network Plasticity for Faster Transfer Learning

Abstract

Restoring Neural Network Plasticity for Faster Transfer Learning

Authors

Abstract

Table of Contents

Figures (1)