Improving Quantization-aware Training of Low-Precision Network via Block Replacement on Full-Precision Counterpart

Chengting Yu; Shu Yang; Fengzhao Zhang; Hanzhi Ma; Aili Wang; Er-Ping Li

Improving Quantization-aware Training of Low-Precision Network via Block Replacement on Full-Precision Counterpart

Chengting Yu, Shu Yang, Fengzhao Zhang, Hanzhi Ma, Aili Wang, Er-Ping Li

TL;DR

The paper tackles aggressive quantization in QAT by addressing gradient mismatch and limited representation at very low bit widths. It introduces the Block-wise Replacement Framework (BWRF), which grafts a fixed full-precision partner into a low-precision network to form intermediate mixed-precision models that guide both forward and backward passes through training. Training optimizes a joint objective that combines task loss with distillation losses from FP and MP branches, providing implicit regularization and improved gradient estimation. Empirically, BWRF delivers state-of-the-art results for 4-, 3-, and 2-bit quantization on ImageNet and CIFAR-10 under uniform quantization and remains compatible with existing QAT pipelines via a concise wrapper.

Abstract

Quantization-aware training (QAT) is a common paradigm for network quantization, in which the training phase incorporates the simulation of the low-precision computation to optimize the quantization parameters in alignment with the task goals. However, direct training of low-precision networks generally faces two obstacles: 1. The low-precision model exhibits limited representation capabilities and cannot directly replicate full-precision calculations, which constitutes a deficiency compared to full-precision alternatives; 2. Non-ideal deviations during gradient propagation are a common consequence of employing pseudo-gradients as approximations in derived quantized functions. In this paper, we propose a general QAT framework for alleviating the aforementioned concerns by permitting the forward and backward processes of the low-precision network to be guided by the full-precision partner during training. In conjunction with the direct training of the quantization model, intermediate mixed-precision models are generated through the block-by-block replacement on the full-precision model and working simultaneously with the low-precision backbone, which enables the integration of quantized low-precision blocks into full-precision networks throughout the training phase. Consequently, each quantized block is capable of: 1. simulating full-precision representation during forward passes; 2. obtaining gradients with improved estimation during backward passes. We demonstrate that the proposed method achieves state-of-the-art results for 4-, 3-, and 2-bit quantization on ImageNet and CIFAR-10. The proposed framework provides a compatible extension for most QAT methods and only requires a concise wrapper for existing codes.

Improving Quantization-aware Training of Low-Precision Network via Block Replacement on Full-Precision Counterpart

TL;DR

Abstract

Improving Quantization-aware Training of Low-Precision Network via Block Replacement on Full-Precision Counterpart

TL;DR

Abstract

Paper Structure

Table of Contents

Figures (5)