Alignment-Aware Quantization for LLM Safety

Sunghyun Wee; Suyoung Kim; Hyeonjin Kim; Kyomin Hwang; Nojun Kwak

Alignment-Aware Quantization for LLM Safety

Sunghyun Wee, Suyoung Kim, Hyeonjin Kim, Kyomin Hwang, Nojun Kwak

TL;DR

This work addresses the risk that standard post-training quantization (PTQ) can erode RLHF-driven safety alignment in large language models. It introduces Alignment-Aware Quantization (AAQ), which incorporates an Alignment-Preserving Contrastive (APC) loss into the PTQ pipeline to actively preserve alignment while achieving aggressive 4-bit quantization ($W4A4$). The APC pull-push objective leverages a safe, fine-tuned model and an unsafe pre-trained reference to guide the quantized model, using top-$K$ filtering to stabilize training on alignment-sensitive outputs. Across diverse model families, AAQ achieves robust safety preservation with minimal sacrifice to other utilities, offering a practical, generalizable path to efficient and trustworthy quantized LLMs. This approach decouples safety from perplexity and demonstrates that alignment-aware quantization can maintain safety without specialized calibration data, enabling safer deployment at scale.

Abstract

Safety and efficiency are both important factors when deploying large language models(LLMs). LLMs are trained to follow human alignment for safety, and post training quantization(PTQ) is applied afterward for efficiency. However, these two objectives are often in conflict, revealing a fundamental flaw in the conventional PTQ paradigm: quantization can turn into a safety vulnerability if it only aims to achieve low perplexity. Models can demonstrate low perplexity yet exhibit significant degradation in alignment with the safety policy, highlighting that perplexity alone is an insufficient and often misleading proxy for model safety. To address this, we propose Alignment-Aware Quantization(AAQ), a novel approach that integrates Alignment-Preserving Contrastive(APC) loss into the PTQ pipeline. Compared to simple reconstruction loss, ours explicitly preserves alignment by encouraging the quantized model to mimic its safe, instruction-tuned model while diverging from the unaligned, pre-trained counterpart. Our method achieves this robust safety alignment without resorting to specialized safety-focused calibration datasets, highlighting its practical utility and broad applicability. AAQ is compatible with standard PTQ techniques and enables robust 4-bit (W4A4) quantization across diverse model families such as LLaMA, Qwen, and Mistral while maintaining safety where previous methods fail. Our work resolves the critical trade-off between efficiency and safety, paving the way toward LLMs that are both efficient and trustworthy. Anonymized code is available in the supplementary material.

Alignment-Aware Quantization for LLM Safety

TL;DR

Abstract

Alignment-Aware Quantization for LLM Safety

TL;DR

Abstract

Paper Structure

Table of Contents

Figures (3)