DSCD: Large Language Model Detoxification with Self-Constrained Decoding
Ming Dong, Jinkui Zhang, Bolong Zheng, Xinhui Tu, Po Hu, Tingting He
TL;DR
DSCD addresses the detoxification challenge in large language models by introducing self-constrained decoding that adjusts next-token distributions using token-level layer signals, without parameter fine-tuning. It defines two operational modes—dynamic (MODE-1) and static (MODE-2) toxic-layer handling—and leverages safety, toxic, and hallucination layers to suppress toxic outputs while preserving fluency. The approach is demonstrated to achieve state-of-the-art detoxification and competitive fluency across multiple open-source models and benchmarks, and it remains plug-and-play when integrated with existing detox methods like DINM and SafeDecoding. These results suggest DSCD offers a scalable, efficient solution for safer LLM deployments with flexible trade-offs between detoxification strength and speed.
Abstract
Detoxification in large language models (LLMs) remains a significant research challenge. Existing decoding detoxification methods are all based on external constraints, which require additional resource overhead and lose generation fluency. This work proposes Detoxification with Self-Constrained Decoding (DSCD), a novel method for LLM detoxification without parameter fine-tuning. DSCD strengthens the inner next-token distribution of the safety layer while weakening that of hallucination and toxic layers during output generation. This effectively diminishes toxicity and enhances output safety. DSCD offers lightweight, high compatibility, and plug-and-play capabilities, readily integrating with existing detoxification methods for further performance improvement. Extensive experiments on representative open-source LLMs and public datasets validate DSCD's effectiveness, demonstrating state-of-the-art (SOTA) performance in both detoxification and generation fluency, with superior efficiency compared to existing methods. These results highlight DSCD's potential as a practical and scalable solution for safer LLM deployments.
