From Evaluation to Defense: Advancing Safety in Video Large Language Models

Yiwei Sun; Peiqi Jiang; Chuanbin Liu; Luohao Lin; Zhiying Lu; Hongtao Xie

From Evaluation to Defense: Advancing Safety in Video Large Language Models

Yiwei Sun, Peiqi Jiang, Chuanbin Liu, Luohao Lin, Zhiying Lu, Hongtao Xie

TL;DR

The paper addresses the overlooked safety risks of video-based LLMs by introducing VideoSafetyBench (VSB-77k), the first large-scale, culturally diverse benchmark for video safety, with evaluation (VSB-Eval) and post-training (VSB-R1-46k) datasets. It quantifies the safety degradation caused by video inputs and proposes a dual-stage defense, VideoSafety-R1, consisting of Alarm Token-Guided Safety Fine-Tuning (AT-SFT) and Safety-Guided GRPO reinforcement learning, achieving substantial gains on VSB-Eval-HH and strong generalization to image-safety benchmarks. The findings demonstrate that active, multi-modal safety reasoning can restore and enhance safety alignment in Video LLMs, marking a shift from passive harm detection to proactive safety reasoning in dynamic multimodal contexts. Overall, the work provides a foundational, scalable framework to study and improve safety in video LLMs with practical benchmarks and post-training methods.

Abstract

While the safety risks of image-based large language models have been extensively studied, their video-based counterparts (Video LLMs) remain critically under-examined. To systematically study this problem, we introduce \textbf{VideoSafetyBench (VSB-77k) - the first large-scale, culturally diverse benchmark for Video LLM safety}, which compromises 77,646 video-query pairs and spans 19 principal risk categories across 10 language communities. \textit{We reveal that integrating video modality degrades safety performance by an average of 42.3\%, exposing systemic risks in multimodal attack exploitation.} To address this vulnerability, we propose \textbf{VideoSafety-R1}, a dual-stage framework achieving unprecedented safety gains through two innovations: (1) Alarm Token-Guided Safety Fine-Tuning (AT-SFT) injects learnable alarm tokens into visual and textual sequences, enabling explicit harm perception across modalities via multitask objectives. (2) Then, Safety-Guided GRPO enhances defensive reasoning through dynamic policy optimization with rule-based rewards derived from dual-modality verification. These components synergize to shift safety alignment from passive harm recognition to active reasoning. The resulting framework achieves a 65.1\% improvement on VSB-Eval-HH, and improves by 59.1\%, 44.3\%, and 15.0\% on the image safety datasets MMBench, VLGuard, and FigStep, respectively. \textit{Our codes are available in the supplementary materials.} \textcolor{red}{Warning: This paper contains examples of harmful language and videos, and reader discretion is recommended.}

From Evaluation to Defense: Advancing Safety in Video Large Language Models

TL;DR

Abstract

From Evaluation to Defense: Advancing Safety in Video Large Language Models

TL;DR

Abstract

Paper Structure

Table of Contents

Figures (12)