Table of Contents
Fetching ...

RAVEN: Robust Advertisement Video Violation Temporal Grounding via Reinforcement Reasoning

Deyi Ji, Yuekui Yang, Haiyang Wu, Shaoping Ma, Tianrun Chen, Lanyun Zhu

TL;DR

RAVEN tackles the challenge of temporally grounding advertisement violations in videos under noisy annotations by integrating curriculum reinforcement learning with a multimodal large language model to achieve emergent reasoning. It introduces a hierarchical reward scheme including Thinking Format, Grounding Format, Temporal IoU, Boundary Alignment, and Category Consistency rewards, and trains in three stages on precise and coarse data using GRPO. The approach yields superior violation category accuracy and precise interval localization on industrial and public benchmarks, with online A/B tests confirming improvements in precision, recall, and localization. RAVEN also demonstrates strong generalization and mitigates catastrophic forgetting, offering a practical path to deploying robust ad policy enforcement in real services.

Abstract

Advertisement (Ad) video violation detection is critical for ensuring platform compliance, but existing methods struggle with precise temporal grounding, noisy annotations, and limited generalization. We propose RAVEN, a novel framework that integrates curriculum reinforcement learning with multimodal large language models (MLLMs) to enhance reasoning and cognitive capabilities for violation detection. RAVEN employs a progressive training strategy, combining precisely and coarsely annotated data, and leverages Group Relative Policy Optimization (GRPO) to develop emergent reasoning abilities without explicit reasoning annotations. Multiple hierarchical sophisticated reward mechanism ensures precise temporal grounding and consistent category prediction. Experiments on industrial datasets and public benchmarks show that RAVEN achieves superior performances in violation category accuracy and temporal interval localization. We also design a pipeline to deploy the RAVEN on the online Ad services, and online A/B testing further validates its practical applicability, with significant improvements in precision and recall. RAVEN also demonstrates strong generalization, mitigating the catastrophic forgetting issue associated with supervised fine-tuning.

RAVEN: Robust Advertisement Video Violation Temporal Grounding via Reinforcement Reasoning

TL;DR

RAVEN tackles the challenge of temporally grounding advertisement violations in videos under noisy annotations by integrating curriculum reinforcement learning with a multimodal large language model to achieve emergent reasoning. It introduces a hierarchical reward scheme including Thinking Format, Grounding Format, Temporal IoU, Boundary Alignment, and Category Consistency rewards, and trains in three stages on precise and coarse data using GRPO. The approach yields superior violation category accuracy and precise interval localization on industrial and public benchmarks, with online A/B tests confirming improvements in precision, recall, and localization. RAVEN also demonstrates strong generalization and mitigates catastrophic forgetting, offering a practical path to deploying robust ad policy enforcement in real services.

Abstract

Advertisement (Ad) video violation detection is critical for ensuring platform compliance, but existing methods struggle with precise temporal grounding, noisy annotations, and limited generalization. We propose RAVEN, a novel framework that integrates curriculum reinforcement learning with multimodal large language models (MLLMs) to enhance reasoning and cognitive capabilities for violation detection. RAVEN employs a progressive training strategy, combining precisely and coarsely annotated data, and leverages Group Relative Policy Optimization (GRPO) to develop emergent reasoning abilities without explicit reasoning annotations. Multiple hierarchical sophisticated reward mechanism ensures precise temporal grounding and consistent category prediction. Experiments on industrial datasets and public benchmarks show that RAVEN achieves superior performances in violation category accuracy and temporal interval localization. We also design a pipeline to deploy the RAVEN on the online Ad services, and online A/B testing further validates its practical applicability, with significant improvements in precision and recall. RAVEN also demonstrates strong generalization, mitigating the catastrophic forgetting issue associated with supervised fine-tuning.
Paper Structure (31 sections, 7 equations, 2 figures, 6 tables)