MVFNet: Multipurpose Video Forensics Network using Multiple Forms of Forensic Evidence

Tai D. Nguyen; Matthew C. Stamm

MVFNet: Multipurpose Video Forensics Network using Multiple Forms of Forensic Evidence

Tai D. Nguyen, Matthew C. Stamm

TL;DR

MVFNet addresses the challenge of detecting and localizing diverse video forgeries without prior knowledge of the manipulation type. It combines spatial forensic residuals, RGB context, temporal forensic residuals, and optical-flow residuals, and processes them with a Multi-Scale Hierarchical Transformer to capture inconsistencies across scales and modalities. The approach introduces new modalities and a targeted pretraining loss, achieving state-of-the-art performance in multi-manipulation scenarios and competitive results against specialized detectors in single-manipulation tasks. The Unified Video Forgery Analysis dataset, along with UVFA-IND, UVFA-OOD, and VideoSham, provides a rigorous benchmark for evaluating generalization to unseen forgeries, highlighting MVFNet’s robustness and practical impact for broad video authentication needs.

Abstract

While videos can be falsified in many different ways, most existing forensic networks are specialized to detect only a single manipulation type (e.g. deepfake, inpainting). This poses a significant issue as the manipulation used to falsify a video is not known a priori. To address this problem, we propose MVFNet - a multipurpose video forensics network capable of detecting multiple types of manipulations including inpainting, deepfakes, splicing, and editing. Our network does this by extracting and jointly analyzing a broad set of forensic feature modalities that capture both spatial and temporal anomalies in falsified videos. To reliably detect and localize fake content of all shapes and sizes, our network employs a novel Multi-Scale Hierarchical Transformer module to identify forensic inconsistencies across multiple spatial scales. Experimental results show that our network obtains state-of-the-art performance in general scenarios where multiple different manipulations are possible, and rivals specialized detectors in targeted scenarios.

MVFNet: Multipurpose Video Forensics Network using Multiple Forms of Forensic Evidence

TL;DR

Abstract

MVFNet: Multipurpose Video Forensics Network using Multiple Forms of Forensic Evidence

TL;DR

Abstract

Paper Structure

Table of Contents

Figures (5)