Multi-Focused Video Group Activities Hashing

Zhongmiao Qi; Yan Jiang; Bolin Zhang; Chong Wang; Lijun Guo; Pengjiang Qian; Jiangbo Qian

Multi-Focused Video Group Activities Hashing

Zhongmiao Qi, Yan Jiang, Bolin Zhang, Chong Wang, Lijun Guo, Pengjiang Qian, Jiangbo Qian

TL;DR

This work introduces STVH, a spatiotemporal interleaved video hashing framework that jointly models object dynamics and group interactions to produce compact $K$-bit hash codes for efficient group activity retrieval. Extending to M-STVH, it adds multi-focused hierarchical fusion and a binary filtering matrix to support activity-focused or visual-focused hashing while reducing storage, using PVF and SGAT to fuse visual and positional cues and a composite loss including $L_{cls}$, $L_q$, $L_CON$, and $L_{recon}$. Experiments on VD, CAD, and CAED demonstrate competitive classification and retrieval performance, with MSF enabling a transition from visual to activity semantics across layers and providing flexible retrieval modes. The method offers practical impact for sports analytics and surveillance by enabling scalable, activity-aware video search with controllable focus and storage efficiency, and opens avenues for cross-camera extension.

Abstract

With the explosive growth of video data in various complex scenarios, quickly retrieving group activities has become an urgent problem. However, many tasks can only retrieve videos focusing on an entire video, not the activity granularity. To solve this problem, we propose a new STVH (spatiotemporal interleaved video hashing) technique for the first time. Through a unified framework, the STVH simultaneously models individual object dynamics and group interactions, capturing the spatiotemporal evolution on both group visual features and positional features. Moreover, in real-life video retrieval scenarios, it may sometimes require activity features, while at other times, it may require visual features of objects. We then further propose a novel M-STVH (multi-focused spatiotemporal video hashing) as an enhanced version to handle this difficult task. The advanced method incorporates hierarchical feature integration through multi-focused representation learning, allowing the model to jointly focus on activity semantics features and object visual features. We conducted comparative experiments on publicly available datasets, and both STVH and M-STVH can achieve excellent results.

Multi-Focused Video Group Activities Hashing

TL;DR

Abstract

Multi-Focused Video Group Activities Hashing

TL;DR

Abstract

Paper Structure

Table of Contents

Figures (12)