Shift and matching queries for video semantic segmentation

Tsubasa Mizuno; Toru Tamaki

Shift and matching queries for video semantic segmentation

Tsubasa Mizuno, Toru Tamaki

TL;DR

Experimental results on CityScapes-VPS and VSPW show significant improvements from the baselines, highlighting the method's effectiveness in enhancing segmentation quality while efficiently reusing pre-trained weights.

Abstract

Video segmentation is a popular task, but applying image segmentation models frame-by-frame to videos does not preserve temporal consistency. In this paper, we propose a method to extend a query-based image segmentation model to video using feature shift and query matching. The method uses a query-based architecture, where decoded queries represent segmentation masks. These queries should be matched before performing the feature shift to ensure that the shifted queries represent the same mask across different frames. Experimental results on CityScapes-VPS and VSPW show significant improvements from the baselines, highlighting the method's effectiveness in enhancing segmentation quality while efficiently reusing pre-trained weights.

Shift and matching queries for video semantic segmentation

TL;DR

Abstract

Shift and matching queries for video semantic segmentation

TL;DR

Abstract

Paper Structure

Table of Contents

Figures (3)