Object-Centric Pretraining via Target Encoder Bootstrapping

Nikola Đukić; Tim Lebailly; Tinne Tuytelaars

Object-Centric Pretraining via Target Encoder Bootstrapping

Nikola Đukić, Tim Lebailly, Tinne Tuytelaars

TL;DR

This paper tackles the upper-bound limitation of object-centric learning imposed by frozen target encoders by introducing OCEBO, a self-distillation framework that pretrains object-centric models from scratch with an EMA-updated target encoder. A key contribution is cross-view patch filtering, which selects informative patches for reconstruction to avoid slot collapse during early training. Empirically, OCEBO trained on COCO-scale data achieves competitive unsupervised object discovery performance compared to models pretrained on hundreds of millions of images and demonstrates scalability with larger COCO-based datasets. The work highlights the importance of injecting object-centric inductive biases into the target encoder and points toward scalable object-centric foundation models, while providing code and pretrained models for reproducibility.

Abstract

Object-centric representation learning has recently been successfully applied to real-world datasets. This success can be attributed to pretrained non-object-centric foundation models, whose features serve as reconstruction targets for slot attention. However, targets must remain frozen throughout the training, which sets an upper bound on the performance object-centric models can attain. Attempts to update the target encoder by bootstrapping result in large performance drops, which can be attributed to its lack of object-centric inductive biases, causing the object-centric model's encoder to drift away from representations useful as reconstruction targets. To address these limitations, we propose Object-CEntric Pretraining by Target Encoder BOotstrapping, a self-distillation setup for training object-centric models from scratch, on real-world data, for the first time ever. In OCEBO, the target encoder is updated as an exponential moving average of the object-centric model, thus explicitly being enriched with object-centric inductive biases introduced by slot attention while removing the upper bound on performance present in other models. We mitigate the slot collapse caused by random initialization of the target encoder by introducing a novel cross-view patch filtering approach that limits the supervision to sufficiently informative patches. When pretrained on 241k images from COCO, OCEBO achieves unsupervised object discovery performance comparable to that of object-centric models with frozen non-object-centric target encoders pretrained on hundreds of millions of images. The code and pretrained models are publicly available at https://github.com/djukicn/ocebo.

Object-Centric Pretraining via Target Encoder Bootstrapping

TL;DR

Abstract

Object-Centric Pretraining via Target Encoder Bootstrapping

TL;DR

Abstract

Paper Structure

Table of Contents

Figures (4)