CO^3: Cooperative Unsupervised 3D Representation Learning for Autonomous Driving

Runjian Chen; Yao Mu; Runsen Xu; Wenqi Shao; Chenhan Jiang; Hang Xu; Zhenguo Li; Ping Luo

CO^3: Cooperative Unsupervised 3D Representation Learning for Autonomous Driving

Runjian Chen, Yao Mu, Runsen Xu, Wenqi Shao, Chenhan Jiang, Hang Xu, Zhenguo Li, Ping Luo

TL;DR

This paper proposes CO^3, namely Cooperative Contrastive Learning and Contextual Shape Prediction, to learn 3D representation for outdoor-scene point clouds in an unsupervised manner and believes CO^3 will facilitate understanding LiDAR point clouds in outdoor scene.

Abstract

Unsupervised contrastive learning for indoor-scene point clouds has achieved great successes. However, unsupervised learning point clouds in outdoor scenes remains challenging because previous methods need to reconstruct the whole scene and capture partial views for the contrastive objective. This is infeasible in outdoor scenes with moving objects, obstacles, and sensors. In this paper, we propose CO^3, namely Cooperative Contrastive Learning and Contextual Shape Prediction, to learn 3D representation for outdoor-scene point clouds in an unsupervised manner. CO^3 has several merits compared to existing methods. (1) It utilizes LiDAR point clouds from vehicle-side and infrastructure-side to build views that differ enough but meanwhile maintain common semantic information for contrastive learning, which are more appropriate than views built by previous methods. (2) Alongside the contrastive objective, shape context prediction is proposed as pre-training goal and brings more task-relevant information for unsupervised 3D point cloud representation learning, which are beneficial when transferring the learned representation to downstream detection tasks. (3) As compared to previous methods, representation learned by CO^3 is able to be transferred to different outdoor scene dataset collected by different type of LiDAR sensors. (4) CO^3 improves current state-of-the-art methods on both Once and KITTI datasets by up to 2.58 mAP. We believe CO^3 will facilitate understanding LiDAR point clouds in outdoor scene.

CO^3: Cooperative Unsupervised 3D Representation Learning for Autonomous Driving

TL;DR

Abstract

Paper Structure (25 sections, 11 equations, 7 figures, 14 tables, 1 algorithm)

This paper contains 25 sections, 11 equations, 7 figures, 14 tables, 1 algorithm.

Introduction
Related Works
Methods
Problem Formulation and Pipeline
Cooperative Contrastive Objective
Contextual Shape Prediction
Experiments
Experiment Setup
Main Results
Ablation Study
Qualitative Experiment
Conclusion and Future Work
Background about View Building in Contrastive Learning
Reconstruction Objective for Task-relevant Information
Datasets Details
...and 10 more sections

Figures (7)

Figure 1: Example views built by different methods in contrastive learning, including (a) previous indoor-scene methods (b) previous outdoor-scene methods and (c) the proposed CO3. Compared to previous methods, CO3 can build two views that differ a lot and share adequate common semantics.
Figure 2: (a) shows the unsupervised 3D representation learning pipeline. (b) presents the performance changes after pre-training with different methods. Our CO3 achieves consistent improvement.
Figure 3: The pipeline of CO3. With vehicle-side and fusion point clouds as inputs, we first process them with the 3D backbone and propose two pre-training objectives: (a) Cooperative Contrastive Loss (b) Contextual Shape Prediction Loss
Figure 4: Two examples of shape context.
Figure 4: Comparison to supervised initialization
...and 2 more figures

Theorems & Definitions (2)

Definition 1
Definition 2

CO^3: Cooperative Unsupervised 3D Representation Learning for Autonomous Driving

TL;DR

Abstract

CO^3: Cooperative Unsupervised 3D Representation Learning for Autonomous Driving

Authors

TL;DR

Abstract

Table of Contents

Figures (7)

Theorems & Definitions (2)