Approaching Outside: Scaling Unsupervised 3D Object Detection from 2D Scene

Authors: Ruiyang Zhang, hu-zhangHu Zhang, hang-yuHang Yu, Zhedong Zheng

Published in European conference on computer vision (ECCV), 2024

Recommended citation: Ruiyang Zhang, Hu Zhang, Hang Yu, Zhedong Zheng, "Approaching Outside: Scaling Unsupervised 3D Object Detection from 2D Scene." European conference on computer vision (ECCV), 2024.
Download PDF: https://zdzheng.xyz/files/2024/ECCV24-Approach.pdf
中文解读: https://www.zhihu.com/question/660698707/answer/3575967153

Code is available at: https://github.com/Ruiyang-061X/LiSe

Abstract: The open-world 3D object detection is to accurately detect objects in unstructured environments with no explicit supervisory signals. This task, given sparse LiDAR point clouds, often results in compromised performance for detecting small or distant objects due to the inherent sparsity and limited spatial resolution. In this paper, we are among the early attempts to integrate LiDAR data with 2D images for open-world 3D detection and introduce a new method, dubbed LiDAR-2D Self-paced Learning (LiSe). We argue that RGB images serve as a valuable complement to LiDAR data, offering precise 2D localization cues, particularly when scarce LiDAR points are available for certain objects. Considering the unique characteristics of both modalities, our framework devises a self-paced learning pipeline that incorporates adaptive sampling and weak model aggregation strategies. The adaptive sampling strategy dynamically tunes the distribution of pseudo labels during training, countering the tendency of models to overfit on easily detected samples, such as nearby and large-sized objects. By doing so, it ensures a balanced learning trajectory across varying object scales and distances. The weak model aggregation component consolidates the strengths of models trained under different pseudo label distributions, culminating in a robust and powerful final model. Experimental evaluations validate the efficacy of our proposed LiSe method, manifesting significant improvements of +7.1% AP-BEV and +3.4% AP-3D on nuScenes compared to existing techniques.

@inproceedings{LiSe,
author = "Zhang, Ruiyang and Zhang, Hu and Yu, Hang and Zheng, Zhedong",
title = "Approaching Outside: Scaling Unsupervised 3D Object Detection from 2D Scene",
abstract = "The open-world 3D object detection is to accurately detect objects in unstructured environments with no explicit supervisory signals. This task, given sparse LiDAR point clouds, often results in compromised performance for detecting small or distant objects due to the inherent sparsity and limited spatial resolution. In this paper, we are among the early attempts to integrate LiDAR data with 2D images for open-world 3D detection and introduce a new method, dubbed LiDAR-2D Self-paced Learning (LiSe). We argue that RGB images serve as a valuable complement to LiDAR data, offering precise 2D localization cues, particularly when scarce LiDAR points are available for certain objects. Considering the unique characteristics of both modalities, our framework devises a self-paced learning pipeline that incorporates adaptive sampling and weak model aggregation strategies. The adaptive sampling strategy dynamically tunes the distribution of pseudo labels during training, countering the tendency of models to overfit on easily detected samples, such as nearby and large-sized objects. By doing so, it ensures a balanced learning trajectory across varying object scales and distances. The weak model aggregation component consolidates the strengths of models trained under different pseudo label distributions, culminating in a robust and powerful final model. Experimental evaluations validate the efficacy of our proposed LiSe method, manifesting significant improvements of +7.1\\% AP-BEV and +3.4\\% AP-3D on nuScenes compared to existing techniques.",
booktitle = "European conference on computer vision (ECCV)",
code = "https://github.com/Ruiyang-061X/LiSe",
url = "https://zdzheng.xyz/files/2024/ECCV24-Approach.pdf",
blog = "https://www.zhihu.com/question/660698707/answer/3575967153",
funding = "SRG2024-00002-FST",
year = "2024" }