HyperAIHyperAI

Command Palette

Search for a command to run...

5 months ago

Sculpting Holistic 3D Representation in Contrastive Language-Image-3D Pre-training

Yipeng Gao; Zeyu Wang; Wei-Shi Zheng; Cihang Xie; Yuyin Zhou

Sculpting Holistic 3D Representation in Contrastive Language-Image-3D Pre-training

Abstract

Contrastive learning has emerged as a promising paradigm for 3D open-world understanding, i.e., aligning point cloud representation to image and text embedding space individually. In this paper, we introduce MixCon3D, a simple yet effective method aiming to sculpt holistic 3D representation in contrastive language-image-3D pre-training. In contrast to point cloud only, we develop the 3D object-level representation from complementary perspectives, e.g., multi-view rendered images with the point cloud. Then, MixCon3D performs language-3D contrastive learning, comprehensively depicting real-world 3D objects and bolstering text alignment. Additionally, we pioneer the first thorough investigation of various training recipes for the 3D contrastive learning paradigm, building a solid baseline with improved performance. Extensive experiments conducted on three representative benchmarks reveal that our method significantly improves over the baseline, surpassing the previous state-of-the-art performance on the challenging 1,156-category Objaverse-LVIS dataset by 5.7%. The versatility of MixCon3D is showcased in applications such as text-to-3D retrieval and point cloud captioning, further evidencing its efficacy in diverse scenarios. The code is available at https://github.com/UCSC-VLAA/MixCon3D.

Code Repositories

ucsc-vlaa/mixcon3d
Official
pytorch

Benchmarks

BenchmarkMethodologyMetrics
zero-shot-transfer-3d-point-cloudMixCon3D-PointBERT
Accuracy (%): 86.8
zero-shot-transfer-3d-point-cloud-2MixCon3D-PointBERT
OBJ_ONLY Accuracy(%): 58.6

Build AI with AI

From idea to launch — accelerate your AI development with free AI co-coding, out-of-the-box environment and best price of GPUs.

AI Co-coding
Ready-to-use GPUs
Best Pricing
Get Started

Hyper Newsletters

Subscribe to our latest updates
We will deliver the latest updates of the week to your inbox at nine o'clock every Monday morning
Powered by MailChimp
Sculpting Holistic 3D Representation in Contrastive Language-Image-3D Pre-training | Papers | HyperAI