Pith. sign in

REVIEW 3 cited by

Joint-MAE: 2D-3D Joint Masked Autoencoders for 3D Point Cloud Pre-training

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2302.14007 v3 pith:BG5Y3JOI submitted 2023-02-27 cs.CV

classification cs.CV
keywords d-3djointjoint-maemaskedpointcloudpre-trainingaccuracy
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Masked Autoencoders (MAE) have shown promising performance in self-supervised learning for both 2D and 3D computer vision. However, existing MAE-style methods can only learn from the data of a single modality, i.e., either images or point clouds, which neglect the implicit semantic and geometric correlation between 2D and 3D. In this paper, we explore how the 2D modality can benefit 3D masked autoencoding, and propose Joint-MAE, a 2D-3D joint MAE framework for self-supervised 3D point cloud pre-training. Joint-MAE randomly masks an input 3D point cloud and its projected 2D images, and then reconstructs the masked information of the two modalities. For better cross-modal interaction, we construct our JointMAE by two hierarchical 2D-3D embedding modules, a joint encoder, and a joint decoder with modal-shared and model-specific decoders. On top of this, we further introduce two cross-modal strategies to boost the 3D representation learning, which are local-aligned attention mechanisms for 2D-3D semantic cues, and a cross-reconstruction loss for 2D-3D geometric constraints. By our pre-training paradigm, Joint-MAE achieves superior performance on multiple downstream tasks, e.g., 92.4% accuracy for linear SVM on ModelNet40 and 86.07% accuracy on the hardest split of ScanObjectNN.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 3 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. StruMamba3D: Exploring Structural Mamba for Self-supervised Point Cloud Representation Learning

    cs.CV 2025-06 conditional novelty 7.0 of 10

    A self-supervised point cloud model that encodes spatial structure into SSM latent states and adapts state-update scale to input length achieves new SOTA on ScanObjectNN and ModelNet40.

  2. Towards More Diverse and Challenging Pre-training for Point Cloud Learning: Self-Supervised Cross Reconstruction with Decoupled Views

    cs.CV 2025-09 conditional novelty 6.0 of 10

    Point-PQAE pre-trains point cloud transformers by cross-reconstructing one randomly cropped and rotated view from another, improving frozen-feature accuracy on ScanObjectNN by up to 7% over Point-MAE.

  3. Harnessing Text-to-Image Diffusion Models for Point Cloud Self-Supervised Learning

    cs.CV 2025-07 conditional novelty 6.0 of 10

    PointSD uses a frozen Stable Diffusion model, conditioned on point clouds through rendered images, to generate training targets for point cloud self-supervised learning.

Pith tools