Pith. sign in

REVIEW 24 cited by

Uni3D: Exploring Unified 3D Representation at Scale

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2310.06773 v1 pith:EBS4EPQA submitted 2023-10-10 cs.CV cs.CL

classification cs.CVcs.CL
keywords uni3drepresentationmodelsscalealignedclassificationexploringfeatures
verification ladder T0 review T1 audit T2 compute T3 formal

Signed reviews

No signed human review yet.

0 comments
read the original abstract

Scaling up representations for images or text has been extensively investigated in the past few years and has led to revolutions in learning vision and language. However, scalable representation for 3D objects and scenes is relatively unexplored. In this work, we present Uni3D, a 3D foundation model to explore the unified 3D representation at scale. Uni3D uses a 2D initialized ViT end-to-end pretrained to align the 3D point cloud features with the image-text aligned features. Via the simple architecture and pretext task, Uni3D can leverage abundant 2D pretrained models as initialization and image-text aligned models as the target, unlocking the great potential of 2D models and scaling-up strategies to the 3D world. We efficiently scale up Uni3D to one billion parameters, and set new records on a broad range of 3D tasks, such as zero-shot classification, few-shot classification, open-world understanding and part segmentation. We show that the strong Uni3D representation also enables applications such as 3D painting and retrieval in the wild. We believe that Uni3D provides a new direction for exploring both scaling up and efficiency of the representation in 3D domain.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 24 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. STAR: A Spatial-Topology Aware Routing Framework for Generalizable 3D Scene Understanding

    cs.CV 2026-08 conditional novelty 6.0 of 10

    STAR adds spatial-topology-aware routing and uncertainty-based expert activation to a 3D mixture-of-experts model, reporting gains of 0.4 to 1.2 mIoU over strong baselines.

  2. CR-Refiner: An Object-Centric Optimal Transport Reranker for Edit-Conditioned 3D Scene Retrieval

    cs.CV 2026-07 conditional novelty 6.0 of 10

    An unbalanced optimal-transport reranker with structural priors and an LLM verifier improves hard-subset 3D scene retrieval, evaluated on the new synthetic 3D-CER benchmark.

  3. CAIRN: Cross-Room 3D Scene Understanding with Topology-Aware Large Multimodal Models

    cs.CV 2026-07 conditional novelty 6.0 of 10

    Topology-aware attention over hierarchical scene graphs lets a 3D-LLM ground, caption, and answer questions across multi-room homes, with large gains on a new HM3D benchmark.

  4. PatchAlign3D: Local Feature Alignment for Dense 3D Shape Understanding

    cs.CV 2026-01 conditional novelty 6.0 of 10

    A feed-forward 3D encoder aligning patch-level point-cloud features with part-name text embeddings achieves state-of-the-art zero-shot 3D part segmentation, surpassing multi-view rendering pipelines by large margins o...

  5. Unifi3D: A Study on 3D Representations for Generation and Reconstruction in a Common Framework

    cs.GR 2025-09 conditional novelty 6.0 of 10

    SDF grids reconstruct best, Dual Octrees score best on automatic generation metrics, but users prefer SDF output, and reconstruction plus compression errors make up a large share of generation error.

  6. ManiFlow: A General Robot Manipulation Policy via Consistency Flow Training

    cs.RO 2025-09 conditional novelty 6.0 of 10

    ManiFlow trains a flow-matching policy with a continuous-time consistency objective and an adaptive cross-attention transformer, enabling dexterous manipulation with 1-2 inference steps and substantially higher succes...

  7. Asymmetric Dual Self-Distillation for 3D Self-Supervised Representation Learning

    cs.CV 2025-06 reject novelty 6.0 of 10

    AsymDSD unifies latent masked point modeling and cross-view invariance self-distillation to learn 3D representations, reporting 90.53% on ScanObjectNN and 93.72% with 930k-shape pretraining.

  8. PoseMaster: A Unified 3D Native Framework for Stylized Pose Generation

    cs.CV 2025-06 conditional novelty 6.0 of 10

    PoseMaster produces a 3D character mesh from one image and a target 3D skeleton, preserving identity and pose in a single unified model, and it outperforms two-stage 2D-to-3D baselines on the VRoid pose canonicalizati...

  9. Scene-R1: Video-Grounded Large Language Models for 3D Scene Reasoning without 3D Annotations

    cs.CV 2025-06 conditional novelty 6.0 of 10

    Video-grounded RL with two-stage 2D grounding and SAM2 lifting enables 3D object localization and QA without dense 3D instance supervision.

  10. UniPre3D: Unified Pre-training of 3D Point Cloud Models with Cross-Modal Gaussian Splatting

    cs.CV 2025-06 conditional novelty 6.0 of 10

    A point cloud pre-training method that uses 3D Gaussian splatting rendering and cross-modal image features to work for both objects and scenes.

  11. Direct3D-S2: Gigascale 3D Generation Made Easy with Spatial Sparse Attention

    cs.CV 2025-05 conditional novelty 6.0 of 10

    Direct3D-S2 uses a new Spatial Sparse Attention mechanism to train a sparse-volume diffusion transformer at 1024^3 resolution on 8 GPUs.

  12. Masked Point-Entity Contrast for Open-Vocabulary 3D Scene Understanding

    cs.CV 2025-04 conditional novelty 6.0 of 10

    A contrastive pre-training method aligns 3D point features with language at the entity level and achieves state-of-the-art open-vocabulary semantic segmentation on ScanNet.

  13. Text-guided Synthetic Geometric Augmentation for Zero-shot 3D Understanding

    cs.CV 2025-01 conditional novelty 6.0 of 10

    Adding Point-E generated, consistency-filtered synthetic point clouds to ShapeNet improves MixCon3D zero-shot 3D classification on Objaverse-LVIS, ScanObjectNN, and ModelNet40.

  14. 3DGraphLLM: Combining Semantic Graphs and Large Language Models for 3D Scene Understanding

    cs.CV 2024-12 conditional novelty 6.0 of 10

    A learnable 3D scene graph representation that injects semantic relation embeddings between objects improves LLM performance on 3D grounding, captioning, and question answering.

  15. ZeroKey: Point-Level Reasoning and Zero-Shot 3D Keypoint Detection from Large Language Models

    cs.CV 2024-12 conditional novelty 6.0 of 10

    ZeroKey detects 3D keypoints on unseen object categories by prompting the Molmo vision-language model on multiple rendered views and aggregating the back-projected points, with no 3D annotations required.

  16. GATE-3D: Geometry-Aware Test-time Adaptive Reranking for Open-Set 3D Shape Retrieval

    cs.CV 2026-07 conditional novelty 5.0 of 10

    A per-query router trained on cross-modal disagreement features blends a geometry residual into appearance-based 3D-shape retrieval, improving mAP@10 by 2.0 points on OS-ESB-core while avoiding always-on fusion degradation.

  17. Fast3D: Accelerating 3D Multi-modal Large Language Models for Efficient 3D Scene Understanding

    cs.CV 2025-07 conditional novelty 5.0 of 10

    Fast3D prunes up to 90% of object-centric visual tokens in 3D MLLMs while preserving about 96.8% of original benchmark performance, using a trained attention predictor and adaptive layer-wise pruning.

  18. Does Your 3D Encoder Really Work? When Pretrain-SFT from 2D VLMs Meets 3D VLMs

    cs.CV 2025-06 conditional novelty 5.0 of 10

    3D scene-centric VLMs underutilize the 3D encoder, overfit to linguistic and answer-frequency shortcuts, and the proposed 3D-RDQA dataset helps expose and mitigate this problem.

  19. From Air to Wear: Personalized 3D Digital Fashion with AR/VR Immersive 3D Sketching

    cs.CV 2025-05 conditional novelty 5.0 of 10

    A VR-sketch-conditioned diffusion model, trained in three stages with curriculum learning and a new 969-pair dataset, generates plausible 3D garments from freehand 3D sketches.

  20. BIP3D: Bridging 2D Images and 3D Perception for Embodied Intelligence

    cs.CV 2024-11 conditional novelty 5.0 of 10

    BIP3D is an image-centric 3D perception model that uses 2D foundation model features with explicit 3D position encoding to beat point-cloud-based methods on the EmbodiedScan detection and grounding benchmarks.

  21. TinyGiantVLM: A Lightweight Vision-Language Architecture for Spatial Reasoning under Resource Constraints

    cs.CV 2025-08 conditional novelty 4.0 of 10

    TinyGiantVLM, a 64M-parameter RGB-D vision-language model with two-phase training, reached 5th place on the AI City Challenge 2025 warehouse spatial reasoning track.

  22. Step1X-3D: Towards High-Fidelity and Controllable Generation of Textured 3D Assets

    cs.CV 2025-05 conditional novelty 4.0 of 10

    An open-source-oriented image-to-3D system combines a VAE-DiT geometry generator and a diffusion texture module, claiming state-of-the-art quality over open-source rivals and near-proprietary performance.

  23. Scaling and Beyond: Advancing Spatial Reasoning in MLLMs Requires New Recipes

    cs.LG 2025-04 conditional novelty 4.0 of 10

    The authors argue that spatial reasoning in multimodal LLMs requires dedicated new recipes in data, architecture, and training objectives, and will not emerge from scaling alone.

  24. Embodied Intelligence for 3D Understanding: A Survey on 3D Scene Question Answering

    cs.CV 2025-02 conditional novelty 4.0 of 10

    A structured survey of 3D Scene Question Answering that categorizes datasets, methods, and metrics and finds a common encoder-fusion-prediction pipeline across approaches.

Pith tools