REVIEW 24 cited by
Uni3D: Exploring Unified 3D Representation at Scale
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
Signed reviews
read the original abstract
Scaling up representations for images or text has been extensively investigated in the past few years and has led to revolutions in learning vision and language. However, scalable representation for 3D objects and scenes is relatively unexplored. In this work, we present Uni3D, a 3D foundation model to explore the unified 3D representation at scale. Uni3D uses a 2D initialized ViT end-to-end pretrained to align the 3D point cloud features with the image-text aligned features. Via the simple architecture and pretext task, Uni3D can leverage abundant 2D pretrained models as initialization and image-text aligned models as the target, unlocking the great potential of 2D models and scaling-up strategies to the 3D world. We efficiently scale up Uni3D to one billion parameters, and set new records on a broad range of 3D tasks, such as zero-shot classification, few-shot classification, open-world understanding and part segmentation. We show that the strong Uni3D representation also enables applications such as 3D painting and retrieval in the wild. We believe that Uni3D provides a new direction for exploring both scaling up and efficiency of the representation in 3D domain.
Forward citations
Cited by 24 Pith papers
-
STAR: A Spatial-Topology Aware Routing Framework for Generalizable 3D Scene Understanding
STAR adds spatial-topology-aware routing and uncertainty-based expert activation to a 3D mixture-of-experts model, reporting gains of 0.4 to 1.2 mIoU over strong baselines.
-
CR-Refiner: An Object-Centric Optimal Transport Reranker for Edit-Conditioned 3D Scene Retrieval
An unbalanced optimal-transport reranker with structural priors and an LLM verifier improves hard-subset 3D scene retrieval, evaluated on the new synthetic 3D-CER benchmark.
-
CAIRN: Cross-Room 3D Scene Understanding with Topology-Aware Large Multimodal Models
Topology-aware attention over hierarchical scene graphs lets a 3D-LLM ground, caption, and answer questions across multi-room homes, with large gains on a new HM3D benchmark.
-
PatchAlign3D: Local Feature Alignment for Dense 3D Shape Understanding
A feed-forward 3D encoder aligning patch-level point-cloud features with part-name text embeddings achieves state-of-the-art zero-shot 3D part segmentation, surpassing multi-view rendering pipelines by large margins o...
-
Unifi3D: A Study on 3D Representations for Generation and Reconstruction in a Common Framework
SDF grids reconstruct best, Dual Octrees score best on automatic generation metrics, but users prefer SDF output, and reconstruction plus compression errors make up a large share of generation error.
-
ManiFlow: A General Robot Manipulation Policy via Consistency Flow Training
ManiFlow trains a flow-matching policy with a continuous-time consistency objective and an adaptive cross-attention transformer, enabling dexterous manipulation with 1-2 inference steps and substantially higher succes...
-
Asymmetric Dual Self-Distillation for 3D Self-Supervised Representation Learning
AsymDSD unifies latent masked point modeling and cross-view invariance self-distillation to learn 3D representations, reporting 90.53% on ScanObjectNN and 93.72% with 930k-shape pretraining.
-
PoseMaster: A Unified 3D Native Framework for Stylized Pose Generation
PoseMaster produces a 3D character mesh from one image and a target 3D skeleton, preserving identity and pose in a single unified model, and it outperforms two-stage 2D-to-3D baselines on the VRoid pose canonicalizati...
-
Scene-R1: Video-Grounded Large Language Models for 3D Scene Reasoning without 3D Annotations
Video-grounded RL with two-stage 2D grounding and SAM2 lifting enables 3D object localization and QA without dense 3D instance supervision.
-
UniPre3D: Unified Pre-training of 3D Point Cloud Models with Cross-Modal Gaussian Splatting
A point cloud pre-training method that uses 3D Gaussian splatting rendering and cross-modal image features to work for both objects and scenes.
-
Direct3D-S2: Gigascale 3D Generation Made Easy with Spatial Sparse Attention
Direct3D-S2 uses a new Spatial Sparse Attention mechanism to train a sparse-volume diffusion transformer at 1024^3 resolution on 8 GPUs.
-
Masked Point-Entity Contrast for Open-Vocabulary 3D Scene Understanding
A contrastive pre-training method aligns 3D point features with language at the entity level and achieves state-of-the-art open-vocabulary semantic segmentation on ScanNet.
-
Text-guided Synthetic Geometric Augmentation for Zero-shot 3D Understanding
Adding Point-E generated, consistency-filtered synthetic point clouds to ShapeNet improves MixCon3D zero-shot 3D classification on Objaverse-LVIS, ScanObjectNN, and ModelNet40.
-
3DGraphLLM: Combining Semantic Graphs and Large Language Models for 3D Scene Understanding
A learnable 3D scene graph representation that injects semantic relation embeddings between objects improves LLM performance on 3D grounding, captioning, and question answering.
-
ZeroKey: Point-Level Reasoning and Zero-Shot 3D Keypoint Detection from Large Language Models
ZeroKey detects 3D keypoints on unseen object categories by prompting the Molmo vision-language model on multiple rendered views and aggregating the back-projected points, with no 3D annotations required.
-
GATE-3D: Geometry-Aware Test-time Adaptive Reranking for Open-Set 3D Shape Retrieval
A per-query router trained on cross-modal disagreement features blends a geometry residual into appearance-based 3D-shape retrieval, improving mAP@10 by 2.0 points on OS-ESB-core while avoiding always-on fusion degradation.
-
Fast3D: Accelerating 3D Multi-modal Large Language Models for Efficient 3D Scene Understanding
Fast3D prunes up to 90% of object-centric visual tokens in 3D MLLMs while preserving about 96.8% of original benchmark performance, using a trained attention predictor and adaptive layer-wise pruning.
-
Does Your 3D Encoder Really Work? When Pretrain-SFT from 2D VLMs Meets 3D VLMs
3D scene-centric VLMs underutilize the 3D encoder, overfit to linguistic and answer-frequency shortcuts, and the proposed 3D-RDQA dataset helps expose and mitigate this problem.
-
From Air to Wear: Personalized 3D Digital Fashion with AR/VR Immersive 3D Sketching
A VR-sketch-conditioned diffusion model, trained in three stages with curriculum learning and a new 969-pair dataset, generates plausible 3D garments from freehand 3D sketches.
-
BIP3D: Bridging 2D Images and 3D Perception for Embodied Intelligence
BIP3D is an image-centric 3D perception model that uses 2D foundation model features with explicit 3D position encoding to beat point-cloud-based methods on the EmbodiedScan detection and grounding benchmarks.
-
TinyGiantVLM: A Lightweight Vision-Language Architecture for Spatial Reasoning under Resource Constraints
TinyGiantVLM, a 64M-parameter RGB-D vision-language model with two-phase training, reached 5th place on the AI City Challenge 2025 warehouse spatial reasoning track.
-
Step1X-3D: Towards High-Fidelity and Controllable Generation of Textured 3D Assets
An open-source-oriented image-to-3D system combines a VAE-DiT geometry generator and a diffusion texture module, claiming state-of-the-art quality over open-source rivals and near-proprietary performance.
-
Scaling and Beyond: Advancing Spatial Reasoning in MLLMs Requires New Recipes
The authors argue that spatial reasoning in multimodal LLMs requires dedicated new recipes in data, architecture, and training objectives, and will not emerge from scaling alone.
-
Embodied Intelligence for 3D Understanding: A Survey on 3D Scene Question Answering
A structured survey of 3D Scene Question Answering that categorizes datasets, methods, and metrics and finds a common encoder-fusion-prediction pipeline across approaches.
Discussion (0). Continue with ORCID to comment.