REVIEW 4 cited by
Orient Anything: Learning Robust Object Orientation Estimation from Rendering 3D Models
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
read the original abstract
Orientation is a key attribute of objects, crucial for understanding their spatial pose and arrangement in images. However, practical solutions for accurate orientation estimation from a single image remain underexplored. In this work, we introduce Orient Anything, the first expert and foundational model designed to estimate object orientation in a single- and free-view image. Due to the scarcity of labeled data, we propose extracting knowledge from the 3D world. By developing a pipeline to annotate the front face of 3D objects and render images from random views, we collect 2M images with precise orientation annotations. To fully leverage the dataset, we design a robust training objective that models the 3D orientation as probability distributions of three angles and predicts the object orientation by fitting these distributions. Besides, we employ several strategies to improve synthetic-to-real transfer. Our model achieves state-of-the-art orientation estimation accuracy in both rendered and real images and exhibits impressive zero-shot ability in various scenarios. More importantly, our model enhances many applications, such as comprehension and generation of complex spatial concepts and 3D object pose adjustment.
Forward citations
Cited by 4 Pith papers
-
GenSpace: Benchmarking Spatially-Aware Image Generation
GenSpace benchmarks spatial awareness in image generation with a 3D reconstruction-based evaluator, showing models struggle with allocentric relations and metric measurements.
-
MarineEVT: Advancing Event-Centric Marine Video Understanding via Visual Tool Reasoning
On MarineEVT, an event-centric 20K-pair marine video QA benchmark, EVT-R1 with tool-integrated RL scores 48.89 average accuracy, 5.22 points above the best untuned open-source VLM and 8.54 above the best tool-using co...
-
UniPose9D: Universal Category-Agnostic Object Pose Estimation
A single category-agnostic model recovers metric 9D object pose from one masked RGB-D observation via point-pair NOCS prediction, flow matching, and adaptive N-hop Kabsch–Umeyama.
-
Disentangling 3D Modeling from Spatial Reasoning
DiSR answers spatial questions by feeding a language model a text summary of metric 3D positions, sizes, and orientations from frozen expert models, reaching top scores on two benchmarks with 59 GPU-hours of training.
Discussion (0). Continue with ORCID to comment.