REVIEW 6 cited by
Anything-3D: Towards Single-view Anything Reconstruction in the Wild
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
read the original abstract
3D reconstruction from a single-RGB image in unconstrained real-world scenarios presents numerous challenges due to the inherent diversity and complexity of objects and environments. In this paper, we introduce Anything-3D, a methodical framework that ingeniously combines a series of visual-language models and the Segment-Anything object segmentation model to elevate objects to 3D, yielding a reliable and versatile system for single-view conditioned 3D reconstruction task. Our approach employs a BLIP model to generate textural descriptions, utilizes the Segment-Anything model for the effective extraction of objects of interest, and leverages a text-to-image diffusion model to lift object into a neural radiance field. Demonstrating its ability to produce accurate and detailed 3D reconstructions for a wide array of objects, \emph{Anything-3D\footnotemark[2]} shows promise in addressing the limitations of existing methodologies. Through comprehensive experiments and evaluations on various datasets, we showcase the merits of our approach, underscoring its potential to contribute meaningfully to the field of 3D reconstruction. Demos and code will be available at \href{https://github.com/Anything-of-anything/Anything-3D}{https://github.com/Anything-of-anything/Anything-3D}.
Forward citations
Cited by 6 Pith papers
-
LISA-3D: Lifting Language-Image Segmentation to 3D via Multi-View Consistency
Geometry-aware LoRA tuning of LISA with differentiable reprojection yields view-consistent masks that lift to better 3D reconstructions through frozen SAM-3D.
-
Towards Fine-grained Interactive Segmentation in Images and Videos
SAM2Refiner adds localization, prompt-retargeting and mask-refinement modules to SAM2, and reports state-of-the-art fine-grained segmentation on four image and two video benchmarks.
-
Pippo: High-Resolution Multi-View Humans from a Single Image
A single-image multi-view diffusion transformer generates 1K-resolution turnaround views of humans, with attention biasing for many views and a new reprojection-error metric.
-
SAMRefiner: Taming Segment Anything Model for Universal Mask Refinement
A prompting scheme that mines points, elastic boxes, and Gaussian-style masks from coarse masks lets SAM refine those masks more accurately than prior refinement tools.
-
SERES: Semantic-aware neural reconstruction from sparse views
A semantic-aware implicit reconstruction method claims 44% and 20% lower Chamfer distance than SparseNeuS and VolRecon, and 69%/68% error reductions as a NeuS/Neuralangelo plugin.
-
scI2CL: Effectively Integrating Single-cell Multi-omics by Intra- and Inter-omics Contrastive Learning
The abstract claims a state-of-the-art single-cell multi-omics integration method with new cell-subtype and trajectory findings, but the supplied full text is a different paper.
Discussion (0). Continue with ORCID to comment.