Pith. sign in

REVIEW 16 cited by

Instant3D: Fast Text-to-3D with Sparse-View Generation and Large Reconstruction Model

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2311.06214 v2 pith:L3ZL5DRK submitted 2023-11-10 cs.CV

classification cs.CV
keywords instant3dmethodsassetsdiffusiondiversefeed-forwardgenerategenerates
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Text-to-3D with diffusion models has achieved remarkable progress in recent years. However, existing methods either rely on score distillation-based optimization which suffer from slow inference, low diversity and Janus problems, or are feed-forward methods that generate low-quality results due to the scarcity of 3D training data. In this paper, we propose Instant3D, a novel method that generates high-quality and diverse 3D assets from text prompts in a feed-forward manner. We adopt a two-stage paradigm, which first generates a sparse set of four structured and consistent views from text in one shot with a fine-tuned 2D text-to-image diffusion model, and then directly regresses the NeRF from the generated images with a novel transformer-based sparse-view reconstructor. Through extensive experiments, we demonstrate that our method can generate diverse 3D assets of high visual quality within 20 seconds, which is two orders of magnitude faster than previous optimization-based methods that can take 1 to 10 hours. Our project webpage: https://jiahao.ai/instant3d/.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 16 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. DGS-LRM: Real-Time Deformable 3D Gaussian Reconstruction From Monocular Videos

    cs.GR 2025-06 conditional novelty 7.0 of 10

    A single feed-forward transformer predicts per-pixel deformable 3D Gaussians with dense scene flow from a posed monocular video, enabling real-time dynamic view synthesis and 3D tracking.

  2. AnimateAnyMesh: A Feed-Forward 4D Foundation Model for Text-Driven Universal Mesh Animation

    cs.CV 2025-06 conditional novelty 7.0 of 10

    A feed-forward VAE plus rectified-flow model animates arbitrary static meshes from text prompts in seconds, with a new 4M-sequence training dataset.

  3. Meshy T2: Fast Native Mesh Generation with Flow Matching

    cs.GR 2026-07 conditional novelty 6.0 of 10

    Single-image native mesh generation runs at interactive speed in Meshy T2 by flow-matching one continuous latent per vertex, then decoding vertices, edge connectivity, and face winding in one pass.

  4. GeoWorldAD: Geometry World Action Model for Autonomous Driving

    cs.RO 2026-07 conditional novelty 6.0 of 10

    Grounding an autonomous-driving action model in ego-aligned multi-scale 3D geometry and latent future-geometry tokens improves NAVSIM closed-loop PDMS/EPDMS over prior geometry- and world-model-based planners.

  5. Muses: Designing, Composing, Generating Nonexistent Fantasy 3D Creatures without Training

    cs.CV 2026-01 conditional novelty 6.0 of 10

    Muses creates new fantasy 3D animals by designing a combined skeleton, fusing voxel parts from separate 3D models along that skeleton, then restyling textures via image editing — with no training.

  6. Let Language Constrain Geometry: Vision-Language Models as Semantic and Spatial Critics for 3D Generation

    cs.CV 2025-11 conditional novelty 6.0 of 10

    A frozen VLM's dual-query Yes/No log-odds act as a differentiable semantic-and-spatial critic, improving alignment and geometry in both SDS-based and feed-forward text-to-3D pipelines.

  7. Make Your MoVe: Make Your 3D Contents by Adapting Multi-View Diffusion Models to External Editing

    cs.CV 2025-08 conditional novelty 6.0 of 10

    A tuning-free dual-pipeline that injects original normal latents into an edited multi-view diffusion stream, preserving geometry during 2D-to-3D appearance editing.

  8. CharacterShot: Controllable and Consistent 4D Character Animation

    cs.CV 2025-08 conditional novelty 6.0 of 10

    A new pipeline generates pose-controlled, view-consistent 4D character animations from one reference image and a 2D pose sequence, backed by a new 13,115-character dataset and benchmark.

  9. DualMat: PBR Material Estimation via Coherent Dual-Path Diffusion

    cs.CV 2025-08 conditional novelty 6.0 of 10

    DualMat is a dual-path diffusion model combining an albedo-optimized pretrained latent path with a material-specialized compact latent path, using feature distillation and rectified flow to estimate PBR materials from...

  10. Can3Tok: Canonical 3D Tokenization and Latent Modeling of Scene-Level 3D Gaussians

    cs.CV 2025-08 conditional novelty 6.0 of 10

    Can3Tok tokenizes scene-level 3D Gaussian splats into canonical latent tokens with normalization and saliency filtering, enabling reconstruction and text/image-to-3D generation.

  11. SeqTex: Generate Mesh Textures in Video Sequence

    cs.CV 2025-07 conditional novelty 6.0 of 10

    SeqTex adapts a pretrained video diffusion model to directly generate complete UV texture maps by jointly predicting four multi-view images and the UV map as a five-frame sequence.

  12. Pro3D-Editor : A Progressive-Views Perspective for Consistent and Precise 3D Editing

    cs.GR 2025-05 conditional novelty 6.0 of 10

    Pro3D-Editor chooses the most editing-salient view, propagates the edit to other key views with per-view LoRA experts, and refines the 3D scene, improving multi-view consistency.

  13. Few-step Flow for 3D Generation via Marginal-Data Transport Distillation

    cs.CV 2025-09 conditional novelty 5.0 of 10

    MDT-dist distills a pretrained 3D flow model into a 1-2 step generator using velocity matching plus velocity distillation, cutting TRELLIS inference from 6.1s to 0.68s while approximately preserving generation quality.

  14. ObjFiller3D: Scaling 3D Object Inpainting to Dense Multi-View Consistency

    cs.CV 2025-08 unverdicted novelty 5.0 of 10

    ObjFiller3D jointly optimizes a dense 360-degree ring of views to inpaint 3D objects with cross-view-consistent textures, reporting higher PSNR and LPIPS than per-view baselines at much lower runtime.

  15. Collaborative Multi-Modal Coding for High-Quality 3D Generation

    cs.CV 2025-08 conditional novelty 5.0 of 10

    TriMM fuses RGB, RGB-D, and point-cloud encoding into a shared triplane latent space and generates 3D assets from a single image with a latent diffusion model.

  16. ShapeLLM-Omni: A Native Multimodal LLM for 3D Generation and Understanding

    cs.CV 2025-06 conditional novelty 5.0 of 10

    ShapeLLM-Omni unifies text, image, and 3D generation and understanding in one autoregressive LLM using discrete 3D tokens and a new 3D-Alpaca training dataset.

Pith tools