Pith. sign in

REVIEW 21 cited by

LGM: Large Multi-View Gaussian Model for High-Resolution 3D Content Creation

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2402.05054 v1 pith:5WOJ34G4 submitted 2024-02-07 cs.CV

classification cs.CV
keywords multi-viewcontentgaussianhigh-resolutionmodelsbackbonecreationgenerate
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

3D content creation has achieved significant progress in terms of both quality and speed. Although current feed-forward models can produce 3D objects in seconds, their resolution is constrained by the intensive computation required during training. In this paper, we introduce Large Multi-View Gaussian Model (LGM), a novel framework designed to generate high-resolution 3D models from text prompts or single-view images. Our key insights are two-fold: 1) 3D Representation: We propose multi-view Gaussian features as an efficient yet powerful representation, which can then be fused together for differentiable rendering. 2) 3D Backbone: We present an asymmetric U-Net as a high-throughput backbone operating on multi-view images, which can be produced from text or single-view image input by leveraging multi-view diffusion models. Extensive experiments demonstrate the high fidelity and efficiency of our approach. Notably, we maintain the fast speed to generate 3D objects within 5 seconds while boosting the training resolution to 512, thereby achieving high-resolution 3D content generation.

Discussion (0). Sign in to comment.

Forward citations

Cited by 21 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. CORGI: Consistency-Aware 3D Dog Reconstruction from a Single Image in the Wild

    cs.CV 2026-07 unverdicted novelty 7.0 of 10

    CORGI reconstructs high-fidelity, animatable 3D dogs from a single in-the-wild image via canonical orbital generation, deformable 3DGS anchored to D-SMAL, and self-supervised generative repair, without 3D supervision.

  2. MVGBench: Comprehensive Benchmark for Multi-view Generation Models

    cs.GR 2025-06 conditional novelty 7.0 of 10

    MVGBench evaluates multi-view generators through self-consistency of 3D reconstructions and uses this protocol to rank 12 models and build a better one.

  3. GeoWorldAD: Geometry World Action Model for Autonomous Driving

    cs.RO 2026-07 conditional novelty 6.0 of 10

    Grounding an autonomous-driving action model in ego-aligned multi-scale 3D geometry and latent future-geometry tokens improves NAVSIM closed-loop PDMS/EPDMS over prior geometry- and world-model-based planners.

  4. Instant NuRec: Feed-Forward 3D Gaussian Reconstruction for Driving Scene Simulation

    cs.GR 2026-07 conditional novelty 6.0 of 10

    A feed-forward model reconstructs a layered, simulation-ready 3D Gaussian world from multi-view driving video in ~1.5 s, with quality approaching per-scene optimized reconstruction.

  5. Hallo4D: Multi-Modal Hallucination Mitigation for Consistent Spatio-Temporal Generation

    cs.CV 2026-07 conditional novelty 6.0 of 10

    Hallo4D uses vision-language models to detect and correct spatial and temporal mistakes in AI-generated 3D and 4D content, improving consistency without retraining the base generators.

  6. PixGS: Pixel-Space Diffusion for Direct 3D Gaussian Splat Generation

    cs.CV 2026-07 unverdicted novelty 6.0 of 10

    PixGS is a single-stage pixel-space diffusion model that directly produces high-quality 3D Gaussian Splats from text or images in ~1s, outperforming multi-stage latent methods on standard benchmarks.

  7. TInR: Exploring Tool-Internalized Reasoning in Large Language Models

    cs.CL 2026-04 unverdicted novelty 6.0 of 10

    TInR-U internalizes tool knowledge into LLMs via bidirectional alignment, supervised fine-tuning, and reinforcement learning, outperforming standard tool-integrated reasoning in both in-domain and out-of-domain evaluations.

  8. Make Your MoVe: Make Your 3D Contents by Adapting Multi-View Diffusion Models to External Editing

    cs.CV 2025-08 conditional novelty 6.0 of 10

    A tuning-free dual-pipeline that injects original normal latents into an edited multi-view diffusion stream, preserving geometry during 2D-to-3D appearance editing.

  9. Gaussian Variation Field Diffusion for High-fidelity Video-to-4D Synthesis

    cs.CV 2025-07 conditional novelty 6.0 of 10

    A video-to-4D model that encodes mesh animations into compact Gaussian variation latents and diffuses them conditioned on the video and a canonical Gaussian splat.

  10. Masks make discriminative models great again!

    cs.CV 2025-07 conditional novelty 6.0 of 10

    Training a single-image 3D Gaussian splat model on visible regions only, using visibility masks from optimized per-scene splats, improves reconstruction quality in visible areas and stays competitive with full-scene models.

  11. Nabla-R2D3: Effective and Efficient 3D Diffusion Alignment with 2D Rewards

    cs.GR 2025-06 conditional novelty 6.0 of 10

    Nabla-R2D3 aligns 3D-native diffusion models with human preferences by backpropagating multi-view 2D reward gradients through the denoising process, improving reward without destroying the pretrained 3D prior.

  12. Render-FM: Feedforward Model for Real-time Photorealistic Volumetric Rendering

    cs.CV 2025-05 conditional novelty 6.0 of 10

    Render-FM predicts 6D Gaussian splatting parameters from a CT volume in a single feedforward pass, enabling real-time rendering with quality comparable to per-scan optimized methods.

  13. A 3D Facial Reconstruction Evaluation Methodology: Comparing Smartphone Scans with Deep Learning Based Methods Using Geometry and Morphometry Criteria

    cs.CV 2025-02 conditional novelty 6.0 of 10

    A new benchmarking framework using geometric morphometrics shows smartphone 3D facial scans preserve shape better than deep learning reconstructions from 2D images.

  14. Large Images are Gaussians: High-Quality Large Image Representation with Levels of 2D Gaussian Splatting

    cs.CV 2025-02 conditional novelty 6.0 of 10

    A two-level 2D Gaussian splatting method with direct covariance optimization fits large images with more Gaussian points and higher PSNR than prior Gaussian-based image representation.

  15. Few-step Flow for 3D Generation via Marginal-Data Transport Distillation

    cs.CV 2025-09 conditional novelty 5.0 of 10

    MDT-dist distills a pretrained 3D flow model into a 1-2 step generator using velocity matching plus velocity distillation, cutting TRELLIS inference from 6.1s to 0.68s while approximately preserving generation quality.

  16. Category-Aware 3D Object Composition with Disentangled Texture and Shape Multi-view Diffusion

    cs.CV 2025-09 conditional novelty 5.0 of 10

    C33D blends a 3D model with an object category by generating a fused front view, then using texture and shape multi-view diffusion plus adaptive inversion to reconstruct a novel, consistent 3D model.

  17. SMPL Normal Map Is All You Need for Single-view Textured Human Reconstruction

    cs.CV 2025-06 conditional novelty 5.0 of 10

    By fine-tuning the LGM object reconstruction model with SMPL normal maps and a normal-map constraint, SEHR reconstructs full clothed 3D humans from one image in a single forward pass.

  18. ShapeLLM-Omni: A Native Multimodal LLM for 3D Generation and Understanding

    cs.CV 2025-06 conditional novelty 5.0 of 10

    ShapeLLM-Omni unifies text, image, and 3D generation and understanding in one autoregressive LLM using discrete 3D tokens and a new 3D-Alpaca training dataset.

  19. ConsistentDreamer: View-Consistent Meshes Through Balanced Multi-View Gaussian Optimization

    cs.CV 2025-02 conditional novelty 5.0 of 10

    Generating six reference views first, then combining closest-view SDS with pixel-level reconstruction and automatic uncertainty weights, improves cross-view consistency in single-image 3D mesh generation.

  20. CA-World: Multi-Object Counterfactual Alignment for Efficient Interactive-Ready Reconstruction

    cs.CV 2026-05 unverdicted novelty 4.0 of 10

    The paper's stated CA-World counterfactual claim is absent from the body, which instead describes the SAM3D-Phys pipeline for multi-object interactive reconstruction and simulation.

  21. SAT: Supervisor Regularization and Animation Augmentation for Two-process Monocular Texture 3D Human Reconstruction

    cs.CV 2025-08 conditional novelty 4.0 of 10

    A two-stage Gaussian-splatting framework with supervisor feature regularization and online animation augmentation improves monocular textured 3D human reconstruction on CustomHuman and THuman3.0.

Pith tools