REVIEW 12 cited by
MeshGPT: Generating Triangle Meshes with Decoder-Only Transformers
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
Signed reviews
read the original abstract
We introduce MeshGPT, a new approach for generating triangle meshes that reflects the compactness typical of artist-created meshes, in contrast to dense triangle meshes extracted by iso-surfacing methods from neural fields. Inspired by recent advances in powerful large language models, we adopt a sequence-based approach to autoregressively generate triangle meshes as sequences of triangles. We first learn a vocabulary of latent quantized embeddings, using graph convolutions, which inform these embeddings of the local mesh geometry and topology. These embeddings are sequenced and decoded into triangles by a decoder, ensuring that they can effectively reconstruct the mesh. A transformer is then trained on this learned vocabulary to predict the index of the next embedding given previous embeddings. Once trained, our model can be autoregressively sampled to generate new triangle meshes, directly generating compact meshes with sharp edges, more closely imitating the efficient triangulation patterns of human-crafted meshes. MeshGPT demonstrates a notable improvement over state of the art mesh generation methods, with a 9% increase in shape coverage and a 30-point enhancement in FID scores across various categories.
Forward citations
Cited by 12 Pith papers
-
GaussianGPT: Towards Autoregressive 3D Gaussian Scene Generation
A causal transformer with 3D RoPE generates vector-quantized 3D Gaussian latent grids autoregressively, enabling unconditional synthesis, completion, and open-ended outpainting of indoor scenes.
-
Unifi3D: A Study on 3D Representations for Generation and Reconstruction in a Common Framework
SDF grids reconstruct best, Dual Octrees score best on automatic generation metrics, but users prefer SDF output, and reconstruction plus compression errors make up a large share of generation error.
-
Auto-Regressive Surface Cutting
SeamGPT generates artist-style mesh cutting seams as auto-regressively predicted quantized 3D line segments, improving UV unwrapping and part segmentation.
-
Vector Representations of Vessel Trees
VeTTA encodes a vascular tree into one vector and recursively decodes it into a geometrically accurate, topologically valid tree, outperforming voxel-based autoencoders on reconstruction metrics.
-
BAG: Body-Aligned 3D Wearable Asset Generation
BAG generates body-aligned 3D wearable assets from a single image by conditioning multi-view diffusion on canonical body XYZ maps and refining alignment with Sim(3) optimization and physics simulation.
-
Don't Mesh with Me: Generating Constructive Solid Geometry Instead of Meshes by Fine-Tuning a Code-Generation LLM
A code-generation LLM fine-tuned on BREP-to-CSG Python scripts can complete simple mechanical geometries from positional input and, for simple cases, from natural language descriptions.
-
GaussianAnything: Interactive Point Cloud Flow Matching For 3D Object Generation
A point-cloud-structured latent space with cascaded flow matching enables high-quality text- and image-conditioned 3D object generation and interactive editing.
-
CNS-Edit++: Category-Agnostic 3D Editing with Coupled Neural Shape Representation
Coupling a global latent code with a 3D feature volume lets off-the-shelf 3D generators perform local semantic edits — copy, delete, resize, mix, and drag — across object categories while preserving unedited regions.
-
FreeMesh: Boosting Mesh Generation with Coordinates Merging
FreeMesh shows that rearranging coordinates into same-axis groups before byte-pair encoding reduces per-token entropy and sequence length, improving point-cloud conditioned mesh generation quality.
-
Reducing the Sensitivity of Neural Physics Simulators to Mesh Topology via Pretraining
Autoencoder pretraining of graph mesh encoders makes neural radar simulators less sensitive to shape-preserving changes in mesh topology.
-
Visual Large Language Models for Generalized and Specialized Applications
This paper reviews and taxonomizes VLLM applications into vision-to-text, vision-to-action, and text-to-vision, adding ethics and future-work discussion.
-
Machine learning for modelling unstructured grid data in computational physics: a review
A broad review of machine learning techniques for modeling unstructured mesh data in computational physics, with a taxonomy, a qualitative comparison, and a list of public benchmarks.
Discussion (0). Continue with ORCID to comment.