Pith. sign in

REVIEW 2 cited by

TV-3DG: Mastering Text-to-3D Customized Generation with Visual Prompt

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2410.21299 v2 pith:QBTJ54NX submitted 2024-10-16 cs.CV

classification cs.CV
keywords generationcustomizedtermvisualdifferenceduringnoiseprompt
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

In recent years, advancements in generative models have significantly expanded the capabilities of text-to-3D generation. Many approaches rely on Score Distillation Sampling (SDS) technology. However, SDS struggles to accommodate multi-condition inputs, such as text and visual prompts, in customized generation tasks. To explore the core reasons, we decompose SDS into a difference term and a classifier-free guidance term. Our analysis identifies the core issue as arising from the difference term and the random noise addition during the optimization process, both contributing to deviations from the target mode during distillation. To address this, we propose a novel algorithm, Classifier Score Matching (CSM), which removes the difference term in SDS and uses a deterministic noise addition process to reduce noise during optimization, effectively overcoming the low-quality limitations of SDS in our customized generation framework. Based on CSM, we integrate visual prompt information with an attention fusion mechanism and sampling guidance techniques, forming the Visual Prompt CSM (VPCSM) algorithm. Furthermore, we introduce a Semantic-Geometry Calibration (SGC) module to enhance quality through improved textual information integration. We present our approach as TV-3DG, with extensive experiments demonstrating its capability to achieve stable, high-quality, customized 3D generation. Project page: \url{https://yjhboy.github.io/TV-3DG}

Discussion (0). Sign in to comment.

Forward citations

Cited by 2 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Empowering Vector Graphics with Consistently Arbitrary Viewing and View-dependent Visibility

    cs.CV 2025-05 conditional novelty 6.0 of 10

    Dream3DVG couples a 3D Gaussian Splatting branch with a 3D vector-graphics branch to generate text-driven sketches and icons that stay consistent across views and cull occluded strokes.

  2. QR-LoRA: Efficient and Disentangled Fine-tuning via QR Decomposition for Customized Generation

    cs.CV 2025-07 conditional novelty 4.0 of 10

    QR-LoRA freezes the QR-decomposed basis of pretrained weights, trains only a residual matrix, and reports halved trainable parameters with improved content-style disentanglement in diffusion models.

Pith tools