Pith. sign in

REVIEW 1 cited by

Viewpoint Textual Inversion: Discovering Scene Representations and 3D View Control in 2D Diffusion Models

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2309.07986 v2 pith:NU5KHMMY submitted 2023-09-14 cs.CV cs.AIcs.LG

classification cs.CVcs.AIcs.LG
keywords viewviewpointdiffusionscenemodelsimagesrepresentationstoken
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Text-to-image diffusion models generate impressive and realistic images, but do they learn to represent the 3D world from only 2D supervision? We demonstrate that yes, certain 3D scene representations are encoded in the text embedding space of models like Stable Diffusion. Our approach, Viewpoint Neural Textual Inversion (ViewNeTI), is to discover 3D view tokens; these tokens control the 3D viewpoint - the rendering pose in a scene - of generated images. Specifically, we train a small neural mapper to take continuous camera viewpoint parameters and predict a view token (a word embedding). This token conditions diffusion generation via cross-attention to produce images with the desired camera viewpoint. Using ViewNeTI as an evaluation tool, we report two findings: first, the text latent space has a continuous view-control manifold for particular 3D scenes; second, we find evidence for a generalized view-control manifold for all scenes. We conclude that since the view token controls the 3D `rendering' viewpoint, there is likely a scene representation embedded in frozen 2D diffusion models. Finally, we exploit the 3D scene representations for 3D vision tasks, namely, view-controlled text-to-image generation, and novel view synthesis from a single image, where our approach sets state-of-the-art for LPIPS. Code available at https://github.com/jmhb0/view_neti

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 1 Pith paper

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. FoundHand: Large-Scale Domain-Specific Learning for Controllable Hand Image Generation

    cs.CV 2024-12 conditional novelty 6.0 of 10

    A 2D-keypoint-conditioned diffusion model trained on a new 10M-image hand dataset enables controllable hand reposing, appearance transfer, novel view synthesis, and zero-shot hand video generation.

Pith tools