Pith. sign in

REVIEW 1 cited by

HanDiffuser: Text-to-Image Generation With Realistic Hand Appearances

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2403.01693 v3 pith:KB4Z75SW submitted 2024-03-04 cs.CV cs.AI

classification cs.CVcs.AI
keywords handgeneratehandiffuserhandsimagesdiffusionfingergenerating
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Text-to-image generative models can generate high-quality humans, but realism is lost when generating hands. Common artifacts include irregular hand poses, shapes, incorrect numbers of fingers, and physically implausible finger orientations. To generate images with realistic hands, we propose a novel diffusion-based architecture called HanDiffuser that achieves realism by injecting hand embeddings in the generative process. HanDiffuser consists of two components: a Text-to-Hand-Params diffusion model to generate SMPL-Body and MANO-Hand parameters from input text prompts, and a Text-Guided Hand-Params-to-Image diffusion model to synthesize images by conditioning on the prompts and hand parameters generated by the previous component. We incorporate multiple aspects of hand representation, including 3D shapes and joint-level finger positions, orientations and articulations, for robust learning and reliable performance during inference. We conduct extensive quantitative and qualitative experiments and perform user studies to demonstrate the efficacy of our method in generating images with high-quality hands.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 1 Pith paper

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Detecting Human Artifacts from Text-to-Image Models

    cs.CV 2024-11 conditional novelty 6.0 of 10

    A new dataset and detector suite localize human artifacts in text-to-image outputs and feed back into generation to reduce them.

Pith tools