Pith. sign in

REVIEW 7 cited by

Face Adapter for Pre-Trained Diffusion Models with Fine-Grained ID and Attribute Control

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2405.12970 v2 pith:JR34V7LU submitted 2024-05-21 cs.CV

classification cs.CV
keywords facemodelsattributecontroldiffusionface-adapterpre-trainedreenactment
verification ladder T0 review T1 audit T2 compute T3 formal

Signed reviews

No signed human review yet.

0 comments
read the original abstract

Current face reenactment and swapping methods mainly rely on GAN frameworks, but recent focus has shifted to pre-trained diffusion models for their superior generation capabilities. However, training these models is resource-intensive, and the results have not yet achieved satisfactory performance levels. To address this issue, we introduce Face-Adapter, an efficient and effective adapter designed for high-precision and high-fidelity face editing for pre-trained diffusion models. We observe that both face reenactment/swapping tasks essentially involve combinations of target structure, ID and attribute. We aim to sufficiently decouple the control of these factors to achieve both tasks in one model. Specifically, our method contains: 1) A Spatial Condition Generator that provides precise landmarks and background; 2) A Plug-and-play Identity Encoder that transfers face embeddings to the text space by a transformer decoder. 3) An Attribute Controller that integrates spatial conditions and detailed attributes. Face-Adapter achieves comparable or even superior performance in terms of motion control precision, ID retention capability, and generation quality compared to fully fine-tuned face reenactment/swapping models. Additionally, Face-Adapter seamlessly integrates with various StableDiffusion models.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 7 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. VideoMaker: Zero-shot Customized Video Generation with the Inherent Force of Video Diffusion Models

    cs.CV 2024-12 conditional novelty 6.0 of 10

    VideoMaker shows that a video diffusion model can itself extract and inject reference-subject features, using reference-frame concatenation and self-attention, achieving state-of-the-art zero-shot customized video generation.

  2. DreamFit: Garment-Centric Human Generation via a Lightweight Anything-Dressing Encoder

    cs.CV 2024-12 conditional novelty 6.0 of 10

    DreamFit generates human images from a garment reference and text by encoding the reference through LoRA-activated layers of a frozen Stable Diffusion UNet and injecting features with adaptive attention.

  3. Beyond Generation: Unlocking Universal Editing via Self-Supervised Fine-Tuning

    cs.CV 2024-12 conditional novelty 6.0 of 10

    UES adds a self-supervised video condition to text-to-video diffusion models, enabling them to edit videos from delta prompts without paired supervision.

  4. DreamID: High-Fidelity and Fast diffusion-based Face Swapping via Triplet ID Group Learning

    cs.CV 2025-04 conditional novelty 5.0 of 10

    DreamID trains a one-step diffusion face swapper with triplet ID groups (source, pseudo target, real target) to jointly optimize identity transfer and attribute preservation.

  5. DynamicFace: High-Quality and Consistent Face Swapping for Image and Video using Composable 3D Facial Priors

    cs.CV 2025-01 conditional novelty 5.0 of 10

    DynamicFace reports a diffusion-based face swapper with four disentangled 3D facial conditions and a temporal TV optimizer, beating prior methods on some FF++ metrics but not on pose or expression.

  6. ArtCrafter: Text-Image Aligning Style Transfer via Embedding Reframing

    cs.CV 2025-01 conditional novelty 5.0 of 10

    ArtCrafter improves text-guided style transfer by extracting style with perceiver attention, aligning image and text embeddings, and blending them through explicit modulation.

  7. HiFiVFS: High Fidelity Video Face Swapping

    cs.CV 2024-11 conditional novelty 5.0 of 10

    HiFiVFS applies SVD to video face swapping with identity-desensitized attribute features and detailed identity tokens, claiming state-of-the-art fidelity and temporal consistency.

Pith tools