Pith. sign in

REVIEW 3 cited by

Model Inversion Attacks Through Target-Specific Conditional Diffusion Models

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2407.11424 v2 pith:RALNPKXC submitted 2024-07-16 cs.CV

classification cs.CV
keywords modeltargetclassifierdiffusionattacksdiff-miinversionmodels
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Model inversion attacks (MIAs) aim to reconstruct private images from a target classifier's training set, thereby raising privacy concerns in AI applications. Previous GAN-based MIAs tend to suffer from inferior generative fidelity due to GAN's inherent flaws and biased optimization within latent space. To alleviate these issues, leveraging on diffusion models' remarkable synthesis capabilities, we propose Diffusion-based Model Inversion (Diff-MI) attacks. Specifically, we introduce a novel target-specific conditional diffusion model (CDM) to purposely approximate target classifier's private distribution and achieve superior accuracy-fidelity balance. Our method involves a two-step learning paradigm. Step-1 incorporates the target classifier into the entire CDM learning under a pretrain-then-finetune fashion, with creating pseudo-labels as model conditions in pretraining and adjusting specified layers with image predictions in fine-tuning. Step-2 presents an iterative image reconstruction method, further enhancing the attack performance through a combination of diffusion priors and target knowledge. Additionally, we propose an improved max-margin loss that replaces the hard max with top-k maxes, fully leveraging feature information and soft labels from the target classifier. Extensive experiments demonstrate that Diff-MI significantly improves generative fidelity with an average decrease of 20\% in FID while maintaining competitive attack accuracy compared to state-of-the-art methods across various datasets and models. Our code is available at: \url{https://github.com/Ouxiang-Li/Diff-MI}.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 3 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Precise, Fast, and Low-cost Concept Erasure in Value Space: Orthogonal Complement Matters

    cs.CV 2024-12 conditional novelty 6.0 of 10

    AdaVD removes target concepts from diffusion models by soft-projecting value vectors away from the target token direction, with a sigmoid threshold that preserves unrelated prompts.

  2. Do Speech Tokens Leak Voiceprints? Speaker Inversion Attacks Against End-to-End Speech Language Models

    cs.SD 2026-07 conditional novelty 5.0 of 10

    Three seconds of speech tokens from Moshi, Higgs3, Kimi-Audio, or Qwen3-Omni allow a trained inversion model to recover speaker embeddings with cosine similarity above 0.70 against a pretrained speaker encoder.

  3. Deep Learning Model Inversion Attacks and Defenses: A Comprehensive Survey

    cs.CR 2025-01 accept novelty 4.0 of 10

    A structured literature review that taxonomizes model inversion attacks and defenses and provides a public resource repository.

Pith tools