Pith. sign in

REVIEW 1 cited by

Enhancing Variational Autoencoders with Smooth Robust Latent Encoding

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2504.17219 v1 pith:G2FP3ZHS submitted 2025-04-24 cs.LG cs.AIcs.CR

classification cs.LGcs.AIcs.CR
keywords robustnessadversarialfidelityimagemodelstraininggenerativelatent
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Variational Autoencoders (VAEs) have played a key role in scaling up diffusion-based generative models, as in Stable Diffusion, yet questions regarding their robustness remain largely underexplored. Although adversarial training has been an established technique for enhancing robustness in predictive models, it has been overlooked for generative models due to concerns about potential fidelity degradation by the nature of trade-offs between performance and robustness. In this work, we challenge this presumption, introducing Smooth Robust Latent VAE (SRL-VAE), a novel adversarial training framework that boosts both generation quality and robustness. In contrast to conventional adversarial training, which focuses on robustness only, our approach smooths the latent space via adversarial perturbations, promoting more generalizable representations while regularizing with originality representation to sustain original fidelity. Applied as a post-training step on pre-trained VAEs, SRL-VAE improves image robustness and fidelity with minimal computational overhead. Experiments show that SRL-VAE improves both generation quality, in image reconstruction and text-guided image editing, and robustness, against Nightshade attacks and image editing attacks. These results establish a new paradigm, showing that adversarial training, once thought to be detrimental to generative models, can instead enhance both fidelity and robustness.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 1 Pith paper

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. LAFR: Efficient Diffusion-based Blind Face Restoration via Latent Codebook Alignment Adapter

    cs.CV 2025-05 reject novelty 4.0 of 10

    LAFR uses a 1024-entry codebook adapter to map low-quality face latents into the high-quality latent space of Stable Diffusion, then LoRA-tunes a pruned UNet on just 600 FFHQ images for blind face restoration.

Pith tools