REVIEW 5 cited by
StyleAlign: Analysis and Applications of Aligned StyleGAN Models
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
read the original abstract
In this paper, we perform an in-depth study of the properties and applications of aligned generative models. We refer to two models as aligned if they share the same architecture, and one of them (the child) is obtained from the other (the parent) via fine-tuning to another domain, a common practice in transfer learning. Several works already utilize some basic properties of aligned StyleGAN models to perform image-to-image translation. Here, we perform the first detailed exploration of model alignment, also focusing on StyleGAN. First, we empirically analyze aligned models and provide answers to important questions regarding their nature. In particular, we find that the child model's latent spaces are semantically aligned with those of the parent, inheriting incredibly rich semantics, even for distant data domains such as human faces and churches. Second, equipped with this better understanding, we leverage aligned models to solve a diverse set of tasks. In addition to image translation, we demonstrate fully automatic cross-domain image morphing. We further show that zero-shot vision tasks may be performed in the child domain, while relying exclusively on supervision in the parent domain. We demonstrate qualitatively and quantitatively that our approach yields state-of-the-art results, while requiring only simple fine-tuning and inversion.
Forward citations
Cited by 5 Pith papers
-
USO: Unified Style and Subject-Driven Generation via Disentangled and Reward Learning
USO trains one DiT model for subject-driven, style-driven, and joint generation by disentangling content and style from triplet data and adding a style-reward objective, claiming SOTA on USO-Bench.
-
OmniConsistency: Learning Style-Agnostic Consistency from Paired Stylization Data
OmniConsistency is a style-agnostic consistency module for Flux that preserves structure and details during stylization with arbitrary LoRAs, reaching GPT-4o-level content consistency.
-
Steering Guidance for Personalized Text-to-Image Diffusion Models
Weight-interpolated null-text weak model in classifier-free guidance improves subject fidelity with minimal text-fidelity loss in personalized text-to-image diffusion.
-
Hybrid Scandium Aluminum Nitride/Silicon Nitride Integrated Photonic Circuits
The abstract reports a low-loss ScAlN/Si3N4 hybrid waveguide, but the full text is an unrelated diffusion-model paper, leaving the photonics claim without supporting evidence.
-
Advancing Facial Stylization through Semantic Preservation Constraint and Pseudo-Paired Supervision
A StyleGAN fine-tuning recipe with semantic preservation and multi-level pseudo-paired supervision yields higher-fidelity facial stylization, plus free multimodal and reference-guided variants.
Discussion (0). Sign in to comment.