REVIEW 4 cited by
StableGarment: Garment-Centric Generation via Stable Diffusion
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
read the original abstract
In this paper, we introduce StableGarment, a unified framework to tackle garment-centric(GC) generation tasks, including GC text-to-image, controllable GC text-to-image, stylized GC text-to-image, and robust virtual try-on. The main challenge lies in retaining the intricate textures of the garment while maintaining the flexibility of pre-trained Stable Diffusion. Our solution involves the development of a garment encoder, a trainable copy of the denoising UNet equipped with additive self-attention (ASA) layers. These ASA layers are specifically devised to transfer detailed garment textures, also facilitating the integration of stylized base models for the creation of stylized images. Furthermore, the incorporation of a dedicated try-on ControlNet enables StableGarment to execute virtual try-on tasks with precision. We also build a novel data engine that produces high-quality synthesized data to preserve the model's ability to follow prompts. Extensive experiments demonstrate that our approach delivers state-of-the-art (SOTA) results among existing virtual try-on methods and exhibits high flexibility with broad potential applications in various garment-centric image generation.
Forward citations
Cited by 4 Pith papers
-
Dress&Dance: Dress up and Dance as You Like It - Technical Preview
A video diffusion framework that unifies text, image, and video conditioning through attention to produce high-resolution virtual try-on videos with reference-driven motion.
-
Stable-Hair v2: Real-World Hair Transfer via Multiple-View Diffusion Model
A multi-view diffusion hair-transfer system transfers a reference hairstyle onto a portrait and renders the edited person from many consistent viewpoints.
-
ID-Cloak: Crafting Identity-Specific Cloaks Against Personalized Text-to-Image Generation
ID-Cloak creates a single universal image perturbation from a few photos that degrades personalized text-to-image generation of that identity across unseen images.
-
CONVERGE: A Multi-Agent Vision-Radio Architecture for xApps
CONVERGE fuses camera and radio sensing inside O-RAN xApps via a multi-agent architecture, reporting under-one-millisecond sensing delay for real-time blockage-driven RAN control.
Discussion (0). Continue with ORCID to comment.