REVIEW 7 cited by
Wear-Any-Way: Manipulable Virtual Try-on via Sparse Correspondence Alignment
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
read the original abstract
This paper introduces a novel framework for virtual try-on, termed Wear-Any-Way. Different from previous methods, Wear-Any-Way is a customizable solution. Besides generating high-fidelity results, our method supports users to precisely manipulate the wearing style. To achieve this goal, we first construct a strong pipeline for standard virtual try-on, supporting single/multiple garment try-on and model-to-model settings in complicated scenarios. To make it manipulable, we propose sparse correspondence alignment which involves point-based control to guide the generation for specific locations. With this design, Wear-Any-Way gets state-of-the-art performance for the standard setting and provides a novel interaction form for customizing the wearing style. For instance, it supports users to drag the sleeve to make it rolled up, drag the coat to make it open, and utilize clicks to control the style of tuck, etc. Wear-Any-Way enables more liberated and flexible expressions of the attires, holding profound implications in the fashion industry.
Forward citations
Cited by 7 Pith papers
-
VTBench: Comprehensive Benchmark Suite Towards Real-World Virtual Try-on Models
VTBench is a multi-dimensional benchmark with novel unpaired metrics and human preference data for evaluating image-based virtual try-on models, though the human-alignment evidence is incomplete.
-
PromptDresser: Improving the Quality and Controllability of Virtual Try-On via Generative Textual Prompt and Prompt-aware Mask
PromptDresser improves text-editable virtual try-on by combining LMM-generated structured captions with a prompt-aware adaptive mask.
-
DiffusionTrend: A Minimalist Approach to Virtual Fashion Try-On
A training-free virtual try-on pipeline that blends DDIM-inverted garment latents into masked model latents, guided by a lightweight CNN apparel mask.
-
FashionComposer: Compositional Fashion Image Generation
A single diffusion framework composes multiple garment and face references into one fashion image using an asset library and subject-binding attention.
-
CatV2TON: Taming Diffusion Transformers for Vision-Based Virtual Try-On with Temporal Concatenation
CatV2TON unifies image and video virtual try-on in one diffusion transformer, using temporal garment-person concatenation and clip-based inference with AdaCN for long, consistent try-on videos.
-
PEMF-VTO: Point-Enhanced Video Virtual Try-on via Mask-free Paradigm
A mask-free video virtual try-on model that uses sparse point correspondences between garment and frames, plus frame-to-frame tracking, to improve garment transfer and temporal coherence.
-
Advancing Myopia To Holism: Fully Contrastive Language-Image Pre-training
Holistic CLIP trains a multi-branch image encoder with multi-to-multi contrastive learning on multiple VLM-generated captions per image and reports consistent gains over one-to-one and one-to-multi CLIP variants.
Discussion (0). Continue with ORCID to comment.