Pith. sign in

REVIEW 4 cited by

MMTryon: Multi-Modal Multi-Reference Control for High-Quality Fashion Generation

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2405.00448 v4 pith:AJJEAZ36 submitted 2024-05-01 cs.CV

classification cs.CV
keywords mmtryontry-onsegmentationexistinggarmentmethodsmulti-referencedependency
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

This paper introduces MMTryon, a multi-modal multi-reference VIrtual Try-ON (VITON) framework, which can generate high-quality compositional try-on results by taking a text instruction and multiple garment images as inputs. Our MMTryon addresses three problems overlooked in prior literature: 1) Support of multiple try-on items. Existing methods are commonly designed for single-item try-on tasks (e.g., upper/lower garments, dresses). 2)Specification of dressing style. Existing methods are unable to customize dressing styles based on instructions (e.g., zipped/unzipped, tuck-in/tuck-out, etc.) 3) Segmentation Dependency. They further heavily rely on category-specific segmentation models to identify the replacement regions, with segmentation errors directly leading to significant artifacts in the try-on results. To address the first two issues, our MMTryon introduces a novel multi-modality and multi-reference attention mechanism to combine the garment information from reference images and dressing-style information from text instructions. Besides, to remove the segmentation dependency, MMTryon uses a parsing-free garment encoder and leverages a novel scalable data generation pipeline to convert existing VITON datasets to a form that allows MMTryon to be trained without requiring any explicit segmentation. Extensive experiments on high-resolution benchmarks and in-the-wild test sets demonstrate MMTryon's superiority over existing SOTA methods both qualitatively and quantitatively. MMTryon's impressive performance on multi-item and style-controllable virtual try-on scenarios and its ability to try on any outfit in a large variety of scenarios from any source image, opens up a new avenue for future investigation in the fashion community.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 4 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Layering Virtual Try-On

    cs.CV 2026-07 conditional novelty 7.0 of 10

    A two-stage diffusion pipeline and new benchmark let virtual try-on add, remove, or swap clothing layers while preserving inner layers, with SOTA results on the new LVTON benchmark and on VITON-HD/DressCode.

  2. WearWow: Native 2K Multi-Garment Virtual Try-On via Adaptive Token Packing and Preference Alignment

    cs.CV 2026-07 conditional novelty 6.0 of 10

    WearWow generates native 2K multi-garment virtual try-on images without masks, using token packing plus dual preference rewards to preserve fabric texture.

  3. Pursuing Temporal-Consistent Video Virtual Try-On via Dynamic Pose Interaction

    cs.CV 2025-05 conditional novelty 6.0 of 10

    DPIDM, a diffusion model with pose-aware spatial and temporal attention plus a temporal attention loss, reports state-of-the-art video virtual try-on and cuts VFID on VVT from 1.280 to 0.506.

  4. TalkFashion: Intelligent Virtual Try-On Assistant Based on Multimodal Large Language Model

    cs.CV 2025-07 conditional novelty 4.0 of 10

    TalkFashion, a text-driven virtual try-on assistant, reports better semantic consistency and visual quality than four baselines on VITON-HD by combining an LLM router, catalog matching, and automatic mask generation.

Pith tools