Pith. sign in

REVIEW 4 cited by

CatV2TON: Taming Diffusion Transformers for Vision-Based Virtual Try-On with Temporal Concatenation

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2501.11325 v1 pith:OXFRQNP4 submitted 2025-01-20 cs.CV cs.AI

classification cs.CVcs.AI
keywords try-onvideocatv2tonimagevirtualacrosstaskstemporal
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Virtual try-on (VTON) technology has gained attention due to its potential to transform online retail by enabling realistic clothing visualization of images and videos. However, most existing methods struggle to achieve high-quality results across image and video try-on tasks, especially in long video scenarios. In this work, we introduce CatV2TON, a simple and effective vision-based virtual try-on (V2TON) method that supports both image and video try-on tasks with a single diffusion transformer model. By temporally concatenating garment and person inputs and training on a mix of image and video datasets, CatV2TON achieves robust try-on performance across static and dynamic settings. For efficient long-video generation, we propose an overlapping clip-based inference strategy that uses sequential frame guidance and Adaptive Clip Normalization (AdaCN) to maintain temporal consistency with reduced resource demands. We also present ViViD-S, a refined video try-on dataset, achieved by filtering back-facing frames and applying 3D mask smoothing for enhanced temporal consistency. Comprehensive experiments demonstrate that CatV2TON outperforms existing methods in both image and video try-on tasks, offering a versatile and reliable solution for realistic virtual try-ons across diverse scenarios.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 4 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. TryOnCrafter: Unleashing Camera Trajectories for Realistic Video Virtual Try-on via a Renderable 4D Try-on Proxy

    cs.CV 2026-06 unverdicted novelty 7.0 of 10

    TryOnCrafter is the first DiT-based framework for camera-controllable video virtual try-on via a renderable 4D try-on proxy distilled from 2D priors into 3DGS avatar animated with SMPL-X.

  2. UniVVT: A Unified End-to-End Framework for High-Fidelity Video Virtual Try-on

    cs.CV 2026-08 conditional novelty 6.0 of 10

    UniVVT reports state-of-the-art video and image virtual try-on by conditioning a diffusion video generator on task tokens from a multimodal language model, with no masks, poses, or warping at inference.

  3. UniMoCa: Unifying Motion and Camera Controls as Visual Proxies for Faithful Human Video Generation

    cs.CV 2026-08 conditional novelty 6.0 of 10

    A visual proxy that renders human motion under the driving camera and overlays camera trajectory markers lets a video diffusion model control both body motion and camera movement from a single visual conditioning space.

  4. DualFit: A Two-Stage Virtual Try-On via Warping and Synthesis

    cs.CV 2025-08 conditional novelty 4.0 of 10

    DualFit combines flow-based warping with a Res-UNet synthesis module guided by a preserved-region image and an inpainting mask, reporting state-of-the-art VITON-HD numbers for detail preservation and realism.

Pith tools