Pith. sign in

Paper Citation Record · LEDGER

CatV2TON: Taming Diffusion Transformers for Vision-Based Virtual Try-On with Temporal Concatenation

As of 10 August 2026, this Paper Citation Record lists 59 of 59 outbound references and 11 inbound Pith citation observations for arXiv:2501.11325.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2501.11325 v1

Coverage vector

measured 59 of 59 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-10T18:28:50.661750Z

measured 70 of 70 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-10T06:31:04.303077+00:00

measured 11 of 11 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-10T19:59:50.198365Z

measured 1 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-08-05T02:28:24.338817Z

Reference resolution

59 of 59 outbound references displayed

  • verified exact1
  • verified fuzzy23
  • unresolved35
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

1
arxiv_reference, observed 2026-08-05T02:28:24.338817Z

Outbound references

Observation 7a561bb7-defa-44f7-b2dc-e9a9832a40ee · outbound

This paper cites Sutherland, Michael Arbel, and Arthur Gretton.

CatV2TON: Taming Diffusion Transformers for Vision-Based Virtual Try-On with Temporal Concatenation Sutherland, Michael Arbel, and Arthur Gretton

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T18:28:51.401929Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-10T18:28:49.460751Z digest=sha256:8b52e3f408151478054a81e833fe032f0526dfa87cd05beae158fef350362375

Observation 345b28ac-bf28-4e90-bda9-df33f20ebccb · outbound

This paper cites an unresolved cited work.

CatV2TON: Taming Diffusion Transformers for Vision-Based Virtual Try-On with Temporal Concatenation Unresolved cited work

Reference 2

Resolution
unresolved
raw_fallback, observed 2026-08-10T18:28:51.387469Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-10T18:28:49.493259Z digest=sha256:6466551ad5960769d24a74a5fae434b47d4ff7ce53c69c41f517e5e4eeb465fb

Observation 80d55974-8f20-481c-a38b-c609b0c365a5 · outbound

This paper cites Quo vadis, action recognition? a new model and the kinetics dataset, 2018.

CatV2TON: Taming Diffusion Transformers for Vision-Based Virtual Try-On with Temporal Concatenation Quo vadis, action recognition? a new model and the kinetics dataset, 2018

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T18:28:51.374420Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-10T18:28:49.496985Z digest=sha256:d3eda2e28299f90c0cb475591b2bc52587a6cfed2fec793ffaf837e4ab93933e

Observation 83289c2f-0b53-42e7-b87e-108b0baf897f · outbound

This paper cites Stable- video: Text-driven consistency-aware diffusion video edit- ing.

CatV2TON: Taming Diffusion Transformers for Vision-Based Virtual Try-On with Temporal Concatenation Stable- video: Text-driven consistency-aware diffusion video edit- ing

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T18:28:51.361501Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-10T18:28:49.500942Z digest=sha256:d7fec0362a3795dfdf4e0edf94d1092069b8edfc54612603c33b67c3f4533db0

Observation 30c02353-98f4-4a12-a692-eb16fa68f509 · outbound

This paper cites Wear-Any-Way: Manipulable Virtual Try-on via Sparse Correspondence Alignment.

CatV2TON: Taming Diffusion Transformers for Vision-Based Virtual Try-On with Temporal Concatenation Wear-Any-Way: Manipulable Virtual Try-on via Sparse Correspondence Alignment

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-10T18:28:49.506731Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T18:28:49.506731Z digest=sha256:0b1e6ce1c96efc720df54a35c184c904dcf922885ccf2eb455349dc7cd427cae

Observation 8db16c5e-2ee9-4648-8220-5249702834db · outbound

This paper cites Viton-hd: High-resolution virtual try-on via misalignment-aware normalization.

CatV2TON: Taming Diffusion Transformers for Vision-Based Virtual Try-On with Temporal Concatenation Viton-hd: High-resolution virtual try-on via misalignment-aware normalization

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T18:28:51.347698Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-10T18:28:49.511978Z digest=sha256:42ac7245abe9511ff6921f803ea74edfc53a122b9cc2238ed563362b4e8fbd97

Observation 60115b94-96f8-4022-a6a4-df2c5ca6aa1a · outbound

This paper cites Improving Diffusion Models for Authentic Virtual Try-on in the Wild.

CatV2TON: Taming Diffusion Transformers for Vision-Based Virtual Try-On with Temporal Concatenation Improving Diffusion Models for Authentic Virtual Try-on in the Wild

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-10T18:28:49.659080Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T18:28:49.659080Z digest=sha256:e2ec81932715987cfe51af8b9971e62d468d79db2ed5bf31b0fc6dabc05e6d60

Observation b39549e5-aba9-4d06-900c-93b1c543c1a2 · outbound

This paper cites CatVTON: Concatenation Is All You Need for Virtual Try-On with Diffusion Models.

CatV2TON: Taming Diffusion Transformers for Vision-Based Virtual Try-On with Temporal Concatenation CatVTON: Concatenation Is All You Need for Virtual Try-On with Diffusion Models

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-10T18:28:49.699961Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T18:28:49.699961Z digest=sha256:a7893fac281ab951c71461008620ab4f3acd9ed51efa4d6b2fd68571d02890b4

Observation 2160f86b-42e1-414a-92a4-6689e6bd94cc · outbound

This paper cites Openmmlab pose estimation tool- box and benchmark.

CatV2TON: Taming Diffusion Transformers for Vision-Based Virtual Try-On with Temporal Concatenation Openmmlab pose estimation tool- box and benchmark

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T18:28:51.327230Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-10T18:28:49.705155Z digest=sha256:13937b34e5d280a3a460bacba62109209cc5ab95820f159a6daef379b933f602

Observation 097613e5-3b38-4c5d-ac12-4ba9056aa5de · outbound

This paper cites Fw-gan: Flow-navigated warping gan for video virtual try-on.

CatV2TON: Taming Diffusion Transformers for Vision-Based Virtual Try-On with Temporal Concatenation Fw-gan: Flow-navigated warping gan for video virtual try-on

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T18:28:51.314341Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-10T18:28:49.709275Z digest=sha256:fb93356734ea99379d5b774833e5d69001a3d0ba288215689dba56307734c568

Observation ea2d2b95-82c0-4f89-ae2d-fe90e7d187ac · outbound

This paper cites Fw-gan: Flow-navigated warping gan for video virtual try-on.

CatV2TON: Taming Diffusion Transformers for Vision-Based Virtual Try-On with Temporal Concatenation Fw-gan: Flow-navigated warping gan for video virtual try-on

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T18:28:51.299948Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-10T18:28:49.713762Z digest=sha256:cbf73b370e3c4bbde7e2a949b4f2ec23974846dc562bb2d720c69cf6a87ca4eb

Observation 2cef41e5-5bd8-4c8b-8584-acaab768cfc7 · outbound

This paper cites Structure and content-guided video synthesis with diffusion models.

CatV2TON: Taming Diffusion Transformers for Vision-Based Virtual Try-On with Temporal Concatenation Structure and content-guided video synthesis with diffusion models

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-10T18:28:49.718317Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T18:28:49.718317Z digest=sha256:52a9de71570eecaa8daea5546865652d3ac3c7d6d2e33f0755a6a7c0695f4ba0

Observation 44438eb5-a6c4-44be-8411-6123a5d24b5f · outbound

This paper cites Vivid: Video virtual try-on using diffusion models,.

CatV2TON: Taming Diffusion Transformers for Vision-Based Virtual Try-On with Temporal Concatenation Vivid: Video virtual try-on using diffusion models,

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-10T18:28:49.722984Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T18:28:49.722984Z digest=sha256:1bd1a579911dd0149ea2bd17697fa6d959b87a369ad5340f7af099b18edcbdd2

Observation b39bafa2-d45b-4bd8-bd74-623147e10097 · outbound

This paper cites Taming the power of diffusion models for high-quality virtual try-on with appearance flow.

CatV2TON: Taming Diffusion Transformers for Vision-Based Virtual Try-On with Temporal Concatenation Taming the power of diffusion models for high-quality virtual try-on with appearance flow

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T18:28:51.271263Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-10T18:28:49.733804Z digest=sha256:c8318dbc51f7c108201b64bf30dec2d27ac31fde501e56e6efa5e523490ac10c

Observation 53693d28-3dd1-4b4a-827a-02c853d67bcb · outbound

This paper cites Densepose: Dense human pose estimation in the wild, 2018.

CatV2TON: Taming Diffusion Transformers for Vision-Based Virtual Try-On with Temporal Concatenation Densepose: Dense human pose estimation in the wild, 2018

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T18:28:51.258679Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-10T18:28:49.821668Z digest=sha256:abc8f971ce0d2559ddf53750d03bcc368acc2b06c700d7e0f5cafaa210cbe233

Observation f7e9cb6c-f0b6-4c02-b8d8-eceac790db01 · outbound

This paper cites WildVidFit: Video Virtual Try-On in the Wild via Image-Based Controlled Diffusion Models.

CatV2TON: Taming Diffusion Transformers for Vision-Based Virtual Try-On with Temporal Concatenation WildVidFit: Video Virtual Try-On in the Wild via Image-Based Controlled Diffusion Models

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-10T18:28:49.900258Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T18:28:49.900258Z digest=sha256:87057571efa4915dc28956c2e50297d28c12257c33ab787112cb05e70f996a67

Observation bfa5eaf7-7c83-42fb-b3c2-4af709ca7905 · outbound

This paper cites Classifier-free diffusion guidance, 2022.

CatV2TON: Taming Diffusion Transformers for Vision-Based Virtual Try-On with Temporal Concatenation Classifier-free diffusion guidance, 2022

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T18:28:51.246948Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-10T18:28:49.905398Z digest=sha256:d647cf72752eba5646efaf9ce0bab19581a8018dda2c3fde9ee1a61a07fe250e

Observation 48317c8a-2fa0-442f-8486-cd1a3f205184 · outbound

This paper cites Classifier-free diffusion guidance, 2022.

CatV2TON: Taming Diffusion Transformers for Vision-Based Virtual Try-On with Temporal Concatenation Classifier-free diffusion guidance, 2022

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-10T18:28:49.909345Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T18:28:49.909345Z digest=sha256:8b0f80ba256a6364bea9cd843e5cd6e6585a1d7258a2467eeaa3a1ea4e69be58

Observation 17535412-defb-49cf-8917-a7bb2154c837 · outbound

This paper cites Denoising diffu- sion probabilistic models, 2020.

CatV2TON: Taming Diffusion Transformers for Vision-Based Virtual Try-On with Temporal Concatenation Denoising diffu- sion probabilistic models, 2020

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T18:28:51.227917Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-10T18:28:49.912930Z digest=sha256:8883424a735205fc7af0535594b58bf386b49ed6e097beb5c1da8acf7c516b3b

Observation 41813d4d-c6dd-48b3-9d5e-54d4b1dcfd20 · outbound

This paper cites Imagen Video: High Definition Video Generation with Diffusion Models.

CatV2TON: Taming Diffusion Transformers for Vision-Based Virtual Try-On with Temporal Concatenation Imagen Video: High Definition Video Generation with Diffusion Models

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-10T18:28:49.916983Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T18:28:49.916983Z digest=sha256:791e842ae0e7729a874f0a1f737c73ede4738353a218f64d867591d55fa9898e

Observation 392ee7f8-6321-4627-9dab-0fcdd1b20bd7 · outbound

This paper cites Animate anyone: Consistent and controllable image- to-video synthesis for character animation.

CatV2TON: Taming Diffusion Transformers for Vision-Based Virtual Try-On with Temporal Concatenation Animate anyone: Consistent and controllable image- to-video synthesis for character animation

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-10T18:28:49.923084Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T18:28:49.923084Z digest=sha256:25cf50c03b170619ecd60c570a73884f0902d35f6ea49e36d2f5fd8ed51c18cd

Observation 190f2057-8fdf-46a0-b722-ca981efbeba4 · outbound

This paper cites Arbitrary style transfer in real-time with adaptive instance normalization, 2017.

CatV2TON: Taming Diffusion Transformers for Vision-Based Virtual Try-On with Temporal Concatenation Arbitrary style transfer in real-time with adaptive instance normalization, 2017

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T18:28:51.207728Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-10T18:28:50.011587Z digest=sha256:0ad0794fddc7739dd3af30f49030e7fa9aa95433033636931caf1f9adcc09a33

Observation d9fcc7b9-80b9-42e4-9323-41abc53f5420 · outbound

This paper cites Cloth- former: Taming video virtual try-on in all module.

CatV2TON: Taming Diffusion Transformers for Vision-Based Virtual Try-On with Temporal Concatenation Cloth- former: Taming video virtual try-on in all module

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T18:28:51.196215Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-10T18:28:50.102385Z digest=sha256:5b53d9a35d30f59c4c30f6a8605c3831234301f771c51e8e406a02e1fef115cf

Observation 6fd2622e-8563-452a-be8c-ffa8b30879fc · outbound

This paper cites Stableviton: Learning semantic correspon- dence with latent diffusion model for virtual try-on, 2023.

CatV2TON: Taming Diffusion Transformers for Vision-Based Virtual Try-On with Temporal Concatenation Stableviton: Learning semantic correspon- dence with latent diffusion model for virtual try-on, 2023

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T18:28:51.185480Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-10T18:28:50.106625Z digest=sha256:7ce4b4eeacf543770b9cae983ab74777b3aa13df844882515d5fe55fbbe691bd

Observation 1509bf5c-3251-49e3-9de3-0d371a840499 · outbound

This paper cites WarpDiffusion: Efficient Diffusion Model for High-Fidelity Virtual Try-on.

CatV2TON: Taming Diffusion Transformers for Vision-Based Virtual Try-On with Temporal Concatenation WarpDiffusion: Efficient Diffusion Model for High-Fidelity Virtual Try-on

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-10T18:28:50.110511Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T18:28:50.110511Z digest=sha256:89145d261ba824c1a40735b0f575d299e45389e61ae902a983e3ba1bf1ddaff0

Observation 02534ac0-df6e-415a-b573-ba7073244cfb · outbound

This paper cites Hunyuan-dit: A powerful multi-resolution diffusion trans- former with fine-grained chinese understanding, 2024.

CatV2TON: Taming Diffusion Transformers for Vision-Based Virtual Try-On with Temporal Concatenation Hunyuan-dit: A powerful multi-resolution diffusion trans- former with fine-grained chinese understanding, 2024

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T18:28:51.174759Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-10T18:28:50.113930Z digest=sha256:73e0b3f5d2235a30fa158f5e620c48c293d2a7100fb5f17822b636be5e98ade2

Observation b8d832fb-b01f-4666-8de0-e2e709f0f8f5 · outbound

This paper cites Decoupled weight decay regularization, 2019.

CatV2TON: Taming Diffusion Transformers for Vision-Based Virtual Try-On with Temporal Concatenation Decoupled weight decay regularization, 2019

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-10T18:28:50.119155Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T18:28:50.119155Z digest=sha256:baab2a2f6bfbc538290d0918f7ecd5b2509ba55b1ce6ab780ce060fed1abc885

Observation 0badffb3-5b22-4eb0-94c4-12156bce61ac · outbound

This paper cites Follow your pose: Pose- guided text-to-video generation using pose-free videos.

CatV2TON: Taming Diffusion Transformers for Vision-Based Virtual Try-On with Temporal Concatenation Follow your pose: Pose- guided text-to-video generation using pose-free videos

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-10T18:28:50.124902Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T18:28:50.124902Z digest=sha256:1ad40c44a5c280476973d71f06ef29896132d3c509918537981ebb8f6e8be78c

Observation 9c84521f-005f-436e-ad1a-a92e06b510b1 · outbound

This paper cites Dreamix: Video Diffusion Models are General Video Editors.

CatV2TON: Taming Diffusion Transformers for Vision-Based Virtual Try-On with Temporal Concatenation Dreamix: Video Diffusion Models are General Video Editors

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-10T18:28:50.186762Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T18:28:50.186762Z digest=sha256:460a1f257910eeb19cde24f7ca8d590894ffb051061d14351bd1950dc8ffafa5

Observation 8062690f-72a3-418a-9d30-85840fa5e618 · outbound

This paper cites Dress code: High- resolution multi-category virtual try-on, 2022.

CatV2TON: Taming Diffusion Transformers for Vision-Based Virtual Try-On with Temporal Concatenation Dress code: High- resolution multi-category virtual try-on, 2022

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T18:28:51.149820Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-10T18:28:50.239005Z digest=sha256:322bbcecb53b4cdb8a3abdf012c1ee6ea376522b4ef31545758b4d403d36dace

Observation 34c4b6d4-f91b-4be5-a98a-7b1bc5974741 · outbound

This paper cites LaDI- VTON: Latent Diffusion Textual-Inversion Enhanced Virtual Try-On.

CatV2TON: Taming Diffusion Transformers for Vision-Based Virtual Try-On with Temporal Concatenation LaDI- VTON: Latent Diffusion Textual-Inversion Enhanced Virtual Try-On

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T18:28:51.138440Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-10T18:28:50.242926Z digest=sha256:87de627276df6c78ca7b86c3c7f1201658231fddf253e57a36329ab183895dd1

Observation 74281462-eed0-4155-8315-e704776d07c8 · outbound

This paper cites Scalable Diffusion Models with Transformers.

CatV2TON: Taming Diffusion Transformers for Vision-Based Virtual Try-On with Temporal Concatenation Scalable Diffusion Models with Transformers

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-10T18:28:50.247193Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T18:28:50.247193Z digest=sha256:078ded4b10bf2f795509990c7293bed9f0306eb528c54eef70666b06e4c2fa14

Observation 3147ecc9-ea73-4fd5-9635-7a61e78326f4 · outbound

This paper cites Controlnext: Powerful and effi- cient control for image and video generation, 2024.

CatV2TON: Taming Diffusion Transformers for Vision-Based Virtual Try-On with Temporal Concatenation Controlnext: Powerful and effi- cient control for image and video generation, 2024

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T18:28:51.128001Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-10T18:28:50.252128Z digest=sha256:8a61688b2624c2b91587d95193cca163c8094434f83e8d8c160349fdc2972358

Observation 7dbbdbae-13ab-4799-bf33-026646ae1c22 · outbound

This paper cites Fatezero: Fus- ing attentions for zero-shot text-based video editing.

CatV2TON: Taming Diffusion Transformers for Vision-Based Virtual Try-On with Temporal Concatenation Fatezero: Fus- ing attentions for zero-shot text-based video editing

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-10T18:28:50.255973Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T18:28:50.255973Z digest=sha256:456a997c5d4eead281a3fac5f8690cdb86203f972cee3d66bd42bdaf8e2ba46c

Observation b13176e9-ea3c-4045-b13c-c0e9928d10cf · outbound

This paper cites Learning transferable visual models from natural language supervision, 2021.

CatV2TON: Taming Diffusion Transformers for Vision-Based Virtual Try-On with Temporal Concatenation Learning transferable visual models from natural language supervision, 2021

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-10T18:28:50.259610Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T18:28:50.259610Z digest=sha256:0cf1f2a9564976e38a5048bfdc0a1efc3124cade5d45aa41b0921a41c67590a3

Observation b7383985-c694-454d-bdcb-84c9e4c5a7ed · outbound

This paper cites pytorch-fid: FID Score for PyTorch.

CatV2TON: Taming Diffusion Transformers for Vision-Based Virtual Try-On with Temporal Concatenation pytorch-fid: FID Score for PyTorch

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-10T18:28:50.263950Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T18:28:50.263950Z digest=sha256:59271b482b6d39f0055eb04ca13a12513e66f248df00e15e4c7cb7c42c3dfd55

Observation 173f627c-e658-4975-a7b2-242854d86e56 · outbound

This paper cites IMAGDressing-v1: Customizable Virtual Dressing.

CatV2TON: Taming Diffusion Transformers for Vision-Based Virtual Try-On with Temporal Concatenation IMAGDressing-v1: Customizable Virtual Dressing

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-10T18:28:50.353671Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T18:28:50.353671Z digest=sha256:214eb1e0a50d55a4d0003cfecd047205920f976c6241f0836e660ceaafaa2a7b

Observation 2f562451-ada4-4466-889f-5e139339e909 · outbound

This paper cites Make-A-Video: Text-to-Video Generation without Text-Video Data.

CatV2TON: Taming Diffusion Transformers for Vision-Based Virtual Try-On with Temporal Concatenation Make-A-Video: Text-to-Video Generation without Text-Video Data

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-10T18:28:50.433894Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T18:28:50.433894Z digest=sha256:a13b6735484e38f9128287572a6967f94da7efb92a990259810c0aa31a6f3b10

Observation 30bbabd4-c3a5-4d3e-b22b-fb5669e2184b · outbound

This paper cites Roformer: Enhanced transformer with rotary position embedding, 2023.

CatV2TON: Taming Diffusion Transformers for Vision-Based Virtual Try-On with Temporal Concatenation Roformer: Enhanced transformer with rotary position embedding, 2023

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-10T18:28:50.440142Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T18:28:50.440142Z digest=sha256:3e54eab3b67ee8fa2607beb0731c7b46c90a8a486afa905ede2ebb90f1b65548

Observation fbbdf039-56e6-4396-b48b-712576acd5a5 · outbound

This paper cites OutfitAnyone: Ultra-high Quality Virtual Try-On for Any Clothing and Any Person.

CatV2TON: Taming Diffusion Transformers for Vision-Based Virtual Try-On with Temporal Concatenation OutfitAnyone: Ultra-high Quality Virtual Try-On for Any Clothing and Any Person

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-10T18:28:50.444165Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T18:28:50.444165Z digest=sha256:274e4af8e3cbffae8244fec839ec4f64bc89bffbb9ea91d69c98538babd87679

Observation 0408ad54-acd8-43cc-b8c4-c44e983d1e0e · outbound

This paper cites an unresolved cited work.

CatV2TON: Taming Diffusion Transformers for Vision-Based Virtual Try-On with Temporal Concatenation Unresolved cited work

Reference 41

Resolution
unresolved
raw_fallback, observed 2026-08-10T18:28:51.092095Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-10T18:28:50.450455Z digest=sha256:e6c329192cd164b37ee76db9410ebedb446e5961dde8a65401b6a17f83ca6bb1

Observation 25349d1d-06a5-4c6a-a09d-8d7278cfddb5 · outbound

This paper cites Improving Virtual Try-On with Garment-focused Diffusion Models.

CatV2TON: Taming Diffusion Transformers for Vision-Based Virtual Try-On with Temporal Concatenation Improving Virtual Try-On with Garment-focused Diffusion Models

Reference 42

Resolution
verified exact
local_arxiv, observed 2026-08-10T18:28:50.893262Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-10T18:28:50.454775Z digest=sha256:c8217d00008817a743602b88775e20a07319981d3bf904f3c3c6b9cbf81dac80

Observation ec12c26f-d431-436e-a304-961b283330aa · outbound

This paper cites MV-VTON: Multi-View Virtual Try-On with Diffusion Models.

CatV2TON: Taming Diffusion Transformers for Vision-Based Virtual Try-On with Temporal Concatenation MV-VTON: Multi-View Virtual Try-On with Diffusion Models

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-10T18:28:50.459347Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T18:28:50.459347Z digest=sha256:14a8300c298a29f119db3fced5ca37e5a9dca7c0182feeff87aee81e775c904a

Observation ee5b1cc9-383b-420f-9104-f8216d3beda9 · outbound

This paper cites StableGarment: Garment-Centric Generation via Stable Diffusion.

CatV2TON: Taming Diffusion Transformers for Vision-Based Virtual Try-On with Temporal Concatenation StableGarment: Garment-Centric Generation via Stable Diffusion

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-10T18:28:50.548037Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T18:28:50.548037Z digest=sha256:f9fea5f935717e41328b38648e67a619a0dc8e0523be4b362fd10bb4aeb18271

Observation 27ae57d2-5c1b-4b77-a5af-d134a1d9c466 · outbound

This paper cites Image quality assessment: from error visibility to structural similarity.

CatV2TON: Taming Diffusion Transformers for Vision-Based Virtual Try-On with Temporal Concatenation Image quality assessment: from error visibility to structural similarity

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-10T18:28:50.605803Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T18:28:50.605803Z digest=sha256:d8232bd556083adb32c8cfe929ea28ee804c2b545f5bd1a5a3bb1ee3587991c8

Observation a1c09288-cd8e-44d1-ac2b-56bc7d25a428 · outbound

This paper cites Tune-a-video: One-shot tuning of image diffusion models for text-to-video generation.

CatV2TON: Taming Diffusion Transformers for Vision-Based Virtual Try-On with Temporal Concatenation Tune-a-video: One-shot tuning of image diffusion models for text-to-video generation

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-10T18:28:50.610056Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T18:28:50.610056Z digest=sha256:27aa3bd878056a5f7cc3afeb6a15fd2d354407741368d65e744971eda81ad379

Observation 6a09cde0-cc96-4e25-91a6-3f4b1052fddb · outbound

This paper cites Pasta-gan++: A versatile framework for high- resolution unpaired virtual try-on, 2022.

CatV2TON: Taming Diffusion Transformers for Vision-Based Virtual Try-On with Temporal Concatenation Pasta-gan++: A versatile framework for high- resolution unpaired virtual try-on, 2022

Reference 47

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T18:28:51.068169Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-10T18:28:50.614193Z digest=sha256:1406917402a21444bc65a5d05d64c4f8a4d5ca4e839565a9b6553fc5303232ae

Observation ee5559a9-7664-4f8b-bbce-db6ec935c765 · outbound

This paper cites Gp- vton: Towards general purpose virtual try-on via collabora- tive local-flow global-parsing learning.

CatV2TON: Taming Diffusion Transformers for Vision-Based Virtual Try-On with Temporal Concatenation Gp- vton: Towards general purpose virtual try-on via collabora- tive local-flow global-parsing learning

Reference 48

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T18:28:51.056818Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-10T18:28:50.618403Z digest=sha256:c5940487d762a91ce71366ed555786394d63b9e94c6faac761214e0c52c65458

Observation da042454-960d-4eef-9618-fc2b2ee33e72 · outbound

This paper cites Easyanimate: 10 A high-performance long video generation method based on transformer architecture.

CatV2TON: Taming Diffusion Transformers for Vision-Based Virtual Try-On with Temporal Concatenation Easyanimate: 10 A high-performance long video generation method based on transformer architecture

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-10T18:28:50.622147Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T18:28:50.622147Z digest=sha256:6718e982be4faf264f34a6aa6698ee5306e137055bd80e5f815147a1d30139d7

Observation 104817aa-6cde-4a2c-9d94-4114c05c8755 · outbound

This paper cites Easyanimate: A high-performance long video generation method based on transformer architecture, 2024.

CatV2TON: Taming Diffusion Transformers for Vision-Based Virtual Try-On with Temporal Concatenation Easyanimate: A high-performance long video generation method based on transformer architecture, 2024

Reference 50

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T18:28:51.045417Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-10T18:28:50.625851Z digest=sha256:f864a2d1a467801cb537c3442ec6e9d1f02ade515744ec437359321281151f20

Observation 43f66974-7fea-4509-afc5-ae5fe05215b9 · outbound

This paper cites OOTDiffusion: Outfitting Fusion based Latent Diffusion for Controllable Virtual Try-on.

CatV2TON: Taming Diffusion Transformers for Vision-Based Virtual Try-On with Temporal Concatenation OOTDiffusion: Outfitting Fusion based Latent Diffusion for Controllable Virtual Try-on

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-10T18:28:50.629353Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T18:28:50.629353Z digest=sha256:10357063e3f50001020f3318a06e5fc2fb666c0c89c57bea76f0894a58557cd1

Observation b57880bb-d5a5-4557-8143-3c26597510ec · outbound

This paper cites Tunnel Try-on: Excavating Spatial-temporal Tunnels for High-quality Virtual Try-on in Videos.

CatV2TON: Taming Diffusion Transformers for Vision-Based Virtual Try-On with Temporal Concatenation Tunnel Try-on: Excavating Spatial-temporal Tunnels for High-quality Virtual Try-on in Videos

Reference 52

Resolution
unresolved
no resolver link, observed 2026-08-10T18:28:50.633254Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T18:28:50.633254Z digest=sha256:28061078c55a4d8fc9bb8b14802b483c9f5cbf22b59653732af0f3d706fc0d1f

Observation 576c1648-c070-4d9e-8f88-dfbbde254bd0 · outbound

This paper cites Magicanimate: Temporally consistent human im- age animation using diffusion model.

CatV2TON: Taming Diffusion Transformers for Vision-Based Virtual Try-On with Temporal Concatenation Magicanimate: Temporally consistent human im- age animation using diffusion model

Reference 53

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T18:28:51.034027Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-10T18:28:50.637380Z digest=sha256:0e8cb912d93d5c07208437b7c4556c6a881a2704e60217c70e0472f1f2062289

Observation f34bcf81-5118-4881-8fe9-f2dd272eb96e · outbound

This paper cites MAGVIT: Masked generative video transformer.

CatV2TON: Taming Diffusion Transformers for Vision-Based Virtual Try-On with Temporal Concatenation MAGVIT: Masked generative video transformer

Reference 54

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T18:28:51.023259Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-10T18:28:50.641081Z digest=sha256:444b36606ba59ad395676899ae8dadf4c4f77bdf32b9d5e055cc3bb0932c8df4

Observation 1d0baee0-cc22-43e2-9b88-496e397cccd7 · outbound

This paper cites Video probabilistic diffusion models in projected latent space.

CatV2TON: Taming Diffusion Transformers for Vision-Based Virtual Try-On with Temporal Concatenation Video probabilistic diffusion models in projected latent space

Reference 55

Resolution
unresolved
no resolver link, observed 2026-08-10T18:28:50.645638Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T18:28:50.645638Z digest=sha256:131647317cc7a8bb953a02f09ff4f617cfe4b8f89aeccab83e6ac0131676f265

Observation 03f3a1a7-f538-487f-a1d0-f89fb5a654d5 · outbound

This paper cites The unreasonable effectiveness of deep features as a perceptual metric.

CatV2TON: Taming Diffusion Transformers for Vision-Based Virtual Try-On with Temporal Concatenation The unreasonable effectiveness of deep features as a perceptual metric

Reference 56

Resolution
unresolved
no resolver link, observed 2026-08-10T18:28:50.650005Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T18:28:50.650005Z digest=sha256:215f6f4ce95d2e1e460d4169efea97c5f0b3d1a21a699f71f2c826d339e68dc3

Observation 13829b45-8a96-42eb-8960-0c107af51299 · outbound

This paper cites MMTryon: Multi-Modal Multi-Reference Control for High-Quality Fashion Generation.

CatV2TON: Taming Diffusion Transformers for Vision-Based Virtual Try-On with Temporal Concatenation MMTryon: Multi-Modal Multi-Reference Control for High-Quality Fashion Generation

Reference 57

Resolution
unresolved
no resolver link, observed 2026-08-10T18:28:50.653764Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T18:28:50.653764Z digest=sha256:da40a4041500efbd7d82cddd3db81eb5af98160aeb8463ec6b9b2e42df490969

Observation 8693fb83-3cbd-4f8c-80ef-f8144476624b · outbound

This paper cites VITON-DiT: Learning In-the-Wild Video Try-On from Human Dance Videos via Diffusion Transformers.

CatV2TON: Taming Diffusion Transformers for Vision-Based Virtual Try-On with Temporal Concatenation VITON-DiT: Learning In-the-Wild Video Try-On from Human Dance Videos via Diffusion Transformers

Reference 58

Resolution
unresolved
no resolver link, observed 2026-08-10T18:28:50.657750Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T18:28:50.657750Z digest=sha256:7843a0e31cc7cae14d3fa617eb7f655fa411c61de99cc42366fec6b4153c93df

Observation 09a7d7ec-8e3c-482e-9ef7-3f50616eaacb · outbound

This paper cites MagicVideo: Efficient Video Generation With Latent Diffusion Models.

CatV2TON: Taming Diffusion Transformers for Vision-Based Virtual Try-On with Temporal Concatenation MagicVideo: Efficient Video Generation With Latent Diffusion Models

Reference 59

Resolution
unresolved
no resolver link, observed 2026-08-10T18:28:50.661750Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T18:28:50.661750Z digest=sha256:ab14466830b8b9a8ec2bd64a26e14d4249fc5be3a21c465fe358467a188c20d8

Pith citing papers

Observation 63f48028-f42b-43c5-b052-a43228267e4e · inbound

DualFit: A Two-Stage Virtual Try-On via Warping and Synthesis cites this paper.

DualFit: A Two-Stage Virtual Try-On via Warping and Synthesis CatV2TON: Taming Diffusion Transformers for Vision-Based Virtual Try-On with Temporal Concatenation

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-05T19:39:09.389257Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T19:39:09.389257Z digest=sha256:e1eb3b4c2e2b0df9af752cfb4e68c9afea506a42920af1db4869f6a39419f7db

Observation 613250bc-8767-4bed-b8e0-0bf4429eccc1 · inbound

RefTon: Reference person shot assist virtual Try-on cites this paper.

RefTon: Reference person shot assist virtual Try-on CatV2TON: Taming Diffusion Transformers for Vision-Based Virtual Try-On with Temporal Concatenation

Reference 15

Resolution
verified exact
arxiv_id, observed 2026-05-18T01:40:36.945734Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-18T01:39:44.746738Z digest=sha256:2e1083a63e8258ce2e32d56c06200dd93d29fcf9384706a271869a6de40d7494

Observation a2edb72d-f071-4f76-968e-086f38f70d1c · inbound

Eevee: Towards Close-up High-resolution Video-based Virtual Try-on cites this paper.

Eevee: Towards Close-up High-resolution Video-based Virtual Try-on CatV2TON: Taming Diffusion Transformers for Vision-Based Virtual Try-On with Temporal Concatenation

Reference 9

Resolution
verified exact
arxiv_id, observed 2026-05-17T06:39:10.514029Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-17T06:34:45.045267Z digest=sha256:24075613f6601cf48b4d7ea830f9ba09baab88466a9356567f5e63a6c0bb80d7

Observation d01d985a-813b-498b-a54f-d63559ce24eb · inbound

The devil is in the details: Enhancing Video Virtual Try-On via Keyframe-Driven Details Injection cites this paper.

The devil is in the details: Enhancing Video Virtual Try-On via Keyframe-Driven Details Injection CatV2TON: Taming Diffusion Transformers for Vision-Based Virtual Try-On with Temporal Concatenation

Reference 9

Resolution
verified exact
arxiv_id, observed 2026-05-16T20:08:22.952453Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-16T20:06:30.109866Z digest=sha256:2b3469ad48914886b0af8f2c625a270e6b8c67dd976c369a6b123cbf374ed090

Observation b33e6942-0c9e-429d-9529-834e0e36c7bf · inbound

TripVVT: A Large-Scale Triplet Dataset and a Coarse-Mask Baseline for In-the-Wild Video Virtual Try-On cites this paper.

TripVVT: A Large-Scale Triplet Dataset and a Coarse-Mask Baseline for In-the-Wild Video Virtual Try-On CatV2TON: Taming Diffusion Transformers for Vision-Based Virtual Try-On with Temporal Concatenation

Reference 10

Resolution
metadata mismatch
arxiv_id, observed 2026-05-09T05:05:12.363779Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-07T05:43:04.644875Z digest=sha256:76d0283b3eaeb9dfc950c9ad2a3fab31219cdb4d2da72d266f80fe6ef0c4dc0f

Observation a6a41355-54d0-40f6-8801-718da040d241 · inbound

OmniTryOn: Video Try-On Anything at Once! cites this paper.

OmniTryOn: Video Try-On Anything at Once! CatV2TON: Taming Diffusion Transformers for Vision-Based Virtual Try-On with Temporal Concatenation

Reference 8

Resolution
verified exact
arxiv_id, observed 2026-07-02T22:17:26.403243Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-06-27T19:00:44.243335Z digest=sha256:f2ac0147aad3d30de734fa12b476809b788a05b3b52b93fc8545cba475cf5162

Observation 06f0404a-c56f-40ad-81fc-9c210b9126de · inbound

TryOnCrafter: Unleashing Camera Trajectories for Realistic Video Virtual Try-on via a Renderable 4D Try-on Proxy cites this paper.

TryOnCrafter: Unleashing Camera Trajectories for Realistic Video Virtual Try-on via a Renderable 4D Try-on Proxy CatV2TON: Taming Diffusion Transformers for Vision-Based Virtual Try-On with Temporal Concatenation

Reference 5

Resolution
metadata mismatch
arxiv_id, observed 2026-07-04T21:00:08.934039Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-06-25T19:20:45.217627Z digest=sha256:13cbaa6696f8397aadaffffabd6a42a9644c1972afc7eba4a00b4afdf3d89739

Observation 475d2fb4-2dbf-4ef1-b6a7-68925d095935 · inbound

TryOnCrafter: Unleashing Camera Trajectories for Realistic Video Virtual Try-on via a Renderable 4D Try-on Proxy cites this paper.

TryOnCrafter: Unleashing Camera Trajectories for Realistic Video Virtual Try-on via a Renderable 4D Try-on Proxy CatV2TON: Taming Diffusion Transformers for Vision-Based Virtual Try-On with Temporal Concatenation

Reference 5

Resolution
unresolved
no resolver link, observed 2026-07-12T12:05:10.800662Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T12:05:10.800662Z digest=sha256:cf920f7fbaac5a3e7a1caede7102c32a129045faaf073344a56a510e22a001d4

Observation 5e1943a3-2872-4aa0-897d-08668255fe95 · inbound

UniMoCa: Unifying Motion and Camera Controls as Visual Proxies for Faithful Human Video Generation cites this paper.

UniMoCa: Unifying Motion and Camera Controls as Visual Proxies for Faithful Human Video Generation CatV2TON: Taming Diffusion Transformers for Vision-Based Virtual Try-On with Temporal Concatenation

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-04T17:55:19.469516Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T17:55:19.469516Z digest=sha256:69cc99093ca59569f3e0dc3b0aada94ad21ea5cd1b4f9c8d3ae8b6dc6d5e639d

Observation 5e46ed22-d538-4b1c-8a72-b8cbaec1f2b9 · inbound

UniVVT: A Unified End-to-End Framework for High-Fidelity Video Virtual Try-on cites this paper.

UniVVT: A Unified End-to-End Framework for High-Fidelity Video Virtual Try-on CatV2TON: Taming Diffusion Transformers for Vision-Based Virtual Try-On with Temporal Concatenation

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-08T00:18:10.195371Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T00:18:10.195371Z digest=sha256:fdaec67f6582ebc88977193285075be91d10984b2c602b5aead95d7689ecfc57

Observation b963571b-dc01-41fe-a345-5c92519dca3e · inbound

GOPI: Generation-Oriented 3D Pose Inference for Furniture Insertion from Single-View RGB-D Indoor Scenes cites this paper.

GOPI: Generation-Oriented 3D Pose Inference for Furniture Insertion from Single-View RGB-D Indoor Scenes CatV2TON: Taming Diffusion Transformers for Vision-Based Virtual Try-On with Temporal Concatenation

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-10T19:59:50.198365Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T19:59:50.198365Z digest=sha256:393d35a7222092f34e53a741ae6b33eafa79684f5723198f09f4f9c10f6172d7