REVIEW 4 major objections 4 minor 96 references
Delay-constrained re-entry governs large-scale brain seizures and other network pathologies
T0 review · 4 major / 4 minor · reviewed 2026-08-05 · deepseek-v4-flash
Pith's one-line read A patient-specific millimetre-scale virtual brain built from diffusion MRI and excitable neural fields shows that realistic cortico-cortical conduction delays are sufficient to generate self-sustaining seizure loops, and that a narrow delay
desk verdict Abstract promises a testable seizure mechanism, but the body is an unrelated virtual try-on paper; unverifiable as submitted. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The carrying object is the delay-constrained re-entrant loop: a wave of excitation that travels from one cortical region to others and returns after a conduction delay long enough to re-excite its origin, so the loop sustains itself. It is embedded in a millimetre-scale virtual brain—a patient-specific reconstruction from diffusion MRI tractography—with excitable neural fields (the Epileptor model) placed on the reconstructed cortical surface and white-matter connectome. The narrow delay-coupling window is the parameter region in conduction delay and coupling strength where the loop neither dies out nor saturates; the paper uses it as the mechanistic link between structural wiring delays and
What would settle it
Measure the same patient's cortico-cortical conduction delays independently (for example, with cortico-cortical evoked potentials from subdural electrodes) and check whether they fall inside the model's narrow delay-coupling window. If the independently measured delays lie outside the window while the 184 recorded seizures still follow the predicted frequency–duration relationship, the delay-window claim is falsified. A second check: deliver biphasic stimulation at the phase the model says aborts re-entry and at the opposite phase; the claim predicts termination only at the model's phase.
Extended reading notes
Core claim
The central claim is that re-entry—a travelling excitation loop that feeds back on itself—arises in the human brain purely from realistic cortico-cortical conduction delays, and that this delay-constrained re-entry is the dynamical mechanism behind large-scale seizure synchrony. In a patient-specific, millimetre-scale virtual brain built from diffusion MRI tractography and embedded with excitable Epileptor neural fields, the authors report that self-sustaining re-entry appears within a narrow delay-coupling window, and that the window's parameter values predict the oscillation frequency and seizure duration of 184 recorded human seizures. They further report that biphasic stimuli timed to a
Load-bearing premise
The load-bearing premise is that the virtual brain's estimates of cortico-cortical conduction delays, derived from diffusion MRI tractography and a velocity mapping, put the patient inside the narrow delay-coupling window that the model requires; if the true delays lie outside it, the claimed mechanism would not govern this patient's seizures.
Editorial extensions
If this is right
- Measurable conduction delays along cortico-cortical pathways become a patient-specific predictor of seizure frequency and duration, since the delay-coupling window is what sets both.
- The virtual brain can serve as an in-silico screening platform: stimulation timing and electrode placement for real epilepsy surgery can be tested against re-entry termination before any intervention.
- Sub-millimetre lesions that cut the re-entrant loop in the model indicate which small white-matter segment would be the surgical target for minimally invasive disconnection.
- Because termination depends on the phase of the seizure cycle, neuromodulation that tracks the patient's intracranial phase could deliver biphasic pulses precisely when they can abort the loop.
- If delay-constrained re-entry is a generic mechanism for large-scale synchrony, the same delay-coupling logic may apply to other paroxysmal brain events, though the paper's direct evidence is in epilepsy.
Reading between the lines
- The narrowness of the delay-coupling window implies a strong testable prediction: patients whose independently measured conduction delays fall outside the window should not show the claimed frequency–duration relationship; measuring delays with cortico-cortical evoked potentials would test this.
- The mechanism generalises naturally to other travelling-wave pathologies such as migraine aura or spreading depolarization, which the paper does not simulate; the same delay-coupling parameter sweep could be run with those wave speeds.
- The 184-seizure match is a quantitative claim; the received text provides only the abstract, so independent replication requires the full protocol for seizure annotation and delay estimation.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. This arXiv record presents an abstract for a computational-neuroscience paper asserting that delay-constrained re-entry in a millimetre-scale virtual brain generates self-sustaining seizure loops, that a narrow delay-coupling window predicts oscillation frequency and seizure duration across 184 recorded seizures, and that precisely timed biphasic stimuli or sub-millimetre virtual lesions can abort re-entry with phase-dependent termination rules validated in intracranial recordings. The full text received, however, is entirely different: 'Voost: A Unified and Scalable Diffusion Transformer for Bidirectional Virtual Try-On and Try-Off' (arXiv:2508.04825v2), a computer-graphics paper on garment transfer. The body contains no Epileptor neural fields, no diffusion MRI tractography, no cortico-cortical delay estimation, no parameter sweeps, no 184-seizure dataset, no stimulation or lesion simulations, and no intracranial validation. Consequently, none of the abstract's central claims is supported by or verifiable from the manuscript as submitted.
Significance. If the abstract's claims were substantiated, the work would be significant: it would provide a patient-specific, mechanism-based account of large-scale seizure dynamics and a testbed for precision neuromodulation and minimally invasive disconnection, with quantitative predictions against a large clinical seizure corpus. The received manuscript, however, contains none of the evidence needed to assess those claims. There are no equations, no model specification, no parameter ranges, no prediction-versus-fit analysis, no error bars, and no clinical annotation protocol; all quantitative content concerns virtual-try-on image quality. The statement that a 'narrow delay-coupling window predicts' 184 seizures cannot be distinguished from post hoc fitting because the window is undefined and the data are absent. I therefore cannot credit any of the claimed findings, and the significance is purely conditional on a body of evidence that is not present here.
major comments (4)
- [Abstract vs. full text] The central claim is unsupported by the received manuscript. The abstract describes an Epileptor-based virtual-brain study, but the full text is a diffusion-transformer paper for virtual try-on/try-off. There is no model equation, no simulation protocol, no tractography pipeline, and no clinical dataset anywhere in the body. Thus every load-bearing assertion in the abstract — 'realistic cortico-cortical delays are sufficient to generate self-sustaining re-entry,' 'narrow delay-coupling window predicts oscillation frequency and seizure duration across 184 recorded seizures,' and 'phase-dependent termination rules validated in intracranial recordings' — is unverifiable from this submission.
- [Full text, §4.5 'Limitations and future work'] The only explicit limitation statement in the manuscript concerns 'precise control over the garment's fit' and future incorporation of body measurements or garment metadata. It makes no reference to seizure dynamics, tractography error, parameter identifiability, or annotation reliability. This is a self-contained signal that the body is a different paper, and it also constitutes a missing-support issue: the abstract's mechanism is presented without any stated limitations, sensitivity analysis, or boundary conditions.
- [Abstract, 'narrow delay-coupling window'] The predictive claim cannot be assessed because the window bounds are never defined and no sweep protocol is provided. In particular, the phrase 'predicts oscillation frequency and seizure duration across 184 recorded seizures' is ambiguous: if the window was selected after inspecting those seizures, the result is a fit, not a prediction. The manuscript must report the pre-specified or cross-validated window, the parameter range searched, the number of subjects, the annotation procedure, and quantitative prediction errors. None of these appear.
- [Full text, absence of methods] There is no description of the diffusion MRI acquisition, tractography, tract-length or velocity mapping, Epileptor coupling terms, numerical integrator, or patient-specific reconstruction. Even under the most charitable reading of the abstract, the measurement chain from MRI to 'realistic cortico-cortical delays' is a fragile premise that cannot be checked. A reader cannot determine whether the model's delays are realistic, whether the 'sub-millimetre virtual lesions' are well defined, or whether the 'biphasic stimuli' are within physiological limits.
minor comments (4)
- [Title and author list] The title and authors on the arXiv record (q-bio.NC, delay-constrained re-entry) do not match the title and authors of the full text (Voost; Lee, Kwak, et al.). The submission cannot be processed as a coherent manuscript until this mismatch is resolved.
- [References] All references in the full text (e.g., [13], [14], [15]) cite computer-vision and diffusion-model papers. None are related to Epileptor models, seizure dynamics, brain network modelling, or intracranial electrophysiology, so the abstract's scientific claims have no literature grounding in the received text.
- [Figures and tables] The figures display clothing images, qualitative try-on/try-off comparisons, user-study screenshots, and failure cases. There are no figures showing virtual brain reconstructions, re-entry loops, parameter sweeps, seizure time series, or intracranial recordings, contrary to what the abstract implies.
- [Data and code availability] No data availability or code availability statement is present. For a paper claiming quantitative validation on 184 clinical seizures, this omission is serious, though it is secondary to the full-text mismatch.
Circularity Check
No circularity demonstrable: the received full text is an unrelated computer-graphics paper, so the claimed derivation chain is absent rather than circular.
full rationale
The supplied full text is the Voost virtual-try-on paper (arXiv:2508.04825v2), not the claimed q-bio.NC manuscript 'Delay-constrained re-entry governs large-scale brain seizures and other network pathologies.' There are no model equations, no Epileptor neural fields, no tractography-derived delay estimates, no delay-coupling window definition, no 184-seizure dataset, and no stimulation or lesion simulations in the received text. Under the hard rule that circularity must be exhibited by quoting the paper's own equations or self-citations, no circular step can be identified. The abstract's sentence 'Systematic parameter sweeps reveal a narrow delay-coupling window that predicts oscillation frequency and seizure duration across 184 recorded seizures' could describe a post-hoc fit, but the received text does not state that the window was selected from those seizures, and no methods are available to verify either reading. Unverifiability is a correctness/evidence concern, not circularity. Therefore the honest finding is no circularity (score 0), with the explicit caveat that the claimed derivation chain is entirely absent from the supplied manuscript.
Assumptions & free parameters
free parameters (3)
- delay-coupling window bounds =
not stated in abstract
- biphasic stimulus timing and amplitude parameters =
not stated in abstract
- Epileptor neural-field parameters (e.g., regional excitability) =
not stated in abstract
assumptions (3)
- domain assumption Diffusion MRI tractography yields cortico-cortical delay estimates that are realistic enough to generate and constrain re-entry
- domain assumption The Epileptor neural field, embedded on the virtual brain, captures the seizure-relevant excitable dynamics
- domain assumption The 184 recorded seizures are correctly identified and annotated for frequency and duration
Cite this review
Pith. "Pith review of Delay-constrained re-entry governs large-scale brain seizures and other network pathologies." pith.science (2026). https://pith.science/paper/6UBVWQPX
@misc{pith2026250804824,
author = {Pith},
title = {Pith review of: Delay-constrained re-entry governs large-scale brain seizures and other network pathologies},
year = {2026},
howpublished = {\url{https://pith.science/paper/6UBVWQPX}},
note = {Machine review of arXiv:2508.04824}
}
read the original abstract
Re-entry of travelling excitation loops is a long-suspected driver of human seizures, yet how such loops arise in patient brain networks -- and how susceptible they are to targeted disruption -- remains unclear. We reconstruct a millimetre-scale virtual brain from diffusion MRI of a drug-resistant epilepsy patient, embed excitable Epileptor neural fields, and show that realistic cortico-cortical delays are sufficient to generate self-sustaining re-entry. Systematic parameter sweeps reveal a narrow delay-coupling window that predicts oscillation frequency and seizure duration across 184 recorded seizures. Precisely timed biphasic stimuli or sub-millimetre virtual lesions abort re-entry in silico, yielding phase-dependent termination rules validated in intracranial recordings. Our framework exposes delay-constrained re-entry as a generic dynamical mechanism for large-scale brain synchrony and provides a patient-specific testbed for precision neuromodulation and minimally invasive disconnection.
Reference graph
Works this paper leans on
-
[1]
Single stage virtual try-on via deformable attention flows
Shuai Bai, Huiling Zhou, Zhikang Li, Chang Zhou, and Hongxia Yang. Single stage virtual try-on via deformable attention flows. InEuropean Conference on Computer Vi- sion (ECCV), 2022. 2
2022
-
[2]
Multimodal garment designer: Human-centric latent diffusion models for fashion image editing
Alberto Baldrati, Davide Morelli, Giuseppe Cartella, Mar- cella Cornia, Marco Bertini, and Rita Cucchiara. Multimodal garment designer: Human-centric latent diffusion models for fashion image editing. InConference on Computer Vision and Pattern Recognition (CVPR), 2023. 2
2023
-
[3]
Demystifying mmd gans.arXiv preprint arXiv:1801.01401, 2018
Mikołaj Bi ´nkowski, Danica J Sutherland, Michael Arbel, and Arthur Gretton. Demystifying mmd gans.arXiv preprint arXiv:1801.01401, 2018. 6
arXiv 2018
-
[4]
Flux.https://github.com/ black-forest-labs/flux, 2024
Black Forest Labs. Flux.https://github.com/ black-forest-labs/flux, 2024. 2, 4
2024
-
[5]
Andreas Blattmann, Tim Dockhorn, Sumith Kulal, Daniel Mendelevitch, Maciej Kilian, Dominik Lorenz, Yam Levi, Zion English, Vikram V oleti, Adam Letts, et al. Stable video diffusion: Scaling latent video diffusion models to large datasets.arXiv preprint arXiv:2311.15127, 2023. 2
arXiv 2023
-
[6]
MasaCtrl: Tuning-free mutual self-attention control for consistent image synthesis and editing
Mingdeng Cao, Xintao Wang, Zhongang Qi, Ying Shan, Xiaohu Qie, and Yinqiang Zheng. MasaCtrl: Tuning-free mutual self-attention control for consistent image synthesis and editing. InConference on Computer Vision and Pattern Recognition (CVPR), 2023. 2
2023
-
[7]
GS- VTON: Controllable 3d virtual try-on with gaussian splat- ting.arXiv, 2024
Yukang Cao, Masoud Hadi, Liang Pan, and Ziwei Liu. GS- VTON: Controllable 3d virtual try-on with gaussian splat- ting.arXiv, 2024. 11
2024
-
[8]
Realtime multi-person 2d pose estimation using part affin- ity fields
Zhe Cao, Tomas Simon, Shih-En Wei, and Yaser Sheikh. Realtime multi-person 2d pose estimation using part affin- ity fields. InConference on Computer Vision and Pattern Recognition (CVPR), 2017. 3
2017
Show all 96 references
-
[9]
Emerg- ing properties in self-supervised vision transformers
Mathilde Caron, Hugo Touvron, Ishan Misra, Herv ´e J´egou, Julien Mairal, Piotr Bojanowski, and Armand Joulin. Emerg- ing properties in self-supervised vision transformers. InIn- ternational Conference on Computer Vision (ICCV), 2021. 2
2021
-
[10]
Wear-any-way: Manip- ulable virtual try-on via sparse correspondence alignment
Mengting Chen, Xi Chen, Zhonghua Zhai, Chen Ju, Xuewen Hong, Jinsong Lan, and Shuai Xiao. Wear-any-way: Manip- ulable virtual try-on via sparse correspondence alignment. In European Conference on Computer Vision (ECCV), 2024. 2
2024
-
[11]
Zero-shot image editing with reference imitation
Xi Chen, Yutong Feng, Mengting Chen, Yiyang Wang, Shi- long Zhang, Yu Liu, Yujun Shen, and Hengshuang Zhao. Zero-shot image editing with reference imitation. InAd- vances in Neural Information Processing Systems (NeurIPS),
-
[12]
Anydoor: Zero-shot object-level im- age customization
Xi Chen, Lianghua Huang, Yu Liu, Yujun Shen, Deli Zhao, and Hengshuang Zhao. Anydoor: Zero-shot object-level im- age customization. InConference on Computer Vision and Pattern Recognition (CVPR), 2024. 2
2024
-
[13]
VITON-HD: High-resolution virtual try-on via misalignment-aware normalization
Seunghwan Choi, Sunghyun Park, Minsoo Lee, and Jaegul Choo. VITON-HD: High-resolution virtual try-on via misalignment-aware normalization. InConference on Com- puter Vision and Pattern Recognition (CVPR), 2021. 2, 5, 8, 10, 15, 17
2021
-
[14]
Improving diffusion models for au- thentic virtual try-on in the wild
Yisol Choi, Sangkyung Kwak, Kyungmin Lee, Hyungwon Choi, and Jinwoo Shin. Improving diffusion models for au- thentic virtual try-on in the wild. InEuropean Conference on Computer Vision (ECCV), 2024. 2, 6, 8
2024
-
[15]
CatVTON: Concatenation is all you need for virtual try-on with diffusion models
Zheng Chong, Xiao Dong, Haoxiang Li, Shiyue Zhang, Wenqing Zhang, Xujie Zhang, Hanqing Zhao, and Xiaodan Liang. CatVTON: Concatenation is all you need for virtual try-on with diffusion models. InInternational Conference on Learning Representations (ICLR), 2025. 2, 3, 6, 8
2025
-
[16]
Patch n’pack: Navit, a vision transformer for any aspect ratio and resolution
Mostafa Dehghani, Basil Mustafa, Josip Djolonga, Jonathan Heek, Matthias Minderer, Mathilde Caron, Andreas Steiner, Joan Puigcerver, Robert Geirhos, Ibrahim M Alabdul- mohsin, et al. Patch n’pack: Navit, a vision transformer for any aspect ratio and resolution. InAdvances in N...
2023
-
[17]
An image is worth 16x16 words: Transformers for image recognition at scale
Alexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn, Xiaohua Zhai, Thomas Unterthiner, Mostafa Dehghani, Matthias Minderer, Georg Heigold, Syl- vain Gelly, Jakob Uszkoreit, and Neil Houlsby. An image is worth 16x16 words: Transformers for image recognition ...
2021
-
[18]
Scaling rectified flow trans- formers for high-resolution image synthesis
Patrick Esser, Sumith Kulal, Andreas Blattmann, Rahim Entezari, Jonas M ¨uller, Harry Saini, Yam Levi, Dominik Lorenz, Axel Sauer, Frederic Boesel, Dustin Podell, Tim Dockhorn, Zion English, Kyle Lacey, Alex Goodwin, Yan- nik Marek, and Robin Rombach. Scaling rectified flow tr...
2024
-
[19]
An image is worth one word: Personalizing text-to-image gen- eration using textual inversion
Rinon Gal, Yuval Alaluf, Yuval Atzmon, Or Patashnik, Amit Haim Bermano, Gal Chechik, and Daniel Cohen-or. An image is worth one word: Personalizing text-to-image gen- eration using textual inversion. InInternational Conference on Learning Representations (ICLR), 2023. 2
2023
-
[20]
Parser-free virtual try-on via distilling appearance flows
Yuying Ge, Yibing Song, Ruimao Zhang, Chongjian Ge, Wei Liu, and Ping Luo. Parser-free virtual try-on via distilling appearance flows. InConference on Computer Vision and Pattern Recognition (CVPR), 2021. 2
2021
-
[21]
Parser-free virtual try-on via distilling ap- pearance flows
Yuying Ge, Yibing Song, Ruimao Zhang, Chongjian Ge, Wei Liu, and Ping Luo. Parser-free virtual try-on via distilling ap- pearance flows. InEuropean Conference on Computer Vision (ECCV), 2021. 2
2021
-
[22]
Generative adversarial networks
Ian Goodfellow, Jean Pouget-Abadie, Mehdi Mirza, Bing Xu, David Warde-Farley, Sherjil Ozair, Aaron Courville, and 11 Yoshua Bengio. Generative adversarial networks. InAd- vances in Neural Information Processing Systems (NeurIPS),
-
[23]
Gemini 2.0 flash - multimodal ai model
Google DeepMind. Gemini 2.0 flash - multimodal ai model. https://gemini.google.com/, 2025. 6, 9
2025
-
[24]
Taming the power of diffu- sion models for high-quality virtual try-on with appearance flow
Junhong Gou, Siyu Sun, Jianfu Zhang, Jianlou Si, Chen Qian, and Liqing Zhang. Taming the power of diffu- sion models for high-quality virtual try-on with appearance flow. InACM International Conference on Multimedia (ACMMM), 2023. 2
2023
-
[25]
Densepose: Dense human pose estimation in the wild
Rıza Alp G ¨uler, Natalia Neverova, and Iasonas Kokkinos. Densepose: Dense human pose estimation in the wild. In Conference on Computer Vision and Pattern Recognition (CVPR), 2018. 3, 15
2018
-
[26]
Animatediff: Animate your personalized text- to-image diffusion models without specific tuning.arXiv preprint arXiv:2307.04725, 2023
Yuwei Guo, Ceyuan Yang, Anyi Rao, Zhengyang Liang, Yaohui Wang, Yu Qiao, Maneesh Agrawala, Dahua Lin, and Bo Dai. Animatediff: Animate your personalized text- to-image diffusion models without specific tuning.arXiv preprint arXiv:2307.04725, 2023. 2
2023 arXiv
-
[27]
Xintong Han, Zuxuan Wu, Zhe Wu, Ruichi Yu, and Larry S. Davis. VITON: An image-based virtual try-on network. InConference on Computer Vision and Pattern Recognition (CVPR), 2018. 2
2018
-
[28]
Clothflow: A flow-based model for clothed person generation
Xintong Han, Xiaojun Hu, Weilin Huang, and Matthew R Scott. Clothflow: A flow-based model for clothed person generation. InConference on Computer Vision and Pattern Recognition (CVPR), 2019. 2
2019
-
[29]
Wildvidfit: Video virtual try-on in the wild via image-based controlled diffusion mod- els
Zijian He, Peixin Chen, Guangrun Wang, Guanbin Li, Philip HS Torr, and Liang Lin. Wildvidfit: Video virtual try-on in the wild via image-based controlled diffusion mod- els. InEuropean Conference on Computer Vision (ECCV),
-
[30]
VTON 360: High-fidelity virtual try-on from any viewing direction
Zijian He, Yuwei Ning, Yipeng Qin, Guangrun Wang, Sibei Yang, Liang Lin, and Guanbin Li. VTON 360: High-fidelity virtual try-on from any viewing direction. InConference on Computer Vision and Pattern Recognition (CVPR), 2025. 11
2025
-
[31]
GANs trained by a two time-scale update rule converge to a local nash equilib- rium
Martin Heusel, Hubert Ramsauer, Thomas Unterthiner, Bernhard Nessler, and Sepp Hochreiter. GANs trained by a two time-scale update rule converge to a local nash equilib- rium. InAdvances in Neural Information Processing Systems (NeurIPS), 2017. 6
2017
-
[32]
Denoising dif- fusion probabilistic models
Jonathan Ho, Ajay Jain, and Pieter Abbeel. Denoising dif- fusion probabilistic models. InAdvances in Neural Informa- tion Processing Systems (NeurIPS), 2020. 3
2020
-
[33]
LoRA: Low-rank adaptation of large language models
Edward J Hu, yelong shen, Phillip Wallis, Zeyuan Allen- Zhu, Yuanzhi Li, Shean Wang, Lu Wang, and Weizhu Chen. LoRA: Low-rank adaptation of large language models. InIn- ternational Conference on Learning Representations (ICLR),
-
[34]
Animate anyone: Consistent and controllable image- to-video synthesis for character animation
Li Hu. Animate anyone: Consistent and controllable image- to-video synthesis for character animation. InConference on Computer Vision and Pattern Recognition (CVPR), 2024. 2
2024
-
[35]
In-context lora for diffusion transformers.arXiv,
Lianghua Huang, Wei Wang, Zhi-Fan Wu, Yupeng Shi, Huanzhang Dou, Chen Liang, Yutong Feng, Yu Liu, and Jin- gren Zhou. In-context lora for diffusion transformers.arXiv,
-
[36]
Do not mask what you do not need to mask: a parser-free virtual try-on
Thibaut Issenhuth, J ´er´emie Mary, and Cl ´ement Calauzenes. Do not mask what you do not need to mask: a parser-free virtual try-on. InEuropean Conference on Computer Vision (ECCV), 2020. 2
2020
-
[37]
Fashion-VDM: Video diffusion model for vir- tual try-on
Johanna Karras, Yingwei Li, Nan Liu, Luyang Zhu, Innfarn Yoo, Andreas Lugmayr, Chris Lee, and Ira Kemelmacher- Shlizerman. Fashion-VDM: Video diffusion model for vir- tual try-on. InIn ACM SIGGRAPH Asia, 2024. 11
2024
-
[38]
A style-based generator architecture for generative adversarial networks
Tero Karras, Samuli Laine, and Timo Aila. A style-based generator architecture for generative adversarial networks. InConference on Computer Vision and Pattern Recognition (CVPR), 2019. 2
2019
-
[39]
Analyzing and improving the image quality of stylegan
Tero Karras, Samuli Laine, Miika Aittala, Janne Hellsten, Jaakko Lehtinen, and Timo Aila. Analyzing and improving the image quality of stylegan. InConference on Computer Vision and Pattern Recognition (CVPR), 2020. 2
2020
-
[40]
Sapiens: Foundation for human vision mod- els.arXiv preprint arXiv:2408.12569, 2024
Rawal Khirodkar, Timur Bagautdinov, Julieta Martinez, Su Zhaoen, Austin James, Peter Selednik, Stuart Anderson, and Shunsuke Saito. Sapiens: Foundation for human vision mod- els.arXiv preprint arXiv:2408.12569, 2024. 15
2024 arXiv
-
[41]
StableVITON: Learning semantic cor- respondence with latent diffusion model for virtual try-on
Jeongho Kim, Guojung Gu, Minho Park, Sunghyun Park, and Jaegul Choo. StableVITON: Learning semantic cor- respondence with latent diffusion model for virtual try-on. InConference on Computer Vision and Pattern Recognition (CVPR), 2024. 2, 6, 8
2024
-
[42]
Berg, Wan-Yen Lo, Piotr Doll ´ar, and Ross Girshick
Alexander Kirillov, Eric Mintun, Nikhila Ravi, Hanzi Mao, Chloe Rolland, Laura Gustafson, Tete Xiao, Spencer White- head, Alexander C. Berg, Wan-Yen Lo, Piotr Doll ´ar, and Ross Girshick. Segment anything.arXiv:2304.02643, 2023. 15
2023 arXiv
-
[43]
Kolors: Effective training of diffusion model for photorealistic text-to-image synthesis.arXiv preprint,
Kolors-Team. Kolors: Effective training of diffusion model for photorealistic text-to-image synthesis.arXiv preprint,
-
[44]
Multi-concept customization of text-to-image diffusion
Nupur Kumari, Bingliang Zhang, Richard Zhang, Eli Shechtman, and Jun-Yan Zhu. Multi-concept customization of text-to-image diffusion. InProceedings of the IEEE/CVF conference on computer vision and pattern recognition, pages 1931–1941, 2023. 16
1931
-
[45]
Vivid-1-to-3: Novel view syn- thesis with video diffusion models
Jeong-gi Kwak, Erqun Dong, Yuhe Jin, Hanseok Ko, Shweta Mahajan, and Kwang Moo Yi. Vivid-1-to-3: Novel view syn- thesis with video diffusion models. InConference on Com- puter Vision and Pattern Recognition (CVPR), 2024. 2
2024
-
[46]
Self- correction for human parsing.IEEE Transactions on Pattern Analysis and Machine Intelligence, 44(6):3260–3271, 2020
Peike Li, Yunqiu Xu, Yunchao Wei, and Yi Yang. Self- correction for human parsing.IEEE Transactions on Pattern Analysis and Machine Intelligence, 44(6):3260–3271, 2020. 15
2020
-
[47]
Few-shot image generation with elastic weight consolida- tion.arXiv preprint arXiv:2012.02780, 2020
Yijun Li, Richard Zhang, Jingwan Lu, and Eli Shechtman. Few-shot image generation with elastic weight consolida- tion.arXiv preprint arXiv:2012.02780, 2020. 16
2012 arXiv
-
[48]
Yaron Lipman, Ricky T. Q. Chen, Heli Ben-Hamu, Maxi- milian Nickel, and Matt Le. Flow matching for generative modeling.arXiv, 2023. 3, 4
2023
-
[49]
Zero-1-to-3: Zero-shot one image to 3d object
Ruoshi Liu, Rundi Wu, Basile Van Hoorick, Pavel Tok- makov, Sergey Zakharov, and Carl V ondrick. Zero-1-to-3: Zero-shot one image to 3d object. InConference on Com- puter Vision and Pattern Recognition (CVPR), 2023. 2
2023
-
[50]
Flow straight and fast: Learning to generate and transfer data with rectified flow.arXiv, 2022
Xingchao Liu, Chengyue Gong, and Qiang Liu. Flow straight and fast: Learning to generate and transfer data with rectified flow.arXiv, 2022. 4 12
2022
-
[51]
Decoupled weight de- cay regularization
Ilya Loshchilov and Frank Hutter. Decoupled weight de- cay regularization. InInternational Conference on Learning Representations (ICLR), 2019. 15
2019
-
[52]
DPM-solver: A fast ode solver for diffusion probabilistic model sampling in around 10 steps
Cheng Lu, Yuhao Zhou, Fan Bao, Jianfei Chen, Chongxuan Li, and Jun Zhu. DPM-solver: A fast ode solver for diffusion probabilistic model sampling in around 10 steps. InNeurIPS,
-
[53]
Controllable person image synthesis with attribute-decomposed gan
Yifang Men, Yiming Mao, Yuning Jiang, Wei-Ying Ma, and Zhouhui Lian. Controllable person image synthesis with attribute-decomposed gan. InConference on Computer Vi- sion and Pattern Recognition (CVPR), 2020. 2
2020
-
[54]
Ben: Using confidence- guided matting for dichotomous image segmentation.arXiv preprint arXiv:2501.06230, 2025
Maxwell Meyer and Jack Spruyt. Ben: Using confidence- guided matting for dichotomous image segmentation.arXiv preprint arXiv:2501.06230, 2025. 15
2025
-
[55]
Dress code: High- resolution multi-category virtual try-on
Davide Morelli, Matteo Fincato, Marcella Cornia, Federico Landi, Fabio Cesari, and Rita Cucchiara. Dress code: High- resolution multi-category virtual try-on. InConference on Computer Vision and Pattern Recognition (CVPR), 2022. 5, 8, 15, 17
2022
-
[56]
LaDI- VTON: Latent Diffusion Textual-Inversion Enhanced Virtual Try-On
Davide Morelli, Alberto Baldrati, Giuseppe Cartella, Mar- cella Cornia, Marco Bertini, and Rita Cucchiara. LaDI- VTON: Latent Diffusion Textual-Inversion Enhanced Virtual Try-On. InACM International Conference on Multimedia (ACMMM), 2023. 2
2023
-
[57]
Vella 1.0.https://vellaml.com/, 2025
OmniousAI. Vella 1.0.https://vellaml.com/, 2025. 6, 9
2025
-
[58]
Chatgpt-4o: Multimodal ai model.https:// openai
OpenAI. Chatgpt-4o: Multimodal ai model.https:// openai . com / index / introducing - 4o - image - generation/, 2025. 6, 9
2025
-
[59]
Maxime Oquab, Timoth ´ee Darcet, Th´eo Moutakanni, Huy V . V o, Marc Szafraniec, Vasil Khalidov, Pierre Fernandez, Daniel HAZIZA, Francisco Massa, Alaaeldin El-Nouby, Mido Assran, Nicolas Ballas, Wojciech Galuba, Russell Howes, Po-Yao Huang, Shang-Wen Li, Ishan Misra, Michael ...
2024
-
[60]
Scalable diffusion models with transformers
William Peebles and Saining Xie. Scalable diffusion models with transformers. InInternational Conference on Computer Vision (ICCV), 2023. 3, 4
2023
-
[61]
Pic copilot - virtual try-on tool.https:// www.piccopilot.com/, 2024
Pic Copilot. Pic copilot - virtual try-on tool.https:// www.piccopilot.com/, 2024. 6, 9
2024
-
[62]
SDXL: Improving latent diffusion models for high-resolution image synthesis
Dustin Podell, Zion English, Kyle Lacey, Andreas Blattmann, Tim Dockhorn, Jonas M ¨uller, Joe Penna, and Robin Rombach. SDXL: Improving latent diffusion models for high-resolution image synthesis. InInternational Con- ference on Learning Representations (ICLR), 2024. 2
2024
-
[63]
Learn- ing transferable visual models from natural language super- vision
Alec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh, Gabriel Goh, Sandhini Agarwal, Girish Sastry, Amanda Askell, Pamela Mishkin, Jack Clark, et al. Learn- ing transferable visual models from natural language super- vision. InInternational Conference on Machine Learning ...
2021
-
[64]
Zero-shot text-to-image generation
Aditya Ramesh, Mikhail Pavlov, Gabriel Goh, Scott Gray, Chelsea V oss, Alec Radford, Mark Chen, and Ilya Sutskever. Zero-shot text-to-image generation. InInternational Confer- ence on Machine Learning (ICML), 2021. 2
2021
-
[65]
Deepspeed: System optimizations enable train- ing deep learning models with over 100 billion parameters
Jeff Rasley, Samyam Rajbhandari, Olatunji Ruwase, and Yuxiong He. Deepspeed: System optimizations enable train- ing deep learning models with over 100 billion parameters. InProceedings of the 26th ACM SIGKDD international con- ference on knowledge discovery & data mining, 2020. 15
2020
-
[66]
High-resolution image syn- thesis with latent diffusion models
Robin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser, and Bj¨orn Ommer. High-resolution image syn- thesis with latent diffusion models. InConference on Com- puter Vision and Pattern Recognition (CVPR), 2022. 2, 3
2022
-
[67]
Dreambooth: Fine tuning text-to-image diffusion models for subject-driven generation
Nataniel Ruiz, Yuanzhen Li, Varun Jampani, Yael Pritch, Michael Rubinstein, and Kfir Aberman. Dreambooth: Fine tuning text-to-image diffusion models for subject-driven generation. InProceedings of the IEEE/CVF conference on computer vision and pattern recognition, 2023. 2
2023
-
[68]
Photorealistic text-to-image diffusion models with deep language understanding
Chitwan Saharia, William Chan, Saurabh Saxena, Lala Li, Jay Whang, Emily L Denton, Kamyar Ghasemipour, Raphael Gontijo Lopes, Burcu Karagol Ayan, Tim Salimans, et al. Photorealistic text-to-image diffusion models with deep language understanding. InAdvances in Neural Informati...
2022
-
[69]
Imagdressing-v1: Cus- tomizable virtual dressing
Fei Shen, Xin Jiang, Xin He, Hu Ye, Cong Wang, Xiaoyu Du, Zechao Li, and Jinhui Tang. Imagdressing-v1: Cus- tomizable virtual dressing. InAAAI Conference on Artificial Intelligence (AAAI), 2025. 2
2025
-
[70]
Real-time single image and video super-resolution using an efficient sub-pixel convolutional neural network
Wenzhe Shi, Jose Caballero, Ferenc Husz ´ar, Johannes Totz, Andrew P Aitken, Rob Bishop, Daniel Rueckert, and Zehan Wang. Real-time single image and video super-resolution using an efficient sub-pixel convolutional neural network. InConference on Computer Vision and Pattern Re...
2016
-
[71]
Motion-i2v: Consistent and controllable image-to-video generation with explicit motion modeling
Xiaoyu Shi, Zhaoyang Huang, Fu-Yun Wang, Weikang Bian, Dasong Li, Yi Zhang, Manyuan Zhang, Ka Chun Cheung, Simon See, Hongwei Qin, et al. Motion-i2v: Consistent and controllable image-to-video generation with explicit motion modeling. InIn ACM SIGGRAPH, 2024. 2
2024
-
[72]
Denois- ing diffusion implicit models
Jiaming Song, Chenlin Meng, and Stefano Ermon. Denois- ing diffusion implicit models. InInternational Conference on Learning Representations (ICLR), 2021. 3
2021
-
[73]
Score-based generative modeling through stochastic differential equa- tions
Yang Song, Jascha Sohl-Dickstein, Diederik P Kingma, Ab- hishek Kumar, Stefano Ermon, and Ben Poole. Score-based generative modeling through stochastic differential equa- tions. InInternational Conference on Learning Represen- tations (ICLR), 2021. 3
2021
-
[74]
Roformer: Enhanced transformer with rotary position embedding.Neurocomputing, 568:127063,
Jianlin Su, Murtadha Ahmed, Yu Lu, Shengfeng Pan, Wen Bo, and Yunfeng Liu. Roformer: Enhanced transformer with rotary position embedding.Neurocomputing, 568:127063,
-
[75]
Ominicontrol: Minimal and uni- versal control for diffusion transformer.arXiv preprint arXiv:2411.15098, 2024
Zhenxiong Tan, Songhua Liu, Xingyi Yang, Qiaochu Xue, and Xinchao Wang. Ominicontrol: Minimal and uni- versal control for diffusion transformer.arXiv preprint arXiv:2411.15098, 2024. 2
2024 arXiv
-
[76]
Emo: Emote portrait alive generating expressive portrait videos with audio2video diffusion model under weak conditions
Linrui Tian, Qi Wang, Bang Zhang, and Liefeng Bo. Emo: Emote portrait alive generating expressive portrait videos with audio2video diffusion model under weak conditions. In European Conference on Computer Vision (ECCV), 2024. 2 13
2024
-
[77]
Attention is all you need
Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszko- reit, Llion Jones, Aidan N Gomez, Łukasz Kaiser, and Illia Polosukhin. Attention is all you need. InAdvances in Neural Information Processing Systems (NeurIPS), 2017. 5
2017
-
[78]
TryOffDiff: Virtual-try-off via high-fidelity gar- ment reconstruction using diffusion models.arXiv, 2024
Riza Velioglu, Petra Bevandic, Robin Chan, and Barbara Hammer. TryOffDiff: Virtual-try-off via high-fidelity gar- ment reconstruction using diffusion models.arXiv, 2024. 2, 3, 6, 9, 10
2024
-
[79]
Sv3d: Novel multi-view syn- thesis and 3d generation from a single image using latent video diffusion
Vikram V oleti, Chun-Han Yao, Mark Boss, Adam Letts, David Pankratz, Dmitry Tochilkin, Christian Laforte, Robin Rombach, and Varun Jampani. Sv3d: Novel multi-view syn- thesis and 3d generation from a single image using latent video diffusion. InEuropean Conference on Computer ...
2024
-
[80]
Toward characteristic- preserving image-based virtual try-on network
Bochao Wang, Huabin Zheng, Xiaodan Liang, Yimin Chen, Liang Lin, and Meng Yang. Toward characteristic- preserving image-based virtual try-on network. InEuropean Conference on Computer Vision (ECCV), 2018. 2
2018
-
[81]
Image quality assessment: from error visibility to structural similarity.IEEE transactions on image processing, 13(4):600–612, 2004
Zhou Wang, Alan C Bovik, Hamid R Sheikh, and Eero P Si- moncelli. Image quality assessment: from error visibility to structural similarity.IEEE transactions on image processing, 13(4):600–612, 2004. 8
2004
-
[82]
TryOf- fAnyone: Tiled cloth generation from a dressed person
Ioannis Xarchakos and Theodoros Koukopoulos. TryOf- fAnyone: Tiled cloth generation from a dressed person. arXiv, 2025. 2, 3, 6, 9, 10
2025
-
[83]
Towards scalable unpaired virtual try-on via patch-routed spatially- adaptive gan
Zhenyu Xie, Zaiyu Huang, Fuwei Zhao, Haoye Dong, Michael Kampffmeyer, and Xiaodan Liang. Towards scalable unpaired virtual try-on via patch-routed spatially- adaptive gan. InAdvances in Neural Information Processing Systems (NeurIPS), 2021. 2
2021
-
[84]
OOTDiffusion: Outfitting fusion based latent diffusion for controllable virtual try-on
Yuhao Xu, Tao Gu, Weifeng Chen, and Chengcai Chen. OOTDiffusion: Outfitting fusion based latent diffusion for controllable virtual try-on. InAAAI Conference on Artificial Intelligence (AAAI), 2025. 2, 8
2025
-
[85]
Towards photo-realistic virtual try-on by adaptively generating-preserving image content
Han Yang, Ruimao Zhang, Xiaobao Guo, Wei Liu, Wang- meng Zuo, and Ping Luo. Towards photo-realistic virtual try-on by adaptively generating-preserving image content. InConference on Computer Vision and Pattern Recognition (CVPR), 2020. 2
2020
-
[86]
IP- Adapter: Text compatible image prompt adapter for text-to- image diffusion models
Hu Ye, Jun Zhang, Sibo Liu, Xiao Han, and Wei Yang. IP- Adapter: Text compatible image prompt adapter for text-to- image diffusion models. 2023. 2
2023
-
[87]
CAT-DM: Controllable accel- erated virtual try-on with diffusion model
Jianhao Zeng, Dan Song, Weizhi Nie, Hongshuo Tian, Tong- tong Wang, and An-An Liu. CAT-DM: Controllable accel- erated virtual try-on with diffusion model. InConference on Computer Vision and Pattern Recognition (CVPR), 2024. 2
2024
-
[88]
Robust-mvton: Learn- ing cross-pose feature alignment and fusion for robust multi- view virtual try-on
Nannan Zhang, Yijiang Li, Dong Du, Zheng Chong, Zheng- wentai Sun, Jianhao Zeng, Yusheng Dai, Zhengyu Xie, Hairui Zhu, and Xiaoguang Han. Robust-mvton: Learn- ing cross-pose feature alignment and fusion for robust multi- view virtual try-on. InConference on Computer Vision and...
2025
-
[89]
Efros, Eli Shecht- man, and Oliver Wang
Richard Zhang, Phillip Isola, Alexei A. Efros, Eli Shecht- man, and Oliver Wang. The unreasonable effectiveness of deep features as a perceptual metric.arXiv, 2018. 8
2018
-
[90]
GP- VTON: Towards general purpose virtual try-on via collabo- rative local-flow global-parsing learning
Xie Zhenyu, Huang Zaiyu, Dong Xin, Zhao Fuwei, Dong Haoye, Zhang Xijin, Zhu Feida, and Liang Xiaodan. GP- VTON: Towards general purpose virtual try-on via collabo- rative local-flow global-parsing learning. InConference on Computer Vision and Pattern Recognition (CVPR), 2023. 2
2023
-
[91]
Learning flow fields in attention for controllable person image generation
Zijian Zhou, Shikun Liu, Xiao Han, Haozhe Liu, Kam Woh Ng, Tian Xie, Yuren Cong, Hang Li, Mengmeng Xu, Juan- Manuel P´erez-R´ua, Aditya Patel, Tao Xiang, Miaojing Shi, and Sen He. Learning flow fields in attention for controllable person image generation. InConference on Compu...
2025
-
[92]
TryOnDiffusion: A tale of two unets
Luyang Zhu, Dawei Yang, Tyler Zhu, Fitsum Reda, William Chan, Chitwan Saharia, Mohammad Norouzi, and Ira Kemelmacher-Shlizerman. TryOnDiffusion: A tale of two unets. InConference on Computer Vision and Pattern Recog- nition (CVPR), 2023. 2 14 pretrained diffusion prior while e...
2023
-
[93]
PAKU”, a dog graphic, or “CREAMSODA
Details on Human Evaluation This section provides details on the user study protocol de- scribed in Sec. 4.4. Participants were shown a reference garment and person image, along with five generated try-on results from our model and four from other state-of-the-art baselines. E...
-
[94]
Failure Cases While our method demonstrates superior and robust perfor- mance across various conditions, there are certain failure cases worth noting (Fig. 16). One common issue occurs when the input mask provided by the user or generated automatically does not fully cover the...
-
[95]
17, 22 which contains additional try-on and try-off results across a wide range of garments and human appearances
Extra Qualitative Results We further showcase the versatility of our model in Fig. 17, 22 which contains additional try-on and try-off results across a wide range of garments and human appearances. Figure 16. Failure cases of our method. (Top) When the input mask, either user-...
-
[96]
More Qualitative Comparisons We present additional qualitative comparisons with state-of- the-art baselines on VITON-HD [13], DressCode [55], and a set of in-the-wild examples gathered from online sources. Fig. 18 shows try-on results on VITON-HD, highlight- ing our model’s ab...
Reviewed August 5, 2026 · model on record in the stance chip above.
Discussion (0). Sign in to comment.