Pith. sign in

REVIEW 4 major objections 2 minor 96 references

Numerical Study of Oblique Detonation Initiation Assisted by Local Energy Deposition

T0 review · 4 major / 2 minor · reviewed 2026-08-15 · deepseek-v4-flash

Pith's one-line read This paper claims that pulsatile local energy deposition sustains oblique detonation on a finite wedge with less than 10% of the average power continuous deposition needs.

desk verdict The submission is unusable: the abstract is a fluid-dynamics study, but the full text is an unrelated AI image-generation paper, so there is nothing to assess. read the letter →

arxiv 2508.08943 v1 pith:GWXMXEKH submitted 2025-08-12 physics.flu-dyn

classification physics.flu-dyn
keywords obliquedetonationwaveenginelocalenergydepositionpulsatileinitiationfinitewedgeplasmaassistance
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

At low flight Mach numbers, a plain wedge may fail to initiate an oblique detonation wave—a combustion front attached to an inclined shock—which would stall combustion in an oblique detonation engine. This paper uses numerical simulation to show that depositing energy locally near the wedge overcomes that failure, and that the deposition can be delivered in short pulses instead of continuously. The central result is that pulsatile deposition sustains on-wedge detonation with an average power below 10% of the continuous-deposition power while keeping the same initiation length. The authors also map how initiation modes change as deposition power or pulse energy grows, and identify a minimum pulse repetition frequency needed for sustainable detonation. A sympathetic reading is that pulse-based plasma assistance could make oblique detonation engines workable under low-Mach, high-altitude conditions that a passive wedge cannot handle.

What carries the argument

The central object is the local energy deposition region placed near the finite wedge, which models the thermal effects of plasma-based initiation assistance. The mechanism being studied is how heat added ahead of or on the wedge couples with the oblique shock to form and sustain a detonation wave. For continuous deposition, the controlling parameter is the deposition power; for pulsatile deposition, the controlling parameters are single-pulse energy and pulse repetition frequency. The spatiotemporal evolution of the primary wave structures supplies the criterion for the minimum repetition frequency, and multi-pulse simulations are the verification step that closes the argument.

What would settle it

A multi-pulse simulation at the claimed minimum repetition frequency and average power, run with a different grid resolution or chemical mechanism, that either fails to sustain the on-wedge detonation or produces a longer initiation length would falsify the claim. The same test could be done experimentally by measuring the initiation length and detonation sustainability for pulsed deposition at an average power of 10% of the continuous threshold.

Watch

Extended reading notes

Core claim

The paper reports that without energy deposition, oblique detonation initiation fails on a finite wedge at low Mach numbers; with either continuous or pulsatile local energy deposition, a sustainable oblique detonation can be established on the wedge. As the continuous deposition power or the single-pulse energy increases, the initiations pass through a sequence of distinct modes, and evolution of the main wave structures under single-pulse deposition reveals a minimum pulse repetition frequency for maintaining the on-wedge detonation. Multi-pulse simulations confirm that this frequency keeps the detonation sustainable. The headline quantitative finding is that pulsatile deposition reaches the same initiation length as continuous deposition while consuming less than 10% of the average power, making pulsed energy the efficient route for initiation assistance on finite wedges under extreme flight conditions.

Load-bearing premise

The load-bearing premise is that the numerical setup—the finite wedge geometry, inflow conditions, and the way energy deposition is modelled—represents real oblique-detonation-engine flight well enough that the 10% power comparison carries over beyond the simulated cases.

Editorial extensions

If this is right

  • If the result transfers to real engines, pulsed plasma deposition could replace continuous deposition for ODW initiation at low Mach numbers, cutting initiation-assistance power demand by more than an order of magnitude.
  • The identified minimum pulse repetition frequency gives engine designers a concrete control parameter for keeping an on-wedge detonation sustainable.
  • The sequential initiation modes with increasing deposition power or pulse energy provide a map for choosing operating points that avoid marginal or failed initiation.
  • Because the same initiation length is preserved at a fraction of the average power, pulsatile assistance can be integrated without lengthening the wedge or the engine.
  • Without any deposition, the finite wedge is predicted to fail initiation at low Mach, so energy assistance becomes a necessary rather than optional component in that flight regime.

Reading between the lines

Editorial extensions of the paper, not claims the author makes directly.

  • The 10% ratio is reported for the simulated conditions; an inference is that the gap between pulsed and continuous power may widen or shrink with deposition location, pulse shape, or wedge angle, since those parameters are not varied in the abstract's headline comparison.
  • The minimum pulse repetition frequency criterion suggests a resonance-like coupling between pulse timing and detonation wave structure, which could be probed directly by measuring initiation length as a function of frequency.
  • An experimental shock-tube test with pulsed laser or plasma deposition on a finite wedge could test whether the simulated 10% average-power advantage survives in a real reacting flow.
  • For high-altitude low-Mach flight, the practical implication is that the engine's initiation system should be specified by time-averaged power plus repetition frequency, not just peak energy.
Share X Bluesky LinkedIn Reddit HN

Signed reviews

No signed human review yet.

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, and a circularity audit.

Referee Report

4 major / 2 minor

Summary. The submission, arXiv:2508.08943, presents an abstract in physics.flu-dyn claiming a numerical study of oblique detonation wave (ODW) initiation assisted by local energy deposition on a finite wedge, asserting that pulsatile energy deposition can sustain on-wedge detonation at less than 10% of the continuous-deposition power while maintaining the same initiation length. However, the full text of the manuscript is an unrelated paper titled "Lay2Story: Extending Diffusion Transformers for Layout-Togglable Story Generation," which is a computer vision manuscript about AI image generation. The body contains no equations, simulation setup, wedge geometry, inflow conditions, energy deposition model, grid resolution, chemistry or turbulence closure, convergence studies, or any results relevant to oblique detonation waves. The central claims of the abstract are therefore entirely unsupported by the submitted manuscript text.

Significance. If the claimed results were present and valid, the finding that pulsatile energy deposition consumes less than 10% of the continuous-deposition power for the same initiation length could be of practical interest for oblique detonation engine initiation at low Mach number and high altitude. The abstract also identifies a physically sensible mode sequence and a minimum repetition frequency estimation procedure using single-pulse simulations followed by multi-pulse verification. However, none of this content is present in the submitted manuscript. Because the scientific content described in the abstract is completely absent, the significance of the work cannot be assessed. The manuscript cannot be evaluated on its merits, and no strength such as reproducible code, machine-checked proofs, or falsifiable predictions can be credited.

major comments (4)
  1. [Full text (entire manuscript)] The full text of the submitted manuscript is the paper "Lay2Story: Extending Diffusion Transformers for Layout-Togglable Story Generation" (arXiv:2508.08949), which is a computer vision paper on diffusion transformers for image generation. This body text shares no content with the physics.flu-dyn abstract on oblique detonation wave initiation. There is no numerical solver description, no governing equations, no wedge geometry, no inflow conditions, no energy deposition model, no grid resolution, no chemistry or turbulence closure, no convergence studies, and no simulation results. The central claim of the abstract — that pulsatile energy deposition sustains on-wedge detonation at less than 10% of the continuous-deposition power with the same initiation length — is asserted without any supporting derivation, data, or methodology in the manuscript.
  2. [Abstract, final sentence] The claim that "sustainable on-wedge detonation can be achieved by pulsatile energy deposition with an average power consumption of less than 10% of that required for continuous energy deposition while maintaining a same initiation length" is a quantitative efficiency comparison that requires precise definitions of "average power consumption," "initiation length," and the parameter sets for both continuous and pulsatile deposition. None of these definitions appear in the manuscript, and no simulation data are presented that could substantiate the 10% figure. The claim is unverifiable from the submitted text.
  3. [Abstract, spatiotemporal evolution sentence] The abstract states that "Analysis of the spatiotemporal evolution of the primary wave structures under single-pulse energy deposition reveals the minimum pulse repetition frequency required for sustainable on-wedge detonation, which is subsequently verified through multi-pulse energy deposition simulations." This describes a two-step procedure that is central to the paper's methodology. The manuscript provides neither the single-pulse analysis nor the multi-pulse verification, nor any description of how the minimum repetition frequency is extracted from the spatiotemporal evolution. The absence of this material places the entire scientific argument outside the submitted text.
  4. [Full text (all sections)] The manuscript contains no limitations section, no error analysis, and no discussion of the validity of the simulation approach. Given that the body text is unrelated to the abstract, the work as submitted is internally inconsistent. No amount of revision to the present text could make the central claim assessable; the submission would require replacement by an entirely different manuscript containing the actual numerical study described in the abstract.
minor comments (2)
  1. [Abstract, first sentence] The abstract refers to "a fixed-angle wedge" and later to "on-wedge initiation," but the body text contains no wedge geometry or coordinate system, so these terms are undefined in the submitted document.
  2. [Abstract, methods sentence] The phrase "plasma-based initiation assistance techniques" suggests a specific physical model of energy deposition, but no plasma model, energy source term, or timescale is described anywhere in the manuscript.

Circularity Check

0 steps flagged · score 0.0 of 10

No circularity can be identified; the submitted full text does not contain the claimed study's methods or results.

full rationale

The abstract reports a numerical study of oblique detonation initiation with local energy deposition, but the provided full text is an unrelated paper on AI-based story image generation (Lay2Story). There is therefore no derivation chain, no equations, no fitted parameters, no self-citations, and no imported uniqueness theorem available to inspect. From the abstract alone, the described procedure—single-pulse analysis yields a minimum pulse repetition frequency, which is then verified in multi-pulse simulations—is an internal consistency check rather than a quantity defined to equal an input. No fitted constant is renamed as a prediction, and no result is forced by definition or by self-citation. The absence of the actual manuscript text and simulation data is a verifiability and integrity concern, not a circularity finding. Per the hard rules, circularity may only be claimed when the paper itself can be quoted to exhibit a specific reduction, so the honest outcome is no circularity, score 0.

Assumptions & free parameters 0 free parameters · 2 assumptions · 0 invented entities

Only the abstract is available for review; the numerical parameters, solver choices, and model constants cannot be enumerated from the abstract alone, so the ledger lists the modeling assumptions we can infer.

assumptions (2)
  • domain assumption The numerical governing equations and combustion chemistry models used in the simulations adequately capture oblique detonation initiation physics.
    Inferred from the abstract; no details on the solver or models are provided, yet the claim relies on simulation fidelity.
  • domain assumption The energy deposition model represents the thermal effect of plasma-based initiation assistance.
    The abstract says the study assesses thermal effects of plasma-based initiation, so the deposition model is assumed to capture the relevant physics.

how reviews work

0 comments
Cite this review

Pith. "Pith review of Numerical Study of Oblique Detonation Initiation Assisted by Local Energy Deposition." pith.science (2026). https://pith.science/paper/GWXMXEKH

@misc{pith2026250808943,
  author       = {Pith},
  title        = {Pith review of: Numerical Study of Oblique Detonation Initiation Assisted by Local Energy Deposition},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/GWXMXEKH}},
  note         = {Machine review of arXiv:2508.08943}
}
read the original abstract

Reliable initiation of oblique detonation waves (ODWs) is crucial for the stable operation of oblique detonation engines (ODEs), especially under flight conditions of low Mach numbers and/or high altitudes. In this case, conventional initiation approaches relying solely on a fixed-angle wedge may engender risks of initiation failure, which necessitates extra initiation assistance measures. In this study, ODW initiation over a finite wedge with local energy deposition is numerically investigated to assess the thermal effects of plasma-based initiation assistance techniques. Particular emphasis is put on the effects of forms and magnitudes of energy deposition on initiation modes and flow field structures of ODWs. The results demonstrate that on-wedge initiation of ODWs fails at a low Mach number without any energy depositions. In contrast, both continuous and pulsatile local energy depositions can effectively initiate ODWs, leading to sustainable detonation on the finite wedge. As continuous energy deposition power or pulsatile single pulse energy increases, several key detonation initiation modes emerge sequentially. Analysis of the spatiotemporal evolution of the primary wave structures under single-pulse energy deposition reveals the minimum pulse repetition frequency required for sustainable on-wedge detonation, which is subsequently verified through multi-pulse energy deposition simulations. Nevertheless, it is found that sustainable on-wedge detonation can be achieved by pulsatile energy deposition with an average power consumption of less than 10% of that required for continuous energy deposition while maintaining a same initiation length, suggesting that the pulsatile one is an efficient way of energy deposition for initiation assistance of ODWs on finite wedges under extreme flight conditions.

Discussion (0). Continue with ORCID to comment.

Reference graph

Works this paper leans on

96 extracted references · 37 canonical work pages

  1. [5]

    Customttt: Motion and ap- pearance customized video generation via test-time training

    Xiuli Bi, Jian Lu, Bo Liu, Xiaodong Cun, Yong Zhang, Weisheng Li, and Bin Xiao. Customttt: Motion and ap- pearance customized video generation via test-time training. In Proceedings of the AAAI Conference on Artificial Intelli- gence, pages 1871–1879, 2025. 2

  2. [6]

    Relactrl: Relevance-guided efficient control for diffusion transformers

    Ke Cao, Jing Wang, Ao Ma, Jiasong Feng, Zhanjie Zhang, Xuanhua He, Shanyuan Liu, Bo Cheng, Dawei Leng, Yuhui Yin, et al. Relactrl: Relevance-guided efficient control for diffusion transformers. arXiv preprint arXiv:2502.14377 ,

  3. [7]

    Gamegen-x: Interactive open-world game video generation

    Haoxuan Che, Xuanhua He, Quande Liu, Cheng Jin, and Hao Chen. Gamegen-x: Interactive open-world game video generation. arXiv preprint arXiv:2411.00769, 2024. 3

  4. [8]

    Pixart- � : Fast training of diffusion transformer for photorealistic text-to-image synthesis

    Junsong Chen, Jincheng Yu, Chongjian Ge, Lewei Yao, Enze Xie, Yue Wu, Zhongdao Wang, James Kwok, Ping Luo, Huchuan Lu, et al. Pixart- � : Fast training of diffusion transformer for photorealistic text-to-image synthesis. arXiv preprint arXiv:2310.00426, 2023. 2, 9, 11

  5. [9]

    Panda-70m: Captioning 70m videos with multiple cross-modality teachers

    Tsai-Shien Chen, Aliaksandr Siarohin, Willi Menapace, Ekaterina Deyneka, Hsiang-wei Chao, Byung Eun Jeon, Yuwei Fang, Hsin-Ying Lee, Jian Ren, Ming-Hsuan Yang, et al. Panda-70m: Captioning 70m videos with multiple cross-modality teachers. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition , pages 13320–13331, 2024. 3

  6. [10]

    Ctr-driven advertising image generation with multimodal large language models

    Xingye Chen, Wei Feng, Zhenbang Du, Weizhen Wang, Yanyin Chen, Haohan Wang, Linkai Liu, Yaoyu Li, Jinyuan Zhao, Yu Li, et al. Ctr-driven advertising image generation with multimodal large language models. In Proceedings of the ACM on Web Conference 2025, pages 2262–2275, 2025. 11

  7. [11]

    Scaling recti- fied flow transformers for high-resolution image synthesis

    Patrick Esser, Sumith Kulal, Andreas Blattmann, Rahim Entezari, Jonas M ¨uller, Harry Saini, Yam Levi, Dominik Lorenz, Axel Sauer, Frederic Boesel, et al. Scaling recti- fied flow transformers for high-resolution image synthesis. In Forty-first international conference on machine learning,

  8. [12]

    FancyVideo: Towards Dynamic and Consistent Video Generation via Cross-frame Textual Guidance

    Jiasong Feng, Ao Ma, Jing Wang, Bo Cheng, Xiaodan Liang, Dawei Leng, and Yuhui Yin. Fancyvideo: Towards dynamic and consistent video generation via cross-frame textual guid- ance. arXiv preprint arXiv:2408.08189, 2024. 3

Show all 96 references
  1. [13]

    Dream- sim: Learning new dimensions of human visual similar- ity using synthetic data

    Stephanie Fu, Netanel Tamir, Shobhita Sundaram, Lucy Chai, Richard Zhang, Tali Dekel, and Phillip Isola. Dream- sim: Learning new dimensions of human visual similar- ity using synthetic data. arXiv preprint arXiv:2306.09344 ,

  2. [14]

    An image is worth one word: Personalizing text-to- image generation using textual inversion

    Rinon Gal, Yuval Alaluf, Yuval Atzmon, Or Patash- nik, Amit H Bermano, Gal Chechik, and Daniel Cohen- Or. An image is worth one word: Personalizing text-to- image generation using textual inversion. arXiv preprint arXiv:2208.01618, 2022. 9

  3. [15]

    Check locate rectify: A training- free layout calibration system for text-to-image generation

    Biao Gong, Siteng Huang, Yutong Feng, Shiwei Zhang, Yuyuan Li, and Yu Liu. Check locate rectify: A training- free layout calibration system for text-to-image generation. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 6624–6634, 2024. 9

  4. [16]

    Variational au- toencoder: An unsupervised model for encoding and decod- ing fmri activity in visual cortex.NeuroImage, 198:125–136,

    Kuan Han, Haiguang Wen, Junxing Shi, Kun-Han Lu, Yizhen Zhang, Di Fu, and Zhongming Liu. Variational au- toencoder: An unsupervised model for encoding and decod- ing fmri activity in visual cortex.NeuroImage, 198:125–136,

  5. [17]

    Anystory: Towards unified single and multiple subject personalization in text-to-image generation

    Junjie He, Yuxiang Tuo, Binghui Chen, Chongyang Zhong, Yifeng Geng, and Liefeng Bo. Anystory: Towards unified single and multiple subject personalization in text-to-image generation. arXiv preprint arXiv:2501.09503, 2025. 2

  6. [18]

    Freeedit: Mask-free reference-based image editing with multi-modal instruction

    Runze He, Kai Ma, Linjiang Huang, Shaofei Huang, Jialin Gao, Xiaoming Wei, Jiao Dai, Jizhong Han, and Si Liu. Freeedit: Mask-free reference-based image editing with multi-modal instruction. arXiv preprint arXiv:2409.18071 ,

  7. [19]

    Plangen: Towards unified layout planning and image generation in auto-regressive vision language models

    Runze He, Bo Cheng, Yuhang Ma, Qingxiang Jia, Shanyuan Liu, Ao Ma, Xiaoyu Wu, Liebucha Wu, Dawei Leng, and Yuhui Yin. Plangen: Towards unified layout planning and image generation in auto-regressive vision language models. arXiv preprint arXiv:2503.10127, 2025. 3

  8. [20]

    Context- aware layout to image generation with enhanced object ap- pearance

    Sen He, Wentong Liao, Michael Ying Yang, Yongxin Yang, Yi-Zhe Song, Bodo Rosenhahn, and Tao Xiang. Context- aware layout to image generation with enhanced object ap- pearance. In Proceedings of the IEEE/CVF conference on computer vision and pattern recognition , pages 15049– 1...

  9. [21]

    Style aligned image generation via shared atten- tion

    Amir Hertz, Andrey V oynov, Shlomi Fruchter, and Daniel Cohen-Or. Style aligned image generation via shared atten- tion. In Proceedings of the IEEE/CVF Conference on Com- puter Vision and Pattern Recognition , pages 4775–4785,

  10. [22]

    Clipscore: A reference-free evaluation met- ric for image captioning

    Jack Hessel, Ari Holtzman, Maxwell Forbes, Ronan Le Bras, and Yejin Choi. Clipscore: A reference-free evaluation met- ric for image captioning. arXiv preprint arXiv:2104.08718,

  11. [23]

    Gans trained by a two time-scale update rule converge to a local nash equilib- rium

    Martin Heusel, Hubert Ramsauer, Thomas Unterthiner, Bernhard Nessler, and Sepp Hochreiter. Gans trained by a two time-scale update rule converge to a local nash equilib- rium. Advances in neural information processing systems , 30, 2017. 7

  12. [24]

    Interactdiffusion: Interaction con- trol in text-to-image diffusion models

    Jiun Tian Hoe, Xudong Jiang, Chee Seng Chan, Yap-Peng Tan, and Weipeng Hu. Interactdiffusion: Interaction con- trol in text-to-image diffusion models. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 6180–6189, 2024. 9

  13. [25]

    Learning disentangled iden- tifiers for action-customized text-to-image generation

    Siteng Huang, Biao Gong, Yutong Feng, Xi Chen, Yuqian Fu, Yu Liu, and Donglin Wang. Learning disentangled iden- tifiers for action-customized text-to-image generation. In Proceedings of the IEEE/CVF Conference on Computer Vi- sion and Pattern Recognition, pages 7797–7806, 2024. 9

  14. [26]

    Reversion: Diffusion-based relation inversion from images

    Ziqi Huang, Tianxing Wu, Yuming Jiang, Kelvin CK Chan, and Ziwei Liu. Reversion: Diffusion-based relation inversion from images. In SIGGRAPH Asia 2024 Conference Papers, pages 1–11, 2024. 9

  15. [27]

    Gpt-4o system card

    Aaron Hurst, Adam Lerer, Adam P Goucher, Adam Perel- man, Aditya Ramesh, Aidan Clark, AJ Ostrow, Akila Weli- hinda, Alan Hayes, Alec Radford, et al. Gpt-4o system card. arXiv preprint arXiv:2410.21276, 2024. 4

  16. [28]

    Res-tuning: A flexible and efficient tuning paradigm via unbinding tuner from backbone

    Zeyinzi Jiang, Chaojie Mao, Ziyuan Huang, Ao Ma, Yiliang Lv, Yujun Shen, Deli Zhao, and Jingren Zhou. Res-tuning: A flexible and efficient tuning paradigm via unbinding tuner from backbone. Advances in Neural Information Processing Systems, 36:42689–42716, 2023. 3

  17. [29]

    Miradata: A large-scale video dataset with long durations and structured captions

    Xuan Ju, Yiming Gao, Zhaoyang Zhang, Ziyang Yuan, Xin- tao Wang, Ailing Zeng, Yu Xiong, Qiang Xu, and Ying Shan. Miradata: A large-scale video dataset with long durations and structured captions. Advances in Neural Information Processing Systems, 37:48955–48970, 2025. 3

  18. [30]

    Story generation with crowdsourced plot graphs

    Boyang Li, Stephen Lee-Urban, George Johnston, and Mark Riedl. Story generation with crowdsourced plot graphs. In Proceedings of the AAAI Conference on Artificial Intelli- gence, pages 598–604, 2013. 2

  19. [31]

    Blip-diffusion: Pre- trained subject representation for controllable text-to-image generation and editing

    Dongxu Li, Junnan Li, and Steven Hoi. Blip-diffusion: Pre- trained subject representation for controllable text-to-image generation and editing. Advances in Neural Information Pro- cessing Systems, 36:30146–30166, 2023. 2, 6, 7, 8

  20. [32]

    Planning and rendering: Towards prod- uct poster generation with diffusion models

    Zhaochen Li, Fengheng Li, Wei Feng, Honghe Zhu, Yaoyu Li, Zheng Zhang, Jingjing Lv, Junjie Shen, Zhangang Lin, Jingping Shao, et al. Planning and rendering: Towards prod- uct poster generation with diffusion models. arXiv preprint arXiv:2312.08822, 2023. 11

  21. [33]

    Photomaker: Customizing re- alistic human photos via stacked id embedding

    Zhen Li, Mingdeng Cao, Xintao Wang, Zhongang Qi, Ming- Ming Cheng, and Ying Shan. Photomaker: Customizing re- alistic human photos via stacked id embedding. In Proceed- ings of the IEEE/CVF conference on computer vision and pattern recognition, pages 8640–8650, 2024. 9

  22. [34]

    Ragar: Retrieval augment person- alized image generation guided by recommendation

    Run Ling, Wenji Wang, Yuting Liu, Guibing Guo, Linying Jiang, and Xingwei Wang. Ragar: Retrieval augment person- alized image generation guided by recommendation. arXiv preprint arXiv:2505.01657, 2025. 2

  23. [35]

    Intelligent grimm-open-ended visual storytelling via latent diffusion models

    Chang Liu, Haoning Wu, Yujie Zhong, Xiaoyun Zhang, Yan- feng Wang, and Weidi Xie. Intelligent grimm-open-ended visual storytelling via latent diffusion models. In Proceed- ings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 6190–6200, 2024. 2, 3, 6, 7, 8

  24. [36]

    Grounding dino: Marrying dino with grounded pre-training for open-set object detection

    Shilong Liu, Zhaoyang Zeng, Tianhe Ren, Feng Li, Hao Zhang, Jie Yang, Qing Jiang, Chunyuan Li, Jianwei Yang, Hang Su, et al. Grounding dino: Marrying dino with grounded pre-training for open-set object detection. In European Conference on Computer Vision , pages 38–55. Springe...

  25. [37]

    Bridge dif- fusion model: Bridge chinese text-to-image diffusion model with english communities

    Shanyuan Liu, Bo Cheng, Yuhang Ma, Liebucha Wu, Ao Ma, Xiaoyu Wu, Dawei Leng, and Yuhui Yin. Bridge dif- fusion model: Bridge chinese text-to-image diffusion model with english communities. In Proceedings of the AAAI Con- ference on Artificial Intelligence, pages 5541–5549, 2025. 2

  26. [38]

    One-prompt-one-story: Free-lunch consistent text-to-image generation using a single prompt

    Tao Liu, Kai Wang, Senmao Li, Joost van de Weijer, Fa- had Shahbaz Khan, Shiqi Yang, Yaxing Wang, Jian Yang, and Ming-Ming Cheng. One-prompt-one-story: Free-lunch consistent text-to-image generation using a single prompt. arXiv preprint arXiv:2501.13554, 2025. 2, 6, 7, 8, 9

  27. [39]

    Recent advances in ood detection: Problems and approaches

    Shuo Lu, Yingsheng Wang, Lijun Sheng, Aihua Zheng, Lingxiao He, and Jian Liang. Recent advances in ood detection: Problems and approaches. arXiv preprint arXiv:2409.11884, 2024. 2

  28. [40]

    Uni-layout: Integrating human feedback in unified layout generation and evaluation

    Shuo Lu, Yanyin Chen, Wei Feng, Jiahao Fan, Fengheng Li, Zheng Zhang, Jingjing Lv, Junjie Shen, Ching Law, and Jian Liang. Uni-layout: Integrating human feedback in unified layout generation and evaluation. arXiv preprint arXiv:2508.02374, 2025. 2

  29. [41]

    Unified multi-modal latent diffusion for joint subject and text conditional image generation.arXiv preprint arXiv:2303.09319, 2023

    Yiyang Ma, Huan Yang, Wenjing Wang, Jianlong Fu, and Jiaying Liu. Unified multi-modal latent diffusion for joint subject and text conditional image generation.arXiv preprint arXiv:2303.09319, 2023. 9

  30. [42]

    Hico: Hierarchical controllable diffu- sion model for layout-to-image generation.Advances in Neu- ral Information Processing Systems , 37:128886–128910,

    Yuhang Ma, Shanyuan Liu, Ao Ma, Xiaoyu Wu, Dawei Leng, and Yuhui Yin. Hico: Hierarchical controllable diffu- sion model for layout-to-image generation.Advances in Neu- ral Information Processing Systems , 37:128886–128910,

  31. [43]

    Storynizor: Consistent story generation via inter-frame synchronized and shuffled id injection

    Yuhang Ma, Wenting Xu, Chaoyi Zhao, Keqiang Sun, Qinfeng Jin, Zeng Zhao, Changjie Fan, and Zhipeng Hu. Storynizor: Consistent story generation via inter-frame synchronized and shuffled id injection. arXiv preprint arXiv:2409.19624, 2024. 2, 3

  32. [44]

    Story-adapter: A training-free iterative framework for long story visualization

    Jiawei Mao, Xiaoke Huang, Yunfei Xie, Yuanqi Chang, Mude Hui, Bingjie Xu, and Yuyin Zhou. Story-adapter: A training-free iterative framework for long story visualization. arXiv preprint arXiv:2410.06244, 2024. 2

  33. [45]

    Lego: Learning to disentangle and invert personalized con- cepts beyond object appearance in text-to-image diffusion models

    Saman Motamed, Danda Pani Paudel, and Luc Van Gool. Lego: Learning to disentangle and invert personalized con- cepts beyond object appearance in text-to-image diffusion models. arXiv preprint arXiv:2311.13833, 2023. 9

  34. [46]

    Openvid-1m: A large-scale high-quality dataset for text-to- video generation

    Kepan Nan, Rui Xie, Penghao Zhou, Tiehan Fan, Zhen- heng Yang, Zhijie Chen, Xiang Li, Jian Yang, and Ying Tai. Openvid-1m: A large-scale high-quality dataset for text-to- video generation. arXiv preprint arXiv:2407.02371 , 2024. 3

  35. [47]

    Scalable diffusion models with transformers

    William Peebles and Saining Xie. Scalable diffusion models with transformers. In Proceedings of the IEEE/CVF inter- national conference on computer vision , pages 4195–4205,

  36. [48]

    Sdxl: Improving latent diffusion mod- els for high-resolution image synthesis

    Dustin Podell, Zion English, Kyle Lacey, Andreas Blattmann, Tim Dockhorn, Jonas M ¨uller, Joe Penna, and Robin Rombach. Sdxl: Improving latent diffusion mod- els for high-resolution image synthesis. arXiv preprint arXiv:2307.01952, 2023. 11

  37. [49]

    Learning transferable visual models from natural language supervi- sion

    Alec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh, Gabriel Goh, Sandhini Agarwal, Girish Sastry, Amanda Askell, Pamela Mishkin, Jack Clark, et al. Learning transferable visual models from natural language supervi- sion. In International conference on machine learning, ...

  38. [50]

    Exploring the limits of transfer learning with a unified text-to-text transformer

    Colin Raffel, Noam Shazeer, Adam Roberts, Katherine Lee, Sharan Narang, Michael Matena, Yanqi Zhou, Wei Li, and Peter J Liu. Exploring the limits of transfer learning with a unified text-to-text transformer. Journal of machine learning research, 21(140):1–67, 2020. 11

  39. [51]

    Hierarchical text-conditional image gener- ation with clip latents

    Aditya Ramesh, Prafulla Dhariwal, Alex Nichol, Casey Chu, and Mark Chen. Hierarchical text-conditional image gener- ation with clip latents. arXiv preprint arXiv:2204.06125, 1 (2):3, 2022. 2

  40. [52]

    High-resolution image synthesis with latent diffusion models

    Robin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser, and Bj ¨orn Ommer. High-resolution image synthesis with latent diffusion models. In Proceedings of the IEEE/CVF conference on computer vision and pattern recognition, pages 10684–10695, 2022. 2, 3, 9, 11

  41. [53]

    U- net: Convolutional networks for biomedical image segmen- tation

    Olaf Ronneberger, Philipp Fischer, and Thomas Brox. U- net: Convolutional networks for biomedical image segmen- tation. In Medical image computing and computer-assisted intervention–MICCAI 2015: 18th international conference, Munich, Germany, October 5-9, 2015, proceedings, pa...

  42. [54]

    Dreambooth: Fine tuning text-to-image diffusion models for subject-driven generation

    Nataniel Ruiz, Yuanzhen Li, Varun Jampani, Yael Pritch, Michael Rubinstein, and Kfir Aberman. Dreambooth: Fine tuning text-to-image diffusion models for subject-driven generation. In Proceedings of the IEEE/CVF conference on computer vision and pattern recognition , pages 2250...

  43. [55]

    We also utilize the FID [23] metrics to as- sess the quality of the generated images

    to remove the image background and replace it with random noise. We also utilize the FID [23] metrics to as- sess the quality of the generated images. Recall@1 mea- sures top-1 text-to-image matching accuracy, while human preference reflects averaged binary ratings from three ...

  44. [56]

    Carvekit: Automated high-quality back- ground removal framework

    Nikita Selin. Carvekit: Automated high-quality back- ground removal framework. https://github.com/ OPHoperHPO/image- background- remove-tool,

  45. [57]

    Eventvad: Training-free event-aware video anomaly detection

    Yihua Shao, Haojin He, Sijie Li, Siyu Chen, Xinwei Long, Fanhu Zeng, Yuxuan Fan, Muyang Zhang, Ziyang Yan, Ao Ma, et al. Eventvad: Training-free event-aware video anomaly detection. arXiv preprint arXiv:2504.13092, 2025. 2

  46. [58]

    Tr-dq: Time-rotation diffusion quantiza- tion

    Yihua Shao, Deyang Lin, Fanhu Zeng, Minxi Yan, Muyang Zhang, Siyu Chen, Yuxuan Fan, Ziyang Yan, Haozhe Wang, Jingcai Guo, et al. Tr-dq: Time-rotation diffusion quantiza- tion. arXiv preprint arXiv:2503.06564, 2025. 11

  47. [59]

    In-context meta lora generation

    Yihua Shao, Minxi Yan, Yang Liu, Siyu Chen, Wenjie Chen, Xinwei Long, Ziyang Yan, Lei Li, Chenyu Zhang, Nicu Sebe, et al. In-context meta lora generation. arXiv preprint arXiv:2501.17635, 2025. 11

  48. [60]

    Storybooth: Training-free multi-subject consistency for improved visual storytelling

    Jaskirat Singh, Junshen K Chen, Jonas K Kohler, and Michael F Cohen. Storybooth: Training-free multi-subject consistency for improved visual storytelling. In The Thir- teenth International Conference on Learning Representa- tions. 2

  49. [61]

    Styledrop: Text-to-image synthesis of any style

    Kihyuk Sohn, Lu Jiang, Jarred Barber, Kimin Lee, Nataniel Ruiz, Dilip Krishnan, Huiwen Chang, Yuanzhen Li, Irfan Essa, Michael Rubinstein, et al. Styledrop: Text-to-image synthesis of any style. Advances in Neural Information Pro- cessing Systems, 36:66860–66889, 2023. 9

  50. [62]

    Instantx flux.1-dev ip-adapter page, 2024

    InstantX Team. Instantx flux.1-dev ip-adapter page, 2024. 2, 6, 7, 8, 9

  51. [63]

    Training-free consis- tent text-to-image generation

    Yoad Tewel, Omri Kaduri, Rinon Gal, Yoni Kasten, Lior Wolf, Gal Chechik, and Yuval Atzmon. Training-free consis- tent text-to-image generation. ACM Transactions on Graph- ics (TOG), 43(4):1–18, 2024. 2, 4, 6, 7, 8, 9

  52. [64]

    Converting video formats with ffmpeg

    Suramya Tomar. Converting video formats with ffmpeg. Linux journal, 2006(146):10, 2006. 3

  53. [65]

    Face0: Instantaneously conditioning a text-to- image model on a face

    Dani Valevski, Danny Lumen, Yossi Matias, and Yaniv Leviathan. Face0: Instantaneously conditioning a text-to- image model on a face. InSIGGRAPH Asia 2023 Conference Papers, pages 1–10, 2023. 9

  54. [66]

    Attention is all you need

    Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszko- reit, Llion Jones, Aidan N Gomez, Łukasz Kaiser, and Illia Polosukhin. Attention is all you need. Advances in neural information processing systems, 30, 2017. 11

  55. [67]

    Is this loss informative? faster text-to-image customization by tracking objective dynamics

    Anton V oronov, Mikhail Khoroshikh, Artem Babenko, and Max Ryabinin. Is this loss informative? faster text-to-image customization by tracking objective dynamics. Advances in Neural Information Processing Systems , 36:37491–37510,

  56. [68]

    p+: Extended textual conditioning in text-to- image generation

    Andrey V oynov, Qinghao Chu, Daniel Cohen-Or, and Kfir Aberman. p+: Extended textual conditioning in text-to- image generation. arXiv preprint arXiv:2303.09522, 2023. 9

  57. [69]

    Qihoo-t2x: An efficiency-focused diffu- sion transformer via proxy tokens for text-to-any-task

    Jing Wang, Ao Ma, Jiasong Feng, Dawei Leng, Yuhui Yin, and Xiaodan Liang. Qihoo-t2x: An efficiency-focused diffu- sion transformer via proxy tokens for text-to-any-task. arXiv e-prints, pages arXiv–2409, 2024. 3

  58. [70]

    Wisa: World simulator assistant for physics-aware text-to-video generation

    Jing Wang, Ao Ma, Ke Cao, Jun Zheng, Zhanjie Zhang, Jiasong Feng, Shanyuan Liu, Yuhang Ma, Bo Cheng, Dawei Leng, et al. Wisa: World simulator assistant for physics-aware text-to-video generation. arXiv preprint arXiv:2503.08153, 2025. 2

  59. [71]

    Videofactory: Swap at- tention in spatiotemporal diffusions for text-to-video gener- ation

    Wenjing Wang, Huan Yang, Zixi Tuo, Huiguo He, Junchen Zhu, Jianlong Fu, and Jiaying Liu. Videofactory: Swap at- tention in spatiotemporal diffusions for text-to-video gener- ation. 2023. 3

  60. [72]

    Spnet: Learning stereo matching with slanted plane aggregation

    Yun Wang, Longguang Wang, Hanyun Wang, and Yulan Guo. Spnet: Learning stereo matching with slanted plane aggregation. IEEE Robotics and Automation Letters , 7(3): 6258–6265, 2022. 11

  61. [73]

    Internvid: A large-scale video-text dataset for multimodal understanding and generation

    Yi Wang, Yinan He, Yizhuo Li, Kunchang Li, Jiashuo Yu, Xin Ma, Xinhao Li, Guo Chen, Xinyuan Chen, Yaohui Wang, et al. Internvid: A large-scale video-text dataset for multimodal understanding and generation. arXiv preprint arXiv:2307.06942, 2023. 3

  62. [74]

    Adstereo: Efficient stereo matching with adaptive downsampling and disparity alignment

    Yun Wang, Kunhong Li, Longguang Wang, Junjie Hu, Dapeng Oliver Wu, and Yulan Guo. Adstereo: Efficient stereo matching with adaptive downsampling and disparity alignment. IEEE Transactions on Image Processing , 2025. 2

  63. [75]

    Learning robust stereo matching in the wild with selective mixture-of-experts

    Yun Wang, Longguang Wang, Chenghao Zhang, Yongjian Zhang, Zhanjie Zhang, Ao Ma, Chenyou Fan, Tin Lun Lam, and Junjie Hu. Learning robust stereo matching in the wild with selective mixture-of-experts. arXiv preprint arXiv:2507.04631, 2025. 11

  64. [76]

    Dualnet: Ro- bust self-supervised stereo matching with pseudo-label su- pervision

    Yun Wang, Jiahao Zheng, Chenghao Zhang, Zhanjie Zhang, Kunhong Li, Yongjian Zhang, and Junjie Hu. Dualnet: Ro- bust self-supervised stereo matching with pseudo-label su- pervision. In Proceedings of the AAAI Conference on Artifi- cial Intelligence, pages 8178–8186, 2025. 2

  65. [77]

    Styleadapter: A unified stylized image generation model

    Zhouxia Wang, Xintao Wang, Liangbin Xie, Zhongang Qi, Ying Shan, Wenping Wang, and Ping Luo. Styleadapter: A unified stylized image generation model. arXiv preprint arXiv:2309.01770, 2023. 9

  66. [78]

    Dropoutgs: Dropping out gaus- sians for better sparse-view rendering

    Yexing Xu, Longguang Wang, Minglin Chen, Sheng Ao, Li Li, and Yulan Guo. Dropoutgs: Dropping out gaus- sians for better sparse-view rendering. In Proceedings of the Computer Vision and Pattern Recognition Conference, pages 701–710, 2025. 11

  67. [79]

    Freestyle layout-to-image synthesis

    Han Xue, Zhiwu Huang, Qianru Sun, Li Song, and Wenjun Zhang. Freestyle layout-to-image synthesis. In Proceedings of the IEEE/CVF conference on computer vision and pattern recognition, pages 14256–14266, 2023. 9

  68. [80]

    Facestudio: Put your face everywhere in seconds

    Yuxuan Yan, Chi Zhang, Rui Wang, Yichao Zhou, Gege Zhang, Pei Cheng, Gang Yu, and Bin Fu. Facestudio: Put your face everywhere in seconds. arXiv preprint arXiv:2312.02663, 2023. 9

  69. [81]

    Seed-story: Multimodal long story generation with large language model

    Shuai Yang, Yuying Ge, Yang Li, Yukang Chen, Yixiao Ge, Ying Shan, and Yingcong Chen. Seed-story: Multimodal long story generation with large language model. arXiv preprint arXiv:2407.08683, 2024. 2, 3, 7, 9

  70. [82]

    Reco: Region-controlled text-to-image genera- tion

    Zhengyuan Yang, Jianfeng Wang, Zhe Gan, Linjie Li, Kevin Lin, Chenfei Wu, Nan Duan, Zicheng Liu, Ce Liu, Michael Zeng, et al. Reco: Region-controlled text-to-image genera- tion. In Proceedings of the IEEE/CVF Conference on Com- puter Vision and Pattern Recognition , pages 14246–14255,

  71. [83]

    Ip- adapter: Text compatible image prompt adapter for text-to- image diffusion models

    Hu Ye, Jun Zhang, Sibo Liu, Xiao Han, and Wei Yang. Ip- adapter: Text compatible image prompt adapter for text-to- image diffusion models. arXiv preprint arXiv:2308.06721,

  72. [84]

    Controlnet-xs: Rethinking the control of text-to- image diffusion models as feedback-control systems

    Denis Zavadski, Johann-Friedrich Feiden, and Carsten Rother. Controlnet-xs: Rethinking the control of text-to- image diffusion models as feedback-control systems. In European Conference on Computer Vision, pages 343–362. Springer, 2024. 2

  73. [85]

    Creatilayout: Siamese multimodal diffusion transformer for creative layout-to-image generation

    Hui Zhang, Dexiang Hong, Tingwei Gao, Yitong Wang, Jie Shao, Xinglong Wu, Zuxuan Wu, and Yu-Gang Jiang. Creatilayout: Siamese multimodal diffusion transformer for creative layout-to-image generation. arXiv preprint arXiv:2412.03859, 2024. 9

  74. [86]

    A survey on personalized content synthesis with diffusion models

    Xulu Zhang, Xiao-Yong Wei, Wengyu Zhang, Jinlin Wu, Zhaoxiang Zhang, Zhen Lei, and Qing Li. A survey on personalized content synthesis with diffusion models. arXiv preprint arXiv:2405.05538, 2024. 9

  75. [87]

    Generative active learning for image synthesis personalization

    Xulu Zhang, Wengyu Zhang, Xiaoyong Wei, Jinlin Wu, Zhaoxiang Zhang, Zhen Lei, and Qing Li. Generative active learning for image synthesis personalization. In Proceedings of the 32nd ACM International Conference on Multimedia , pages 10669–10677, 2024. 9

  76. [88]

    Towards highly realistic artistic style transfer via stable diffusion with step-aware and layer-aware prompt

    Zhanjie Zhang, Quanwei Zhang, Huaizhong Lin, Wei Xing, Juncheng Mo, Shuaicheng Huang, Jinheng Xie, Guangyuan Li, Junsheng Luan, Lei Zhao, et al. Towards highly realistic artistic style transfer via stable diffusion with step-aware and layer-aware prompt. In Proceedings of the ...

  77. [89]

    Artbank: Artistic style transfer with pre-trained diffusion model and implicit style prompt bank

    Zhanjie Zhang, Quanwei Zhang, Wei Xing, Guangyuan Li, Lei Zhao, Jiakai Sun, Zehua Lan, Junsheng Luan, Yiling Huang, and Huaizhong Lin. Artbank: Artistic style transfer with pre-trained diffusion model and implicit style prompt bank. In Proceedings of the AAAI Conference on Art...

  78. [90]

    Lgast: Towards high- quality arbitrary style transfer with local–global style learn- ing

    Zhanjie Zhang, Yuxiang Li, Ruichen Xia, Mengyuan Yang, Yun Wang, Lei Zhao, and Wei Xing. Lgast: Towards high- quality arbitrary style transfer with local–global style learn- ing. Neurocomputing, 623:129434, 2025. 2

  79. [91]

    U- stydit: Ultra-high quality artistic style transfer using diffu- sion transformers

    Zhanjie Zhang, Ao Ma, Ke Cao, Jing Wang, Shanyuan Liu, Yuhang Ma, Bo Cheng, Dawei Leng, and Yuhui Yin. U- stydit: Ultra-high quality artistic style transfer using diffu- sion transformers. arXiv preprint arXiv:2503.08157, 2025. 2

  80. [92]

    Spast: Arbitrary style trans- fer with style priors via pre-trained large-scale model.Neural Networks, page 107556, 2025

    Zhanjie Zhang, Quanwei Zhang, Junsheng Luan, Mengyuan Yang, Yun Wang, and Lei Zhao. Spast: Arbitrary style trans- fer with style priors via pre-trained large-scale model.Neural Networks, page 107556, 2025. 2

  81. [93]

    Vectorsketcher: Learning to create a vector-based free-hand sketch

    Zhanjie Zhang, Quanwei Zhang, Junsheng Luan, Mengyuan Yang, Yun Wang, and Lei Zhao. Vectorsketcher: Learning to create a vector-based free-hand sketch. Engineering Appli- cations of Artificial Intelligence, 156:111005, 2025. 2

  82. [94]

    Image generation from layout

    Bo Zhao, Lili Meng, Weidong Yin, and Leonid Sigal. Image generation from layout. In Proceedings of the IEEE/CVF conference on computer vision and pattern recognition , pages 8584–8593, 2019. 9

  83. [95]

    Uni-controlnet: All-in-one control to text-to-image diffusion models

    Shihao Zhao, Dongdong Chen, Yen-Chun Chen, Jianmin Bao, Shaozhe Hao, Lu Yuan, and Kwan-Yee K Wong. Uni-controlnet: All-in-one control to text-to-image diffusion models. Advances in Neural Information Processing Sys- tems, 36:11127–11150, 2023. 2

  84. [96]

    Layoutdiffusion: Controllable diffu- sion model for layout-to-image generation

    Guangcong Zheng, Xianpan Zhou, Xuewei Li, Zhongang Qi, Ying Shan, and Xi Li. Layoutdiffusion: Controllable diffu- sion model for layout-to-image generation. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pat- tern Recognition, pages 22490–22499, 2023. 9

  85. [97]

    Enhanc- ing detail preservation for customized text-to-image gen- eration: A regularization-free approach

    Yufan Zhou, Ruiyi Zhang, Tong Sun, and Jinhui Xu. Enhanc- ing detail preservation for customized text-to-image gen- eration: A regularization-free approach. arXiv preprint arXiv:2305.13579, 2023. 9

  86. [98]

    Storydiffusion: Consistent self- attention for long-range image and video generation

    Yupeng Zhou, Daquan Zhou, Ming-Ming Cheng, Jiashi Feng, and Qibin Hou. Storydiffusion: Consistent self- attention for long-range image and video generation. Ad- vances in Neural Information Processing Systems , 37: 110315–110340, 2025. 2, 6, 7, 8, 9

  87. [99]

    Storymaker: Towards holistic consistent characters in text-to-image generation

    Zhengguang Zhou, Jing Li, Huaxia Li, Nemo Chen, and Xu Tang. Storymaker: Towards holistic consistent characters in text-to-image generation. arXiv preprint arXiv:2409.12576,

  88. [2023]

    Accessed: March 6, 2025. 7

Pith tools

Reviewed August 15, 2026 · model on record in the stance chip above.