Pith. sign in

REVIEW 4 major objections 5 minor 297 references

Deep Neural Networks Inspired by Differential Equations

T0 review · 4 major / 5 minor · reviewed 2026-08-04 · deepseek-v4-flash

Pith's one-line read This review claims differential equations are the unifying design principle behind modern deep learning, from ResNets to diffusion models.

desk verdict Survey of DE-inspired networks with a useful taxonomy but unreliable benchmark tables that need correction before it can serve as a reference. read the letter →

arxiv 2510.09685 v2 pith:PTQXEN66 submitted 2025-10-09 cs.LG cs.AIcs.CVcs.NAmath.NA

classification cs.LGcs.AIcs.CVcs.NAmath.NA MSC 68T0760H1065L06
keywords deeplearningdifferentialequationsneuralODEsstochasticresidualnetworksflowmatchingdiffusionmodelssurvey
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

The paper argues that ordinary and stochastic differential equations provide a unifying language for deep learning: feed-forward networks can be read as numerical discretizations of ODEs and PDEs, regularization tricks such as dropout and noise injection can be derived as SDE discretizations, and modern generative models—from Neural ODEs to flow matching to diffusion models—are best understood as estimating the drift and diffusion terms of a continuous dynamical system. The authors organize dozens of architectures into a taxonomy based on equation type (first-order, higher-order, PDE, SDE) and dynamics-modeling paradigm, and they compile benchmark tables on CIFAR-10 and CIFAR-100 to compare the families. If the survey is accurate, it gives new researchers a map of the field and a quantitative starting point for choosing among equation-inspired architectures.

What carries the argument

The central organizing device is the discretization map: a network layer x_{n+1} = x_n + f(x_n) is identified with an Euler step of dx/dt = v(x,t); multistep and Runge–Kutta blocks are identified with higher-order integrators; stochastic layers (dropout, noise injection, stochastic depth) are identified with Euler–Maruyama steps of an SDE; and generative models are identified with the Liouville equation (for deterministic flow) or Fokker–Planck equation (for diffusion). This mapping lets numerical-analysis knowledge (stability, order of convergence) be imported into network design, and lets stochastic-calculus tools (reverse SDEs, score functions) be imported into regularization and sampling

What would settle it

A concrete check: verify the two CIFAR-100 entries for ResNet and FitResNet. The text says ResNet reaches 75.46% and FitResNet 76.63%, but Table 5 lists 72.24% and 72.34%; also, the text cites PDE CIFAR-100 values 80.55% and 80.34% that do not appear in Table 5. If the discrepancy cannot be resolved, the survey's quantitative comparisons are unreliable.

Watch

Extended reading notes

Core claim

The paper's central claim is that differential equations are not merely an analogy for deep learning but a design principle: every residual network, stochastic regularizer, and diffusion sampler can be placed on a spectrum whose endpoints are deterministic ODEs and stochastic SDEs. Concretely, the authors assert that ResNets are forward-Euler discretizations of an ODE; higher-order networks (LM-ResNets, RKCNNs) correspond to multistep or Runge–Kutta integrators; PDE-based CNNs encode spatial derivatives; dropout, Gaussian noise injection, and stochastic depth have SDE limits; and generative models solve either the Liouville equation (flow matching) or the Fokker–Planck equation (diffusion).

Load-bearing premise

The survey's comparative value rests on the numbers in Tables 5 and 6 faithfully transcribing the cited papers; if the transcriptions are inconsistent or wrong, the taxonomy still stands but the quantitative story does not.

Editorial extensions

If this is right

  • If the correspondence is taken seriously, stability of a network can be analyzed by checking Lipschitz or contraction conditions on the vector field, promising mitigation of vanishing or exploding gradients.
  • Higher-order numerical schemes should yield networks that train more accurately; the tables support this trend, with RKCNN outperforming plain ResNet on both CIFAR-10 and CIFAR-100.
  • The same SDE formalism that explains dropout can justify new regularization schemes by choosing different noise processes, extending beyond Gaussian assumptions to jump or Lévy noise.
  • Diffusion models and flow matching become two ends of one continuum, so solvers developed for one transfer to the other; probability-flow ODEs and DDIM already demonstrate such a transfer.
  • Generative models can trade stochastic sampling for deterministic one-step generation, as Rectified Flow and MeanFlow show, reaching competitive FID with a single function evaluation.
  • The Fokker–Planck framing suggests that the noise schedule in diffusion is the diffusion coefficient of an SDE, so optimizing the schedule is equivalent to choosing the SDE's diffusion term in a carefully designed way.

Reading between the lines

Editorial extensions of the paper, not claims the author makes directly.

  • The discretization map suggests a design recipe: any new numerical integrator for ODEs or SDEs is a candidate neural architecture, and any stochastic regularizer is a candidate SDE discretization—this predicts that the field will continue importing numerical-analysis results faster than inventing new network blocks.
  • The benchmark tables, if accurate, imply that the performance gap between plain ResNets and PDE or higher-order networks is partly a statement about integration order; a fair reader could test whether the same gap holds on larger datasets or with modern training schedules, which the survey does not report.
  • Because the survey groups stochastic regularizers with SDEs, it implicitly predicts that regularization strength and diffusion coefficient are the same dial; this could be tested by measuring generalization as a function of an annealing schedule in both settings.
  • The survey's taxonomy could serve as a generative catalog: every combination of equation type (ODE/PDE/SDE), discretization scheme (Euler/RK/multistep), and stochastic component (additive/multiplicative noise) corresponds to a potential new architecture, and the paper's tables mark which cells have already been explored.
Share X Bluesky LinkedIn Reddit HN

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, and a circularity audit.

Referee Report

4 major / 5 minor

Summary. The paper is a survey that organizes deep-learning architectures, regularization techniques, and generative models around ODE/PDE/SDE interpretations. It proposes a taxonomy, gives mathematical formulations in a series of tables, and provides numerical comparisons (classification accuracy on CIFAR-10/100 and FID/NFE on CIFAR-10) drawn from the literature. The abstract's central claim is that the paper provides an 'extensive review' and 'numerical comparisons of these models to illustrate their characteristics and performance.' Since the paper is purely a review, the value of the manuscript rests entirely on the accuracy of its taxonomy, its transcriptions of external results, and the internal consistency of its narrative.

Significance. If the taxonomy and the numerical tables were faithful to the cited sources, the survey could be a useful entry point for researchers seeking a bird's-eye view of differential-equation-inspired deep learning. The paper covers a broad set of methods, includes code links, and draws useful distinctions (e.g., first-order vs. higher-order ODE discretizations, forward vs. reverse SDE design). It also makes no original derivation or experimental claim, which means the potential contribution is reference value rather than new results. That reference value is currently undermined by the data integrity and citation problems detailed below; if these are corrected, the manuscript could serve as a useful, if not comprehensive, survey.

major comments (4)
  1. [§5.1 and Table 5] The numerical comparisons in the text contradict Table 5. In §5.1, the text states that ResNets achieve 75.46% on CIFAR-100 and FitResNet 76.63%, but Table 5 lists 72.24% and 72.34% for the same rows. The text also claims that Parabolic and Hyperbolic PDE models achieve 80.55% and 80.34% on CIFAR-100, yet Table 5 has no CIFAR-100 entries for the Parabolic PDE/Hyperbolic PDE rows; those rows report only CIFAR-10 values. This is not a formatting slip: the text uses these numbers to argue that higher-order and PDE-guided models outperform ResNets, so the inconsistency leaves the central comparison unsupported.
  2. [Table 5 (RevNets and FractalNet rows)] Table 5 contains implausible and mismatched entries. The RevNets row reports 86.31% CIFAR-100 accuracy with 1.79M parameters; the cited RevNets paper (Gomez et al., NIPS 2017) does not report such a number, and the value is far above any known RevNets result. The FractalNet row cites reference [34] (Cho et al., AAAI 2024), but the text and the reference list identify FractalNet as [133] (Larsson et al., ICLR 2017); [34] is a different paper. These are not isolated typos but symptoms that Table 5 has not been systematically verified against its sources. Since Table 5 is the empirical backbone of §5.1, this is a load-bearing reliability issue.
  3. [§5.2 and Table 6] The generative-model comparison in Table 6 conflates incomparable settings without any caveat. For example, flow matching is listed as 'NIPs 2019', but the reference [147] is actually an ICLR 2023 paper. More importantly, the FID/NFE numbers come from different papers using different model sizes, training budgets, and evaluation protocols (e.g., EDM reports FID 1.97 with 55.7M params, DPM-Solver reports 2.69 with 35.7M params, and several rows have no param count). The text's claim of a 'clear progression' from high-step stochastic sampling to single-step deterministic optimization is therefore not supported by a controlled comparison; it is a compilation of disparate results. The survey should either include a methodological note on how numbers were collected and whether they are directly comparable, or explicitly warn the reader otherwise.
  4. [General (methodology of the numerical review)] The paper does not state how the numbers in Tables 5 and 6 were selected, transcribed, or verified. For a survey whose abstract advertises 'numerical comparisons', this is more than an omission: the reader has no way to decide which of the conflicting values (text vs. table, or table vs. original paper) is trustworthy. Given the documented discrepancies, the current manuscript cannot serve as a reliable reference. A revision would need to include a data-verification protocol and, ideally, a table of the exact source (section/table/line) for each reported number.
minor comments (5)
  1. [Table 6] Typo 'Recitied Flow' should be 'Rectified Flow'; also use 'NeurIPS' consistently instead of 'NIPs'.
  2. [§2.2.3] The acronym 'ODESs' is unusual and likely should be 'ODEs' or a spelled-out 'systems of ODEs'.
  3. [Figures 1 and 2] The cross-references use Roman numerals ('Sec Ⅲ.A', 'Sec Ⅳ.A'); these should be Arabic ('Sec. 3.1', 'Sec. 4.1').
  4. [Title page/footer] The ACM Reference Format block and the copyright line indicate 2018, while the arXiv submission is dated 2025; this inconsistency should be resolved. Also, the conference placeholder 'Conference acronym ’XX' is unfilled.
  5. [Reference [147]] Year mismatch: reference [147] is an ICLR 2023 paper, but Table 6 lists the same work as 'NIPs 2019'. Similarly, reference [34] is not the FractalNet paper, creating a citation conflict with Table 5.

Circularity Check

0 steps flagged · score 0.0 of 10

No circularity: survey with no derived predictions; internal table/text inconsistencies are accuracy issues, not circular reasoning.

full rationale

The paper is a review/taxonomy of DDE- and SDE-inspired neural network architectures and generative models. It makes no derivation chain, no fitted parameters called predictions, and no theoretical claim that is justified by its own prior results. The central claim is descriptive: that the paper provides a structured review and numerical comparisons. Those comparisons in Section 5 (Tables 5 and 6) are transcriptions of externally published results (e.g., ResNets [90], LM-ResNets [164], PDE-CNNs [211], DDPMs [97], DEIS [285]), not fitted inputs renamed as predictions. The main risk identified by the reader is an internal inconsistency between the Section 5.1 text (ResNet 75.46% CIFAR-100, FitResNet 76.63%, PDE CIFAR-100 80.55/80.34) and Table 5 (72.24/72.34, and PDE CIFAR-100 64.8/64.9/65.4). That is a correctness/transcription concern, not a circularity concern: the survey does not use its tables to justify itself. No self-citation is load-bearing, no uniqueness theorem is imported, and no ansatz is smuggled in via citation. The paper is not circular; it is a survey whose value depends on the accuracy of its external benchmarks, which is outside the scope of circularity analysis.

Assumptions & free parameters 0 free parameters · 3 assumptions · 0 invented entities

As a review, the paper introduces no free parameters or invented entities. The central content restates prior work; the only assumptions are the accuracy of the compiled numbers and the validity of the taxonomy.

assumptions (3)
  • domain assumption The cited papers' reported accuracies and FID scores are correctly transcribed.
    The survey's comparative tables (Tables 5 and 6) rely entirely on numbers from external papers; if any transcription is inaccurate, the comparison is invalid. Internal discrepancies suggest this assumption may be violated.
  • domain assumption The taxonomy categories (first-order ODEs, higher-order ODEs, DE systems, PDEs, SDE regularization, etc.) are well-defined and collectively exhaustive.
    The paper's organization assumes these categories cleanly partition the field; if the categories overlap or omit important work, the review's framework is misleading.
  • standard math Standard ODE/SDE theory (existence, uniqueness, Fokker-Planck, reverse-time SDE, etc.) is correct.
    Background preliminaries in Section 2 rely on standard results; no novel mathematical claims are made.

how reviews work

0 comments
Cite this review

Pith. "Pith review of Deep Neural Networks Inspired by Differential Equations." pith.science (2026). https://pith.science/paper/PTQXEN66

@misc{pith2026251009685,
  author       = {Pith},
  title        = {Pith review of: Deep Neural Networks Inspired by Differential Equations},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/PTQXEN66}},
  note         = {Machine review of arXiv:2510.09685}
}
read the original abstract

Deep learning has become a pivotal technology in fields such as computer vision, scientific computing, and dynamical systems, significantly advancing these disciplines. However, neural Networks persistently face challenges related to theoretical understanding, interpretability, and generalization. To address these issues, researchers are increasingly adopting a differential equations perspective to propose a unified theoretical framework and systematic design methodologies for neural networks. In this paper, we provide an extensive review of deep neural network architectures and dynamic modeling methods inspired by differential equations. We specifically examine deep neural network models and deterministic dynamical network constructs based on ordinary differential equations (ODEs), as well as regularization techniques and stochastic dynamical network models informed by stochastic differential equations (SDEs). We present numerical comparisons of these models to illustrate their characteristics and performance. Finally, we explore promising research directions in integrating differential equations with deep learning to offer new insights for developing intelligent computational methods that boast enhanced interpretability and generalization capabilities.

Figures

Figures reproduced from arXiv: 2510.09685 by the authors.

Figure 1
Figure 1. Overview of DDE-driven neural network architectures (left, Sec. [PITH_FULL_IMAGE:figures/full_fig_p006_1.png] view at source ↗
Figure 2
Figure 2. Overview of ODE-guided deterministic dynamics (left, Sec. [PITH_FULL_IMAGE:figures/full_fig_p007_2.png] view at source ↗
Figure 3
Figure 3. Schematics of representative network architectures inspired by discretizations of DEs. The diagrams illustrate the computational [PITH_FULL_IMAGE:figures/full_fig_p009_3.png] view at source ↗

Discussion (0). Continue with ORCID to comment.

Reference graph

Works this paper leans on

297 extracted references · 4 canonical work pages

  1. [34]

    Woojin Cho, Seunghyeon Cho, Hyundong Jin, Jinsung Jeon, Kookjin Lee, Sanghyun Hong, Dongeun Lee, Jonghyun Choi, and Noseong Park. 2024. Operator-learning-inspired modeling of neural ordinary differential equations. InProceedings of the AAAI Conference on Artificial Intelligence, Vol. 38. 11543–11551

  2. [289]

    Yichi Zhang, Yici Yan, Alex Schwing, and Zhizhen Zhao. 2025. Towards Hierarchical Rectified Flow. InInternational Conference on Learning Representations

  3. [295]

    Yuanzhi Zhu, Xingchao Liu, and Qiang Liu. 2024. SlimFlow: Training Smaller One-Step Diffusion Models with Rectified Flow. InEuropean Conference on Computer Vision. 342–359

  4. [133]

    Gustav Larsson, Michael Maire, and Gregory Shakhnarovich. 2017. FractalNet: Ultra-Deep Neural Networks without Residuals. InInternational Conference on Learning Representations

  5. [147]

    Yaron Lipman, Ricky TQ Chen, Heli Ben-Hamu, Maximilian Nickel, and Matt Le. 2023. Flow matching for generative modeling. InInternational Conference on Learning Representations

  6. [1]

    Abien Fred Agarap. 2018. Deep learning using rectified linear units (relu).arXiv preprint arXiv:1803.08375(2018)

  7. [2]

    Ravi Aggarwal, Viknesh Sounderajah, Guy Martin, Daniel SW Ting, Alan Karthikesalingam, Dominic King, Hutan Ashrafian, and Ara Darzi. 2021. Diagnostic accuracy of deep learning in medical imaging: a systematic review and meta-analysis.NPJ digital medicine4, 1 (2021), 65

  8. [3]

    Michael Samuel Albergo and Eric Vanden-Eijnden. 2023. Building Normalizing Flows with Stochastic Interpolants. InInternational Conference on Learning Representations

Show all 297 references
  1. [4]

    Tobias Alt, Karl Schrader, Matthias Augustin, Pascal Peter, and Joachim Weickert. 2023. Connections between numerical algorithms for PDEs and neural networks.Journal of Mathematical Imaging and Vision65, 1 (2023), 185–208

  2. [5]

    Brian DO Anderson. 1982. Reverse-time diffusion equation models.Stochastic Processes and their Applications12, 3 (1982), 313–326

  3. [6]

    Ludwig Arnold. 1974. Stochastic differential equations.New York2 (1974), 2

  4. [7]

    Jimmy Ba and Brendan Frey. 2013. Adaptive dropout for training deep neural networks. InAdvances in Neural Information Processing Systems, Vol. 26

  5. [8]

    Bakary Badjie, Jose Cecilio, and Antonio Casimiro. 2024. Adversarial attacks and countermeasures on image classification-based deep learning models in autonomous driving systems: A systematic review.Comput. Surveys57, 1 (2024), 1–52

  6. [9]

    Pierre Baldi and Peter J Sadowski. 2013. Understanding dropout. InAdvances in Neural Information Processing Systems, Vol. 26

  7. [10]

    Chayan Banerjee, Kien Nguyen, Clinton Fookes, and Karniadakis George. 2024. Physics-Informed Computer Vision: A Review and Perspectives. Comput. Surveys57, 1, Article 17 (Oct. 2024), 38 pages. doi:10.1145/3689037

  8. [11]

    Georgios Batzolis, Jan Stanczuk, Carola-Bibiane Schönlieb, and Christian Etmann. 2021. Conditional image generation with score-based diffusion models.arXiv preprint arXiv:2111.13606(2021)

  9. [12]

    Jens Behrmann, Will Grathwohl, Ricky TQ Chen, David Duvenaud, and Jörn-Henrik Jacobsen. 2019. Invertible residual networks. InInternational Conference on Machine Learning, Vol. 97. 573–582

  10. [13]

    Mikhail Belkin, Daniel Hsu, Siyuan Ma, and Soumik Mandal. 2019. Reconciling modern machine-learning practice and the classical bias–variance trade-off.Proceedings of the National Academy of Sciences116, 32 (2019), 15849–15854. arXiv:https://www.pnas.org/doi/pdf/10.1073/pnas.19...

  11. [14]

    Marin Biloš, Johanna Sommer, Syama Sundar Rangapuram, Tim Januschowski, and Stephan Günnemann. 2021. Neural flows: Efficient alternative to neural ODEs. InAdvances in Neural Information Processing Systems, Vol. 34. 21325–21337

  12. [15]

    Olivier Bousquet and Andre Elisseeff. 2002. Stability and generalization.Journal of Machine Learning Research2 (2002), 499–526

  13. [16]

    1983.Differential equations and their applications

    Martin Braun and Martin Golubitsky. 1983.Differential equations and their applications. Vol. 4. Springer

  14. [17]

    2016.Numerical methods for ordinary differential equations

    John Charles Butcher. 2016.Numerical methods for ordinary differential equations. John Wiley & Sons

  15. [18]

    Lei Cai, Jingyang Gao, and Di Zhao. 2020. A review of the application of deep learning in medical image classification and segmentation.Annals of Translational Medicine8, 11 (2020), 713

  16. [19]

    Andrew Campbell, Joe Benton, Valentin De Bortoli, Tom Rainforth, George Deligiannidis, and Arnaud Doucet. 2022. A continuous time framework for discrete denoising models. InAdvances in Neural Information Processing Systems. 28266–28279

  17. [20]

    2021.Understanding gaussian noise injections in neural networks

    Alexander Camuto. 2021.Understanding gaussian noise injections in neural networks. Ph. D. Dissertation. University of Oxford

  18. [21]

    Alexander Camuto, Matthew Willetts, Umut Simsekli, Stephen J Roberts, and Chris C Holmes. 2020. Explicit regularisation in gaussian noise injections. InAdvances in Neural Information Processing Systems, Vol. 33. 16603–16614

  19. [22]

    Fei Cao, Kimball Johnston, Thomas Laurent, Justin Le, and Sébastien Motsch. 2025. Generative diffusion models from a PDE perspective.arXiv preprint arXiv:2501.17054(2025)

  20. [23]

    Hanqun Cao, Cheng Tan, Zhangyang Gao, Yilun Xu, Guangyong Chen, Pheng-Ann Heng, and Stan Z Li. 2024. A survey on generative diffusion models.IEEE Transactions on Knowledge and Data Engineering(2024)

  21. [24]

    Yu Cao, Jingrun Chen, Yixin Luo, and Xiang Zhou. 2023. Exploring the optimal choice for generative processes in diffusion models: Ordinary vs stochastic differential equations. InAdvances in Neural Information Processing Systems, Vol. 36. 33420–33468

  22. [25]

    Junyi Chai, Hao Zeng, Anming Li, and Eric W.T. Ngai. 2021. Deep learning in computer vision: A critical review of emerging techniques and application scenarios.Machine Learning with Applications6 (2021), 100134. doi:10.1016/j.mlwa.2021.100134

  23. [26]

    Bo Chang, Lili Meng, Eldad Haber, Lars Ruthotto, David Begert, and Elliot Holtham. 2018. Reversible architectures for arbitrarily deep residual neural networks. InProceedings of the AAAI conference on artificial intelligence, Vol. 32

  24. [27]

    Chunlei Chen, Peng Zhang, Huixiang Zhang, Jiangyan Dai, Yugen Yi, Huihui Zhang, and Yonghui Zhang. 2020. Deep Learning on Computational-Resource-Limited Platforms: A Survey.Mobile Information Systems2020, 1 (2020), 8454327. arXiv:https://onlinelibrary.wiley.com/doi/pdf/10.1155...

  25. [28]

    Ricky T. Q. Chen and Yaron Lipman. 2024. Flow Matching on General Geometries. InInternational Conference on Learning Representations

  26. [29]

    Ricky T. Q. Chen, Yulia Rubanova, Jesse Bettencourt, and David K Duvenaud. 2018. Neural Ordinary Differential Equations. InAdvances in Neural Information Processing Systems, Vol. 31

  27. [30]

    YangQuan Chen, Ivo Petras, and Dingyu Xue. 2009. Fractional order control-a tutorial. In2009 American control conference. IEEE, 1397–1411

  28. [31]

    Chaoran Cheng, Jiahan Li, Jiajun Fan, and Ge Liu. 2025. 𝛼-Flow: A Unified Framework for Continuous-State Discrete Flow Matching Models. arXiv preprint arXiv:2504.10283(2025). Manuscript submitted to ACM Deep Neural Networks Inspired by Differential Equations 27

  29. [32]

    Chaoran Cheng, Jiahan Li, Jian Peng, and Ge Liu. 2024. Categorical flow matching on statistical manifolds. InAdvances in Neural Information Processing Systems, Vol. 37. 54787–54819

  30. [33]

    Jer-Shiou Chiou and Yen-Hsien Lee. 2009. Jump dynamics and volatility: Oil and the stock markets.Energy34, 6 (2009), 788–796. doi:10.1016/j. energy.2009.02.011

  31. [35]

    Krzysztof M Choromanski, Jared Quincy Davis, Valerii Likhosherstov, Xingyou Song, Jean-Jacques Slotine, Jacob Varley, Honglak Lee, Adrian Weller, and Vikas Sindhwani. 2020. Ode to an ODE. InAdvances in Neural Information Processing Systems, Vol. 33. 3338–3350

  32. [36]

    Cecılia Coelho, M Fernanda P Costa, and Luis L Ferrás. 2025. Neural fractional differential equations.Applied Mathematical Modelling144 (2025), 116060

  33. [37]

    Antoine Collas, Ce Ju, Nicolas Salvy, and Bertrand Thirion. 2025. Riemannian Flow Matching for Brain Connectivity Matrices via Pullback Geometry.arXiv preprint arXiv:2505.18193(2025)

  34. [38]

    Ronan Collobert and Jason Weston. 2008. A unified architecture for natural language processing: deep neural networks with multitask learning. In International Conference on Machine Learning(Helsinki, Finland). 160–167

  35. [39]

    Florinel-Alin Croitoru, Vlad Hondru, Radu Tudor Ionescu, and Mubarak Shah. 2023. Diffusion models in vision: A survey.IEEE Transactions on Pattern Analysis and Machine Intelligence45, 9 (2023), 10850–10869

  36. [40]

    Dominik Csiba and Peter Richtárik. 2018. Importance sampling for minibatches.Journal of Machine Learning Research19, 27 (2018), 1–21

  37. [41]

    Qinpeng Cui, Xinyi Zhang, Qiqi Bao, and Qingmin Liao. 2025. Elucidating the solution space of extended reverse-time SDE for diffusion models. In 2025 IEEE/CVF Winter Conference on Applications of Computer Vision (W ACV). 243–252

  38. [42]

    Wenjun Cui, Qiyu Kang, Xuhao Li, Kai Zhao, Wee Peng Tay, Weihua Deng, and Yidong Li. 2025. Neural Variable-Order Fractional Differential Equation Networks. InProceedings of the AAAI Conference on Artificial Intelligence, Vol. 39. 16109–16117

  39. [43]

    Wenjun Cui, Honglei Zhang, Haoyu Chu, Pipi Hu, and Yidong Li. 2023. On robustness of neural ODEs image classifiers.Information Sciences632 (2023), 576–593

  40. [44]

    Daems, Rembert and Opper, Manfred and Crevecoeur, Guillaume and Birdal, Tolga. 2025. Efficient training of neural SDEs using stochastic optimal control. InESANN 2025 : 33rd European Symposium on Artificial Neural Networks, Computational Intelligence and Machine Learning, Proce...

  41. [45]

    Quan Dao, Hao Phung, Binh Nguyen, and Anh Tran. 2023. Flow matching in latent space.arXiv preprint arXiv:2307.08698(2023)

  42. [46]

    Kalyan Das, Jiming Jiang, and JNK Rao. 2004. Mean squared error of empirical predictor.The Annals of Statistics32, 2 (2004), 818–840

  43. [47]

    Oscar Davis, Samuel Kessler, Mircea Petrache, Ismail Ceylan, Michael Bronstein, and Joey Bose. 2024. Fisher flow matching for generative modeling over discrete data. InAdvances in Neural Information Processing Systems, Vol. 37. 139054–139084

  44. [48]

    Arturo De Marinis, Nicola Guglielmi, Stefano Sicilia, and Francesco Tudisco. 2025. Stability of neural ODEs by a control over the expansivity of their flows.arXiv preprint arXiv:2501.10740(2025)

  45. [49]

    2018.Deep learning in natural language processing

    Li Deng and Yang Liu. 2018.Deep learning in natural language processing. Springer

  46. [50]

    Teo Deveney, Jan Stanczuk, Lisa Kreusser, Chris Budd, and Carola-Bibiane Schönlieb. 2025. Closing the ODE–SDE gap in score-based diffusion models through the Fokker–Planck equation.Philosophical Transactions A383, 2298 (2025), 20240503

  47. [51]

    Omar Dhifallah and Yitong Lu. 2021. On the Inherent Regularization Effects of Noise Injection During Training. InInternational Conference on Machine Learning. 2676–2686

  48. [52]

    1989.Introduction to electric circuits

    Richard C Dorf. 1989.Introduction to electric circuits. John Wiley & Sons

  49. [53]

    Finale Doshi-Velez and Been Kim. 2017. Towards a rigorous science of interpretable machine learning.arXiv preprint arXiv:1702.08608(2017)

  50. [54]

    Weitao Du, He Zhang, Tao Yang, and Yuanqi Du. 2023. A flexible diffusion model. InInternational Conference on Machine Learning. 8678–8696

  51. [55]

    Emilien Dupont, Arnaud Doucet, and Yee Whye Teh. 2019. Augmented neural odes. InAdvances in Neural Information Processing Systems, Vol. 32

  52. [56]

    Michael B Elowitz, Arnold J Levine, Eric D Siggia, and Peter S Swain. 2002. Stochastic gene expression in a single cell.Science297, 5584 (2002), 1183–1186

  53. [57]

    Jonathan Ephrath, Moshe Eliasof, Lars Ruthotto, Eldad Haber, and Eran Treister. 2020. LeanConvNets: low-cost yet effective convolutional neural networks.IEEE Journal of Selected Topics in Signal Processing14, 4 (2020), 894–904

  54. [58]

    2009.Applied delay differential equations

    Thomas Erneux. 2009.Applied delay differential equations. Springer

  55. [59]

    2022.Partial differential equations

    Lawrence C Evans. 2022.Partial differential equations. Vol. 19. American Mathematical Society

  56. [60]

    Angela Fan, Edouard Grave, and Armand Joulin. 2020. Reducing Transformer Depth on Demand with Structured Dropout. InInternational Conference on Learning Representations

  57. [61]

    Chris Finlay, Jörn-Henrik Jacobsen, Levon Nurbekyan, and Adam Oberman. 2020. How to train your neural ODE: the world of Jacobian and kinetic regularization. InInternational Conference on Machine Learning. 3154–3164

  58. [62]

    Kevin Frans, Danijar Hafner, Sergey Levine, and Pieter Abbeel. 2025. One Step Diffusion via Shortcut Models. InInternational Conference on Learning Representations

  59. [63]

    2020.A course on rough paths

    Peter K Friz and Martin Hairer. 2020.A course on rough paths. Springer. Manuscript submitted to ACM 28 Liu et al

  60. [64]

    Yuxiang Fu, Qi Yan, Lele Wang, Ke Li, and Renjie Liao. 2025. Moflow: One-step flow matching for human trajectory forecasting via implicit maximum likelihood estimation based distillation. InProceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition. 17282–17293

  61. [65]

    Yarin Gal and Zoubin Ghahramani. 2016. A theoretically grounded application of dropout in recurrent neural networks. InAdvances in Neural Information Processing Systems, Vol. 29

  62. [66]

    Lucio Galeati and Fabian A Harang. 2022. Regularization of multiplicative SDEs through additive noise.The Annals of Applied Probability32, 5 (2022), 3930–3963

  63. [67]

    1985.Handbook of stochastic methods

    Crispin W Gardiner et al. 1985.Handbook of stochastic methods. Vol. 3. springer Berlin

  64. [68]

    Xavier Gastaldi. 2017. Shake-shake regularization of 3-branch residual networks. InICLR Workshop

  65. [69]

    Itai Gat, Tal Remez, Neta Shaul, Felix Kreuk, Ricky TQ Chen, Gabriel Synnaeve, Yossi Adi, and Yaron Lipman. 2024. Discrete flow matching. In Advances in Neural Information Processing Systems, Vol. 37. 133345–133385

  66. [70]

    Zhengyang Geng, Mingyang Deng, Xingjian Bai, J Zico Kolter, and Kaiming He. 2025. Mean flows for one-step generative modeling.arXiv preprint arXiv:2505.13447(2025)

  67. [71]

    Golnaz Ghiasi, Tsung-Yi Lin, and Quoc V Le. 2018. Dropblock: A regularization method for convolutional networks. InAdvances in Neural Information Processing Systems, Vol. 31

  68. [72]

    Arnab Ghosh, Harkirat Behl, Emilien Dupont, Philip Torr, and Vinay Namboodiri. 2020. Steer: Simple temporal regularization for neural ode. In Advances in Neural Information Processing Systems, Vol. 33. 14831–14843

  69. [73]

    Patryk Gierjatowicz, Marc Sabate-Vidales, David Šiška, Lukasz Szpruch, and Žan Žurič. 2020. Robust pricing and hedging via neural SDEs.arXiv preprint arXiv:2007.04154(2020)

  70. [74]

    Aidan N Gomez, Mengye Ren, Raquel Urtasun, and Roger B Grosse. 2017. The reversible residual network: Backpropagation without storing activations. InAdvances in Neural Information Processing Systems, Vol. 30

  71. [75]

    Aidan N Gomez, Ivan Zhang, Siddhartha Rao Kamalakara, Divyam Madaan, Kevin Swersky, Yarin Gal, and Geoffrey E Hinton. 2019. Learning sparse networks using targeted dropout.arXiv preprint arXiv:1905.13678(2019)

  72. [76]

    Martin Gonzalez, Nelson Fernandez, Thuy Vinh Dinh Tran, Elies Gherbi, Hatem Hajri, and Nader Masmoudi. 2023. SEEDS: Exponential SDE Solvers for Fast High-Quality Sampling from Diffusion Models. InNeural Information Processing Systems

  73. [77]

    2016.Deep learning

    Ian Goodfellow, Yoshua Bengio, and Aaron Courville. 2016.Deep learning. MIT press

  74. [78]

    Ian Goodfellow, Jean Pouget-Abadie, Mehdi Mirza, Bing Xu, David Warde-Farley, Sherjil Ozair, Aaron Courville, and Yoshua Bengio. 2014. Generative adversarial nets. InAdvances in Neural Information Processing Systems. 2672–2680

  75. [79]

    HS Greenside and E Helfand. 1981. Numerical integration of stochastic differential equations—II.Bell System Technical Journal60, 8 (1981), 1927–1940

  76. [80]

    Samuel Greydanus, Misko Dzamba, and Jason Yosinski. 2019. Hamiltonian neural networks. InAdvances in Neural Information Processing Systems, Vol. 32

  77. [81]

    Sorin Grigorescu, Bogdan Trasnea, Tiberiu Cocias, and Gigel Macesanu. 2020. A survey of deep learning techniques for autonomous driving. Journal of Field Robotics37, 3 (2020), 362–386. arXiv:https://onlinelibrary.wiley.com/doi/pdf/10.1002/rob.21918 doi:10.1002/rob.21918

  78. [82]

    Pengsheng Guo and Alex Schwing. 2025. Variational Rectified Flow Matching. InInternational Conference on Machine Learning

  79. [83]

    Yuan Guo, Jingyu Kong, Yu Wang, and Yuping Duan. 2025. Take the Bull by the Horns: Learning to Segment Hard Samples. InProceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition. 15642–15652

  80. [84]

    Abhinav Gupta and Pierre FJ Lermusiaux. 2023. Generalized neural closure models with interpretability.Scientific Reports13, 1 (2023), 10634

  81. [85]

    Eldad Haber and Lars Ruthotto. 2017. Stable architectures for deep neural networks.Inverse problems34, 1 (2017), 014004

  82. [86]

    Jun Han and Claudio Moraga. 1995. The influence of the sigmoid function parameters on the speed of backpropagation learning. InFrom Natural to Artificial Neural Computation. 195–201

  83. [87]

    Sicong Han, Chenhao Lin, Chao Shen, Qian Wang, and Xiaohong Guan. 2023. Interpreting Adversarial Examples in Deep Learning: A Review. Comput. Surveys55, 14s, Article 328 (July 2023), 38 pages. doi:10.1145/3594869

  84. [88]

    Soufiane Hayou and Fadhel Ayed. 2021. Regularization in resnet with stochastic depth. InAdvances in Neural Information Processing Systems, Vol. 34. 15464–15474

  85. [89]

    Chunming He, Yuqi Shen, Chengyu Fang, Fengyang Xiao, Longxiang Tang, Yulun Zhang, Wangmeng Zuo, Zhenhua Guo, and Xiu Li. 2025. Diffusion Models in Low-Level Vision: A Survey.IEEE Transactions on Pattern Analysis and Machine Intelligence47, 6 (2025), 4630–4651

  86. [90]

    Kaiming He, Xiangyu Zhang, Shaoqing Ren, and Jian Sun. 2016. Deep residual learning for image recognition. InProceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition. 770–778

  87. [91]

    Xiangyu He, Zitao Mo, Peisong Wang, Yang Liu, Mingyuan Yang, and Jian Cheng. 2019. Ode-inspired network design for single image super- resolution. InProceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition. 1732–1741

  88. [92]

    Zhezhi He, Adnan Siraj Rakin, and Deliang Fan. 2019. Parametric noise injection: Trainable randomness to improve deep neural network robustness against adversarial attack. InProceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition. 588–597

  89. [93]

    Calypso Herrera, Florian Krach, and Josef Teichmann. 2021. Neural Jump Ordinary Differential Equations: Consistent Continuous-Time Prediction and Filtering. InInternational Conference on Learning Representations

  90. [94]

    Steven L Heston. 1993. A closed-form solution for options with stochastic volatility with applications to bond and currency options.The review of financial studies6, 2 (1993), 327–343. Manuscript submitted to ACM Deep Neural Networks Inspired by Differential Equations 29

  91. [95]

    1985.Methods of mathematical physics

    David Hilbert. 1985.Methods of mathematical physics. CUP Archive

  92. [96]

    Geoffrey E Hinton, Nitish Srivastava, Alex Krizhevsky, Ilya Sutskever, and Ruslan R Salakhutdinov. 2012. Improving neural networks by preventing co-adaptation of feature detectors.arXiv preprint arXiv:1207.0580(2012)

  93. [97]

    Jonathan Ho, Ajay Jain, and Pieter Abbeel. 2020. Denoising diffusion probabilistic models. InAdvances in Neural Information Processing Systems, Vol. 33. 6840–6851

  94. [98]

    Sepp Hochreiter and Jürgen Schmidhuber. 1997. Long short-term memory.Neural Computation9, 8 (1997), 1735–1780

  95. [99]

    Peter Holderrieth, Marton Havasi, Jason Yim, Neta Shaul, Itai Gat, Tommi Jaakkola, Brian Karrer, Ricky T. Q. Chen, and Yaron Lipman. 2025. Generator Matching: Generative modeling with arbitrary Markov processes. InInternational Conference on Learning Representations

  96. [100]

    Md Tanzib Hosain, Jamin Rahman Jim, MF Mridha, and Md Mohsin Kabir. 2024. Explainable AI approaches in deep learning: Advancements, applications and challenges.Computers and Electrical Engineering117 (2024), 109246

  97. [101]

    Xiaoyang Hou, Tian Zhu, Milong Ren, Dongbo Bu, Xin Gao, Chunming Zhang, and Shiwei Sun. 2024. Improving Molecular Graph Generation with Flow Matching and Optimal Transport. InNeurIPS 2024 Workshop on AI for New Drug Modalities

  98. [102]

    Xing Hua, Haodong Chen, Qianqian Duan, Danfeng Hong, Ruijiao Li, Huiliang Shang, Linghua Jiang, Haima Yang, and Dawei Zhang. 2025. A Comprehensive Review of Diffusion Models in Smart Agriculture: Progress, Applications, and Challenges.arXiv preprint arXiv:2507.18376(2025)

  99. [103]

    Gao Huang, Zhuang Liu, Laurens Van Der Maaten, and Kilian Q Weinberger. 2017. Densely connected convolutional networks. InProceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition. 4700–4708

  100. [104]

    Gao Huang, Yu Sun, Zhuang Liu, Daniel Sedra, and Kilian Q Weinberger. 2016. Deep networks with stochastic depth. InEuropean Conference on Computer Vision. 646–661

  101. [105]

    Xingchang Huang, Corentin Salaun, Cristina Vasconcelos, Christian Theobalt, Cengiz Oztireli, and Gurprit Singh. 2024. Blue noise for diffusion models. InACM SIGGRAPH 2024 conference papers. 1–11

  102. [106]

    Yichuan Huang. 2025. Model Selection for Diffusion Coefficient Estimation in SDE. (Jan. 2025). https://hal.science/hal-04901917 This paper focuses on the estimation of diffusion coefficients in stochastic differential equations using high-frequency data

  103. [107]

    Sergey Ioffe and Christian Szegedy. 2015. Batch normalization: Accelerating deep network training by reducing internal covariate shift. In International Conference on Machine Learning. 448–456

  104. [108]

    Zacharia Issa, Blanka Horvath, Maud Lemercier, and Cristopher Salvi. 2023. Non-adversarial training of Neural SDEs with signature kernel scores. InAdvances in Neural Information Processing Systems, Vol. 36. 11102–11126

  105. [109]

    Jörn-Henrik Jacobsen, Arnold Smeulders, and Edouard Oyallon. 2018. i-RevNet: Deep Invertible Networks. InInternational Conference on Learning Representations

  106. [110]

    Arnulf Jentzen and Michael Röckner. 2015. A Milstein scheme for SPDEs.Foundations of Computational Mathematics15 (2015), 313–362

  107. [111]

    Junteng Jia and Austin R Benson. 2019. Neural jump stochastic differential equations. InAdvances in Neural Information Processing Systems, Vol. 32

  108. [112]

    Ahmed Joudal, Youness El Moutaouakil, Ahmed El Hassani, and Youness El Moutaouakil. 2022. An Adaptive Drop Method for Deep Neural Networks Regularization.Knowledge-Based Systems241 (2022), 108011

  109. [113]

    Anil Kag and Venkatesh Saligrama. 2022. Condensing cnns with partial differential equations. InProceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition. 610–619

  110. [114]

    Guoliang Kang, Jun Li, and Dacheng Tao. 2017. Shakeout: A new approach to regularized deep neural network training.IEEE Transactions on Pattern Analysis and Machine Intelligence40, 5 (2017), 1245–1258

  111. [115]

    Qiyu Kang, Yang Song, Qinxu Ding, and Wee Peng Tay. 2021. Stable neural ode with lyapunov-stable equilibrium points for defending against adversarial attacks. InAdvances in Neural Information Processing Systems, Vol. 34. 14925–14937

  112. [116]

    Kacper Kapuńniak, Peter Potaptchik, Teodora Reu, Leo Zhang, Alexander Tong, Michael Bronstein, Avishek Joey Bose, and Francesco Di Giovanni

  113. [117]

    Tero Karras, Miika Aittala, Timo Aila, and Samuli Laine. 2022. Elucidating the Design Space of Diffusion-Based Generative Models. InAdvances in Neural Information Processing Systems, Vol. 35. 26565–26577

  114. [118]

    Jacob Kelly, Jesse Bettencourt, Matthew J Johnson, and David K Duvenaud. 2020. Learning differential equations that are easy to solve. InAdvances in Neural Information Processing Systems, Vol. 33. 4370–4380

  115. [119]

    2021.On neural differential equations

    P Kidger. 2021.On neural differential equations. Ph. D. Dissertation. University of Oxford

  116. [120]

    Patrick Kidger, James Foster, Xuechen Li, and Terry J Lyons. 2021. Neural sdes as infinite-dimensional gans. InInternational Conference on Machine Learning, Vol. 139. 5453–5463

  117. [121]

    Patrick Kidger, James Foster, Xuechen Chen Li, and Terry Lyons. 2021. Efficient and accurate gradients for neural sdes. InAdvances in Neural Information Processing Systems, Vol. 34. 18747–18761

  118. [122]

    Patrick Kidger, James Morrill, James Foster, and Terry Lyons. 2020. Neural controlled differential equations for irregular time series. InAdvances in Neural Information Processing Systems, Vol. 33. 6696–6707

  119. [123]

    AA Kilbas and JJ Trujillo. 2001. Differential equations of fractional order: methods results and problem—I.Applicable Analysis78, 1-2 (2001), 153–192

  120. [124]

    Jongseon Kim, Hyungjoon Kim, HyunGi Kim, Dongjun Lee, and Sungroh Yoon. 2025. A comprehensive survey of deep learning for time series forecasting: architectural diversity and open challenges.Artificial Intelligence Review58, 7 (2025), 1–95. Manuscript submitted to ACM 30 Liu et al

  121. [125]

    Durk P Kingma, Tim Salimans, and Max Welling. 2015. Variational dropout and the local reparameterization trick. InAdvances in Neural Information Processing Systems, Vol. 28

  122. [126]

    Lin Kong, Wei Sun, Fanhua Shang, Yuanyuan Liu, and Hongying Liu. 2022. HNO: High-Order Numerical Architecture for ODE-Inspired Deep Unfolding Networks. InProceedings of the AAAI Conference on Artificial Intelligence, Vol. 36. 7220–7228

  123. [127]

    Zhifeng Kong and Wei Ping. 2021. On Fast Sampling of Diffusion Probabilistic Models. InICML Workshop on Invertible Neural Networks, Normalizing Flows, and Explicit Likelihood Models

  124. [128]

    Nikita Kornilov, Petr Mokrov, Alexander Gasnikov, and Aleksandr Korotin. 2024. Optimal flow matching: Learning straight trajectories in just one step. InAdvances in Neural Information Processing Systems, Vol. 37. 104180–104204

  125. [129]

    Alex Krizhevsky, Ilya Sutskever, and Geoffrey E Hinton. 2012. ImageNet classification with deep convolutional neural networks. InAdvances in Neural Information Processing Systems. 1097–1105

  126. [130]

    Anders Krogh and John A Hertz. 1992. A simple weight decay can improve generalization. InAdvances in Neural Information Processing Systems, Vol. 4

  127. [131]

    Siddharth Krishna Kumar. 2017. On weight initialization in deep neural networks.arXiv preprint arXiv:1704.08863(2017)

  128. [132]

    Alex Labach, Hojjat Salehinejad, and Shahrokh Valaee. 2019. Survey of dropout methods for deep neural networks.arXiv preprint arXiv:1904.13310 (2019)

  129. [134]

    Quoc V Le, Jiquan Ngiam, Adam Coates, Abhik Lahiri, Bobby Prochnow, and Andrew Y Ng. 2011. On optimization methods for deep learning. In International Conference on Machine Learning. 265–272

  130. [135]

    Yann LeCun, Yoshua Bengio, and Geoffrey Hinton. 2015. Deep learning.Nature521, 7553 (2015), 436–444

  131. [136]

    Sangyun Lee, Zinan Lin, and Giulia Fanti. 2024. Improving the training of rectified flows. InAdvances in Neural Information Processing Systems, Vol. 37. 63082–63109

  132. [137]

    Mikko Lehtimäki, Lassi Paunonen, and Marja-Leena Linne. 2022. Accelerating neural odes using model order reduction.IEEE Transactions on Neural Networks and Learning Systems35, 1 (2022), 519–531

  133. [138]

    Ao Li, Wei Fang, Hongbo Zhao, Le Lu, Ge Yang, and Minfeng Xu. 2025. MaRS: A Fast Sampler for Mean Reverting Diffusion based on ODE and SDE Solvers.arXiv preprint arXiv:2502.07856(2025)

  134. [139]

    Mufan Li, Mihai Nica, and Dan Roy. 2022. The neural covariance SDE: Shaped infinite depth-and-width networks at initialization. InAdvances in Neural Information Processing Systems, Vol. 35. 10795–10808

  135. [140]

    Ruichuan Li, Qiyou Sun, Xinkai Ding, Yisheng Zhang, Wentao Yuan, and Tong Wu. 2022. Review of Flow-Matching Technology for Hydraulic Systems.Processes10, 12 (2022). doi:10.3390/pr10122482

  136. [141]

    Shufan Li, Konstantinos Kallidromitis, Akash Gokul, Zichun Liao, Yusuke Kato, Kazuki Kozuka, and Aditya Grover. 2025. Omniflow: Any-to-any generation with multi-modal rectified flows. InProceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition. 13178–13188

  137. [142]

    Shan Li and Yiu-Ming Liu. 2016. Whiteout: Gaussian adaptive noise regularization in deep neural networks.arXiv preprint arXiv:1612.01490(2016)

  138. [143]

    Xuechen Li, Ting-Kam Leonard Wong, Ricky TQ Chen, and David Duvenaud. 2020. Scalable gradients for stochastic differential equations. In International Conference on Artificial Intelligence and Statistics, Vol. 108. 3870–3882

  139. [144]

    Zifeng Lian, Xiaojun Jing, Xiaohan Wang, Hai Huang, Youheng Tan, and Yuanhao Cui. 2016. DropConnect regularization method with sparsity constraint for neural networks.Chinese Journal of Electronics25, 1 (2016), 152–158

  140. [145]

    Kang Liao, Zongsheng Yue, Zhouxia Wang, and Chen Change Loy. [n. d.]. Denoising as Adaptation: Noise-Space Domain Adaptation for Image Restoration. InThe Thirteenth International Conference on Learning Representations

  141. [146]

    Alec J Linot, Joshua W Burby, Qi Tang, Prasanna Balaprakash, Michael D Graham, and Romit Maulik. 2023. Stabilized neural ordinary differential equations for long-time forecasting of dynamical systems.J. Comput. Phys.474 (2023), 111838

  142. [148]

    Yaron Lipman, Marton Havasi, Peter Holderrieth, Neta Shaul, Matt Le, Brian Karrer, Ricky TQ Chen, David Lopez-Paz, Heli Ben-Hamu, and Itai Gat. 2024. Flow matching guide and code.arXiv preprint arXiv:2412.06264(2024)

  143. [149]

    Enshu Liu, Xuefei Ning, Yu Wang, and Zinan Lin. 2025. Distilled Decoding 1: One-step Sampling of Image Auto-regressive Models with Flow Matching. InThe Thirteenth International Conference on Learning Representations. https://openreview.net/forum?id=zKlFXV87Pp

  144. [150]

    Guan-Horng Liu, Tianrong Chen, and Evangelos Theodorou. 2021. Second-order neural ode optimizer. InAdvances in Neural Information Processing Systems, Vol. 34. 25267–25279

  145. [151]

    Luping Liu, Yi Ren, Zhijie Lin, and Zhou Zhao. 2022. Pseudo Numerical Methods for Diffusion Models on Manifolds. InInternational Conference on Learning Representations

  146. [152]

    Qiang Liu. 2022. Rectified flow: A marginal preserving approach to optimal transport.arXiv preprint arXiv:2209.14577(2022)

  147. [153]

    Weibo Liu, Zidong Wang, Xiaohui Liu, Nianyin Zeng, Yurong Liu, and Fuad E. Alsaadi. 2017. A survey of deep neural network architectures and their applications.Neurocomputing234 (2017), 11–26. doi:10.1016/j.neucom.2016.12.038

  148. [154]

    Xingchao Liu, Chengyue Gong, and qiang liu. 2023. Flow Straight and Fast: Learning to Generate and Transfer Data with Rectified Flow. In International Conference on Learning Representations. Manuscript submitted to ACM Deep Neural Networks Inspired by Differential Equations 31

  149. [155]

    Xuanqing Liu, Tesi Xiao, Si Si, Qin Cao, Sanjiv Kumar, and Cho-Jui Hsieh. 2019. Neural sde: Stabilizing neural ode networks with stochastic noise. arXiv preprint arXiv:1906.02355(2019)

  150. [156]

    Xuanqing Liu, Tesi Xiao, Si Si, Qin Cao, Sanjiv Kumar, and Cho-Jui Hsieh. 2020. How does noise help robustness? explanation and exploration under the neural sde framework. InProceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition. 282–290

  151. [157]

    Xingchao Liu, Xiwen Zhang, Jianzhu Ma, Jian Peng, et al . 2023. Instaflow: One step is enough for high-quality diffusion-based text-to-image generation. InInternational Conference on Learning Representations

  152. [158]

    Yan Liu, Tao Jiang, Rui Li, Lingling Yuan, Marcin Grzegorzek, Chen Li, and Xiaoyan Li. 2025. A state-of-the-art review of diffusion model applications for microscopic image and micro-alike image analysis.Frontiers in Medicine12 (2025), 1551894

  153. [159]

    Andreas Look, Melih Kandemir, Barbara Rakitsch, and Jan Peters. 2022. A deterministic approximation to neural SDEs.IEEE Transactions on Pattern Analysis and Machine Intelligence45, 4 (2022), 4023–4037

  154. [160]

    Aaron Lou and Stefano Ermon. 2023. Reflected diffusion models. InInternational Conference on Machine Learning. 22675–22701

  155. [161]

    Cheng Lu, Yuhao Zhou, Fan Bao, Jianfei Chen, Chongxuan Li, and Jun Zhu. 2022. Dpm-solver: A fast ode solver for diffusion probabilistic model sampling in around 10 steps. InAdvances in Neural Information Processing Systems, Vol. 35. 5775–5787

  156. [162]

    Cheng Lu, Yuhao Zhou, Fan Bao, Jianfei Chen, Chongxuan Li, and Jun Zhu. 2025. Dpm-solver++: Fast solver for guided sampling of diffusion probabilistic models.Machine Intelligence Research(2025), 1–22

  157. [163]

    Siqi Lu, Fengxu Guan, Hanyu Zhang, and Haitao Lai. 2023. Speed-up ddpm for real-time underwater image enhancement.IEEE Transactions on Circuits and Systems for Video Technology34, 5 (2023), 3576–3588

  158. [164]

    Yiping Lu, Aoxiao Zhong, Quanzheng Li, and Bin Dong. 2018. Beyond finite layer neural networks: Bridging deep architectures and numerical differential equations. InInternational Conference on Machine Learning, Vol. 80. 3276–3285

  159. [165]

    Ziwei Luo, Fredrik K Gustafsson, Zheng Zhao, Jens Sjölund, and Thomas B Schön. 2023. Image restoration with mean-reverting stochastic differential equations. InInternational Conference on Machine Learning. 23045–23066

  160. [166]

    Zhengbo Luo, Zitang Sun, Weilian Zhou, Zizhang Wu, and Sei-ichiro Kamata. 2022. Rethinking ResNets: improved stacking strategies with high-order schemes for image classification.Complex & Intelligent Systems8, 4 (2022), 3395–3407

  161. [167]

    Zhiyuan Ma, Ruixun Liu, Sixian Liu, Jianjun Li, and Bowen Zhou. 2025. Flow Diverse and Efficient: Learning Momentum Flow Matching via Stochastic Velocity Field Sampling.arXiv preprint arXiv:2506.08796(2025)

  162. [168]

    Zhiyuan Ma, Yuzhu Zhang, Guoli Jia, Liangliang Zhao, Yichao Ma, Mingjie Ma, Gaofeng Liu, Kaiyan Zhang, Ning Ding, Jianjun Li, et al. 2025. Efficient diffusion models: A comprehensive survey from principles to practices.IEEE Transactions on Pattern Analysis and Machine Intellig...

  163. [169]

    Shie Mannor, Dori Peleg, and Reuven Rubinstein. 2005. The cross entropy method for classification. InInternational Conference on Machine Learning. 561–568

  164. [170]

    Andre Martins and Ramon Astudillo. 2016. From softmax to sparsemax: A sparse model of attention and multi-label classification. InInternational Conference on Machine Learning. 1614–1623

  165. [171]

    Gisiro Maruyama. 1955. Continuous Markov processes and stochastic equations.Rendiconti del Circolo Matematico di Palermo4 (1955), 48–90

  166. [172]

    Stefano Massaroli, Michael Poli, Michelangelo Bin, Jinkyoo Park, Atsushi Yamashita, and Hajime Asama. 2020. Stable neural flows.arXiv preprint arXiv:2003.08063(2020)

  167. [173]

    Stefano Massaroli, Michael Poli, Jinkyoo Park, Atsushi Yamashita, and Hajime Asama. 2020. Dissecting neural odes. InAdvances in Neural Information Processing Systems, Vol. 33. 3952–3963

  168. [174]

    Sam McCallum and James Foster. 2024. Efficient, accurate and stable gradients for neural odes.arXiv preprint arXiv:2410.11648(2024)

  169. [175]

    Brett A Melbourne and Alan Hastings. 2008. Extinction risk depends strongly on factors contributing to stochasticity.Nature454, 7200 (2008), 100–103

  170. [176]

    2013.Numerical integration of stochastic differential equations

    Grigorii Noikhovich Milstein. 2013.Numerical integration of stochastic differential equations. Vol. 313. Springer Science & Business Media

  171. [177]

    Dmytro Mishkin and Jiri Matas. 2015. All you need is a good init.arXiv preprint arXiv:1511.06422(2015)

  172. [178]

    Reza Moradi, Reza Berangi, and Behrouz Minaei. 2020. A survey of regularization strategies for deep models.Artificial Intelligence Review53, 6 (2020), 3947–3986

  173. [179]

    Pietro Morerio, Jacopo Cavazza, Riccardo Volpi, René Vidal, and Vittorio Murino. 2017. Curriculum dropout. InProceedings of the IEEE/CVF International Conference on Computer Vision. 3544–3552

  174. [180]

    James Morrill, Cristopher Salvi, Patrick Kidger, and James Foster. 2021. Neural rough differential equations for long time series. InInternational Conference on Machine Learning. 7829–7838

  175. [181]

    Moser, Arundhati S

    Brian B. Moser, Arundhati S. Shanbhag, Federico Raue, Stanislav Frolov, Sebastian Palacio, and Andreas Dengel. 2025. Diffusion Models, Image Super-Resolution, and Everything: A Survey.IEEE Transactions on Neural Networks and Learning Systems36, 7 (2025), 11793–11813. doi:10.11...

  176. [182]

    de Albuquerque

    Khan Muhammad, Amin Ullah, Jaime Lloret, Javier Del Ser, and Victor Hugo C. de Albuquerque. 2021. Deep Learning for Safe Autonomous Driving: Current Challenges and Future Directions.IEEE Transactions on Intelligent Transportation Systems22, 7 (2021), 4316–4336. doi:10.1109/ TI...

  177. [183]

    2007.Mathematical biology: I

    James D Murray. 2007.Mathematical biology: I. An introduction. Vol. 17. Springer Science & Business Media. Manuscript submitted to ACM 32 Liu et al

  178. [184]

    Meenal V Narkhede, Prashant P Bartakke, and Mukul S Sutaone. 2022. A review on weight initialization strategies for neural networks.Artificial Intelligence Review55, 1 (2022), 291–322

  179. [185]

    Ho Huu Nghia Nguyen, Tan Nguyen, Huyen Vo, Stanley Osher, and Thieu Vo. 2022. Improving neural ordinary differential equations with nesterov’s accelerated gradient method. InAdvances in Neural Information Processing Systems, Vol. 35. 7712–7726

  180. [186]

    Shen Nie, Hanzhong Allan Guo, Cheng Lu, Yuhao Zhou, Chenyu Zheng, and Chongxuan Li. 2024. The blessing of randomness: Sde beats ode in general diffusion-based image editing. InInternational Conference on Learning Representations

  181. [187]

    Hao Niu, Yuxiang Zhou, Xiaohao Yan, Jun Wu, Yuncheng Shen, Zhang Yi, and Junjie Hu. 2024. On the applications of neural ordinary differential equations in medical image analysis.Artificial Intelligence Review57, 9 (2024), 236

  182. [188]

    Hyeonwoo Noh, Tackgeun You, Jonghwan Mun, and Bohyung Han. 2017. Regularizing deep neural networks by noise: Its interpretation and optimization. InAdvances in Neural Information Processing Systems, Vol. 30

  183. [189]

    Alexander Norcliffe, Cristian Bodnar, Ben Day, Jacob Moss, and Pietro Liò. 2021. Neural ODE Processes. InInternational Conference on Learning Representations

  184. [190]

    Alexander Norcliffe, Cristian Bodnar, Ben Day, Nikola Simidjievski, and Pietro Liò. 2020. On second order behaviour in augmented neural odes. In Advances in Neural Information Processing Systems, Vol. 33. 5911–5921

  185. [191]

    YongKyung Oh, Seungsu Kam, Jonghun Lee, Dong-Young Lim, Sungil Kim, and Alex Bui. 2025. Comprehensive review of neural differential equations for time series analysis.arXiv preprint arXiv:2502.09885(2025)

  186. [192]

    YongKyung Oh, Dongyoung Lim, and Sungil Kim. 2024. Stable Neural Stochastic Differential Equations in Analyzing Irregular Time Series Data. In International Conference on Learning Representations

  187. [193]

    2013.Stochastic differential equations: an introduction with applications

    Bernt Oksendal. 2013.Stochastic differential equations: an introduction with applications. Springer Science & Business Media

  188. [194]

    Antonio Orvieto, Anant Raj, Hans Kersting, and Francis Bach. 2023. Explicit regularization in overparametrized models via noise injection. In International Conference on Artificial Intelligence and Statistics. 7265–7287

  189. [195]

    Otter, Julian R

    Daniel W. Otter, Julian R. Medina, and Jugal K. Kalita. 2021. A Survey of the Usages of Deep Learning for Natural Language Processing.IEEE Transactions on Neural Networks and Learning Systems32, 2 (2021), 604–624. doi:10.1109/TNNLS.2020.2979670

  190. [196]

    Ambar Pal, Connor Lane, René Vidal, and Benjamin D Haeffele. 2020. On the regularization properties of structured dropout. InProceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition. 7671–7679

  191. [197]

    Dogyun Park, Sojin Lee, Sihyeon Kim, Taehoon Lee, Youngjoon Hong, and Hyunwoo J. Kim. 2024. Constant Acceleration Flow. InAdvances in Neural Information Processing Systems, Vol. 37. 90030–90060

  192. [198]

    Sung Woo Park, Hyomin Kim, Kyungjae Lee, and Junseok Kwon. 2022. Riemannian neural SDE: learning stochastic representations on manifolds. InAdvances in Neural Information Processing Systems, Vol. 35. 1434–1444

  193. [199]

    Yura Perugachi-Diaz, Jakub Tomczak, and Sandjai Bhulai. 2021. Invertible densenets with concatenated lipswish. InAdvances in Neural Information Processing Systems, Vol. 34. 17246–17257

  194. [200]

    2007.Stochastic partial differential equations with Lévy noise: An evolution equation approach

    Szymon Peszat and Jerzy Zabczyk. 2007.Stochastic partial differential equations with Lévy noise: An evolution equation approach. Vol. 113. Cambridge University Press

  195. [201]

    Tomaso Poggio, Hrushikesh Mhaskar, Lorenzo Rosasco, Brando Miranda, and Qianli Liao. 2017. Why and when can deep-but not shallow-networks avoid the curse of dimensionality: a review.International Journal of Automation and Computing14, 5 (2017), 503–519

  196. [202]

    2018.Mathematical theory of optimal processes

    Lev Semenovich Pontryagin. 2018.Mathematical theory of optimal processes. Routledge

  197. [203]

    Anichur Rahman, Tanoy Debnath, Dipanjali Kundu, Md Saikat Islam Khan, Airin Afroj Aishi, Sadia Sazzad, Mohammad Sayduzzaman, and Shahab S Band. 2024. Machine learning and deep learning-based approach in smart healthcare: Recent advances, applications, challenges and opportunit...

  198. [204]

    Steven J Rennie, Vaibhava Goel, and Samuel Thomas. 2014. Annealed dropout training of deep networks. In2014 IEEE Spoken Language Technology Workshop (SLT). IEEE, 159–164

  199. [205]

    Pierre Harvey Richemond, Sander Dieleman, and Arnaud Doucet. 2023. Categorical SDEs with Simplex Diffusion. InICML 2023 Workshop: Sampling and Optimization in Discrete Space

  200. [206]

    Hannes Risken. 1989. Fokker-planck equation. InThe Fokker-Planck equation: methods of solution and applications. Springer, 63–95

  201. [207]

    1996.Fokker-Planck Equation

    Hannes Risken. 1996.Fokker-Planck Equation. Springer Berlin Heidelberg, Berlin, Heidelberg, 63–95. doi:10.1007/978-3-642-61544-3_4

  202. [208]

    Ivan Dario Jimenez Rodriguez, Aaron Ames, and Yisong Yue. 2022. Lyanet: A lyapunov framework for training neural odes. InInternational Conference on Machine Learning. 18687–18703

  203. [209]

    Yulia Rubanova, Ricky TQ Chen, and David K Duvenaud. 2019. Latent ordinary differential equations for irregularly-sampled time series. In Advances in Neural Information Processing Systems, Vol. 32

  204. [210]

    Cynthia Rudin. 2019. Stop explaining black box machine learning models for high stakes decisions and use interpretable models instead.Nature Machine Intelligence1, 5 (2019), 206–215

  205. [211]

    Lars Ruthotto and Eldad Haber. 2020. Deep neural networks motivated by partial differential equations.Journal of Mathematical Imaging and Vision62, 3 (2020), 352–364

  206. [212]

    Vardan R Sahakyan, Vahagn G Melkonyan, Gor A Gharagyozyan, and Arman S Avetisyan. 2023. Enhancing Image Recognition with Pre-Defined Convolutional Layers Based on PDEs.Programming and Computer Software49, 3 (2023), 192–197. Manuscript submitted to ACM Deep Neural Networks Insp...

  207. [213]

    Imrus Salehin and Dae-Ki Kang. 2023. A review on dropout regularization approaches for deep neural networks within the scholarly domain. Electronics12, 14 (2023), 3106. https://www.mdpi.com/2079-9292/12/14/3106

  208. [214]

    Hojjat Salehinejad and Shahrokh Valaee. 2021. Edropout: Energy-based dropout and pruning of deep neural networks.IEEE Transactions on Neural Networks and Learning Systems33, 10 (2021), 5279–5292

  209. [215]

    Cristopher Salvi, Maud Lemercier, and Andris Gerasimovics. 2022. Neural stochastic pdes: Resolution-invariant learning of continuous spatiotem- poral dynamics. InAdvances in Neural Information Processing Systems, Vol. 35. 1333–1344

  210. [216]

    Johannes Schusterbauer, Ming Gui, Frank Fundel, and Björn Ommer. 2025. Diff2Flow: Training Flow Matching Models via Diffusion Model Alignment. InProceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition. 28347–28357

  211. [217]

    2005.The numerical solution of ordinary and partial differential equations

    Granville Sewell. 2005.The numerical solution of ordinary and partial differential equations. Vol. 75. John Wiley & Sons

  212. [218]

    Dario Shariatian, Umut Simsekli, and Alain Durmus. 2024. Denoising L\’evy Probabilistic Models.arXiv preprint arXiv:2407.18609(2024)

  213. [219]

    Dinggang Shen, Guorong Wu, and Heung-Il Suk. 2017. Deep learning in medical image analysis.Annual review of biomedical engineering19, 1 (2017), 221–248

  214. [220]

    Macheng Shen and Chen Cheng. 2025. Neural SDEs as a Unified Approach to Continuous-Domain Sequence Modeling.arXiv preprint arXiv:2501.18871(2025)

  215. [221]

    Cristina Silvano, Daniele Ielmini, Fabrizio Ferrandi, Leandro Fiorin, Serena Curzel, Luca Benini, Francesco Conti, Angelo Garofalo, Cristian Zambelli, Enrico Calore, Sebastiano Schifano, Maurizio Palesi, Giuseppe Ascia, Davide Patti, Nicola Petra, Davide De Caro, Luciano Lavag...

  216. [222]

    Saurabh Singh, Derek Hoiem, and David Forsyth. 2016. Swapout: Learning an ensemble of deep architectures. InAdvances in Neural Information Processing Systems, Vol. 29

  217. [223]

    Singh, Lipo Wang, Sukrit Gupta, Haveesh Goli, Parasuraman Padmanabhan, and Balázs Gulyás

    Satya P. Singh, Lipo Wang, Sukrit Gupta, Haveesh Goli, Parasuraman Padmanabhan, and Balázs Gulyás. 2020. 3D Deep Learning on Medical Images: A Review.Sensors20, 18 (2020). doi:10.3390/s20185097

  218. [224]

    Bart MN Smets, Jim Portegies, Erik J Bekkers, and Remco Duits. 2023. PDE-based group equivariant convolutional neural networks.Journal of Mathematical Imaging and Vision65, 1 (2023), 209–239

  219. [225]

    Luke Snow and Vikram Krishnamurthy. 2025. Efficient Neural SDE Training using Wiener-Space Cubature.arXiv preprint arXiv:2502.12395(2025)

  220. [226]

    Jascha Sohl-Dickstein, Eric Weiss, Niru Maheswaranathan, and Surya Ganguli. 2015. Deep unsupervised learning using nonequilibrium thermody- namics. InInternational Conference on Machine Learning. 2256–2265

  221. [227]

    Jiaming Song, Chenlin Meng, and Stefano Ermon. 2021. Denoising Diffusion Implicit Models. InInternational Conference on Learning Representations

  222. [228]

    Yang Song and Stefano Ermon. 2019. Generative modeling by estimating gradients of the data distribution. InAdvances in Neural Information Processing Systems, Vol. 32

  223. [229]

    Yang Song and Stefano Ermon. 2020. Improved Techniques for Training Score-Based Generative Models. InAdvances in Neural Information Processing Systems, H. Larochelle, M. Ranzato, R. Hadsell, M.F. Balcan, and H. Lin (Eds.), Vol. 33. Curran Associates, Inc., 12438–12448

  224. [230]

    Yang Song, Jascha Sohl-Dickstein, Diederik P Kingma, Abhishek Kumar, Stefano Ermon, and Ben Poole. 2021. Score-based generative modeling through stochastic differential equations. InInternational Conference on Learning Representations

  225. [231]

    Yang Song, Jascha Sohl-Dickstein, Diederik P Kingma, Abhishek Kumar, Stefano Ermon, and Ben Poole. 2021. Score-Based Generative Modeling through Stochastic Differential Equations. InInternational Conference on Learning Representations

  226. [232]

    Nitish Srivastava, Geoffrey Hinton, Alex Krizhevsky, Ilya Sutskever, and Ruslan Salakhutdinov. 2014. Dropout: a simple way to prevent neural networks from overfitting.The Journal of Machine Learning Research15, 1 (2014), 1929–1958

  227. [233]

    2007.Partial differential equations: An introduction

    Walter A Strauss. 2007.Partial differential equations: An introduction. John Wiley & Sons

  228. [234]

    Qi Sun, Yunzhe Tao, and Qiang Du. 2018. Stochastic training of residual networks: a differential equation viewpoint.arXiv preprint arXiv:1812.00174 (2018)

  229. [235]

    Ruo-Yu Sun. 2020. Optimization for deep learning: An overview.Journal of the Operations Research Society of China8, 2 (2020), 249–294

  230. [236]

    Vivienne Sze, Yu-Hsin Chen, Joel Emer, Amr Suleiman, and Zhengdong Zhang. 2017. Hardware for machine learning: Challenges and opportunities. In2017 IEEE Custom Integrated Circuits Conference (CICC). 1–8. doi:10.1109/CICC.2017.7993626

  231. [237]

    2022.Computer vision: algorithms and applications

    Richard Szeliski. 2022.Computer vision: algorithms and applications. Springer Nature

  232. [238]

    Sasha Targ, Diogo Almeida, and Kevin Lyman. 2016. Resnet in resnet: Generalizing residual architectures.arXiv preprint arXiv:1603.08029(2016)

  233. [239]

    Jonathan Tompson, Ross Goroshin, Arjun Jain, Yann LeCun, and Christoph Bregler. 2015. Efficient object localization using convolutional networks. InProceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition. 648–656

  234. [240]

    Maria Trigka and Elias Dritsas. 2025. A comprehensive survey of deep learning approaches in image processing.Sensors25, 2 (2025), 531

  235. [241]

    Aaron Tuor, Jan Drgona, and Draguna Vrabie. 2019. Constrained Neural Ordinary Differential Equations with Stability Guarantees. InICLR 2020 Workshop on Integration of Deep Neural Models and Differential Equations

  236. [242]

    Belinda Tzen and Maxim Raginsky. 2019. Neural stochastic differential equations: Deep latent gaussian models in the diffusion limit.arXiv preprint arXiv:1905.09883(2019)

  237. [243]

    Anwaar Ulhaq and Naveed Akhtar. 2022. Efficient diffusion models for vision: A survey.arXiv preprint arXiv:2210.09292(2022)

  238. [244]

    Hamit Taner Ünal and Fatih Başçiftçi. 2022. Evolutionary design of neural network architectures: a review of three decades of research.Artificial Intelligence Review55, 3 (2022), 1723–1802. Manuscript submitted to ACM 34 Liu et al

  239. [245]

    Nicolaas G Van Kampen. 1976. Stochastic differential equations.Physics Reports24, 3 (1976), 171–228

  240. [246]

    Jente Vandersanden, Sascha Holl, Xingchang Huang, and Gurprit Singh. 2025. Edge-preserving noise for diffusion models. InICLR 2025 Workshop on Deep Generative Model in Machine Learning: Theory, Principle and Efficacy

  241. [247]

    Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N Gomez, Łukasz Kaiser, and Illia Polosukhin. 2017. Attention is all you need. InAdvances in Neural Information Processing Systems. 5998–6008

  242. [248]

    Andreas Veit, Neil Alldrin, Gal Chechik, Ivan Krasin, Abhinav Gupta, and Serge Belongie. 2017. Learning from noisy large-scale datasets with minimal supervision. InProceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition. 839–847

  243. [249]

    Vikram Voleti, Christopher Pal, and Adam M Oberman. 2022. Score-based Denoising Diffusion with Non-Isotropic Gaussian Noise Models. In NeurIPS 2022 Workshop on Score-Based Methods

  244. [250]

    Athanasios Voulodimos, Nikolaos Doulamis, Anastasios Doulamis, and Eftychios Protopapadakis. 2018. Deep Learning for Computer Vision: A Brief Review.Computational Intelligence and Neuroscience2018, 1 (2018), 7068349. arXiv:https://onlinelibrary.wiley.com/doi/pdf/10.1155/2018/7...

  245. [251]

    Li Wan, Matthew Zeiler, Sixin Zhang, Yann Le Cun, and Rob Fergus. 2013. Regularization of neural networks using dropconnect. InInternational Conference on Machine Learning. 1058–1066

  246. [252]

    Fu-Yun Wang, Ling Yang, Zhaoyang Huang, Mengdi Wang, and Hongsheng Li. 2025. Rectified Diffusion: Straightness Is Not Your Need in Rectified Flow. InInternational Conference on Learning Representations

  247. [253]

    Luran Wang, Chaoran Cheng, Yizhen Liao, Yanru Qu, and Ge Liu. 2025. Training Free Guided Flow-Matching with Optimal Control. InInternational Conference on Learning Representations

  248. [254]

    Shengbo Wang, Jose Blanchet, and Peter Glynn. 2024. An Efficient High-dimensional Gradient Estimator for Stochastic Differential Equations. In Advances in Neural Information Processing Systems, Vol. 37. 88045–88090

  249. [255]

    Sida Wang and Christopher Manning. 2013. Fast dropout training. InInternational Conference on Machine Learning. 118–126

  250. [256]

    Tangjun Wang, Chenglong Bao, and Zuoqiang Shi. 2025. Convection-Diffusion Equation: A Theoretically Certified Framework for Neural Networks. IEEE Transactions on Pattern Analysis and Machine Intelligence47, 5 (2025), 4170–4182. doi:10.1109/TPAMI.2025.3540310

  251. [257]

    Ee Weinan. 2017. A proposal on machine learning via dynamical systems.Communications in Mathematics and Statistics5, 1 (2017), 1–11

  252. [258]

    Wilamowski

    Bogdan M. Wilamowski. 2009. Neural network architectures and learning algorithms.IEEE Industrial Electronics Magazine3, 4 (2009), 56–63. doi:10.1109/MIE.2009.934790

  253. [259]

    Haibing Wu and Xiaodong Gu. 2015. Max-pooling dropout for regularization of convolutional neural networks. InNeural Information Processing: 22nd International Conference, ICONIP 2015, Istanbul, Turkey, November 9-12, 2015, Proceedings, Part I 22. Springer, 46–54

  254. [260]

    Yuchen Wu, Yuxin Chen, and Yuting Wei. 2024. Stochastic runge-kutta methods: Provable acceleration of diffusion models.arXiv preprint arXiv:2410.04760(2024)

  255. [261]

    Hedi Xia, Vai Suliafu, Hangjie Ji, Tan Nguyen, Andrea Bertozzi, Stanley Osher, and Bao Wang. 2021. Heavy ball neural ordinary differential equations. InAdvances in Neural Information Processing Systems, Vol. 34. 18646–18659

  256. [262]

    Saining Xie, Ross Girshick, Piotr Dollár, Zhuowen Tu, and Kaiming He. 2017. Aggregated residual transformations for deep neural networks. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition. 1492–1500

  257. [263]

    2010.Differential equations for engineers

    Wei-Chau Xie. 2010.Differential equations for engineers. Cambridge university press

  258. [264]

    Kaiwen Xue, Yuhao Zhou, Shen Nie, Xu Min, Xiaolu Zhang, JUN ZHOU, and Chongxuan Li. 2024. Unifying Bayesian Flow Networks and Diffusion Models through Stochastic Differential Equations. InInternational Conference on Machine Learning

  259. [265]

    Shuchen Xue, Mingyang Yi, Weijian Luo, Shifeng Zhang, Jiacheng Sun, Zhenguo Li, and Zhi-Ming Ma. 2023. Sa-solver: Stochastic adams solver for fast sampling of diffusion models. InAdvances in Neural Information Processing Systems, Vol. 36. 77632–77674

  260. [266]

    Yu Yamada, Masaki Iwamura, Takuya Akiba, and Koichi Kise. 2019. ShakeDrop regularization for deep residual learning.IEEE Access7 (2019), 186126–186136

  261. [267]

    Yoshihiro Yamada, Masakazu Iwamura, and Koichi Kise. 2016. Deep Pyramidal Residual Networks with Separated Stochastic Depth.CoRR abs/1612.01230 (2016). http://arxiv.org/abs/1612.01230

  262. [268]

    Hanshu YAN, Jiawei DU, Vincent TAN, and Jiashi FENG. 2020. On Robustness of Neural Ordinary Differential Equations. InInternational Conference on Learning Representations

  263. [269]

    Hanshu Yan, Xingchao Liu, Jiachun Pan, Jun Hao Liew, Qiang Liu, and Jiashi Feng. 2024. Perflow: Piecewise rectified flow as universal plug-and-play accelerator. InAdvances in Neural Information Processing Systems, Vol. 37. 78630–78652

  264. [270]

    Ling Yang, Zhilong Zhang, Yang Song, Shenda Hong, Runsheng Xu, Yue Zhao, Wentao Zhang, Bin Cui, and Ming-Hsuan Yang. 2023. Diffusion models: A comprehensive survey of methods and applications.Comput. Surveys56, 4 (2023), 1–39

  265. [271]

    Ling Yang, Zixiang Zhang, Zhilong Zhang, Xingchao Liu, Minkai Xu, Wentao Zhang, Chenlin Meng, Stefano Ermon, and Bin Cui. 2024. Consistency Flow Matching: Defining Straight Flows with Velocity Consistency.CoRRabs/2407.02398 (2024). https://doi.org/10.48550/arXiv.2407.02398

  266. [272]

    Yuan-Chih Yang and Hung-Hsuan Chen. 2025. Dynamic DropConnect: Enhancing Neural Network Robustness Through Adaptive Edge Dropping Strategies. InData Science: Foundations and Applications. 110–121

  267. [273]

    Nanyang Ye, Linfeng Cao, Liujia Yang, Ziqing Zhang, Zhicheng Fang, Qinying Gu, and Guang-Zhong Yang. 2023. Improving the robustness of analog deep neural networks through a Bayes-optimized noise injection approach.Communications Engineering2, 1 (2023), 25. Manuscript submitted...

  268. [274]

    Eun Bi Yoon, Keehun Park, Sungwoong Kim, and Sungbin Lim. 2023. Score-based generative models with Lévy processes. InAdvances in Neural Information Processing Systems, Vol. 36. 40694–40707

  269. [275]

    Fisher Yu, Ari Seff, Yinda Zhang, Shuran Song, Thomas Funkhouser, and Jianxiong Xiao. 2015. Lsun: Construction of a large-scale image dataset using deep learning with humans in the loop.arXiv preprint arXiv:1506.03365(2015)

  270. [276]

    Hu Yu, Jie Huang, Kaiwen Zheng, and Feng Zhao. 2023. High-quality image dehazing with diffusion model.arXiv preprint arXiv:2308.11949(2023)

  271. [277]

    Ye Yuan, Jiaming Song, Umar Iqbal, Arash Vahdat, and Jan Kautz. 2023. Physdiff: Physics-guided human motion diffusion model. InProceedings of the IEEE/CVF international conference on computer vision. 16010–16021

  272. [278]

    Sergey Zagoruyko and Nikos Komodakis. 2016. Wide Residual Networks. InBritish Machine Vision Conference

  273. [279]

    Hong Zhang, Ying Liu, and Romit Maulik. 2025. Semi-implicit neural ordinary differential equations. InProceedings of the AAAI Conference on Artificial Intelligence, Vol. 39. 22416–22424

  274. [280]

    Jianxin Zhang, Josh Viktorov, Doosan Jung, and Emily Pitler. 2024. Efficient training of neural stochastic differential equations by matching finite dimensional distributions. InInternational Conference on Learning Representations

  275. [281]

    Jianxin Zhang, Josh Viktorov, Doosan Jung, and Emily Pitler. 2025. Efficient Training of Neural Stochastic Differential Equations by Matching Finite Dimensional Distributions. InInternational Conference on Learning Representations

  276. [282]

    Kai Zhang, Wangmeng Zuo, Yunjin Chen, Deyu Meng, and Lei Zhang. 2017. Beyond a Gaussian Denoiser: Residual Learning of Deep CNN for Image Denoising.IEEE Transactions on Image Processing26, 7 (2017), 3142–3155. doi:10.1109/TIP.2017.2662206

  277. [283]

    Linan Zhang and Hayden Schaeffer. 2020. Forward stability of ResNet and its variants.Journal of Mathematical Imaging and Vision62 (2020), 328–351

  278. [284]

    Qinsheng Zhang and Yongxin Chen. 2021. Diffusion normalizing flow. InAdvances in Neural Information Processing Systems, Vol. 34. 16280–16291

  279. [285]

    Qinsheng Zhang and Yongxin Chen. 2022. Fast Sampling of Diffusion Models with Exponential Integrator. InNeurIPS 2022 Workshop on Score-Based Methods

  280. [286]

    Qinsheng Zhang, Molei Tao, and Yongxin Chen. 2023. gDDIM: Generalized denoising diffusion implicit models. InInternational Conference on Learning Representations

  281. [287]

    Xingcheng Zhang, Zhizhong Li, Chen Change Loy, and Dahua Lin. 2017. Polynet: A pursuit of structural diversity in very deep networks. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition. 718–726

  282. [288]

    Yichi Zhang, Yici Yan, Alex Schwing, and Zhizhen Zhao. 2025. Hierarchical Rectified Flow Matching with Mini-Batch Couplings.arXiv preprint arXiv:2507.13350(2025)

  283. [290]

    Hongjue Zhao, Yuchen Wang, Hairong Qi, Zijie Huang, Han Zhao, Lui Sha, and Huajie Shao. 2025. Accelerating Neural ODEs: A Variational Formulation-based Approach. InInternational Conference on Learning Representations

  284. [291]

    Wenqing Zheng, Jiyang Xie, Xian Sun, and Zhanyu Ma. 2022. Structured Dropconnect for Uncertainty Inference in Image Classification. In2022 IEEE International Conference on Image Processing (ICIP). IEEE, 366–370

  285. [292]

    Wangchunshu Zhou, Tao Ge, Furu Wei, Ming Zhou, and Ke Xu. 2020. Scheduled DropHead: A Regularization Method for Transformer Models. In Findings of the Association for Computational Linguistics: EMNLP 2020. 1971–1980

  286. [293]

    Zhenyu Zhou, Defang Chen, Can Wang, and Chun Chen. 2024. Fast ode-based sampling for diffusion models in around 5 steps. InProceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition. 7777–7786

  287. [294]

    Mai Zhu, Bo Chang, and Chong Fu. 2023. Convolutional neural networks combined with Runge–Kutta methods.Neural Computing and Applications 35, 2 (2023), 1629–1643

  288. [296]

    Yixuan Zhu, Wenliang Zhao, Ao Li, Yansong Tang, Jie Zhou, and Jiwen Lu. 2024. FlowIE: Efficient Image Enhancement via Rectified Flow. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition. 13–22. Manuscript submitted to ACM

  289. [2025]

    InNeural Information Processing Systems

    Metric flow matching for smooth interpolations on the data manifold. InNeural Information Processing Systems. Article 4291, 32 pages

Pith tools

Reviewed August 4, 2026 · model on record in the stance chip above.