Pith. sign in

REVIEW 3 major objections 3 minor 161 references

Towards Adaptive External Communication in Autonomous Vehicles: A Conceptual Design Framework

T0 review · 3 major / 3 minor · reviewed 2026-08-05 · deepseek-v4-flash

Pith's one-line read Adaptive car-to-pedestrian displays need a three-layer design framework so the vehicle's message can change with who is around and what is happening.

desk verdict A plausible three-layer framework for adaptive eHMIs, but the only part we can actually check is the abstract; the provided full text is a different paper, so the central utility claim is unverified. read the letter →

arxiv 2508.12518 v1 pith:NUI62EGP submitted 2025-08-17 cs.HC

classification cs.HC
keywords adaptiveeHMIexternalhuman-machineinterfaceautonomousvehiclesdesignframeworkcyber-physicalsystemroadactorcommunicationscalabilityinclusivity
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

The paper proposes that external communication from autonomous vehicles should be adaptive: the interface should change what it says, how it says it, and when it says it depending on who is around and what is happening. It offers a three-layer conceptual framework—Input, Processing, Output—structured by a cyber-physical system lens, so that designers and researchers can think about adaptive eHMIs systematically. The claim is that this structure gives a common tool for designing, analysing, and comparing adaptive communication strategies, and that doing so can address scalability and inclusivity problems left unsolved by current reactive eHMIs. This matters because autonomous vehicles will share streets with pedestrians, cyclists, and other drivers who have very different needs and abilities. If the framework is right, it supplies a starting structure for turning context awareness into communication choices.

What carries the argument

The central organizing object is the three-layer Input–Processing–Output decomposition, using the cyber-physical system as a structuring lens. It separates perception (what the vehicle detects about road actors and context), decision (how the system chooses communication), and expression (how the interface outputs the message), so that each layer can vary independently and be analysed for its contribution to communication success.

What would settle it

Take two groups of designers: one uses the framework, one does not; ask both to design an eHMI for the same mixed-traffic scenario, then measure pedestrian comprehension and crossing decisions in a simulator across child, adult, older adult, and cyclist actors in day and night contexts. If framework-designed interfaces are not reliably better, the claim that the framework helps design, analyse, and assess adaptive eHMIs fails.

Watch

Extended reading notes

Core claim

The central discovery is a conceptual ordering rather than a measured result: adaptive eHMI design can be decomposed into what the system detects (Input), how it decides to communicate (Processing), and how it presents the communication (Output). The paper argues that most current eHMIs are reactive—fixed messages that do not adjust—and that treating the interface as part of a cyber-physical loop makes variability across road actors and contexts a design parameter rather than an afterthought. Developed through theory-led abstraction and expert discussion, the framework is offered as a structured tool for researchers and designers to design, analyse, and assess adaptive communication strategi

Load-bearing premise

The load-bearing premise is that expert-derived categories of road actors and contexts are complete enough that a designer can map any real situation onto Input–Processing–Output choices, and that this mapping improves real eHMIs; the paper offers no user data to establish that.

Editorial extensions

If this is right

  • eHMI design becomes a configurable pipeline rather than a fixed display, so a single interface can vary message content, timing, and modality by road actor and context.
  • Researchers can attribute communication outcomes to specific layers, making studies of adaptive eHMIs easier to compare and reproduce.
  • Scalability is treated as a design requirement: the system does not need a new fixed message for every situation if the Processing layer can generate appropriate outputs from detected inputs.
  • Inclusivity becomes part of the Processing layer's job, since adaptation must account for differences in perception, literacy, and mobility among road actors.
  • Wider adoption of adaptive eHMIs would confront design teams with ethical choices about sensing people and deciding who receives extra cautionary information.

Reading between the lines

Editorial extensions of the paper, not claims the author makes directly.

  • If the framework is right, a concrete prediction follows that the paper does not test: a system designed through Input–Processing–Output will communicate better than a static interface in mixed-traffic settings with varied actors and contexts.
  • The cyber-physical framing implies evaluation should cover the entire perception–decision–output loop; a display that fails because the vehicle misdetected an actor is still an adaptive-eHMI failure under this view.
  • One testable extension is to turn the framework into an audit matrix (road actor × context × message type) and use it to identify gaps in existing eHMI proposals.
  • Note on the supplied record: the full text under this title is a different manuscript on stochastic optimal control; it neither elaborates nor tests the eHMI framework, so the framework's support rests on the abstract's description of theory-led abstraction and expert discussion.
Share X Bluesky LinkedIn Reddit HN

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, and a circularity audit.

Referee Report

3 major / 3 minor

Summary. The manuscript arXiv:2508.12518 presents, in its abstract, a conceptual design framework for adaptive external human-machine interfaces (eHMIs) for autonomous vehicles. The framework is described as having three layers—Input, Processing, Output—developed through theory-led abstraction and expert discussion, and is claimed to help researchers and designers systematically design, analyse, and assess adaptive communication strategies, resolving longstanding limitations in eHMI research. The full text supplied for review, however, is arXiv:2508.12511v2, a paper on trust-region stochastic optimal control and measure transport. It contains no description of eHMIs, no Input/Processing/Output framework, and no expert-discussion methodology. Consequently, the development of the framework, its definitions, its application, and its claimed benefits cannot be checked from the submitted manuscript.

Significance. If properly developed and validated, a three-layer framework for adaptive eHMIs could provide a useful structuring vocabulary for a fragmented research area, especially with attention to scalability and inclusivity. The conceptual decomposition into Input, Processing, and Output is a reasonable starting point, and the abstract's emphasis on dynamic adjustment to road actors and context is timely. However, the current manuscript provides no evidence that the framework is descriptively adequate or practically useful: there is no worked example, no user or design study, no systematic classification of existing eHMI implementations, and no comparison with alternative taxonomies. The significance therefore remains entirely hypothetical at this stage.

major comments (3)
  1. [Full text / Abstract] The supplied full text is arXiv:2508.12511v2, a paper on trust-region stochastic optimal control, not the eHMI manuscript announced in the abstract. This mismatch removes the body of the paper under review: the Input/Processing/Output framework, the theory-led abstraction, the expert discussion, and the claims about resolving longstanding limitations have no supporting text. This is load-bearing because no reviewer can verify that the framework is derived as claimed or that the layers are coherently defined. The paper must be resubmitted with the correct full text.
  2. [Abstract, central utility claim] The claim that the framework 'helps researchers and designers think systematically' and provides a 'structured tool to design, analyse, and assess' adaptive eHMIs is unsupported. Theory-led abstraction and expert discussion can generate a taxonomy, but they do not establish that the taxonomy maps cleanly onto real road-actor variability or that it is practically useful. No user study, design exercise, or systematic classification of existing eHMI implementations is reported. A concrete test would be to apply the framework to existing eHMI designs spanning multi-party interactions, occluded pedestrians, group behaviour, and time-varying contexts, and to show that it represents them without omission or conflation.
  3. [Abstract, 'resolving longstanding limitations'] The abstract does not identify which longstanding limitations in eHMI research are resolved, nor the mechanism by which the framework resolves them. This makes the claim unfalsifiable. The authors should enumerate the limitations and map each to a specific feature of the Input/Processing/Output layers, ideally demonstrating the mapping on known hard cases such as group behaviour, occluded road users, and context-dependent communication norms.
minor comments (3)
  1. [Abstract / Definitions] The term 'adaptive' is used without a precise definition. It could mean reactive parameter changes, learned adaptation, or policy-level adjustment. The introduction of the correct full text should define this explicitly.
  2. [References] Once the correct full text is provided, the paper should engage with existing eHMI surveys and taxonomies to position the contribution within the field; the current abstract provides no related-work grounding.
  3. [Notation] The labels Input, Processing, and Output need explicit definitions and examples. Without the correct body, terms like 'what the system detects' are ambiguous and cannot be operationalized.

Circularity Check

0 steps flagged · score 0.0 of 10

No significant circularity: the abstract presents an expert-derived taxonomy whose utility is unverified, but no derivation or prediction reduces to its inputs.

full rationale

The abstract makes no quantitative prediction, fits no parameter, and contains no equation in which an output is equal to an input by construction. The proposed Input/Processing/Output framework is a structuring lens (a standard detect-decide-act decomposition), and the claim that it 'helps researchers and designers think systematically' is a stated design hypothesis grounded in 'theory-led abstraction and expert discussion.' That claim is unsupported by user studies or deployed-system evidence, but lack of empirical validation is a correctness risk, not circularity. There are no load-bearing self-citations, no imported uniqueness theorems, and no fitted values renamed as predictions. The provided full text is a different arXiv paper (2508.12511 on trust-region stochastic optimal control), so it cannot corroborate or contradict the eHMI abstract; this is a document-integrity / input-mismatch issue rather than circular reasoning. Honest finding: no circularity identified in the abstract's derivation chain.

Assumptions & free parameters 0 free parameters · 3 assumptions · 1 invented entities

Only the abstract of the target paper was provided; the full text corresponds to a different preprint. The framework as described introduces no free parameters, but depends on domain assumptions about adaptivity and the sufficiency of expert-derived categories.

assumptions (3)
  • domain assumption Adaptive eHMIs are a desirable or necessary direction for autonomous vehicle external communication.
    The entire framework presupposes that adaptivity is the right goal; the abstract states it as motivation, not as a finding.
  • domain assumption The cyber-physical system lens provides a valid structuring of eHMI design.
    The framework 'uses the cyber-physical system as a structuring lens'; no derivation or evidence for this choice is given.
  • domain assumption Theory-led abstraction and expert discussion yield a framework that is useful for design, analysis, and assessment.
    The abstract describes the method of development but offers no empirical validation of the resulting framework.
invented entities (1)
  • Three-layer adaptive eHMI framework (Input, Processing, Output)
    purpose: Structures how adaptive external human-machine interfaces sense, decide, and communicate
    The framework is a conceptual construct introduced by the paper; no falsifiable handle or independent validation is described in the abstract.

how reviews work

0 comments
Cite this review

Pith. "Pith review of Towards Adaptive External Communication in Autonomous Vehicles: A Conceptual Design Framework." pith.science (2026). https://pith.science/paper/NUI62EGP

@misc{pith2026250812518,
  author       = {Pith},
  title        = {Pith review of: Towards Adaptive External Communication in Autonomous Vehicles: A Conceptual Design Framework},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/NUI62EGP}},
  note         = {Machine review of arXiv:2508.12518}
}
read the original abstract

External Human-Machine Interfaces (eHMIs) are key to facilitating interaction between autonomous vehicles and external road actors, yet most remain reactive and do not account for scalability and inclusivity. This paper introduces a conceptual design framework for adaptive eHMIs-interfaces that dynamically adjust communication as road actors vary and context shifts. Using the cyber-physical system as a structuring lens, the framework comprises three layers: Input (what the system detects), Processing (how the system decides), and Output (how the system communicates). Developed through theory-led abstraction and expert discussion, the framework helps researchers and designers think systematically about adaptive eHMIs and provides a structured tool to design, analyse, and assess adaptive communication strategies. We show how such systems may resolve longstanding limitations in eHMI research while raising new ethical and technical considerations.

Discussion (0). Sign in to comment.

Reference graph

Works this paper leans on

161 extracted references · 31 canonical work pages

  1. [1]

    Abdolmaleki, R

    A. Abdolmaleki, R. Lioutikov, J. R. Peters, N. Lau, L. Pualo Reis, and G. Neumann. Model- based relative entropy stochastic search.Advances in Neural Information Processing Systems, 28, 2015

  2. [2]

    Abdolmaleki, J

    A. Abdolmaleki, J. T. Springenberg, J. Degrave, S. Bohez, Y . Tassa, D. Belov, N. Heess, and M. Riedmiller. Relative entropy regularized policy iteration.arXiv preprint arXiv:1812.02256, 2018

  3. [3]

    Abdolmaleki, J

    A. Abdolmaleki, J. T. Springenberg, Y . Tassa, R. Munos, N. Heess, and M. Riedmiller. Maximum a posteriori policy optimisation.arXiv preprint arXiv:1806.06920, 2018

  4. [4]

    Achiam, D

    J. Achiam, D. Held, A. Tamar, and P. Abbeel. Constrained policy optimization. InInternational conference on machine learning, pages 22–31. PMLR, 2017

  5. [5]

    Akhound-Sadegh, J

    T. Akhound-Sadegh, J. Lee, A. J. Bose, V . De Bortoli, A. Doucet, M. M. Bronstein, D. Beaini, S. Ravanbakhsh, K. Neklyudov, and A. Tong. Progressive inference-time annealing of diffusion models for sampling from Boltzmann densities.arXiv preprint arXiv:2506.16471, 2025

  6. [6]

    Akhound-Sadegh, J

    T. Akhound-Sadegh, J. Rector-Brooks, A. J. Bose, S. Mittal, P. Lemos, C.-H. Liu, M. Sendera, S. Ravanbakhsh, G. Gidel, Y . Bengio, et al. Iterated denoising energy matching for sampling from Boltzmann densities.arXiv preprint arXiv:2402.06121, 2024

  7. [7]

    Akrour, J

    R. Akrour, J. Pajarinen, J. Peters, and G. Neumann. Projections for approximate policy iteration algorithms. InInternational Conference on Machine Learning, pages 181–190. PMLR, 2019

  8. [8]

    M. S. Albergo and E. Vanden-Eijnden. NETS: A non-equilibrium transport sampler.arXiv preprint arXiv:2410.02711, 2024

Show all 161 references
  1. [9]

    Arenz, P

    O. Arenz, P. Dahlinger, Z. Ye, M. V olpp, and G. Neumann. A unified perspective on natural gradient variational inference with Gaussian mixture models.arXiv preprint arXiv:2209.11533, 2022

  2. [10]

    Arenz, M

    O. Arenz, M. Zhong, and G. Neumann. Trust-region variational inference with Gaussian mixture models.Journal of Machine Learning Research, 21(163):1–60, 2020

  3. [11]

    C. Beck, S. Becker, P. Grohs, N. Jaafari, and A. Jentzen. Solving the Kolmogorov PDE by means of deep learning.Journal of Scientific Computing, 88:1–28, 2021

  4. [12]

    Becker, N

    P. Becker, N. Freymuth, S. Thilges, F. Otto, and G. Neumann. Troll: Trust regions improve reinforcement learning for large language models.arXiv preprint arXiv:2510.03817, 2025

  5. [13]

    Bellman.Dynamic programming

    R. Bellman.Dynamic programming. Princeton University Press, 1957

  6. [14]

    Berner, M

    J. Berner, M. Dablander, and P. Grohs. Numerically solving parametric families of high- dimensional Kolmogorov partial differential equations via deep learning.Advances in Neural Information Processing Systems, 33:16615–16627, 2020

  7. [15]

    Berner, L

    J. Berner, L. Richter, M. Sendera, J. Rector-Brooks, and N. Malkin. From discrete-time policies to continuous-time diffusion samplers: Asymptotic equivalences and faster training. arXiv preprint arXiv:2501.06148, 2025

  8. [16]

    Berner, L

    J. Berner, L. Richter, and K. Ullrich. An optimal control perspective on diffusion-based generative modeling.Transactions on Machine Learning Research, 2024

  9. [17]

    Black, M

    K. Black, M. Janner, Y . Du, I. Kostrikov, and S. Levine. Training diffusion models with reinforcement learning. InThe Twelfth International Conference on Learning Representations, 2024

  10. [18]

    Blessing, J

    D. Blessing, J. Berner, L. Richter, and G. Neumann. Underdamped diffusion bridges with appli- cations to sampling. InThe Thirteenth International Conference on Learning Representations, 2025

  11. [19]

    Blessing, X

    D. Blessing, X. Jia, J. Esslinger, F. Vargas, and G. Neumann. Beyond ELBOs: A large-scale evaluation of variational methods for sampling.arXiv preprint arXiv:2406.07423, 2024. 11

  12. [20]

    Blessing, X

    D. Blessing, X. Jia, and G. Neumann. End-to-end learning of Gaussian mixture priors for diffusion sampler.arXiv preprint arXiv:2503.00524, 2025

  13. [21]

    P. G. Bolhuis, D. Chandler, C. Dellago, and P. L. Geissler. Transition path sampling: Throwing ropes over rough mountain passes, in the dark.Annual review of physical chemistry, 53(1):291– 318, 2002

  14. [22]

    Bradbury, R

    J. Bradbury, R. Frostig, P. Hawkins, M. J. Johnson, C. Leary, D. Maclaurin, G. Necula, A. Paszke, J. VanderPlas, S. Wanderman-Milne, et al. Jax: Autograd and xla.Astrophysics Source Code Library, pages ascl–2111, 2021

  15. [23]

    Brekelmans, V

    R. Brekelmans, V . Masrani, F. Wood, G. V . Steeg, and A. Galstyan. All in the exponential fam- ily: Bregman duality in thermodynamic variational inference.arXiv preprint arXiv:2007.00642, 2020

  16. [24]

    R. P. Brent. An algorithm with guaranteed convergence for finding a zero of a function.The computer journal, 14(4):422–425, 1971

  17. [25]

    Celik, Z

    O. Celik, Z. Li, D. Blessing, G. Li, D. Palanicek, J. Peters, G. Chalvatzaki, and G. Neu- mann. DIME: Diffusion-based maximum entropy reinforcement learning.arXiv preprint arXiv:2502.02316, 2025

  18. [26]

    J. Chen, L. Richter, J. Berner, D. Blessing, G. Neumann, and A. Anandkumar. Sequential controlled Langevin diffusions. InThe Thirteenth International Conference on Learning Representations, 2025

  19. [27]

    Chetrite and H

    R. Chetrite and H. Touchette. Variational and optimal control representations of condi- tioned and driven processes.Journal of Statistical Mechanics: Theory and Experiment, 2015(12):P12001, 2015

  20. [28]

    J. Choi, Y . Chen, M. Tao, and G.-H. Liu. Non-equilibrium annealed adjoint sampler.arXiv preprint arXiv:2506.18165, 2025

  21. [29]

    Clark, P

    K. Clark, P. Vicol, K. Swersky, and D. J. Fleet. Directly fine-tuning diffusion models on differentiable rewards. InThe Twelfth International Conference on Learning Representations, 2024

  22. [30]

    A. R. Conn, N. I. Gould, and P. L. Toint.Trust region methods. SIAM, 2000

  23. [31]

    G. E. Crooks. Measuring thermodynamic length.Physical Review Letters, 99(10):100602, 2007

  24. [32]

    M. Cuturi. Sinkhorn distances: Lightspeed computation of optimal transport.Advances in neural information processing systems, 26, 2013

  25. [33]

    Cuturi, L

    M. Cuturi, L. Meng-Papaxanthos, Y . Tian, C. Bunne, G. Davis, and O. Teboul. Optimal trans- port tools (OTT): A jax toolbox for all things Wasserstein.arXiv preprint arXiv:2201.12324, 2022

  26. [34]

    P. Dai Pra. A stochastic control approach to reciprocal diffusion processes.Applied mathemat- ics and Optimization, 23(1):313–329, 1991

  27. [35]

    Dai Pra, L

    P. Dai Pra, L. Meneghini, and W. J. Runggaldier. Connections between stochastic control and dynamic games.Mathematics of Control, Signals and Systems, 9:303–326, 1996

  28. [36]

    A. Das, D. C. Rose, J. P. Garrahan, and D. T. Limmer. Reinforcement learning of rare diffusive dynamics.The Journal of Chemical Physics, 155(13), 2021

  29. [37]

    De Bortoli, J

    V . De Bortoli, J. Thornton, J. Heng, and A. Doucet. Diffusion Schrödinger bridge with applications to score-based generative modeling.Advances in Neural Information Processing Systems, 34:17695–17709, 2021

  30. [38]

    Dellago, P

    C. Dellago, P. G. Bolhuis, and D. Chandler. Efficient transition path sampling: Application to lennard-jones cluster rearrangements.The Journal of chemical physics, 108(22):9236–9245, 1998. 12

  31. [39]

    K. Didi, F. Vargas, S. V . Mathis, V . Dutordoir, E. Mathieu, U. J. Komorowska, and P. Lio. A framework for conditional diffusion modelling with applications in motif scaffolding for protein design.arXiv preprint arXiv:2312.09236, 2023

  32. [40]

    Z. Ding, Y . Jiao, X. Lu, Z. Yang, and C. Yuan. Sampling via Föllmer flow.arXiv preprint arXiv:2311.03660, 2023

  33. [41]

    Domingo-Enrich

    C. Domingo-Enrich. A taxonomy of loss functions for stochastic optimal control.arXiv preprint arXiv:2410.00345, 2024

  34. [42]

    Domingo-Enrich, M

    C. Domingo-Enrich, M. Drozdzal, B. Karrer, and R. T. Q. Chen. Adjoint matching: Fine-tuning flow and diffusion generative models with memoryless stochastic optimal control. InThe Thirteenth International Conference on Learning Representations, 2025

  35. [43]

    Domingo-Enrich, J

    C. Domingo-Enrich, J. Han, B. Amos, J. Bruna, and R. T. Q. Chen. Stochastic optimal control matching. InThe Thirty-eighth Annual Conference on Neural Information Processing Systems, 2024

  36. [44]

    Doucet, W

    A. Doucet, W. Grathwohl, A. G. d. G. Matthews, and H. Strathmann. Score-based diffusion meets annealed importance sampling. InAdvances in Neural Information Processing Systems, 2022

  37. [45]

    Y . Du, M. Plainer, R. Brekelmans, C. Duan, F. Noe, C. P. Gomes, A. Aspuru-Guzik, and K. Neklyudov. Doob’s lagrangian: A sample-efficient variational approach to transition path sampling. InThe Thirty-eighth Annual Conference on Neural Information Processing Systems

  38. [46]

    Y . Du, M. Plainer, R. Brekelmans, C. Duan, F. Noe, C. P. Gomes, A. Aspuru-Guzik, and K. Neklyudov. Doob’s lagrangian: A sample-efficient variational approach to transition path sampling. InThe Thirty-eighth Annual Conference on Neural Information Processing Systems, 2024

  39. [47]

    Erives, B

    E. Erives, B. Jing, P. Holderrieth, and T. Jaakkola. Continuously tempered diffusion samplers. InFrontiers in Probabilistic Inference: Learning meets Sampling, 2025

  40. [48]

    Y . Fan, O. Watkins, Y . Du, H. Liu, M. Ryu, C. Boutilier, P. Abbeel, M. Ghavamzadeh, K. Lee, and K. Lee. Dpok: Reinforcement learning for fine-tuning text-to-image diffusion models. arXiv preprint arXiv:2305.16381, 2023

  41. [49]

    M. F. Faulkner and S. Livingstone. Sampling algorithms in statistical physics: a guide for statistics and machine learning.Statistical Science, 39(1):137–164, 2024

  42. [50]

    Fleming and R

    W. Fleming and R. Rishel.Deterministic and Stochastic Optimal Control. Applications of mathematics. Springer, 1975

  43. [51]

    W. H. Fleming and H. M. Soner.Controlled Markov processes and viscosity solutions, volume 25. Springer Science & Business Media, 2006

  44. [52]

    S. Fu, N. Tamir, S. Sundaram, L. Chai, R. Zhang, T. Dekel, and P. Isola. Dreamsim: Learning new dimensions of human visual similarity using synthetic data.arXiv preprint arXiv:2306.09344, 2023

  45. [53]

    Geffner and J

    T. Geffner and J. Domke. MCMC variational inference via uncorrected hamiltonian annealing. Advances in Neural Information Processing Systems, 34:639–651, 2021

  46. [54]

    Geffner and J

    T. Geffner and J. Domke. Langevin diffusion variational inference. InInternational Conference on Artificial Intelligence and Statistics, pages 576–593. PMLR, 2023

  47. [55]

    Gelman, J

    A. Gelman, J. Carlin, H. Stern, D. Dunson, A. Vehtari, and D. Rubin.Bayesian Data Analysis, Third Edition. Chapman & Hall/CRC Texts in Statistical Science. Taylor & Francis, 2013

  48. [56]

    Grenioux, M

    L. Grenioux, M. Noble, and M. Gabrié. Improving the evaluation of samplers on multi-modal targets.arXiv preprint arXiv:2504.08916, 2025

  49. [57]

    Gritsaev, N

    T. Gritsaev, N. Morozov, K. Tamogashev, D. Tiapkin, S. Samsonov, A. Naumov, D. Vetrov, and N. Malkin. Adaptive destruction processes for diffusion samplers.arXiv preprint arXiv:2506.01541, 2025. 13

  50. [58]

    W. Guo, M. Tao, and Y . Chen. Complexity analysis of normalizing constant estimation: from Jarzynski equality to annealed importance sampling and beyond.arXiv preprint arXiv:2502.04575, 2025

  51. [59]

    J. Han, A. Jentzen, and W. E. Solving high-dimensional partial differential equations using deep learning.Proceedings of the National Academy of Sciences, 115(34):8505–8510, 2018

  52. [60]

    Hartmann, O

    C. Hartmann, O. Kebiri, L. Neureither, and L. Richter. Variational approach to rare event simulation using least-squares regression.Chaos: An Interdisciplinary Journal of Nonlinear Science, 29(6), 2019

  53. [61]

    Hartmann and L

    C. Hartmann and L. Richter. Nonasymptotic bounds for suboptimal importance sampling. SIAM/ASA Journal on Uncertainty Quantification, 12(2):309–346, 2024

  54. [62]

    Hartmann, L

    C. Hartmann, L. Richter, C. Schütte, and W. Zhang. Variational characterization of free energy: Theory and algorithms.Entropy, 19(11), 2017

  55. [63]

    Hartmann and C

    C. Hartmann and C. Schütte. Efficient rare event simulation by optimal nonequilibrium forcing. Journal of Statistical Mechanics: Theory and Experiment, 2012(11):P11004, 2012

  56. [64]

    Havens, B

    A. Havens, B. K. Miller, B. Yan, C. Domingo-Enrich, A. Sriram, B. Wood, D. Levine, B. Hu, B. Amos, B. Karrer, et al. Adjoint sampling: Highly scalable diffusion samplers via adjoint matching.arXiv preprint arXiv:2504.11713, 2025

  57. [65]

    J. He, W. Chen, M. Zhang, D. Barber, and J. M. Hernández-Lobato. Training neural samplers with reverse diffusive kl divergence.arXiv preprint arXiv:2410.12456, 2024

  58. [66]

    J. He, Y . Du, F. Vargas, Y . Wang, C. P. Gomes, J. M. Hernández-Lobato, and E. Vanden-Eijnden. FEAT: Free energy estimators with adaptive transport.arXiv preprint arXiv:2504.11516, 2025

  59. [67]

    J. He, Y . Du, F. Vargas, D. Zhang, S. Padhy, R. OuYang, C. Gomes, and J. M. Hernández- Lobato. No trick, no treat: Pursuits and challenges towards simulation-free training of neural samplers.arXiv preprint arXiv:2502.06685, 2025

  60. [68]

    Hendrycks and K

    D. Hendrycks and K. Gimpel. Gaussian error linear units (gelus).arXiv preprint arXiv:1606.08415, 2016

  61. [69]

    Hénin, T

    J. Hénin, T. Lelièvre, M. R. Shirts, O. Valsson, and L. Delemotte. Enhanced sampling methods for molecular dynamics simulations.arXiv preprint arXiv:2202.04164, 2022

  62. [70]

    Hertrich and R

    J. Hertrich and R. Gruhlke. Importance corrected neural jko sampling.arXiv preprint arXiv:2407.20444, 2024

  63. [71]

    Hessel, A

    J. Hessel, A. Holtzman, M. Forbes, R. L. Bras, and Y . Choi. Clipscore: A reference-free evaluation metric for image captioning.arXiv preprint arXiv:2104.08718, 2021

  64. [72]

    Holderrieth, M

    P. Holderrieth, M. S. Albergo, and T. Jaakkola. Leaps: A discrete neural sampler via locally equivariant networks.arXiv preprint arXiv:2502.10843, 2025

  65. [73]

    Holdijk, Y

    L. Holdijk, Y . Du, F. Hooft, P. Jaini, B. Ensing, and M. Welling. Stochastic optimal control for collective variable free sampling of molecular transition paths. InThirty-seventh Conference on Neural Information Processing Systems, 2023

  66. [74]

    Holdijk, Y

    L. Holdijk, Y . Du, P. Jaini, F. Hooft, B. Ensing, and M. Welling. Path integral stochastic optimal control for sampling transition paths. InICML 2022 2nd AI for Science Workshop, 2022

  67. [75]

    Huang, Y

    J. Huang, Y . Jiao, L. Kang, X. Liao, J. Liu, and Y . Liu. Schrödinger-Föllmer sampler: sampling without ergodicity.arXiv preprint arXiv:2106.10880, 2021

  68. [76]

    Huang, H

    X. Huang, H. Dong, Y . Hao, Y . Ma, and T. Zhang. Monte Carlo sampling without isoperimetry: A reverse diffusion approach.arXiv preprint arXiv:2307.02037, 2023

  69. [77]

    Izrailev, S

    S. Izrailev, S. Stepaniants, B. Isralewitz, D. Kosztin, H. Lu, F. Molnar, W. Wriggers, and K. Schulten. Steered molecular dynamics. InComputational Molecular Dynamics: Chal- lenges, Methods, Ideas: Proceedings of the 2nd International Symposium on Algorithms for Macromolecular...

  70. [78]

    H. J. Kappen and H. C. Ruiz. Adaptive importance sampling for control and inference.Journal of Statistical Physics, 162(5):1244–1266, 2016

  71. [79]

    M. Kim, S. Choi, T. Yun, E. Bengio, L. Feng, J. Rector-Brooks, S. Ahn, J. Park, N. Malkin, and Y . Bengio. Adaptive teachers for amortized samplers.arXiv preprint arXiv:2410.01432, 2024

  72. [80]

    M. Kim, K. Seong, D. Woo, S. Ahn, and M. Kim. On scalable and efficient training of diffusion samplers.arXiv preprint arXiv:2505.19552, 2025

  73. [81]

    D. P. Kingma. Adam: A method for stochastic optimization.arXiv preprint arXiv:1412.6980, 2014

  74. [82]

    G.-H. Liu, J. Choi, Y . Chen, B. K. Miller, and R. T. Chen. Adjoint schrödinger bridge sampler. arXiv preprint arXiv:2506.22565, 2025

  75. [83]

    J. Liu, G. Liu, J. Liang, Y . Li, J. Liu, X. Wang, P. Wan, D. Zhang, and W. Ouyang. Flow-grpo: Training flow matching models via online rl, 2025

  76. [84]

    Z. Liu, T. Z. Xiao, W. Liu, Y . Bengio, and D. Zhang. Efficient diversity-preserving diffusion alignment via gradient-informed GFlowNets, 2025

  77. [85]

    W. Meng, Q. Zheng, Y . Shi, and G. Pan. An off-policy trust region policy optimization method with monotonic improvement guarantee for deep reinforcement learning.IEEE Transactions on Neural Networks and Learning Systems, 33(5):2223–2235, 2021

  78. [86]

    L. I. Midgley, V . Stimper, G. N. Simm, B. Schölkopf, and J. M. Hernández-Lobato. Flow annealed importance sampling bootstrap.arXiv preprint arXiv:2208.01893, 2022

  79. [87]

    R. M. Neal. Probabilistic inference using Markov chain Monte Carlo methods. 1993

  80. [88]

    Noble, L

    M. Noble, L. Grenioux, M. Gabrié, and A. O. Durmus. Learned reference-based diffusion sampling for multi-modal distributions.arXiv preprint arXiv:2410.19449, 2024

  81. [89]

    Nüsken and L

    N. Nüsken and L. Richter. Solving high-dimensional Hamilton–Jacobi–Bellman PDEs using neural networks: perspectives from the theory of controlled diffusions and measures on path space.Partial differential equations and applications, 2:1–48, 2021

  82. [90]

    F. Otto, P. Becker, N. A. Vien, H. C. Ziesche, and G. Neumann. Differentiable trust region layers for deep reinforcement learning.arXiv preprint arXiv:2101.09207, 2021

  83. [91]

    OuYang, B

    R. OuYang, B. Qiang, and J. M. Hernández-Lobato. BNEM: A Boltzmann sampler based on bootstrapped noised energy matching.arXiv preprint arXiv:2409.09787, 2024

  84. [92]

    Pajarinen, H

    J. Pajarinen, H. L. Thai, R. Akrour, J. Peters, and G. Neumann. Compatible natural gradient policy search.Machine Learning, 108:1443–1466, 2019

  85. [93]

    M. Pavon. Stochastic control and nonequilibrium thermodynamical systems.Applied Mathe- matics and Optimization, 19(1):187–202, 1989

  86. [94]

    M. Pavon. On local entropy, stochastic control and deep neural networks.arXiv preprint arXiv:2204.13049, 2022

  87. [95]

    Peters, K

    J. Peters, K. Mulling, and Y . Altun. Relative entropy policy search. InProceedings of the AAAI Conference on Artificial Intelligence, volume 24, pages 1607–1612, 2010

  88. [96]

    Pham.Continuous-time Stochastic Control and Optimization with Financial Applications

    H. Pham.Continuous-time Stochastic Control and Optimization with Financial Applications. Stochastic Modelling and Applied Probability. Springer Berlin Heidelberg, 2009

  89. [97]

    Richter.Solving high-dimensional PDEs, approximation of path space measures and importance sampling of diffusions

    L. Richter.Solving high-dimensional PDEs, approximation of path space measures and importance sampling of diffusions. PhD thesis, BTU Cottbus-Senftenberg, 2021

  90. [98]

    Richter and J

    L. Richter and J. Berner. Robust SDE-based variational formulations for solving linear PDEs via deep learning. InInternational Conference on Machine Learning, pages 18649–18666. PMLR, 2022. 15

  91. [99]

    Richter and J

    L. Richter and J. Berner. Improved sampling via learned diffusions. InInternational Conference on Learning Representations, 2024

  92. [100]

    Richter, A

    L. Richter, A. Boustati, N. Nüsken, F. Ruiz, and O. D. Akyildiz. VarGrad: A low-variance gradient estimator for variational inference.Advances in Neural Information Processing Systems, 33:13481–13492, 2020

  93. [101]

    Richter, L

    L. Richter, L. Sallandt, and N. Nüsken. Solving high-dimensional parabolic PDEs using the tensor train format. InInternational Conference on Machine Learning, pages 8998–9009. PMLR, 2021

  94. [102]

    Richter, L

    L. Richter, L. Sallandt, and N. Nüsken. From continuous-time formulations to discretization schemes: tensor trains and robust regression for bsdes and parabolic pdes.Journal of Machine Learning Research, 25(248):1–40, 2024

  95. [103]

    Rissanen, R

    S. Rissanen, R. OuYang, J. He, W. Chen, M. Heinonen, A. Solin, and J. M. Hernández-Lobato. Progressive tempering sampler with diffusion.arXiv preprint arXiv:2506.05231, 2025

  96. [104]

    Rombach, A

    R. Rombach, A. Blattmann, D. Lorenz, P. Esser, and B. Ommer. High-resolution image synthesis with latent diffusion models. InProceedings of the IEEE/CVF conference on computer vision and pattern recognition, pages 10684–10695, 2022

  97. [105]

    D. C. Rose, J. F. Mair, and J. P. Garrahan. A reinforcement learning approach to rare trajectory sampling.New Journal of Physics, 23(1):013013, 2021

  98. [106]

    R. Y . Rubinstein and D. P. Kroese.The cross-entropy method: a unified approach to combi- natorial optimization, Monte-Carlo simulation and machine learning. Springer Science & Business Media, 2013

  99. [107]

    Sabate Vidales, D

    M. Sabate Vidales, D. Šiška, and L. Szpruch. Unbiased deep solvers for linear parametric PDEs.Applied Mathematical Finance, 28(4):299–329, 2021

  100. [108]

    Salamon and R

    P. Salamon and R. S. Berry. Thermodynamic length and dissipated availability.Physical Review Letters, 51(13):1127, 1983

  101. [109]

    Sanokowski, W

    S. Sanokowski, W. Berghammer, M. Ennemoser, H. P. Wang, S. Hochreiter, and S. Lehner. Scalable discrete diffusion samplers: Combinatorial optimization and statistical physics.arXiv preprint arXiv:2502.08696, 2025

  102. [110]

    Sanokowski, L

    S. Sanokowski, L. Gruber, C. Bartmann, S. Hochreiter, and S. Lehner. Rethinking losses for diffusion bridge samplers.arXiv preprint arXiv:2506.10982, 2025

  103. [111]

    Schopmans and P

    H. Schopmans and P. Friederich. Temperature-annealed boltzmann generators.arXiv preprint arXiv:2501.19077, 2025

  104. [112]

    Schulman, S

    J. Schulman, S. Levine, P. Abbeel, M. Jordan, and P. Moritz. Trust region policy optimization. InInternational conference on machine learning, pages 1889–1897. PMLR, 2015

  105. [113]

    Schulman, F

    J. Schulman, F. Wolski, P. Dhariwal, A. Radford, and O. Klimov. Proximal policy optimization algorithms.arXiv preprint arXiv:1707.06347, 2017

  106. [114]

    Sendera, M

    M. Sendera, M. Kim, S. Mittal, P. Lemos, L. Scimeca, J. Rector-Brooks, A. Adam, Y . Bengio, and N. Malkin. Improved off-policy training of diffusion samplers.Advances in Neural Information Processing Systems, 37:81016–81045, 2024

  107. [115]

    Seong, S

    K. Seong, S. Park, S. Kim, W. Y . Kim, and S. Ahn. Transition path sampling with improved off-policy training of diffusion path samplers.arXiv preprint arXiv:2405.19961, 2024

  108. [116]

    Z. Shi, L. Yu, T. Xie, and C. Zhang. Diffusion-PINN sampler.arXiv preprint arXiv:2410.15336, 2024

  109. [117]

    A. N. Singh, A. Das, and D. T. Limmer. Variational path sampling of rare dynamical events. Annual Review of Physical Chemistry, 76, 2025

  110. [118]

    A. N. Singh and D. T. Limmer. Variational deep learning of equilibrium transition path ensembles.The Journal of Chemical Physics, 159(2), 2023. 16

  111. [119]

    J. Sun, J. Berner, L. Richter, M. Zeinhofer, J. Müller, K. Azizzadenesheli, and A. Anand- kumar. Dynamical measure transport and neural PDE solvers for sampling.arXiv preprint arXiv:2407.07873, 2024

  112. [120]

    Y . Sun, D. Wierstra, T. Schaul, and J. Schmidhuber. Efficient natural evolution strategies. In Proceedings of the 11th Annual conference on Genetic and evolutionary computation, pages 539–546, 2009

  113. [121]

    S. Syed, A. Bouchard-Côté, K. Chern, and A. Doucet. Optimised annealed sequential Monte Carlo samplers.arXiv preprint arXiv:2408.12057, 2024

  114. [122]

    C. B. Tan, A. J. Bose, C. Lin, L. Klein, M. M. Bronstein, and A. Tong. Scalable equilibrium sampling with sequential Boltzmann generators.arXiv preprint arXiv:2502.18462, 2025

  115. [123]

    H. Y . Tan, S. Osher, and W. Li. Noise-free sampling algorithms via regularized Wasserstein proximals.arXiv preprint arXiv:2308.14945, 2023

  116. [124]

    Tancik, P

    M. Tancik, P. Srinivasan, B. Mildenhall, S. Fridovich-Keil, N. Raghavan, U. Singhal, R. Ra- mamoorthi, J. Barron, and R. Ng. Fourier features let networks learn high frequency functions in low dimensional domains.Advances in neural information processing systems, 33:7537– 7547, 2020

  117. [125]

    Thalmeier, H

    D. Thalmeier, H. J. Kappen, S. Totaro, and V . Gómez. Adaptive smoothing for path integral control.Journal of Machine Learning Research, 21(191):1–37, 2020

  118. [126]

    A. Thin, N. Kotelevskii, A. Durmus, E. Moulines, M. Panov, and A. Doucet. Monte Carlo variational auto-encoders. InInternational Conference on Machine Learning, 2021

  119. [127]

    Tzen and M

    B. Tzen and M. Raginsky. Theoretical guarantees for sampling and inference in generative models with latent diffusions. InConference on Learning Theory, pages 3084–3114. PMLR, 2019

  120. [128]

    Uehara, Y

    M. Uehara, Y . Zhao, K. Black, E. Hajiramezanali, G. Scalia, N. L. Diamant, A. M. Tseng, T. Biancalani, and S. Levine. Fine-tuning of continuous-time diffusion models as entropy- regularized control.arXiv preprint arXiv:2402.15194, 2024

  121. [129]

    Van Handel

    R. Van Handel. Stochastic calculus, filtering, and stochastic control.Course notes., URL http://www. princeton. edu/rvan/acm217/ACM217. pdf, 14, 2007

  122. [130]

    Vanden-Eijnden et al

    E. Vanden-Eijnden et al. Transition-path theory and path-finding algorithms for the study of rare events.Annual review of physical chemistry, 61:391–420, 2010

  123. [131]

    Vargas, W

    F. Vargas, W. Grathwohl, and A. Doucet. Denoising diffusion samplers.arXiv preprint arXiv:2302.13834, 2023

  124. [132]

    Vargas, A

    F. Vargas, A. Ovsianas, D. Fernandes, M. Girolami, N. D. Lawrence, and N. Nüsken. Bayesian learning via neural Schrödinger–Föllmer flows.Statistics and Computing, 33(1):3, 2023

  125. [133]

    Vargas, S

    F. Vargas, S. Padhy, D. Blessing, and N. Nüsken. Transport meets variational inference: Controlled Monte Carlo diffusions. InThe Twelfth International Conference on Learning Representations, 2024

  126. [134]

    Venkatraman, M

    S. Venkatraman, M. Jain, L. Scimeca, M. Kim, M. Sendera, M. Hasan, L. Rowe, S. Mittal, P. Lemos, E. Bengio, et al. Amortizing intractable inference in diffusion models for vision, language, and control.arXiv preprint arXiv:2405.20971, 2024

  127. [135]

    von Klitzing, D

    C. von Klitzing, D. Blessing, H. Schopmans, P. Friederich, and G. Neumann. Learning boltzmann generators via constrained mass transport.arXiv preprint arXiv:2510.18460, 2025

  128. [136]

    C. Wang, X. Zhang, K. Cui, W. Zhao, Y . Guan, and T. Yu. Importance weighted score matching for diffusion samplers with enhanced mode coverage.arXiv preprint arXiv:2505.19431, 2025

  129. [137]

    Wierstra, T

    D. Wierstra, T. Schaul, T. Glasmachers, Y . Sun, J. Peters, and J. Schmidhuber. Natural evolution strategies.The Journal of Machine Learning Research, 15(1):949–980, 2014. 17

  130. [138]

    H. Wu, J. Köhler, and F. Noé. Stochastic normalizing flows.Advances in neural information processing systems, 33:5933–5944, 2020

  131. [139]

    L. Wu, Y . Han, C. A. Naesseth, and J. P. Cunningham. Reverse diffusion Sequential Monte Carlo samplers.arXiv preprint arXiv:2508.05926, 2025

  132. [140]

    X. Wu, Y . Hao, K. Sun, Y . Chen, F. Zhu, R. Zhao, and H. Li. Human preference score v2: A solid benchmark for evaluating human preferences of text-to-image synthesis.arXiv preprint arXiv:2306.09341, 2023

  133. [141]

    Y . Wu, E. Mansimov, R. B. Grosse, S. Liao, and J. Ba. Scalable trust-region method for deep reinforcement learning using Kronecker-factored approximation.Advances in neural information processing systems, 30, 2017

  134. [142]

    H. Xu, J. Xuan, G. Zhang, and J. Lu. Trust region policy optimization via entropy regularization for kullback–leibler divergence constraint.Neurocomputing, 589:127716, 2024

  135. [143]

    J. Xu, X. Liu, Y . Wu, Y . Tong, Q. Li, M. Ding, J. Tang, and Y . Dong. Imagereward: Learning and evaluating human preferences for text-to-image generation. InThirty-seventh Conference on Neural Information Processing Systems, 2023

  136. [144]

    J. Yan, H. Touchette, and G. M. Rotskoff. Learning nonequilibrium control forces to character- ize dynamical phase transitions.Physical Review E, 105(2):024115, 2022

  137. [145]

    T.-Y . Yang, J. Rosca, K. Narasimhan, and P. J. Ramadge. Projection-based constrained policy optimization.arXiv preprint arXiv:2010.03152, 2020

  138. [146]

    S. Yoon, H. Hwang, H. Jeong, D. K. Shin, C.-S. Park, S. Kweon, and F. C. Park. Value gradient sampler: Sampling as sequential decision making.arXiv preprint arXiv:2502.13280, 2025

  139. [147]

    Zhang, R

    D. Zhang, R. T. Chen, C.-H. Liu, A. Courville, and Y . Bengio. Diffusion generative flow samplers: Improving learning signals through partial trajectory optimization. InThe Twelfth International Conference on Learning Representations, 2024

  140. [148]

    Zhang, Y

    D. Zhang, Y . Zhang, J. Gu, R. Zhang, J. Susskind, N. Jaitly, and S. Zhai. Improving GFlowNets for text-to-image diffusion alignment.arXiv preprint arXiv:2406.00633, 2024

  141. [149]

    Zhang, K

    G. Zhang, K. Hsu, J. Li, C. Finn, and R. Grosse. Differentiable annealed importance sampling and the perils of gradient noise. InAdvances in Neural Information Processing Systems, 2021

  142. [150]

    Zhang, P

    L. Zhang, P. Potaptchik, G. Deligiannidis, A. Doucet, H.-D. Dau, and S. Syed. Generalised parallel tempering: Flexible replica exchange via flows and diffusions. InFrontiers in Proba- bilistic Inference: Learning meets Sampling, 2025

  143. [151]

    Zhang, P

    L. Zhang, P. Potaptchik, J. He, Y . Du, A. Doucet, F. Vargas, H.-D. Dau, and S. Syed. Acceler- ated parallel tempering via neural transports.arXiv preprint arXiv:2502.10328, 2025

  144. [152]

    Zhang and Y

    Q. Zhang and Y . Chen. Path Integral Sampler: a stochastic control approach for sampling. In International Conference on Learning Representations, 2022

  145. [153]

    Zhang, H

    W. Zhang, H. Wang, C. Hartmann, M. Weber, and C. Schütte. Applications of the cross-entropy method to importance sampling and optimal control of diffusions.SIAM Journal on Scientific Computing, 36(6):A2654–A2672, 2014

  146. [154]

    Zhang, L

    X. Zhang, L. Wang, J. Helwig, Y . Luo, C. Fu, Y . Xie, M. Liu, Y . Lin, Z. Xu, K. Yan, et al. Artificial intelligence for science in quantum, atomistic, and continuum systems.arXiv preprint arXiv:2307.08423, 2023

  147. [155]

    M. Zhou, J. Han, and J. Lu. Actor-critic method for high dimensional static Hamilton–Jacobi– Bellman partial differential equations based on neural networks.SIAM Journal on Scientific Computing, 43(6):A4043–A4066, 2021

  148. [156]

    reward hacking

    Y . Zhu, W. Guo, J. Choi, G.-H. Liu, Y . Chen, and M. Tao. Mdns: Masked diffusion neural sampler via stochastic optimal control, 2025. 18 Appendix A Assumptions and auxiliary results 21 A.1 Additional notation . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . ....

  149. [157]

    [84, 148] consider alternative algorithms that learn the value functions

    considers GRPO for flow matching fine-tuning. [84, 148] consider alternative algorithms that learn the value functions. Diffusion-based sampling from unnormalized densities.Early work on sampling from unnormal- ized densities based on a Schrödinger-Föllmer diffusions dates bac...

  150. [158]

    (Connection between solution and value function) The solution can be written as u∗ =�σ ⊤�V

  151. [159]

    the uncontrolled path measurePsatisfies dQ dP (X) = e−W(X,0) Z(X 0) withZ(X 0) =E e−W(X,0) |X0 .(38)

    (Optimal change of measure) The Radon-Nikodym derivative of the optimal path measure Q w.r.t. the uncontrolled path measurePsatisfies dQ dP (X) = e−W(X,0) Z(X 0) withZ(X 0) =E e−W(X,0) |X0 .(38)

  152. [160]

    (PDE for value function) The value function V is the solution to the Hamilton-Jacobi-Bellman (HJB) equation (∂t +L)V(x, t)− 1 2 ∥(σ⊤∇V)(x, t)∥2 +f(x, t) = 0, V(x, T) =g(x),(39) 24 where L:= 1 2 � d i,j=1(σσ ⊤)ij∂x� ∂x� + � d i=1 bi∂x� denotes the infinitesimal generator of the...

  153. [161]

    log dP u��� dP u (X u��� ) 2# − E log dP u��� dP u (X u��� ) 2 (94b) =E

    (Estimator for value function) For every (x, t)�� d �[0, T] the value function can be written as V(x, t) =�log� � e−W(X,t)��Xt =x � ,whereXis the solution of the uncontrolled SDE in(34). Combining the expressions for u∗ and V in Thm. D.1, we directly obtain the path integral r...

Pith tools

Reviewed August 5, 2026 · model on record in the stance chip above.