REVIEW 3 major objections 5 minor 85 references
Time-ordered free energy in correlated quantum systems: An agentic approach
T0 review · 3 major / 5 minor · reviewed 2026-08-15 · deepseek-v4-flash
Pith's one-line read The time-ordered free energy of a quantum sequence is attainable by a belief-based dynamic-programming agent, and the loss from causality is exactly an entropy gap the paper calls causal dissipation.
desk verdict A genuinely new measure and method for online quantum work extraction, with a load-bearing but fixable overclaim in Theorem 1 about discretization exactness. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The load-bearing object is the belief state, the posterior distribution over the hidden states of the hidden Markov model conditioned on the agent's action-observation history. Because this belief is a sufficient statistic for the past, the original history-dependent strategy can be replaced by a belief-dependent policy without loss of expected work. The paper then applies backward dynamic programming over a finite discretization of the continuous belief simplex, using the reward formula for expected work as the non-equilibrium free energy of the expected state minus a relative-entropy mismatch, and it shows the action search can be narrowed to eigenbases. The final identity that carries the thermodynamic conclusion is the entropy formula for causal dissipation, which connects the work deficit to the Shannon entropy of the work record and the von Neumann entropy of the final reduced state.
What would settle it
For a fixed perturbed-coin source, compare the dynamic-programming value on K_DP at increasing grid resolutions with a dense-grid or analytic optimum over the continuous belief simplex; if the DP value converges to a number strictly below the continuous optimum for some transition probability, state overlap, and horizon L, then the exact equivalence claimed in Theorem 1 fails.
Extended reading notes
Core claim
The central claim is that the TOFE, defined as the maximum over all causal strategies of the expected cumulative extracted work, is not merely a formal benchmark: it is achievable by a policy that maps the agent's Bayesian belief about the hidden Markov source directly to an extraction action, and this optimal policy can be computed by backward dynamic programming in O(L) time for fixed action and belief sets. The paper further proves that the work deficit of the TOFE relative to the global non-equilibrium free energy is the causal dissipation, the minimum over policies of the accumulated Shannon entropy of observed work outcomes plus the von Neumann entropy of the final conditional state minus the entropy of the entire multi-time state. Operationally, sequential work extraction is equivalent to sequential quantum measurement, so the deficit is exactly the entropy increase produced by measuring the stream one system at a time. The paper illustrates the mechanism on a perturbed-coin qubit source, where the optimal policy sacrifices immediate work to buy predictive information when the source is predictable but not classical.
Load-bearing premise
The load-bearing premise is that maximizing over the finite discretized belief set K_DP gives exactly the same value as maximizing over the continuous belief simplex in the definition of the TOFE, and the paper does not prove a finite-discretization error bound.
Editorial extensions
If this is right
- A causally constrained agent with no persistent quantum memory can provably attain the TOFE, making the benchmark operationally meaningful for online energy harvesters.
- Maximizing long-term work generally requires deliberately choosing mismatched protocols at early steps to gain predictive information; greedy local optimization is suboptimal in predictable but non-classical regimes.
- The gap between unconstrained and causal work equals causal dissipation, and for two time steps it reduces to quantum discord, so the thermodynamic cost of causality is an information-theoretic quantity.
- Causal dissipation can have zero asymptotic rate even when finite-time dissipation is nonzero: in deterministic limits the rate gap vanishes.
- The algorithm's runtime is O(L) for fixed action and belief sets, avoiding exponential scaling in the sequence length.
Reading between the lines
- If the discretization gap in Theorem 1 can be quantified, the same dynamic-programming scheme would give certified lower bounds on the TOFE rather than merely approximate values.
- The L=2 reduction of causal dissipation to quantum discord suggests that, for longer sequences, causal dissipation may be the multi-time analogue of measurement-induced disturbance; testing whether it matches a known multi-partite discord would link these results to the broader discord literature.
- Because the TOFE is defined relative to a fixed action set, a natural next question is how the maximum grows when the action set is enlarged or when the agent is allowed a bounded quantum memory; the paper does not analyze that trade-off.
- The non-degenerate Hamiltonian case is left with an unresolved control cost, so the clean entropy identity for causal dissipation would only survive if that control cost is accounted for elsewhere.
Signed reviews
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper studies sequential work extraction from temporally correlated quantum states emitted by a hidden Markov source, under an online agent constrained by causality and lacking persistent quantum memory. It defines the time-ordered free energy (TOFE) as the maximum cumulative expected work over causal strategies, proposes a dynamic programming reduction to a belief-state MDP, and proves an identity expressing the work deficit relative to the global nonequilibrium free energy as a causal dissipation in terms of Shannon and von Neumann entropies. The framework is illustrated on a perturbed-coin model, with numerical evidence for finite horizons L=3 and L=4. The main derivations are self-contained: belief-state sufficiency, backward induction for the finite discretized MDP, an entropy-balance argument for the causal dissipation identity, and an epsilon-argument showing that belief memory updates can be implemented with arbitrarily small thermodynamic cost.
Significance. If the central claims hold, the paper provides an operational, causally constrained notion of free energy for temporal quantum sequences, a linear-time method (for fixed discretization) to compute optimal online extraction policies, and a quantitative information-theoretic measure of the cost of causality. The causal dissipation identity usefully extends discord-like quantities to multi-time quantum processes and connects thermodynamic resource theories with POMDP and computational-mechanics methods. The paper also contains a strong technical contribution in Appendix J, where near-reversible memory updates are shown to be implementable with arbitrarily small work penalty. The main weakness is the gap between the continuous TOFE definition and the finite-grid algorithm, which currently makes the advertised 'provably optimal' claim and the exact causal dissipation identity overstated.
major comments (3)
- [Policy optimization / Theorem 1 / Algorithm 1] The second sentence of Theorem 1 overstates what is proved. Eq. (2) maximizes over all causal strategies with a continuous belief simplex and a continuous action set, while Algorithm 1 solves a finite MDP over N grid points K_DP and M actions. Appendix D proves backward induction only for this finite discretized MDP; the cited convergence results [37,38] are asymptotic approximation guarantees, not exactness certificates for finite N. Algorithm 1 also does not specify how a Bayesian update that lands outside K_DP is represented (e.g., projection or interpolation). In general, the computed value is a lower bound on the TOFE of Eq. (2), not the maximum itself, so the reported f_TO, the hierarchy (10), and the numerical checks in Fig. 5 concern the discretized quantity until an explicit error bound is supplied. This gap also undermines the abstract's claim of a 'provably optimal agent strategy.'
- [Causal dissipation (Theorem 2, Eq. (13)) and Fig. 5] The exact causal dissipation identity inherits the same discretization problem. Eq. (13) is stated as an equality involving the continuous TOFE, but the minimization over policies Lambda is implemented in practice over the finite grid of Algorithm 1; if the grid misses the optimal belief trajectory, the entropy difference in Eq. (13) is not the true delta of Eq. (11). The numerical confirmation in Fig. 5 compares a simulated work deficit obtained with a 300-point grid and 300 actions against a causal dissipation obtained by differential evolution, so it does not resolve the exactness question. The theorem needs either a finite-grid error bound for both sides of Eq. (13) or a reformulation that makes the discretized object explicit.
- [Abstract and complexity claims] The statement that the method has time complexity O(L) is only valid for fixed discretization sizes N and M. Since the algorithm does not bound the error relative to the continuous TOFE, the sizes N and M may need to grow with L or with the desired accuracy, and no complexity estimate in terms of N, M, and L is given. The abstract's 'time complexity that scales linearly with sequence length' is therefore potentially misleading; please qualify it as the per-step cost for a fixed discretized MDP and state the dependence on grid size.
minor comments (5)
- [Eq. (2)] Please replace 'max' in Eq. (2) by 'sup' if the maximum may not be attained over the continuous strategy space, or justify attainment; otherwise the definition is formally problematic.
- [Appendix G] The claim that equal diagonal entries of xi_k in two different bases imply identical future statistics is insufficient, because future statistics depend on the diagonal entries of each sigma(x), not only on the mixture xi_k. Since Theorem 3 fixes a basis, this sentence is unnecessary for the proof and should be corrected or removed.
- [Algorithm 1] Please specify the projection or rounding rule for updated beliefs not contained in K_DP, or use the continuous belief update with interpolation, so that the transition probabilities in the finite-state DP are well defined.
- [Fig. 5 and Appendix I.3] Please report quantitative agreement (e.g., maximum absolute difference between the simulated deficit and the causal dissipation) and include a convergence check over the discretization sizes N and M, since both plotted quantities are approximate.
- [General formatting] Minor typos and formatting: 'free energy(TOFE)' is missing a space, and the axis labels in Figs. 2, 3, and 5 appear garbled ('0 1 /2 1'); please correct these in revision.
Circularity Check
No circular derivation; the DP optimality and causal-dissipation theorems are self-contained, with only minor non-load-bearing self-citation.
full rationale
The derivation chain is not circular. Eq. (2) defines the TOFE as a maximization over causal strategies; Appendix C proves, rather than assumes, that the belief state is a sufficient statistic and that maximizing over belief-dependent policies reproduces the same value, yielding Eq. (C21). Appendix D proves backward-induction optimality for the finite-state MDP by a standard interchange of maximum and expectation, which is a theorem and not a restatement of Eq. (2). Theorem 2's causal-dissipation identity is derived from Eq. (6), Eq. (12), and Corollary 3.1, where the trace identity tr[~rho ln rho*] = -H(W|h) follows from the relative-entropy minimization in Theorem 3; no parameter is fitted to a target outcome, and no equation is defined in terms of the quantity it is meant to predict. The only author-overlap citations are [20] for the LO baseline and the action-space parametrization, neither of which carries the central DP or dissipation proofs; the convergence citations [37,38] are external POMDP approximation results. The finite-discretization gap in Theorem 1 - Algorithm 1 solves a grid MDP while Eq. (2) optimizes over a continuous strategy space - is a genuine correctness or approximation concern, since the cited results give convergence guarantees rather than finite-N exactness, but it is not a circular reduction: the computed value is a lower bound, not a fitted surrogate presented as a prediction. Overall, the central claims have independent mathematical content, so the score reflects only the minor, non-load-bearing self-citation.
Assumptions & free parameters
free parameters (2)
- Belief discretization size N =
300 intervals
- Action discretization size M =
300 rotation angles over [0,2pi)
assumptions (4)
- domain assumption Thermal operations resource theory with a degenerate system Hamiltonian
- domain assumption Rank-1 work outcomes identify the projected eigenstate
- domain assumption Full prior knowledge of the hidden Markov source model
- ad hoc to paper Finite discretization of the belief simplex inherits the continuous optimal value
Cite this review
Pith. "Pith review of Time-ordered free energy in correlated quantum systems: An agentic approach." pith.science (2026). https://pith.science/paper/5XXV3JLH
@misc{pith2026260812942,
author = {Pith},
title = {Pith review of: Time-ordered free energy in correlated quantum systems: An agentic approach},
year = {2026},
howpublished = {\url{https://pith.science/paper/5XXV3JLH}},
note = {Machine review of arXiv:2608.12942}
}
read the original abstract
How much work can an agent extract from a temporal sequence of quantum states when it can only operate online under causal constraints---deciding which energy extraction method to use with knowledge of what it has observed before? Here, we study this problem in the context of quantum state sequences that are potentially non-Markovian---generated by some underlying hidden Markov machine that the agent cannot observe. Using techniques from dynamic programming and computational mechanics, we present a method to identify the provably optimal agent strategy, with time complexity that scales linearly with sequence length. This motivates us to introduce the maximum work such agents can extract---time-ordered free energy(TOFE)---as a fundamental measure of free energy available in a temporally correlated quantum system subject to causal considerations.
Figures
Figures from the paper (2 more)
Reference graph
Works this paper leans on
-
[1]
On the decrease of entropy in a thermody- namic system by the intervention of intelligent beings
Leo Szilard. On the decrease of entropy in a thermody- namic system by the intervention of intelligent beings. Behavioral Science, 9(4):301–310, 1964
1964
-
[2]
Thermodynamics of prediction.Physi- cal review letters, 109(12):120604, 2012
Susanne Still, David A Sivak, Anthony J Bell, and Gavin E Crooks. Thermodynamics of prediction.Physi- cal review letters, 109(12):120604, 2012
2012
-
[3]
Maximizing free energy gain
Artemy Kolchinsky, Iman Marvian, Can Gokler, Zi-Wen Liu, Peter Shor, Oles Shtanko, Kevin Thompson, David Wolpert, and Seth Lloyd. Maximizing free energy gain. Entropy, 27(1):91, 2025
2025
-
[4]
Thermodynamics+ natural selection= bayesian inference.arXiv preprint arXiv:2511.17641, 2025
Seth Lloyd. Thermodynamics+ natural selection= bayesian inference.arXiv preprint arXiv:2511.17641, 2025
-
[5]
Work and information processing in a solvable model of maxwell’s demon.Proceedings of the National Academy of Sciences, 109(29):11641–11645, 2012
Dibyendu Mandal and Christopher Jarzynski. Work and information processing in a solvable model of maxwell’s demon.Proceedings of the National Academy of Sciences, 109(29):11641–11645, 2012
2012
-
[6]
Identifying functional thermodynamics in autonomous maxwellian ratchets.New Journal of Physics, 18(2):023049, 2016
Alexander B Boyd, Dibyendu Mandal, and James P Crutchfield. Identifying functional thermodynamics in autonomous maxwellian ratchets.New Journal of Physics, 18(2):023049, 2016
2016
-
[7]
Alexander B Boyd, Dibyendu Mandal, and James P Crutchfield. Leveraging environmental correlations: The thermodynamics of requisite variety.Journal of Statisti- cal Physics, 167(6):1555–1585, 2017
work page 2017
-
[8]
Thermodynamics of complexity and pat- tern manipulation.Physical Review E, 95(4):042140, 2017
Andrew JP Garner, Jayne Thompson, Vlatko Vedral, and Mile Gu. Thermodynamics of complexity and pat- tern manipulation.Physical Review E, 95(4):042140, 2017
work page 2017
Show all 85 references
-
[9]
Thermodynamics of modularity: Structural costs beyond the landauer bound.Physical Review X, 8(3):031036, 2018
Alexander B Boyd, Dibyendu Mandal, and James P Crutchfield. Thermodynamics of modularity: Structural costs beyond the landauer bound.Physical Review X, 8(3):031036, 2018
2018
-
[10]
The fundamental thermodynamic bounds on finite models.Chaos: An Interdisciplinary Journal of Nonlinear Science, 31(6), 2021
Andrew JP Garner. The fundamental thermodynamic bounds on finite models.Chaos: An Interdisciplinary Journal of Nonlinear Science, 31(6), 2021
2021
-
[11]
Thermodynamic overfitting and general- ization: energetics of predictive intelligence.New Journal of Physics, 27(6):063901, 2025
Alexander B Boyd, James P Crutchfield, Mile Gu, and Felix C Binder. Thermodynamic overfitting and general- ization: energetics of predictive intelligence.New Journal of Physics, 27(6):063901, 2025
2025
-
[12]
Maximal work extraction from finite quantum systems.Europhysics Letters, 67(4):565, 2004
Armen E Allahverdyan, Roger Balian, and Th M Nieuwenhuizen. Maximal work extraction from finite quantum systems.Europhysics Letters, 67(4):565, 2004
2004
-
[13]
Truly work-like work extraction via a single-shot analysis.Nature communications, 4(1):1925, 2013
Johan ˚Aberg. Truly work-like work extraction via a single-shot analysis.Nature communications, 4(1):1925, 2013
1925
-
[14]
Resource theory of quantum states out of thermal equi- librium.Physical review letters, 111(25):250404, 2013
Fernando GSL Brandao, Micha l Horodecki, Jonathan Oppenheim, Joseph M Renes, and Robert W Spekkens. Resource theory of quantum states out of thermal equi- librium.Physical review letters, 111(25):250404, 2013
2013
-
[15]
Fundamen- tal limitations for quantum and nanoscale thermodynam- ics.Nature communications, 4(1):2059, 2013
Micha l Horodecki and Jonathan Oppenheim. Fundamen- tal limitations for quantum and nanoscale thermodynam- ics.Nature communications, 4(1):2059, 2013
2013
-
[16]
Work extraction and thermodynamics for individual quantum systems.Nature communications, 5(1):4185, 2014
Paul Skrzypczyk, Anthony J Short, and Sandu Popescu. Work extraction and thermodynamics for individual quantum systems.Nature communications, 5(1):4185, 2014
2014
-
[17]
Work extraction from unknown quantum sources.Physical Re- view Letters, 130(21):210401, 2023
Dominik ˇSafr´ anek, Dario Rosa, and Felix C Binder. Work extraction from unknown quantum sources.Physical Re- view Letters, 130(21):210401, 2023
2023
-
[18]
Optical quantum memory.Nature photonics, 3(12):706–714, 2009
Alexander I Lvovsky, Barry C Sanders, and Wolfgang Tittel. Optical quantum memory.Nature photonics, 3(12):706–714, 2009
2009
-
[19]
Quantum memories: emerging applications and recent advances.Journal of modern optics, 63(20):2005–2028, 2016
Khabat Heshami, Duncan G England, Peter C Humphreys, Philip J Bustard, Victor M Acosta, Joshua Nunn, and Benjamin J Sussman. Quantum memories: emerging applications and recent advances.Journal of modern optics, 63(20):2005–2028, 2016
2005
-
[20]
Engines for predictive work extraction from memoryful quantum stochastic processes.Quan- tum, 7:1203, 2023
Ruo Cheng Huang, Paul M Riechers, Mile Gu, and Varun Narasimhachar. Engines for predictive work extraction from memoryful quantum stochastic processes.Quan- tum, 7:1203, 2023
2023
-
[21]
Quantum thermodynamics: A dynamical viewpoint.Entropy, 15(6):2100–2128, 2013
Ronnie Kosloff. Quantum thermodynamics: A dynamical viewpoint.Entropy, 15(6):2100–2128, 2013
2013
-
[22]
Catalytic coherence.Physical review let- ters, 113(15):150402, 2014
Johan ˚Aberg. Catalytic coherence.Physical review let- ters, 113(15):150402, 2014
2014
-
[23]
The resource theory of informational nonequilibrium in ther- modynamics.Physics Reports, 583:1–58, 2015
Gilad Gour, Markus P M¨ uller, Varun Narasimhachar, Robert W Spekkens, and Nicole Yunger Halpern. The resource theory of informational nonequilibrium in ther- modynamics.Physics Reports, 583:1–58, 2015
2015
-
[24]
Description of quantum coherence in thermodynamic processes requires constraints beyond free energy.Na- ture communications, 6(1):6383, 2015
Matteo Lostaglio, David Jennings, and Terry Rudolph. Description of quantum coherence in thermodynamic processes requires constraints beyond free energy.Na- ture communications, 6(1):6383, 2015
2015
-
[25]
The extraction of work from quantum coherence.New Journal of Physics, 18(2):023045, 2016
Kamil Korzekwa, Matteo Lostaglio, Jonathan Oppen- heim, and David Jennings. The extraction of work from quantum coherence.New Journal of Physics, 18(2):023045, 2016
2016
-
[26]
Thermodynamic resource theories, non-commutativity and maximum entropy principles.New Journal of Physics, 19(4):043008, 2017
Matteo Lostaglio, David Jennings, and Terry Rudolph. Thermodynamic resource theories, non-commutativity and maximum entropy principles.New Journal of Physics, 19(4):043008, 2017
2017
-
[27]
Resource theory for work and heat.Physical Re- view A, 96(5):052112, 2017
Carlo Sparaciari, Jonathan Oppenheim, and Tobias Fritz. Resource theory for work and heat.Physical Re- view A, 96(5):052112, 2017
2017
-
[28]
Initial-state dependence of thermodynamic dissipation for any quantum process
Paul M Riechers and Mile Gu. Initial-state dependence of thermodynamic dissipation for any quantum process. Physical Review E, 103(4):042145, 2021
2021
-
[29]
Quantum state-agnostic work extraction (almost) without dissipation.arXiv preprint arXiv:2505.09456, 2025
Josep Lumbreras, Ruo Cheng Huang, Yanglin Hu, Mile Gu, and Marco Tomamichel. Quantum state-agnostic work extraction (almost) without dissipation.arXiv preprint arXiv:2505.09456, 2025
2025 arXiv
-
[30]
Probability, frequency and reasonable expectation.American journal of physics, 14(1):1–13, 1946
Richard T Cox. Probability, frequency and reasonable expectation.American journal of physics, 14(1):1–13, 1946
1946
-
[31]
Cambridge university press, 2003
Edwin T Jaynes.Probability theory: The logic of science. Cambridge university press, 2003
2003
-
[32]
Exact complexity: The spectral decomposi- tion of intrinsic computation.Physics Letters A, 380(9- 10):998–1002, 2016
James P Crutchfield, Christopher J Ellison, and Paul M Riechers. Exact complexity: The spectral decomposi- tion of intrinsic computation.Physics Letters A, 380(9- 10):998–1002, 2016
2016
-
[33]
Optimal control of markov processes with incomplete state information i.Journal of mathe- matical analysis and applications, 10:174–205, 1965
Karl Johan ˚Astr¨ om. Optimal control of markov processes with incomplete state information i.Journal of mathe- matical analysis and applications, 10:174–205, 1965
1965
-
[34]
The opti- mal control of partially observable markov processes over a finite horizon.Operations research, 21(5):1071–1088, 1973
Richard D Smallwood and Edward J Sondik. The opti- mal control of partially observable markov processes over a finite horizon.Operations research, 21(5):1071–1088, 1973
1973
-
[35]
Die messung quantenmechanischer operatoren.Zeitschrift f¨ ur Physik A Hadrons and nuclei, 133(1):101–108, 1952
Eugene P Wigner. Die messung quantenmechanischer operatoren.Zeitschrift f¨ ur Physik A Hadrons and nuclei, 133(1):101–108, 1952
1952
-
[36]
Measurement of quantum mechanical operators.Physical Review, 120(2):622, 1960
Huzihiro Araki and Mutsuo M Yanase. Measurement of quantum mechanical operators.Physical Review, 120(2):622, 1960. 6
1960
-
[37]
Computationally feasible bounds for partially observed markov decision processes.Operations research, 39(1):162–175, 1991
William S Lovejoy. Computationally feasible bounds for partially observed markov decision processes.Operations research, 39(1):162–175, 1991
1991
-
[38]
Value-function approximations for partially observable markov decision processes.Journal of artificial intelligence research, 13:33–94, 2000
Milos Hauskrecht. Value-function approximations for partially observable markov decision processes.Journal of artificial intelligence research, 13:33–94, 2000
2000
-
[39]
MIT press Cambridge, 1998
Richard S Sutton, Andrew G Barto, et al.Reinforce- ment learning: An introduction, volume 1. MIT press Cambridge, 1998
1998
-
[40]
On the use of non- stationary policies for stationary infinite-horizon markov decision processes.Advances in Neural Information Pro- cessing Systems, 25, 2012
Bruno Scherrer and Boris Lesner. On the use of non- stationary policies for stationary infinite-horizon markov decision processes.Advances in Neural Information Pro- cessing Systems, 25, 2012
2012
-
[41]
Quantum discord and maxwell’s demons.Physical Review A, 67(1):012320, 2003
Wojciech Hubert Zurek. Quantum discord and maxwell’s demons.Physical Review A, 67(1):012320, 2003
2003
-
[42]
Quantum dis- cord, local operations, and maxwell’s demons.Physi- cal Review A—Atomic, Molecular, and Optical Physics, 81(6):062103, 2010
Aharon Brodutch and Daniel R Terno. Quantum dis- cord, local operations, and maxwell’s demons.Physi- cal Review A—Atomic, Molecular, and Optical Physics, 81(6):062103, 2010
2010
-
[43]
Black box work ex- traction and composite hypothesis testing.Physical Re- view Letters, 133(25):250401, 2024
Kaito Watanabe and Ryuji Takagi. Black box work ex- traction and composite hypothesis testing.Physical Re- view Letters, 133(25):250401, 2024
2024
-
[44]
Universal work ex- traction in quantum thermodynamics.Nature Commu- nications, 17(1):1857, 2026
Kaito Watanabe and Ryuji Takagi. Universal work ex- traction in quantum thermodynamics.Nature Commu- nications, 17(1):1857, 2026
2026
-
[45]
Causal asymmetry in a quantum world.Physical Review X, 8(3):031013, 2018
Jayne Thompson, Andrew JP Garner, John R Mahoney, James P Crutchfield, Vlatko Vedral, and Mile Gu. Causal asymmetry in a quantum world.Physical Review X, 8(3):031013, 2018
2018
-
[46]
Causal asymmetry of classical and quantum autonomous agents.arXiv preprint arXiv:2309.13572, 2023
Spiros Kechrimparis, Mile Gu, and Hyukjoon Kwon. Causal asymmetry of classical and quantum autonomous agents.arXiv preprint arXiv:2309.13572, 2023
2023 arXiv
-
[47]
Quantum coherence, time-translation symmetry, and thermodynamics.Physical review X, 5(2):021001, 2015
Matteo Lostaglio, Kamil Korzekwa, David Jennings, and Terry Rudolph. Quantum coherence, time-translation symmetry, and thermodynamics.Physical review X, 5(2):021001, 2015
2015
-
[48]
Woods and Micha l Horodecki
Mischa P. Woods and Micha l Horodecki. Autonomous quantum devices: When are they realizable without ad- ditional thermodynamic costs?Phys. Rev. X, 13:011016, Feb 2023
2023
-
[49]
Inferring statis- tical complexity.Physical review letters, 63(2):105–108, 1989
James P Crutchfield, Karl Young, et al. Inferring statis- tical complexity.Physical review letters, 63(2):105–108, 1989
1989
-
[50]
Shalizi and James P
Cosma R. Shalizi and James P. Crutchfield. Computa- tional Mechanics: Pattern and Prediction, Structure and Simplicity.Journal of Statistical Physics, 104(3-4):817– 879, 2001
2001
-
[51]
John Wiley & Sons, 2014
Martin L Puterman.Markov decision processes: discrete stochastic dynamic programming. John Wiley & Sons, 2014
2014
-
[52]
Prediction, retrodiction, and the amount of information stored in the present.Journal of Statistical Physics, 136(6):1005–1034, 2009
Christopher J Ellison, John R Mahoney, and James P Crutchfield. Prediction, retrodiction, and the amount of information stored in the present.Journal of Statistical Physics, 136(6):1005–1034, 2009
2009
-
[53]
Exact syn- chronization for finite-state sources.Journal of Statistical Physics, 145(5):1181–1201, 2011
Nicholas F Travers and James P Crutchfield. Exact syn- chronization for finite-state sources.Journal of Statistical Physics, 145(5):1181–1201, 2011
2011
-
[54]
Spectral sim- plicity of apparent complexity
Paul M Riechers and James P Crutchfield. Spectral sim- plicity of apparent complexity. ii. exact complexities and complexity spectra.Chaos: An Interdisciplinary Journal of Nonlinear Science, 28(3), 2018
2018
-
[55]
Cite- seer, 1965
Richard Bellman, Robert E Kalaba, et al.Dynamic pro- gramming and modern control theory, volume 81. Cite- seer, 1965
1965
-
[56]
Cambridge university press, 2010
Michael A Nielsen and Isaac L Chuang.Quantum compu- tation and quantum information. Cambridge university press, 2010
2010
-
[57]
Mc- Graw Hill, 2021
Walter Rudin.Principles of mathematical analysis. Mc- Graw Hill, 2021. Appendix A: Energy Non-degenerate Hamiltonian In the main body, we have discussed the application of theρ ∗-ideal protocol for a degenerate Hamiltonian. We mentioned that for a degenerate Hamiltonian, this ...
2021
-
[58]
The agent’s belief state is a sufficient statistic for the entire history of actions and observations
-
[59]
The belief trajectoryK 0:L−1 is a deterministic function of the trajectory over physical realizationsX 1:L, actions A1:L, and observationsW 1:L
-
[60]
The Belief State as a Sufficient Statistic We will first prove that belief states are statistically sufficient for making any prediction. Formally, we want to show that Pr(Xt|W1:t−1,A 1:t−1,µ 0) = Pr(Xt|Kt−1,K 0 =µ 0).(C2) This tells us that any information that the past obser...
-
[61]
max a∈A JL(KL−1,aL) K0 =η (i) # =I L−1 +E
Formal Equivalence While the technique of replacing observation histories with belief states is standard in POMDP literature [34, 51], we provide an explicit derivation below adapted to our thermodynamic framework, directly linking the global multi-time state to the agent’s va...
-
[62]
Non-stationary phase At timet= 2, which is the last step of the process, the value functionV 3 is 0 regardless of the belief state. Hence, the action taken att= 2 for a DP agent will minimize the expected dissipation, a∗ 2(K1 =η (i)) = argmax a(j)∈A − 1 β D(ξη(i)∥ρ(j)) .(F1) T...
-
[63]
The expected distribution of the future is uniform regardless of the current belief
Whenp= 0.5, where the process becomes a purely random process. The expected distribution of the future is uniform regardless of the current belief. 16
-
[64]
Whenr= 0, the classical limit where all states are orthogonal; one can measure along the eigenbasis to obtain perfect knowledge
-
[65]
Whenr= 1, a trivial limit where the process emits identical states; hence, regardless of action, the expectation of the future is the same. Interestingly, when the boundary effect completely disappears even for the finite-horizon case, the optimal policy for all time steps bec...
-
[66]
We first clarify some notational differences
Equivalence between work deficit and causal dissipation In this subsection, we will prove Theorem 2. We first clarify some notational differences. In Eq. (I7), the causal dissipation is defined with respect to ˜HL = (ΠQ1,O 1,···,Π QL,OL), while in Theorem 2, it is defined with...
-
[67]
Mathematical properties of causal dissipation Here, we try to establish some mathematical properties of causal dissipation
-
[68]
The sequence of adaptive measurements/operations can be modeled as a single global unital CPTP dephasing channel Φ acting on the firstL−1 subsystems
Positivity:δ(Q −→1:L)≥0. The sequence of adaptive measurements/operations can be modeled as a single global unital CPTP dephasing channel Φ acting on the firstL−1 subsystems. The von Neumann entropy of the resulting block-diagonal state evaluates exactly to the Shannon entropy...
-
[69]
We then rewrite the entropy rate with this quantity
since both states appear with probability 1/2 ; here,H 2 is the binary entropy. We then rewrite the entropy rate with this quantity. h := lim L→∞ 1 LS(ρ(1:L)) = lim L→∞ χL L + 1 2LS(ρ0) + 1 2LS(ρ1).(I29) Since the Holevo information is bounded, this results in h= lim L→∞ 1 2LS...
-
[70]
5, we demonstrate the agreement between the simulated work deficit and the causal dissipation for both L= 3 andL= 4
Numerical Simulation In Fig. 5, we demonstrate the agreement between the simulated work deficit and the causal dissipation for both L= 3 andL= 4. Note that the lack of comparison for a longer time horizon is strictly a restriction due to the brute force optimization in calcula...
-
[71]
The emitted quantum state can take on either|ϕ 0⟩=|0⟩or|ϕ 1⟩= √r|0⟩+ √1−r|1⟩
-
[72]
Belief states:K={(1/2 +ϵ,1/2−ϵ)|ϵ∈[−1/2,1/2]}; we divide the parameterϵinto 300 equal-sized intervals
-
[73]
After the optimal policy is found, we apply it by initializing the agent at belief stateK 0 =µ 0 =π= (1/2,1/2) of the perturbed coin process in Fig
Actions:Aconsists of all bases spanned by|ϕ 0⟩and|ϕ 1⟩, taking the form ofA={|ψ θ⟩= cos θ 2|0⟩+sin θ 2|1⟩|θ∈ [0,2π)}; we also divide the parameterθinto 300 equal sized intervals. After the optimal policy is found, we apply it by initializing the agent at belief stateK 0 =µ 0 =...
-
[74]
For a chosen basis, define pt,η,i :=⟨ψ t,η,i|ξ η|ψt,η,i⟩.(J7) The optimal target spectrum for this fixed basis isp t,η
Spectral tagging Letξ η be the expected state conditioned on the beliefη. For a chosen basis, define pt,η,i :=⟨ψ t,η,i|ξ η|ψt,η,i⟩.(J7) The optimal target spectrum for this fixed basis isp t,η. The unperturbed optimal target state thus takes the form: ρt,η = dX i=1 pt,η,i|ψt,η...
-
[75]
the probability of every branch conditioned on every emitted stateσ (x)
-
[76]
the Bayesian posterior conditioned on that branch
-
[77]
the probability distribution over future beliefs
-
[78]
Proof.The perturbed and unperturbed target states share exactly the same eigenbasis
the future bases selected by the original policy. Proof.The perturbed and unperturbed target states share exactly the same eigenbasis. Branch probabilities depend exclusively on this basis, independently of the target’s eigenvalues, as shown in Eq. (J2). Therefore, the likelih...
-
[79]
Work penalty for tagging Having established the injectivity of the update map, we must quantify the average work deficit incurred by introducing this tagging perturbation. Recall that the expected work extracted from the expected stateξusing a protocol tailored forρis given by...
-
[80]
It uses the same measurement basis as the original policy at all time steps and beliefs
-
[81]
The induced branch probabilities, Bayesian beliefs, and future basis choices match the untagged policy perfectly
-
[82]
Every memory update can be implemented via a unitary on the memory and current battery at zero additional thermodynamic cost
-
[83]
The memory can be returned to its initial state by sequentially applying the inverse update unitaries in reverse temporal order
-
[84]
Proof.We choose α= 1−exp −β∆ L .(J22) At every time step and belief, construct the tagged spectrum using Lemma 4
Its cumulative expected work extraction is at most∆below the theoretical optimum of the original policy. Proof.We choose α= 1−exp −β∆ L .(J22) At every time step and belief, construct the tagged spectrum using Lemma 4. Points 1 and 2 follow from Lemma 5, while point 3 follows ...
-
[85]
We first cast theNdiscrete belief states asNmutually orthogonal quantum states{|i⟩ M}N i=1, each of which would be mapped to different points on the probability simplex
Operational Justification To clarify how these sequential updates can be executed and eventually uncomputed without incurring a Landauer erasure cost, we construct an explicit operational model. We first cast theNdiscrete belief states asNmutually orthogonal quantum states{|i⟩...
Reviewed August 15, 2026 · model on record in the stance chip above.
Discussion (0). Continue with ORCID to comment.