Pith. sign in

REVIEW 3 major objections 5 minor 85 references

Time-ordered free energy in correlated quantum systems: An agentic approach

T0 review · 3 major / 5 minor · reviewed 2026-08-15 · deepseek-v4-flash

Pith's one-line read The time-ordered free energy of a quantum sequence is attainable by a belief-based dynamic-programming agent, and the loss from causality is exactly an entropy gap the paper calls causal dissipation.

desk verdict A genuinely new measure and method for online quantum work extraction, with a load-bearing but fixable overclaim in Theorem 1 about discretization exactness. read the letter →

arxiv 2608.12942 v1 pith:5XXV3JLH submitted 2026-08-13 quant-ph physics.comp-ph

classification quant-phphysics.comp-ph
keywords time-orderedfreeenergycausaldissipationdynamicprogrammingquantumworkextractionhiddenMarkovmodelsbeliefstatessequentialmeasurementthermodynamicsofprediction
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

This paper tries to establish that a causally constrained agent, one that can act on each quantum system as it arrives and keeps no persistent quantum memory, can still extract the maximum possible sequential work from a correlated quantum state stream. That maximum is captured by a quantity the paper introduces, the time-ordered free energy (TOFE), and the paper claims a dynamic-programming algorithm finds the optimal policy in time linear in the sequence length for any fixed action set. The central result is an exact accounting of what causality costs: the shortfall between the unconstrained non-equilibrium free energy and the TOFE equals a quantity the paper calls causal dissipation, expressed as an entropy difference. A sympathetic reader would care because it turns the intuitive trade-off between immediate energy harvest and predictive information into a provable optimality statement with a concrete computational route.

What carries the argument

The load-bearing object is the belief state, the posterior distribution over the hidden states of the hidden Markov model conditioned on the agent's action-observation history. Because this belief is a sufficient statistic for the past, the original history-dependent strategy can be replaced by a belief-dependent policy without loss of expected work. The paper then applies backward dynamic programming over a finite discretization of the continuous belief simplex, using the reward formula for expected work as the non-equilibrium free energy of the expected state minus a relative-entropy mismatch, and it shows the action search can be narrowed to eigenbases. The final identity that carries the thermodynamic conclusion is the entropy formula for causal dissipation, which connects the work deficit to the Shannon entropy of the work record and the von Neumann entropy of the final reduced state.

What would settle it

For a fixed perturbed-coin source, compare the dynamic-programming value on K_DP at increasing grid resolutions with a dense-grid or analytic optimum over the continuous belief simplex; if the DP value converges to a number strictly below the continuous optimum for some transition probability, state overlap, and horizon L, then the exact equivalence claimed in Theorem 1 fails.

Watch

Extended reading notes

Core claim

The central claim is that the TOFE, defined as the maximum over all causal strategies of the expected cumulative extracted work, is not merely a formal benchmark: it is achievable by a policy that maps the agent's Bayesian belief about the hidden Markov source directly to an extraction action, and this optimal policy can be computed by backward dynamic programming in O(L) time for fixed action and belief sets. The paper further proves that the work deficit of the TOFE relative to the global non-equilibrium free energy is the causal dissipation, the minimum over policies of the accumulated Shannon entropy of observed work outcomes plus the von Neumann entropy of the final conditional state minus the entropy of the entire multi-time state. Operationally, sequential work extraction is equivalent to sequential quantum measurement, so the deficit is exactly the entropy increase produced by measuring the stream one system at a time. The paper illustrates the mechanism on a perturbed-coin qubit source, where the optimal policy sacrifices immediate work to buy predictive information when the source is predictable but not classical.

Load-bearing premise

The load-bearing premise is that maximizing over the finite discretized belief set K_DP gives exactly the same value as maximizing over the continuous belief simplex in the definition of the TOFE, and the paper does not prove a finite-discretization error bound.

Editorial extensions

If this is right

  • A causally constrained agent with no persistent quantum memory can provably attain the TOFE, making the benchmark operationally meaningful for online energy harvesters.
  • Maximizing long-term work generally requires deliberately choosing mismatched protocols at early steps to gain predictive information; greedy local optimization is suboptimal in predictable but non-classical regimes.
  • The gap between unconstrained and causal work equals causal dissipation, and for two time steps it reduces to quantum discord, so the thermodynamic cost of causality is an information-theoretic quantity.
  • Causal dissipation can have zero asymptotic rate even when finite-time dissipation is nonzero: in deterministic limits the rate gap vanishes.
  • The algorithm's runtime is O(L) for fixed action and belief sets, avoiding exponential scaling in the sequence length.

Reading between the lines

Editorial extensions of the paper, not claims the author makes directly.

  • If the discretization gap in Theorem 1 can be quantified, the same dynamic-programming scheme would give certified lower bounds on the TOFE rather than merely approximate values.
  • The L=2 reduction of causal dissipation to quantum discord suggests that, for longer sequences, causal dissipation may be the multi-time analogue of measurement-induced disturbance; testing whether it matches a known multi-partite discord would link these results to the broader discord literature.
  • Because the TOFE is defined relative to a fixed action set, a natural next question is how the maximum grows when the action set is enlarged or when the agent is allowed a bounded quantum memory; the paper does not analyze that trade-off.
  • The non-degenerate Hamiltonian case is left with an unresolved control cost, so the clean entropy identity for causal dissipation would only survive if that control cost is accounted for elsewhere.
Share X Bluesky LinkedIn Reddit HN

Signed reviews

No signed human review yet.

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, and a circularity audit.

Referee Report

3 major / 5 minor

Summary. The paper studies sequential work extraction from temporally correlated quantum states emitted by a hidden Markov source, under an online agent constrained by causality and lacking persistent quantum memory. It defines the time-ordered free energy (TOFE) as the maximum cumulative expected work over causal strategies, proposes a dynamic programming reduction to a belief-state MDP, and proves an identity expressing the work deficit relative to the global nonequilibrium free energy as a causal dissipation in terms of Shannon and von Neumann entropies. The framework is illustrated on a perturbed-coin model, with numerical evidence for finite horizons L=3 and L=4. The main derivations are self-contained: belief-state sufficiency, backward induction for the finite discretized MDP, an entropy-balance argument for the causal dissipation identity, and an epsilon-argument showing that belief memory updates can be implemented with arbitrarily small thermodynamic cost.

Significance. If the central claims hold, the paper provides an operational, causally constrained notion of free energy for temporal quantum sequences, a linear-time method (for fixed discretization) to compute optimal online extraction policies, and a quantitative information-theoretic measure of the cost of causality. The causal dissipation identity usefully extends discord-like quantities to multi-time quantum processes and connects thermodynamic resource theories with POMDP and computational-mechanics methods. The paper also contains a strong technical contribution in Appendix J, where near-reversible memory updates are shown to be implementable with arbitrarily small work penalty. The main weakness is the gap between the continuous TOFE definition and the finite-grid algorithm, which currently makes the advertised 'provably optimal' claim and the exact causal dissipation identity overstated.

major comments (3)
  1. [Policy optimization / Theorem 1 / Algorithm 1] The second sentence of Theorem 1 overstates what is proved. Eq. (2) maximizes over all causal strategies with a continuous belief simplex and a continuous action set, while Algorithm 1 solves a finite MDP over N grid points K_DP and M actions. Appendix D proves backward induction only for this finite discretized MDP; the cited convergence results [37,38] are asymptotic approximation guarantees, not exactness certificates for finite N. Algorithm 1 also does not specify how a Bayesian update that lands outside K_DP is represented (e.g., projection or interpolation). In general, the computed value is a lower bound on the TOFE of Eq. (2), not the maximum itself, so the reported f_TO, the hierarchy (10), and the numerical checks in Fig. 5 concern the discretized quantity until an explicit error bound is supplied. This gap also undermines the abstract's claim of a 'provably optimal agent strategy.'
  2. [Causal dissipation (Theorem 2, Eq. (13)) and Fig. 5] The exact causal dissipation identity inherits the same discretization problem. Eq. (13) is stated as an equality involving the continuous TOFE, but the minimization over policies Lambda is implemented in practice over the finite grid of Algorithm 1; if the grid misses the optimal belief trajectory, the entropy difference in Eq. (13) is not the true delta of Eq. (11). The numerical confirmation in Fig. 5 compares a simulated work deficit obtained with a 300-point grid and 300 actions against a causal dissipation obtained by differential evolution, so it does not resolve the exactness question. The theorem needs either a finite-grid error bound for both sides of Eq. (13) or a reformulation that makes the discretized object explicit.
  3. [Abstract and complexity claims] The statement that the method has time complexity O(L) is only valid for fixed discretization sizes N and M. Since the algorithm does not bound the error relative to the continuous TOFE, the sizes N and M may need to grow with L or with the desired accuracy, and no complexity estimate in terms of N, M, and L is given. The abstract's 'time complexity that scales linearly with sequence length' is therefore potentially misleading; please qualify it as the per-step cost for a fixed discretized MDP and state the dependence on grid size.
minor comments (5)
  1. [Eq. (2)] Please replace 'max' in Eq. (2) by 'sup' if the maximum may not be attained over the continuous strategy space, or justify attainment; otherwise the definition is formally problematic.
  2. [Appendix G] The claim that equal diagonal entries of xi_k in two different bases imply identical future statistics is insufficient, because future statistics depend on the diagonal entries of each sigma(x), not only on the mixture xi_k. Since Theorem 3 fixes a basis, this sentence is unnecessary for the proof and should be corrected or removed.
  3. [Algorithm 1] Please specify the projection or rounding rule for updated beliefs not contained in K_DP, or use the continuous belief update with interpolation, so that the transition probabilities in the finite-state DP are well defined.
  4. [Fig. 5 and Appendix I.3] Please report quantitative agreement (e.g., maximum absolute difference between the simulated deficit and the causal dissipation) and include a convergence check over the discretization sizes N and M, since both plotted quantities are approximate.
  5. [General formatting] Minor typos and formatting: 'free energy(TOFE)' is missing a space, and the axis labels in Figs. 2, 3, and 5 appear garbled ('0 1 /2 1'); please correct these in revision.

Circularity Check

0 steps flagged · score 2.0 of 10

No circular derivation; the DP optimality and causal-dissipation theorems are self-contained, with only minor non-load-bearing self-citation.

full rationale

The derivation chain is not circular. Eq. (2) defines the TOFE as a maximization over causal strategies; Appendix C proves, rather than assumes, that the belief state is a sufficient statistic and that maximizing over belief-dependent policies reproduces the same value, yielding Eq. (C21). Appendix D proves backward-induction optimality for the finite-state MDP by a standard interchange of maximum and expectation, which is a theorem and not a restatement of Eq. (2). Theorem 2's causal-dissipation identity is derived from Eq. (6), Eq. (12), and Corollary 3.1, where the trace identity tr[~rho ln rho*] = -H(W|h) follows from the relative-entropy minimization in Theorem 3; no parameter is fitted to a target outcome, and no equation is defined in terms of the quantity it is meant to predict. The only author-overlap citations are [20] for the LO baseline and the action-space parametrization, neither of which carries the central DP or dissipation proofs; the convergence citations [37,38] are external POMDP approximation results. The finite-discretization gap in Theorem 1 - Algorithm 1 solves a grid MDP while Eq. (2) optimizes over a continuous strategy space - is a genuine correctness or approximation concern, since the cited results give convergence guarantees rather than finite-N exactness, but it is not a circular reduction: the computed value is a lower bound, not a fitted surrogate presented as a prediction. Overall, the central claims have independent mathematical content, so the score reflects only the minor, non-load-bearing self-citation.

Assumptions & free parameters 2 free parameters · 4 assumptions · 0 invented entities

The central claim depends on standard resource-theoretic assumptions, the rank-1 measurement assumption, full HMM knowledge, and the unproven discretization-to-continuum equivalence. The only hand-chosen numerical parameters are the discretization sizes. TOFE and causal dissipation are new defined quantities, not postulated physical entities, so they are not listed as invented entities.

free parameters (2)
  • Belief discretization size N = 300 intervals
    Numerical simulations divide the one-dimensional belief simplex into 300 bins. The DP optimality theorem is stated for a given finite set of belief states, so this discretization controls how close the computed value is to the continuous TOFE, but no error bound is reported.
  • Action discretization size M = 300 rotation angles over [0,2pi)
    Numerical simulations restrict actions to 300 target-state eigenbases. The theorem is stated for a fixed action set, so the computed optimal policy is optimal only within this finite action set, not over all possible density-matrix targets.
assumptions (4)
  • domain assumption Thermal operations resource theory with a degenerate system Hamiltonian
    The framework assumes standard thermal operations and an energy-degenerate H_Q so that every basis is an energy eigenbasis and thermal operations can extract all non-equilibrium free energy. Non-degenerate Hamiltonians are deferred to Appendix A, where control costs remain unresolved.
  • domain assumption Rank-1 work outcomes identify the projected eigenstate
    The main derivations assume each observed work value uniquely identifies the eigenstate onto which the system is projected. Appendix B provides a perturbation argument to restore injectivity, but this is an added assumption required for Bayesian observability.
  • domain assumption Full prior knowledge of the hidden Markov source model
    The DP policy and the Bayesian belief update require exact knowledge of the HMM transition matrices and emitted quantum states. The Discussion explicitly acknowledges this as a key assumption and leaves model learning to future work.
  • ad hoc to paper Finite discretization of the belief simplex inherits the continuous optimal value
    Theorem 1 and Algorithm 1 optimize over a finite set K_DP, while Eq. (2) defines TOFE over all continuous causal strategies. The paper cites asymptotic convergence results for discretized POMDP value functions but gives no finite-N exactness proof, yet the theorem claims mathematical equivalence.

how reviews work

0 comments
Cite this review

Pith. "Pith review of Time-ordered free energy in correlated quantum systems: An agentic approach." pith.science (2026). https://pith.science/paper/5XXV3JLH

@misc{pith2026260812942,
  author       = {Pith},
  title        = {Pith review of: Time-ordered free energy in correlated quantum systems: An agentic approach},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/5XXV3JLH}},
  note         = {Machine review of arXiv:2608.12942}
}
read the original abstract

How much work can an agent extract from a temporal sequence of quantum states when it can only operate online under causal constraints---deciding which energy extraction method to use with knowledge of what it has observed before? Here, we study this problem in the context of quantum state sequences that are potentially non-Markovian---generated by some underlying hidden Markov machine that the agent cannot observe. Using techniques from dynamic programming and computational mechanics, we present a method to identify the provably optimal agent strategy, with time complexity that scales linearly with sequence length. This motivates us to introduce the maximum work such agents can extract---time-ordered free energy(TOFE)---as a fundamental measure of free energy available in a temporally correlated quantum system subject to causal considerations.

Figures

Figures reproduced from arXiv: 2608.12942 by the authors.

Figure 1
Figure 1. (a) The perturbed coin process is an example [PITH_FULL_IMAGE:figures/full_fig_p002_1.png] view at source ↗
Figure 2
Figure 2. Comparison of asymptotic work-extraction [PITH_FULL_IMAGE:figures/full_fig_p004_2.png] view at source ↗
Figure 3
Figure 3. Relation between how much energy an agent is [PITH_FULL_IMAGE:figures/full_fig_p005_3.png] view at source ↗
Figures from the paper (2 more)
Figure 4
Figure 4. Figure 4: Graph of reward and dissipation, conditioned on the belief state [PITH_FULL_IMAGE:figures/full_fig_p016_4.png]
Figure 5
Figure 5. Figure 5: The gap between the non-equilibrium free energy and the TOFE (panels (a) and (c)) shows good [PITH_FULL_IMAGE:figures/full_fig_p025_5.png]

Discussion (0). Continue with ORCID to comment.

Reference graph

Works this paper leans on

85 extracted references · 70 canonical work pages

  1. [1]

    On the decrease of entropy in a thermody- namic system by the intervention of intelligent beings

    Leo Szilard. On the decrease of entropy in a thermody- namic system by the intervention of intelligent beings. Behavioral Science, 9(4):301–310, 1964

  2. [2]

    Thermodynamics of prediction.Physi- cal review letters, 109(12):120604, 2012

    Susanne Still, David A Sivak, Anthony J Bell, and Gavin E Crooks. Thermodynamics of prediction.Physi- cal review letters, 109(12):120604, 2012

  3. [3]

    Maximizing free energy gain

    Artemy Kolchinsky, Iman Marvian, Can Gokler, Zi-Wen Liu, Peter Shor, Oles Shtanko, Kevin Thompson, David Wolpert, and Seth Lloyd. Maximizing free energy gain. Entropy, 27(1):91, 2025

  4. [4]

    Thermodynamics+ natural selection= bayesian inference.arXiv preprint arXiv:2511.17641, 2025

    Seth Lloyd. Thermodynamics+ natural selection= bayesian inference.arXiv preprint arXiv:2511.17641, 2025

  5. [5]

    Work and information processing in a solvable model of maxwell’s demon.Proceedings of the National Academy of Sciences, 109(29):11641–11645, 2012

    Dibyendu Mandal and Christopher Jarzynski. Work and information processing in a solvable model of maxwell’s demon.Proceedings of the National Academy of Sciences, 109(29):11641–11645, 2012

  6. [6]

    Identifying functional thermodynamics in autonomous maxwellian ratchets.New Journal of Physics, 18(2):023049, 2016

    Alexander B Boyd, Dibyendu Mandal, and James P Crutchfield. Identifying functional thermodynamics in autonomous maxwellian ratchets.New Journal of Physics, 18(2):023049, 2016

  7. [7]

    Leveraging environmental correlations: The thermodynamics of requisite variety.Journal of Statisti- cal Physics, 167(6):1555–1585, 2017

    Alexander B Boyd, Dibyendu Mandal, and James P Crutchfield. Leveraging environmental correlations: The thermodynamics of requisite variety.Journal of Statisti- cal Physics, 167(6):1555–1585, 2017

  8. [8]

    Thermodynamics of complexity and pat- tern manipulation.Physical Review E, 95(4):042140, 2017

    Andrew JP Garner, Jayne Thompson, Vlatko Vedral, and Mile Gu. Thermodynamics of complexity and pat- tern manipulation.Physical Review E, 95(4):042140, 2017

Show all 85 references
  1. [9]

    Thermodynamics of modularity: Structural costs beyond the landauer bound.Physical Review X, 8(3):031036, 2018

    Alexander B Boyd, Dibyendu Mandal, and James P Crutchfield. Thermodynamics of modularity: Structural costs beyond the landauer bound.Physical Review X, 8(3):031036, 2018

  2. [10]

    The fundamental thermodynamic bounds on finite models.Chaos: An Interdisciplinary Journal of Nonlinear Science, 31(6), 2021

    Andrew JP Garner. The fundamental thermodynamic bounds on finite models.Chaos: An Interdisciplinary Journal of Nonlinear Science, 31(6), 2021

  3. [11]

    Thermodynamic overfitting and general- ization: energetics of predictive intelligence.New Journal of Physics, 27(6):063901, 2025

    Alexander B Boyd, James P Crutchfield, Mile Gu, and Felix C Binder. Thermodynamic overfitting and general- ization: energetics of predictive intelligence.New Journal of Physics, 27(6):063901, 2025

  4. [12]

    Maximal work extraction from finite quantum systems.Europhysics Letters, 67(4):565, 2004

    Armen E Allahverdyan, Roger Balian, and Th M Nieuwenhuizen. Maximal work extraction from finite quantum systems.Europhysics Letters, 67(4):565, 2004

  5. [13]

    Truly work-like work extraction via a single-shot analysis.Nature communications, 4(1):1925, 2013

    Johan ˚Aberg. Truly work-like work extraction via a single-shot analysis.Nature communications, 4(1):1925, 2013

  6. [14]

    Resource theory of quantum states out of thermal equi- librium.Physical review letters, 111(25):250404, 2013

    Fernando GSL Brandao, Micha l Horodecki, Jonathan Oppenheim, Joseph M Renes, and Robert W Spekkens. Resource theory of quantum states out of thermal equi- librium.Physical review letters, 111(25):250404, 2013

  7. [15]

    Fundamen- tal limitations for quantum and nanoscale thermodynam- ics.Nature communications, 4(1):2059, 2013

    Micha l Horodecki and Jonathan Oppenheim. Fundamen- tal limitations for quantum and nanoscale thermodynam- ics.Nature communications, 4(1):2059, 2013

  8. [16]

    Work extraction and thermodynamics for individual quantum systems.Nature communications, 5(1):4185, 2014

    Paul Skrzypczyk, Anthony J Short, and Sandu Popescu. Work extraction and thermodynamics for individual quantum systems.Nature communications, 5(1):4185, 2014

  9. [17]

    Work extraction from unknown quantum sources.Physical Re- view Letters, 130(21):210401, 2023

    Dominik ˇSafr´ anek, Dario Rosa, and Felix C Binder. Work extraction from unknown quantum sources.Physical Re- view Letters, 130(21):210401, 2023

  10. [18]

    Optical quantum memory.Nature photonics, 3(12):706–714, 2009

    Alexander I Lvovsky, Barry C Sanders, and Wolfgang Tittel. Optical quantum memory.Nature photonics, 3(12):706–714, 2009

  11. [19]

    Quantum memories: emerging applications and recent advances.Journal of modern optics, 63(20):2005–2028, 2016

    Khabat Heshami, Duncan G England, Peter C Humphreys, Philip J Bustard, Victor M Acosta, Joshua Nunn, and Benjamin J Sussman. Quantum memories: emerging applications and recent advances.Journal of modern optics, 63(20):2005–2028, 2016

  12. [20]

    Engines for predictive work extraction from memoryful quantum stochastic processes.Quan- tum, 7:1203, 2023

    Ruo Cheng Huang, Paul M Riechers, Mile Gu, and Varun Narasimhachar. Engines for predictive work extraction from memoryful quantum stochastic processes.Quan- tum, 7:1203, 2023

  13. [21]

    Quantum thermodynamics: A dynamical viewpoint.Entropy, 15(6):2100–2128, 2013

    Ronnie Kosloff. Quantum thermodynamics: A dynamical viewpoint.Entropy, 15(6):2100–2128, 2013

  14. [22]

    Catalytic coherence.Physical review let- ters, 113(15):150402, 2014

    Johan ˚Aberg. Catalytic coherence.Physical review let- ters, 113(15):150402, 2014

  15. [23]

    The resource theory of informational nonequilibrium in ther- modynamics.Physics Reports, 583:1–58, 2015

    Gilad Gour, Markus P M¨ uller, Varun Narasimhachar, Robert W Spekkens, and Nicole Yunger Halpern. The resource theory of informational nonequilibrium in ther- modynamics.Physics Reports, 583:1–58, 2015

  16. [24]

    Description of quantum coherence in thermodynamic processes requires constraints beyond free energy.Na- ture communications, 6(1):6383, 2015

    Matteo Lostaglio, David Jennings, and Terry Rudolph. Description of quantum coherence in thermodynamic processes requires constraints beyond free energy.Na- ture communications, 6(1):6383, 2015

  17. [25]

    The extraction of work from quantum coherence.New Journal of Physics, 18(2):023045, 2016

    Kamil Korzekwa, Matteo Lostaglio, Jonathan Oppen- heim, and David Jennings. The extraction of work from quantum coherence.New Journal of Physics, 18(2):023045, 2016

  18. [26]

    Thermodynamic resource theories, non-commutativity and maximum entropy principles.New Journal of Physics, 19(4):043008, 2017

    Matteo Lostaglio, David Jennings, and Terry Rudolph. Thermodynamic resource theories, non-commutativity and maximum entropy principles.New Journal of Physics, 19(4):043008, 2017

  19. [27]

    Resource theory for work and heat.Physical Re- view A, 96(5):052112, 2017

    Carlo Sparaciari, Jonathan Oppenheim, and Tobias Fritz. Resource theory for work and heat.Physical Re- view A, 96(5):052112, 2017

  20. [28]

    Initial-state dependence of thermodynamic dissipation for any quantum process

    Paul M Riechers and Mile Gu. Initial-state dependence of thermodynamic dissipation for any quantum process. Physical Review E, 103(4):042145, 2021

  21. [29]

    Quantum state-agnostic work extraction (almost) without dissipation.arXiv preprint arXiv:2505.09456, 2025

    Josep Lumbreras, Ruo Cheng Huang, Yanglin Hu, Mile Gu, and Marco Tomamichel. Quantum state-agnostic work extraction (almost) without dissipation.arXiv preprint arXiv:2505.09456, 2025

  22. [30]

    Probability, frequency and reasonable expectation.American journal of physics, 14(1):1–13, 1946

    Richard T Cox. Probability, frequency and reasonable expectation.American journal of physics, 14(1):1–13, 1946

  23. [31]

    Cambridge university press, 2003

    Edwin T Jaynes.Probability theory: The logic of science. Cambridge university press, 2003

  24. [32]

    Exact complexity: The spectral decomposi- tion of intrinsic computation.Physics Letters A, 380(9- 10):998–1002, 2016

    James P Crutchfield, Christopher J Ellison, and Paul M Riechers. Exact complexity: The spectral decomposi- tion of intrinsic computation.Physics Letters A, 380(9- 10):998–1002, 2016

  25. [33]

    Optimal control of markov processes with incomplete state information i.Journal of mathe- matical analysis and applications, 10:174–205, 1965

    Karl Johan ˚Astr¨ om. Optimal control of markov processes with incomplete state information i.Journal of mathe- matical analysis and applications, 10:174–205, 1965

  26. [34]

    The opti- mal control of partially observable markov processes over a finite horizon.Operations research, 21(5):1071–1088, 1973

    Richard D Smallwood and Edward J Sondik. The opti- mal control of partially observable markov processes over a finite horizon.Operations research, 21(5):1071–1088, 1973

  27. [35]

    Die messung quantenmechanischer operatoren.Zeitschrift f¨ ur Physik A Hadrons and nuclei, 133(1):101–108, 1952

    Eugene P Wigner. Die messung quantenmechanischer operatoren.Zeitschrift f¨ ur Physik A Hadrons and nuclei, 133(1):101–108, 1952

  28. [36]

    Measurement of quantum mechanical operators.Physical Review, 120(2):622, 1960

    Huzihiro Araki and Mutsuo M Yanase. Measurement of quantum mechanical operators.Physical Review, 120(2):622, 1960. 6

  29. [37]

    Computationally feasible bounds for partially observed markov decision processes.Operations research, 39(1):162–175, 1991

    William S Lovejoy. Computationally feasible bounds for partially observed markov decision processes.Operations research, 39(1):162–175, 1991

  30. [38]

    Value-function approximations for partially observable markov decision processes.Journal of artificial intelligence research, 13:33–94, 2000

    Milos Hauskrecht. Value-function approximations for partially observable markov decision processes.Journal of artificial intelligence research, 13:33–94, 2000

  31. [39]

    MIT press Cambridge, 1998

    Richard S Sutton, Andrew G Barto, et al.Reinforce- ment learning: An introduction, volume 1. MIT press Cambridge, 1998

  32. [40]

    On the use of non- stationary policies for stationary infinite-horizon markov decision processes.Advances in Neural Information Pro- cessing Systems, 25, 2012

    Bruno Scherrer and Boris Lesner. On the use of non- stationary policies for stationary infinite-horizon markov decision processes.Advances in Neural Information Pro- cessing Systems, 25, 2012

  33. [41]

    Quantum discord and maxwell’s demons.Physical Review A, 67(1):012320, 2003

    Wojciech Hubert Zurek. Quantum discord and maxwell’s demons.Physical Review A, 67(1):012320, 2003

  34. [42]

    Quantum dis- cord, local operations, and maxwell’s demons.Physi- cal Review A—Atomic, Molecular, and Optical Physics, 81(6):062103, 2010

    Aharon Brodutch and Daniel R Terno. Quantum dis- cord, local operations, and maxwell’s demons.Physi- cal Review A—Atomic, Molecular, and Optical Physics, 81(6):062103, 2010

  35. [43]

    Black box work ex- traction and composite hypothesis testing.Physical Re- view Letters, 133(25):250401, 2024

    Kaito Watanabe and Ryuji Takagi. Black box work ex- traction and composite hypothesis testing.Physical Re- view Letters, 133(25):250401, 2024

  36. [44]

    Universal work ex- traction in quantum thermodynamics.Nature Commu- nications, 17(1):1857, 2026

    Kaito Watanabe and Ryuji Takagi. Universal work ex- traction in quantum thermodynamics.Nature Commu- nications, 17(1):1857, 2026

  37. [45]

    Causal asymmetry in a quantum world.Physical Review X, 8(3):031013, 2018

    Jayne Thompson, Andrew JP Garner, John R Mahoney, James P Crutchfield, Vlatko Vedral, and Mile Gu. Causal asymmetry in a quantum world.Physical Review X, 8(3):031013, 2018

  38. [46]

    Causal asymmetry of classical and quantum autonomous agents.arXiv preprint arXiv:2309.13572, 2023

    Spiros Kechrimparis, Mile Gu, and Hyukjoon Kwon. Causal asymmetry of classical and quantum autonomous agents.arXiv preprint arXiv:2309.13572, 2023

  39. [47]

    Quantum coherence, time-translation symmetry, and thermodynamics.Physical review X, 5(2):021001, 2015

    Matteo Lostaglio, Kamil Korzekwa, David Jennings, and Terry Rudolph. Quantum coherence, time-translation symmetry, and thermodynamics.Physical review X, 5(2):021001, 2015

  40. [48]

    Woods and Micha l Horodecki

    Mischa P. Woods and Micha l Horodecki. Autonomous quantum devices: When are they realizable without ad- ditional thermodynamic costs?Phys. Rev. X, 13:011016, Feb 2023

  41. [49]

    Inferring statis- tical complexity.Physical review letters, 63(2):105–108, 1989

    James P Crutchfield, Karl Young, et al. Inferring statis- tical complexity.Physical review letters, 63(2):105–108, 1989

  42. [50]

    Shalizi and James P

    Cosma R. Shalizi and James P. Crutchfield. Computa- tional Mechanics: Pattern and Prediction, Structure and Simplicity.Journal of Statistical Physics, 104(3-4):817– 879, 2001

  43. [51]

    John Wiley & Sons, 2014

    Martin L Puterman.Markov decision processes: discrete stochastic dynamic programming. John Wiley & Sons, 2014

  44. [52]

    Prediction, retrodiction, and the amount of information stored in the present.Journal of Statistical Physics, 136(6):1005–1034, 2009

    Christopher J Ellison, John R Mahoney, and James P Crutchfield. Prediction, retrodiction, and the amount of information stored in the present.Journal of Statistical Physics, 136(6):1005–1034, 2009

  45. [53]

    Exact syn- chronization for finite-state sources.Journal of Statistical Physics, 145(5):1181–1201, 2011

    Nicholas F Travers and James P Crutchfield. Exact syn- chronization for finite-state sources.Journal of Statistical Physics, 145(5):1181–1201, 2011

  46. [54]

    Spectral sim- plicity of apparent complexity

    Paul M Riechers and James P Crutchfield. Spectral sim- plicity of apparent complexity. ii. exact complexities and complexity spectra.Chaos: An Interdisciplinary Journal of Nonlinear Science, 28(3), 2018

  47. [55]

    Cite- seer, 1965

    Richard Bellman, Robert E Kalaba, et al.Dynamic pro- gramming and modern control theory, volume 81. Cite- seer, 1965

  48. [56]

    Cambridge university press, 2010

    Michael A Nielsen and Isaac L Chuang.Quantum compu- tation and quantum information. Cambridge university press, 2010

  49. [57]

    Mc- Graw Hill, 2021

    Walter Rudin.Principles of mathematical analysis. Mc- Graw Hill, 2021. Appendix A: Energy Non-degenerate Hamiltonian In the main body, we have discussed the application of theρ ∗-ideal protocol for a degenerate Hamiltonian. We mentioned that for a degenerate Hamiltonian, this ...

  50. [58]

    The agent’s belief state is a sufficient statistic for the entire history of actions and observations

  51. [59]

    The belief trajectoryK 0:L−1 is a deterministic function of the trajectory over physical realizationsX 1:L, actions A1:L, and observationsW 1:L

  52. [60]

    The Belief State as a Sufficient Statistic We will first prove that belief states are statistically sufficient for making any prediction. Formally, we want to show that Pr(Xt|W1:t−1,A 1:t−1,µ 0) = Pr(Xt|Kt−1,K 0 =µ 0).(C2) This tells us that any information that the past obser...

  53. [61]

    max a∈A JL(KL−1,aL) K0 =η (i) # =I L−1 +E

    Formal Equivalence While the technique of replacing observation histories with belief states is standard in POMDP literature [34, 51], we provide an explicit derivation below adapted to our thermodynamic framework, directly linking the global multi-time state to the agent’s va...

  54. [62]

    Non-stationary phase At timet= 2, which is the last step of the process, the value functionV 3 is 0 regardless of the belief state. Hence, the action taken att= 2 for a DP agent will minimize the expected dissipation, a∗ 2(K1 =η (i)) = argmax a(j)∈A − 1 β D(ξη(i)∥ρ(j)) .(F1) T...

  55. [63]

    The expected distribution of the future is uniform regardless of the current belief

    Whenp= 0.5, where the process becomes a purely random process. The expected distribution of the future is uniform regardless of the current belief. 16

  56. [64]

    Whenr= 0, the classical limit where all states are orthogonal; one can measure along the eigenbasis to obtain perfect knowledge

  57. [65]

    Whenr= 1, a trivial limit where the process emits identical states; hence, regardless of action, the expectation of the future is the same. Interestingly, when the boundary effect completely disappears even for the finite-horizon case, the optimal policy for all time steps bec...

  58. [66]

    We first clarify some notational differences

    Equivalence between work deficit and causal dissipation In this subsection, we will prove Theorem 2. We first clarify some notational differences. In Eq. (I7), the causal dissipation is defined with respect to ˜HL = (ΠQ1,O 1,···,Π QL,OL), while in Theorem 2, it is defined with...

  59. [67]

    Mathematical properties of causal dissipation Here, we try to establish some mathematical properties of causal dissipation

  60. [68]

    The sequence of adaptive measurements/operations can be modeled as a single global unital CPTP dephasing channel Φ acting on the firstL−1 subsystems

    Positivity:δ(Q −→1:L)≥0. The sequence of adaptive measurements/operations can be modeled as a single global unital CPTP dephasing channel Φ acting on the firstL−1 subsystems. The von Neumann entropy of the resulting block-diagonal state evaluates exactly to the Shannon entropy...

  61. [69]

    We then rewrite the entropy rate with this quantity

    since both states appear with probability 1/2 ; here,H 2 is the binary entropy. We then rewrite the entropy rate with this quantity. h := lim L→∞ 1 LS(ρ(1:L)) = lim L→∞ χL L + 1 2LS(ρ0) + 1 2LS(ρ1).(I29) Since the Holevo information is bounded, this results in h= lim L→∞ 1 2LS...

  62. [70]

    5, we demonstrate the agreement between the simulated work deficit and the causal dissipation for both L= 3 andL= 4

    Numerical Simulation In Fig. 5, we demonstrate the agreement between the simulated work deficit and the causal dissipation for both L= 3 andL= 4. Note that the lack of comparison for a longer time horizon is strictly a restriction due to the brute force optimization in calcula...

  63. [71]

    The emitted quantum state can take on either|ϕ 0⟩=|0⟩or|ϕ 1⟩= √r|0⟩+ √1−r|1⟩

  64. [72]

    Belief states:K={(1/2 +ϵ,1/2−ϵ)|ϵ∈[−1/2,1/2]}; we divide the parameterϵinto 300 equal-sized intervals

  65. [73]

    After the optimal policy is found, we apply it by initializing the agent at belief stateK 0 =µ 0 =π= (1/2,1/2) of the perturbed coin process in Fig

    Actions:Aconsists of all bases spanned by|ϕ 0⟩and|ϕ 1⟩, taking the form ofA={|ψ θ⟩= cos θ 2|0⟩+sin θ 2|1⟩|θ∈ [0,2π)}; we also divide the parameterθinto 300 equal sized intervals. After the optimal policy is found, we apply it by initializing the agent at belief stateK 0 =µ 0 =...

  66. [74]

    For a chosen basis, define pt,η,i :=⟨ψ t,η,i|ξ η|ψt,η,i⟩.(J7) The optimal target spectrum for this fixed basis isp t,η

    Spectral tagging Letξ η be the expected state conditioned on the beliefη. For a chosen basis, define pt,η,i :=⟨ψ t,η,i|ξ η|ψt,η,i⟩.(J7) The optimal target spectrum for this fixed basis isp t,η. The unperturbed optimal target state thus takes the form: ρt,η = dX i=1 pt,η,i|ψt,η...

  67. [75]

    the probability of every branch conditioned on every emitted stateσ (x)

  68. [76]

    the Bayesian posterior conditioned on that branch

  69. [77]

    the probability distribution over future beliefs

  70. [78]

    Proof.The perturbed and unperturbed target states share exactly the same eigenbasis

    the future bases selected by the original policy. Proof.The perturbed and unperturbed target states share exactly the same eigenbasis. Branch probabilities depend exclusively on this basis, independently of the target’s eigenvalues, as shown in Eq. (J2). Therefore, the likelih...

  71. [79]

    Work penalty for tagging Having established the injectivity of the update map, we must quantify the average work deficit incurred by introducing this tagging perturbation. Recall that the expected work extracted from the expected stateξusing a protocol tailored forρis given by...

  72. [80]

    It uses the same measurement basis as the original policy at all time steps and beliefs

  73. [81]

    The induced branch probabilities, Bayesian beliefs, and future basis choices match the untagged policy perfectly

  74. [82]

    Every memory update can be implemented via a unitary on the memory and current battery at zero additional thermodynamic cost

  75. [83]

    The memory can be returned to its initial state by sequentially applying the inverse update unitaries in reverse temporal order

  76. [84]

    Proof.We choose α= 1−exp −β∆ L .(J22) At every time step and belief, construct the tagged spectrum using Lemma 4

    Its cumulative expected work extraction is at most∆below the theoretical optimum of the original policy. Proof.We choose α= 1−exp −β∆ L .(J22) At every time step and belief, construct the tagged spectrum using Lemma 4. Points 1 and 2 follow from Lemma 5, while point 3 follows ...

  77. [85]

    We first cast theNdiscrete belief states asNmutually orthogonal quantum states{|i⟩ M}N i=1, each of which would be mapped to different points on the probability simplex

    Operational Justification To clarify how these sequential updates can be executed and eventually uncomputed without incurring a Landauer erasure cost, we construct an explicit operational model. We first cast theNdiscrete belief states asNmutually orthogonal quantum states{|i⟩...

Pith tools

Reviewed August 15, 2026 · model on record in the stance chip above.