Pith. sign in

REVIEW 3 major objections 5 minor 28 references

Exploratory Utility Maximization Problem with Tsallis Entropy

T0 review · 3 major / 5 minor · reviewed 2026-08-09 · deepseek-v4-flash

Pith's one-line read Tsallis exploration can break Merton well-posedness unless cost scales with wealth

desk verdict The β=1 Gaussian analysis and the ill-posedness discussion are genuine contributions, but the β=3 semicircle ODE has a factor error and the headline result does not currently solve its own HJB. read the letter →

arxiv 2502.01269 v1 pith:7Z6PPZHZ submitted 2025-02-03 cs.LG q-fin.MF

classification cs.LGq-fin.MF MSC 91G1093E2060H3068T05
keywords reinforcementlearningMertonproblemTsallisentropyutilitymaximizationwell-posednessexploratoryHJBequationWignersemicircledistributionGaussianexploration
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

This paper studies what happens to the classical Merton portfolio problem when an agent must explore, so the control is randomized and a Tsallis entropy term rewards spread in action choice. It reports that exploration can make the problem ill-posed: with a time-dependent exploration weight $\lambda(t)$, the value is infinite for CRRA parameter $0

What carries the argument

The carrying object is the exploratory HJB equation with the Tsallis entropy regularizer, together with the dimensionality reduction that follows from the ansatz $v(t,w)=f(t)w^p/p$. Because $\lambda(t,w)=\gamma w^p$ shares the exponent $p$ with the utility, every term in the HJB equation has the same homogeneity in $w$, turning the PDE into an ODE for $f$: $y'=ay+b\log y+c$ for $\beta=1$ and $y'=ay+b\sqrt{y}+c$ for $\beta=3$. The sign pattern of these ODEs determines whether the solution exists up to the horizon, which is exactly what decides well-posedness and whether the semi-closed-form optimal policies are valid.

What would settle it

Simulate Gaussian exploratory policies $\pi^n_s\sim N(0,n^2)$ with time-dependent temperature and $0<p<1$: Proposition 2.1 predicts the reward diverges as $n\to\infty$, so observing bounded rewards would refute the ill-posedness claim. For the wealth-scaled case, numerically solve the ODE for $f$ over a grid of parameters with $0<p<1$; existing at the horizon for all parameter choices supports the paper, while a finite-time blow-up below $T$ would refute it.

Watch

Extended reading notes

Core claim

The central claim is that the exploratory utility maximization problem with CRRA utility and Tsallis entropy is not automatically well-posed; over-exploration can push the value function to infinity. For $\lambda$ depending only on time, the problem is ill-posed for $0<p<1$ and well-posed for $p\le 0$, with the closed-form Gaussian policy available only for $p=0$. With primary temperature $\lambda(t,w)=\gamma w^p$, the exploratory HJB equation becomes homogeneous and reduces to an ODE. For $\beta=1$ the optimal policy is Gaussian with mean $(\mu-r)/(\sigma^2(1-p))$, the Merton mean, and variance $\gamma/((1-p)\sigma^2 f(t))$; for $\beta=3$ it is a Wigner semicircle on a compact interval around the same mean, equivalent to a scaled $\mathrm{Beta}(3/2,3/2)$ distribution. For $0<p<1$ the problem is always well-posed, while for $p<0$ well-posedness depends on whether the ODE solution survives to the horizon, and explicit examples show the value can become infinite.

Load-bearing premise

The results rest on choosing the exploration weight as $\lambda(t,w)=\gamma w^p$, proportional to the $p$-th power of wealth; if that scaling does not model the true exploration cost, the well-posedness guarantees and the Gaussian and semicircle optimal policies do not follow.

Editorial extensions

If this is right

  • If $0<p<1$, the exploratory problem is always well-posed under $\lambda=\gamma w^p$ for both $\beta=1$ and $\beta=3$, so the learning objective is finite and the optimal policy has a semi-closed form.
  • The mean of the optimal exploratory policy coincides with the classical Merton ratio, so entropy-regularized exploration does not shift the target allocation, only adds noise around it.
  • As $\gamma\to 0^+$, the value function converges locally uniformly to the classical value, the relative exploration cost and the policy variance go to zero, and the optimal policy converges weakly to a point mass at the Merton strategy.
  • For $p<0$, some parameter choices make the value function infinite before the horizon, so exploration can destroy the learning problem entirely and the agent must detect and avoid those regimes.

Reading between the lines

Editorial extensions of the paper, not claims the author makes directly.

  • An extension not pursued in the paper: the ill-posedness for time-dependent $\lambda$ suggests a general design rule that in portfolio RL, entropy bonuses must scale with wealth according to the utility curvature, otherwise the algorithm optimizes exploration noise rather than terminal wealth.
  • One could test a natural generalization where $\lambda(t,w)=\gamma(t)w^p$ with time-varying $\gamma$; the homogeneity argument would survive, but the ODE becomes non-autonomous and the well-posedness thresholds would shift.
  • The compact-support semicircle policy for $\beta=3$ implies bounded sampled actions, meaning the model gives an explicit bound on leverage whenever the market parameters are trusted.
  • The polynomial parameterization of $f$ indicates that in richer models without a closed form, the same actor-critic scheme would likely need a more expressive function approximator, which the paper leaves open.
Share X Bluesky LinkedIn Reddit HN

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, and a circularity audit.

Referee Report

3 major / 5 minor

Summary. The paper studies the exploratory version of Merton's optimal investment problem with CRRA utility, adding a Tsallis-entropy regularizer to induce exploration. It first shows that a purely time-dependent temperature makes the problem ill-posed for 0<p<1, then proposes a wealth-dependent temperature λ(t,w)=γw^p that makes the exploratory HJB equation homogeneous. For β=1 (Shannon entropy) the authors derive a Gaussian optimal distributional policy, reduce the HJB to a one-dimensional ODE for f(t), characterize well-posedness, prove convergence to the classical Merton solution as γ→0, and design an actor-critic RL algorithm with numerical experiments. For β=3 they claim the optimal policy is a Wigner semicircle distribution and give a semiclosed-form value function through ODE (24). The appendix extends the results to multiple assets. The main advertised novelties are the ill-posedness findings, the semicircle exploration distribution, and the convergence results.

Significance. If correct, the β=1 analysis would be a useful contribution to continuous-time RL for utility maximization, and the β=3 semicircle example would be a distinctive non-Gaussian exploration result. The paper also provides a verification theorem for β=1, an explicit convergence statement, and a working numerical algorithm; these are concrete strengths. However, the β=3 example is one of the two central cases advertised in the abstract and introduction, and its derivation contains a load-bearing algebraic error. Because the claimed semiclosed-form solution does not actually solve the exploratory HJB equation, the Wigner-semicircle and well-posedness conclusions for β=3 are not supported as written. The error appears fixable by reworking Section 3.2 and Appendix 7.6, but the paper is not publishable in its current form.

major comments (3)
  1. [Section 3.2, Eq. (24)] Equation (24) is algebraically inconsistent with the exploratory HJB equation (9). Substituting the semicircular policy (26) and the moments (23) into (9) and dividing by w^p gives f'/p + [r + (μ−r)^2/(2σ^2(1−p))] f − (1/(2π)) sqrt(3γσ^2(1−p)) sqrt(f) + γ/2 = 0. The coefficient of sqrt(f) in (24) is instead p sqrt(3(1−p)σ^2γ/(2π)), which differs from the correct coefficient by the factor p sqrt(2π). For p=1/3, σ=0.5, γ=0.3, μ=0.2, r=0, f=1, the left-hand side of (24) is approximately −0.0101 rather than zero. Consequently Theorem 3.2's claimed value function is not a solution of the HJB, and Proposition 3.5, which relies on (24), inherits the error.
  2. [Appendix 7.6, Proposition 7.1] The multi-asset β=3 ODE in Proposition 7.1 does not reduce to the corrected single-asset ODE when d=1. Evaluating K at d=1 gives pK = p/(4√π) sqrt(3γσ^2(1−p)), whereas the corrected single-asset coefficient is p/(2π) sqrt(3γσ^2(1−p)). These are unequal, so the appendix's dimensional reduction is internally inconsistent. This must be reconciled after the Section 3.2 derivation is corrected.
  3. [Section 3.2, displayed formula for φ(t,w)] The expression for φ(t,w) after the normalization condition is not homogeneous of degree p in wealth: the term √6 γ w^p A/π contains w^{2p} because A = ½σ^2(1−p)w^p f, and the term (μ−r)^2 w^p f^2/(2σ^2(1−p)) has an incorrect dependence on f. Since φ appears inside the square root in (22) together with terms of degree p, this displayed formula cannot be correct. The normalization integral (15) must be recomputed carefully. In addition, Proposition 3.5 and Theorem 3.2 are stated without proofs; after correcting the ODE, those proofs need to be supplied.
minor comments (5)
  1. [Lemma 3.1] The proof of Lemma 3.1 is omitted with the note 'The proof is easy. We omit it.' Since this lemma underpins the whole ODE analysis for β=1, the authors should provide at least a sketch of the proof.
  2. [Equation (9)] In the displayed exploratory HJB equation, the drift term inside the integral is written as (u−r)w v_w u; this should be (μ−r)w v_w u. The same typo appears to be present in the text around (9).
  3. [Appendix 7.6, Proposition 7.1] The first bullet of Proposition 7.1 states 'When β=0', but Tsallis entropy is defined for β>1 and β=1 only. This should presumably read 'When β=1'.
  4. [Algorithm 1] The algorithm samples ε_i ∼ N(0, I_{Ñ}), but the control u is scalar in the one-asset setting; the notation should be N(0,1) or the dimension of the normal distribution should be clarified.
  5. [Abstract and Section 3] The phrase 'fully characterize their well-posedness' is stronger than what is shown: the characterization is specific to the chosen temperature λ(t,w)=γw^p and to β=1 and β=3, not to the general exploratory utility problem. The wording should be adjusted to reflect the conditional scope.

Circularity Check

0 steps flagged · score 0.0 of 10

No significant circularity: results follow from an explicitly stated temperature ansatz and HJB calculations; the flagged β=3 issue is an algebraic error, not a circular step.

full rationale

The derivation chain is self-contained conditional on the explicitly stated modeling assumption λ(t,w)=γw^p. Section 2.2 motivates this choice by homogeneity: 'we infer that the candidate primary temperature function λ(·,·) is related to wealth and consider the case that λ(t,w)=γw^p.' This is a transparent ansatz, not a hidden fit or a prediction equivalent to an input. The main results are obtained by solving the exploratory HJB equation (9) with the separation ansatz v=f(t)w^p/p; the Gaussian (β=1) and semicircle (β=3) forms, and the equality of the mean with the classical Merton ratio, follow from the Euler-Lagrange equation and the homogeneity ansatz, not from imposing the target. The ill-posedness and well-posedness conclusions are proved from explicit control sequences and ODE analysis. The paper cites prior work by Wang-Zhou, Jia-Zhou, and Bo et al. for the exploratory framework and entropy regularizer, but no load-bearing conclusion reduces to a self-citation by the present authors. The β=3 coefficient discrepancy noted by the skeptic would be an algebraic error in the derived ODE, not a circularity, and therefore does not affect this verdict.

Assumptions & free parameters 3 free parameters · 7 assumptions · 0 invented entities

The central theoretical claim rests on the ad hoc choice of a wealth-dependent temperature function λ=γw^p, the moment-matched exploratory wealth dynamics, and standard ODE/calculus tools. No new entities are introduced; γ is a model parameter, and the RL basis coefficients are fit only in the numerical section.

free parameters (3)
  • γ (exploration temperature coefficient) = none in theory; set to 0.3 in experiments
    Introduced as the wealth-scaling coefficient in λ(t,w)=γw^p; it controls the exploration variance and is not derived from data in the theoretical part.
  • θ_i (basis coefficients for f^θ(t)) = learned values reported for a specific experiment in Section 5
    Fitted in the RL algorithm to approximate the true ODE solution f(t); do not enter the theoretical claim.
  • φ_1, φ_2 (policy parameters) = learned values in Table 2, e.g., φ1 → (μ-r)/σ^2
    Fitted in the actor-critic algorithm; used only in the numerical section to represent the Gaussian policy.
assumptions (7)
  • ad hoc to paper Primary temperature function λ(t,w)=γw^p with same p as the CRRA utility exponent
    Chosen so the HJB equation is homogeneous and dimension reduction works; the well-posedness and optimal-policy results are conditional on this scaling. Stated in Section 2.2.
  • domain assumption Exploratory wealth dynamics given by matching first and second moments (equation (4))
    Adopted from Wang-Zhou [26] and derived heuristically in Appendix 7.1; all later HJB derivations use this dynamics.
  • domain assumption Value function is C^{1,2} on the region where f>0 and the verification theorem applies
    Needed for Itô's formula and dynamic programming; the paper assumes this in Theorem 3.1 and the p<0 ill-posed cases.
  • domain assumption Market is complete with a GBM risky asset and constant risk-free rate
    Standard Merton setup from equation (1)-(2); restricts scope to complete markets.
  • standard math Tsallis entropy definition H_β(z)=(z-z^β)/(β-1) for β>1 and Shannon for β=1
    Definition used as the regularizer; properties of the entropy are not proved in the paper.
  • standard math Existence, uniqueness, and extendability of ODE solutions (Picard theorem, comparison theorem, Arzelà-Ascoli)
    Used in Propositions 3.2, 3.5, and Lemma 3.2 without proof.
  • standard math First-variation optimality of the density π over P(R) via the Euler-Lagrange equation
    Used in Section 3 to derive the form of the optimal density; requires interchange of sup and integral that is not fully justified.

how reviews work

0 comments
Cite this review

Pith. "Pith review of Exploratory Utility Maximization Problem with Tsallis Entropy." pith.science (2026). https://pith.science/paper/7Z6PPZHZ

@misc{pith2026250201269,
  author       = {Pith},
  title        = {Pith review of: Exploratory Utility Maximization Problem with Tsallis Entropy},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/7Z6PPZHZ}},
  note         = {Machine review of arXiv:2502.01269}
}
read the original abstract

We study expected utility maximization problem with constant relative risk aversion utility function in a complete market under the reinforcement learning framework. To induce exploration, we introduce the Tsallis entropy regularizer, which generalizes the commonly used Shannon entropy. Unlike the classical Merton's problem, which is always well-posed and admits closed-form solutions, we find that the utility maximization exploratory problem is ill-posed in certain cases, due to over-exploration. With a carefully selected primary temperature function, we investigate two specific examples, for which we fully characterize their well-posedness and provide semi-closed-form solutions. It is interesting to find that one example has the well-known Gaussian distribution as the optimal strategy, while the other features the rare Wigner semicircle distribution, which is equivalent to a scaled Beta distribution. The means of the two optimal exploratory policies coincide with that of the classical counterpart. In addition, we examine the convergence of the value function and optimal exploratory strategy as the exploration vanishes. Finally, we design a reinforcement learning algorithm and conduct numerical experiments to demonstrate the advantages of reinforcement learning.

Figures

Figures reproduced from arXiv: 2502.01269 by the authors.

Figure 1
Figure 1. Numerical Solutions of the ODE under Different Parameter Selections when [PITH_FULL_IMAGE:figures/full_fig_p014_1.png] view at source ↗
Figure 2
Figure 2. Numerical Solutions of the ODE under Different Parameter Selections when [PITH_FULL_IMAGE:figures/full_fig_p015_2.png] view at source ↗
Figure 3
Figure 3. dynamics of φ φ2 converge to the true values quickly. Especially for φ1, it converges after 1000 iterations. It is worth mentioning that the entire experiment took less than half a minute. Next, we examine the approximation of f θ to f. We conduct five experi￾ments, with the results shown in [PITH_FULL_IMAGE:figures/full_fig_p025_3.png] view at source ↗
Figures from the paper (2 more)
Figure 4
Figure 4. Figure 4: approximation of f the true value. For φ2, when iterations end, the learned value has not yet converged for small γ. This suggests that we should carefully consider the exploration weight. 6 Conclusion In this paper, we study the utility maximization problem with the C…
Figure 5
Figure 5. Figure 5: Performance under different γ a semicircular distribution defined on a compact support. This demonstrates that Tsallis entropy excels in scenarios with prevalent non-Gaussian, heavy-tailed behav￾ior on compact support. The results of multiple assets are deferred in App…

Discussion (0). Continue with ORCID to comment.

Reference graph

Works this paper leans on

28 extracted references · 25 canonical work pages

  1. [1]

    L. Bo, Y. Huang, X. Yu, and T. Zhang, Continuous-time q-learning for jump-diffusion models under Tsallis entropy,arXiv preprint arXiv:2407.03888 (2024)

  2. [2]

    Boyce, R.C

    W.E. Boyce, R.C. DiPrima, and D.B. Meade, Elementary differential equa- tions, John Wiley & Sons(2017)

  3. [3]

    M. Dai, Y. Dong, Y. Jia, and X.Y. Zhou, Learning Merton’s Strategies in an Incomplete Market: Recursive Entropy Regularization and Biased Gaussian Exploration, arXiv preprint arXiv:2312.11797(2023)

  4. [4]

    M. Dai, Y. Sun, Z.Q. Xu, and X.Y. Zhou, Learning to Optimally Stop a Diffusion Process,Available at SSRN(2024)

  5. [5]

    Dai and F

    M. Dai and F. Yi, Finite-horizon optimal investment with transaction costs: A parabolic double obstacle problem,Journal of Differential Equations246 (2009), 1445–1469

  6. [6]

    Dong, Randomized optimal stopping problem in continuous time and rein- forcement learning algorithm,SIAM Journal on Control and Optimization62 (2024), 1590–1614

    Y. Dong, Randomized optimal stopping problem in continuous time and rein- forcement learning algorithm,SIAM Journal on Control and Optimization62 (2024), 1590–1614

  7. [7]

    Donnelly and S

    R. Donnelly and S. Jaimungal, Exploratory control with Tsallis entropy for latent factor models,SIAM Journal on Financial Mathematics15 (2024), 26– 53

  8. [8]

    Duffie and L.G

    D. Duffie and L.G. Epstein, Stochastic differential utility,Econometrica: Jour- nal of the Econometric Society(1992), 353–394

Show all 28 references
  1. [9]

    Elie and N

    R. Elie and N. Touzi, Optimal lifetime consumption and investment under a drawdown constraint,Finance and Stochastics12 (2008), 299–330

  2. [10]

    Epstein and S.E

    L.G. Epstein and S.E. Zin, Substitution, risk aversion and the temporal be- havior of consumption and asset returns: A theoretical framework,Handbook of the fundamentals of financial decision making: Part i(2013), 207–239

  3. [11]

    J. Guo, X. Han, and H. Wang, Exploratory mean-variance portfolio selection with Choquet regularizers,arXiv preprint arXiv:2307.03026(2023). 36

  4. [12]

    X. Han, R. Wang, and X.Y. Zhou, Choquet regularization for continuous-time reinforcement learning,SIAM Journal on Control and Optimization61 (2023), 2777–2801

  5. [13]

    Jia and X.Y

    Y. Jia and X.Y. Zhou, Policy evaluation and temporal-difference learning in continuoustimeandspace: Amartingaleapproach, Journal of Machine Learn- ing Research23 (2022), 1–55

  6. [14]

    Jia and X.Y

    Y. Jia and X.Y. Zhou, Policy gradient and actor-critic learning in continu- ous time and space: Theory and algorithms, Journal of Machine Learning Research 23 (2022), 1–50

  7. [15]

    Jia and X.Y

    Y. Jia and X.Y. Zhou, q-Learning in continuous time,Journal of Machine Learning Research24 (2023), 1–61

  8. [16]

    Jiang, D

    R. Jiang, D. Saunders, and C. Weng, The reinforcement learning Kelly strat- egy,Quantitative Finance22 (2022), 1445–1464

  9. [17]

    Kong, A short course in ordinary differential equations,Springer (2014)

    Q. Kong, A short course in ordinary differential equations,Springer (2014)

  10. [18]

    Lim and Y.H

    B.H. Lim and Y.H. Shin, Optimal investment, consumption and retirement decision with disutility and borrowing constraints,Quantitative Finance 11 (2011), 1581–1592

  11. [19]

    Liu and M

    H. Liu and M. Loewenstein, Optimal portfolio selection with transaction costs and finite horizons,The Review of Financial Studies15 (2002), 805–835

  12. [20]

    Merton, Optimum consumption and portfolio rules in a continuous-time model, Stochastic optimization models in finance(1975), 621–661

    R.C. Merton, Optimum consumption and portfolio rules in a continuous-time model, Stochastic optimization models in finance(1975), 621–661

  13. [21]

    Shreve and H.M

    S.E. Shreve and H.M. Soner, Optimal investment and consumption with trans- action costs,The Annals of Applied Probability(1994), 609–692

  14. [22]

    Sutton, Reinforcement learning: An introduction, A Bradford Book (2018)

    R.S. Sutton, Reinforcement learning: An introduction, A Bradford Book (2018)

  15. [23]

    Szepesvári, Algorithms for reinforcement learning,Springer Nature(2022)

    C. Szepesvári, Algorithms for reinforcement learning,Springer Nature(2022)

  16. [24]

    Tang, Y.P

    W. Tang, Y.P. Zhang, and X.Y. Zhou, Exploratory HJB equations and their convergence, SIAM Journal on Control and Optimization 60 (2022), 3191– 3216. 37

  17. [25]

    Tsallis, Possible generalization of Boltzmann-Gibbs statistics,Journal of Statistical Physics52 (1988), 479–487

    C. Tsallis, Possible generalization of Boltzmann-Gibbs statistics,Journal of Statistical Physics52 (1988), 479–487

  18. [26]

    H. Wang, T. Zariphopoulou, and X.Y. Zhou, Reinforcement learning in con- tinuous time and space: A stochastic control approach,Journal of Machine Learning Research21 (2020), 1–34

  19. [27]

    Wang and X.Y

    H. Wang and X.Y. Zhou, Continuous-time mean–variance portfolio selection: A reinforcement learning framework,Mathematical Finance30 (2020), 1273– 1308

  20. [28]

    Wu and L

    B. Wu and L. Li, Reinforcement learning for continuous-time mean-variance portfolio selection in a regime-switching market,Journal of Economic Dynam- ics and Control158 (2024), 104787. 38

Pith tools

Reviewed August 9, 2026 · model on record in the stance chip above.