Pith. sign in

REVIEW 3 cited by

Recent Advances in Reinforcement Learning in Finance

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2112.04553 v4 pith:HLP3OALS submitted 2021-12-08 q-fin.MF cs.LGq-fin.CPq-fin.TR

classification q-fin.MFcs.LGq-fin.CPq-fin.TR
keywords datafinancealgorithmsapproachesassumptionsfinancialmodelamount
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

The rapid changes in the finance industry due to the increasing amount of data have revolutionized the techniques on data processing and data analysis and brought new theoretical and computational challenges. In contrast to classical stochastic control theory and other analytical approaches for solving financial decision-making problems that heavily reply on model assumptions, new developments from reinforcement learning (RL) are able to make full use of the large amount of financial data with fewer model assumptions and to improve decisions in complex financial environments. This survey paper aims to review the recent developments and use of RL approaches in finance. We give an introduction to Markov decision processes, which is the setting for many of the commonly used RL approaches. Various algorithms are then introduced with a focus on value and policy based methods that do not require any model assumptions. Connections are made with neural networks to extend the framework to encompass deep RL algorithms. Our survey concludes by discussing the application of these RL algorithms in a variety of decision-making problems in finance, including optimal execution, portfolio optimization, option pricing and hedging, market making, smart order routing, and robo-advising.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 3 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Reinforcement-Learning Portfolio Allocation with Dynamic Embedding of Market Information

    q-fin.PM 2025-01 conditional novelty 7.0 of 10

    DERL, a dynamic-embedding RL framework, beats value/equal-weighted portfolios and an MLP predict-then-optimize baseline on 1993-2022 U.S. large-cap returns, chiefly in high-volatility periods.

  2. Differentially Private Policy Gradient

    cs.LG 2025-01 conditional novelty 6.0 of 10

    A differentially private policy gradient method that frames DP clipping as a trust-region choice and demonstrates competitive returns on deep RL and RLHF tasks, with the caveat that harder tasks use weak privacy budgets.

  3. Constrained Policy Optimization with Cantelli-Bounded Value-at-Risk

    cs.LG 2026-01 unverdicted novelty 5.0 of 10

    VaR-CPO approximates non-differentiable VaR constraints via Cantelli's inequality to enable safe, sample-efficient policy optimization with zero training violations in feasible environments.

Pith tools