Pith. sign in

Paper Citation Record · LEDGER

Finite-Sample Convergence Bounds for Trust Region Policy Optimization in Mean-Field Games

As of 16 August 2026, this Paper Citation Record lists 99 of 99 outbound references and 1 inbound Pith citation observation for arXiv:2505.22781.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2505.22781 v1

Coverage vector

measured 99 of 99 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-07T13:10:27.705945Z

measured 100 of 100 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-16T06:30:59.297886+00:00

measured 1 of 1 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-05-13T01:17:41.135172Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-05-13T01:27:02.760603Z

Reference resolution

99 of 99 outbound references displayed

  • verified exact2
  • verified fuzzy54
  • unresolved43
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation bf637fce-5932-47bb-a84d-fccde22410ae · outbound

This paper cites and Capuzzo-Dolcetta, I.

Finite-Sample Convergence Bounds for Trust Region Policy Optimization in Mean-Field Games and Capuzzo-Dolcetta, I

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-07T13:10:12.196564Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:10:12.196564Z digest=sha256:ec3869c7134fcddc11c687beaeb1d16524839fa47e2250113e54e8b85cf26694

Observation 4a6670e5-2aa5-40cc-b372-5d783f349faa · outbound

This paper cites and Porretta, A.

Finite-Sample Convergence Bounds for Trust Region Policy Optimization in Mean-Field Games and Porretta, A

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-07T13:10:12.329482Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:10:12.329482Z digest=sha256:f81184912b79b4ac075583409e9b9e3b172f576a9511febf651f31c525d05dcc

Observation 453d78a0-a8be-4b55-9a06-4050d75d6c03 · outbound

This paper cites Mean field games: numerical methods for the planning problem.

Finite-Sample Convergence Bounds for Trust Region Policy Optimization in Mean-Field Games Mean field games: numerical methods for the planning problem

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-07T13:10:12.491661Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:10:12.491661Z digest=sha256:2149c7fa7aaff07808bbb603e2af3ddd3a6c41cf5c03c6e07e38c2b087641518

Observation ccbebb40-1486-421b-adef-e6689a6eff64 · outbound

This paper cites Mean field games and applications: Numerical aspects.

Finite-Sample Convergence Bounds for Trust Region Policy Optimization in Mean-Field Games Mean field games and applications: Numerical aspects

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-07T13:10:12.619752Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:10:12.619752Z digest=sha256:e7506bfc7410354701e9178cfdfcc9e64bdf8948a62ef2cbea2d8ddc607accb6

Observation 50fb8c21-04b9-48c5-9ee9-c258f2db1281 · outbound

This paper cites an unresolved cited work.

Finite-Sample Convergence Bounds for Trust Region Policy Optimization in Mean-Field Games Unresolved cited work

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-07T13:10:12.806388Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:10:12.806388Z digest=sha256:5b4251c98567bb0e72a1f1b11b785a826a0d66ec5baba982ab67a022b56e97a8

Observation 99b8884a-b888-4366-8840-6503594e9d34 · outbound

This paper cites An extended mean field game for storage in smart grids.

Finite-Sample Convergence Bounds for Trust Region Policy Optimization in Mean-Field Games An extended mean field game for storage in smart grids

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-07T13:10:13.017801Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:10:13.017801Z digest=sha256:4a10980b0840b4e4ec34b13e3590343e9c5ccd1680e2bf02e76fd0c481d6740d

Observation 5c2ea739-69c9-4425-b6bb-f1b752ee86ad · outbound

This paper cites Regularization of the policy updates for stabilizing mean field games.

Finite-Sample Convergence Bounds for Trust Region Policy Optimization in Mean-Field Games Regularization of the policy updates for stabilizing mean field games

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-07T13:10:13.221565Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:10:13.221565Z digest=sha256:461ba89887e7805ae608a75380387b42726008d074fc0f404bcd0c7664d4de6b

Observation 68a8815d-4929-4184-b623-1ef5e7930c55 · outbound

This paper cites D., and Saldi, N.

Finite-Sample Convergence Bounds for Trust Region Policy Optimization in Mean-Field Games D., and Saldi, N

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-07T13:10:13.373064Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:10:13.373064Z digest=sha256:f3403d4749bdb18788cc301c6ea98a22e2dfa51fc6420d0015de761af41c508f

Observation 5f18d09d-74d5-4ad8-aa92-cd8971e47d3c · outbound

This paper cites D., and Saldi, N.

Finite-Sample Convergence Bounds for Trust Region Policy Optimization in Mean-Field Games D., and Saldi, N

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-07T13:10:13.491777Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:10:13.491777Z digest=sha256:85112aac559f92094496bbaa9f1f8284ddd16076a5889a8a8601de321daf1df6

Observation 29ad1f48-6d67-45d6-8117-15ccfb0cfe63 · outbound

This paper cites Mean-field sampling for cooperative multi-agent reinforcement learning.

Finite-Sample Convergence Bounds for Trust Region Policy Optimization in Mean-Field Games Mean-field sampling for cooperative multi-agent reinforcement learning

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-07T13:10:13.621544Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:10:13.621544Z digest=sha256:af6676d65bce6489cb52de4b2a2229e27fd5b56403b95a20646b3e3cc4676b20

Observation 570de5b3-c6e6-4ed4-84d7-b2cd89a7bb11 · outbound

This paper cites Unified Reinforcement Q-Learning for Mean Field Game and Control Problems.

Finite-Sample Convergence Bounds for Trust Region Policy Optimization in Mean-Field Games Unified Reinforcement Q-Learning for Mean Field Game and Control Problems

Reference 11

Resolution
verified exact
local_arxiv, observed 2026-08-07T13:10:28.517786Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-08-07T13:10:13.774546Z digest=sha256:70d50006b210f440409bad004a0e5dac39ad69b0c0decacb3a1ae3eca14eaed5

Observation a115e582-f4c1-4f97-874b-d709fb4e066e · outbound

This paper cites Convergence of Multi-Scale Reinforcement Q-Learning Algorithms for Mean Field Game and Control Problems.

Finite-Sample Convergence Bounds for Trust Region Policy Optimization in Mean-Field Games Convergence of Multi-Scale Reinforcement Q-Learning Algorithms for Mean Field Game and Control Problems

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-07T13:10:13.932048Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:10:13.932048Z digest=sha256:0aa21a65583fe2557217f071c0e653c49e4e0da963178e807c20ec23e9a68a27

Observation d9faeb28-e1ab-46f7-9b5b-4d91e7518fbd · outbound

This paper cites On solutions of mean field games with ergodic cost.

Finite-Sample Convergence Bounds for Trust Region Policy Optimization in Mean-Field Games On solutions of mean field games with ergodic cost

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-07T13:10:14.074540Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:10:14.074540Z digest=sha256:a28a14a3f3de4ba2d4fcd7a4ba5506701ed54fde375da4bb297494b3fc46ec82

Observation 7c38864c-386d-4df4-b653-44aafd25fbe2 · outbound

This paper cites Lipschitz continuity in model-based reinforcement learning.

Finite-Sample Convergence Bounds for Trust Region Policy Optimization in Mean-Field Games Lipschitz continuity in model-based reinforcement learning

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-07T13:10:14.243790Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:10:14.243790Z digest=sha256:f46546fe9b0c16e8a7c97f0fd6447ab82b47f5c42d96243c97ce669333aa621a

Observation d1c12000-e29d-4a60-afc9-9eed044892e0 · outbound

This paper cites Inapproximability of np-complete variants of nash equilibrium.

Finite-Sample Convergence Bounds for Trust Region Policy Optimization in Mean-Field Games Inapproximability of np-complete variants of nash equilibrium

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-07T13:10:14.377895Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:10:14.377895Z digest=sha256:92f893b99f37a587341b47abcf3ce2f1f4984849b9d3e5350be2e1f00d97839d

Observation cb65bbb7-a39b-440f-8e8e-874cfdae2ee1 · outbound

This paper cites M., Munos, R., and Kappen, H.

Finite-Sample Convergence Bounds for Trust Region Policy Optimization in Mean-Field Games M., Munos, R., and Kappen, H

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-07T13:10:14.540084Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:10:14.540084Z digest=sha256:9160b5fde9b5b375e524ee9b991150d8070d409bed36731f2fecb5ddd3905e77

Observation d527bfc3-31dc-47f3-9bcb-745d3b78ebf6 · outbound

This paper cites and Priuli, F.

Finite-Sample Convergence Bounds for Trust Region Policy Optimization in Mean-Field Games and Priuli, F

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-07T13:10:14.688053Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:10:14.688053Z digest=sha256:6969c4ce6daad2220b718ef4615eab70602b40e8081cce57325c8b061197fb5f

Observation 0bb83870-e706-48a0-8f24-a83d215f526a · outbound

This paper cites A mean-field game model of electricity market dynamics.

Finite-Sample Convergence Bounds for Trust Region Policy Optimization in Mean-Field Games A mean-field game model of electricity market dynamics

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-07T13:10:14.886881Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:10:14.886881Z digest=sha256:3f14da479af15f2fa25a28fc3d821d8c5d2f615556ab73f08fef215078f1db61

Observation 7a13ea33-c593-4ce7-b35b-3ddd47b16e40 · outbound

This paper cites and Hesse, S.

Finite-Sample Convergence Bounds for Trust Region Policy Optimization in Mean-Field Games and Hesse, S

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-07T13:10:15.037956Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:10:15.037956Z digest=sha256:7dfa9c2209bdf0880f54ed584f8f55c1c536ccc12bc521ca6d62a88fa1ddaa1d

Observation 7d1e2b1d-41c2-481a-9346-d99ba8da6420 · outbound

This paper cites First-order methods in optimization.

Finite-Sample Convergence Bounds for Trust Region Policy Optimization in Mean-Field Games First-order methods in optimization

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-07T13:10:15.203972Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:10:15.203972Z digest=sha256:9a380018bf7811d4d2242b2c50bc511a9c56fbbc787282a2adb66953f14dca5f

Observation be756859-2ab8-4cf2-a5ca-494054b2bbe8 · outbound

This paper cites and Russo, D.

Finite-Sample Convergence Bounds for Trust Region Policy Optimization in Mean-Field Games and Russo, D

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-07T13:10:15.404955Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:10:15.404955Z digest=sha256:0f59ad7e625d58c61b1397e4f7bd4cd96ea2737e719dbc13c3c85fe46e2ad9ef

Observation 48e4c65e-cc51-4081-b65b-5e4d78f9ff8f · outbound

This paper cites A., Ortega, P.

Finite-Sample Convergence Bounds for Trust Region Policy Optimization in Mean-Field Games A., Ortega, P

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-07T13:10:15.613780Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:10:15.613780Z digest=sha256:9b4c4102b3037c0f90d6c52005dc6edc1b79695956066043707ff36ea2a0849b

Observation 6bd91a86-fb64-40ac-8366-1dfc1730f3f9 · outbound

This paper cites The master equation and the convergence problem in mean field games:(ams-201).

Finite-Sample Convergence Bounds for Trust Region Policy Optimization in Mean-Field Games The master equation and the convergence problem in mean field games:(ams-201)

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-07T13:10:15.729213Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:10:15.729213Z digest=sha256:53cb821507853c6c4768b794d7560a96285e9ce32c216aa32b0e157304ab4d67

Observation 7dce5b72-226e-41ef-bbc0-72f17a6825ad · outbound

This paper cites and Lauri \`e re, M.

Finite-Sample Convergence Bounds for Trust Region Policy Optimization in Mean-Field Games and Lauri \`e re, M

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:10:45.706097Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-08-07T13:10:15.885470Z digest=sha256:e39daee44965faafba92e4460d95a032674ec91663bb8eb3fbd8e122a23a3fd1

Observation 9096057d-4782-48ed-a291-3e6a5577120b · outbound

This paper cites Mean Field Games and Systemic Risk.

Finite-Sample Convergence Bounds for Trust Region Policy Optimization in Mean-Field Games Mean Field Games and Systemic Risk

Reference 25

Resolution
verified exact
local_arxiv, observed 2026-08-07T13:10:28.223000Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-08-07T13:10:16.063803Z digest=sha256:630ae9e81171c0b8e225306a11fe78ad1015b70f5346386c6bbb8c58d15ee9c3

Observation db6ebada-1185-4011-b4e7-f9bebad60ca4 · outbound

This paper cites Probabilistic theory of mean field games with applications I-II.

Finite-Sample Convergence Bounds for Trust Region Policy Optimization in Mean-Field Games Probabilistic theory of mean field games with applications I-II

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:10:45.413359Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-08-07T13:10:16.251055Z digest=sha256:bece6fb0b582985355031e9c0d789c8417a928ba3f6132616fc323bb322383db

Observation 25424109-0fc2-4849-905e-af23cd3e2c59 · outbound

This paper cites Numerical method for fbsdes of mckean--vlasov type.

Finite-Sample Convergence Bounds for Trust Region Policy Optimization in Mean-Field Games Numerical method for fbsdes of mckean--vlasov type

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:10:45.138057Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-08-07T13:10:16.373548Z digest=sha256:fccf809002a50b1d45900e76c9b5697d8b986d73385ff8fbb0d4a8137fe8f11c

Observation 6b9d0135-894c-417e-83a5-c4d104926724 · outbound

This paper cites an unresolved cited work.

Finite-Sample Convergence Bounds for Trust Region Policy Optimization in Mean-Field Games Unresolved cited work

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-07T13:10:16.529488Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:10:16.529488Z digest=sha256:0b45235a7666ea937f658c88828da636b0e03e168562fc8092cae2ad4edae16d

Observation 2e93f1e8-f4a9-473f-a040-9bb9bc1dec24 · outbound

This paper cites and Koeppl, H.

Finite-Sample Convergence Bounds for Trust Region Policy Optimization in Mean-Field Games and Koeppl, H

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:10:44.865814Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-08-07T13:10:16.744269Z digest=sha256:6309f954d237f640c3b3ae9301df8c7fdb2931efabda62f73758ac3efa9ba181

Observation 27bbb1b0-3aba-4b92-9e01-73ad00a0db6d · outbound

This paper cites and Koeppl, H.

Finite-Sample Convergence Bounds for Trust Region Policy Optimization in Mean-Field Games and Koeppl, H

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:10:44.565375Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-08-07T13:10:16.855232Z digest=sha256:7a2dd7deed822e833f4001a7cf721cecf9f219e84a256afbdf33aa03c73bd201

Observation dec370c1-d457-45b3-bb97-7d9744f31cb0 · outbound

This paper cites A mean field game analysis of sir dynamics with vaccination.

Finite-Sample Convergence Bounds for Trust Region Policy Optimization in Mean-Field Games A mean field game analysis of sir dynamics with vaccination

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:10:44.298047Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-08-07T13:10:17.038129Z digest=sha256:5dbec3c3533b675fff49b26d33110a5a77d314775ce729b53685f5d20f029185

Observation dd77e7cb-60ec-45b4-9aab-1a478328e33f · outbound

This paper cites and Touzi, N.

Finite-Sample Convergence Bounds for Trust Region Policy Optimization in Mean-Field Games and Touzi, N

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:10:43.918293Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-08-07T13:10:17.232176Z digest=sha256:7e19599d54870350f6f233779b1990ea375aefb6e93c6f353d2c26455662a3de

Observation f47c2e38-69af-4291-97e3-dfd0e973e66b · outbound

This paper cites and Silva, F.

Finite-Sample Convergence Bounds for Trust Region Policy Optimization in Mean-Field Games and Silva, F

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:10:43.646217Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-08-07T13:10:17.431580Z digest=sha256:c087f48ac12ecb0e323f824e5639e6e7df654c5c19910d305b0f46dafe3b057e

Observation 97469166-6922-49d8-8dd6-e7e430524278 · outbound

This paper cites N-player games and mean field games of moderate interactions.

Finite-Sample Convergence Bounds for Trust Region Policy Optimization in Mean-Field Games N-player games and mean field games of moderate interactions

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:10:43.356158Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-08-07T13:10:17.587562Z digest=sha256:2de7806b84960d2cd4f930f17c27ae712bd78249966c7bab1adfab4cdecd286e

Observation 023d22f7-2603-4386-8fbe-f1dde2ae9152 · outbound

This paper cites Counterfactual multi-agent policy gradients.

Finite-Sample Convergence Bounds for Trust Region Policy Optimization in Mean-Field Games Counterfactual multi-agent policy gradients

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:10:43.097433Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-08-07T13:10:17.719823Z digest=sha256:878d04d0db54e88d660c70ea20c2dbef94f27b1833b0f2b6e1e369faa8f582a3

Observation 79bbc963-3836-4e96-9a35-0f8520672c3c · outbound

This paper cites Convergence of adaptive and interacting markov chain monte carlo algorithms.

Finite-Sample Convergence Bounds for Trust Region Policy Optimization in Mean-Field Games Convergence of adaptive and interacting markov chain monte carlo algorithms

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:10:42.793131Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-08-07T13:10:17.855496Z digest=sha256:266fff0340f97b4cbddafc0c1854a9204b5359120c88c15a2f786e8898ba2141

Observation 03d71b27-7be9-4f46-a629-b5e3f2106beb · outbound

This paper cites Taming the noise in reinforcement learning via soft updates.

Finite-Sample Convergence Bounds for Trust Region Policy Optimization in Mean-Field Games Taming the noise in reinforcement learning via soft updates

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:10:42.494859Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-08-07T13:10:18.009623Z digest=sha256:23a1c100ac1a885e576e620e63ba102263c2879cd3b035bcbeace288d938f497

Observation b1b27f0a-1f7f-4f2d-9c0f-95401d9ff0e1 · outbound

This paper cites A theory of regularized markov decision processes.

Finite-Sample Convergence Bounds for Trust Region Policy Optimization in Mean-Field Games A theory of regularized markov decision processes

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-07T13:10:18.193450Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:10:18.193450Z digest=sha256:6b13125b4a58db3ea3afcce428e35949ffdf8a1e20459a30387dfc2452f99981

Observation 78fb16c5-db85-4513-951f-93cf8e0bbccb · outbound

This paper cites Concave Utility Reinforcement Learning: the Mean-Field Game Viewpoint.

Finite-Sample Convergence Bounds for Trust Region Policy Optimization in Mean-Field Games Concave Utility Reinforcement Learning: the Mean-Field Game Viewpoint

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-07T13:10:18.346868Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:10:18.346868Z digest=sha256:362a770cc729e29fa25874671ff9a224a41520a850eb96e25eca12a91f6a6425

Observation d69e59ea-a95b-4796-8bf2-d7fbc8088ae4 · outbound

This paper cites Numerical resolution of mckean-vlasov fbsdes using neural networks.

Finite-Sample Convergence Bounds for Trust Region Policy Optimization in Mean-Field Games Numerical resolution of mckean-vlasov fbsdes using neural networks

Reference 40

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:10:42.256776Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-08-07T13:10:18.546112Z digest=sha256:19c0ba0ab3fa7003340350ad0cb56b4e00f299609927cbbb7e64b8b4ec9b5582

Observation 3a1a6c36-559e-48d7-94c7-a7e450b1a6bf · outbound

This paper cites and Diepold, K.

Finite-Sample Convergence Bounds for Trust Region Policy Optimization in Mean-Field Games and Diepold, K

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-07T13:10:18.689464Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:10:18.689464Z digest=sha256:a0f6d2a1bf80b5d7c7dea2537fd3de1915516a1660fa0e2126e51e200a175045

Observation 5bfd2c8a-fd34-4a50-ae8f-e06c4b1236c4 · outbound

This paper cites Learning mean-field games.

Finite-Sample Convergence Bounds for Trust Region Policy Optimization in Mean-Field Games Learning mean-field games

Reference 42

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:10:41.972833Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-08-07T13:10:18.864106Z digest=sha256:ee13efe238da7b2b86af72027f5c61dd1a84c63bacb1588797b313801f1d14ba

Observation ccd5ec80-907a-4e95-8d0e-98e07be6c675 · outbound

This paper cites A general framework for learning mean-field games.

Finite-Sample Convergence Bounds for Trust Region Policy Optimization in Mean-Field Games A general framework for learning mean-field games

Reference 43

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:10:41.639343Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-08-07T13:10:19.028275Z digest=sha256:d52622072364e5abfd86b89e4faa09f73628c9cda9a705d1ca6021836df080b6

Observation e65f3309-8d85-4952-bc78-5a6fae295914 · outbound

This paper cites Reinforcement learning with deep energy-based policies.

Finite-Sample Convergence Bounds for Trust Region Policy Optimization in Mean-Field Games Reinforcement learning with deep energy-based policies

Reference 44

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:10:41.336299Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-08-07T13:10:19.165381Z digest=sha256:867cbdca040dc4f74e99b9345fa9f3f8d4494b13cf7239e7ee264f2301bdf6d5

Observation 5b074396-72b5-4276-ba77-c170a92c8aa2 · outbound

This paper cites J., Liaw, C., Plan, Y., and Randhawa, S.

Finite-Sample Convergence Bounds for Trust Region Policy Optimization in Mean-Field Games J., Liaw, C., Plan, Y., and Randhawa, S

Reference 45

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:10:40.989906Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-08-07T13:10:19.339607Z digest=sha256:1010eb71f47074be24f16340d0ca9997eeba51b2a43064c027a10d2740890cb9

Observation 94f56b99-100a-4721-8c09-56180e6a5ef6 · outbound

This paper cites an unresolved cited work.

Finite-Sample Convergence Bounds for Trust Region Policy Optimization in Mean-Field Games Unresolved cited work

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-07T13:10:19.489673Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:10:19.489673Z digest=sha256:ae825b36249d666c7e03d9cf5b51ee51add116d05befa2c5d45e95729850fa9b

Observation f8a7b82f-be21-40f2-bce0-cb5781e8892a · outbound

This paper cites an unresolved cited work.

Finite-Sample Convergence Bounds for Trust Region Policy Optimization in Mean-Field Games Unresolved cited work

Reference 47

Resolution
unresolved
raw_fallback, observed 2026-08-07T13:10:40.725380Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-08-07T13:10:19.640667Z digest=sha256:dd89abb16b77b38bce7f46458f3ab65719fcd311693e0669c7f6ae99a0705f73

Observation 0a887376-c623-4d30-a983-83a98eab0cb3 · outbound

This paper cites E., and Malham \'e , R.

Finite-Sample Convergence Bounds for Trust Region Policy Optimization in Mean-Field Games E., and Malham \'e , R

Reference 48

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:10:40.511051Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-08-07T13:10:19.736199Z digest=sha256:82ff75b48d682723ea0e878474b09f657b29f3d502941f24c986e992bbef96b6

Observation 18f08f07-8014-474d-ad13-e41ec9bc9bf9 · outbound

This paper cites P., and Caines, P.

Finite-Sample Convergence Bounds for Trust Region Policy Optimization in Mean-Field Games P., and Caines, P

Reference 49

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:10:40.194629Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-08-07T13:10:19.880740Z digest=sha256:21792a36d80f48312a138768685c5b5e224868c40968075f4e170c146b08d88e

Observation 39a6dc44-f5cd-4cd2-a412-8b19c3e48907 · outbound

This paper cites P., and Caines, P.

Finite-Sample Convergence Bounds for Trust Region Policy Optimization in Mean-Field Games P., and Caines, P

Reference 50

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:10:39.850916Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-08-07T13:10:20.085274Z digest=sha256:ef084e0d0ca441de4d12770a45d102841af13af70c1bebe72c09f81c9ec9212f

Observation db69e7c9-770d-46b8-9d34-f9b69d01db74 · outbound

This paper cites P., and Caines, P.

Finite-Sample Convergence Bounds for Trust Region Policy Optimization in Mean-Field Games P., and Caines, P

Reference 51

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:10:39.591521Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-08-07T13:10:20.238426Z digest=sha256:63664d0c11fc575cbe713daa235e1c7beecc1ca95cca7c5632c6c3292fcab063

Observation 0868478e-fd85-46bb-9d00-5a1b1ab40a13 · outbound

This paper cites and Sha, F.

Finite-Sample Convergence Bounds for Trust Region Policy Optimization in Mean-Field Games and Sha, F

Reference 52

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:10:39.225351Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-08-07T13:10:20.405360Z digest=sha256:7c743f27b290f6f5997e7f2a3ad34742f85a798f196c0d69d150ab45d6ce9094

Observation 9f637947-4b9e-4903-a1cf-004bc8ed7e5d · outbound

This paper cites and Langford, J.

Finite-Sample Convergence Bounds for Trust Region Policy Optimization in Mean-Field Games and Langford, J

Reference 53

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:10:38.952013Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-08-07T13:10:20.536988Z digest=sha256:b89c62994c732bd1074c12c8cb515630c7dbd70e35cdb07c7c044d5d734b53b2

Observation d2cf9ab8-375c-4972-aa89-067c533187b7 · outbound

This paper cites an unresolved cited work.

Finite-Sample Convergence Bounds for Trust Region Policy Optimization in Mean-Field Games Unresolved cited work

Reference 54

Resolution
unresolved
raw_fallback, observed 2026-08-07T13:10:38.682202Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-08-07T13:10:20.673691Z digest=sha256:d1aaf8ff158e7ef01737eeb80457c36877edc8c9800fe6ddb464462d731cba73

Observation cbab3140-898f-4cc1-af1f-5adf410ed2c3 · outbound

This paper cites and Singh, S.

Finite-Sample Convergence Bounds for Trust Region Policy Optimization in Mean-Field Games and Singh, S

Reference 55

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:10:38.365483Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-08-07T13:10:20.845556Z digest=sha256:4fb796278584013f2d35dae309fc588c5bfe508272419ba6e6ef9807a7d0aa6d

Observation a733ded7-9ecf-4b7d-a69f-896fc7515207 · outbound

This paper cites A., and Peters, J.

Finite-Sample Convergence Bounds for Trust Region Policy Optimization in Mean-Field Games A., and Peters, J

Reference 56

Resolution
unresolved
no resolver link, observed 2026-08-07T13:10:21.019331Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:10:21.019331Z digest=sha256:86d85a040122213820b9ebf2061dadac5824247cd02f89278b601934e120d221

Observation b216d78f-32db-4a42-9f27-166e20c2548e · outbound

This paper cites and Zariphopoulou, T.

Finite-Sample Convergence Bounds for Trust Region Policy Optimization in Mean-Field Games and Zariphopoulou, T

Reference 57

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:10:38.039128Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-08-07T13:10:21.140695Z digest=sha256:4613d1782e79d9b5241d77466fa33f4393158bbfaa478762c07f69cda407b928

Observation 6137cd43-9af3-45e5-b098-e70a40ca2750 · outbound

This paper cites and Lions, P.-L.

Finite-Sample Convergence Bounds for Trust Region Policy Optimization in Mean-Field Games and Lions, P.-L

Reference 58

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:10:37.682067Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-08-07T13:10:21.319382Z digest=sha256:2d43d6f435f3866a735e0f75db016d71dbd7c21b6a8b6e8d3036478e9bb1a8fc

Observation 00e46286-6eca-46ee-98e3-ab074020cf14 · outbound

This paper cites and Lions, P.-L.

Finite-Sample Convergence Bounds for Trust Region Policy Optimization in Mean-Field Games and Lions, P.-L

Reference 59

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:10:37.358005Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-08-07T13:10:21.475442Z digest=sha256:ba49134dff6a4d48ce38bd9b18b1e5af3b7bf4a414895b2adc054dda05fb4ddb

Observation dc7ff6ff-efdf-4703-ab72-e57f9deb2496 · outbound

This paper cites and Lions, P.-L.

Finite-Sample Convergence Bounds for Trust Region Policy Optimization in Mean-Field Games and Lions, P.-L

Reference 60

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:10:37.110259Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-08-07T13:10:21.678545Z digest=sha256:2ee345e54f98b816d951aa9a8ddb9a1bb69ea1da1c0df4e8b56e289e11417949

Observation df7cf107-2d2d-46dd-ad91-acc63b4c49c2 · outbound

This paper cites Learning in Mean Field Games: A Survey.

Finite-Sample Convergence Bounds for Trust Region Policy Optimization in Mean-Field Games Learning in Mean Field Games: A Survey

Reference 61

Resolution
unresolved
no resolver link, observed 2026-08-07T13:10:21.805809Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:10:21.805809Z digest=sha256:e0a5a471a79fcdeb4865592cff0fb3a213df6ddb2da787ceddb1aa25840c326e

Observation caa7464b-ca39-4d2a-b1a3-52c81f97e0b2 · outbound

This paper cites and Tankov, P.

Finite-Sample Convergence Bounds for Trust Region Policy Optimization in Mean-Field Games and Tankov, P

Reference 62

Resolution
unresolved
no resolver link, observed 2026-08-07T13:10:21.939715Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:10:21.939715Z digest=sha256:5cb7c6727d046cf3480c37bf470f663caa5709b95db26036875d3af0ab7690ee

Observation b4ef1a75-91d8-4d15-95cf-65c2afa50d05 · outbound

This paper cites G., and Castro, P.

Finite-Sample Convergence Bounds for Trust Region Policy Optimization in Mean-Field Games G., and Castro, P

Reference 63

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:10:36.778709Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-08-07T13:10:22.118395Z digest=sha256:765e47ea88351de83fb85433cd4734bb57811e5481c6b526bf12412b39413759

Observation 061924ee-e8d9-42d4-a570-edba6260118a · outbound

This paper cites A mean-field game approach to cloud resource management with function approximation.

Finite-Sample Convergence Bounds for Trust Region Policy Optimization in Mean-Field Games A mean-field game approach to cloud resource management with function approximation

Reference 64

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:10:36.518742Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-08-07T13:10:22.294882Z digest=sha256:adf9c8f7f5da8dd02041a5413159563b5296f973bfe439870a90068c0a1c1cd0

Observation e4d9343c-f81b-4127-ac4e-b3612d920a9b · outbound

This paper cites I., Fern \'a ndez-Gaucherand, E., Hern \'a ndez-Hernandez, D., Coraluppi, S., and Fard, P.

Finite-Sample Convergence Bounds for Trust Region Policy Optimization in Mean-Field Games I., Fern \'a ndez-Gaucherand, E., Hern \'a ndez-Hernandez, D., Coraluppi, S., and Fard, P

Reference 65

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:10:36.242074Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-08-07T13:10:22.422162Z digest=sha256:31c8860500bba44501c71f286d8b5ca8c30a6a3581b00e6ae4092db654cf4a62

Observation 62b0348e-6afa-4d28-a1a0-b1e1d296efc5 · outbound

This paper cites J., and Le Fort-Piat, N.

Finite-Sample Convergence Bounds for Trust Region Policy Optimization in Mean-Field Games J., and Le Fort-Piat, N

Reference 66

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:10:35.938947Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-08-07T13:10:22.542113Z digest=sha256:7b6b2743795e5265bcb40630644a813f4bb7c7f659b414a015a5c51a43bf43cb

Observation c2192078-c026-4f09-9341-aeb0ddfb1cd0 · outbound

This paper cites On the global convergence rates of softmax policy gradient methods.

Finite-Sample Convergence Bounds for Trust Region Policy Optimization in Mean-Field Games On the global convergence rates of softmax policy gradient methods

Reference 67

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:10:35.674906Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-08-07T13:10:22.629063Z digest=sha256:75254b2110ffc1419edc2ba68cd3e0e5f5eeb8085800833cd8e22a11c69753aa

Observation 179377ee-779a-40c0-adfe-684e4f89a49c · outbound

This paper cites Asynchronous methods for deep reinforcement learning.

Finite-Sample Convergence Bounds for Trust Region Policy Optimization in Mean-Field Games Asynchronous methods for deep reinforcement learning

Reference 68

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:10:35.407930Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-08-07T13:10:22.697117Z digest=sha256:061d6b2aaaa646fc85b02496bd86cdbfcc4d21e0582ec98447a9df52be7effef

Observation deee8e41-6131-4519-83d0-0d81307edf98 · outbound

This paper cites Equilibrium points in n-person games.

Finite-Sample Convergence Bounds for Trust Region Policy Optimization in Mean-Field Games Equilibrium points in n-person games

Reference 69

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:10:35.017098Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-08-07T13:10:22.825724Z digest=sha256:b2201a91d0283f9bf0ebb517005683b0b01a7086275de2a9e7fb40d332e78f43

Observation 9459171a-5cd1-4b5f-a64c-ace63906c4e2 · outbound

This paper cites Non-cooperative games.

Finite-Sample Convergence Bounds for Trust Region Policy Optimization in Mean-Field Games Non-cooperative games

Reference 70

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:10:34.652398Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-08-07T13:10:22.945284Z digest=sha256:fa4ab80adcce5e69fffae6e5cfba956ca0a6b3c0007be74cf37ff110a0d2e120

Observation 8797b717-c186-4857-a0cb-d49e42ae35a9 · outbound

This paper cites A unified view of entropy-regularized Markov decision processes.

Finite-Sample Convergence Bounds for Trust Region Policy Optimization in Mean-Field Games A unified view of entropy-regularized Markov decision processes

Reference 71

Resolution
unresolved
no resolver link, observed 2026-08-07T13:10:23.087809Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:10:23.087809Z digest=sha256:d94811c3d98d19989d99d5a09ad736f970aef82e90ca32c9eaf12b4e1e1c9bbd

Observation d45e711e-b08b-4bf2-b8aa-78fad369ea65 · outbound

This paper cites Combining policy gradient and q-learning.

Finite-Sample Convergence Bounds for Trust Region Policy Optimization in Mean-Field Games Combining policy gradient and q-learning

Reference 72

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:10:34.305505Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-08-07T13:10:23.258768Z digest=sha256:8125a3f1f1faae057517dc59a1f00cfdb20af384f95cb32fbe8336953d86c314

Observation 6a6656f1-46f9-4cf2-a0f6-3db7528d17c0 · outbound

This paper cites Scaling mean field games by online mirror descent.

Finite-Sample Convergence Bounds for Trust Region Policy Optimization in Mean-Field Games Scaling mean field games by online mirror descent

Reference 73

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:10:33.964017Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-08-07T13:10:23.407117Z digest=sha256:e869ff9d150daace48c15b00e77318e940952a703feb35bc2ac12045ac939be7

Observation d341e81a-6d43-4e45-9261-cf9318c72e8f · outbound

This paper cites Fictitious play for mean field games: C ontinuous time analysis and applications.

Finite-Sample Convergence Bounds for Trust Region Policy Optimization in Mean-Field Games Fictitious play for mean field games: C ontinuous time analysis and applications

Reference 74

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:10:33.667286Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-08-07T13:10:23.669181Z digest=sha256:a34a69d7ba5a915f84352b821f8fb5a9d0da70bc654a57c6ffa0d34b36f75cb9

Observation 9804e0d7-ba49-4ec8-b59b-1efd1be482c1 · outbound

This paper cites Mean Field Games Flock! The Reinforcement Learning Way.

Finite-Sample Convergence Bounds for Trust Region Policy Optimization in Mean-Field Games Mean Field Games Flock! The Reinforcement Learning Way

Reference 75

Resolution
unresolved
no resolver link, observed 2026-08-07T13:10:23.894935Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:10:23.894935Z digest=sha256:1fb73672a54abe76d0a0b46a8c8f5244fe017213ef42d546779dd8088a4abeac

Observation ed45f950-7f0f-4a09-9c65-7cf997065bfc · outbound

This paper cites Generalization in mean field games by learning master policies.

Finite-Sample Convergence Bounds for Trust Region Policy Optimization in Mean-Field Games Generalization in mean field games by learning master policies

Reference 76

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:10:33.421759Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-08-07T13:10:24.088295Z digest=sha256:c9f82e16332ce71e6e71c107bea8011554a87b8dc60037f06cc3d2b1f90037ee

Observation 6f072f44-c0b5-4a69-907c-fc5f15848068 · outbound

This paper cites Relative entropy policy search.

Finite-Sample Convergence Bounds for Trust Region Policy Optimization in Mean-Field Games Relative entropy policy search

Reference 77

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:10:33.137548Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-08-07T13:10:24.280777Z digest=sha256:2c0d8f94f3f07f17a9024db65355294114f1b1c41c6afde9e65fa2b96dd0c43d

Observation a6d8c63e-2826-4cb1-8ef4-c27bdf46da08 · outbound

This paper cites an unresolved cited work.

Finite-Sample Convergence Bounds for Trust Region Policy Optimization in Mean-Field Games Unresolved cited work

Reference 78

Resolution
unresolved
raw_fallback, observed 2026-08-07T13:10:32.954863Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-08-07T13:10:24.415989Z digest=sha256:70c389e2278f8dc8a1969fff02a818736120193dcea5548f1726673266e3b553

Observation b27f4cb1-4c6d-4247-9f29-14222d27e36e · outbound

This paper cites Risk-averse dynamic programming for markov decision processes.

Finite-Sample Convergence Bounds for Trust Region Policy Optimization in Mean-Field Games Risk-averse dynamic programming for markov decision processes

Reference 79

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:10:32.785228Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-08-07T13:10:24.501979Z digest=sha256:7b62f067b16bfc2971b8e17c781538abcbd6e0b2f7a0cbad56da20f5abda627b

Observation 78eb7089-8b69-4d91-8dea-1905199a3f2a · outbound

This paper cites Discrete-time average-cost mean-field games on polish spaces.

Finite-Sample Convergence Bounds for Trust Region Policy Optimization in Mean-Field Games Discrete-time average-cost mean-field games on polish spaces

Reference 80

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:10:32.563548Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-08-07T13:10:24.662129Z digest=sha256:bdb5ece792fbc6db086802a7d655ad242535b14507938ad1d5364e53a9879292

Observation a06817dc-94c0-40e9-b3b1-aced4a8a7c3e · outbound

This paper cites The StarCraft Multi-Agent Challenge.

Finite-Sample Convergence Bounds for Trust Region Policy Optimization in Mean-Field Games The StarCraft Multi-Agent Challenge

Reference 81

Resolution
unresolved
no resolver link, observed 2026-08-07T13:10:24.777756Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:10:24.777756Z digest=sha256:8c8d2d11bb762210d904b578b8688c5f4a8e0757caecdf492d6bc4cb1e92c115

Observation 27eef3f0-dde0-4a7e-9fbb-adfc6193a832 · outbound

This paper cites and Geist, M.

Finite-Sample Convergence Bounds for Trust Region Policy Optimization in Mean-Field Games and Geist, M

Reference 82

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:10:32.297226Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-08-07T13:10:24.904569Z digest=sha256:2121764af980260193a142a4d37fccd3abd152aa788262f67e98b70aca833919

Observation 17cde2f0-fc39-4990-83b9-63c2af12ed7f · outbound

This paper cites Trust region policy optimization.

Finite-Sample Convergence Bounds for Trust Region Policy Optimization in Mean-Field Games Trust region policy optimization

Reference 83

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:10:32.042431Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-08-07T13:10:25.085428Z digest=sha256:5a8aeaa9a8402ea56c6512c8036ad483904d46c5f485a59c3cd463170fee48b2

Observation e163f9de-fd01-4a78-be95-55c051baff8d · outbound

This paper cites Safe, Multi-Agent, Reinforcement Learning for Autonomous Driving.

Finite-Sample Convergence Bounds for Trust Region Policy Optimization in Mean-Field Games Safe, Multi-Agent, Reinforcement Learning for Autonomous Driving

Reference 84

Resolution
unresolved
no resolver link, observed 2026-08-07T13:10:25.197393Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:10:25.197393Z digest=sha256:c4faec34858840f8eaee264f7feca72b88286d6183ba2632d2a7dbc925501df2

Observation 62a70fbc-ee7d-4dfb-82e9-fedcc03f6342 · outbound

This paper cites Adaptive trust region policy optimization: Global convergence and faster rates for regularized mdps.

Finite-Sample Convergence Bounds for Trust Region Policy Optimization in Mean-Field Games Adaptive trust region policy optimization: Global convergence and faster rates for regularized mdps

Reference 85

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:10:31.863235Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-08-07T13:10:25.356684Z digest=sha256:2ce906960b5694e615da43edd2d89f0df5b58141703130200b01a16d4dae16e8

Observation c61aa626-fc8a-479d-9489-b171f6d3fac2 · outbound

This paper cites Near-optimal time and sample complexities for solving markov decision processes with a generative model.

Finite-Sample Convergence Bounds for Trust Region Policy Optimization in Mean-Field Games Near-optimal time and sample complexities for solving markov decision processes with a generative model

Reference 86

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:10:31.502983Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-08-07T13:10:25.492393Z digest=sha256:284074f370e86cea14536a5989a30a384ecb97834d6ab13b30dcbfb5ed206951

Observation c285cf1f-31d3-422b-9690-aae4aa618c9a · outbound

This paper cites and Barto, A.

Finite-Sample Convergence Bounds for Trust Region Policy Optimization in Mean-Field Games and Barto, A

Reference 87

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:10:31.332007Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-08-07T13:10:25.663184Z digest=sha256:6934528072b8bf12d753e7d677ddc4e61c83ed80b663f24d63cd6ccb82ec50ee

Observation bd5dc256-aa0c-4404-8f4f-3f0e0ae72acd · outbound

This paper cites an unresolved cited work.

Finite-Sample Convergence Bounds for Trust Region Policy Optimization in Mean-Field Games Unresolved cited work

Reference 88

Resolution
unresolved
raw_fallback, observed 2026-08-07T13:10:31.146473Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-08-07T13:10:25.789076Z digest=sha256:3d5ca3662a8444747fceb0a98c48892e5aae6ffa7b0d5264cb22fb6e17a7f779

Observation 8f602fc9-608c-4b01-8d32-15962ae3e3ba · outbound

This paper cites Algorithms for reinforcement learning.

Finite-Sample Convergence Bounds for Trust Region Policy Optimization in Mean-Field Games Algorithms for reinforcement learning

Reference 89

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:10:30.866770Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-08-07T13:10:25.976783Z digest=sha256:ca93a4149b8d94c6eb1b2d26f7eeaea0ce698fc33d20d8993d4862dbfbbada8b

Observation d97a3575-cbc2-4494-a256-039da6b3be71 · outbound

This paper cites and Zhou, X.

Finite-Sample Convergence Bounds for Trust Region Policy Optimization in Mean-Field Games and Zhou, X

Reference 90

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:10:30.584759Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-08-07T13:10:26.108732Z digest=sha256:edef46f74ee16becf12cb25886e1f7258f62312e73ee45df7e9ee8e20200067a

Observation d23c640d-f2cb-4332-a07e-46bc412993b2 · outbound

This paper cites Deep learning-based numerical methods for high-dimensional parabolic partial differential equations and backward stochastic differential equations.

Finite-Sample Convergence Bounds for Trust Region Policy Optimization in Mean-Field Games Deep learning-based numerical methods for high-dimensional parabolic partial differential equations and backward stochastic differential equations

Reference 91

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:10:30.349052Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-08-07T13:10:26.241320Z digest=sha256:5e1a2993a75e2cc3415f99f6f250d7a5c3f8123fc49783283eabd5198e80a9d4

Observation 14ce55e6-3de5-4025-be1d-644c4611101b · outbound

This paper cites an unresolved cited work.

Finite-Sample Convergence Bounds for Trust Region Policy Optimization in Mean-Field Games Unresolved cited work

Reference 92

Resolution
unresolved
raw_fallback, observed 2026-08-07T13:10:29.996593Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-08-07T13:10:26.416453Z digest=sha256:2e77dad74b0199e98b949fbe854fd2fcf8ca38f6507e1c3ee760848be414058e

Observation c4948045-4cdc-45ec-9ad2-856d0e229c0e · outbound

This paper cites an unresolved cited work.

Finite-Sample Convergence Bounds for Trust Region Policy Optimization in Mean-Field Games Unresolved cited work

Reference 93

Resolution
unresolved
raw_fallback, observed 2026-08-07T13:10:29.720378Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-08-07T13:10:26.554078Z digest=sha256:1d8b243d0ecd85b50b219ea6726eee12ef3ef17450a3df7f884bb02299713695

Observation 97283845-d13d-4399-bf3c-36ce0d4ffa19 · outbound

This paper cites Policy mirror ascent for efficient and independent learning in mean field games.

Finite-Sample Convergence Bounds for Trust Region Policy Optimization in Mean-Field Games Policy mirror ascent for efficient and independent learning in mean field games

Reference 94

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:10:29.386478Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-08-07T13:10:26.764903Z digest=sha256:dc957a241f37c62d7e9df6ec1c4feaa51ae51cca814e25b9dd7f4a6b4621146b

Observation e4890fa3-9eba-4808-83e8-edca356f2b54 · outbound

This paper cites an unresolved cited work.

Finite-Sample Convergence Bounds for Trust Region Policy Optimization in Mean-Field Games Unresolved cited work

Reference 95

Resolution
unresolved
raw_fallback, observed 2026-08-07T13:10:29.189716Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-08-07T13:10:27.017211Z digest=sha256:4130fa1fb4457cc8f7967002fc804342d037ca6c5e997dc1f3e8c23c144234fd

Observation 83382493-16e4-41c4-bcf8-b5ca657e6f61 · outbound

This paper cites Multi-agent reinforcement learning: A selective overview of theories and algorithms.

Finite-Sample Convergence Bounds for Trust Region Policy Optimization in Mean-Field Games Multi-agent reinforcement learning: A selective overview of theories and algorithms

Reference 96

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:10:29.016638Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-08-07T13:10:27.275330Z digest=sha256:d44a317acdbbe1c0420f0cc5f3e7f8fbb9b5cf44c2d64b1f22b14ccc82a39502

Observation a2658e27-0b40-4d5f-b58c-a51f579f4b39 · outbound

This paper cites an unresolved cited work.

Finite-Sample Convergence Bounds for Trust Region Policy Optimization in Mean-Field Games Unresolved cited work

Reference 97

Resolution
unresolved
no resolver link, observed 2026-08-07T13:10:27.440566Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:10:27.440566Z digest=sha256:4cd61550358771075c955a5e9a95184651d44380b36871a8b2756e9d3d69a512

Observation 010bacfd-047a-4a62-9ea7-646781f58239 · outbound

This paper cites D., Bagnell, J.

Finite-Sample Convergence Bounds for Trust Region Policy Optimization in Mean-Field Games D., Bagnell, J

Reference 98

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:10:28.811262Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-08-07T13:10:27.525688Z digest=sha256:a09686b7c613e0f4a4bcd0629ad0d94d4cc451d369c333e3f431cb46ad5e781d

Observation 26a31921-354f-45b5-b123-62d2532794ca · outbound

This paper cites write newline.

Finite-Sample Convergence Bounds for Trust Region Policy Optimization in Mean-Field Games write newline

Reference 99

Resolution
unresolved
no resolver link, observed 2026-08-07T13:10:27.705945Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:10:27.705945Z digest=sha256:3f391576dc6b283795bca2b08b7cfadc47eeffd9e654ef28d75588b626cb45c7

Pith citing papers

Observation d4fce0ae-be86-4a46-ab5c-2eb9ce8dd9b7 · inbound

Towards Model-Free Learning in Dynamic Population Games: An Application to Karma Economies cites this paper.

Towards Model-Free Learning in Dynamic Population Games: An Application to Karma Economies Finite-Sample Convergence Bounds for Trust Region Policy Optimization in Mean-Field Games

Reference 27

Resolution
verified exact
arxiv_id, observed 2026-05-13T01:27:02.762569Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-05-13T01:17:41.135172Z digest=sha256:bcc22b123f17cb64e94e745efb6343f14f5a1b0a38f6ad09152b37b39717bfee