Pith. sign in

Paper Citation Record · LEDGER

Finite-Sample Convergence Bounds for Trust Region Policy Optimization in Mean-Field Games

As of 9 August 2026, this Paper Citation Record lists 99 of 99 outbound references and 1 inbound Pith citation observation for arXiv:2505.22781.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2505.22781 v1

Coverage vector

measured 99 of 99 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-07T13:10:27.705945Z

measured 100 of 100 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-08T06:32:00.761636+00:00

measured 1 of 1 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-05-13T01:17:41.135172Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-05-13T01:27:02.760603Z

Reference resolution

99 of 99 outbound references displayed

  • verified exact2
  • verified fuzzy54
  • unresolved43
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation bf637fce-5932-47bb-a84d-fccde22410ae · outbound

This paper cites and Capuzzo-Dolcetta, I.

Finite-Sample Convergence Bounds for Trust Region Policy Optimization in Mean-Field Games and Capuzzo-Dolcetta, I

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-07T13:10:12.196564Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:10:12.196564Z digest=sha256:e59781bcc82ad2fe985fdf3e139c30b0a0eae95e06ad28c1d2361907364a0cf4

Observation 4a6670e5-2aa5-40cc-b372-5d783f349faa · outbound

This paper cites and Porretta, A.

Finite-Sample Convergence Bounds for Trust Region Policy Optimization in Mean-Field Games and Porretta, A

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-07T13:10:12.329482Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:10:12.329482Z digest=sha256:083913b70451116fa183c8b7504ae22c530c0a366a988218b2d4a982dec51a51

Observation 453d78a0-a8be-4b55-9a06-4050d75d6c03 · outbound

This paper cites Mean field games: numerical methods for the planning problem.

Finite-Sample Convergence Bounds for Trust Region Policy Optimization in Mean-Field Games Mean field games: numerical methods for the planning problem

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-07T13:10:12.491661Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:10:12.491661Z digest=sha256:3e093bb7e1f89ddbacba40a9d0ed20d443756d03f1003a4bc060db7843e4d18f

Observation ccbebb40-1486-421b-adef-e6689a6eff64 · outbound

This paper cites Mean field games and applications: Numerical aspects.

Finite-Sample Convergence Bounds for Trust Region Policy Optimization in Mean-Field Games Mean field games and applications: Numerical aspects

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-07T13:10:12.619752Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:10:12.619752Z digest=sha256:51987169be3809e00bd4aaedfd4799b4a108b9be7cb1947eeea6ac9a145fc1a4

Observation 50fb8c21-04b9-48c5-9ee9-c258f2db1281 · outbound

This paper cites an unresolved cited work.

Finite-Sample Convergence Bounds for Trust Region Policy Optimization in Mean-Field Games Unresolved cited work

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-07T13:10:12.806388Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:10:12.806388Z digest=sha256:fcc3aaec843741b5046fc58e8cb3309fbb52c2734227119c585eb9ac8a57a923

Observation 99b8884a-b888-4366-8840-6503594e9d34 · outbound

This paper cites An extended mean field game for storage in smart grids.

Finite-Sample Convergence Bounds for Trust Region Policy Optimization in Mean-Field Games An extended mean field game for storage in smart grids

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-07T13:10:13.017801Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:10:13.017801Z digest=sha256:44a0899d4700c2931c01db6453d6859970bf425708d1956a335e92be46ca742d

Observation 5c2ea739-69c9-4425-b6bb-f1b752ee86ad · outbound

This paper cites Regularization of the policy updates for stabilizing mean field games.

Finite-Sample Convergence Bounds for Trust Region Policy Optimization in Mean-Field Games Regularization of the policy updates for stabilizing mean field games

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-07T13:10:13.221565Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:10:13.221565Z digest=sha256:752891d3aed2e33107a96279d9d1355b7197003f3e7da1a6f570edec7e55b2dc

Observation 68a8815d-4929-4184-b623-1ef5e7930c55 · outbound

This paper cites D., and Saldi, N.

Finite-Sample Convergence Bounds for Trust Region Policy Optimization in Mean-Field Games D., and Saldi, N

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-07T13:10:13.373064Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:10:13.373064Z digest=sha256:4c34b60368c4cbf553c9969d77900a47aa84236ecf2e6117ccdd8f0fa0b906d3

Observation 5f18d09d-74d5-4ad8-aa92-cd8971e47d3c · outbound

This paper cites D., and Saldi, N.

Finite-Sample Convergence Bounds for Trust Region Policy Optimization in Mean-Field Games D., and Saldi, N

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-07T13:10:13.491777Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:10:13.491777Z digest=sha256:957b7683ab457c6e6d607a2510a32775a07e4da64e835123cff28dc6273504d6

Observation 29ad1f48-6d67-45d6-8117-15ccfb0cfe63 · outbound

This paper cites Mean-field sampling for cooperative multi-agent reinforcement learning.

Finite-Sample Convergence Bounds for Trust Region Policy Optimization in Mean-Field Games Mean-field sampling for cooperative multi-agent reinforcement learning

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-07T13:10:13.621544Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:10:13.621544Z digest=sha256:7c37bd849556d7b141dc511c1522b6e35f0e442bb2ab3444504d8c1d216108b1

Observation 570de5b3-c6e6-4ed4-84d7-b2cd89a7bb11 · outbound

This paper cites Unified Reinforcement Q-Learning for Mean Field Game and Control Problems.

Finite-Sample Convergence Bounds for Trust Region Policy Optimization in Mean-Field Games Unified Reinforcement Q-Learning for Mean Field Game and Control Problems

Reference 11

Resolution
verified exact
local_arxiv, observed 2026-08-07T13:10:28.517786Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-07T13:10:13.774546Z digest=sha256:fad378e46a277f0ac293051d409a785fd03d464aabdd127c1f885284f008b473

Observation a115e582-f4c1-4f97-874b-d709fb4e066e · outbound

This paper cites Convergence of Multi-Scale Reinforcement Q-Learning Algorithms for Mean Field Game and Control Problems.

Finite-Sample Convergence Bounds for Trust Region Policy Optimization in Mean-Field Games Convergence of Multi-Scale Reinforcement Q-Learning Algorithms for Mean Field Game and Control Problems

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-07T13:10:13.932048Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:10:13.932048Z digest=sha256:3dd49144f8e69989d21f90e3fc307223b0c7c9f76fb20e19af38eefc1901e14e

Observation d9faeb28-e1ab-46f7-9b5b-4d91e7518fbd · outbound

This paper cites On solutions of mean field games with ergodic cost.

Finite-Sample Convergence Bounds for Trust Region Policy Optimization in Mean-Field Games On solutions of mean field games with ergodic cost

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-07T13:10:14.074540Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:10:14.074540Z digest=sha256:7b9330529f3e9c0948f7acae60395b3f5f1409c42411ab97ab3ca898e720bdcc

Observation 7c38864c-386d-4df4-b653-44aafd25fbe2 · outbound

This paper cites Lipschitz continuity in model-based reinforcement learning.

Finite-Sample Convergence Bounds for Trust Region Policy Optimization in Mean-Field Games Lipschitz continuity in model-based reinforcement learning

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-07T13:10:14.243790Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:10:14.243790Z digest=sha256:8d91251da7387d26dacb78679062daeba0b9bcaa4d03340e5f85ae3a747608cf

Observation d1c12000-e29d-4a60-afc9-9eed044892e0 · outbound

This paper cites Inapproximability of np-complete variants of nash equilibrium.

Finite-Sample Convergence Bounds for Trust Region Policy Optimization in Mean-Field Games Inapproximability of np-complete variants of nash equilibrium

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-07T13:10:14.377895Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:10:14.377895Z digest=sha256:07477e1b22e15b121b1125856b936d95a49f4a53af8aa740693025868bfc0d71

Observation cb65bbb7-a39b-440f-8e8e-874cfdae2ee1 · outbound

This paper cites M., Munos, R., and Kappen, H.

Finite-Sample Convergence Bounds for Trust Region Policy Optimization in Mean-Field Games M., Munos, R., and Kappen, H

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-07T13:10:14.540084Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:10:14.540084Z digest=sha256:0d61d9774996956665ebc365558ee5ab2e54ade06327149cd3c18cefda5eb43c

Observation d527bfc3-31dc-47f3-9bcb-745d3b78ebf6 · outbound

This paper cites and Priuli, F.

Finite-Sample Convergence Bounds for Trust Region Policy Optimization in Mean-Field Games and Priuli, F

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-07T13:10:14.688053Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:10:14.688053Z digest=sha256:582b0b971178db658e38ae0cf26e382af23a89d816f555d39412b587c7858cec

Observation 0bb83870-e706-48a0-8f24-a83d215f526a · outbound

This paper cites A mean-field game model of electricity market dynamics.

Finite-Sample Convergence Bounds for Trust Region Policy Optimization in Mean-Field Games A mean-field game model of electricity market dynamics

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-07T13:10:14.886881Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:10:14.886881Z digest=sha256:e68459a62d0817a61197d7f47e731a96c83e3e02f43a54795a2cb40262fca1de

Observation 7a13ea33-c593-4ce7-b35b-3ddd47b16e40 · outbound

This paper cites and Hesse, S.

Finite-Sample Convergence Bounds for Trust Region Policy Optimization in Mean-Field Games and Hesse, S

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-07T13:10:15.037956Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:10:15.037956Z digest=sha256:76e3d1fa4cfd0349c3eb583676f4f797afd95235cec4ad83c2f91ca7b2bc368e

Observation 7d1e2b1d-41c2-481a-9346-d99ba8da6420 · outbound

This paper cites First-order methods in optimization.

Finite-Sample Convergence Bounds for Trust Region Policy Optimization in Mean-Field Games First-order methods in optimization

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-07T13:10:15.203972Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:10:15.203972Z digest=sha256:30542a93da86bc803a0b07e6e75a91ddd1c6287415bfebde477b9c54c042d34b

Observation be756859-2ab8-4cf2-a5ca-494054b2bbe8 · outbound

This paper cites and Russo, D.

Finite-Sample Convergence Bounds for Trust Region Policy Optimization in Mean-Field Games and Russo, D

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-07T13:10:15.404955Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:10:15.404955Z digest=sha256:c5267005a1dca5209e46d223ec89f2a09f8dd8ae15ac05965d31bd6812820e1a

Observation 48e4c65e-cc51-4081-b65b-5e4d78f9ff8f · outbound

This paper cites A., Ortega, P.

Finite-Sample Convergence Bounds for Trust Region Policy Optimization in Mean-Field Games A., Ortega, P

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-07T13:10:15.613780Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:10:15.613780Z digest=sha256:95f28d5b2434a1d0bc3046f85f328340a4e7de6a6b45e4f56f79350400762ede

Observation 6bd91a86-fb64-40ac-8366-1dfc1730f3f9 · outbound

This paper cites The master equation and the convergence problem in mean field games:(ams-201).

Finite-Sample Convergence Bounds for Trust Region Policy Optimization in Mean-Field Games The master equation and the convergence problem in mean field games:(ams-201)

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-07T13:10:15.729213Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:10:15.729213Z digest=sha256:3cfe437dc494c52110d72d7ea30def4dd06e67f42641908f9a200a335e0798ce

Observation 7dce5b72-226e-41ef-bbc0-72f17a6825ad · outbound

This paper cites and Lauri \`e re, M.

Finite-Sample Convergence Bounds for Trust Region Policy Optimization in Mean-Field Games and Lauri \`e re, M

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:10:45.706097Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-07T13:10:15.885470Z digest=sha256:d5bc98e53acb2c36c8d0289b913f1636cc713974969d30c4ddb41ac252c63af4

Observation 9096057d-4782-48ed-a291-3e6a5577120b · outbound

This paper cites Mean Field Games and Systemic Risk.

Finite-Sample Convergence Bounds for Trust Region Policy Optimization in Mean-Field Games Mean Field Games and Systemic Risk

Reference 25

Resolution
verified exact
local_arxiv, observed 2026-08-07T13:10:28.223000Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-07T13:10:16.063803Z digest=sha256:b2789071474a0588c58500e552720db62415b5ad4cc3164d70e0459f6121dd71

Observation db6ebada-1185-4011-b4e7-f9bebad60ca4 · outbound

This paper cites Probabilistic theory of mean field games with applications I-II.

Finite-Sample Convergence Bounds for Trust Region Policy Optimization in Mean-Field Games Probabilistic theory of mean field games with applications I-II

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:10:45.413359Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-07T13:10:16.251055Z digest=sha256:09594ffcf409896952d3aec14a289a09d8156549ea539370bb0a884c989c487b

Observation 25424109-0fc2-4849-905e-af23cd3e2c59 · outbound

This paper cites Numerical method for fbsdes of mckean--vlasov type.

Finite-Sample Convergence Bounds for Trust Region Policy Optimization in Mean-Field Games Numerical method for fbsdes of mckean--vlasov type

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:10:45.138057Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-07T13:10:16.373548Z digest=sha256:55ca212346109da719f5c8c07b5192929d273720527adb6ec9cfbcccaa594bf0

Observation 6b9d0135-894c-417e-83a5-c4d104926724 · outbound

This paper cites an unresolved cited work.

Finite-Sample Convergence Bounds for Trust Region Policy Optimization in Mean-Field Games Unresolved cited work

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-07T13:10:16.529488Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:10:16.529488Z digest=sha256:f2b0792fd0c53504efadf6c5c359bf0920ab818ee07d5be114752191ec1b2ed1

Observation 2e93f1e8-f4a9-473f-a040-9bb9bc1dec24 · outbound

This paper cites and Koeppl, H.

Finite-Sample Convergence Bounds for Trust Region Policy Optimization in Mean-Field Games and Koeppl, H

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:10:44.865814Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-07T13:10:16.744269Z digest=sha256:5a0e4730b343e2ef053bbbb4f24700985058657d0c9196ae0eba0f6b483585f3

Observation 27bbb1b0-3aba-4b92-9e01-73ad00a0db6d · outbound

This paper cites and Koeppl, H.

Finite-Sample Convergence Bounds for Trust Region Policy Optimization in Mean-Field Games and Koeppl, H

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:10:44.565375Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-07T13:10:16.855232Z digest=sha256:b3b97b3f5c595df12a8a338af1b5ae54e0f53d8c32ca09e56918473c4152c595

Observation dec370c1-d457-45b3-bb97-7d9744f31cb0 · outbound

This paper cites A mean field game analysis of sir dynamics with vaccination.

Finite-Sample Convergence Bounds for Trust Region Policy Optimization in Mean-Field Games A mean field game analysis of sir dynamics with vaccination

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:10:44.298047Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-07T13:10:17.038129Z digest=sha256:40b2c7fc903a81ec4f0119a76e108e6f8a04443e22d68f13021b66f14c3ff0f7

Observation dd77e7cb-60ec-45b4-9aab-1a478328e33f · outbound

This paper cites and Touzi, N.

Finite-Sample Convergence Bounds for Trust Region Policy Optimization in Mean-Field Games and Touzi, N

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:10:43.918293Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-07T13:10:17.232176Z digest=sha256:df25f6783c321f511ca1ee303bce89cadc901a779f698a8a91f4947355d766fa

Observation f47c2e38-69af-4291-97e3-dfd0e973e66b · outbound

This paper cites and Silva, F.

Finite-Sample Convergence Bounds for Trust Region Policy Optimization in Mean-Field Games and Silva, F

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:10:43.646217Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-07T13:10:17.431580Z digest=sha256:e4a5931d89f410427131936a5f25fa5735231ab84c6f89b6b043c92374a6e5d8

Observation 97469166-6922-49d8-8dd6-e7e430524278 · outbound

This paper cites N-player games and mean field games of moderate interactions.

Finite-Sample Convergence Bounds for Trust Region Policy Optimization in Mean-Field Games N-player games and mean field games of moderate interactions

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:10:43.356158Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-07T13:10:17.587562Z digest=sha256:828c47a77ff26f2c2fff04e0d1117a0f0e854f790211190e30ced8c2a6756208

Observation 023d22f7-2603-4386-8fbe-f1dde2ae9152 · outbound

This paper cites Counterfactual multi-agent policy gradients.

Finite-Sample Convergence Bounds for Trust Region Policy Optimization in Mean-Field Games Counterfactual multi-agent policy gradients

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:10:43.097433Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-07T13:10:17.719823Z digest=sha256:0da2ffa3982e4a218a5ce6ea1865fdc532ad980e489ac2a392fe0611b4d99e42

Observation 79bbc963-3836-4e96-9a35-0f8520672c3c · outbound

This paper cites Convergence of adaptive and interacting markov chain monte carlo algorithms.

Finite-Sample Convergence Bounds for Trust Region Policy Optimization in Mean-Field Games Convergence of adaptive and interacting markov chain monte carlo algorithms

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:10:42.793131Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-07T13:10:17.855496Z digest=sha256:fb0e46b2b7ca4c6d497b88bf3a00bdafdb7c332865b12225b01302eeebd70914

Observation 03d71b27-7be9-4f46-a629-b5e3f2106beb · outbound

This paper cites Taming the noise in reinforcement learning via soft updates.

Finite-Sample Convergence Bounds for Trust Region Policy Optimization in Mean-Field Games Taming the noise in reinforcement learning via soft updates

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:10:42.494859Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-07T13:10:18.009623Z digest=sha256:1e2ac2d130f45c2c9de6a34358df7740a06984190c9e5d457124c34bff717916

Observation b1b27f0a-1f7f-4f2d-9c0f-95401d9ff0e1 · outbound

This paper cites A theory of regularized markov decision processes.

Finite-Sample Convergence Bounds for Trust Region Policy Optimization in Mean-Field Games A theory of regularized markov decision processes

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-07T13:10:18.193450Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:10:18.193450Z digest=sha256:9f1e682f6ea819c6172d08fcbc7420de8d8bd09f168ed3773bf599184046d5b3

Observation 78fb16c5-db85-4513-951f-93cf8e0bbccb · outbound

This paper cites Concave Utility Reinforcement Learning: the Mean-Field Game Viewpoint.

Finite-Sample Convergence Bounds for Trust Region Policy Optimization in Mean-Field Games Concave Utility Reinforcement Learning: the Mean-Field Game Viewpoint

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-07T13:10:18.346868Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:10:18.346868Z digest=sha256:d4c02233a169632bec02a0d0b7f15ed33584135418d645d552268e13cdbfff9d

Observation d69e59ea-a95b-4796-8bf2-d7fbc8088ae4 · outbound

This paper cites Numerical resolution of mckean-vlasov fbsdes using neural networks.

Finite-Sample Convergence Bounds for Trust Region Policy Optimization in Mean-Field Games Numerical resolution of mckean-vlasov fbsdes using neural networks

Reference 40

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:10:42.256776Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-07T13:10:18.546112Z digest=sha256:aa0daab867d6ce7f23ba8cdd0f26a3a73f375c8f150de27a78a93a8255b9954a

Observation 3a1a6c36-559e-48d7-94c7-a7e450b1a6bf · outbound

This paper cites and Diepold, K.

Finite-Sample Convergence Bounds for Trust Region Policy Optimization in Mean-Field Games and Diepold, K

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-07T13:10:18.689464Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:10:18.689464Z digest=sha256:295de81571772d7d22f46fe2f5ddacfa89cea65558f2e326d18cb08fa810272a

Observation 5bfd2c8a-fd34-4a50-ae8f-e06c4b1236c4 · outbound

This paper cites Learning mean-field games.

Finite-Sample Convergence Bounds for Trust Region Policy Optimization in Mean-Field Games Learning mean-field games

Reference 42

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:10:41.972833Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-07T13:10:18.864106Z digest=sha256:2de31495f1a6b5409fe86575ff5816f06a836a248349bbeb51159981cac8ebbf

Observation ccd5ec80-907a-4e95-8d0e-98e07be6c675 · outbound

This paper cites A general framework for learning mean-field games.

Finite-Sample Convergence Bounds for Trust Region Policy Optimization in Mean-Field Games A general framework for learning mean-field games

Reference 43

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:10:41.639343Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-07T13:10:19.028275Z digest=sha256:937e5fa76e1b31cb4699d88171ef2cf874cdbd5a3718f11b5e1f9f0285ff3988

Observation e65f3309-8d85-4952-bc78-5a6fae295914 · outbound

This paper cites Reinforcement learning with deep energy-based policies.

Finite-Sample Convergence Bounds for Trust Region Policy Optimization in Mean-Field Games Reinforcement learning with deep energy-based policies

Reference 44

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:10:41.336299Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-07T13:10:19.165381Z digest=sha256:f5e0828258ea485af0704f26b9b6f104046f037a4010dd2db6936961e2de667e

Observation 5b074396-72b5-4276-ba77-c170a92c8aa2 · outbound

This paper cites J., Liaw, C., Plan, Y., and Randhawa, S.

Finite-Sample Convergence Bounds for Trust Region Policy Optimization in Mean-Field Games J., Liaw, C., Plan, Y., and Randhawa, S

Reference 45

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:10:40.989906Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-07T13:10:19.339607Z digest=sha256:a285ea809e7b8b49b7c2b8b48401c114c1e7cbac001d5a22485243f38847f53d

Observation 94f56b99-100a-4721-8c09-56180e6a5ef6 · outbound

This paper cites an unresolved cited work.

Finite-Sample Convergence Bounds for Trust Region Policy Optimization in Mean-Field Games Unresolved cited work

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-07T13:10:19.489673Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:10:19.489673Z digest=sha256:277e5aebd1adddf198eabdebcb222afa19f6b1bcc2c85b49a9710f903f8b0280

Observation f8a7b82f-be21-40f2-bce0-cb5781e8892a · outbound

This paper cites an unresolved cited work.

Finite-Sample Convergence Bounds for Trust Region Policy Optimization in Mean-Field Games Unresolved cited work

Reference 47

Resolution
unresolved
raw_fallback, observed 2026-08-07T13:10:40.725380Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-07T13:10:19.640667Z digest=sha256:f2289ce9e9cdc1d20aaa608fb4e7b1bcf33106673d12276238265d3cfa118d21

Observation 0a887376-c623-4d30-a983-83a98eab0cb3 · outbound

This paper cites E., and Malham \'e , R.

Finite-Sample Convergence Bounds for Trust Region Policy Optimization in Mean-Field Games E., and Malham \'e , R

Reference 48

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:10:40.511051Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-07T13:10:19.736199Z digest=sha256:c4ac5caccabafe4c2404b9df42bbf6953d2bb8c5aef6055c6d65bb472aef692d

Observation 18f08f07-8014-474d-ad13-e41ec9bc9bf9 · outbound

This paper cites P., and Caines, P.

Finite-Sample Convergence Bounds for Trust Region Policy Optimization in Mean-Field Games P., and Caines, P

Reference 49

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:10:40.194629Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-07T13:10:19.880740Z digest=sha256:cabfbaeb1eded73d45eafa18109ab9df8ad5654a6e9a4a08ad179f701ecb17d0

Observation 39a6dc44-f5cd-4cd2-a412-8b19c3e48907 · outbound

This paper cites P., and Caines, P.

Finite-Sample Convergence Bounds for Trust Region Policy Optimization in Mean-Field Games P., and Caines, P

Reference 50

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:10:39.850916Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-07T13:10:20.085274Z digest=sha256:d74aa9f16570df8faad25f25f5c33a4425d8a9baa219ff490333b4917c695720

Observation db69e7c9-770d-46b8-9d34-f9b69d01db74 · outbound

This paper cites P., and Caines, P.

Finite-Sample Convergence Bounds for Trust Region Policy Optimization in Mean-Field Games P., and Caines, P

Reference 51

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:10:39.591521Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-07T13:10:20.238426Z digest=sha256:27dc9b944e0adb88806970622b6d3e62195a15bbb3577cdf88f19d28e081eff5

Observation 0868478e-fd85-46bb-9d00-5a1b1ab40a13 · outbound

This paper cites and Sha, F.

Finite-Sample Convergence Bounds for Trust Region Policy Optimization in Mean-Field Games and Sha, F

Reference 52

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:10:39.225351Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-07T13:10:20.405360Z digest=sha256:98ce28103a64e325ed3f527e8a166b1bea490934479985fc206f646ac115cafc

Observation 9f637947-4b9e-4903-a1cf-004bc8ed7e5d · outbound

This paper cites and Langford, J.

Finite-Sample Convergence Bounds for Trust Region Policy Optimization in Mean-Field Games and Langford, J

Reference 53

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:10:38.952013Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-07T13:10:20.536988Z digest=sha256:2beb7f04db2f0bcda1b150eb743ea30d7fbc3b40680080c56b9bb86a91a16107

Observation d2cf9ab8-375c-4972-aa89-067c533187b7 · outbound

This paper cites an unresolved cited work.

Finite-Sample Convergence Bounds for Trust Region Policy Optimization in Mean-Field Games Unresolved cited work

Reference 54

Resolution
unresolved
raw_fallback, observed 2026-08-07T13:10:38.682202Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-07T13:10:20.673691Z digest=sha256:186a3429781b8d921eee925557b39e8d4ab82ad5b8a956db8800f10fa7b4c761

Observation cbab3140-898f-4cc1-af1f-5adf410ed2c3 · outbound

This paper cites and Singh, S.

Finite-Sample Convergence Bounds for Trust Region Policy Optimization in Mean-Field Games and Singh, S

Reference 55

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:10:38.365483Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-07T13:10:20.845556Z digest=sha256:d52610d357a231f179c55b04ed70e198a2b3ccb53c00d214515b73553585b741

Observation a733ded7-9ecf-4b7d-a69f-896fc7515207 · outbound

This paper cites A., and Peters, J.

Finite-Sample Convergence Bounds for Trust Region Policy Optimization in Mean-Field Games A., and Peters, J

Reference 56

Resolution
unresolved
no resolver link, observed 2026-08-07T13:10:21.019331Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:10:21.019331Z digest=sha256:7b10a9c22e3635eb518ea1954b158b6baa67a1bc8f2999a6958c143644089a40

Observation b216d78f-32db-4a42-9f27-166e20c2548e · outbound

This paper cites and Zariphopoulou, T.

Finite-Sample Convergence Bounds for Trust Region Policy Optimization in Mean-Field Games and Zariphopoulou, T

Reference 57

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:10:38.039128Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-07T13:10:21.140695Z digest=sha256:4f1f66eaf3b6181810af88f63df0759a3b9b7698214527c8bd36f0688c7d21de

Observation 6137cd43-9af3-45e5-b098-e70a40ca2750 · outbound

This paper cites and Lions, P.-L.

Finite-Sample Convergence Bounds for Trust Region Policy Optimization in Mean-Field Games and Lions, P.-L

Reference 58

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:10:37.682067Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-07T13:10:21.319382Z digest=sha256:d7d7ccfeb4c5b0df76092bbdee5431cc03eb87ca1a70bb7a6d9927159c439f2c

Observation 00e46286-6eca-46ee-98e3-ab074020cf14 · outbound

This paper cites and Lions, P.-L.

Finite-Sample Convergence Bounds for Trust Region Policy Optimization in Mean-Field Games and Lions, P.-L

Reference 59

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:10:37.358005Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-07T13:10:21.475442Z digest=sha256:e039e344d14b16c8e367c8c9a9fb3d13865e074ce46d142c42193fe0d944a19d

Observation dc7ff6ff-efdf-4703-ab72-e57f9deb2496 · outbound

This paper cites and Lions, P.-L.

Finite-Sample Convergence Bounds for Trust Region Policy Optimization in Mean-Field Games and Lions, P.-L

Reference 60

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:10:37.110259Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-07T13:10:21.678545Z digest=sha256:6c7f6adf4b0c60a0ca4c448b2201cc1357edf86adb49c7896696944a36ea3aee

Observation df7cf107-2d2d-46dd-ad91-acc63b4c49c2 · outbound

This paper cites Learning in Mean Field Games: A Survey.

Finite-Sample Convergence Bounds for Trust Region Policy Optimization in Mean-Field Games Learning in Mean Field Games: A Survey

Reference 61

Resolution
unresolved
no resolver link, observed 2026-08-07T13:10:21.805809Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:10:21.805809Z digest=sha256:274be21aa4717101ecc158ccf1d0957d5fc3202077f313960104be8e34b4bc2b

Observation caa7464b-ca39-4d2a-b1a3-52c81f97e0b2 · outbound

This paper cites and Tankov, P.

Finite-Sample Convergence Bounds for Trust Region Policy Optimization in Mean-Field Games and Tankov, P

Reference 62

Resolution
unresolved
no resolver link, observed 2026-08-07T13:10:21.939715Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:10:21.939715Z digest=sha256:428ba328b1dd0ad001351e52ec213db53338250fa999d16e5f03ff5512013ccf

Observation b4ef1a75-91d8-4d15-95cf-65c2afa50d05 · outbound

This paper cites G., and Castro, P.

Finite-Sample Convergence Bounds for Trust Region Policy Optimization in Mean-Field Games G., and Castro, P

Reference 63

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:10:36.778709Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-07T13:10:22.118395Z digest=sha256:c190fbfd8c142d9386667deecfa0c9e861ca791d4f51d9ec5c10acee6658661d

Observation 061924ee-e8d9-42d4-a570-edba6260118a · outbound

This paper cites A mean-field game approach to cloud resource management with function approximation.

Finite-Sample Convergence Bounds for Trust Region Policy Optimization in Mean-Field Games A mean-field game approach to cloud resource management with function approximation

Reference 64

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:10:36.518742Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-07T13:10:22.294882Z digest=sha256:790ef1d4971445fbc859dfc86e7c5dea91778d9841264a7d1de2372bb839f548

Observation e4d9343c-f81b-4127-ac4e-b3612d920a9b · outbound

This paper cites I., Fern \'a ndez-Gaucherand, E., Hern \'a ndez-Hernandez, D., Coraluppi, S., and Fard, P.

Finite-Sample Convergence Bounds for Trust Region Policy Optimization in Mean-Field Games I., Fern \'a ndez-Gaucherand, E., Hern \'a ndez-Hernandez, D., Coraluppi, S., and Fard, P

Reference 65

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:10:36.242074Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-07T13:10:22.422162Z digest=sha256:cf1b093307e9491d907a29e4e3b934a9a804032c323c6c74fecb87ff8f9771b3

Observation 62b0348e-6afa-4d28-a1a0-b1e1d296efc5 · outbound

This paper cites J., and Le Fort-Piat, N.

Finite-Sample Convergence Bounds for Trust Region Policy Optimization in Mean-Field Games J., and Le Fort-Piat, N

Reference 66

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:10:35.938947Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-07T13:10:22.542113Z digest=sha256:684feb2018697e982c98d2752a2ef4078086e7d4955ba53ab6692d23fc394573

Observation c2192078-c026-4f09-9341-aeb0ddfb1cd0 · outbound

This paper cites On the global convergence rates of softmax policy gradient methods.

Finite-Sample Convergence Bounds for Trust Region Policy Optimization in Mean-Field Games On the global convergence rates of softmax policy gradient methods

Reference 67

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:10:35.674906Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-07T13:10:22.629063Z digest=sha256:c04d37d9f451e44c1d8cc6b52b50b3c23c9f93634cad2ef0d82eb1c3c0bd1b7c

Observation 179377ee-779a-40c0-adfe-684e4f89a49c · outbound

This paper cites Asynchronous methods for deep reinforcement learning.

Finite-Sample Convergence Bounds for Trust Region Policy Optimization in Mean-Field Games Asynchronous methods for deep reinforcement learning

Reference 68

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:10:35.407930Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-07T13:10:22.697117Z digest=sha256:7fe4567b9d4c980b66fc91ccd68d5cdbe22020402c67424407198e1bf38c5e0d

Observation deee8e41-6131-4519-83d0-0d81307edf98 · outbound

This paper cites Equilibrium points in n-person games.

Finite-Sample Convergence Bounds for Trust Region Policy Optimization in Mean-Field Games Equilibrium points in n-person games

Reference 69

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:10:35.017098Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-07T13:10:22.825724Z digest=sha256:f7d878eaba28270703858eb99528eb84b600aaddcc58a57b28cfb9a535f1b111

Observation 9459171a-5cd1-4b5f-a64c-ace63906c4e2 · outbound

This paper cites Non-cooperative games.

Finite-Sample Convergence Bounds for Trust Region Policy Optimization in Mean-Field Games Non-cooperative games

Reference 70

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:10:34.652398Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-07T13:10:22.945284Z digest=sha256:46dec04f614168b2b1a5b588133ec14c2be7091ca2684af61941f612f2c8c60b

Observation 8797b717-c186-4857-a0cb-d49e42ae35a9 · outbound

This paper cites A unified view of entropy-regularized Markov decision processes.

Finite-Sample Convergence Bounds for Trust Region Policy Optimization in Mean-Field Games A unified view of entropy-regularized Markov decision processes

Reference 71

Resolution
unresolved
no resolver link, observed 2026-08-07T13:10:23.087809Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:10:23.087809Z digest=sha256:acaa07234090bd3bb6438ef995418dd9bd6b64e7cfd827c73232046191c87db7

Observation d45e711e-b08b-4bf2-b8aa-78fad369ea65 · outbound

This paper cites Combining policy gradient and q-learning.

Finite-Sample Convergence Bounds for Trust Region Policy Optimization in Mean-Field Games Combining policy gradient and q-learning

Reference 72

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:10:34.305505Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-07T13:10:23.258768Z digest=sha256:59402eebc921f2fca71c20fa3379ab00989751386ca2200a10b3ed137d001219

Observation 6a6656f1-46f9-4cf2-a0f6-3db7528d17c0 · outbound

This paper cites Scaling mean field games by online mirror descent.

Finite-Sample Convergence Bounds for Trust Region Policy Optimization in Mean-Field Games Scaling mean field games by online mirror descent

Reference 73

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:10:33.964017Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-07T13:10:23.407117Z digest=sha256:be736e83201965ea25249b98680be664988d918a1e62a5a9a6a84b1df0a68f87

Observation d341e81a-6d43-4e45-9261-cf9318c72e8f · outbound

This paper cites Fictitious play for mean field games: C ontinuous time analysis and applications.

Finite-Sample Convergence Bounds for Trust Region Policy Optimization in Mean-Field Games Fictitious play for mean field games: C ontinuous time analysis and applications

Reference 74

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:10:33.667286Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-07T13:10:23.669181Z digest=sha256:89bb888bc1bf3eed26be924f5a3330f4282e506ff47128a450847949429472a5

Observation 9804e0d7-ba49-4ec8-b59b-1efd1be482c1 · outbound

This paper cites Mean Field Games Flock! The Reinforcement Learning Way.

Finite-Sample Convergence Bounds for Trust Region Policy Optimization in Mean-Field Games Mean Field Games Flock! The Reinforcement Learning Way

Reference 75

Resolution
unresolved
no resolver link, observed 2026-08-07T13:10:23.894935Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:10:23.894935Z digest=sha256:7b32d7f7c65406baa5032f7216eb46dc3170a3a2540610aa735965e2a443619e

Observation ed45f950-7f0f-4a09-9c65-7cf997065bfc · outbound

This paper cites Generalization in mean field games by learning master policies.

Finite-Sample Convergence Bounds for Trust Region Policy Optimization in Mean-Field Games Generalization in mean field games by learning master policies

Reference 76

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:10:33.421759Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-07T13:10:24.088295Z digest=sha256:776743f1118d9fe8e3f9d3fb82c04f118ac3b833baf81f4d2456f9843b1a897d

Observation 6f072f44-c0b5-4a69-907c-fc5f15848068 · outbound

This paper cites Relative entropy policy search.

Finite-Sample Convergence Bounds for Trust Region Policy Optimization in Mean-Field Games Relative entropy policy search

Reference 77

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:10:33.137548Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-07T13:10:24.280777Z digest=sha256:43a59041ed1fccebcfe88c9a7fa3a07fc5fa92d69641c708a6baf4a3d2ef02bd

Observation a6d8c63e-2826-4cb1-8ef4-c27bdf46da08 · outbound

This paper cites an unresolved cited work.

Finite-Sample Convergence Bounds for Trust Region Policy Optimization in Mean-Field Games Unresolved cited work

Reference 78

Resolution
unresolved
raw_fallback, observed 2026-08-07T13:10:32.954863Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-07T13:10:24.415989Z digest=sha256:7c0c3e7e9c02c4279b98b2ced8f7faa483d87bb191d42714be69d8a089e9bb8f

Observation b27f4cb1-4c6d-4247-9f29-14222d27e36e · outbound

This paper cites Risk-averse dynamic programming for markov decision processes.

Finite-Sample Convergence Bounds for Trust Region Policy Optimization in Mean-Field Games Risk-averse dynamic programming for markov decision processes

Reference 79

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:10:32.785228Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-07T13:10:24.501979Z digest=sha256:734a49155777a139b82407042e59e436714f686b0d664fce58372d5c8c013950

Observation 78eb7089-8b69-4d91-8dea-1905199a3f2a · outbound

This paper cites Discrete-time average-cost mean-field games on polish spaces.

Finite-Sample Convergence Bounds for Trust Region Policy Optimization in Mean-Field Games Discrete-time average-cost mean-field games on polish spaces

Reference 80

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:10:32.563548Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-07T13:10:24.662129Z digest=sha256:a27fe780d44d4a33f1f4bc9787332591436b640ac658c6c74f6783b7173ec534

Observation a06817dc-94c0-40e9-b3b1-aced4a8a7c3e · outbound

This paper cites The StarCraft Multi-Agent Challenge.

Finite-Sample Convergence Bounds for Trust Region Policy Optimization in Mean-Field Games The StarCraft Multi-Agent Challenge

Reference 81

Resolution
unresolved
no resolver link, observed 2026-08-07T13:10:24.777756Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:10:24.777756Z digest=sha256:62eaf5811b7bf3cec76845e29132deb66d6683cd6a6e74bcbe96eef1e3bd1dbb

Observation 27eef3f0-dde0-4a7e-9fbb-adfc6193a832 · outbound

This paper cites and Geist, M.

Finite-Sample Convergence Bounds for Trust Region Policy Optimization in Mean-Field Games and Geist, M

Reference 82

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:10:32.297226Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-07T13:10:24.904569Z digest=sha256:6af14de85e0144ac27f83f2f2178d8ab92f2fec51117735474b90b27d80f14a8

Observation 17cde2f0-fc39-4990-83b9-63c2af12ed7f · outbound

This paper cites Trust region policy optimization.

Finite-Sample Convergence Bounds for Trust Region Policy Optimization in Mean-Field Games Trust region policy optimization

Reference 83

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:10:32.042431Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-07T13:10:25.085428Z digest=sha256:ffa34d12e9c7b1277ca2c7037f1cd8859c1f490238d5c5ad5fb82b3916a0ad06

Observation e163f9de-fd01-4a78-be95-55c051baff8d · outbound

This paper cites Safe, Multi-Agent, Reinforcement Learning for Autonomous Driving.

Finite-Sample Convergence Bounds for Trust Region Policy Optimization in Mean-Field Games Safe, Multi-Agent, Reinforcement Learning for Autonomous Driving

Reference 84

Resolution
unresolved
no resolver link, observed 2026-08-07T13:10:25.197393Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:10:25.197393Z digest=sha256:2b41953b006c78a8c8eed53b923e8e205d86117925f77048cb4fd070fa8efb8f

Observation 62a70fbc-ee7d-4dfb-82e9-fedcc03f6342 · outbound

This paper cites Adaptive trust region policy optimization: Global convergence and faster rates for regularized mdps.

Finite-Sample Convergence Bounds for Trust Region Policy Optimization in Mean-Field Games Adaptive trust region policy optimization: Global convergence and faster rates for regularized mdps

Reference 85

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:10:31.863235Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-07T13:10:25.356684Z digest=sha256:e5674509f19e8d52960ba6970cb851bec65e9ecda97b24c92d076f4cbfd6b9c1

Observation c61aa626-fc8a-479d-9489-b171f6d3fac2 · outbound

This paper cites Near-optimal time and sample complexities for solving markov decision processes with a generative model.

Finite-Sample Convergence Bounds for Trust Region Policy Optimization in Mean-Field Games Near-optimal time and sample complexities for solving markov decision processes with a generative model

Reference 86

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:10:31.502983Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-07T13:10:25.492393Z digest=sha256:9a180660483a3c650b9c27d239e5407157e8002bb9a5bfd48e6ec4387d7b7364

Observation c285cf1f-31d3-422b-9690-aae4aa618c9a · outbound

This paper cites and Barto, A.

Finite-Sample Convergence Bounds for Trust Region Policy Optimization in Mean-Field Games and Barto, A

Reference 87

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:10:31.332007Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-07T13:10:25.663184Z digest=sha256:468432079ac9fd486f2edd381848d4422350e9a117e84781c6261e997b038d2a

Observation bd5dc256-aa0c-4404-8f4f-3f0e0ae72acd · outbound

This paper cites an unresolved cited work.

Finite-Sample Convergence Bounds for Trust Region Policy Optimization in Mean-Field Games Unresolved cited work

Reference 88

Resolution
unresolved
raw_fallback, observed 2026-08-07T13:10:31.146473Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-07T13:10:25.789076Z digest=sha256:1096fa3fdbd633897405477e848f9faa2c66162c2071ce1284a7a19874cd67d3

Observation 8f602fc9-608c-4b01-8d32-15962ae3e3ba · outbound

This paper cites Algorithms for reinforcement learning.

Finite-Sample Convergence Bounds for Trust Region Policy Optimization in Mean-Field Games Algorithms for reinforcement learning

Reference 89

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:10:30.866770Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-07T13:10:25.976783Z digest=sha256:260093d17f033f7bbedd781a9709b916659fac8dec1fbfbbe6380585fefc2f47

Observation d97a3575-cbc2-4494-a256-039da6b3be71 · outbound

This paper cites and Zhou, X.

Finite-Sample Convergence Bounds for Trust Region Policy Optimization in Mean-Field Games and Zhou, X

Reference 90

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:10:30.584759Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-07T13:10:26.108732Z digest=sha256:13fc6864c4eb8d810b58c7341da07738b13e660a4edda75bf687992b7301713b

Observation d23c640d-f2cb-4332-a07e-46bc412993b2 · outbound

This paper cites Deep learning-based numerical methods for high-dimensional parabolic partial differential equations and backward stochastic differential equations.

Finite-Sample Convergence Bounds for Trust Region Policy Optimization in Mean-Field Games Deep learning-based numerical methods for high-dimensional parabolic partial differential equations and backward stochastic differential equations

Reference 91

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:10:30.349052Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-07T13:10:26.241320Z digest=sha256:7ca5f3d4feeec887ac33988d5d663e628f88cef7d1601b578b5bb3d49dd644a7

Observation 14ce55e6-3de5-4025-be1d-644c4611101b · outbound

This paper cites an unresolved cited work.

Finite-Sample Convergence Bounds for Trust Region Policy Optimization in Mean-Field Games Unresolved cited work

Reference 92

Resolution
unresolved
raw_fallback, observed 2026-08-07T13:10:29.996593Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-07T13:10:26.416453Z digest=sha256:8f3c7003f79b8e70fcbeb2460d7526f97e9be96dbda46d541970a96a3cbb7fdc

Observation c4948045-4cdc-45ec-9ad2-856d0e229c0e · outbound

This paper cites an unresolved cited work.

Finite-Sample Convergence Bounds for Trust Region Policy Optimization in Mean-Field Games Unresolved cited work

Reference 93

Resolution
unresolved
raw_fallback, observed 2026-08-07T13:10:29.720378Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-07T13:10:26.554078Z digest=sha256:acba269f50cb22686b72f4ce6b593179c9fd0ba765d8ad983f64fbe76a74969d

Observation 97283845-d13d-4399-bf3c-36ce0d4ffa19 · outbound

This paper cites Policy mirror ascent for efficient and independent learning in mean field games.

Finite-Sample Convergence Bounds for Trust Region Policy Optimization in Mean-Field Games Policy mirror ascent for efficient and independent learning in mean field games

Reference 94

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:10:29.386478Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-07T13:10:26.764903Z digest=sha256:39258262ea016416e9414a6c6d8c773dc91191f34e3de1e5e38de997a0ec4111

Observation e4890fa3-9eba-4808-83e8-edca356f2b54 · outbound

This paper cites an unresolved cited work.

Finite-Sample Convergence Bounds for Trust Region Policy Optimization in Mean-Field Games Unresolved cited work

Reference 95

Resolution
unresolved
raw_fallback, observed 2026-08-07T13:10:29.189716Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-07T13:10:27.017211Z digest=sha256:9561d0c055592f74ba952d3f46842861184b8663799c9391f830076a64a6f1a3

Observation 83382493-16e4-41c4-bcf8-b5ca657e6f61 · outbound

This paper cites Multi-agent reinforcement learning: A selective overview of theories and algorithms.

Finite-Sample Convergence Bounds for Trust Region Policy Optimization in Mean-Field Games Multi-agent reinforcement learning: A selective overview of theories and algorithms

Reference 96

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:10:29.016638Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-07T13:10:27.275330Z digest=sha256:455a59542749b4c6bf0499022db371315f6406da1d02c9bdb8fcf7e82bc77629

Observation a2658e27-0b40-4d5f-b58c-a51f579f4b39 · outbound

This paper cites an unresolved cited work.

Finite-Sample Convergence Bounds for Trust Region Policy Optimization in Mean-Field Games Unresolved cited work

Reference 97

Resolution
unresolved
no resolver link, observed 2026-08-07T13:10:27.440566Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:10:27.440566Z digest=sha256:426433d89ff8e73f8b9be4bed71b8d7afd4435baa51a1f6a5ca9dcfc2875261a

Observation 010bacfd-047a-4a62-9ea7-646781f58239 · outbound

This paper cites D., Bagnell, J.

Finite-Sample Convergence Bounds for Trust Region Policy Optimization in Mean-Field Games D., Bagnell, J

Reference 98

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:10:28.811262Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-07T13:10:27.525688Z digest=sha256:46da0021c94eff1a973c18219b732e20e614efff5903d26ef7a9698aae0383ba

Observation 26a31921-354f-45b5-b123-62d2532794ca · outbound

This paper cites write newline.

Finite-Sample Convergence Bounds for Trust Region Policy Optimization in Mean-Field Games write newline

Reference 99

Resolution
unresolved
no resolver link, observed 2026-08-07T13:10:27.705945Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:10:27.705945Z digest=sha256:d4ea16d413dfe56943482c46c46217191e0832cd842ba15ddfb89f402ec61bbb

Pith citing papers

Observation d4fce0ae-be86-4a46-ab5c-2eb9ce8dd9b7 · inbound

Towards Model-Free Learning in Dynamic Population Games: An Application to Karma Economies cites this paper.

Towards Model-Free Learning in Dynamic Population Games: An Application to Karma Economies Finite-Sample Convergence Bounds for Trust Region Policy Optimization in Mean-Field Games

Reference 27

Resolution
verified exact
arxiv_id, observed 2026-05-13T01:27:02.762569Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-13T01:17:41.135172Z digest=sha256:4d7287ff491e6be24b465c8a4a7acf73547f93366498f5e9951415dda22a378b