Pith. sign in

Paper Citation Record · LEDGER

The Actor-Critic Update Order Matters for PPO in Federated Reinforcement Learning

As of 10 August 2026, this Paper Citation Record lists 35 of 35 outbound references and 2 inbound Pith citation observations for arXiv:2506.01261.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2506.01261 v1

Coverage vector

measured 35 of 35 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-07T11:54:26.316448Z

measured 37 of 37 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-09T06:31:02.800959+00:00

measured 2 of 2 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-06-28T11:44:53.211503Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-02T01:36:25.635046Z

Reference resolution

35 of 35 outbound references displayed

  • verified exact0
  • verified fuzzy22
  • unresolved11
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch2

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation f866580d-3a6a-4ac7-b845-022d9dddbfc3 · outbound

This paper cites Fitted q-iteration in continuous action-space mdps.

The Actor-Critic Update Order Matters for PPO in Federated Reinforcement Learning Fitted q-iteration in continuous action-space mdps

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:54:32.085703Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T11:54:23.122203Z digest=sha256:65ae474c59b63a15c5e2cf118c1e8f9dfc31d275202c06d0665e485e1845156e

Observation 7ad5d9ca-845d-4ea8-8c06-04e49da4417b · outbound

This paper cites S., and Guin, S.

The Actor-Critic Update Order Matters for PPO in Federated Reinforcement Learning S., and Guin, S

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-07T11:54:23.284017Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:54:23.284017Z digest=sha256:034c51779dd8c1ad2f241407de622497059b092eb8789726ee551ce0175160a7

Observation c672ef05-d8e8-4300-b67b-2cfaae3b0fc8 · outbound

This paper cites OpenAI Gym.

The Actor-Critic Update Order Matters for PPO in Federated Reinforcement Learning OpenAI Gym

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-07T11:54:23.368455Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:54:23.368455Z digest=sha256:557674752524375e79e99a1760bbb40cfdba451de3a87af8e1b07b7704c3169d

Observation c7cacfe5-bcd0-4e18-b0f7-f160f1699d2f · outbound

This paper cites D., and Wang, Z.

The Actor-Critic Update Order Matters for PPO in Federated Reinforcement Learning D., and Wang, Z

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:54:31.888111Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T11:54:23.514345Z digest=sha256:789ae3b776e1a1812cf0162591fdcd80f13edbb79553b8a58a1a7f066542b717

Observation c8deb5e1-1c7a-4d43-b690-9ae93ad6d974 · outbound

This paper cites W., Hilton, J., Klimov, O., and Schulman, J.

The Actor-Critic Update Order Matters for PPO in Federated Reinforcement Learning W., Hilton, J., Klimov, O., and Schulman, J

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:54:31.641212Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T11:54:23.606041Z digest=sha256:40403f897a1f385d87ca5ffbcc45fedac7458338be559be550b42f6c1d93518d

Observation 300c19b8-a938-4320-b4dd-c41e32cb8cff · outbound

This paper cites Linear off-policy actor-critic.

The Actor-Critic Update Order Matters for PPO in Federated Reinforcement Learning Linear off-policy actor-critic

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:54:31.431771Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T11:54:23.695877Z digest=sha256:b2f08b7392f3964abe980a656d0636a37cde1013033aea940b73fc7615176739

Observation c0ff3ad0-e46a-4700-a7cc-47aed292b5ad · outbound

This paper cites A theoretical analysis of deep q-learning.

The Actor-Critic Update Order Matters for PPO in Federated Reinforcement Learning A theoretical analysis of deep q-learning

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:54:31.246319Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T11:54:23.790666Z digest=sha256:cbea79dc5b3f056f4318dddd4c3ff7881e22fc9190e688d874145cc49e0e6b38

Observation 3a257f89-2ad1-41ae-9a14-5468bd6c9ba5 · outbound

This paper cites an unresolved cited work.

The Actor-Critic Update Order Matters for PPO in Federated Reinforcement Learning Unresolved cited work

Reference 8

Resolution
unresolved
raw_fallback, observed 2026-08-07T11:54:31.047177Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T11:54:23.873294Z digest=sha256:eaccdc1e99533c845bbfc74e969a186e967c2e1b2da64527313b919277d8bd1c

Observation 2e2979de-4976-4bf6-87ac-ef722abf575a · outbound

This paper cites Error propagation for approximate policy and value iteration.

The Actor-Critic Update Order Matters for PPO in Federated Reinforcement Learning Error propagation for approximate policy and value iteration

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:54:30.864069Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T11:54:23.956416Z digest=sha256:d2aaae57a0720d40ba70583788df0f9a1a73f026c1eb6a99a4ac2438c5806e17

Observation 8dbfb1cd-09cf-482d-acd3-d80dfdbb919e · outbound

This paper cites Reinforcement learning with deep energy- based policies.

The Actor-Critic Update Order Matters for PPO in Federated Reinforcement Learning Reinforcement learning with deep energy- based policies

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:54:30.668546Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T11:54:24.009912Z digest=sha256:397b9391a74bd582fecccdd7ff29f58ef16f9c5ac1d65c7dbcc7f63926353263

Observation 358e5bf4-13f5-43e1-9ef0-c75e466949be · outbound

This paper cites Federated reinforcement learning with environment heterogeneity.

The Actor-Critic Update Order Matters for PPO in Federated Reinforcement Learning Federated reinforcement learning with environment heterogeneity

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:54:30.469003Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T11:54:24.059749Z digest=sha256:aca10f770780745640e0026796d5ed07599a7edd8a2040078b4e580b3c096ab9

Observation 12327596-a491-4a38-8cb4-3196ee54b005 · outbound

This paper cites P., Kale, S., Mohri, M., Reddi, S., Stich, S., and Suresh, A.

The Actor-Critic Update Order Matters for PPO in Federated Reinforcement Learning P., Kale, S., Mohri, M., Reddi, S., Stich, S., and Suresh, A

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:54:30.225430Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T11:54:24.138911Z digest=sha256:a8aaef4d15c5ea7cd11f1719ac8e6f29a5e9937848291e0e3abd9c32745f7d15

Observation 66c651a9-28fd-4962-84c3-d6b1b7bb6311 · outbound

This paper cites and Tsitsiklis, J.

The Actor-Critic Update Order Matters for PPO in Federated Reinforcement Learning and Tsitsiklis, J

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:54:30.016740Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T11:54:24.191908Z digest=sha256:ce21382aaa54646e9adc3ac2e8d80c64ab1f9ecb672d2043d5d0e0f83281aa76

Observation ba962be4-b9d1-4ccc-86b7-75adc426ddac · outbound

This paper cites K., Zaheer, M., Sanjabi, M., Talwalkar, A., and Smith, V.

The Actor-Critic Update Order Matters for PPO in Federated Reinforcement Learning K., Zaheer, M., Sanjabi, M., Talwalkar, A., and Smith, V

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:54:29.781855Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T11:54:24.283564Z digest=sha256:3deed534f31876f7430e3c415e2b700446ae8031cc9ff6d3290b8c6ac21b2194

Observation a6420aa0-3851-418a-b662-b8c9ab552015 · outbound

This paper cites On the convergence of fedavg on non-iid data.

The Actor-Critic Update Order Matters for PPO in Federated Reinforcement Learning On the convergence of fedavg on non-iid data

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:54:29.555582Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T11:54:24.398859Z digest=sha256:a6bf83732f7df3644bded98a9a30fe8b864a11819cd445220c35e83962c25d9b

Observation 34a3dda1-b350-45d5-bab2-6d89dd56ba48 · outbound

This paper cites Neural trust region/proximal policy optimization attains globally optimal policy.

The Actor-Critic Update Order Matters for PPO in Federated Reinforcement Learning Neural trust region/proximal policy optimization attains globally optimal policy

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:54:29.346041Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T11:54:24.532711Z digest=sha256:f891e20f2f05f24f5630be4a45b9c243157252e8c64e6525f554de24f86d40f4

Observation 65f24c31-4e0e-473a-ab9e-ebeb3b62479b · outbound

This paper cites an unresolved cited work.

The Actor-Critic Update Order Matters for PPO in Federated Reinforcement Learning Unresolved cited work

Reference 17

Resolution
unresolved
raw_fallback, observed 2026-08-07T11:54:29.152442Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T11:54:24.593013Z digest=sha256:85983196883196ea6e745759b9c554858e9667196c4313badfd9fb15cb39289f

Observation 709cd9a4-1602-451f-85c3-1e542b3155fc · outbound

This paper cites On the global convergence rates of softmax policy gradient methods.

The Actor-Critic Update Order Matters for PPO in Federated Reinforcement Learning On the global convergence rates of softmax policy gradient methods

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:54:28.968600Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T11:54:24.703231Z digest=sha256:c224210778ced7d725d2e0ba0a6fa5a917a971c6c1dc36286c5f08fb9e01ba59

Observation b5d387ed-4cd7-43e2-8d98-86763ca50cc6 · outbound

This paper cites A., Veness, J., Bellemare, M.

The Actor-Critic Update Order Matters for PPO in Federated Reinforcement Learning A., Veness, J., Bellemare, M

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:54:28.764589Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T11:54:24.807776Z digest=sha256:458165472dc21273501b147997b8cb335716cebfe3ec3246200398d21f63554b

Observation 29a9c553-d78e-438d-bf62-611abe8ae074 · outbound

This paper cites and Szepesvári, C.

The Actor-Critic Update Order Matters for PPO in Federated Reinforcement Learning and Szepesvári, C

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:54:28.584788Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T11:54:24.918651Z digest=sha256:0ce13012d1d36fcc8cb96d0ce81a048e1d9554d81c9614db8281076a9f13a407

Observation 9718f0b3-1123-4839-9bce-329c0def2396 · outbound

This paper cites Planet dump retrieved from https://planet.osm.org.

The Actor-Critic Update Order Matters for PPO in Federated Reinforcement Learning Planet dump retrieved from https://planet.osm.org

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:54:28.421351Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T11:54:25.028428Z digest=sha256:f7bdf49c08d1437b90ad41b4c4128f0ac6b16e16b9df5c032731064c20a14c69

Observation 12f480a1-5c6c-4313-b891-fa47167bf3b7 · outbound

This paper cites Trust Region Policy Optimization.

The Actor-Critic Update Order Matters for PPO in Federated Reinforcement Learning Trust Region Policy Optimization

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-07T11:54:25.098661Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:54:25.098661Z digest=sha256:89eb492b5ecd22fc787cf9dd65f7eccc596c8a3ee2e4143abe2a0064cfe37713

Observation 9292f581-f4b7-4a79-858f-3cf4e5080240 · outbound

This paper cites Proximal Policy Optimization Algorithms.

The Actor-Critic Update Order Matters for PPO in Federated Reinforcement Learning Proximal Policy Optimization Algorithms

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-07T11:54:25.165690Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:54:25.165690Z digest=sha256:b0ccdaddf5d045349839677049bb6b9707c166e737a3feb5812f2ca76b5e0a2c

Observation 12536ad8-a079-4510-876b-58edb0fc1fb0 · outbound

This paper cites an unresolved cited work.

The Actor-Critic Update Order Matters for PPO in Federated Reinforcement Learning Unresolved cited work

Reference 24

Resolution
unresolved
raw_fallback, observed 2026-08-07T11:54:28.208016Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T11:54:25.271960Z digest=sha256:424f6196cd6de611713e7349e42e84cce557348e388b5beb2c0872323bd6ca1b

Observation 54b8813e-da05-4c67-8138-8a79fcc5c049 · outbound

This paper cites S., McAllester, D., Singh, S., and Mansour, Y.

The Actor-Critic Update Order Matters for PPO in Federated Reinforcement Learning S., McAllester, D., Singh, S., and Mansour, Y

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:54:27.978909Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T11:54:25.369311Z digest=sha256:574d23687890c46f9ea2da822b9e45b64f7002badb45086b94c4bca06d97b6fb

Observation fd1f72d0-98e4-4f32-9d04-6696078872f9 · outbound

This paper cites Boosted fitted q-iteration.

The Actor-Critic Update Order Matters for PPO in Federated Reinforcement Learning Boosted fitted q-iteration

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:54:27.778136Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T11:54:25.449432Z digest=sha256:e8adf1e346b40f347a28366fcdda5d30171f48520f506775491f14be66b21f82

Observation 5eb2fc9f-e882-4ecd-927b-9ba6abec0656 · outbound

This paper cites Deep reinforcement learning with double q-learning.

The Actor-Critic Update Order Matters for PPO in Federated Reinforcement Learning Deep reinforcement learning with double q-learning

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-07T11:54:25.514852Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:54:25.514852Z digest=sha256:fc4208bc096119fcd30471f0602e84c6cdd35cc4b6e686405f7cf778ff401f7b

Observation 014e2b7f-8963-441d-95b9-07541ce6162d · outbound

This paper cites E., Srivastava, S., Tuia, D., and Falcão, A.

The Actor-Critic Update Order Matters for PPO in Federated Reinforcement Learning E., Srivastava, S., Tuia, D., and Falcão, A

Reference 28

Resolution
metadata mismatch
raw_fallback, observed 2026-08-07T11:54:26.944638Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T11:54:25.614183Z digest=sha256:c93a22f1b10067e8c8d5a3180b8c05d3673129c6a7b1b48a63818852b1ff1d4c

Observation 27b574ad-db67-4b56-b86b-4083bd2da289 · outbound

This paper cites L., Kheterpal, N., Jang, K., Wu, C., Wu, F., Liaw, R., Liang, E., and Bayen, A.

The Actor-Critic Update Order Matters for PPO in Federated Reinforcement Learning L., Kheterpal, N., Jang, K., Wu, C., Wu, F., Liaw, R., Liang, E., and Bayen, A

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:54:27.583825Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T11:54:25.722659Z digest=sha256:52fdb10fece2f5263dbef4ac3e8190b25031462ba43582e390eba99a9efdfe9d

Observation 935a9aa8-71f1-4b99-8138-d8b4689fdcc5 · outbound

This paper cites Neural Policy Gradient Methods: Global Optimality and Rates of Convergence.

The Actor-Critic Update Order Matters for PPO in Federated Reinforcement Learning Neural Policy Gradient Methods: Global Optimality and Rates of Convergence

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-07T11:54:25.833975Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:54:25.833975Z digest=sha256:e564f588e761d90aa13a29d95bbd680df972da0527578a9a16d0e77294264f5d

Observation 98bad912-d5dd-449b-b01d-99871ff1812b · outbound

This paper cites P., and Kakade, S.

The Actor-Critic Update Order Matters for PPO in Federated Reinforcement Learning P., and Kakade, S

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:54:27.391913Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T11:54:25.909375Z digest=sha256:04158d756018beae6bb85b2df8a1860d0d7d134bd67b8b3d16ae4e6dd5239617

Observation 39660861-fed3-4826-b7ea-26838d4e7be3 · outbound

This paper cites and Song, S.

The Actor-Critic Update Order Matters for PPO in Federated Reinforcement Learning and Song, S

Reference 32

Resolution
metadata mismatch
raw_fallback, observed 2026-08-07T11:54:26.697142Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T11:54:26.039203Z digest=sha256:6b97e35093e954d0ca67e9c65896dcf5d91b8bd4372f964763d53461f7b69b2a

Observation 8b3badc4-4a1d-4f43-83be-6a93508fab2d · outbound

This paper cites A General Approach to Adding Differential Privacy to Iterative Training Procedures.

The Actor-Critic Update Order Matters for PPO in Federated Reinforcement Learning A General Approach to Adding Differential Privacy to Iterative Training Procedures

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-07T11:54:26.115849Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:54:26.115849Z digest=sha256:2efc77093180c6ab2b1feebb0a0c31c6fa976d19cf5c7bb6018c7f9489b88e95

Observation 852b693b-da84-4ba9-8f6e-ee8146aed99f · outbound

This paper cites Federated natural policy gradient and actor critic methods for multi-task reinforcement learning.

The Actor-Critic Update Order Matters for PPO in Federated Reinforcement Learning Federated natural policy gradient and actor critic methods for multi-task reinforcement learning

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:54:27.168863Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T11:54:26.220188Z digest=sha256:080be24a50022941188db2016b745a75cff2b443f8060a3da896906671739353

Observation 722dd00c-e413-424a-b058-98232ef9a157 · outbound

This paper cites Federated Learning with Non-IID Data.

The Actor-Critic Update Order Matters for PPO in Federated Reinforcement Learning Federated Learning with Non-IID Data

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-07T11:54:26.316448Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:54:26.316448Z digest=sha256:f5e8b79c190d1d2be23b4030ba471dfc14e4c997c5aec2b749df6f141c07f4b7

Pith citing papers

Observation b32d63e9-5e95-4638-a392-534a5ecc3069 · inbound

Collaborative Yet Personalized Policy Training: Single-Timescale Federated Actor-Critic cites this paper.

Collaborative Yet Personalized Policy Training: Single-Timescale Federated Actor-Critic The Actor-Critic Update Order Matters for PPO in Federated Reinforcement Learning

Reference 48

Resolution
verified exact
arxiv_id, observed 2026-05-15T02:08:29.340835Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-15T02:07:23.862809Z digest=sha256:77cdc379a724c0c64193c72a280ffaf9b7a4ed94b4fcc869238da1b3c2bd6d5c

Observation 1377c879-6903-497f-b020-dba0efa9ead2 · inbound

FGRPO: Federated GRPO with Adaptive Aggregation on Non-IID Data cites this paper.

FGRPO: Federated GRPO with Adaptive Aggregation on Non-IID Data The Actor-Critic Update Order Matters for PPO in Federated Reinforcement Learning

Reference 32

Resolution
metadata mismatch
arxiv_id, observed 2026-07-02T01:36:25.637603Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-06-28T11:44:53.211503Z digest=sha256:ebb46645fed827e150b2df54c7b094d09ef38a31409f37609259f7a5bc24738d