Pith. sign in

Paper Citation Record · LEDGER

The Actor-Critic Update Order Matters for PPO in Federated Reinforcement Learning

As of 22 August 2026, this Paper Citation Record lists 35 of 35 outbound references and 2 inbound Pith citation observations for arXiv:2506.01261.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2506.01261 v1

Coverage vector

measured 35 of 35 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-07T11:54:26.316448Z

measured 37 of 37 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-21T06:32:19.484+00:00

measured 2 of 2 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-06-28T11:44:53.211503Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-02T01:36:25.635046Z

Reference resolution

35 of 35 outbound references displayed

  • verified exact0
  • verified fuzzy22
  • unresolved11
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch2

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation f866580d-3a6a-4ac7-b845-022d9dddbfc3 · outbound

This paper cites Fitted q-iteration in continuous action-space mdps.

The Actor-Critic Update Order Matters for PPO in Federated Reinforcement Learning Fitted q-iteration in continuous action-space mdps

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:54:32.085703Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-07T11:54:23.122203Z digest=sha256:ef4ff05e6808a0e6b5bdfd62ec0367a840100b58f776e966682141b40b9b762d

Observation 7ad5d9ca-845d-4ea8-8c06-04e49da4417b · outbound

This paper cites S., and Guin, S.

The Actor-Critic Update Order Matters for PPO in Federated Reinforcement Learning S., and Guin, S

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-07T11:54:23.284017Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:54:23.284017Z digest=sha256:ab193291eb476cbfbf42e753a51019c0d3ec2b537ea53aa287cb40c23fd6193e

Observation c672ef05-d8e8-4300-b67b-2cfaae3b0fc8 · outbound

This paper cites OpenAI Gym.

The Actor-Critic Update Order Matters for PPO in Federated Reinforcement Learning OpenAI Gym

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-07T11:54:23.368455Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:54:23.368455Z digest=sha256:6e721183cd5d262e607b32a4b61faef9a0aa84e2fca1e38bbad104454792e4bc

Observation c7cacfe5-bcd0-4e18-b0f7-f160f1699d2f · outbound

This paper cites D., and Wang, Z.

The Actor-Critic Update Order Matters for PPO in Federated Reinforcement Learning D., and Wang, Z

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:54:31.888111Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-07T11:54:23.514345Z digest=sha256:12a1fab092249a15abd8148562271b89e3d581f0795a63a9787bcf2157cd4a71

Observation c8deb5e1-1c7a-4d43-b690-9ae93ad6d974 · outbound

This paper cites W., Hilton, J., Klimov, O., and Schulman, J.

The Actor-Critic Update Order Matters for PPO in Federated Reinforcement Learning W., Hilton, J., Klimov, O., and Schulman, J

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:54:31.641212Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-07T11:54:23.606041Z digest=sha256:4d5af8c325b7b31ba68603f68ab60f404958ffd1ebf2a0d62774dc0dcbd35f57

Observation 300c19b8-a938-4320-b4dd-c41e32cb8cff · outbound

This paper cites Linear off-policy actor-critic.

The Actor-Critic Update Order Matters for PPO in Federated Reinforcement Learning Linear off-policy actor-critic

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:54:31.431771Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-07T11:54:23.695877Z digest=sha256:a96f59b759368a1e88ea4c24a4298d9c19007f1dd4d4119749eb2d4a8be6653e

Observation c0ff3ad0-e46a-4700-a7cc-47aed292b5ad · outbound

This paper cites A theoretical analysis of deep q-learning.

The Actor-Critic Update Order Matters for PPO in Federated Reinforcement Learning A theoretical analysis of deep q-learning

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:54:31.246319Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-07T11:54:23.790666Z digest=sha256:14295a81697ca7e5ae9ca9735591edd2ccfcd7e4898aa8ec6fe6634ccc4a361f

Observation 3a257f89-2ad1-41ae-9a14-5468bd6c9ba5 · outbound

This paper cites an unresolved cited work.

The Actor-Critic Update Order Matters for PPO in Federated Reinforcement Learning Unresolved cited work

Reference 8

Resolution
unresolved
raw_fallback, observed 2026-08-07T11:54:31.047177Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-07T11:54:23.873294Z digest=sha256:c26aebf6a131fcb43ea2fcb765b4378e48114d38e75743c76827419677d921ef

Observation 2e2979de-4976-4bf6-87ac-ef722abf575a · outbound

This paper cites Error propagation for approximate policy and value iteration.

The Actor-Critic Update Order Matters for PPO in Federated Reinforcement Learning Error propagation for approximate policy and value iteration

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:54:30.864069Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-07T11:54:23.956416Z digest=sha256:84f2aaa2bf38062c927fd9408b608cb1f7a752ce782438e2077952f8343176aa

Observation 8dbfb1cd-09cf-482d-acd3-d80dfdbb919e · outbound

This paper cites Reinforcement learning with deep energy- based policies.

The Actor-Critic Update Order Matters for PPO in Federated Reinforcement Learning Reinforcement learning with deep energy- based policies

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:54:30.668546Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-07T11:54:24.009912Z digest=sha256:f6a933c2f8a486a3dc52c03c49fcc2a6985011d49b509993545c112c5338245a

Observation 358e5bf4-13f5-43e1-9ef0-c75e466949be · outbound

This paper cites Federated reinforcement learning with environment heterogeneity.

The Actor-Critic Update Order Matters for PPO in Federated Reinforcement Learning Federated reinforcement learning with environment heterogeneity

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:54:30.469003Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-07T11:54:24.059749Z digest=sha256:1393599d459e56cc4f9e05caa8402f252ba56fcc47b1180f6d2d188b2f165859

Observation 12327596-a491-4a38-8cb4-3196ee54b005 · outbound

This paper cites P., Kale, S., Mohri, M., Reddi, S., Stich, S., and Suresh, A.

The Actor-Critic Update Order Matters for PPO in Federated Reinforcement Learning P., Kale, S., Mohri, M., Reddi, S., Stich, S., and Suresh, A

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:54:30.225430Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-07T11:54:24.138911Z digest=sha256:942d028d056b5def119eaf213ef996288cadee0b770705de1d01a04a79982007

Observation 66c651a9-28fd-4962-84c3-d6b1b7bb6311 · outbound

This paper cites and Tsitsiklis, J.

The Actor-Critic Update Order Matters for PPO in Federated Reinforcement Learning and Tsitsiklis, J

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:54:30.016740Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-07T11:54:24.191908Z digest=sha256:1cbdcb7ac46e394548adc9844baf606f5e72c772ad95eabdad6e5096485dc084

Observation ba962be4-b9d1-4ccc-86b7-75adc426ddac · outbound

This paper cites K., Zaheer, M., Sanjabi, M., Talwalkar, A., and Smith, V.

The Actor-Critic Update Order Matters for PPO in Federated Reinforcement Learning K., Zaheer, M., Sanjabi, M., Talwalkar, A., and Smith, V

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:54:29.781855Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-07T11:54:24.283564Z digest=sha256:4d21b66137a8f72731f7c7cec596055890a3dd15a7405f82ffb66400b7102596

Observation a6420aa0-3851-418a-b662-b8c9ab552015 · outbound

This paper cites On the convergence of fedavg on non-iid data.

The Actor-Critic Update Order Matters for PPO in Federated Reinforcement Learning On the convergence of fedavg on non-iid data

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:54:29.555582Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-07T11:54:24.398859Z digest=sha256:7d089da2ca0ad0fb465faac9caed62922fd1981dc6ab55f8070804ab0aa3ca17

Observation 34a3dda1-b350-45d5-bab2-6d89dd56ba48 · outbound

This paper cites Neural trust region/proximal policy optimization attains globally optimal policy.

The Actor-Critic Update Order Matters for PPO in Federated Reinforcement Learning Neural trust region/proximal policy optimization attains globally optimal policy

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:54:29.346041Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-07T11:54:24.532711Z digest=sha256:e3639631e46f7150973c374995b1c4b29176bbf3ef8ba712cbcdebd4409f6ce4

Observation 65f24c31-4e0e-473a-ab9e-ebeb3b62479b · outbound

This paper cites an unresolved cited work.

The Actor-Critic Update Order Matters for PPO in Federated Reinforcement Learning Unresolved cited work

Reference 17

Resolution
unresolved
raw_fallback, observed 2026-08-07T11:54:29.152442Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-07T11:54:24.593013Z digest=sha256:c1fe2184962e7c89c78317c2a7a625e014e5bae3d6af958a1914711c5a069709

Observation 709cd9a4-1602-451f-85c3-1e542b3155fc · outbound

This paper cites On the global convergence rates of softmax policy gradient methods.

The Actor-Critic Update Order Matters for PPO in Federated Reinforcement Learning On the global convergence rates of softmax policy gradient methods

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:54:28.968600Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-07T11:54:24.703231Z digest=sha256:e21f539ca8c4170cbc7421ea82ae66cea974f49394f0f75df251250ccc7dcf93

Observation b5d387ed-4cd7-43e2-8d98-86763ca50cc6 · outbound

This paper cites A., Veness, J., Bellemare, M.

The Actor-Critic Update Order Matters for PPO in Federated Reinforcement Learning A., Veness, J., Bellemare, M

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:54:28.764589Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-07T11:54:24.807776Z digest=sha256:4def49a9afbf56ea6f2821a5b13b105dfb255d177b4410e109af66cb1d2ddce7

Observation 29a9c553-d78e-438d-bf62-611abe8ae074 · outbound

This paper cites and Szepesvári, C.

The Actor-Critic Update Order Matters for PPO in Federated Reinforcement Learning and Szepesvári, C

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:54:28.584788Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-07T11:54:24.918651Z digest=sha256:aae6be19c22dc73135ddf1a547d39e5ce4e5d8c7cb0a84ab57e92ac1fd9cfb72

Observation 9718f0b3-1123-4839-9bce-329c0def2396 · outbound

This paper cites Planet dump retrieved from https://planet.osm.org.

The Actor-Critic Update Order Matters for PPO in Federated Reinforcement Learning Planet dump retrieved from https://planet.osm.org

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:54:28.421351Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-07T11:54:25.028428Z digest=sha256:55112ff93dc36d009577466bf77c64af3851d16d1eac01f3ee6941b97193ca22

Observation 12f480a1-5c6c-4313-b891-fa47167bf3b7 · outbound

This paper cites Trust Region Policy Optimization.

The Actor-Critic Update Order Matters for PPO in Federated Reinforcement Learning Trust Region Policy Optimization

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-07T11:54:25.098661Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:54:25.098661Z digest=sha256:744ec1d352a9145ef8b7a77dfe700f19d2ae73d55ab4d6243598a546e7c2b909

Observation 9292f581-f4b7-4a79-858f-3cf4e5080240 · outbound

This paper cites Proximal Policy Optimization Algorithms.

The Actor-Critic Update Order Matters for PPO in Federated Reinforcement Learning Proximal Policy Optimization Algorithms

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-07T11:54:25.165690Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:54:25.165690Z digest=sha256:0aeb55b4bc58e54871ef986ad15894c3e944a8cb67bda3064761c7857abdfad2

Observation 12536ad8-a079-4510-876b-58edb0fc1fb0 · outbound

This paper cites an unresolved cited work.

The Actor-Critic Update Order Matters for PPO in Federated Reinforcement Learning Unresolved cited work

Reference 24

Resolution
unresolved
raw_fallback, observed 2026-08-07T11:54:28.208016Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-07T11:54:25.271960Z digest=sha256:514e7045197995fbde61d2591826c92be74e26bbc76fc82536f41ac41faf584b

Observation 54b8813e-da05-4c67-8138-8a79fcc5c049 · outbound

This paper cites S., McAllester, D., Singh, S., and Mansour, Y.

The Actor-Critic Update Order Matters for PPO in Federated Reinforcement Learning S., McAllester, D., Singh, S., and Mansour, Y

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:54:27.978909Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-07T11:54:25.369311Z digest=sha256:665276b6c2495ad66ebc4926a83c701c8e4395a3e90ac25fdd034cbed06b5586

Observation fd1f72d0-98e4-4f32-9d04-6696078872f9 · outbound

This paper cites Boosted fitted q-iteration.

The Actor-Critic Update Order Matters for PPO in Federated Reinforcement Learning Boosted fitted q-iteration

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:54:27.778136Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-07T11:54:25.449432Z digest=sha256:17a7a4b8270cefa8c317ed694d429eea404b8cbcda49429a151e91c9f8040066

Observation 5eb2fc9f-e882-4ecd-927b-9ba6abec0656 · outbound

This paper cites Deep reinforcement learning with double q-learning.

The Actor-Critic Update Order Matters for PPO in Federated Reinforcement Learning Deep reinforcement learning with double q-learning

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-07T11:54:25.514852Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:54:25.514852Z digest=sha256:88790f44a2de3f91d030ade8574977d29fbb3edd8863b127e2f88c70916cf7db

Observation 014e2b7f-8963-441d-95b9-07541ce6162d · outbound

This paper cites E., Srivastava, S., Tuia, D., and Falcão, A.

The Actor-Critic Update Order Matters for PPO in Federated Reinforcement Learning E., Srivastava, S., Tuia, D., and Falcão, A

Reference 28

Resolution
metadata mismatch
raw_fallback, observed 2026-08-07T11:54:26.944638Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-07T11:54:25.614183Z digest=sha256:e5ca2dee16fd83a21197b198cf6ad94c01778a95be006e90aba83e1b05d2fccc

Observation 27b574ad-db67-4b56-b86b-4083bd2da289 · outbound

This paper cites L., Kheterpal, N., Jang, K., Wu, C., Wu, F., Liaw, R., Liang, E., and Bayen, A.

The Actor-Critic Update Order Matters for PPO in Federated Reinforcement Learning L., Kheterpal, N., Jang, K., Wu, C., Wu, F., Liaw, R., Liang, E., and Bayen, A

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:54:27.583825Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-07T11:54:25.722659Z digest=sha256:5ae6973dcad0e28b605a8145b29c5b5d84b92d0d8eb384beba48804f20994a3b

Observation 935a9aa8-71f1-4b99-8138-d8b4689fdcc5 · outbound

This paper cites Neural Policy Gradient Methods: Global Optimality and Rates of Convergence.

The Actor-Critic Update Order Matters for PPO in Federated Reinforcement Learning Neural Policy Gradient Methods: Global Optimality and Rates of Convergence

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-07T11:54:25.833975Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:54:25.833975Z digest=sha256:3381b1fe7b27420a7cb2daabb2a83c35790903b492aea5d8c9e0f46d6b362387

Observation 98bad912-d5dd-449b-b01d-99871ff1812b · outbound

This paper cites P., and Kakade, S.

The Actor-Critic Update Order Matters for PPO in Federated Reinforcement Learning P., and Kakade, S

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:54:27.391913Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-07T11:54:25.909375Z digest=sha256:092bb7c445f540010808b325a2f0d20b3db5536206574f1372639dec71c13cd7

Observation 39660861-fed3-4826-b7ea-26838d4e7be3 · outbound

This paper cites and Song, S.

The Actor-Critic Update Order Matters for PPO in Federated Reinforcement Learning and Song, S

Reference 32

Resolution
metadata mismatch
raw_fallback, observed 2026-08-07T11:54:26.697142Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-07T11:54:26.039203Z digest=sha256:d39d5a68abffdb3686809a4be1e96e31821e1ec25a3a27d07a0f391878cf4a4f

Observation 8b3badc4-4a1d-4f43-83be-6a93508fab2d · outbound

This paper cites A General Approach to Adding Differential Privacy to Iterative Training Procedures.

The Actor-Critic Update Order Matters for PPO in Federated Reinforcement Learning A General Approach to Adding Differential Privacy to Iterative Training Procedures

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-07T11:54:26.115849Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:54:26.115849Z digest=sha256:9d6b4089e5a74d8f32f163e95272744ddc5d12f7f58fcbe4ba1f12c8d1ce5822

Observation 852b693b-da84-4ba9-8f6e-ee8146aed99f · outbound

This paper cites Federated natural policy gradient and actor critic methods for multi-task reinforcement learning.

The Actor-Critic Update Order Matters for PPO in Federated Reinforcement Learning Federated natural policy gradient and actor critic methods for multi-task reinforcement learning

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:54:27.168863Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-07T11:54:26.220188Z digest=sha256:0607f82c19f8e2e561fdb5147649d5fac37f2820b26dfbdd8159f4bf8b91e6e5

Observation 722dd00c-e413-424a-b058-98232ef9a157 · outbound

This paper cites Federated Learning with Non-IID Data.

The Actor-Critic Update Order Matters for PPO in Federated Reinforcement Learning Federated Learning with Non-IID Data

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-07T11:54:26.316448Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:54:26.316448Z digest=sha256:f7eb3252edf32b241025d5fb189cfdc80adfb6b02bc2b232cedbafae6876e5c9

Pith citing papers

Observation b32d63e9-5e95-4638-a392-534a5ecc3069 · inbound

Collaborative Yet Personalized Policy Training: Single-Timescale Federated Actor-Critic cites this paper.

Collaborative Yet Personalized Policy Training: Single-Timescale Federated Actor-Critic The Actor-Critic Update Order Matters for PPO in Federated Reinforcement Learning

Reference 48

Resolution
verified exact
arxiv_id, observed 2026-05-15T02:08:29.340835Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-05-15T02:07:23.862809Z digest=sha256:58a688a61a2bf09bf36c090cbfe38ac78604466cc0494bce0678ec60b5d98613

Observation 1377c879-6903-497f-b020-dba0efa9ead2 · inbound

FGRPO: Federated GRPO with Adaptive Aggregation on Non-IID Data cites this paper.

FGRPO: Federated GRPO with Adaptive Aggregation on Non-IID Data The Actor-Critic Update Order Matters for PPO in Federated Reinforcement Learning

Reference 32

Resolution
metadata mismatch
arxiv_id, observed 2026-07-02T01:36:25.637603Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-06-28T11:44:53.211503Z digest=sha256:0403a00e6f0e7a982ad12cd037dbd40a0a0c8c9b8931e400212ad32b35ca8f89