Pith. sign in

Paper Citation Record · LEDGER

Action Dependency Graphs for Globally Optimal Coordinated Reinforcement Learning

As of 21 August 2026, this Paper Citation Record lists 51 of 51 outbound references and 0 inbound Pith citation observations for arXiv:2506.00797.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2506.00797 v1

Coverage vector

measured 51 of 51 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-07T12:09:13.560082Z

measured 51 of 51 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-20T06:33:59.587034+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

51 of 51 outbound references displayed

  • verified exact4
  • verified fuzzy33
  • unresolved13
  • parse uncertain0
  • malformed identifier1
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 8b6bd020-82aa-42ab-88a0-e05f2d1d3abe · outbound

This paper cites Multi-agent reinforcement learning: A selective overview of theories and algorithms.

Action Dependency Graphs for Globally Optimal Coordinated Reinforcement Learning Multi-agent reinforcement learning: A selective overview of theories and algorithms

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:09:21.135416Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-07T12:09:08.649500Z digest=sha256:ad0e43c91ad704a8505671e03f8cac5020fd9ba1e33a55a4c23362f2026a4460

Observation 51024266-e33e-4c57-97d6-a3db8b1a2638 · outbound

This paper cites A review of cooperative multi-agent deep reinforce- ment learning.

Action Dependency Graphs for Globally Optimal Coordinated Reinforcement Learning A review of cooperative multi-agent deep reinforce- ment learning

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-07T12:09:08.748019Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:09:08.748019Z digest=sha256:392119916fabc6bb86fbd77b67fc2848a7239d678575a6b00bfd154a7cfd2d1d

Observation c7f3603e-4208-4639-9314-9223ee5d0c63 · outbound

This paper cites Revisiting some common practices in cooperative multi-agent reinforcement learning.

Action Dependency Graphs for Globally Optimal Coordinated Reinforcement Learning Revisiting some common practices in cooperative multi-agent reinforcement learning

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:09:20.954678Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-07T12:09:08.860682Z digest=sha256:aa7d609d19f54e0f21a509a07bbf0f059de9a4c197cbd9ea920e52c6e5842868

Observation 89109b44-7e27-47d0-9b75-268f5356a050 · outbound

This paper cites Towards global optimality in cooperative marl with sequential transformation.

Action Dependency Graphs for Globally Optimal Coordinated Reinforcement Learning Towards global optimality in cooperative marl with sequential transformation

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:09:20.771005Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-07T12:09:09.003537Z digest=sha256:21c2ae90a9d60af6a8727d5cf2820dc8d976e9e2a271a4e34d6fab1e70e1b62a

Observation e38ff12b-715f-4e7a-bca9-b8c61d99d06f · outbound

This paper cites an unresolved cited work.

Action Dependency Graphs for Globally Optimal Coordinated Reinforcement Learning Unresolved cited work

Reference 5

Resolution
unresolved
raw_fallback, observed 2026-08-07T12:09:20.579717Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-07T12:09:09.155414Z digest=sha256:9331afb9082c288b557e2dcc66cea22a2fe3ce1d25aa550d2e7224a7880b599c

Observation f87454b7-397d-457d-91dc-cdd452778d19 · outbound

This paper cites Multiagent reinforcement learning: Rollout and policy iteration.

Action Dependency Graphs for Globally Optimal Coordinated Reinforcement Learning Multiagent reinforcement learning: Rollout and policy iteration

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:09:20.397911Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-07T12:09:09.249278Z digest=sha256:66a0f0b59a7b001604321cea4d1772466f5cf835fbf5ce579a23554e9bbc9b57

Observation 2e812c2d-9698-4dc1-bf85-6c04095e31bb · outbound

This paper cites Context-aware bayesian network actor-critic methods for cooperative multi-agent reinforcement learning.

Action Dependency Graphs for Globally Optimal Coordinated Reinforcement Learning Context-aware bayesian network actor-critic methods for cooperative multi-agent reinforcement learning

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:09:20.209812Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-07T12:09:09.341192Z digest=sha256:033fe0aa2b02ee3c29d45657731b9e3cad56d8524a69770f1d3e376e4892a3ce

Observation 1c7acf73-a58e-4307-afc1-2ee4bc6c28aa · outbound

This paper cites Coordinated reinforcement learning.

Action Dependency Graphs for Globally Optimal Coordinated Reinforcement Learning Coordinated reinforcement learning

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:09:19.999451Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-07T12:09:09.429599Z digest=sha256:f83e26ecd4db05d8e4d0bd0b3bf4ddbd97ebcecf049a742149911fb04edfaf6b

Observation 5b1cc5d4-84a6-4c0c-9624-aff77fee7ada · outbound

This paper cites More Centralized Training, Still Decentralized Execution: Multi-Agent Conditional Policy Factorization.

Action Dependency Graphs for Globally Optimal Coordinated Reinforcement Learning More Centralized Training, Still Decentralized Execution: Multi-Agent Conditional Policy Factorization

Reference 9

Resolution
verified exact
local_arxiv, observed 2026-08-07T12:09:14.299095Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-07T12:09:09.530585Z digest=sha256:f71f485e6359b306b472fb6f8a72a50733ecbe305f213d6f0d14cba86783e3d9

Observation 32ccfa27-6f00-40a1-a13a-276a6a831e68 · outbound

This paper cites Multi-agent reinforcement learning: Independent vs.

Action Dependency Graphs for Globally Optimal Coordinated Reinforcement Learning Multi-agent reinforcement learning: Independent vs

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:09:19.760972Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-07T12:09:09.679161Z digest=sha256:1703f64e8fb6182670a691953d0f801862c6128c788844f2256b1a6ddeedcd01

Observation 5fdb052c-3b61-4b2b-a13e-ed0f424ccbd0 · outbound

This paper cites Value- decomposition networks for cooperative multi-agent learning based on team reward.

Action Dependency Graphs for Globally Optimal Coordinated Reinforcement Learning Value- decomposition networks for cooperative multi-agent learning based on team reward

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-07T12:09:09.757254Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:09:09.757254Z digest=sha256:cdcb75d43b821da654180f26966fac6b6f4f747a7367a06949ce2ec754fc174a

Observation 617e1732-bcb8-4416-9fa7-b10ea31c513f · outbound

This paper cites Qmix: Monotonic value function factorisation for deep multi-agent reinforcement learning.

Action Dependency Graphs for Globally Optimal Coordinated Reinforcement Learning Qmix: Monotonic value function factorisation for deep multi-agent reinforcement learning

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-07T12:09:09.871952Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:09:09.871952Z digest=sha256:ecf8dd90d122be213df9e971d8eafe850d314cc426c66995d493f8bf495b0676

Observation 9f9d4401-e4c2-4462-ab8f-6f2f0b0c0310 · outbound

This paper cites Qtran: Learning to factorize with transformation for cooperative multi-agent reinforcement learning.

Action Dependency Graphs for Globally Optimal Coordinated Reinforcement Learning Qtran: Learning to factorize with transformation for cooperative multi-agent reinforcement learning

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:09:19.533967Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-07T12:09:10.006495Z digest=sha256:65b8ebe885996982c7c23de66d1cbc72874d8ec5971f37fe1a8d3fcc5c4237d4

Observation 2d436245-855b-47d1-a796-f8c1f5101110 · outbound

This paper cites Multi-agent actor-critic for mixed cooperative-competitive environments.

Action Dependency Graphs for Globally Optimal Coordinated Reinforcement Learning Multi-agent actor-critic for mixed cooperative-competitive environments

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:09:19.388235Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-07T12:09:10.158960Z digest=sha256:73eaf79754f6d7c9701a6e7e23e662cc6bfbe4452a252f7ed58affcd523e226b

Observation 2c22b13d-e5f9-4fc1-980a-0f6363672d32 · outbound

This paper cites Counterfactual multi-agent policy gradients.

Action Dependency Graphs for Globally Optimal Coordinated Reinforcement Learning Counterfactual multi-agent policy gradients

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-07T12:09:10.294956Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:09:10.294956Z digest=sha256:5bfa5f5ce8fc6a797e4f68bb1dcf851530c195b27393ec571c69075ab53b6097

Observation 7f009bdb-7411-4286-8e57-61636f1a08b4 · outbound

This paper cites Actor-attention-critic for multi-agent reinforcement learning.

Action Dependency Graphs for Globally Optimal Coordinated Reinforcement Learning Actor-attention-critic for multi-agent reinforcement learning

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:09:19.168581Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-07T12:09:10.443532Z digest=sha256:f22acfc6a17d798d20315a12688da2d4db4f275f6babaa2c59bd0aeb08882e5d

Observation da0dc977-e59a-4176-aeaa-e2452537fbca · outbound

This paper cites The surprising effectiveness of ppo in cooperative multi-agent games.

Action Dependency Graphs for Globally Optimal Coordinated Reinforcement Learning The surprising effectiveness of ppo in cooperative multi-agent games

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:09:18.942059Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-07T12:09:10.589233Z digest=sha256:8a02cddf457a0b86c6c7abfa4e5bd040a04080772626fa6d1b5e552aa29184d7

Observation 1b428e06-5328-44be-9ed5-0917852f530a · outbound

This paper cites Deep coordination graphs.

Action Dependency Graphs for Globally Optimal Coordinated Reinforcement Learning Deep coordination graphs

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:09:18.754897Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-07T12:09:10.709772Z digest=sha256:da51965889aecd621e0f8fa370a8d7c67cf1b3280ec81560cfc38a314b2ed32e

Observation 6349482b-11c0-43e9-9a51-0b679ed6eea2 · outbound

This paper cites Deep implicit coordination graphs for multi-agent reinforcement learning.

Action Dependency Graphs for Globally Optimal Coordinated Reinforcement Learning Deep implicit coordination graphs for multi-agent reinforcement learning

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:09:18.551846Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-07T12:09:10.773040Z digest=sha256:b363232e0d73b7a111c505078f7d68bfecd839831ff2a9107bbe63ab70bf3554

Observation 2eecfbe2-19f6-4f33-94bd-62d8861d7bf3 · outbound

This paper cites Context-aware sparse deep coordination graphs.

Action Dependency Graphs for Globally Optimal Coordinated Reinforcement Learning Context-aware sparse deep coordination graphs

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-07T12:09:10.869588Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:09:10.869588Z digest=sha256:00ddd45931f2aa7c136d27622206a928a24798be3c995822398a2c8fd3d4d5e9

Observation aa0c4210-4186-4ffe-a81b-fc3e9971270e · outbound

This paper cites Analysing factoriza- tions of action-value networks for cooperative multi-agent reinforcement learning.

Action Dependency Graphs for Globally Optimal Coordinated Reinforcement Learning Analysing factoriza- tions of action-value networks for cooperative multi-agent reinforcement learning

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:09:18.377599Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-07T12:09:10.972434Z digest=sha256:253b0efc77300c40f79212e049302a1069aaf64b382207fb70a02c08e4c78d7a

Observation b2e9d5bf-c29b-4406-83a2-9d2fb17d526e · outbound

This paper cites Biasing coevolutionary search for optimal multiagent behaviors.

Action Dependency Graphs for Globally Optimal Coordinated Reinforcement Learning Biasing coevolutionary search for optimal multiagent behaviors

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:09:18.216281Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-07T12:09:11.048003Z digest=sha256:fb2a0868f18660838edc8d49c8dee8f4dccf3818bedc8af63e12a9d283fa5526

Observation f88da3a5-629c-41da-935e-3766dd477d76 · outbound

This paper cites Bounded approximate decentralised coordination via the max-sum algorithm.

Action Dependency Graphs for Globally Optimal Coordinated Reinforcement Learning Bounded approximate decentralised coordination via the max-sum algorithm

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:09:18.090154Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-07T12:09:11.130000Z digest=sha256:cbe94c5046d7173d7b58892305f2432be06a51224a99d76457f6064bfa330e09

Observation 908f2428-3f57-4636-bf13-0dbf972ef113 · outbound

This paper cites Nonserial dynamic programming.

Action Dependency Graphs for Globally Optimal Coordinated Reinforcement Learning Nonserial dynamic programming

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:09:17.865044Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-07T12:09:11.226850Z digest=sha256:8774e6bb445113bcceeba0111f4633c681598d2fe328874f82565e7418f738fe

Observation 184c00e3-a0d2-49d9-acc3-b51f6c95b328 · outbound

This paper cites Gcs: Graph-based coordination strategy for multi-agent rein- forcement learning.

Action Dependency Graphs for Globally Optimal Coordinated Reinforcement Learning Gcs: Graph-based coordination strategy for multi-agent rein- forcement learning

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:09:17.665808Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-07T12:09:11.302675Z digest=sha256:c4a65f6f9868776ad715d40fde8764c212ebf9a6a26660702643df7ddb80931b

Observation ad4dc34a-6841-4ad4-a88e-b29c27c1e97a · outbound

This paper cites Ace: Cooperative multi-agent q-learning with bidirectional action-dependency.

Action Dependency Graphs for Globally Optimal Coordinated Reinforcement Learning Ace: Cooperative multi-agent q-learning with bidirectional action-dependency

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:09:17.505016Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-07T12:09:11.375942Z digest=sha256:e4d07879910ddb5679fc581f5bd23761f1c913f570c98d076d8ece333eb83dce

Observation ad3ff415-1ea2-4460-ad54-4dd744c33bcf · outbound

This paper cites Backpropagation through agents.

Action Dependency Graphs for Globally Optimal Coordinated Reinforcement Learning Backpropagation through agents

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:09:17.302075Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-07T12:09:11.463197Z digest=sha256:c1dc94de10f01d009f239d4cb7fae8c52e3aca4613d6029460f7a37668f8b7bf

Observation 0ad99584-ad15-4aad-8479-709a1fd67805 · outbound

This paper cites Group-aware coordination graph for multi-agent reinforce- ment learning.

Action Dependency Graphs for Globally Optimal Coordinated Reinforcement Learning Group-aware coordination graph for multi-agent reinforce- ment learning

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:09:17.124913Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-07T12:09:11.548731Z digest=sha256:3c95389b08e844bfc99c5bdfb0e4f965c6304faa0e8309e5f85e81724c8ef6ab

Observation d7e21359-9f50-4aa0-a801-56994194dd44 · outbound

This paper cites Is Centralized Training with Decentralized Execution Framework Centralized Enough for MARL?.

Action Dependency Graphs for Globally Optimal Coordinated Reinforcement Learning Is Centralized Training with Decentralized Execution Framework Centralized Enough for MARL?

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-07T12:09:11.648515Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:09:11.648515Z digest=sha256:6aca8f64b2131218a53347d50aa9b2778a3d0f36dd4424df5763ec9854493688

Observation e55f56e1-7572-4113-8bbd-39aa00d71bb6 · outbound

This paper cites Multi-agent reinforcement learning is a sequence modeling problem.

Action Dependency Graphs for Globally Optimal Coordinated Reinforcement Learning Multi-agent reinforcement learning is a sequence modeling problem

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:09:16.911860Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-07T12:09:11.719251Z digest=sha256:529d0671db7a00ab8f35e98ce390abaebc121599933116d0ef93dc161311b12b

Observation 0e09f4b3-627c-4f57-b282-b16fd3ef2fd5 · outbound

This paper cites Coordinated multi-agent reinforcement learning in net- worked distributed POMDPs.

Action Dependency Graphs for Globally Optimal Coordinated Reinforcement Learning Coordinated multi-agent reinforcement learning in net- worked distributed POMDPs

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:09:16.673642Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-07T12:09:11.819710Z digest=sha256:c7ab8d25a4a3a3f6a2efed4050264a9bf5251070e22e7b1522f4314f8b1b3da5

Observation 809ce0f2-1670-46a6-ab0e-8f3084295b37 · outbound

This paper cites Learning to coordinate with coordination graphs in repeated single-stage multi-agent decision problems.

Action Dependency Graphs for Globally Optimal Coordinated Reinforcement Learning Learning to coordinate with coordination graphs in repeated single-stage multi-agent decision problems

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:09:16.457907Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-07T12:09:11.885075Z digest=sha256:54751a322cd4fb86a21436e3a43fa9d46e7dc031945b680cb039849d87815a13

Observation 02f366e8-e52a-4426-9378-aa4310f7ef90 · outbound

This paper cites Coordinated Reinforcement Learning for Optimizing Mobile Networks.

Action Dependency Graphs for Globally Optimal Coordinated Reinforcement Learning Coordinated Reinforcement Learning for Optimizing Mobile Networks

Reference 33

Resolution
verified exact
local_arxiv, observed 2026-08-07T12:09:14.106967Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-07T12:09:11.957490Z digest=sha256:895fe3b0db4a5dfe70edb197fb6d0404918edc042bc8d483b523521ff09fb479

Observation 41654430-94ba-4132-96aa-30a723418a60 · outbound

This paper cites Understanding Value Decomposition Algorithms in Deep Cooperative Multi-Agent Reinforcement Learning.

Action Dependency Graphs for Globally Optimal Coordinated Reinforcement Learning Understanding Value Decomposition Algorithms in Deep Cooperative Multi-Agent Reinforcement Learning

Reference 34

Resolution
verified exact
local_arxiv, observed 2026-08-07T12:09:13.963852Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-07T12:09:12.080207Z digest=sha256:779e4300c3407bba5b9b14cf86639d8fb8dca53347f897e7b6cc957856a74015

Observation 1f1f0e6a-b47a-4210-8dba-072c3981eac2 · outbound

This paper cites On the global convergence rates of decentralized softmax gradient play in markov potential games.

Action Dependency Graphs for Globally Optimal Coordinated Reinforcement Learning On the global convergence rates of decentralized softmax gradient play in markov potential games

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:09:16.257303Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-07T12:09:12.157464Z digest=sha256:ec8a15b236824897cd39d674aa16cdc900c37089692f34714417ead138838478

Observation 1da1cc19-7669-4d0a-8ca0-1869eb7b544d · outbound

This paper cites Trust region policy optimisation in multi-agent reinforcement learning.

Action Dependency Graphs for Globally Optimal Coordinated Reinforcement Learning Trust region policy optimisation in multi-agent reinforcement learning

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:09:16.081145Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-07T12:09:12.226182Z digest=sha256:d1a469630a1acc820d6f9d7cbea21fb67bd539852df80408250b5b3396518d4b

Observation 69ee98c4-8091-4c8c-aced-d371d6ed6dfa · outbound

This paper cites On minmax theorems for multiplayer games.

Action Dependency Graphs for Globally Optimal Coordinated Reinforcement Learning On minmax theorems for multiplayer games

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:09:15.898241Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-07T12:09:12.275355Z digest=sha256:c86712d31f226a218d1e537d52aea75d12e443411383036cbcc735189190e009

Observation c780cd43-1d7f-4464-98d5-af134bf54847 · outbound

This paper cites Team theory and person-by-person optimization with binary decisions.

Action Dependency Graphs for Globally Optimal Coordinated Reinforcement Learning Team theory and person-by-person optimization with binary decisions

Reference 38

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:09:15.732224Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-07T12:09:12.360426Z digest=sha256:39a348c9ccbb7e6a4a2eccfdf7f718a5391022f922be2bad029583e355a70f81

Observation 850a2fe5-a037-4794-a047-ec8be72c6304 · outbound

This paper cites Collaborative multiagent reinforcement learning by payoff propagation.

Action Dependency Graphs for Globally Optimal Coordinated Reinforcement Learning Collaborative multiagent reinforcement learning by payoff propagation

Reference 39

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:09:15.527030Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-07T12:09:12.491947Z digest=sha256:3b362ad7b71e85d05474a930d6a85ddbb26d6e820d76f315edf6a1abd6a1b112

Observation 1ed0f93c-a933-458c-b352-e2eefac6ca96 · outbound

This paper cites Reinforcement learning: An introduction.

Action Dependency Graphs for Globally Optimal Coordinated Reinforcement Learning Reinforcement learning: An introduction

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-07T12:09:12.638776Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:09:12.638776Z digest=sha256:2d4732d16501502f8afc1cd8af7b328300971f40ffe29a89b9a29be587e27a7c

Observation 3c90702e-bb78-468b-ae25-b015cb39909e · outbound

This paper cites Microscopic traffic simulation using sumo.

Action Dependency Graphs for Globally Optimal Coordinated Reinforcement Learning Microscopic traffic simulation using sumo

Reference 41

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:09:15.312623Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-07T12:09:12.701283Z digest=sha256:acbbc5d9aa57f26e80b0d44e50d4706b6b6ff989cc6df398d7303f8935dc2235

Observation 84d1ec4b-c5e0-423b-8129-c892d0aa8b05 · outbound

This paper cites an unresolved cited work.

Action Dependency Graphs for Globally Optimal Coordinated Reinforcement Learning Unresolved cited work

Reference 42

Resolution
unresolved
raw_fallback, observed 2026-08-07T12:09:15.146793Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-07T12:09:12.775891Z digest=sha256:c9e638df8be4ed938dd06e86bb2565fcb1f04e3e15311b345be67034900a1be0

Observation e963249d-7e11-47ae-b019-9e3359cf0bd2 · outbound

This paper cites The StarCraft Multi-Agent Challenge.

Action Dependency Graphs for Globally Optimal Coordinated Reinforcement Learning The StarCraft Multi-Agent Challenge

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-07T12:09:12.878586Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:09:12.878586Z digest=sha256:40d99a4826bee9548b3248ed244148cb405488673ab8c473181e6c10baba74c9

Observation ddcf55f2-8984-4bd4-9119-914bb90b057b · outbound

This paper cites Soft actor-critic: Off- policy maximum entropy deep reinforcement learning with a stochastic actor.

Action Dependency Graphs for Globally Optimal Coordinated Reinforcement Learning Soft actor-critic: Off- policy maximum entropy deep reinforcement learning with a stochastic actor

Reference 44

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:09:14.987869Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-07T12:09:12.993195Z digest=sha256:1c308d3e78f8b27f9ac90fcdb5f7a2973d7f295cf2e077e327eb02c7e700f205

Observation 770ba9b0-04db-4ea9-960a-d0f43d179fc6 · outbound

This paper cites Multi-agent Reinforcement Learning for Networked System Control.

Action Dependency Graphs for Globally Optimal Coordinated Reinforcement Learning Multi-agent Reinforcement Learning for Networked System Control

Reference 45

Resolution
verified exact
local_arxiv, observed 2026-08-07T12:09:13.790402Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-07T12:09:13.071784Z digest=sha256:c471ac32d2eb6e688c01b853beef4d5744f53f896af0f783b6d7b04063e8e130

Observation 792e14a1-b117-4f07-97aa-11490eed0c67 · outbound

This paper cites Proximal Policy Optimization Algorithms.

Action Dependency Graphs for Globally Optimal Coordinated Reinforcement Learning Proximal Policy Optimization Algorithms

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-07T12:09:13.162796Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:09:13.162796Z digest=sha256:b37ef3743975d8d149305b9f5267669088bb8b971c3a45348fa5fdab2735d23e

Observation 31c27e78-34fa-4be7-ad14-d95b43c19ca2 · outbound

This paper cites Abstract dynamic programming.

Action Dependency Graphs for Globally Optimal Coordinated Reinforcement Learning Abstract dynamic programming

Reference 47

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:09:14.827639Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-07T12:09:13.215863Z digest=sha256:e5968dff89cfa0bece6dfe0a82e0f3009086858d9a2cdd2c19bbddee503bb600

Observation 2f839a35-c33c-41fc-b476-2f4ed18c165e · outbound

This paper cites Dynamical model of traffic congestion and numerical simulation.

Action Dependency Graphs for Globally Optimal Coordinated Reinforcement Learning Dynamical model of traffic congestion and numerical simulation

Reference 48

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:09:14.696229Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-07T12:09:13.294728Z digest=sha256:7af00fe8376c90459ea71861a0fc0d1db87fb254a7006e63df4626f77f5642ae

Observation 4bc9af4d-727e-4618-85e9-d07eb361547d · outbound

This paper cites High-Dimensional Continuous Control Using Generalized Advantage Estimation.

Action Dependency Graphs for Globally Optimal Coordinated Reinforcement Learning High-Dimensional Continuous Control Using Generalized Advantage Estimation

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-07T12:09:13.360881Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:09:13.360881Z digest=sha256:3d431b4ecb424bcb233beb91116eb421a4297b5880f3e07fa0e6021f6e76dcb1

Observation 6bc46274-8ff6-4695-8499-a18a4dfa8c4f · outbound

This paper cites Heterogeneous-Agent Mirror Learning: A Continuum of Solutions to Cooperative MARL.

Action Dependency Graphs for Globally Optimal Coordinated Reinforcement Learning Heterogeneous-Agent Mirror Learning: A Continuum of Solutions to Cooperative MARL

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-07T12:09:13.464761Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:09:13.464761Z digest=sha256:d9cd617da79d32f814ab16126b3d6b9491ce287b6d9977491b83982a056de3a1

Observation 6276f5eb-db5b-4c6f-8332-e34c28c3607f · outbound

This paper cites Albrecht.

Action Dependency Graphs for Globally Optimal Coordinated Reinforcement Learning Albrecht

Reference 51

Resolution
malformed identifier
raw_fallback, observed 2026-08-07T12:09:14.491389Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-07T12:09:13.560082Z digest=sha256:d3ca50143dc599d7eb0fb4ad0ce5aa67196fe105459906c54b8d571da9f36dda

Pith citing papers

No inbound Pith citation observations are available.