Pith. sign in

Paper Citation Record · LEDGER

Action Dependency Graphs for Globally Optimal Coordinated Reinforcement Learning

As of 9 August 2026, this Paper Citation Record lists 51 of 51 outbound references and 0 inbound Pith citation observations for arXiv:2506.00797.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2506.00797 v1

Coverage vector

measured 51 of 51 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-07T12:09:13.560082Z

measured 51 of 51 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-09T06:31:02.800959+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

51 of 51 outbound references displayed

  • verified exact4
  • verified fuzzy33
  • unresolved13
  • parse uncertain0
  • malformed identifier1
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 8b6bd020-82aa-42ab-88a0-e05f2d1d3abe · outbound

This paper cites Multi-agent reinforcement learning: A selective overview of theories and algorithms.

Action Dependency Graphs for Globally Optimal Coordinated Reinforcement Learning Multi-agent reinforcement learning: A selective overview of theories and algorithms

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:09:21.135416Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T12:09:08.649500Z digest=sha256:54bb04281ae4f30e758c7d1f15e4b6f91fd37996ad9ce8a2805fed2fe6080a3a

Observation 51024266-e33e-4c57-97d6-a3db8b1a2638 · outbound

This paper cites A review of cooperative multi-agent deep reinforce- ment learning.

Action Dependency Graphs for Globally Optimal Coordinated Reinforcement Learning A review of cooperative multi-agent deep reinforce- ment learning

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-07T12:09:08.748019Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:09:08.748019Z digest=sha256:5969d9af7ccf8c370a746afc170a7631ff9dda2f33e9f97b0d30459dc6f48142

Observation c7f3603e-4208-4639-9314-9223ee5d0c63 · outbound

This paper cites Revisiting some common practices in cooperative multi-agent reinforcement learning.

Action Dependency Graphs for Globally Optimal Coordinated Reinforcement Learning Revisiting some common practices in cooperative multi-agent reinforcement learning

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:09:20.954678Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T12:09:08.860682Z digest=sha256:097aec8bfc560bc3dc7d726856631366558f77fd5aa88669cf3e1368f1976a9d

Observation 89109b44-7e27-47d0-9b75-268f5356a050 · outbound

This paper cites Towards global optimality in cooperative marl with sequential transformation.

Action Dependency Graphs for Globally Optimal Coordinated Reinforcement Learning Towards global optimality in cooperative marl with sequential transformation

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:09:20.771005Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T12:09:09.003537Z digest=sha256:36ae16fbc33b357908cee92b0df0b8f4aec64d39ff0832bc9fe667832b4e618c

Observation e38ff12b-715f-4e7a-bca9-b8c61d99d06f · outbound

This paper cites an unresolved cited work.

Action Dependency Graphs for Globally Optimal Coordinated Reinforcement Learning Unresolved cited work

Reference 5

Resolution
unresolved
raw_fallback, observed 2026-08-07T12:09:20.579717Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T12:09:09.155414Z digest=sha256:d51804f924386727a1a049967e81925dcda1aa76bed873f29e0305ef85df2ede

Observation f87454b7-397d-457d-91dc-cdd452778d19 · outbound

This paper cites Multiagent reinforcement learning: Rollout and policy iteration.

Action Dependency Graphs for Globally Optimal Coordinated Reinforcement Learning Multiagent reinforcement learning: Rollout and policy iteration

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:09:20.397911Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T12:09:09.249278Z digest=sha256:843ed90084b50a2f0a6c17b4f56906a96a79ad473a2c681c942ec53f1bb3421b

Observation 2e812c2d-9698-4dc1-bf85-6c04095e31bb · outbound

This paper cites Context-aware bayesian network actor-critic methods for cooperative multi-agent reinforcement learning.

Action Dependency Graphs for Globally Optimal Coordinated Reinforcement Learning Context-aware bayesian network actor-critic methods for cooperative multi-agent reinforcement learning

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:09:20.209812Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T12:09:09.341192Z digest=sha256:2244e3a3cb218141d6cdc0e6647eee546720680821e498009c5f5f6a46361672

Observation 1c7acf73-a58e-4307-afc1-2ee4bc6c28aa · outbound

This paper cites Coordinated reinforcement learning.

Action Dependency Graphs for Globally Optimal Coordinated Reinforcement Learning Coordinated reinforcement learning

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:09:19.999451Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T12:09:09.429599Z digest=sha256:aca3916c85ec000ac3bf5bab381d3ea1a2eddafa0b3fdbb73ff85dc8792e5138

Observation 5b1cc5d4-84a6-4c0c-9624-aff77fee7ada · outbound

This paper cites More Centralized Training, Still Decentralized Execution: Multi-Agent Conditional Policy Factorization.

Action Dependency Graphs for Globally Optimal Coordinated Reinforcement Learning More Centralized Training, Still Decentralized Execution: Multi-Agent Conditional Policy Factorization

Reference 9

Resolution
verified exact
local_arxiv, observed 2026-08-07T12:09:14.299095Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T12:09:09.530585Z digest=sha256:7cf0b6e1ceb168fb923494808b950da4c52b66345ad1bf9431d9eff7081fb24f

Observation 32ccfa27-6f00-40a1-a13a-276a6a831e68 · outbound

This paper cites Multi-agent reinforcement learning: Independent vs.

Action Dependency Graphs for Globally Optimal Coordinated Reinforcement Learning Multi-agent reinforcement learning: Independent vs

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:09:19.760972Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T12:09:09.679161Z digest=sha256:48d4753687629574db092081fdc5561519b6b29997514a5a037b6a671274edfd

Observation 5fdb052c-3b61-4b2b-a13e-ed0f424ccbd0 · outbound

This paper cites Value- decomposition networks for cooperative multi-agent learning based on team reward.

Action Dependency Graphs for Globally Optimal Coordinated Reinforcement Learning Value- decomposition networks for cooperative multi-agent learning based on team reward

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-07T12:09:09.757254Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:09:09.757254Z digest=sha256:89e8b418cafe1f22414b2005dbfb558427cce994f7579c6f29ab6cffc80259ab

Observation 617e1732-bcb8-4416-9fa7-b10ea31c513f · outbound

This paper cites Qmix: Monotonic value function factorisation for deep multi-agent reinforcement learning.

Action Dependency Graphs for Globally Optimal Coordinated Reinforcement Learning Qmix: Monotonic value function factorisation for deep multi-agent reinforcement learning

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-07T12:09:09.871952Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:09:09.871952Z digest=sha256:8fc72a78eb70378bcd60172e7e1bbeabef64ec45e5ddac699e51af384e88aec3

Observation 9f9d4401-e4c2-4462-ab8f-6f2f0b0c0310 · outbound

This paper cites Qtran: Learning to factorize with transformation for cooperative multi-agent reinforcement learning.

Action Dependency Graphs for Globally Optimal Coordinated Reinforcement Learning Qtran: Learning to factorize with transformation for cooperative multi-agent reinforcement learning

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:09:19.533967Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T12:09:10.006495Z digest=sha256:18de6880e5483c21e1e74adf21734a817c2ed3069c538479443aabe868eb94fd

Observation 2d436245-855b-47d1-a796-f8c1f5101110 · outbound

This paper cites Multi-agent actor-critic for mixed cooperative-competitive environments.

Action Dependency Graphs for Globally Optimal Coordinated Reinforcement Learning Multi-agent actor-critic for mixed cooperative-competitive environments

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:09:19.388235Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T12:09:10.158960Z digest=sha256:994368a94db1018e3a417131d4990369650b1c7c79e571502a0933d5efa6bc0f

Observation 2c22b13d-e5f9-4fc1-980a-0f6363672d32 · outbound

This paper cites Counterfactual multi-agent policy gradients.

Action Dependency Graphs for Globally Optimal Coordinated Reinforcement Learning Counterfactual multi-agent policy gradients

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-07T12:09:10.294956Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:09:10.294956Z digest=sha256:c3d137bd09f83c938a05bc9e89ac6cd6631e06f75c4c1d3c325003d287c116fd

Observation 7f009bdb-7411-4286-8e57-61636f1a08b4 · outbound

This paper cites Actor-attention-critic for multi-agent reinforcement learning.

Action Dependency Graphs for Globally Optimal Coordinated Reinforcement Learning Actor-attention-critic for multi-agent reinforcement learning

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:09:19.168581Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T12:09:10.443532Z digest=sha256:49bf1fb54443178c81f4196da48e7dab403d43b67c803efd2e7ee9cc04381b2a

Observation da0dc977-e59a-4176-aeaa-e2452537fbca · outbound

This paper cites The surprising effectiveness of ppo in cooperative multi-agent games.

Action Dependency Graphs for Globally Optimal Coordinated Reinforcement Learning The surprising effectiveness of ppo in cooperative multi-agent games

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:09:18.942059Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T12:09:10.589233Z digest=sha256:7253c126e9a02fdd5b269544a4b74a0980b78dd71e1422e5e800656cb6b11aea

Observation 1b428e06-5328-44be-9ed5-0917852f530a · outbound

This paper cites Deep coordination graphs.

Action Dependency Graphs for Globally Optimal Coordinated Reinforcement Learning Deep coordination graphs

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:09:18.754897Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T12:09:10.709772Z digest=sha256:723e9d5eebadf5d4478ed556f2d3a71fe5c853e41fba86c5d4662ff164eec616

Observation 6349482b-11c0-43e9-9a51-0b679ed6eea2 · outbound

This paper cites Deep implicit coordination graphs for multi-agent reinforcement learning.

Action Dependency Graphs for Globally Optimal Coordinated Reinforcement Learning Deep implicit coordination graphs for multi-agent reinforcement learning

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:09:18.551846Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T12:09:10.773040Z digest=sha256:f5398dee0f724e6561a2c4028f2c3028cc2908c26dcefd4509ba39a7a3c32c01

Observation 2eecfbe2-19f6-4f33-94bd-62d8861d7bf3 · outbound

This paper cites Context-aware sparse deep coordination graphs.

Action Dependency Graphs for Globally Optimal Coordinated Reinforcement Learning Context-aware sparse deep coordination graphs

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-07T12:09:10.869588Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:09:10.869588Z digest=sha256:bb209006e2e4a4a6c6d73680c9eb95a36deff5875bbaae053ca2339ce287ed8b

Observation aa0c4210-4186-4ffe-a81b-fc3e9971270e · outbound

This paper cites Analysing factoriza- tions of action-value networks for cooperative multi-agent reinforcement learning.

Action Dependency Graphs for Globally Optimal Coordinated Reinforcement Learning Analysing factoriza- tions of action-value networks for cooperative multi-agent reinforcement learning

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:09:18.377599Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T12:09:10.972434Z digest=sha256:25b3de0bce6f51c052089e21d874e13f7b74003eb7fccb05c9723a1429c2ff5e

Observation b2e9d5bf-c29b-4406-83a2-9d2fb17d526e · outbound

This paper cites Biasing coevolutionary search for optimal multiagent behaviors.

Action Dependency Graphs for Globally Optimal Coordinated Reinforcement Learning Biasing coevolutionary search for optimal multiagent behaviors

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:09:18.216281Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T12:09:11.048003Z digest=sha256:c17e030cd392e3bbae095fa8b4f6247b347e1735aaa561d47eef960bf7e372bc

Observation f88da3a5-629c-41da-935e-3766dd477d76 · outbound

This paper cites Bounded approximate decentralised coordination via the max-sum algorithm.

Action Dependency Graphs for Globally Optimal Coordinated Reinforcement Learning Bounded approximate decentralised coordination via the max-sum algorithm

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:09:18.090154Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T12:09:11.130000Z digest=sha256:782f997492906345298ce4f195427a364eba9ab70d20c45b89a9b0d2285e41b4

Observation 908f2428-3f57-4636-bf13-0dbf972ef113 · outbound

This paper cites Nonserial dynamic programming.

Action Dependency Graphs for Globally Optimal Coordinated Reinforcement Learning Nonserial dynamic programming

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:09:17.865044Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T12:09:11.226850Z digest=sha256:05abe7099b9a4ec7ac78554be33879b82f206a07c6e37d480792fb9fe40208fb

Observation 184c00e3-a0d2-49d9-acc3-b51f6c95b328 · outbound

This paper cites Gcs: Graph-based coordination strategy for multi-agent rein- forcement learning.

Action Dependency Graphs for Globally Optimal Coordinated Reinforcement Learning Gcs: Graph-based coordination strategy for multi-agent rein- forcement learning

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:09:17.665808Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T12:09:11.302675Z digest=sha256:25a58edef013e2e0d1c6e5448da1c66f83b8d84aace0bc7f5eccc82624316065

Observation ad4dc34a-6841-4ad4-a88e-b29c27c1e97a · outbound

This paper cites Ace: Cooperative multi-agent q-learning with bidirectional action-dependency.

Action Dependency Graphs for Globally Optimal Coordinated Reinforcement Learning Ace: Cooperative multi-agent q-learning with bidirectional action-dependency

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:09:17.505016Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T12:09:11.375942Z digest=sha256:70de4f7a6860856a552a54f41005ed9db89779ce8dfc4577ba28843a03408cb9

Observation ad3ff415-1ea2-4460-ad54-4dd744c33bcf · outbound

This paper cites Backpropagation through agents.

Action Dependency Graphs for Globally Optimal Coordinated Reinforcement Learning Backpropagation through agents

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:09:17.302075Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T12:09:11.463197Z digest=sha256:07acc17eee9b111cb09d29a9a14a8bed490b1127e71a2420fada1712eece5b48

Observation 0ad99584-ad15-4aad-8479-709a1fd67805 · outbound

This paper cites Group-aware coordination graph for multi-agent reinforce- ment learning.

Action Dependency Graphs for Globally Optimal Coordinated Reinforcement Learning Group-aware coordination graph for multi-agent reinforce- ment learning

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:09:17.124913Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T12:09:11.548731Z digest=sha256:325b2c6c127de985ab5029f2cc66de0638ccbf7e0316887f5f039c4b20eb145f

Observation d7e21359-9f50-4aa0-a801-56994194dd44 · outbound

This paper cites Is Centralized Training with Decentralized Execution Framework Centralized Enough for MARL?.

Action Dependency Graphs for Globally Optimal Coordinated Reinforcement Learning Is Centralized Training with Decentralized Execution Framework Centralized Enough for MARL?

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-07T12:09:11.648515Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:09:11.648515Z digest=sha256:22722af74d8d706a92f6fef081509af1485e012de5328cf52ea48b579e577968

Observation e55f56e1-7572-4113-8bbd-39aa00d71bb6 · outbound

This paper cites Multi-agent reinforcement learning is a sequence modeling problem.

Action Dependency Graphs for Globally Optimal Coordinated Reinforcement Learning Multi-agent reinforcement learning is a sequence modeling problem

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:09:16.911860Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T12:09:11.719251Z digest=sha256:1e99e6c8e860e59ee1a2e5c7bfa73e26751303b1bafecba4bdeddfc240f28124

Observation 0e09f4b3-627c-4f57-b282-b16fd3ef2fd5 · outbound

This paper cites Coordinated multi-agent reinforcement learning in net- worked distributed POMDPs.

Action Dependency Graphs for Globally Optimal Coordinated Reinforcement Learning Coordinated multi-agent reinforcement learning in net- worked distributed POMDPs

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:09:16.673642Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T12:09:11.819710Z digest=sha256:a79d8d381b3de86b530a1cce216a5ffc59d1127fae53dde30d8be660f5fe934c

Observation 809ce0f2-1670-46a6-ab0e-8f3084295b37 · outbound

This paper cites Learning to coordinate with coordination graphs in repeated single-stage multi-agent decision problems.

Action Dependency Graphs for Globally Optimal Coordinated Reinforcement Learning Learning to coordinate with coordination graphs in repeated single-stage multi-agent decision problems

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:09:16.457907Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T12:09:11.885075Z digest=sha256:6202d18440d9baed2933631cae1ffed17ef18fb17e487977c05b4b18679a9949

Observation 02f366e8-e52a-4426-9378-aa4310f7ef90 · outbound

This paper cites Coordinated Reinforcement Learning for Optimizing Mobile Networks.

Action Dependency Graphs for Globally Optimal Coordinated Reinforcement Learning Coordinated Reinforcement Learning for Optimizing Mobile Networks

Reference 33

Resolution
verified exact
local_arxiv, observed 2026-08-07T12:09:14.106967Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T12:09:11.957490Z digest=sha256:a6870a6e3114266937de8a4248db316975c6cae5bd58e47f42def79ae17e9ac2

Observation 41654430-94ba-4132-96aa-30a723418a60 · outbound

This paper cites Understanding Value Decomposition Algorithms in Deep Cooperative Multi-Agent Reinforcement Learning.

Action Dependency Graphs for Globally Optimal Coordinated Reinforcement Learning Understanding Value Decomposition Algorithms in Deep Cooperative Multi-Agent Reinforcement Learning

Reference 34

Resolution
verified exact
local_arxiv, observed 2026-08-07T12:09:13.963852Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T12:09:12.080207Z digest=sha256:466904701f88a590d6f07d074c6ea910fc1f892d66198f0728456f612e228504

Observation 1f1f0e6a-b47a-4210-8dba-072c3981eac2 · outbound

This paper cites On the global convergence rates of decentralized softmax gradient play in markov potential games.

Action Dependency Graphs for Globally Optimal Coordinated Reinforcement Learning On the global convergence rates of decentralized softmax gradient play in markov potential games

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:09:16.257303Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T12:09:12.157464Z digest=sha256:159cc8c800979a69f2e09498b261db52231a4f849ab511143e27029fff0eb9eb

Observation 1da1cc19-7669-4d0a-8ca0-1869eb7b544d · outbound

This paper cites Trust region policy optimisation in multi-agent reinforcement learning.

Action Dependency Graphs for Globally Optimal Coordinated Reinforcement Learning Trust region policy optimisation in multi-agent reinforcement learning

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:09:16.081145Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T12:09:12.226182Z digest=sha256:a04193e99825a011a37495e8048be4d2d45a5ad78acd1f459a5ff9e13273050a

Observation 69ee98c4-8091-4c8c-aced-d371d6ed6dfa · outbound

This paper cites On minmax theorems for multiplayer games.

Action Dependency Graphs for Globally Optimal Coordinated Reinforcement Learning On minmax theorems for multiplayer games

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:09:15.898241Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T12:09:12.275355Z digest=sha256:81f444109b73d5f85dab78fc9e0afba1e9f1a754aadb7b99bef5060ef9b9540e

Observation c780cd43-1d7f-4464-98d5-af134bf54847 · outbound

This paper cites Team theory and person-by-person optimization with binary decisions.

Action Dependency Graphs for Globally Optimal Coordinated Reinforcement Learning Team theory and person-by-person optimization with binary decisions

Reference 38

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:09:15.732224Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T12:09:12.360426Z digest=sha256:4c6d2ee192615ddef08fa44f2dc20bad4f215eda906f7e7bc8c9009280eb71a1

Observation 850a2fe5-a037-4794-a047-ec8be72c6304 · outbound

This paper cites Collaborative multiagent reinforcement learning by payoff propagation.

Action Dependency Graphs for Globally Optimal Coordinated Reinforcement Learning Collaborative multiagent reinforcement learning by payoff propagation

Reference 39

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:09:15.527030Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T12:09:12.491947Z digest=sha256:805239615feb0e1e77baa049d25162f596f19dbdbb97fb40e9e67271f90117b7

Observation 1ed0f93c-a933-458c-b352-e2eefac6ca96 · outbound

This paper cites Reinforcement learning: An introduction.

Action Dependency Graphs for Globally Optimal Coordinated Reinforcement Learning Reinforcement learning: An introduction

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-07T12:09:12.638776Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:09:12.638776Z digest=sha256:86730d85791a9ab139d768234bb05de19ed74d1b420253fa089cb47047a5f9ab

Observation 3c90702e-bb78-468b-ae25-b015cb39909e · outbound

This paper cites Microscopic traffic simulation using sumo.

Action Dependency Graphs for Globally Optimal Coordinated Reinforcement Learning Microscopic traffic simulation using sumo

Reference 41

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:09:15.312623Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T12:09:12.701283Z digest=sha256:a82955385bfcdbd49c8a09e6781c810a5aa66d444b93eb669aa8bb64455f2d4c

Observation 84d1ec4b-c5e0-423b-8129-c892d0aa8b05 · outbound

This paper cites an unresolved cited work.

Action Dependency Graphs for Globally Optimal Coordinated Reinforcement Learning Unresolved cited work

Reference 42

Resolution
unresolved
raw_fallback, observed 2026-08-07T12:09:15.146793Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T12:09:12.775891Z digest=sha256:49904768190a4a79688247185144e4a241d6c335d72b376bb81784dcf20f8f30

Observation e963249d-7e11-47ae-b019-9e3359cf0bd2 · outbound

This paper cites The StarCraft Multi-Agent Challenge.

Action Dependency Graphs for Globally Optimal Coordinated Reinforcement Learning The StarCraft Multi-Agent Challenge

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-07T12:09:12.878586Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:09:12.878586Z digest=sha256:486bc694a4d57c46f1a504e262197ab60e5d9a87fbfda25cccc6a3cb5078b064

Observation ddcf55f2-8984-4bd4-9119-914bb90b057b · outbound

This paper cites Soft actor-critic: Off- policy maximum entropy deep reinforcement learning with a stochastic actor.

Action Dependency Graphs for Globally Optimal Coordinated Reinforcement Learning Soft actor-critic: Off- policy maximum entropy deep reinforcement learning with a stochastic actor

Reference 44

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:09:14.987869Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T12:09:12.993195Z digest=sha256:f9668304106ad3d595b4d964a288b393990fe76c3ee0b908d67c876e64cbc46d

Observation 770ba9b0-04db-4ea9-960a-d0f43d179fc6 · outbound

This paper cites Multi-agent Reinforcement Learning for Networked System Control.

Action Dependency Graphs for Globally Optimal Coordinated Reinforcement Learning Multi-agent Reinforcement Learning for Networked System Control

Reference 45

Resolution
verified exact
local_arxiv, observed 2026-08-07T12:09:13.790402Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T12:09:13.071784Z digest=sha256:377b79f4148a9c1d8a0825d9daa5071c1456d84c6dec83053bd6a89cf583c7df

Observation 792e14a1-b117-4f07-97aa-11490eed0c67 · outbound

This paper cites Proximal Policy Optimization Algorithms.

Action Dependency Graphs for Globally Optimal Coordinated Reinforcement Learning Proximal Policy Optimization Algorithms

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-07T12:09:13.162796Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:09:13.162796Z digest=sha256:85fa3188c1319c609ff8c6d0bcddc21f8a60aa1c73ed6270a77096609fd81175

Observation 31c27e78-34fa-4be7-ad14-d95b43c19ca2 · outbound

This paper cites Abstract dynamic programming.

Action Dependency Graphs for Globally Optimal Coordinated Reinforcement Learning Abstract dynamic programming

Reference 47

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:09:14.827639Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T12:09:13.215863Z digest=sha256:8b04012ae2e1441f7c6a9ec8f6a579d3893c9be070c8c5d105c0cbd9090e3797

Observation 2f839a35-c33c-41fc-b476-2f4ed18c165e · outbound

This paper cites Dynamical model of traffic congestion and numerical simulation.

Action Dependency Graphs for Globally Optimal Coordinated Reinforcement Learning Dynamical model of traffic congestion and numerical simulation

Reference 48

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:09:14.696229Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T12:09:13.294728Z digest=sha256:c92108f0035e1765bf3685d3f47942ffdde7258408971fa39a48f513faf4d999

Observation 4bc9af4d-727e-4618-85e9-d07eb361547d · outbound

This paper cites High-Dimensional Continuous Control Using Generalized Advantage Estimation.

Action Dependency Graphs for Globally Optimal Coordinated Reinforcement Learning High-Dimensional Continuous Control Using Generalized Advantage Estimation

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-07T12:09:13.360881Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:09:13.360881Z digest=sha256:24a01b3645bbab1b2326b2a2774f47ad1760294e890c4f9410cef373d25e5c7b

Observation 6bc46274-8ff6-4695-8499-a18a4dfa8c4f · outbound

This paper cites Heterogeneous-Agent Mirror Learning: A Continuum of Solutions to Cooperative MARL.

Action Dependency Graphs for Globally Optimal Coordinated Reinforcement Learning Heterogeneous-Agent Mirror Learning: A Continuum of Solutions to Cooperative MARL

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-07T12:09:13.464761Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:09:13.464761Z digest=sha256:99b72318b5008f5a42a0b44aa25e14e097bbc084dacd4ff29e4536a9f026d666

Observation 6276f5eb-db5b-4c6f-8332-e34c28c3607f · outbound

This paper cites Albrecht.

Action Dependency Graphs for Globally Optimal Coordinated Reinforcement Learning Albrecht

Reference 51

Resolution
malformed identifier
raw_fallback, observed 2026-08-07T12:09:14.491389Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T12:09:13.560082Z digest=sha256:094b90446e8504d3da64ed76e89c416bf444457098e177a95726d9d9c6e6564d

Pith citing papers

No inbound Pith citation observations are available.