Pith. sign in

Paper Citation Record · LEDGER

The challenge of hidden gifts in multi-agent reinforcement learning

As of 18 August 2026, this Paper Citation Record lists 32 of 32 outbound references and 0 inbound Pith citation observations for arXiv:2505.20579.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2505.20579 v7

Coverage vector

measured 32 of 32 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-07T13:57:40.245927Z

measured 32 of 32 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-18T06:34:40.430872+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

32 of 32 outbound references displayed

  • verified exact2
  • verified fuzzy16
  • unresolved14
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 0785d4bb-e775-4cfb-8e2c-d5f46fb58dfd · outbound

This paper cites LOQA : Learning with opponent q-learning awareness.

The challenge of hidden gifts in multi-agent reinforcement learning LOQA : Learning with opponent q-learning awareness

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:57:41.973000Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-08-07T13:57:38.263636Z digest=sha256:7ea110bb544f0cfc76c777af79818d0439a2183525c7c4d91fa51cbcc4a4d599

Observation 65b50dc3-88b8-429a-8651-2d1ae047486d · outbound

This paper cites Unifying temporal and structural credit assignment problems.

The challenge of hidden gifts in multi-agent reinforcement learning Unifying temporal and structural credit assignment problems

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:57:41.845941Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-08-07T13:57:38.403146Z digest=sha256:ce9679f881710a5998a6f53a66e6b65d9c7edf3bb236b7ea7a9eaab1016065d3

Observation 45148112-d0c6-4aa3-becf-9b96e2203caf · outbound

This paper cites Understanding the impact of entropy on policy optimization.

The challenge of hidden gifts in multi-agent reinforcement learning Understanding the impact of entropy on policy optimization

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-07T13:57:38.517182Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:57:38.517182Z digest=sha256:3dd54cecab2f9eb368febb4bf68f3140780737f1e8d9c1af0e87c8c09f6eaed6

Observation 717fd1f9-4ef1-4ca7-92c3-7639a377b6b5 · outbound

This paper cites Effective choice in the prisoner's dilemma.

The challenge of hidden gifts in multi-agent reinforcement learning Effective choice in the prisoner's dilemma

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:57:41.762392Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-08-07T13:57:38.596201Z digest=sha256:e9ce0da2917a22cbb73d7466ff2c9da816a447a43a87aabc1dfffe3ed32dde68

Observation 545c2036-bae3-4537-a921-a7b2ee20f4d5 · outbound

This paper cites Manitokanac.

The challenge of hidden gifts in multi-agent reinforcement learning Manitokanac

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:57:41.704296Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-08-07T13:57:38.723112Z digest=sha256:0f7e7cbd876dc2fc00da62b181c413fbcf93d7b936ebc8dd2541a1b587653642

Observation 878a3ef2-3b7d-4462-b2c7-3f02961b23a5 · outbound

This paper cites The theory of dynamic programming.

The challenge of hidden gifts in multi-agent reinforcement learning The theory of dynamic programming

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-07T13:57:38.847739Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:57:38.847739Z digest=sha256:bbeba802764778ce19bd933a1fb9f26b3eb4d45d65fe0cd25f4136777a1b57b2

Observation fa69cbd5-ecc0-40ef-98f3-07dacafc609b · outbound

This paper cites Prisoner's dilemma; a study in conflict and cooperation.

The challenge of hidden gifts in multi-agent reinforcement learning Prisoner's dilemma; a study in conflict and cooperation

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:57:41.652409Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-08-07T13:57:38.955323Z digest=sha256:aba2400a3a8ac97abd2e7cef472bc83c5d9355803447ac068565ac196591103e

Observation fecfc178-cc40-4f6c-ad19-b731c3caee36 · outbound

This paper cites Decision transformer: Reinforcement learning via sequence modeling.

The challenge of hidden gifts in multi-agent reinforcement learning Decision transformer: Reinforcement learning via sequence modeling

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-07T13:57:39.056790Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:57:39.056790Z digest=sha256:6f5c27512f53b758437ec1631ca72188e001f8ddbc041b803fc29443e8d00e64

Observation c395c7bc-213f-4ad1-ab39-50dc2bea790a · outbound

This paper cites Minigrid & miniworld: Modular & customizable reinforcement learning environments for goal-oriented tasks.

The challenge of hidden gifts in multi-agent reinforcement learning Minigrid & miniworld: Modular & customizable reinforcement learning environments for goal-oriented tasks

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-07T13:57:39.091824Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:57:39.091824Z digest=sha256:a2c18f57bf841fd79aea8188f36c0dd6071550058da69c6a42e42ff5876de3be

Observation 440b0c45-833b-4302-9b72-75deec581e6f · outbound

This paper cites Learning phrase representations using RNN encoder -- decoder for statistical machine translation.

The challenge of hidden gifts in multi-agent reinforcement learning Learning phrase representations using RNN encoder -- decoder for statistical machine translation

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-07T13:57:39.129254Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:57:39.129254Z digest=sha256:29ec884b97e20ce71e0d38ec8501e2c6cad601ddc09a330d99a85637a49bc389

Observation 1429d519-a76c-4b38-8964-b27f8e51bc77 · outbound

This paper cites Hypothetical minds: Scaffolding theory of mind for multi-agent tasks with large language models.

The challenge of hidden gifts in multi-agent reinforcement learning Hypothetical minds: Scaffolding theory of mind for multi-agent tasks with large language models

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:57:41.604374Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-08-07T13:57:39.183179Z digest=sha256:6eb5d2ab0eb6850d23b1c4875a219bf9bcd7ff49b317b9f299f952ba31f0e51b

Observation ee237fee-8184-496b-b614-bec4548cb90b · outbound

This paper cites Maximum entropy RL (provably) solves some robust RL problems.

The challenge of hidden gifts in multi-agent reinforcement learning Maximum entropy RL (provably) solves some robust RL problems

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:57:41.525145Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-08-07T13:57:39.240202Z digest=sha256:06373cd1fc64ecf0015b155ecf6456d50ada79762040042e1a42d81d599df35d

Observation edcda3d2-97cb-44e2-b238-7b58dad96676 · outbound

This paper cites Counterfactual multi-agent policy gradients.

The challenge of hidden gifts in multi-agent reinforcement learning Counterfactual multi-agent policy gradients

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-07T13:57:39.298552Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:57:39.298552Z digest=sha256:2b3ea85fe195a932afea5e3e3e7e82e36d1bca204242be3216f28cb292514137

Observation 1d218df6-f17b-4151-b312-7a51599cfd3e · outbound

This paper cites Learning with Opponent-Learning Awareness.

The challenge of hidden gifts in multi-agent reinforcement learning Learning with Opponent-Learning Awareness

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-07T13:57:39.358832Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:57:39.358832Z digest=sha256:f93f893eaf19ce7aaa8df171ac6119a9769df510f8c3861a8d61ab4b2b7ae8fa

Observation fa5fa247-7d8a-4ce1-90c4-4e7525eabf33 · outbound

This paper cites Structural credit assignment in neural networks using reinforcement learning.

The challenge of hidden gifts in multi-agent reinforcement learning Structural credit assignment in neural networks using reinforcement learning

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:57:41.422159Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-08-07T13:57:39.409064Z digest=sha256:dabae78d9b2fbc98c0ddb452fccf2d0391ab467fb2cebddaf064bfcedd8d4629

Observation 23447fde-0e28-4662-8aaf-02fb51e13281 · outbound

This paper cites Soft actor-critic: Off-policy maximum entropy deep reinforcement learning with a stochastic actor.

The challenge of hidden gifts in multi-agent reinforcement learning Soft actor-critic: Off-policy maximum entropy deep reinforcement learning with a stochastic actor

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-07T13:57:39.454659Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:57:39.454659Z digest=sha256:7c5489603b0b2a1cadcd597e809ec9da455231e215384e08d533487496928802

Observation c9e0a6a7-9510-41ae-9122-827df185a5cf · outbound

This paper cites Optimizing agent behavior over long time scales by transporting value.

The challenge of hidden gifts in multi-agent reinforcement learning Optimizing agent behavior over long time scales by transporting value

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:57:41.293111Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-08-07T13:57:39.511697Z digest=sha256:3655a260bb54a9485891d4d8fd03d1ad42d7a3aabbe7da9ebc3fcbffe74d9af5

Observation 0e6b197b-fa05-4ba0-96bc-c5aeff909f22 · outbound

This paper cites Stateful active facilitator: Coordination and environmental heterogeneity in cooperative multi-agent reinforcement learning.

The challenge of hidden gifts in multi-agent reinforcement learning Stateful active facilitator: Coordination and environmental heterogeneity in cooperative multi-agent reinforcement learning

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:57:41.188200Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-08-07T13:57:39.568734Z digest=sha256:2d1567e566e7b8ec9d85f0db05c5d2a28a5515f58d2414e15ca0cb05586bf783

Observation 38f3c47b-f42f-45cf-add0-5cde06837423 · outbound

This paper cites Maven: Multi-agent variational exploration.

The challenge of hidden gifts in multi-agent reinforcement learning Maven: Multi-agent variational exploration

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-07T13:57:39.609339Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:57:39.609339Z digest=sha256:b68f5960897d6dddf4151d576a1db40c01e1164ce4b7a1114de29f0117c5d900

Observation 5778cf27-9108-4fa6-9e27-766e4d1956f5 · outbound

This paper cites Multi-agent cooperation through learning-aware policy gradients.

The challenge of hidden gifts in multi-agent reinforcement learning Multi-agent cooperation through learning-aware policy gradients

Reference 20

Resolution
verified exact
local_arxiv, observed 2026-08-07T13:57:40.513690Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-08-07T13:57:39.665521Z digest=sha256:73ea0437975cfea9bf3f672836e960cd08ee8bc6b8dcf99c31b562319cf5d0f4

Observation 86740e3a-9d60-4325-aeb7-1832e2994eb9 · outbound

This paper cites Equilibrium points in n-person games.

The challenge of hidden gifts in multi-agent reinforcement learning Equilibrium points in n-person games

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-07T13:57:39.695724Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:57:39.695724Z digest=sha256:e2f41a4d3b70b23771d82361239eef3af0b2e9c5356478bb55a2f36bd8997925

Observation 5cf85eeb-2d4b-4a4e-b747-7bc456e0db6b · outbound

This paper cites When do transformers shine in rl? decoupling memory from credit assignment.

The challenge of hidden gifts in multi-agent reinforcement learning When do transformers shine in rl? decoupling memory from credit assignment

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:57:41.109096Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-08-07T13:57:39.752780Z digest=sha256:779738834d4c6f2c7a539ae09842e41400178f1086f42f6440a0c17a5f9251c5

Observation 55335269-7cf6-46a5-9871-7febce7d37f4 · outbound

This paper cites Monotonic value function factorisation for deep multi-agent reinforcement learning.

The challenge of hidden gifts in multi-agent reinforcement learning Monotonic value function factorisation for deep multi-agent reinforcement learning

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-07T13:57:39.809487Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:57:39.809487Z digest=sha256:97b663eb1ac11ebc4accab1f3fa086f3f7fa2a4a18d00285d9cfe0f47eb5a72e

Observation 451b9fd3-8fe7-4219-8eaa-1f1103f85942 · outbound

This paper cites Proximal Policy Optimization Algorithms.

The challenge of hidden gifts in multi-agent reinforcement learning Proximal Policy Optimization Algorithms

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-07T13:57:39.857334Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:57:39.857334Z digest=sha256:ace17be061fb5aba66154a0de087c07129574efe720c23b7f6cf292cc3412f33

Observation e3603498-d8ec-4585-ae22-691096eaa45e · outbound

This paper cites Agent-Time Attention for Sparse Rewards Multi-Agent Reinforcement Learning.

The challenge of hidden gifts in multi-agent reinforcement learning Agent-Time Attention for Sparse Rewards Multi-Agent Reinforcement Learning

Reference 25

Resolution
verified exact
local_arxiv, observed 2026-08-07T13:57:40.393765Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-08-07T13:57:39.909858Z digest=sha256:2773ac97f07e6adf9c7246bda667c57ce5493ed202d406b33729ba63f4d471a8

Observation a839fd47-f4ca-41e4-84a2-aa7124976a93 · outbound

This paper cites Qtran: Learning to factorize with transformation for cooperative multi-agent reinforcement learning.

The challenge of hidden gifts in multi-agent reinforcement learning Qtran: Learning to factorize with transformation for cooperative multi-agent reinforcement learning

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:57:41.008477Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-08-07T13:57:39.956054Z digest=sha256:c50c7abd3e0b53e169550964bf67bf84d00c14ca53088f09057e709568712739

Observation db7404eb-e5c0-4b34-a830-334337701ea0 · outbound

This paper cites PufferLib: Making Reinforcement Learning Libraries and Environments Play Nice.

The challenge of hidden gifts in multi-agent reinforcement learning PufferLib: Making Reinforcement Learning Libraries and Environments Play Nice

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-07T13:57:40.035699Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:57:40.035699Z digest=sha256:ef66b46a9456629ea364c7af4d9bc8a8a3569d2d5f732187b686cdc7382a1f81

Observation 91fd2b73-84ff-4050-9654-cc1a7d4f5e5a · outbound

This paper cites Value-decomposition networks for cooperative multi-agent learning.

The challenge of hidden gifts in multi-agent reinforcement learning Value-decomposition networks for cooperative multi-agent learning

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:57:40.936484Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-08-07T13:57:40.079910Z digest=sha256:f2f7c8565f863da56d245dd787d1a7bc14678847e88df7c5bf39e2856977044b

Observation 6588f00d-0d2f-4415-a0a9-d4590e78ca2a · outbound

This paper cites Policy gradient methods for reinforcement learning with function approximation.

The challenge of hidden gifts in multi-agent reinforcement learning Policy gradient methods for reinforcement learning with function approximation

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-07T13:57:40.131961Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:57:40.131961Z digest=sha256:d77e401e2f6efb3bffd6c78c45e1de5f71288683bd76e7264ce9642a5d9e1f8c

Observation 32731990-74ff-40fb-8c52-8d65ba72f489 · outbound

This paper cites Learning sequences of actions in collectives of autonomous agents.

The challenge of hidden gifts in multi-agent reinforcement learning Learning sequences of actions in collectives of autonomous agents

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:57:40.819723Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-08-07T13:57:40.173465Z digest=sha256:290138f28eb4a0a9531d4995fab80914a963a431c5a1327f94ee454b50962407

Observation 4873364c-44db-423f-8ed7-6303da79cba4 · outbound

This paper cites Cola: consistent learning with opponent-learning awareness.

The challenge of hidden gifts in multi-agent reinforcement learning Cola: consistent learning with opponent-learning awareness

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:57:40.720516Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-08-07T13:57:40.206889Z digest=sha256:c8320a7f9c845e23c46c969be30ca48200f514583da6306fa0a7690c3162991d

Observation 4e88e8fe-79cc-47f6-9a09-8e48e4f66aa1 · outbound

This paper cites The surprising effectiveness of ppo in cooperative multi-agent games.

The challenge of hidden gifts in multi-agent reinforcement learning The surprising effectiveness of ppo in cooperative multi-agent games

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:57:40.611299Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-08-07T13:57:40.245927Z digest=sha256:0015ad29c1cffebcf6103e9158ba69838f0cae5b5f4a21481bf0da0bc21640c3

Pith citing papers

No inbound Pith citation observations are available.