Pith. sign in

Paper Citation Record · LEDGER

The challenge of hidden gifts in multi-agent reinforcement learning

As of 14 August 2026, this Paper Citation Record lists 32 of 32 outbound references and 0 inbound Pith citation observations for arXiv:2505.20579.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2505.20579 v7

Coverage vector

measured 32 of 32 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-07T13:57:40.245927Z

measured 32 of 32 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-14T06:32:32.682623+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

32 of 32 outbound references displayed

  • verified exact2
  • verified fuzzy16
  • unresolved14
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 0785d4bb-e775-4cfb-8e2c-d5f46fb58dfd · outbound

This paper cites LOQA : Learning with opponent q-learning awareness.

The challenge of hidden gifts in multi-agent reinforcement learning LOQA : Learning with opponent q-learning awareness

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:57:41.973000Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-08-07T13:57:38.263636Z digest=sha256:27f3214c1001523c7b1720eee0a1b23f8f4aba60e28333027cfed85b809bedc3

Observation 65b50dc3-88b8-429a-8651-2d1ae047486d · outbound

This paper cites Unifying temporal and structural credit assignment problems.

The challenge of hidden gifts in multi-agent reinforcement learning Unifying temporal and structural credit assignment problems

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:57:41.845941Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-08-07T13:57:38.403146Z digest=sha256:5bbf313c7acc3ad2a748141908489b94e6ba58640fc24e7bda4094750cfeaab1

Observation 45148112-d0c6-4aa3-becf-9b96e2203caf · outbound

This paper cites Understanding the impact of entropy on policy optimization.

The challenge of hidden gifts in multi-agent reinforcement learning Understanding the impact of entropy on policy optimization

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-07T13:57:38.517182Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:57:38.517182Z digest=sha256:3dd54cecab2f9eb368febb4bf68f3140780737f1e8d9c1af0e87c8c09f6eaed6

Observation 717fd1f9-4ef1-4ca7-92c3-7639a377b6b5 · outbound

This paper cites Effective choice in the prisoner's dilemma.

The challenge of hidden gifts in multi-agent reinforcement learning Effective choice in the prisoner's dilemma

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:57:41.762392Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-08-07T13:57:38.596201Z digest=sha256:9caa5862073536cb417af2d195419c532841e24a36fe87f5e0ad809ac74f99e0

Observation 545c2036-bae3-4537-a921-a7b2ee20f4d5 · outbound

This paper cites Manitokanac.

The challenge of hidden gifts in multi-agent reinforcement learning Manitokanac

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:57:41.704296Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-08-07T13:57:38.723112Z digest=sha256:7d4ee2910cf19d9fd49c2e0d8753002ef3fde59a984809cd37a010dc4b4fafcf

Observation 878a3ef2-3b7d-4462-b2c7-3f02961b23a5 · outbound

This paper cites The theory of dynamic programming.

The challenge of hidden gifts in multi-agent reinforcement learning The theory of dynamic programming

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-07T13:57:38.847739Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:57:38.847739Z digest=sha256:bbeba802764778ce19bd933a1fb9f26b3eb4d45d65fe0cd25f4136777a1b57b2

Observation fa69cbd5-ecc0-40ef-98f3-07dacafc609b · outbound

This paper cites Prisoner's dilemma; a study in conflict and cooperation.

The challenge of hidden gifts in multi-agent reinforcement learning Prisoner's dilemma; a study in conflict and cooperation

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:57:41.652409Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-08-07T13:57:38.955323Z digest=sha256:7cdbb0a20003acc04d93d0406adb3967175d2bea7d5be43092f299e3d218e92f

Observation fecfc178-cc40-4f6c-ad19-b731c3caee36 · outbound

This paper cites Decision transformer: Reinforcement learning via sequence modeling.

The challenge of hidden gifts in multi-agent reinforcement learning Decision transformer: Reinforcement learning via sequence modeling

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-07T13:57:39.056790Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:57:39.056790Z digest=sha256:6f5c27512f53b758437ec1631ca72188e001f8ddbc041b803fc29443e8d00e64

Observation c395c7bc-213f-4ad1-ab39-50dc2bea790a · outbound

This paper cites Minigrid & miniworld: Modular & customizable reinforcement learning environments for goal-oriented tasks.

The challenge of hidden gifts in multi-agent reinforcement learning Minigrid & miniworld: Modular & customizable reinforcement learning environments for goal-oriented tasks

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-07T13:57:39.091824Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:57:39.091824Z digest=sha256:a2c18f57bf841fd79aea8188f36c0dd6071550058da69c6a42e42ff5876de3be

Observation 440b0c45-833b-4302-9b72-75deec581e6f · outbound

This paper cites Learning phrase representations using RNN encoder -- decoder for statistical machine translation.

The challenge of hidden gifts in multi-agent reinforcement learning Learning phrase representations using RNN encoder -- decoder for statistical machine translation

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-07T13:57:39.129254Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:57:39.129254Z digest=sha256:29ec884b97e20ce71e0d38ec8501e2c6cad601ddc09a330d99a85637a49bc389

Observation 1429d519-a76c-4b38-8964-b27f8e51bc77 · outbound

This paper cites Hypothetical minds: Scaffolding theory of mind for multi-agent tasks with large language models.

The challenge of hidden gifts in multi-agent reinforcement learning Hypothetical minds: Scaffolding theory of mind for multi-agent tasks with large language models

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:57:41.604374Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-08-07T13:57:39.183179Z digest=sha256:9be87b6fffe7070b9664884c9cbb1c1bcc47029ad5ae5aecff63429afcf62f43

Observation ee237fee-8184-496b-b614-bec4548cb90b · outbound

This paper cites Maximum entropy RL (provably) solves some robust RL problems.

The challenge of hidden gifts in multi-agent reinforcement learning Maximum entropy RL (provably) solves some robust RL problems

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:57:41.525145Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-08-07T13:57:39.240202Z digest=sha256:2a39ccf4038cc46fd7e47b66a0f10b6acc128ee91f8133710a79b540eb149da4

Observation edcda3d2-97cb-44e2-b238-7b58dad96676 · outbound

This paper cites Counterfactual multi-agent policy gradients.

The challenge of hidden gifts in multi-agent reinforcement learning Counterfactual multi-agent policy gradients

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-07T13:57:39.298552Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:57:39.298552Z digest=sha256:2b3ea85fe195a932afea5e3e3e7e82e36d1bca204242be3216f28cb292514137

Observation 1d218df6-f17b-4151-b312-7a51599cfd3e · outbound

This paper cites Learning with Opponent-Learning Awareness.

The challenge of hidden gifts in multi-agent reinforcement learning Learning with Opponent-Learning Awareness

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-07T13:57:39.358832Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:57:39.358832Z digest=sha256:fca71e6db909024a2ce79c6ade92fd59350a475b05797fe01671b70186715caf

Observation fa5fa247-7d8a-4ce1-90c4-4e7525eabf33 · outbound

This paper cites Structural credit assignment in neural networks using reinforcement learning.

The challenge of hidden gifts in multi-agent reinforcement learning Structural credit assignment in neural networks using reinforcement learning

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:57:41.422159Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-08-07T13:57:39.409064Z digest=sha256:7205e6b57c2040fdda9031dd9c15cb9f510521507d3a856ed0587a6be10d472a

Observation 23447fde-0e28-4662-8aaf-02fb51e13281 · outbound

This paper cites Soft actor-critic: Off-policy maximum entropy deep reinforcement learning with a stochastic actor.

The challenge of hidden gifts in multi-agent reinforcement learning Soft actor-critic: Off-policy maximum entropy deep reinforcement learning with a stochastic actor

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-07T13:57:39.454659Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:57:39.454659Z digest=sha256:7c5489603b0b2a1cadcd597e809ec9da455231e215384e08d533487496928802

Observation c9e0a6a7-9510-41ae-9122-827df185a5cf · outbound

This paper cites Optimizing agent behavior over long time scales by transporting value.

The challenge of hidden gifts in multi-agent reinforcement learning Optimizing agent behavior over long time scales by transporting value

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:57:41.293111Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-08-07T13:57:39.511697Z digest=sha256:b5217c6e6c96276a0af3b2298742747511dc492e7d7ce6bebc9485da93f9b099

Observation 0e6b197b-fa05-4ba0-96bc-c5aeff909f22 · outbound

This paper cites Stateful active facilitator: Coordination and environmental heterogeneity in cooperative multi-agent reinforcement learning.

The challenge of hidden gifts in multi-agent reinforcement learning Stateful active facilitator: Coordination and environmental heterogeneity in cooperative multi-agent reinforcement learning

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:57:41.188200Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-08-07T13:57:39.568734Z digest=sha256:0f3fa559324f1792aa0d07da5a1c005374bfc2589e1ef1c7b940a7ad436eaa80

Observation 38f3c47b-f42f-45cf-add0-5cde06837423 · outbound

This paper cites Maven: Multi-agent variational exploration.

The challenge of hidden gifts in multi-agent reinforcement learning Maven: Multi-agent variational exploration

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-07T13:57:39.609339Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:57:39.609339Z digest=sha256:b68f5960897d6dddf4151d576a1db40c01e1164ce4b7a1114de29f0117c5d900

Observation 5778cf27-9108-4fa6-9e27-766e4d1956f5 · outbound

This paper cites Multi-agent cooperation through learning-aware policy gradients.

The challenge of hidden gifts in multi-agent reinforcement learning Multi-agent cooperation through learning-aware policy gradients

Reference 20

Resolution
verified exact
local_arxiv, observed 2026-08-07T13:57:40.513690Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-08-07T13:57:39.665521Z digest=sha256:5100b8112d0ddebeda52b27d283a7fd131a0f27010f10c0c3a38a4d19ebbe055

Observation 86740e3a-9d60-4325-aeb7-1832e2994eb9 · outbound

This paper cites Equilibrium points in n-person games.

The challenge of hidden gifts in multi-agent reinforcement learning Equilibrium points in n-person games

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-07T13:57:39.695724Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:57:39.695724Z digest=sha256:e2f41a4d3b70b23771d82361239eef3af0b2e9c5356478bb55a2f36bd8997925

Observation 5cf85eeb-2d4b-4a4e-b747-7bc456e0db6b · outbound

This paper cites When do transformers shine in rl? decoupling memory from credit assignment.

The challenge of hidden gifts in multi-agent reinforcement learning When do transformers shine in rl? decoupling memory from credit assignment

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:57:41.109096Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-08-07T13:57:39.752780Z digest=sha256:45a3fe5409cf4d9b22ff9adc619486c4789d5fec08e2caafb9c842c4aff8416e

Observation 55335269-7cf6-46a5-9871-7febce7d37f4 · outbound

This paper cites Monotonic value function factorisation for deep multi-agent reinforcement learning.

The challenge of hidden gifts in multi-agent reinforcement learning Monotonic value function factorisation for deep multi-agent reinforcement learning

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-07T13:57:39.809487Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:57:39.809487Z digest=sha256:97b663eb1ac11ebc4accab1f3fa086f3f7fa2a4a18d00285d9cfe0f47eb5a72e

Observation 451b9fd3-8fe7-4219-8eaa-1f1103f85942 · outbound

This paper cites Proximal Policy Optimization Algorithms.

The challenge of hidden gifts in multi-agent reinforcement learning Proximal Policy Optimization Algorithms

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-07T13:57:39.857334Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:57:39.857334Z digest=sha256:1b178b333987561c1e0a6e0769d058a0fe0b936db3043152bb49bc7ea3ac0805

Observation e3603498-d8ec-4585-ae22-691096eaa45e · outbound

This paper cites Agent-Time Attention for Sparse Rewards Multi-Agent Reinforcement Learning.

The challenge of hidden gifts in multi-agent reinforcement learning Agent-Time Attention for Sparse Rewards Multi-Agent Reinforcement Learning

Reference 25

Resolution
verified exact
local_arxiv, observed 2026-08-07T13:57:40.393765Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-08-07T13:57:39.909858Z digest=sha256:afae37bc91865bdf1898ef4f82340febb177156203f51dd9a7b189e0976cd0d0

Observation a839fd47-f4ca-41e4-84a2-aa7124976a93 · outbound

This paper cites Qtran: Learning to factorize with transformation for cooperative multi-agent reinforcement learning.

The challenge of hidden gifts in multi-agent reinforcement learning Qtran: Learning to factorize with transformation for cooperative multi-agent reinforcement learning

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:57:41.008477Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-08-07T13:57:39.956054Z digest=sha256:7517c3f49ad2245893a01fe9731650f142afe5fd016c5a7f27c4b33403381565

Observation db7404eb-e5c0-4b34-a830-334337701ea0 · outbound

This paper cites PufferLib: Making Reinforcement Learning Libraries and Environments Play Nice.

The challenge of hidden gifts in multi-agent reinforcement learning PufferLib: Making Reinforcement Learning Libraries and Environments Play Nice

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-07T13:57:40.035699Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:57:40.035699Z digest=sha256:c054298199e4775aac614171657ba1df811cac0a0de0d88d48544f879be47ecc

Observation 91fd2b73-84ff-4050-9654-cc1a7d4f5e5a · outbound

This paper cites Value-decomposition networks for cooperative multi-agent learning.

The challenge of hidden gifts in multi-agent reinforcement learning Value-decomposition networks for cooperative multi-agent learning

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:57:40.936484Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-08-07T13:57:40.079910Z digest=sha256:a91d57905eec8621e536b17420a6566217f8890e58638e46fe33a1e19289980b

Observation 6588f00d-0d2f-4415-a0a9-d4590e78ca2a · outbound

This paper cites Policy gradient methods for reinforcement learning with function approximation.

The challenge of hidden gifts in multi-agent reinforcement learning Policy gradient methods for reinforcement learning with function approximation

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-07T13:57:40.131961Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:57:40.131961Z digest=sha256:d77e401e2f6efb3bffd6c78c45e1de5f71288683bd76e7264ce9642a5d9e1f8c

Observation 32731990-74ff-40fb-8c52-8d65ba72f489 · outbound

This paper cites Learning sequences of actions in collectives of autonomous agents.

The challenge of hidden gifts in multi-agent reinforcement learning Learning sequences of actions in collectives of autonomous agents

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:57:40.819723Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-08-07T13:57:40.173465Z digest=sha256:4ac3c3ad294b0497dd03c6ed00d2131243639e2769ae3f043d8f8d1a63e58859

Observation 4873364c-44db-423f-8ed7-6303da79cba4 · outbound

This paper cites Cola: consistent learning with opponent-learning awareness.

The challenge of hidden gifts in multi-agent reinforcement learning Cola: consistent learning with opponent-learning awareness

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:57:40.720516Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-08-07T13:57:40.206889Z digest=sha256:3489edba3e28b7ad089b4f2d69e352156a3f4ec575016db162e6f0e09bf45f56

Observation 4e88e8fe-79cc-47f6-9a09-8e48e4f66aa1 · outbound

This paper cites The surprising effectiveness of ppo in cooperative multi-agent games.

The challenge of hidden gifts in multi-agent reinforcement learning The surprising effectiveness of ppo in cooperative multi-agent games

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:57:40.611299Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-08-07T13:57:40.245927Z digest=sha256:e81bf02c3e1b802a6bee3992860ec27b7c4a6bfbf9598a37232adee1f246f97e

Pith citing papers

No inbound Pith citation observations are available.