Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-10T23:04:37.750141Z
Paper Citation Record · LEDGER
As of 11 August 2026, this Paper Citation Record lists 62 of 62 outbound references and 1 inbound Pith citation observation for arXiv:2501.00160.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-10T23:04:37.750141Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-11T06:34:44.6726+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-08-03T17:56:58.486176Z
A source-named dated measurement, never combined with another source.
Source: cited_works
62 of 62 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation c82d0212-b734-4ce6-8b5c-1de19bd6988e · outbound
Deterministic Model of Incremental Multi-Agent Boltzmann Q-Learning: Transient Cooperation, Metastability, and Oscillations Sutton and Andrew G
Reference 1
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation ede5c600-0596-4f40-916d-c7d11cdb3cba · outbound
Reference 2
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation fb0ee503-4963-4f04-ab85-80a77ddcb118 · outbound
Deterministic Model of Incremental Multi-Agent Boltzmann Q-Learning: Transient Cooperation, Metastability, and Oscillations A neural substrate of prediction and reward
Reference 3
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 2fe1db84-9e11-4071-a921-51e8eccedcf2 · outbound
Deterministic Model of Incremental Multi-Agent Boltzmann Q-Learning: Transient Cooperation, Metastability, and Oscillations Reinforcement learning: the good, the bad and the ugly
Reference 4
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 069be5b4-5c8c-4d98-8e7a-2872ba4acde2 · outbound
Deterministic Model of Incremental Multi-Agent Boltzmann Q-Learning: Transient Cooperation, Metastability, and Oscillations Read Montague
Reference 5
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f5710a1e-1e26-4a50-a54d-1c7fcb1720a7 · outbound
Deterministic Model of Incremental Multi-Agent Boltzmann Q-Learning: Transient Cooperation, Metastability, and Oscillations Playing Atari with Deep Reinforcement Learning
Reference 6
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6d0e6ee3-847e-4b95-9d40-4f8a78af7e18 · outbound
Deterministic Model of Incremental Multi-Agent Boltzmann Q-Learning: Transient Cooperation, Metastability, and Oscillations Human-level control through deep reinforcement learning
Reference 7
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 87b19751-182e-4e01-bd46-3be7295d6753 · outbound
Deterministic Model of Incremental Multi-Agent Boltzmann Q-Learning: Transient Cooperation, Metastability, and Oscillations Mastering the game of go with deep neural networks and tree search
Reference 8
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 90890e8b-073a-4732-9a24-9a4103ab86fe · outbound
Deterministic Model of Incremental Multi-Agent Boltzmann Q-Learning: Transient Cooperation, Metastability, and Oscillations Open Problems in Cooperative AI
Reference 9
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2f41f710-0c76-4ef8-8854-f531335c1e30 · outbound
Deterministic Model of Incremental Multi-Agent Boltzmann Q-Learning: Transient Cooperation, Metastability, and Oscillations Cooperative ai: machines must learn to find common ground, 2021
Reference 10
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation d0490090-1490-4591-8b81-8c9391b46e1e · outbound
Deterministic Model of Incremental Multi-Agent Boltzmann Q-Learning: Transient Cooperation, Metastability, and Oscillations Albrecht, Filippos Christianos, and Lukas Sch¨ afer
Reference 11
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 66e493d2-750a-4f41-8a3f-b583ca55f98b · outbound
Deterministic Model of Incremental Multi-Agent Boltzmann Q-Learning: Transient Cooperation, Metastability, and Oscillations Multi-agent reinforcement learning: independent vs
Reference 12
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 074fdb90-989e-4979-9108-594560fbbae3 · outbound
Deterministic Model of Incremental Multi-Agent Boltzmann Q-Learning: Transient Cooperation, Metastability, and Oscillations A Survey of Learning in Multiagent Environments: Dealing with Non-Stationarity
Reference 13
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5927b641-37e8-4527-a776-0eb37471a74f · outbound
Deterministic Model of Incremental Multi-Agent Boltzmann Q-Learning: Transient Cooperation, Metastability, and Oscillations Independent reinforcement learners in cooperative markov games: a survey regarding coordination problems
Reference 14
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 888d0c23-4265-4fc3-9858-4795c1980cd8 · outbound
Deterministic Model of Incremental Multi-Agent Boltzmann Q-Learning: Transient Cooperation, Metastability, and Oscillations A survey and critique of multiagent deep reinforcement learning
Reference 15
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 431129d0-3e96-4b6d-aeea-44cd422c9dfd · outbound
Deterministic Model of Incremental Multi-Agent Boltzmann Q-Learning: Transient Cooperation, Metastability, and Oscillations Benchmarking Multi-Agent Deep Reinforcement Learning Algorithms in Cooperative Tasks
Reference 16
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0687e0cd-9be0-450a-bcb2-dda5084f9e2d · outbound
Deterministic Model of Incremental Multi-Agent Boltzmann Q-Learning: Transient Cooperation, Metastability, and Oscillations Learning through reinforcement and replicator dynam- ics
Reference 17
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 755cd89e-5898-46fc-b5c3-189ffc5fdf46 · outbound
Deterministic Model of Incremental Multi-Agent Boltzmann Q-Learning: Transient Cooperation, Metastability, and Oscillations A selection-mutation model for q-learning in multi-agent systems
Reference 18
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation e0a297b2-13ba-4172-a790-f3315cba42ae · outbound
Reference 19
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 05adf438-9b55-4cb9-8e4a-8826f82021a7 · outbound
Reference 20
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 608361af-33bd-4ec0-ab50-0325dfeb9be6 · outbound
Reference 21
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 818a8dbe-0bb5-4302-9f5a-d6726841646b · outbound
Deterministic Model of Incremental Multi-Agent Boltzmann Q-Learning: Transient Cooperation, Metastability, and Oscillations Individual q-learning in normal form games
Reference 22
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 7c5e8bd7-72b2-4ec2-b3fb-f4fd290f5608 · outbound
Deterministic Model of Incremental Multi-Agent Boltzmann Q-Learning: Transient Cooperation, Metastability, and Oscillations Reinforcement learning dynamics in social dilemmas
Reference 23
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation eb3d957f-51e8-4d4a-88f3-1209bb6600f9 · outbound
Deterministic Model of Incremental Multi-Agent Boltzmann Q-Learning: Transient Cooperation, Metastability, and Oscillations Learning and equilibrium
Reference 24
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 53a100fc-dd1d-4a89-8e68-6773734a60c6 · outbound
Deterministic Model of Incremental Multi-Agent Boltzmann Q-Learning: Transient Cooperation, Metastability, and Oscillations Intrinsic noise in game dynamical learning
Reference 25
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 66a06a5b-048b-4f66-a589-3eb444863253 · outbound
Deterministic Model of Incremental Multi-Agent Boltzmann Q-Learning: Transient Cooperation, Metastability, and Oscillations A theoretical analysis of temporal difference learning in the iterated prisoner’s dilemma game
Reference 26
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation d596418a-fa44-4095-b7e4-b2ac82b9d517 · outbound
Deterministic Model of Incremental Multi-Agent Boltzmann Q-Learning: Transient Cooperation, Metastability, and Oscillations Classes of multiagent q-learning dynamics with epsilon-greedy exploration
Reference 27
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation a1b6a395-0b89-4a1e-bb55-6b32f1bdfa33 · outbound
Deterministic Model of Incremental Multi-Agent Boltzmann Q-Learning: Transient Cooperation, Metastability, and Oscillations Numerical analysis of a reinforcement learning model with the dynamic aspiration level in the iterated prisoner’s dilemma
Reference 28
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 7dbba3b5-f465-4bde-8d67-01301685e1b5 · outbound
Deterministic Model of Incremental Multi-Agent Boltzmann Q-Learning: Transient Cooperation, Metastability, and Oscillations Cycles of cooperation and defection in imperfect learning
Reference 29
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f1953e47-2139-4c63-bb1c-109378b2a46e · outbound
Deterministic Model of Incremental Multi-Agent Boltzmann Q-Learning: Transient Cooperation, Metastability, and Oscillations Dynamics of boltzmann q learning in two-player two-action games
Reference 30
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 6c92b88c-3c2f-4dde-9e7b-347a99363090 · outbound
Deterministic Model of Incremental Multi-Agent Boltzmann Q-Learning: Transient Cooperation, Metastability, and Oscillations Continuous strategy replicator dynamics for multi-agent q-learning
Reference 31
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation c6866ed2-8ffd-48fc-b9ab-ea5b82e58334 · outbound
Reference 32
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 32690eba-1bac-4a08-8285-be07bfd8d0ea · outbound
Deterministic Model of Incremental Multi-Agent Boltzmann Q-Learning: Transient Cooperation, Metastability, and Oscillations Evolutionary dynamics of multi-agent learning: A survey
Reference 33
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation c8f3080a-3b9c-44d5-8592-53aca54a9702 · outbound
Deterministic Model of Incremental Multi-Agent Boltzmann Q-Learning: Transient Cooperation, Metastability, and Oscillations Unresolved cited work
Reference 34
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0b931376-2dbd-4507-9242-c7bdc6a05b25 · outbound
Deterministic Model of Incremental Multi-Agent Boltzmann Q-Learning: Transient Cooperation, Metastability, and Oscillations Donges, and J¨ urgen Kurths
Reference 35
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation bc80f7e8-be7f-49e0-80b8-70a8c21ad044 · outbound
Deterministic Model of Incremental Multi-Agent Boltzmann Q-Learning: Transient Cooperation, Metastability, and Oscillations Modelling the dynamics of multiagent q-learning in repeated symmetric games: a mean field theoretic approach
Reference 36
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation d170359c-ecde-4b7a-b828-3f5ba27c1ee3 · outbound
Deterministic Model of Incremental Multi-Agent Boltzmann Q-Learning: Transient Cooperation, Metastability, and Oscillations Dynamical systems as a level of cognitive analysis of multi-agent learning
Reference 37
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation d5c323d6-1f11-4456-b750-81e19428dd89 · outbound
Deterministic Model of Incremental Multi-Agent Boltzmann Q-Learning: Transient Cooperation, Metastability, and Oscillations The Dynamics of Q-learning in Population Games: a Physics-Inspired Continuity Equation Model
Reference 38
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 464b5458-f6ed-43d9-a852-0e76a6dc80f2 · outbound
Deterministic Model of Incremental Multi-Agent Boltzmann Q-Learning: Transient Cooperation, Metastability, and Oscillations A formal model for multiagent q-learning dynamics on regular graphs
Reference 39
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 83891b36-205b-4e5b-8b72-46159c69cfd1 · outbound
Deterministic Model of Incremental Multi-Agent Boltzmann Q-Learning: Transient Cooperation, Metastability, and Oscillations Modeling the effects of environmental and perceptual uncertainty using deterministic reinforcement learning dynamics with partial observability
Reference 40
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation ae5d1519-d53c-4637-a811-f02ce349dffd · outbound
Deterministic Model of Incremental Multi-Agent Boltzmann Q-Learning: Transient Cooperation, Metastability, and Oscillations Exploration-exploitation in multi-agent learning: Catastrophe theory meets game theory
Reference 41
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation b204136b-54d9-4b1b-bf93-c084026b2a45 · outbound
Deterministic Model of Incremental Multi-Agent Boltzmann Q-Learning: Transient Cooperation, Metastability, and Oscillations Frequency adjusted multi-agent q-learning
Reference 42
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation dde5802c-95a4-4165-9a28-757e77cfda30 · outbound
Deterministic Model of Incremental Multi-Agent Boltzmann Q-Learning: Transient Cooperation, Metastability, and Oscillations Evolutionary Multi-agent Reinforcement Learning in Group Social Dilemmas
Reference 43
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 007617c1-8e8b-4f8c-ab5a-06becd7747ec · outbound
Deterministic Model of Incremental Multi-Agent Boltzmann Q-Learning: Transient Cooperation, Metastability, and Oscillations Sandholm and Robert H
Reference 44
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation ba7a4dca-6d42-4315-b8ab-b6ca5c79b24e · outbound
Deterministic Model of Incremental Multi-Agent Boltzmann Q-Learning: Transient Cooperation, Metastability, and Oscillations Faq-learning in matrix games: Demonstrating convergence near nash equilibria, and bifurcation of attractors in the battle of sexes
Reference 45
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation cf2886f0-3d29-4440-84c1-78048f471843 · outbound
Deterministic Model of Incremental Multi-Agent Boltzmann Q-Learning: Transient Cooperation, Metastability, and Oscillations L´ evy noise promotes cooperation in the prisoner’s dilemma game with reinforcement learning
Reference 46
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation da0760aa-9940-4be1-b097-a0450dcfaa2c · outbound
Deterministic Model of Incremental Multi-Agent Boltzmann Q-Learning: Transient Cooperation, Metastability, and Oscillations Limiting dynamics for q-learning with memory one in symmetric two-player, two-action games
Reference 47
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 2c497bf0-d8f4-45a3-b1dd-2ed07bd7b9fc · outbound
Deterministic Model of Incremental Multi-Agent Boltzmann Q-Learning: Transient Cooperation, Metastability, and Oscillations Q-learners can provably collude in the iterated prisoner’s dilemma
Reference 48
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 87679894-bb4d-4df0-be0a-d0ff9041c445 · outbound
Deterministic Model of Incremental Multi-Agent Boltzmann Q-Learning: Transient Cooperation, Metastability, and Oscillations Symmetric equilibrium of multi-agent reinforcement learning in repeated prisoner’s dilemma
Reference 49
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 7f307846-458d-4acd-9eab-b0e27043a061 · outbound
Deterministic Model of Incremental Multi-Agent Boltzmann Q-Learning: Transient Cooperation, Metastability, and Oscillations Q-learning in two-player two-action games
Reference 50
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation af5294dc-6586-4fa4-bb30-47d8b12cac0a · outbound
Deterministic Model of Incremental Multi-Agent Boltzmann Q-Learning: Transient Cooperation, Metastability, and Oscillations Melioration learning in iterated public goods games: The impact of exploratory noise
Reference 51
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 6dcea1a7-2004-4da3-8d4d-1e2f3c91b413 · outbound
Deterministic Model of Incremental Multi-Agent Boltzmann Q-Learning: Transient Cooperation, Metastability, and Oscillations Rein- forcement learning and decision making in monkeys during a competitive game
Reference 52
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation b524ab3f-a1ab-4c81-a39a-8e307a7893e1 · outbound
Deterministic Model of Incremental Multi-Agent Boltzmann Q-Learning: Transient Cooperation, Metastability, and Oscillations Valuation of uncertain and delayed rewards in primate prefrontal cortex
Reference 53
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation b408edeb-f28f-4075-a9c9-697d34c90c7e · outbound
Deterministic Model of Incremental Multi-Agent Boltzmann Q-Learning: Transient Cooperation, Metastability, and Oscillations Unresolved cited work
Reference 54
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 4420efb1-5bf7-463b-983d-58b89ed118c1 · outbound
Deterministic Model of Incremental Multi-Agent Boltzmann Q-Learning: Transient Cooperation, Metastability, and Oscillations Batch Reinforcement Learning, pages 45–73
Reference 55
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d390adb9-6282-47e6-955c-0976dca97129 · outbound
Deterministic Model of Incremental Multi-Agent Boltzmann Q-Learning: Transient Cooperation, Metastability, and Oscillations Quantal response equilibria for normal form games
Reference 56
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation b01182b5-d586-484e-8dfb-5fefa770ff00 · outbound
Deterministic Model of Incremental Multi-Agent Boltzmann Q-Learning: Transient Cooperation, Metastability, and Oscillations High-stakes failures of backward induction
Reference 57
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 00597c54-f9fb-4af8-8538-a57afa102217 · outbound
Deterministic Model of Incremental Multi-Agent Boltzmann Q-Learning: Transient Cooperation, Metastability, and Oscillations Timing of transients: quantifying reaching times and transient behavior in complex systems
Reference 58
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 334dd36e-5558-4080-b6d5-f283cc63ff50 · outbound
Deterministic Model of Incremental Multi-Agent Boltzmann Q-Learning: Transient Cooperation, Metastability, and Oscillations Unresolved cited work
Reference 59
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation f480cf6d-f0b7-465a-b8ee-bc4823072742 · outbound
Deterministic Model of Incremental Multi-Agent Boltzmann Q-Learning: Transient Cooperation, Metastability, and Oscillations ISBN 1581136838
Reference 2003
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 302cb385-d681-445d-a95a-a63c593f0215 · outbound
Deterministic Model of Incremental Multi-Agent Boltzmann Q-Learning: Transient Cooperation, Metastability, and Oscillations URL https://link.aps.org/doi/10.1103/ PhysRevLett.103.198702
Reference 2009
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 56916ade-5a51-445c-a2c6-1b176da410fa · outbound
Deterministic Model of Incremental Multi-Agent Boltzmann Q-Learning: Transient Cooperation, Metastability, and Oscillations URL http://dx.doi.org/10.1016/j.physd.2005
Reference 2789
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation f8fcddfd-3b42-4b8a-b4e6-acceba4a25ad · inbound
An Agent-Centric Dynamical Systems Perspective on Multi-Agent Reinforcement Learning Deterministic Model of Incremental Multi-Agent Boltzmann Q-Learning: Transient Cooperation, Metastability, and Oscillations
Reference 23
Source-reported events for the cited work
Unavailable: canonical work link unavailable.