Pith. sign in

Paper Citation Record · LEDGER

Metric-Gradient Projection for Stable Multi-Agent Policy Learning

As of 8 August 2026, this Paper Citation Record lists 20 of 20 outbound references and 2 inbound Pith citation observations for arXiv:2605.18809.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2605.18809 v1

Coverage vector

measured 20 of 20 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-05-20T22:13:00.107205Z

measured 22 of 22 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-08T06:32:00.761636+00:00

measured 2 of 2 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-06T00:45:45.078825Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-08-06T00:45:45.612357Z

Reference resolution

20 of 20 outbound references displayed

  • verified exact13
  • verified fuzzy5
  • unresolved1
  • parse uncertain0
  • malformed identifier1
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 431d1d1c-af84-48bb-8dd3-dc82d76142c2 · outbound

This paper cites Markov games as a framework for multi-agent reinforcement learning.

Metric-Gradient Projection for Stable Multi-Agent Policy Learning Markov games as a framework for multi-agent reinforcement learning

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-05-20T22:13:47.550963Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-20T22:13:00.107205Z digest=sha256:6c8d9593cdcf29cd10267ead3b324b34f2e45fa79bfc2eb55b0b5c008291cc00

Observation 5f6b9916-7d09-45dc-8f47-b98d89f3e164 · outbound

This paper cites Learning to collaborate with unknown agents in the absence of reward.

Metric-Gradient Projection for Stable Multi-Agent Policy Learning Learning to collaborate with unknown agents in the absence of reward

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-05-20T22:13:47.542787Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-20T22:13:00.107205Z digest=sha256:ab35526ce5d6f62b5ce548d1e242b2891bb188d6c08d2fb86b759a748911ce00

Observation cc9f3ea4-086c-4107-b949-0480238b3f36 · outbound

This paper cites Modeling Other Players with Bayesian Beliefs for Games with Incomplete Information.

Metric-Gradient Projection for Stable Multi-Agent Policy Learning Modeling Other Players with Bayesian Beliefs for Games with Incomplete Information

Reference 3

Resolution
verified exact
arxiv_id, observed 2026-05-20T22:13:46.730580Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-20T22:13:00.107205Z digest=sha256:3258e28cc2097777b69bcf3742b4165f4c6c5ceab36e732572828b64e617ba09

Observation cf33e756-3e32-44b9-aa43-1f5a4befbeda · outbound

This paper cites Lipschitz Lifelong Monte Carlo Tree Search for Mastering Non-Stationary Tasks.

Metric-Gradient Projection for Stable Multi-Agent Policy Learning Lipschitz Lifelong Monte Carlo Tree Search for Mastering Non-Stationary Tasks

Reference 4

Resolution
verified exact
arxiv_id, observed 2026-05-20T22:13:46.755847Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-20T22:13:00.107205Z digest=sha256:5284532901759bec83edf847714437d1a7b6fa3b198d90a76c6ad92c343c65a3

Observation 7a72cee9-0251-4351-a65b-890e3433b57d · outbound

This paper cites Geometry of drifting mdps with path-integral stability certificates.

Metric-Gradient Projection for Stable Multi-Agent Policy Learning Geometry of drifting mdps with path-integral stability certificates

Reference 5

Resolution
verified exact
arxiv_id, observed 2026-05-20T22:13:46.745527Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-20T22:13:00.107205Z digest=sha256:0beb1ec821ba4c85c71571c3a0cec86d274f074ecc6cb0d6399e8977e5c24843

Observation d1cf2afa-db78-4c34-b2ee-0f2738b3b7f6 · outbound

This paper cites High-Dimensional Continuous Control Using Generalized Advantage Estimation.

Metric-Gradient Projection for Stable Multi-Agent Policy Learning High-Dimensional Continuous Control Using Generalized Advantage Estimation

Reference 6

Resolution
verified exact
local_arxiv, observed 2026-05-20T22:13:46.740818Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-20T22:13:00.107205Z digest=sha256:2197e1199675d5166d5a74f34854e0b9d6a416b580859905c66c7d0ed6408a89

Observation 71da555e-6cdf-4216-9200-4bdc9f925f19 · outbound

This paper cites Value-Decomposition Networks For Cooperative Multi-Agent Learning.

Metric-Gradient Projection for Stable Multi-Agent Policy Learning Value-Decomposition Networks For Cooperative Multi-Agent Learning

Reference 7

Resolution
verified exact
local_arxiv, observed 2026-05-20T22:13:46.750682Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-20T22:13:00.107205Z digest=sha256:7300c17126705f664bbb2371d8081d3dc57c575739f4d26d98f6bafdfa8c0016

Observation d59117f1-7687-457b-a8a9-845e873dae32 · outbound

This paper cites A Variational Inequality Perspective on Generative Adversarial Networks.

Metric-Gradient Projection for Stable Multi-Agent Policy Learning A Variational Inequality Perspective on Generative Adversarial Networks

Reference 8

Resolution
verified exact
arxiv_id, observed 2026-05-20T22:13:46.735866Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-20T22:13:00.107205Z digest=sha256:adb1358257138180ea826e72021732102fc0b50f1052335c71597708e8cb75e1

Observation 00259185-6ade-4ca5-a807-c3042e6f934a · outbound

This paper cites an unresolved cited work.

Metric-Gradient Projection for Stable Multi-Agent Policy Learning Unresolved cited work

Reference 9

Resolution
unresolved
raw_fallback, observed 2026-05-20T22:13:47.536600Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-20T22:13:00.107205Z digest=sha256:136f013ecb8337fae57fc5c12e300e73cbe77941309e0c82c09daa4838217dd3

Observation 62bb4fd6-ead0-4037-a239-64bbf3ad5d45 · outbound

This paper cites neurips.cc/paper/2005/hash/9752d873fa71c19dc602bf2a0696f9b5-Abstract.html.

Metric-Gradient Projection for Stable Multi-Agent Policy Learning neurips.cc/paper/2005/hash/9752d873fa71c19dc602bf2a0696f9b5-Abstract.html

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-05-20T22:13:47.530040Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-20T22:13:00.107205Z digest=sha256:f1d675dad2da072c8405ea712cd0a2d2b7aca154d49a3958e9741e86db1d5406

Observation 7fe04ec0-4364-45d5-9891-d4b31e0832d8 · outbound

This paper cites Optimistic mirror descent in saddle-point problems: Going the extra (gradient) mile.

Metric-Gradient Projection for Stable Multi-Agent Policy Learning Optimistic mirror descent in saddle-point problems: Going the extra (gradient) mile

Reference 11

Resolution
verified exact
local_arxiv, observed 2026-05-20T22:13:46.714040Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-20T22:13:00.107205Z digest=sha256:05a985c9b83c0d3c6e49bdb2147ea06dfcd726ab89a945bdc6a2b32daa8dea05

Observation 76497578-d2c3-4c27-9e54-8c5b78121797 · outbound

This paper cites Training GANs with Optimism.

Metric-Gradient Projection for Stable Multi-Agent Policy Learning Training GANs with Optimism

Reference 12

Resolution
verified exact
local_arxiv, observed 2026-05-20T22:13:46.735399Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-20T22:13:00.107205Z digest=sha256:c276dd7432573424d74abf822d707dd210c3b654fb0641c11f70344dff1ee154

Observation c1475e4b-3f28-4bc2-bfe8-f5199a3f1f71 · outbound

This paper cites Algorithms, graph theory, and linear equations in laplacian matrices.

Metric-Gradient Projection for Stable Multi-Agent Policy Learning Algorithms, graph theory, and linear equations in laplacian matrices

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-05-20T22:13:47.528832Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-20T22:13:00.107205Z digest=sha256:4a24f10dfc7d94891f801a2f99d9e388a2769173703124e2de8bf09da85bc3cf

Observation 8ba9d677-3c5b-4b4e-9b5d-219586756a36 · outbound

This paper cites Discrete calculus: Applied analysis on graphs for computational science, leo grady, jonathan polimeni, springer (2010), $129.00, isbn: 978-1-84996-289-6.

Metric-Gradient Projection for Stable Multi-Agent Policy Learning Discrete calculus: Applied analysis on graphs for computational science, leo grady, jonathan polimeni, springer (2010), $129.00, isbn: 978-1-84996-289-6

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-05-20T22:13:47.515900Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-20T22:13:00.107205Z digest=sha256:68a5f0a1e2830a4e8be9d51cac1b6619cd53d18b1b2d52e7ed3b9583feacd806

Observation d321031c-eb48-4861-b417-ee43ed033c99 · outbound

This paper cites Cochain perspectives on temporal-difference signals for learning beyond markov dynamics.arXiv preprint arXiv:2602.06939, 2026b.

Metric-Gradient Projection for Stable Multi-Agent Policy Learning Cochain perspectives on temporal-difference signals for learning beyond markov dynamics.arXiv preprint arXiv:2602.06939, 2026b

Reference 15

Resolution
verified exact
arxiv_id, observed 2026-05-20T22:13:46.725417Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-20T22:13:00.107205Z digest=sha256:b36fb2efdaf47eb5edd047d2cd79a86617f8ef3af47a2baad1cab0411174f314

Observation d6bf5b80-299b-4bec-804d-3ff3c5d4d63a · outbound

This paper cites Melting Pot 2.0.

Metric-Gradient Projection for Stable Multi-Agent Policy Learning Melting Pot 2.0

Reference 16

Resolution
verified exact
arxiv_id, observed 2026-05-20T22:13:46.731145Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-20T22:13:00.107205Z digest=sha256:f7a327d41ebca187f030e7431fa17d7ca236611d912c5b238df1300da2b3bd02

Observation 8be89070-a4b1-43b4-9bfb-a0c50b7d81b2 · outbound

This paper cites Agent alpha: Tree search unifying generation, exploration and evaluation for computer-use agents.arXiv preprint arXiv:2602.02995.

Metric-Gradient Projection for Stable Multi-Agent Policy Learning Agent alpha: Tree search unifying generation, exploration and evaluation for computer-use agents.arXiv preprint arXiv:2602.02995

Reference 17

Resolution
verified exact
arxiv_id, observed 2026-05-20T22:13:46.739801Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-20T22:13:00.107205Z digest=sha256:bab4231e35c66889b5d422a8b4aeb898f94de67fa2c630fe4b409db19817bae9

Observation bf6547c0-7d3a-4068-8e95-5c9bc24a9ec3 · outbound

This paper cites NonZero: Interaction-Guided Exploration for Multi-Agent Monte Carlo Tree Search.

Metric-Gradient Projection for Stable Multi-Agent Policy Learning NonZero: Interaction-Guided Exploration for Multi-Agent Monte Carlo Tree Search

Reference 18

Resolution
verified exact
local_arxiv, observed 2026-05-20T22:13:46.709264Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-20T22:13:00.107205Z digest=sha256:36136c4d784335cd8644ab00aa1f87b4f6f1bcb5c062f0446992c4193e2e796c

Observation 7fe68f37-42e7-4e88-93bd-679eaf6c85e7 · outbound

This paper cites Lisfc-search: Lifelong search for network sfc optimization under non-stationary drifts.arXiv preprint arXiv:2602.14360, 2026a.

Metric-Gradient Projection for Stable Multi-Agent Policy Learning Lisfc-search: Lifelong search for network sfc optimization under non-stationary drifts.arXiv preprint arXiv:2602.14360, 2026a

Reference 19

Resolution
verified exact
arxiv_id, observed 2026-05-20T22:13:46.744924Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-20T22:13:00.107205Z digest=sha256:57495edd5f25e6f87cd2d72bcc279871abdeb7b7da3064cf278712ed459a1074

Observation 24334fc3-f4fa-41d3-9a93-1f7445636fe1 · outbound

This paper cites First, cyclic matrix games, including Rock–Paper–Scissors and generalized cyclic games on∆K ×∆ K, expose cyclic interaction dynamics in a familiar simplex geometry.

Metric-Gradient Projection for Stable Multi-Agent Policy Learning First, cyclic matrix games, including Rock–Paper–Scissors and generalized cyclic games on∆K ×∆ K, expose cyclic interaction dynamics in a familiar simplex geometry

Reference 20

Resolution
malformed identifier
raw_fallback, observed 2026-05-20T22:13:47.525688Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-20T22:13:00.107205Z digest=sha256:51787f63bcebceb89b72af0e99b344e29d582059046f45694b9ca92cd1a78b6d

Pith citing papers

Observation 32377c4b-0ede-4256-9a68-3ab688476fde · inbound

FlowEdit: Information-Theoretic Control of LLM Reasoning Flows for Ill-posed Problems Involving Conflicts cites this paper.

FlowEdit: Information-Theoretic Control of LLM Reasoning Flows for Ill-posed Problems Involving Conflicts Metric-Gradient Projection for Stable Multi-Agent Policy Learning

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-02T10:39:41.077110Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T10:39:41.077110Z digest=sha256:2d0fd67787a59907ae82dc2b0dddf2baf9dc52945f7996a571a0420f79b84aef

Observation 41ada73a-9ebb-4a48-bad2-b28bc9ef4772 · inbound

Learning Not to Optimize: Physics-Informed Action-Space Reshaping for Intent-Based Network Control cites this paper.

Learning Not to Optimize: Physics-Informed Action-Space Reshaping for Intent-Based Network Control Metric-Gradient Projection for Stable Multi-Agent Policy Learning

Reference 46

Resolution
verified exact
local_arxiv, observed 2026-08-06T00:45:45.617991Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T00:45:45.078825Z digest=sha256:9adb9de630ad82b6f9b3d52be11bf2d0413fac600028b1be37156e85c5382b31