Pith. sign in

Paper Citation Record · LEDGER

Trust Region Policy Optimisation in Multi-Agent Reinforcement Learning

As of 10 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 18 inbound Pith citation observations for arXiv:2109.11251.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2109.11251 v2

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 18 of 18 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-10T06:31:04.303077+00:00

measured 18 of 18 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-08T21:35:36.313216Z

measured 1 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-08-05T02:28:24.338817Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

83
arxiv_reference, observed 2026-08-05T02:28:24.338817Z

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation f5224185-f41a-4c17-a900-bfd7ca74fecd · inbound

An Extended Benchmarking of Multi-Agent Reinforcement Learning Algorithms in Complex Fully Cooperative Tasks cites this paper.

An Extended Benchmarking of Multi-Agent Reinforcement Learning Algorithms in Complex Fully Cooperative Tasks Trust Region Policy Optimisation in Multi-Agent Reinforcement Learning

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-08T21:35:36.313216Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T21:35:36.313216Z digest=sha256:09b520f2800590d57ba1f36dd8c83e9631f5d14dcdeedbf941f312d8f83e4155

Observation 364f7221-7d8a-4d5f-8aa8-8232c2eda30f · inbound

Low-Rank Agent-Specific Adaptation (LoRASA) for Multi-Agent Policy Learning cites this paper.

Low-Rank Agent-Specific Adaptation (LoRASA) for Multi-Agent Policy Learning Trust Region Policy Optimisation in Multi-Agent Reinforcement Learning

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-08T18:50:49.824318Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T18:50:49.824318Z digest=sha256:1dee9c4270fb7c5c72e0f722ac93e4b3fb161c7e889e7695b0a32ba5f707b276

Observation 1cff9fb9-eb41-4802-a372-4046bfc8210e · inbound

Light Aircraft Game : Basic Implementation and training results analysis cites this paper.

Light Aircraft Game : Basic Implementation and training results analysis Trust Region Policy Optimisation in Multi-Agent Reinforcement Learning

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-07T00:23:07.654898Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:23:07.654898Z digest=sha256:7dcae8961b9b099471f9d43ad1aa4371b3195b58e0fb4483323cb039a4364b1e

Observation b9f7b077-a641-4f1e-b536-eba274abbe39 · inbound

Improving monotonic optimization in heterogeneous multi-agent reinforcement learning with optimal marginal deterministic policy gradient cites this paper.

Improving monotonic optimization in heterogeneous multi-agent reinforcement learning with optimal marginal deterministic policy gradient Trust Region Policy Optimisation in Multi-Agent Reinforcement Learning

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-06T17:47:37.242154Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T17:47:37.242154Z digest=sha256:2abfc86300a61abaa32cebbf502d47eb4fce201ca2d7246cf17de5d59aa70d5e

Observation ecd55247-3864-45aa-92d3-e5006119891a · inbound

Centralized Adaptive Sampling for Reliable Co-Training of Independent Multi-Agent Policies cites this paper.

Centralized Adaptive Sampling for Reliable Co-Training of Independent Multi-Agent Policies Trust Region Policy Optimisation in Multi-Agent Reinforcement Learning

Reference 47

Resolution
metadata mismatch
arxiv_id, observed 2026-05-19T01:06:57.409927Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-19T01:05:17.086399Z digest=sha256:c73a00427b714441d3eeca76db1bcb6b3e44f7e711c28466f452e851f7da968f

Observation 924df6fc-5af8-4b3f-8a73-96966cc68a29 · inbound

Multi-Agent Trust Region Policy Optimisation: A Joint Constraint Approach cites this paper.

Multi-Agent Trust Region Policy Optimisation: A Joint Constraint Approach Trust Region Policy Optimisation in Multi-Agent Reinforcement Learning

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-05T20:31:02.434225Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T20:31:02.434225Z digest=sha256:12a12ab21df39c1df195cf6cf86018bbca5a404a0dfe4eac5c330a39ea7c9d61

Observation ce1cbd52-a50d-43df-9d3c-f212f995948f · inbound

TABX: A High-Throughput Sandbox Battle Simulator for Multi-Agent Reinforcement Learning cites this paper.

TABX: A High-Throughput Sandbox Battle Simulator for Multi-Agent Reinforcement Learning Trust Region Policy Optimisation in Multi-Agent Reinforcement Learning

Reference 9

Resolution
verified exact
arxiv_id, observed 2026-05-25T07:45:29.793674Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-25T07:41:48.634351Z digest=sha256:2ee09cd69aeeb5d4ab8c275d066ed4da16411978e2575fa26706383d91b48dfd

Observation acc98b05-d732-4449-9d73-9c2f0350a4cc · inbound

TABX: A High-Throughput Sandbox Battle Simulator for Multi-Agent Reinforcement Learning cites this paper.

TABX: A High-Throughput Sandbox Battle Simulator for Multi-Agent Reinforcement Learning Trust Region Policy Optimisation in Multi-Agent Reinforcement Learning

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-03T05:37:38.440274Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T05:37:38.440274Z digest=sha256:f3d3e87643895f898cff7d503848110f930cf70c6c5f844994ca386d5076c760

Observation 1d29723a-b446-4a5e-9b8e-cae77fc18570 · inbound

Topology-Driven Anti-Entanglement Control for Soft Robots cites this paper.

Topology-Driven Anti-Entanglement Control for Soft Robots Trust Region Policy Optimisation in Multi-Agent Reinforcement Learning

Reference 12

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T15:41:28.058511Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-09T19:28:39.040470Z digest=sha256:d903f74cd01dc8709bee1cba1bda6bcf59c5a2e218ad985c678f1ab71a48f42d

Observation b778fc2f-accf-435a-8f96-be38f0240d6c · inbound

Robust Instruction Compliance in Cooperative Multi-Agent Reinforcement Learning cites this paper.

Robust Instruction Compliance in Cooperative Multi-Agent Reinforcement Learning Trust Region Policy Optimisation in Multi-Agent Reinforcement Learning

Reference 122

Resolution
verified exact
arxiv_id, observed 2026-06-30T22:15:05.903774Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-06-30T22:11:35.277901Z digest=sha256:29de71a07def3bec791dfe44feaa7f18ddece577e329701f397afa0477f65889

Observation 401883ea-8851-4bfa-9b3c-1bfcc7509581 · inbound

Self-Supervised On-Policy Reinforcement Learning via Contrastive Proximal Policy Optimisation cites this paper.

Self-Supervised On-Policy Reinforcement Learning via Contrastive Proximal Policy Optimisation Trust Region Policy Optimisation in Multi-Agent Reinforcement Learning

Reference 10

Resolution
verified exact
arxiv_id, observed 2026-05-14T20:12:54.902012Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-14T20:12:07.920183Z digest=sha256:28492318a1e604e0837587abb245677b7d1fe112be1324ffbaee49309ab7139f

Observation 1a6b8c58-8480-4861-9f48-66da0251baad · inbound

TeamTR: Trust-Region Fine-Tuning for Multi-Agent LLM Coordination cites this paper.

TeamTR: Trust-Region Fine-Tuning for Multi-Agent LLM Coordination Trust Region Policy Optimisation in Multi-Agent Reinforcement Learning

Reference 60

Resolution
metadata mismatch
arxiv_id, observed 2026-05-19T18:02:42.323932Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-05-19T18:01:06.649723Z digest=sha256:7b8f237fec3e2cc4b0fb5e57eddf7f9fd41ce71009909c9570774b54e9552372

Observation 113eb189-443b-40bf-9fcc-d7de7a1804dc · inbound

Phi-Actor-Critic: Steering General-Sum Games to Pareto-Efficient Correlated Equilibria cites this paper.

Phi-Actor-Critic: Steering General-Sum Games to Pareto-Efficient Correlated Equilibria Trust Region Policy Optimisation in Multi-Agent Reinforcement Learning

Reference 40

Resolution
metadata mismatch
arxiv_id, observed 2026-07-03T08:47:50.453357Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-06-27T10:42:17.852996Z digest=sha256:89586e1db2cf738c8ec5bf90c32ddc3147bdffa4758de2c5528dd8993060949b

Observation bbbcca0b-11cc-4c9c-93b3-d697a1a39799 · inbound

Reference-Free Heterogeneous Multi-Agent Reinforcement Learning for Grid-Friendly Tie-Line Power Shaping in Industrial Microgrids cites this paper.

Reference-Free Heterogeneous Multi-Agent Reinforcement Learning for Grid-Friendly Tie-Line Power Shaping in Industrial Microgrids Trust Region Policy Optimisation in Multi-Agent Reinforcement Learning

Reference 34

Resolution
metadata mismatch
arxiv_id, observed 2026-07-04T20:50:11.104720Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-06-25T19:36:51.509710Z digest=sha256:7e8a53d84d2665741c7f3aff0f2753c6e9c500ce484eeee4dae98bd948d797c3

Observation 235d54c4-9e76-462f-9e21-80b4af10140d · inbound

Action-Factored Multi-Agent Reinforcement Learning for Scalable Quantum Device Tuning cites this paper.

Action-Factored Multi-Agent Reinforcement Learning for Scalable Quantum Device Tuning Trust Region Policy Optimisation in Multi-Agent Reinforcement Learning

Reference 63

Resolution
unresolved
no resolver link, observed 2026-07-13T03:08:01.590659Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-13T03:08:01.590659Z digest=sha256:563da3e49ccf2f33d89dcf7b11899b5c4b196d54fe54ccbd39e4bfc52b7c5280

Observation 8d46f04a-1a2d-4c1b-87ba-8bdd023b8cd3 · inbound

Heterogeneous Multi-Agent Reinforcement Learning for Radio Resource Management under Coupled Finite-Horizon Constraints cites this paper.

Heterogeneous Multi-Agent Reinforcement Learning for Radio Resource Management under Coupled Finite-Horizon Constraints Trust Region Policy Optimisation in Multi-Agent Reinforcement Learning

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-04T21:47:09.010614Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T21:47:09.010614Z digest=sha256:b3f23ef0c79614a2db1dc6813634717d60a8d917b883fa088ce6af7fdb7d942d

Observation 748a4a37-ff95-438e-836e-6ee1bd1fd78b · inbound

Multi-Agent Transformer for Queue-Level XR Traffic Scheduling in TSN Networks cites this paper.

Multi-Agent Transformer for Queue-Level XR Traffic Scheduling in TSN Networks Trust Region Policy Optimisation in Multi-Agent Reinforcement Learning

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-08T15:03:59.487746Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T15:03:59.487746Z digest=sha256:c0532e8be93a3ba036a4da70f742f2e06bc69ec94414d43f7e97879749546148

Observation 2218156e-1c60-4456-9cef-55a246be9aff · inbound

Multi-Agent Reinforcement Learning for Online Traffic Scheduling in Time-Sensitive Application cites this paper.

Multi-Agent Reinforcement Learning for Online Traffic Scheduling in Time-Sensitive Application Trust Region Policy Optimisation in Multi-Agent Reinforcement Learning

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-08T14:55:44.155987Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T14:55:44.155987Z digest=sha256:0920a38c3e06370f4fefcbf7de3c6f090c2c7232dbce809a3281dae7b955d5c7