Pith. sign in

Paper Citation Record · LEDGER

Trust Region Policy Optimisation in Multi-Agent Reinforcement Learning

As of 8 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 17 inbound Pith citation observations for arXiv:2109.11251.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2109.11251 v2

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 17 of 17 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-08T06:32:00.761636+00:00

measured 17 of 17 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-08T18:50:49.824318Z

measured 1 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-08-05T02:28:24.338817Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

83
arxiv_reference, observed 2026-08-05T02:28:24.338817Z

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation 364f7221-7d8a-4d5f-8aa8-8232c2eda30f · inbound

Low-Rank Agent-Specific Adaptation (LoRASA) for Multi-Agent Policy Learning cites this paper.

Low-Rank Agent-Specific Adaptation (LoRASA) for Multi-Agent Policy Learning Trust Region Policy Optimisation in Multi-Agent Reinforcement Learning

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-08T18:50:49.824318Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T18:50:49.824318Z digest=sha256:b4df908640c38cda2aec89f72dcf02f47885d220754e97d9d72215c4515915f6

Observation 1cff9fb9-eb41-4802-a372-4046bfc8210e · inbound

Light Aircraft Game : Basic Implementation and training results analysis cites this paper.

Light Aircraft Game : Basic Implementation and training results analysis Trust Region Policy Optimisation in Multi-Agent Reinforcement Learning

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-07T00:23:07.654898Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:23:07.654898Z digest=sha256:dbde5e4a4856a63fdae36663317c2a2827c572e534d8e0f3f7fb33cc62690838

Observation b9f7b077-a641-4f1e-b536-eba274abbe39 · inbound

Improving monotonic optimization in heterogeneous multi-agent reinforcement learning with optimal marginal deterministic policy gradient cites this paper.

Improving monotonic optimization in heterogeneous multi-agent reinforcement learning with optimal marginal deterministic policy gradient Trust Region Policy Optimisation in Multi-Agent Reinforcement Learning

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-06T17:47:37.242154Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T17:47:37.242154Z digest=sha256:67fa610064ac1bf003b40914f5c58f88938f7b3a052159b4b719913778f968cc

Observation ecd55247-3864-45aa-92d3-e5006119891a · inbound

Centralized Adaptive Sampling for Reliable Co-Training of Independent Multi-Agent Policies cites this paper.

Centralized Adaptive Sampling for Reliable Co-Training of Independent Multi-Agent Policies Trust Region Policy Optimisation in Multi-Agent Reinforcement Learning

Reference 47

Resolution
metadata mismatch
arxiv_id, observed 2026-05-19T01:06:57.409927Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-19T01:05:17.086399Z digest=sha256:22e206a97fbd18ff25277d94a6c9ab813b07053a0cef530da61b7b0cb712441f

Observation 924df6fc-5af8-4b3f-8a73-96966cc68a29 · inbound

Multi-Agent Trust Region Policy Optimisation: A Joint Constraint Approach cites this paper.

Multi-Agent Trust Region Policy Optimisation: A Joint Constraint Approach Trust Region Policy Optimisation in Multi-Agent Reinforcement Learning

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-05T20:31:02.434225Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T20:31:02.434225Z digest=sha256:357688c4c8399f9c6baeb3e15e1180bf999aeb36fcdf3a45a1c64d1e8672d291

Observation ce1cbd52-a50d-43df-9d3c-f212f995948f · inbound

TABX: A High-Throughput Sandbox Battle Simulator for Multi-Agent Reinforcement Learning cites this paper.

TABX: A High-Throughput Sandbox Battle Simulator for Multi-Agent Reinforcement Learning Trust Region Policy Optimisation in Multi-Agent Reinforcement Learning

Reference 9

Resolution
verified exact
arxiv_id, observed 2026-05-25T07:45:29.793674Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-25T07:41:48.634351Z digest=sha256:a096fe9b9e4ff04a13aaef2e056206f3009f6751746b34527f4d489315d470c6

Observation acc98b05-d732-4449-9d73-9c2f0350a4cc · inbound

TABX: A High-Throughput Sandbox Battle Simulator for Multi-Agent Reinforcement Learning cites this paper.

TABX: A High-Throughput Sandbox Battle Simulator for Multi-Agent Reinforcement Learning Trust Region Policy Optimisation in Multi-Agent Reinforcement Learning

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-03T05:37:38.440274Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T05:37:38.440274Z digest=sha256:55991b9b9526d7a9a1b7ae957964439b3bc8d4494fd4f8c9b02b575ef7021a24

Observation 1d29723a-b446-4a5e-9b8e-cae77fc18570 · inbound

Topology-Driven Anti-Entanglement Control for Soft Robots cites this paper.

Topology-Driven Anti-Entanglement Control for Soft Robots Trust Region Policy Optimisation in Multi-Agent Reinforcement Learning

Reference 12

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T15:41:28.058511Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-09T19:28:39.040470Z digest=sha256:fc30952b3a268d0d7413c8f54ee17813d735cd9ea25b07b3ad49ef2804d54051

Observation b778fc2f-accf-435a-8f96-be38f0240d6c · inbound

Robust Instruction Compliance in Cooperative Multi-Agent Reinforcement Learning cites this paper.

Robust Instruction Compliance in Cooperative Multi-Agent Reinforcement Learning Trust Region Policy Optimisation in Multi-Agent Reinforcement Learning

Reference 122

Resolution
verified exact
arxiv_id, observed 2026-06-30T22:15:05.903774Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-06-30T22:11:35.277901Z digest=sha256:76717fc796c39d565a03a01576ea2e1d2b60ab4014f23df8f331495cf5faf49e

Observation 401883ea-8851-4bfa-9b3c-1bfcc7509581 · inbound

Self-Supervised On-Policy Reinforcement Learning via Contrastive Proximal Policy Optimisation cites this paper.

Self-Supervised On-Policy Reinforcement Learning via Contrastive Proximal Policy Optimisation Trust Region Policy Optimisation in Multi-Agent Reinforcement Learning

Reference 10

Resolution
verified exact
arxiv_id, observed 2026-05-14T20:12:54.902012Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-14T20:12:07.920183Z digest=sha256:481b19edceea95d9f15e6f9aed6b6e3b27c162bfc0314e6ca0db661e3f7331c8

Observation 1a6b8c58-8480-4861-9f48-66da0251baad · inbound

TeamTR: Trust-Region Fine-Tuning for Multi-Agent LLM Coordination cites this paper.

TeamTR: Trust-Region Fine-Tuning for Multi-Agent LLM Coordination Trust Region Policy Optimisation in Multi-Agent Reinforcement Learning

Reference 60

Resolution
metadata mismatch
arxiv_id, observed 2026-05-19T18:02:42.323932Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-05-19T18:01:06.649723Z digest=sha256:323bb2353dab7f606fc2dc58440d23a3f8ff2e148017fb8a9b80293cbed9e270

Observation 113eb189-443b-40bf-9fcc-d7de7a1804dc · inbound

Phi-Actor-Critic: Steering General-Sum Games to Pareto-Efficient Correlated Equilibria cites this paper.

Phi-Actor-Critic: Steering General-Sum Games to Pareto-Efficient Correlated Equilibria Trust Region Policy Optimisation in Multi-Agent Reinforcement Learning

Reference 40

Resolution
metadata mismatch
arxiv_id, observed 2026-07-03T08:47:50.453357Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-06-27T10:42:17.852996Z digest=sha256:c00867ed6879226b2810de1c26dca666dd46b8202723e322e88994d28a74bff1

Observation bbbcca0b-11cc-4c9c-93b3-d697a1a39799 · inbound

Reference-Free Heterogeneous Multi-Agent Reinforcement Learning for Grid-Friendly Tie-Line Power Shaping in Industrial Microgrids cites this paper.

Reference-Free Heterogeneous Multi-Agent Reinforcement Learning for Grid-Friendly Tie-Line Power Shaping in Industrial Microgrids Trust Region Policy Optimisation in Multi-Agent Reinforcement Learning

Reference 34

Resolution
metadata mismatch
arxiv_id, observed 2026-07-04T20:50:11.104720Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-06-25T19:36:51.509710Z digest=sha256:825b00349fefd77f12e09163cf72341453807d71d7fa0d067d8664a16683f7d4

Observation 235d54c4-9e76-462f-9e21-80b4af10140d · inbound

Action-Factored Multi-Agent Reinforcement Learning for Scalable Quantum Device Tuning cites this paper.

Action-Factored Multi-Agent Reinforcement Learning for Scalable Quantum Device Tuning Trust Region Policy Optimisation in Multi-Agent Reinforcement Learning

Reference 63

Resolution
unresolved
no resolver link, observed 2026-07-13T03:08:01.590659Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-13T03:08:01.590659Z digest=sha256:4eca52da32d356ebe6508e28ffad7f9baf2ca6d3f22dec0959c3e15d7b05ce11

Observation 8d46f04a-1a2d-4c1b-87ba-8bdd023b8cd3 · inbound

Heterogeneous Multi-Agent Reinforcement Learning for Radio Resource Management under Coupled Finite-Horizon Constraints cites this paper.

Heterogeneous Multi-Agent Reinforcement Learning for Radio Resource Management under Coupled Finite-Horizon Constraints Trust Region Policy Optimisation in Multi-Agent Reinforcement Learning

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-04T21:47:09.010614Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T21:47:09.010614Z digest=sha256:6c0cc5bffa1e916885443ae820ff32df88e086203d5a8f38f0d37186f72eed9e

Observation 748a4a37-ff95-438e-836e-6ee1bd1fd78b · inbound

Multi-Agent Transformer for Queue-Level XR Traffic Scheduling in TSN Networks cites this paper.

Multi-Agent Transformer for Queue-Level XR Traffic Scheduling in TSN Networks Trust Region Policy Optimisation in Multi-Agent Reinforcement Learning

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-08T15:03:59.487746Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T15:03:59.487746Z digest=sha256:39f70cd0b3db7a99fe9e68f82f2ee95bcc570709a0999b1134a2c5106d728ec1

Observation 2218156e-1c60-4456-9cef-55a246be9aff · inbound

Multi-Agent Reinforcement Learning for Online Traffic Scheduling in Time-Sensitive Application cites this paper.

Multi-Agent Reinforcement Learning for Online Traffic Scheduling in Time-Sensitive Application Trust Region Policy Optimisation in Multi-Agent Reinforcement Learning

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-08T14:55:44.155987Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T14:55:44.155987Z digest=sha256:7c3623814b4eb4ca001e87762987531b5915563fbca4ab799d8cfc77ded0e08c