Pith. sign in

Paper Citation Record · LEDGER

Scalable Policy Optimization for Networked Multi-Agent Reinforcement Learning with Continuous State-Action Spaces

As of 17 August 2026, this Paper Citation Record lists 34 of 34 outbound references and 0 inbound Pith citation observations for arXiv:2607.18554.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2607.18554 v1

Coverage vector

measured 34 of 34 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-01T15:07:45.671697Z

measured 34 of 34 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-17T06:30:58.91139+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

34 of 34 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved34
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation db074900-9149-4f9e-9ff3-efba1aaf7f94 · outbound

This paper cites Multi-agent reinforcement learning: A selective overview of theories and algorithms,.

Scalable Policy Optimization for Networked Multi-Agent Reinforcement Learning with Continuous State-Action Spaces Multi-agent reinforcement learning: A selective overview of theories and algorithms,

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-01T15:07:44.131604Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T15:07:44.131604Z digest=sha256:758be22f59c153643ae9db096bac142b4c1aff592e7055d3c47745a28cb283e4

Observation 4d0640fd-9165-401e-a914-f4e93c376b17 · outbound

This paper cites Stability constrained reinforcement learning for decentralized real-time voltage control,.

Scalable Policy Optimization for Networked Multi-Agent Reinforcement Learning with Continuous State-Action Spaces Stability constrained reinforcement learning for decentralized real-time voltage control,

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-01T15:07:44.242509Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T15:07:44.242509Z digest=sha256:3f1301b7062c57073f8b3271642db102fa5d5c7b5ebb98137dfcd76e4d9b88e3

Observation 98d8d835-e27f-40d5-9f1e-f373ce4ec53b · outbound

This paper cites Scalable reinforcement learning of localized policies for multi-agent networked systems,.

Scalable Policy Optimization for Networked Multi-Agent Reinforcement Learning with Continuous State-Action Spaces Scalable reinforcement learning of localized policies for multi-agent networked systems,

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-01T15:07:44.336581Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T15:07:44.336581Z digest=sha256:60a0bd260f473a16b1f974ea6bc8e677a89bd307f4d14ba8affd1aec9247e511

Observation 39abe093-75f4-43d7-bcd1-2ed504277f55 · outbound

This paper cites Scalable reinforcement learning for multiagent networked sys- tems,.

Scalable Policy Optimization for Networked Multi-Agent Reinforcement Learning with Continuous State-Action Spaces Scalable reinforcement learning for multiagent networked sys- tems,

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-01T15:07:44.428470Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T15:07:44.428470Z digest=sha256:ebe07b4c25941ebf4625bca30c1990e26f47884b70532931745503b3257cd16e

Observation 9b0bebf4-cc51-4cac-8dcc-91045062ae00 · outbound

This paper cites Multi-agent reinforcement learning in stochastic networked systems,.

Scalable Policy Optimization for Networked Multi-Agent Reinforcement Learning with Continuous State-Action Spaces Multi-agent reinforcement learning in stochastic networked systems,

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-01T15:07:44.504402Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T15:07:44.504402Z digest=sha256:a792aa7c0350d43062f75a72d70ae2921a3d806999db8a3ff97bed1f97d90e57

Observation 26d11f3e-74f4-45f7-b721-4fbdb9af915f · outbound

This paper cites Global convergence of localized policy iteration in networked multi-agent reinforcement learning,.

Scalable Policy Optimization for Networked Multi-Agent Reinforcement Learning with Continuous State-Action Spaces Global convergence of localized policy iteration in networked multi-agent reinforcement learning,

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-01T15:07:44.587308Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T15:07:44.587308Z digest=sha256:122ee1c942adba105d312108978de9fde584bb25f751a97d545f4eb4dfbdff2e

Observation a5fd7b01-b3ea-470a-983e-9fb19504aaac · outbound

This paper cites Fully decentralized multi-agent reinforcement learning with networked agents,.

Scalable Policy Optimization for Networked Multi-Agent Reinforcement Learning with Continuous State-Action Spaces Fully decentralized multi-agent reinforcement learning with networked agents,

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-01T15:07:44.672839Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T15:07:44.672839Z digest=sha256:c9c6943937b17f5be499ea3125548a8858dda84f64e2a84d98f3c0e2dce4d34d

Observation 380f43a1-8d15-4e7e-a29a-f779e003e649 · outbound

This paper cites Finite-time analysis of dis- tributed td (0) with linear function approximation on multi-agent rein- forcement learning,.

Scalable Policy Optimization for Networked Multi-Agent Reinforcement Learning with Continuous State-Action Spaces Finite-time analysis of dis- tributed td (0) with linear function approximation on multi-agent rein- forcement learning,

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-01T15:07:44.749424Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T15:07:44.749424Z digest=sha256:d5a3fed5ba525552c0c1c20c55229a8167851f4802661fc14f015bbcf04a40b0

Observation 80fce664-357f-49ca-ac5e-0e9859a7e637 · outbound

This paper cites Decentralized online convex optimization in networked systems,.

Scalable Policy Optimization for Networked Multi-Agent Reinforcement Learning with Continuous State-Action Spaces Decentralized online convex optimization in networked systems,

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-01T15:07:44.816692Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T15:07:44.816692Z digest=sha256:3846fc73284ad5d0eef93f2ba0a0ed72a46a632ccdff89ece1bf11a8ebb41dd2

Observation ff360be4-3fe0-43b1-a6e6-7082e426bde6 · outbound

This paper cites Multi-agent Reinforcement Learning for Networked System Control.

Scalable Policy Optimization for Networked Multi-Agent Reinforcement Learning with Continuous State-Action Spaces Multi-agent Reinforcement Learning for Networked System Control

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-01T15:07:44.900715Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T15:07:44.900715Z digest=sha256:f1b9b6e0b2901cc2e12df196fd76872f65eb73b10a6c29ac2526d9825610b1ce

Observation ff10ef29-0db5-445e-b3dd-d0e651b15034 · outbound

This paper cites Random features for large-scale kernel machines,.

Scalable Policy Optimization for Networked Multi-Agent Reinforcement Learning with Continuous State-Action Spaces Random features for large-scale kernel machines,

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-01T15:07:44.968846Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T15:07:44.968846Z digest=sha256:6d5ceb0f7a03791c228f81246f92f7467868f11e0f406d711581ad7b6c6c21b4

Observation 45b6321f-8983-43f4-8e68-0aa60bec6e0b · outbound

This paper cites Random features for ker- nel approximation: A survey on algorithms, theory, and beyond,.

Scalable Policy Optimization for Networked Multi-Agent Reinforcement Learning with Continuous State-Action Spaces Random features for ker- nel approximation: A survey on algorithms, theory, and beyond,

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-01T15:07:45.043151Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T15:07:45.043151Z digest=sha256:22ae4b04ab755a488a3ab4990d378a0b126a4e9a9e8e07b2404fd26b4e06d4b8

Observation 0758319c-58e1-4bb2-ae17-8568332d969e · outbound

This paper cites Scalable spectral representations for multi-agent reinforcement learning in network MDPs.

Scalable Policy Optimization for Networked Multi-Agent Reinforcement Learning with Continuous State-Action Spaces Scalable spectral representations for multi-agent reinforcement learning in network MDPs

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-01T15:07:45.122994Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T15:07:45.122994Z digest=sha256:e0140f07aaa187db73cc2e3b652ddec8793b3e264f6e3271baffae62482c9d1e

Observation af3c1833-2880-4063-92fd-cfa01126defc · outbound

This paper cites Linear least-squares algorithms for temporal difference learning,.

Scalable Policy Optimization for Networked Multi-Agent Reinforcement Learning with Continuous State-Action Spaces Linear least-squares algorithms for temporal difference learning,

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-01T15:07:45.209477Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T15:07:45.209477Z digest=sha256:93bb0a5eac9bf87f53a44a8e0dfef734a8adef1bfc1905980411f4ce0dcb5852

Observation 820e9772-ffdf-4788-abfe-c477b2a7d58c · outbound

This paper cites Technical update: Least-squares temporal difference learning,.

Scalable Policy Optimization for Networked Multi-Agent Reinforcement Learning with Continuous State-Action Spaces Technical update: Least-squares temporal difference learning,

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-01T15:07:45.298257Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T15:07:45.298257Z digest=sha256:b7f31f556720c3dd7b8239839b00d2a564c0fe27ef776bdc4e0463be1a7c445f

Observation 9c84af58-9a4d-4e35-8308-b8ae758afa4d · outbound

This paper cites Finite-sample analysis of least-squares policy iteration,.

Scalable Policy Optimization for Networked Multi-Agent Reinforcement Learning with Continuous State-Action Spaces Finite-sample analysis of least-squares policy iteration,

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-01T15:07:45.391730Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T15:07:45.391730Z digest=sha256:b7d04446fee1a70314c85a9bc8c1f5741cc3dbbddc7e4c0cc1c9c61dead04319

Observation bf0b56bf-d20b-4ebe-8711-5e4887d30db3 · outbound

This paper cites A finite time analysis of temporal difference learning with linear function approximation,.

Scalable Policy Optimization for Networked Multi-Agent Reinforcement Learning with Continuous State-Action Spaces A finite time analysis of temporal difference learning with linear function approximation,

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-01T15:07:45.455637Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T15:07:45.455637Z digest=sha256:cec7cd4153560a63722e89de8e55320e95751b917b54debe856b1bfb2aef50cc

Observation fc623c16-c7c9-42c7-b88d-12c72fc7b6bf · outbound

This paper cites An introduction to matrix concentration inequalities,.

Scalable Policy Optimization for Networked Multi-Agent Reinforcement Learning with Continuous State-Action Spaces An introduction to matrix concentration inequalities,

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-01T15:07:45.493186Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T15:07:45.493186Z digest=sha256:edcd651a598e5ef42b28bd0124f716626b2b34e818d27834e098089d52b2d69b

Observation 8a738021-d68d-4b74-aa2f-2c70b18a138e · outbound

This paper cites Vershynin,High-dimensional probability: An introduction with ap- plications in data science.

Scalable Policy Optimization for Networked Multi-Agent Reinforcement Learning with Continuous State-Action Spaces Vershynin,High-dimensional probability: An introduction with ap- plications in data science

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-01T15:07:45.534618Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T15:07:45.534618Z digest=sha256:965bd36c578a9aff8a2f99155604c3ecb6519671ed0a1920c2e3c0e63cd01eb7

Observation 9912ec25-11ff-43f0-b13c-6c977f9c24e1 · outbound

This paper cites Optimum bounds for the distributions of martingales in banach spaces,.

Scalable Policy Optimization for Networked Multi-Agent Reinforcement Learning with Continuous State-Action Spaces Optimum bounds for the distributions of martingales in banach spaces,

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-01T15:07:45.609186Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T15:07:45.609186Z digest=sha256:28a3a16f3718bb7a709978468c4c1231823ecb16a746f8dfa0bcaebfa68ef3ae

Observation 02ba7837-d616-490e-971e-35b18668207b · outbound

This paper cites an unresolved cited work.

Scalable Policy Optimization for Networked Multi-Agent Reinforcement Learning with Continuous State-Action Spaces Unresolved cited work

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-01T15:07:45.625579Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T15:07:45.625579Z digest=sha256:c2fad7f5848c532d3987e6b346211c34f5e9d1ac0208185d6a726127843fed33

Observation 7dd116c8-5efc-4694-ad44-d110294d6535 · outbound

This paper cites Policy gradi- ent methods for reinforcement learning with function approximation,.

Scalable Policy Optimization for Networked Multi-Agent Reinforcement Learning with Continuous State-Action Spaces Policy gradi- ent methods for reinforcement learning with function approximation,

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-01T15:07:45.629412Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T15:07:45.629412Z digest=sha256:7fd387f2bbfca13d1de005ccf0e8f3a06ca20c7c96fa312e274177cfab478890

Observation f6d9d64c-2c40-4010-a9b0-ec8248b72543 · outbound

This paper cites Lower bounds for non-convex stochastic optimization,.

Scalable Policy Optimization for Networked Multi-Agent Reinforcement Learning with Continuous State-Action Spaces Lower bounds for non-convex stochastic optimization,

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-01T15:07:45.633092Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T15:07:45.633092Z digest=sha256:576c047c534cd00873dcfa2cfff8d2275fd355c263bbaf84e472b7009ec9b755

Observation fdf0f9db-840d-4561-9b15-0bcb01c2a9bd · outbound

This paper cites On the theory of policy gradient methods: Optimality, approximation, and distribution shift,.

Scalable Policy Optimization for Networked Multi-Agent Reinforcement Learning with Continuous State-Action Spaces On the theory of policy gradient methods: Optimality, approximation, and distribution shift,

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-01T15:07:45.636986Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T15:07:45.636986Z digest=sha256:f0c25d54608a2453560633aef34975813ddae9c25e1ff90244d8bbfa79dfb368

Observation 17da105f-f405-4446-bce6-e922b25e8719 · outbound

This paper cites Multi-agent actor-critic for mixed cooperative-competitive environ- ments,.

Scalable Policy Optimization for Networked Multi-Agent Reinforcement Learning with Continuous State-Action Spaces Multi-agent actor-critic for mixed cooperative-competitive environ- ments,

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-01T15:07:45.640729Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T15:07:45.640729Z digest=sha256:361458a6186b24e4aecdf686177565f1cf75342a4925aae8b2538372316e6aeb

Observation 1ac1c5a9-6649-41a3-8111-6a498f25c506 · outbound

This paper cites Counterfactual multi-agent policy gradients,.

Scalable Policy Optimization for Networked Multi-Agent Reinforcement Learning with Continuous State-Action Spaces Counterfactual multi-agent policy gradients,

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-01T15:07:45.644189Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T15:07:45.644189Z digest=sha256:f83d492ddf65c8563a361a6011053e7e6febc489a2d5054b5f92c0c3b805c655

Observation 473a9151-3d76-4f46-a4b3-2dbc7785c245 · outbound

This paper cites Actor-attention-critic for multi-agent reinforcement learning,.

Scalable Policy Optimization for Networked Multi-Agent Reinforcement Learning with Continuous State-Action Spaces Actor-attention-critic for multi-agent reinforcement learning,

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-01T15:07:45.647674Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T15:07:45.647674Z digest=sha256:9eddf63a7d55598f79f28a27b875e4da6296d776d2381ab9f4d8183d44136898

Observation b786be04-134f-4043-9584-1e38448ba09a · outbound

This paper cites The surprising effectiveness of ppo in cooperative multi-agent games,.

Scalable Policy Optimization for Networked Multi-Agent Reinforcement Learning with Continuous State-Action Spaces The surprising effectiveness of ppo in cooperative multi-agent games,

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-01T15:07:45.651434Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T15:07:45.651434Z digest=sha256:d4fe6855da95c5e090eb55d0976cf09243dd535ce9bc381afced0fcd607bbed7

Observation a368e43b-aa00-4bd3-bb82-6b0019cdf913 · outbound

This paper cites Near-optimal distributed linear-quadratic regulator for networked systems,.

Scalable Policy Optimization for Networked Multi-Agent Reinforcement Learning with Continuous State-Action Spaces Near-optimal distributed linear-quadratic regulator for networked systems,

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-01T15:07:45.654811Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T15:07:45.654811Z digest=sha256:b8e484ba9bce811b86fdc03eea9a207a41e16f3a5a84875eb66669caff6db1ad

Observation 11be9423-6ad9-4224-baf8-250cb9548d5b · outbound

This paper cites Network reconfiguration in distribution systems for loss reduction and load balancing,.

Scalable Policy Optimization for Networked Multi-Agent Reinforcement Learning with Continuous State-Action Spaces Network reconfiguration in distribution systems for loss reduction and load balancing,

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-01T15:07:45.658193Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T15:07:45.658193Z digest=sha256:7a7fcf213506b582c277365df7e271b7a6174256c8e29aeacfbff8f074401304

Observation 67f10cbf-2aa6-477b-9827-8c23e0c86bdc · outbound

This paper cites The description of a random field by means of conditional probabilities and conditions of its regularity,.

Scalable Policy Optimization for Networked Multi-Agent Reinforcement Learning with Continuous State-Action Spaces The description of a random field by means of conditional probabilities and conditions of its regularity,

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-01T15:07:45.661726Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T15:07:45.661726Z digest=sha256:5ae7db696a4865098423f16f106130d56015649ac8fd59340a0b2dd6365042a1

Observation cb46e0a5-6875-44b3-ba4b-4267b90c209c · outbound

This paper cites Can local particle filters beat the curse of dimensionality?.

Scalable Policy Optimization for Networked Multi-Agent Reinforcement Learning with Continuous State-Action Spaces Can local particle filters beat the curse of dimensionality?

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-01T15:07:45.665084Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T15:07:45.665084Z digest=sha256:9d0f7a661bda38da0c9b7be4239ad42ff92275bb82383f51052aceb982f110bf

Observation dde6708c-d21c-42d2-afc5-192f05407b19 · outbound

This paper cites Mini-batch stochastic approx- imation methods for nonconvex stochastic composite optimization,.

Scalable Policy Optimization for Networked Multi-Agent Reinforcement Learning with Continuous State-Action Spaces Mini-batch stochastic approx- imation methods for nonconvex stochastic composite optimization,

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-01T15:07:45.668392Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T15:07:45.668392Z digest=sha256:0dbaa13c3ee8ca58d77baf8d51fc2e954bf50388f527cabb3030eb84610b4dd2

Observation e0bcf005-760d-4027-9f54-2ae8649d016b · outbound

This paper cites Stochastic first-and zeroth-order methods for nonconvex stochastic programming,.

Scalable Policy Optimization for Networked Multi-Agent Reinforcement Learning with Continuous State-Action Spaces Stochastic first-and zeroth-order methods for nonconvex stochastic programming,

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-01T15:07:45.671697Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T15:07:45.671697Z digest=sha256:25a116940c483e37b81694ee18650c5d49a645a1e191377b97034f94c58fb368

Pith citing papers

No inbound Pith citation observations are available.