Pith. sign in

Paper Citation Record · LEDGER

Informed Asymmetric Actor-Critic: Leveraging Privileged Signals Beyond Full-State Access

As of 17 August 2026, this Paper Citation Record lists 33 of 33 outbound references and 0 inbound Pith citation observations for arXiv:2509.26000.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2509.26000 v3

Coverage vector

measured 33 of 33 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-04T13:43:32.011996Z

measured 33 of 33 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-17T06:30:58.91139+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

33 of 33 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved32
  • parse uncertain0
  • malformed identifier1
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation d1895590-a8ae-41da-a642-52a358684f77 · outbound

This paper cites Reinforcement learning for HVAC control in intelligent buildings: A technical and conceptual review.

Informed Asymmetric Actor-Critic: Leveraging Privileged Signals Beyond Full-State Access Reinforcement learning for HVAC control in intelligent buildings: A technical and conceptual review

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-04T13:43:28.830256Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T13:43:28.830256Z digest=sha256:ab28b39301be7e9931ab6ba578461cd2892c15c7fbb462d65420c62882c8cf2d

Observation cce86ea8-b16e-4a46-ac6c-6a0c3c6cb978 · outbound

This paper cites MicroPPO: Safe power flow management in decentralized micro-grids with proximal policy optimization.

Informed Asymmetric Actor-Critic: Leveraging Privileged Signals Beyond Full-State Access MicroPPO: Safe power flow management in decentralized micro-grids with proximal policy optimization

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-04T13:43:28.873203Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T13:43:28.873203Z digest=sha256:71a0cae6f3a1d8232eb7574665d0c42394d096f59549f5b429074c76c4ee4d9f

Observation 1bb75d7a-ee11-48a0-a78e-4610715bd47f · outbound

This paper cites Deep reinforcement learning solutions for energy microgrids management.

Informed Asymmetric Actor-Critic: Leveraging Privileged Signals Beyond Full-State Access Deep reinforcement learning solutions for energy microgrids management

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-04T13:43:28.949222Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T13:43:28.949222Z digest=sha256:09fe5769ccbfc7ad936e66f9213e8bf612db75d608a5af6e6053c60ae141933a

Observation 52133c76-122a-4296-a6ce-e07a8dc78929 · outbound

This paper cites Deep reinforcement learning framework for autonomous driving.Electronic Imaging, 2017:70–76, 01 2017.

Informed Asymmetric Actor-Critic: Leveraging Privileged Signals Beyond Full-State Access Deep reinforcement learning framework for autonomous driving.Electronic Imaging, 2017:70–76, 01 2017

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-04T13:43:29.014886Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T13:43:29.014886Z digest=sha256:aef0a35e4fa78be9b01d7355150392894f60b854823042b4bcdf0f3fbb5ee73c

Observation 9718af9b-a604-44be-b832-bafe48757bc4 · outbound

This paper cites Deep reinforcement learning for robotics: A survey of real-world successes.

Informed Asymmetric Actor-Critic: Leveraging Privileged Signals Beyond Full-State Access Deep reinforcement learning for robotics: A survey of real-world successes

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-04T13:43:29.134741Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T13:43:29.134741Z digest=sha256:9ab8c002c043863ffb27d5934159cbaf6b6dfe9a29d216f2d752eddd171f6bad

Observation 97e8ab1a-4cbd-431d-a0b1-0ef4da00ab7b · outbound

This paper cites Planning and acting in partially observable stochastic domains.Artificial intelligence, 101(1-2):99–134, 1998.

Informed Asymmetric Actor-Critic: Leveraging Privileged Signals Beyond Full-State Access Planning and acting in partially observable stochastic domains.Artificial intelligence, 101(1-2):99–134, 1998

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-04T13:43:29.253019Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T13:43:29.253019Z digest=sha256:b1172f3da2b806b8ed5656c10f789fdc032d47e2274bb715725a8fafefb9fe7b

Observation 6cf44693-a62c-464e-b551-f53110d02e18 · outbound

This paper cites Deep recurrent Q-learning for partially observable MDPs.

Informed Asymmetric Actor-Critic: Leveraging Privileged Signals Beyond Full-State Access Deep recurrent Q-learning for partially observable MDPs

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-04T13:43:29.320270Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T13:43:29.320270Z digest=sha256:7565a2787803885cbf537b651f98419f0a6264a53c8640168f03f8d43f0ab449

Observation e9029117-8f2c-4fc5-8ab4-e44ec99a76fc · outbound

This paper cites Learning deep neural network policies with continuous memory states.

Informed Asymmetric Actor-Critic: Leveraging Privileged Signals Beyond Full-State Access Learning deep neural network policies with continuous memory states

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-04T13:43:29.457909Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T13:43:29.457909Z digest=sha256:bb6bd754c98961d581829555194c965f784b9b795721b928720099629f31edf6

Observation d87faf64-147a-4028-b1cf-33390191b749 · outbound

This paper cites Asymmetric Actor Critic for Image-Based Robot Learning.

Informed Asymmetric Actor-Critic: Leveraging Privileged Signals Beyond Full-State Access Asymmetric Actor Critic for Image-Based Robot Learning

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-04T13:43:29.538519Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T13:43:29.538519Z digest=sha256:aee81265976509e243eb36f819654e6df1299629937948405522cc85762272f6

Observation 89f14fee-3b2c-4d77-bcdd-3934d7075e4e · outbound

This paper cites Unbiased asymmetric reinforcement learning under partial observability.

Informed Asymmetric Actor-Critic: Leveraging Privileged Signals Beyond Full-State Access Unbiased asymmetric reinforcement learning under partial observability

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-04T13:43:29.624030Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T13:43:29.624030Z digest=sha256:afb6837f18b3ff853a9175fa9a592049f4d3420dbfe46c7f4d3a0f44bb459227

Observation 4473a43d-2153-4e3b-acc8-4ff7754032d2 · outbound

This paper cites On Improving Deep Reinforcement Learning for POMDPs.

Informed Asymmetric Actor-Critic: Leveraging Privileged Signals Beyond Full-State Access On Improving Deep Reinforcement Learning for POMDPs

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-04T13:43:29.684346Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T13:43:29.684346Z digest=sha256:6f1c0a30db1c8ce83a547e6b71291d7b1a6da5c6b582505dd339782213c81d1d

Observation 39256772-0beb-4366-8859-be6985cac341 · outbound

This paper cites Recurrent policy gradients.

Informed Asymmetric Actor-Critic: Leveraging Privileged Signals Beyond Full-State Access Recurrent policy gradients

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-04T13:43:29.771032Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T13:43:29.771032Z digest=sha256:962e131467e0b22c78f64fc669f211255e6ee47684b4c502dd64fd79506c3b13

Observation 68af1ac4-061f-427c-a9f2-2485903e3fda · outbound

This paper cites Reinforcement learning with long short-term memory.

Informed Asymmetric Actor-Critic: Leveraging Privileged Signals Beyond Full-State Access Reinforcement learning with long short-term memory

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-04T13:43:29.924982Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T13:43:29.924982Z digest=sha256:8b89a3e66fcfff2c8a3aa6b360a6cc541774d1e33b198accc4ef1205b65f8e8e

Observation 9e44a7f0-07de-4c95-b98c-eb67bccf8d3d · outbound

This paper cites Recurrent natural policy gradient for POMDPs.

Informed Asymmetric Actor-Critic: Leveraging Privileged Signals Beyond Full-State Access Recurrent natural policy gradient for POMDPs

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-04T13:43:29.996009Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T13:43:29.996009Z digest=sha256:fd9bcc3e7969f4390f5a50c7edcec2854ee918fd92f750a7cfaffa555f91a342

Observation e961f44d-52cf-4de8-99ab-6c98522680bb · outbound

This paper cites Bridging State and History Representations: Understanding Self-Predictive RL.

Informed Asymmetric Actor-Critic: Leveraging Privileged Signals Beyond Full-State Access Bridging State and History Representations: Understanding Self-Predictive RL

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-04T13:43:30.088955Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T13:43:30.088955Z digest=sha256:dd8dc34730cd43a9ac2fb070f12202709df5b014167c4c735ea7b6178198b827

Observation 13f3e427-d2ce-4bff-a5a2-6c6914a49a09 · outbound

This paper cites Approximate information state for approximate planning and reinforcement learning in partially observed systems.

Informed Asymmetric Actor-Critic: Leveraging Privileged Signals Beyond Full-State Access Approximate information state for approximate planning and reinforcement learning in partially observed systems

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-04T13:43:30.236286Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T13:43:30.236286Z digest=sha256:f9134b521540aab78524c225e31d9f8be40eaee690061e344a85eb1a91a656bd

Observation f78ed1fe-7a1a-4aab-87f0-3d21cb5b36a6 · outbound

This paper cites Data-driven planning via imitation learning.The International Journal of Robotics Research, 37(13-14):1632–1672, 2018.

Informed Asymmetric Actor-Critic: Leveraging Privileged Signals Beyond Full-State Access Data-driven planning via imitation learning.The International Journal of Robotics Research, 37(13-14):1632–1672, 2018

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-04T13:43:30.344287Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T13:43:30.344287Z digest=sha256:ea8682c0bfe6bc20e0801ce76a37db384dc5a0357552914db3603a3ca19a6045

Observation 83bbfdd3-b98b-4170-83d9-57e9e99b6eee · outbound

This paper cites Robust asymmetric learning in POMDPs.

Informed Asymmetric Actor-Critic: Leveraging Privileged Signals Beyond Full-State Access Robust asymmetric learning in POMDPs

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-04T13:43:30.423812Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T13:43:30.423812Z digest=sha256:9c5f564b2bcbbf5c95044a28ef6bab2bb962ce2f589d0a271011275ff27ccadc

Observation d643a3bb-2e4c-49ea-a50e-808bcfaac549 · outbound

This paper cites Informed POMDP: Leveraging additional information in model-based RL.Reinforcement Learning Journal, 2024.

Informed Asymmetric Actor-Critic: Leveraging Privileged Signals Beyond Full-State Access Informed POMDP: Leveraging additional information in model-based RL.Reinforcement Learning Journal, 2024

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-04T13:43:30.529384Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T13:43:30.529384Z digest=sha256:21ce381b2668bbc2fd12d95cd3acc1b0b2666a4761598c80fcc8fa99e4c4b931

Observation 69114514-c576-442a-b49e-0d5c1a44bd22 · outbound

This paper cites The wasserstein believer: Learning belief updates for partially observable environments through reliable latent space models.

Informed Asymmetric Actor-Critic: Leveraging Privileged Signals Beyond Full-State Access The wasserstein believer: Learning belief updates for partially observable environments through reliable latent space models

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-04T13:43:30.639403Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T13:43:30.639403Z digest=sha256:5ff620a9b9348d3a22b7ff1384be9a8d47e7d0139d786bb519c7d871bcb90122

Observation 8be63fca-ba09-4327-bf7f-7fbaeab8df18 · outbound

This paper cites Hu, James Springer, Oleh Rybkin, and Dinesh Jayaraman.

Informed Asymmetric Actor-Critic: Leveraging Privileged Signals Beyond Full-State Access Hu, James Springer, Oleh Rybkin, and Dinesh Jayaraman

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-04T13:43:30.683356Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T13:43:30.683356Z digest=sha256:6dad91b2ea54087cecab0ea88e939190537dd40e718c88febee8b892965777f7

Observation 55edacc7-bfc8-4618-829b-275b50df21ec · outbound

This paper cites Neural policy gradient methods: Global optimality and rates of convergence.

Informed Asymmetric Actor-Critic: Leveraging Privileged Signals Beyond Full-State Access Neural policy gradient methods: Global optimality and rates of convergence

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-04T13:43:30.729903Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T13:43:30.729903Z digest=sha256:48424b3fa41971555a8edefc97218f65336853bba4db3141164cbee99f551b80

Observation f0442c94-7ca7-42a4-a4ae-0edf1834d53a · outbound

This paper cites A theoretical justification for asymmetric actor-critic algorithms.

Informed Asymmetric Actor-Critic: Leveraging Privileged Signals Beyond Full-State Access A theoretical justification for asymmetric actor-critic algorithms

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-04T13:43:30.853496Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T13:43:30.853496Z digest=sha256:be585d983466c8017bd6de256bc9d9f573be4f9348d3a4caa6c1c4a23e7be5af

Observation cf3f0c19-26e5-4169-998c-7d190c0e4778 · outbound

This paper cites A hilbert space embedding for distributions.

Informed Asymmetric Actor-Critic: Leveraging Privileged Signals Beyond Full-State Access A hilbert space embedding for distributions

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-04T13:43:30.999386Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T13:43:30.999386Z digest=sha256:ca0012f1b825cc4f4169b46b995c6fdb91ba1e67a2de2fb4fd63a747b5c471fb

Observation dd101662-1b6f-4c81-b239-024beef257c0 · outbound

This paper cites Measuring statistical dependence with hilbert-schmidt norms.

Informed Asymmetric Actor-Critic: Leveraging Privileged Signals Beyond Full-State Access Measuring statistical dependence with hilbert-schmidt norms

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-04T13:43:31.088783Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T13:43:31.088783Z digest=sha256:a5f05f7432b96ca3fe9645e734b9586628f066723e5dc63e48413de841e90dfd

Observation ebde8a40-81ea-41f8-85cf-df3cd3d1741c · outbound

This paper cites A measure-theoretic approach to kernel conditional mean embeddings.

Informed Asymmetric Actor-Critic: Leveraging Privileged Signals Beyond Full-State Access A measure-theoretic approach to kernel conditional mean embeddings

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-04T13:43:31.199136Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T13:43:31.199136Z digest=sha256:ee3a658f231ad2624a43d19de9fa4430b1cccd29b0afe1c639e52690afe7c472

Observation e15a151e-9ea8-4e62-942d-c5b01d3735be · outbound

This paper cites On overfitting and asymptotic bias in batch reinforcement learning with partial observability.

Informed Asymmetric Actor-Critic: Leveraging Privileged Signals Beyond Full-State Access On overfitting and asymptotic bias in batch reinforcement learning with partial observability

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-04T13:43:31.347183Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T13:43:31.347183Z digest=sha256:d175732bc99b349c829396eeca46af15b4cf10db1306db4dcdb3116199acd013

Observation d6a1de3f-9aba-4d51-a38d-fa71277dc52e · outbound

This paper cites Solving large POMDPs using real time dynamic programming.

Informed Asymmetric Actor-Critic: Leveraging Privileged Signals Beyond Full-State Access Solving large POMDPs using real time dynamic programming

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-04T13:43:31.453617Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T13:43:31.453617Z digest=sha256:48d6418b60ffbac1e492bc37f1b7dc4d338d815f58a20249c7516cce970ecb18

Observation 21ba80b6-33b5-4c57-8a09-c92cff800ce1 · outbound

This paper cites gym-pomdps: Gym environments from POMDP files.https://github.

Informed Asymmetric Actor-Critic: Leveraging Privileged Signals Beyond Full-State Access gym-pomdps: Gym environments from POMDP files.https://github

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-04T13:43:31.539400Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T13:43:31.539400Z digest=sha256:72cb13140a210d3ba1b08e418c33aab666499da5c7545ad9de3189392210275f

Observation f264b523-9d61-4d30-8bb0-4f1527b2c5ba · outbound

This paper cites POMDP Robot Domains.

Informed Asymmetric Actor-Critic: Leveraging Privileged Signals Beyond Full-State Access POMDP Robot Domains

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-04T13:43:31.621646Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T13:43:31.621646Z digest=sha256:47e5fb498d872b7e88d3848cace7c4ab76c8da2e195f1c216f6fa7664618d494

Observation 77eeb484-d2fb-4b49-b37e-f5cf836fc87b · outbound

This paper cites Multi-agent reinforcement learning with directed exploration and selective memory reuse.

Informed Asymmetric Actor-Critic: Leveraging Privileged Signals Beyond Full-State Access Multi-agent reinforcement learning with directed exploration and selective memory reuse

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-04T13:43:31.677707Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T13:43:31.677707Z digest=sha256:e3083581f9f1bd65692ce3b1d8953bd4be27b64e10d5720f8b2b9b42d84f4df4

Observation e2adb9da-7b1f-47a2-bde3-b735d2dcdc22 · outbound

This paper cites gym-gridverse: Gridworld domains for fully and partially observable settings.https://github.com/abaisero/gym-gridverse, 2021.

Informed Asymmetric Actor-Critic: Leveraging Privileged Signals Beyond Full-State Access gym-gridverse: Gridworld domains for fully and partially observable settings.https://github.com/abaisero/gym-gridverse, 2021

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-04T13:43:31.827191Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T13:43:31.827191Z digest=sha256:bd8e6027d829d7144e6f95e73f8bb31f15e300f16cf9a46b9b0ff2b4db5de6ad

Observation 9554b710-679a-4493-8cd2-6f2e36c9363f · outbound

This paper cites asym-porl: Asymmetric methods for partially observable reinforcement learning.

Informed Asymmetric Actor-Critic: Leveraging Privileged Signals Beyond Full-State Access asym-porl: Asymmetric methods for partially observable reinforcement learning

Reference 33

Resolution
malformed identifier
no resolver link, observed 2026-08-04T13:43:32.011996Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T13:43:32.011996Z digest=sha256:fb42ae40b07b0a347dcb19d697c38887bce3ec9b0f044904752a1ee12b88acc0

Pith citing papers

No inbound Pith citation observations are available.