Pith. sign in

Paper Citation Record · LEDGER

Convergence of regularized agent-state-based Q-learning in POMDPs

As of 24 August 2026, this Paper Citation Record lists 47 of 47 outbound references and 0 inbound Pith citation observations for arXiv:2508.21314.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2508.21314 v2

Coverage vector

measured 47 of 47 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-05T14:31:38.795970Z

measured 47 of 47 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-23T06:30:58.430688+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

47 of 47 outbound references displayed

  • verified exact0
  • verified fuzzy41
  • unresolved6
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 33b9f85e-83ba-409e-91d2-23c77c4164d8 · outbound

This paper cites Optimal control of Markov processes with incomplete state information I,.

Convergence of regularized agent-state-based Q-learning in POMDPs Optimal control of Markov processes with incomplete state information I,

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T14:31:47.937314Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-05T14:31:34.726619Z digest=sha256:23dcb8fc121666c680eb9644968042ee5cdb021a8c4991fa271dd1c82352d751

Observation dbbb163e-c714-4370-a962-f4a5485fde5b · outbound

This paper cites The optimal control of partially observable Markov processes over a finite horizon,.

Convergence of regularized agent-state-based Q-learning in POMDPs The optimal control of partially observable Markov processes over a finite horizon,

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T14:31:47.834699Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-05T14:31:34.806924Z digest=sha256:c324db08aa67ff378e2ba652f111da40588a1a404afdb7aa92fd1b2ac9ec3e81

Observation d77d49cb-4264-422f-b0f9-d3a53976a919 · outbound

This paper cites Approximate information state for approximate planning and reinforcement learning in partially observed systems,.

Convergence of regularized agent-state-based Q-learning in POMDPs Approximate information state for approximate planning and reinforcement learning in partially observed systems,

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T14:31:47.697682Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-05T14:31:34.892758Z digest=sha256:fe25ecbe4434d6a3ebc000d07133deb6a3b2f79a870a3ff36f5ed17d5fd930d8

Observation e999cc6a-1015-4d36-921d-d4476f78ce11 · outbound

This paper cites Deep recurrent Q-learning for partially observable MDPs.

Convergence of regularized agent-state-based Q-learning in POMDPs Deep recurrent Q-learning for partially observable MDPs

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T14:31:47.553437Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-05T14:31:34.956485Z digest=sha256:ea58ebeb7b149a5853b1112892c77eaa6cc2f32192908f171def3e535894eede

Observation 159ff1fe-54b4-4789-9574-1d09c12456cc · outbound

This paper cites Deep variational reinforcement learning for POMDPs,.

Convergence of regularized agent-state-based Q-learning in POMDPs Deep variational reinforcement learning for POMDPs,

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T14:31:47.378652Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-05T14:31:35.023211Z digest=sha256:df2fc532a135ca8c9c50ce13638b1c0a5bfc1a6e64783f0b5843caf68cccabfd

Observation df196e73-64fd-43da-aa25-4a7373d4756a · outbound

This paper cites On Improving Deep Reinforcement Learning for POMDPs.

Convergence of regularized agent-state-based Q-learning in POMDPs On Improving Deep Reinforcement Learning for POMDPs

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-05T14:31:35.087409Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T14:31:35.087409Z digest=sha256:a2c7589108395b7facbf857e037515902c6276b8d37d80fc494b4d8508a83842

Observation c3d8e18d-0db7-49a9-bafb-4eb16fa66946 · outbound

This paper cites Memory-based deep reinforcement learning for POMDPs,.

Convergence of regularized agent-state-based Q-learning in POMDPs Memory-based deep reinforcement learning for POMDPs,

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T14:31:47.181372Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-05T14:31:35.154351Z digest=sha256:8d185430cd50e65b7b2365857c3120bfe4ced49accbfb6fde2443d0e785b4626

Observation 4a2642f5-2397-4d12-acc7-5abeb1c96664 · outbound

This paper cites Simple agent, complex envi- ronment: Efficient reinforcement learning with agent states,.

Convergence of regularized agent-state-based Q-learning in POMDPs Simple agent, complex envi- ronment: Efficient reinforcement learning with agent states,

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T14:31:47.035610Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-05T14:31:35.246813Z digest=sha256:301fe065b9b0e0f0dd72542e74d01d888a154fc3d1c41ea5a7dc734df1a3dfbb

Observation c9d00860-670b-49b2-8744-f1c80c843f50 · outbound

This paper cites Agent-state based policies in POMDPs: Beyond belief-state MDPs,.

Convergence of regularized agent-state-based Q-learning in POMDPs Agent-state based policies in POMDPs: Beyond belief-state MDPs,

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T14:31:46.868522Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-05T14:31:35.313591Z digest=sha256:2122e3db9f6d234859e71a80ca467929ae50bc58de4fdecbd5fed77b69d02869

Observation c3688d29-c638-4798-aae7-069a54b5d3aa · outbound

This paper cites Reinforcement learning algorithm for partially observable Markov decision problems,.

Convergence of regularized agent-state-based Q-learning in POMDPs Reinforcement learning algorithm for partially observable Markov decision problems,

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T14:31:46.655382Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-05T14:31:35.394078Z digest=sha256:89f0df6b8935aad958c3229e9fb489208fd1a501a0c884c960fbffc7081e2893

Observation b638296f-62e1-47f6-945b-fb580ccef064 · outbound

This paper cites Convergence of finite memory Q learning for POMDPs and near optimality of learned policies under filter stability,.

Convergence of regularized agent-state-based Q-learning in POMDPs Convergence of finite memory Q learning for POMDPs and near optimality of learned policies under filter stability,

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T14:31:46.505820Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-05T14:31:35.465534Z digest=sha256:dd399462a734b0c33cfb3f9d842decf990614765cb47538e0a5bc02a3279ba5e

Observation 5f1c3ac5-5f2b-4344-9903-cbf5a49b5271 · outbound

This paper cites Periodic agent-state based Q- learning for POMDPs,.

Convergence of regularized agent-state-based Q-learning in POMDPs Periodic agent-state based Q- learning for POMDPs,

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T14:31:46.315477Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-05T14:31:35.530699Z digest=sha256:cc1335fbff08749fc4e70eaef39099dbe7a923d7d03b8a75d47d51ce90e2d028

Observation ede0e294-ec60-42b2-8ef0-664e944b95b9 · outbound

This paper cites Q-learning for stochastic control under general information structures and non-markovian environments,.

Convergence of regularized agent-state-based Q-learning in POMDPs Q-learning for stochastic control under general information structures and non-markovian environments,

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T14:31:46.092609Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-05T14:31:35.595069Z digest=sha256:6f7a9a0be2609ae9959fe070c1326874cac69dc1203b6e108c6a50f4c3c0b109

Observation 27d0ba31-a0cc-42fd-92bd-10fb265bd4f2 · outbound

This paper cites Reinforcement learning in non-Markovian environments,.

Convergence of regularized agent-state-based Q-learning in POMDPs Reinforcement learning in non-Markovian environments,

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T14:31:45.943350Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-05T14:31:35.677050Z digest=sha256:a66ffdfb9eecb8820c7b1eb372883abe429271e04b1e53b0b603b955452b7ed5

Observation 4b90a2be-c0ae-463a-94f0-7b2268e34eb4 · outbound

This paper cites On actor-critic algorithms,.

Convergence of regularized agent-state-based Q-learning in POMDPs On actor-critic algorithms,

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T14:31:45.692776Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-05T14:31:35.765001Z digest=sha256:03cc70e2c8428ea84721e1333f254856b48baa3ddd7a40616ade85c94f6afb17

Observation 54d6498f-8d3a-4649-9795-36c8460fc974 · outbound

This paper cites Proximal Policy Optimization Algorithms.

Convergence of regularized agent-state-based Q-learning in POMDPs Proximal Policy Optimization Algorithms

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-05T14:31:35.857345Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T14:31:35.857345Z digest=sha256:1af88f5bebc0bd167e985705a3e7ece7015a5ae3996d5d88d73bcaeac87a2741

Observation 6f54d08a-858c-4b4a-b90f-52c8076f127f · outbound

This paper cites Soft actor-critic: Off- policy maximum entropy deep reinforcement learning with a stochastic actor,.

Convergence of regularized agent-state-based Q-learning in POMDPs Soft actor-critic: Off- policy maximum entropy deep reinforcement learning with a stochastic actor,

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T14:31:45.541911Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-05T14:31:35.945532Z digest=sha256:46d96153e0fd53ae90a838d1c7a133278cc40883bb6afbca7e7f4f1463fa873b

Observation a5c8d233-b142-4b9c-868f-66b4c9e65581 · outbound

This paper cites Under- standing the impact of entropy on policy optimization,.

Convergence of regularized agent-state-based Q-learning in POMDPs Under- standing the impact of entropy on policy optimization,

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T14:31:45.364516Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-05T14:31:36.011272Z digest=sha256:7d7a393ddab1300939152f8da3df7cec0b35984ca412362f41a10fc2f4997b85

Observation df1c5df4-e07a-464b-9283-9e5d00a0a50c · outbound

This paper cites Deep reinforcement learning that matters,.

Convergence of regularized agent-state-based Q-learning in POMDPs Deep reinforcement learning that matters,

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T14:31:45.141508Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-05T14:31:36.109162Z digest=sha256:3d386817b1fce5fc72a920a392e77724f2e4f5dd2b2653a1f57e7ca9a79c7a87

Observation 80329930-38b3-4885-89ff-c0b9b9139e3f · outbound

This paper cites Relative entropy policy search,.

Convergence of regularized agent-state-based Q-learning in POMDPs Relative entropy policy search,

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T14:31:44.996455Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-05T14:31:36.204865Z digest=sha256:fe93d862fd042e7bebb178ab12801769be5bd823832863b25698f71785f83924

Observation 2b2236a6-2014-46c0-9a35-d4773476fc40 · outbound

This paper cites Trust region policy optimization,.

Convergence of regularized agent-state-based Q-learning in POMDPs Trust region policy optimization,

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T14:31:44.767473Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-05T14:31:36.289776Z digest=sha256:55e2605fa6e05a4212af034ae952ff1c56eab5b71a05fb666a52facf11edf889

Observation d4ace202-8d60-4877-9cce-60afb331329e · outbound

This paper cites A unified view of entropy- regularized Markov decision processes,.

Convergence of regularized agent-state-based Q-learning in POMDPs A unified view of entropy- regularized Markov decision processes,

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T14:31:44.583680Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-05T14:31:36.379160Z digest=sha256:5be0ca02fb6725842ab136ccb977f4f2b14ca168c0675e1c473454e94f7d09b8

Observation f8f17d5c-6661-4d4c-83be-9bea98b722b4 · outbound

This paper cites A theory of regularized Markov decision processes,.

Convergence of regularized agent-state-based Q-learning in POMDPs A theory of regularized Markov decision processes,

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T14:31:44.369959Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-05T14:31:36.429516Z digest=sha256:95d40885cb41e725eedd58cc3a6a41b636b966cfcd7e83acff49b16c37cfbc08

Observation e46cf0c7-e872-4fe8-a4e1-1c1fa766e711 · outbound

This paper cites Learning latent dynamics for planning from pixels,.

Convergence of regularized agent-state-based Q-learning in POMDPs Learning latent dynamics for planning from pixels,

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T14:31:44.123291Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-05T14:31:36.495999Z digest=sha256:78848d0f4ad60ef6dedca805f4420bd862c2a9685f41961ff8ef734f66a4367b

Observation 21b2282f-babd-462a-9f4d-ed48a9b90bc4 · outbound

This paper cites Solar: Deep structured representations for model-based reinforcement learning,.

Convergence of regularized agent-state-based Q-learning in POMDPs Solar: Deep structured representations for model-based reinforcement learning,

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T14:31:43.900185Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-05T14:31:36.562816Z digest=sha256:6102d2ed30107c3648c327c455172b7e66288cbd321162103c9eb88e08ec0bc2

Observation f8683189-dbd0-49f4-8ad1-d2f855b17ca7 · outbound

This paper cites Dream to control: Learning behaviors by latent imagination,.

Convergence of regularized agent-state-based Q-learning in POMDPs Dream to control: Learning behaviors by latent imagination,

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T14:31:43.687181Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-05T14:31:36.626433Z digest=sha256:1d089e9bf3a89451a429a27e4ec9fbd8f3dabf158ed051a7b39971c44666c21f

Observation 544e148c-7e2a-44be-afec-00c55961a26b · outbound

This paper cites Bridging state and history representations: Understanding self-predictive RL,.

Convergence of regularized agent-state-based Q-learning in POMDPs Bridging state and history representations: Understanding self-predictive RL,

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T14:31:43.482457Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-05T14:31:36.695055Z digest=sha256:0c298b790f1f6478830b50fe9bcb2036ee78b54c9c7b07adef3135ce3bacd655

Observation e597574b-5bf1-44e1-9e16-55694b7ce7b8 · outbound

This paper cites Entropy-regularized Point-based Value Iteration.

Convergence of regularized agent-state-based Q-learning in POMDPs Entropy-regularized Point-based Value Iteration

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-05T14:31:36.781916Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T14:31:36.781916Z digest=sha256:9a0322e9c96e8d58ae4b545db70ff61f69a59c5f12f46744662544bc9f39e449

Observation b2bf9b10-891e-4cd9-bba0-2bff1acb2561 · outbound

This paper cites DESPOT: Online POMDP planning with regularization,.

Convergence of regularized agent-state-based Q-learning in POMDPs DESPOT: Online POMDP planning with regularization,

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T14:31:43.322394Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-05T14:31:36.867927Z digest=sha256:77d6a86352215cce1257e5750057d9b8c1ae6e9250c29fc524efacf33e24d484

Observation e19dd16f-0ff1-430d-a5d8-04bc53bb48a3 · outbound

This paper cites Smoother entropy for active state trajectory estimation and obfuscation in POMDPs,.

Convergence of regularized agent-state-based Q-learning in POMDPs Smoother entropy for active state trajectory estimation and obfuscation in POMDPs,

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T14:31:43.069938Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-05T14:31:36.936524Z digest=sha256:1e8de8fd9b5b1dd108809c8ea4db42c02f49de85d020f6a0238d117e97f9b8b9

Observation a4468ba5-39a4-49b0-9778-870181cf0d07 · outbound

This paper cites The limits of pure exploration in POMDPs: When the observation entropy is enough,.

Convergence of regularized agent-state-based Q-learning in POMDPs The limits of pure exploration in POMDPs: When the observation entropy is enough,

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T14:31:42.908793Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-05T14:31:37.017812Z digest=sha256:b864d33326e58a90090fbc4fb8721f6b60d3fc244fb58b7b058f70fbdcec43fd

Observation c184507a-1e85-4b33-8a98-74a804bb044f · outbound

This paper cites an unresolved cited work.

Convergence of regularized agent-state-based Q-learning in POMDPs Unresolved cited work

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-05T14:31:37.084415Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T14:31:37.084415Z digest=sha256:d291245665512127929076b97e956bb4c5d38e19ae5b563937c2868d20ede176

Observation 18d25e0a-98c9-422f-947b-135b42902d99 · outbound

This paper cites Hiriart-Urruty and C.

Convergence of regularized agent-state-based Q-learning in POMDPs Hiriart-Urruty and C

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T14:31:42.707089Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-05T14:31:37.159941Z digest=sha256:444267bb57b585a3495b633c0593738c75f93d2dd0ec0bb81fd2a21af5f7ad8e

Observation a55d2cd3-1b4c-4363-89e4-79d9a2276489 · outbound

This paper cites Differentiable dynamic programming for structured prediction and attention,.

Convergence of regularized agent-state-based Q-learning in POMDPs Differentiable dynamic programming for structured prediction and attention,

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T14:31:42.503351Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-05T14:31:37.237763Z digest=sha256:f97d0f585fec5c1a7da2b0344a3a2a0b0b91d59cec276882ab64b63aa7bffcf0

Observation 7cfe5ea7-ef0b-483e-b378-ccb4976e26bd · outbound

This paper cites Sequential decomposition of sequential dynamic teams: applications to real-time communication and networked control sys- tems,.

Convergence of regularized agent-state-based Q-learning in POMDPs Sequential decomposition of sequential dynamic teams: applications to real-time communication and networked control sys- tems,

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T14:31:42.331616Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-05T14:31:37.323069Z digest=sha256:cbd1bcaf0e7e0e6b4e89ecc67acace80a1de571d1e753814a712ce321c984621

Observation 4de4b806-c3dd-4c6d-9232-57f89e0772f5 · outbound

This paper cites Reinforcement learning with deep energy-based policies,.

Convergence of regularized agent-state-based Q-learning in POMDPs Reinforcement learning with deep energy-based policies,

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T14:31:42.068652Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-05T14:31:37.412286Z digest=sha256:d9245eac3e8a4a7e331458551b664309e54091f9d34ea4b98f64273903e7769a

Observation 040ab679-c9b4-4c4e-a0d3-edf9b0ad9552 · outbound

This paper cites A stochastic approximation method,.

Convergence of regularized agent-state-based Q-learning in POMDPs A stochastic approximation method,

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T14:31:41.902865Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-05T14:31:37.525757Z digest=sha256:8679a6068d96c2bba37597d69c6d6970536d5a6878fe04e50198724a65bc430d

Observation 3ab540f2-df98-4d98-92f6-2610b567d596 · outbound

This paper cites Q-learning,.

Convergence of regularized agent-state-based Q-learning in POMDPs Q-learning,

Reference 38

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T14:31:41.716577Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-05T14:31:37.615856Z digest=sha256:b80872b50f4bcbb50006906beeb9582713365999c99ae4091868829b07117ee7

Observation 119f1e27-9cfe-4cb5-b36f-b3e55cf620c2 · outbound

This paper cites Asynchronous stochastic approximation and Q- learning,.

Convergence of regularized agent-state-based Q-learning in POMDPs Asynchronous stochastic approximation and Q- learning,

Reference 39

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T14:31:41.485248Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-05T14:31:37.652967Z digest=sha256:372cb4ecf319c9c1c57b694dc53276868b87ac84d8fc49544548a9236f14d98d

Observation 59d2640a-5d45-44aa-b9ff-1e4605571daf · outbound

This paper cites Learning without state- estimation in partially observable Markovian decision processes,.

Convergence of regularized agent-state-based Q-learning in POMDPs Learning without state- estimation in partially observable Markovian decision processes,

Reference 40

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T14:31:41.237200Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-05T14:31:37.681926Z digest=sha256:c2f0f6465051858a51a91b7ca6247d2d2178a4d40e29014307cb9864dbff4ec5

Observation f091b0cf-ee18-4f4a-bb48-fc5ff88244e9 · outbound

This paper cites Gradient-based algorithms for zeroth- order optimization,.

Convergence of regularized agent-state-based Q-learning in POMDPs Gradient-based algorithms for zeroth- order optimization,

Reference 41

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T14:31:40.858054Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-05T14:31:37.802366Z digest=sha256:2211cb21dd54ed8830be27ce51368f4b6ee66fcdf36d7c7808e7d89a4bb87ddb

Observation df21a8b9-1592-4b0e-a279-2b28286661cd · outbound

This paper cites X a∈A πt(a | Zt+1)Qt(Zt+1, a) − Ω(πt(· |Zt+1))− X a∈A π⋆(a | Zt+1)Qµ(Zt+1, a) + Ω(π⋆(· |Zt+1)) # (a) ≤ γ.

Convergence of regularized agent-state-based Q-learning in POMDPs X a∈A πt(a | Zt+1)Qt(Zt+1, a) − Ω(πt(· |Zt+1))− X a∈A π⋆(a | Zt+1)Qµ(Zt+1, a) + Ω(π⋆(· |Zt+1)) # (a) ≤ γ

Reference 42

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T14:31:40.652015Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-05T14:31:37.911109Z digest=sha256:8543409d0b008d2d8a1ac30323d9d233050aeedc3b92b6b358850addc7ff51a0

Observation 8d962d91-cbdd-4605-aec2-ea7843db0b2c · outbound

This paper cites an unresolved cited work.

Convergence of regularized agent-state-based Q-learning in POMDPs Unresolved cited work

Reference 43

Resolution
unresolved
raw_fallback, observed 2026-08-05T14:31:40.410414Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-05T14:31:38.053757Z digest=sha256:1ff6fe921f2d5c54dc036ea0b2cc2c054f1c4a4630cadab13e785953fc4e31be

Observation 861e7c8a-51fa-45db-8e7c-2d55f55dcccb · outbound

This paper cites future states.

Convergence of regularized agent-state-based Q-learning in POMDPs future states

Reference 44

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T14:31:40.070206Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-05T14:31:38.266308Z digest=sha256:1566d215874e6cb5f8fe552b33d110e73948da14671bec482d6b8954c8d13923

Observation c5a98ed7-d2b8-4cf5-b4cb-df55b3d781f7 · outbound

This paper cites X a∈A πℓ t (a | Zt+1)QJℓ+1K t (Zt+1, a) − Ω(πℓ t (· |Zt+1))− X a∈A πℓ,⋆(a | Zt+1)QJℓ+1K µ (Zt+1, a) + Ω(πℓ,⋆(· |Zt+1)) # (a) ≤ γ.

Convergence of regularized agent-state-based Q-learning in POMDPs X a∈A πℓ t (a | Zt+1)QJℓ+1K t (Zt+1, a) − Ω(πℓ t (· |Zt+1))− X a∈A πℓ,⋆(a | Zt+1)QJℓ+1K µ (Zt+1, a) + Ω(πℓ,⋆(· |Zt+1)) # (a) ≤ γ

Reference 45

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T14:31:39.710166Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-05T14:31:38.420084Z digest=sha256:7379f53cd5593eafd4891a7c23de594ce8bde206a3a7517f79ae466877621fed

Observation 8f606a3a-2faf-468d-add3-72e98cb0c788 · outbound

This paper cites an unresolved cited work.

Convergence of regularized agent-state-based Q-learning in POMDPs Unresolved cited work

Reference 46

Resolution
unresolved
raw_fallback, observed 2026-08-05T14:31:39.414915Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-05T14:31:38.616959Z digest=sha256:d90c96999209e53af59db1824282c845f04f16ddf7a64041623abb159b675bd9

Observation f9b99d8f-132e-4671-ad10-b449dc842323 · outbound

This paper cites These two cases must be considered separately.

Convergence of regularized agent-state-based Q-learning in POMDPs These two cases must be considered separately

Reference 47

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T14:31:39.077876Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-05T14:31:38.795970Z digest=sha256:0dd19d331c923301b3460b8969d0bc0795c8ac593ecb84315a04e4898437191a

Pith citing papers

No inbound Pith citation observations are available.