Pith. sign in

Paper Citation Record · LEDGER

Monte Carlo Beam Search for Actor-Critic Reinforcement Learning in Continuous Control

As of 22 August 2026, this Paper Citation Record lists 24 of 24 outbound references and 0 inbound Pith citation observations for arXiv:2505.09029.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2505.09029 v1

Coverage vector

measured 24 of 24 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-15T21:45:35.150536Z

measured 24 of 24 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-22T06:32:14.747728+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

24 of 24 outbound references displayed

  • verified exact0
  • verified fuzzy14
  • unresolved10
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation b214486c-e58f-4253-8f3b-f0f7564568bf · outbound

This paper cites ”Reinforcement learning: An introduction.

Monte Carlo Beam Search for Actor-Critic Reinforcement Learning in Continuous Control ”Reinforcement learning: An introduction

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T21:45:35.547841Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-15T21:45:35.040358Z digest=sha256:2d5c64c3ff6b6da79beb001c654b78a65e1e4a80dc19b5cc0f37c6a4d817807e

Observation cae4ee93-2c40-4817-881e-53a88a3a888f · outbound

This paper cites ”Reinforcement learning algorithms: A brief survey.” Expert Systems with Applications 231 (2023): 120495.

Monte Carlo Beam Search for Actor-Critic Reinforcement Learning in Continuous Control ”Reinforcement learning algorithms: A brief survey.” Expert Systems with Applications 231 (2023): 120495

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T21:45:35.533228Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-15T21:45:35.045488Z digest=sha256:d1f2ff82421e1c16b8beaa204e398b551d10e76a45c8db5cdb82e8f664024f5a

Observation 9e7cba20-a200-4130-adee-723c7eba24cd · outbound

This paper cites ”Using reinforcement learning for load testing of video games.” Proceedings of the 44th international conference on software engineering.

Monte Carlo Beam Search for Actor-Critic Reinforcement Learning in Continuous Control ”Using reinforcement learning for load testing of video games.” Proceedings of the 44th international conference on software engineering

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T21:45:35.517836Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-15T21:45:35.050151Z digest=sha256:fbb94b3cb053a106f3ddab3c75d22622a9c7ef987054e6699371e912f9477cec

Observation 6c3251f6-180d-4827-8cb1-ca76e7cd2591 · outbound

This paper cites ”A comprehensive survey of research towards AI-enabled unmanned aerial systems in pre-, active-, and post-wildfire management.” Information Fusion (2024): 102369.

Monte Carlo Beam Search for Actor-Critic Reinforcement Learning in Continuous Control ”A comprehensive survey of research towards AI-enabled unmanned aerial systems in pre-, active-, and post-wildfire management.” Information Fusion (2024): 102369

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T21:45:35.503172Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-15T21:45:35.054810Z digest=sha256:5ce0a2e0aa248a39a06c9842c951fc66b02c5991e3f6589eb0595a2e5f6f9edb

Observation 30fd499a-75dc-4bcf-b16c-6153aecc0ae8 · outbound

This paper cites an unresolved cited work.

Monte Carlo Beam Search for Actor-Critic Reinforcement Learning in Continuous Control Unresolved cited work

Reference 5

Resolution
unresolved
raw_fallback, observed 2026-08-15T21:45:35.488101Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-15T21:45:35.060042Z digest=sha256:012734cf6deacdfb940a29f2ac9f187af4dc4c47c3d93f9a55de7423a070eaa6

Observation 0184d00b-5bf5-4fab-ac08-c62caccdaf01 · outbound

This paper cites MuJoCo Playground.

Monte Carlo Beam Search for Actor-Critic Reinforcement Learning in Continuous Control MuJoCo Playground

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-15T21:45:35.064770Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:45:35.064770Z digest=sha256:ce5e9a575297e2a8ef002a1b52a907fc192e3da02b4ae27933820893e7e55f9f

Observation af29866b-345d-4b86-af39-4cf23f056e9f · outbound

This paper cites ”Continuous control actions learning and adaptation for robotic manipulation through reinforcement learning.” Autonomous Robots 46.3 (2022): 483-498.

Monte Carlo Beam Search for Actor-Critic Reinforcement Learning in Continuous Control ”Continuous control actions learning and adaptation for robotic manipulation through reinforcement learning.” Autonomous Robots 46.3 (2022): 483-498

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T21:45:35.472673Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-15T21:45:35.070259Z digest=sha256:0f2adafee53aaf470d123df1d767261d7898d22dc9f798bcc7c453362798d796

Observation b5f35a32-d0f8-405b-92c3-84355e64a4ed · outbound

This paper cites ”Deep deterministic policy gradient algorithm: A systematic review.” Heliyon (2024).

Monte Carlo Beam Search for Actor-Critic Reinforcement Learning in Continuous Control ”Deep deterministic policy gradient algorithm: A systematic review.” Heliyon (2024)

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T21:45:35.455950Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-15T21:45:35.074744Z digest=sha256:fd2525b45bb341ed265b35b1cf6810541075b1d0130404767595f956bec89995

Observation 503de0fe-97cc-423e-8862-ca89dfbd3465 · outbound

This paper cites ”Addressing function approximation error in actor-critic methods.” International conference on machine learning.

Monte Carlo Beam Search for Actor-Critic Reinforcement Learning in Continuous Control ”Addressing function approximation error in actor-critic methods.” International conference on machine learning

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T21:45:35.439305Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-15T21:45:35.079309Z digest=sha256:d931061dee4a1e4f0713baed6bc9464584d3ceede9a665c11a78fea69ab08dd0

Observation 2020f5a2-5dac-4227-b83e-ced049cbbba8 · outbound

This paper cites ”Stable-baselines3: Reliable reinforcement learning implementations.” Journal of machine learning research 22.268 (2021): 1-8.

Monte Carlo Beam Search for Actor-Critic Reinforcement Learning in Continuous Control ”Stable-baselines3: Reliable reinforcement learning implementations.” Journal of machine learning research 22.268 (2021): 1-8

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T21:45:35.421564Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-15T21:45:35.083839Z digest=sha256:0b114b9147f1603eb52bf312a5073c92ca03e5bfe80872335eda9ed1c6e72ec8

Observation c9a48528-1cfa-4590-a1cd-9d4ee9ef4617 · outbound

This paper cites Eyes on the Environment: AI-Driven Analysis for Fire and Smoke Classification, Segmentation, and Detection.

Monte Carlo Beam Search for Actor-Critic Reinforcement Learning in Continuous Control Eyes on the Environment: AI-Driven Analysis for Fire and Smoke Classification, Segmentation, and Detection

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-15T21:45:35.088192Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:45:35.088192Z digest=sha256:93ef47a35b9b1650d5d563d869bd6d0cdabd529d379bd702e8c82972eead68c4

Observation 0ce57820-9eae-4921-ad16-ec3d785ba5c2 · outbound

This paper cites Deep Reinforcement Learning Hands-On: Apply modern RL methods, with deep Q-networks, value iteration, policy gradients, TRPO, AlphaGo Zero and more.

Monte Carlo Beam Search for Actor-Critic Reinforcement Learning in Continuous Control Deep Reinforcement Learning Hands-On: Apply modern RL methods, with deep Q-networks, value iteration, policy gradients, TRPO, AlphaGo Zero and more

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T21:45:35.404273Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-15T21:45:35.092822Z digest=sha256:c805b596014f0fb737e46465f7b65331826c7cc53e7d670ace98af45eef4cad4

Observation eda5f139-b497-4678-85be-b4cc802f87e9 · outbound

This paper cites ”AlphaZero.” Deep Reinforce- ment Learning: Fundamentals, Research and Applications (2020): 391- 415.

Monte Carlo Beam Search for Actor-Critic Reinforcement Learning in Continuous Control ”AlphaZero.” Deep Reinforce- ment Learning: Fundamentals, Research and Applications (2020): 391- 415

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T21:45:35.387544Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-15T21:45:35.097469Z digest=sha256:883f5e3c81b2b0778134cb3676e216af5cf52c886686d88fec71d8fef4be2f71

Observation 866b4b9a-3db6-4e01-aa06-335aa0eef24b · outbound

This paper cites Automatic Prompt Optimization with "Gradient Descent" and Beam Search.

Monte Carlo Beam Search for Actor-Critic Reinforcement Learning in Continuous Control Automatic Prompt Optimization with "Gradient Descent" and Beam Search

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-15T21:45:35.101965Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:45:35.101965Z digest=sha256:bfb1648626ba18acbffba17c10f07ab81bff58c210aa33f496b48bc54ceec68d

Observation c8ccdd19-8117-4251-94a3-3953d7dbfa9b · outbound

This paper cites ”Simulation-guided beam search for neural com- binatorial optimization.” Advances in Neural Information Processing Systems 35 (2022): 8760-8772.

Monte Carlo Beam Search for Actor-Critic Reinforcement Learning in Continuous Control ”Simulation-guided beam search for neural com- binatorial optimization.” Advances in Neural Information Processing Systems 35 (2022): 8760-8772

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T21:45:35.371794Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-15T21:45:35.106933Z digest=sha256:248f2f4afe9ef55475f1293fdb3a7aa1dafff80c6de90f9f63198c565fa0e98a

Observation 43d36889-2ea1-4806-b25b-77e502231a41 · outbound

This paper cites ”Monte Carlo tree search: A review of recent modifications and applications.” Artificial Intelligence Review 56.3 (2023): 2497-2562.

Monte Carlo Beam Search for Actor-Critic Reinforcement Learning in Continuous Control ”Monte Carlo tree search: A review of recent modifications and applications.” Artificial Intelligence Review 56.3 (2023): 2497-2562

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T21:45:35.356588Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-15T21:45:35.111345Z digest=sha256:2239237a575f7f999f7b7312e4604c895dec6fa7890115322aea33f273226305

Observation ae813dbc-8b2a-44ac-b6b5-de339479dfe9 · outbound

This paper cites Monte Carlo Tree Search Boosts Reasoning via Iterative Preference Learning.

Monte Carlo Beam Search for Actor-Critic Reinforcement Learning in Continuous Control Monte Carlo Tree Search Boosts Reasoning via Iterative Preference Learning

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-15T21:45:35.115772Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:45:35.115772Z digest=sha256:52c6b240ff6dc96345ca157daf4849c6e97020115a873bbf28054f65e3f8c0ed

Observation 0315eb13-f23d-43cc-b771-cc935a2af31d · outbound

This paper cites ”Beyond greedy search: Tracking by multi-agent reinforcement learning-based beam search.” IEEE Transactions on Image Processing 31 (2022): 6239-6254.

Monte Carlo Beam Search for Actor-Critic Reinforcement Learning in Continuous Control ”Beyond greedy search: Tracking by multi-agent reinforcement learning-based beam search.” IEEE Transactions on Image Processing 31 (2022): 6239-6254

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T21:45:35.341152Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-15T21:45:35.120482Z digest=sha256:8b673c47de778ba114e979cec2c26cc23f777050784751905c10d034825fb8f6

Observation b68f52c9-d7ec-4cfe-9954-f33d97a001dd · outbound

This paper cites ”RL Baselines3 Zoo.” GitHub repository (2020).

Monte Carlo Beam Search for Actor-Critic Reinforcement Learning in Continuous Control ”RL Baselines3 Zoo.” GitHub repository (2020)

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T21:45:35.323678Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-15T21:45:35.125457Z digest=sha256:74b902c85005b3da24624a3e108c623f5db1e067b610057ef900172d8abdc314

Observation d34cd3fe-9512-440a-9675-026d56de3c96 · outbound

This paper cites OpenAI Gym.

Monte Carlo Beam Search for Actor-Critic Reinforcement Learning in Continuous Control OpenAI Gym

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-15T21:45:35.130093Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:45:35.130093Z digest=sha256:23b34c2ca65c7c683e8d5bbd1b102e9c7a79aac2ec9a6940919ceef80e3be685

Observation a7f1d474-7272-4cd3-8e9b-0d075b2ae5aa · outbound

This paper cites Soft Actor-Critic Algorithms and Applications.

Monte Carlo Beam Search for Actor-Critic Reinforcement Learning in Continuous Control Soft Actor-Critic Algorithms and Applications

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-15T21:45:35.134938Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:45:35.134938Z digest=sha256:f53b3f4a396ccbaa75a22df24cf355c2a6c9b979c5e0ce7b0e6f204e93f5e1bf

Observation 845e944e-000f-4355-ab96-3c7494df062e · outbound

This paper cites A2C is a special case of PPO.

Monte Carlo Beam Search for Actor-Critic Reinforcement Learning in Continuous Control A2C is a special case of PPO

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-15T21:45:35.140853Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:45:35.140853Z digest=sha256:80302c26009452271f58e1f1659c0d3426cf16e50620f33e08edac19a22fa1b3

Observation 6dbc1c6b-53ec-47bd-be34-3cf5cfb4fa7a · outbound

This paper cites Proximal Policy Optimization Algorithms.

Monte Carlo Beam Search for Actor-Critic Reinforcement Learning in Continuous Control Proximal Policy Optimization Algorithms

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-15T21:45:35.146067Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:45:35.146067Z digest=sha256:4080a55a42fd82cab70710cd78a612061dc45f490cadf32a0be2c4308648bdc1

Observation ff361cf6-bfb1-4f83-8c13-456afcea3f27 · outbound

This paper cites VisionGPT: LLM-Assisted Real-Time Anomaly Detection for Safe Visual Navigation.

Monte Carlo Beam Search for Actor-Critic Reinforcement Learning in Continuous Control VisionGPT: LLM-Assisted Real-Time Anomaly Detection for Safe Visual Navigation

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-15T21:45:35.150536Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:45:35.150536Z digest=sha256:0830c0e5ed7bd07d75fa8b89338bf0c8f03bbfe7c921100b420e5de6b6a8c866

Pith citing papers

No inbound Pith citation observations are available.