Pith. sign in

Paper Citation Record · LEDGER

Bidirectional Soft Actor-Critic: Leveraging Forward and Reverse KL Divergence for Efficient Reinforcement Learning

As of 15 August 2026, this Paper Citation Record lists 46 of 46 outbound references and 0 inbound Pith citation observations for arXiv:2506.01639.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2506.01639 v1

Coverage vector

measured 46 of 46 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-07T11:42:34.550668Z

measured 46 of 46 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-15T06:32:42.880941+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

46 of 46 outbound references displayed

  • verified exact0
  • verified fuzzy35
  • unresolved11
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 037ab0cc-4bfe-4cb9-83c5-532dc0ba7cdc · outbound

This paper cites Deriving and improving cma-es with information geometric trust regions.

Bidirectional Soft Actor-Critic: Leveraging Forward and Reverse KL Divergence for Efficient Reinforcement Learning Deriving and improving cma-es with information geometric trust regions

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:42:35.449301Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=arxiv_source observed=2026-08-07T11:42:33.434436Z digest=sha256:5d3661ad877f437fbff272df510f80627608290f7f86f290d3c17c920d0f3041

Observation e99c98b3-4569-4f1f-8ab5-30ed82433c2b · outbound

This paper cites Maximum a posteriori policy optimisation.

Bidirectional Soft Actor-Critic: Leveraging Forward and Reverse KL Divergence for Efficient Reinforcement Learning Maximum a posteriori policy optimisation

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:42:35.438011Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=arxiv_source observed=2026-08-07T11:42:33.454596Z digest=sha256:b86ebcf8a4b8e00cf59e8e7eb08c76c2ed16b1b27555d87ff4ff56b84dee6f64

Observation edc3fbf0-fb53-47d0-ad32-f868c9c1eef6 · outbound

This paper cites Maximum a Posteriori Policy Optimisation.

Bidirectional Soft Actor-Critic: Leveraging Forward and Reverse KL Divergence for Efficient Reinforcement Learning Maximum a Posteriori Policy Optimisation

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-07T11:42:33.475700Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T11:42:33.475700Z digest=sha256:33cc9f33615856bc398df4717dde36a6337023a53d8c3c5d1030eb97161235df

Observation 669f855a-fb1f-4345-8f34-9131c79842e1 · outbound

This paper cites Improved soft actor-critic: Mixing prioritized off-policy samples with on-policy experiences.

Bidirectional Soft Actor-Critic: Leveraging Forward and Reverse KL Divergence for Efficient Reinforcement Learning Improved soft actor-critic: Mixing prioritized off-policy samples with on-policy experiences

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:42:35.427993Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=arxiv_source observed=2026-08-07T11:42:33.497612Z digest=sha256:e8dd6d1f77c85c5513bef5060faaa17c5dfc3a741d381e1f01e97178e5214795

Observation 998ab124-4c61-4dbe-9271-1092e56717df · outbound

This paper cites Sukhatme.

Bidirectional Soft Actor-Critic: Leveraging Forward and Reverse KL Divergence for Efficient Reinforcement Learning Sukhatme

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:42:35.418179Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=arxiv_source observed=2026-08-07T11:42:33.519432Z digest=sha256:4a950905b73ff12fbb38201ec2ee6ca1b1b5feff37127c62e26532269eb718f3

Observation a76f75fc-75d7-4c36-86b2-3f2e8582a2bd · outbound

This paper cites Box2d: A 2d physics engine for games.

Bidirectional Soft Actor-Critic: Leveraging Forward and Reverse KL Divergence for Efficient Reinforcement Learning Box2d: A 2d physics engine for games

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:42:35.408208Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=arxiv_source observed=2026-08-07T11:42:33.548555Z digest=sha256:448e3d4b5221e4d2baaaeaadccc7a8ff94cf9882dafedcc99a076d05bafd70f0

Observation 91e8e127-d4ff-421f-950a-dd767d09f0a5 · outbound

This paper cites Greedification operators for policy optimization: Investigating forward and reverse kl divergences.

Bidirectional Soft Actor-Critic: Leveraging Forward and Reverse KL Divergence for Efficient Reinforcement Learning Greedification operators for policy optimization: Investigating forward and reverse kl divergences

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-07T11:42:33.569756Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T11:42:33.569756Z digest=sha256:6ea1efc1c6efb532792f54b9cb90b4d7507daee93c08a6838721fdb977356104

Observation c3becc7a-769b-42c3-9abb-f51c1c418bb5 · outbound

This paper cites Using expectation-maximization for reinforcement learning.

Bidirectional Soft Actor-Critic: Leveraging Forward and Reverse KL Divergence for Efficient Reinforcement Learning Using expectation-maximization for reinforcement learning

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:42:35.392907Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=arxiv_source observed=2026-08-07T11:42:33.606210Z digest=sha256:5546f78644806915a120eb406d99ce37d5df426e6617cd73c63dc9ab8e252d3f

Observation a15b1794-9e53-4f58-a320-1a8848be737a · outbound

This paper cites Soft actor-critic for navigation of mobile robots.

Bidirectional Soft Actor-Critic: Leveraging Forward and Reverse KL Divergence for Efficient Reinforcement Learning Soft actor-critic for navigation of mobile robots

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:42:35.383303Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=arxiv_source observed=2026-08-07T11:42:33.642675Z digest=sha256:8d7d0531a6804a37745a576565683e7f4e6e6e1c852f92bcb230369b8fb3de61

Observation 4c72111c-d7f2-4dc9-991f-b248de50cd2f · outbound

This paper cites Distributional soft actor-critic: Off-policy reinforcement learning for addressing value estimation errors.

Bidirectional Soft Actor-Critic: Leveraging Forward and Reverse KL Divergence for Efficient Reinforcement Learning Distributional soft actor-critic: Off-policy reinforcement learning for addressing value estimation errors

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:42:35.374373Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=arxiv_source observed=2026-08-07T11:42:33.671122Z digest=sha256:5eb8b33b457e651f51c31666d140cab6c27e5af581456b23d4cb898c587d8279

Observation 9e1a854b-3c7a-4867-a536-e8a699640a70 · outbound

This paper cites Distributional soft actor-critic: Off-policy reinforcement learning for addressing value estimation errors.

Bidirectional Soft Actor-Critic: Leveraging Forward and Reverse KL Divergence for Efficient Reinforcement Learning Distributional soft actor-critic: Off-policy reinforcement learning for addressing value estimation errors

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:42:35.364809Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=arxiv_source observed=2026-08-07T11:42:33.692111Z digest=sha256:5f6cb86af53da3c9ec290be4a000df4780431888decefdda3c6782595bc55507

Observation a0d0cc48-fe5e-45cc-9910-719dcbb4657c · outbound

This paper cites Virel: A variational inference framework for reinforcement learning.

Bidirectional Soft Actor-Critic: Leveraging Forward and Reverse KL Divergence for Efficient Reinforcement Learning Virel: A variational inference framework for reinforcement learning

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:42:35.355409Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=arxiv_source observed=2026-08-07T11:42:33.713449Z digest=sha256:670f62024eaf11c97ab791adddf2aec8fc1132f9be66929781d970d568985c3e

Observation 3032c6c8-423b-4b27-b522-9467f96829ad · outbound

This paper cites Brax - a differentiable physics engine for large scale rigid body simulation.

Bidirectional Soft Actor-Critic: Leveraging Forward and Reverse KL Divergence for Efficient Reinforcement Learning Brax - a differentiable physics engine for large scale rigid body simulation

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:42:35.345980Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=arxiv_source observed=2026-08-07T11:42:33.734844Z digest=sha256:ed858d3d5ff4a3a7c248e1c7f22a063bae69a8cef613d0d610b9f7a2795b01ec

Observation cfe213d7-5062-4740-bba8-ddca2d989ba3 · outbound

This paper cites Iq-learn: Inverse soft-q learning for imitation.

Bidirectional Soft Actor-Critic: Leveraging Forward and Reverse KL Divergence for Efficient Reinforcement Learning Iq-learn: Inverse soft-q learning for imitation

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:42:35.336235Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=arxiv_source observed=2026-08-07T11:42:33.756905Z digest=sha256:0822876babce5eb291eca264c03dae86986884e16729d9639bf4d3d5664e7655

Observation 589bee85-2e5e-4e9a-9106-8d32a3887ca8 · outbound

This paper cites Reinforcement learning with deep energy-based policies.

Bidirectional Soft Actor-Critic: Leveraging Forward and Reverse KL Divergence for Efficient Reinforcement Learning Reinforcement learning with deep energy-based policies

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:42:35.326371Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=arxiv_source observed=2026-08-07T11:42:33.779818Z digest=sha256:e35348facb2316dc8b9c3428d038767f0f30ea4210dfeb400d6a0ee5e1f6fcc8

Observation 13f7922a-3653-4d92-8f01-029b477554f0 · outbound

This paper cites Soft actor-critic: Off-policy maximum entropy deep reinforcement learning with a stochastic actor.

Bidirectional Soft Actor-Critic: Leveraging Forward and Reverse KL Divergence for Efficient Reinforcement Learning Soft actor-critic: Off-policy maximum entropy deep reinforcement learning with a stochastic actor

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:42:35.317218Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=arxiv_source observed=2026-08-07T11:42:33.801739Z digest=sha256:801b1c965c0745563900e4606b5df2c57aa6b6b917050873c6591f0682d1eb3a

Observation 09af56c0-619f-4464-b928-8b977479ce74 · outbound

This paper cites Soft Actor-Critic Algorithms and Applications.

Bidirectional Soft Actor-Critic: Leveraging Forward and Reverse KL Divergence for Efficient Reinforcement Learning Soft Actor-Critic Algorithms and Applications

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-07T11:42:33.823457Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T11:42:33.823457Z digest=sha256:dd778d251e044815862b47079c537dda8b2ddb3e592e9c2106b15369a73cb65b

Observation 07c4f6a2-5767-458f-b939-a5c2a23099d9 · outbound

This paper cites The curse of dimensionality for numerical integration of smooth functions.

Bidirectional Soft Actor-Critic: Leveraging Forward and Reverse KL Divergence for Efficient Reinforcement Learning The curse of dimensionality for numerical integration of smooth functions

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:42:35.307790Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=arxiv_source observed=2026-08-07T11:42:33.844042Z digest=sha256:17b66e5f115061c53745d9c96161bde90b38a7533474d0d895987ba2ecefe98a

Observation 6f661f2b-2f62-430a-b6b9-17b221abea52 · outbound

This paper cites Graph soft actor–critic reinforcement learning for large-scale distributed multirobot coordination.

Bidirectional Soft Actor-Critic: Leveraging Forward and Reverse KL Divergence for Efficient Reinforcement Learning Graph soft actor–critic reinforcement learning for large-scale distributed multirobot coordination

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:42:35.297605Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=arxiv_source observed=2026-08-07T11:42:33.874324Z digest=sha256:89f2af9f61f581673444ebb8f08f55482fd7ab7c7213d7e3fdb5e16e1c900bf2

Observation c1ada51c-5566-4cae-8b05-ee776a377e36 · outbound

This paper cites Cleanrl: High-quality single-file implementations of deep reinforcement learning algorithms.

Bidirectional Soft Actor-Critic: Leveraging Forward and Reverse KL Divergence for Efficient Reinforcement Learning Cleanrl: High-quality single-file implementations of deep reinforcement learning algorithms

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:42:35.288646Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=arxiv_source observed=2026-08-07T11:42:33.903554Z digest=sha256:e0ca85f816cd3ee2f6dbc6b475f67e39e7dcb247bb919f3eee2df0a54ab5e974

Observation 9ee33d03-f8d9-4270-a3e6-bdcd64224198 · outbound

This paper cites Accelerating reinforcement learning with value-conditional state entropy exploration.

Bidirectional Soft Actor-Critic: Leveraging Forward and Reverse KL Divergence for Efficient Reinforcement Learning Accelerating reinforcement learning with value-conditional state entropy exploration

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:42:35.278864Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=arxiv_source observed=2026-08-07T11:42:33.931033Z digest=sha256:0d29298173811c1acc30015ad1740644f2e8452a5b7e664f36643d6c8e96ce9a

Observation fff43b0d-d505-4b4c-bb31-b943f546a04b · outbound

This paper cites Optimistic reinforcement learning by forward kullback--leibler divergence optimization.

Bidirectional Soft Actor-Critic: Leveraging Forward and Reverse KL Divergence for Efficient Reinforcement Learning Optimistic reinforcement learning by forward kullback--leibler divergence optimization

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:42:35.268488Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=arxiv_source observed=2026-08-07T11:42:33.952445Z digest=sha256:4003acc8d2ab6750a6e58a26feccbe0307beefb57f0a11f9b7ede67361b6a232

Observation f4940d73-37ee-4291-af6e-d085bf9bb700 · outbound

This paper cites Conservative q-learning for offline reinforcement learning.

Bidirectional Soft Actor-Critic: Leveraging Forward and Reverse KL Divergence for Efficient Reinforcement Learning Conservative q-learning for offline reinforcement learning

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:42:35.258046Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=arxiv_source observed=2026-08-07T11:42:33.973438Z digest=sha256:79b8c34a2acd08cb2b1d1893c751a0c9ca9ccbfc53b5da417a70c7e2f1455ebb

Observation af218bc2-7795-4bd9-99b4-11c5e994d2ce · outbound

This paper cites Continuous control with deep reinforcement learning.

Bidirectional Soft Actor-Critic: Leveraging Forward and Reverse KL Divergence for Efficient Reinforcement Learning Continuous control with deep reinforcement learning

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-07T11:42:33.995064Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T11:42:33.995064Z digest=sha256:817b8c77c9afd84a8053eecadb9a009ad64915374e39d58afa400ab816f3c6ff

Observation 7a4ee0e9-c39c-4524-871a-32c1d5b62299 · outbound

This paper cites Constrained variational policy optimization for safe reinforcement learning.

Bidirectional Soft Actor-Critic: Leveraging Forward and Reverse KL Divergence for Efficient Reinforcement Learning Constrained variational policy optimization for safe reinforcement learning

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:42:35.248022Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=arxiv_source observed=2026-08-07T11:42:34.017398Z digest=sha256:a7c4ef6aeea03fe1f296c740fa0bef797841c3c54cff3d60e08c58aaad9c2722

Observation ef99ea56-bcce-4636-9b5a-a19d99e946b8 · outbound

This paper cites Algorithm 145: Adaptive numerical integration by simpson's rule.

Bidirectional Soft Actor-Critic: Leveraging Forward and Reverse KL Divergence for Efficient Reinforcement Learning Algorithm 145: Adaptive numerical integration by simpson's rule

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:42:35.238937Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=arxiv_source observed=2026-08-07T11:42:34.042879Z digest=sha256:569f16210dd1e24e61281176a62bdcb145bb0da0ef8703c3cdd06bc1ee2287a8

Observation 31cb562f-2668-4bdc-998a-82683a8e3736 · outbound

This paper cites On principled entropy exploration in policy optimization.

Bidirectional Soft Actor-Critic: Leveraging Forward and Reverse KL Divergence for Efficient Reinforcement Learning On principled entropy exploration in policy optimization

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:42:35.227828Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=arxiv_source observed=2026-08-07T11:42:34.064724Z digest=sha256:5abe59738a8eac72a25ad32b2f1b2071b2c5a818805eefdb69c87fc229026826

Observation a4ce579e-74e0-4780-b0b3-f6b3ebf5c88f · outbound

This paper cites Importance sampling techniques for policy optimization.

Bidirectional Soft Actor-Critic: Leveraging Forward and Reverse KL Divergence for Efficient Reinforcement Learning Importance sampling techniques for policy optimization

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:42:35.217571Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=arxiv_source observed=2026-08-07T11:42:34.095121Z digest=sha256:2b51c44dc317b3b94bbfff0c9ed598fe0e73b998554aa8017423a496f5a3d347

Observation ebb7a7e3-ef59-4277-9405-574668b10395 · outbound

This paper cites Human-level control through deep reinforcement learning.

Bidirectional Soft Actor-Critic: Leveraging Forward and Reverse KL Divergence for Efficient Reinforcement Learning Human-level control through deep reinforcement learning

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-07T11:42:34.147276Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T11:42:34.147276Z digest=sha256:e195db2bf259ada16df73e63bcb7b58475e78990eec59a22f3d4ebaa1a03bf20

Observation 9198cf18-ceb7-4813-9cd9-0f3804966b78 · outbound

This paper cites Asynchronous methods for deep reinforcement learning.

Bidirectional Soft Actor-Critic: Leveraging Forward and Reverse KL Divergence for Efficient Reinforcement Learning Asynchronous methods for deep reinforcement learning

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:42:35.201661Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=arxiv_source observed=2026-08-07T11:42:34.170202Z digest=sha256:5cb3efed5a89d54863c4e7d1ccfe81015fb547aa85ee99bc12c88c1bf312262a

Observation c4b94ee6-79d0-4931-8ac2-ba475ccf4576 · outbound

This paper cites Improving Policy Gradient by Exploring Under-appreciated Rewards.

Bidirectional Soft Actor-Critic: Leveraging Forward and Reverse KL Divergence for Efficient Reinforcement Learning Improving Policy Gradient by Exploring Under-appreciated Rewards

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-07T11:42:34.196122Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T11:42:34.196122Z digest=sha256:38387235dfe224cb6a7c04eea376347d795b5b90a2d55b25ccf58af4ff8b4ee1

Observation c225e167-9421-4220-b4f7-27df15930022 · outbound

This paper cites Stable-baselines3: Reliable reinforcement learning implementations.

Bidirectional Soft Actor-Critic: Leveraging Forward and Reverse KL Divergence for Efficient Reinforcement Learning Stable-baselines3: Reliable reinforcement learning implementations

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-07T11:42:34.217785Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T11:42:34.217785Z digest=sha256:84102f1dba4de8749910f54bcf82102f752504961cc16b5f7dcfb244ed3293e2

Observation 69f38043-f28e-468a-a69e-d6efaccbcd41 · outbound

This paper cites Trust region policy optimization.

Bidirectional Soft Actor-Critic: Leveraging Forward and Reverse KL Divergence for Efficient Reinforcement Learning Trust region policy optimization

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:42:35.182420Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=arxiv_source observed=2026-08-07T11:42:34.238985Z digest=sha256:9c8ae9fc30a709524adf49dc23165ea9095cf10c4cfdc61ccae893b51706890a

Observation a83300ff-a461-4fae-ab21-f7d0736b5e1d · outbound

This paper cites Proximal Policy Optimization Algorithms.

Bidirectional Soft Actor-Critic: Leveraging Forward and Reverse KL Divergence for Efficient Reinforcement Learning Proximal Policy Optimization Algorithms

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-07T11:42:34.261078Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T11:42:34.261078Z digest=sha256:239ce6ec902c971e21f9f7483a6f8d0af3f6da08ff215f9bb7c048126db605df

Observation 1513d567-505b-47af-9015-2f7aca0c2c7d · outbound

This paper cites Monte carlo sampling methods.

Bidirectional Soft Actor-Critic: Leveraging Forward and Reverse KL Divergence for Efficient Reinforcement Learning Monte carlo sampling methods

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:42:35.134269Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=arxiv_source observed=2026-08-07T11:42:34.285541Z digest=sha256:aae1d7d8884fb3d51b4adfd99e743de06d7fad9b800287e0b290cc8c3ca503a7

Observation 10f56de9-2b04-4e8c-af7e-3c56428a2a02 · outbound

This paper cites V-mpo: On-policy maximum a posteriori policy optimization for discrete and continuous control.

Bidirectional Soft Actor-Critic: Leveraging Forward and Reverse KL Divergence for Efficient Reinforcement Learning V-mpo: On-policy maximum a posteriori policy optimization for discrete and continuous control

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:42:35.056253Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=arxiv_source observed=2026-08-07T11:42:34.312110Z digest=sha256:73af38457acb490201873f2c20779fd17a063efa8ea560edf5fdabea0145c147

Observation 884cde22-32a5-463b-b479-c70cad6a28a7 · outbound

This paper cites Value-Decomposition Networks For Cooperative Multi-Agent Learning.

Bidirectional Soft Actor-Critic: Leveraging Forward and Reverse KL Divergence for Efficient Reinforcement Learning Value-Decomposition Networks For Cooperative Multi-Agent Learning

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-07T11:42:34.333612Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T11:42:34.333612Z digest=sha256:de4374c175067c716f126d6618e7526b5baf9a75e4a20b2362d676f2ead1cdfd

Observation a0e0493e-d1ca-4cdc-96c2-5fbe9793ae0a · outbound

This paper cites Sutton and Andrew G.

Bidirectional Soft Actor-Critic: Leveraging Forward and Reverse KL Divergence for Efficient Reinforcement Learning Sutton and Andrew G

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-07T11:42:34.371554Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T11:42:34.371554Z digest=sha256:590f24b0f30bf2a271f1a2c6c23017325c8f95652b6c9e9546e954d67dc2a128

Observation e1a785fa-c6f7-4a64-910c-c4b7f60424ba · outbound

This paper cites Mujoco: A physics engine for model-based control.

Bidirectional Soft Actor-Critic: Leveraging Forward and Reverse KL Divergence for Efficient Reinforcement Learning Mujoco: A physics engine for model-based control

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-07T11:42:34.393604Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T11:42:34.393604Z digest=sha256:d70be7f557efc7e89bedfe6a3585c6b2f9673243b4649df0b3b0c5806340ebe5

Observation 4872ae71-a7d7-4117-b3db-df8724e4ae0b · outbound

This paper cites Probabilistic inference for solving discrete and continuous state markov decision processes.

Bidirectional Soft Actor-Critic: Leveraging Forward and Reverse KL Divergence for Efficient Reinforcement Learning Probabilistic inference for solving discrete and continuous state markov decision processes

Reference 40

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:42:34.998747Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=arxiv_source observed=2026-08-07T11:42:34.414592Z digest=sha256:ffb65f5d0dd2c104d13526fcb1877534842871f39cc2707b4d2425a820d474ca

Observation e2e48a29-57ef-40ea-b092-25ed2df07db4 · outbound

This paper cites Self-play reinforcement learning guides protein engineering.

Bidirectional Soft Actor-Critic: Leveraging Forward and Reverse KL Divergence for Efficient Reinforcement Learning Self-play reinforcement learning guides protein engineering

Reference 41

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:42:34.939905Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=arxiv_source observed=2026-08-07T11:42:34.436726Z digest=sha256:fedcd6a81de3740f545baabc296fb393b43e545a75035714f20e4bf69e2772fc

Observation 1b92f031-66b6-4596-a6af-efae17983045 · outbound

This paper cites Scalable trust-region method for deep reinforcement learning using kronecker-factored approximation.

Bidirectional Soft Actor-Critic: Leveraging Forward and Reverse KL Divergence for Efficient Reinforcement Learning Scalable trust-region method for deep reinforcement learning using kronecker-factored approximation

Reference 42

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:42:34.897254Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=arxiv_source observed=2026-08-07T11:42:34.458763Z digest=sha256:77ebc795f408de6f8880e4e66e7df3449c3a7b4926931dcc19513686f18a5eb5

Observation d012f224-7276-4e29-92e3-93359c316e22 · outbound

This paper cites Monocular vision approach for soft actor-critic based car-following strategy in adaptive cruise control.

Bidirectional Soft Actor-Critic: Leveraging Forward and Reverse KL Divergence for Efficient Reinforcement Learning Monocular vision approach for soft actor-critic based car-following strategy in adaptive cruise control

Reference 43

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:42:34.839762Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=arxiv_source observed=2026-08-07T11:42:34.480133Z digest=sha256:3176ca43b9bcafd9dc726513236d272486087eb1e7effe1790f7b623e544afbd

Observation 1bdd5f55-a73a-40ce-a820-4a8fafe4d91e · outbound

This paper cites Wcsac: Worst-case soft actor critic for safety-constrained reinforcement learning.

Bidirectional Soft Actor-Critic: Leveraging Forward and Reverse KL Divergence for Efficient Reinforcement Learning Wcsac: Worst-case soft actor critic for safety-constrained reinforcement learning

Reference 44

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:42:34.785426Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=arxiv_source observed=2026-08-07T11:42:34.501065Z digest=sha256:46e5cc4fe8cf92d422876e674ba75188e0d30155c8b5f3ebd1c2c97d6c78c81b

Observation 84d7064f-7693-4298-be96-8accd733b318 · outbound

This paper cites Maximum entropy inverse reinforcement learning.

Bidirectional Soft Actor-Critic: Leveraging Forward and Reverse KL Divergence for Efficient Reinforcement Learning Maximum entropy inverse reinforcement learning

Reference 45

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:42:34.731567Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=arxiv_source observed=2026-08-07T11:42:34.523235Z digest=sha256:c0f837e43766aeb5f04aca5d8159f49fbf9661057179c47c94357c08661148cd

Observation 7cb4a608-aa9c-4c54-b024-3d59f93e49df · outbound

This paper cites Wasserstein gradient flows for optimizing gaussian mixture policies.

Bidirectional Soft Actor-Critic: Leveraging Forward and Reverse KL Divergence for Efficient Reinforcement Learning Wasserstein gradient flows for optimizing gaussian mixture policies

Reference 46

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:42:34.688869Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=arxiv_source observed=2026-08-07T11:42:34.550668Z digest=sha256:dd39005f3759c97303003dd47170acf4b059fd2c889f2a41f98ceed9a6752548

Pith citing papers

No inbound Pith citation observations are available.