Pith. sign in

Paper Citation Record · LEDGER

Bidirectional Soft Actor-Critic: Leveraging Forward and Reverse KL Divergence for Efficient Reinforcement Learning

As of 15 August 2026, this Paper Citation Record lists 46 of 46 outbound references and 0 inbound Pith citation observations for arXiv:2506.01639.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2506.01639 v1

Coverage vector

measured 46 of 46 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-07T11:42:34.550668Z

measured 46 of 46 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-14T06:32:32.682623+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

46 of 46 outbound references displayed

  • verified exact0
  • verified fuzzy35
  • unresolved11
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 037ab0cc-4bfe-4cb9-83c5-532dc0ba7cdc · outbound

This paper cites Deriving and improving cma-es with information geometric trust regions.

Bidirectional Soft Actor-Critic: Leveraging Forward and Reverse KL Divergence for Efficient Reinforcement Learning Deriving and improving cma-es with information geometric trust regions

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:42:35.449301Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-08-07T11:42:33.434436Z digest=sha256:04a5d13c3a3fdd67e159dbbb55c132fdbd07e1ebce911862b8c44a4528cd4e0e

Observation e99c98b3-4569-4f1f-8ab5-30ed82433c2b · outbound

This paper cites Maximum a posteriori policy optimisation.

Bidirectional Soft Actor-Critic: Leveraging Forward and Reverse KL Divergence for Efficient Reinforcement Learning Maximum a posteriori policy optimisation

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:42:35.438011Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-08-07T11:42:33.454596Z digest=sha256:39bcf75de9ed012d553c01ea10b41e424d9dcf2234c8145ca06040b99518c810

Observation edc3fbf0-fb53-47d0-ad32-f868c9c1eef6 · outbound

This paper cites Maximum a Posteriori Policy Optimisation.

Bidirectional Soft Actor-Critic: Leveraging Forward and Reverse KL Divergence for Efficient Reinforcement Learning Maximum a Posteriori Policy Optimisation

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-07T11:42:33.475700Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T11:42:33.475700Z digest=sha256:923b8b5e5b70ff5e7481c3f54d23115936d450fea38848f379ca656505718547

Observation 669f855a-fb1f-4345-8f34-9131c79842e1 · outbound

This paper cites Improved soft actor-critic: Mixing prioritized off-policy samples with on-policy experiences.

Bidirectional Soft Actor-Critic: Leveraging Forward and Reverse KL Divergence for Efficient Reinforcement Learning Improved soft actor-critic: Mixing prioritized off-policy samples with on-policy experiences

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:42:35.427993Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-08-07T11:42:33.497612Z digest=sha256:563f5c9ebc1ee0c1fc66b0beab408524ea62f05f2f8c4a20db2e9e5a93964574

Observation 998ab124-4c61-4dbe-9271-1092e56717df · outbound

This paper cites Sukhatme.

Bidirectional Soft Actor-Critic: Leveraging Forward and Reverse KL Divergence for Efficient Reinforcement Learning Sukhatme

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:42:35.418179Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-08-07T11:42:33.519432Z digest=sha256:10aff305ca55c6f3b075680b4bcf7512d812a2f71edbc6b9362fee0d941e4923

Observation a76f75fc-75d7-4c36-86b2-3f2e8582a2bd · outbound

This paper cites Box2d: A 2d physics engine for games.

Bidirectional Soft Actor-Critic: Leveraging Forward and Reverse KL Divergence for Efficient Reinforcement Learning Box2d: A 2d physics engine for games

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:42:35.408208Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-08-07T11:42:33.548555Z digest=sha256:ca8ad955004fd9565fa0472a3ae29aa08d318a6b57b027a0bfd837b21c014361

Observation 91e8e127-d4ff-421f-950a-dd767d09f0a5 · outbound

This paper cites Greedification operators for policy optimization: Investigating forward and reverse kl divergences.

Bidirectional Soft Actor-Critic: Leveraging Forward and Reverse KL Divergence for Efficient Reinforcement Learning Greedification operators for policy optimization: Investigating forward and reverse kl divergences

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-07T11:42:33.569756Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T11:42:33.569756Z digest=sha256:6ea1efc1c6efb532792f54b9cb90b4d7507daee93c08a6838721fdb977356104

Observation c3becc7a-769b-42c3-9abb-f51c1c418bb5 · outbound

This paper cites Using expectation-maximization for reinforcement learning.

Bidirectional Soft Actor-Critic: Leveraging Forward and Reverse KL Divergence for Efficient Reinforcement Learning Using expectation-maximization for reinforcement learning

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:42:35.392907Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-08-07T11:42:33.606210Z digest=sha256:d7341ef40cce788a44344738eb5f492e264e2af8df34996bf938decca4feda30

Observation a15b1794-9e53-4f58-a320-1a8848be737a · outbound

This paper cites Soft actor-critic for navigation of mobile robots.

Bidirectional Soft Actor-Critic: Leveraging Forward and Reverse KL Divergence for Efficient Reinforcement Learning Soft actor-critic for navigation of mobile robots

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:42:35.383303Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-08-07T11:42:33.642675Z digest=sha256:bdc58b1b8ace34f3dfe4b6b7e9a53d9c9490aa91be79454e386c940400f499a1

Observation 4c72111c-d7f2-4dc9-991f-b248de50cd2f · outbound

This paper cites Distributional soft actor-critic: Off-policy reinforcement learning for addressing value estimation errors.

Bidirectional Soft Actor-Critic: Leveraging Forward and Reverse KL Divergence for Efficient Reinforcement Learning Distributional soft actor-critic: Off-policy reinforcement learning for addressing value estimation errors

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:42:35.374373Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-08-07T11:42:33.671122Z digest=sha256:027d92d87700059d6b5fcecd51204dbf3675f47f8c5c8e4b93519d97479ebe8c

Observation 9e1a854b-3c7a-4867-a536-e8a699640a70 · outbound

This paper cites Distributional soft actor-critic: Off-policy reinforcement learning for addressing value estimation errors.

Bidirectional Soft Actor-Critic: Leveraging Forward and Reverse KL Divergence for Efficient Reinforcement Learning Distributional soft actor-critic: Off-policy reinforcement learning for addressing value estimation errors

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:42:35.364809Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-08-07T11:42:33.692111Z digest=sha256:78208772680b0d016e169e90e90a2724a1ff402fe54b1d2238b1247df7e20386

Observation a0d0cc48-fe5e-45cc-9910-719dcbb4657c · outbound

This paper cites Virel: A variational inference framework for reinforcement learning.

Bidirectional Soft Actor-Critic: Leveraging Forward and Reverse KL Divergence for Efficient Reinforcement Learning Virel: A variational inference framework for reinforcement learning

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:42:35.355409Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-08-07T11:42:33.713449Z digest=sha256:55d1155fa4765013a9fedf00ffa2bb4f4ad959dbe45338fe14cf1aa9516f8332

Observation 3032c6c8-423b-4b27-b522-9467f96829ad · outbound

This paper cites Brax - a differentiable physics engine for large scale rigid body simulation.

Bidirectional Soft Actor-Critic: Leveraging Forward and Reverse KL Divergence for Efficient Reinforcement Learning Brax - a differentiable physics engine for large scale rigid body simulation

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:42:35.345980Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-08-07T11:42:33.734844Z digest=sha256:6493f6151e0b4ab8a1e94bdd9f58c7c0ccca6f5d39b48c41b81012d6045c3345

Observation cfe213d7-5062-4740-bba8-ddca2d989ba3 · outbound

This paper cites Iq-learn: Inverse soft-q learning for imitation.

Bidirectional Soft Actor-Critic: Leveraging Forward and Reverse KL Divergence for Efficient Reinforcement Learning Iq-learn: Inverse soft-q learning for imitation

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:42:35.336235Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-08-07T11:42:33.756905Z digest=sha256:195c351d58443ab3a7278d1dc3cd1a5b647519eb0686dd2d9db8f9f9c930793b

Observation 589bee85-2e5e-4e9a-9106-8d32a3887ca8 · outbound

This paper cites Reinforcement learning with deep energy-based policies.

Bidirectional Soft Actor-Critic: Leveraging Forward and Reverse KL Divergence for Efficient Reinforcement Learning Reinforcement learning with deep energy-based policies

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:42:35.326371Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-08-07T11:42:33.779818Z digest=sha256:d72e1343702fafe4e52e55948d79c059c7d86988a947197b81df14d3f5579458

Observation 13f7922a-3653-4d92-8f01-029b477554f0 · outbound

This paper cites Soft actor-critic: Off-policy maximum entropy deep reinforcement learning with a stochastic actor.

Bidirectional Soft Actor-Critic: Leveraging Forward and Reverse KL Divergence for Efficient Reinforcement Learning Soft actor-critic: Off-policy maximum entropy deep reinforcement learning with a stochastic actor

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:42:35.317218Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-08-07T11:42:33.801739Z digest=sha256:0a191bc406f2513dc9ae347c70d63a9aadd2a56f3cf884168995ccfb87edf3d1

Observation 09af56c0-619f-4464-b928-8b977479ce74 · outbound

This paper cites Soft Actor-Critic Algorithms and Applications.

Bidirectional Soft Actor-Critic: Leveraging Forward and Reverse KL Divergence for Efficient Reinforcement Learning Soft Actor-Critic Algorithms and Applications

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-07T11:42:33.823457Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T11:42:33.823457Z digest=sha256:dd778d251e044815862b47079c537dda8b2ddb3e592e9c2106b15369a73cb65b

Observation 07c4f6a2-5767-458f-b939-a5c2a23099d9 · outbound

This paper cites The curse of dimensionality for numerical integration of smooth functions.

Bidirectional Soft Actor-Critic: Leveraging Forward and Reverse KL Divergence for Efficient Reinforcement Learning The curse of dimensionality for numerical integration of smooth functions

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:42:35.307790Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-08-07T11:42:33.844042Z digest=sha256:3094a9ef9ac2d9bdf64638f07283ccd88083049a0110a45aaac5865b8feec209

Observation 6f661f2b-2f62-430a-b6b9-17b221abea52 · outbound

This paper cites Graph soft actor–critic reinforcement learning for large-scale distributed multirobot coordination.

Bidirectional Soft Actor-Critic: Leveraging Forward and Reverse KL Divergence for Efficient Reinforcement Learning Graph soft actor–critic reinforcement learning for large-scale distributed multirobot coordination

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:42:35.297605Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-08-07T11:42:33.874324Z digest=sha256:7bf084af3f7b9cc4f865d0cf58e6b8517ff264bf9b629abb072d0a4ebdf0b18d

Observation c1ada51c-5566-4cae-8b05-ee776a377e36 · outbound

This paper cites Cleanrl: High-quality single-file implementations of deep reinforcement learning algorithms.

Bidirectional Soft Actor-Critic: Leveraging Forward and Reverse KL Divergence for Efficient Reinforcement Learning Cleanrl: High-quality single-file implementations of deep reinforcement learning algorithms

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:42:35.288646Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-08-07T11:42:33.903554Z digest=sha256:706662f26fd99d5517bac3090191bffc631facd3ef1e0f4465ba3d710409e2c1

Observation 9ee33d03-f8d9-4270-a3e6-bdcd64224198 · outbound

This paper cites Accelerating reinforcement learning with value-conditional state entropy exploration.

Bidirectional Soft Actor-Critic: Leveraging Forward and Reverse KL Divergence for Efficient Reinforcement Learning Accelerating reinforcement learning with value-conditional state entropy exploration

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:42:35.278864Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-08-07T11:42:33.931033Z digest=sha256:5af69523a39307e3b1cfe00f6f4ee4c03746573fc97285d698cd848b8f597607

Observation fff43b0d-d505-4b4c-bb31-b943f546a04b · outbound

This paper cites Optimistic reinforcement learning by forward kullback--leibler divergence optimization.

Bidirectional Soft Actor-Critic: Leveraging Forward and Reverse KL Divergence for Efficient Reinforcement Learning Optimistic reinforcement learning by forward kullback--leibler divergence optimization

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:42:35.268488Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-08-07T11:42:33.952445Z digest=sha256:16295e415c3b1d8e9d106d56b8083f8c6b3644af8e3ab821a4065a7776f100d5

Observation f4940d73-37ee-4291-af6e-d085bf9bb700 · outbound

This paper cites Conservative q-learning for offline reinforcement learning.

Bidirectional Soft Actor-Critic: Leveraging Forward and Reverse KL Divergence for Efficient Reinforcement Learning Conservative q-learning for offline reinforcement learning

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:42:35.258046Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-08-07T11:42:33.973438Z digest=sha256:b85af60215142ba3a812b1f8cb5c39aa420ccdcd997e08f4f78c0d8116561eb7

Observation af218bc2-7795-4bd9-99b4-11c5e994d2ce · outbound

This paper cites Continuous control with deep reinforcement learning.

Bidirectional Soft Actor-Critic: Leveraging Forward and Reverse KL Divergence for Efficient Reinforcement Learning Continuous control with deep reinforcement learning

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-07T11:42:33.995064Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T11:42:33.995064Z digest=sha256:817b8c77c9afd84a8053eecadb9a009ad64915374e39d58afa400ab816f3c6ff

Observation 7a4ee0e9-c39c-4524-871a-32c1d5b62299 · outbound

This paper cites Constrained variational policy optimization for safe reinforcement learning.

Bidirectional Soft Actor-Critic: Leveraging Forward and Reverse KL Divergence for Efficient Reinforcement Learning Constrained variational policy optimization for safe reinforcement learning

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:42:35.248022Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-08-07T11:42:34.017398Z digest=sha256:50a5bc0bee095fcbbe79ad14b2ca77b2a852926ad23612b3ce50c3ce32400fbb

Observation ef99ea56-bcce-4636-9b5a-a19d99e946b8 · outbound

This paper cites Algorithm 145: Adaptive numerical integration by simpson's rule.

Bidirectional Soft Actor-Critic: Leveraging Forward and Reverse KL Divergence for Efficient Reinforcement Learning Algorithm 145: Adaptive numerical integration by simpson's rule

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:42:35.238937Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-08-07T11:42:34.042879Z digest=sha256:8cf4300af9026fb9c569d9425d26ced39cfa6d0ee69e0f65aa4013c7d3fab1a0

Observation 31cb562f-2668-4bdc-998a-82683a8e3736 · outbound

This paper cites On principled entropy exploration in policy optimization.

Bidirectional Soft Actor-Critic: Leveraging Forward and Reverse KL Divergence for Efficient Reinforcement Learning On principled entropy exploration in policy optimization

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:42:35.227828Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-08-07T11:42:34.064724Z digest=sha256:f87a29ea546e737679a4d8db13ca16334016139a49ab9e088a2be4ed57cb9727

Observation a4ce579e-74e0-4780-b0b3-f6b3ebf5c88f · outbound

This paper cites Importance sampling techniques for policy optimization.

Bidirectional Soft Actor-Critic: Leveraging Forward and Reverse KL Divergence for Efficient Reinforcement Learning Importance sampling techniques for policy optimization

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:42:35.217571Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-08-07T11:42:34.095121Z digest=sha256:dfb818a191307bd373d74527d14503b29dc341b028e47341c675840178b6096c

Observation ebb7a7e3-ef59-4277-9405-574668b10395 · outbound

This paper cites Human-level control through deep reinforcement learning.

Bidirectional Soft Actor-Critic: Leveraging Forward and Reverse KL Divergence for Efficient Reinforcement Learning Human-level control through deep reinforcement learning

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-07T11:42:34.147276Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T11:42:34.147276Z digest=sha256:e195db2bf259ada16df73e63bcb7b58475e78990eec59a22f3d4ebaa1a03bf20

Observation 9198cf18-ceb7-4813-9cd9-0f3804966b78 · outbound

This paper cites Asynchronous methods for deep reinforcement learning.

Bidirectional Soft Actor-Critic: Leveraging Forward and Reverse KL Divergence for Efficient Reinforcement Learning Asynchronous methods for deep reinforcement learning

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:42:35.201661Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-08-07T11:42:34.170202Z digest=sha256:5d6882af975a79dc5e22e7bc27d13379d83a084a18aedb8551f7b83da18b29e6

Observation c4b94ee6-79d0-4931-8ac2-ba475ccf4576 · outbound

This paper cites Improving Policy Gradient by Exploring Under-appreciated Rewards.

Bidirectional Soft Actor-Critic: Leveraging Forward and Reverse KL Divergence for Efficient Reinforcement Learning Improving Policy Gradient by Exploring Under-appreciated Rewards

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-07T11:42:34.196122Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T11:42:34.196122Z digest=sha256:38387235dfe224cb6a7c04eea376347d795b5b90a2d55b25ccf58af4ff8b4ee1

Observation c225e167-9421-4220-b4f7-27df15930022 · outbound

This paper cites Stable-baselines3: Reliable reinforcement learning implementations.

Bidirectional Soft Actor-Critic: Leveraging Forward and Reverse KL Divergence for Efficient Reinforcement Learning Stable-baselines3: Reliable reinforcement learning implementations

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-07T11:42:34.217785Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T11:42:34.217785Z digest=sha256:84102f1dba4de8749910f54bcf82102f752504961cc16b5f7dcfb244ed3293e2

Observation 69f38043-f28e-468a-a69e-d6efaccbcd41 · outbound

This paper cites Trust region policy optimization.

Bidirectional Soft Actor-Critic: Leveraging Forward and Reverse KL Divergence for Efficient Reinforcement Learning Trust region policy optimization

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:42:35.182420Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-08-07T11:42:34.238985Z digest=sha256:022bd54703e2c40847cd4f54ff51308385ba00edec51b7511b61f16af315a835

Observation a83300ff-a461-4fae-ab21-f7d0736b5e1d · outbound

This paper cites Proximal Policy Optimization Algorithms.

Bidirectional Soft Actor-Critic: Leveraging Forward and Reverse KL Divergence for Efficient Reinforcement Learning Proximal Policy Optimization Algorithms

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-07T11:42:34.261078Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T11:42:34.261078Z digest=sha256:239ce6ec902c971e21f9f7483a6f8d0af3f6da08ff215f9bb7c048126db605df

Observation 1513d567-505b-47af-9015-2f7aca0c2c7d · outbound

This paper cites Monte carlo sampling methods.

Bidirectional Soft Actor-Critic: Leveraging Forward and Reverse KL Divergence for Efficient Reinforcement Learning Monte carlo sampling methods

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:42:35.134269Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-08-07T11:42:34.285541Z digest=sha256:9b4dff667de9b4aec60b12cee80c5130a1abb8da4ee2867b4e45cf38084062ed

Observation 10f56de9-2b04-4e8c-af7e-3c56428a2a02 · outbound

This paper cites V-mpo: On-policy maximum a posteriori policy optimization for discrete and continuous control.

Bidirectional Soft Actor-Critic: Leveraging Forward and Reverse KL Divergence for Efficient Reinforcement Learning V-mpo: On-policy maximum a posteriori policy optimization for discrete and continuous control

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:42:35.056253Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-08-07T11:42:34.312110Z digest=sha256:f824b45057df124ea058ea1b025fd2b8e095b2749a38a1c7f6bd97fbcc08d33e

Observation 884cde22-32a5-463b-b479-c70cad6a28a7 · outbound

This paper cites Value-Decomposition Networks For Cooperative Multi-Agent Learning.

Bidirectional Soft Actor-Critic: Leveraging Forward and Reverse KL Divergence for Efficient Reinforcement Learning Value-Decomposition Networks For Cooperative Multi-Agent Learning

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-07T11:42:34.333612Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T11:42:34.333612Z digest=sha256:de4374c175067c716f126d6618e7526b5baf9a75e4a20b2362d676f2ead1cdfd

Observation a0e0493e-d1ca-4cdc-96c2-5fbe9793ae0a · outbound

This paper cites Sutton and Andrew G.

Bidirectional Soft Actor-Critic: Leveraging Forward and Reverse KL Divergence for Efficient Reinforcement Learning Sutton and Andrew G

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-07T11:42:34.371554Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T11:42:34.371554Z digest=sha256:590f24b0f30bf2a271f1a2c6c23017325c8f95652b6c9e9546e954d67dc2a128

Observation e1a785fa-c6f7-4a64-910c-c4b7f60424ba · outbound

This paper cites Mujoco: A physics engine for model-based control.

Bidirectional Soft Actor-Critic: Leveraging Forward and Reverse KL Divergence for Efficient Reinforcement Learning Mujoco: A physics engine for model-based control

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-07T11:42:34.393604Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T11:42:34.393604Z digest=sha256:d70be7f557efc7e89bedfe6a3585c6b2f9673243b4649df0b3b0c5806340ebe5

Observation 4872ae71-a7d7-4117-b3db-df8724e4ae0b · outbound

This paper cites Probabilistic inference for solving discrete and continuous state markov decision processes.

Bidirectional Soft Actor-Critic: Leveraging Forward and Reverse KL Divergence for Efficient Reinforcement Learning Probabilistic inference for solving discrete and continuous state markov decision processes

Reference 40

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:42:34.998747Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-08-07T11:42:34.414592Z digest=sha256:07ac7f793a4f3da1be68d06ee1a1b824f7c7491eb721e2d921923c727086782f

Observation e2e48a29-57ef-40ea-b092-25ed2df07db4 · outbound

This paper cites Self-play reinforcement learning guides protein engineering.

Bidirectional Soft Actor-Critic: Leveraging Forward and Reverse KL Divergence for Efficient Reinforcement Learning Self-play reinforcement learning guides protein engineering

Reference 41

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:42:34.939905Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-08-07T11:42:34.436726Z digest=sha256:562f1b50d55ad0a5f6d8655c00712b302d887b64350963be053d8745c01db6d7

Observation 1b92f031-66b6-4596-a6af-efae17983045 · outbound

This paper cites Scalable trust-region method for deep reinforcement learning using kronecker-factored approximation.

Bidirectional Soft Actor-Critic: Leveraging Forward and Reverse KL Divergence for Efficient Reinforcement Learning Scalable trust-region method for deep reinforcement learning using kronecker-factored approximation

Reference 42

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:42:34.897254Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-08-07T11:42:34.458763Z digest=sha256:97b610c56f855b1d67cbd8a6bf833e900fa78493519d6150b410f037990d9653

Observation d012f224-7276-4e29-92e3-93359c316e22 · outbound

This paper cites Monocular vision approach for soft actor-critic based car-following strategy in adaptive cruise control.

Bidirectional Soft Actor-Critic: Leveraging Forward and Reverse KL Divergence for Efficient Reinforcement Learning Monocular vision approach for soft actor-critic based car-following strategy in adaptive cruise control

Reference 43

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:42:34.839762Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-08-07T11:42:34.480133Z digest=sha256:f9201facf2b3ec21e69b153392e5548ee156ff93c2a7920283969466ca39ff9b

Observation 1bdd5f55-a73a-40ce-a820-4a8fafe4d91e · outbound

This paper cites Wcsac: Worst-case soft actor critic for safety-constrained reinforcement learning.

Bidirectional Soft Actor-Critic: Leveraging Forward and Reverse KL Divergence for Efficient Reinforcement Learning Wcsac: Worst-case soft actor critic for safety-constrained reinforcement learning

Reference 44

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:42:34.785426Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-08-07T11:42:34.501065Z digest=sha256:14828b3ca3d4ef87d73ca85dff141a2e73ee85190ab020d33b6a7800131d00df

Observation 84d7064f-7693-4298-be96-8accd733b318 · outbound

This paper cites Maximum entropy inverse reinforcement learning.

Bidirectional Soft Actor-Critic: Leveraging Forward and Reverse KL Divergence for Efficient Reinforcement Learning Maximum entropy inverse reinforcement learning

Reference 45

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:42:34.731567Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-08-07T11:42:34.523235Z digest=sha256:840022e8d961c327a12f260c954db3aaf78eaceb5b223aa5cda5ff96faee742a

Observation 7cb4a608-aa9c-4c54-b024-3d59f93e49df · outbound

This paper cites Wasserstein gradient flows for optimizing gaussian mixture policies.

Bidirectional Soft Actor-Critic: Leveraging Forward and Reverse KL Divergence for Efficient Reinforcement Learning Wasserstein gradient flows for optimizing gaussian mixture policies

Reference 46

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:42:34.688869Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-08-07T11:42:34.550668Z digest=sha256:4a48194d6efd4d95376de27d9ad742c359c68bffbd40c773581671b9613e73c0

Pith citing papers

No inbound Pith citation observations are available.