Pith. sign in

Paper Citation Record · LEDGER

Average-Reward Soft Actor-Critic

As of 19 August 2026, this Paper Citation Record lists 24 of 24 outbound references and 2 inbound Pith citation observations for arXiv:2501.09080.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2501.09080 v2

Coverage vector

measured 24 of 24 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-10T20:13:57.717789Z

measured 26 of 26 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-19T06:32:44.657259+00:00

measured 2 of 2 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-05-15T08:41:52.119405Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-05-15T08:45:19.217416Z

Reference resolution

24 of 24 outbound references displayed

  • verified exact5
  • verified fuzzy6
  • unresolved13
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation f0146d23-1cdb-4608-9a76-fdb7bbe79c2a · outbound

This paper cites Image Augmentation Is All You Need: Regularizing Deep Reinforcement Learning from Pixels.

Average-Reward Soft Actor-Critic Image Augmentation Is All You Need: Regularizing Deep Reinforcement Learning from Pixels

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-10T20:13:57.551979Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T20:13:57.551979Z digest=sha256:83b116c8469569e322318fbf1f43c8972c2b5ca4653f2dcbf557397ba2fc3197

Observation f2a6e11c-6dc1-4c5a-b861-10be3ea7f69c · outbound

This paper cites Stochastic first-order methods for average-reward Markov decision processes.

Average-Reward Soft Actor-Critic Stochastic first-order methods for average-reward Markov decision processes

Reference 9

Resolution
verified exact
local_arxiv, observed 2026-08-10T20:13:58.151359Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-10T20:13:57.565005Z digest=sha256:060af0a77ce62cc7645cd34445003312bc37140739a2a4efa658b2894d0a18ff

Observation 6c73f66b-967c-4f76-8a4d-f29c5bb2f753 · outbound

This paper cites Lillicrap, Jonathan J.

Average-Reward Soft Actor-Critic Lillicrap, Jonathan J

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T20:13:58.778193Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-10T20:13:57.574195Z digest=sha256:8e0bbbe4e1f836d97c4ac8ce7f2174e18271c51c5c8b3b1c80442953d3cc6442

Observation 1b6fd44b-c44c-4368-894a-37436948a5c3 · outbound

This paper cites Discounted Reinforcement Learning Is Not an Optimization Problem.

Average-Reward Soft Actor-Critic Discounted Reinforcement Learning Is Not an Optimization Problem

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-10T20:13:57.580973Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T20:13:57.580973Z digest=sha256:cb2fdfc95f809e9912ebdcda320a3d5594d4a0dbb75475925a0dfd336ab3343e

Observation 66e598d2-f89e-43f0-9f97-8c10a4339abe · outbound

This paper cites Controllability-Aware Unsupervised Skill Discovery.

Average-Reward Soft Actor-Critic Controllability-Aware Unsupervised Skill Discovery

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-10T20:13:57.602182Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T20:13:57.602182Z digest=sha256:66afa6cceccc9b92d84c88a799d824690715e07f91c64c5b144ccc1790f2f46a

Observation 4054c8d6-6ce6-4337-a5a7-ceaf8b0314e0 · outbound

This paper cites Relative entropy and free energy dualities: Con- nections to path integral and kl control.

Average-Reward Soft Actor-Critic Relative entropy and free energy dualities: Con- nections to path integral and kl control

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T20:13:58.746542Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-10T20:13:57.655247Z digest=sha256:17ec70a5159fb0d4a46ab04847c10a044754e79159f140525e62701cf8a83958

Observation 2dcb0544-e426-4be8-bd63-58865362d663 · outbound

This paper cites Behavior Regularized Offline Reinforcement Learning.

Average-Reward Soft Actor-Critic Behavior Regularized Offline Reinforcement Learning

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-10T20:13:57.685681Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T20:13:57.685681Z digest=sha256:091d6b8ebd022b7c24a2fa6fba0f420b6752e6efc44e6ecf3e5f08ae1e89f7df

Observation 992949d2-b9ae-474e-a18b-77742679f4f9 · outbound

This paper cites Efficient Reinforcement Learning with Large Language Model Priors.

Average-Reward Soft Actor-Critic Efficient Reinforcement Learning with Large Language Model Priors

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-10T20:13:57.693280Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T20:13:57.693280Z digest=sha256:4909666f2afcc96cd1c6fe3b4df4058fda6d0245d074776d44a185cb0d49c1b0

Observation da70aefb-c78e-47c8-8757-d9ca2eed2c29 · outbound

This paper cites Finite sample analysis of average-reward TD learning and Q-learning.

Average-Reward Soft Actor-Critic Finite sample analysis of average-reward TD learning and Q-learning

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T20:13:58.665375Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-10T20:13:57.700341Z digest=sha256:10797219c2ff2133649200d123f845ed1b561cb376d8da0c8b0ca0c0f300cd2d

Observation f5849538-2447-48e1-8bb2-cd64b643977e · outbound

This paper cites reward scale.

Average-Reward Soft Actor-Critic reward scale

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T20:13:58.613200Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-10T20:13:57.717789Z digest=sha256:38c4b6e8668b4291f319f84c0d645bb27c3b80b4dcaba2a2d8807884997b0257

Observation 255b3875-7684-45df-b7fc-a09b7b8aa4b3 · outbound

This paper cites Reward Tweaking: Maximizing the Total Reward While Planning for Short Horizons.

Average-Reward Soft Actor-Critic Reward Tweaking: Maximizing the Total Reward While Planning for Short Horizons

Reference 1999

Resolution
verified exact
local_arxiv, observed 2026-08-10T20:13:57.873267Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-10T20:13:57.638508Z digest=sha256:e286e7a1d23dd2ace6943ad2f0607fa7f1624318f457a5380796a20531b3e5eb

Observation 1d97c46c-cecf-417d-b851-c37836d391bf · outbound

This paper cites Dense dynamics-aware reward synthesis: Integrating prior experience with demonstrations.

Average-Reward Soft Actor-Critic Dense dynamics-aware reward synthesis: Integrating prior experience with demonstrations

Reference 2003

Resolution
verified exact
raw_fallback, observed 2026-08-10T20:13:58.392412Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-10T20:13:57.536401Z digest=sha256:0e0920d5a756ef84c38001eec45d2031b02b8c570233ab8449af27840eaee20c

Observation 2ce61fd6-8b07-45bb-aece-fa51b7471a3f · outbound

This paper cites Mujoco: A physics engine for model-based control.

Average-Reward Soft Actor-Critic Mujoco: A physics engine for model-based control

Reference 2009

Resolution
unresolved
no resolver link, observed 2026-08-10T20:13:57.667433Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T20:13:57.667433Z digest=sha256:8dfab008fb9028c767aa2f4359b47e6e2126dee2d9a1e4d0aec3fff97f69dcb3

Observation da48ec9a-95f8-4a39-876d-784452b9a9a8 · outbound

This paper cites ∞X k=1 r(st+k, at+k) − 1 β log π(at+k|st+k) π0(at+k|st+k) − θπ # , Qπ(st+1, at+1) = r(st+1, at+1) − θπ + E p,π.

Average-Reward Soft Actor-Critic ∞X k=1 r(st+k, at+k) − 1 β log π(at+k|st+k) π0(at+k|st+k) − θπ # , Qπ(st+1, at+1) = r(st+1, at+1) − θπ + E p,π

Reference 2010

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T20:13:58.641196Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-10T20:13:57.708275Z digest=sha256:2420567a77f16e942ba6f14b9cf09b3a1e2726d70a903a7a40f1a1d572550824

Observation 711b2ad6-1711-4dfa-8ed3-b693e82bc4ae · outbound

This paper cites Proximal Policy Optimization Algorithms.

Average-Reward Soft Actor-Critic Proximal Policy Optimization Algorithms

Reference 2016

Resolution
unresolved
no resolver link, observed 2026-08-10T20:13:57.618379Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T20:13:57.618379Z digest=sha256:b76e671ff2feb9829e538e26ee00a0e1539f205e51bad22159214b62cc200ef1

Observation caafb751-ac41-488d-84a2-853866f4051f · outbound

This paper cites Soft Actor-Critic Algorithms and Applications.

Average-Reward Soft Actor-Critic Soft Actor-Critic Algorithms and Applications

Reference 2017

Resolution
unresolved
no resolver link, observed 2026-08-10T20:13:57.506033Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T20:13:57.506033Z digest=sha256:c2ac8a092c1d2e41d06f3abd6f41dcb454fc4f99af37ab76755a8ab84b0acb1a

Observation b7e82671-4a27-4b1f-b938-ed127239bce1 · outbound

This paper cites RVI-SAC: Average Reward Off-Policy Deep Reinforcement Learning.

Average-Reward Soft Actor-Critic RVI-SAC: Average Reward Off-Policy Deep Reinforcement Learning

Reference 2018

Resolution
verified exact
local_arxiv, observed 2026-08-10T20:13:58.447879Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-10T20:13:57.513390Z digest=sha256:a9789a222048b79f14dd59342b4c38a8a0d4741089d7238c136c1434901e662c

Observation 94bd5da1-694f-4bd3-85f0-935f9e338448 · outbound

This paper cites A unified view of entropy-regularized Markov decision processes.

Average-Reward Soft Actor-Critic A unified view of entropy-regularized Markov decision processes

Reference 2019

Resolution
unresolved
no resolver link, observed 2026-08-10T20:13:57.593356Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T20:13:57.593356Z digest=sha256:6d75dfeae4a5df2afdcbbb574ce87cfa6b1bc98cff2dff5e06a3e78ec07759af

Observation 1d17c5b1-a7cd-4e07-989b-bc2d03398a94 · outbound

This paper cites Reinforcement Learning and Control as Probabilistic Inference: Tutorial and Review.

Average-Reward Soft Actor-Critic Reinforcement Learning and Control as Probabilistic Inference: Tutorial and Review

Reference 2020

Resolution
unresolved
no resolver link, observed 2026-08-10T20:13:57.558208Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T20:13:57.558208Z digest=sha256:4fcce5d7b26a399aa955db72060b1e5e440030bc2a8017e96c1e5901ee5a12b1

Observation 71ced4dc-16f7-4a88-a920-be7660fe61c2 · outbound

This paper cites Dueling network architectures for deep reinforcement learning.

Average-Reward Soft Actor-Critic Dueling network architectures for deep reinforcement learning

Reference 2021

Resolution
unresolved
no resolver link, observed 2026-08-10T20:13:57.676411Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T20:13:57.676411Z digest=sha256:99adcf802a0b03ef03fee23572dbcc9418eaeb9d6a9215ef366a0ef884fe4573

Observation 97ac2ec5-a628-4b86-a794-32b9c58f7522 · outbound

This paper cites Making Reinforcement Learning Work on Swimmer.

Average-Reward Soft Actor-Critic Making Reinforcement Learning Work on Swimmer

Reference 2022

Resolution
verified exact
local_arxiv, observed 2026-08-10T20:13:58.551293Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-10T20:13:57.499060Z digest=sha256:93e742d36a4708c44e74c5e1663343f2c423fcc9faa60e0d6a0700bd98e7db5a

Observation b4e59e9f-f86a-4533-bde0-f99391d668fa · outbound

This paper cites Prioritized Experience Replay.

Average-Reward Soft Actor-Critic Prioritized Experience Replay

Reference 2023

Resolution
unresolved
no resolver link, observed 2026-08-10T20:13:57.609321Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T20:13:57.609321Z digest=sha256:33a18a6939f0cf2f8ec2b9505305253591c5b1c801643d0a285e43a9660df593

Observation db8c2468-c875-4db1-b9ea-441555c33241 · outbound

This paper cites The dependence of effective planning horizon on model accuracy.

Average-Reward Soft Actor-Critic The dependence of effective planning horizon on model accuracy

Reference 2024

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T20:13:58.807736Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-10T20:13:57.528380Z digest=sha256:fa9ec01abc22c34edd85b8e311afb9740d23be83ec56601290d616770eaee854

Observation 4d84dc87-68c2-43ee-b53f-4d0c58b22abf · outbound

This paper cites What Matters In On-Policy Reinforcement Learning? A Large-Scale Empirical Study.

Average-Reward Soft Actor-Critic What Matters In On-Policy Reinforcement Learning? A Large-Scale Empirical Study

Reference 2025

Resolution
unresolved
no resolver link, observed 2026-08-10T20:13:57.486495Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T20:13:57.486495Z digest=sha256:c4b91b7babf99c7bd3eb4c65d3261f0c1844f0482f163d3f4444af417ba12732

Pith citing papers

Observation 2f999352-14a3-4199-b542-99f5bf01bf69 · inbound

Learning Adaptive Parameter Policies for Nonlinear Bayesian Filtering cites this paper.

Learning Adaptive Parameter Policies for Nonlinear Bayesian Filtering Average-Reward Soft Actor-Critic

Reference 31

Resolution
verified exact
arxiv_id, observed 2026-05-15T08:45:19.219042Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-05-15T08:41:52.119405Z digest=sha256:f8fd6c4e31437fac6fd5978c32492bc92a2c2f882d99e49c52402de483974b1b

Observation 3fb8d06e-83b7-44b8-bd36-ffba1af6ba03 · inbound

When Policy Entropy Constraint Fails: Preserving Diversity in Flow-based RLHF via Perceptual Entropy cites this paper.

When Policy Entropy Constraint Fails: Preserving Diversity in Flow-based RLHF via Perceptual Entropy Average-Reward Soft Actor-Critic

Reference 1

Resolution
metadata mismatch
arxiv_id, observed 2026-05-13T05:52:22.146444Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-05-13T05:50:13.653022Z digest=sha256:0e58f8ac3eeb2644e7e9af61633866eadf4205eddc88d253fe5830c70b5196de