Pith. sign in

Paper Citation Record · LEDGER

Soft Actor-Critic for Discrete Action Settings

As of 15 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 30 inbound Pith citation observations for arXiv:1910.07207.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
1910.07207 v2

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 30 of 30 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-14T06:32:32.682623+00:00

measured 30 of 30 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-12T18:52:14.128108Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-04T19:40:05.986446Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation 67f86b76-e5df-433a-9d5b-4d957e1c4d7f · inbound

AMAGO-2: Breaking the Multi-Task Barrier in Meta-Reinforcement Learning with Transformers cites this paper.

AMAGO-2: Breaking the Multi-Task Barrier in Meta-Reinforcement Learning with Transformers Soft Actor-Critic for Discrete Action Settings

Reference 76

Resolution
unresolved
no resolver link, observed 2026-08-12T18:52:14.128108Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T18:52:14.128108Z digest=sha256:0aafe014d08f420fcd3ff738f965c89eba4f51ad80f913c8eda7694a9d101b02

Observation ea992a80-c1bd-4a3e-9791-c4d206ad2e72 · inbound

Supervised Learning-enhanced Multi-Group Actor Critic for Live Stream Allocation in Feed cites this paper.

Supervised Learning-enhanced Multi-Group Actor Critic for Live Stream Allocation in Feed Soft Actor-Critic for Discrete Action Settings

Reference 2019

Resolution
unresolved
no resolver link, observed 2026-08-12T10:52:50.944478Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T10:52:50.944478Z digest=sha256:c5664145629138d39b07b56b922615491da6653155a49635f397fb12cadf6b31

Observation 6202d432-773c-4208-affb-58f1ba6d27dc · inbound

Dream to Drive with Predictive Individual World Model cites this paper.

Dream to Drive with Predictive Individual World Model Soft Actor-Critic for Discrete Action Settings

Reference 67

Resolution
unresolved
no resolver link, observed 2026-08-10T11:08:53.322150Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T11:08:53.322150Z digest=sha256:c3b8139e1882f381def03f3c16a0039725b364ace4187626d97bb4319dfd2ad3

Observation ebce658c-c843-4fdb-a824-5f93eff0308e · inbound

Decorrelated Soft Actor-Critic for Efficient Deep Reinforcement Learning cites this paper.

Decorrelated Soft Actor-Critic for Efficient Deep Reinforcement Learning Soft Actor-Critic for Discrete Action Settings

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-09T21:16:56.312627Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T21:16:56.312627Z digest=sha256:e3a3224abdd4a43aaccdd5ba2bf42053754c813d81dfd7a64b39e77d85fe7ef6

Observation 297a5d88-0833-4fa6-9b98-c16ed13bda4e · inbound

Behavior-Regularized Diffusion Policy Optimization for Offline Reinforcement Learning cites this paper.

Behavior-Regularized Diffusion Policy Optimization for Offline Reinforcement Learning Soft Actor-Critic for Discrete Action Settings

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-08T21:39:25.525912Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T21:39:25.525912Z digest=sha256:1433ed9ca2f44ff21c0e4d2f7db1b3faf50bb2a0802ed87ab441a29fe9df52fa

Observation 5e975619-5ca9-4778-966f-6b1b82faa63c · inbound

SLAC: Safe and Efficient Real-Robot Reinforcement Learning via Unsupervised Simulation Pre-Training cites this paper.

SLAC: Safe and Efficient Real-Robot Reinforcement Learning via Unsupervised Simulation Pre-Training Soft Actor-Critic for Discrete Action Settings

Reference 60

Resolution
unresolved
no resolver link, observed 2026-08-07T10:53:10.179732Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:53:10.179732Z digest=sha256:4c0bf225387301533d71a5b1d5b43ee508df98dead6d9b0a1f6ad1c00b0cd1f6

Observation c81fe141-3c17-4314-a79f-5993e0643127 · inbound

Hierarchical Learning-Enhanced MPC for Safe Crowd Navigation with Heterogeneous Constraints cites this paper.

Hierarchical Learning-Enhanced MPC for Safe Crowd Navigation with Heterogeneous Constraints Soft Actor-Critic for Discrete Action Settings

Reference 52

Resolution
malformed identifier
no resolver link, observed 2026-08-07T04:47:14.935311Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:47:14.935311Z digest=sha256:fd86f7e48c48d9e17c6fe60a52b624852770d3d739136c634e3342b6373d60fd

Observation 7b6c6c6b-ad19-4b2e-8a59-d6bf083e2595 · inbound

Goal-conditioned Hierarchical Reinforcement Learning for Sample-efficient and Safe Autonomous Driving at Intersections cites this paper.

Goal-conditioned Hierarchical Reinforcement Learning for Sample-efficient and Safe Autonomous Driving at Intersections Soft Actor-Critic for Discrete Action Settings

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-06T23:49:14.797853Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:49:14.797853Z digest=sha256:908d8146ce6b2c054da5c0b4b61a3b346faf7b87216568e6d7847f77bb83efcd

Observation 7a9a86bb-cec3-4ae3-a691-68ba02f6f199 · inbound

Multi-Agent Reinforcement Learning for Inverse Design in Photonic Integrated Circuits cites this paper.

Multi-Agent Reinforcement Learning for Inverse Design in Photonic Integrated Circuits Soft Actor-Critic for Discrete Action Settings

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-06T23:21:19.534851Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T23:21:19.534851Z digest=sha256:d757ae57210347b20b24d8d1b964830fa168bce608fc156e95e23f699f3f5b02

Observation af61d2fc-ea81-4b3b-a899-4dd8e4a604de · inbound

Learning To Communicate Over An Unknown Shared Network cites this paper.

Learning To Communicate Over An Unknown Shared Network Soft Actor-Critic for Discrete Action Settings

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-06T19:09:05.218895Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:09:05.218895Z digest=sha256:19f92c159800d42e9b8116161377ee7882892ec39daf09d8c745088453771299

Observation 9de1f464-3e53-4755-8767-6d82f0270712 · inbound

Efficient Multi-Task Reinforcement Learning with Cross-Task Policy Guidance cites this paper.

Efficient Multi-Task Reinforcement Learning with Cross-Task Policy Guidance Soft Actor-Critic for Discrete Action Settings

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-06T19:06:26.517932Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:06:26.517932Z digest=sha256:e1b50c6cede96615f0f2c5a44a073f8dd14f6b7d0f5d4b8efe079a1a5db634b1

Observation 50fb1378-d1f6-4cb1-9e71-21cdd1439884 · inbound

Relative Entropy Pathwise Policy Optimization cites this paper.

Relative Entropy Pathwise Policy Optimization Soft Actor-Critic for Discrete Action Settings

Reference 3

Resolution
metadata mismatch
arxiv_id, observed 2026-05-19T04:22:04.036090Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-05-19T04:19:15.018380Z digest=sha256:b3c1f882c1a9b10b8e590eda447988419030dde90584c49920d7bb7e9c594690

Observation 8a37a432-c78c-4615-a09a-5be87034d878 · inbound

Personalized Exercise Recommendation with Semantically-Grounded Knowledge Tracing cites this paper.

Personalized Exercise Recommendation with Semantically-Grounded Knowledge Tracing Soft Actor-Critic for Discrete Action Settings

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-06T17:24:29.206067Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T17:24:29.206067Z digest=sha256:69190e79cb08f54defea3fb089848e7a57524473b5ca0f85faaf4afd5f5ccc3c

Observation 45a43ee0-7ece-4498-97a2-fe9732ceeb15 · inbound

DOA: A Degeneracy Optimization Agent with Adaptive Pose Compensation Capability based on Deep Reinforcement Learning cites this paper.

DOA: A Degeneracy Optimization Agent with Adaptive Pose Compensation Capability based on Deep Reinforcement Learning Soft Actor-Critic for Discrete Action Settings

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-06T14:11:03.043402Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T14:11:03.043402Z digest=sha256:883c00f704d119fe269cdd48d99ec2174dcf1e2e2e3706ee001aaf700759f71c

Observation 5ea9f523-b5a7-47d8-9b2a-551b021f2ce8 · inbound

Dissecting Discrete Soft Actor-Critic: Limitations and Principled Alternatives cites this paper.

Dissecting Discrete Soft Actor-Critic: Limitations and Principled Alternatives Soft Actor-Critic for Discrete Action Settings

Reference 6

Resolution
metadata mismatch
arxiv_id, observed 2026-05-18T17:06:39.804243Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-05-18T17:05:36.100114Z digest=sha256:51ae30f67d6d783a3a25be5433fc45871e7345f2f73909c167485cd8ed290959

Observation fd073855-f9c5-42f0-b93a-14c5a3161913 · inbound

Generalizable Pareto-Optimal Offloading with Reinforcement Learning in Mobile Edge Computing cites this paper.

Generalizable Pareto-Optimal Offloading with Reinforcement Learning in Mobile Edge Computing Soft Actor-Critic for Discrete Action Settings

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-05T15:45:59.079116Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T15:45:59.079116Z digest=sha256:367ad29a73f89494605e6e2a03c78769612de461646e25afcacd79c492d383e7

Observation 6cd3e863-f66f-4f17-a3a8-b8df328c222a · inbound

MacroNav: Multi-Task Context Representation Learning Enables Efficient Navigation in Unknown Environments cites this paper.

MacroNav: Multi-Task Context Representation Learning Enables Efficient Navigation in Unknown Environments Soft Actor-Critic for Discrete Action Settings

Reference 23

Resolution
metadata mismatch
arxiv_id, observed 2026-05-18T01:10:33.889159Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-05-18T01:10:12.195443Z digest=sha256:22b027b6e84093d334af8c52f2993f8b286821a3b3eb2e76a36640c77ccdf85a

Observation 16501943-0d04-4c7d-ad33-cb1a8e2a7d5e · inbound

R2PS: Worst-Case Robust Real-Time Pursuit Strategies under Partial Observability cites this paper.

R2PS: Worst-Case Robust Real-Time Pursuit Strategies under Partial Observability Soft Actor-Critic for Discrete Action Settings

Reference 2

Resolution
metadata mismatch
arxiv_id, observed 2026-05-17T20:20:11.613626Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-05-17T20:19:24.841090Z digest=sha256:e73143816977565cce720db283bbd90c405512f637ec9f775427322e19d170ba

Observation b59e7892-1632-4b36-9df7-20360aefeeee · inbound

Emotion Entanglement and Bayesian Inference for Multi-Dimensional Emotion Understanding cites this paper.

Emotion Entanglement and Bayesian Inference for Multi-Dimensional Emotion Understanding Soft Actor-Critic for Discrete Action Settings

Reference 27

Resolution
unresolved
no resolver link, observed 2026-07-13T14:51:03.626575Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-13T14:51:03.626575Z digest=sha256:eb4a9a28b91bbf8e12cccae5f35c2968cbf471b4faf3613173c77396a2ec11df

Observation 3f2a1f31-9eaf-4298-bd9e-47f1e69b87a8 · inbound

Don't Forget the Critic: Value-Based Data Rehearsal for Multi-Cyclic Continual Reinforcement Learning cites this paper.

Don't Forget the Critic: Value-Based Data Rehearsal for Multi-Cyclic Continual Reinforcement Learning Soft Actor-Critic for Discrete Action Settings

Reference 9

Resolution
verified exact
arxiv_id, observed 2026-05-22T07:16:12.944273Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-05-22T07:15:32.024198Z digest=sha256:a138440c3b2acf562e66694e21d7105c012cabf488c0cbe767392247aa86f07f

Observation 4ce9cad1-9684-465b-a774-bee3b9e4f106 · inbound

Retry Policy Gradients in Continuous Action Spaces cites this paper.

Retry Policy Gradients in Continuous Action Spaces Soft Actor-Critic for Discrete Action Settings

Reference 2

Resolution
metadata mismatch
arxiv_id, observed 2026-07-02T12:56:57.023856Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-06-28T01:44:37.492572Z digest=sha256:9c6eea1d4fec70865e6e38d5479caca3bf6276bf66b2160752cad7855811f500

Observation caf878c9-6fa1-4d62-9296-9b5c8914068b · inbound

Your GFlowNet Secretly Learns an Optimal Transport Plan cites this paper.

Your GFlowNet Secretly Learns an Optimal Transport Plan Soft Actor-Critic for Discrete Action Settings

Reference 81

Resolution
verified exact
arxiv_id, observed 2026-07-02T12:06:55.527589Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-06-28T02:35:08.323804Z digest=sha256:b9ad1d345450965d802a0d00959c459ee1f1e41e6a3f6bebd82c3168daa0909c

Observation 21e4be6e-28a0-418b-a2a5-a72dab9f0226 · inbound

Back to the Familiar Future: Failure Recovery for VLA Policies via Pre-Imagined Milestone Selection cites this paper.

Back to the Familiar Future: Failure Recovery for VLA Policies via Pre-Imagined Milestone Selection Soft Actor-Critic for Discrete Action Settings

Reference 29

Resolution
verified exact
arxiv_id, observed 2026-07-03T01:17:30.764532Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-06-27T16:42:49.740745Z digest=sha256:d56cfb57587411700b5cfefb15933594679350f59c88de5261d49606879cfcd5

Observation bfa06eaf-6f94-443d-af2f-9c75b1f544cf · inbound

Event-Driven Reinforcement Learning Enables Long-Horizon Control in Semiconductor Fabrication cites this paper.

Event-Driven Reinforcement Learning Enables Long-Horizon Control in Semiconductor Fabrication Soft Actor-Critic for Discrete Action Settings

Reference 60

Resolution
metadata mismatch
arxiv_id, observed 2026-07-03T04:07:36.639368Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-06-27T14:14:38.297507Z digest=sha256:b2e2f36b67ddf8100cd3d284d8103f65255d8a7e7e36d7786107cab51e823f93

Observation 1ac1056f-70e0-4e87-92f9-c9aba62a0250 · inbound

Modularized Reinforcement Learning on LLMs: From MDP Creation to Exploration and Learning cites this paper.

Modularized Reinforcement Learning on LLMs: From MDP Creation to Exploration and Learning Soft Actor-Critic for Discrete Action Settings

Reference 31

Resolution
verified exact
arxiv_id, observed 2026-07-04T07:59:40.643542Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-06-26T12:15:08.304150Z digest=sha256:22898bc8be94ba387497deb9352c28787cffd664c434c66ff199f4f082190e55

Observation 18408309-491c-402f-9e55-7c8bf0147e3e · inbound

FactorLibrary: From Polynomials to Circuits via Recursive Subgoals cites this paper.

FactorLibrary: From Polynomials to Circuits via Recursive Subgoals Soft Actor-Critic for Discrete Action Settings

Reference 51

Resolution
metadata mismatch
arxiv_id, observed 2026-07-04T19:40:05.988216Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-06-25T21:12:22.718409Z digest=sha256:80307025561effc472977d8c25883ed17a958701ac12a8cf92d673fa64d02a49

Observation cc54dc00-345d-4228-9582-bdbcec8f327d · inbound

ACPO: Agent-Chained Policy Optimization for Multi-Agent Reinforcement Learning cites this paper.

ACPO: Agent-Chained Policy Optimization for Multi-Agent Reinforcement Learning Soft Actor-Critic for Discrete Action Settings

Reference 2020

Resolution
unresolved
no resolver link, observed 2026-08-02T09:39:05.935795Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T09:39:05.935795Z digest=sha256:c864fb822de7f37fbe35c9a17b2482bec313995e67cb00a4248b9a71fc7b828c

Observation 406d8430-2d0c-4105-9d0e-a070b86eef4b · inbound

Practical Graph Optimisation and AI-Driven Models for Active Directory Security Hardening cites this paper.

Practical Graph Optimisation and AI-Driven Models for Active Directory Security Hardening Soft Actor-Critic for Discrete Action Settings

Reference 345

Resolution
unresolved
no resolver link, observed 2026-08-01T06:15:18.525217Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T06:15:18.525217Z digest=sha256:aa3b22866a4ca3335a237492324a6acde22826b8654dc6772bb6a2d7c8fdec2b

Observation 5d729623-6732-4a19-b31b-0a77d438c5f2 · inbound

Deep Reinforcement Learning: From First Principles to Reasoning Models cites this paper.

Deep Reinforcement Learning: From First Principles to Reasoning Models Soft Actor-Critic for Discrete Action Settings

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-04T01:16:00.255389Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T01:16:00.255389Z digest=sha256:13b54b7b0ed19a46cea2f124c302cf03c3a1828d14dae81abfbc5fd0157bba34

Observation 641735f2-9b57-48e6-98be-efa0f519e538 · inbound

Efficient Real-World Online Reinforcement Learning for Robot Manipulation via Centralized Training and Critic Decomposition cites this paper.

Efficient Real-World Online Reinforcement Learning for Robot Manipulation via Centralized Training and Critic Decomposition Soft Actor-Critic for Discrete Action Settings

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-11T11:20:20.061953Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T11:20:20.061953Z digest=sha256:7d877c8fdc8fe2d7d1189105784313e62f1de09d1de479b88ad9303e78a9facd