Pith. sign in

Paper Citation Record · LEDGER

Uncertainty Weighted Actor-Critic for Offline Reinforcement Learning

As of 8 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 7 inbound Pith citation observations for arXiv:2105.08140.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2105.08140 v1

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 7 of 7 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-08T06:32:00.761636+00:00

measured 7 of 7 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-02T08:40:52.852816Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-06-29T08:43:15.128113Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation cd89333e-74bc-4e73-9bc8-3c96b01c2613 · inbound

Towards Efficient and Expressive Offline RL via Flow-Anchored Noise-conditioned Q-Learning cites this paper.

Towards Efficient and Expressive Offline RL via Flow-Anchored Noise-conditioned Q-Learning Uncertainty Weighted Actor-Critic for Offline Reinforcement Learning

Reference 33

Resolution
verified exact
arxiv_id, observed 2026-05-11T08:56:00.734820Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-05-10T16:25:25.739019Z digest=sha256:d01498be500df9827e1b7ba234d8152fef4a0eb6df9dadb2c7e9a9fb3afc5f40

Observation 9c4d26f6-1093-42de-b6f9-f4cb7be3b953 · inbound

Beyond Penalization: Diffusion-based Out-of-Distribution Detection and Selective Regularization in Offline Reinforcement Learning cites this paper.

Beyond Penalization: Diffusion-based Out-of-Distribution Detection and Selective Regularization in Offline Reinforcement Learning Uncertainty Weighted Actor-Critic for Offline Reinforcement Learning

Reference 10

Resolution
verified exact
arxiv_id, observed 2026-05-12T07:41:50.547178Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-05-12T02:17:25.783688Z digest=sha256:cee1531c2d7122c531da110ed5481e6ad017fe4e2054d97b0614c8de3ddf8ce1

Observation 1273e49c-7c95-4b3f-84d3-f000ba58e8fb · inbound

UNIQ: Conformal Calibration for Adaptive Conservatism in Offline Reinforcement Learning cites this paper.

UNIQ: Conformal Calibration for Adaptive Conservatism in Offline Reinforcement Learning Uncertainty Weighted Actor-Critic for Offline Reinforcement Learning

Reference 30

Resolution
verified exact
arxiv_id, observed 2026-06-29T08:43:15.129743Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-06-29T08:39:34.884726Z digest=sha256:cdbca352b7c14acdd6c4d91a96af0a8550e77a17c3c697336288016c61321a98

Observation b02f5f93-5f63-4162-94f2-c034f54f4e04 · inbound

Multi-Turn On-Policy Distillation with Prefix Replay cites this paper.

Multi-Turn On-Policy Distillation with Prefix Replay Uncertainty Weighted Actor-Critic for Offline Reinforcement Learning

Reference 176

Resolution
unresolved
no resolver link, observed 2026-07-11T13:53:36.775836Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-11T13:53:36.775836Z digest=sha256:3a1ea4c35069bc395df580375f9254b69052b60203b54c569653deca45a73e5f

Observation 3e3ed705-f1f9-4f08-9ccb-495cb2c68033 · inbound

Multi-Turn On-Policy Distillation with Prefix Replay cites this paper.

Multi-Turn On-Policy Distillation with Prefix Replay Uncertainty Weighted Actor-Critic for Offline Reinforcement Learning

Reference 177

Resolution
unresolved
no resolver link, observed 2026-08-02T08:40:52.852816Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-02T08:40:52.852816Z digest=sha256:d9c4340cef9dc3f20c3aeda26ccee93fdd63ab722154f9e86729745a6bb0082b

Observation a074f1fd-fa65-4776-a518-e65e3bfced2f · inbound

Collaborative Weighting with Pessimistic Critic for Mitigating Overestimation in Off-Policy Reinforcement Learning cites this paper.

Collaborative Weighting with Pessimistic Critic for Mitigating Overestimation in Off-Policy Reinforcement Learning Uncertainty Weighted Actor-Critic for Offline Reinforcement Learning

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-01T14:19:58.343808Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T14:19:58.343808Z digest=sha256:c09e1be423feeb873d55f3db4e477cc5ef12a0040cb3277f4f44477f0815ea1e

Observation 9b78a919-2db4-40d0-8f21-70201db89ad9 · inbound

Uncertainty-Guided LLM Semantic Augmentation for Heterogeneous Treatment Effect Estimation cites this paper.

Uncertainty-Guided LLM Semantic Augmentation for Heterogeneous Treatment Effect Estimation Uncertainty Weighted Actor-Critic for Offline Reinforcement Learning

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-01T12:41:27.456257Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T12:41:27.456257Z digest=sha256:7812feb0fda229a1ccd3f039aa4676bd7d5ea515351d231537bf9d22b7b500d7