Pith. sign in

Paper Citation Record · LEDGER

Uncertainty Weighted Actor-Critic for Offline Reinforcement Learning

As of 9 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 7 inbound Pith citation observations for arXiv:2105.08140.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2105.08140 v1

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 7 of 7 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-09T06:31:02.800959+00:00

measured 7 of 7 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-02T08:40:52.852816Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-06-29T08:43:15.128113Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation cd89333e-74bc-4e73-9bc8-3c96b01c2613 · inbound

Towards Efficient and Expressive Offline RL via Flow-Anchored Noise-conditioned Q-Learning cites this paper.

Towards Efficient and Expressive Offline RL via Flow-Anchored Noise-conditioned Q-Learning Uncertainty Weighted Actor-Critic for Offline Reinforcement Learning

Reference 33

Resolution
verified exact
arxiv_id, observed 2026-05-11T08:56:00.734820Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-05-10T16:25:25.739019Z digest=sha256:1100f1517f38c0863325fa4ea32e11675585e9f5cb0b50eba61d78a93b9954c1

Observation 9c4d26f6-1093-42de-b6f9-f4cb7be3b953 · inbound

Beyond Penalization: Diffusion-based Out-of-Distribution Detection and Selective Regularization in Offline Reinforcement Learning cites this paper.

Beyond Penalization: Diffusion-based Out-of-Distribution Detection and Selective Regularization in Offline Reinforcement Learning Uncertainty Weighted Actor-Critic for Offline Reinforcement Learning

Reference 10

Resolution
verified exact
arxiv_id, observed 2026-05-12T07:41:50.547178Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-05-12T02:17:25.783688Z digest=sha256:c95eed06617d1525e86cea1c20fa68c7e559427ec30f1b2ce94d266dc0cbbd0e

Observation 1273e49c-7c95-4b3f-84d3-f000ba58e8fb · inbound

UNIQ: Conformal Calibration for Adaptive Conservatism in Offline Reinforcement Learning cites this paper.

UNIQ: Conformal Calibration for Adaptive Conservatism in Offline Reinforcement Learning Uncertainty Weighted Actor-Critic for Offline Reinforcement Learning

Reference 30

Resolution
verified exact
arxiv_id, observed 2026-06-29T08:43:15.129743Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-06-29T08:39:34.884726Z digest=sha256:f6dee8fe8eeb533a0a0e5a113f13e7bf2bd4689b9eebe4c22dd08cf6d00e9dc5

Observation b02f5f93-5f63-4162-94f2-c034f54f4e04 · inbound

Multi-Turn On-Policy Distillation with Prefix Replay cites this paper.

Multi-Turn On-Policy Distillation with Prefix Replay Uncertainty Weighted Actor-Critic for Offline Reinforcement Learning

Reference 176

Resolution
unresolved
no resolver link, observed 2026-07-11T13:53:36.775836Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-11T13:53:36.775836Z digest=sha256:3a1ea4c35069bc395df580375f9254b69052b60203b54c569653deca45a73e5f

Observation 3e3ed705-f1f9-4f08-9ccb-495cb2c68033 · inbound

Multi-Turn On-Policy Distillation with Prefix Replay cites this paper.

Multi-Turn On-Policy Distillation with Prefix Replay Uncertainty Weighted Actor-Critic for Offline Reinforcement Learning

Reference 177

Resolution
unresolved
no resolver link, observed 2026-08-02T08:40:52.852816Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-02T08:40:52.852816Z digest=sha256:d9c4340cef9dc3f20c3aeda26ccee93fdd63ab722154f9e86729745a6bb0082b

Observation a074f1fd-fa65-4776-a518-e65e3bfced2f · inbound

Collaborative Weighting with Pessimistic Critic for Mitigating Overestimation in Off-Policy Reinforcement Learning cites this paper.

Collaborative Weighting with Pessimistic Critic for Mitigating Overestimation in Off-Policy Reinforcement Learning Uncertainty Weighted Actor-Critic for Offline Reinforcement Learning

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-01T14:19:58.343808Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T14:19:58.343808Z digest=sha256:c09e1be423feeb873d55f3db4e477cc5ef12a0040cb3277f4f44477f0815ea1e

Observation 9b78a919-2db4-40d0-8f21-70201db89ad9 · inbound

Uncertainty-Guided LLM Semantic Augmentation for Heterogeneous Treatment Effect Estimation cites this paper.

Uncertainty-Guided LLM Semantic Augmentation for Heterogeneous Treatment Effect Estimation Uncertainty Weighted Actor-Critic for Offline Reinforcement Learning

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-01T12:41:27.456257Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T12:41:27.456257Z digest=sha256:d4233c0106f48159af90fef2b86ddb27de0127d60b793f88103f6ed13f128eb3