Pith. sign in

Paper Citation Record · LEDGER

Dealing with Sparse Rewards in Reinforcement Learning

As of 13 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 8 inbound Pith citation observations for arXiv:1910.09281.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
1910.09281 v2

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 8 of 8 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-13T06:32:02.005865+00:00

measured 8 of 8 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-11T15:45:23.242099Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-04T15:29:55.594907Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation 337f6a73-6a5b-45ff-b647-f0f198d752fb · inbound

Noise-Resilient Symbolic Regression with Dynamic Gating Reinforcement Learning cites this paper.

Noise-Resilient Symbolic Regression with Dynamic Gating Reinforcement Learning Dealing with Sparse Rewards in Reinforcement Learning

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-10T22:40:16.221035Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T22:40:16.221035Z digest=sha256:fad60d4f8af4ccbf6c8d325a61ccd9c0204c79716eb82465a03f09c7441cdc53

Observation bfb38035-4bbb-48cc-85b1-f078aa19bae0 · inbound

From Sparse to Dense: Toddler-inspired Reward Transition in Goal-Oriented Reinforcement Learning cites this paper.

From Sparse to Dense: Toddler-inspired Reward Transition in Goal-Oriented Reinforcement Learning Dealing with Sparse Rewards in Reinforcement Learning

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-10T04:37:01.861960Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T04:37:01.861960Z digest=sha256:0ff9024bec53662dd3aa41f8bdc3b2df5e912cecdbe91ace06197eb9cd24d1a7

Observation 3dcc38c7-143c-4d6c-9845-a260871e95c9 · inbound

VARD: Efficient and Dense Fine-Tuning for Diffusion Models with Value-based RL cites this paper.

VARD: Efficient and Dense Fine-Tuning for Diffusion Models with Value-based RL Dealing with Sparse Rewards in Reinforcement Learning

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-07T15:18:52.713790Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:18:52.713790Z digest=sha256:e6118ce88a054f5fe247f4a0344ef3f0f91149a660171e843d69e0fa2f0e5009

Observation dc2e7ed3-6589-478f-8935-61ef15aaacef · inbound

SCAR: Shapley Credit Assignment for More Efficient RLHF cites this paper.

SCAR: Shapley Credit Assignment for More Efficient RLHF Dealing with Sparse Rewards in Reinforcement Learning

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-07T14:02:49.225204Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:02:49.225204Z digest=sha256:fea426da04c1fe757fb53b0c3d0f0d6b3741aa5752b6b98417cfef5e47e0a008

Observation ce1b4e95-c112-4510-8d56-4328a4dbab63 · inbound

Do Less, Achieve More: Do We Need Every-Step Optimization for RL Fine-tuning of Diffusion Models? cites this paper.

Do Less, Achieve More: Do We Need Every-Step Optimization for RL Fine-tuning of Diffusion Models? Dealing with Sparse Rewards in Reinforcement Learning

Reference 13

Resolution
verified exact
arxiv_id, observed 2026-05-20T19:48:57.311552Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-05-20T19:48:17.049547Z digest=sha256:7fce5a81f8df38909f467b64472c07c6ec7fdb09fc13fc1a360177338d8590ab

Observation 824aeed7-4275-48ab-ba4a-539b19db6c86 · inbound

Mesh-RL: Coupled subgrid reinforcement learning cites this paper.

Mesh-RL: Coupled subgrid reinforcement learning Dealing with Sparse Rewards in Reinforcement Learning

Reference 28

Resolution
verified exact
arxiv_id, observed 2026-07-04T15:29:55.596964Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-06-26T01:37:31.834106Z digest=sha256:7fadc7ec09e64b0bf57948ab58aa70be91549564a4dd9af87647ce87de1bf898

Observation d234f2db-73f8-4d88-8010-ee730468adda · inbound

Learning Gait-Aware Quadruped Locomotion with Temporal Logic Specifications cites this paper.

Learning Gait-Aware Quadruped Locomotion with Temporal Logic Specifications Dealing with Sparse Rewards in Reinforcement Learning

Reference 30

Resolution
verified exact
arxiv_id, observed 2026-07-02T12:06:55.368456Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-07-02T11:59:04.103002Z digest=sha256:610d9edd4f075714ff88e3f179c2fe41de0cfb888620886c291dce9cf051dc2a

Observation 11b4206c-2cfd-4ab9-a6c8-5e136347aefd · inbound

STAIR: Effective Incident Response Using an End-to-End Agentic Planning Framework cites this paper.

STAIR: Effective Incident Response Using an End-to-End Agentic Planning Framework Dealing with Sparse Rewards in Reinforcement Learning

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-11T15:45:23.242099Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T15:45:23.242099Z digest=sha256:8b40eb49b0237376da36cfb5df2d9118f3c845fd96acb04de4fecc24a1fdac53