Pith. sign in

Paper Citation Record · LEDGER

Dealing with Sparse Rewards in Reinforcement Learning

As of 14 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 8 inbound Pith citation observations for arXiv:1910.09281.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
1910.09281 v2

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 8 of 8 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-13T06:32:02.005865+00:00

measured 8 of 8 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-11T15:45:23.242099Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-04T15:29:55.594907Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation 337f6a73-6a5b-45ff-b647-f0f198d752fb · inbound

Noise-Resilient Symbolic Regression with Dynamic Gating Reinforcement Learning cites this paper.

Noise-Resilient Symbolic Regression with Dynamic Gating Reinforcement Learning Dealing with Sparse Rewards in Reinforcement Learning

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-10T22:40:16.221035Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T22:40:16.221035Z digest=sha256:75227247492e968a0c0cc40ce83eb17a14049c16e8f1b947dfea1f2e284539f9

Observation bfb38035-4bbb-48cc-85b1-f078aa19bae0 · inbound

From Sparse to Dense: Toddler-inspired Reward Transition in Goal-Oriented Reinforcement Learning cites this paper.

From Sparse to Dense: Toddler-inspired Reward Transition in Goal-Oriented Reinforcement Learning Dealing with Sparse Rewards in Reinforcement Learning

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-10T04:37:01.861960Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T04:37:01.861960Z digest=sha256:cc5de91c21125a09500a517dec2dc1572d847d31a055206ec8014879fa31e067

Observation 3dcc38c7-143c-4d6c-9845-a260871e95c9 · inbound

VARD: Efficient and Dense Fine-Tuning for Diffusion Models with Value-based RL cites this paper.

VARD: Efficient and Dense Fine-Tuning for Diffusion Models with Value-based RL Dealing with Sparse Rewards in Reinforcement Learning

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-07T15:18:52.713790Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:18:52.713790Z digest=sha256:d687e62ffafa678b337d869d81a753c95ba8f8fcb2d11a30fafbd40540c93bfe

Observation dc2e7ed3-6589-478f-8935-61ef15aaacef · inbound

SCAR: Shapley Credit Assignment for More Efficient RLHF cites this paper.

SCAR: Shapley Credit Assignment for More Efficient RLHF Dealing with Sparse Rewards in Reinforcement Learning

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-07T14:02:49.225204Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:02:49.225204Z digest=sha256:306fe189dfa52c0b55606000ab1bfbd0147b2e3d38c240756e007afa3f31dd7f

Observation ce1b4e95-c112-4510-8d56-4328a4dbab63 · inbound

Do Less, Achieve More: Do We Need Every-Step Optimization for RL Fine-tuning of Diffusion Models? cites this paper.

Do Less, Achieve More: Do We Need Every-Step Optimization for RL Fine-tuning of Diffusion Models? Dealing with Sparse Rewards in Reinforcement Learning

Reference 13

Resolution
verified exact
arxiv_id, observed 2026-05-20T19:48:57.311552Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-05-20T19:48:17.049547Z digest=sha256:0e9543e407ad659e3df41a80e4d8b23eb792a06f7cdb672f4bf38395a8861eff

Observation 824aeed7-4275-48ab-ba4a-539b19db6c86 · inbound

Mesh-RL: Coupled subgrid reinforcement learning cites this paper.

Mesh-RL: Coupled subgrid reinforcement learning Dealing with Sparse Rewards in Reinforcement Learning

Reference 28

Resolution
verified exact
arxiv_id, observed 2026-07-04T15:29:55.596964Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-06-26T01:37:31.834106Z digest=sha256:0e2bae746553e2e98cce1f77d0d39b91b40dd804b931c799dd628c269e593988

Observation d234f2db-73f8-4d88-8010-ee730468adda · inbound

Learning Gait-Aware Quadruped Locomotion with Temporal Logic Specifications cites this paper.

Learning Gait-Aware Quadruped Locomotion with Temporal Logic Specifications Dealing with Sparse Rewards in Reinforcement Learning

Reference 30

Resolution
verified exact
arxiv_id, observed 2026-07-02T12:06:55.368456Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-07-02T11:59:04.103002Z digest=sha256:493d35ea7893171fd45fafbcad25df7b6edc33d6938d7fa8e421b951c96aae12

Observation 11b4206c-2cfd-4ab9-a6c8-5e136347aefd · inbound

STAIR: Effective Incident Response Using an End-to-End Agentic Planning Framework cites this paper.

STAIR: Effective Incident Response Using an End-to-End Agentic Planning Framework Dealing with Sparse Rewards in Reinforcement Learning

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-11T15:45:23.242099Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T15:45:23.242099Z digest=sha256:6697ff18ee7976abc4d2e5eb7e23734f7ff5d304d88b01febdd2d2a456c19973