Pith. sign in

Paper Citation Record · LEDGER

Language Reward Modulation for Pretraining Reinforcement Learning

As of 10 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 7 inbound Pith citation observations for arXiv:2308.12270.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2308.12270 v1

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 7 of 7 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-10T06:31:04.303077+00:00

measured 7 of 7 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-10T11:39:56.118369Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-03T12:38:07.182042Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation b5079cf3-34dc-4d7f-bafa-83893cb2a5af · inbound

Improving Vision-Language-Action Model with Online Reinforcement Learning cites this paper.

Improving Vision-Language-Action Model with Online Reinforcement Learning Language Reward Modulation for Pretraining Reinforcement Learning

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-10T11:39:56.118369Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T11:39:56.118369Z digest=sha256:9607c54d0fe804a4b148b0f66b91fb9682d32326842a10aab0aac7948c19460e

Observation bda703ef-bb9d-483f-997b-2915c27a2829 · inbound

Preference VLM: Leveraging VLMs for Scalable Preference-Based Reinforcement Learning cites this paper.

Preference VLM: Leveraging VLMs for Scalable Preference-Based Reinforcement Learning Language Reward Modulation for Pretraining Reinforcement Learning

Reference 2004

Resolution
unresolved
no resolver link, observed 2026-08-09T14:52:27.228163Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T14:52:27.228163Z digest=sha256:9d6088116e8297715592cc3a44704ada4f2f42b6db6aa351e2081a3a2a4c44a1

Observation f9ad4aef-6a45-4b4e-a471-bad9b521ad4b · inbound

World Model Implanting for Test-time Adaptation of Embodied Agents cites this paper.

World Model Implanting for Test-time Adaptation of Embodied Agents Language Reward Modulation for Pretraining Reinforcement Learning

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-05T10:35:25.217617Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T10:35:25.217617Z digest=sha256:c9fb755988732ebda31f62ed04a9747caf77ba2d1adf26aaa76ad0d533a7be6b

Observation 1408bd9c-6cc0-4a5c-976e-adf6f535fd0e · inbound

Exploratory Retrieval-Augmented Planning For Continual Embodied Instruction Following cites this paper.

Exploratory Retrieval-Augmented Planning For Continual Embodied Instruction Following Language Reward Modulation for Pretraining Reinforcement Learning

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-04T21:06:11.957361Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T21:06:11.957361Z digest=sha256:768a40d929d38a6decaa709e53e45efb9d738f556652d078acdbb73adec777af

Observation 64bcc828-ea4b-4093-93a6-e0bc5277943e · inbound

Robometer: Scaling General-Purpose Robotic Reward Models via Trajectory Comparisons cites this paper.

Robometer: Scaling General-Purpose Robotic Reward Models via Trajectory Comparisons Language Reward Modulation for Pretraining Reinforcement Learning

Reference 46

Resolution
verified exact
arxiv_id, observed 2026-05-15T18:00:12.601521Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-15T17:59:52.365630Z digest=sha256:ee9c6f887e642256dfac8387ad2a34d64ce798b894dfc12f4d9c37e7877b1f5c

Observation c7651e20-adac-4601-9eb8-77299d18c095 · inbound

CoRe: Combined Rewards with Vision-Language Model Feedback for Preference-Aligned Reinforcement Learning cites this paper.

CoRe: Combined Rewards with Vision-Language Model Feedback for Preference-Aligned Reinforcement Learning Language Reward Modulation for Pretraining Reinforcement Learning

Reference 70

Resolution
verified exact
arxiv_id, observed 2026-07-03T12:38:07.183723Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-07-03T12:29:40.316336Z digest=sha256:c2adf3c0eb8b0d88f19a0ead1a273fd4bf892ec6a4ba1ee58b9667eead0071d2

Observation 5b32704d-960c-4be5-b95b-8601d7dc8a62 · inbound

Weights or Skills? A Survey of Robot-Learning Techniques: from Action-Predicting Weights to Robots that Write their Own Skills cites this paper.

Weights or Skills? A Survey of Robot-Learning Techniques: from Action-Predicting Weights to Robots that Write their Own Skills Language Reward Modulation for Pretraining Reinforcement Learning

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-04T19:45:26.283153Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T19:45:26.283153Z digest=sha256:d2aeef14f14c14ed132c65188fdbf12de8155073ea2197b1d0c90c4923fce44c