Pith. sign in

Paper Citation Record · LEDGER

Reward-Conditioned Policies

As of 14 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 21 inbound Pith citation observations for arXiv:1912.13465.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
1912.13465 v1

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 21 of 21 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-14T06:32:32.682623+00:00

measured 21 of 21 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-11T15:12:40.476298Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-04T16:39:58.107559Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation ae86ddc8-8c94-4f20-907b-78862ec38bcf · inbound

Decision Transformer: Reinforcement Learning via Sequence Modeling cites this paper.

Decision Transformer: Reinforcement Learning via Sequence Modeling Reward-Conditioned Policies

Reference 44

Resolution
verified exact
arxiv_id, observed 2026-05-18T15:11:11.125644Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-05-18T15:11:11.056013Z digest=sha256:f3898174223fa91e52a2e514df7817487ed576826e5d09a88d0755f27633a0d6

Observation b9a32467-594f-4c1b-ae53-362f28907590 · inbound

Is Conditional Generative Modeling all you need for Decision-Making? cites this paper.

Is Conditional Generative Modeling all you need for Decision-Making? Reward-Conditioned Policies

Reference 208

Resolution
metadata mismatch
arxiv_id, observed 2026-05-15T15:35:10.839866Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-05-15T15:35:10.593969Z digest=sha256:046765f89f4e1ff37675b77c007928c2eb7bc97ab96721cd70e39b7b0ebb496c

Observation d1f2439e-4514-4507-86fd-7b8d5b984794 · inbound

Are Expressive Models Truly Necessary for Offline RL? cites this paper.

Are Expressive Models Truly Necessary for Offline RL? Reward-Conditioned Policies

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-11T15:12:40.476298Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T15:12:40.476298Z digest=sha256:00c926d5b3b924467dc685403e9d6aee78a1576a9d844a1c0b6bd79daf895722

Observation 0124c1ed-df2a-44cc-8b23-59c16ca05f04 · inbound

A Provable Approach for End-to-End Safe Reinforcement Learning cites this paper.

A Provable Approach for End-to-End Safe Reinforcement Learning Reward-Conditioned Policies

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-07T13:30:22.843583Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:30:22.843583Z digest=sha256:4afc73d51205faf72b63ef9fc551e6fc564c145a89b21841552b06fc3ae51475

Observation b68cdd13-872c-4b13-9aac-37ecd00af636 · inbound

Diffusion Guidance Is a Controllable Policy Improvement Operator cites this paper.

Diffusion Guidance Is a Controllable Policy Improvement Operator Reward-Conditioned Policies

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-07T12:56:28.274947Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:56:28.274947Z digest=sha256:51f516153699964e34efef6ce6e5a9d395a0617c83e44698151f85230bb69bf4

Observation 6c5875f2-64b1-4600-ae59-09958a242422 · inbound

How to Provably Improve Return Conditioned Supervised Learning? cites this paper.

How to Provably Improve Return Conditioned Supervised Learning? Reward-Conditioned Policies

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-07T05:19:06.527500Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:19:06.527500Z digest=sha256:3bdba05f4fc1c6d8fda3f1a28b1ec0f7d6d75a2af778128d79af59e6d1800aef

Observation 0e7ef5bb-d240-44de-ad57-dd99ce411702 · inbound

Behavioral Exploration: Learning to Explore via In-Context Adaptation cites this paper.

Behavioral Exploration: Learning to Explore via In-Context Adaptation Reward-Conditioned Policies

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-06T18:15:45.032209Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:15:45.032209Z digest=sha256:0c9b4f7a587244b1d1fee7195a9a47fbb8dcdb0ca6b0cd487cbd0cdc0f619a1f

Observation 62d58338-dc62-4d35-90de-29ccdc81b61c · inbound

EBaReT: Expert-guided Bag Reward Transformer for Auto Bidding cites this paper.

EBaReT: Expert-guided Bag Reward Transformer for Auto Bidding Reward-Conditioned Policies

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-06T15:22:49.982393Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:22:49.982393Z digest=sha256:d843f208acbff0e6fd0ee5a0195b65960ac9c360472375b1ed228ba2a3e917d0

Observation a8c825fa-143b-4850-afc9-d45175fd01bb · inbound

Generative Sequential Notification Optimization via Multi-Objective Decision Transformers cites this paper.

Generative Sequential Notification Optimization via Multi-Objective Decision Transformers Reward-Conditioned Policies

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-05T11:39:57.863900Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T11:39:57.863900Z digest=sha256:8a2c05ab0cff698bfb184bf57cd5574677139957be55657c4de84899c307f919

Observation 334c2d12-73fa-4d3b-9df4-d5254ff8fd4c · inbound

$\pi^{*}_{0.6}$: a VLA That Learns From Experience cites this paper.

$\pi^{*}_{0.6}$: a VLA That Learns From Experience Reward-Conditioned Policies

Reference 49

Resolution
verified exact
arxiv_id, observed 2026-05-12T10:34:59.416893Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-05-12T10:34:59.134604Z digest=sha256:ace54cbae63328ae4ee3e068d21e675744f9bc91fa08950f2d5253a10f91cdab

Observation 48ede48a-9dcb-4b78-b6c7-64d207282156 · inbound

RISE: Self-Improving Robot Policy with Compositional World Model cites this paper.

RISE: Self-Improving Robot Policy with Compositional World Model Reward-Conditioned Policies

Reference 48

Resolution
verified exact
arxiv_id, observed 2026-05-16T02:30:31.951919Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-05-16T02:28:37.997148Z digest=sha256:699f49d7526baeb9ed7b4f453ce4f16e2dac0b59e723247722d3cf81d1aed6df

Observation 7305dd02-e6d0-4416-8ef4-1330b44edccf · inbound

QHyer: Q-conditioned Hybrid Attention-mamba Transformer for Offline Goal-conditioned RL cites this paper.

QHyer: Q-conditioned Hybrid Attention-mamba Transformer for Offline Goal-conditioned RL Reward-Conditioned Policies

Reference 77

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T04:30:58.404146Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-05-11T01:17:48.643521Z digest=sha256:b9f79392b6e939bef699a2785adee9417541a13415494d3856b60b97b5318b42

Observation 206569dd-0496-457f-a49c-c357571fcf4e · inbound

When Evidence is Sparse: Weakly Supervised Early Failure Alerting in Dialogs and LLM-Agent Trajectories cites this paper.

When Evidence is Sparse: Weakly Supervised Early Failure Alerting in Dialogs and LLM-Agent Trajectories Reward-Conditioned Policies

Reference 4

Resolution
metadata mismatch
arxiv_id, observed 2026-07-02T08:26:48.331761Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-06-28T06:02:44.836106Z digest=sha256:fa1e6b8ba7181edcaba846e20c2a1bf70c1ad9daccf0f809705f4df4784f64c3

Observation cec724fc-b32a-4f56-ad42-1c676b21e041 · inbound

Neuro-Symbolic Injection of LTLf Constraints in Autoregressive Reinforcement Learning Policies cites this paper.

Neuro-Symbolic Injection of LTLf Constraints in Autoregressive Reinforcement Learning Policies Reward-Conditioned Policies

Reference 16

Resolution
verified exact
arxiv_id, observed 2026-07-02T21:57:25.919416Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-06-27T19:23:32.966503Z digest=sha256:aad0866ea050bbdcae9b2c3289e6243045035bf8c11effd168b4e1f7cdec47c1

Observation 907118e1-1e59-4f08-95b3-58b922d30a9c · inbound

Test-Time Gradient Guidance of Flow Policies in Reinforcement Learning cites this paper.

Test-Time Gradient Guidance of Flow Policies in Reinforcement Learning Reward-Conditioned Policies

Reference 30

Resolution
verified exact
arxiv_id, observed 2026-07-03T04:17:36.855518Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-06-27T14:05:01.073951Z digest=sha256:770310eddd99fba2061a48d1f26b1f3d30308b6755df5fa2c9d5fd5db66a73d8

Observation 82d3c53f-7b78-4f99-b6e0-775de5af61fd · inbound

FlowR2A: Learning Reward-to-Action Distribution for Multimodal Driving Planning cites this paper.

FlowR2A: Learning Reward-to-Action Distribution for Multimodal Driving Planning Reward-Conditioned Policies

Reference 25

Resolution
verified exact
arxiv_id, observed 2026-07-04T16:39:58.109405Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-06-26T00:19:49.473294Z digest=sha256:d4f06750b6a5edf7ad112112c5f6bda053e2f1215a67fcee195e1415dcaed9c3

Observation 811a19ca-7731-4ab8-ac6a-f0cc53ddadf3 · inbound

Freeform Preference Learning for Robotic Manipulation cites this paper.

Freeform Preference Learning for Robotic Manipulation Reward-Conditioned Policies

Reference 20

Resolution
verified exact
arxiv_id, observed 2026-07-01T10:55:41.968818Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-07-01T04:58:28.971536Z digest=sha256:38a60fbbd960e38156fb0f3c76c021c1224e69e132fc90df3fcbc1eecc4f1708

Observation 1e258ca9-2236-406d-a98f-3f22e7ff378c · inbound

Freeform Preference Learning for Robotic Manipulation cites this paper.

Freeform Preference Learning for Robotic Manipulation Reward-Conditioned Policies

Reference 20

Resolution
unresolved
no resolver link, observed 2026-07-14T16:55:18.028851Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T16:55:18.028851Z digest=sha256:53915e5a0f7b9b1db4a0105fc3102a686c5a01966b6a73d2608a3e52e9da2f62

Observation 09de577c-114a-4c8a-97bf-3d2defe065d2 · inbound

Freeform Preference Learning for Robotic Manipulation cites this paper.

Freeform Preference Learning for Robotic Manipulation Reward-Conditioned Policies

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-04T02:32:20.187685Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T02:32:20.187685Z digest=sha256:c1cc69180a25515ef08891e95f91eb94e83e545330e6c7e7bf3ad61c5463b1fc

Observation c4c07eb1-5087-43d5-bfda-0fb8f9e1680a · inbound

SwanTale: Unified Multi-Speaker Speech and Audio Generation for Instruct and Zero-Shot Tasks cites this paper.

SwanTale: Unified Multi-Speaker Speech and Audio Generation for Instruct and Zero-Shot Tasks Reward-Conditioned Policies

Reference 2019

Resolution
unresolved
no resolver link, observed 2026-08-04T16:29:24.589310Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T16:29:24.589310Z digest=sha256:6e9205070f2ba5b6fc4438da27d6a992729d5fb1a4662ab46bcbb6f9384bbebe

Observation 34254d69-a337-42cf-ae02-7b91a8c217d5 · inbound

SwanTale: Unified Multi-Speaker Speech and Audio Generation for Instruct and Zero-Shot Tasks cites this paper.

SwanTale: Unified Multi-Speaker Speech and Audio Generation for Instruct and Zero-Shot Tasks Reward-Conditioned Policies

Reference 2019

Resolution
unresolved
no resolver link, observed 2026-08-07T00:14:45.648164Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:14:45.648164Z digest=sha256:121567d74caed7e1dd012f5e037245467f787ef4fe6f238e58aef241be35caa0