Pith. sign in

Paper Citation Record · LEDGER

Dueling RL: Reinforcement Learning with Trajectory Preferences

As of 13 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 15 inbound Pith citation observations for arXiv:2111.04850.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2111.04850 v3

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 15 of 15 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-13T06:32:02.005865+00:00

measured 15 of 15 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-11T10:21:33.356599Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-04T13:09:51.333293Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation 102e7f13-4d9e-4e7b-b86e-24dd6689a6e8 · inbound

Online Learning from Strategic Human Feedback in LLM Fine-Tuning cites this paper.

Online Learning from Strategic Human Feedback in LLM Fine-Tuning Dueling RL: Reinforcement Learning with Trajectory Preferences

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-11T10:21:33.356599Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T10:21:33.356599Z digest=sha256:c1a1afe23e3faa93a9629a806759094f6a6c26e5a533fc49cecdbfef5a2efa52

Observation dd549d22-6eaf-4d51-b6c9-afd1c0b4758b · inbound

Can We Generate Images with CoT? Let's Verify and Reinforce Image Generation Step by Step cites this paper.

Can We Generate Images with CoT? Let's Verify and Reinforce Image Generation Step by Step Dueling RL: Reinforcement Learning with Trajectory Preferences

Reference 56

Resolution
unresolved
no resolver link, observed 2026-08-10T15:32:46.438395Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T15:32:46.438395Z digest=sha256:b58628e0e5b78452e4c3658734ea6c04648c0c34ae43ff913d22026cec2b91d6

Observation 24eb2a2f-3a01-4507-98d8-ef79d8fa9179 · inbound

Leveraging Sparsity for Sample-Efficient Preference Learning: A Theoretical Perspective cites this paper.

Leveraging Sparsity for Sample-Efficient Preference Learning: A Theoretical Perspective Dueling RL: Reinforcement Learning with Trajectory Preferences

Reference 62

Resolution
unresolved
no resolver link, observed 2026-08-10T00:11:59.266988Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T00:11:59.266988Z digest=sha256:97547d128c30a2a586ef2823e18bbf51c44542ed4d465c950b7189c06e78fa26

Observation d4263822-9942-4e16-8e67-8b2d01b6219e · inbound

BRiTE: Bootstrapping Reinforced Thinking Process to Enhance Language Model Reasoning cites this paper.

BRiTE: Bootstrapping Reinforced Thinking Process to Enhance Language Model Reasoning Dueling RL: Reinforcement Learning with Trajectory Preferences

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-09T22:20:13.781528Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T22:20:13.781528Z digest=sha256:a271d707eb6a72c45a742bf72b0461e2ac9b16df05f2c873462c1cc34dd27257

Observation d9904036-a23e-4ef2-b25c-ac1113470fd4 · inbound

A Unified Theoretical Analysis of Private and Robust Offline Alignment: from RLHF to DPO cites this paper.

A Unified Theoretical Analysis of Private and Robust Offline Alignment: from RLHF to DPO Dueling RL: Reinforcement Learning with Trajectory Preferences

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-07T15:19:09.343297Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:19:09.343297Z digest=sha256:3b5163f9fd606adbc42478bd52f1dc43ed7b0af8b9a2e46118a2e6d0c80622f6

Observation d1835f1e-c6e8-4cb5-9847-3189e666cdff · inbound

Outcome-Based Online Reinforcement Learning: Algorithms and Fundamental Limits cites this paper.

Outcome-Based Online Reinforcement Learning: Algorithms and Fundamental Limits Dueling RL: Reinforcement Learning with Trajectory Preferences

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-07T14:10:42.530610Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T14:10:42.530610Z digest=sha256:9f0f2badec6d0047df41877537a72607de52d34522d6b1710465e3a450bf8893

Observation 0c62cbf0-9062-4a4a-99ed-a51ca34dc1f6 · inbound

On the optimization dynamics of RLVR: Gradient gap and step size thresholds cites this paper.

On the optimization dynamics of RLVR: Gradient gap and step size thresholds Dueling RL: Reinforcement Learning with Trajectory Preferences

Reference 15

Resolution
verified exact
arxiv_id, observed 2026-05-18T08:36:07.354937Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-05-18T08:34:36.543874Z digest=sha256:20b2a0e113c98bc9a95ad3b36577812d5fbc950a22f16b2d5f5fa8f13669375d

Observation defe716c-5c46-480f-a793-7748213f5d3a · inbound

OPRIDE: Offline Preference-based Reinforcement Learning via In-Dataset Exploration cites this paper.

OPRIDE: Offline Preference-based Reinforcement Learning via In-Dataset Exploration Dueling RL: Reinforcement Learning with Trajectory Preferences

Reference 36

Resolution
metadata mismatch
arxiv_id, observed 2026-05-15T21:40:21.050286Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=arxiv_source observed=2026-05-15T21:38:53.792221Z digest=sha256:44b01396d18a989a8fb37de73423f95aee4b3270cd481d2bc998f10c6d489412

Observation 07ba240b-ae6f-4e58-b5e9-813337454882 · inbound

Online KL-Regularized Reinforcement Learning with Function Approximation under Misspecification cites this paper.

Online KL-Regularized Reinforcement Learning with Function Approximation under Misspecification Dueling RL: Reinforcement Learning with Trajectory Preferences

Reference 11

Resolution
metadata mismatch
arxiv_id, observed 2026-07-02T12:06:55.881408Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-06-28T02:31:11.200818Z digest=sha256:396eb153024b3a827e724612e7ea5c7b55e5bde548f188f32c86c102cd6588c6

Observation 5a19d89e-b388-4be5-b1cd-d27469ec0e8d · inbound

Online KL-Regularized Reinforcement Learning with Function Approximation under Misspecification cites this paper.

Online KL-Regularized Reinforcement Learning with Function Approximation under Misspecification Dueling RL: Reinforcement Learning with Trajectory Preferences

Reference 11

Resolution
unresolved
no resolver link, observed 2026-07-14T18:23:21.126826Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T18:23:21.126826Z digest=sha256:e70db6724b44954dcda1caa99098713b1b1770a1cce4b4d6e65c51ffe2352b5a

Observation 781d6f84-532b-4e3a-83e9-be6b92614edf · inbound

Finding Stationary Points by Comparisons cites this paper.

Finding Stationary Points by Comparisons Dueling RL: Reinforcement Learning with Trajectory Preferences

Reference 41

Resolution
metadata mismatch
arxiv_id, observed 2026-07-04T13:09:51.334817Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-06-26T05:26:07.720218Z digest=sha256:7ea443a65bd632f1dbeb0f16264e201520d479fdc69b00de145f1af98f397c5d

Observation 244e8fb4-b71c-45f9-8612-a7f08767163e · inbound

Preference-Based Reward Learning under Partial Observability with Inexact Dynamics cites this paper.

Preference-Based Reward Learning under Partial Observability with Inexact Dynamics Dueling RL: Reinforcement Learning with Trajectory Preferences

Reference 14

Resolution
verified exact
arxiv_id, observed 2026-06-30T14:44:45.876543Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-06-30T05:19:16.647402Z digest=sha256:0f4b1c6def4b9d4cb211abf5b3ff06c0c2bdb06dbcb9e0e4699f1d5d9f489955

Observation ad2d9dad-f5ce-435c-8795-433203f26192 · inbound

SPLC: Social Preference Learning for Crowd Robot Navigation cites this paper.

SPLC: Social Preference Learning for Crowd Robot Navigation Dueling RL: Reinforcement Learning with Trajectory Preferences

Reference 31

Resolution
metadata mismatch
arxiv_id, observed 2026-07-03T12:18:06.219993Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-07-03T12:10:03.492401Z digest=sha256:da7ec58c263e69ddc1cad704d27e699c773bf59eab7374abca5192a2c93ab8ae

Observation 09abf6a9-e868-4936-8a26-b281ca942b1e · inbound

Multi-Turn On-Policy Distillation with Prefix Replay cites this paper.

Multi-Turn On-Policy Distillation with Prefix Replay Dueling RL: Reinforcement Learning with Trajectory Preferences

Reference 264

Resolution
unresolved
no resolver link, observed 2026-07-11T13:53:36.775836Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-11T13:53:36.775836Z digest=sha256:6eeaf0b2c7e10e0a1619264e6410a163384eda97ae51102cd70e9e9b2dfb80b1

Observation 83ae8a47-bfda-4f03-ad99-e042ad21549f · inbound

Multi-Turn On-Policy Distillation with Prefix Replay cites this paper.

Multi-Turn On-Policy Distillation with Prefix Replay Dueling RL: Reinforcement Learning with Trajectory Preferences

Reference 265

Resolution
unresolved
no resolver link, observed 2026-08-02T08:41:03.234561Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-02T08:41:03.234561Z digest=sha256:4c300022cb1137630f1c3748b4a02585f96cb7e96edf6d30fb5c5d38efd8cb87