Pith. sign in

Paper Citation Record · LEDGER

Benchmarks and Algorithms for Offline Preference-Based Reward Learning

As of 9 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 7 inbound Pith citation observations for arXiv:2301.01392.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2301.01392 v1

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 7 of 7 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-09T06:31:02.800959+00:00

measured 7 of 7 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-07T13:45:00.273913Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-03T12:18:06.214750Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation 876b151e-b9d3-4fb6-92b3-4596e38486f6 · inbound

Square$\chi$PO: Differentially Private and Robust $\chi^2$-Preference Optimization in Offline Direct Alignment cites this paper.

Square$\chi$PO: Differentially Private and Robust $\chi^2$-Preference Optimization in Offline Direct Alignment Benchmarks and Algorithms for Offline Preference-Based Reward Learning

Reference 66

Resolution
unresolved
no resolver link, observed 2026-08-07T13:45:00.273913Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:45:00.273913Z digest=sha256:eb1a608d19a2ef797d9e30ca0c9f0ba7318b5ebc9fac96f4430aa09a9f53c264

Observation 3931ac7b-a868-4eec-9af1-bf714c641b6c · inbound

CLARIFY: Contrastive Preference Reinforcement Learning for Untangling Ambiguous Queries cites this paper.

CLARIFY: Contrastive Preference Reinforcement Learning for Untangling Ambiguous Queries Benchmarks and Algorithms for Offline Preference-Based Reward Learning

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-07T12:13:00.672841Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:13:00.672841Z digest=sha256:c99de19515edd15497b797fe3affead25e9692638333b35f1ae8ac74a85342be

Observation ffb01e19-fb44-4a5f-a819-c5e26406ef07 · inbound

Residual Reward Models for Preference-based Reinforcement Learning cites this paper.

Residual Reward Models for Preference-based Reinforcement Learning Benchmarks and Algorithms for Offline Preference-Based Reward Learning

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-06T21:17:39.261586Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:17:39.261586Z digest=sha256:f9d0f8ec61a86a5182609030ac0438fe3b9730b873fa515e2ef6dd0776eaa976

Observation 1454ab01-6b28-4d9b-b892-9307e16442ce · inbound

OPRIDE: Offline Preference-based Reinforcement Learning via In-Dataset Exploration cites this paper.

OPRIDE: Offline Preference-based Reinforcement Learning via In-Dataset Exploration Benchmarks and Algorithms for Offline Preference-Based Reward Learning

Reference 44

Resolution
verified exact
arxiv_id, observed 2026-05-15T21:40:21.002736Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-05-15T21:38:53.792221Z digest=sha256:86e533ec2b355f893a88e6792a980bb6bf232590942632b571b258b7951e10b7

Observation 62244405-13a5-4f60-aefa-2ef2d81d2be4 · inbound

SPLC: Social Preference Learning for Crowd Robot Navigation cites this paper.

SPLC: Social Preference Learning for Crowd Robot Navigation Benchmarks and Algorithms for Offline Preference-Based Reward Learning

Reference 34

Resolution
metadata mismatch
arxiv_id, observed 2026-07-03T12:18:06.216403Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-07-03T12:10:03.492401Z digest=sha256:9d85eae310bd6c69054aa956c1b4f6e132cf90401c53cbef9926766e660e1f8c

Observation 83171e37-a26e-4543-b3c8-9e6034f87ddb · inbound

Multi-Turn On-Policy Distillation with Prefix Replay cites this paper.

Multi-Turn On-Policy Distillation with Prefix Replay Benchmarks and Algorithms for Offline Preference-Based Reward Learning

Reference 281

Resolution
unresolved
no resolver link, observed 2026-07-11T13:53:36.775836Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-11T13:53:36.775836Z digest=sha256:e6d2372f16b0120f0ea12f3f9c6f1a8ecf1091ae0cb395afd3d8db8cff0aeb7a

Observation 3bfe8de6-3dbf-4c42-b71c-7908cd313927 · inbound

Multi-Turn On-Policy Distillation with Prefix Replay cites this paper.

Multi-Turn On-Policy Distillation with Prefix Replay Benchmarks and Algorithms for Offline Preference-Based Reward Learning

Reference 282

Resolution
unresolved
no resolver link, observed 2026-08-02T08:41:05.320411Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-02T08:41:05.320411Z digest=sha256:d4e4b4c0357d83f258021e2287fb0b6531aa84b1ef499be17d8abd81b7abd76f