Pith. sign in

Paper Citation Record · LEDGER

Is RLHF More Difficult than Standard RL?

As of 10 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 10 inbound Pith citation observations for arXiv:2306.14111.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2306.14111 v2

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 10 of 10 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-10T06:31:04.303077+00:00

measured 10 of 10 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-08T19:40:42.054396Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-04T13:09:51.330383Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation a64d54ee-d14d-4af9-a98e-368242dbd36e · inbound

Design Considerations in Offline Preference-based RL cites this paper.

Design Considerations in Offline Preference-based RL Is RLHF More Difficult than Standard RL?

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-08T19:40:42.054396Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T19:40:42.054396Z digest=sha256:e51d6642a78fed4910f0b45a597a4e25a0b14e253c950ec03c1bfaf72d06ed67

Observation f0ca88e5-076b-4d1f-974d-5f4c209eb44a · inbound

Outcome-Based Online Reinforcement Learning: Algorithms and Fundamental Limits cites this paper.

Outcome-Based Online Reinforcement Learning: Algorithms and Fundamental Limits Is RLHF More Difficult than Standard RL?

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-07T14:10:42.886665Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T14:10:42.886665Z digest=sha256:2987df85d919a7612ae360027d8e5b10fa9d87c481b3976f7d13ac2d8f436181

Observation bc6e284f-0f30-4db7-b4a7-e9324e6af1ae · inbound

Square$\chi$PO: Differentially Private and Robust $\chi^2$-Preference Optimization in Offline Direct Alignment cites this paper.

Square$\chi$PO: Differentially Private and Robust $\chi^2$-Preference Optimization in Offline Direct Alignment Is RLHF More Difficult than Standard RL?

Reference 81

Resolution
unresolved
no resolver link, observed 2026-08-07T13:45:02.561160Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:45:02.561160Z digest=sha256:7287729b85c9d81f99b4ea435d0bf719ce759efabd81efa6cd0c181ef8dbf3d7

Observation 68c03123-55da-4bc6-a574-bbf61d4e86f5 · inbound

Sign-SZPO: Provable Preference-based Reinforcement Learning with an Unknown Link Function cites this paper.

Sign-SZPO: Provable Preference-based Reinforcement Learning with an Unknown Link Function Is RLHF More Difficult than Standard RL?

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-07T11:20:12.854603Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T11:20:12.854603Z digest=sha256:0955a210c500e2d24c4c98241c639e34dba12f86d1f07acc700eebbe7a5c84c3

Observation a076c2a1-5fe0-4a7e-ae2c-54027c45ddc9 · inbound

On the optimization dynamics of RLVR: Gradient gap and step size thresholds cites this paper.

On the optimization dynamics of RLVR: Gradient gap and step size thresholds Is RLHF More Difficult than Standard RL?

Reference 23

Resolution
verified exact
arxiv_id, observed 2026-05-18T08:36:07.373729Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-18T08:34:36.543874Z digest=sha256:afca5adf47f0f93d9afd723b1b2c4c1c7f6447a89d4cec614dcecd95fb93a5a5

Observation 6befe8be-e0f8-434d-84d1-92d2498a0dd2 · inbound

OPRIDE: Offline Preference-based Reinforcement Learning via In-Dataset Exploration cites this paper.

OPRIDE: Offline Preference-based Reinforcement Learning via In-Dataset Exploration Is RLHF More Difficult than Standard RL?

Reference 47

Resolution
verified exact
arxiv_id, observed 2026-05-15T21:40:20.961449Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-05-15T21:38:53.792221Z digest=sha256:688a1da5827695405f826a7aea63b9f17ea410a78bcf0602ccf1e72e4474c095

Observation 52f4111c-04ef-40a4-8821-5420410760d2 · inbound

Convex Optimization for Alignment and Preference Learning on a Single GPU cites this paper.

Convex Optimization for Alignment and Preference Learning on a Single GPU Is RLHF More Difficult than Standard RL?

Reference 46

Resolution
verified exact
arxiv_id, observed 2026-05-25T05:05:23.015101Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-05-25T05:01:31.560963Z digest=sha256:8ae8918b79575f897b837c27f5b562f240f2a07b77927897003175bc1e3ab154

Observation 5604856b-fb3b-4589-a8f2-4086f6b651d6 · inbound

Online KL-Regularized Reinforcement Learning with Function Approximation under Misspecification cites this paper.

Online KL-Regularized Reinforcement Learning with Function Approximation under Misspecification Is RLHF More Difficult than Standard RL?

Reference 16

Resolution
metadata mismatch
arxiv_id, observed 2026-07-02T12:06:55.886189Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-06-28T02:31:11.200818Z digest=sha256:fd524a2948d929bc857a5b4b58010d27d3c6e6d425ecb565296505f5964224c2

Observation 6d598499-4ca3-44e8-90fc-58b233a99fd8 · inbound

Online KL-Regularized Reinforcement Learning with Function Approximation under Misspecification cites this paper.

Online KL-Regularized Reinforcement Learning with Function Approximation under Misspecification Is RLHF More Difficult than Standard RL?

Reference 16

Resolution
unresolved
no resolver link, observed 2026-07-14T18:23:21.126826Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T18:23:21.126826Z digest=sha256:d46dfee90e9fb29d4a98020319241c78bc738abedd1d3eb8b4e5a6815709661a

Observation b36b1692-97d3-438d-b55e-ac2f3ea1231f · inbound

Finding Stationary Points by Comparisons cites this paper.

Finding Stationary Points by Comparisons Is RLHF More Difficult than Standard RL?

Reference 45

Resolution
metadata mismatch
arxiv_id, observed 2026-07-04T13:09:51.331914Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-06-26T05:26:07.720218Z digest=sha256:7bbe20dfaf6633e0c95afe41e4a89e41de0546e6d934d41912b1dcd27e3a3764