Pith. sign in

Paper Citation Record · LEDGER

Sample-Efficient Reinforcement Learning from Human Feedback via Information-Directed Sampling

As of 12 August 2026, this Paper Citation Record lists 12 of 12 outbound references and 1 inbound Pith citation observation for arXiv:2502.05434.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2502.05434 v3

Coverage vector

measured 12 of 12 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-08T19:28:41.615055Z

measured 13 of 13 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-12T06:34:41.77262+00:00

measured 1 of 1 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-07T14:01:05.568274Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-08-07T14:01:08.101786Z

Reference resolution

12 of 12 outbound references displayed

  • verified exact0
  • verified fuzzy6
  • unresolved6
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 301c011c-aaae-4ba8-a33f-c7d68292fcee · outbound

This paper cites an unresolved cited work.

Sample-Efficient Reinforcement Learning from Human Feedback via Information-Directed Sampling Unresolved cited work

Reference 1

Resolution
unresolved
raw_fallback, observed 2026-08-08T19:28:42.418904Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-08T19:28:41.321150Z digest=sha256:d1531ce00f86f569a1733724aa353131d626d09668edfcf7008fd71333f1d68a

Observation 16cf7854-2d9d-4833-9a00-279d84c2eae4 · outbound

This paper cites an unresolved cited work.

Sample-Efficient Reinforcement Learning from Human Feedback via Information-Directed Sampling Unresolved cited work

Reference 2

Resolution
unresolved
raw_fallback, observed 2026-08-08T19:28:42.409655Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-08T19:28:41.325333Z digest=sha256:795f61b8d17052906014c083213a494328f12b05f1b61995f9aed5bb91386069

Observation c5ece5db-d921-4e3e-8e88-15709e1805c9 · outbound

This paper cites an unresolved cited work.

Sample-Efficient Reinforcement Learning from Human Feedback via Information-Directed Sampling Unresolved cited work

Reference 3

Resolution
unresolved
raw_fallback, observed 2026-08-08T19:28:42.263004Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-08T19:28:41.329421Z digest=sha256:c4af708c4ec6871c57756007c2301159d4dbbaa1197ec5d5545e58c05321df8b

Observation 7f48ac4f-09c5-4e78-8c4a-96a53d977655 · outbound

This paper cites an unresolved cited work.

Sample-Efficient Reinforcement Learning from Human Feedback via Information-Directed Sampling Unresolved cited work

Reference 4

Resolution
unresolved
raw_fallback, observed 2026-08-08T19:28:42.065511Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-08T19:28:41.332925Z digest=sha256:388c6f7cc82ac6d3dece93bc17ff20db116704cc577e9a013f658ceb2f4b9a1c

Observation 5006d36b-43e7-411c-99bf-a58c8d7a1fb6 · outbound

This paper cites For the first term in Eqn.(A.3), using the basic fact thatA − λB/2 ≤ A2/2λB for B, λ≥ 0, we have Et V eE ∗ t 1,π∗ E (st.

Sample-Efficient Reinforcement Learning from Human Feedback via Information-Directed Sampling For the first term in Eqn.(A.3), using the basic fact thatA − λB/2 ≤ A2/2λB for B, λ≥ 0, we have Et V eE ∗ t 1,π∗ E (st

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T19:28:42.001984Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-08T19:28:41.336619Z digest=sha256:5358b1ad3e0eb688565d15c31ac4fdfc63e06414a7490b0765bf2e78d5cf5766

Observation 52c36661-91b5-4215-8ada-ab2a6923595c · outbound

This paper cites an unresolved cited work.

Sample-Efficient Reinforcement Learning from Human Feedback via Information-Directed Sampling Unresolved cited work

Reference 6

Resolution
unresolved
raw_fallback, observed 2026-08-08T19:28:41.991949Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-08T19:28:41.340457Z digest=sha256:825d149c0dffd2a5c04d2104455bb26f6dc61a9a47d0ebce79342029627b83f6

Observation 022b7e68-6ada-4e95-bfc5-46b5307b7e49 · outbound

This paper cites (A.4) 21 where we introduce the tool ofinformation ratio Γ πt TS t for ease of analysis.

Sample-Efficient Reinforcement Learning from Human Feedback via Information-Directed Sampling (A.4) 21 where we introduce the tool ofinformation ratio Γ πt TS t for ease of analysis

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T19:28:41.981743Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-08T19:28:41.345195Z digest=sha256:0998057db8f43f7b5aaaea9ed7c9e63629d747df11462191ecb86490e34e3158

Observation e3e53918-6140-4950-99db-7c2cc4ed6cdd · outbound

This paper cites However, Lemma C.1 can only be applied to handle the difference between two value functions with the same policy and different environments, while inV eE ∗ t 1,π∗ E (st.

Sample-Efficient Reinforcement Learning from Human Feedback via Information-Directed Sampling However, Lemma C.1 can only be applied to handle the difference between two value functions with the same policy and different environments, while inV eE ∗ t 1,π∗ E (st

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T19:28:41.969658Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-08T19:28:41.489650Z digest=sha256:a119dde690b4c74514f03328e8eb0ebe455d275131d102addd521a1e2a4a5ecb

Observation 292e9b55-a7f7-4a87-8aa1-1acea5ebc1c3 · outbound

This paper cites unifying.

Sample-Efficient Reinforcement Learning from Human Feedback via Information-Directed Sampling unifying

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T19:28:41.893703Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-08T19:28:41.532321Z digest=sha256:08d6b437f886627a09571140760e873184a485aadc753a96aef0041c07097feb

Observation bd7d0244-7710-47ef-bcb9-7162a8cc4884 · outbound

This paper cites an unresolved cited work.

Sample-Efficient Reinforcement Learning from Human Feedback via Information-Directed Sampling Unresolved cited work

Reference 10

Resolution
unresolved
raw_fallback, observed 2026-08-08T19:28:41.686031Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-08T19:28:41.592222Z digest=sha256:10a8c45473515e1d5df7ec46a8ec2f90144b35af7b2aa66a1a08339ebdd47272

Observation fa6fa785-4cdc-49de-b1cc-a9a01032805b · outbound

This paper cites integrated.

Sample-Efficient Reinforcement Learning from Human Feedback via Information-Directed Sampling integrated

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T19:28:41.654315Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-08T19:28:41.610662Z digest=sha256:cf49c1963245c6b4c349a084606d04a8e800b5b856e05bf8313a8934b88058c0

Observation 650aaa96-bb07-43e6-84a4-86b62ce39d17 · outbound

This paper cites Since K(ϵ) ≤ ( 3H 2 ϵ )SAH · ( 6H 2√ S ϵ )H · ( 3H ϵ )SAH , 27 we have log(K(ϵ)) ≤ SAH log 3H 2 ϵ + H log 6H 2√ S ϵ + SAH log 3H ϵ ≤ 3SAH log 6H 2√ S ϵ.

Sample-Efficient Reinforcement Learning from Human Feedback via Information-Directed Sampling Since K(ϵ) ≤ ( 3H 2 ϵ )SAH · ( 6H 2√ S ϵ )H · ( 3H ϵ )SAH , 27 we have log(K(ϵ)) ≤ SAH log 3H 2 ϵ + H log 6H 2√ S ϵ + SAH log 3H ϵ ≤ 3SAH log 6H 2√ S ϵ

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T19:28:41.644467Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-08T19:28:41.615055Z digest=sha256:128bcd0f5a66c3f249252aaf91fc3bab7499a76e0d82ac14148fa95cf4637447

Pith citing papers

Observation 301e7a58-e6e9-434e-9935-0ccf2266b1bf · inbound

Learning a Pessimistic Reward Model in RLHF cites this paper.

Learning a Pessimistic Reward Model in RLHF Sample-Efficient Reinforcement Learning from Human Feedback via Information-Directed Sampling

Reference 2022

Resolution
verified exact
local_arxiv, observed 2026-08-07T14:01:08.191682Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-07T14:01:05.568274Z digest=sha256:286b83de66138af873f14aff596ef88f0dcdaeb0215a25a0e1889b7ed47a8a0c