Pith. sign in

Paper Citation Record · LEDGER

Sample-Efficient Reinforcement Learning from Human Feedback via Information-Directed Sampling

As of 10 August 2026, this Paper Citation Record lists 12 of 12 outbound references and 1 inbound Pith citation observation for arXiv:2502.05434.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2502.05434 v3

Coverage vector

measured 12 of 12 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-08T19:28:41.615055Z

measured 13 of 13 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-09T06:31:02.800959+00:00

measured 1 of 1 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-07T14:01:05.568274Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-08-07T14:01:08.101786Z

Reference resolution

12 of 12 outbound references displayed

  • verified exact0
  • verified fuzzy6
  • unresolved6
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 301c011c-aaae-4ba8-a33f-c7d68292fcee · outbound

This paper cites an unresolved cited work.

Sample-Efficient Reinforcement Learning from Human Feedback via Information-Directed Sampling Unresolved cited work

Reference 1

Resolution
unresolved
raw_fallback, observed 2026-08-08T19:28:42.418904Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-08T19:28:41.321150Z digest=sha256:5d51c40438a7c1d08a9d99b8df9c3c08de75f55d358173ba691b25dacff5f950

Observation 16cf7854-2d9d-4833-9a00-279d84c2eae4 · outbound

This paper cites an unresolved cited work.

Sample-Efficient Reinforcement Learning from Human Feedback via Information-Directed Sampling Unresolved cited work

Reference 2

Resolution
unresolved
raw_fallback, observed 2026-08-08T19:28:42.409655Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-08T19:28:41.325333Z digest=sha256:eae8d8958bf09146e0c27aa9895bf9e1b2f840cf8fc1f7d846f65885ca32c7d8

Observation c5ece5db-d921-4e3e-8e88-15709e1805c9 · outbound

This paper cites an unresolved cited work.

Sample-Efficient Reinforcement Learning from Human Feedback via Information-Directed Sampling Unresolved cited work

Reference 3

Resolution
unresolved
raw_fallback, observed 2026-08-08T19:28:42.263004Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-08T19:28:41.329421Z digest=sha256:67efe49082f8578d50e782a861987140649455b8a8015bfd13a7ff0d38952cfd

Observation 7f48ac4f-09c5-4e78-8c4a-96a53d977655 · outbound

This paper cites an unresolved cited work.

Sample-Efficient Reinforcement Learning from Human Feedback via Information-Directed Sampling Unresolved cited work

Reference 4

Resolution
unresolved
raw_fallback, observed 2026-08-08T19:28:42.065511Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-08T19:28:41.332925Z digest=sha256:30ff351e093cea61835cf07bc3039f86793bf0117a206b180c0bd4515e56bd85

Observation 5006d36b-43e7-411c-99bf-a58c8d7a1fb6 · outbound

This paper cites For the first term in Eqn.(A.3), using the basic fact thatA − λB/2 ≤ A2/2λB for B, λ≥ 0, we have Et V eE ∗ t 1,π∗ E (st.

Sample-Efficient Reinforcement Learning from Human Feedback via Information-Directed Sampling For the first term in Eqn.(A.3), using the basic fact thatA − λB/2 ≤ A2/2λB for B, λ≥ 0, we have Et V eE ∗ t 1,π∗ E (st

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T19:28:42.001984Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-08T19:28:41.336619Z digest=sha256:e2b14c7c92d88c5665381791eef91a18d3c9f728d99073d9442506e7442b875b

Observation 52c36661-91b5-4215-8ada-ab2a6923595c · outbound

This paper cites an unresolved cited work.

Sample-Efficient Reinforcement Learning from Human Feedback via Information-Directed Sampling Unresolved cited work

Reference 6

Resolution
unresolved
raw_fallback, observed 2026-08-08T19:28:41.991949Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-08T19:28:41.340457Z digest=sha256:9b083331899eb337467286cd598b96912f9fba6ba315ad2be7e58d3fe601cfe8

Observation 022b7e68-6ada-4e95-bfc5-46b5307b7e49 · outbound

This paper cites (A.4) 21 where we introduce the tool ofinformation ratio Γ πt TS t for ease of analysis.

Sample-Efficient Reinforcement Learning from Human Feedback via Information-Directed Sampling (A.4) 21 where we introduce the tool ofinformation ratio Γ πt TS t for ease of analysis

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T19:28:41.981743Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-08T19:28:41.345195Z digest=sha256:7db104f24d80bce7904b412d71bffb7044adfb022e999a2e750543f37deddf2d

Observation e3e53918-6140-4950-99db-7c2cc4ed6cdd · outbound

This paper cites However, Lemma C.1 can only be applied to handle the difference between two value functions with the same policy and different environments, while inV eE ∗ t 1,π∗ E (st.

Sample-Efficient Reinforcement Learning from Human Feedback via Information-Directed Sampling However, Lemma C.1 can only be applied to handle the difference between two value functions with the same policy and different environments, while inV eE ∗ t 1,π∗ E (st

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T19:28:41.969658Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-08T19:28:41.489650Z digest=sha256:faf764601899879f5085a0b6da670f75d969cd3784696c3d718da4daf4237204

Observation 292e9b55-a7f7-4a87-8aa1-1acea5ebc1c3 · outbound

This paper cites unifying.

Sample-Efficient Reinforcement Learning from Human Feedback via Information-Directed Sampling unifying

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T19:28:41.893703Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-08T19:28:41.532321Z digest=sha256:e4d8b79babd369e76d020e6caeaa3752fa8456fb7de7febac97cb0c4e44b2102

Observation bd7d0244-7710-47ef-bcb9-7162a8cc4884 · outbound

This paper cites an unresolved cited work.

Sample-Efficient Reinforcement Learning from Human Feedback via Information-Directed Sampling Unresolved cited work

Reference 10

Resolution
unresolved
raw_fallback, observed 2026-08-08T19:28:41.686031Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-08T19:28:41.592222Z digest=sha256:bfd163967064cf7c94e6dc9c78610daac91b74df27145a7c6084dbb2c598e35f

Observation fa6fa785-4cdc-49de-b1cc-a9a01032805b · outbound

This paper cites integrated.

Sample-Efficient Reinforcement Learning from Human Feedback via Information-Directed Sampling integrated

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T19:28:41.654315Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-08T19:28:41.610662Z digest=sha256:aa8d8543b4d4273e814b59ffadd123fc9bab4e745a437995f13380a3a4f139c8

Observation 650aaa96-bb07-43e6-84a4-86b62ce39d17 · outbound

This paper cites Since K(ϵ) ≤ ( 3H 2 ϵ )SAH · ( 6H 2√ S ϵ )H · ( 3H ϵ )SAH , 27 we have log(K(ϵ)) ≤ SAH log 3H 2 ϵ + H log 6H 2√ S ϵ + SAH log 3H ϵ ≤ 3SAH log 6H 2√ S ϵ.

Sample-Efficient Reinforcement Learning from Human Feedback via Information-Directed Sampling Since K(ϵ) ≤ ( 3H 2 ϵ )SAH · ( 6H 2√ S ϵ )H · ( 3H ϵ )SAH , 27 we have log(K(ϵ)) ≤ SAH log 3H 2 ϵ + H log 6H 2√ S ϵ + SAH log 3H ϵ ≤ 3SAH log 6H 2√ S ϵ

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T19:28:41.644467Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-08T19:28:41.615055Z digest=sha256:c4c6f287d44e87fad8f906f19830aa4571924c7acc71a8dfe778590a24ef164e

Pith citing papers

Observation 301e7a58-e6e9-434e-9935-0ccf2266b1bf · inbound

Learning a Pessimistic Reward Model in RLHF cites this paper.

Learning a Pessimistic Reward Model in RLHF Sample-Efficient Reinforcement Learning from Human Feedback via Information-Directed Sampling

Reference 2022

Resolution
verified exact
local_arxiv, observed 2026-08-07T14:01:08.191682Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T14:01:05.568274Z digest=sha256:4338519ffbead753636c0855746fb4eebf79c7d2459aa9009d75ee5cbbf0f211