Pith. sign in

Paper Citation Record · LEDGER

Self-Exploring Language Models: Active Preference Elicitation for Online Alignment

As of 20 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 17 inbound Pith citation observations for arXiv:2405.19332.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2405.19332 v3

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 17 of 17 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-20T06:33:59.587034+00:00

measured 17 of 17 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-15T20:53:39.902673Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-01T00:35:10.408592Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation 3dc9bcdb-6823-45d8-aaaf-6cd36051aa42 · inbound

DSTC: Direct Preference Learning with Only Self-Generated Tests and Code to Improve Code LMs cites this paper.

DSTC: Direct Preference Learning with Only Self-Generated Tests and Code to Improve Code LMs Self-Exploring Language Models: Active Preference Elicitation for Online Alignment

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-12T17:05:15.224095Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T17:05:15.224095Z digest=sha256:6558adfcb0424406415825ff2bc7d780824362fee4b0cae3bcb06b8823c5fe72

Observation 770af17c-b057-4d2b-8d01-158888e17854 · inbound

Online Learning from Strategic Human Feedback in LLM Fine-Tuning cites this paper.

Online Learning from Strategic Human Feedback in LLM Fine-Tuning Self-Exploring Language Models: Active Preference Elicitation for Online Alignment

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-11T10:21:33.311936Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T10:21:33.311936Z digest=sha256:604cf4f3d895012a9080afbe39b3f7de1eb32579ef1400e6ad90e777d6131414

Observation 7173bdc4-e1e1-48a7-85ae-4dbfe67fac49 · inbound

Understanding the Logic of Direct Preference Alignment through Logic cites this paper.

Understanding the Logic of Direct Preference Alignment through Logic Self-Exploring Language Models: Active Preference Elicitation for Online Alignment

Reference 1975

Resolution
unresolved
no resolver link, observed 2026-08-11T05:24:21.075426Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T05:24:21.075426Z digest=sha256:1137c53b646fcef31a394dbfa1448cd505d163542ae6bd752f5b42a81ea8e84c

Observation 8589d06f-41f9-4b44-9136-39958e26ad64 · inbound

Online Preference Alignment for Language Models via Count-based Exploration cites this paper.

Online Preference Alignment for Language Models via Count-based Exploration Self-Exploring Language Models: Active Preference Elicitation for Online Alignment

Reference 73

Resolution
unresolved
no resolver link, observed 2026-08-10T16:58:02.671392Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T16:58:02.671392Z digest=sha256:f43f7f86f5944e266c0e0992f784282a2b4b4e7eb4bfd9df5e1fea54dc5d19e9

Observation 6626d737-9273-44dd-8a96-e0c78418a57c · inbound

BRiTE: Bootstrapping Reinforced Thinking Process to Enhance Language Model Reasoning cites this paper.

BRiTE: Bootstrapping Reinforced Thinking Process to Enhance Language Model Reasoning Self-Exploring Language Models: Active Preference Elicitation for Online Alignment

Reference 69

Resolution
unresolved
no resolver link, observed 2026-08-09T22:20:13.892177Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T22:20:13.892177Z digest=sha256:edc3977c77c078e68aa657a9edd5bee61ad447b1432c639545f99214167bb195

Observation c09f1447-2d90-4e96-bf86-e8d9d5c36f80 · inbound

PILAF: Optimal Human Preference Sampling for Reward Modeling cites this paper.

PILAF: Optimal Human Preference Sampling for Reward Modeling Self-Exploring Language Models: Active Preference Elicitation for Online Alignment

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-08T23:03:53.293003Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T23:03:53.293003Z digest=sha256:938076d543866588a0c7ea6065b747186203789fe2a0a084fbb6e657f0f9e26e

Observation 59137204-5150-4843-93cc-d2b788d199f2 · inbound

Outcome-Based Online Reinforcement Learning: Algorithms and Fundamental Limits cites this paper.

Outcome-Based Online Reinforcement Learning: Algorithms and Fundamental Limits Self-Exploring Language Models: Active Preference Elicitation for Online Alignment

Reference 52

Resolution
unresolved
no resolver link, observed 2026-08-07T14:10:43.634354Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T14:10:43.634354Z digest=sha256:7c1431fc15c6338f57fe9df94a04098b5de1509cda34bee185e5de670dea5bb4

Observation 9a3474c4-bb56-4ffb-89d4-bf52cc09386e · inbound

Learning a Pessimistic Reward Model in RLHF cites this paper.

Learning a Pessimistic Reward Model in RLHF Self-Exploring Language Models: Active Preference Elicitation for Online Alignment

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-07T14:01:07.063483Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:01:07.063483Z digest=sha256:bb427de2cdd6a9ad27269cf5e833586ea26c8c84be63a0b9a232866fcd2292a9

Observation 6247927f-d8b1-4abb-a09b-43bc9f775eb8 · inbound

Mutual-Taught for Co-adapting Policy and Reward Models cites this paper.

Mutual-Taught for Co-adapting Policy and Reward Models Self-Exploring Language Models: Active Preference Elicitation for Online Alignment

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-15T20:53:39.902673Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T20:53:39.902673Z digest=sha256:8d03e411cc6b8c7135b4e3a56f71182894c860a00a1131b45793b7b9c526d260

Observation 95de08ac-82bd-4bc9-aeae-fa829c9d921b · inbound

MEMETRON: Metaheuristic Mechanisms for Test-time Response Optimization of Large Language Models cites this paper.

MEMETRON: Metaheuristic Mechanisms for Test-time Response Optimization of Large Language Models Self-Exploring Language Models: Active Preference Elicitation for Online Alignment

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-07T05:10:47.640424Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:10:47.640424Z digest=sha256:8f2817b093a7b1cb0e30b494dcaf0b72fd14f81a08ba38051b3989db3f71f5e2

Observation 9fac93e6-af53-45fd-9ed7-fd92bd07a6d4 · inbound

Exploration from a Primal-Dual Lens: Value-Incentivized Actor-Critic Methods for Sample-Efficient Online RL cites this paper.

Exploration from a Primal-Dual Lens: Value-Incentivized Actor-Critic Methods for Sample-Efficient Online RL Self-Exploring Language Models: Active Preference Elicitation for Online Alignment

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-06T22:16:41.348358Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:16:41.348358Z digest=sha256:ff4e6e19006edf721fb765ad7aab0d3257f7ec95133d0a0fa06238bb8615602f

Observation 243ea628-e798-41c4-acb3-3fda90368e0b · inbound

SharedRep-RLHF: A Shared Representation Approach to RLHF with Diverse Preferences cites this paper.

SharedRep-RLHF: A Shared Representation Approach to RLHF with Diverse Preferences Self-Exploring Language Models: Active Preference Elicitation for Online Alignment

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-05T10:54:13.592468Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T10:54:13.592468Z digest=sha256:c3dc21ae13511c87d19a196c88d97f23d313fb3cf5052b95c0ff3f4ab4777573

Observation 07e30d42-0401-4a3e-98d7-9f776f3d8e3e · inbound

Outcome-based Exploration for LLM Reasoning cites this paper.

Outcome-based Exploration for LLM Reasoning Self-Exploring Language Models: Active Preference Elicitation for Online Alignment

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-04T22:59:14.599880Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T22:59:14.599880Z digest=sha256:bb8175769cdcfb76eb04e0c4b207e8c3845a2cc22cc5cb94a3dae0b8e82a259c

Observation 22b91954-b84e-4ac6-a24d-098a88669763 · inbound

Data-dependent Exploration for Online Reinforcement Learning from Human Feedback cites this paper.

Data-dependent Exploration for Online Reinforcement Learning from Human Feedback Self-Exploring Language Models: Active Preference Elicitation for Online Alignment

Reference 90

Resolution
verified exact
arxiv_id, observed 2026-05-11T17:51:09.202379Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-05-08T17:01:04.571087Z digest=sha256:d1838b31138f2872c81c602b21f9e969b3aec58c7f0e57106b18d60f6baed85e

Observation 982bf544-798f-40b6-b81c-1003b1591837 · inbound

Data-dependent Exploration for Online Reinforcement Learning from Human Feedback cites this paper.

Data-dependent Exploration for Online Reinforcement Learning from Human Feedback Self-Exploring Language Models: Active Preference Elicitation for Online Alignment

Reference 36

Resolution
verified exact
arxiv_id, observed 2026-07-01T00:35:10.410391Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-07-01T00:27:39.321158Z digest=sha256:fad0d1b857f643a6e3e66bcfa75b3ebb7af288030f60f039a3a6fa53b5bd4a91

Observation 757cb0ef-18f2-4480-a604-30e6e4b799b9 · inbound

Recall Isn't Enough: Bounding Commitments in Personalized Language Systems cites this paper.

Recall Isn't Enough: Bounding Commitments in Personalized Language Systems Self-Exploring Language Models: Active Preference Elicitation for Online Alignment

Reference 12

Resolution
metadata mismatch
arxiv_id, observed 2026-05-20T17:33:36.647908Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-20T17:29:55.760720Z digest=sha256:0841689a2b80275b920bfda0ba5757309b81bd3e64a3f6e6c5fd72b29945559d

Observation 04e9207b-5337-4d7f-af8e-b06cb0c08b39 · inbound

Spectral Souping: A Unified Framework for Online Preference Alignment cites this paper.

Spectral Souping: A Unified Framework for Online Preference Alignment Self-Exploring Language Models: Active Preference Elicitation for Online Alignment

Reference 36

Resolution
verified exact
arxiv_id, observed 2026-05-21T07:59:50.921026Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-21T07:54:56.356555Z digest=sha256:d6a8d800bc440d6694b90b116208f0d60f7ecb59d54682d51e105843ed4afbf6