Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links
Paper Citation Record · LEDGER
As of 20 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 17 inbound Pith citation observations for arXiv:2405.19332.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-20T06:33:59.587034+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-08-15T20:53:39.902673Z
A source-named dated measurement, never combined with another source.
Source: arxiv_reference, observed 2026-07-01T00:35:10.408592Z
0 of 0 outbound references displayed
External citation measurements
No source-named external measurement is stored.
No outbound reference observations are available for this paper version.
Observation 3dc9bcdb-6823-45d8-aaaf-6cd36051aa42 · inbound
DSTC: Direct Preference Learning with Only Self-Generated Tests and Code to Improve Code LMs Self-Exploring Language Models: Active Preference Elicitation for Online Alignment
Reference 46
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 770af17c-b057-4d2b-8d01-158888e17854 · inbound
Online Learning from Strategic Human Feedback in LLM Fine-Tuning Self-Exploring Language Models: Active Preference Elicitation for Online Alignment
Reference 15
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7173bdc4-e1e1-48a7-85ae-4dbfe67fac49 · inbound
Understanding the Logic of Direct Preference Alignment through Logic Self-Exploring Language Models: Active Preference Elicitation for Online Alignment
Reference 1975
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8589d06f-41f9-4b44-9136-39958e26ad64 · inbound
Online Preference Alignment for Language Models via Count-based Exploration Self-Exploring Language Models: Active Preference Elicitation for Online Alignment
Reference 73
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6626d737-9273-44dd-8a96-e0c78418a57c · inbound
BRiTE: Bootstrapping Reinforced Thinking Process to Enhance Language Model Reasoning Self-Exploring Language Models: Active Preference Elicitation for Online Alignment
Reference 69
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c09f1447-2d90-4e96-bf86-e8d9d5c36f80 · inbound
PILAF: Optimal Human Preference Sampling for Reward Modeling Self-Exploring Language Models: Active Preference Elicitation for Online Alignment
Reference 43
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 59137204-5150-4843-93cc-d2b788d199f2 · inbound
Outcome-Based Online Reinforcement Learning: Algorithms and Fundamental Limits Self-Exploring Language Models: Active Preference Elicitation for Online Alignment
Reference 52
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9a3474c4-bb56-4ffb-89d4-bf52cc09386e · inbound
Learning a Pessimistic Reward Model in RLHF Self-Exploring Language Models: Active Preference Elicitation for Online Alignment
Reference 27
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6247927f-d8b1-4abb-a09b-43bc9f775eb8 · inbound
Mutual-Taught for Co-adapting Policy and Reward Models Self-Exploring Language Models: Active Preference Elicitation for Online Alignment
Reference 36
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 95de08ac-82bd-4bc9-aeae-fa829c9d921b · inbound
MEMETRON: Metaheuristic Mechanisms for Test-time Response Optimization of Large Language Models Self-Exploring Language Models: Active Preference Elicitation for Online Alignment
Reference 5
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9fac93e6-af53-45fd-9ed7-fd92bd07a6d4 · inbound
Exploration from a Primal-Dual Lens: Value-Incentivized Actor-Critic Methods for Sample-Efficient Online RL Self-Exploring Language Models: Active Preference Elicitation for Online Alignment
Reference 19
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 243ea628-e798-41c4-acb3-3fda90368e0b · inbound
SharedRep-RLHF: A Shared Representation Approach to RLHF with Diverse Preferences Self-Exploring Language Models: Active Preference Elicitation for Online Alignment
Reference 35
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 07e30d42-0401-4a3e-98d7-9f776f3d8e3e · inbound
Outcome-based Exploration for LLM Reasoning Self-Exploring Language Models: Active Preference Elicitation for Online Alignment
Reference 36
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 22b91954-b84e-4ac6-a24d-098a88669763 · inbound
Data-dependent Exploration for Online Reinforcement Learning from Human Feedback Self-Exploring Language Models: Active Preference Elicitation for Online Alignment
Reference 90
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.
Observation 982bf544-798f-40b6-b81c-1003b1591837 · inbound
Data-dependent Exploration for Online Reinforcement Learning from Human Feedback Self-Exploring Language Models: Active Preference Elicitation for Online Alignment
Reference 36
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.
Observation 757cb0ef-18f2-4480-a604-30e6e4b799b9 · inbound
Recall Isn't Enough: Bounding Commitments in Personalized Language Systems Self-Exploring Language Models: Active Preference Elicitation for Online Alignment
Reference 12
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.
Observation 04e9207b-5337-4d7f-af8e-b06cb0c08b39 · inbound
Spectral Souping: A Unified Framework for Online Preference Alignment Self-Exploring Language Models: Active Preference Elicitation for Online Alignment
Reference 36
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.