Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links
Paper Citation Record · LEDGER
As of 10 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 62 inbound Pith citation observations for arXiv:2305.10425.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-09T06:31:02.800959+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-08-09T22:20:13.896255Z
A source-named dated measurement, never combined with another source.
Source: arxiv_reference, observed 2026-07-04T11:09:46.433136Z
0 of 0 outbound references displayed
External citation measurements
No source-named external measurement is stored.
No outbound reference observations are available for this paper version.
Observation 1722cd2e-8a89-48d6-ba18-325f40167411 · inbound
Self-Rewarding Language Models SLiC-HF: Sequence Likelihood Calibration with Human Feedback
Reference 124
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 5e276e74-323e-4467-a46d-42e1c255fd92 · inbound
KTO: Model Alignment as Prospect Theoretic Optimization SLiC-HF: Sequence Likelihood Calibration with Human Feedback
Reference 24
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 34a278a0-618b-4598-9c74-270390f42fdb · inbound
Inference Scaling Laws: An Empirical Analysis of Compute-Optimal Inference for Problem-Solving with Language Models SLiC-HF: Sequence Likelihood Calibration with Human Feedback
Reference 237
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 75b6c38f-8957-4816-bbcb-19d670c555ee · inbound
Agent Q: Advanced Reasoning and Learning for Autonomous AI Agents SLiC-HF: Sequence Likelihood Calibration with Human Feedback
Reference 210
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 7a0c075c-9a59-4218-876e-38c6449f733a · inbound
Preference Goal Tuning: Post-Training as Latent Control for Frozen Policies SLiC-HF: Sequence Likelihood Calibration with Human Feedback
Reference 58
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation fb60df69-5da4-4d2d-bc3d-6c928af3da1a · inbound
BRiTE: Bootstrapping Reinforced Thinking Process to Enhance Language Model Reasoning SLiC-HF: Sequence Likelihood Calibration with Human Feedback
Reference 70
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 049fc6f9-58a1-4cf3-831d-abe63ed3fee8 · inbound
A Checks-and-Balances Framework for Context-Aware Ethical AI Alignment SLiC-HF: Sequence Likelihood Calibration with Human Feedback
Reference 75
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a6eccc06-70ba-46d1-a3e0-5158e1fbea3c · inbound
On Almost Surely Safe Alignment of Large Language Models at Inference-Time SLiC-HF: Sequence Likelihood Calibration with Human Feedback
Reference 23
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 52741b9a-7c71-44bf-85a1-fe1cb7775f6b · inbound
Reveal the Mystery of DPO: The Connection between DPO and RL Algorithms SLiC-HF: Sequence Likelihood Calibration with Human Feedback
Reference 38
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ff55ae22-d549-4c47-bb36-1e8685b45293 · inbound
LLM Alignment as Retriever Optimization: An Information Retrieval Perspective SLiC-HF: Sequence Likelihood Calibration with Human Feedback
Reference 34
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2380c056-2542-4170-8d4c-6ccb870716e3 · inbound
How Humans Help LLMs: Assessing and Incentivizing Human Preference Annotators SLiC-HF: Sequence Likelihood Calibration with Human Feedback
Reference 91
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 3c00348f-ff9c-439e-a0ce-a04c3d03c991 · inbound
Design Considerations in Offline Preference-based RL SLiC-HF: Sequence Likelihood Calibration with Human Feedback
Reference 19
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f423684f-2bfc-493e-9f0c-f0ec5536991d · inbound
Principled Data Selection for Alignment: The Hidden Risks of Difficult Examples SLiC-HF: Sequence Likelihood Calibration with Human Feedback
Reference 89
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 59175734-fc8a-42e0-9c75-894c1cea7c81 · inbound
Reinforcement Learning from Human Feedback SLiC-HF: Sequence Likelihood Calibration with Human Feedback
Reference 193
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation f2cf9eaa-5d02-454e-91db-cc41a6877c14 · inbound
Shallow Preference Signals: Large Language Model Aligns Even Better with Truncated Data? SLiC-HF: Sequence Likelihood Calibration with Human Feedback
Reference 25
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 87bd096e-8ccb-4955-abd4-69db5d72d6bc · inbound
Incentivizing High-Quality Human Annotations with Golden Questions SLiC-HF: Sequence Likelihood Calibration with Human Feedback
Reference 37
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 0f25a33c-c908-4cf8-a2fd-d6b57a926c3f · inbound
Square$\chi$PO: Differentially Private and Robust $\chi^2$-Preference Optimization in Offline Direct Alignment SLiC-HF: Sequence Likelihood Calibration with Human Feedback
Reference 96
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0b8a432e-e368-4c56-b05f-05153a48c8e2 · inbound
Thompson Sampling in Online RLHF with General Function Approximation SLiC-HF: Sequence Likelihood Calibration with Human Feedback
Reference 58
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 821c1dfc-b893-4828-acbd-f8130eb23105 · inbound
On Symmetric Losses for Robust Policy Optimization with Noisy Preferences SLiC-HF: Sequence Likelihood Calibration with Human Feedback
Reference 41
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation faf8bccb-9cd4-44ef-92c4-992109603475 · inbound
Sign-SZPO: Provable Preference-based Reinforcement Learning with an Unknown Link Function SLiC-HF: Sequence Likelihood Calibration with Human Feedback
Reference 40
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e928c41e-609e-4942-909e-de9dc75b6c28 · inbound
BPO: Revisiting Preference Modeling in Direct Preference Optimization SLiC-HF: Sequence Likelihood Calibration with Human Feedback
Reference 28
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 96a4e02e-511b-44ed-9d04-e4057629f929 · inbound
Explicit Preference Optimization: No Need for an Implicit Reward Model SLiC-HF: Sequence Likelihood Calibration with Human Feedback
Reference 55
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1dc57259-667c-4d60-b5b1-f7b8a9949c6b · inbound
ConfPO: Exploiting Policy Model Confidence for Critical Token Selection in Preference Optimization SLiC-HF: Sequence Likelihood Calibration with Human Feedback
Reference 46
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ad212efc-b520-4d1a-86e3-86662d50493f · inbound
On Monotonicity in AI Alignment SLiC-HF: Sequence Likelihood Calibration with Human Feedback
Reference 47
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9520ea74-8889-43c0-bd1b-372e99a61b44 · inbound
Debiasing Online Preference Learning via Preference Feature Preservation SLiC-HF: Sequence Likelihood Calibration with Human Feedback
Reference 47
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 586052fa-c7b1-42f5-b7d7-8e467d6e9ebc · inbound
Value-Free Policy Optimization via Reward Partitioning SLiC-HF: Sequence Likelihood Calibration with Human Feedback
Reference 19
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7b97a484-904c-4c66-8d4a-e5b1f170f0c2 · inbound
Optimising Language Models for Downstream Tasks: A Post-Training Perspective SLiC-HF: Sequence Likelihood Calibration with Human Feedback
Reference 274
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 665991a4-8154-4b0a-a47d-ecd103f3e2eb · inbound
Implicit Reward as the Bridge: A Unified View of SFT and DPO Connections SLiC-HF: Sequence Likelihood Calibration with Human Feedback
Reference 38
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 75cfad5b-8491-433a-a1af-4a7ef81d8a73 · inbound
Improving Consistency in Vehicle Trajectory Prediction Through Preference Optimization SLiC-HF: Sequence Likelihood Calibration with Human Feedback
Reference 28
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1036f3d1-a9c8-4f11-bc1c-117c7058c47c · inbound
LLM Hypnosis: Exploiting User Feedback for Unauthorized Knowledge Injection to All Users SLiC-HF: Sequence Likelihood Calibration with Human Feedback
Reference 21
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 6b833df4-92a0-4977-88b7-29c850e08e7f · inbound
DPO Unchained: Your Training Algorithm is Secretly Disentangled in Human Choice Theory (and its Loss' Convexity is Dispensable) SLiC-HF: Sequence Likelihood Calibration with Human Feedback
Reference 67
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d00b723e-d4c4-46eb-9e1d-42352f6ba700 · inbound
Phi-Ground Tech Report: Advancing Perception in GUI Grounding SLiC-HF: Sequence Likelihood Calibration with Human Feedback
Reference 70
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5fd91298-c1f0-4ab7-ae4c-bcc8c72fe381 · inbound
Enhancing Small LLM Alignment through Margin-Based Objective Modifications under Resource Constraints SLiC-HF: Sequence Likelihood Calibration with Human Feedback
Reference 18
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9727fb8b-240f-40a7-a8a8-5838076e0ff6 · inbound
Failure Modes of Maximum Entropy RLHF SLiC-HF: Sequence Likelihood Calibration with Human Feedback
Reference 60
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation dca28929-3147-44b8-80dc-00812939bbea · inbound
Adaptive Margin RLHF via Preference over Preferences SLiC-HF: Sequence Likelihood Calibration with Human Feedback
Reference 40
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6d45353d-31f5-45da-842c-4dfe6de29502 · inbound
Alignment-Aware Decoding SLiC-HF: Sequence Likelihood Calibration with Human Feedback
Reference 12
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1409158b-53b4-4237-b651-722e6ea4b7a9 · inbound
POPI: Personalizing LLMs via Optimized Natural Language Preference Inference SLiC-HF: Sequence Likelihood Calibration with Human Feedback
Reference 54
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation c4a76026-141a-4769-8298-c6042f584ffc · inbound
Representation-Guided Parameter-Efficient LLM Unlearning SLiC-HF: Sequence Likelihood Calibration with Human Feedback
Reference 209
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 0783e4e0-6a46-4245-9168-026877db7356 · inbound
Mind the Gap: Structure-Aware Consistency in Preference Learning SLiC-HF: Sequence Likelihood Calibration with Human Feedback
Reference 63
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation b7f30dcf-b92d-47e0-9d7e-b27dc150c79b · inbound
Anomaly-Preference Image Generation SLiC-HF: Sequence Likelihood Calibration with Human Feedback
Reference 24
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 2eb63ced-9368-49ee-a8e6-807028c53ccc · inbound
Anomaly-Preference Image Generation SLiC-HF: Sequence Likelihood Calibration with Human Feedback
Reference 24
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 1e1e8f3b-289b-407d-b2a8-157ed91bcedc · inbound
Anomaly-Preference Image Generation SLiC-HF: Sequence Likelihood Calibration with Human Feedback
Reference 8
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 7fa8adaa-2ecf-40ef-bb7e-38081db26d93 · inbound
Data-dependent Exploration for Online Reinforcement Learning from Human Feedback SLiC-HF: Sequence Likelihood Calibration with Human Feedback
Reference 94
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 754f1ef1-1628-4e0b-b4f3-c6b36a820e0c · inbound
Data-dependent Exploration for Online Reinforcement Learning from Human Feedback SLiC-HF: Sequence Likelihood Calibration with Human Feedback
Reference 40
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 75f08799-ea08-4a97-9d8d-68b5a3dcbaff · inbound
CROP: Expert-Aligned Image Cropping via Compositional Reasoning and Optimizing Preference SLiC-HF: Sequence Likelihood Calibration with Human Feedback
Reference 12
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 188984cc-6ede-4346-90f0-ba835b7ebe66 · inbound
Conditional Equivalence of DPO and RLHF: Implicit Assumption, Failure Modes, and Provable Alignment SLiC-HF: Sequence Likelihood Calibration with Human Feedback
Reference 35
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation fde0b802-f605-407f-921a-0821d188ab28 · inbound
Token-weighted Direct Preference Optimization with Attention SLiC-HF: Sequence Likelihood Calibration with Human Feedback
Reference 22
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation b6439581-fc26-4740-9979-7b9337f814fe · inbound
S-SPPO: Semantic-Calibrated Self-Play Preference Optimization SLiC-HF: Sequence Likelihood Calibration with Human Feedback
Reference 22
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 496bb23e-2733-4db9-a00b-e336c57b3c5e · inbound
P$^2$-DPO: Grounding Hallucination in Perceptual Processing via Calibration Direct Preference Optimization SLiC-HF: Sequence Likelihood Calibration with Human Feedback
Reference 119
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 72de1279-8562-4fca-b409-83c5916d3e1e · inbound
MARD: Mirror-Augmented Reasoning Distillation for Mechanism-Level Drug-Drug Interaction Prediction SLiC-HF: Sequence Likelihood Calibration with Human Feedback
Reference 88
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 06574a1f-879b-48d7-b34f-95e68bfb9f58 · inbound
Emo-LiPO: Listwise Preference Optimization for Fine-Grained Emotion Intensity Control in LLM-based Text-to-Speech SLiC-HF: Sequence Likelihood Calibration with Human Feedback
Reference 31
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation e3deae7e-03de-46fb-9a72-0ded20e4c8be · inbound
The Hitchhiker's Guide to Agentic AI: From Foundations to Systems SLiC-HF: Sequence Likelihood Calibration with Human Feedback
Reference 194
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation b1c3db5a-96df-4011-a5a9-6bf7ac7609b3 · inbound
The Hitchhiker's Guide to Agentic AI: From Foundations to Systems SLiC-HF: Sequence Likelihood Calibration with Human Feedback
Reference 194
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8fc7b84b-5fce-481d-8dc9-a489f7d6da51 · inbound
Vision-driven Preference Synthesis for Mitigating Hallucinations in VLMs SLiC-HF: Sequence Likelihood Calibration with Human Feedback
Reference 56
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 8051384f-8274-450f-b576-80cc0349ec02 · inbound
Your Data Manifold is Secretly a Reward Model: Shell-LCC for Text-to-Video Generation SLiC-HF: Sequence Likelihood Calibration with Human Feedback
Reference 48
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation cccff79b-33e6-4cc0-bb99-7ab96de8cd7b · inbound
Your Data Manifold is Secretly a Reward Model: Shell-LCC for Text-to-Video Generation SLiC-HF: Sequence Likelihood Calibration with Human Feedback
Reference 48
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1470d54a-8ab3-4f29-b8b4-1ffcaa06c8cf · inbound
Unbiased Alignment for Large Language Models with Noisy Preferences SLiC-HF: Sequence Likelihood Calibration with Human Feedback
Reference 12
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c5def73b-34c0-4807-b171-8bde50c6a6f6 · inbound
Multi-Turn On-Policy Distillation with Prefix Replay SLiC-HF: Sequence Likelihood Calibration with Human Feedback
Reference 254
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation af3df9e8-9307-4cec-bf01-c72357f13696 · inbound
Multi-Turn On-Policy Distillation with Prefix Replay SLiC-HF: Sequence Likelihood Calibration with Human Feedback
Reference 255
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation fe93fbcb-ec48-420e-a923-cb85d810674c · inbound
Every Sample Counts: Supervised Fine-Tuning of Language Models with Pointwise Constraints SLiC-HF: Sequence Likelihood Calibration with Human Feedback
Reference 35
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8921611f-4095-440e-86e1-14e4d0f270dd · inbound
Normalized Rewards for Preference Optimization SLiC-HF: Sequence Likelihood Calibration with Human Feedback
Reference 24
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation bf7d0bd4-da96-4800-b1c7-887d73d5295c · inbound
Test-Time Scaling via Error Localization SLiC-HF: Sequence Likelihood Calibration with Human Feedback
Reference 173
Source-reported events for the cited work
Unavailable: canonical work link unavailable.