Pith. sign in

Paper Citation Record · LEDGER

SLiC-HF: Sequence Likelihood Calibration with Human Feedback

As of 10 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 62 inbound Pith citation observations for arXiv:2305.10425.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2305.10425 v1

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 62 of 62 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-09T06:31:02.800959+00:00

measured 62 of 62 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-09T22:20:13.896255Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-04T11:09:46.433136Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation 1722cd2e-8a89-48d6-ba18-325f40167411 · inbound

Self-Rewarding Language Models cites this paper.

Self-Rewarding Language Models SLiC-HF: Sequence Likelihood Calibration with Human Feedback

Reference 124

Resolution
verified exact
arxiv_id, observed 2026-05-13T12:01:42.528194Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-05-13T12:01:42.290502Z digest=sha256:e9a01f94ba622b2dc91993301face0e2cf4166cb6c681321051edc967d8aa000

Observation 5e276e74-323e-4467-a46d-42e1c255fd92 · inbound

KTO: Model Alignment as Prospect Theoretic Optimization cites this paper.

KTO: Model Alignment as Prospect Theoretic Optimization SLiC-HF: Sequence Likelihood Calibration with Human Feedback

Reference 24

Resolution
verified exact
arxiv_id, observed 2026-05-12T12:17:53.546101Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-12T12:17:53.478052Z digest=sha256:533a095b055c9348015d2dae8f8e6b2942605878994e810be42e1f5ffb3bf85e

Observation 34a278a0-618b-4598-9c74-270390f42fdb · inbound

Inference Scaling Laws: An Empirical Analysis of Compute-Optimal Inference for Problem-Solving with Language Models cites this paper.

Inference Scaling Laws: An Empirical Analysis of Compute-Optimal Inference for Problem-Solving with Language Models SLiC-HF: Sequence Likelihood Calibration with Human Feedback

Reference 237

Resolution
metadata mismatch
arxiv_id, observed 2026-05-18T06:38:37.016336Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-05-18T06:38:36.517935Z digest=sha256:fd1f70ffb1186e615629fb5447f2fbea36c9d958cec29d21bf53459466b314f6

Observation 75b6c38f-8957-4816-bbcb-19d670c555ee · inbound

Agent Q: Advanced Reasoning and Learning for Autonomous AI Agents cites this paper.

Agent Q: Advanced Reasoning and Learning for Autonomous AI Agents SLiC-HF: Sequence Likelihood Calibration with Human Feedback

Reference 210

Resolution
metadata mismatch
arxiv_id, observed 2026-05-20T09:42:04.390390Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-05-20T09:41:59.979595Z digest=sha256:18c77807fe76c2aa8ff0f6caf1f924687f88ed4dc1508312ae700fa26a2e6644

Observation 7a0c075c-9a59-4218-876e-38c6449f733a · inbound

Preference Goal Tuning: Post-Training as Latent Control for Frozen Policies cites this paper.

Preference Goal Tuning: Post-Training as Latent Control for Frozen Policies SLiC-HF: Sequence Likelihood Calibration with Human Feedback

Reference 58

Resolution
verified exact
arxiv_id, observed 2026-05-23T08:22:44.243462Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-05-23T08:20:05.898025Z digest=sha256:fdbe8f75fcbd6094c44d72b8401b74ab5e33e22e8ccc259d9a9e0bc460a54f7a

Observation fb60df69-5da4-4d2d-bc3d-6c928af3da1a · inbound

BRiTE: Bootstrapping Reinforced Thinking Process to Enhance Language Model Reasoning cites this paper.

BRiTE: Bootstrapping Reinforced Thinking Process to Enhance Language Model Reasoning SLiC-HF: Sequence Likelihood Calibration with Human Feedback

Reference 70

Resolution
unresolved
no resolver link, observed 2026-08-09T22:20:13.896255Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T22:20:13.896255Z digest=sha256:78ae7c0effc418669f7b235c3f2b8c6dee18b597f347c29842022348fadebced

Observation 049fc6f9-58a1-4cf3-831d-abe63ed3fee8 · inbound

A Checks-and-Balances Framework for Context-Aware Ethical AI Alignment cites this paper.

A Checks-and-Balances Framework for Context-Aware Ethical AI Alignment SLiC-HF: Sequence Likelihood Calibration with Human Feedback

Reference 75

Resolution
unresolved
no resolver link, observed 2026-08-09T20:09:54.404741Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T20:09:54.404741Z digest=sha256:594a58d2255f4f05a163757f20a5e6f87931ae17a9b64d90a6b1dbd4334743d0

Observation a6eccc06-70ba-46d1-a3e0-5158e1fbea3c · inbound

On Almost Surely Safe Alignment of Large Language Models at Inference-Time cites this paper.

On Almost Surely Safe Alignment of Large Language Models at Inference-Time SLiC-HF: Sequence Likelihood Calibration with Human Feedback

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-09T16:18:40.583638Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T16:18:40.583638Z digest=sha256:6d2385cc09a046e8303d870b58167210f4626017c2edd2427e641a7028625160

Observation 52741b9a-7c71-44bf-85a1-fe1cb7775f6b · inbound

Reveal the Mystery of DPO: The Connection between DPO and RL Algorithms cites this paper.

Reveal the Mystery of DPO: The Connection between DPO and RL Algorithms SLiC-HF: Sequence Likelihood Calibration with Human Feedback

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-09T06:04:01.644525Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T06:04:01.644525Z digest=sha256:186105b3d7af57edde208a073e1d58ee8a54ffe30fda9ed8e613d01ad6e4ab34

Observation ff55ae22-d549-4c47-bb36-1e8685b45293 · inbound

LLM Alignment as Retriever Optimization: An Information Retrieval Perspective cites this paper.

LLM Alignment as Retriever Optimization: An Information Retrieval Perspective SLiC-HF: Sequence Likelihood Calibration with Human Feedback

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-09T04:08:52.221810Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T04:08:52.221810Z digest=sha256:12c1a5b37925e99b6fb588fffd8172177aad6f62832c7b9821edd618774da4aa

Observation 2380c056-2542-4170-8d4c-6ccb870716e3 · inbound

How Humans Help LLMs: Assessing and Incentivizing Human Preference Annotators cites this paper.

How Humans Help LLMs: Assessing and Incentivizing Human Preference Annotators SLiC-HF: Sequence Likelihood Calibration with Human Feedback

Reference 91

Resolution
verified exact
arxiv_id, observed 2026-05-23T03:57:29.564192Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-05-23T03:56:18.703995Z digest=sha256:a815ce360b4ff6c87144597591a464fe312e7577c9307d76e87e3593a9f03c09

Observation 3c00348f-ff9c-439e-a0ce-a04c3d03c991 · inbound

Design Considerations in Offline Preference-based RL cites this paper.

Design Considerations in Offline Preference-based RL SLiC-HF: Sequence Likelihood Calibration with Human Feedback

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-08T19:40:42.069597Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T19:40:42.069597Z digest=sha256:3bdef3d0a8d6f13aad6c7e28a9160ba924a80e58d1ced1bb06641a6eb8c586f1

Observation f423684f-2bfc-493e-9f0c-f0ec5536991d · inbound

Principled Data Selection for Alignment: The Hidden Risks of Difficult Examples cites this paper.

Principled Data Selection for Alignment: The Hidden Risks of Difficult Examples SLiC-HF: Sequence Likelihood Calibration with Human Feedback

Reference 89

Resolution
unresolved
no resolver link, observed 2026-08-08T11:57:12.067249Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T11:57:12.067249Z digest=sha256:21625e9b8369ed6e9009b7198aeb9a6890615a46f5f5509fe47e051e18bd1f9a

Observation 59175734-fc8a-42e0-9c75-894c1cea7c81 · inbound

Reinforcement Learning from Human Feedback cites this paper.

Reinforcement Learning from Human Feedback SLiC-HF: Sequence Likelihood Calibration with Human Feedback

Reference 193

Resolution
verified exact
arxiv_id, observed 2026-05-22T19:32:01.402266Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-22T19:27:40.991325Z digest=sha256:80eeec783b9544b82ef9f69230d3465d5282fdaee025e26709c7df7747dea685

Observation f2cf9eaa-5d02-454e-91db-cc41a6877c14 · inbound

Shallow Preference Signals: Large Language Model Aligns Even Better with Truncated Data? cites this paper.

Shallow Preference Signals: Large Language Model Aligns Even Better with Truncated Data? SLiC-HF: Sequence Likelihood Calibration with Human Feedback

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-07T15:15:46.299308Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:15:46.299308Z digest=sha256:1c6df0e18e454aac57f9cbc642143cae3319ad6cb460d1dff33a02b4cee30dce

Observation 87bd096e-8ccb-4955-abd4-69db5d72d6bc · inbound

Incentivizing High-Quality Human Annotations with Golden Questions cites this paper.

Incentivizing High-Quality Human Annotations with Golden Questions SLiC-HF: Sequence Likelihood Calibration with Human Feedback

Reference 37

Resolution
verified exact
arxiv_id, observed 2026-05-19T13:42:19.370731Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-19T13:41:26.730528Z digest=sha256:49e4e9e13a6f0e6ac182d342296039e1f4be977f04880aa90ea374c9484b5910

Observation 0f25a33c-c908-4cf8-a2fd-d6b57a926c3f · inbound

Square$\chi$PO: Differentially Private and Robust $\chi^2$-Preference Optimization in Offline Direct Alignment cites this paper.

Square$\chi$PO: Differentially Private and Robust $\chi^2$-Preference Optimization in Offline Direct Alignment SLiC-HF: Sequence Likelihood Calibration with Human Feedback

Reference 96

Resolution
unresolved
no resolver link, observed 2026-08-07T13:45:04.554366Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:45:04.554366Z digest=sha256:97caa8f62f744750211014c3d7ec217d2f3741e1b5c45f79f7a25a36cd7305fe

Observation 0b8a432e-e368-4c56-b05f-05153a48c8e2 · inbound

Thompson Sampling in Online RLHF with General Function Approximation cites this paper.

Thompson Sampling in Online RLHF with General Function Approximation SLiC-HF: Sequence Likelihood Calibration with Human Feedback

Reference 58

Resolution
unresolved
no resolver link, observed 2026-08-07T12:43:50.251391Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:43:50.251391Z digest=sha256:09794b3969b0b6e3e06a22770692251220262d04c58272546bbf34ea707fdcae

Observation 821c1dfc-b893-4828-acbd-f8130eb23105 · inbound

On Symmetric Losses for Robust Policy Optimization with Noisy Preferences cites this paper.

On Symmetric Losses for Robust Policy Optimization with Noisy Preferences SLiC-HF: Sequence Likelihood Calibration with Human Feedback

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-07T12:27:57.177745Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:27:57.177745Z digest=sha256:c664f38c82cf8aefeb2a68505f2fdcb9f650d32bfbb8c05171c7b6a0a2408a83

Observation faf8bccb-9cd4-44ef-92c4-992109603475 · inbound

Sign-SZPO: Provable Preference-based Reinforcement Learning with an Unknown Link Function cites this paper.

Sign-SZPO: Provable Preference-based Reinforcement Learning with an Unknown Link Function SLiC-HF: Sequence Likelihood Calibration with Human Feedback

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-07T11:20:13.316647Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T11:20:13.316647Z digest=sha256:9492c770341551b35ed2a006e2449d665473005bd02309eb8b2ed61223e388be

Observation e928c41e-609e-4942-909e-de9dc75b6c28 · inbound

BPO: Revisiting Preference Modeling in Direct Preference Optimization cites this paper.

BPO: Revisiting Preference Modeling in Direct Preference Optimization SLiC-HF: Sequence Likelihood Calibration with Human Feedback

Reference 28

Resolution
malformed identifier
no resolver link, observed 2026-08-07T11:07:35.619682Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:07:35.619682Z digest=sha256:7ee5ba3ce0e567462e7d9a2dd77e8e28f6b2e4e828e3eacd74d43cb32392831b

Observation 96a4e02e-511b-44ed-9d04-e4057629f929 · inbound

Explicit Preference Optimization: No Need for an Implicit Reward Model cites this paper.

Explicit Preference Optimization: No Need for an Implicit Reward Model SLiC-HF: Sequence Likelihood Calibration with Human Feedback

Reference 55

Resolution
unresolved
no resolver link, observed 2026-08-07T05:40:18.036861Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T05:40:18.036861Z digest=sha256:93859e3049fc73521f17c488528e12be29f6af85f76f5649a5ff88e709bf16ff

Observation 1dc57259-667c-4d60-b5b1-f7b8a9949c6b · inbound

ConfPO: Exploiting Policy Model Confidence for Critical Token Selection in Preference Optimization cites this paper.

ConfPO: Exploiting Policy Model Confidence for Critical Token Selection in Preference Optimization SLiC-HF: Sequence Likelihood Calibration with Human Feedback

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-07T05:14:45.342641Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T05:14:45.342641Z digest=sha256:91936b82b69679518c2af11dfecc972cc7978d63a00c03ccfdde17d1c6829025

Observation ad212efc-b520-4d1a-86e3-86662d50493f · inbound

On Monotonicity in AI Alignment cites this paper.

On Monotonicity in AI Alignment SLiC-HF: Sequence Likelihood Calibration with Human Feedback

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-07T05:08:05.322977Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:08:05.322977Z digest=sha256:2fb45e6cc54ed563d912f6463fb6c20cf545d647b29aa3c28ec5c37f4edde43f

Observation 9520ea74-8889-43c0-bd1b-372e99a61b44 · inbound

Debiasing Online Preference Learning via Preference Feature Preservation cites this paper.

Debiasing Online Preference Learning via Preference Feature Preservation SLiC-HF: Sequence Likelihood Calibration with Human Feedback

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-07T06:05:32.720004Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T06:05:32.720004Z digest=sha256:dcaf4ce801f2950da45b325c2360148d7123d3c1e6f9952c65734a6ebad46feb

Observation 586052fa-c7b1-42f5-b7d7-8e467d6e9ebc · inbound

Value-Free Policy Optimization via Reward Partitioning cites this paper.

Value-Free Policy Optimization via Reward Partitioning SLiC-HF: Sequence Likelihood Calibration with Human Feedback

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-07T00:31:49.293738Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:31:49.293738Z digest=sha256:97cc7d7103ccb052b10bad5fb5415b25f52d1d25b2069dd3b4e5b866fa186c41

Observation 7b97a484-904c-4c66-8d4a-e5b1f170f0c2 · inbound

Optimising Language Models for Downstream Tasks: A Post-Training Perspective cites this paper.

Optimising Language Models for Downstream Tasks: A Post-Training Perspective SLiC-HF: Sequence Likelihood Calibration with Human Feedback

Reference 274

Resolution
unresolved
no resolver link, observed 2026-08-06T22:44:44.816609Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:44:44.816609Z digest=sha256:130ff763ea09457e41ca2b260bfde55e39ed6046c6aa2d5814199107f17d6d2f

Observation 665991a4-8154-4b0a-a47d-ecd103f3e2eb · inbound

Implicit Reward as the Bridge: A Unified View of SFT and DPO Connections cites this paper.

Implicit Reward as the Bridge: A Unified View of SFT and DPO Connections SLiC-HF: Sequence Likelihood Calibration with Human Feedback

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-07T00:54:49.868787Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:54:49.868787Z digest=sha256:33ee35adc07db445ea6586cbd57161f23e2916cdca851145169e15c589246295

Observation 75cfad5b-8491-433a-a1af-4a7ef81d8a73 · inbound

Improving Consistency in Vehicle Trajectory Prediction Through Preference Optimization cites this paper.

Improving Consistency in Vehicle Trajectory Prediction Through Preference Optimization SLiC-HF: Sequence Likelihood Calibration with Human Feedback

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-06T20:34:37.335053Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:34:37.335053Z digest=sha256:4e989c88b805891fd530c84be076507c03eb13474413636f44b181ba9da31ca9

Observation 1036f3d1-a9c8-4f11-bc1c-117c7058c47c · inbound

LLM Hypnosis: Exploiting User Feedback for Unauthorized Knowledge Injection to All Users cites this paper.

LLM Hypnosis: Exploiting User Feedback for Unauthorized Knowledge Injection to All Users SLiC-HF: Sequence Likelihood Calibration with Human Feedback

Reference 21

Resolution
verified exact
arxiv_id, observed 2026-05-19T06:02:07.792836Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-19T05:58:17.452837Z digest=sha256:60f4061ff1dc005b2153662bc056ee7ec2de2676e6a18b0a38e5f0091272a68c

Observation 6b833df4-92a0-4977-88b7-29c850e08e7f · inbound

DPO Unchained: Your Training Algorithm is Secretly Disentangled in Human Choice Theory (and its Loss' Convexity is Dispensable) cites this paper.

DPO Unchained: Your Training Algorithm is Secretly Disentangled in Human Choice Theory (and its Loss' Convexity is Dispensable) SLiC-HF: Sequence Likelihood Calibration with Human Feedback

Reference 67

Resolution
unresolved
no resolver link, observed 2026-08-06T18:52:38.808890Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T18:52:38.808890Z digest=sha256:587ac57bcebf950ba49ff2113abd95761f737ae1c3fc90733e364294355fb57f

Observation d00b723e-d4c4-46eb-9e1d-42352f6ba700 · inbound

Phi-Ground Tech Report: Advancing Perception in GUI Grounding cites this paper.

Phi-Ground Tech Report: Advancing Perception in GUI Grounding SLiC-HF: Sequence Likelihood Calibration with Human Feedback

Reference 70

Resolution
unresolved
no resolver link, observed 2026-08-06T10:28:56.119117Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T10:28:56.119117Z digest=sha256:1dad1ef3da3e1d9c711bcbb20182e03f1ad4d77ee74d8067c9436f65b0b5e038

Observation 5fd91298-c1f0-4ab7-ae4c-bcc8c72fe381 · inbound

Enhancing Small LLM Alignment through Margin-Based Objective Modifications under Resource Constraints cites this paper.

Enhancing Small LLM Alignment through Margin-Based Objective Modifications under Resource Constraints SLiC-HF: Sequence Likelihood Calibration with Human Feedback

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-05T21:33:14.871864Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T21:33:14.871864Z digest=sha256:105556d8c6b1e600f1e062d5b7ccff0f301500c7efefa125d141572481155c87

Observation 9727fb8b-240f-40a7-a8a8-5838076e0ff6 · inbound

Failure Modes of Maximum Entropy RLHF cites this paper.

Failure Modes of Maximum Entropy RLHF SLiC-HF: Sequence Likelihood Calibration with Human Feedback

Reference 60

Resolution
verified exact
arxiv_id, observed 2026-05-18T14:02:39.779960Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-05-18T14:02:11.084514Z digest=sha256:3fb2cd095b3d7c00310b926c348464baf65e88266f950cd5e7cc868e27df70a2

Observation dca28929-3147-44b8-80dc-00812939bbea · inbound

Adaptive Margin RLHF via Preference over Preferences cites this paper.

Adaptive Margin RLHF via Preference over Preferences SLiC-HF: Sequence Likelihood Calibration with Human Feedback

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-04T14:52:23.843777Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T14:52:23.843777Z digest=sha256:1bcb25b750fe3c35e1085d29a7de6d1a90c40a1f0c090a1cfb8023720fa787a1

Observation 6d45353d-31f5-45da-842c-4dfe6de29502 · inbound

Alignment-Aware Decoding cites this paper.

Alignment-Aware Decoding SLiC-HF: Sequence Likelihood Calibration with Human Feedback

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-04T13:40:15.787091Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T13:40:15.787091Z digest=sha256:5fb249c5ef50a42ca9757bb5cc585aa872c042a62cb96d23a6eaf0bf39dea5ae

Observation 1409158b-53b4-4237-b651-722e6ea4b7a9 · inbound

POPI: Personalizing LLMs via Optimized Natural Language Preference Inference cites this paper.

POPI: Personalizing LLMs via Optimized Natural Language Preference Inference SLiC-HF: Sequence Likelihood Calibration with Human Feedback

Reference 54

Resolution
verified exact
arxiv_id, observed 2026-05-18T05:42:24.390403Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-18T05:41:58.231139Z digest=sha256:2faa5285f47a0f8f4b1fa0cd2561163c0159e4b895ebb45e858c6243b25bf16d

Observation c4a76026-141a-4769-8298-c6042f584ffc · inbound

Representation-Guided Parameter-Efficient LLM Unlearning cites this paper.

Representation-Guided Parameter-Efficient LLM Unlearning SLiC-HF: Sequence Likelihood Calibration with Human Feedback

Reference 209

Resolution
metadata mismatch
arxiv_id, observed 2026-05-10T06:06:19.162389Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-05-10T06:01:46.885030Z digest=sha256:99eb4493f638c944d8f7bda6e5ecb6108dd8fc9a0c379bc59ebcfee97a75aefa

Observation 0783e4e0-6a46-4245-9168-026877db7356 · inbound

Mind the Gap: Structure-Aware Consistency in Preference Learning cites this paper.

Mind the Gap: Structure-Aware Consistency in Preference Learning SLiC-HF: Sequence Likelihood Calibration with Human Feedback

Reference 63

Resolution
verified exact
arxiv_id, observed 2026-05-12T10:16:28.593236Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-05-07T06:41:18.311986Z digest=sha256:a348fd36d67f522f353facf08243c497e73743c7078df022fff44b2860f62430

Observation b7f30dcf-b92d-47e0-9d7e-b27dc150c79b · inbound

Anomaly-Preference Image Generation cites this paper.

Anomaly-Preference Image Generation SLiC-HF: Sequence Likelihood Calibration with Human Feedback

Reference 24

Resolution
metadata mismatch
arxiv_id, observed 2026-05-09T06:15:37.409124Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-05-08T18:45:20.508200Z digest=sha256:03e0e8df96186b773983843709ede391fb9bc40425335b3cee76fb7c30220115

Observation 2eb63ced-9368-49ee-a8e6-807028c53ccc · inbound

Anomaly-Preference Image Generation cites this paper.

Anomaly-Preference Image Generation SLiC-HF: Sequence Likelihood Calibration with Human Feedback

Reference 24

Resolution
metadata mismatch
arxiv_id, observed 2026-05-21T00:43:53.717724Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-05-21T00:39:31.730684Z digest=sha256:b5ac90729d9439f6e9ab421613d9e531625838bb66df96f3cd1bab16cda6ba4b

Observation 1e1e8f3b-289b-407d-b2a8-157ed91bcedc · inbound

Anomaly-Preference Image Generation cites this paper.

Anomaly-Preference Image Generation SLiC-HF: Sequence Likelihood Calibration with Human Feedback

Reference 8

Resolution
verified exact
arxiv_id, observed 2026-07-01T00:45:12.283055Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-07-01T00:36:26.957598Z digest=sha256:db49bc717607b9216c5eebdfcb9f7333d453e7b0d43e3ef41c60e57caf49833d

Observation 7fa8adaa-2ecf-40ef-bb7e-38081db26d93 · inbound

Data-dependent Exploration for Online Reinforcement Learning from Human Feedback cites this paper.

Data-dependent Exploration for Online Reinforcement Learning from Human Feedback SLiC-HF: Sequence Likelihood Calibration with Human Feedback

Reference 94

Resolution
verified exact
arxiv_id, observed 2026-05-11T17:51:09.182781Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-05-08T17:01:04.571087Z digest=sha256:813a51da0c38eef3fae848b1ea96fc6bbe69735055e1c52191e6f2c29e1d9d5a

Observation 754f1ef1-1628-4e0b-b4f3-c6b36a820e0c · inbound

Data-dependent Exploration for Online Reinforcement Learning from Human Feedback cites this paper.

Data-dependent Exploration for Online Reinforcement Learning from Human Feedback SLiC-HF: Sequence Likelihood Calibration with Human Feedback

Reference 40

Resolution
verified exact
arxiv_id, observed 2026-07-01T00:35:10.407105Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-07-01T00:27:39.321158Z digest=sha256:47e373d2c1a3d1886317f024b77a67d8647368e7dd05af69325d4ef3c1531464

Observation 75f08799-ea08-4a97-9d8d-68b5a3dcbaff · inbound

CROP: Expert-Aligned Image Cropping via Compositional Reasoning and Optimizing Preference cites this paper.

CROP: Expert-Aligned Image Cropping via Compositional Reasoning and Optimizing Preference SLiC-HF: Sequence Likelihood Calibration with Human Feedback

Reference 12

Resolution
verified exact
arxiv_id, observed 2026-05-14T21:38:00.841961Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-14T21:36:26.958957Z digest=sha256:aaba69b1d2a6efefd6f5d5de763ddc33ebe12980d3984e04fdd78f0f7fdca55a

Observation 188984cc-6ede-4346-90f0-ba835b7ebe66 · inbound

Conditional Equivalence of DPO and RLHF: Implicit Assumption, Failure Modes, and Provable Alignment cites this paper.

Conditional Equivalence of DPO and RLHF: Implicit Assumption, Failure Modes, and Provable Alignment SLiC-HF: Sequence Likelihood Calibration with Human Feedback

Reference 35

Resolution
metadata mismatch
arxiv_id, observed 2026-05-21T05:03:58.124603Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-05-21T04:59:41.187000Z digest=sha256:f9b3784db1c30f36bdc89ec7ba64a3f7b4135e2120cbbf5f215547cc493c062f

Observation fde0b802-f605-407f-921a-0821d188ab28 · inbound

Token-weighted Direct Preference Optimization with Attention cites this paper.

Token-weighted Direct Preference Optimization with Attention SLiC-HF: Sequence Likelihood Calibration with Human Feedback

Reference 22

Resolution
verified exact
arxiv_id, observed 2026-05-22T07:06:11.750006Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-05-22T07:05:23.189975Z digest=sha256:87e2e6928ee78238f996d023be96e657547f3da47560ba1bdd18e5c3f02086c9

Observation b6439581-fc26-4740-9979-7b9337f814fe · inbound

S-SPPO: Semantic-Calibrated Self-Play Preference Optimization cites this paper.

S-SPPO: Semantic-Calibrated Self-Play Preference Optimization SLiC-HF: Sequence Likelihood Calibration with Human Feedback

Reference 22

Resolution
verified exact
arxiv_id, observed 2026-07-01T22:56:19.435524Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-06-28T14:58:33.167891Z digest=sha256:536c9b08e49d17cf0969afc26d0dfcafd6cace18566c7035c018f7e8767d6e2c

Observation 496bb23e-2733-4db9-a00b-e336c57b3c5e · inbound

P$^2$-DPO: Grounding Hallucination in Perceptual Processing via Calibration Direct Preference Optimization cites this paper.

P$^2$-DPO: Grounding Hallucination in Perceptual Processing via Calibration Direct Preference Optimization SLiC-HF: Sequence Likelihood Calibration with Human Feedback

Reference 119

Resolution
verified exact
arxiv_id, observed 2026-07-02T02:06:27.135124Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-06-28T11:13:37.291834Z digest=sha256:db2a314aed39ad12c769666ca65e75844267a80f35e44725d04cdc92766d4203

Observation 72de1279-8562-4fca-b409-83c5916d3e1e · inbound

MARD: Mirror-Augmented Reasoning Distillation for Mechanism-Level Drug-Drug Interaction Prediction cites this paper.

MARD: Mirror-Augmented Reasoning Distillation for Mechanism-Level Drug-Drug Interaction Prediction SLiC-HF: Sequence Likelihood Calibration with Human Feedback

Reference 88

Resolution
metadata mismatch
arxiv_id, observed 2026-07-03T10:58:03.338464Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-06-27T09:44:38.871250Z digest=sha256:288c45d91c00d45975cc8b2989c5199480ebe9f6cbe5d06d28c486de6f2833a0

Observation 06574a1f-879b-48d7-b34f-95e68bfb9f58 · inbound

Emo-LiPO: Listwise Preference Optimization for Fine-Grained Emotion Intensity Control in LLM-based Text-to-Speech cites this paper.

Emo-LiPO: Listwise Preference Optimization for Fine-Grained Emotion Intensity Control in LLM-based Text-to-Speech SLiC-HF: Sequence Likelihood Calibration with Human Feedback

Reference 31

Resolution
verified exact
arxiv_id, observed 2026-07-03T16:08:37.812263Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-06-27T06:01:06.123832Z digest=sha256:ea8ec4181556635dd0cca794c0fb76e41111fbf5d87194707320b16983af55b3

Observation e3deae7e-03de-46fb-9a72-0ded20e4c8be · inbound

The Hitchhiker's Guide to Agentic AI: From Foundations to Systems cites this paper.

The Hitchhiker's Guide to Agentic AI: From Foundations to Systems SLiC-HF: Sequence Likelihood Calibration with Human Feedback

Reference 194

Resolution
verified exact
arxiv_id, observed 2026-07-04T11:09:46.434801Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-06-26T08:09:57.542558Z digest=sha256:c36fcfd1897816642eeed74070c6d047ea105f2d36f6269a04a82d35b510fbb9

Observation b1c3db5a-96df-4011-a5a9-6bf7ac7609b3 · inbound

The Hitchhiker's Guide to Agentic AI: From Foundations to Systems cites this paper.

The Hitchhiker's Guide to Agentic AI: From Foundations to Systems SLiC-HF: Sequence Likelihood Calibration with Human Feedback

Reference 194

Resolution
unresolved
no resolver link, observed 2026-08-02T10:27:18.422406Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T10:27:18.422406Z digest=sha256:34b55ea603c323fc85e39ec0ec2432bf54ccf9d2977c146e43fea9ca20a646be

Observation 8fc7b84b-5fce-481d-8dc9-a489f7d6da51 · inbound

Vision-driven Preference Synthesis for Mitigating Hallucinations in VLMs cites this paper.

Vision-driven Preference Synthesis for Mitigating Hallucinations in VLMs SLiC-HF: Sequence Likelihood Calibration with Human Feedback

Reference 56

Resolution
verified exact
arxiv_id, observed 2026-07-01T15:35:47.677633Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-06-30T01:22:16.176398Z digest=sha256:3c76184c6409f9873222db45460e675817f92cab2b5b4732aa6789136682e9d5

Observation 8051384f-8274-450f-b576-80cc0349ec02 · inbound

Your Data Manifold is Secretly a Reward Model: Shell-LCC for Text-to-Video Generation cites this paper.

Your Data Manifold is Secretly a Reward Model: Shell-LCC for Text-to-Video Generation SLiC-HF: Sequence Likelihood Calibration with Human Feedback

Reference 48

Resolution
metadata mismatch
arxiv_id, observed 2026-06-30T08:14:26.741596Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-06-30T06:07:44.343539Z digest=sha256:faee78c9fc7927d545c85266fc0bd8a011fed44f8d962750196bc1108c62009b

Observation cccff79b-33e6-4cc0-bb99-7ab96de8cd7b · inbound

Your Data Manifold is Secretly a Reward Model: Shell-LCC for Text-to-Video Generation cites this paper.

Your Data Manifold is Secretly a Reward Model: Shell-LCC for Text-to-Video Generation SLiC-HF: Sequence Likelihood Calibration with Human Feedback

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-02T09:37:24.616723Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T09:37:24.616723Z digest=sha256:cff887847b41e76c708fda90ed41ba03b881d3fca39020da595567a5150d6038

Observation 1470d54a-8ab3-4f29-b8b4-1ffcaa06c8cf · inbound

Unbiased Alignment for Large Language Models with Noisy Preferences cites this paper.

Unbiased Alignment for Large Language Models with Noisy Preferences SLiC-HF: Sequence Likelihood Calibration with Human Feedback

Reference 12

Resolution
unresolved
no resolver link, observed 2026-07-12T03:50:30.002103Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T03:50:30.002103Z digest=sha256:60cddd4b01e17c8e68b3c4cafd17ec2c7f5eeeda5e06184027bcbc746f2127e6

Observation c5def73b-34c0-4807-b171-8bde50c6a6f6 · inbound

Multi-Turn On-Policy Distillation with Prefix Replay cites this paper.

Multi-Turn On-Policy Distillation with Prefix Replay SLiC-HF: Sequence Likelihood Calibration with Human Feedback

Reference 254

Resolution
unresolved
no resolver link, observed 2026-07-11T13:53:36.775836Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-11T13:53:36.775836Z digest=sha256:88022c81dd8d2d57d4184f69cb174fdd500e4507a750e63af71875771c93a0cf

Observation af3df9e8-9307-4cec-bf01-c72357f13696 · inbound

Multi-Turn On-Policy Distillation with Prefix Replay cites this paper.

Multi-Turn On-Policy Distillation with Prefix Replay SLiC-HF: Sequence Likelihood Calibration with Human Feedback

Reference 255

Resolution
unresolved
no resolver link, observed 2026-08-02T08:41:01.968095Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-02T08:41:01.968095Z digest=sha256:ed40a6ba5be4399a4aaa57a42602c435db672c1e5713da8386c36b1dc59ff010

Observation fe93fbcb-ec48-420e-a923-cb85d810674c · inbound

Every Sample Counts: Supervised Fine-Tuning of Language Models with Pointwise Constraints cites this paper.

Every Sample Counts: Supervised Fine-Tuning of Language Models with Pointwise Constraints SLiC-HF: Sequence Likelihood Calibration with Human Feedback

Reference 35

Resolution
unresolved
no resolver link, observed 2026-07-13T05:26:06.829382Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-13T05:26:06.829382Z digest=sha256:0dcfb2db4b7144c2f0c4d2d7b523fc1581f8dda9b624a138b94e6e16d25402b8

Observation 8921611f-4095-440e-86e1-14e4d0f270dd · inbound

Normalized Rewards for Preference Optimization cites this paper.

Normalized Rewards for Preference Optimization SLiC-HF: Sequence Likelihood Calibration with Human Feedback

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-02T10:01:59.868237Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T10:01:59.868237Z digest=sha256:3ea2e5970597b52dbcca13c8002e37d6c3119281e99d841d80e64b066fdd9386

Observation bf7d0bd4-da96-4800-b1c7-887d73d5295c · inbound

Test-Time Scaling via Error Localization cites this paper.

Test-Time Scaling via Error Localization SLiC-HF: Sequence Likelihood Calibration with Human Feedback

Reference 173

Resolution
unresolved
no resolver link, observed 2026-08-01T07:28:35.893789Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T07:28:35.893789Z digest=sha256:7f309aa547ee833344d6cb18fe677faf5f177516a2f763a08c4cbcb573699fd6