Pith. sign in

Paper Citation Record · LEDGER

RLHF from Heterogeneous Feedback via Personalization and Preference Aggregation

As of 9 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 15 inbound Pith citation observations for arXiv:2405.00254.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2405.00254 v2

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 15 of 15 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-09T06:31:02.800959+00:00

measured 15 of 15 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-07T14:31:05.229990Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-03T18:28:48.418933Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation 310957cd-3dfd-4981-9928-d4c190692c08 · inbound

MOSLIM:Align with diverse preferences in prompts through reward classification cites this paper.

MOSLIM:Align with diverse preferences in prompts through reward classification RLHF from Heterogeneous Feedback via Personalization and Preference Aggregation

Reference 2022

Resolution
unresolved
no resolver link, observed 2026-08-07T14:31:05.229990Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:31:05.229990Z digest=sha256:b76f867890d28ed35c8b8ee0c3a5c2ace349ccc52c5318ce8fc6d2b6902e284f

Observation a3775ccd-ef83-4694-be14-831abd0f9ccf · inbound

Distortion of AI Alignment: Does Preference Optimization Optimize for Preferences? cites this paper.

Distortion of AI Alignment: Does Preference Optimization Optimize for Preferences? RLHF from Heterogeneous Feedback via Personalization and Preference Aggregation

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-07T12:49:09.533133Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:49:09.533133Z digest=sha256:d7dd385b46668b39cc4152574f65c2bbec0aa88ba834d88ad647a650e973ac83

Observation 88259934-527d-46dc-b40b-0a6c150812d2 · inbound

Theoretical Tensions in RLHF: Reconciling Empirical Success with Inconsistencies in Social Choice Theory cites this paper.

Theoretical Tensions in RLHF: Reconciling Empirical Success with Inconsistencies in Social Choice Theory RLHF from Heterogeneous Feedback via Personalization and Preference Aggregation

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-07T01:04:31.735488Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T01:04:31.735488Z digest=sha256:f90db0449b4c87bf2831fad15a7ace56672dac54ecd5657258b421b090d18128

Observation 14f7efb8-29bc-47d7-8750-2cda7b46eed0 · inbound

Uncertainty Quantification for Ranking with Heterogeneous Preferences cites this paper.

Uncertainty Quantification for Ranking with Heterogeneous Preferences RLHF from Heterogeneous Feedback via Personalization and Preference Aggregation

Reference 57

Resolution
unresolved
no resolver link, observed 2026-08-05T12:13:14.351163Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T12:13:14.351163Z digest=sha256:a41ab774ecb61d6b4643a7308c9881fc82f3411ef9db9c318cdd91c6e8cce13e

Observation 239942ba-700d-42ff-ba00-a3e876b9be0d · inbound

T-POP: Test-Time Personalization with Online Preference Feedback cites this paper.

T-POP: Test-Time Personalization with Online Preference Feedback RLHF from Heterogeneous Feedback via Personalization and Preference Aggregation

Reference 2022

Resolution
unresolved
no resolver link, observed 2026-08-04T13:52:12.190762Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T13:52:12.190762Z digest=sha256:d100981c4aeede7350748da08f83a8d8c5543ce7bcc5645cd153152181629658

Observation fae086ed-9e72-481f-b3f7-ab60576aacda · inbound

Collaborative and Efficient Fine-tuning: Leveraging Task Similarity cites this paper.

Collaborative and Efficient Fine-tuning: Leveraging Task Similarity RLHF from Heterogeneous Feedback via Personalization and Preference Aggregation

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-03T03:49:40.071442Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T03:49:40.071442Z digest=sha256:3da26ddc7b12832b1082d626eb24ee03ef73eeb0a2b6713a0606fdcafbb8ef23

Observation b952dc43-dab5-4a02-ba32-3f648395960e · inbound

Context Engineering: A Practitioner Methodology for Structured Human-AI Collaboration cites this paper.

Context Engineering: A Practitioner Methodology for Structured Human-AI Collaboration RLHF from Heterogeneous Feedback via Personalization and Preference Aggregation

Reference 5

Resolution
unresolved
no resolver link, observed 2026-07-13T10:27:28.560210Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-13T10:27:28.560210Z digest=sha256:80f0d35c3953c4c4b197ca6a885b055d92372b3da6b53652fc65adc9b0429e31

Observation 7e3424c1-875a-4fa4-8a9c-49bc081ee610 · inbound

Curated Synthetic Data Doesn't Have to Collapse: A Theoretical Study of Generative Retraining with Pluralistic Preferences cites this paper.

Curated Synthetic Data Doesn't Have to Collapse: A Theoretical Study of Generative Retraining with Pluralistic Preferences RLHF from Heterogeneous Feedback via Personalization and Preference Aggregation

Reference 117

Resolution
verified exact
arxiv_id, observed 2026-05-11T02:30:54.490353Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-05-11T02:30:14.693348Z digest=sha256:05c9233cb3c70d5128c50847b4a61e81fed08fc307808fdd724c070efbfd937f

Observation 9aa0016b-0d44-44a2-b1bb-81773a671c81 · inbound

Curated Synthetic Data Doesn't Have to Collapse: A Theoretical Study of Generative Retraining with Pluralistic Preferences cites this paper.

Curated Synthetic Data Doesn't Have to Collapse: A Theoretical Study of Generative Retraining with Pluralistic Preferences RLHF from Heterogeneous Feedback via Personalization and Preference Aggregation

Reference 1

Resolution
verified exact
arxiv_id, observed 2026-07-01T13:35:45.953490Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-06-30T23:07:51.683581Z digest=sha256:b83fa0d84b08893b056c46f537ab1247d580deab0629a2ce54ae4f40e3a31847

Observation 6c2f6044-ea51-484e-b2e7-792df8212b65 · inbound

In-Context Reward Adaptation for Robust Preference Modeling cites this paper.

In-Context Reward Adaptation for Robust Preference Modeling RLHF from Heterogeneous Feedback via Personalization and Preference Aggregation

Reference 10

Resolution
verified exact
arxiv_id, observed 2026-06-29T08:23:15.196063Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-06-29T08:19:55.177440Z digest=sha256:0056c240461ec4a3072536c8f5431b56ccda87eb8fe90bda4f3466c793c70578

Observation f6755436-d58a-473e-bcd9-f504a92a7e35 · inbound

Inverse Reinforcement Learning without an Optimal Demonstrator: A Feasible Reward Set Approach cites this paper.

Inverse Reinforcement Learning without an Optimal Demonstrator: A Feasible Reward Set Approach RLHF from Heterogeneous Feedback via Personalization and Preference Aggregation

Reference 4

Resolution
verified exact
arxiv_id, observed 2026-06-29T00:02:49.757731Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-06-28T23:58:10.258001Z digest=sha256:9a05e11824280806bfe9910da2a62786d1f4007406ed5a4f324e74446ab51df5

Observation 6798f687-da7b-45f6-9998-02404fd32324 · inbound

PAFO: Pareto Fairness Optimization for Personalized Reward Modeling cites this paper.

PAFO: Pareto Fairness Optimization for Personalized Reward Modeling RLHF from Heterogeneous Feedback via Personalization and Preference Aggregation

Reference 38

Resolution
verified exact
arxiv_id, observed 2026-07-02T20:57:23.632083Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-06-27T20:00:05.900814Z digest=sha256:67cbc925ff608b11c335ed93dcf3643b5031b7a9b5e5436ddf6bcd83e08447dc

Observation 93c26ac0-237b-4446-9be5-d9e50eaf5f3d · inbound

Hidden Consensus:Preference-Validity Compression in Human Feedback cites this paper.

Hidden Consensus:Preference-Validity Compression in Human Feedback RLHF from Heterogeneous Feedback via Personalization and Preference Aggregation

Reference 2

Resolution
metadata mismatch
arxiv_id, observed 2026-07-03T05:07:39.038708Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-06-27T13:25:24.761133Z digest=sha256:a382be2ec124bda21913edda475b5fb999ddc4ce3fc6407aa7a6d415f1d3084a

Observation ccad351d-7b56-4eb5-8122-4fc3ec027d10 · inbound

CoPersona: Collaborative Persona Graphs for Robust LLM Personalization cites this paper.

CoPersona: Collaborative Persona Graphs for Robust LLM Personalization RLHF from Heterogeneous Feedback via Personalization and Preference Aggregation

Reference 42

Resolution
verified exact
arxiv_id, observed 2026-07-03T18:28:48.420560Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-07-03T18:20:53.930801Z digest=sha256:651ee731c73f6c9ffccd7b3cd6af20f46e2d8a6992052ff94a388a1d2b49d9d7

Observation 7d7103a2-8431-4882-8504-9b91fd8f6d1e · inbound

Training with (Swap) Regret Loss in a Single-Layer Self-Attention Model: A Case Study on the Probability Simplex cites this paper.

Training with (Swap) Regret Loss in a Single-Layer Self-Attention Model: A Case Study on the Probability Simplex RLHF from Heterogeneous Feedback via Personalization and Preference Aggregation

Reference 83

Resolution
unresolved
no resolver link, observed 2026-07-31T23:51:56.233532Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-31T23:51:56.233532Z digest=sha256:2d7d003e48e1fd15eb33023a7cd75a298d8f8064aa2b132eef091ed29c34af4a