Pith. sign in

Paper Citation Record · LEDGER

Distributional Preference Learning: Understanding and Accounting for Hidden Context in RLHF

As of 18 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 16 inbound Pith citation observations for arXiv:2312.08358.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2312.08358 v2

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 16 of 16 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-18T06:34:40.430872+00:00

measured 16 of 16 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-12T18:15:15.978088Z

measured 1 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-08-05T02:28:24.338817Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

2
arxiv_reference, observed 2026-08-05T02:28:24.338817Z

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation 1f529d69-7010-4884-8f41-2a12335c311a · inbound

Mapping out the Space of Human Feedback for Reinforcement Learning: A Conceptual Framework cites this paper.

Mapping out the Space of Human Feedback for Reinforcement Learning: A Conceptual Framework Distributional Preference Learning: Understanding and Accounting for Hidden Context in RLHF

Reference 171

Resolution
unresolved
no resolver link, observed 2026-08-12T18:15:15.978088Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T18:15:15.978088Z digest=sha256:e8ebaba316a98581572d35f2100835027cf808e0855a1d9ce6d6ce5ba6b8e093

Observation 23783648-074d-49db-887b-bd1fdb950640 · inbound

Beyond the Binary: Capturing Diverse Preferences With Reward Regularization cites this paper.

Beyond the Binary: Capturing Diverse Preferences With Reward Regularization Distributional Preference Learning: Understanding and Accounting for Hidden Context in RLHF

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-11T22:06:51.247150Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T22:06:51.247150Z digest=sha256:b3475800a468ace93712af24d09445b416299acbcac05080987a3b1c3ecef881

Observation c0531dae-9d5b-44ea-8ec1-4e8c62143004 · inbound

Test-Time Alignment via Hypothesis Reweighting cites this paper.

Test-Time Alignment via Hypothesis Reweighting Distributional Preference Learning: Understanding and Accounting for Hidden Context in RLHF

Reference 54

Resolution
verified exact
arxiv_id, observed 2026-05-23T06:57:40.535424Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-05-23T06:55:54.051821Z digest=sha256:a9506ca0b46b800cd3ed6383a9a684a28a59b7ecc228865f777ce3054ed1842d

Observation 60d3818d-8a1a-445f-bb9a-ecb28acf6327 · inbound

Clone-Robust AI Alignment cites this paper.

Clone-Robust AI Alignment Distributional Preference Learning: Understanding and Accounting for Hidden Context in RLHF

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-10T20:15:11.833637Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T20:15:11.833637Z digest=sha256:38d6c9f8ef75064b4f1effa54c126ebade01601bf18eafa28f16b3cfc952adcf

Observation d07cd496-0aa3-4874-aa36-cd583d2aca4a · inbound

Examining Alignment of Large Language Models through Representative Heuristics: The Case of Political Stereotypes cites this paper.

Examining Alignment of Large Language Models through Representative Heuristics: The Case of Political Stereotypes Distributional Preference Learning: Understanding and Accounting for Hidden Context in RLHF

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-10T15:18:19.167631Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T15:18:19.167631Z digest=sha256:61ddcafe1c4a52cb8f6c5a1b2af1abc3ede2f41e4c856e239040625200c26d75

Observation a4dada43-e749-4918-a3d3-c68069742f65 · inbound

Jackpot! Alignment as a Maximal Lottery cites this paper.

Jackpot! Alignment as a Maximal Lottery Distributional Preference Learning: Understanding and Accounting for Hidden Context in RLHF

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-09T20:51:59.197859Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T20:51:59.197859Z digest=sha256:85e7950d345d63589e04f3e239093d5dfbaa644bc26ef07e9cc20390813f2163

Observation 809a079d-4dea-4418-bbbf-009af9095b07 · inbound

CTR-Driven Advertising Image Generation with Multimodal Large Language Models cites this paper.

CTR-Driven Advertising Image Generation with Multimodal Large Language Models Distributional Preference Learning: Understanding and Accounting for Hidden Context in RLHF

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-09T10:21:04.459437Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T10:21:04.459437Z digest=sha256:d41d2234c9bd127742e10a837652b5f61db430582c210124f1e367bcd9c6c00b

Observation 5d19d2f2-1c7b-4921-a9fd-08f0ea002ba9 · inbound

Distortion of AI Alignment: Does Preference Optimization Optimize for Preferences? cites this paper.

Distortion of AI Alignment: Does Preference Optimization Optimize for Preferences? Distributional Preference Learning: Understanding and Accounting for Hidden Context in RLHF

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-07T12:49:10.543390Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:49:10.543390Z digest=sha256:647ae5612f20eab916c4e99ba949f2f41339b9178f2d006cb1f54e55f149e52b

Observation 699e2fc7-0744-4fe0-9bf4-189127a6bc72 · inbound

Theoretical Tensions in RLHF: Reconciling Empirical Success with Inconsistencies in Social Choice Theory cites this paper.

Theoretical Tensions in RLHF: Reconciling Empirical Success with Inconsistencies in Social Choice Theory Distributional Preference Learning: Understanding and Accounting for Hidden Context in RLHF

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-07T01:04:32.028901Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T01:04:32.028901Z digest=sha256:3135681df89c32542dfd637ecc2cc33722b8235331c74c773181c3300314e3dd

Observation 129db947-a33a-4513-93e2-67f981934587 · inbound

Inverse Reinforcement Learning Meets Large Language Model Post-Training: Basics, Advances, and Opportunities cites this paper.

Inverse Reinforcement Learning Meets Large Language Model Post-Training: Basics, Advances, and Opportunities Distributional Preference Learning: Understanding and Accounting for Hidden Context in RLHF

Reference 85

Resolution
unresolved
no resolver link, observed 2026-08-06T16:34:25.197977Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T16:34:25.197977Z digest=sha256:d5be7295f1a2744e4ff94af36bece526ad814a13311b188c406a2e09e18fb8e2

Observation e20a2f2a-12a6-4e2e-8637-5519f87425ea · inbound

Active Query Selection for Crowd-Based Reinforcement Learning cites this paper.

Active Query Selection for Crowd-Based Reinforcement Learning Distributional Preference Learning: Understanding and Accounting for Hidden Context in RLHF

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-05T16:02:26.007783Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T16:02:26.007783Z digest=sha256:a76f869a17ee40653b21bed8bb5fcb8955e9f549e0cab636d4e36c7922190c92

Observation 1abf0f85-8f73-4ef6-801e-8190f4448a73 · inbound

RLHF May Not Reflect Genuine Preferences cites this paper.

RLHF May Not Reflect Genuine Preferences Distributional Preference Learning: Understanding and Accounting for Hidden Context in RLHF

Reference 9

Resolution
verified exact
arxiv_id, observed 2026-05-16T08:40:46.322374Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-16T08:39:08.880486Z digest=sha256:0bdceb97fcc71b98d60a0d85f4e0c7ecdbb1b294d9d493beeec24f0a70223a31

Observation 4fc5fe13-7443-4442-8716-83ee117809d9 · inbound

Efficient Personalization of Generative User Interfaces cites this paper.

Efficient Personalization of Generative User Interfaces Distributional Preference Learning: Understanding and Accounting for Hidden Context in RLHF

Reference 89

Resolution
verified exact
arxiv_id, observed 2026-05-11T08:00:57.559506Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-10T16:52:34.799158Z digest=sha256:c3b15581297d7e7ea0b0b94a7ca531a4df6cba070eb77280d76a20f46d23f305

Observation 06cf540a-7373-4691-b261-15b869a81079 · inbound

Relative Principals, Pluralistic Alignment, and the Structural Value Alignment Problem cites this paper.

Relative Principals, Pluralistic Alignment, and the Structural Value Alignment Problem Distributional Preference Learning: Understanding and Accounting for Hidden Context in RLHF

Reference 71

Resolution
verified exact
arxiv_id, observed 2026-05-09T23:04:17.696412Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-09T23:02:34.564375Z digest=sha256:37ce30b64a5e0178ca5f2de32e319e86736ec5a9bf85a63bf1f889bff1efc9ff

Observation a4e556fa-f511-4cbe-b4cd-41ca01a10e4c · inbound

Personalization Meets Safety:Mechanisms,Risks,and Mitigations in Personalized LLMs cites this paper.

Personalization Meets Safety:Mechanisms,Risks,and Mitigations in Personalized LLMs Distributional Preference Learning: Understanding and Accounting for Hidden Context in RLHF

Reference 186

Resolution
metadata mismatch
arxiv_id, observed 2026-06-27T16:51:06.005505Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-06-27T16:49:14.243931Z digest=sha256:564dea00e01c80ad1959a463296a1d98dfe34db7e6ee65bd7f02c43197d047ea

Observation 7c1484f3-54f5-4e27-b71b-48038f27c286 · inbound

The Periodic Table of LLM Reasoning: A Structured Survey of Reasoning Paradigms, Methods, and Failure Modes cites this paper.

The Periodic Table of LLM Reasoning: A Structured Survey of Reasoning Paradigms, Methods, and Failure Modes Distributional Preference Learning: Understanding and Accounting for Hidden Context in RLHF

Reference 215

Resolution
verified exact
arxiv_id, observed 2026-07-03T05:57:41.259581Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-06-27T12:59:51.091008Z digest=sha256:016cf8aaad68d23b1371e631ce783d63662cf6e45cb66759b8b62527e69c6904