Pith. sign in

Paper Citation Record · LEDGER

Procedural Fairness Failures in RLHF from Preference Averaging

As of 19 August 2026, this Paper Citation Record lists 8 of 8 outbound references and 0 inbound Pith citation observations for arXiv:2608.10126.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2608.10126 v1

Coverage vector

measured 8 of 8 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-14T04:15:38.248174Z

measured 8 of 8 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-19T06:32:44.657259+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

8 of 8 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved7
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch1

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation c92d9153-955f-4409-8c5f-ff86cac16482 · outbound

This paper cites Fine-Tuning Language Models from Human Preferences.

Procedural Fairness Failures in RLHF from Preference Averaging Fine-Tuning Language Models from Human Preferences

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-14T04:15:38.192422Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-14T04:15:38.192422Z digest=sha256:40957b8ddd015bcf60a2f57bd4b8ac368c1d091084559ba85e983e19b5465d43

Observation d21df635-6ffb-4a7e-96ce-dc14de00ba1c · outbound

This paper cites Training language models to follow instructions with human feedback.

Procedural Fairness Failures in RLHF from Preference Averaging Training language models to follow instructions with human feedback

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-14T04:15:38.198111Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-14T04:15:38.198111Z digest=sha256:d503a6144f55d80b4efb97bf07886492e3d5cdc4e9abe17bb55c5e7d38114aa2

Observation e3766b2a-9706-466f-a019-2b51725111a8 · outbound

This paper cites $\textit{Ab initio}$ dynamical mean-field theory with natural orbitals renormalization group impurity solver: Formalism and applications.

Procedural Fairness Failures in RLHF from Preference Averaging $\textit{Ab initio}$ dynamical mean-field theory with natural orbitals renormalization group impurity solver: Formalism and applications

Reference 3

Resolution
metadata mismatch
local_arxiv, observed 2026-08-14T04:15:38.406721Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-08-14T04:15:38.207974Z digest=sha256:0c841a285fd0a46b2194a5e692d8db60c1ff2eaf9f75b299f2fca78f62cb5323

Observation ec071893-9b6d-49e0-9113-6bd73884dc47 · outbound

This paper cites Latent Preference Coding: Aligning Large Language Models via Discrete Latent Codes.

Procedural Fairness Failures in RLHF from Preference Averaging Latent Preference Coding: Aligning Large Language Models via Discrete Latent Codes

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-14T04:15:38.214421Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-14T04:15:38.214421Z digest=sha256:6e5c27d1fe11c3f9b4367a65ed54d07f53c008efd5aba9e9cb92713418b6da0a

Observation 006d9a92-e3d2-4474-ae74-e729b55f641a · outbound

This paper cites Controllable Preference Optimization: Toward Controllable Multi-Objective Alignment.

Procedural Fairness Failures in RLHF from Preference Averaging Controllable Preference Optimization: Toward Controllable Multi-Objective Alignment

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-14T04:15:38.222855Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-14T04:15:38.222855Z digest=sha256:46058f91dcd8a8be485fbc98b9e7f80c6cf57630fc0b519ed576a5a33b8aaf37

Observation e19d68a7-45bd-4d12-a5aa-d51ef1e1a873 · outbound

This paper cites A Survey on Personalized and Pluralistic Preference Alignment in Large Language Models.

Procedural Fairness Failures in RLHF from Preference Averaging A Survey on Personalized and Pluralistic Preference Alignment in Large Language Models

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-14T04:15:38.236029Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-14T04:15:38.236029Z digest=sha256:82292c009ce41295ca4027621cc80535020547749f26f64acd1c4f969669704a

Observation ca990641-9310-4004-a49f-2cf44770d3a6 · outbound

This paper cites an unresolved cited work.

Procedural Fairness Failures in RLHF from Preference Averaging Unresolved cited work

Reference 7

Resolution
unresolved
raw_fallback, observed 2026-08-14T04:15:38.473629Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-08-14T04:15:38.242008Z digest=sha256:2fa3af8577dd61991ffc5cb00746d486ee03514c34ac6714de52771314cf5372

Observation 00a91046-ce5b-4259-9251-58c85057b1f6 · outbound

This paper cites Personalisation within bounds: A risk taxonomy and policy framework for the alignment of large language models with personalised feedback.

Procedural Fairness Failures in RLHF from Preference Averaging Personalisation within bounds: A risk taxonomy and policy framework for the alignment of large language models with personalised feedback

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-14T04:15:38.248174Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-14T04:15:38.248174Z digest=sha256:be3ccb6f9bdd7e97d03dbeb0ba21dea8c14638081c7e40e700c9e33d8fc8b3e1

Pith citing papers

No inbound Pith citation observations are available.