Pith. sign in

Paper Citation Record · LEDGER

Procedural Fairness Failures in RLHF from Preference Averaging

As of 18 August 2026, this Paper Citation Record lists 8 of 8 outbound references and 0 inbound Pith citation observations for arXiv:2608.10126.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2608.10126 v1

Coverage vector

measured 8 of 8 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-14T04:15:38.248174Z

measured 8 of 8 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-18T06:34:40.430872+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

8 of 8 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved7
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch1

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation c92d9153-955f-4409-8c5f-ff86cac16482 · outbound

This paper cites Fine-Tuning Language Models from Human Preferences.

Procedural Fairness Failures in RLHF from Preference Averaging Fine-Tuning Language Models from Human Preferences

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-14T04:15:38.192422Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-14T04:15:38.192422Z digest=sha256:8fd0d64c03d3f2e9c683f7e305ed9586c41ff4b5785f7f710a1f7d5ab9696422

Observation d21df635-6ffb-4a7e-96ce-dc14de00ba1c · outbound

This paper cites Training language models to follow instructions with human feedback.

Procedural Fairness Failures in RLHF from Preference Averaging Training language models to follow instructions with human feedback

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-14T04:15:38.198111Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-14T04:15:38.198111Z digest=sha256:1da42cdb42dd8d253752a87b09340a8802a9124fd72b677d3bb893b518eae41f

Observation e3766b2a-9706-466f-a019-2b51725111a8 · outbound

This paper cites $\textit{Ab initio}$ dynamical mean-field theory with natural orbitals renormalization group impurity solver: Formalism and applications.

Procedural Fairness Failures in RLHF from Preference Averaging $\textit{Ab initio}$ dynamical mean-field theory with natural orbitals renormalization group impurity solver: Formalism and applications

Reference 3

Resolution
metadata mismatch
local_arxiv, observed 2026-08-14T04:15:38.406721Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-08-14T04:15:38.207974Z digest=sha256:a53ed70646bc0ece8d93040f9c3cec5d8d9c7741547c62a012ba38ac0dda71aa

Observation ec071893-9b6d-49e0-9113-6bd73884dc47 · outbound

This paper cites Latent Preference Coding: Aligning Large Language Models via Discrete Latent Codes.

Procedural Fairness Failures in RLHF from Preference Averaging Latent Preference Coding: Aligning Large Language Models via Discrete Latent Codes

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-14T04:15:38.214421Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-14T04:15:38.214421Z digest=sha256:a7fe4e5646ce6392600e3e1a2f6e82cfddf6555736f0c49fd7feba92d22d2032

Observation 006d9a92-e3d2-4474-ae74-e729b55f641a · outbound

This paper cites Controllable Preference Optimization: Toward Controllable Multi-Objective Alignment.

Procedural Fairness Failures in RLHF from Preference Averaging Controllable Preference Optimization: Toward Controllable Multi-Objective Alignment

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-14T04:15:38.222855Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-14T04:15:38.222855Z digest=sha256:1cb2d49bd79746ea9e8d95bf33df8d4b96e487b9b6c234cd2dfc40262dabbfae

Observation e19d68a7-45bd-4d12-a5aa-d51ef1e1a873 · outbound

This paper cites A Survey on Personalized and Pluralistic Preference Alignment in Large Language Models.

Procedural Fairness Failures in RLHF from Preference Averaging A Survey on Personalized and Pluralistic Preference Alignment in Large Language Models

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-14T04:15:38.236029Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-14T04:15:38.236029Z digest=sha256:4107411d06e6103669419131ea9788486b8a8c03052379be2517cf34d17316f4

Observation ca990641-9310-4004-a49f-2cf44770d3a6 · outbound

This paper cites an unresolved cited work.

Procedural Fairness Failures in RLHF from Preference Averaging Unresolved cited work

Reference 7

Resolution
unresolved
raw_fallback, observed 2026-08-14T04:15:38.473629Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-08-14T04:15:38.242008Z digest=sha256:2217f81d20ca37dbdd0a82f260a4ec85a66495eed5dd9ce765ffc6e9e6198732

Observation 00a91046-ce5b-4259-9251-58c85057b1f6 · outbound

This paper cites Personalisation within bounds: A risk taxonomy and policy framework for the alignment of large language models with personalised feedback.

Procedural Fairness Failures in RLHF from Preference Averaging Personalisation within bounds: A risk taxonomy and policy framework for the alignment of large language models with personalised feedback

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-14T04:15:38.248174Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-14T04:15:38.248174Z digest=sha256:1d97af0378c4603e92554cc1e87042033b85aaf4d0174cda1b51f137310d2e63

Pith citing papers

No inbound Pith citation observations are available.