Pith. sign in

Paper Citation Record · LEDGER

The Alignment Ceiling: Objective Mismatch in Reinforcement Learning from Human Feedback

As of 18 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 11 inbound Pith citation observations for arXiv:2311.00168.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2311.00168 v2

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 11 of 11 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-18T06:34:40.430872+00:00

measured 11 of 11 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-16T11:44:51.027151Z

measured 1 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-08-05T02:28:24.338817Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

3
arxiv_reference, observed 2026-08-05T02:28:24.338817Z

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation 681a28bd-bd91-48d0-9441-54926090c9a0 · inbound

A Roadmap to Pluralistic Alignment cites this paper.

A Roadmap to Pluralistic Alignment The Alignment Ceiling: Objective Mismatch in Reinforcement Learning from Human Feedback

Reference 195

Resolution
verified exact
arxiv_id, observed 2026-05-16T14:37:53.473020Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-05-16T14:37:53.279275Z digest=sha256:3de3aaa815d8a2b72a89d0811139a7bf694142c4027c63d76fe78cb81dc2283b

Observation 6e4466c0-7a78-42d1-8a35-87b0fad540b6 · inbound

Drowning in Documents: Consequences of Scaling Reranker Inference cites this paper.

Drowning in Documents: Consequences of Scaling Reranker Inference The Alignment Ceiling: Objective Mismatch in Reinforcement Learning from Human Feedback

Reference 53

Resolution
unresolved
no resolver link, observed 2026-08-12T18:13:43.336223Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T18:13:43.336223Z digest=sha256:2159012465e01593f4e053d6f0f353839cf8d3cc399d5f930f03ddb5e93a7e2c

Observation 18fe0b8c-1553-4484-875b-0d72e477049e · inbound

Disentangling Length Bias In Preference Learning Via Response-Conditioned Modeling cites this paper.

Disentangling Length Bias In Preference Learning Via Response-Conditioned Modeling The Alignment Ceiling: Objective Mismatch in Reinforcement Learning from Human Feedback

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-09T17:46:29.206702Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T17:46:29.206702Z digest=sha256:11643b326b8d9998f68c52d1f90b43ca1cf948e4455b35827123de27be9f06db

Observation b5762376-55e6-4638-8078-b0f1ab623c98 · inbound

Establishing Reliability Metrics for Reward Models in Large Language Models cites this paper.

Establishing Reliability Metrics for Reward Models in Large Language Models The Alignment Ceiling: Objective Mismatch in Reinforcement Learning from Human Feedback

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-16T11:44:51.027151Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T11:44:51.027151Z digest=sha256:40439d71956c0f4ab6bc1daf33859f75f355d95ed943a9651ed3cace0b6261c8

Observation 2d3a95c6-e250-4a0b-8ff2-59dfa125a0ef · inbound

Bias Fitting to Mitigate Length Bias of Reward Model in RLHF cites this paper.

Bias Fitting to Mitigate Length Bias of Reward Model in RLHF The Alignment Ceiling: Objective Mismatch in Reinforcement Learning from Human Feedback

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-15T20:31:01.230482Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:31:01.230482Z digest=sha256:bb20c1515130aadd0f85a5f3f0fae6df16227d4a59f450bebd7104df08100750

Observation 92b21f53-dc94-49df-ad26-ea6491b33297 · inbound

Reward Hacking in the Era of Large Models: Mechanisms, Emergent Misalignment, Challenges cites this paper.

Reward Hacking in the Era of Large Models: Mechanisms, Emergent Misalignment, Challenges The Alignment Ceiling: Objective Mismatch in Reinforcement Learning from Human Feedback

Reference 60

Resolution
metadata mismatch
arxiv_id, observed 2026-05-10T14:00:28.451656Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-10T13:58:53.430492Z digest=sha256:345c6ba7316537eedc13d39178bba1d85c99f6afa9280d41b3c957676e64ef83

Observation d37aecc7-4f51-4543-9e9a-8999ad36b06b · inbound

Towards Order Fairness: Mitigating LLMs Order Sensitivity through Dual Group Advantage Optimization cites this paper.

Towards Order Fairness: Mitigating LLMs Order Sensitivity through Dual Group Advantage Optimization The Alignment Ceiling: Objective Mismatch in Reinforcement Learning from Human Feedback

Reference 46

Resolution
verified exact
arxiv_id, observed 2026-05-13T07:37:29.854732Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-05-13T07:32:58.404947Z digest=sha256:783b89df9910fc6a253b81df03dc4209eb9364b0d662f8e852d371402658720c

Observation e0b4ec5b-9979-40ce-b916-2db4a4c0f7ba · inbound

In-Context Reward Adaptation for Robust Preference Modeling cites this paper.

In-Context Reward Adaptation for Robust Preference Modeling The Alignment Ceiling: Objective Mismatch in Reinforcement Learning from Human Feedback

Reference 7

Resolution
verified exact
arxiv_id, observed 2026-06-29T08:23:15.182550Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-06-29T08:19:55.177440Z digest=sha256:ec1a96e5e8d54fbaa25b096d77f192c49b442bde596201135060615c1610e6f7

Observation 53c5fe0a-59ab-4df1-88b6-f88cf772a7f3 · inbound

What Do People Actually Want From AI? Mapping Preference Plurality cites this paper.

What Do People Actually Want From AI? Mapping Preference Plurality The Alignment Ceiling: Objective Mismatch in Reinforcement Learning from Human Feedback

Reference 55

Resolution
verified exact
arxiv_id, observed 2026-06-28T01:41:29.695819Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-06-28T01:32:04.660400Z digest=sha256:db44a3b3d9ed2c671eaf888f1f3ab6dbf922b776fffc10f1e5c30eaacd0978bc

Observation 65ac556c-c160-40ef-a9ca-dfe763c34ee8 · inbound

Proxy Reward Internalization and Mechanistic Exploitation: A Learned Precursor to Reward Hacking and Its Generalization cites this paper.

Proxy Reward Internalization and Mechanistic Exploitation: A Learned Precursor to Reward Hacking and Its Generalization The Alignment Ceiling: Objective Mismatch in Reinforcement Learning from Human Feedback

Reference 228

Resolution
verified exact
arxiv_id, observed 2026-07-03T01:37:30.492205Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-06-27T16:26:34.918099Z digest=sha256:2219c846717a4a07e1f22c2abe51a6f656d43e4503e7af7a3e7d9e17873d17b1

Observation 82d42bf6-2cca-4297-ae8f-a32229b733be · inbound

Exponential Reward Weighting for Fine-Tuning Generative Recommenders under Sparse and Noisy Feedback cites this paper.

Exponential Reward Weighting for Fine-Tuning Generative Recommenders under Sparse and Noisy Feedback The Alignment Ceiling: Objective Mismatch in Reinforcement Learning from Human Feedback

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-05T00:19:14.540226Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T00:19:14.540226Z digest=sha256:48aa358ed089c9ea3d089fa6beb383acc25825076479c1a6f9d033747ed00858