Pith. sign in

Paper Citation Record · LEDGER

Bayesian Reward Models for LLM Alignment

As of 14 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 11 inbound Pith citation observations for arXiv:2402.13210.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2402.13210 v2

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 11 of 11 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-14T06:32:32.682623+00:00

measured 11 of 11 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-12T17:41:13.413554Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-03T00:37:30.066691Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation 906ad3da-1dc1-450e-b802-35b45e0d3319 · inbound

Why you don't overfit, and don't need Bayes if you only train for one epoch cites this paper.

Why you don't overfit, and don't need Bayes if you only train for one epoch Bayesian Reward Models for LLM Alignment

Reference 2009

Resolution
unresolved
no resolver link, observed 2026-08-12T17:41:13.413554Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T17:41:13.413554Z digest=sha256:b711e0ad9f185ed2ae853988b5f8ff314ad87fca321da8dcdf33ddda731f8030

Observation 96f0f89b-53c5-4a16-919a-2d875d3b71d2 · inbound

Solving the Inverse Alignment Problem for Efficient RLHF cites this paper.

Solving the Inverse Alignment Problem for Efficient RLHF Bayesian Reward Models for LLM Alignment

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-11T15:55:14.495738Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T15:55:14.495738Z digest=sha256:7b58309350eb6522a7e43dbd7c083be99f1d93da18918501b3764d3b07681531

Observation 06152cda-8b23-4307-9751-c981a0d929f3 · inbound

When Can Proxies Improve the Sample Complexity of Preference Learning? cites this paper.

When Can Proxies Improve the Sample Complexity of Preference Learning? Bayesian Reward Models for LLM Alignment

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-11T10:41:24.092589Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T10:41:24.092589Z digest=sha256:4c52cca1f6e965819a71ff2e27e4e4b268a5928f2f925e7c206d79ce3c383023

Observation 161a171d-292a-4401-bc16-1ec505d985f0 · inbound

Many of Your DPOs are Secretly One: Attempting Unification Through Mutual Information cites this paper.

Many of Your DPOs are Secretly One: Attempting Unification Through Mutual Information Bayesian Reward Models for LLM Alignment

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-10T22:32:56.546850Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T22:32:56.546850Z digest=sha256:2d91e2a8776c1604d0f05840bfdfc802ab7c98d072ca1286e40cec270b7de277

Observation 66514782-85be-4daf-b14d-38c05902923e · inbound

Optimising Language Models for Downstream Tasks: A Post-Training Perspective cites this paper.

Optimising Language Models for Downstream Tasks: A Post-Training Perspective Bayesian Reward Models for LLM Alignment

Reference 260

Resolution
unresolved
no resolver link, observed 2026-08-06T22:44:44.755499Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:44:44.755499Z digest=sha256:b4cfd9507bd4448274dccd5b4e33bc0c4127a35c6ffa384e8e5f4ea871d2ea97

Observation e343fe7e-4f12-44f1-84b3-6cf25f16acf5 · inbound

Aligned Query Expansion: Efficient Query Expansion for Information Retrieval through LLM Alignment cites this paper.

Aligned Query Expansion: Efficient Query Expansion for Information Retrieval through LLM Alignment Bayesian Reward Models for LLM Alignment

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-06T17:24:39.524257Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T17:24:39.524257Z digest=sha256:7a30765370ddd8a648efec081747feaae566322e8c999bde48f349946e53ffa8

Observation 7cf9db51-0e08-48f4-90c8-a2c8f098b12a · inbound

Mitigating Think-Answer Mismatch in LLM Reasoning Through Noise-Aware Advantage Reweighting cites this paper.

Mitigating Think-Answer Mismatch in LLM Reasoning Through Noise-Aware Advantage Reweighting Bayesian Reward Models for LLM Alignment

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-05T23:10:24.529784Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T23:10:24.529784Z digest=sha256:ce76f8548a89ce51dbaa32ad3b2ea7a7a9f673deb886cad45e5c5020f2846211

Observation 4440dea5-165d-4523-8693-b37bfd7be1bc · inbound

The Alignment Auditor: A Bayesian Framework for Verifying and Refining LLM Objectives cites this paper.

The Alignment Auditor: A Bayesian Framework for Verifying and Refining LLM Objectives Bayesian Reward Models for LLM Alignment

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-04T11:17:36.758976Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T11:17:36.758976Z digest=sha256:91044fb34980b7f9b3285c8997fa1396c15212a61f9155f785562f5a12bf8d6c

Observation 80a23c92-a12d-4d19-8a64-564e89829535 · inbound

The Hidden Bias of Process Reward Models:PRISM for Rewarding the Right Reasoning cites this paper.

The Hidden Bias of Process Reward Models:PRISM for Rewarding the Right Reasoning Bayesian Reward Models for LLM Alignment

Reference 20

Resolution
verified exact
arxiv_id, observed 2026-07-03T00:37:30.068171Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-06-27T17:07:06.227106Z digest=sha256:415ea99a732b84c7d684cc77bb74659c5f5e7874573399fb03d42519803d85e8

Observation fab0b4f4-7813-40ed-972f-2290c1d1aab7 · inbound

BaRA: Bayesian Adaptive Rank Allocation for Parameter-Efficient Fine-Tuning cites this paper.

BaRA: Bayesian Adaptive Rank Allocation for Parameter-Efficient Fine-Tuning Bayesian Reward Models for LLM Alignment

Reference 12

Resolution
metadata mismatch
arxiv_id, observed 2026-06-30T08:04:28.489458Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-06-30T07:55:13.502149Z digest=sha256:dda954a464a37976deb2cd0df5210a5354c591542c871017b4429e5682d46a5d

Observation 388f28e2-382c-41b3-b71c-84ee765bf8a4 · inbound

ConceptSMILE: Auditing the Trustworthiness of Concept-Based Explainable AI cites this paper.

ConceptSMILE: Auditing the Trustworthiness of Concept-Based Explainable AI Bayesian Reward Models for LLM Alignment

Reference 19

Resolution
unresolved
no resolver link, observed 2026-07-13T01:29:02.336552Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-13T01:29:02.336552Z digest=sha256:e60ca4783a9f5348b75bef9f2c018a2f6775626f5f2d93c20ee94ce78555febf