Pith. sign in

Paper Citation Record · LEDGER

Rethinking Bradley-Terry Models in Preference-Based Reward Modeling: Foundations, Theory, and Alternatives

As of 11 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 13 inbound Pith citation observations for arXiv:2411.04991.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2411.04991 v2

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 13 of 13 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-10T06:31:04.303077+00:00

measured 13 of 13 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-10T22:51:52.736901Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-06-28T19:32:35.329195Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation 23fbc6e4-de51-40d2-819c-4fc4d70ace23 · inbound

An Overview and Discussion on Using Large Language Models for Implementation Generation of Solutions to Open-Ended Problems cites this paper.

An Overview and Discussion on Using Large Language Models for Implementation Generation of Solutions to Open-Ended Problems Rethinking Bradley-Terry Models in Preference-Based Reward Modeling: Foundations, Theory, and Alternatives

Reference 147

Resolution
unresolved
no resolver link, observed 2026-08-10T22:51:52.736901Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T22:51:52.736901Z digest=sha256:2dbaf64dd79c495f67f6e2759a9685e9c6388f7238de4eb66294775d9f79cf4c

Observation 15458f65-86ad-44d6-81ba-f999ed6230ce · inbound

Leveraging Sparsity for Sample-Efficient Preference Learning: A Theoretical Perspective cites this paper.

Leveraging Sparsity for Sample-Efficient Preference Learning: A Theoretical Perspective Rethinking Bradley-Terry Models in Preference-Based Reward Modeling: Foundations, Theory, and Alternatives

Reference 84

Resolution
unresolved
no resolver link, observed 2026-08-10T00:11:59.381133Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T00:11:59.381133Z digest=sha256:0a28d2c3e9f2b6015fdeb568b64f5824143f7961650db7a3d356562738cec034

Observation e4c1d78b-fcf4-4cdd-a783-98b08dbba329 · inbound

Reviving The Classics: Active Reward Modeling in Large Language Model Alignment cites this paper.

Reviving The Classics: Active Reward Modeling in Large Language Model Alignment Rethinking Bradley-Terry Models in Preference-Based Reward Modeling: Foundations, Theory, and Alternatives

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-09T11:47:17.611380Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T11:47:17.611380Z digest=sha256:158abbbb9aecff12b220a67314302c9921cfc7951150961dd60e6a941dd720d3

Observation 0376cf34-9ce7-4b62-a313-8f94310bd832 · inbound

Reusing Embeddings: Reproducible Reward Model Research in Large Language Model Alignment without GPUs cites this paper.

Reusing Embeddings: Reproducible Reward Model Research in Large Language Model Alignment without GPUs Rethinking Bradley-Terry Models in Preference-Based Reward Modeling: Foundations, Theory, and Alternatives

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-09T11:32:47.943542Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T11:32:47.943542Z digest=sha256:c9c43c439faf6d6a9f9becb53569971bfacd592383d538855dbc9c225d2dcb38

Observation 49292bd6-b8fe-4ec8-9aa2-a8a55676064d · inbound

How Humans Help LLMs: Assessing and Incentivizing Human Preference Annotators cites this paper.

How Humans Help LLMs: Assessing and Incentivizing Human Preference Annotators Rethinking Bradley-Terry Models in Preference-Based Reward Modeling: Foundations, Theory, and Alternatives

Reference 83

Resolution
verified exact
arxiv_id, observed 2026-05-23T03:57:29.667863Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-05-23T03:56:18.703995Z digest=sha256:6f9f080e2099e30aadf14c95262f1dbbae7f597ab0a1bd0837b00380b894686a

Observation cf465eb7-e721-48fe-b32b-d83ee0b6d455 · inbound

Game Theory Meets Large Language Models: A Systematic Survey with Taxonomy and New Frontiers cites this paper.

Game Theory Meets Large Language Models: A Systematic Survey with Taxonomy and New Frontiers Rethinking Bradley-Terry Models in Preference-Based Reward Modeling: Foundations, Theory, and Alternatives

Reference 144

Resolution
unresolved
no resolver link, observed 2026-08-07T22:50:33.178988Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T22:50:33.178988Z digest=sha256:6ab5872195d1329311c1b5f43b44aa3780be00e45036ded001faeb8e57641031

Observation 4dd5aed5-b748-4f5e-8d73-f93a350c6de0 · inbound

OpenReview Should be Protected and Leveraged as a Community Asset for Research in the Era of Large Language Models cites this paper.

OpenReview Should be Protected and Leveraged as a Community Asset for Research in the Era of Large Language Models Rethinking Bradley-Terry Models in Preference-Based Reward Modeling: Foundations, Theory, and Alternatives

Reference 137

Resolution
unresolved
no resolver link, observed 2026-08-07T14:31:46.865446Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T14:31:46.865446Z digest=sha256:f2e8b233f0917102230634fa19729d7c99159225397fb958422bdab74eb1c4c6

Observation a13a5ca8-7485-485b-9904-bfad94f57cc4 · inbound

AMoPO: Adaptive Multi-objective Preference Optimization without Reward Models and Reference Models cites this paper.

AMoPO: Adaptive Multi-objective Preference Optimization without Reward Models and Reference Models Rethinking Bradley-Terry Models in Preference-Based Reward Modeling: Foundations, Theory, and Alternatives

Reference 1027

Resolution
unresolved
no resolver link, observed 2026-08-07T05:48:04.364900Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:48:04.364900Z digest=sha256:2872f84efb78436b5899429183263c606bc2e605e7e6afbaa454d1fe7d4ce665

Observation 1b8a5cf1-704d-4ee7-9611-ac9b10123401 · inbound

Inverse Reinforcement Learning Meets Large Language Model Post-Training: Basics, Advances, and Opportunities cites this paper.

Inverse Reinforcement Learning Meets Large Language Model Post-Training: Basics, Advances, and Opportunities Rethinking Bradley-Terry Models in Preference-Based Reward Modeling: Foundations, Theory, and Alternatives

Reference 90

Resolution
unresolved
no resolver link, observed 2026-08-06T16:34:25.213454Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T16:34:25.213454Z digest=sha256:02d1b7589986652834413edc17bd13f6ccfe75a67da84626effd73e8efbadee6

Observation faaa401d-b7f0-4ec2-b1aa-a54f8cd4e031 · inbound

Multimodal LLMs as Customized Reward Models for Text-to-Image Generation cites this paper.

Multimodal LLMs as Customized Reward Models for Text-to-Image Generation Rethinking Bradley-Terry Models in Preference-Based Reward Modeling: Foundations, Theory, and Alternatives

Reference 55

Resolution
unresolved
no resolver link, observed 2026-08-06T12:54:37.374599Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T12:54:37.374599Z digest=sha256:ab96a1cc6097574522982c00925a6e355a7200699c46515af3c5db3b6420b27c

Observation 4e368313-1aaa-4300-b279-2afba233b418 · inbound

Curated Synthetic Data Doesn't Have to Collapse: A Theoretical Study of Generative Retraining with Pluralistic Preferences cites this paper.

Curated Synthetic Data Doesn't Have to Collapse: A Theoretical Study of Generative Retraining with Pluralistic Preferences Rethinking Bradley-Terry Models in Preference-Based Reward Modeling: Foundations, Theory, and Alternatives

Reference 81

Resolution
verified exact
arxiv_id, observed 2026-05-11T02:30:54.518269Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-05-11T02:30:14.693348Z digest=sha256:dabc117b7122e33379598db6c0930881cf779ae717ba37021c5036b4709f58b9

Observation 4cf75416-2bca-4a70-b028-30d595d534e8 · inbound

Large Language Models Should Learn Personalized Rather Than Aggregated Human Preferences cites this paper.

Large Language Models Should Learn Personalized Rather Than Aggregated Human Preferences Rethinking Bradley-Terry Models in Preference-Based Reward Modeling: Foundations, Theory, and Alternatives

Reference 27

Resolution
verified exact
arxiv_id, observed 2026-06-28T19:32:35.330785Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-06-28T19:03:47.751245Z digest=sha256:4782f67639fb0eed640a96ff59b05813b7652d3326583f4def0595472248201d

Observation 59dad6fc-750f-4375-a283-f13f9e6a4c58 · inbound

From Feasibility to Desirability: Plan, Learn, Adapt (PLA) Framework for Personalized On-Device Itinerary Generation cites this paper.

From Feasibility to Desirability: Plan, Learn, Adapt (PLA) Framework for Personalized On-Device Itinerary Generation Rethinking Bradley-Terry Models in Preference-Based Reward Modeling: Foundations, Theory, and Alternatives

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-01T23:02:53.068235Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T23:02:53.068235Z digest=sha256:1c5c0f3bfddce95867ccbda36f9797d7e576b523fe6e63e790c00bcf3c0f9b49