Pith. sign in

Paper Citation Record · LEDGER

Beyond the Binary: Capturing Diverse Preferences With Reward Regularization

As of 13 August 2026, this Paper Citation Record lists 57 of 57 outbound references and 1 inbound Pith citation observation for arXiv:2412.03822.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2412.03822 v1

Coverage vector

measured 57 of 57 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-11T22:06:51.356046Z

measured 58 of 58 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-13T06:32:02.005865+00:00

measured 1 of 1 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-06-28T01:32:04.660400Z

measured 1 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-08-05T02:28:24.338817Z

Reference resolution

57 of 57 outbound references displayed

  • verified exact0
  • verified fuzzy12
  • unresolved44
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch1

External citation measurements

0
arxiv_reference, observed 2026-08-05T02:28:24.338817Z

Outbound references

Observation 99e1cee2-2870-4e1b-8b6b-e677a898dcd2 · outbound

This paper cites Training language models to follow instructions with human feedback.

Beyond the Binary: Capturing Diverse Preferences With Reward Regularization Training language models to follow instructions with human feedback

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-11T22:06:51.044240Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T22:06:51.044240Z digest=sha256:a2d7d675710ad82a0dee8929e4c4bce09229a19bcc5404cb05fe0403ebc27c3c

Observation cfa3e659-91c5-40bb-8b5b-5b9a6682aea9 · outbound

This paper cites Training a Helpful and Harmless Assistant with Reinforcement Learning from Human Feedback.

Beyond the Binary: Capturing Diverse Preferences With Reward Regularization Training a Helpful and Harmless Assistant with Reinforcement Learning from Human Feedback

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-11T22:06:51.050083Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T22:06:51.050083Z digest=sha256:bb02a2c4e5d27033d37a81f855b654150a80034acd31ffa0bcf676e7997b6396

Observation db57cba3-db7b-466e-b44f-8a45c1b5ce2f · outbound

This paper cites Llama 2: Open Foundation and Fine-Tuned Chat Models.

Beyond the Binary: Capturing Diverse Preferences With Reward Regularization Llama 2: Open Foundation and Fine-Tuned Chat Models

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-11T22:06:51.056099Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T22:06:51.056099Z digest=sha256:8c848470646c176599d39f23f74a44e473949fe8e5a636cb1ade09877f1fe286

Observation 2d6f3939-1a8e-41c5-941e-743146953699 · outbound

This paper cites GPT-4 Technical Report.

Beyond the Binary: Capturing Diverse Preferences With Reward Regularization GPT-4 Technical Report

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-11T22:06:51.061791Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T22:06:51.061791Z digest=sha256:bd42daab6c16325338d1df449482f9fd45cda8fe56b963b8d422ce349aee07f0

Observation 22760f65-963c-4a7c-98d9-e29b8cac525e · outbound

This paper cites How Far Are We From AGI: Are LLMs All We Need?.

Beyond the Binary: Capturing Diverse Preferences With Reward Regularization How Far Are We From AGI: Are LLMs All We Need?

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-11T22:06:51.067338Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T22:06:51.067338Z digest=sha256:8c80fd0790ae6d4b1f03c8eeb684e0539ef1acc0d647425f595bff7138833a25

Observation 97a45ed3-b108-43d3-b22c-bbd68b0d9935 · outbound

This paper cites Learning to summarize with human feedback.

Beyond the Binary: Capturing Diverse Preferences With Reward Regularization Learning to summarize with human feedback

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-11T22:06:51.073149Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T22:06:51.073149Z digest=sha256:32a20079eb1d57bd3c153888f8fd18c61f25204358a7d8aa10f71838b8f24890

Observation b2752897-a48e-4fcd-8676-bd76aaf0634e · outbound

This paper cites Direct preference optimization: Your language model is secretly a reward model.

Beyond the Binary: Capturing Diverse Preferences With Reward Regularization Direct preference optimization: Your language model is secretly a reward model

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-11T22:06:51.079176Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T22:06:51.079176Z digest=sha256:4fe5fa9ce6e2fde60d271415cf6fdd3430ee52332d3f4eeacab19b5bf8602cee

Observation 802c7432-bfa4-4aa4-bf4b-795b6f5085a5 · outbound

This paper cites SimPO: Simple Preference Optimization with a Reference-Free Reward.

Beyond the Binary: Capturing Diverse Preferences With Reward Regularization SimPO: Simple Preference Optimization with a Reference-Free Reward

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-11T22:06:51.083942Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T22:06:51.083942Z digest=sha256:a59d059dd35fe267c9c525dbdbf9294117fca367b734f300451be29eea80bc10

Observation 21b18e6a-5855-4e9d-8e84-92dcd08b8b93 · outbound

This paper cites The History and Risks of Reinforcement Learning and Human Feedback.

Beyond the Binary: Capturing Diverse Preferences With Reward Regularization The History and Risks of Reinforcement Learning and Human Feedback

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-11T22:06:51.089208Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T22:06:51.089208Z digest=sha256:6d86c6170b4097e9fd1d6f93657bf01ee6e60308981c48d681463d8b35dcf981

Observation b79f6e8f-9af0-405d-870f-bf3606118fb6 · outbound

This paper cites The past, present and better future of feedback learning in large language models for subjective human preferences and values.

Beyond the Binary: Capturing Diverse Preferences With Reward Regularization The past, present and better future of feedback learning in large language models for subjective human preferences and values

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T22:06:52.377687Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-11T22:06:51.094552Z digest=sha256:59207362c254a5cbffddf7829ba7d417c2ca79c1932f6291a13aae6d00170d32

Observation 388ff225-b08c-4b3a-8bcf-c0d955670635 · outbound

This paper cites Artificial Intelligence, Values and Alignment.

Beyond the Binary: Capturing Diverse Preferences With Reward Regularization Artificial Intelligence, Values and Alignment

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-11T22:06:51.099490Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T22:06:51.099490Z digest=sha256:f33f2d7f2c77c57dae3aa8aa0eeadd7121ed6ccbc0cd82c015bee9b55e5d3266

Observation 028d699f-485d-4526-867a-7036adcfa59f · outbound

This paper cites Artificial Intelligence, Humanistic Ethics.

Beyond the Binary: Capturing Diverse Preferences With Reward Regularization Artificial Intelligence, Humanistic Ethics

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T22:06:52.359427Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-11T22:06:51.105715Z digest=sha256:b80c7547c1104ee5383eb8c7fcbff99f698e5d4bf4f3c511736b8e2953dac279

Observation 721b89fe-31a3-4c83-a215-fb417b6944ce · outbound

This paper cites Beyond Preferences in AI Alignment.

Beyond the Binary: Capturing Diverse Preferences With Reward Regularization Beyond Preferences in AI Alignment

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-11T22:06:51.116195Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T22:06:51.116195Z digest=sha256:fe1cda4e6ea3f8ab3860175308120f960afbc633b6aba7783f79eb6d1898fc31

Observation 80e5b85b-db90-48d5-af76-5386352a2606 · outbound

This paper cites Meta Community Forum: Re- sults Analysis.

Beyond the Binary: Capturing Diverse Preferences With Reward Regularization Meta Community Forum: Re- sults Analysis

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T22:06:52.341743Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-11T22:06:51.121237Z digest=sha256:ccb62fb9c262aa29243f8e04235107230b30d5404b31312b68a3e094c685c9cb

Observation e15e84d2-04ba-4681-b23e-f62263bc02dc · outbound

This paper cites STELA: a community-centred approach to norm elicitation for AI alignment.

Beyond the Binary: Capturing Diverse Preferences With Reward Regularization STELA: a community-centred approach to norm elicitation for AI alignment

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-11T22:06:51.125962Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T22:06:51.125962Z digest=sha256:a5df9e57600d42be1abe0e2f2a25165e8fd2abcae3e2dd3993f64c26986979b6

Observation 59936292-d71c-4b0e-b898-45c3c8c5b2e2 · outbound

This paper cites Constitutional AI: Harmlessness from AI Feedback.

Beyond the Binary: Capturing Diverse Preferences With Reward Regularization Constitutional AI: Harmlessness from AI Feedback

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-11T22:06:51.130971Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T22:06:51.130971Z digest=sha256:adfbe6041b8dbecd2a608f15ec31876477dc15b5e5d329c372a28f875586db04

Observation db993abc-2c68-4cd6-8ce3-5b35f2fc6b72 · outbound

This paper cites Improving alignment of dialogue agents via targeted human judgements.

Beyond the Binary: Capturing Diverse Preferences With Reward Regularization Improving alignment of dialogue agents via targeted human judgements

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-11T22:06:51.136182Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T22:06:51.136182Z digest=sha256:f19b1a87c99a3218b88f2115aff3ab5c18435a95f286389924bb61ff861e07ba

Observation 99f768e8-7b04-482e-aa2b-8ad5b81e2964 · outbound

This paper cites Collective constitutional ai: Aligning a language model with public input.

Beyond the Binary: Capturing Diverse Preferences With Reward Regularization Collective constitutional ai: Aligning a language model with public input

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-11T22:06:51.141942Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T22:06:51.141942Z digest=sha256:d1d80f427acdaf406647378465c623974986be5df323ef7c935341a3464f37d4

Observation 433669fe-1be5-4bf8-8e0f-5eb8d9ca4342 · outbound

This paper cites Introducing Meta Llama 3: The most capable openly available LLM to date, April.

Beyond the Binary: Capturing Diverse Preferences With Reward Regularization Introducing Meta Llama 3: The most capable openly available LLM to date, April

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T22:06:52.311965Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-11T22:06:51.147147Z digest=sha256:fdf7f0e59e660ef433110c45b0d2e11d410215b7fda0643704112936132bc732

Observation 14a7b8f0-b84a-4342-aa2d-8adb9e2c6fb8 · outbound

This paper cites MMToM-QA: Multimodal Theory of Mind Question Answering.

Beyond the Binary: Capturing Diverse Preferences With Reward Regularization MMToM-QA: Multimodal Theory of Mind Question Answering

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-11T22:06:51.157581Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T22:06:51.157581Z digest=sha256:82da499691ea680a2f8fb9f2d284f1ea3d31236ef75d0977e3b86887ef2176e1

Observation fdf7d294-b607-472a-bee8-ccc13d0ce288 · outbound

This paper cites ChatGPT’s weekly users have doubled in less than a year.

Beyond the Binary: Capturing Diverse Preferences With Reward Regularization ChatGPT’s weekly users have doubled in less than a year

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T22:06:52.277815Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-11T22:06:51.162693Z digest=sha256:d5971b16f0611495ec8e6a54ceb8ebc9636e860697f0f126f179313ce8e3c87a

Observation 3bf9c8eb-cb4a-4d79-9527-d85eb53a5e0e · outbound

This paper cites LaMDA: Language Models for Dialog Applications.

Beyond the Binary: Capturing Diverse Preferences With Reward Regularization LaMDA: Language Models for Dialog Applications

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-11T22:06:51.167430Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T22:06:51.167430Z digest=sha256:dc8d1e8fe8dc3d98f794926b4cacd61616d8be1c9e1bf5560fc1e8d34fffa7ab

Observation afd6c360-f8af-41b0-aee8-fa567431ebfd · outbound

This paper cites DICES Dataset: Diversity in Conversational AI Evaluation for Safety.

Beyond the Binary: Capturing Diverse Preferences With Reward Regularization DICES Dataset: Diversity in Conversational AI Evaluation for Safety

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-11T22:06:51.172816Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T22:06:51.172816Z digest=sha256:45d41ddd013f1f0ba92a692dda964de22b79c66260971037fe69c6a9a1876523

Observation 77e596d9-6cbc-4ec4-8aa7-bcae2448e9e5 · outbound

This paper cites Pretraining language models with human preferences.

Beyond the Binary: Capturing Diverse Preferences With Reward Regularization Pretraining language models with human preferences

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-11T22:06:51.177999Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T22:06:51.177999Z digest=sha256:18e9f6c582a417bb4507d68ea21c266d28de2d4eee82fc56f8c9cd93189711f3

Observation a78eb4c6-df25-4f00-80ef-efc87adc010b · outbound

This paper cites alignment.

Beyond the Binary: Capturing Diverse Preferences With Reward Regularization alignment

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T22:06:52.249075Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-11T22:06:51.182842Z digest=sha256:ad91b8b6d3869308c04a56093affbc0d94f7995a2cd694df803ad770ac29bd4f

Observation 8686eea0-48ae-458c-86a8-cf66d1761691 · outbound

This paper cites Whose Language Counts as High Quality? Measuring Language Ideologies in Text Data Selection.

Beyond the Binary: Capturing Diverse Preferences With Reward Regularization Whose Language Counts as High Quality? Measuring Language Ideologies in Text Data Selection

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-11T22:06:51.187916Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T22:06:51.187916Z digest=sha256:1c435b844e1d6a7e5512acbfe381c2b2b92a3c7b7ee75ebc18c7f1ea13f0dc38

Observation 8e117d74-1c97-478c-af8f-2e7605ad7c3f · outbound

This paper cites The PRISM Alignment Dataset: What Participatory, Representative and Individualised Human Feedback Reveals About the Subjective and Multicultural Alignment of Large Language Models.

Beyond the Binary: Capturing Diverse Preferences With Reward Regularization The PRISM Alignment Dataset: What Participatory, Representative and Individualised Human Feedback Reveals About the Subjective and Multicultural Alignment of Large Language Models

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-11T22:06:51.192853Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T22:06:51.192853Z digest=sha256:96855f14fb1f513dab87a5ce1f2e7b78601ee5574853f620b19900646fb3c52f

Observation a28e7541-daf6-4fef-8bc8-079c0e99f380 · outbound

This paper cites The Cultural Psychology of Large Language Models: Is ChatGPT a Holistic or Analytic Thinker?.

Beyond the Binary: Capturing Diverse Preferences With Reward Regularization The Cultural Psychology of Large Language Models: Is ChatGPT a Holistic or Analytic Thinker?

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-11T22:06:51.198392Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T22:06:51.198392Z digest=sha256:f6a0820476cc72c0c70f28c378d700c5f8a1964b6e1bf2950c66d488281d9764

Observation da0a89d7-7cf0-4605-8415-302fdd4ef571 · outbound

This paper cites Diverging preferences: When do annotators disagree and do models know? arXiv preprint arXiv:2410.14632, 2024.

Beyond the Binary: Capturing Diverse Preferences With Reward Regularization Diverging preferences: When do annotators disagree and do models know? arXiv preprint arXiv:2410.14632, 2024

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-11T22:06:51.203705Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T22:06:51.203705Z digest=sha256:a6f9472c6ce9114033ed19461140917827bbe9442726bfb7ef71eff6c6289bd4

Observation 381a39c0-e688-4706-86aa-ce211ba081a8 · outbound

This paper cites Personalized Soups: Personalized Large Language Model Alignment via Post-hoc Parameter Merging.

Beyond the Binary: Capturing Diverse Preferences With Reward Regularization Personalized Soups: Personalized Large Language Model Alignment via Post-hoc Parameter Merging

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-11T22:06:51.209034Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T22:06:51.209034Z digest=sha256:518841a02df43a1cb354afbe32f2ea67d879d264415edb45969e017a93753bf6

Observation f071d434-42d9-4bc0-85d9-6610397831de · outbound

This paper cites Personalized Language Modeling from Personalized Human Feedback.

Beyond the Binary: Capturing Diverse Preferences With Reward Regularization Personalized Language Modeling from Personalized Human Feedback

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-11T22:06:51.214914Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T22:06:51.214914Z digest=sha256:7e9c85194c84c3af0e658451d6941effea5ec7cb35fb669b979a7cb3a3c1dc25

Observation 1fed5acb-3a7c-47bd-853b-55d90103f283 · outbound

This paper cites Personalizing Reinforcement Learning from Human Feedback with Variational Preference Learning.

Beyond the Binary: Capturing Diverse Preferences With Reward Regularization Personalizing Reinforcement Learning from Human Feedback with Variational Preference Learning

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-11T22:06:51.220449Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T22:06:51.220449Z digest=sha256:55b7b04f956b11b48682041c923b55bfe866ae54a523a202d86935127e3c08fa

Observation 3b78df44-c63b-4647-8235-1d4c0525780a · outbound

This paper cites Rewarded soups: towards pareto-optimal alignment by interpolating weights fine-tuned on diverse rewards.Advances in Neural Information Processing Systems, 36, 2024.

Beyond the Binary: Capturing Diverse Preferences With Reward Regularization Rewarded soups: towards pareto-optimal alignment by interpolating weights fine-tuned on diverse rewards.Advances in Neural Information Processing Systems, 36, 2024

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T22:06:52.232281Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-11T22:06:51.225923Z digest=sha256:68ca070032720d728aa6e789544d1635439452e28528726516a0736f36a66c31

Observation ecf77a7b-1b82-40c2-9009-138632ef1698 · outbound

This paper cites Fine-tuning language models to find agreement among humans with diverse preferences.

Beyond the Binary: Capturing Diverse Preferences With Reward Regularization Fine-tuning language models to find agreement among humans with diverse preferences

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-11T22:06:51.230994Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T22:06:51.230994Z digest=sha256:cf39d61d8631ad854ac37c176015b85c4f96166781be9888f2b3031a222b4be2

Observation 2d01beb3-7a78-4607-acee-3dd57276c533 · outbound

This paper cites MaxMin-RLHF: Alignment with Diverse Human Preferences.

Beyond the Binary: Capturing Diverse Preferences With Reward Regularization MaxMin-RLHF: Alignment with Diverse Human Preferences

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-11T22:06:51.235932Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T22:06:51.235932Z digest=sha256:6aa5d56ae5da9faf4f907298c4f12a2757578cb25e56e56ea886ab1f65102e00

Observation 340b4c90-159c-45f6-87e5-8086a3184a9c · outbound

This paper cites Social Choice Should Guide AI Alignment in Dealing with Diverse Human Feedback.

Beyond the Binary: Capturing Diverse Preferences With Reward Regularization Social Choice Should Guide AI Alignment in Dealing with Diverse Human Feedback

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-11T22:06:51.241253Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T22:06:51.241253Z digest=sha256:b481df83751b1ac52afe3b167f6ea4261fd25551639b0b3ae5537e90307277a5

Observation 23783648-074d-49db-887b-bd1fdb950640 · outbound

This paper cites Distributional Preference Learning: Understanding and Accounting for Hidden Context in RLHF.

Beyond the Binary: Capturing Diverse Preferences With Reward Regularization Distributional Preference Learning: Understanding and Accounting for Hidden Context in RLHF

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-11T22:06:51.247150Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T22:06:51.247150Z digest=sha256:c64e6b8d1ff3d4dcd6dae125a7a37a7a6e25a6eaded21feb77c7cc4204936987

Observation 08fcd4e1-cc17-4c7d-a100-2797d16e8a90 · outbound

This paper cites Aligning Crowd Feedback via Distributional Preference Reward Modeling.

Beyond the Binary: Capturing Diverse Preferences With Reward Regularization Aligning Crowd Feedback via Distributional Preference Reward Modeling

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-11T22:06:51.252782Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T22:06:51.252782Z digest=sha256:2f532a370f5ae45df6877dc31398a69d8de27e5be6a329c5b890fdc941c72935

Observation 76c581c7-7fd2-4bd6-9bfb-61de741da16d · outbound

This paper cites Uncertainty-aware Reward Model: Teaching Reward Models to Know What is Unknown.

Beyond the Binary: Capturing Diverse Preferences With Reward Regularization Uncertainty-aware Reward Model: Teaching Reward Models to Know What is Unknown

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-11T22:06:51.258372Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T22:06:51.258372Z digest=sha256:92fa808266f46f27a2f5027641a9bc90a00353d8460ccc190bdd19bc6dbde21c

Observation dc4ed378-108a-4b3d-b1da-7c481dec6011 · outbound

This paper cites WebGPT: Browser-assisted question-answering with human feedback.

Beyond the Binary: Capturing Diverse Preferences With Reward Regularization WebGPT: Browser-assisted question-answering with human feedback

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-11T22:06:51.264538Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T22:06:51.264538Z digest=sha256:b89acc2a61304560625f36cc556289d8ccaba3cd772a5c9aa20539bc944ca691

Observation 27d43019-d72c-437d-bd35-d10a9dc4b073 · outbound

This paper cites synthetic-instruct-gptj-pairwise (revision cc92d8d),.

Beyond the Binary: Capturing Diverse Preferences With Reward Regularization synthetic-instruct-gptj-pairwise (revision cc92d8d),

Reference 41

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T22:06:52.203555Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-11T22:06:51.271386Z digest=sha256:0d3282a7acbef5489978f58f2718008077434a848d29a406a78780724e5ec2fc

Observation 2ade0842-68ab-4012-b155-8a9a76cbb5d4 · outbound

This paper cites Ultrafeedback: Boosting language models with high-quality feedback, 2023.

Beyond the Binary: Capturing Diverse Preferences With Reward Regularization Ultrafeedback: Boosting language models with high-quality feedback, 2023

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-11T22:06:51.285345Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T22:06:51.285345Z digest=sha256:e2adede38384c9dc1fa7194b0346d2dfd807c92b2d7a6e76f3f8a5f652716be5

Observation 06542e24-3659-4559-ab7f-0bd9add0c4b6 · outbound

This paper cites The Llama 3 Herd of Models.

Beyond the Binary: Capturing Diverse Preferences With Reward Regularization The Llama 3 Herd of Models

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-11T22:06:51.291078Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T22:06:51.291078Z digest=sha256:bea55be52c09df9b2d5224d380f198c57c0831facf459be2074203f32910107a

Observation ac75e5fa-82e2-4ce3-ae80-ffa97cef70bc · outbound

This paper cites Transformers: State- of-the-art natural language processing.

Beyond the Binary: Capturing Diverse Preferences With Reward Regularization Transformers: State- of-the-art natural language processing

Reference 44

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T22:06:52.156056Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-11T22:06:51.296646Z digest=sha256:03bd6d3a4997267f97d970419ab46fd18e3427a73ef8b9c626262c3416aa04f1

Observation 396228b2-8607-415e-a327-5bec2dd767af · outbound

This paper cites Holistic Evaluation of Language Models.

Beyond the Binary: Capturing Diverse Preferences With Reward Regularization Holistic Evaluation of Language Models

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-11T22:06:51.303089Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T22:06:51.303089Z digest=sha256:6c6dfd10be285891a992b0df4532f0fd91752317765ad28323b5f2c2775a8ba8

Observation d7bab6bc-55b5-4a19-87c5-021f937b027b · outbound

This paper cites Argyle, Ethan C.

Beyond the Binary: Capturing Diverse Preferences With Reward Regularization Argyle, Ethan C

Reference 46

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T22:06:52.137895Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-11T22:06:51.308510Z digest=sha256:c9039b178a94d7a642f062c252273496fa4869dad198ac935abd7ae6922bba80

Observation 06bdd5c2-9118-4f2d-9ee2-a7dbd851d771 · outbound

This paper cites Are large language models good annotators? In Proceedings on, pages 38–48.

Beyond the Binary: Capturing Diverse Preferences With Reward Regularization Are large language models good annotators? In Proceedings on, pages 38–48

Reference 47

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T22:06:52.120199Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-11T22:06:51.322035Z digest=sha256:562f3df3462e161409e697919b1644950a53a2a83c6c49e5b0e8045556e05858

Observation ff1e3223-68a9-40a2-bd97-2f351a502601 · outbound

This paper cites Automated social science: Language models as scientist and subjects.

Beyond the Binary: Capturing Diverse Preferences With Reward Regularization Automated social science: Language models as scientist and subjects

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-11T22:06:51.327536Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T22:06:51.327536Z digest=sha256:af56ebaacd60733146416b1f9f81cc30c77ff15078eae5486a106afc041b16ba

Observation 7bf040e1-683c-4a15-8c3e-9923cb3bf3b1 · outbound

This paper cites Large language models that replace human participants can harmfully misportray and flatten identity groups.

Beyond the Binary: Capturing Diverse Preferences With Reward Regularization Large language models that replace human participants can harmfully misportray and flatten identity groups

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-11T22:06:51.333381Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T22:06:51.333381Z digest=sha256:21833760b3d536a18367a61fb81518ead7923348f779a157adfaeca657ce0910

Observation a1068a39-73d5-4815-8761-172725cef2d9 · outbound

This paper cites Out of One, Many: Using Language Models to Simulate Human Samples.

Beyond the Binary: Capturing Diverse Preferences With Reward Regularization Out of One, Many: Using Language Models to Simulate Human Samples

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-11T22:06:51.315201Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T22:06:51.315201Z digest=sha256:40c195f5f8796386dfeb0d638e4f5fc2a72367e929ab3e767441f149468dd3be

Observation 6da6ab26-ac28-4126-945e-024afc033184 · outbound

This paper cites Stevie Bergman, Jennifer Chien, Mark Díaz, Seliem El-Sayed, Jaylen Pittman, Shakir Mohamed, and Kevin R.

Beyond the Binary: Capturing Diverse Preferences With Reward Regularization Stevie Bergman, Jennifer Chien, Mark Díaz, Seliem El-Sayed, Jaylen Pittman, Shakir Mohamed, and Kevin R

Reference 51

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T22:06:52.077496Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-11T22:06:51.344513Z digest=sha256:e65faf0e8faa3391e78ea140052d57bc8da84b444bee6af1514ef07da2f2fbbc

Observation 39f114fe-2bd6-4ba9-9612-161e00d9ce6a · outbound

This paper cites A Roadmap to Pluralistic Alignment.

Beyond the Binary: Capturing Diverse Preferences With Reward Regularization A Roadmap to Pluralistic Alignment

Reference 52

Resolution
unresolved
no resolver link, observed 2026-08-11T22:06:51.356046Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T22:06:51.356046Z digest=sha256:a86f768432b732b03fc3d6b6201005d9cf0cdbe9468e3fe8fab7ce8438b0d54d

Observation c3760899-b7e0-433e-b121-187ef371e613 · outbound

This paper cites Whose opinions do language models reflect? In International Conference on Machine Learning, pages 29971–30004.

Beyond the Binary: Capturing Diverse Preferences With Reward Regularization Whose opinions do language models reflect? In International Conference on Machine Learning, pages 29971–30004

Reference 54

Resolution
unresolved
no resolver link, observed 2026-08-11T22:06:51.339375Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T22:06:51.339375Z digest=sha256:6cab823bb0e993e2dd3785c11751ce8bea07c53ef08c536196ee59e41f1f1cc0

Observation 91f673de-6b3d-4e78-b56a-ffc829469bca · outbound

This paper cites The illusion of artificial inclusion.

Beyond the Binary: Capturing Diverse Preferences With Reward Regularization The illusion of artificial inclusion

Reference 56

Resolution
metadata mismatch
local_arxiv, observed 2026-08-11T22:06:51.462091Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-11T22:06:51.350182Z digest=sha256:dad489a6d3a22e1be4bf9179b297d2798819454855effda13d883a03fafa7e0c

Observation 8d6231ee-7d13-4e0b-b006-28e9bb0d1b04 · outbound

This paper cites doi: 10.1162/daed_a_01912.

Beyond the Binary: Capturing Diverse Preferences With Reward Regularization doi: 10.1162/daed_a_01912

Reference 2022

Resolution
unresolved
no resolver link, observed 2026-08-11T22:06:51.110809Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T22:06:51.110809Z digest=sha256:70b56c5c510c20d95ffc4bf44da70b667b39e24e595f7cd7aa5cf793f7513fa1

Observation 6de0a1df-d639-40e3-bb12-c5a52b5baf6e · outbound

This paper cites an unresolved cited work.

Beyond the Binary: Capturing Diverse Preferences With Reward Regularization Unresolved cited work

Reference 2023

Resolution
unresolved
raw_fallback, observed 2026-08-11T22:06:52.185515Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-11T22:06:51.277039Z digest=sha256:62d6bf708068504254d1ace7d9c152b34d47c8e81fdef6a7f01cb618cc97a016

Observation 14ef19e5-51d4-4da7-abc4-c3fa5effd67b · outbound

This paper cites an unresolved cited work.

Beyond the Binary: Capturing Diverse Preferences With Reward Regularization Unresolved cited work

Reference 2024

Resolution
unresolved
raw_fallback, observed 2026-08-11T22:06:52.294454Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-11T22:06:51.152203Z digest=sha256:c47c7e25271ea64242e249c976210f180095786148d6818751424f0206468fc1

Pith citing papers

Observation 7bad1392-2f41-4fc1-94b2-27e6f266a4c2 · inbound

What Do People Actually Want From AI? Mapping Preference Plurality cites this paper.

What Do People Actually Want From AI? Mapping Preference Plurality Beyond the Binary: Capturing Diverse Preferences With Reward Regularization

Reference 67

Resolution
verified exact
arxiv_id, observed 2026-06-28T01:41:29.823462Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-06-28T01:32:04.660400Z digest=sha256:272db6aaebaadf82f94ed5e3be318e739595ab88a8b16a40d23f5b02b10a43d1