Pith. sign in

Paper Citation Record · LEDGER

Reward Modeling with Ordinal Feedback: Wisdom of the Crowd

As of 15 August 2026, this Paper Citation Record lists 50 of 50 outbound references and 0 inbound Pith citation observations for arXiv:2411.12843.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2411.12843 v1

Coverage vector

measured 50 of 50 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-12T17:15:47.969050Z

measured 50 of 50 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-15T06:32:42.880941+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

50 of 50 outbound references displayed

  • verified exact0
  • verified fuzzy2
  • unresolved48
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 4e161b2c-34e3-4133-a110-a5e585e1f317 · outbound

This paper cites , " * write output.state after.block = add.period write newline.

Reward Modeling with Ordinal Feedback: Wisdom of the Crowd , " * write output.state after.block = add.period write newline

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-12T17:15:47.761360Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T17:15:47.761360Z digest=sha256:bb1c0a8dd4bf8524e8d70bc00e0c7fe8275f165998a64ce4e4caf765754ef65b

Observation a25c7f74-9667-4f33-b6fb-4c84c54ce535 · outbound

This paper cites write newline.

Reward Modeling with Ordinal Feedback: Wisdom of the Crowd write newline

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-12T17:15:47.767020Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T17:15:47.767020Z digest=sha256:c6d55600d5468f303feffd19f5676ac4da5720242d3a474d860dd33b21e498c2

Observation 081e54fb-3acf-418c-be61-60cf8a58c732 · outbound

This paper cites an unresolved cited work.

Reward Modeling with Ordinal Feedback: Wisdom of the Crowd Unresolved cited work

Reference 3

Resolution
unresolved
raw_fallback, observed 2026-08-12T17:15:48.570039Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=arxiv_source observed=2026-08-12T17:15:47.771511Z digest=sha256:5ad5a355f9b3497defb1567253c977def3293337b37f763ea8c6cca340ca1088

Observation 3691b473-c2fe-40e0-9867-c3f02393bc96 · outbound

This paper cites Direct Preference Optimization with an Offset.

Reward Modeling with Ordinal Feedback: Wisdom of the Crowd Direct Preference Optimization with an Offset

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-12T17:15:47.775855Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T17:15:47.775855Z digest=sha256:e2c2e4c19069afef8322b978ffdfc72c988e150ae0479282489289bbae505b0e

Observation ec021049-0eab-49c7-b308-445dbea10f05 · outbound

This paper cites A General Language Assistant as a Laboratory for Alignment.

Reward Modeling with Ordinal Feedback: Wisdom of the Crowd A General Language Assistant as a Laboratory for Alignment

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-12T17:15:47.780377Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T17:15:47.780377Z digest=sha256:d234c423c382ceb950d4a7acd54c51d36bb660a6a527f9d11db201d6330da2c1

Observation bf53b892-14f2-425c-a9a3-ab687e1ea5ec · outbound

This paper cites an unresolved cited work.

Reward Modeling with Ordinal Feedback: Wisdom of the Crowd Unresolved cited work

Reference 6

Resolution
unresolved
raw_fallback, observed 2026-08-12T17:15:48.560838Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=arxiv_source observed=2026-08-12T17:15:47.784982Z digest=sha256:09d8e296cacaa772add6168b6dc2b396f80384534c2bf08bc3e1e3f70db6267e

Observation b6c3f1fc-d1d5-437e-89cb-e94123dd7724 · outbound

This paper cites an unresolved cited work.

Reward Modeling with Ordinal Feedback: Wisdom of the Crowd Unresolved cited work

Reference 7

Resolution
unresolved
raw_fallback, observed 2026-08-12T17:15:48.551619Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=arxiv_source observed=2026-08-12T17:15:47.788433Z digest=sha256:529993032de81d4f55bc7809e0e6be212483f71f9f20aff75bfea11847aaeb3a

Observation 40253885-eda4-42bb-897d-6536a1eb4854 · outbound

This paper cites Training a Helpful and Harmless Assistant with Reinforcement Learning from Human Feedback.

Reward Modeling with Ordinal Feedback: Wisdom of the Crowd Training a Helpful and Harmless Assistant with Reinforcement Learning from Human Feedback

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-12T17:15:47.793228Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T17:15:47.793228Z digest=sha256:82a2e3638fae51f55e392f3eb71748756de02a0cc78380823ff6798b9e94e2db

Observation 6d254da8-a3c7-41f1-97e3-76f4752b190d · outbound

This paper cites an unresolved cited work.

Reward Modeling with Ordinal Feedback: Wisdom of the Crowd Unresolved cited work

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-12T17:15:47.797841Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T17:15:47.797841Z digest=sha256:a4d6209cb77f40e1b5685d83390438cedc5cc464bb4a54b992fffa826f8dba5c

Observation 2ef17b6b-f8b7-4189-ab8b-30825ccca466 · outbound

This paper cites an unresolved cited work.

Reward Modeling with Ordinal Feedback: Wisdom of the Crowd Unresolved cited work

Reference 10

Resolution
unresolved
raw_fallback, observed 2026-08-12T17:15:48.535395Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=arxiv_source observed=2026-08-12T17:15:47.801861Z digest=sha256:b29bdf930aa2bb26a4e01be46ea4f6917fcd8c135c30b5b85815f03e68ebea43

Observation 3d06af9e-fbb3-4650-ac11-200bed482f14 · outbound

This paper cites MaxMin-RLHF: Alignment with Diverse Human Preferences.

Reward Modeling with Ordinal Feedback: Wisdom of the Crowd MaxMin-RLHF: Alignment with Diverse Human Preferences

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-12T17:15:47.805929Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T17:15:47.805929Z digest=sha256:e971bb8c1712878921793b0aaa93f034374214f2b5e76dbb1ad2a96a79847896

Observation 70a8b507-2a26-45ac-9373-33f3a7df35a0 · outbound

This paper cites an unresolved cited work.

Reward Modeling with Ordinal Feedback: Wisdom of the Crowd Unresolved cited work

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-12T17:15:47.810658Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T17:15:47.810658Z digest=sha256:c28b7a10f849526a4521532578f261c888ab11b36d17ec9e19ac1116db821f32

Observation d935d9db-872b-4b7f-96cc-af7637e00f58 · outbound

This paper cites u rnkranz, Eyke H \.

Reward Modeling with Ordinal Feedback: Wisdom of the Crowd u rnkranz, Eyke H \

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T17:15:48.523675Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=arxiv_source observed=2026-08-12T17:15:47.814466Z digest=sha256:652fe0a30127d38478877e424d307b1880946f6b14ca83ae05368cc745107a86

Observation 1a394c47-48e2-4352-b8af-e43e3388513e · outbound

This paper cites On the Weaknesses of Reinforcement Learning for Neural Machine Translation.

Reward Modeling with Ordinal Feedback: Wisdom of the Crowd On the Weaknesses of Reinforcement Learning for Neural Machine Translation

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-12T17:15:47.818566Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T17:15:47.818566Z digest=sha256:7b3181868f5d9c82874838c2719a1eecdfee17c7569e906fcf40de9e8197b310

Observation 152af994-d48c-4e32-975f-2259477455ae · outbound

This paper cites an unresolved cited work.

Reward Modeling with Ordinal Feedback: Wisdom of the Crowd Unresolved cited work

Reference 15

Resolution
unresolved
raw_fallback, observed 2026-08-12T17:15:48.512901Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=arxiv_source observed=2026-08-12T17:15:47.823139Z digest=sha256:b12b0f033a60a4ea38e7b632bb4c82c87dfa3093b54830d657d2c09e34bf76e4

Observation b4e678f4-cd4b-47b1-be74-b7f539c50a42 · outbound

This paper cites UltraFeedback: Boosting Language Models with Scaled AI Feedback.

Reward Modeling with Ordinal Feedback: Wisdom of the Crowd UltraFeedback: Boosting Language Models with Scaled AI Feedback

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-12T17:15:47.827057Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T17:15:47.827057Z digest=sha256:64fef5dd705dbcc3cd5c547ca7ddc39e3b2e9f7073b1c91966c6899d6be77fa8

Observation c3dc8528-77f9-4e5b-b3b4-fd185d323166 · outbound

This paper cites an unresolved cited work.

Reward Modeling with Ordinal Feedback: Wisdom of the Crowd Unresolved cited work

Reference 17

Resolution
unresolved
raw_fallback, observed 2026-08-12T17:15:48.502602Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=arxiv_source observed=2026-08-12T17:15:47.831232Z digest=sha256:2b411f58443a3a4f29e05c0ec2892c010b99ae993a1baf59a9ea050eefacb4f6

Observation b83db4eb-5179-401b-a927-dd8a7d208015 · outbound

This paper cites RAFT: Reward rAnked FineTuning for Generative Foundation Model Alignment.

Reward Modeling with Ordinal Feedback: Wisdom of the Crowd RAFT: Reward rAnked FineTuning for Generative Foundation Model Alignment

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-12T17:15:47.835813Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T17:15:47.835813Z digest=sha256:6ea2299c587416932359cbc1b969bac213a0e4e9c51e96d59ba94fa57169fcf1

Observation d77b52b3-ffcf-4168-ada9-63c5d20a9e18 · outbound

This paper cites AlignDiff: Aligning Diverse Human Preferences via Behavior-Customisable Diffusion Model.

Reward Modeling with Ordinal Feedback: Wisdom of the Crowd AlignDiff: Aligning Diverse Human Preferences via Behavior-Customisable Diffusion Model

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-12T17:15:47.840044Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T17:15:47.840044Z digest=sha256:4de5ca7f4abc86528005a6d69452151c93c1ba2cc115c2167867a962a50fafaf

Observation 3af6c8b8-c15d-43ce-8191-56798173b0c4 · outbound

This paper cites KTO: Model Alignment as Prospect Theoretic Optimization.

Reward Modeling with Ordinal Feedback: Wisdom of the Crowd KTO: Model Alignment as Prospect Theoretic Optimization

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-12T17:15:47.848222Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T17:15:47.848222Z digest=sha256:c3eb8bffcada871d3f8ce8f4d8928f3bcf01fcc38c579af0446012cdabdc0fcf

Observation b902db48-cb28-4ae0-84b9-70ae807cc43f · outbound

This paper cites an unresolved cited work.

Reward Modeling with Ordinal Feedback: Wisdom of the Crowd Unresolved cited work

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-12T17:15:47.852311Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T17:15:47.852311Z digest=sha256:37b6e12ab23d717a17785775ead277e4b800020a5e73a84240257f9014a32438

Observation 940f65f2-140c-471d-b43c-2e6c68a55b92 · outbound

This paper cites Distilling the Knowledge in a Neural Network.

Reward Modeling with Ordinal Feedback: Wisdom of the Crowd Distilling the Knowledge in a Neural Network

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-12T17:15:47.856138Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T17:15:47.856138Z digest=sha256:522eb2b392bc4ecbfb40289ebb154115a8f6372095f47efe714cfeee91d73301

Observation 9a96e8ad-c492-4abe-bda4-17336cf459b9 · outbound

This paper cites AI Alignment: A Comprehensive Survey.

Reward Modeling with Ordinal Feedback: Wisdom of the Crowd AI Alignment: A Comprehensive Survey

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-12T17:15:47.859637Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T17:15:47.859637Z digest=sha256:5442c458f8437e9afa405f6744ce24612b905676982582b0440843470f4b1926

Observation 2561aac9-bb22-4ecd-b954-2056a00fcb8d · outbound

This paper cites an unresolved cited work.

Reward Modeling with Ordinal Feedback: Wisdom of the Crowd Unresolved cited work

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-12T17:15:47.863766Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T17:15:47.863766Z digest=sha256:e917a1967ec73f14b97a5e7862658e09bd64cbdf46639f7f94e1f9dbd7728122

Observation 940af647-dbf0-44da-84ef-cbae7ac48c00 · outbound

This paper cites Smith, Hannaneh Hajishirzi.

Reward Modeling with Ordinal Feedback: Wisdom of the Crowd Smith, Hannaneh Hajishirzi

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T17:15:48.492158Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=arxiv_source observed=2026-08-12T17:15:47.867933Z digest=sha256:13f9e5e20753d7913093d7e9e30688ff55a022dad5020a65dafb87679b22dac0

Observation 5ff87c1e-2a71-461b-9b7c-9b5836edc896 · outbound

This paper cites Skywork-Reward: Bag of Tricks for Reward Modeling in LLMs.

Reward Modeling with Ordinal Feedback: Wisdom of the Crowd Skywork-Reward: Bag of Tricks for Reward Modeling in LLMs

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-12T17:15:47.872874Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T17:15:47.872874Z digest=sha256:49815b9f129ed966ed0d7ff36fbdaf889cb7897ac1038af015668e0da35d6cd7

Observation 92c73e22-97ef-45f7-855b-f8d2ae6cf105 · outbound

This paper cites Reward Learning From Preference With Ties.

Reward Modeling with Ordinal Feedback: Wisdom of the Crowd Reward Learning From Preference With Ties

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-12T17:15:47.876946Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T17:15:47.876946Z digest=sha256:5a5ea21c29656e1e6a5d26b15af67538fca92104e292b5f13c0bd139733cd9cf

Observation 0bb642c8-3734-47cb-8b39-7bd02e952bdb · outbound

This paper cites LiPO: Listwise Preference Optimization through Learning-to-Rank.

Reward Modeling with Ordinal Feedback: Wisdom of the Crowd LiPO: Listwise Preference Optimization through Learning-to-Rank

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-12T17:15:47.880969Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T17:15:47.880969Z digest=sha256:b877473df2a2aff0523b2d2dc0466a156739c101e5826232ea8c1db3f6d4720f

Observation c5b7fa9b-491e-42b9-b450-131bf43bd4af · outbound

This paper cites Statistical Rejection Sampling Improves Preference Optimization.

Reward Modeling with Ordinal Feedback: Wisdom of the Crowd Statistical Rejection Sampling Improves Preference Optimization

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-12T17:15:47.885158Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T17:15:47.885158Z digest=sha256:6d7b73583e3892faea9ec99e4d312f58513afc5afb0a3ec9ed8b3df65d9c6849

Observation 87abdf8f-3f4b-4d97-bd42-e8bc9a804c22 · outbound

This paper cites The Llama 3 Herd of Models.

Reward Modeling with Ordinal Feedback: Wisdom of the Crowd The Llama 3 Herd of Models

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-12T17:15:47.889376Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T17:15:47.889376Z digest=sha256:ff1ea3b5cbb101bc78b86880721b5cd1d56f2f2ad5d35eb723110c7b583d8d36

Observation 93e065f2-3805-4137-a6b4-2d950aae6d9a · outbound

This paper cites an unresolved cited work.

Reward Modeling with Ordinal Feedback: Wisdom of the Crowd Unresolved cited work

Reference 32

Resolution
unresolved
raw_fallback, observed 2026-08-12T17:15:48.481699Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=arxiv_source observed=2026-08-12T17:15:47.893254Z digest=sha256:6a47d8703bc3f249fc8eac646f0348be2c5224e684bffc631e727cbf8691d8a3

Observation d7f4663a-1ec0-45f7-b058-b44373011881 · outbound

This paper cites an unresolved cited work.

Reward Modeling with Ordinal Feedback: Wisdom of the Crowd Unresolved cited work

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-12T17:15:47.898069Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T17:15:47.898069Z digest=sha256:a5f159222b897d83bc9349e428b8b7b5ada1977eb6ac358b4fea1e575b79608d

Observation 6c9964f0-f602-4602-9447-f84738a543bf · outbound

This paper cites an unresolved cited work.

Reward Modeling with Ordinal Feedback: Wisdom of the Crowd Unresolved cited work

Reference 34

Resolution
unresolved
raw_fallback, observed 2026-08-12T17:15:48.464869Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=arxiv_source observed=2026-08-12T17:15:47.902599Z digest=sha256:968038dcfc4b25da89837b83de6edc927c3574d6974e5f8deba4c042694aa7c8

Observation 68a67b04-bd87-4310-8196-92aa9276446d · outbound

This paper cites an unresolved cited work.

Reward Modeling with Ordinal Feedback: Wisdom of the Crowd Unresolved cited work

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-12T17:15:47.905917Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T17:15:47.905917Z digest=sha256:e995836639d59137ae3107dd4dca16c4236e644342e05e6691a9ab29253a80c8

Observation a5988650-d0e3-4e86-b80f-e694e68a69ba · outbound

This paper cites an unresolved cited work.

Reward Modeling with Ordinal Feedback: Wisdom of the Crowd Unresolved cited work

Reference 36

Resolution
unresolved
raw_fallback, observed 2026-08-12T17:15:48.448474Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=arxiv_source observed=2026-08-12T17:15:47.909873Z digest=sha256:11b5297ade389696a23900fe4d8bd3074ab0ca8b73ac7a6290aab1a7c495ca06

Observation d3072f0c-8627-4e49-a563-cc20e1536aff · outbound

This paper cites an unresolved cited work.

Reward Modeling with Ordinal Feedback: Wisdom of the Crowd Unresolved cited work

Reference 37

Resolution
unresolved
raw_fallback, observed 2026-08-12T17:15:48.438091Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=arxiv_source observed=2026-08-12T17:15:47.913664Z digest=sha256:9b54b0b8121f6d914daa9bfcbd85683347ab1ba92f58074b88a86ad72fab1f45

Observation b0922306-e35a-4b56-a400-cf09f76f5d87 · outbound

This paper cites an unresolved cited work.

Reward Modeling with Ordinal Feedback: Wisdom of the Crowd Unresolved cited work

Reference 38

Resolution
unresolved
raw_fallback, observed 2026-08-12T17:15:48.428510Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=arxiv_source observed=2026-08-12T17:15:47.917263Z digest=sha256:6c0b11a570f6e3ddfbcd882d122f0cf8ff247780b9145efb5ea82b3f8ee87cc9

Observation 3f8bd9b5-5eb3-496b-a319-81820c7e1d00 · outbound

This paper cites Gemma 2: Improving Open Language Models at a Practical Size.

Reward Modeling with Ordinal Feedback: Wisdom of the Crowd Gemma 2: Improving Open Language Models at a Practical Size

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-12T17:15:47.921096Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T17:15:47.921096Z digest=sha256:b750ff11888d26cf7456aac67f00628d15bff575ff49ff70d48a78526c272932

Observation 250473cb-99de-4577-8adf-89c68256c17d · outbound

This paper cites Beyond Reverse KL: Generalizing Direct Preference Optimization with Diverse Divergence Constraints.

Reward Modeling with Ordinal Feedback: Wisdom of the Crowd Beyond Reverse KL: Generalizing Direct Preference Optimization with Diverse Divergence Constraints

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-12T17:15:47.925208Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T17:15:47.925208Z digest=sha256:7d955d90600f429c33df4f63d55fae4cbb004af00320e4fa6ea110e3bab43159

Observation 0c1c3830-a08e-4dee-aea8-068e398d2bfe · outbound

This paper cites Arithmetic Control of LLMs for Diverse User Preferences: Directional Preference Alignment with Multi-Objective Rewards.

Reward Modeling with Ordinal Feedback: Wisdom of the Crowd Arithmetic Control of LLMs for Diverse User Preferences: Directional Preference Alignment with Multi-Objective Rewards

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-12T17:15:47.929415Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T17:15:47.929415Z digest=sha256:d657dd2a9bfa311c703d37229c2df24b2114ef087af21e9588d79cace5220978

Observation 6586546e-2a46-48dc-8144-e1e528e25074 · outbound

This paper cites Interpretable Preferences via Multi-Objective Reward Modeling and Mixture-of-Experts.

Reward Modeling with Ordinal Feedback: Wisdom of the Crowd Interpretable Preferences via Multi-Objective Reward Modeling and Mixture-of-Experts

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-12T17:15:47.933909Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T17:15:47.933909Z digest=sha256:d3b105d36924476ad5ea4139f2e451357559c8c759d78b570164c2b3f04d8b51

Observation 5e5792f0-9ebf-48ac-bfaf-dcd252318112 · outbound

This paper cites HelpSteer2: Open-source dataset for training top-performing reward models.

Reward Modeling with Ordinal Feedback: Wisdom of the Crowd HelpSteer2: Open-source dataset for training top-performing reward models

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-12T17:15:47.938224Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T17:15:47.938224Z digest=sha256:0e98d6fc1b89c2cab95fdcc6e39dcd6bc2c8edaf3c0701a7f27f2d622a019533

Observation 2cb5f4f7-c7cd-4489-a6f2-2162d6cc3936 · outbound

This paper cites Contrastive Preference Optimization: Pushing the Boundaries of LLM Performance in Machine Translation.

Reward Modeling with Ordinal Feedback: Wisdom of the Crowd Contrastive Preference Optimization: Pushing the Boundaries of LLM Performance in Machine Translation

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-12T17:15:47.942379Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T17:15:47.942379Z digest=sha256:af2dd8d4f8cdbe569c2117b9b16f81ae82a1dafa0fad9d8a977fe8d9db33668f

Observation 74005b54-252c-498a-b0b3-0efe579b79e0 · outbound

This paper cites an unresolved cited work.

Reward Modeling with Ordinal Feedback: Wisdom of the Crowd Unresolved cited work

Reference 45

Resolution
unresolved
raw_fallback, observed 2026-08-12T17:15:48.417739Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=arxiv_source observed=2026-08-12T17:15:47.946180Z digest=sha256:f7f9771dc5390e9cb5827660654806a3757004c6efcbe0bbace427bfcbb4886c

Observation eec568d8-294d-4e6f-ab4a-d7e74b41319c · outbound

This paper cites an unresolved cited work.

Reward Modeling with Ordinal Feedback: Wisdom of the Crowd Unresolved cited work

Reference 46

Resolution
unresolved
raw_fallback, observed 2026-08-12T17:15:48.406353Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=arxiv_source observed=2026-08-12T17:15:47.950140Z digest=sha256:be7e99ec384cf0558d931cb9bc566360af1ee2964dbd31a1546e91f4c87bd5e6

Observation 14559811-d16a-449b-a098-a619feddcf6f · outbound

This paper cites an unresolved cited work.

Reward Modeling with Ordinal Feedback: Wisdom of the Crowd Unresolved cited work

Reference 47

Resolution
unresolved
raw_fallback, observed 2026-08-12T17:15:48.394128Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=arxiv_source observed=2026-08-12T17:15:47.954036Z digest=sha256:eca201b0b09460d0fe81ae49cef2136579ca06b80e050dd12e1b53112a571a5f

Observation ba846ef6-2e07-4e29-a80a-bb825c1e5a78 · outbound

This paper cites Token-level Direct Preference Optimization.

Reward Modeling with Ordinal Feedback: Wisdom of the Crowd Token-level Direct Preference Optimization

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-12T17:15:47.958038Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T17:15:47.958038Z digest=sha256:6bbe6f26d714754843795bbb9434a524766db38b67482c496b133e70dee8ff61

Observation 9a74b634-c129-40ca-b1e4-d533e99c7df5 · outbound

This paper cites SLiC-HF: Sequence Likelihood Calibration with Human Feedback.

Reward Modeling with Ordinal Feedback: Wisdom of the Crowd SLiC-HF: Sequence Likelihood Calibration with Human Feedback

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-12T17:15:47.961445Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T17:15:47.961445Z digest=sha256:0ed1e511b9aa7003a65d95ca0d93e381637e0f4b934c11e34227b699d7497c5e

Observation b4cbc0b9-fd3d-459d-8dee-dc44595f53ad · outbound

This paper cites Rethinking Soft Labels for Knowledge Distillation: A Bias-Variance Tradeoff Perspective.

Reward Modeling with Ordinal Feedback: Wisdom of the Crowd Rethinking Soft Labels for Knowledge Distillation: A Bias-Variance Tradeoff Perspective

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-12T17:15:47.965644Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T17:15:47.965644Z digest=sha256:53a08c71d501ff9f256bd91010bd307635bd3d6786e5af7f93c545af28a3927b

Observation 9af073f4-65d2-49b9-b1f3-9963811425cb · outbound

This paper cites Fine-Tuning Language Models from Human Preferences.

Reward Modeling with Ordinal Feedback: Wisdom of the Crowd Fine-Tuning Language Models from Human Preferences

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-12T17:15:47.969050Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T17:15:47.969050Z digest=sha256:37a7badf12473c60b9a26f58436b580fac4162f025a6fc8f7ee1830beecab746

Pith citing papers

No inbound Pith citation observations are available.