Pith. sign in

Paper Citation Record · LEDGER

Two Minds Better Than One: Collaborative Reward Modeling for LLM Alignment

As of 17 August 2026, this Paper Citation Record lists 80 of 80 outbound references and 1 inbound Pith citation observation for arXiv:2505.10597.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2505.10597 v2

Coverage vector

measured 80 of 80 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-15T21:19:41.527363Z

measured 81 of 81 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-17T06:30:58.91139+00:00

measured 1 of 1 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-05-10T08:32:38.883019Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-05-10T08:32:51.695159Z

Reference resolution

80 of 80 outbound references displayed

  • verified exact0
  • verified fuzzy46
  • unresolved34
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 3030dd8e-2d36-4d2c-8741-5ed4fca3c232 · outbound

This paper cites Training a Helpful and Harmless Assistant with Reinforcement Learning from Human Feedback.

Two Minds Better Than One: Collaborative Reward Modeling for LLM Alignment Training a Helpful and Harmless Assistant with Reinforcement Learning from Human Feedback

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-15T21:19:41.263687Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:19:41.263687Z digest=sha256:49d9ca356b149ab692662d63a8b90bc4fe5505695666f393fb8a09ff16bee793

Observation 7ca5174f-0ea4-406a-a676-b463e0ca9104 · outbound

This paper cites Training language models to follow instructions with human feedback.

Two Minds Better Than One: Collaborative Reward Modeling for LLM Alignment Training language models to follow instructions with human feedback

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-15T21:19:41.268318Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:19:41.268318Z digest=sha256:c82c48dadabac81edef05ebd2bd42a7cdc336d4231c44acde4698630b2491070

Observation fa577b07-d0a0-418e-97dd-183650ecf221 · outbound

This paper cites A Comprehensive Survey of Reward Models: Taxonomy, Applications, Challenges, and Future.

Two Minds Better Than One: Collaborative Reward Modeling for LLM Alignment A Comprehensive Survey of Reward Models: Taxonomy, Applications, Challenges, and Future

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-15T21:19:41.271849Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:19:41.271849Z digest=sha256:46c44c62b1d34b27fc93fde4fcc43feefe1788c04a9d5857bba0fe8d9dbb28b6

Observation 00e2721e-cb01-40d5-ae25-630bdfb76a8c · outbound

This paper cites A Unified Pairwise Framework for RLHF: Bridging Generative Reward Modeling and Policy Optimization.

Two Minds Better Than One: Collaborative Reward Modeling for LLM Alignment A Unified Pairwise Framework for RLHF: Bridging Generative Reward Modeling and Policy Optimization

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-15T21:19:41.275259Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:19:41.275259Z digest=sha256:593cb95c02e98c8b1c54f2eb7b8ece1e54dbbdbdb606c32c75fe35b4febbc440

Observation 92402796-48dc-41fc-9aea-a29af5247491 · outbound

This paper cites Better Process Supervision with Bi-directional Rewarding Signals.

Two Minds Better Than One: Collaborative Reward Modeling for LLM Alignment Better Process Supervision with Bi-directional Rewarding Signals

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-15T21:19:41.278716Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:19:41.278716Z digest=sha256:442b75ad0673986166923b8ab6292aa36c981587f1085946511eb67a2372f557

Observation 4d239447-3e69-4824-bafa-4fad15c5596c · outbound

This paper cites Reward Function Design in Reinforcement Learning, pages 25–33.

Two Minds Better Than One: Collaborative Reward Modeling for LLM Alignment Reward Function Design in Reinforcement Learning, pages 25–33

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T21:19:42.248402Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-15T21:19:41.282169Z digest=sha256:8d853fc508975db42ebd9ee45dd32bfc52dc05f8e3ddab84ac07a0e298db3561

Observation be8753fc-b9cb-4731-9515-b7a1cb6c1dab · outbound

This paper cites The Effects of Reward Misspecification: Mapping and Mitigating Misaligned Models.

Two Minds Better Than One: Collaborative Reward Modeling for LLM Alignment The Effects of Reward Misspecification: Mapping and Mitigating Misaligned Models

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-15T21:19:41.285668Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:19:41.285668Z digest=sha256:f6f07c49f5f8df48a2a829a014976b9c64eeb996b1e66ee5dd7ed4a976a83de5

Observation 3289ef25-83b8-40ca-baa2-b91abff2fa82 · outbound

This paper cites Gpt-4 technical report, 2024.

Two Minds Better Than One: Collaborative Reward Modeling for LLM Alignment Gpt-4 technical report, 2024

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-15T21:19:41.288969Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:19:41.288969Z digest=sha256:8e6af15953fae5f0953f28f1cbccdf7a21829cef0b30918b9e06f55820f98e74

Observation 21ad8af1-130a-4f9f-bb6f-4d01402cdb91 · outbound

This paper cites Secrets of RLHF in Large Language Models Part II: Reward Modeling.

Two Minds Better Than One: Collaborative Reward Modeling for LLM Alignment Secrets of RLHF in Large Language Models Part II: Reward Modeling

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-15T21:19:41.292618Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:19:41.292618Z digest=sha256:4bf3ae1f49af05555100293a95caee299e0368e40959050e408327277fee06e3

Observation 1c8d58a6-9201-4524-b469-2666a30a531d · outbound

This paper cites Skywork-reward: Bag of tricks for reward modeling in llms, October 2024.

Two Minds Better Than One: Collaborative Reward Modeling for LLM Alignment Skywork-reward: Bag of tricks for reward modeling in llms, October 2024

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T21:19:42.235122Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-15T21:19:41.296080Z digest=sha256:72d0fe8e01778ee04e6803204b0eece8c4828d556a8c10d21ea84ef5415925d2

Observation ba77e19c-df9a-46d0-81a4-eeb13a0ea201 · outbound

This paper cites RMB: Comprehensively benchmarking reward models in LLM alignment.

Two Minds Better Than One: Collaborative Reward Modeling for LLM Alignment RMB: Comprehensively benchmarking reward models in LLM alignment

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T21:19:42.227695Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-15T21:19:41.300521Z digest=sha256:95c080a0953a50438e7df9dddc26174346eead5654f5eebcbc96e9c1fb4f83e5

Observation 082da1a0-2646-43a7-a0a0-e67ca2f656d6 · outbound

This paper cites Helpsteer 2: Open-source dataset for training top-performing reward models.

Two Minds Better Than One: Collaborative Reward Modeling for LLM Alignment Helpsteer 2: Open-source dataset for training top-performing reward models

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T21:19:42.219817Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-15T21:19:41.304537Z digest=sha256:e2198b44ed6b88b44861b2f8be742c8e49501f10e6ef77aa9be89d041612db18

Observation 60ff893e-87d5-4441-bd2e-aab7d1ffb549 · outbound

This paper cites Judging llm-as-a-judge with mt-bench and chatbot arena.

Two Minds Better Than One: Collaborative Reward Modeling for LLM Alignment Judging llm-as-a-judge with mt-bench and chatbot arena

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T21:19:42.211777Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-15T21:19:41.307570Z digest=sha256:9549fb75eb3357df8c563c941b393f774a00939734b80a5aaa85c8e1b6c3954b

Observation 38fba78c-52a5-4611-8f9d-0db0dae359b5 · outbound

This paper cites Impact of preference noise on the alignment performance of generative language models.

Two Minds Better Than One: Collaborative Reward Modeling for LLM Alignment Impact of preference noise on the alignment performance of generative language models

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T21:19:42.203441Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-15T21:19:41.310840Z digest=sha256:d69bac8a0b244eda1647f9926d93fb5feeeee2c21e16a6890a4389956ba084ac

Observation 53221f72-75e5-41ce-849e-7ecbc4f10ae5 · outbound

This paper cites Improving reinforcement learning from human feedback using contrastive rewards, March 2024.

Two Minds Better Than One: Collaborative Reward Modeling for LLM Alignment Improving reinforcement learning from human feedback using contrastive rewards, March 2024

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T21:19:42.195356Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-15T21:19:41.313689Z digest=sha256:e295aecf3dfd94c08e7d10cd17c4d806b7c8f45c628f0a8fda1d504dc46f58ba

Observation 988419a2-f71a-4481-8f66-eeacda53a897 · outbound

This paper cites Goal misgeneralization in deep reinforcement learning.

Two Minds Better Than One: Collaborative Reward Modeling for LLM Alignment Goal misgeneralization in deep reinforcement learning

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T21:19:42.186985Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-15T21:19:41.316987Z digest=sha256:1f5214e3dfef68aaacf52c02fe47d7203544ead03c2838dc01d808e4d5222052

Observation fde9f6c0-a8b0-4a72-8da0-9a1a444cf8ee · outbound

This paper cites Scaling laws for reward model overoptimization.

Two Minds Better Than One: Collaborative Reward Modeling for LLM Alignment Scaling laws for reward model overoptimization

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T21:19:42.178575Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-15T21:19:41.320021Z digest=sha256:9d887c7eea24bc01f413b8a24848b2744d8b66267125316b5250a2bfe092cfcc

Observation f48b4cb1-5d4a-40ed-9ea3-02446c1a41dc · outbound

This paper cites Improving discriminative capability of reward models in rlhf using contrastive learning.

Two Minds Better Than One: Collaborative Reward Modeling for LLM Alignment Improving discriminative capability of reward models in rlhf using contrastive learning

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T21:19:42.170632Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-15T21:19:41.323132Z digest=sha256:c64f8b11acd0e62378b3dbbdb384dd0563a69dbb86c2ba70b0f6b15a61700556

Observation d58fa8a7-78b0-417d-8794-d3fe9ad0c387 · outbound

This paper cites Reward Generalization in RLHF: A Topological Perspective.

Two Minds Better Than One: Collaborative Reward Modeling for LLM Alignment Reward Generalization in RLHF: A Topological Perspective

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-15T21:19:41.326063Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:19:41.326063Z digest=sha256:fcfdc35199ae949ed09ac5d2e9f37e053a4734b2354d6babf48c0f8814f737ae

Observation 2a2abe0b-2f93-4fdb-9f83-54a3fb57ff3b · outbound

This paper cites A note on dpo with noisy preferences & relationship to ipo, 2023.

Two Minds Better Than One: Collaborative Reward Modeling for LLM Alignment A note on dpo with noisy preferences & relationship to ipo, 2023

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-15T21:19:41.329460Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:19:41.329460Z digest=sha256:4bff43e2b439dda25b8196c23c47e9759dd171c6a690beb727489c6bc556a678

Observation 16d62cf5-fd55-4a94-83e5-bc4e82b705ba · outbound

This paper cites Provably robust dpo: aligning language models with noisy feedback.

Two Minds Better Than One: Collaborative Reward Modeling for LLM Alignment Provably robust dpo: aligning language models with noisy feedback

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T21:19:42.156044Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-15T21:19:41.332351Z digest=sha256:82af9576ec09b3720358a0010ae30be71d77b2ec953e654bcf08fe24b6080112

Observation 18218c88-4406-4f6b-8e3e-56db5d47e382 · outbound

This paper cites Towards Robust Alignment of Language Models: Distributionally Robustifying Direct Preference Optimization.

Two Minds Better Than One: Collaborative Reward Modeling for LLM Alignment Towards Robust Alignment of Language Models: Distributionally Robustifying Direct Preference Optimization

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-15T21:19:41.335604Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:19:41.335604Z digest=sha256:19123c4b528079b59942dbad646b947052b4e5a6e8c3ab5ee5d863b5368a1908

Observation 4157e457-ac76-4f1e-8d66-a8849691436d · outbound

This paper cites ROPO: Robust Preference Optimization for Large Language Models.

Two Minds Better Than One: Collaborative Reward Modeling for LLM Alignment ROPO: Robust Preference Optimization for Large Language Models

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-15T21:19:41.338848Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:19:41.338848Z digest=sha256:cb85b50ea0b012a21fa5a363525542deef8e0ff12a11eaef4c04d95fea7fcfb7

Observation 96fc20ea-3b02-4aca-b22f-12ad442ef022 · outbound

This paper cites Rank analysis of incomplete block designs: I.

Two Minds Better Than One: Collaborative Reward Modeling for LLM Alignment Rank analysis of incomplete block designs: I

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-15T21:19:41.342349Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:19:41.342349Z digest=sha256:9062ef9be1b4bddbb42a89bf03348b94e35ca59fe440c96e9b31d5f6accbfa66

Observation 9ff9cb68-246c-472b-a123-41e13de42a5b · outbound

This paper cites Inform: Mitigating reward hacking in rlhf via information-theoretic reward modeling.

Two Minds Better Than One: Collaborative Reward Modeling for LLM Alignment Inform: Mitigating reward hacking in rlhf via information-theoretic reward modeling

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-15T21:19:41.345277Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:19:41.345277Z digest=sha256:76252a332c41deb94886da5592c91d26b9e767082a98101c07da6b63ab702c2e

Observation b8ccced3-09c5-4b76-b69f-9e5cb79926c2 · outbound

This paper cites Direct preference optimization: Your language model is secretly a reward model.

Two Minds Better Than One: Collaborative Reward Modeling for LLM Alignment Direct preference optimization: Your language model is secretly a reward model

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T21:19:42.137207Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-15T21:19:41.348463Z digest=sha256:e2a4eae7f8714926b30c976a885c3bd648ecbe023f53782bbf246f53b52d937f

Observation 8ef22d1c-5541-41d8-a330-f0e54ccc63c1 · outbound

This paper cites Confirma- tion bias in human reinforcement learning: Evidence from counterfactual feedback processing.

Two Minds Better Than One: Collaborative Reward Modeling for LLM Alignment Confirma- tion bias in human reinforcement learning: Evidence from counterfactual feedback processing

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T21:19:42.128912Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-15T21:19:41.351616Z digest=sha256:04a68043ac912ff6a9b2633f8364642f20568b38009472d765964e680186ed2b

Observation ba666142-e16b-42e1-9f92-c3b045e2efdb · outbound

This paper cites Pseudo- labeling and confirmation bias in deep semi-supervised learning.

Two Minds Better Than One: Collaborative Reward Modeling for LLM Alignment Pseudo- labeling and confirmation bias in deep semi-supervised learning

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T21:19:42.119261Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-15T21:19:41.354598Z digest=sha256:e1ee54b996a2c67b5bf5cf59007f08cba6646fe13c0b7ca921be6b892c7d6442

Observation cabc74fc-328b-4ccb-9ac2-c5c7083b72f4 · outbound

This paper cites Zephyr: Direct distillation of lm alignment.

Two Minds Better Than One: Collaborative Reward Modeling for LLM Alignment Zephyr: Direct distillation of lm alignment

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T21:19:42.109684Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-15T21:19:41.357680Z digest=sha256:bc5a246afa3796b431ee636d817001c4739a4be99b615e99a8c995799f081a62

Observation cf523574-1ae8-40ac-9f1a-f61880f0780e · outbound

This paper cites Smith, and Hannaneh Hajishirzi.

Two Minds Better Than One: Collaborative Reward Modeling for LLM Alignment Smith, and Hannaneh Hajishirzi

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-15T21:19:41.361790Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:19:41.361790Z digest=sha256:509e879955319b18d92eb687cc604cb810f921554e3c791ed1ec97778471f433

Observation b15471eb-e204-4329-b44c-4f008bcf0341 · outbound

This paper cites Rlhf workflow: From reward modeling to online rlhf, 2024.

Two Minds Better Than One: Collaborative Reward Modeling for LLM Alignment Rlhf workflow: From reward modeling to online rlhf, 2024

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-15T21:19:41.365456Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:19:41.365456Z digest=sha256:75a98325865ddcc791ad0f6fc2d110ed3d3937b3191be3262b7023e92e9fd941

Observation f2e51a55-b2c5-4d1c-9d37-1647e826b773 · outbound

This paper cites Vicuna: An open-source chatbot impressing gpt-4 with 90%* chatgpt quality.

Two Minds Better Than One: Collaborative Reward Modeling for LLM Alignment Vicuna: An open-source chatbot impressing gpt-4 with 90%* chatgpt quality

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-15T21:19:41.368380Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:19:41.368380Z digest=sha256:4bce5e9eeb017ff354f0902ab33772f450aabfe98d60252e095b8d2b0f1e56bb

Observation d782e6c3-2cd6-47a9-abcb-73ed3b3fe353 · outbound

This paper cites WizardLM: Empowering large pre-trained language models to follow complex instructions.

Two Minds Better Than One: Collaborative Reward Modeling for LLM Alignment WizardLM: Empowering large pre-trained language models to follow complex instructions

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-15T21:19:41.371363Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:19:41.371363Z digest=sha256:4b08ac9a8fa8351640fca4d1d4cd46d4fe493138f2233fdabdcc351a6d208ca9

Observation 54fb0898-f685-49e7-ab98-c4c3a1911a67 · outbound

This paper cites Secrets of RLHF in Large Language Models Part I: PPO.

Two Minds Better Than One: Collaborative Reward Modeling for LLM Alignment Secrets of RLHF in Large Language Models Part I: PPO

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-15T21:19:41.374445Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:19:41.374445Z digest=sha256:253285232f72886919add47a9402e9a91aa9415a09111ca25a135e8326af4991

Observation b550525c-d12b-4d96-b34b-92ca68151aa5 · outbound

This paper cites RM-bench: Benchmarking reward models of language models with subtlety and style.

Two Minds Better Than One: Collaborative Reward Modeling for LLM Alignment RM-bench: Benchmarking reward models of language models with subtlety and style

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-15T21:19:41.378271Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:19:41.378271Z digest=sha256:1389dbac6143e370aefac66dc8ba7eed1645af23194b19022440f9f590939ef5

Observation 8ff587ba-c6d5-4423-8122-eff6a4e55f94 · outbound

This paper cites From Crowdsourced Data to High-Quality Benchmarks: Arena-Hard and BenchBuilder Pipeline.

Two Minds Better Than One: Collaborative Reward Modeling for LLM Alignment From Crowdsourced Data to High-Quality Benchmarks: Arena-Hard and BenchBuilder Pipeline

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-15T21:19:41.381195Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:19:41.381195Z digest=sha256:763894a95ff3091451cbc20ed24c8f417a5300dd64fe60e339b371a6e9c04aa3

Observation 5003dbc6-318f-4334-a41e-5cfc9bc965f0 · outbound

This paper cites Weak-to-strong generalization: Eliciting strong capabilities with weak supervision.

Two Minds Better Than One: Collaborative Reward Modeling for LLM Alignment Weak-to-strong generalization: Eliciting strong capabilities with weak supervision

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T21:19:42.080708Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-15T21:19:41.384743Z digest=sha256:d7bba5520d349cd536a0a114099c9c9234c47b3bc346512888037f20b35ac3f2

Observation 8a0e4840-f774-4ab0-9b3e-0d7ffa901a37 · outbound

This paper cites Weak-to-strong prefer- ence optimization: Stealing reward from weak aligned model.

Two Minds Better Than One: Collaborative Reward Modeling for LLM Alignment Weak-to-strong prefer- ence optimization: Stealing reward from weak aligned model

Reference 38

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T21:19:42.073280Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-15T21:19:41.387657Z digest=sha256:a1c9a9f02a2297101811676554630d77b8e34e5bfdc538677995e892b4b422c1

Observation eca98603-0169-41fa-8fa7-63551354f620 · outbound

This paper cites Adversarial Training of Reward Models.

Two Minds Better Than One: Collaborative Reward Modeling for LLM Alignment Adversarial Training of Reward Models

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-15T21:19:41.390489Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:19:41.390489Z digest=sha256:48aff87b5b59225e96fc17e796ae5b6932aa36ed17c9223e971eca84b94f97e8

Observation c3ad53a6-8e84-4009-8174-79ef207ada62 · outbound

This paper cites Defining and charac- terizing reward gaming.

Two Minds Better Than One: Collaborative Reward Modeling for LLM Alignment Defining and charac- terizing reward gaming

Reference 40

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T21:19:42.065101Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-15T21:19:41.393496Z digest=sha256:2f898d37c4d3d294e66a8f66c75019fa8b7e17d3367bd1215f1581235fdd136c

Observation b4d11414-6125-41d7-a0b6-b2b415bb0752 · outbound

This paper cites Reward Shaping to Mitigate Reward Hacking in RLHF.

Two Minds Better Than One: Collaborative Reward Modeling for LLM Alignment Reward Shaping to Mitigate Reward Hacking in RLHF

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-15T21:19:41.396471Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:19:41.396471Z digest=sha256:bd7faf66d9ec208f977ff9a5059df28fca88fa91b255e623bd00dcc7ed501360

Observation bff2a27b-8fc5-41b3-b048-998c8d611d2c · outbound

This paper cites Odin: disentangled reward mitigates hacking in rlhf.

Two Minds Better Than One: Collaborative Reward Modeling for LLM Alignment Odin: disentangled reward mitigates hacking in rlhf

Reference 42

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T21:19:42.057629Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-15T21:19:41.399094Z digest=sha256:d81ea464900a4cc66ef76c386f3b24e3a46c9bb1f6307d41816dceb5cf5f928d

Observation a467c576-3136-40ee-849e-12726229087b · outbound

This paper cites Taming Overconfidence in LLMs: Reward Calibration in RLHF.

Two Minds Better Than One: Collaborative Reward Modeling for LLM Alignment Taming Overconfidence in LLMs: Reward Calibration in RLHF

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-15T21:19:41.402572Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:19:41.402572Z digest=sha256:4515059300fe502562ec23286a73d76c27f760db5a5f840c45faf95faac6b044

Observation 6c73bc40-df62-4342-a705-47d6cbc6fc6c · outbound

This paper cites Overcom- ing reward overoptimization via adversarial policy optimization with lightweight uncertainty estimation, July 2024.

Two Minds Better Than One: Collaborative Reward Modeling for LLM Alignment Overcom- ing reward overoptimization via adversarial policy optimization with lightweight uncertainty estimation, July 2024

Reference 44

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T21:19:42.049800Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-15T21:19:41.406480Z digest=sha256:f52e6ffd519e3613c4169addef3ab04188a54887ae2054c40eec33d788855975

Observation eb3802f2-ce7b-42ab-a53c-98f3a8f4f09c · outbound

This paper cites The energy loss phenomenon in rlhf: A new perspective on mitigating reward hacking, February 2025.

Two Minds Better Than One: Collaborative Reward Modeling for LLM Alignment The energy loss phenomenon in rlhf: A new perspective on mitigating reward hacking, February 2025

Reference 45

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T21:19:42.041660Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-15T21:19:41.409426Z digest=sha256:8f0533dfae0c1205e078335970c1ad35b6be765a9445800cf257a0f9127c743a

Observation 1c655f76-79e9-47ad-b6d7-7eb3abed0f85 · outbound

This paper cites Reward model ensembles help mitigate overoptimizatio.

Two Minds Better Than One: Collaborative Reward Modeling for LLM Alignment Reward model ensembles help mitigate overoptimizatio

Reference 46

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T21:19:42.034242Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-15T21:19:41.412385Z digest=sha256:4972605cafa43c449fe7504666d4ed21f1e73b5ac38be1b86f71bf60188f6672

Observation 52e9f12e-fc13-45a0-885e-f9273de244d7 · outbound

This paper cites Helping or herding? reward model ensembles mitigate but do not eliminate reward hacking.

Two Minds Better Than One: Collaborative Reward Modeling for LLM Alignment Helping or herding? reward model ensembles mitigate but do not eliminate reward hacking

Reference 47

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T21:19:42.026701Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-15T21:19:41.416160Z digest=sha256:48a8e549d6d9e8b6d584f5c07b1973c514bfa431f93913c92e6c26359accdb16

Observation 100c5a23-f993-43dc-8f28-a38a16053b51 · outbound

This paper cites Reward-robust rlhf in llms, October 2024.

Two Minds Better Than One: Collaborative Reward Modeling for LLM Alignment Reward-robust rlhf in llms, October 2024

Reference 48

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T21:19:42.019019Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-15T21:19:41.419845Z digest=sha256:2d9f5968aca09d260649a6ccb750cb88b4345ac18752ec73c61a1cc70c05b698

Observation 6b50f217-957a-487a-96e1-3f8ebff88faf · outbound

This paper cites Uncertainty-penalized reinforcement learning from human feedback with diverse reward lora ensembles, December 2023.

Two Minds Better Than One: Collaborative Reward Modeling for LLM Alignment Uncertainty-penalized reinforcement learning from human feedback with diverse reward lora ensembles, December 2023

Reference 49

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T21:19:42.011300Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-15T21:19:41.422374Z digest=sha256:94c1d380887340485106e0fbc84947e6b38d040e422f4f019a9f8b6c69e24e69

Observation 6b62b461-8298-4123-b4e8-66c177a4693e · outbound

This paper cites Interpretable prefer- ences via multi-objective reward modeling and mixture-of-experts.

Two Minds Better Than One: Collaborative Reward Modeling for LLM Alignment Interpretable prefer- ences via multi-objective reward modeling and mixture-of-experts

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-15T21:19:41.425267Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:19:41.425267Z digest=sha256:0048387be3d025959dfab6e03e5bee9b80b709867c1fcfbba0ddd965f12fa705

Observation 57244e02-3310-442d-959f-f0d6e96fe115 · outbound

This paper cites RRM: Robust Reward Model Training Mitigates Reward Hacking.

Two Minds Better Than One: Collaborative Reward Modeling for LLM Alignment RRM: Robust Reward Model Training Mitigates Reward Hacking

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-15T21:19:41.428176Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:19:41.428176Z digest=sha256:ea4723e5ca9692594413e9a8e5631669decc4edb6ff9711054027affa6dc0b84

Observation f8052a73-4469-45a1-af82-b118f48354af · outbound

This paper cites A general theoretical paradigm to understand learning from human preferences.

Two Minds Better Than One: Collaborative Reward Modeling for LLM Alignment A general theoretical paradigm to understand learning from human preferences

Reference 52

Resolution
unresolved
no resolver link, observed 2026-08-15T21:19:41.430957Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:19:41.430957Z digest=sha256:27bd599390a7e38591a629a7ea82fdf225cf674c121baed6ea14d6740d71fc5a

Observation 94c0d798-619d-468d-a22b-5d7c26826fb4 · outbound

This paper cites KTO: Model Alignment as Prospect Theoretic Optimization.

Two Minds Better Than One: Collaborative Reward Modeling for LLM Alignment KTO: Model Alignment as Prospect Theoretic Optimization

Reference 53

Resolution
unresolved
no resolver link, observed 2026-08-15T21:19:41.434016Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:19:41.434016Z digest=sha256:892fa98c58dd7fda4207c2a54bea91fa7ec29f684edce30f5d3b948d9b45cbf7

Observation 951e9b08-4d2f-40db-9183-2f049dbd036b · outbound

This paper cites Orpo: Monolithic preference optimization without reference model.

Two Minds Better Than One: Collaborative Reward Modeling for LLM Alignment Orpo: Monolithic preference optimization without reference model

Reference 54

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T21:19:41.993237Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-15T21:19:41.437862Z digest=sha256:3135d29c3e891ffc5edf655921a163e08b8d66d975b9a13d3501bc432e26125d

Observation 5be45d8c-f341-4446-a2aa-549e0898254b · outbound

This paper cites Simpo: Simple preference optimization with a reference-free reward.

Two Minds Better Than One: Collaborative Reward Modeling for LLM Alignment Simpo: Simple preference optimization with a reference-free reward

Reference 55

Resolution
unresolved
no resolver link, observed 2026-08-15T21:19:41.440975Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:19:41.440975Z digest=sha256:7f2e0aca15d6ec7eb3682b36ac7371b709f60bbd0ce179baef078cf26169c366

Observation 0020e4f9-2fa6-4555-a8a1-c7af9ab23f5b · outbound

This paper cites Is dpo superior to ppo for llm alignment? a comprehensive study.

Two Minds Better Than One: Collaborative Reward Modeling for LLM Alignment Is dpo superior to ppo for llm alignment? a comprehensive study

Reference 56

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T21:19:41.978016Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-15T21:19:41.443647Z digest=sha256:ca9ea2a1dd6663efe8d7e3249cb498be89f5849a3056cdc3743ca32c78273cda

Observation 3b7cb61e-1722-40ea-97b9-284057dc51b6 · outbound

This paper cites Smith, Yejin Choi, and Hannaneh Hajishirzi.

Two Minds Better Than One: Collaborative Reward Modeling for LLM Alignment Smith, Yejin Choi, and Hannaneh Hajishirzi

Reference 57

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T21:19:41.968885Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-15T21:19:41.448068Z digest=sha256:36d207007e47c3e280b546c6026d60c30c7e7023f9d228e13f5800907c8da4a4

Observation c9c7c98f-214b-4e3c-9397-7cc677826137 · outbound

This paper cites Direct Language Model Alignment from Online AI Feedback.

Two Minds Better Than One: Collaborative Reward Modeling for LLM Alignment Direct Language Model Alignment from Online AI Feedback

Reference 58

Resolution
unresolved
no resolver link, observed 2026-08-15T21:19:41.451174Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:19:41.451174Z digest=sha256:cf3f3cc41e0169e1f0f1f9820d292d35b3d1c22cfa59af2f3ce9cfc394e5e1a7

Observation 5d220a40-9759-4e58-a761-1d70cbc63392 · outbound

This paper cites DPO-Shift: Shifting the Distribution of Direct Preference Optimization.

Two Minds Better Than One: Collaborative Reward Modeling for LLM Alignment DPO-Shift: Shifting the Distribution of Direct Preference Optimization

Reference 59

Resolution
unresolved
no resolver link, observed 2026-08-15T21:19:41.454892Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:19:41.454892Z digest=sha256:5e83c48edfe39b381a2bb6717459eceb009b6b3279a36557cd558e390ba620c7

Observation 9557c46c-3769-4070-9a87-bd0f00de5dd9 · outbound

This paper cites Understanding generalization of preference optimization under noisy feedback.

Two Minds Better Than One: Collaborative Reward Modeling for LLM Alignment Understanding generalization of preference optimization under noisy feedback

Reference 60

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T21:19:41.959084Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-15T21:19:41.458038Z digest=sha256:e7506455d9f9d2295132c886afcc5d4c2ed21dd5c7d079ade748b51db453913c

Observation 90ca6071-d427-4f0c-a23d-4cdaa5e807cb · outbound

This paper cites Robust reinforcement learning from corrupted human feedback.

Two Minds Better Than One: Collaborative Reward Modeling for LLM Alignment Robust reinforcement learning from corrupted human feedback

Reference 61

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T21:19:41.949483Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-15T21:19:41.461332Z digest=sha256:391d35adb61a11f020d1648c770a5bb91c26a271c21231db5218ac7ed077c7d6

Observation 19fd781e-d323-47f1-9ff3-051373fba9cc · outbound

This paper cites Theory of games and economic behavior, 60th- anniversary, 2007.

Two Minds Better Than One: Collaborative Reward Modeling for LLM Alignment Theory of games and economic behavior, 60th- anniversary, 2007

Reference 62

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T21:19:41.941040Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-15T21:19:41.465070Z digest=sha256:60be15175be4e5aeef4646337f774098b5637da1171af855a34624ba7901a6a7

Observation 24cdd193-71e1-4a6d-aec3-67c2dbc7e614 · outbound

This paper cites Multiagent systems: Algorithmic, game-theoretic, and logical foundations.

Two Minds Better Than One: Collaborative Reward Modeling for LLM Alignment Multiagent systems: Algorithmic, game-theoretic, and logical foundations

Reference 63

Resolution
unresolved
no resolver link, observed 2026-08-15T21:19:41.468046Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:19:41.468046Z digest=sha256:fadaf916dbbf3e3b12826b199e8c686488e48d1ea2cc80824ba72520c15a33aa

Observation f30cc94a-75ce-452a-a767-917d0e9a53e8 · outbound

This paper cites [Yes] " is generally preferable to.

Two Minds Better Than One: Collaborative Reward Modeling for LLM Alignment [Yes] " is generally preferable to

Reference 64

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T21:19:41.926924Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-15T21:19:41.472095Z digest=sha256:1b073633b27e56496c48adb470d542da3788a5f5f62f499826efe7b6ee74c66f

Observation 587f0f17-064a-44f2-adff-90fc31e59d3d · outbound

This paper cites Guidelines: • The answer NA means that the abstract and introduction do not include the claims made in the paper.

Two Minds Better Than One: Collaborative Reward Modeling for LLM Alignment Guidelines: • The answer NA means that the abstract and introduction do not include the claims made in the paper

Reference 65

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T21:19:41.918496Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-15T21:19:41.475861Z digest=sha256:c3f269d07764cc85a6ee1350df723faa272ad9a1b0fd8e98e22431328e0729e7

Observation 2587aa68-18fb-45d6-b3f9-7186901ca14d · outbound

This paper cites Limitations.

Two Minds Better Than One: Collaborative Reward Modeling for LLM Alignment Limitations

Reference 66

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T21:19:41.910990Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-15T21:19:41.479196Z digest=sha256:15041f19e18153e139c4b8bab7ae4b6a088da4926b7c53548172f6c0c02f5252

Observation 0ae68a1d-93c7-4ca9-bf0f-efb0910d40e3 · outbound

This paper cites Guidelines: • The answer NA means that the paper does not include theoretical results.

Two Minds Better Than One: Collaborative Reward Modeling for LLM Alignment Guidelines: • The answer NA means that the paper does not include theoretical results

Reference 67

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T21:19:41.902892Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-15T21:19:41.483442Z digest=sha256:8b71a3ee89cc420e4dcaab9e3c460867b130cb22c8e139d56ede1b3906ce0e73

Observation 45575863-75fc-491c-b97b-d3828cd3544c · outbound

This paper cites Guidelines: • The answer NA means that the paper does not include experiments.

Two Minds Better Than One: Collaborative Reward Modeling for LLM Alignment Guidelines: • The answer NA means that the paper does not include experiments

Reference 68

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T21:19:41.894349Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-15T21:19:41.486539Z digest=sha256:da4270d606eff225dbaeac80ca08852dc5c6814346e2a12effa9fa326d761540

Observation 28b3ac97-92c3-44fd-808d-72b200a62c21 · outbound

This paper cites Guidelines: • The answer NA means that paper does not include experiments requiring code.

Two Minds Better Than One: Collaborative Reward Modeling for LLM Alignment Guidelines: • The answer NA means that paper does not include experiments requiring code

Reference 69

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T21:19:41.884914Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-15T21:19:41.490836Z digest=sha256:58143a1d78631617fba122742a2e1b7aa392314cdf211749e2f7233b8ba3bb1b

Observation ca42c146-6f32-41b1-a962-51a0703b86ee · outbound

This paper cites Guidelines: • The answer NA means that the paper does not include experiments.

Two Minds Better Than One: Collaborative Reward Modeling for LLM Alignment Guidelines: • The answer NA means that the paper does not include experiments

Reference 70

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T21:19:41.876158Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-15T21:19:41.494510Z digest=sha256:e0fdef453f29dd054f666d1295c691c0980432ce7d20be3ec9b9449806ae1aef

Observation 6e4a06e4-3921-4646-909d-6cfebcadef07 · outbound

This paper cites Guidelines: • The answer NA means that the paper does not include experiments.

Two Minds Better Than One: Collaborative Reward Modeling for LLM Alignment Guidelines: • The answer NA means that the paper does not include experiments

Reference 71

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T21:19:41.866146Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-15T21:19:41.497449Z digest=sha256:c77afbcad06b1f44567e1c64cda45cf2de57126647793cbf6f51270e30a3e1ae

Observation e0aa576d-be3d-41bc-a356-eb2cdc6abde9 · outbound

This paper cites Guidelines: • The answer NA means that the paper does not include experiments.

Two Minds Better Than One: Collaborative Reward Modeling for LLM Alignment Guidelines: • The answer NA means that the paper does not include experiments

Reference 72

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T21:19:41.855662Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-15T21:19:41.500484Z digest=sha256:ef56e324a58b2173fd5d57de94381ceed1719a293bf4c1da416b246d1fe86767

Observation c43dba8f-bd3f-4f05-ad03-59ed9980f531 · outbound

This paper cites Guidelines: • The answer NA means that the authors have not reviewed the NeurIPS Code of Ethics.

Two Minds Better Than One: Collaborative Reward Modeling for LLM Alignment Guidelines: • The answer NA means that the authors have not reviewed the NeurIPS Code of Ethics

Reference 73

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T21:19:41.843927Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-15T21:19:41.505139Z digest=sha256:a7d312a3d290f7132ee469214194ef1af71fe1f3703c72253f4bf1102190096e

Observation 76be5127-3bb9-4fb1-bdfe-168239b3e32c · outbound

This paper cites an unresolved cited work.

Two Minds Better Than One: Collaborative Reward Modeling for LLM Alignment Unresolved cited work

Reference 74

Resolution
unresolved
raw_fallback, observed 2026-08-15T21:19:41.828840Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-15T21:19:41.508480Z digest=sha256:7de16059cee9b5bf5c296659d2b2607c101b953a7705cec57968e1f85036e60a

Observation 187b16d1-8534-48d7-ba67-1575806c6f2b · outbound

This paper cites an unresolved cited work.

Two Minds Better Than One: Collaborative Reward Modeling for LLM Alignment Unresolved cited work

Reference 75

Resolution
unresolved
raw_fallback, observed 2026-08-15T21:19:41.811120Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-15T21:19:41.512195Z digest=sha256:f81fb7d5f24a28b0f1c3b587441f8515e341bc9fdeb9ee5d097a414c75d6a934

Observation a29dce32-996e-4b09-9b85-e560dff6dfb2 · outbound

This paper cites Guidelines: • The answer NA means that the paper does not use existing assets.

Two Minds Better Than One: Collaborative Reward Modeling for LLM Alignment Guidelines: • The answer NA means that the paper does not use existing assets

Reference 76

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T21:19:41.792568Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-15T21:19:41.515056Z digest=sha256:704b69f9c4ea8498074d4e9cdabee372622642528b2c2ee9df31fbb1d4b1919f

Observation 27395a9e-b895-4c14-88eb-bfb3742071e2 · outbound

This paper cites • Researchers should communicate the details of the dataset/code/model as part of their submissions via structured templates.

Two Minds Better Than One: Collaborative Reward Modeling for LLM Alignment • Researchers should communicate the details of the dataset/code/model as part of their submissions via structured templates

Reference 77

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T21:19:41.772923Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-15T21:19:41.517980Z digest=sha256:4ff89ae3deedaf324f91a585717ac4b518ad50b323693ba07279f5c88326fc1f

Observation 4ec17451-33bb-4202-ad24-41121e2d824b · outbound

This paper cites 32 Guidelines: • The answer NA means that the paper does not involve crowdsourcing nor research with human subjects.

Two Minds Better Than One: Collaborative Reward Modeling for LLM Alignment 32 Guidelines: • The answer NA means that the paper does not involve crowdsourcing nor research with human subjects

Reference 78

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T21:19:41.754738Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-15T21:19:41.520830Z digest=sha256:9d1ed5cec993de571970e822125b97ed2eeef0b0dba7ac0e39f5663a88f53caa

Observation 009638ff-9737-402e-ab8d-ae28df7ecc55 · outbound

This paper cites Guidelines: • The answer NA means that the paper does not involve crowdsourcing nor research with human subjects.

Two Minds Better Than One: Collaborative Reward Modeling for LLM Alignment Guidelines: • The answer NA means that the paper does not involve crowdsourcing nor research with human subjects

Reference 79

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T21:19:41.737498Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-15T21:19:41.524216Z digest=sha256:6f9da34f722ce496f37304bf3d4959a21c91ec38f238ee3e578c062f6e6c8027

Observation 767017fa-c7a9-451f-8aa4-13ab29e25803 · outbound

This paper cites Answer: [NA] Justification: LLM is used only for editing, or formatting and does not impact the core methodology, scientific rigorousness.

Two Minds Better Than One: Collaborative Reward Modeling for LLM Alignment Answer: [NA] Justification: LLM is used only for editing, or formatting and does not impact the core methodology, scientific rigorousness

Reference 80

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T21:19:41.722866Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-15T21:19:41.527363Z digest=sha256:1715ea032eae86eab859e54d38d72014b8d90f92849f2048d8694a0f23e0c10e

Pith citing papers

Observation b431773f-cf7b-4d93-9727-f0dfce4a409b · inbound

AgentV-RL: Scaling Reward Modeling with Agentic Verifier cites this paper.

AgentV-RL: Scaling Reward Modeling with Agentic Verifier Two Minds Better Than One: Collaborative Reward Modeling for LLM Alignment

Reference 6

Resolution
verified exact
arxiv_id, observed 2026-05-10T08:32:51.696658Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-05-10T08:32:38.883019Z digest=sha256:80308f3e0c88d38ec68d367e61630fed3cbb62f70a049a833fc1f85af7eeef55