Pith. sign in

Paper Citation Record · LEDGER

RRC: Unlocking Generative Reward Models in LLM Reinforcement Learning via Ranking-Based Reward Construction

As of 11 August 2026, this Paper Citation Record lists 77 of 77 outbound references and 0 inbound Pith citation observations for arXiv:2608.06310.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2608.06310 v1

Coverage vector

measured 77 of 77 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-07T05:49:51.597998Z

measured 77 of 77 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-11T06:34:44.6726+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

77 of 77 outbound references displayed

  • verified exact0
  • verified fuzzy62
  • unresolved15
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation b12cabb1-843c-474d-a18d-35c787b36ab9 · outbound

This paper cites Advances in Neural Information Processing Systems , volume=.

RRC: Unlocking Generative Reward Models in LLM Reinforcement Learning via Ranking-Based Reward Construction Advances in Neural Information Processing Systems , volume=

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:49:52.743820Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=arxiv_source observed=2026-08-07T05:49:45.244621Z digest=sha256:cf7f102a4d38c1d995811741865d34949ae5b70ea84058b88e6bff6138e499fd

Observation 22df6d52-8ba6-4011-9982-cd97d2cc0ea1 · outbound

This paper cites Proceedings of the AAAI Conference on Artificial Intelligence , volume=.

RRC: Unlocking Generative Reward Models in LLM Reinforcement Learning via Ranking-Based Reward Construction Proceedings of the AAAI Conference on Artificial Intelligence , volume=

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:49:52.729886Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=arxiv_source observed=2026-08-07T05:49:45.347881Z digest=sha256:c37a3705105a53cc319f5ebf10da43b49f19cc556e074990914a9f044dee411f

Observation 874c2201-b413-4803-bccf-15fc82642af2 · outbound

This paper cites Unified Reward Model for Multimodal Understanding and Generation.

RRC: Unlocking Generative Reward Models in LLM Reinforcement Learning via Ranking-Based Reward Construction Unified Reward Model for Multimodal Understanding and Generation

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-07T05:49:45.440896Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T05:49:45.440896Z digest=sha256:7faff5705fdb7835ea92ccb6acc937f804fd77916126b79579928461fa35bb98

Observation 872aecdc-9cd9-4a4f-a1a1-b8c05bff6acf · outbound

This paper cites Advances in neural information processing systems , volume=.

RRC: Unlocking Generative Reward Models in LLM Reinforcement Learning via Ranking-Based Reward Construction Advances in neural information processing systems , volume=

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:49:52.714983Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=arxiv_source observed=2026-08-07T05:49:45.553185Z digest=sha256:d866bf2726e431c2fc6c1f23f1f9b4244c48b9f3acdd09b5f2d7e7595ac45c56

Observation f2875dae-92c1-47ac-88bf-3db3247f8185 · outbound

This paper cites Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers) , pages=.

RRC: Unlocking Generative Reward Models in LLM Reinforcement Learning via Ranking-Based Reward Construction Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers) , pages=

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:49:52.699831Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=arxiv_source observed=2026-08-07T05:49:45.641778Z digest=sha256:a589fe5601e5704214ff84a00620fa205b9c66515cee645844a7f8d906bb86fa

Observation e834e349-dfe1-4c35-b6a7-699bc4d1b514 · outbound

This paper cites WorldPM: Scaling Human Preference Modeling.

RRC: Unlocking Generative Reward Models in LLM Reinforcement Learning via Ranking-Based Reward Construction WorldPM: Scaling Human Preference Modeling

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-07T05:49:45.723481Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T05:49:45.723481Z digest=sha256:632df4b2af19a40d853003b8f0768e0f5e470d654bbea7d64dc5ac61b42dde81

Observation e1eae25c-98dc-46f4-bd31-67057d34c166 · outbound

This paper cites Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers) , pages=.

RRC: Unlocking Generative Reward Models in LLM Reinforcement Learning via Ranking-Based Reward Construction Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers) , pages=

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:49:52.684677Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=arxiv_source observed=2026-08-07T05:49:45.828879Z digest=sha256:284a4c2b5aebcad41383dda2ea8b3e1f77280d8f100f55611bf08b7800ce9bcb

Observation 5ac76aea-1516-4274-aa1a-c86127d9e040 · outbound

This paper cites Reward is enough: Llms are in-context reinforcement learners , volume =.

RRC: Unlocking Generative Reward Models in LLM Reinforcement Learning via Ranking-Based Reward Construction Reward is enough: Llms are in-context reinforcement learners , volume =

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:49:52.669302Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=arxiv_source observed=2026-08-07T05:49:45.901475Z digest=sha256:4c3fe642dd7b156f67cba8bd6fdf7a18cceac224b0f6c92c281d5ec4940c2797

Observation d9c87546-9603-4be6-9e86-b1c393da434a · outbound

This paper cites an unresolved cited work.

RRC: Unlocking Generative Reward Models in LLM Reinforcement Learning via Ranking-Based Reward Construction Unresolved cited work

Reference 9

Resolution
unresolved
raw_fallback, observed 2026-08-07T05:49:52.654417Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=arxiv_source observed=2026-08-07T05:49:45.971684Z digest=sha256:ce9931559609f1d876ba5125e6818504b9df6b1beee25c59584738ea79c7f39f

Observation 88ac1e2a-c41b-44a4-821e-b08628422215 · outbound

This paper cites Scaling laws for neural language models , volume =.

RRC: Unlocking Generative Reward Models in LLM Reinforcement Learning via Ranking-Based Reward Construction Scaling laws for neural language models , volume =

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:49:52.640904Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=arxiv_source observed=2026-08-07T05:49:46.075861Z digest=sha256:682f654a5b100100b34979f981ecd365044278fb0495421a4a7f8c5e8311e78d

Observation 3c3453f5-e0ce-45b3-8d12-65c133728908 · outbound

This paper cites Large language models are not fair evaluators , year =.

RRC: Unlocking Generative Reward Models in LLM Reinforcement Learning via Ranking-Based Reward Construction Large language models are not fair evaluators , year =

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:49:52.626147Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=arxiv_source observed=2026-08-07T05:49:46.128006Z digest=sha256:a54ef25e434cb340873627c942f633b1487c7e5b9857b5b6c9bb098ce33f2b0f

Observation a26e3229-0fc7-499a-9efc-b1b29c362500 · outbound

This paper cites Theoretical and empirical evaluation of data reduction for exact Kemeny rank aggregation , volume =.

RRC: Unlocking Generative Reward Models in LLM Reinforcement Learning via Ranking-Based Reward Construction Theoretical and empirical evaluation of data reduction for exact Kemeny rank aggregation , volume =

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:49:52.610894Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=arxiv_source observed=2026-08-07T05:49:46.200418Z digest=sha256:3d8f09c0c216ebc634aee26c35b92fbe3202d88da666b50232ebe51d410d77a9

Observation abc954f5-9411-4fb6-88fb-fb668c94d847 · outbound

This paper cites Advances in Neural Information Processing Systems , volume=.

RRC: Unlocking Generative Reward Models in LLM Reinforcement Learning via Ranking-Based Reward Construction Advances in Neural Information Processing Systems , volume=

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:49:52.596995Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=arxiv_source observed=2026-08-07T05:49:46.264080Z digest=sha256:c9fde0fd71a201fe6e18acd84ea2ce1b3b9a0fdcd3bd99bf93595ce0f0266733

Observation a288283c-417e-43f7-b702-0afebe30db57 · outbound

This paper cites Improved parameterized algorithms for the Kemeny aggregation problem , year =.

RRC: Unlocking Generative Reward Models in LLM Reinforcement Learning via Ranking-Based Reward Construction Improved parameterized algorithms for the Kemeny aggregation problem , year =

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:49:52.581671Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=arxiv_source observed=2026-08-07T05:49:46.349746Z digest=sha256:dd9c3a629ef8ce8710edf6840742219622fbd3976131274e0fb9ce6a9e9eb207

Observation a3233059-6b6a-45b9-9c8d-52dbd3cd32d7 · outbound

This paper cites Are we done with mmlu? , year =.

RRC: Unlocking Generative Reward Models in LLM Reinforcement Learning via Ranking-Based Reward Construction Are we done with mmlu? , year =

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:49:52.566449Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=arxiv_source observed=2026-08-07T05:49:46.436711Z digest=sha256:74cb514bcb6dc1aa07c816010476e81628b246483ff33a5d5ee4ab1f461a28ba

Observation 9dddf58f-edbd-4289-ae91-9ba7ba432359 · outbound

This paper cites Length-controlled alpacaeval: A simple debiasing of automatic evaluators , year =.

RRC: Unlocking Generative Reward Models in LLM Reinforcement Learning via Ranking-Based Reward Construction Length-controlled alpacaeval: A simple debiasing of automatic evaluators , year =

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:49:52.550882Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=arxiv_source observed=2026-08-07T05:49:46.503362Z digest=sha256:01df86ca9de590c1bc96f4ab8e2fcccf2f6432a04f3b3e9f4e583ead1933e56d

Observation 3fd36120-3b59-402d-8147-bfbd4713ea10 · outbound

This paper cites Let's Verify Step by Step , year =.

RRC: Unlocking Generative Reward Models in LLM Reinforcement Learning via Ranking-Based Reward Construction Let's Verify Step by Step , year =

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:49:52.536480Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=arxiv_source observed=2026-08-07T05:49:46.594152Z digest=sha256:0e3bea3982302feea095163d76e43e78e07a3976eeae771db2042df4c87d57bb

Observation cfcb2e39-e491-480b-a1b7-85d637b189bc · outbound

This paper cites Gpqa: A graduate-level google-proof q&a benchmark , year =.

RRC: Unlocking Generative Reward Models in LLM Reinforcement Learning via Ranking-Based Reward Construction Gpqa: A graduate-level google-proof q&a benchmark , year =

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:49:52.522331Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=arxiv_source observed=2026-08-07T05:49:46.638322Z digest=sha256:3e7088033e7696437b35b0ca136e8d4e7d38affb121267dde1932982f81b2894

Observation 0fa929c0-9b64-4521-a979-12c017d6a46f · outbound

This paper cites From crowdsourced data to high-quality benchmarks: Arena-hard and benchbuilder pipeline , volume =.

RRC: Unlocking Generative Reward Models in LLM Reinforcement Learning via Ranking-Based Reward Construction From crowdsourced data to high-quality benchmarks: Arena-hard and benchbuilder pipeline , volume =

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:49:52.508036Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=arxiv_source observed=2026-08-07T05:49:46.727147Z digest=sha256:6b6840389b1fcdf57839ddec0ac88833a1ffed74c0e93f9648acab097d74e729

Observation 3fe14d06-9186-4c56-9830-380a238aacd7 · outbound

This paper cites Wildbench: Benchmarking llms with challenging tasks from real users in the wild , volume =.

RRC: Unlocking Generative Reward Models in LLM Reinforcement Learning via Ranking-Based Reward Construction Wildbench: Benchmarking llms with challenging tasks from real users in the wild , volume =

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:49:52.493263Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=arxiv_source observed=2026-08-07T05:49:46.762252Z digest=sha256:229260c4ee95c8fa57d1a54b1b93bbc5654405835acd17cd7c3417f5fd6b4eee

Observation 3e82869b-6700-4aab-a177-234bab8f057f · outbound

This paper cites Hashimoto , howpublished =.

RRC: Unlocking Generative Reward Models in LLM Reinforcement Learning via Ranking-Based Reward Construction Hashimoto , howpublished =

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:49:52.478353Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=arxiv_source observed=2026-08-07T05:49:46.900604Z digest=sha256:238a43047345627fff083e843b5e22425618711965964d35864faccd5d49dbd7

Observation 355a6e70-0d2d-43f0-89b1-a2fda7ff1101 · outbound

This paper cites SimPO: Simple Preference Optimization with a Reference-Free Reward , year =.

RRC: Unlocking Generative Reward Models in LLM Reinforcement Learning via Ranking-Based Reward Construction SimPO: Simple Preference Optimization with a Reference-Free Reward , year =

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:49:52.462913Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=arxiv_source observed=2026-08-07T05:49:47.007901Z digest=sha256:bc6ef8948201828e8973c2c1c2fcdd7b7121e2fa6262a4d9a4a9c8deaea86e84

Observation 19117320-1377-4f01-b117-5727f531c0b1 · outbound

This paper cites HelpSteer3-Preference: Open Human-Annotated Preference Data across Diverse Tasks and Languages , volume =.

RRC: Unlocking Generative Reward Models in LLM Reinforcement Learning via Ranking-Based Reward Construction HelpSteer3-Preference: Open Human-Annotated Preference Data across Diverse Tasks and Languages , volume =

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:49:52.448710Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=arxiv_source observed=2026-08-07T05:49:47.133870Z digest=sha256:95e67f0497fef36f88e009d8e42485c206fa8a63b5647690173c3250b34e2ac1

Observation 9d034b94-468a-43be-adde-ba6bbeade5ea · outbound

This paper cites Qwen2 technical report , volume =.

RRC: Unlocking Generative Reward Models in LLM Reinforcement Learning via Ranking-Based Reward Construction Qwen2 technical report , volume =

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:49:52.431704Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=arxiv_source observed=2026-08-07T05:49:47.251853Z digest=sha256:8ebd1d7d94590d26c14acf1299adefcb8ac6d05f58aa72ef7b811b4ca5edb583

Observation 365a72b4-5c00-4d89-bb4f-1ce13ea0d5f8 · outbound

This paper cites The llama 3 herd of models , volume =.

RRC: Unlocking Generative Reward Models in LLM Reinforcement Learning via Ranking-Based Reward Construction The llama 3 herd of models , volume =

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:49:52.415158Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=arxiv_source observed=2026-08-07T05:49:47.388928Z digest=sha256:be02bd8739d8c4177ea27252be0223e7931825c2f70cd07aef9a90bf6a8aac31

Observation e5370f44-3ffc-4801-8fbe-d55fc5530a3a · outbound

This paper cites Rrhf: Rank responses to align language models with human feedback without tears , year =.

RRC: Unlocking Generative Reward Models in LLM Reinforcement Learning via Ranking-Based Reward Construction Rrhf: Rank responses to align language models with human feedback without tears , year =

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:49:52.400745Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=arxiv_source observed=2026-08-07T05:49:47.473399Z digest=sha256:b271df903898e5c56049a79103fd0b6323f0dc3f917ea61ec96681e761128916

Observation 85fedc31-14ce-4255-b22b-5d720dbc7908 · outbound

This paper cites A computational study of the Kemeny rule for preference aggregation , year =.

RRC: Unlocking Generative Reward Models in LLM Reinforcement Learning via Ranking-Based Reward Construction A computational study of the Kemeny rule for preference aggregation , year =

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:49:52.385608Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=arxiv_source observed=2026-08-07T05:49:47.596594Z digest=sha256:754ff2b22742ecb1a79c02b3fb837d4191e798f19cb1d76dc86dd99b9806333d

Observation 59d360b3-9b3d-456f-8839-3bdac42719d9 · outbound

This paper cites Judgebench: A benchmark for evaluating llm-based judges , year =.

RRC: Unlocking Generative Reward Models in LLM Reinforcement Learning via Ranking-Based Reward Construction Judgebench: A benchmark for evaluating llm-based judges , year =

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:49:52.371174Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=arxiv_source observed=2026-08-07T05:49:47.725278Z digest=sha256:ecc1d04b5acfac74a437dd8d0e0689214d34f784d3528e69d82e83473801a0ee

Observation 76e7c7a9-8a8b-4d97-8ccb-726744604d13 · outbound

This paper cites Rm-bench: Benchmarking reward models of language models with subtlety and style , year =.

RRC: Unlocking Generative Reward Models in LLM Reinforcement Learning via Ranking-Based Reward Construction Rm-bench: Benchmarking reward models of language models with subtlety and style , year =

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:49:52.356201Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=arxiv_source observed=2026-08-07T05:49:47.767357Z digest=sha256:6cb672c8339d5e37c4afb9af3795eea2e315d69966a22af4448fd19bed5f2cd1

Observation fa601f3e-928d-4972-b376-ebb6eb3ad227 · outbound

This paper cites Proximal policy optimization algorithms , year =.

RRC: Unlocking Generative Reward Models in LLM Reinforcement Learning via Ranking-Based Reward Construction Proximal policy optimization algorithms , year =

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-07T05:49:47.826230Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T05:49:47.826230Z digest=sha256:335e07c0626b2bae0b4d9c91600e4560165053de2b52e2690056a1f6e105683f

Observation 9ff36b12-62a6-450f-bea7-5ed06b613eb1 · outbound

This paper cites Rank analysis of incomplete block designs: I.

RRC: Unlocking Generative Reward Models in LLM Reinforcement Learning via Ranking-Based Reward Construction Rank analysis of incomplete block designs: I

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:49:52.331929Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=arxiv_source observed=2026-08-07T05:49:47.882955Z digest=sha256:b159ec6d392d7c2e82c876f22a2fe01fe8f3dde73ac31d0aadca3135c99fc203

Observation 85e72548-907a-4bea-85fa-b08ebf5e0fd7 · outbound

This paper cites Reinforcement Learning with Verifiable Rewards: GRPO's Effective Loss, Dynamics, and Success Amplification , year =.

RRC: Unlocking Generative Reward Models in LLM Reinforcement Learning via Ranking-Based Reward Construction Reinforcement Learning with Verifiable Rewards: GRPO's Effective Loss, Dynamics, and Success Amplification , year =

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:49:52.317716Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=arxiv_source observed=2026-08-07T05:49:47.917984Z digest=sha256:c3f7e6afd1d8299cb4f5823e21620519b54f34c8f8ecabb6a4e3c5b7caba4312

Observation 8a42ff62-ad78-4892-9759-5b79eb8f1844 · outbound

This paper cites Language models that think, chat better , year =.

RRC: Unlocking Generative Reward Models in LLM Reinforcement Learning via Ranking-Based Reward Construction Language models that think, chat better , year =

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:49:52.303692Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=arxiv_source observed=2026-08-07T05:49:47.979978Z digest=sha256:f40e30a9f4b73760231e2460e00bafe1f71f5488444c80b84de7e5e6675bec0a

Observation ef34f03e-4bb5-40be-849c-abb9c80bb355 · outbound

This paper cites Dissecting Long Reasoning Models: An Empirical Study , year =.

RRC: Unlocking Generative Reward Models in LLM Reinforcement Learning via Ranking-Based Reward Construction Dissecting Long Reasoning Models: An Empirical Study , year =

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:49:52.290057Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=arxiv_source observed=2026-08-07T05:49:48.040116Z digest=sha256:c4bad7cad9eafb87ca0388aefd9e52bcbf0adc03e6620f0ac018d92e6dcdd139

Observation defa5713-cebe-4c74-93d7-09fa59f6532f · outbound

This paper cites Reinforcement learning with verifiable rewards implicitly incentivizes correct reasoning in base llms , year =.

RRC: Unlocking Generative Reward Models in LLM Reinforcement Learning via Ranking-Based Reward Construction Reinforcement learning with verifiable rewards implicitly incentivizes correct reasoning in base llms , year =

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:49:52.275348Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=arxiv_source observed=2026-08-07T05:49:48.103572Z digest=sha256:807283b7f0d6005e755e0b638f2bd6a30600afa9ddb2919d624aa04f516ba05a

Observation b4c28ff7-c18d-4f2b-ac58-2b744eb86fe8 · outbound

This paper cites an unresolved cited work.

RRC: Unlocking Generative Reward Models in LLM Reinforcement Learning via Ranking-Based Reward Construction Unresolved cited work

Reference 36

Resolution
unresolved
raw_fallback, observed 2026-08-07T05:49:52.261521Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=arxiv_source observed=2026-08-07T05:49:48.168385Z digest=sha256:d283b112c3bfd2d80776ae0cb3fd3db640ab422c7d90038bf3f331bae81aa406

Observation 1b3741e8-8e3a-45a2-9816-bb41d5a2b51c · outbound

This paper cites Pre-Trained Policy Discriminators are General Reward Models , year =.

RRC: Unlocking Generative Reward Models in LLM Reinforcement Learning via Ranking-Based Reward Construction Pre-Trained Policy Discriminators are General Reward Models , year =

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:49:52.247207Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=arxiv_source observed=2026-08-07T05:49:48.209023Z digest=sha256:d082905f114821382a1f2df65c795e19f1fa34d7d5cd6c277848226caf11e3db

Observation ea53c487-dc4c-4eeb-ba1a-bdbd04c02b01 · outbound

This paper cites Dynamic Reward Adjustment in Multi-Reward Reinforcement Learning for Counselor Reflection Generation , year =.

RRC: Unlocking Generative Reward Models in LLM Reinforcement Learning via Ranking-Based Reward Construction Dynamic Reward Adjustment in Multi-Reward Reinforcement Learning for Counselor Reflection Generation , year =

Reference 38

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:49:52.231862Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=arxiv_source observed=2026-08-07T05:49:48.289493Z digest=sha256:dd934f517987ebc79337f4d71758b47383be61f178ca9a336c1e75b0c3c3100d

Observation 91696147-46b0-43a5-866b-90ec5bcbbe09 · outbound

This paper cites Contrastive Preference Optimization: Pushing the Boundaries of.

RRC: Unlocking Generative Reward Models in LLM Reinforcement Learning via Ranking-Based Reward Construction Contrastive Preference Optimization: Pushing the Boundaries of

Reference 39

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:49:52.217493Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=arxiv_source observed=2026-08-07T05:49:48.361116Z digest=sha256:e6514ae3b8d1b4331ce2d67a5a191082665321fbf5f3ea919379298b2374b9ff

Observation e2cfe807-ba89-40b6-af9b-ae97d316bef8 · outbound

This paper cites From system 1 to system 2: A survey of reasoning large language models , year =.

RRC: Unlocking Generative Reward Models in LLM Reinforcement Learning via Ranking-Based Reward Construction From system 1 to system 2: A survey of reasoning large language models , year =

Reference 40

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:49:52.202329Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=arxiv_source observed=2026-08-07T05:49:48.425028Z digest=sha256:1191e2f4627215db18b0fdbf9b3bfbb23c3e8c47b4e4bb9dd27d21171232868a

Observation 42037de2-a9df-499d-9cc1-42189ddef0e3 · outbound

This paper cites an unresolved cited work.

RRC: Unlocking Generative Reward Models in LLM Reinforcement Learning via Ranking-Based Reward Construction Unresolved cited work

Reference 41

Resolution
unresolved
raw_fallback, observed 2026-08-07T05:49:52.187097Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=arxiv_source observed=2026-08-07T05:49:48.464368Z digest=sha256:19854c9baee2e2e142d3aacab5edf13cbd847fca780b65497cedefac0c7187a4

Observation b566abb0-1aa6-4453-b6c6-97ae51c03c6d · outbound

This paper cites Prior constraints-based reward model training for aligning large language models , year =.

RRC: Unlocking Generative Reward Models in LLM Reinforcement Learning via Ranking-Based Reward Construction Prior constraints-based reward model training for aligning large language models , year =

Reference 42

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:49:52.171582Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=arxiv_source observed=2026-08-07T05:49:48.548703Z digest=sha256:905c394446cb73e9376d618a1cc4aaf0acaaf8fd5acdbd8345299160185e7d58

Observation de7f27d8-4851-4122-94fc-73e36ae6b8bd · outbound

This paper cites Improving In-Context Learning via Sequentially Selection and Preference Alignment for Few-Shot Aspect-Based Sentiment Analysis , year =.

RRC: Unlocking Generative Reward Models in LLM Reinforcement Learning via Ranking-Based Reward Construction Improving In-Context Learning via Sequentially Selection and Preference Alignment for Few-Shot Aspect-Based Sentiment Analysis , year =

Reference 43

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:49:52.156190Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=arxiv_source observed=2026-08-07T05:49:48.602152Z digest=sha256:e95890c363b87ae1af35ac6f15dc0fecdaf78d3717c8040429c86e1af1628a57

Observation 119bb83d-96b8-4f36-a699-fe18ba729caa · outbound

This paper cites Qwen-audio: Advancing universal audio understanding via unified large-scale audio-language models , year =.

RRC: Unlocking Generative Reward Models in LLM Reinforcement Learning via Ranking-Based Reward Construction Qwen-audio: Advancing universal audio understanding via unified large-scale audio-language models , year =

Reference 44

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:49:52.141183Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=arxiv_source observed=2026-08-07T05:49:48.696057Z digest=sha256:6fd255cb6fe79430d9365e1a675fcc9c0881594416d04b6e3bd0e7e8793e6c7d

Observation 38c72927-3206-40a4-98de-0a66f7aa5a28 · outbound

This paper cites Manning and Stefano Ermon and Chelsea Finn , booktitle =.

RRC: Unlocking Generative Reward Models in LLM Reinforcement Learning via Ranking-Based Reward Construction Manning and Stefano Ermon and Chelsea Finn , booktitle =

Reference 45

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:49:52.126244Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=arxiv_source observed=2026-08-07T05:49:48.746417Z digest=sha256:37bfe32ff4d6fe71afc9e8bc7ee9aece5ee7652c4cf11a8fb7c3f695c174e7fd

Observation 5f6a0495-ddc0-49fa-8cfd-e859a86ded4f · outbound

This paper cites Discriminative Reranking for Neural Machine Translation , year =.

RRC: Unlocking Generative Reward Models in LLM Reinforcement Learning via Ranking-Based Reward Construction Discriminative Reranking for Neural Machine Translation , year =

Reference 46

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:49:52.111466Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=arxiv_source observed=2026-08-07T05:49:48.800135Z digest=sha256:9dee823d7b9054a3760f52ea53a9835d7c97970b032ed7fbc7837669ea757153

Observation e4bcb13c-a52a-4763-809d-d513742ce78b · outbound

This paper cites Dapo: An open-source llm reinforcement learning system at scale , year =.

RRC: Unlocking Generative Reward Models in LLM Reinforcement Learning via Ranking-Based Reward Construction Dapo: An open-source llm reinforcement learning system at scale , year =

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-07T05:49:48.894739Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T05:49:48.894739Z digest=sha256:df1a715f3455b0378427db7376ff6784c0537b78612ad6b4542431f8ee1d25cd

Observation 5b28d81b-3420-4344-b60c-6f5b0b42fc86 · outbound

This paper cites Deepseekmath: Pushing the limits of mathematical reasoning in open language models , year =.

RRC: Unlocking Generative Reward Models in LLM Reinforcement Learning via Ranking-Based Reward Construction Deepseekmath: Pushing the limits of mathematical reasoning in open language models , year =

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-07T05:49:48.962879Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T05:49:48.962879Z digest=sha256:197091f29f020ba524bb8bdc85cc0a0effb99553f17972db9c6272505d821d0b

Observation 057e7e02-d2c4-4063-a2e4-9bb69bf4c57a · outbound

This paper cites Generative reward modeling via synthetic criteria preference learning , year =.

RRC: Unlocking Generative Reward Models in LLM Reinforcement Learning via Ranking-Based Reward Construction Generative reward modeling via synthetic criteria preference learning , year =

Reference 49

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:49:52.077498Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=arxiv_source observed=2026-08-07T05:49:49.006947Z digest=sha256:c00e16517a19d15b8e5e42bb774f64f87c47c8be4031ea8409ff5d89eb295344

Observation 6efafa89-5f2b-4bb9-9ea4-e4d74a342fb2 · outbound

This paper cites Unified multimodal chain-of-thought reward model through reinforcement fine-tuning , year =.

RRC: Unlocking Generative Reward Models in LLM Reinforcement Learning via Ranking-Based Reward Construction Unified multimodal chain-of-thought reward model through reinforcement fine-tuning , year =

Reference 50

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:49:52.062854Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=arxiv_source observed=2026-08-07T05:49:49.013516Z digest=sha256:b4327b10db8f18381dc97c431f65783790a4752a925b061d88906f75191d3940

Observation d03689ab-3acf-429a-b7ab-fc5322fbf02f · outbound

This paper cites Rm-r1: Reward modeling as reasoning , year =.

RRC: Unlocking Generative Reward Models in LLM Reinforcement Learning via Ranking-Based Reward Construction Rm-r1: Reward modeling as reasoning , year =

Reference 51

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:49:52.048217Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=arxiv_source observed=2026-08-07T05:49:49.112945Z digest=sha256:45e79df9a2d34535b36a38afbee1424e68a09a4b02817861a9798137f091c663

Observation cce62ad7-7a71-4235-b790-53db65196ea3 · outbound

This paper cites Reward reasoning model , year =.

RRC: Unlocking Generative Reward Models in LLM Reinforcement Learning via Ranking-Based Reward Construction Reward reasoning model , year =

Reference 52

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:49:52.033753Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=arxiv_source observed=2026-08-07T05:49:49.193651Z digest=sha256:b5554d1a24a8dac85e8e0a3d42a7af1a4e224a7e62bd39973e8d2abb4e273f93

Observation 2b9f8ed9-f566-4459-9c84-dc3dbcae46fc · outbound

This paper cites GRAM-R ^2 : Self-Training Generative Foundation Reward Models for Reward Reasoning , year =.

RRC: Unlocking Generative Reward Models in LLM Reinforcement Learning via Ranking-Based Reward Construction GRAM-R ^2 : Self-Training Generative Foundation Reward Models for Reward Reasoning , year =

Reference 53

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:49:52.019979Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=arxiv_source observed=2026-08-07T05:49:49.259871Z digest=sha256:8edc7263cd6de64966f61519593f8adc653c12d4ad1a2152511b5fdcf6a61060

Observation fe4d15e3-bc62-408c-851f-2489c4f00019 · outbound

This paper cites GRAM: A Generative Foundation Reward Model for Reward Generalization , year =.

RRC: Unlocking Generative Reward Models in LLM Reinforcement Learning via Ranking-Based Reward Construction GRAM: A Generative Foundation Reward Model for Reward Generalization , year =

Reference 54

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:49:52.003423Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=arxiv_source observed=2026-08-07T05:49:49.377175Z digest=sha256:a7cd26e40dabd0abe01df9c01036d51b9100104dd992d9c817b267c47258e070

Observation fc0e704b-c1bf-4185-b0f1-5273ea7b7dab · outbound

This paper cites Inference-time scaling for generalist reward modeling , year =.

RRC: Unlocking Generative Reward Models in LLM Reinforcement Learning via Ranking-Based Reward Construction Inference-time scaling for generalist reward modeling , year =

Reference 55

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:49:51.988506Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=arxiv_source observed=2026-08-07T05:49:49.520356Z digest=sha256:343b88042a5d403ecc9c49f3499e80d6951c22bf8a5eb003848b07a3f9f6a48f

Observation 5a571064-f449-4e59-bd29-3d741065ed9a · outbound

This paper cites Reward Model Ensembles Help Mitigate Overoptimization , year =.

RRC: Unlocking Generative Reward Models in LLM Reinforcement Learning via Ranking-Based Reward Construction Reward Model Ensembles Help Mitigate Overoptimization , year =

Reference 56

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:49:51.971964Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=arxiv_source observed=2026-08-07T05:49:49.726109Z digest=sha256:b5ee0db28bcbb6699fadf422b1afa86d5b2536be517a8806e1f3a6020458c55c

Observation 8b55f41c-0c8f-4224-a776-77118909d455 · outbound

This paper cites Skywork-Reward-V2: Scaling Preference Data Curation via Human-AI Synergy , year =.

RRC: Unlocking Generative Reward Models in LLM Reinforcement Learning via Ranking-Based Reward Construction Skywork-Reward-V2: Scaling Preference Data Curation via Human-AI Synergy , year =

Reference 57

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:49:51.956432Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=arxiv_source observed=2026-08-07T05:49:49.855495Z digest=sha256:5b64937d9b271108aad2cc865f6075284792d001280841bce121b3de30eef5d6

Observation 0c1242af-9be7-489e-b125-e1785b84d53a · outbound

This paper cites Rovrm: A robust visual reward model optimized via auxiliary textual preference data , year =.

RRC: Unlocking Generative Reward Models in LLM Reinforcement Learning via Ranking-Based Reward Construction Rovrm: A robust visual reward model optimized via auxiliary textual preference data , year =

Reference 58

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:49:51.942367Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=arxiv_source observed=2026-08-07T05:49:49.961442Z digest=sha256:65901c3f8e53559e42e15e73241effb5c066ea164872f8b383b2cd8548350e39

Observation b1d32292-7cd2-4a90-bbde-433e37e585c5 · outbound

This paper cites Specialist or Generalist? Instruction Tuning for Specific.

RRC: Unlocking Generative Reward Models in LLM Reinforcement Learning via Ranking-Based Reward Construction Specialist or Generalist? Instruction Tuning for Specific

Reference 59

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:49:51.926107Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=arxiv_source observed=2026-08-07T05:49:50.136217Z digest=sha256:4586a57cacd69bacdcdf148dbf8f12193d7dcbae55cab5f5c9980d727b4baa44

Observation ea9c6137-14cb-4ddb-8dad-99dcf3f0622d · outbound

This paper cites Unveiling the Generalization Power of Fine-Tuned Large Language Models , year =.

RRC: Unlocking Generative Reward Models in LLM Reinforcement Learning via Ranking-Based Reward Construction Unveiling the Generalization Power of Fine-Tuned Large Language Models , year =

Reference 60

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:49:51.911106Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=arxiv_source observed=2026-08-07T05:49:50.196702Z digest=sha256:c4740010dc9c37bced0c0cf4eb15f562dd3e66726f6ca693a7bd03d78fa8465b

Observation aef914da-2aac-4694-baa9-bfca318d3932 · outbound

This paper cites Deepseek-r1: Incentivizing reasoning capability in llms via reinforcement learning , year =.

RRC: Unlocking Generative Reward Models in LLM Reinforcement Learning via Ranking-Based Reward Construction Deepseek-r1: Incentivizing reasoning capability in llms via reinforcement learning , year =

Reference 61

Resolution
unresolved
no resolver link, observed 2026-08-07T05:49:50.258977Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T05:49:50.258977Z digest=sha256:f63f79a881ca47b10f3c59fe715ac13f95229a3160bff352dac4d1602e9919f8

Observation 823d99c6-f758-4b11-a3bd-c88f0f608d21 · outbound

This paper cites Chi and Quoc V.

RRC: Unlocking Generative Reward Models in LLM Reinforcement Learning via Ranking-Based Reward Construction Chi and Quoc V

Reference 62

Resolution
unresolved
no resolver link, observed 2026-08-07T05:49:50.354120Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T05:49:50.354120Z digest=sha256:f8d8ab5dec317e1174d08c7dbc36532ce20e22d3e318b59bd65de0e64fb28674

Observation bf973903-1f16-46b4-ac31-9a8109a64c7e · outbound

This paper cites Scaling instruction-finetuned language models , year =.

RRC: Unlocking Generative Reward Models in LLM Reinforcement Learning via Ranking-Based Reward Construction Scaling instruction-finetuned language models , year =

Reference 63

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:49:51.873938Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=arxiv_source observed=2026-08-07T05:49:50.450939Z digest=sha256:fb817ff8b0b3c6b3816058504994cd97a7df2f34e35edb827d22d3716b4abece

Observation 22f12c7c-1afd-4d0a-bd3a-4081930d89c9 · outbound

This paper cites ArXiv preprint , title =.

RRC: Unlocking Generative Reward Models in LLM Reinforcement Learning via Ranking-Based Reward Construction ArXiv preprint , title =

Reference 64

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:49:51.856887Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=arxiv_source observed=2026-08-07T05:49:50.550071Z digest=sha256:f33b8f548c718f69e22c5b035e3aa1421dd695b11af72fa8335b8b3f6168311b

Observation bc442f6e-970a-4c0d-ac49-e85517058b0d · outbound

This paper cites Generative verifiers: Reward modeling as next-token prediction , year =.

RRC: Unlocking Generative Reward Models in LLM Reinforcement Learning via Ranking-Based Reward Construction Generative verifiers: Reward modeling as next-token prediction , year =

Reference 65

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:49:51.841712Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=arxiv_source observed=2026-08-07T05:49:50.632006Z digest=sha256:c842c02f5895d90ed69a1a0f358fa46bde3c1796312fa6ed06a26bee2f34c4b9

Observation 9680b05a-c1c1-4367-81c7-ec56d5854582 · outbound

This paper cites Foundations of large language models , year =.

RRC: Unlocking Generative Reward Models in LLM Reinforcement Learning via Ranking-Based Reward Construction Foundations of large language models , year =

Reference 66

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:49:51.825586Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=arxiv_source observed=2026-08-07T05:49:50.787906Z digest=sha256:0e48822c625b08e29a8cd029e363d0db850ffa9cb49cff73f2c18cf9de78803e

Observation d3e6dc32-ce32-44c9-ae07-f5a2d468a4ab · outbound

This paper cites Step-level verifier-guided hybrid test-time scaling for large language models , year =.

RRC: Unlocking Generative Reward Models in LLM Reinforcement Learning via Ranking-Based Reward Construction Step-level verifier-guided hybrid test-time scaling for large language models , year =

Reference 67

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:49:51.808971Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=arxiv_source observed=2026-08-07T05:49:50.928739Z digest=sha256:898840d2b68ffc8daeac1e9c4399b1d10e5f8fdd08c8364032634ef60804c457

Observation f1499bb0-3728-49e7-ba81-b386ef8fedb2 · outbound

This paper cites s1: Simple test-time scaling , year =.

RRC: Unlocking Generative Reward Models in LLM Reinforcement Learning via Ranking-Based Reward Construction s1: Simple test-time scaling , year =

Reference 68

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:49:51.791856Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=arxiv_source observed=2026-08-07T05:49:51.061742Z digest=sha256:d7f5b6dd9185a08b67ac68f44e9dfe1f87e7809fdb649eabbe6bb7bc9e71eec1

Observation 6962bda8-8b9a-482d-885d-ab4b684d4b79 · outbound

This paper cites Ziegler and Ryan Lowe and Chelsea Voss and Alec Radford and Dario Amodei and Paul F.

RRC: Unlocking Generative Reward Models in LLM Reinforcement Learning via Ranking-Based Reward Construction Ziegler and Ryan Lowe and Chelsea Voss and Alec Radford and Dario Amodei and Paul F

Reference 69

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:49:51.773679Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=arxiv_source observed=2026-08-07T05:49:51.131325Z digest=sha256:bb5b95f88f42a46d48694e758fc749bf184d17481f5ca664e032fb2564a8d38f

Observation 2e27ccf8-bc49-44e0-8a17-e6ad72a664fb · outbound

This paper cites Christiano and Jan Leike and Tom B.

RRC: Unlocking Generative Reward Models in LLM Reinforcement Learning via Ranking-Based Reward Construction Christiano and Jan Leike and Tom B

Reference 70

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:49:51.757551Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=arxiv_source observed=2026-08-07T05:49:51.254831Z digest=sha256:73abf57b2182eb08f35c0eeeaaef6bcad3834be4fdeebfd52044c228d5cb46a9

Observation 8475e862-f0f7-479e-9626-806ee73707bc · outbound

This paper cites Pku-saferlhf: Towards multi-level safety alignment for llms with human preference , year =.

RRC: Unlocking Generative Reward Models in LLM Reinforcement Learning via Ranking-Based Reward Construction Pku-saferlhf: Towards multi-level safety alignment for llms with human preference , year =

Reference 71

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:49:51.739933Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=arxiv_source observed=2026-08-07T05:49:51.270033Z digest=sha256:3481b5e475c49e13c123eedb2eb31029add431d64907a9e4489e6abf73c2439b

Observation c4d27562-bd7a-497e-9384-c35c7d3df281 · outbound

This paper cites Training a helpful and harmless assistant with reinforcement learning from human feedback , year =.

RRC: Unlocking Generative Reward Models in LLM Reinforcement Learning via Ranking-Based Reward Construction Training a helpful and harmless assistant with reinforcement learning from human feedback , year =

Reference 72

Resolution
unresolved
no resolver link, observed 2026-08-07T05:49:51.396419Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T05:49:51.396419Z digest=sha256:f353746fa57fb2000fbfab15468ea51f451dccae5250218d89ac581cfe7f4932

Observation 6a7abe3c-2669-44aa-a261-06f63631e3f5 · outbound

This paper cites Hybrid alignment training for large language models , year =.

RRC: Unlocking Generative Reward Models in LLM Reinforcement Learning via Ranking-Based Reward Construction Hybrid alignment training for large language models , year =

Reference 73

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:49:51.714830Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=arxiv_source observed=2026-08-07T05:49:51.461846Z digest=sha256:52ae17b27977c116193d6afda4ca93f5278a8932f59a2e07716054c51f2a067d

Observation 42a3a541-e686-4a72-8b1d-e169d7cf4070 · outbound

This paper cites an unresolved cited work.

RRC: Unlocking Generative Reward Models in LLM Reinforcement Learning via Ranking-Based Reward Construction Unresolved cited work

Reference 74

Resolution
unresolved
raw_fallback, observed 2026-08-07T05:49:51.698270Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=arxiv_source observed=2026-08-07T05:49:51.540240Z digest=sha256:909be107f61a57a5d50a03e91425d298b90f3abf6cf88d33777e566ce53e8f23

Observation 4fe81ff3-0c90-4e39-99a0-995292c5cf9f · outbound

This paper cites Scaling Learning Algorithms Towards.

RRC: Unlocking Generative Reward Models in LLM Reinforcement Learning via Ranking-Based Reward Construction Scaling Learning Algorithms Towards

Reference 75

Resolution
unresolved
no resolver link, observed 2026-08-07T05:49:51.589012Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T05:49:51.589012Z digest=sha256:603f5221d66987783c5344cdce8d6cc4256e575f2bc9937a310f32674feb5f20

Observation 85bf7a5a-2af4-417c-ad49-306fdf7e79fd · outbound

This paper cites and Osindero, Simon and Teh, Yee Whye , journal =.

RRC: Unlocking Generative Reward Models in LLM Reinforcement Learning via Ranking-Based Reward Construction and Osindero, Simon and Teh, Yee Whye , journal =

Reference 76

Resolution
unresolved
no resolver link, observed 2026-08-07T05:49:51.593444Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T05:49:51.593444Z digest=sha256:6716bc91799bf7ed228c0950854dc7625cf763e2c66712508b92972669867834

Observation 88a5e574-b9ad-4d50-83de-e772ac4b4c88 · outbound

This paper cites 2016 , publisher=.

RRC: Unlocking Generative Reward Models in LLM Reinforcement Learning via Ranking-Based Reward Construction 2016 , publisher=

Reference 77

Resolution
unresolved
no resolver link, observed 2026-08-07T05:49:51.597998Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T05:49:51.597998Z digest=sha256:cfb501059f3c26e92f5a323302eb38d2f09dff21095942d6da7f8a3c550567f8

Pith citing papers

No inbound Pith citation observations are available.