Pith. sign in

Paper Citation Record · LEDGER

Reward Reasoning Model

As of 18 August 2026, this Paper Citation Record lists 87 of 87 outbound references and 15 inbound Pith citation observations for arXiv:2505.14674.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2505.14674 v1

Coverage vector

measured 87 of 87 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-07T15:34:53.874909Z

measured 102 of 102 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-18T06:34:40.430872+00:00

measured 15 of 15 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-07T12:15:37.528799Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-04T06:49:38.201535Z

Reference resolution

87 of 87 outbound references displayed

  • verified exact1
  • verified fuzzy18
  • unresolved68
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 2b5dc76e-1feb-4983-8214-fa8e19d1cf8a · outbound

This paper cites GPT-4 Technical Report.

Reward Reasoning Model GPT-4 Technical Report

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-07T15:34:50.218433Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:34:50.218433Z digest=sha256:d76a12f3ec9668463b2c103eb6fe4355a6557f5a6f374c55cd76e195ba2d983f

Observation ef46c4eb-d007-43dd-bcac-af4fe3fbe2f0 · outbound

This paper cites Big-Math: A Large-Scale, High-Quality Math Dataset for Reinforcement Learning in Language Models.

Reward Reasoning Model Big-Math: A Large-Scale, High-Quality Math Dataset for Reinforcement Learning in Language Models

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-07T15:34:50.264050Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:34:50.264050Z digest=sha256:f6e36ab426a0b793dfc331768db71eff2ce8b68bda3b3c396999c214ba5cceca

Observation 10d6e181-a23c-4326-a1ea-b681e34819c8 · outbound

This paper cites Atla Selene Mini: A General Purpose Evaluation Model.

Reward Reasoning Model Atla Selene Mini: A General Purpose Evaluation Model

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-07T15:34:50.298845Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:34:50.298845Z digest=sha256:ca135a593fe888bc02ae5c553ef6e469b85d6687b2007e37434fd55dc6b612cc

Observation 324dd137-dbb8-4b26-ad1d-bfb5723149e7 · outbound

This paper cites The claude 3 model family: Opus, sonnet, haiku.Claude-3 Model Card, 1:1, 2024.

Reward Reasoning Model The claude 3 model family: Opus, sonnet, haiku.Claude-3 Model Card, 1:1, 2024

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-07T15:34:50.335857Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:34:50.335857Z digest=sha256:e9e5362cb78948e9805472bd9bbf23addfc7a8cb0d7a61f48da8071b46f9f340

Observation abb5fcfe-f89d-4712-9a0a-c907f47bcfec · outbound

This paper cites an unresolved cited work.

Reward Reasoning Model Unresolved cited work

Reference 5

Resolution
verified exact
doi, observed 2026-08-07T15:34:54.013847Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T15:34:50.375780Z digest=sha256:e944d26d6facc2fc7b86c818928d222b415a740994b188a0a5adbf1e2947c9d3

Observation 47b45be1-fd2f-4c65-b61d-eb409a76aef7 · outbound

This paper cites Training a Helpful and Harmless Assistant with Reinforcement Learning from Human Feedback.

Reward Reasoning Model Training a Helpful and Harmless Assistant with Reinforcement Learning from Human Feedback

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-07T15:34:50.406624Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:34:50.406624Z digest=sha256:2cf4f31ef8331b8a359998bb7109d7196c778ec261e837bb91129291520fa2fa

Observation f6cb77be-ec9f-4514-991a-2a89a49521dc · outbound

This paper cites Constitutional AI: Harmlessness from AI Feedback.

Reward Reasoning Model Constitutional AI: Harmlessness from AI Feedback

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-07T15:34:50.446245Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:34:50.446245Z digest=sha256:604232387995e4a289907b49365ff60c23c2649e8c6cda1bf5fd3406d02453c9

Observation c9e244ad-e186-419b-a6b6-3427b84bc06e · outbound

This paper cites Large Language Monkeys: Scaling Inference Compute with Repeated Sampling.

Reward Reasoning Model Large Language Monkeys: Scaling Inference Compute with Repeated Sampling

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-07T15:34:50.491218Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:34:50.491218Z digest=sha256:38165484833a31c71165dc3bf586941c56c1a417dadda5058ae2b611e4296aee

Observation 7d2dc476-2b95-4d48-8f9e-b5fce7fb796a · outbound

This paper cites Language models are few-shot learners.Advances in neural information processing systems, 33:1877–1901, 2020.

Reward Reasoning Model Language models are few-shot learners.Advances in neural information processing systems, 33:1877–1901, 2020

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-07T15:34:50.535585Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:34:50.535585Z digest=sha256:fe553607e87e9a21f7dd5f391d4fcbdc2055a60349f732673deffefff6ca9c85

Observation b2be30d7-3095-4200-b422-7a1795082215 · outbound

This paper cites Sparks of artificial general intelligence: Early experiments with GPT-4, 2023.

Reward Reasoning Model Sparks of artificial general intelligence: Early experiments with GPT-4, 2023

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-07T15:34:50.592474Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:34:50.592474Z digest=sha256:78830afdb02501b8fe1222356c745d1680fdee3964dc16727dd97b8afffdaca9

Observation f46c7d44-57de-479f-95fe-2c2c9e4dd482 · outbound

This paper cites CompassJudger-1: All-in-one Judge Model Helps Model Evaluation and Evolution.

Reward Reasoning Model CompassJudger-1: All-in-one Judge Model Helps Model Evaluation and Evolution

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-07T15:34:50.638608Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:34:50.638608Z digest=sha256:dd5b19d6924504d017cd67b4f7111d94069888a1debc78853d04df07f4e70bed

Observation ecaf28c3-59e1-4e5f-9512-4f3d20541476 · outbound

This paper cites Judgelrm: Large reasoning models as a judge.arXiv preprint arXiv:2504.00050, 2025.

Reward Reasoning Model Judgelrm: Large reasoning models as a judge.arXiv preprint arXiv:2504.00050, 2025

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-07T15:34:50.684754Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:34:50.684754Z digest=sha256:3e98e85dbff90a3f71e4bbad865fb4b0b4a701f899519b2570d2d69c8f210529

Observation a6981fcf-1d7b-42ff-b33b-3da231f42faf · outbound

This paper cites Seal: Steerable reasoning calibration of large language models for free.

Reward Reasoning Model Seal: Steerable reasoning calibration of large language models for free

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-07T15:34:50.725027Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:34:50.725027Z digest=sha256:a877a9c140a76996e6a02f7241238ea7015e9c2d51e37e00a54f6397f151c726

Observation 89dbad2b-12a3-4097-97aa-2fe0d4550f18 · outbound

This paper cites Rm-r1: Reward modeling as reasoning.arXiv preprint arXiv:2505.02387, 2025.

Reward Reasoning Model Rm-r1: Reward modeling as reasoning.arXiv preprint arXiv:2505.02387, 2025

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-07T15:34:50.777215Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:34:50.777215Z digest=sha256:60b5c13173025b526f9096263e62746813e68b4406b9734dedb612cdb02ccf3f

Observation 4325b17c-a6d6-41a3-89e8-0f630abe60e5 · outbound

This paper cites Christiano, Jan Leike, Tom B.

Reward Reasoning Model Christiano, Jan Leike, Tom B

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-07T15:34:50.819870Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:34:50.819870Z digest=sha256:79d118144f8ac38c4dab1113c4feb28d3ded32bddd6f89288d943e50e34bb4fe

Observation 1c864530-7a15-455c-a7f2-e29d36148472 · outbound

This paper cites Training Verifiers to Solve Math Word Problems.

Reward Reasoning Model Training Verifiers to Solve Math Word Problems

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-07T15:34:50.852224Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:34:50.852224Z digest=sha256:ff5bbb6c45e96992d11c3f0d274bb2f3dc85bf17bce066590f2fea707bc227a4

Observation 99b94b53-9660-4767-8b08-eb18dde2642a · outbound

This paper cites Elo.The Rating of Chessplayers, Past and Present.

Reward Reasoning Model Elo.The Rating of Chessplayers, Past and Present

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-07T15:34:50.904944Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:34:50.904944Z digest=sha256:cd12f7e7d38835f3495c22eb8538bbf1badfd16b32d4184703076e9490b1d3f8

Observation 7e5287c2-99f5-44e1-9322-1a6eb4ea874c · outbound

This paper cites Gonzalez, and Ion Stoica.

Reward Reasoning Model Gonzalez, and Ion Stoica

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-07T15:34:50.948093Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:34:50.948093Z digest=sha256:be5a7c24cce777561fe4e8cc40c7df3e1de53d5e95e535665776af1500c0a420

Observation 5f88e40d-533c-4b66-bd39-e43d86ce9f77 · outbound

This paper cites On Designing Effective RL Reward at Training Time for LLM Reasoning.

Reward Reasoning Model On Designing Effective RL Reward at Training Time for LLM Reasoning

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-07T15:34:50.993826Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:34:50.993826Z digest=sha256:743008cea63b0f5f57cd1df31c27ae25256153a9fea7c3b21ba47fea0872cd49

Observation 033762ca-1f77-4f47-a0ad-dbf9d7e01395 · outbound

This paper cites Scaling laws for reward model overoptimization.

Reward Reasoning Model Scaling laws for reward model overoptimization

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-07T15:34:51.045534Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:34:51.045534Z digest=sha256:3d55d9f679e2d39298a6111220349664513a369d5168065ac93340ca72522c82

Observation ca1e8f79-71e1-4408-afda-e93bfc3593fb · outbound

This paper cites A Survey on LLM-as-a-Judge.

Reward Reasoning Model A Survey on LLM-as-a-Judge

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-07T15:34:51.098716Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:34:51.098716Z digest=sha256:0960d27c6035a06513c352856d8237d10729fd4fefcb6863c4c7f647f344af5c

Observation 9251b517-e7ab-4e26-8341-e109762f8faf · outbound

This paper cites DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning.

Reward Reasoning Model DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-07T15:34:51.142179Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:34:51.142179Z digest=sha256:1a38ad66984176b17dc89dea0ea15cdcf3e231ce6a8f5538ca1adf13de9a67bf

Observation f8aff6e2-39e4-47fe-9c9e-541080cb20fd · outbound

This paper cites Mcranker: Generating diverse criteria on-the-fly to improve pointwise llm rankers.

Reward Reasoning Model Mcranker: Generating diverse criteria on-the-fly to improve pointwise llm rankers

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-07T15:34:51.178507Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:34:51.178507Z digest=sha256:f0ade4cf41d1dc7437c7bbc7c8fae06652307d1c869f0e6013c47d636dfe7f0f

Observation 8b2953dc-7c4b-4ecd-a612-2d49d768aa75 · outbound

This paper cites Skywork open reasoner series.

Reward Reasoning Model Skywork open reasoner series

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-07T15:34:51.224822Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:34:51.224822Z digest=sha256:48c336aef1889a2d53da984bdb1ee8cc44a4852aecf6efdbbd3e36b191325952

Observation e6099a2b-3bf3-4ae4-aec6-0aec91939039 · outbound

This paper cites Measuring mathematical problem solving with the MATH dataset.

Reward Reasoning Model Measuring mathematical problem solving with the MATH dataset

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:34:55.958482Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T15:34:51.259139Z digest=sha256:e2eec498eb8f2d00dfda32b7e141f2a09803fc5a9ec6d6452e9f214ba797d6b1

Observation f12a8b94-3cf5-4ae0-851e-f9a5556d8a5d · outbound

This paper cites GPT-4o System Card.

Reward Reasoning Model GPT-4o System Card

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-07T15:34:51.300610Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:34:51.300610Z digest=sha256:4b3802199c61fe67834155f427578f57003c68f5c5c774dde4f4f0469cc1a358

Observation 31bd0b84-c20e-4a88-b147-8b2dd8f52a19 · outbound

This paper cites OpenAI o1 System Card.

Reward Reasoning Model OpenAI o1 System Card

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-07T15:34:51.337869Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:34:51.337869Z digest=sha256:afde70bed2fdfbf6680a22660b5ae1f20d38421f6271860f6f51f6d818261e23

Observation 536d8b20-32fd-46c3-a843-3794f23bc031 · outbound

This paper cites LLM-blender: Ensembling large language models with pairwise ranking and generative fusion.

Reward Reasoning Model LLM-blender: Ensembling large language models with pairwise ranking and generative fusion

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-07T15:34:51.370094Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:34:51.370094Z digest=sha256:217ef2b3ff25ac0567581554e85c25ca652fa01d026dc0e0efbe53dca5d9e31a

Observation cd8e1793-3b22-41ac-acd8-1db43cb79403 · outbound

This paper cites SWE-bench: Can language models resolve real-world github issues? InThe Twelfth International Conference on Learning Representations, 2024.

Reward Reasoning Model SWE-bench: Can language models resolve real-world github issues? InThe Twelfth International Conference on Learning Representations, 2024

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-07T15:34:51.410454Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:34:51.410454Z digest=sha256:f6da871e850116fca2f3d767dda69a82c0bd6e5517a1abb4534444914b9f0860

Observation 825d3a07-6a2b-4c64-aec2-07b755b4daa0 · outbound

This paper cites Scaling Laws for Neural Language Models.

Reward Reasoning Model Scaling Laws for Neural Language Models

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-07T15:34:51.450665Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:34:51.450665Z digest=sha256:135b3586aa25f3cdbe0f4bf43ac8d0dfad0d2ae7e56ddc6fae64c09be22ed942

Observation 8afb08ed-a533-41f1-a60d-44f323f602d8 · outbound

This paper cites Maple: Multi-modal prompt learning.

Reward Reasoning Model Maple: Multi-modal prompt learning

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-07T15:34:51.493581Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:34:51.493581Z digest=sha256:c13b3b4659988244d1af259c3bb3cf3287695e2a7b7ac7228ba1ee772bfe1e0a

Observation 95386e55-d1bd-4352-9799-9053ebb635a1 · outbound

This paper cites Prometheus 2: An open source language model specialized in evaluating other language models.

Reward Reasoning Model Prometheus 2: An open source language model specialized in evaluating other language models

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-07T15:34:51.527785Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:34:51.527785Z digest=sha256:5c3b2b14f7c48db6a3be1c88792052754898318eb63115eb0f9098725593ad3f

Observation ce635670-f63c-447b-afd0-6a86ceaab1fe · outbound

This paper cites Tulu 3: Pushing Frontiers in Open Language Model Post-Training.

Reward Reasoning Model Tulu 3: Pushing Frontiers in Open Language Model Post-Training

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-07T15:34:51.610704Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:34:51.610704Z digest=sha256:af732a9f28f0d689ea805163fe5930a56ec7344f2a6e33dbe3301c0cd2213408

Observation 1ce3f9a5-3861-4ee6-b727-1c15d0e0512d · outbound

This paper cites RewardBench: Evaluating Reward Models for Language Modeling.

Reward Reasoning Model RewardBench: Evaluating Reward Models for Language Modeling

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-07T15:34:51.641108Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:34:51.641108Z digest=sha256:2c64115e27d61a0067dc4fc2f27f6ca4c7507d24a44e8412f4a648f0555b0666

Observation d59bc306-acfb-42b5-9cef-d49b096d8c08 · outbound

This paper cites Generative judge for evaluating alignment.

Reward Reasoning Model Generative judge for evaluating alignment

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:34:55.936539Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T15:34:51.675833Z digest=sha256:5fc0d5bc3e079e5202b6dc5e302e26aec9f487e091451a58efefb0e6e932917b

Observation 97e93a02-a458-4cf8-bb3d-1fa1e9b98994 · outbound

This paper cites From Crowdsourced Data to High-Quality Benchmarks: Arena-Hard and BenchBuilder Pipeline.

Reward Reasoning Model From Crowdsourced Data to High-Quality Benchmarks: Arena-Hard and BenchBuilder Pipeline

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-07T15:34:51.745476Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:34:51.745476Z digest=sha256:61524741c5652545b2c21fbab53f1be9ee991e00f694518a668755339ae4b01b

Observation d0dc8f84-b9f5-42c8-8c22-50f1fcd63e53 · outbound

This paper cites an unresolved cited work.

Reward Reasoning Model Unresolved cited work

Reference 37

Resolution
unresolved
raw_fallback, observed 2026-08-07T15:34:55.926236Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T15:34:51.708821Z digest=sha256:33969cba601e6e1304bf2edebe1d969857b1cd96cc37d1bbca6bf5c3c822bba3

Observation 17a41f16-cc95-4c35-847c-587a961c0c0d · outbound

This paper cites Rouge: A package for automatic evaluation of summaries.

Reward Reasoning Model Rouge: A package for automatic evaluation of summaries

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-07T15:34:51.842603Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:34:51.842603Z digest=sha256:326bb4a520160118c2782e9b8e6b80112c588f2600bc2143cd08bd10a33fa670

Observation 6ef38318-1a13-4659-b948-b3841db37b33 · outbound

This paper cites Let’s verify step by step.

Reward Reasoning Model Let’s verify step by step

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-07T15:34:51.791581Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:34:51.791581Z digest=sha256:2591b3d3be9708633dfa23a7ceb864aee33be3cdcf9ca496a188fd8e6dddd24c

Observation 10f0424d-d52f-4e3b-bd01-22350bb450f2 · outbound

This paper cites PairJudge RM: Perform Best-of-N Sampling with Knockout Tournament.

Reward Reasoning Model PairJudge RM: Perform Best-of-N Sampling with Knockout Tournament

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-07T15:34:51.960784Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:34:51.960784Z digest=sha256:241466d1688af54df248e97d2256eac774b745a5d7d6792ab7ecd91f10a1eb3e

Observation 1944545c-20ca-4cd3-a1b0-57e105fd9972 · outbound

This paper cites Skywork-Reward: Bag of Tricks for Reward Modeling in LLMs.

Reward Reasoning Model Skywork-Reward: Bag of Tricks for Reward Modeling in LLMs

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-07T15:34:51.910884Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:34:51.910884Z digest=sha256:b0aad5f7183832598124a4c229d70afb7ca4970b7a32f2a9e0e78ef95557fcc7

Observation 0f772cdb-2858-463c-9082-4001e982e021 · outbound

This paper cites General-reasoner: Advancing llm reasoning across all domains.

Reward Reasoning Model General-reasoner: Advancing llm reasoning across all domains

Reference 42

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:34:55.902751Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T15:34:52.052327Z digest=sha256:982d6bbf60743ba044deb293ebaa6780e055aebcaf23c25ac29447198f182bd5

Observation fbbf5d0d-d666-4fea-beba-18ae91c32152 · outbound

This paper cites Inference-time scaling for generalist reward modeling.arXiv preprint arXiv:2504.02495, 2025.

Reward Reasoning Model Inference-time scaling for generalist reward modeling.arXiv preprint arXiv:2504.02495, 2025

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-07T15:34:51.997784Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:34:51.997784Z digest=sha256:978a179f383555972fbf770945e11cb0e4e615bc24dc7a0a80d9fee63f5de658

Observation ca842ffc-23a6-46a5-a14a-e0cee2affeb2 · outbound

This paper cites Training language models to follow instructions with human feedback.

Reward Reasoning Model Training language models to follow instructions with human feedback

Reference 44

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:34:55.891726Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T15:34:52.134440Z digest=sha256:a1090e85701e614a43fa9a3ba575c2e93b1f14ad6db3892f6fb545eeba544632

Observation 04e333ca-9936-4e0a-b645-56707c172df2 · outbound

This paper cites Generative Reward Models.

Reward Reasoning Model Generative Reward Models

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-07T15:34:52.079730Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:34:52.079730Z digest=sha256:7ac2e1ce72078d50b7537a8e3161e6472ca5587f551de496e59f6ca82e612e63

Observation e156917e-cdb5-489a-9e9f-32173197839d · outbound

This paper cites Bleu: a method for automatic evaluation of machine translation.

Reward Reasoning Model Bleu: a method for automatic evaluation of machine translation

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-07T15:34:52.199462Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:34:52.199462Z digest=sha256:9a7636a78275fc44c61702bfd830c9171f55820ea15ed717cbbea37e6b328a8b

Observation 2e9deba9-8e08-489e-9407-3eeaba038f60 · outbound

This paper cites Training language models to follow instructions with human feedback.Advances in neural information processing systems, 35:27730–27744, 2022.

Reward Reasoning Model Training language models to follow instructions with human feedback.Advances in neural information processing systems, 35:27730–27744, 2022

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-07T15:34:52.171617Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:34:52.171617Z digest=sha256:6466dfcd43e0ab9e49477b60c6103053cb2e69986b774726905a0b9abb9748c1

Observation d40c1b1c-1b3f-4874-8d8c-5e9cad63712b · outbound

This paper cites Qwen2.5 Technical Report.

Reward Reasoning Model Qwen2.5 Technical Report

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-07T15:34:52.271529Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:34:52.271529Z digest=sha256:119fe8da54b1cfe1d1555b15e095c07b80ae9df1d3af22569cad411a69c6f277

Observation 7f077f48-f53e-4671-8c10-968b40c4aba1 · outbound

This paper cites OffsetBias: Leveraging debiased data for tuning evaluators.

Reward Reasoning Model OffsetBias: Leveraging debiased data for tuning evaluators

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-07T15:34:52.231093Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:34:52.231093Z digest=sha256:6de2797fa6c71cf6e904a81d2fcd8a3cba76a3fee3a1c4a019d10bd3af439388

Observation 6d207513-7fda-4cb8-bf09-5dedac3090c0 · outbound

This paper cites Bradley Knox, Chelsea Finn, and Scott Niekum.

Reward Reasoning Model Bradley Knox, Chelsea Finn, and Scott Niekum

Reference 50

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:34:55.866191Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T15:34:52.342892Z digest=sha256:274acf7238af494f8e3f353ac05653a1c14595740c60affb39510238a9b45ad1

Observation 93b2115a-a6af-4994-afc3-20a88c48f4dd · outbound

This paper cites Direct preference optimization: Your language model is secretly a reward model.

Reward Reasoning Model Direct preference optimization: Your language model is secretly a reward model

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-07T15:34:52.301877Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:34:52.301877Z digest=sha256:ebc6a16e19e466b0a43127b9ead74d22df9ced94facaa1cca33eaf5d5c9f2682

Observation 1b28ebea-f75d-48d4-8953-78e11cbea415 · outbound

This paper cites DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models.

Reward Reasoning Model DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models

Reference 52

Resolution
unresolved
no resolver link, observed 2026-08-07T15:34:52.426178Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:34:52.426178Z digest=sha256:6bb399ced2a4954feacf8de0ee91100fcf053972095ca263cebea55434cb3e3a

Observation a5b06b04-de8e-479d-85e1-a53711d9db2d · outbound

This paper cites The limits of automatic summarisation according to rouge.

Reward Reasoning Model The limits of automatic summarisation according to rouge

Reference 53

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:34:55.856216Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T15:34:52.383058Z digest=sha256:2ee18dc4c983c30796e2fc6a31f0abed5391ad3e4a631ebfbbd2be4f0a451d00

Observation a9d9763e-c9e5-4662-920b-503fffe91314 · outbound

This paper cites Skywork critic model se- ries.

Reward Reasoning Model Skywork critic model se- ries

Reference 54

Resolution
unresolved
no resolver link, observed 2026-08-07T15:34:52.496371Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:34:52.496371Z digest=sha256:18a58d28336a8b4957b7e618be407cf2717db480b08bb9b27ddf77cc92e8d62b

Observation c60e5510-ccfa-4c84-8520-879502e2d277 · outbound

This paper cites HybridFlow: A Flexible and Efficient RLHF Framework.

Reward Reasoning Model HybridFlow: A Flexible and Efficient RLHF Framework

Reference 55

Resolution
unresolved
no resolver link, observed 2026-08-07T15:34:52.464502Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:34:52.464502Z digest=sha256:05277789325838a299dfdf0b82a95b7aaaf62c22a348578bc9789f2b97f304d6

Observation 8540d3f5-3be1-4a25-9b67-af2314e4b64a · outbound

This paper cites LLaMA: Open and Efficient Foundation Language Models.

Reward Reasoning Model LLaMA: Open and Efficient Foundation Language Models

Reference 56

Resolution
unresolved
no resolver link, observed 2026-08-07T15:34:52.573985Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:34:52.573985Z digest=sha256:acb4f064ebdc23c73c694a73d8965172d2ee94853c2a7dd88bd7ce9158a9e895

Observation cfef8121-af75-407d-85ed-ec9b6453c90e · outbound

This paper cites Scaling LLM test- time compute optimally can be more effective than scaling parameters for reasoning.

Reward Reasoning Model Scaling LLM test- time compute optimally can be more effective than scaling parameters for reasoning

Reference 57

Resolution
unresolved
no resolver link, observed 2026-08-07T15:34:52.538743Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:34:52.538743Z digest=sha256:73ff59cd8c0ef5b9003da8b88061cc9d9f5366b3068c748b03c78e47a65a4fcc

Observation a812a115-e4e8-4789-8eca-a48adb5d8948 · outbound

This paper cites Secrets of RLHF in Large Language Models Part II: Reward Modeling.

Reward Reasoning Model Secrets of RLHF in Large Language Models Part II: Reward Modeling

Reference 58

Resolution
unresolved
no resolver link, observed 2026-08-07T15:34:52.624148Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:34:52.624148Z digest=sha256:ad07cda78140bca82bce6e8187416c338d3ec922d0c6fd1a91e34f3b71f5987e

Observation 9e4f17b6-25e3-441e-a397-57d7508cd05e · outbound

This paper cites Foundational autoraters: Taming large language models for better automatic evaluation.

Reward Reasoning Model Foundational autoraters: Taming large language models for better automatic evaluation

Reference 59

Resolution
unresolved
no resolver link, observed 2026-08-07T15:34:52.593804Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:34:52.593804Z digest=sha256:9a29fa1f60273b2c8b0925b02499689cc66813bba62c7a14e0520aa1f74c0a4c

Observation 72abce9f-4137-4ad2-b582-5787a5b1cfe9 · outbound

This paper cites Pandalm: An automatic evaluation benchmark for llm instruction tuning optimization.

Reward Reasoning Model Pandalm: An automatic evaluation benchmark for llm instruction tuning optimization

Reference 60

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:34:55.834388Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T15:34:52.701742Z digest=sha256:66ad304ac1224602a47c45ee324f7d33cb7f8c85ce5b6f3c2b60d279295369ba

Observation d720deb4-efae-4f39-b636-0d2064001d37 · outbound

This paper cites Self-Taught Evaluators.

Reward Reasoning Model Self-Taught Evaluators

Reference 61

Resolution
unresolved
no resolver link, observed 2026-08-07T15:34:52.667559Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:34:52.667559Z digest=sha256:a977da56f8b953a6c176cd0e5a23c49dbf47df940e7c55ffff32c5527354539d

Observation 3b22891b-1cb1-4eb9-8f68-fa2c7b066dd8 · outbound

This paper cites HelpSteer: Multi-attribute Helpfulness Dataset for SteerLM.

Reward Reasoning Model HelpSteer: Multi-attribute Helpfulness Dataset for SteerLM

Reference 62

Resolution
unresolved
no resolver link, observed 2026-08-07T15:34:52.819560Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:34:52.819560Z digest=sha256:126fb5bf4f3669c78c7e9c3ec18ab161d206a1ef0211903d250f4f8257256667

Observation 341ee96e-1efe-4a1d-bfac-48453a6ff084 · outbound

This paper cites an unresolved cited work.

Reward Reasoning Model Unresolved cited work

Reference 63

Resolution
unresolved
raw_fallback, observed 2026-08-07T15:34:55.823373Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T15:34:52.737656Z digest=sha256:79d3e777fc10e65f6a111691d320f99ce7ddaa57e8845ce79d44d935bb4cab5e

Observation 34bbcde9-445e-44e9-a73d-b4270e62c790 · outbound

This paper cites Reinforcement learning for reasoning in large language models with one training example.

Reward Reasoning Model Reinforcement learning for reasoning in large language models with one training example

Reference 64

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:34:55.813269Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T15:34:52.773786Z digest=sha256:b9f89de269f79e9df6637589b5bae688e294ecdbae3afa12376a43121034c947

Observation ba7f0d83-0f17-420d-866d-a61a987a3d3a · outbound

This paper cites J1: Incentivizing thinking in llm-as-a-judge via reinforcement learning.arXiv preprint arXiv:2505.10320, 2025.

Reward Reasoning Model J1: Incentivizing thinking in llm-as-a-judge via reinforcement learning.arXiv preprint arXiv:2505.10320, 2025

Reference 65

Resolution
unresolved
no resolver link, observed 2026-08-07T15:34:52.940879Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:34:52.940879Z digest=sha256:a1a2cd5167246ec4464560af8509666ee3088366e13b9122f46c2a517f105b06

Observation 514a7db1-078e-4cde-b405-c3bca95521a6 · outbound

This paper cites Zhang, Makesh Narsimhan Sreedhar, and Oleksii Kuchaiev.

Reward Reasoning Model Zhang, Makesh Narsimhan Sreedhar, and Oleksii Kuchaiev

Reference 66

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:34:55.802137Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T15:34:52.862712Z digest=sha256:b0ee34a82609d2614b528cae48e66f2cce584651226a5052fa364d2f303377c2

Observation a79a40e6-edfa-40cc-b0df-2f7e4ae87bf4 · outbound

This paper cites Chi, Quoc V Le, and Denny Zhou.

Reward Reasoning Model Chi, Quoc V Le, and Denny Zhou

Reference 67

Resolution
unresolved
no resolver link, observed 2026-08-07T15:34:52.909489Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:34:52.909489Z digest=sha256:86e7053cc75956c9d451a41262f9382323347bd3baa59bdd8263cbfcda1abf74

Observation 6245467e-b388-4761-9d3c-06b0d0775aed · outbound

This paper cites Inference scaling laws: An empirical analysis of compute-optimal inference for LLM problem-solving.

Reward Reasoning Model Inference scaling laws: An empirical analysis of compute-optimal inference for LLM problem-solving

Reference 68

Resolution
unresolved
no resolver link, observed 2026-08-07T15:34:53.021374Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:34:53.021374Z digest=sha256:5899458f6a44394e1685009837ff1d0f0cf356611af07295ca5c6f1523fefac6

Observation 3914fb9b-10c8-4674-9571-b66f9dc9e801 · outbound

This paper cites Metametrics: Calibrating metrics for generation tasks using human preferences.

Reward Reasoning Model Metametrics: Calibrating metrics for generation tasks using human preferences

Reference 69

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:34:55.784727Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T15:34:52.996374Z digest=sha256:ef1f54e099d79457352f88e2004e7060f1cd86280b7706579bb7550a6b3b0d0e

Observation 84ed587f-2345-4098-8b28-592c52a1dbb8 · outbound

This paper cites Meta-Rewarding Language Models: Self-Improving Alignment with LLM-as-a-Meta-Judge.

Reward Reasoning Model Meta-Rewarding Language Models: Self-Improving Alignment with LLM-as-a-Meta-Judge

Reference 70

Resolution
unresolved
no resolver link, observed 2026-08-07T15:34:53.003907Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:34:53.003907Z digest=sha256:d956071008381214c142d8ec34ceb78f7770a3bbf47326128150a016940a4e54

Observation ed3d3685-5069-4f6d-a549-19afa6aca6af · outbound

This paper cites Learning LLM-as-a-judge for preference alignment.

Reward Reasoning Model Learning LLM-as-a-judge for preference alignment

Reference 71

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:34:55.759404Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T15:34:53.201510Z digest=sha256:859bca0cefac4de5fe17f1b568c0e69d7c7c4d4307917791b6f1484edac26355

Observation 24c0ae7f-4b60-4b0e-9cc0-07b08648a264 · outbound

This paper cites Qwen2 Technical Report.

Reward Reasoning Model Qwen2 Technical Report

Reference 72

Resolution
unresolved
no resolver link, observed 2026-08-07T15:34:53.085306Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:34:53.085306Z digest=sha256:c22befb66a0599ee198577311dc06f986cdf610c4cdc0d2fd69836063a630a3a

Observation 85e6dfe3-c4c2-442f-bd24-53cb550d790f · outbound

This paper cites Qwen2.5-Math Technical Report: Toward Mathematical Expert Model via Self-Improvement.

Reward Reasoning Model Qwen2.5-Math Technical Report: Toward Mathematical Expert Model via Self-Improvement

Reference 73

Resolution
unresolved
no resolver link, observed 2026-08-07T15:34:53.137028Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:34:53.137028Z digest=sha256:22fd5f3f32ae4c9aee69541f8facc39e3a0ce29f7032be4d69af784ca985f618

Observation de494628-796f-48eb-8e34-1871d3509850 · outbound

This paper cites Self-rewarding language models.

Reward Reasoning Model Self-rewarding language models

Reference 74

Resolution
unresolved
no resolver link, observed 2026-08-07T15:34:53.349374Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:34:53.349374Z digest=sha256:506fd5084c34ad3c6c8707754cc4532a92af145de0d72acaf51e6ac45f4f7218

Observation a6fdcaf4-4f08-469b-9515-792345129f24 · outbound

This paper cites DAPO: An Open-Source LLM Reinforcement Learning System at Scale.

Reward Reasoning Model DAPO: An Open-Source LLM Reinforcement Learning System at Scale

Reference 75

Resolution
unresolved
no resolver link, observed 2026-08-07T15:34:53.259571Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:34:53.259571Z digest=sha256:acdab6c6a36045fdeac6204064bb3f9a02d5e48d66099ddce4d31d4550e1d2cf

Observation e64558c4-476e-4df2-888d-8be1b6db038b · outbound

This paper cites Self-generated critiques boost reward modeling for language models.

Reward Reasoning Model Self-generated critiques boost reward modeling for language models

Reference 76

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:34:55.740256Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T15:34:53.306879Z digest=sha256:f47bef2b2fd045b9190329f0661d6161ac3e66228f1ca14f8b2c443e849766a2

Observation 3868ed14-945a-4658-936a-d99a0c28f3ae · outbound

This paper cites Gonzalez, and Ion Stoica.

Reward Reasoning Model Gonzalez, and Ion Stoica

Reference 77

Resolution
unresolved
no resolver link, observed 2026-08-07T15:34:53.509079Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:34:53.509079Z digest=sha256:0a4a3b602f22e08e7d021f1202a71445460037096c81bcd67ca86d486028fd61

Observation bf5a87e0-1d19-497d-8395-e141e26afb38 · outbound

This paper cites Generative verifiers: Reward modeling as next-token prediction.

Reward Reasoning Model Generative verifiers: Reward modeling as next-token prediction

Reference 78

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:34:55.714358Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T15:34:53.404561Z digest=sha256:fba05f7a8ff42b1d3116fce9622eecea94f4382797024396c2498dc8921e54e3

Observation a5b784e3-5bd9-4f6e-949f-953ce030b548 · outbound

This paper cites The Lessons of Developing Process Reward Models in Mathematical Reasoning.

Reward Reasoning Model The Lessons of Developing Process Reward Models in Mathematical Reasoning

Reference 79

Resolution
unresolved
no resolver link, observed 2026-08-07T15:34:53.466706Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:34:53.466706Z digest=sha256:04fe73135d2031081165584c538be3a5073b648816f281b3319e95a05c1aa267

Observation 09dd05a2-446f-4202-8241-8fb9c98a7b3c · outbound

This paper cites JudgeLM: Fine-tuned large language mod- els are scalable judges.

Reward Reasoning Model JudgeLM: Fine-tuned large language mod- els are scalable judges

Reference 80

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:34:55.686012Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T15:34:53.658787Z digest=sha256:a492be5b683fe649bf6be01593dd860b85521840bbeb6c44188d3590a61a5e8e

Observation 9df9a508-d550-4afe-bdfb-f965df1da3ef · outbound

This paper cites Judging llm-as-a-judge with mt-bench and chatbot arena.Advances in Neural Information Processing Systems, 36:46595–46623, 2023.

Reward Reasoning Model Judging llm-as-a-judge with mt-bench and chatbot arena.Advances in Neural Information Processing Systems, 36:46595–46623, 2023

Reference 81

Resolution
unresolved
no resolver link, observed 2026-08-07T15:34:53.545855Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:34:53.545855Z digest=sha256:30b2a2bec201b8598417dccc021c42c711cb7b192ea4d425687fdb8997b3b72c

Observation e82255d5-fafd-40d0-857b-3536af06b0d7 · outbound

This paper cites A Comprehensive Survey of Reward Models: Taxonomy, Applications, Challenges, and Future.

Reward Reasoning Model A Comprehensive Survey of Reward Models: Taxonomy, Applications, Challenges, and Future

Reference 82

Resolution
unresolved
no resolver link, observed 2026-08-07T15:34:53.601769Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:34:53.601769Z digest=sha256:340a6b8a86e7354771eeaab05e3adc36cfb5ab24b08c6275257508c7d3b87f46

Observation f22bab45-4e75-432a-8778-e7f7e5105b1a · outbound

This paper cites Partially Adhered.

Reward Reasoning Model Partially Adhered

Reference 85

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:34:55.466419Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T15:34:53.781321Z digest=sha256:a87801d1ac13dedbdb6f122b34ce4436678256fc79bdfebbd4b06f5aaac7c6eb

Observation 5065b370-aea8-49c4-9f0b-5c6e0df169ea · outbound

This paper cites Useful but Incomplete.

Reward Reasoning Model Useful but Incomplete

Reference 86

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:34:55.250151Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T15:34:53.828691Z digest=sha256:4cf8630d051ab02b9fa718afd8880862b72b639506c318c63c987f10d24331e1

Observation 51c7af70-e049-4117-9136-927f002999ed · outbound

This paper cites Not Detailed.

Reward Reasoning Model Not Detailed

Reference 87

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:34:54.938238Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T15:34:53.874909Z digest=sha256:095b40ce317722eff13aa64ebd0a70afa1be586b10c0c08c91ba23ce5d9008be

Observation 7be62865-2f07-45af-bfcd-045c54341bcd · outbound

This paper cites doi: 10.18653/v1/2024.emnlp-main.248.

Reward Reasoning Model doi: 10.18653/v1/2024.emnlp-main.248

Reference 2024

Resolution
unresolved
no resolver link, observed 2026-08-07T15:34:51.573946Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:34:51.573946Z digest=sha256:f3b43a3e6495bd274fe6db9e210ad6ca17bdff55cca4e5fefbabdc5807469172

Observation 6be746c7-c633-497b-844a-2731a542c577 · outbound

This paper cites \boxed{Assistant 1}.

Reward Reasoning Model \boxed{Assistant 1}

Reference 2025

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:34:55.665928Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T15:34:53.714483Z digest=sha256:f8b4de31e03e41718c9b79d19ec72e00363bcb1cc3c3e39bf7b4b792fe55d46f

Pith citing papers

Observation aab485b9-07c5-43db-a0cd-505f144055de · inbound

Tournament of Prompts: Evolving LLM Instructions Through Structured Debates and Elo Ratings cites this paper.

Tournament of Prompts: Evolving LLM Instructions Through Structured Debates and Elo Ratings Reward Reasoning Model

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-07T12:15:37.528799Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:15:37.528799Z digest=sha256:c3720baf1e13b3ebccb95f3b90ef3d6b51a736bdbb34d1b71f56474e3755e6d6

Observation 04ec0a88-7b73-41a4-92c7-b46d6cf7ab9b · inbound

Rubrics as Rewards: Reinforcement Learning Beyond Verifiable Domains cites this paper.

Rubrics as Rewards: Reinforcement Learning Beyond Verifiable Domains Reward Reasoning Model

Reference 12

Resolution
verified exact
arxiv_id, observed 2026-05-13T06:07:56.760173Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-13T06:07:56.678339Z digest=sha256:39a8e07af0fe5ba345d7a16b9c3606ec6d7130741fe56c1de56e6b820a4f56aa

Observation a2818eef-f5ce-4311-b868-782e9556d367 · inbound

GenSelect: A Generative Approach to Best-of-N cites this paper.

GenSelect: A Generative Approach to Best-of-N Reward Reasoning Model

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-06T14:51:26.527647Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T14:51:26.527647Z digest=sha256:9c8fe7e0122f40b678fe69b721b2aec611dee708933d61834271af71285a1c09

Observation cf18a61e-8d83-4445-8171-7d29f8bc4780 · inbound

VRPRM: Process Reward Modeling via Visual Reasoning cites this paper.

VRPRM: Process Reward Modeling via Visual Reasoning Reward Reasoning Model

Reference 3

Resolution
verified exact
arxiv_id, observed 2026-05-22T12:21:30.995972Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-22T12:20:17.430881Z digest=sha256:91848f3cae1e971e25b98a22b2a187d2624ce8fe3896f15525326b503b2724f1

Observation a5c71336-4ebc-41fc-83b2-a6491ddc19d0 · inbound

VRPRM: Process Reward Modeling via Visual Reasoning cites this paper.

VRPRM: Process Reward Modeling via Visual Reasoning Reward Reasoning Model

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-06T04:27:05.207869Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T04:27:05.207869Z digest=sha256:71d96e5a1854abf49d8b51ce380540174061c34e56de357800027c137687835a

Observation 25a4676c-98de-4856-8bc4-f4660d2fd46d · inbound

Learning to Refine: Self-Refinement of Parallel Reasoning in LLMs cites this paper.

Learning to Refine: Self-Refinement of Parallel Reasoning in LLMs Reward Reasoning Model

Reference 10

Resolution
verified exact
arxiv_id, observed 2026-05-18T21:16:50.971419Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-05-18T21:16:15.703057Z digest=sha256:4709d5131c6dc9b2bd2810ac4cc4929a78e1cc79af27ee5844c0ae53fd381d9e

Observation 589bb886-b043-4291-8a39-8c5ed8c35d8c · inbound

A Survey of Reinforcement Learning for Large Reasoning Models cites this paper.

A Survey of Reinforcement Learning for Large Reasoning Models Reward Reasoning Model

Reference 173

Resolution
metadata mismatch
arxiv_id, observed 2026-05-18T00:05:31.631151Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-05-18T00:02:24.352947Z digest=sha256:eea31ff342f83f76ce4d98dcd068f8bfac133c223e73d35284a15bd2e3eb2327

Observation c3900aa1-9feb-460d-9ceb-9b19a2dc5c8d · inbound

Reinforcement Learning Meets Large Language Models: A Survey of Advancements and Applications Across the LLM Lifecycle cites this paper.

Reinforcement Learning Meets Large Language Models: A Survey of Advancements and Applications Across the LLM Lifecycle Reward Reasoning Model

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-04T16:07:29.589806Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T16:07:29.589806Z digest=sha256:5d3c04ad663142805efcb89e3df218786c3916b2b5a4a8e4d76dc4684b38858d

Observation 5ee2d70b-a2a2-46cf-ad4e-9e71f89c5c87 · inbound

PaTaRM: Bridging Pairwise and Pointwise Signals via Preference-Aware Task-Adaptive Reward Modeling cites this paper.

PaTaRM: Bridging Pairwise and Pointwise Signals via Preference-Aware Task-Adaptive Reward Modeling Reward Reasoning Model

Reference 6

Resolution
metadata mismatch
arxiv_id, observed 2026-05-18T03:20:49.334817Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-18T03:15:57.744706Z digest=sha256:d9f0dee5e7cb7badcee6f4434822eb8d9a9f10ed59a1f2d4c86dd034d4a7738c

Observation c37af3d7-af75-4336-af67-24164bd603c7 · inbound

OpenReward: Learning to Reward Long-form Agentic Tasks via Reinforcement Learning cites this paper.

OpenReward: Learning to Reward Long-form Agentic Tasks via Reinforcement Learning Reward Reasoning Model

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-04T07:43:24.151892Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T07:43:24.151892Z digest=sha256:61cc5ba59bef807ccbb92201f57040e94ab6114497a8f6b3cf0910677e8f844e

Observation 7a625218-dee7-4b6e-b395-3d97d857f5b3 · inbound

AI Can Learn Scientific Taste cites this paper.

AI Can Learn Scientific Taste Reward Reasoning Model

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-02T18:14:50.408504Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T18:14:50.408504Z digest=sha256:98b928f254f2a352a12014f30b69f88ba6a30eb680bb7e6632f34dee42e48206

Observation e63bc315-3599-4d2c-9fd4-993a8f2154b0 · inbound

Trust Region On-Policy Distillation cites this paper.

Trust Region On-Policy Distillation Reward Reasoning Model

Reference 155

Resolution
verified exact
arxiv_id, observed 2026-07-01T20:56:13.571928Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-06-28T17:38:50.313305Z digest=sha256:62236a57f444e758205935b59a195cab72a57b1623d5698c683e8bf229ec870e

Observation ad125d65-8338-4ff8-b3a2-25cb045d814e · inbound

Biological Reasoning-Informed Regression for Interpretable Regulatory DNA Activity Prediction cites this paper.

Biological Reasoning-Informed Regression for Interpretable Regulatory DNA Activity Prediction Reward Reasoning Model

Reference 19

Resolution
verified exact
arxiv_id, observed 2026-07-02T22:17:26.268285Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-06-27T19:01:43.892354Z digest=sha256:a1345c5ae3ca5e6052f6e1994d57043b3f2283bf3a6940711ae02ae0c21c5113

Observation d329b7c8-0292-4432-947f-e065d7525ebf · inbound

Counsel: A Meta-Evaluation Dataset for Agentic Tasks cites this paper.

Counsel: A Meta-Evaluation Dataset for Agentic Tasks Reward Reasoning Model

Reference 26

Resolution
verified exact
arxiv_id, observed 2026-07-04T06:49:38.204971Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-06-26T14:07:59.446478Z digest=sha256:fcff22ef3fd1828b679e3530f2752f6b4d2033b8c9da4008354ecddfa995554e

Observation 12a1c266-d55d-436a-b877-aadc29b49771 · inbound

TAPAS: Throughput-adaptive Perception for Autonomous Systems cites this paper.

TAPAS: Throughput-adaptive Perception for Autonomous Systems Reward Reasoning Model

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-01T18:24:04.760985Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T18:24:04.760985Z digest=sha256:62a6264eeba4cfa69b20a7a0b2f7766d9deb99f76534b863295698fef1aff3a4