Pith. sign in

Paper Citation Record · LEDGER

Think-RM: Enabling Long-Horizon Reasoning in Generative Reward Models

As of 15 August 2026, this Paper Citation Record lists 49 of 49 outbound references and 8 inbound Pith citation observations for arXiv:2505.16265.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2505.16265 v1

Coverage vector

measured 49 of 49 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-07T15:09:56.857674Z

measured 57 of 57 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-14T06:32:32.682623+00:00

measured 8 of 8 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-10T23:01:35.734381Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-01T20:56:13.265666Z

Reference resolution

49 of 49 outbound references displayed

  • verified exact0
  • verified fuzzy8
  • unresolved41
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation c58bb891-0435-43ff-9642-7742c8b15d5b · outbound

This paper cites Training a Helpful and Harmless Assistant with Reinforcement Learning from Human Feedback.

Think-RM: Enabling Long-Horizon Reasoning in Generative Reward Models Training a Helpful and Harmless Assistant with Reinforcement Learning from Human Feedback

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-07T15:09:52.682509Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:09:52.682509Z digest=sha256:bf6dc39e5c506c675120bc6a82e8c945d6aed65abe2e1ed4863ecce6fd015991

Observation 59bfa91b-2909-4427-9ccd-c67d9ec13674 · outbound

This paper cites Training language models to follow instructions with human feedback.Advances in neural information processing systems, 35:27730–27744, 2022.

Think-RM: Enabling Long-Horizon Reasoning in Generative Reward Models Training language models to follow instructions with human feedback.Advances in neural information processing systems, 35:27730–27744, 2022

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-07T15:09:52.737506Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:09:52.737506Z digest=sha256:1beadf84dea8251b0b66fcdb8db5018f6f29f3f76bc0c939933eec49e530c4ea

Observation f07ae36f-c76b-4d7c-8c1f-f222743a2dc6 · outbound

This paper cites Self-instruct: Aligning language models with self-generated in- structions.

Think-RM: Enabling Long-Horizon Reasoning in Generative Reward Models Self-instruct: Aligning language models with self-generated in- structions

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-07T15:09:52.808601Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:09:52.808601Z digest=sha256:476e5e2f7c0f9af4f52f6edef21481a8fcf7c5e9105010627ab5ad6f8d735dc8

Observation 43a330cc-1fed-4d3c-9827-4ac2971ef4f5 · outbound

This paper cites Direct preference optimization: Your language model is secretly a reward model.

Think-RM: Enabling Long-Horizon Reasoning in Generative Reward Models Direct preference optimization: Your language model is secretly a reward model

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-07T15:09:52.870219Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:09:52.870219Z digest=sha256:1e1b84a9ac52b2ea58ac9d865bf8a9622936d599f623f676c67de25427b0a62e

Observation d72f5654-f8a9-48f7-8ff6-9ad313727da2 · outbound

This paper cites Solving math word problems with process- and outcome-based feedback.

Think-RM: Enabling Long-Horizon Reasoning in Generative Reward Models Solving math word problems with process- and outcome-based feedback

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-07T15:09:52.975342Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:09:52.975342Z digest=sha256:d338404bcee6ec400f412d0e988fdde0a12104af9b0a34f2d1e27a81b2993b3a

Observation 923b156a-bd74-4bff-985b-e3b4802d60f3 · outbound

This paper cites Let’s verify step by step.

Think-RM: Enabling Long-Horizon Reasoning in Generative Reward Models Let’s verify step by step

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-07T15:09:53.051311Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:09:53.051311Z digest=sha256:27d164ce830a60bf89f764e153e77ca008d5e697454fab4e0efd2bcad9fe04e8

Observation a518297f-2337-45c0-949c-ee65ff909fff · outbound

This paper cites Math-Shepherd: Verify and Reinforce LLMs Step-by-step without Human Annotations.

Think-RM: Enabling Long-Horizon Reasoning in Generative Reward Models Math-Shepherd: Verify and Reinforce LLMs Step-by-step without Human Annotations

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-07T15:09:53.131349Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:09:53.131349Z digest=sha256:85b6ee851df114e2b534e62d4b721d4a3c84a99906abdce47836ed4ec2fa63c9

Observation d19f54ef-db2a-4b5b-867e-f353c9648dd4 · outbound

This paper cites Teaching Large Language Models to Reason with Reinforcement Learning.

Think-RM: Enabling Long-Horizon Reasoning in Generative Reward Models Teaching Large Language Models to Reason with Reinforcement Learning

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-07T15:09:53.173407Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:09:53.173407Z digest=sha256:8c8a6af37e2bfcf3629e9c4d3311b705bcbf26edfe3b64bd336e76209b5c5170

Observation 95b87412-9d1f-44a2-84a9-81f256e99745 · outbound

This paper cites Beavertails: Towards improved safety alignment of LLM via a human-preference dataset.

Think-RM: Enabling Long-Horizon Reasoning in Generative Reward Models Beavertails: Towards improved safety alignment of LLM via a human-preference dataset

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:09:59.615816Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-07T15:09:53.261459Z digest=sha256:06b29233e4c1f2ea82c72ba030d5b2e8cb8b50e8f5617326d1b39262b5f50ef2

Observation b8da442f-c411-4031-8c19-ccdcebc0bbe0 · outbound

This paper cites Safe rlhf: Safe reinforcement learning from human feedback.

Think-RM: Enabling Long-Horizon Reasoning in Generative Reward Models Safe rlhf: Safe reinforcement learning from human feedback

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-07T15:09:53.325000Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:09:53.325000Z digest=sha256:7ef152eafd27a2194ff80ce48202d937805cfc47a908ee441fd7a77dd7737e48

Observation 1be1e9ba-9aab-4af3-9a66-735e7c64dbff · outbound

This paper cites Rule Based Rewards for Language Model Safety.

Think-RM: Enabling Long-Horizon Reasoning in Generative Reward Models Rule Based Rewards for Language Model Safety

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-07T15:09:53.390130Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:09:53.390130Z digest=sha256:2b3b95d68695d67e14c8b5a7aa7872f028fdd9824bd97c04308ad84b0f265c80

Observation 6a542a66-92af-4f80-9787-089b613587b3 · outbound

This paper cites Deep reinforcement learning from human preferences.Advances in neural information processing systems, 30, 2017.

Think-RM: Enabling Long-Horizon Reasoning in Generative Reward Models Deep reinforcement learning from human preferences.Advances in neural information processing systems, 30, 2017

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-07T15:09:53.452403Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:09:53.452403Z digest=sha256:0e43cd69659c718249f2d0d72da44d273346ece7fe167106c4008e832a6d336d

Observation 9714b577-6cec-468a-b8b7-8700d92a58cb · outbound

This paper cites Ziegler, Ryan Lowe, Chelsea V oss, Alec Radford, Dario Amodei, and Paul F.

Think-RM: Enabling Long-Horizon Reasoning in Generative Reward Models Ziegler, Ryan Lowe, Chelsea V oss, Alec Radford, Dario Amodei, and Paul F

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:09:59.370846Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-07T15:09:53.539620Z digest=sha256:89aff7bc6037141a93adf968984a72399ff85d54971429b4aed7f27ab9fb6307

Observation f9169d0a-e4b5-4e96-b42f-62d01dd8571c · outbound

This paper cites Helping or Herding? Reward Model Ensembles Mitigate but do not Eliminate Reward Hacking.

Think-RM: Enabling Long-Horizon Reasoning in Generative Reward Models Helping or Herding? Reward Model Ensembles Mitigate but do not Eliminate Reward Hacking

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-07T15:09:53.635786Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:09:53.635786Z digest=sha256:ada03749c679705f91fc3f296a2f097a185d97fb1f315cb0b082141e917d0d31

Observation 2e3a2de9-5eb9-4149-98b7-fb67fb6bd51f · outbound

This paper cites Scaling laws for reward model overoptimization.

Think-RM: Enabling Long-Horizon Reasoning in Generative Reward Models Scaling laws for reward model overoptimization

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-07T15:09:53.734461Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:09:53.734461Z digest=sha256:2baeee37475acf1dd8f1e138740ab9f2f6708005fc04c8afb57e60a6eb2e6027

Observation 9cb827c2-9926-4f6e-8717-aabe5f2532a0 · outbound

This paper cites Direct Language Model Alignment from Online AI Feedback.

Think-RM: Enabling Long-Horizon Reasoning in Generative Reward Models Direct Language Model Alignment from Online AI Feedback

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-07T15:09:53.852749Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:09:53.852749Z digest=sha256:4e8bbc7dd79a6c32d3a7116ec630ba8ad4b743370ed78f14996d3e7979cdc64c

Observation 31781822-1fb8-43b9-b342-af1f9554021d · outbound

This paper cites RRM: Robust Reward Model Training Mitigates Reward Hacking.

Think-RM: Enabling Long-Horizon Reasoning in Generative Reward Models RRM: Robust Reward Model Training Mitigates Reward Hacking

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-07T15:09:53.954441Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:09:53.954441Z digest=sha256:e9d222d8c26477fc65964159ce413b5753bf8ec8d447d1f9a1a350486aeae37b

Observation aa3293c5-7248-418f-b6d0-259228f7611e · outbound

This paper cites Improving Reward Models with Synthetic Critiques.

Think-RM: Enabling Long-Horizon Reasoning in Generative Reward Models Improving Reward Models with Synthetic Critiques

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-07T15:09:54.077941Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:09:54.077941Z digest=sha256:baabdb53a76a7247410bfc9868e6b36b43aa5aae31e93f9f3d155f9b6955a975

Observation 73873e99-f590-43c3-8f88-36992110828b · outbound

This paper cites Critique-out-Loud Reward Models.

Think-RM: Enabling Long-Horizon Reasoning in Generative Reward Models Critique-out-Loud Reward Models

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-07T15:09:54.211281Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:09:54.211281Z digest=sha256:cab29d41197a0bc85935c4fe766d63ef7e03d235e1df7879d0972d193105ba90

Observation 02278e09-b24d-478b-8dae-36186b32ead0 · outbound

This paper cites Generative Reward Models.

Think-RM: Enabling Long-Horizon Reasoning in Generative Reward Models Generative Reward Models

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-07T15:09:54.317164Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:09:54.317164Z digest=sha256:f241e5034e94a021c8641668c22d5e69c3a1c1d10de5992a15fa748c067b7687

Observation faf6a9ce-0c08-4f47-8de0-a77a1dec0b40 · outbound

This paper cites Self-Generated Critiques Boost Reward Modeling for Language Models.

Think-RM: Enabling Long-Horizon Reasoning in Generative Reward Models Self-Generated Critiques Boost Reward Modeling for Language Models

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-07T15:09:54.445894Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:09:54.445894Z digest=sha256:c39346945453b9e08463ea4e3ac27e71e4587b025c4d62eafd07b18226441273

Observation ab3d36a6-1f52-434c-a252-04f94a95dbe2 · outbound

This paper cites Generative verifiers: Reward modeling as next-token prediction.

Think-RM: Enabling Long-Horizon Reasoning in Generative Reward Models Generative verifiers: Reward modeling as next-token prediction

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:09:59.140897Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-07T15:09:54.584180Z digest=sha256:b7a7a17ea9a77b68b0b38d852cca2d370970d26dcdffc64ba11aac6309d66d07

Observation ca8cef68-5cf3-4bfc-9169-7caea7e48c55 · outbound

This paper cites Learning to reason with llms.OpenAI Blog, 2024.

Think-RM: Enabling Long-Horizon Reasoning in Generative Reward Models Learning to reason with llms.OpenAI Blog, 2024

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-07T15:09:54.673142Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:09:54.673142Z digest=sha256:7fb98799c02efcab6cf25980be4b58d2e23844d930f46d33fae0f39a2c4900a3

Observation c2ca967b-49ae-483d-a002-177acfefe3f3 · outbound

This paper cites DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning.

Think-RM: Enabling Long-Horizon Reasoning in Generative Reward Models DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-07T15:09:54.768046Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:09:54.768046Z digest=sha256:a26ed78063d3526034d6332c30ccb19796017b4d402d64bf2542834ce65bf32c

Observation d056e02f-8513-4fec-955c-7b8b41f8f23e · outbound

This paper cites s1: Simple test-time scaling.

Think-RM: Enabling Long-Horizon Reasoning in Generative Reward Models s1: Simple test-time scaling

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-07T15:09:54.919059Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:09:54.919059Z digest=sha256:71f6cbed1d9388b3d1478ac0aec4d45bbb475695c5a16f41a4234c14cc52b669

Observation d21f3082-9dce-4316-a9f9-2f06c95f2fcf · outbound

This paper cites LLaMA: Open and Efficient Foundation Language Models.

Think-RM: Enabling Long-Horizon Reasoning in Generative Reward Models LLaMA: Open and Efficient Foundation Language Models

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-07T15:09:55.000866Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:09:55.000866Z digest=sha256:09c75ef067fd79bffa9efdd0b7d49c5e270d6068eb7c4e8916e40bb2f95249eb

Observation b3c836f0-fc4e-4263-b699-052ca9d012b6 · outbound

This paper cites The Llama 3 Herd of Models.

Think-RM: Enabling Long-Horizon Reasoning in Generative Reward Models The Llama 3 Herd of Models

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-07T15:09:55.146505Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:09:55.146505Z digest=sha256:17650c606b55da9f8f62c4bf5a7ea800a8f8f297da071bc2c89b08aedd965ad3

Observation d6f6d54b-6ec2-4a63-9475-00a3d0475250 · outbound

This paper cites RM-bench: Benchmarking reward models of language models with subtlety and style.

Think-RM: Enabling Long-Horizon Reasoning in Generative Reward Models RM-bench: Benchmarking reward models of language models with subtlety and style

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-07T15:09:55.283280Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:09:55.283280Z digest=sha256:ef8a0c84af25c63fcdc6e378aec6a325fe7e217d52b87d9f3cfb021991f45055

Observation 00e257ad-035a-4abd-8868-118ac98e178b · outbound

This paper cites Rank analysis of incomplete block designs: I.

Think-RM: Enabling Long-Horizon Reasoning in Generative Reward Models Rank analysis of incomplete block designs: I

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-07T15:09:55.398260Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:09:55.398260Z digest=sha256:cba28821ec680143b0c7e61679b0ca9fd5e11521213f9a4915ad306a53afe629

Observation 1f30d9f2-2268-4864-8233-10613ff33b72 · outbound

This paper cites GPT-4 Technical Report.

Think-RM: Enabling Long-Horizon Reasoning in Generative Reward Models GPT-4 Technical Report

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-07T15:09:55.488038Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:09:55.488038Z digest=sha256:cf08cf78b00b50c6ddff9198237804ad66dc37bd66bba877c19740f828446cf8

Observation 82111220-12b3-46b0-8fc2-b0f3719e97c0 · outbound

This paper cites Qwen2.5 Technical Report.

Think-RM: Enabling Long-Horizon Reasoning in Generative Reward Models Qwen2.5 Technical Report

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-07T15:09:55.548298Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:09:55.548298Z digest=sha256:9c90b1974fa5a04083c378034bac98aae1ce3edebadacadfa26f1206a1cd2691

Observation 315b6cf5-c3f8-4782-a3d8-6976fb258f4a · outbound

This paper cites Judging llm-as-a-judge with mt-bench and chatbot arena.Advances in Neural Information Processing Systems, 36:46595–46623, 2023.

Think-RM: Enabling Long-Horizon Reasoning in Generative Reward Models Judging llm-as-a-judge with mt-bench and chatbot arena.Advances in Neural Information Processing Systems, 36:46595–46623, 2023

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-07T15:09:55.607193Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:09:55.607193Z digest=sha256:a0a401efb07727639e9ddf2737142bea503461388a404dbe6604aa9978558779

Observation 405f8c7c-4d30-4e61-a0bd-a7d3597ff81a · outbound

This paper cites xFinder: Large Language Models as Automated Evaluators for Reliable Evaluation.

Think-RM: Enabling Long-Horizon Reasoning in Generative Reward Models xFinder: Large Language Models as Automated Evaluators for Reliable Evaluation

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-07T15:09:55.688065Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:09:55.688065Z digest=sha256:67c028f991e3931d9569ceb196226c2ffee9328f801907b1489750fa408e9c9f

Observation ea497a77-b5f1-46f0-a748-54aa1e3ede75 · outbound

This paper cites Chain of thought prompting elicits reasoning in large language models.

Think-RM: Enabling Long-Horizon Reasoning in Generative Reward Models Chain of thought prompting elicits reasoning in large language models

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:09:58.870900Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-07T15:09:55.756393Z digest=sha256:0dfe1e7c36a49f36b587cd51a6e62544977ae6a6b163c04316e4d5ab87c5ab6f

Observation 4755a606-83be-4f69-839a-7730e7adbc10 · outbound

This paper cites Self-Consistency Improves Chain of Thought Reasoning in Language Models.

Think-RM: Enabling Long-Horizon Reasoning in Generative Reward Models Self-Consistency Improves Chain of Thought Reasoning in Language Models

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-07T15:09:55.832125Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:09:55.832125Z digest=sha256:66575f3f235ec64254053d79e37cae3d5d99fe860472d22727570f3d3c34af1d

Observation 95e5ed7f-22b3-4919-bbdc-19ac14b174d8 · outbound

This paper cites Tree of thoughts: Deliberate problem solving with large language models.Ad- vances in neural information processing systems, 36:11809–11822, 2023.

Think-RM: Enabling Long-Horizon Reasoning in Generative Reward Models Tree of thoughts: Deliberate problem solving with large language models.Ad- vances in neural information processing systems, 36:11809–11822, 2023

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-07T15:09:55.913418Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:09:55.913418Z digest=sha256:ec16874455a94f60e7f18c95325646e40ef65b7101230bee19d31d2c99226844

Observation f8eedf22-3a99-45bf-8183-efc9e4022c63 · outbound

This paper cites Introducing openai o3 and o4-mini.OpenAI Blog, 2025.

Think-RM: Enabling Long-Horizon Reasoning in Generative Reward Models Introducing openai o3 and o4-mini.OpenAI Blog, 2025

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:09:58.678509Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-07T15:09:55.990614Z digest=sha256:e52956781e3b1d3e66c85b0cc8b5bda4da1acbd8ac3114390e1c0aaabfc853c9

Observation 91ec9f65-4299-4d48-995d-13cef4560e11 · outbound

This paper cites Qwq-32b: Embracing the power of reinforcement learning, March 2025.

Think-RM: Enabling Long-Horizon Reasoning in Generative Reward Models Qwq-32b: Embracing the power of reinforcement learning, March 2025

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-07T15:09:56.047246Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:09:56.047246Z digest=sha256:ce19920f83a7bd33ae42815df721cb12e5a21e9cdf59dc66977cbb60f6b1d80b

Observation fef25670-483a-49d5-8d98-06f2cf72795a · outbound

This paper cites Grok 3 beta — the age of reasoning agents.xAI Blog, 2025.

Think-RM: Enabling Long-Horizon Reasoning in Generative Reward Models Grok 3 beta — the age of reasoning agents.xAI Blog, 2025

Reference 39

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:09:58.514382Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-07T15:09:56.122255Z digest=sha256:cc3afffc34423e2d2a0a0409bda59eae9cb320a431fb144571b95e53dcb9868d

Observation 1f226cff-6000-4620-a171-35e03b575a8b · outbound

This paper cites Helpsteer2-preference: Complementing ratings with prefer- ences.

Think-RM: Enabling Long-Horizon Reasoning in Generative Reward Models Helpsteer2-preference: Complementing ratings with prefer- ences

Reference 40

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:09:58.371923Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-07T15:09:56.198046Z digest=sha256:01bcb20f26d53d0ae747cbad001fe9f21b733138024bfb88c5f5238afbbe4925

Observation 491e81c0-9044-40e1-ae49-84b255cc435f · outbound

This paper cites Proximal Policy Optimization Algorithms.

Think-RM: Enabling Long-Horizon Reasoning in Generative Reward Models Proximal Policy Optimization Algorithms

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-07T15:09:56.275062Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:09:56.275062Z digest=sha256:9305caa91d1addb6b693796440c5707d3a5c99ec507bd8247617bfa4abb0e1f2

Observation f9970fc8-6749-4400-8c26-6721eeff6864 · outbound

This paper cites High-Dimensional Continuous Control Using Generalized Advantage Estimation.

Think-RM: Enabling Long-Horizon Reasoning in Generative Reward Models High-Dimensional Continuous Control Using Generalized Advantage Estimation

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-07T15:09:56.314481Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:09:56.314481Z digest=sha256:48e934d5bca1b9327a49b9c0e7e94dce4b9cd77aa9444b6b436d34e37d38d22b

Observation 2c5bc38d-4ee2-48c6-b992-dabbb58dff82 · outbound

This paper cites HelpSteer3: Human-Annotated Feedback and Edit Data to Empower Inference-Time Scaling in Open-Ended General-Domain Tasks.

Think-RM: Enabling Long-Horizon Reasoning in Generative Reward Models HelpSteer3: Human-Annotated Feedback and Edit Data to Empower Inference-Time Scaling in Open-Ended General-Domain Tasks

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-07T15:09:56.386550Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:09:56.386550Z digest=sha256:7f7f33bd83693c1eae24a8dacea217a31cbb1c3cf955d1c864ee71bfe573bb8c

Observation c77ff6d2-21a3-4b1c-b875-fabd16585d9a · outbound

This paper cites RewardBench: Evaluating Reward Models for Language Modeling.

Think-RM: Enabling Long-Horizon Reasoning in Generative Reward Models RewardBench: Evaluating Reward Models for Language Modeling

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-07T15:09:56.474821Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:09:56.474821Z digest=sha256:826182e7c9156532f9e47f0984ab7f799c05d218bf34420f2696df392c436609

Observation ff40040a-91c2-4b89-acc5-7b8adba6ef66 · outbound

This paper cites Length-controlled alpacaeval: A simple debiasing of automatic evaluators.

Think-RM: Enabling Long-Horizon Reasoning in Generative Reward Models Length-controlled alpacaeval: A simple debiasing of automatic evaluators

Reference 45

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:09:58.236959Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-07T15:09:56.552982Z digest=sha256:bba092e2affe5a906043940d4b81b49e5bdc3f34930d989b003986bb0c466cce

Observation f22dc59e-be31-420c-83ff-cf2f8ae7d59d · outbound

This paper cites OpenRLHF: An Easy-to-use, Scalable and High-performance RLHF Framework.

Think-RM: Enabling Long-Horizon Reasoning in Generative Reward Models OpenRLHF: An Easy-to-use, Scalable and High-performance RLHF Framework

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-07T15:09:56.640514Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:09:56.640514Z digest=sha256:3c0045efa19413853b7f52357a77fd2b64ecf2fa1a9bfc516c9cdafa7f253364

Observation 855b4e30-ca65-46b9-8145-2edb3eeeaab0 · outbound

This paper cites HybridFlow: A Flexible and Efficient RLHF Framework.

Think-RM: Enabling Long-Horizon Reasoning in Generative Reward Models HybridFlow: A Flexible and Efficient RLHF Framework

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-07T15:09:56.727030Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:09:56.727030Z digest=sha256:932966ede3d0200a41b5a7a9588a3efecc36d163322097218315fafa561e4e50

Observation 3f293609-f5a0-4863-ace7-1a3fb58bb197 · outbound

This paper cites Adam: A Method for Stochastic Optimization.

Think-RM: Enabling Long-Horizon Reasoning in Generative Reward Models Adam: A Method for Stochastic Optimization

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-07T15:09:56.803561Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:09:56.803561Z digest=sha256:c6c39187d9dceb0706d0d2feeb1cd86c7f65b901101d5c3f2dc87f85481c5529

Observation 6a916405-9325-47ba-adc5-217306390a62 · outbound

This paper cites Decoupled Weight Decay Regularization.

Think-RM: Enabling Long-Horizon Reasoning in Generative Reward Models Decoupled Weight Decay Regularization

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-07T15:09:56.857674Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:09:56.857674Z digest=sha256:ba55fecb9cdee9ba8c6b36ec91a5143cd59e8b852c5995f1ac088352da2a92ae

Pith citing papers

Observation 09ba2b4f-5426-456b-8e8d-e3b90054b011 · inbound

VRPRM: Process Reward Modeling via Visual Reasoning cites this paper.

VRPRM: Process Reward Modeling via Visual Reasoning Think-RM: Enabling Long-Horizon Reasoning in Generative Reward Models

Reference 4

Resolution
verified exact
arxiv_id, observed 2026-05-22T12:21:30.982861Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-05-22T12:20:17.430881Z digest=sha256:58b6b3ab63a598b43348a5278d37d18af640ea71154391d424b6c16176951213

Observation d56e645e-ecad-49e2-beb1-9b4a70e64d7c · inbound

VRPRM: Process Reward Modeling via Visual Reasoning cites this paper.

VRPRM: Process Reward Modeling via Visual Reasoning Think-RM: Enabling Long-Horizon Reasoning in Generative Reward Models

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-06T04:27:05.232691Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T04:27:05.232691Z digest=sha256:a5093920f1f259b24bd238a73d26f02a9c9419649f257faafaeeeb5399fba8c6

Observation a4e76bdc-41c2-4d56-a2ec-563800eacb62 · inbound

The Landscape of Agentic Reinforcement Learning for LLMs: A Survey cites this paper.

The Landscape of Agentic Reinforcement Learning for LLMs: A Survey Think-RM: Enabling Long-Horizon Reasoning in Generative Reward Models

Reference 266

Resolution
metadata mismatch
arxiv_id, observed 2026-05-18T19:21:48.839803Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-05-18T19:19:36.427337Z digest=sha256:c28b6d1d03788d9963f30e6c391fc543475b674b20b9d4a677e82b2b17d4e442

Observation d543c8c8-7392-4f7c-a4c7-18eb6997aa46 · inbound

A Survey of Reinforcement Learning for Large Reasoning Models cites this paper.

A Survey of Reinforcement Learning for Large Reasoning Models Think-RM: Enabling Long-Horizon Reasoning in Generative Reward Models

Reference 196

Resolution
metadata mismatch
arxiv_id, observed 2026-05-18T00:05:31.521389Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-05-18T00:02:24.352947Z digest=sha256:295b6b9ff6ffc9bc54e0b5cbe8192d0a921e1d0ca7c28a74648e11ba70c2dc91

Observation 6a7b75e1-8549-4c77-a49b-f46d21f74ffc · inbound

Leveraging Verifier-Based Reinforcement Learning in Image Editing cites this paper.

Leveraging Verifier-Based Reinforcement Learning in Image Editing Think-RM: Enabling Long-Horizon Reasoning in Generative Reward Models

Reference 26

Resolution
verified exact
arxiv_id, observed 2026-05-12T10:06:27.574471Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-05-07T08:00:33.307429Z digest=sha256:11717951e657f076d83497b98833855a46c9682d9bc3a6636e46fc1ce9d5affa

Observation 7c2feb22-2a50-40a3-a289-dbdae529330b · inbound

Leveraging Verifier-Based Reinforcement Learning in Image Editing cites this paper.

Leveraging Verifier-Based Reinforcement Learning in Image Editing Think-RM: Enabling Long-Horizon Reasoning in Generative Reward Models

Reference 26

Resolution
verified exact
arxiv_id, observed 2026-05-21T09:14:06.015503Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-05-21T09:11:02.183133Z digest=sha256:7c0a831d2ecd211d4fd93017fb2e4a274bab40cc86b7b7d16260e8f1ee79247c

Observation 1c86b97d-afc4-47f6-b8d9-8d25caf5fcec · inbound

Trust Region On-Policy Distillation cites this paper.

Trust Region On-Policy Distillation Think-RM: Enabling Long-Horizon Reasoning in Generative Reward Models

Reference 153

Resolution
verified exact
arxiv_id, observed 2026-07-01T20:56:13.267006Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-06-28T17:38:50.313305Z digest=sha256:88ee3b8716d30ad1cf8abeb9a4950f600d9b2a1ee97ff73e3739669f4bf3feda

Observation abcf2ec5-4a78-4c06-90b6-c33f2e7f18b6 · inbound

The Horizon Gap: Planning, Memory, Execution, Training, and Evaluation for Long-Horizon LLM Agents cites this paper.

The Horizon Gap: Planning, Memory, Execution, Training, and Evaluation for Long-Horizon LLM Agents Think-RM: Enabling Long-Horizon Reasoning in Generative Reward Models

Reference 93

Resolution
unresolved
no resolver link, observed 2026-08-10T23:01:35.734381Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T23:01:35.734381Z digest=sha256:a9e677b0978332d4f2702470d2c77797f4732cc4cce6ae7a149ba3aaa0501971