Pith. sign in

Paper Citation Record · LEDGER

OpenReward: Learning to Reward Long-form Agentic Tasks via Reinforcement Learning

As of 10 August 2026, this Paper Citation Record lists 19 of 19 outbound references and 2 inbound Pith citation observations for arXiv:2510.24636.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2510.24636 v3

Coverage vector

measured 19 of 19 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-04T07:43:26.153787Z

measured 21 of 21 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-10T06:31:04.303077+00:00

measured 2 of 2 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-04T15:12:55.978703Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-07-02T03:26:29.328767Z

Reference resolution

19 of 19 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved19
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 4017fda2-d60e-4662-ab92-ced0d4c72468 · outbound

This paper cites LitSearch: A Retrieval Benchmark for Scientific Literature Search.

OpenReward: Learning to Reward Long-form Agentic Tasks via Reinforcement Learning LitSearch: A Retrieval Benchmark for Scientific Literature Search

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-04T07:43:23.166612Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T07:43:23.166612Z digest=sha256:c9fba17317cd026cc87120131edb1bc68c81cba374f9465011fe14ee5d53c391

Observation a669659d-eb68-4d3f-b5f3-51fdfb80fb1c · outbound

This paper cites UltraFeedback: Boosting Language Models with Scaled AI Feedback.

OpenReward: Learning to Reward Long-form Agentic Tasks via Reinforcement Learning UltraFeedback: Boosting Language Models with Scaled AI Feedback

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-04T07:43:23.842667Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T07:43:23.842667Z digest=sha256:9156fba1d6e858bca6a6bb2f27071aca81f086f460c2668b54049d2374396143

Observation 0d203271-3140-4f3f-baaa-ff3b3defe4d4 · outbound

This paper cites A Survey on LLM-as-a-Judge.

OpenReward: Learning to Reward Long-form Agentic Tasks via Reinforcement Learning A Survey on LLM-as-a-Judge

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-04T07:43:23.975882Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T07:43:23.975882Z digest=sha256:76c611fe157a7fb3bfcc24641267032e2f64a3ba269a4322bc3975b876d1862c

Observation c37af3d7-af75-4336-af67-24164bd603c7 · outbound

This paper cites Reward Reasoning Model.

OpenReward: Learning to Reward Long-form Agentic Tasks via Reinforcement Learning Reward Reasoning Model

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-04T07:43:24.151892Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T07:43:24.151892Z digest=sha256:c5f409b1761b252e06726ec40839765fe20a8d814b89079b99a2514b4d4b65c1

Observation 4536026d-9fbc-432f-ab63-f788f4e274dd · outbound

This paper cites GPT-4o System Card.

OpenReward: Learning to Reward Long-form Agentic Tasks via Reinforcement Learning GPT-4o System Card

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-04T07:43:24.321802Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T07:43:24.321802Z digest=sha256:d5777b536d7d8a177215500e651a7f249f964bfc7ba27d42de4fc98b221f3bd0

Observation 9c9888e2-16e0-4a62-a3a7-75b701758423 · outbound

This paper cites Search-R1: Training LLMs to Reason and Leverage Search Engines with Reinforcement Learning.

OpenReward: Learning to Reward Long-form Agentic Tasks via Reinforcement Learning Search-R1: Training LLMs to Reason and Leverage Search Engines with Reinforcement Learning

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-04T07:43:24.471106Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T07:43:24.471106Z digest=sha256:26cdf2cc2b3d7a7140f1e01d7e8097f91de2bebc5fdd1ca4547f5726de1b1bea

Observation bb48b874-e55b-4f2c-8572-c3f6a9ed6cb4 · outbound

This paper cites RewardBench: Evaluating Reward Models for Language Modeling.

OpenReward: Learning to Reward Long-form Agentic Tasks via Reinforcement Learning RewardBench: Evaluating Reward Models for Language Modeling

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-04T07:43:24.665067Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T07:43:24.665067Z digest=sha256:9611b8dc0463d0a383cbff2a1d88b46b7ce45669ca6459016480f8d0130bac5f

Observation 2ec46496-2ed5-4882-87c3-f8db1312b954 · outbound

This paper cites LLMs-as-Judges: A Comprehensive Survey on LLM-based Evaluation Methods.

OpenReward: Learning to Reward Long-form Agentic Tasks via Reinforcement Learning LLMs-as-Judges: A Comprehensive Survey on LLM-based Evaluation Methods

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-04T07:43:24.808467Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T07:43:24.808467Z digest=sha256:3819872dd8dcc23e60b2c273057e787b4a744ab628b5f6d0cd112e74598e76e0

Observation 76c6d061-bfe2-4bd6-bed9-7a813641a7a9 · outbound

This paper cites DeepSeek-V3 Technical Report.

OpenReward: Learning to Reward Long-form Agentic Tasks via Reinforcement Learning DeepSeek-V3 Technical Report

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-04T07:43:24.944019Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T07:43:24.944019Z digest=sha256:bef3ea963e503865ebeccdf3c8c7d176f90b597d255f904dfeb1bd73f633905d

Observation 0825a5ca-74e1-48e8-97be-1945e1419bf0 · outbound

This paper cites ColBERTv2: Effective and Efficient Retrieval via Lightweight Late Interaction.

OpenReward: Learning to Reward Long-form Agentic Tasks via Reinforcement Learning ColBERTv2: Effective and Efficient Retrieval via Lightweight Late Interaction

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-04T07:43:25.107788Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T07:43:25.107788Z digest=sha256:8a3b12c7285f8e96d4b4abc91c97e13277760a0be54bca3611bbe821245c8e10

Observation df821748-0ef0-4ed6-a211-7d8dc675b0d9 · outbound

This paper cites Lianmin Zheng, Wei-Lin Chiang, Ying Sheng, Siyuan Zhuang, Zhanghao Wu, Yonghao Zhuang, Zi Lin, Zhuohan Li, Dacheng Li, Eric Xing, et al.

OpenReward: Learning to Reward Long-form Agentic Tasks via Reinforcement Learning Lianmin Zheng, Wei-Lin Chiang, Ying Sheng, Siyuan Zhuang, Zhanghao Wu, Yonghao Zhuang, Zi Lin, Zhuohan Li, Dacheng Li, Eric Xing, et al

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-04T07:43:25.470188Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T07:43:25.470188Z digest=sha256:0e8b35c7d4cac6d769ede779cc7733966a42f9ea73b6787a43d9875f3f4abdf1

Observation 620c362c-7cd2-42bc-ab82-327944e133cc · outbound

This paper cites DeepResearcher: Scaling Deep Research via Reinforcement Learning in Real-world Environments.

OpenReward: Learning to Reward Long-form Agentic Tasks via Reinforcement Learning DeepResearcher: Scaling Deep Research via Reinforcement Learning in Real-world Environments

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-04T07:43:25.595789Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T07:43:25.595789Z digest=sha256:35b71cdc725e82dd8e8c57608dfad98616d3f3ff3a87f4f6e17a92efb513914a

Observation 5494b6d3-cc63-4ce5-aa65-7cff8bc3b62f · outbound

This paper cites DeepReview: Improving LLM-based Paper Review with Human-like Deep Thinking Process.

OpenReward: Learning to Reward Long-form Agentic Tasks via Reinforcement Learning DeepReview: Improving LLM-based Paper Review with Human-like Deep Thinking Process

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-04T07:43:25.717056Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T07:43:25.717056Z digest=sha256:5b5f3421b86ebd4c49245bc7a4d7bb718a4d35cf25ccc1dfaca70a2ed9dc8616

Observation 9154d8a7-6529-4d91-a2b9-cbb063127e87 · outbound

This paper cites TTRL: Test-Time Reinforcement Learning.

OpenReward: Learning to Reward Long-form Agentic Tasks via Reinforcement Learning TTRL: Test-Time Reinforcement Learning

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-04T07:43:25.904052Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T07:43:25.904052Z digest=sha256:507b43c5615070f3938f17dded649fd6d25947b9daa139e6930eb0de12ae28ea

Observation d1803f90-09a4-43e3-9cd1-ab8b8054d950 · outbound

This paper cites <search>.

OpenReward: Learning to Reward Long-form Agentic Tasks via Reinforcement Learning <search>

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-04T07:43:26.153787Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T07:43:26.153787Z digest=sha256:3bd76112931b9473c1d6b306e56201789e1c89c13bec967b26d36c552b9500a2

Observation 0e72a6e3-9c4c-430a-b864-1cc7972b1e72 · outbound

This paper cites DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models.

OpenReward: Learning to Reward Long-form Agentic Tasks via Reinforcement Learning DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models

Reference 2021

Resolution
unresolved
no resolver link, observed 2026-08-04T07:43:25.323549Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T07:43:25.323549Z digest=sha256:6b25cd4e92375e843f894d1029557868cc9ba31d309b2b81f836564eb4e0576f

Observation fe62aaab-6e99-4cff-a2d1-4f62d6b61ad2 · outbound

This paper cites Gemini 2.5: Pushing the Frontier with Advanced Reasoning, Multimodality, Long Context, and Next Generation Agentic Capabilities.

OpenReward: Learning to Reward Long-form Agentic Tasks via Reinforcement Learning Gemini 2.5: Pushing the Frontier with Advanced Reasoning, Multimodality, Long Context, and Next Generation Agentic Capabilities

Reference 2023

Resolution
unresolved
no resolver link, observed 2026-08-04T07:43:23.625509Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T07:43:23.625509Z digest=sha256:1c664413e9fd62e7ca2b3308e1752d944487f0099c7ef2f5d1be2a967624eaa3

Observation 277f7038-62bf-4de2-9117-c86f581cb9d5 · outbound

This paper cites Judgelrm: Large reasoning models as a judge.arXiv preprint arXiv:2504.00050, 2025a.

OpenReward: Learning to Reward Long-form Agentic Tasks via Reinforcement Learning Judgelrm: Large reasoning models as a judge.arXiv preprint arXiv:2504.00050, 2025a

Reference 2024

Resolution
unresolved
no resolver link, observed 2026-08-04T07:43:23.483080Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T07:43:23.483080Z digest=sha256:c8d23fe48e1065d8b4e4f21092ece3c8d441df3d56ffee130a32aacd38c44018

Observation 6210a4bb-f318-477e-b054-a9acfed1291d · outbound

This paper cites TravelAgent: An AI Assistant for Personalized Travel Planning.

OpenReward: Learning to Reward Long-form Agentic Tasks via Reinforcement Learning TravelAgent: An AI Assistant for Personalized Travel Planning

Reference 2025

Resolution
unresolved
no resolver link, observed 2026-08-04T07:43:23.344859Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T07:43:23.344859Z digest=sha256:c1ec5c87e8f84c838eb40a56b251dbdc7886a3b88603f922fec2846060598a7e

Pith citing papers

Observation 2fa5c8cd-2e1d-4666-a6ab-e196f6207750 · inbound

AgenticRL: Self-Refining Agentic Reinforcement Learning for Vision-Conditioned UAV Navigation cites this paper.

AgenticRL: Self-Refining Agentic Reinforcement Learning for Vision-Conditioned UAV Navigation OpenReward: Learning to Reward Long-form Agentic Tasks via Reinforcement Learning

Reference 23

Resolution
verified exact
local_arxiv, observed 2026-07-02T03:26:29.330841Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-06-28T10:03:54.185019Z digest=sha256:cd0708580ad76fb28a4312d7092d3395b31e4f1d8f6b083bcd193d4eb6c37c75

Observation f31f01ec-4405-47ac-808b-e0a1c6ec45da · inbound

Fetch-then-Explore: Decoupling Selection from Extraction over a Persistent Workspace for Search Agents cites this paper.

Fetch-then-Explore: Decoupling Selection from Extraction over a Persistent Workspace for Search Agents OpenReward: Learning to Reward Long-form Agentic Tasks via Reinforcement Learning

Reference 174

Resolution
unresolved
no resolver link, observed 2026-08-04T15:12:55.978703Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T15:12:55.978703Z digest=sha256:7701205be84c4a48874531b412cb842fa7d4204b4cb632c669e2228f0becf259