Pith. sign in

Paper Citation Record · LEDGER

Beyond Verifiable Rewards: Scaling Reinforcement Learning for Language Models to Unverifiable Data

As of 9 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 7 inbound Pith citation observations for arXiv:2503.19618.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2503.19618 v2

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 7 of 7 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-09T06:31:02.800959+00:00

measured 7 of 7 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-07T13:35:54.275667Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-03T20:38:55.982366Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation c5e30f52-8d0b-4008-a8f2-fc41b25e47ad · inbound

Reinforcing General Reasoning without Verifiers cites this paper.

Reinforcing General Reasoning without Verifiers Beyond Verifiable Rewards: Scaling Reinforcement Learning for Language Models to Unverifiable Data

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-07T13:35:54.275667Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:35:54.275667Z digest=sha256:07cb2fea3e9d17a69ef3e3f3320c2395f7d0da8863e8c95de07ff207b59b6582

Observation ed179084-8be9-4cb2-abeb-95217cf7a310 · inbound

Direct Reasoning Optimization: Token-Level Reasoning Reflectivity Meets Rubric Gates for Unverifiable Tasks cites this paper.

Direct Reasoning Optimization: Token-Level Reasoning Reflectivity Meets Rubric Gates for Unverifiable Tasks Beyond Verifiable Rewards: Scaling Reinforcement Learning for Language Models to Unverifiable Data

Reference 27

Resolution
metadata mismatch
arxiv_id, observed 2026-05-19T09:52:14.141843Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-19T09:48:56.990745Z digest=sha256:650aca85f301a37c8a84e1ac0c38262f6c86fa7f6fe1f96be122fe5a6c695c00

Observation 163a6772-576c-499a-86c3-32ddb7dde046 · inbound

Compute as Teacher: Turning Inference Compute Into Reference-Free Supervision cites this paper.

Compute as Teacher: Turning Inference Compute Into Reference-Free Supervision Beyond Verifiable Rewards: Scaling Reinforcement Learning for Language Models to Unverifiable Data

Reference 23

Resolution
metadata mismatch
arxiv_id, observed 2026-05-18T15:46:34.079688Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-18T15:45:09.730804Z digest=sha256:73992b9ba2606074266ea99c5645082b6ef23da8b99a37917fcbee8488adafce

Observation 9793575a-5edd-46e4-a648-7609fd8e5b4f · inbound

Coupled Variational Reinforcement Learning for Language Model General Reasoning cites this paper.

Coupled Variational Reinforcement Learning for Language Model General Reasoning Beyond Verifiable Rewards: Scaling Reinforcement Learning for Language Models to Unverifiable Data

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-03T16:43:57.074776Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T16:43:57.074776Z digest=sha256:a6166fb208bd0a0c5ba210fc83b17a245010c8e6156343117ed08c118e95addf

Observation ac37b85a-00d4-4752-a6e6-1bca5813bced · inbound

Trust Your Memory: Verifiable Control of Smart Homes through Reinforcement Learning with Multi-dimensional Rewards cites this paper.

Trust Your Memory: Verifiable Control of Smart Homes through Reinforcement Learning with Multi-dimensional Rewards Beyond Verifiable Rewards: Scaling Reinforcement Learning for Language Models to Unverifiable Data

Reference 17

Resolution
verified exact
arxiv_id, observed 2026-05-11T08:06:01.140815Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-10T16:49:00.343580Z digest=sha256:b7564927da9386ed694cd1fd1edf9c6c05da5953fc1c7fe6711016cb73ddb451

Observation 4f3fdcbb-411b-4dd3-b2fa-138136a6315d · inbound

Likelihood scoring for continuations of mathematical text: a self-supervised benchmark with tests for shortcut vulnerabilities cites this paper.

Likelihood scoring for continuations of mathematical text: a self-supervised benchmark with tests for shortcut vulnerabilities Beyond Verifiable Rewards: Scaling Reinforcement Learning for Language Models to Unverifiable Data

Reference 36

Resolution
verified exact
arxiv_id, observed 2026-05-19T17:22:41.904395Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-19T17:20:39.969447Z digest=sha256:6246570c122ceb235884b093fe038f8025a59651568db8ee5ebffb2d33d732a1

Observation 98a01425-21af-4464-8e2a-77a6619d7ce1 · inbound

From Reasoning Traces to Reusable Modules: Understanding Compositional Generalization in Language Model Reasoning cites this paper.

From Reasoning Traces to Reusable Modules: Understanding Compositional Generalization in Language Model Reasoning Beyond Verifiable Rewards: Scaling Reinforcement Learning for Language Models to Unverifiable Data

Reference 63

Resolution
metadata mismatch
arxiv_id, observed 2026-07-03T20:38:55.983717Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-06-27T01:13:11.483599Z digest=sha256:aaa4399824a0975769db46def5ebd493ae20e1bc68af61cb775853b405c7664f