Pith. sign in

Paper Citation Record · LEDGER

Process Reward Model with Q-Value Rankings

As of 11 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 13 inbound Pith citation observations for arXiv:2410.11287.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2410.11287 v2

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 13 of 13 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-11T06:34:44.6726+00:00

measured 13 of 13 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-10T15:49:44.016691Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-04T13:59:51.866161Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation 48f26aaa-f025-40cb-8a7f-01a2f6e712ee · inbound

Inference-Time Scaling for Diffusion Models beyond Scaling Denoising Steps cites this paper.

Inference-Time Scaling for Diffusion Models beyond Scaling Denoising Steps Process Reward Model with Q-Value Rankings

Reference 44

Resolution
verified exact
arxiv_id, observed 2026-05-20T11:45:17.597534Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-20T11:45:17.473970Z digest=sha256:2a8d63240de4dc2cd116cd71b14b9cca791b4774e0043191e59bf25c4da68a56

Observation 71dbc219-e964-4d7b-b9f2-442416720d72 · inbound

Coarse-to-Fine Process Reward Modeling for Mathematical Reasoning cites this paper.

Coarse-to-Fine Process Reward Modeling for Mathematical Reasoning Process Reward Model with Q-Value Rankings

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-10T15:49:44.016691Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T15:49:44.016691Z digest=sha256:204d0ab337653f01118f5f2e342b5ec790e441952d21dc3c3c08d3d3cfe29741

Observation 8c3df4bf-21af-4c82-87ea-c8af838614f8 · inbound

Can 1B LLM Surpass 405B LLM? Rethinking Compute-Optimal Test-Time Scaling cites this paper.

Can 1B LLM Surpass 405B LLM? Rethinking Compute-Optimal Test-Time Scaling Process Reward Model with Q-Value Rankings

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-08T14:40:35.988761Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T14:40:35.988761Z digest=sha256:409a4b1ee1ff31a417931c6a89107a0922bffefbbbbb6bb1e1da54738a593251

Observation 7443a436-cbb5-447d-b931-7f039b861a59 · inbound

Video-R1: Reinforcing Video Reasoning in MLLMs cites this paper.

Video-R1: Reinforcing Video Reasoning in MLLMs Process Reward Model with Q-Value Rankings

Reference 21

Resolution
verified exact
arxiv_id, observed 2026-05-12T09:43:00.397141Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-12T09:43:00.208065Z digest=sha256:36ab63e66c715742e1c7f1cf86a2d320ce6f109c2cb378564ebd2629eec4d3ce

Observation 9e0e3dc3-97e1-4e1e-b2b4-f66c1636120f · inbound

Shallow Preference Signals: Large Language Model Aligns Even Better with Truncated Data? cites this paper.

Shallow Preference Signals: Large Language Model Aligns Even Better with Truncated Data? Process Reward Model with Q-Value Rankings

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-07T15:15:46.413629Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:15:46.413629Z digest=sha256:9ea03ed5df964b1371bbce1db5f1b60f905c72bf5c16a072a46e65ecc92591b8

Observation d879184b-5a0e-4540-b615-95c2dd7eec5d · inbound

FreePRM: Training Process Reward Models Without Ground Truth Process Labels cites this paper.

FreePRM: Training Process Reward Models Without Ground Truth Process Labels Process Reward Model with Q-Value Rankings

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-07T11:06:41.397499Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:06:41.397499Z digest=sha256:34534db6002234f6b25f73261de19618b0e46e63c9f8b5db29db01865809e1d4

Observation 8e28e85a-a52d-48c4-81a8-bcdde58098ea · inbound

BMMR: A Large-Scale Bilingual Multimodal Multi-Discipline Reasoning Dataset cites this paper.

BMMR: A Large-Scale Bilingual Multimodal Multi-Discipline Reasoning Dataset Process Reward Model with Q-Value Rankings

Reference 60

Resolution
unresolved
no resolver link, observed 2026-08-06T20:15:52.788301Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:15:52.788301Z digest=sha256:e0f0b3e6edf13ec0223cf8d285174a8fe5c6e934cf80db332dd5a08ed9d6ff1b

Observation 1dff3c38-1b04-47eb-a071-2298d56aea64 · inbound

CoLD: Counterfactually-Guided Length Debiasing for Process Reward Models in Mathematical Reasoning cites this paper.

CoLD: Counterfactually-Guided Length Debiasing for Process Reward Models in Mathematical Reasoning Process Reward Model with Q-Value Rankings

Reference 9

Resolution
verified exact
arxiv_id, observed 2026-05-21T23:25:45.336123Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-21T23:24:43.556606Z digest=sha256:e9b4893a64d489fe101d1cc42c57c2d556b94482723fc1779fb96794c2de821b

Observation 6c701981-9b3a-4eff-8e56-2ae0d2d47507 · inbound

GM-PRM: A Generative Multimodal Process Reward Model for Multimodal Mathematical Reasoning cites this paper.

GM-PRM: A Generative Multimodal Process Reward Model for Multimodal Mathematical Reasoning Process Reward Model with Q-Value Rankings

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-06T00:59:52.140332Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T00:59:52.140332Z digest=sha256:84c09c3fd3b728054cb4382850544f918d8f3d9f5208dc3d5870f5faf5f390a6

Observation 50027bf8-af34-475d-ac06-6d1b5a69ec1e · inbound

GEAR: Granularity-Adaptive Advantage Reweighting for LLM Agents via Self-Distillation cites this paper.

GEAR: Granularity-Adaptive Advantage Reweighting for LLM Agents via Self-Distillation Process Reward Model with Q-Value Rankings

Reference 22

Resolution
verified exact
arxiv_id, observed 2026-05-13T06:47:26.714002Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-13T06:44:30.069444Z digest=sha256:9fd7deb41e974759e21ae815d380289853c99662c9e0ffebe306e3195c999bdf

Observation afac84e0-9a3c-4970-9a04-08d9b45a8c40 · inbound

GEAR: Granularity-Adaptive Advantage Reweighting for LLM Agents via Self-Distillation cites this paper.

GEAR: Granularity-Adaptive Advantage Reweighting for LLM Agents via Self-Distillation Process Reward Model with Q-Value Rankings

Reference 22

Resolution
verified exact
arxiv_id, observed 2026-05-15T05:59:48.297668Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-15T05:59:23.949124Z digest=sha256:d6c518d220ea86fec89ef49982a20b10272dd89f218fd2b57d9f12baf178b435

Observation 23d205d3-bf76-4552-b38b-6417f0701040 · inbound

From Failure to Feedback: Group Revision Unlocks Hard Cases in Object-Level Grounding cites this paper.

From Failure to Feedback: Group Revision Unlocks Hard Cases in Object-Level Grounding Process Reward Model with Q-Value Rankings

Reference 42

Resolution
verified exact
arxiv_id, observed 2026-05-20T18:43:38.791764Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-20T18:39:11.904941Z digest=sha256:efe8b8e6123fd60d6e4aeb458da43239171297a35736e6404c11dd6b03bd51a6

Observation 53b48197-71da-42f8-9961-2a058b7e843c · inbound

Reinforcement Learning without Ground-Truth Solutions can Improve LLMs cites this paper.

Reinforcement Learning without Ground-Truth Solutions can Improve LLMs Process Reward Model with Q-Value Rankings

Reference 50

Resolution
verified exact
arxiv_id, observed 2026-07-04T13:59:51.867763Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=arxiv_source observed=2026-06-26T04:47:47.691913Z digest=sha256:1550a37fdea76d722a5a2930b581b6cfa6be99f73b26a1c3e7f35a571cf1b9b9