Pith. sign in

Paper Citation Record · LEDGER

StepWiser: Stepwise Generative Judges for Wiser Reasoning

As of 10 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 11 inbound Pith citation observations for arXiv:2508.19229.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2508.19229 v2

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 11 of 11 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-10T06:31:04.303077+00:00

measured 11 of 11 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-04T13:15:47.940922Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-05-21T22:40:43.196250Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation cae2534f-3fe8-4a48-b7cf-a13dee1c0617 · inbound

The Landscape of Agentic Reinforcement Learning for LLMs: A Survey cites this paper.

The Landscape of Agentic Reinforcement Learning for LLMs: A Survey StepWiser: Stepwise Generative Judges for Wiser Reasoning

Reference 164

Resolution
verified exact
arxiv_id, observed 2026-05-18T19:21:48.431511Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-18T19:19:36.427337Z digest=sha256:d90f384cecb0ef1cb4ef91fca0c12bd7f32a118918da3f0f2199ea8cc9921efb

Observation 16d53c6d-a008-4682-bbcc-524ba55d0895 · inbound

Beyond Correctness: Harmonizing Process and Outcome Rewards through RL Training cites this paper.

Beyond Correctness: Harmonizing Process and Outcome Rewards through RL Training StepWiser: Stepwise Generative Judges for Wiser Reasoning

Reference 31

Resolution
verified exact
arxiv_id, observed 2026-05-21T22:40:43.198972Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-05-21T22:38:57.833414Z digest=sha256:59017ecbfe7d9a82f488b9f11bdc34e0d44d1177a3449c38e11709e7fa8b244d

Observation c38c875f-867f-4278-bf55-448f101bb6dc · inbound

Simultaneous Multi-objective Alignment Across Verifiable and Non-verifiable Rewards cites this paper.

Simultaneous Multi-objective Alignment Across Verifiable and Non-verifiable Rewards StepWiser: Stepwise Generative Judges for Wiser Reasoning

Reference 67

Resolution
unresolved
no resolver link, observed 2026-08-04T13:15:47.940922Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T13:15:47.940922Z digest=sha256:bb8d415219335d77f2d92a8a31057119a6db4cc21dd57bfa9a8e7fa311b505df

Observation e665e297-36b3-4ea9-9532-fc39ece9bd3c · inbound

ReProbe: Efficient Test-Time Scaling of Multi-Step Reasoning by Probing Internal States of Large Language Models cites this paper.

ReProbe: Efficient Test-Time Scaling of Multi-Step Reasoning by Probing Internal States of Large Language Models StepWiser: Stepwise Generative Judges for Wiser Reasoning

Reference 4

Resolution
metadata mismatch
arxiv_id, observed 2026-05-18T00:20:32.081580Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-18T00:19:01.211080Z digest=sha256:f6bec3f6a7ff92c3b1e0fd2cf21474895a45ab8b3712a4a0839a56883932ca1b

Observation 7c94dc5b-54a7-489e-9f51-e93e7a6c470a · inbound

P-Check: Advancing Personalized Reward Model via Learning to Generate Dynamic Checklist cites this paper.

P-Check: Advancing Personalized Reward Model via Learning to Generate Dynamic Checklist StepWiser: Stepwise Generative Judges for Wiser Reasoning

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-03T12:29:43.974215Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T12:29:43.974215Z digest=sha256:809d058556b42b7a70466616f87249d4992f86252951551c040870cf52823adb

Observation ced92e37-a2d6-40d8-843e-b5d2ff7ce3ca · inbound

Reward Modeling for Reinforcement Learning-Based LLM Reasoning: Design, Challenges, and Evaluation cites this paper.

Reward Modeling for Reinforcement Learning-Based LLM Reasoning: Design, Challenges, and Evaluation StepWiser: Stepwise Generative Judges for Wiser Reasoning

Reference 114

Resolution
unresolved
no resolver link, observed 2026-08-03T03:04:44.848131Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T03:04:44.848131Z digest=sha256:f6065beaa1492fc1746fb998494eeaa429dee60eb08090b049273b836358b187

Observation 82e585c9-85f7-4920-8bd1-2fb38c21ed7c · inbound

StoSignSGD: Unbiased Structural Stochasticity Fixes SignSGD for Training Large Language Models cites this paper.

StoSignSGD: Unbiased Structural Stochasticity Fixes SignSGD for Training Large Language Models StepWiser: Stepwise Generative Judges for Wiser Reasoning

Reference 43

Resolution
verified exact
arxiv_id, observed 2026-05-10T12:15:22.276900Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-10T12:10:44.802059Z digest=sha256:334fccf6a3d80754954e37c5b5fb6d8b12c449d06abd139e161789ce45c93d74

Observation 3c895cc7-747d-4f29-b959-cbd6f8e399d4 · inbound

Process Rewards with Learned Reliability cites this paper.

Process Rewards with Learned Reliability StepWiser: Stepwise Generative Judges for Wiser Reasoning

Reference 65

Resolution
verified exact
arxiv_id, observed 2026-05-19T14:53:06.962668Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-19T14:51:13.966538Z digest=sha256:21f8ae236140dfb9e0ae7a4fae397313835399149c44ca7c800bccee9af0f1fd

Observation 2529183e-44ad-4f8f-82db-2d7c03b11040 · inbound

ReCrit: Transition-Aware Reinforcement Learning for Scientific Critic Reasoning cites this paper.

ReCrit: Transition-Aware Reinforcement Learning for Scientific Critic Reasoning StepWiser: Stepwise Generative Judges for Wiser Reasoning

Reference 40

Resolution
verified exact
arxiv_id, observed 2026-05-20T22:53:49.405350Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-20T22:51:56.666980Z digest=sha256:5e369f43ede6eb88de19f1d292da53bf73d88955ff263a39f316e4b59ecc983b

Observation c6967dfb-2fe5-4f0a-9d44-3f19d4dcf38d · inbound

The Hidden Signal of Verifier Strictness: Controlling and Improving Step-Wise Verification via Selective Latent Steering cites this paper.

The Hidden Signal of Verifier Strictness: Controlling and Improving Step-Wise Verification via Selective Latent Steering StepWiser: Stepwise Generative Judges for Wiser Reasoning

Reference 23

Resolution
verified exact
arxiv_id, observed 2026-05-21T06:54:01.042670Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-21T06:51:56.556213Z digest=sha256:24daf2516d517b7e2338348a600775544613f35bb5dd7ad6162436f1a9f7d676

Observation eb4010a3-4fbb-4df1-9b8b-caddea7b863e · inbound

Trajectories That Segment Themselves: Agent-Declared Boundaries as a Training Unit cites this paper.

Trajectories That Segment Themselves: Agent-Declared Boundaries as a Training Unit StepWiser: Stepwise Generative Judges for Wiser Reasoning

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-04T09:44:32.967941Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T09:44:32.967941Z digest=sha256:4411782f676d6f757b21bc9b891f3b3a9a73383ac75d125d875307bb2eb3b81f