Pith. sign in

Paper Citation Record · LEDGER

Improving Reward Models with Synthetic Critiques

As of 9 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 3 inbound Pith citation observations for arXiv:2405.20850.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2405.20850 v2

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 3 of 3 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-08T06:32:00.761636+00:00

measured 3 of 3 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-07T15:09:54.077941Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-05-13T06:07:56.734008Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation aa3293c5-7248-418f-b6d0-259228f7611e · inbound

Think-RM: Enabling Long-Horizon Reasoning in Generative Reward Models cites this paper.

Think-RM: Enabling Long-Horizon Reasoning in Generative Reward Models Improving Reward Models with Synthetic Critiques

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-07T15:09:54.077941Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:09:54.077941Z digest=sha256:ae4a417196967cdc97d352c912c907bf588ee6d7b00f9324a0d7b7de15f6f115

Observation 6033c67d-e2d2-4948-892c-2582aa79a115 · inbound

Rubrics as Rewards: Reinforcement Learning Beyond Verifiable Domains cites this paper.

Rubrics as Rewards: Reinforcement Learning Beyond Verifiable Domains Improving Reward Models with Synthetic Critiques

Reference 36

Resolution
verified exact
arxiv_id, observed 2026-05-13T06:07:56.735973Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-13T06:07:56.678339Z digest=sha256:3a91e5dda012861c0f53bd329df2741111c3576eeeb085a6ff15d17c112266ce

Observation 666e019f-72e8-4a95-9c53-e87a7e1cefd1 · inbound

TRPrompt: Bootstrapping Query-Aware Prompt Optimization from Textual Rewards cites this paper.

TRPrompt: Bootstrapping Query-Aware Prompt Optimization from Textual Rewards Improving Reward Models with Synthetic Critiques

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-06T14:35:46.596042Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T14:35:46.596042Z digest=sha256:9c522a53f699855772950f6b7fdfc32e7be639fa2f0b180e32235f0fbdbe96c2