Pith. sign in

Paper Citation Record · LEDGER

Trust Region Preference Approximation: A simple and stable reinforcement learning algorithm for LLM reasoning

As of 6 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 6 inbound Pith citation observations for arXiv:2504.04524.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2504.04524 v2

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 6 of 6 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-05T06:32:48.257954+00:00

measured 6 of 6 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-04T06:47:16.873682Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-04T08:39:42.243396Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation 8a53ea4e-2d10-467b-868c-15af3475e4ae · inbound

Learning to Reason under Off-Policy Guidance cites this paper.

Learning to Reason under Off-Policy Guidance Trust Region Preference Approximation: A simple and stable reinforcement learning algorithm for LLM reasoning

Reference 46

Resolution
verified exact
arxiv_id, observed 2026-05-15T23:17:02.863067Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-15T23:17:02.701393Z digest=sha256:32fee00b84938d24400b0267a855fa5c5c46f387dc90c85e814d291aa9ba6513

Observation 18f62dfc-7d9f-4739-b673-f9e314a89751 · inbound

Variance-Aware Baselines and Adaptive Learning Rates for Reinforcement Learning with Verifiable Rewards cites this paper.

Variance-Aware Baselines and Adaptive Learning Rates for Reinforcement Learning with Verifiable Rewards Trust Region Preference Approximation: A simple and stable reinforcement learning algorithm for LLM reasoning

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-03T19:38:22.422879Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T19:38:22.422879Z digest=sha256:dfb1181018b529ad71d7094417830fd845847bfb7667f632e37a1a1e5feb9c0b

Observation d354e495-9260-4fa8-a6ce-32bd312f9be2 · inbound

Variance-Aware Baselines and Adaptive Learning Rates for Reinforcement Learning with Verifiable Rewards cites this paper.

Variance-Aware Baselines and Adaptive Learning Rates for Reinforcement Learning with Verifiable Rewards Trust Region Preference Approximation: A simple and stable reinforcement learning algorithm for LLM reasoning

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-04T06:47:16.873682Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T06:47:16.873682Z digest=sha256:14bd7cbef885e55f75a7fa183754749a951a2a77f8621cbf77892f893f647f72

Observation fccb151c-0cc3-4f93-83c2-cf182470efc2 · inbound

Revisiting Reinforcement Learning with Verifiable Rewards from a Contrastive Perspective cites this paper.

Revisiting Reinforcement Learning with Verifiable Rewards from a Contrastive Perspective Trust Region Preference Approximation: A simple and stable reinforcement learning algorithm for LLM reasoning

Reference 18

Resolution
verified exact
arxiv_id, observed 2026-05-14T20:22:54.809006Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-14T20:21:26.640493Z digest=sha256:38623b191c975573cc30183a31d15a6618cccad394fa6f29548cda4cdd04e5a4

Observation 19fedf20-ab9c-4e66-b1ab-2428440dc814 · inbound

Revisiting Reinforcement Learning with Verifiable Rewards from a Contrastive Perspective cites this paper.

Revisiting Reinforcement Learning with Verifiable Rewards from a Contrastive Perspective Trust Region Preference Approximation: A simple and stable reinforcement learning algorithm for LLM reasoning

Reference 18

Resolution
verified exact
arxiv_id, observed 2026-05-20T21:43:45.426051Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-20T21:42:49.452347Z digest=sha256:b072d3d452e94bb3b6e11ca266eb1fa98cd425cb051390a9805b36508bcb672c

Observation 7ef419e3-2fbe-40db-8337-6893563729c0 · inbound

Curriculum Reinforcement Learning Can Incentivize Reasoning Capacity in LLMs Beyond the Base Model cites this paper.

Curriculum Reinforcement Learning Can Incentivize Reasoning Capacity in LLMs Beyond the Base Model Trust Region Preference Approximation: A simple and stable reinforcement learning algorithm for LLM reasoning

Reference 27

Resolution
verified exact
arxiv_id, observed 2026-07-04T08:39:42.245665Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-06-26T11:10:56.755558Z digest=sha256:8dc24eaea5d03e3b84876386fe2d939c6c5910888ccf3c69940e150e593a2669