Pith. sign in

Paper Citation Record · LEDGER

Offline Reinforcement Learning for LLM Multi-Step Reasoning

As of 10 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 14 inbound Pith citation observations for arXiv:2412.16145.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2412.16145 v2

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 14 of 14 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-09T06:31:02.800959+00:00

measured 14 of 14 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-09T22:20:13.837391Z

measured 1 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-08-05T02:28:24.338817Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

1
arxiv_reference, observed 2026-08-05T02:28:24.338817Z

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation 5f1be69d-e707-456a-8242-d80ec7109be4 · inbound

BRiTE: Bootstrapping Reinforced Thinking Process to Enhance Language Model Reasoning cites this paper.

BRiTE: Bootstrapping Reinforced Thinking Process to Enhance Language Model Reasoning Offline Reinforcement Learning for LLM Multi-Step Reasoning

Reference 55

Resolution
unresolved
no resolver link, observed 2026-08-09T22:20:13.837391Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T22:20:13.837391Z digest=sha256:a7492fa604188251d8f7279d87806c2b306a964eea4ac2c3bbcd130626ec9c7c

Observation 3b599f16-a305-4fc8-9721-66e3ce66ec3c · inbound

PIPA: Preference Alignment as Prior-Informed Statistical Estimation cites this paper.

PIPA: Preference Alignment as Prior-Informed Statistical Estimation Offline Reinforcement Learning for LLM Multi-Step Reasoning

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-08T18:10:53.436669Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T18:10:53.436669Z digest=sha256:1b4a2c18c0f6e52ba6e6ffb509006d67a6999326f44139bccdfd9784c10aebfe

Observation 40d8a2c3-061f-4d2d-ab59-133b144a810e · inbound

Response-Level Rewards Are All You Need for Online Reinforcement Learning in LLMs: A Mathematical Perspective cites this paper.

Response-Level Rewards Are All You Need for Online Reinforcement Learning in LLMs: A Mathematical Perspective Offline Reinforcement Learning for LLM Multi-Step Reasoning

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-07T11:32:30.615635Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:32:30.615635Z digest=sha256:13aea5cbe4e7545d367ff44cae86e59d4c79f8cc6064573c2ce38677a3cd2a7d

Observation 3c787b02-2301-4e54-b8de-08a800880ea4 · inbound

Unleashing Embodied Task Planning Ability in LLMs via Reinforcement Learning cites this paper.

Unleashing Embodied Task Planning Ability in LLMs via Reinforcement Learning Offline Reinforcement Learning for LLM Multi-Step Reasoning

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-06T21:55:04.280984Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T21:55:04.280984Z digest=sha256:023080d1d0c8dfb8cacdebf180a7219eecaeba7defd7691ada33460f18007a58

Observation 1ebdf16a-4c76-47fc-970b-4ff90b9ad66e · inbound

A Technical Survey of Reinforcement Learning Techniques for Large Language Models cites this paper.

A Technical Survey of Reinforcement Learning Techniques for Large Language Models Offline Reinforcement Learning for LLM Multi-Step Reasoning

Reference 132

Resolution
unresolved
no resolver link, observed 2026-08-06T19:59:37.039893Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:59:37.039893Z digest=sha256:bdf3e08c27754feae0c44b269a7bda560245d7524a0b4af8688286a73b80a532

Observation a9ca7058-d9b0-4a0d-ba6d-d67f668e1201 · inbound

Think Clearly: Improving Reasoning via Redundant Token Pruning cites this paper.

Think Clearly: Improving Reasoning via Redundant Token Pruning Offline Reinforcement Learning for LLM Multi-Step Reasoning

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-07T00:23:18.742674Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:23:18.742674Z digest=sha256:f4b98a1b0b3befaa738bf85864f90e438c19e546be4db445c495026190737279

Observation 7e0152b2-7e3e-4348-be31-8bd04f407e02 · inbound

Reinforcement Learning in hyperbolic space for multi-step reasoning cites this paper.

Reinforcement Learning in hyperbolic space for multi-step reasoning Offline Reinforcement Learning for LLM Multi-Step Reasoning

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-06T15:24:27.704354Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:24:27.704354Z digest=sha256:89713ad0728c1d9d5fee9001b2ad84581cd84846c6b6daf8e7aadaa1a182d268

Observation a2612bc7-e4b2-45f7-82e7-425310cab47c · inbound

On the optimization dynamics of RLVR: Gradient gap and step size thresholds cites this paper.

On the optimization dynamics of RLVR: Gradient gap and step size thresholds Offline Reinforcement Learning for LLM Multi-Step Reasoning

Reference 21

Resolution
verified exact
arxiv_id, observed 2026-05-18T08:36:07.334653Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-18T08:34:36.543874Z digest=sha256:d722e8326ff3adde0a43279c5be651200435d4618f1d13833d66d8824c590200

Observation 2c55574d-7af5-489b-80b5-00c7710b88e2 · inbound

Pramana: Fine-Tuning Large Language Models for Epistemic Reasoning through Navya-Nyaya cites this paper.

Pramana: Fine-Tuning Large Language Models for Epistemic Reasoning through Navya-Nyaya Offline Reinforcement Learning for LLM Multi-Step Reasoning

Reference 7

Resolution
verified exact
arxiv_id, observed 2026-05-15T22:00:21.063340Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-15T21:58:40.973772Z digest=sha256:0f84b52ca20d12642e46d0101564d640af6de04c4cad31d7041b2224b3220be0

Observation 6440f138-b617-4260-9f74-7f4622f8fb3b · inbound

Trust Region On-Policy Distillation cites this paper.

Trust Region On-Policy Distillation Offline Reinforcement Learning for LLM Multi-Step Reasoning

Reference 257

Resolution
verified exact
arxiv_id, observed 2026-07-01T20:56:13.653343Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-06-28T17:38:50.313305Z digest=sha256:599728b655b621bf743e1a5e643f042324d4e8331a574a3c4c5b6c90bbcc2cbb

Observation 6583c324-0b76-40b0-b47f-248c6f4e67b0 · inbound

The Periodic Table of LLM Reasoning: A Structured Survey of Reasoning Paradigms, Methods, and Failure Modes cites this paper.

The Periodic Table of LLM Reasoning: A Structured Survey of Reasoning Paradigms, Methods, and Failure Modes Offline Reinforcement Learning for LLM Multi-Step Reasoning

Reference 240

Resolution
metadata mismatch
arxiv_id, observed 2026-06-27T13:00:56.038173Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-06-27T12:59:51.091008Z digest=sha256:e40efb206676b647aff76823d10c5e8eeb83c6e72280832a45ed02de42762c48

Observation 4f89117f-84fb-4cf7-b597-b2bc031fceba · inbound

Modularized Reinforcement Learning on LLMs: From MDP Creation to Exploration and Learning cites this paper.

Modularized Reinforcement Learning on LLMs: From MDP Creation to Exploration and Learning Offline Reinforcement Learning for LLM Multi-Step Reasoning

Reference 212

Resolution
verified exact
arxiv_id, observed 2026-07-04T07:59:40.096925Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-06-26T12:15:08.304150Z digest=sha256:5836fe199ca38495013f49c884b3fa398c6e512425f088838b1bef1dff1eb864

Observation 14a4ff1a-7006-4230-be33-2d08996d25f1 · inbound

LeAct: Learning to Reason from Expert Actions cites this paper.

LeAct: Learning to Reason from Expert Actions Offline Reinforcement Learning for LLM Multi-Step Reasoning

Reference 52

Resolution
unresolved
no resolver link, observed 2026-08-01T06:32:23.624774Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T06:32:23.624774Z digest=sha256:85380a50cb5a24a56531e6efb71c9d8a8b064ee6b41179c716c195d64c29b072

Observation 8b8a284a-2592-458f-a08e-8596a8c151ca · inbound

CRAFT: Learn the Schema, Execute the Plan cites this paper.

CRAFT: Learn the Schema, Execute the Plan Offline Reinforcement Learning for LLM Multi-Step Reasoning

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-02T10:18:39.144202Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T10:18:39.144202Z digest=sha256:85925a40be86272e90a3e9dcf73014fc17e53d71e26c2bc3c70b9e723d8860d4