Pith. sign in

Paper Citation Record · LEDGER

Distilled Reinforcement Learning for LLM Post-training

As of 9 August 2026, this Paper Citation Record lists 37 of 37 outbound references and 0 inbound Pith citation observations for arXiv:2607.17247.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2607.17247 v1

Coverage vector

measured 37 of 37 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-01T18:39:41.431601Z

measured 37 of 37 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-09T06:31:02.800959+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

37 of 37 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved37
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation d3639001-7ee8-425c-8e05-8ff8b063389a · outbound

This paper cites On-policy distillation of language models: Learning from self- generated mistakes.

Distilled Reinforcement Learning for LLM Post-training On-policy distillation of language models: Learning from self- generated mistakes

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-01T18:39:38.757408Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T18:39:38.757408Z digest=sha256:8b8a42de5b417210ecb8f68f05f9cda7b414295650ede301794819164e5ea4dc

Observation 66aa2a9b-e845-4840-8ca0-79cba8c1e59a · outbound

This paper cites Reasoning with Exploration: An Entropy Perspective.

Distilled Reinforcement Learning for LLM Post-training Reasoning with Exploration: An Entropy Perspective

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-01T18:39:39.011321Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T18:39:39.011321Z digest=sha256:667f8b578faf7f7000dc6b994b288fbf1b046cfac93e4963d54ce48e0a750b84

Observation f747b952-059c-4908-b455-ed0d5199652c · outbound

This paper cites an unresolved cited work.

Distilled Reinforcement Learning for LLM Post-training Unresolved cited work

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-01T18:39:41.431601Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T18:39:41.431601Z digest=sha256:79547da1f5a6925133b0b1f321f929681e9a200c5022b694fd94f51a6b7da1d5

Observation f308f626-b3e0-4513-abfa-ec7e4c1916f1 · outbound

This paper cites Revisiting On-Policy Distillation: Empirical Failure Modes and Simple Fixes.

Distilled Reinforcement Learning for LLM Post-training Revisiting On-Policy Distillation: Empirical Failure Modes and Simple Fixes

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-01T18:39:39.225995Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T18:39:39.225995Z digest=sha256:1156dc356ef1403ca5d61db5b1f7f0d0e3236132d7f3dcd90c0a5111aadd7202

Observation 77c158ab-115d-45df-a6df-efabde58cab0 · outbound

This paper cites ChatGLM: A Family of Large Language Models from GLM-130B to GLM-4 All Tools.

Distilled Reinforcement Learning for LLM Post-training ChatGLM: A Family of Large Language Models from GLM-130B to GLM-4 All Tools

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-01T18:39:39.287381Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T18:39:39.287381Z digest=sha256:65c92e24f6fada10b4ddd53d8413be541d4864fb7c0f94f44d5d3d5a284d9a5a

Observation 90a82ab3-edb9-4610-a60c-e8a8955ea2aa · outbound

This paper cites Minillm: Knowledge distillation of large language models.

Distilled Reinforcement Learning for LLM Post-training Minillm: Knowledge distillation of large language models

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-01T18:39:39.349107Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T18:39:39.349107Z digest=sha256:bac1312d4920d0a5ebc59d359b81035fd56e557a98ab380010024d13f1c4eacf

Observation 41a80cd6-60be-4fa5-8f4f-2fd9e34166d0 · outbound

This paper cites DeepSeek-Coder: When the Large Language Model Meets Programming -- The Rise of Code Intelligence.

Distilled Reinforcement Learning for LLM Post-training DeepSeek-Coder: When the Large Language Model Meets Programming -- The Rise of Code Intelligence

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-01T18:39:39.387426Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T18:39:39.387426Z digest=sha256:c790f5aa66fb7d444a884c1fa8d0d5643d1e90c35b9d9cf92ff755618a4c4588

Observation 4270e065-4d31-421d-975b-a6971cb08ef9 · outbound

This paper cites DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning.

Distilled Reinforcement Learning for LLM Post-training DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-01T18:39:39.453355Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T18:39:39.453355Z digest=sha256:74fe6e09aaa2f9eb681e23ff3d6643be69e6a03e3f73db632c6c9a8b8736c886

Observation 9e52171a-0e32-40c1-80d6-66c8516014d0 · outbound

This paper cites Sequence-level knowledge distillation.

Distilled Reinforcement Learning for LLM Post-training Sequence-level knowledge distillation

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-01T18:39:39.613393Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T18:39:39.613393Z digest=sha256:11c9732a3d8c593d4d5d0ae2adbbe0f8774cc78c1bf3b55b2c4905887769a7c2

Observation 5d2a8488-18fd-4831-9eac-02fb84f50a6f · outbound

This paper cites TreePO: Bridging the Gap of Policy Optimization and Efficacy and Inference Efficiency with Heuristic Tree-based Modeling.

Distilled Reinforcement Learning for LLM Post-training TreePO: Bridging the Gap of Policy Optimization and Efficacy and Inference Efficiency with Heuristic Tree-based Modeling

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-01T18:39:39.765212Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T18:39:39.765212Z digest=sha256:32324d40ecb08a848d1ff2fac71447a1f0183798b235c0c52c6bb98bdba6111c

Observation d66b7e15-bd33-47c8-b307-b7f66b3528e6 · outbound

This paper cites Let's Verify Step by Step.

Distilled Reinforcement Learning for LLM Post-training Let's Verify Step by Step

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-01T18:39:39.857107Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T18:39:39.857107Z digest=sha256:467b59baab3ced9704e1b52b740bb5bc80b4ce66035223670617d5a8548966ef

Observation 25e39112-4a92-4e04-897e-81a633265030 · outbound

This paper cites Length-unbiased sequence policy optimization: Revealing and controlling response length variation in rlvr.arXiv preprint arXiv:2602.05261,.

Distilled Reinforcement Learning for LLM Post-training Length-unbiased sequence policy optimization: Revealing and controlling response length variation in rlvr.arXiv preprint arXiv:2602.05261,

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-01T18:39:39.976099Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T18:39:39.976099Z digest=sha256:1e6f1109b872964082a378c0b90018003e59c61602c880e258e74ce9bef6a7da

Observation 0a52b37e-3c9d-4f12-9be6-b2f196349785 · outbound

This paper cites Improve Mathematical Reasoning in Language Models by Automated Process Supervision.

Distilled Reinforcement Learning for LLM Post-training Improve Mathematical Reasoning in Language Models by Automated Process Supervision

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-01T18:39:40.059010Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T18:39:40.059010Z digest=sha256:f2c97bb6de3b752d14e4daac5a01cc606e6d7008d4ebbbf78ae68c0a3d1b12dc

Observation b8d209c0-ab3f-41e1-81fa-19c0cc639450 · outbound

This paper cites Proximal Policy Optimization Algorithms.

Distilled Reinforcement Learning for LLM Post-training Proximal Policy Optimization Algorithms

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-01T18:39:40.146868Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T18:39:40.146868Z digest=sha256:63a725447b4ac179a3cacbf36b43f48adfdd9e459e5e19aa7991534d483442bb

Observation 35476f98-1e81-4179-a447-6f7e5039932d · outbound

This paper cites Kimi k1.5: Scaling Reinforcement Learning with LLMs.

Distilled Reinforcement Learning for LLM Post-training Kimi k1.5: Scaling Reinforcement Learning with LLMs

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-01T18:39:40.268643Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T18:39:40.268643Z digest=sha256:08e153f85ecd83094ac6eee56e961b90d52114b3f679e4784f3640152bad0f4a

Observation e4da1c71-9642-46c4-8a6e-806cf355410e · outbound

This paper cites Llama 2: Open Foundation and Fine-Tuned Chat Models.

Distilled Reinforcement Learning for LLM Post-training Llama 2: Open Foundation and Fine-Tuned Chat Models

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-01T18:39:40.318444Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T18:39:40.318444Z digest=sha256:ba94724c5112ce7e141c059a8630de97f81f7d1c96522a6c88cda3464634b4ff

Observation ac733956-c1a5-465e-a790-f3fbbf16fb9a · outbound

This paper cites SCOPE-RL: Stable and Quantitative Control of Policy Entropy in RL Post-Training.

Distilled Reinforcement Learning for LLM Post-training SCOPE-RL: Stable and Quantitative Control of Policy Entropy in RL Post-Training

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-01T18:39:40.386157Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T18:39:40.386157Z digest=sha256:b8112860495c9f707051034136824fd6a054ff16ffd2275b6f9c608da71c524c

Observation ecf09c03-2927-4a36-9243-d9819033cdab · outbound

This paper cites Implicit Compression Regularization: Concise Reasoning via Internal Shorter Distributions in RL Post-Training.

Distilled Reinforcement Learning for LLM Post-training Implicit Compression Regularization: Concise Reasoning via Internal Shorter Distributions in RL Post-Training

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-01T18:39:40.459071Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T18:39:40.459071Z digest=sha256:3747fdfc264a2b536555a19108f735ecf6e40b45abbc9e5dd885a5f7b973ce3a

Observation 823550af-5e8d-4066-89dc-f6741deba2b2 · outbound

This paper cites Lightning OPD: Efficient Post-Training for Large Reasoning Models with Offline On-Policy Distillation.

Distilled Reinforcement Learning for LLM Post-training Lightning OPD: Efficient Post-Training for Large Reasoning Models with Offline On-Policy Distillation

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-01T18:39:40.517367Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T18:39:40.517367Z digest=sha256:41ed6354a193e546426b6b3b83c401cc78d2ac4d09fa8f6a5824137f9dcc7395

Observation e7b7c17c-c775-4d1a-b2e9-a3a38fbc143d · outbound

This paper cites On the generalization of sft: A reinforcement learning perspective with reward rectification.arXiv preprint arXiv:2508.05629, 2025b.

Distilled Reinforcement Learning for LLM Post-training On the generalization of sft: A reinforcement learning perspective with reward rectification.arXiv preprint arXiv:2508.05629, 2025b

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-01T18:39:40.619539Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T18:39:40.619539Z digest=sha256:6c74fca5a813a72fcb80d372affc3d00d4bffb3518604dc42378b6f2a02b90e4

Observation 723986cb-e5e4-4411-94a7-bf89bd911262 · outbound

This paper cites Qwen2.5 Technical Report.

Distilled Reinforcement Learning for LLM Post-training Qwen2.5 Technical Report

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-01T18:39:40.700245Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T18:39:40.700245Z digest=sha256:b28e4a01e4972ed2aa3f6a28cd411abf57caca8376dc0b7b299448e85293fad8

Observation 3a83a283-4a92-4c3b-82db-f4231eb2b2c1 · outbound

This paper cites Qwen3 Technical Report.

Distilled Reinforcement Learning for LLM Post-training Qwen3 Technical Report

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-01T18:39:40.772979Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T18:39:40.772979Z digest=sha256:222a7be801b414e71a405cc8f0f98d35baf79d03a2cbc609f2dc2e93282cf22f

Observation 8acea5ca-4c80-4681-bac4-47206eb3b3dd · outbound

This paper cites Learning beyond Teacher: Generalized On-Policy Distillation with Reward Extrapolation.

Distilled Reinforcement Learning for LLM Post-training Learning beyond Teacher: Generalized On-Policy Distillation with Reward Extrapolation

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-01T18:39:40.846363Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T18:39:40.846363Z digest=sha256:29e39d9ab1623507341d2e0f235b3c7004517bce664db440b2e59bfa9c1c3589

Observation cb87e227-e96b-448c-96c4-342c584747ee · outbound

This paper cites Shorterbetter: Guiding reasoning models to find optimal inference length for efficient reasoning.arXiv preprint arXiv:2504.21370,.

Distilled Reinforcement Learning for LLM Post-training Shorterbetter: Guiding reasoning models to find optimal inference length for efficient reasoning.arXiv preprint arXiv:2504.21370,

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-01T18:39:40.916252Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T18:39:40.916252Z digest=sha256:c17bcb880ec318e595bd4ebc15164aba2b203e6ef5d202d8dc5da2d9553da958

Observation b9216b88-e480-4cc5-add2-b25348642673 · outbound

This paper cites DAPO: An Open-Source LLM Reinforcement Learning System at Scale.

Distilled Reinforcement Learning for LLM Post-training DAPO: An Open-Source LLM Reinforcement Learning System at Scale

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-01T18:39:41.004081Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T18:39:41.004081Z digest=sha256:01d0db885bd259657cc4826cf2dd8aa92cf611f2ff11a16dab2b3d6a3e6b3a71

Observation 866ed5c2-5541-4e61-b98d-380337bf4304 · outbound

This paper cites Does Reinforcement Learning Really Incentivize Reasoning Capacity in LLMs Beyond the Base Model?.

Distilled Reinforcement Learning for LLM Post-training Does Reinforcement Learning Really Incentivize Reasoning Capacity in LLMs Beyond the Base Model?

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-01T18:39:41.081442Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T18:39:41.081442Z digest=sha256:78ec69bc7d397ce0e540c9d01d4b6b71d314b2fdb1b1fb529505bd90f59f8dc1

Observation ee81150d-a934-48b1-8599-f78542a401ea · outbound

This paper cites DPO Meets PPO: Reinforced Token Optimization for RLHF.

Distilled Reinforcement Learning for LLM Post-training DPO Meets PPO: Reinforced Token Optimization for RLHF

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-01T18:39:41.185696Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T18:39:41.185696Z digest=sha256:1f5868e0b7398ebd86a1c5f4b13b5a79f9768d69a65eb0e6f5eb63ea24c049dc

Observation 419e3584-e959-4a1f-b782-c0a758894060 · outbound

This paper cites Consequently, the training procedure cannot directly determine whether the teacher is capable of solving a given problem.

Distilled Reinforcement Learning for LLM Post-training Consequently, the training procedure cannot directly determine whether the teacher is capable of solving a given problem

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-01T18:39:41.299106Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T18:39:41.299106Z digest=sha256:a2d6f4bf59c6e6f2650f3069ec1fab77955b705c333e798cd92e8e422959594c

Observation 74263fe0-6d6e-4c38-9125-3c5ac7ae9bf6 · outbound

This paper cites an unresolved cited work.

Distilled Reinforcement Learning for LLM Post-training Unresolved cited work

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-01T18:39:41.371264Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T18:39:41.371264Z digest=sha256:d6a1ef4b26a65923fb705d066a953928463df242abf53d8a4b6c537c5c2c897b

Observation a1d7b461-44bc-4f7a-b38b-99ed442eba31 · outbound

This paper cites Entropy-Aware On-Policy Distillation of Language Models.

Distilled Reinforcement Learning for LLM Post-training Entropy-Aware On-Policy Distillation of Language Models

Reference 2015

Resolution
unresolved
no resolver link, observed 2026-08-01T18:39:39.531202Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T18:39:39.531202Z digest=sha256:833a899c2d59fb825ad2364d40548771c854984fdc1b71737e0d6d8f4318d69e

Observation e6b2c343-f33b-4bcb-8a45-fe710ceef832 · outbound

This paper cites DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models.

Distilled Reinforcement Learning for LLM Post-training DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models

Reference 2017

Resolution
unresolved
no resolver link, observed 2026-08-01T18:39:40.211343Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T18:39:40.211343Z digest=sha256:512bda3bf39d7e926ad80092ca7bd0e042e99dbfa0b5b6d17cbb2bc5931b5bee

Observation 20a53cbf-ad2a-43d3-a965-8e0fa4be83f1 · outbound

This paper cites Process Reinforcement through Implicit Rewards.

Distilled Reinforcement Learning for LLM Post-training Process Reinforcement through Implicit Rewards

Reference 2021

Resolution
unresolved
no resolver link, observed 2026-08-01T18:39:39.130446Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T18:39:39.130446Z digest=sha256:693ac7abb57da32a8edbfc8915237a93836f993e4734a0ee7b9a620619966308

Observation 406fe001-ec82-45d6-8d51-a99bcb72e033 · outbound

This paper cites Rethinking On-Policy Distillation of Large Language Models: Phenomenology, Mechanism, and Recipe.

Distilled Reinforcement Learning for LLM Post-training Rethinking On-Policy Distillation of Large Language Models: Phenomenology, Mechanism, and Recipe

Reference 2022

Resolution
unresolved
no resolver link, observed 2026-08-01T18:39:39.657391Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T18:39:39.657391Z digest=sha256:6d950de06b46b0f6b77b6b37bebb6dc99c79ef9f67d26ab275dfdacb917fb762

Observation fcdb8a70-309e-4690-9fbb-1dfee0c572ca · outbound

This paper cites DeepSeek-V3 Technical Report.

Distilled Reinforcement Learning for LLM Post-training DeepSeek-V3 Technical Report

Reference 2023

Resolution
unresolved
no resolver link, observed 2026-08-01T18:39:39.919563Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T18:39:39.919563Z digest=sha256:1ee235fe927a6e5ead37a7db1dd25b260f05588db05134e1ccd26c8eaba1b2d0

Observation a927e4a9-c254-4971-943a-61e6a14053fb · outbound

This paper cites MathArena: Evaluating LLMs on Uncontaminated Math Competitions.

Distilled Reinforcement Learning for LLM Post-training MathArena: Evaluating LLMs on Uncontaminated Math Competitions

Reference 2024

Resolution
unresolved
no resolver link, observed 2026-08-01T18:39:38.885385Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T18:39:38.885385Z digest=sha256:f95e1935611304bfa1ed297b51b362388e7731e55c0721bffad1def3a41b79ee

Observation b401463d-7b58-4d3d-9f8f-8c43135e25df · outbound

This paper cites MathArena: Evaluating LLMs on Uncontaminated Math Competitions.

Distilled Reinforcement Learning for LLM Post-training MathArena: Evaluating LLMs on Uncontaminated Math Competitions

Reference 2025

Resolution
unresolved
no resolver link, observed 2026-08-01T18:39:38.938706Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T18:39:38.938706Z digest=sha256:54ccc8ccff6edf4669555692554956f1ea4397f67ba4bd445a41517e3fe7f0df

Observation 99b7d37f-60a2-4deb-9aa7-6557f3e84074 · outbound

This paper cites Training Verifiers to Solve Math Word Problems.

Distilled Reinforcement Learning for LLM Post-training Training Verifiers to Solve Math Word Problems

Reference 2026

Resolution
unresolved
no resolver link, observed 2026-08-01T18:39:39.071891Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T18:39:39.071891Z digest=sha256:72d32087d33b5c0c971bc4522026ffa5f8f96e36b80bce99092c54a38ff932da

Pith citing papers

No inbound Pith citation observations are available.