Pith. sign in

Paper Citation Record · LEDGER

Distilled Reinforcement Learning for LLM Post-training

As of 14 August 2026, this Paper Citation Record lists 37 of 37 outbound references and 0 inbound Pith citation observations for arXiv:2607.17247.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2607.17247 v1

Coverage vector

measured 37 of 37 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-01T18:39:41.431601Z

measured 37 of 37 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-14T06:32:32.682623+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

37 of 37 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved37
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation d3639001-7ee8-425c-8e05-8ff8b063389a · outbound

This paper cites On-policy distillation of language models: Learning from self- generated mistakes.

Distilled Reinforcement Learning for LLM Post-training On-policy distillation of language models: Learning from self- generated mistakes

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-01T18:39:38.757408Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T18:39:38.757408Z digest=sha256:c9cec5fe47f3a91583a6085044065ef5ddc9d8a95953d88b5676f9e73766d052

Observation 66aa2a9b-e845-4840-8ca0-79cba8c1e59a · outbound

This paper cites Reasoning with Exploration: An Entropy Perspective.

Distilled Reinforcement Learning for LLM Post-training Reasoning with Exploration: An Entropy Perspective

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-01T18:39:39.011321Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T18:39:39.011321Z digest=sha256:b528285140a1f1d26b7ab74fff61cdc82ad6692f0c758f9cfe9d5f5cc729d64e

Observation f747b952-059c-4908-b455-ed0d5199652c · outbound

This paper cites an unresolved cited work.

Distilled Reinforcement Learning for LLM Post-training Unresolved cited work

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-01T18:39:41.431601Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T18:39:41.431601Z digest=sha256:d0feb54929211dfff5a88e54e91ff16dd91005ecfac14c6ec86a9aac56a439af

Observation f308f626-b3e0-4513-abfa-ec7e4c1916f1 · outbound

This paper cites Revisiting On-Policy Distillation: Empirical Failure Modes and Simple Fixes.

Distilled Reinforcement Learning for LLM Post-training Revisiting On-Policy Distillation: Empirical Failure Modes and Simple Fixes

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-01T18:39:39.225995Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T18:39:39.225995Z digest=sha256:bcdd7284db39bad8c6f9c259c12f2a854c50c5e2b2740733d341947fd52ec9f2

Observation 77c158ab-115d-45df-a6df-efabde58cab0 · outbound

This paper cites ChatGLM: A Family of Large Language Models from GLM-130B to GLM-4 All Tools.

Distilled Reinforcement Learning for LLM Post-training ChatGLM: A Family of Large Language Models from GLM-130B to GLM-4 All Tools

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-01T18:39:39.287381Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T18:39:39.287381Z digest=sha256:e61bea86f523aac0faf25765f931691e76b623931c03be34b198fb2ac8c42f48

Observation 90a82ab3-edb9-4610-a60c-e8a8955ea2aa · outbound

This paper cites Minillm: Knowledge distillation of large language models.

Distilled Reinforcement Learning for LLM Post-training Minillm: Knowledge distillation of large language models

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-01T18:39:39.349107Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T18:39:39.349107Z digest=sha256:9b19292d950f785464ad05988399f5615f68cc5eb6ba0129ecc83c7d3900b9a6

Observation 41a80cd6-60be-4fa5-8f4f-2fd9e34166d0 · outbound

This paper cites DeepSeek-Coder: When the Large Language Model Meets Programming -- The Rise of Code Intelligence.

Distilled Reinforcement Learning for LLM Post-training DeepSeek-Coder: When the Large Language Model Meets Programming -- The Rise of Code Intelligence

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-01T18:39:39.387426Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T18:39:39.387426Z digest=sha256:15eb8d91d7a9b8ef26fa55172da61524ccd8ffe50c47de9e1e848972076368fe

Observation 4270e065-4d31-421d-975b-a6971cb08ef9 · outbound

This paper cites DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning.

Distilled Reinforcement Learning for LLM Post-training DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-01T18:39:39.453355Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T18:39:39.453355Z digest=sha256:913d10e5f815a802b4ac7b5c946ba506ad1225e20c271d6297c04036694cfb94

Observation 9e52171a-0e32-40c1-80d6-66c8516014d0 · outbound

This paper cites Sequence-level knowledge distillation.

Distilled Reinforcement Learning for LLM Post-training Sequence-level knowledge distillation

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-01T18:39:39.613393Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T18:39:39.613393Z digest=sha256:19f3584a4aa4823a8dab7667dffcfc9ecc0a53edbb00431eac98faf0fb71c116

Observation 5d2a8488-18fd-4831-9eac-02fb84f50a6f · outbound

This paper cites TreePO: Bridging the Gap of Policy Optimization and Efficacy and Inference Efficiency with Heuristic Tree-based Modeling.

Distilled Reinforcement Learning for LLM Post-training TreePO: Bridging the Gap of Policy Optimization and Efficacy and Inference Efficiency with Heuristic Tree-based Modeling

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-01T18:39:39.765212Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T18:39:39.765212Z digest=sha256:d6a2ae62924936ab316fa00135fdda27237a506b063b7a09b974ac673590c166

Observation d66b7e15-bd33-47c8-b307-b7f66b3528e6 · outbound

This paper cites Let's Verify Step by Step.

Distilled Reinforcement Learning for LLM Post-training Let's Verify Step by Step

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-01T18:39:39.857107Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T18:39:39.857107Z digest=sha256:6d1044416648240e59ba6d1b205048bc4b30f5e44e4ae3cc58d967eda19b93d3

Observation 25e39112-4a92-4e04-897e-81a633265030 · outbound

This paper cites Length-unbiased sequence policy optimization: Revealing and controlling response length variation in rlvr.arXiv preprint arXiv:2602.05261,.

Distilled Reinforcement Learning for LLM Post-training Length-unbiased sequence policy optimization: Revealing and controlling response length variation in rlvr.arXiv preprint arXiv:2602.05261,

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-01T18:39:39.976099Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T18:39:39.976099Z digest=sha256:fb9b693d02a405a2ca0296a785aabcb8b783eb740f1c71ddae7f684e66043a04

Observation 0a52b37e-3c9d-4f12-9be6-b2f196349785 · outbound

This paper cites Improve Mathematical Reasoning in Language Models by Automated Process Supervision.

Distilled Reinforcement Learning for LLM Post-training Improve Mathematical Reasoning in Language Models by Automated Process Supervision

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-01T18:39:40.059010Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T18:39:40.059010Z digest=sha256:580f83cc9149484a0493637e83c3edbaff30e98c7d52bde72c7be6d8c9c9e8ae

Observation b8d209c0-ab3f-41e1-81fa-19c0cc639450 · outbound

This paper cites Proximal Policy Optimization Algorithms.

Distilled Reinforcement Learning for LLM Post-training Proximal Policy Optimization Algorithms

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-01T18:39:40.146868Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T18:39:40.146868Z digest=sha256:0eaeb67f563f028cd27db4eba5f3f53277179edbc0073434d3e9fc7a16239994

Observation 35476f98-1e81-4179-a447-6f7e5039932d · outbound

This paper cites Kimi k1.5: Scaling Reinforcement Learning with LLMs.

Distilled Reinforcement Learning for LLM Post-training Kimi k1.5: Scaling Reinforcement Learning with LLMs

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-01T18:39:40.268643Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T18:39:40.268643Z digest=sha256:02c614243eac2cf5ade1726686cafb4972335c3d0dc34c358b4264139bbd2d85

Observation e4da1c71-9642-46c4-8a6e-806cf355410e · outbound

This paper cites Llama 2: Open Foundation and Fine-Tuned Chat Models.

Distilled Reinforcement Learning for LLM Post-training Llama 2: Open Foundation and Fine-Tuned Chat Models

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-01T18:39:40.318444Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T18:39:40.318444Z digest=sha256:e5dd6875bc8259115d704a0120a9c3418e761688919fcc06ad44be8bc604e729

Observation ac733956-c1a5-465e-a790-f3fbbf16fb9a · outbound

This paper cites SCOPE-RL: Stable and Quantitative Control of Policy Entropy in RL Post-Training.

Distilled Reinforcement Learning for LLM Post-training SCOPE-RL: Stable and Quantitative Control of Policy Entropy in RL Post-Training

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-01T18:39:40.386157Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T18:39:40.386157Z digest=sha256:9dbe26c97731543a1d337b552cfbd90750f88126b3a8a7bca024893425242db4

Observation ecf09c03-2927-4a36-9243-d9819033cdab · outbound

This paper cites Implicit Compression Regularization: Concise Reasoning via Internal Shorter Distributions in RL Post-Training.

Distilled Reinforcement Learning for LLM Post-training Implicit Compression Regularization: Concise Reasoning via Internal Shorter Distributions in RL Post-Training

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-01T18:39:40.459071Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T18:39:40.459071Z digest=sha256:70553ab0dca8cb0ecb04499af6e84dfa1c9d38f040d922cdcb4abe8bc4ff891c

Observation 823550af-5e8d-4066-89dc-f6741deba2b2 · outbound

This paper cites Lightning OPD: Efficient Post-Training for Large Reasoning Models with Offline On-Policy Distillation.

Distilled Reinforcement Learning for LLM Post-training Lightning OPD: Efficient Post-Training for Large Reasoning Models with Offline On-Policy Distillation

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-01T18:39:40.517367Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T18:39:40.517367Z digest=sha256:c3e4b461f6158a18719a9c30dd838cfbe9f01780cef450ffbd6080c2e158f519

Observation e7b7c17c-c775-4d1a-b2e9-a3a38fbc143d · outbound

This paper cites On the generalization of sft: A reinforcement learning perspective with reward rectification.arXiv preprint arXiv:2508.05629, 2025b.

Distilled Reinforcement Learning for LLM Post-training On the generalization of sft: A reinforcement learning perspective with reward rectification.arXiv preprint arXiv:2508.05629, 2025b

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-01T18:39:40.619539Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T18:39:40.619539Z digest=sha256:9efbd11b1d6f95f983b1694610606f41debdf76432b8f33f3db51c0753d47dc6

Observation 723986cb-e5e4-4411-94a7-bf89bd911262 · outbound

This paper cites Qwen2.5 Technical Report.

Distilled Reinforcement Learning for LLM Post-training Qwen2.5 Technical Report

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-01T18:39:40.700245Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T18:39:40.700245Z digest=sha256:b154dedef2376c670aa0d8661100f344eead7049a1e7c74fa4f65a57540ad544

Observation 3a83a283-4a92-4c3b-82db-f4231eb2b2c1 · outbound

This paper cites Qwen3 Technical Report.

Distilled Reinforcement Learning for LLM Post-training Qwen3 Technical Report

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-01T18:39:40.772979Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T18:39:40.772979Z digest=sha256:0eb5a3c04041193697a9118d142a8b89280b0746a4048b7bd6b44625f6c949de

Observation 8acea5ca-4c80-4681-bac4-47206eb3b3dd · outbound

This paper cites Learning beyond Teacher: Generalized On-Policy Distillation with Reward Extrapolation.

Distilled Reinforcement Learning for LLM Post-training Learning beyond Teacher: Generalized On-Policy Distillation with Reward Extrapolation

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-01T18:39:40.846363Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T18:39:40.846363Z digest=sha256:43ab67ed59c583957fb16e0f107f633516fee63ba4d7e53c11add1e20f876168

Observation cb87e227-e96b-448c-96c4-342c584747ee · outbound

This paper cites Shorterbetter: Guiding reasoning models to find optimal inference length for efficient reasoning.arXiv preprint arXiv:2504.21370,.

Distilled Reinforcement Learning for LLM Post-training Shorterbetter: Guiding reasoning models to find optimal inference length for efficient reasoning.arXiv preprint arXiv:2504.21370,

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-01T18:39:40.916252Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T18:39:40.916252Z digest=sha256:ee4d59516ef32be77211b2e839c14debe3ab50cf2db5b45ff7b0d9f26a61ceba

Observation b9216b88-e480-4cc5-add2-b25348642673 · outbound

This paper cites DAPO: An Open-Source LLM Reinforcement Learning System at Scale.

Distilled Reinforcement Learning for LLM Post-training DAPO: An Open-Source LLM Reinforcement Learning System at Scale

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-01T18:39:41.004081Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T18:39:41.004081Z digest=sha256:6a080bfed44ed7579fd44880acf83bbfe165db0d16d8c22cbe0c30e23d2f81ed

Observation 866ed5c2-5541-4e61-b98d-380337bf4304 · outbound

This paper cites Does Reinforcement Learning Really Incentivize Reasoning Capacity in LLMs Beyond the Base Model?.

Distilled Reinforcement Learning for LLM Post-training Does Reinforcement Learning Really Incentivize Reasoning Capacity in LLMs Beyond the Base Model?

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-01T18:39:41.081442Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T18:39:41.081442Z digest=sha256:68d01ba5bac32ed40409c39f2037e0410af3ce266e058c01cfd861246d32d1bb

Observation ee81150d-a934-48b1-8599-f78542a401ea · outbound

This paper cites DPO Meets PPO: Reinforced Token Optimization for RLHF.

Distilled Reinforcement Learning for LLM Post-training DPO Meets PPO: Reinforced Token Optimization for RLHF

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-01T18:39:41.185696Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T18:39:41.185696Z digest=sha256:168d77e166810f7bca784b8d294f8ee19ac3d10e06949d1b02110816518a4452

Observation 419e3584-e959-4a1f-b782-c0a758894060 · outbound

This paper cites Consequently, the training procedure cannot directly determine whether the teacher is capable of solving a given problem.

Distilled Reinforcement Learning for LLM Post-training Consequently, the training procedure cannot directly determine whether the teacher is capable of solving a given problem

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-01T18:39:41.299106Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T18:39:41.299106Z digest=sha256:50a2fbb7dfd1aa014a53cae83492642c582377f6e75de3a4ffaf6a311272e913

Observation 74263fe0-6d6e-4c38-9125-3c5ac7ae9bf6 · outbound

This paper cites an unresolved cited work.

Distilled Reinforcement Learning for LLM Post-training Unresolved cited work

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-01T18:39:41.371264Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T18:39:41.371264Z digest=sha256:a772e619ec98faebb836f4429ab0fd195da05208c28ea52bcf4d52196c873736

Observation a1d7b461-44bc-4f7a-b38b-99ed442eba31 · outbound

This paper cites Entropy-Aware On-Policy Distillation of Language Models.

Distilled Reinforcement Learning for LLM Post-training Entropy-Aware On-Policy Distillation of Language Models

Reference 2015

Resolution
unresolved
no resolver link, observed 2026-08-01T18:39:39.531202Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T18:39:39.531202Z digest=sha256:0b04f545e8c9baebf5819244bb1278ba7ed2ceac862d5cf83c3623e429f5514d

Observation e6b2c343-f33b-4bcb-8a45-fe710ceef832 · outbound

This paper cites DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models.

Distilled Reinforcement Learning for LLM Post-training DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models

Reference 2017

Resolution
unresolved
no resolver link, observed 2026-08-01T18:39:40.211343Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T18:39:40.211343Z digest=sha256:970bd7ee993b55e5ada4dc00a6fb03eb87b40fa067f4f39d6300627f9c7fc866

Observation 20a53cbf-ad2a-43d3-a965-8e0fa4be83f1 · outbound

This paper cites Process Reinforcement through Implicit Rewards.

Distilled Reinforcement Learning for LLM Post-training Process Reinforcement through Implicit Rewards

Reference 2021

Resolution
unresolved
no resolver link, observed 2026-08-01T18:39:39.130446Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T18:39:39.130446Z digest=sha256:ec6d715ced0927dc7b89f008ee409d15fa57e8b451f7687572b9b4af3860d3ae

Observation 406fe001-ec82-45d6-8d51-a99bcb72e033 · outbound

This paper cites Rethinking On-Policy Distillation of Large Language Models: Phenomenology, Mechanism, and Recipe.

Distilled Reinforcement Learning for LLM Post-training Rethinking On-Policy Distillation of Large Language Models: Phenomenology, Mechanism, and Recipe

Reference 2022

Resolution
unresolved
no resolver link, observed 2026-08-01T18:39:39.657391Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T18:39:39.657391Z digest=sha256:2776bde0e9aabcae3561f847b5137269c18e575cc484ab06cdbc4961b7b83a29

Observation fcdb8a70-309e-4690-9fbb-1dfee0c572ca · outbound

This paper cites DeepSeek-V3 Technical Report.

Distilled Reinforcement Learning for LLM Post-training DeepSeek-V3 Technical Report

Reference 2023

Resolution
unresolved
no resolver link, observed 2026-08-01T18:39:39.919563Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T18:39:39.919563Z digest=sha256:f05813917cc3d5ee37fa411513e43e4d2fdb4122f5edba29b6ef495ad7f154d8

Observation a927e4a9-c254-4971-943a-61e6a14053fb · outbound

This paper cites MathArena: Evaluating LLMs on Uncontaminated Math Competitions.

Distilled Reinforcement Learning for LLM Post-training MathArena: Evaluating LLMs on Uncontaminated Math Competitions

Reference 2024

Resolution
unresolved
no resolver link, observed 2026-08-01T18:39:38.885385Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T18:39:38.885385Z digest=sha256:73b4bc5802ef038204b9b8a45adfe0d9284be1c8cd49c5459bc768c0583b73cf

Observation b401463d-7b58-4d3d-9f8f-8c43135e25df · outbound

This paper cites MathArena: Evaluating LLMs on Uncontaminated Math Competitions.

Distilled Reinforcement Learning for LLM Post-training MathArena: Evaluating LLMs on Uncontaminated Math Competitions

Reference 2025

Resolution
unresolved
no resolver link, observed 2026-08-01T18:39:38.938706Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T18:39:38.938706Z digest=sha256:a50c4f2aabec6ab6573ffb1328a5ea58d9813876fcd12ea004fe2448bdc8d848

Observation 99b7d37f-60a2-4deb-9aa7-6557f3e84074 · outbound

This paper cites Training Verifiers to Solve Math Word Problems.

Distilled Reinforcement Learning for LLM Post-training Training Verifiers to Solve Math Word Problems

Reference 2026

Resolution
unresolved
no resolver link, observed 2026-08-01T18:39:39.071891Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T18:39:39.071891Z digest=sha256:72f67f739f05ecbb294577cb851fdf2bb221732aeb21cbbbc8ce397899a6cfbc

Pith citing papers

No inbound Pith citation observations are available.