Pith. sign in

Paper Citation Record · LEDGER

ReST-RL: Achieving Accurate Code Reasoning of LLMs with Optimized Self-Training and Decoding

As of 19 August 2026, this Paper Citation Record lists 39 of 39 outbound references and 4 inbound Pith citation observations for arXiv:2508.19576.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2508.19576 v2

Coverage vector

measured 39 of 39 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-05T15:45:25.874749Z

measured 43 of 43 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-19T06:32:44.657259+00:00

measured 4 of 4 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-15T14:38:34.436702Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-03T14:08:21.670108Z

Reference resolution

39 of 39 outbound references displayed

  • verified exact0
  • verified fuzzy11
  • unresolved28
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation a60a4c92-264c-447b-8b38-be13e1b5183b · outbound

This paper cites Gpt-4 technical report, 2024.

ReST-RL: Achieving Accurate Code Reasoning of LLMs with Optimized Self-Training and Decoding Gpt-4 technical report, 2024

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-05T15:45:24.345030Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T15:45:24.345030Z digest=sha256:0fa08d60e250f63bbcca6c499ea180a1af91be32bedc171ac22282f7d4b0dbb4

Observation bd05cd66-c6bb-414c-a6d5-e445557a6857 · outbound

This paper cites an unresolved cited work.

ReST-RL: Achieving Accurate Code Reasoning of LLMs with Optimized Self-Training and Decoding Unresolved cited work

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-05T15:45:24.394829Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T15:45:24.394829Z digest=sha256:11fe376d3c8c906de854125712b34bb02da24122a1e9cbcf8f1c6c0c54aa50d0

Observation 0464326c-2220-4f6c-909e-d1c51896f0ff · outbound

This paper cites The lessons of developing process reward models in mathematical reasoning, 2025.

ReST-RL: Achieving Accurate Code Reasoning of LLMs with Optimized Self-Training and Decoding The lessons of developing process reward models in mathematical reasoning, 2025

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T15:45:30.424911Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-05T15:45:24.445026Z digest=sha256:fcdb91dea33cd23baffb9ef4e087601d5acdbf87c892e1cc6cb67721ec263de1

Observation 285765fa-6c82-49c7-89c2-f592990cd7fd · outbound

This paper cites Measuring coding challenge competence with apps, 2021.

ReST-RL: Achieving Accurate Code Reasoning of LLMs with Optimized Self-Training and Decoding Measuring coding challenge competence with apps, 2021

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-05T15:45:24.466089Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T15:45:24.466089Z digest=sha256:39a6fd74f8a30bac47a5ba981aa25ebed31974a3a0e3e88dee77f43c6b914b3b

Observation b14037e4-5565-4d98-8b12-1f5a4cf810d1 · outbound

This paper cites Bigcodebench: Benchmarking code generation with diverse function calls and complex instructions, 2025.

ReST-RL: Achieving Accurate Code Reasoning of LLMs with Optimized Self-Training and Decoding Bigcodebench: Benchmarking code generation with diverse function calls and complex instructions, 2025

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T15:45:30.176357Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-05T15:45:24.507007Z digest=sha256:75a3b465b99221f9963efb9d517c679af1540f77d74d47028beb5ebdc32fb5f2

Observation 6de4234e-f4a4-43c3-99af-0c1adfb6f309 · outbound

This paper cites Wizardmath: Empowering mathematical reasoning for large language models via reinforced evol-instruct, 2025.

ReST-RL: Achieving Accurate Code Reasoning of LLMs with Optimized Self-Training and Decoding Wizardmath: Empowering mathematical reasoning for large language models via reinforced evol-instruct, 2025

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-05T15:45:24.557391Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T15:45:24.557391Z digest=sha256:46c3e5a94cd3969a8a4ebb38eb318eb6756d0bc2d6ef8f96071e7202cfc34b65

Observation ab22b356-6564-4fd2-9f72-eacc65650263 · outbound

This paper cites an unresolved cited work.

ReST-RL: Achieving Accurate Code Reasoning of LLMs with Optimized Self-Training and Decoding Unresolved cited work

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-05T15:45:24.606921Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T15:45:24.606921Z digest=sha256:996424c53093a4d1b400fbdcff51eedae5c3abb60e8fe1a624c7474160f2006d

Observation 92664116-49b8-468f-9888-29660062a672 · outbound

This paper cites Proximal policy optimization algorithms, 2017.

ReST-RL: Achieving Accurate Code Reasoning of LLMs with Optimized Self-Training and Decoding Proximal policy optimization algorithms, 2017

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T15:45:29.850809Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-05T15:45:24.634749Z digest=sha256:6dda76000cfc99647ca9c3959971835e920030c738c4d05124ba3fc06b0fd5ff

Observation b633ee1d-19fb-4e55-bc88-399855cef95b · outbound

This paper cites Reinforced self-training (rest) for language modeling, 2023.

ReST-RL: Achieving Accurate Code Reasoning of LLMs with Optimized Self-Training and Decoding Reinforced self-training (rest) for language modeling, 2023

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-05T15:45:24.665314Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T15:45:24.665314Z digest=sha256:aa10294e1563ee06cccfd4823ef29d265f58e0ad25b2fc26b6516ed3cba34d88

Observation 8b07e88e-f82b-4015-879d-84ab74b7dec5 · outbound

This paper cites Direct preference optimization: Your language model is secretly a reward model.

ReST-RL: Achieving Accurate Code Reasoning of LLMs with Optimized Self-Training and Decoding Direct preference optimization: Your language model is secretly a reward model

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-05T15:45:24.691732Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T15:45:24.691732Z digest=sha256:88c427bb12f2b8f2f7dd9efe1852928f95dad367e4a763aa02e3b3fafed3b29e

Observation 7ac86a5a-13e6-422a-a657-8efefc8693bb · outbound

This paper cites an unresolved cited work.

ReST-RL: Achieving Accurate Code Reasoning of LLMs with Optimized Self-Training and Decoding Unresolved cited work

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-05T15:45:24.715154Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T15:45:24.715154Z digest=sha256:b69aa18cc1e7899488f4a7203b743d656bbb9f81c6219f7cbb1c6adb57d94793

Observation 87b836da-7df4-457d-a007-3dbb733fc341 · outbound

This paper cites Beyond Human Data: Scaling Self-Training for Problem-Solving with Language Models.

ReST-RL: Achieving Accurate Code Reasoning of LLMs with Optimized Self-Training and Decoding Beyond Human Data: Scaling Self-Training for Problem-Solving with Language Models

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-05T15:45:24.746306Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T15:45:24.746306Z digest=sha256:7e06bb029208b4b2fafecced32198bcf74d33f6269706b495a6d396945747624

Observation 395e6487-9a00-42b5-816e-1352a9c5a287 · outbound

This paper cites Rest-mcts*: Llm self-training via process reward guided tree search, 2024.

ReST-RL: Achieving Accurate Code Reasoning of LLMs with Optimized Self-Training and Decoding Rest-mcts*: Llm self-training via process reward guided tree search, 2024

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-05T15:45:24.772773Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T15:45:24.772773Z digest=sha256:cf356371bfc3182e8e352f700ef3723946b536681ca316b3255fb9967e415748

Observation 26a75e80-eec5-40d5-bc22-5b826471279b · outbound

This paper cites Solving math word problems with process- and outcome-based feedback, 2022.

ReST-RL: Achieving Accurate Code Reasoning of LLMs with Optimized Self-Training and Decoding Solving math word problems with process- and outcome-based feedback, 2022

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-05T15:45:24.807225Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T15:45:24.807225Z digest=sha256:3c5b248595c5fe64eee5639b0c239b7d0336309389ef5200af57b62f4fa7fa17

Observation cf8b6baa-1348-419d-a785-65e77a65bd9a · outbound

This paper cites Let’s verify step by step, 2023.

ReST-RL: Achieving Accurate Code Reasoning of LLMs with Optimized Self-Training and Decoding Let’s verify step by step, 2023

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-05T15:45:24.854247Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T15:45:24.854247Z digest=sha256:17090f527cd78d60962d0288bb59f4724c6b89da37b6a63c136bcf085312e0a3

Observation b2298171-678a-429f-b3ab-d93bede8b177 · outbound

This paper cites Don’t throw away your value model! generating more preferable text with value-guided monte-carlo tree search decoding, 2024.

ReST-RL: Achieving Accurate Code Reasoning of LLMs with Optimized Self-Training and Decoding Don’t throw away your value model! generating more preferable text with value-guided monte-carlo tree search decoding, 2024

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T15:45:29.124983Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-05T15:45:24.894475Z digest=sha256:36ce59733b71688e14e885849f14be4cffc5406678b40302f48b7e0665b364bd

Observation 7d9bf68a-649b-4cf0-9b89-28b967dd121c · outbound

This paper cites Dapo: An open-source llm reinforcement learning system at scale, 2025.

ReST-RL: Achieving Accurate Code Reasoning of LLMs with Optimized Self-Training and Decoding Dapo: An open-source llm reinforcement learning system at scale, 2025

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-05T15:45:24.908108Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T15:45:24.908108Z digest=sha256:747bd26573aba9474cb102e27d198aa8362627b74dbf9b0c6b58d9ef52eb84d4

Observation d026e5d3-7815-40cd-b6c9-57440a7e51fd · outbound

This paper cites R1-vl: Learning to reason with multimodal large language models via step-wise group relative policy optimization, 2025.

ReST-RL: Achieving Accurate Code Reasoning of LLMs with Optimized Self-Training and Decoding R1-vl: Learning to reason with multimodal large language models via step-wise group relative policy optimization, 2025

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T15:45:28.884827Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-05T15:45:24.935044Z digest=sha256:202b288a4ac6b4d9d0392795b6e7dedbb054544545b268bf4131f5456b98604c

Observation 67be6a1d-670a-4911-b0be-c06630bba609 · outbound

This paper cites an unresolved cited work.

ReST-RL: Achieving Accurate Code Reasoning of LLMs with Optimized Self-Training and Decoding Unresolved cited work

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-05T15:45:24.974749Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T15:45:24.974749Z digest=sha256:e7bbdb4800b6e5f7bd0ac7b3838dee3dd2e07d86cdaad955699954de08f7c4c0

Observation 4483e15b-d22a-450a-ae42-64843050f91f · outbound

This paper cites Mankowitz, Esme Sutherland Robson, Pushmeet Kohli, Nando de Freitas, Koray Kavukcuoglu, and Oriol Vinyals.

ReST-RL: Achieving Accurate Code Reasoning of LLMs with Optimized Self-Training and Decoding Mankowitz, Esme Sutherland Robson, Pushmeet Kohli, Nando de Freitas, Koray Kavukcuoglu, and Oriol Vinyals

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T15:45:28.654742Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-05T15:45:25.024763Z digest=sha256:d115f13234ffe77c4f3c385d50bbe7b1d9d4774c236454df61f93744153b94a6

Observation 7426c0c7-f201-4fe8-9947-e31e6a755f36 · outbound

This paper cites Code with codeqwen1.5, April 2024.

ReST-RL: Achieving Accurate Code Reasoning of LLMs with Optimized Self-Training and Decoding Code with codeqwen1.5, April 2024

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T15:45:28.494850Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-05T15:45:25.057163Z digest=sha256:aa23819250ae09c6dd1c9a9fa460122427ed0c8d825cfdf062ac1f890e71263b

Observation e00ea81a-0836-4ae5-9e35-a39b47ab048e · outbound

This paper cites OpenCodeInterpreter: Integrating Code Generation with Execution and Refinement.

ReST-RL: Achieving Accurate Code Reasoning of LLMs with Optimized Self-Training and Decoding OpenCodeInterpreter: Integrating Code Generation with Execution and Refinement

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-05T15:45:25.104749Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T15:45:25.104749Z digest=sha256:61c9044debafa9ffec54816f0cf8467fcc85a30bcfe6dd5dfc44d3c1e70dd491

Observation ced27b90-3a83-484a-a60f-dcb6b9027441 · outbound

This paper cites Tenenbaum, and Chuang Gan.

ReST-RL: Achieving Accurate Code Reasoning of LLMs with Optimized Self-Training and Decoding Tenenbaum, and Chuang Gan

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T15:45:28.355103Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-05T15:45:25.145037Z digest=sha256:fa0344bf0a859828c6cea242c26ed2a5c082c2f636d84b87f91d3f8666823b73

Observation 92ee3d93-802c-4791-8843-9a1f350c15cd · outbound

This paper cites Jiang, Jia Deng, Stella Biderman, and Sean Welleck.

ReST-RL: Achieving Accurate Code Reasoning of LLMs with Optimized Self-Training and Decoding Jiang, Jia Deng, Stella Biderman, and Sean Welleck

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T15:45:28.224747Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-05T15:45:25.195191Z digest=sha256:68cb2e8da04215a059f2976d99348ffa9d1f23f2dfe820412ab8cb29fb8a9639

Observation db52f004-bdc6-46db-98bf-c841ce605ba3 · outbound

This paper cites an unresolved cited work.

ReST-RL: Achieving Accurate Code Reasoning of LLMs with Optimized Self-Training and Decoding Unresolved cited work

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-05T15:45:25.245194Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T15:45:25.245194Z digest=sha256:68a6c52fcdf2349f0cc8647446fc05cb5a002193d7c952ee4fcd012ad5f0b6d7

Observation efde76ed-bddc-498b-bb3a-9ca6434be0e5 · outbound

This paper cites Pspo*: An effective process-supervised policy optimization for reasoning alignment, 2024.

ReST-RL: Achieving Accurate Code Reasoning of LLMs with Optimized Self-Training and Decoding Pspo*: An effective process-supervised policy optimization for reasoning alignment, 2024

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T15:45:27.986574Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-05T15:45:25.284748Z digest=sha256:8e6fc586775ac438f198b008791068b58ad728f839aaf24d13c2ebe9e643d043

Observation 4dff5136-fc75-4f6a-8de6-66ce1bea5707 · outbound

This paper cites an unresolved cited work.

ReST-RL: Achieving Accurate Code Reasoning of LLMs with Optimized Self-Training and Decoding Unresolved cited work

Reference 27

Resolution
unresolved
raw_fallback, observed 2026-08-05T15:45:27.674899Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-05T15:45:25.334751Z digest=sha256:b8e486697b285b06b7d3de39ca55fd75d3c9ff2fa9d68ccfcf9686fe3ec3319b

Observation 633b31e6-0ab6-414b-aa06-415a3df86dc0 · outbound

This paper cites Chain-of-thought prompting elicits reasoning in large language models, 2023.

ReST-RL: Achieving Accurate Code Reasoning of LLMs with Optimized Self-Training and Decoding Chain-of-thought prompting elicits reasoning in large language models, 2023

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-05T15:45:25.384748Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T15:45:25.384748Z digest=sha256:543cd4192ca1c2e8db5db87ef8b2d186b173a76fa09414b8376bb88ebad375a0

Observation 3afc0a37-acd2-4b08-8782-f6ee69c42127 · outbound

This paper cites Griffiths, Yuan Cao, and Karthik Narasimhan.

ReST-RL: Achieving Accurate Code Reasoning of LLMs with Optimized Self-Training and Decoding Griffiths, Yuan Cao, and Karthik Narasimhan

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-05T15:45:25.434824Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T15:45:25.434824Z digest=sha256:957c1d37edee7768bc123e0a02f370abdcabca800535fe467f76d073b2f196d5

Observation b9d8e6d4-9b13-4c45-a38e-0ce90ee69bf4 · outbound

This paper cites Sra-mcts: Self-driven reasoning augmentation with monte carlo tree search for code generation, 2025.

ReST-RL: Achieving Accurate Code Reasoning of LLMs with Optimized Self-Training and Decoding Sra-mcts: Self-driven reasoning augmentation with monte carlo tree search for code generation, 2025

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T15:45:27.054841Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-05T15:45:25.475053Z digest=sha256:f46339d8efcfe55afbc46deec335fe5e9d59c55f8de4b86538e77a66d025019f

Observation f14061ea-d91f-4861-b37e-73c7d3245ce1 · outbound

This paper cites Accessing gpt-4 level mathematical olympiad solutions via monte carlo tree self-refine with llama-3 8b, 2024.

ReST-RL: Achieving Accurate Code Reasoning of LLMs with Optimized Self-Training and Decoding Accessing gpt-4 level mathematical olympiad solutions via monte carlo tree self-refine with llama-3 8b, 2024

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-05T15:45:25.524747Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T15:45:25.524747Z digest=sha256:0da50d9a78ec51270623bc4e9599c4cfc216d234453013226da36fffc76f20fc

Observation 283cc7ff-73ce-46b9-a5c6-f315b223cce8 · outbound

This paper cites Program synthesis with large language models, 2021.

ReST-RL: Achieving Accurate Code Reasoning of LLMs with Optimized Self-Training and Decoding Program synthesis with large language models, 2021

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-05T15:45:25.574819Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T15:45:25.574819Z digest=sha256:236baa1bbaeb3c2a96db3140b7f7b257aca409e8b24af2ad5fba14ffdefb551d

Observation 3c654c56-76c3-4b7a-9920-237fa938e6ff · outbound

This paper cites Qwen2.5-Coder Technical Report.

ReST-RL: Achieving Accurate Code Reasoning of LLMs with Optimized Self-Training and Decoding Qwen2.5-Coder Technical Report

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-05T15:45:25.624748Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T15:45:25.624748Z digest=sha256:3aacd01447c1d378d612c31cca18fd0027d7c6bb262087118f251665f834c52f

Observation 629afcb8-3aff-44ef-b900-cd5457763673 · outbound

This paper cites Qwen3 technical report, 2025.

ReST-RL: Achieving Accurate Code Reasoning of LLMs with Optimized Self-Training and Decoding Qwen3 technical report, 2025

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-05T15:45:25.674750Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T15:45:25.674750Z digest=sha256:49bcb8f9df76a791f3522c85fe3c7ab6fec132b4ae8fa4938696a69f845263b5

Observation d069aee4-119a-4538-a5cb-537ee583cb65 · outbound

This paper cites The llama 3 herd of models, 2024.

ReST-RL: Achieving Accurate Code Reasoning of LLMs with Optimized Self-Training and Decoding The llama 3 herd of models, 2024

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-05T15:45:25.717705Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T15:45:25.717705Z digest=sha256:cdf4a69680c97baeaff408c6a3e6c5ff0083f206e717700143d048f2e800e036

Observation 87a5e739-ada5-49ec-b3d4-c7812f70237e · outbound

This paper cites Ds-1000: A natural and reliable benchmark for data science code generation, 2022.

ReST-RL: Achieving Accurate Code Reasoning of LLMs with Optimized Self-Training and Decoding Ds-1000: A natural and reliable benchmark for data science code generation, 2022

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-05T15:45:25.764827Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T15:45:25.764827Z digest=sha256:8bdae41c6f22cbff81313f0df11635358bc1cc8ea2b46414e058a38799b5b486

Observation 0b083143-5493-4f29-95ae-d9af24999b87 · outbound

This paper cites Is your code generated by chatgpt really correct? rigorous evaluation of large language models for code generation, 2023.

ReST-RL: Achieving Accurate Code Reasoning of LLMs with Optimized Self-Training and Decoding Is your code generated by chatgpt really correct? rigorous evaluation of large language models for code generation, 2023

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-05T15:45:25.814825Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T15:45:25.814825Z digest=sha256:faf1170e4eb19541baae0d218861e51cb9b4225996fd7cd74e587e6caee8d486

Observation c177dc3c-5724-40b6-80a9-cea292e5f0cb · outbound

This paper cites Reward hacking in reinforcement learning.lilianweng.github.io, Nov 2024.

ReST-RL: Achieving Accurate Code Reasoning of LLMs with Optimized Self-Training and Decoding Reward hacking in reinforcement learning.lilianweng.github.io, Nov 2024

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-05T15:45:25.854752Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T15:45:25.854752Z digest=sha256:ab12aff4d613bf19b056badd280ff589945f14a81d3d968e8371d80461d847ec

Observation c5d66e67-fb19-425b-a5e7-c3f3d9d21913 · outbound

This paper cites Skywork-Reward: Bag of Tricks for Reward Modeling in LLMs.

ReST-RL: Achieving Accurate Code Reasoning of LLMs with Optimized Self-Training and Decoding Skywork-Reward: Bag of Tricks for Reward Modeling in LLMs

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-05T15:45:25.874749Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T15:45:25.874749Z digest=sha256:d8803a0b391c188e282ed670375c47f51bd3c59f3b0072e782736a10709b623e

Pith citing papers

Observation faac2c7a-0710-4a7d-ae45-b324f3875e95 · inbound

MARS$^2$: Scaling Multi-Agent Tree Search via Reinforcement Learning for Code Generation cites this paper.

MARS$^2$: Scaling Multi-Agent Tree Search via Reinforcement Learning for Code Generation ReST-RL: Achieving Accurate Code Reasoning of LLMs with Optimized Self-Training and Decoding

Reference 35

Resolution
verified exact
arxiv_id, observed 2026-05-10T11:30:18.824443Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-05-10T11:27:28.245835Z digest=sha256:12b2826ff2254891197ac9f32441b509f1333b5c7007b592001e9e3d83ef681f

Observation ecddc6aa-ea6c-481f-914b-ec867be66e78 · inbound

Mental-R1: Aligning LLM Reasoning for Mental Health Assessment cites this paper.

Mental-R1: Aligning LLM Reasoning for Mental Health Assessment ReST-RL: Achieving Accurate Code Reasoning of LLMs with Optimized Self-Training and Decoding

Reference 67

Resolution
verified exact
arxiv_id, observed 2026-07-03T14:08:21.671873Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-06-27T07:15:00.398111Z digest=sha256:1c404711676b79cfc2a13739490de66a55002a21dfa344af3d5c25fb91712893

Observation d6f58a78-07b7-4087-9e46-9547b07c6fd2 · inbound

Continual Learning in Transition cites this paper.

Continual Learning in Transition ReST-RL: Achieving Accurate Code Reasoning of LLMs with Optimized Self-Training and Decoding

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-07T12:23:56.891313Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:23:56.891313Z digest=sha256:82f479357ac59a813912cad2f836173bb97a7afd2d4e2281a7d6ddf5029ea952

Observation 487856c3-96c1-421e-9208-665406c3484f · inbound

Continual Learning in Transition cites this paper.

Continual Learning in Transition ReST-RL: Achieving Accurate Code Reasoning of LLMs with Optimized Self-Training and Decoding

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-15T14:38:34.436702Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T14:38:34.436702Z digest=sha256:796ed3eb189a259e3b0e7659bd99dd990437da4e3e74cede3a67cd2591ef5dd7