Pith. sign in

Paper Citation Record · LEDGER

ReST-RL: Achieving Accurate Code Reasoning of LLMs with Optimized Self-Training and Decoding

As of 10 August 2026, this Paper Citation Record lists 39 of 39 outbound references and 3 inbound Pith citation observations for arXiv:2508.19576.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2508.19576 v2

Coverage vector

measured 39 of 39 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-05T15:45:25.874749Z

measured 42 of 42 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-10T06:31:04.303077+00:00

measured 3 of 3 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-07T12:23:56.891313Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-03T14:08:21.670108Z

Reference resolution

39 of 39 outbound references displayed

  • verified exact0
  • verified fuzzy11
  • unresolved28
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation a60a4c92-264c-447b-8b38-be13e1b5183b · outbound

This paper cites Gpt-4 technical report, 2024.

ReST-RL: Achieving Accurate Code Reasoning of LLMs with Optimized Self-Training and Decoding Gpt-4 technical report, 2024

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-05T15:45:24.345030Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T15:45:24.345030Z digest=sha256:29649b8f12d03739421a8f3c8a05025ac7c749069ac45e8f4508b26e06cc5399

Observation bd05cd66-c6bb-414c-a6d5-e445557a6857 · outbound

This paper cites an unresolved cited work.

ReST-RL: Achieving Accurate Code Reasoning of LLMs with Optimized Self-Training and Decoding Unresolved cited work

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-05T15:45:24.394829Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T15:45:24.394829Z digest=sha256:20dda98932c1992efd9eaa95d7e9977a05165573a2c67cefd3420907a172d6b1

Observation 0464326c-2220-4f6c-909e-d1c51896f0ff · outbound

This paper cites The lessons of developing process reward models in mathematical reasoning, 2025.

ReST-RL: Achieving Accurate Code Reasoning of LLMs with Optimized Self-Training and Decoding The lessons of developing process reward models in mathematical reasoning, 2025

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T15:45:30.424911Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-05T15:45:24.445026Z digest=sha256:77e319293f9b3d7d68d34f276c8388383fd8dd73eaa087c9ffc6b9c231910226

Observation 285765fa-6c82-49c7-89c2-f592990cd7fd · outbound

This paper cites Measuring coding challenge competence with apps, 2021.

ReST-RL: Achieving Accurate Code Reasoning of LLMs with Optimized Self-Training and Decoding Measuring coding challenge competence with apps, 2021

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-05T15:45:24.466089Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T15:45:24.466089Z digest=sha256:3c6a4b914bc0f90dc4da0b51f8a919764e8e4a5072bf09fae8cbd5a08dcda3b6

Observation b14037e4-5565-4d98-8b12-1f5a4cf810d1 · outbound

This paper cites Bigcodebench: Benchmarking code generation with diverse function calls and complex instructions, 2025.

ReST-RL: Achieving Accurate Code Reasoning of LLMs with Optimized Self-Training and Decoding Bigcodebench: Benchmarking code generation with diverse function calls and complex instructions, 2025

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T15:45:30.176357Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-05T15:45:24.507007Z digest=sha256:12769521a13b58a380dbd75d7d67c45d4c17f1deccd0b063797f918e0b557687

Observation 6de4234e-f4a4-43c3-99af-0c1adfb6f309 · outbound

This paper cites Wizardmath: Empowering mathematical reasoning for large language models via reinforced evol-instruct, 2025.

ReST-RL: Achieving Accurate Code Reasoning of LLMs with Optimized Self-Training and Decoding Wizardmath: Empowering mathematical reasoning for large language models via reinforced evol-instruct, 2025

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-05T15:45:24.557391Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T15:45:24.557391Z digest=sha256:f8f5f1cdb28f14e6ee6a8584bae0c23a9a93550dec6118dfb5431fd114cd0320

Observation ab22b356-6564-4fd2-9f72-eacc65650263 · outbound

This paper cites an unresolved cited work.

ReST-RL: Achieving Accurate Code Reasoning of LLMs with Optimized Self-Training and Decoding Unresolved cited work

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-05T15:45:24.606921Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T15:45:24.606921Z digest=sha256:ee2f095af564a72e5aff0c1abd61bab85f1db9b580644e8c4da22ed1209b3fe2

Observation 92664116-49b8-468f-9888-29660062a672 · outbound

This paper cites Proximal policy optimization algorithms, 2017.

ReST-RL: Achieving Accurate Code Reasoning of LLMs with Optimized Self-Training and Decoding Proximal policy optimization algorithms, 2017

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T15:45:29.850809Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-05T15:45:24.634749Z digest=sha256:e782f186005c1e8354e60f985ba3733c9cb8286aec1daaac2e6008db7f601660

Observation b633ee1d-19fb-4e55-bc88-399855cef95b · outbound

This paper cites Reinforced self-training (rest) for language modeling, 2023.

ReST-RL: Achieving Accurate Code Reasoning of LLMs with Optimized Self-Training and Decoding Reinforced self-training (rest) for language modeling, 2023

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-05T15:45:24.665314Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T15:45:24.665314Z digest=sha256:3912b4c5b393803bbd10c1fc24d91568cb2bec993bc7d08b405938c6c5d4400e

Observation 8b07e88e-f82b-4015-879d-84ab74b7dec5 · outbound

This paper cites Direct preference optimization: Your language model is secretly a reward model.

ReST-RL: Achieving Accurate Code Reasoning of LLMs with Optimized Self-Training and Decoding Direct preference optimization: Your language model is secretly a reward model

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-05T15:45:24.691732Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T15:45:24.691732Z digest=sha256:183ee113c2bdb4b4463bb41acce3aea40cb75e1e6964af3cd75390e485c4b807

Observation 7ac86a5a-13e6-422a-a657-8efefc8693bb · outbound

This paper cites an unresolved cited work.

ReST-RL: Achieving Accurate Code Reasoning of LLMs with Optimized Self-Training and Decoding Unresolved cited work

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-05T15:45:24.715154Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T15:45:24.715154Z digest=sha256:35188a074f34c6511ee256a26b06505284fd7fe647daa0b6a47e3ecd7c358738

Observation 87b836da-7df4-457d-a007-3dbb733fc341 · outbound

This paper cites Beyond Human Data: Scaling Self-Training for Problem-Solving with Language Models.

ReST-RL: Achieving Accurate Code Reasoning of LLMs with Optimized Self-Training and Decoding Beyond Human Data: Scaling Self-Training for Problem-Solving with Language Models

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-05T15:45:24.746306Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T15:45:24.746306Z digest=sha256:9844d2e9030833ea1641a57e6e5bd71752738d0f4aaebb53e15bcfce2e561ebb

Observation 395e6487-9a00-42b5-816e-1352a9c5a287 · outbound

This paper cites Rest-mcts*: Llm self-training via process reward guided tree search, 2024.

ReST-RL: Achieving Accurate Code Reasoning of LLMs with Optimized Self-Training and Decoding Rest-mcts*: Llm self-training via process reward guided tree search, 2024

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-05T15:45:24.772773Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T15:45:24.772773Z digest=sha256:cac4e1aa0d60ae28bd2483ffc3ced12eec45ee13b3f09b74cf30abe6ca77ef49

Observation 26a75e80-eec5-40d5-bc22-5b826471279b · outbound

This paper cites Solving math word problems with process- and outcome-based feedback, 2022.

ReST-RL: Achieving Accurate Code Reasoning of LLMs with Optimized Self-Training and Decoding Solving math word problems with process- and outcome-based feedback, 2022

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-05T15:45:24.807225Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T15:45:24.807225Z digest=sha256:b7729990dd720afbbb82cc9ec14c192b2e95b27ce433c66f77c5d838f1aa346c

Observation cf8b6baa-1348-419d-a785-65e77a65bd9a · outbound

This paper cites Let’s verify step by step, 2023.

ReST-RL: Achieving Accurate Code Reasoning of LLMs with Optimized Self-Training and Decoding Let’s verify step by step, 2023

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-05T15:45:24.854247Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T15:45:24.854247Z digest=sha256:495faae68c1295baf3078fcf7b98a06b02d16504b74b126be092703eff5b8073

Observation b2298171-678a-429f-b3ab-d93bede8b177 · outbound

This paper cites Don’t throw away your value model! generating more preferable text with value-guided monte-carlo tree search decoding, 2024.

ReST-RL: Achieving Accurate Code Reasoning of LLMs with Optimized Self-Training and Decoding Don’t throw away your value model! generating more preferable text with value-guided monte-carlo tree search decoding, 2024

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T15:45:29.124983Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-05T15:45:24.894475Z digest=sha256:a7394815b29fc93618412f22bf57ffb489de5d618004fe4c4bdfecb610b198cb

Observation 7d9bf68a-649b-4cf0-9b89-28b967dd121c · outbound

This paper cites Dapo: An open-source llm reinforcement learning system at scale, 2025.

ReST-RL: Achieving Accurate Code Reasoning of LLMs with Optimized Self-Training and Decoding Dapo: An open-source llm reinforcement learning system at scale, 2025

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-05T15:45:24.908108Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T15:45:24.908108Z digest=sha256:8724eef21d24b11d1014f752781b4d560ad448b7826a66ab52b440874c131abc

Observation d026e5d3-7815-40cd-b6c9-57440a7e51fd · outbound

This paper cites R1-vl: Learning to reason with multimodal large language models via step-wise group relative policy optimization, 2025.

ReST-RL: Achieving Accurate Code Reasoning of LLMs with Optimized Self-Training and Decoding R1-vl: Learning to reason with multimodal large language models via step-wise group relative policy optimization, 2025

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T15:45:28.884827Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-05T15:45:24.935044Z digest=sha256:466a9cec3871d8d75459dd1b7ea716996514cd14cd2a2165376291de7648e986

Observation 67be6a1d-670a-4911-b0be-c06630bba609 · outbound

This paper cites an unresolved cited work.

ReST-RL: Achieving Accurate Code Reasoning of LLMs with Optimized Self-Training and Decoding Unresolved cited work

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-05T15:45:24.974749Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T15:45:24.974749Z digest=sha256:3290e0fdd4a0f45bd875b6a3e249321c7f0864f0380efca52d2b7b8afce0fc1a

Observation 4483e15b-d22a-450a-ae42-64843050f91f · outbound

This paper cites Mankowitz, Esme Sutherland Robson, Pushmeet Kohli, Nando de Freitas, Koray Kavukcuoglu, and Oriol Vinyals.

ReST-RL: Achieving Accurate Code Reasoning of LLMs with Optimized Self-Training and Decoding Mankowitz, Esme Sutherland Robson, Pushmeet Kohli, Nando de Freitas, Koray Kavukcuoglu, and Oriol Vinyals

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T15:45:28.654742Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-05T15:45:25.024763Z digest=sha256:d94fa743dfb308df02653f45592654125310fcf3943a84001482953fad0dfef1

Observation 7426c0c7-f201-4fe8-9947-e31e6a755f36 · outbound

This paper cites Code with codeqwen1.5, April 2024.

ReST-RL: Achieving Accurate Code Reasoning of LLMs with Optimized Self-Training and Decoding Code with codeqwen1.5, April 2024

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T15:45:28.494850Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-05T15:45:25.057163Z digest=sha256:2006fd412a537238c6760cbed2b3c8dc39d93d17a7b4db39466ed79098ea8b52

Observation e00ea81a-0836-4ae5-9e35-a39b47ab048e · outbound

This paper cites OpenCodeInterpreter: Integrating Code Generation with Execution and Refinement.

ReST-RL: Achieving Accurate Code Reasoning of LLMs with Optimized Self-Training and Decoding OpenCodeInterpreter: Integrating Code Generation with Execution and Refinement

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-05T15:45:25.104749Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T15:45:25.104749Z digest=sha256:18bc1f8259b9b25d02d5fe9a428a4e1a596fdd6ca901502ba97c8018ef9d8981

Observation ced27b90-3a83-484a-a60f-dcb6b9027441 · outbound

This paper cites Tenenbaum, and Chuang Gan.

ReST-RL: Achieving Accurate Code Reasoning of LLMs with Optimized Self-Training and Decoding Tenenbaum, and Chuang Gan

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T15:45:28.355103Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-05T15:45:25.145037Z digest=sha256:a5b0f69eae7f68a7fcd3f0b9aea8d4ed61d662eea430d7eb91103022e50d2312

Observation 92ee3d93-802c-4791-8843-9a1f350c15cd · outbound

This paper cites Jiang, Jia Deng, Stella Biderman, and Sean Welleck.

ReST-RL: Achieving Accurate Code Reasoning of LLMs with Optimized Self-Training and Decoding Jiang, Jia Deng, Stella Biderman, and Sean Welleck

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T15:45:28.224747Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-05T15:45:25.195191Z digest=sha256:357580d3d1d9515d9d4e7bf0eeec79dae7c4f70f828c879f5dab3549118d1812

Observation db52f004-bdc6-46db-98bf-c841ce605ba3 · outbound

This paper cites an unresolved cited work.

ReST-RL: Achieving Accurate Code Reasoning of LLMs with Optimized Self-Training and Decoding Unresolved cited work

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-05T15:45:25.245194Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T15:45:25.245194Z digest=sha256:4b3180af65284497f238d1c29388200f0cde4d92950ac9dcfa9f31a97768cb1c

Observation efde76ed-bddc-498b-bb3a-9ca6434be0e5 · outbound

This paper cites Pspo*: An effective process-supervised policy optimization for reasoning alignment, 2024.

ReST-RL: Achieving Accurate Code Reasoning of LLMs with Optimized Self-Training and Decoding Pspo*: An effective process-supervised policy optimization for reasoning alignment, 2024

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T15:45:27.986574Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-05T15:45:25.284748Z digest=sha256:edde2bdca107db2ea2bf2f7c904302652b509a03ba2974fa57ed2b85f6ceba9d

Observation 4dff5136-fc75-4f6a-8de6-66ce1bea5707 · outbound

This paper cites an unresolved cited work.

ReST-RL: Achieving Accurate Code Reasoning of LLMs with Optimized Self-Training and Decoding Unresolved cited work

Reference 27

Resolution
unresolved
raw_fallback, observed 2026-08-05T15:45:27.674899Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-05T15:45:25.334751Z digest=sha256:0637a97d82fe3ad763db436a53269502f0483980b5658051e29da1a320c1e972

Observation 633b31e6-0ab6-414b-aa06-415a3df86dc0 · outbound

This paper cites Chain-of-thought prompting elicits reasoning in large language models, 2023.

ReST-RL: Achieving Accurate Code Reasoning of LLMs with Optimized Self-Training and Decoding Chain-of-thought prompting elicits reasoning in large language models, 2023

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-05T15:45:25.384748Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T15:45:25.384748Z digest=sha256:ee944c6637d66667df339d2a5203f13e8ae8510e54dd3577287ab95db8c7f192

Observation 3afc0a37-acd2-4b08-8782-f6ee69c42127 · outbound

This paper cites Griffiths, Yuan Cao, and Karthik Narasimhan.

ReST-RL: Achieving Accurate Code Reasoning of LLMs with Optimized Self-Training and Decoding Griffiths, Yuan Cao, and Karthik Narasimhan

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-05T15:45:25.434824Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T15:45:25.434824Z digest=sha256:0de5794f590d0e4db36f844550254938efd8bd747bb64f625201bdafd6ad6067

Observation b9d8e6d4-9b13-4c45-a38e-0ce90ee69bf4 · outbound

This paper cites Sra-mcts: Self-driven reasoning augmentation with monte carlo tree search for code generation, 2025.

ReST-RL: Achieving Accurate Code Reasoning of LLMs with Optimized Self-Training and Decoding Sra-mcts: Self-driven reasoning augmentation with monte carlo tree search for code generation, 2025

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T15:45:27.054841Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-05T15:45:25.475053Z digest=sha256:0ce0b41e4eac9cf1b86cbb9182dbf7807029e7cee60913542145a8e933627de9

Observation f14061ea-d91f-4861-b37e-73c7d3245ce1 · outbound

This paper cites Accessing gpt-4 level mathematical olympiad solutions via monte carlo tree self-refine with llama-3 8b, 2024.

ReST-RL: Achieving Accurate Code Reasoning of LLMs with Optimized Self-Training and Decoding Accessing gpt-4 level mathematical olympiad solutions via monte carlo tree self-refine with llama-3 8b, 2024

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-05T15:45:25.524747Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T15:45:25.524747Z digest=sha256:78da79f68a18fde1fe5f9a7fd6bf636679bb1c6e30d93c4c18e0e6c66758eb9c

Observation 283cc7ff-73ce-46b9-a5c6-f315b223cce8 · outbound

This paper cites Program synthesis with large language models, 2021.

ReST-RL: Achieving Accurate Code Reasoning of LLMs with Optimized Self-Training and Decoding Program synthesis with large language models, 2021

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-05T15:45:25.574819Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T15:45:25.574819Z digest=sha256:860dfa793099756172990a18177644db7dda1ff23697ba469a2cd49d296b9f51

Observation 3c654c56-76c3-4b7a-9920-237fa938e6ff · outbound

This paper cites Qwen2.5-Coder Technical Report.

ReST-RL: Achieving Accurate Code Reasoning of LLMs with Optimized Self-Training and Decoding Qwen2.5-Coder Technical Report

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-05T15:45:25.624748Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T15:45:25.624748Z digest=sha256:a666f4e519eded7c0bee66bd6adb78a48bdb6f934e86fb097ba57c0aaf20d932

Observation 629afcb8-3aff-44ef-b900-cd5457763673 · outbound

This paper cites Qwen3 technical report, 2025.

ReST-RL: Achieving Accurate Code Reasoning of LLMs with Optimized Self-Training and Decoding Qwen3 technical report, 2025

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-05T15:45:25.674750Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T15:45:25.674750Z digest=sha256:36703438fef067cde4f87d911f87119bbd614f15fffd310b77b3c4c64537de48

Observation d069aee4-119a-4538-a5cb-537ee583cb65 · outbound

This paper cites The llama 3 herd of models, 2024.

ReST-RL: Achieving Accurate Code Reasoning of LLMs with Optimized Self-Training and Decoding The llama 3 herd of models, 2024

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-05T15:45:25.717705Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T15:45:25.717705Z digest=sha256:c676d091382f03613437310f7eb7ab6e3808b32a6a7de4cca0deaba71244969d

Observation 87a5e739-ada5-49ec-b3d4-c7812f70237e · outbound

This paper cites Ds-1000: A natural and reliable benchmark for data science code generation, 2022.

ReST-RL: Achieving Accurate Code Reasoning of LLMs with Optimized Self-Training and Decoding Ds-1000: A natural and reliable benchmark for data science code generation, 2022

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-05T15:45:25.764827Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T15:45:25.764827Z digest=sha256:432c69e1eba5bc0e4fbc6806b7639adacbcb5c07449c7a4f77e917b29cffc1ab

Observation 0b083143-5493-4f29-95ae-d9af24999b87 · outbound

This paper cites Is your code generated by chatgpt really correct? rigorous evaluation of large language models for code generation, 2023.

ReST-RL: Achieving Accurate Code Reasoning of LLMs with Optimized Self-Training and Decoding Is your code generated by chatgpt really correct? rigorous evaluation of large language models for code generation, 2023

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-05T15:45:25.814825Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T15:45:25.814825Z digest=sha256:dbf5145338285468d843adf68300fde86ef0f6aa1ca1a261149b947eaf6c9101

Observation c177dc3c-5724-40b6-80a9-cea292e5f0cb · outbound

This paper cites Reward hacking in reinforcement learning.lilianweng.github.io, Nov 2024.

ReST-RL: Achieving Accurate Code Reasoning of LLMs with Optimized Self-Training and Decoding Reward hacking in reinforcement learning.lilianweng.github.io, Nov 2024

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-05T15:45:25.854752Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T15:45:25.854752Z digest=sha256:40e28a59187178f4b6f1c65a03e03313c5ae1f0cfa3b63b5d829ee46c3af2773

Observation c5d66e67-fb19-425b-a5e7-c3f3d9d21913 · outbound

This paper cites Skywork-Reward: Bag of Tricks for Reward Modeling in LLMs.

ReST-RL: Achieving Accurate Code Reasoning of LLMs with Optimized Self-Training and Decoding Skywork-Reward: Bag of Tricks for Reward Modeling in LLMs

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-05T15:45:25.874749Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T15:45:25.874749Z digest=sha256:58fcd0c5daf70c611b8a031390007f51549c9426c6a4a9443a39f14722ac6594

Pith citing papers

Observation faac2c7a-0710-4a7d-ae45-b324f3875e95 · inbound

MARS$^2$: Scaling Multi-Agent Tree Search via Reinforcement Learning for Code Generation cites this paper.

MARS$^2$: Scaling Multi-Agent Tree Search via Reinforcement Learning for Code Generation ReST-RL: Achieving Accurate Code Reasoning of LLMs with Optimized Self-Training and Decoding

Reference 35

Resolution
verified exact
arxiv_id, observed 2026-05-10T11:30:18.824443Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-05-10T11:27:28.245835Z digest=sha256:cac54f47953df5c32fb64391f0f0b54e4c40863e39caadee1b035677cdef1d49

Observation ecddc6aa-ea6c-481f-914b-ec867be66e78 · inbound

Mental-R1: Aligning LLM Reasoning for Mental Health Assessment cites this paper.

Mental-R1: Aligning LLM Reasoning for Mental Health Assessment ReST-RL: Achieving Accurate Code Reasoning of LLMs with Optimized Self-Training and Decoding

Reference 67

Resolution
verified exact
arxiv_id, observed 2026-07-03T14:08:21.671873Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-06-27T07:15:00.398111Z digest=sha256:092c0015c57478314d4019a949971fcbec17415e7ff1e21a70c5abdde3177c37

Observation d6f58a78-07b7-4087-9e46-9547b07c6fd2 · inbound

Continual Learning in Transition cites this paper.

Continual Learning in Transition ReST-RL: Achieving Accurate Code Reasoning of LLMs with Optimized Self-Training and Decoding

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-07T12:23:56.891313Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:23:56.891313Z digest=sha256:8a662e04b9fd30300641950cae5e510e58929639a716b0cbeae84e30c8270e13