Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-05T15:45:25.874749Z
Paper Citation Record · LEDGER
As of 10 August 2026, this Paper Citation Record lists 39 of 39 outbound references and 3 inbound Pith citation observations for arXiv:2508.19576.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-05T15:45:25.874749Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-10T06:31:04.303077+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-08-07T12:23:56.891313Z
A source-named dated measurement, never combined with another source.
Source: arxiv_reference, observed 2026-07-03T14:08:21.670108Z
39 of 39 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation a60a4c92-264c-447b-8b38-be13e1b5183b · outbound
ReST-RL: Achieving Accurate Code Reasoning of LLMs with Optimized Self-Training and Decoding Gpt-4 technical report, 2024
Reference 1
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation bd05cd66-c6bb-414c-a6d5-e445557a6857 · outbound
ReST-RL: Achieving Accurate Code Reasoning of LLMs with Optimized Self-Training and Decoding Unresolved cited work
Reference 2
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0464326c-2220-4f6c-909e-d1c51896f0ff · outbound
ReST-RL: Achieving Accurate Code Reasoning of LLMs with Optimized Self-Training and Decoding The lessons of developing process reward models in mathematical reasoning, 2025
Reference 3
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 285765fa-6c82-49c7-89c2-f592990cd7fd · outbound
ReST-RL: Achieving Accurate Code Reasoning of LLMs with Optimized Self-Training and Decoding Measuring coding challenge competence with apps, 2021
Reference 4
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b14037e4-5565-4d98-8b12-1f5a4cf810d1 · outbound
ReST-RL: Achieving Accurate Code Reasoning of LLMs with Optimized Self-Training and Decoding Bigcodebench: Benchmarking code generation with diverse function calls and complex instructions, 2025
Reference 5
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 6de4234e-f4a4-43c3-99af-0c1adfb6f309 · outbound
ReST-RL: Achieving Accurate Code Reasoning of LLMs with Optimized Self-Training and Decoding Wizardmath: Empowering mathematical reasoning for large language models via reinforced evol-instruct, 2025
Reference 6
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ab22b356-6564-4fd2-9f72-eacc65650263 · outbound
ReST-RL: Achieving Accurate Code Reasoning of LLMs with Optimized Self-Training and Decoding Unresolved cited work
Reference 7
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 92664116-49b8-468f-9888-29660062a672 · outbound
ReST-RL: Achieving Accurate Code Reasoning of LLMs with Optimized Self-Training and Decoding Proximal policy optimization algorithms, 2017
Reference 8
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation b633ee1d-19fb-4e55-bc88-399855cef95b · outbound
ReST-RL: Achieving Accurate Code Reasoning of LLMs with Optimized Self-Training and Decoding Reinforced self-training (rest) for language modeling, 2023
Reference 9
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8b07e88e-f82b-4015-879d-84ab74b7dec5 · outbound
ReST-RL: Achieving Accurate Code Reasoning of LLMs with Optimized Self-Training and Decoding Direct preference optimization: Your language model is secretly a reward model
Reference 10
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7ac86a5a-13e6-422a-a657-8efefc8693bb · outbound
ReST-RL: Achieving Accurate Code Reasoning of LLMs with Optimized Self-Training and Decoding Unresolved cited work
Reference 11
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 87b836da-7df4-457d-a007-3dbb733fc341 · outbound
ReST-RL: Achieving Accurate Code Reasoning of LLMs with Optimized Self-Training and Decoding Beyond Human Data: Scaling Self-Training for Problem-Solving with Language Models
Reference 12
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 395e6487-9a00-42b5-816e-1352a9c5a287 · outbound
ReST-RL: Achieving Accurate Code Reasoning of LLMs with Optimized Self-Training and Decoding Rest-mcts*: Llm self-training via process reward guided tree search, 2024
Reference 13
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 26a75e80-eec5-40d5-bc22-5b826471279b · outbound
ReST-RL: Achieving Accurate Code Reasoning of LLMs with Optimized Self-Training and Decoding Solving math word problems with process- and outcome-based feedback, 2022
Reference 14
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation cf8b6baa-1348-419d-a785-65e77a65bd9a · outbound
ReST-RL: Achieving Accurate Code Reasoning of LLMs with Optimized Self-Training and Decoding Let’s verify step by step, 2023
Reference 15
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b2298171-678a-429f-b3ab-d93bede8b177 · outbound
ReST-RL: Achieving Accurate Code Reasoning of LLMs with Optimized Self-Training and Decoding Don’t throw away your value model! generating more preferable text with value-guided monte-carlo tree search decoding, 2024
Reference 16
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 7d9bf68a-649b-4cf0-9b89-28b967dd121c · outbound
ReST-RL: Achieving Accurate Code Reasoning of LLMs with Optimized Self-Training and Decoding Dapo: An open-source llm reinforcement learning system at scale, 2025
Reference 17
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d026e5d3-7815-40cd-b6c9-57440a7e51fd · outbound
ReST-RL: Achieving Accurate Code Reasoning of LLMs with Optimized Self-Training and Decoding R1-vl: Learning to reason with multimodal large language models via step-wise group relative policy optimization, 2025
Reference 18
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 67be6a1d-670a-4911-b0be-c06630bba609 · outbound
ReST-RL: Achieving Accurate Code Reasoning of LLMs with Optimized Self-Training and Decoding Unresolved cited work
Reference 19
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4483e15b-d22a-450a-ae42-64843050f91f · outbound
ReST-RL: Achieving Accurate Code Reasoning of LLMs with Optimized Self-Training and Decoding Mankowitz, Esme Sutherland Robson, Pushmeet Kohli, Nando de Freitas, Koray Kavukcuoglu, and Oriol Vinyals
Reference 20
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 7426c0c7-f201-4fe8-9947-e31e6a755f36 · outbound
ReST-RL: Achieving Accurate Code Reasoning of LLMs with Optimized Self-Training and Decoding Code with codeqwen1.5, April 2024
Reference 21
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation e00ea81a-0836-4ae5-9e35-a39b47ab048e · outbound
ReST-RL: Achieving Accurate Code Reasoning of LLMs with Optimized Self-Training and Decoding OpenCodeInterpreter: Integrating Code Generation with Execution and Refinement
Reference 22
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ced27b90-3a83-484a-a60f-dcb6b9027441 · outbound
ReST-RL: Achieving Accurate Code Reasoning of LLMs with Optimized Self-Training and Decoding Tenenbaum, and Chuang Gan
Reference 23
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 92ee3d93-802c-4791-8843-9a1f350c15cd · outbound
ReST-RL: Achieving Accurate Code Reasoning of LLMs with Optimized Self-Training and Decoding Jiang, Jia Deng, Stella Biderman, and Sean Welleck
Reference 24
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation db52f004-bdc6-46db-98bf-c841ce605ba3 · outbound
ReST-RL: Achieving Accurate Code Reasoning of LLMs with Optimized Self-Training and Decoding Unresolved cited work
Reference 25
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation efde76ed-bddc-498b-bb3a-9ca6434be0e5 · outbound
ReST-RL: Achieving Accurate Code Reasoning of LLMs with Optimized Self-Training and Decoding Pspo*: An effective process-supervised policy optimization for reasoning alignment, 2024
Reference 26
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 4dff5136-fc75-4f6a-8de6-66ce1bea5707 · outbound
ReST-RL: Achieving Accurate Code Reasoning of LLMs with Optimized Self-Training and Decoding Unresolved cited work
Reference 27
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 633b31e6-0ab6-414b-aa06-415a3df86dc0 · outbound
ReST-RL: Achieving Accurate Code Reasoning of LLMs with Optimized Self-Training and Decoding Chain-of-thought prompting elicits reasoning in large language models, 2023
Reference 28
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3afc0a37-acd2-4b08-8782-f6ee69c42127 · outbound
ReST-RL: Achieving Accurate Code Reasoning of LLMs with Optimized Self-Training and Decoding Griffiths, Yuan Cao, and Karthik Narasimhan
Reference 29
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b9d8e6d4-9b13-4c45-a38e-0ce90ee69bf4 · outbound
ReST-RL: Achieving Accurate Code Reasoning of LLMs with Optimized Self-Training and Decoding Sra-mcts: Self-driven reasoning augmentation with monte carlo tree search for code generation, 2025
Reference 30
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation f14061ea-d91f-4861-b37e-73c7d3245ce1 · outbound
ReST-RL: Achieving Accurate Code Reasoning of LLMs with Optimized Self-Training and Decoding Accessing gpt-4 level mathematical olympiad solutions via monte carlo tree self-refine with llama-3 8b, 2024
Reference 31
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 283cc7ff-73ce-46b9-a5c6-f315b223cce8 · outbound
ReST-RL: Achieving Accurate Code Reasoning of LLMs with Optimized Self-Training and Decoding Program synthesis with large language models, 2021
Reference 32
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3c654c56-76c3-4b7a-9920-237fa938e6ff · outbound
ReST-RL: Achieving Accurate Code Reasoning of LLMs with Optimized Self-Training and Decoding Qwen2.5-Coder Technical Report
Reference 33
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 629afcb8-3aff-44ef-b900-cd5457763673 · outbound
ReST-RL: Achieving Accurate Code Reasoning of LLMs with Optimized Self-Training and Decoding Qwen3 technical report, 2025
Reference 34
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d069aee4-119a-4538-a5cb-537ee583cb65 · outbound
ReST-RL: Achieving Accurate Code Reasoning of LLMs with Optimized Self-Training and Decoding The llama 3 herd of models, 2024
Reference 35
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 87a5e739-ada5-49ec-b3d4-c7812f70237e · outbound
ReST-RL: Achieving Accurate Code Reasoning of LLMs with Optimized Self-Training and Decoding Ds-1000: A natural and reliable benchmark for data science code generation, 2022
Reference 36
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0b083143-5493-4f29-95ae-d9af24999b87 · outbound
ReST-RL: Achieving Accurate Code Reasoning of LLMs with Optimized Self-Training and Decoding Is your code generated by chatgpt really correct? rigorous evaluation of large language models for code generation, 2023
Reference 37
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c177dc3c-5724-40b6-80a9-cea292e5f0cb · outbound
ReST-RL: Achieving Accurate Code Reasoning of LLMs with Optimized Self-Training and Decoding Reward hacking in reinforcement learning.lilianweng.github.io, Nov 2024
Reference 38
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c5d66e67-fb19-425b-a5e7-c3f3d9d21913 · outbound
ReST-RL: Achieving Accurate Code Reasoning of LLMs with Optimized Self-Training and Decoding Skywork-Reward: Bag of Tricks for Reward Modeling in LLMs
Reference 39
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation faac2c7a-0710-4a7d-ae45-b324f3875e95 · inbound
MARS$^2$: Scaling Multi-Agent Tree Search via Reinforcement Learning for Code Generation ReST-RL: Achieving Accurate Code Reasoning of LLMs with Optimized Self-Training and Decoding
Reference 35
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation ecddc6aa-ea6c-481f-914b-ec867be66e78 · inbound
Mental-R1: Aligning LLM Reasoning for Mental Health Assessment ReST-RL: Achieving Accurate Code Reasoning of LLMs with Optimized Self-Training and Decoding
Reference 67
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation d6f58a78-07b7-4087-9e46-9547b07c6fd2 · inbound
Continual Learning in Transition ReST-RL: Achieving Accurate Code Reasoning of LLMs with Optimized Self-Training and Decoding
Reference 10
Source-reported events for the cited work
Unavailable: canonical work link unavailable.