Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-07T00:31:38.586086Z
Paper Citation Record · LEDGER
As of 9 August 2026, this Paper Citation Record lists 58 of 58 outbound references and 11 inbound Pith citation observations for arXiv:2506.13923.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-07T00:31:38.586086Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-09T06:31:02.800959+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-08-04T10:53:08.661292Z
A source-named dated measurement, never combined with another source.
Source: arxiv_reference, observed 2026-07-03T20:48:56.210532Z
58 of 58 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation 93c69185-51db-403f-b91c-1b1ce5ffd1d1 · outbound
Adaptive Guidance Accelerates Reinforcement Learning of Reasoning Models OpenAI o1 System Card
Reference 1
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a8988f1e-0c2e-4284-bd35-823af1320e74 · outbound
Adaptive Guidance Accelerates Reinforcement Learning of Reasoning Models DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning
Reference 2
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b7ba5a2e-91c5-42a3-9966-0dd733e5b83c · outbound
Adaptive Guidance Accelerates Reinforcement Learning of Reasoning Models Tulu 3: Pushing Frontiers in Open Language Model Post-Training
Reference 3
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation aeed9b1a-ca68-44e7-a2ec-7aa18495b3a8 · outbound
Adaptive Guidance Accelerates Reinforcement Learning of Reasoning Models Teaching large language models to reason with reinforcement learning
Reference 4
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d0963d46-c208-49b6-a62c-dfbab7fe6557 · outbound
Adaptive Guidance Accelerates Reinforcement Learning of Reasoning Models Does Reinforcement Learning Really Incentivize Reasoning Capacity in LLMs Beyond the Base Model?
Reference 5
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 13e84987-c960-42bc-9126-4f976a084cc3 · outbound
Adaptive Guidance Accelerates Reinforcement Learning of Reasoning Models Unresolved cited work
Reference 6
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation e30394ea-1a98-4aed-b777-7263413883c9 · outbound
Adaptive Guidance Accelerates Reinforcement Learning of Reasoning Models Reinforced Self-Training (ReST) for Language Modeling
Reference 7
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 689b6931-b901-48bb-96f4-67616561938a · outbound
Adaptive Guidance Accelerates Reinforcement Learning of Reasoning Models Openai’s reinforcement fine-tuning research program, 2024
Reference 8
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation c1892acf-dbe5-43b7-95c3-e78b068ceaea · outbound
Adaptive Guidance Accelerates Reinforcement Learning of Reasoning Models Be- yond human data: Scaling self-training for problem-solving with language models.Transactions on Machine Learning Research, 2024
Reference 9
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 77cc6ec1-5c4b-4636-94df-9c093f623c76 · outbound
Adaptive Guidance Accelerates Reinforcement Learning of Reasoning Models V-star: Training verifiers for self-taught reasoners
Reference 10
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation c2bb22dc-bb45-4d6b-ab4c-61751a999089 · outbound
Adaptive Guidance Accelerates Reinforcement Learning of Reasoning Models ReST- MCTS*: LLM self-training via process reward guided tree search
Reference 11
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation b2a05105-dcec-47c0-9c68-e37529ec670d · outbound
Adaptive Guidance Accelerates Reinforcement Learning of Reasoning Models Continuous control with deep reinforcement learning
Reference 12
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b0c55e14-5657-40b7-af77-ff320e39af2b · outbound
Adaptive Guidance Accelerates Reinforcement Learning of Reasoning Models Qwen2.5 Technical Report
Reference 13
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 503ce4ba-3497-4e35-aff9-16f65a6e7112 · outbound
Adaptive Guidance Accelerates Reinforcement Learning of Reasoning Models Measuring Mathematical Problem Solving With the MATH Dataset
Reference 15
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a94a7577-dd41-45a5-befa-721617b21833 · outbound
Adaptive Guidance Accelerates Reinforcement Learning of Reasoning Models Aime 2024
Reference 16
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 27f7e066-ca93-4f55-bde1-e8f978f7cb64 · outbound
Adaptive Guidance Accelerates Reinforcement Learning of Reasoning Models Aime 2025
Reference 17
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 5a8cccb6-bcf6-462a-af42-818c0d348387 · outbound
Adaptive Guidance Accelerates Reinforcement Learning of Reasoning Models Amc 2023
Reference 18
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 07087427-359a-4e9b-9f33-1c67cb2dfdfa · outbound
Adaptive Guidance Accelerates Reinforcement Learning of Reasoning Models Gpqa: A graduate-level google-proof q&a benchmark
Reference 19
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c05f9dc9-279f-4538-b816-a2bec136b49b · outbound
Adaptive Guidance Accelerates Reinforcement Learning of Reasoning Models OlympiadBench: A Challenging Benchmark for Promoting AGI with Olympiad-Level Bilingual Multimodal Scientific Problems
Reference 20
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation beb9fd64-2327-4d99-9065-6f3e2b8f317e · outbound
Adaptive Guidance Accelerates Reinforcement Learning of Reasoning Models DeepSeek-Coder: When the Large Language Model Meets Programming -- The Rise of Code Intelligence
Reference 21
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2577b0f8-e45f-4c7a-a99f-bf593af79254 · outbound
Adaptive Guidance Accelerates Reinforcement Learning of Reasoning Models Livecodebench: Holistic and contamination free evaluation of large language models for code.arXiv preprint, 2024
Reference 22
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 098d9765-40b6-4420-900c-835487736344 · outbound
Adaptive Guidance Accelerates Reinforcement Learning of Reasoning Models Evaluating Large Language Models Trained on Code
Reference 23
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 35b5e4f3-a3a6-487f-a86d-414cb3cac3d3 · outbound
Adaptive Guidance Accelerates Reinforcement Learning of Reasoning Models Open r1: A fully open reproduction of deepseek-r1, January 2025
Reference 24
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a9ff5d5f-0123-4e33-b672-8e7b2dd1c244 · outbound
Adaptive Guidance Accelerates Reinforcement Learning of Reasoning Models Dapo: An open-source llm reinforcement learning system at scale, 2025
Reference 25
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c7094ab0-9a3c-441e-b0b6-b7e5c9eb3c5e · outbound
Adaptive Guidance Accelerates Reinforcement Learning of Reasoning Models Fine-Tuning Language Models from Human Preferences
Reference 26
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 97b7a885-7d76-403d-8754-9278b13936d2 · outbound
Adaptive Guidance Accelerates Reinforcement Learning of Reasoning Models Unresolved cited work
Reference 27
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation a9c0a123-1458-4ec2-ad46-403029fe2937 · outbound
Adaptive Guidance Accelerates Reinforcement Learning of Reasoning Models LMRL Gym: Benchmarks for Multi-Turn Reinforcement Learning with Language Models
Reference 28
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c63c58de-bbae-4044-892f-a21618e9f075 · outbound
Adaptive Guidance Accelerates Reinforcement Learning of Reasoning Models Archer: Training language model agents via hierarchical multi-turn rl
Reference 29
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 9a9491a7-4fca-4fb7-b1ee-481b3c08470d · outbound
Adaptive Guidance Accelerates Reinforcement Learning of Reasoning Models Aligning large multimodal models with factually augmented rlhf
Reference 30
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 63cc29ad-0a40-48b2-b947-c9ab58369e60 · outbound
Adaptive Guidance Accelerates Reinforcement Learning of Reasoning Models Kimi k1.5: Scaling Reinforcement Learning with LLMs
Reference 31
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 62593bd3-d18c-4b33-b0fb-2ab464d70a43 · outbound
Adaptive Guidance Accelerates Reinforcement Learning of Reasoning Models Training Verifiers to Solve Math Word Problems
Reference 32
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a56e454b-536e-4593-a6ea-861ee4fba24f · outbound
Adaptive Guidance Accelerates Reinforcement Learning of Reasoning Models Program Synthesis with Large Language Models
Reference 33
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 43b36156-2531-46f0-b9dd-afe892cd943c · outbound
Adaptive Guidance Accelerates Reinforcement Learning of Reasoning Models RLEF: Grounding Code LLMs in Execution Feedback with Reinforcement Learning
Reference 34
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0e997cca-1852-4654-a328-b4b1db39ef8c · outbound
Adaptive Guidance Accelerates Reinforcement Learning of Reasoning Models SFT Memorizes, RL Generalizes: A Comparative Study of Foundation Model Post-training
Reference 35
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 386ef46f-11a8-45a7-9bcd-9bf7868f5487 · outbound
Adaptive Guidance Accelerates Reinforcement Learning of Reasoning Models Scaling LLM test-time com- pute optimally can be more effective than scaling parameters for reasoning
Reference 36
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3f71c7fa-5d4b-4ef8-a912-a461dd82f61c · outbound
Adaptive Guidance Accelerates Reinforcement Learning of Reasoning Models Echo Chamber: RL Post-training Amplifies Behaviors Learned in Pretraining
Reference 37
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9fb8845d-8e9c-4e18-981d-6214dbf83a0d · outbound
Adaptive Guidance Accelerates Reinforcement Learning of Reasoning Models Assessing diversity collapse in reasoning
Reference 38
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation fee6f11f-2095-4421-a497-0c9994c674eb · outbound
Adaptive Guidance Accelerates Reinforcement Learning of Reasoning Models Direct preference optimization: Your language model is secretly a reward model
Reference 39
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 699988d1-d4a5-403a-b834-fa3f35c6c9b9 · outbound
Adaptive Guidance Accelerates Reinforcement Learning of Reasoning Models Tapered Off-Policy REINFORCE: Stable and efficient reinforcement learning for LLMs
Reference 40
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation eaf1f6ed-4b5c-439f-ad65-cf6251c95323 · outbound
Adaptive Guidance Accelerates Reinforcement Learning of Reasoning Models Learning to Reason under Off-Policy Guidance
Reference 41
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3294b235-12d3-4aea-af15-30508f14116f · outbound
Adaptive Guidance Accelerates Reinforcement Learning of Reasoning Models HybridFlow: A Flexible and Efficient RLHF Framework
Reference 42
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2cbc1157-af60-4496-8d47-c2d7ac995883 · outbound
Adaptive Guidance Accelerates Reinforcement Learning of Reasoning Models Unresolved cited work
Reference 43
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 0b287539-eb83-4218-bd02-e499e193e097 · outbound
Adaptive Guidance Accelerates Reinforcement Learning of Reasoning Models Unresolved cited work
Reference 44
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 3ed7525f-09d6-411a-9a85-7468a3918160 · outbound
Adaptive Guidance Accelerates Reinforcement Learning of Reasoning Models Unresolved cited work
Reference 45
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation b6d942a2-c234-4217-98f2-f56e1763451c · outbound
Adaptive Guidance Accelerates Reinforcement Learning of Reasoning Models Unresolved cited work
Reference 46
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 2bc9a944-19ff-4b71-bdeb-cc3b8a69008b · outbound
Adaptive Guidance Accelerates Reinforcement Learning of Reasoning Models A hint to the problem is provided below: [HINT_START]
Reference 47
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 85b570b4-9fb5-485c-b190-aab6de5316ce · outbound
Adaptive Guidance Accelerates Reinforcement Learning of Reasoning Models Think about how the identity (a²/4) - (4/a²) might be used as a building block for factoring the larger expression
Reference 48
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation e13d6c72-369a-437f-ad8c-5d16f7844303 · outbound
Adaptive Guidance Accelerates Reinforcement Learning of Reasoning Models Ask yourself if a difference of powers or a recognizable factorization pattern might help connect the two parts of the expression
Reference 49
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation a818a75b-c756-4365-9c30-8c403835893b · outbound
Adaptive Guidance Accelerates Reinforcement Learning of Reasoning Models How can this substitution simplify the structure of the problem? ,→ ,→
Reference 50
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation ea984d5a-4322-41c4-93a6-29d8d3507596 · outbound
Adaptive Guidance Accelerates Reinforcement Learning of Reasoning Models Unresolved cited work
Reference 51
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 5ecf4a51-a6e2-4c78-aa2a-0f184f08c3db · outbound
Adaptive Guidance Accelerates Reinforcement Learning of Reasoning Models Use these observations to guide your step-by-step approach toward the final simplified result
Reference 52
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation c4566f0d-4b2d-4f7c-a2b6-124fc1a5569b · outbound
Adaptive Guidance Accelerates Reinforcement Learning of Reasoning Models Unresolved cited work
Reference 53
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation d504db0d-c856-45e5-804c-5eb4622496fe · outbound
Adaptive Guidance Accelerates Reinforcement Learning of Reasoning Models Unresolved cited work
Reference 54
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation b03f974b-9501-4cfc-9c19-0e6a97c4b804 · outbound
Adaptive Guidance Accelerates Reinforcement Learning of Reasoning Models What does that imply about multiplying the side length? ,→ ,→
Reference 55
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation e8681102-13af-4343-9de8-61d3a7d873a2 · outbound
Adaptive Guidance Accelerates Reinforcement Learning of Reasoning Models Ensure each step follows from the properties of a square.,→
Reference 56
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation a5312c55-30df-4e3f-bd80-b29e1640dc0b · outbound
Adaptive Guidance Accelerates Reinforcement Learning of Reasoning Models using the hint
Reference 57
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 79e676a7-60b4-4be4-b507-194276321478 · outbound
Adaptive Guidance Accelerates Reinforcement Learning of Reasoning Models This example shows how a concise, domain-specific hint can redirect the model’s reasoning and correct a systematic geometric error
Reference 58
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 9efcd0cd-68c1-4bd7-bcb7-2855da000ad5 · outbound
Adaptive Guidance Accelerates Reinforcement Learning of Reasoning Models <think>\n {thoughts} </think>\n\
Reference 59
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation ae143971-845e-4b2c-a967-2259fae6503a · inbound
EvoCoT: Overcoming the Exploration Bottleneck in Reinforcement Learning Adaptive Guidance Accelerates Reinforcement Learning of Reasoning Models
Reference 11
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation a5a87b9b-3829-498b-a3b9-57b85d3d4566 · inbound
Position: The Hidden Costs and Measurement Gaps of Reinforcement Learning with Verifiable Rewards Adaptive Guidance Accelerates Reinforcement Learning of Reasoning Models
Reference 25
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 11411b44-0b68-42f9-aeb6-eab51a369da2 · inbound
TaoSR-AGRL: Adaptive Guided Reinforcement Learning Framework for E-commerce Search Relevance Adaptive Guidance Accelerates Reinforcement Learning of Reasoning Models
Reference 23
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c4ef3d49-52e8-4f36-943f-efa6b001fcff · inbound
Selective Off-Policy Reference Tuning with Plan Guidance Adaptive Guidance Accelerates Reinforcement Learning of Reasoning Models
Reference 18
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation f6f35c7f-8dc4-4247-968c-0dd26bff82cf · inbound
Selective Off-Policy Reference Tuning with Plan Guidance Adaptive Guidance Accelerates Reinforcement Learning of Reasoning Models
Reference 18
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 2cf22fff-26b7-4055-817f-97f6907687bc · inbound
Learning Agentic Policy from Action Guidance Adaptive Guidance Accelerates Reinforcement Learning of Reasoning Models
Reference 42
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 155a57b7-919a-4bda-89cb-84696c4310ee · inbound
Hide to Guide: Learning via Semantic Masking Adaptive Guidance Accelerates Reinforcement Learning of Reasoning Models
Reference 14
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 9ca26ac3-6aa4-4d9d-99a9-9af34efbb577 · inbound
Zone of Proximal Policy Optimization: Teacher in Prompts, Not Gradients Adaptive Guidance Accelerates Reinforcement Learning of Reasoning Models
Reference 94
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 8c9260d2-455f-405f-93e9-c5f57909abad · inbound
It Takes 8 Tokens: Weak-to-Strong Off-Policy RL via Auxiliary Branches Adaptive Guidance Accelerates Reinforcement Learning of Reasoning Models
Reference 24
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 994c961a-99d7-4a9e-8378-21d7cd9ceb46 · inbound
LoRA Scaffolded Policy Optimization (LSPO): A Sampling-Time Low-Rank Scaffold for Recovering Reinforcement-Learning Gradient on Zero-Reward Cliff Prompts Adaptive Guidance Accelerates Reinforcement Learning of Reasoning Models
Reference 16
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 44f73572-76ad-4cac-a237-d43b714ec9a0 · inbound
Beacon: Knowing When and How to Perform Agentic Visual Reasoning Adaptive Guidance Accelerates Reinforcement Learning of Reasoning Models
Reference 18
Source-reported events for the cited work
Unavailable: canonical work link unavailable.