Pith. sign in

Paper Citation Record · LEDGER

Adaptive Guidance Accelerates Reinforcement Learning of Reasoning Models

As of 9 August 2026, this Paper Citation Record lists 58 of 58 outbound references and 11 inbound Pith citation observations for arXiv:2506.13923.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2506.13923 v2

Coverage vector

measured 58 of 58 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-07T00:31:38.586086Z

measured 69 of 69 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-09T06:31:02.800959+00:00

measured 11 of 11 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-04T10:53:08.661292Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-03T20:48:56.210532Z

Reference resolution

58 of 58 outbound references displayed

  • verified exact0
  • verified fuzzy20
  • unresolved38
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 93c69185-51db-403f-b91c-1b1ce5ffd1d1 · outbound

This paper cites OpenAI o1 System Card.

Adaptive Guidance Accelerates Reinforcement Learning of Reasoning Models OpenAI o1 System Card

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-07T00:31:33.630703Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:31:33.630703Z digest=sha256:4f6a0bc2d7ba1afd4fd27d9f503087006ce3cdb96957ec55636c937775c3beb5

Observation a8988f1e-0c2e-4284-bd35-823af1320e74 · outbound

This paper cites DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning.

Adaptive Guidance Accelerates Reinforcement Learning of Reasoning Models DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-07T00:31:33.673149Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:31:33.673149Z digest=sha256:abf3acf8aae62017c72048422240cd0dfa9f423090adf8d63e32e45a0cd9efe1

Observation b7ba5a2e-91c5-42a3-9966-0dd733e5b83c · outbound

This paper cites Tulu 3: Pushing Frontiers in Open Language Model Post-Training.

Adaptive Guidance Accelerates Reinforcement Learning of Reasoning Models Tulu 3: Pushing Frontiers in Open Language Model Post-Training

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-07T00:31:33.763546Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:31:33.763546Z digest=sha256:f955fa08f83f9362253e56cea22991a31f38602019e9e2a94a97f60dd0b01bb8

Observation aeed9b1a-ca68-44e7-a2ec-7aa18495b3a8 · outbound

This paper cites Teaching large language models to reason with reinforcement learning.

Adaptive Guidance Accelerates Reinforcement Learning of Reasoning Models Teaching large language models to reason with reinforcement learning

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-07T00:31:33.824990Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:31:33.824990Z digest=sha256:169718502b6de9e66734617094e0ad0e4527cd831f8fcd3e5dcdfd4c194bb058

Observation d0963d46-c208-49b6-a62c-dfbab7fe6557 · outbound

This paper cites Does Reinforcement Learning Really Incentivize Reasoning Capacity in LLMs Beyond the Base Model?.

Adaptive Guidance Accelerates Reinforcement Learning of Reasoning Models Does Reinforcement Learning Really Incentivize Reasoning Capacity in LLMs Beyond the Base Model?

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-07T00:31:33.883507Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:31:33.883507Z digest=sha256:371fcb9b9ffab23ba78972d69a5bd401a3e06566cc60419293cc60f3674de1f7

Observation 13e84987-c960-42bc-9126-4f976a084cc3 · outbound

This paper cites an unresolved cited work.

Adaptive Guidance Accelerates Reinforcement Learning of Reasoning Models Unresolved cited work

Reference 6

Resolution
unresolved
raw_fallback, observed 2026-08-07T00:31:42.968767Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T00:31:33.976587Z digest=sha256:c095fab2cc93102530f6fe5c8a8810a70a50c006d6df8d97f925255fb09b2e17

Observation e30394ea-1a98-4aed-b777-7263413883c9 · outbound

This paper cites Reinforced Self-Training (ReST) for Language Modeling.

Adaptive Guidance Accelerates Reinforcement Learning of Reasoning Models Reinforced Self-Training (ReST) for Language Modeling

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-07T00:31:34.071107Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:31:34.071107Z digest=sha256:42dd5df1aa62b15e9ac582eecb8a7aa8aa446f7520ca93c115fe1cc83c4743bd

Observation 689b6931-b901-48bb-96f4-67616561938a · outbound

This paper cites Openai’s reinforcement fine-tuning research program, 2024.

Adaptive Guidance Accelerates Reinforcement Learning of Reasoning Models Openai’s reinforcement fine-tuning research program, 2024

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:31:42.742679Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T00:31:34.138304Z digest=sha256:7b4f8e33168bbff8b9c9aa0e3c157538ee314dc2d0cd418846936bb086bf543f

Observation c1892acf-dbe5-43b7-95c3-e78b068ceaea · outbound

This paper cites Be- yond human data: Scaling self-training for problem-solving with language models.Transactions on Machine Learning Research, 2024.

Adaptive Guidance Accelerates Reinforcement Learning of Reasoning Models Be- yond human data: Scaling self-training for problem-solving with language models.Transactions on Machine Learning Research, 2024

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:31:42.566202Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T00:31:34.205487Z digest=sha256:c8d01e115225ce79f35300a0448e156707834e0f5f782db081a4784392d8c931

Observation 77cc6ec1-5c4b-4636-94df-9c093f623c76 · outbound

This paper cites V-star: Training verifiers for self-taught reasoners.

Adaptive Guidance Accelerates Reinforcement Learning of Reasoning Models V-star: Training verifiers for self-taught reasoners

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:31:42.501094Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T00:31:34.323233Z digest=sha256:bc6e7064ae01ab558c68ff842500df714b3f7c13b6502a26fb3f5a9c29ebdc14

Observation c2bb22dc-bb45-4d6b-ab4c-61751a999089 · outbound

This paper cites ReST- MCTS*: LLM self-training via process reward guided tree search.

Adaptive Guidance Accelerates Reinforcement Learning of Reasoning Models ReST- MCTS*: LLM self-training via process reward guided tree search

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:31:42.266560Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T00:31:34.410205Z digest=sha256:e6ca6445d0dbc2f235830f411a28d619304fedd17b57fc51dcd874575156834d

Observation b2a05105-dcec-47c0-9c68-e37529ec670d · outbound

This paper cites Continuous control with deep reinforcement learning.

Adaptive Guidance Accelerates Reinforcement Learning of Reasoning Models Continuous control with deep reinforcement learning

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-07T00:31:34.508482Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:31:34.508482Z digest=sha256:3ec29cc65ed4a2f0fb434baa29f37074cf67506cbc8268a31be1366c0e7bf88d

Observation b0c55e14-5657-40b7-af77-ff320e39af2b · outbound

This paper cites Qwen2.5 Technical Report.

Adaptive Guidance Accelerates Reinforcement Learning of Reasoning Models Qwen2.5 Technical Report

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-07T00:31:34.635249Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:31:34.635249Z digest=sha256:0d455790633bd45b067f3392f123c920dc17320717d2f8029f33b88019d89360

Observation 503ce4ba-3497-4e35-aff9-16f65a6e7112 · outbound

This paper cites Measuring Mathematical Problem Solving With the MATH Dataset.

Adaptive Guidance Accelerates Reinforcement Learning of Reasoning Models Measuring Mathematical Problem Solving With the MATH Dataset

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-07T00:31:34.851692Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:31:34.851692Z digest=sha256:73af62a64cc4c46ff469f376f488e5f22cd5747c282261fdce2e12efd56a6a5d

Observation a94a7577-dd41-45a5-befa-721617b21833 · outbound

This paper cites Aime 2024.

Adaptive Guidance Accelerates Reinforcement Learning of Reasoning Models Aime 2024

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:31:42.081687Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T00:31:34.924828Z digest=sha256:222033952e24f50fa617520a5bc54c1c0f2e9b3b42155ed6abb507257b31ed1d

Observation 27f7e066-ca93-4f55-bde1-e8f978f7cb64 · outbound

This paper cites Aime 2025.

Adaptive Guidance Accelerates Reinforcement Learning of Reasoning Models Aime 2025

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:31:42.010708Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T00:31:35.021740Z digest=sha256:342e5880cc4369483b29e5cf5317fd6fd08f5e7df1062f82347f3de423d5d831

Observation 5a8cccb6-bcf6-462a-af42-818c0d348387 · outbound

This paper cites Amc 2023.

Adaptive Guidance Accelerates Reinforcement Learning of Reasoning Models Amc 2023

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:31:41.885367Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T00:31:35.145729Z digest=sha256:6bf15a40808e1d674d944a156c7c926a2609398368671c9572be1f249865f776

Observation 07087427-359a-4e9b-9f33-1c67cb2dfdfa · outbound

This paper cites Gpqa: A graduate-level google-proof q&a benchmark.

Adaptive Guidance Accelerates Reinforcement Learning of Reasoning Models Gpqa: A graduate-level google-proof q&a benchmark

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-07T00:31:35.213884Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:31:35.213884Z digest=sha256:565584635a2c35b0edbc8abc15388d997ac60ea7b73e2cf756f7f9195d5e4b5f

Observation c05f9dc9-279f-4538-b816-a2bec136b49b · outbound

This paper cites OlympiadBench: A Challenging Benchmark for Promoting AGI with Olympiad-Level Bilingual Multimodal Scientific Problems.

Adaptive Guidance Accelerates Reinforcement Learning of Reasoning Models OlympiadBench: A Challenging Benchmark for Promoting AGI with Olympiad-Level Bilingual Multimodal Scientific Problems

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-07T00:31:35.273503Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:31:35.273503Z digest=sha256:1a3a59e7a2d56942569f66512541a1b76ea6f6bb41acb2dd16114457bc4ff2c7

Observation beb9fd64-2327-4d99-9065-6f3e2b8f317e · outbound

This paper cites DeepSeek-Coder: When the Large Language Model Meets Programming -- The Rise of Code Intelligence.

Adaptive Guidance Accelerates Reinforcement Learning of Reasoning Models DeepSeek-Coder: When the Large Language Model Meets Programming -- The Rise of Code Intelligence

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-07T00:31:35.407856Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:31:35.407856Z digest=sha256:8cea4854cbdf92962877b1d71dd1e25731e46010355b9e071d72bd5650b01590

Observation 2577b0f8-e45f-4c7a-a99f-bf593af79254 · outbound

This paper cites Livecodebench: Holistic and contamination free evaluation of large language models for code.arXiv preprint, 2024.

Adaptive Guidance Accelerates Reinforcement Learning of Reasoning Models Livecodebench: Holistic and contamination free evaluation of large language models for code.arXiv preprint, 2024

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:31:41.787072Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T00:31:35.477463Z digest=sha256:5a66992e2b466fa889ca9c069d49755e5eec1f9927c5537e6d49dde9378961dd

Observation 098d9765-40b6-4420-900c-835487736344 · outbound

This paper cites Evaluating Large Language Models Trained on Code.

Adaptive Guidance Accelerates Reinforcement Learning of Reasoning Models Evaluating Large Language Models Trained on Code

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-07T00:31:35.563471Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:31:35.563471Z digest=sha256:a3dc1c56d153f19391d5c08f16fd9ccec55b38ac9a2f6a7a30293c4a8ead5f50

Observation 35b5e4f3-a3a6-487f-a86d-414cb3cac3d3 · outbound

This paper cites Open r1: A fully open reproduction of deepseek-r1, January 2025.

Adaptive Guidance Accelerates Reinforcement Learning of Reasoning Models Open r1: A fully open reproduction of deepseek-r1, January 2025

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-07T00:31:35.647369Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:31:35.647369Z digest=sha256:32da7f98244f913aa549eb2ee3c58b4a05d0e1c70a605c8d1fb4e251cde1f209

Observation a9ff5d5f-0123-4e33-b672-8e7b2dd1c244 · outbound

This paper cites Dapo: An open-source llm reinforcement learning system at scale, 2025.

Adaptive Guidance Accelerates Reinforcement Learning of Reasoning Models Dapo: An open-source llm reinforcement learning system at scale, 2025

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-07T00:31:35.732089Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:31:35.732089Z digest=sha256:73cf1d3acc8b034da9a2c6da078ddc5826b5d2a07749be3c1100e987c4460fac

Observation c7094ab0-9a3c-441e-b0b6-b7e5c9eb3c5e · outbound

This paper cites Fine-Tuning Language Models from Human Preferences.

Adaptive Guidance Accelerates Reinforcement Learning of Reasoning Models Fine-Tuning Language Models from Human Preferences

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-07T00:31:35.831960Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:31:35.831960Z digest=sha256:90fd090f35dda4a662e91db780df5761a969ac9e4a629332f9bb36473ef8363b

Observation 97b7a885-7d76-403d-8754-9278b13936d2 · outbound

This paper cites an unresolved cited work.

Adaptive Guidance Accelerates Reinforcement Learning of Reasoning Models Unresolved cited work

Reference 27

Resolution
unresolved
raw_fallback, observed 2026-08-07T00:31:41.628336Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T00:31:35.904313Z digest=sha256:e40a768def504e12970a279de212bbd8655abf4d916a569ce9ebd195dba23369

Observation a9c0a123-1458-4ec2-ad46-403029fe2937 · outbound

This paper cites LMRL Gym: Benchmarks for Multi-Turn Reinforcement Learning with Language Models.

Adaptive Guidance Accelerates Reinforcement Learning of Reasoning Models LMRL Gym: Benchmarks for Multi-Turn Reinforcement Learning with Language Models

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-07T00:31:36.028658Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:31:36.028658Z digest=sha256:0ffa287224e8781cfde829739a7953a75ff3dca15acc79578ad645a421b32336

Observation c63c58de-bbae-4044-892f-a21618e9f075 · outbound

This paper cites Archer: Training language model agents via hierarchical multi-turn rl.

Adaptive Guidance Accelerates Reinforcement Learning of Reasoning Models Archer: Training language model agents via hierarchical multi-turn rl

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:31:41.511850Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T00:31:36.112436Z digest=sha256:618ae47c213ced2d6b10a3390a8b08bcce127bc018be89d7168aeb288edd7b1d

Observation 9a9491a7-4fca-4fb7-b1ee-481b3c08470d · outbound

This paper cites Aligning large multimodal models with factually augmented rlhf.

Adaptive Guidance Accelerates Reinforcement Learning of Reasoning Models Aligning large multimodal models with factually augmented rlhf

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:31:41.386519Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T00:31:36.165728Z digest=sha256:58ac117fb4bf8138b901bcfbc733ee97beef2f911efccb3efc4d5168ad273c28

Observation 63cc29ad-0a40-48b2-b947-c9ab58369e60 · outbound

This paper cites Kimi k1.5: Scaling Reinforcement Learning with LLMs.

Adaptive Guidance Accelerates Reinforcement Learning of Reasoning Models Kimi k1.5: Scaling Reinforcement Learning with LLMs

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-07T00:31:36.250869Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:31:36.250869Z digest=sha256:0520ea563c831f5978ff0a21d8733ecb8ed04c0bf6fe7de7b097fb4d9ab67e4b

Observation 62593bd3-d18c-4b33-b0fb-2ab464d70a43 · outbound

This paper cites Training Verifiers to Solve Math Word Problems.

Adaptive Guidance Accelerates Reinforcement Learning of Reasoning Models Training Verifiers to Solve Math Word Problems

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-07T00:31:36.352872Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:31:36.352872Z digest=sha256:2b2a467aa4b947e9aa5e70d2633533f1396040c416378682051da3a95303832d

Observation a56e454b-536e-4593-a6ea-861ee4fba24f · outbound

This paper cites Program Synthesis with Large Language Models.

Adaptive Guidance Accelerates Reinforcement Learning of Reasoning Models Program Synthesis with Large Language Models

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-07T00:31:36.490405Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:31:36.490405Z digest=sha256:4be4d738458d62fcc94bcedf29d21bc74cdd27578f98ea24ecd1d03fa4591dbb

Observation 43b36156-2531-46f0-b9dd-afe892cd943c · outbound

This paper cites RLEF: Grounding Code LLMs in Execution Feedback with Reinforcement Learning.

Adaptive Guidance Accelerates Reinforcement Learning of Reasoning Models RLEF: Grounding Code LLMs in Execution Feedback with Reinforcement Learning

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-07T00:31:36.562765Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:31:36.562765Z digest=sha256:b62051157a42f4ce9d7af060ba88589f6478fc572c2c773f245c4f08098d4690

Observation 0e997cca-1852-4654-a328-b4b1db39ef8c · outbound

This paper cites SFT Memorizes, RL Generalizes: A Comparative Study of Foundation Model Post-training.

Adaptive Guidance Accelerates Reinforcement Learning of Reasoning Models SFT Memorizes, RL Generalizes: A Comparative Study of Foundation Model Post-training

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-07T00:31:36.601300Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:31:36.601300Z digest=sha256:622329e15a2457b39fc6951f5d342eff20c55c9b2e1bc55a6efd06bfcf414255

Observation 386ef46f-11a8-45a7-9bcd-9bf7868f5487 · outbound

This paper cites Scaling LLM test-time com- pute optimally can be more effective than scaling parameters for reasoning.

Adaptive Guidance Accelerates Reinforcement Learning of Reasoning Models Scaling LLM test-time com- pute optimally can be more effective than scaling parameters for reasoning

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-07T00:31:36.691691Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:31:36.691691Z digest=sha256:81edc37fa589ff0f350de92f5de6cd7618fefd53ba0021340ec850a4f770849e

Observation 3f71c7fa-5d4b-4ef8-a912-a461dd82f61c · outbound

This paper cites Echo Chamber: RL Post-training Amplifies Behaviors Learned in Pretraining.

Adaptive Guidance Accelerates Reinforcement Learning of Reasoning Models Echo Chamber: RL Post-training Amplifies Behaviors Learned in Pretraining

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-07T00:31:36.795096Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:31:36.795096Z digest=sha256:c5c547d20883039460511a138094100ca46a2314224ff73b8b7ec6b60ae7507c

Observation 9fb8845d-8e9c-4e18-981d-6214dbf83a0d · outbound

This paper cites Assessing diversity collapse in reasoning.

Adaptive Guidance Accelerates Reinforcement Learning of Reasoning Models Assessing diversity collapse in reasoning

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-07T00:31:36.881955Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:31:36.881955Z digest=sha256:1bbed6f69ef18ef08f0e84d6e311ae64359e52b312ccee01710215ec26586e7b

Observation fee6f11f-2095-4421-a497-0c9994c674eb · outbound

This paper cites Direct preference optimization: Your language model is secretly a reward model.

Adaptive Guidance Accelerates Reinforcement Learning of Reasoning Models Direct preference optimization: Your language model is secretly a reward model

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-07T00:31:36.976302Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:31:36.976302Z digest=sha256:5df40053a378e609bf88429d8c74207abb86d876d431a18ae63ce808d56a9ee3

Observation 699988d1-d4a5-403a-b834-fa3f35c6c9b9 · outbound

This paper cites Tapered Off-Policy REINFORCE: Stable and efficient reinforcement learning for LLMs.

Adaptive Guidance Accelerates Reinforcement Learning of Reasoning Models Tapered Off-Policy REINFORCE: Stable and efficient reinforcement learning for LLMs

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-07T00:31:37.043476Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:31:37.043476Z digest=sha256:a349c49499469b7c44b2b7d264f62244df90f0b5eb7af0be4f23edf79863d769

Observation eaf1f6ed-4b5c-439f-ad65-cf6251c95323 · outbound

This paper cites Learning to Reason under Off-Policy Guidance.

Adaptive Guidance Accelerates Reinforcement Learning of Reasoning Models Learning to Reason under Off-Policy Guidance

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-07T00:31:37.128523Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:31:37.128523Z digest=sha256:7eec3b8f70d16b845b7684083c5d812998039307304ba46f0751ea3e87a84569

Observation 3294b235-12d3-4aea-af15-30508f14116f · outbound

This paper cites HybridFlow: A Flexible and Efficient RLHF Framework.

Adaptive Guidance Accelerates Reinforcement Learning of Reasoning Models HybridFlow: A Flexible and Efficient RLHF Framework

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-07T00:31:37.204240Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:31:37.204240Z digest=sha256:3c17f1cecf98aac425ead089e3107b70f7e0c14e455e858e51550d15c5584c48

Observation 2cbc1157-af60-4496-8d47-c2d7ac995883 · outbound

This paper cites an unresolved cited work.

Adaptive Guidance Accelerates Reinforcement Learning of Reasoning Models Unresolved cited work

Reference 43

Resolution
unresolved
raw_fallback, observed 2026-08-07T00:31:41.227473Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T00:31:37.313993Z digest=sha256:a637ae223d85117dc1c0e9f7276a25c6d7d8da767dfcfc5ff1ef37b214e3ca7d

Observation 0b287539-eb83-4218-bd02-e499e193e097 · outbound

This paper cites an unresolved cited work.

Adaptive Guidance Accelerates Reinforcement Learning of Reasoning Models Unresolved cited work

Reference 44

Resolution
unresolved
raw_fallback, observed 2026-08-07T00:31:41.116242Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T00:31:37.393926Z digest=sha256:3a5e75d211cc897586c399c0897f42774748711820a797b861514fcafbe5e6e7

Observation 3ed7525f-09d6-411a-9a85-7468a3918160 · outbound

This paper cites an unresolved cited work.

Adaptive Guidance Accelerates Reinforcement Learning of Reasoning Models Unresolved cited work

Reference 45

Resolution
unresolved
raw_fallback, observed 2026-08-07T00:31:40.987925Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T00:31:37.478330Z digest=sha256:a00bff5c11f487d155f8fab457a81057de75a2fe18a395b5eacf5133420a645a

Observation b6d942a2-c234-4217-98f2-f56e1763451c · outbound

This paper cites an unresolved cited work.

Adaptive Guidance Accelerates Reinforcement Learning of Reasoning Models Unresolved cited work

Reference 46

Resolution
unresolved
raw_fallback, observed 2026-08-07T00:31:40.908596Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T00:31:37.558393Z digest=sha256:638071e578fb9d6d1860b3a2fbddda6a7dbcafb99d98ab85d259663bcf831b96

Observation 2bc9a944-19ff-4b71-bdeb-cc3b8a69008b · outbound

This paper cites A hint to the problem is provided below: [HINT_START].

Adaptive Guidance Accelerates Reinforcement Learning of Reasoning Models A hint to the problem is provided below: [HINT_START]

Reference 47

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:31:40.776609Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T00:31:37.607764Z digest=sha256:c97fb34b2f5d922d8666bdd1564731691661565988962b7daa5c07780e0e7a1e

Observation 85b570b4-9fb5-485c-b190-aab6de5316ce · outbound

This paper cites Think about how the identity (a²/4) - (4/a²) might be used as a building block for factoring the larger expression.

Adaptive Guidance Accelerates Reinforcement Learning of Reasoning Models Think about how the identity (a²/4) - (4/a²) might be used as a building block for factoring the larger expression

Reference 48

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:31:40.646643Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T00:31:37.657446Z digest=sha256:5287bfd93ae6e76f7b86b4af7c63d11f8bd6db8ae790c83fc8e80ec126aa72d3

Observation e13d6c72-369a-437f-ad8c-5d16f7844303 · outbound

This paper cites Ask yourself if a difference of powers or a recognizable factorization pattern might help connect the two parts of the expression.

Adaptive Guidance Accelerates Reinforcement Learning of Reasoning Models Ask yourself if a difference of powers or a recognizable factorization pattern might help connect the two parts of the expression

Reference 49

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:31:40.472946Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T00:31:37.765087Z digest=sha256:db05778d4868c574291b66fd39e2a8845ab48efa10b751bd24da9d9d8f68f6da

Observation a818a75b-c756-4365-9c30-8c403835893b · outbound

This paper cites How can this substitution simplify the structure of the problem? ,→ ,→.

Adaptive Guidance Accelerates Reinforcement Learning of Reasoning Models How can this substitution simplify the structure of the problem? ,→ ,→

Reference 50

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:31:40.307652Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T00:31:37.850950Z digest=sha256:deb08743dcec419a62bca684237d87c7b8281892056cb465ddc8ce8b98c95564

Observation ea984d5a-4322-41c4-93a6-29d8d3507596 · outbound

This paper cites an unresolved cited work.

Adaptive Guidance Accelerates Reinforcement Learning of Reasoning Models Unresolved cited work

Reference 51

Resolution
unresolved
raw_fallback, observed 2026-08-07T00:31:40.185924Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T00:31:37.940379Z digest=sha256:89b0a5129be057f93f1ead0352f36c72c19aeeb60adfae37ec2662b15f014f84

Observation 5ecf4a51-a6e2-4c78-aa2a-0f184f08c3db · outbound

This paper cites Use these observations to guide your step-by-step approach toward the final simplified result.

Adaptive Guidance Accelerates Reinforcement Learning of Reasoning Models Use these observations to guide your step-by-step approach toward the final simplified result

Reference 52

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:31:40.020913Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T00:31:38.063163Z digest=sha256:8cb9e6a608948180fa69758d2c021cd6ffed536c22e0526f902ec50ea61115cf

Observation c4566f0d-4b2d-4f7c-a2b6-124fc1a5569b · outbound

This paper cites an unresolved cited work.

Adaptive Guidance Accelerates Reinforcement Learning of Reasoning Models Unresolved cited work

Reference 53

Resolution
unresolved
raw_fallback, observed 2026-08-07T00:31:39.842436Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T00:31:38.128400Z digest=sha256:b3198021915b2bae2d29ed7a28703e093a65891ddfbcda3a8180a8e31e5d7850

Observation d504db0d-c856-45e5-804c-5eb4622496fe · outbound

This paper cites an unresolved cited work.

Adaptive Guidance Accelerates Reinforcement Learning of Reasoning Models Unresolved cited work

Reference 54

Resolution
unresolved
raw_fallback, observed 2026-08-07T00:31:39.762269Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T00:31:38.201603Z digest=sha256:dd82ec8350efb068d4a385615bf46ab18da1695e4b4429e3476d2478b8f3d66c

Observation b03f974b-9501-4cfc-9c19-0e6a97c4b804 · outbound

This paper cites What does that imply about multiplying the side length? ,→ ,→.

Adaptive Guidance Accelerates Reinforcement Learning of Reasoning Models What does that imply about multiplying the side length? ,→ ,→

Reference 55

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:31:39.645444Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T00:31:38.255892Z digest=sha256:1306fef9380eed8c5d799e5e4ecca1ff0e3dd46d257492c8fa246c44dab63b17

Observation e8681102-13af-4343-9de8-61d3a7d873a2 · outbound

This paper cites Ensure each step follows from the properties of a square.,→.

Adaptive Guidance Accelerates Reinforcement Learning of Reasoning Models Ensure each step follows from the properties of a square.,→

Reference 56

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:31:39.495423Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T00:31:38.345981Z digest=sha256:5a01a6bc55f27a45094eaaab04f86ab97fae1ff7fe2425be2c25efaa8f8d7b51

Observation a5312c55-30df-4e3f-bd80-b29e1640dc0b · outbound

This paper cites using the hint.

Adaptive Guidance Accelerates Reinforcement Learning of Reasoning Models using the hint

Reference 57

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:31:39.367706Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T00:31:38.429016Z digest=sha256:2a99d773ccf2aa6dd43d801bd49a2b87cc408c2367cfc023ebcab533b91802da

Observation 79e676a7-60b4-4be4-b507-194276321478 · outbound

This paper cites This example shows how a concise, domain-specific hint can redirect the model’s reasoning and correct a systematic geometric error.

Adaptive Guidance Accelerates Reinforcement Learning of Reasoning Models This example shows how a concise, domain-specific hint can redirect the model’s reasoning and correct a systematic geometric error

Reference 58

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:31:39.190927Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T00:31:38.509431Z digest=sha256:7dffc3753a8b3a743a64306c5e03fd611de0f700490f5eab4a3495676dd69730

Observation 9efcd0cd-68c1-4bd7-bcb7-2855da000ad5 · outbound

This paper cites <think>\n {thoughts} </think>\n\.

Adaptive Guidance Accelerates Reinforcement Learning of Reasoning Models <think>\n {thoughts} </think>\n\

Reference 59

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:31:39.011456Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T00:31:38.586086Z digest=sha256:f64ccf1cdf6859eaf65d88bc869c89498270dff94122ba8beec7615c07c9a6af

Pith citing papers

Observation ae143971-845e-4b2c-a967-2259fae6503a · inbound

EvoCoT: Overcoming the Exploration Bottleneck in Reinforcement Learning cites this paper.

EvoCoT: Overcoming the Exploration Bottleneck in Reinforcement Learning Adaptive Guidance Accelerates Reinforcement Learning of Reasoning Models

Reference 11

Resolution
verified exact
arxiv_id, observed 2026-05-18T23:56:55.008684Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-18T23:56:47.274169Z digest=sha256:63627276c79196a629a622702b6b8cc92debe1f2fa59a50c7bb8b32b92e89445

Observation a5a87b9b-3829-498b-a3b9-57b85d3d4566 · inbound

Position: The Hidden Costs and Measurement Gaps of Reinforcement Learning with Verifiable Rewards cites this paper.

Position: The Hidden Costs and Measurement Gaps of Reinforcement Learning with Verifiable Rewards Adaptive Guidance Accelerates Reinforcement Learning of Reasoning Models

Reference 25

Resolution
metadata mismatch
arxiv_id, observed 2026-05-18T14:26:28.367388Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-18T14:24:48.666197Z digest=sha256:3e88be1af8ce8fce978063202841a2fa1b517f97fc7f969494fb0083ec60e5ad

Observation 11411b44-0b68-42f9-aeb6-eab51a369da2 · inbound

TaoSR-AGRL: Adaptive Guided Reinforcement Learning Framework for E-commerce Search Relevance cites this paper.

TaoSR-AGRL: Adaptive Guided Reinforcement Learning Framework for E-commerce Search Relevance Adaptive Guidance Accelerates Reinforcement Learning of Reasoning Models

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-04T10:53:08.661292Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T10:53:08.661292Z digest=sha256:677de1dd6bc1c225acfb5c200c1cf1fa8cc4cf4cc9192052af55a66582e6178c

Observation c4ef3d49-52e8-4f36-943f-efa6b001fcff · inbound

Selective Off-Policy Reference Tuning with Plan Guidance cites this paper.

Selective Off-Policy Reference Tuning with Plan Guidance Adaptive Guidance Accelerates Reinforcement Learning of Reasoning Models

Reference 18

Resolution
metadata mismatch
arxiv_id, observed 2026-05-13T01:47:05.199208Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-13T01:28:18.615371Z digest=sha256:c9e407547b379e414c8eb54c0d605642c18ea5561f51590ca31392bc0777c855

Observation f6f35c7f-8dc4-4247-968c-0dd26bff82cf · inbound

Selective Off-Policy Reference Tuning with Plan Guidance cites this paper.

Selective Off-Policy Reference Tuning with Plan Guidance Adaptive Guidance Accelerates Reinforcement Learning of Reasoning Models

Reference 18

Resolution
metadata mismatch
arxiv_id, observed 2026-05-14T21:22:59.342700Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-14T21:20:24.066520Z digest=sha256:0364310ef78dca54584274193e5f9c850c893d3fc2970b2890e1527cfcf23cd3

Observation 2cf22fff-26b7-4055-817f-97f6907687bc · inbound

Learning Agentic Policy from Action Guidance cites this paper.

Learning Agentic Policy from Action Guidance Adaptive Guidance Accelerates Reinforcement Learning of Reasoning Models

Reference 42

Resolution
metadata mismatch
arxiv_id, observed 2026-05-13T05:07:17.751601Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-13T05:02:49.206053Z digest=sha256:62ad4db0c36922949bf064c62de8f37dfeadb358721179b352c858e8c6013792

Observation 155a57b7-919a-4bda-89cb-84696c4310ee · inbound

Hide to Guide: Learning via Semantic Masking cites this paper.

Hide to Guide: Learning via Semantic Masking Adaptive Guidance Accelerates Reinforcement Learning of Reasoning Models

Reference 14

Resolution
metadata mismatch
arxiv_id, observed 2026-06-30T12:34:39.107566Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-06-30T12:16:12.108715Z digest=sha256:69870f55d6170b1637f6b731c9bbf819ab36e33d3988092e1da99374265c67f5

Observation 9ca26ac3-6aa4-4d9d-99a9-9af34efbb577 · inbound

Zone of Proximal Policy Optimization: Teacher in Prompts, Not Gradients cites this paper.

Zone of Proximal Policy Optimization: Teacher in Prompts, Not Gradients Adaptive Guidance Accelerates Reinforcement Learning of Reasoning Models

Reference 94

Resolution
metadata mismatch
arxiv_id, observed 2026-07-03T20:48:56.211961Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-06-27T01:08:52.981296Z digest=sha256:b86ca2ef4abc1436f7917cf1a4a1f829d5ac9d58a60dd213bfc6cc89256855c3

Observation 8c9260d2-455f-405f-93e9-c5f57909abad · inbound

It Takes 8 Tokens: Weak-to-Strong Off-Policy RL via Auxiliary Branches cites this paper.

It Takes 8 Tokens: Weak-to-Strong Off-Policy RL via Auxiliary Branches Adaptive Guidance Accelerates Reinforcement Learning of Reasoning Models

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-02T14:45:22.917496Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-02T14:45:22.917496Z digest=sha256:1881513d804f24007f512ee4c68bc6221a5296ccf6bce20dfd8c3e6aebc61031

Observation 994c961a-99d7-4a9e-8378-21d7cd9ceb46 · inbound

LoRA Scaffolded Policy Optimization (LSPO): A Sampling-Time Low-Rank Scaffold for Recovering Reinforcement-Learning Gradient on Zero-Reward Cliff Prompts cites this paper.

LoRA Scaffolded Policy Optimization (LSPO): A Sampling-Time Low-Rank Scaffold for Recovering Reinforcement-Learning Gradient on Zero-Reward Cliff Prompts Adaptive Guidance Accelerates Reinforcement Learning of Reasoning Models

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-01T01:28:56.693276Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T01:28:56.693276Z digest=sha256:1520d407d0934d466f607930633908ec9ac7fc095dd531e1a1d0471379bc0836

Observation 44f73572-76ad-4cac-a237-d43b714ec9a0 · inbound

Beacon: Knowing When and How to Perform Agentic Visual Reasoning cites this paper.

Beacon: Knowing When and How to Perform Agentic Visual Reasoning Adaptive Guidance Accelerates Reinforcement Learning of Reasoning Models

Reference 18

Resolution
unresolved
no resolver link, observed 2026-07-31T02:45:28.621215Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-31T02:45:28.621215Z digest=sha256:1aa8621f67bf365cb60116715ef3196282a2ac31638d008a809a3daf2c6f8222