Pith. sign in

Paper Citation Record · LEDGER

Scaling Up RL: Unlocking Diverse Reasoning in LLMs via Prolonged Training

As of 19 August 2026, this Paper Citation Record lists 33 of 33 outbound references and 10 inbound Pith citation observations for arXiv:2507.12507.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2507.12507 v1

Coverage vector

measured 33 of 33 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-06T16:51:43.099838Z

measured 43 of 43 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-19T06:32:44.657259+00:00

measured 10 of 10 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-15T18:21:08.555405Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-01T20:56:13.554863Z

Reference resolution

33 of 33 outbound references displayed

  • verified exact0
  • verified fuzzy14
  • unresolved19
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 01702ac7-f90e-41df-ad01-e3366eed0b87 · outbound

This paper cites OpenAI o1 System Card.

Scaling Up RL: Unlocking Diverse Reasoning in LLMs via Prolonged Training OpenAI o1 System Card

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-06T16:51:41.308878Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T16:51:41.308878Z digest=sha256:22fe9685be4a774a5a65f120527832bd6b72480beaa63fcc40cec022aec1ec5a

Observation b91c1ff4-f3e1-4e84-836e-c7ef764722fb · outbound

This paper cites DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning.

Scaling Up RL: Unlocking Diverse Reasoning in LLMs via Prolonged Training DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-06T16:51:41.379398Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T16:51:41.379398Z digest=sha256:4ef27d8f55d8d67d94528669a4a440cf668cdee8adf6171f282bf520eb1e6e25

Observation 9cd6446e-a56c-4d76-832e-a6d630ed29b0 · outbound

This paper cites Tang, Manan Roongta, Colin Cai, Jef- frey Luo, Li Erran Li, Raluca Ada Popa, and Ion Stoica.

Scaling Up RL: Unlocking Diverse Reasoning in LLMs via Prolonged Training Tang, Manan Roongta, Colin Cai, Jef- frey Luo, Li Erran Li, Raluca Ada Popa, and Ion Stoica

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T16:51:44.536753Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-06T16:51:41.478402Z digest=sha256:380850abdc3de56288a2e756e331141d6c96beb377870a10e6920d84b8cb96c1

Observation ae443173-a672-4bb6-87ec-464a477d3bef · outbound

This paper cites Dapo: An open-source llm reinforcement learning system at scale, 2025.

Scaling Up RL: Unlocking Diverse Reasoning in LLMs via Prolonged Training Dapo: An open-source llm reinforcement learning system at scale, 2025

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-06T16:51:41.605352Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T16:51:41.605352Z digest=sha256:544e43be123bca9c4137f6f5b583cc68b02951dcf19ba3678dbf7294e315ac2d

Observation d72de7c8-7f04-4e43-b906-ee8f2bc20d39 · outbound

This paper cites Process Reinforcement through Implicit Rewards.

Scaling Up RL: Unlocking Diverse Reasoning in LLMs via Prolonged Training Process Reinforcement through Implicit Rewards

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-06T16:51:41.675432Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T16:51:41.675432Z digest=sha256:f03736c14b9809084d681ac25fb6e0cf9c753413ca3279578cff32a3928cc5ea

Observation e3eded81-4d23-4ab7-abbf-2568a743311f · outbound

This paper cites Scp-116k: A high-quality problem-solution dataset and a generalized pipeline for automated extraction in the higher education science domain, 2025.

Scaling Up RL: Unlocking Diverse Reasoning in LLMs via Prolonged Training Scp-116k: A high-quality problem-solution dataset and a generalized pipeline for automated extraction in the higher education science domain, 2025

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-06T16:51:41.754532Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T16:51:41.754532Z digest=sha256:0d5ff4b3cd9193994042280fb780d368af182e6f1672e1d7102fd4449b9911a4

Observation b7202fe5-07b3-4c21-8477-6aec33c4b1fc · outbound

This paper cites Reasoning gym: Reasoning environments for reinforcement learning with verifiable rewards, 2025.

Scaling Up RL: Unlocking Diverse Reasoning in LLMs via Prolonged Training Reasoning gym: Reasoning environments for reinforcement learning with verifiable rewards, 2025

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T16:51:44.440037Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-06T16:51:41.839172Z digest=sha256:7aaa500bb63dfebe5d87b1b83930173b4c37eaf730cac867d9252f5b0b52395d

Observation e2d45004-e28f-4f90-b895-4a4bf75b0781 · outbound

This paper cites an unresolved cited work.

Scaling Up RL: Unlocking Diverse Reasoning in LLMs via Prolonged Training Unresolved cited work

Reference 8

Resolution
unresolved
raw_fallback, observed 2026-08-06T16:51:44.361140Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-06T16:51:41.927847Z digest=sha256:f2a857601e16d99c91444f71c80c8924780035c6b5104fdee4a25c43d748e91e

Observation 6c2e0656-9b38-4af2-8d61-d9e6fb88c713 · outbound

This paper cites Instruction-following evaluation for large language models, 2023.

Scaling Up RL: Unlocking Diverse Reasoning in LLMs via Prolonged Training Instruction-following evaluation for large language models, 2023

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-06T16:51:41.986503Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T16:51:41.986503Z digest=sha256:ef0c9dee7c869b74b854883ec0efdbbd8c9a5d3a6b9a27301c1d194c882b885d

Observation 73910816-2529-43d4-ba57-7b29a62e8611 · outbound

This paper cites DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models.

Scaling Up RL: Unlocking Diverse Reasoning in LLMs via Prolonged Training DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-06T16:51:42.072510Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T16:51:42.072510Z digest=sha256:c8882f96d7cd1c5b711c74faf9419a630b30d8a260aaa37369f08de52b50945c

Observation 38a9ed1a-7369-4098-9a0b-75cbc62aa494 · outbound

This paper cites Proximal policy optimization algorithms, 2017.

Scaling Up RL: Unlocking Diverse Reasoning in LLMs via Prolonged Training Proximal policy optimization algorithms, 2017

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-06T16:51:42.132236Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T16:51:42.132236Z digest=sha256:16dfe2120f5707a0666f76ad5d43430e14de25fd35dc91e47814765c3a5cfd55

Observation 44a62020-97b4-44bf-b019-693d73a66d65 · outbound

This paper cites Approximating KL Divergence.

Scaling Up RL: Unlocking Diverse Reasoning in LLMs via Prolonged Training Approximating KL Divergence

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T16:51:44.270150Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-06T16:51:42.192659Z digest=sha256:6acc2bcda732c27297c6b6253a7f1e0544983669c382f0c6dc2e8bf0ff0cc6e5

Observation e9f94756-acd8-4fa3-9f76-d3a664121d0b · outbound

This paper cites Deepcoder: A fully open-source 14b coder at o3-mini level.

Scaling Up RL: Unlocking Diverse Reasoning in LLMs via Prolonged Training Deepcoder: A fully open-source 14b coder at o3-mini level

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T16:51:44.190811Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-06T16:51:42.236452Z digest=sha256:f87ef10fbb5101d769f7476609a312af9029f8f4214ae30d9e701adf2797070b

Observation 5036a461-add3-44f6-971b-82c784a0fff8 · outbound

This paper cites Vapo: Efficient and reliable reinforcement learning for advanced reasoning tasks, 2025.

Scaling Up RL: Unlocking Diverse Reasoning in LLMs via Prolonged Training Vapo: Efficient and reliable reinforcement learning for advanced reasoning tasks, 2025

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T16:51:44.120691Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-06T16:51:42.281094Z digest=sha256:941dc55e6fadde3aca2378e0e90b18612899a5a7bc9b72e5b46810538f1a4303

Observation 06296c41-5472-4ca1-a519-acb330813ae0 · outbound

This paper cites Skywork open reasoner series.https://capricious-hydrogen-41c .notion.site/Skywork-Open-Reaonser- Series-1d0bc9ae823a80459b46c149e4f51680, 2025.

Scaling Up RL: Unlocking Diverse Reasoning in LLMs via Prolonged Training Skywork open reasoner series.https://capricious-hydrogen-41c .notion.site/Skywork-Open-Reaonser- Series-1d0bc9ae823a80459b46c149e4f51680, 2025

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T16:51:44.066802Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-06T16:51:42.312913Z digest=sha256:861c0aab5d76f73576c9d1c3aac8d2d053f22298f4f0f81d067b02aacaf12dc4

Observation 9fb672b9-d13d-4e17-97ff-e2bb93fea91c · outbound

This paper cites Hybridflow: A flexible and efficient rlhf framework.

Scaling Up RL: Unlocking Diverse Reasoning in LLMs via Prolonged Training Hybridflow: A flexible and efficient rlhf framework

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T16:51:43.928142Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-06T16:51:42.366724Z digest=sha256:a5cd2e27b59b3ba2986f17e61f6b7ec18ec64d175f27491a75f3181f3ead9a77

Observation 1f6e9583-fd27-417e-ae19-f98d530e3e09 · outbound

This paper cites Decoupled weight decay regularization, 2019.

Scaling Up RL: Unlocking Diverse Reasoning in LLMs via Prolonged Training Decoupled weight decay regularization, 2019

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-06T16:51:42.428615Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T16:51:42.428615Z digest=sha256:3e821b6dd5260d66e18a263b0f078618dec81a163958bb7ddcab25454379315a

Observation 4371198a-3b07-42c5-a489-2af2baca32d0 · outbound

This paper cites 7b model and 8k examples: Emerging reasoning with reinforcement learning is both effective and efficient.https://hkust-nlp.notion.site/ simplerl-reason, 2025.

Scaling Up RL: Unlocking Diverse Reasoning in LLMs via Prolonged Training 7b model and 8k examples: Emerging reasoning with reinforcement learning is both effective and efficient.https://hkust-nlp.notion.site/ simplerl-reason, 2025

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T16:51:43.806690Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-06T16:51:42.481817Z digest=sha256:1c32c1ab55736fba44677caf3d4af3dc112e818f05ff6ad7812151c0def940b1

Observation 9ec5ac18-d477-4187-bdce-9003b7eb8312 · outbound

This paper cites American invitational mathematics examination - aime.

Scaling Up RL: Unlocking Diverse Reasoning in LLMs via Prolonged Training American invitational mathematics examination - aime

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T16:51:43.711163Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-06T16:51:42.511720Z digest=sha256:e837541aebbcd81041646c4a2ba4186480a82b51e4a5434ca8a37e5f83feefda

Observation 62f2ad05-1b60-44e6-b401-05a5ffa833e7 · outbound

This paper cites American invitational mathematics examination - aime.

Scaling Up RL: Unlocking Diverse Reasoning in LLMs via Prolonged Training American invitational mathematics examination - aime

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T16:51:43.608952Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-06T16:51:42.546714Z digest=sha256:f7c922f8aacf9df92cedf75be4891dd6b67ae7068af74eb3c43ac042c6ff658a

Observation ea63ccc2-ada3-4e43-b86d-47132b004ac6 · outbound

This paper cites American mathematics competition - amc.

Scaling Up RL: Unlocking Diverse Reasoning in LLMs via Prolonged Training American mathematics competition - amc

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T16:51:43.543312Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-06T16:51:42.576951Z digest=sha256:b862e15fdb1e17c262ac1c03ec28be97f71c68b685072a29290d5e38d4f9746c

Observation c7028bb1-2669-48a3-b27d-c54cbecb0cb6 · outbound

This paper cites Measuring mathematical problem solving with the math dataset, 2021.

Scaling Up RL: Unlocking Diverse Reasoning in LLMs via Prolonged Training Measuring mathematical problem solving with the math dataset, 2021

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-06T16:51:42.599645Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T16:51:42.599645Z digest=sha256:08df93d84a4900c82b0285faa42cf09e7a25d1aa8f45c959023366366d9020f6

Observation bde776d0-36a9-4c99-a458-dfee812f1963 · outbound

This paper cites Solving quantitative reasoning problems with language models, 2022.

Scaling Up RL: Unlocking Diverse Reasoning in LLMs via Prolonged Training Solving quantitative reasoning problems with language models, 2022

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-06T16:51:42.633994Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T16:51:42.633994Z digest=sha256:e35af49f99b2aa3a8512bf3a0465a3e96023b1d6192053a115c027cf533976a6

Observation 58383f93-4bd2-4994-9394-11a5d15a3604 · outbound

This paper cites Olympiadbench: A challenging benchmark for promoting agi with olympiad-level bilingual multimodal scientific problems, 2024.

Scaling Up RL: Unlocking Diverse Reasoning in LLMs via Prolonged Training Olympiadbench: A challenging benchmark for promoting agi with olympiad-level bilingual multimodal scientific problems, 2024

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-06T16:51:42.657244Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T16:51:42.657244Z digest=sha256:aa3619097a0d4f92b7a173d7664573cf30a72e3b6dfef3fd931353cf918967ae

Observation 43299953-156f-4e44-91f1-fd7833a838ea · outbound

This paper cites Measuring coding challenge competence with apps, 2021.

Scaling Up RL: Unlocking Diverse Reasoning in LLMs via Prolonged Training Measuring coding challenge competence with apps, 2021

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-06T16:51:42.719030Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T16:51:42.719030Z digest=sha256:2f299fd5d65748ffa49d641e6736bfcee97f21b2fb553c5f58da39661abd1e76

Observation 31cdb266-0bbe-4b21-ba0e-157ff5892789 · outbound

This paper cites Mankowitz, Esme Sutherland Robson, Pushmeet Kohli, Nando de Freitas, Koray Kavukcuoglu, and Oriol Vinyals.

Scaling Up RL: Unlocking Diverse Reasoning in LLMs via Prolonged Training Mankowitz, Esme Sutherland Robson, Pushmeet Kohli, Nando de Freitas, Koray Kavukcuoglu, and Oriol Vinyals

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-06T16:51:42.778292Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T16:51:42.778292Z digest=sha256:bb89b8cb63b025002ba8e21df9ac988bb543f12d652522f405d167643a99e290

Observation 12403a07-828b-4577-9316-ef7bc464d2d8 · outbound

This paper cites Taco: Topics in algorithmic code generation dataset, 2023.

Scaling Up RL: Unlocking Diverse Reasoning in LLMs via Prolonged Training Taco: Topics in algorithmic code generation dataset, 2023

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-06T16:51:42.831655Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T16:51:42.831655Z digest=sha256:ad73661fcc633dca69e413ee0e0623bf45cc43e29057350879c3c747a88cb9a2

Observation 98a14543-e3ae-4cb7-bccf-8b9d918c770c · outbound

This paper cites Is your code generated by chatGPT really correct? rigorous evaluation of large language models for code generation.

Scaling Up RL: Unlocking Diverse Reasoning in LLMs via Prolonged Training Is your code generated by chatGPT really correct? rigorous evaluation of large language models for code generation

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T16:51:43.418470Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-06T16:51:42.871217Z digest=sha256:6cee687d2e177a852eb7821b5216027831bec0106d9bcf5e751b461cef58bef8

Observation ae44f13a-84c5-4e58-82bc-276933a6febe · outbound

This paper cites Livecodebench: Holistic and contamination free evaluation of large language models for code, 2024.

Scaling Up RL: Unlocking Diverse Reasoning in LLMs via Prolonged Training Livecodebench: Holistic and contamination free evaluation of large language models for code, 2024

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-06T16:51:42.887053Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T16:51:42.887053Z digest=sha256:e6c0fdec7c3d61fe4edaf09d206f4bc6ca7cbdafd321f96bc864fbb547116e0e

Observation 202e4379-5d88-4739-af38-9804f644e500 · outbound

This paper cites an unresolved cited work.

Scaling Up RL: Unlocking Diverse Reasoning in LLMs via Prolonged Training Unresolved cited work

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-06T16:51:42.908581Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T16:51:42.908581Z digest=sha256:5353c8e6878f95d4a095b24a8b6939ae7f80ac9d7e649f610bc0edba093452c6

Observation e0df78d3-4be0-4360-8fef-94fdf4fc1630 · outbound

This paper cites Online difficulty filtering for reasoning oriented reinforcement learning, 2025.

Scaling Up RL: Unlocking Diverse Reasoning in LLMs via Prolonged Training Online difficulty filtering for reasoning oriented reinforcement learning, 2025

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-06T16:51:42.984504Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T16:51:42.984504Z digest=sha256:31629e9142ee1191d62964f78fa2ec80f25e7482d7de10a8d6a12a5046f18b05

Observation bba4a0ec-2391-4fa3-a991-b13230e25b7a · outbound

This paper cites Gonzalez, Hao Zhang, and Ion Stoica.

Scaling Up RL: Unlocking Diverse Reasoning in LLMs via Prolonged Training Gonzalez, Hao Zhang, and Ion Stoica

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T16:51:43.341473Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-06T16:51:43.036697Z digest=sha256:fca42a2c2b877816cefa9ae4cc7c486320c34f49500252842419a6250df0bb43

Observation 3573b1db-f3f2-4d53-8cdc-f1b92d7d675d · outbound

This paper cites The curious case of neural text degeneration, 2020.

Scaling Up RL: Unlocking Diverse Reasoning in LLMs via Prolonged Training The curious case of neural text degeneration, 2020

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T16:51:43.283087Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-06T16:51:43.099838Z digest=sha256:d2ca637e459af188efcc302b5173c0f869ddf2eb4359b1570b1a4a5c46f6b0d2

Pith citing papers

Observation fa18ef4b-7416-40ea-a7b3-0d7e23549957 · inbound

Revisiting LLM Reasoning via Information Bottleneck cites this paper.

Revisiting LLM Reasoning via Information Bottleneck Scaling Up RL: Unlocking Diverse Reasoning in LLMs via Prolonged Training

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-15T18:21:08.555405Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T18:21:08.555405Z digest=sha256:397d48f5e58a16ed24ffa79cd9d3cc51cb2c7b11efe02fd52b05aa043fabb635

Observation 74b92fd2-80fe-457e-8dfc-f7d363cc0b3b · inbound

Scaling DRL for Decision Making: A Survey on Data, Network, and Training Budget Strategies cites this paper.

Scaling DRL for Decision Making: A Survey on Data, Network, and Training Budget Strategies Scaling Up RL: Unlocking Diverse Reasoning in LLMs via Prolonged Training

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-06T04:39:04.024925Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T04:39:04.024925Z digest=sha256:48b4a90f9acd16f31a2c91e36ee2c3ca590dca7688444bee6cff50965bb121b1

Observation da2f79da-b1fe-4b5c-bc8a-065e3102426e · inbound

Learning to Reason Efficiently with Discounted Reinforcement Learning cites this paper.

Learning to Reason Efficiently with Discounted Reinforcement Learning Scaling Up RL: Unlocking Diverse Reasoning in LLMs via Prolonged Training

Reference 2020

Resolution
unresolved
no resolver link, observed 2026-08-04T07:59:53.046706Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T07:59:53.046706Z digest=sha256:79e810450976c8474aaa66e9beeaa10ee9b8f36176812b2a181c7967654b5861

Observation 3310297e-1a88-4046-94cf-424c1b424a51 · inbound

Cognitive Pivot Points and Visual Anchoring: Unveiling and Rectifying Hallucinations in Multimodal Reasoning Models cites this paper.

Cognitive Pivot Points and Visual Anchoring: Unveiling and Rectifying Hallucinations in Multimodal Reasoning Models Scaling Up RL: Unlocking Diverse Reasoning in LLMs via Prolonged Training

Reference 36

Resolution
verified exact
arxiv_id, observed 2026-05-11T09:25:59.026410Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-05-10T16:03:15.222571Z digest=sha256:e928ea330528a48553b4079b8d83c9fd8ab639dbb8ec960394eabff0ad9cbe4a

Observation ddab7b49-4cd9-47a9-8686-07871448f7bd · inbound

Cognitive Pivot Points and Visual Anchoring: Unveiling and Rectifying Hallucinations in Multimodal Reasoning Models cites this paper.

Cognitive Pivot Points and Visual Anchoring: Unveiling and Rectifying Hallucinations in Multimodal Reasoning Models Scaling Up RL: Unlocking Diverse Reasoning in LLMs via Prolonged Training

Reference 37

Resolution
unresolved
no resolver link, observed 2026-07-12T22:48:45.647588Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T22:48:45.647588Z digest=sha256:756fbd4e118802d62210d05d490ac380a59d09d99b3207515d09b03a6a89b78c

Observation 19e525d1-700f-4d79-b634-ea8bad51050a · inbound

Generalization in LLM Problem Solving: The Case of the Shortest Path cites this paper.

Generalization in LLM Problem Solving: The Case of the Shortest Path Scaling Up RL: Unlocking Diverse Reasoning in LLMs via Prolonged Training

Reference 32

Resolution
verified exact
arxiv_id, observed 2026-05-10T10:39:38.040670Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-05-10T10:37:45.355872Z digest=sha256:fb56f52c8d8b58ca5221475362493cc265ff3fbf4345a540cda5abfd23622fd3

Observation e5e3e00a-54c4-4420-88fc-74d3b10c6268 · inbound

DISA: Offline Importance Sampling for Distribution-Matching LLM-RL cites this paper.

DISA: Offline Importance Sampling for Distribution-Matching LLM-RL Scaling Up RL: Unlocking Diverse Reasoning in LLMs via Prolonged Training

Reference 17

Resolution
verified exact
arxiv_id, observed 2026-05-20T15:13:24.864105Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-05-20T15:11:43.235574Z digest=sha256:fd5bf6629c2936b35864c18f07d282a34ea41df8b3c5de7837caff644650726a

Observation 087dd9df-be03-46a5-a491-625313e27927 · inbound

ARES: Automated Rubric Synthesis for Scalable LLM Reinforcement Learning cites this paper.

ARES: Automated Rubric Synthesis for Scalable LLM Reinforcement Learning Scaling Up RL: Unlocking Diverse Reasoning in LLMs via Prolonged Training

Reference 11

Resolution
verified exact
arxiv_id, observed 2026-05-25T04:35:21.594882Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-05-25T04:31:37.641598Z digest=sha256:2b0728e546bb5cb9721ec13dbc9b568ec14bd40dd214dd9710e330bf90b7870a

Observation 85b190e6-9899-4bb1-95db-2d694c32d13a · inbound

ARES: Automated Rubric Synthesis for Scalable LLM Reinforcement Learning cites this paper.

ARES: Automated Rubric Synthesis for Scalable LLM Reinforcement Learning Scaling Up RL: Unlocking Diverse Reasoning in LLMs via Prolonged Training

Reference 11

Resolution
verified exact
arxiv_id, observed 2026-06-30T16:24:55.362639Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-06-30T16:21:40.582248Z digest=sha256:3e1ed6479ce948aeda6c70fc4758f67cc323b84fe2344d9a146db35b630e27c6

Observation 373a9187-b9f6-4da7-8a24-f1581a8da92d · inbound

Trust Region On-Policy Distillation cites this paper.

Trust Region On-Policy Distillation Scaling Up RL: Unlocking Diverse Reasoning in LLMs via Prolonged Training

Reference 226

Resolution
verified exact
arxiv_id, observed 2026-07-01T20:56:13.556324Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-06-28T17:38:50.313305Z digest=sha256:e744d000c125c86cec9a491c80b6a9d06bf7455a539d785a11f85c15a6d799fd