Pith. sign in

Paper Citation Record · LEDGER

Stable Reinforcement Learning for Efficient Reasoning

As of 18 August 2026, this Paper Citation Record lists 35 of 35 outbound references and 7 inbound Pith citation observations for arXiv:2505.18086.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2505.18086 v1

Coverage vector

measured 35 of 35 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-07T14:38:47.682695Z

measured 42 of 42 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-17T06:30:58.91139+00:00

measured 7 of 7 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-06T12:09:39.636062Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-01T22:46:20.251611Z

Reference resolution

35 of 35 outbound references displayed

  • verified exact0
  • verified fuzzy3
  • unresolved32
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation c9a7d3d8-c24b-4338-a122-6e5ccd4fa911 · outbound

This paper cites Scaling LLM Test-Time Compute Optimally can be More Effective than Scaling Model Parameters.

Stable Reinforcement Learning for Efficient Reasoning Scaling LLM Test-Time Compute Optimally can be More Effective than Scaling Model Parameters

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-07T14:38:42.693276Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:38:42.693276Z digest=sha256:7a1b0ea43b8807b1e948ef3b12298f8ee3b35fe965302c5713b256390d0e345a

Observation cbe93923-50d0-41cd-9438-741854849ba4 · outbound

This paper cites Scaling Laws for Neural Language Models.

Stable Reinforcement Learning for Efficient Reasoning Scaling Laws for Neural Language Models

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-07T14:38:42.839047Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:38:42.839047Z digest=sha256:3ea341e6003ab0ad29f1a0a80e27b22c18e1280739e45522ce1e3d33a41ab2ce

Observation 3ca8bd60-b550-43a9-b38e-1fa1042351e7 · outbound

This paper cites DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning.

Stable Reinforcement Learning for Efficient Reasoning DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-07T14:38:42.993274Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:38:42.993274Z digest=sha256:3e1089cb65d4c9cc9060b964371be21b7a4007919fc60195868d5e97118435c4

Observation ecb8034f-13b8-4298-b5b8-19ee820e5f47 · outbound

This paper cites Qwen3 Technical Report.

Stable Reinforcement Learning for Efficient Reasoning Qwen3 Technical Report

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-07T14:38:43.173424Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:38:43.173424Z digest=sha256:60a1bfae4ee925d9ec939c8e22c0309065832336e6403c6530e75a93216329ca

Observation ad775c01-2cc7-44cc-8f75-7f57e29df208 · outbound

This paper cites DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models.

Stable Reinforcement Learning for Efficient Reasoning DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-07T14:38:43.302144Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:38:43.302144Z digest=sha256:5564c62add72c00e7dbe8c3c5d782ca0f56b8f10b8d9962170c8c3ad13ebe22e

Observation 513424c0-8561-4ecf-946a-cf6e63e53b55 · outbound

This paper cites Training a Helpful and Harmless Assistant with Reinforcement Learning from Human Feedback.

Stable Reinforcement Learning for Efficient Reasoning Training a Helpful and Harmless Assistant with Reinforcement Learning from Human Feedback

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-07T14:38:43.477652Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:38:43.477652Z digest=sha256:f50c92c3ed7e94cae5ca8704c29f6053d865755203703a851f341b8bb2583e31

Observation a1f10c50-543c-4187-9aba-4dc66166503e · outbound

This paper cites Training language models to follow instructions with human feedback.

Stable Reinforcement Learning for Efficient Reasoning Training language models to follow instructions with human feedback

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-07T14:38:43.607566Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:38:43.607566Z digest=sha256:ec7fcd7c492a9a7564d0b670160869c02fea798871db3e93f21509901ec61cc2

Observation 5d224063-e1c8-4345-b429-4bca1e0e5344 · outbound

This paper cites Proximal Policy Optimization Algorithms.

Stable Reinforcement Learning for Efficient Reasoning Proximal Policy Optimization Algorithms

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-07T14:38:43.731606Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:38:43.731606Z digest=sha256:99f125f52f509b3f21ac0b8e38f336d8b907e27029fbe8a67fa6d6971f5542b6

Observation 945cf57e-8403-4d7f-b180-3496e2dbec1a · outbound

This paper cites Group robust preference optimization in reward-free RLHF.

Stable Reinforcement Learning for Efficient Reasoning Group robust preference optimization in reward-free RLHF

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:38:49.661260Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-07T14:38:43.873464Z digest=sha256:00fcc85345305740094e60c27251fa81b5391d7a1ba0d999849528fe39bcad37

Observation 2de9ee18-9ab1-4ab1-a720-4f5ab0da98dd · outbound

This paper cites Towards Large Reasoning Models: A Survey of Reinforced Reasoning with Large Language Models.

Stable Reinforcement Learning for Efficient Reasoning Towards Large Reasoning Models: A Survey of Reinforced Reasoning with Large Language Models

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-07T14:38:44.107498Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:38:44.107498Z digest=sha256:2ee7aaa06b999c0c4fcf851bbd04c93d9772b29ec6a295fbac7ffa93fa190558

Observation 09bed22b-4186-4cc2-9e0c-4b544f4540bf · outbound

This paper cites Chain-of-Thought Prompting Elicits Reasoning in Large Language Models.

Stable Reinforcement Learning for Efficient Reasoning Chain-of-Thought Prompting Elicits Reasoning in Large Language Models

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-07T14:38:44.256624Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:38:44.256624Z digest=sha256:ef7359400509cdfd35fd3f43482a95640917187f8bcbc8654fdcf5f8a15f3ae7

Observation a05546bc-dc99-4b26-8b01-9c0784e78b85 · outbound

This paper cites Do NOT Think That Much for 2+3=? On the Overthinking of o1-Like LLMs.

Stable Reinforcement Learning for Efficient Reasoning Do NOT Think That Much for 2+3=? On the Overthinking of o1-Like LLMs

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-07T14:38:44.377434Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:38:44.377434Z digest=sha256:d7db42567f4f408c43a08d7d79aee2da0b4b1d5d0b8497ec4207f208d3eb80d8

Observation b91f5d9f-a288-40d2-8523-5518f8dd2466 · outbound

This paper cites The Danger of Overthinking: Examining the Reasoning-Action Dilemma in Agentic Tasks.

Stable Reinforcement Learning for Efficient Reasoning The Danger of Overthinking: Examining the Reasoning-Action Dilemma in Agentic Tasks

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-07T14:38:44.636364Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:38:44.636364Z digest=sha256:df5eff5f4bed8d6432020ad1a00162bec0f556efe6cb3d06b7531209375c2d14

Observation 1c8b41cd-4bc4-4d55-a372-8f305df9f6cb · outbound

This paper cites When More is Less: Understanding Chain-of-Thought Length in LLMs.

Stable Reinforcement Learning for Efficient Reasoning When More is Less: Understanding Chain-of-Thought Length in LLMs

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-07T14:38:44.762699Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:38:44.762699Z digest=sha256:19d5fef034e155f4ff7ec801b0d837995fd27d1054acc9842e4f9ef65ab10387

Observation 88487cad-08c9-4ba8-96d8-177b0f0632c0 · outbound

This paper cites Dynamic early exit in reasoning models, 2025.

Stable Reinforcement Learning for Efficient Reasoning Dynamic early exit in reasoning models, 2025

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-07T14:38:44.901523Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:38:44.901523Z digest=sha256:03ea8b31fa3ccfad6ebcbad3547bdb563557f3016c3898096e24ad0670a4181f

Observation e6ebe74d-386b-4f74-9ac0-e03020456fba · outbound

This paper cites Kimi k1.5: Scaling Reinforcement Learning with LLMs.

Stable Reinforcement Learning for Efficient Reasoning Kimi k1.5: Scaling Reinforcement Learning with LLMs

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-07T14:38:45.030627Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:38:45.030627Z digest=sha256:c0b0088598047f284b50dc1db348f96f8476925fc28bf792a2d62a98194fb399

Observation aaaf58df-ce8c-448f-a445-8adae971953e · outbound

This paper cites S-GRPO: Early Exit via Reinforcement Learning in Reasoning Models.

Stable Reinforcement Learning for Efficient Reasoning S-GRPO: Early Exit via Reinforcement Learning in Reasoning Models

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-07T14:38:45.299309Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:38:45.299309Z digest=sha256:ed7c37eb12e74614a4377a485d76863529c23a3f9653d09cd9852a3bf89fc61d

Observation 2ab6bda3-e684-4120-8286-37f5ceb9ed7b · outbound

This paper cites Training Verifiers to Solve Math Word Problems.

Stable Reinforcement Learning for Efficient Reasoning Training Verifiers to Solve Math Word Problems

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-07T14:38:45.504178Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:38:45.504178Z digest=sha256:466022a9266dedf83b98082412a01d1dc0f0fa8f10597e5653ae4ea918e7fa4d

Observation 43b2069d-6689-4980-978e-fa6673d4508a · outbound

This paper cites GPQA: A Graduate-Level Google-Proof Q&A Benchmark.

Stable Reinforcement Learning for Efficient Reasoning GPQA: A Graduate-Level Google-Proof Q&A Benchmark

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-07T14:38:45.622242Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:38:45.622242Z digest=sha256:e023faf843e7f6fbe5ec3ff74d732f43ed1cf5c875a8e21a8d7e03088ca20deb

Observation f63de3ea-d04c-4cd4-ac5d-e3a980e3cb62 · outbound

This paper cites Aime problems and solutions.

Stable Reinforcement Learning for Efficient Reasoning Aime problems and solutions

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:38:49.078178Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-07T14:38:45.747265Z digest=sha256:5bf1a16af2bd29fd8ed77e7228dcf821e01191ac7756745d00551a939470ffad

Observation f898c5e0-3baa-4776-a942-9b7e0c2f03a1 · outbound

This paper cites Amc 2023, 2024.

Stable Reinforcement Learning for Efficient Reasoning Amc 2023, 2024

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-07T14:38:45.879870Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:38:45.879870Z digest=sha256:4ace642b57b79757bcacd23465d3dcebe855c4d14008bf446738beed8075ace0

Observation af2ae48e-fb79-4463-9aec-7e25a8fbac07 · outbound

This paper cites Measuring mathematical problem solving with the math dataset,.

Stable Reinforcement Learning for Efficient Reasoning Measuring mathematical problem solving with the math dataset,

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-07T14:38:46.047459Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:38:46.047459Z digest=sha256:ba3f196a26042b6ca27b2a1d519e3a45de35224191a63b5c00a8362acd10d259

Observation 37b165fe-3c43-458e-999e-5afc1c52a00f · outbound

This paper cites Learning to reason with llms.

Stable Reinforcement Learning for Efficient Reasoning Learning to reason with llms

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:38:48.564499Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-07T14:38:46.312282Z digest=sha256:990c3a55f7f4c28f4802dcd341c2810b1ae21b6e6ba6bf38474840797ed7b8f4

Observation 71d0c005-ec14-442e-885a-eded273fb183 · outbound

This paper cites Training language models to follow instructions with human feedback.

Stable Reinforcement Learning for Efficient Reasoning Training language models to follow instructions with human feedback

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-07T14:38:46.432494Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:38:46.432494Z digest=sha256:5626eb2c00ae64a388421cf25c464041320b8e503a50a02bf71b3ec52f774446

Observation 0da1767a-63d3-4bbe-acac-d4186ae8de5c · outbound

This paper cites On Designing Effective RL Reward at Training Time for LLM Reasoning.

Stable Reinforcement Learning for Efficient Reasoning On Designing Effective RL Reward at Training Time for LLM Reasoning

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-07T14:38:46.529177Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:38:46.529177Z digest=sha256:442befddb590959e2b3c83243e932ec0d120db99e697588fa9a4fbc657133570

Observation 143a71ca-8286-4c91-9cea-81e9229fdd6f · outbound

This paper cites Tulu 3: Pushing Frontiers in Open Language Model Post-Training.

Stable Reinforcement Learning for Efficient Reasoning Tulu 3: Pushing Frontiers in Open Language Model Post-Training

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-07T14:38:46.709657Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:38:46.709657Z digest=sha256:8207438ca4e57e19859dbdb3840348d7cb3d89c2d345c6fd2d0c8bb15b8c3076

Observation ff791d8f-0479-4963-a508-1360be495822 · outbound

This paper cites SimpleRL-Zoo: Investigating and Taming Zero Reinforcement Learning for Open Base Models in the Wild.

Stable Reinforcement Learning for Efficient Reasoning SimpleRL-Zoo: Investigating and Taming Zero Reinforcement Learning for Open Base Models in the Wild

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-07T14:38:46.831392Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:38:46.831392Z digest=sha256:c783ff42c5849083f1f0172dc426140d33533a4615368b21ed9415ab439f1904

Observation 8f6d38db-9d42-4123-b1f9-235377011f6b · outbound

This paper cites Light-R1: Curriculum SFT, DPO and RL for Long COT from Scratch and Beyond.

Stable Reinforcement Learning for Efficient Reasoning Light-R1: Curriculum SFT, DPO and RL for Long COT from Scratch and Beyond

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-07T14:38:46.973811Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:38:46.973811Z digest=sha256:89752acc1876d981ee99cde64f4a7d1e4954adb23786e7e3d7f7fec11d171a60

Observation 13461ca3-1766-4edb-a8b9-c4016710b253 · outbound

This paper cites Fastcurl: Curriculum reinforcement learning with progressive context extension for efficient training r1-like reasoning models, 2025.

Stable Reinforcement Learning for Efficient Reasoning Fastcurl: Curriculum reinforcement learning with progressive context extension for efficient training r1-like reasoning models, 2025

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-07T14:38:47.135822Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:38:47.135822Z digest=sha256:eb8992935ecac9a81450d4c499b3a80adf14ad28c99048f22d93ea8f137dd9f7

Observation ca3c6a0b-18fe-4b0e-a2cf-36a820f24d15 · outbound

This paper cites Training language models to reason efficiently.

Stable Reinforcement Learning for Efficient Reasoning Training language models to reason efficiently

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-07T14:38:47.237133Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:38:47.237133Z digest=sha256:6038fe30638790783489db0676d7ad31387e68c95acb48307cbb746b91090e86

Observation 616be23c-0f78-40a6-8881-8a637368c55a · outbound

This paper cites DeepMath-103K: A Large-Scale, Challenging, Decontaminated, and Verifiable Mathematical Dataset for Advancing Reasoning.

Stable Reinforcement Learning for Efficient Reasoning DeepMath-103K: A Large-Scale, Challenging, Decontaminated, and Verifiable Mathematical Dataset for Advancing Reasoning

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-07T14:38:47.403222Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:38:47.403222Z digest=sha256:65ed2794b38747c7d8e3549590cd84463cf9d7c9e5af817196568d6e0d5f1c0d

Observation 445767f5-5ddd-4cc7-9d3a-4bb0c17866a8 · outbound

This paper cites Adam: A Method for Stochastic Optimization.

Stable Reinforcement Learning for Efficient Reasoning Adam: A Method for Stochastic Optimization

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-07T14:38:47.561040Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:38:47.561040Z digest=sha256:44faa4ae80f8f53ae53ce4c5b8dace05095921fd338a33b49debe06976191b7f

Observation 456d2ebb-02e8-4cfb-b296-56ce53d3e3d9 · outbound

This paper cites ThinkPrune: Pruning Long Chain-of-Thought of LLMs via Reinforcement Learning.

Stable Reinforcement Learning for Efficient Reasoning ThinkPrune: Pruning Long Chain-of-Thought of LLMs via Reinforcement Learning

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-07T14:38:47.682695Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:38:47.682695Z digest=sha256:0ca62adaec197c5f64c3a17f29dac6ff43bbc3123e97bd29339a88e3ce758564

Observation ac5f1b6b-6f50-4cac-89de-0faf5272030e · outbound

This paper cites Measuring Mathematical Problem Solving With the MATH Dataset.

Stable Reinforcement Learning for Efficient Reasoning Measuring Mathematical Problem Solving With the MATH Dataset

Reference 2021

Resolution
unresolved
no resolver link, observed 2026-08-07T14:38:46.155558Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:38:46.155558Z digest=sha256:ee8b81abe2bbcbd72b0b6a0952cc72f3b12e5bc884391b8b06357cab9f1c5bbf

Observation df1278cb-0481-40ff-81f4-9e0813881153 · outbound

This paper cites an unresolved cited work.

Stable Reinforcement Learning for Efficient Reasoning Unresolved cited work

Reference 2024

Resolution
unresolved
raw_fallback, observed 2026-08-07T14:38:49.453422Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-07T14:38:43.956922Z digest=sha256:9e39b461ca6c82a3f98dc87e5fd74a460b0542b8ee31439f89714e86ebb6d908

Pith citing papers

Observation d005cb89-ec4f-4219-824f-e9cb83aa415a · inbound

Secure Tug-of-War (SecTOW): Iterative Defense-Attack Training with Reinforcement Learning for Multimodal Model Security cites this paper.

Secure Tug-of-War (SecTOW): Iterative Defense-Attack Training with Reinforcement Learning for Multimodal Model Security Stable Reinforcement Learning for Efficient Reasoning

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-06T12:09:39.636062Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T12:09:39.636062Z digest=sha256:744e30c6d1cb9391f0a11796d3e56775d3d5135ab10c96ca7b62b1f86adea6a9

Observation 83c4f67c-abf3-4632-be6f-9ddc3cb4bbdb · inbound

Gradient Extrapolation-Based Policy Optimization cites this paper.

Gradient Extrapolation-Based Policy Optimization Stable Reinforcement Learning for Efficient Reasoning

Reference 5

Resolution
verified exact
arxiv_id, observed 2026-05-11T03:55:56.726774Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-05-11T02:07:08.792030Z digest=sha256:c3c7dc5ad17e5ce5d7658415cee49b6f59e819e4718e030a3d947eb77225e22c

Observation 32eafdbe-ed52-49a4-b269-446d843604d6 · inbound

CLORE: Content-Level Optimization for Reasoning Efficiency cites this paper.

CLORE: Content-Level Optimization for Reasoning Efficiency Stable Reinforcement Learning for Efficient Reasoning

Reference 8

Resolution
verified exact
arxiv_id, observed 2026-05-22T05:51:08.263012Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-05-22T05:50:23.111591Z digest=sha256:7593848c62eaa67f95517a97d7ed55c72f979ec85f6c8be482c70ac22e7e50b3

Observation 6967ca88-cf8c-4aac-a953-bb8e16f81ff5 · inbound

Learning When Not to Act: Mitigating Tool Abuse in Agentic Reinforcement Learning cites this paper.

Learning When Not to Act: Mitigating Tool Abuse in Agentic Reinforcement Learning Stable Reinforcement Learning for Efficient Reasoning

Reference 28

Resolution
metadata mismatch
arxiv_id, observed 2026-07-01T22:46:20.253672Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-06-28T14:59:57.983338Z digest=sha256:f1a7a436127f7423655de53b963dd0d5611005d7e3507054e96566ecc8109114

Observation 2163f506-be2c-4659-b578-8de3e92f0899 · inbound

Length Penalties Make Chain-of-Thought Less Monitorable cites this paper.

Length Penalties Make Chain-of-Thought Less Monitorable Stable Reinforcement Learning for Efficient Reasoning

Reference 54

Resolution
unresolved
no resolver link, observed 2026-07-14T15:45:54.532529Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-14T15:45:54.532529Z digest=sha256:93301e47346163fd2900a14970414b249891ac729d06d1f9fb546aa958b68aac

Observation af84df61-620c-4d7d-b0a9-b19de84984bb · inbound

Length Penalties Make Chain-of-Thought Less Monitorable cites this paper.

Length Penalties Make Chain-of-Thought Less Monitorable Stable Reinforcement Learning for Efficient Reasoning

Reference 55

Resolution
unresolved
no resolver link, observed 2026-08-02T08:06:10.779080Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-02T08:06:10.779080Z digest=sha256:eb21e32e6a38e8083c7e87e554dae8de9d104d3e729ca8fb895aac8c95871df0

Observation ef30e1dc-c058-4e94-95c1-46f0948119e0 · inbound

Length Penalties Make Chain-of-Thought Less Monitorable cites this paper.

Length Penalties Make Chain-of-Thought Less Monitorable Stable Reinforcement Learning for Efficient Reasoning

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-04T04:30:27.449308Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T04:30:27.449308Z digest=sha256:f35ae5a8535de408a03a94c40fcee21927dd75778881ae07a193b7a9ed9df90e