Pith. sign in

Paper Citation Record · LEDGER

Stable Reinforcement Learning for Efficient Reasoning

As of 8 August 2026, this Paper Citation Record lists 35 of 35 outbound references and 7 inbound Pith citation observations for arXiv:2505.18086.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2505.18086 v1

Coverage vector

measured 35 of 35 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-07T14:38:47.682695Z

measured 42 of 42 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-08T06:32:00.761636+00:00

measured 7 of 7 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-06T12:09:39.636062Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-01T22:46:20.251611Z

Reference resolution

35 of 35 outbound references displayed

  • verified exact0
  • verified fuzzy3
  • unresolved32
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation c9a7d3d8-c24b-4338-a122-6e5ccd4fa911 · outbound

This paper cites Scaling LLM Test-Time Compute Optimally can be More Effective than Scaling Model Parameters.

Stable Reinforcement Learning for Efficient Reasoning Scaling LLM Test-Time Compute Optimally can be More Effective than Scaling Model Parameters

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-07T14:38:42.693276Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:38:42.693276Z digest=sha256:108eb597928cf8e30b211dce9f3f8bb4e6e859bfc464e5c79a4ebf76565400e8

Observation cbe93923-50d0-41cd-9438-741854849ba4 · outbound

This paper cites Scaling Laws for Neural Language Models.

Stable Reinforcement Learning for Efficient Reasoning Scaling Laws for Neural Language Models

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-07T14:38:42.839047Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:38:42.839047Z digest=sha256:72821120a055704861ce3adee1b93f9aa67a5c35c60b36d47d4a6d50feccb372

Observation 3ca8bd60-b550-43a9-b38e-1fa1042351e7 · outbound

This paper cites DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning.

Stable Reinforcement Learning for Efficient Reasoning DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-07T14:38:42.993274Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:38:42.993274Z digest=sha256:ee857f26971e34660a3492909193bf0dc47c2f288478108edec6229be04ec1a5

Observation ecb8034f-13b8-4298-b5b8-19ee820e5f47 · outbound

This paper cites Qwen3 Technical Report.

Stable Reinforcement Learning for Efficient Reasoning Qwen3 Technical Report

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-07T14:38:43.173424Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:38:43.173424Z digest=sha256:898b4604e5c23d94a545fe359bc2c78ea84a3fc1b3697492c56c5854eefaaa74

Observation ad775c01-2cc7-44cc-8f75-7f57e29df208 · outbound

This paper cites DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models.

Stable Reinforcement Learning for Efficient Reasoning DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-07T14:38:43.302144Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:38:43.302144Z digest=sha256:2570114833257ccf2d131cd93fe36ead37bbabd50edf9e06abb7c3916335b355

Observation 513424c0-8561-4ecf-946a-cf6e63e53b55 · outbound

This paper cites Training a Helpful and Harmless Assistant with Reinforcement Learning from Human Feedback.

Stable Reinforcement Learning for Efficient Reasoning Training a Helpful and Harmless Assistant with Reinforcement Learning from Human Feedback

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-07T14:38:43.477652Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:38:43.477652Z digest=sha256:27908a72f8913cd245edec48ab05a51449def0ad636cef4608dc3f972a1753b8

Observation a1f10c50-543c-4187-9aba-4dc66166503e · outbound

This paper cites Training language models to follow instructions with human feedback.

Stable Reinforcement Learning for Efficient Reasoning Training language models to follow instructions with human feedback

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-07T14:38:43.607566Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:38:43.607566Z digest=sha256:fe25f77485edd61d41cd6666877315452922a30ffc40397266b1d4b83c6db42f

Observation 5d224063-e1c8-4345-b429-4bca1e0e5344 · outbound

This paper cites Proximal Policy Optimization Algorithms.

Stable Reinforcement Learning for Efficient Reasoning Proximal Policy Optimization Algorithms

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-07T14:38:43.731606Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:38:43.731606Z digest=sha256:b3906ca2a4eb3f2ce43f6e77a3a700def716ba6106d6401fed8df469469726eb

Observation 945cf57e-8403-4d7f-b180-3496e2dbec1a · outbound

This paper cites Group robust preference optimization in reward-free RLHF.

Stable Reinforcement Learning for Efficient Reasoning Group robust preference optimization in reward-free RLHF

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:38:49.661260Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T14:38:43.873464Z digest=sha256:9becdbfc5c8871b0d3c6d29191512ce13e63c65703db738ce59db945b32a81a3

Observation 2de9ee18-9ab1-4ab1-a720-4f5ab0da98dd · outbound

This paper cites Towards Large Reasoning Models: A Survey of Reinforced Reasoning with Large Language Models.

Stable Reinforcement Learning for Efficient Reasoning Towards Large Reasoning Models: A Survey of Reinforced Reasoning with Large Language Models

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-07T14:38:44.107498Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:38:44.107498Z digest=sha256:9bf62e900c2e77bb3a2857380e2f6f6d2cbfbd6f8a72ff3ee8a6e309f2b8695d

Observation 09bed22b-4186-4cc2-9e0c-4b544f4540bf · outbound

This paper cites Chain-of-Thought Prompting Elicits Reasoning in Large Language Models.

Stable Reinforcement Learning for Efficient Reasoning Chain-of-Thought Prompting Elicits Reasoning in Large Language Models

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-07T14:38:44.256624Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:38:44.256624Z digest=sha256:bd4e804bdece57e2399ecd1c1b31a629500c852d1ed0fe40c6cb364090f0b671

Observation a05546bc-dc99-4b26-8b01-9c0784e78b85 · outbound

This paper cites Do NOT Think That Much for 2+3=? On the Overthinking of o1-Like LLMs.

Stable Reinforcement Learning for Efficient Reasoning Do NOT Think That Much for 2+3=? On the Overthinking of o1-Like LLMs

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-07T14:38:44.377434Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:38:44.377434Z digest=sha256:8ff86cefeea9ec87c2385ae240baaa347385835e521feab5f15eb8fb6ee3e8d5

Observation b91f5d9f-a288-40d2-8523-5518f8dd2466 · outbound

This paper cites The Danger of Overthinking: Examining the Reasoning-Action Dilemma in Agentic Tasks.

Stable Reinforcement Learning for Efficient Reasoning The Danger of Overthinking: Examining the Reasoning-Action Dilemma in Agentic Tasks

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-07T14:38:44.636364Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:38:44.636364Z digest=sha256:3077a1824cd471a0b26f387d76b277decc03b65d4eaa8d447d48df6923940002

Observation 1c8b41cd-4bc4-4d55-a372-8f305df9f6cb · outbound

This paper cites When More is Less: Understanding Chain-of-Thought Length in LLMs.

Stable Reinforcement Learning for Efficient Reasoning When More is Less: Understanding Chain-of-Thought Length in LLMs

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-07T14:38:44.762699Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:38:44.762699Z digest=sha256:4009ffa5e1e009a622e1cec290b998bb2f10bd19b83e09c9c358b3f212b0802b

Observation 88487cad-08c9-4ba8-96d8-177b0f0632c0 · outbound

This paper cites Dynamic early exit in reasoning models, 2025.

Stable Reinforcement Learning for Efficient Reasoning Dynamic early exit in reasoning models, 2025

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-07T14:38:44.901523Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:38:44.901523Z digest=sha256:4208be767590cfd6efc42feda3e7eb0d2167bcf4852a4e674ddc0d9785fef6fd

Observation e6ebe74d-386b-4f74-9ac0-e03020456fba · outbound

This paper cites Kimi k1.5: Scaling Reinforcement Learning with LLMs.

Stable Reinforcement Learning for Efficient Reasoning Kimi k1.5: Scaling Reinforcement Learning with LLMs

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-07T14:38:45.030627Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:38:45.030627Z digest=sha256:d735acedcfd99016047a25c698052e48bf29c7ab3bbb9862c66b5144795e8783

Observation aaaf58df-ce8c-448f-a445-8adae971953e · outbound

This paper cites S-GRPO: Early Exit via Reinforcement Learning in Reasoning Models.

Stable Reinforcement Learning for Efficient Reasoning S-GRPO: Early Exit via Reinforcement Learning in Reasoning Models

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-07T14:38:45.299309Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:38:45.299309Z digest=sha256:e3bef73bfdf2814374358bb138919d0ec20754cc1f7671092c5900422cc99555

Observation 2ab6bda3-e684-4120-8286-37f5ceb9ed7b · outbound

This paper cites Training Verifiers to Solve Math Word Problems.

Stable Reinforcement Learning for Efficient Reasoning Training Verifiers to Solve Math Word Problems

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-07T14:38:45.504178Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:38:45.504178Z digest=sha256:02c5bbb240c530604af8bf7a83c7668e98b1fd6e8806f21957fe160e301364f6

Observation 43b2069d-6689-4980-978e-fa6673d4508a · outbound

This paper cites GPQA: A Graduate-Level Google-Proof Q&A Benchmark.

Stable Reinforcement Learning for Efficient Reasoning GPQA: A Graduate-Level Google-Proof Q&A Benchmark

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-07T14:38:45.622242Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:38:45.622242Z digest=sha256:ee78e05fec92097a4e9703961ad4c4c1eea8e367db3fb8e417ba6aa56ed6eee3

Observation f63de3ea-d04c-4cd4-ac5d-e3a980e3cb62 · outbound

This paper cites Aime problems and solutions.

Stable Reinforcement Learning for Efficient Reasoning Aime problems and solutions

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:38:49.078178Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T14:38:45.747265Z digest=sha256:3f4b9d0702f13cff2fa199de7d1412d16423fef4f30f81bdfbda4b55bfa1d9a3

Observation f898c5e0-3baa-4776-a942-9b7e0c2f03a1 · outbound

This paper cites Amc 2023, 2024.

Stable Reinforcement Learning for Efficient Reasoning Amc 2023, 2024

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-07T14:38:45.879870Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:38:45.879870Z digest=sha256:bc3607a9b3a8f842c82066898322d7c8e42c71b2807bc74ae6b764dcb4ded541

Observation af2ae48e-fb79-4463-9aec-7e25a8fbac07 · outbound

This paper cites Measuring mathematical problem solving with the math dataset,.

Stable Reinforcement Learning for Efficient Reasoning Measuring mathematical problem solving with the math dataset,

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-07T14:38:46.047459Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:38:46.047459Z digest=sha256:6c3b95dcf5e88a97ab060babdc277554c5a22d418a7179dbee1256f66480d0cd

Observation 37b165fe-3c43-458e-999e-5afc1c52a00f · outbound

This paper cites Learning to reason with llms.

Stable Reinforcement Learning for Efficient Reasoning Learning to reason with llms

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:38:48.564499Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T14:38:46.312282Z digest=sha256:6eb0637b3103d81790a1fe24274bfc9ca84bb71312b9411ee35cc7fdfb9982bc

Observation 71d0c005-ec14-442e-885a-eded273fb183 · outbound

This paper cites Training language models to follow instructions with human feedback.

Stable Reinforcement Learning for Efficient Reasoning Training language models to follow instructions with human feedback

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-07T14:38:46.432494Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:38:46.432494Z digest=sha256:79f7f18b1b3b45c0307bbb67429bf226171aeebbffa7170edcab350f7db4a118

Observation 0da1767a-63d3-4bbe-acac-d4186ae8de5c · outbound

This paper cites On Designing Effective RL Reward at Training Time for LLM Reasoning.

Stable Reinforcement Learning for Efficient Reasoning On Designing Effective RL Reward at Training Time for LLM Reasoning

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-07T14:38:46.529177Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:38:46.529177Z digest=sha256:9a5d9dc942fa065ed991cb596f1f2219064967ac6cbba398da483ad3068371ae

Observation 143a71ca-8286-4c91-9cea-81e9229fdd6f · outbound

This paper cites Tulu 3: Pushing Frontiers in Open Language Model Post-Training.

Stable Reinforcement Learning for Efficient Reasoning Tulu 3: Pushing Frontiers in Open Language Model Post-Training

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-07T14:38:46.709657Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:38:46.709657Z digest=sha256:4731256f4a0ca5f9fb20e52959e7b09dd3e16595d2d98b240fc0427850744dab

Observation ff791d8f-0479-4963-a508-1360be495822 · outbound

This paper cites SimpleRL-Zoo: Investigating and Taming Zero Reinforcement Learning for Open Base Models in the Wild.

Stable Reinforcement Learning for Efficient Reasoning SimpleRL-Zoo: Investigating and Taming Zero Reinforcement Learning for Open Base Models in the Wild

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-07T14:38:46.831392Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:38:46.831392Z digest=sha256:b41bfeeebd48c86fbc6d9679be8b1c302bb4303adc783e19f742afc4a86b1201

Observation 8f6d38db-9d42-4123-b1f9-235377011f6b · outbound

This paper cites Light-R1: Curriculum SFT, DPO and RL for Long COT from Scratch and Beyond.

Stable Reinforcement Learning for Efficient Reasoning Light-R1: Curriculum SFT, DPO and RL for Long COT from Scratch and Beyond

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-07T14:38:46.973811Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:38:46.973811Z digest=sha256:b07b923da23502c8b30d3cae0c0a03b00a2c07072e9a1854c3d82758895cec67

Observation 13461ca3-1766-4edb-a8b9-c4016710b253 · outbound

This paper cites Fastcurl: Curriculum reinforcement learning with progressive context extension for efficient training r1-like reasoning models, 2025.

Stable Reinforcement Learning for Efficient Reasoning Fastcurl: Curriculum reinforcement learning with progressive context extension for efficient training r1-like reasoning models, 2025

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-07T14:38:47.135822Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:38:47.135822Z digest=sha256:dfb12c67803ccbe747199a22f19d060d9634250d618651eb214320488bc2b7c3

Observation ca3c6a0b-18fe-4b0e-a2cf-36a820f24d15 · outbound

This paper cites Training language models to reason efficiently.

Stable Reinforcement Learning for Efficient Reasoning Training language models to reason efficiently

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-07T14:38:47.237133Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:38:47.237133Z digest=sha256:d4c832d5c7c27147129a11fa6ca17850b291c50e7149166d297b16eff4c3aeba

Observation 616be23c-0f78-40a6-8881-8a637368c55a · outbound

This paper cites DeepMath-103K: A Large-Scale, Challenging, Decontaminated, and Verifiable Mathematical Dataset for Advancing Reasoning.

Stable Reinforcement Learning for Efficient Reasoning DeepMath-103K: A Large-Scale, Challenging, Decontaminated, and Verifiable Mathematical Dataset for Advancing Reasoning

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-07T14:38:47.403222Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:38:47.403222Z digest=sha256:82a820065c0b70e6d590db2679267665d18a902bd57bee2414aeb6cb6ea54082

Observation 445767f5-5ddd-4cc7-9d3a-4bb0c17866a8 · outbound

This paper cites Adam: A Method for Stochastic Optimization.

Stable Reinforcement Learning for Efficient Reasoning Adam: A Method for Stochastic Optimization

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-07T14:38:47.561040Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:38:47.561040Z digest=sha256:4f99dfd61b45e6eb9274e4162502ea3c7b523cafac39fd630aaa5f57db5e6593

Observation 456d2ebb-02e8-4cfb-b296-56ce53d3e3d9 · outbound

This paper cites ThinkPrune: Pruning Long Chain-of-Thought of LLMs via Reinforcement Learning.

Stable Reinforcement Learning for Efficient Reasoning ThinkPrune: Pruning Long Chain-of-Thought of LLMs via Reinforcement Learning

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-07T14:38:47.682695Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:38:47.682695Z digest=sha256:1d040253797894e193a7a4d055305737d0aec900fefa4ed8d5d343011c632c0d

Observation ac5f1b6b-6f50-4cac-89de-0faf5272030e · outbound

This paper cites Measuring Mathematical Problem Solving With the MATH Dataset.

Stable Reinforcement Learning for Efficient Reasoning Measuring Mathematical Problem Solving With the MATH Dataset

Reference 2021

Resolution
unresolved
no resolver link, observed 2026-08-07T14:38:46.155558Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:38:46.155558Z digest=sha256:fe8d970075d1e4626e02bd2df9e79c8f39616a4a4b3e74bc5724f0eb2a181b01

Observation df1278cb-0481-40ff-81f4-9e0813881153 · outbound

This paper cites an unresolved cited work.

Stable Reinforcement Learning for Efficient Reasoning Unresolved cited work

Reference 2024

Resolution
unresolved
raw_fallback, observed 2026-08-07T14:38:49.453422Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T14:38:43.956922Z digest=sha256:7e6bc6ebf4ec8fd1b9d71085000e9242f1913eab1d71d53d2aac8ec5d6648251

Pith citing papers

Observation d005cb89-ec4f-4219-824f-e9cb83aa415a · inbound

Secure Tug-of-War (SecTOW): Iterative Defense-Attack Training with Reinforcement Learning for Multimodal Model Security cites this paper.

Secure Tug-of-War (SecTOW): Iterative Defense-Attack Training with Reinforcement Learning for Multimodal Model Security Stable Reinforcement Learning for Efficient Reasoning

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-06T12:09:39.636062Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T12:09:39.636062Z digest=sha256:0304d027475723d3759b1fdc9abef205dbe2249eb324276a055b5330d1b2afd4

Observation 83c4f67c-abf3-4632-be6f-9ddc3cb4bbdb · inbound

Gradient Extrapolation-Based Policy Optimization cites this paper.

Gradient Extrapolation-Based Policy Optimization Stable Reinforcement Learning for Efficient Reasoning

Reference 5

Resolution
verified exact
arxiv_id, observed 2026-05-11T03:55:56.726774Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-05-11T02:07:08.792030Z digest=sha256:815180a778f1a5014beed1a4f5ef767da5a76f0c1a10be52491be6ff8b0eafe9

Observation 32eafdbe-ed52-49a4-b269-446d843604d6 · inbound

CLORE: Content-Level Optimization for Reasoning Efficiency cites this paper.

CLORE: Content-Level Optimization for Reasoning Efficiency Stable Reinforcement Learning for Efficient Reasoning

Reference 8

Resolution
verified exact
arxiv_id, observed 2026-05-22T05:51:08.263012Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-22T05:50:23.111591Z digest=sha256:7610a49eb5c7d2b37691c70ce6f0649de3bc03a35e8a5b702fdb4040074fbb58

Observation 6967ca88-cf8c-4aac-a953-bb8e16f81ff5 · inbound

Learning When Not to Act: Mitigating Tool Abuse in Agentic Reinforcement Learning cites this paper.

Learning When Not to Act: Mitigating Tool Abuse in Agentic Reinforcement Learning Stable Reinforcement Learning for Efficient Reasoning

Reference 28

Resolution
metadata mismatch
arxiv_id, observed 2026-07-01T22:46:20.253672Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-06-28T14:59:57.983338Z digest=sha256:55212e397692a91dc45778b241cb55e3a6450f741b7d6fddd86274247f251b8b

Observation 2163f506-be2c-4659-b578-8de3e92f0899 · inbound

Length Penalties Make Chain-of-Thought Less Monitorable cites this paper.

Length Penalties Make Chain-of-Thought Less Monitorable Stable Reinforcement Learning for Efficient Reasoning

Reference 54

Resolution
unresolved
no resolver link, observed 2026-07-14T15:45:54.532529Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-14T15:45:54.532529Z digest=sha256:abb294b3fedfdf4151a0e71a5ad58b41052bae49d05517176aa72b0ce290f52f

Observation af84df61-620c-4d7d-b0a9-b19de84984bb · inbound

Length Penalties Make Chain-of-Thought Less Monitorable cites this paper.

Length Penalties Make Chain-of-Thought Less Monitorable Stable Reinforcement Learning for Efficient Reasoning

Reference 55

Resolution
unresolved
no resolver link, observed 2026-08-02T08:06:10.779080Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-02T08:06:10.779080Z digest=sha256:22143b294dbac420bc198fd3b95a606c59f15c8ffd3a985a86404cf4c67755c1

Observation ef30e1dc-c058-4e94-95c1-46f0948119e0 · inbound

Length Penalties Make Chain-of-Thought Less Monitorable cites this paper.

Length Penalties Make Chain-of-Thought Less Monitorable Stable Reinforcement Learning for Efficient Reasoning

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-04T04:30:27.449308Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T04:30:27.449308Z digest=sha256:3387c7935a516cf17ef71a04bb7f86d2bf6d3ab999265e1d4747c714fd614ded