Pith. sign in

Paper Citation Record · LEDGER

Learn to Reason Efficiently with Adaptive Length-based Reward Shaping

As of 18 August 2026, this Paper Citation Record lists 33 of 33 outbound references and 16 inbound Pith citation observations for arXiv:2505.15612.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2505.15612 v1

Coverage vector

measured 33 of 33 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-07T15:18:04.508009Z

measured 49 of 49 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-18T06:34:40.430872+00:00

measured 16 of 16 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-06T21:45:08.292468Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-04T09:19:43.881659Z

Reference resolution

33 of 33 outbound references displayed

  • verified exact0
  • verified fuzzy5
  • unresolved28
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 28e84fc4-fc33-4282-ba75-0814213730e5 · outbound

This paper cites L1: Controlling How Long A Reasoning Model Thinks With Reinforcement Learning.

Learn to Reason Efficiently with Adaptive Length-based Reward Shaping L1: Controlling How Long A Reasoning Model Thinks With Reinforcement Learning

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-07T15:17:59.467282Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:17:59.467282Z digest=sha256:8b3660b24fea33b3a937e7a0cd02515f3b7b557f0473026f35b8123d89e6c4ea

Observation 1cdb23a8-2f9d-4a88-b874-010ff7f73628 · outbound

This paper cites Arora and A.

Learn to Reason Efficiently with Adaptive Length-based Reward Shaping Arora and A

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-07T15:17:59.568317Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:17:59.568317Z digest=sha256:2b60ea7fd2eb10f5ec547c8bbd59ee3ba1e68cb54b86938766475bca127818ea

Observation eee27329-6959-4997-b431-f94e55efbdc8 · outbound

This paper cites Do NOT Think That Much for 2+3=? On the Overthinking of o1-Like LLMs.

Learn to Reason Efficiently with Adaptive Length-based Reward Shaping Do NOT Think That Much for 2+3=? On the Overthinking of o1-Like LLMs

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-07T15:17:59.667035Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:17:59.667035Z digest=sha256:a7c6478a47cac25fb7f74c7b95a7a5df2caf3f63400cd3367b50c968c5da8fb8

Observation ad411e10-4157-4f26-81fe-82d48ecf5652 · outbound

This paper cites DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning.

Learn to Reason Efficiently with Adaptive Length-based Reward Shaping DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-07T15:17:59.798415Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:17:59.798415Z digest=sha256:15bc9e20032ca8deec0d2f7aa34e7e39d2536649d8b31fcd819df700b455100f

Observation 95b32e92-0252-450b-8771-e9bd405a767c · outbound

This paper cites Gandhi, A.

Learn to Reason Efficiently with Adaptive Length-based Reward Shaping Gandhi, A

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:18:07.683827Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T15:17:59.894964Z digest=sha256:731ee642b1e266214442cddd31394e04dca71805185ea1bfb47dda0cd9a8fa2e

Observation 3d3cf30e-68a5-4ad6-8741-267b5a3249cb · outbound

This paper cites Training Large Language Models to Reason in a Continuous Latent Space.

Learn to Reason Efficiently with Adaptive Length-based Reward Shaping Training Large Language Models to Reason in a Continuous Latent Space

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-07T15:18:00.003763Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:18:00.003763Z digest=sha256:8bd48a6deddd018c1456c0f6dfd36e69cf01626f7e1dab0e586a0215855d8039

Observation 7b40c05d-5a62-4179-9807-fb0ab96c2062 · outbound

This paper cites an unresolved cited work.

Learn to Reason Efficiently with Adaptive Length-based Reward Shaping Unresolved cited work

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-07T15:18:00.102621Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:18:00.102621Z digest=sha256:22ae1223fcd4aac075ae940fae7d4fe495fb59da0e8a0f3028f6f5a8f2ac09bb

Observation af46f05c-9118-4372-a39c-0345ca05e32a · outbound

This paper cites Measuring Massive Multitask Language Understanding.

Learn to Reason Efficiently with Adaptive Length-based Reward Shaping Measuring Massive Multitask Language Understanding

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-07T15:18:00.199909Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:18:00.199909Z digest=sha256:8266fcec131baeac01730272f88089f8c9cd6d78cd1a5bed05067e2d7cd6c454

Observation b241c72e-f219-4882-a4c3-24328accca6a · outbound

This paper cites Hendrycks, C.

Learn to Reason Efficiently with Adaptive Length-based Reward Shaping Hendrycks, C

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:18:07.412863Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T15:18:00.275045Z digest=sha256:da47f935a8659f8ebc6eae5311133082bd9f66c145218a364e47594b81c8240b

Observation b913ec90-d3c6-4440-b6c0-178707ad69a4 · outbound

This paper cites ThinkPrune: Pruning Long Chain-of-Thought of LLMs via Reinforcement Learning.

Learn to Reason Efficiently with Adaptive Length-based Reward Shaping ThinkPrune: Pruning Long Chain-of-Thought of LLMs via Reinforcement Learning

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-07T15:18:00.357524Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:18:00.357524Z digest=sha256:ed8341103ac396abaaadb9c5421be86443e341a55d8c8a2f3ff679ab2ce5b67c

Observation fd11b564-bd11-4335-97e6-3cf7d9d1af9d · outbound

This paper cites Kimi k1.5: Scaling Reinforcement Learning with LLMs.

Learn to Reason Efficiently with Adaptive Length-based Reward Shaping Kimi k1.5: Scaling Reinforcement Learning with LLMs

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-07T15:18:00.496831Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:18:00.496831Z digest=sha256:332ab30b2428b1f41bfb2f3ef88df960165e729510049669561ebacb6339f828

Observation 954cf2c1-8e1d-48fd-b239-df1e25e27ff8 · outbound

This paper cites an unresolved cited work.

Learn to Reason Efficiently with Adaptive Length-based Reward Shaping Unresolved cited work

Reference 12

Resolution
unresolved
raw_fallback, observed 2026-08-07T15:18:06.995684Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T15:18:00.595825Z digest=sha256:3eaa9110998da45059826391e00b3efda1cca43b5f530658d08cfbb08ba40618

Observation 1d8b928e-392b-4957-aeb1-1c5fee32c0aa · outbound

This paper cites s1: Simple test-time scaling.

Learn to Reason Efficiently with Adaptive Length-based Reward Shaping s1: Simple test-time scaling

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-07T15:18:00.725959Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:18:00.725959Z digest=sha256:6b336fb69903b47daee8a1595b979dc40dc0afe0967e8f6ef16d3923840f0fbd

Observation 288839eb-69ee-4030-a74f-913083a2f4e8 · outbound

This paper cites Self-Training Elicits Concise Reasoning in Large Language Models.

Learn to Reason Efficiently with Adaptive Length-based Reward Shaping Self-Training Elicits Concise Reasoning in Large Language Models

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-07T15:18:00.832579Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:18:00.832579Z digest=sha256:fc3af2389efe4bd565cd0301a143d389d99fbde46ab2cabb1fd5fcf5171a0a7e

Observation d730ab99-18ba-4bbf-b24a-fe68527707d4 · outbound

This paper cites OpenAI o1 System Card.

Learn to Reason Efficiently with Adaptive Length-based Reward Shaping OpenAI o1 System Card

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-07T15:18:00.974743Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:18:00.974743Z digest=sha256:2a3ccd5d150f9af44941882843ad455c33146abb7a2f406d6aeff46f910476c3

Observation 4c36ddf6-0e6d-4bf7-b3dd-4867468d26a6 · outbound

This paper cites an unresolved cited work.

Learn to Reason Efficiently with Adaptive Length-based Reward Shaping Unresolved cited work

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-07T15:18:01.079263Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:18:01.079263Z digest=sha256:2d59a0bea30ed9a64edf2e8ffa835302429c6195528b67490b6aee7ff548977b

Observation 5d08d321-e896-4279-95d7-dd0755e96a78 · outbound

This paper cites GPQA: A Graduate-Level Google-Proof Q&A Benchmark.

Learn to Reason Efficiently with Adaptive Length-based Reward Shaping GPQA: A Graduate-Level Google-Proof Q&A Benchmark

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-07T15:18:01.206010Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:18:01.206010Z digest=sha256:87c6a4ef871487effc5d9185db7891e85c4c6dbff5f5651216a5f1c485f522c1

Observation b27b57e0-8994-4bc7-874c-23c3118b2407 · outbound

This paper cites an unresolved cited work.

Learn to Reason Efficiently with Adaptive Length-based Reward Shaping Unresolved cited work

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-07T15:18:01.330227Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:18:01.330227Z digest=sha256:409bfabf9068a4a54fa71ef9a57767a77016c57cab53e2f013c61d282a0966fd

Observation fcb88eb2-64be-4b8b-a1f7-82935f1ad521 · outbound

This paper cites Proximal Policy Optimization Algorithms.

Learn to Reason Efficiently with Adaptive Length-based Reward Shaping Proximal Policy Optimization Algorithms

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-07T15:18:01.437191Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:18:01.437191Z digest=sha256:90334b5d2792484a3411cc551c750a7474f0f8adf9c46b16ec83936c452ce737

Observation 35381598-08c1-420a-bc5c-2d6a5cc66f97 · outbound

This paper cites DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models.

Learn to Reason Efficiently with Adaptive Length-based Reward Shaping DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-07T15:18:01.525978Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:18:01.525978Z digest=sha256:59bb9136097e6a126857156e17c130948315226c731d505004fe67650712e9ae

Observation fc3775df-960a-4b0a-baaf-26fe0d688ef5 · outbound

This paper cites HybridFlow: A Flexible and Efficient RLHF Framework.

Learn to Reason Efficiently with Adaptive Length-based Reward Shaping HybridFlow: A Flexible and Efficient RLHF Framework

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-07T15:18:01.639730Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:18:01.639730Z digest=sha256:835785740fa0dc01ebf983a3f58109ccaed784cebe43e1c1903b9a1be41a0dcd

Observation 474c8158-6c00-4b95-a7e6-e9f2fa52867b · outbound

This paper cites Harnessing the Reasoning Economy: A Survey of Efficient Reasoning for Large Language Models.

Learn to Reason Efficiently with Adaptive Length-based Reward Shaping Harnessing the Reasoning Economy: A Survey of Efficient Reasoning for Large Language Models

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-07T15:18:01.733588Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:18:01.733588Z digest=sha256:893c0a364b5800d6242f3b5941c71cc686163a23e275423494b5ff9309042bdd

Observation 2a519d24-8e85-4981-8899-c72fa7b103f2 · outbound

This paper cites an unresolved cited work.

Learn to Reason Efficiently with Adaptive Length-based Reward Shaping Unresolved cited work

Reference 23

Resolution
unresolved
raw_fallback, observed 2026-08-07T15:18:06.594553Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T15:18:01.847962Z digest=sha256:81abcedaabef404e91e6238bd72cf0f59d605f84d20cb5d6316d480aea606e55

Observation bcae8ac7-cbd8-45d8-8249-037a2be47198 · outbound

This paper cites an unresolved cited work.

Learn to Reason Efficiently with Adaptive Length-based Reward Shaping Unresolved cited work

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-07T15:18:01.984634Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:18:01.984634Z digest=sha256:885eec31c45747a96eec056a56488c6954f232d01f3ddb0492dd520141ba365b

Observation 515b6d5c-4d55-4420-a9a5-dfc1e798df75 · outbound

This paper cites an unresolved cited work.

Learn to Reason Efficiently with Adaptive Length-based Reward Shaping Unresolved cited work

Reference 25

Resolution
unresolved
raw_fallback, observed 2026-08-07T15:18:06.205324Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T15:18:02.101816Z digest=sha256:f4e539cddc0f43b536fdf8f779c8f743683c37010bdeac86e3979bf496a369e8

Observation 20148de8-f039-471e-bae1-d6fbfbd23660 · outbound

This paper cites an unresolved cited work.

Learn to Reason Efficiently with Adaptive Length-based Reward Shaping Unresolved cited work

Reference 26

Resolution
unresolved
raw_fallback, observed 2026-08-07T15:18:05.874359Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T15:18:02.267130Z digest=sha256:37cbf817f8f9a4175f4bb1db24e3bc15e6e83a29dbfd301fbaa93cd0113aea60

Observation 41384598-14da-41dc-ae07-b941ab3cd98d · outbound

This paper cites DAPO: An Open-Source LLM Reinforcement Learning System at Scale.

Learn to Reason Efficiently with Adaptive Length-based Reward Shaping DAPO: An Open-Source LLM Reinforcement Learning System at Scale

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-07T15:18:02.954245Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:18:02.954245Z digest=sha256:1fca6041844e8e89eb7d70f8ab4c161fbaef20ed266788fb6ec8f3664e4c7dd8

Observation 5dfbb9aa-1fd5-4817-b7d3-6c317d6bb337 · outbound

This paper cites SimpleRL-Zoo: Investigating and Taming Zero Reinforcement Learning for Open Base Models in the Wild.

Learn to Reason Efficiently with Adaptive Length-based Reward Shaping SimpleRL-Zoo: Investigating and Taming Zero Reinforcement Learning for Open Base Models in the Wild

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-07T15:18:03.643174Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:18:03.643174Z digest=sha256:0acd5961553891ec3312a61768929ee2a4d2a4ae66b8fc5e333bd86df72ca148

Observation a01d5355-f674-45c6-923f-75fd472e93ab · outbound

This paper cites <think>...</think>.

Learn to Reason Efficiently with Adaptive Length-based Reward Shaping <think>...</think>

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:18:05.676753Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T15:18:04.240695Z digest=sha256:772566d1677a950129964d3f35d97a0533626e41409c70736d217f30b34d4a2c

Observation 577e7356-dde9-457a-becc-c5fcc438a479 · outbound

This paper cites Wait, subtracting a negative is like adding the positive, so that would be 3 + 6, which is 9.

Learn to Reason Efficiently with Adaptive Length-based Reward Shaping Wait, subtracting a negative is like adding the positive, so that would be 3 + 6, which is 9

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:18:05.438832Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T15:18:04.346749Z digest=sha256:3cc840242686985fef2ab5d8ffb1f8caee5567af4ef4677cf4e25fa43223dac5

Observation 01a5be31-685e-4785-a785-45382d5c0368 · outbound

This paper cites Calculate \( f(-1) \):\[f(-1) = \frac{3(-1) - 2}{-1 - 2} = \frac{-3 - 2}{-3} = \frac{-5}{-3} = \frac{5}{3}\]3.

Learn to Reason Efficiently with Adaptive Length-based Reward Shaping Calculate \( f(-1) \):\[f(-1) = \frac{3(-1) - 2}{-1 - 2} = \frac{-3 - 2}{-3} = \frac{-5}{-3} = \frac{5}{3}\]3

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:18:05.150909Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T15:18:04.508009Z digest=sha256:33885ab996786a9eb866e65010b2170b673555fc2491e3d5f5395becac7d4f4b

Observation e37737d0-0717-4936-91d9-e1195dd1ca07 · outbound

This paper cites Qwen2.5-Math Technical Report: Toward Mathematical Expert Model via Self-Improvement.

Learn to Reason Efficiently with Adaptive Length-based Reward Shaping Qwen2.5-Math Technical Report: Toward Mathematical Expert Model via Self-Improvement

Reference 2024

Resolution
unresolved
no resolver link, observed 2026-08-07T15:18:02.183957Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:18:02.183957Z digest=sha256:646bb5987fa33cc7a96b25a9853537faa9ae338a7c2e834d218eab4e0c88bec2

Observation a80ecce7-12be-4b1d-ab64-ae428b6d2493 · outbound

This paper cites Demystifying Long Chain-of-Thought Reasoning in LLMs.

Learn to Reason Efficiently with Adaptive Length-based Reward Shaping Demystifying Long Chain-of-Thought Reasoning in LLMs

Reference 2025

Resolution
unresolved
no resolver link, observed 2026-08-07T15:18:02.380737Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:18:02.380737Z digest=sha256:2f16158ee61b54d616b5ca4d35e91708e902344fc6bd9f4fad77fa43ce109068

Pith citing papers

Observation 8215d2d0-315b-4add-9aa8-a54da7b0598d · inbound

Stop Overthinking: A Survey on Efficient Reasoning for Large Language Models cites this paper.

Stop Overthinking: A Survey on Efficient Reasoning for Large Language Models Learn to Reason Efficiently with Adaptive Length-based Reward Shaping

Reference 116

Resolution
verified exact
arxiv_id, observed 2026-05-14T01:29:57.440468Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-14T01:29:56.480020Z digest=sha256:619f0be2133d80421b62f8442696711b8f535b236e85dd03bfea4b560b6c4096

Observation e2806cf4-e1a3-4291-8d88-34263ede61be · inbound

Garbage In, Reasoning Out? Why Benchmark Scores are Unreliable and What to Do About It cites this paper.

Garbage In, Reasoning Out? Why Benchmark Scores are Unreliable and What to Do About It Learn to Reason Efficiently with Adaptive Length-based Reward Shaping

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-06T21:45:08.292468Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T21:45:08.292468Z digest=sha256:64c1b821540be652ed3d7de843f7cd90d73664c39926e0389970707ea0050651

Observation 1cb0dbfa-7cc7-4266-9804-628e6c489f8c · inbound

Reconsidering Overthinking: Penalizing Internal and External Redundancy in CoT Reasoning cites this paper.

Reconsidering Overthinking: Penalizing Internal and External Redundancy in CoT Reasoning Learn to Reason Efficiently with Adaptive Length-based Reward Shaping

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-06T05:14:37.167029Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T05:14:37.167029Z digest=sha256:a8feed99f3abf175dbc656f2790b36182d0994aeb4f9507fa3c3895b1ab203b8

Observation 6db48e9a-15f4-4c36-8b2f-f582bbe38f7d · inbound

GDPO: Group reward-Decoupled Normalization Policy Optimization for Multi-reward RL Optimization cites this paper.

GDPO: Group reward-Decoupled Normalization Policy Optimization for Multi-reward RL Optimization Learn to Reason Efficiently with Adaptive Length-based Reward Shaping

Reference 1

Resolution
verified exact
arxiv_id, observed 2026-05-15T05:31:55.906712Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-15T05:31:55.864438Z digest=sha256:b4e2ab5ff1db52e2ee245772a613f6da56941e45edbf6f71c2efe1c4cfe11f95

Observation 57e59cd1-dfe9-492f-b79f-d81a8bceeacb · inbound

Numerically Optimizing Shortcuts to Adiabaticity: A Hybrid Control Strategy cites this paper.

Numerically Optimizing Shortcuts to Adiabaticity: A Hybrid Control Strategy Learn to Reason Efficiently with Adaptive Length-based Reward Shaping

Reference 25

Resolution
unresolved
no resolver link, observed 2026-07-13T14:28:35.916911Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-13T14:28:35.916911Z digest=sha256:befa24302fb54e66f62acb94927c8aa8570232dadce8b11b71374256f31061cf

Observation d1d3c71b-ce7b-40f6-9879-0ea2168f9d51 · inbound

T$^2$PO: Uncertainty-Guided Exploration Control for Stable Multi-Turn Agentic Reinforcement Learning cites this paper.

T$^2$PO: Uncertainty-Guided Exploration Control for Stable Multi-Turn Agentic Reinforcement Learning Learn to Reason Efficiently with Adaptive Length-based Reward Shaping

Reference 11

Resolution
metadata mismatch
arxiv_id, observed 2026-05-09T05:45:22.135373Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-08T19:36:00.359351Z digest=sha256:7d7a186ccacf06ffc7a75ed00db1934b9145bfe00a6dd816a599ba82d54e35b6

Observation decffc5e-5dc7-49d2-9828-01f11c1e8ba2 · inbound

Post Reasoning: Improving the Performance of Non-Thinking Models at No Cost cites this paper.

Post Reasoning: Improving the Performance of Non-Thinking Models at No Cost Learn to Reason Efficiently with Adaptive Length-based Reward Shaping

Reference 244

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T20:06:09.969293Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-05-08T10:19:08.451445Z digest=sha256:19496e63726838a1c16dcc76ce477af30f95d3ff640c032468e12698c527428f

Observation 621aa321-98df-4cf1-8d43-93fc91d81e1d · inbound

Implicit Compression Regularization: Concise Reasoning via Internal Shorter Distributions in RL Post-Training cites this paper.

Implicit Compression Regularization: Concise Reasoning via Internal Shorter Distributions in RL Post-Training Learn to Reason Efficiently with Adaptive Length-based Reward Shaping

Reference 21

Resolution
verified exact
arxiv_id, observed 2026-05-11T04:36:00.027949Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-11T01:12:20.864362Z digest=sha256:b8021a22769b59ed6b7e9ca6760ebd6a28146068147463b4a117d4ae1a37bce5

Observation 1bd3322e-6d0f-47ed-aae6-2d7d2b22e475 · inbound

AIPO: Learning to Reason from Active Interaction cites this paper.

AIPO: Learning to Reason from Active Interaction Learn to Reason Efficiently with Adaptive Length-based Reward Shaping

Reference 42

Resolution
verified exact
arxiv_id, observed 2026-05-12T08:06:31.370255Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-12T01:17:28.124867Z digest=sha256:9e74e29ee1a3f59d7a7777827bea02ff30d1283f146e22c20c9c71795f754b01

Observation 0b3bda08-724c-48d2-8257-5898f42dd122 · inbound

AIPO: Learning to Reason from Active Interaction cites this paper.

AIPO: Learning to Reason from Active Interaction Learn to Reason Efficiently with Adaptive Length-based Reward Shaping

Reference 42

Resolution
verified exact
arxiv_id, observed 2026-05-19T18:07:42.254204Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-19T18:07:27.492419Z digest=sha256:51cd78d729c311c640218a9f55c254392f0594aff70c666ea6ad66fa463b019a

Observation 0f8f4104-ab31-4fd6-a7e2-35d3b24f9e7b · inbound

LEAD: Length-Efficient Adaptive and Dynamic Reasoning for Large Language Models cites this paper.

LEAD: Length-Efficient Adaptive and Dynamic Reasoning for Large Language Models Learn to Reason Efficiently with Adaptive Length-based Reward Shaping

Reference 19

Resolution
verified exact
arxiv_id, observed 2026-05-12T02:31:16.776375Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-12T02:30:42.407934Z digest=sha256:72b28818ff40db89b888d31e1098f5c8533fd82f066028e95cfb69decc45d72d

Observation 9473f81e-b1c1-4bbf-b0f4-d52050b385e9 · inbound

CLORE: Content-Level Optimization for Reasoning Efficiency cites this paper.

CLORE: Content-Level Optimization for Reasoning Efficiency Learn to Reason Efficiently with Adaptive Length-based Reward Shaping

Reference 31

Resolution
verified exact
arxiv_id, observed 2026-05-22T05:51:08.161581Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-22T05:50:23.111591Z digest=sha256:f00ab0cb478d647b96951900a685605b11fd7c776d365ec53661ccdeeeaa5ca3

Observation 5f010621-c287-4485-9c86-3830856bcd32 · inbound

DVAO: Dynamic Variance-adaptive Advantage Optimization for Multi-reward Reinforcement Learning cites this paper.

DVAO: Dynamic Variance-adaptive Advantage Optimization for Multi-reward Reinforcement Learning Learn to Reason Efficiently with Adaptive Length-based Reward Shaping

Reference 14

Resolution
verified exact
arxiv_id, observed 2026-06-29T21:33:59.250975Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-06-29T21:29:31.723326Z digest=sha256:556942aa3d833f6a7611e7ac7f92234216ba333a25fe2a2dbbaa57a5cc6642a0

Observation c6993438-c84d-4363-b3cd-d682e4cf7a33 · inbound

SLAT: Segment-Level Adaptive Trimming for Efficient CoT Reasoning cites this paper.

SLAT: Segment-Level Adaptive Trimming for Efficient CoT Reasoning Learn to Reason Efficiently with Adaptive Length-based Reward Shaping

Reference 16

Resolution
verified exact
arxiv_id, observed 2026-06-28T22:32:44.426841Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-06-28T22:27:30.783923Z digest=sha256:7c2cc4e28ff7abd253a3d0498d87ae8011807d139c6919a59774ce603870316a

Observation d4422fa3-a96f-48cf-aaa0-9b986e39183d · inbound

CARE: Competence-Aware Reward Shaping for Adaptive Reasoning Length in Video-MLLMs cites this paper.

CARE: Competence-Aware Reward Shaping for Adaptive Reasoning Length in Video-MLLMs Learn to Reason Efficiently with Adaptive Length-based Reward Shaping

Reference 7

Resolution
verified exact
arxiv_id, observed 2026-07-04T03:29:31.421204Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-06-26T17:56:34.203135Z digest=sha256:74bcc716b99fd3370d082606d7a76781c7c7c58016365f8dd90c12660e4b8047

Observation 5c542a3d-e8a9-416e-a70b-6da1d6e9e7af · inbound

Beyond Penalizing Mistakes: Stabilizing Efficiency Training in Large Reasoning Models via Adaptive Correct-Only Rewards cites this paper.

Beyond Penalizing Mistakes: Stabilizing Efficiency Training in Large Reasoning Models via Adaptive Correct-Only Rewards Learn to Reason Efficiently with Adaptive Length-based Reward Shaping

Reference 18

Resolution
metadata mismatch
arxiv_id, observed 2026-07-04T09:19:43.883814Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-06-26T10:12:30.692295Z digest=sha256:68007c22bd65561af2407a39c13ca6f58c4ee7ab62f18dac60394ac03217b911