Pith. sign in

Paper Citation Record · LEDGER

Learn to Reason Efficiently with Adaptive Length-based Reward Shaping

As of 13 August 2026, this Paper Citation Record lists 33 of 33 outbound references and 16 inbound Pith citation observations for arXiv:2505.15612.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2505.15612 v1

Coverage vector

measured 33 of 33 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-07T15:18:04.508009Z

measured 49 of 49 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-13T06:32:02.005865+00:00

measured 16 of 16 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-06T21:45:08.292468Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-04T09:19:43.881659Z

Reference resolution

33 of 33 outbound references displayed

  • verified exact0
  • verified fuzzy5
  • unresolved28
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 28e84fc4-fc33-4282-ba75-0814213730e5 · outbound

This paper cites L1: Controlling How Long A Reasoning Model Thinks With Reinforcement Learning.

Learn to Reason Efficiently with Adaptive Length-based Reward Shaping L1: Controlling How Long A Reasoning Model Thinks With Reinforcement Learning

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-07T15:17:59.467282Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:17:59.467282Z digest=sha256:715530acce8795c7966f50587d2f831a63151a3b96957f5980a1ac5f86b5e2f7

Observation 1cdb23a8-2f9d-4a88-b874-010ff7f73628 · outbound

This paper cites Arora and A.

Learn to Reason Efficiently with Adaptive Length-based Reward Shaping Arora and A

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-07T15:17:59.568317Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:17:59.568317Z digest=sha256:5a211f9c4b7152ff1dd5e62fddc11f25b85c8f5d034761ecc0782ee5853fcd44

Observation eee27329-6959-4997-b431-f94e55efbdc8 · outbound

This paper cites Do NOT Think That Much for 2+3=? On the Overthinking of o1-Like LLMs.

Learn to Reason Efficiently with Adaptive Length-based Reward Shaping Do NOT Think That Much for 2+3=? On the Overthinking of o1-Like LLMs

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-07T15:17:59.667035Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:17:59.667035Z digest=sha256:3d6cdf7788145bf2714ee80b303d8ddcef35e5a51df7a984f801e73bf85a7abd

Observation ad411e10-4157-4f26-81fe-82d48ecf5652 · outbound

This paper cites DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning.

Learn to Reason Efficiently with Adaptive Length-based Reward Shaping DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-07T15:17:59.798415Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:17:59.798415Z digest=sha256:2104a906297699671642cda9a64cd7c861f6b857ffc4e08bd55fc44d24ee58bc

Observation 95b32e92-0252-450b-8771-e9bd405a767c · outbound

This paper cites Gandhi, A.

Learn to Reason Efficiently with Adaptive Length-based Reward Shaping Gandhi, A

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:18:07.683827Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-07T15:17:59.894964Z digest=sha256:8a479536dde7adaad90a2e8a4ae8fb8344792f3cde90cc117cdc2379f06657ae

Observation 3d3cf30e-68a5-4ad6-8741-267b5a3249cb · outbound

This paper cites Training Large Language Models to Reason in a Continuous Latent Space.

Learn to Reason Efficiently with Adaptive Length-based Reward Shaping Training Large Language Models to Reason in a Continuous Latent Space

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-07T15:18:00.003763Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:18:00.003763Z digest=sha256:e558c35f6babf66c29939626e68ecc5302a666261e54ebe5294670c2cf1eb530

Observation 7b40c05d-5a62-4179-9807-fb0ab96c2062 · outbound

This paper cites an unresolved cited work.

Learn to Reason Efficiently with Adaptive Length-based Reward Shaping Unresolved cited work

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-07T15:18:00.102621Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:18:00.102621Z digest=sha256:615cdaa509770651012aedebc606d7d58521bdad8d76e5225e149f52b3ff84aa

Observation af46f05c-9118-4372-a39c-0345ca05e32a · outbound

This paper cites Measuring Massive Multitask Language Understanding.

Learn to Reason Efficiently with Adaptive Length-based Reward Shaping Measuring Massive Multitask Language Understanding

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-07T15:18:00.199909Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:18:00.199909Z digest=sha256:7c3171b7b9f414b2073d28185d5d685ba651c4a12a33ef67d77cbf5dd36a10b5

Observation b241c72e-f219-4882-a4c3-24328accca6a · outbound

This paper cites Hendrycks, C.

Learn to Reason Efficiently with Adaptive Length-based Reward Shaping Hendrycks, C

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:18:07.412863Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-07T15:18:00.275045Z digest=sha256:e5862420b63be74a693875b9326f61fad11dd89b0d4511dd56a1e04745d0aa01

Observation b913ec90-d3c6-4440-b6c0-178707ad69a4 · outbound

This paper cites ThinkPrune: Pruning Long Chain-of-Thought of LLMs via Reinforcement Learning.

Learn to Reason Efficiently with Adaptive Length-based Reward Shaping ThinkPrune: Pruning Long Chain-of-Thought of LLMs via Reinforcement Learning

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-07T15:18:00.357524Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:18:00.357524Z digest=sha256:ec9c6a69d2de999dcf9873f7900e3caf91005c7ba627366c8e98b16631d4ee84

Observation fd11b564-bd11-4335-97e6-3cf7d9d1af9d · outbound

This paper cites Kimi k1.5: Scaling Reinforcement Learning with LLMs.

Learn to Reason Efficiently with Adaptive Length-based Reward Shaping Kimi k1.5: Scaling Reinforcement Learning with LLMs

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-07T15:18:00.496831Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:18:00.496831Z digest=sha256:d2af668a20eece2613430372aee8a62637260af42fec7c7f83c7cd9cd9fc32da

Observation 954cf2c1-8e1d-48fd-b239-df1e25e27ff8 · outbound

This paper cites an unresolved cited work.

Learn to Reason Efficiently with Adaptive Length-based Reward Shaping Unresolved cited work

Reference 12

Resolution
unresolved
raw_fallback, observed 2026-08-07T15:18:06.995684Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-07T15:18:00.595825Z digest=sha256:dfb4d2e8ab0c4c71591b833140f1f4b838dcd3f0c30e29b8d2f4774f3dc22be2

Observation 1d8b928e-392b-4957-aeb1-1c5fee32c0aa · outbound

This paper cites s1: Simple test-time scaling.

Learn to Reason Efficiently with Adaptive Length-based Reward Shaping s1: Simple test-time scaling

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-07T15:18:00.725959Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:18:00.725959Z digest=sha256:5245f4767a65b9b220bd6484778db793ad897c3af74d829de3304d05cf12c953

Observation 288839eb-69ee-4030-a74f-913083a2f4e8 · outbound

This paper cites Self-Training Elicits Concise Reasoning in Large Language Models.

Learn to Reason Efficiently with Adaptive Length-based Reward Shaping Self-Training Elicits Concise Reasoning in Large Language Models

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-07T15:18:00.832579Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:18:00.832579Z digest=sha256:1eaf9f3db019b890129f699a8712a5726b593000e2ec544f1869db67d940714c

Observation d730ab99-18ba-4bbf-b24a-fe68527707d4 · outbound

This paper cites OpenAI o1 System Card.

Learn to Reason Efficiently with Adaptive Length-based Reward Shaping OpenAI o1 System Card

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-07T15:18:00.974743Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:18:00.974743Z digest=sha256:351aa53e523ab81c4f6ef6e42a4367921382ca6d10ca84d0d8c1928ffad92f53

Observation 4c36ddf6-0e6d-4bf7-b3dd-4867468d26a6 · outbound

This paper cites an unresolved cited work.

Learn to Reason Efficiently with Adaptive Length-based Reward Shaping Unresolved cited work

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-07T15:18:01.079263Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:18:01.079263Z digest=sha256:feb056854414624c5fe67b6c22373a3e4a626fb3cbd28e89b123a94da3fe2f10

Observation 5d08d321-e896-4279-95d7-dd0755e96a78 · outbound

This paper cites GPQA: A Graduate-Level Google-Proof Q&A Benchmark.

Learn to Reason Efficiently with Adaptive Length-based Reward Shaping GPQA: A Graduate-Level Google-Proof Q&A Benchmark

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-07T15:18:01.206010Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:18:01.206010Z digest=sha256:82d710f9b4cc3ffe03d3224722de72414c5dca2d246406b4488ecb3f983cb1b2

Observation b27b57e0-8994-4bc7-874c-23c3118b2407 · outbound

This paper cites an unresolved cited work.

Learn to Reason Efficiently with Adaptive Length-based Reward Shaping Unresolved cited work

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-07T15:18:01.330227Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:18:01.330227Z digest=sha256:4429a963fd7ad9c33810942f70502664bd69bc143fa4bea21de875620ba8ac79

Observation fcb88eb2-64be-4b8b-a1f7-82935f1ad521 · outbound

This paper cites Proximal Policy Optimization Algorithms.

Learn to Reason Efficiently with Adaptive Length-based Reward Shaping Proximal Policy Optimization Algorithms

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-07T15:18:01.437191Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:18:01.437191Z digest=sha256:caa8d643949d9ee9e8630af4df1144a1ad3ce770d6e9de61aaf0706e20bf1fa0

Observation 35381598-08c1-420a-bc5c-2d6a5cc66f97 · outbound

This paper cites DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models.

Learn to Reason Efficiently with Adaptive Length-based Reward Shaping DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-07T15:18:01.525978Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:18:01.525978Z digest=sha256:383e6bc3e3cabef8cff45c713e3d2cfa98696374aab6c2c6e36ce588ea488024

Observation fc3775df-960a-4b0a-baaf-26fe0d688ef5 · outbound

This paper cites HybridFlow: A Flexible and Efficient RLHF Framework.

Learn to Reason Efficiently with Adaptive Length-based Reward Shaping HybridFlow: A Flexible and Efficient RLHF Framework

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-07T15:18:01.639730Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:18:01.639730Z digest=sha256:58a0c239c7ca8b2ba944a6bf13291f4b73b82b0ed8414e534e9a19ab3d486768

Observation 474c8158-6c00-4b95-a7e6-e9f2fa52867b · outbound

This paper cites Harnessing the Reasoning Economy: A Survey of Efficient Reasoning for Large Language Models.

Learn to Reason Efficiently with Adaptive Length-based Reward Shaping Harnessing the Reasoning Economy: A Survey of Efficient Reasoning for Large Language Models

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-07T15:18:01.733588Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:18:01.733588Z digest=sha256:eef82fcdd9de3c83a584aa6bb18455694273c84acae5a2a54293d57851e23c3f

Observation 2a519d24-8e85-4981-8899-c72fa7b103f2 · outbound

This paper cites an unresolved cited work.

Learn to Reason Efficiently with Adaptive Length-based Reward Shaping Unresolved cited work

Reference 23

Resolution
unresolved
raw_fallback, observed 2026-08-07T15:18:06.594553Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-07T15:18:01.847962Z digest=sha256:0e75538a63368d3136b65a6e8a21969acdf2350abf9a742ac471119077c0b0dc

Observation bcae8ac7-cbd8-45d8-8249-037a2be47198 · outbound

This paper cites an unresolved cited work.

Learn to Reason Efficiently with Adaptive Length-based Reward Shaping Unresolved cited work

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-07T15:18:01.984634Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:18:01.984634Z digest=sha256:7befa55a23d1bf715c1888e94644a16946dc845f1164c68d8ca92fb0c27028e5

Observation 515b6d5c-4d55-4420-a9a5-dfc1e798df75 · outbound

This paper cites an unresolved cited work.

Learn to Reason Efficiently with Adaptive Length-based Reward Shaping Unresolved cited work

Reference 25

Resolution
unresolved
raw_fallback, observed 2026-08-07T15:18:06.205324Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-07T15:18:02.101816Z digest=sha256:7cae02966b8ba4b76073f2a3c440188b98947103f7afd0599643fb532a091299

Observation 20148de8-f039-471e-bae1-d6fbfbd23660 · outbound

This paper cites an unresolved cited work.

Learn to Reason Efficiently with Adaptive Length-based Reward Shaping Unresolved cited work

Reference 26

Resolution
unresolved
raw_fallback, observed 2026-08-07T15:18:05.874359Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-07T15:18:02.267130Z digest=sha256:e0f251c332363e03fd1da125b64d398f46d8e89b82f202b5131dfae804624e8a

Observation 41384598-14da-41dc-ae07-b941ab3cd98d · outbound

This paper cites DAPO: An Open-Source LLM Reinforcement Learning System at Scale.

Learn to Reason Efficiently with Adaptive Length-based Reward Shaping DAPO: An Open-Source LLM Reinforcement Learning System at Scale

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-07T15:18:02.954245Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:18:02.954245Z digest=sha256:bd6b6721e2be9a0280ba19273a3caef05a7d8b8aa7c82929bc6adadf6f1dd3f5

Observation 5dfbb9aa-1fd5-4817-b7d3-6c317d6bb337 · outbound

This paper cites SimpleRL-Zoo: Investigating and Taming Zero Reinforcement Learning for Open Base Models in the Wild.

Learn to Reason Efficiently with Adaptive Length-based Reward Shaping SimpleRL-Zoo: Investigating and Taming Zero Reinforcement Learning for Open Base Models in the Wild

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-07T15:18:03.643174Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:18:03.643174Z digest=sha256:e40771885f328dbb317de2fcaae42492521d00f1a821d56e7f16c945aedaf3eb

Observation a01d5355-f674-45c6-923f-75fd472e93ab · outbound

This paper cites <think>...</think>.

Learn to Reason Efficiently with Adaptive Length-based Reward Shaping <think>...</think>

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:18:05.676753Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-07T15:18:04.240695Z digest=sha256:5d4a2408498f93afb394a2a634b822903b39865588b482c4bdf66ded64e8ad93

Observation 577e7356-dde9-457a-becc-c5fcc438a479 · outbound

This paper cites Wait, subtracting a negative is like adding the positive, so that would be 3 + 6, which is 9.

Learn to Reason Efficiently with Adaptive Length-based Reward Shaping Wait, subtracting a negative is like adding the positive, so that would be 3 + 6, which is 9

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:18:05.438832Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-07T15:18:04.346749Z digest=sha256:0cbb2247a4c6b5d76cfd705e5ece7774902ac85138194908cba39afda8af8aad

Observation 01a5be31-685e-4785-a785-45382d5c0368 · outbound

This paper cites Calculate \( f(-1) \):\[f(-1) = \frac{3(-1) - 2}{-1 - 2} = \frac{-3 - 2}{-3} = \frac{-5}{-3} = \frac{5}{3}\]3.

Learn to Reason Efficiently with Adaptive Length-based Reward Shaping Calculate \( f(-1) \):\[f(-1) = \frac{3(-1) - 2}{-1 - 2} = \frac{-3 - 2}{-3} = \frac{-5}{-3} = \frac{5}{3}\]3

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:18:05.150909Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-07T15:18:04.508009Z digest=sha256:dac81614be5b4d604a6eb589d6be6afad34e20efddcc89ffe56fbb425be6f65c

Observation e37737d0-0717-4936-91d9-e1195dd1ca07 · outbound

This paper cites Qwen2.5-Math Technical Report: Toward Mathematical Expert Model via Self-Improvement.

Learn to Reason Efficiently with Adaptive Length-based Reward Shaping Qwen2.5-Math Technical Report: Toward Mathematical Expert Model via Self-Improvement

Reference 2024

Resolution
unresolved
no resolver link, observed 2026-08-07T15:18:02.183957Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:18:02.183957Z digest=sha256:4fff5d69575b1e3981677befd94adbd5dc35780fab75a6a66c83c212d4c2c779

Observation a80ecce7-12be-4b1d-ab64-ae428b6d2493 · outbound

This paper cites Demystifying Long Chain-of-Thought Reasoning in LLMs.

Learn to Reason Efficiently with Adaptive Length-based Reward Shaping Demystifying Long Chain-of-Thought Reasoning in LLMs

Reference 2025

Resolution
unresolved
no resolver link, observed 2026-08-07T15:18:02.380737Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:18:02.380737Z digest=sha256:17c6bdaaec36cdbb090b28aa827dadfd1a9174d9209d1c9424aa27ddd955a54b

Pith citing papers

Observation 8215d2d0-315b-4add-9aa8-a54da7b0598d · inbound

Stop Overthinking: A Survey on Efficient Reasoning for Large Language Models cites this paper.

Stop Overthinking: A Survey on Efficient Reasoning for Large Language Models Learn to Reason Efficiently with Adaptive Length-based Reward Shaping

Reference 116

Resolution
verified exact
arxiv_id, observed 2026-05-14T01:29:57.440468Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-05-14T01:29:56.480020Z digest=sha256:d91f5a28124940e8ba8a8913e8f09f2379471a17b49960a95a02e595fb9e93dc

Observation e2806cf4-e1a3-4291-8d88-34263ede61be · inbound

Garbage In, Reasoning Out? Why Benchmark Scores are Unreliable and What to Do About It cites this paper.

Garbage In, Reasoning Out? Why Benchmark Scores are Unreliable and What to Do About It Learn to Reason Efficiently with Adaptive Length-based Reward Shaping

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-06T21:45:08.292468Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T21:45:08.292468Z digest=sha256:26cc08f22b00232a19be98772f846802d1f8596616f21970af4a7a827d48f9f1

Observation 1cb0dbfa-7cc7-4266-9804-628e6c489f8c · inbound

Reconsidering Overthinking: Penalizing Internal and External Redundancy in CoT Reasoning cites this paper.

Reconsidering Overthinking: Penalizing Internal and External Redundancy in CoT Reasoning Learn to Reason Efficiently with Adaptive Length-based Reward Shaping

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-06T05:14:37.167029Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T05:14:37.167029Z digest=sha256:4c06fd02ea388521da63f0107f9da52ef547823981fe38c058bd40645f547091

Observation 6db48e9a-15f4-4c36-8b2f-f582bbe38f7d · inbound

GDPO: Group reward-Decoupled Normalization Policy Optimization for Multi-reward RL Optimization cites this paper.

GDPO: Group reward-Decoupled Normalization Policy Optimization for Multi-reward RL Optimization Learn to Reason Efficiently with Adaptive Length-based Reward Shaping

Reference 1

Resolution
verified exact
arxiv_id, observed 2026-05-15T05:31:55.906712Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-05-15T05:31:55.864438Z digest=sha256:70ff255d3691f691539ab33fb128dda2f495c33dcae3c2a55930e82001be39e9

Observation 57e59cd1-dfe9-492f-b79f-d81a8bceeacb · inbound

Numerically Optimizing Shortcuts to Adiabaticity: A Hybrid Control Strategy cites this paper.

Numerically Optimizing Shortcuts to Adiabaticity: A Hybrid Control Strategy Learn to Reason Efficiently with Adaptive Length-based Reward Shaping

Reference 25

Resolution
unresolved
no resolver link, observed 2026-07-13T14:28:35.916911Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-13T14:28:35.916911Z digest=sha256:f43696178befa56ba5a16b7dab68d89eaee9ed3adb314bf7b3fb27b3b58a9245

Observation d1d3c71b-ce7b-40f6-9879-0ea2168f9d51 · inbound

T$^2$PO: Uncertainty-Guided Exploration Control for Stable Multi-Turn Agentic Reinforcement Learning cites this paper.

T$^2$PO: Uncertainty-Guided Exploration Control for Stable Multi-Turn Agentic Reinforcement Learning Learn to Reason Efficiently with Adaptive Length-based Reward Shaping

Reference 11

Resolution
metadata mismatch
arxiv_id, observed 2026-05-09T05:45:22.135373Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-05-08T19:36:00.359351Z digest=sha256:4c292fdb4297ab3d154ac98860a3053280d65f7737284acd7db7ddd357eea4e9

Observation decffc5e-5dc7-49d2-9828-01f11c1e8ba2 · inbound

Post Reasoning: Improving the Performance of Non-Thinking Models at No Cost cites this paper.

Post Reasoning: Improving the Performance of Non-Thinking Models at No Cost Learn to Reason Efficiently with Adaptive Length-based Reward Shaping

Reference 244

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T20:06:09.969293Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=arxiv_source observed=2026-05-08T10:19:08.451445Z digest=sha256:ec2b9caf2335c30b3e6a101085ca70fb033e7aaba374366249547028c2900a37

Observation 621aa321-98df-4cf1-8d43-93fc91d81e1d · inbound

Implicit Compression Regularization: Concise Reasoning via Internal Shorter Distributions in RL Post-Training cites this paper.

Implicit Compression Regularization: Concise Reasoning via Internal Shorter Distributions in RL Post-Training Learn to Reason Efficiently with Adaptive Length-based Reward Shaping

Reference 21

Resolution
verified exact
arxiv_id, observed 2026-05-11T04:36:00.027949Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-05-11T01:12:20.864362Z digest=sha256:e9267089bf94daf03a148eaf70c74562eb9ddb25f6646378e25e9c62357502bb

Observation 1bd3322e-6d0f-47ed-aae6-2d7d2b22e475 · inbound

AIPO: Learning to Reason from Active Interaction cites this paper.

AIPO: Learning to Reason from Active Interaction Learn to Reason Efficiently with Adaptive Length-based Reward Shaping

Reference 42

Resolution
verified exact
arxiv_id, observed 2026-05-12T08:06:31.370255Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-05-12T01:17:28.124867Z digest=sha256:ac703e33e6b1a9e2eb7b077d03f9ae1d74bc5819c515ef59ad804e3c63fb90a8

Observation 0b3bda08-724c-48d2-8257-5898f42dd122 · inbound

AIPO: Learning to Reason from Active Interaction cites this paper.

AIPO: Learning to Reason from Active Interaction Learn to Reason Efficiently with Adaptive Length-based Reward Shaping

Reference 42

Resolution
verified exact
arxiv_id, observed 2026-05-19T18:07:42.254204Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-05-19T18:07:27.492419Z digest=sha256:8ff4f33d2f567da2e5225aa3be89dd53a085dc86bc872ce6648ac9e8d468a30d

Observation 0f8f4104-ab31-4fd6-a7e2-35d3b24f9e7b · inbound

LEAD: Length-Efficient Adaptive and Dynamic Reasoning for Large Language Models cites this paper.

LEAD: Length-Efficient Adaptive and Dynamic Reasoning for Large Language Models Learn to Reason Efficiently with Adaptive Length-based Reward Shaping

Reference 19

Resolution
verified exact
arxiv_id, observed 2026-05-12T02:31:16.776375Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-05-12T02:30:42.407934Z digest=sha256:49f153848dd1f86167e083423f0d6c2f4cc707e6db2a1154ba4dbf9663f54311

Observation 9473f81e-b1c1-4bbf-b0f4-d52050b385e9 · inbound

CLORE: Content-Level Optimization for Reasoning Efficiency cites this paper.

CLORE: Content-Level Optimization for Reasoning Efficiency Learn to Reason Efficiently with Adaptive Length-based Reward Shaping

Reference 31

Resolution
verified exact
arxiv_id, observed 2026-05-22T05:51:08.161581Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-05-22T05:50:23.111591Z digest=sha256:902508e209071ca68ab6ddb962b4295796a1967425f7071eab124f433b5e8f43

Observation 5f010621-c287-4485-9c86-3830856bcd32 · inbound

DVAO: Dynamic Variance-adaptive Advantage Optimization for Multi-reward Reinforcement Learning cites this paper.

DVAO: Dynamic Variance-adaptive Advantage Optimization for Multi-reward Reinforcement Learning Learn to Reason Efficiently with Adaptive Length-based Reward Shaping

Reference 14

Resolution
verified exact
arxiv_id, observed 2026-06-29T21:33:59.250975Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-06-29T21:29:31.723326Z digest=sha256:bb92ce36c65af2aa43e864251dc9f8edc3b20849978d0817a39fea55a5f1d833

Observation c6993438-c84d-4363-b3cd-d682e4cf7a33 · inbound

SLAT: Segment-Level Adaptive Trimming for Efficient CoT Reasoning cites this paper.

SLAT: Segment-Level Adaptive Trimming for Efficient CoT Reasoning Learn to Reason Efficiently with Adaptive Length-based Reward Shaping

Reference 16

Resolution
verified exact
arxiv_id, observed 2026-06-28T22:32:44.426841Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-06-28T22:27:30.783923Z digest=sha256:d0f9cdae0e6060361b535ed5753d3138bd9878458e8de7a714569fa81ec66901

Observation d4422fa3-a96f-48cf-aaa0-9b986e39183d · inbound

CARE: Competence-Aware Reward Shaping for Adaptive Reasoning Length in Video-MLLMs cites this paper.

CARE: Competence-Aware Reward Shaping for Adaptive Reasoning Length in Video-MLLMs Learn to Reason Efficiently with Adaptive Length-based Reward Shaping

Reference 7

Resolution
verified exact
arxiv_id, observed 2026-07-04T03:29:31.421204Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-06-26T17:56:34.203135Z digest=sha256:f0f16fd6653a9c075994940fc07046a600405aceff65d922e975c4571c0d06a6

Observation 5c542a3d-e8a9-416e-a70b-6da1d6e9e7af · inbound

Beyond Penalizing Mistakes: Stabilizing Efficiency Training in Large Reasoning Models via Adaptive Correct-Only Rewards cites this paper.

Beyond Penalizing Mistakes: Stabilizing Efficiency Training in Large Reasoning Models via Adaptive Correct-Only Rewards Learn to Reason Efficiently with Adaptive Length-based Reward Shaping

Reference 18

Resolution
metadata mismatch
arxiv_id, observed 2026-07-04T09:19:43.883814Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=arxiv_source observed=2026-06-26T10:12:30.692295Z digest=sha256:6a74fc5ac554aed8c946a09751b30389c62e4eaaf27007d16debd8440f060b3e