Pith. sign in

Paper Citation Record · LEDGER

Thinking Fast and Right: Balancing Accuracy and Reasoning Length with Adaptive Rewards

As of 13 August 2026, this Paper Citation Record lists 35 of 35 outbound references and 8 inbound Pith citation observations for arXiv:2505.18298.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2505.18298 v1

Coverage vector

measured 35 of 35 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-07T14:37:03.196744Z

measured 43 of 43 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-13T06:32:02.005865+00:00

measured 8 of 8 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-06T21:45:10.270039Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-04T09:19:43.854814Z

Reference resolution

35 of 35 outbound references displayed

  • verified exact0
  • verified fuzzy4
  • unresolved31
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 442ab1f7-9fc6-484f-bc92-061aefa9a772 · outbound

This paper cites L1: Controlling How Long A Reasoning Model Thinks With Reinforcement Learning.

Thinking Fast and Right: Balancing Accuracy and Reasoning Length with Adaptive Rewards L1: Controlling How Long A Reasoning Model Thinks With Reinforcement Learning

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-07T14:37:00.874264Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:37:00.874264Z digest=sha256:b223934d3e99bb811f232a1784e36470edf4d91d4bbf07444b77f03ce9e18664

Observation 8f53ce4f-0aef-4a8a-8e9f-3731ce5fe379 · outbound

This paper cites Training language models to reason efficiently.arXiv preprint arXiv:2502.04463, 2025.

Thinking Fast and Right: Balancing Accuracy and Reasoning Length with Adaptive Rewards Training language models to reason efficiently.arXiv preprint arXiv:2502.04463, 2025

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-07T14:37:01.003276Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:37:01.003276Z digest=sha256:9d4763602f5c4eaf4c7e5da7c078cc7e89018200d270d6a4e75e70234ec4f043

Observation 72a9f013-a6a3-47ba-a1df-020837f1e845 · outbound

This paper cites Do NOT Think That Much for 2+3=? On the Overthinking of o1-Like LLMs.

Thinking Fast and Right: Balancing Accuracy and Reasoning Length with Adaptive Rewards Do NOT Think That Much for 2+3=? On the Overthinking of o1-Like LLMs

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-07T14:37:01.107838Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:37:01.107838Z digest=sha256:ce485d92c0b05428ccb38cf2a427494f803e253d038508c53ecce3747e861054

Observation b3663651-1041-422a-858d-3edb87b86a86 · outbound

This paper cites Training Verifiers to Solve Math Word Problems.

Thinking Fast and Right: Balancing Accuracy and Reasoning Length with Adaptive Rewards Training Verifiers to Solve Math Word Problems

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-07T14:37:01.174456Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:37:01.174456Z digest=sha256:3e5b99f055a9141bec211a724d270b26e72db7008dccde0d9179d9258808b63b

Observation 51308257-6ebd-4863-a821-4f87adf34c01 · outbound

This paper cites competitive programming platform.https://codeforces.com/, 2025.

Thinking Fast and Right: Balancing Accuracy and Reasoning Length with Adaptive Rewards competitive programming platform.https://codeforces.com/, 2025

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:37:04.503936Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-07T14:37:01.215547Z digest=sha256:cefabc9f8883fbe1121bd9353819fb59a1f8328edf9c1887473db429233a7cfe

Observation c970b6b3-d015-41fc-ba95-5c2ab9399b22 · outbound

This paper cites Efficient reasoning models: A survey.arXiv preprint arXiv:2504.10903, 2025.

Thinking Fast and Right: Balancing Accuracy and Reasoning Length with Adaptive Rewards Efficient reasoning models: A survey.arXiv preprint arXiv:2504.10903, 2025

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-07T14:37:01.267450Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:37:01.267450Z digest=sha256:06349d883095b4918efb006173eae7c2762042302aff5d0667f002d227489b84

Observation 5ee06ac6-e9c8-4669-a56a-d1e83aa8f32b · outbound

This paper cites Omni-MATH: A Universal Olympiad Level Mathematic Benchmark For Large Language Models.

Thinking Fast and Right: Balancing Accuracy and Reasoning Length with Adaptive Rewards Omni-MATH: A Universal Olympiad Level Mathematic Benchmark For Large Language Models

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-07T14:37:01.324509Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:37:01.324509Z digest=sha256:e291fdc613446e274efb6f3cfe4668bf2835b78a1907b92745614cc8553689ff

Observation 38674f00-a0af-40af-b4f9-bd9e328a6162 · outbound

This paper cites Scaling laws for reward model overoptimization.

Thinking Fast and Right: Balancing Accuracy and Reasoning Length with Adaptive Rewards Scaling laws for reward model overoptimization

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-07T14:37:01.397009Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:37:01.397009Z digest=sha256:7ae3109f0654aa2d414fd3134d17bd16c0b8c81909f50f7bc0efac994153d266

Observation 8832b3e4-85ff-4a45-9030-0a56a9b4c4ce · outbound

This paper cites DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning.

Thinking Fast and Right: Balancing Accuracy and Reasoning Length with Adaptive Rewards DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-07T14:37:01.478805Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:37:01.478805Z digest=sha256:b1cb3bc12ade9a8bf4b8b0337063dd7575e476e31dd24c52de7f8131c92c5fe6

Observation b90ca9c3-e075-477f-9356-016a9a549240 · outbound

This paper cites Token-Budget-Aware LLM Reasoning.

Thinking Fast and Right: Balancing Accuracy and Reasoning Length with Adaptive Rewards Token-Budget-Aware LLM Reasoning

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-07T14:37:01.537382Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:37:01.537382Z digest=sha256:d4755f9ed8573e07c628825401beac1dc067045fb661256799f4736b7bdbd64f

Observation be7cfda0-09a0-4455-874b-d240ddc9084d · outbound

This paper cites Olympiadbench: A challenging benchmark for promoting agi with olympiad-level bilingual multimodal scientific problems.

Thinking Fast and Right: Balancing Accuracy and Reasoning Length with Adaptive Rewards Olympiadbench: A challenging benchmark for promoting agi with olympiad-level bilingual multimodal scientific problems

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-07T14:37:01.603480Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:37:01.603480Z digest=sha256:196b419d3943ed94c40158a0c710b5771bd4aa2e9f70e3e7440d63744c2beadc

Observation dcef9361-7a71-4c42-88a1-def13466c23d · outbound

This paper cites Measuring Mathematical Problem Solving With the MATH Dataset.

Thinking Fast and Right: Balancing Accuracy and Reasoning Length with Adaptive Rewards Measuring Mathematical Problem Solving With the MATH Dataset

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-07T14:37:01.676325Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:37:01.676325Z digest=sha256:d2e5e8eea23fff4be6696119aec664f4a499d9cd31875f131687fc9bbbfaf524

Observation 664a0074-5786-45a7-a606-0116e12b066e · outbound

This paper cites ThinkPrune: Pruning Long Chain-of-Thought of LLMs via Reinforcement Learning.

Thinking Fast and Right: Balancing Accuracy and Reasoning Length with Adaptive Rewards ThinkPrune: Pruning Long Chain-of-Thought of LLMs via Reinforcement Learning

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-07T14:37:01.733717Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:37:01.733717Z digest=sha256:c0dd94de0fb49b2eef97597742c39772a5e58fa626a7e8977d3789c2096cd37b

Observation 5856107e-8d78-40cd-af8d-bca9abfc7d6f · outbound

This paper cites C3ot: Generating shorter chain-of- thought without compromising effectiveness.

Thinking Fast and Right: Balancing Accuracy and Reasoning Length with Adaptive Rewards C3ot: Generating shorter chain-of- thought without compromising effectiveness

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-07T14:37:01.791923Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:37:01.791923Z digest=sha256:cdacc1df243af0ba2a8238a22369e5d0c4db3e195c6bae3be0787736f7fea24d

Observation 7ab5e3d9-f52c-45d2-b9af-3ab106d5cec7 · outbound

This paper cites Training Language Models to Self-Correct via Reinforcement Learning.

Thinking Fast and Right: Balancing Accuracy and Reasoning Length with Adaptive Rewards Training Language Models to Self-Correct via Reinforcement Learning

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-07T14:37:01.878223Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:37:01.878223Z digest=sha256:09f127153ef1e264f0a6f84a8fb2b31a3ca6bd7c588d19609fbef367f3524414

Observation 31e03873-e1f0-4215-94d9-6e52937bf939 · outbound

This paper cites Let’s verify step by step.

Thinking Fast and Right: Balancing Accuracy and Reasoning Length with Adaptive Rewards Let’s verify step by step

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-07T14:37:01.939588Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:37:01.939588Z digest=sha256:20c069fb804614747decfec85290f4407d68a5c5c138ea19c893b9141837df83

Observation be441b8d-1204-4559-8db8-fb50648a4f2a · outbound

This paper cites O1-Pruner: Length-Harmonizing Fine-Tuning for O1-Like Reasoning Pruning.

Thinking Fast and Right: Balancing Accuracy and Reasoning Length with Adaptive Rewards O1-Pruner: Length-Harmonizing Fine-Tuning for O1-Like Reasoning Pruning

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-07T14:37:01.996062Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:37:01.996062Z digest=sha256:def65fdacd62c6cbf2a1af81e3e13835eb2508fd977c7df28b2ac6b005402518

Observation 1076a454-ec4d-4c8c-b4cb-24c664984679 · outbound

This paper cites Tang, Manan Roongta, Colin Cai, Jeffrey Luo, Tianjun Zhang, Li Erran Li, Raluca Ada Popa, and Ion Stoica.

Thinking Fast and Right: Balancing Accuracy and Reasoning Length with Adaptive Rewards Tang, Manan Roongta, Colin Cai, Jeffrey Luo, Tianjun Zhang, Li Erran Li, Raluca Ada Popa, and Ion Stoica

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:37:04.310465Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-07T14:37:02.066745Z digest=sha256:6a5bd8fa534b6d580771aa529fe127a7bf3ad589c1c8f0045b62246468009ed5

Observation 505ccb87-69d6-4fee-9cd0-1cc0f6d00f67 · outbound

This paper cites Imitate, Explore, and Self-Improve: A Reproduction Report on Slow-thinking Reasoning Systems.

Thinking Fast and Right: Balancing Accuracy and Reasoning Length with Adaptive Rewards Imitate, Explore, and Self-Improve: A Reproduction Report on Slow-thinking Reasoning Systems

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-07T14:37:02.150424Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:37:02.150424Z digest=sha256:54a17bc2e8b35096ad875874be4544e4de5a341810c31dd42b33793742fb7608

Observation 8e1c8307-a8e8-4f91-836a-98a20af4269f · outbound

This paper cites Self-Training Elicits Concise Reasoning in Large Language Models.

Thinking Fast and Right: Balancing Accuracy and Reasoning Length with Adaptive Rewards Self-Training Elicits Concise Reasoning in Large Language Models

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-07T14:37:02.225034Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:37:02.225034Z digest=sha256:9418a6ba6971e172fde1f9c963fa48f386a6a706c474baf55ed100b7357c32cf

Observation abf27e81-4979-46ec-bfe9-c90a609f7654 · outbound

This paper cites Concise Thoughts: Impact of Output Length on LLM Reasoning and Cost.

Thinking Fast and Right: Balancing Accuracy and Reasoning Length with Adaptive Rewards Concise Thoughts: Impact of Output Length on LLM Reasoning and Cost

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-07T14:37:02.280834Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:37:02.280834Z digest=sha256:856aeb67dd65bf4380fab20a37a59c09b9182c0343b134e0537001acd0200b91

Observation 6b90efee-0b7f-4959-addf-336cbe67c12a · outbound

This paper cites Learning to reason with llms.https://openai.com/index/learning-to-reason-with-llms/, 2024.

Thinking Fast and Right: Balancing Accuracy and Reasoning Length with Adaptive Rewards Learning to reason with llms.https://openai.com/index/learning-to-reason-with-llms/, 2024

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:37:04.201016Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-07T14:37:02.360443Z digest=sha256:caa75a22d0e30767fd6326d0ab6fe6f0806b16162f9979b907327ae8e06d98ce

Observation 60242fa7-0d39-445e-a632-bc67c3386528 · outbound

This paper cites Gpqa: A graduate-level google-proof q&a benchmark.

Thinking Fast and Right: Balancing Accuracy and Reasoning Length with Adaptive Rewards Gpqa: A graduate-level google-proof q&a benchmark

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-07T14:37:02.400867Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:37:02.400867Z digest=sha256:c1347161c7c4ea905b14a9e8cea3b789074fa281ab70373833763250b579a41b

Observation 08aef909-c430-4afe-9729-9393712e7ad7 · outbound

This paper cites Code Llama: Open Foundation Models for Code.

Thinking Fast and Right: Balancing Accuracy and Reasoning Length with Adaptive Rewards Code Llama: Open Foundation Models for Code

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-07T14:37:02.480031Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:37:02.480031Z digest=sha256:7d159864c9674839ef7587c106575b17953bb69acec532434f8a3762996c3be7

Observation 158a1cca-3fa1-41c1-9c31-f17173af4c38 · outbound

This paper cites DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models.

Thinking Fast and Right: Balancing Accuracy and Reasoning Length with Adaptive Rewards DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-07T14:37:02.554391Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:37:02.554391Z digest=sha256:0a651d2d3c4b49cf74b1b8ead133aae996057ef11508a3ba841d6ef6cfbf7116

Observation 0f4b5653-1520-4686-9390-b95c16d4e23b · outbound

This paper cites Between Underthinking and Overthinking: An Empirical Study of Reasoning Length and correctness in LLMs.

Thinking Fast and Right: Balancing Accuracy and Reasoning Length with Adaptive Rewards Between Underthinking and Overthinking: An Empirical Study of Reasoning Length and correctness in LLMs

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-07T14:37:02.625061Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:37:02.625061Z digest=sha256:b8ea3de24e97cf5609a5edb9b61a265baa3a99d8e4fe7270dfc6e3bd99cdb7f4

Observation 8c121087-817b-458e-85a1-5ec109884a4b · outbound

This paper cites Stop Overthinking: A Survey on Efficient Reasoning for Large Language Models.

Thinking Fast and Right: Balancing Accuracy and Reasoning Length with Adaptive Rewards Stop Overthinking: A Survey on Efficient Reasoning for Large Language Models

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-07T14:37:02.693985Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:37:02.693985Z digest=sha256:3a603e17d6212f97271793e90aaaa404402b9730c2750db5f55ee7239c6a5e79

Observation f35abcbc-a4b1-4f97-bd70-f2bc9589c7f4 · outbound

This paper cites Qwq-32b-preview.https://qwenlm.github.io/blog/qwq-32b-preview/, 2025.

Thinking Fast and Right: Balancing Accuracy and Reasoning Length with Adaptive Rewards Qwq-32b-preview.https://qwenlm.github.io/blog/qwq-32b-preview/, 2025

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:37:04.090914Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-07T14:37:02.767940Z digest=sha256:2f835a0f95867b180eb3d01ceb89f1c119ca96b227e4f25dbe4db69d909a1e30

Observation 8ee41af2-69d6-45d3-bbe2-f7952706419e · outbound

This paper cites Harnessing the Reasoning Economy: A Survey of Efficient Reasoning for Large Language Models.

Thinking Fast and Right: Balancing Accuracy and Reasoning Length with Adaptive Rewards Harnessing the Reasoning Economy: A Survey of Efficient Reasoning for Large Language Models

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-07T14:37:02.822409Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:37:02.822409Z digest=sha256:70d29d8732353d3f79c1cf605a924a3dc585e1c8a083c0c66872a673663550df

Observation 91d8a4e6-f946-45b6-a219-4417b9846b2b · outbound

This paper cites Self-Consistency Improves Chain of Thought Reasoning in Language Models.

Thinking Fast and Right: Balancing Accuracy and Reasoning Length with Adaptive Rewards Self-Consistency Improves Chain of Thought Reasoning in Language Models

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-07T14:37:02.885543Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:37:02.885543Z digest=sha256:fc4de6e95d5eaba3af1e0359279f2f676d625ac0e696aa2c87016ce74270fcd8

Observation 9d748111-dfab-45c2-b6bb-fd5012ffa94b · outbound

This paper cites Chain-of-thought prompting elicits reasoning in large language models.

Thinking Fast and Right: Balancing Accuracy and Reasoning Length with Adaptive Rewards Chain-of-thought prompting elicits reasoning in large language models

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-07T14:37:02.944987Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:37:02.944987Z digest=sha256:0ef32408b714f3e68607ff71becd7892a19ad331148f7814dd5f3bdf784282d6

Observation b3f68f8b-fa53-436f-93dc-151615958a5f · outbound

This paper cites Unlocking Efficient Long-to-Short LLM Reasoning with Model Merging.

Thinking Fast and Right: Balancing Accuracy and Reasoning Length with Adaptive Rewards Unlocking Efficient Long-to-Short LLM Reasoning with Model Merging

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-07T14:37:03.004066Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:37:03.004066Z digest=sha256:0ad641407ed1fa0df6467a83b71099fa2078263ccf92e1dfe381915cbd37d5a4

Observation 1bd1b620-86a1-489e-bec5-64558b313831 · outbound

This paper cites Tokenskip: Controllable chain-of-thought compression in llms.arXiv preprint arXiv:2502.12067, 2025.

Thinking Fast and Right: Balancing Accuracy and Reasoning Length with Adaptive Rewards Tokenskip: Controllable chain-of-thought compression in llms.arXiv preprint arXiv:2502.12067, 2025

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-07T14:37:03.060411Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:37:03.060411Z digest=sha256:3ec7153c78e1d97185524d6506a54af2dab12797bbc35142ad2307a5973eebe9

Observation 56280a21-99c1-43fe-94e3-d410c4ac50c2 · outbound

This paper cites Towards thinking-optimal scaling of test-time compute for llm reasoning.arXiv preprint arXiv:2502.18080, 2025.

Thinking Fast and Right: Balancing Accuracy and Reasoning Length with Adaptive Rewards Towards thinking-optimal scaling of test-time compute for llm reasoning.arXiv preprint arXiv:2502.18080, 2025

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-07T14:37:03.121217Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:37:03.121217Z digest=sha256:0e6be1fa8c6552471a45cfcd9e89495a86b4f8b8f1f7873fd8875cd5f2edd6ed

Observation 68153ac2-28be-4582-9590-89a703dc1273 · outbound

This paper cites Shorterbetter: Guiding reasoning models to find optimal inference length for efficient reasoning.arXiv preprint arXiv:2504.21370, 2025.

Thinking Fast and Right: Balancing Accuracy and Reasoning Length with Adaptive Rewards Shorterbetter: Guiding reasoning models to find optimal inference length for efficient reasoning.arXiv preprint arXiv:2504.21370, 2025

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-07T14:37:03.196744Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:37:03.196744Z digest=sha256:2ea705f47367ba47ba32823e1882f64bd15f0d01d870d178bdd4cb413f990475

Pith citing papers

Observation d52c04d0-32e9-4730-b36a-609bd4a0a2d3 · inbound

Do Thinking Tokens Help or Trap? Towards More Efficient Large Reasoning Model cites this paper.

Do Thinking Tokens Help or Trap? Towards More Efficient Large Reasoning Model Thinking Fast and Right: Balancing Accuracy and Reasoning Length with Adaptive Rewards

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-06T21:45:10.270039Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T21:45:10.270039Z digest=sha256:aa7b8261fe67fa28b37b171ae28687dd722a2ba0c0ded4c54ec5963241c564d8

Observation 02d91076-4512-4bc6-ad22-2425111373bc · inbound

Towards Concise and Adaptive Thinking in Large Reasoning Models: A Survey cites this paper.

Towards Concise and Adaptive Thinking in Large Reasoning Models: A Survey Thinking Fast and Right: Balancing Accuracy and Reasoning Length with Adaptive Rewards

Reference 168

Resolution
unresolved
no resolver link, observed 2026-08-06T17:54:17.672973Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T17:54:17.672973Z digest=sha256:a1f2e06a8ea477c0b3fb6bfc8674778ed9cccc5ae60054b27107da55ecd1e842

Observation e800aca3-effb-423a-8206-db125821c2c5 · inbound

Learning to Reason Efficiently with Discounted Reinforcement Learning cites this paper.

Learning to Reason Efficiently with Discounted Reinforcement Learning Thinking Fast and Right: Balancing Accuracy and Reasoning Length with Adaptive Rewards

Reference 2010

Resolution
unresolved
no resolver link, observed 2026-08-04T07:59:53.678805Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T07:59:53.678805Z digest=sha256:fd778abd47d180796f9742beafd016d215f52be4829b9ddef9b6fab573d7aef4

Observation 56fffa5f-907e-49f8-a5be-9506a292801d · inbound

GDPO: Group reward-Decoupled Normalization Policy Optimization for Multi-reward RL Optimization cites this paper.

GDPO: Group reward-Decoupled Normalization Policy Optimization for Multi-reward RL Optimization Thinking Fast and Right: Balancing Accuracy and Reasoning Length with Adaptive Rewards

Reference 45

Resolution
verified exact
arxiv_id, observed 2026-05-15T05:31:56.036833Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-05-15T05:31:55.864438Z digest=sha256:1d109945fedea5be49afa97d0b5ac2b700bd6c6c83a84fcc700ceb4ae30d4a87

Observation 4369648a-efbc-44ca-b123-55488db6a1dc · inbound

Implicit Compression Regularization: Concise Reasoning via Internal Shorter Distributions in RL Post-Training cites this paper.

Implicit Compression Regularization: Concise Reasoning via Internal Shorter Distributions in RL Post-Training Thinking Fast and Right: Balancing Accuracy and Reasoning Length with Adaptive Rewards

Reference 25

Resolution
verified exact
arxiv_id, observed 2026-05-11T04:35:59.701953Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-05-11T01:12:20.864362Z digest=sha256:0696e2def7fc00cf856658538475195ed564e4d2c919846d9b5d29c7151ffd13

Observation f1e542d4-bdda-4a54-853a-dcf3a992cc5b · inbound

CLORE: Content-Level Optimization for Reasoning Efficiency cites this paper.

CLORE: Content-Level Optimization for Reasoning Efficiency Thinking Fast and Right: Balancing Accuracy and Reasoning Length with Adaptive Rewards

Reference 42

Resolution
verified exact
arxiv_id, observed 2026-05-22T05:51:08.229808Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-05-22T05:50:23.111591Z digest=sha256:62323aecccd389d365a358f4c2d8acd7e4dbd6cf44c74f29bf757dd246a191b9

Observation bb273491-ab63-42ba-8e1b-99c1e0dfdacf · inbound

CARE: Competence-Aware Reward Shaping for Adaptive Reasoning Length in Video-MLLMs cites this paper.

CARE: Competence-Aware Reward Shaping for Adaptive Reasoning Length in Video-MLLMs Thinking Fast and Right: Balancing Accuracy and Reasoning Length with Adaptive Rewards

Reference 6

Resolution
verified exact
arxiv_id, observed 2026-07-04T03:29:31.439181Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-06-26T17:56:34.203135Z digest=sha256:53b4f94219c4ce7955506a8820bf168ada3ce9a1aa66a26a0ef1a5699fe8e904

Observation 3782feab-adce-4d78-806e-cc9f1647c702 · inbound

Beyond Penalizing Mistakes: Stabilizing Efficiency Training in Large Reasoning Models via Adaptive Correct-Only Rewards cites this paper.

Beyond Penalizing Mistakes: Stabilizing Efficiency Training in Large Reasoning Models via Adaptive Correct-Only Rewards Thinking Fast and Right: Balancing Accuracy and Reasoning Length with Adaptive Rewards

Reference 20

Resolution
metadata mismatch
arxiv_id, observed 2026-07-04T09:19:43.856293Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=arxiv_source observed=2026-06-26T10:12:30.692295Z digest=sha256:bc930d10473ea731a6c5bb3f6c799fcc74723f2e88a47bad1f64d0d5bc6d63c1