Pith. sign in

Paper Citation Record · LEDGER

Thinking Fast and Right: Balancing Accuracy and Reasoning Length with Adaptive Rewards

As of 20 August 2026, this Paper Citation Record lists 35 of 35 outbound references and 8 inbound Pith citation observations for arXiv:2505.18298.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2505.18298 v1

Coverage vector

measured 35 of 35 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-07T14:37:03.196744Z

measured 43 of 43 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-20T06:33:59.587034+00:00

measured 8 of 8 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-06T21:45:10.270039Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-04T09:19:43.854814Z

Reference resolution

35 of 35 outbound references displayed

  • verified exact0
  • verified fuzzy4
  • unresolved31
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 442ab1f7-9fc6-484f-bc92-061aefa9a772 · outbound

This paper cites L1: Controlling How Long A Reasoning Model Thinks With Reinforcement Learning.

Thinking Fast and Right: Balancing Accuracy and Reasoning Length with Adaptive Rewards L1: Controlling How Long A Reasoning Model Thinks With Reinforcement Learning

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-07T14:37:00.874264Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:37:00.874264Z digest=sha256:b00ef1d56f297432528911fa8cc77d987a14c1bbc281a34be7475c1e7a912c50

Observation 8f53ce4f-0aef-4a8a-8e9f-3731ce5fe379 · outbound

This paper cites Training language models to reason efficiently.arXiv preprint arXiv:2502.04463, 2025.

Thinking Fast and Right: Balancing Accuracy and Reasoning Length with Adaptive Rewards Training language models to reason efficiently.arXiv preprint arXiv:2502.04463, 2025

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-07T14:37:01.003276Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:37:01.003276Z digest=sha256:1ad2fa300208db12b48e5bb069a8155bee4276ab6733d105a318ecb0f5eac5ae

Observation 72a9f013-a6a3-47ba-a1df-020837f1e845 · outbound

This paper cites Do NOT Think That Much for 2+3=? On the Overthinking of o1-Like LLMs.

Thinking Fast and Right: Balancing Accuracy and Reasoning Length with Adaptive Rewards Do NOT Think That Much for 2+3=? On the Overthinking of o1-Like LLMs

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-07T14:37:01.107838Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:37:01.107838Z digest=sha256:d86c139afd74e5893b7a51fc06893c4eae1b91d016ce582b365c824734517cd4

Observation b3663651-1041-422a-858d-3edb87b86a86 · outbound

This paper cites Training Verifiers to Solve Math Word Problems.

Thinking Fast and Right: Balancing Accuracy and Reasoning Length with Adaptive Rewards Training Verifiers to Solve Math Word Problems

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-07T14:37:01.174456Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:37:01.174456Z digest=sha256:e69c48f405b766f21ceaeb31dd23a229e87a965873f6e6544d8bd62b2e4baf66

Observation 51308257-6ebd-4863-a821-4f87adf34c01 · outbound

This paper cites competitive programming platform.https://codeforces.com/, 2025.

Thinking Fast and Right: Balancing Accuracy and Reasoning Length with Adaptive Rewards competitive programming platform.https://codeforces.com/, 2025

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:37:04.503936Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-07T14:37:01.215547Z digest=sha256:2a4ea2e39c8f60d561639e088d4b0b3b4c8ba9fbfa902889c1e3be723df971dd

Observation c970b6b3-d015-41fc-ba95-5c2ab9399b22 · outbound

This paper cites Efficient reasoning models: A survey.arXiv preprint arXiv:2504.10903, 2025.

Thinking Fast and Right: Balancing Accuracy and Reasoning Length with Adaptive Rewards Efficient reasoning models: A survey.arXiv preprint arXiv:2504.10903, 2025

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-07T14:37:01.267450Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:37:01.267450Z digest=sha256:29f36d801e2929384bf5972f15897bce35cc6891718122bcc57ecaa28f913700

Observation 5ee06ac6-e9c8-4669-a56a-d1e83aa8f32b · outbound

This paper cites Omni-MATH: A Universal Olympiad Level Mathematic Benchmark For Large Language Models.

Thinking Fast and Right: Balancing Accuracy and Reasoning Length with Adaptive Rewards Omni-MATH: A Universal Olympiad Level Mathematic Benchmark For Large Language Models

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-07T14:37:01.324509Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:37:01.324509Z digest=sha256:1b9f9714f2522768e8e5d86a32c0be0ba17e93bc9e26171615d613c99a73f5df

Observation 38674f00-a0af-40af-b4f9-bd9e328a6162 · outbound

This paper cites Scaling laws for reward model overoptimization.

Thinking Fast and Right: Balancing Accuracy and Reasoning Length with Adaptive Rewards Scaling laws for reward model overoptimization

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-07T14:37:01.397009Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:37:01.397009Z digest=sha256:33e7291e639338f4d3c1503224c5771d3453cb378b3d5e385988a370fcc8496a

Observation 8832b3e4-85ff-4a45-9030-0a56a9b4c4ce · outbound

This paper cites DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning.

Thinking Fast and Right: Balancing Accuracy and Reasoning Length with Adaptive Rewards DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-07T14:37:01.478805Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:37:01.478805Z digest=sha256:6119ca1365d78a69e9015de72285caaa6bd543817581361d41fba354534175d5

Observation b90ca9c3-e075-477f-9356-016a9a549240 · outbound

This paper cites Token-Budget-Aware LLM Reasoning.

Thinking Fast and Right: Balancing Accuracy and Reasoning Length with Adaptive Rewards Token-Budget-Aware LLM Reasoning

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-07T14:37:01.537382Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:37:01.537382Z digest=sha256:6bed9fe7078d80990773a03079ce93f73b6a8fb3e60e54c4b062081ffe58770b

Observation be7cfda0-09a0-4455-874b-d240ddc9084d · outbound

This paper cites Olympiadbench: A challenging benchmark for promoting agi with olympiad-level bilingual multimodal scientific problems.

Thinking Fast and Right: Balancing Accuracy and Reasoning Length with Adaptive Rewards Olympiadbench: A challenging benchmark for promoting agi with olympiad-level bilingual multimodal scientific problems

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-07T14:37:01.603480Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:37:01.603480Z digest=sha256:c3fcef494c3bb19e35036b6c678edc6b530c987f68f99ad68967b22922561da3

Observation dcef9361-7a71-4c42-88a1-def13466c23d · outbound

This paper cites Measuring Mathematical Problem Solving With the MATH Dataset.

Thinking Fast and Right: Balancing Accuracy and Reasoning Length with Adaptive Rewards Measuring Mathematical Problem Solving With the MATH Dataset

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-07T14:37:01.676325Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:37:01.676325Z digest=sha256:37f13ecdb47503702856a19498943d55acc579bc3c2c6260a86722acba23908c

Observation 664a0074-5786-45a7-a606-0116e12b066e · outbound

This paper cites ThinkPrune: Pruning Long Chain-of-Thought of LLMs via Reinforcement Learning.

Thinking Fast and Right: Balancing Accuracy and Reasoning Length with Adaptive Rewards ThinkPrune: Pruning Long Chain-of-Thought of LLMs via Reinforcement Learning

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-07T14:37:01.733717Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:37:01.733717Z digest=sha256:83cf6877e6e5a2bfbf5566ab1395ceca667d2dda61f62f23c5b27324e1c4a9aa

Observation 5856107e-8d78-40cd-af8d-bca9abfc7d6f · outbound

This paper cites C3ot: Generating shorter chain-of- thought without compromising effectiveness.

Thinking Fast and Right: Balancing Accuracy and Reasoning Length with Adaptive Rewards C3ot: Generating shorter chain-of- thought without compromising effectiveness

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-07T14:37:01.791923Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:37:01.791923Z digest=sha256:d2713df5ce032d475b923f4b2ff2bc2dc49a923d3ef96a5f4549b10f9492b089

Observation 7ab5e3d9-f52c-45d2-b9af-3ab106d5cec7 · outbound

This paper cites Training Language Models to Self-Correct via Reinforcement Learning.

Thinking Fast and Right: Balancing Accuracy and Reasoning Length with Adaptive Rewards Training Language Models to Self-Correct via Reinforcement Learning

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-07T14:37:01.878223Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:37:01.878223Z digest=sha256:cd1374df800524e507a3e241456d6bbc2f92303940002fd8457c6744e2398252

Observation 31e03873-e1f0-4215-94d9-6e52937bf939 · outbound

This paper cites Let’s verify step by step.

Thinking Fast and Right: Balancing Accuracy and Reasoning Length with Adaptive Rewards Let’s verify step by step

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-07T14:37:01.939588Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:37:01.939588Z digest=sha256:73d21cce192c32c8c712984ce02c11ec0e3141930ccff621ef2abc5284a797eb

Observation be441b8d-1204-4559-8db8-fb50648a4f2a · outbound

This paper cites O1-Pruner: Length-Harmonizing Fine-Tuning for O1-Like Reasoning Pruning.

Thinking Fast and Right: Balancing Accuracy and Reasoning Length with Adaptive Rewards O1-Pruner: Length-Harmonizing Fine-Tuning for O1-Like Reasoning Pruning

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-07T14:37:01.996062Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:37:01.996062Z digest=sha256:c613738166fa8a2f075599544f9b848056c76b483f636bae64f92a89e2410115

Observation 1076a454-ec4d-4c8c-b4cb-24c664984679 · outbound

This paper cites Tang, Manan Roongta, Colin Cai, Jeffrey Luo, Tianjun Zhang, Li Erran Li, Raluca Ada Popa, and Ion Stoica.

Thinking Fast and Right: Balancing Accuracy and Reasoning Length with Adaptive Rewards Tang, Manan Roongta, Colin Cai, Jeffrey Luo, Tianjun Zhang, Li Erran Li, Raluca Ada Popa, and Ion Stoica

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:37:04.310465Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-07T14:37:02.066745Z digest=sha256:3f6e4cffc6798f6fb2ebf3086f661e0346c68f761812aa33c7f245de4ba291bd

Observation 505ccb87-69d6-4fee-9cd0-1cc0f6d00f67 · outbound

This paper cites Imitate, Explore, and Self-Improve: A Reproduction Report on Slow-thinking Reasoning Systems.

Thinking Fast and Right: Balancing Accuracy and Reasoning Length with Adaptive Rewards Imitate, Explore, and Self-Improve: A Reproduction Report on Slow-thinking Reasoning Systems

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-07T14:37:02.150424Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:37:02.150424Z digest=sha256:1b3a206ed501b611d35204947bc81a858f8feaf09f9c593c32976501c884ee5c

Observation 8e1c8307-a8e8-4f91-836a-98a20af4269f · outbound

This paper cites Self-Training Elicits Concise Reasoning in Large Language Models.

Thinking Fast and Right: Balancing Accuracy and Reasoning Length with Adaptive Rewards Self-Training Elicits Concise Reasoning in Large Language Models

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-07T14:37:02.225034Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:37:02.225034Z digest=sha256:d1a62983d957ef30991422fa103bb626437a9af04d20ea9bcad8e8e41bebe023

Observation abf27e81-4979-46ec-bfe9-c90a609f7654 · outbound

This paper cites Concise Thoughts: Impact of Output Length on LLM Reasoning and Cost.

Thinking Fast and Right: Balancing Accuracy and Reasoning Length with Adaptive Rewards Concise Thoughts: Impact of Output Length on LLM Reasoning and Cost

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-07T14:37:02.280834Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:37:02.280834Z digest=sha256:6c75b73fc0aa09f9028ad36b54984f016f97f7c199e8a0a3c18d0ddef6660596

Observation 6b90efee-0b7f-4959-addf-336cbe67c12a · outbound

This paper cites Learning to reason with llms.https://openai.com/index/learning-to-reason-with-llms/, 2024.

Thinking Fast and Right: Balancing Accuracy and Reasoning Length with Adaptive Rewards Learning to reason with llms.https://openai.com/index/learning-to-reason-with-llms/, 2024

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:37:04.201016Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-07T14:37:02.360443Z digest=sha256:90aa0741d6d72ef8202c86c5b9479e02c8ddf60a41a2d51b09758ab3e5e59a98

Observation 60242fa7-0d39-445e-a632-bc67c3386528 · outbound

This paper cites Gpqa: A graduate-level google-proof q&a benchmark.

Thinking Fast and Right: Balancing Accuracy and Reasoning Length with Adaptive Rewards Gpqa: A graduate-level google-proof q&a benchmark

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-07T14:37:02.400867Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:37:02.400867Z digest=sha256:028f6cbc89071a5de36719952f82168b7b34cd4277f22ba17da7145cc00f6bca

Observation 08aef909-c430-4afe-9729-9393712e7ad7 · outbound

This paper cites Code Llama: Open Foundation Models for Code.

Thinking Fast and Right: Balancing Accuracy and Reasoning Length with Adaptive Rewards Code Llama: Open Foundation Models for Code

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-07T14:37:02.480031Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:37:02.480031Z digest=sha256:aa19d826be705e3183044e4c4d46ff47c4cb46c89b435f64d9bbb7f3c3657de2

Observation 158a1cca-3fa1-41c1-9c31-f17173af4c38 · outbound

This paper cites DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models.

Thinking Fast and Right: Balancing Accuracy and Reasoning Length with Adaptive Rewards DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-07T14:37:02.554391Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:37:02.554391Z digest=sha256:703f5aefcbac5357135fa8e5401a6cf17df46ee6407ae96d180d95b63169e923

Observation 0f4b5653-1520-4686-9390-b95c16d4e23b · outbound

This paper cites Between Underthinking and Overthinking: An Empirical Study of Reasoning Length and correctness in LLMs.

Thinking Fast and Right: Balancing Accuracy and Reasoning Length with Adaptive Rewards Between Underthinking and Overthinking: An Empirical Study of Reasoning Length and correctness in LLMs

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-07T14:37:02.625061Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:37:02.625061Z digest=sha256:c71c2b8c5c519ebf28768289a9b3430b48d19772e90e3c7dcb97d355956804b6

Observation 8c121087-817b-458e-85a1-5ec109884a4b · outbound

This paper cites Stop Overthinking: A Survey on Efficient Reasoning for Large Language Models.

Thinking Fast and Right: Balancing Accuracy and Reasoning Length with Adaptive Rewards Stop Overthinking: A Survey on Efficient Reasoning for Large Language Models

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-07T14:37:02.693985Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:37:02.693985Z digest=sha256:8d8b094bed85e0a853cc8cb067624e6a49f4841c4a0232f2e566be704e235726

Observation f35abcbc-a4b1-4f97-bd70-f2bc9589c7f4 · outbound

This paper cites Qwq-32b-preview.https://qwenlm.github.io/blog/qwq-32b-preview/, 2025.

Thinking Fast and Right: Balancing Accuracy and Reasoning Length with Adaptive Rewards Qwq-32b-preview.https://qwenlm.github.io/blog/qwq-32b-preview/, 2025

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:37:04.090914Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-07T14:37:02.767940Z digest=sha256:3d7a4f062d4798a3705df96f8e3349a3b53f5d1d91bb22ba7557d20f6d604d10

Observation 8ee41af2-69d6-45d3-bbe2-f7952706419e · outbound

This paper cites Harnessing the Reasoning Economy: A Survey of Efficient Reasoning for Large Language Models.

Thinking Fast and Right: Balancing Accuracy and Reasoning Length with Adaptive Rewards Harnessing the Reasoning Economy: A Survey of Efficient Reasoning for Large Language Models

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-07T14:37:02.822409Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:37:02.822409Z digest=sha256:bfdad87b6ed855bc2670b78a906bd09c73fbb5d6d5f8a09e193dd43406bd79b5

Observation 91d8a4e6-f946-45b6-a219-4417b9846b2b · outbound

This paper cites Self-Consistency Improves Chain of Thought Reasoning in Language Models.

Thinking Fast and Right: Balancing Accuracy and Reasoning Length with Adaptive Rewards Self-Consistency Improves Chain of Thought Reasoning in Language Models

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-07T14:37:02.885543Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:37:02.885543Z digest=sha256:fb92b17192882531e5c31de1e01b5e6d385fd974686b85fc8a5b07deca3c5481

Observation 9d748111-dfab-45c2-b6bb-fd5012ffa94b · outbound

This paper cites Chain-of-thought prompting elicits reasoning in large language models.

Thinking Fast and Right: Balancing Accuracy and Reasoning Length with Adaptive Rewards Chain-of-thought prompting elicits reasoning in large language models

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-07T14:37:02.944987Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:37:02.944987Z digest=sha256:27b34c67a6c6694c9a4e45916ed77ad3b85b8302c7b38aae3859ffba9a38ac16

Observation b3f68f8b-fa53-436f-93dc-151615958a5f · outbound

This paper cites Unlocking Efficient Long-to-Short LLM Reasoning with Model Merging.

Thinking Fast and Right: Balancing Accuracy and Reasoning Length with Adaptive Rewards Unlocking Efficient Long-to-Short LLM Reasoning with Model Merging

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-07T14:37:03.004066Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:37:03.004066Z digest=sha256:59cd7a30ad226186c6cf6d95c8333af29931437d14c4758e3e96114d2f03b958

Observation 1bd1b620-86a1-489e-bec5-64558b313831 · outbound

This paper cites Tokenskip: Controllable chain-of-thought compression in llms.arXiv preprint arXiv:2502.12067, 2025.

Thinking Fast and Right: Balancing Accuracy and Reasoning Length with Adaptive Rewards Tokenskip: Controllable chain-of-thought compression in llms.arXiv preprint arXiv:2502.12067, 2025

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-07T14:37:03.060411Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:37:03.060411Z digest=sha256:43e455757487fb81b597acc26e5cb670af8dd3289310aebb28d61b24ec5407da

Observation 56280a21-99c1-43fe-94e3-d410c4ac50c2 · outbound

This paper cites Towards thinking-optimal scaling of test-time compute for llm reasoning.arXiv preprint arXiv:2502.18080, 2025.

Thinking Fast and Right: Balancing Accuracy and Reasoning Length with Adaptive Rewards Towards thinking-optimal scaling of test-time compute for llm reasoning.arXiv preprint arXiv:2502.18080, 2025

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-07T14:37:03.121217Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:37:03.121217Z digest=sha256:7be8b65dc3df9e4f4d4f610ed0df6ad91b94c3a76820eed834e4362d4d22fc12

Observation 68153ac2-28be-4582-9590-89a703dc1273 · outbound

This paper cites Shorterbetter: Guiding reasoning models to find optimal inference length for efficient reasoning.arXiv preprint arXiv:2504.21370, 2025.

Thinking Fast and Right: Balancing Accuracy and Reasoning Length with Adaptive Rewards Shorterbetter: Guiding reasoning models to find optimal inference length for efficient reasoning.arXiv preprint arXiv:2504.21370, 2025

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-07T14:37:03.196744Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:37:03.196744Z digest=sha256:8052aaf362637c8b168eff1fb2c5309bc5ad19dac2691f56772f3a4b4bfce6ab

Pith citing papers

Observation d52c04d0-32e9-4730-b36a-609bd4a0a2d3 · inbound

Do Thinking Tokens Help or Trap? Towards More Efficient Large Reasoning Model cites this paper.

Do Thinking Tokens Help or Trap? Towards More Efficient Large Reasoning Model Thinking Fast and Right: Balancing Accuracy and Reasoning Length with Adaptive Rewards

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-06T21:45:10.270039Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T21:45:10.270039Z digest=sha256:62aa7536765c529e7690c7a487d04dd3b7640cf5637e5967de6d0635e3c88b1a

Observation 02d91076-4512-4bc6-ad22-2425111373bc · inbound

Towards Concise and Adaptive Thinking in Large Reasoning Models: A Survey cites this paper.

Towards Concise and Adaptive Thinking in Large Reasoning Models: A Survey Thinking Fast and Right: Balancing Accuracy and Reasoning Length with Adaptive Rewards

Reference 168

Resolution
unresolved
no resolver link, observed 2026-08-06T17:54:17.672973Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T17:54:17.672973Z digest=sha256:7a5faccc227cddbbc2aebf70aa090b1ea1d036b7a178fbf7fd9873c9abe5e678

Observation e800aca3-effb-423a-8206-db125821c2c5 · inbound

Learning to Reason Efficiently with Discounted Reinforcement Learning cites this paper.

Learning to Reason Efficiently with Discounted Reinforcement Learning Thinking Fast and Right: Balancing Accuracy and Reasoning Length with Adaptive Rewards

Reference 2010

Resolution
unresolved
no resolver link, observed 2026-08-04T07:59:53.678805Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T07:59:53.678805Z digest=sha256:91720c63f5b3ef1531a88fcda1c6a8fb5632ef0ef7a367ac32bbafd8a3654daf

Observation 56fffa5f-907e-49f8-a5be-9506a292801d · inbound

GDPO: Group reward-Decoupled Normalization Policy Optimization for Multi-reward RL Optimization cites this paper.

GDPO: Group reward-Decoupled Normalization Policy Optimization for Multi-reward RL Optimization Thinking Fast and Right: Balancing Accuracy and Reasoning Length with Adaptive Rewards

Reference 45

Resolution
verified exact
arxiv_id, observed 2026-05-15T05:31:56.036833Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-15T05:31:55.864438Z digest=sha256:15c48069e61c73a7eed472226fd4dca3372fc5901f22060ab4835207935d8ba7

Observation 4369648a-efbc-44ca-b123-55488db6a1dc · inbound

Implicit Compression Regularization: Concise Reasoning via Internal Shorter Distributions in RL Post-Training cites this paper.

Implicit Compression Regularization: Concise Reasoning via Internal Shorter Distributions in RL Post-Training Thinking Fast and Right: Balancing Accuracy and Reasoning Length with Adaptive Rewards

Reference 25

Resolution
verified exact
arxiv_id, observed 2026-05-11T04:35:59.701953Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-11T01:12:20.864362Z digest=sha256:dcf9e6cbb87b39051d46e7207e1c40eec42112188930caaae175370f4d26903f

Observation f1e542d4-bdda-4a54-853a-dcf3a992cc5b · inbound

CLORE: Content-Level Optimization for Reasoning Efficiency cites this paper.

CLORE: Content-Level Optimization for Reasoning Efficiency Thinking Fast and Right: Balancing Accuracy and Reasoning Length with Adaptive Rewards

Reference 42

Resolution
verified exact
arxiv_id, observed 2026-05-22T05:51:08.229808Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-22T05:50:23.111591Z digest=sha256:425592da9ecf4427b6466ec70248176669e96bc77ea317e782ff6ba55315b120

Observation bb273491-ab63-42ba-8e1b-99c1e0dfdacf · inbound

CARE: Competence-Aware Reward Shaping for Adaptive Reasoning Length in Video-MLLMs cites this paper.

CARE: Competence-Aware Reward Shaping for Adaptive Reasoning Length in Video-MLLMs Thinking Fast and Right: Balancing Accuracy and Reasoning Length with Adaptive Rewards

Reference 6

Resolution
verified exact
arxiv_id, observed 2026-07-04T03:29:31.439181Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-06-26T17:56:34.203135Z digest=sha256:e29bd4a17a25dc63fa53d11d661445a28d97ea3755f1bdc2aec46a7b53640f7e

Observation 3782feab-adce-4d78-806e-cc9f1647c702 · inbound

Beyond Penalizing Mistakes: Stabilizing Efficiency Training in Large Reasoning Models via Adaptive Correct-Only Rewards cites this paper.

Beyond Penalizing Mistakes: Stabilizing Efficiency Training in Large Reasoning Models via Adaptive Correct-Only Rewards Thinking Fast and Right: Balancing Accuracy and Reasoning Length with Adaptive Rewards

Reference 20

Resolution
metadata mismatch
arxiv_id, observed 2026-07-04T09:19:43.856293Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-06-26T10:12:30.692295Z digest=sha256:e93952c598951c10f53e7d2b95cf10e7d829462a44aee9b50e607e71c3855d36