Pith. sign in

Paper Citation Record · LEDGER

ConciseRL: Conciseness-Guided Reinforcement Learning for Efficient Reasoning Models

As of 20 August 2026, this Paper Citation Record lists 19 of 19 outbound references and 5 inbound Pith citation observations for arXiv:2505.17250.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2505.17250 v1

Coverage vector

measured 19 of 19 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-07T14:53:55.003012Z

measured 24 of 24 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-19T06:32:44.657259+00:00

measured 5 of 5 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-06T17:53:46.148845Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-05-14T01:29:56.674717Z

Reference resolution

19 of 19 outbound references displayed

  • verified exact0
  • verified fuzzy9
  • unresolved10
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation cfee40c1-21b0-481e-a4b0-4bba34439fe1 · outbound

This paper cites an unresolved cited work.

ConciseRL: Conciseness-Guided Reinforcement Learning for Efficient Reasoning Models Unresolved cited work

Reference 2

Resolution
unresolved
raw_fallback, observed 2026-08-07T14:53:57.694772Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-07T14:53:54.060057Z digest=sha256:40f0d3b80a1eb4796e7fa80dc18e017fac7eceb06b30e8b88bda2bc5a463fbf2

Observation 01ab6025-3b14-4b08-b130-730fa9bcf27c · outbound

This paper cites ThinkPrune: Pruning Long Chain-of-Thought of LLMs via Reinforcement Learning.

ConciseRL: Conciseness-Guided Reinforcement Learning for Efficient Reasoning Models ThinkPrune: Pruning Long Chain-of-Thought of LLMs via Reinforcement Learning

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-07T14:53:53.592950Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:53:53.592950Z digest=sha256:01621ec06162eeccab3c3383330d891cddb8e5807bdb972ee523cb75f010690e

Observation 252a0634-5e80-4cfa-a6cd-97bf4f22afc5 · outbound

This paper cites Woosuk Kwon, Zhuohan Li, Siyuan Zhuang, Ying Sheng, Lianmin Zheng, Cody Hao Yu, Joseph Gon- zalez, Hao Zhang, and Ion Stoica.

ConciseRL: Conciseness-Guided Reinforcement Learning for Efficient Reasoning Models Woosuk Kwon, Zhuohan Li, Siyuan Zhuang, Ying Sheng, Lianmin Zheng, Cody Hao Yu, Joseph Gon- zalez, Hao Zhang, and Ion Stoica

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-07T14:53:53.689958Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:53:53.689958Z digest=sha256:fb9b48714f7b11ca26bf3e4f094ee5ff3d248ffac50ec9ab0f723f4a95ccbaf8

Observation 8c3d54fa-02f2-4d5e-bf42-34806e629466 · outbound

This paper cites Imitate, Explore, and Self-Improve: A Reproduction Report on Slow-thinking Reasoning Systems.

ConciseRL: Conciseness-Guided Reinforcement Learning for Efficient Reasoning Models Imitate, Explore, and Self-Improve: A Reproduction Report on Slow-thinking Reasoning Systems

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-07T14:53:53.857132Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:53:53.857132Z digest=sha256:99d07e84c0fecac1dd27e0601229d179c3f797451af471652b4e2652a2b66674

Observation b5a47764-a058-4ec4-8052-a3e12af3bfa9 · outbound

This paper cites ConciseRL.

ConciseRL: Conciseness-Guided Reinforcement Learning for Efficient Reasoning Models ConciseRL

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:53:57.458987Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-07T14:53:54.205339Z digest=sha256:2b0876a31a4da01a325e6d6272838c9de2b1a818df9a952635f922ab7a6cb13c

Observation c5ef0484-65a6-41c6-8e8d-02944d49bae0 · outbound

This paper cites an unresolved cited work.

ConciseRL: Conciseness-Guided Reinforcement Learning for Efficient Reasoning Models Unresolved cited work

Reference 10

Resolution
unresolved
raw_fallback, observed 2026-08-07T14:53:57.282534Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-07T14:53:54.369927Z digest=sha256:f98e4b26a8e992cd184df484ae2aeea6e5d980acd18c3bc699fbf518afd826c0

Observation 20bdaeef-e216-4000-b71c-7f0103bf6dd1 · outbound

This paper cites Check \( G > B \): 26 > 9, which is true.

ConciseRL: Conciseness-Guided Reinforcement Learning for Efficient Reasoning Models Check \( G > B \): 26 > 9, which is true

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:53:57.187488Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-07T14:53:54.425961Z digest=sha256:465b7c5ae1db6c55fa83e71e0e960eddbf51c3ea15cd3e54a84d678ed46270bc

Observation a9186dec-a6d0-4fbb-91ca-e14c666e256b · outbound

This paper cites an unresolved cited work.

ConciseRL: Conciseness-Guided Reinforcement Learning for Efficient Reasoning Models Unresolved cited work

Reference 12

Resolution
unresolved
raw_fallback, observed 2026-08-07T14:53:56.941051Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-07T14:53:54.507187Z digest=sha256:a1df1e8aa6eeb00307f4c4f0d6cd126c1cab5c9634ea580c625d2ba0a9bf5380

Observation 1ce1d444-ea9c-4a2e-9192-a7ed8fb43740 · outbound

This paper cites [...] The number of boys at the meeting is \boxed{9}.

ConciseRL: Conciseness-Guided Reinforcement Learning for Efficient Reasoning Models [...] The number of boys at the meeting is \boxed{9}

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:53:56.624923Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-07T14:53:54.593658Z digest=sha256:e597f6262e769d96d802a1edb3e0fe016e6c949e55c0653af7319d9d37e6bb9d

Observation f354b64d-04c4-465c-b955-9926bc508fe6 · outbound

This paper cites ConciseRL.

ConciseRL: Conciseness-Guided Reinforcement Learning for Efficient Reasoning Models ConciseRL

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:53:56.445256Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-07T14:53:54.659295Z digest=sha256:c0f84d3a15fc5f36d2411afe8f0bb36550fe3361eac855b8def698e230b75b5e

Observation 379006c1-2b78-4445-b93d-1e8292bb6424 · outbound

This paper cites ConciseRL.

ConciseRL: Conciseness-Guided Reinforcement Learning for Efficient Reasoning Models ConciseRL

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:53:56.194770Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-07T14:53:54.698751Z digest=sha256:f7d5bf415d3e6cfc5d9ad7da96418c06a57c31f6b773e82640f382c137c03f4c

Observation ff73a4f6-48e2-46ce-8dae-8f96ea294006 · outbound

This paper cites ConciseRL.

ConciseRL: Conciseness-Guided Reinforcement Learning for Efficient Reasoning Models ConciseRL

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:53:56.028054Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-07T14:53:54.772420Z digest=sha256:e2b34a3a527d9c4da290ba6a9e1f069bff87172aa489ddead24c778381755bc2

Observation 6470cf79-7523-4904-9497-390c3011a55c · outbound

This paper cites Thus, the result of the computation is \(\boxed{\dfrac{5}{9}}\).

ConciseRL: Conciseness-Guided Reinforcement Learning for Efficient Reasoning Models Thus, the result of the computation is \(\boxed{\dfrac{5}{9}}\)

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:53:55.833060Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-07T14:53:54.844826Z digest=sha256:ff968c7ed671db15d708b18544ac02c5762d371ba4e40a4d1c8041480deed6ae

Observation e8d56cd0-ca21-46c6-b46e-4dcaebbab35b · outbound

This paper cites ConciseRL.

ConciseRL: Conciseness-Guided Reinforcement Learning for Efficient Reasoning Models ConciseRL

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:53:55.678609Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-07T14:53:54.926865Z digest=sha256:6893e3f230dfb8d3aabbf857460fcaeef5ebc54a78b8166fecad8ef4b8fa3bf8

Observation bb41f141-21cd-45da-a624-b1436980f04b · outbound

This paper cites The full reasoning traces are available at https://github.com/RazvanDu/ConciseRL.

ConciseRL: Conciseness-Guided Reinforcement Learning for Efficient Reasoning Models The full reasoning traces are available at https://github.com/RazvanDu/ConciseRL

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:53:55.546967Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-07T14:53:55.003012Z digest=sha256:8674f39d0decf134a21f2325bc52f8b86938a6704d01b2e02b5f662c6c4022ef

Observation 877a46cf-69d8-4232-87e8-fdd4b282a5b4 · outbound

This paper cites Training Verifiers to Solve Math Word Problems.

ConciseRL: Conciseness-Guided Reinforcement Learning for Efficient Reasoning Models Training Verifiers to Solve Math Word Problems

Reference 2021

Resolution
unresolved
no resolver link, observed 2026-08-07T14:53:53.407543Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:53:53.407543Z digest=sha256:0103b9ced11f193b6a83e6ada174c6e52220d7c10f747f54d7283d647235840e

Observation 3f8b6ab6-9178-46d5-b918-eb26621a3c06 · outbound

This paper cites Demystifying Long Chain-of-Thought Reasoning in LLMs.

ConciseRL: Conciseness-Guided Reinforcement Learning for Efficient Reasoning Models Demystifying Long Chain-of-Thought Reasoning in LLMs

Reference 2023

Resolution
unresolved
no resolver link, observed 2026-08-07T14:53:53.957988Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:53:53.957988Z digest=sha256:0d2ffe6b744282497c5aa43c5eaba23642acee7df83d523979faf0d2a290471f

Observation 0e811332-6d86-4d06-b2aa-270eec8efef0 · outbound

This paper cites O1-Pruner: Length-Harmonizing Fine-Tuning for O1-Like Reasoning Pruning.

ConciseRL: Conciseness-Guided Reinforcement Learning for Efficient Reasoning Models O1-Pruner: Length-Harmonizing Fine-Tuning for O1-Like Reasoning Pruning

Reference 2024

Resolution
unresolved
no resolver link, observed 2026-08-07T14:53:53.745836Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:53:53.745836Z digest=sha256:18571cfbb2d9956cd43838e02fb4c056d677200ea084f71cd6a6cd65d8439161

Observation 33e62ba4-65ba-4b43-8254-ea5e0011726d · outbound

This paper cites DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning.

ConciseRL: Conciseness-Guided Reinforcement Learning for Efficient Reasoning Models DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning

Reference 2025

Resolution
unresolved
no resolver link, observed 2026-08-07T14:53:53.478299Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:53:53.478299Z digest=sha256:68c8df113d5a18495a2801658969d86c55a2b9636fa5034f9706f7251db75ca9

Pith citing papers

Observation b77ec013-62b2-4281-94a6-722c927958dd · inbound

Stop Overthinking: A Survey on Efficient Reasoning for Large Language Models cites this paper.

Stop Overthinking: A Survey on Efficient Reasoning for Large Language Models ConciseRL: Conciseness-Guided Reinforcement Learning for Efficient Reasoning Models

Reference 40

Resolution
verified exact
arxiv_id, observed 2026-05-14T01:29:56.678764Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-05-14T01:29:56.480020Z digest=sha256:4dab4810c16bd7f7ab585de26cfe9fdda548fabf9e17c358c686338853d97849

Observation 47f83da4-5200-47fe-8120-9b2d8a75cc12 · inbound

Towards Concise and Adaptive Thinking in Large Reasoning Models: A Survey cites this paper.

Towards Concise and Adaptive Thinking in Large Reasoning Models: A Survey ConciseRL: Conciseness-Guided Reinforcement Learning for Efficient Reasoning Models

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-06T17:53:46.148845Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T17:53:46.148845Z digest=sha256:e5801f9595c4f09965374168cd6405c058a597dbc5c11f1c0155a3d2590f3ff6

Observation 9b43aca3-afee-4244-acb4-276249572b40 · inbound

Learning to Reason Efficiently with Discounted Reinforcement Learning cites this paper.

Learning to Reason Efficiently with Discounted Reinforcement Learning ConciseRL: Conciseness-Guided Reinforcement Learning for Efficient Reasoning Models

Reference 2025

Resolution
unresolved
no resolver link, observed 2026-08-04T07:59:52.455824Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T07:59:52.455824Z digest=sha256:a8c47e8ea69aed38b07a9d653ffa7a1189243799625b304fd99c16840479f0de

Observation d204015b-7dc1-438e-b5ab-c0d07dbc4a03 · inbound

Post Reasoning: Improving the Performance of Non-Thinking Models at No Cost cites this paper.

Post Reasoning: Improving the Performance of Non-Thinking Models at No Cost ConciseRL: Conciseness-Guided Reinforcement Learning for Efficient Reasoning Models

Reference 217

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T20:06:09.173076Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-05-08T10:19:08.451445Z digest=sha256:653d275a135ceed258225ebe4ec5adf6c6bfadcf22d40dc50eccf99aed7bb08d

Observation 8c70757a-9ef0-403b-b493-63747d8b6beb · inbound

Contrastive On-Policy Distillation cites this paper.

Contrastive On-Policy Distillation ConciseRL: Conciseness-Guided Reinforcement Learning for Efficient Reasoning Models

Reference 2024

Resolution
unresolved
no resolver link, observed 2026-08-01T13:40:38.524310Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T13:40:38.524310Z digest=sha256:005d3b4591354f16fd733f8ca1aac71e20bcaee295a1f481875717ce8421f486