Pith. sign in

Paper Citation Record · LEDGER

EvoThink: Evolving Thinking in Large Reasoning Models via Self-Pruning and Aha-Moment Preference Optimization

As of 9 August 2026, this Paper Citation Record lists 22 of 22 outbound references and 0 inbound Pith citation observations for arXiv:2607.19962.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2607.19962 v1

Coverage vector

measured 22 of 22 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-01T11:17:40.289683Z

measured 22 of 22 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-09T06:31:02.800959+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

22 of 22 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved22
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation cb855382-5745-4d62-98f3-66af94cc6b88 · outbound

This paper cites Sketch-of-thought: Efficient LLM rea- soning with adaptive cognitive-inspired sketching.

EvoThink: Evolving Thinking in Large Reasoning Models via Self-Pruning and Aha-Moment Preference Optimization Sketch-of-thought: Efficient LLM rea- soning with adaptive cognitive-inspired sketching

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-01T11:17:40.195703Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T11:17:40.195703Z digest=sha256:bc09bf05a67cb87d7781f6062aa225f0be2381d1d8aa431814d89919dd6522f6

Observation 5b0eb710-8d34-4a72-aadd-14e25293efe4 · outbound

This paper cites Optimizing Length Compression in Large Reasoning Models.

EvoThink: Evolving Thinking in Large Reasoning Models via Self-Pruning and Aha-Moment Preference Optimization Optimizing Length Compression in Large Reasoning Models

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-01T11:17:40.210209Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T11:17:40.210209Z digest=sha256:5d3cd97865c306a7896ebdf7e3f333773255b606e1f53f872ef35cda6411fd28

Observation d1a0ecba-065e-4ee2-9176-c567615d9f1c · outbound

This paper cites The Danger of Overthinking: Examining the Reasoning-Action Dilemma in Agentic Tasks.

EvoThink: Evolving Thinking in Large Reasoning Models via Self-Pruning and Aha-Moment Preference Optimization The Danger of Overthinking: Examining the Reasoning-Action Dilemma in Agentic Tasks

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-01T11:17:40.215084Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T11:17:40.215084Z digest=sha256:be136dba0fd62812d47da7e15a7e84ef3ac438747ca4697c95d5c38077214b2e

Observation f07cc2cf-ff5d-46e0-af9f-dc1de20e1401 · outbound

This paper cites DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning.

EvoThink: Evolving Thinking in Large Reasoning Models via Self-Pruning and Aha-Moment Preference Optimization DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-01T11:17:40.220113Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T11:17:40.220113Z digest=sha256:fc7138be6d6e91c0ed825de6ab3c92f957f3332349fe602ca905f86ef6c7cfba

Observation 85cb031e-917b-4801-a3ca-3083afc79230 · outbound

This paper cites Token- budget-aware llm reasoning.

EvoThink: Evolving Thinking in Large Reasoning Models via Self-Pruning and Aha-Moment Preference Optimization Token- budget-aware llm reasoning

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-01T11:17:40.224760Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T11:17:40.224760Z digest=sha256:a4872a5ecc7c936119573f8b41371442cd709b4641e903fdd8d547584ba6c216

Observation 40955301-8fd3-4990-85b5-e3de72f17479 · outbound

This paper cites Measuring mathemat- ical problem solving with the math dataset.

EvoThink: Evolving Thinking in Large Reasoning Models via Self-Pruning and Aha-Moment Preference Optimization Measuring mathemat- ical problem solving with the math dataset

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-01T11:17:40.229399Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T11:17:40.229399Z digest=sha256:5d86ca21e16966aa69376bec5efac3f48ea8f809fa2ce3a51396e3d16d4d19c0

Observation 6b2bb410-a8d6-4927-bb03-a392866cb7eb · outbound

This paper cites TACO: Topics in Algorithmic COde generation dataset.

EvoThink: Evolving Thinking in Large Reasoning Models via Self-Pruning and Aha-Moment Preference Optimization TACO: Topics in Algorithmic COde generation dataset

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-01T11:17:40.237963Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T11:17:40.237963Z digest=sha256:8c76a8537ee06f75a30f5973416f2b3d6ecc51ee9e248e49958b1b7a65825b7b

Observation cda6ec71-ed17-418c-86d5-1055d6e3b994 · outbound

This paper cites O1-Pruner: Length-Harmonizing Fine-Tuning for O1-Like Reasoning Pruning.

EvoThink: Evolving Thinking in Large Reasoning Models via Self-Pruning and Aha-Moment Preference Optimization O1-Pruner: Length-Harmonizing Fine-Tuning for O1-Like Reasoning Pruning

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-01T11:17:40.247507Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T11:17:40.247507Z digest=sha256:f1ff5340435a13a480aabab115b9cb6eeede3c970724d2b7d4455e62b30f204f

Observation 7fdbc693-ab41-4104-9c81-3107b4d03b6b · outbound

This paper cites Reasoning Models Can Be Effective Without Thinking.

EvoThink: Evolving Thinking in Large Reasoning Models via Self-Pruning and Aha-Moment Preference Optimization Reasoning Models Can Be Effective Without Thinking

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-01T11:17:40.252219Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T11:17:40.252219Z digest=sha256:694b7ea65aa64f4d904a1dd2ce167597af9520636ffb7ee66cb7ccc1a1ef5f19

Observation 8d2a359f-7229-4e3d-86bb-070bc164659f · outbound

This paper cites s1: Simple test-time scaling.

EvoThink: Evolving Thinking in Large Reasoning Models via Self-Pruning and Aha-Moment Preference Optimization s1: Simple test-time scaling

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-01T11:17:40.256620Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T11:17:40.256620Z digest=sha256:e8f83c0d4f93b0218300bd9e17339ecb3bb5e588526474d8891bf9f275d52b98

Observation 3e1b069d-8157-49fc-b506-eadc9e2da592 · outbound

This paper cites DynaThink: Fast or slow? a dynamic decision-making framework for large language models.

EvoThink: Evolving Thinking in Large Reasoning Models via Self-Pruning and Aha-Moment Preference Optimization DynaThink: Fast or slow? a dynamic decision-making framework for large language models

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-01T11:17:40.260565Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T11:17:40.260565Z digest=sha256:22ef911b9b9a25ff39d4b66417932b912274f7f38f960d32bd717ee7d95572e4

Observation 302028a2-a39b-41e6-9985-742e17983424 · outbound

This paper cites Direct preference optimization: Your lan- guage model is secretly a reward model.Advances in neural information processing systems, 36:53728–53741,.

EvoThink: Evolving Thinking in Large Reasoning Models via Self-Pruning and Aha-Moment Preference Optimization Direct preference optimization: Your lan- guage model is secretly a reward model.Advances in neural information processing systems, 36:53728–53741,

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-01T11:17:40.264612Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T11:17:40.264612Z digest=sha256:60dfd988580551a247a9227159c7b07413b4f30313d64be67bdef1b274f08878

Observation 23dbda9e-4e0d-4b86-b021-aec00b4b198b · outbound

This paper cites DAST: Difficulty-adaptive slow-thinking for large reasoning mod- els.

EvoThink: Evolving Thinking in Large Reasoning Models via Self-Pruning and Aha-Moment Preference Optimization DAST: Difficulty-adaptive slow-thinking for large reasoning mod- els

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-01T11:17:40.268664Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T11:17:40.268664Z digest=sha256:cb0e2705b6d5aecad24da73f94f551c45fa4869379419bb363e771b9e9bd4305

Observation 30e244d6-2465-44ff-871a-4d845bddc279 · outbound

This paper cites Hybridflow: A flexible and efficient rlhf framework.

EvoThink: Evolving Thinking in Large Reasoning Models via Self-Pruning and Aha-Moment Preference Optimization Hybridflow: A flexible and efficient rlhf framework

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-01T11:17:40.272856Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T11:17:40.272856Z digest=sha256:a660a9ff442e5f3740a3f67f5b00e4adc4a8d310a5cc2e66ed36bcdd6073ccaa

Observation e7a233f3-2f83-4bfc-bda6-5be9a1c1717c · outbound

This paper cites Dualformer: Controllable fast and slow thinking by learning with ran- domized reasoning traces.

EvoThink: Evolving Thinking in Large Reasoning Models via Self-Pruning and Aha-Moment Preference Optimization Dualformer: Controllable fast and slow thinking by learning with ran- domized reasoning traces

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-01T11:17:40.276719Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T11:17:40.276719Z digest=sha256:dbb2e290a8bdf48ea7c004bd288ef2e952638c1884e1bc76e18c41e30c5e87cc

Observation c3ba12e7-a5dc-432a-a7e5-b3dc1b750f5b · outbound

This paper cites Stop Overthinking: A Survey on Efficient Reasoning for Large Language Models.

EvoThink: Evolving Thinking in Large Reasoning Models via Self-Pruning and Aha-Moment Preference Optimization Stop Overthinking: A Survey on Efficient Reasoning for Large Language Models

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-01T11:17:40.280865Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T11:17:40.280865Z digest=sha256:704ea14121feeb91fb1720bcbbf9485b485ca1f6fd62d8e1baf75721e725d3b9

Observation 810116fd-8204-4e41-ae84-f2225e54129a · outbound

This paper cites Kimi k1.5: Scaling Reinforcement Learning with LLMs.

EvoThink: Evolving Thinking in Large Reasoning Models via Self-Pruning and Aha-Moment Preference Optimization Kimi k1.5: Scaling Reinforcement Learning with LLMs

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-01T11:17:40.285036Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T11:17:40.285036Z digest=sha256:0aa6704a2a97abe61a52d5e8c1a0e6b70f78b340b70bacbc8fa666c2ec92343b

Observation c3aaf4ff-c0b0-4533-bbfe-1c9a075eb16c · outbound

This paper cites Qwen2.5 Technical Report.

EvoThink: Evolving Thinking in Large Reasoning Models via Self-Pruning and Aha-Moment Preference Optimization Qwen2.5 Technical Report

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-01T11:17:40.289683Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T11:17:40.289683Z digest=sha256:05e8da2a3c18a49bd6dc38f980a82c892e50c0d2378f158092b3c55ad7b35ccf

Observation 0d0c6014-004c-4ee5-af98-36920e3e0d16 · outbound

This paper cites ThinkPrune: Pruning Long Chain-of-Thought of LLMs via Reinforcement Learning.

EvoThink: Evolving Thinking in Large Reasoning Models via Self-Pruning and Aha-Moment Preference Optimization ThinkPrune: Pruning Long Chain-of-Thought of LLMs via Reinforcement Learning

Reference 2021

Resolution
unresolved
no resolver link, observed 2026-08-01T11:17:40.233585Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T11:17:40.233585Z digest=sha256:d52d6f80a10f13750fbdf69767974fa6572c3236ef763d9fe8f652bd5503408d

Observation 41d58f43-6f89-4885-aa6a-76213e83e346 · outbound

This paper cites ThinkSwitcher: When to think hard, when to think fast.

EvoThink: Evolving Thinking in Large Reasoning Models via Self-Pruning and Aha-Moment Preference Optimization ThinkSwitcher: When to think hard, when to think fast

Reference 2023

Resolution
unresolved
no resolver link, observed 2026-08-01T11:17:40.243121Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T11:17:40.243121Z digest=sha256:0bdef7624fd39fc1dafccb27048f87bdce0716a335bf0b972ad24934210e6214

Observation 4746e66a-7f39-47a5-98a7-f34fbf85603e · outbound

This paper cites The overthinker’s DIET: Cutting token calories with DIfficulty-aware training.

EvoThink: Evolving Thinking in Large Reasoning Models via Self-Pruning and Aha-Moment Preference Optimization The overthinker’s DIET: Cutting token calories with DIfficulty-aware training

Reference 2024

Resolution
unresolved
no resolver link, observed 2026-08-01T11:17:40.205564Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T11:17:40.205564Z digest=sha256:e681f988799582e102025900de8560058714b06a7f68a9e9421a44742c42a795

Observation 4682635f-18ea-4e05-af52-36b22b0d716f · outbound

This paper cites Do NOT Think That Much for 2+3=? On the Overthinking of o1-Like LLMs.

EvoThink: Evolving Thinking in Large Reasoning Models via Self-Pruning and Aha-Moment Preference Optimization Do NOT Think That Much for 2+3=? On the Overthinking of o1-Like LLMs

Reference 2025

Resolution
unresolved
no resolver link, observed 2026-08-01T11:17:40.200772Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T11:17:40.200772Z digest=sha256:a8032929ca872c06c7be5959c991459166e4892acce14e59f5a4988d6fab075a

Pith citing papers

No inbound Pith citation observations are available.