Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-15T22:11:34.020610Z
Paper Citation Record · LEDGER
As of 16 August 2026, this Paper Citation Record lists 36 of 36 outbound references and 7 inbound Pith citation observations for arXiv:2505.07961.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-15T22:11:34.020610Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-16T06:30:59.297886+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-08-15T19:16:03.077573Z
A source-named dated measurement, never combined with another source.
Source: arxiv_reference, observed 2026-05-20T19:43:44.121220Z
36 of 36 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation 8d396b3e-ca1c-42b2-80c0-76af3d123a3b · outbound
Making Small Language Models Efficient Reasoners: Intervention, Supervision, Reinforcement L1: Controlling How Long A Reasoning Model Thinks With Reinforcement Learning
Reference 1
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c87e41e7-106a-4d93-954e-18bc23e143de · outbound
Making Small Language Models Efficient Reasoners: Intervention, Supervision, Reinforcement The Surprising Effectiveness of Test-Time Training for Few-Shot Learning
Reference 2
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 50380b6f-a70b-4f9a-94cc-0d5ef0ec6303 · outbound
Making Small Language Models Efficient Reasoners: Intervention, Supervision, Reinforcement Precise Length Control in Large Language Models
Reference 3
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 734d0927-4c52-4467-a238-670416fd82ec · outbound
Making Small Language Models Efficient Reasoners: Intervention, Supervision, Reinforcement FrugalGPT: How to Use Large Language Models While Reducing Cost and Improving Performance
Reference 4
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e2b6c686-0811-4902-80a7-94087760d863 · outbound
Making Small Language Models Efficient Reasoners: Intervention, Supervision, Reinforcement GPTQ: Accurate Post-Training Quantization for Generative Pre-trained Transformers
Reference 5
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a60b0daa-05dc-4f83-95f0-d20e549d7e22 · outbound
Making Small Language Models Efficient Reasoners: Intervention, Supervision, Reinforcement Gemini 2.0 flash thinking mode (gemini-2.0f lash-thinking-exp-1219), 2024
Reference 6
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation 58b1928c-096f-4a7d-aa3d-e4ce852ead83 · outbound
Making Small Language Models Efficient Reasoners: Intervention, Supervision, Reinforcement Test-time training provably improves transformers as in-context learners
Reference 7
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation df98ec15-9c2c-45de-8b48-d13c6d332cd3 · outbound
Making Small Language Models Efficient Reasoners: Intervention, Supervision, Reinforcement DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning
Reference 8
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ae9df6fc-5089-4b4e-af9e-b27fbdbcda68 · outbound
Making Small Language Models Efficient Reasoners: Intervention, Supervision, Reinforcement Language Model Cascades: Token-level uncertainty and beyond
Reference 9
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 06e46da3-438e-4daa-978b-9dfe39d5fd7a · outbound
Making Small Language Models Efficient Reasoners: Intervention, Supervision, Reinforcement ThinkPrune: Pruning Long Chain-of-Thought of LLMs via Reinforcement Learning
Reference 10
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d77aac5a-79c2-46f2-ac1c-d4b8ee6c0d13 · outbound
Making Small Language Models Efficient Reasoners: Intervention, Supervision, Reinforcement OpenAI o1 System Card
Reference 11
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9dc70092-78b9-4b5b-afde-f8834b657351 · outbound
Making Small Language Models Efficient Reasoners: Intervention, Supervision, Reinforcement Fast inference from transformers via speculative decoding
Reference 12
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation 2e33c4ad-7933-458f-a900-2c504879f806 · outbound
Making Small Language Models Efficient Reasoners: Intervention, Supervision, Reinforcement Autobalance: Optimized loss functions for imbalanced data
Reference 13
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation 20a9a835-c90b-4c5e-806a-38350e17314d · outbound
Making Small Language Models Efficient Reasoners: Intervention, Supervision, Reinforcement Small models struggle to learn from strong reasoners
Reference 14
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 622e7f77-58bb-4f78-bce7-d89a22e1bf4f · outbound
Making Small Language Models Efficient Reasoners: Intervention, Supervision, Reinforcement Let’s verify step by step
Reference 15
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b1cd46e9-9114-4316-91df-8934f2b3eb7c · outbound
Making Small Language Models Efficient Reasoners: Intervention, Supervision, Reinforcement Awq: Activation-aware weight quantization for on-device llm compression and acceleration
Reference 16
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c5093dcf-123c-4933-8cfa-b9efb290d97e · outbound
Making Small Language Models Efficient Reasoners: Intervention, Supervision, Reinforcement Tang, Manan Roongta, Colin Cai, Jeffrey Luo, Tianjun Zhang, Li Erran Li, Raluca Ada Popa, and Ion Stoica
Reference 17
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation 289ba2ba-3319-49a3-bb7a-f9d0643dcc30 · outbound
Making Small Language Models Efficient Reasoners: Intervention, Supervision, Reinforcement Long-tail learning via logit adjustment
Reference 18
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 283a70c5-b00c-4a19-bb2a-62ee448c36a9 · outbound
Making Small Language Models Efficient Reasoners: Intervention, Supervision, Reinforcement s1: Simple test-time scaling
Reference 19
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b2e8ccc2-6e5a-41c1-ab96-fd265de3eab9 · outbound
Making Small Language Models Efficient Reasoners: Intervention, Supervision, Reinforcement Gpt-4o mini: advancing cost-efficient intelligence, 2024
Reference 20
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation e972ac4f-18f0-4cdd-8a9d-7a15da8a7927 · outbound
Making Small Language Models Efficient Reasoners: Intervention, Supervision, Reinforcement Gpqa: A graduate-level google-proof q&a benchmark
Reference 21
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2e45beb1-1e63-4b61-af85-f924f23096fc · outbound
Making Small Language Models Efficient Reasoners: Intervention, Supervision, Reinforcement Scaling Test-Time Compute Without Verification or RL is Suboptimal
Reference 22
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3b883b6d-825e-40e2-8a28-c751a8bc144a · outbound
Making Small Language Models Efficient Reasoners: Intervention, Supervision, Reinforcement Outrageously Large Neural Networks: The Sparsely-Gated Mixture-of-Experts Layer
Reference 23
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation bf1e8d8b-aeff-4d37-80c8-d49962fd5640 · outbound
Making Small Language Models Efficient Reasoners: Intervention, Supervision, Reinforcement Scaling LLM Test-Time Compute Optimally can be More Effective than Scaling Model Parameters
Reference 24
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 515e1ab3-531f-4550-b5f1-783b83f88ef4 · outbound
Making Small Language Models Efficient Reasoners: Intervention, Supervision, Reinforcement Scalable Chain of Thoughts via Elastic Reasoning
Reference 25
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e0a603d5-c187-43de-afcc-3bcf18c85c01 · outbound
Making Small Language Models Efficient Reasoners: Intervention, Supervision, Reinforcement Qwen2.5 Technical Report
Reference 26
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 974af263-8d69-4ade-92f7-82521c49154f · outbound
Making Small Language Models Efficient Reasoners: Intervention, Supervision, Reinforcement Qwen2.5-Math Technical Report: Toward Mathematical Expert Model via Self-Improvement
Reference 27
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c09474c1-4f25-499c-831b-50a3951f0c2a · outbound
Making Small Language Models Efficient Reasoners: Intervention, Supervision, Reinforcement Towards thinking-optimal scaling of test-time compute for llm reasoning
Reference 28
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1a7ccb58-5918-45cb-b07c-a8493b5f62cb · outbound
Making Small Language Models Efficient Reasoners: Intervention, Supervision, Reinforcement Following Length Constraints in Instructions
Reference 29
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 723c00b6-029d-4125-817b-5881c65d56c4 · outbound
Making Small Language Models Efficient Reasoners: Intervention, Supervision, Reinforcement Selective attention: Enhancing transformer through principled context control
Reference 30
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation 341dd39d-09de-457a-9e35-50c36a47f104 · outbound
Making Small Language Models Efficient Reasoners: Intervention, Supervision, Reinforcement Efficient contextual LLM cascades through budget-constrained policy learning
Reference 31
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation 95e6457c-71a2-4a0b-ab20-f7906668edc1 · outbound
Making Small Language Models Efficient Reasoners: Intervention, Supervision, Reinforcement Class-attribute priors: adapting optimization to heterogeneity and fairness objective
Reference 32
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation e4933300-2e03-433c-a78c-bcb55c7094e3 · outbound
Making Small Language Models Efficient Reasoners: Intervention, Supervision, Reinforcement Pleaseanalyze the following text and determine if there are any meaningless repetitions of identical sentences
Reference 33
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation cf0b46b3-7ea4-4360-939d-7f887e96ca61 · outbound
Making Small Language Models Efficient Reasoners: Intervention, Supervision, Reinforcement If the model prematurely generates the end-of-thinking token before reaching the desired length, the token is removed, and generation continues until the target length is reached
Reference 34
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation 9d808209-4fd3-4a47-bac1-8d66ce1414bc · outbound
Making Small Language Models Efficient Reasoners: Intervention, Supervision, Reinforcement Unresolved cited work
Reference 35
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation 7f619b3d-b0d2-4d11-b6b4-f3d34b2dbe8a · outbound
Making Small Language Models Efficient Reasoners: Intervention, Supervision, Reinforcement 2k” and “4k
Reference 36
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation a90e90b1-f664-4c52-a947-aabcf031f41f · inbound
Stop Overthinking: A Survey on Efficient Reasoning for Large Language Models Making Small Language Models Efficient Reasoners: Intervention, Supervision, Reinforcement
Reference 240
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation 6820098a-2d10-4b21-ae17-00996f34dbae · inbound
BREAD: Branched Rollouts from Expert Anchors Bridge SFT & RL for Reasoning Making Small Language Models Efficient Reasoners: Intervention, Supervision, Reinforcement
Reference 43
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation caafa4cf-1572-4ec2-90a7-2cb382e35772 · inbound
Bayesian Social Deduction with Graph-Informed Language Models Making Small Language Models Efficient Reasoners: Intervention, Supervision, Reinforcement
Reference 76
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation 8c401132-526a-4f28-975e-e06499a9d6a7 · inbound
Post Reasoning: Improving the Performance of Non-Thinking Models at No Cost Making Small Language Models Efficient Reasoners: Intervention, Supervision, Reinforcement
Reference 233
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation 6788d978-46b9-445a-9287-e00caccf6347 · inbound
VSPO: Vector-Steered Policy Optimization for Behavioral Control Making Small Language Models Efficient Reasoners: Intervention, Supervision, Reinforcement
Reference 32
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation 8f3b2d51-b066-417b-b03b-9934dd5d82f7 · inbound
OS-Pruner: Pruning Chains-of-Thought of Reasoning Models via Optimal Stopping Making Small Language Models Efficient Reasoners: Intervention, Supervision, Reinforcement
Reference 25
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6a2788d7-a87f-4422-ad06-7d008e98f891 · inbound
From Trajectories to Prefixes: Reusing Teacher Trajectories via Replayed Prefixes and Online Continuation Making Small Language Models Efficient Reasoners: Intervention, Supervision, Reinforcement
Reference 4
Source-reported events for the cited work
Unavailable: canonical work link unavailable.