Pith. sign in

Paper Citation Record · LEDGER

Making Small Language Models Efficient Reasoners: Intervention, Supervision, Reinforcement

As of 16 August 2026, this Paper Citation Record lists 36 of 36 outbound references and 7 inbound Pith citation observations for arXiv:2505.07961.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2505.07961 v3

Coverage vector

measured 36 of 36 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-15T22:11:34.020610Z

measured 43 of 43 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-16T06:30:59.297886+00:00

measured 7 of 7 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-15T19:16:03.077573Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-05-20T19:43:44.121220Z

Reference resolution

36 of 36 outbound references displayed

  • verified exact0
  • verified fuzzy10
  • unresolved25
  • parse uncertain0
  • malformed identifier1
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 8d396b3e-ca1c-42b2-80c0-76af3d123a3b · outbound

This paper cites L1: Controlling How Long A Reasoning Model Thinks With Reinforcement Learning.

Making Small Language Models Efficient Reasoners: Intervention, Supervision, Reinforcement L1: Controlling How Long A Reasoning Model Thinks With Reinforcement Learning

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-15T22:11:33.877772Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T22:11:33.877772Z digest=sha256:c557551197f34f7a1781e1e46ae4a5017f94e9d7e0ec067f7f7f5f0c845cd62a

Observation c87e41e7-106a-4d93-954e-18bc23e143de · outbound

This paper cites The Surprising Effectiveness of Test-Time Training for Few-Shot Learning.

Making Small Language Models Efficient Reasoners: Intervention, Supervision, Reinforcement The Surprising Effectiveness of Test-Time Training for Few-Shot Learning

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-15T22:11:33.883525Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T22:11:33.883525Z digest=sha256:9524db905d04be236b05501dd911169575b8f678723bf9ecacdaf893299078c1

Observation 50380b6f-a70b-4f9a-94cc-0d5ef0ec6303 · outbound

This paper cites Precise Length Control in Large Language Models.

Making Small Language Models Efficient Reasoners: Intervention, Supervision, Reinforcement Precise Length Control in Large Language Models

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-15T22:11:33.887907Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T22:11:33.887907Z digest=sha256:0eaa8912261e664bfb0e2613d01d864648dbd55c931dd7be4ce2ab3b8d6b741f

Observation 734d0927-4c52-4467-a238-670416fd82ec · outbound

This paper cites FrugalGPT: How to Use Large Language Models While Reducing Cost and Improving Performance.

Making Small Language Models Efficient Reasoners: Intervention, Supervision, Reinforcement FrugalGPT: How to Use Large Language Models While Reducing Cost and Improving Performance

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-15T22:11:33.892890Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T22:11:33.892890Z digest=sha256:5b61429fccfd49abfc4215e05c3237b093b2b39d425488806283265df5a5bbea

Observation e2b6c686-0811-4902-80a7-94087760d863 · outbound

This paper cites GPTQ: Accurate Post-Training Quantization for Generative Pre-trained Transformers.

Making Small Language Models Efficient Reasoners: Intervention, Supervision, Reinforcement GPTQ: Accurate Post-Training Quantization for Generative Pre-trained Transformers

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-15T22:11:33.897078Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T22:11:33.897078Z digest=sha256:e82b5dc69ad0b7797962f439d0b2257fd35bbc4293d7f52117629bcb1946f285

Observation a60b0daa-05dc-4f83-95f0-d20e549d7e22 · outbound

This paper cites Gemini 2.0 flash thinking mode (gemini-2.0f lash-thinking-exp-1219), 2024.

Making Small Language Models Efficient Reasoners: Intervention, Supervision, Reinforcement Gemini 2.0 flash thinking mode (gemini-2.0f lash-thinking-exp-1219), 2024

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T22:11:34.641536Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-15T22:11:33.901046Z digest=sha256:13f050b9b96153e276bf57e955b97202584456f390d8c45dc7b8212cf88eb27e

Observation 58b1928c-096f-4a7d-aa3d-e4ce852ead83 · outbound

This paper cites Test-time training provably improves transformers as in-context learners.

Making Small Language Models Efficient Reasoners: Intervention, Supervision, Reinforcement Test-time training provably improves transformers as in-context learners

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-15T22:11:33.905805Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T22:11:33.905805Z digest=sha256:a29440451e9d0e4c8c3888745e7c201fc4c89518353463a65a3fc5c418e22144

Observation df98ec15-9c2c-45de-8b48-d13c6d332cd3 · outbound

This paper cites DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning.

Making Small Language Models Efficient Reasoners: Intervention, Supervision, Reinforcement DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-15T22:11:33.910188Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T22:11:33.910188Z digest=sha256:df05596c3f5fca1e9763a35c62054da78142db9682a3e4c12f6c997425889b67

Observation ae9df6fc-5089-4b4e-af9e-b27fbdbcda68 · outbound

This paper cites Language Model Cascades: Token-level uncertainty and beyond.

Making Small Language Models Efficient Reasoners: Intervention, Supervision, Reinforcement Language Model Cascades: Token-level uncertainty and beyond

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-15T22:11:33.914295Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T22:11:33.914295Z digest=sha256:456c2118d0aa9b6f64227c20f9edea1092c65c485197a97952f210f754003bcd

Observation 06e46da3-438e-4daa-978b-9dfe39d5fd7a · outbound

This paper cites ThinkPrune: Pruning Long Chain-of-Thought of LLMs via Reinforcement Learning.

Making Small Language Models Efficient Reasoners: Intervention, Supervision, Reinforcement ThinkPrune: Pruning Long Chain-of-Thought of LLMs via Reinforcement Learning

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-15T22:11:33.918046Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T22:11:33.918046Z digest=sha256:4dc4a13a1c1c7052c800c7246e53caf73df0d95a2a86b84ca58da817e1f2dd21

Observation d77aac5a-79c2-46f2-ac1c-d4b8ee6c0d13 · outbound

This paper cites OpenAI o1 System Card.

Making Small Language Models Efficient Reasoners: Intervention, Supervision, Reinforcement OpenAI o1 System Card

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-15T22:11:33.921887Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T22:11:33.921887Z digest=sha256:fbe33183dba889d553987d628269bba6e7acc971f51e370f2760d9c9f5c19ab4

Observation 9dc70092-78b9-4b5b-afde-f8834b657351 · outbound

This paper cites Fast inference from transformers via speculative decoding.

Making Small Language Models Efficient Reasoners: Intervention, Supervision, Reinforcement Fast inference from transformers via speculative decoding

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T22:11:34.629933Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-15T22:11:33.925186Z digest=sha256:8649d3ac7c8bd4c5ef902749e2628444fa9c96ca7182234f7f0b168b0173be8a

Observation 2e33c4ad-7933-458f-a900-2c504879f806 · outbound

This paper cites Autobalance: Optimized loss functions for imbalanced data.

Making Small Language Models Efficient Reasoners: Intervention, Supervision, Reinforcement Autobalance: Optimized loss functions for imbalanced data

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T22:11:34.618115Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-15T22:11:33.928576Z digest=sha256:9173c06427ef828bc8709da3c4bae0f506539f1ce1f6b4c222a1f7d20fe0ad16

Observation 20a9a835-c90b-4c5e-806a-38350e17314d · outbound

This paper cites Small models struggle to learn from strong reasoners.

Making Small Language Models Efficient Reasoners: Intervention, Supervision, Reinforcement Small models struggle to learn from strong reasoners

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-15T22:11:33.931708Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T22:11:33.931708Z digest=sha256:3aed0ee366093c0b5596f1a88117c3dacc0d24db0b3f819e3327f6c2de574717

Observation 622e7f77-58bb-4f78-bce7-d89a22e1bf4f · outbound

This paper cites Let’s verify step by step.

Making Small Language Models Efficient Reasoners: Intervention, Supervision, Reinforcement Let’s verify step by step

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-15T22:11:33.935195Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T22:11:33.935195Z digest=sha256:374734ccb32db993837430362d4924e84303db81f63b18c090d331991b92d3f6

Observation b1cd46e9-9114-4316-91df-8934f2b3eb7c · outbound

This paper cites Awq: Activation-aware weight quantization for on-device llm compression and acceleration.

Making Small Language Models Efficient Reasoners: Intervention, Supervision, Reinforcement Awq: Activation-aware weight quantization for on-device llm compression and acceleration

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-15T22:11:33.938798Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T22:11:33.938798Z digest=sha256:c17c7bf51fb29b2fcb1cd2fda0ac47bb288ebb4f54ec03795ccfe02db9dc51d9

Observation c5093dcf-123c-4933-8cfa-b9efb290d97e · outbound

This paper cites Tang, Manan Roongta, Colin Cai, Jeffrey Luo, Tianjun Zhang, Li Erran Li, Raluca Ada Popa, and Ion Stoica.

Making Small Language Models Efficient Reasoners: Intervention, Supervision, Reinforcement Tang, Manan Roongta, Colin Cai, Jeffrey Luo, Tianjun Zhang, Li Erran Li, Raluca Ada Popa, and Ion Stoica

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T22:11:34.594816Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-15T22:11:33.942182Z digest=sha256:a385a600e697fc81205b06a7b390bdb002dc47172b1ee0c625c49a25451aa3fe

Observation 289ba2ba-3319-49a3-bb7a-f9d0643dcc30 · outbound

This paper cites Long-tail learning via logit adjustment.

Making Small Language Models Efficient Reasoners: Intervention, Supervision, Reinforcement Long-tail learning via logit adjustment

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-15T22:11:33.946059Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T22:11:33.946059Z digest=sha256:1b55a656bab8de09ba3335e4df560c6b505640aa7d411040e5f61d6e23ca4124

Observation 283a70c5-b00c-4a19-bb2a-62ee448c36a9 · outbound

This paper cites s1: Simple test-time scaling.

Making Small Language Models Efficient Reasoners: Intervention, Supervision, Reinforcement s1: Simple test-time scaling

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-15T22:11:33.949614Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T22:11:33.949614Z digest=sha256:f128dd45ee7d5ec2228585e002c762834538b3e4511bdd185020800f10be7ead

Observation b2e8ccc2-6e5a-41c1-ab96-fd265de3eab9 · outbound

This paper cites Gpt-4o mini: advancing cost-efficient intelligence, 2024.

Making Small Language Models Efficient Reasoners: Intervention, Supervision, Reinforcement Gpt-4o mini: advancing cost-efficient intelligence, 2024

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T22:11:34.577740Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-15T22:11:33.953453Z digest=sha256:2bf2a879df947fe308b412b32c46cf2ef6f33486e1a47dfdb30e093d8a071fef

Observation e972ac4f-18f0-4cdd-8a9d-7a15da8a7927 · outbound

This paper cites Gpqa: A graduate-level google-proof q&a benchmark.

Making Small Language Models Efficient Reasoners: Intervention, Supervision, Reinforcement Gpqa: A graduate-level google-proof q&a benchmark

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-15T22:11:33.957798Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T22:11:33.957798Z digest=sha256:927628cffd51aaeef015c69e0029916e6a0626ebfbd354b93934fc6060ab2ae3

Observation 2e45beb1-1e63-4b61-af85-f924f23096fc · outbound

This paper cites Scaling Test-Time Compute Without Verification or RL is Suboptimal.

Making Small Language Models Efficient Reasoners: Intervention, Supervision, Reinforcement Scaling Test-Time Compute Without Verification or RL is Suboptimal

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-15T22:11:33.962269Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T22:11:33.962269Z digest=sha256:ba6892ea25cee9351ae2d158fd7702a536e84edc021b612c7fc80cce7e6223c8

Observation 3b883b6d-825e-40e2-8a28-c751a8bc144a · outbound

This paper cites Outrageously Large Neural Networks: The Sparsely-Gated Mixture-of-Experts Layer.

Making Small Language Models Efficient Reasoners: Intervention, Supervision, Reinforcement Outrageously Large Neural Networks: The Sparsely-Gated Mixture-of-Experts Layer

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-15T22:11:33.967068Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T22:11:33.967068Z digest=sha256:7fcccee71aa1cfb3b5fd1b3a385e222caa7091100a2f85e8add01c31a1de5fbf

Observation bf1e8d8b-aeff-4d37-80c8-d49962fd5640 · outbound

This paper cites Scaling LLM Test-Time Compute Optimally can be More Effective than Scaling Model Parameters.

Making Small Language Models Efficient Reasoners: Intervention, Supervision, Reinforcement Scaling LLM Test-Time Compute Optimally can be More Effective than Scaling Model Parameters

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-15T22:11:33.971434Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T22:11:33.971434Z digest=sha256:db6d93dafa416ca1ca2f58802ba00c64c8e88241b3c6427ff87eafc308879a1d

Observation 515e1ab3-531f-4550-b5f1-783b83f88ef4 · outbound

This paper cites Scalable Chain of Thoughts via Elastic Reasoning.

Making Small Language Models Efficient Reasoners: Intervention, Supervision, Reinforcement Scalable Chain of Thoughts via Elastic Reasoning

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-15T22:11:33.975998Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T22:11:33.975998Z digest=sha256:cf2da55e92d6869ac6f316c604a4b1da7b5c199310e70927ff8ddb5e80da92d1

Observation e0a603d5-c187-43de-afcc-3bcf18c85c01 · outbound

This paper cites Qwen2.5 Technical Report.

Making Small Language Models Efficient Reasoners: Intervention, Supervision, Reinforcement Qwen2.5 Technical Report

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-15T22:11:33.979744Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T22:11:33.979744Z digest=sha256:c1128cd71278a07a028bb5f6779e20c74f0d1a4ba592037b08ab9fb54cf98807

Observation 974af263-8d69-4ade-92f7-82521c49154f · outbound

This paper cites Qwen2.5-Math Technical Report: Toward Mathematical Expert Model via Self-Improvement.

Making Small Language Models Efficient Reasoners: Intervention, Supervision, Reinforcement Qwen2.5-Math Technical Report: Toward Mathematical Expert Model via Self-Improvement

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-15T22:11:33.983505Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T22:11:33.983505Z digest=sha256:7db5a156fb214615b0a730b0bf5a964e9e9fe4aff2c7bbc295879e9f6c828be4

Observation c09474c1-4f25-499c-831b-50a3951f0c2a · outbound

This paper cites Towards thinking-optimal scaling of test-time compute for llm reasoning.

Making Small Language Models Efficient Reasoners: Intervention, Supervision, Reinforcement Towards thinking-optimal scaling of test-time compute for llm reasoning

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-15T22:11:33.987395Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T22:11:33.987395Z digest=sha256:e53506b262eee3196f681dd9b884a6ae90dd8c19f4327b7648a2a0e7031b685a

Observation 1a7ccb58-5918-45cb-b07c-a8493b5f62cb · outbound

This paper cites Following Length Constraints in Instructions.

Making Small Language Models Efficient Reasoners: Intervention, Supervision, Reinforcement Following Length Constraints in Instructions

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-15T22:11:33.991234Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T22:11:33.991234Z digest=sha256:7eb375cb262a71b8bd8f4ccf5b0b359a59eef970c1ec92b56dec87e5b2da5a76

Observation 723c00b6-029d-4125-817b-5881c65d56c4 · outbound

This paper cites Selective attention: Enhancing transformer through principled context control.

Making Small Language Models Efficient Reasoners: Intervention, Supervision, Reinforcement Selective attention: Enhancing transformer through principled context control

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T22:11:34.560114Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-15T22:11:33.995774Z digest=sha256:b379cff580e2e05eac211c5d1fb3e2a856dc12bb461b1838a64e418ac36db584

Observation 341dd39d-09de-457a-9e35-50c36a47f104 · outbound

This paper cites Efficient contextual LLM cascades through budget-constrained policy learning.

Making Small Language Models Efficient Reasoners: Intervention, Supervision, Reinforcement Efficient contextual LLM cascades through budget-constrained policy learning

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T22:11:34.547918Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-15T22:11:33.999567Z digest=sha256:2d7ce887006228543477ce85197aaaa21d985ea31c6782caee2d404f778914aa

Observation 95e6457c-71a2-4a0b-ab20-f7906668edc1 · outbound

This paper cites Class-attribute priors: adapting optimization to heterogeneity and fairness objective.

Making Small Language Models Efficient Reasoners: Intervention, Supervision, Reinforcement Class-attribute priors: adapting optimization to heterogeneity and fairness objective

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T22:11:34.535486Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-15T22:11:34.003119Z digest=sha256:df97856a136c07f45a41490fa98f9ba47daab0ff622e55d677a28e5b73fccc17

Observation e4933300-2e03-433c-a78c-bcb55c7094e3 · outbound

This paper cites Pleaseanalyze the following text and determine if there are any meaningless repetitions of identical sentences.

Making Small Language Models Efficient Reasoners: Intervention, Supervision, Reinforcement Pleaseanalyze the following text and determine if there are any meaningless repetitions of identical sentences

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T22:11:34.524261Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-15T22:11:34.007830Z digest=sha256:497e8c14731d3c375782961111d586bb71d4d8e7d2e991edd78b6cefb80c6a20

Observation cf0b46b3-7ea4-4360-939d-7f887e96ca61 · outbound

This paper cites If the model prematurely generates the end-of-thinking token before reaching the desired length, the token is removed, and generation continues until the target length is reached.

Making Small Language Models Efficient Reasoners: Intervention, Supervision, Reinforcement If the model prematurely generates the end-of-thinking token before reaching the desired length, the token is removed, and generation continues until the target length is reached

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T22:11:34.512528Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-15T22:11:34.011237Z digest=sha256:3b6cc8a341f47336566372ef9638b42c338464e1b9334544f581dcb41d99a129

Observation 9d808209-4fd3-4a47-bac1-8d66ce1414bc · outbound

This paper cites an unresolved cited work.

Making Small Language Models Efficient Reasoners: Intervention, Supervision, Reinforcement Unresolved cited work

Reference 35

Resolution
unresolved
raw_fallback, observed 2026-08-15T22:11:34.499056Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-15T22:11:34.016188Z digest=sha256:ffc69002cb5c8d9312f5a452225d7302edb31bfdcff7a644306b2ee26c42b020

Observation 7f619b3d-b0d2-4d11-b6b4-f3d34b2dbe8a · outbound

This paper cites 2k” and “4k.

Making Small Language Models Efficient Reasoners: Intervention, Supervision, Reinforcement 2k” and “4k

Reference 36

Resolution
malformed identifier
raw_fallback, observed 2026-08-15T22:11:34.111991Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-15T22:11:34.020610Z digest=sha256:77d670f6a692b25e330f9df9d09c5df7e57e978a00c42cda84649b0b50f5614a

Pith citing papers

Observation a90e90b1-f664-4c52-a947-aabcf031f41f · inbound

Stop Overthinking: A Survey on Efficient Reasoning for Large Language Models cites this paper.

Stop Overthinking: A Survey on Efficient Reasoning for Large Language Models Making Small Language Models Efficient Reasoners: Intervention, Supervision, Reinforcement

Reference 240

Resolution
verified exact
arxiv_id, observed 2026-05-14T01:29:57.422076Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-05-14T01:29:56.480020Z digest=sha256:d7d1994c3cac146d5ca618df0cd53cdeeb0efa7cf6ca3b38465a348dd76ed2d5

Observation 6820098a-2d10-4b21-ae17-00996f34dbae · inbound

BREAD: Branched Rollouts from Expert Anchors Bridge SFT & RL for Reasoning cites this paper.

BREAD: Branched Rollouts from Expert Anchors Bridge SFT & RL for Reasoning Making Small Language Models Efficient Reasoners: Intervention, Supervision, Reinforcement

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-15T19:16:03.077573Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T19:16:03.077573Z digest=sha256:bf63f4308a55be644a8a5776ee714ec7f3f11f6b6d7b22ba040e29a2e71f8dc5

Observation caafa4cf-1572-4ec2-90a7-2cb382e35772 · inbound

Bayesian Social Deduction with Graph-Informed Language Models cites this paper.

Bayesian Social Deduction with Graph-Informed Language Models Making Small Language Models Efficient Reasoners: Intervention, Supervision, Reinforcement

Reference 76

Resolution
verified exact
arxiv_id, observed 2026-05-19T07:32:09.410714Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-05-19T07:27:20.027179Z digest=sha256:bb54267c8442d254de0e853d4f993302c82967a1eb3f44cb4dc5eac2d145aab6

Observation 8c401132-526a-4f28-975e-e06499a9d6a7 · inbound

Post Reasoning: Improving the Performance of Non-Thinking Models at No Cost cites this paper.

Post Reasoning: Improving the Performance of Non-Thinking Models at No Cost Making Small Language Models Efficient Reasoners: Intervention, Supervision, Reinforcement

Reference 233

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T20:06:09.979128Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-05-08T10:19:08.451445Z digest=sha256:232b403438240f3c3c4d99124c56a1d3af8fda1dabdaae57ebb74b3ee52b4457

Observation 6788d978-46b9-445a-9287-e00caccf6347 · inbound

VSPO: Vector-Steered Policy Optimization for Behavioral Control cites this paper.

VSPO: Vector-Steered Policy Optimization for Behavioral Control Making Small Language Models Efficient Reasoners: Intervention, Supervision, Reinforcement

Reference 32

Resolution
verified exact
arxiv_id, observed 2026-05-20T19:43:44.123476Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-05-20T19:39:55.294398Z digest=sha256:70e4d60e1ce45acc4e1c9002503523ca7bcc9b12e2ad5b54d9e3a39c5d4bc385

Observation 8f3b2d51-b066-417b-b03b-9934dd5d82f7 · inbound

OS-Pruner: Pruning Chains-of-Thought of Reasoning Models via Optimal Stopping cites this paper.

OS-Pruner: Pruning Chains-of-Thought of Reasoning Models via Optimal Stopping Making Small Language Models Efficient Reasoners: Intervention, Supervision, Reinforcement

Reference 25

Resolution
unresolved
no resolver link, observed 2026-07-14T07:09:11.486851Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T07:09:11.486851Z digest=sha256:c15008120d4cf00594e70c1c6a7d8470994eb54dc66709caf839fe1be6779ce5

Observation 6a2788d7-a87f-4422-ad06-7d008e98f891 · inbound

From Trajectories to Prefixes: Reusing Teacher Trajectories via Replayed Prefixes and Online Continuation cites this paper.

From Trajectories to Prefixes: Reusing Teacher Trajectories via Replayed Prefixes and Online Continuation Making Small Language Models Efficient Reasoners: Intervention, Supervision, Reinforcement

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-02T08:56:34.604372Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T08:56:34.604372Z digest=sha256:a1680d67ce253f7fac7769b9d388e559431e25e7edbf1e5112f49016de125ac3