Pith. sign in

Paper Citation Record · LEDGER

Enhancing Reasoning Capabilities in SLMs with Reward Guided Dataset Distillation

As of 19 August 2026, this Paper Citation Record lists 61 of 61 outbound references and 0 inbound Pith citation observations for arXiv:2507.00054.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2507.00054 v1

Coverage vector

measured 61 of 61 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-06T22:45:04.656219Z

measured 61 of 61 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-19T06:32:44.657259+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

61 of 61 outbound references displayed

  • verified exact4
  • verified fuzzy32
  • unresolved25
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation f9f7112a-3f5c-4430-89ce-ec48996e0545 · outbound

This paper cites , Aneja, J.

Enhancing Reasoning Capabilities in SLMs with Reward Guided Dataset Distillation , Aneja, J

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:45:11.959401Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-08-06T22:44:59.832083Z digest=sha256:23ee125968b8575340cdc4b26d9a4f4be86adb6f3c47b90b4e423b40f7d59609

Observation 5334a726-edd3-4b45-a7fd-d1a4b28a4136 · outbound

This paper cites , Vieillard, N.

Enhancing Reasoning Capabilities in SLMs with Reward Guided Dataset Distillation , Vieillard, N

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:45:11.756821Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-08-06T22:44:59.887421Z digest=sha256:38f70dc352f06099acb59dbba10deb04a77768257536b55428c6381d2ed9dfa2

Observation 7976eb97-c850-406e-8b1a-15efc9efb035 · outbound

This paper cites , Vieillard, N.

Enhancing Reasoning Capabilities in SLMs with Reward Guided Dataset Distillation , Vieillard, N

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:45:11.433341Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-08-06T22:44:59.969850Z digest=sha256:446c2b5bbe2405d335160520f9c735e3c02f11f285d57cea9529423f562d9b53

Observation 963a99e5-8b87-4347-bc6c-33ec5f7679fa · outbound

This paper cites Towards Understanding Distilled Reasoning Models: A Representational Approach.

Enhancing Reasoning Capabilities in SLMs with Reward Guided Dataset Distillation Towards Understanding Distilled Reasoning Models: A Representational Approach

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-06T22:45:00.068336Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T22:45:00.068336Z digest=sha256:92cc87fd2ab717106f93c4624edc94fdfdcddea0c05c3932f27c8e6c7160f224

Observation 9813e50b-31b0-44d5-a629-a293fc9c17f5 · outbound

This paper cites , Sastre, I.

Enhancing Reasoning Capabilities in SLMs with Reward Guided Dataset Distillation , Sastre, I

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:45:11.259853Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-08-06T22:45:00.163772Z digest=sha256:cbfa2a4d639b8b4f3913c5879c96b2140e5816fdd185673af4ff184ac90f693b

Observation ee832064-5677-498b-93ca-96c24d0a7770 · outbound

This paper cites Smaller, Weaker, Yet Better: Training LLM Reasoners via Compute-Optimal Sampling.

Enhancing Reasoning Capabilities in SLMs with Reward Guided Dataset Distillation Smaller, Weaker, Yet Better: Training LLM Reasoners via Compute-Optimal Sampling

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-06T22:45:00.253407Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T22:45:00.253407Z digest=sha256:6d64f5e3e72629a225c7a48f7a62d958d163da58518d0e656f7f0c046cf5a467

Observation 7ca5a211-4a5a-4bd4-8f10-23e74975e1db · outbound

This paper cites , Peebles, B.

Enhancing Reasoning Capabilities in SLMs with Reward Guided Dataset Distillation , Peebles, B

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:45:11.016657Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-08-06T22:45:00.323553Z digest=sha256:e61193284e42d9a27cde44fa766be793e3de8ab6f32f41a6d9af1c3c5db9d03a

Observation 8585932b-95fd-4e64-8168-4765ac6904cd · outbound

This paper cites Training Verifiers to Solve Math Word Problems.

Enhancing Reasoning Capabilities in SLMs with Reward Guided Dataset Distillation Training Verifiers to Solve Math Word Problems

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-06T22:45:00.383119Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T22:45:00.383119Z digest=sha256:6db19918ba0e4ef248f73e5c52656e57895285d7b791330d5290728a5e2058eb

Observation ae03c8ce-0b48-4fad-9d78-2c9f0baa78ed · outbound

This paper cites \ Ngo, C.

Enhancing Reasoning Capabilities in SLMs with Reward Guided Dataset Distillation \ Ngo, C

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:45:10.872019Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-08-06T22:45:00.432702Z digest=sha256:72055c046d4cd7e30199c8781771702a54f92813c05ba89d9d5dc1be3489685b

Observation b0065cb9-5d07-4763-8d75-6ddc406cc9e4 · outbound

This paper cites , See, A.

Enhancing Reasoning Capabilities in SLMs with Reward Guided Dataset Distillation , See, A

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:45:10.714160Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-08-06T22:45:00.499454Z digest=sha256:452822561666b813e14ad7982f91e13edd0239ab1bb2538049010677d4c789a7

Observation fda81fd6-744d-48b2-a6d0-d940a7757342 · outbound

This paper cites , Feng, B.

Enhancing Reasoning Capabilities in SLMs with Reward Guided Dataset Distillation , Feng, B

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:45:10.515033Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-08-06T22:45:00.562770Z digest=sha256:a5b190fbd98dc0411e1b0291beb2b51df1431d9d8489b914b6860094191294d0

Observation 88c34840-15c4-4128-afbb-d71d4410be5d · outbound

This paper cites an unresolved cited work.

Enhancing Reasoning Capabilities in SLMs with Reward Guided Dataset Distillation Unresolved cited work

Reference 12

Resolution
unresolved
raw_fallback, observed 2026-08-06T22:45:10.354570Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-08-06T22:45:00.644914Z digest=sha256:80cf8928345dd3332fabbb5d7be4072a1f96588268f7ecd8eba07b39e71f6cb4

Observation 8a236ce7-78d9-45d5-9e9a-3c7e6eb7b7cd · outbound

This paper cites an unresolved cited work.

Enhancing Reasoning Capabilities in SLMs with Reward Guided Dataset Distillation Unresolved cited work

Reference 13

Resolution
unresolved
raw_fallback, observed 2026-08-06T22:45:10.118810Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-08-06T22:45:00.717936Z digest=sha256:0b5cf22adb0d0d3730b46a1ef9914bf105e4ed92ccda0906e8944c555dd501ee

Observation 8af77e78-1dab-42c3-bf88-2b07b2c7012b · outbound

This paper cites Advantage-Guided Distillation for Preference Alignment in Small Language Models.

Enhancing Reasoning Capabilities in SLMs with Reward Guided Dataset Distillation Advantage-Guided Distillation for Preference Alignment in Small Language Models

Reference 14

Resolution
verified exact
local_arxiv, observed 2026-08-06T22:45:05.664522Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-08-06T22:45:00.774715Z digest=sha256:171512ff0d3a4eea5d32b48645fe55f058610bd88259be59b72126c646e05d16

Observation ba453a7e-76aa-4268-970c-cc687a2a4b18 · outbound

This paper cites MiniLLM: On-Policy Distillation of Large Language Models.

Enhancing Reasoning Capabilities in SLMs with Reward Guided Dataset Distillation MiniLLM: On-Policy Distillation of Large Language Models

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-06T22:45:00.840611Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T22:45:00.840611Z digest=sha256:fe784951fe62d828137e3e8ffd9f5e30ba3a27abab9ecf938fb1820056bc3583

Observation ec3ea647-c2f7-4d1b-bdb7-e7e74b0b621e · outbound

This paper cites , Zhang, L L.

Enhancing Reasoning Capabilities in SLMs with Reward Guided Dataset Distillation , Zhang, L L

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:45:09.899083Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-08-06T22:45:00.916470Z digest=sha256:31e537f768be4660a7e1239a59773654b4d0a7f3c25fdc7e8ae1eab04cd40697

Observation 3a330c4d-5b4f-4e1d-91db-ba0f3d6215db · outbound

This paper cites DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning.

Enhancing Reasoning Capabilities in SLMs with Reward Guided Dataset Distillation DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-06T22:45:00.969331Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T22:45:00.969331Z digest=sha256:6ab8cfd27c7d43679696cf51bd9e071da77b7765ded01135537511816b67ddcb

Observation 2cc95be0-a373-4d5d-89c2-100f15211594 · outbound

This paper cites , Vinyals, O.

Enhancing Reasoning Capabilities in SLMs with Reward Guided Dataset Distillation , Vinyals, O

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:45:09.681927Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-08-06T22:45:01.034051Z digest=sha256:a8e26a3d923e4fc2dc3cd06548f09d1c0e5835a4cd5a8789b0abeb2f3a64265f

Observation 2988ec95-7570-4805-86fe-20b740f166c6 · outbound

This paper cites , Borgeaud, S.

Enhancing Reasoning Capabilities in SLMs with Reward Guided Dataset Distillation , Borgeaud, S

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:45:09.484507Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-08-06T22:45:01.104215Z digest=sha256:29eb4d67d2982f87ed1d169245bd318ba19e5a90318431c24786e8a4c3d06fe6

Observation 38769f70-6b26-4119-a44c-9bc87c2dcc74 · outbound

This paper cites , Li, C L.

Enhancing Reasoning Capabilities in SLMs with Reward Guided Dataset Distillation , Li, C L

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:45:09.258428Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-08-06T22:45:01.208161Z digest=sha256:2be88c1f8a7009b942f3a1498e2bac91dca87e5640d064f024c31061852400d4

Observation b2223c7a-1689-484d-a335-981e059a7d3c · outbound

This paper cites , Yin, Y.

Enhancing Reasoning Capabilities in SLMs with Reward Guided Dataset Distillation , Yin, Y

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:45:09.049575Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-08-06T22:45:01.282141Z digest=sha256:82a66caa0ebe9d45edf42574daa18b1cfc738cf7bb34a3c57b5170e2b85a0d6d

Observation c63aaa27-aeda-4fdf-b195-07a0ea7ca4b8 · outbound

This paper cites MindStar: Enhancing Math Reasoning in Pre-trained LLMs at Inference Time.

Enhancing Reasoning Capabilities in SLMs with Reward Guided Dataset Distillation MindStar: Enhancing Math Reasoning in Pre-trained LLMs at Inference Time

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-06T22:45:01.362923Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T22:45:01.362923Z digest=sha256:5eb7b077442159af73fa5e62009fa97eabf0b462d5cbb7bc1d1067f2c8679a30

Observation ba9ae507-d70d-49b0-9c49-f1052c19a884 · outbound

This paper cites \ Rush, A M.

Enhancing Reasoning Capabilities in SLMs with Reward Guided Dataset Distillation \ Rush, A M

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:45:08.841836Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-08-06T22:45:01.446884Z digest=sha256:327aea7bc1ed58d90ba6a31402d81987a36cee8342fcf16e1a33a60111a5ab72

Observation d76b30f7-97b5-49f2-8da9-86c67609c337 · outbound

This paper cites DistiLLM-2: A Contrastive Approach Boosts the Distillation of LLMs.

Enhancing Reasoning Capabilities in SLMs with Reward Guided Dataset Distillation DistiLLM-2: A Contrastive Approach Boosts the Distillation of LLMs

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-06T22:45:01.511764Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T22:45:01.511764Z digest=sha256:82d2bcfeb8252e5fcee3aeffaa1dfa25b731cefdf8084740c186664d7543bca8

Observation bf1372cb-2c2d-4815-b264-30df58d36e6d · outbound

This paper cites DistiLLM: Towards Streamlined Distillation for Large Language Models.

Enhancing Reasoning Capabilities in SLMs with Reward Guided Dataset Distillation DistiLLM: Towards Streamlined Distillation for Large Language Models

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-06T22:45:01.602702Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T22:45:01.602702Z digest=sha256:33994fde460b53daa19a08cf87ae1c5f3268594bb61132fb584337e68852c511

Observation bbfeb791-a91d-4c43-9cde-01ea480d017f · outbound

This paper cites , Fang, L.

Enhancing Reasoning Capabilities in SLMs with Reward Guided Dataset Distillation , Fang, L

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:45:08.621314Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-08-06T22:45:01.662995Z digest=sha256:5426734da06091de6390b95eab81b24c265d95314ce0d9483a923edb9442d684

Observation 1f40c54b-8e47-4293-9125-d763d5183b0d · outbound

This paper cites Winning Big with Small Models: Knowledge Distillation vs. Self-Training for Reducing Hallucination in Product QA Agents.

Enhancing Reasoning Capabilities in SLMs with Reward Guided Dataset Distillation Winning Big with Small Models: Knowledge Distillation vs. Self-Training for Reducing Hallucination in Product QA Agents

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-06T22:45:01.734181Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T22:45:01.734181Z digest=sha256:7e87487efd1c626d797005ca04b28f72f5b5cb9e12892330a660ebc0cccb4a16

Observation 1f1635b6-be41-4bcb-aa4b-a089978da199 · outbound

This paper cites GSM-Plus: A Comprehensive Benchmark for Evaluating the Robustness of LLMs as Mathematical Problem Solvers.

Enhancing Reasoning Capabilities in SLMs with Reward Guided Dataset Distillation GSM-Plus: A Comprehensive Benchmark for Evaluating the Robustness of LLMs as Mathematical Problem Solvers

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-06T22:45:01.803122Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T22:45:01.803122Z digest=sha256:266373102a8cee7f47f596757f92ffc9b3b8e89393afbf0e1735ac05a7385ad4

Observation a41021f4-9434-410f-86af-7fde9090a21d · outbound

This paper cites Evolving Knowledge Distillation with Large Language Models and Active Learning.

Enhancing Reasoning Capabilities in SLMs with Reward Guided Dataset Distillation Evolving Knowledge Distillation with Large Language Models and Active Learning

Reference 29

Resolution
verified exact
local_arxiv, observed 2026-08-06T22:45:05.399923Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-08-06T22:45:01.899416Z digest=sha256:aca60fdbbe336c9631e479e0273158a442f5e1d4b27fe735ffb4555ff03505a6

Observation c9427546-55e8-4e8e-81c2-4f74cca695e4 · outbound

This paper cites , Chen, C.

Enhancing Reasoning Capabilities in SLMs with Reward Guided Dataset Distillation , Chen, C

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:45:08.535787Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-08-06T22:45:01.987456Z digest=sha256:c96250b9cb5ea02ca96fe03af037e2816610ceeaee325807f0cb2291fb8231d4

Observation 663344f9-7e30-47e0-89bc-9ac38933d316 · outbound

This paper cites , Yang, Z.

Enhancing Reasoning Capabilities in SLMs with Reward Guided Dataset Distillation , Yang, Z

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:45:08.365808Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-08-06T22:45:02.050079Z digest=sha256:7a63248b0ade9a338910fded2847d8835fc307fe98cb5d059390e002c0a3b520

Observation 335802d2-6c7b-402c-8a4f-1c2f5861e9a2 · outbound

This paper cites , Dani, J.

Enhancing Reasoning Capabilities in SLMs with Reward Guided Dataset Distillation , Dani, J

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:45:08.178769Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-08-06T22:45:02.176949Z digest=sha256:b4b81144aae2dc644cef02b632cb62cbd5cd2050d14849f5f83f7da001779791

Observation f6df0b60-2933-495f-86bb-c081bc56fa71 · outbound

This paper cites an unresolved cited work.

Enhancing Reasoning Capabilities in SLMs with Reward Guided Dataset Distillation Unresolved cited work

Reference 33

Resolution
unresolved
raw_fallback, observed 2026-08-06T22:45:08.012500Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-08-06T22:45:02.296854Z digest=sha256:984725b5ff90d73ad2e972094aec937661c51a0424be957c982c0b5d5a0a1330

Observation 71d9bb01-9f8b-4ccc-a41a-b8716ccb4285 · outbound

This paper cites , Lerer, A.

Enhancing Reasoning Capabilities in SLMs with Reward Guided Dataset Distillation , Lerer, A

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:45:07.780048Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-08-06T22:45:02.381998Z digest=sha256:06f4c211dbb405bab64d628c23e7434ad45c5b74a6f9b2e50b386f024666b530

Observation 710f5486-52fa-4b6d-8121-0e832a6133de · outbound

This paper cites Progressive distillation induces an implicit curriculum.

Enhancing Reasoning Capabilities in SLMs with Reward Guided Dataset Distillation Progressive distillation induces an implicit curriculum

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-06T22:45:02.463708Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T22:45:02.463708Z digest=sha256:d74f900457d9dc06e6174cd8192171484d8f211c452f5fc74efb1c5313da81fe

Observation 0d86e769-ece0-4b0e-8736-d57c1f90975d · outbound

This paper cites , Kim, D.

Enhancing Reasoning Capabilities in SLMs with Reward Guided Dataset Distillation , Kim, D

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:45:07.585952Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-08-06T22:45:02.552158Z digest=sha256:1b84429cdd1defd2cca64e9e8fd1784dd8177eb7d7e92293d8477e5067298895

Observation 17d3ac87-dd4d-41a7-bd6d-017ad7752bd1 · outbound

This paper cites \ Xie, S.

Enhancing Reasoning Capabilities in SLMs with Reward Guided Dataset Distillation \ Xie, S

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:45:07.347692Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-08-06T22:45:02.621187Z digest=sha256:7a045d88df02d4bf5c71399d1582cc9d33830686f4e264ccf7c21db814bf577e

Observation dedb6313-b2d1-43cd-89cb-46c0ddf3e22b · outbound

This paper cites CourseGPT-zh: an Educational Large Language Model Based on Knowledge Distillation Incorporating Prompt Optimization.

Enhancing Reasoning Capabilities in SLMs with Reward Guided Dataset Distillation CourseGPT-zh: an Educational Large Language Model Based on Knowledge Distillation Incorporating Prompt Optimization

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-06T22:45:02.697368Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T22:45:02.697368Z digest=sha256:7397fc7b7909fefb6a7a3971fc6f890c4d9994c4e0ec474ca0383b1e8a3a2c0e

Observation ab12e167-1817-48ba-a32c-9941bc621f32 · outbound

This paper cites DistilBERT, a distilled version of BERT: smaller, faster, cheaper and lighter.

Enhancing Reasoning Capabilities in SLMs with Reward Guided Dataset Distillation DistilBERT, a distilled version of BERT: smaller, faster, cheaper and lighter

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-06T22:45:02.785656Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T22:45:02.785656Z digest=sha256:8fee1fb3fc1ec573a6c8493492575872681ad2ae24bb9dee67e83554fb34965d

Observation 870979c2-4e34-441d-8bad-e60b4bcbc6c5 · outbound

This paper cites Spurious Rewards: Rethinking Training Signals in RLVR.

Enhancing Reasoning Capabilities in SLMs with Reward Guided Dataset Distillation Spurious Rewards: Rethinking Training Signals in RLVR

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-06T22:45:02.881915Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T22:45:02.881915Z digest=sha256:b55cf3c8df013e99254d899afa4e09cfb0e7775de934c16e9dc8b1e0a710ce47

Observation f2c60319-8fc3-427f-901d-8f0079bc5860 · outbound

This paper cites , Wang, P.

Enhancing Reasoning Capabilities in SLMs with Reward Guided Dataset Distillation , Wang, P

Reference 41

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:45:07.207353Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-08-06T22:45:02.953539Z digest=sha256:41c016e64ae46637c95eb5b242406061dc74e797b8a265ef9786f77cf0da5d44

Observation 98e0f490-a8be-44cd-9845-3e4ee7a4db71 · outbound

This paper cites The Curse of Recursion: Training on Generated Data Makes Models Forget.

Enhancing Reasoning Capabilities in SLMs with Reward Guided Dataset Distillation The Curse of Recursion: Training on Generated Data Makes Models Forget

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-06T22:45:03.029566Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T22:45:03.029566Z digest=sha256:42b5be15cb595ec259624e6dfcfc05d43deaf5c2db36f7f0171819e6e95a6960

Observation 13b7323d-6a46-4c54-9d5e-0eae71e6578a · outbound

This paper cites Gemma 3 Technical Report.

Enhancing Reasoning Capabilities in SLMs with Reward Guided Dataset Distillation Gemma 3 Technical Report

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-06T22:45:03.122039Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T22:45:03.122039Z digest=sha256:1e727c539cac40cb3c0c07c07419e422f4720c4743d397ef05454bfd8757fbec

Observation 48454789-9132-4ff1-9fe3-3a143de98d2c · outbound

This paper cites , Riviere, M.

Enhancing Reasoning Capabilities in SLMs with Reward Guided Dataset Distillation , Riviere, M

Reference 44

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:45:07.097328Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-08-06T22:45:03.183890Z digest=sha256:eeb9fa0cfa9f7cd3f6ff5269f2f363418cd1f4a56cfded3bd07db2b2af80a5cb

Observation e8406170-4e51-474e-9ba0-a26b4c29b939 · outbound

This paper cites , Han, Y.

Enhancing Reasoning Capabilities in SLMs with Reward Guided Dataset Distillation , Han, Y

Reference 45

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:45:06.972873Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-08-06T22:45:03.261446Z digest=sha256:d4f20bd0581f234d8e50ce40a9aabc47987312525d7e28911885b50c24c68e11

Observation 8a34b152-2464-440c-a5d0-9a7c4e36895b · outbound

This paper cites On Teacher Hacking in Language Model Distillation.

Enhancing Reasoning Capabilities in SLMs with Reward Guided Dataset Distillation On Teacher Hacking in Language Model Distillation

Reference 46

Resolution
verified exact
local_arxiv, observed 2026-08-06T22:45:05.113351Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-08-06T22:45:03.331296Z digest=sha256:6c9f9efaa793c92f4237dbca3919bdd433d0ebb5b29cf1d3264b15d43544e2f4

Observation 210ff84e-986d-4660-8fd8-a7be4c99448c · outbound

This paper cites Who Taught You That? Tracing Teachers in Model Distillation.

Enhancing Reasoning Capabilities in SLMs with Reward Guided Dataset Distillation Who Taught You That? Tracing Teachers in Model Distillation

Reference 47

Resolution
verified exact
local_arxiv, observed 2026-08-06T22:45:04.943544Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-08-06T22:45:03.441504Z digest=sha256:5643a8fcb27de99e3428b5387cedf54f0e892582328836cc632cf96e765c295a

Observation 42e753af-4bd2-485d-819e-d73112561ade · outbound

This paper cites , Deng, Y.

Enhancing Reasoning Capabilities in SLMs with Reward Guided Dataset Distillation , Deng, Y

Reference 48

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:45:06.848770Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-08-06T22:45:03.513056Z digest=sha256:14b314ff8c5b762179bf5508258e6d9b30306eadd26ad58bedd013b98e7d57a3

Observation 5fc159c6-a4f4-462a-8797-6b4e79b3f834 · outbound

This paper cites A Comprehensive Survey of Small Language Models in the Era of Large Language Models: Techniques, Enhancements, Applications, Collaboration with LLMs, and Trustworthiness.

Enhancing Reasoning Capabilities in SLMs with Reward Guided Dataset Distillation A Comprehensive Survey of Small Language Models in the Era of Large Language Models: Techniques, Enhancements, Applications, Collaboration with LLMs, and Trustworthiness

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-06T22:45:03.646532Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T22:45:03.646532Z digest=sha256:7f171416622953d01d39565b1d8daa12c9ca8fa5174e9e263167e592a28f2e86

Observation 41d95ac2-3c60-4166-a078-0a33e91b1192 · outbound

This paper cites , Zhu, J Y.

Enhancing Reasoning Capabilities in SLMs with Reward Guided Dataset Distillation , Zhu, J Y

Reference 50

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:45:06.722937Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-08-06T22:45:03.749143Z digest=sha256:b2f335f7b8f82aca6cb0a632569cbc925ce46e6104b9b5484cbb7be8f64d9494

Observation ba61adf8-9c15-4502-b0db-7ff5881e1db4 · outbound

This paper cites an unresolved cited work.

Enhancing Reasoning Capabilities in SLMs with Reward Guided Dataset Distillation Unresolved cited work

Reference 51

Resolution
unresolved
raw_fallback, observed 2026-08-06T22:45:06.608963Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-08-06T22:45:03.850600Z digest=sha256:c3142aa4ae06b26d57d7e2eb4ac18a0478738b19c4d036581f525a6245abd578

Observation d3fcc0b3-7961-4740-aadb-991067663937 · outbound

This paper cites , Wang, X.

Enhancing Reasoning Capabilities in SLMs with Reward Guided Dataset Distillation , Wang, X

Reference 52

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:45:06.488819Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-08-06T22:45:03.920673Z digest=sha256:bf71938f34be4a30995211b4ffd8d33db4bc018b1e2b7308d989bb3519fa0f06

Observation 8eafc91f-9730-40aa-b099-dbda10b62673 · outbound

This paper cites , Bai, H.

Enhancing Reasoning Capabilities in SLMs with Reward Guided Dataset Distillation , Bai, H

Reference 53

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:45:06.376947Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-08-06T22:45:03.977383Z digest=sha256:aafb9a0b0b3811421280c491f28e1c9b0dffe3c89893e96b3a53daed3e74794e

Observation 6290ed46-791e-4ba6-aeac-df36eb0ac94b · outbound

This paper cites Speculative Knowledge Distillation: Bridging the Teacher-Student Gap Through Interleaved Sampling.

Enhancing Reasoning Capabilities in SLMs with Reward Guided Dataset Distillation Speculative Knowledge Distillation: Bridging the Teacher-Student Gap Through Interleaved Sampling

Reference 55

Resolution
unresolved
no resolver link, observed 2026-08-06T22:45:04.144603Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T22:45:04.144603Z digest=sha256:38277a6f4357be5e7ecbb9ecbbe42ad790040059844141586a8717bca711b27c

Observation c9a0cc89-1f21-4f07-be33-6c3063688a6b · outbound

This paper cites Qwen2.5 Technical Report.

Enhancing Reasoning Capabilities in SLMs with Reward Guided Dataset Distillation Qwen2.5 Technical Report

Reference 56

Resolution
unresolved
no resolver link, observed 2026-08-06T22:45:04.222611Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T22:45:04.222611Z digest=sha256:8350d76a715ce19c1c8828f8d728de87b511f34d31580d94e73ef6287f3c3bc4

Observation 5f3339e7-b708-4899-86bf-8386abf8fd21 · outbound

This paper cites Demystifying Long Chain-of-Thought Reasoning in LLMs.

Enhancing Reasoning Capabilities in SLMs with Reward Guided Dataset Distillation Demystifying Long Chain-of-Thought Reasoning in LLMs

Reference 57

Resolution
unresolved
no resolver link, observed 2026-08-06T22:45:04.300085Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T22:45:04.300085Z digest=sha256:50eb80fcea295a4a1ef3be165cd211ef3d65d6bb7b53cf5d0b6f579a189ae184

Observation 012f0bf2-3f3d-406b-b1c2-e8c678c75d5e · outbound

This paper cites , Wang, C.

Enhancing Reasoning Capabilities in SLMs with Reward Guided Dataset Distillation , Wang, C

Reference 58

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:45:06.275447Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-08-06T22:45:04.408916Z digest=sha256:84e7c233e120bc36596219b801d13f314871fd478a8cecc377c962f5bec6627c

Observation c1d9a273-3749-4648-b408-bc667bbe660f · outbound

This paper cites , Zhu, R.

Enhancing Reasoning Capabilities in SLMs with Reward Guided Dataset Distillation , Zhu, R

Reference 59

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:45:06.151692Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-08-06T22:45:04.485680Z digest=sha256:3078d900b1ed170d8f294c2acbee00376c4acc9798cd45945c23d636a4dab49b

Observation d6cefc1b-c2e3-4129-a043-5a2f8a170d5d · outbound

This paper cites , Zhu, R.

Enhancing Reasoning Capabilities in SLMs with Reward Guided Dataset Distillation , Zhu, R

Reference 60

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:45:06.021904Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-08-06T22:45:04.548728Z digest=sha256:40259f1d57a172423a33d21ce08c66e530d15d415e83f74459c2707ab10ecf18

Observation 5b25e6b2-6caf-47bb-b638-48ef19505fc7 · outbound

This paper cites , Shen, J.

Enhancing Reasoning Capabilities in SLMs with Reward Guided Dataset Distillation , Shen, J

Reference 61

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:45:05.904741Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-08-06T22:45:04.612738Z digest=sha256:857f16cb3a7eeacdd1751c9c3689b86018938b7a22edbfc7c33d3e26dc8a5073

Observation 5dae83d3-89e7-4d02-807a-4b57195fbd38 · outbound

This paper cites Distill Not Only Data but Also Rewards: Can Smaller Language Models Surpass Larger Ones?.

Enhancing Reasoning Capabilities in SLMs with Reward Guided Dataset Distillation Distill Not Only Data but Also Rewards: Can Smaller Language Models Surpass Larger Ones?

Reference 62

Resolution
unresolved
no resolver link, observed 2026-08-06T22:45:04.656219Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T22:45:04.656219Z digest=sha256:cfaae9855564ca91a70896bb2d6b5704ec5926d39963bd2900c7cabe858e1280

Pith citing papers

No inbound Pith citation observations are available.