Pith. sign in

Paper Citation Record · LEDGER

Enhancing Reasoning Capabilities in SLMs with Reward Guided Dataset Distillation

As of 8 August 2026, this Paper Citation Record lists 61 of 61 outbound references and 0 inbound Pith citation observations for arXiv:2507.00054.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2507.00054 v1

Coverage vector

measured 61 of 61 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-06T22:45:04.656219Z

measured 61 of 61 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-08T06:32:00.761636+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

61 of 61 outbound references displayed

  • verified exact4
  • verified fuzzy32
  • unresolved25
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation f9f7112a-3f5c-4430-89ce-ec48996e0545 · outbound

This paper cites , Aneja, J.

Enhancing Reasoning Capabilities in SLMs with Reward Guided Dataset Distillation , Aneja, J

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:45:11.959401Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-06T22:44:59.832083Z digest=sha256:4b39990cd6de024df27f83a7606ff7141274b5ccacdf342a909fb9156a4479e8

Observation 5334a726-edd3-4b45-a7fd-d1a4b28a4136 · outbound

This paper cites , Vieillard, N.

Enhancing Reasoning Capabilities in SLMs with Reward Guided Dataset Distillation , Vieillard, N

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:45:11.756821Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-06T22:44:59.887421Z digest=sha256:f3e01ada123c67cd4899e9b3d6c739e68550c268903c247a853fd97c0befb260

Observation 7976eb97-c850-406e-8b1a-15efc9efb035 · outbound

This paper cites , Vieillard, N.

Enhancing Reasoning Capabilities in SLMs with Reward Guided Dataset Distillation , Vieillard, N

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:45:11.433341Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-06T22:44:59.969850Z digest=sha256:5f16a04e308366fa85db984969b4991d10d7df266b048ebac40ebf3d37d280a5

Observation 963a99e5-8b87-4347-bc6c-33ec5f7679fa · outbound

This paper cites Towards Understanding Distilled Reasoning Models: A Representational Approach.

Enhancing Reasoning Capabilities in SLMs with Reward Guided Dataset Distillation Towards Understanding Distilled Reasoning Models: A Representational Approach

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-06T22:45:00.068336Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T22:45:00.068336Z digest=sha256:1a2b94a5938edb5d1965f505a167c7f40f17f712dbac110fbe2a8f38eca46812

Observation 9813e50b-31b0-44d5-a629-a293fc9c17f5 · outbound

This paper cites , Sastre, I.

Enhancing Reasoning Capabilities in SLMs with Reward Guided Dataset Distillation , Sastre, I

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:45:11.259853Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-06T22:45:00.163772Z digest=sha256:b09ef389fb90c6f36e1c922c9d29844af5dc90c4c2c67251739608f835d8ab48

Observation ee832064-5677-498b-93ca-96c24d0a7770 · outbound

This paper cites Smaller, Weaker, Yet Better: Training LLM Reasoners via Compute-Optimal Sampling.

Enhancing Reasoning Capabilities in SLMs with Reward Guided Dataset Distillation Smaller, Weaker, Yet Better: Training LLM Reasoners via Compute-Optimal Sampling

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-06T22:45:00.253407Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T22:45:00.253407Z digest=sha256:10934d0949996e8611b9593414ae251ce89898634f5285afe04d2e0471e22150

Observation 7ca5a211-4a5a-4bd4-8f10-23e74975e1db · outbound

This paper cites , Peebles, B.

Enhancing Reasoning Capabilities in SLMs with Reward Guided Dataset Distillation , Peebles, B

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:45:11.016657Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-06T22:45:00.323553Z digest=sha256:459d1ad5c94a059f25ba90b25875684ed8de7011675485ef471ecde7805a6a98

Observation 8585932b-95fd-4e64-8168-4765ac6904cd · outbound

This paper cites Training Verifiers to Solve Math Word Problems.

Enhancing Reasoning Capabilities in SLMs with Reward Guided Dataset Distillation Training Verifiers to Solve Math Word Problems

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-06T22:45:00.383119Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T22:45:00.383119Z digest=sha256:0f963ad8822e26d1cbc96422e1e73c610b084615043ce841d67e82ccbaeb90ee

Observation ae03c8ce-0b48-4fad-9d78-2c9f0baa78ed · outbound

This paper cites \ Ngo, C.

Enhancing Reasoning Capabilities in SLMs with Reward Guided Dataset Distillation \ Ngo, C

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:45:10.872019Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-06T22:45:00.432702Z digest=sha256:ae071ca97ae0e4a24653d797d06893538abd35e8fe5d481256c6bdf254387f5d

Observation b0065cb9-5d07-4763-8d75-6ddc406cc9e4 · outbound

This paper cites , See, A.

Enhancing Reasoning Capabilities in SLMs with Reward Guided Dataset Distillation , See, A

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:45:10.714160Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-06T22:45:00.499454Z digest=sha256:9cf836780ccd1e7f442c212f760ceac1f7ff5862813e5800599f56dc4897ef26

Observation fda81fd6-744d-48b2-a6d0-d940a7757342 · outbound

This paper cites , Feng, B.

Enhancing Reasoning Capabilities in SLMs with Reward Guided Dataset Distillation , Feng, B

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:45:10.515033Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-06T22:45:00.562770Z digest=sha256:1e34cf87ef900ea8756df4ede43fad046b30d21448960a6e8e1d3c7a55d5642a

Observation 88c34840-15c4-4128-afbb-d71d4410be5d · outbound

This paper cites an unresolved cited work.

Enhancing Reasoning Capabilities in SLMs with Reward Guided Dataset Distillation Unresolved cited work

Reference 12

Resolution
unresolved
raw_fallback, observed 2026-08-06T22:45:10.354570Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-06T22:45:00.644914Z digest=sha256:29f4e083ae563a2d1bd1f26152db30998b16a2198210c3f9dcf9392b4382dc28

Observation 8a236ce7-78d9-45d5-9e9a-3c7e6eb7b7cd · outbound

This paper cites an unresolved cited work.

Enhancing Reasoning Capabilities in SLMs with Reward Guided Dataset Distillation Unresolved cited work

Reference 13

Resolution
unresolved
raw_fallback, observed 2026-08-06T22:45:10.118810Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-06T22:45:00.717936Z digest=sha256:8dad00b22dfaf5ab6734b1d8b73c9cf0ea358f95feb03d61212bfa5937410229

Observation 8af77e78-1dab-42c3-bf88-2b07b2c7012b · outbound

This paper cites Advantage-Guided Distillation for Preference Alignment in Small Language Models.

Enhancing Reasoning Capabilities in SLMs with Reward Guided Dataset Distillation Advantage-Guided Distillation for Preference Alignment in Small Language Models

Reference 14

Resolution
verified exact
local_arxiv, observed 2026-08-06T22:45:05.664522Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-06T22:45:00.774715Z digest=sha256:db90c55bfe9cd0b392bafed1b17c806658e9c15da4149165bc2ddb002d406d76

Observation ba453a7e-76aa-4268-970c-cc687a2a4b18 · outbound

This paper cites MiniLLM: On-Policy Distillation of Large Language Models.

Enhancing Reasoning Capabilities in SLMs with Reward Guided Dataset Distillation MiniLLM: On-Policy Distillation of Large Language Models

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-06T22:45:00.840611Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T22:45:00.840611Z digest=sha256:67f282d663404dbfa92c065978e3176eaacf93f9943b09594532fb809795402e

Observation ec3ea647-c2f7-4d1b-bdb7-e7e74b0b621e · outbound

This paper cites , Zhang, L L.

Enhancing Reasoning Capabilities in SLMs with Reward Guided Dataset Distillation , Zhang, L L

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:45:09.899083Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-06T22:45:00.916470Z digest=sha256:1033f2af9e08c2c4e7b7f73af26a02f61d315f475bbcad0ea6c28271f0456e12

Observation 3a330c4d-5b4f-4e1d-91db-ba0f3d6215db · outbound

This paper cites DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning.

Enhancing Reasoning Capabilities in SLMs with Reward Guided Dataset Distillation DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-06T22:45:00.969331Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T22:45:00.969331Z digest=sha256:70f194743530221567c97104c897a53b7fc6b6c1833f58dd3d26129e22934fbb

Observation 2cc95be0-a373-4d5d-89c2-100f15211594 · outbound

This paper cites , Vinyals, O.

Enhancing Reasoning Capabilities in SLMs with Reward Guided Dataset Distillation , Vinyals, O

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:45:09.681927Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-06T22:45:01.034051Z digest=sha256:ec21ebf0a5d31a630ad26608ce9283dfe64fd47eb59c188a5beaff2b84d865b5

Observation 2988ec95-7570-4805-86fe-20b740f166c6 · outbound

This paper cites , Borgeaud, S.

Enhancing Reasoning Capabilities in SLMs with Reward Guided Dataset Distillation , Borgeaud, S

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:45:09.484507Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-06T22:45:01.104215Z digest=sha256:a6dd4c3cd3c3a193798e7c6938497fc749bd0a8ac906ed01833c44fcdb155446

Observation 38769f70-6b26-4119-a44c-9bc87c2dcc74 · outbound

This paper cites , Li, C L.

Enhancing Reasoning Capabilities in SLMs with Reward Guided Dataset Distillation , Li, C L

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:45:09.258428Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-06T22:45:01.208161Z digest=sha256:6a1314595505627a1a394cd0adf9504be3803d2f364cdb5515756d92fa3015dd

Observation b2223c7a-1689-484d-a335-981e059a7d3c · outbound

This paper cites , Yin, Y.

Enhancing Reasoning Capabilities in SLMs with Reward Guided Dataset Distillation , Yin, Y

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:45:09.049575Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-06T22:45:01.282141Z digest=sha256:33a197ddfa68bca822dfa0fe69f1769b041214cf2e9e1fb0bfbf2fd48cc9c8e9

Observation c63aaa27-aeda-4fdf-b195-07a0ea7ca4b8 · outbound

This paper cites MindStar: Enhancing Math Reasoning in Pre-trained LLMs at Inference Time.

Enhancing Reasoning Capabilities in SLMs with Reward Guided Dataset Distillation MindStar: Enhancing Math Reasoning in Pre-trained LLMs at Inference Time

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-06T22:45:01.362923Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T22:45:01.362923Z digest=sha256:b1d6887c35bbff9d35ebe76b7090e03f63155ffada84355b561feb9c77580ebe

Observation ba9ae507-d70d-49b0-9c49-f1052c19a884 · outbound

This paper cites \ Rush, A M.

Enhancing Reasoning Capabilities in SLMs with Reward Guided Dataset Distillation \ Rush, A M

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:45:08.841836Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-06T22:45:01.446884Z digest=sha256:fbfc945ffb971b9a334a47aa67f11f4db9a2bc9fe37c5c5e751e5eb13864164c

Observation d76b30f7-97b5-49f2-8da9-86c67609c337 · outbound

This paper cites DistiLLM-2: A Contrastive Approach Boosts the Distillation of LLMs.

Enhancing Reasoning Capabilities in SLMs with Reward Guided Dataset Distillation DistiLLM-2: A Contrastive Approach Boosts the Distillation of LLMs

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-06T22:45:01.511764Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T22:45:01.511764Z digest=sha256:0fc74332e37574a3389a4a0b9ae4ca18e58e569c54f47133a9e99755bedddcdb

Observation bf1372cb-2c2d-4815-b264-30df58d36e6d · outbound

This paper cites DistiLLM: Towards Streamlined Distillation for Large Language Models.

Enhancing Reasoning Capabilities in SLMs with Reward Guided Dataset Distillation DistiLLM: Towards Streamlined Distillation for Large Language Models

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-06T22:45:01.602702Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T22:45:01.602702Z digest=sha256:4a0ecc3c561178599ff89d929c5998f88affcc22046821820c2bcee862278ea8

Observation bbfeb791-a91d-4c43-9cde-01ea480d017f · outbound

This paper cites , Fang, L.

Enhancing Reasoning Capabilities in SLMs with Reward Guided Dataset Distillation , Fang, L

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:45:08.621314Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-06T22:45:01.662995Z digest=sha256:14c3d3c3fcd82684891cb868fb70b9def7f7efe94df1fcdf51d798fff6c9f758

Observation 1f40c54b-8e47-4293-9125-d763d5183b0d · outbound

This paper cites Winning Big with Small Models: Knowledge Distillation vs. Self-Training for Reducing Hallucination in Product QA Agents.

Enhancing Reasoning Capabilities in SLMs with Reward Guided Dataset Distillation Winning Big with Small Models: Knowledge Distillation vs. Self-Training for Reducing Hallucination in Product QA Agents

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-06T22:45:01.734181Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T22:45:01.734181Z digest=sha256:f1f20016d8ed08d8350a027b923b99ed16f05029444122cd28083c57fa7dbd22

Observation 1f1635b6-be41-4bcb-aa4b-a089978da199 · outbound

This paper cites GSM-Plus: A Comprehensive Benchmark for Evaluating the Robustness of LLMs as Mathematical Problem Solvers.

Enhancing Reasoning Capabilities in SLMs with Reward Guided Dataset Distillation GSM-Plus: A Comprehensive Benchmark for Evaluating the Robustness of LLMs as Mathematical Problem Solvers

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-06T22:45:01.803122Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T22:45:01.803122Z digest=sha256:218c0aea0a4e47e8f95ee2e1ffce93fd5c278fb97225e13ff7288702e677185a

Observation a41021f4-9434-410f-86af-7fde9090a21d · outbound

This paper cites Evolving Knowledge Distillation with Large Language Models and Active Learning.

Enhancing Reasoning Capabilities in SLMs with Reward Guided Dataset Distillation Evolving Knowledge Distillation with Large Language Models and Active Learning

Reference 29

Resolution
verified exact
local_arxiv, observed 2026-08-06T22:45:05.399923Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-06T22:45:01.899416Z digest=sha256:283dda2fdb9f14b59cf05e4a1f69d43319b9baf336c980dc12aa02a9917924e9

Observation c9427546-55e8-4e8e-81c2-4f74cca695e4 · outbound

This paper cites , Chen, C.

Enhancing Reasoning Capabilities in SLMs with Reward Guided Dataset Distillation , Chen, C

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:45:08.535787Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-06T22:45:01.987456Z digest=sha256:41caf52b4008b12b67aa1a8aa1a832306b605fc6ad3b1b9412afd50ad56c1122

Observation 663344f9-7e30-47e0-89bc-9ac38933d316 · outbound

This paper cites , Yang, Z.

Enhancing Reasoning Capabilities in SLMs with Reward Guided Dataset Distillation , Yang, Z

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:45:08.365808Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-06T22:45:02.050079Z digest=sha256:63e313ef08eb5fc831c869289d5011c6d3776fed2971bed427df3a9e4000d86a

Observation 335802d2-6c7b-402c-8a4f-1c2f5861e9a2 · outbound

This paper cites , Dani, J.

Enhancing Reasoning Capabilities in SLMs with Reward Guided Dataset Distillation , Dani, J

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:45:08.178769Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-06T22:45:02.176949Z digest=sha256:2912d8682a536ffa4895757f15977f3dc68af480a033beae13dd28e30bf313b1

Observation f6df0b60-2933-495f-86bb-c081bc56fa71 · outbound

This paper cites an unresolved cited work.

Enhancing Reasoning Capabilities in SLMs with Reward Guided Dataset Distillation Unresolved cited work

Reference 33

Resolution
unresolved
raw_fallback, observed 2026-08-06T22:45:08.012500Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-06T22:45:02.296854Z digest=sha256:90b9b5f0dcc2c3a4a12086ed22c76412867cf4047645627cdb689ba00e0969ed

Observation 71d9bb01-9f8b-4ccc-a41a-b8716ccb4285 · outbound

This paper cites , Lerer, A.

Enhancing Reasoning Capabilities in SLMs with Reward Guided Dataset Distillation , Lerer, A

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:45:07.780048Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-06T22:45:02.381998Z digest=sha256:1669a7d407d0192a45f85d7f8ab66089d939332f8c606df7fca05aeafd6ea7be

Observation 710f5486-52fa-4b6d-8121-0e832a6133de · outbound

This paper cites Progressive distillation induces an implicit curriculum.

Enhancing Reasoning Capabilities in SLMs with Reward Guided Dataset Distillation Progressive distillation induces an implicit curriculum

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-06T22:45:02.463708Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T22:45:02.463708Z digest=sha256:a82506a9e5fddc2420f8a687b606e3d90348e7e19a3ed3a51ef67b7c87c5dc13

Observation 0d86e769-ece0-4b0e-8736-d57c1f90975d · outbound

This paper cites , Kim, D.

Enhancing Reasoning Capabilities in SLMs with Reward Guided Dataset Distillation , Kim, D

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:45:07.585952Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-06T22:45:02.552158Z digest=sha256:8d92b2bf4f241cfccd1f6948cc2d3efeefbd2bebfd3000a857490658961fd5c4

Observation 17d3ac87-dd4d-41a7-bd6d-017ad7752bd1 · outbound

This paper cites \ Xie, S.

Enhancing Reasoning Capabilities in SLMs with Reward Guided Dataset Distillation \ Xie, S

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:45:07.347692Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-06T22:45:02.621187Z digest=sha256:450cd2832c1ed4a442895af2ae8ab01df49a9b21f851d1c8b0f6fc2027bf7de6

Observation dedb6313-b2d1-43cd-89cb-46c0ddf3e22b · outbound

This paper cites CourseGPT-zh: an Educational Large Language Model Based on Knowledge Distillation Incorporating Prompt Optimization.

Enhancing Reasoning Capabilities in SLMs with Reward Guided Dataset Distillation CourseGPT-zh: an Educational Large Language Model Based on Knowledge Distillation Incorporating Prompt Optimization

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-06T22:45:02.697368Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T22:45:02.697368Z digest=sha256:e9c19bc7b5e497e1ede0aa37b9d37bb7c0d0847870c6a212e165b9172b04e62d

Observation ab12e167-1817-48ba-a32c-9941bc621f32 · outbound

This paper cites DistilBERT, a distilled version of BERT: smaller, faster, cheaper and lighter.

Enhancing Reasoning Capabilities in SLMs with Reward Guided Dataset Distillation DistilBERT, a distilled version of BERT: smaller, faster, cheaper and lighter

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-06T22:45:02.785656Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T22:45:02.785656Z digest=sha256:183cd636bd0b9f874d30314b872fd97f0a37268760331cbcd29ee624aa596e28

Observation 870979c2-4e34-441d-8bad-e60b4bcbc6c5 · outbound

This paper cites Spurious Rewards: Rethinking Training Signals in RLVR.

Enhancing Reasoning Capabilities in SLMs with Reward Guided Dataset Distillation Spurious Rewards: Rethinking Training Signals in RLVR

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-06T22:45:02.881915Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T22:45:02.881915Z digest=sha256:a92196c8ce02768d8874449f0cf929e4dc88bf1af4166c01f4c97e354cf5ba19

Observation f2c60319-8fc3-427f-901d-8f0079bc5860 · outbound

This paper cites , Wang, P.

Enhancing Reasoning Capabilities in SLMs with Reward Guided Dataset Distillation , Wang, P

Reference 41

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:45:07.207353Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-06T22:45:02.953539Z digest=sha256:127c58e63df6c359111a3f4aef442ae820c830a4b060e41aaa7aadffea2f75b3

Observation 98e0f490-a8be-44cd-9845-3e4ee7a4db71 · outbound

This paper cites The Curse of Recursion: Training on Generated Data Makes Models Forget.

Enhancing Reasoning Capabilities in SLMs with Reward Guided Dataset Distillation The Curse of Recursion: Training on Generated Data Makes Models Forget

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-06T22:45:03.029566Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T22:45:03.029566Z digest=sha256:b31e3bf39ce61c3063274aa0579b5d52bb3a73cb95c6ca51f13b74a1c80dbe9c

Observation 13b7323d-6a46-4c54-9d5e-0eae71e6578a · outbound

This paper cites Gemma 3 Technical Report.

Enhancing Reasoning Capabilities in SLMs with Reward Guided Dataset Distillation Gemma 3 Technical Report

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-06T22:45:03.122039Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T22:45:03.122039Z digest=sha256:beb1cc818f97fa1eb44c85806990010d2fa34b800e22aae9e55310dfe9adeb31

Observation 48454789-9132-4ff1-9fe3-3a143de98d2c · outbound

This paper cites , Riviere, M.

Enhancing Reasoning Capabilities in SLMs with Reward Guided Dataset Distillation , Riviere, M

Reference 44

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:45:07.097328Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-06T22:45:03.183890Z digest=sha256:c0d2f169956ff06df0a2f0d706e8b5cd3ed219945233b0cfedaa5eb7cfcaa869

Observation e8406170-4e51-474e-9ba0-a26b4c29b939 · outbound

This paper cites , Han, Y.

Enhancing Reasoning Capabilities in SLMs with Reward Guided Dataset Distillation , Han, Y

Reference 45

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:45:06.972873Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-06T22:45:03.261446Z digest=sha256:7d98c372deaa2b44f8e54239a8c949fd78938cd77847b6abba2a44e15a36091e

Observation 8a34b152-2464-440c-a5d0-9a7c4e36895b · outbound

This paper cites On Teacher Hacking in Language Model Distillation.

Enhancing Reasoning Capabilities in SLMs with Reward Guided Dataset Distillation On Teacher Hacking in Language Model Distillation

Reference 46

Resolution
verified exact
local_arxiv, observed 2026-08-06T22:45:05.113351Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-06T22:45:03.331296Z digest=sha256:fcce4763d0f7e0e0eab56d9dd3b0dd2a2837a96dfd4a8ed3a4edde31c5ab81ae

Observation 210ff84e-986d-4660-8fd8-a7be4c99448c · outbound

This paper cites Who Taught You That? Tracing Teachers in Model Distillation.

Enhancing Reasoning Capabilities in SLMs with Reward Guided Dataset Distillation Who Taught You That? Tracing Teachers in Model Distillation

Reference 47

Resolution
verified exact
local_arxiv, observed 2026-08-06T22:45:04.943544Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-06T22:45:03.441504Z digest=sha256:f0b598c40a1ad7425b3351bd3dbacb698cc13047fce3f0e814d836f4ac9108e4

Observation 42e753af-4bd2-485d-819e-d73112561ade · outbound

This paper cites , Deng, Y.

Enhancing Reasoning Capabilities in SLMs with Reward Guided Dataset Distillation , Deng, Y

Reference 48

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:45:06.848770Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-06T22:45:03.513056Z digest=sha256:7d4bd23a38e3a274bf351d756ee80ab978fbfd8d13043b55d32a630aefc4bfdb

Observation 5fc159c6-a4f4-462a-8797-6b4e79b3f834 · outbound

This paper cites A Comprehensive Survey of Small Language Models in the Era of Large Language Models: Techniques, Enhancements, Applications, Collaboration with LLMs, and Trustworthiness.

Enhancing Reasoning Capabilities in SLMs with Reward Guided Dataset Distillation A Comprehensive Survey of Small Language Models in the Era of Large Language Models: Techniques, Enhancements, Applications, Collaboration with LLMs, and Trustworthiness

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-06T22:45:03.646532Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T22:45:03.646532Z digest=sha256:09027f31e48f7982dc867d24aa3862a532772b7514530608e77708980953a5f7

Observation 41d95ac2-3c60-4166-a078-0a33e91b1192 · outbound

This paper cites , Zhu, J Y.

Enhancing Reasoning Capabilities in SLMs with Reward Guided Dataset Distillation , Zhu, J Y

Reference 50

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:45:06.722937Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-06T22:45:03.749143Z digest=sha256:6c73949a98a0a5e8a9a42931a0b54c17c02844647c99328a42c19f37998fe9e6

Observation ba61adf8-9c15-4502-b0db-7ff5881e1db4 · outbound

This paper cites an unresolved cited work.

Enhancing Reasoning Capabilities in SLMs with Reward Guided Dataset Distillation Unresolved cited work

Reference 51

Resolution
unresolved
raw_fallback, observed 2026-08-06T22:45:06.608963Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-06T22:45:03.850600Z digest=sha256:a77106f8987c9b4607d72386c92409d5cc02136112c329a783cd3990f35f18be

Observation d3fcc0b3-7961-4740-aadb-991067663937 · outbound

This paper cites , Wang, X.

Enhancing Reasoning Capabilities in SLMs with Reward Guided Dataset Distillation , Wang, X

Reference 52

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:45:06.488819Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-06T22:45:03.920673Z digest=sha256:52dc4fe82e3549114c733f25e549a77264ca7825072548c80d58332cc24f031e

Observation 8eafc91f-9730-40aa-b099-dbda10b62673 · outbound

This paper cites , Bai, H.

Enhancing Reasoning Capabilities in SLMs with Reward Guided Dataset Distillation , Bai, H

Reference 53

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:45:06.376947Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-06T22:45:03.977383Z digest=sha256:26d9aff816eff95c4c0c0c2dc15eec2ad94abae013021a273ed430fa959f18b8

Observation 6290ed46-791e-4ba6-aeac-df36eb0ac94b · outbound

This paper cites Speculative Knowledge Distillation: Bridging the Teacher-Student Gap Through Interleaved Sampling.

Enhancing Reasoning Capabilities in SLMs with Reward Guided Dataset Distillation Speculative Knowledge Distillation: Bridging the Teacher-Student Gap Through Interleaved Sampling

Reference 55

Resolution
unresolved
no resolver link, observed 2026-08-06T22:45:04.144603Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T22:45:04.144603Z digest=sha256:eba31324aa4fb2c0ca86c250f095c54d4e01a2209c02595d3ec3a2b31b28e63d

Observation c9a0cc89-1f21-4f07-be33-6c3063688a6b · outbound

This paper cites Qwen2.5 Technical Report.

Enhancing Reasoning Capabilities in SLMs with Reward Guided Dataset Distillation Qwen2.5 Technical Report

Reference 56

Resolution
unresolved
no resolver link, observed 2026-08-06T22:45:04.222611Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T22:45:04.222611Z digest=sha256:24f21ff2a56db7433787e472154800d79a68e805ed2d1f9e2b55666c2346834b

Observation 5f3339e7-b708-4899-86bf-8386abf8fd21 · outbound

This paper cites Demystifying Long Chain-of-Thought Reasoning in LLMs.

Enhancing Reasoning Capabilities in SLMs with Reward Guided Dataset Distillation Demystifying Long Chain-of-Thought Reasoning in LLMs

Reference 57

Resolution
unresolved
no resolver link, observed 2026-08-06T22:45:04.300085Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T22:45:04.300085Z digest=sha256:5735c85e3701dd3f337e5ab4765b68a162f1a9068d0749f4533645d816f6e13d

Observation 012f0bf2-3f3d-406b-b1c2-e8c678c75d5e · outbound

This paper cites , Wang, C.

Enhancing Reasoning Capabilities in SLMs with Reward Guided Dataset Distillation , Wang, C

Reference 58

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:45:06.275447Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-06T22:45:04.408916Z digest=sha256:72c2f967029c6fc98913fbedd2b34e89b1272dd2df87217cf1f6f0bcab7bcb55

Observation c1d9a273-3749-4648-b408-bc667bbe660f · outbound

This paper cites , Zhu, R.

Enhancing Reasoning Capabilities in SLMs with Reward Guided Dataset Distillation , Zhu, R

Reference 59

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:45:06.151692Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-06T22:45:04.485680Z digest=sha256:8131b08a0025d0d129978fe67da183b4eff78932b4797eb448d14da0daf76daf

Observation d6cefc1b-c2e3-4129-a043-5a2f8a170d5d · outbound

This paper cites , Zhu, R.

Enhancing Reasoning Capabilities in SLMs with Reward Guided Dataset Distillation , Zhu, R

Reference 60

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:45:06.021904Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-06T22:45:04.548728Z digest=sha256:b495e5ac6e39d4f34248b5e8745bcaedd0bb6ab2745b8b14589b1ba8027a1656

Observation 5b25e6b2-6caf-47bb-b638-48ef19505fc7 · outbound

This paper cites , Shen, J.

Enhancing Reasoning Capabilities in SLMs with Reward Guided Dataset Distillation , Shen, J

Reference 61

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:45:05.904741Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-06T22:45:04.612738Z digest=sha256:b430f6a974920ba98c932d906af38a297ba045f919ff29d058b2b53674760807

Observation 5dae83d3-89e7-4d02-807a-4b57195fbd38 · outbound

This paper cites Distill Not Only Data but Also Rewards: Can Smaller Language Models Surpass Larger Ones?.

Enhancing Reasoning Capabilities in SLMs with Reward Guided Dataset Distillation Distill Not Only Data but Also Rewards: Can Smaller Language Models Surpass Larger Ones?

Reference 62

Resolution
unresolved
no resolver link, observed 2026-08-06T22:45:04.656219Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T22:45:04.656219Z digest=sha256:f2e652c2064f263423f4b712ade8b32530cee144da17dce07d6a93d63abe6f7a

Pith citing papers

No inbound Pith citation observations are available.