Pith. sign in

Paper Citation Record · LEDGER

Online Knowledge Distillation with Reward Guidance

As of 12 August 2026, this Paper Citation Record lists 64 of 64 outbound references and 0 inbound Pith citation observations for arXiv:2505.18952.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2505.18952 v1

Coverage vector

measured 64 of 64 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-07T14:27:49.447695Z

measured 64 of 64 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-12T06:34:41.77262+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

64 of 64 outbound references displayed

  • verified exact0
  • verified fuzzy32
  • unresolved32
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 59addfc3-77f6-42dc-9b13-99b28a498885 · outbound

This paper cites Improved algorithms for linear stochastic bandits.

Online Knowledge Distillation with Reward Guidance Improved algorithms for linear stochastic bandits

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-07T14:27:49.172936Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:27:49.172936Z digest=sha256:b6e7b4411d2344f84b4513a316e248508e67d175985310064a5c83dd094bb097

Observation b69b458e-c723-4647-bb7d-58d1371ed75a · outbound

This paper cites Gpt-4 technical report.

Online Knowledge Distillation with Reward Guidance Gpt-4 technical report

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-07T14:27:49.179006Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:27:49.179006Z digest=sha256:699f92ae29828dff65e19cf24dd8daed7dca17ed042b27500b7d1b920427c745

Observation 66c8fd9c-96b9-4e6a-9be4-f58ce783a0ba · outbound

This paper cites On-policy distillation of language models: Learning from self-generated mistakes.

Online Knowledge Distillation with Reward Guidance On-policy distillation of language models: Learning from self-generated mistakes

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-07T14:27:49.184236Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:27:49.184236Z digest=sha256:da0899254e053f5b073e15b32863ee3782ed377cd6d78cd634f7b7ffd9f16645

Observation 6d333c0d-b3e6-439f-b780-8e2532147483 · outbound

This paper cites PaLM 2 Technical Report.

Online Knowledge Distillation with Reward Guidance PaLM 2 Technical Report

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-07T14:27:49.189344Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:27:49.189344Z digest=sha256:17bc5930d6930387a2037e42e38c5f15239ed13cf4a0dd9683d8f51e9027b929

Observation 2282d970-697f-4073-aa62-40c41c0c9510 · outbound

This paper cites Claude 3 family.

Online Knowledge Distillation with Reward Guidance Claude 3 family

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:27:50.139409Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-07T14:27:49.193617Z digest=sha256:d4dc3e7357f5a3b314d97804c24d6a21d3256460d4968bdf1c54e083c9fd7fac

Observation d52f406c-99ba-418d-b2c5-263817954980 · outbound

This paper cites Gpt-4 is openai’s most advanced system, producing safer and more useful responses, 2024.

Online Knowledge Distillation with Reward Guidance Gpt-4 is openai’s most advanced system, producing safer and more useful responses, 2024

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:27:50.127285Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-07T14:27:49.208501Z digest=sha256:80551c30f6b5334acf6224e55381804ae5f4b2d03689f4d57d8f1c06b4e6e1a3

Observation 710f5ccd-f9e6-472b-989a-e103c6851d94 · outbound

This paper cites Value-Incentivized Preference Optimization: A Unified Approach to Online and Offline RLHF.

Online Knowledge Distillation with Reward Guidance Value-Incentivized Preference Optimization: A Unified Approach to Online and Offline RLHF

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-07T14:27:49.213678Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:27:49.213678Z digest=sha256:e074c7bbd006ca00ea70289a34d25b4631a92a5a9be479071a7a5c64a813d0c6

Observation e953fc06-285d-4d15-baab-417c2c7dc0e6 · outbound

This paper cites Knowledge distillation of black-box large language models, 2024.

Online Knowledge Distillation with Reward Guidance Knowledge distillation of black-box large language models, 2024

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:27:50.113522Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-07T14:27:49.219066Z digest=sha256:e908979fae27798c582005f7df8e0dbf2a2dc868ec3786cb1dece83a1600ad54

Observation 72cfc732-41d6-4d40-8a45-21290303f112 · outbound

This paper cites Information-theoretic considerations in batch reinforcement learning.

Online Knowledge Distillation with Reward Guidance Information-theoretic considerations in batch reinforcement learning

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:27:50.100612Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-07T14:27:49.223575Z digest=sha256:374b8f6b81ca85e85d3f77cf42ae1daf53f95ff0fe8f6f79708e0baff0cc4ada

Observation 00ef125b-bfbd-47d0-a74b-e3139451a397 · outbound

This paper cites Distilling knowledge learned in bert for text generation.

Online Knowledge Distillation with Reward Guidance Distilling knowledge learned in bert for text generation

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:27:50.087106Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-07T14:27:49.227911Z digest=sha256:2956a0eb62c21fafaf9ef31a5c71db3573d32504cf79b7e70b8c4cdf6e23d3d0

Observation 4208ee3d-3827-491b-b60f-b8609a9c7a79 · outbound

This paper cites Gonzalez, Ion Stoica, and Eric P.

Online Knowledge Distillation with Reward Guidance Gonzalez, Ion Stoica, and Eric P

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-07T14:27:49.232858Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:27:49.232858Z digest=sha256:a4c16ddf510258f2b8f66259790a31a0dea7b7fa8d100976ead886d2ac8d6ac2

Observation 34ae95c8-2171-4fa9-8927-572d0ea0e9d7 · outbound

This paper cites Deep reinforcement learning from human preferences.

Online Knowledge Distillation with Reward Guidance Deep reinforcement learning from human preferences

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-07T14:27:49.236919Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:27:49.236919Z digest=sha256:cf090940ed4b31b27209845462a59eb2e2f744d61e1b5d36434c8e13f48e5ec7

Observation ca2004ef-5cdb-4b16-8c46-a47a8af4582d · outbound

This paper cites Think you have Solved Question Answering? Try ARC, the AI2 Reasoning Challenge.

Online Knowledge Distillation with Reward Guidance Think you have Solved Question Answering? Try ARC, the AI2 Reasoning Challenge

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-07T14:27:49.240934Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:27:49.240934Z digest=sha256:95becf7294aedfa39f4bab85dee1e17b1e9e1acc3a0ed51dfd16e3e5a587240b

Observation 98b21e71-a03c-41e5-af1b-ba32a9930d78 · outbound

This paper cites Training Verifiers to Solve Math Word Problems.

Online Knowledge Distillation with Reward Guidance Training Verifiers to Solve Math Word Problems

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-07T14:27:49.244981Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:27:49.244981Z digest=sha256:802f46c84712440647f9fb401758c0442522f150c623848ef91716435033864b

Observation fdbadbfd-da79-4e23-a822-dc15546d48e4 · outbound

This paper cites Free dolly: Introducing the world’s first truly open instruction-tuned llm, 2023.

Online Knowledge Distillation with Reward Guidance Free dolly: Introducing the world’s first truly open instruction-tuned llm, 2023

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-07T14:27:49.249279Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:27:49.249279Z digest=sha256:e38564ce650b9ef2d03d1ef830b406a63d567a51924f6eb337408e77084faa0f

Observation f3317dc0-9d1d-4283-97f0-72bd436f7dcc · outbound

This paper cites Ultrafeedback: Boosting language models with scaled ai feedback.

Online Knowledge Distillation with Reward Guidance Ultrafeedback: Boosting language models with scaled ai feedback

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:27:50.052465Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-07T14:27:49.253183Z digest=sha256:0525567c2dbda858e0ad9a8bf4c24fd0dd270e68fd2e116887c3646170175d28

Observation 4d41c677-2518-404b-a14c-a747c279317b · outbound

This paper cites Stochastic linear optimization under bandit feedback.

Online Knowledge Distillation with Reward Guidance Stochastic linear optimization under bandit feedback

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:27:50.040022Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-07T14:27:49.257789Z digest=sha256:b2a5451eb8913ba8d4cf02cef0806f888e79f41afed18c5fe76f4935fe6ea7c6

Observation ce12f250-4f44-48c1-ad61-7200fb8b2439 · outbound

This paper cites Openllama: An open reproduction of llama.

Online Knowledge Distillation with Reward Guidance Openllama: An open reproduction of llama

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:27:50.028443Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-07T14:27:49.262377Z digest=sha256:72d15e2abb2ee511bd11e6960f65a83ab53d01287aa2729563eddf49c105d4d0

Observation ad3fb493-2ba3-4271-a22d-85192fab2dc5 · outbound

This paper cites Minillm: Knowledge distillation of large language models.

Online Knowledge Distillation with Reward Guidance Minillm: Knowledge distillation of large language models

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-07T14:27:49.266807Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:27:49.266807Z digest=sha256:f480821139f19ff94f150231cb48d3b1ea724ba0dbf93a7bc67058660bca406a

Observation 49e448cb-6a11-4e17-ae8d-5ecd6eb7f0f3 · outbound

This paper cites Direct Language Model Alignment from Online AI Feedback.

Online Knowledge Distillation with Reward Guidance Direct Language Model Alignment from Online AI Feedback

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-07T14:27:49.270567Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:27:49.270567Z digest=sha256:465125d19cb2a24ad00a037cd4b8d78a7b8795544bbace59c131a1c65db76cdf

Observation 3afeb083-cbd3-4ad9-8b15-2882493f894d · outbound

This paper cites Measuring massive multitask language understanding.

Online Knowledge Distillation with Reward Guidance Measuring massive multitask language understanding

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-07T14:27:49.275247Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:27:49.275247Z digest=sha256:5e07e202f9d07ae2140518b67362f26287060ce777e441e41e1e94b450fbdda2

Observation afd52b7c-3055-4d4d-84ab-0bea56822713 · outbound

This paper cites Distilling the Knowledge in a Neural Network.

Online Knowledge Distillation with Reward Guidance Distilling the Knowledge in a Neural Network

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-07T14:27:49.279060Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:27:49.279060Z digest=sha256:04d206350d307eab51f4e54e7aaa5595e551e50966fab54493e633d130be6b67

Observation 04deee70-79ff-4e23-b213-0866dca0cf22 · outbound

This paper cites Unnatural instructions: Tuning language models with (almost) no human labor.

Online Knowledge Distillation with Reward Guidance Unnatural instructions: Tuning language models with (almost) no human labor

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:27:50.001512Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-07T14:27:49.283639Z digest=sha256:77847bcbc19ebac5c3c91962e84bedb9f634e49944b1cf9bc6dac5b0cb1f90ae

Observation 095018b2-f406-4c49-a671-cd95e0e668e3 · outbound

This paper cites Distilling step-by-step! outperforming larger language models with less training data and smaller model sizes.

Online Knowledge Distillation with Reward Guidance Distilling step-by-step! outperforming larger language models with less training data and smaller model sizes

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:27:49.983865Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-07T14:27:49.287480Z digest=sha256:ae50ad45154cad816b48cdaf4cd7c575a9b8a4f246b272f932fecbece11b8a32

Observation ddd20eab-7613-46db-b578-35c42979b48d · outbound

This paper cites Adversarial moment-matching distillation of large language models.

Online Knowledge Distillation with Reward Guidance Adversarial moment-matching distillation of large language models

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:27:49.962918Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-07T14:27:49.295167Z digest=sha256:d8d2231bd810a41cbe2e5efc0548dda023ffdf46de13af9d62456cc1cfbe6335

Observation 9d3ce41b-3151-416d-944b-76ebfa0b15a2 · outbound

This paper cites an unresolved cited work.

Online Knowledge Distillation with Reward Guidance Unresolved cited work

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-07T14:27:49.291106Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:27:49.291106Z digest=sha256:ceda3aca4870e3b38a86f41fa382e65019d8ea255a8e7a388757d9124151f2e6

Observation b7de2377-19f5-4805-8bd6-fb1023fece96 · outbound

This paper cites Sequence-level knowledge distillation.

Online Knowledge Distillation with Reward Guidance Sequence-level knowledge distillation

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:27:49.935337Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-07T14:27:49.303395Z digest=sha256:72154705b5a7df15efb8aca32563fc80a36184f663c573ec575bf7be76d9c6a7

Observation f367807d-13ac-45ef-b61b-4af49f4a7255 · outbound

This paper cites Tinybert: Distilling bert for natural language understanding.

Online Knowledge Distillation with Reward Guidance Tinybert: Distilling bert for natural language understanding

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:27:49.948970Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-07T14:27:49.299174Z digest=sha256:fe6eb6bf0f986889998eedb5a8c57a43be7f366200adfced7c2cca5368731022

Observation 5fbb818f-063b-466e-85ef-e368a5800469 · outbound

This paper cites Direct Preference Knowledge Distillation for Large Language Models.

Online Knowledge Distillation with Reward Guidance Direct Preference Knowledge Distillation for Large Language Models

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-07T14:27:49.311523Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:27:49.311523Z digest=sha256:37af5fb081fd27770db3285dc659fc5622f22238c50068214a382910812d5a4a

Observation 28d439b1-116d-4074-9e5d-17e9966d8356 · outbound

This paper cites Distillm: Towards streamlined distillation for large language models.

Online Knowledge Distillation with Reward Guidance Distillm: Towards streamlined distillation for large language models

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:27:49.923252Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-07T14:27:49.307693Z digest=sha256:bd816eb9d579d984e5a252ca7b5493a2b27c9bbdea43608bfbd32ee9d0d90f6b

Observation 59d05b33-e740-40c0-bf04-d1503d336845 · outbound

This paper cites Autoregressive knowledge distillation through imitation learning.

Online Knowledge Distillation with Reward Guidance Autoregressive knowledge distillation through imitation learning

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:27:49.899697Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-07T14:27:49.319541Z digest=sha256:aeb53149e1d021578f7d549c80d4d840866f5dddaf4ce90c0e7b14ec6b8c960f

Observation a062182a-381a-4113-a7ca-2b9a8f9fb9fa · outbound

This paper cites Openorca: An open dataset of gpt augmented flan reasoning traces, 2023.

Online Knowledge Distillation with Reward Guidance Openorca: An open dataset of gpt augmented flan reasoning traces, 2023

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:27:49.911217Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-07T14:27:49.315912Z digest=sha256:8404161f6dbdfb763770bd2e5da3e5fed1bbaeb55d4e77e682605d3f4a5ce38e

Observation 2a94db03-1fd9-44d4-b52f-aaccfd86e1d6 · outbound

This paper cites TinyGSM: achieving >80% on GSM8k with small language models.

Online Knowledge Distillation with Reward Guidance TinyGSM: achieving >80% on GSM8k with small language models

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-07T14:27:49.328172Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:27:49.328172Z digest=sha256:d47c396db7422267eb3baaa5a639b6810f9f755b2f51993597527ad4fdb9dd46

Observation 0b1527a5-863c-4975-9caf-eb50e87ce8ad · outbound

This paper cites Rouge: A package for automatic evaluation of summaries.

Online Knowledge Distillation with Reward Guidance Rouge: A package for automatic evaluation of summaries

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:27:49.887728Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-07T14:27:49.324021Z digest=sha256:fe7ee3dff7d1a3fab381655158f734dd633c658fb2018ca43b1625afc5e06b93

Observation d77b095e-aa7a-44bf-bc36-b8993ea832d3 · outbound

This paper cites Training language models to follow instructions with human feedback.

Online Knowledge Distillation with Reward Guidance Training language models to follow instructions with human feedback

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-07T14:27:49.335177Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:27:49.335177Z digest=sha256:ef4c184bd2b3836b3bd4c0cd47bc8b8a946067eddcc47bf1a435b4aa53fafecd

Observation c400cc17-38ff-4af9-9e18-eb952fb59898 · outbound

This paper cites Orca: Progressive Learning from Complex Explanation Traces of GPT-4.

Online Knowledge Distillation with Reward Guidance Orca: Progressive Learning from Complex Explanation Traces of GPT-4

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-07T14:27:49.331916Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:27:49.331916Z digest=sha256:81811862d5748825a708aedef4f1a3cfd19ec08e011aaf1ec850714c2a2cca8a

Observation b82c11d5-5432-4c83-9329-afdfa5b23c4e · outbound

This paper cites Linearly parameterized bandits.

Online Knowledge Distillation with Reward Guidance Linearly parameterized bandits

Reference 38

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:27:49.866336Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-07T14:27:49.343111Z digest=sha256:9d388e7cbe1036ca676158fc51340be539c9b39835b7c354934375cb33f948aa

Observation d098173b-6f47-467a-becf-6decefae2354 · outbound

This paper cites Iterative Reasoning Preference Optimization.

Online Knowledge Distillation with Reward Guidance Iterative Reasoning Preference Optimization

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-07T14:27:49.338434Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:27:49.338434Z digest=sha256:c35bb8de869278191573e46650c967605ba4768e5e3bbde33bd769e7a19b018d

Observation 55a182ff-d4d1-48db-b5c8-ccd0febaee66 · outbound

This paper cites Hybrid rl: Using both offline and online data can make rl efficient.

Online Knowledge Distillation with Reward Guidance Hybrid rl: Using both offline and online data can make rl efficient

Reference 40

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:27:49.852660Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-07T14:27:49.351084Z digest=sha256:d5f9ffd821aea1a05e8d726ec8169a5c7e92c3c4b8cf7d8d9846c1db710d0e1a

Observation cf70783c-9ce3-47e1-a898-1a34128b8653 · outbound

This paper cites Proximal Policy Optimization Algorithms.

Online Knowledge Distillation with Reward Guidance Proximal Policy Optimization Algorithms

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-07T14:27:49.346880Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:27:49.346880Z digest=sha256:7c41e7de11996c61f698153f683bdcc0dc56b1d41d21003db5367c8ef9a33e71

Observation 6f562761-1290-4e74-8baf-5206bf01bc0a · outbound

This paper cites Patient knowledge distillation for bert model compression.

Online Knowledge Distillation with Reward Guidance Patient knowledge distillation for bert model compression

Reference 42

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:27:49.827919Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-07T14:27:49.361088Z digest=sha256:d34b88a7038ccd6d471abdde23917f99f5301d25fdc14354d78982fe6124a348

Observation 9f0d88d0-0d9f-4a1c-8562-0d5b05c465e6 · outbound

This paper cites Learning to summarize with human feedback.

Online Knowledge Distillation with Reward Guidance Learning to summarize with human feedback

Reference 43

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:27:49.840182Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-07T14:27:49.356124Z digest=sha256:6b85e2944ae3d334fd667d5b7b463ea298353a53e775445284a3c043dbc67141

Observation 8bf6598a-24a4-4fb3-a270-8dd9a7bb1ea4 · outbound

This paper cites Of moments and match- ing: A game-theoretic framework for closing the imitation gap.

Online Knowledge Distillation with Reward Guidance Of moments and match- ing: A game-theoretic framework for closing the imitation gap

Reference 44

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:27:49.803933Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-07T14:27:49.369280Z digest=sha256:055210d988a212478f8cd74ae6e5a8eb86db87b5f2c9acd183d5af3f27223527

Observation 56170899-96d4-424b-831d-609f23fc7305 · outbound

This paper cites Challenging big-bench tasks and whether chain-of-thought can solve them.

Online Knowledge Distillation with Reward Guidance Challenging big-bench tasks and whether chain-of-thought can solve them

Reference 45

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:27:49.816354Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-07T14:27:49.365451Z digest=sha256:2348248b2fe68c0ed395381bef33f3a0bf7ebd641fdd8027da718da1faeca7ca

Observation e2523298-82ef-452f-aff5-ac7a8ab597cb · outbound

This paper cites Llama 2: Open Foundation and Fine-Tuned Chat Models.

Online Knowledge Distillation with Reward Guidance Llama 2: Open Foundation and Fine-Tuned Chat Models

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-07T14:27:49.379150Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:27:49.379150Z digest=sha256:0af460217714ca7fa339d804a94b5ab874021048a626ccf02a934f4e4d296dfb

Observation 79e79d0a-3303-4470-a879-5698d3738d7c · outbound

This paper cites Hashimoto.

Online Knowledge Distillation with Reward Guidance Hashimoto

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-07T14:27:49.374653Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:27:49.374653Z digest=sha256:d58889617e8da9aa7f07958b2ed0fb0313957437a55dd80186420fdff6c9fba4

Observation bbe4ff12-43fe-4232-bcd4-e15dd3add8bc · outbound

This paper cites Minilm: Deep self-attention distillation for task-agnostic compression of pre-trained transformers.

Online Knowledge Distillation with Reward Guidance Minilm: Deep self-attention distillation for task-agnostic compression of pre-trained transformers

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-07T14:27:49.387131Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:27:49.387131Z digest=sha256:7dbc76c9b6609166e1f34fc2c3d252a91051c72a787c132752d8e9319c4627d4

Observation 30b4d144-9a8c-4fbf-88ff-0a0d427ae1d2 · outbound

This paper cites Selective knowledge distillation for neural machine translation.

Online Knowledge Distillation with Reward Guidance Selective knowledge distillation for neural machine translation

Reference 49

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:27:49.783300Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-07T14:27:49.382987Z digest=sha256:8d02475f84706971aa79caff1f41cbe322de0f06093a568347ea57bafce8ee4f

Observation 7f0a7e71-d80b-4079-9c28-82e5807ba4a4 · outbound

This paper cites Self-instruct: Aligning language models with self-generated in- structions.

Online Knowledge Distillation with Reward Guidance Self-instruct: Aligning language models with self-generated in- structions

Reference 50

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:27:49.751385Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-07T14:27:49.395277Z digest=sha256:d2d2217c86f5532832ebdbf53dbbc93931be059bb76d189ce5416e94890155ef

Observation ea241061-40c4-49a8-be59-4fa3a0749da7 · outbound

This paper cites Smith, Daniel Khashabi, and Hannaneh Hajishirzi.

Online Knowledge Distillation with Reward Guidance Smith, Daniel Khashabi, and Hannaneh Hajishirzi

Reference 51

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:27:49.764237Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-07T14:27:49.391263Z digest=sha256:33930ba06741640c75c0f288efbc3de0c8b48eb51f989af87568256996ebee6b

Observation a7fd4a4a-8696-4e12-878d-91962d7ac8ec · outbound

This paper cites f-divergence minimization for sequence-level knowledge distillation.

Online Knowledge Distillation with Reward Guidance f-divergence minimization for sequence-level knowledge distillation

Reference 52

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:27:49.726555Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-07T14:27:49.403615Z digest=sha256:cff8790ea6b4540b2e9073edff633cd03aa5b38f8a72a831a4981ea7459de85e

Observation 9193a10d-b238-4dfa-b57a-18285898ac5b · outbound

This paper cites Super-naturalinstructions: Generalization via declarative instructions on 1600+ nlp tasks.

Online Knowledge Distillation with Reward Guidance Super-naturalinstructions: Generalization via declarative instructions on 1600+ nlp tasks

Reference 53

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:27:49.739753Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-07T14:27:49.399435Z digest=sha256:e14afa9ca63d2ccda2b35527bfea7645e36994d78081690d08a47df012bda82b

Observation 31f262d2-7e67-4908-bc9b-f3c41959ec52 · outbound

This paper cites Qwen2 Technical Report.

Online Knowledge Distillation with Reward Guidance Qwen2 Technical Report

Reference 54

Resolution
unresolved
no resolver link, observed 2026-08-07T14:27:49.412002Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:27:49.412002Z digest=sha256:61802fcce6ae6f9d57e57bc2f8b7daf147007df9557318c73cb45b47f3edd226

Observation 76c67b11-d524-4152-8e7a-03e5b443adc0 · outbound

This paper cites Iterative preference learning from human feedback: Bridging theory and practice for rlhf under kl-constraint.

Online Knowledge Distillation with Reward Guidance Iterative preference learning from human feedback: Bridging theory and practice for rlhf under kl-constraint

Reference 55

Resolution
unresolved
no resolver link, observed 2026-08-07T14:27:49.408130Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:27:49.408130Z digest=sha256:bac582fcaf053fe9c03e1177e20b73c03a6e69f550e962ab88b09759451df001

Observation c770c7b2-3bab-4173-b871-f675889b37eb · outbound

This paper cites Provable offline preference-based reinforcement learning.

Online Knowledge Distillation with Reward Guidance Provable offline preference-based reinforcement learning

Reference 56

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:27:49.694198Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-07T14:27:49.420593Z digest=sha256:d86dc661311aeb44a3c4bccc9de6487c6d0f6fd8e956ab602c3e6b56e575e25f

Observation 72392bfe-04c8-4e01-94bd-8013d2fedae8 · outbound

This paper cites Online iterative reinforcement learning from human feedback with general preference model.

Online Knowledge Distillation with Reward Guidance Online iterative reinforcement learning from human feedback with general preference model

Reference 57

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:27:49.707419Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-07T14:27:49.415999Z digest=sha256:ce8e3698dd9a112756be5361f30587aca4a1db1046178d0bb4b274d6d4d6eae9

Observation 10a3c533-330a-43cd-8bf4-351f10f6a8fd · outbound

This paper cites Plad: Preference-based large language model distillation with pseudo-preference pairs.

Online Knowledge Distillation with Reward Guidance Plad: Preference-based large language model distillation with pseudo-preference pairs

Reference 58

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:27:49.682195Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-07T14:27:49.428701Z digest=sha256:6e4ba9de560eab797b0123bc8151fdcf3098b28fa32369b27a53cd59f01dd18f

Observation a52340ce-2c6a-4184-ad7a-2a7213c3799b · outbound

This paper cites TinyLlama: An Open-Source Small Language Model.

Online Knowledge Distillation with Reward Guidance TinyLlama: An Open-Source Small Language Model

Reference 59

Resolution
unresolved
no resolver link, observed 2026-08-07T14:27:49.424641Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:27:49.424641Z digest=sha256:c1aa4e3a2c6afefe625956d5d4876ea75eca43228b247612ba1a69b48c391f82

Observation 84784f89-109a-40bb-bd27-586a98d19ecc · outbound

This paper cites Judging llm-as-a-judge with mt-bench and chatbot arena.

Online Knowledge Distillation with Reward Guidance Judging llm-as-a-judge with mt-bench and chatbot arena

Reference 60

Resolution
unresolved
no resolver link, observed 2026-08-07T14:27:49.436721Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:27:49.436721Z digest=sha256:4167bfe21be3d915a7db3016d66f136e7261c588a18f92c918261eb271542d79

Observation 3f336086-d72a-4518-ae58-2840ed8c28f1 · outbound

This paper cites Mathematical analysis of machine learning algorithms.

Online Knowledge Distillation with Reward Guidance Mathematical analysis of machine learning algorithms

Reference 61

Resolution
unresolved
no resolver link, observed 2026-08-07T14:27:49.432511Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:27:49.432511Z digest=sha256:53c9facac8a9562b7f1edf1f7b9afd133c79ae116e00267b34070141510ff57f

Observation 8e007a6a-5f47-487e-b2f2-221b56409059 · outbound

This paper cites Starling-7b: Improving llm helpfulness & harmlessness with rlaif, November 2023.

Online Knowledge Distillation with Reward Guidance Starling-7b: Improving llm helpfulness & harmlessness with rlaif, November 2023

Reference 62

Resolution
unresolved
no resolver link, observed 2026-08-07T14:27:49.443540Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:27:49.443540Z digest=sha256:21fbbf9a5cbf82a4e1366e88a1c8e256f41b6aa636c95cea4e425d264b4ae3e3

Observation 7d12c72a-b00d-46dc-b77e-05f94e5f29cd · outbound

This paper cites Agieval: A human-centric benchmark for evaluating foundation models.

Online Knowledge Distillation with Reward Guidance Agieval: A human-centric benchmark for evaluating foundation models

Reference 63

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:27:49.654744Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-07T14:27:49.440054Z digest=sha256:ccc0e1c8121e2ccd534607d8aad5605b1ed6762779ad3f36b350d73f06b73246

Observation 1dd11a70-1af2-4288-9ece-00b737f417f7 · outbound

This paper cites Fine-Tuning Language Models from Human Preferences.

Online Knowledge Distillation with Reward Guidance Fine-Tuning Language Models from Human Preferences

Reference 65

Resolution
unresolved
no resolver link, observed 2026-08-07T14:27:49.447695Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:27:49.447695Z digest=sha256:7d7981de8fab66f52f7edf6a41fa17ce8938d03171ec059dbe541a6fd146683e

Observation ad94eeee-5986-4c0f-956d-b8a04ca31f6f · outbound

This paper cites Measurements of nematic susceptibility with phase sensitive nuclear magnetic resonance in pulsed strain fields.

Online Knowledge Distillation with Reward Guidance Measurements of nematic susceptibility with phase sensitive nuclear magnetic resonance in pulsed strain fields

Reference 2023

Resolution
unresolved
no resolver link, observed 2026-08-07T14:27:49.203879Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:27:49.203879Z digest=sha256:64da8f8952c83c6dd1916a8aa937ec6ec1a4ba10abed1c37e110e4f4936b7617

Pith citing papers

No inbound Pith citation observations are available.