Pith. sign in

Paper Citation Record · LEDGER

Online Knowledge Distillation with Reward Guidance

As of 11 August 2026, this Paper Citation Record lists 64 of 64 outbound references and 0 inbound Pith citation observations for arXiv:2505.18952.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2505.18952 v1

Coverage vector

measured 64 of 64 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-07T14:27:49.447695Z

measured 64 of 64 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-11T06:34:44.6726+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

64 of 64 outbound references displayed

  • verified exact0
  • verified fuzzy32
  • unresolved32
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 59addfc3-77f6-42dc-9b13-99b28a498885 · outbound

This paper cites Improved algorithms for linear stochastic bandits.

Online Knowledge Distillation with Reward Guidance Improved algorithms for linear stochastic bandits

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-07T14:27:49.172936Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:27:49.172936Z digest=sha256:c10c23dd893bda7f368bca943f2c1c3b353b9ebfceabda8fe6c78565e91e7ce4

Observation b69b458e-c723-4647-bb7d-58d1371ed75a · outbound

This paper cites Gpt-4 technical report.

Online Knowledge Distillation with Reward Guidance Gpt-4 technical report

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-07T14:27:49.179006Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:27:49.179006Z digest=sha256:627112ee41b26103912f7d1536110c6aec78e90d5f5e265224c78675b059da7a

Observation 66c8fd9c-96b9-4e6a-9be4-f58ce783a0ba · outbound

This paper cites On-policy distillation of language models: Learning from self-generated mistakes.

Online Knowledge Distillation with Reward Guidance On-policy distillation of language models: Learning from self-generated mistakes

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-07T14:27:49.184236Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:27:49.184236Z digest=sha256:353051956f1c952a4e5d328b31753d3a56d6e3337c247689e89eadbd67d368d9

Observation 6d333c0d-b3e6-439f-b780-8e2532147483 · outbound

This paper cites PaLM 2 Technical Report.

Online Knowledge Distillation with Reward Guidance PaLM 2 Technical Report

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-07T14:27:49.189344Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:27:49.189344Z digest=sha256:d4cf3688ab4f3a21d939265629a039f4d16edadb5a36723286b6b0073f084840

Observation 2282d970-697f-4073-aa62-40c41c0c9510 · outbound

This paper cites Claude 3 family.

Online Knowledge Distillation with Reward Guidance Claude 3 family

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:27:50.139409Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-07T14:27:49.193617Z digest=sha256:475fd775f7359fac5480cb1c7cc53e7a0788d45694be2c8e7e54f677ff6a6151

Observation d52f406c-99ba-418d-b2c5-263817954980 · outbound

This paper cites Gpt-4 is openai’s most advanced system, producing safer and more useful responses, 2024.

Online Knowledge Distillation with Reward Guidance Gpt-4 is openai’s most advanced system, producing safer and more useful responses, 2024

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:27:50.127285Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-07T14:27:49.208501Z digest=sha256:d8a182f1dc291e71e3b2e2391d7c7f8525351b8901348e3aab02c5988101ef0b

Observation 710f5ccd-f9e6-472b-989a-e103c6851d94 · outbound

This paper cites Value-Incentivized Preference Optimization: A Unified Approach to Online and Offline RLHF.

Online Knowledge Distillation with Reward Guidance Value-Incentivized Preference Optimization: A Unified Approach to Online and Offline RLHF

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-07T14:27:49.213678Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:27:49.213678Z digest=sha256:fc93f7516791db078311319ab2a120520c78225c2883ea2abe678776d30cd563

Observation e953fc06-285d-4d15-baab-417c2c7dc0e6 · outbound

This paper cites Knowledge distillation of black-box large language models, 2024.

Online Knowledge Distillation with Reward Guidance Knowledge distillation of black-box large language models, 2024

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:27:50.113522Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-07T14:27:49.219066Z digest=sha256:a871775e95fc84ca06f08b91e855e74819ef4d7efc7469c7cb8b6346f153ab01

Observation 72cfc732-41d6-4d40-8a45-21290303f112 · outbound

This paper cites Information-theoretic considerations in batch reinforcement learning.

Online Knowledge Distillation with Reward Guidance Information-theoretic considerations in batch reinforcement learning

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:27:50.100612Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-07T14:27:49.223575Z digest=sha256:0bf6f4e56f073af4508f8333f0ccf9bb12fa0947de0e714d76f9d30b86dbdb42

Observation 00ef125b-bfbd-47d0-a74b-e3139451a397 · outbound

This paper cites Distilling knowledge learned in bert for text generation.

Online Knowledge Distillation with Reward Guidance Distilling knowledge learned in bert for text generation

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:27:50.087106Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-07T14:27:49.227911Z digest=sha256:3e4794857462516a19880478e933a3774137e9b2b572ce940b761d0509786b69

Observation 4208ee3d-3827-491b-b60f-b8609a9c7a79 · outbound

This paper cites Gonzalez, Ion Stoica, and Eric P.

Online Knowledge Distillation with Reward Guidance Gonzalez, Ion Stoica, and Eric P

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-07T14:27:49.232858Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:27:49.232858Z digest=sha256:3af5f3bac0313fe0be03a2e67eab5963ca9e129c9d558ee1bbba486302fbbe00

Observation 34ae95c8-2171-4fa9-8927-572d0ea0e9d7 · outbound

This paper cites Deep reinforcement learning from human preferences.

Online Knowledge Distillation with Reward Guidance Deep reinforcement learning from human preferences

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-07T14:27:49.236919Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:27:49.236919Z digest=sha256:93b2089f7ac7d2c96898d97623b4f6ba94f1cf842c22972ca747751002fa11dc

Observation ca2004ef-5cdb-4b16-8c46-a47a8af4582d · outbound

This paper cites Think you have Solved Question Answering? Try ARC, the AI2 Reasoning Challenge.

Online Knowledge Distillation with Reward Guidance Think you have Solved Question Answering? Try ARC, the AI2 Reasoning Challenge

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-07T14:27:49.240934Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:27:49.240934Z digest=sha256:eac59e0a6fc55b9ffec2c0bee19847e1369d6effc2c6999427e4d590705a9fd2

Observation 98b21e71-a03c-41e5-af1b-ba32a9930d78 · outbound

This paper cites Training Verifiers to Solve Math Word Problems.

Online Knowledge Distillation with Reward Guidance Training Verifiers to Solve Math Word Problems

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-07T14:27:49.244981Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:27:49.244981Z digest=sha256:d79d605e3200c07b9894ace138fbe0f83002eed019d3bcd35c89ef7b55a9fc10

Observation fdbadbfd-da79-4e23-a822-dc15546d48e4 · outbound

This paper cites Free dolly: Introducing the world’s first truly open instruction-tuned llm, 2023.

Online Knowledge Distillation with Reward Guidance Free dolly: Introducing the world’s first truly open instruction-tuned llm, 2023

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-07T14:27:49.249279Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:27:49.249279Z digest=sha256:bf798721a435e2cb1525ffc0cfbb5aa889b2a57fde7bdb29a2f75085beec54ec

Observation f3317dc0-9d1d-4283-97f0-72bd436f7dcc · outbound

This paper cites Ultrafeedback: Boosting language models with scaled ai feedback.

Online Knowledge Distillation with Reward Guidance Ultrafeedback: Boosting language models with scaled ai feedback

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:27:50.052465Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-07T14:27:49.253183Z digest=sha256:c3fd886f3ca396509feb9aa7c9d1635b712a2ff799682ec98b1476dbdbf03632

Observation 4d41c677-2518-404b-a14c-a747c279317b · outbound

This paper cites Stochastic linear optimization under bandit feedback.

Online Knowledge Distillation with Reward Guidance Stochastic linear optimization under bandit feedback

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:27:50.040022Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-07T14:27:49.257789Z digest=sha256:3630c2551eee58ba4e754e51bbed77629c5deb99bd5e6e4021d3e35e16ed47ad

Observation ce12f250-4f44-48c1-ad61-7200fb8b2439 · outbound

This paper cites Openllama: An open reproduction of llama.

Online Knowledge Distillation with Reward Guidance Openllama: An open reproduction of llama

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:27:50.028443Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-07T14:27:49.262377Z digest=sha256:ede685a79d7dc1449b1cdcc0f4c08794d380ed72b039a9dffdebb0f1c8a56f55

Observation ad3fb493-2ba3-4271-a22d-85192fab2dc5 · outbound

This paper cites Minillm: Knowledge distillation of large language models.

Online Knowledge Distillation with Reward Guidance Minillm: Knowledge distillation of large language models

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-07T14:27:49.266807Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:27:49.266807Z digest=sha256:affef133f73d5c3be86830fbdd203f16c3f5ac614a302c14ff1119298bafeb9b

Observation 49e448cb-6a11-4e17-ae8d-5ecd6eb7f0f3 · outbound

This paper cites Direct Language Model Alignment from Online AI Feedback.

Online Knowledge Distillation with Reward Guidance Direct Language Model Alignment from Online AI Feedback

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-07T14:27:49.270567Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:27:49.270567Z digest=sha256:15cd9bffaff7bd058046bbe9b313a2d9fd2bc9febfc7a0a2f7b8c79faa7724fb

Observation 3afeb083-cbd3-4ad9-8b15-2882493f894d · outbound

This paper cites Measuring massive multitask language understanding.

Online Knowledge Distillation with Reward Guidance Measuring massive multitask language understanding

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-07T14:27:49.275247Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:27:49.275247Z digest=sha256:d86da549fb64911e0912b15dec25a658c0280bc6d8c8b51149fa809f54818bcf

Observation afd52b7c-3055-4d4d-84ab-0bea56822713 · outbound

This paper cites Distilling the Knowledge in a Neural Network.

Online Knowledge Distillation with Reward Guidance Distilling the Knowledge in a Neural Network

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-07T14:27:49.279060Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:27:49.279060Z digest=sha256:e799223b568421636e06c97f927c3e1aee3e381f8b035100c23c0345dba4401b

Observation 04deee70-79ff-4e23-b213-0866dca0cf22 · outbound

This paper cites Unnatural instructions: Tuning language models with (almost) no human labor.

Online Knowledge Distillation with Reward Guidance Unnatural instructions: Tuning language models with (almost) no human labor

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:27:50.001512Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-07T14:27:49.283639Z digest=sha256:c8e5429485ac4782795a64a5067dae0c5817f69e7e72ecf1a8c0124f8ea3a11a

Observation 095018b2-f406-4c49-a671-cd95e0e668e3 · outbound

This paper cites Distilling step-by-step! outperforming larger language models with less training data and smaller model sizes.

Online Knowledge Distillation with Reward Guidance Distilling step-by-step! outperforming larger language models with less training data and smaller model sizes

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:27:49.983865Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-07T14:27:49.287480Z digest=sha256:e3ef3ae38685b0af15e6d8ff7eee15ddcd1dc8ee2d5513cebcca21508207274d

Observation ddd20eab-7613-46db-b578-35c42979b48d · outbound

This paper cites Adversarial moment-matching distillation of large language models.

Online Knowledge Distillation with Reward Guidance Adversarial moment-matching distillation of large language models

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:27:49.962918Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-07T14:27:49.295167Z digest=sha256:a7549e1dbd8457958b5beaae38533738f3a6cb032cd582c8e63375fc3fabb4c8

Observation 9d3ce41b-3151-416d-944b-76ebfa0b15a2 · outbound

This paper cites an unresolved cited work.

Online Knowledge Distillation with Reward Guidance Unresolved cited work

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-07T14:27:49.291106Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:27:49.291106Z digest=sha256:8fdc3423a2e65e736a74d60322d701b0d15d9d433ad5ce3fd07ae3195bed5251

Observation b7de2377-19f5-4805-8bd6-fb1023fece96 · outbound

This paper cites Sequence-level knowledge distillation.

Online Knowledge Distillation with Reward Guidance Sequence-level knowledge distillation

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:27:49.935337Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-07T14:27:49.303395Z digest=sha256:42e38039cb8bac9960d7377aceb66d87096cf9a33a1692125e2703c255a7cd7e

Observation f367807d-13ac-45ef-b61b-4af49f4a7255 · outbound

This paper cites Tinybert: Distilling bert for natural language understanding.

Online Knowledge Distillation with Reward Guidance Tinybert: Distilling bert for natural language understanding

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:27:49.948970Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-07T14:27:49.299174Z digest=sha256:eaf3ead8ddd07fc97778ad5d652ecc325d1bfd6b39583096ace4afa4aca71d27

Observation 5fbb818f-063b-466e-85ef-e368a5800469 · outbound

This paper cites Direct Preference Knowledge Distillation for Large Language Models.

Online Knowledge Distillation with Reward Guidance Direct Preference Knowledge Distillation for Large Language Models

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-07T14:27:49.311523Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:27:49.311523Z digest=sha256:d2576d5324e3919b180476d7ec64a05d633c8ef37cf4b53f8361ba8db5ed4d45

Observation 28d439b1-116d-4074-9e5d-17e9966d8356 · outbound

This paper cites Distillm: Towards streamlined distillation for large language models.

Online Knowledge Distillation with Reward Guidance Distillm: Towards streamlined distillation for large language models

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:27:49.923252Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-07T14:27:49.307693Z digest=sha256:d9d006297aec64a4aa4a45728803ce8175cde32da9ed9d8f72e53eb015eeffd8

Observation 59d05b33-e740-40c0-bf04-d1503d336845 · outbound

This paper cites Autoregressive knowledge distillation through imitation learning.

Online Knowledge Distillation with Reward Guidance Autoregressive knowledge distillation through imitation learning

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:27:49.899697Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-07T14:27:49.319541Z digest=sha256:061f85f84c7405c1351b56592dc0ece76da1d6d08ab6dc454d4dd68981bd78d8

Observation a062182a-381a-4113-a7ca-2b9a8f9fb9fa · outbound

This paper cites Openorca: An open dataset of gpt augmented flan reasoning traces, 2023.

Online Knowledge Distillation with Reward Guidance Openorca: An open dataset of gpt augmented flan reasoning traces, 2023

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:27:49.911217Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-07T14:27:49.315912Z digest=sha256:5c126c2b69da79de4ef36a1111f9d4c91479216e48591dba97385e84fd2c24c3

Observation 2a94db03-1fd9-44d4-b52f-aaccfd86e1d6 · outbound

This paper cites TinyGSM: achieving >80% on GSM8k with small language models.

Online Knowledge Distillation with Reward Guidance TinyGSM: achieving >80% on GSM8k with small language models

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-07T14:27:49.328172Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:27:49.328172Z digest=sha256:02f15c467cee8f7fe291c94f48ab06af7584f688d49f67c3d7d1d603644f7c3f

Observation 0b1527a5-863c-4975-9caf-eb50e87ce8ad · outbound

This paper cites Rouge: A package for automatic evaluation of summaries.

Online Knowledge Distillation with Reward Guidance Rouge: A package for automatic evaluation of summaries

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:27:49.887728Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-07T14:27:49.324021Z digest=sha256:ef51a418ae03afd481751a7077bea15a9a236d06a543975d93a1ede02d233c2f

Observation d77b095e-aa7a-44bf-bc36-b8993ea832d3 · outbound

This paper cites Training language models to follow instructions with human feedback.

Online Knowledge Distillation with Reward Guidance Training language models to follow instructions with human feedback

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-07T14:27:49.335177Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:27:49.335177Z digest=sha256:da8282e17f3ff5d137b7ae25b19901c1f891e3e902b8c82dcabf956ac6f10d11

Observation c400cc17-38ff-4af9-9e18-eb952fb59898 · outbound

This paper cites Orca: Progressive Learning from Complex Explanation Traces of GPT-4.

Online Knowledge Distillation with Reward Guidance Orca: Progressive Learning from Complex Explanation Traces of GPT-4

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-07T14:27:49.331916Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:27:49.331916Z digest=sha256:6c832c55afaa62be8271edcbe06cbe75a44e2b11ff20fd743ad912df50940e4d

Observation b82c11d5-5432-4c83-9329-afdfa5b23c4e · outbound

This paper cites Linearly parameterized bandits.

Online Knowledge Distillation with Reward Guidance Linearly parameterized bandits

Reference 38

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:27:49.866336Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-07T14:27:49.343111Z digest=sha256:6652de63d1cf0bd015505c9fd0d39e53c7581a22afe6180178a91876ccc759dd

Observation d098173b-6f47-467a-becf-6decefae2354 · outbound

This paper cites Iterative Reasoning Preference Optimization.

Online Knowledge Distillation with Reward Guidance Iterative Reasoning Preference Optimization

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-07T14:27:49.338434Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:27:49.338434Z digest=sha256:0cf07fcac8167ee95080be8310b87b8e52a081d778fdf65fcebf576b66309535

Observation 55a182ff-d4d1-48db-b5c8-ccd0febaee66 · outbound

This paper cites Hybrid rl: Using both offline and online data can make rl efficient.

Online Knowledge Distillation with Reward Guidance Hybrid rl: Using both offline and online data can make rl efficient

Reference 40

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:27:49.852660Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-07T14:27:49.351084Z digest=sha256:d17a3314ee2cf64483a1b2498ca11d8ad542cff0308a0bf6830d7737218ce101

Observation cf70783c-9ce3-47e1-a898-1a34128b8653 · outbound

This paper cites Proximal Policy Optimization Algorithms.

Online Knowledge Distillation with Reward Guidance Proximal Policy Optimization Algorithms

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-07T14:27:49.346880Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:27:49.346880Z digest=sha256:8ad9651ca4842cf530cd13340d4224b9f2a01c10b2453d7e68aef9db4698b8ad

Observation 6f562761-1290-4e74-8baf-5206bf01bc0a · outbound

This paper cites Patient knowledge distillation for bert model compression.

Online Knowledge Distillation with Reward Guidance Patient knowledge distillation for bert model compression

Reference 42

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:27:49.827919Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-07T14:27:49.361088Z digest=sha256:5f6b9feaf96b3188c66ec55860cd1d13791960d7bd9a895f57d7b39321a10a39

Observation 9f0d88d0-0d9f-4a1c-8562-0d5b05c465e6 · outbound

This paper cites Learning to summarize with human feedback.

Online Knowledge Distillation with Reward Guidance Learning to summarize with human feedback

Reference 43

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:27:49.840182Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-07T14:27:49.356124Z digest=sha256:d73da14a42ba28394e3948d1979336ca0845ee764065bcc7cb96ec9748c83677

Observation 8bf6598a-24a4-4fb3-a270-8dd9a7bb1ea4 · outbound

This paper cites Of moments and match- ing: A game-theoretic framework for closing the imitation gap.

Online Knowledge Distillation with Reward Guidance Of moments and match- ing: A game-theoretic framework for closing the imitation gap

Reference 44

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:27:49.803933Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-07T14:27:49.369280Z digest=sha256:c81aea83bda13db82e68fdcba60e98f9fccc426216daffa7fb6d0ad7f7719c6d

Observation 56170899-96d4-424b-831d-609f23fc7305 · outbound

This paper cites Challenging big-bench tasks and whether chain-of-thought can solve them.

Online Knowledge Distillation with Reward Guidance Challenging big-bench tasks and whether chain-of-thought can solve them

Reference 45

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:27:49.816354Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-07T14:27:49.365451Z digest=sha256:29f7568dcd0167ba04e500012625f8be7bb7e8f2f538e19eaa14c2ae48fe1a8d

Observation e2523298-82ef-452f-aff5-ac7a8ab597cb · outbound

This paper cites Llama 2: Open Foundation and Fine-Tuned Chat Models.

Online Knowledge Distillation with Reward Guidance Llama 2: Open Foundation and Fine-Tuned Chat Models

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-07T14:27:49.379150Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:27:49.379150Z digest=sha256:df0c62d8fa1af83d4edf9784063a9fe91b0c0ae2f0c138d962fd40b76fa8c817

Observation 79e79d0a-3303-4470-a879-5698d3738d7c · outbound

This paper cites Hashimoto.

Online Knowledge Distillation with Reward Guidance Hashimoto

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-07T14:27:49.374653Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:27:49.374653Z digest=sha256:004f34a015345017f5b0f6d40bf563a733ffe98a4321c2059c5509cd49d43b92

Observation bbe4ff12-43fe-4232-bcd4-e15dd3add8bc · outbound

This paper cites Minilm: Deep self-attention distillation for task-agnostic compression of pre-trained transformers.

Online Knowledge Distillation with Reward Guidance Minilm: Deep self-attention distillation for task-agnostic compression of pre-trained transformers

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-07T14:27:49.387131Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:27:49.387131Z digest=sha256:f0b1cfd9182b39550be009fd13b7f7a7a4e6acb35edbfd22e0b45f7882692923

Observation 30b4d144-9a8c-4fbf-88ff-0a0d427ae1d2 · outbound

This paper cites Selective knowledge distillation for neural machine translation.

Online Knowledge Distillation with Reward Guidance Selective knowledge distillation for neural machine translation

Reference 49

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:27:49.783300Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-07T14:27:49.382987Z digest=sha256:78276c372d8ddcf88e211af571d7bda67ef7355823f9735e6b40f5a65c43fa81

Observation 7f0a7e71-d80b-4079-9c28-82e5807ba4a4 · outbound

This paper cites Self-instruct: Aligning language models with self-generated in- structions.

Online Knowledge Distillation with Reward Guidance Self-instruct: Aligning language models with self-generated in- structions

Reference 50

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:27:49.751385Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-07T14:27:49.395277Z digest=sha256:08bfb1eadaad1eb2b82437cb9a12af005a2db34bdd4de48529a11f3fd328cf92

Observation ea241061-40c4-49a8-be59-4fa3a0749da7 · outbound

This paper cites Smith, Daniel Khashabi, and Hannaneh Hajishirzi.

Online Knowledge Distillation with Reward Guidance Smith, Daniel Khashabi, and Hannaneh Hajishirzi

Reference 51

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:27:49.764237Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-07T14:27:49.391263Z digest=sha256:cacf227cfcf805e000173251044681e37660afb1af0d89ea1d381ab8ec145d6a

Observation a7fd4a4a-8696-4e12-878d-91962d7ac8ec · outbound

This paper cites f-divergence minimization for sequence-level knowledge distillation.

Online Knowledge Distillation with Reward Guidance f-divergence minimization for sequence-level knowledge distillation

Reference 52

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:27:49.726555Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-07T14:27:49.403615Z digest=sha256:333198c44e42a8aabd11d6a2a91a1807c9aa8c9a0a80959e74e21c96ac806eb1

Observation 9193a10d-b238-4dfa-b57a-18285898ac5b · outbound

This paper cites Super-naturalinstructions: Generalization via declarative instructions on 1600+ nlp tasks.

Online Knowledge Distillation with Reward Guidance Super-naturalinstructions: Generalization via declarative instructions on 1600+ nlp tasks

Reference 53

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:27:49.739753Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-07T14:27:49.399435Z digest=sha256:fc753ad9ca3f8909212c4541c9568f9bef07714d7bb0cf75fc3c3f6c5d337088

Observation 31f262d2-7e67-4908-bc9b-f3c41959ec52 · outbound

This paper cites Qwen2 Technical Report.

Online Knowledge Distillation with Reward Guidance Qwen2 Technical Report

Reference 54

Resolution
unresolved
no resolver link, observed 2026-08-07T14:27:49.412002Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:27:49.412002Z digest=sha256:8b6c179b9d0891eb1db34d5826e97c52c6c247b009e06623df8dcd50c5be6ec8

Observation 76c67b11-d524-4152-8e7a-03e5b443adc0 · outbound

This paper cites Iterative preference learning from human feedback: Bridging theory and practice for rlhf under kl-constraint.

Online Knowledge Distillation with Reward Guidance Iterative preference learning from human feedback: Bridging theory and practice for rlhf under kl-constraint

Reference 55

Resolution
unresolved
no resolver link, observed 2026-08-07T14:27:49.408130Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:27:49.408130Z digest=sha256:45466df429ecf1a500fb918f0fe30a89f34833341af336fb5103c0a8d451138a

Observation c770c7b2-3bab-4173-b871-f675889b37eb · outbound

This paper cites Provable offline preference-based reinforcement learning.

Online Knowledge Distillation with Reward Guidance Provable offline preference-based reinforcement learning

Reference 56

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:27:49.694198Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-07T14:27:49.420593Z digest=sha256:1b7c70a1b428f42bad799e4cc5d0a404a4b103030ee0fb3d463fa55644102131

Observation 72392bfe-04c8-4e01-94bd-8013d2fedae8 · outbound

This paper cites Online iterative reinforcement learning from human feedback with general preference model.

Online Knowledge Distillation with Reward Guidance Online iterative reinforcement learning from human feedback with general preference model

Reference 57

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:27:49.707419Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-07T14:27:49.415999Z digest=sha256:5fe7173db4aabb112fb9023edae2bce334955ed45ed0ba4087c70deccd70ed87

Observation 10a3c533-330a-43cd-8bf4-351f10f6a8fd · outbound

This paper cites Plad: Preference-based large language model distillation with pseudo-preference pairs.

Online Knowledge Distillation with Reward Guidance Plad: Preference-based large language model distillation with pseudo-preference pairs

Reference 58

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:27:49.682195Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-07T14:27:49.428701Z digest=sha256:726ef1d7dd2526939509690ff26b138b05dd25434f79a90a796c76f7949669d9

Observation a52340ce-2c6a-4184-ad7a-2a7213c3799b · outbound

This paper cites TinyLlama: An Open-Source Small Language Model.

Online Knowledge Distillation with Reward Guidance TinyLlama: An Open-Source Small Language Model

Reference 59

Resolution
unresolved
no resolver link, observed 2026-08-07T14:27:49.424641Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:27:49.424641Z digest=sha256:b52f5b6fcfa4504bf40911e271374823356e460fca5fe7016d9cc0d33e6833f2

Observation 84784f89-109a-40bb-bd27-586a98d19ecc · outbound

This paper cites Judging llm-as-a-judge with mt-bench and chatbot arena.

Online Knowledge Distillation with Reward Guidance Judging llm-as-a-judge with mt-bench and chatbot arena

Reference 60

Resolution
unresolved
no resolver link, observed 2026-08-07T14:27:49.436721Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:27:49.436721Z digest=sha256:c7e21ad0555d3df630768323c8472ae320360dd235e332be7e685716f164de4e

Observation 3f336086-d72a-4518-ae58-2840ed8c28f1 · outbound

This paper cites Mathematical analysis of machine learning algorithms.

Online Knowledge Distillation with Reward Guidance Mathematical analysis of machine learning algorithms

Reference 61

Resolution
unresolved
no resolver link, observed 2026-08-07T14:27:49.432511Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:27:49.432511Z digest=sha256:c979ee1f96b92c1dbf70497f805bd1a528a5705bbf76ad4cefbcdaadb7d6c71c

Observation 8e007a6a-5f47-487e-b2f2-221b56409059 · outbound

This paper cites Starling-7b: Improving llm helpfulness & harmlessness with rlaif, November 2023.

Online Knowledge Distillation with Reward Guidance Starling-7b: Improving llm helpfulness & harmlessness with rlaif, November 2023

Reference 62

Resolution
unresolved
no resolver link, observed 2026-08-07T14:27:49.443540Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:27:49.443540Z digest=sha256:a03cd5354cbf9bd83a8ad9e842ef61ecaa6beb5c9d6f1f9f59c9ded66f884267

Observation 7d12c72a-b00d-46dc-b77e-05f94e5f29cd · outbound

This paper cites Agieval: A human-centric benchmark for evaluating foundation models.

Online Knowledge Distillation with Reward Guidance Agieval: A human-centric benchmark for evaluating foundation models

Reference 63

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:27:49.654744Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-07T14:27:49.440054Z digest=sha256:db22d23a8a9c084fffee2beecd8eb9f9ae3153854585de37ae41a67f2d0d0b81

Observation 1dd11a70-1af2-4288-9ece-00b737f417f7 · outbound

This paper cites Fine-Tuning Language Models from Human Preferences.

Online Knowledge Distillation with Reward Guidance Fine-Tuning Language Models from Human Preferences

Reference 65

Resolution
unresolved
no resolver link, observed 2026-08-07T14:27:49.447695Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:27:49.447695Z digest=sha256:dac8df21e65c21334cb0b2609da38f4b8d8842ff941148c66452d2dd79788faa

Observation ad94eeee-5986-4c0f-956d-b8a04ca31f6f · outbound

This paper cites Measurements of nematic susceptibility with phase sensitive nuclear magnetic resonance in pulsed strain fields.

Online Knowledge Distillation with Reward Guidance Measurements of nematic susceptibility with phase sensitive nuclear magnetic resonance in pulsed strain fields

Reference 2023

Resolution
unresolved
no resolver link, observed 2026-08-07T14:27:49.203879Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:27:49.203879Z digest=sha256:7667db59f7bccc828651bd3326d3c1c5f2c0de0b0b9cf6a146b69e6d5632fe30

Pith citing papers

No inbound Pith citation observations are available.