Pith. sign in

Paper Citation Record · LEDGER

Well-Read Students Learn Better: On the Importance of Pre-training Compact Models

As of 8 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 25 inbound Pith citation observations for arXiv:1908.08962.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
1908.08962 v2

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 25 of 25 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-08T06:32:00.761636+00:00

measured 25 of 25 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-07T15:05:15.374000Z

measured 1 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-08-05T02:28:24.338817Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

429
arxiv_reference, observed 2026-08-05T02:28:24.338817Z

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation df108f19-e6b9-407a-9853-abd23d00e31d · inbound

ALBERT: A Lite BERT for Self-supervised Learning of Language Representations cites this paper.

ALBERT: A Lite BERT for Self-supervised Learning of Language Representations Well-Read Students Learn Better: On the Importance of Pre-training Compact Models

Reference 34

Resolution
verified exact
arxiv_id, observed 2026-05-13T12:26:58.123489Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-13T12:26:58.015594Z digest=sha256:e2acc34aed1f10215a601ad0bd8d7a69c12cce17becca136df74e039a3baefbf

Observation 1c9ae76d-aa94-46de-9d6f-6451b731c5b4 · inbound

DistilBERT, a distilled version of BERT: smaller, faster, cheaper and lighter cites this paper.

DistilBERT, a distilled version of BERT: smaller, faster, cheaper and lighter Well-Read Students Learn Better: On the Importance of Pre-training Compact Models

Reference 45

Resolution
verified exact
arxiv_id, observed 2026-05-11T05:02:34.182192Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-05-11T05:02:34.076464Z digest=sha256:deb266037c7669c77babd9d9568c26bf39cf03ed4bac8f58bba9039acd45399b

Observation 46bed6a4-fbd8-40fe-a3c1-6aa47f22cc38 · inbound

Knowledge Distillation in Iterative Generative Models for Improved Sampling Speed cites this paper.

Knowledge Distillation in Iterative Generative Models for Improved Sampling Speed Well-Read Students Learn Better: On the Importance of Pre-training Compact Models

Reference 41

Resolution
verified exact
arxiv_id, observed 2026-05-17T03:56:19.927545Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-17T03:56:19.837017Z digest=sha256:5dbe2a3eb4d27fdca0398cca13792e62e9bd25c1d24bf63430374644b0e78991

Observation f621dc27-ae8f-43be-b1bf-d9caff9b78c8 · inbound

DiffuSeq: Sequence to Sequence Text Generation with Diffusion Models cites this paper.

DiffuSeq: Sequence to Sequence Text Generation with Diffusion Models Well-Read Students Learn Better: On the Importance of Pre-training Compact Models

Reference 9

Resolution
verified exact
arxiv_id, observed 2026-05-20T06:50:38.320346Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-20T06:50:38.258692Z digest=sha256:779ab86eef259880eb8591cd7b58ef5f29ccc441ed5ec8c44a9ac067f2ec9903

Observation 5ec53084-21e3-433c-a1f3-684ac4d331f3 · inbound

Discrepancies are Virtue: Weak-to-Strong Generalization through Lens of Intrinsic Dimension cites this paper.

Discrepancies are Virtue: Weak-to-Strong Generalization through Lens of Intrinsic Dimension Well-Read Students Learn Better: On the Importance of Pre-training Compact Models

Reference 22

Resolution
verified exact
arxiv_id, observed 2026-05-23T03:25:20.537624Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-23T03:24:07.851782Z digest=sha256:95906435b56de00cc1bfee9ef1189aee54894f8b32eb373dce3e803902fe44d6

Observation f79a4b8a-b26f-42a1-b1b4-ab7a545d4d09 · inbound

X-ARES: A Comprehensive Framework for Assessing Audio Encoder Performance cites this paper.

X-ARES: A Comprehensive Framework for Assessing Audio Encoder Performance Well-Read Students Learn Better: On the Importance of Pre-training Compact Models

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-07T15:05:15.374000Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:05:15.374000Z digest=sha256:618b9ec73aa7cd6c6604ae046903a2caf7bd6550a6399fac000686027c584595

Observation 1cf6706a-a972-4eb0-899f-44906e722334 · inbound

ReqBrain: Task-Specific Instruction Tuning of LLMs for AI-Assisted Requirements Generation cites this paper.

ReqBrain: Task-Specific Instruction Tuning of LLMs for AI-Assisted Requirements Generation Well-Read Students Learn Better: On the Importance of Pre-training Compact Models

Reference 67

Resolution
unresolved
no resolver link, observed 2026-08-07T14:48:06.850072Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:48:06.850072Z digest=sha256:3b8213ac2d55840b59eb23dcab5acb23e20b8cb80f5506ae4e14c3e50d77da0c

Observation 8a704e81-ed1e-4f6a-b3fc-de2c013d0472 · inbound

Clustering and Median Aggregation Improve Differentially Private Inference cites this paper.

Clustering and Median Aggregation Improve Differentially Private Inference Well-Read Students Learn Better: On the Importance of Pre-training Compact Models

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-07T10:47:28.375567Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:47:28.375567Z digest=sha256:471a2ff499eb2280946db569cadb9c85bd63c2e4b2528dacadd84d2f93cff090

Observation 45c3ebdf-44e2-4bb3-8b52-ab475807b07d · inbound

Basis Transformers for Multi-Task Tabular Regression cites this paper.

Basis Transformers for Multi-Task Tabular Regression Well-Read Students Learn Better: On the Importance of Pre-training Compact Models

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-07T05:50:18.871863Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:50:18.871863Z digest=sha256:8d5173242f4d47b4ef91a75b670cd3a6cad1faaa2d35c5d0c6367407329fd65b

Observation b7083735-83d9-4ea7-b284-1bb1e179a483 · inbound

GenRecal: Generation after Recalibration from Large to Small Vision-Language Models cites this paper.

GenRecal: Generation after Recalibration from Large to Small Vision-Language Models Well-Read Students Learn Better: On the Importance of Pre-training Compact Models

Reference 100

Resolution
unresolved
no resolver link, observed 2026-08-06T23:57:26.565601Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:57:26.565601Z digest=sha256:6fe9d11ea0cf65734c3f216d0619965fd2db3451f4f1ab4fcda88e06a3fb66bd

Observation 2275fc31-732d-48c8-8c54-46dc8e0b0da4 · inbound

QFFN-BERT: An Empirical Study of Depth, Performance, and Data Efficiency in Hybrid Quantum-Classical Transformers cites this paper.

QFFN-BERT: An Empirical Study of Depth, Performance, and Data Efficiency in Hybrid Quantum-Classical Transformers Well-Read Students Learn Better: On the Importance of Pre-training Compact Models

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-06T20:36:45.789946Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:36:45.789946Z digest=sha256:37202dc0997f87601026eb1688c8504fe7384905e50e01ae9a0bfefd17d5c274

Observation 36da8803-3ba9-4b50-ad02-c6a6d6676f75 · inbound

The Impact of Background Speech on Interruption Detection in Collaborative Groups cites this paper.

The Impact of Background Speech on Interruption Detection in Collaborative Groups Well-Read Students Learn Better: On the Importance of Pre-training Compact Models

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-06T18:49:30.548612Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:49:30.548612Z digest=sha256:0319add1bac86268ba3e1d8d27424c6d376ecbefa24fc732c350c1bf4258c705

Observation 0200a562-5331-4b94-b02d-47c6bea20251 · inbound

Semantic Convergence: Investigating Shared Representations Across Scaled LLMs cites this paper.

Semantic Convergence: Investigating Shared Representations Across Scaled LLMs Well-Read Students Learn Better: On the Importance of Pre-training Compact Models

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-06T15:39:46.566529Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:39:46.566529Z digest=sha256:72b967fe5c4cee24a117cd7bdf4e2549ae6490d0a3ec04abf144142be97a08bd

Observation 6888f70e-d5f5-49e2-9b47-a3bf6254f2f2 · inbound

Compressed Models are NOT Trust-equivalent to Their Large Counterparts cites this paper.

Compressed Models are NOT Trust-equivalent to Their Large Counterparts Well-Read Students Learn Better: On the Importance of Pre-training Compact Models

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-05T19:00:30.108034Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T19:00:30.108034Z digest=sha256:1151018a42b502a1bb5a0d7b33d2ea713d22f7332826df16d96e40aa1f68418b

Observation 8642ce73-9232-43de-adca-746c0bc8e8b0 · inbound

Mitigating Attention Localization in Small Scale: Self-Attention Refinement via One-step Belief Propagation cites this paper.

Mitigating Attention Localization in Small Scale: Self-Attention Refinement via One-step Belief Propagation Well-Read Students Learn Better: On the Importance of Pre-training Compact Models

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-04T22:30:12.153879Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T22:30:12.153879Z digest=sha256:c284760273bbf6fbad7a844702641e4466fb47f6a721e413ef117c3e20571010

Observation 3b23346f-ddd7-47a2-838f-b8cc05b610ce · inbound

Learn from A Rationalist: Distilling Intermediate Interpretable Rationales cites this paper.

Learn from A Rationalist: Distilling Intermediate Interpretable Rationales Well-Read Students Learn Better: On the Importance of Pre-training Compact Models

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-03T06:36:47.013517Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-03T06:36:47.013517Z digest=sha256:6bdc09150f68c9135351991add11eb0db026cee14d94e56124df943a22ae31b1

Observation 2380fa43-c316-4fd3-95ac-7337dde7ea55 · inbound

Transporting Task Vectors across Different Architectures without Training cites this paper.

Transporting Task Vectors across Different Architectures without Training Well-Read Students Learn Better: On the Importance of Pre-training Compact Models

Reference 2024

Resolution
unresolved
no resolver link, observed 2026-08-02T23:43:34.290725Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T23:43:34.290725Z digest=sha256:15752a1699e1ab013e3cffd180487f768c7d11f64ecbce08f5edd7f29db6c4d3

Observation b0f4d5e1-631f-4c05-a16f-886417ac6a32 · inbound

LLM-ODE: Data-driven Discovery of Dynamical Systems with Large Language Models cites this paper.

LLM-ODE: Data-driven Discovery of Dynamical Systems with Large Language Models Well-Read Students Learn Better: On the Importance of Pre-training Compact Models

Reference 35

Resolution
verified exact
arxiv_id, observed 2026-05-15T06:35:09.663608Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-15T06:30:29.297712Z digest=sha256:41fd3e046c5d62c5361c40228c8925c2c0cb6e572ff3fa9dbf5a6c60e14dd16b

Observation 53ce16ed-f4f4-4578-951e-6507980dc056 · inbound

Spectrum-Adaptive Generalization Bounds for Trained Deep Transformers cites this paper.

Spectrum-Adaptive Generalization Bounds for Trained Deep Transformers Well-Read Students Learn Better: On the Importance of Pre-training Compact Models

Reference 27

Resolution
verified exact
arxiv_id, observed 2026-05-11T04:35:58.818603Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-05-11T01:13:25.378895Z digest=sha256:134393197dc15d8fd048b6b16837843f92729eef89140e3999d774351068a5a7

Observation c11bb07a-c02e-4f2a-beba-32b1c119e058 · inbound

Emergent Communication between Heterogeneous Visual Agents through Decentralized Learning cites this paper.

Emergent Communication between Heterogeneous Visual Agents through Decentralized Learning Well-Read Students Learn Better: On the Importance of Pre-training Compact Models

Reference 37

Resolution
verified exact
arxiv_id, observed 2026-05-13T06:42:25.642869Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-13T06:38:48.402207Z digest=sha256:d7f59db5f5c2fcaa5b8a8b8ce3c74030c968e1ed82a090860c6ad294a9705dd4

Observation a9ea690a-c066-4608-8fe2-7cce1f3c4b50 · inbound

PortBERT: Navigating the Depths of Portuguese Language Models cites this paper.

PortBERT: Navigating the Depths of Portuguese Language Models Well-Read Students Learn Better: On the Importance of Pre-training Compact Models

Reference 27

Resolution
metadata mismatch
arxiv_id, observed 2026-07-01T23:06:20.160364Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-06-28T14:43:28.401687Z digest=sha256:1641bee9538d65d33786531404d5952bf9c636017e488bc61371c03d8870088f

Observation b12765b1-cbca-4ba6-a02b-d1d5ba1b9dcc · inbound

DIPBox: A Multi-scale Testing Framework for Tracking Dataset Regeneration cites this paper.

DIPBox: A Multi-scale Testing Framework for Tracking Dataset Regeneration Well-Read Students Learn Better: On the Importance of Pre-training Compact Models

Reference 68

Resolution
verified exact
arxiv_id, observed 2026-07-04T06:39:38.256949Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-06-26T14:15:19.553221Z digest=sha256:1ef9193065b336f08fce0736f68aea74c437e7e1fcdf567a0b80f2672b82774d

Observation ca8b2d3f-9ec1-4b31-b6cc-096821f93aea · inbound

Geometric and Information Compression of Representations in Deep Learning cites this paper.

Geometric and Information Compression of Representations in Deep Learning Well-Read Students Learn Better: On the Importance of Pre-training Compact Models

Reference 49

Resolution
metadata mismatch
arxiv_id, observed 2026-07-04T06:09:36.862161Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-06-26T14:51:57.746071Z digest=sha256:6f65dac61f9408b768a18467866bbea1bf984ff8408262de3a20502d83e7ca56

Observation 8f78a028-d1f0-41c2-9611-f9ce741e7327 · inbound

Geometry-Aware Online Scheduling for LLM Serving: From Theoretical Bound to System Practice cites this paper.

Geometry-Aware Online Scheduling for LLM Serving: From Theoretical Bound to System Practice Well-Read Students Learn Better: On the Importance of Pre-training Compact Models

Reference 19

Resolution
verified exact
arxiv_id, observed 2026-07-04T08:39:41.972877Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-06-26T11:17:35.860286Z digest=sha256:a53220f4e639dba30bf6b034d20cb7f6f9e55b0478215c52bf745d0235b61d95

Observation 0c0743ae-b6b1-426a-9ec2-4a42530d799f · inbound

Team DACTYL at PAN 2026: Bayesian Data Mixing and Empirical X-risk Minimization for AI-text Detection cites this paper.

Team DACTYL at PAN 2026: Bayesian Data Mixing and Empirical X-risk Minimization for AI-text Detection Well-Read Students Learn Better: On the Importance of Pre-training Compact Models

Reference 53

Resolution
unresolved
no resolver link, observed 2026-08-01T18:14:43.084276Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T18:14:43.084276Z digest=sha256:343854991ab1264f224635ce63bc6c0505f539504d5d88c186489a66acaa2886