Pith. sign in

Paper Citation Record · LEDGER

Scale Efficiently: Insights from Pre-training and Fine-tuning Transformers

As of 8 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 16 inbound Pith citation observations for arXiv:2109.10686.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2109.10686 v2

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 16 of 16 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-08T06:32:00.761636+00:00

measured 16 of 16 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-08T16:20:38.370936Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-03T20:58:58.624838Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation f5c8e044-8d91-4c13-8ae6-394d6e6722d2 · inbound

ST-MoE: Designing Stable and Transferable Sparse Expert Models cites this paper.

ST-MoE: Designing Stable and Transferable Sparse Expert Models Scale Efficiently: Insights from Pre-training and Fine-tuning Transformers

Reference 203

Resolution
verified exact
arxiv_id, observed 2026-05-12T23:14:25.884221Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-05-12T23:14:25.431471Z digest=sha256:17eba526db338cbf0b1be42c7e80a0b81111096950aa0bbffee48462e9056fc7

Observation 7cbb63ad-5a56-42f2-8554-9bf072061909 · inbound

BloombergGPT: A Large Language Model for Finance cites this paper.

BloombergGPT: A Large Language Model for Finance Scale Efficiently: Insights from Pre-training and Fine-tuning Transformers

Reference 114

Resolution
verified exact
arxiv_id, observed 2026-05-13T23:19:46.822979Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-05-13T23:19:46.231145Z digest=sha256:1253a449a4d43a276a09f41392aaf058ef0b504c29bd6927061047c930a33e94

Observation 5bdfa959-e310-4485-baac-dd5456fe053f · inbound

Scaling Data-Constrained Language Models cites this paper.

Scaling Data-Constrained Language Models Scale Efficiently: Insights from Pre-training and Fine-tuning Transformers

Reference 113

Resolution
verified exact
arxiv_id, observed 2026-05-18T01:35:21.396172Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-18T01:35:21.150772Z digest=sha256:f8694f9d2c9c1900e30fd4e3d9d698012ca16a0f7a057085a80d202c4dbcc644

Observation 4ac6baf2-3cc7-48da-a810-1e916a881b4e · inbound

Chronos: Learning the Language of Time Series cites this paper.

Chronos: Learning the Language of Time Series Scale Efficiently: Insights from Pre-training and Fine-tuning Transformers

Reference 83

Resolution
verified exact
arxiv_id, observed 2026-05-13T08:27:23.392866Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-05-13T08:27:23.298009Z digest=sha256:76849550527d4048aec0f2526897bac5dda4707d334c5f49d79a09da9b1ff57b

Observation f22ca0db-c68d-4623-9a52-b12d456b3865 · inbound

PiKE: Adaptive Data Mixing for Large-Scale Multi-Task Learning Under Low Gradient Conflicts cites this paper.

PiKE: Adaptive Data Mixing for Large-Scale Multi-Task Learning Under Low Gradient Conflicts Scale Efficiently: Insights from Pre-training and Fine-tuning Transformers

Reference 58

Resolution
unresolved
no resolver link, observed 2026-08-08T16:20:38.370936Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T16:20:38.370936Z digest=sha256:1b55e950012b8443b8d15e8d29a84d99a8b5cf32cca56bcf3efec049fc9e3242

Observation ceb5ec8e-0c97-4438-a924-013435778864 · inbound

MUDDFormer: Breaking Residual Bottlenecks in Transformers via Multiway Dynamic Dense Connections cites this paper.

MUDDFormer: Breaking Residual Bottlenecks in Transformers via Multiway Dynamic Dense Connections Scale Efficiently: Insights from Pre-training and Fine-tuning Transformers

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-07T22:34:08.126213Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T22:34:08.126213Z digest=sha256:274a676a3c8c6db34d8716f81dbac6d8b06e2c8df5915782ec0437045fe1ac9f

Observation 4b3b7d66-e4a5-42fb-8d59-7cbebfbb1b7c · inbound

An Efficient Private GPT Never Autoregressively Decodes cites this paper.

An Efficient Private GPT Never Autoregressively Decodes Scale Efficiently: Insights from Pre-training and Fine-tuning Transformers

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-07T15:32:18.828359Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T15:32:18.828359Z digest=sha256:62b0bb87f7e48296502ef5d49be55ed0def191f8f1a061004fbfd1b8ae482398

Observation 61d5f9ca-7202-40c5-8fdd-4c14bb08c6b2 · inbound

Progressive Scaling Visual Object Tracking cites this paper.

Progressive Scaling Visual Object Tracking Scale Efficiently: Insights from Pre-training and Fine-tuning Transformers

Reference 71

Resolution
unresolved
no resolver link, observed 2026-08-07T14:11:11.182755Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:11:11.182755Z digest=sha256:dee3c79e3ab3497a3324af1dd1f00f4af90ce71b1c014c6d2cb49f2220de6cd7

Observation bd6fc5f9-c955-458c-9330-f74c042996d9 · inbound

Text-to-LoRA: Instant Transformer Adaption cites this paper.

Text-to-LoRA: Instant Transformer Adaption Scale Efficiently: Insights from Pre-training and Fine-tuning Transformers

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-07T06:05:29.762633Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T06:05:29.762633Z digest=sha256:785641e7ab2141efdcf72defb10e748985a57d4ea58a23eaa4786694d86b9d39

Observation 7f3d18c2-6685-4b5f-a35a-2ac7f7ab6ca8 · inbound

Random Initialization Can't Catch Up: The Advantage of Language Model Transfer for Time Series Forecasting cites this paper.

Random Initialization Can't Catch Up: The Advantage of Language Model Transfer for Time Series Forecasting Scale Efficiently: Insights from Pre-training and Fine-tuning Transformers

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-07T04:16:22.020141Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T04:16:22.020141Z digest=sha256:51d3e084a1804eea87f5f19c0240db335e9ee1206da6c253a6a233966ec38cdb

Observation 348a92a7-7bae-4d58-bf8f-fbcd81cf3ceb · inbound

Can Interpretation Predict Behavior on Unseen Data? cites this paper.

Can Interpretation Predict Behavior on Unseen Data? Scale Efficiently: Insights from Pre-training and Fine-tuning Transformers

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-06T19:10:30.753766Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:10:30.753766Z digest=sha256:8ad01513db1b18e5002bac7d666afe5f7c090b6e424706ca91940615ad1fda40

Observation 080ef9b7-cddb-4152-bcc5-4922789c9383 · inbound

On the Fitness Landscape in the $NK$ Model cites this paper.

On the Fitness Landscape in the $NK$ Model Scale Efficiently: Insights from Pre-training and Fine-tuning Transformers

Reference 91

Resolution
unresolved
no resolver link, observed 2026-08-05T19:29:00.389785Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T19:29:00.389785Z digest=sha256:a6cd18dc8470660dd6d1113eeefdbc50fa327116adb80330747ee58da69f3206

Observation 573986e9-fc42-4169-9db4-cb78c5561e47 · inbound

Large-Scale AI and Foundation Models for Neuroscience: A Comprehensive Review cites this paper.

Large-Scale AI and Foundation Models for Neuroscience: A Comprehensive Review Scale Efficiently: Insights from Pre-training and Fine-tuning Transformers

Reference 67

Resolution
unresolved
no resolver link, observed 2026-08-04T09:14:46.158438Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T09:14:46.158438Z digest=sha256:04f713e821185597d6536c2021ffa50e7247ac6c5734fc19cf2c52dfb12922c0

Observation ad6a4eb1-b724-46c8-9c8d-59507364db17 · inbound

Scaling Laws Meet Model Architecture: Toward Inference-Efficient LLMs cites this paper.

Scaling Laws Meet Model Architecture: Toward Inference-Efficient LLMs Scale Efficiently: Insights from Pre-training and Fine-tuning Transformers

Reference 39

Resolution
verified exact
arxiv_id, observed 2026-05-18T05:30:55.097385Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-18T05:30:11.389756Z digest=sha256:de787560c7ac1177af1214f1f372c9a43ab11431b51a6631a7821bd33c804f0f

Observation 15cfc1ff-6bfc-4712-9e17-e679d507a60f · inbound

Large-scale Codec Avatars: The Unreasonable Effectiveness of Large-scale Avatar Pretraining cites this paper.

Large-scale Codec Avatars: The Unreasonable Effectiveness of Large-scale Avatar Pretraining Scale Efficiently: Insights from Pre-training and Fine-tuning Transformers

Reference 59

Resolution
verified exact
arxiv_id, observed 2026-05-13T21:33:18.336782Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-13T21:32:39.946578Z digest=sha256:34c59e5b9cf1416aaef7b0223db23478f4e400b241b83e7a3ab7bff562fdd8d5

Observation 93b62383-9748-40ad-aebb-e83093829b24 · inbound

Variable-Width Transformers cites this paper.

Variable-Width Transformers Scale Efficiently: Insights from Pre-training and Fine-tuning Transformers

Reference 39

Resolution
metadata mismatch
arxiv_id, observed 2026-07-03T20:58:58.626372Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-06-27T00:57:22.872902Z digest=sha256:d2cdb739abed7115e0504fbced87944bb8eb9c6fee1f5fd6ffd72486b858ad21