Pith. sign in

Paper Citation Record · LEDGER

Mixture of Training: Recombining Small-Scale Scaffolded Pretraining Runs into a Larger Language Model

As of 22 August 2026, this Paper Citation Record lists 30 of 30 outbound references and 0 inbound Pith citation observations for arXiv:2608.13277.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2608.13277 v1

Coverage vector

measured 30 of 30 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-14T14:53:28.952395Z

measured 30 of 30 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-22T06:32:14.747728+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

30 of 30 outbound references displayed

  • verified exact0
  • verified fuzzy9
  • unresolved21
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 98109fd7-6f8b-40cb-b188-b0a7ed773738 · outbound

This paper cites Scaling Learning Algorithms Towards.

Mixture of Training: Recombining Small-Scale Scaffolded Pretraining Runs into a Larger Language Model Scaling Learning Algorithms Towards

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-14T14:53:28.565114Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-14T14:53:28.565114Z digest=sha256:ccf178e4c859687fd6cfa0e8fb52db723482261fda7a89e4f5514d3386445c4a

Observation 8dbfa3d5-3abe-47bc-91b7-33d60a19f9fc · outbound

This paper cites and Osindero, Simon and Teh, Yee Whye , journal =.

Mixture of Training: Recombining Small-Scale Scaffolded Pretraining Runs into a Larger Language Model and Osindero, Simon and Teh, Yee Whye , journal =

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-14T14:53:28.569461Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-14T14:53:28.569461Z digest=sha256:1ff82ac1c40c339c47b6a7a155bdd4599b4a14e399bed2cefc3f769dbbb85132

Observation d9809a49-a920-4f2d-aaf9-b2d1ace22bea · outbound

This paper cites 2016 , publisher=.

Mixture of Training: Recombining Small-Scale Scaffolded Pretraining Runs into a Larger Language Model 2016 , publisher=

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-14T14:53:28.574174Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-14T14:53:28.574174Z digest=sha256:76f005ab75658d55f0d32bfacda4f50c8826a4041378437c8af9f584e5e69122

Observation f4b26545-0aba-4d8e-b207-040213d58538 · outbound

This paper cites Scaling Laws for Neural Language Models.

Mixture of Training: Recombining Small-Scale Scaffolded Pretraining Runs into a Larger Language Model Scaling Laws for Neural Language Models

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-14T14:53:28.578151Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-14T14:53:28.578151Z digest=sha256:1d40cd5942406f602c99afbefd95af66884a865754912657e28f769a73e075bb

Observation 60d5944c-e46d-4c57-8a01-fa4b810cc6d7 · outbound

This paper cites 2022 , eprint=.

Mixture of Training: Recombining Small-Scale Scaffolded Pretraining Runs into a Larger Language Model 2022 , eprint=

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-14T14:53:28.607908Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-14T14:53:28.607908Z digest=sha256:9e392dddd901391cef3bef4687d0fb54b98c1bbae3879fc3ca48d504386f6508

Observation 57dd31bf-f372-40a8-aa36-7e8140d730d5 · outbound

This paper cites 2025 , eprint=.

Mixture of Training: Recombining Small-Scale Scaffolded Pretraining Runs into a Larger Language Model 2025 , eprint=

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T14:53:29.551475Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=arxiv_source observed=2026-08-14T14:53:28.634888Z digest=sha256:05a1d2414108f852fe469452d775c18f3e68985fc97e4fd0a3b6cee8e3ec5d93

Observation ea31c015-7823-468e-b558-74bde54c95c2 · outbound

This paper cites 2023 , eprint=.

Mixture of Training: Recombining Small-Scale Scaffolded Pretraining Runs into a Larger Language Model 2023 , eprint=

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-14T14:53:28.711059Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-14T14:53:28.711059Z digest=sha256:ef3b44ad6b146d5e781eed4c74343e24e51ffe78e4429694346c162b19c6bfd3

Observation 4de13d72-2e35-43eb-a6c1-01df85b31d50 · outbound

This paper cites 2020 , eprint=.

Mixture of Training: Recombining Small-Scale Scaffolded Pretraining Runs into a Larger Language Model 2020 , eprint=

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-14T14:53:28.752336Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-14T14:53:28.752336Z digest=sha256:aa8e53cabc3811c097e57cce8c9450fa2cc11b8b6847381c6db3b2970c9a1dd9

Observation bd4cc876-0765-4bbc-bf72-8170333a7469 · outbound

This paper cites 2022 , editor =.

Mixture of Training: Recombining Small-Scale Scaffolded Pretraining Runs into a Larger Language Model 2022 , editor =

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-14T14:53:28.755667Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-14T14:53:28.755667Z digest=sha256:32d039d34844203c00beed0b87c092e0113ef9360bd3675c728689871ee249fe

Observation 1bb9491f-5c44-4b53-9652-71c82b67cdca · outbound

This paper cites Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR) , month =.

Mixture of Training: Recombining Small-Scale Scaffolded Pretraining Runs into a Larger Language Model Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR) , month =

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T14:53:29.526389Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=arxiv_source observed=2026-08-14T14:53:28.759838Z digest=sha256:71f7af661bd1f99bd2cc20454b7950ebc75782e024bd49560c25abf2c6a1effb

Observation 9fcb1430-c893-4a7b-a565-e5e18a7b021a · outbound

This paper cites Routing Networks and the Challenges of Modular and Compositional Computation.

Mixture of Training: Recombining Small-Scale Scaffolded Pretraining Runs into a Larger Language Model Routing Networks and the Challenges of Modular and Compositional Computation

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-14T14:53:28.763874Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-14T14:53:28.763874Z digest=sha256:1d14798cee6baa86676ccf9e7f72fad42ad55f65be6f615ec2dd90ee27b95a95

Observation c80b79f3-cb58-4040-8048-0be01b59df24 · outbound

This paper cites 2024 , eprint=.

Mixture of Training: Recombining Small-Scale Scaffolded Pretraining Runs into a Larger Language Model 2024 , eprint=

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-14T14:53:28.768724Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-14T14:53:28.768724Z digest=sha256:7793b0e84de13d44faaa8d93f663b8d078a82ddb8953bd58bf8a80ccffd78753

Observation 8998f33e-0bee-4572-8c62-72568bff7f1d · outbound

This paper cites Gemma: Open Models Based on Gemini Research and Technology.

Mixture of Training: Recombining Small-Scale Scaffolded Pretraining Runs into a Larger Language Model Gemma: Open Models Based on Gemini Research and Technology

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-14T14:53:28.772018Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-14T14:53:28.772018Z digest=sha256:1910bc27052f07db1843494bdc3345ef13f8c5fee315d652c1d978cf7193afd3

Observation 3aef1a72-e43d-4464-9615-9cdf9654a132 · outbound

This paper cites Distill , year =.

Mixture of Training: Recombining Small-Scale Scaffolded Pretraining Runs into a Larger Language Model Distill , year =

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-14T14:53:28.775615Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-14T14:53:28.775615Z digest=sha256:c54341f2c62b13241d6f14a104b7234ade37c12790ef055cd6f737d3aa085865

Observation 0ad97f43-1212-435a-8428-8331de054987 · outbound

This paper cites 2025 , eprint=.

Mixture of Training: Recombining Small-Scale Scaffolded Pretraining Runs into a Larger Language Model 2025 , eprint=

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T14:53:29.511405Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=arxiv_source observed=2026-08-14T14:53:28.778762Z digest=sha256:efd63eb9705d36a7bf0476772ca0eaebc2ae0e15b73987b0581cc0aa301a6842

Observation 95693b35-454c-4dd4-b555-fa24d7b8c296 · outbound

This paper cites 2022 , eprint=.

Mixture of Training: Recombining Small-Scale Scaffolded Pretraining Runs into a Larger Language Model 2022 , eprint=

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-14T14:53:28.782353Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-14T14:53:28.782353Z digest=sha256:28d7acfdca8f22cad6540e588a73f60724e76730296bd0d3ddbd319f4f8c7683

Observation 7ec37f5a-8f8a-4761-abd8-267d527b58ad · outbound

This paper cites bert2 BERT : Towards Reusable Pretrained Language Models.

Mixture of Training: Recombining Small-Scale Scaffolded Pretraining Runs into a Larger Language Model bert2 BERT : Towards Reusable Pretrained Language Models

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-14T14:53:28.787436Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-14T14:53:28.787436Z digest=sha256:27aec673aaf8b03d7a1701607558f840f0966254b98c4350df1ffd44d9e25830

Observation 911d241e-5440-49f3-adad-26074b55ca2a · outbound

This paper cites Revisiting Model Stitching to Compare Neural Representations.

Mixture of Training: Recombining Small-Scale Scaffolded Pretraining Runs into a Larger Language Model Revisiting Model Stitching to Compare Neural Representations

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-14T14:53:28.825001Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-14T14:53:28.825001Z digest=sha256:a656a05f9ae035f046641c9690e568bee23df0a623a3462ac76e490531afe281

Observation 8974cb8d-5f80-41bd-ac65-faedce62ac31 · outbound

This paper cites 2023 , eprint=.

Mixture of Training: Recombining Small-Scale Scaffolded Pretraining Runs into a Larger Language Model 2023 , eprint=

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T14:53:29.380917Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=arxiv_source observed=2026-08-14T14:53:28.890106Z digest=sha256:69dc64795be47f3351267d8533c42c98d6e29c0699d0f11ed222ef48992c4243

Observation 6ffa92f1-a89e-4ef2-a459-d30c012bc284 · outbound

This paper cites Efficient Large-Scale Distributed Training of Conditional Maximum Entropy Models , url =.

Mixture of Training: Recombining Small-Scale Scaffolded Pretraining Runs into a Larger Language Model Efficient Large-Scale Distributed Training of Conditional Maximum Entropy Models , url =

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T14:53:29.261207Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=arxiv_source observed=2026-08-14T14:53:28.916907Z digest=sha256:86770b117b4b6b83217026457da742694fce03f9e0b677e22537a790b890d66b

Observation 7b5f18d1-4900-4ca4-8629-ac78eac32f24 · outbound

This paper cites Finetuned Language Models Are Zero-Shot Learners.

Mixture of Training: Recombining Small-Scale Scaffolded Pretraining Runs into a Larger Language Model Finetuned Language Models Are Zero-Shot Learners

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-14T14:53:28.919861Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-14T14:53:28.919861Z digest=sha256:a2537cb2b8d99072941a93243a1860d127b43bca86d49efa375dd38cd8d21fff

Observation 118a7fe0-25b7-4844-9d6f-4861088b1b81 · outbound

This paper cites 2022 , eprint=.

Mixture of Training: Recombining Small-Scale Scaffolded Pretraining Runs into a Larger Language Model 2022 , eprint=

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-14T14:53:28.923926Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-14T14:53:28.923926Z digest=sha256:04266cf0d16911ac4052c8288cf563bd96662b7791f2355e0bd224abed016d83

Observation b0b9d635-7744-49ad-be65-ee5b67080821 · outbound

This paper cites 2022 , eprint=.

Mixture of Training: Recombining Small-Scale Scaffolded Pretraining Runs into a Larger Language Model 2022 , eprint=

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-14T14:53:28.928129Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-14T14:53:28.928129Z digest=sha256:b7171983ba824ac612c70a1618f8f5f51c235c11386886ca4e3515b5f29dee4a

Observation a2686474-1188-493b-95da-65c1bf6241a1 · outbound

This paper cites 2022 , eprint=.

Mixture of Training: Recombining Small-Scale Scaffolded Pretraining Runs into a Larger Language Model 2022 , eprint=

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-14T14:53:28.930690Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-14T14:53:28.930690Z digest=sha256:f2a930116d5cf14ff9d3679b7c9bac996d1a8f66f8b5e4e38da2f0aed2e0da0f

Observation 1a8ef394-e7e8-4f03-adaf-c0e050621e9d · outbound

This paper cites Efficient Large-Scale Language Model Training on GPU Clusters Using Megatron-LM.

Mixture of Training: Recombining Small-Scale Scaffolded Pretraining Runs into a Larger Language Model Efficient Large-Scale Language Model Training on GPU Clusters Using Megatron-LM

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-14T14:53:28.934067Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-14T14:53:28.934067Z digest=sha256:39b56818e73ee33d0dde878ce6052d1a82bd690589e4d058a79cf7b6c7150645

Observation 1ae7bd33-5974-4081-a7a4-cb45b1d14471 · outbound

This paper cites 2026 , eprint=.

Mixture of Training: Recombining Small-Scale Scaffolded Pretraining Runs into a Larger Language Model 2026 , eprint=

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T14:53:29.233856Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=arxiv_source observed=2026-08-14T14:53:28.938634Z digest=sha256:c38e4a968803f2b162568eb12f920503b08adab24eaef0046c1473b33962efcd

Observation 0ac3cd42-437f-4c24-a7e7-57e746b52d4b · outbound

This paper cites The Twelfth International Conference on Learning Representations , year =.

Mixture of Training: Recombining Small-Scale Scaffolded Pretraining Runs into a Larger Language Model The Twelfth International Conference on Learning Representations , year =

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T14:53:29.224560Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=arxiv_source observed=2026-08-14T14:53:28.942835Z digest=sha256:e33e3676da56ef27548f23f6bc77435b81b571178a24f0ad41941c78d6fcdf90

Observation 452539a7-ce9d-4dbc-b6df-63e6e4441c1e · outbound

This paper cites Stacking Your Transformers: A Closer Look at Model Growth for Efficient.

Mixture of Training: Recombining Small-Scale Scaffolded Pretraining Runs into a Larger Language Model Stacking Your Transformers: A Closer Look at Model Growth for Efficient

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T14:53:29.206022Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=arxiv_source observed=2026-08-14T14:53:28.946202Z digest=sha256:44c395748116db6b9fd5060f82fe7ebedbbfb1aa13d034d6f431a4dc0393aad0

Observation af331a98-4586-494d-8ef1-4f6823dc5c4d · outbound

This paper cites Proceedings of the IEEE/CVF International Conference on Computer Vision , pages =.

Mixture of Training: Recombining Small-Scale Scaffolded Pretraining Runs into a Larger Language Model Proceedings of the IEEE/CVF International Conference on Computer Vision , pages =

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T14:53:29.190185Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=arxiv_source observed=2026-08-14T14:53:28.949493Z digest=sha256:8d9f56a37aea408ee8d06aba06896535073f0ba6aa6817cda09ebdaa8eba2a17

Observation f21f8c71-d134-4461-9602-267260c4bca0 · outbound

This paper cites an unresolved cited work.

Mixture of Training: Recombining Small-Scale Scaffolded Pretraining Runs into a Larger Language Model Unresolved cited work

Reference 30

Resolution
unresolved
raw_fallback, observed 2026-08-14T14:53:29.094967Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=arxiv_source observed=2026-08-14T14:53:28.952395Z digest=sha256:a1ed92570099348de14b88eb6d4ad6921b42989fb28a3eab0055450871ea3e61

Pith citing papers

No inbound Pith citation observations are available.