Pith. sign in

Paper Citation Record · LEDGER

Pre-Training LLMs on a budget: A comparison of three optimizers

As of 11 August 2026, this Paper Citation Record lists 40 of 40 outbound references and 3 inbound Pith citation observations for arXiv:2507.08472.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2507.08472 v2

Coverage vector

measured 40 of 40 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-06T18:22:23.931757Z

measured 43 of 43 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-10T06:31:04.303077+00:00

measured 3 of 3 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-07-31T11:19:35.239826Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-05-10T12:15:22.206452Z

Reference resolution

40 of 40 outbound references displayed

  • verified exact0
  • verified fuzzy15
  • unresolved25
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 7a2ea393-f05a-43d5-a9c0-434ecaa518c0 · outbound

This paper cites sign SGD with majority vote is communication efficient and fault tolerant.

Pre-Training LLMs on a budget: A comparison of three optimizers sign SGD with majority vote is communication efficient and fault tolerant

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:22:26.282195Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-06T18:22:21.483030Z digest=sha256:6d29a3e6569261f2547fbc7e310bfad24c84c74811fb088ad3d58c2209dd90c0

Observation 4fb7e08d-9d50-4d3c-869b-4e0a89bbbd10 · outbound

This paper cites Pythia: A Suite for Analyzing Large Language Models Across Training and Scaling.

Pre-Training LLMs on a budget: A comparison of three optimizers Pythia: A Suite for Analyzing Large Language Models Across Training and Scaling

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-06T18:22:21.496384Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T18:22:21.496384Z digest=sha256:34f2ad982d1bb3e0dcfcf4d24f06f8ff1690098fb3695d1550c7d713602dafbc

Observation cd89f21a-c756-47b7-b3a5-5b92a4e5978f · outbound

This paper cites Language models are few-shot learners.

Pre-Training LLMs on a budget: A comparison of three optimizers Language models are few-shot learners

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-06T18:22:21.506585Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T18:22:21.506585Z digest=sha256:a06e4dea5fc393acb3096f21f3db2e3d20799bc7fcf01c1c2fef90bdb8928b60

Observation ab00a737-6491-4258-8b39-6a8453ca908d · outbound

This paper cites Symbolic discovery of optimization algorithms.

Pre-Training LLMs on a budget: A comparison of three optimizers Symbolic discovery of optimization algorithms

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:22:26.188949Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-06T18:22:21.516425Z digest=sha256:9ebc643eac8a2d29aca402133e372d850a81c3d88f4d70d94c6bb710e4e1689a

Observation b138bef4-ea8b-4b2e-b0ba-148afe2a5154 · outbound

This paper cites Think you have Solved Question Answering? Try ARC, the AI2 Reasoning Challenge.

Pre-Training LLMs on a budget: A comparison of three optimizers Think you have Solved Question Answering? Try ARC, the AI2 Reasoning Challenge

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-06T18:22:21.526060Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T18:22:21.526060Z digest=sha256:aa98c93c7a5490f95446d0ddd358d4112bb23949092d553e60553c78feeb508b

Observation 318d535f-1e5d-4bdd-9b46-383f247151c7 · outbound

This paper cites BTLM-3B-8K: 7B Parameter Performance in a 3B Parameter Model.

Pre-Training LLMs on a budget: A comparison of three optimizers BTLM-3B-8K: 7B Parameter Performance in a 3B Parameter Model

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-06T18:22:21.535485Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T18:22:21.535485Z digest=sha256:c085e7ef6ff7d047de8758305063c5959646f47e943321263c7e4124190c8081

Observation 589650ec-6d73-4964-9628-2395376204ee · outbound

This paper cites A framework for few-shot language model evaluation, 07 2024.

Pre-Training LLMs on a budget: A comparison of three optimizers A framework for few-shot language model evaluation, 07 2024

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-06T18:22:21.545685Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T18:22:21.545685Z digest=sha256:b1b91d1855f4ae92a0d8a7e9557a6d892245524e1b701b19a5f6af0dc60e3f3e

Observation 2f589883-c4bb-4b03-a752-766fe9ab3b97 · outbound

This paper cites OLMES: A Standard for Language Model Evaluations.

Pre-Training LLMs on a budget: A comparison of three optimizers OLMES: A Standard for Language Model Evaluations

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-06T18:22:21.555619Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T18:22:21.555619Z digest=sha256:e880149d1a1484197e05e046c4b757e40634b206481d25d0a2c7ca0c9e614a83

Observation dab9aca4-5267-47b8-9945-a66bd0f6c6c1 · outbound

This paper cites Measuring massive multitask language understanding.

Pre-Training LLMs on a budget: A comparison of three optimizers Measuring massive multitask language understanding

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-06T18:22:21.580428Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T18:22:21.580428Z digest=sha256:b3eb7883e1b43b93ff8aef53d2bc7490d737c679bb2dc80eaac1463f7b395920

Observation 718cde59-c1d3-4d8c-9802-b0ba220e4bd2 · outbound

This paper cites Rae, and Laurent Sifre.

Pre-Training LLMs on a budget: A comparison of three optimizers Rae, and Laurent Sifre

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-06T18:22:21.590334Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T18:22:21.590334Z digest=sha256:a074da0e9ebf0cd55331e74e1cf49dc2edadce007a7e363ccc9522b099918e7d

Observation 4b1c014c-1ea5-4000-ab3c-92f0fae0f339 · outbound

This paper cites On the Parameterization of Second-Order Optimization Effective Towards the Infinite Width.

Pre-Training LLMs on a budget: A comparison of three optimizers On the Parameterization of Second-Order Optimization Effective Towards the Infinite Width

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-06T18:22:21.602995Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T18:22:21.602995Z digest=sha256:68f6fe4cca99cc38e07e7533f7f3ebcb0d6656e8bf6cf14abd7c61e4458c1c6c

Observation 97449433-a998-45a4-86f9-76aa52746a28 · outbound

This paper cites No train no gain: Revisiting efficient training algorithms for transformer-based language models.

Pre-Training LLMs on a budget: A comparison of three optimizers No train no gain: Revisiting efficient training algorithms for transformer-based language models

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:22:26.059920Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-06T18:22:21.712776Z digest=sha256:87341edf11199a656301d572360bb023f9d7afd2f9b2ee223210f609d7c421a6

Observation 297ef640-4e68-4fd3-87d3-48d6b58f84c6 · outbound

This paper cites A method for stochastic optimization.

Pre-Training LLMs on a budget: A comparison of three optimizers A method for stochastic optimization

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:22:25.918161Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-06T18:22:21.768846Z digest=sha256:ce305c309889d74d364c5e969d3d8fc8491cdc755cf94a9c3b78756a97298a5d

Observation c50b93d0-4f3b-4381-92cc-9206bc9fab8d · outbound

This paper cites Asam: Adaptive sharpness-aware minimization for scale-invariant learning of deep neural networks.

Pre-Training LLMs on a budget: A comparison of three optimizers Asam: Adaptive sharpness-aware minimization for scale-invariant learning of deep neural networks

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:22:25.776470Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-06T18:22:21.835815Z digest=sha256:6b7f76caa822d8d78700895209fef0f550a4b950b3e7bfca1f373147d4dd24dc

Observation 405688fe-4eae-4972-8b32-0ae8ffec189b · outbound

This paper cites ROPE : Reading order equivariant positional encoding for graph-based document information extraction.

Pre-Training LLMs on a budget: A comparison of three optimizers ROPE : Reading order equivariant positional encoding for graph-based document information extraction

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:22:25.616982Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-06T18:22:21.935665Z digest=sha256:da6fb5c592dd0e222e4f0f4aeaa32478bda93aac374b66d82c69e90b87531383

Observation f702a233-b0e3-4ee1-9389-32653e13a488 · outbound

This paper cites an unresolved cited work.

Pre-Training LLMs on a budget: A comparison of three optimizers Unresolved cited work

Reference 16

Resolution
unresolved
raw_fallback, observed 2026-08-06T18:22:25.497357Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-06T18:22:22.013963Z digest=sha256:f40bf2e7f182bebe06282975cef62019e2581b546d9bcb85c90c0c2270fc369b

Observation 96cc9c6d-e052-43a9-a3e7-940e0e07b170 · outbound

This paper cites An Empirical Study of $\mu$P Learning Rate Transfer.

Pre-Training LLMs on a budget: A comparison of three optimizers An Empirical Study of $\mu$P Learning Rate Transfer

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-06T18:22:22.099491Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T18:22:22.099491Z digest=sha256:f4c91496b2c0b4bdd3008bc7e93d4bd91f86dc2773c774781e6d96acc01c5722

Observation ec7119fa-76ab-4581-be0e-0493269ea6ca · outbound

This paper cites Sophia: A scalable stochastic second-order optimizer for language model pre-training.

Pre-Training LLMs on a budget: A comparison of three optimizers Sophia: A scalable stochastic second-order optimizer for language model pre-training

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:22:25.349057Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-06T18:22:22.233634Z digest=sha256:5473f44f6365d5afc62309cdd9682a73d41c4107559aac28ccfefd1a7c3e96b5

Observation 1c3f85f4-b2ee-4d14-a27e-94a65b1cbd47 · outbound

This paper cites Decoupled Weight Decay Regularization.

Pre-Training LLMs on a budget: A comparison of three optimizers Decoupled Weight Decay Regularization

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-06T18:22:22.322826Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T18:22:22.322826Z digest=sha256:0e840f2bd777f0883ccac0bea04cc427a3e3b79cc1bfebded623c7fdaade04a4

Observation 616636a0-c779-4e6e-be13-096e23376066 · outbound

This paper cites Scaling data-constrained language models.

Pre-Training LLMs on a budget: A comparison of three optimizers Scaling data-constrained language models

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-06T18:22:22.396965Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T18:22:22.396965Z digest=sha256:02a3a3545d7dd589b85b54940b60d2f087df40954d4b580b4da12c10adb777a2

Observation cdd2d457-dc2a-4937-8cad-2bd12a46ee15 · outbound

This paper cites Language models are unsupervised multitask learners.

Pre-Training LLMs on a budget: A comparison of three optimizers Language models are unsupervised multitask learners

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-06T18:22:22.479151Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T18:22:22.479151Z digest=sha256:ef17de46c0e4076fa5521851700f8dc062cb2aa2647cf110a698035bbf05d4e7

Observation f3cfcba2-07bc-456a-a1dd-b56b60f81b68 · outbound

This paper cites A modified A dam algorithm for deep neural network optimization.

Pre-Training LLMs on a budget: A comparison of three optimizers A modified A dam algorithm for deep neural network optimization

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:22:25.218513Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-06T18:22:22.557255Z digest=sha256:b22ac6f6ab4cb89fddc24fbbc3927e3f0bbf1aa6c41fc2699e573bed9eef30aa

Observation f782298b-0f15-41d1-8d96-eaaeea223815 · outbound

This paper cites An overview of gradient descent optimization algorithms.

Pre-Training LLMs on a budget: A comparison of three optimizers An overview of gradient descent optimization algorithms

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-06T18:22:22.619341Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T18:22:22.619341Z digest=sha256:9125098688cfea8b75951ad9d6455635fe9243fbdcc0cbe8d5c6cf2f1343bb6a

Observation 4d23e4df-9c67-47f1-b7d5-d39e12080c38 · outbound

This paper cites Adafactor: Adaptive learning rates with sublinear memory cost.

Pre-Training LLMs on a budget: A comparison of three optimizers Adafactor: Adaptive learning rates with sublinear memory cost

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:22:25.139815Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-06T18:22:22.702979Z digest=sha256:5a57304ee13b3016b12d2a62266bfae53f4f42b51bd1f7f51e5d82faa9a686f6

Observation 1de602ec-89cb-49b7-bdc8-a716687951ec · outbound

This paper cites SlimPajama: A 627B token cleaned and deduplicated version of RedPajama.

Pre-Training LLMs on a budget: A comparison of three optimizers SlimPajama: A 627B token cleaned and deduplicated version of RedPajama

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-06T18:22:22.820809Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T18:22:22.820809Z digest=sha256:008bbf8fc6b204066065d30e23ff688a86c6376b1525a8e869fae36462864802

Observation c1df782e-021a-4b02-bf3f-58d33eddde44 · outbound

This paper cites Spike no more: Stabilizing the pre-training of large language models, 2025.

Pre-Training LLMs on a budget: A comparison of three optimizers Spike no more: Stabilizing the pre-training of large language models, 2025

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:22:24.996150Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-06T18:22:22.890439Z digest=sha256:3fe9d71967361627276a189af01b9af917fc09493fd222250accb394b66f2928

Observation f3481f1d-0cd6-4369-a2bf-aa8047fea6d1 · outbound

This paper cites LLaMA: Open and Efficient Foundation Language Models.

Pre-Training LLMs on a budget: A comparison of three optimizers LLaMA: Open and Efficient Foundation Language Models

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-06T18:22:23.020109Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T18:22:23.020109Z digest=sha256:3dd695e450c01fcf1e33c60a767c41ce2111f99e925ce858977a6b34878bff27

Observation b9ed3627-872a-4194-bb55-ecae015a44b2 · outbound

This paper cites Evolution and role of optimizers in training deep learning models, 2024.

Pre-Training LLMs on a budget: A comparison of three optimizers Evolution and role of optimizers in training deep learning models, 2024

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:22:24.868709Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-06T18:22:23.098112Z digest=sha256:36a1f895b40a6c6200128692ccc7d66dc5e8f6317c2c7ef813e342809896e42d

Observation 9e4f294a-5d8d-4cab-a6d6-faeb9a94718c · outbound

This paper cites Ranger21: a synergistic deep learning optimizer.

Pre-Training LLMs on a budget: A comparison of three optimizers Ranger21: a synergistic deep learning optimizer

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-06T18:22:23.166969Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T18:22:23.166969Z digest=sha256:d35f327a5362211384fc554f648ea01a87138ce0efe4c2248445fe18d11e3c65

Observation 6f250c12-6b80-4fca-9c58-2be7d5e6a03e · outbound

This paper cites Adan: Adaptive nesterov momentum algorithm for faster optimizing deep models.

Pre-Training LLMs on a budget: A comparison of three optimizers Adan: Adaptive nesterov momentum algorithm for faster optimizing deep models

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:22:24.767914Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-06T18:22:23.243728Z digest=sha256:018aceead177387f67d1410408a044c810df30011ffd3096977e61f46c96a6c7

Observation c83071cc-073c-4e99-b631-ab1a1292e651 · outbound

This paper cites Tuning large neural networks via zero-shot hyperparameter transfer.

Pre-Training LLMs on a budget: A comparison of three optimizers Tuning large neural networks via zero-shot hyperparameter transfer

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:22:24.654484Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-06T18:22:23.319784Z digest=sha256:23b0d8652484482a1c3acacefae328dc79b55df5f822342579e5e59b2f7ecac0

Observation 38cdee5c-1ba7-41cd-a7cb-2e6166473774 · outbound

This paper cites Hellaswag: Can a machine really finish your sentence? In Proceedings of the 57th Annual Meeting of the Association for Computational Linguistics, 2019.

Pre-Training LLMs on a budget: A comparison of three optimizers Hellaswag: Can a machine really finish your sentence? In Proceedings of the 57th Annual Meeting of the Association for Computational Linguistics, 2019

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-06T18:22:23.370940Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T18:22:23.370940Z digest=sha256:9b51fd70920e85096b7ab3e7b6a767ea5af58a40f78f9d763b737fb964e40db3

Observation 71c1903a-373d-4d1c-b18a-7bf22963b90f · outbound

This paper cites Adam-mini: Use Fewer Learning Rates To Gain More.

Pre-Training LLMs on a budget: A comparison of three optimizers Adam-mini: Use Fewer Learning Rates To Gain More

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-06T18:22:23.426489Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T18:22:23.426489Z digest=sha256:56a47bc9ef6304c25826bdc221c674e495bc305a5829435f650a459a31880a01

Observation e7156432-f7cf-4697-b702-72cfd2c05fc4 · outbound

This paper cites Improved adam optimizer for deep neural networks.

Pre-Training LLMs on a budget: A comparison of three optimizers Improved adam optimizer for deep neural networks

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:22:24.506253Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-06T18:22:23.522210Z digest=sha256:16b2e604ef99304d90694cb37ebdbd110848cfe3e9583264f54f70fdd9243572

Observation 6f168c4e-d217-4171-9806-b2a140876725 · outbound

This paper cites an unresolved cited work.

Pre-Training LLMs on a budget: A comparison of three optimizers Unresolved cited work

Reference 35

Resolution
unresolved
raw_fallback, observed 2026-08-06T18:22:24.359342Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-06T18:22:23.554627Z digest=sha256:cb112f581b98a272a6f87308d7c614b954d2b32f89cee892c8210e00b103b3e7

Observation 6b5ec9fb-1624-45c9-a126-83da3fa1f79f · outbound

This paper cites Adabelief optimizer: Adapting stepsizes by the belief in observed gradients.

Pre-Training LLMs on a budget: A comparison of three optimizers Adabelief optimizer: Adapting stepsizes by the belief in observed gradients

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:22:24.258442Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-06T18:22:23.592781Z digest=sha256:7d16efe4155bdae62bb2f3c2ab6093c3a4d6c8eefa4a49ea68902f31107d2baf

Observation fe6525a8-af09-40f1-99a7-af0421227267 · outbound

This paper cites write newline.

Pre-Training LLMs on a budget: A comparison of three optimizers write newline

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-06T18:22:23.691536Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T18:22:23.691536Z digest=sha256:d7fb49634ac5c4abf0f3ed80979448aeecab0fbe893ef5cfec34e8a7eb953c70

Observation 6062aa24-20ae-455b-85c0-f3a9a36eddc4 · outbound

This paper cites @esa (Ref.

Pre-Training LLMs on a budget: A comparison of three optimizers @esa (Ref

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-06T18:22:23.785237Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T18:22:23.785237Z digest=sha256:800e31b0c9a2f4f114dee9d7201323a4a9a840af66c7c67df69d2e0f0039e9de

Observation 6b64a7ad-d192-4238-aff7-d1f3aaf472e4 · outbound

This paper cites an unresolved cited work.

Pre-Training LLMs on a budget: A comparison of three optimizers Unresolved cited work

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-06T18:22:23.885318Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T18:22:23.885318Z digest=sha256:151c602aa6934208442a67a29c33445031a9f96d2023c1380e68f8ca50fa84b6

Observation b8b538c4-8e24-424d-9cfa-7ef785a12a42 · outbound

This paper cites an unresolved cited work.

Pre-Training LLMs on a budget: A comparison of three optimizers Unresolved cited work

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-06T18:22:23.931757Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T18:22:23.931757Z digest=sha256:f120de3928ded411c3deba2aa3fa1256bdb4e066ef1e4ed46f512ea59648e6c3

Pith citing papers

Observation 89f747dd-bff9-48d0-b28f-70f65461bca6 · inbound

StoSignSGD: Unbiased Structural Stochasticity Fixes SignSGD for Training Large Language Models cites this paper.

StoSignSGD: Unbiased Structural Stochasticity Fixes SignSGD for Training Large Language Models Pre-Training LLMs on a budget: A comparison of three optimizers

Reference 35

Resolution
verified exact
arxiv_id, observed 2026-05-10T12:15:22.208322Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-10T12:10:44.802059Z digest=sha256:056434e98491ab069ebdde762cff3b587ff532195e2a820c7bc50183eaddcfbe

Observation 2391eeff-d65c-459e-99ec-1f5d71f89933 · inbound

OmniOpt: Taxonomy, Geometry, and Benchmarking of Modern Optimizers cites this paper.

OmniOpt: Taxonomy, Geometry, and Benchmarking of Modern Optimizers Pre-Training LLMs on a budget: A comparison of three optimizers

Reference 96

Resolution
unresolved
no resolver link, observed 2026-07-11T22:10:49.683444Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-11T22:10:49.683444Z digest=sha256:f381f6a34752d14f1e0796c075c05fd9a818b719dbb0272766129e4315a32c34

Observation 6feb931c-ffda-4945-bc51-44caea724920 · inbound

From Data to Device: ELMOD An Efficient German-First 2.7B Language Model for Mobile Inference cites this paper.

From Data to Device: ELMOD An Efficient German-First 2.7B Language Model for Mobile Inference Pre-Training LLMs on a budget: A comparison of three optimizers

Reference 32

Resolution
unresolved
no resolver link, observed 2026-07-31T11:19:35.239826Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-31T11:19:35.239826Z digest=sha256:218f22360ce86e09965ed2e5985795355cdf4f8300094a9a481acd638bf86d0e