Pith. sign in

Paper Citation Record · LEDGER

Optimal Embedding Learning Rate in LLMs: The Effect of Vocabulary Size

As of 18 August 2026, this Paper Citation Record lists 38 of 38 outbound references and 3 inbound Pith citation observations for arXiv:2506.15025.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2506.15025 v1

Coverage vector

measured 38 of 38 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-15T19:54:42.614806Z

measured 41 of 41 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-17T06:30:58.91139+00:00

measured 3 of 3 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-06-30T17:31:32.533941Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-06-30T17:34:57.593618Z

Reference resolution

38 of 38 outbound references displayed

  • verified exact0
  • verified fuzzy17
  • unresolved21
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation b3f8b846-6e14-4517-be05-f687fba975d8 · outbound

This paper cites u-$\mu$P: The Unit-Scaled Maximal Update Parametrization.

Optimal Embedding Learning Rate in LLMs: The Effect of Vocabulary Size u-$\mu$P: The Unit-Scaled Maximal Update Parametrization

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-15T19:54:42.444889Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T19:54:42.444889Z digest=sha256:8084646b33c697110c3d9ad013a86770ff5cc6aecf345497308b8d32e56b1c8f

Observation b4fc21cb-ea38-4af4-afab-63a2f63c7b5e · outbound

This paper cites Depthwise Hyperparameter Transfer in Residual Networks: Dynamics and Scaling Limit.

Optimal Embedding Learning Rate in LLMs: The Effect of Vocabulary Size Depthwise Hyperparameter Transfer in Residual Networks: Dynamics and Scaling Limit

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-15T19:54:42.450340Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T19:54:42.450340Z digest=sha256:d1cdb3ed9f4f31b8c7ad24300d2db342701576365239842704dfc841e6244b96

Observation 4e9517dd-5fdd-4a8e-91e5-5d50efc79d17 · outbound

This paper cites On Lazy Training in Differentiable Programming.

Optimal Embedding Learning Rate in LLMs: The Effect of Vocabulary Size On Lazy Training in Differentiable Programming

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-15T19:54:42.454722Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T19:54:42.454722Z digest=sha256:cd73a0f6a9b67412a809df9ef4599fdda1802826d266f057030a98647ba83707

Observation 2e495ec2-61aa-4a99-b5c0-c7a317109251 · outbound

This paper cites Infinite- width limit of deep linear neural networks.

Optimal Embedding Learning Rate in LLMs: The Effect of Vocabulary Size Infinite- width limit of deep linear neural networks

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-15T19:54:42.459894Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T19:54:42.459894Z digest=sha256:f6da7a3877ea489cd27a713b1469c09e4353bf3f25bb3717890bc089d22f3d2a

Observation 52002095-c746-4854-844c-d90658e74767 · outbound

This paper cites Alemi, Roman Novak, Peter J.

Optimal Embedding Learning Rate in LLMs: The Effect of Vocabulary Size Alemi, Roman Novak, Peter J

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T19:54:43.129312Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-15T19:54:42.464142Z digest=sha256:8a803a3d695085e9f4fca4b6fc36032feb4b7969ae905575f3ee99efa005981c

Observation 82f197d9-4eb3-444a-83c2-647fe2f7f6c1 · outbound

This paper cites On the infinite-depth limit of finite-width neural networks.

Optimal Embedding Learning Rate in LLMs: The Effect of Vocabulary Size On the infinite-depth limit of finite-width neural networks

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T19:54:43.115696Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-15T19:54:42.473690Z digest=sha256:1654ed7f89fefe2227d4c69f5aa311a99a41a11ccdb25117fbd7ae674e650e60

Observation 87e1e4de-8cd6-4997-a736-fe3820039dec · outbound

This paper cites On the impact of the ac- tivation function on deep neural networks training.

Optimal Embedding Learning Rate in LLMs: The Effect of Vocabulary Size On the impact of the ac- tivation function on deep neural networks training

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T19:54:43.102018Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-15T19:54:42.477648Z digest=sha256:9a8c970fdef7ed14da64f0672feaa5da5e6d2ebefc07466bc85cb66f26842305

Observation 2f285087-ba6c-4b17-946f-12894b8c479c · outbound

This paper cites Stable resnet.

Optimal Embedding Learning Rate in LLMs: The Effect of Vocabulary Size Stable resnet

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T19:54:43.087750Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-15T19:54:42.482191Z digest=sha256:bdac151c21a8ffabb8e2a05479ae8f88aaa11cc07b40c3ef44ee7c35127d4158

Observation b5cf8226-42e4-4f65-8ba2-6adf96c288c4 · outbound

This paper cites Delving Deep into Rectifiers: Surpassing Human-Level Performance on ImageNet Classification.

Optimal Embedding Learning Rate in LLMs: The Effect of Vocabulary Size Delving Deep into Rectifiers: Surpassing Human-Level Performance on ImageNet Classification

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-15T19:54:42.485996Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T19:54:42.485996Z digest=sha256:7c64a1358261c57a159b920c871ebc8a5cfa51fe887c80fe3583c8f4c1c674b2

Observation fd4661cb-4347-492a-aca5-915c0d87ec2a · outbound

This paper cites Neural Tangent Kernel: Convergence and Generalization in Neural Networks.

Optimal Embedding Learning Rate in LLMs: The Effect of Vocabulary Size Neural Tangent Kernel: Convergence and Generalization in Neural Networks

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-15T19:54:42.490672Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T19:54:42.490672Z digest=sha256:abe337356e4c895a9406db9c89571b35d3ac04d67cf9a840cd5b475b1ab6d704

Observation 590b4620-eaa3-47ac-be14-15a5c3c5747e · outbound

This paper cites Muon: An optimizer for hidden layers in neural networks,.

Optimal Embedding Learning Rate in LLMs: The Effect of Vocabulary Size Muon: An optimizer for hidden layers in neural networks,

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T19:54:43.070446Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-15T19:54:42.495390Z digest=sha256:a98170236b414f67eea7835fabf8c7dbc181a851eba516bf9d0f1f53c2a5eb58

Observation 0f979c3d-4478-4ecb-953e-6ca8318b1ec3 · outbound

This paper cites Kingma and Jimmy Ba.

Optimal Embedding Learning Rate in LLMs: The Effect of Vocabulary Size Kingma and Jimmy Ba

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-15T19:54:42.503330Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T19:54:42.503330Z digest=sha256:7a7426d5b466594a71240ee55fccd33b9ce3424f64a7881f737852665e34f601

Observation d8abe900-b9de-495f-bf1e-7662454d1e7f · outbound

This paper cites an unresolved cited work.

Optimal Embedding Learning Rate in LLMs: The Effect of Vocabulary Size Unresolved cited work

Reference 13

Resolution
unresolved
raw_fallback, observed 2026-08-15T19:54:43.056873Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-15T19:54:42.499441Z digest=sha256:ca4a835b6e0798347e61ef80ae66a94f7ed8d8f5083cb0f971d1747c7afeef85

Observation 3075e622-a266-4e46-b146-b163ba8bfff7 · outbound

This paper cites The llama 3 herd of models, 2024.

Optimal Embedding Learning Rate in LLMs: The Effect of Vocabulary Size The llama 3 herd of models, 2024

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-15T19:54:42.516190Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T19:54:42.516190Z digest=sha256:d73ef9826909b6f21ce78e9735b6dfcde0f320be1503deb54cef1b967f2f1e90

Observation 230ffc16-b8eb-421f-8d85-49e6a96e7ec7 · outbound

This paper cites Pointer sen- tinel mixture models, 2016.

Optimal Embedding Learning Rate in LLMs: The Effect of Vocabulary Size Pointer sen- tinel mixture models, 2016

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T19:54:43.028154Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-15T19:54:42.521875Z digest=sha256:8bdc57eb3aa2343be48f180f829a997ea53e0d65848ada45b32ac5f649c43293

Observation 4976277b-3b6d-46c9-ada3-c88d92d87a5f · outbound

This paper cites An Empirical Study of $\mu$P Learning Rate Transfer.

Optimal Embedding Learning Rate in LLMs: The Effect of Vocabulary Size An Empirical Study of $\mu$P Learning Rate Transfer

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-15T19:54:42.511780Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T19:54:42.511780Z digest=sha256:be50f0f1d5b5e3e06f6c7f30aad074d65ab0bf2933ffdf6e07c1ad52fccef56a

Observation a6df0e87-6b3a-459e-9930-6d91c199d3ae · outbound

This paper cites Poole, S.

Optimal Embedding Learning Rate in LLMs: The Effect of Vocabulary Size Poole, S

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T19:54:43.015329Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-15T19:54:42.530857Z digest=sha256:3d2ace480886cb8dcab021c8de05fcedf58536930b82efdba999c18d5d480528

Observation 07b7ee97-7df3-4b56-9441-1ac76142e1c4 · outbound

This paper cites Schoenholz, J.

Optimal Embedding Learning Rate in LLMs: The Effect of Vocabulary Size Schoenholz, J

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T19:54:43.001690Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-15T19:54:42.535008Z digest=sha256:d4b2e3199090a8ac6d4f6cfe3051e6593f90b2adb37fd855d9350e75a6dd746c

Observation 19eda179-91e6-419b-b004-cbde3f0c47a5 · outbound

This paper cites an unresolved cited work.

Optimal Embedding Learning Rate in LLMs: The Effect of Vocabulary Size Unresolved cited work

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-15T19:54:42.526095Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T19:54:42.526095Z digest=sha256:9787491bf562a44d6140aeb6d512f8643b430578a7a1b23ecd1a1af354072337

Observation 60dcbd05-82f5-42d8-827a-7668986fbb75 · outbound

This paper cites The Falcon Series of Open Language Models.

Optimal Embedding Learning Rate in LLMs: The Effect of Vocabulary Size The Falcon Series of Open Language Models

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-15T19:54:42.544080Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T19:54:42.544080Z digest=sha256:c0ada17521e8f31ca113013185641a38b0482e4d157f056317604d990716000b

Observation 2d3fc3b8-693c-4e3e-9bc2-7b06c8d81e88 · outbound

This paper cites Gemma 3 technical report, 2025.

Optimal Embedding Learning Rate in LLMs: The Effect of Vocabulary Size Gemma 3 technical report, 2025

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-15T19:54:42.548770Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T19:54:42.548770Z digest=sha256:7f1250696457c482e838bd5c10ca0fcabaa9925bcb9ead1900e36739dfb6c3c4

Observation f736ab14-e453-4cf0-811d-2a28c79e8fe8 · outbound

This paper cites Scaling Laws with Vocabulary: Larger Models Deserve Larger Vocabularies.

Optimal Embedding Learning Rate in LLMs: The Effect of Vocabulary Size Scaling Laws with Vocabulary: Larger Models Deserve Larger Vocabularies

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-15T19:54:42.539112Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T19:54:42.539112Z digest=sha256:111644c4c74b46666b63e84b08627cb7afd497bf269894af782f25343c93d52c

Observation ea63cfd8-d429-48e1-af50-f8099316e3ec · outbound

This paper cites Scaling Limits of Wide Neural Networks with Weight Sharing: Gaussian Process Behavior, Gradient Independence, and Neural Tangent Kernel Derivation.

Optimal Embedding Learning Rate in LLMs: The Effect of Vocabulary Size Scaling Limits of Wide Neural Networks with Weight Sharing: Gaussian Process Behavior, Gradient Independence, and Neural Tangent Kernel Derivation

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-15T19:54:42.557483Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T19:54:42.557483Z digest=sha256:9f9d536576561c3b95db5d6d4f99f56073fd081f3c57070fd844da64c3bed061

Observation 9469a6ce-f8dc-40e0-b5ca-7656daaba731 · outbound

This paper cites Tensor Programs V: Tuning Large Neural Networks via Zero-Shot Hyperparameter Transfer.

Optimal Embedding Learning Rate in LLMs: The Effect of Vocabulary Size Tensor Programs V: Tuning Large Neural Networks via Zero-Shot Hyperparameter Transfer

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-15T19:54:42.562346Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T19:54:42.562346Z digest=sha256:28094cc91df2835ff3e40a3f400c85d9ca17b6823e30b051050225eb464fc82d

Observation ec43db71-939e-494a-b99c-f35d63d9c7a3 · outbound

This paper cites Phi-3 Technical Report: A Highly Capable Language Model Locally on Your Phone.

Optimal Embedding Learning Rate in LLMs: The Effect of Vocabulary Size Phi-3 Technical Report: A Highly Capable Language Model Locally on Your Phone

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-15T19:54:42.552969Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T19:54:42.552969Z digest=sha256:875ea4a1e9c8f3024a6d95530b4db24db69f952e7f6826c7f8683e4a550f96c5

Observation 1d57d817-22eb-4cb8-a62c-a0f2eb0de283 · outbound

This paper cites How does critical batch size scale in pre-training?,.

Optimal Embedding Learning Rate in LLMs: The Effect of Vocabulary Size How does critical batch size scale in pre-training?,

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T19:54:42.981917Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-15T19:54:42.571451Z digest=sha256:59e951eecaeb804815b3e789eff2593aafbd2ffd5d70cbd69d8e9103a6bf0893

Observation 23af0345-ba61-4677-a09e-a194ec2b363b · outbound

This paper cites Selected Studies of the Principle of Relative Frequency in Language.

Optimal Embedding Learning Rate in LLMs: The Effect of Vocabulary Size Selected Studies of the Principle of Relative Frequency in Language

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T19:54:42.969342Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-15T19:54:42.579612Z digest=sha256:ae5094daffe1f9266212e5711da2c8a846b699e9fa779f444c5337c67f179531

Observation 301056cd-5a09-4409-b131-1c21484cc905 · outbound

This paper cites Tensor Programs VI: Feature Learning in Infinite-Depth Neural Networks.

Optimal Embedding Learning Rate in LLMs: The Effect of Vocabulary Size Tensor Programs VI: Feature Learning in Infinite-Depth Neural Networks

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-15T19:54:42.567261Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T19:54:42.567261Z digest=sha256:4b5cf5ff548cd906db7755f9bdb29dea3ad1658a75459dec4ef0849c9c699fdf

Observation 69665084-8ddd-457a-aa2c-ee73b03ebf63 · outbound

This paper cites Because Wj∼N (0,Im) and Mj =⟨v,Wj⟩, for anyi∈ [m] the pair (Mj,Wji) is jointly Gaussian with correlationρ := vi ∥v∥.

Optimal Embedding Learning Rate in LLMs: The Effect of Vocabulary Size Because Wj∼N (0,Im) and Mj =⟨v,Wj⟩, for anyi∈ [m] the pair (Mj,Wji) is jointly Gaussian with correlationρ := vi ∥v∥

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T19:54:42.956995Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-15T19:54:42.585134Z digest=sha256:87d14c79cf3e2a94e6651efd775a073c262d0494639f69691a667179a1326c66

Observation f171a23a-e471-4f34-8e9d-ba48be49c6c9 · outbound

This paper cites Still conditional on v, sign(Mj) is±1 with equal probability, independent of the magnitude of Wj.

Optimal Embedding Learning Rate in LLMs: The Effect of Vocabulary Size Still conditional on v, sign(Mj) is±1 with equal probability, independent of the magnitude of Wj

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T19:54:42.943251Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-15T19:54:42.590226Z digest=sha256:65ac113d4041d7c60f01e1f5ad836c2adff2628117ad7b83c24932ae0e868565

Observation a4358c98-7dd9-49d2-b588-4b01e82d6cf0 · outbound

This paper cites Because the Yj’s are conditionally independent, Ev Cov(X|v) = Ev h dX j=1 Cov(Yj|v) i =d Im− Ev µ(v)µ(v)⊤ | {z } = 2 πmIm = d− 2d πm Im.

Optimal Embedding Learning Rate in LLMs: The Effect of Vocabulary Size Because the Yj’s are conditionally independent, Ev Cov(X|v) = Ev h dX j=1 Cov(Yj|v) i =d Im− Ev µ(v)µ(v)⊤ | {z } = 2 πmIm = d− 2d πm Im

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T19:54:42.930270Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-15T19:54:42.597097Z digest=sha256:bef1da06e4b3aed99ec38c778932af22945d7ecf4ffd117898629133424bc3db

Observation 31f911b0-7376-4e10-a45a-49d7fdd94672 · outbound

This paper cites From Step1, E[X|v] =dµ (v), so Covv E[X|v] =d2 Covv µ(v) =d2 2 π Covv v ∥v∥ = 2d2 πmIm.

Optimal Embedding Learning Rate in LLMs: The Effect of Vocabulary Size From Step1, E[X|v] =dµ (v), so Covv E[X|v] =d2 Covv µ(v) =d2 2 π Covv v ∥v∥ = 2d2 πmIm

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T19:54:42.917670Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-15T19:54:42.601434Z digest=sha256:211f673e0149c4b8dbbe8117a3e4762b87ea9c3da28b52a27e1fa280a84a4755

Observation 99f7d6eb-a9a8-4b1c-8f78-015dd947a521 · outbound

This paper cites Conditioning on E, each entry of E⊤M a centered Gaussian variable, hence E[S(E⊤M)|E] = 0.

Optimal Embedding Learning Rate in LLMs: The Effect of Vocabulary Size Conditioning on E, each entry of E⊤M a centered Gaussian variable, hence E[S(E⊤M)|E] = 0

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T19:54:42.903671Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-15T19:54:42.606351Z digest=sha256:86aba21143bb86743b6b701baf5db73af512818cc24d4a7a56128ca8f8e89c7b

Observation e1760d3f-2b56-4eaa-be4c-e3aca666982c · outbound

This paper cites Fix a column index k.

Optimal Embedding Learning Rate in LLMs: The Effect of Vocabulary Size Fix a column index k

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T19:54:42.888341Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-15T19:54:42.610552Z digest=sha256:a3bc71316d2dfe81331433e4a9592e512c5ef13fff05e689403eb364b0c151f6

Observation ef577664-d8da-4277-9f3e-194bb3efff91 · outbound

This paper cites Different columns ofM (differentk) are independent, so Cov(X) is diagonal and each coordinate variance is the same as above.

Optimal Embedding Learning Rate in LLMs: The Effect of Vocabulary Size Different columns ofM (differentk) are independent, so Cov(X) is diagonal and each coordinate variance is the same as above

Reference 38

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T19:54:42.871604Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-15T19:54:42.614806Z digest=sha256:46d435649369d25cf8bcfbdd0c9c48d2e16c842dd897cfbb2f010a924f976648

Observation 52994ea5-ed58-4778-9f43-0c7fd4e17cd0 · outbound

This paper cites Adam: A Method for Stochastic Optimization.

Optimal Embedding Learning Rate in LLMs: The Effect of Vocabulary Size Adam: A Method for Stochastic Optimization

Reference 2017

Resolution
unresolved
no resolver link, observed 2026-08-15T19:54:42.507645Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T19:54:42.507645Z digest=sha256:30ec554fbea41edec1fb4cf8d59ec810b33956882f449de6e39cca4a47c9841a

Observation 3c8dd48f-a720-4c59-839c-43c18097175c · outbound

This paper cites Scaling Exponents Across Parameterizations and Optimizers.

Optimal Embedding Learning Rate in LLMs: The Effect of Vocabulary Size Scaling Exponents Across Parameterizations and Optimizers

Reference 2024

Resolution
unresolved
no resolver link, observed 2026-08-15T19:54:42.469149Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T19:54:42.469149Z digest=sha256:51dd35860063efb1b7108c16f282b83733d809ddbe75e194c819370b1ff75ae5

Observation 08d43654-9d72-4310-8e60-28d09a1f0868 · outbound

This paper cites How Does Critical Batch Size Scale in Pre-training?.

Optimal Embedding Learning Rate in LLMs: The Effect of Vocabulary Size How Does Critical Batch Size Scale in Pre-training?

Reference 2025

Resolution
unresolved
no resolver link, observed 2026-08-15T19:54:42.575268Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T19:54:42.575268Z digest=sha256:f2146e96973c347ea77a0be1f217102bd79c0ea9666782895be7577ef48a3027

Pith citing papers

Observation b12fee4f-c0fe-4629-9cfd-5c377a29cff1 · inbound

Quantifying Hyperparameter Transfer and the Importance of Embedding Layer Learning Rate cites this paper.

Quantifying Hyperparameter Transfer and the Importance of Embedding Layer Learning Rate Optimal Embedding Learning Rate in LLMs: The Effect of Vocabulary Size

Reference 17

Resolution
verified exact
arxiv_id, observed 2026-05-21T05:03:57.956920Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-05-21T05:01:27.451161Z digest=sha256:ca6295070c74cf6d4e8ff40e70ca420ea54061ef5508db0e05478d37e0f2d1b3

Observation 00fd0d3b-1820-4c62-ad19-72b38dc51853 · inbound

One LR Doesn't Fit All: Heavy-Tail Guided Layerwise Learning Rates for LLMs cites this paper.

One LR Doesn't Fit All: Heavy-Tail Guided Layerwise Learning Rates for LLMs Optimal Embedding Learning Rate in LLMs: The Effect of Vocabulary Size

Reference 3

Resolution
verified exact
arxiv_id, observed 2026-05-22T07:44:42.593237Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-05-22T07:44:16.677054Z digest=sha256:e9fe653601031921ce383f0ba2b2832a28555c0a07e246365ebad0431c0d696e

Observation bc462715-4679-4150-bef0-3cee690c782f · inbound

One LR Doesn't Fit All: Heavy-Tail Guided Layerwise Learning Rates for LLMs cites this paper.

One LR Doesn't Fit All: Heavy-Tail Guided Layerwise Learning Rates for LLMs Optimal Embedding Learning Rate in LLMs: The Effect of Vocabulary Size

Reference 3

Resolution
verified exact
arxiv_id, observed 2026-06-30T17:34:57.595450Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-06-30T17:31:32.533941Z digest=sha256:afae182926dd0a8a1e0bf8db4a377d99835d600352fdc83af4e8b999fdcdd5a6