Pith. sign in

Paper Citation Record · LEDGER

Optimal Embedding Learning Rate in LLMs: The Effect of Vocabulary Size

As of 16 August 2026, this Paper Citation Record lists 38 of 38 outbound references and 3 inbound Pith citation observations for arXiv:2506.15025.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2506.15025 v1

Coverage vector

measured 38 of 38 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-15T19:54:42.614806Z

measured 41 of 41 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-16T06:30:59.297886+00:00

measured 3 of 3 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-06-30T17:31:32.533941Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-06-30T17:34:57.593618Z

Reference resolution

38 of 38 outbound references displayed

  • verified exact0
  • verified fuzzy17
  • unresolved21
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation b3f8b846-6e14-4517-be05-f687fba975d8 · outbound

This paper cites u-$\mu$P: The Unit-Scaled Maximal Update Parametrization.

Optimal Embedding Learning Rate in LLMs: The Effect of Vocabulary Size u-$\mu$P: The Unit-Scaled Maximal Update Parametrization

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-15T19:54:42.444889Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T19:54:42.444889Z digest=sha256:a9f40996ccfd3ef89e6ee4d440446dfa759778b549cc7718949e94be312f1b75

Observation b4fc21cb-ea38-4af4-afab-63a2f63c7b5e · outbound

This paper cites Depthwise Hyperparameter Transfer in Residual Networks: Dynamics and Scaling Limit.

Optimal Embedding Learning Rate in LLMs: The Effect of Vocabulary Size Depthwise Hyperparameter Transfer in Residual Networks: Dynamics and Scaling Limit

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-15T19:54:42.450340Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T19:54:42.450340Z digest=sha256:784431b9c389d0e9bc3b60899e87ffd3c14e2a1772aeb89f5a45bcc96233e60a

Observation 4e9517dd-5fdd-4a8e-91e5-5d50efc79d17 · outbound

This paper cites On Lazy Training in Differentiable Programming.

Optimal Embedding Learning Rate in LLMs: The Effect of Vocabulary Size On Lazy Training in Differentiable Programming

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-15T19:54:42.454722Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T19:54:42.454722Z digest=sha256:cd73a0f6a9b67412a809df9ef4599fdda1802826d266f057030a98647ba83707

Observation 2e495ec2-61aa-4a99-b5c0-c7a317109251 · outbound

This paper cites Infinite- width limit of deep linear neural networks.

Optimal Embedding Learning Rate in LLMs: The Effect of Vocabulary Size Infinite- width limit of deep linear neural networks

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-15T19:54:42.459894Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T19:54:42.459894Z digest=sha256:f6da7a3877ea489cd27a713b1469c09e4353bf3f25bb3717890bc089d22f3d2a

Observation 52002095-c746-4854-844c-d90658e74767 · outbound

This paper cites Alemi, Roman Novak, Peter J.

Optimal Embedding Learning Rate in LLMs: The Effect of Vocabulary Size Alemi, Roman Novak, Peter J

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T19:54:43.129312Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-15T19:54:42.464142Z digest=sha256:932ec7665effb3563a293d4b719a76fefce02dc6bea7df870669f2a7f7b394af

Observation 82f197d9-4eb3-444a-83c2-647fe2f7f6c1 · outbound

This paper cites On the infinite-depth limit of finite-width neural networks.

Optimal Embedding Learning Rate in LLMs: The Effect of Vocabulary Size On the infinite-depth limit of finite-width neural networks

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T19:54:43.115696Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-15T19:54:42.473690Z digest=sha256:6e74f3319d98731ab644ba08c617c0956eda633ce21c54914fdf6afec63630cd

Observation 87e1e4de-8cd6-4997-a736-fe3820039dec · outbound

This paper cites On the impact of the ac- tivation function on deep neural networks training.

Optimal Embedding Learning Rate in LLMs: The Effect of Vocabulary Size On the impact of the ac- tivation function on deep neural networks training

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T19:54:43.102018Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-15T19:54:42.477648Z digest=sha256:d0d6c1c67bc5ee505b8309a0cdbe73100a372cb2bba0c7c1eac8f823f04cbfd2

Observation 2f285087-ba6c-4b17-946f-12894b8c479c · outbound

This paper cites Stable resnet.

Optimal Embedding Learning Rate in LLMs: The Effect of Vocabulary Size Stable resnet

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T19:54:43.087750Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-15T19:54:42.482191Z digest=sha256:9843cf72c4cc43157c55e19681d0b4f1b5808542938ddedab1bbdf36e278c3ac

Observation b5cf8226-42e4-4f65-8ba2-6adf96c288c4 · outbound

This paper cites Delving Deep into Rectifiers: Surpassing Human-Level Performance on ImageNet Classification.

Optimal Embedding Learning Rate in LLMs: The Effect of Vocabulary Size Delving Deep into Rectifiers: Surpassing Human-Level Performance on ImageNet Classification

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-15T19:54:42.485996Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T19:54:42.485996Z digest=sha256:7c64a1358261c57a159b920c871ebc8a5cfa51fe887c80fe3583c8f4c1c674b2

Observation fd4661cb-4347-492a-aca5-915c0d87ec2a · outbound

This paper cites Neural Tangent Kernel: Convergence and Generalization in Neural Networks.

Optimal Embedding Learning Rate in LLMs: The Effect of Vocabulary Size Neural Tangent Kernel: Convergence and Generalization in Neural Networks

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-15T19:54:42.490672Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T19:54:42.490672Z digest=sha256:abe337356e4c895a9406db9c89571b35d3ac04d67cf9a840cd5b475b1ab6d704

Observation 590b4620-eaa3-47ac-be14-15a5c3c5747e · outbound

This paper cites Muon: An optimizer for hidden layers in neural networks,.

Optimal Embedding Learning Rate in LLMs: The Effect of Vocabulary Size Muon: An optimizer for hidden layers in neural networks,

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T19:54:43.070446Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-15T19:54:42.495390Z digest=sha256:edd1c0c7000bc11436c809ebcc171358b1252a47277567351004440233879693

Observation 0f979c3d-4478-4ecb-953e-6ca8318b1ec3 · outbound

This paper cites Kingma and Jimmy Ba.

Optimal Embedding Learning Rate in LLMs: The Effect of Vocabulary Size Kingma and Jimmy Ba

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-15T19:54:42.503330Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T19:54:42.503330Z digest=sha256:7a7426d5b466594a71240ee55fccd33b9ce3424f64a7881f737852665e34f601

Observation d8abe900-b9de-495f-bf1e-7662454d1e7f · outbound

This paper cites an unresolved cited work.

Optimal Embedding Learning Rate in LLMs: The Effect of Vocabulary Size Unresolved cited work

Reference 13

Resolution
unresolved
raw_fallback, observed 2026-08-15T19:54:43.056873Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-15T19:54:42.499441Z digest=sha256:5679b2fc55949250ad90cdd0aeffa6b971817953360a3cae40d1f9f26cb9062a

Observation 3075e622-a266-4e46-b146-b163ba8bfff7 · outbound

This paper cites The llama 3 herd of models, 2024.

Optimal Embedding Learning Rate in LLMs: The Effect of Vocabulary Size The llama 3 herd of models, 2024

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-15T19:54:42.516190Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T19:54:42.516190Z digest=sha256:d73ef9826909b6f21ce78e9735b6dfcde0f320be1503deb54cef1b967f2f1e90

Observation 230ffc16-b8eb-421f-8d85-49e6a96e7ec7 · outbound

This paper cites Pointer sen- tinel mixture models, 2016.

Optimal Embedding Learning Rate in LLMs: The Effect of Vocabulary Size Pointer sen- tinel mixture models, 2016

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T19:54:43.028154Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-15T19:54:42.521875Z digest=sha256:9a6b9f3fea048628e67696e7ea1767186569d377fab826fa8f3931ed5a64e053

Observation 4976277b-3b6d-46c9-ada3-c88d92d87a5f · outbound

This paper cites An Empirical Study of $\mu$P Learning Rate Transfer.

Optimal Embedding Learning Rate in LLMs: The Effect of Vocabulary Size An Empirical Study of $\mu$P Learning Rate Transfer

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-15T19:54:42.511780Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T19:54:42.511780Z digest=sha256:6ddb575d3fb296509b44db53ab627a0edd4f3ae9dea64ace2adc1521f1897ca4

Observation a6df0e87-6b3a-459e-9930-6d91c199d3ae · outbound

This paper cites Poole, S.

Optimal Embedding Learning Rate in LLMs: The Effect of Vocabulary Size Poole, S

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T19:54:43.015329Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-15T19:54:42.530857Z digest=sha256:daba7d7d734d5ebb6b703b0c3681a5d7be0ac38f8845b4ac85f08a6f747f4f00

Observation 07b7ee97-7df3-4b56-9441-1ac76142e1c4 · outbound

This paper cites Schoenholz, J.

Optimal Embedding Learning Rate in LLMs: The Effect of Vocabulary Size Schoenholz, J

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T19:54:43.001690Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-15T19:54:42.535008Z digest=sha256:e43c38963598d8f5d5336783c2154944e7745a8cc6da3b642e20b7ac5c59d681

Observation 19eda179-91e6-419b-b004-cbde3f0c47a5 · outbound

This paper cites an unresolved cited work.

Optimal Embedding Learning Rate in LLMs: The Effect of Vocabulary Size Unresolved cited work

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-15T19:54:42.526095Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T19:54:42.526095Z digest=sha256:9787491bf562a44d6140aeb6d512f8643b430578a7a1b23ecd1a1af354072337

Observation 60dcbd05-82f5-42d8-827a-7668986fbb75 · outbound

This paper cites The Falcon Series of Open Language Models.

Optimal Embedding Learning Rate in LLMs: The Effect of Vocabulary Size The Falcon Series of Open Language Models

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-15T19:54:42.544080Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T19:54:42.544080Z digest=sha256:c0ada17521e8f31ca113013185641a38b0482e4d157f056317604d990716000b

Observation 2d3fc3b8-693c-4e3e-9bc2-7b06c8d81e88 · outbound

This paper cites Gemma 3 technical report, 2025.

Optimal Embedding Learning Rate in LLMs: The Effect of Vocabulary Size Gemma 3 technical report, 2025

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-15T19:54:42.548770Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T19:54:42.548770Z digest=sha256:7f1250696457c482e838bd5c10ca0fcabaa9925bcb9ead1900e36739dfb6c3c4

Observation f736ab14-e453-4cf0-811d-2a28c79e8fe8 · outbound

This paper cites Scaling Laws with Vocabulary: Larger Models Deserve Larger Vocabularies.

Optimal Embedding Learning Rate in LLMs: The Effect of Vocabulary Size Scaling Laws with Vocabulary: Larger Models Deserve Larger Vocabularies

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-15T19:54:42.539112Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T19:54:42.539112Z digest=sha256:111644c4c74b46666b63e84b08627cb7afd497bf269894af782f25343c93d52c

Observation ea63cfd8-d429-48e1-af50-f8099316e3ec · outbound

This paper cites Scaling Limits of Wide Neural Networks with Weight Sharing: Gaussian Process Behavior, Gradient Independence, and Neural Tangent Kernel Derivation.

Optimal Embedding Learning Rate in LLMs: The Effect of Vocabulary Size Scaling Limits of Wide Neural Networks with Weight Sharing: Gaussian Process Behavior, Gradient Independence, and Neural Tangent Kernel Derivation

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-15T19:54:42.557483Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T19:54:42.557483Z digest=sha256:9f9d536576561c3b95db5d6d4f99f56073fd081f3c57070fd844da64c3bed061

Observation 9469a6ce-f8dc-40e0-b5ca-7656daaba731 · outbound

This paper cites Tensor Programs V: Tuning Large Neural Networks via Zero-Shot Hyperparameter Transfer.

Optimal Embedding Learning Rate in LLMs: The Effect of Vocabulary Size Tensor Programs V: Tuning Large Neural Networks via Zero-Shot Hyperparameter Transfer

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-15T19:54:42.562346Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T19:54:42.562346Z digest=sha256:28094cc91df2835ff3e40a3f400c85d9ca17b6823e30b051050225eb464fc82d

Observation ec43db71-939e-494a-b99c-f35d63d9c7a3 · outbound

This paper cites Phi-3 Technical Report: A Highly Capable Language Model Locally on Your Phone.

Optimal Embedding Learning Rate in LLMs: The Effect of Vocabulary Size Phi-3 Technical Report: A Highly Capable Language Model Locally on Your Phone

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-15T19:54:42.552969Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T19:54:42.552969Z digest=sha256:da07485c98d7fd399c0f7aecc92c2afcdfed46e97c29ba7d60d6526bd3bfefc1

Observation 1d57d817-22eb-4cb8-a62c-a0f2eb0de283 · outbound

This paper cites How does critical batch size scale in pre-training?,.

Optimal Embedding Learning Rate in LLMs: The Effect of Vocabulary Size How does critical batch size scale in pre-training?,

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T19:54:42.981917Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-15T19:54:42.571451Z digest=sha256:1eb9e98c7dc5b18ff2671a8c967e769076212087089136ba88c1e19c54433607

Observation 23af0345-ba61-4677-a09e-a194ec2b363b · outbound

This paper cites Selected Studies of the Principle of Relative Frequency in Language.

Optimal Embedding Learning Rate in LLMs: The Effect of Vocabulary Size Selected Studies of the Principle of Relative Frequency in Language

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T19:54:42.969342Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-15T19:54:42.579612Z digest=sha256:67916160faf976bafefc97f075c3a9d94845272b1e295951b98de9780c22364e

Observation 301056cd-5a09-4409-b131-1c21484cc905 · outbound

This paper cites Tensor Programs VI: Feature Learning in Infinite-Depth Neural Networks.

Optimal Embedding Learning Rate in LLMs: The Effect of Vocabulary Size Tensor Programs VI: Feature Learning in Infinite-Depth Neural Networks

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-15T19:54:42.567261Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T19:54:42.567261Z digest=sha256:4b5cf5ff548cd906db7755f9bdb29dea3ad1658a75459dec4ef0849c9c699fdf

Observation 69665084-8ddd-457a-aa2c-ee73b03ebf63 · outbound

This paper cites Because Wj∼N (0,Im) and Mj =⟨v,Wj⟩, for anyi∈ [m] the pair (Mj,Wji) is jointly Gaussian with correlationρ := vi ∥v∥.

Optimal Embedding Learning Rate in LLMs: The Effect of Vocabulary Size Because Wj∼N (0,Im) and Mj =⟨v,Wj⟩, for anyi∈ [m] the pair (Mj,Wji) is jointly Gaussian with correlationρ := vi ∥v∥

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T19:54:42.956995Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-15T19:54:42.585134Z digest=sha256:50d40cbad42631894806730b5a4ab19337d7a84a30ded5f21b97079f4df1cec2

Observation f171a23a-e471-4f34-8e9d-ba48be49c6c9 · outbound

This paper cites Still conditional on v, sign(Mj) is±1 with equal probability, independent of the magnitude of Wj.

Optimal Embedding Learning Rate in LLMs: The Effect of Vocabulary Size Still conditional on v, sign(Mj) is±1 with equal probability, independent of the magnitude of Wj

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T19:54:42.943251Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-15T19:54:42.590226Z digest=sha256:6b1d2866a47436faaf579df218bd1b2e65a4a6471e4f900d5fb01d2b718b735f

Observation a4358c98-7dd9-49d2-b588-4b01e82d6cf0 · outbound

This paper cites Because the Yj’s are conditionally independent, Ev Cov(X|v) = Ev h dX j=1 Cov(Yj|v) i =d Im− Ev µ(v)µ(v)⊤ | {z } = 2 πmIm = d− 2d πm Im.

Optimal Embedding Learning Rate in LLMs: The Effect of Vocabulary Size Because the Yj’s are conditionally independent, Ev Cov(X|v) = Ev h dX j=1 Cov(Yj|v) i =d Im− Ev µ(v)µ(v)⊤ | {z } = 2 πmIm = d− 2d πm Im

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T19:54:42.930270Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-15T19:54:42.597097Z digest=sha256:a4d3a9f93756bc855e5524a2b037c3543a7f11312a62a10afa81c671ac6e78f3

Observation 31f911b0-7376-4e10-a45a-49d7fdd94672 · outbound

This paper cites From Step1, E[X|v] =dµ (v), so Covv E[X|v] =d2 Covv µ(v) =d2 2 π Covv v ∥v∥ = 2d2 πmIm.

Optimal Embedding Learning Rate in LLMs: The Effect of Vocabulary Size From Step1, E[X|v] =dµ (v), so Covv E[X|v] =d2 Covv µ(v) =d2 2 π Covv v ∥v∥ = 2d2 πmIm

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T19:54:42.917670Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-15T19:54:42.601434Z digest=sha256:765fe5e9de4af724b493dd8d001883431469ef8cd9c508d4cb30bc34cf7110f5

Observation 99f7d6eb-a9a8-4b1c-8f78-015dd947a521 · outbound

This paper cites Conditioning on E, each entry of E⊤M a centered Gaussian variable, hence E[S(E⊤M)|E] = 0.

Optimal Embedding Learning Rate in LLMs: The Effect of Vocabulary Size Conditioning on E, each entry of E⊤M a centered Gaussian variable, hence E[S(E⊤M)|E] = 0

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T19:54:42.903671Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-15T19:54:42.606351Z digest=sha256:e67806f990117798e8f4302483534565881219db9c1575789752534a0fb67065

Observation e1760d3f-2b56-4eaa-be4c-e3aca666982c · outbound

This paper cites Fix a column index k.

Optimal Embedding Learning Rate in LLMs: The Effect of Vocabulary Size Fix a column index k

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T19:54:42.888341Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-15T19:54:42.610552Z digest=sha256:5813d9c4511e114ab5795fb47c24171c8ce24206d426a3a12640cf07d1c26590

Observation ef577664-d8da-4277-9f3e-194bb3efff91 · outbound

This paper cites Different columns ofM (differentk) are independent, so Cov(X) is diagonal and each coordinate variance is the same as above.

Optimal Embedding Learning Rate in LLMs: The Effect of Vocabulary Size Different columns ofM (differentk) are independent, so Cov(X) is diagonal and each coordinate variance is the same as above

Reference 38

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T19:54:42.871604Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-15T19:54:42.614806Z digest=sha256:0f0d4c4f4743a53c043ecc0f5c75312fe481e407c81e93672ba8166ee2b94376

Observation 52994ea5-ed58-4778-9f43-0c7fd4e17cd0 · outbound

This paper cites Adam: A Method for Stochastic Optimization.

Optimal Embedding Learning Rate in LLMs: The Effect of Vocabulary Size Adam: A Method for Stochastic Optimization

Reference 2017

Resolution
unresolved
no resolver link, observed 2026-08-15T19:54:42.507645Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T19:54:42.507645Z digest=sha256:65cb2cd9e558eec1f1626d9f6acf8ce7a29052649072bf9ffad96a184207b13e

Observation 3c8dd48f-a720-4c59-839c-43c18097175c · outbound

This paper cites Scaling Exponents Across Parameterizations and Optimizers.

Optimal Embedding Learning Rate in LLMs: The Effect of Vocabulary Size Scaling Exponents Across Parameterizations and Optimizers

Reference 2024

Resolution
unresolved
no resolver link, observed 2026-08-15T19:54:42.469149Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T19:54:42.469149Z digest=sha256:51dd35860063efb1b7108c16f282b83733d809ddbe75e194c819370b1ff75ae5

Observation 08d43654-9d72-4310-8e60-28d09a1f0868 · outbound

This paper cites How Does Critical Batch Size Scale in Pre-training?.

Optimal Embedding Learning Rate in LLMs: The Effect of Vocabulary Size How Does Critical Batch Size Scale in Pre-training?

Reference 2025

Resolution
unresolved
no resolver link, observed 2026-08-15T19:54:42.575268Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T19:54:42.575268Z digest=sha256:f2146e96973c347ea77a0be1f217102bd79c0ea9666782895be7577ef48a3027

Pith citing papers

Observation b12fee4f-c0fe-4629-9cfd-5c377a29cff1 · inbound

Quantifying Hyperparameter Transfer and the Importance of Embedding Layer Learning Rate cites this paper.

Quantifying Hyperparameter Transfer and the Importance of Embedding Layer Learning Rate Optimal Embedding Learning Rate in LLMs: The Effect of Vocabulary Size

Reference 17

Resolution
verified exact
arxiv_id, observed 2026-05-21T05:03:57.956920Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-05-21T05:01:27.451161Z digest=sha256:85116861b603c2a2004ac23699c5576b7796c6e29714672cc06497d671b25c52

Observation 00fd0d3b-1820-4c62-ad19-72b38dc51853 · inbound

One LR Doesn't Fit All: Heavy-Tail Guided Layerwise Learning Rates for LLMs cites this paper.

One LR Doesn't Fit All: Heavy-Tail Guided Layerwise Learning Rates for LLMs Optimal Embedding Learning Rate in LLMs: The Effect of Vocabulary Size

Reference 3

Resolution
verified exact
arxiv_id, observed 2026-05-22T07:44:42.593237Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-05-22T07:44:16.677054Z digest=sha256:d1180e6c841883c6a560db8eda7a8a35e946796ec98bb666d306a1278c3be17d

Observation bc462715-4679-4150-bef0-3cee690c782f · inbound

One LR Doesn't Fit All: Heavy-Tail Guided Layerwise Learning Rates for LLMs cites this paper.

One LR Doesn't Fit All: Heavy-Tail Guided Layerwise Learning Rates for LLMs Optimal Embedding Learning Rate in LLMs: The Effect of Vocabulary Size

Reference 3

Resolution
verified exact
arxiv_id, observed 2026-06-30T17:34:57.595450Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-06-30T17:31:32.533941Z digest=sha256:9bd75f0547d608d055ce1613e55deeaf573c8e347d40479d4f40ca6e85ee2aad