Pith. sign in

Paper Citation Record · LEDGER

Is your batch size the problem? Revisiting the Adam-SGD gap in language modeling

As of 21 August 2026, this Paper Citation Record lists 30 of 30 outbound references and 6 inbound Pith citation observations for arXiv:2506.12543.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2506.12543 v1

Coverage vector

measured 30 of 30 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-07T00:50:30.649181Z

measured 36 of 36 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-21T06:32:19.484+00:00

measured 6 of 6 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-04T09:38:48.446865Z

measured 1 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-08-05T02:28:24.338817Z

Reference resolution

30 of 30 outbound references displayed

  • verified exact2
  • verified fuzzy2
  • unresolved25
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch1

External citation measurements

0
arxiv_reference, observed 2026-08-05T02:28:24.338817Z

Outbound references

Observation a21bd626-5831-4ece-acb5-d4840f36f7b9 · outbound

This paper cites Disentangling Adaptive Gradient Methods from Learning Rates.

Is your batch size the problem? Revisiting the Adam-SGD gap in language modeling Disentangling Adaptive Gradient Methods from Learning Rates

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-07T00:50:29.766980Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:50:29.766980Z digest=sha256:25ec27d20d0027be953f85d4b91f014005362da059860743f586a05239ed0550

Observation 49d4b4c9-e468-4c4f-b591-c01e47235889 · outbound

This paper cites As before, the gap decreases the longer we train, and SGD can eventually outperform Adam.

Is your batch size the problem? Revisiting the Adam-SGD gap in language modeling As before, the gap decreases the longer we train, and SGD can eventually outperform Adam

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:50:31.048280Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-07T00:50:30.645921Z digest=sha256:e81cf3a5cc4df167207ac0f9deb7e999162ffa2a743ec2e0ca1060ae21faf7cf

Observation 42de01fd-13d3-4930-a9bc-9a2017e833c3 · outbound

This paper cites On the Interaction of Batch Noise, Adaptivity, and Compression, under $(L_0,L_1)$-Smoothness: An SDE Approach.

Is your batch size the problem? Revisiting the Adam-SGD gap in language modeling On the Interaction of Batch Noise, Adaptivity, and Compression, under $(L_0,L_1)$-Smoothness: An SDE Approach

Reference 5

Resolution
verified exact
local_arxiv, observed 2026-08-07T00:50:30.995734Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-07T00:50:30.238100Z digest=sha256:103e1a09d93ad1addd87b250ee60fe756bea8fad5ac9a06ca6396d28d987d230

Observation bbf646a2-602d-4105-bf75-187a47e1043a · outbound

This paper cites Scaling Laws and Compute-Optimal Training Beyond Fixed Training Durations.

Is your batch size the problem? Revisiting the Adam-SGD gap in language modeling Scaling Laws and Compute-Optimal Training Beyond Fixed Training Durations

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-07T00:50:30.555421Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:50:30.555421Z digest=sha256:e33441a73e7b4e6ba711050ea5b7b8f5844f4cdbdf4991052c712623d39d0000

Observation faba8271-5cc9-4dd9-a404-e1ec57e7b9b5 · outbound

This paper cites How Does Adaptive Optimization Impact Local Neural Network Geometry?.

Is your batch size the problem? Revisiting the Adam-SGD gap in language modeling How Does Adaptive Optimization Impact Local Neural Network Geometry?

Reference 10

Resolution
metadata mismatch
local_arxiv, observed 2026-08-07T00:50:30.948721Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-07T00:50:30.563102Z digest=sha256:97daf7dcdd128baef8482f3d21b6195930563cc8a871f938c40fe82a9cc36de8

Observation 3ed2fc01-02dd-4f7e-831e-dbf28261e37e · outbound

This paper cites Adam: A Method for Stochastic Optimization.

Is your batch size the problem? Revisiting the Adam-SGD gap in language modeling Adam: A Method for Stochastic Optimization

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-07T00:50:30.570637Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:50:30.570637Z digest=sha256:43c547cb8a94a3f60ace2bcb06a0e35d78258860de05b3b2189e4c30bac13c94

Observation 93b8ace5-b1e8-41e7-976f-b4ee1f3a29da · outbound

This paper cites Heavy-Tailed Class Imbalance and Why Adam Outperforms Gradient Descent on Language Models.

Is your batch size the problem? Revisiting the Adam-SGD gap in language modeling Heavy-Tailed Class Imbalance and Why Adam Outperforms Gradient Descent on Language Models

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-07T00:50:30.584903Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:50:30.584903Z digest=sha256:0a269074d97cd10cdb4a199c967e90dd8c5c8e98dd76a0e0b6529a2e9c6b4795

Observation e4f2a746-919f-4dfd-9dad-d458551d844e · outbound

This paper cites Muon is Scalable for LLM Training.

Is your batch size the problem? Revisiting the Adam-SGD gap in language modeling Muon is Scalable for LLM Training

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-07T00:50:30.594639Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:50:30.594639Z digest=sha256:cad3a8de1e56a62d6727d306066d38e88bb18a581f9963c57a06de315e347415

Observation cf808663-47b5-4c1a-9afc-7e27f355ba34 · outbound

This paper cites On the SDEs and Scaling Rules for Adaptive Gradient Algorithms.

Is your batch size the problem? Revisiting the Adam-SGD gap in language modeling On the SDEs and Scaling Rules for Adaptive Gradient Algorithms

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-07T00:50:30.606742Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:50:30.606742Z digest=sha256:5906248637667abd8ea398cf53ede6033a97e1380163528374d4196917a1ca7d

Observation 3c264e59-9a2c-40dd-b301-bfb40bb9ea61 · outbound

This paper cites Orvieto and R.

Is your batch size the problem? Revisiting the Adam-SGD gap in language modeling Orvieto and R

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-07T00:50:30.612924Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:50:30.612924Z digest=sha256:e8e310e8b465ce4d3f6133a4e69dc21adab99779e59f55bb92d06c3127c1ec07

Observation c90daf0d-ba2c-4c9c-8d1c-cfba25a472cb · outbound

This paper cites Toward Understanding Why Adam Converges Faster Than SGD for Transformers.

Is your batch size the problem? Revisiting the Adam-SGD gap in language modeling Toward Understanding Why Adam Converges Faster Than SGD for Transformers

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-07T00:50:30.615997Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:50:30.615997Z digest=sha256:ed40148292b6b8fb4e4d8e2fa35d130b11f543d8bcc06c667421a47c80965339

Observation 69e7fee2-9ad5-4733-b8d8-8f7515dc23f4 · outbound

This paper cites arXiv:2502.00213 [cs].

Is your batch size the problem? Revisiting the Adam-SGD gap in language modeling arXiv:2502.00213 [cs]

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-07T00:50:30.624491Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:50:30.624491Z digest=sha256:cee64cfaeb648eac4098b430f0e7730ce2d3131af6fb3a1d1a2811642279968a

Observation ff2174f1-3d25-4dac-8c73-a60ed35c049d · outbound

This paper cites Adam Exploits $\ell_\infty$-geometry of Loss Landscape via Coordinate-wise Adaptivity.

Is your batch size the problem? Revisiting the Adam-SGD gap in language modeling Adam Exploits $\ell_\infty$-geometry of Loss Landscape via Coordinate-wise Adaptivity

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-07T00:50:30.627480Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:50:30.627480Z digest=sha256:7dd68d800c1782803c026a7d236fd7a1b17a649f6f0795279c4c0bbee6c76f9a

Observation 37a5bffc-f69f-4701-ab14-eb9cf9c9450d · outbound

This paper cites How Does Critical Batch Size Scale in Pre-training?.

Is your batch size the problem? Revisiting the Adam-SGD gap in language modeling How Does Critical Batch Size Scale in Pre-training?

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-07T00:50:30.630441Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:50:30.630441Z digest=sha256:c959f0cb5d0dfc11e0d6d959048ad407414c9c4dd028fa18b9e553d1c06c71b1

Observation f7dcaf01-f68f-4aad-a8d0-86ec4ad73dc9 · outbound

This paper cites Why Transformers Need Adam: A Hessian Perspective.

Is your batch size the problem? Revisiting the Adam-SGD gap in language modeling Why Transformers Need Adam: A Hessian Perspective

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-07T00:50:30.633781Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:50:30.633781Z digest=sha256:c934db75cb8215341385e5cc068d438766baa585bfd18860f0389bd15498d872

Observation d6d36155-b1fa-4519-8f90-785c1d93fa03 · outbound

This paper cites Deconstructing What Makes a Good Optimizer for Language Models.

Is your batch size the problem? Revisiting the Adam-SGD gap in language modeling Deconstructing What Makes a Good Optimizer for Language Models

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-07T00:50:30.636520Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:50:30.636520Z digest=sha256:ac3ce7f48a5a1c62145a8fd25bcc5f448200f6ef899af1fd8b860c5922552363

Observation a976dd82-e9e3-4506-b999-0f76593a6436 · outbound

This paper cites an unresolved cited work.

Is your batch size the problem? Revisiting the Adam-SGD gap in language modeling Unresolved cited work

Reference 27

Resolution
unresolved
raw_fallback, observed 2026-08-07T00:50:31.065026Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-07T00:50:30.639985Z digest=sha256:f452833a232fade5b487a0cb69d577a06af7b5d83af9b1017191d8a1c4cd604a

Observation 08b0145f-91e0-48b2-ae3e-3a5593e047d9 · outbound

This paper cites [2024], and uses the codebase of Orvieto and Gower [2025].

Is your batch size the problem? Revisiting the Adam-SGD gap in language modeling [2024], and uses the codebase of Orvieto and Gower [2025]

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:50:31.039339Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-07T00:50:30.649181Z digest=sha256:a5047c396f516afd3805786b138094da7a080c1964c05a24ca0b401c4f2c3396

Observation f34f1c0d-9218-4ca8-aec6-2a8131c3d5ec · outbound

This paper cites an unresolved cited work.

Is your batch size the problem? Revisiting the Adam-SGD gap in language modeling Unresolved cited work

Reference 1024

Resolution
unresolved
raw_fallback, observed 2026-08-07T00:50:31.056492Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-07T00:50:30.642968Z digest=sha256:fdd09d3799020d8fac52cfa42c3249eb8552eec36310124eb76ba903d938828d

Observation f176b608-2526-4295-95c2-1538cb7575ee · outbound

This paper cites Transformers without Tears: Improving the Normalization of Self-Attention.

Is your batch size the problem? Revisiting the Adam-SGD gap in language modeling Transformers without Tears: Improving the Normalization of Self-Attention

Reference 1986

Resolution
unresolved
no resolver link, observed 2026-08-07T00:50:30.609473Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:50:30.609473Z digest=sha256:4b212740693cc527640d1de2f336cd6bdecc7367d5426ca4399b49ca13524091

Observation 6799d912-962d-42ed-bda7-de130b44a4a0 · outbound

This paper cites How to Fine-Tune Vision Models with SGD.

Is your batch size the problem? Revisiting the Adam-SGD gap in language modeling How to Fine-Tune Vision Models with SGD

Reference 2014

Resolution
verified exact
local_arxiv, observed 2026-08-07T00:50:30.926338Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-07T00:50:30.576108Z digest=sha256:9e3a150e7d5a00599f6c278c41f9021dfd67f1bb124ce9207eb40d097c2018b0

Observation c597f601-3722-465a-851c-2e2ab7b14dcb · outbound

This paper cites The Llama 3 Herd of Models.

Is your batch size the problem? Revisiting the Adam-SGD gap in language modeling The Llama 3 Herd of Models

Reference 2017

Resolution
unresolved
no resolver link, observed 2026-08-07T00:50:30.548431Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:50:30.548431Z digest=sha256:d11c4626a75ffbe16cc1fbc3dc717aa1824850d783b77d71afe68051ebf79adb

Observation c4d8fd51-e618-4b44-9367-f940273b414c · outbound

This paper cites signSGD: Compressed Optimisation for Non-Convex Problems.

Is your batch size the problem? Revisiting the Adam-SGD gap in language modeling signSGD: Compressed Optimisation for Non-Convex Problems

Reference 2018

Resolution
unresolved
no resolver link, observed 2026-08-07T00:50:29.980013Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:50:29.980013Z digest=sha256:7f7634ab6d9073bbf04de29867847c579cd0e997d441d8fe68c3c0f8994a0248

Observation ef191c64-b70d-4e05-a88f-27d993caa9c9 · outbound

This paper cites Practical Efficiency of Muon for Pretraining.

Is your batch size the problem? Revisiting the Adam-SGD gap in language modeling Practical Efficiency of Muon for Pretraining

Reference 2019

Resolution
unresolved
no resolver link, observed 2026-08-07T00:50:30.618825Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:50:30.618825Z digest=sha256:8e78567a7c2999ef8d2499a34db6c730f632ee17ad6508ac8d9f01673f0eb5b9

Observation 4f04113b-7173-4087-a8fc-853e94d5a18b · outbound

This paper cites Accurate, Large Minibatch SGD: Training ImageNet in 1 Hour.

Is your batch size the problem? Revisiting the Adam-SGD gap in language modeling Accurate, Large Minibatch SGD: Training ImageNet in 1 Hour

Reference 2020

Resolution
unresolved
no resolver link, observed 2026-08-07T00:50:30.464304Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:50:30.464304Z digest=sha256:839b5f9a4e432953d6ae9eb01c725598a001abc327b402b4d562e32af8fdd41a

Observation 72b5ec18-fdd9-4bf8-b615-731d1aa0231f · outbound

This paper cites SGDR: Stochastic Gradient Descent with Warm Restarts.

Is your batch size the problem? Revisiting the Adam-SGD gap in language modeling SGDR: Stochastic Gradient Descent with Warm Restarts

Reference 2021

Resolution
unresolved
no resolver link, observed 2026-08-07T00:50:30.603923Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:50:30.603923Z digest=sha256:4c7940106f8ab0c7b39a2d4b31505a178ef52c7ece9b5af32b8d9edf12ae4df0

Observation a959df74-d31d-472b-b5f2-150c36d047b8 · outbound

This paper cites An Image is Worth 16x16 Words: Transformers for Image Recognition at Scale.

Is your batch size the problem? Revisiting the Adam-SGD gap in language modeling An Image is Worth 16x16 Words: Transformers for Image Recognition at Scale

Reference 2022

Resolution
unresolved
no resolver link, observed 2026-08-07T00:50:30.348876Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:50:30.348876Z digest=sha256:174742c0c79824ca1b4b00f7ed684bd4434c2672a0b1808cb803821d6c02d3a4

Observation 6cec26bc-75cf-40f6-8c17-d216703fe868 · outbound

This paper cites GPT-NeoX-20B: An Open-Source Autoregressive Language Model.

Is your batch size the problem? Revisiting the Adam-SGD gap in language modeling GPT-NeoX-20B: An Open-Source Autoregressive Language Model

Reference 2023

Resolution
unresolved
no resolver link, observed 2026-08-07T00:50:30.096081Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:50:30.096081Z digest=sha256:33b1e87c5aa92c9b0abb404da17e4158b787f9940d50e4d5d890c7d5b39ea019

Observation 3c56be0f-8610-4504-8510-4e9111b6f761 · outbound

This paper cites Linear attention is (maybe) all you need (to understand transformer optimization).

Is your batch size the problem? Revisiting the Adam-SGD gap in language modeling Linear attention is (maybe) all you need (to understand transformer optimization)

Reference 2024

Resolution
unresolved
no resolver link, observed 2026-08-07T00:50:29.838100Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:50:29.838100Z digest=sha256:1297a2de77607319d0c2984ec1e39af95932ca97e4b526b1b76c972a5bf5251b

Observation 0da4f519-d61b-47a1-8033-3430a73bbbff · outbound

This paper cites GLU Variants Improve Transformer.

Is your batch size the problem? Revisiting the Adam-SGD gap in language modeling GLU Variants Improve Transformer

Reference 2025

Resolution
unresolved
no resolver link, observed 2026-08-07T00:50:30.621589Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:50:30.621589Z digest=sha256:8334bfbc102d8efaf8508ea17915023b7e7834032691b474451d766f711d5d5d

Pith citing papers

Observation 34bb4621-ccf2-4792-bf5c-1e3ccfebbc71 · inbound

Seesaw: Accelerating Training by Balancing Learning Rate and Batch Size Scheduling cites this paper.

Seesaw: Accelerating Training by Balancing Learning Rate and Batch Size Scheduling Is your batch size the problem? Revisiting the Adam-SGD gap in language modeling

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-04T09:38:48.446865Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T09:38:48.446865Z digest=sha256:cf897411e482f143370a333530864dfc257ed256f6e6808532db3b52f11eeb8f

Observation 0d414ffe-531d-4c19-a57a-892ac77a246f · inbound

How to Scale Mixture-of-Experts: From muP to the Maximally Scale-Stable Parameterization cites this paper.

How to Scale Mixture-of-Experts: From muP to the Maximally Scale-Stable Parameterization Is your batch size the problem? Revisiting the Adam-SGD gap in language modeling

Reference 41

Resolution
metadata mismatch
arxiv_id, observed 2026-05-15T04:49:44.656518Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=arxiv_source observed=2026-05-15T04:45:20.091598Z digest=sha256:62d17486269f4032fb46fa00c747ba84485a0cf29ce13757e24f3cf96d5cc76f

Observation 01bd3cbc-d651-4a8c-bf5a-ad2e192dce3d · inbound

Revisiting the Adam-SGD Gap in LLM Pre-Training: The Role of Large Effective Learning Rates cites this paper.

Revisiting the Adam-SGD Gap in LLM Pre-Training: The Role of Large Effective Learning Rates Is your batch size the problem? Revisiting the Adam-SGD gap in language modeling

Reference 9

Resolution
verified exact
arxiv_id, observed 2026-05-20T13:38:19.427606Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-05-20T13:34:09.129379Z digest=sha256:fc676d4fc0b6a1bc0a0629ab6a71bc1fa07a96d950d88fa8650af8c83a27a077

Observation 53449721-9569-4766-a9a1-050f989768bf · inbound

Why SGD is not Brownian Motion: A New Perspective on Stochastic Dynamics cites this paper.

Why SGD is not Brownian Motion: A New Perspective on Stochastic Dynamics Is your batch size the problem? Revisiting the Adam-SGD gap in language modeling

Reference 235

Resolution
verified exact
arxiv_id, observed 2026-05-22T08:11:17.363127Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=arxiv_source observed=2026-05-22T08:06:52.309619Z digest=sha256:81701a410279df8f10439e731a7cb71379eabb2bfa075e34966a81f907450c67

Observation 3aa9feae-5033-46d6-8a2e-f130184609b1 · inbound

A Theoretical and Experimental Study of a Novel Adaptive Learning Algorithm cites this paper.

A Theoretical and Experimental Study of a Novel Adaptive Learning Algorithm Is your batch size the problem? Revisiting the Adam-SGD gap in language modeling

Reference 11

Resolution
verified exact
arxiv_id, observed 2026-06-29T09:23:16.439125Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-06-29T09:18:31.526770Z digest=sha256:d235602db29da24f0421c5a7a08e1beba3a78cb62faa059302c2fefd1f1e32dd

Observation e9e7170e-6db5-4851-b0e2-92c553ad4910 · inbound

Quantifying the Agreement Between Data-Influence and Data-Similarity to Understand LLM Behavior cites this paper.

Quantifying the Agreement Between Data-Influence and Data-Similarity to Understand LLM Behavior Is your batch size the problem? Revisiting the Adam-SGD gap in language modeling

Reference 48

Resolution
verified exact
arxiv_id, observed 2026-06-26T08:49:14.928649Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=arxiv_source observed=2026-06-26T08:45:34.884703Z digest=sha256:8c53a86d021074f3a435d16e5812e40886ef0394ff2e0f09465f36e601ecb4f3