Pith. sign in

Paper Citation Record · LEDGER

Is your batch size the problem? Revisiting the Adam-SGD gap in language modeling

As of 9 August 2026, this Paper Citation Record lists 30 of 30 outbound references and 6 inbound Pith citation observations for arXiv:2506.12543.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2506.12543 v1

Coverage vector

measured 30 of 30 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-07T00:50:30.649181Z

measured 36 of 36 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-09T06:31:02.800959+00:00

measured 6 of 6 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-04T09:38:48.446865Z

measured 1 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-08-05T02:28:24.338817Z

Reference resolution

30 of 30 outbound references displayed

  • verified exact2
  • verified fuzzy2
  • unresolved25
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch1

External citation measurements

0
arxiv_reference, observed 2026-08-05T02:28:24.338817Z

Outbound references

Observation a21bd626-5831-4ece-acb5-d4840f36f7b9 · outbound

This paper cites Disentangling Adaptive Gradient Methods from Learning Rates.

Is your batch size the problem? Revisiting the Adam-SGD gap in language modeling Disentangling Adaptive Gradient Methods from Learning Rates

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-07T00:50:29.766980Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:50:29.766980Z digest=sha256:9fa1f3a9e7d71c91558191fe5bd463efaaf8055f9a78ecea06a039321ce89934

Observation 49d4b4c9-e468-4c4f-b591-c01e47235889 · outbound

This paper cites As before, the gap decreases the longer we train, and SGD can eventually outperform Adam.

Is your batch size the problem? Revisiting the Adam-SGD gap in language modeling As before, the gap decreases the longer we train, and SGD can eventually outperform Adam

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:50:31.048280Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T00:50:30.645921Z digest=sha256:7bd80877c3a9e1d3785268b729e4f6cc9393d91c04eb4e6b29ee6e0201db5c71

Observation 42de01fd-13d3-4930-a9bc-9a2017e833c3 · outbound

This paper cites On the Interaction of Batch Noise, Adaptivity, and Compression, under $(L_0,L_1)$-Smoothness: An SDE Approach.

Is your batch size the problem? Revisiting the Adam-SGD gap in language modeling On the Interaction of Batch Noise, Adaptivity, and Compression, under $(L_0,L_1)$-Smoothness: An SDE Approach

Reference 5

Resolution
verified exact
local_arxiv, observed 2026-08-07T00:50:30.995734Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T00:50:30.238100Z digest=sha256:a1ee0cb31338ae2a32f628924f60ed644f797082d49b17da59d934203f99658f

Observation bbf646a2-602d-4105-bf75-187a47e1043a · outbound

This paper cites Scaling Laws and Compute-Optimal Training Beyond Fixed Training Durations.

Is your batch size the problem? Revisiting the Adam-SGD gap in language modeling Scaling Laws and Compute-Optimal Training Beyond Fixed Training Durations

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-07T00:50:30.555421Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:50:30.555421Z digest=sha256:58c51b399e0aa688ecce55d34e76f496ae4c26476aff57f7e295313f46f3e227

Observation faba8271-5cc9-4dd9-a404-e1ec57e7b9b5 · outbound

This paper cites How Does Adaptive Optimization Impact Local Neural Network Geometry?.

Is your batch size the problem? Revisiting the Adam-SGD gap in language modeling How Does Adaptive Optimization Impact Local Neural Network Geometry?

Reference 10

Resolution
metadata mismatch
local_arxiv, observed 2026-08-07T00:50:30.948721Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T00:50:30.563102Z digest=sha256:d66a296508b10ad59e1c52002a92a9601515f6c8d810ccd1d57e259e7bdb11ab

Observation 3ed2fc01-02dd-4f7e-831e-dbf28261e37e · outbound

This paper cites Adam: A Method for Stochastic Optimization.

Is your batch size the problem? Revisiting the Adam-SGD gap in language modeling Adam: A Method for Stochastic Optimization

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-07T00:50:30.570637Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:50:30.570637Z digest=sha256:5e31468f014b0d41723fb219ad14a35da87c85087b676f355153abf7ea4d9705

Observation 93b8ace5-b1e8-41e7-976f-b4ee1f3a29da · outbound

This paper cites Heavy-Tailed Class Imbalance and Why Adam Outperforms Gradient Descent on Language Models.

Is your batch size the problem? Revisiting the Adam-SGD gap in language modeling Heavy-Tailed Class Imbalance and Why Adam Outperforms Gradient Descent on Language Models

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-07T00:50:30.584903Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:50:30.584903Z digest=sha256:3161ad8ac57f8710205e7d94598a97f4c7a7043e6d109df9019f26ba97bd0046

Observation e4f2a746-919f-4dfd-9dad-d458551d844e · outbound

This paper cites Muon is Scalable for LLM Training.

Is your batch size the problem? Revisiting the Adam-SGD gap in language modeling Muon is Scalable for LLM Training

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-07T00:50:30.594639Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:50:30.594639Z digest=sha256:9fb02496d586bdd27aa9f784bf77ada28837d1533c112b53e3559ad39c1d24e8

Observation cf808663-47b5-4c1a-9afc-7e27f355ba34 · outbound

This paper cites On the SDEs and Scaling Rules for Adaptive Gradient Algorithms.

Is your batch size the problem? Revisiting the Adam-SGD gap in language modeling On the SDEs and Scaling Rules for Adaptive Gradient Algorithms

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-07T00:50:30.606742Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:50:30.606742Z digest=sha256:b6b1c6d6453a5c96867c811f35f97fc42ca896820373f7e7e59638c2b109c2b3

Observation 3c264e59-9a2c-40dd-b301-bfb40bb9ea61 · outbound

This paper cites Orvieto and R.

Is your batch size the problem? Revisiting the Adam-SGD gap in language modeling Orvieto and R

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-07T00:50:30.612924Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:50:30.612924Z digest=sha256:4157a0fcd76f6de8ea5148d394406bb8cb9552324ea58e31ff221dcbd9321465

Observation c90daf0d-ba2c-4c9c-8d1c-cfba25a472cb · outbound

This paper cites Toward Understanding Why Adam Converges Faster Than SGD for Transformers.

Is your batch size the problem? Revisiting the Adam-SGD gap in language modeling Toward Understanding Why Adam Converges Faster Than SGD for Transformers

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-07T00:50:30.615997Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:50:30.615997Z digest=sha256:298562269e5cec567ff60c0cb840adb87644fa3f207089bac9fa21105828106c

Observation 69e7fee2-9ad5-4733-b8d8-8f7515dc23f4 · outbound

This paper cites arXiv:2502.00213 [cs].

Is your batch size the problem? Revisiting the Adam-SGD gap in language modeling arXiv:2502.00213 [cs]

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-07T00:50:30.624491Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:50:30.624491Z digest=sha256:599868c82c042351c5491b39aa206bfe720b04808529b0bae56b09d28ff0ea56

Observation ff2174f1-3d25-4dac-8c73-a60ed35c049d · outbound

This paper cites Adam Exploits $\ell_\infty$-geometry of Loss Landscape via Coordinate-wise Adaptivity.

Is your batch size the problem? Revisiting the Adam-SGD gap in language modeling Adam Exploits $\ell_\infty$-geometry of Loss Landscape via Coordinate-wise Adaptivity

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-07T00:50:30.627480Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:50:30.627480Z digest=sha256:36177fe775e0db124d2c0bd605d9b4792a30f45e6e91ae18f24c247681ba48dc

Observation 37a5bffc-f69f-4701-ab14-eb9cf9c9450d · outbound

This paper cites How Does Critical Batch Size Scale in Pre-training?.

Is your batch size the problem? Revisiting the Adam-SGD gap in language modeling How Does Critical Batch Size Scale in Pre-training?

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-07T00:50:30.630441Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:50:30.630441Z digest=sha256:ba5a046878978aebbd4dcbca8594e3dbb5e638a8e9159fa1cde8ae4de4ac34fa

Observation f7dcaf01-f68f-4aad-a8d0-86ec4ad73dc9 · outbound

This paper cites Why Transformers Need Adam: A Hessian Perspective.

Is your batch size the problem? Revisiting the Adam-SGD gap in language modeling Why Transformers Need Adam: A Hessian Perspective

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-07T00:50:30.633781Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:50:30.633781Z digest=sha256:ca088a4ee09050b28696a50fa495db3621d0f1abb0e685f27bee8125cb1d32ac

Observation d6d36155-b1fa-4519-8f90-785c1d93fa03 · outbound

This paper cites Deconstructing What Makes a Good Optimizer for Language Models.

Is your batch size the problem? Revisiting the Adam-SGD gap in language modeling Deconstructing What Makes a Good Optimizer for Language Models

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-07T00:50:30.636520Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:50:30.636520Z digest=sha256:2d350675019d0b3cccecbc17e7ec15407e08f106bb447e3e99d26925ee6b5eca

Observation a976dd82-e9e3-4506-b999-0f76593a6436 · outbound

This paper cites an unresolved cited work.

Is your batch size the problem? Revisiting the Adam-SGD gap in language modeling Unresolved cited work

Reference 27

Resolution
unresolved
raw_fallback, observed 2026-08-07T00:50:31.065026Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T00:50:30.639985Z digest=sha256:948096aaece0daec274703994fceaef97d66e444cb9a307af7fde5fdc90b6bc6

Observation 08b0145f-91e0-48b2-ae3e-3a5593e047d9 · outbound

This paper cites [2024], and uses the codebase of Orvieto and Gower [2025].

Is your batch size the problem? Revisiting the Adam-SGD gap in language modeling [2024], and uses the codebase of Orvieto and Gower [2025]

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:50:31.039339Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T00:50:30.649181Z digest=sha256:90f32acf093876e0b6a62638a40c8332badfeccdd5e413797bb4af153f9675f6

Observation f34f1c0d-9218-4ca8-aec6-2a8131c3d5ec · outbound

This paper cites an unresolved cited work.

Is your batch size the problem? Revisiting the Adam-SGD gap in language modeling Unresolved cited work

Reference 1024

Resolution
unresolved
raw_fallback, observed 2026-08-07T00:50:31.056492Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T00:50:30.642968Z digest=sha256:295607e026f50189dcc4cdb4b8d65bc6e79190f4986c70164f70bd091825bd41

Observation f176b608-2526-4295-95c2-1538cb7575ee · outbound

This paper cites Transformers without Tears: Improving the Normalization of Self-Attention.

Is your batch size the problem? Revisiting the Adam-SGD gap in language modeling Transformers without Tears: Improving the Normalization of Self-Attention

Reference 1986

Resolution
unresolved
no resolver link, observed 2026-08-07T00:50:30.609473Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:50:30.609473Z digest=sha256:961deaac2e7da1969981429719fd5039d7cc13b0e8e3e0fd8cc475f9ff91e7e6

Observation 6799d912-962d-42ed-bda7-de130b44a4a0 · outbound

This paper cites How to Fine-Tune Vision Models with SGD.

Is your batch size the problem? Revisiting the Adam-SGD gap in language modeling How to Fine-Tune Vision Models with SGD

Reference 2014

Resolution
verified exact
local_arxiv, observed 2026-08-07T00:50:30.926338Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T00:50:30.576108Z digest=sha256:365553f811eba1eedbfdf75e14a8fdff70a35f47c0923d69a5869891dc0f27a4

Observation c597f601-3722-465a-851c-2e2ab7b14dcb · outbound

This paper cites The Llama 3 Herd of Models.

Is your batch size the problem? Revisiting the Adam-SGD gap in language modeling The Llama 3 Herd of Models

Reference 2017

Resolution
unresolved
no resolver link, observed 2026-08-07T00:50:30.548431Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:50:30.548431Z digest=sha256:88d00bdcc59039c96b3424e237012a0f606ab52508aeb85853ba2c0f64c3e34a

Observation c4d8fd51-e618-4b44-9367-f940273b414c · outbound

This paper cites signSGD: Compressed Optimisation for Non-Convex Problems.

Is your batch size the problem? Revisiting the Adam-SGD gap in language modeling signSGD: Compressed Optimisation for Non-Convex Problems

Reference 2018

Resolution
unresolved
no resolver link, observed 2026-08-07T00:50:29.980013Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:50:29.980013Z digest=sha256:44e1f9384fa791f310725115923dbf166fba845553cee3383eb3861fa2f31873

Observation ef191c64-b70d-4e05-a88f-27d993caa9c9 · outbound

This paper cites Practical Efficiency of Muon for Pretraining.

Is your batch size the problem? Revisiting the Adam-SGD gap in language modeling Practical Efficiency of Muon for Pretraining

Reference 2019

Resolution
unresolved
no resolver link, observed 2026-08-07T00:50:30.618825Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:50:30.618825Z digest=sha256:c5498a1e20e3240eec157720d556150089db6cdbd3c0ef8cf48147338eda96ce

Observation 4f04113b-7173-4087-a8fc-853e94d5a18b · outbound

This paper cites Accurate, Large Minibatch SGD: Training ImageNet in 1 Hour.

Is your batch size the problem? Revisiting the Adam-SGD gap in language modeling Accurate, Large Minibatch SGD: Training ImageNet in 1 Hour

Reference 2020

Resolution
unresolved
no resolver link, observed 2026-08-07T00:50:30.464304Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:50:30.464304Z digest=sha256:514555182a21c0ccda6672de8edd906b072368efe1df4d8af323d680fc7dac44

Observation 72b5ec18-fdd9-4bf8-b615-731d1aa0231f · outbound

This paper cites SGDR: Stochastic Gradient Descent with Warm Restarts.

Is your batch size the problem? Revisiting the Adam-SGD gap in language modeling SGDR: Stochastic Gradient Descent with Warm Restarts

Reference 2021

Resolution
unresolved
no resolver link, observed 2026-08-07T00:50:30.603923Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:50:30.603923Z digest=sha256:903795e9c017dfed83be26fb010d856bb40b083c23b0a98fbbd63bb82356a4a0

Observation a959df74-d31d-472b-b5f2-150c36d047b8 · outbound

This paper cites An Image is Worth 16x16 Words: Transformers for Image Recognition at Scale.

Is your batch size the problem? Revisiting the Adam-SGD gap in language modeling An Image is Worth 16x16 Words: Transformers for Image Recognition at Scale

Reference 2022

Resolution
unresolved
no resolver link, observed 2026-08-07T00:50:30.348876Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:50:30.348876Z digest=sha256:e39aa8f9198d7b70be9f1c4606688a5fad2a76d4417ca7c9f9678dc157cccede

Observation 6cec26bc-75cf-40f6-8c17-d216703fe868 · outbound

This paper cites GPT-NeoX-20B: An Open-Source Autoregressive Language Model.

Is your batch size the problem? Revisiting the Adam-SGD gap in language modeling GPT-NeoX-20B: An Open-Source Autoregressive Language Model

Reference 2023

Resolution
unresolved
no resolver link, observed 2026-08-07T00:50:30.096081Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:50:30.096081Z digest=sha256:836562ec093ef9de722036a33c415deb443cd96ea10a09c563e9dbcbdb9edd18

Observation 3c56be0f-8610-4504-8510-4e9111b6f761 · outbound

This paper cites Linear attention is (maybe) all you need (to understand transformer optimization).

Is your batch size the problem? Revisiting the Adam-SGD gap in language modeling Linear attention is (maybe) all you need (to understand transformer optimization)

Reference 2024

Resolution
unresolved
no resolver link, observed 2026-08-07T00:50:29.838100Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:50:29.838100Z digest=sha256:c52660246d58f76dc2a47bbfe60798fd88108a10848f08b74b7f11208547f623

Observation 0da4f519-d61b-47a1-8033-3430a73bbbff · outbound

This paper cites GLU Variants Improve Transformer.

Is your batch size the problem? Revisiting the Adam-SGD gap in language modeling GLU Variants Improve Transformer

Reference 2025

Resolution
unresolved
no resolver link, observed 2026-08-07T00:50:30.621589Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:50:30.621589Z digest=sha256:688d1eed323e03f6259b0b63c4c85baffc927caa1bc2ff81f155a08407fc15a3

Pith citing papers

Observation 34bb4621-ccf2-4792-bf5c-1e3ccfebbc71 · inbound

Seesaw: Accelerating Training by Balancing Learning Rate and Batch Size Scheduling cites this paper.

Seesaw: Accelerating Training by Balancing Learning Rate and Batch Size Scheduling Is your batch size the problem? Revisiting the Adam-SGD gap in language modeling

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-04T09:38:48.446865Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T09:38:48.446865Z digest=sha256:48924771408f83c9ce2263f3deabfb312157e18718e4a74deb4877d444e211af

Observation 0d414ffe-531d-4c19-a57a-892ac77a246f · inbound

How to Scale Mixture-of-Experts: From muP to the Maximally Scale-Stable Parameterization cites this paper.

How to Scale Mixture-of-Experts: From muP to the Maximally Scale-Stable Parameterization Is your batch size the problem? Revisiting the Adam-SGD gap in language modeling

Reference 41

Resolution
metadata mismatch
arxiv_id, observed 2026-05-15T04:49:44.656518Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-05-15T04:45:20.091598Z digest=sha256:749e70741d272fb680c6ce445b74f0a294bb966cef8c1b86dce8ba52ba57dd2b

Observation 01bd3cbc-d651-4a8c-bf5a-ad2e192dce3d · inbound

Revisiting the Adam-SGD Gap in LLM Pre-Training: The Role of Large Effective Learning Rates cites this paper.

Revisiting the Adam-SGD Gap in LLM Pre-Training: The Role of Large Effective Learning Rates Is your batch size the problem? Revisiting the Adam-SGD gap in language modeling

Reference 9

Resolution
verified exact
arxiv_id, observed 2026-05-20T13:38:19.427606Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-20T13:34:09.129379Z digest=sha256:b5a2c159b413df333680f4a1512598cf09a022a26b9ce0ee7b0b3e7ac46f55c7

Observation 53449721-9569-4766-a9a1-050f989768bf · inbound

Why SGD is not Brownian Motion: A New Perspective on Stochastic Dynamics cites this paper.

Why SGD is not Brownian Motion: A New Perspective on Stochastic Dynamics Is your batch size the problem? Revisiting the Adam-SGD gap in language modeling

Reference 235

Resolution
verified exact
arxiv_id, observed 2026-05-22T08:11:17.363127Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-05-22T08:06:52.309619Z digest=sha256:bd91ac5703ae018df99faf4532a13d7156aaf58d283a5260cb756d45d3b516ef

Observation 3aa9feae-5033-46d6-8a2e-f130184609b1 · inbound

A Theoretical and Experimental Study of a Novel Adaptive Learning Algorithm cites this paper.

A Theoretical and Experimental Study of a Novel Adaptive Learning Algorithm Is your batch size the problem? Revisiting the Adam-SGD gap in language modeling

Reference 11

Resolution
verified exact
arxiv_id, observed 2026-06-29T09:23:16.439125Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-06-29T09:18:31.526770Z digest=sha256:71443736b97ef7f97e626537bbde3377be86eed868783a44fbbf1f7b8b81c972

Observation e9e7170e-6db5-4851-b0e2-92c553ad4910 · inbound

Quantifying the Agreement Between Data-Influence and Data-Similarity to Understand LLM Behavior cites this paper.

Quantifying the Agreement Between Data-Influence and Data-Similarity to Understand LLM Behavior Is your batch size the problem? Revisiting the Adam-SGD gap in language modeling

Reference 48

Resolution
verified exact
arxiv_id, observed 2026-06-26T08:49:14.928649Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-06-26T08:45:34.884703Z digest=sha256:b19dbccdf8d615efd8d9e27eb9b3e18974652f3572cb209578d098d40f823e2b