Pith. sign in

Paper Citation Record · LEDGER

Optimizer-Model Consistency: Full Finetuning with the Same Optimizer as Pretraining Forgets Less

As of 20 August 2026, this Paper Citation Record lists 52 of 52 outbound references and 3 inbound Pith citation observations for arXiv:2605.06654.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2605.06654 v1

Coverage vector

measured 52 of 52 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-05-08T12:00:49.127471Z

measured 55 of 55 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-20T06:33:59.587034+00:00

measured 3 of 3 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-04T04:19:38.424519Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-07-02T11:56:55.992114Z

Reference resolution

52 of 52 outbound references displayed

  • verified exact41
  • verified fuzzy4
  • unresolved2
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch5

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation c2207734-b6f6-42fc-a872-15b9fd029f84 · outbound

This paper cites arXiv preprint arXiv:2512.16928 , year=.

Optimizer-Model Consistency: Full Finetuning with the Same Optimizer as Pretraining Forgets Less arXiv preprint arXiv:2512.16928 , year=

Reference 1

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T19:26:08.621624Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-08T12:00:49.127471Z digest=sha256:9bcf205fbca239a9b516192815b4bc70e49002f52bc9049ab898beec8dcfc361

Observation 1a26534a-2a18-4038-9238-482c3f75d6d6 · outbound

This paper cites The Geometry of Sign Gradient Descent.

Optimizer-Model Consistency: Full Finetuning with the Same Optimizer as Pretraining Forgets Less The Geometry of Sign Gradient Descent

Reference 2

Resolution
verified exact
arxiv_id, observed 2026-05-11T19:26:08.613799Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-08T12:00:49.127471Z digest=sha256:b03db7576482ff0f9731f21b2218f8fe622cff12c23be737cc49428b502f9fb0

Observation 7f285ec0-7c51-4989-94a2-5e67b5bff8f7 · outbound

This paper cites Old Optimizer, New Norm: An Anthology.

Optimizer-Model Consistency: Full Finetuning with the Same Optimizer as Pretraining Forgets Less Old Optimizer, New Norm: An Anthology

Reference 3

Resolution
verified exact
arxiv_id, observed 2026-05-16T07:27:53.041705Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-08T12:00:49.127471Z digest=sha256:2fe8af101e658717a3f1e785d13f037250b52b9eaddc047983f3a41a0e0a4150

Observation ad4e9f03-6c4e-478a-826b-f9b537b6122c · outbound

This paper cites LoRA Learns Less and Forgets Less.

Optimizer-Model Consistency: Full Finetuning with the Same Optimizer as Pretraining Forgets Less LoRA Learns Less and Forgets Less

Reference 4

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T19:26:08.668232Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-08T12:00:49.127471Z digest=sha256:d30fdcdf0f8dcc83e62b04a57dca951c90c1628fb2d3e085ab34deae4f68bb04

Observation 2c37b115-b48a-4f6b-83e9-cfa5e1166495 · outbound

This paper cites Why Gradients Rapidly Increase Near the End of Training.

Optimizer-Model Consistency: Full Finetuning with the Same Optimizer as Pretraining Forgets Less Why Gradients Rapidly Increase Near the End of Training

Reference 5

Resolution
verified exact
arxiv_id, observed 2026-05-11T19:26:08.465260Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-08T12:00:49.127471Z digest=sha256:ba9cc1481f5de0b93466dce6363cbd605441912141fae84ee271a8f2a6495f11

Observation 9aa17456-7f4c-4915-8828-f9adb80ba8cb · outbound

This paper cites The Llama 3 Herd of Models.

Optimizer-Model Consistency: Full Finetuning with the Same Optimizer as Pretraining Forgets Less The Llama 3 Herd of Models

Reference 6

Resolution
verified exact
local_arxiv, observed 2026-05-11T19:26:08.588850Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-08T12:00:49.127471Z digest=sha256:ff2aba715858702f6dcde5424cc1a1e27044c19e41a72f03ae8dac0dc559edcf

Observation 7d23d21a-41e1-4567-924e-045438fda472 · outbound

This paper cites Gaussian Error Linear Units (GELUs).

Optimizer-Model Consistency: Full Finetuning with the Same Optimizer as Pretraining Forgets Less Gaussian Error Linear Units (GELUs)

Reference 7

Resolution
verified exact
local_arxiv, observed 2026-05-11T19:26:08.454674Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-08T12:00:49.127471Z digest=sha256:9eb4cdfd4d154eddb5121de7299392e73c724048029bcb890e19fc29dbd3aadf

Observation 75507299-eb8a-4078-83fb-a22a249ceb3b · outbound

This paper cites Measuring Forgetting of Memorized Training Examples.

Optimizer-Model Consistency: Full Finetuning with the Same Optimizer as Pretraining Forgets Less Measuring Forgetting of Memorized Training Examples

Reference 8

Resolution
verified exact
arxiv_id, observed 2026-05-11T19:26:08.593021Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-08T12:00:49.127471Z digest=sha256:5559439ea3d77ed9fe486a44ad8f33d31c1a7be41975f8881315a27c62deaa5d

Observation 7ee54a71-d2e5-47f8-929b-96ac55cfdafa · outbound

This paper cites Provable Complexity Improvement of AdaGrad over SGD: Upper and Lower Bounds in Stochastic Non-Convex Optimization.

Optimizer-Model Consistency: Full Finetuning with the Same Optimizer as Pretraining Forgets Less Provable Complexity Improvement of AdaGrad over SGD: Upper and Lower Bounds in Stochastic Non-Convex Optimization

Reference 9

Resolution
verified exact
arxiv_id, observed 2026-05-11T19:26:08.606390Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-08T12:00:49.127471Z digest=sha256:19b84739faae6c24be510fe8139247bfe80e102182c3711ce8359d72313bff4e

Observation 08cf0d11-a3bd-41b6-be38-ff7898efc49b · outbound

This paper cites Kimi K2.5: Visual Agentic Intelligence.

Optimizer-Model Consistency: Full Finetuning with the Same Optimizer as Pretraining Forgets Less Kimi K2.5: Visual Agentic Intelligence

Reference 10

Resolution
verified exact
local_arxiv, observed 2026-05-11T19:26:08.580519Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-08T12:00:49.127471Z digest=sha256:629705f1d1145b43e006bc5aa6e48b1419f3e80633515c0fd4abb4a269cc50b5

Observation 36fc57b1-e1fc-445b-9dbc-effb24ae7b67 · outbound

This paper cites Adam: A Method for Stochastic Optimization.

Optimizer-Model Consistency: Full Finetuning with the Same Optimizer as Pretraining Forgets Less Adam: A Method for Stochastic Optimization

Reference 11

Resolution
verified exact
local_arxiv, observed 2026-05-11T19:26:08.584831Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-08T12:00:49.127471Z digest=sha256:b632a6b7e5f103c8fc02f6b0380cf76eeebb775fa041297fef83d7b93d901af7

Observation f777e85e-7854-414a-a13e-ad3cc895b8ce · outbound

This paper cites Noise Is Not the Main Factor Behind the Gap Between SGD and Adam on Transformers, but Sign Descent Might Be.

Optimizer-Model Consistency: Full Finetuning with the Same Optimizer as Pretraining Forgets Less Noise Is Not the Main Factor Behind the Gap Between SGD and Adam on Transformers, but Sign Descent Might Be

Reference 12

Resolution
verified exact
arxiv_id, observed 2026-05-11T19:26:08.565537Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-08T12:00:49.127471Z digest=sha256:1e2e6e24a264e94d1dc0f80a730e4e199ffccc234c9dbc62552d0a5908f3891a

Observation 4ac87229-543a-4618-82ed-a76d92920cdc · outbound

This paper cites Muon is Scalable for LLM Training.

Optimizer-Model Consistency: Full Finetuning with the Same Optimizer as Pretraining Forgets Less Muon is Scalable for LLM Training

Reference 13

Resolution
verified exact
arxiv_id, observed 2026-05-11T23:02:52.981675Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-08T12:00:49.127471Z digest=sha256:f748a55055b3050ba6a9bfecbc58b42ce48c616c36730545ee34a1f639a549b5

Observation 20218c8c-a913-4081-90e6-74a40eeb76b6 · outbound

This paper cites AdaGrad under Anisotropic Smoothness.

Optimizer-Model Consistency: Full Finetuning with the Same Optimizer as Pretraining Forgets Less AdaGrad under Anisotropic Smoothness

Reference 14

Resolution
verified exact
arxiv_id, observed 2026-05-11T19:26:08.577041Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-08T12:00:49.127471Z digest=sha256:ce70db0ee1da8afd3d0e660406b128e899addf4a52daea2e968f39a80306f58e

Observation 96edba34-090c-4e5c-b12d-8427fd30a166 · outbound

This paper cites SGDR: Stochastic Gradient Descent with Warm Restarts.

Optimizer-Model Consistency: Full Finetuning with the Same Optimizer as Pretraining Forgets Less SGDR: Stochastic Gradient Descent with Warm Restarts

Reference 15

Resolution
verified exact
local_arxiv, observed 2026-05-11T19:26:08.573133Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-08T12:00:49.127471Z digest=sha256:635fb508df5a8b58eb8969001193974fc15b75c62b92fceabfbea262dbcec759

Observation 4f098732-d3b8-4da8-b2ed-62c4bafda5c2 · outbound

This paper cites Decoupled Weight Decay Regularization.

Optimizer-Model Consistency: Full Finetuning with the Same Optimizer as Pretraining Forgets Less Decoupled Weight Decay Regularization

Reference 16

Resolution
verified exact
local_arxiv, observed 2026-05-11T19:26:08.600508Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-08T12:00:49.127471Z digest=sha256:9602de469edeab1818a326793741a206524f69319a9f9a89b6f635bd65edf4f4

Observation f4527207-05ea-4dd8-9507-941d1a833de2 · outbound

This paper cites Sculpting Subspaces: Constrained Full Fine-Tuning in LLMs for Continual Learning.

Optimizer-Model Consistency: Full Finetuning with the Same Optimizer as Pretraining Forgets Less Sculpting Subspaces: Constrained Full Fine-Tuning in LLMs for Continual Learning

Reference 17

Resolution
verified exact
arxiv_id, observed 2026-05-11T19:26:08.634575Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-08T12:00:49.127471Z digest=sha256:2f5e4c5a20594a75d5e41efa0f1279baad148954b43c24a52408506b98d3ccfb

Observation 1e85877e-44bc-4d15-a1a8-5749c9a97cc2 · outbound

This paper cites Unbiased gradient low-rank projection.arXiv preprint arXiv:2510.17802.

Optimizer-Model Consistency: Full Finetuning with the Same Optimizer as Pretraining Forgets Less Unbiased gradient low-rank projection.arXiv preprint arXiv:2510.17802

Reference 18

Resolution
verified exact
arxiv_id, observed 2026-05-11T19:26:08.558293Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-08T12:00:49.127471Z digest=sha256:b4f02c7eec032ce7b42c5122ebfb80df26b244357629c972da889ed38d451b52

Observation 8de20694-9b62-4668-9ea0-79d4fddf3e94 · outbound

This paper cites Training Deep Learning Models with Norm-Constrained LMOs.

Optimizer-Model Consistency: Full Finetuning with the Same Optimizer as Pretraining Forgets Less Training Deep Learning Models with Norm-Constrained LMOs

Reference 19

Resolution
verified exact
arxiv_id, observed 2026-05-21T21:22:37.653978Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-08T12:00:49.127471Z digest=sha256:d9e234f6ca874afea5accf30073cbc2bdbc1b34fce7fbeecb55a045bc3610c72

Observation f1f02a4f-4515-4026-a6d2-bd867ff84826 · outbound

This paper cites icarl: Incre- mental classifier and representation learning.

Optimizer-Model Consistency: Full Finetuning with the Same Optimizer as Pretraining Forgets Less icarl: Incre- mental classifier and representation learning

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-05-26T13:47:25.028824Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-08T12:00:49.127471Z digest=sha256:b983b739b18dbc59a313e6860109e11275b422dd100f0d187196a81985869316

Observation 7a950eec-14aa-4e47-ab6c-b8982a1f4b85 · outbound

This paper cites On the Convergence of Adam and Beyond.

Optimizer-Model Consistency: Full Finetuning with the Same Optimizer as Pretraining Forgets Less On the Convergence of Adam and Beyond

Reference 21

Resolution
verified exact
arxiv_id, observed 2026-05-11T19:26:08.480366Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-08T12:00:49.127471Z digest=sha256:f50ea13639284c9175a793a702d41956030db6df3fbf82ef9bc30a22ee9c3815

Observation e69b531a-65e8-493c-9425-606bceeda5fb · outbound

This paper cites (How) Learning Rates Regulate Catastrophic Overtraining.

Optimizer-Model Consistency: Full Finetuning with the Same Optimizer as Pretraining Forgets Less (How) Learning Rates Regulate Catastrophic Overtraining

Reference 22

Resolution
verified exact
local_arxiv, observed 2026-05-11T19:26:08.516618Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-08T12:00:49.127471Z digest=sha256:066fed1cbd0bc5ff70a6ee30d1d679227330f5b1e28aac0fcbbf44a34c199659

Observation c241f457-713a-41f8-9d3e-7a495da4d34b · outbound

This paper cites Benchmarking Optimizers for Large Language Model Pretraining.

Optimizer-Model Consistency: Full Finetuning with the Same Optimizer as Pretraining Forgets Less Benchmarking Optimizers for Large Language Model Pretraining

Reference 23

Resolution
verified exact
arxiv_id, observed 2026-05-11T19:26:08.543311Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-08T12:00:49.127471Z digest=sha256:220db84680636e73b16acea7c059e6f4ff05ae3345011007773aadd4ca130eab

Observation fd97bb7d-23c4-4df2-b005-67030e28bb84 · outbound

This paper cites GLU Variants Improve Transformer.

Optimizer-Model Consistency: Full Finetuning with the Same Optimizer as Pretraining Forgets Less GLU Variants Improve Transformer

Reference 24

Resolution
verified exact
local_arxiv, observed 2026-05-11T19:26:08.625847Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-08T12:00:49.127471Z digest=sha256:19ad7c13be709d61b8655b82f9f8c57daf1f59aad1499fee561c4e4112e6cba6

Observation ee1c7cdd-d492-4dec-a458-30197356ce85 · outbound

This paper cites Lora vs full fine-tuning: An illusion of equivalence.

Optimizer-Model Consistency: Full Finetuning with the Same Optimizer as Pretraining Forgets Less Lora vs full fine-tuning: An illusion of equivalence

Reference 25

Resolution
verified exact
arxiv_id, observed 2026-05-11T19:26:08.547227Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-08T12:00:49.127471Z digest=sha256:f5d4736805a0905a638b6f80f6c21c65501ef2b26e45efb8fd7670f478c9646d

Observation b97e7a15-a4d7-4248-9177-2ad6489cbcff · outbound

This paper cites Overtrained Language Models Are Harder to Fine-Tune.

Optimizer-Model Consistency: Full Finetuning with the Same Optimizer as Pretraining Forgets Less Overtrained Language Models Are Harder to Fine-Tune

Reference 26

Resolution
verified exact
arxiv_id, observed 2026-05-11T19:26:08.539517Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-08T12:00:49.127471Z digest=sha256:695d3c0e27cb04686356c6a313faea0e24b9d8446e33e9d01b776ccc2caeb040

Observation ab44dce8-7b8b-44cb-902e-4aa2bd1a0bc3 · outbound

This paper cites Less Regret via Online Conditioning.

Optimizer-Model Consistency: Full Finetuning with the Same Optimizer as Pretraining Forgets Less Less Regret via Online Conditioning

Reference 27

Resolution
verified exact
arxiv_id, observed 2026-05-11T19:26:08.508615Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-08T12:00:49.127471Z digest=sha256:f92495f134fa895cd07345a944d662ed1c40135188c286bbb5be1c967e709ce2

Observation 1e3575ad-ab31-4b6d-8858-e9664cfd7ddf · outbound

This paper cites ArXiv Preprint: 2511.00674 , Year =.

Optimizer-Model Consistency: Full Finetuning with the Same Optimizer as Pretraining Forgets Less ArXiv Preprint: 2511.00674 , Year =

Reference 28

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T19:26:08.610339Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-08T12:00:49.127471Z digest=sha256:8b13349bf5a225e265e7d187e3821c753f1e85c29d2ee1e5e2aa9e169adf53e6

Observation d8f36dc5-1a7a-4d95-a955-201b4894fdc2 · outbound

This paper cites Llama 2: Open Foundation and Fine-Tuned Chat Models.

Optimizer-Model Consistency: Full Finetuning with the Same Optimizer as Pretraining Forgets Less Llama 2: Open Foundation and Fine-Tuned Chat Models

Reference 29

Resolution
verified exact
local_arxiv, observed 2026-05-11T19:26:08.512185Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-08T12:00:49.127471Z digest=sha256:19057d2c4c70b41d3fc818d58556ce1d64026debb8466ccd7e52f368a4550b6d

Observation 29e05265-12eb-40cc-95fd-de75bd8b3692 · outbound

This paper cites SOAP: Improving and Stabilizing Shampoo using Adam.

Optimizer-Model Consistency: Full Finetuning with the Same Optimizer as Pretraining Forgets Less SOAP: Improving and Stabilizing Shampoo using Adam

Reference 30

Resolution
verified exact
arxiv_id, observed 2026-05-15T01:02:28.631723Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-08T12:00:49.127471Z digest=sha256:507b1b92f7cce0b7ec382273555e6816b88006170b3bab2a78008ebbdfd1a9a8

Observation 6a34c718-50c3-41dc-a453-89bed6c500bc · outbound

This paper cites Muon outperforms adam in tail-end associative memory learning.

Optimizer-Model Consistency: Full Finetuning with the Same Optimizer as Pretraining Forgets Less Muon outperforms adam in tail-end associative memory learning

Reference 31

Resolution
verified exact
arxiv_id, observed 2026-05-11T19:26:08.554861Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-08T12:00:49.127471Z digest=sha256:8c090d51634914ae3b7cc649d31e8b905fb2c754c4ba8a0e1c9324433d3342b3

Observation 8347914b-4c3a-4e8e-aefe-e7d261dbc576 · outbound

This paper cites Magicoder: Empowering Code Generation with OSS-Instruct.

Optimizer-Model Consistency: Full Finetuning with the Same Optimizer as Pretraining Forgets Less Magicoder: Empowering Code Generation with OSS-Instruct

Reference 32

Resolution
verified exact
arxiv_id, observed 2026-05-11T19:26:08.655977Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-08T12:00:49.127471Z digest=sha256:650800bdffd4a13933a303887fa6b21fd4b76e9e8f9aec26fb8c2b692d7f7178

Observation fbb0ff62-ebc5-4ca2-8e0d-a4c89742f877 · outbound

This paper cites Fantastic Pretraining Optimizers and Where to Find Them.

Optimizer-Model Consistency: Full Finetuning with the Same Optimizer as Pretraining Forgets Less Fantastic Pretraining Optimizers and Where to Find Them

Reference 33

Resolution
verified exact
arxiv_id, observed 2026-05-11T19:26:08.470520Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-08T12:00:49.127471Z digest=sha256:9b770f715d997ea2cad34f46153e88915c45f1fe4c456a89b5cb7337f95e247d

Observation 804a2766-31e4-49f3-92f5-fb3a064bf02b · outbound

This paper cites Structured Preconditioners in Adaptive Optimization: A Unified Analysis.

Optimizer-Model Consistency: Full Finetuning with the Same Optimizer as Pretraining Forgets Less Structured Preconditioners in Adaptive Optimization: A Unified Analysis

Reference 34

Resolution
verified exact
arxiv_id, observed 2026-05-11T19:26:08.660083Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-08T12:00:49.127471Z digest=sha256:8db6d018c1f6c2f8eccb033fc52c013b8f55a85f6631eaec986f491ccfc55c9c

Observation ccac7eb6-8554-4f1d-ba41-a2562d30ea4f · outbound

This paper cites Controlled llm training on spectral sphere.

Optimizer-Model Consistency: Full Finetuning with the Same Optimizer as Pretraining Forgets Less Controlled llm training on spectral sphere

Reference 35

Resolution
verified exact
arxiv_id, observed 2026-05-11T19:26:08.531367Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-08T12:00:49.127471Z digest=sha256:e925e7b7a23875bf688d3fbb6af0a4b0ab4c464c079d822820e26e2b97ce8e85

Observation c4701b1e-aca3-42f1-8db0-da32010a60c0 · outbound

This paper cites On the width scaling of neural optimizers under matrix operator norms i: Row/column normalization and hyperparameter transfer.arXiv preprint arXiv:2603.09952.

Optimizer-Model Consistency: Full Finetuning with the Same Optimizer as Pretraining Forgets Less On the width scaling of neural optimizers under matrix operator norms i: Row/column normalization and hyperparameter transfer.arXiv preprint arXiv:2603.09952

Reference 36

Resolution
verified exact
arxiv_id, observed 2026-05-11T19:26:08.475919Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-08T12:00:49.127471Z digest=sha256:fd1acd16dfa2ca8c8ddbbb4d3a4482c518ff8e5a5c412b09690959f7a0770aa8

Observation 300a9d68-6a8e-4151-aa35-e275248e36ae · outbound

This paper cites Qwen3 Technical Report.

Optimizer-Model Consistency: Full Finetuning with the Same Optimizer as Pretraining Forgets Less Qwen3 Technical Report

Reference 37

Resolution
verified exact
local_arxiv, observed 2026-05-11T19:26:08.629870Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-08T12:00:49.127471Z digest=sha256:9619617e31922eab00a067cc36d10c7764281e38b6a4bfd5fe0828cd7dfb0dff

Observation 982ace06-fbed-4f02-b303-9cc6e2c4cdfa · outbound

This paper cites A Spectral Condition for Feature Learning.

Optimizer-Model Consistency: Full Finetuning with the Same Optimizer as Pretraining Forgets Less A Spectral Condition for Feature Learning

Reference 38

Resolution
verified exact
arxiv_id, observed 2026-05-11T19:26:08.596926Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-08T12:00:49.127471Z digest=sha256:40e5c417ae6b1de0145e133c5fedc204232a53aa44787847eeab1d147f3a8235

Observation d01d0395-cad3-4f9b-8c96-e98959813614 · outbound

This paper cites Large Batch Optimization for Deep Learning: Training BERT in 76 minutes.

Optimizer-Model Consistency: Full Finetuning with the Same Optimizer as Pretraining Forgets Less Large Batch Optimization for Deep Learning: Training BERT in 76 minutes

Reference 39

Resolution
verified exact
arxiv_id, observed 2026-05-21T21:39:00.104164Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-08T12:00:49.127471Z digest=sha256:a336cd214387bc544d9f5b4b0233b6a4ae17567086de3312610b0fd1632b2e80

Observation 06f3d469-ea31-425a-b541-7c6f1cd48023 · outbound

This paper cites StoSignSGD: Unbiased Structural Stochasticity Fixes SignSGD for Training Large Language Models.

Optimizer-Model Consistency: Full Finetuning with the Same Optimizer as Pretraining Forgets Less StoSignSGD: Unbiased Structural Stochasticity Fixes SignSGD for Training Large Language Models

Reference 40

Resolution
verified exact
local_arxiv, observed 2026-05-11T19:26:08.562081Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-08T12:00:49.127471Z digest=sha256:49ee35571e5342f3ab350cb7aa7288153032b7340bd13403606da0369899a718

Observation f9e4fb2e-0950-4211-b854-5b50d1359a24 · outbound

This paper cites MetaMath: Bootstrap Your Own Mathematical Questions for Large Language Models.

Optimizer-Model Consistency: Full Finetuning with the Same Optimizer as Pretraining Forgets Less MetaMath: Bootstrap Your Own Mathematical Questions for Large Language Models

Reference 41

Resolution
verified exact
arxiv_id, observed 2026-05-13T10:07:53.977132Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-08T12:00:49.127471Z digest=sha256:e3783288140bcc9695b3147b386de57f2ee09ffa0b8c0c248c684db98d8a3097

Observation 538e25c1-2441-4600-8506-41ad000ad01b · outbound

This paper cites MARS: Unleashing the Power of Variance Reduction for Training Large Models.

Optimizer-Model Consistency: Full Finetuning with the Same Optimizer as Pretraining Forgets Less MARS: Unleashing the Power of Variance Reduction for Training Large Models

Reference 42

Resolution
verified exact
arxiv_id, observed 2026-05-11T19:26:08.551157Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-08T12:00:49.127471Z digest=sha256:bb0acb69050c1e626bd566aba0870393cf522cbe9aead69fc87287c3b4f866a8

Observation 7998c599-731f-4125-89ab-d025e16f3398 · outbound

This paper cites ADADELTA: An Adaptive Learning Rate Method.

Optimizer-Model Consistency: Full Finetuning with the Same Optimizer as Pretraining Forgets Less ADADELTA: An Adaptive Learning Rate Method

Reference 43

Resolution
verified exact
arxiv_id, observed 2026-05-11T19:26:08.535605Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-08T12:00:49.127471Z digest=sha256:79a2803889c70ddb6d52491f99683adb7b4574658edd81c9584324fef3c74d1a

Observation a21fa863-90da-41e5-9bfa-ccd8459b71d7 · outbound

This paper cites GLM-4.5: Agentic, Reasoning, and Coding (ARC) Foundation Models.

Optimizer-Model Consistency: Full Finetuning with the Same Optimizer as Pretraining Forgets Less GLM-4.5: Agentic, Reasoning, and Coding (ARC) Foundation Models

Reference 44

Resolution
verified exact
local_arxiv, observed 2026-05-11T19:26:08.642431Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-08T12:00:49.127471Z digest=sha256:7966fb06e2be9f4a4778b9657d9db62a0d0b7fd9d97f7bf432de993497f5da44

Observation 2b360052-3f06-406c-91f1-0c3317a26c5a · outbound

This paper cites Understanding deep learning requires rethinking generalization.

Optimizer-Model Consistency: Full Finetuning with the Same Optimizer as Pretraining Forgets Less Understanding deep learning requires rethinking generalization

Reference 45

Resolution
verified exact
arxiv_id, observed 2026-05-13T11:56:40.372790Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-08T12:00:49.127471Z digest=sha256:e0e51b87c0d5132cd40a63d18cf3657f39f490107ee2a18d9d768ff8eb405291

Observation b7dff74f-a846-48eb-8e76-8b915e9acffc · outbound

This paper cites Why Transformers Need Adam: A Hessian Perspective.

Optimizer-Model Consistency: Full Finetuning with the Same Optimizer as Pretraining Forgets Less Why Transformers Need Adam: A Hessian Perspective

Reference 46

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T19:26:08.489778Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-08T12:00:49.127471Z digest=sha256:14869d3b249372497c374a9c290d72250145ad76e66a73b70e70ab830aa7dc9f

Observation 9d9eb6cb-2630-4c13-a323-20ebf4406cca · outbound

This paper cites LlamaFactory: Unified Efficient Fine-Tuning of 100+ Language Models.

Optimizer-Model Consistency: Full Finetuning with the Same Optimizer as Pretraining Forgets Less LlamaFactory: Unified Efficient Fine-Tuning of 100+ Language Models

Reference 47

Resolution
metadata mismatch
local_arxiv, observed 2026-05-11T19:26:08.526331Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-08T12:00:49.127471Z digest=sha256:47f446de4bc2c87aac0da37a3bb950cfd509a0df575afecd9ef0e07697cc7ced

Observation 34dcb554-30af-42cc-b424-357df4da4d06 · outbound

This paper cites an unresolved cited work.

Optimizer-Model Consistency: Full Finetuning with the Same Optimizer as Pretraining Forgets Less Unresolved cited work

Reference 48

Resolution
unresolved
raw_fallback, observed 2026-05-26T13:47:25.031255Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-08T12:00:49.127471Z digest=sha256:79a87b7520268e8855a0b2d4d54525d37f4d6e0cc0260ead138b19114aa92a6c

Observation 54569520-ba6a-4626-91be-97757a02b5b6 · outbound

This paper cites Here, LoRA rank 64 shows a greater forgetting compared to rank 256, mainly because rank 256 diverges for lr=5e-4, and a smaller learning rate leads to less forgetting.

Optimizer-Model Consistency: Full Finetuning with the Same Optimizer as Pretraining Forgets Less Here, LoRA rank 64 shows a greater forgetting compared to rank 256, mainly because rank 256 diverges for lr=5e-4, and a smaller learning rate leads to less forgetting

Reference 49

Resolution
verified fuzzy
raw_fallback, observed 2026-05-26T13:47:25.025446Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-08T12:00:49.127471Z digest=sha256:3b9aa480ba9d26db963b9ed8ab6660a6731b99a8ea36db7981b1fdb7258327da

Observation 581d67f9-75c1-41ff-b6eb-d653bb622d9a · outbound

This paper cites an unresolved cited work.

Optimizer-Model Consistency: Full Finetuning with the Same Optimizer as Pretraining Forgets Less Unresolved cited work

Reference 50

Resolution
unresolved
raw_fallback, observed 2026-05-26T13:47:25.019542Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-08T12:00:49.127471Z digest=sha256:bd8db2dfc2d9dc940add9692c5e86bd23ab8c6a0e0928f02cd094763fcc225c5

Observation 2469ed4d-f60e-4a5b-8514-9601ba82e64b · outbound

This paper cites B.2 More Detailed Activation Plots In Figures 10 and 11, we present the average activation sparsity of detailed modules and specific ac- tivation splits.

Optimizer-Model Consistency: Full Finetuning with the Same Optimizer as Pretraining Forgets Less B.2 More Detailed Activation Plots In Figures 10 and 11, we present the average activation sparsity of detailed modules and specific ac- tivation splits

Reference 51

Resolution
verified fuzzy
raw_fallback, observed 2026-05-26T13:47:25.016966Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-08T12:00:49.127471Z digest=sha256:5834f740617008c7301b544ed054cdeb91918add958d2b48d78985b7588e1218

Observation 57ef95fa-30bb-4711-b3c2-176177269826 · outbound

This paper cites 23 •Ifα∈(α 1,∞], based on Lemma 2 and Assumption 4, it holds that ∥∆W∥ α1,β∗ ≤ ∥∆W∥ α,β∗ andE h ∥x∥2 α1 i = Θ E h ∥x∥2 α i.

Optimizer-Model Consistency: Full Finetuning with the Same Optimizer as Pretraining Forgets Less 23 •Ifα∈(α 1,∞], based on Lemma 2 and Assumption 4, it holds that ∥∆W∥ α1,β∗ ≤ ∥∆W∥ α,β∗ andE h ∥x∥2 α1 i = Θ E h ∥x∥2 α i

Reference 52

Resolution
verified fuzzy
raw_fallback, observed 2026-05-26T13:47:25.022703Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-08T12:00:49.127471Z digest=sha256:762394158ff15d9ca2738310c469c12920f17448b4dab78a69bd0d1acab20de7

Pith citing papers

Observation 0dcd802d-c6f2-4675-a64f-0c12f2d4c537 · inbound

Double Preconditioning (DoPr): Optimization for Test-Time Performance, not Validation Loss cites this paper.

Double Preconditioning (DoPr): Optimization for Test-Time Performance, not Validation Loss Optimizer-Model Consistency: Full Finetuning with the Same Optimizer as Pretraining Forgets Less

Reference 185

Resolution
metadata mismatch
local_arxiv, observed 2026-07-02T11:56:55.993275Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-06-28T02:35:39.845487Z digest=sha256:31cc0a6b866637fc67f2db5d0056af299545111be75cfbf3c8e0bdb2d0201482

Observation 259292d8-88cf-4435-a378-1f4e34a0f944 · inbound

When Does Muon Help Agentic Reinforcement Learning? cites this paper.

When Does Muon Help Agentic Reinforcement Learning? Optimizer-Model Consistency: Full Finetuning with the Same Optimizer as Pretraining Forgets Less

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-01T21:13:39.201865Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T21:13:39.201865Z digest=sha256:0e4fb3ab34d9405cc90d82f52e75810d9938759499a16641ca0138ccbe324834

Observation 31f5b1af-31b0-4aac-9ee3-6833d9bf061e · inbound

When Does Muon Help Agentic Reinforcement Learning? cites this paper.

When Does Muon Help Agentic Reinforcement Learning? Optimizer-Model Consistency: Full Finetuning with the Same Optimizer as Pretraining Forgets Less

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-04T04:19:38.424519Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T04:19:38.424519Z digest=sha256:4bc7893cb6c2ef6da7ed1e12cb3dd5d608fdbbb16fce9e0a5515b984fd8d4af7