Pith. sign in

Paper Citation Record · LEDGER

Memory-Efficient 4-bit Preconditioned Stochastic Optimization

As of 13 August 2026, this Paper Citation Record lists 67 of 67 outbound references and 0 inbound Pith citation observations for arXiv:2412.10663.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2412.10663 v2

Coverage vector

measured 67 of 67 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-11T15:50:24.396837Z

measured 67 of 67 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-13T06:32:02.005865+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

67 of 67 outbound references displayed

  • verified exact0
  • verified fuzzy37
  • unresolved29
  • parse uncertain0
  • malformed identifier1
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 3baabe3b-ea9d-4129-8036-9a93820a5ae4 · outbound

This paper cites Disentangling Adaptive Gradient Methods from Learning Rates.

Memory-Efficient 4-bit Preconditioned Stochastic Optimization Disentangling Adaptive Gradient Methods from Learning Rates

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-11T15:50:24.038813Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T15:50:24.038813Z digest=sha256:1510af93af1072595634d5e6478c374a34df3e9d3ceb05c2301e985afc8e647f

Observation d1f9b074-f827-40ff-a6c8-66c8a94d6cc3 · outbound

This paper cites Qsgd: Communication-efficient sgd via gra- dient quantization and encoding.

Memory-Efficient 4-bit Preconditioned Stochastic Optimization Qsgd: Communication-efficient sgd via gra- dient quantization and encoding

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T15:50:25.681430Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-11T15:50:24.044650Z digest=sha256:0de73fa689550192acd48ce1d63998e4f4a8d8d72f9f2eeeb0394a0521c716f2

Observation ef0786a3-b958-43c2-87fa-15b60fbfcef7 · outbound

This paper cites Scalable Second Order Optimization for Deep Learning.

Memory-Efficient 4-bit Preconditioned Stochastic Optimization Scalable Second Order Optimization for Deep Learning

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-11T15:50:24.049744Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T15:50:24.049744Z digest=sha256:2ce3c6df00c2d0cf8c93a8bc6d881e8341e126b7d4821bd7fb2f73558bd03ec0

Observation 54b4f2df-400d-4138-a2fa-08bb085dd44d · outbound

This paper cites Stochas- tic approximations and differential inclusions.

Memory-Efficient 4-bit Preconditioned Stochastic Optimization Stochas- tic approximations and differential inclusions

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T15:50:25.660354Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-11T15:50:24.055011Z digest=sha256:6464d9a6528db392c3da40a0c36ecf43dcb9311a6e0386ce270b071ea3c2e3c4

Observation 2f77c16d-6a36-455e-a837-ce2365108578 · outbound

This paper cites Conservative set valued fields, automatic differentiation, stochastic gradient methods and deep learning.

Memory-Efficient 4-bit Preconditioned Stochastic Optimization Conservative set valued fields, automatic differentiation, stochastic gradient methods and deep learning

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T15:50:25.643518Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-11T15:50:24.059985Z digest=sha256:226d2cee3bae1b8d580aeb05546f5569331a3e1026088cff7e268d3742d30783

Observation 5e71a2e1-7644-48ae-ac6a-f08c8ff71905 · outbound

This paper cites Stochastic approximation: a dynamical sys- tems viewpoint.

Memory-Efficient 4-bit Preconditioned Stochastic Optimization Stochastic approximation: a dynamical sys- tems viewpoint

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T15:50:25.618021Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-11T15:50:24.065354Z digest=sha256:6709a952b61e954df3a3905499d05df918bee6caf81a2527563cf52ac9f0341c

Observation a2e85daa-e140-4eaa-afc7-71be80c143eb · outbound

This paper cites Language Models are Few-Shot Learners.

Memory-Efficient 4-bit Preconditioned Stochastic Optimization Language Models are Few-Shot Learners

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-11T15:50:24.070170Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T15:50:24.070170Z digest=sha256:f738744c107c1f93f8e1423e058bc9da714544c59852cdefd99f9f46f297bb8d

Observation 0c3beb52-263d-4181-97d1-8e84ea942ce5 · outbound

This paper cites Lower bounds for finding stationary points ii: first- order methods.

Memory-Efficient 4-bit Preconditioned Stochastic Optimization Lower bounds for finding stationary points ii: first- order methods

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T15:50:25.596438Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-11T15:50:24.075803Z digest=sha256:40bfbfce3ed567bd12ec3a3487a4b122a4a4025037f22715cd9b5af757e16aee

Observation 219db309-1fd4-4e00-86d8-e886a1485330 · outbound

This paper cites Optimization and nonsmooth analysis.

Memory-Efficient 4-bit Preconditioned Stochastic Optimization Optimization and nonsmooth analysis

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T15:50:25.571868Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-11T15:50:24.080337Z digest=sha256:bf29ee7429042aaf2c8c84ccb3faea2ec5d0c20b50c8804b441266a243370c41

Observation 044afc28-ef1d-4aa3-a969-8aab1f5ce6d8 · outbound

This paper cites AutoAugment: Learning Augmentation Policies from Data.

Memory-Efficient 4-bit Preconditioned Stochastic Optimization AutoAugment: Learning Augmentation Policies from Data

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-11T15:50:24.085329Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T15:50:24.085329Z digest=sha256:0642bb69b7990393c5e539dcb29ed32d1abd2f8d66bf6a3ed3be810c5a090045

Observation 4c7b912a-605b-4cc0-9da7-d29d372c7f5b · outbound

This paper cites Randaugment: Practical automated data augmen- tation with a reduced search space.

Memory-Efficient 4-bit Preconditioned Stochastic Optimization Randaugment: Practical automated data augmen- tation with a reduced search space

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T15:50:25.551165Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-11T15:50:24.091242Z digest=sha256:7686ad2f8a84f76b05e59d92324d7eba944760eab7e332339ea13a6a807a47ac

Observation 6d53f9a6-1289-4727-bab1-ffa30862597d · outbound

This paper cites Pathological sub- gradient dynamics.

Memory-Efficient 4-bit Preconditioned Stochastic Optimization Pathological sub- gradient dynamics

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T15:50:25.525584Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-11T15:50:24.095983Z digest=sha256:00b514a5f5c16ade91a13dbc66e0ddd0a8cd42a5ca8b407f0c77869e7e678f6b

Observation 8103e450-7622-49ed-a284-260752fa330a · outbound

This paper cites Stochastic subgradient method converges on tame functions.

Memory-Efficient 4-bit Preconditioned Stochastic Optimization Stochastic subgradient method converges on tame functions

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T15:50:25.502286Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-11T15:50:24.102733Z digest=sha256:ad336d4e49dc46dbbc46979aa0d3b929d844cb568258750ef88e7321afb6b315

Observation 74e2a6d3-c45b-44a7-9f0a-62fc56fe814f · outbound

This paper cites Imagenet: A large-scale hierarchical image database.

Memory-Efficient 4-bit Preconditioned Stochastic Optimization Imagenet: A large-scale hierarchical image database

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-11T15:50:24.110579Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T15:50:24.110579Z digest=sha256:ca45abe3fd4e9b12ec1e2234495d39faaaa523934480f611a362c92683ce1736

Observation c3a1db67-6fec-43e1-bcb1-1f138d582f8d · outbound

This paper cites 8-bit Optimizers via Block-wise Quantization.

Memory-Efficient 4-bit Preconditioned Stochastic Optimization 8-bit Optimizers via Block-wise Quantization

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-11T15:50:24.116174Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T15:50:24.116174Z digest=sha256:2b766c5e832824b2d906eaf8a6f9d4782d686cced499eb40432d9cc7e0d9859c

Observation cb132225-4516-47d1-9830-d914a8a3fa9c · outbound

This paper cites An image is worth 16x16 words: Trans- formers for image recognition at scale.

Memory-Efficient 4-bit Preconditioned Stochastic Optimization An image is worth 16x16 words: Trans- formers for image recognition at scale

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-11T15:50:24.123963Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T15:50:24.123963Z digest=sha256:31ebc0a40e697636015d23a5922b8b51be29c7b33df77565392b880419cc658e

Observation f36936a6-d005-4d6f-8c47-47a1a693c816 · outbound

This paper cites Adaptive sub- gradient methods for online learning and stochastic opti- mization.

Memory-Efficient 4-bit Preconditioned Stochastic Optimization Adaptive sub- gradient methods for online learning and stochastic opti- mization

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T15:50:25.456438Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-11T15:50:24.131062Z digest=sha256:44a25f104ad72c432b6d76e93102fe54a7854c035223f9e784560bfeb1011c3d

Observation 83b72dc2-c9a3-4805-a166-c81408a46899 · outbound

This paper cites Stochastic methods for com- posite and weakly convex optimization problems.

Memory-Efficient 4-bit Preconditioned Stochastic Optimization Stochastic methods for com- posite and weakly convex optimization problems

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T15:50:25.437387Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-11T15:50:24.137144Z digest=sha256:028866d71d320948b4bf8e48dde05201a324fa4fff12be45b6d30a0f8f672b0e

Observation 83881c3b-dbff-440b-b03b-4a1640ce0314 · outbound

This paper cites A survey of quan- tization methods for efficient neural network inference.

Memory-Efficient 4-bit Preconditioned Stochastic Optimization A survey of quan- tization methods for efficient neural network inference

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T15:50:25.415832Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-11T15:50:24.142495Z digest=sha256:1ac8232295414dc7c504562c1830003be3df2335f76dd5d2be178f4451db7b5c

Observation a85be752-776d-4b3b-9302-cf9ca20aa6e7 · outbound

This paper cites Practi- cal quasi-newton methods for training deep neural networks.

Memory-Efficient 4-bit Preconditioned Stochastic Optimization Practi- cal quasi-newton methods for training deep neural networks

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T15:50:25.396650Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-11T15:50:24.147386Z digest=sha256:9041272e1c4b759962fb5c05b6ab4968b2004ac29c411bf34ac6c7d7a5d1f0a5

Observation 4f6e5407-8681-490f-b325-faeedf0713f0 · outbound

This paper cites A schur–newton method for the matrixp th root and its inverse.SIAM Journal on Matrix Analysis and Applications , 28(3):788–804, 2006.

Memory-Efficient 4-bit Preconditioned Stochastic Optimization A schur–newton method for the matrixp th root and its inverse.SIAM Journal on Matrix Analysis and Applications , 28(3):788–804, 2006

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T15:50:25.378818Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-11T15:50:24.152394Z digest=sha256:6cf995fe79dfc87f9373abe3ffe534aeaaf818fbe3c7010637c58e02c4afd370

Observation 6e4e80f8-2677-4a8e-9401-40c5f8321e56 · outbound

This paper cites Shampoo: Preconditioned stochastic tensor optimization.

Memory-Efficient 4-bit Preconditioned Stochastic Optimization Shampoo: Preconditioned stochastic tensor optimization

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T15:50:25.362689Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-11T15:50:24.157728Z digest=sha256:c4dec9d71508de8d9ccb4fcbf93bcb2201d10513518df3dba6ed63980f3df5c0

Observation 5240ece3-6e13-42db-a26f-e792f39bf890 · outbound

This paper cites Deep residual learning for image recognition.

Memory-Efficient 4-bit Preconditioned Stochastic Optimization Deep residual learning for image recognition

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T15:50:25.345719Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-11T15:50:24.162986Z digest=sha256:fd0e1f256633b238d0e1485e74c6452fa738b8312a220a05c3ad25afa05c2976

Observation 0e2e08b5-2303-43f5-aa24-323fde4246d3 · outbound

This paper cites Training Compute-Optimal Large Language Models.

Memory-Efficient 4-bit Preconditioned Stochastic Optimization Training Compute-Optimal Large Language Models

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-11T15:50:24.168423Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T15:50:24.168423Z digest=sha256:ce698e89b3d5b3d3f8d8ce117410bf6ccc183e48dc11cbf0e163f373b2f815de

Observation d15b6993-2368-44b3-aed2-a014844042e7 · outbound

This paper cites Scaling Laws for Neural Language Models.

Memory-Efficient 4-bit Preconditioned Stochastic Optimization Scaling Laws for Neural Language Models

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-11T15:50:24.174002Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T15:50:24.174002Z digest=sha256:a078089907f02554d3dd7091c69a2a1086d61a5ad6199fe4f1203486acb6f4d8

Observation 284003b4-e4c2-43ec-a958-e14673caa2df · outbound

This paper cites Adam: A Method for Stochastic Optimization.

Memory-Efficient 4-bit Preconditioned Stochastic Optimization Adam: A Method for Stochastic Optimization

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-11T15:50:24.179731Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T15:50:24.179731Z digest=sha256:5d795bdee25d2628662eb7244abe8f4a1b63158d85038cb1111d1d5e1462c51b

Observation 508fc494-7374-4f12-8b47-ed211b1650cc · outbound

This paper cites Learning multiple layers of features from tiny images.

Memory-Efficient 4-bit Preconditioned Stochastic Optimization Learning multiple layers of features from tiny images

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-11T15:50:24.185289Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T15:50:24.185289Z digest=sha256:4fbbdcb3a8b05f0331b0807a9a01a98856e4b826ebfa91d0ce30a20e588ee88e

Observation 4abdb69c-71f3-47c5-b718-afe0724d1c5e · outbound

This paper cites Imagenet classification with deep convolutional neural net- works.

Memory-Efficient 4-bit Preconditioned Stochastic Optimization Imagenet classification with deep convolutional neural net- works

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-11T15:50:24.190866Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T15:50:24.190866Z digest=sha256:71eec19cc0db47032c5ea6f6a4eb48000b53449f941b247b8593bb6d50e93a60

Observation c4002d35-b1a8-414e-9d1b-0928c4dbfaa7 · outbound

This paper cites Tiny imagenet visual recognition challenge.

Memory-Efficient 4-bit Preconditioned Stochastic Optimization Tiny imagenet visual recognition challenge

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-11T15:50:24.196055Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T15:50:24.196055Z digest=sha256:85b774b009c41990af1f5e0a8646963fd7a12f5d17de58f9bef92ca2f2fd0651

Observation 3e485a9b-809e-4cde-8b23-87d1fe47ac09 · outbound

This paper cites Biobert: a pre-trained biomedical language representation model for biomedical text mining.

Memory-Efficient 4-bit Preconditioned Stochastic Optimization Biobert: a pre-trained biomedical language representation model for biomedical text mining

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T15:50:25.284735Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-11T15:50:24.200658Z digest=sha256:4fe0a47ff6f9305c2739e502dd2b29f270fc999a250a00c50c5aed5082b10337

Observation 6c4c0482-d15c-4101-8629-360cf7d432ea · outbound

This paper cites Vision Transformer for Small-Size Datasets.

Memory-Efficient 4-bit Preconditioned Stochastic Optimization Vision Transformer for Small-Size Datasets

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-11T15:50:24.205713Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T15:50:24.205713Z digest=sha256:ff2ab4ced0b002551978f73448633f4e60f641fcbf7c571247588ac39d4243f0

Observation 71d89b49-da1a-47fb-9fe4-ca8648433408 · outbound

This paper cites Memory efficient optimizers with 4-bit states.

Memory-Efficient 4-bit Preconditioned Stochastic Optimization Memory efficient optimizers with 4-bit states

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T15:50:25.268482Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-11T15:50:24.210780Z digest=sha256:1eee9aad9fd3a5fc734e13108d90e72db766f5fe6f99fdf9efcbcf589c8f0de0

Observation bd284cb7-e48f-44f4-91c2-705c9cac06ff · outbound

This paper cites ReLoRA: High-Rank Training Through Low-Rank Updates.

Memory-Efficient 4-bit Preconditioned Stochastic Optimization ReLoRA: High-Rank Training Through Low-Rank Updates

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-11T15:50:24.215832Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T15:50:24.215832Z digest=sha256:72d43b65de560ee2a904699163c3231145aa60a87780ee23095ab287e9346373

Observation 1229df2f-866f-4300-8857-8c7fff14e9b3 · outbound

This paper cites Swin transformer: Hierarchical vision transformer using shifted windows.

Memory-Efficient 4-bit Preconditioned Stochastic Optimization Swin transformer: Hierarchical vision transformer using shifted windows

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T15:50:25.251290Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-11T15:50:24.221264Z digest=sha256:4e6d0a98b996ae74baa6b1b777677b86ead1087930faaabe591c774459e17f9f

Observation d549d709-24d2-44c8-9b15-3d8a98ff84fc · outbound

This paper cites Decoupled weight de- cay regularization.

Memory-Efficient 4-bit Preconditioned Stochastic Optimization Decoupled weight de- cay regularization

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T15:50:25.235801Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-11T15:50:24.227137Z digest=sha256:a3a9dc953214648681660feb53ec5736a19cc6d24d6216d348f139948b53beca

Observation 530cdb4b-f617-4bc3-a572-5ddbbaa0ceb3 · outbound

This paper cites Optimizing neural net- works with kronecker-factored approximate curvature.

Memory-Efficient 4-bit Preconditioned Stochastic Optimization Optimizing neural net- works with kronecker-factored approximate curvature

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T15:50:25.217693Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-11T15:50:24.232474Z digest=sha256:e2e30e78506f8120ce01fe165f790f387d20ed27ee65d03aadc47d6f6f5fd7a3

Observation c6ddb572-8caa-4109-8cd5-b596d7cece6a · outbound

This paper cites A New Perspective on Shampoo's Preconditioner.

Memory-Efficient 4-bit Preconditioned Stochastic Optimization A New Perspective on Shampoo's Preconditioner

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-11T15:50:24.237399Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T15:50:24.237399Z digest=sha256:95e57da871440d1d6f8875ed9b21a53a8086088e45e5ce9344b4acf5bae2fb00

Observation 491079d2-2698-4a27-9360-c1acb16f9854 · outbound

This paper cites Learning transferable visual models from natural language supervi- sion.

Memory-Efficient 4-bit Preconditioned Stochastic Optimization Learning transferable visual models from natural language supervi- sion

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-11T15:50:24.243153Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T15:50:24.243153Z digest=sha256:db6baa568ccf557c13b071ae5b919e18925d7befcfaa19682e956e7c95952960

Observation 8a9e99e1-4006-4404-9ccb-b7571a4def38 · outbound

This paper cites Exploring the limits of transfer learning with a unified text-to-text transformer.

Memory-Efficient 4-bit Preconditioned Stochastic Optimization Exploring the limits of transfer learning with a unified text-to-text transformer

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-11T15:50:24.248004Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T15:50:24.248004Z digest=sha256:0741949521304421efe801f7dab963366bb79d71dbc194026e04445464536bf4

Observation a77dfca3-3598-4f5e-927d-6fbfa004e624 · outbound

This paper cites Ef21: A new, simpler, theoretically better, and practically faster error feedback.

Memory-Efficient 4-bit Preconditioned Stochastic Optimization Ef21: A new, simpler, theoretically better, and practically faster error feedback

Reference 40

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T15:50:25.171369Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-11T15:50:24.252774Z digest=sha256:67ca81a24558766d9c4794689191e5a8a9438c994cb0f6ec579b998624365eb1

Observation 82b856b9-e98d-4d56-afe7-03f7d4924448 · outbound

This paper cites A stochastic approxima- tion method.

Memory-Efficient 4-bit Preconditioned Stochastic Optimization A stochastic approxima- tion method

Reference 41

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T15:50:25.151379Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-11T15:50:24.257589Z digest=sha256:bbc77aa4d8767e06f8b60941ca28fb4b5c3fb74fc2a46616dd20ae81fa93ac1f

Observation 9c3e612e-aed1-405f-99e7-9ff27088862d · outbound

This paper cites 1-bit stochastic gradient descent and its application to data- parallel distributed training of speech dnns.

Memory-Efficient 4-bit Preconditioned Stochastic Optimization 1-bit stochastic gradient descent and its application to data- parallel distributed training of speech dnns

Reference 42

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T15:50:25.134140Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-11T15:50:24.262743Z digest=sha256:c1e6b173813a010f4d779f2158cac458c4f73c13c45b9cd73ffb171f7448f00a

Observation 4da20aa6-8ab7-4c1e-8e21-8c47e36fb61a · outbound

This paper cites A Distributed Data-Parallel PyTorch Implementation of the Distributed Shampoo Optimizer for Training Neural Networks At-Scale.

Memory-Efficient 4-bit Preconditioned Stochastic Optimization A Distributed Data-Parallel PyTorch Implementation of the Distributed Shampoo Optimizer for Training Neural Networks At-Scale

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-11T15:50:24.267978Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T15:50:24.267978Z digest=sha256:682595f1dc51372a72f1dd0974ec805b2edcab54005d606cfe30e52160b763ac

Observation ff676b04-675c-4a94-8bdb-b603b6613c71 · outbound

This paper cites Very Deep Convolutional Networks for Large-Scale Image Recognition.

Memory-Efficient 4-bit Preconditioned Stochastic Optimization Very Deep Convolutional Networks for Large-Scale Image Recognition

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-11T15:50:24.272997Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T15:50:24.272997Z digest=sha256:25e8de6b845a4ffae0cd67afe4d49d96273f361878dcd7d163504a385ff29f3f

Observation 94ae2857-c901-4213-8b4a-5a185fbe7f38 · outbound

This paper cites On the importance of initialization and momentum in deep learning.

Memory-Efficient 4-bit Preconditioned Stochastic Optimization On the importance of initialization and momentum in deep learning

Reference 45

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T15:50:25.113071Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-11T15:50:24.277637Z digest=sha256:df0cadcb09e133952bc89f10be741b17cb4b17a9272554fc1ffec1347e144208

Observation 11f7b997-8f6e-4413-8fd2-6457c0d2fdb8 · outbound

This paper cites 1-bit adam: Communication efficient large- scale training with adam’s convergence speed.

Memory-Efficient 4-bit Preconditioned Stochastic Optimization 1-bit adam: Communication efficient large- scale training with adam’s convergence speed

Reference 46

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T15:50:25.092633Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-11T15:50:24.282445Z digest=sha256:bc773eec951bdd6176c3f2d20c061d0f9e644d750b603e9cf81a4e8e21018b17

Observation d4ea303b-9cd2-427d-8ae7-3aa68cf20ed3 · outbound

This paper cites Lecture 6.5- rmsprop: Divide the gradient by a running average of its re- cent magnitude.

Memory-Efficient 4-bit Preconditioned Stochastic Optimization Lecture 6.5- rmsprop: Divide the gradient by a running average of its re- cent magnitude

Reference 47

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T15:50:25.070539Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-11T15:50:24.287286Z digest=sha256:ccd5ccfafdf953689b482d0ec0f19351601a72e2d264f589560585b85a30a2e1

Observation 83951d8a-9aa4-41c5-b9a4-294e56bef14b · outbound

This paper cites Llama 2: Open Foundation and Fine-Tuned Chat Models.

Memory-Efficient 4-bit Preconditioned Stochastic Optimization Llama 2: Open Foundation and Fine-Tuned Chat Models

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-11T15:50:24.293143Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T15:50:24.293143Z digest=sha256:e4a16e56af1891346da2a349ee532dd9bf91fec6e26c5b6f000f9411402534c1

Observation 9a9287f3-272b-40e3-a161-a985cca76413 · outbound

This paper cites Powersgd: Practical low-rank gradient compression for dis- tributed optimization.

Memory-Efficient 4-bit Preconditioned Stochastic Optimization Powersgd: Practical low-rank gradient compression for dis- tributed optimization

Reference 49

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T15:50:25.045328Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-11T15:50:24.298708Z digest=sha256:631889a19e00b915a98211d15e449fd38ba8ca532fcbc13d81b1bf02cb08c357

Observation bd7348a7-9db4-4a20-bbd2-d939d6120adf · outbound

This paper cites SOAP: Improving and Stabilizing Shampoo using Adam.

Memory-Efficient 4-bit Preconditioned Stochastic Optimization SOAP: Improving and Stabilizing Shampoo using Adam

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-11T15:50:24.303608Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T15:50:24.303608Z digest=sha256:9ba417f43ed16e84576312e7c91cfa422f81f81db586a650c0da126c0139bd1b

Observation 2d1e864b-319e-4c3e-a327-d32a25ef9dbe · outbound

This paper cites 4-bit Shampoo for Memory-Efficient Network Training.

Memory-Efficient 4-bit Preconditioned Stochastic Optimization 4-bit Shampoo for Memory-Efficient Network Training

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-11T15:50:24.308787Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T15:50:24.308787Z digest=sha256:4de350fceb795f19c09d3ba05faa21540d7d600f8cef26921d75ffdfb1744f64

Observation 2c9988b1-a79c-4bd4-b47f-251750f56fd6 · outbound

This paper cites Terngrad: Ternary gradients to reduce communication in distributed deep learning.

Memory-Efficient 4-bit Preconditioned Stochastic Optimization Terngrad: Ternary gradients to reduce communication in distributed deep learning

Reference 52

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T15:50:25.025285Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-11T15:50:24.313567Z digest=sha256:d050de165260020f4193c9aab081338f6080f2ffddccee9576f0ac99039f3646

Observation 0d779a2b-6e52-42c8-98de-36510fe533fe · outbound

This paper cites ResNet strikes back: An improved training procedure in timm.

Memory-Efficient 4-bit Preconditioned Stochastic Optimization ResNet strikes back: An improved training procedure in timm

Reference 53

Resolution
unresolved
no resolver link, observed 2026-08-11T15:50:24.318482Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T15:50:24.318482Z digest=sha256:a27a128576117c53758569ece2d0b518ad6f79450f5b2e82e1646fb4712c9236

Observation 837dac85-0c35-4d77-8720-c24c9d328772 · outbound

This paper cites BloombergGPT: A Large Language Model for Finance.

Memory-Efficient 4-bit Preconditioned Stochastic Optimization BloombergGPT: A Large Language Model for Finance

Reference 54

Resolution
unresolved
no resolver link, observed 2026-08-11T15:50:24.323992Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T15:50:24.323992Z digest=sha256:5505cdb59aa876cd4f9521f0a08207d9f199f6dcbf244e9dc3069686f6612743

Observation 9f2669cf-3312-4c83-8134-f4730221d530 · outbound

This paper cites Crystal graph convolu- tional neural networks for an accurate and interpretable pre- diction of material properties.

Memory-Efficient 4-bit Preconditioned Stochastic Optimization Crystal graph convolu- tional neural networks for an accurate and interpretable pre- diction of material properties

Reference 55

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T15:50:25.007259Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-11T15:50:24.330016Z digest=sha256:66795f805151cacab2f43503fe53eb2d552a55d33edb90d78ea9ba1728921332

Observation b532440f-c87c-43b1-9ed8-218200bdd7f0 · outbound

This paper cites LoCo: Low-Bit Communication Adaptor for Large-scale Model Training.

Memory-Efficient 4-bit Preconditioned Stochastic Optimization LoCo: Low-Bit Communication Adaptor for Large-scale Model Training

Reference 56

Resolution
unresolved
no resolver link, observed 2026-08-11T15:50:24.335113Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T15:50:24.335113Z digest=sha256:c13d90dd567b20ea3265c643cc175a6826e0ce97c21ba9bfb40f67ab4edf402a

Observation 0514f0bb-b1c6-4db2-a352-774fcb21dfc7 · outbound

This paper cites Adan: Adaptive nesterov momentum algo- rithm for faster optimizing deep models.

Memory-Efficient 4-bit Preconditioned Stochastic Optimization Adan: Adaptive nesterov momentum algo- rithm for faster optimizing deep models

Reference 57

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T15:50:24.990483Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-11T15:50:24.341648Z digest=sha256:875b03cefe54f0e8bc905c09adf9212da9570b28b4514c7ec188c267b2d067d0

Observation 6c659bcd-1664-45b7-a297-0384a93ec693 · outbound

This paper cites Zeroquant: Ef- ficient and affordable post-training quantization for large- scale transformers.

Memory-Efficient 4-bit Preconditioned Stochastic Optimization Zeroquant: Ef- ficient and affordable post-training quantization for large- scale transformers

Reference 58

Resolution
unresolved
no resolver link, observed 2026-08-11T15:50:24.347114Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T15:50:24.347114Z digest=sha256:ca9df71892a942171ea2b86f6ddb31969006577b8b8372d1eca09c5583fc1a0c

Observation e45e018e-1712-48c0-b86c-642f11f8746d · outbound

This paper cites A general regret bound of preconditioned gradient method for dnn training.

Memory-Efficient 4-bit Preconditioned Stochastic Optimization A general regret bound of preconditioned gradient method for dnn training

Reference 59

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T15:50:24.960505Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-11T15:50:24.352290Z digest=sha256:8a446a410a2aff83b43d01fdf4c8351ddb750f1e87cc559cb5bf1023f97d10a3

Observation 103ed47c-d57b-4d5c-a7ec-531d6af374ad · outbound

This paper cites Cutmix: Regu- larization strategy to train strong classifiers with localizable features.

Memory-Efficient 4-bit Preconditioned Stochastic Optimization Cutmix: Regu- larization strategy to train strong classifiers with localizable features

Reference 60

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T15:50:24.943136Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-11T15:50:24.358057Z digest=sha256:97dc98a4411d30c195713ad6e9834ec91e55141160a682ebc2e750544d58c5c9

Observation 167cede4-e469-48ae-89f0-4b21c3efc5be · outbound

This paper cites mixup: Beyond empirical risk minimiza- tion.

Memory-Efficient 4-bit Preconditioned Stochastic Optimization mixup: Beyond empirical risk minimiza- tion

Reference 61

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T15:50:24.926154Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-11T15:50:24.363316Z digest=sha256:07dd3d681249e91964efeddc4276a62fde2403d6d2f3569c56c2d5ec61c48cdf

Observation f5e70d96-6696-41f9-abb7-21b535705d11 · outbound

This paper cites Why are adaptive methods good for attention mod- els? Advances in Neural Information Processing Systems , 33:15383–15393, 2020.

Memory-Efficient 4-bit Preconditioned Stochastic Optimization Why are adaptive methods good for attention mod- els? Advances in Neural Information Processing Systems , 33:15383–15393, 2020

Reference 62

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T15:50:24.908110Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-11T15:50:24.368363Z digest=sha256:18808c103d2bde2ff01e691ea925ee77d61752521620b9d0e73dcba90d11d11c

Observation 9948481d-ec79-41c1-98c9-912879ad5dea · outbound

This paper cites GaLore: Memory-Efficient LLM Training by Gradient Low-Rank Projection.

Memory-Efficient 4-bit Preconditioned Stochastic Optimization GaLore: Memory-Efficient LLM Training by Gradient Low-Rank Projection

Reference 63

Resolution
unresolved
no resolver link, observed 2026-08-11T15:50:24.373990Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T15:50:24.373990Z digest=sha256:e08631ddb7ff055ef4550d3a9298a9599194ff99dc0059c07963289d90fc96cf

Observation 3f7081ec-f753-40ea-8b75-b1bdf44ee7bb · outbound

This paper cites Random erasing data augmentation.

Memory-Efficient 4-bit Preconditioned Stochastic Optimization Random erasing data augmentation

Reference 64

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T15:50:24.890889Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-11T15:50:24.379527Z digest=sha256:d051c4f74be4bdd38897610e7b64ff938ed68ea9baf006cdf96f969a2bebd2a5

Observation e543c5ac-48f1-47ae-ba25-6244b883ca8a · outbound

This paper cites Definition B.2.

Memory-Efficient 4-bit Preconditioned Stochastic Optimization Definition B.2

Reference 65

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T15:50:24.871998Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-11T15:50:24.385998Z digest=sha256:9e1261dd7c517975373751985d984abafe395338d3a81918d8b4e9d029994816

Observation 2c018be2-19f1-4e67-9b6c-2a4afb0248bc · outbound

This paper cites an unresolved cited work.

Memory-Efficient 4-bit Preconditioned Stochastic Optimization Unresolved cited work

Reference 66

Resolution
unresolved
raw_fallback, observed 2026-08-11T15:50:24.853790Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-11T15:50:24.391447Z digest=sha256:fcc27aed3199605f0e416152bc03fad91e56de821dbbfa42f63c4d0828fe2831

Observation cacbd0aa-ca71-4ff6-9f9e-6952fac1849c · outbound

This paper cites Here NMi is the normal space of Mi.

Memory-Efficient 4-bit Preconditioned Stochastic Optimization Here NMi is the normal space of Mi

Reference 67

Resolution
malformed identifier
raw_fallback, observed 2026-08-11T15:50:24.834907Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-11T15:50:24.396837Z digest=sha256:15c7dffecb85e1a2555d0e1e1cf22ce7da89e2a53d4247ee445803aacaf3b9f6

Pith citing papers

No inbound Pith citation observations are available.