Pith. sign in

Paper Citation Record · LEDGER

Taming LLMs by Scaling Learning Rates with Gradient Grouping

As of 9 August 2026, this Paper Citation Record lists 92 of 92 outbound references and 2 inbound Pith citation observations for arXiv:2506.01049.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2506.01049 v1

Coverage vector

measured 92 of 92 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-07T11:57:29.278658Z

measured 94 of 94 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-09T06:31:02.800959+00:00

measured 2 of 2 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-07-11T22:10:49.683444Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-05-11T16:41:06.415113Z

Reference resolution

92 of 92 outbound references displayed

  • verified exact1
  • verified fuzzy4
  • unresolved87
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 0963deb6-c74b-4b8c-80b6-ca00a44c6488 · outbound

This paper cites online" 'onlinestring :=.

Taming LLMs by Scaling Learning Rates with Gradient Grouping online" 'onlinestring :=

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-07T11:57:28.960786Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T11:57:28.960786Z digest=sha256:76e642ffdd973fb997147a9a0a0246d86bb89ddb843d67a3cb5a848adefcba69

Observation 551f9183-6011-4774-a690-29566546a9c1 · outbound

This paper cites write newline.

Taming LLMs by Scaling Learning Rates with Gradient Grouping write newline

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-07T11:57:28.965573Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T11:57:28.965573Z digest=sha256:e7c63447966578966789082b48b0666783629bd6b68d72e3bc9e20760a0ff279

Observation b2f78582-7e3d-46ff-ad52-36f938fb7e84 · outbound

This paper cites Qwen-VL: A Versatile Vision-Language Model for Understanding, Localization, Text Reading, and Beyond.

Taming LLMs by Scaling Learning Rates with Gradient Grouping Qwen-VL: A Versatile Vision-Language Model for Understanding, Localization, Text Reading, and Beyond

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-07T11:57:28.969116Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T11:57:28.969116Z digest=sha256:6245df8f678e5315c2a5944cc76e05409d4734a3e581070803ec9315c007a7fd

Observation 0746c1e7-aa97-4e50-b269-2fbc6807cd5e · outbound

This paper cites LoRA-XS: Low-Rank Adaptation with Extremely Small Number of Parameters.

Taming LLMs by Scaling Learning Rates with Gradient Grouping LoRA-XS: Low-Rank Adaptation with Extremely Small Number of Parameters

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-07T11:57:28.972955Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T11:57:28.972955Z digest=sha256:74a42776c14862dcea6ffc50e47356309fabf736a0c32647d1ebfc4fbc4f3bf8

Observation f31865e9-46bc-4f4e-83af-6ca0a2eb761f · outbound

This paper cites an unresolved cited work.

Taming LLMs by Scaling Learning Rates with Gradient Grouping Unresolved cited work

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-07T11:57:28.976631Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T11:57:28.976631Z digest=sha256:a2e4db99f11e586a99fa857780f51b6cb8d5e7c922d21dd57a73568f8f4ae923

Observation ff8d07a4-c54e-4f38-97eb-367bb43c0c7f · outbound

This paper cites an unresolved cited work.

Taming LLMs by Scaling Learning Rates with Gradient Grouping Unresolved cited work

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-07T11:57:28.980071Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T11:57:28.980071Z digest=sha256:292bc20b68a6cb0d9535164941ff55c4bd6a7274c17d1aba2da8054cc6ba16ec

Observation 52fac5ca-1489-47af-8830-ce9d1cc63cc6 · outbound

This paper cites an unresolved cited work.

Taming LLMs by Scaling Learning Rates with Gradient Grouping Unresolved cited work

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-07T11:57:28.983339Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T11:57:28.983339Z digest=sha256:2746d88f0931474b484055885505918ad5c9f61717d2d5b17ae32cea88067868

Observation acd9c9c6-d16c-4306-a6d0-9b5881e8ea8c · outbound

This paper cites LLaVA-KD: A Framework of Distilling Multimodal Large Language Models.

Taming LLMs by Scaling Learning Rates with Gradient Grouping LLaVA-KD: A Framework of Distilling Multimodal Large Language Models

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-07T11:57:28.986563Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T11:57:28.986563Z digest=sha256:245441eac81045de87d9bd6b034039f91598af9b0452ac489a0a0bd3a6c5bbd0

Observation f2b3a661-29ae-442f-9be4-8049a765a487 · outbound

This paper cites an unresolved cited work.

Taming LLMs by Scaling Learning Rates with Gradient Grouping Unresolved cited work

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-07T11:57:28.991360Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T11:57:28.991360Z digest=sha256:9ab0321685adc6589188be904c601f22d70e17a47958778b777dd61c2baadea7

Observation ad7b1bb8-15bc-46ae-9929-e2f44bc0ca30 · outbound

This paper cites an unresolved cited work.

Taming LLMs by Scaling Learning Rates with Gradient Grouping Unresolved cited work

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-07T11:57:28.995330Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T11:57:28.995330Z digest=sha256:8510b00bff90162d34edc1151c6d6f0aad44165bc971d5ee205735a29cdc1a3c

Observation 3cd80351-bc8c-47b5-abf1-f6c9500b5ca5 · outbound

This paper cites an unresolved cited work.

Taming LLMs by Scaling Learning Rates with Gradient Grouping Unresolved cited work

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-07T11:57:28.998658Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T11:57:28.998658Z digest=sha256:07a12bb730ce47a81b6d618ab9765182fbde44272253a34711e19ff91946c5e2

Observation e68aeed3-1427-4dc0-bee3-9e33ee9beed7 · outbound

This paper cites an unresolved cited work.

Taming LLMs by Scaling Learning Rates with Gradient Grouping Unresolved cited work

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-07T11:57:29.002055Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T11:57:29.002055Z digest=sha256:c3825e2f6c81203816612dff7a458d6937e0afb7d9607a68be5fb0090a5524c4

Observation 692be308-d7b0-470f-9cb0-d27506ec7c40 · outbound

This paper cites BoolQ: Exploring the Surprising Difficulty of Natural Yes/No Questions.

Taming LLMs by Scaling Learning Rates with Gradient Grouping BoolQ: Exploring the Surprising Difficulty of Natural Yes/No Questions

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-07T11:57:29.005938Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T11:57:29.005938Z digest=sha256:4e6b7cdbfe186d4b3ee1adcf24024ae5c6d94dac92097617a0c2c6c373c69f13

Observation e73f06e3-5284-43c7-a3a7-e651d2fa2cef · outbound

This paper cites Think you have Solved Question Answering? Try ARC, the AI2 Reasoning Challenge.

Taming LLMs by Scaling Learning Rates with Gradient Grouping Think you have Solved Question Answering? Try ARC, the AI2 Reasoning Challenge

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-07T11:57:29.010001Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T11:57:29.010001Z digest=sha256:3f62d4d7e9405e811b1ea6e5ec2ee3b5cade027cb81032d2242251c9da4d2429

Observation 4038c6b4-e5b0-41c2-acfa-ba023fc04a99 · outbound

This paper cites InstructBLIP: Towards General-purpose Vision-Language Models with Instruction Tuning.

Taming LLMs by Scaling Learning Rates with Gradient Grouping InstructBLIP: Towards General-purpose Vision-Language Models with Instruction Tuning

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-07T11:57:29.013438Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T11:57:29.013438Z digest=sha256:d9cf57e2f415549f73ee085623c1374b9087e4ba43f49c70bc8942d0b18b7680

Observation 9bc5028c-58e4-4e6e-b142-8b2c4c3e1e08 · outbound

This paper cites an unresolved cited work.

Taming LLMs by Scaling Learning Rates with Gradient Grouping Unresolved cited work

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-07T11:57:29.016941Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T11:57:29.016941Z digest=sha256:de239a634c2c98aa5e28da0a53c52275a3a13b104be306a4791bedd6a7d3671d

Observation 3d09eed9-f3f1-4f1b-b336-868922556f8a · outbound

This paper cites an unresolved cited work.

Taming LLMs by Scaling Learning Rates with Gradient Grouping Unresolved cited work

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-07T11:57:29.020514Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T11:57:29.020514Z digest=sha256:30f26d8a5765ba55ca8da365fdc97f04a2ba13419ea3ad9c339af3e2d1a84eea

Observation 2b3ff932-08e5-4c17-91a2-f03696222cde · outbound

This paper cites an unresolved cited work.

Taming LLMs by Scaling Learning Rates with Gradient Grouping Unresolved cited work

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-07T11:57:29.024107Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T11:57:29.024107Z digest=sha256:e6a0944002b929e94883077ee0aafed243f9e8b5f7f1310fb4169ec96dae81b5

Observation 82a835d2-ab35-4e35-8390-b26a6bc7eb62 · outbound

This paper cites an unresolved cited work.

Taming LLMs by Scaling Learning Rates with Gradient Grouping Unresolved cited work

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-07T11:57:29.027322Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T11:57:29.027322Z digest=sha256:29816dc2c998891376d057b9535e8b643a1df7c102d110c7ed39aa70cadbd032

Observation 971fbb92-8263-4497-9930-f1671856488c · outbound

This paper cites an unresolved cited work.

Taming LLMs by Scaling Learning Rates with Gradient Grouping Unresolved cited work

Reference 20

Resolution
unresolved
raw_fallback, observed 2026-08-07T11:57:30.364643Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-07T11:57:29.030485Z digest=sha256:24550f074e5025e10195d432b1b254bdf4e47f7cc5377027b6ddadbb92138f0e

Observation c48c3115-6596-43e0-b15e-8798df9e7736 · outbound

This paper cites an unresolved cited work.

Taming LLMs by Scaling Learning Rates with Gradient Grouping Unresolved cited work

Reference 21

Resolution
unresolved
raw_fallback, observed 2026-08-07T11:57:30.352712Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-07T11:57:29.033741Z digest=sha256:34a67626a8370db8663f8f30e2c441d6364c476f66bf2f3521d02662b18f5dab

Observation fd64ccbf-f79b-4857-82be-e64792012b92 · outbound

This paper cites LoRA+: Efficient Low Rank Adaptation of Large Models.

Taming LLMs by Scaling Learning Rates with Gradient Grouping LoRA+: Efficient Low Rank Adaptation of Large Models

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-07T11:57:29.037629Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T11:57:29.037629Z digest=sha256:576394c008a5eadcb4ea785a7c34edbf3c7d0ea4f8a3ae6633efce997e07c8d3

Observation 7292eecc-2457-4f45-a096-18a097d6434f · outbound

This paper cites an unresolved cited work.

Taming LLMs by Scaling Learning Rates with Gradient Grouping Unresolved cited work

Reference 23

Resolution
unresolved
raw_fallback, observed 2026-08-07T11:57:30.339368Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-07T11:57:29.040979Z digest=sha256:6a7e5153bf826de436882292452f39aec1a1a2aca795d52ed87d3c6592264105

Observation 165ad793-8fe0-459a-8b08-99fa783084ed · outbound

This paper cites an unresolved cited work.

Taming LLMs by Scaling Learning Rates with Gradient Grouping Unresolved cited work

Reference 24

Resolution
unresolved
raw_fallback, observed 2026-08-07T11:57:30.328100Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-07T11:57:29.044311Z digest=sha256:219e878894f207ea42727ef96bb812892ddc2f1ee8ddfafb305131cbf005ae74

Observation d79f9994-2fba-40fe-89ff-79a511e26d96 · outbound

This paper cites LoRA: Low-Rank Adaptation of Large Language Models.

Taming LLMs by Scaling Learning Rates with Gradient Grouping LoRA: Low-Rank Adaptation of Large Language Models

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-07T11:57:29.047501Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T11:57:29.047501Z digest=sha256:dec443d2e4c0d1f4ae2fc5ca368aab7d8f745825b88cf7f2d90a54392e4d831e

Observation 05cda9ed-9b95-4933-ad63-6ddbaff80b92 · outbound

This paper cites LLM-Adapters: An Adapter Family for Parameter-Efficient Fine-Tuning of Large Language Models.

Taming LLMs by Scaling Learning Rates with Gradient Grouping LLM-Adapters: An Adapter Family for Parameter-Efficient Fine-Tuning of Large Language Models

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-07T11:57:29.051249Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T11:57:29.051249Z digest=sha256:28174c70f694acfd65d1fb60f87127c30e8237446acedede606a95fa1e90cb9c

Observation 9dd4f4fe-d38b-496a-ae49-bc58b853c8da · outbound

This paper cites SPAM: Spike-Aware Adam with Momentum Reset for Stable LLM Training.

Taming LLMs by Scaling Learning Rates with Gradient Grouping SPAM: Spike-Aware Adam with Momentum Reset for Stable LLM Training

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-07T11:57:29.054680Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T11:57:29.054680Z digest=sha256:24ca61263964103192e0a26169355b32765a8d498f087d24c07f0ec5ae814ee9

Observation 331539f7-7b6a-463e-93b9-846c81e04baf · outbound

This paper cites an unresolved cited work.

Taming LLMs by Scaling Learning Rates with Gradient Grouping Unresolved cited work

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-07T11:57:29.058208Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T11:57:29.058208Z digest=sha256:c78fc26d048b848ef0c100dda4f4e9793207c5f982994033cab356769b783839

Observation cbab1350-ed83-40e4-9c6d-b1b7bb62ffb6 · outbound

This paper cites Exploring Low Rank Training of Deep Neural Networks.

Taming LLMs by Scaling Learning Rates with Gradient Grouping Exploring Low Rank Training of Deep Neural Networks

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-07T11:57:29.061428Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T11:57:29.061428Z digest=sha256:00b68c4add307ecf0dac4937c50e6ea7b13bb43b2f36d63152819325ec34b2af

Observation 0f00894e-3412-4e42-903f-4742cd1e9907 · outbound

This paper cites an unresolved cited work.

Taming LLMs by Scaling Learning Rates with Gradient Grouping Unresolved cited work

Reference 30

Resolution
unresolved
raw_fallback, observed 2026-08-07T11:57:30.307740Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-07T11:57:29.064748Z digest=sha256:4906716e52f9760e884a960d4a47ad176570e7fec8cb42bcb515adf2ce18050c

Observation c83b9a68-badd-459f-8bb2-6f2024d8bbd8 · outbound

This paper cites Kingma and Jimmy Ba.

Taming LLMs by Scaling Learning Rates with Gradient Grouping Kingma and Jimmy Ba

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:57:30.296503Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-07T11:57:29.067938Z digest=sha256:ac26dc5883571cfd86d24c8288f5eb94403edb736bea273daadc9dae1f9d5a1e

Observation ae32cc3f-32fa-495d-8410-ca4a74769a87 · outbound

This paper cites o pf, Yannic Kilcher, Dimitri Von R \.

Taming LLMs by Scaling Learning Rates with Gradient Grouping o pf, Yannic Kilcher, Dimitri Von R \

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:57:30.285335Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-07T11:57:29.071052Z digest=sha256:19e5279fc8a41c0caccbe9f0a40ccf2892609d431a66607d06c3fe115ee33d74

Observation 05593c1e-0953-4188-becd-72abe6222c2d · outbound

This paper cites SEED-Bench: Benchmarking Multimodal LLMs with Generative Comprehension.

Taming LLMs by Scaling Learning Rates with Gradient Grouping SEED-Bench: Benchmarking Multimodal LLMs with Generative Comprehension

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-07T11:57:29.074356Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T11:57:29.074356Z digest=sha256:031c02bf1ae307d04d6862cdc80ff4fd03e513a0aa3a06589900fbc2c858c6a0

Observation 109676e1-f73c-4df9-a5d3-778c96f7cd5f · outbound

This paper cites LoRAP: Transformer Sub-Layers Deserve Differentiated Structured Compression for Large Language Models.

Taming LLMs by Scaling Learning Rates with Gradient Grouping LoRAP: Transformer Sub-Layers Deserve Differentiated Structured Compression for Large Language Models

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-07T11:57:29.078239Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T11:57:29.078239Z digest=sha256:51b7b321435d12eefb16c4eea86f516216530f9e30f1364449509ac341791a2d

Observation 10c826ac-215f-4559-8d2c-002b3bdd6f5a · outbound

This paper cites an unresolved cited work.

Taming LLMs by Scaling Learning Rates with Gradient Grouping Unresolved cited work

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-07T11:57:29.081480Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T11:57:29.081480Z digest=sha256:fd62eb51ba4b3003e99a9c2432a6821f8ec670f3724bc1cfcc1719253eefee89

Observation 84c01ecb-d986-44d9-846c-1c85806a28e3 · outbound

This paper cites Surge Phenomenon in Optimal Learning Rate and Batch Size Scaling.

Taming LLMs by Scaling Learning Rates with Gradient Grouping Surge Phenomenon in Optimal Learning Rate and Batch Size Scaling

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-07T11:57:29.084798Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T11:57:29.084798Z digest=sha256:9d9967f8a836cb0fef4a100e7ec9d517c70956ca846334e73b8e1d155be1661b

Observation 88980bb8-475c-4efc-8aa0-b0dffb53b9f5 · outbound

This paper cites Unveiling the Backbone-Optimizer Coupling Bias in Visual Representation Learning.

Taming LLMs by Scaling Learning Rates with Gradient Grouping Unveiling the Backbone-Optimizer Coupling Bias in Visual Representation Learning

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-07T11:57:29.088289Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T11:57:29.088289Z digest=sha256:f10fdc81eb788023d5fa3b3007441a682a7189199cd20febb5d0b056e88fe13a

Observation 364bbec4-e582-4037-b2b8-ce6cc5bc78cc · outbound

This paper cites an unresolved cited work.

Taming LLMs by Scaling Learning Rates with Gradient Grouping Unresolved cited work

Reference 38

Resolution
unresolved
raw_fallback, observed 2026-08-07T11:57:30.264637Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-07T11:57:29.091805Z digest=sha256:2d3a46c34a496c986ef2371485207d9658f051ee6279127b3ecf8dd0e8a71f76

Observation 7cdfcf04-d675-4a18-854c-468ef8462bf5 · outbound

This paper cites an unresolved cited work.

Taming LLMs by Scaling Learning Rates with Gradient Grouping Unresolved cited work

Reference 39

Resolution
unresolved
raw_fallback, observed 2026-08-07T11:57:30.252122Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-07T11:57:29.095039Z digest=sha256:deefdd15778c749a2b0e14d5a957591837e2ced24d7cfa2b6b40d41a1d7c35b4

Observation dc732852-44aa-4a48-9b37-8a32aae9c42d · outbound

This paper cites an unresolved cited work.

Taming LLMs by Scaling Learning Rates with Gradient Grouping Unresolved cited work

Reference 40

Resolution
unresolved
raw_fallback, observed 2026-08-07T11:57:30.239884Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-07T11:57:29.098283Z digest=sha256:3b24aff27d7361a4d15ab530ef657e0f8f13e5a5e0cf977b6e7b5bed8342c8fd

Observation f23a6cc7-1999-48cb-a9f3-eea0fd23118a · outbound

This paper cites an unresolved cited work.

Taming LLMs by Scaling Learning Rates with Gradient Grouping Unresolved cited work

Reference 41

Resolution
unresolved
raw_fallback, observed 2026-08-07T11:57:30.228139Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-07T11:57:29.101419Z digest=sha256:f8b773d6ab1f0156727f0e64271f696c7e31fc153fbcc7b3493af3cddf427576

Observation 58996f95-d9c1-4333-a044-f7b8b12fb5d0 · outbound

This paper cites an unresolved cited work.

Taming LLMs by Scaling Learning Rates with Gradient Grouping Unresolved cited work

Reference 42

Resolution
unresolved
raw_fallback, observed 2026-08-07T11:57:30.216422Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-07T11:57:29.104867Z digest=sha256:6c1ef2969b8960e467fce203be819d49a3007e5876b6fa35dd65c19be23a4716

Observation d965385b-4881-4497-a283-25194e44450c · outbound

This paper cites MoE-LLaVA: Mixture of Experts for Large Vision-Language Models.

Taming LLMs by Scaling Learning Rates with Gradient Grouping MoE-LLaVA: Mixture of Experts for Large Vision-Language Models

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-07T11:57:29.108339Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T11:57:29.108339Z digest=sha256:5179e4cf643e6c8363612ae76b764f01d1bd9972c671aee214f9751aa2a9a8a6

Observation f77f1f80-dc91-4ecc-b75a-d070ee67d29a · outbound

This paper cites an unresolved cited work.

Taming LLMs by Scaling Learning Rates with Gradient Grouping Unresolved cited work

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-07T11:57:29.111921Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T11:57:29.111921Z digest=sha256:62f34dfa0f79b31512b99e690b589eeb12ba14609f1f4c3cfbcaf313365fbae5

Observation f249484d-dc3b-4482-b5f3-a6ae44cf3c29 · outbound

This paper cites an unresolved cited work.

Taming LLMs by Scaling Learning Rates with Gradient Grouping Unresolved cited work

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-07T11:57:29.115370Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T11:57:29.115370Z digest=sha256:7ffe7bdcd201ccdad4a584a3270c0ec1be555e5beba29028ae7d34b7379ce3c9

Observation 01d72d3a-a244-4dcc-9acc-ac26c3dc937d · outbound

This paper cites Muon is Scalable for LLM Training.

Taming LLMs by Scaling Learning Rates with Gradient Grouping Muon is Scalable for LLM Training

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-07T11:57:29.118553Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T11:57:29.118553Z digest=sha256:27985919a71422afc5bad0792fb6c5166fdb0359891953fd2bc6a120c8230cd7

Observation c88cdf3a-b3a5-47fd-9c7a-756791a6b415 · outbound

This paper cites an unresolved cited work.

Taming LLMs by Scaling Learning Rates with Gradient Grouping Unresolved cited work

Reference 47

Resolution
unresolved
raw_fallback, observed 2026-08-07T11:57:30.188577Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-07T11:57:29.122075Z digest=sha256:1abc5b67ae27018311fa715ab5b128f43c236e0ce297e2a16c39b702cdddbcc1

Observation bd877286-b983-49ef-a3a1-6eeba05ac596 · outbound

This paper cites an unresolved cited work.

Taming LLMs by Scaling Learning Rates with Gradient Grouping Unresolved cited work

Reference 48

Resolution
unresolved
raw_fallback, observed 2026-08-07T11:57:30.176770Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-07T11:57:29.125346Z digest=sha256:d92b56c0bcc1aa86a09c5ea166c249c85b70a4532c164691cca0d17d162175b8

Observation 60d1ffe9-8272-4615-84eb-d64e913ea568 · outbound

This paper cites DoRA: Weight-Decomposed Low-Rank Adaptation.

Taming LLMs by Scaling Learning Rates with Gradient Grouping DoRA: Weight-Decomposed Low-Rank Adaptation

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-07T11:57:29.129514Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T11:57:29.129514Z digest=sha256:c955a34d0e7e56e1a067607858c2c8c1d5c57973f049247aa95e4aa54ad725af

Observation cc653f9c-4569-4eb6-86cf-d8074085a207 · outbound

This paper cites an unresolved cited work.

Taming LLMs by Scaling Learning Rates with Gradient Grouping Unresolved cited work

Reference 50

Resolution
unresolved
raw_fallback, observed 2026-08-07T11:57:30.164748Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-07T11:57:29.132995Z digest=sha256:b0625f5afe7c384dac6083861e23c2c133929ec78e1757b14889ca4eecb0fa7c

Observation e9786330-16e2-47b7-8b01-43fb2e0337be · outbound

This paper cites an unresolved cited work.

Taming LLMs by Scaling Learning Rates with Gradient Grouping Unresolved cited work

Reference 51

Resolution
unresolved
raw_fallback, observed 2026-08-07T11:57:30.153328Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-07T11:57:29.136203Z digest=sha256:1f2b4227ad1c753d06d878691eec9fb1410440025834ebad36adbe607b32de71

Observation 476339c5-5912-4f3c-bb0e-e5fb722a7511 · outbound

This paper cites an unresolved cited work.

Taming LLMs by Scaling Learning Rates with Gradient Grouping Unresolved cited work

Reference 52

Resolution
unresolved
raw_fallback, observed 2026-08-07T11:57:30.141876Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-07T11:57:29.139446Z digest=sha256:7609efab906b3a7fc2081b4c427dabc30229f54ea4edb1f51ba45ab4bed03635

Observation b931a06c-dd32-4450-988e-2683564428c6 · outbound

This paper cites an unresolved cited work.

Taming LLMs by Scaling Learning Rates with Gradient Grouping Unresolved cited work

Reference 53

Resolution
unresolved
raw_fallback, observed 2026-08-07T11:57:30.129665Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-07T11:57:29.142621Z digest=sha256:15d5df643a5723640e2d26a46eb03e01a50d0b6344ec03b7949cbe6a7c78e24d

Observation 16288887-825e-4553-aa28-8fd0e6c42175 · outbound

This paper cites an unresolved cited work.

Taming LLMs by Scaling Learning Rates with Gradient Grouping Unresolved cited work

Reference 54

Resolution
unresolved
raw_fallback, observed 2026-08-07T11:57:30.118083Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-07T11:57:29.145820Z digest=sha256:0bcc698eb998f4b327dfa5370968ce9202b07dee0607f434f7fcf7c46a474abc

Observation d1cd5d36-7d70-46d1-9dfc-ce586ea37d32 · outbound

This paper cites an unresolved cited work.

Taming LLMs by Scaling Learning Rates with Gradient Grouping Unresolved cited work

Reference 55

Resolution
unresolved
raw_fallback, observed 2026-08-07T11:57:30.107042Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-07T11:57:29.149567Z digest=sha256:c76952ac773b4a0d1647b1766d4e24cac9bcb14f3b4e7bff309dbe6711481c0b

Observation dfa9f3e3-acae-45c6-8047-415b0567a2ef · outbound

This paper cites CAME: Confidence-guided Adaptive Memory Efficient Optimization.

Taming LLMs by Scaling Learning Rates with Gradient Grouping CAME: Confidence-guided Adaptive Memory Efficient Optimization

Reference 56

Resolution
unresolved
no resolver link, observed 2026-08-07T11:57:29.152833Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T11:57:29.152833Z digest=sha256:8630635ae883d9c98f8a0aa3842c550f8bccc021a89289ae35b64e485cbaefc7

Observation 28962728-0aa8-47cb-bd0e-ce6920dfe53a · outbound

This paper cites Visual Perception by Large Language Model's Weights.

Taming LLMs by Scaling Learning Rates with Gradient Grouping Visual Perception by Large Language Model's Weights

Reference 57

Resolution
unresolved
no resolver link, observed 2026-08-07T11:57:29.156262Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T11:57:29.156262Z digest=sha256:a443cad9233efef65eb8e40558b13f2c29d49d1fd1f1b873913334b3ff898512

Observation 897b4096-df09-470e-8495-9f998b50dd0e · outbound

This paper cites an unresolved cited work.

Taming LLMs by Scaling Learning Rates with Gradient Grouping Unresolved cited work

Reference 58

Resolution
unresolved
raw_fallback, observed 2026-08-07T11:57:30.095539Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-07T11:57:29.159763Z digest=sha256:866b8ceab7e28e105e0a9afe8dce5339b183d78dc721d27bcaf4b43615e27737

Observation 3a9a2b64-3069-4a7c-ae52-e76ccfec22df · outbound

This paper cites Can a Suit of Armor Conduct Electricity? A New Dataset for Open Book Question Answering.

Taming LLMs by Scaling Learning Rates with Gradient Grouping Can a Suit of Armor Conduct Electricity? A New Dataset for Open Book Question Answering

Reference 59

Resolution
unresolved
no resolver link, observed 2026-08-07T11:57:29.163047Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T11:57:29.163047Z digest=sha256:c21b0d15bb100eb3dcbe3ff8023c2959b94eac4c10411e1fa655fe38c5307dfe

Observation 432c66dc-b5e1-4d55-91c0-e6a82faaa734 · outbound

This paper cites A Theory on Adam Instability in Large-Scale Machine Learning.

Taming LLMs by Scaling Learning Rates with Gradient Grouping A Theory on Adam Instability in Large-Scale Machine Learning

Reference 60

Resolution
unresolved
no resolver link, observed 2026-08-07T11:57:29.166603Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T11:57:29.166603Z digest=sha256:d2aa0501dbdeaa2c8093cac4884e603c02a443df546403381917cf047a966257

Observation 21b291a7-59a5-4f8b-8bc7-13a0211d6ec5 · outbound

This paper cites EDoRA: Efficient Weight-Decomposed Low-Rank Adaptation via Singular Value Decomposition.

Taming LLMs by Scaling Learning Rates with Gradient Grouping EDoRA: Efficient Weight-Decomposed Low-Rank Adaptation via Singular Value Decomposition

Reference 61

Resolution
verified exact
local_arxiv, observed 2026-08-07T11:57:29.517503Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-07T11:57:29.170521Z digest=sha256:398b7d5f9e0fde00c304c046894f5319943db919376b70ae2e8809f06dfcc0d9

Observation 7c681ca2-5a6a-42e7-b647-7079d06871f0 · outbound

This paper cites an unresolved cited work.

Taming LLMs by Scaling Learning Rates with Gradient Grouping Unresolved cited work

Reference 62

Resolution
unresolved
raw_fallback, observed 2026-08-07T11:57:30.082626Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-07T11:57:29.174065Z digest=sha256:46f55520e58269d8894b9c4021e72b0b9397d4a5df50882e1c7c91475d2627bb

Observation b89d3d72-25b6-4554-8823-3b457dc481da · outbound

This paper cites an unresolved cited work.

Taming LLMs by Scaling Learning Rates with Gradient Grouping Unresolved cited work

Reference 63

Resolution
unresolved
no resolver link, observed 2026-08-07T11:57:29.177371Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T11:57:29.177371Z digest=sha256:157e9c65030fd92f37666dc19c26851e974adaa880fef47545f90351aa4667f9

Observation 63d0ecf8-4439-4f80-b8f9-a4d88e9b00a6 · outbound

This paper cites Reddi, Satyen Kale, and Surinder Kumar.

Taming LLMs by Scaling Learning Rates with Gradient Grouping Reddi, Satyen Kale, and Surinder Kumar

Reference 64

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:57:30.062417Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-07T11:57:29.180610Z digest=sha256:05f53430109ffc8076d783cfb0fda3c6be5cf590452d38e0ffb1c944cc290a76

Observation 89ed90fc-8074-4108-bc77-7327446f2e31 · outbound

This paper cites an unresolved cited work.

Taming LLMs by Scaling Learning Rates with Gradient Grouping Unresolved cited work

Reference 65

Resolution
unresolved
no resolver link, observed 2026-08-07T11:57:29.183773Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T11:57:29.183773Z digest=sha256:31fe0c135ed4b9a6fbff84db412218a858253e58a06b6f2e1eb83bf8cfa0ff80

Observation 10022d71-56be-4848-a29a-443a275c82c6 · outbound

This paper cites SocialIQA: Commonsense Reasoning about Social Interactions.

Taming LLMs by Scaling Learning Rates with Gradient Grouping SocialIQA: Commonsense Reasoning about Social Interactions

Reference 66

Resolution
unresolved
no resolver link, observed 2026-08-07T11:57:29.187200Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T11:57:29.187200Z digest=sha256:7e52b910222c50f82b3bdd121fb88484771e1146967b3f206423b6d16927ebdf

Observation 35116b00-3bc9-47c6-9fa4-a821ae2ae174 · outbound

This paper cites an unresolved cited work.

Taming LLMs by Scaling Learning Rates with Gradient Grouping Unresolved cited work

Reference 67

Resolution
unresolved
raw_fallback, observed 2026-08-07T11:57:30.041991Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-07T11:57:29.190488Z digest=sha256:1821a7ab02a543a6c16e083e15d65ca3140a01f20390f3bf1029ed62765e8f48

Observation 2e966505-a288-4ba3-b062-3fa242178992 · outbound

This paper cites Adafactor: Adaptive Learning Rates with Sublinear Memory Cost.

Taming LLMs by Scaling Learning Rates with Gradient Grouping Adafactor: Adaptive Learning Rates with Sublinear Memory Cost

Reference 68

Resolution
unresolved
no resolver link, observed 2026-08-07T11:57:29.193839Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T11:57:29.193839Z digest=sha256:65e0181cd9eac55377497c39d6fd6c789935530e02b90097c39a8121d39989a6

Observation e5acd949-8072-4ba1-91a2-67b0f7a43217 · outbound

This paper cites LLaVA-MoD: Making LLaVA Tiny via MoE Knowledge Distillation.

Taming LLMs by Scaling Learning Rates with Gradient Grouping LLaVA-MoD: Making LLaVA Tiny via MoE Knowledge Distillation

Reference 69

Resolution
unresolved
no resolver link, observed 2026-08-07T11:57:29.197574Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T11:57:29.197574Z digest=sha256:9ffaa7971dfa008c49784f2fd627b34ced87bc3f9e284890e542194205eb606b

Observation e7cbdbc6-d00e-48ee-bb36-3159467da3d8 · outbound

This paper cites an unresolved cited work.

Taming LLMs by Scaling Learning Rates with Gradient Grouping Unresolved cited work

Reference 70

Resolution
unresolved
no resolver link, observed 2026-08-07T11:57:29.200848Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T11:57:29.200848Z digest=sha256:2c7b4d0a69ae27d11170255562df94875c34bad822cff9f3b6982b82e2deba89

Observation 35eca4f5-bb6e-41e3-a6af-5695ac55c152 · outbound

This paper cites Sinha and Michael P.

Taming LLMs by Scaling Learning Rates with Gradient Grouping Sinha and Michael P

Reference 71

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:57:30.021951Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-07T11:57:29.204136Z digest=sha256:5f64cdb4dd51442bfc05db80c778cb1e68d60eb3e2ea7f1bac5966f52a38bf6e

Observation 05858b3f-1b4d-414c-91b4-fcfbaf86820a · outbound

This paper cites an unresolved cited work.

Taming LLMs by Scaling Learning Rates with Gradient Grouping Unresolved cited work

Reference 72

Resolution
unresolved
no resolver link, observed 2026-08-07T11:57:29.207378Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T11:57:29.207378Z digest=sha256:425439811a3c533d0b7c2002e944756c00bb92ef73caf11654aee81d562fe0b2

Observation c41112a3-43f6-4036-b58a-eefa42ff0149 · outbound

This paper cites an unresolved cited work.

Taming LLMs by Scaling Learning Rates with Gradient Grouping Unresolved cited work

Reference 73

Resolution
unresolved
no resolver link, observed 2026-08-07T11:57:29.210678Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T11:57:29.210678Z digest=sha256:df6f792a23dbcdbd4bd6dfe09b3759eda34fcde2985febe42f37de4138c0a3ca

Observation 93957090-2b1a-4fb8-b488-492be99e021e · outbound

This paper cites SOAP: Improving and Stabilizing Shampoo using Adam.

Taming LLMs by Scaling Learning Rates with Gradient Grouping SOAP: Improving and Stabilizing Shampoo using Adam

Reference 74

Resolution
unresolved
no resolver link, observed 2026-08-07T11:57:29.214022Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T11:57:29.214022Z digest=sha256:0033d0e46a6cdcc25f95931a770d472ce8450fd7d89dddc2618a557d3e3aa453

Observation 67f2c45a-9ca7-4b76-84dc-c6dcdb8ce9d1 · outbound

This paper cites GLUE: A Multi-Task Benchmark and Analysis Platform for Natural Language Understanding.

Taming LLMs by Scaling Learning Rates with Gradient Grouping GLUE: A Multi-Task Benchmark and Analysis Platform for Natural Language Understanding

Reference 75

Resolution
unresolved
no resolver link, observed 2026-08-07T11:57:29.217387Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T11:57:29.217387Z digest=sha256:792dae9bf05df9d5bb7adca5934f16dcf3b5906852abccbce6ce8ffcbf9270b5

Observation a56bb22a-4cd7-401b-a945-b401ac0a3a33 · outbound

This paper cites an unresolved cited work.

Taming LLMs by Scaling Learning Rates with Gradient Grouping Unresolved cited work

Reference 76

Resolution
unresolved
raw_fallback, observed 2026-08-07T11:57:29.992788Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-07T11:57:29.220927Z digest=sha256:45de1bc657bd38227a9a108b1893b7f7daba693d1a20ee276a30715647f2b1fe

Observation 58806685-fab1-4a62-8eef-0b1953f973ce · outbound

This paper cites Qwen2.5-Omni Technical Report.

Taming LLMs by Scaling Learning Rates with Gradient Grouping Qwen2.5-Omni Technical Report

Reference 77

Resolution
unresolved
no resolver link, observed 2026-08-07T11:57:29.224160Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T11:57:29.224160Z digest=sha256:d937ac7ec1e1e24d8559cdfefea3b0c59829721a072b08209bb430a2ef3dc31d

Observation 27195656-a9b3-4348-bf49-d5c7179c9221 · outbound

This paper cites an unresolved cited work.

Taming LLMs by Scaling Learning Rates with Gradient Grouping Unresolved cited work

Reference 78

Resolution
unresolved
raw_fallback, observed 2026-08-07T11:57:29.980801Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-07T11:57:29.227542Z digest=sha256:559d8e4da8f87c6726875d251fa8173f7c1ce445121fe7d66cb8f8caba79c90b

Observation 40064f04-3bae-490b-bee2-4066438baa9a · outbound

This paper cites A Survey on Multimodal Large Language Models.

Taming LLMs by Scaling Learning Rates with Gradient Grouping A Survey on Multimodal Large Language Models

Reference 79

Resolution
unresolved
no resolver link, observed 2026-08-07T11:57:29.230732Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T11:57:29.230732Z digest=sha256:d579fda656f474fd54c1d11a9aaf9c9185f2faf38b313229056f101c9b406c89

Observation e0699f0b-c24d-4f2e-b5d5-575773268cff · outbound

This paper cites an unresolved cited work.

Taming LLMs by Scaling Learning Rates with Gradient Grouping Unresolved cited work

Reference 80

Resolution
unresolved
raw_fallback, observed 2026-08-07T11:57:29.967560Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-07T11:57:29.233998Z digest=sha256:bcd131a388cd1c412f4aab7f97fe65a79b98217467bc85aea94d4761901f3634

Observation 488b4aeb-24ea-4312-9f2b-73aa3153362d · outbound

This paper cites an unresolved cited work.

Taming LLMs by Scaling Learning Rates with Gradient Grouping Unresolved cited work

Reference 81

Resolution
unresolved
raw_fallback, observed 2026-08-07T11:57:29.955405Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-07T11:57:29.237565Z digest=sha256:bd449be7d34f8a5b81a48d13eeeedfe727115258e66559a8a380982e6ee6aed1

Observation 29cfd7c2-10a0-46a8-859c-59632a011363 · outbound

This paper cites an unresolved cited work.

Taming LLMs by Scaling Learning Rates with Gradient Grouping Unresolved cited work

Reference 82

Resolution
unresolved
raw_fallback, observed 2026-08-07T11:57:29.943249Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-07T11:57:29.241519Z digest=sha256:4a3b13d066b3c46a308146e1604fc69a3107406525be8a4f455a51b95610fe84

Observation ca378888-4d58-4a7c-b368-cd91497d0e09 · outbound

This paper cites HellaSwag: Can a Machine Really Finish Your Sentence?.

Taming LLMs by Scaling Learning Rates with Gradient Grouping HellaSwag: Can a Machine Really Finish Your Sentence?

Reference 83

Resolution
unresolved
no resolver link, observed 2026-08-07T11:57:29.245272Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T11:57:29.245272Z digest=sha256:869e910bc38be12b95bb31824a18350366e5a5d2251c62c7c584822664115afd

Observation 9a396661-7194-4d09-9c32-792aaac1e40c · outbound

This paper cites Parameter-Efficient Fine-Tuning for Foundation Models.

Taming LLMs by Scaling Learning Rates with Gradient Grouping Parameter-Efficient Fine-Tuning for Foundation Models

Reference 84

Resolution
unresolved
no resolver link, observed 2026-08-07T11:57:29.248704Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T11:57:29.248704Z digest=sha256:ffd491dcc6513ac9649d7541328ce89ea60cc8da10dacf580dae10d6ed62b5fe

Observation 210d5578-f9c6-4cd4-8a05-e7b376f11d57 · outbound

This paper cites an unresolved cited work.

Taming LLMs by Scaling Learning Rates with Gradient Grouping Unresolved cited work

Reference 85

Resolution
unresolved
raw_fallback, observed 2026-08-07T11:57:29.932025Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-07T11:57:29.252216Z digest=sha256:8b07122a8340fc995d9ce2b652d525fe8ef37b7f8d3d3abe2d99111aaee2c960

Observation 08551d6d-0321-4715-9421-bc9c26773f70 · outbound

This paper cites Adam-mini: Use Fewer Learning Rates To Gain More.

Taming LLMs by Scaling Learning Rates with Gradient Grouping Adam-mini: Use Fewer Learning Rates To Gain More

Reference 86

Resolution
unresolved
no resolver link, observed 2026-08-07T11:57:29.255600Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T11:57:29.255600Z digest=sha256:ee85f0d612b58d0bbe016dfdf1128fe7d9fbf94f3c8c88139a8e45eb7b3100eb

Observation 5d65b0bc-6e39-43ed-8c87-c595703282c0 · outbound

This paper cites GaLore: Memory-Efficient LLM Training by Gradient Low-Rank Projection.

Taming LLMs by Scaling Learning Rates with Gradient Grouping GaLore: Memory-Efficient LLM Training by Gradient Low-Rank Projection

Reference 87

Resolution
unresolved
no resolver link, observed 2026-08-07T11:57:29.259287Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T11:57:29.259287Z digest=sha256:d7db0a9074032077edba9317b3efed1703ba4f26b7ff8b0c77add10eeb873b07

Observation b19eaa3e-d997-4475-af1e-df2e6a124e0c · outbound

This paper cites Deconstructing What Makes a Good Optimizer for Language Models.

Taming LLMs by Scaling Learning Rates with Gradient Grouping Deconstructing What Makes a Good Optimizer for Language Models

Reference 88

Resolution
unresolved
no resolver link, observed 2026-08-07T11:57:29.263146Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T11:57:29.263146Z digest=sha256:052a1bcab52ccd8c1ddd0aadf4109aacd6a64f5058c32c94deaef2018cf1443f

Observation fcc700fc-9253-4936-a883-bf2c1addffe4 · outbound

This paper cites TinyLLaVA: A Framework of Small-scale Large Multimodal Models.

Taming LLMs by Scaling Learning Rates with Gradient Grouping TinyLLaVA: A Framework of Small-scale Large Multimodal Models

Reference 89

Resolution
unresolved
no resolver link, observed 2026-08-07T11:57:29.266772Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T11:57:29.266772Z digest=sha256:a77c764e0dd9a56f1392b47dd6ffef87af6e0cde5bec42db8aff34217f2d2261

Observation 4ec620a7-60ca-46d6-8a67-437c76a0961b · outbound

This paper cites APOLLO: SGD-like Memory, AdamW-level Performance.

Taming LLMs by Scaling Learning Rates with Gradient Grouping APOLLO: SGD-like Memory, AdamW-level Performance

Reference 90

Resolution
unresolved
no resolver link, observed 2026-08-07T11:57:29.271257Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T11:57:29.271257Z digest=sha256:f896d62f9be59ce70a085cbde4e6da4e1f4954872cee2d70c6f415634c09c76e

Observation b4fcc1d0-1de3-4fe9-9690-5d3932797f3c · outbound

This paper cites Transformers without Normalization.

Taming LLMs by Scaling Learning Rates with Gradient Grouping Transformers without Normalization

Reference 91

Resolution
unresolved
no resolver link, observed 2026-08-07T11:57:29.274935Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T11:57:29.274935Z digest=sha256:3010bab315e40d31275d7d40ddfee0f79569cfa7fe4a17ad83e5fe6d0f5d44cc

Observation 52272ecf-36fd-481c-87d1-ff47692244fe · outbound

This paper cites an unresolved cited work.

Taming LLMs by Scaling Learning Rates with Gradient Grouping Unresolved cited work

Reference 92

Resolution
unresolved
raw_fallback, observed 2026-08-07T11:57:29.920689Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-07T11:57:29.278658Z digest=sha256:5a399d7dd3fa4d6ad143620a416b5b50615b94fc221c572f7a6688c94fdc25b1

Pith citing papers

Observation 39b37602-eb7a-4980-b4a0-f8e3a6ae7eb7 · inbound

Revealing Modular Gradient Noise Imbalance in LLMs: Calibrating Adam via Signal-to-Noise Ratio cites this paper.

Revealing Modular Gradient Noise Imbalance in LLMs: Calibrating Adam via Signal-to-Noise Ratio Taming LLMs by Scaling Learning Rates with Gradient Grouping

Reference 18

Resolution
verified exact
arxiv_id, observed 2026-05-11T16:41:06.419261Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-09T15:39:51.611115Z digest=sha256:f458e712928cdf4c63a48b8f8eb1009b6e1e163be5f84397f2a0213efa5b48f4

Observation b1e36740-728e-47b1-8590-2f2aeec73c55 · inbound

OmniOpt: Taxonomy, Geometry, and Benchmarking of Modern Optimizers cites this paper.

OmniOpt: Taxonomy, Geometry, and Benchmarking of Modern Optimizers Taming LLMs by Scaling Learning Rates with Gradient Grouping

Reference 59

Resolution
unresolved
no resolver link, observed 2026-07-11T22:10:49.683444Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-11T22:10:49.683444Z digest=sha256:5277e4e07f4adbd50eecc3d6be53a93009c4f402887cdf24c1e299950e394107