Pith. sign in

Paper Citation Record · LEDGER

Tensor Programs V: Tuning Large Neural Networks via Zero-Shot Hyperparameter Transfer

As of 10 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 65 inbound Pith citation observations for arXiv:2203.03466.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2203.03466 v2

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 65 of 65 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-09T06:31:02.800959+00:00

measured 65 of 65 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-09T16:30:40.480336Z

measured 1 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-08-05T02:28:24.338817Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

22
arxiv_reference, observed 2026-08-05T02:28:24.338817Z

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation 49ebe1bb-df9a-4815-97d4-dcafa6184419 · inbound

The Falcon Series of Open Language Models cites this paper.

The Falcon Series of Open Language Models Tensor Programs V: Tuning Large Neural Networks via Zero-Shot Hyperparameter Transfer

Reference 3

Resolution
metadata mismatch
arxiv_id, observed 2026-05-16T09:46:09.828773Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-05-16T09:46:09.701440Z digest=sha256:f52706fdfd235eb66f761add3ee1c81ac41baf6075fe5e67cb5966490a2d9a22

Observation c444c54f-998f-4ff6-8b4f-73b341ea2453 · inbound

MiniCPM: Unveiling the Potential of Small Language Models with Scalable Training Strategies cites this paper.

MiniCPM: Unveiling the Potential of Small Language Models with Scalable Training Strategies Tensor Programs V: Tuning Large Neural Networks via Zero-Shot Hyperparameter Transfer

Reference 46

Resolution
verified exact
arxiv_id, observed 2026-05-13T18:00:53.492235Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-13T18:00:53.389420Z digest=sha256:4adfc8b70c0fd97385c605377a8c3dfea3c850996ac876dd5ae5cdba095e33c9

Observation deffac94-4f6c-4eac-b5e7-54e88ba42668 · inbound

Quantum Machine Learning: A Hands-on Tutorial for Machine Learning Practitioners and Researchers cites this paper.

Quantum Machine Learning: A Hands-on Tutorial for Machine Learning Practitioners and Researchers Tensor Programs V: Tuning Large Neural Networks via Zero-Shot Hyperparameter Transfer

Reference 297

Resolution
unresolved
no resolver link, observed 2026-08-09T16:30:40.480336Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T16:30:40.480336Z digest=sha256:c0daf48ed9e1c9cb87de2f57910437b96d7ab9729c9aaeb7ed6e45fff339e837

Observation f5e1b34c-27a7-47fb-828b-3d361a3956a8 · inbound

Peri-LN: Revisiting Normalization Layer in the Transformer Architecture cites this paper.

Peri-LN: Revisiting Normalization Layer in the Transformer Architecture Tensor Programs V: Tuning Large Neural Networks via Zero-Shot Hyperparameter Transfer

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-09T11:23:41.028667Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T11:23:41.028667Z digest=sha256:e4a5fce036b5e3318de1c4761fd673f5849590b89f450cb340fb0411496a04bb

Observation d732df23-9454-40e7-b4ec-9883a72a9547 · inbound

Learning Real-World Action-Video Dynamics with Heterogeneous Masked Autoregression cites this paper.

Learning Real-World Action-Video Dynamics with Heterogeneous Masked Autoregression Tensor Programs V: Tuning Large Neural Networks via Zero-Shot Hyperparameter Transfer

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-08T22:55:57.885107Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T22:55:57.885107Z digest=sha256:d4abe5c0fcebe132caafd721d7aa74b48a79e375606d2783cff4e301290447be

Observation 04e73110-0a67-4513-93e1-ad65f68f6e77 · inbound

Training Deep Learning Models with Norm-Constrained LMOs cites this paper.

Training Deep Learning Models with Norm-Constrained LMOs Tensor Programs V: Tuning Large Neural Networks via Zero-Shot Hyperparameter Transfer

Reference 213

Resolution
metadata mismatch
arxiv_id, observed 2026-05-21T21:22:37.131238Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-05-21T21:22:36.870292Z digest=sha256:e3ee41140f6c36858afe604fa86e44a34fdd51666d5eea0102b1c82b3e3f54b2

Observation f0f72734-4604-44e2-8874-80dae20de07e · inbound

Adaptive kernel predictors from feature-learning infinite limits of neural networks cites this paper.

Adaptive kernel predictors from feature-learning infinite limits of neural networks Tensor Programs V: Tuning Large Neural Networks via Zero-Shot Hyperparameter Transfer

Reference 63

Resolution
unresolved
no resolver link, observed 2026-08-08T11:19:07.043513Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T11:19:07.043513Z digest=sha256:6f26ea8c739d4f5b5172d7495f1ab10278093e1730d1d2b1d13e8134c12ca09f

Observation fa20308c-6c33-4e18-b5bd-0f87e03f4ec8 · inbound

Dense Local Dependencies Induce Attention-Logit Explosion and Training Instability During Long-Sequence Transformer Training cites this paper.

Dense Local Dependencies Induce Attention-Logit Explosion and Training Instability During Long-Sequence Transformer Training Tensor Programs V: Tuning Large Neural Networks via Zero-Shot Hyperparameter Transfer

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-07T15:19:28.124012Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T15:19:28.124012Z digest=sha256:ebd76c20efa6e699935bbccee7e98ade14b887a719d52f5f41961ea37840f156

Observation 697b68ac-94da-4602-88c9-c39d7c8cf124 · inbound

Eigenspectrum Analysis of Neural Networks without Aspect Ratio Bias cites this paper.

Eigenspectrum Analysis of Neural Networks without Aspect Ratio Bias Tensor Programs V: Tuning Large Neural Networks via Zero-Shot Hyperparameter Transfer

Reference 61

Resolution
unresolved
no resolver link, observed 2026-08-07T06:00:51.918076Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T06:00:51.918076Z digest=sha256:855e7c05d0674b286eeaef6e5d3c40486c28775ecf30e4af4cae01e312dcb524

Observation 52abb96c-b5eb-434a-a4a2-e9fb7259ebac · inbound

NeurIPS 2025 E2LM Competition : Early Training Evaluation of Language Models cites this paper.

NeurIPS 2025 E2LM Competition : Early Training Evaluation of Language Models Tensor Programs V: Tuning Large Neural Networks via Zero-Shot Hyperparameter Transfer

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-07T05:34:24.066290Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:34:24.066290Z digest=sha256:000bbb0aa5433de994b1030c7625b8a0a2195980d61eb516e9906d05b3db09c5

Observation 40a7be92-1fac-4728-bf5c-e254c7922771 · inbound

MiniCPM4: Ultra-Efficient LLMs on End Devices cites this paper.

MiniCPM4: Ultra-Efficient LLMs on End Devices Tensor Programs V: Tuning Large Neural Networks via Zero-Shot Hyperparameter Transfer

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-07T05:31:22.209465Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:31:22.209465Z digest=sha256:6b389ffd6ac21ec67979ee29f038f183dd5f32cfe9d467b3e75652e06042146c

Observation 72c0ebd4-46d4-481f-a0f6-93c6d391586c · inbound

Scaling Transformers for Time Series Forecasting: Do Pretrained Large Models Outperform Small-Scale Alternatives? cites this paper.

Scaling Transformers for Time Series Forecasting: Do Pretrained Large Models Outperform Small-Scale Alternatives? Tensor Programs V: Tuning Large Neural Networks via Zero-Shot Hyperparameter Transfer

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-06T23:11:53.406334Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:11:53.406334Z digest=sha256:f35e6ffb3f2a5b60da6d2e9b772a546f4291c1fd6a8da92e0d4f9184f40fdd20

Observation 9ce7b782-41b0-4714-80a4-1790ebce2626 · inbound

Decoupled Relative Learning Rate Schedules cites this paper.

Decoupled Relative Learning Rate Schedules Tensor Programs V: Tuning Large Neural Networks via Zero-Shot Hyperparameter Transfer

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-06T20:14:11.879402Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T20:14:11.879402Z digest=sha256:a8e05ae976a91e119df32602dd3ad485d990a74eace563f06f2f685c637ff178

Observation a9fd95c1-acac-4044-b814-f73f22274a83 · inbound

SingLoRA: Low Rank Adaptation Using a Single Matrix cites this paper.

SingLoRA: Low Rank Adaptation Using a Single Matrix Tensor Programs V: Tuning Large Neural Networks via Zero-Shot Hyperparameter Transfer

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-06T19:38:16.774558Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:38:16.774558Z digest=sha256:9a70b266c1f7f6dfcd7806e5a357623790786c828299fad8cf3b061f3d2ca62f

Observation 7c6d03a7-2de4-4203-a7a9-4740bcf55b6a · inbound

Simple Convergence Proof of Adam From a Sign-like Descent Perspective cites this paper.

Simple Convergence Proof of Adam From a Sign-like Descent Perspective Tensor Programs V: Tuning Large Neural Networks via Zero-Shot Hyperparameter Transfer

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-06T19:25:24.952647Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:25:24.952647Z digest=sha256:150c4772b4e6351088955fe9992fdf1667b8036b04087dee31dc43acbcd885c6

Observation 3427ea84-fa32-4cde-859f-82d61e64520b · inbound

Tree-Structured Parzen Estimator Can Solve Black-Box Combinatorial Optimization More Efficiently cites this paper.

Tree-Structured Parzen Estimator Can Solve Black-Box Combinatorial Optimization More Efficiently Tensor Programs V: Tuning Large Neural Networks via Zero-Shot Hyperparameter Transfer

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-06T18:44:32.764549Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T18:44:32.764549Z digest=sha256:14a6d84eb29c89fd107a6a08eb29b69d149abd2e2da20e87f4281fa40ec65648

Observation a97201bc-28ae-4c43-ab1a-f42024abd097 · inbound

BlockFFN: Towards End-Side Acceleration-Friendly Mixture-of-Experts with Chunk-Level Activation Sparsity cites this paper.

BlockFFN: Towards End-Side Acceleration-Friendly Mixture-of-Experts with Chunk-Level Activation Sparsity Tensor Programs V: Tuning Large Neural Networks via Zero-Shot Hyperparameter Transfer

Reference 56

Resolution
unresolved
no resolver link, observed 2026-08-06T18:20:00.461750Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T18:20:00.461750Z digest=sha256:a3e575ce66f1543a12a8d42020e58bd3fcd28c684b3db9a59d49c0709ece566c

Observation dc378ec0-4c39-468d-b6b5-76884dd7a08f · inbound

Sub-Scaling Laws: On the Role of Data Density and Training Strategies in LLMs cites this paper.

Sub-Scaling Laws: On the Role of Data Density and Training Strategies in LLMs Tensor Programs V: Tuning Large Neural Networks via Zero-Shot Hyperparameter Transfer

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-06T17:56:44.138484Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T17:56:44.138484Z digest=sha256:a23257fd39a9cd81c68e1d9f75bfd46c2ff60a9e6c58c5cbc6e6372b50a880f4

Observation 34726096-14df-4e98-91c6-80915f8efca6 · inbound

Language Models Improve When Pretraining Data Matches Target Tasks cites this paper.

Language Models Improve When Pretraining Data Matches Target Tasks Tensor Programs V: Tuning Large Neural Networks via Zero-Shot Hyperparameter Transfer

Reference 114

Resolution
unresolved
no resolver link, observed 2026-08-06T16:53:15.816131Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T16:53:15.816131Z digest=sha256:29551c66f2eb641995d6b3d4999187d9f7f85944e594616cf329b737cc9726e2

Observation 483f52a0-8dc1-42fe-bf33-4d97f6bd4c49 · inbound

Geometry of Neural Reinforcement Learning in Continuous State and Action Spaces cites this paper.

Geometry of Neural Reinforcement Learning in Continuous State and Action Spaces Tensor Programs V: Tuning Large Neural Networks via Zero-Shot Hyperparameter Transfer

Reference 135

Resolution
unresolved
no resolver link, observed 2026-08-06T13:21:59.434212Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T13:21:59.434212Z digest=sha256:7561bf6f2d5be28e01edc5621b310f1692636f90a32bff4fd8a0adce0722f6b9

Observation 0743b662-ab65-4238-a029-20da4a2bda73 · inbound

Falcon-H1: A Family of Hybrid-Head Language Models Redefining Efficiency and Performance cites this paper.

Falcon-H1: A Family of Hybrid-Head Language Models Redefining Efficiency and Performance Tensor Programs V: Tuning Large Neural Networks via Zero-Shot Hyperparameter Transfer

Reference 117

Resolution
unresolved
no resolver link, observed 2026-08-06T11:44:05.207827Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T11:44:05.207827Z digest=sha256:72620bef55505b72f420a3c8f2956dacf2ad503cfe766e13636a5eb113850d0b

Observation 81b581f8-cae7-4270-9ef3-5eca51b731dd · inbound

FM4NPP: A Scaling Foundation Model for Nuclear and Particle Physics cites this paper.

FM4NPP: A Scaling Foundation Model for Nuclear and Particle Physics Tensor Programs V: Tuning Large Neural Networks via Zero-Shot Hyperparameter Transfer

Reference 52

Resolution
unresolved
no resolver link, observed 2026-08-05T20:50:54.584556Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T20:50:54.584556Z digest=sha256:e97a655812c5a58af1db422dbb2cbc18d86cc1fd35a6bb3ed2aef5b7603c0fc7

Observation 0af2bd64-eb77-4913-9748-358a415f6665 · inbound

Customizing the Inductive Biases of Softmax Attention using Structured Matrices cites this paper.

Customizing the Inductive Biases of Softmax Attention using Structured Matrices Tensor Programs V: Tuning Large Neural Networks via Zero-Shot Hyperparameter Transfer

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-04T21:33:58.882869Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T21:33:58.882869Z digest=sha256:9637b1b2864e3cf87094bd3a0bf06958ca0d9bcf74a0d27b078e4f376b8aa474

Observation ab5518ee-fc94-4220-a46e-72af56d69016 · inbound

Why Low-Precision Transformer Training Fails: An Analysis on Flash Attention cites this paper.

Why Low-Precision Transformer Training Fails: An Analysis on Flash Attention Tensor Programs V: Tuning Large Neural Networks via Zero-Shot Hyperparameter Transfer

Reference 32

Resolution
verified exact
arxiv_id, observed 2026-05-18T10:02:31.664367Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-18T10:01:56.131253Z digest=sha256:b7ff0410261e73de64c2dfa2f2f7f9115090c56e4a5db9890862da7e4b0678ae

Observation fd660cdf-3628-46a6-b947-574615fd1ace · inbound

Scaling depth capacity via zero/one-layer model expansion cites this paper.

Scaling depth capacity via zero/one-layer model expansion Tensor Programs V: Tuning Large Neural Networks via Zero-Shot Hyperparameter Transfer

Reference 7

Resolution
malformed identifier
no resolver link, observed 2026-08-03T23:39:21.740885Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T23:39:21.740885Z digest=sha256:048f9d9e498a359dfa0184767e655a3e96c3fca6b2f1592a078a68b84a651157

Observation bb368725-c50d-45c7-8f62-05e13dca0e1a · inbound

Deep Delta Learning cites this paper.

Deep Delta Learning Tensor Programs V: Tuning Large Neural Networks via Zero-Shot Hyperparameter Transfer

Reference 16

Resolution
verified exact
arxiv_id, observed 2026-05-16T17:48:11.502720Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-16T17:46:54.521991Z digest=sha256:3216a2d87643077aae2c833a4dcfadb5d974b9b76f52c67b7cc2d72d489410ac

Observation f573360b-9b5b-4e98-adca-cbf73a1091b1 · inbound

Deep Delta Learning cites this paper.

Deep Delta Learning Tensor Programs V: Tuning Large Neural Networks via Zero-Shot Hyperparameter Transfer

Reference 2019

Resolution
unresolved
no resolver link, observed 2026-08-03T13:08:24.859118Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T13:08:24.859118Z digest=sha256:58c13e2b5c256e66e29c4ef67343d6a490da19a93c2be9f0d339cb551e0d4b1e

Observation 8b000a58-37e8-4306-b5c5-7c2ae9ac8dca · inbound

SPARKLING: Balancing Signal Preservation and Symmetry Breaking for Width-Progressive Learning cites this paper.

SPARKLING: Balancing Signal Preservation and Symmetry Breaking for Width-Progressive Learning Tensor Programs V: Tuning Large Neural Networks via Zero-Shot Hyperparameter Transfer

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-03T05:27:12.183400Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T05:27:12.183400Z digest=sha256:5c7c5e999c95ae31e957a27dd84782144a7005ca68c8507f55f1a22efd5e1ad8

Observation 3214c141-a4e9-4554-9b9e-31637c460081 · inbound

Theory of Optimal Learning Rate Schedules and Scaling Laws for a Random Feature Model cites this paper.

Theory of Optimal Learning Rate Schedules and Scaling Laws for a Random Feature Model Tensor Programs V: Tuning Large Neural Networks via Zero-Shot Hyperparameter Transfer

Reference 22

Resolution
metadata mismatch
arxiv_id, observed 2026-05-16T07:00:43.308266Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-16T06:58:38.927268Z digest=sha256:93046d6bb9c7668f42e988ac435f6edaaafeeeaaabab11f821e89a875e5fcb68

Observation 67f86891-955d-4825-a120-c12d2d5ae5e3 · inbound

Spectral Condition for $\mu$P under Width-Depth Scaling cites this paper.

Spectral Condition for $\mu$P under Width-Depth Scaling Tensor Programs V: Tuning Large Neural Networks via Zero-Shot Hyperparameter Transfer

Reference 45

Resolution
metadata mismatch
arxiv_id, observed 2026-05-15T18:06:25.451488Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-15T18:03:31.202134Z digest=sha256:99369d2a0b3fd090e951dc10702c662bb1fdbff6429a240c384cd2b59fdbb8fd

Observation d1182c22-33cd-4657-afeb-7823b8e9844c · inbound

Rethinking Language Model Scaling under Transferable Hypersphere Optimization cites this paper.

Rethinking Language Model Scaling under Transferable Hypersphere Optimization Tensor Programs V: Tuning Large Neural Networks via Zero-Shot Hyperparameter Transfer

Reference 26

Resolution
metadata mismatch
arxiv_id, observed 2026-05-14T21:53:02.414734Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-14T21:51:14.678941Z digest=sha256:21f285aa401bebcd4d52ba65d8a898cef4bcc5fc4d67971da2ec49e7a665f9b7

Observation e90e9a0c-317b-4ffb-b84e-5ab5945b74d3 · inbound

MixAtlas: Uncertainty-aware Data Mixture Optimization for Multimodal LLM Midtraining cites this paper.

MixAtlas: Uncertainty-aware Data Mixture Optimization for Multimodal LLM Midtraining Tensor Programs V: Tuning Large Neural Networks via Zero-Shot Hyperparameter Transfer

Reference 21

Resolution
metadata mismatch
arxiv_id, observed 2026-05-13T20:08:12.645597Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-13T20:07:29.544613Z digest=sha256:bd03e183fdbcf32f64bdeba26c9997a5dc37d2bed74d88583e7ef530eb93816e

Observation eecc8385-3b1a-4813-964e-1f18c989a4de · inbound

There Will Be a Scientific Theory of Deep Learning cites this paper.

There Will Be a Scientific Theory of Deep Learning Tensor Programs V: Tuning Large Neural Networks via Zero-Shot Hyperparameter Transfer

Reference 93

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T15:21:08.817339Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-05-09T20:11:17.616190Z digest=sha256:d6dbff63fbef55a7cd5bd2d49a20d1cb6654ccdc9b4eacd444e7469a3495143e

Observation 54776053-ae13-4975-8faf-474c8e3bcb47 · inbound

Feature Starvation as Geometric Instability in Sparse Autoencoders cites this paper.

Feature Starvation as Geometric Instability in Sparse Autoencoders Tensor Programs V: Tuning Large Neural Networks via Zero-Shot Hyperparameter Transfer

Reference 39

Resolution
verified exact
arxiv_id, observed 2026-05-08T20:09:09.887485Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-08T16:31:38.468164Z digest=sha256:a362dea32a280e752dd21a94bf890571cc27720cf21713f8e5092f58b15b3188

Observation 74854a07-b373-4121-8df6-0403657081e1 · inbound

OrScale: Orthogonalised Optimization with Layer-Wise Trust-Ratio Scaling cites this paper.

OrScale: Orthogonalised Optimization with Layer-Wise Trust-Ratio Scaling Tensor Programs V: Tuning Large Neural Networks via Zero-Shot Hyperparameter Transfer

Reference 20

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T03:05:55.500915Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-11T02:47:08.766380Z digest=sha256:1e2a17291e2bff431825bf723386fb4a424e9f7a111ee6ade3b763552c1bd123

Observation b34976b3-b3b6-439e-bef8-05741a924671 · inbound

Spectral Dynamics in Deep Networks: Feature Learning, Outlier Escape, and Learning Rate Transfer cites this paper.

Spectral Dynamics in Deep Networks: Feature Learning, Outlier Escape, and Learning Rate Transfer Tensor Programs V: Tuning Large Neural Networks via Zero-Shot Hyperparameter Transfer

Reference 21

Resolution
verified exact
arxiv_id, observed 2026-05-11T03:05:53.394859Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-11T03:02:52.833353Z digest=sha256:dabb9d8058fa25b6cc458843c657dd3fc3ba1ff67335fd8c5d7f38f20eab1038

Observation 72b3403e-7121-4245-9d98-d3e8d63d4b50 · inbound

Spectral Dynamics in Deep Networks: Feature Learning, Outlier Escape, and Learning Rate Transfer cites this paper.

Spectral Dynamics in Deep Networks: Feature Learning, Outlier Escape, and Learning Rate Transfer Tensor Programs V: Tuning Large Neural Networks via Zero-Shot Hyperparameter Transfer

Reference 21

Resolution
verified exact
arxiv_id, observed 2026-05-22T10:26:24.110548Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-22T10:25:54.649302Z digest=sha256:98b24ac652280b56e305ecb36eebc20f87c3bb7d34f6d47cc54543809fda6995

Observation 9f20b3bb-8813-4816-bbe7-768aa250bd67 · inbound

Sparse Layers are Critical to Scaling Looped Language Models cites this paper.

Sparse Layers are Critical to Scaling Looped Language Models Tensor Programs V: Tuning Large Neural Networks via Zero-Shot Hyperparameter Transfer

Reference 18

Resolution
metadata mismatch
arxiv_id, observed 2026-05-12T06:21:28.372242Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-12T04:19:12.696712Z digest=sha256:031ee56eb54fa7f16dcf8c587a5ffe7f793aa212b7c75f7ed1cb885a1b6f7a72

Observation 8e1563b3-a639-4f80-b3cb-96f1a24bc552 · inbound

Sparse Layers are Critical to Scaling Looped Language Models cites this paper.

Sparse Layers are Critical to Scaling Looped Language Models Tensor Programs V: Tuning Large Neural Networks via Zero-Shot Hyperparameter Transfer

Reference 18

Resolution
metadata mismatch
arxiv_id, observed 2026-07-01T07:35:29.134236Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-07-01T07:26:37.559459Z digest=sha256:8d61852937a6742f4e069e60778f6eb5ce6f56417a1608aa892d143d7bb5e229

Observation 223ee6e0-6e30-416c-bd99-f75729fde684 · inbound

Intrinsic Muon: Spectral Optimization on Riemannian Matrix Manifolds cites this paper.

Intrinsic Muon: Spectral Optimization on Riemannian Matrix Manifolds Tensor Programs V: Tuning Large Neural Networks via Zero-Shot Hyperparameter Transfer

Reference 67

Resolution
metadata mismatch
arxiv_id, observed 2026-05-12T05:36:26.119308Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-12T05:07:39.558349Z digest=sha256:c952272881c0cfa5f21fee7be99608f7108d36f287d9a94c5b641fd50a828b1c

Observation bbd3b439-e241-4331-b0e6-ac119083b36a · inbound

Mix, Don't Tune: Bilingual Pre-Training Outperforms Hyperparameter Search in Data-Constrained Settings cites this paper.

Mix, Don't Tune: Bilingual Pre-Training Outperforms Hyperparameter Search in Data-Constrained Settings Tensor Programs V: Tuning Large Neural Networks via Zero-Shot Hyperparameter Transfer

Reference 27

Resolution
metadata mismatch
arxiv_id, observed 2026-05-14T20:19:27.406229Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-14T20:17:26.661595Z digest=sha256:eb4870bb5d01ec1749ad6ca3fe837611168b011d132ba5c867aca8e0a661cd88

Observation 137d2869-5d3a-495f-be56-6ce14f851813 · inbound

How to Scale Mixture-of-Experts: From muP to the Maximally Scale-Stable Parameterization cites this paper.

How to Scale Mixture-of-Experts: From muP to the Maximally Scale-Stable Parameterization Tensor Programs V: Tuning Large Neural Networks via Zero-Shot Hyperparameter Transfer

Reference 116

Resolution
metadata mismatch
arxiv_id, observed 2026-05-15T04:49:44.699169Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-05-15T04:45:20.091598Z digest=sha256:b9840f1baf151408bf4d07021e266115e62ed58e9fdffa5355a54c662316e56c

Observation ea5e42c0-568f-4e84-8ab1-52c434573633 · inbound

GQA-{\mu}P: The maximal parameterization update for grouped query attention cites this paper.

GQA-{\mu}P: The maximal parameterization update for grouped query attention Tensor Programs V: Tuning Large Neural Networks via Zero-Shot Hyperparameter Transfer

Reference 19

Resolution
verified exact
arxiv_id, observed 2026-05-19T16:37:39.825511Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-19T16:35:40.231293Z digest=sha256:cf70c41b3d4c1c960253eacbb4824fbf1a30d174a42ace0aa02e0397ec89d527

Observation 179ecece-900d-401e-86c9-a166afde2f41 · inbound

Simply Stabilizing the Loop via Fully Looped Transformer cites this paper.

Simply Stabilizing the Loop via Fully Looped Transformer Tensor Programs V: Tuning Large Neural Networks via Zero-Shot Hyperparameter Transfer

Reference 3

Resolution
verified exact
arxiv_id, observed 2026-07-01T14:05:47.263372Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-06-30T22:15:47.909334Z digest=sha256:e6563ffa45683e7d2e0870ed77c56447c7b739572a65a7f861ecf151362a9147

Observation 7a499c47-f64b-49f4-a754-0378549c499c · inbound

Block-Based Double Decoders cites this paper.

Block-Based Double Decoders Tensor Programs V: Tuning Large Neural Networks via Zero-Shot Hyperparameter Transfer

Reference 16

Resolution
metadata mismatch
arxiv_id, observed 2026-05-20T22:09:07.439320Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-20T22:05:25.465669Z digest=sha256:6e70cafa2d00bf0cb90da8eeff6be1800f4b89b446834d95812bf34dbc5e4968

Observation e40a3231-baa4-4812-8478-49d6dfdc0cdb · inbound

Block-Based Double Decoders cites this paper.

Block-Based Double Decoders Tensor Programs V: Tuning Large Neural Networks via Zero-Shot Hyperparameter Transfer

Reference 16

Resolution
metadata mismatch
arxiv_id, observed 2026-07-01T14:15:47.113144Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-06-30T22:10:23.224012Z digest=sha256:99d21c512400e92ebb87f23f736a4d6fa3879d0313378e52bb1307f18d99f230

Observation 8c8492ca-96bf-428a-b8bd-ec0762b8846e · inbound

Unified Neural Scaling Laws cites this paper.

Unified Neural Scaling Laws Tensor Programs V: Tuning Large Neural Networks via Zero-Shot Hyperparameter Transfer

Reference 29

Resolution
metadata mismatch
arxiv_id, observed 2026-06-29T23:44:03.344613Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-06-29T22:56:43.393302Z digest=sha256:1794ca48b3008495eed729362007cc051fdc501bd5290c1e6ece0ed9aad376dd

Observation d0fdc726-c70c-4df4-9a57-e1a44dd9b2bc · inbound

MuCon: Clipped Muon Updates for LLM Training cites this paper.

MuCon: Clipped Muon Updates for LLM Training Tensor Programs V: Tuning Large Neural Networks via Zero-Shot Hyperparameter Transfer

Reference 4

Resolution
verified exact
arxiv_id, observed 2026-06-29T20:03:56.395584Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-06-29T19:56:47.120210Z digest=sha256:814420e558f8baf24ce5e69d9f48357b1a5e35656aa618e57b507d017bae2c3a

Observation f3b1de82-4626-42fd-9ac3-884dae64da54 · inbound

Spectral Reach: Understanding Neural Scaling as Progress into the Spectral Tail cites this paper.

Spectral Reach: Understanding Neural Scaling as Progress into the Spectral Tail Tensor Programs V: Tuning Large Neural Networks via Zero-Shot Hyperparameter Transfer

Reference 2025

Resolution
unresolved
no resolver link, observed 2026-08-04T04:58:48.901353Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T04:58:48.901353Z digest=sha256:d27d0527755971bb7e5b4d3bfec840a8f99be2c71e01f8ec3bc73e208e04557a

Observation de0c3d54-50eb-47db-8ec4-798875eb26da · inbound

On the Scaling of PEFT: Towards Million Personal Models of Trillion Parameters cites this paper.

On the Scaling of PEFT: Towards Million Personal Models of Trillion Parameters Tensor Programs V: Tuning Large Neural Networks via Zero-Shot Hyperparameter Transfer

Reference 36

Resolution
metadata mismatch
arxiv_id, observed 2026-07-01T22:16:16.642255Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-06-28T15:32:19.456119Z digest=sha256:8bee531778578e797a5255bd63b46389c5307d9b0e668a4f80949d214ae77acf

Observation 7ae09695-3ace-4ddd-970f-0f69a7e1452b · inbound

Unlocking Feature Learning in Gated Delta Networks at Scale cites this paper.

Unlocking Feature Learning in Gated Delta Networks at Scale Tensor Programs V: Tuning Large Neural Networks via Zero-Shot Hyperparameter Transfer

Reference 14

Resolution
verified exact
arxiv_id, observed 2026-07-02T02:06:26.424178Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-06-28T11:18:55.669013Z digest=sha256:befe6820846528dd4bac9e757926d18e90d764a0336d79b9b86dcc19eb37b42e

Observation aee8d983-2224-4d30-94c8-a2bf7f7e5443 · inbound

Spectral Scaling Laws of Muon cites this paper.

Spectral Scaling Laws of Muon Tensor Programs V: Tuning Large Neural Networks via Zero-Shot Hyperparameter Transfer

Reference 14

Resolution
metadata mismatch
arxiv_id, observed 2026-07-02T02:06:26.383110Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-06-28T11:19:05.939340Z digest=sha256:a5f3e9bc832b35b494af74c7b1dd0881583527c2d088e3335c1dcefb579562e1

Observation de8c3ed4-101f-4eba-bafc-97276805c1e0 · inbound

Double Preconditioning (DoPr): Optimization for Test-Time Performance, not Validation Loss cites this paper.

Double Preconditioning (DoPr): Optimization for Test-Time Performance, not Validation Loss Tensor Programs V: Tuning Large Neural Networks via Zero-Shot Hyperparameter Transfer

Reference 102

Resolution
metadata mismatch
arxiv_id, observed 2026-07-02T11:56:56.103926Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-06-28T02:35:39.845487Z digest=sha256:e8e0c0d513275d2e5c24a075e120f5aca3f0bc12fd9c779ac871ee4e219b095f

Observation 344187bf-160c-4d78-a7e9-d048c05d745f · inbound

On the Residual Scaling of Looped Transformers: Stability and Transferability cites this paper.

On the Residual Scaling of Looped Transformers: Stability and Transferability Tensor Programs V: Tuning Large Neural Networks via Zero-Shot Hyperparameter Transfer

Reference 31

Resolution
malformed identifier
arxiv_id, observed 2026-07-03T21:08:58.378584Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-06-27T00:53:51.358774Z digest=sha256:0f35ec1d2986968aaae7be535b26c5efa888a4f17cef08ea803cd0e2438cb020

Observation 14a3a96b-c447-4fd7-b4ff-5620e21c9628 · inbound

LLM Evolution as an Industry-Scale Ecosystem: A Lifecycle Perspective on Continual Learning cites this paper.

LLM Evolution as an Industry-Scale Ecosystem: A Lifecycle Perspective on Continual Learning Tensor Programs V: Tuning Large Neural Networks via Zero-Shot Hyperparameter Transfer

Reference 125

Resolution
verified exact
arxiv_id, observed 2026-07-03T16:48:39.965275Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-06-27T05:02:18.347642Z digest=sha256:5c2cf5d8b1ee4e209b5cc27c358900b14512e7adb32856bfca006d82fcb2b901

Observation ad024d55-dee2-4067-9e3c-0b8a1ac4c6bd · inbound

Improving Neural Network Training by Decoupling the Magnitude and Direction of Weight Vectors cites this paper.

Improving Neural Network Training by Decoupling the Magnitude and Direction of Weight Vectors Tensor Programs V: Tuning Large Neural Networks via Zero-Shot Hyperparameter Transfer

Reference 145

Resolution
verified exact
arxiv_id, observed 2026-07-04T20:30:07.815811Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-06-25T20:05:09.179627Z digest=sha256:28a94223cd34e9d7d7e727d9adbbb1577f142e4c1db998cc689e127851f00f83

Observation d0ce3013-56c9-42f3-8993-87acf7fe0c60 · inbound

Improving Neural Network Training by Decoupling the Magnitude and Direction of Weight Vectors cites this paper.

Improving Neural Network Training by Decoupling the Magnitude and Direction of Weight Vectors Tensor Programs V: Tuning Large Neural Networks via Zero-Shot Hyperparameter Transfer

Reference 88

Resolution
unresolved
no resolver link, observed 2026-08-02T10:14:12.944746Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-02T10:14:12.944746Z digest=sha256:6c33534338a70f1c8461450af28889a2faf5bb289b0edf5da6a6d41af2fbf140

Observation 323bfc35-c383-4d2e-84ee-2b2faada5d28 · inbound

Geometric Dyson Brownian Motions and the Free Log-Normal Limit for a Non-Square Gaussian Matrix Product cites this paper.

Geometric Dyson Brownian Motions and the Free Log-Normal Limit for a Non-Square Gaussian Matrix Product Tensor Programs V: Tuning Large Neural Networks via Zero-Shot Hyperparameter Transfer

Reference 47

Resolution
verified exact
arxiv_id, observed 2026-07-01T12:45:44.992051Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-07-01T01:42:14.145227Z digest=sha256:08ae81531b7227a87bf7f4fdb60089c76da37c7317a85ab717167825c251e38d

Observation cd3b708c-ec2c-4b61-bcfe-e8f332637497 · inbound

Geometric Dyson Brownian Motions and the Free Log-Normal Limit for a Non-Square Gaussian Matrix Product cites this paper.

Geometric Dyson Brownian Motions and the Free Log-Normal Limit for a Non-Square Gaussian Matrix Product Tensor Programs V: Tuning Large Neural Networks via Zero-Shot Hyperparameter Transfer

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-02T09:36:52.548930Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T09:36:52.548930Z digest=sha256:9a0a6188e3142ae65526e9db281172810ba66e1a07522c0c7caaabc220bafd82

Observation e15db8af-203a-4c7f-9dfb-862f7c416441 · inbound

The Role of Rigor in Artificial Intelligence cites this paper.

The Role of Rigor in Artificial Intelligence Tensor Programs V: Tuning Large Neural Networks via Zero-Shot Hyperparameter Transfer

Reference 49

Resolution
unresolved
no resolver link, observed 2026-07-12T16:22:39.179665Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T16:22:39.179665Z digest=sha256:31dcac26069681fe76c482e88ba2fe0b19d8586277028a6d192ae17351d9f99f

Observation 6143a722-f2b9-467f-bd1d-61bb7bc87469 · inbound

Index SLM Technical Report cites this paper.

Index SLM Technical Report Tensor Programs V: Tuning Large Neural Networks via Zero-Shot Hyperparameter Transfer

Reference 24

Resolution
unresolved
no resolver link, observed 2026-07-14T14:51:11.146058Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T14:51:11.146058Z digest=sha256:1572089db3fd79f24d86c8d7372b7da3ed4c405210cc07331e387a0760569e78

Observation 38cbf6cf-6290-4515-aa64-80b98712a8be · inbound

Scale Weight Decay and Train Better cites this paper.

Scale Weight Decay and Train Better Tensor Programs V: Tuning Large Neural Networks via Zero-Shot Hyperparameter Transfer

Reference 41

Resolution
unresolved
no resolver link, observed 2026-07-30T12:53:41.041793Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-30T12:53:41.041793Z digest=sha256:7d04ebcea982c64dec43ac77417036e3629b10cc117ccbec699855a88849639f

Observation 4feb1af7-058a-4472-a586-d05d034f1dd2 · inbound

Bridging Compute- and Data-Optimal Pretraining cites this paper.

Bridging Compute- and Data-Optimal Pretraining Tensor Programs V: Tuning Large Neural Networks via Zero-Shot Hyperparameter Transfer

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-01T03:02:01.193632Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T03:02:01.193632Z digest=sha256:f440dcf348fbcb4679c07b77ac12afb19ef7bb32c9d777ba706006dc878f95ed

Observation f9cc9133-94d9-4698-a52f-d1cec4bca3cb · inbound

Chimera: Designing and Chinchilla-Scaling Hybrid Visual Diffusion Transformers cites this paper.

Chimera: Designing and Chinchilla-Scaling Hybrid Visual Diffusion Transformers Tensor Programs V: Tuning Large Neural Networks via Zero-Shot Hyperparameter Transfer

Reference 93

Resolution
unresolved
no resolver link, observed 2026-07-31T02:16:08.253839Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-31T02:16:08.253839Z digest=sha256:ac9e0f76555f6cf56807d71b0e9c14888c33b213c0fe933eb93fb1b3a2a5f5cb

Observation 3d27a86f-b256-4327-8e34-4a6e96894d2e · inbound

Sign compression for Muon: SignMuon, MuonSign, and the Limits of Error Feedback cites this paper.

Sign compression for Muon: SignMuon, MuonSign, and the Limits of Error Feedback Tensor Programs V: Tuning Large Neural Networks via Zero-Shot Hyperparameter Transfer

Reference 2021

Resolution
unresolved
no resolver link, observed 2026-08-03T02:09:02.591486Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T02:09:02.591486Z digest=sha256:6200de88c924d371933d2ea4a525100c9964247b5484eef4ccc80ab2e0d5c135