Pith. sign in

Paper Citation Record · LEDGER

Tensor Programs V: Tuning Large Neural Networks via Zero-Shot Hyperparameter Transfer

As of 10 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 65 inbound Pith citation observations for arXiv:2203.03466.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2203.03466 v2

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 65 of 65 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-10T06:31:04.303077+00:00

measured 65 of 65 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-09T16:30:40.480336Z

measured 1 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-08-05T02:28:24.338817Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

22
arxiv_reference, observed 2026-08-05T02:28:24.338817Z

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation 49ebe1bb-df9a-4815-97d4-dcafa6184419 · inbound

The Falcon Series of Open Language Models cites this paper.

The Falcon Series of Open Language Models Tensor Programs V: Tuning Large Neural Networks via Zero-Shot Hyperparameter Transfer

Reference 3

Resolution
metadata mismatch
arxiv_id, observed 2026-05-16T09:46:09.828773Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-05-16T09:46:09.701440Z digest=sha256:bd2d68c3cafb59a2b2ad83852364037b265230d7bf969ea56d75f538a1c8e9c0

Observation c444c54f-998f-4ff6-8b4f-73b341ea2453 · inbound

MiniCPM: Unveiling the Potential of Small Language Models with Scalable Training Strategies cites this paper.

MiniCPM: Unveiling the Potential of Small Language Models with Scalable Training Strategies Tensor Programs V: Tuning Large Neural Networks via Zero-Shot Hyperparameter Transfer

Reference 46

Resolution
verified exact
arxiv_id, observed 2026-05-13T18:00:53.492235Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-13T18:00:53.389420Z digest=sha256:6e07837a0644bd1c2946c50189505ed15be5c25f012a3971dc0f9f10e64c0802

Observation deffac94-4f6c-4eac-b5e7-54e88ba42668 · inbound

Quantum Machine Learning: A Hands-on Tutorial for Machine Learning Practitioners and Researchers cites this paper.

Quantum Machine Learning: A Hands-on Tutorial for Machine Learning Practitioners and Researchers Tensor Programs V: Tuning Large Neural Networks via Zero-Shot Hyperparameter Transfer

Reference 297

Resolution
unresolved
no resolver link, observed 2026-08-09T16:30:40.480336Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T16:30:40.480336Z digest=sha256:c0daf48ed9e1c9cb87de2f57910437b96d7ab9729c9aaeb7ed6e45fff339e837

Observation f5e1b34c-27a7-47fb-828b-3d361a3956a8 · inbound

Peri-LN: Revisiting Normalization Layer in the Transformer Architecture cites this paper.

Peri-LN: Revisiting Normalization Layer in the Transformer Architecture Tensor Programs V: Tuning Large Neural Networks via Zero-Shot Hyperparameter Transfer

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-09T11:23:41.028667Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T11:23:41.028667Z digest=sha256:e4a5fce036b5e3318de1c4761fd673f5849590b89f450cb340fb0411496a04bb

Observation d732df23-9454-40e7-b4ec-9883a72a9547 · inbound

Learning Real-World Action-Video Dynamics with Heterogeneous Masked Autoregression cites this paper.

Learning Real-World Action-Video Dynamics with Heterogeneous Masked Autoregression Tensor Programs V: Tuning Large Neural Networks via Zero-Shot Hyperparameter Transfer

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-08T22:55:57.885107Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T22:55:57.885107Z digest=sha256:d4abe5c0fcebe132caafd721d7aa74b48a79e375606d2783cff4e301290447be

Observation 04e73110-0a67-4513-93e1-ad65f68f6e77 · inbound

Training Deep Learning Models with Norm-Constrained LMOs cites this paper.

Training Deep Learning Models with Norm-Constrained LMOs Tensor Programs V: Tuning Large Neural Networks via Zero-Shot Hyperparameter Transfer

Reference 213

Resolution
metadata mismatch
arxiv_id, observed 2026-05-21T21:22:37.131238Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-05-21T21:22:36.870292Z digest=sha256:b9cc8f294213ab7200e092138c3cdabef0f11194fdb4f2e070fd383d9c1b3422

Observation f0f72734-4604-44e2-8874-80dae20de07e · inbound

Adaptive kernel predictors from feature-learning infinite limits of neural networks cites this paper.

Adaptive kernel predictors from feature-learning infinite limits of neural networks Tensor Programs V: Tuning Large Neural Networks via Zero-Shot Hyperparameter Transfer

Reference 63

Resolution
unresolved
no resolver link, observed 2026-08-08T11:19:07.043513Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T11:19:07.043513Z digest=sha256:6f26ea8c739d4f5b5172d7495f1ab10278093e1730d1d2b1d13e8134c12ca09f

Observation fa20308c-6c33-4e18-b5bd-0f87e03f4ec8 · inbound

Dense Local Dependencies Induce Attention-Logit Explosion and Training Instability During Long-Sequence Transformer Training cites this paper.

Dense Local Dependencies Induce Attention-Logit Explosion and Training Instability During Long-Sequence Transformer Training Tensor Programs V: Tuning Large Neural Networks via Zero-Shot Hyperparameter Transfer

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-07T15:19:28.124012Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T15:19:28.124012Z digest=sha256:ebd76c20efa6e699935bbccee7e98ade14b887a719d52f5f41961ea37840f156

Observation 697b68ac-94da-4602-88c9-c39d7c8cf124 · inbound

Eigenspectrum Analysis of Neural Networks without Aspect Ratio Bias cites this paper.

Eigenspectrum Analysis of Neural Networks without Aspect Ratio Bias Tensor Programs V: Tuning Large Neural Networks via Zero-Shot Hyperparameter Transfer

Reference 61

Resolution
unresolved
no resolver link, observed 2026-08-07T06:00:51.918076Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T06:00:51.918076Z digest=sha256:855e7c05d0674b286eeaef6e5d3c40486c28775ecf30e4af4cae01e312dcb524

Observation 52abb96c-b5eb-434a-a4a2-e9fb7259ebac · inbound

NeurIPS 2025 E2LM Competition : Early Training Evaluation of Language Models cites this paper.

NeurIPS 2025 E2LM Competition : Early Training Evaluation of Language Models Tensor Programs V: Tuning Large Neural Networks via Zero-Shot Hyperparameter Transfer

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-07T05:34:24.066290Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:34:24.066290Z digest=sha256:000bbb0aa5433de994b1030c7625b8a0a2195980d61eb516e9906d05b3db09c5

Observation 40a7be92-1fac-4728-bf5c-e254c7922771 · inbound

MiniCPM4: Ultra-Efficient LLMs on End Devices cites this paper.

MiniCPM4: Ultra-Efficient LLMs on End Devices Tensor Programs V: Tuning Large Neural Networks via Zero-Shot Hyperparameter Transfer

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-07T05:31:22.209465Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:31:22.209465Z digest=sha256:6b389ffd6ac21ec67979ee29f038f183dd5f32cfe9d467b3e75652e06042146c

Observation 72c0ebd4-46d4-481f-a0f6-93c6d391586c · inbound

Scaling Transformers for Time Series Forecasting: Do Pretrained Large Models Outperform Small-Scale Alternatives? cites this paper.

Scaling Transformers for Time Series Forecasting: Do Pretrained Large Models Outperform Small-Scale Alternatives? Tensor Programs V: Tuning Large Neural Networks via Zero-Shot Hyperparameter Transfer

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-06T23:11:53.406334Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:11:53.406334Z digest=sha256:24d802f3f8ececb6fd98bd4cfbd05c7ab821c373d9d3e5b2fbd51a516905a480

Observation 9ce7b782-41b0-4714-80a4-1790ebce2626 · inbound

Decoupled Relative Learning Rate Schedules cites this paper.

Decoupled Relative Learning Rate Schedules Tensor Programs V: Tuning Large Neural Networks via Zero-Shot Hyperparameter Transfer

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-06T20:14:11.879402Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T20:14:11.879402Z digest=sha256:a8e05ae976a91e119df32602dd3ad485d990a74eace563f06f2f685c637ff178

Observation a9fd95c1-acac-4044-b814-f73f22274a83 · inbound

SingLoRA: Low Rank Adaptation Using a Single Matrix cites this paper.

SingLoRA: Low Rank Adaptation Using a Single Matrix Tensor Programs V: Tuning Large Neural Networks via Zero-Shot Hyperparameter Transfer

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-06T19:38:16.774558Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:38:16.774558Z digest=sha256:9a70b266c1f7f6dfcd7806e5a357623790786c828299fad8cf3b061f3d2ca62f

Observation 7c6d03a7-2de4-4203-a7a9-4740bcf55b6a · inbound

Simple Convergence Proof of Adam From a Sign-like Descent Perspective cites this paper.

Simple Convergence Proof of Adam From a Sign-like Descent Perspective Tensor Programs V: Tuning Large Neural Networks via Zero-Shot Hyperparameter Transfer

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-06T19:25:24.952647Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:25:24.952647Z digest=sha256:150c4772b4e6351088955fe9992fdf1667b8036b04087dee31dc43acbcd885c6

Observation 3427ea84-fa32-4cde-859f-82d61e64520b · inbound

Tree-Structured Parzen Estimator Can Solve Black-Box Combinatorial Optimization More Efficiently cites this paper.

Tree-Structured Parzen Estimator Can Solve Black-Box Combinatorial Optimization More Efficiently Tensor Programs V: Tuning Large Neural Networks via Zero-Shot Hyperparameter Transfer

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-06T18:44:32.764549Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T18:44:32.764549Z digest=sha256:14a6d84eb29c89fd107a6a08eb29b69d149abd2e2da20e87f4281fa40ec65648

Observation a97201bc-28ae-4c43-ab1a-f42024abd097 · inbound

BlockFFN: Towards End-Side Acceleration-Friendly Mixture-of-Experts with Chunk-Level Activation Sparsity cites this paper.

BlockFFN: Towards End-Side Acceleration-Friendly Mixture-of-Experts with Chunk-Level Activation Sparsity Tensor Programs V: Tuning Large Neural Networks via Zero-Shot Hyperparameter Transfer

Reference 56

Resolution
unresolved
no resolver link, observed 2026-08-06T18:20:00.461750Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T18:20:00.461750Z digest=sha256:a3e575ce66f1543a12a8d42020e58bd3fcd28c684b3db9a59d49c0709ece566c

Observation dc378ec0-4c39-468d-b6b5-76884dd7a08f · inbound

Sub-Scaling Laws: On the Role of Data Density and Training Strategies in LLMs cites this paper.

Sub-Scaling Laws: On the Role of Data Density and Training Strategies in LLMs Tensor Programs V: Tuning Large Neural Networks via Zero-Shot Hyperparameter Transfer

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-06T17:56:44.138484Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T17:56:44.138484Z digest=sha256:a23257fd39a9cd81c68e1d9f75bfd46c2ff60a9e6c58c5cbc6e6372b50a880f4

Observation 34726096-14df-4e98-91c6-80915f8efca6 · inbound

Language Models Improve When Pretraining Data Matches Target Tasks cites this paper.

Language Models Improve When Pretraining Data Matches Target Tasks Tensor Programs V: Tuning Large Neural Networks via Zero-Shot Hyperparameter Transfer

Reference 114

Resolution
unresolved
no resolver link, observed 2026-08-06T16:53:15.816131Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T16:53:15.816131Z digest=sha256:29551c66f2eb641995d6b3d4999187d9f7f85944e594616cf329b737cc9726e2

Observation 483f52a0-8dc1-42fe-bf33-4d97f6bd4c49 · inbound

Geometry of Neural Reinforcement Learning in Continuous State and Action Spaces cites this paper.

Geometry of Neural Reinforcement Learning in Continuous State and Action Spaces Tensor Programs V: Tuning Large Neural Networks via Zero-Shot Hyperparameter Transfer

Reference 135

Resolution
unresolved
no resolver link, observed 2026-08-06T13:21:59.434212Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T13:21:59.434212Z digest=sha256:42b58c848cf68d2f01d49b1433361ed1bbae6ab94a43ca0615cc72f63fec1262

Observation 0743b662-ab65-4238-a029-20da4a2bda73 · inbound

Falcon-H1: A Family of Hybrid-Head Language Models Redefining Efficiency and Performance cites this paper.

Falcon-H1: A Family of Hybrid-Head Language Models Redefining Efficiency and Performance Tensor Programs V: Tuning Large Neural Networks via Zero-Shot Hyperparameter Transfer

Reference 117

Resolution
unresolved
no resolver link, observed 2026-08-06T11:44:05.207827Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T11:44:05.207827Z digest=sha256:72620bef55505b72f420a3c8f2956dacf2ad503cfe766e13636a5eb113850d0b

Observation 81b581f8-cae7-4270-9ef3-5eca51b731dd · inbound

FM4NPP: A Scaling Foundation Model for Nuclear and Particle Physics cites this paper.

FM4NPP: A Scaling Foundation Model for Nuclear and Particle Physics Tensor Programs V: Tuning Large Neural Networks via Zero-Shot Hyperparameter Transfer

Reference 52

Resolution
unresolved
no resolver link, observed 2026-08-05T20:50:54.584556Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T20:50:54.584556Z digest=sha256:e97a655812c5a58af1db422dbb2cbc18d86cc1fd35a6bb3ed2aef5b7603c0fc7

Observation 0af2bd64-eb77-4913-9748-358a415f6665 · inbound

Customizing the Inductive Biases of Softmax Attention using Structured Matrices cites this paper.

Customizing the Inductive Biases of Softmax Attention using Structured Matrices Tensor Programs V: Tuning Large Neural Networks via Zero-Shot Hyperparameter Transfer

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-04T21:33:58.882869Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T21:33:58.882869Z digest=sha256:9637b1b2864e3cf87094bd3a0bf06958ca0d9bcf74a0d27b078e4f376b8aa474

Observation ab5518ee-fc94-4220-a46e-72af56d69016 · inbound

Why Low-Precision Transformer Training Fails: An Analysis on Flash Attention cites this paper.

Why Low-Precision Transformer Training Fails: An Analysis on Flash Attention Tensor Programs V: Tuning Large Neural Networks via Zero-Shot Hyperparameter Transfer

Reference 32

Resolution
verified exact
arxiv_id, observed 2026-05-18T10:02:31.664367Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-18T10:01:56.131253Z digest=sha256:df411ee958174f950430ddf81ec7d03bbd0b03a2b8d0dc7617b081cb0bc0c4e0

Observation fd660cdf-3628-46a6-b947-574615fd1ace · inbound

Scaling depth capacity via zero/one-layer model expansion cites this paper.

Scaling depth capacity via zero/one-layer model expansion Tensor Programs V: Tuning Large Neural Networks via Zero-Shot Hyperparameter Transfer

Reference 7

Resolution
malformed identifier
no resolver link, observed 2026-08-03T23:39:21.740885Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T23:39:21.740885Z digest=sha256:048f9d9e498a359dfa0184767e655a3e96c3fca6b2f1592a078a68b84a651157

Observation bb368725-c50d-45c7-8f62-05e13dca0e1a · inbound

Deep Delta Learning cites this paper.

Deep Delta Learning Tensor Programs V: Tuning Large Neural Networks via Zero-Shot Hyperparameter Transfer

Reference 16

Resolution
verified exact
arxiv_id, observed 2026-05-16T17:48:11.502720Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-16T17:46:54.521991Z digest=sha256:9d503c89b08d9f675f43cbe7df31a36875bf8194ad8022aa414688e18074d4c0

Observation f573360b-9b5b-4e98-adca-cbf73a1091b1 · inbound

Deep Delta Learning cites this paper.

Deep Delta Learning Tensor Programs V: Tuning Large Neural Networks via Zero-Shot Hyperparameter Transfer

Reference 2019

Resolution
unresolved
no resolver link, observed 2026-08-03T13:08:24.859118Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T13:08:24.859118Z digest=sha256:58c13e2b5c256e66e29c4ef67343d6a490da19a93c2be9f0d339cb551e0d4b1e

Observation 8b000a58-37e8-4306-b5c5-7c2ae9ac8dca · inbound

SPARKLING: Balancing Signal Preservation and Symmetry Breaking for Width-Progressive Learning cites this paper.

SPARKLING: Balancing Signal Preservation and Symmetry Breaking for Width-Progressive Learning Tensor Programs V: Tuning Large Neural Networks via Zero-Shot Hyperparameter Transfer

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-03T05:27:12.183400Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T05:27:12.183400Z digest=sha256:5c7c5e999c95ae31e957a27dd84782144a7005ca68c8507f55f1a22efd5e1ad8

Observation 3214c141-a4e9-4554-9b9e-31637c460081 · inbound

Theory of Optimal Learning Rate Schedules and Scaling Laws for a Random Feature Model cites this paper.

Theory of Optimal Learning Rate Schedules and Scaling Laws for a Random Feature Model Tensor Programs V: Tuning Large Neural Networks via Zero-Shot Hyperparameter Transfer

Reference 22

Resolution
metadata mismatch
arxiv_id, observed 2026-05-16T07:00:43.308266Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-16T06:58:38.927268Z digest=sha256:5af5d78028723f8b296ebe354af8c2857504e1816d3d6a4d650a4a15ffa3db03

Observation 67f86891-955d-4825-a120-c12d2d5ae5e3 · inbound

Spectral Condition for $\mu$P under Width-Depth Scaling cites this paper.

Spectral Condition for $\mu$P under Width-Depth Scaling Tensor Programs V: Tuning Large Neural Networks via Zero-Shot Hyperparameter Transfer

Reference 45

Resolution
metadata mismatch
arxiv_id, observed 2026-05-15T18:06:25.451488Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-15T18:03:31.202134Z digest=sha256:622c11e58e506ccc8e5294b03a57155a0383044643baed8297aadf037175943d

Observation d1182c22-33cd-4657-afeb-7823b8e9844c · inbound

Rethinking Language Model Scaling under Transferable Hypersphere Optimization cites this paper.

Rethinking Language Model Scaling under Transferable Hypersphere Optimization Tensor Programs V: Tuning Large Neural Networks via Zero-Shot Hyperparameter Transfer

Reference 26

Resolution
metadata mismatch
arxiv_id, observed 2026-05-14T21:53:02.414734Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-14T21:51:14.678941Z digest=sha256:e9d2e4f8a44bd8308665eaf656445d83b98edbc3ce2bce341abd35969764e57f

Observation e90e9a0c-317b-4ffb-b84e-5ab5945b74d3 · inbound

MixAtlas: Uncertainty-aware Data Mixture Optimization for Multimodal LLM Midtraining cites this paper.

MixAtlas: Uncertainty-aware Data Mixture Optimization for Multimodal LLM Midtraining Tensor Programs V: Tuning Large Neural Networks via Zero-Shot Hyperparameter Transfer

Reference 21

Resolution
metadata mismatch
arxiv_id, observed 2026-05-13T20:08:12.645597Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-13T20:07:29.544613Z digest=sha256:a3c5a47d5927a9ae52774842e758f864577033caf373cbf675c7e6f165c91193

Observation eecc8385-3b1a-4813-964e-1f18c989a4de · inbound

There Will Be a Scientific Theory of Deep Learning cites this paper.

There Will Be a Scientific Theory of Deep Learning Tensor Programs V: Tuning Large Neural Networks via Zero-Shot Hyperparameter Transfer

Reference 93

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T15:21:08.817339Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-05-09T20:11:17.616190Z digest=sha256:774daf028b0064ae6c08afb8931e1989b4b34b2092d718c38a4a15b371b5b1f3

Observation 54776053-ae13-4975-8faf-474c8e3bcb47 · inbound

Feature Starvation as Geometric Instability in Sparse Autoencoders cites this paper.

Feature Starvation as Geometric Instability in Sparse Autoencoders Tensor Programs V: Tuning Large Neural Networks via Zero-Shot Hyperparameter Transfer

Reference 39

Resolution
verified exact
arxiv_id, observed 2026-05-08T20:09:09.887485Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-08T16:31:38.468164Z digest=sha256:2cc4f262c2a501519ffeb389d103598c28c60e2f1189203608e9d77cb29f59e3

Observation 74854a07-b373-4121-8df6-0403657081e1 · inbound

OrScale: Orthogonalised Optimization with Layer-Wise Trust-Ratio Scaling cites this paper.

OrScale: Orthogonalised Optimization with Layer-Wise Trust-Ratio Scaling Tensor Programs V: Tuning Large Neural Networks via Zero-Shot Hyperparameter Transfer

Reference 20

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T03:05:55.500915Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-11T02:47:08.766380Z digest=sha256:bc098bd13712998c111e137687c7e5f980ec97faffdd13b9fbb6bba936b9cbdf

Observation b34976b3-b3b6-439e-bef8-05741a924671 · inbound

Spectral Dynamics in Deep Networks: Feature Learning, Outlier Escape, and Learning Rate Transfer cites this paper.

Spectral Dynamics in Deep Networks: Feature Learning, Outlier Escape, and Learning Rate Transfer Tensor Programs V: Tuning Large Neural Networks via Zero-Shot Hyperparameter Transfer

Reference 21

Resolution
verified exact
arxiv_id, observed 2026-05-11T03:05:53.394859Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-11T03:02:52.833353Z digest=sha256:8e962b78412e9ddd20b02d62d3e5483b2c17a6003ae4bacf9a782fd8d1074f88

Observation 72b3403e-7121-4245-9d98-d3e8d63d4b50 · inbound

Spectral Dynamics in Deep Networks: Feature Learning, Outlier Escape, and Learning Rate Transfer cites this paper.

Spectral Dynamics in Deep Networks: Feature Learning, Outlier Escape, and Learning Rate Transfer Tensor Programs V: Tuning Large Neural Networks via Zero-Shot Hyperparameter Transfer

Reference 21

Resolution
verified exact
arxiv_id, observed 2026-05-22T10:26:24.110548Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-22T10:25:54.649302Z digest=sha256:8d3ac5fb2355c187ee1eb169d3b38a9f3fd6f4d88362a8861014d9f211922a3c

Observation 9f20b3bb-8813-4816-bbe7-768aa250bd67 · inbound

Sparse Layers are Critical to Scaling Looped Language Models cites this paper.

Sparse Layers are Critical to Scaling Looped Language Models Tensor Programs V: Tuning Large Neural Networks via Zero-Shot Hyperparameter Transfer

Reference 18

Resolution
metadata mismatch
arxiv_id, observed 2026-05-12T06:21:28.372242Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-12T04:19:12.696712Z digest=sha256:c044a5c4aec3631028e475c3cd069ff1adfa0cb9a7969a6b0d1cbf4a862587a9

Observation 8e1563b3-a639-4f80-b3cb-96f1a24bc552 · inbound

Sparse Layers are Critical to Scaling Looped Language Models cites this paper.

Sparse Layers are Critical to Scaling Looped Language Models Tensor Programs V: Tuning Large Neural Networks via Zero-Shot Hyperparameter Transfer

Reference 18

Resolution
metadata mismatch
arxiv_id, observed 2026-07-01T07:35:29.134236Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-07-01T07:26:37.559459Z digest=sha256:34e6632b41e6f6c6f74a6fd86ec2c8b1ffca90c7ec8f40c97021108c1b6fffd8

Observation 223ee6e0-6e30-416c-bd99-f75729fde684 · inbound

Intrinsic Muon: Spectral Optimization on Riemannian Matrix Manifolds cites this paper.

Intrinsic Muon: Spectral Optimization on Riemannian Matrix Manifolds Tensor Programs V: Tuning Large Neural Networks via Zero-Shot Hyperparameter Transfer

Reference 67

Resolution
metadata mismatch
arxiv_id, observed 2026-05-12T05:36:26.119308Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-12T05:07:39.558349Z digest=sha256:f9a356adfda2533e77e47bf143a02d5ea6510a6d81c428e1b04678fd9128fcb6

Observation bbd3b439-e241-4331-b0e6-ac119083b36a · inbound

Mix, Don't Tune: Bilingual Pre-Training Outperforms Hyperparameter Search in Data-Constrained Settings cites this paper.

Mix, Don't Tune: Bilingual Pre-Training Outperforms Hyperparameter Search in Data-Constrained Settings Tensor Programs V: Tuning Large Neural Networks via Zero-Shot Hyperparameter Transfer

Reference 27

Resolution
metadata mismatch
arxiv_id, observed 2026-05-14T20:19:27.406229Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-14T20:17:26.661595Z digest=sha256:47231d4b13c7a97890d28e0f7f956bc30894dddb1699aee679985f499fc94b0e

Observation 137d2869-5d3a-495f-be56-6ce14f851813 · inbound

How to Scale Mixture-of-Experts: From muP to the Maximally Scale-Stable Parameterization cites this paper.

How to Scale Mixture-of-Experts: From muP to the Maximally Scale-Stable Parameterization Tensor Programs V: Tuning Large Neural Networks via Zero-Shot Hyperparameter Transfer

Reference 116

Resolution
metadata mismatch
arxiv_id, observed 2026-05-15T04:49:44.699169Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-05-15T04:45:20.091598Z digest=sha256:ec00937bc959d6ed600132510793a083e6d14bcaf4ecade0fa0dfff9aa6f1339

Observation ea5e42c0-568f-4e84-8ab1-52c434573633 · inbound

GQA-{\mu}P: The maximal parameterization update for grouped query attention cites this paper.

GQA-{\mu}P: The maximal parameterization update for grouped query attention Tensor Programs V: Tuning Large Neural Networks via Zero-Shot Hyperparameter Transfer

Reference 19

Resolution
verified exact
arxiv_id, observed 2026-05-19T16:37:39.825511Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-19T16:35:40.231293Z digest=sha256:85577d7b0a30693d3cd8b1e02ac2d4abe7cb8ab505be35a9a0d37e9bbaee026c

Observation 179ecece-900d-401e-86c9-a166afde2f41 · inbound

Simply Stabilizing the Loop via Fully Looped Transformer cites this paper.

Simply Stabilizing the Loop via Fully Looped Transformer Tensor Programs V: Tuning Large Neural Networks via Zero-Shot Hyperparameter Transfer

Reference 3

Resolution
verified exact
arxiv_id, observed 2026-07-01T14:05:47.263372Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-06-30T22:15:47.909334Z digest=sha256:2b178bfcd0bedba0690adf095dcffcbbdf7c037cf7fb12f76804fa7ff98c1b1b

Observation 7a499c47-f64b-49f4-a754-0378549c499c · inbound

Block-Based Double Decoders cites this paper.

Block-Based Double Decoders Tensor Programs V: Tuning Large Neural Networks via Zero-Shot Hyperparameter Transfer

Reference 16

Resolution
metadata mismatch
arxiv_id, observed 2026-05-20T22:09:07.439320Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-20T22:05:25.465669Z digest=sha256:83dfb084bb75afaf5bf4a1a7c1eb001568fcc57a54fb32d50c876f496f6fabe0

Observation e40a3231-baa4-4812-8478-49d6dfdc0cdb · inbound

Block-Based Double Decoders cites this paper.

Block-Based Double Decoders Tensor Programs V: Tuning Large Neural Networks via Zero-Shot Hyperparameter Transfer

Reference 16

Resolution
metadata mismatch
arxiv_id, observed 2026-07-01T14:15:47.113144Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-06-30T22:10:23.224012Z digest=sha256:b8185ca9b0f3de069f17bde233c54f63d4cc06e27631cc78a7ff4b3e6184b2de

Observation 8c8492ca-96bf-428a-b8bd-ec0762b8846e · inbound

Unified Neural Scaling Laws cites this paper.

Unified Neural Scaling Laws Tensor Programs V: Tuning Large Neural Networks via Zero-Shot Hyperparameter Transfer

Reference 29

Resolution
metadata mismatch
arxiv_id, observed 2026-06-29T23:44:03.344613Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-06-29T22:56:43.393302Z digest=sha256:403140b53aad7274bb7aed4aa89e0c8cd402aeb6cdd17e767fdd944b60d0ee8a

Observation d0fdc726-c70c-4df4-9a57-e1a44dd9b2bc · inbound

MuCon: Clipped Muon Updates for LLM Training cites this paper.

MuCon: Clipped Muon Updates for LLM Training Tensor Programs V: Tuning Large Neural Networks via Zero-Shot Hyperparameter Transfer

Reference 4

Resolution
verified exact
arxiv_id, observed 2026-06-29T20:03:56.395584Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-06-29T19:56:47.120210Z digest=sha256:29e38dd21745bd6828ffca0220eceb3a70cfbcd5810c20d5609197ddb4f1c39a

Observation f3b1de82-4626-42fd-9ac3-884dae64da54 · inbound

Spectral Reach: Understanding Neural Scaling as Progress into the Spectral Tail cites this paper.

Spectral Reach: Understanding Neural Scaling as Progress into the Spectral Tail Tensor Programs V: Tuning Large Neural Networks via Zero-Shot Hyperparameter Transfer

Reference 2025

Resolution
unresolved
no resolver link, observed 2026-08-04T04:58:48.901353Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T04:58:48.901353Z digest=sha256:d27d0527755971bb7e5b4d3bfec840a8f99be2c71e01f8ec3bc73e208e04557a

Observation de0c3d54-50eb-47db-8ec4-798875eb26da · inbound

On the Scaling of PEFT: Towards Million Personal Models of Trillion Parameters cites this paper.

On the Scaling of PEFT: Towards Million Personal Models of Trillion Parameters Tensor Programs V: Tuning Large Neural Networks via Zero-Shot Hyperparameter Transfer

Reference 36

Resolution
metadata mismatch
arxiv_id, observed 2026-07-01T22:16:16.642255Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-06-28T15:32:19.456119Z digest=sha256:e438033209aa9a4a13d13cee030e33829a82c53befd921a5d11f4ebb40a70504

Observation 7ae09695-3ace-4ddd-970f-0f69a7e1452b · inbound

Unlocking Feature Learning in Gated Delta Networks at Scale cites this paper.

Unlocking Feature Learning in Gated Delta Networks at Scale Tensor Programs V: Tuning Large Neural Networks via Zero-Shot Hyperparameter Transfer

Reference 14

Resolution
verified exact
arxiv_id, observed 2026-07-02T02:06:26.424178Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-06-28T11:18:55.669013Z digest=sha256:67ec4207f1af6ebdd17f3028e2f2614efd7b232a6ae7d14d830d5b0c7fc4e4e7

Observation aee8d983-2224-4d30-94c8-a2bf7f7e5443 · inbound

Spectral Scaling Laws of Muon cites this paper.

Spectral Scaling Laws of Muon Tensor Programs V: Tuning Large Neural Networks via Zero-Shot Hyperparameter Transfer

Reference 14

Resolution
metadata mismatch
arxiv_id, observed 2026-07-02T02:06:26.383110Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-06-28T11:19:05.939340Z digest=sha256:74cf54e6d3cfd88f7847ea0350ae6b097347cab25ec44fb46d7cecc95cb438aa

Observation de8c3ed4-101f-4eba-bafc-97276805c1e0 · inbound

Double Preconditioning (DoPr): Optimization for Test-Time Performance, not Validation Loss cites this paper.

Double Preconditioning (DoPr): Optimization for Test-Time Performance, not Validation Loss Tensor Programs V: Tuning Large Neural Networks via Zero-Shot Hyperparameter Transfer

Reference 102

Resolution
metadata mismatch
arxiv_id, observed 2026-07-02T11:56:56.103926Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-06-28T02:35:39.845487Z digest=sha256:01293426fbb53dafc2dc515ae5f2c765cbc0dd98924aa05748c4e8697db30661

Observation 344187bf-160c-4d78-a7e9-d048c05d745f · inbound

On the Residual Scaling of Looped Transformers: Stability and Transferability cites this paper.

On the Residual Scaling of Looped Transformers: Stability and Transferability Tensor Programs V: Tuning Large Neural Networks via Zero-Shot Hyperparameter Transfer

Reference 31

Resolution
malformed identifier
arxiv_id, observed 2026-07-03T21:08:58.378584Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-06-27T00:53:51.358774Z digest=sha256:b6b9a82ce70bbc903c00e276969162e475f8c92428c89bf24c3192c47291a3f4

Observation 14a3a96b-c447-4fd7-b4ff-5620e21c9628 · inbound

LLM Evolution as an Industry-Scale Ecosystem: A Lifecycle Perspective on Continual Learning cites this paper.

LLM Evolution as an Industry-Scale Ecosystem: A Lifecycle Perspective on Continual Learning Tensor Programs V: Tuning Large Neural Networks via Zero-Shot Hyperparameter Transfer

Reference 125

Resolution
verified exact
arxiv_id, observed 2026-07-03T16:48:39.965275Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-06-27T05:02:18.347642Z digest=sha256:ee0fe84683613dea838cc8199ae8cd1ac521f3f8315299e285cad301c97f2b14

Observation ad024d55-dee2-4067-9e3c-0b8a1ac4c6bd · inbound

Improving Neural Network Training by Decoupling the Magnitude and Direction of Weight Vectors cites this paper.

Improving Neural Network Training by Decoupling the Magnitude and Direction of Weight Vectors Tensor Programs V: Tuning Large Neural Networks via Zero-Shot Hyperparameter Transfer

Reference 145

Resolution
verified exact
arxiv_id, observed 2026-07-04T20:30:07.815811Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-06-25T20:05:09.179627Z digest=sha256:be40f2f326018b26fa2d8c9f318b2200c42f3441ee46b33f0b89c89dde330242

Observation d0ce3013-56c9-42f3-8993-87acf7fe0c60 · inbound

Improving Neural Network Training by Decoupling the Magnitude and Direction of Weight Vectors cites this paper.

Improving Neural Network Training by Decoupling the Magnitude and Direction of Weight Vectors Tensor Programs V: Tuning Large Neural Networks via Zero-Shot Hyperparameter Transfer

Reference 88

Resolution
unresolved
no resolver link, observed 2026-08-02T10:14:12.944746Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-02T10:14:12.944746Z digest=sha256:6c33534338a70f1c8461450af28889a2faf5bb289b0edf5da6a6d41af2fbf140

Observation 323bfc35-c383-4d2e-84ee-2b2faada5d28 · inbound

Geometric Dyson Brownian Motions and the Free Log-Normal Limit for a Non-Square Gaussian Matrix Product cites this paper.

Geometric Dyson Brownian Motions and the Free Log-Normal Limit for a Non-Square Gaussian Matrix Product Tensor Programs V: Tuning Large Neural Networks via Zero-Shot Hyperparameter Transfer

Reference 47

Resolution
verified exact
arxiv_id, observed 2026-07-01T12:45:44.992051Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-07-01T01:42:14.145227Z digest=sha256:b77885dfe45e47d83013a7db54b7d300c8b6c414c9df668bf13e94c35198bd4a

Observation cd3b708c-ec2c-4b61-bcfe-e8f332637497 · inbound

Geometric Dyson Brownian Motions and the Free Log-Normal Limit for a Non-Square Gaussian Matrix Product cites this paper.

Geometric Dyson Brownian Motions and the Free Log-Normal Limit for a Non-Square Gaussian Matrix Product Tensor Programs V: Tuning Large Neural Networks via Zero-Shot Hyperparameter Transfer

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-02T09:36:52.548930Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T09:36:52.548930Z digest=sha256:9a0a6188e3142ae65526e9db281172810ba66e1a07522c0c7caaabc220bafd82

Observation e15db8af-203a-4c7f-9dfb-862f7c416441 · inbound

The Role of Rigor in Artificial Intelligence cites this paper.

The Role of Rigor in Artificial Intelligence Tensor Programs V: Tuning Large Neural Networks via Zero-Shot Hyperparameter Transfer

Reference 49

Resolution
unresolved
no resolver link, observed 2026-07-12T16:22:39.179665Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T16:22:39.179665Z digest=sha256:31dcac26069681fe76c482e88ba2fe0b19d8586277028a6d192ae17351d9f99f

Observation 6143a722-f2b9-467f-bd1d-61bb7bc87469 · inbound

Index SLM Technical Report cites this paper.

Index SLM Technical Report Tensor Programs V: Tuning Large Neural Networks via Zero-Shot Hyperparameter Transfer

Reference 24

Resolution
unresolved
no resolver link, observed 2026-07-14T14:51:11.146058Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T14:51:11.146058Z digest=sha256:1572089db3fd79f24d86c8d7372b7da3ed4c405210cc07331e387a0760569e78

Observation 38cbf6cf-6290-4515-aa64-80b98712a8be · inbound

Scale Weight Decay and Train Better cites this paper.

Scale Weight Decay and Train Better Tensor Programs V: Tuning Large Neural Networks via Zero-Shot Hyperparameter Transfer

Reference 41

Resolution
unresolved
no resolver link, observed 2026-07-30T12:53:41.041793Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-30T12:53:41.041793Z digest=sha256:7d04ebcea982c64dec43ac77417036e3629b10cc117ccbec699855a88849639f

Observation 4feb1af7-058a-4472-a586-d05d034f1dd2 · inbound

Bridging Compute- and Data-Optimal Pretraining cites this paper.

Bridging Compute- and Data-Optimal Pretraining Tensor Programs V: Tuning Large Neural Networks via Zero-Shot Hyperparameter Transfer

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-01T03:02:01.193632Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T03:02:01.193632Z digest=sha256:f440dcf348fbcb4679c07b77ac12afb19ef7bb32c9d777ba706006dc878f95ed

Observation f9cc9133-94d9-4698-a52f-d1cec4bca3cb · inbound

Chimera: Designing and Chinchilla-Scaling Hybrid Visual Diffusion Transformers cites this paper.

Chimera: Designing and Chinchilla-Scaling Hybrid Visual Diffusion Transformers Tensor Programs V: Tuning Large Neural Networks via Zero-Shot Hyperparameter Transfer

Reference 93

Resolution
unresolved
no resolver link, observed 2026-07-31T02:16:08.253839Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-31T02:16:08.253839Z digest=sha256:ac9e0f76555f6cf56807d71b0e9c14888c33b213c0fe933eb93fb1b3a2a5f5cb

Observation 3d27a86f-b256-4327-8e34-4a6e96894d2e · inbound

Sign compression for Muon: SignMuon, MuonSign, and the Limits of Error Feedback cites this paper.

Sign compression for Muon: SignMuon, MuonSign, and the Limits of Error Feedback Tensor Programs V: Tuning Large Neural Networks via Zero-Shot Hyperparameter Transfer

Reference 2021

Resolution
unresolved
no resolver link, observed 2026-08-03T02:09:02.591486Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T02:09:02.591486Z digest=sha256:6200de88c924d371933d2ea4a525100c9964247b5484eef4ccc80ab2e0d5c135